Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books / PDF Originals
Riley, Hobson. Mathematical methods for physics and engineering (2ed., 2002)(1253s)_MPt_-1
PDF · 1253 pages · 6.6 MB
Open PDF file
Published textbook by Riley and Hobson, not Phil's own work, kept in a folder of math methods book downloads. The contents list covers algebra, calculus, complex numbers, series, partial differentiation, multiple integrals, vectors and matrices, normal modes, vector calculus, Fourier series and transforms, ordinary and partial differential equations, and Sturm-Liouville eigenfunction methods. Later chapters beyond the first 19 are not shown in the extracted text.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Contents
Preface to the second edition xix
Preface to the first edition xxi
1 Preliminary algebra 1
1.1 Simple functions and equations 1
Polynomial equations; factorisation; properties of roots
1.2 Trigonometric identities 10
Single angle; compound-angles; double- and half-angle identities
1.3 Coordinate geometry 151.4 Partial fractions 18
Complications and special cases; complex roots; repeated roots
1.5 Binomial expansion 251.6 Properties of binomial coefficients 271.7 Some particular methods of proof 30
Methods of proof; by induction; by contradiction; necessary and sufficient
conditions
1.8 Exercises 36
1.9 Hints and answers 39
2 Preliminary calculus 42
2.1 Differentiation 42
Differentiation from first principles; products; the chain rule; quotients;
implicit differentiation; logarithmic differentiation; Leibniz’ theorem; specialpoints of a function; theorems of differentiation
v
CONTENTS
2.2 Integration 60
Integration from first principles; the inverse of differentiation; integration
by inspection; sinusoidal functions; logarithmic integration; integrationusing partial fractions; substitution method; integration by parts; reduction
formulae; infinite and improper integrals; plane polar coordinates; integral
inequalities; applications of integration
2.3 Exercises 77
2.4 Hints and answers 82
3 Complex numbers and hyperbolic functions 86
3.1 The need for complex numbers 86
3.2 Manipulation of complex numbers 88
Addition and subtraction; modulus and argument; multiplication; complex
conjugate; division
3.3 Polar representation of complex numbers 95
Multiplication and division in polar form
3.4 de Moivre’s theorem 98
trigonometric identities; finding the nth roots of unity; solving polynomial
equations
3.5 Complex logarithms and complex powers 102
3.6 Applications to differentiation and integration 1043.7 Hyperbolic functions 105
Definitions; hyperbolic–trigonometric analogies; identities of hyperbolic
functions; solving hyperbolic equations ; inverses of hyperbolic functions;
calculus of hyperbolic functions
3.8 Exercises 112
3.9 Hints and answers 116
4 Series and limits 118
4.1 Series 1184.2 Summation of series 119
Arithmetic series; geometric series; arithmetico-geometric series; the
difference method; series involving natural numbers; transformation of series
4.3 Convergence of infinite series 127
Absolute and conditional convergence; convergence of a series containing
only real positive terms; alternating series test
4.4 Operations with series 134
4.5 Power series 134
Convergence of power series; operations with power series
4.6 Taylor series 139
Taylor’s theorem; approximation errors in Taylor series; standard Maclaurin
series
vi
CONTENTS
4.7 Evaluation of limits 144
4.8 Exercises 147
4.9 Hints and answers 152
5 Partial differentiation 154
5.1 Definition of the partial derivative 154
5.2 The total differential and total derivative 156
5.3 Exact and inexact differentials 158
5.4 Useful theorems of partial differentiation 160
5.5 The chain rule 160
5.6 Change of variables 161
5.7 Taylor’s theorem for many-variable functions 163
5.8 Stationary values of many-variable functions 165
5.9 Stationary values under constraints 170
5.10 Envelopes 1765.11 Thermodynamic relations 179
5.12 Differentiation of integrals 181
5.13 Exercises 182
5.14 Hints and answers 188
6 Multiple integrals 190
6.1 Double integrals 190
6.2 Triple integrals 193
6.3 Applications of multiple integrals 194
Areas and volumes; masses, centres of mass and centroids; Pappus’
theorems; moments of inertia; mean values of functions
6.4 Change of variables in multiple integrals 202
Change of variables in double integrals; evaluation of the integral I=R∞
−∞e−x2dx; change of variables in triple integrals; general properties of
Jacobians
6.5 Exercises 210
6.6 Hints and answers 214
7 Vector algebra 216
7.1 Scalars and vectors 216
7.2 Addition and subtraction of vectors 217
7.3 Multiplication by a scalar 218
7.4 Basis vectors and components 221
7.5 Magnitude of a vector 222
7.6 Multiplication of vectors 223
Scalar product; vector product; scalar triple product; vector triple product
vii
CONTENTS
7.7 Equations of lines, planes and spheres 230
Equation of a line; equation of a plane
7.8 Using vectors to find distances 233
Point to line; point to plane; line to line; line to plane
7.9 Reciprocal vectors 237
7.10 Exercises 2387.11 Hints and answers 244
8 Matrices and vector spaces 246
8.1 Vector spaces 247
Basis vectors; the inner product; some useful inequalities
8.2 Linear operators 252
Properties of linear operators
8.3 Matrices 254
Matrix addition and multiplication by a scalar; multiplication of matrices
8.4 Basic matrix algebra 2558.5 Functions of matrices 2608.6 The transpose of a matrix 2608.7 The complex and Hermitian conjugates of a matrix 2618.8 The trace of a matrix 2638.9 The determinant of a matrix 264
Properties of determinants
8.10 The inverse of a matrix 2688.11 The rank of a matrix 2728.12 Special types of square matrix 273
Diagonal; symmetric and antisymmetri c; orthogonal; Hermitian; unitary;
normal
8.13 Eigenvectors and eigenvalues 277
Of a normal matrix; of Hermitian and anti-Hermitian matrices; of a unitarymatrix; of a general square matrix
8.14 Determination of eigenvalues and eigenvectors 285
Degenerate eigenvalues
8.15 Change of basis and similarity transformations 288
8.16 Diagonalisation of matrices 2908.17 Quadratic and Hermitian forms 293
The stationary properties of the eigenvectors; quadratic surfaces
8.18 Simultaneous linear equations 297
Nsimultaneous linear equations in Nunknowns
8.19 Exercises 312
8.20 Hints and answers 319
viii
CONTENTS
9 Normal modes 322
9.1 Typical oscillatory systems 3239.2 Symmetry and normal modes 3289.3 Rayleigh–Ritz method 333
9.4 Exercises 335
9.5 Hints and answers 338
10 Vector calculus 340
10.1 Differentiation of vectors 340
Composite vector expressions; differential of a vector
10.2 Integration of vectors 345
10.3 Space curves 346
10.4 Vector functions of several arguments 35010.5 Surfaces 35110.6 Scalar and vector fields 35310.7 Vector operators 353
Gradient of a scalar field; divergence of a vector field; curlof a vector field
10.8 Vector operator formulae 360
Vector operators acting on sums and products; combinations of grad, div
andcurl
10.9 Cylindrical and spherical polar coordinates 363
Cylindrical polar coordinates; spherical polar coordinates
10.10 General curvilinear coordinates 370
10.11 Exercises 37510.12 Hints and answers 381
11 Line, surface and volume integrals 383
11.1 Line integrals 383
Evaluating line integrals; physical examples of line integrals; line integrals
with respect to a scalar
11.2 Connectivity of regions 389
11.3 Green’s theorem in a plane 390
11.4 Conservative fields and potentials 393
11.5 Surface integrals 395
Evaluating surface integrals; vector areas of surfaces; physical examples of
surface integrals
11.6 Volume integrals 402
Volumes of three-dimensional regions
11.7 Integral forms for grad, div and curl 404
11.8 Divergence theorem and related theorems 407
Green’s theorems; other related integral theorems; physical applications of
the divergence theorem
ix
CONTENTS
11.9 Stokes’ theorem and related theorems 412
Related integral theorems; physical applications of Stokes’ theorem
11.10 Exercises 415
11.11 Hints and answers 420
12 Fourier series 421
12.1 The Dirichlet conditions 42112.2 The Fourier coefficients 42312.3 Symmetry considerations 42512.4 Discontinuous functions 42612.5 Non-periodic functions 428
12.6 Integration and differentiation 430
12.7 Complex Fourier series 43012.8 Parseval’s theorem 43212.9 Exercises 43312.10 Hints and answers 437
13 Integral transforms 439
13.1 Fourier transforms 439
The uncertainty principle; Fraunhofer diffraction; the Dirac δ-function;
relation of the δ-function to Fourier transforms; properties of Fourier
transforms; odd and even functions; c onvolution and de convolution;
correlation functions and energy spectra; Parseval’s theorem; Fouriertransforms in higher dimensions
13.2 Laplace transforms 459
Laplace transforms of derivatives and integrals; other properties of Laplacetransforms
13.3 Concluding remarks 465
13.4 Exercises 46613.5 Hints and answers 472
14 First-order ordinary differential equations 474
14.1 General form of solution 475
14.2 First-degree first-order equations 476
Separable-variable equations; exact equations; inexact equations: integrat-
ing factors; linear equations; homogene ous equations; isobaric equations;
Bernoulli’s equation; mi scellaneous equations
14.3 Higher-degree first-order equations 486
Equations soluble for p;f o r x;f o r y; Clairaut’s equation
14.4 Exercises 490
14.5 Hints and answers 494
x
CONTENTS
15 Higher-order ordinary differential equations 496
15.1 Linear equations with constant coefficients 498
Finding the complementary function yc(x); finding the particular integral
yp(x); constructing the general solution yc(x)+yp(x); linear recurrence
relations; Laplace transform method
15.2 Linear equations with variable coefficients 509
The Legendre and Euler linear equations; exact equations; partiallyknown complementary function; variation of parameters; Green’s functions;canonical form for second-order equations
15.3 General ordinary differential equations 524
Dependent variable absent; independent variable absent; non-linear exactequations; isobaric or homogeneous equations; equations homogeneous in x
oryalone; equations having y=Ae
xas a solution
15.4 Exercises 529
15.5 Hints and answers 535
16 Series solutions of ordinary differential equations 537
16.1 Second-order linear ordinary differential equations 537
Ordinary and singular points
16.2 Series solutions about an ordinary point 54116.3 Series solutions about a regular singular point 544
Distinct roots not differing by an integer; repeated root of the indicial
equation; distinct roots differing by an integer
16.4 Obtaining a second solution 549
The Wronskian method; the derivative method; series form of the secondsolution
16.5 Polynomial solutions 554
16.6 Legendre’s equation 555
General solution for integer /lscript; properties of Legendre polynomials
16.7 Bessel’s equation 564
General solution for non-integer ν; general solution for integer ν; properties
of Bessel functions
16.8 General remarks 575
16.9 Exercises 575
16.10 Hints and answers 579
17 Eigenfunction methods for differential equations 581
17.1 Sets of functions 583
Some useful inequalities
17.2 Adjoint and Hermitian operators 587
xi
CONTENTS
17.3 The properties of Hermitian operators 588
Reality of the eigenvalues; orthogonalit y of the eigenfunctions; construction
of real eigenfunctions
17.4 Sturm–Liouville equations 591
Valid boundary conditions; putting an equation into Sturm–Liouville form
17.5 Examples of Sturm–Liouville equations 593
Legendre’s equation; the associated Legendre equation; Bessel’s equation;
the simple harmonic equation; Hermite’s equation; Laguerre’s equation;Chebyshev’s equation
17.6 Superposition of eigenfunctions: Green’s functions 597
17.7 A useful generalisation 601
17.8 Exercises 602
17.9 Hints and answers 606
18 Partial differential equations: general and particular solutions 608
18.1 Important partial differential equations 609
The wave equation; the diffusion equation; Laplace’s equation; Poisson’s
equation; Schr ¨odinger’s equation
18.2 General form of solution 613
18.3 General and particular solutions 614
First-order equations; inhomogeneous e quations and problems; second-order
equations
18.4 The wave equation 626
18.5 The diffusion equation 62818.6 Characteristics and the existence of solutions 632
First-order equations; second-order equations
18.7 Uniqueness of solutions 63818.8 Exercises 64018.9 Hints and answers 644
19 Partial differential equations: separation of variables
and other methods 646
19.1 Separation of variables: the general method 646
19.2 Superposition of separated solutions 65019.3 Separation of variables in polar coordinates 658
Laplace’s equation in polar coordinates; spherical harmonics; other
equations in polar coordinates; solution by expansion; separation ofvariables in inhomogeneous equations
19.4 Integral transform methods 681
19.5 Inhomogeneous problems – Green’s functions 686
Similarities with Green’s function for ordinary differential equations; general
boundary-value problems; Dirichlet problems; Neumann problems
xii
CONTENTS
19.6 Exercises 702
19.7 Hints and answers 708
20 Complex variables 710
20.1 Functions of a complex variable 711
20.2 The Cauchy–Riemann relations 71320.3 Power series in a complex variable 71620.4 Some elementary functions 71820.5 Multivalued functions and branch cuts 72120.6 Singularities and zeroes of complex functions 72320.7 Complex potentials 725
20.8 Conformal transformations 730
20.9 Applications of conformal transformations 73520.10 Complex integrals 73820.11 Cauchy’s theorem 74220.12 Cauchy’s integral formula 74520.13 Taylor and Laurent series 74720.14 Residue theorem 752
20.15 Location of zeroes 754
20.16 Integrals of sinusoidal functions 75820.17 Some infinite integrals 75920.18 Integrals of multivalued functions 76220.19 Summation of series 76420.20 Inverse Laplace transform 76520.21 Exercises 768
20.22 Hints and answers 773
21 Tensors 776
21.1 Some notation 77721.2 Change of basis 77821.3 Cartesian tensors 77921.4 First- and zero-order Cartesian tensors 781
21.5 Second- and higher-order Cartesian tensors 784
21.6 The algebra of tensors 78721.7 The quotient law 78821.8 The tensors δ
ijand/epsilon1ijk 790
21.9 Isotropic tensors 79321.10 Improper rotations and pseudotensors 79521.11 Dual tensors 798
21.12 Physical applications of tensors 799
21.13 Integral theorems for tensors 80321.14 Non-Cartesian coordinates 804
xiii
CONTENTS
21.15 The metric tensor 806
21.16 General coordinate transformations and tensors 80921.17 Relative tensors 81221.18 Derivatives of basis vectors and Christoffel symbols 814
21.19 Covariant differentiation 817
21.20 Vector operators in tensor form 82021.21 Absolute derivatives along curves 82421.22 Geodesics 82521.23 Exercises 82621.24 Hints and answers 831
22 Calculus of variations 834
22.1 The Euler–Lagrange equation 83522.2 Special cases 836
Fdoes not contain yexplicitly; Fdoes not contain xexplicitly
22.3 Some extensions 840
Several dependent variables; several i ndependent variables; higher-order
derivatives; variable end-points
22.4 Constrained variation 844
22.5 Physical variational principles 846
Fermat’s principle in optics; Hamilton’s principle in mechanics
22.6 General eigenvalue problems 84922.7 Estimation of eigenvalues and eigenfunctions 85122.8 Adjustment of parameters 85422.9 Exercises 856
22.10 Hints and answers 860
23 Integral equations 862
23.1 Obtaining an integral equation from a differential equation 86223.2 Types of integral equation 86323.3 Operator notation and the existence of solutions 86423.4 Closed-form solutions 865
Separable kernels; integral transform methods; differentiation
23.5 Neumann series 87223.6 Fredholm theory 87423.7 Schmidt–Hilbert theory 87523.8 Exercises 87823.9 Hints and answers 882
24 Group theory 883
24.1 Groups 883
Definition of a group; further examples of groups
xiv
CONTENTS
24.2 Finite groups 891
24.3 Non-Abelian groups 89424.4 Permutation groups 89824.5 Mappings between groups 901
24.6 Subgroups 903
24.7 Subdividing a group 905
Equivalence relations and classes; congruence and cosets; conjugates and
classes
24.8 Exercises 912
24.9 Hints and answers 915
25 Representation theory 918
25.1 Dipole moments of molecules 91925.2 Choosing an appropriate formalism 92025.3 Equivalent representations 92625.4 Reducibility of a representation 92825.5 The orthogonality theorem for irreducible representations 93225.6 Characters 934
Orthogonality property of characters
25.7 Counting irreps using characters 937
Summation rules for irreps
25.8 Construction of a character table 94225.9 Group nomenclature 94425.10 Product representations 94525.11 Physical applications of group theory 947
Bonding in molecules; matrix elemen ts in quantum mechanics; degeneracy
of normal modes; breaking of degeneracies
25.12 Exercises 955
25.13 Hints and answers 959
26 Probability 961
26.1 Venn diagrams 961
26.2 Probability 966
Axioms and theorems; conditional probability; Bayes’ theorem
26.3 Permutations and combinations 975
26.4 Random variables and distributions 981
Discrete random variables; continuous random variables
26.5 Properties of distributions 985
Mean; mode and median; variance; higher moments; higher central moments
26.6 Functions of random variables 99226.7 Generating functions 999
Probability generating functions; moment generating functions
xv
CONTENTS
26.8 Important discrete distributions 1009
Binomial; hypergeometric; Poisson; Poisson approximation to the binomial
distribution; multiple Poisson distributions
26.9 Important continuous distributions 1021
Gaussian; Gaussian approximation to the binomial distribution; Gaussianapproximation to the Poisson distribution; multiple Gaussian; exponential;uniform
26.10 The central limit theorem 1036
26.11 Joint distributions 1038
Discrete bivariate; continuous bivariate; conditional; marginal
26.12 Properties of joint distributions 1041
Expectation values; variance; covariance and correlation
26.13 Generating functions for joint distributions 104726.14 Transformation of variables in joint distributions 104826.15 Important joint distributions 1049
Multinominal; multivariate Gaussian; transformation of variables in multi-
variate distributions
26.16 Exercises 1053
26.17 Hints and answers 1061
27 Statistics 1064
27.1 Experiments, samples and populations 106427.2 Sample statistics 1065
Averages; variance and standard deviation; moments; covariance and
correlation
27.3 Estimators and sampling distributions 1072
Consistency, bias and efficiency; Fisher’s inequality; standard errors;confidence limits
27.4 Some basic estimators 1086
Mean; variance; standard deviation; moments; covariance and correlation
27.5 Maximum-likelihood method 1097
ML estimator; transformation invariance and bias; efficiency; errors and
confidence limits; Bayesian interpretation; large Nbehaviour; extended
maximum-likelihood
27.6 The method of least squares 1113
Linear least squares; non-linear least squares
27.7 Hypothesis testing 1119
Simple and composite hypotheses; statistical tests; Neyman-Pearson;generalised likelihood-ratio; Student’s t;F i s h e r ’ s F; goodness-of-fit
27.8 Exercises 1140
27.9 Hints and answers 1145
xvi
CONTENTS
28 Numerical methods 1148
28.1 Algebraic and transcendental equations 1149
Rearrangement of the equation; linea r interpolation; binary chopping;
Newton–Raphson method
28.2 Convergence of iteration schemes 1156
28.3 Simultaneous linear equations 1158
Gaussian elimination; Gauss–Seidel iteration; tridiagonal matrices
28.4 Numerical integration 1164
Trapezium rule; Simpson’s rule; Gaussian integration; Monte-Carlo methods
28.5 Finite differences 117928.6 Differential equations 1180
Difference equations; Taylor series solutions; prediction and correction;
Runge–Kutta methods; isoclines
28.7 Higher-order equations 1188
28.8 Partial differential equations 119028.9 Exercises 119328.10 Hints and answers 1198
Appendix Gamma, beta and error functions 1201
A1.1 The gamma function 1201A1.2 The beta function 1203A1.3 The error function 1204
Index 1206
xvii
Preface to the second edition
Since the publication of the first edition of this book, we have, both through
teaching the material it covers and as a result of receiving helpful comments fromcolleagues, become aware of the desirability of changes in a number of areas.The most important of these is the fact that the mathematical preparation ofcurrent senior college and university entrants is now less than it used to be. To
match this, we have decided to include a preliminary chapter covering areas such
as polynomial equations, trigonometric identities, coordinate geometry, partialfractions, binomial expansions, necessary and sufficient conditions, and proof byinduction and contradiction.
Whilst the general level of what is included in this second edition has not
been raised, some areas have been expanded to take in topics we now feel were
not adequately covered in the first. In particular, increased attention has been
given to non-square sets of simultaneous linear equations and their associatedmatrices. We hope that this more extended treatment, together with the inclusionof singular value matrix decomposition will make the material of more practicaluse to engineering students. In the same spirit, an elementary treatment of linearrecurrence relations has been included. The topic of normal modes has now beengiven a small chapter of its own, though the links to matrices on the one hand,
and to representation theory on the other, have not been lost.
Elsewhere, the presentation of probability and statistics has been reorganised to
give the two aspects more nearly equal weights. The early part of the probabilitychapter has been rewritten in order to present a more coherent developmentbased on Boolean algebra, the fundamental axioms of probability theory andthe properties of intersections and unions. Whilst this is somewhat more formalthan previously, we think that it has not reduced the accessibility of these topics
and hope that it has increased it. The scope of the chapter has been somewhat
extended to include all physically important distributions and an introduction tocumulants.
xix
PREFACE TO THE SECOND EDITION
Statistics now occupies a substantial chapter of its own, one that includes
systematic discussions of estimators and their efficiency, sample distributions,andt-a n d F-tests for comparing means and variances. Other new topics are
applications of the chi-squared distribution, maximum-likelihood parameter es-
timation and least-squares fitting. In other chapters we have added material on
the following topics: curvature, envelopes, curve-sketching, more refined numer-ical methods for differential equations, and the elements of integration usingmonte-carlo techniques.
Over the last four years we have received somewhat mixed feedback about
the number of exercises to include at the ends of the various chapters. Afterconsideration, we decided to increase it substantially, partly to correspond to the
additional topics covered in the text, but mainly to give both students and their
teachers a wider choice. There are now nearly eight hundred such exercises, manywith several parts. An even more vexed question is that of whether or not toprovide hints and answers to all of the exercises, or just to ‘the odd-numbered’ones, as is the normal practice for textbooks in the United States, thus makingthe remainder more suitable for setting as homework. In the end, we decided thathints and outline solutions should be provided for all the exercises, in order to
facilitate independent study while leaving the details of the calculation as a task
for the student.
In conclusion we hope that this edition will be thought by its users to be
‘heading in the right direction’ and would like to place on record our thanks toall who have helped to bring about the changes and adjustments. Naturally, thosecolleagues who have noted errors or ambiguities in the first edition and broughtthem to our attention figure high on the list, as do the staff at The Cambridge
University Press. In particular, we are grateful to Dave Green for continued L
ATEX
advice, Susan Parkinson for copy-editing the 2nd edition with her usual keen eyefor detail and flair for crafting coherent prose, and Alison Woollatt for once againturning our basic L
ATEX into a beautifully typeset book. Our thanks go to all of
them, though of course we accept full responsibility for any remaining errors orambiguities, of which, as with any new publication, there are bound to be some.
On a more personal note, KFR again wishes to thank his wife Penny for her
unwavering support, not only in his academic and tutorial work, but also in their
joint efforts to convert time at the bridge table into ‘green points’ on their record.MPH is once more indebted to his wife, Becky, and his mother, Pat, for theirtireless support and encouragement above and beyond the call of duty. MPHdedicates his contribution to this book to the memory of his father, RonaldLeonard Hobson, whose gentle kindness, patient understanding and unbreakablespirit made all things seem possible.
Ken Riley, Michael Hobson
Cambridge, 2002
xx
Preface to the first edition
A knowledge of mathematical methods is important for an increasing number of
university and college courses, particularly in physics, engineering and chemistry,but also in more general science. Students embarking on such courses come fromdiverse mathematical backgrounds, and their core knowledge varies considerably.We have therefore decided to write a textbook that assumes knowledge only ofmaterial that can be expected to be familiar to all the current generation of
students starting physical science courses at university. In the United Kingdom
this corresponds to the standard of Mathematics A-level, whereas in the UnitedStates the material assumed is that which would normally be covered at juniorcollege.
Starting from this level, the first six chapters cover a collection of topics
with which the reader may already be familiar, but which are here extended
and applied to typical problems encountered by first-year university students.
They are aimed at providing a common base of general techniques used inthe development of the remaining chapters. Students who have had additionalpreparation, such as Further Mathematics at A-level, will find much of thismaterial straightforward.
Following these opening chapters, the remainder of the book is intended to
cover at least that mathematical material which an undergraduate in the physical
sciences might encounter up to the end of his or her course. The book is also
appropriate for those beginning graduate study with a mathematical content, andnaturally much of the material forms parts of courses for mathematics students.Furthermore, the text should provide a useful reference for research workers.
The general aim of the book is to present a topic in three stages. The first
stage is a qualitative introduction, wherever possible from a physical point of
view. The second is a more formal presentation, although we have deliberately
avoided strictly mathematical questions such as the existence of limits, uniformconvergence, the interchanging of integration and summation orders, etc. on the
xxi
PREFACE TO THE FIRST EDITION
grounds that ‘this is the real world; it must behave reasonably’. Finally a worked
example is presented, often drawn from familiar situations in physical scienceand engineering. These examples have generally been fully worked, since, inthe authors’ experience, partially worked examples are unpopular with students.
Only in a few cases, where trivial algebraic manipulation is involved, or where
repetition of the main text would result, has an example been left as an exercisefor the reader. Nevertheless, a number of exercises also appear at the end of eachchapter, and these should give the reader ample opportunity to test his or herunderstanding. Hints and answers to these exercises are also provided.
With regard to the presentation of the mathematics, it has to be accepted that
many equations (especially partial differential equations) can be written more
compactly by using subscripts, e.g. u
xyfor a second partial derivative, instead of
the more familiar ∂2u/∂x∂y , and that this certainly saves typographical space.
However, for many students, the labour of mentally unpacking such equationsis sufficiently great that it is not possible to think of an equation’s physicalinterpretation at the same time. Consequently, wherever possible we have decidedto write out such expressions in their more obvious but longer form.
During the writing of this book we have received much help and encouragement
from various colleagues at the Cavendish Laboratory, Clare College, Trinity Hall
and Peterhouse. In particular, we would like to thank Peter Scheuer, whosecomments and general enthusiasm proved invaluable in the early stages. Forreading sections of the manuscript, for pointing out misprints and for numeroususeful comments, we thank many of our students and colleagues at the Universityof Cambridge. We are especially grateful to Chris Doran, John Huber, GarthLeder, Tom K ¨orner and, not least, Mike Stobbs, who, sadly, died before the book
was completed. We also extend our thanks to the University of Cambridge and
the Cavendish teaching staff, whose examination questions and lecture hand-outshave collectively provided the basis for some of the examples included. Of course,any errors and ambiguities remaining are entirely the responsibility of the authors,and we would be most grateful to have them brought to our attention.
We are indebted to Dave Green for a great deal of advice concerning typesetting
in L
ATEX and to Andrew Lovatt for various other computing tips. Our thanks
also go to Anja Visser and Grac ¸a Rocha for enduring many hours of (sometimes
heated) debate. At Cambridge University Press, we are very grateful to our editorAdam Black for his help and patience and to Alison Woollatt for her experttypesetting of such a complicated text. We also thank our copy-editor SusanParkinson for many useful suggestions that have undoubtedly improved the styleof the book.
Finally, on a personal note, KFR wishes to thank his wife Penny, not only for
a long and happy marriage, but also for her support and understanding during
his recent illness – and when things have not gone too well at the bridge table!MPH is indebted both to Rebecca Morris and to his parents for their tireless
xxii
PREFACE TO THE FIRST EDITION
support and patience, and for their unending supplies of tea. SJB is grateful to
Anthony Gritten for numerous relaxing discussions about J.S.Bach, to SusannahTicciati for her patience and understanding, and to Kate Isaak for her calminglate-night e-mails from the USA.
Ken Riley, Michael Hobson and Stephen Bence
Cambridge, 1997
xxiii
1
Preliminary algebra
This opening chapter reviews the basic algebra of which a working knowledge is
presumed in the rest of the book. Many students will be familiar with much, ifnot all, of it, but recent changes in what is studied during secondary educationmean that it cannot be taken for granted that they will already have a masteryof all the topics presented here. The reader may assess which areas need furtherstudy or revision by attempting the exercises at the end of the chapter. The mainareas covered are polynomial equations and the related topic of partial fractions,
curve sketching, coordinate geometry, trigonometric identities and the notions of
proof by induction or contradiction.
1.1 Simple functions and equations
It is normal practice when starting the mathematical investigation of a physical
problem to assign an algebraic symbol to the quantity whose value is sought, eithernumerically or as an explicit algebraic expression. For the sake of definiteness, inthis chapter we will use xto denote this quantity most of the time. Subsequent
steps in the analysis involve applying a combination of known laws, consistency
conditions and (possibly) given constraints to derive one or more equationssatisfied by x. These equations may take many forms, ranging from a simple
polynomial equation to, say, a partial differential equation with several boundaryconditions. Some of the more complicated possibilities are treated in the laterchapters of this book, but for the present we will be concerned with techniquesfor the solution of relatively straightforward algebraic equations.
1.1.1 Polynomials and polynomial equations
Firstly we consider the simplest type of equation, a polynomial equation in which
apolynomial expression in x, denoted by f(x), is set equal to zero and thereby
1
PRELIMINARY ALGEBRA
forms an equation which is satisfied by particular values of x; these values are
called the rootsof the equation.
f(x)=anxn+an−1xn−1+···+a1x+a0=0. (1.1)
Here nis an integer >0, called the degree of both the polynomial and the
equation, and the known coefficients a0,a1,...,a nare real quantities with an/negationslash=0 .
Equations such as (1.1) arise frequently in physical problems, the coefficients ai
being determined by the physical properties of the system under study. What is
needed is to find some or all of the roots solutions of (1.1), i.e. the x-values, αk,,
that satisfy f(αk)=0 ;h e r e kis an index that, as we shall see later, can take up to
ndifferent values, i.e. k=1,2,...,n. The roots of the polynomial equations can
equally well be described as the zeroes of the polynomial. When they are real,
they correspond to the points at which a graph of f(x)c r o s s e st h e x-axis. Roots
that are complex (see chapter 3) do not have such a graphical interpretation.
For polynomial equations containing powers of xgreater tha x4general meth-
ods do not exist for obtaining explicit expressions for the roots αk.E v e nf o r
n=3a n d n= 4 the prescriptions for obtaining the roots are sufficiently compli-
cated that it is usually preferable to obtain exact or approximate values by othermethods. Only for n=1a n d n= 2 can closed-form solutions be given. These
results will be well known to the reader, but they are given here for the sake of
completeness. For n= 1, (1.1) reduces to the linear equation
a
1x+a0= 0; (1.2)
the solution (root) is α1=−a0/a1.F o r n= 2, (1.1) reduces to the quadratic
equation
a2x2+a1x+a0= 0; (1.3)
the two roots α1andα2are given by
α1,2=−a1±radicalBig
a2
1−4a2a0
2a2. (1.4)
When discussing specifically quadratic equations, as opposed to more general
polynomial equations, it is usual to write the equation in one of the two notations
ax2+bx+c=0,a x2+2bx+c=0, (1.5)
with respective explicit pairs of solutions
α1,2=−b±√
b2−4ac
2a,α 1,2=−b±√
b2−ac
a. (1.6)
Of course, these two notations are entirely equivalent and the only important
2
1.1 SIMPLE FUNCTIONS AND EQUATIONS
point is to associate each form of answer with the corresponding form of equation;
most people keep to one form, to avoid any possible confusion.
If the value of the quantity appearing under the square root sign is positive
then both roots are real; if it is negative then the roots form a complex conjugate
pair, i.e. they are of the form p±iqwith pandqreal (see chapter 3); if it has
zero value then the two roots are equal and special considerations usually arise.
Thus linear and quadratic equations can be dealt with in a cut-and-dried way.
We now turn to methods for obtaining partial information about the roots ofhigher-degree polynomial equations. In some circumstances the knowledge thatan equation has a root lying in a certain range, or that it has no real roots at all,is all that is actually required. For example, in the design of electronic circuits
it is necessary to know whether the current in a proposed circuit will break
into spontaneous oscillation. To test this, it is sufficient to establish whether acertain polynomial equation, whose coefficients are determined by the physicalparameters of the circuit, has a root with a positive real part (see chapter 3);complete determination of all the roots is not needed for this purpose. If thecomplete set of roots of a polynomial equation is required, it can usually beobtained to any desired accuracy by numerical methods such as those described
in chapter 28.
There is no explicit step-by-step approach to finding the roots of a general
polynomial equation such as (1.1). In most cases analytic methods yield onlyinformation about the roots, rather than their exact values. To explain the relevant
techniques we will consider a particular example, ‘thinking aloud’ on paper andexpanding on special points about methods and lines of reasoning. In moreroutine situations such comment would be absent and the whole process briefer
and more tightly focussed.
Example: the cubic case
Let us investigate the roots of the equation
g(x)=4 x
3+3x2−6x−1 = 0 (1.7)
or, in an alternative phrasing, investigate the zeroes of g(x). We note first of all
that this is a cubic equation. It can be seen that for xlarge and positive g(x)
will be large and positive and equally that for xlarge and negative g(x) will
be large and negative. Therefore, intuitively (or, more formally, by continuity)g(x) must cross the x-axis at least once and so g(x) = 0 must have at least one
real root. Furthermore, it can be shown that if f(x)i sa n nth-degree polynomial
then the graph of f(x) must cross the x-axis an even or odd number of times
asxvaries between −∞and +∞, according to whether nitself is even or odd.
Thus a polynomial of odd degree always has at least one real root, but one of
even degree may have no real root. A small complication, discussed later in thissection, occurs when repeated roots arise.
3
PRELIMINARY ALGEBRA
Having established that g(x) = 0, equation(1.7), has at least one real root, we
may ask how many real roots it could have. To answer this we need one of the
fundamental theorems of algebra, mentioned above:
Annth-degree polynomial equation has exactly nroots.
It should be noted that this does not imply that there are nrealroots (only that
there are not more than n); some of the roots may be of the form p+iq.
To make the above theorem plausible and to see what is meant by repeated
roots, let us suppose that the nth-degree polynomial equation f(x) = 0, (1.1), has
rroots α1,α2,...,α rconsidered distinct for the moment. That is, we suppose that
f(αk)=0f o r k=1,2,...,r,s ot h a t f(x) vanishes only when xis equal to one of
thervalues αk. But the same can be said for the function
F(x)=A(x−α1)(x−α2)···(x−αr), (1.8)
in which Ais a non-zero constant; F(x) can clearly be multiplied out to form a
polynomial expression.
We now call upon a second fudamental result in algebra: that if two polynomial
functions f(x)a n d F(x) have equal values for allvalues of x, then their coefficients
are equal on a term-by-term basis. In other words, we can equate the coefficients
of each and every power of xin the two expressions; in particular we can equate
the coefficients of the highest power of x. From this we have Axr≡anxnand
thus that r=nandA=an.A sris both equal to nand to the number of roots
off(x) = 0, we conclude that the nth-degree polynomial f(x)=0h a s nroots.
(Although this line of reasoning may make the theorem plausible, it does notconstitute a proof since we have not shown that it is permissible to write f(x)i n
the form of equation (1.8).)
We next note that the condition f(α
k)=0f o r k=1,2,...,r, could also be met
if (1.8) were replaced by
F(x)=A(x−α1)m1(x−α2)m2···(x−αr)mr, (1.9)
with A=an. In (1.9) the mkare integers ≥1 and are known as the multiplicities
of the roots, mkbeing the multiplicity of αk. Expanding the right-hand side (RHS)
leads to a polynomial of degree m1+m2+···+mr. This sum must be equal to n.
Thus, if any of the mkis greater than unity then the number of distinct roots, r,
is less than n; the total number of roots remains at n, but one or more of the αk
counts more than once. For example, the equation
F(x)=A(x−α1)2(x−α2)3(x−α3)(x−α4)=0
has exactly seven roots, α1being a double root and α2a triple root, whilst α3and
α4are unrepeated ( simple )r o o t s .
We can now say that our particular equation (1.7) has either one or three real
roots but in the latter case it may be that not all the roots are distinct. To decide
4
1.1 SIMPLE FUNCTIONS AND EQUATIONS
xxφ1(x) φ2(x)
β1 β1β2
β2
Figure 1.1 Two curves φ1(x)a n d φ2(x), both with zero derivatives at the
same values of x, but with different numbers of real solutions to φi(x)=0 .
how many real roots the equation has, we need to anticipate two ideas from the
next chapter. The first of these is the notion of the derivative of a function, andthe second is a result known as Rolle’s theorem.
Thederivative f
/prime(x) of a function f(x) measures the slope of the tangent to
the graph of f(x) at that value of x(see figure 2.1 in the next chapter). For
the moment, the reader with no prior knowledge of calculus is asked to acceptthat the derivative of ax
nisnaxn−1, so that the derivative g/prime(x)o ft h ec u r v e
g(x)=4 x3+3x2−6x−1i sg i v e nb y g/prime(x)=1 2 x2+6x−6. Similar expressions
for the derivatives of other polynomials are used later in this chapter.
Rolle’s theorem states that, if f(x) has equal values at two different values of
xthen at some point between these two x-values its derivative is equal to zero;
i.e. the tangent to its graph is parallel to the x-axis at that point (see figure 2.2).
Having briefly mentioned the derivative of a function and Rolle’s theorem, we
now use them to etablish whether g(x) has one or three real zeroes. If g(x)=0
does have three real roots αk,i . e . g(αk)=0f o r k=1,2,3, then it follows from
Rolle’s theorem that between any consecutive pair of them (say α1andα2)t h e r e
must be some real value of xat which g/prime(x) = 0. Similarly, there must be a further
zero of g/prime(x) lying between α2andα3. Thus a necessary condition for three real
roots of g(x)=0i st h a t g/prime(x) = 0 itself has two real roots.
However, this condition on the number of roots of g/prime(x) = 0, whilst necessary,
is not sufficient to guarantee three real roots of g(x) = 0. This can be seen by
inspecting the cubic curves in figure 1.1. For each of the two functions φ1(x)a n d
φ2(x), the derivative is equal to zero at both x=β1andx=β2. Clearly, though,
φ2(x) = 0 has three real roots whilst φ1(x) = 0 has only one. It is easy to see that
the crucial difference is that φ1(β1)a n d φ1(β2) have the same sign, whilst φ2(β1)
andφ2(β2) have opposite signs.
5
PRELIMINARY ALGEBRA
It will be apparent that for some cubic equations, φ(x)=0s a y , φ/prime(x)e q u a l s
zero at a value of xfor which φ(x) is also zero. Then the graph of φ(x)j u s t
touches the x-axis and there may appear to be only two roots. However, when
this happens the value of xso found is, in fact, a double real root of the cubic
(corresponding to one of the mkin (1.9) having the value 2) and must be counted
twice when determining the number of real roots.
Finally, then, we are in a position to decide the number of real roots of the
equation
g(x)=4 x3+3x2−6x−1=0 .
The equation g/prime(x)=0 ,w i t h g/prime(x)=1 2 x2+6x−6, is a quadratic equation with
explicit solutions †
β1,2=−3±√
9+7 2
12,
so that β1=−1a n d β2=1/2. The corresponding values of g(x)a r e g(β1)=4a n d
g(β2)=−11/4, which are of opposite sign. This indicates that 4 x3+3x2−6x−1=0
has three real roots, one lying in the range −1<x<1
2and the others one on
each side of that range.
The techniques we have developed above have been used to tackle a cubic
equation, but they can be applied to polynomial equations f(x)=0o fd e g r e e
greater than 3. However, much of the analysis centres around the equation
f/prime(x) = 0 and this, itself, being then a polynomial equation of degree 3 or more
either has no closed-form general solution or one that is complicated to evaluate.Thus the amount of information that can be obtained about the roots of f(x)=0
is correspondingly reduced.
A more general case
To illustrate what can (and cannot) be done in the more general case we now
investigate as far as possible the real roots of
f(x)=x
7+5x6+x4−x3+x2−2=0 .
The following points can be made.
(i) This is a seventh-degree polynomial equation; therefore the number of
r e a lr o o t si s1 ,3 ,5o r7 .
(ii)f(0) is negative whilst f(∞)=+∞, so there must be at least one positive
root.
†The two roots β1,β2are written as β1,2. By convention β1refers to the upper symbol in ±,β2to
the lower symbol.
6
1.1 SIMPLE FUNCTIONS AND EQUATIONS
(iii) The equation f/prime(x) = 0 can be written as x(7x5+3 0x4+4x2−3x+2 )=0
and thus x= 0 is a solution. The derivative of f/prime(x), denoted by f/prime/prime(x),
equals 42 x5+ 150 x4+1 2x2−6x+2 . T h a t f/prime(x) is zero whilst f/prime/prime(x)i s
positive at x= 0 indicates (subsection 2.1.8 ) that f(x) has a minimum
there. This, together with the facts that f(0) is negative and f(∞)=∞,
implies that the total number of real roots to the right of x= 0 must be
odd. Since the total number of real roots must be odd, the number to theleft must be even (0, 2, 4 or 6).
This is about all that can be deduced by simple analytic methods in this case,
although some further progress can be made in the ways indicated in exercise 1.3.
There are, in fact, more sophisticated tests that examine the relative signs of
successive terms in an equation such as (1.1), and in quantities derived from
them, to place limits on the numbers and positions of roots. But they are not
prerequisites for the remainder of this book and will not be pursued further here.
We conclude this section with a worked example which demonstrates that the
practical application of the ideas developed so far can be both short and decisive.IFor what values of k,i fa n y ,d o e s
f(x)=x3−3x2+6x+k=0
have three real roots?
Firstly study the equation f/prime(x)=0 ,i . e .3 x2−6x+ 6 = 0. This is a quadratic equation
but, using (1.6), because 62<4×3×6, it can have no real roots. Therefore, it follows
immediately that f(x) has no turning points, i.e. no maximum or minimum; consequently
f(x) = 0 cannot have more than one real root, whatever the value of k.
J
1.1.2 Factorising polynomials
In the previous subsection we saw how a polynomial with rgiven distinct zeroes
αkcould be constructed as the product of factors containing those zeroes,
f(x)=an(x−α1)m1(x−α2)m2···(x−αr)mr
=anxn+an−1xn−1+···+a1x+a0, (1.10)
with m1+m2+···+mr=n, the degree of the polynomial. It will cause no loss of
generality in what follows to suppose that all the zeroes are simple, i.e. all mk=1
andr=n, and this we will do.
Sometimes it is desirable to be able to reverse this process, in particular when
one exact zero has been found by some method and the remaining zeroes are tobe investigated. Suppose that we have located one zero, α; it is then possible to
write (1.10) as
f(x)=(x−α)f
1(x), (1.11)
7
PRELIMINARY ALGEBRA
where f1(x) is a polynomial of degree n−1. How can we find f1(x)? The procedure
is much more complicated to describe in a general form than to carry out foran equation with given numerical coefficients a
i. If such manipulations are too
complicated to be carried out mentally, they could be laid out along the lines of
an algebraic ‘long division’ sum. However, a more compact form of calculation
is as follows. Write f1(x)a s
f1(x)=bn−1xn−1+bn−2xn−2+bn−3xn−3+···+b1x+b0.
Substitution of this form into (1.11) and subsequent comparison of the coefficients
ofxpforp=n,n−1,..., 1, 0 with those in the second line of (1.10) generates
the series of equations
bn−1=an,
bn−2−αbn−1=an−1,
bn−3−αbn−2=an−2,
...
b0−αb1=a1,
−αb0=a0.
These can be solved successively for the bj, starting either from the top or from
the bottom of the series. In either case the final equation used serves as a check;if it is not satisfied, at least one mistake has been made in the computation –orαis not a zero of f(x) = 0. We now illustrate this procedure with a worked
example.IDetermine by inspection the simple roots of the equation
f(x)=3 x4−x3−10x2−2x+4=0
and hence, by factorisation, find the rest of its roots.
From the pattern of coefficients it can be seen that x=−1 is a solution to the equation.
We therefore write
f(x)=(x+1 ) ( b3x3+b2x2+b1x+b0),
where
b3=3,
b2+b3=−1,
b1+b2=−10,
b0+b1=−2,
b0=4.
These equations give b3=3,b2=−4,b1=−6,b0= 4 (check) and so
f(x)=(x+1 )f1(x)=(x+ 1)(3 x3−4x2−6x+4 ).
8
1.1 SIMPLE FUNCTIONS AND EQUATIONS
We now note that f1(x)=0i f xis set equal to 2. Thus x−2i saf a c t o ro f f1(x), which
therefore can be written as
f1(x)=(x−2)f2(x)=(x−2)(c2x2+c1x+c0)
with
c2=3,
c1−2c2=−4,
c0−2c1=−6,
−2c0=4.
These equations determine f2(x)a s3 x2+2x−2. Since f2(x) = 0 is a quadratic equation,
its solutions can be written explicitly as
x=−1±√1+6
3.
Thus the four roots of f(x)=0a r e−1,2,1
3(−1+√7) and1
3(−1−√7).
J
1.1.3 Properties of roots
From the fact that a polynomial equation can be written in any of the alternative
forms
f(x)=anxn+an−1xn−1+···+a1x+a0=0,
f(x)=an(x−α1)m1(x−α2)m2···(x−αr)mr=0,
f(x)=an(x−α1)(x−α2)···(x−αn)=0 ,
it follows that it must be possible to express the coefficients aiin terms of the
roots αk. To take the most obvious example, comparison of the constant terms
(formally the coefficient of x0) in the first and third expressions shows that
an(−α1)(−α2)···(−αn)=a0,
or, using the product notation,
nproductdisplay
k=1αk=(−1)na0
an. (1.12)
Only slightly less obvious is a result obtained by comparing the coefficients of
xn−1in the same two expressions of the polynomial:
nsummationdisplay
k=1αk=−an−1
an. (1.13)
Comparing the coefficients of other powers of xyields further results, though
they are of less general use than the two just given. One such, which the readermay wish to derive, is
nsummationdisplay
j=1nsummationdisplay
k>jαjαk=an−2
an. (1.14)
9
PRELIMINARY ALGEBRA
In the case of a quadratic equation these root properties are used sufficiently
often that they are worth stating explicitly, as follows. If the roots of the quadraticequation ax
2+bx+c=0a r e α1andα2then
α1+α2=−b
a,
α1α2=c
a.
If the alternative standard form for the quadratic is used, bis replaced by 2 bin
both the equation and the first of these results.IFind a cubic equation whose roots are −4,3and5.
From results (1.12) – (1.14) we can compute that, arbitrarily setting a3=1 ,
−a2=3X
k=1αk=4,a 1=3X
j=13X
k>jαjαk=−17,a 0=(−1)33Y
k=1αk=6 0.
Thus a possible cubic equation is x3+(−4)x2+(−17)x+(60) = 0. Of course, any multiple
ofx3−4x2−17x+ 60 = 0 will do just as well.
J
1.2 Trigonometric identities
So many of the applications of mathematics to physics and engineering are
concerned with periodic, and in particular sinusoidal, behaviour that a sure and
ready handling of the corresponding mathematical functions is an essential skill.Even situations with no obvious periodicity are often expressed in terms ofperiodic functions for the purposes of analysis. Later in this book whole chaptersare devoted to developing the techniques involved, but as a necessary prerequisitewe here establish (or remind the reader of) some standard identities with which heor she should be fully familiar, so that the manipulation of expressions containing
sinusoids becomes automatic and reliable. So as to emphasise the angular nature
of the argument of a sinusoid we will denote it in this section by θrather than x.
1.2.1 Single-angle identities
We give without proof the basic identity satisfied by the sinusoidal functions sin θ
and cos θ,n a m e l y
cos
2θ+s i n2θ=1. (1.15)
If sin θand cos θhave been defined geometrically in terms of the coordinates of
a point on a circle, a reference to the name of Pythagoras will suffice to establish
this result. If they have been defined by means of series (with θexpressed in
radians) then the reader should refer to Euler’s equation (3.23) on page 96, andnote that e
iθhas unit modulus if θis real.
10
1.2 TRIGONOMETRIC IDENTITIES
xy
x/primey/prime
OABP
TNR
M
Figure 1.2 Illustration of the compound-angle identities. Refer to the main
text for details.
Other standard single-angle formulae derived from (1.15) by dividing through
by various powers of sin θand cos θare
1+t a n2θ=s e c2θ. (1.16)
cot2θ+1=c o s e c2θ. (1.17)
1.2.2 Compound-angle identities
The basis for building expressions for the sinusoidal functions of compound
angles are those for the sum and difference of just two angles, since all othercases can be built up from these, in principle. Later we will see that a study of
complex numbers can provide a more efficient approach in some cases.
To prove the basic formulae for the sine and cosine of a compound angle
A+Bin terms of the sines and cosines of AandB, we consider the construction
shown in figure 1.2. It shows two sets of axes, OxyandOx
/primey/prime, with a common
origin but rotated with respect to each other through an angle A. The point
Plies on the unit circle centred on the common origin Oand has coordinates
cos(A+B),sin(A+B) with respect to the axes Oxyand coordinates cos B,sinB
with respect to the axes Ox/primey/prime.
Parallels to the axes Oxy(dotted lines) and Ox/primey/prime(broken lines) have been
drawn through P. Further parallels ( MRandRN)t ot h e Ox/primey/primeaxes have been
11
PRELIMINARY ALGEBRA
drawn through R, the point (0 ,sin(A+B)) in the Oxysystem. That all the angles
marked with the symbol •are equal to Afollows from the simple geometry of
right-angled triangles and crossing lines.
We now determine the coordinates of Pin terms of lengths in the figure,
expressing those lengths in terms of both sets of coordinates:
(i) cos B=x/prime=TN+NP=MR+NP
=ORsinA+RPcosA= sin( A+B)s i nA+c o s ( A+B)cosA;
(ii) sin B=y/prime=OM−TM=OM−NR
=ORcosA−RPsinA= sin( A+B)cosA−cos(A+B)s i nA.
Now, if equation (i) is multiplied by sin Aand added to equation (ii) multiplied
by cos A, the result is
sinAcosB+c o s AsinB= sin( A+B)(sin2A+c o s2A)=s i n ( A+B).
Similarly, if equation (ii) is multiplied by sin Aand subtracted from equation (i)
multiplied by cos A, the result is
cosAcosB−sinAsinB=c o s ( A+B)(cos2A+s i n2A)=c o s ( A+B).
Corresponding graphically based results can be derived for the sines and cosines
of the difference of two angles; however, they are more easily obtained by settingBto−Bin the previous results and remembering that sin Bbecomes−sinB
whilst cos Bis unchanged. The four results may be summarised by
sin(A±B)=s i n AcosB±cosAsinB (1.18)
cos(A±B)=c o s AcosB∓sinAsinB. (1.19)
Standard results can be deduced from these by setting one of the two angles
equal to πor to π/2:
sin(π−θ)=s i n θ, cos(π−θ)=−cosθ,sinparenleftbig
1
2π−θparenrightbig
(1.20)
=c o s θ, cosparenleftbig1
2π−θparenrightbig
=s i n θ, (1.21)
From these basic results many more can be derived. An immediate deduction,
obtained by taking the ratio of the two equations (1.18) and (1.19) and thendividing both the numerator and denominator of this ratio by cos AcosB,i s
tan(A±B)=tanA±tanB
1∓tanAtanB. (1.22)
One application of this result is a test for whether two lines on a graph
are orthogonal (perpendicular); more generally, it determines the angle between
them. The standard notation for a straight-line graph is y=mx+c,i nw h i c h m
is the slope of the graph and cis its intercept on the y-axis. It should be noted
that the slope mis also the tangent of the angle the line makes with the x-axis.
12
1.2 TRIGONOMETRIC IDENTITIES
Consequently the angle θ12between two such straight-line graphs is equal to the
difference in the angles they individually make with the x-axis, and the tangent
of that angle is given by (1.22):
tanθ12=tanθ1−tanθ2
1+t a n θ1tanθ2=m1−m2
1+m1m2. (1.23)
For the lines to be orthogonal we must have θ12=π/2, i.e. the final fraction on
the RHS of the above equation must equal ∞,a n ds o
m1m2=−1. (1.24)
A kind of inversion of equations (1.18) and (1.19) enables the sum or difference
of two sines or cosines to be expressed as the product of two sinusoids; theprocedure is typified by the following. Adding together the expressions given by(1.18) for sin( A+B) and sin( A−B) yields
sin(A+B)+s i n ( A−B)=2s i n AcosB.
If we now write A+B=CandA−B=D, this becomes
sinC+s i n D=2s i nparenleftbiggC+D
2parenrightbigg
cosparenleftbiggC−D
2parenrightbigg
. (1.25)
In a similar way each of the following equations can be derived:
sinC−sinD=2c o sparenleftbiggC+D
2parenrightbigg
sinparenleftbiggC−D
2parenrightbigg
, (1.26)
cosC+c o s D=2c o sparenleftbiggC+D
2parenrightbigg
cosparenleftbiggC−D
2parenrightbigg
, (1.27)
cosC−cosD=−2s i nparenleftbiggC+D
2parenrightbigg
sinparenleftbiggC−D
2parenrightbigg
. (1.28)
The minus sign on the right of the last of these equations should be noted; it may
help to avoid overlooking this ‘oddity’ to recall that if C>D then cos C<cosD.
1.2.3 Double- and half-angle identities
Double-angle and half-angle identities are needed so often in practical calculations
that they should be committed to memory by any physical scientist. They can beobtained by setting Bequal to Ain results (1.18) and (1.19). When this is done,
13
PRELIMINARY ALGEBRA
and use made of equation (1.15), the following results are obtained:
sin2θ=2s i n θcosθ, (1.29)
cos2θ=c o s2θ−sin2θ
=2c o s2θ−1
=1−2s i n2θ, (1.30)
tan2θ=2t a n θ
1−tan2θ. (1.31)
A further set of identities enables sinusoidal functions of θto be expressed
as polynomial functions of a variable t=t a n ( θ/2). They are not used in their
primary role until the next chapter, but we give a derivation of them here forreference.
Ift=t a n ( θ/2), then it follows from (1.16) that 1+ t
2=s e c2(θ/2) and cos( θ/2) =
(1 +t2)−1/2, whilst sin( θ/2) = t(1 +t2)−1/2. Now, using (1.29) and (1.30), we may
write:
sinθ=2s i nθ
2cosθ
2=2t
1+t2, (1.32)
cosθ=c o s2θ
2−sin2θ
2=1−t2
1+t2, (1.33)
tanθ=2t
1−t2. (1.34)
It can be further shown that the derivative of θwith respect to ttakes the
algebraic form 2 /(1 + t2). This completes a package of results that enables
expressions involving sinusoids, particularly when they appear as integrands, to
be cast in more convenient algebraic forms. The proof of the derivative propertyand examples of use of the above results are given in subsection (2.2.7).
We conclude this section with a worked example which is of such a commonly
occurring form that it might be considered a standard procedure.ISolve for θthe equation
asinθ+bcosθ=k,
where a, bandkare given real quantities.
To solve this equation we make use of result (1.18) by setting a=Kcosφandb=Ksinφ
for suitable values of Kandφ. We then have
k=Kcosφsinθ+Ksinφcosθ=Ksin(θ+φ),
with
K2=a2+b2and φ=t a n−1b
a.
Whether φlies in 0≤φ≤πor in−π<φ< 0 has to be determined by the individual
signs of aandb. The solution is thus
θ=s i n−1
/k
K
/
−φ,
14
1.3 COORDINATE GEOMETRY
with Kandφas given above. Notice that there is no re al solution to the original equation
if|k|>|K|=(a2+b2)1/2.
J
1.3 Coordinate geometry
We have already mentioned the standard form for a straight-line graph, namely
y=mx+c, (1.35)
representing a linear relationship between the independent variable xand the
dependent variable y.T h es l o p e mis equal to the tangent of the angle the line
makes with the x-axis whilst cis the intercept on the y-axis.
An alternative form for the equation of a straight line is
ax+by+k=0, (1.36)
to which (1.35) is clearly connected by
m=−a
band c=−k
b.
This form treats xandyon a more symmetrical basis, the intercepts on the two
axes being−k/aand−k/brespectively.
A power relationship between two variables, i.e. one of the form y=Axn,c a n
also be cast into straight-line form by taking the logarithms of both sides. Whilstit is normal in mathematical work to use natural logarithms (to base e, written
lnx), for practical investigations logarithms to base 10 are often employed. In
either case the form is the same, but it needs to be remembered which has beenused when recovering the value of Afrom fitted data. In the mathematical (base
e) form, the power relationship becomes
lny=nlnx+l nA. (1.37)
Now the slope gives the power n, whilst the intercept on the ln yaxis is ln A,
which yields A, either by exponentiation or by taking antilogarithms.
The other standard coordinate forms of two-dimensional curves that students
should know and recognise are those concerned with the conic sections – so called
because they can all be obtained by taking suitable sections across a (double)cone. Because the conic sections can take many different orientations and scalingstheir general form is complex,
Ax
2+By2+Cxy+Dx+Ey+F=0, (1.38)
but each can be represented by one of four generic forms, an ellipse, a parabola, a
hyperbola or, the degenerate form, a pair of straight lines. If they are reduced to
15
PRELIMINARY ALGEBRA
their standard representations, in which axes of symmetry are made to coincide
with the coordinate axes, the first three take the forms
(x−α)2
a2+(y−β)2
b2= 1 (ellipse), (1.39)
(y−β)2=4a(x−α) (parabola), (1.40)
(x−α)2
a2−(y−β)2
b2= 1 (hyperbola). (1.41)
Here, ( α, β) gives the position of the ‘centre’ of the curve, usually taken as
the origin (0 ,0) when this does not conflict with any imposed conditions. The
parabola equation given is that for a curve symmetric about a line parallel tothex-axis. For one symmetrical about a parallel to the y-axis the equation would
read ( x−α)
2=4a(y−β).
Of course, the circle is the special case of an ellipse in which b=aand the
equation takes the form
(x−α)2+(y−β)2=a2. (1.42)
The distinguishing characteristic of this equation is that when it is expressed in
the form (1.38) the coefficients of x2andy2are equal and that of xyis zero; this
property is not changed by any reorientation or scaling and so acts to identify a
general conic as a circle.
Definitions of the conic sections in terms of geometrical properties are also
available; for example, a parabola can be defined as the locus of a point thatis always at the same distance from a given straight line (the directrix )a si ti s
from a given point (the focus). When these properties are expressed in Cartesian
coordinates the above equations are obtained. For a circle, the defining propertyis that all points on the curve are a distance afrom ( α, β); (1.42) expresses this
requirement very directly. In the following worked example we derive the equation
for a parabola.IFind the equation of a parabola that has the line x=−aas its directrix and the point
(a,0)as its focus.
Figure 1.3 shows the situation in Cartesian coordinates. Expressing the defining requirement
thatPNandPFare equal in length gives
(x+a)=[ ( x−a)2+y2]1/2⇒(x+a)2=(x−a)2+y2
which, on expansion of the squared terms, immediately gives y2=4ax. This is (1.40) with
αandβboth set equal to zero.
J
Although the algebra is more complicated, the same method can be used to
derive the equations for the ellipse and the hyperbola. In these cases the distance
from the fixed point is a definite fraction, e, known as the eccentricity ,o ft h e
distance from the fixed line. For an ellipse 0 <e< 1, for a circle e=0 ,a n df o ra
hyperbola e>1. The parabola corresponds to the case e=1 .
16
1.3 COORDINATE GEOMETRY
xy
OP
FN
x=−a(a,0)(x, y)
Figure 1.3 Construction of a parabola using the point ( a,0) as the focus and
the line x=−aas the directrix.
The values of aandb(with a≥b) in equation (1.39) for an ellipse are related
toethrough
e2=a2−b2
a2
and give the lengths of the semi-axes of the ellipse. If the ellipse is centred on
the origin, i.e. α=β= 0, then the focus is ( −ae,0) and the directrix is the line
x=−a/e.
For each conic section curve, although we have two variables, xandy,t h e ya r e
not independent, since if one is given then the other can be determined. However,
determining ywhen xis given, say, involves solving a quadratic equation on each
occasion, and so it is convenient to have parametric representations of the curves.
A parametric representation allows each point on a curve to be associated witha unique value of a single parameter t. The simplest parametric representations
for the conic sections are as given below, though that for the hyperbola useshyperbolic functions, not formally introduced until chapter 3. That they do givevalid parameterizations can be verified by substituting them into the standard
forms (1.39) – (1.41); in each case the standard form is reduced to an algebraic
or trigonometric identity.
x=α+acosφ,y=β+bsinφ(ellipse),
x=α+at
2, y=β+2at (parabola),
x=α+acoshφ,y=β+bsinhφ(hyperbola).
As a final example illustrating several topics from this section we now prove
17
PRELIMINARY ALGEBRA
the well-known result that the angle subtended by a diameter at any point on a
circle is a right angle.ITaking the diameter to be the line joining Q=(−a,0)andR=(a,0)and the point Pto
be any point on the circle x2+y2=a2, prove that angle QP R is a right angle.
IfPis the point ( x, y), the slope of the line QPis
m1=y−0
x−(−a)=y
x+a.
That of RPis
m2=y−0
x−(a)=y
x−a.
Thus
m1m2=y2
x2−a2.
But, since Pis on the circle, y2=a2−x2and consequently m1m2=−1. From result (1.24)
this implies that QPandRPare orthogonal and that QP Ris therefore a right angle. Note
that this is true for anypoint Pon the circle.
J
1.4 Partial fractions
In subsequent chapters, and in particular when we come to study integration
in chapter 2, we will need to express a function f(x) that is the ratio of two
polynomials in a more manageable form. To remove some potential complexity
from our discussion we will assume that all the coefficients in the polynomialsare real, although this is not an essential simplification.
The behaviour of f(x) is crucially determined by the location of the zeroes of
its denominator, i.e. if f(x) is written as f(x)=g(x)/h(x) where both g(x)a n d
h(x) are polynomials †,t h e n f(x) changes extremely rapidly when xis close to
those values α
ithat are the roots of h(x) = 0. To make such behaviour explicit,
we write f(x) as a sum of terms such as A/(x−α)n,i nw h i c h Ais a constant, αis
one of the αithat satisfy h(αi)=0a n d nis a positive integer. Writing a function
in this way is known as expressing it in partial fractions .
Suppose, for the sake of definiteness, that we wish to express the function
f(x)=4x+2
x2+3x+2
†It is assumed that the ratio has been reduced so that g(x)a n d h(x) do not contain any common
factors, i.e. there is no value of xthat makes both vanish at the same time. We may also assume
without any loss of generality that the coefficient of the highest power of xinh(x) has been made
equal to unity, if necessary, by dividing both numerator and denominator by the coefficient of this
highest power.
18
1.4 PARTIAL FRACTIONS
in partial fractions, i.e. to write it as
f(x)=g(x)
h(x)=4x+2
x2+3x+2=A1
(x−α1)n1+A2
(x−α2)n2+···.
(1.43)
The first question that arises is that of how many terms there should be on
the right-hand side (RHS). Although some complications occur when h(x)h a s
repeated roots (these are considered below) it is clear that f(x) only becomes
infinite at the twovalues of x,α1andα2,t h a tm a k e h(x) = 0. Consequently the
RHS can only become infinite at the same two values of xand therefore contains
only two partial fractions – these are the ones shown explicitly. This argumentcan be trivially extended (again temporarily ignoring the possibility of repeatedroots of h(x)) to show that if h(x) is a polynomial of degree nthen there should be
nterms on the RHS, each containing a different root α
iof the equation h(αi)=0 .
A second general question concerns the appropriate values of the ni.T h i si s
answered by putting the RHS over a common denominator, which will clearly
have to be the product ( x−α1)n1(x−α2)n2···. Comparison of the highest power
ofxin this new RHS with the same power in h(x)s h o w st h a t n1+n2+···=n.
This result holds whether or not h(x) = 0 has repeated roots and, although we
do not give a rigorous proof, strongly suggests the correct conclusions that:
•The number of terms on the RHS is equal to the number of distinct roots of
h(x) = 0, each term having a different root αiin its denominator ( x−αi)ni;
•Ifαiis a multiple root of h(x) = 0 then the value to be assigned to niin (1.43) is
that of miwhen h(x) is written in the product form (1.9). Further, as discussed
on p. 23, Aihas to be replaced by a polynomial of degree mi−1.This is also
formally true for non-repeated roots, since then both miandniare equal to
unity.
Returning to our specific example we note that the denominator h(x)h a sz e r o e s
atx=α1=−1a n d x=α2=−2; these x-values are the simple (non-repeated)
roots of h(x) = 0. Thus the partial fraction expansion will be of the form
4x+2
x2+3x+2=A1
x+1+A2
x+2. (1.44)
We now list several methods available for determining the coefficients A1and
A2. We also remind the reader that, as with all the explicit examples and techniques
described, these methods are to be considered as models for the handling of any
ratio of polynomials, with or without characteristics which makes it a specialcase.
(i) The RHS can be put over a common denominator, in this case ( x+1)(x+2),
and then the coefficients of the various powers of xcan be equated in the
19
PRELIMINARY ALGEBRA
numerators on both sides of the equation. This leads to
4x+2= A1(x+2 )+ A2(x+1 ),
4=A1+A22=2 A1+A2.
Solving the simultaneous equations for A1andA2gives A1=−2a n d
A2=6.
(ii) A second method is to substitute two (or more generally n) different
values of xinto each side of (1.44) and so obtain two (or n) simultaneous
equations for the two (or n)c o n s t a n t s Ai. To justify this practical way of
proceeding it is necessary, strictly speaking, to appeal to method (i) above,
which establishes that there are unique values for A1andA2valid for
all values of x. It is normally very convenient to take zero as one of the
values of x, but of course any set will do. Suppose in the present case that
we use the values x=0a n d x= 1 and substitute in (1.44). The resulting
equations are
2
2=A1
1+A2
2,
6
6=A1
2+A2
3,
which on solution give A1=−2a n d A2= 6, as before. The reader can
easily verify that any other pair of values for x(except for a pair that
includes α1orα2) gives the same values for A1andA2.
(iii) The very reason why method (ii) fails if xis chosen as one of the roots
αiofh(x) = 0 can be made the basis for determining the values of the Ai
corresponding to non-multiple roots without having to solve simultaneous
equations. The method is conceptually more difficult than the other meth-ods presented here, and needs results from the theory of complex variables(chapter 20) to justify it. However, we give a practical ‘cookbook’ recipefor determining the coefficients.
(a) To determine the coefficient A
k, imagine the denominator h(x)
written as the product ( x−α1)(x−α2)···(x−αn), with any m-fold
repeated root giving rise to mfactors in parentheses.
(b) Now set xequal to αkand evaluate the expression obtained after
omitting the factor that reads αk−αk.
(c) Divide the value so obtained into g(αk); the result is the required
coefficient Ak.
For our specific example we find that in step (a) that h(x)=(x+1 ) ( x+2 )
and that in evaluating A1step (b) yields −1 + 2 = 1. Since g(−1) =
4(−1) + 2 =−2, step (c) gives A1as (−2)/(1), i.e in agreement with our
other evaluations. In a similar way A2is evaluated as ( −6)/(−1) = 6.
20
1.4 PARTIAL FRACTIONS
Thus any one of the methods listed above shows that
4x+2
x2+3x+2=−2
x+1+6
x+2.
The best method to use in any particular circumstance will depend on the
complexity, in terms of the degrees of the polynomials and the multiplicities ofthe roots of the denominator, of the function being considered and, to someextent, on the individual inclinations of the student; some prefer lengthy butstraightforward solution of simultaneous equations, whilst others feel more athome carrying shorter but more abstract calculations in their heads.
1.4.1 Complications and special cases
Having established the basic method for partial fractions, we now show, through
further worked examples, how some complications are dealt with by extensions
to the procedure. These extensions are introduced one at a time, but of course in
any practical application more than one may be involved.
The degree of the numerator is greater than or equal to that of the denominator
Although we have not specifically mentioned the fact, it will be apparent from
trying to apply method (i) of the previous subsection to such a case, that if the
degree of the numerator ( m) is not less than that of the denominator ( n) then the
ratio of two polynomials cannot be expressed in partial fractions.
To get round this difficulty it is necessary to start by dividing the denominator
h(x) into the numerator g(x) to obtain a further polynomial, which we will denote
bys(x), together with a function t(x)t h a t isa ratio of two polynomials for which
the degree of the numerator is less than that of the denominator. The functiont(x)cantherefore be expanded in partial fractions. As a formula,
f(x)=g(x)
h(x)=s(x)+t(x)≡s(x)+r(x)
h(x). (1.45)
It is apparent that the polynomial r(x)i st h e remainder obtained when g(x)i s
divided by h(x), and, in general, will be a polynomial of degree n−1. It is also
clear that the polynomial s(x) will be of degree m−n. Again, the actual division
process can be set out as an algebraic long division sum but is probably moreeasily handled by writing (1.45) in the form
g(x)=s(x)h(x)+r(x) (1.46)
or, more explicitly, as
g(x)=(s
m−nxm−n+sm−n−1xm−n−1+···+s0)h(x)+(rn−1xn−1+rn−2xn−2+···+r0)
(1.47)
and then equating coefficients.
21
PRELIMINARY ALGEBRA
We illustrate this procedure with the following worked example.IFind the partial fraction decomposition of the function
f(x)=x3+3x2+2x+1
x2−x−6.
Since the degree of the numerator is 3 and that of the denominator is 2, a preliminary
long division is necessary. The polynomial s(x) resulting from the division will have degree
3−2 = 1 and the remainder r(x) will be of degree 2 −1 = 1 (or less). Thus we write
x3+3x2+2x+1=( s1x+s0)(x2−x−6) + ( r1x+r0).
From equating the coefficients of the various powers of xon the two sides of the equation,
starting with the highest, we now obtain the simultaneous equations
1=s1,
3=s0−s1,
2=−s0−6s1+r1,
1=−6s0+r0.
These are readily solved, in the given order, to yield s1=1 , s0=4 , r1=1 2a n d r0= 25.
Thus f(x) can be written as
f(x)=x+4+12x+2 5
x2−x−6.
The last term can now be decomposed into partial fractions as previously. The zeroes of
the denominator are at x=3a n d x=−2 and the application of any method from the
previous subsection yields the respective constants as A1=1 21
5andA2=−1
5. Thus the
final partial fraction decomposition of f(x)i s
x+4+61
5(x−3)−1
5(x+2 ).
J
Factors of the form a2+x2in the denominator
We have so far assumed that the roots of h(x) = 0, needed for the factorisation of
the denominator of f(x), can always be found. In principle they always can but
in some cases they are not real. Consider, for example, attempting to express inpartial fractions a polynomial ratio whose denominator is h(x)=x
3−x2+2x−2.
Clearly x= 1 gives a zero of h(x), and so a first factorisation is ( x−1)(x2+2 ) .
However we cannot make any further progress because the factor x2+ 2 cannot
be expressed as ( x−α)(x−β)f o ra n yr e a l αandβ.
Complex numbers are introduced later in this book (chapter 3) and, when the
reader has studied them, he or she may wish to justify the procedure set outbelow. It can be shown to be equivalent to that already given, but the zeroes ofh(x) are now allowed to be complex and terms that are complex conjugates of
each other are combined to leave only real terms.
Since quadratic factors of the form a
2+x2that appear in h(x) cannot be reduced
to the product of two linear factors, partial fraction expansions including themneed to have numerators in the corresponding terms that are not simply constants
22
1.4 PARTIAL FRACTIONS
Aibut linear functions of x,i . e .o ft h ef o r m Bix+Ci. Thus, in the expansion,
linear terms (first-degree polynomials) in the denominator have constants (zero-degree polynomials) in their numerators, whilst quadratic terms (second-degreepolynomials) in the denominator have linear terms (first-degree polynomials) in
their numerators. As a symbolic formula, the partial fraction expansion of
g(x)
(x−α1)(x−α2)···(x−αp)(x2+a2
1)(x2+a2
2)···(x2+a2q)
should take the form
A1
x−α1+A2
x−α2+···+Ap
x−αp+B1x+C1
x2+a2
1+B2x+C2
x2+a2
2+···+Bqx+Cq
x2+a2q.
Of course, the degree of g(x) must be less than p+2q; if it is not, an initial
division must be carried out as demonstrated earlier.
Repeated factors in the denominator
Consider trying (incorrectly) to expand
f(x)=x−4
(x+1 ) ( x−2)2
in partial fraction form as follows:
x−4
(x+1 ) ( x−2)2=A1
x+1+A2
(x−2)2.
Multiplying both sides of this supposed equality by ( x+1 ) ( x−2)2produces an
equation whose LHS is linear in x, whilst its RHS is quadratic. This is clearly
wrong and so an expansion in the above form cannot be valid. The correction wemust make is very similar to that needed in the previous subsection, namely that
since ( x−2)
2is a quadratic polynomial the numerator of the term containing it
must be a first-degree polynomial, and not simply a constant.
The correct form for the part of the expansion containing the doubly repeated
root is therefore ( Bx+C)/(x−2)2. Using this form and either of methods (i) and
(ii) for determining the constants gives the full partial fraction expansion as
x−4
(x+1 ) ( x−2)2=−5
9(x+1 )+5x−16
9(x−2)2,
as the reader may verify.
Since any term of the form ( Bx+C)/(x−α)2can be written as
B(x−α)+C+Bα
(x−α)2=B
x−α+C+Bα
(x−α)2,
and similarly for multiply repeated roots, an alternative form for the part of the
partial fraction expansion containing a repeated root αis
D1
x−α+D2
(x−α)2+···+Dp
(x−α)p. (1.48)
23
PRELIMINARY ALGEBRA
In this form, all x-dependence has disappeared from the numerators but at the
expense of p−1 additional terms; the total number of constants to be determined
remains unchanged, as it must.
When describing possible methods of determining the constants in a partial
fraction expansion, we noted that method (iii), p. 20, which avoids the need to
solve simultaneous equations, is restricted to terms involving non-repeated roots.
In fact, it can be applied in repeated-root situations, when the expansion is putin the form (1.48), but only to find the constant in the term involving the largestinverse power of x−α,i . e .D
pin (1.48).
We conclude this section with a more protracted worked example that contains
all three of the complications discussed.IResolve the following expression F(x)into partial fractions:
F(x)=x5−2x4−x3+5x2−46x+ 100
(x2+6 ) ( x−2)2.
We note that the degree of the denominator (4) is not greater than that of the numerator
(5), and so we must start by dividing the latter by the former. It follows, from the differencein degrees and the coefficients of the highest powers in each, that the result will be a linearexpression s
1x+s0with the coefficient s1equal to 1. Thus the numerator of F(x)m u s tb e
expressible as
(x+s0)(x4−4x3+1 0x2−24x+ 24) + ( r3x3+r2x2+r1x+r0),
where the second factor in parentheses is the denominator of F(x) written as a polynomial.
Equating the coefficients of x4gives−2=−4+s0and fixes s0as 2. Equating the coefficients
of powers less than 4 gives equations involving the coefficients rias follows:
−1=−8+1 0+ r3,
5=−24 + 20 + r2,
−46 = 24−48 + r1,
100 = 48 + r0.
Thus the remainder polynomial r(x) can be constructed and F(x) written as
F(x)=x+2+−3x3+9x2−22x+5 2
(x2+6 ) ( x−2)2≡x+2+ f(x).
The polynomial ratio f(x) can now be expressed in partial fraction form, noting that its
denominator contains both a term of the form x2+a2and a repeated root. Thus
f(x)=Bx+C
x2+6+D1
x−2+D2
(x−2)2.
We could now put the RHS of this equation over the common denominator ( x2+6)(x−2)2
and find B,C,D 1andD2by equating coefficients of powers of x. It is quicker, however,
to use methods (iii) and (ii). Method (iii) gives D2as (−24 + 36−44 + 52) /(4 + 6) = 2.
We choose to evaluate the other coeffi cients by method (ii), and setting x=0 , x=1a n d
24
1.5 BINOMIAL EXPANSION
x=−1 gives respectively
52
24=C
6−D1
2+2
4,
36
7=B+C
7−D1+2,
86
63=C−B
7−D1
3+2
9.
These equations reduce to
4C−12D1=4 0,
B+C−7D1=2 2,
−9B+9C−21D1=7 2,
with solution B=0 , C=1 , D1=−3.
Thus, finally, we may re-write the original expression F(x) in partial fractions as
F(x)=x+2+1
x2+6−3
x−2+2
(x−2)2.
J
1.5 Binomial expansion
Earlier in this chapter we were led to consider functions containing powers of
the sum or difference of two terms, e.g. ( x−α)m. Later in this book we will find
numerous occasions on which we wish to write such a product of repeated factorsas a polynomial in xor, more generally, as a sum of terms each of which contains
powers of xandαseparately, as opposed to a power of their sum or difference.
To make the discussion general and the result applicable to a wide variety of
situations, we will consider the general expansion of f(x)=(x+y)
n,w h e r e xand
ymay stand for constants, variables or functions and, for the time being, nis a
positive integer. It may not be obvious what form the general expansion takesbut some idea can be obtained by carrying out the multiplication explicitly forsmall values of n. Thus we obtain successively
(x+y)
1=x+y,
(x+y)2=(x+y)(x+y)=x2+2xy+y2,
(x+y)3=(x+y)(x2+2xy+y2)=x3+3x2y+3xy2+y3,
(x+y)4=(x+y)(x3+3x2y+3xy2+y3)=x4+4x3y+6x2y2+4xy3+y4.
This does not establish a general formula, but the regularity of the terms in
the expansions and the suggestion of a pattern in the coefficients indicate that a
general formula for power nwill have n+ 1 terms, that the powers of xandyin
every term will add up to nand that the coefficients of the first and last terms
will be unity whilst those of the second and penultimate terms will be n.
25
PRELIMINARY ALGEBRA
In fact, the general expression, the binomial expansion for power n,i sg i v e nb y
(x+y)n=k=nsummationdisplay
k=0nCkxn−kyk, (1.49)
wherenCkis called the binomial coefficient and is expressed in terms of factorial
functions by n!/[k!(n−k)!]. Clearly, simply to make such a statement does not
constitute proof of its validity, but, as we will see in subsection 1.5.2, (1.49) canbeproved using a method called induction. Before turning to that proof, we
investigate some of the elementary properties of the binomial coefficients.
1.5.1 Binomial coefficients
As stated above, the binomial coefficients are defined by
nCk≡n!
k!(n−k)!≡parenleftbiggn
kparenrightbigg
for 0≤k≤n, (1.50)
where in the second identity we give a common alternative notation fornCk.
Obvious properties include
(i)nC0=nCn=1,
(ii)nC1=nCn−1=n,
(iii)nCk=nCn−k.
We note that, for any given n, the largest coefficient in the binomial expansion is
the middle one ( k=n/2) if nis even; the middle two coffficients ( k=1
2(n±1))
are equal largest if nis odd. Somewhat less obvious is the result
nCk+nCk−1=n!
k!(n−k)!+n!
(k−1)!(n−k+1 ) !
=n![(n+1−k)+k]
k!(n+1−k)!
=(n+1 ) !
k!(n+1−k)!=n+1Ck. (1.51)
An equivalent statement, in which khas been redefined as k+1 ,i s
nCk+nCk+1=n+1Ck+1. (1.52)
1.5.2 Proof of the binomial expansion
We are now in a position to prove the binomial expansion (1.49). In doing so, we
introduce the reader to a procedure applicable to certain types of problems and
known as the method of induction . The method is discussed much more fully in
subsection 1.7.1.
We start by assuming that(1.49) is true for some positive integer n=N.W e
26
1.6 PROPERTIES OF BINOMIAL COEFFICIENTS
now proceed to show that this implies that it must also be true for n=N+1 ,a s
follows:
(x+y)N+1=(x+y)Nsummationdisplay
k=0NCkxN−kyk
=Nsummationdisplay
k=0NCkxN+1−kyk+Nsummationdisplay
k=0NCkxN−kyk+1
=Nsummationdisplay
k=0NCkxN+1−kyk+N+1summationdisplay
j=1NCj−1x(N+1)−jyj,
where in the first line we have used the assumption and in the third line have
moved the second summation index, by unity by writing k+1= j.W en o w
separate off the first term of the first sum,NC0xN+1, and write it asN+1C0xN+1;
we can do this since, as noted in (i) following (1.50),nC0=1f o re v e r y n. Similarly,
the last term of the second summation can be replaced byN+1CN+1yN+1.
The remaining terms of each of the two summations are now written together,
with the summation index denoted by kin both terms. Thus
(x+y)N+1=N+1C0xN+1+Nsummationdisplay
k=1parenleftbigNCk+NCk−1parenrightbig
x(N+1)−kyk+N+1CN+1yN+1
=N+1C0xN+1+Nsummationdisplay
k=1N+1Ckx(N+1)−kyk+N+1CN+1yN+1
=N+1summationdisplay
k=0N+1Ckx(N+1)−kyk.
In going from the first to the second line we have used result (1.51). Now we
observe that the final overall equation is just the original assumed result (1.49)
but with n=N+ 1. Thus it has been shown that if the binomial expansion is
assumed to be true for n=N,t h e ni tc a nb e proved to be true for n=N+1 .B u t
it holds trivially for n= 1, and therefore for n= 2 also. By the same token it is
valid for n=3,4,..., and hence is established for all positive integers n.
1.6 Properties of binomial coefficients
1.6.1 Identities involving binomial coefficients
There are many identities involving the binomial coefficients that can be derived
directly from their definition, and yet more that follow from their appearance in
the binomial expansion. Only the most elementary ones, given earlier, are worth
committing to memory but, as illustrations, we now derive two results involvingsums of binomial coefficients.
27
PRELIMINARY ALGEBRA
The first is a further application of the method of induction. Consider the
proposal that, for any n≥1a n d k≥0,
n−1summationdisplay
s=0k+sCk=n+kCk+1. (1.53)
Notice that here n, the number of terms in the sum, is the parameter that varies,
kis a fixed parameter, whilst sis a summation index and does not appear on the
RHS of the equation.
Now we suppose that this statement about the value of the sum of the binomial
coefficientskCk,k+1Ck,...,k+n−1Ckis true for n=N. We next write down a series
with an extra term and determine the implications of the supposition for the newseries:
N+1−1summationdisplay
s=0k+sCk=N−1summationdisplay
s=0k+sCk+k+NCk
=N+kCk+1+N+kCk
=N+k+1Ck+1.
But this is just proposal (1.53) with nnow set equal to N+ 1. To obtain the last
line, we have used (1.52), with nset equal to N+k.
It only remains to consider the case n= 1, when the summation only contains
one term and (1.53) reduces to
kCk=1+kCk+1.
This is trivially valid for any ksince both sides are equal to unity, thus completing
the proof of (1.53) for all positive integers n.
The second result, which gives a formula for combining terms from two sets
of binomial coefficients in a particular way (a kind of ‘convolution’, for readerswho are already familiar with this term), is derived by applying the binomialexpansion directly to the identity
(x+y)
p(x+y)q≡(x+y)p+q.
Written in terms of binomial expansions, this reads
psummationdisplay
s=0pCsxp−sysqsummationdisplay
t=0qCtxq−tyt=p+qsummationdisplay
r=0p+qCrxp+q−ryr.
We now equate coefficients of xp+q−ryron the two sides of the equation, noting
that on the LHS all combinations of sandtsuch that s+t=rcontribute. This
gives as an identity that
rsummationdisplay
t=0pCr−tqCt=p+qCr=rsummationdisplay
t=0pCtqCr−t. (1.54)
28
1.6 PROPERTIES OF BINOMIAL COEFFICIENTS
We have specifically included the second equality to emphasise the symmetrical
nature of the relationship with respect to pandq.
Further identities involving the coefficients can be obtained by giving xandy
special values in the defining equation (1.49) for the expansion. If both are setequal to unity then we obtain (using the alternative notation so as to producefamiliarity with it)
parenleftbiggn
0parenrightbigg
+parenleftbiggn
1parenrightbigg
+parenleftbiggn
2parenrightbigg
+···+parenleftbiggn
nparenrightbigg
=2
n, (1.55)
whilst setting x=1a n d y=−1 yields
parenleftbiggn
0parenrightbigg
−parenleftbiggn
1parenrightbigg
+parenleftbiggn
2parenrightbigg
−···+(−1)nparenleftbiggn
nparenrightbigg
=0. (1.56)
1.6.2 Negative and non-integral values of n
Up till now we have restricted nin the binomial expansion to be a positive
integer. Negative values can be accommodated, but only at the cost of an infiniteseries of terms rather than the finite one represented by (1.49). For reasons thatare intuitively sensible and will be discussed in more detail in chapter 4, very
often we require an expansion in which, at least ultimately, successive terms in
the infinite series decrease in magnitude. For this reason, if x>y we consider
(x+y)
−m,w h e r e mitself is a positive integer, in the form
(x+y)n=(x+y)−m=x−mparenleftBig
1+y
xparenrightBig−m
.
Since the ratio y/xis less than unity, terms containing higher powers of it will be
small in magnitude, whilst raising the unit term to any power will not affect itsmagnitude. If y>x the roles of the two must be interchanged.
We can now state, but will not explicitly prove, the form of the binomial
expansion appropriate to negative values of n(nequal to−m):
(x+y)
n=(x+y)−m=x−m∞summationdisplay
k=0−mCkparenleftBigy
xparenrightBigk
, (1.57)
where the hitherto undefined quantity−mCk, which appears to involve factorials
of negative numbers, is given by
−mCk=(−1)km(m+1 )···(m+k−1)
k!=(−1)k(m+k−1)!
(m−1)!k!=(−1)km+k−1Ck.
(1.58)
The binomial coefficient on the extreme right of this equation has its normal
meaning and is well defined since m+k−1≥k.
Thus we have a definition of binomial coefficients for negative integer values
ofnin terms of those for positive n. The connection between the two may not
29
PRELIMINARY ALGEBRA
be obvious, but they are both formed in the same way in terms of recurrence
relations. Whatever the sign of n, the series of coefficientsnCkcan be generated
by starting withnC0= 1 and using the recurrence relation
nCk+1=n−k
k+1nCk. (1.59)
The difference is that for positive integer nthe series terminates when k=n,
whereas for negative nthere is no such termination – in line with the infinite
series of terms in the corresponding expansion.
Finally we note that, in fact, equation (1.59) generates the appropriate coef-
ficients for all values of n, positive or negative, integer or non-integer, with the
obvious exception of the case in which x=−yandnis negative. For non-integer
nthe expansion does not terminate, even if nis positive.
1.7 Some particular methods of proof
Much of the mathematics used by physicists and engineers is concerned with
obtaining a particular value, formula or function from a given set of data andstated conditions. However, just as it is essential in physics to formulate the basiclaws and so be able to set boundaries on what can or cannot happen, so itis important in mathematics to be able to state general propositions about theoutcomes that are or are not possible. To this end one attempts to establishtheorems that state in as general a way as possible mathematical results that
apply to particular types of situation. We conclude this introductory chapter by
describing two methods that can sometimes be used to prove particular classesof theorems.
The two general methods of proof are known as proof by induction (which
has already been met in this chapter) and proof by contradiction. They share
the common characteristic that at an early stage in the proof an assumptionis made that a particular (unproven) statement is true; the consequences ofthat assumption are then explored. In an inductive proof the conclusion isreached that the assumption is self-consistent and has other equally consistentbut broader implications, which are then applied to establish the general validityof the assumption. A proof by contradiction, however, establishes an internal
inconsistency and thus shows that the assumption is unsustainable; the natural
consequence of this is that the negative of the assumption is established as true.
Later in this book use will be made of these methods of proof to explore new
territory, e.g. to examine the properties of vector spaces, matrices and groups.
However, at this stage we will draw our illustrative and test examples from earliersections of this chapter and other topics in elementary algebra and number theory.
30
1.7 SOME PARTICULAR METHODS OF PROOF
1.7.1 Proof by induction
The proof of the binomial expansion given in subsection 1.5.2 and the identity
established in subsection 1.6.1 have already shown the way in which an inductiveproof is carried through. They also indicated the main limitation of the method,namely that only an initially supposed result can be proved. Thus the methodof induction is of no use for deducing a previously unknown result; a putative
equation or result has to be arrived at by some other means, usually by noticing
patterns or by trial and error using simple values of the variables involved. It
will also be clear that propositions that can be proved by induction are limitedto those containing a parameter that takes a range of integer values (usuallyinfinite).
For a proposition involving a parameter n, the five steps in a proof using
induction are as follows.
(i) Formulate the supposed result for general n.
(ii) Suppose (i) to be true for n=N(or more generally for all values of
n≤N;s e eb e l o w ) ,w h e r e Nis restricted to lie in the stated range.
(iii) Show, using only proven results and supposition (ii), that proposition (i)
is true for n=N+1 .
(iv) Demonstrate directly, and without any assumptions, that proposition (i) is
true when ntakes the lowest value in its range.
(v) It then follows from (iii) and (iv) that the proposition is valid for all values
ofnin the stated range.
(It should be noted that, although many proofs at stage (iii) require the validity
of the proposition only for n=N, some require it for all nless than or equal to N
– hence the form of inequality given in parentheses in the stage (ii) assumption.)
To illustrate further the method of induction, we now apply it to two worked
examples; the first concerns the sum of the squares of the first nnatural numbers.IProve that the sum of the squares of the first nnatural numbers is given by
nX
r=1r2=1
6n(n+ 1)(2 n+1 ). (1.60)
As previously we start by assuming the result is true for n=N. Then it follows that
N+1X
r=1r2=NX
r=1r2+(N+1 )2
=1
6N(N+ 1)(2 N+1 )+( N+1 )2
=1
6(N+1 ) [ N(2N+1 )+6 N+6 ]
=1
6(N+ 1)[(2 N+3 ) ( N+2 ) ]
=1
6(N+1 ) [ ( N+1 )+1 ] [ 2 ( N+1 )+1 ] .
31
PRELIMINARY ALGEBRA
This is precisely the original assumption, but with Nreplaced by N+ 1. To complete the
proof we only have to verify (1.60) for n= 1. This is trivially done and establishes the
result for all positive n. The same and related results are obtained by a different method
in subsection 4.2.5.
J
Our second example is somewhat more complex and involves two nested proofs
by induction: whilst trying to establish the main result by induction, we find that
we are faced with a second proposition which itself requires an inductive proof.IShow that Q(n)=n4+2n3+2n2+nis divisible by 6 (without remainder) for all positive
integer values of n.
Again we start by assuming the result is true for some particular value Nofn, whilst
noting that it is trivially true for n= 0. We next examine Q(N+ 1), writing each of its
terms as a binomial expansion:
Q(N+1 )=( N+1 )4+2 (N+1 )3+2 (N+1 )2+(N+1 )
=(N4+4N3+6N2+4N+1 )+2 ( N3+3N2+3N+1 )
+2 (N2+2N+1 )+( N+1 )
=(N4+2N3+2N2+N)+( 4 N3+1 2N2+1 4N+6 ).
Now, by our assumption, the group of terms within the first parentheses in the last line
is divisible by 6 and clearly so are the terms 12 N2and 6 within the second parentheses.
Thus it comes down to deciding whether 4 N3+1 4Ni sd i v i s i b l eb y6–o re q u i v a l e n t l y ,
whether R(N)=2 N3+7Nis divisible by 3.
To settle this latter question we try using a second inductive proof and assume that
R(N)isdivisible by 3 for N=M, whilst again noting that the proposition is trivially true
forN=M= 0. This time we examine R(M+1 ) :
R(M+1 )=2 ( M+1 )3+7 (M+1 )
=2 (M3+3M2+3M+1 )+7 ( M+1 )
=( 2M3+7M)+3 ( 2 M2+2M+3 )
By assumption, the first group of terms in the last line is divisible by 3 and the second
group is patently so. We thus conclude that R(N) is divisible by 3 for all N≥M,a n d
taking M= 0 shows that it is divisible by 3 for all N.
We can now return to the main proposition and conclude that since R(N)=2 N3+7N
is divisible by 3, 4 N3+1 2N2+1 4N+ 6 is divisible by 6. This in turn establishes that the
divisibility of Q(N+ 1) by 6 follows from the assumption that Q(N) divides by 6. Since
Q(0) clearly divides by 6, the proposition in the question is established for all values of n.
J
1.7.2 Proof by contradiction
The second general line of proof, but again one that is normally only useful when
the result is already suspected, is proof by contradiction. The questions it canattempt to answer are only those that can be expressed in a proposition that
is either true or false. Clearly, it could be argued that any mathematical result
can be so expressed but, if the proposition is no more than a guess, the chancesof success are negligible. Valid propositions containing even modest formulae
32
1.7 SOME PARTICULAR METHODS OF PROOF
are either the result of true inspiration or, much more normally, yet another
reworking of an old chestnut!
The essence of the method is to exploit the fact that mathematics is required
to be self-consistent, so that, for example, two calculations of the same quantity,starting from the same given data but proceeding by different methods, must givethe same answer. Equally, it must not be possible to follow a line of reasoning anddraw a conclusion that contradicts either the input data or any other conclusionbased upon the same data.
It is this requirement on which the method of proof by contradiction is based.
The crux of the method is to assume that the proposition to be proved isnottrue, and then use this incorrect assumption and ‘watertight’ reasoning to
draw a conclusion that contradicts the assumption. The only way out of the
self-contradiction is then to conclude that the assumption was indeed false andtherefore that the proposition is true.
It must be emphasised that once a (false) contrary assumption has been made,
every subsequent conclusion in the argument mustfollow of necessity. Proof by
contradiction fails if at any stage we have to admit ‘this may or may not bethe case’. That is, each step in the argument must be a necessary consequence of
results that precede it (taken together with the assumption), rather than simply apossible consequence.
It should also be added that if no contradiction can be found using sound
reasoning based on the assumption then no conclusion can be drawn about eitherthe proposition or its negative and some other approach must be tried.
We illustrate the general method with an example in which the mathematical
reasoning is straightforward so that attention can be focussed on the structure ofthe proof.IA rational number ris a fraction r=p/qin which pandqare integers with qpositive.
Further, ris expressed in its lowest terms, any integer common factor of pandqhaving
been divided out.
Prove that the square root of an integer mcannot be a rational number, unless the square
root itself is an integer.
We begin by supposing that the stated result is nottrue and that we canwrite an equation
√m=r=p
qfor integers m, p, q with q/negationslash=1.
It then follows that p2=mq2.B u t ,s i n c e ris expressed in its lowest terms, pandq,a n d
hence p2andq2, have no factors in common whilst mis an integer. This is only possible
ifq=1a n d p2=m. This conclusion contradicts the requirement that q/negationslash=1a n ds ol e a d s
to the conclusion that it was wrong to suppose that√mcan be expressed as a non-integer
rational number. This completes the proof of the statement in the question.
J
Our second worked example, also taken from elementary number theory,
involves slightly more complicated mathematical reasoning but again exhibits thestructure associated with this type of proof.
33
PRELIMINARY ALGEBRAIThe prime integers piare labelled in ascending order, thus p1=1,p2=2,p5=7,etc.
Show that there is no largest prime number.
Assume, on the contrary, that there is a largest prime and let it be pN.C o n s i d e rn o wt h e
number qformed by multiplying together all the primes from p1topNand then adding
one to the product, i.e.
q=p1p2···pN+1.
By our assumption pNis the largest prime, and so no number can have a prime factor
greater than this. However, for every prime pi(i=1,2,...,N ) the quotient q/p ihas the
form Mi+( 1/pi)w i t h Mian integer and 1 /pinon-integer. This means that q/p icannot be
an integer and so picannot be a divisor of q.
Since qis not divisible by any of the (assumed) finite set of primes, it must be itself
ap r i m e .A s qis also clearly greater than pN, we have a contradiction. Thus it follows
that our assumption that there is a largest prime integer must be false, and so it has beenproved that there is no largest prime integer.
It should be noted that the given construction for qdoes not generate all the primes
that actually exist (e.g. for N=3,q= 7 rather than the next actual prime value of 5, is
found), but this does not matter for the purposes of our proof by contradiction.J
1.7.3 Necessary and sufficient conditions
As the final topic in this introductory chapter, we consider briefly the notion
of, and distinction between, necessary and sufficient conditions in the contextof proving a mathematical proposition. In ordinary English the distinction iswell defined, and that distinction is maintained in mathematics. However, in
the authors’ experience students tend to overlook it and assume (wrongly) that,
having proved that the validity of proposition Aimplies the truth of proposition
B, it follows by ‘reversing the argument’ that the validity of Bautomatically
implies that of A.
As an example, let proposition Abe that an integer Nis divisible without
remainder by 6, and proposition Bbe that Nis divisible without remainder by
2. Clearly, if Ais true then it follows that Bis true, i.e. Ais a sufficient condition
forB; it is not however a necessary condition, as is trivially shown by taking N
as 8. Conversely, the same value of Nshows that whilst the validity of Bis a
necessary condition for Ato hold, it is not sufficient.
An alternative terminology to ‘necessary’ and ‘sufficient’ often employed by
mathematicians is that of ‘if’ and ‘only if’, particularly in the combination ‘if andonly if’ which is usually written as IFF or denoted by a double-headed arrow⇐⇒. The equivalent statements can be summarised by
AifBA is true if Bis true or B=⇒A,
Bis a sufficient condition for AB =⇒A,
Aonly if BA is true only if Bis true or A=⇒B,
Bis a necessary consequence of AA =⇒B,
34
1.7 SOME PARTICULAR METHODS OF PROOF
AIFFBA is true if and only if Bis true or B⇐⇒ A,
AandBnecessarily imply each other B⇐⇒ A.
Although at this stage in the book we are able to employ for illustrative purposes
only simple and fairly obvious results, the following example is given as a modelof how necessary and sufficient conditions should be proved. The essential pointis that for the second part of the proof (whether it be the ‘necessary’ part or the
‘sufficient’ part) one needs to start again from scratch; more often than not, the
lines of the second part of the proof will notbe simply those of the first written
in reverse order.IProve that ( A) a function f(x)is a quadratic polynomial with zeroes at x=2andx=3
if and only if ( B) the function f(x)has the form λ(x2−5x+6)withλa non-zero constant.
(1) Assume A,i . e .t h a t f(x)isa quadratic polynomial with zeroes at x=2a n d x=3 .L e t
its form be ax2+bx+cwith a/negationslash= 0. Then we have
4a+2b+c=0,
9a+3b+c=0,
and subtraction shows that 5 a+b=0a n d b=−5a. Substitution of this into the first of
the above equations gives c=−4a−2b=−4a+1 0a=6a. Thus, it follows that
f(x)=a(x2−5x+6 ) w i t h a/negationslash=0,
and establishes the ‘ Aonly if B’ part of the stated result.
(2) Now assume that f(x)hasthe form λ(x2−5x+6 )w i t h λa non-zero constant. Firstly
we note that f(x) is a quadratic polynomial, and so it only remains to prove that its
zeroes occur at x=2a n d x=3 .C o n s i d e r f(x) = 0, which, after dividing through by the
non-zero constant λ,g i v e s
x2−5x+6=0 .
We proceed by using a technique known as completing the square , for the purposes of
illustration, although the factorisation of the above equation should be clear to the reader.Thus we write
x
2−5x+(5
2)2−(5
2)2+6=0 ,
(x−5
2)2=1
4,
x−5
2=±1
2.
The two roots of f(x) = 0 are therefore x=2a n d x=3 ;t h e s e x-values give the zeroes
off(x). This establishes the second (‘ AifB’) part of the result. Thus we have shown
that the assumption of either condition implies the validity of the other and the proof iscomplete.J
It should be noted that the propositions have to be carefully and precisely
formulated. If, for example, the word ‘quadratic’ were omitted from A, statement
Bwould still be a sufficient condition for Abut not a necessary one, since f(x)
could then be x3−4x2+x+6a n d Awould not require B. Omitting the constant
λfrom the stated form of f(x)i nBhas the same effect. Conversely, if Awere to
state that f(x)=3 ( x−2)(x−3) then Bwould be a necessary condition for Abut
not a sufficient one.
35
PRELIMINARY ALGEBRA
1.8 Exercises
Polynomial equations
1.1 Continue the investigation of equation (1.7), namely
g(x)=4 x3+3x2−6x−1,
as follows.
(a) Make a table of values of g(x) for integer values of xbetween−2a n d2 .U s e
it and the information derived in the text to draw a graph and so determinethe roots of g(x) = 0 as accurately as possible.
(b) Find one accurate root of g(x) = 0 by inspection and hence determine precise
values for the other two roots.
(c) Show that f(x)=4 x
3+3x2−6x−k= 0 has only one real root unless
−5≤k≤7
4.
1.2 Determine how the number of real roots of the equation
g(x)=4 x3−17x2+1 0x+k=0
depends upon k. Are there any cases for which the equation has exactly two
distinct real roots?
1.3 Continue the analysis of the polynomial equation
f(x)=x7+5x6+x4−x3+x2−2=0 ,
investigated in subsection 1.1.1, as follows.
(a) By writing the fifth-degree polynom ial appearing in the expression for f/prime(x)
in the form 7 x5+3 0x4+a(x−b)2+c, show that there is in fact only one
positive root of f(x)=0 .
(b) By evaluating f(1),f(0) and f(−1), and by inspecting the form of f(x)f o r
negative values of x, determine what you can about the positions of the real
roots of f(x)=0 .
1.4 Given that x=2i so n er o o to f
g(x)=2 x4+4x3−9x2−11x−6=0 ,
use factorisation to determine how many real roots it has.
1.5 Construct the quadratic equations that have the following pairs of roots: (a)
−6,−3; (b) 0 ,4; (c) 2 ,2; (d) 3 + 2 i,3−2i,w h e r e i2=−1.
1.6 Use the results of (i) equation (1.13), (ii) equation (1.12) and (iii) equation (1.14)
to prove that if the roots of 3 x3−x2−10x+8=0a r e α1,α2andα3then
(a)α−1
1+α−1
2+α−1
3=5/4,
(b)α2
1+α2
2+α2
3=6 1/9,
(c)α3
1+α3
2+α3
3=−125/27.
(d) Convince yourself that eliminating (say) α2andα3from (i), (ii) and (iii) does
notgive a simple explicit way of finding α1.
Trigonometric identities
1.7 Prove that
cosπ
12=√3+1
2√2
by considering
36
1.8 EXERCISES
(a) the sum of the sines of π/3a n d π/6,
(b) the sine of the sum of π/3a n d π/4.
1.8 (a) Use the fact that sin( π/6) = 1 /2 to prove that tan( π/12) = 2−√3.
(b) Use the result of (a) to show further that tan( π/24) = q(2−q)w h e r e
q2=2+√3.
1.9 Find the real solutions of
(a) 3 sin θ−4cos θ=2,
(b) 4 sin θ+3c o s θ=6,
(c) 12 sin θ−5c os θ=−6.
1.10 If s= sin( π/8), prove that
8s4−8s2+1=0 ,
and hence show that s=[ ( 2−√2)/4]1/2.
1.11 Find all the solutions of
sinθ+s i n4 θ=s i n2 θ+s i n3 θ
that lie in the range −π<θ≤π. What is the multiplicity of the solution θ=0 ?
Coordinate geometry
1.12 Obtain in the form (1.38) the equations that describe the following:
(a) a circle of radius 5 with its centre at (1 ,−1);
(b) the line 2 x+3y+ 4 = 0 and the line orthogonal to it which passes through
(1,1);
(c) an ellipse of eccentricity 0 .6 with centre (1 ,1) and its major axis of length 10
parallel to the y-axis.
1.13 Determine the forms of the conic sections described by the following equations:
(a)x2+y2+6x+8y=0 ;
(b) 9 x2−4y2−54x−16y+2 9=0 ;
(c) 2 x2+2y2+5xy−4x+y−6=0 ;
(d)x2+y2+2xy−8x+8y=0.
1.14 For the ellipse
x2
a2+y2
b2=1
with eccentricity e, the two points ( −ae,0) and ( ae,0) are known as its foci. Show
that the sum of the distances from anypoint on the ellipse to the foci is 2 a.( T h e
constancy of the sum of the distances from two fixed points can be used as analternative defining property of an ellipse.)
Partial fractions
1.15 Resolve the following into partial fractions using the three methods given in
section 1.4, verifying that the same decomposition is obtained by each method:
(a)2x+1
x2+3x−10, (b)4
x2−3x.
1.16 Express the following in partial fraction form:
(a)2x3−5x+1
x2−2x−8, (b)x2+x−1
x2+x−2.
37
PRELIMINARY ALGEBRA
1.17 Rearrange the following functions in partial fraction form:
(a)x−6
x3−x2+4x−4, (b)x3+3x2+x+1 9
x4+1 0x2+9.
1.18 Resolve the following into partial fractions in such a way that xdoes not appear
in any numerator:
(a)2x2+x+1
(x−1)2(x+3 ),(b)x2−2
x3+8x2+1 6x,(c)x3−x−1
(x+3 )3(x+1 ).
Binomial expansion
1.19 Evaluate those of the following that are defined: (a)5C3,( b )3C5,( c )−5C3,( d )
−3C5.
1.20 Use a binomial expansion to evaluate 1 /√
4.2 to five places of decimals, and
compare it with the accurate answer obtained using a calculator.
Proof by induction and contradiction
1.21 Prove by induction that
nX
r=1r=1
2n(n+1 ) a n dnX
r=1r3=1
4n2(n+1 )2.
1.22 Prove by induction that
1+r+r2+···+rk+···+rn=1−rn+1
1−r.
1.23 Prove that 32n+7 ,w h e r e nis a non-negative integer, is divisible by 8.
1.24 If a sequence of terms unsatisfies the recurrence relation un+1=( 1−x)un+nx
with u1= 0 then show, using induction, that for n≥1
un=1
x[nx−1+( 1−x)n].
1.25 Prove by induction that
nX
r=11
2rtan
/θ
2r
/
=1
2ncot
/θ
2n
/
−cotθ.
1.26 The quantities aiin this exercise are all positive real numbers.
(a) Show that
a1a2≤
/a1+a2
2
/2
.
(b) Hence prove by induction on mthat
a1a2···ap≤
/a1+a2+···+ap
p
/p
,
where p=2mwith ma positive integer. Note that each increase of mby
unity doubles the number of factors in the product.
1.27 Establish the values of kfor which the binomial coefficientpCkis divisible by p
when pis a prime number. Use your result and the method of induction to prove
thatnp−nis divisible by pfor all integers nand all prime numbers p. Deduce
thatn5−nis divisible by 30 for any integer n.
38
1.9 HINTS AND ANSWERS
1.28 An arithmetic progression of integers anis one in which an=a0+nd,w h e r e a0
anddare integers and ntakes successive values 0 ,1,2,....
(a) Show that if any one term of the progression is the cube of an integer then
so are infinitely many others.
(b) Show that no cube of an integer can be expressed as 7 n+5 for some positive
integer n.
1.29 Prove, by the method of contradiction, that the equation
xn+an−1xn−1+···+a1x+a0=0,
in which all the coefficients aiare integers, cannot have a rational root, unless
that root is an integer. Deduce that any integral root must be a divisor of a0and
hence find all rational roots of
(a)x4+6x3+4x2+5x+4=0 ,
(b)x4+5x3+2x2−10x+6=0 .
Necessary and sufficient conditions
1.30 Prove that the equation ax2+bx+c=0 ,i nw h i c h a,bandcare real and a>0,
has two real distinct solutions IFF b2>4ac.
1.31 For the real variable x, show that a sufficient, but not necessary, condition for
f(x)=x(x+ 1)(2 x+ 1) to be divisible by 6 is that xis an integer.
1.32 Given that at least one of aandb, and at least one of candd, are non-zero,
show that ad=bcis both a necessary and sufficient condition for the equations
ax+by=0,
cx+dy=0,
to have a solution in which at least one of xandyis non-zero.
1.33 The coefficients aiin the polynomial Q(x)=a4x4+a3x3+a2x2+a1xare all
integers. Show that Q(n) is divisible by 24 for all integers n≥0 if and only if all
the following conditions are satisfied:(i) 2a
4+a3is divisible by 4;
(ii)a4+a2is divisible by 12;
(iii)a4+a3+a2+a1is divisible by 24.
1.9 Hints and answers
1.1 (b) The roots are 1 ,1
8(−7+√
33) =−0.1569,1
8(−7−√
33) =−1.593. (c)−5a n d
7
4are the values of kthat make f(−1) and f(1
2) equal to zero.
1.2 Three distinct roots if −43
27<k<75
4; two distinct roots, one at x=1
3,i fk=−43
27;
two distinct roots, one at x=5
2,i fk=75
4.
1.3 (a) a=4,b=3
8andc=23
6are all positive. Therefore f/prime(x)>0f o ra l l x>0.
(b)f(1) = 5, f(0) =−2a n d f(−1) = 5, and so there is at least one root in each
of the ranges 0 <x< 1a n d−1<x< 0. (x7+5x6)+(x4−x3)+(x2−2)
is positive definite for −5<x<−√2. There are therefore no roots in this
range, but there must be one to the left of x=−5.
1.4 g(x)=( x−2)(x+ 3)(2 x2+2x+ 1). The quadratic has complex roots and so
g(x)=0h a so n l yt w or e a lr o o t s .
1.5 (a) x2+9x+1 8=0 ;( b ) x2−4x=0 ;( c ) x2−4x+4=0 ;( d ) x2−6x+1 3=0 .
1.6 (a) Divide (iii) by (ii). (b) Consider (i)2−2(iii). (c) Consider (i)3−3(i)(iii)+3(ii).
1.7 (a) Use sin( π/4) = 1 /√2. (b) Use results (1.20) and (1.20).
1.8 (a) Use (1.32). (b) Use (1.34) and show that q4+1=4 q2.
39
PRELIMINARY ALGEBRA
1.9 (a) 1 .339. (b) No solution because 62>42+32.( c )−0.0849.
1.10 Use the formula for sin(2 π/8) and square both sides. sin2(π/4) = 1 /2.
1.11 Show that the equation is equivalent to sin(5 θ/2)sin( θ)sin(θ/2) = 0.
Solutions are −4π/5,−2π/5,0,2π/5,4π/5,π. Its multiplicity is 3.
1.12 (a) x2+y2−2x+2y−23 = 0 .
(b) The orthogonal line is 3 x−2y−1 = 0. The pair of lines has equation
6x2−6y2+5xy+1 0x−11y−4=0 .
(c) The minor axis has length 8. The ellipse has equation
25x2+1 6y2−50x−32y−359 = 0.
1.13 (a) A circle of radius 5 centred on ( −3,−4).
(b) A hyperbola with ‘centre’ (3 ,−2) and ‘semi-axes’ 2 and 3.
(c) The expression factorises into two lines, x+2y−3=0a n d2 x+y+2=0 .
(d) Write the expression as ( x+y)2=8 (x−y) to see that it represents a parabola
passing through the origin with the line x+y= 0 as its axis of symmetry.
1.14 Show that y2can be replaced by a2−x2−a2e2+x2e2and that the two lengths
area+exanda−ex.
1.15 (a)5
7(x−2)+9
7(x+5 ), (b)−4
3x+4
3(x−3).
1.16 (a) 2 x+4+109
6(x−4)+5
6(x+2 ), (b) 1−1
3(x+2 )+1
3(x−1).
1.17 (a)x+2
x2+4−1
x−1, (b)x+1
x2+9+2
x2+1.
1.18 (a)1
(x−1)2+1
(x−1)+1
(x+3 ).
(b)−1
8x+9
8(x+4 )−7
2(x+4 )2.
(c)1
8
/
−1
(x+1 )+9
(x+3 )−54
(x+3 )2+100
(x+3 )3
/
.
1.19 (a) 10, (b) not defined, (c) −35, (d)−21.
1.20 Write it as1
2( 1+0 .05)−1/2and evaluate−1/2Ckup to k= 3. The approximate and
accurate values agree to five places of decimals, both giving 0 .48795.
1.23 Write 32nas 8m−7.
1.25 Use the half-angle formulae of equations (1.32) to (1.34) to relate functions of
θ/2kto those of θ/2k+1.
1.26 (a) Consider ( a1−a2)2≥0. (b) Write a1+···+ap=Aandap+1+···+ap+p=B
and use result (a) to replace the product ABwith an expression involving the
sumA+B. Note that 2 p=2m+1.
1.27 Divisible for k=1,2,...,p−1.Expand ( n+1 )pasnp+
Pp−1
1pCknk+ 1. Apply
the stated result for p= 5. Note that n5−n=n(n−1)(n+1 ) ( n2+ 1); the product
of any three consecutive integers must divide by both 2 and 3.
1.28 (a) Suppose aN=a0+Nd=m3is the largest cube; then consider ( m+d)3.
(b) Suppose that 7 N+5= m3. Show that ( m−7)3differs from this by a multiple
of 7. Deduce that q3must have the form 7 n+5f o rs o m e qin 0≤q≤7.
Show explicitly that this is not so. Note. It is not sufficient to carry out the
explicit valuations and rely on the construct from part (a).
1.29 By assuming x=p/qwith q/negationslash= 1, show that a fraction −pn/qis equal to an
integer an−1pn−1+···+a1pqn−2+a0qn−1. This is a contradiction and is only
resolved if q= 1 and the root is an integer.
(a) The only possible candidates are ±1,±2,±4. None is a root.
(b) The only possible candidates are ±1,±2,±3,±6. Only−3i sar o o t .
40
1.9 HINTS AND ANSWERS
1.30 (i) Show that the equation can be reformulated as
a
/
x+b
2a
/2
=b2−4ac
4a.
(ii) If the real distinct solutions are αandβ, show that b=−(α+β)a,a n d
c=αβa. Then consider the inequality 0 <(α−β)2=(α+β)2−4αβ.
1.31 f(x) can be written as x(x+1 ) ( x+2 )+ x(x+1 ) ( x−1). Each term consists of
the product of three consecutive integers of which one must therefore divide by
2 and (a different) one by 3. Thus each term separately divides by 6 and sotherefore does f(x). Note that if xis the root of 2 x
3+3x2+x−24 = 0 that lies
near the non-integer value x=1.826 then x(x+ 1)(2 x+ 1) = 24 and therefore
divides by 6.
1.32 (i) If x/negationslash= 0, multiply the first equation by dand the second by band subtract.
Ify/negationslash= 0, multiply by cand arespectively instead. (ii) Suppose a/negationslash=0a n d
c/negationslash= 0. Whilst ensuring that no possible division by zero occurs, deduce that
the equations are consistent, with solution x=−(b/a)y=−(d/c)yfor arbitrary
non-zero y.
1.33 Note that, e.g., the condition for 6 a4+a3to be divisible by 4 is the same as the
condition for 2 a4+a3to be divisible by 4.
For the necessary (only if) part of the proof set n=1,2,3 and take integer
combinations of the resulting equations.For the sufficient (if) part of the proof use the stated conditions to prove theproposition by induction. Note that n
3−nis divisible by 6 and that n2+3nis
even.
41
2
Preliminary calculus
This chapter is concerned with the formalism of probably the most widely used
mathematical technique in the physical sciences, namely the calculus. The chapter
divides into two sections. The first deals with the process of differentiation and the
second with its inverse process, integration. The material covered is essential forthe remainder of the book and serves as a reference. Readers who have previouslystudied these topics should ensure familiarity by looking at the worked examplesin the main text and by attempting the exercises at the end of the chapter.
2.1 Differentiation
Differentiation is the process of determining how quickly or slowly a function
varies, as the quantity on which it depends, its argument , is changed. More
specifically it is the procedure for obtaining an expression (numerical or algebraic)for the rate of change of the function with respect to its argument. Familiarexamples of rates of change include acceleration (the rate of change of velocity)and chemical reaction rate (the rate of change of chemical composition). Both
acceleration and reaction rate give a measure of the change of a quantity with
respect to time. However, differentiation may also be applied to changes withrespect to other quantities, for example the change in pressure with respect to achange in temperature.
Although it will not be apparent from what we have said so far, differentiation
is in fact a limiting process, that is, it deals only with the infinitesimal change inone quantity resulting from an infinitesimal change in another.
2.1.1 Differentiation from first principles
Let us consider a function f(x) that depends on only one variable x, together with
numerical constants, for example, f(x)=3 x
2orf(x)=s i n xorf(x)=2+3 /x.
42
2.1 DIFFERENTIATION
A
P
xf(x)
x+∆xf(x+∆x)
∆f
θ∆x
Figure 2.1 The graph of a function f(x) showing that the gradient of the
function at P,g i v e nb yt a n θ, is approximately equal to ∆ f/∆x.
Figure 2.1 shows an example of such a function. Near any particular point,
P, the value of the function changes by an amount ∆ f,s a y ,a s xchanges
by a small amount ∆ x. The slope of the tangent to the graph of f(x)a t P
is then approximately ∆ f/∆x, and the change in the value of the function is
∆f=f(x+∆x)−f(x). In order to calculate the true value of the gradient, or
first derivative , of the function at P, we must let ∆ xbecome infinitesimally small.
We therefore define the first derivative of f(x)a s
f/prime(x)≡df(x)
dx≡lim
∆x→0f(x+∆x)−f(x)
∆x, (2.1)
provided that the limit exists. The limit will depend in almost all cases on the
value of x. If the limit does exist at a point x=athen the function is said to be
differentiable at a; otherwise it is said to be non-differentiable at a.T h ef o r m a l
concept of a limit and its existence or non-existence is discussed in chapter 4; forpresent purposes we will adopt an intuitive approach.
In the definition (2.1), we allow ∆ xto tend to zero from either positive or
negative values and require the same limit to be obtained in both cases. Afunction that is differentiable at ais necessarily continuous at a(there must be
no jump in the value of the function at a), though the converse is not necessarily
true. This latter assertion is illustrated in figure 2.1: the function is continuousat the ‘kink’ Abut the two limits of the gradient as ∆ xtends to zero from
43
PRELIMINARY CALCULUS
positive or negative values are different and so the function is not differentiable
atA.
It should be clear from the above discussion that near the point Pwe may
approximate the change in the value of the function, ∆ f, that results from a small
change ∆ xinxby
∆f≈df(x)
dx∆x. (2.2)
As one would expect, the approximation improves as the value of ∆ xis reduced.
In the limit in which the change ∆ xbecomes infinitesimally small, we denote it
by the differential dx, and (2.2) reads
df=df(x)
dxdx. (2.3)
Thisequality relates the infinitesimal change in the function, df, to the infinitesimal
change dxthat causes it.
So far we have discussed only the first derivative of a function. However, we
can also define the second derivative as the gradient of the gradient of a function.
Again we use the definition (2.1) but now with f(x) replaced by f/prime(x). Hence the
second derivative is defined by
f/prime/prime(x)≡lim
∆x→0f/prime(x+∆x)−f/prime(x)
∆x, (2.4)
provided that the limit exists. A physical example of a second derivative is the
second derivative of the distance travelled by a particle with respect to time. Sincethe first derivative of distance travelled gives the particle’s velocity, the secondderivative gives its acceleration.
We can continue in this manner, the nth derivative of the function f(x)b e i n g
defined by
f
(n)(x)≡lim
∆x→0f(n−1)(x+∆x)−f(n−1)(x)
∆x. (2.5)
It should be noted that with this notation f/prime(x)≡f(1)(x),f/prime/prime(x)≡f(2)(x), etc., and
that formally f(0)(x)≡f(x).
All this should be familiar to the reader, though perhaps not with such formal
definitions. The following example shows the differentiation of f(x)=x2from first
principles. In practice, however, it is desirable simply to remember the derivatives
of standard functions; the techniques given in the remainder of this section canbe applied to find more complicated derivatives.
44
2.1 DIFFERENTIATIONIFind from first principles the derivative with respect to xoff(x)=x2.
Using the definition (2.1),
f/prime(x) = lim
∆x→0f(x+∆x)−f(x)
∆x
= lim
∆x→0(x+∆x)2−x2
∆x
= lim
∆x→02x∆x+( ∆x)2
∆x
= lim
∆x→0(2x+∆x).
As ∆ xtends to zero, 2 x+∆xtends towards 2 x, hence
f/prime(x)=2 x.
J
Derivatives of other functions can be obtained in the same way. The derivatives
of some simple functions are listed below (note that ais a constant):
d
dx(xn)=nxn−1,d
dx(eax)=aeax,d
dx(lnax)=1
x,
d
dx(sinax)=acosax,d
dx(cosax)=−asinax,d
dx(secax)=asecaxtanax,
d
dx(tanax)=asec2ax,d
dx(cosec ax)=−acosec axcotax,
d
dx(cotax)=−acosec2ax,d
dxparenleftBig
sin−1x
aparenrightBig
=1√
a2−x2,
d
dxparenleftBig
cos−1x
aparenrightBig
=−1√
a2−x2,d
dxparenleftBig
tan−1x
aparenrightBig
=a
a2+x2.
Differentiation from first principles emphasises the definition of a derivative as
the gradient of a function. However, for most practical purposes, returning to the
definition (2.1) is time consuming and does not aid our understanding. Instead, as
mentioned above, we employ a number of techniques, which use the derivativeslisted above as ‘building blocks’, to evaluate the derivatives of more complicatedfunctions than hitherto encountered. S ubsections 2.1.2–2.1.7 develop the methods
required.
2.1.2 Differentiation of products
As a first example of the differentiation of a more complicated function, we
consider finding the derivative of a function f(x) that can be written as the
product of two other functions of x,n a m e l y f(x)=u(x)v(x). For example, if
f(x)= x
3sinxthen we might take u(x)= x3and v(x)=s i n x. Clearly the
45
PRELIMINARY CALCULUS
separation is not unique. (In the given example, possible alternative break-ups
would be u(x)=x2,v(x)=xsinx,o re v e n u(x)=x4tanx,v(x)=x−1cosx.)
The purpose of the separation is to split the function into two (or more) parts,
of which we know the derivatives (or at least we can evaluate these derivatives
more easily than that of the whole). We would gain little, however, if we did
not know the relationship between the derivative of fand those of uandv.
Fortunately, they are very simply related, as we shall now show.
Since f(x) is written as the product u(x)v(x), it follows that
f(x+∆x)−f(x)=u(x+∆x)v(x+∆x)−u(x)v(x)
=u(x+∆x)[v(x+∆x)−v(x)] + [ u(x+∆x)−u(x)]v(x).
From the definition of a derivative (2.1),
df
dx= lim
∆x→0f(x+∆x)−f(x)
∆x
= lim
∆x→0braceleftbigg
u(x+∆x)bracketleftbiggv(x+∆x)−v(x)
∆xbracketrightbigg
+bracketleftbiggu(x+∆x)−u(x)
∆xbracketrightbigg
v(x)bracerightbigg
.
In the limit ∆ x→0, the factors in square brackets become dv/dx anddu/dx
(by the definitions of these quantities) and u(x+∆x) simply becomes u(x).
Consequently we obtain
df
dx=d
dx[u(x)v(x)] =u(x)dv(x)
dx+du(x)
dxv(x). (2.6)
In primed notation and without writing the argument xexplicitly, (2.6) is stated
concisely as
f/prime=(uv)/prime=uv/prime+u/primev. (2.7)
This is a general result obtained without making any assumptions about the
specific forms f,uandv, other than that f(x)=u(x)v(x). In words, the result
reads as follows. The derivative of the product of two functions is equal to the
first function times the derivative of the second plus the second function times thederivative of the first .IFind the derivative with respect to xoff(x)=x3sinx.
Using the product rule, (2.6),
d
dx(x3sinx)=x3d
dx(sinx)+d
dx(x3)si nx
=x3cosx+3x2sinx.
J
The product rule may readily be extended to the product of three or more
functions. Considering the function
f(x)=u(x)v(x)w(x) (2.8)
46
2.1 DIFFERENTIATION
and using (2.6), we obtain, as before omitting the argument,
df
dx=ud
dx(vw)+du
dxvw.
Using (2.6) again to expand the first term on the RHS gives the complete result
d
dx(uvw)=uvdw
dx+udv
dxw+du
dxvw (2.9)
or
(uvw)/prime=uvw/prime+uv/primew+u/primevw. (2.10)
It is readily apparent that this can be extended to products containing any number
nof factors; the expression for the derivative will then consist of nterms with
the prime appearing in successive terms on each of the nfactors in turn. This is
probably the easiest way to recall the product rule.
2.1.3 The chain rule
Products are just one type of complicated function that we may encounter in
differentiation. Another is the function of a function, e.g. f(x)=( 3+ x2)3=u(x)3,
where u(x)=3+ x2.I f∆ f,∆uand ∆ xare small finite quantities, it follows that
∆f
∆x=∆f
∆u∆u
∆x;
As the quantities become infinitesimally small we obtain
df
dx=df
dudu
dx. (2.11)
This is the chain rule , which we must apply when differentiating a function of a
function.IFind the derivative with respect to xoff(x)=( 3+ x2)3.
Rewriting the function as f(x)=u3,w h e r e u(x)=3+ x2, and applying (2.11) we find
df
dx=3u2du
dx=3u2d
dx(3 +x2)=3 u2×2x=6x(3 +x2)2.
J
Similarly, the derivative with respect to xoff(x)=1 /v(x)m a yb eo b t a i n e db y
rewriting the function as f(x)=v−1and applying (2.11):
df
dx=−v−2dv
dx=−1
v2dv
dx. (2.12)
The chain rule is also useful for calculating the derivative of a function fwith
respect to xwhen both xandfare written in terms of a variable (or parameter),
sayt.
47
PRELIMINARY CALCULUSIFind the derivative with respect to xoff(t)=2 at,w h e r e x=at2.
We could of course substitute for ta n dt h e nd i ff e r e n t i a t e fas a function of x, but in this
case it is quicker to use
df
dx=df
dtdt
dx=2a1
2at=1
t,
where we have used the fact that
dt
dx=
/dx
dt
/−1
.
J
2.1.4 Differentiation of quotients
Applying (2.6) for the derivative of a product to a function f(x)=u(x)[1/v(x)],
we may obtain the derivative of the quotient of two factors. Thus
f/prime=parenleftBigu
vparenrightBig/prime
=uparenleftbigg1
vparenrightbigg/prime
+u/primeparenleftbigg1
vparenrightbigg
=uparenleftbigg
−v/prime
v2parenrightbigg
+u/prime
v,
where (2.12) has been used to evaluate (1 /v)/prime. This can now be rearranged into
the more convenient and memorisable form
f/prime=parenleftBigu
vparenrightBig/prime
=vu/prime−uv/prime
v2. (2.13)
This can be expressed in words as the derivative of a quotient is equal to the bottom
times the derivative of the top minus the top times the derivative of the bottom, all
over the bottom squared .IFind the derivative with respect to xoff(x)=s i n x/x.
Using (2.13) with u(x)=s i n x,v(x)=xand hence u/prime(x)=c o s x,v/prime(x) = 1, we find
f/prime(x)=xcosx−sinx
x2=cosx
x−sinx
x2.
J
2.1.5 Implicit differentiation
So far we have only differentiated functions written in the form y=f(x).
However, we may not always be presented with a relationship in this simpleform. As an example consider the relation x
3−3xy+y3=2 .I nt h i sc a s ei ti s
not possible to rearrange the equation to give yas a function of x. Nevertheless,
by differentiating term by term with respect to x(implicit differentiation ), we can
fi n dt h ed e r i v a t i v eo f y.
48
2.1 DIFFERENTIATIONIFind dy/dx ifx3−3xy+y3=2.
Differentiating each term in the equation with respect to xwe obtain
d
dx(x3)−d
dx(3xy)+d
dx(y3)=d
dx(2),
⇒3x2−
/
3xdy
dx+3y
/
+3y2dy
dx=0,
where the derivative of 3 xyhas been found using the product rule. Hence, rearranging for
dy/dx ,
dy
dx=y−x2
y2−x.
Note that dy/dx is a function of both xandyand cannot be expressed as a function of x
only.
J
2.1.6 Logarithmic differentiation
In circumstances in which the variable with respect to which we are differentiating
is an exponent, taking logarithms and then differentiating implicitly is the simplestway to find the derivative.IFind the derivative with respect to xofy=ax.
To find the required derivative we first take logarithms and then differentiate implicitly:
lny=l nax=xlna⇒1
ydy
dx=l na.
Now, rearranging and substituting for y, we find
dy
dx=ylna=axlna.
J
2.1.7 Leibniz’ theorem
We have discussed already how to find the derivative of a product of two or
more functions. We now consider Leibniz’ theorem , which gives the corresponding
results for the higher derivatives of products.
Consider again the function f(x)=u(x)v(x). We know from the product rule
thatf/prime=uv/prime+u/primev. Using the rule once more for each of the products, we obtain
f/prime/prime=(uv/prime/prime+u/primev/prime)+(u/primev/prime+u/prime/primev)
=uv/prime/prime+2u/primev/prime+u/prime/primev.
Similarly, differentiating twice more gives
f/prime/prime/prime=uv/prime/prime/prime+3u/primev/prime/prime+3u/prime/primev/prime+u/prime/prime/primev,
f(4)=uv(4)+4u/primev/prime/prime/prime+6u/prime/primev/prime/prime+4u/prime/prime/primev/prime+u(4)v.
49
PRELIMINARY CALCULUS
The pattern emerging is clear and strongly suggests that the results generalise to
f(n)=nsummationdisplay
r=0n!
r!(n−r)!u(r)v(n−r)=nsummationdisplay
r=0nCru(r)v(n−r), (2.14)
where the fraction n!/[r!(n−r)!] is identified with the binomial coefficientnCr
(see chapter 1). To prove that this is so, we use the method of induction as follows.
Assume that (2.14) is valid for nequal to some integer N.T h e n
f(N+1)=Nsummationdisplay
r=0NCrd
dxparenleftbig
u(r)v(N−r)parenrightbig
=Nsummationdisplay
r=0NCr[u(r)v(N−r+1)+u(r+1)v(N−r)]
=Nsummationdisplay
s=0NCsu(s)v(N+1−s)+N+1summationdisplay
s=1NCs−1u(s)v(N+1−s),
where we have substituted summation index sforrin the first summation, and
forr+ 1 in the second. Now, from our earlier discussion of binomial coefficients,
equation (1.51), we have
NCs+NCs−1=N+1Cs
and so, after separating out the first term of the first summation and the last
term of the second, obtain
f(N+1)=NC0u(0)v(N+1)+Nsummationdisplay
s=1N+1Csu(s)v(N+1−s)+NCNu(N+1)v(0).
ButNC0=1=N+1C0andNCN=1=N+1CN+1, and so we may write
f(N+1)=N+1C0u(0)v(N+1)+Nsummationdisplay
s=1N+1Csu(s)v(N+1−s)+N+1CN+1u(N+1)v(0)
=N+1summationdisplay
s=0N+1Csu(s)v(N+1−s).
This is just (2.14) with nset equal to N+ 1. Thus, assuming the validity of (2.14)
forn=Nimplies its validity for n=N+ 1. However, when n=1e q u a t i o n
(2.14) is simply the product rule, and this we have already proved directly. These
results taken together establish the validity of (2.14) for all nand prove Leibniz’
theorem.
50
2.1 DIFFERENTIATION
Q
A
BCf(x)
xS
Figure 2.2 A graph of a function, f(x), showing how differentiation corre-
sponds to finding the gradient of the function at a particular point. Points B,
QandSare stationary points (see text).IFind the third derivative of the function f(x)=x3sinx.
Using (2.14) we immediately find
f/prime/prime/prime(x)=6s i n x+3 ( 6 x)cosx+3 ( 3 x2)(−sinx)+x3(−cosx)
=3 ( 2−3x2)si nx+x(18−x2)c osx.
J
2.1.8 Special points of a function
We have interpreted the derivative of a function as the gradient of the function
at the relevant point (figure 2.1). If the gradient is zero at some point then thefunction is said to have a stationary point there. Clearly, in graphical terms, this
corresponds to a horizontal tangent to the graph at that point.
Stationary points may be divided into three categories and an example of each
is shown in figure 2.2. Point Bis said to be a minimum since the function increases
in value in both directions away from it. Point Qis said to be a maximum since the
function decreases in both directions away from it. Note that Bis not the overall
minimum value of the function and Qis not the overall maximum; rather, they
are a local minimum and a local maximum. The third type of stationary point isthestationary point of inflection ,S. In this case the function falls in the positive
x-direction and rises in the negative x-direction so that Sis neither a maximum
nor a minimum. Nevertheless, the gradient of the function is zero at S,i . e .t h e
graph of the function is flat there, and this justifies our calling it a stationary
point. Of course, a point at which the gradient of the function is zero but the
function rises in the positive x-direction and falls in the negative x-direction is
also a stationary point of inflection.
51
PRELIMINARY CALCULUS
The above distinction between the three types of stationary point has been
made rather descriptively. However, it is possible to define and distinguish sta-tionary points mathematically. From their definition as points of zero gradient,all stationary points must be characterised by df/dx = 0. In the case of the
minimum, B, the slope, i.e. df/dx , changes from negative at Ato positive at C
through zero at B. Thus df/dx is increasing and so the second derivative d
2f/dx2
must be positive. Conversely, at the maximum, Q, we must have that d2f/dx2is
negative.
It is less obvious, but intuitively reasonable, that at S,d2f/dx2is zero. This may
be inferred from the following observations. To the left of Sthe curve is concave
upwards so that df/dx is increasing with xand hence d2f/dx2>0. To the right
ofS, however, the curve is concave downwards so that df/dx is decreasing with
xand hence d2f/dx2<0.
In summary, at a stationary point df/dx =0a n d
(i) for a minimum, d2f/dx2>0,
(ii) for a maximum, d2f/dx2<0,
(iii) for a stationary point of inflection, d2f/dx2=0a n d d2f/dx2changes sign
through the point.
In case (iii), a stationary point of inflection, in order that d2f/dx2changes sign
through the point we normally require d3f/dx3/negationslash= 0 at that point. This simple
rule can fail for some functions, however, and in general if the first non-vanishingderivative of f(x) at the stationary point is f
(n)then if nis even the point is a
maximum or minimum and if nis odd the point is a stationary point of inflection.
This may be seen from the Taylor expansion (see equation (4.17)) of the function
about the stationary point, but it is not proved here.IFind the positions and natures of the stationary points of the function
f(x)=2 x3−3x2−36x+2.
The first criterion for a stationary point is that df/dx = 0, and hence we set
df
dx=6x2−6x−36 = 0 ,
from which we obtain
(x−3)(x+2 )=0 .
Hence the stationary points are at x=3a n d x=−2. To determine the nature of the
stationary point we must evaluate d2f/dx2:
d2f
dx2=1 2x−6.
52
2.1 DIFFERENTIATION
Gf(x)
x
Figure 2.3 The graph of a function f(x) that has a general point of inflection
at the point G.
Now, we examine each stationary point in turn. For x=3 , d2f/dx2= 30. Since this is
positive, we conclude that x= 3 is a minimum. Similarly, for x=−2,d2f/dx2=−30 and
sox=−2 is a maximum.
J
So far we have concentrated on stationary points, which are defined to have
df/dx = 0. We have found that at a stationary point of inflection d2f/dx2is
also zero and changes sign. This naturally leads us to consider points at whichd
2f/dx2is zero and changes sign but at which df/dx isnot, in general, zero. Such
points are called general points of inflection or simply points of inflection . Clearly,
a stationary point of inflection is a special case for which df/dx is also zero.
At a general point of inflection the graph of the function changes from beingconcave upwards to concave downwards (or vice versa), but the tangent to thecurve at this point need not be horizontal. A typical example of a general pointof inflection is shown in figure 2.3.
The determination of the stationary points of a function, together with the
identification of its zeroes, infinities and possible asymptotes, is usually sufficientto enable a graph of the function showing most of its significant features to besketched. Some examples for the reader to try are included in the exercises at the
end of this chapter.
2.1.9 Curvature of a function
In the previous section we saw that at a point of inflection of the function
f(x), the second derivative d
2f/dx2changes sign and passes through zero. The
corresponding graph of fshows an inversion of its curvature at the point of
inflection. We now develop a more quantitative measure of the curvature of afunction (or its graph), which is applicable at general points and not just in theneighbourhood of a point of inflection.
As in figure 2.1, let θbe the angle made with the x-axis by the tangent at a
53
PRELIMINARY CALCULUS
C
PQρ
θ θ+∆θ∆θ
xf(x)
Figure 2.4 Two neighbouring tangents to the curve f(x) whose slopes differ
by ∆ θ. The angular separation of the corresponding radii of the circle of
curvature is also ∆ θ.
point Pon the curve f=f(x), with tan θ=df/dx evaluated at P. Now consider
also the tangent at a neighbouring point Qon the curve, and suppose that it
makes an angle θ+∆θwith the x-axis, as illustrated in figure 2.4.
It follows that the corresponding normals at PandQ, which are perpendicular
to the respective tangents, also intersect at an angle ∆ θ. Furthermore, their point
of intersection, Cin the figure, will be the position of the centre of a circle that
approximates the arc PQ, at least to the extent of having the same tangents at
the extremities of the arc. This circle is called the circle of curvature .
For a finite arc PQ, the lengths of CPandCQwill not, in general, be equal,
as they would be if f=f(x)werein fact the equation of a circle. But, as Q
is allowed to tend to P,i . e .a s∆ θ→0, they do become equal, their common
value being ρ, the radius of the circle, known as the radius of curvature . It follows
immediately that the curve and the circle of curvature have a common tangentatPand lie on the same side of it. The reciprocal of the radius of curvature, ρ
−1,
defines the curvature of the function f(x) at the point P.
The radius of curvature can be defined more mathematically as follows. The
length ∆ sof arc PQis approximately equal to ρ∆θand, in the limit ∆ θ→0, this
relationship defines ρas
ρ= lim
∆θ→0∆s
∆θ=ds
dθ. (2.15)
It should be noted that, as sincreases, θmay increase or decrease according to
whether the curve is locally concave upwards (i.e. shaped as if it were near a
minimum in f(x)) or concave downwards. This is reflected in the sign of ρ,w h i c h
therefore also indicates the position of the curve (and of the circle of curvature)
54
2.1 DIFFERENTIATION
relative to the common tangent, above or below. Thus a negative value of ρ
indicates that the curve is locally concave downwards and that the tangent liesabove the curve.
We next obtain an expression for ρ, not in terms of sandθbut in terms
ofxandf(x). The expression, though somewhat cumbersome, follows from the
defining equation (2.15), the defining property of θthat tan θ=df/dx≡f
/primeand
the fact that the rate of change of arc length with xis given by
ds
dx=bracketleftBigg
1+parenleftbiggdf
dxparenrightbigg2bracketrightBigg1/2
. (2.16)
This last result, simply quoted here, is proved more formally in subsection 2.2.13.
From the chain rule (2.11) it follows that
ρ=ds
dθ=ds
dxdx
dθ. (2.17)
Differentiating both sides of tan θ=df/dx with respect to xgives
sec2θdθ
dx=d2f
dx2≡f/prime/prime,
from which, using sec2θ=1+t a n2θ=1+( f/prime)2,w ec a no b t a i n dx/dθ as
dx
dθ=1+t a n2θ
f/prime/prime=1+(f/prime)2
f/prime/prime. (2.18)
Substituting (2.16) and (2.18) into (2.17) then yields the final expression for ρ,
ρ=bracketleftbig
1+(f/prime)2bracketrightbig3/2
f/prime/prime. (2.19)
It should be noted that the quantity in brackets is always positive and that
its3
2th root is also taken as positive. The sign of ρis thus solely determined by
that of d2f/dx2, in line with our previous discussion relating the sign to whether
the curve is concave or convex upwards. If, as happens at a point of inflection,d
2f/dx2is zero then ρis formally infinite and the curvature of f(x)i sz e r o .A s
d2f/dx2changes sign on passing through zero, both the local tangent and the
circle of curvature change from their initial positions to the opposite side of thecurve.
55
PRELIMINARY CALCULUSIShow that the radius of curvature at the point (x, y)on the ellipse
x2
a2+y2
b2=1
has magnitude (a4y2+b4x2)3/2/(a4b4)and the opposite sign to y. Check the special case
b=a, for which the ellipse becomes a circle.
Differentiating the equation of the ellipse with respect to xgives
2x
a2+2y
b2dy
dx=0
and so
dy
dx=−b2x
a2y.
A second differentiation, using (2.13), then yields
d2y
dx2=−b2
a2
/y−xy/prime
y2
/
=−b4
a2y3
/y2
b2+x2
a2
/
=−b4
a2y3,
where we have used the fact that ( x, y) lies on the ellipse. We note that d2y/dx2,a n d
hence ρ, has the opposite sign to y3, and hence to y. Substituting in (2.19) gives for the
magnitude of the radius of curvature
|ρ|=
/////
/
1+b4x2/(a4y2)
/3/2
−b4/(a2y3)
/////=(a4y2+b4x2)3/2
a4b4.
For the special case b=a,|ρ|reduces to a−2(y2+x2)3/2and, since x2+y2=a2,t h i si n
turn gives|ρ|=a, as expected.
J
The discussion in this section has been confined to the behaviour of curves
that lie in one plane; examples of the application of curvature to the bending ofloaded beams and to particle orbits under the influence of a central forces can befound in the exercises at the ends of later chapters. A more general treatment ofcurvature in three dimensions is given in section 10.3, where a vector approach isadopted.
2.1.10 Theorems of differentiation
Rolle’s theorem
Rolle’s theorem (figure 2.5) states that if a function f(x) is continuous in the
range a≤x≤c, is differentiable in the range a<x<c and satisfies f(a)=f(c)
then for at least one point x=b,w h e r e a<b<c ,f
/prime(b) = 0. Thus Rolle’s
theorem states that for a well-behaved (continuous and differentiable) functionthat has the same value at two points either there is at least one stationary pointbetween those points or the function is a constant between them. The validity of
the theorem is immediately apparent from figure 2.5 and a full analytic proof will
not be given. The theorem is used in deriving the mean value theorem, which wenow discuss.
56
2.1 DIFFERENTIATION
a b cf(x)
x
Figure 2.5 The graph of a function f(x), showing that if f(a)=f(c)t h e na t
one point at least between x=aandx=cthe graph has zero gradient.
a b cC
Af(a)f(x)
xf(c)
Figure 2.6 The graph of a function f(x); at some point x=bit has the same
gradient as the line AC.
Mean value theorem
The mean value theorem (figure 2.6) states that if a function f(x) is continuous
in the range a≤x≤cand differentiable in the range a<x<c then
f/prime(b)=f(c)−f(a)
c−a, (2.20)
for at least one value bwhere a<b<c . Thus the mean value theorem states
that for a well-behaved function the gradient of the line joining two points on thecurve is equal to the slope of the tangent to the curve for at least one interveningpoint.
The proof of the mean value theorem is found by examination of figure 2.6, as
follows. The equation of the line ACis
g(x)=f(a)+(x−a)f(c)−f(a)
c−a,
57
PRELIMINARY CALCULUS
and hence the difference between the curve and the line is
h(x)=f(x)−g(x)=f(x)−f(a)−(x−a)f(c)−f(a)
c−a.
Since the curve and the line intersect at AandC,h(x) = 0 at both of these points.
Hence, by an application of Rolle’s theorem, h/prime(x) = 0 for at least one point b
between AandC. Differentiating our expression for h(x), we find
h/prime(x)=f/prime(x)−f(c)−f(a)
c−a,
and hence at b,w h e r e h/prime(x)=0 ,
f/prime(b)=f(c)−f(a)
c−a.
Applications of Rolle’s theorem and the mean value theorem
Since the validity of Rolle’s theorem is intuitively obvious, given the conditions
imposed on f(x), it will not be surprising that the problems that can be solved
by applications of the theorem alone are relatively simple ones. Nevertheless we
will illustrate it with the following example.IWhat semi-quantitative results can be deduced by applying Rolle’s theorem to the follow-
ing functions f(x),w i t h aandcchosen so that f(a)=f(c)=0?( i )sinx, (ii)cosx, (iii)
x2−3x+2,( i v ) x2+7x+3,( v )2x3−9x2−24x+k.
(i) If the consecutive values of xthat make sin x=0a r e α1,α2,...(actually x=nπ,f o r
any integer n) then Rolle’s theorem implies that the derivative of sin x,n a m e l yc o s x,h a s
at least one zero lying between each pair of values αiandαi+1.
(ii) In an exactly similar way, we conclude that the derivative of cos x,n a m e l y−sinx,
has at least one zero lying between consecutive pairs of zeroes of cos x.T h e s et w o
results taken together (but neither separately) imply that sin xand cos xhave interleaving
zeroes.
(iii) For f(x)=x2−3x+2 ,f(a)=f(c)=0i f aandcare taken as 1 and 2 respectively.
Rolle’s theorem then implies that f/prime(x)=2 x−3=0h a sas o l u t i o n x=bwith bin the
range 1 <b< 2. This is obviously so, since b=3/2.
(iv) With f(x)= x2+7x+ 3, the theorem tells us that if there are two roots of
x2+7x+ 3 = 0 then they have the root of f/prime(x)=2 x+ 7 = 0 lying between them. Thus
any (real) roots of x2+7x+ 3 = 0 lie on either side of x=−7/2. The actual roots are
(−7±√
37)/2.
(v) If f(x)=2 x3−9x2−24x+kthen f/prime(x) = 0 is the equation 6 x2−18x−24 = 0,
which has solutions x=−1a n d x=4 .C o n s e q u e n t l y ,i f α1andα2are two different roots
off(x)=0t h e na tl e a s to n eo f −1 and 4 must lie in the open interval α1toα2.I f ,a si s
the case for a certain range of values of k,f(x) = 0 has three roots, α1,α2andα3,t h e n
α1<−1<α2<4<α3.
58
2.1 DIFFERENTIATION
In each case, as might be expected, the applic ation of Rolle’s theorem does no more than
focus attention on particular ranges of values; it does not yield precise answers.
J
Direct verification of the mean value theorem is straightforward when it is
applied to simple functions. For example, if f(x)=x2, it states that there is a
value bin the interval a<b<c such that
c2−a2=f(c)−f(a)=(c−a)f/prime(b)=(c−a)2b.
This is clearly so, since b=(a+c)/2 satisfies the relevant criteria.
As a slightly more complicated example we may consider a cubic equation, say
f(x)=x3+2x2+4x−6 = 0, between two specified values of x,s a y1a n d2 .I n
this case we need to verify that there is a value of xlying in the range 1 <x< 2
that satisfies
18−1=f(2)−f(1) = (2−1)f/prime(x)=1 ( 3 x2+4x+4 ).
This is easily done, either by evaluating 3 x2+4x+4−17 at x=1a n da t x=2a n d
checking that the values have opposite signs or by solving 3 x2+4x+4−17 = 0
and showing that one of the roots lies in the stated interval.
The following applications of the mean value theorem establish some general
inequalities for two common functions.IDetermine inequalities satisfied by lnxandsinxfor suitable ranges of the real variable x.
Since for positive values of its argument the derivative of ln xisx−1, the mean value
theorem gives us
lnc−lna
c−a=1
b
for some bin 0<a<b<c . Further, since a<b<c implies that c−1<b−1<a−1,w e
have
1
c<lnc−lna
c−a<1
a,
or, multiplying through by c−aand writing c/a=xwhere x>1,
1−1
x<lnx<x−1.
Applying the mean value theorem to sin xshows that
sinc−sina
c−a=c o s b
for some blying between aandc.I faandcare restricted to lie in the range 0 ≤a<c≤π,
in which the cosine function is monotonically decreasing (i.e. there are no turning points),we can deduce that
cosc<sinc−sina
c−a<cosa.
J
59
PRELIMINARY CALCULUS
abf(x)
x
Figure 2.7 An integral as the area under a curve.
2.2 Integration
The notion of an integral as the area under a curve will be familiar to the reader.
In figure 2.7, in which the solid line is a plot of a function f(x), the shaded area
represents the quantity denoted by
I=integraldisplayb
af(x)dx. (2.21)
This expression is known as the definite integral off(x)b e t w e e nt h e lower limit
x=aand the upper limit x=b,a n d f(x) is called the integrand .
2.2.1 Integration from first principles
The definition of an integral as the area under a curve is not a formal definition,
but one that can be readily visualised. The formal definition of Iinvolves
subdividing the finite interval a≤x≤binto a large number of subintervals, by
defining intermediate points ξisuch that a=ξ0<ξ1<ξ2<···<ξ n=b,a n d
then forming the sum
S=nsummationdisplay
i=1f(xi)(ξi−ξi−1), (2.22)
where xiis an arbitrary point that lies in the range ξi−1≤xi≤ξi(see figure 2.8).
If now nis allowed to tend to infinity in any way whatsoever, subject only to the
restriction that the length of every subinterval ξi−1toξitends to zero, then S
might, or might not, tend to a unique limit, I. If it does then the definite integral
off(x) between aandbis defined as having the value I. If no unique limit exists
the integral is undefined. For continuous functions and a finite interval a≤x≤b
the existence of a unique limit is assured and the integral is guaranteed to exist.
60
2.2 INTEGRATION
abx1x2 x3 x4 x5 ξ1 ξ2 ξ3 ξ4f(x)
x
Figure 2.8 The evaluation of a definite integral by subdividing the interval
a≤x≤binto subintervals.IEvaluate from first principles the integral I=
Rb
0x2dx.
We first approximate the area under the curve y=x2between 0 and bbynrectangles of
equal width h. If we take the value at the lower end of each subinterval (in the limit of an
infinite number of subintervals we could equally well have chosen the value at the upperend) to give the height of the corresponding rectangle, then the area of the kth rectangle
will be ( kh)
2h=k2h3. The total area is thus
A=n−1X
k=0k2h3=(h3)1
6n(n−1)(2n−1),
where we have used the expression for the sum of the squares of the natural numbers
derived in subsection 1.7.1. Now h=b/nand so
A=
/b3
n3
/n
6(n−1)(2n−1) =b3
6
/
1−1
n
//
2−1
n
/
.
Asn→∞,A→b3/3, which is thus the value Iof the integral.
J
Some straightforward properties of definite integrals that are almost self-evident
are:
integraldisplayb
a0dx=0,integraldisplaya
af(x)dx=0, (2.23)
integraldisplayc
af(x)dx=integraldisplayb
af(x)dx+integraldisplayc
bf(x)dx, (2.24)
integraldisplayb
a[f(x)+g(x)]dx=integraldisplayb
af(x)dx+integraldisplayb
ag(x)dx. (2.25)
61
PRELIMINARY CALCULUS
Combining (2.23) and (2.24) with cset equal to ashows that
integraldisplayb
af(x)dx=−integraldisplaya
bf(x)dx. (2.26)
2.2.2 Integration as the inverse of differentiation
The definite integral has been defined as the area under a curve between two
fixed limits. Let us now consider the integral
F(x)=integraldisplayx
af(u)du (2.27)
in which the lower limit aremains fixed but the upper limit xis now variable. It
will be noticed that this is essentially a restatement of (2.21), but that the variable
xin the integrand has been replaced by a new variable u. It is conventional to
rename the dummy variable in the integrand in this way in order that the same
variable does not appear in both the integrand and the integration limits.
It is apparent from (2.27) that F(x) is a continuous function of x, but at first
glance the definition of an integral as the area under a curve does not connect withour assertion that integration is the inverse process to differentiation. However,
by considering the integral (2.27) and using the elementary property (2.24), we
obtain
F(x+∆x)=integraldisplay
x+∆x
af(u)du
=integraldisplayx
af(u)du+integraldisplayx+∆x
xf(u)du
=F(x)+integraldisplayx+∆x
xf(u)du.
Rearranging and dividing through by ∆ xyields
F(x+∆x)−F(x)
∆x=1
∆xintegraldisplayx+∆x
xf(u)du.
Letting ∆ x→0 and using (2.1) we find that the LHS becomes dF/dx ,w h e r e a s
the RHS becomes f(x). The latter conclusion follows because when ∆ xis small
the value of the integral on the RHS is approximately f(x)∆x, and in the limit
∆x→0 no approximation is involved. Thus
dF(x)
dx=f(x), (2.28)
or, substituting for F(x) from (2.27),
d
dxbracketleftbiggintegraldisplayx
af(u)dubracketrightbigg
=f(x).
62
2.2 INTEGRATION
From the last two equations it is clear that integration can be considered as
the inverse of differentiation. However, we see from the above analysis that thelower limit ais arbitrary and so differentiation does not have a unique inverse.
Any function F(x) obeying (2.28) is called an indefinite integral off(x), though
any two such functions can differ by at most an arbitrary additive constant. Since
the lower limit is arbitrary, it is usual to write
F(x)=integraldisplay
x
f(u)du (2.29)
and explicitly include the arbitrary constant only when evaluating F(x). The
evaluation is conventionally written in the form
integraldisplay
f(x)dx=F(x)+c (2.30)
where cis called the constant of integration . It will be noticed that, in the absence
of any integration limits, we use the same symbol for the arguments of both f
andF. This can be confusing, but is sufficiently common practice that the reader
needs to become familiar with it.
We also note that the definite integral of f(x) between the fixed limits x=a
andx=bc a nb ew r i t t e ni nt e r m so f F(x). From (2.27) we have
integraldisplayb
af(x)dx=integraldisplayb
x0f(x)dx−integraldisplaya
x0f(x)dx
=F(b)−F(a), (2.31)
where x0isanythird fixed point. Using the notation F/prime(x)=dF/dx ,w em a y
rewrite (2.28) as F/prime(x)=f(x), and so express (2.31) as
integraldisplayb
aF/prime(x)dx=F(b)−F(a)≡[F]b
a.
In contrast to differentiation, where repeated applications of the product rule
and/or the chain rule will always give the required derivative, it is not alwayspossible to find the integral of an arbitrary function. Indeed, in most real phys-ical problems exact integration cannot be performed and we have to revert tonumerical approximations. Despite this cautionary note, it is in fact possible tointegrate many simple functions and the following subsections introduce the most
common types. Many of the techniques will be familiar to the reader and so are
summarised by example.
2.2.3 Integration by inspection
The simplest method of integrating a function is by inspection. Some of the more
elementary functions have well-known integrals that should be remembered. Thereader will notice that these integrals are precisely the inverses of the derivatives
63
PRELIMINARY CALCULUS
found near the end of subsection 2.1.1. A few are presented below, using the form
given in (2.30).
integraldisplay
ad x=ax+c,integraldisplay
axndx=axn+1
n+1+c,
integraldisplay
eaxdx=eax
a+c,integraldisplaya
xdx=alnx+c,
integraldisplay
acosbx dx =asinbx
b+c,integraldisplay
asinbx dx =−acosbx
b+c,
integraldisplay
atanbx dx=−aln(cos bx)
b+c,integraldisplay
acosbxsinnbx dx=asinn+1bx
b(n+1 )+c,
integraldisplaya
a2+x2dx=t a n−1parenleftBigx
aparenrightBig
+c,integraldisplay
asinbxcosnbx dx =−acosn+1bx
b(n+1 )+c,
integraldisplay−1√
a2−x2dx=c o s−1parenleftBigx
aparenrightBig
+c,integraldisplay1√
a2−x2dx=s i n−1parenleftBigx
aparenrightBig
+c,
where the integrals that depend on nare valid for all n/negationslash=−1a n dw h e r e aandb
are constants. In the two final results |x|≤a.
2.2.4 Integration of sinusoidal functions
Integrals of the typeintegraltext
sinnxd xandintegraltext
cosnxd xmay be found by using trigono-
metric expansions. Two methods are applicable, one for odd nand the other for
even n. They are best illustrated by example.IEvaluate the integral I=
R
sin5xd x.
Rewriting the integral as a product of sin xa n da ne v e np o w e ro fs i n x, and then using
the relation sin2x=1−cos2xyields
I=
Z
sin4xsinxd x
=
Z
(1−cos2x)2sinxd x
=
Z
(1−2cos2x+c o s4x)sinxd x
=
Z
(sinx−2si nxcos2x+s i n xcos4x)dx
=−cosx+2
3cos3x−1
5cos5x+c,
where the integration has been carried out using the results of subsection 2.2.3.
J
64
2.2 INTEGRATIONIEvaluate the integral I=
R
cos4xd x.
Rewriting the integral as a power of cos2xand then using the double-angle formula
cos2x=1
2(1 + cos2 x) yields
I=
Z
(cos2x)2dx=
Z
/1+c o s2 x
2
/2
dx
=
Z
1
4(1 + 2 cos2 x+c o s22x)dx.
Using the double-angle formula again we may write cos22x=1
2(1 + cos 4 x), and hence
I=
Z/1
4+1
2cos 2x+1
8(1 + cos 4 x)
/
dx
=1
4x+1
4sin 2x+1
8x+1
32sin4x+c
=3
8x+1
4sin 2x+1
32sin 4x+c.
J
2.2.5 Logarithmic integration
Integrals for which the integrand may be written as a fraction in which the
numerator is the derivative of the denominator may be evaluated using
integraldisplayf/prime(x)
f(x)dx=l nf(x)+c. (2.32)
This follows directly from the differentiation of a logarithm as a function of a
function (see subsection 2.1.3).IEvaluate the integral
I=
Z6x2+2c o s x
x3+s i n xdx.
We note first that the numerator can be factorised to give 2(3 x2+c o s x) ,a n dt h e nt h a t
the quantity in brackets is the derivative of the denominator. Hence
I=2
Z3x2+c o s x
x3+s i n xdx=2l n ( x3+s i n x)+c.
J
2.2.6 Integration using partial fractions
The method of partial fractions was discussed at some length in section 1.4, but
in essence consists of the manipulation of a fraction (here the integrand) in such
a way that it can be written as the sum of two or more simpler fractions. Againwe illustrate the method by an example.
65
PRELIMINARY CALCULUSIEvaluate the integral
I=
Z1
x2+xdx.
We note that the denominator factorises to give x(x+1 ) .H e n c e
I=
Z1
x(x+1 )dx.
We now separate the fraction into two partial fractions and integrate directly:
I=
Z
/1
x−1
x+1
/
dx=l nx−ln(x+1 )+ c=l n
/x
x+1
/
+c.
J
2.2.7 Integration by substitution
Sometimes it is possible to make a substitution of variables that turns a com-
plicated integral into a simpler one, which can then be integrated by a standard
method. There are many useful substitutions and knowing which to use is a matter
of experience. We now present a few examples of particularly useful substitutions.IEvaluate the integral
I=
Z1√
1−x2dx.
Making the substitution x=s i n u, we note that dx=c o s ud u, and hence
I=
Z1√
1−sin2ucosud u=
Z1√
cos2ucosud u=
Z
du=u+c.
Now substituting back for u,
I=s i n−1x+c.
This corresponds to one of the results given in subsection 2.2.3.
J
Another particular example of integration by substitution is afforded by inte-
grals of the form
I=integraldisplay1
a+bcosxdx or I=integraldisplay1
a+bsinxdx. (2.33)
In these cases making the substitution t=t a n ( x/2) yields integrals that can be
solved more easily than the originals. Formulae expressing sin xand cos xin
terms of twere derived in equations (1.32) and (1.33) [see p. 14], but before we
can use them we must relate dxtodtas follows.
66
2.2 INTEGRATION
Since
dt
dx=1
2sec2x
2=1
2parenleftBig
1+t a n2x
2parenrightBig
=1+t2
2,
the required relationship is
dx=2
1+t2dt. (2.34)IEvaluate the integral
I=
Z2
1+3c o s xdx.
Rewriting cos xin terms of tand using (2.34) yields
I=
Z2
1+3
/
(1−t2)(1 + t2)−1
/
/2
1+t2
/
dt
=
Z2(1 + t2)
1+t2+3 ( 1−t2)
/2
1+t2
/
dt
=
Z2
2−t2dt=
Z2
(√
2−t)(√
2+t)dt
=
Z1√
2
/1√
2−t+1√
2+t
/
dt
=−1√
2ln(√
2−t)+1√
2ln(√
2+t)+c
=1√
2ln
/"√
2+t a n( x/2)√
2−tan (x/2)
/#
+c.
J
Integrals of a similar form to (2.33), but involving sin 2 x,c o s2 x,t a n 2 x,s i n2x,
cos2xor tan2xinstead of cos xand sin x, should be evaluated by using the
substitution t=t a n x.I nt h i sc a s e
sinx=t√
1+t2,cosx=1√
1+t2and dx=dt
1+t2.(2.35)
A final example of the evaluation of integrals using substitution is the method
of completing the square (cf. subsection 1.7.3).
67
PRELIMINARY CALCULUSIEvaluate the integral
I=
Z1
x2+4x+7dx.
We can write the integral in the form
I=
Z1
(x+2 )2+3dx.
Substituting y=x+ 2, we find dy=dxand hence
I=
Z1
y2+3dy,
Hence, by comparison with the table of standard integrals (see subsection 2.2.3)
I=√
3
3tan−1
/y√
3
/
+c=√
3
3tan−1
/x+2√
3
/
+c.
J
2.2.8 Integration by parts
Integration by parts is the integration analogy of product differentiation. The
principle is to break down a complicated function into two functions, at least one
of which can be integrated by inspection. The method in fact relies on the resultfor the differentiation of a product. Recalling from (2.6) that
d
dx(uv)=udv
dx+du
dxv,
where uandvare functions of x, we now integrate to find
uv=integraldisplay
udv
dxdx+integraldisplaydu
dxvd x .
Rearranging into the standard form for integration by parts gives
integraldisplay
udv
dxdx=uv−integraldisplaydu
dxvd x . (2.36)
Integration by parts is often remembered for practical purposes in the form
the integral of a product of two functions is equal to {the first times the integral of
the second}minus the integral of {the derivative of the first times the integral of
the second}. Here, uis ‘the first’ and dv/dx is ‘the second’; clearly the integral v
of ‘the second’ must be determinable by inspection.IEvaluate the integral I=
R
xsinxdx.
In the notation given above, we identify xwith uand sin xwith dv/dx. Hence v=−cosx
anddu/dx = 1 and so using (2.36)
I=x(−cosx)−
Z
(1)(−cosx)dx=−xcosx+s i n x+c.
J
68
2.2 INTEGRATION
The separation of the functions is not always so apparent, as is illustrated by
the following example.IEvaluate the integral I=
R
x3e−x2dx.
Firstly we rewrite the integral as
I=
Z
x2
/
xe−x2
/
dx.
Now, using the notation given above, we identify x2with uandxe−x2with dv/dx. Hence
v=−1
2e−x2anddu/dx =2x,s ot h a t
I=−1
2x2e−x2−
Z
(−x)e−x2dx=−1
2x2e−x2−1
2e−x2+c.
J
A trick that is sometimes useful is to take ‘1’ as one factor of the product, as
is illustrated by the following example.IEvaluate the integral I=
R
lnxd x.
Firstly we rewrite the integral as
I=
Z
(lnx)1dx.
Now, using the notation above, we identify ln xwith uand 1 with dv/dx. Hence we have
v=xanddu/dx =1/x,a n ds o
I=( l n x)(x)−
Z
/1
x
/
xd x=xlnx−x+c.
J
It is sometimes necessary to integrate by parts more than once. In doing so,
we may occasionally re-encounter the original integral I.I ns u c hc a s e sw ec a n
obtain a linear algebraic equation for Ithat can be solved to obtain its value.IEvaluate the integral I=
R
eaxcosbx dx.
Integrating by parts, taking eaxas the first function, we find
I=eax
/sinbx
b
/
−
Z
aeax
/sinbx
b
/
dx,
where, for convenience, we have omitted the constant of integration. Integrating by parts
a second time,
I=eax
/sinbx
b
/
−aeax
/−cosbx
b2
/
+
Z
a2eax
/−cosbx
b2
/
dx.
Notice that the integral on the RHS is just −a2/b2times the original integral I. Thus
I=eax
/1
bsinbx+a
b2cosbx
/
−a2
b2I.
69
PRELIMINARY CALCULUS
Rearranging this expression to obtain Iexplicitly and including the constant of integration
we find
I=eax
a2+b2(bsinbx+acosbx)+c. (2.37)
Another method of evaluating this integral, using the exponential of a complex number,
is given in section 3.6.
J
2.2.9 Reduction formulae
Integration using reduction formulae is a process that involves first evaluating a
simple integral and then, in stages, using it to find a more complicated integral.IUsing integration by parts, find a relationship between InandIn−1where
In=
Z1
0(1−x3)ndx
andnis any positive integer. Hence evaluate I2=
R1
0(1−x3)2dx.
Writing the integrand as a product and separating the integral into two we find
In=
Z1
0(1−x3)(1−x3)n−1dx
=
Z1
0(1−x3)n−1dx−
Z1
0x3(1−x3)n−1dx.
The first term on the RHS is clearly In−1and so, writing the integrand in the second term
on the RHS as a product,
In=In−1−
Z1
0(x)x2(1−x3)n−1dx.
Integrating by parts we find
In=In−1+
hx
3n(1−x3)n
i1
0−
Z1
01
3n(1−x3)ndx
=In−1+0−1
3nIn,
which on rearranging gives
In=3n
3n+1In−1.
We now have a relation connecting successive integrals. Hence, if we can evaluate I0,w e
can find I1,I2etc. Evaluating I0is trivial:
I0=
Z1
0(1−x3)0dx=
Z1
0dx=[x]1
0=1.
Hence
I1=(3×1)
(3×1) + 1×1=3
4,I 2=(3×2)
(3×2) + 1×3
4=9
14.
Although the first few Incould be evaluated by direct multiplication, this becomes tedious
for integrals containing higher values of n; these are therefore best evaluated using the
reduction formula.
J
70
2.2 INTEGRATION
2.2.10 Infinite and improper integrals
The definition of an integral given previously does not allow for cases in which
either of the limits of integration is infinite (an infinite integral )o rf o rc a s e s
in which f(x) is infinite in some part of the range (an improper integral ), e.g.
f(x)=( 2−x)−1/4near the point x= 2. Nevertheless, modification of the
definition of an integral gives infinite and improper integrals each a meaning.
In the case of an integral I=integraltextb
af(x)dx, the infinite integral, in which btends
to∞, is defined by
I=integraldisplay∞
af(x)dx= lim
b→∞integraldisplayb
af(x)dx= lim
b→∞F(b)−F(a).
As previously, F(x) is the indefinite integral of f(x) and lim b→∞F(b)m e a n st h e
limit (or value) that F(b) approaches as b→∞;i ti se v a l u a t e d aftercalculating
the integral. The formal concept of a limit will be introduced in chapter 4.IEvaluate the integral
I=
Z∞
0x
(x2+a2)2dx.
Integrating, we find F(x)=−1
2(x2+a2)−1+cand so
I= lim
b→∞
/−1
2(b2+a2)
/
−
/−1
2a2
/
=1
2a2.
J
For the case of improper integrals, we adopt the approach of excluding the
unbounded range from the integral. For example, if the integrand f(x) is infinite
atx=c(say), a≤c≤bthen
integraldisplayb
af(x)dx= lim
δ→0integraldisplayc−δ
af(x)dx+ lim
/epsilon1→0integraldisplayb
c+/epsilon1f(x)dx.IEvaluate the integral I=
R2
0(2−x)−1/4dx.
Integrating directly,
I= lim
/epsilon1→0
/
−4
3(2−x)3/4
/2−/epsilon1
0= lim
/epsilon1→0
/
−4
3/epsilon13/4
/
+4
323/4=
/;4
3
/
23/4.
J
2.2.11 Integration in plane polar coordinates
In plane polar coordinates ρ, φ, a curve is defined by its distance ρfrom the
origin as a function of the angle φbetween the line joining a point on the curve
to the origin and the x-axis, i.e. ρ=ρ(φ). The area of an element is given by
71
PRELIMINARY CALCULUS
dA
ρ(φ)ρ(φ+dφ)ρd φ
xy
OBC
Figure 2.9 Finding the area of a sector OBC defined by the curve ρ(φ)a n d
the radii OB,OC, at angles to the x-axis φ1,φ2respectively.
dA=1
2ρ2dφ, as illustrated in figure 2.9, and hence the total area between two
angles φ1andφ2is given by
A=integraldisplayφ2
φ11
2ρ2dφ. (2.38)
An immediate observation is that the area of a circle of radius ais given by
A=integraldisplay2π
01
2a2dφ=bracketleftbig1
2a2φbracketrightbig2π
0=πa2.IThe equation in polar coordinates of an ellipse with semi-axes aandbis
1
ρ2=cos2φ
a2+sin2φ
b2.
Find the area Aof the ellipse.
Using (2.38) and symmetry, we have
A=1
2
Z2π
0a2b2
b2cos2φ+a2sin2φdφ=2a2b2
Zπ/2
01
b2cos2φ+a2sin2φdφ.
To evaluate this integral we write t=t a n φand use (2.35):
A=2a2b2
Z∞
01
b2+a2t2dt=2b2
Z∞
01
(b/a)2+t2dt.
Finally, from the list of standard integrals (see subsection 2.2.3),
A=2b2
/1
(b/a)tan−1t
(b/a)
/∞
0=2ab
/π
2−0
/
=πab.
J
72
2.2 INTEGRATION
2.2.12 Integral inequalities
Consider the functions f(x),φ1(x)a n d φ2(x) such that φ1(x)≤f(x)≤φ2(x)f o r
allxin the range a≤x≤b. It immediately follows that
integraldisplayb
aφ1(x)dx≤integraldisplayb
af(x)dx≤integraldisplayb
aφ2(x)dx, (2.39)
which gives us a way of estimating an integral that is difficult to evaluate explicitly.IShow that the value of the integral
I=
Z1
01
(1 +x2+x3)1/2dx
lies between 0.810and0.882.
We note that for xin the range 0 ≤x≤1, 0≤x3≤x2. Hence
(1 +x2)1/2≤(1 +x2+x3)1/2≤(1 + 2 x2)1/2,
and so
1
(1 +x2)1/2≥1
(1 +x2+x3)1/2≥1
(1 + 2 x2)1/2.
Consequently,Z1
01
(1 +x2)1/2dx≥
Z1
01
(1 +x2+x3)1/2dx≥
Z1
01
(1 + 2 x2)1/2dx,
from which we obtainh
ln(x+
p
1+x2)
i1
0≥I≥
/
1√
2ln
/
x+
q
1
2+x2
//1
0
0.8814≥I≥0.8105
0.882≥I≥0.810.
In the last line the calculated values have been rounded to three significant figures,
one rounded up and the other rounded down so that the proved inequality cannot beunknowingly made invalid.J
2.2.13 Applications of integration
Mean value of a function
The mean value mof a function between two limits aandbis defined by
m=1
b−aintegraldisplayb
af(x)dx. (2.40)
The mean value may be thought of as the height of the rectangle that has the
same area (over the same interval) as the area under the curve f(x). This is
illustrated in figure 2.10.
73
PRELIMINARY CALCULUS
mf(x)
b x a
Figure 2.10 The mean value mof a function.IFind the mean value mof the function f(x)=x2between the limits x=2andx=4.
Using (2.40),
m=1
4−2
Z4
2x2dx=1
2
/x3
3
/4
2=1
2
/43
3−23
3
/
=28
3.
J
Finding the length of a curve
Finding the area between a curve and certain straight lines provides one example
of the use of integration. Another is in finding the length of a curve. If a curveis defined by y=f(x) then the distance along the curve, ∆ s, that corresponds to
small changes ∆ xand ∆ yinxandyis given by
∆s≈radicalbig
(∆x)2+( ∆y)2; (2.41)
this follows directly from Pythagoras’ theorem (see figure 2.11). Dividing (2.41)
through by ∆ xand letting ∆ x→0w eo b t a i n †
ds
dx=radicalBigg
1+parenleftbiggdy
dxparenrightbigg2
.
Clearly the total length sof the curve between the points x=aandx=bis then
given by integrating both sides of the equation:
s=integraldisplayb
aradicalBigg
1+parenleftbiggdy
dxparenrightbigg2
dx. (2.42)
†Instead of considering small changes ∆ xand ∆ yand letting these tend to zero, we could have
derived (2.41) by considering infinitesimal changes dxanddyfrom the start. After writing ( ds)2=
(dx)2+(dy)2, (2.41) may be deduced by using the formal device of dividing through by dx. Although
not mathematically rigorous, this method is often used and generally leads to the correct result.
74
2.2 INTEGRATION
∆x∆y∆sy=f(x)
xf(x)
Figure 2.11 The distance moved along a curve, ∆ s, corresponding to the
small changes ∆ xand ∆ y.
In plane polar coordinates,
ds=radicalbig
(dr)2+(rd φ)2⇒ s=integraldisplayr2
r1radicalBigg
1+r2parenleftbiggdφ
drparenrightbigg2
dr.
(2.43)IFind the length of the curve y=x3/2from x=0tox=2.
Using (2.42) and noting that dy/dx =3
2√x, the length sof the curve is given by
s=
Z2
0
q
1+9
4xd x
=
h
2
3
/;4
9
//;
1+9
4x
/3/2
i2
0=8
27
h/;
1+9
4x
/3/2
i2
0
=8
27
h/;11
2
/3/2−1
i
.
J
Surfaces of revolution
Consider the surface Sformed by rotating the curve y=f(x) about the x-axis
(see figure 2.12). The surface area of the ‘collar’ formed by rotating an element
of the curve, ds, about the x- a x i si s2 πy ds, and hence the total surface area is
S=integraldisplayb
a2πy ds.
Since ( ds)2=(dx)2+(dy)2from (2.41), the total surface area between the planes
x=aandx=bis
S=integraldisplayb
a2πyradicalBigg
1+parenleftbiggdy
dxparenrightbigg2
dx. (2.44)
75
PRELIMINARY CALCULUS
Sy
bVf(x)
dx
x ads
Figure 2.12 The surface and volume of revolution for the curve y=f(x).IFind the surface area of a cone formed by rotating about the x-axis the line y=2x
between x=0andx=h.
Using (2.44), the surface area is given by
S=
Zh
0(2π)2x
s
1+
/d
dx(2x)
/2
dx
=
Zh
04πx
/;
1+22
/1/2dx=
Zh
04√
5πxdx
=
h
2√
5πx2
ih
0=2√
5π(h2−0) = 2√
5πh2.
J
We note that a surface of revolution may also be formed by rotating a line
about the y-axis. In this case the surface area between y=aandy=bis
S=integraldisplayb
a2πxradicalBigg
1+parenleftbiggdx
dyparenrightbigg2
dy. (2.45)
Volumes of revolution
The volume Venclosed by rotating the curve y=f(x) about the x-axis can also
be found (see figure 2.12). The volume of the disc between xandx+dxis given
bydV=πy2dx.Hence the total volume between x=aandx=bis
V=integraldisplayb
aπy2dx. (2.46)
76
2.3 EXERCISESIFind the volume of a cone enclosed by the surface formed by rotating about the x-axis
the line y=2xbetween x=0andx=h.
Using (2.46), the volume is given by
V=
Zh
0π(2x)2dx=
Zh
04πx2dx
=
/4
3πx3
/h
0=4
3π(h3−0) =4
3πh3.
J
As before, it is also possible to form a volume of revolution by rotating a curve
about the y-axis. In this case the volume enclosed between y=aandy=bis
V=integraldisplayb
aπx2dy. (2.47)
2.3 Exercises
2.1 Obtain the following derivatives from first principles:
(a) the first derivative of 3 x+4 ;
(b) the first, second and third derivatives of x2+x;
(c) the first derivative of sin x.
2.2 Find from first principles the first derivative of ( x+3)2and compare your answer
with that obtained using the chain rule.
2.3 Find the first derivatives of
(a)x2expx,( b )2s i n xcosx,( c )s i n2 x,( d )xsinax,
(e) (exp ax)(sinax)tan−1ax,( f )l n ( xa+x−a),
(g) ln( ax+a−x), (h) xx.
2.4 Find the first derivatives of
(a)x/(a+x)2,( b ) x/(1−x)1/2,( c )t a n x,a ss i n x/cosx,
(d) (3 x2+2x+1 )/(8x2−4x+2 ) .
2.5 Use result (2.12) to find the first derivatives of
(a) (2 x+3 )−3,( b )s e c2x, (c) cosech33x,( d )1 /lnx,( e )1 /[sin−1(x/a)].
2.6 Show that the function y(x)=e x p (−|x|) defined by
y(x)=
/8/>/</>/:expx forx<0,
1f o r x=0,
exp(−x)f o r x>0,
isnotdifferentiable at x= 0. Consider the limiting process for both ∆ x>0a n d
∆x<0.
2.7 Find dy/dx ifx=(t−2)/(t+2 )a n d y=2t/(t+1 )f o r−∞<t<∞. Show that
it is always non-negative, and make use of this result in sketching the curve of y
as a function of x.
2.8 If 2 y+s i n y+5= x4+4x3+2π, show that dy/dx =1 6w h e n x=1 .
2.9 Find the second derivative of y(x)=c o s [ ( π/2)−ax]. Now set a=1a n dv e r i f y
that the result is the same as that obtained by first setting a= 1 and simplifying
y(x) before differentiating.
77
PRELIMINARY CALCULUS
2.10 The function y(x) is defined by y(x)=( 1+ xm)n.
(a) Use the chain rule to show that the first derivative of yisnmxm−1(1 +xm)n−1.
(b) The binomial expansion (see section 1.5) of (1 + z)nis
(1 +z)n=1+ nz+n(n−1)
2!z2+···+n(n−1)···(n−r+1 )
r!zr+···.
Keeping only the terms of zeroth and first order in dx, apply this result twice
to derive result (a) from first principles.
(c) Expand yin a series of powers of xbefore differentiating term by term.
Show that the result is the series obtained by expanding the answer givenfordy/dx in (a).
2.11 Show by differentiation and substi tution that the differential equation
4x
2d2y
dx2−4xdy
dx+( 4x2+3 )y=0
has a solution of the form y(x)=xnsinx, and find the value of n.
2.12 Find the positions and natures of the stationary points of the following functions:
(a)x3−3x+3 ;( b ) x3−3x2+3x;( c )x3+3x+3 ;
(d) sin axwith a/negationslash=0 ;( e ) x5+x3;( f )x5−x3.
2.13 Show that the lowest value taken by the function 3 x4+4x3−12x2+6i s−26.
2.14 By finding their stationary points and examining their general forms, determine
the range of values that each of the following functions y(x) can take. In each
case make a sketch-graph incorporating the features you have identified.
(a)y(x)=(x−1)/(x2+2x+6 ) .
(b)y(x)=1 /(4 + 3 x−x2).
(c)y(x)=( 8s i n x)/(15 + 8 tan2x).
2.15 Show that y(x)=xa2xexpx2has no stationary points other than x=0 ,i f
exp(−√
2)<a< exp(√
2).
2.16 The curve 4 y3=a2(x+3y) can be parameterised as x=acos 3θ,y=acosθ.
(a) Obtain expressions for dy/dx (i) by implicit differentiation and (ii) in param-
eterised form. Verify that they are equivalent.
(b) Show that the only point of inflection occurs at the origin. Is it a stationary
point of inflection?
(c) Use the information gained in (a) and (b) to sketch the curve, paying
particular attention to its shape near the points ( −a, a/2) and ( a,−a/2) and
to its slope at the ‘end points’ ( a, a)a n d(−a,−a).
2.17 The parametric equations for the motion of a charged particle released from rest
in electric and magnetic fields at right angles to each other take the forms
x=a(θ−sinθ),y =a(1−cosθ).
Show that the tangent to the curve has slope cot( θ/2). Use this result at a few
calculated values of xandyto sketch the form of the particle’s trajectory.
2.18 Show that the maximum curvature on the catenary y(x)=acosh( x/a)i s1/a.Y o u
will need some of the results about hyperbolic functions stated in subsection 3.7.6.
2.19 The curve whose equation is x2/3+y2/3=a2/3for positive xandyand which
is completed by its symmetric reflections in both axes is known as an astroid.Sketch it and show that its radius of curvature in the first quadrant is 3( axy)
1/3.
2.20 A two-dimensional coordinate system useful for orbit problems is the tangential-
polar coordinate system (figure 2.13). In this system a curve is defined by r,t h e
distance from a fixed point Oto a general point Pof the curve, and p,t h e
78
2.3 EXERCISES
OC
PQρ
ρ
rr+∆rc
pp+∆p
Figure 2.13 The coordinate system described in exercise 2.20.
perpendicular distance from Oto the tangent to the curve at P. By proceeding
as indicated below, show that the radius of curvature at Pc a nb ew r i t t e ni nt h e
form ρ=r dr/dp .
Consider two neighbouring points PandQon the curve. The normals to the
curve through those points meet at C, with (in the limit Q→P)CP=CQ=ρ.
Apply the cosine rule to triangles OPC andOQC to obtain two expressions for
c2,o n ei nt e r m so f randpand the other in terms of r+∆randp+∆p.B y
equating them and letting Q→Pdeduce the stated result.
2.21 Use Leibniz’ theorem to find
(a) the second derivative of cos xsin2x,
(b) the third derivative of sin xlnx,
(c) the fourth derivative of (2 x3+3x2+x+2 )e x p2 x.
2.22 If y=e x p (−x2), show that dy/dx =−2xyand hence, by applying Leibniz’
theorem, prove that for n≥1
y(n+1)+2xy(n)+2ny(n−1)=0.
2.23 (a) By considering its properties near x= 1, show that f(x)=5 x4−11x3+
26x2−44x+ 24 takes negative values for some range of x.
(b) Show that f(x)=t a n x−xcannot be negative for 0 ≤x≤π/2, and deduce
thatg(x)=x−1sinxdecreases monotonically in the same range.
2.24 Determine what can be learned from applying Rolle’s theorem to the following
functions f(x): (a) ex;( b ) x2+6x;( c )2 x2+3x+1 ; ( d ) 2 x2+3x+2 ; ( e )
2x3−21x2+6 0x+k.( f )I f k=−45 in (e), show that x=3i so n er o o to f
f(x) = 0, find the other roots, and verify that the conclusions from (e) are
satisfied.
2.25 By applying Rolle’s theorem to xnsinnx,w h e r e nis an arbitrary positive integer,
show that tan nx+x=0h a sas o l u t i o n α1with 0 <α 1<π / n . Apply the
theorem a second time to obtain the non sensical result that there is a real α2in
0<α2<π / n , such that cos2(nα2)=−n2. Explain why this incorrect result arises.
2.26 Use the mean value theorem to establish bounds
(a) for−ln(1−y), by considering ln xin the range 0 <1−y<x< 1,
(b) for ey−1, by considering ex−1 in the range 0 <x<y .
79
PRELIMINARY CALCULUS
2.27 For the function y(x)=x2exp(−x) obtain a simple relationship between yand
dy/dx and then, by applying Leibniz’ theorem, prove that
xy(n+1)+(n+x−2)y(n)+ny(n−1)=0.
2.28 Use Rolle’s theorem to deduce that if the equation f(x) = 0 has a repeated root
x1then x1is also a root of the equation f/prime(x)=0 .
(a) Apply this result to the ‘standard’ quadratic equation ax2+bx+c=0 ,t o
show that the condition for equal roots is b2=4ac.
(b) Find all the roots of f(x)=x3+4x2−3x−18 = 0, given that one of them
is a repeated root.
(c) The equation f(x)=x4+4x3+7x2+6x+2 = 0 has a repeated integer root.
How many real roots does it have altogether?
2.29 Show that the curve x3+y3−12x−8y−16 = 0 touches the x-axis.
2.30 Find the following indefinite integrals:
(a)
R
(4 +x2)−1dx;( b )
R
(8 + 2 x−x2)−1/2dxfor 2≤x≤4;
(c)
R
(1 + sin θ)−1dθ;( d )
R
(x√
1−x)−1dxfor 0 <x≤1.
2.31 Find the indefinite integrals Jof the following ratios of polynomials:
(a) ( x+3 )/(x2+x−2);
(b) ( x3+5x2+8x+ 12) /(2x2+1 0x+ 12);
(c) (3 x2+2 0x+ 28) /(x2+6x+9 ) ;
(d)x3/(a8+x8).
2.32 Express x2(ax+b)−1as the sum of powers of xand another integrable term, and
hence evaluateZb/a
0x2
ax+bdx.
2.33 Find the integral Jof (ax2+bx+c)−1,w i t h a/negationslash= 0, distinguishing between the
cases (i) b2>4ac, (ii)b2<4ac, and (iii) b2=4ac.
2.34 Use logarithmic integration to find the indefinite integrals Jof the following:
(a) sin2 x/(1 + 4 sin2x);
(b)ex/(ex−e−x);
(c) (1 + xlnx)/(xlnx);
(d) [ x(xn+an)]−1.
2.35 Find the derivative of f(x)=( 1+s i n x)/cosxand hence determine the indefinite
integral Jof sec x.
2.36 Find the indefinite integrals Jof the following functions involving sinusoids:
(a) cos5x−cos3x;
(b) (1−cosx)/(1 + cos x);
(c) cos xsinx/(1 + cos x);
(d) sec2x/(1−tan2x).
2.37 By making the substitution x=acos2θ+bsin2θ, evaluate the de finite integrals
Jbetween limits aandb(>a) of the following functions:
(a) [( x−a)(b−x)]−1/2;
(b) [( x−a)(b−x)]1/2;
(c) [( x−a)/(b−x)]1/2.
80
2.3 EXERCISES
2.38 Determine whether the following integrals exist and, where they do, evaluate
them:
(a)
Z∞
0exp(−λx)dx;( b )
Z∞
−∞x
(x2+a2)2dx;
(c)
Z∞
11
x+1dx;( d )
Z1
01
x2dx;
(e)
Zπ/2
0cotθd θ;( f )
Z1
0x
(1−x2)1/2dx.
2.39 Use integration by parts to evaluate the following:
(a)
Zy
0x2sinxd x;( b )
Zy
1xlnxd x;
(c)
Zy
0sin−1xd x;( d )
Zy
1ln(a2+x2)/x2dx.
2.40 Show, by each of the following methods, that the indefinite integral Jofx3/(x+
1)1/2is
J=2
35(5x3−6x2+8x−16)(x+1 )1/2+c.
(a) by using repeated integration by parts.
(b) by setting x+1= u2and determining dJ/du as (dJ/dx )(dx/du ).
2.41 The gamma function Γ( n) is defined for all n>−1b y
Γ(n+1 )=
Z∞
0xne−xdx.
Find a recurrence relation connecting Γ( n+1 )a n dΓ ( n).
(a) Deduce (i) the value of Γ( n+1)when nis a non-negative integer and (ii) the
value of Γ
/;7
2
/
,g i v e nt h a tΓ
/;1
2
/
=√π.
(b) Now, taking factorial mforanymto be defined by m!=Γ ( m+ 1), evaluate/;
−3
2
/
!.
2.42 Define J(m, n), for non-negative integers mandn, by the integral
J(m, n)=
Zπ/2
0cosmθsinnθd θ .
(a) Evaluate J(0,0),J(0,1),J(1,0),J(1,1),J(m,1),J(1,n).
(b) Using integration by parts prove that, for mandnboth >0,
J(m, n)=m−1
m+nJ(m−2,n)a n d J(m, n)=n−1
m+nJ(m, n−2).
(c) Evaluate (i) J(5,3), (ii) J(6,5), (iii) J(4,8).
2.43 By integrating by parts twice, prove that Inas defined in the first equality below
for positive integers nhas the value given in the second equality.
In=
Zπ/2
0sinnθcosθd θ=n−sin(nπ/2)
n2−1.
2.44 Evaluate the following definite integrals:
(a)
R∞
0xe−xdx;( b )
R1
0
/
(x3+1 )/(x4+4x+1 )
/
dx;
(c)
Rπ/2
0[a+(a−1)cos θ]−1dθwith a>1
2;( d )
R∞
−∞(x2+6x+ 18)−1dx.
81
PRELIMINARY CALCULUS
2.45 If Jris the integralZ∞
0xrexp(−x2)dx
show that
(a)J2r+1=(r!)/2,
(b)J2r=2−r(2r−1)(2r−3)···(5)(3)(1) J0.
2.46 (a) Find positive constants a,bsuch that ax≤sinx≤bxfor 0≤x≤π/2. Use
this inequality to find (to two significant figures) upper and lower boundsfor the integral
I=
Zπ/2
0(1 + sin x)1/2dx.
(b) Use the substitution t=t a n ( x/2) to evaluate Iexactly.
2.47 By noting that for 0 ≤η≤1,η1/2≥η3/4≥η, prove that
2
3≤1
a5/2
Za
0(a2−x2)3/4dx≤π
4.
2.48 Show that the total length of the astroid x2/3+y2/3=a2/3,w h i c hc a nb e
parameterised as x=acos3θ,y=asin3θ,i s6a.
2.49 By noting that sinh x<1
2ex<coshx,a n dt h a t1+ z2<(1 +z)2forz>0, show
that for x>0, the length Lof the curve y=1
2exmeasured from the origin
satisfies the inequalities sinh x<L<x +s i n h x.
2.50 The equation of a cardioid in plane polar coordinates is
ρ=a(1−sinφ).
Sketch the curve and find (i) its area, (ii) its total length, (iii) the surface area of
the solid formed by rotating the cardioi d about its axis of symmetry and (iv) the
volume of the same solid.
2.4 Hints and answers
2.1 (a) 3; (b) 2 x+ 1, 2, 0; (c) cos x.
2.2 2 x+6 .
2.3 (a) ( x2+2x)exp x;( b )2 ( c o s2x−sin2x)=2 c o s 2 x;( c )2 c o s 2 x;( d )s i n ax+
axcosax;
(e) (aexpax)[(sin ax+c o s ax)tan−1ax+( s i n ax)(1 + a2x2)−1];
(f) [a(xa−x−a)]/[x(xa+x−a)]; (g) [( ax−a−x)lna]/(ax+a−x); (h) (1 + ln x)xx.
2.4 (a) ( a−x)(a+x)−3;( b )( 1−x/2)(1−x)−3/2;( c )s e c2x;
(d) (−7x2−x+ 2)(4 x2−2x+1 )−2.
2.5 (a) −6(2x+3 )−4;( b )2 s e c2xtanx;( c )−9cosec h33xcoth3 x;
(d)−x−1(lnx)−2;( e )−(a2−x2)−1/2[sin−1(x/a)]−2.
2.6 The two limits are −1( f o r∆ x>0) and +1 (for ∆ x<0) and are not equal.
2.7 ( t+2 )2/[2(t+1 )2].
2.8 y=πatx=1 .
2.9−sinxin both cases.
2.10 (b) Write 1 + ( x+∆x)mas 1 + xm(1 + ∆ x/x)m; (c) in the general terms of the two
series, the indices randsare related by r=s±1.
2.11 The required conditions are 8 n−4=0a n d4 n2−8n+ 3 = 0; both are satisfied
byn=1
2.
82
2.4 HINTS AND ANSWERS
−15−10−5 5 10 15
−0.4−0.20.20.4
(a)−3−2−11 2 3456
−0.8−0.40.40.8
(b)
0
−0.20.2
π 2π 3π
(c)
Figure 2.14 The solutions to exercise 2.14.
2.12 (a) Minimum at x= 1, maximum at x=−1; (b) inflection at x=1 ;( c )n o
stationary points; (d) x=(n+1
2)π/a, maximum for neven, minimum for nodd;
(e) inflection at x= 0; (f) inflection at x= 0, maximum at x=−(3
5)1/2, minimum
at (3
5)1/2.
2.13−26 at x=−2; other stationary values are 6 at x= 0 and 1 at x=1 .
2.14 See figure 2.14(a)–(c).
(a)y(1) = 0; no infinities; minimum y(−2) =−1
2, maximum y(4) =1
10;−1
2≤
y≤1
10.
(b) No zeroes; y(−1) =±∞,y(4) =±∞; minimum y(3
2)=4
25;y<0o ry≥4
25.
(c) Periodic with period 2 π.W i t h i n0 ≤x≤π, symmetry about x=π/2.
Within 0≤x≤2π, antisymmetry about x=π; zeroes at x=nπand
x=( 2m+1)π/2; no infinities; other stationary points at x=c o s−1(±2/√
7);
|y|≤8/(7√
21).
2.15 Use logarithmic differentiation. Set dy/dx = 0, obtaining 2 x2+2xlna+1=0 .
2.16 (a) (i) a2/(12y2−3a2), (ii) (12 cos2θ−3)−1.( b )N o , dy/dx =−1/3. (c) Vertical
tangents when y=±a/2;dy/dx =1/9a ty=±a.
2.17 See figure 2.15.2.18 First show that ρ=y
2/a.
2.19dy
dx=−
/y
x
/1/3
;d2y
dx2=a2/3
3x4/3y1/3.
2.20 For example, OC2=ρ2+r2−2pρ, where use has been made of the fact that
rcosOPC =p.
2.21 (a) 2(2 −9cos2x)sinx;( b )( 2 x−3−3x−1)sinx−(3x−2+l nx)c osx;( c )8 ( 4 x3+
30x2+6 2x+ 38)exp2 x.
83
PRELIMINARY CALCULUS
πa 2πa2a
xy
Figure 2.15 The solution to exercise 2.17.
2.23 (a) f(1) = 0 whilst f/prime(1)/negationslash=0a n ds o f(x)m u s tb en e g a t i v ei ns o m er e g i o nw i t h
x= 1 as an endpoint.
(b)f/prime(x)=t a n2x>0a n d f(0) = 0; g/prime(x)=(−cosx)(tan x−x)/x2,which is
never positive in the range.
2.24 (a) Any two consecutive roots of ex= 0 have another root of ex= 0 lying between
them; thus there is at most one root of ex= 0 (formally −∞). (b) The root of
2x+ 6 = 0 lies in the range −6<x< 0. (c) Any roots of f(x) = 0 (actually −1
and−1
2) lie on either side of x=−3
4. (d) As in (c), but there are no real roots.
More generally, if there are two values of xthat give 2 x2+3x+kequal values
then they lie one on each side of x=−3
4.( e )f/prime(x)=6 x2−42x+60=0hasroots
2 and 5. Therefore, if f(x) = 0 has three real roots αithen α1<2<α2<5<α3.
(f) The other roots are1
4(15±√
105).
2.25 The false result arises because tan nxis not differentiable at x=π/(2n), which
lies in the range 0 <x<π / n , and so the conditions for applying Rolle’s theorem
are not satisfied.
2.26 (a) y<−ln(1−y)<y /(1−y); (b) y<ey−1<y ey.
2.27 xd y/ d x =( 2−x)y.
2.28 (a) Show that x=−b/(2a).
(b) Possible repeated roots are −3a n d1
3;o n l y−3s a t i s fi e s f(x)=0 .F a c t o r i s e
f(x)a s( x+3 )2(x−b), giving b=2a n d x= 2 as the third root.
(c)f/prime(x) = 0 has the integer solution x=−1 (by inspection); f(x) factorises as
the product ( x+1)2(x2+2x+2) and hence f(x) = 0 has only two (coincident)
real roots.
2.29 By implicit differentiation, y/prime(x)=( 3 x2−12)/(8−3y2), giving y/prime(±2) = 0. Since
y(2) = 4 and y(−2) = 0, the curve touches the x-axis at the point ( −2,0).
2.30 (a) [tan−1(x/2)]/2; (b) sin−1[(x−1)/3]; (c)−2[1 + tan( θ/2)]−1; (d) put y=
(1−x)1/2,l n
/
[1−(1−x)1/2]/[1 + (1−x)1/2]
/
.
2.31 (a) Express in partial fractions; J=1
3ln[(x−1)4/(x+2 ) ]+ c.
(b) Divide the numerator by the denominator and express the remainder in
partial fractions; J=x2/4+4l n ( x+2 )−3l n(x+3 )+ c.
(c) After division of the numerator by the denominator the remainder can be
expressed as 2( x+3 )−1−5(x+3 )−2;J=3x+2l n ( x+3 )+5 ( x+3 )−1+c.
(d) Set x4=u;J=( 4a4)−1tan−1(x4/a4)+c.
2.32 Express as ( x/a)−(b/a2)+(b/a)2(ax+b)−1;(b2/a3)(ln2−1
2).
84
2.4 HINTS AND ANSWERS
2.33 Writing b2−4acas ∆2>0, or 4 ac−b2as ∆/prime2>0:
(i) ∆−1ln[(2ax+b−∆)/(2ax+b+∆ ) ]+ k;
(ii) 2∆/prime−1tan−1[(2ax+b)/∆/prime]+k;
(iii)−2(2ax+b)−1+k.
2.34 (a) J=1
4ln(1 + 4sin2x)+c.
(b) Multiply numerator and denominator by ex;J=1
2ln(e2x−1) +c.
(c) First divide the numerator by the denominator. J=x+l n ( l n x)+c.
(d) Multiply numerator and denominator by xn−1,a n dt h e ns e t xn=u.
J=(nan)−1ln[xn/(xn+an)] +c.
2.35 f/prime(x)=( 1+s i n x)/cos2x=f(x)se cx;J=l n ( f(x)) +c=l n ( s e c x+t a n x)+c.
2.36 (a) Show cos4x−cos2x=s i n4x−sin2x;J=1
5sin5x−1
3sin3x+c.
(b) Either write the numerator and denominator in terms of sinusoidal functions
ofx/2 or make the substitution t=t a n ( x/2);J=2t a n ( x/2)−x+c.
(c) Substitute t=t a n ( x/2);J=2l n ( c o s ( x/2))−2c os2(x/2) +c.
(d) Either set tan x=uor show that the integrand is sec2 xand use the result of
exercise 2.35. J=1
2ln(sec2 x+t a n2 x)+c=1
2ln[(1 + tan x)/(1−tanx)] +c.
2.37 (a) π;( b ) π(b−a)2/8; (c) π(b−a)/2.
2.38 (a) Yes, for λ>0, value λ−1; (b) yes, value 0; (c) no, ln(1 + R)→∞asR→∞;
(d) no, /epsilon1−1→∞as/epsilon1→0; (e) no, ln(sin θ)→−∞ asθ→0; (f) yes, value 1.
2.39 (a) (2 −y2)c osy+2ysiny−2; (b) [( y2lny)/2] + [(1−y2)/4];
(c)ysin−1y+( 1−y2)1/2−1;
(d) ln( a2+1 )−(1/y)ln(a2+y2)+( 2 /a)[tan−1(y/a)−tan−1(1/a)].
2.40 (b) dJ/du =2 (u2−1)3.
2.41 Γ( n+1 )= nΓ(n); (a) (i) n!, (ii) 15√π/8; (b)−2√π.
2.42 (a) π/2, 1, 1, 1 /2, 1/(m+1 ) ,1 /(n+1 ) .
(b) Write the initial integrand as cosm−1θsinnθcosθ, and later rewrite sinn+2θ
as sinnθ(1−cos2θ).
(c) (i) 1 /24, (ii) 8 /693, (iii) 7 π/2048.
2.44 (a) 1; (b) (ln6) /4; (c)
/
2t a n−1[(2a−1)−1/2]
/ /
(2a−1)1/2;( d ) π/3.
2.46 (a) a=2/π,b=1 ;2
3[(1+π
2)3/2−1]>I>π
3(23/2−1), 2.08>I> 1.91; (b) I=2 .
2.47 Set η=1−(x/a)2.
2.49 L=
Rx
0
/;
1+1
4exp 2 x
/1/2dx.
2.50 Note that to avoid any possible double counting, integrals should be taken from
π/2t o3 π/2 and symmetry used for scaling up. The integrands (and infinitesimals)
should be as indicated, with ρ/primedenoting dρ/dφ :
(i) (ρ2/2)dφ,3πa2/2; (ii) 2( ρ/prime2+ρ2)1/2dφ,8a;
(iii) 2 πρcosφ(ρ/prime2+ρ2)1/2dφ,3 2πa2/5;
(iv)πρ2cos2φd(ρsinφ), 8πa3/3.
85
3
Complex numbers and
hyperbolic functions
This chapter is concerned with the representation and manipulation of complex
numbers. Complex numbers pervade this book, underscoring their wide appli-cation in the mathematics of the physical sciences. The application of complexnumbers to the description of physical systems is left until later chapters and
only the basic tools are presented here.
3.1 The need for complex numbers
Although complex numbers occur in many branches of mathematics, they arise
most directly out of solving polynomial equations. We examine a specific quadratic
equation as an example.
Consider the quadratic equation
z
2−4z+5=0 . (3.1)
Equation (3.1) has two solutions, z1andz2, such that
(z−z1)(z−z2)=0 . (3.2)
Using the familiar formula for the roots of a quadratic equation, (1.4), the
solutions z1andz2, written in brief as z1,2,a r e
z1,2=4±radicalbig
(−4)2−4(1×5)
2
=2±√
−4
2. (3.3)
Both solutions contain the square root of a negative number. However, it is not
true to say that there are no solutions to the quadratic equation. The fundamental
theorem of algebra states that a quadratic equation will always have two solutions
and these are in fact given by (3.3). The second term on the RHS of (3.3) iscalled an imaginary term since it contains the square root of a negative number;
86
3.1 THE NEED FOR COMPLEX NUMBERS
11
22
33
445
zf(z)
Figure 3.1 The function f(z)=z2−4z+5 .
the first term is called a realterm. The full solution is the sum of a real term
and an imaginary term and is called a complex number . A plot of the function
f(z)=z2−4z+ 5 is shown in figure 3.1. It will be seen that the plot does not
intersect the z-axis, corresponding to the fact that the equation f(z)=0h a sn o
purely real solutions.
The choice of the symbol zfor the quadratic variable was not arbitrary; the
conventional representation of a complex number is z,w h e r e zis the sum of a
real part xanditimes an imaginary part y,i . e .
z=x+iy,
where iis used to denote the square root of −1. The real part xand the imaginary
part yare usually denoted by Re zand Im zrespectively. We note at this point
that some physical scientists, engineers in particular, use jinstead of i. However,
for consistency, we will use ithroughout this book.
I no u rp a r t i c u l a re x a m p l e ,√
−4=2√
−1=2 i, and hence the two solutions of
(3.1) are
z1,2=2±2i
2=2±i.
Thus here x=2a n d y=±1.
For compactness a complex number is sometimes written in the form
z=(x, y),
where the components of zmay be thought of as coordinates in an xy-plot. Such
a plot is called an Argand diagram and is a common representation of complex
numbers; an example is shown in figure 3.2.
87
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
RezImz
z=x+iy
xy
Figure 3.2 The Argand diagram.
Our particular example of a quadratic equation may be generalised readily to
polynomials whose highest power (degree) is greater than 2, e.g. cubic equations(degree 3), quartic equations (degree 4) and so on. For a general polynomial f(z),
of degree n, the fundamental theorem of algebra states that the equation f(z)=0
will have exactly nsolutions. We will examine cases of higher-degree equations
in subsection 3.4.3.
The remainder of this chapter deals with: the algebra and manipulation of
complex numbers; their polar representation, which has advantages in manycircumstances; complex exponentials and logarithms; the use of complex numbersin finding the roots of polynomial equations; and hyperbolic functions.
3.2 Manipulation of complex numbers
This section considers basic complex number manipulation. Some analogy may
be drawn with vector manipulation (see chapter 7) but this section stands aloneas an introduction.
3.2.1 Addition and subtraction
The addition of two complex numbers, z
1and z2, in general gives another
complex number. The real components and the imaginary components are added
separately and in a like manner to the familiar addition of real numbers:
z1+z2=(x1+iy1)+(x2+iy2)=(x1+x2)+i(y1+y2),
88
3.2 MANIPULATION OF COMPLEX NUMBERS
RezImz
z1z2z1+z2
Figure 3.3 The addition of two complex numbers.
or in component notation
z1+z2=(x1,y1)+(x2,y2)=(x1+x2,y1+y2).
The Argand representation of the addition of two complex numbers is shown in
figure 3.3.
By straightforward application of the commutativity and associativity of the
real and imaginary parts separately, we can show that the addition of complexnumbers is itself commutative and associative, i.e.
z
1+z2=z2+z1,
z1+(z2+z3)=(z1+z2)+z3.
Thus it is immaterial in what order complex numbers are added.ISum the complex numbers 1+2 i,3−4i,−2+i.
Summing the real terms we obtain
1+3−2=2 ,
and summing the imaginary terms we obtain
2i−4i+i=−i.
Hence
(1 + 2 i)+( 3−4i)+(−2+i)=2−i.
J
The subtraction of complex numbers is very similar to their addition. As in the
case of real numbers, if two identical complex numbers are subtracted then theresult is zero.
89
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
RezImz
|z|
xy
argz
Figure 3.4 The modulus and argument of a complex number.
3.2.2 Modulus and argument
The modulus of the complex number zis denoted by |z|and is defined as
|z|=radicalbig
x2+y2. (3.4)
Hence the modulus of the complex number is the distance of the corresponding
point from the origin in the Argand diagram, as may be seen in figure 3.4.
The argument of the complex number zis denoted by arg zand is defined as
argz=t a n−1parenleftBigy
xparenrightBig
. (3.5)
Thus arg zis the angle that the line joining the origin to zon the Argand diagram
makes with the positive x-axis. The anticlockwise direction is taken to be positive
by convention. The angle arg zis shown in figure 3.4. Account must be taken
of the signs of xandyindividually in determining in which quadrant arg zlies.
Thus, for example, if xandyare both negative then arg zlies in the range
−π<argz<−π/2 rather than in the first quadrant (0 <argz<π / 2), though
both cases give the same value for the ratio of ytox.IFind the modulus and the argument of the complex number z=2−3i.
Using (3.4), the modulus is given by
|z|=
p
22+(−3)2=√
13.
Using (3.5), the argument is given by
argz=t a n−1
/;
−3
2
/
.
The two angles whose tangents equal −1.5a r e−0.9828rad and 2 .1588rad. Since x=2a n d
y=−3,zclearly lies in the fourth quadrant; therefore arg z=−0.9828 is the appropriate
answer.
J
90
3.2 MANIPULATION OF COMPLEX NUMBERS
3.2.3 Multiplication
Complex numbers may be multiplied together and in general give a complex
number as the result. The product of two complex numbers z1andz2is found
by multiplying them out in full and remembering that i2=−1, i.e.
z1z2=(x1+iy1)(x2+iy2)
=x1x2+ix1y2+iy1x2+i2y1y2
=(x1x2−y1y2)+i(x1y2+y1x2). (3.6)IMultiply the complex numbers z1=3+2 iandz2=−1−4i.
By direct multiplication we find
z1z2=( 3+2 i)(−1−4i)
=−3−2i−12i−8i2
=5−14i.
J (3.7)
The multiplication of complex numbers is both commutative and associative,
i.e.
z1z2=z2z1, (3.8)
(z1z2)z3=z1(z2z3). (3.9)
The product of two complex numbers also has the simple properties
|z1z2|=|z1||z2|, (3.10)
arg(z1z2)=a r g z1+a r g z2. (3.11)
These relations are derived in subsection 3.3.1.IVerify that (3.10) holds for the product of z1=3+2 iandz2=−1−4i.
From (3.7)
|z1z2|=|5−14i|=
p
52+(−14)2=√
221.
We also find
|z1|=
p
32+22=√
13,
|z2|=
p
(−1)2+(−4)2=√
17,
and hence
|z1||z2|=√
13√
17 =√
221 =|z1z2|.
J
We now examine the effect on a complex number zof multiplying it by ±1
and±i. These four multipliers have modulus unity and we can see immediately
from (3.10) that multiplying zby another complex number of unit modulus gives
a product with the same modulus as z. We can also see from (3.11) that if
91
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
RezImz
iz
−izz
−z
Figure 3.5 Multiplication of a complex number by ±1a n d±i.
we multiply zby a complex number, the argument of the product is the sum
of the argument of zand the argument of the multiplier. Hence multiplying
zby unity (which has argument zero) leaves zunchanged in both modulus
and argument, i.e. zis completely unaltered by the operation. Multiplying by
−1 (which has argument π) leads to rotation, through an angle π, of the line
joining the origin to zin the Argand diagram. Similarly, multiplication by ior
−ilead to corresponding rotations of π/2o r−π/2 respectively. This geometrical
interpretation of multiplication is shown in figure 3.5.IUsing the geometrical interpretation of multiplication by i, find the product i(1−i).
The complex number 1 −ihas argument −π/4 and modulus√
2. Thus, using (3.10) and
(3.11), its product with ihas argument + π/4 and unchanged modulus√
2. The complex
number with modulus√
2 and argument + π/4i s1+ iand so
i(1−i)=1+ i,
as is easily verified by direct multiplication.
J
The division of two complex numbers is similar to their multiplication but
requires the notion of the complex conjugate (see the following subsection) andso discussion is postponed until subsection 3.2.5.
3.2.4 Complex conjugate
Ifzhas the convenient form x+iythen the complex conjugate, denoted by z
∗,
may be found simply by changing the sign of the imaginary part, i.e. if z=x+iy
then z∗=x−iy. More generally, we may define the complex conjugate of zas
the (complex) number having the same magnitude as zthat when multiplied by
zleaves a real result, i.e. there is no imaginary component in the product.
92
3.2 MANIPULATION OF COMPLEX NUMBERS
RezImz
z=x+iy
xy
−yz∗=x−iy
Figure 3.6 The complex conjugate as a mirror image in the real axis.
I nt h ec a s ew h e r e zc a nb ew r i t t e ni nt h ef o r m x+iyit is easily verified, by
direct multiplication of the components, that the product zz∗gives a real result:
zz∗=(x+iy)(x−iy)=x2−ixy+ixy−i2y2=x2+y2=|z|2.
Complex conjugation corresponds to a reflection of zin the real axis of the
Argand diagram, as may be seen in figure 3.6.IFind the complex conjugate of z=a+2i+3ib.
The complex number is written in the standard form
z=a+i(2 + 3 b);
then, replacing iby−i,w eo b t a i n
z∗=a−i(2 + 3 b).
J
In some cases, however, it may not be simple to rearrange the expression for
zinto the standard form x+iy. Nevertheless, given two complex numbers, z1
andz2, it is straightforward to show that the complex conjugate of their sum
(or difference) is equal to the sum (or difference) of their complex conjugates, i.e.(z
1±z2)∗=z∗
1±z∗
2. Similarly, it may be shown that the complex conjugate of the
product (or quotient) of z1andz2is equal to the product (or quotient) of their
complex conjugates, i.e. ( z1z2)∗=z∗
1z∗
2and ( z1/z2)∗=z∗
1/z∗
2.
Using these results, it can be deduced that, no matter how complicated the
expression, its complex conjugate may always be found by replacing every iby
−i. To apply this rule, however, we must always ensure that all complex parts are
first written out in full, so that no i’s are hidden.
93
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONSIFind the complex conjugate of the complex number z=w(3y+2ix)where w=x+5i.
Although we do not discuss complex powers until section 3.5, the simple rule given above
still enables us to find the complex conjugate of z.
In this case witself contains real and imaginary components and so must be written
out in full, i.e.
z=w3y+2ix=(x+5i)3y+2ix.
Now we can replace each iby−ito obtain
z∗=(x−5i)(3y−2ix).
It can be shown that the product zz∗is real, as required.
J
The following properties of the complex conjugate are easily proved and others
may be derived from them. If z=x+iythen
(z∗)∗=z, (3.12)
z+z∗=2R e z=2x, (3.13)
z−z∗=2iImz=2iy, (3.14)
z
z∗=parenleftbiggx2−y2
x2+y2parenrightbigg
+iparenleftbigg2xy
x2+y2parenrightbigg
. (3.15)
The derivation of this last relation relies on the results of the following subsection.
3.2.5 Division
The division of two complex numbers z1andz2bears some similarity to their
multiplication. Writing the quotient in component form we obtain
z1
z2=x1+iy1
x2+iy2. (3.16)
In order to separate the real and imaginary components of the quotient, we
multiply both numerator and denominator by the complex conjugate of the
denominator. By definition, this process will leave the denominator as a realquantity. Equation (3.16) gives
z
1
z2=(x1+iy1)(x2−iy2)
(x2+iy2)(x2−iy2)=(x1x2+y1y2)+i(x2y1−x1y2)
x2
2+y2
2
=x1x2+y1y2
x2
2+y2
2+ix2y1−x1y2
x2
2+y2
2.
Hence we have separated the quotient into real and imaginary components, as
required.
In the special case where z2=z∗
1,s ot h a t x2=x1andy2=−y1, the general
result reduces to (3.15).
94
3.3 POLAR REPRESENTATION OF COMPLEX NUMBERSIExpress zin the form x+iy,w h e n
z=3−2i
−1+4 i.
Multiplying numerator and denominator by the complex conjugate of the denominator
we obtain
z=(3−2i)(−1−4i)
(−1+4 i)(−1−4i)=−11−10i
17
=−11
17−10
17i.
J
In analogy to (3.10) and (3.11), which describe the multiplication of two
complex numbers, the following relations apply to division:
vextendsinglevextendsinglevextendsinglevextendsinglez1
z2vextendsinglevextendsinglevextendsinglevextendsingle=|z1|
|z2|, (3.17)
argparenleftbiggz1
z2parenrightbigg
=a r g z1−argz2. (3.18)
The proof of these relations is left until subsection 3.3.1.
3.3 Polar representation of complex numbers
Although considering a complex number as the sum of a real and an imaginary
part is often useful, sometimes the polar representation proves easier to manipulate.
This makes use of the complex exponential function, which is defined by
ez=e x p z≡1+z+z2
2!+z3
3!+···. (3.19)
Strictly speaking it is the function exp zthat is defined by (3.19). The number e
is the value of exp(1), i.e. it is just a number. However, it may be shown that ez
and exp zare equivalent when zis real and rational and mathematicians then
define their equivalence for irrational and complex z. For the purposes of this
book we will not concern ourselves further with this mathematical nicety but,rather, assume that (3.19) is valid for all z. We also note that, using (3.19), by
multiplying together the appropriate series we may show that (see chapter 20)
e
z1ez2=ez1+z2, (3.20)
which is analogous to the familiar result for exponentials of real numbers.
95
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
RezImz
r
θ
xy z=reiθ
Figure 3.7 The polar representation of a complex number.
From (3.19), it immediately follows that for z=iθ,θreal,
eiθ=1+ iθ−θ2
2!−iθ3
3!+··· (3.21)
=1−θ2
2!+θ4
4!−···+iparenleftbigg
θ−θ3
3!+θ5
5!−···parenrightbigg
, (3.22)
and hence that
eiθ=c o s θ+isinθ, (3.23)
where the last equality follows from the series expansions of trigonometric func-
tions (see subsection 4.6.3). This last relationship is called Euler’s equation . It also
follows from (3.23) that
einθ=c o s nθ+isinnθ
for all n. From Euler’s equation (3.23) and using figure 3.7 we deduce that
reiθ=r(cosθ+isinθ)
=x+iy.
Thus a complex number may be represented in the polar form
z=reiθ. (3.24)
Referring again to figure 3.7, we can identify rwith|z|andθwith arg z.T h e
simplicity of the representation of the modulus and argument is one of the mainreasons for using the polar representation. The angle θlies conventionally in the
range−π<θ≤π, but, since rotation by θis the same as rotation by 2 nπ+θ,
where nis any integer,
re
iθ≡rei(θ+2nπ).
96
3.3 POLAR REPRESENTATION OF COMPLEX NUMBERS
RezImz
r1r2ei(θ1+θ2)
r2eiθ2
r1eiθ1
Figure 3.8 The multiplication of two complex numbers. In this case r1and
r2are both greater than unity.
The algebra of the polar representation is different from that of the real and
imaginary component representation, though, of course, the results are identical.
Some operations prove much easier in the polar representation, others much more
complicated. The best representation for a particular problem must be determinedby the manipulation required.
3.3.1 Multiplication and division in polar form
Multiplication and division in polar form are particularly simple. The product of
z
1=r1eiθ1andz2=r2eiθ2is given by
z1z2=r1eiθ1r2eiθ2
=r1r2ei(θ1+θ2). (3.25)
The relations |z1z2|=|z1||z2|and arg( z1z2)=a r g z1+a r g z2follow immediately.
An example of the multiplication of two complex numbers is shown in figure 3.8.
Division is equally simple in polar form; the quotient of z1andz2is given by
z1
z2=r1eiθ1
r2eiθ2=r1
r2ei(θ1−θ2). (3.26)
The relations |z1/z2|=|z1|/|z2|and arg( z1/z2)=a r g z1−argz2are again imme-
97
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
RezImz
r1
r2ei(θ1−θ2)r2eiθ2r1eiθ1
Figure 3.9 The division of two complex numbers. As in the previous figure,
r1andr2are both greater than unity.
diately apparent. The division of two complex numbers in polar form is shown
in figure 3.9.
3.4 de Moivre’s theorem
We now derive an extremely important theorem. Sinceparenleftbig
eiθparenrightbign=einθ, we have
(cosθ+isinθ)n=c o s nθ+isinnθ, (3.27)
where the identity einθ=c o s nθ+isinnθfollows from the series definition of
einθ(see (3.21)). This result is called de Moivre’s theorem a n di so f t e nu s e di nt h e
manipulation of complex numbers. The theorem is valid for all nwhether real,
imaginary or complex.
There are numerous applications of de Moivre’s theorem but this section
examines just three: proofs of trigonometric identities; finding the nth roots of
unity; and solving complex equations.
3.4.1 Trigonometric identities
The use of de Moivre’s theorem in finding trigonometric identities is best illus-
trated by example. We consider the expression of a multiple-angle function interms of a polynomial in the single-angle function, and its converse.
98
3.4 DE MOIVRE’S THEOREMIExpress sin3θandcos 3θi nt e r m so fp o w e r so f cosθandsinθ.
Using de Moivre’s theorem,
cos 3θ+isin 3θ=( c o s θ+isinθ)3
=( c o s3θ−3c os θsin2θ)+i(3 sin θcos2θ−sin3θ). (3.28)
We can equate the real and imaginary coefficients separately, i.e.
cos 3θ=c o s3θ−3cos θsin2θ
=4c o s3θ−3c os θ (3.29)
and
sin 3θ=3s i n θcos2θ−sin3θ
=3s i n θ−4si n3θ.
J
This method can clearly be applied to finding power expansions of cos nθand
sinnθfor any positive integer n.
The converse process uses the following properties of z=eiθ,
zn+1
zn=2c o s nθ, (3.30)
zn−1
zn=2isinnθ. (3.31)
These equalities follow from simple applications of de Moivre’s theorem, i.e.
zn+1
zn=( c o s θ+isinθ)n+( c o s θ+isinθ)−n
=c o s nθ+isinnθ+c o s (−nθ)+isin(−nθ)
=c o s nθ+isinnθ+c o s nθ−isinnθ
=2c o s nθ
and
zn−1
zn=( c o s θ+isinθ)n−(cosθ+isinθ)−n
=c o s nθ+isinnθ−cosnθ+isinnθ
=2isinnθ.
In the particular case where n=1 ,
z+1
z=eiθ+e−iθ=2c o s θ, (3.32)
z−1
z=eiθ−e−iθ=2isinθ. (3.33)
99
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONSIFind an expression for cos3θin terms of cos 3θandcosθ.
Using (3.32),
cos3θ=1
23
/
z+1
z
/3
=1
8
/
z3+3z+3
z+1
z3
/
=1
8
/
z3+1
z3
/
+3
8
/
z+1
z
/
.
Now using (3.30) and (3.32), we find
cos3θ=1
4cos 3θ+3
4cosθ.
J
This result happens to be a simple rearrangement of (3.29), but cases involving
larger values of nare better handled using this direct method than by rearranging
polynomial expansions of multiple-angle functions.
3.4.2 Finding the nth roots of unity
The equation z2= 1 has the familiar solutions z=±1. However, now that
we have introduced the concept of complex numbers we can solve the generalequation z
n= 1. Recalling the fundamental theorem of algebra, we know that
the equation has nsolutions. In order to proceed we rewrite the equation as
zn=e2ikπ,
where kis any integer. Now taking the nth root of each side of the equation we
find
z=e2ikπ/n.
Hence, the solutions of zn=1a r e
z1,2,...,n=1,e2iπ/n, ..., e2i(n−1)π/n,
corresponding to the values 0 ,1,2,...,n−1f o r k. Larger integer values of kdo
not give new solutions, since the roots already listed are simply cyclically repeatedfork=n, n+1,n+2 ,e t c .IFind the solutions to the equation z3=1.
By applying the above method we find
z=e2ikπ/3.
Hence the three solutions are z1=e0i=1 ,z2=e2iπ/3,z3=e4iπ/3. We note that, as expected,
the next solution, for which k=3 ,g i v e s z4=e6iπ/3=1= z1, so that there are only three
separate solutions.
J
100
3.4 DE MOIVRE’S THEOREM
RezImz
e−2iπ/3e2iπ/3
2π/32π/3
1
Figure 3.10 The solutions of z3=1 .
Not surprisingly, given that |z3|=|z|3from (3.10), all the roots of unity have
unit modulus, i.e. they all lie on a circle in the Argand diagram of unit radius.The three roots are shown in figure 3.10.
The cube roots of unity are often written 1, ωandω
2. The properties ω3=1
and 1 + ω+ω2= 0 are easily proved.
3.4.3 Solving polynomial equations
A third application of de Moivre’s theorem is to the solution of polynomial
equations. Complex equations in the form of a polynomial relationship must first
be solved for zin a similar fashion to the method for finding the roots of real
polynomial equations. Then the complex roots of zmay be found.ISolve the equation z6−z5+4z4−6z3+2z2−8z+8=0 .
We first factorise to give
(z3−2)(z2+4 ) ( z−1) = 0 .
Hence z3=2o r z2=−4o r z= 1. The solutions to the quadratic equation are z=±2i;
to find the complex cube roots, we first write the equation in the form
z3=2=2 e2ikπ,
where kis any integer. If we now take the cube root, we get
z=21/3e2ikπ/3.
101
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
To avoid the duplication of solutions, we use the fact that −π<argz≤πand find
z1=21/3,
z2=21/3e2πi/3=21/3
/
−1
2+√
3
2i
/!
,
z3=21/3e−2πi/3=21/3
/
−1
2−√
3
2i
/!
.
The complex numbers z1,z2andz3,t o g e t h e rw i t h z4=2i,z5=−2iandz6=1a r et h e
solutions to the original polynomial equation.
As expected from the fundamental theorem of algebra, we find that the total number
of complex roots (six, in this case) is equal to the largest power of zin the polynomial.
J
A useful result is that the roots of a polynomial with real coefficients occur in
conjugate pairs (i.e. if z1is a root, then z∗
1is a second distinct root, unless z1is
real). This may be proved as follows. Let the polynomial equation of which zis
ar o o tb e
anzn+an−1zn−1+···+a1z+a0=0.
Taking the complex conjugate of this equation,
a∗
n(z∗)n+a∗
n−1(z∗)n−1+···+a∗
1z∗+a∗
0=0.
But the anare real, and so z∗satisfies
an(z∗)n+an−1(z∗)n−1+···+a1z∗+a0=0,
and is also a root of the original equation.
3.5 Complex logarithms and complex powers
The concept of a complex exponential has already been introduced in section 3.3,
where it was assumed that the definition of an exponential as a series was validfor complex numbers as well as for real numbers. Similarly we can define thelogarithm of a complex number and we can use complex numbers as exponents.
Let us denote the natural logarithm of a complex number zbyw=L n z,w h e r e
the notation Ln will be explained shortly. Thus, wmust satisfy
z=e
w.
Using (3.20), we see that
z1z2=ew1ew2=ew1+w2,
and taking logarithms of both sides we find
Ln(z1z2)=w1+w2=L n z1+L n z2, (3.34)
which shows that the familiar rule for the logarithm of the product of two real
numbers also holds for complex numbers.
102
3.5 COMPLEX LOGARITHMS AND COMPLEX POWERS
We may use (3.34) to investigate further the properties of Ln z. We have already
noted that the argument of a complex number is multivalued, i.e. arg z=θ+2nπ,
where nis any integer. Thus, in polar form, the complex number zshould strictly
be written as
z=rei(θ+2nπ).
Taking the logarithm of both sides, and using (3.34), we find
Lnz=l nr+i(θ+2nπ), (3.35)
where ln ris the natural logarithm of the real positive quantity ra n ds oi s
written normally. Thus from (3.35) we see that Ln zis itself multivalued. To avoid
this multivalued behaviour it is conventional to define another function ln z,t h e
principal value of Ln z, which is obtained from Ln zby restricting the argument
ofzto lie in the range −π<θ≤π.IEvaluate Ln(−i).
By rewriting −ias a complex exponential, we find
Ln(−i)=L n
/
ei(−π/2+2nπ)
/
=i(−π/2+2 nπ),
where nis any integer. Hence Ln( −i)=−iπ/2,3iπ/2,. . .. We note that ln( −i), the
principal value of Ln( −i), is given by ln( −i)=−iπ/2.
J
Ifzandtare both complex numbers then the zth power of tis defined by
tz=ezLnt.
Since Ln tis multivalued, so too is this definition.ISimplify the expression z=i−2i.
Firstly we take the logarithm of both sides of the equation to give
Lnz=−2iLni.
Now inverting the process we find
eLnz=z=e−2iLni.
We can write i=ei(π/2+2nπ),where nis any integer, and hence
Lni=L n
h
ei(π/2+2nπ)
i
=i
/;
π/2+2 nπ
/
.
We can now simplify zto give
i−2i=e−2i×i(π/2+2nπ)
=e(π+4nπ),
which, perhaps surprisingly, is a real quantity rather than a complex one.
J
Complex powers and the logarithms of complex numbers are discussed further
in chapter 20.
103
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
3.6 Applications to differentiation and integration
We can use the exponential form of a complex number together with de Moivre’s
theorem (see section 3.4) to simplify the differentiation of trigonometric functions.IFind the derivative with respect to xofe3xcos 4x.
We could differentiate this function straightforwardly using the product rule (see subsec-
tion 2.1.2). However, an alternative method in this case is to use a complex exponential.Let us consider the complex number
z=e
3x(cos4 x+isin4x)=e3xe4ix=e(3+4 i)x,
where we have used de Moivre’s theorem to rewrite the trigonometric functions as a com-
plex exponential. This complex number has e3xcos 4xas its real part. Now, differentiating
zwith respect to xwe obtain
dz
dx=( 3+4 i)e(3+4 i)x=( 3+4 i)e3x(cos4 x+isin 4x), (3.36)
where we have again used de Moivre’s theorem. Equating real parts we then find
d
dx
/;
e3xcos 4x
/
=e3x(3cos 4 x−4si n4 x).
By equating the imaginary parts of (3.36), we also obtain, as a bonus,
d
dx
/;
e3xsin 4x
/
=e3x(4cos4 x+3s i n4 x).
J
In a similar way the complex exponential can be used to evaluate integrals
containing trigonometric and exponential functions.IEvaluate the integral I=
R
eaxcosbx dx.
Let us consider the integrand as the real part of the complex number
eax(cosbx+isinbx)=eaxeibx=e(a+ib)x,
where we use de Moivre’s theorem to rewrite the trigonometric functions as a complex
exponential. Integrating we findZ
e(a+ib)xdx=e(a+ib)x
a+ib+c
=(a−ib)e(a+ib)x
(a−ib)(a+ib)+c
=eax
a2+b2
/;
aeibx−ibeibx
/
+c, (3.37)
where the constant of integration cis in general complex. Denoting this constant by
c=c1+ic2and equating real parts in (3.37) we obtain
I=
Z
eaxcosbx dx=eax
a2+b2(acosbx+bsinbx)+c1,
which agrees with result (2.37) found using integration by parts. Equating imaginary parts
in (3.37) we obtain, as a bonus,
J=
Z
eaxsinbx dx=eax
a2+b2(asinbx−bcosbx)+c2.
J
104
3.7 HYPERBOLIC FUNCTIONS
3.7 Hyperbolic functions
Thehyperbolic functions are the complex analogues of the trigonometric functions.
The analogy may not be immediately apparent and their definitions may appearat first to be somewhat arbitrary. However, careful examination of their propertiesreveals the purpose of the definitions. For instance, their close relationship with
the trigonometric functions, both in their identities and in their calculus, means
that many of the familiar properties of trigonometric functions can also be appliedto the hyperbolic functions. Further, hyperbolic functions occur regularly, and sogiving them special names is a notational convenience.
3.7.1 Definitions
The two fundamental hyperbolic functions are cosh xand sinh x, which, as their
names suggest, are the hyperbolic equivalents of cos xand sin x. They are defined
by the following relations:
coshx=
1
2(ex+e−x), (3.38)
sinhx=1
2(ex−e−x). (3.39)
Note that cosh xis an even function and sinh xis an odd function. By analogy
with the trigonometric functions, the remaining hyperbolic functions are
tanh x=sinhx
coshx=ex−e−x
ex+e−x, (3.40)
sechx=1
coshx=2
ex+e−x, (3.41)
cosech x=1
sinhx=2
ex−e−x, (3.42)
cothx=1
tanh x=ex+e−x
ex−e−x. (3.43)
All the hyperbolic functions above have been defined in terms of the real variable
x. However, this was simply so that they may be plotted (see figures 3.11–3.13);
the definitions are equally valid for any complex number z.
3.7.2 Hyperbolic–trigonometric analogies
In the previous subsections we have alluded to the analogy between trigonometric
and hyperbolic functions. Here, we discuss the close relationship between the twogroups of functions.
Recalling (3.32) and (3.33) we find
cosix=
1
2(ex+e−x),
sinix=1
2i(ex−e−x).
105
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
sech xcoshx
x1234
−1 −21 2
Figure 3.11 Graphs of cosh xand sech x.
cosech xcosech x
sinhx
x24
−2
−4−1 −21 2
Figure 3.12 Graphs of sinh xand cosech x.
Hence, by the definitions given in the previous subsection,
coshx=c o s ix, (3.44)
isinhx=s i n ix, (3.45)
cosx=c o s h ix, (3.46)
isinx= sinh ix. (3.47)
These useful equations make the relationship between hyperbolic and trigono-
106
3.7 HYPERBOLIC FUNCTIONS
cothxcothx
tanh x
x24
−2
−4−1 −21 2
Figure 3.13 Graphs of tanh xand coth x.
metric functions transparent. The similarity in their calculus is discussed further
in subsection 3.7.6.
3.7.3 Identities of hyperbolic functions
The analogies between trigonometric functions and hyperbolic functions having
been established, we should not be surprised that all the trigonometric identitiesalso hold for hyperbolic functions, with the following modification. Whereversin
2xoccurs it must be replaced by −sinh2x, and vice versa. Note that this
replacement is necessary even if the sin2xis hidden, e.g. tan2x=s i n2x/cos2x
and so must be replaced by ( −sinh2x/cosh2x)=−tanh2x.IFind the hyperbolic identity analogous to cos2x+s i n2x=1.
Using the rules stated above cos2xmust be replaced by cosh2x,a n ds i n2xmust be replaced
by−sinh2x, and so the identity becomes
cosh2x−sinh2x=1.
This can be verified by direct substitution, using the definitions of cosh xand sinh x;s e e
(3.38) and (3.39).
J
Some other identities that can be proved in a similar way are
sech2x=1−tanh2x, (3.48)
cosech2x=c o t h2x−1, (3.49)
sinh2 x=2s i n h xcoshx, (3.50)
cosh2 x=c o s h2x+ sinh2x. (3.51)
107
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
3.7.4 Solving hyperbolic equations
When we are presented with a hyperbolic equation to solve, we may proceed
by analogy with the solution of trigonometric equations. However, it is almostalways easier to express the equation directly in terms of exponentials.ISolve the hyperbolic equation coshx−5si nh x−5=0.
Substituting the definitions of the hyperbolic functions we obtain
1
2(ex+e−x)−5
2(ex−e−x)−5=0 .
Rearranging, and then multiplying through by −ex, gives in turn
−2ex+3e−x−5=0
and
2e2x+5ex−3=0 .
N o ww ec a nf a c t o r i s ea n ds o l v e :
(2ex−1)(ex+3 )=0 .
Thus ex=1/2o r ex=−3. Hence x=−ln 2 or x=l n (−3). The interpretation of the
logarithm of a negative number has been discussed in section 3.5.
J
3.7.5 Inverses of hyperbolic functions
Just like trigonometric functions, hyperbolic functions have inverses. If y=
coshxthen x=c o s h−1y, which serves as a definition of the inverse. By using
the fundamental definitions of hyperbolic functions, we can find closed-formexpressions for their inverses. This is best illustrated by example.IFind a closed-form expression fo r the inverse hyperbolic function y=s i n h−1x.
First we write xas a function of y,i . e .
y=s i n h−1x⇒x=s i n h y.
Now, since cosh y=1
2(ey+e−y) and sinh y=1
2(ey−e−y),
ey=c o s h y+s i n h y
=
q
1+s i n h2y+s i n h y
ey=
p
1+x2+x,
and hence
y=l n (
p
1+x2+x).
J
In a similar fashion it can be shown that
cosh−1x=l n (√
x2−1+x).
108
3.7 HYPERBOLIC FUNCTIONS
sech−1xsech−1x
cosh−1xcosh−1x
x24
−2
−43 412
Figure 3.14 Graphs of cosh−1xand sech−1x.IFind a closed-form expression fo r the inverse hyperbolic function y=t a n h−1x.
First we write xas a function of y,i . e .
y=t a n h−1x⇒ x=t a n h y.
Now, using the definition of tanh yand rearranging, we find
x=ey−e−y
ey+e−y⇒ (x+1 )e−y=( 1−x)ey.
Thus, it follows that
e2y=1+x
1−x⇒ ey=
r
1+x
1−x,
y=l n
r
1+x
1−x,
tanh−1x=1
2ln
/1+x
1−x
/
.
J
Graphs of the inverse hyperbolic functions are given in figures 3.14–3.16.
3.7.6 Calculus of hyperbolic functions
Just as the identities of hyperbolic functions closely follow those of their trigono-
metric counterparts, so their calculus is similar. The derivatives of the two basic
109
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
sinh−1x
cosech−1xcosech−1x
x24
−2
−4−1 −2 12
Figure 3.15 Graphs of sinh−1xand cosech−1x.
coth−1xcoth−1xtanh−1x
x24
−2
−4−1 −21 2
Figure 3.16 Graphs of tanh−1xand coth−1x.
hyperbolic functions are
d
dx(coshx)= sinh x, (3.52)
d
dx(sinhx)=c o s h x. (3.53)
These may be deduced by considering the definitions.
110
3.7 HYPERBOLIC FUNCTIONSIVerify the relation (d/dx)c osh x=s i n h x.
Using the definition of cosh x,
coshx=1
2(ex+e−x),
and differentiating directly, we find
d
dx(coshx)=1
2(ex−e−x)
=s i n h x.
J
Clearly the integrals of the fundamental hyperbolic functions are also defined
by these relations. The derivatives of the remaining hyperbolic functions can bederived by product differentiation and are presented below only for complete-ness.
d
dx(tanh x)=s e c h2x, (3.54)
d
dx(sech x)=−sech xtanh x, (3.55)
d
dx(cosech x)=−cosech xcothx, (3.56)
d
dx(cothx)=−cosech2x. (3.57)
The inverse hyperbolic functions also have derivatives, which are given by the
following:
d
dxparenleftBig
cosh−1x
aparenrightBig
=1√
x2−a2, (3.58)
d
dxparenleftBig
sinh−1x
aparenrightBig
=1√
x2+a2, (3.59)
d
dxparenleftBig
tanh−1x
aparenrightBig
=a
a2−x2,forx2<a2, (3.60)
d
dxparenleftBig
coth−1x
aparenrightBig
=−a
x2−a2,forx2>a2. (3.61)
These may be derived from the logarithmic form of the inverse (see subsec-
tion 3.7.5).
111
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONSIEvaluate (d/dx)si nh−1xusing the logarithmic form of the inverse.
From the results of section 3.7.5,
d
dx
/;
sinh−1x
/
=d
dx
h
ln
/
x+
p
x2+1
/i
=1
x+√
x2+1
/
1+x√
x2+1
/
=1
x+√
x2+1
/ √
x2+1+ x√
x2+1
/!
=1√
x2+1.
J
3.8 Exercises
3.1 Two complex numbers zandware given by z=3+4 iandw=2−i.O na n
Argand diagram plot
(a)z+w,( b ) w−z,( c )wz,( d ) z/w,
(e)z∗w+w∗z,( f )w2,( g )l n z,( h )( 1+ z+w)1/2.
3.2 By considering the real and imaginary parts of the product eiθeiφprove the
standard formulae for cos( θ+φ) and sin( θ+φ).
3.3 By writing π/12 = ( π/3)−(π/4) and considering eiπ/12, evaluate cot( π/12).
3.4 Find the locus in the complex z-plane of points that satisfy the following equa-
tions.
(a)z−c=ρ
/1+it
1−it
/
,where cis complex, ρis real and tis a real parameter
that varies in the range −∞<t<∞.
(b)z=a+bt+ct2,i nw h i c h tis a real parameter and a,b,a n d care complex
numbers with b/creal.
3.5 Evaluate
(a) Re(exp 2 iz), (b) Im(cosh2z), (c) (−1+√
3i)1/2,
(d)|exp(i1/2)|, (e) exp( i3), (f) Im(2i+3), (g) ii,( h )l n [ (√
3+i)3].
3.6 Find the equations in terms of xandyof the sets of points in the Argand
diagram that satisfy the following:
(a) Re z2=I m z2;
(b) (Im z2)/z2=−i;
(c) arg[ z/(z−1)] = π/2.
3.7 Show that the locus of all points z=x+iyin the complex plane that satisfy
|z−ia|=λ|z+ia|,λ > 0,
is a circle of radius |2λ/(1−λ2)|acentred on the point z=ia[(1 + λ2)/(1−λ2)].
Sketch the circles for a few typical values of λ, including λ<1,λ>1a n d λ=1 .
3.8 The two sets of points z=a,z=b,z=c,a n d z=A,z=B,z=Care
the corners of two similar triangles in the Argand diagram. Express in terms ofa, b, . . . , C
112
3.8 EXERCISES
(a) the equalities of corresponding angles, and
(b) the constant ratio of corresponding sides,
in the two triangles.
By noting that any complex quantity can be expressed as
z=|z|exp(iargz),
deduce that
a(B−C)+b(C−A)+c(A−B)=0 .
3.9 For the real constant afind the loci of all points z=x+iyin the complex plane
that satisfy
(a) Re
/
ln
/z−ia
z+ia
//
=c, c>0,
(b) Im
/
ln
/z−ia
z+ia
//
=k,0 ≤k≤π/2.
Identify the two families of curves and verify that in case (b) all curves pass
through the two points ±ia.
3.10 The most general type of transformation between one Argand diagram, in the
z-plane, and another, in the Z-plane, that gives one and only one value of Zfor
each value of z(and conversely) is known as the general bilinear transformation
and takes the form
z=aZ+b
cZ+d.
(a) Confirm that the transformation from the Z-plane to the z-plane is also a
general bilinear transformation.
(b) Recalling that the equation of a circle can be written in the form////z−z1
z−z2
////=λ, λ /negationslash=1,
show that the general bilinear transformation transforms circles into circles
(or straight lines). What is the condition that z1,z2andλmust satisfy if the
transformed circle is to be a straight line?
3.11 Sketch the parts of the Argand diagram in which
(a) Re z2<0,|z1/2|≤2,
(b) 0≤argz∗≤π/2,
(c)|expz3|→0a s|z|→∞ .
What is the area of the region in which all three conditions are satisfied?
3.12 Denote the nth roots of unity by 1, ωn,ω2
n,...,ωn−1
n.
(a) Prove that
(i)n−1X
r=0ωr
n=0,(ii)n−1Y
r=0ωr
n=(−1)n+1.
(b) Express x2+y2+z2−yz−zx−xyas the product of two factors, each linear
inx,yandz, with coefficients dependent on the third roots of unity (and
those of the xterms arbitrarily taken as real).
113
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
3.13 Prove that x2m+1−a2m+1,w h e r e mis an integer ≥1 ,c a nb ew r i t t e na s
x2m+1−a2m+1=(x−a)mY
r=1
/
x2−2axcos
/2πr
2m+1
/
+a2
/
.
3.14 The complex position vectors of two parallel interacting equal fluid vortices
moving with their axes of rotation always perpendicular to the z-plane are z1
andz2. The equations governing their motions are
dz∗
1
dt=−i
z1−z2,dz∗
2
dt=−i
z2−z1.
Deduce that (a) z1+z2,( b )|z1−z2|and (c)|z1|2+|z2|2are all constant in time,
and hence describe the motion geometrically.
3.15 Solve the equation
z7−4z6+6z5−6z4+6z3−12z2+8z+4=0 ,
(a) by examining the effect of setting z3e q u a lt o2 ,a n d
(b) by factorising and using the binomial expansion of ( z+a)4.
Plot the seven roots of the equation on an Argand plot, exemplifying that complex
roots of a polynomial equation always occur in conjugate pairs if the polynomialhas real coefficients.
3.16 The polynomial f(z) is defined by
f(z)=z
5−6z4+1 5z3−34z2+3 6z−48.
(a) Show that the equation f(z) = 0 has roots of the form z=λiwhere λis real,
and hence factorize f(z).
(b) Show further that the cubic factor of f(z) can be written in the form
(z+a)3+b,w h e r e aandbare real, and hence solve the equation f(z)=0
completely.
3.17 The binomial expansion of (1 + x)n, discussed in chapter 1, can be written for a
positive integer nas
(1 +x)n=nX
r=0nCrxr,
wherenCr=n!/[r!(n−r)!].
(a) Use de Moivre’s theorem to show that the sum
S1(n)=nC0−nC2+nC4−···+(−1)mnC2m,n−1≤2m≤n,
has the value 2n/2cos(nπ/4).
(b) Derive a similar result for the sum
S2(n)=nC1−nC3+nC5−···+(−1)mnC2m+1,n−1≤2m+1≤n,
and verify it for the cases n=6 ,7a n d8 .
3.18 By considering (1 + exp iθ)n, prove that
nX
r=0nCrcosnθ=2ncosn(θ/2)cos( nθ/2),
nX
r=0nCrsinnθ=2ncosn(θ/2)sin( nθ/2),
wherenCr=n!/[r!(n−r)!].
114
3.8 EXERCISES
3.19 Use de Moivre’s theorem with n= 4 to prove that
cos 4θ=8c o s4θ−8cos2θ+1,
and deduce that
cosπ
8=
/
2+√
2
4
/!1/2
.
3.20 Express sin4θentirely in terms of the trigonometric functions of multiple angles
and deduce that its average value over a complete cycle is3
8.
3.21 Use de Moivre’s theorem to prove that
tan5θ=t5−10t3+5t
5t4−10t2+1,
where t=t a n θ. Deduce the values of tan( nπ/10) for n=1 ,2 ,3 ,4 .
3.22 (a) Prove that
coshx−coshy=2s i n h
/x+y
2
/
sinh
/x−y
2
/
.
(b) Prove that, if y=s i n h−1x,
(x2+1 )d2y
dx2+xdy
dx=0.
3.23 Determine the conditions under which the equation
acoshx+bsinhx=c, c > 0,
has zero, one, or two real solutions for x. What is the solution if a2=c2+b2?
3.24 (a) Solve cosh x=s i n h x+2 s e c h x.
(b) Show that the real solution xof tanh x=c o s e c h xc a nb ew r i t t e ni nt h e
form x=l n ( u+√u). Find an explicit value for u.
(c) Evaluate tanh xwhen xis the real solution of cosh2 x=2c o s h x.
3.25 Express sinh4xin terms of hyperbolic cosines of multiples of x, and hence solve
2cosh4 x−8cosh2 x+5=0 .
3.26 In the theory of special relativity, the relationship between the position and time
coordinates of an event as measured in two frames of reference that have parallelx-axes can be expressed in terms of hyperbolic functions. If the coordinates are x
andtin one frame and x
/primeandt/primein the other then the relationship take the form
x/prime=xcoshφ−ctsinhφ,
ct/prime=−xsinhφ+ctcoshφ.
Express xandctin terms of x/prime,ct/primeandφand show that
x2−(ct)2=(x/prime)2−(ct/prime)2.
3.27 A closed barrel has as its curved surface that obtained by rotating about the
x-axis the part of the curve
y=a[2−cosh( x/a)]
lying in the range −b≤x≤b. Show that the total surface area Aof the barrel
is given by
A=πa[9a−8aexp(−b/a)+aexp(−2b/a)−2b].
115
COMPLEX NUMBERS AND HYPERBOLIC FUNCTIONS
3.28 The principal value of the logarithmic function of a complex variable is defined
to have its argument in the range −π<argz≤π.B yw r i t i n g z=t a n win terms
of exponentials show that
tan−1z=1
2iln
/1+iz
1−iz
/
.
Use this result to evaluate
tan−1
/
2√
3−3i
7
/!
.
3.9 Hints and answers
3.1 (a) 5 + 3 i;( b )−1−5i;( c )1 0+5 i;( d )2 /5+1 1 i/5; (e) 4; (f) 3 −4i;
(g) 5 + i[tan−1(4/3) + 2 nπ]; (h)±(2.521 + 0 .595i).
3.3 2 +√
3.
3.4 (a) Set t=t a n θwith−π/2<θ<π / 2. The equation becomes z−c=ρe2iθ.
The locus is a circle, centre c,r a d i u s ρ.
(b) Eliminate the tterm between xand y. Note that the coefficient of tis
proportional to Im( b/c). The locus is a straight line (Im k)[x−Re(a)] =
(Rek)[y−Im(a)], where k=borc.
3.5 (a) exp( −2y)cos2 x;( b )( s i n2 ysinh2 x)/2; (c)√
2e x p ( πi/3) or√
2e x p ( 4 πi/3);
(d) exp(1 /√
2) or exp(−1/√
2); (e) 0 .540−0.841i;( f )8s i n ( l n2 )=5 .11;
(g) exp(−π/2−2πn); (h) ln 8 + i(2n+1/2)π.
3.6 (a) y=(±√
2−1)x;( b ) x=±y;( c )t h eh a l fo ft h ec i r c l e( x−1
2)2+y2=1
4that
lies in y<0.
3.7 Starting from |x+iy−ia|=λ|x+iy+ia|, show that the coefficients of xandy
are equal, and write the equation in the form x2+(y−α)2=r2.
3.8 (a) arg[( b−a)/(c−a)] = arg[( B−A)/(C−A)].
(b)|(b−a)|/|(c−a)|=|(B−A)|/|(C−A)|.
3.9 (a) Circles enclosing z=−ia,w i t h λ=e x p c>1.
(b) The condition is that arg[( z−ia)/(z+ia)] = k. This can be rearranged to
givea(z+z∗)=k(a2−|z|2), which becomes in x, ycoordinates the equation
of a circle with centre ( −a/k,0) and radius a(1 +k−2)1/2.
3.10 (a) Z=(−dz+b)/(cz−a).
(b)|(Z−Z1)/(Z−Z2)|=Λ ,w i t h Z1,2given by setting z=z1,2in the result in
(a);|a−cz1|=λ|a−cz2|.
3.11 All three conditions are satisfied in 3 π/2≤θ≤7π/4,|z|≤4; area = 2 π.
3.12 (a) Express ωn−1 as a product of factors like ( ω−ωr
n) and examine the
coefficients of (i) ωn−1and (ii) ω0.
(b) ( x+ω3y+ω2
3z)(x+ω2
3y+ω3z).
3.13 Denoting exp[2 πi/(2m+ 1)] by Ω, express x2m+1−a2m+1as a product of factors
like ( x−aΩr) and then combine those containing Ωrand Ω2m+1−r.U s et h ef a c t
that Ω2m+1=1 .
3.14 (b) Differentiate ( z1−z2)(z∗
1−z∗
2). (c) Write 2 |z1|2+2|z2|2as|z1+z2|2−|z1−z2|2.
Circular motion about a fixed point with the vortices at the opposite ends of adiameter.
3.15 The roots are 2
1/3exp(2 πni/3) for n=0,1,2; 1±31/4;1±31/4i.
3.16 (a) The vanishing of the real and imaginary parts of f(λi)r e q u i r e s( λ2=3o r8
3)
and ( λ2= 0 or 3 or 12); hence λ2=3a n d f(z)=(z2+3)(z3−6z2+12z−16).
(b)a=−2,b=−8. The roots are ±i√
3, 4, 1±i√
3.
116
3.9 HINTS AND ANSWERS
3.17 (b) S2(n)=2n/2sin(nπ/4).S2(6) =−8,S2(7) =−8,S2(8) = 0.
3.18 Write 1 + cos θand sin θin terms of θ/2.
3.20 (cos4 θ)/8−(cos 2 θ)/2+3 /8.
3.21 Show that cos5 θ=1 6 c5−20c3+5c,w h e r e c=c o s θ, and correspondingly for
sin 5θ.U s ec o s−2θ=1+t a n2θ. The four required values are
[(5−√
20)/5]1/2,( 5−√
20)1/2,[ ( 5+√
20)/5]1/2,( 5+√
20)1/2.
3.23 Reality of the root(s) requires c2+b2≥a2anda+b>0. With these conditions,
there are two roots if a2>b2, but only one if b2>a2.
Fora2=c2+b2,x=1
2ln[(a−b)/(a+b)].
3.24 (a) ln(1 /√
3); (b) (1 +√
5)/2; (c)±(12)1/4/(√
3+1 ) .
3.25 Reduce the equation to 16sinh4x= 1, yielding x=±0.481.
3.26 The same expressions but with φreplaced by −φare obtained.
3.27 Show that ds=( c o s h x/a)dx;
curved surface area = πa2[8sinh( b/a)−sinh(2 b/a)]−2πab.
3.28 π/6−iln√
2.
117
4
Series and limits
4.1 Series
Many examples exist in the physical sciences of situations where we are presented
with a sum of terms to evaluate. For example, we may wish to add the contributions
from successive slits in a diffraction grating to find the total light intensity at a
particular point behind the grating.
A series may have either a finite or infinite number of terms. In either case, the
sum of the first Nterms of a series (often called a partial sum) is written
SN=u1+u2+u3+···+uN,
where the terms of the series un,n=1,2,3,...,N are numbers, that may in
general be complex. If the terms are complex then SNwill in general be complex
also, and we can write SN=XN+iYN,w h e r e XNandYNare the partial sums of
the real and imaginary parts of each term separately and are therefore real. If a
series has only Nterms then the partial sum SNis of course the sum of the series.
Sometimes we may encounter series where each term depends on some variable,x, say. In this case the partial sum of the series will depend on the value assumed
byx. For example, consider the infinite series
S(x)=1+ x+x
2
2!+x3
3!+···.
This is an example of a power series; these are discussed in more detail in
section 4.5. It is in fact the Maclaurin expansion of exp x(see subsection 4.6.3).
Therefore S(x)=e x p xand, of course, varies according to the value of the
variable x. A series might just as easily depend on a complex variable z.
A general, random sequence of numbers can be described as a series and a sum
of the terms found. However, for cases of practical interest, there will usually be
118
4.2 SUMMATION OF SERIES
some sort of relationship between successive terms. For example, if the nth term
o fas e r i e si sg i v e nb y
un=1
2n,
forn=1,2,3,...,N then the sum of the first Nterms will be
SN=Nsummationdisplay
n=1un=1
2+1
4+1
8+···+1
2N. (4.1)
It is clear that the sum of a finite number of terms is always finite, provided
that each term is itself finite. It is often of practical interest, however, to consider
the sum of a series with an infinite number of finite terms. The sum of an
infinite number of terms is best defined by first considering the partial sumof the first Nterms, S
N. If the value of the partial sum SNtends to a finite
limit, S,a sNtends to infinity, then the series is said to converge and its sum
is given by the limit S. In other words, the sum of an infinite series is given
by
S= lim
N→∞SN,
provided the limit exists. For complex infinite series, if SNapproaches a limit
S=X+iYasN→∞, this means that XN→XandYN→Yseparately, i.e.
the real and imaginary parts of the series are each convergent series with sumsXandYrespectively.
However, not all infinite series have finite sums. As N→∞, the value of the
partial sum S
Nmay diverge: it may approach + ∞or−∞, or oscillate finitely
or infinitely. Moreover, for a series where each term depends on some variable,its convergence can depend on the value assumed by the variable. Whether aninfinite series converges, diverges or oscillates has important implications whendescribing physical systems. Methods for d etermining whether a series converges
are discussed in section 4.3.
4.2 Summation of series
It is often necessary to find the sum of a finite series or a convergent infinite
series. We now describe arithmetic, geometric and arithmetico-geometric series,
which are particularly common and for which the sums are easily found. Other
methods that can sometimes be used to sum more complicated series are discussedbelow.
119
SERIES AND LIMITS
4.2.1 Arithmetic series
Anarithmetic series has the characteristic that the difference between successive
terms is constant. The sum of a general arithmetic series is written
SN=a+(a+d)+(a+2d)+···+[a+(N−1)d]=N−1summationdisplay
n=0(a+nd).
Rewriting the series in the opposite order and adding this term by term to the
original expression for SN, we find
SN=N
2[a+a+(N−1)d]=N
2(first term + last term) . (4.2)
If an infinite number of such terms are added the series will increase (or decrease)
indefinitely; that is to say, it diverges.ISum the integers between 1and1000inclusive.
This is an arithmetic series with a=1 ,d=1a n d N= 1000. Therefore, using (4.2) we find
SN=1000
2(1 + 1000) = 500500 ,
which can be checked directly only with considerable effort.
J
4.2.2 Geometric series
Equation (4.1) is a particular example of a geometric series , which has the
characteristic that the ratio of successive terms is a constant (one-half in thiscase). The sum of a geometric series is in general written
S
N=a+ar+ar2+···+arN−1=N−1summationdisplay
n=0arn,
where ais a constant and ris the ratio of successive terms, the common ratio .T h e
sum may be evaluated by considering SNandrSN:
SN=a+ar+ar2+ar3+···+arN−1,
rSN=ar+ar2+ar3+ar4+···+arN.
If we now subtract the second equation from the first we obtain
(1−r)SN=a−arN,
and hence
SN=a(1−rN)
1−r. (4.3)
120
4.2 SUMMATION OF SERIES
For a series with an infinite number of terms and |r|<1, we have lim N→∞rN=0 ,
and the sum tends to the limit
S=a
1−r. (4.4)
In (4.1), r=1
2,a=1
2,a n ds o S=1 .F o r|r|≥1, however, the series either diverges
or oscillates.IConsider a ball that drops from a height of 27 mand on each bounce retains only a third
of its kinetic energy; thus after one bounce it will return to a height of 9m,a f t e rt w o
bounces to 3m, and so on. Find the total distance travelled between the first bounce and
theMth bounce.
The total distance travelled between the first bounce and the Mth bounce is given by the
sum of M−1t e r m s :
SM−1=2(9+3+1+ ···)=2M−2X
m=09
3m
forM> 1, where the factor of 2 is included to allow for both the upward and the
downward journey. Inside the parentheses we clearly have a geometric series with firstterm 9 and common ratio 1 /3 and hence the distance is given by (4.3), i.e.
S
M−1=2×9
h
1−
/;1
3
/M−1
i
1−1
3=2 7
h
1−
/;1
3
/M−1
i
,
where the number of terms Nin (4.3) has been replaced by M−1.
J
4.2.3 Arithmetico-geometric series
An arithmetico-geometric series, as its name suggests, is a combined arithmetic
and geometric series. It has the general form
SN=a+(a+d)r+(a+2d)r2+···+[a+(N−1)d]rN−1=N−1summationdisplay
n=0(a+nd)rn,
and can be summed, in a similar way to a pure geometric series, by multiplying
byrand subtracting the result from the original series to obtain
(1−r)SN=a+rd+r2d+···+rN−1d−[a+(N−1)d]rN.
Using the expression for the sum of a geometric series (4.3) and rearranging, we
find
SN=a−[a+(N−1)d]rN
1−r+rd(1−rN−1)
(1−r)2.
For an infinite series with |r|<1, lim N→∞rN= 0 as in the previous subsection,
and the sum tends to the limit
S=a
1−r+rd
(1−r)2. (4.5)
As for a geometric series, if |r|≥1 then the series either diverges or oscillates.
121
SERIES AND LIMITSISum the series
S=2+5
2+8
22+11
23+···.
This is an infinite arithmetico-geometric series with a=2 , d=3a n d r=1/2. Therefore,
from (4.5), we obtain S= 10.
J
4.2.4 The difference method
The difference method is sometimes useful in summing series that are more
complicated than the examples discussed above. Let us consider the general series
Nsummationdisplay
n=1un=u1+u2+···+uN.
If the terms of the series, un, can be expressed in the form
un=f(n)−f(n−1)
for some function f(n) then its (partial) sum is given by
SN=Nsummationdisplay
n=1un=f(N)−f(0).
This can be shown as follows. The sum is given by
SN=u1+u2+···+uN
and since un=f(n)−f(n−1), it may be rewritten
SN=[f(1)−f(0)] + [ f(2)−f(1)] + ···+[f(N)−f(N−1)].
By cancelling terms we see that
SN=f(N)−f(0).IEvaluate the sum
NX
n=11
n(n+1 ).
Using partial fractions we find
un=−
/1
n+1−1
n
/
.
Hence un=f(n)−f(n−1) with f(n)=−1/(n+1 ) ,a n ds ot h es u mi sg i v e nb y
SN=f(N)−f(0) =−1
N+1+1=N
N+1.
J
122
4.2 SUMMATION OF SERIES
The difference method may be easily extended to evaluate sums in which each
term can be expressed in the form
un=f(n)−f(n−m), (4.6)
where mis an integer. By writing out the sum to Nterms with each term expressed
in this form, and cancelling terms in pairs as before, we find
SN=msummationdisplay
k=1f(N−k+1 )−msummationdisplay
k=1f(1−k).IEvaluate the sum
NX
n=11
n(n+2 ).
Using partial fractions we find
un=−
/1
2(n+2 )−1
2n
/
.
Hence un=f(n)−f(n−2) with f(n)=−1/[2(n+ 2)], and so the sum is given by
SN=f(N)+f(N−1)−f(0)−f(−1) =3
4−1
2
/1
N+2+1
N+1
/
.
J
In fact the difference method is quite flexible and may be used to evaluate
sums even when each term cannot be expressed as in (4.6). The method still relies,however, on being able to write u
nin terms of a single function such that most
terms in the sum cancel, leaving only a few terms at the beginning and the end.This is best illustrated by an example.IEvaluate the sum
NX
n=11
n(n+1 ) ( n+2 ).
Using partial fractions we find
un=1
2(n+2 )−1
n+1+1
2n.
Hence un=f(n)−2f(n−1) +f(n−2) with f(n)=1 /[2(n+ 2)]. If we write out the sum,
expressing each term unin this form, we find that most terms cancel and the sum is given
by
SN=f(N)−f(N−1)−f(0) + f(−1) =1
4+1
2
/1
N+2−1
N+1
/
.
J
123
SERIES AND LIMITS
4.2.5 Series involving natural numbers
Series consisting of the natural numbers 1, 2, 3, ..., or the square or cube of these
numbers, occur frequently and deserve a special mention. Let us first consider
the sum of the first Nnatural numbers,
SN=1+2+3+ ···+N=Nsummationdisplay
n=1n.
This is clearly an arithmetic series with first term a= 1 and common difference
d= 1. Therefore, from (4.2), SN=1
2N(N+1 ) .
Next, we consider the sum of the squares of the first Nnatural numbers:
SN=12+22+32+...+N2=Nsummationdisplay
n=1n2,
which may be evaluated using the difference method. The nth term in the series
isun=n2, which we need to express in the form f(n)−f(n−1) for some function
f(n). Consider the function
f(n)=n(n+ 1)(2 n+1 )⇒ f(n−1) = ( n−1)n(2n−1).
For this function f(n)−f(n−1) = 6 n2,a n ds ow ec a nw r i t e
un=1
6[f(n)−f(n−1)].
Therefore, by the difference method,
SN=1
6[f(N)−f(0)] =1
6N(N+ 1)(2 N+1 ).
Finally, we calculate the sum of the cubes of the first Nnatural numbers,
SN=13+23+33+···+N3=Nsummationdisplay
n=1n3,
again using the difference method. Consider the function
f(n)=[n(n+1 ) ]2⇒ f(n−1) = [( n−1)n]2,
for which f(n)−f(n−1) = 4 n3. Therefore we can write the general nth term of
the series as
un=1
4[f(n)−f(n−1)],
and using the difference method we find
SN=1
4[f(N)−f(0)] =1
4N2(N+1 )2.
Note that this is the square of the sum of the natural numbers, i.e.
Nsummationdisplay
n=1n3=parenleftBiggNsummationdisplay
n=1nparenrightBigg2
.
124
4.2 SUMMATION OF SERIESISum the series
NX
n=1(n+1 ) ( n+3 ).
Thenth term in this series is
un=(n+1 ) ( n+3 )= n2+4n+3,
and therefore we can write
NX
n=1(n+1 ) ( n+3 )=NX
n=1(n2+4n+3 )
=NX
n=1n2+4NX
n=1n+NX
n=13
=1
6N(N+ 1)(2 N+1 )+4×1
2N(N+1 )+3 N
=1
6N(2N2+1 5N+ 31) .
J
4.2.6 Transformation of series
A complicated series may sometimes be summed by transforming it into a
familiar series for which we already know the sum, perhaps a geometric series
or the Maclaurin expansion of a simple function (see subsection 4.6.3). Varioustechniques are useful, and deciding which one to use in any given case is a matterof experience. We now discuss a few of the more common methods.
The differentiation or integration of a series is often useful in transforming an
apparently intractable series into a more familiar one. If we wish to differentiateor integrate a series that already depends on some variable then we may do soin a straightforward manner.ISum the series
S(x)=x4
3(0!)+x5
4(1!)+x6
5(2!)+···.
Dividing both sides by xwe obtain
S(x)
x=x3
3(0!)+x4
4(1!)+x5
5(2!)+···,
which is easily differentiated to give
d
dx
/S(x)
x
/
=x2
0!+x3
1!+x4
2!+x5
3!+···.
Recalling the Maclaurin expansion of exp xgiven in subsection 4.6.3, we recognise that
the RHS is equal to x2expx. Having done so, we can now integrate both sides to obtain
S(x)/x=
Z
x2expxd x .
125
SERIES AND LIMITS
Integrating the RHS by parts we find
S(x)/x=x2expx−2xexpx+2e x p x+c,
where the value of the constant of integration ccan be fixed by the requirement that
S(x)/x=0a t x= 0. Thus we find that c=−2, and that the sum is given by
S(x)=x3expx−2x2expx+2xexpx−2x.
J
Often, however, we require the sum of a series that does not depend on a
variable. In this case, in order that we may differentiate or integrate the series,we define a function of some variable xsuch that the value of this function is
equal to the sum of the series for some particular value of x(usually at x=1 ) .ISum the series
S=1+2
2+3
22+4
23+···.
Let us begin by defining the function
f(x)=1+2 x+3x2+4x3+···,
so that the sum S=f(1/2). Integrating this function we obtainZ
f(x)dx=x+x2+x3+···,
which we recognise as an infinite geometric series with first term a=xand common ratio
r=x. Therefore, from (4.4), we find that the sum of this series is x/(1−x). In other wordsZ
f(x)dx=x
1−x,
so that f(x)i sg i v e nb y
f(x)=d
dx
/x
1−x
/
=1
(1−x)2.
The sum of the original series is therefore S=f(1/2) = 4.
J
Aside from differentiation and integration, an appropriate substitution can
sometimes transform a series into a more familiar form. In particular, series withterms that contain trigonometric functions can often be summed by the use ofcomplex exponentials.ISum the series
S(θ)=1+c o s θ+cos 2θ
2!+cos 3θ
3!+···.
Replacing the cosine terms with a complex exponential, we obtain
S(θ)=R e
/
1+e x p iθ+exp2 iθ
2!+exp 3 iθ
3!+···
/
=R e
/
1+e x p iθ+(expiθ)2
2!+(expiθ)3
3!+···
/
.
126
4.3 CONVERGENCE OF INFINITE SERIES
Again using the Maclaurin expansion of exp xgiven in subsection 4.6.3, we notice that
S(θ) = Re [exp(exp iθ)] = Re [exp(cos θ+isinθ)]
=R e|:{[exp(cos θ)][exp( isinθ)]}= [exp(cos θ)]Re [exp( isinθ)]
= [exp(cos θ)][cos(sin θ)].
J
4.3 Convergence of infinite series
Although the sums of some commonly occurring infinite series may be found,
the sum of a general infinite series is usually difficult to calculate. Nevertheless,it is often useful to know whether the partial sum of such a series converges toa limit, even if the limit cannot be found explicitly. As mentioned at the end of
section 4.1, if we allow Nto tend to infinity, the partial sum
S
N=Nsummationdisplay
n=1un
of a series may tend to a definite limit (i.e. the sum Sof the series), or increase
or decrease without limit, or oscillate finitely or infinitely.
To investigate the convergence of any given series, it is useful to have available
a number of tests and theorems of general applicability. We discuss them below;
some we will merely state, since once they have been stated they become almostself-evident, but are no less useful for that.
4.3.1 Absolute and conditional convergence
Let us first consider some general points concerning the convergence, or otherwise,
of an infinite series. In general an infinite seriessummationtextu
ncan have complex terms,
and in the special case of a real series the terms can be positive or negative. Fromany such series, however, we can always construct another seriessummationtext|u
n|in which
each term is simply the modulus of the corresponding term in the original series.Then each term in the new series will be a positive real number.
If the seriessummationtext|u
n|converges thensummationtextunalso converges, andsummationtextunis said to be
absolutely convergent , i.e. the series formed by the absolute values is convergent.
For an absolutely convergent series, the terms may be reordered without affectingthe convergence of the series. However, ifsummationtext|u
n|diverges whilstsummationtextunconverges
thensummationtextunis said to be conditionally convergent . For a conditionally convergent
series, rearranging the order of the terms can affect the behaviour of the sumand, hence, whether the series converges or diverges. In fact, a theorem dueto Riemann shows that, by a suitable rearrangement, a conditionally convergent
series may be made to converge to any arbitrary limit, or to diverge, or to oscillate
finitely or infinitely! Of course, if the original seriessummationtextu
nconsists only of positive
real terms and converges then automatically it is absolutely convergent.
127
SERIES AND LIMITS
4.3.2 Convergence of a series containing only real positive terms
As discussed above, in order to test for the absolute convergence of a seriessummationtextun, we first construct the corresponding seriessummationtext|un|that consists only of real
positive terms. Therefore in this subsection we will restrict our attention to seriesof this type.
We discuss below some tests that may be used to investigate the convergence of
such a series. Before doing so, however, we note the following crucial consideration .
In all the tests for, or discussions of, the convergence of a series, it is not whathappens in the first ten, or the first thousand, or the first million terms (or anyother finite number of terms) that matters, but what happens ultimately .
Preliminary test
A necessary but not sufficient condition for a series of real positive termssummationtextu
n
to be convergent is that the term untends to zero as ntends to infinity, i.e. we
require
lim
n→∞un=0.
If this condition is not satisfied then the series must diverge. Even if it is satisfied,
however, the series may still diverge, and further testing is required.
Comparison test
The comparison test is the most basic test for convergence. Let us consider two
seriessummationtextunandsummationtextvnand suppose that we know the latter to be convergent (by
some earlier analysis, for example). Then, if each term unin the first series is less
than or equal to the corresponding term vnin the second series, for all ngreater
than some fixed number Nwhich will vary from series to series, then the original
seriessummationtextunis also convergent. In other words, ifsummationtextvnis convergent and
un≤vnforn>N ,
thensummationtextunconverges.
However, ifsummationtextvndiverges and un≥vnfor all ngreater than some fixed number
thensummationtextundiverges.IDetermine whether the following series converges:
∞X
n=11
n!+1=1
2+1
3+1
7+1
25+···. (4.7)
Let us compare this series with the series
∞X
n=01
n!=1
0!+1
1!+1
2!+1
3!+···=2+1
2!+1
3!+···, (4.8)
128
4.3 CONVERGENCE OF INFINITE SERIES
which is merely the series obtained by setting x= 1 in the Maclaurin expansion of exp x
(see subsection 4.6.3), i.e.
exp(1) = e=1+1
1!+1
2!+1
3!+···.
Clearly this second series is convergent, since it consists of only positive terms and has a
finite sum. Thus, since each term unin the series (4.7) is less than the corresponding term
1/n! in (4.8), we conclude from the comparison test that (4.7) is also convergent.
J
D’Alembert’s ratio test
The ratio test determines whether a series converges by comparing the relative
magnitude of successive terms. If we consider a seriessummationtextunand set
ρ= lim
n→∞parenleftbiggun+1
unparenrightbigg
, (4.9)
then if ρ<1 the series is convergent; if ρ>1 the series is divergent; if ρ=1
then the behaviour of the series is undetermined by this test.
To prove this we observe that if the limit (4.9) is less than unity, i.e. ρ<1t h e n
we can find a value rin the range ρ<r< 1 and a value Nsuch that
un+1
un<r ,
for all n>N . Now the terms unof the series that follow uNare
uN+1,u N+2,u N+3, ...,
and each of these is less than the corresponding term of
ruN,r2uN,r3uN, ... . (4.10)
However, the terms of (4.10) are those of a geometric series with a common
ratio rthat is less than unity. This geometric series consequently converges and
therefore, by the comparison test discussed above, so must the original seriessummationtextun. An analogous argument may be used to prove the divergent case when
ρ>1.IDetermine whether the following series converges:
∞X
n=01
n!=1
0!+1
1!+1
2!+1
3!+···=2+1
2!+1
3!+···.
As mentioned in the previous example, this series may be obtained by setting x=1i nt h e
Maclaurin expansion of exp x, and hence we know already that it converges and has the
sum exp(1) = e. Nevertheless, we may use the ratio test to confirm that it converges.
Using (4.9), we have
ρ= lim
n→∞
/n!
(n+1 ) !
/
= lim
n→∞
/1
n+1
/
= 0 (4.11)
and since ρ<1, the series converges, as expected.
J
129
SERIES AND LIMITS
Ratio comparison test
As its name suggests, the ratio comparison test is a combination of the ratio and
comparison tests. Let us consider the two seriessummationtextunandsummationtextvnand assume that
we know the latter to be convergent. It may be shown that if
un+1
un≤vn+1
vn
for all ngreater than some fixed value Nthensummationtextunis also convergent.
Similarly, if
un+1
un≥vn+1
vn
for all sufficiently large n,a n dsummationtextvndiverges thensummationtextunalso diverges.IDetermine whether the following series converges:
∞X
n=11
(n!)2=1+1
22+1
62+···.
In this case the ratio of successive terms, as ntends to infinity, is given by
R= lim
n→∞
/n!
(n+1 ) !
/2
= lim
n→∞
/1
n+1
/2
,
which is less than the ratio seen in (4.11). Hence, by the ratio comparison test, the series
converges. (It is clear that this series could also be found to be convergent using the ratiotest.)J
Quotient test
The quotient test may also be considered as a combination of the ratio and
comparison tests. Let us again consider the two seriessummationtextunandsummationtextvn, and define
ρas the limit
ρ= lim
n→∞parenleftbiggun
vnparenrightbigg
. (4.12)
Then, it can be shown that:
(i) if ρ/negationslash= 0 but is finite thensummationtextunandsummationtextvneither both converge or both
diverge;
(ii) if ρ=0a n dsummationtextvnconverges thensummationtextunconverges;
(iii) if ρ=∞andsummationtextvndiverges thensummationtextundiverges.
130
4.3 CONVERGENCE OF INFINITE SERIESIGiven that the series
P∞
n=11/ndiverges, determine whether the following series converges:
∞X
n=14n2−n−3
n3+2n. (4.13)
If we set un=( 4n2−n−3)/(n3+2n)a n d vn=1/nthen the limit (4.12) becomes
ρ= lim
n→∞
/(4n2−n−3)/(n3+2n)
1/n
/
= lim
n→∞
/4n3−n2−3n
n3+2n
/
=4.
Since ρis finite but non-zero and
Pvndiverges, from (i) above
Punmust also diverge.
J
Integral test
The integral test is an extremely powerful means of investigating the convergence
of a seriessummationtextun. Suppose that there exists a function f(x) which monotonically
decreases for xgreater than some fixed value x0and for which f(n)=un,i . e .t h e
value of the function at integer values of xis equal to the corresponding term
in the series under investigation. Then it can be shown that, if the limit of the
integral
lim
N→∞integraldisplayN
f(x)dx
exists, the seriessummationtextunis convergent. Otherwise the series diverges. Note that the
integral defined here has no lower limit; the test is sometimes stated with lowerlimit of unity for the integral, but this can lead to unnecessary difficulties.IDetermine whether the following series converges:
∞X
n=11
(n−3/2)2=4+4+4
9+4
25+···.
Let us consider the function f(x)=(x−3/2)−2. Clearly f(n)=unandf(x) monotonically
decreases for x>3/2. Applying the integral test, we consider
lim
N→∞
ZN1
(x−3/2)2dx= lim
N→∞
/−1
N−3/2
/
=0.
Since the limit exists the series converges. Note, however, that if we had included a lower
limit of unity in the integral then we would have run into problems, since the integranddiverges at x=3/2.J
The integral test is also useful for examining the convergence of the Riemann
zeta series. This is a special series that occurs regularly and is of the form
∞summationdisplay
n=11
np.
It converges for p>1 and diverges if p≤1. These convergence criteria may be
derived as follows.
131
SERIES AND LIMITS
Using the integral test, we consider
lim
N→∞integraldisplayN1
xpdx= lim
N→∞parenleftbiggN1−p
1−pparenrightbigg
,
and it is obvious that the limit tends to zero for p>1a n dt o∞forp≤1.
Cauchy’s root test
Cauchy’s root test may be useful in testing for convergence, especially if the nth
terms of the series contains an nth power. If we define the limit
ρ= lim
n→∞(un)1/n,
then it may be proved that the seriessummationtextunconverges if ρ<1. If ρ>1 then the
series diverges. Its behaviour is undetermined if ρ=1 .IDetermine whether the following series converges:
∞X
n=1
/1
n
/n
=1+1
4+1
27+···.
Using Cauchy’s root test, we find
ρ= lim
n→∞
/1
n
/
=0,
and hence the series converges.
J
Grouping terms
We now consider the Riemann zeta series, mentioned above, with an alternative
proof of its convergence that uses the method of grouping terms. In general thereare better ways of determining convergence, but the grouping method may beused if it is not immediately obvious how to approach a problem by a better
method.
First consider the case where p>1 and group the terms in the series as follows:
S
N=1
1p+parenleftbigg1
2p+1
3pparenrightbigg
+parenleftbigg1
4p+···+1
7pparenrightbigg
+···.
Now we can see that each bracket of this series is less than each term of the
geometric series
SN=1
1p+2
2p+4
4p+···.
This geometric series has common ratio r=parenleftbig1
2parenrightbigp−1; therefore r<1s i n c e p>1,
and so the geometric series converges. Then the comparison test shows that theRiemann zeta series also converges for p>1.
132
4.3 CONVERGENCE OF INFINITE SERIES
The divergence of the Riemann zeta series for p≤1 can be seen by first
considering the case p= 1. The series is
SN=1+1
2+1
3+1
4+···,
which does notconverge, as may be seen by bracketing the terms of the series in
groups in the following way:
SN=Nsummationdisplay
n=1un=1+parenleftbigg1
2parenrightbigg
+parenleftbigg1
3+1
4parenrightbigg
+parenleftbigg1
5+1
6+1
7+1
8parenrightbigg
+···.
The sum of the terms in each bracket is ≥1
2and, since as many such groupings
can be made as we wish, it is clear that SNincreases indefinitely as Nis increased.
Now returning to the case of the Riemann zeta series for p<1, we note that
each term in the series is greater than the corresponding one in the series forwhich p=1 .I no t h e rw o r d s1 /n
p>1/nforn>1,p<1. The comparison test
then shows us that the Riemann zeta series will diverge for all p≤1.
4.3.3 Alternating series test
The tests discussed in the last subsection have been concerned with determining
whether the series of real positive termssummationtext|un|converges, and so whethersummationtextun
is absolutely convergent. Nevertheless, it is sometimes useful to consider whether
a series is merely convergent rather than absolutely convergent. This is especially
true for series containing an infinite number of both positive and negative terms.
In particular, we will consider the convergence of series in which the positive andnegative terms alternate, i.e. an alternating series .
An alternating series can be written as
∞summationdisplay
n=1(−1)n+1un=u1−u2+u3−u4+u5−···,
with all un≥0. Such a series can be shown to converge provided (i) un→0a s
n→∞and (ii) un<u n−1for all n>N for some finite N. If these conditions are
not met then the series oscillates.
To prove this, suppose for definiteness that Nis odd and consider the series
starting at uN. The sum of its first 2 mterms is
S2m=(uN−uN+1)+(uN+2−uN+3)+···+(uN+2m−2−uN+2m−1).
By condition (ii) above, all the parentheses are positive, and so S2mincreases as
mincreases. We can also write, however,
S2m=uN−(uN+1−uN+2)−···−(uN+2m−3−uN+2m−2)−uN+2m−1,
and since each parenthesis is positive, we must have S2m<u N. Thus, since S2m
133
SERIES AND LIMITS
is always less than uNfor all mandun→0a s n→∞, the alternating series
converges. It is clear that an analogous proof can be constructed in the casewhere Nis even.IDetermine whether the following series converges:
∞X
n=1(−1)n+11
n=1−1
2+1
3−···.
This alternating series clearly satisfies conditions (i) and (ii) above and hence converges.
However, as shown above by the method of grouping terms, the corresponding series withall positive terms is divergent.J
4.4 Operations with series
Simple operations with series are fairly intuitive, and we discuss them here only
for completeness. The following points apply to both finite and infinite seriesunless otherwise stated.
(i) Ifsummationtextu
n=Sthensummationtextkun=kSwhere kis any constant.
(ii) Ifsummationtextun=Sandsummationtextvn=Tthensummationtext(un+vn)=S+T.
(iii) Ifsummationtextun=Sthen a+summationtextun=a+S. A simple extension of this trivial result
shows that the removal or insertion of a finite number of terms anywherein a series does not affect its convergence.
(iv) If the infinite seriessummationtextu
nandsummationtextvnare both absolutely convergent then
the seriessummationtextwn,w h e r e
wn=u1vn+u2vn−1+···+unv1,
is also absolutely convergent. The seriessummationtextwnis called the Cauchy product
of the two original series. Furthermore, ifsummationtextunconverges to the sum S
andsummationtextvnconverges to the sum Tthensummationtextwnconverges to the sum ST.
(v) It is not true in general that term-by-term differentiation or integration of
a series will result in a new series with the same convergence properties.
4.5 Power series
A power series has the form
P(x)=a0+a1x+a2x2+a3x3+···,
where a0,a1,a2,a3etc. are constants. Such series regularly occur in physics and
engineering and are useful because, for |x|<1, the later terms in the series may
become very small and be discarded. For example the series
P(x)=1+ x+x2+x3+···,
134
4.5 POWER SERIES
although in principle infinitely long, in practice may be simplified if xhappens to
have a value small compared with unity. To see this note that P(x)f o r x=0.1
has the following values: 1, if just one term is taken into account; 1.1, for twoterms; 1.11, for three terms; 1.111, for four terms, etc. If the quantity that it
represents can only be measured with an accuracy of two decimal places, then all
but the first three terms may be ignored, i.e. when x=0.1o rl e s s
P(x)=1+ x+x
2+O ( x3)≈1+x+x2.
This sort of approximation is often used to simplify equations into manageable
forms. It may seem imprecise at first but is perfectly acceptable insofar as itmatches the experimental accuracy that can be achieved.
The symbols O and ≈used above need some further explanation. They are
used to compare the behaviour of two functions when a variable upon which
both functions depend tends to a particular limit, usually zero or infinity (and
obvious from the context). For two functions f(x)a n d g(x), with gpositive, the
formal definitions of the above symbols are as follows:
(i) If there exists a constant ksuch that|f|≤kgas the limit is approached
then f=O ( g).
(ii) If as the limit of xis approached f/gtends to a limit l,w h e r e l/negationslash=0 ,t h e n
f≈lg. The statement f≈gmeans that the ratio of the two sides tends
to unity.
4.5.1 Convergence of power series
The convergence or otherwise of power series is a crucial consideration in practical
terms. For example, if we are to use a power series as an approximation, it is
clearly important that it tends to the precise answer as more and more terms ofthe approximation are taken. Consider the general power series
P(x)=a
0+a1x+a2x2+···.
Using d’Alembert’s ratio test (see subsection 4.3.2), we see that P(x)c o n v e r g e s
absolutely if
ρ= lim
n→∞vextendsinglevextendsinglevextendsinglevextendsinglea
n+1
anxvextendsinglevextendsinglevextendsinglevextendsingle=|x|lim
n→∞vextendsinglevextendsinglevextendsinglevextendsinglea
n+1
anvextendsinglevextendsinglevextendsinglevextendsingle<1.
Thus the convergence of P(x) depends upon the value of x,i . e .t h e r ei s ,i ng e n e r a l ,
a range of values of xfor which P(x)c o n v e r g e s ,a n interval of convergence .N o t e
that at the limits of this range ρ= 1, and so the series may converge or diverge.
The convergence of the series at the end-points may be determined by substituting
these values of xinto the power series P(x) and testing the resulting series using
any applicable method (discussed in section 4.3).
135
SERIES AND LIMITSIDetermine the range of values of xfor which the following power series converges:
P(x)=1+2 x+4x2+8x3+···.
By using the interval-of-convergence method discussed above,
ρ= lim
n→∞
////2n+1
2nx
////=|2x|,
and hence the power series will converge for |x|<1/2. Examining the end-points of the
interval separately, we find
P(1/2 )=1+1+1+ ···,
P(−1/2) = 1−1+1−···.
Obviously P(1/2) diverges, while P(−1/2) oscillates. Therefore P(x) is not convergent at
either end-point of the region but is convergent for −1<x< 1.
J
The convergence of power series may be extended to the case where the
parameter zis complex. For the power series
P(z)=a0+a1z+a2z2+···,
we find that P(z)c o n v e r g e si f
ρ= lim
n→∞vextendsinglevextendsinglevextendsinglevextendsinglea
n+1
anzvextendsinglevextendsinglevextendsinglevextendsingle=|z|lim
n→∞vextendsinglevextendsinglevextendsinglevextendsinglea
n+1
anvextendsinglevextendsinglevextendsinglevextendsingle<1.
We therefore have a range in |z|for which P(z)c o n v e r g e s ,i . e . P(z)c o n v e r g e s
for values of zlying within a circle in the Argand diagram (in this case centred
on the origin of the Argand diagram). The radius of the circle is called the
radius of convergence :i fzlies inside the circle, the series will converge whereas
ifzlies outside the circle, the series will diverge; if, though, zlies on the circle
then the convergence must be tested using another method. Clearly the radius ofconvergence Ris given by 1 /R= lim
n→∞|an+1/an|.IDetermine the range of values of zfor which the following complex power series converges:
P(z)=1−z
2+z2
4−z3
8+···.
We find that ρ=|z/2|, which shows that P(z)c o n v e r g e sf o r |z|<2. Therefore the circle
of convergence in the Argand diagram is centred on the origin and has a radius R=2 .
On this circle we must test the conve rgence by substituting the value of zintoP(z)a n d
considering the resulting series. On the circle of convergence we can write z=2 e x p iθ.
Substituting this into P(z), we obtain
P(z)=1−2exp iθ
2+4e x p2 iθ
4−···
=1−expiθ+[ e x p iθ]2−···,
which is a complex infinite geometric series with first term a= 1 and common ratio
136
4.5 POWER SERIES
r=−expiθ. Therefore, on the the circle of convergence we have
P(z)=1
1+e x p iθ.
Unless θ=πthis is a finite complex number, and so P(z) converges at all points on the
circle|z|= 2 except at θ=π(i.e.z=−2), where it diverges. Note that P(z) is just the
binomial expansion of (1 + z/2)−1, for which it is obvious that z=−2 is a singular point.
In general, for power series expansions of complex functions about a given point in thecomplex plane, the circle of convergence extends as far as the nearest singular point. This
is discussed further in chapter 20.J
Note that the centre of the circle of convergence does not necessarily lie at the
origin. For example, applying the ratio test to the complex power series
P(z)=1+z−1
2+(z−1)2
4+(z−1)3
8+···,
we find that for it to converge we require |(z−1)/2|<1. Thus the series converges
forzlying within a circle of radius 2 centred on the point (1,0) in the Argand
diagram.
4.5.2 Operations with power series
The following rules are useful when manipulating power series; they apply to
power series in a real or complex variable.
( i )I ft w op o w e rs e r i e s P(x)a n d Q(x) have regions of convergence that overlap
to some extent then the series produced by taking the sum, the difference or theproduct of P(x)a n d Q(x) converges in the common region.
(ii) If two power series P(x)a n d Q(x) converge for all values of xthen one
series may be substituted into the other to give a third series, which also convergesf o ra l lv a l u e so f x. For example, consider the power series expansions of sin xand
e
xgiven below in subsection 4.6.3,
sinx=x−x3
3!+x5
5!−x7
7!+···
ex=1+ x+x2
2!+x3
3!+x4
4!+···,
both of which converge for all values of x. Substituting the series for sin xinto
that for exwe obtain
esinx=1+ x+x2
2!−3x4
4!−8x5
5!+···,
which also converges for all values of x.
If, however, either of the power series P(x)a n d Q(x) has only a limited region
of convergence, or if they both do so, then further care must be taken when
substituting one series into the other. For example, suppose Q(x)c o n v e r g e sf o r
allx, but P(x)o n l yc o n v e r g e sf o r xwithin a finite range. We may substitute
137
SERIES AND LIMITS
Q(x)i n t o P(x)t oo b t a i n P(Q(x)), but we must be careful since the value of Q(x)
may lie outside the region of convergence for P(x), with the consequence that the
resulting series P(Q(x)) does not converge.
(iii) If a power series P(x) converges for a particular range of xthen the series
obtained by differentiating every term and the series obtained by integrating every
term also converge in this range.
This is easily seen for the power series
P(x)=a0+a1x+a2x2+···,
which converges if |x|<limn→∞|an/an+1|≡k. The series obtained by differenti-
ating P(x) with respect to xis given by
dP
dx=a1+2a2x+3a3x2+···
and converges if
|x|<lim
n→∞vextendsinglevextendsinglevextendsinglevextendsinglenan
(n+1 )an+1vextendsinglevextendsinglevextendsinglevextendsingle=k.
Similarly the series obtained by integrating P(x) term by term,
integraldisplay
P(x)dx=a0x+a1x2
2+a2x3
3+···,
converges if
|x|<lim
n→∞vextendsinglevextendsinglevextendsinglevextendsingle(n+2 )an
(n+1 )an+1vextendsinglevextendsinglevextendsinglevextendsingle=k.
So, series resulting from differentiation or integration have the same interval of
convergence as the original series. However, even if the original series convergesat either end-point of the interval, it is not necessarily the case that the new serieswill do so. These new series must be tested separately at the end-points in orderto determine whether they converge there. Note that although power series maybe integrated or differentiated without altering their interval of convergence, thisis not true for series in general.
It is also worth noting that differentiating or integrating a power series term
by term within its interval of convergence is equivalent to differentiating orintegrating the function it represents. For example, consider the power seriesexpansion of sin x,
sinx=x−x
3
3!+x5
5!−x7
7!+···, (4.14)
which converges for all values of x. If we differentiate term by term, the series
becomes
1−x2
2!+x4
4!−x6
6!+···,
which is the series expansion of cos x,a sw ee x p e c t .
138
4.6 TAYLOR SERIES
4.6 Taylor series
Taylor’s theorem provides a way of expressing a function as a power series in x,
known as a Taylor series , but it can be applied only to those functions that are
continuous and differentiable within the x-range of interest.
4.6.1 Taylor’s theorem
Suppose that we have a function f(x) that we wish to express as a power series
inx−aabout the point x=a. We shall assume that, in a given x-range, f(x)
is a continuous, single-valued function of xhaving continuous derivatives with
respect to x, denoted by f/prime(x),f/prime/prime(x) and so on, up to and including f(n−1)(x). We
shall also assume that f(n)(x) exists in this range.
From the equation following (2.31) we may write
integraldisplaya+h
af/prime(x)dx=f(a+h)−f(a),
where a,a+hare neighbouring values of x. Rearranging this equation, we may
express the value of the function at x=a+hin terms of its value at aby
f(a+h)=f(a)+integraldisplaya+h
af/prime(x)dx. (4.15)
Afirst approximation forf(a+h) may be obtained by substituting f/prime(a)f o r
f/prime(x) in (4.15), to obtain
f(a+h)≈f(a)+hf/prime(a).
This approximation is shown graphically in figure 4.1. We may write this first
approximation in terms of xandaas
f(x)≈f(a)+(x−a)f/prime(a),
and, in a similar way,
f/prime(x)≈f/prime(a)+(x−a)f/prime/prime(a)
f/prime/prime(x)≈f/prime/prime(a)+(x−a)f/prime/prime/prime(a),
and so on. Substituting for f/prime(x) in (4.15), we obtain the second approximation :
f(a+h)≈f(a)+integraldisplaya+h
a[f/prime(a)+(x−a)f/prime/prime(a)]dx
≈f(a)+hf/prime(a)+h2
2f/prime/prime(a).
We may repeat this procedure as often as we like (so long as the derivatives
139
SERIES AND LIMITS
f(a)f(x)
a a+hxPQ
R
θhf/prime(a)
h
Figure 4.1 The first-order Taylor series approximation to a function f(x).
The slope of the function at P,i . e .t a n θ,e q u a l s f/prime(a). Thus the value of the
function at Q,f(a+h), is approximated by the ordinate of R,f(a)+hf/prime(a).
off(x) exist) to obtain higher-order approximations to f(a+h); we find the
(n−1)th-order approximation †to be
f(a+h)≈f(a)+hf/prime(a)+h2
2!f/prime/prime(a)+···+hn−1
(n−1)!f(n−1)(a). (4.16)
As might have been anticipated, the error associated with approximating f(a+h)
by this ( n−1)th-order power series is of the order of the next term in the series.
This error or remainder can be shown to be given by
Rn(h)=hn
n!f(n)(ξ),
for some ξthat lies in the range [ a, a+h]. Taylor’s theorem then states that we
may write the equality
f(a+h)=f(a)+hf/prime(a)+h2
2!f/prime/prime(a)+···+h(n−1)
(n−1)!f(n−1)(a)+Rn(h).
(4.17)
The theorem may also be written in a form suitable for finding f(x)g i v e n
the value of the function and its relevant derivatives at x=a, by substituting
†The order of the approximation is simply the highest power of hin the series. Note, though, that
the (n−1)th-order approximation contains nterms.
140
4.6 TAYLOR SERIES
x=a+hin the above expression. It then reads
f(x)=f(a)+(x−a)f/prime(a)+(x−a)2
2!f/prime/prime(a)+···+(x−a)n−1
(n−1)!f(n−1)(a)+Rn(x),
(4.18)
where the remainder now takes the form
Rn(x)=(x−a)n
n!f(n)(ξ),
andξlies in the range [ a, x]. Each of the formulae (4.17), (4.18) gives us the
Taylor expansion of the function about the point x=a. A special case occurs
when a= 0. Such Taylor expansions, about x= 0, are called Maclaurin series .
Taylor’s theorem is also valid without significant modification for functions
of a complex variable (see chapter 20). The extension of Taylor’s theorem tofunctions of more than one variable is given in chapter 5.
For a function to be expressible as an infinite power series we require it to be
infinitely differentiable and the remainder term R
nto tend to zero as ntends to
infinity, i.e. lim n→∞Rn= 0. In this case the infinite power series will represent the
function within the interval of convergence of the series.IExpand f(x)=s i n xas a Maclaurin series, i.e. about x=0.
We must first verify that sin xmay indeed be represented by an infinite power series. It is
easily shown that the nth derivative of f(x)i sg i v e nb y
f(n)(x)=s i n
/
x+nπ
2
/
.
Therefore the remainder after expanding f(x)a sa n( n−1)th-order polynomial about
x= 0 is given by
Rn(x)=xn
n!sin
/
ξ+nπ
2
/
,
where ξlies in the range [0 ,x]. Since the modulus of the sine term is always less than or
equal to unity, we can write |Rn(x)|<|xn|/n!. For any particular value of x,s a y x=c,
Rn(c)→0a s n→∞. Hence lim n→∞Rn(x) = 0, and so sin xcan be represented by an
infinite Maclaurin series.
Evaluating the function and its derivatives at x=0w eo b t a i n
f(0) = sin 0 = 0 ,
f/prime(0) = sin( π/2) = 1 ,
f/prime/prime(0) = sin π=0,
f/prime/prime/prime(0) = sin(3 π/2) =−1,
and so on. Therefore, the Maclaurin series expansion of sin xis given by
sinx=x−x3
3!+x5
5!−···.
Note that, as expected, since sin xis an odd function, its power series expansion contains
only odd powers of x.
J
141
SERIES AND LIMITS
We may follow a similar procedure to obtain a Taylor series about an arbitrary
point x=a.IExpand f(x)=c o s xas a Taylor series about x=π/3.
As in the above example, it is easily shown that the nth derivative of f(x)i sg i v e nb y
f(n)(x)=c o s
/
x+nπ
2
/
.
Therefore the remainder after expanding f(x)a sa n( n−1)th-order polynomial about
x=π/3i sg i v e nb y
Rn(x)=(x−π/3)n
n!cos
/
ξ+nπ
2
/
,
where ξlies in the range [ π/3,x]. The modulus of the cosine term is always less than or
equal to unity, and so |Rn(x)|<|(x−π/3)n|/n!. As in the previous example, lim n→∞Rn(x)=
0 for any particular value of x,a n ds oc o s xcan be represented by an infinite Taylor series
about x=π/3.
Evaluating the function and its derivatives at x=π/3w eo b t a i n
f(π/3) = cos( π/3) = 1 /2,
f/prime(π/3) = cos(5 π/6) =−√
3/2,
f/prime/prime(π/3) = cos(4 π/3) =−1/2,
and so on. Thus the Taylor series expansion of cos xabout x=π/3i sg i v e nb y
cosx=1
2−√
3
2
/;
x−π/3
/
−1
2
/;
x−π/3
/2
2!+···.
J
4.6.2 Approximation errors in Taylor series
In the previous subsection we saw how to represent a function f(x) by an infinite
power series, which is exactly equal to f(x)f o ra l l xwithin the interval of
convergence of the series. However, in physical problems we usually do not wantto have to sum an infinite number of terms, but prefer to use only a finite numberof terms in the Taylor series to approximate the function in some given range
ofx. In this case it is desirable to know what is the maximum possible error
associated with the approximation.
As given in (4.18), a function f(x) can be represented by a finite ( n−1)th-order
power series together with a remainder term such that
f(x)=f(a)+(x−a)f
/prime(a)+(x−a)2
2!f/prime/prime(a)+···+(x−a)n−1
(n−1)!f(n−1)(a)+Rn(x),
where
Rn(x)=(x−a)n
n!f(n)(ξ)
andξlies in the range [ a, x].Rn(x) is the remainder term, and represents the error
in approximating f(x)b yt h ea b o v e( n−1)th-order power series. Since the exact
142
4.6 TAYLOR SERIES
value of ξthat satisfies the expression for Rn(x) is not known, an upper limit on
the error may be found by differentiating Rn(x)w i t hr e s p e c tt o ξand equating
the derivative to zero in the usual way for finding maxima.IExpand f(x)=c o s xas a Taylor series about x=0and find the error associated with
using the approximation to evaluate cos(0 .5)if only the first two non-vanishing terms are
taken. (Note that the Taylor expansions of trigonometrical functions are only valid forangles measured in radians.)
Evaluating the function and its derivatives at x= 0, we find
f(0) = cos0 = 1 ,
f/prime(0) =−sin 0 = 0 ,
f/prime/prime(0) =−cos 0 =−1,
f/prime/prime/prime(0) = sin0 = 0 .
So, for small |x|, we find from (4.18)
cosx≈1−x2
2.
Note that since cos xis an even function, its power series expansion contains only even
powers of x. Therefore, in order to estimate the error in this approximation, we must
consider the term in x4, which is the next in the series. The required derivative is f(4)(x)
a n dt h i si s( b yc h a n c e )e q u a lt oc o s x. Thus, adding in the remainder term R4(x), we find
cosx=1−x2
2+x4
4!cosξ,
where ξlies in the range [0 ,x]. Thus, the maximum possible error is x4/4!, since cos ξ
cannot exceed unity. If x=0.5, taking just the first two terms yields cos(0 .5)≈0.875 with
a predicted error of less than 0 .00260. In fact cos(0 .5) = 0 .87758 to 5 decimal places. Thus,
to this accuracy, the true error is 0.00258, an error of about 0.3%.
J
4.6.3 Standard Maclaurin series
It is often useful to have a readily available table of Maclaurin series for standard
elementary functions, and therefore these are listed below.
sinx=x−x3
3!+x5
5!−x7
7!+···for−∞<x<∞,
cosx=1−x2
2!+x4
4!−x6
6!+···for−∞<x<∞,
tan−1x=x−x3
3+x5
5−x7
7+···for−1<x< 1,
ex=1+ x+x2
2!+x3
3!+x4
4!+···for−∞<x<∞,
ln(1 + x)=x−x2
2+x3
3−x4
4+···for−1<x≤1,
(1 +x)n=1+ nx+n(n−1)x2
2!+n(n−1)(n−2)x3
3!+···for−∞<x<∞.
143
SERIES AND LIMITS
These can all be derived by straightforward application of Taylor’s theorem to
the expansion of a function about x=0 .
4.7 Evaluation of limits
The idea of the limit of a function f(x)a sxapproaches a value ais fairly intuitive,
though a strict definition exists and is stated below. In many cases, the limit of
the function as xapproaches awill be simply the value f(a), but in others this is
not so. Firstly, the function may be undefined at x=a, as, for example, when
f(x)=sinx
x,
which takes the value 0 /0a t x= 0. However, the limit as xapproaches zero
does exist and can be evaluated as unity using l’H ˆopital’s rule below. Another
possibility is that even if f(x) is defined at x=aits value may not be equal to the
limiting value lim x→af(x). This can occur for a discontinuous function at a point
of discontinuity. The strict definition of a limit is that iflimx→af(x)=lthen
for any number /epsilon1however small, it must be possible to find a number ηsuch that
|f(x)−l|</epsilon1whenever|x−a|<η.In other words, as xbecomes arbitrarily close to
a,f(x) becomes arbitrarily close to its limit, l. To remove any ambiguity, it should
be stated that, in general, the number ηwill depend on both /epsilon1and the form of f(x).
The following observations are often useful in finding the limit of a function.
(i) A limit may be ±∞. For example as x→0, 1/x2→∞ .
(ii) A limit may be approached from below or above and the value may be
different in each case. For example consider the function f(x)=t a n x.A sxtends
toπ/2f r o mb e l o w f(x)→∞, but if the limit is approached from above then
f(x)→−∞ . Another way of writing this is
lim
x→π
2−tanx=∞, lim
x→π
2+tanx=−∞.
(iii) It may ease the evaluation of limits if the function under consideration is
split into a sum, product or quotient. Provided each of the limits exists, the rulesfor evaluating such limits are as follows.
(a) lim
x→a{f(x)+g(x)}= lim
x→af(x) + lim
x→ag(x).
(b) lim
x→a{f(x)g(x)}= lim
x→af(x) lim
x→ag(x).
(c) lim
x→af(x)
g(x)=limx→af(x)
limx→ag(x), provided that
the numerator and denominator are
not both equal to zero or infinity.
Examples of cases (a)–(c) are discussed below.
144
4.7 EVALUATION OF LIMITSIEvaluate the limits
lim
x→1(x2+2x3),lim
x→0(xcosx), lim
x→π/2sinx
x.
Using (a) above,
lim
x→1(x2+2x3) = lim
x→1x2+ lim
x→12x3=3.
Using (b),
lim
x→0(xcosx) = lim
x→0xlim
x→0cosx=0×1=0 .
Using (c),
lim
x→π/2sinx
x=limx→π/2sinx
limx→π/2x=1
π/2=2
π.
J
(iv) Limits of functions of xthat contain exponents that themselves depend on
xcan often be found by taking logarithms.IEvaluate the limit
lim
x→∞
/
1−a2
x2
/x2
.
Let us define
y=
/
1−a2
x2
/x2
and consider the logarithm of the required limit, i.e.
lim
x→∞lny= lim
x→∞
/
x2ln
/
1−a2
x2
//
.
Using the Maclaurin series for ln(1 + x) given in subsection 4.6.3, we can expand the
logarithm as a series and obtain
lim
x→∞lny= lim
x→∞
/
x2
/
−a2
x2−a4
2x4+···
//
=−a2.
Therefore, since lim x→∞lny=−a2it follows that lim x→∞y=e x p (−a2).
J
(v) L’H ˆopital’s rule may be used; it is an extension of (iii)(c) above. In cases
where both numerator and denominator are zero or both are infinite, furtherconsideration of the limit must follow. Let us first consider lim
x→af(x)/g(x),
where f(a)=g(a) = 0. Expanding the numerator and denominator as Taylor
series we obtain
f(x)
g(x)=f(a)+(x−a)f/prime(a)+[ ( x−a)2/2!]f/prime/prime(a)+···
g(a)+(x−a)g/prime(a)+[ ( x−a)2/2!]g/prime/prime(a)+···.
However, f(a)=g(a)=0s o
f(x)
g(x)=f/prime(a)+[ ( x−a)/2!]f/prime/prime(a)+···
g/prime(a)+[ ( x−a)/2!]g/prime/prime(a)+···.
145
SERIES AND LIMITS
Therefore we find
lim
x→af(x)
g(x)=f/prime(a)
g/prime(a),
provided f/prime(a)a n d g/prime(a) are not themselves both equal to zero. If, however,
f/prime(a)a n d g/prime(a)areboth zero then the same process can be applied to the ratio
f/prime(x)/g/prime(x) to yield
lim
x→af(x)
g(x)=f/prime/prime(a)
g/prime/prime(a),
provided that at least one of f/prime/prime(a)a n d g/prime/prime(a) is non-zero. If the original limit does
exist then it can be found by repeating the process as many times as is necessaryfor the ratio of corresponding nth derivatives not to be of the indeterminate form
0/0, i.e.
lim
x→af(x)
g(x)=f(n)(a)
g(n)(a).IEvaluate the limit
lim
x→0sinx
x.
We first note that if x= 0, both numerator and denominator are zero. Thus we apply
l’Hˆopital’s rule: differentiating, we obtain
lim
x→0(sinx/x) = lim
x→0(cosx/1) = 1 .
J
So far we have only considered the case where f(a)=g(a)=0 .F o rt h ec a s e
where f(a)=g(a)=∞we may still apply l’H ˆopital’s rule by writing
lim
x→af(x)
g(x)= lim
x→a1/g(x)
1/f(x),
w h i c hi sn o wo ft h ef o r m0 /0a t x=a. Note also that l’H ˆopital’s rule is still
valid for finding limits as x→∞,i . e .w h e n a=∞. This is easily shown by letting
y=1/xas follows:
lim
x→∞f(x)
g(x)= lim
y→0f(1/y)
g(1/y)
= lim
y→0−f/prime(1/y)/y2
−g/prime(1/y)/y2
= lim
y→0f/prime(1/y)
g/prime(1/y)
= lim
x→∞f/prime(x)
g/prime(x).
146
4.8 EXERCISES
Summary of methods for evaluating limits
To find the limit of a continuous function f(x) at a point x=a, simply substitute
the value ainto the function noting that0
∞=0a n dt h a t∞
0=∞.T h eo n l y
difficulty occurs when either of the expressions0
0or∞
∞results. In this case
differentiate top and bottom and try again. Continue differentiating until the topand bottom limits are no longer both zero or both infinity. If the undeterminedform 0×∞occurs then it can always be rewritten as
0
0or∞
∞.
4.8 Exercises
4.1 Sum the even numbers between 1000 and 2000 inclusive.
4.2 If you invest £1000 on the first day of each year, and interest is paid at 5% on
your balance at the end of each year, how much money do you have after 25
years?
4.3 How does the convergence of the series
∞X
n=r(n−r)!
n!
depend on the integer r?
4.4 Show that for testing the convergence of of the series
x+y+x2+y2+x3+y3+···,
where 0 <x<y< 1, the D’Alembert ratio test fails but the Cauchy root test is
successful.
4.5 Find the sum SNof the first Nterms of the following series, and hence determine
whether the series are convergent, divergent, or oscillatory:
(a)∞X
n=1ln
/n+1
n
/
,(b)∞X
n=0(−2)n,(c)∞X
n=1(−1)n+1n
3n.
4.6 By grouping and rearranging terms of the absolutely convergent series
S=∞X
n=11
n2,
show that
So=∞X
nodd1
n2=3S
4.
4.7 Use the difference method to sum the series
NX
n=22n−1
2n2(n−1)2.
147
SERIES AND LIMITS
4.8 The N+ 1 complex numbers ωmare given by ωm=e x p ( 2 πim/N )f o r m=
0,1,2,... ,N .
(a) Evaluate the following:
(i)NX
m=0ωm,(ii)NX
m=0ω2
m,(iii)NX
m=0ωmxm.
(b) Use these results to evaluate
(i)NX
m=0
/
cos
/2πm
N
/
−cos
/4πm
N
//
,(ii)3X
m=02msin
/2πm
3
/
.
4.9 Prove that
cosθ+c o s ( θ+α)+···+c o s ( θ+nα)=sin1
2(n+1 )α
sin1
2αcos(θ+1
2nα).
4.10 Determine whether the following series converge ( θand pare positive real
numbers):
(a)∞X
n=12si nnθ
n(n+1 ),(b)∞X
n=12
n2,(c)∞X
n=11
2n1/2,
(d)∞X
n=2(−1)n(n2+1 )1/2
nlnn,(e)∞X
n=1np
n!.
4.11 Find the real values of xfor which the following are series convergent:
(a)∞X
n=1xn
n+1,(b)∞X
n=1(sinx)n,(c)∞X
n=1nx,
(d)∞X
n=1enx,(e)∞X
n=2(lnn)x.
4.12 Determine whether the following series are convergent:
(a)∞X
n=1n1/2
(n+1 )1/2,(b)∞X
n=1n2
n!,(c)∞X
n=1(lnn)n
nn/2,(d)∞X
n=1nn
n!.
4.13 Determine whether the following series are absolutely convergent, convergent or
oscillatory:
(a)∞X
n=1(−1)n
n5/2,(b)∞X
n=1(−1)n(2n+1 )
n,(c)∞X
n=0(−1)n|x|n
n!,
(d)∞X
n=0(−1)n
n2+3n+2,(e)∞X
n=1(−1)n2n
n1/2.
4.14 Determine the positive values of xfor which the following series converges:
∞X
n=1xn/2e−n
n.
148
4.8 EXERCISES
4.15 Prove that
∞X
n=2ln
/nr+(−1)n
nr
/
is absolutely convergent for r= 2, but only conditionally convergent for r=1 .
4.16 An extension to the proof of the integral test (subsection 4.3.2) shows that, if f(x)
is positive, continuous and monotonically decreasing, for x≥1, and the series
f(1) + f(2) +···is convergent, then its sum does not exceed f(1) + L,w h e r e L
is the integralZ∞
1f(x)dx.
Use this result to show that the sum ζ(p) of the Riemann zeta series
Pn−p,w i t h
p>1, is not greater than p/(p−1).
4.17 Demonstrate that rearranging the order of its terms can make a condition-
ally convergent series converge to a different limit by considering the seriesP(−1)n+1n−1=l n2=0 .693. Rearrange the series as
S=1
1+1
3−1
2+1
5+1
7−1
4+1
9+1
11−1
6+1
13+···
and group each set of three successive terms. Show that the series can then be
written
∞X
m=18m−3
2m(4m−3)(4m−1),
which is convergent (by comparison with
Pn−2) and contains only positive
terms. Evaluate the first of these and hence deduce that Sis not equal to ln 2.
4.18 Illustrate result (iv) of section 4.4 about Cauchy products by considering the
double summation
S=∞X
n=1nX
r=11
r2(n+1−r)3.
By examining the points in the nr-plane over which the double summation is to
be carried out, show that Scan be written as
S=∞X
n=r∞X
r=11
r2(n+1−r)3.
Deduce that S≤3.
4.19 A Fabry–P ´erot interferometer consists of two parallel heavily silvered glass plates;
light enters normally to the plates, an d undergoes repeated reflections between
them, with a small transmitted fraction emerging at each reflection. Find theintensity|B|
2of the emerging wave, where
B=A(1−r)∞X
n=0rneinφ,
with randφreal.
149
SERIES AND LIMITS
4.20 Identify the series
∞X
n=1(−1)n+1x2n
(2n−1)!,
and then by integration and differentiation deduce the values Sof the following
series,
(a)∞X
n=1(−1)n+1n2
(2n)!,( b )∞X
n=1(−1)n+1n
(2n+1 ) !,
(c)∞X
n=1(−1)n+1nπ2n
4n(2n−1)!,( d )∞X
n=0(−1)n(n+1 )
(2n)!.
4.21 Starting from the Maclaurin series for cos x, show that
(cosx)−2=1+ x2+2x4
3+···.
Deduce the first three terms in the Maclaurin series for tan x.
4.22 Find the Maclaurin series for
(a)l n
/1+x
1−x
/
,(b)(x2+4 )−1,(c)s i n2x.
4.23 If f(x)=s i n h−1x,a n di t s nth derivative f(n)(x) is written as Pn(x)/(1 +x2)n−1/2,
where Pn(x) is a polynomial (of order n−1), show that the Pn(x)s a t i s f yt h e
recurrence relation
Pn+1(x)=( 1+ x2)P/prime
n(x)−(2n−1)xPn(x).
Hence generate the coefficients necessary to express sinh−1xas a Maclaurin series
up to terms in x5.
4.24 Find the first three non-zero terms in the Maclaurin series for the following
functions:
(a) (x2+9 )−1/2,(b) ln[(2 + x)3], (c) exp(sin x),
(d) ln(cos x), (e) exp[−(x−a)−2],(f) tan−1x.
4.25 By using the logarithmic series, prove that if aandbare positive and nearly
equal then
lna
b/similarequal2(a−b)
a+b.
Show that the error in this approximation is about 2( a−b)3/[3(a+b)3].
4.26 Determine whether the following functions f(x) are (i) continuous, and (ii)
differentiable at x=0 :
(a)f(x)=e x p (−|x|);
(b)f(x)=( 1−cosx)/x2forx/negationslash=0 , f(0) =1
2;
(c)f(x)=xsin(1/x)f o r x/negationslash=0 , f(0) = 0;
(d)f(x)=[ 4−x2], where [ y] denotes the integer part of y.
4.27 Find the limit as x→0o f[√
1+xm−√
1−xm]/xn,i nw h i c h mandnare positive
integers.
4.28 Evaluate the following limits:
(a) lim
x→0sin3x
sinhx, (b) lim
x→0tanx−tanh x
sinhx−x,
(c) lim
x→0tanx−x
cosx−1, (d) lim
x→0
/cosec x
x3−sinhx
x5
/
.
150
4.8 EXERCISES
4.29 Find the limits of the following functions:
(a)x3+x2−5x−2
2x3−7x2+4x+4,a s x→0,x→∞andx→2;
(b)sinx−xcoshx
sinhx−x,a s x→0;
(c)
Zπ/2
x
/ycosy−siny
y2
/
dy,a s x→0.
4.30 Use Taylor expansions to three terms to find approximations to (a)4√
17, and
(b)3√
26.
4.31 Using a first-order Taylor expansion about x=x0, show that a better approxi-
mation than x0to the solution of the equation
f(x)=s i n x+t a n x=2
is given by x=x0+h,w h e r e
h=2−f(x0)
cosx0+s e c2x0.
(a) Use this procedure twice to find the solution of f(x) = 2 to six significant
figures, given that it is close to x=0.9.
(b) Use the result in (a) to deduce, to the same degree of accuracy, one solution
of the quartic equation
y4−4y3+4y2+4y−4=0 .
4.32 Evaluate
lim
x→0
/1
x3
/
cosec x−1
x−x
6
//
.
4.33 In quantum theory, a system of osc illators, each of fundamental frequency ν,
interacting at temperature Thas an average energy ¯Egiven by
¯E=
P∞
n=0nhνe−nxP∞
n=0e−nx,
where x=hν/kT ,handkbeing the Planck and Boltzmann constants respectively.
Prove that both series converge, evaluate their sums, and show that at hightemperatures ¯E≈kTwhilst at low temperatures ¯E≈hνexp(−hν/kT ).
4.34 In a very simple model of a crystal, point-like atomic ions are regularly spaced
along an infinite one-dimensional row with spacing R. Alternate ions carry equal
and opposite charges ±e. The potential energy of the ith ion in the electric field
due to the jth ion is
q
iqj
4π/epsilon10rij,
where qkis the charge on the kth ion and rijis the distance between the ith and
jth ions.
Write down a series giving the total contribution Viof the ith ion to the overall
potential energy . Show that the series converges, and, if Viis written as
Vi=αe2
4π/epsilon10R,
find a closed-form expression for α, the Madelung constant for this (unrealistic)
lattice.
151
SERIES AND LIMITS
4.35 One of the factors contributing to the hi gh relative permittivi ty of water to static
electric fields is the permanent electric dipole moment pof the water molecule. In
an external field Ethe dipoles tend to line up with the field, but they do not do
so completely because of thermal agitation at the temperature Tof the water. A
classical (non-quantum) calculation using the Boltzmann distribution shows thatthe average polarisability per molecule αis given by
α=p
E(coth x−x−1),
where x=pE/kT andkis the Boltzmann constant.
At ordinary temperatures, even with high field strengths (104Vm−1or more),
x/lessmuch1. By making suitable series expansions of the hyperbolic functions involved,
show that α=p2/3kTto an accuracy of about one part in 15 x−2.
4.36 In quantum theory a certain method (the Born approximation) gives the (so-
called) amplitude f(θ) for the scattering of a particle of mass mthrough an angle
θby a uniform potential well of depth V0and radius b(i.e. the potential energy
of the particle is −V0within a sphere of radius band zero elsewhere) as
f(θ)=2mV0/~2K3(sinKb−KbcosKb).
Here /~is the Planck constant divided by 2 π, the energy of the particle is /~2k2/2m
andKis 2ksin(θ/2).
Use l’H ˆopital’s rule to evaluate the amplitude at low energies, i.e. when kand
hence Ktend to zero, and so determine the low-energy total cross-section.
(Note: the differential cross-section is given by |f(θ)|2a n dt h et o t a lc r o s s - s e c t i o n
by the integral of this over all solid angles, i.e. 2 π
Rπ
0|f(θ)|2sinθd θ.)
4.9 Hints and answers
4.1 2
P1000
500n= 751500.
4.2 Ar(rn−1)/(r−1) =£50 113 .
4.3 Divergent for r≤1; convergent for r≥2.
4.4 The ratio of successive terms oscillates between 0 and ∞asn→∞;un≤
(yn/2)1/n<1.
4.5 (a) ln( N+ 1), divergent; (b) [1 −(−2)n]/3, oscillates infinitely; (c) Add SN/3t o
theSNseries;3
16[1−(−3)−N]+3
4N(−3)−N−1,c o n v e r g e n tt o3
16.
4.6 Write all terms of the form (2 m)−2as1
4m−2; their sum is clearly1
4S.
4.7 (1 −N−2)/2.
4.8 (a) (i) 2 for N= 1; 1 otherwise. (ii) 2 for N=1 ;3f o r N= 2; 1 otherwise.
(iii) 1 + xforN=1 ;[ 1−xN+1exp(2 πi/N)]/[1−xexp(2 πi/N)] otherwise.
(b) (i) Consider Re( ωm−ω2
m);−2f o r N= 2; 0 otherwise. (ii) Consider Im(2mωm);
−√
3.
4.9 Sum the geometric series with rth term exp[ i(θ+rα)]. Its real part is
{cosθ−cos[(n+1 )α+θ]−cos(θ−α)+c o s ( θ+nα)}/4sin2(α/2),
which can be reduced to the given answer.
4.10 (a) Convergent, compare with
Pn−1(n+1 )−1; (b) convergent, ratio test; (c)
divergent, compare with
Pn−1; (d) convergent, alternating signs; (e) convergent,
ratio test.
152
4.9 HINTS AND ANSWERS
4.11 (a) −1≤x<1; (b) all xexcept x=( 2n±1)π/2; (c) x<−1; (d) x<0; (e)
always divergent. Clearly divergent for x>−1. For−X=x<−1, consider
∞X
k=1MkX
n=Mk−1+11
(lnMk)X,
where ln Mk=kand note that Mk−Mk−1=e−1(e−1)Mk; hence show that the
series diverges.
4.12 (a) Divergent, undoes not tend to 0. (b) Convergen t, ratio test. (c) Convergent,
root test. (d) Divergent, ratio tends to e,o rundoes not tend to 0.
4.13 (a) Absolutely convergent, compare wit h exercise 4.10(b). (b) Oscillates infinitely.
(c) Absolutely convergent for all x. (d) Absolutely convergent; use partial frac-
tions. (e) Oscillates infinitely.
4.14 x<e2, by the root test.
4.15 Divide the series into two series, nodd and neven. For r= 2 both are absolutely
convergent, by comparison with
Pn−2.F o r r=1n e i t h e rs e r i e si sc o n v e r g e n t ,
by comparison with
Pn−1. However, the sum of the two is convergent, by the
alternating sign test or by showing that the terms cancel in pairs.
4.17 The first term has value 0.833 and all other terms are positive.4.18 The original summation ran along lines parallel to the r-axis; replace it with
one running along lines parallel to the n-axis. Write n+1−r=sand deduce
thatS=ζ(2)ζ(3), where ζ(p) is the Riemann zeta functionPn−p.U s et h er e s u l t
proved in exercise 4.16 to give the stated conclusion.
4.19|A|2(1−r)2/(1 +r2−2rcosφ).
4.20 xsinx.
(a) Differentiate once; set x=1 . S=( s i n1+c o s1 ) /4=0 .345.
(b) Integrate once; set x=1 . S=( s i n1−cos 1) /2=0 .151.
(c) Differentiate once; set x=π/2.S=π/4=0 .785.
(d) Differentiate twice; set s=n−1a n d x=1 . S=( 2c o s1−sin1) /2=0 .120.
4.21 Use the binomial expansion and collect terms up to x4. Integrate both sides of
the displayed equation. tan x=x+x3/3+2 x5/15 +···.
4.22 (a)
X
nodd2xn
n;( b )∞X
n=0(−1)n
4
/x
2
/2n
;( c )∞X
n=1(−1)n+1(2x)2n
2(2n)!.
4.23 For example, P5(x)=2 4 x4−72x2+9 .s i n h−1x=x−x3/6+3 x5/40−···.
4.24 (a) [1 −(x2/18)−(3x4/648)] /3. (b) ln 8 + 3 x/2−3x2/8. (c) 1 + x+x2/2. (d)
−x2/2−x4/12−x6/45. (e) exp(−a−2){1−2x/a3−x2[(3/a4)−(2/a6)]}.( f )x−
x3/3+x5/5.
4.25 Set a=D+δandb=D−δand use the expansion for ln(1 ±δ/D).
4.26 (i) (a), (b) and (c) are continuous. (ii) Only (b) is differentiable.4.27 The limit is 0 for m>n ,1f o r m=n,a n d∞form<n .
4.28 (a) 3, (b) 4, (c) 0, (d)
1
90.
4.29 (a) −1
2,1
2,∞;( b )−4; (c)−1+2 /π.
4.30 (a) Expand f(x)=x1/4about x0= 16; approximation 2.030518, actual 2.030543.
(b) Expand f(x)=x1/3about x0= 27; approximation 2.962506, actual 2.962496.
4.31 (a) First approximation 0.886452; second approximation 0.886287. (b) Set y=
sinxand re-express f(x) = 2 as a polynomial equation. y= sin(0 .886287) =
0.774730.
4.32 7 /360.
4.33 E=hν[exp( hν/kT )−1]−1.
4.34 α=−2ln2.
4.36 f(θ)=2 mV0b3/3 /~2(i.e. independent of θ); 4π(2mV0b3/3 /~2)2.
153
5
Partial differentiation
In chapter 2, we discussed functions fof only one variable x, which were usually
written f(x). Certain constants and parameters may also have appeared in the
definition of f,e . g . f(x)=ax+2 contains the constant 2 and the parameter a, but
only xwas considered as a variable and only the derivatives f(n)(x)=dnf/dxn
were defined.
However, we may equally well consider functions that depend on more than one
variable, e.g. the function f(x, y)=x2+3xy, which depends on the two variables
xandy. For any pair of values x, y, the function f(x, y) has a well-defined value,
e.g.f(2,3) = 22. This notion can clearly be extended to functions dependent on
more than two variables. For the n-variable case, we write f(x1,x2,...,x n)f o r
a function that depends on the variables x1,x2,...,x n.W h e n n=2 , x1andx2
correspond to the variables xandyused above.
Functions of one variable, like f(x), can be represented by a graph on a
plane sheet of paper, and it is apparent that functions of two variables can,with little effort, be represented by a surface in three-dimensional space. Thus,
we may also picture f(x, y) as describing the variation of height with position
in a mountainous landscape. Functions of many variables, however, are usuallyvery difficult to visualise and so the preliminary discussion in this chapter willconcentrate on functions of just two variables.
5.1 Definition of the partial derivative
It is clear that a function f(x, y) of two variables will have a gradient in all
directions in the xy-plane. A general expression for this rate of change can be
found and will be discussed in the next section. However, we first consider the
simpler case of finding the rate of change of f(x, y) in the positive x-a n d y-
directions. These rates of change are called the partial derivatives with respect
154
5.1 DEFINITION OF THE PARTIAL DERIVATIVE
toxandyrespectively, and they are extremely important in a wide range of
physical applications.
For a function of two variables f(x, y) we may define the derivative with respect
tox, for example, by saying that it is that for a one-variable function when yis
held fixed and treated as a constant. To signify that a derivative is with respect
tox, but at the same time to recognize that a derivative with respect to yalso
exists, the former is denoted by ∂f/∂x and is the partial derivative of f(x, y)with
respect to x. Similarly, the partial derivative of fwith respect to yis denoted by
∂f/∂y .
To define formally the partial derivative of f(x, y) with respect to x, we have
∂f
∂x= lim
∆x→0f(x+∆x, y)−f(x, y)
∆x, (5.1)
provided that the limit exists. This is much the same as for the derivative of a
one-variable function. The other partial derivative of f(x, y) is similarly defined
as a limit (provided it exists):
∂f
∂y= lim
∆y→0f(x, y+∆y)−f(x, y)
∆y. (5.2)
It is common practice in connection with partial derivatives of functions
involving more than one variable to indicate those variables that are held constant
by writing them as subscripts to the derivative symbol. Thus, the partial derivatives
defined in (5.1) and (5.2) would be written respectively as
parenleftbigg∂f
∂xparenrightbigg
yandparenleftbigg∂f
∂yparenrightbigg
x.
In this form, the subscript shows explicitly which variable is to be kept constant.
A more compact notation for these partial derivatives is fxandfy. However, it is
extremely important when using partial derivatives to remember which variablesare being held constant and it is wise to write out the partial derivative in explicit
form if there is any possibility of confusion.
The extension of the definitions (5.1), (5.2) to the general n-variable case is
straightforward and can be formally written as
∂f(x
1,x2,...,x n)
∂xi= lim
∆xi→0[f(x1,x2,...,x i+∆xi,...,x n)−f(x1,x2,...,x i,...,x n)]
∆xi,
provided that the limit exists.
Just as for one-variable functions, second (and higher) partial derivatives may
be defined in a similar way. For a two-variable function f(x, y)t h e ya r e
∂
∂xparenleftbigg∂f
∂xparenrightbigg
=∂2f
∂x2=fxx,∂
∂yparenleftbigg∂f
∂yparenrightbigg
=∂2f
∂y2=fyy,
∂
∂xparenleftbigg∂f
∂yparenrightbigg
=∂2f
∂x∂y=fxy,∂
∂yparenleftbigg∂f
∂xparenrightbigg
=∂2f
∂y∂x=fyx.
155
PARTIAL DIFFERENTIATION
Only three of the second derivatives are independent since the relation
∂2f
∂x∂y=∂2f
∂y∂x,
is always obeyed, provided that the second partial derivatives are continuous
at the point in question. This relation often proves useful as a labour-savingdevice when evaluating second partial derivatives. It can also be shown that fora function of nvariables, f(x
1,x2,...,x n), under the same conditions,
∂2f
∂xi∂xj=∂2f
∂xj∂xi.IFind the first and second partial derivatives of the function
f(x, y)=2 x3y2+y3.
The first partial derivatives are
∂f
∂x=6x2y2,∂f
∂y=4x3y+3y2,
and the second partial derivatives are
∂2f
∂x2=1 2xy2,∂2f
∂y2=4x3+6y,∂2f
∂x∂y=1 2x2y,∂2f
∂y∂x=1 2x2y,
the last two being equal, as expected.
J
5.2 The total differential and total derivative
Having defined the (first) partial derivatives of a function f(x, y), which give the
rate of change of falong the positive x-a n d y-axes, we consider next the rate of
change of f(x, y) in an arbitrary direction. Suppose that we make simultaneous
small changes ∆ xinxand ∆ yinyand that, as a result, fchanges to f+∆f.
Then we must have
∆f=f(x+∆x, y+∆y)−f(x, y)
=f(x+∆x, y+∆y)−f(x, y+∆y)+f(x, y+∆y)−f(x, y)
=bracketleftbiggf(x+∆x, y+∆y)−f(x, y+∆y)
∆xbracketrightbigg
∆x+bracketleftbiggf(x, y+∆y)−f(x, y)
∆ybracketrightbigg
∆y.
(5.3)
In the last line we note that the quantities in brackets are very similar to those
involved in the definitions of partial derivatives (5.1), (5.2). For them to be strictlyequal to the partial derivatives, ∆ xand ∆ ywould need to be infinitesimally small.
But even for finite (but not too large) ∆ xand ∆ ythe approximate formula
∆f≈∂f(x, y)
∂x∆x+∂f(x, y)
∂y∆y, (5.4)
156
5.2 THE TOTAL DIFFERENTIAL AND TOTAL DERIVATIVE
can be obtained. It will be noticed that the first bracket in (5.3) actually approxi-
mates to ∂f(x, y+∆y)/∂xbut that this has been replaced by ∂f(x, y)/∂xin (5.4).
This approximation clearly has the same degree of validity as that which replacesthe bracket by the partial derivative.
How valid an approximation (5.4) is to (5.3) depends not only on how small
∆xand ∆ yare but also on the magnitudes of higher partial derivatives; this is
discussed further in section 5.7 in the context of Taylor series for functions of
more than one variable. Nevertheless, letting the small changes ∆ xand ∆ yin
(5.4) become infinitesimal, we can define the total differential dfof the function
f(x, y), without any approximation, as
df=∂f
∂xdx+∂f
∂ydy. (5.5)
Equation (5.5) can be extended to the case of a function of nvariables,
f(x1,x2,...,x n);
df=∂f
∂x1dx1+∂f
∂x2dx2+···+∂f
∂xndxn. (5.6)IFind the total differential of the function f(x, y)=yexp(x+y).
Evaluating the first partial derivatives, we find
∂f
∂x=yexp(x+y),∂f
∂y=e x p ( x+y)+yexp(x+y).
Applying (5.5), we then find that the total differential is given by
df=[yexp(x+y)]dx+[ ( 1+ y)e x p( x+y)]dy.
J
In some situations, despite the fact that several variables xi,i=1,2,...,n,
appear to be involved, effectively only one of them is. This occurs if there aresubsidiary relationships constraining all the x
ito have values dependent on the
value of one of them, say x1. These relationships may be represented by equations
that are typically of the form
xi=xi(x1),i =2,3,...,n . (5.7)
In principle fcan then be expressed as a function of x1alone by substituting
from (5.7) for x2,x3,...,x n, and then the total derivative (or simply the derivative)
offwith respect to x1is obtained by ordinary differentiation.
Alternatively, (5.6) can be used to give
df
dx1=∂f
∂x1+parenleftbigg∂f
∂x2parenrightbiggdx2
dx1+···+parenleftbigg∂f
∂xnparenrightbiggdxn
dx1. (5.8)
It should be noted that the LHS of this equation is the total derivative df/dx 1,
whilst the partial derivative ∂f/∂x 1forms only a part of the RHS. In evaluating
157
PARTIAL DIFFERENTIATION
this partial derivative account must be taken only of explicit appearances of x1in
the function f,a n dnoallowance must be made for the knowledge that changing
x1necessarily changes x2,x3,...,x n. The contribution from these latter changes is
precisely that of the remaining terms on the RHS of (5.8). Naturally, what has
been shown using x1in the above argument applies equally well to any other of
thexi, with the appropriate consequent changes.IFind the total derivative of f(x, y)=x2+3xywith respect to x, given that y=s i n−1x.
We can see immediately that
∂f
∂x=2x+3y,∂f
∂y=3x,dy
dx=1
(1−x2)1/2
and so, using (5.8) with x1=xandx2=y,
df
dx=2x+3y+3x1
(1−x2)1/2
=2x+3s i n−1x+3x
(1−x2)1/2.
Obviously the same expression would have resulted if we had substituted for yfrom the
start, but the above method often produces results with reduced calculation, particularly
in more complicated examples.
J
5.3 Exact and inexact differentials
In the last section we discussed how to find the total differential of a function, i.e.
its infinitesimal change in an arbitrary direction, in terms of its gradients ∂f/∂x
and∂f/∂y in the x-a n d y- directions (see (5.5)). Sometimes, however, we wish
to reverse the process and find the function fthat differentiates to give a known
differential. Usually, finding such functions relies on inspection and experience.
As an example, it is easy to see that the function whose differential is df=
xd y+yd xis simply f(x, y)=xy+c,w h e r e cis a constant. Differentials such as
this, which integrate directly, are called exact differentials , whereas those that do
not are inexact differentials . For example, xd y+3yd xis not the straightforward
differential of any function (see below). Inexact differentials can be made exact,
however, by multiplying through by a suitable function called an integratingfactor. This is discussed further in subsection 14.2.3.IShow that the differential xd y+3yd xis inexact.
On the one hand, if we integrate with respect to xwe conclude that f(x, y)=3 xy+g(y),
where g(y) is any function of y. On the other hand, if we integrate with respect to ywe
conclude that f(x, y)=xy+h(x)w h e r e h(x) is any function of x. These conclusions are
inconsistent for any and every choice of g(y)a n d h(x), and therefore the differential is
inexact.
J
158
5.3 EXACT AND INEXACT DIFFERENTIALS
It is naturally of interest to investigate which properties of a differential make
it exact. Consider the general differential containing two variables,
df=A(x, y)dx+B(x, y)dy.
We see that
∂f
∂x=A(x, y),∂f
∂y=B(x, y)
and, using the property fxy=fyx, we therefore require
∂A
∂y=∂B
∂x. (5.9)
This is in fact both a necessary and a sufficient condition for the differential to
be exact.IUsing (5.9) show that xd y+3yd xis inexact.
In the above notation, A(x, y)=3 yandB(x, y)=xand so
∂A
∂y=3,∂B
∂x=1.
As these are not equal it follows that the differential is inexact.
J
Determining whether a differential containing many variable x1,x2,...,x nis
exact is a simple extension of the above. A differential containing many variablescan be written in general as
df= nsummationdisplay
i=1gi(x1,x2,...,x n)dxi
and will be exact if
∂gi
∂xj=∂gj
∂xifor all pairs i, j. (5.10)
There will be1
2n(n−1) such relationships to be satisfied.IShow that
(y+z)dx+xd y+xd z
is an exact differential.
In this case, g1(x, y, z)=y+z,g2(x, y, z)=x,g3(x, y, z)=xand hence ∂g1/∂y=1=
∂g2/∂x,∂g3/∂x=1= ∂g1/∂z,∂g2/∂z=0= ∂g3/∂y; therefore, from (5.10), the differential
is exact. As mentioned above, it is sometimes possible to show that a differential is exactsimply by finding by inspection the function from which it originates. In this example, itcan be seen easily that f(x, y, z)=x(y+z)+c.J
159
PARTIAL DIFFERENTIATION
5.4 Useful theorems of partial differentiation
So far our discussion has centred on a function f(x, y) dependent on two variables,
xandy. Equally, however, we could have expressed xas a function of fandy,
oryas a function of fandx. To emphasise the point that all the variables are
of equal standing, we now replace fbyz. This does not imply that x,yandz
are coordinate positions (though they might be). Since xis a function of yandz,
it follows that
dx=parenleftbigg∂x
∂yparenrightbigg
zdy+parenleftbigg∂x
∂zparenrightbigg
ydz (5.11)
and similarly, since y=y(x, z),
dy=parenleftbigg∂y
∂xparenrightbigg
zdx+parenleftbigg∂y
∂zparenrightbigg
xdz. (5.12)
We may now substitute (5.12) into (5.11) to obtain
dx=parenleftbigg∂x
∂yparenrightbigg
zparenleftbigg∂y
∂xparenrightbigg
zdx+bracketleftBiggparenleftbigg∂x
∂yparenrightbigg
zparenleftbigg∂y
∂zparenrightbigg
x+parenleftbigg∂x
∂zparenrightbigg
ybracketrightBigg
dz. (5.13)
Now if we hold zconstant, so that dz= 0, we obtain the reciprocity relation
parenleftbigg∂x
∂yparenrightbigg
z=parenleftbigg∂y
∂xparenrightbigg−1
z,
which holds provided both partial derivatives exist and neither is equal to zero.
Note, further, that this relationship only holds when the variable being keptconstant, in this case z, is the same on both sides of the equation.
Alternatively we can put dx= 0 in (5.13). Then the contents of the square
brackets also equal zero, and we obtain the cyclic relation
parenleftbigg∂y
∂zparenrightbigg
xparenleftbigg∂z
∂xparenrightbigg
yparenleftbigg∂x
∂yparenrightbigg
z=−1,
which holds unless any of the derivatives vanish. In deriving this result we have
used the reciprocity relation to replace ( ∂x/∂z )−1
yby (∂z/∂x )y.
5.5 The chain rule
So far we have discussed the differentiation of a function f(x, y) with respect to
its variables xandy.W en o wc o n s i d e rt h ec a s ew h e r e xandyare themselves
functions of another variable, say u. If we wish to find the derivative df/du ,
we could simply substitute in f(x, y) the expressions for x(u)a n d y(u)a n dt h e n
differentiate the resulting function of u. Such substitution will quickly give the
desired answer in simple cases, but in more complicated examples it is easier tomake use of the total differentials described in the previous section.
160
5.6 CHANGE OF VARIABLES
From equation (5.5) the total differential of f(x, y)i sg i v e nb y
df=∂f
∂xdx+∂f
∂ydy,
but we now note that by using the formal device of dividing through by duthis
immediately implies
df
du=∂f
∂xdx
du+∂f
∂ydy
du, (5.14)
which is called the chain rule for partial differentiation. This expression provides
a direct method for calculating the total derivative of fwith respect to uand is
particularly useful when an equation is expressed in a parametric form.IGiven that x(u)=1+ auandy(u)=bu3, find the rate of change of f(x, y)=xe−ywith
respect to u.
As discussed above, this problem could be addressed by substituting for xandyto obtain
fas a function only of uand then differentiating with respect to u. However, using (5.14)
directly we obtain
df
du=(e−y)a+(−xe−y)3bu2,
which on substituting for xandygives
df
du=e−bu3(a−3bu2−3bau3).
J
Equation (5.14) is an example of the chain rule for a function of two variables
each of which depends on a single variable. The chain rule may be extended to
functions of many variables, each of which is itself a function of a variable u,i . e .
f(x1,x2,x3,...,x n), with xi=xi(u). In this case the chain rule gives
df
du=nsummationdisplay
i=1∂f
∂xidxi
du=∂f
∂x1dx1
du+∂f
∂x2dx2
du+···+∂f
∂xndxn
du. (5.15)
5.6 Change of variables
It is sometimes necessary or desirable to make a change of variables during the
course of an analysis, and consequently to have to change an equation expressedin one set of variables into an equation using another set. The same situation arisesif a function fdepends on one set of variables x
i,s ot h a t f=f(x1,x2,...,x n),
but the xiare given in terms of a further set of variables ujby the equations
xi=xi(u1,u2,...,u m). (5.16)
Thexion the right of this equation is a function (of the uj) whilst the xion the
left is the value of that function. For each different value of i,t h e xion the right
161
PARTIAL DIFFERENTIATION
xy
ρ
φ
Figure 5.1 The relationship between Cartesian and plane cylindrical polar
coordinates.
will be a different function of the uj. In this case the chain rule (5.15) becomes
∂f
∂uj=nsummationdisplay
i=1∂f
∂xi∂xi
∂uj,j =1,2,...,m , (5.17)
and is said to express a change of variables . In general the number of variables
in each set need not be equal, i.e. mneed not equal n, but if both the xiand the
uiare sets of independent variables then m=n.IPlane polar coordinates, ρandφ, and Cartesian coordinates, xandy, are related by the
expressions
x=ρcosφ, y =ρsinφ,
as can be seen from figure 5.1. An arbitrary function f(x, y)can be re-expressed as a
function g(ρ, φ). Transform the expression
∂2f
∂x2+∂2f
∂y2
into one in ρandφ.
We first note that ρ2=x2+y2,φ=t a n−1(y/x). We can now write down the four partial
derivatives
∂ρ
∂x=x
(x2+y2)1/2=c o s φ,∂φ
∂x=−(y/x2)
1+(y/x)2=−sinφ
ρ,
∂ρ
∂y=y
(x2+y2)1/2=s i n φ,∂φ
∂y=1/x
1+(y/x)2=cosφ
ρ.
Thus, from (5.17), we may write
∂
∂x=c o s φ∂
∂ρ−sinφ
ρ∂
∂φ,∂
∂y=s i n φ∂
∂ρ+cosφ
ρ∂
∂φ.
162
5.7 TAYLOR’S THEOREM FOR MANY-VARIABLE FUNCTIONS
Now it is only a matter of writing
∂2f
∂x2=∂
∂x
/∂f
∂x
/
=∂
∂x
/∂
∂x
/
f
=
/
cosφ∂
∂ρ−sinφ
ρ∂
∂φ
//
cosφ∂
∂ρ−sinφ
ρ∂
∂φ
/
g
=
/
cosφ∂
∂ρ−sinφ
ρ∂
∂φ
//
cosφ∂g
∂ρ−sinφ
ρ∂g
∂φ
/
=c o s2φ∂2g
∂ρ2+2c os φsinφ
ρ2∂g
∂φ−2c os φsinφ
ρ∂2g
∂φ∂ρ
+sin2φ
ρ∂g
∂ρ+sin2φ
ρ2∂2g
∂φ2
and a similar expression for ∂2f/∂y2,
∂2f
∂y2=
/
sinφ∂
∂ρ+cosφ
ρ∂
∂φ
//
sinφ∂
∂ρ+cosφ
ρ∂
∂φ
/
g
=s i n2φ∂2g
∂ρ2−2cos φsinφ
ρ2∂g
∂φ+2cos φsinφ
ρ∂2g
∂φ∂ρ
+cos2φ
ρ∂g
∂ρ+cos2φ
ρ2∂2g
∂φ2.
When these two expressions are added together the change of variables is complete and
we obtain
∂2f
∂x2+∂2f
∂y2=∂2g
∂ρ2+1
ρ∂g
∂ρ+1
ρ2∂2g
∂φ2.
J
5.7 Taylor’s theorem for many-variable functions
We have already introduced Taylor’s theorem for a function f(x)o fo n ev a r i a b l e ,
in section 4.6. In an analogous way, the Taylor expansion of a function f(x, y)o f
two variables is given by
f(x, y)=f(x0,y0)+∂f
∂x∆x+∂f
∂y∆y
+1
2!bracketleftbigg∂2f
∂x2(∆x)2+2∂2f
∂x∂y∆x∆y+∂2f
∂y2(∆y)2bracketrightbigg
+···,
(5.18)
where ∆ x=x−x0and ∆ y=y−y0, and all the derivatives are to be evaluated
at (x0,y0).
163
PARTIAL DIFFERENTIATIONIFind the Taylor expansion, up to quadratic terms in x−2andy−3,o ff(x, y)=yexpxy
about the point x=2,y=3.
We first evaluate the required partial derivatives of the function, i.e.
∂f
∂x=y2expxy,∂f
∂y=e x p xy+xyexpxy,
∂2f
∂x2=y3expxy,∂2f
∂y2=2xexpxy+x2yexpxy,
∂2f
∂x∂y=2yexpxy+xy2expxy.
Using (5.18), the Taylor expansion of a two-variable function, we find
f(x, y)≈e6
n
3+9 ( x−2) + 7( y−3)
+(2!)−1
/
27(x−2)2+ 48( x−2)(y−3) + 16( y−3)2
/
o
.
J
It will be noticed that the terms in (5.18) containing first derivatives can be
written as
∂f
∂x∆x+∂f
∂y∆y=parenleftbigg
∆x∂
∂x+∆y∂
∂yparenrightbigg
f(x, y),
where both sides of this relation should be evaluated at the point ( x0,y0). Similarly
the terms in (5.18) containing second derivatives can be written as
1
2!bracketleftbigg∂2f
∂x2(∆x)2+2∂2f
∂x∂y∆x∆y+∂2f
∂y2(∆y)2bracketrightbigg
=1
2!parenleftbigg
∆x∂
∂x+∆y∂
∂yparenrightbigg2
f(x, y),
(5.19)
where it is understood that the partial derivatives resulting from squaring the
expression in parentheses act only on f(x, y) and its derivatives, and not on ∆ x
or ∆y; again both sides of (5.19) should be evaluated at ( x0,y0). It can be shown
that the higher-order terms of the Taylor expansion of f(x, y) can be written in
an analogous way, and that we may write the full Taylor series as
f(x, y)=∞summationdisplay
n=01
n!bracketleftbiggparenleftbigg
∆x∂
∂x+∆y∂
∂yparenrightbiggn
f(x, y)bracketrightbigg
x0,y0
where, as indicated, all the terms on the RHS are to be evaluated at ( x0,y0).
The most general form of Taylor’s theorem, for a function f(x1,x2,...,x n)o fn
variables, is a simple extension of the above. Although it is not necessary to do
so, we may think of the xias coordinates in n-dimensional space and write the
function as f(x), where xis a vector from the origin to ( x1,x2,...,x n). Taylor’s
164
5.8 STATIONARY VALUES OF MANY-VARIABLE FUNCTIONS
theorem then becomes
f(x)=f(x0)+summationdisplay
i∂f
∂xi∆xi+1
2!summationdisplay
isummationdisplay
j∂2f
∂xi∂xj∆xi∆xj+···,
(5.20)
where ∆ xi=xi−xi0and the partial derivatives are evaluated at ( x10,x20,...,x n0).
For completeness, we note that in this case the full Taylor series can be writtenin the form
f(x)=∞summationdisplay
n=01
n!bracketleftbig
(∆x·∇)nf(x)bracketrightbig
x=x0,
where∇is the vector differential operator del, to be discussed in chapter 10.
5.8 Stationary values of many-variable functions
The idea of the stationary points of a function of just one variable has already
been discussed in subsection 2.1.8. We recall that the function f(x) has a stationary
point at x=x0if its gradient df/dx is zero at that point. A function may have
any number of stationary points, and their nature, i.e. whether they are maxima,minima or stationary points of inflection, is determined by the value of the secondderivative at the point. A stationary point is
(i) a minimum if d
2f/dx2>0;
(ii) a maximum if d2f/dx2<0;
(iii) a stationary point of inflection if d2f/dx2= 0 and changes sign through
the point.
We now consider the stationary points of functions of more than one variable;
we will see that partial differential analysis is ideally suited to the determination
of the position and nature of such points. It is helpful to consider first the case
of a function of just two variables but, even in this case, the general situationis more complex than that for a function of one variable, as can be seen fromfigure 5.2.
This figure shows part of a three-dimensional model of a function f(x, y). At
positions PandBthere are a peak and a bowl respectively or, more mathemati-
cally, a local maximum and a local minimum. At position Sthe gradient in any
direction is zero but the situation is complicated, since a section parallel to theplane x= 0 would show a maximum, but one parallel to the plane y=0w o u l d
show a minimum. A point such as Sis known as a saddle point . The orientation
of the ‘saddle’ in the xy-plane is irrelevant; it is as shown in the figure solely for
ease of discussion. For any saddle point the function increases in some directionsaway from the point but decreases in other directions.
165
PARTIAL DIFFERENTIATION
PS
y
xB
Figure 5.2 Stationary points of a function of two variables. A minimum
occurs at B, a maximum at Pand a saddle point at S.
For functions of two variables, such as the one shown, it should be clear that a
necessary condition for a stationary point (maximum, minimum or saddle point)
to occur is that
∂f
∂x=0 a n d∂f
∂y=0. (5.21)
The vanishing of the partial derivatives in directions parallel to the axes is enough
to ensure that the partial derivative in any arbitrary direction is also zero. The
latter can be considered as the superposition of two contributions, one alongeach axis; since both contributions are zero, so is the partial derivative in thearbitrary direction. This may be made more precise by considering the totaldifferential
df=∂f
∂xdx+∂f
∂ydy.
Using (5.21) we see that although the infinitesimal changes dxanddycan be
chosen independently the change in the value of the infinitesimal function dfis
always zero at a stationary point.
We now turn our attention to determining the nature of a stationary point of
a function of two variables, i.e. whether it is a maximum, a minimum or a saddlepoint. By analogy with the one-variable case we see that ∂
2f/∂x2and∂2f/∂y2
must both be positive for a minimum and both be negative for a maximum.
However these are not sufficient conditions since they could also be obeyed at
complicated saddle points. What is important for a minimum (or maximum) is
that the second partial derivative must be positive (or negative) in alldirections,
not just the x-a n d y- directions.
166
5.8 STATIONARY VALUES OF MANY-VARIABLE FUNCTIONS
To establish just what constitutes sufficient conditions we first note that, since
fis a function of two variables and ∂f/∂x =∂f/∂y = 0, a Taylor expansion of
the type (5.18) about the stationary point yields
f(x, y)−f(x0,y0)≈1
2!bracketleftbig
(∆x)2fxx+2 ∆ x∆yfxy+( ∆y)2fyybracketrightbig
,
where ∆ x=x−x0and ∆ y=y−y0and where the partial derivatives have been
written in more compact notation. Rearranging the contents of the bracket as
the weighted sum of two squares, we find
f(x, y)−f(x0,y0)≈1
2bracketleftBigg
fxxparenleftbigg
∆x+fxy∆y
fxxparenrightbigg2
+( ∆y)2parenleftBigg
fyy−f2
xy
fxxparenrightBiggbracketrightBigg
.
(5.22)
For a minimum, we require (5.22) to be positive for all ∆ xand ∆ y, and hence
fxx>0a n d fyy−(f2
xy/fxx)>0. Given the first constraint, the second can be
written fxxfyy>f2
xy. Similarly for a maximum we require (5.22) to be negative,
and hence fxx<0a n d fxxfyy>f2
xy. For minima and maxima, symmetry requires
thatfyyobeys the same criteria as fxx. When (5.22) is negative (or zero) for some
values of ∆ xand ∆ ybut positive (or zero) for others, we have a saddle point. In
this case fxxfyy<f2
xy. In summary, all stationary points have fx=fy=0a n d
they may be classified further as
(i) minima if both fxxandfyyare positive andf2
xy<f xxfyy,
(ii) maxima if both fxxandfyyare negative andf2
xy<f xxfyy,
(iii) saddle points if fxxandfyyhave opposite signs orf2
xy>f xxfyy.
Note, however, that if f2
xy=fxxfyythen f(x, y)−f(x0,y0) can be written in one
of the four forms
±1
2parenleftBig
∆x|fxx|1/2±∆y|fyy|1/2parenrightBig2
.
For some choice of the ratio ∆ y/∆xthis expression has zero value showing
that, for a displacement from the stationary point in this particular direction,f(x
0+∆x, y0+∆y) does not differ from f(x0,y0) to second order in ∆ xand
∆y; in such situations further investigation is required. In particular, if fxx,fyy
and fxyare all zero then the Taylor expansion has to be taken to a higher
order. As examples, such extended investigations would show that the function
f(x, y)=x4+y4has a mimimum at the origin but that g(x, y)=x4+y3has a
saddle point there.
167
PARTIAL DIFFERENTIATIONIShow that the function f(x, y)=x3exp(−x2−y2)has a maximum at the point (
p
3/2,0),
a minimum at (−
p
3/2,0)and a stationary point at the origin whose nature cannot be
determined by the above procedures.
Setting the first two partial derivatives to zero to locate the stationary points, we find
∂f
∂x=( 3x2−2x4)e x p(−x2−y2)=0 , (5.23)
∂f
∂y=−2yx3exp(−x2−y2)=0 . (5.24)
For (5.24) to be satisfied we require x=0o r y= 0 and for (5.23) to be satisfied we require
x=0o r x=±
p
3/2. Hence the stationary points are at (0 ,0), (
p
3/2,0) and (−
p
3/2,0).
We now find the second partial derivatives:
fxx=( 4x5−14x3+6x)exp(−x2−y2)
fyy=x3(4y2−2)exp(−x2−y2)
fxy=2x2y(2x2−3) exp(−x2−y2).
We then substitute the pairs of values of xandyfor each stationary point and find that
at (0,0)
fxx=0,f yy=0,f xy=0
and at (±
p
3/2,0)
fxx=∓6
p
3/2exp(−3/2),f yy=∓3
p
3/2exp(−3/2),f xy=0.
Hence, applying criteria (i)–(iii) above, we find that (0 ,0) is an undetermined stationary
point, (
p
3/2,0) is a maximum and ( −
p
3/2,0) is a minimum. The function is shown in
figure 5.3.
J
Determining the nature of stationary points for functions of a general number
of variables is considerably more difficult and requires a knowledge of the
eigenvectors and eigenvalues of matrices. Although these are not discussed until
chapter 8, we present the analysis here for completeness. The remainder of thissection can therefore be omitted on a first reading.
For a function of nreal variables, f(x
1,x2,...,x n), we require that, at all
stationary points,
∂f
∂xi=0 f o ra l l xi.
In order to determine the nature of a stationary point, we must expand the
function as a Taylor series about the point. Recalling the Taylor expansion (5.20)for a function of nvariables, we see that
∆f=f(x)−f(x
0)≈1
2summationdisplay
isummationdisplay
j∂2f
∂xi∂xj∆xi∆xj. (5.25)
168
5.8 STATIONARY VALUES OF MANY-VARIABLE FUNCTIONS
minimum
xy
−1 1 2 3−22
−2 −3−0.2
−0.40.20.4
0
00maximum
Figure 5.3 The function f(x, y)=x3exp(−x2−y2).
If we define the matrix Mto have elements given by
Mij=∂2f
∂xi∂xj,
then we can rewrite (5.25) as
∆f=1
2∆xTM∆x, (5.26)
where ∆ xis the column vector with the ∆ xias its components and ∆ xTis its
transpose. Since Mis real and symmetric it has nreal eigenvalues λrand n
orthogonal eigenvectors er, which after suitable normalisation satisfy
Mer=λrer, eT
res=δrs,
where the Kronecker delta , written δrs, equals unity for r=sand equals zero
otherwise. These eigenvectors form a basis set for the n-dimensional space and
we can therefore expand ∆ xin terms of them, obtaining
∆x=summationdisplay
rarer,
169
PARTIAL DIFFERENTIATION
where the arare coefficients dependent upon ∆ x. Substituting this into (5.26), we
find
∆f=1
2∆xTM∆x=1
2summationdisplay
rλra2
r.
Now, for the stationary point to be a minimum, we require ∆ f=1
2summationtext
rλra2
r>0
for all sets of values of the ar, and therefore all the eigenvalues of Mto be
greater than zero. Conversely, for a maximum we require ∆ f=1
2summationtext
rλra2
r<0,
and therefore all the eigenvalues of Mto be less than zero. If the eigenvalues have
mixed signs, then we have a saddle point. Note that the test may fail if some or
all of the eigenvalues are equal to zero and all the non-zero ones have the samesign.IDerive the conditions for maxima, minima and saddle points for a function of two real
variables, using the above analysis.
For a two-variable function the matrix Mis given by
M=
/
fxxfxy
fyxfyy
/
.
Therefore its eigenvalues satisfy the equation////fxx−λf xy
fxy fyy−λ
////=0.
Hence
(fxx−λ)(fyy−λ)−f2
xy=0
⇒ fxxfyy−(fxx+fyy)λ+λ2−f2
xy=0
⇒ 2λ=(fxx+fyy)±
q
(fxx+fyy)2−4(fxxfyy−f2
xy),
which by rearrangement of the terms under the square root gives
2λ=(fxx+fyy)±
q
(fxx−fyy)2+4f2
xy.
Now, that Mis real and symmetric implies that its eigenvalues are real, and so for both
eigenvalues to be positive (corresponding to a minimum), we require fxxandfyypositive
and also
fxx+fyy>
q
(fxx+fyy)2−4(fxxfyy−f2
xy),
⇒ fxxfyy−f2
xy>0.
A similar procedure will find the criteria for maxima and saddle points.
J
5.9 Stationary values under constraints
In the previous section we looked at the problem of finding stationary values of
a function of two or more variables when all the variables may be independently
170
5.9 STATIONARY VALUES UNDER CONSTRAINTS
varied. However, it is often the case in physical problems that not all the vari-
ables used to describe a situation are in fact independent, i.e. some relationshipbetween the variables must be satisfied. For example, if we walk through a hillylandscape and we are constrained to walk along a path, we will never reach
the highest peak on the landscape, unless the path happens to take us to it.
Nevertheless, we can still find the highest point that we have reached during ourjourney.
We first discuss the case of a function of just two variables. Let us consider
finding the maximum value of the differentiable function f(x, y) subject to the
constraint g(x, y)=c,w h e r e cis a constant. In the above analogy, f(x, y)m i g h t
represent the height of the land above sea-level in some hilly region, whilst
g(x, y)=cis the equation of the path along which we walk.
We could, of course, use the constraint g(x, y)=cto substitute for xoryin
f(x, y), thereby obtaining a new function of only one variable whose stationary
points could be found using the methods discussed in subsection 2.1.8. However,
such a procedure can involve a lot of algebra and becomes very tedious for func-
tions of more than two variables. A more direct method for solving such problemsis the method of Lagrange undetermined multipliers , which we now discuss.
To maximise fwe require
df=∂f
∂xdx+∂f
∂ydy=0.
Ifdxanddywere independent, we could conclude fx=0= fy. However, here
they are not independent, but constrained because gis constant:
dg=∂g
∂xdx+∂g
∂ydy=0.
Multiplying dgby an as yet unknown number λand adding it to dfwe obtain
d(f+λg)=parenleftbigg∂f
∂x+λ∂g
∂xparenrightbigg
dx+parenleftbigg∂f
∂y+λ∂g
∂yparenrightbigg
dy=0,
where λis called a Lagrange undetermined multiplier . In this equation dxanddy
are to be independent and arbitrary; we must therefore choose λsuch that
∂f
∂x+λ∂g
∂x=0, (5.27)
∂f
∂y+λ∂g
∂y=0. (5.28)
These equations, together with the constraint g(x, y)=c, are sufficient to find the
three unknowns, i.e. λand the values of xandyat the stationary point.
171
PARTIAL DIFFERENTIATIONIThe temperature of a point (x, y)on a unit circle is given by T(x, y)=1+ xy.F i n dt h e
temperature of the two hottest points on the circle.
We need to maximise T(x, y) subject to the constraint x2+y2= 1. Applying (5.27) and
(5.28), we obtain
y+2λx=0, (5.29)
x+2λy=0. (5.30)
These results, together with the original constraint x2+y2= 1, provide three simultaneous
equations that may be solved for λ,xandy.
From (5.29) and (5.30) we find λ=±1/2, which in turn implies that y=∓x. Remem-
bering that x2+y2= 1, we find that
y=x⇒x=±1√
2,y =±1√
2
y=−x⇒x=∓1√
2,y =±1√
2.
We have not yet determined which of these stationary points are maxima and which are
minima. In this simple case, we need only substitute the four pairs of x-a n d y- values into
T(x, y)=1+ xyto find that the maximum temperature on the unit circle is Tmax=3/2a t
the points y=x=±1/√
2.
J
The method of Lagrange multipliers can be used to find the stationary points of
functions of more than two variables, subject to several constraints, provided thatthe number of constraints is smaller than the number of variables. For example,if we wish to find the stationary points of f(x, y, z) subject to the constraints
g(x, y, z)=c
1andh(x, y, z)=c2,w h e r e c1andc2are constants, then we proceed
as above, obtaining
∂
∂x(f+λg+µh)=∂f
∂x+λ∂g
∂x+µ∂h
∂x=0,
∂
∂y(f+λg+µh)=∂f
∂y+λ∂g
∂y+µ∂h
∂y=0, (5.31)
∂
∂z(f+λg+µh)=∂f
∂z+λ∂g
∂z+µ∂h
∂z=0.
We may now solve these three equations, together with the two constraints, to
giveλ,µ,x,yandz.
172
5.9 STATIONARY VALUES UNDER CONSTRAINTSIFind the stationary points of f(x, y, z)=x3+y3+z3subject to the following constraints:
(i)g(x, y, z)=x2+y2+z2=1;
(ii)g(x, y, z)=x2+y2+z2=1andh(x, y, z)=x+y+z=0.
Case (i). Since there is only one constraint in this case, we need only introduce a single
Lagrange multiplier to obtain
∂
∂x(f+λg)=3 x2+2λx=0,
∂
∂y(f+λg)=3 y2+2λy=0, (5.32)
∂
∂z(f+λg)=3 z2+2λz=0.
These equations are highly symmetrical and clearly have the solution x=y=z=−2λ/3.
Using the constraint x2+y2+z2= 1 we find λ=±√
3/2 and so stationary points occur
at
x=y=z=±1√
3. (5.33)
In solving the three equations (5.32) in this way, however, we have implicitly assumed
thatx,yandzare non-zero. However, it is clear from (5.32) that any of these values can
equal zero, with the exception of the case x=y=z= 0 since this is prohibited by the
constraint x2+y2+z2= 1. We must consider the other cases separately.
Ifx= 0, for example, we require
3y2+2λy=0,
3z2+2λz=0,
y2+z2=1.
Clearly, we require λ/negationslash= 0, otherwise these equations are inconsistent. If neither ynor
zis zero we find y=−2λ/3= za n df r o mt h et h i r de q u a t i o nw er e q u i r e y=z=
±1/√
2. If y= 0, however, then z=±1 and, similarly, if z=0t h e n y=±1. Thus the
stationary points having x=0a r e( 0 ,0,±1), (0 ,±1,0) and (0 ,±1/√
2,±1/√
2). A similar
procedure can be followed for the cases y=0a n d z= 0 respectively and, in addition
to those already obtained, we find the stationary points ( ±1,0,0), (±1/√
2,0,±1/√
2) and
(±1/√
2,±1/√
2,0).
Case (ii). We now have two constraints and must therefore introduce two Lagrange
multipliers to obtain (cf. (5.31))
∂
∂x(f+λg+µh)=3 x2+2λx+µ=0, (5.34)
∂
∂y(f+λg+µh)=3 y2+2λy+µ=0, (5.35)
∂
∂z(f+λg+µh)=3 z2+2λz+µ=0. (5.36)
These equations are again highly symmetr ical and the simplest way to proceed is to
subtract (5.35) from (5.34) to obtain
3(x2−y2)+2λ(x−y)=0
⇒ 3(x+y)(x−y)+2λ(x−y)=0 . (5.37)
This equation is clearly satisfied if x=y; then, from the second constraint, x+y+z=0 ,
173
PARTIAL DIFFERENTIATION
we find z=−2x. Substituting these values into the first constraint, x2+y2+z2=1 ,w e
obtain
x=±1√
6,y =±1√
6,z =∓2√
6. (5.38)
Because of the high degree of symmetry amongst the equations (5.34)–(5.36), we may obtain
by inspection two further relations analogous to (5.37), one containing the variables y,z
and the other the variables x, z. Assuming y=zin the first relation and x=zin the
second, we find the stationary points
x=±1√
6,y =∓2√
6,z =±1√
6(5.39)
and
x=∓2√
6,y =±1√
6,z =±1√
6. (5.40)
We note that in finding the stationary points (5.38)–(5.40) we did not need to evaluate the
Lagrange multipliers λandµexplicitly. This is not always the case, however, and in some
problems it may be simpler to begin by finding the values of these multipliers.
Returning to (5.37) we must now consider the case where x/negationslash=y; then we find
3(x+y)+2λ=0. (5.41)
However, in obtaining the stationary points (5.39), (5.40), we did notassume x=ybut
only required y=zandx=zrespectively. It is clear that x/negationslash=yat these stationary points,
and it can be shown that they do indeed satisfy (5.41). Similarly, several stationary pointsfor which x/negationslash=zory/negationslash=zhave already been found.
Thus we need to consider further only two cases: ( a)x=y=z,a n d( b)x,yandzare
all different. The first is clearly prohibited by the constraint x+y+z= 0. For the second
case, (5.41) must be satisfied, together with the analogous equations containing y,zand
x, zrespectively, i.e.
3(x+y)+2λ=0,
3(y+z)+2λ=0,
3(x+z)+2λ=0.
Adding these three equations together and using the constraint x+y+z= 0 we find λ=0 .
However, for λ= 0 the equations are inconsistent for non-zero x,yandz. Therefore all
the stationary points have already been found and are given by (5.38)–(5.40).J
The method may be extended to functions of any number nof variables
subject to any smaller number mof constraints. This means that effectively there
aren−mindependent variables and, as mentioned above, we could solve by
substitution and then by the methods of the previous section. However, for large
nthis becomes cumbersome and the use of Lagrange undetermined multipliers is
a useful simplification.
174
5.9 STATIONARY VALUES UNDER CONSTRAINTSIA system contains a very large number Nof particles, each of which can be in any of R
energy levels with a corresponding energy Ei,i=1,2,...,R . The number of particles in the
ith level is niand the total energy of the system is a constant, E. Find the distribution of
particles amongst the energy level s that maximises the expression
P=N!
n1!n2!···nR!,
subject to the constraints that both the number of particles and the total energy remain
constant, i.e.
g=N−RX
i=1ni=0 a n d h=E−RX
i=1niEi=0.
The way in which we proceed is as follows. In order to maximise P, we must minimise
its denominator (since the numerator is fixed). Minimising the denominator is the same asminimising the logarithm of the denominator, i.e.
f=l n(n
1!n2!···nR!)=l n(n1!)+l n(n2!)+···+l n(nR!).
Using Stirling’s approximation, ln (n!)≈nlnn−n, we find that
f=n1lnn1+n2lnn2+···+nRlnnR−(n1+n2+···+nR)
=
/ RX
i=1nilnni
/!
−N.
It has been assumed here that, for the desired distribution, all the niare large. Thus, we
now have a function fsubject to two constraints, g=0a n d h= 0, and we can apply the
Lagrange method, obtaining (cf. (5.31))
∂f
∂n1+λ∂g
∂n1+µ∂h
∂n1=0,
∂f
∂n2+λ∂g
∂n2+µ∂h
∂n2=0,
...
∂f
∂nR+λ∂g
∂nR+µ∂h
∂nR=0.
Since all these equations are alike, we consider the general case
∂f
∂nk+λ∂g
∂nk+µ∂h
∂nk=0,
fork=1,2,...,R . Substituting the functions f,gandhinto this relation we find
nk
nk+l nnk+λ(−1) +µ(−Ek)=0 ,
which can be rearranged to give
lnnk=µEk+λ−1,
and hence
nk=CexpµEk.
175
PARTIAL DIFFERENTIATION
We now have the general form for the distribution of particles amongst energy levels, but
in order to determine the two constants µ,Cwe recall that
RX
k=1CexpµEk=N
and
RX
k=1CE kexpµEk=E.
This is known as the Boltzmann distribution and is a well-known result from statistical
mechanics.
J
5.10 Envelopes
As noted at the start of this chapter, many of the functions with which the
physicists, chemists and engineers have to deal contain, in addition to constantsand one or more variables, quantities that are normally considered as parametersof the system under study. Such parameters may, for example, represent thecapacitance of a capacitor, the length of a rod, or the mass of a particle –
quantities that are normally taken as fixed for any particular physical set-up.
The corresponding variables may well be time, currents, charges, positions andvelocities. However, the parameters could be varied and in this section we study
the effects of doing so; in particular we study how the form of dependence ofone variable on another, typically y=y(x), is affected when the value of a
parameter is changed in a smooth and continuous way. In effect, we are makingthe parameter into an additional variable.
As a particular parameter, which we denote by α, is varied over its permitted
range, the shape of the plot of yagainst xwill change, usually, but not always,
in a smooth and continuous way. For example, if the muzzle speed vof a shell
fired from a gun is increased through a range of values then its height–distancetrajectories will be a series of curves with a common starting point that areessentially just magnified copies of the original; furthermore the curves do notcross each other. However, if the muzzle speed is kept constant but θ, the angle
of elevation of the gun, is increased through a series of values, the corresponding
trajectories do not vary in a monotonic way. When θhas been increased beyond
45
◦the resulting trajectory does cross some of the trajectories corresponding
toθ<45◦. The trajectories all lie within a curve that touches each individual
trajectory at one point. Such a curve is called the envelope to the set of trajectory
solutions; it is to the study of such envelopes that this section is devoted.
For our general discussion of envelopes we will consider an equation of the
form f=f(x, y, α ) = 0. A function of three Cartesian variables, f=f(x, y, α ),
is defined at all points in xyα-space, whereas f=f(x, y, α )=0i sa surface in
this space. A plane of constant α, which is parallel to the xy-plane, cuts such
176
5.10 ENVELOPES
PP1
f(x, y, α 1)=0 f(x, y, α 1+h)=0
Figure 5.4 Two neighbouring curves in the xy-plane of the family f(x, y, α)=
0 intersecting at P.F o rfi x e d α1, the point P1is the limiting position of Pas
h→0. As α1is varied, P1delineates the envelope of the family (broken line).
a surface in a curve. Thus different values of the parameter αcorrespond to
different curves, which can be plotted in the xy-plane. We now investigate how
theenvelope equation for such a family of curves is obtained.
5.10.1 Envelope equations
Suppose f(x, y, α 1)=0a n d f(x, y, α 1+h) = 0 are two neighbouring curves of a
family for which the parameter αdiffers by a small amount h. Let them intersect
at the point Pwith coordinates x, y, as shown in figure 5.4. Then the envelope,
indicated by the broken line in the figure, touches f(x, y, α 1) = 0 at the point P1,
which is defined as the limiting position of Pwhen α1is fixed but h→0. The
full envelope is the curve traced out by P1asα1changes to generate successive
members of the family of curves. Of course, for any finite h,f(x, y, α 1+h)=0i s
one of these curves and the envelope touches it at the point P2.
We are now going to apply Rolle’s theorem, see subsection 2.1.10, with the
parameter αas the independent variable and xandyfixed as constants. In this
context, the two curves in figure 5.4 can be thought of as the projections onto thexy-plane of the planar curves in which the surface f=f(x, y, α ) = 0 meets the
planes α=α
1andα=α1+h.
Along the normal to the page that passes through P,a sαchanges from α1
toα1+hthe value of f=f(x, y, α ) will depart from zero, because the normal
meets the surface f=f(x, y, α )=0o n l ya t α=α1and at α=α1+h. However,
at these end points the values of f=f(x, y, α ) will both be zero, and therefore
equal. This allows us to apply Rolle’s theorem and so to conclude that for someθin the range 0 ≤θ≤1 the partial derivative ∂f(x, y, α
1+θh)/∂αis zero. When
177
PARTIAL DIFFERENTIATION
his made arbitrarily small, so that P→P1, the three defining equations reduce
to two and define the envelope point P1:
f(x, y, α 1)=0 a n d∂f(x, y, α 1)
∂α=0, (5.42)
In (5.42) both the function and the gradient are evaluated at α=α1. The equation
of the envelope g(x, y) = 0 is found by eliminating α1between the two equations.
As a simple example we will now solve the problem which when posed mathe-
matically reads ‘calculate the envelope appropriate to the family of straight lines
in the xy-plane whose points of intersection with the coordinate axes are a fixed
distance apart’. In more ordinary language, the problem is about a ladder leaningagainst a wall.IA ladder of length Lcan be stood on level ground and leant at any angle against a vertical
wall. Find the equation of the curve bounding the vertical area that can be accessed fromthe ladder.
We take the ground and the wall as the x-a n d y-axes respectively. If the foot of the ladder
isafrom the foot of the wall and the top is babove the ground, the straight-line equation
of the ladder is
x
a+y
b=1,
where aandbare connected by a2+b2=L2. Expressed in standard form with only one
independent parameter, a, the equation becomes
f(x, y, a)=x
a+y
(L2−a2)1/2−1=0 . (5.43)
Now, differentiating (5.43) with respect to aand setting the derivative ∂f/∂a equal to
zero gives
−x
a2+ay
(L2−a2)3/2=0 ;
from which it follows that
a=Lx1/3
(x2/3+y2/3)1/2and ( L2−a2)1/2=Ly1/3
(x2/3+y2/3)1/2.
Eliminating aby substituting these values into (5.43) gives, for the equation of the
envelope of all possible positions on the ladder,
x2/3+y2/3=L2/3.
This is the equation of an astroid (mentioned in exercise 2.19), and, together with the wall
and the ground, marks the boundary of the vertical area that can be accessed by (theshoes of) a person standing on the ladder.J
Other examples, drawn from both geometry and and the physical sciences, are
considered in the exercises at the end of this chapter. The shell trajectory problem
discussed earlier in this section is solved there, but in the guise of a questionabout the water bell of an ornamental fountain.
178
5.11 THERMODYNAMIC RELATIONS
5.11 Thermodynamic relations
Thermodynamic relations provide a useful set of physical examples of partial
differentiation. The relations we will derive are called Maxwell’s thermodynamic
relations . They express relationships between four thermodynamic quantities de-
scribing a unit mass of a substance. The quantities are the pressure P, the volume
V, the thermodynamic temperature Tand the entropy Sof the substance. These
four quantities are not independent; any two of them can be varied indepen-dently, but the other two are then determined. The first law of thermodynamicsmay be expressed as
dU=Td S−Pd V , (5.44)
where Uis the internal energy of the substance. Essentially this is a conservation
of energy equation, but we shall concern ourselves, not with the physics, but ratherwith the use of partial differentials to relate the four basic quantities discussedabove. The method involves writing a total differential, dUsay, in terms of the
differentials of two variables, say XandY, thus
dU=parenleftbigg∂U
∂Xparenrightbigg
YdX+parenleftbigg∂U
∂Yparenrightbigg
XdY , (5.45)
and then using the relationship
∂2U
∂X∂Y=∂2U
∂Y ∂X
to obtain the required Maxwell relation. The variables XandYare to be chosen
from P,V,TandS.IShow that (∂T/∂V )S=−(∂P/∂S )V.
Here the two variables that have to be held constant, in turn, happen to be those whose
differentials appear on the RHS of (5.44). And so, taking XasSandYasVin (5.45), we
have
Td S−Pd V=dU=
/∂U
∂S
/
VdS+
/∂U
∂V
/
SdV,
and find directly that/∂U
∂S
/
V=T and
/∂U
∂V
/
S=−P.
Differentiating the first expression with respect to Vand the second with respect to S,a n d
using
∂2U
∂V∂S=∂2U
∂S∂V,
we find the Maxwell relation/∂T
∂V
/
S=−
/∂P
∂S
/
V.
J
179
PARTIAL DIFFERENTIATIONIShow that (∂S/∂V )T=(∂P/∂T )V.
Applying (5.45) to dS, with independent variables VandT, we find
dU=Td S−Pd V=T
//∂S
∂V
/
TdV+
/∂S
∂T
/
VdT
/
−PdV.
Similarly applying (5.45) to dU, we find
dU=
/∂U
∂V
/
TdV+
/∂U
∂T
/
VdT.
Thus, equating partial derivatives,/∂U
∂V
/
T=T
/∂S
∂V
/
T−Pand
/∂U
∂T
/
V=T
/∂S
∂T
/
V.
But, since
∂2U
∂T∂V=∂2U
∂V∂T,i.e.∂
∂T
/∂U
∂V
/
T=∂
∂V
/∂U
∂T
/
V,
it follows that/∂S
∂V
/
T+T∂2S
∂T∂V−
/∂P
∂T
/
V=∂
∂V
/
T
/∂S
∂T
/
V
/
T=T∂2S
∂V∂T.
Thus finally we get the Maxwell relation/∂S
∂V
/
T=
/∂P
∂T
/
V.
J
The above derivation is rather cumbersome, however, and a useful trick that
can simplify the working is to define a new function, called a potential .T h e
internal energy Udiscussed above is one example of a potential but three others
are commonly defined and they are described below.IShow that (∂S/∂V )T=(∂P/∂T )Vby considering the potential U−ST.
We first consider the differential d(U−ST). From (5.5), we obtain
d(U−ST)=dU−SdT−TdS=−SdT−PdV
when use is made of (5.44). We rewrite U−STasFfor convenience of notation; Fis
called the Helmholtz potential . Thus
dF=−SdT−PdV,
and it follows that/∂F
∂T
/
V=−S and
/∂F
∂V
/
T=−P.
Using these results together with
∂2F
∂T∂V=∂2F
∂V∂T,
we can see immediately that/∂S
∂V
/
T=
/∂P
∂T
/
V,
which is the same Maxwell relation as before.
J
180
5.12 DIFFERENTIATION OF INTEGRALS
Although the Helmholtz potential has other uses, in this context it has simply
provided a means for a quick derivation of the Maxwell relation. The otherMaxwell relations can be derived similarly by using two other potentials, theenthalpy ,H=U+PV,a n dt h e Gibbs free energy ,G=U+PV−ST(see
exercise 5.25).
5.12 Differentiation of integrals
We conclude this chapter with a discussion of the differentiation of integrals. Let
us consider the indefinite integral (cf. equation (2.30))
F(x, t)=integraldisplay
f(x, t)dt,
from which it follows immediately that
∂F(x, t)
∂t=f(x, t).
Assuming that the second partial derivatives of F(x, t) are continuous, we have
∂2F(x, t)
∂t∂x=∂2F(x, t)
∂x∂t,
a n ds ow ec a nw r i t e
∂
∂tbracketleftbigg∂F(x, t)
∂xbracketrightbigg
=∂
∂xbracketleftbigg∂F(x, t)
∂tbracketrightbigg
=∂f(x, t)
∂x.
Integrating this equation with respect to tthen gives
∂F(x, t)
∂x=integraldisplay∂f(x, t)
∂xdt. (5.46)
Now consider the definite integral
I(x)=integraldisplayt=v
t=uf(x, t)dt
=F(x, v)−F(x, u),
where uandvare constants. Differentiating this integral with respect to x,a n d
using (5.46), we see that
dI(x)
dx=∂F(x, v)
∂x−∂F(x, u)
∂x
=integraldisplayv∂f(x, t)
∂xdt−integraldisplayu∂f(x, t)
∂xdt
=integraldisplayv
u∂f(x, t)
∂xdt.
This is Leibnitz’ rule for differentiating integrals, and basically it states that for
181
PARTIAL DIFFERENTIATION
constant limits of integration the order of integration and differentiation can be
reversed.
In the more general case where the limits of the integral are themselves functions
ofx, it follows immediately that
I(x)=integraldisplayt=v(x)
t=u(x)f(x, t)dt
=F(x, v(x))−F(x, u(x)),
which yields the partial derivatives
∂I
∂v=f(x, v(x)),∂I
∂u=−f(x, u(x)).
Consequently
dI
dx=parenleftbigg∂I
∂vparenrightbiggdv
dx+parenleftbigg∂I
∂uparenrightbiggdu
dx+∂I
∂x
=f(x, v(x))dv
dx−f(x, u(x))du
dx+∂
∂xintegraldisplayv(x)
u(x)f(x, t)dt
=f(x, v(x))dv
dx−f(x, u(x))du
dx+integraldisplayv(x)
u(x)∂f(x, t)
∂xdt, (5.47)
where the partial derivative with respect to xin the last term has been taken
inside the integral sign using (5.46). This procedure is valid because u(x)a n d v(x)
are being held constant in this term.IFind the derivative with respect to xof the integral
I(x)=
Zx2
xsinxt
tdt.
Applying (5.47), we see that
dI
dx=sinx3
x2(2x)−sinx2
x(1) +
Zx2
xtcosxt
tdt
=2si nx3
x−sinx2
x+
/sinxt
x
/x2
x
=3sinx3
x−2sinx2
x
=1
x(3sin x3−2si nx2).
J
5.13 Exercises
5.1 (a) Find all the first partial derivatives of the following functions f(x, y): (i) x2y,
(ii)x2+y2+ 4, (iii) sin( x/y), (iv) tan−1(y/x), (v) r(x, y, z)=(x2+y2+z2)1/2.
182
5.13 EXERCISES
(b) For (i), (ii) and (v), find ∂2f/∂x2,∂2f/∂y2,∂2f/∂x∂y.
(c) For (iv) verify that ∂2f/∂x∂y =∂2f/∂y∂x.
5.2 Determine which of the following are exact differentials:
(a) (3 x+2 )yd x+x(x+1 )dy,
(b)ytanxd x+xtanyd y,
(c)y2(lnx+1 )dx+2xylnxd y,
(d)y2(lnx+1 )dy+2xylnxd x,
(e) [ x/(x2+y2)]dy−[y/(x2+y2)]dx.
5.3 Show that the differential
df=x2dy−(y2+xy)dx
is not exact, but that dg=(xy2)−1dfis exact.
5.4 (a) Show that
df=y(1 +x−x2)dx+x(x+1 )dy
is not an exact differential.
(b) Find the differential equation that a function g(x) must satisfy if dφ=g(x)df
is to be an exact differential. Verify that g(x)=e−xis a solution of this
equation and deduce the form of φ(x, y).
5.5 The equation 3 y=z3+3xzdefines zimplicitly as a function of xandy. Evaluate
all three second partial derivatives of zwith respect to xand/or y.V e r i f yt h a t z
is a solution of
x∂2z
∂y2+∂2z
∂x2=0.
5.6 A possible equation of state for a gas takes the form
pV=RTexp
/
−α
VRT
/
,
in which αandRare constants. Calculate expressions for/∂p
∂V
/
T,
/∂V
∂T
/
p,
/∂T
∂p
/
V,
and show that their product is −1, as stated in section 5.4.
5.7 The function G(t) is defined by
G(t)=F(x, y)=x2+y2+3xy,
where x(t)=at2andy(t)=2 at. Use the chain rule to find the values of ( x, y)a t
which G(t) has stationary values as a function of t. Do any of them correspond
to the stationary points of F(x, y) as a function of xandy?
5.8 In the xy-plane, new coordinates sandtare defined by
s=1
2(x+y),t =1
2(x−y).
Transform the equation
∂2φ
∂x2−∂2φ
∂y2=0
into the new coordinates and deduce that its general solution can be written
φ(x, y)=f(x+y)+g(x−y),
where f(u)a n d g(v) are arbitrary functions of uandvrespectively.
183
PARTIAL DIFFERENTIATION
5.9 The function f(x, y) satisfies the differential equation
y∂f
∂x+x∂f
∂y=0.
By changing to new variables u=x2−y2andv=2xy, show that fis, in fact, a
function of x2−y2only.
5.10 If x=eucosθandy=eusinθ, show that
∂2φ
∂u2+∂2φ
∂θ2=(x2+y2)
/∂2f
∂x2+∂2f
∂y2
/
,
where f(x, y)=φ(u, θ).
5.11 Find and evaluate the maxima, minima and saddle points of the function
f(x, y)=xy(x2+y2−1).
5.12 Show that
f(x, y)=x3−12xy+4 8x+by2,b /negationslash=0,
has two, one, or zero stationary points according to whether |b|is less than, equal
to, or greater than 3.
5.13 Locate the stationary points of the function
f(x, y)=(x2−2y2)exp[−(x2+y2)/a2],
where ais a non-zero constant.
Sketch the function along the x-a n d y- axes and hence identify the nature and
values of the stationary points.
5.14 Find the stationary points of the function
f(x, y)=x3+xy2−12x−y2
and identify their nature.
5.15 Find the stationary values of
f(x, y)=4 x2+4y2+x4−6x2y2+y4
and classify them as maxima, minima or saddle points. Make a rough sketch of
the contours of fin the quarter plane x, y≥0.
5.16 The temperature of a point ( x, y, z) on the unit sphere is given by
T(x, y, z)=1+ xy+yz.
By using the method of Lagrange multipliers find the temperature of the hottest
point on the sphere.
5.17 A rectangular parallelepiped has all eight vertices on the ellipsoid
x2+3y2+3z2=1.
Using the symmetry of the parallelepiped about each of the planes x=0 ,
y=0 , z= 0, write down the surface area of the parallelepiped in terms of
the coordinates of the vertex that lies in the octant x, y, z≥0. Hence find the
maximum value of the surface area of such a parallelepiped.
5.18 Two horizontal corridors, 0 ≤x≤awith y≥0, and 0≤y≤bwith x≥0, meet
at right angles. Find the length Lof the longest ladder (considered as a stick)
that may be carried horizontally around the corner.
5.19 A barn is to be constructed with a uniform cross-sectional area Athroughout
its length. The cross-section is to be a rectangle of wall height h(fixed) and
width w, surmounted by an isosceles triangular roof that makes an angle θwith
184
5.13 EXERCISES
the horizontal. The cost of construction is αper unit height of wall and βper
unit (slope) length of roof. Show that, irrespective of the values of αandβ,t o
minimise costs wshould be chosen to satisfy the equation
w4=1 6A(A−wh),
andθmade such that 2tan2 θ=w/h.
5.20 Show that the envelope of all concentric ellipses that have their axes along the
x-a n d y-coordinate axes and that have the sum of their semi-axes equal to a
constant Lis the same curve (an astroid) as that found in the worked example
in section 5.10.
5.21 Find the area of the region covered by points on the lines
x
a+y
b=1,
where the sum of any line’s intercepts on the coordinate axes is fixed and equal
toc.
5.22 Prove that the envelope of the circles whose diameters are those chords of a
given circle that pass through a fixed point on its circumference, is the cardioid
r=a(1 + cos θ).
Here ais the radius of the given circle and ( r,θ) are the polar coordinates of the
envelope. Take as the system parameter the angle φbetween a chord and the
polar axis from which θis measured.
5.23 A water feature contains a spray head at water level at the centre of a round
basin. The head is in the form of a hemisphere with many evenly distributedsmall holes in it, and through which water spurts out at the same speed v
0in all
directions.
(a) What is the shape of the ‘water bell’ so formed?
(b) What must be the minimum diameter of the bowl if no water is to be lost?
5.24 In order to make a focussing mirror that concentrates parallel axial rays to one
spot (or conversely forms a parallel beam from a point source) a parabolic shapeshould be adopted. If a mirror that is part of a circular cylinder or sphere wereused, the light would be spread out along a curve. This curve is known as acaustic and is the envelope of the rays reflected from the mirror. Denoting by θ
the angle which a typical incident axial ray makes with the normal to the mirrorat the place where it is reflected, the geometry of reflection (the angle of incidenceequals the angle of reflection) is shown in figure 5.5.
Show that a parametric specification of the caustic is
x=Rcosθ/;1
2+s i n2θ
/
,y =Rsin3θ,
where Ris the radius of curvature of the mirror. The curve is, in fact, part of an
epicycloid.
5.25 By considering the differential
dG=d(U+PV−ST),
where Gis the Gibbs free energy, Pthe pressure, Vthe volume, Sthe entropy
andTthe temperature of a system, and given further that
dU=TdS−PdV,
derive a Maxwell relation connecting ( ∂V/∂T )Pand ( ∂S/∂P )T.
185
PARTIAL DIFFERENTIATION
OR
xy
θθ
2θ
Figure 5.5 The reflecting mirror discussed in exercise 5.24.
5.26 Functions P(V,T),U(V,T)a n d S(V,T) are related by
TdS=dU+PdV,
where the symbols have the same meaning as in the previous question. Pis
known from experiment to have the form
P=T4
3+T
V,
in appropriate units. If
U=αVT4+βT,
where α,β, are constants (or at least do not depend on T,V), deduce that α
must have a specific value but βmay have any value. Find the corresponding
form of S.
5.27 As in the previous two exercises on the thermodynamics of a simple gas, the
quantity dS=T−1(dU+PdV) is an exact differential. Use this to prove that/∂U
∂V
/
T=T
/∂P
∂T
/
V−P.
In the van der Waals model of a gas, Pobeys the equation
P=RT
V−b−a
V2,
where R,aandbare constants. Further, in the limit V→∞, the form of U
becomes U=cT ,where cis another constant. Find the complete expression for
U(V,T).
5.28 The entropy S(H,T), the magnetisation M(H,T) and the internal energy U(H,T)
of a magnetic salt placed in a magnetic field of strength Hat temperature Tare
connected by the equation
TdS=dU−HdM.
186
5.13 EXERCISES
By considering d(U−TS−HM), or otherwise, prove that/∂M
∂T
/
H=
/∂S
∂H
/
T.
For a particular salt
M(H,T)=M0[1−exp(−αH/T )].
Show that, at a fixed temperature, if the applied field is increased from zero to
a strength such that the magnetization of the salt is3
4M0then the salt’s entropy
decreases by an amount
M0
4α(3−ln 4).
5.29 Using the results of section 5.12, evaluate the integral
I(y)=
Z∞
0e−xysinx
xdx.
Hence show that
J=
Z∞
0sinx
xdx=π
2.
5.30 The integralZ∞
−∞e−αx2dx
has the value ( π/α)1/2. Use this result to evaluate
J(n)=
Z∞
−∞x2ne−x2dx,
where nis a positive integer. Express your answer in terms of factorials.
5.31 The function f(x) is differentiable and f(0) = 0. A second function g(y) is defined
by
g(y)=
Zy
0f(x)dx√y−x.
Prove that
dg
dy=
Zy
0df
dxdx√y−x.
For the case f(x)=xn, prove that
dng
dyn=2 (n!)√y.
5.32 The functions f(x, t)a n d F(x) are defined by
f(x, t)=e−xt,
F(x)=
Zx
0f(x, t)dt.
Verify by explicit calculation that
dF
dx=f(x, x)+
Zx
0∂f(x, t)
∂xdt.
187
PARTIAL DIFFERENTIATION
5.33 If
I(α)=
Z1
0xα−1
lnxdx, α > −1,
what is the value of I(0)? Show that
d
dαxα=xαlnx,
and deduce that
d
dαI(α)=1
α+1.
Hence prove that I(α)=l n ( 1+ α).
5.34 Find the derivative with respect to xof the integral
I(x)=
Z3x
xexpxt dt.
5.35 The function G(t, ξ) is defined for 0 ≤t≤πby
G(t, ξ)=
/(
−costsinξ forξ≤t,
−sintcosξ forξ>t .
Show that the function x(t) defined by
x(t)=
Zπ
0G(t, ξ)f(ξ)dξ
satisfies the equation
d2x
dt2+x=f(t)
foranyarbitrary (continuous) function f(t). Show further that x(0) =
[dx/dt]x=π= 0, again for any f(t), but that the value of x(π) does depend
upon the form of f(t).
(The function G(t, ξ) is an example of a Green’s function, an important
concept in the solution of differential equations and one studied extensively inlater chapters.)
5.14 Hints and answers
5.1 (a) (i) 2 xy, x2; (ii) 2 x,2y; (iii) y−1cos(x/y),(−x/y2)cos( x/y);
(iv)−y/(x2+y2),x /(x2+y2); (v) x/r,y/r,z/r.
(b) (i) 2 y,0,2x; (ii) 2 ,2,0; (v) ( y2+z2)r−3,(x2+z2)r−3,−xyr−3.
(c) Both second derivatives are equal to ( y2−x2)(x2+y2)−2.
5.2 Only (c) and (e).5.3 2 x/negationslash=−2y−x.Forgboth sides of equation (5.9) equal y
−2.
5.4 (a) 1 + x−x2/negationslash=2x+1 .( b ) g/prime=−g.φ(x, y)=x(x+1 )ye−x+k.
5.5 ∂2z/∂x2=2xz(z2+x)−3,∂2z/∂x∂y =(z2−x)(z2+x)−3,∂2z/∂y2=−2z(z2+x)−3.
5.6 The equation is most easily differentiated in the form ln p+l nV−lnR−lnT=
−α/(VRT).p(α−VRT)/(V2RT);V(α+VRT)/[T(VRT−α)];VRT2/[p(α+
VRT)].
5.7 (0 ,0),(a/4,−a)a n d( 1 6 a,−8a). Only the saddle point at (0 ,0).
5.8 The transformed equation is ∂2ψ/∂t∂s = 0 where ψ(s, t)=φ(x, y).
5.9 The transformed equation is 2( x2+y2)∂f/∂v = 0; hence fdoes not depend on v.
5.10 Write ∂/∂uand∂/∂θin terms of x,y,∂/∂x and∂/∂y using (5.17). The terms
that cancel when ∂2φ/∂u2and∂2φ/∂θ2are added together are ±[x(∂f/∂x )+
y(∂f/∂y )+2xy(∂2f/∂x∂y )].
188
5.14 HINTS AND ANSWERS
5.11 Maxima equal to 1 /8a t±(1/2,−1/2), minima equal to −1/8a t±(1/2,1/2),
saddle points equalling 0 at (0 ,0), (0 ,±1), (±1,0).
5.12 From ∂f/∂y =0,y=6x/b. Substitute this into ∂f/∂x = 0 to obtain a quadratic
equation for x.
5.13 Maxima equal to a2e−1at (±a,0), minima equal to −2a2e−1at (0 ,±a), saddle
point equalling 0 at (0 ,0).
5.14 Maximum equal to 16 at ( −2,0), minimum equal to −16 at (2 ,0), saddle points
equalling−11 at (1 ,±3).
5.15 Minimum at (0 ,0); saddle points at ( ±1,±1).
5.161+√
2√
2at±
/
1
2,√
2
2,1
2
/
.
5.17 Lagrange multiplier method gives z=y=x/2f o rm a x i m a la r e ao f4 .
5.18 Put the ends of the ladder at ( a+ξ,0) and (0 ,b+η)a n dr e q u i r e( a, b)t ob eo n
the ladder. L=(a2/3+b2/3)3/2.
5.19 The cost always includes 2 αhwhich can be ignored in the optimisation. With
Lagrange multiplier λ,s i nθ=λw/(4β)a n d βsecθ−1
2λwtanθ=λh, leading to
the stated results.
5.20 If the semi-axis in the x-direction is a,t h e n x2/y2=a3/(L−a)3for the envelope.
5.21 The envelope of lines x/a+y/(c−a)−1 = 0, as avaries, is√x+√y=√c.A r e a
=c2/6.
5.22 The equation of a typical circle is r=2acosφcos(θ−φ). The envelope condition
gives φ=θ/2.
5.23 (a) Using α=c o t θ,w h e r e θis the initial angle a jet makes with the vertical, the
equation is f(z,ρ,α)=z−ρα+[gρ2(1+α2)/(2v2
0)], and setting ∂f/∂α = 0 gives
α=v2
0/(gρ). The water bell has a parabolic profile z=v2
0/(2g)−gρ2/(2v2
0).
(b) Setting z= 0 gives the minimum diameter as 2 v2
0/g.
5.24 The reflected ray has equation y=t a n 2 θ(x−Rsinθ/sin 2θ). Put this into
the standard form f(x, y, θ) = 0 and eliminate yorxfrom this equation and
∂f/∂θ =0 .
5.25 Show that ( ∂G/∂P )T=Vand ( ∂G/∂T )P=−S. From each result obtain an
expression for ∂2G/∂T∂P and equate these, giving ( ∂V/∂T )P=−(∂S/∂P )T.
5.26 Establish that ( ∂U/∂V )T=T(∂S/∂V )T−Pand that ( ∂U/∂T )V=T(∂S/∂T )V.
Equate expressions for ∂2S/∂T∂V and hence show α=1 .I n t e g r a t e( ∂S/∂V )T
and ( ∂S/∂T )Vto show that S=4T3V/3+l n V+βlnT+c.
5.27 Find expressions for ( ∂S/∂V )Tand ( ∂S/∂T )V, and equate ∂2S/∂V∂T with
∂2S/∂T∂V .U(V,T)=cT−aV−1.
5.28 Show that dF=d(U−TS−HM)=−Sd T−Md H and find two expressions for
∂2F/∂H∂T . Establish that ( ∂S/∂H )T=−M0αHT−2exp(−αH/T ) and integrate
with respect to H.
5.29 dI/dy =−Im[
R∞
0exp(−xy+ix)dx]=−1/(1 +y2). Integrate dI/dy from 0 to∞.
I(∞)=0a n d I(0) = J.
5.30 Differentiate both the integral and its value ntimes with respect to αand then
setα= 1. Note that 1 3 5 ···(2n−1) = (2 n)!/(2nn!).J(n)=( 2 n)!√π/(4nn!).
5.31 Integrate the RHS of the equation by parts before differentiating with respect
toy. Repeated application of the method establishes the result for all orders of
derivative.
5.32 Both sides of the equation equal e−x2+x−2(e−x2−1).
5.33 I(0) = 0; use Leibniz’ rule.
5.34 (6 −x−2)e x p( 3 x2)−(2−x−2)e x p x2.
5.35 Write x(t)=−cost
Rt
0sinξf(ξ)dξ−sint
Rπ
tcosξf(ξ)dξand differentiate each
term as a product to obtain dx/dt.O b t a i n d2x/dt2in a similar way. Note
that integrals that have equal lower and upper limits have value zero. x(π)=Rπ
0sinξf(ξ)dξ.
189
6
Multiple integrals
For functions of several variables, just as we may consider derivatives with respect
to two or more of them, so may the integral of the function with respect to morethan one variable be formed. The formal definitions of such multiple integrals areextensions of that for a single variable, discussed in chapter 2. We first discussdouble and triple integrals and illustrate some of their applications. We thenconsider changing the variables in multiple integrals and discuss some generalproperties of Jacobians.
6.1 Double integrals
For an integral involving two variables – a double integral – we have a function,
f(x, y) say, to be integrated with respect to xandybetween certain limits. These
limits can usually be represented by a closed curve Cbounding a region Rin the
xy-plane. Following the discussion of single integrals given in chapter 2, let us
divide the region RintoNsubregions ∆ R
pof area ∆ Ap,p=1,2,...,N , and let
(xp,yp) be any point in subregion ∆ Rp. Now consider the sum
S=Nsummationdisplay
p=1f(xp,yp)∆Ap,
and let N→∞ as each of the areas ∆ Ap→0. If the sum Stends to a unique
limit, I, then this is called the double integral of f(x, y)over the region Rand is
written
I=integraldisplay
Rf(x, y)dA, (6.1)
where dAstands for the element of area in the xy-plane. By choosing the
subregions to be small rectangles each of area ∆ A=∆x∆y, and letting both ∆ x
190
6.1 DOUBLE INTEGRALS
V
U
C
TSdxdy
RdA=dxdyy
d
c
a b x
Figure 6.1 A simple curve Cin the xy-plane, enclosing a region R.
and ∆ y→0 ,w ec a na l s ow r i t et h ei n t e g r a la s
I=integraldisplayintegraldisplay
Rf(x, y)dx dy, (6.2)
where we have written out the element of area explicitly as the product of the
two coordinate differentials (see figure 6.1).
Some authors use a single integration symbol whatever the dimension of the
integral; others use as many symbols as the dimension. In different circumstancesboth have their advantages. We will adopt the convention used in (6.1) and (6.2),that as many integration symbols will be used as differentials explicitly written.
The form (6.2) gives us a clue as to how we may proceed in the evaluation
of a double integral. Referring to figure 6.1, the limits on the integration may
be written as an equation c(x, y) = 0 giving the boundary curve C. However, an
explicit statement of the limits can be written in two distinct ways.
One way of evaluating the integral is first to sum up the contributions from
the small rectangular elemental areas into horizontal strips of width dy(as shown
in the figure) and then to combine these horizontal strips to cover the region R.
In this case, we write
I=integraldisplay
y=d
y=cbraceleftbiggintegraldisplayx=x2(y)
x=x1(y)f(x, y)dxbracerightbigg
dy, (6.3)
where x=x1(y)a n d x=x2(y) are the equations of the curves TSV andTUV
respectively. This expression indicates that first f(x, y) is to be integrated with
respect to x(treating yas a constant) between the values x=x1(y)a n d x=x2(y)
and then the result, considered as a function of y,i st ob ei n t e g r a t e db e t w e e nt h e
limits y=candy=d. Thus the double integral is evaluated by expressing it in
terms of two single integrals called iterated (orrepeated ) integrals.
An alternative way of evaluating the integral, however, is first to sum up the
191
MULTIPLE INTEGRALS
contributions from the elemental rectangles arranged into vertical strips and then
to combine these vertical strips to cover the region R.W et h e nw r i t e
I=integraldisplayx=b
x=abraceleftbiggintegraldisplayy=y2(x)
y=y1(x)f(x, y)dybracerightbigg
dx, (6.4)
where y=y1(x)a n d y=y2(x) are the equations of the curves STU andSVU
respectively. In going to (6.4) from (6.3), we have essentially interchanged the
order of integration.
In the discussion above we assumed that the curve Cwas such that any
line parallel to either the x-o r y-axis intersected Cat most twice. In general,
provided f(x, y) is continuous everywhere in Rand the boundary curve Chas this
simple shape, the same result is obtained irrespective of the order of integration.In cases where the region Rhas a more complicated shape, it can usually be
subdivided into smaller simpler regions R
1,R2etc. that satisfy this criterion. The
double integral over Ris then merely the sum of the double integrals over the
subregions.IEvaluate the double integral
I=
ZZ
Rx2yd xd y ,
where Ris the triangular area bounded by the lines x=0,y=0andx+y=1. Reverse
the order of integration and demonstrate that the same result is obtained.
The area of integration is shown in figure 6.2. Suppose we choose to carry out theintegration with respect to yfirst. With xfixed, the range of yis 0 to 1−x.W ec a n
therefore write
I=
Zx=1
x=0
/Zy=1−x
y=0x2yd y
/
dx
=
Zx=1
x=0
/x2y2
2
/y=1−x
y=0dx=
Z1
0x2(1−x)2
2dx=1
60.
Alternatively, we may choose to perform the integration with respect to xfirst. With y
fixed, the range of xis 0 to 1−y, so we have
I=
Zy=1
y=0
/Zx=1−y
x=0x2yd x
/
dy
=
Zy=1
y=0
/x3y
3
/x=1−y
x=0dx=
Z1
0(1−y)3y
3dy=1
60.
As expected, we obtain the same result irrespective of the order of integration.
J
We may avoid the use of braces in expressions such as (6.3) and (6.4) by writing
(6.4), for example, as
I=integraldisplayb
adxintegraldisplayy2(x)
y1(x)dy f(x, y),
where it is understood that each integral symbol acts on everything to its right,
192
6.2 TRIPLE INTEGRALS
y
11
dy
00dx xx+y=1
R
Figure 6.2 The triangular region whose sides are the axes x=0 , y=0a n d
the line x+y=1 .
and that the order of integration is from right to left. So, in this example, the
integrand f(x, y) is first to be integrated with respect to ya n dt h e nw i t hr e s p e c t
tox. With the double integral expressed in this way, we will no longer write the
independent variables explicitly in the limits of integration, since the differentialof the variable with respect to which we are integrating is always adjacent to therelevant integral sign.
Using the order of integration in (6.3), we could also write the double integral as
I=integraldisplay
d
cdyintegraldisplayx2(y)
x1(y)dx f(x, y).
Occasionally, however, the interchange of the order of integration in a double
integral is not permissible, as it yields a different result. For example, difficultiesmight arise if the region Rwere unbounded with some of the limits are infi-
nite, though, in many cases involving infinite limits the same result is obtainedwhichever order of integration is used. Difficulties can also occur if the integrandf(x, y) has any discontinuities in the region Ror on its boundary C.
6.2 Triple integrals
The above discussion for double integrals can easily be extended to triple integrals.
Consider the function f(x, y, z) defined in a closed three-dimensional region R.
Proceeding as we did for double integrals, let us divide the region Rinto N
subregions ∆ R
pof volume ∆ Vp,p=1,2,...,N ,a n dl e t( xp,yp,zp) be any point in
the subregion ∆ Rp. Now we form the sum
S=Nsummationdisplay
p=1f(xp,yp,zp)∆Vp,
193
MULTIPLE INTEGRALS
and let N→∞as each of the volumes ∆ Vp→0. If the sum Stends to a unique
limit, I, then this is called the triple integral of f(x, y, z)over the region Rand is
written
I=integraldisplay
Rf(x, y, z)dV, (6.5)
where dVstands for the element of volume. By choosing the subregions to be
small cuboids, each of volume ∆ V=∆x∆y∆z, and proceeding to the limit, we
c a na l s ow r i t et h ei n t e g r a la s
I=integraldisplayintegraldisplayintegraldisplay
Rf(x, y, z)dx dy dz, (6.6)
where we have written out the element of volume explicitly as the product of the
three coordinate differentials. Extending the discussion of double integrals, wemay write triple integrals as three iterated integrals, for example,
I=integraldisplay
x2
x1dxintegraldisplayy2(x)
y1(x)dyintegraldisplayz2(x,y)
z1(x,y)dz f(x, y, z),
where the limits on each of the integrals describe the values that x,yandztake
on the boundary of the region R. As for double integrals, in most cases the order
of integration does not affect the value of the integral.
We can extend these ideas to define multiple integrals of higher dimensionality
in a similar way.
6.3 Applications of multiple integrals
Multiple integrals have many uses in the physical sciences, since there are numer-
ous physical quantities which can be written in terms of them. We now discuss a
few of the more common examples.
6.3.1 Areas and volumes
Multiple integrals are often used in finding areas and volumes. For example, the
integral
A=integraldisplay
RdA=integraldisplayintegraldisplay
Rdx dy
is simply equal to the area of the region R. Similarly, if we consider the surface
z=f(x, y) in three-dimensional Cartesian coordinates then the volume under this
surface that stands vertically above the region Ris given by the integral
V=integraldisplay
Rzd A=integraldisplayintegraldisplay
Rf(x, y)dx dy,
where volumes above the xy-plane are counted as positive, and those below as
negative.
194
6.3 APPLICATIONS OF MULTIPLE INTEGRALS
z
c
dx
a
xdz
dyb ydV=dx dy dz
Figure 6.3 The tetrahedron bounded by the coordinate surfaces and the
plane x/a+y/b+z/c= 1 is divided up into vertical slabs, the slabs into
columns and the columns into small boxes.IFind the volume of the tetrahedron bounded by the three coordinate surfaces x=0,y=0
andz=0and the plane x/a+y/b+z/c=1.
Referring to figure 6.3, the elemental volume of the shaded region is given by dV=zd xd y ,
and we must integrate over the triangular region Rin the xy-plane whose sides are x=0 ,
y=0a n d y=b−bx/a. The total volume of the tetrahedron is therefore given by
V=
ZZ
Rzd xd y =
Za
0dx
Zb−bx/a
0dy c
/
1−y
b−x
a
/
=c
Za
0dx
/
y−y2
2b−xy
a
/y=b−bx/a
y=0
=c
Za
0dx
/bx2
2a2−bx
a+b
2
/
=abc
6.
J
Alternatively, we can write the volume of a three-dimensional region Ras
V=integraldisplay
RdV=integraldisplayintegraldisplayintegraldisplay
Rdx dy dz, (6.7)
where the only difficulty occurs in setting the correct limits on each of the
integrals. For the above example, writing the volume in this way corresponds to
dividing the tetrahedron into elemental boxes of volume dx dy dz (as shown in
figure 6.3); integration over zthen adds up the boxes to form the shaded column
in the figure. The limits of integration are z=0t o z=cparenleftbig
1−y/b−x/aparenrightbig
,a n d
195
MULTIPLE INTEGRALS
the total volume of the tetrahedron is given by
V=integraldisplaya
0dxintegraldisplayb−bx/a
0dyintegraldisplayc(1−y/b−x/a)
0dz, (6.8)
which clearly gives the same result as above. This method is illustrated further in
the following example.IFind the volume of the region bounded by the paraboloid z=x2+y2and the plane
z=2y.
The required region is shown in figure 6.4. In order to write the volume of the region in
the form (6.7), we must deduce the limits on each of the integrals. Since the integrationscan be performed in any order, let us first divide the region into vertical slabs of thicknessdyperpendicular to the y-axis, and then as shown in the figure we cut each slab into
horizontal strips of height dz, and each strip into elemental boxes of volume dV=dx dy dz .
Integrating first with respect to x(adding up the elemental boxes to get a horizontal strip),
the limits on xarex=−p
z−y2tox=
p
z−y2. Now integrating with respect to z
(adding up the strips to form a vertical slab) the limits on zarez=y2toz=2y. Finally,
integrating with respect to y(adding up the slabs to obtain the required region), the limits
onyarey=0a n d y= 2, the solutions of the simultaneous equations z=02+y2and
z=2y. So the volume of the region is
V=
Z2
0dy
Z2y
y2dz
Z√
z−y2
−√
z−y2dx=
Z2
0dy
Z2y
y2dz2
p
z−y2
=
Z2
0dy
/4
3(z−y2)3/2
/z=2y
z=y2=
Z2
0dy4
3(2y−y2)3/2.
The integral over ymay be evaluated straightforwa rdly by making the substitution y=
1+s i n u, and gives V=π/2.
J
In general, when calculating the volume (area) of a region, the volume (area)
elements need not be small boxes as in the previous example, but may be of anyconvenient shape. They are usually chosen to make the evaluation of the integralas simple as possible.
6.3.2 Masses, centres of mass and centroids
It is sometimes necessary to calculate the mass of a given object having a non-
uniform density. Symbolically, this mass is given simply by
M=integraldisplay
dM,
where dMis the element of mass and the integral is taken over the extent of the
object. For a solid three-dimensional body the element of mass is just dM=ρd V,
where dVis an element of volume and ρis the variable density. For a laminar
body (i.e. a uniform sheet of material) the element of mass is dM=σd A,w h e r e
σis the mass per unit area of the body and dAis an area element. Finally, for
a body in the form of a thin wire we have dM=λd s,w h e r e λis the mass per
196
6.3 APPLICATIONS OF MULTIPLE INTEGRALS
z
02 y
dV=dx dy dzz=2y
z=x2+y2
x
Figure 6.4 The region bounded by the paraboloid z=x2+y2and the plane
z=2yis divided into vertical slabs, the slabs into horizontal strips and the
strips into boxes.
unit length and dsis an element of arc length along the wire. When evaluating
the required integral, we are free to divide up the body into mass elements in
the most convenient way, provided that over each mass element the density isapproximately constant.IFind the mass of the tetrahedron bounded by the three coordinate surfaces and the plane
x/a+y/b+z/c=1, if its density is given by ρ(x, y, z)=ρ0(1 +x/a).
From (6.8), we can immediately write down the mass of the tetrahedron as
M=
Z
Rρ0
/
1+x
a
/
dV=
Za
0dx ρ 0
/
1+x
a
/
Zb−bx/a
0dy
Zc(1−y/b−x/a)
0dz,
where we have taken the density outside the integrations with respect to zandysince it
depends only on x. Therefore the integrations with respect to zandyproceed exactly as
they did when finding the volume of the tetrahedron, and we have
M=cρ0
Za
0dx
/
1+x
a
/
/bx2
2a2−bx
a+b
2
/
. (6.9)
We could have arrived at (6.9) more directly by dividing the tetrahedron into triangular
slabs of thickness dxperpendicular to the x-axis (see figure 6.3), each of which is of
constant density, since ρdepends on xalone. A slab at a position xhas volume dV=
1
2c(1−x/a)(b−bx/a)dxand mass dM=ρd V=ρ0(1 + x/a)dV. Integrating over xwe
again obtain (6.9). This integral is easily evaluated and gives M=5
24abcρ 0.
J
197
MULTIPLE INTEGRALS
The coordinates of the centre of mass of a solid or laminar body may also be
written as multiple integrals. The centre of mass of a body has coordinates ¯x,¯y,
¯z) given by the three equations
¯xintegraldisplay
dM=integraldisplay
xd M
¯yintegraldisplay
dM=integraldisplay
yd M
¯zintegraldisplay
dM=integraldisplay
zd M ,
where again dMis an element of mass as described above, x,y,zare the
coordinates of the centre of mass of the element dMand the integrals are taken
over the entire body. Obviously, for any body that lies entirely in, or is symmetricalabout, the xy-plane (say), we immediately have ¯z= 0. For completeness, we note
that the three equations above can be written as the single vector equation (seechapter 7)
¯r=1
Mintegraldisplay
rdM,
where ¯ris the position vector of the body’s centre of mass with respect to the
origin, ris the position vector of the centre of mass of the element dMand
M=integraltext
dMis the total mass of the body. As previously, we may divide the body
into the most convenient mass elements for evaluating the necessary integrals,provided each mass element is of constant density.
We further note that the coordinates of the centroid of a body are defined as
those that its centre of mass would have if the body had uniform density.IFind the centre of mass of the solid hemisphere bounded by the surfaces x2+y2+z2=a2
and the xy-plane, assuming that it has a uniform density ρ.
Referring to figure 6.5, we know from symmetry that the centre of mass must lie on
thez-axis. Let us divide the hemisphere into volume elements that are circular slabs of
thickness dzparallel to the xy-plane. For a slab at a height z, the mass of the element is
dM=ρd V=ρπ(a2−z2)dz. Integrating over z, we find that the z-coordinate of the centre
of mass of the hemisphere is given by
¯z
Za
0ρπ(a2−z2)dz=
Za
0zρπ(a2−z2)dz.
The integrals are easily evaluated and give ¯z=3a/8. Since the hemisphere is of uniform
density, this is also the position of its centroid.
J
6.3.3 Pappus’ theorems
The theorems of Pappus (which are about seventeen centuries old) relate centroids
to volumes of revolution and areas of surfaces, discussed in chapter 2, and can beuseful for finding one quantity given another that may be calculated more easily.
198
6.3 APPLICATIONS OF MULTIPLE INTEGRALS
z
xya
a
a√
a2−z2
dz
Figure 6.5 The solid hemisphere bounded by the surfaces x2+y2+z2=a2
and the xy-plane.
A
yy
xdA
¯y
Figure 6.6 An area Ain the xy-plane, which may be rotated about the x-axis
to form a volume of revolution.
If a plane area is rotated about an axis that does not intersect it then the solid
so generated is called a volume of revolution .Pappus’ first theorem states that the
volume of such a solid is given by the plane area Amultiplied by the distance
moved by its centroid (see figure 6.6). This may be proved by considering thedefinition of the centroid of the plane area as the position of the centre of massif the density is uniform, so that
¯y=1
Aintegraldisplay
yd A .
Now the volume generated by rotating the plane area about the x- a x i si sg i v e nb y
V=integraldisplay
2πy dA =2π¯yA,
which is the area multiplied by the distance moved by the centroid.
199
MULTIPLE INTEGRALS
y
yds
¯y
x
Figure 6.7 A curve in the xy-plane, which may be rotated about the x-axis
to form a surface of revolution.
Pappus’ second theorem states that if a plane curve is rotated about a coplanar
axis that does not intersect it then the area of the surface of revolution so generated
is given by the length of the curve Lmultiplied by the distance moved by its
centroid (see figure 6.7). This may be proved in a similar manner to the first
theorem by considering the definition of the centroid of a plane curve,
¯y=1
Lintegraldisplay
yd s ,
and noting that the surface area generated is given by
S=integraldisplay
2πy ds=2π¯yL,
which is equal to the length of the curve multiplied by the distance moved by its
centroid.IA semicircular uniform lamina is freely suspended from one of its corners. Show that its
straight edge makes an angle of 23.0◦with the vertical.
Referring to figure 6.8, the suspended lamina will have its centre of gravity Cvertically
below the suspension point and its straight edge will make an angle θ=t a n−1(d/a)w i t h
the vertical, where 2 ais the diameter of the semicircle and dis the distance of its centre
of mass from the diameter.
Since rotating the lamina about the diameter generates a sphere of volume4
3πa3, Pappus’
first theorem requires that
4
3πa3=2π×d×1
2πa2.
Hence d=4
3a/πandθ=t a n−1(4
3π)=2 3 .0◦.
J
200
6.3 APPLICATIONS OF MULTIPLE INTEGRALS
aθ
dC
Figure 6.8 Suspending a semicircular lamina from one of its corners.
6.3.4 Moments of inertia
For problems in rotational mechanics it is often necessary to calculate the moment
of inertia of a body about a given axis. This is defined by the multiple integral
I=integraldisplay
l2dM,
where lis the distance of a mass element dMfrom the axis. We may again choose
mass elements convenient for evaluating the integral. In this case, however, inaddition to elements of constant density we require all parts of each element tobe at approximately the same distance from the axis about which the moment of
inertia is required.IFind the moment of inertia of a uniform rectangular lamina of mass Mwith sides aand
babout one of the sides of length b.
Referring to figure 6.9, we wish to calculate the moment of inertia about the y-axis.
We therefore divide the rectangular lamina into elemental strips parallel to the y-axis of
width dx. The mass of such a strip is dM=σbdx,w h e r e σis the mass per unit area of
the lamina. The moment of inertia of a strip at a distance xfrom the y-axis is simply
dI=x2dM=σbx2dx. The total moment of inertia of the lamina about the y-axis is
therefore
I=
Za
0σbx2dx=σba3
3.
Since the total mass of the lamina is M=σab,w ec a nw r i t e I=1
3Ma2.
J
201
MULTIPLE INTEGRALS
y
xb
dx adM=σbdx
Figure 6.9 A uniform rectangular lamina of mass Mwith sides aandbcan
be divided into vertical strips.
6.3.5 Mean values of functions
In chapter 2 we discussed average values for functions of a single variable. This
is easily extended to functions of several variables. Let us consider, for example,a function f(x, y) defined in some region Rof the xy-plane. Then the average
value ¯fof the function is given by
¯fintegraldisplay
RdA=integraldisplay
Rf(x, y)dA. (6.10)
This definition is easily extended to three (and higher) dimensions; if a function
f(x, y, z) is defined in some three-dimensional region of space Rthen the average
value ¯fof the function is given by
¯fintegraldisplay
RdV=integraldisplay
Rf(x, y, z)dV. (6.11)IA tetrahedron is bounded by the three coordinate surfaces and the plane x/a+y/b+z/c=
1and has density ρ(x, y, z)=ρ0(1 +x/a). Find the average value of the density.
From (6.11), the average value of the density is given by
¯ρ
Z
RdV=
Z
Rρ(x, y, z)dV.
Now the integral on the LHS is just the volume of the tetrahedron, which we found in
subsection 6.3.1 to be V=1
6abc, and the integral on the RHS is its mass M=5
24abcρ 0,
calculated in subsection 6.3.2. Therefore ¯ρ=M/V =5
4ρ0.
J
6.4 Change of variables in multiple integrals
It often happens that, either because of the form of the integrand involved or
because of the boundary shape of the region of integration, it is desirable to
202
6.4 CHANGE OF VARIABLES IN MULTIPLE INTEGRALS
y
xu=c o n s t a n t
v=c o n s t a n t
NM
L
KR
C
Figure 6.10 A region of integration Roverlaid with a grid formed by the
family of curves u=c o n s t a n ta n d v= constant. The parallelogram KLMN
defines the area element dAuv.
express a multiple integral in terms of a new set of variables. We now consider
h o wt od ot h i s .
6.4.1 Change of variables in double integrals
Let us begin by examining the change of variables in a double integral. Suppose
that we require to change an integral
I=integraldisplayintegraldisplay
Rf(x, y)dx dy,
in terms of coordinates xandy, into one expressed in new coordinates uandv,
given in terms of xandyby differentiable equations u=u(x, y)a n d v=v(x, y)
with inverses x=x(u, v)a n d y=y(u, v). The region Rin the xy-plane and the
curve Cthat bounds it will become a new region R/primeand a new boundary C/primein
theuv-plane, and so we must change the limits of integration accordingly. Also,
the function f(x, y) becomes a new function g(u, v)o ft h en e wc o o r d i n a t e s .
Now the part of the integral that requires most consideration is the area element.
In the xy-plane the element is the rectangular area dAxy=dx dy generated by
constructing a grid of straight lines parallel to the x-a n d y- axes respectively.
Our task is to determine the corresponding area element in the uv-coordinates. In
general the corresponding element dAuvwill not be the same shape as dAxy, but
this does not matter since all elements are infinitesimally small and the value ofthe integrand is considered constant over them. Since the sides of the area element
are infinitesimal, dA
uvwill in general have the shape of a parallelogram. We can
find the connection between dAxyanddAuvby considering the grid formed by the
family of curves u=c o n s t a n ta n d v= constant, as shown in figure 6.10. Since v
203
MULTIPLE INTEGRALS
is constant along the line element KL, the latter has components ( ∂x/∂u )duand
(∂y/∂u )duin the directions of the x-a n d y-axes respectively. Similarly, since u
is constant along the line element KN, the latter has corresponding components
(∂x/∂v )dvand ( ∂y/∂v )dv. Using the result for the area of a parallelogram given
in chapter 7, we find that the area of the parallelogram KLMN is given by
dAuv=vextendsinglevextendsinglevextendsinglevextendsingle∂x
∂udu∂y
∂vdv−∂x
∂vdv∂y
∂uduvextendsinglevextendsinglevextendsinglevextendsingle
=vextendsinglevextendsinglevextendsinglevextendsingle∂x
∂u∂y
∂v−∂x
∂v∂y
∂uvextendsinglevextendsinglevextendsinglevextendsingledu dv.
Defining the Jacobian ofx,ywith respect to u,vas
J=∂(x, y)
∂(u, v)≡∂x
∂u∂y
∂v−∂x
∂v∂y
∂u,
we have
dAuv=vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y)
∂(u, v)vextendsinglevextendsinglevextendsinglevextendsingledu dv.
The reader acquainted with determinants will notice that the Jacobian can also
be written as the 2 ×2 determinant
J=∂(x, y)
∂(u, v)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂u∂y
∂u
∂x
∂v∂y
∂vvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
Such determinants can in general be evaluated using the methods of chapter 8.
So, in summary, the relationship between the size of the area element generated
bydx,dyand the size of the corresponding area element generated by du,dvis
dx dy =vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y)
∂(u, v)vextendsinglevextendsinglevextendsinglevextendsingledu dv.
This equality should be taken as meaning that when transforming from coordi-
nates x, yto coordinates u, v, the area element dx dy should be replaced by the
expression on the RHS of the above equality. Of course, the Jacobian can, andin general will, vary over the region of integration. We may express the doubleintegral in either coordinate system as
I=integraldisplayintegraldisplay
Rf(x, y)dx dy =integraldisplayintegraldisplay
R/primeg(u, v)vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y)
∂(u, v)vextendsinglevextendsinglevextendsinglevextendsingledu dv. (6.12)
When evaluating the integral in the new coordinate system, it is usually advisable
to sketch the region of integration R/primein the uv-plane.
204
6.4 CHANGE OF VARIABLES IN MULTIPLE INTEGRALSIEvaluate the double integral
I=
ZZ
R
/
a+
p
x2+y2
/
dx dy,
where Ris the region bounded by the circle x2+y2=a2.
In Cartesian coordinates, the integral may be written
I=
Za
−adx
Z√
a2−x2
−√
a2−x2dy
/
a+
p
x2+y2
/
,
and can be calculated directly. However, because of the circular boundary of the integration
region, a change of variables to plane polar coordinates ρ,φis indicated. The relationship
between Cartesian and plane polar coordinates is given by x=ρcosφandy=ρsinφ.
Using (6.12) we can therefore write
I=
ZZ
R/prime(a+ρ)
////∂(x, y)
∂(ρ, φ)
////dρ dφ,
where R/primeis the rectangular region in the ρφ-plane whose sides are ρ=0 , ρ=a,φ=0
andφ=2π. The Jacobian is easily calculated, and we obtain
J=∂(x, y)
∂(ρ, φ)=
////cosφ sinφ
−ρsinφρcosφ
////=ρ(cos2φ+s i n2φ)=ρ.
So the relationship between the area elements in Cartesian and in plane polar coordinates is
dx dy =ρd ρd φ .
Therefore, when expressed in plane polar coordinates, the integral is given by
I=
ZZ
R/prime(a+ρ)ρd ρd φ
=
Z2π
0dφ
Za
0dρ(a+ρ)ρ=2π
/aρ2
2+ρ3
3
/a
0=5πa3
3.
J
6.4.2 Evaluation of the integral I=integraltext∞
−∞e−x2dx
By making a judicious change of variables, it is sometimes possible to evaluate
an integral that would be intractable otherwise. An important example of this
method is provided by the evaluation of the integral
I=integraldisplay∞
−∞e−x2dx.
Its value may be found by first constructing I2, as follows:
I2=integraldisplay∞
−∞e−x2dxintegraldisplay∞
−∞e−y2dy=integraldisplay∞
−∞dxintegraldisplay∞
−∞dy e−(x2+y2)
=integraldisplayintegraldisplay
Re−(x2+y2)dx dy,
205
MULTIPLE INTEGRALS
a
a
−a−ay
x
Figure 6.11 The regions used to illustrate the convergence properties of the
integral I(a)=
Ra
−ae−x2dxasa→∞.
where the region Ris the whole xy-plane. Then, transforming to plane polar
coordinates, we find
I2=integraldisplayintegraldisplay
R/primee−ρ2ρd ρd φ =integraldisplay2π
0dφintegraldisplay∞
0dρ ρe−ρ2=2πbracketleftBig
−1
2e−ρ2bracketrightBig∞
0=π.
Therefore the original integral is given by I=√π. Because the integrand is an
even function of x, it follows that the value of the integral from 0 to ∞is simply√π/2.
We note, however, that unlike in all the previous examples, the regions of
integration RandR/primeare both infinite in extent (i.e. unbounded). It is therefore
prudent to derive this result more rigorously; this we do by considering theintegral
I(a)=integraldisplay
a
−ae−x2dx.
We then have
I2(a)=integraldisplayintegraldisplay
Re−(x2+y2)dx dy,
where Ri st h es q u a r eo fs i d e2 acentred on the origin. Referring to figure 6.11,
since the integrand is always positive the value of the integral taken over the
square lies between the value of the integral taken over the region bounded by
the inner circle of radius aand the value of the integral taken over the outer
circle of radius√
2a. Transforming to plane polar coordinates as above, we may
206
6.4 CHANGE OF VARIABLES IN MULTIPLE INTEGRALS
z
xyCR
T
SPQu=c1v=c2
w=c3
Figure 6.12 A three-dimensional region of integration R, showing an el-
ement of volume in u, v, w coordinates formed by the coordinate surfaces
u=c o n s t a n t , v=c o n s t a n t , w=c o n s t a n t .
evaluate the integrals over the inner and outer circles respectively, and we find
πparenleftBig
1−e−a2parenrightBig
<I2(a)<πparenleftBig
1−e−2a2parenrightBig
.
Taking the limit a→∞, we find I2(a)→π.T h e r e f o r e I=√πas we found
previously. We use this result in the discussion of the normal distribution inchapter 26.
6.4.3 Change of variables in triple integrals
A change of variable in a triple integral follows the same general lines as that for
a double integral. Suppose we wish to change variables from x, y, z tou, v, w .
In the x, y, z coordinates the element of volume is a cuboid of sides dx, dy, dz
and volume dV
xyz=dx dy dz . If, however, we divide up the total volume into
infinitesimal elements by constructing a grid formed from the coordinate surfacesu=c o n s t a n t , v=c o n s t a n ta n d w= constant, then the element of volume dV
uvw
in the new coordinates will have the shape of a parallelepiped whose faces are the
coordinate surfaces and whose edges are the curves formed by the intersections
of these surfaces (see figure 6.12). Along the line element PQthe coordinates v
andware constant, and so PQhas components of ( ∂x/∂u )du,(∂y/∂u )duand
207
MULTIPLE INTEGRALS
(∂z/∂u )duin the direction of the x-,y-a n d z- axes respectively. The components
of the line elements PSandSTare found by replacing ubyvandwrespectively.
The expression for the volume of a parallelepiped in terms of the components
of its edges with respect to the x-,y-a n d z-axes is given in chapter 7. Using this,
we find that the element of volume in u, v, w coordinates is given by
dVuvw=vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y, z)
∂(u, v, w)vextendsinglevextendsinglevextendsinglevextendsingledu dv dw,
where the Jacobian of x, y, z with respect to u, v, w is a short-hand for a 3 ×3
determinant:
∂(x, y, z)
∂(u, v, w)≡vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂u∂y
∂u∂z
∂u
∂x
∂v∂y
∂v∂z
∂v
∂x
∂w∂y
∂w∂z
∂wvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
So, in summary, the relationship between the elemental volumes in multiple
integrals formulated in the two coordinate systems is given in Jacobian form by
dx dy dz =vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y, z)
∂(u, v, w)vextendsinglevextendsinglevextendsinglevextendsingledu dv dw,
and we can write a triple integral in either set of coordinates as
I=integraldisplayintegraldisplayintegraldisplay
Rf(x, y, z)dx dy dz =integraldisplayintegraldisplayintegraldisplay
R/primeg(u, v, w)vextendsinglevextendsinglevextendsinglevextendsingle∂(x, y, z)
∂(u, v, w)vextendsinglevextendsinglevextendsinglevextendsingledu dv dw.IFind an expression for a volume element in sphe rical polar coordinates, and hence calcu-
late the moment of inertia about a diameter of a uniform sphere of radius aand mass M.
Spherical polar coordinates r,θ,φ are defined by
x=rsinθcosφ, y =rsinθsinφ, z =rcosθ
(and are discussed fully in chapter 10). The required Jacobian is therefore
J=∂(x, y, z)
∂(r,θ,φ)=
//////sinθcosφ sinθsinφ cosθ
rcosθcosφr cosθsinφ−rsinθ
−rsinθsinφrsinθcosφ 0
//////.
The determinant is most easily evaluated by expanding it with respect to the last column
(see chapter 8), which gives
J=c o s θ(r2sinθcosθ)+rsinθ(rsin2θ)
=r2sinθ(cos2θ+s i n2θ)=r2sinθ.
Therefore the volume element in spherical polar coordinates is given by
dV=∂(x, y, z)
∂(r,θ,φ)dr dθ dφ =r2sinθd rd θd φ ,
which agrees with the result given in chapter 10.
208
6.4 CHANGE OF VARIABLES IN MULTIPLE INTEGRALS
If we place the sphere with its centre at the origin of an x, y, z coordinate system then
its moment of inertia about the z-axis (which is, of course, a diameter of the sphere) is
I=
Z/;
x2+y2
/
dM=ρ
Z/;
x2+y2
/
dV,
where the integral is taken over the sphere, and ρis the density. Using spherical polar
coordinates, we can write this as
I=ρ
ZZZ
V
/;
r2sin2θ
/
r2sinθd rd θd φ
=ρ
Z2π
0dφ
Zπ
0dθsin3θ
Za
0dr r4
=ρ×2π×4
3×1
5a5=8
15πa5ρ.
Since the mass of the sphere is M=4
3πa3ρ, the moment of inertia can also be written as
I=2
5Ma2.
J
6.4.4 General properties of Jacobians
Although we will not prove it, the general result for a change of coordinates in
ann-dimensional integral from a set xito a set yj(where iandjboth run from
1t on)i s
dx1dx2···dxn=vextendsinglevextendsinglevextendsinglevextendsingle∂(x
1,x2,...,x n)
∂(y1,y2,...,y n)vextendsinglevextendsinglevextendsinglevextendsingledy
1dy2···dyn,
where the n-dimensional Jacobian can be written as an n×ndeterminant (see
chapter 8) in an analogous way to the two- and three-dimensional cases.
For readers who already have sufficient familiarity with matrices (see chapter 8)
and their properties, a fairly compact proof of some useful general properties
of Jacobians can be given as follows. Other readers should turn straight to theresults (6.16) and (6.17) and return to the proof at some later time.
Consider three sets of variables x
i,yiandzi, with irunning from 1 to nfor
each set. From the chain rule in partial differentiation (see (5.17)), we know that
∂xi
∂zj=nsummationdisplay
k=1∂xi
∂yk∂yk
∂zj. (6.13)
Now let A,Band Cbe the matrices whose ijth elements are ∂xi/∂y j,∂yi/∂z jand
∂xi/∂z jrespectively. We can then write (6.13) as the matrix product
cij=nsummationdisplay
k=1aikbkj or C=AB. (6.14)
We may now use the general result for the determinant of the product of two
matrices, namely |AB|=|A||B|, and recall that the Jacobian
Jxy=∂(x1,...,x n)
∂(y1,...,y n)=|A|, (6.15)
209
MULTIPLE INTEGRALS
and similarly for JyzandJxz. On taking the determinant of (6.14), we therefore
obtain
Jxz=JxyJyz
or, in the usual notation,
∂(x1,...,x n)
∂(z1,...,z n)=∂(x1,...,x n)
∂(y1,...,y n)∂(y1,...,y n)
∂(z1,...,z n). (6.16)
As a special case, if the set zii st a k e nt ob ei d e n t i c a lt ot h es e t xi,a n dt h e
obvious result Jxx= 1 is used, we obtain
JxyJyx=1
or, in the usual notation,
∂(x1,...,x n)
∂(y1,...,y n)=bracketleftbigg∂(y1,...,y n)
∂(x1,...,x n)bracketrightbigg−1
. (6.17)
The similarity between the properties of Jacobians and those of derivatives is
apparent, and to some extent is suggested by the notation. We further note from(6.15) that since |A|=|A
T|,w h e r e ATis the transpose of A, we can interchange the
rows and columns in the determinantal form of the Jacobian without changingits value.
6.5 Exercises
6.1 Sketch the curved wedge bounded by the surfaces y2=4ax,x+z=aandz=0 ,
and hence calculate its volume V.
6.2 Evaluate the volume integral of x2+y2+z2over the rectangular parallelepiped
bounded by the six surfaces x=±a,y=±b,z=±c.
6.3 Find the volume integral of x2yover the tetrahedral volume bounded by the
planes x=0 , y=0 , z=0 ,a n d x+y+z=1 .
6.4 Evaluate the surface integral of f(x, y) over the rectangle 0 ≤x≤a,0≤y≤b
for the functions
(a)f(x, y)=x
x2+y2, (b)f(x, y)=(b−y+x)−3/2.
6.5 (a) Prove that the area of the ellipse
x2
a2+y2
b2=1
isπab.
(b) Use this result to obtain an expression for the volume of a slice of thickness
dzof the ellipsoid
x2
a2+y2
b2+z2
c2=1.
Hence show that the volume of the ellipsoid is 4 πabc/ 3.
210
6.5 EXERCISES
6.6 The function
Ψ(r)=A
/
2−Zr
a
/
e−Zr/2a
gives the form of the quantum mechanical wavefunction representing the electron
in a hydrogen-like atom of atomic number Zwhen the electron is in its first
allowed spherically symmetric excited state. Here ris the usual spherical polar
coordinate, but, because of the spherical symmetry, the coordinates θandφdo
not appear explicitly in Ψ. Determine the value that A(assumed real) must have
if the wavefunction is to be correctly normalised, i.e. the volume integral of |Ψ|2
over all space is equal to unity.
6.7 In quantum mechanics the electron in a hydrogen atom in some particular state
is described by a wavefunction Ψ, which is such that |Ψ|2dVis the probability of
finding the electron in the infinitesimal volume dV. In spherical polar coordinates
Ψ=Ψ ( r,θ,φ)a n d dV=r2sinθd rd θd φ . Two such states are described by
Ψ1=
/1
4π
/1/2
/1
a0
/3/2
2e−r/a0,
Ψ2=−
/3
8π
/1/2
sinθeiφ
/1
2a0
/3/2re−r/2a0
a0√
3.
(a) Show that each Ψ iis normalised, i.e. the integral over all space
R
|Ψ|2dVis
equal to unity – physically, this means that the electron must be somewhere.
(b) The (so-called) dipole matrix element between the states 1 and 2 is given by
the integral
px=
Z
Ψ∗
1qrsinθcosφΨ2dV,
where qis the charge on the electron. Prove that pxhas the value −27qa0/35.
6.8 A planar figure is formed from uniform wire and consists of two semicircular
arcs, each with its own closing diameter, joined so as to form a letter ‘B’. Thefigure is freely suspended from its top left-hand corner. Show that the straightedge of the figure makes an angle θwith the vertical given by tan θ=( 2+ π)
−1.
6.9 A certain torus has a circular vertical cross-section of radius acentred on a
horizontal circle of radius c(>a).
(a) Find the volume Vand surface area Aof the torus, and show that they can
be written as
V=π2
4(r2
o−r2
i)(ro−ri),A =π2(r2
o−r2
i),
where roandroare respectively the outer and inner radii of the torus.
(b) Show that a vertical circular cylinder of radius c, coaxial with the torus,
divides Ain the ratio
πc+2a:πc−2a.
6.10 A thin uniform circular disc has mass Mand radius a.
(a) Prove that its moment of inertia about an axis perpendicular to its plane
and passing through its centre is1
2Ma2.
(b) Prove that the moment of inertia of the same disc about a diameter is1
4Ma2.
This is an example of the general result for planar bodies that the moment of
inertia of the body about an axis perpendicular to the plane is equal to the sum
211
MULTIPLE INTEGRALS
of the moments of inertia about two perpendicular axes lying in the plane: in an
obvious notation
Iz=
Z
r2dm=
Z
(x2+y2)dm=
Z
x2dm+
Z
y2dm=Iy+Ix.
6.11 In some applications in mechanics the moment of inertia of a body about a
single point (as opposed to about an axis) is needed. The moment of inertia I
about the origin of a uniform solid body of density ρis given by the volume
integral
I=
Z
V(x2+y2+z2)ρd V.
Show that the moment of inertia of a right circular cylinder of radius a,l e n g t h
2b,a n dm a s s Mabout its centre is
M
/a2
2+b2
3
/
.
6.12 The shape of an axially symmetric hard-boiled egg, of uniform density ρ0,i s
given in spherical polar coordinates by r=a(2−cosθ), where θis measured
from the axis of symmetry.
(a) Prove that the mass Mof the egg is M=40
3πρ0a3.
(b) Prove that the egg’s moment of inertia about its axis of symmetry is342
175Ma2.
6.13 In spherical polar coordinates r, θ, φ the element of volume for a body that
is symmetrical about the polar axis is dV=2πr2sinθd rd θ , whilst its element
of surface area is 2 πrsinθ[(dr)2+r2(dθ)2]1/2. A particular surface is defined by
r=2acosθ,w h e r e ais a constant, and 0 ≤θ≤π/2. Find its total surface area
and the volume it encloses, and hence identify the surface.
6.14 By expressing both the integrand and the surface element in spherical polar
coordinates, show that the surface integralZx2
x2+y2dS
over the surface x2+y2=z2,0≤z≤1, has the value π/√2.
6.15 By transforming to cylindrical polar coordinates, evaluate the integral
I=
Z Z Z
ln(x2+y2)dx dy dz
over the interior of the conical region x2+y2≤z2,0≤z≤1.
6.16 Sketch the two families of curves
y2=4u(u−x),y2=4v(v+x),
where uandvare parameters.
By transforming to the uv-plane evaluate the integral of y/(x2+y2)1/2over
that part of the quadrant x>0,y>0 bounded by the lines x=0 , y=0a n d
the curve y2=4a(a−x).
6.17 By making two successive simple changes of variables, evaluate
I=
Z Z Z
x2dx dy dz
over the ellipsoidal region
x2
a2+y2
b2+z2
c2≤1.
212
6.5 EXERCISES
6.18 Sketch the domain of integration for the integral
I=
Z1
0
Z1/y
x=yy3
xexp[y2(x2+x−2)]dx dy
and characterise its boundaries in terms of new variables u=xyandv=y/x.
Show that the Jacobian for the change from ( x, y)t o( u, v)i se q u a lt o( 2 v)−1,a n d
hence evaluate I.
6.19 Sketch that part of the region 0 ≤x,0≤y≤π/2 which is bounded by the
curves x=0 , y=0 ,s i n h xcosy=1a n dc o s h xsiny= 1. By making a suitable
change of variables, evaluate the integral
I=
Z Z
(sinh2x+c o s2y)si nh2 xsin 2yd xd y
over the bounded sub-region.
6.20 Define a coordinate system u, vwhose origin coincides with that of the usual
x, ysystem and whose u-axis coincides with the x-axis, whilst the v-axis makes
an angle αwith it. By considering the integral I=
R
exp(−r2)dA,w h e r e ris the
radial distance from the origin, over the area defined by 0 ≤u<∞,0≤v<∞,
prove thatZ∞
0
Z∞
0exp(−u2−v2−2uvcosα)du dv=α
2si nα.
6.21 As stated in section 5.11, the first law of thermodynamics can be expressed as
dU=TdS−PdV.
By calculating and equating ∂2U/∂Y ∂X and∂2U/∂X∂Y ,w h e r e XandYare an
unspecified pair of variables (drawn from P,V,T andS), prove that
∂(S,T)
∂(X,Y)=∂(V,P)
∂(X,Y).
Using the properties of Jacobians, deduce that
∂(S,T)
∂(V,P)=1.
6.22 The distances of the variable point P, which has coordinates x, y, z, from the fixed
points (0 ,0,1) and (0 ,0,−1) are denoted by uandvrespectively. New variables
ξ,η,φ are defined by
ξ=1
2(u+v),η =1
2(u−v),
andφis the angle between the plane y= 0 and the plane containing the three
points. Prove that the Jacobian ∂(ξ,η,φ )/∂(x, y, z) has the value ( ξ2−η2)−1and
thatZ Z Z
all space(u−v)2
uvexp
/
−u+v
2
/
dx dy dz =32π
3e.
6.23 This is a more difficult question about ‘volumes’ in an increasing number of
dimensions.
(a) Let Rbe a real positive number and define Kmby
Km=
ZR
−R
/;
R2−x2
/mdx.
Show, using integration by parts, that Kmsatisfies the recurrence relation
(2m+1 )Km=2mR2Km−1.
213
MULTIPLE INTEGRALS
(b) For integer n, define In=KnandJn=Kn+1/2. Evaluate I0andJ0directly
and hence prove that
In=22n+1(n!)2R2n+1
(2n+1 ) !and Jn=π(2n+1 ) ! R2n+2
22n+1n!(n+1 ) !.
(c) A sequence of functions Vn(R) is defined by
V0(R)=1 ,
Vn(R)=
ZR
−RVn−1
/√
R2−x2
/
dx, n≥1.
Prove by induction that
V2n(R)=πnR2n
n!,V 2n+1(R)=πn22n+1n!R2n+1
(2n+1 ) !.
(d) For interest,
(i) show that V2n+2(1)<V 2n(1) and V2n+1(1)<V 2n−1(1) for all n≥3;
(ii) hence, by explicitly writing out Vk(R)f o r1≤k≤8 (say), show that the
‘volume’ of the totally symmetric solid of unit radius is a maximum in
five dimensions.
6.6 Hints and answers
6.1 For integration in the order z,y,x the limits are (0 ,a−x),(−√
4ax,√
4ax),(0,a).
For integration in the order y,x,z the limits are ( −√
4ax,√
4ax),(0,a−z),(0,a).
V=1 6a3/15.
6.2 8 abc(a2+b2+c2)/3.
6.3 1 /360.
6.4 (a) Integrate by parts to obtain ( b/2)ln[1 + ( a/b)2]+atan−1(b/a);
(b) 4[ a1/2+b1/2−(a+b)1/2].
6.5 (a) Evaluate
R
2b[1−(x/a)2]1/2dxby setting x=acosφ;
(b)dV=π×a[1−(z/c)2]1/2×b[1−(z/c)2]1/2dz.
6.6 A=±(Z/a)3/2/√
32π.
6.8 If one of the semicircles has radius a, Pappus’ second theorem shows that its
centre of gravity /angbracketleftx/angbracketrightis 2a/πfrom the centre of the circle of which it is half. For
the whole figure, /angbracketleftx/angbracketright=4a/(2π+4 ) .
6.9 (a) V=2πc×πa2andA=2πa×2πc. Setting ro=c+aandri=c−agives the
stated results. (b) See hint for previous exercise.
6.10 (b) Evaluate
R
2(a2−x2)1/2x2(M/πa2)dxby setting x=acosφ.
6.11 Transform to cylindrical polar coordinates.6.12 (a) Show that dz=2asinθ(cosθ−1)dθ.W r i t i n gc o s θascto save space, the
integrand is 2 πρ
0a3(1−c2)(1−c)(2−c)2dcover the range −1≤c≤1.
(b) The integrand is πρ0a5(1−c2)2(1−c)(2−c)4dc.
6.13 4 πa2,4πa3/3, a sphere.
6.14 The coordinate ranges are 0 ≤r≤√2a n d0≤φ≤2π,w i t h θ=π/4. The
integrand for the randφintegrations is ( rcos2φ)/√2.
6.15 The volume element is ρd φd ρd z . The integrand for the final z-integration is
given by 2 π[(z2lnz)−(z2/2)];I=−5π/9.
6.16 Jacobian = ( u/v)1/2+(v/u)1/2;a r e ai n uv-plane is the triangle bounded by v=0 ,
u=v,u=a;i n t e g r a l= a2.
6.17 Set ξ=x/a,η=y/b,ζ=z/cto map the ellipsoid onto the unit sphere, and then
change from ( ξ,η,ζ) coordinates to spherical polar coordinates; I=4πa3bc/15.
214
6.6 HINTS AND ANSWERS
6.18 The boundaries of the three-sided region are u=v=0,v=1a n d u=1 .
I=(e−1)2/8.
6.19 Set u=s i n h xcosy,v=c o s h xsiny;Jxy,uv=( s i n h2x+cos2y)−1and the integrand
reduces to 4 uvover the region 0 ≤u≤1, 0≤v≤1;I=1.
6.20 x=vcosα+u,y=vsinα.J a c o b i a n=s i n α.
I=(α/2π)
R
exp(−r2)dAover all space.
6.21 Terms such as T∂2S/∂Y ∂X cancel in pairs. Use equations (6.17) and (6.16).
6.22 Note that uv=(ξ2−η2). The ranges for the new variables are 1 ≤ξ<∞,
−1≤η≤1, 0≤φ≤2π.
6.23 (d)(ii) 2, π,4π/3,π2/2, 8π2/15,π3/6, 16π3/105,π4/24.
215
7
Vector algebra
This chapter introduces space vectors and their manipulation. Firstly we deal
with the description and algebra of vectors and then we consider how vectorsmay be used to describe lines and planes and finally we look at the practical useof vectors in finding distances. Much use of vectors will be made in subsequentchapters; this chapter gives only some basic rules.
7.1 Scalars and vectors
The simplest kind of physical quantity is one that can be completely specified by
its magnitude, a single number, together with the units in which it is measured.Such a quantity is called a scalar and examples include temperature, time and
density.
Avector is a quantity that requires both a magnitude ( ≥0) and a direction in
space to specify it completely; we may think of it as an arrow in space. A familiarexample is force, which has a magnitude (strength) measured in newtons and adirection of application. The large number of vectors that are used to describethe physical world include velocity, displacement, momentum and electric field.Vectors are also used to describe quantities such as angular momentum andsurface elements (a surface element has an area and a direction defined by the
normal to its tangent plane); in such cases their definitions may seem somewhat
arbitrary (though in fact they are standard) and not as physically intuitive as forvectors such as force. A vector is denoted by bold type, the convention of thisbook, or by underlining, the latter being much used in handwritten work.
This chapter considers basic vector algebra and illustrates just how powerful
vector analysis can be. All the techniques are presented for three-dimensionalspace but most can be readily extended to more dimensions.
Throughout the book we will represent vectors in diagrams as a line together
with an arrowhead. We will make no distinction between an arrowhead at the
216
7.2 ADDITION AND SUBTRACTION OF VECTORS
aa
b
b a+bb+a
Figure 7.1 Addition of two vectors showing the commutation relation. We
make no distinction between an arrowhead at the end of the line and one
along the line’s length, but rather use that which gives the clearer diagram.
end of the line or one along the line’s length but, rather, use that which gives the
clearer diagram. Furthermore, even though we are considering three-dimensionalvectors, we have to draw them in the plane of the paper. It should not be assumedthat vectors drawn thus are coplanar, unless this is explicitly stated.
7.2 Addition and subtraction of vectors
Theresultant orvector sum of two displacement vectors is the displacement vector
that results from performing first one and then the other displacement, as shownin figure 7.1; this process is known as vector addition. However, the principleof addition has physical meaning for vector quantities other than displacements;
for example, if two forces act on the same body then the resultant force acting
on the body is the vector sum of the two. The addition of vectors only makesphysical sense if they are of a like kind, for example if they are both forcesacting in three dimensions. It may be seen from figure 7.1 that vector addition iscommutative, i.e.
a+b=b+a. (7.1)
The generalisation of this procedure to the addition of three (or more) vectors is
clear and leads to the associativity property of addition (see figure 7.2), e.g.
a+(b+c)=(a+b)+c. (7.2)
Thus, it is immaterial in what order any number of vectors are added.
The subtraction of two vectors is very similar to their addition (see figure 7.3),
that is,
a−b=a+(−b)
where−bis a vector of equal magnitude but exactly opposite direction to vector b.
217
VECTOR ALGEBRA
aa
ab
bb
ccc
a+(b+c)
(a+b)+cb+cb+c
a+b
a+b
Figure 7.2 Addition of three vectors showing the associativity relation.
−b
ba
aa−b
Figure 7.3 Subtraction of two vectors.
The subtraction of two equal vectors yields the zero vector, 0, which has zero
magnitude and no associated direction.
7.3 Multiplication by a scalar
Multiplication of a vector by a scalar (not to be confused with the ‘scalar
product’, to be discussed in subsection 7.6.1) gives a vector in the same direction
as the original but of a proportional magnitude. This can be seen in figure 7.4.The scalar may be positive, negative or zero. It can also be complex in someapplications. Clearly, when the scalar is negative we obtain a vector pointingin the opposite direction to the original vector. Multiplication by a scalar isassociative, commutative and distributiv e over addition. These properties may be
summarised for arbitrary vectors aandband arbitrary scalars λandµby
(λµ)a=λ(µa)=µ(λa), (7.3)
λ(a+b)=λa+λb, (7.4)
(λ+µ)a=λa+µa. (7.5)
218
7.3 MULTIPLICATION BY A SCALAR
a
aλ
Figure 7.4 Scalar multiplication of a vector (for λ>1).
OAB
P
ab
pµ
λ
Figure 7.5 An illustration of the ratio theorem. The point Pdivides the line
segment ABin the ratio λ:µ.
Having defined the operations of addition, subtraction and multiplication by a
scalar, we can now use vectors to solve simple problems in geometry.IA point Pdivides a line segment ABin the ratio λ:µ(see figure 7.5). If the position
vectors of the points AandBareaandbrespectively, find the position vector of the
point P.
As is conventional for vector geometry problems, we denote the vector from the point A
to the point BbyAB. If the position vectors of the points AandB, relative to some origin
O,a r eaandb, it should be clear that AB=b−a.
Now, from figure 7.5 we see that one possible way of reaching the point Pfrom Ois
first to go from OtoAand to go along the line ABfor a distance equal to the the fraction
λ/(λ+µ) of its total length. We may express this in terms of vectors as
OP=p=a+λ
λ+µAB
=a+λ
λ+µ(b−a)
=
/
1−λ
λ+µ
/
a+λ
λ+µb
=µ
λ+µa+λ
λ+µb, (7.6)
which expresses the position vector of the point Pin terms of those of AandB.W ew o u l d ,
of course, obtain the same result by considering the path from OtoBand then to P.
J
219
VECTOR ALGEBRA
OA
BC
DE
F G
a
bc
Figure 7.6 The centroid of a triangle. The triangle is defined by the points A,
BandCthat have position vectors a,bandc. The broken lines CD,BE,AF
connect the vertices of the triangle to the mid-points of the opposite sides;these lines intersect at the centroid Gof the triangle.
The result (7.6) is a version of the ratio theorem and we may use it in solving
more complicated problems.IThe vertices of triangle ABC have position vectors a,bandcrelative to some origin O
(see figure 7.6). Find the position vector of the centroid Gof the triangle.
From figure 7.6, the points DandEbisect the lines ABandACrespectively. Thus from
the ratio theorem (7.6), with λ=µ=1/2, the position vectors of DandErelative to the
origin are
d=1
2a+1
2b,
e=1
2a+1
2c.
Using the ratio theorem again, we may write the position vector of a general point on the
lineCDthat divides the line in the ratio λ:( 1−λ)a s
r=( 1−λ)c+λd,
=( 1−λ)c+1
2λ(a+b), (7.7)
where we have expressed din terms of aandb. Similarly, the position vector of a general
point on the line BEcan be expressed as
r=( 1−µ)b+µe,
=( 1−µ)b+1
2µ(a+c). (7.8)
Thus, at the intersection of the lines CDandBEwe require, from (7.7), (7.8),
(1−λ)c+1
2λ(a+b)=( 1−µ)b+1
2µ(a+c).
By equating the coefficents of the vectors a,b,cwe find
λ=µ,1
2λ=1−µ, 1−λ=1
2µ.
220
7.4 BASIS VECTORS AND COMPONENTS
These equations are consistent and have the solution λ=µ=2/3. Substituting these
values into either (7.7) or (7.8) we find that the position vector of the centroid Gis given
by
g=1
3(a+b+c).
J
7.4 Basis vectors and components
Given any three different vectors e1,e2ande3, which do not all lie in a plane,
it is possible, in three-dimensional space, to write any other vector in terms ofscalar multiples of them:
a=a
1e1+a2e2+a3e3. (7.9)
The three vectors e1,e2ande3are said to form a basis(for the three-dimensional
space); the scalars a1,a2anda3, which may be positive, negative or zero, are
called the components of the vector awith respect to this basis. We say that the
vector has been resolved into components.
Most often we shall use basis vectors that are mutually perpendicular, for ease
of manipulation, though this is not necessary. In general, a basis set must
(i) have as many basis vectors as the number of dimensions (in more formal
language, the basis vectors must span the space) and
(ii) be such that no basis vector may be described as a sum of the others, or,
more formally, the basis vectors must be linearly independent . Putting this
mathematically, in Ndimensions, we require
c1e1+c2e2+···+cNeN/negationslash=0,
for any set of coefficients c1,c2,...,c Nexcept c1=c2=···=cN=0 .
In this chapter we will only consider vectors in three dimensions; higher dimen-
sionality can be achieved by simple extension.
If we wish to label points in space using a Cartesian coordinate system ( x, y, z),
we may introduce the unit vectors i,jandk, which point along the positive x-,
y-a n d z- axes respectively. A vector amay then be written as a sum of three
vectors, each parallel to a different coordinate axis:
a=axi+ayj+azk. (7.10)
A vector in three-dimensional space thus requires three components to describe
fully both its direction and its magnitude. A displacement in space may be
t h o u g h to fa st h es u mo fd i s p l a c e m e n t sa l o n gt h e x-,y-a n d z- directions (see
figure 7.7). For brevity, the components of a vector awith respect to a particular
coordinate system are sometimes written in the form ( ax,ay,az). Note that the
221
VECTOR ALGEBRA
ijk
axiayj
azka
Figure 7.7 A Cartesian basis set. The vector ais the sum of axi,ayjandazk.
basis vectors i,jandkmay themselves be represented by (1 ,0,0), (0 ,1,0) and
(0,0,1) respectively.
We can consider the addition and subtraction of vectors in terms of their
components. The sum of two vectors aandbis found by simply adding their
components, i.e.
a+b=axi+ayj+azk+bxi+byj+bzk
=(ax+bx)i+(ay+by)j+(az+bz)k (7.11)
and their difference by subtracting them,
a−b=axi+ayj+azk−(bxi+byj+bzk)
=(ax−bx)i+(ay−by)j+(az−bz)k. (7.12)ITwo particles have velocities v1=i+3j+6kandv2=i−2krespectively. Find the
velocity uof the second particle relative to the first.
The required relative velocity is given by
u=v2−v1=( 1−1)i+( 0−3)j+(−2−6)k
=−3j−8k.
J
7.5 Magnitude of a vector
The magnitude of the vector ais denoted by |a|ora. In terms of its components
in three-dimensional Cartesian coordinates, the magnitude of ais given by
a≡|a|=radicalBig
a2x+a2y+a2z. (7.13)
Hence, the magnitude of a vector is a measure of its length. Such an analogy is
useful for displacement vectors but magnitude is better described, for example, by
222
7.6 MULTIPLICATION OF VECTORS
‘strength’ for vectors such as force or by ‘speed’ for velocity vectors. For instance,
in the previous example, the speed of the second particle relative to the first isgiven by
u=|u|=radicalbig
(−3)2+(−8)2=√
73.
A vector whose magnitude equals unity is called a unit vector. The unit vector
in the direction ais usually notated ˆaand may be evaluated as
ˆa=a
|a|. (7.14)
The unit vector is a useful concept because a vector written as λˆathen has mag-
nitude λand direction ˆa. Thus magnitude and direction are explicitly separated.
7.6 Multiplication of vectors
We have already considered multiplying a vector by a scalar. Now we consider
the concept of multiplying one vector by another vector. It is not immediatelyobvious what the product of two vectors represents and in fact two productsare commonly defined, the scalar product and the vector product . As their names
imply, the scalar product of two vectors is just a number, whereas the vectorproduct is itself a vector. Although neither the scalar nor the vector product
is what we might normally think of as a product, their use is widespread and
numerous examples will be described elsewhere in this book.
7.6.1 Scalar product
The scalar product (or dot product) of two vectors aandbis denoted by a·b
and is given by
a·b≡|a||b|cosθ,0≤θ≤π, (7.15)
where θis the angle between the two vectors, placed ‘tail to tail’ or ‘head to head’.
Thus, the value of the scalar product a·bequals the magnitude of amultiplied
by the projection of bontoa(see figure 7.8).
From (7.15) we see that the scalar product has the particularly useful property
that
a·b= 0 (7.16)
is a necessary and sufficient condition for ato be perpendicular to b(unless either
of them is zero). It should be noted in particular that the Cartesian basis vectorsi,jandk, being mutually orthogonal unit vectors, satisfy the equations
i·i=j·j=k·k=1, (7.17)
i·j=j·k=k·i=0. (7.18)
223
VECTOR ALGEBRA
ab
Oθ
bcosθ
Figure 7.8 The projection of bonto the direction of aisbcosθ. The scalar
product of aandbisabcosθ.
Examples of scalar products arise naturally throughout physics and in partic-
ular in connection with energy. Perhaps the simplest is the work done F·rin
moving the point of application of a constant force Fthrough a displacement r;
notice that, as expected, if the displacement is perpendicular to the direction of
the force then F·r= 0 and no work is done. A second simple example is afforded
by the potential energy −m·Bof a magnetic dipole, represented in strength and
orientation by a vector m, placed in an external magnetic field B.
As the name implies, the scalar product has a magnitude but no direction. The
scalar product is commutative and distributive over addition:
a·b=b·a (7.19)
a·(b+c)=a·b+a·c. (7.20)IFour points A, B, C, D are positioned such that the line ADis perpendicular to BCand
BDis perpendicular to AC. Show that CDis perpendicular to AB.
Let us denote the position vectors of the points A, B, C, D bya,b,c,drespectively. As
the four points are not coplanar it is difficult to draw a helpful diagram of the situation,but this is not a drawback when vector methods are used. We start by noting that, sinceAD⊥BC, we have from (7.16) that
(d−a)·(c−b)=0 .
Similarly, since BD⊥AC,
(d−b)·(c−a)=0 .
Combining these two equations we find
(d−a)·(c−b)=(d−b)·(c−a),
which, on mutliplying out the parentheses, gives
d·c−a·c−d·b+a·b=d·c−b·c−d·a+b·a.
Cancelling terms that appear on both sides and rearranging yields
d·b−d·a−c·b+c·a=0,
which simplifies to give
(d−c)·(b−a)=0 .
From (7.16), we see that this implies that CDis perpendicular to AB.J
224
7.6 MULTIPLICATION OF VECTORS
If we introduce a set of basis vectors that are mutually orthogonal, such as i,j,
k, we can write the components of a vector a, with respect to that basis, in terms
of the scalar product of awith each of the basis vectors, i.e. ax=a·i,ay=a·jand
az=a·k. In terms of the components ax,ayandazthe scalar product is given by
a·b=(axi+ayj+azk)·(bxi+byj+bzk)=axbx+ayby+azbz, (7.21)
where the cross terms such as axi·byjare zero because the basis vectors are
mutually perpendicular; see equation (7.18). It should be clear from (7.15) thatthe value of a·bhas a geometrical definition and that this value is independent
of the actual basis vectors used.IFind the angle between the vectors a=i+2j+3kandb=2i+3j+4k.
From (7.15) the cosine of the angle θbetween aandbis given by
cosθ=a·b
|a||b|.
From (7.21) the scalar product a·bhas the value
a·b=1×2+2×3+3×4=2 0 ,
and from (7.13)the lengths of the vectors are
|a|=
p
12+22+32=√
14 and |b|=
p
22+32+42=√
29.
Thus,
cosθ=20√
14√
29≈0.9926⇒ θ=0.12 rad .
J
We can see from the expressions (7.15), (7.21) for the scalar product that if θ
is the angle between aandbthen
cosθ=ax
abx
b+ay
aby
b+az
abz
b
where ax/a,ay/aandaz/aare called the direction cosines ofa, since they give the
cosine of the angle made by awith each of the basis vectors. Similarly bx/b,by/b
andbz/bare the direction cosines of b.
If we take the scalar product of any vector awith itself then clearly θ=0a n d
from (7.15) we have
a·a=|a|2.
Thus the magnitude of acan be written in a coordinate-independent form as
|a|=√a·a.
Finally, we note that the scalar product may be extended to vectors with
complex components if it is redefined as
a·b=a∗
xbx+a∗
yby+a∗
zbz,
where the asterisk represents the operation of complex conjugation. To accom-
225
VECTOR ALGEBRA
θa×b
ab
Figure 7.9 The vector product. The vectors a,banda×bform a right-handed
set.
modate this extension the commutation property (7.19) must be modified to
read
a·b=(b·a)∗. (7.22)
In particular it should be noted that ( λa)·b=λ∗a·b,whereas a·(λb)=λa·b.
However, the magnitude of a complex vector is still given by |a|=√a·a,s i n c e
a·ais always real.
7.6.2 Vector product
The vector product (or cross product) of two vectors aandbis denoted by a×b
and is defined to be a vector of magnitude |a||b|sinθin a direction perpendicular
to both aandb;
|a×b|=|a||b|sinθ.
The direction is found by ‘rotating’ aintobthrough the smallest possible angle.
The sense of rotation is that of a right-handed screw which moves forward inthe direction a×b(see figure 7.9). Again, θis the angle between the two vectors
placed ‘tail to tail’ or ‘head to head’. With this definition a,banda×bform a
right-handed set. A more directly usable description of the relative directions ina vector product is provided by a right hand whose first two fingers and thumbare held to be as nearly mutually perpendicular as possible. If the first finger is
pointed in the direction of the first vector and the second finger in the direction
of the second vector, then the thumb gives the direction of the vector product.
The vector product is distributive over addition, but anticommutative andnon-
associative :
(a+b)×c=(a×c)+(b×c), (7.23)
b×a=−(a×b), (7.24)
(a×b)×c/negationslash=a×(b×c). (7.25)
226
7.6 MULTIPLICATION OF VECTORS
θ
ORP
F
r
Figure 7.10 The moment of the force Fabout Oisr×F. The cross represents
the direction of r×F, which is perpendicularly into the plane of the paper.
From its definition, we see that the vector product has the very useful property
that if a×b=0thenais parallel or antiparallel to b(unless either of them is
zero). We also note that
a×a=0. (7.26)IShow that if a=b+λc,for some scalar λ,t h e n a×c=b×c.
From (7.23) we have
a×c=(b+λc)×c=b×c+λc×c.
However, from (7.26), c×c=0and so
a×c=b×c. (7.27)
We note in passing that the fact that (7.27) is satisfied does notimply that a=b.
J
An example of the use of the vector product is that of finding the area, A,o f
a parallelogram with sides aandb, using the formula
A=|a×b|. (7.28)
Another example is afforded by considering a force Facting through a point R,
whose vector position relative to the origin Oisr(see figure 7.10). Its moment
ortorque about Ois the strength of the force times the perpendicular distance
OP, which numerically is just Frsinθ, i.e. the magnitude of r×F.Furthermore,
the sense of the moment is clockwise about an axis through Othat points
perpendicularly into the plane of the paper (the axis is represented by a crossin the figure). Thus the moment is completely represented by the vector r×F,
in both magnitude and spatial sense. It should be noted that the same vectorproduct is obtained wherever the point Ris chosen, so long as it lies on the line
of action of F.
Similarly, if a solid body is rotating about some axis that passes through the
origin, with an angular velocity ωthen we can describe this rotation by a vector
ωthat has magnitude ωand points along the axis of rotation. The direction of ω
227
VECTOR ALGEBRA
is the forward direction of a right-handed screw rotating in the same sense as the
body. The velocity of any point in the body with position vector ris then given
byv=ω×r.
Since the basis vectors i,j,kare mutually perpendicular unit vectors, forming
a right-handed set, their vector products are easily seen to be
i×i=j×j=k×k=0, (7.29)
i×j=−j×i=k, (7.30)
j×k=−k×j=i, (7.31)
k×i=−i×k=j. (7.32)
Using these relations, it is straightforward to show that the vector product of two
general vectors aandbis given in terms of their components with respect to the
basis set i,j,k,b y
a×b=(aybz−azby)i+(azbx−axbz)j+(axby−aybx)k. (7.33)
For the reader who is familiar with determinants (see chapter 8), we record that
this can also be written as
a×b=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleijk
a
xayaz
bxbybzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
That the cross product a×bis perpendicular to both aandbcan be verified
in component form by forming its dot products with each of the two vectors and
showing that it is zero in both cases.IFind the area Aof the parallelogram with sides a=i+2j+3kandb=4i+5j+6k.
The vector product a×bis given in component form by
a×b=( 2×6−3×5)i+( 3×4−1×6)j+( 1×5−2×4)k
=−3i+6j−3k.
Thus the area of the parallelogram is
A=|a×b|=
p
(−3)2+62+(−3)2=√
54.
J
7.6.3 Scalar triple product
Now that we have defined the scalar and vector products, we can extend our
discussion to define products of three vectors. Again, there are two possibilities,thescalar triple product and the vector triple product .
228
7.6 MULTIPLICATION OF VECTORS
θOφP
abcv
Figure 7.11 The triple scalar product gives the volume of a parallelepiped.
The scalar triple product is denoted by
[a,b,c]≡a·(b×c)
and, as its name suggests, it is just a number. It is most simply interpreted as the
volume of a parallelepiped whose edges are given by a,bandc(see figure 7.11).
The vector v=a×bis perpendicular to the base of the solid and has magnitude
v=absinθ, i.e. the area of the base. Further, v·c=vccosφ. Thus, since ccosφ
=OPis the vertical height of the parallelepiped, it is clear that ( a×b)·c=a r e a
of the base ×perpendicular height = volume. It follows that, if the vectors a,b
andcare coplanar, a·(b×c)=0 .
Expressed in terms of the components of each vector with respect to the
Cartesian basis set i,j,kthe scalar triple product is
a·(b×c)=ax(bycz−bzcy)+ay(bzcx−bxcz)+az(bxcy−bycx),
(7.34)
which can also be written as a determinant:
a·(b×c)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea
xayaz
bxbybz
cxcyczvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
By writing the vectors in component form, it can be shown that
a·(b×c)=(a×b)·c,
so that the dot and cross symbols can be interchanged without changing the result.
More generally, the triple scalar product is unchanged under cyclic permutation
of the vectors a,b,c. Other permutations simply give the negative of the original
triple scalar product. These results can be summarised by
[a,b,c]=[b,c,a]=[c,a,b]=−[a,c,b]=−[b,a,c]=−[c,b,a]. (7.35)
229
VECTOR ALGEBRAIFind the volume Vof the parallelepiped with sides a=i+2j+3k,b=4i+5j+6kand
c=7i+8j+1 0k.
We have already found that a×b=−3i+6j−3k, in subsection 7.6.2. Hence the volume
of the parallelepiped is given by
V=|a·(b×c)|=|(a×b)·c|
=|(−3i+6j−3k)·(7i+8j+1 0k)|
=|(−3)(7) + (6)(8) + ( −3)(10)|=3.
J
Another useful formula involving both the scalar and vector products is La-
grange’s identity (see exercise 7.9), i.e.
(a×b)·(c×d)≡(a·c)(b·d)−(a·d)(b·c). (7.36)
7.6.4 Vector triple product
By the vector triple product of three vectors a,b,cwe mean the vector a×(b×c).
Clearly, a×(b×c) is perpendicular to aand lies in the plane of bandcand so
can be expressed in terms of them (see (7.37) below). We note, from (7.25), thatthe vector triple product is not associative, i.e. a×(b×c)/negationslash=(a×b)×c.
Two useful formulae involving the vector triple product are
a×(b×c)=(a·c)b−(a·b)c, (7.37)
(a×b)×c=(a·c)b−(b·c)a, (7.38)
which may be derived by writing each vector in component form (see exercise 7.8).
It can also be shown that for any three vectors a,b,c,
a×(b×c)+b×(c×a)+c×(a×b)=0.
7.7 Equations of lines, planes and spheres
Now that we have described the basic algebra of vectors, we can apply the results
to a variety of problems, the first of which is to find the equation of a line in
vector form.
7.7.1 Equation of a line
Consider the line passing through the fixed point Awith position vector aand
having a direction b(see figure 7.12). It is clear that the position vector rof a
general point Ron the line can be written as
r=a+λb, (7.39)
230
7.7 EQUATIONS OF LINES, PLANES AND SPHERES
OAR
ab
r
Figure 7.12 The equation of a line. The vector bis in the direction ARand
λbis the vector from AtoR.
since Rcan be reached by starting from O, going along the translation vector
ato the point Aon the line and then adding some multiple λbof the vector b.
Different values of λgive different points Ron the line.
Taking the components of (7.39), we see that the equation of the line can also
b ew r i t t e ni nt h ef o r m
x−ax
bx=y−ay
by=z−az
bz=c o n s t a n t . (7.40)
Taking the vector product of (7.39) with band remembering that b×b=0,g i v e s
an alternative equation for the line
(r−a)×b=0.
We may also find the equation of the line that passes through two fixed points
AandCwith position vectors aandc.S i n c e ACis given by c−a, the position
vector of a general point on the line is
r=a+λ(c−a).
7.7.2 Equation of a plane
The equation of a plane through a point Awith position vector aand perpendic-
ular to a unit position vector ˆn(see figure 7.13) is
(r−a)·ˆn= 0; (7.41)
this follows since the vector joining Ato a general point Rwith position vector r
isr−a;rwill lie in the plane if this vector is perpendicular to the normal to the
plane. Rewriting (7.41) as r·ˆn=a·ˆn, we see that the equation of the plane may
also be expressed in the form r·ˆn=d, or in component form as
lx+my+nz=d, (7.42)
231
VECTOR ALGEBRA
Od aˆn
rAR
Figure 7.13 The equation of the plane is ( r−a)·ˆn=0 .
where the unit normal to the plane is ˆn=li+mj+nkandd=a·ˆnis the
perpendicular distance of the plane from the origin.
The equation of a plane containing points a,bandcis
r=a+λ(b−a)+µ(c−a).
This is apparent because starting from the point ain the plane, all other points
may be reached by moving a distance along each of two (non-parallel) directionsin the plane. Two such directions are given by b−aandc−a.I tc a nb es h o w n
that the equation of this plane may also be written in the more symmetrical form
r=αa+βb+γc,
where α+β+γ=1 .IFind the direction of the line of intersection of the planes x+3y−z=5 and
2x−2y+4z=3.
The two planes have normal vectors n1=i+3j−kandn2=2i−2j+4k.I ti sc l e a r
that these are not parallel vectors and so the planes must intersect along some line. Thedirection pof this line must be parallel to both planes and hence perpendicular to both
normals. Therefore
p=n
1×n2
= [(3)(4)−(−2)(−1)]i+[ (−1)(2)−(1)(4)] j+ [(1)(−2)−(3)(2)] k
=1 0i−6j−8k.
J
7.7.3 Equation of a sphere
Clearly, the defining property of a sphere is that all points on it are equidistant
from a fixed point in space and that the common distance is equal to the radius
232
7.8 USING VECTORS TO FIND DISTANCES
of the sphere. This is easily expressed in vector notation as
|r−c|2=(r−c)·(r−c)=a2, (7.43)
where cis the position vector of the centre of the sphere and ais its radius.IFind the radius ρof the circle that is the intersection of the plane ˆn·r=pand the sphere
of radius acentred on the point with position vector c.
The equation of the sphere is
|r−c|2=a2, (7.44)
and that of the circle of intersection is
|r−b|2=ρ2, (7.45)
where ris restricted to lie in the plane and bis the position of the circle’s centre.
Asblies on the plane whose normal is ˆn, the vector b−cmust be parallel to ˆn,i . e .
b−c=λˆnfor some λ. Further, by Pythagoras, we must have ρ2+|b−c|2=a2. Thus
λ2=a2−ρ2.
Writing b=c+
p
a2−ρ2ˆnand substituting in (7.45) gives
r2−2r·
/
c+
p
a2−ρ2ˆn
/
+c2+2 (c·ˆn)
p
a2−ρ2+a2−ρ2=ρ2,
whilst, on expansion, (7.44) becomes
r2−2r·c+c2=a2.
Subtracting these last two equations, using ˆn·r=pand simplifying yields
p−c·ˆn=
p
a2−ρ2.
On rearrangement, this gives ρas
p
a2−(p−c·ˆn)2, which places obvious geometrical
constraints on the values a,c,ˆnandpcan take if a real intersection between the sphere
and the plane is to occur.
J
7.8 Using vectors to find distances
This section deals with the practical application of vectors to finding distances.
Some of these problems are extremely cumbersome in component form, but theyall reduce to neat solutions when general vectors, with no explicit basis set,are used. These examples show the power of vectors in simplifying geometricalproblems.
7.8.1 Distance from a point to a line
Figure 7.14 shows a line having direction bthat passes through a point Awhose
position vector is a. To find the minimum distance dof the line from a point P
whose position vector is p, we must solve the right-angled triangle shown. We see
thatd=|p−a|sinθ; so, from the definition of the vector product, it follows that
d=|(p−a)׈b|.
233
VECTOR ALGEBRA
OAP
θp−a d
abp
Figure 7.14 The minimum distance from a point to a line.IFind the minimum distance from the point Pwith coordinates (1,2,1)to the line r=a+λb,
where a=i+j+kandb=2i−j+3k.
Comparison with (7.39) shows that the line passes through the point (1 ,1,1) and has
direction 2 i−j+3k. The unit vector in this direction is
ˆb=1√
14(2i−j+3k).
The position vector of Pisp=i+2j+kand we find
(p−a)׈b=1√
14[j×(2i−3j+3k)]
=1√
14(3i−2k).
Thus the minimum distance from the line to the point Pisd=
p
13/14.
J
7.8.2 Distance from a point to a plane
The minimum distance dfrom a point Pwhose position vector is pto the plane
defined by ( r−a)·ˆn= 0 may be deduced by finding any vector from Pto the
plane and then determining its component in the normal direction. This is shownin figure 7.15. Consider the vector a−p, which is a particular vector from Pto
the plane. Its component normal to the plane, and hence its distance from the
plane, is given by
d=(a−p)·ˆn, (7.46)
where the sign of ddepends on which side of the plane Pis situated.
234
7.8 USING VECTORS TO FIND DISTANCES
OP
d
apˆn
Figure 7.15 The minimum distance dfrom a point to a plane.IFind the distance from the point Pwith coordinates (1,2,3)to the plane that contains the
points A,BandChaving coordinates (0,1,0),(2,3,1)and(5,7,2).
Let us denote the position vectors of the points A, B, C bya,b,c. Two vectors in the
plane are
b−a=2i+2j+k and c−a=5i+6j+2k,
and hence a vector normal to the plane is
n=( 2i+2j+k)×(5i+6j+2k)=−2i+j+2k,
and its unit normal is
ˆn=n
|n|=1
3(−2i+j+2k).
Denoting the position vector of Pbyp, the minimum distance from the plane to Pis
given by
d=(a−p)·ˆn
=(−i−j−3k)·1
3(−2i+j+2k)
=2
3−1
3−2=−5
3.
If we take Pto be the origin O, then we find d=1
3, i.e. a positive quantity. It follows from
this that the original point Pwith coordinates (1 ,2,3), for which dwas negative, is on the
opposite side of the plane from the origin.
J
7.8.3 Distance from a line to a line
Consider two lines in the directions aandb, as shown in figure 7.16. Since a×b
is by definition perpendicular to both aandb, the unit vector normal to both
these lines is
ˆn=a×b
|a×b|.
235
VECTOR ALGEBRA
O Q
Pˆnq
ab
p
Figure 7.16 The minimum distance from one line to another.
Ifpandqare the position vectors of any two points PandQon different lines
then the vector connecting them is p−q. Thus, the minimum distance dbetween
the lines is this vector’s component along the unit normal, i.e.
d=|(p−q)·ˆn|.IA line is inclined at equal angles to the x-,y- and z- axes and passes through the origin.
Another line passes through the points (1,2,4)and(0,0,1). Find the minimum distance
between the two lines.
The first line is given by
r1=λ(i+j+k),
and the second by
r2=k+µ(i+2j+3k).
Hence a vector normal to both lines is
n=(i+j+k)×(i+2j+3k)=i−2j+k,
and the unit normal is
ˆn=1√
6(i−2j+k).
A vector between the two lines is, for example, the one connecting the points (0 ,0,0)
and (0 ,0,1), which is simply k. Thus it follows that the minimum distance between the
two lines is
d=1√
6|k·(i−2j+k)|=1√
6.
J
7.8.4 Distance from a line to a plane
Let us consider the line r=a+λb. This line will intersect any plane to which it
is not parallel. Thus, if a plane has a normal ˆnthen the minimum distance from
236
7.9 RECIPROCAL VECTORS
the line to the plane is zero unless
b·ˆn=0,
in which case the distance, d, will be
d=|(a−r)·ˆn|,
where ris any point in the plane.IA line is given by r=a+λb,w h e r e a=i+2j+3kandb=4i+5j+6k.F i n dt h e
coordinates of the point Pat which the line intersects the plane
x+2y+3z=6.
A vector normal to the plane is
n=i+2j+3k,
from which we find that b·n/negationslash= 0. Thus the line does indeed intersect the plane. To find
the point of intersection we merely substitute the x-,y-a n d z- values of a general point
on the line into the equation of the plane obtaining
1+4 λ+2 ( 2+5 λ)+3 ( 3+6 λ)=6⇒ 14 + 32 λ=6.
This gives λ=−1
4, which we may substitute into the equation for the line to obtain
x=1−1
4(4) = 0, y=2−1
4(5) =3
4andz=3−1
4(6) =3
2. Thus the point of intersection is
(0,3
4,3
2).
J
7.9 Reciprocal vectors
The final section of this chapter introduces the concept of reciprocal vectors,
which have particular uses in crystallography.
The two sets of vectors a,b,canda/prime,b/prime,c/primeare called reciprocal sets if
a·a/prime=b·b/prime=c·c/prime= 1 (7.47)
and
a/prime·b=a/prime·c=b/prime·a=b/prime·c=c/prime·a=c/prime·b=0. (7.48)
It can be verified (see exercise 7.19) that the reciprocal vectors of a,bandcare
given by
a/prime=b×c
a·(b×c), (7.49)
b/prime=c×a
a·(b×c), (7.50)
c/prime=a×b
a·(b×c), (7.51)
where a·(b×c)/negationslash= 0. In other words, reciprocal vectors only exist if a,bandcare
237
VECTOR ALGEBRA
not coplanar. Moreover, if a,bandcare mutually orthogonal unit vectors then
a/prime=a,b/prime=bandc/prime=c, so that the two systems of vectors are identical.IConstruct the reciprocal vectors of a=2i,b=j+k,c=i+k.
First we evaluate the triple scalar product:
a·(b×c)=2i·[(j+k)×(i+k)]
=2i·(i+j−k)=2 .
Now we find the reciprocal vectors:
a/prime=1
2(j+k)×(i+k)=1
2(i+j−k),
b/prime=1
2(i+k)×2i=j,
c/prime=1
2(2i)×(j+k)=−j+k.
It is easily verified that these reciprocal vectors satisfy their defining properties (7.47),
(7.48).
J
We may also use the concept of reciprocal vectors to define the components of a
vector awith respect to basis vectors e1,e2,e3that are not mutually orthogonal.
If the basis vectors are of unit length and mutually orthogonal, such as the
Cartesian basis vectors i,j,k, then (see the text preceeding (7.21)) acan be
w r i t t e ni nt h ef o r m
a=(a·i)i+(a·j)j+(a·k)k.
If the basis is not orthonormal, however, then this is no longer true. Nevertheless,
we may write the components of awith respect to a non-orthonormal basis
e1,e2,e3in terms of its reciprocal basis vectors e/prime
1,e/prime
2,e/prime
3, which are defined as in
(7.49)–(7.51). If we let
a=a1e1+a2e2+a3e3,
then the scalar product a·e/prime
1is given by
a·e/prime
1=a1e1·e/prime
1+a2e2·e/prime
1+a3e3·e/prime
1=a1,
where we have used the relations (7.48). Similarly, a2=a·e/prime
2anda3=a·e/prime
3;s o
now
a=(a·e/prime
1)e1+(a·e/prime
2)e2+(a·e/prime
3)e3. (7.52)
7.10 Exercises
7.1 Which of the following statements about general vectors a,bandcare true?
(a)c·(a×b)=(b×a)·c.
(b)a×(b×c)=(a×b)×c.
(c)a×(b×c)=(a·c)b−(a·b)c.
(d)d=λa+µbimplies ( a×b)·d=0.
(e)a×c=b×cimplies c·a−c·b=c|a−b|.
(f) (a×b)×(c×b)=b[b·(c×a)].
238
7.10 EXERCISES
7.2 A unit cell of diamond is a cube of side Awith carbon atoms at each corner, at
the centre of each face and, in addition, displaced by1
4A(i+j+k)f r o me a c ho f
the previously mentioned ones, where i,j,kare unit vectors along the cube axes.
One corner of the cube is taken as the origin of coordinates. What are the vectorsjoining the atom at
1
4A(i+j+k) to its four nearest neighbours? Determine the
angle between the carbon bonds in diamond.
7.3 Identify the following surfaces:
(a)|r|=k;( b )r·u=l;( c )r·u=m|r|for−1≤m≤+1;
(d)|r−(r·u)u|=n.
Here k,l,mandnare fixed scalars and uis a fixed unit vector.
7.4 Find the angle between the position vectors to the points (3 ,−4,0) and (−2,1,0)
and find the direction cosines of a vector perpendicular to both.
7.5 A, B, C andDare the four corners, in order, of one face of a cube of side 2
units. The opposite face has corners E,F,G andH,w i t h AE, BF, CG andDHas
parallel edges of the cube. The centre Oo ft h ec u b ei st a k e na st h eo r i g i na n dt h e
x-,y-a n d z-axes are parallel to AD,AEandABrespectively. Find the following:
(a) the angle between the face diagonal AFand the body diagonal AG;
(b) the equation of the plane through Bthat is parallel to the plane CGE;
(c) the perpendicular distance from the centre Jof the face BCGF to the plane
OCG;
(d) the volume of the tetrahedron JOCG .
7.6 Use vector methods to prove that the lines joining the mid-points of the opposite
edges of a tetrahedron OABC meet at a point and that this point bisects each of
the lines.
7.7 The edges OP,OQand ORof a tetrahedron OPQR are vectors p,qandr
respectively, where p=2i+4j,q=2i−j+3kandr=4i−2j+5k. Show that
OPis perpendicular to the plane containing OQR. Express the volume of the
tetrahedron in terms of p,qandrand hence calculate the volume.
7.8 Prove, by writing it out in component form, that
(a×b)×c=(a·c)b−(b·c)a,
and deduce the result, stated in (7.25), that the operation of forming the vector
product is non-associative.
7.9 Prove Lagrange’s identity, i.e.
(a×b)·(c×d)=(a·c)(b·d)−(a·d)(b·c).
7.10 For four arbitrary vectors a,b,candd, evaluate
(a×b)×(c×d)
in two different ways and so prove that
a[b,c,d]−b[c,d,a]+c[d,a,b]−d[a,b,c]=0 .
Show that this reduces to the normal Cartesian representation of the vector d,
i.e.dxi+dyj+dzkifa,bandcare taken as i,jandk, the Cartesian base vectors.
7.11 Show that the points (1 ,0,1), (1 ,1,0) and (1 ,−3,4) lie on a straight line. Give the
equation of the line in the form
r=a+λb.
7.12 The plane P1contains the points A,Band C, which have position vectors
a=−3i+2j,b=7i+2jandc=2i+3j+2krespectively. Plane P2passes through
Aand is orthogonal to the line BC, whilst plane P3passes through Band is
orthogonal to the line AC. Find the coordinates of r, the point of intersection of
the three planes.
239
VECTOR ALGEBRA
7.13 Two planes have non-parallel unit normals ˆnandˆmand their closest distances
from the origin are λandµrespectively. Find the vector equation of their line of
intersection in the form r=νp+a.
7.14 Two fixed points, AandB, in three-dimensional space have position vectors a
andb. Identify the plane Pgiven by
(a−b)·r=1
2(a2−b2),
where aandbare the magnitudes of aandb.
Show also that the equation
(a−r)·(b−r)=0
describes a sphere Sof radius|a−b|/2. Deduce that the intersection of Pand
Sis also the intersection of two spheres, centred on AandBand each of radius
|a−b|/√2.
7.15 Let O,A,BandCbe four points with position vectors 0,a,bandc, and denote
byg=λa+µb+νcthe position of the centre of the sphere on which they all lie.
(a) Prove that λ,µandνsimultaneously satisfy
(a·a)λ+(a·b)µ+(a·c)ν=1
2a2
and two other similar equations.
(b) By making a change of origin, find the centre and radius of the sphere on
which the points p=3i+j−2k,q=4i+3j−3k,r=7i−3kands=6i+j−k
all lie.
7.16 The vectors a,bandcare coplanar and related by
λa+µb+νc=0,
where λ,µ,νare not all zero. Show that the condition for the points with position
vectors αa,βbandγcto be collinear is
λ
α+µ
β+ν
γ=0.
7.17 (a) Show that the line of intersection of the planes x+2y+3z=0a n d
3x+2y+z= 0 is equally inclined to the x-a n d z-a x e sa n dm a k e sa na n g l e
cos−1(−2/√
6) with the y-axis.
(b) Find the perpendicular distance between one corner of a unit cube and the
major diagonal not passing through it.
7.18 Four points Xi(i=1,2,3,4), taken for simplicity as all lying within the octant
x, y, z≥0, have position vectors xi. Convince yourself that vector xnlies within
the sector of space defined by the other three vectors if
max
over i
/
min
over j/negationslash=i
/xi·xj
|xi||xj|
//
=n,
i.e. if nequals that value of ifor which the largest of the set of angles which xi
makes with the other vectors is the lowest. Determine whether any of the four
points with coordinates
X1=( 3,2,2),X 2=( 2,3,1),X 3=( 2,1,3),X 4=( 3,0,3)
lies within the tetrahedron defined by the origin and the other three points.
7.19 The vectors a,bandcare not coplanar. The vectors a/prime,b/primeandc/primeare the
240
7.10 EXERCISES
d
aa
b
c
Figure 7.17 A face-centred cubic crystal.
associated reciprocal vectors. Verify that the expressions (7.49)–(7.51) define a set
of reciprocal vectors a/prime,b/primeandc/primewith the following properties:
(a)a/prime·a=b/prime·b=c/prime·c=1 ;
(b)a/prime·b=a/prime·c=b/prime·aetc = 0;
(c) [a/prime,b/prime,c/prime]=1 /[a,b,c];
(d)a=(b/prime×c/prime)/[a/prime,b/prime,c/prime].
7.20 Three non-coplanar vectors a,bandc, have as their respective reciprocal vectors
the set a/prime,b/primeandc/prime. Show that the normal to the plane containing the points
k−1a,l−1bandm−1cis in the direction of the vector ka/prime+lb/prime+mc/prime.
7.21 In a crystal with a face-centred cubic structure, the basic cell can be taken as a
cube of edge awith its centre at the origin of coordinates and its edges parallel
to the Cartesian coordinate axes; atoms are sited at the eight corners and at thecentre of each face. However, other basic cells are possible. One is the rhomboidshown in figure 7.17, which has the three vectors b,canddas edges.
(a) Show that the volume of the rhomboid is one-quarter that of the cube.
(b) Show that the angles between pairs of edges of the rhomboid are 60
◦and that
the corresponding angles between pairs of edges of the rhomboid defined bythe reciprocal vectors to b,c,dare each 109 .5
◦. (This rhomboid can be used
as the basic cell of a body-centred cubic structure, more easily visualised asa cube with an atom at each corner and one at its centre.)
(c) In order to use the Bragg formula, 2 dsinθ=nλ, for the scattering of X-rays
by a crystal, it is necessary to know the perpendicular distance dbetween
successive planes of atoms; for a given crystal structure, dhas a particular
value for each set of planes considered. For the face-centred cubic structurefind the distance between successive planes with normals in the k,i+jand
i+j+kdirections.
7.22 In subsection 7.6.2 we showed how the moment or torque of a force about an axis
could be represented by a vector in the direction of the axis. The magnitude ofthe vector gives the size of the moment and the sign of the vector gives the sense.Similar representations can be used for angular velocities and angular momenta.
(a) The magnitude of the angular momentum about the origin of a particle of
mass mmoving with velocity von a path that is a perpendicular distance d
241
VECTOR ALGEBRA
from the origin is given by m|v|d. Show that if ris the position of the particle
then the vector J=r×mvrepresents the angular momentum.
(b) Now consider a rigid collection of particles (or a solid body) rotating about
an axis through the origin, the angular velocity of the collection beingrepresented by ω.
(i) Show that the velocity of the ith particle is
v
i=ω×ri
and that the total angular momentum Jis
J=
X
imi[r2
iω−(ri·ω)ri].
(ii) Show further that the component of Jalong the axis of rotation can
be written as Iω,w h e r e I, the moment of inertia of the collection
about the axis or rotation, is given by
I=
X
imiρ2
i.
Interpret ρigeometrically.
(iii) Prove that the total kinetic energy of the particles is1
2Iω2.
7.23 By proceeding as indicated below, prove the parallel axis theorem ,w h i c hs t a t e s
that, for a body of mass M, the moment of inertia Iabout any axis is related to
the corresponding moment of inertia I0about a parallel axis that passes through
the centre of mass of the body by
I=I0+Ma2
⊥,
where a⊥is the perpendicular distance between the two axes. Note that I0can
be written asZ
(ˆn×r)·(ˆn×r)dm,
where ris the vector position, relative to the centre of mass, of the infinitesimal
mass dmandˆnis a unit vector in the direction of the axis of rotation. Write a
similar expression for Iin which ris replaced by r/prime=r−a,w h e r e ais the vector
position of any point on the axis to which Irefers. Use Lagrange’s identity and
the fact that
R
rdm=0(by the definition of the centre of mass) to establish the
result.
7.24 Without carrying out any further integration, use the results of the previous
exercise, the worked example in subsection 6.3.4 and exercise 6.10 to prove thatthe moment of inertia of a uniform rectangular lamina, of mass Mand sides a
andb, about an axis perpendicular to its plane and passing through the point
(αa/2,βb /2), with−1≤α, β≤1,is
M
12[a2(1 + 3 α2)+b2(1 + 3 β2)].
7.25 Define a set of (non-orthogonal) base vectors a=j+k,b=i+kandc=i+j.
(a) Establish their reciprocal vectors and hence express the vectors p=3i−2j+k,
q=i+4jandr=−2i+j+kin terms of the base vectors a,bandc.
(b) Verify that the scalar product p·qhas the same value, −5, when evaluated
using either set of components.
242
7.10 EXERCISES
f= 200 Hz
ω= 400 πs−1V0cosωtV1V2
V3V4
R1=5 0ΩR2
I1I2
I3L
C=1 0 µF
Figure 7.18 An oscillatory electric c ircuit. The power supply has angular
frequency ω=2πf= 400 πs−1.
7.26 Systems that can be modelled as damped harmonic oscillators are widespread;
pendulum clocks, car shock absorbers, tuning circuits in television sets and radios,and collective electron motions in plasmas and metals are just a few examples.
In all these cases, one or more variables describing the system obey(s) an
equation of the form
¨x+2γ˙x+ω
2
0x=Pcosωt,
where ˙x=dx/dt, etc. and the inclusion of the factor 2 is conventional. In the
steady state (i.e. after the effects of any in itial displacement or velocity have been
damped out) the solution of the equation takes the form
x(t)=Acos(ωt+φ).
By expressing each term in the form Bcos(ωt+/epsilon1) and representing it by a vector
of magnitude Bmaking an angle /epsilon1with the x-axis, draw a closed vector diagram,
att= 0, say, that is equivalent to the equation.
(a) Convince yourself that whatever the value of ω(>0)φmust be negative
(−π<φ≤0) and that
φ=t a n−1
/−2γω
ω2
0−ω2
/
.
(b) Obtain an expression for Ain terms of P,ω0andω.
7.27 According to alternating current theory, the currents and voltages in the compo-
nents of the circuit shown in figure 7.18 are determined by Kirchhoff’s laws andthe relationships
I
1=V1
R1,I 2=V2
R2,I 3=iω CV 3,V 4=iω LI 2.
The factor i=√
−1 in the expression for I3indicates that the phase of I3is 90◦
ahead of V3. Similarly the phase of V4is 90◦ahead of I2.
Measurement shows that V3has an amplitude of 0 .661V0and a phase of
+13.4◦relative to that of the power supply. Taking V0= 1 V and using a series
of vector plots for voltages and currents (they could all be on the same plot ifsuitable scales were chosen), determine all unknown currents and voltages andfind values for the inductance of Land the resistance of R
2. (Scales of 1 cm =
0.1 V for voltages and 1 cm = 1 mA for currents are convenient.)
243
VECTOR ALGEBRA
7.11 Hints and answers
7.1 (c), (d) and (e).
7.2 In units of1
4Athe vectors are −i−j−k,i+j−k,i−j+k,−i+j+k;
cos−1(−1
3) = 109 .5◦.
7.3 (a) A sphere of radius kcentred on the origin; (b) a plane with its normal in the
direction of uand a distance lfrom the origin; (c) a cone with its axis parallel to
uand semiangle cos−1m; (d) a circular cylinder of radius nwith its axis parallel
tou.
7.4 cos−1(−2/√
5) = 153 .4◦;0 ,0 ,1 .
7.5 (a) cos−1
p
2/3; (b) z−x=2 ;( c )1 /√2; (d)1
31
2(c×g)·j=1
3.
7.6 With an obvious notation, the mid-points of OAandBCarea/2a n d( b+c)/2;
the mid-point of the line joining them is ( a+b+c)/2. The same result is obtained
forOBandAC,a n df o r OCandAB.
7.7 Show that q×ris parallel to p; volume =1
3
/1
2(q×r)·p
/
=5
3.
7.9 Note that ( a×b)·(c×d)=d·[(a×b)×c] and use the result from the previous
question.
7.10 Consider ( a×b)×[(c×d)] as λa+µband [( a×b)]×(c×d)a sλ/primec+µ/primedusing
t h er e s u l to fe x e r c i s e7 . 8 .
7.11 Show that the position vectors of the points are linearly dependent; r=a+λb
where a=i+kandb=−j+k.
7.12 The conditions are ( r−a)·[(b−a)×(c−a)] = 0, ( r−a)·(b−c)=0a n d
(r−b)·(c−a) = 0; the point of intersection is r=2i+7j+1 0k.
7.13 Show that pmust have the direction ˆn׈mand write aasxˆn+yˆm. By obtaining a
pair of simultaneous equations for xandy, prove that x=(λ−µˆn·ˆm)/[1−(ˆn·ˆm)2]
and that y=(µ−λˆn·ˆm)/[1−(ˆn·ˆm)2].
7.14 Pis the plane orthogonal to the line joining AandBand equidistant from them.
Sis|r−c|2=(|a−b|/2)2,w h e r e c=(a+b)/2. Add and subtract the equations
forPandSand arrange the resulting equations in the form |r−d|2=R2.
7.15 (a) Note that |a−g|2=R2=|0−g|2, leading to a·a=2a·g.
(b) Make pthe new origin and solve the three simultaneous linear equations to
obtain λ=5/18,µ=1 0 /18,ν=−3/18, giving g=2i−kand a sphere of
radius√5 centred on (5 ,1,−3).
7.16 For collinearity, γc=θαa+( 1−θ)βbfor some θ.
7.17 (a) Find two points on both planes, say (0 ,0,0) and (1 ,−2,1), and hence determine
the direction cosines of the line of intersection; (b) (2
3)1/2.
7.18 The scalar products sij(i=1,2,3;j>i) between pairs of unit vectors are 0.907,
0.907, 0.857; 0.714, 0.567; 0.945. Thus i= 1 has the highest minimum ( s14=0.857)
and so only X1could meet the condition. The plane containing X2,X3andX4
isx+y+z−6=0 ; s i n c e3+2+2 −6>0,X1lies outside the tetrahedron
OX2X3X4. None of the points meets the condition.
7.19 For (c) and (d), use the result of exercise 7.8 to evaluate ( c×a)×(a×b).
7.20 The normal is in the direction ( l−1b−k−1a)×(m−1c−k−1a).
7.21 (b) b/prime=a−1(−i+j+k),c/prime=a−1(i−j+k),d/prime=a−1(i+j−k); (c) a/2 for direction
k; successive planes through (0 ,0,0) and ( a/2,0,0) give a spacing of a/√
8f o r
direction i+j; successive planes through ( −a/2,0,0) and ( a/2,0,0) give a spacing
ofa/√
3 for direction i+j+k.
7.22 (a) Check both magnitude and rotational sense. (b)(i) Use the result of exercise
7.8 to evaluate ri×mi(ω×ri). (ii) Form ( J·ω)/ω;ρiis the distance of the
ith particle from the axis of rotation. (iii) use Lagrange’s identity to evaluate
(ω×ri)·(ω×ri).
7.23 Note that a2−(ˆn·a)2=a2
⊥.
7.24 The moment of inertia about an axis through the centre of the rectangle and
perpendicular to its plane is1
12M(a2+b2).
244
7.11 HINTS AND ANSWERS
7.26φ2φ12γωA
2γωAω2A
ω2Aω2
0A
ω2
0AP
Figure 7.19 The vector diagram for the equation in exercise 7.26.
7.25 p=−2a+3b,q=3
2a−3
2b+5
2candr=2a−b−c. Remember that a·a=b·b=
c·c=2a n d a·b=a·c=b·c=1 .
See figure 7.19 and recall that −cosθ=c o s ( θ+π)a n d−sinθ=c o s ( θ+π/2).
(a) With φ1>0, no matter what value ωtakes, the possible resultants (broken
arrows) can never equal P.W i t h φ2<0, closure of the quadrilateral is possible.
(b)A=P[(ω2
0−ω2)2+4γ2ω2]−1/2.
7.27 With currents in units of mA/ |V0|. Voltages in units of V0:
I1=( 7.76,−23.2◦),I2=( 1 4 .36,−50.8◦),I3=( 8.30,103.4◦);
V1=( 0.388,−23.2◦),V2=( 0.287,−50.8◦),V4=( 0.596,39.2◦);
L= 33 mH, R2=2 0Ω .
245
8
Matrices and vector spaces
In the previous chapter we defined a vector as a geometrical object which has
both a magnitude and a direction and which may be thought of as an arrow fixedin our familiar three-dimensional space, a space which, if we need to, we defineby reference to, say, the fixed stars. This geometrical definition of a vector is bothuseful and important since it is independent of any coordinate system with which
we choose to label points in space.
In most specific applications, however, it is necessary at some stage to choose
a coordinate system and to break down a vector into its component vectors in
the directions of increasing coordinate values. Thus for a particular Cartesiancoordinate system (for example) the component vectors of a vector awill be a
xi,
ayjandazkand the complete vector will be
a=axi+ayj+azk. (8.1)
Although we have so far considered only real three-dimensional space, we may
extend our notion of a vector to more abstract spaces, which in general canhave an arbitrary number of dimensions N. We may still think of such a vector
as an ‘arrow’ in this abstract space, so that it is again independent of any ( N-
dimensional) coordinate system with which we choose to label the space. As an
example of such a space, which, though abstract, has very practical applications,
we may consider the description of a mechanical or electrical system. If the stateof a system is uniquely specified by assigning values to a set of Nvariables,
which may be angles or currents, for example, then that state can be representedby a vector in an N-dimensional space, the vector having those values as its
components.
In this chapter we first discuss general vector spaces and their properties. We
then go on to discuss the transformation of one vector into another by a linear
operator. This leads naturally to the concept of a matrix , a two-dimensional array
of numbers. The properties of matrices are then discussed and we conclude with
246
8.1 VECTOR SPACES
a discussion of how to use these properties to solve systems of linear equations.
The application of matrices to the study of oscillations in physical systems ist a k e nu pi nc h a p t e r9 .
8.1 Vector spaces
A set of objects (vectors) a,b,c,...is said to form a linear vector space Vif:
(i) the set is closed under commutative and associative addition, so that
a+b=b+a, (8.2)
(a+b)+c=a+(b+c); (8.3)
(ii) the set is closed under multiplication by a scalar (any complex number) to
form a new vector λa, the operation being both distributive and associative
so that
λ(a+b)=λa+λb, (8.4)
(λ+µ)a=λa+µa, (8.5)
λ(µa)=(λµ)a, (8.6)
where λandµare arbitrary scalars;
(iii) there exists a null vector 0such that a+0=afor all a;
(iv) multiplication by unity leaves any vector unchanged, i.e. 1 ×a=a;
(v) all vectors have a corresponding negative vector −asuch that a+(−a)=0.
It follows from (8.5) with λ=1a n d µ=−1t h a t−ais the same vector as
(−1)×a.
We note that if we restrict all scalars to be real then we obtain a real vector
space (an example of which is our familiar three-dimensional space); otherwise,
in general, we obtain a complex vector space . We note that it is common to use the
terms ‘vector space’ and ‘space’, instead of the more formal ‘linear vector space’.
Thespanof a set of vectors a,b,...,sis defined as the set of all vectors that
may be written as a linear sum of the original set, i.e. all vectors
x=αa+βb+···+σs (8.7)
that result from the infinite number of possible values of the (in general complex)
scalars α ,β,...,σ .I fxin (8.7) is equal to 0for some choice of α ,β,...,σ (notall
zero), i.e. if
αa+βb+···+σs=0, (8.8)
then the set of vectors a,b,...,s,i ss a i dt ob e linearly dependent .I ns u c has e t
at least one vector is redundant, since it can be expressed as a linear sum of
247
MATRICES AND VECTOR SPACES
the others. If, however, (8.8) is not satisfied by anyset of coefficients (other than
the trivial case in which all the coefficients are zero) then the vectors are linearly
independent , and no vector in the set can be expressed as a linear sum of the
others.
If, in a given vector space, there exist sets of Nlinearly independent vectors,
but no set of N+ 1 linearly independent vectors, then the vector space is said to
beN-dimensional. (In this chapter we will limit our discussion to vector spaces of
finite dimensionality; spaces of infinite dimensionality are discussed in chapter 17.)
8.1.1 Basis vectors
IfVis an N-dimensional vector space then anyset of Nlinearly independent
vectors e1,e2,...,eNforms a basisforV.I fxis an arbitrary vector lying in Vthen
the set of N+ 1 vectors x,e1,e2,...,eN, must be linearly dependent and therefore
such that
αe1+βe2+···+σeN+χx=0, (8.9)
where the coefficients α ,β,...,χ are not all equal to 0, and in particular χ/negationslash=0 .
Rearranging (8.9) we may write xas a linear sum of the vectors eias follows:
x=x1e1+x2e2+···+xNeN=Nsummationdisplay
i=1xiei, (8.10)
for some set of coefficients xithat are simply related to the original coefficients,
e.g.x1=−α/χ,x2=−β/χ,e t c .S i n c ea n y xlying in the span of Vcan be
expressed in terms of the basis orbase vectors ei, the latter are said to form
acomplete set. The coefficients xiare the components ofxwith respect to the
ei-basis. These components are unique , since if both
x=Nsummationdisplay
i=1xieiand x=Nsummationdisplay
i=1yiei,
then
Nsummationdisplay
i=1(xi−yi)ei=0, (8.11)
which, since the eiare linearly independent, has only the solution xi=yifor all
i=1,2,...,N .
From the above discussion we see that anyset of Nlinearly independent
vectors can form a basis for an N-dimensional space. If we choose a different set
e/prime
i,i=1,...,N then we can write xas
x=x/prime
1e/prime1+x/prime
2e/prime2+···+x/prime
Ne/primeN=Nsummationdisplay
i=1x/prime
ie/primei. (8.12)
248
8.1 VECTOR SPACES
We reiterate that the vector x(a geometrical entity) is independent of the basis
– it is only the components of xthat depend on the basis. We note, however,
that given a set of vectors u1,u2,...,uM,w h e r e M/negationslash=N,i na n N-dimensional
vector space, then either there exists a vector that cannot be expressed as a
linear combination of the uior, for some vector that can be so expressed, the
components are not unique.
8.1.2 The inner product
We may usefully add to the description of vectors in a vector space by defining
theinner product of two vectors, denoted in general by /angbracketlefta|b/angbracketright, which is a scalar
function of aandb. The scalar or dot product, a·b≡|a||b|cosθ,o fv e c t o r s
in real three-dimensional space (where θis the angle between the vectors), was
introduced in the last chapter and is an example of an inner product. In effect thenotion of an inner product /angbracketlefta|b/angbracketrightis a generalisation of the dot product to more
abstract vector spaces. Alternative notations for /angbracketlefta|b/angbracketrightare (a,b), or simply a·b.
The inner product has the following properties:
(i)/angbracketlefta|b/angbracketright=/angbracketleftb|a/angbracketright
∗,
(ii)/angbracketlefta|λb+µc/angbracketright=λ/angbracketlefta|b/angbracketright+µ/angbracketlefta|c/angbracketright.
We note that in general, for a complex vector space, (i) and (ii) imply that
/angbracketleftλa+µb|c/angbracketright=λ∗/angbracketlefta|c/angbracketright+µ∗/angbracketleftb|c/angbracketright, (8.13)
/angbracketleftλa|µb/angbracketright=λ∗µ/angbracketlefta|b/angbracketright. (8.14)
Following the analogy with the dot product in three-dimensional real space,
two vectors in a general vector space are defined to be orthogonal if/angbracketlefta|b/angbracketright=0 .
Similarly, the norm of a vector ais given by /bardbla/bardbl=/angbracketlefta|a/angbracketright1/2and is clearly a
generalisation of the length or modulus |a|of a vector ain three-dimensional
space. In a general vector space /angbracketlefta|a/angbracketrightcan be positive or negative; however, we
shall be primarily concerned with spaces in which /angbracketlefta|a/angbracketright≥0 and which are thus
said to have a positive semi-definite norm .I ns u c has p a c e /angbracketlefta|a/angbracketright= 0 implies a=0.
Let us now introduce into our N-dimensional vector space a basis ˆe1,ˆe2,...,ˆeN
that has the desirable property of being orthonormal (the basis vectors are mutually
orthogonal and each has unit norm), i.e. a basis that has the property
/angbracketleftˆei|ˆej/angbracketright=δij. (8.15)
Here δijis the Kronecker delta symbol (of which we say more in chapter 21) and
has the properties
δij=braceleftBigg
1f o r i=j,
0f o r i/negationslash=j.
249
MATRICES AND VECTOR SPACES
In the above basis we may express any two vectors aandbas
a=Nsummationdisplay
i=1aiˆeiand b=Nsummationdisplay
i=1biˆei.
Furthermore, in such an orthonormal basis we have, for any a,
/angbracketleftˆej|a/angbracketright=Nsummationdisplay
i=1/angbracketleftˆej|aiˆei/angbracketright=Nsummationdisplay
i=1ai/angbracketleftˆej|ˆei/angbracketright=aj. (8.16)
Thus the components of aare given by ai=/angbracketleftˆei|a/angbracketright. Note that this is nottrue
unless the basis is orthonormal. We can write the inner product of aandbin
terms of their components in an orthonormal basis as
/angbracketlefta|b/angbracketright=/angbracketlefta1ˆe1+a2ˆe2+···+aNˆeN|b1ˆe1+b2ˆe2+···+bNˆeN/angbracketright
=Nsummationdisplay
i=1a∗
ibi/angbracketleftˆei|ˆei/angbracketright+Nsummationdisplay
i=1Nsummationdisplay
j/negationslash=ia∗
ibj/angbracketleftˆei|ˆej/angbracketright
=Nsummationdisplay
i=1a∗
ibi,
where the second equality follows from (8.14) and the third from (8.15). This is
clearly a generalisation of the expression (7.21) for the dot product of vectors inthree-dimensional space.
We may generalise the above to the case where the base vectors e
1,e2,...,eN
arenotorthonormal (or orthogonal). In general we can define the N2numbers
Gij=/angbracketleftei|ej/angbracketright. (8.17)
Then, if a=summationtextN
i=1aieiandb=summationtextN
i=1biei, the inner product of aandbis given by
/angbracketlefta|b/angbracketright=angbracketleftBiggNsummationdisplay
i=1aieivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleNsummationdisplay
j=1bjejangbracketrightBigg
=Nsummationdisplay
i=1Nsummationdisplay
j=1a∗
ibj/angbracketleftei|ej/angbracketright
=Nsummationdisplay
i=1Nsummationdisplay
j=1a∗
iGijbj. (8.18)
We further note that from (8.17) and the properties of the inner product we
require Gij=G∗
ji. This in turn ensures that /bardbla/bardbl=/angbracketlefta|a/angbracketrightis real, since then
/angbracketlefta|a/angbracketright∗=Nsummationdisplay
i=1Nsummationdisplay
j=1aiG∗
ija∗j=Nsummationdisplay
j=1Nsummationdisplay
i=1a∗
jGjiai=/angbracketlefta|a/angbracketright.
250
8.1 VECTOR SPACES
8.1.3 Some useful inequalities
For a set of objects (vectors) forming a linear vector space in which /angbracketlefta|a/angbracketright≥0f o r
alla, the following inequalities are often useful.
(i)Schwarz’s inequality is the most basic result and states that
|/angbracketlefta|b/angbracketright|≤/bardbla/bardbl/bardblb/bardbl, (8.19)
where the equality holds when ais a scalar multiple of b,i . e .w h e n a=λb.
It is important here to distinguish between the absolute value of a scalar,
|λ|,a n dt h e normof a vector, /bardbla/bardbl. Schwarz’s inequality may be proved by
considering
/bardbla+λb/bardbl2=/angbracketlefta+λb|a+λb/angbracketright
=/angbracketlefta|a/angbracketright+λ/angbracketlefta|b/angbracketright+λ∗/angbracketleftb|a/angbracketright+λλ∗/angbracketleftb|b/angbracketright.
If we write /angbracketlefta|b/angbracketrightas|/angbracketlefta|b/angbracketright|eiαthen
/bardbla+λb/bardbl2=/bardbla/bardbl2+|λ|2/bardblb/bardbl2+λ|/angbracketlefta|b/angbracketright|eiα+λ∗|/angbracketlefta|b/angbracketright|e−iα.
However,/bardbla+λb/bardbl2≥0f o ra l l λ,s ow em a yc h o o s e λ=re−iαand require
that, for all r,
0≤/bardbla+λb/bardbl2=/bardbla/bardbl2+r2/bardblb/bardbl2+2r|/angbracketlefta|b/angbracketright|.
This means that the quadratic equation in rformed by setting the RHS
equal to zero must have no real roots. This, in turn, implies that
4|/angbracketlefta|b/angbracketright|2≤4/bardbla/bardbl2/bardblb/bardbl2,
which, on taking the square root (all factors are necessarily positive) of
both sides, gives Schwarz’s inequality.
(ii) The triangle inequality states that
/bardbla+b/bardbl≤/bardbla/bardbl+/bardblb/bardbl (8.20)
and may be derived from the properties of the inner product and Schwarz’s
inequality as follows. Let us first consider
/bardbla+b/bardbl2=/bardbla/bardbl2+/bardblb/bardbl2+2R e/angbracketlefta|b/angbracketright≤/bardbla/bardbl2+/bardblb/bardbl2+2|/angbracketlefta|b/angbracketright|.
Using Schwarz’s inequality we then have
/bardbla+b/bardbl2≤/bardbla/bardbl2+/bardblb/bardbl2+2/bardbla/bardbl/bardblb/bardbl=(/bardbla/bardbl+/bardblb/bardbl)2,
which, on taking the square root, gives the triangle inequality (8.20).
(iii)Bessel’s inequality requires the introduction of an orthonormal basis ˆei,
i=1,2,...,N into the N-dimensional vector space; it states that
/bardbla/bardbl2≥summationdisplay
i|/angbracketleftˆei|a/angbracketright|2, (8.21)
251
MATRICES AND VECTOR SPACES
where the equality holds if the sum includes all Nbasis vectors. If not
all the basis vectors are included in the sum then the inequality results(though of course the equality remains if those basis vectors omitted allhave a
i= 0). Bessel’s inequality can also be written
/angbracketlefta|a/angbracketright≥summationdisplay
i|ai|2,
where the aiare the components of ain the orthonormal basis. From (8.16)
these are given by ai=/angbracketleftˆei|a/angbracketright. The above may be proved by considering
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea−summationdisplay
i/angbracketleftˆei|a/angbracketrightˆeivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle2
=angbracketleftBig
a−summationdisplay
i/angbracketleftˆei|a/angbracketrightˆeivextendsinglevextendsinglevextendsinglea−summationdisplay
j/angbracketleftˆej|a/angbracketrightˆejangbracketrightBig
.
Expanding out the inner product and using /angbracketleftˆei|a/angbracketright∗=/angbracketlefta|ˆei/angbracketright,w eo b t a i n
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea−summationdisplay
i/angbracketleftˆei|a/angbracketrightˆeivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle2
=/angbracketlefta|a/angbracketright−2summationdisplay
i/angbracketlefta|ˆei/angbracketright/angbracketleftˆei|a/angbracketright+summationdisplay
isummationdisplay
j/angbracketlefta|ˆei/angbracketright/angbracketleftˆej|a/angbracketright/angbracketleftˆei|ˆej/angbracketright.
Now/angbracketleftˆei|ˆej/angbracketright=δij, since the basis is orthonormal, and so we find
0≤vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea−summationdisplay
i/angbracketleftˆei|a/angbracketrightˆeivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle2
=/bardbla/bardbl2−summationdisplay
i|/angbracketleftˆei|a/angbracketright|2,
which is Bessel’s inequality.
We take this opportunity to mention also
(iv) the parallelogram equality
/bardbla+b/bardbl2+/bardbla−b/bardbl2=2parenleftbig
/bardbla/bardbl2+/bardblb/bardbl2parenrightbig
, (8.22)
which may be proved straightforwardly from the properties of the inner
product.
8.2 Linear operators
We now discuss the action of linear operators on vectors in a vector space. A
linear operator Aassociates with every vector xanother vector
y=Ax,
in such a way that, for two vectors aandb,
A(λa+µb)=λAa+µAb,
where λ,µare scalars. We say that A‘operates’ on xto give the vector y.W e
note that the action of Aisindependent of any basis or coordinate system and
252
8.2 LINEAR OPERATORS
may be thought of as ‘transforming’ one geometrical entity (i.e. a vector) into
another.
If we now introduce a basis ei,i=1,2,...,N , into our vector space then the
action of Aon each of the basis vectors is to produce a linear combination of
the latter; this may be written as
Aej=Nsummationdisplay
i=1Aijei, (8.23)
where Aijis the ith component of the vector Aejin this basis; collectively the
numbers Aijare called the components of the linear operator in the ei-basis. In
this basis we can express the relation y=Axin component form as
y=Nsummationdisplay
i=1yiei=A
Nsummationdisplay
j=1xjej
=Nsummationdisplay
j=1xjNsummationdisplay
i=1Aijei,
and hence, in purely component form, in this basis we have
yi=Nsummationdisplay
j=1Aijxj. (8.24)
If we had chosen a different basis e/prime
i, in which the components of x,yandA
arex/prime
i,y/prime
iandA/prime
ijrespectively then the geometrical relationship y=Axwould be
represented in this new basis by
y/prime
i=Nsummationdisplay
j=1A/prime
ijx/primej.
We have so far assumed that the vector yis in the same vector space as
x.I f ,h o w e v e r , ybelongs to a different vector space, which may in general be
M-dimensional ( M/negationslash=N) then the above analysis needs a slight modification. By
introducing a basis set fi,i=1,2,...,M , into the vector space to which ybelongs
we may generalise (8.23) as
Aej=Msummationdisplay
i=1Aijfi,
where the components Aijof the linear operator Arelate to both of the bases ej
andfi.
253
MATRICES AND VECTOR SPACES
8.2.1 Properties of linear operators
Ifxis a vector and AandBare two linear operators then it follows that
(A+B)x=Ax+Bx,
(λA)x=λ(Ax),
(AB)x=A(Bx),
where in the last equality we see that the action of two linear operators in
succession is associative. The product of two linear operators is not in generalcommutative, however, so that in general ABx/negationslash=BAx. In an obvious way we
define the null (or zero) and identity operators by
Ox=0and Ix=x,
for any vector xin our vector space. Two operators AandBare equal if
Ax=Bxfor all vectors x. Finally, if there exists an operator A
−1such that
AA−1=A−1A=I
thenA−1is the inverse ofA. Some linear operators do not possess an inverse
and are called singular , whilst those operators that do have an inverse are termed
non-singular .
8.3 Matrices
We have seen that in a particular basis eiboth vectors and linear operators
can be described in terms of their components with respect to the basis. Thesecomponents may be displayed as an array of numbers called a matrix .I ng e n e r a l ,
if a linear operator Atransforms vectors from an N-dimensional vector space,
for which we choose a basis e
j,j=1,2,...,N , into vectors belonging to an
M-dimensional vector space, with basis fi,i=1,2,...,M , then we may represent
the operator Aby the matrix
A=
A11 A12... A 1N
A21 A22... A 2N
............
A
M1AM2... A MN
. (8.25)
Thematrix elements Aijare the components of the linear operator with respect
to the bases ejandfi; the component Aijof the linear operator appears in the
ith row and jth column of the matrix. The array has Mrows and Ncolumns
a n di st h u sc a l l e da n M×Nmatrix. If the dimensions of the two vector spaces
are the same, i.e. M=N(for example, if they are the same vector space) then we
may represent Aby an N×Norsquare matrix of order N. The component Aij,
which in general may be complex, is also denoted by ( A)ij.
254
8.4 BASIC MATRIX ALGEBRA
In a similar way we may denote a vector xin terms of its components xiin a
basisei,i=1,2,...,N , by the array
x=
x1
x2
...
xN
,
which is a special case of (8.25) and is called a column matrix (or conventionally,
and slightly confusingly, a column vector or even just a vector – strictly speaking
the term ‘vector’ refers to the geometrical entity x). The column matrix xcan also
be written as
x=(x1x2··· xN)T,
which is the transpose of arow matrix (see section 8.6).
We note that in a different basis e/prime
ithe vector xwould be represented by a
different column matrix containing the components x/prime
iin the new basis, i.e.
x/prime=
x/prime
1
x/prime2
...
x/prime
N
.
Thus, we use xand x/primeto denote different column matrices which, in different bases
eiande/prime
i, represent the samevector x.I nm a n yt e x t s ,h o w e v e r ,t h i sd i s t i n c t i o ni s
not made and x(rather than x) is equated to the corresponding column matrix; if
we regard xas the geometrical entity, however, this can be misleading and so we
explicitly make the distinction. A similar argument follows for linear operators;the same linear operator Ais described in different bases by different matrices A
and A
/prime, containing different matrix elements.
8.4 Basic matrix algebra
The basic algebra of matrices may be deduced from the properties of the linear
operators that they represent. In a given basis the action of two linear operatorsAandBon an arbitrary vector x(see the beginning of subsection 8.2.1), when
written in terms of components using (8.24), is given by
summationdisplay
j(A+B)ijxj=summationdisplay
jAijxj+summationdisplay
jBijxj,
summationdisplay
j(λA)ijxj=λsummationdisplay
jAijxj,
summationdisplay
j(AB)ijxj=summationdisplay
kAik(Bx)k=summationdisplay
jsummationdisplay
kAikBkjxj.
255
MATRICES AND VECTOR SPACES
Now, since xis arbitrary, we can immediately deduce the way in which matrices
are added or multiplied, i.e.
(A+B)ij=Aij+Bij, (8.26)
(λA)ij=λAij, (8.27)
(AB)ij=summationdisplay
kAikBkj. (8.28)
We note that a matrix element may, in general, be complex. We now discuss
matrix addition and multiplication in more detail.
8.4.1 Matrix addition and multiplication by a scalar
From (8.26) we see that the sum of two matrices, S=A+B, is the matrix whose
elements are given by
Sij=Aij+Bij
for every pair of subscripts i, j, with i=1,2,...,M and j=1,2,...,N .F o r
example, if Aand Bare 2×3 matrices then S=A+Bis given by
parenleftbiggS11S12S13
S21S22S23parenrightbigg
=parenleftbiggA11A12A13
A21A22A23parenrightbigg
+parenleftbiggB11B12B13
B21B22B23parenrightbigg
=parenleftbiggA11+B11A12+B12A13+B13
A21+B21A22+B22A23+B23parenrightbigg
. (8.29)
Clearly, for the sum of two matrices to have any meaning, the matrices must have
the same dimensions, i.e. both be M×Nmatrices.
From definition (8.29) it follows that A+B=B+Aand that the sum of a
number of matrices can be written unambiguously without bracketting, i.e. matrixaddition is commutative andassociative .
The difference of two matrices is defined by direct analogy with addition. The
matrix D=A−Bhas elements
D
ij=Aij−Bij,fori=1,2,...,M ,j=1,2,...,N . (8.30)
From (8.27) the product of a matrix Awith a scalar λis the matrix with
elements λAij, for example
λparenleftbiggA11A12A13
A21A22A23parenrightbigg
=parenleftbiggλA11λA12λA13
λA21λA22λA23parenrightbigg
. (8.31)
Multiplication by a scalar is distributive and associative.
256
8.4 BASIC MATRIX ALGEBRAIThe matrices A,Band Care given by
A=
/
2−1
31
/
, B=
/
10
0−2
/
, C=
/
−21
−11
/
.
Find the matrix D=A+2B−C.
D=
/
2−1
31
/
+2
/
100−2
/
−
/
−21
−11
/
=
/
2+2×1−(−2)−1+2×0−1
3+2×0−(−1) 1 + 2×(−2)−1
/
=
/
6−2
4−4
/
.
J
From the above considerations we see that the set of all, in general complex,
M×Nmatrices (with fixed MandN) forms a linear vector space of dimension
MN. One basis for the space is the set of M×Nmatrices E(p,q)with the property
that E(p,q)
ij=1i f i=pandj=qwhilst E(p,q)
ij= 0 for all other values of iand
j, i.e. each matrix has only one non-zero entry, which equals unity. Here the pair
(p, q) is simply a label that picks out a particular one of the matrices E(p,q),t h e
total number of which is MN.
8.4.2 Multiplication of matrices
Let us consider again the ‘transformation’ of one vector into another, y=Ax,
which, from (8.24), may be described in terms of components with respect to aparticular basis as
y
i=Nsummationdisplay
j=1Aijxjfori=1,2,...,M. (8.32)
Writing this in matrix form as y=Axwe have
y
1
y2
...
yM
=
A
11 A12... A 1N
A21 A22... A 2N
............
AM1AM2... A MN
x1
x2
...
xN
(8.33)
where we have highlighted with boxes the components used to calculate the
element y
2: using (8.32) for i=2 ,
y2=A21x1+A22x2+···+A2NxN.
All the other components yiare calculated similarly.
If instead we operate with Aon a basis vector ejhaving all components zero
257
MATRICES AND VECTOR SPACES
except for the jth, which equals unity, then we find
Aej=
A
11 A12... A 1N
A21 A22... A 2N
............
A
M1AM2... A MN
0
0
...
1
...
0
=
A
1j
A2j
...
AMj
,
and so confirm our identification of the matrix element A
ijas the ith component
ofAejin this basis.
From (8.28) we can extend our discussion to the product of two matrices
P=AB,w h e r e Pis the matrix of the quantities formed by the operation of
the rows of Aon the columns of B, treating each column of Bin turn as the
vector xrepresented in component form in (8.32). It is clear that, for this to be
a meaningful definition, the number of columns in Amust equal the number of
rows in B. Thus the product ABof an M×Nmatrix Awith an N×Rmatrix B
is itself an M×Rmatrix P,w h e r e
Pij=Nsummationdisplay
k=1AikBkjfori=1,2,...,M ,j=1,2,...,R.
For example, P=ABmay be written in matrix form
parenleftBigg
P11 P12
P21 P22parenrightBigg
=parenleftbiggA11A12A13
A21A22A23parenrightbigg
B11B12
B21B22
B31B32
where
P
11=A11B11+A12B21+A13B31,
P21=A21B11+A22B21+A23B31,
P12=A11B12+A12B22+A13B32,
P22=A21B12+A22B22+A23B32.
Multiplication of more than two matrices follows naturally and is associative.
So, for example,
A(BC)≡(AB)C, (8.34)
provided, of course, that all the products are defined.
As mentioned above, if Ais an M×Nmatrix and Bis an N×Mmatrix then
two product matrices are possible, i.e.
P=AB and Q=BA.
258
8.4 BASIC MATRIX ALGEBRA
These are clearly not the same, since Pis an M×Mmatrix whilst Qis an
N×Nmatrix. Thus, particular care must be taken to write matrix products in
the intended order; P=ABbut Q=BA. We note in passing that A2means AA,
A3means A(AA)=( AA)Aetc. Even if both Aand Bare square, in general
AB/negationslash=BA, (8.35)
i.e. the multiplication of matrices is not, in general, commutative.IEvaluate P=ABand Q=BAwhere
A=
/0/@32−1
03 21−34
/1A, B=
/0/@2−23
110321
/1A.
As we saw for the 2 ×2 case above, the element Pijof the matrix P=ABis found by
mentally taking the ‘scalar product’ of the ith row of Awith the jth column of B.F o r
example, P11=3×2+2×1+(−1)×3=5 , P12=3×(−2) + 2×1+(−1)×2=−6, etc.
Thus
P=AB=
/0/@32−1
03 2
1−34
/1A
/0/@2−23
110
321
/1A=
/0
/@
5−68
97 2
1 137
/1A,
and, similarly,
Q=BA=
/0/@2−23
110
321
/1A
/0/@32−1
03 2
1−34
/1A=
/0
/@
9−11 6
35 1
10 9 5
/1A.
These results illustrate that, in general, two matrices do not commute.
J
The property that matrix multiplication is distributive over addition, i.e. that
(A+B)C=AC+BC (8.36)
and
C(A+B)=CA+CB, (8.37)
follows directly from its definition.
8.4.3 The null and identity matrices
Both the null matrix and the identity matrix are frequently encountered, and we
take this opportunity to introduce them briefly, leaving their uses until later. Thenullorzeromatrix 0has all elements equal to zero, and so its properties are
A0=0=0A,
A+0=0+A=A.
259
MATRICES AND VECTOR SPACES
Theidentity matrix Ihas the property
AI=IA=A.
It is clear that, in order for the above products to be defined, the identity matrix
must be square. The N×Nidentity matrix (often denoted by IN)h a st h ef o r m
IN=
10 ···0
01...
......0
0··· 01
.
8.5 Functions of matrices
If a matrix Aissquare then, as mentioned above, one can define powers ofAis a
straightforward way. For example A2=AA,A3=AAA, or in the general case
An=AA···A (ntimes) ,
where nis a positive integer. Having defined powers of a square matrix A,w e
may construct functions ofAof the form
S=summationdisplay
nanAn,
where the akare simple scalars and the number of terms in the summation may
be finite or infinite. In the case where the sum has an infinite number of terms,the sum has meaning only if it converges. A common example of such a functionis the exponential of a matrix, which is defined by
exp A=
∞summationdisplay
n=0An
n!. (8.38)
This definition can, in turn, be used to define other functions such as sin Aand
cosA.
8.6 The transpose of a matrix
We have seen that the components of a linear operator in a given coordinate sys-
tem can be written in the form of a matrix A. We will also find it useful, however,
to consider the different (but clearly related) matrix formed by interchanging the
rows and columns of A. The matrix is called the transpose ofAand is denoted
byAT.
260
8.7 THE COMPLEX AND HERMITIAN CONJUGATES OF A MATRIXIFind the transpose of the matrix
A=
/
312
041
/
.
By interchanging the rows and columns of Awe immediately obtain
AT=
/0/@30
1421
/1A.
J
It is obvious that if Ais an M×Nmatrix then its transpose ATis aN×M
matrix. As mentioned in section 8.3, the transpose of a column matrix is a
row matrix and vice versa. An important use of column and row matrices isin the representation of the inner product of two real vectors in terms of theircomponents in a given basis. This notion is discussed fully in the next section,where it is extended to complex vectors.
The transpose of the product of two matrices, ( AB)
T, is given by the product
of their transposes taken in the reverse order, i.e.
(AB)T=BTAT. (8.39)
This is proved as follows:
(AB)T
ij=(AB)ji=summationdisplay
kAjkBki
=summationdisplay
k(AT)kj(BT)ik=summationdisplay
k(BT)ik(AT)kj=(BTAT)ij,
and the proof can be extended to the product of several matrices to give
(ABC···G)T=GT···CTBTAT.
8.7 The complex and Hermitian conjugates of a matrix
Two further matrices that can be derived from a given general M×Nmatrix
are the complex conjugate , denoted by A∗,a n dt h e Hermitian conjugate , denoted
byA†.
The complex conjugate of a matrix Ais the matrix obtained by taking the
complex conjugate of each of the elements of A,i . e .
(A∗)ij=(Aij)∗.
Obviously if a matrix is real(i.e. it contains only real elements) then A∗=A.
261
MATRICES AND VECTOR SPACESIFind the complex conjugate of the matrix
A=
/
12 3 i
1+i10
/
.
By taking the complex conjugate of each element we obtain immediately
A∗=
/
12−3i
1−i10
/
.
J
The Hermitian conjugate, or adjoint ,o fam a t r i x Ais the transpose of its
complex conjugate, or equivalently, the complex conjugate of its transpose, i.e.
A†=(A∗)T=(AT)∗.
We note that if Ais real (and so A∗=A)t h e n A†=AT, and taking the Hermitian
conjugate is equivalent to taking the transpose. Following the previous line ofargument for the transpose of the product of several matrices, the Hermitianconjugate of such a product can be shown to be given by
(AB···G)
†=G†···B†A†. (8.40)IFind the Hermitian conjugate of the matrix
A=
/
12 3 i
1+i10
/
.
Taking the complex conjugate of Aand then forming the transpose we find
A†=
/0/@11−i
21
−3i0
/1A.
We obtain the same result, of course, if we first take the transpose of Aand then take the
complex conjugate.
J
An important use of the Hermitian conjugate (or transpose in the real case)
is in connection with the inner product of two vectors. Suppose that in a givenorthonormal basis the vectors aandbmay be represented by the column matrices
a=
a
1
a2
...
aN
and b=
b1
b2
...
bN
. (8.41)
Taking the Hermitian conjugate of a, to give a row matrix, and multiplying (on
262
8 . 8T H ET R A C EO FAM A T R I X
the right) by bwe obtain
a†b=(a∗
1a∗2···a∗
N)
b1
b2
...
bN
=Nsummationdisplay
i=1a∗
ibi, (8.42)
which is the expression for the inner product /angbracketlefta|b/angbracketrightin that basis. We note that for
real vectors (8.42) reduces to aTb=summationtextN
i=1aibi.
If the basis eiisnotorthonormal, so that, in general,
/angbracketleftei|ej/angbracketright=Gij/negationslash=δij,
then, from (8.18), the scalar product of aandbin terms of their components with
respect to this basis is given by
/angbracketlefta|b/angbracketright=Nsummationdisplay
i=1Nsummationdisplay
j=1a∗
iGijbj=a†Gb,
where Gis the N×Nmatrix with elements Gij.
8.8 The trace of a matrix
For a given matrix A, in the previous two sections we have considered various
other matrices that can be derived from it. However, sometimes one wishes toderive a single number from a matrix. The simplest example is the trace(orspur)
of a square matrix, which is denoted by Tr A. This quantity is defined as the sum
of the diagonal elements of the matrix,
TrA=A
11+A22+···+ANN=Nsummationdisplay
i=1Aii. (8.43)
It clear that the trace is a linear operation so that, for example,
Tr(A±B)=T r A±TrB.
A very useful property of traces is that the trace of the product of two matrices
is independent of the order of their multiplication; this results holds whether ornot the matrices commute and is proved as follows:
TrAB= Nsummationdisplay
i=1(AB)ii=Nsummationdisplay
i=1Nsummationdisplay
j=1AijBji=Nsummationdisplay
i=1Nsummationdisplay
j=1BjiAij=Nsummationdisplay
j=1(BA)jj=T r BA.
(8.44)
The result can be extended to the product of several matrices. For example, from
(8.44), we immediately find
TrABC=T r BCA=T r CAB,
263
MATRICES AND VECTOR SPACES
which shows that the trace of a product is invariant under cyclic permutations of
the matrices in the product. Other easily derived properties of the trace are, forexample, Tr A
T=T r Aand Tr A†=( T r A)∗.
8.9 The determinant of a matrix
For a given matrix A, the determinant det A(like the trace) is a single number (or
algebraic expression) that depends upon the elements of A. Also like the trace,
the determinant is defined only for square matrices. If, for example, Ais a 3×3
matrix then its determinant, of order 3, is denoted by
detA=|A|=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA
11A12A13
A21A22A23
A31A32A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (8.45)
In order to calculate the value of a determinant, we first need to introduce
the notions of the minor and the cofactor of an element of a matrix. (We
shall see that we can use the cofactors to write an order-3 determinant as theweighted sum of three order-2 determinants, thereby simplifying its evaluation.)The minor M
ijof the element Aijof an N×Nmatrix Ais the determinant of
the (N−1)×(N−1) matrix obtained by removing all the elements of the ith
row and jth column of A; the associated cofactor, Cij, is found by multiplying
the minor by ( −1)i+j.IFind the cofactor of the element A23of the matrix
A=
/0/@A11A12A13
A21A22A23
A31A32A33
/1
A
.
Removing all the elements of the second row and third column of Aand forming the
determinant of the remaining terms gives the minor
M23=
//
//
A11A12
A31A32
//
//
.
Multiplying the minor by ( −1)2+3=(−1)5=−1g i v e s
C23=−
////A11A12
A31A32
/
///
.
J
We now define a determinant as the sum of the products of the elements of any
row or column and their corresponding cofactors ,e . g . A21C21+A22C22+A23C23or
A13C13+A23C23+A33C33.S u c has u mi sc a l l e da Laplace expansion . For example,
in the first of these expansions, using the elements of the second row of the
264
8.9 THE DETERMINANT OF A MATRIX
determinant defined by (8.45) and their corresponding cofactors, we write |A|as
the Laplace expansion
|A|=A21(−1)(2+1)M21+A22(−1)(2+2)M22+A23(−1)(2+3)M23
=−A21vextendsinglevextendsinglevextendsinglevextendsingleA
12A13
A32A33vextendsinglevextendsinglevextendsinglevextendsingle+A
22vextendsinglevextendsinglevextendsinglevextendsingleA
11A13
A31A33vextendsinglevextendsinglevextendsinglevextendsingle−A
23vextendsinglevextendsinglevextendsinglevextendsingleA
11A12
A31A32vextendsinglevextendsinglevextendsinglevextendsingle.
We will see later that the value of the determinant is independent of the row
or column chosen. Of course, we have not yet determined the value of |A|but,
rather, written it as the weighted sum of three determinants of order 2. However,applying again the definition of a determinant, we can evaluate each of theorder-2 determinants.IEvaluate the determinant////A12A13
A32A33
////.
By considering the products of the elements of the first row in the determinant, and their
corresponding cofactors, we find///
/
A12A13
A32A33
///
/
=A12(−1)(1+1)|A33|+A13(−1)(1+2)|A32|
=A12A33−A13A32,
where the values of the order-1 determinants |A33|and|A32|are defined to be A33andA32
respectively. It must be remembered that the determinant is notthe same as the modulus,
e.g. det (−2) =|−2|=−2, not 2.
J
We can now combine all the above results to show that the value of the
determinant (8.45) is given by
|A|=−A21(A12A33−A13A32)+A22(A11A33−A13A31)
−A23(A11A32−A12A31) (8.46)
=A11(A22A33−A23A32)+A12(A23A31−A21A33)
+A13(A21A32−A22A31), (8.47)
where the final expression gives the form in which the determinant is usually
remembered and is the form that is obtained immediately by considering theLaplace expansion using the first row of the determinant. The last equality, which
essentially rearranges a Laplace expansion using the second row into one using
the first row, supports our assertion that the value of the determinant is unaffectedby which row or column is chosen for the expansion.
265
MATRICES AND VECTOR SPACESISuppose the rows of a real 3×3matrix Aare interpreted as the components in a given
basis of three (three-component) vectors a,bandc. Show that one can write the determinant
ofAas
|A|=a·(b×c).
If one writes the rows of Aas the components in a given basis of three vectors a,bandc,
we have from (8.47) that
|A|=
//////a1a2a3
b1b2b3
c1c2c3
//////=a1(b2c3−b3c2)+a2(b3c1−b1c3)+a3(b1c2−b2c1).
From expression (7.34) for the scalar triple product given in subsection 7.6.3, it follows
that we may write the determinant as
|A|=a·(b×c). (8.48)
In other words, |A|is the volume of the parallelepiped defined by the vectors a,band
c. (One could equally well interpret the columns of the matrix Aas the components of
three vectors, and result (8.48) would still hold.) This result provides a more memorable(and more meaningful) expression than (8.47) for the value of a 3 ×3 determinant. Indeed,
using this geometrical inte rpretation, we see immediately that, if the vectors a
1,a2,a3are
not linearly independent then the value of the determinant vanishes: |A|=0 .
J
The evaluation of determinants of order greater than 3 follows the same general
method as that presented above, in that it relies on successively reducing the orderof the determinant by writing it as a Laplace expansion. Thus, a determinantof order 4 is first written as a sum of four determinants of order 3, which
are then evaluated using the above method. For higher-order determinants, one
cannot write down directly a simple geometrical expression for |A|analogous to
that given in (8.48). Nevertheless, it is still true that if the rows or columns oftheN×Nmatrix Aare interpreted as the components in a given basis of N
(N-component) vectors a
1,a2,...,aN, then the determinant |A|vanishes if these
vectors are not all linearly independent.
8.9.1 Properties of determinants
A number of properties of determinants follow straightforwardly from the defini-
tion of det A; their use will often reduce the labour of evaluating a determinant.
We present them here without specific proofs, though they all follow readily fromthe alternative form for a determinant, given in equation (21.28) on page 791,and expressed in terms of the Levi–Civita symbol /epsilon1
ijk(see exercise 21.9).
(i)Determinant of the transpose . The transpose matrix AT(which, we recall,
is obtained by interchanging the rows and columns of A) has the same
determinant as Aitself, i.e.
|AT|=|A|. (8.49)
266
8.9 THE DETERMINANT OF A MATRIX
It follows that anytheorem established for the rows of Awill apply to the
columns as well, and vice versa.
(ii)Determinant of the complex and Hermitian conjugate . It is clear that the
matrix A∗obtained by taking the complex conjugate of each element of A
has the determinant |A∗|=|A|∗. Combining this result with (8.49), we find
that
|A†|=|(A∗)T|=|A∗|=|A|∗. (8.50)
(iii)Interchanging two rows or two columns . If two rows (columns) of Aare
interchanged, its determinant changes sign but is unaltered in magnitude.
(iv)Removing factors . If all the elements of a single row (column) of Ahave
a common factor, λ, then this factor may be removed; the value of the
determinant is given by the product of the remaining determinant and λ.
Clearly this implies that if all the elements of any row (column) are zerothen|A|= 0. It also follows that if every element of the N×Nmatrix A
is multiplied by a constant factor λthen
|λA|=λ
N|A|. (8.51)
(v)Identical rows or columns . If any two rows (columns) of Aare identical or
are multiples of one another, then it can be shown that |A|=0 .
(vi)Adding a constant multiple of one row (column) to another . The determinant
of a matrix is unchanged in value by adding to the elements of one row(column) any fixed multiple of the elements of another row (column).
(vii)Determinant of a product .I fAand Bare square matrices of the same order
then
|AB|=|A||B|=|BA|. (8.52)
A simple extension of this property gives, for example,
|AB···G|=|A||B|···|G|=|A||G|···|B|=|A···GB|,
which shows that the determinant is invariant to cyclic permutations of
the matrices in the product.
There is no explicit procedure for using the above results in the evaluation of
any given determinant, and judging the quickest route to an answer is a matterof experience. A general guide is to try to reduce all terms but one in a row orcolumn to zero and hence in effect to obtain a determinant of smaller size. The
steps taken in evaluating the determinant in the example below are certainly not
the fastest, but they have been chosen in order to illustrate the use of most of theproperties listed above.
267
MATRICES AND VECTOR SPACESIEvaluate the determinant
|A|=
//
///
//
1023
01−21
3−34−2
−21−2−1
//
///
//
.
Taking a factor 2 out of the third column and then adding the second column to the third
gives
|A|=2
///////1013
01−11
3−32−2
−21−1−1
///////=2
///////1013
01013−3−1−2
−21 0 −1
///////.
Subtracting the second column from the fourth gives
|A|=2
///////1013
01003−3−11
−21 0 −2
///////.
We now note that the second row has only one non-zero element and so the determinant
may conveniently be written as a Laplace expansion, i.e.
|A|=2×1×(−1)2+2
///
///
113
3−11
−20−2
///
///
=2
///
///
404
3−11
−20−2
///
///
,
where the last equality follows by adding the second row to the first. It can now be seen
that the first row is minus twice the third, and so the value of the determinant is zero, byproperty (v) above.J
8.10 The inverse of a matrix
Our first use of determinants will be in defining the inverse of a matrix. If we
were dealing with ordinary numbers we would consider the relation P=ABas
equivalent to B=P/A, provided that A/negationslash= 0. However, if A,Band Pare matrices
then this notation does not have an obvious meaning. What we really want toknow is whether an explicit formula for Bcan be obtained in terms of Aand
P. It will be shown that this is possible for those cases in which |A|/negationslash=0 .A
square matrix whose determinant is zero is called a singular matrix; otherwise it
isnon-singular . We will show that if Ais non-singular we can define a matrix,
denoted by A
−1and called the inverse ofA, which has the property that if AB=P
then B=A−1P.I nw o r d s , Bcan be obtained by multiplying Pfrom the left by
A−1. Analogously, if Bis non-singular then, by multiplication from the right,
A=PB−1.
It is clear that
AI=A⇒ I=A−1A, (8.53)
268
8.10 THE INVERSE OF A MATRIX
where Iis the unit matrix, and so A−1A=I=AA−1. These statements are
equivalent to saying that if we first multiply a matrix, Bsay, by Aand then
multiply by the inverse A−1, we end up with the matrix we started with, i.e.
A−1AB=B. (8.54)
This justifies our use of the term inverse. It is also clear that the inverse is only
defined for square matrices.
So far we have only defined what we mean by the inverse of a matrix. Actually
finding the inverse of a matrix Amay be carried out in a number of ways. We will
show that one method is to construct first the matrix Ccontaining the cofactors
of the elements of A, as discussed in the last subsection. Then the required inverse
A−1can be found by forming the transpose of Cand dividing by the determinant
ofA. Thus the elements of the inverse A−1are given by
(A−1)ik=(C)T
ik
|A|=Cki
|A|. (8.55)
That this procedure does indeed result in the inverse may be seen by considering
the components of A−1A,i . e .
(A−1A)ij=summationdisplay
k(A−1)ik(A)kj=summationdisplay
kCki
|A|Akj=|A|
|A|δij. (8.56)
The last equality in (8.56) relies on the property
summationdisplay
kCkiAkj=|A|δij; (8.57)
this can be proved by considering the matrix A/primeobtained from the original matrix
Awhen the ith column of Ais replaced by one of the other columns, say the jth.
Thus A/primeis a matrix with two identical columns and so has zero determinant.
However, replacing the ith column by another does not change the cofactors Cki
of the elements in the ith column, which are therefore the same in Aand A/prime.
Recalling the Laplace expansion of a determinant, i.e.
|A|=summationdisplay
kAkiCki,
we obtain
0=|A/prime|=summationdisplay
kA/prime
kiC/prime
ki=summationdisplay
kAkjCki,i/negationslash=j,
which together with the Laplace expansion itself may be summarised by (8.57).
It is immediately obvious from (8.55) that the inverse of a matrix is not defined
if the matrix is singular (i.e. if |A|=0 ) .
269
MATRICES AND VECTOR SPACESIFind the inverse of the matrix
A=
/0/@243
1−2−2
−33 2
/1A.
We first determine |A|:
|A|=2 [−2(2)−(−2)3] + 4[(−2)(−3)−(1)(2)] + 3[(1)(3) −(−2)(−3)]
=1 1. (8.58)
This is non-zero and so an inverse matrix can be constructed. To do this we need the
matrix of the cofactors, C, and hence CT. We find
C=
/0/@24−3
11 3−18
−27−8
/1A and CT=
/0/@21−2
41 37
−3−18−8
/1A,
and hence
A−1=CT
|A|=1
11
/0/@21−2
41 37
−3−18−8
/1A.
J (8.59)
For a 2×2 matrix, the inverse has a particularly simple form. If the matrix is
A=parenleftbiggA11A12
A21A22parenrightbigg
then its determinant |A|is given by |A|=A11A22−A12A21,a n dt h em a t r i xo f
cofactors is
C=parenleftbiggA22−A21
−A12 A11parenrightbigg
.
Thus the inverse of Ais given by
A−1=CT
|A|=1
A11A22−A12A21parenleftbiggA22−A12
−A21 A11parenrightbigg
. (8.60)
It can be seen that the transposed matrix of cofactors for a 2 ×2m a t r i xi st h e
same as the matrix formed by swapping the elements on the leading diagonal
(A11andA22) and changing the signs of the other two elements ( A12andA21).
This is completely general for a 2 ×2 matrix and is easy to remember.
The following are some further useful properties related to the inverse matrix
270
8.10 THE INVERSE OF A MATRIX
and may be straightforwardly derived.
(i) ( A−1)−1=A.
(ii) ( AT)−1=(A−1)T.
(iii) ( A†)−1=(A−1)†.
(iv) ( AB)−1=B−1A−1.
(v) ( AB···G)−1=G−1···B−1A−1.IProve the properties (i)–(v) stated above.
We begin by writing down the fundamental expression defining the inverse of a non-
singular square matrix A:
AA−1=I=A−1A. (8.61)
Property (i): This follows immediately from the expression (8.61).
Property (ii): Taking the transpose of each expression in (8.61) gives
(AA−1)T=IT=(A−1A)T.
Using the result (8.39) for the transpose of a product of matrices and noting that IT=I,
we find
(A−1)TAT=I=AT(A−1)T.
However, from (8.61), this implies ( A−1)T=(AT)−1and hence proves result (ii) above.
Property (iii): This may be proved in an analogous way to property (ii), by replacing
the transposes in (ii) by Hermitian conjugates and using the result (8.40) for the Hermitianconjugate of a product of matrices.
Property (iv): Using (8.61), we may write
(AB)(AB)
−1=I=(AB)−1(AB),
From the left-hand equality it follo ws, by multiplying on the left by A−1,t h a t
A−1AB(AB)−1=A−1Iand hence B(AB)−1=A−1.
Now multiplying on the left by B−1gives
B−1B(AB)−1=B−1A−1,
and hence the stated result.
Property (v): Finally, result (iv) may extended to case (iv) in a straightforward manner.
For example, using result (iv) twice we find
(ABC)−1=(BC)−1A−1=C−1B−1A−1.
J
We conclude this section by noting that the determinant |A−1|of the inverse
matrix can be expressed very simply in terms of the determinant |A|of the matrix
itself. Again we start with the fundamental expression (8.61). Then, using the
property (8.52) for the determinant of a product, we find
|AA−1|=|A||A−1|=|I|.
It is straightforward to show by Laplace expansion that |I|= 1, and so we arrive
at the useful result
|A−1|=1
|A|. (8.62)
271
MATRICES AND VECTOR SPACES
8.11 The rank of a matrix
Therankof a general M×Nmatrix is an important concept, particularly in
the solution of sets of simultaneous linear equations, to be discussed in the nextsection, and we now discuss it in some detail. Like the trace and determinant,
the rank of matrix Ais a single number (or algebraic expression) that depends
on the elements of A. Unlike the trace and determinant, however, the rank of a
matrix can be defined even when Ais not square. As we shall see, there are two
equivalent definitions of the rank of a general matrix.
Firstly, the rank of a matrix may be defined in terms of the linear independence
of vectors. Suppose that the columns of an M×Nmatrix are interpreted as
the components in a given basis of N(M-component) vectors v
1,v2,...,vN,a s
follows:
A=
↑↑ ↑
v1v2... vN
↓↓ ↓
.
Then the rankofA, denoted by rank Aor by R(A), is defined as the number
oflinearly independent vectors in the set v1,v2,...,vN, and equals the dimension
of the vector space spanned by those vectors. Alternatively, we may consider therows of Ato contain the components in a given basis of the M(N-component)
vectors w
1,w2,...,wMas follows:
A=
← w1→
← w2→
...
← wM→
.
It may then be shown †that the rank of Ais also equal to the number of
linearly independent vectors in the set w1,w2,...,wM. From this definition it is
should be clear that the rank of Ais unaffected by the exchange of two rows
(or two columns) or by the multiplication of a row (or column) by a constant.
Furthermore, suppose that a constant multiple of one row (column) is added to
another row (column): for example, we might replace the row wibywi+cwj.
This also has no effect on the number of linearly independent rows and so leavesthe rank of Aunchanged. We may use these properties to evaluate the rank of a
given matrix.
A second (equivalent) definition may be given of the rank of a matrix and uses
the concept of submatrices . A submatrix of Ais any matrix that can be formed
from the elements of Aby ignoring one. or more than one, row or column. It
†For a fuller discussion, see, for example, Modern Mathematical Methods for Physicists and Engineers ,
chapter 6, C. D. Cantrell (Cambridge University Press).
272
8.12 SPECIAL TYPES OF SQUARE MATRIX
may be shown that the rank of a general M×Nmatrix is equal to the size of
the largest square submatrix of Awhose determinant is non-zero. Therefore, if a
matrix Ahas an r×rsubmatrix Swith|S|/negationslash= 0, but no ( r+1)×(r+1) submatrix
with non-zero determinant then the rank of the matrix is r. From either definition
it is clear that the rank of Ais less than or equal to the smaller of MandN.IDetermine the rank of the matrix
A=
/0/@110−2
202 2
413 1
/1A.
The largest possible square submatrices of Amust be of dimension 3 ×3. Clearly, A
possesses four such submatrices, the determinants of which are given by//////110
202413
//////=0,
//////11−2
20 241 1
//////=0,//////10−2
22 243 1
//////=0,
//////10−2
02 213 1
//////=0.
(In each case the determinant may be evaluated as described in subsection 8.9.1.)
The next largest square submatrices of Aare of dimension 2 ×2. Consider, for example,
the 2×2 submatrix formed by ignoring the third row and the third and fourth columns
ofA; this has determinant////11
20
////=1×0−2×1=−2.
Since its determinant is non-zero, Ais of rank 2 and we need not consider any other 2 ×2
submatrices.
J
In the special case in which the matrix Ais asquare N×Nmatrix, by comparing
either of the above definitions of rank with our discussion of determinants insection 8.9, we see that |A|= 0 unless the rank of AisN. In other words, Ais
singular unless R(A)=N.
8.12 Special types of square matrix
Matrices that are square, i.e. N×N, are very common in physical applications.
We now consider some special forms of square matrix that are of particularimportance.
8.12.1 Diagonal matrices
The unit matrix, which we have already encountered, is an example of a diagonal
matrix. Such matrices are characterised by having non-zero elements only on the
273
MATRICES AND VECTOR SPACES
leading diagonal , i.e. only elements Aijwith i=jmay be non-zero. For example,
A=
10 0
02 000−3
,
is a 3×3 diagonal matrix. Such a matrix is often denoted by A= diag (1 ,2,−3).
By performing a Laplace expansion, it is easily shown that the determinant of anN×Ndiagonal matrix is equal to the product of the diagonal elements. Thus, if
the matrix has the form A=d i a g ( A
11,A22,...,A NN)t h e n
|A|=A11A22···ANN. (8.63)
Moreover, it is also straightforward to show that the inverse of Ais also a
diagonal matrix given by
A−1= diagparenleftbigg1
A11,1
A22,...,1
ANNparenrightbigg
.
Finally, we note that, if two matrices Aand Barebothdiagonal then they have
the useful property that their product is commutative:
AB=BA.
This is nottrue for matrices in general.
8.12.2 Lower and upper triangular matrices
A square matrix Ais called lower triangular if all the elements above the principal
diagonal are zero. For example, the general form for a 3 ×3 lower triangular
matrix is
A=
A1100
A21A220
A31A32A33
,
where the elements Aijmay be zero or non-zero. Similarly an upper triangular
square matrix is one for which all the elements below the principal diagonal are
zero. The general 3 ×3 form is thus
A=
A11A12A13
0A22A23
00 A33
.
By performing a Laplace expansion, it is straightforward to show that, in the
general N×Ncase, the determinant of an upper or lower triangular matrix is
equal to the product of its diagonal elements,
|A|=A11A22···ANN. (8.64)
274
8.12 SPECIAL TYPES OF SQUARE MATRIX
Clearly result (8.63) for diagonal matrices is a special case of this result. Moreover,
it may be shown that the inverse of a non-singular lower (upper) triangular matrixis also lower (upper) triangular.
8.12.3 Symmetric and antisymmetric matrices
A square matrix Aof order Nwith the property A=A
Tis said to be symmetric .
Similarly a matrix for which A=−ATis said to be anti-orskew-symmetric
and its diagonal elements a11,a22,...,a NNare necessarily zero. Moreover, if Ais
(anti-)symmetric then so too is its inverse A−1. This is easily proved by noting
that if A=±ATthen
(A−1)T=(AT)−1=±A−1.
Any N×Nmatrix Acan be written as the sum of a symmetric and an
antisymmetric matrix, since we may write
A=1
2(A+AT)+1
2(A−AT)=B+C,
where clearly B=BTand C=−CT. The matrix Bis therefore called the
symmetric part of A,a n d Cis the antisymmetric part.IIfAis an N×Nantisymmetric matrix, show that |A|=0ifNis odd.
IfAis antisymmetric then AT=−A. Using the properties of determinants (8.49) and
(8.51), we have
|A|=|AT|=|−A|=(−1)N|A|.
Thus, if Nis odd then |A|=−|A|, which implies that |A|=0 .
J
8.12.4 Orthogonal matrices
A non-singular matrix with the property that its transpose is also its inverse,
AT=A−1, (8.65)
is called an orthogonal matrix . It follows immediately that the inverse of an
orthogonal matrix is also orthogonal, since
(A−1)T=(AT)−1=(A−1)−1.
Moreover, since for an orthogonal matrix ATA=I, we have
|ATA|=|AT||A|=|A|2=|I|=1.
Thus the determinant of an orthogonal matrix must be |A|=±1.
An orthogonal matrix represents, in a particular basis, a linear operator that
leaves the norms (lengths) of real vectors unchanged, as we will now show.
275
MATRICES AND VECTOR SPACES
Suppose that y=Axis represented in some coordinate system by the matrix
equation y=Ax;t h e n/angbracketlefty|y/angbracketrightis given in this coordinate system by
yTy=xTATAx=xTx.
Hence/angbracketlefty|y/angbracketright=/angbracketleftx|x/angbracketright, showing that the action of a linear operator represented by
an orthogonal matrix does not change the norm of a real vector.
8.12.5 Hermitian and anti-Hermitian matrices
AnHermitian matrix is one that satisfies A=A†,w h e r e A†is the Hermitian
conjugate discussed in section 8.7. Similarly if A†=−A,t h e n Ais called anti-
Hermitian . A real (anti-)symmetric matrix is a special case of an (anti-)Hermitian
matrix, in which all the elements of the matrix are real. Also, if Ais an (anti-
)Hermitian matrix then so too is its inverse A−1,s i n c e
(A−1)†=(A†)−1=±A−1.
Any N×Nmatrix Acan be written as the sum of an Hermitian matrix and
an anti-Hermitian matrix, since
A=1
2(A+A†)+1
2(A−A†)=B+C,
where clearly B=B†and C=−C†. The matrix Bis called the Hermitian part of
A,a n d Cis called the anti-Hermitian part.
8.12.6 Unitary matrices
Aunitary matrix Ais defined as one for which
A†=A−1. (8.66)
Clearly, if Ais real then A†=AT, showing that a real orthogonal matrix is a
special case of a unitary matrix, one in which all the elements are real. We note
that the inverse A−1of a unitary is also unitary, since
(A−1)†=(A†)−1=(A−1)−1.
Moreover, since for a unitary matrix A†A=I, we have
|A†A|=|A†||A|=|A|∗|A|=|I|=1.
Thus the determinant of a unitary matrix has unit modulus.
A unitary matrix represents, in a particular basis, a linear operator that leaves
the norms (lengths) of complex vectors unchanged. If y=Axis represented in
some coordinate system by the matrix equation y=Axthen/angbracketlefty|y/angbracketrightis given in this
coordinate system by
y†y=x†A†Ax=x†x.
276
8.13 EIGENVECTORS AND EIGENVALUES
Hence/angbracketlefty|y/angbracketright=/angbracketleftx|x/angbracketright, showing that the action of the linear operator represented by
a unitary matrix does not change the norm of a complex vector. The action of aunitary matrix on a complex column matrix thus parallels that of an orthogonalmatrix acting on a real column matrix.
8.12.7 Normal matrices
A final important set of special matrices consists of the normal matrices, for which
AA
†=A†A,
i.e. a normal matrix is one that commutes with its Hermitian conjugate.
We can easily show that Hermitian matrices and unitary matrices (or symmetric
matrices and orthogonal matrices in the real case) are examples of normalmatrices. For an Hermitian matrix, A=A
†and so
AA†=AA=A†A.
Similarly, for a unitary matrix, A−1=A†and so
AA†=AA−1=A−1A=A†A.
Finally, we note that, if Ais normal then so too is its inverse A−1,s i n c e
A−1(A−1)†=A−1(A†)−1=(A†A)−1=(AA†)−1=(A†)−1A−1=(A−1)†A−1.
This broad class of matrices is important in the discussion of eigenvectors and
eigenvalues in the next section.
8.13 Eigenvectors and eigenvalues
Suppose that a linear operator Atransforms vectors xin an N-dimensional
vector space into other vectors Axin the same space. The possibility then arises
that there exist vectors xeach of which is transformed by Ainto a multiple of
itself. Such vectors would have to satisfy
Ax=λx. (8.67)
Any non-zero vector xthat satisfies (8.67) for some value of λis called an
eigenvector of the linear operator A,a n d λis called the corresponding eigenvalue .
As will be discussed below, in general the operator AhasNindependent
eigenvectors xi, with eigenvalues λi.T h e λiare not necessarily all distinct.
If we choose a particular basis in the vector space, we can write (8.67) in terms
of the components of Aandxwith respect to this basis as the matrix equation
Ax=λx, (8.68)
where Ais an N×Nmatrix. The column matrices xthat satisfy (8.68) obviously
277
MATRICES AND VECTOR SPACES
represent the eigenvectors xofAin our chosen coordinate system. Convention-
ally, these column matrices are also referred to as the eigenvectors of the matrix
A.†Clearly, if xis an eigenvector of A(with some eigenvalue λ) then any scalar
multiple µxis also an eigenvector with the same eigenvalue. We therefore often
usenormalised eigenvectors, for which
x†x=1
(note that x†xcorresponds to the inner product /angbracketleftx|x/angbracketrightin our basis). Any eigen-
vector xcan be normalised by dividing all its components by the scalar ( x†x)1/2.
As will be seen, the problem of finding the eigenvalues and corresponding
eigenvectors of a square matrix Aplays an important role in many physical
investigations. Throughout this chapter we denote the ith eigenvector of a square
matrix Abyxiand the corresponding eigenvalue by λi. This superscript notation
for eigenvectors is used to avoid any confusion with components.IA non-singular matrix Ahas eigenvalues λiand eigenvectors xi. Find the eigenvalues and
eigenvectors of the inverse matrix A−1.
The eigenvalues and eigenvectors of Asatisfy
Axi=λixi.
Left-multiplying both sid es of this equation by A−1, we find
A−1Axi=λiA−1xi.
Since A−1A=I, on rearranging we obtain
A−1xi=1
λixi.
Thus, we see that A−1has the same eigenvectors xias does A, but the corresponding
eigenvalues are 1 /λi.
J
In the remainder of this section we will discuss some useful results concerning
the eigenvectors and eigenvalues of certain special (though commonly occurring)square matrices. The results will be established for matrices whose elements maybe complex; the corresponding properties for real matrices may be obtained asspecial cases.
8.13.1 Eigenvectors and eigenvalues of a normal matrix
In subsection 8.12.7 we defined a normal matrix Aas one that commutes with its
Hermitian conjugate, so that
A
†A=AA†.
†In this context, when referring to linear combinations of eigenvectors xwe will normally use the
term ‘vector’.
278
8.13 EIGENVECTORS AND EIGENVALUES
We also showed that both Hermitian and unitary matrices (or symmetric and
orthogonal matrices in the real case) are examples of normal matrices. We nowdiscuss the properties of the eigenvectors and eigenvalues of a normal matrix.
Ifxis an eigenvector of a normal matrix Awith corresponding eigenvalue λ
then Ax=λx, or equivalently,
(A−λI)x=0. (8.69)
Denoting B=A−λI, (8.69) becomes Bx=0and, taking the Hermitian conjugate,
we also have
(Bx)
†=x†B†=0. (8.70)
From (8.69) and (8.70) we then have
x†B†Bx=0. (8.71)
However, the product B†Bis given by
B†B=(A−λI)†(A−λI)=( A†−λ∗I)(A−λI)=A†A−λ∗A−λA†+λλ∗.
Now since Ais normal, AA†=A†A(see subsection 8.12.7) and so
B†B=AA†−λ∗A−λA†+λλ∗=(A−λI)(A−λI)†=BB†,
and hence Bis also normal. From (8.71) we then find
x†B†Bx=x†BB†x=(B†x)†B†x=0,
from which we obtain
B†x=(A†−λ∗I)x=0.
Therefore, for a normal matrix A,the eigenvalues of A†are the complex conjugates
of the eigenvalues of A.
Let us now consider two eigenvectors xiand xjof a normal matrix Acorre-
sponding to two different eigenvalues λiandλj. We then have
Axi=λixi, (8.72)
Axj=λjxj. (8.73)
Multiplying (8.73) on the left by ( xi)†we obtain
(xi)†Axj=λj(xi)†xj. (8.74)
However, on the LHS of (8.74) we have
(xi)†A=(A†xi)†=(λ∗
ixi)†=λi(xi)†, (8.75)
where we have used (8.40) and the property just proved for a normal matrix to
279
MATRICES AND VECTOR SPACES
write A†xi=λ∗
ixi. From (8.74) and (8.75) we have
(λi−λj)(xi)†xj=0. (8.76)
Thus, ifλi/negationslash=λjthe eigenvectors xiand xjmust be orthogonal ,i . e .( xi)†xj=0.
It follows immediately from (8.76) that if all Neigenvalues of a normal matrix
Aare distinct then all Neigenvectors of Aare mutually orthogonal. If, however,
two or more eigenvalues are the same then further consideration is required. Aneigenvalue corresponding to two or more different eigenvectors (i.e. they are notsimply multiples of one another) is said to be degenerate . Suppose that λ
1isk-fold
degenerate, i.e.
Axi=λ1xifori=1,2,...,k, (8.77)
but that it is different from any of λk+1,λk+2, etc. Then any linear combination
of these xiis also an eigenvector with eigenvalue λ1,s i n c e ,f o r z=summationtextk
i=1cixi,
Az≡Aksummationdisplay
i=1cixi=ksummationdisplay
i=1ciAxi=ksummationdisplay
i=1ciλ1xi=λ1z. (8.78)
If the xidefined in (8.77) are not already mutually orthogonal then we can
construct new eigenvectors zithat are orthogonal by the following procedure:
z1=x1,
z2=x2−bracketleftBig
(ˆz1)†x2bracketrightBig
ˆz1,
z3=x3−bracketleftBig
(ˆz2)†x3bracketrightBig
ˆz2−bracketleftBig
(ˆz1)†x3bracketrightBig
ˆz1,
...
zk=xk−bracketleftBig
(ˆzk−1)†xkbracketrightBig
ˆzk−1−···−bracketleftBig
(ˆz1)†xkbracketrightBig
ˆz1.
In this procedure, known as Gram–Schmidt orthogonalisation , each new eigen-
vector ziis normalised to give the unit vector ˆzibefore proceeding to the construc-
tion of the next one (the normalisation is carried out by dividing each element ofthe vector z
iby [( zi)†zi]1/2). Note that each factor in brackets ( ˆzm)†xnis a scalar
product and thus only a number. It follows that, as shown in (8.78), each vectorz
iso constructed is an eigenvector of Awith eigenvalue λ1and will remain so
on normalisation. It is straightforward to check that, provided the previous neweigenvectors have been normalised as prescribed, each z
iis orthogonal to all its
predecessors. (In practice, however, the method is laborious and the example in
subsection 8.14.1 gives a less rigorous but considerably quicker way.)
Therefore, even if Ahas some degenerate eigenvalues we can by construction
obtain a set of Nmutually orthogonal eigenvectors. Moreover, it may be shown
(although the proof is beyond the scope of this book) that these eigenvectorsarecomplete in that they form a basis for the N-dimensional vector space. As
280
8.13 EIGENVECTORS AND EIGENVALUES
a result any arbitrary vector ycan be expressed as a linear combination of the
eigenvectors xi:
y=Nsummationdisplay
i=1aixi, (8.79)
where ai=(xi)†y. Thus, the eigenvectors form an orthogonal basis for the vector
space. By normalising the eigenvectors so that ( xi)†xi=1t h i sb a s i si sm a d e
orthonormal.IShow that a normal matrix Acan be written in terms of its eigenvalues λiand orthogonal
eigenvectors xias
A=NX
i=1λixi(xi)†. (8.80)
The key to proving the validity of (8.80) is to show that both sides of the expression give
the same result when acting on an arbitary vector y.S i n c e Ais normal, we may expand y
in terms of the eigenvectors xi, as shown in (8.79). Thus, we have
Ay=ANX
i=1aixi=NX
i=1aiλixi.
Alternatively, the action of the RHS of (8.80) on yis given by
NX
i=1λixi(xi)†y=NX
i=1aiλixi,
since ai=(xi)†y. We see that the two expressions for the action of each side of (8.80) on y
are identical, which implies that this relationship is indeed correct.
J
8.13.2 Eigenvectors and eigenvalues of Hermitian and anti-Hermitian matrices
For a normal matrix we showed that if Ax=λxthen A†x=λ∗x. However, if Ais
also Hermitian, A=A†, it follows necessarily that λ=λ∗. Thus, the eigenvalues
of an Hermitian matrix are real, a result which may be proved directly.IProve that the eigenvalues of an Hermitian matrix are real.
For any particular eigenvector xi, we take the Hermitian conjugate of Axi=λixito give
(xi)†A†=λ∗
i(xi)†. (8.81)
Using A†=A,s i n c e Ais Hermitian, and multiplying on the right by xi,w eo b t a i n
(xi)†Axi=λ∗
i(xi)†xi. (8.82)
But multiplying Axi=λixithrough on the left by ( xi)†gives
(xi)†Axi=λi(xi)†xi.
281
MATRICES AND VECTOR SPACES
Subtracting this from (8.82) yields
0=(λ∗
i−λi)(xi)†xi.
But ( xi)†xiis the modulus squared of the non-zero vector xiand is thus non-zero. Hence
λ∗
imust equal λiand thus be real. The same argument can be used to show that the
eigenvalues of a real symmetric matrix are themselves real.
J
The importance of the above result will be apparent to any student of quantum
mechanics. In quantum mechanics the eigenvalues of operators correspond to
measured values of observable quantities, e.g. energy, angular momentum, parity
and so on, and these clearly must be real. If we use Hermitian operators toformulate the theories of quantum mechanics, the above property guaranteesphysically meaningful results.
Since an Hermitian matrix is also a normal matrix, its eigenvectors are orthog-
onal (or can be made so using the Gram–Schmidt orthogonalisation procedure).Alternatively we can prove the orthogonality of the eigenvectors directly.IProve that the eigenvectors corresponding to di fferent eigenvalues of an Hermitian matrix
are orthogonal.
Consider two unequal eigenvalues λiandλjand their corresponding eigenvectors satisfying
Axi=λixi, (8.83)
Axj=λjxj. (8.84)
Taking the Hermitian conjugate of (8.83) we find ( xi)†A†=λ∗
i(xi)†. Multiplying this on the
right by xjwe obtain
(xi)†A†xj=λ∗
i(xi)†xj,
and similarly multiplying (8.84) through on the left by ( xi)†we find
(xi)†Axj=λj(xi)†xj.
Then, since A†=A, the two left-hand sides are equal and, because the λiare real, on
subtraction we obtain
0=(λi−λj)(xi)†xj.
Finally we note that λi/negationslash=λjand so ( xi)†xj=0, i.e. the eigenvectors xiand xjare
orthogonal.
J
In the case where some of the eigenvalues are equal, further justification of the
orthogonality of the eigenvectors is needed. The Gram–Schmidt orthogonalisa-tion procedure discussed above provides a proof of, and a means of achieving,orthogonality. The general method has already been described and we will notrepeat it here.
We may also consider the properties of the eigenvalues and eigenvectors of an
anti-Hermitian matrix, for which A
†=−Aand thus
AA†=A(−A)=(−A)A=A†A.
Therefore matrices that are anti-Hermitian are also normal and so have mutu-
ally orthogonal eigenvectors. The properties of the eigenvalues are also simply
282
8.13 EIGENVECTORS AND EIGENVALUES
deduced, since if Ax=λxthen
λ∗x=A†x=−Ax=−λx.
Hence λ∗=−λand so λmust be pure imaginary (orzero). In a similar manner
to that used for Hermitian matrices, these properties may be proved directly.
8.13.3 Eigenvectors and eigenvalues of a unitary matrix
A unitary matrix satisfies A†=A−1and is also a normal matrix, with mutually
orthogonal eigenvectors. To investigate the eigenvalues of a unitary matrix, we
note that if Ax=λxthen
x†x=x†A†Ax=λ∗λx†x,
and we deduce that λλ∗=|λ|2= 1. Thus, the eigenvalues of a unitary matrix
have unit modulus.
8.13.4 Eigenvectors and eigenvalues of a general square matrix
When an N×Nmatrix is not normal there are no general properties of its
eigenvalues and eigenvectors; in general it is not possible to find any orthogonalset of Neigenvectors or even to find pairsof orthogonal eigenvectors (except
by chance in some cases). While the Nnon-orthogonal eigenvectors are usually
linearly independent and hence form a basis for the N-dimensional vector space,
this is not necessarily so. It may be shown (although we will not prove it) that anyN×Nmatrix with distinct eigenvalues has Nlinearly independent eigenvectors,
which therefore form a basis for the N-dimensional vector space. If a general
square matrix has degenerate eigenvalues, however, then it may or may not haveNlinearly independent eigenvectors. A matrix whose eigenvectors are not linearly
independent is said to be defective .
8.13.5 Simultaneous eigenvectors
We may now ask under what conditions two different normal matrices can have
a common set of eigenvectors. The result – that they do so if, and only if, theycommute – has profound significance for the foundations of quantum mechanics.
To prove this important result let Aand Bbe two N×Nnormal matrices and
x
ibe the ith eigenvector of Acorresponding to eigenvalue λi,i . e .
Axi=λixifor i=1,2,... ,N.
For the present we assume that the eigenvalues are all different.
(i) First suppose that Aand Bcommute. Now consider
ABxi=BAxi=Bλixi=λiBxi,
283
MATRICES AND VECTOR SPACES
where we have used the commutativity for the first equality and the eigenvector
property for the second. It follows that A(Bxi)=λi(Bxi) and thus that Bxiis an
eigenvector of Acorresponding to eigenvalue λi. But the eigenvector solutions of
(A−λiI)xi=0are unique to within a scale factor, and we therefore conclude that
Bxi=µixi
for some scale factor µi. However, this is just an eigenvector equation for Band
shows that xiis an eigenvector of B, in addition to being an eigenvector of A.B y
reversing the roles of Aand B, it also follows that every eigenvector of Bis an
eigenvector of A. Thus the two sets of eigenvectors are identical.
(ii) Now suppose that Aand Bhave all their eigenvectors in common, a typical
one xisatisfying both
Axi=λixiand Bxi=µixi.
As the eigenvectors span the N-dimensional vector space, any arbitrary vector x
in the space can be written as a linear combination of the eigenvectors,
x=Nsummationdisplay
i=1cixi.
Now consider both
ABx=ABNsummationdisplay
i=1cixi=ANsummationdisplay
i=1ciµixi=Nsummationdisplay
i=1ciλiµixi,
and
BAx=BANsummationdisplay
i=1cixi=BNsummationdisplay
i=1ciλixi=Nsummationdisplay
i=1ciµiλixi.
It follows that ABxand BAxare the same for any arbitrary xand hence that
(AB−BA)x=0
for all x.T h a ti s , Aand Bcommute .
This completes the proof that a necessary and sufficient condition for two
normal matrices to have a set of eigenvectors in common is that they commute.It should be noted that if an eigenvalue of A, say, is degenerate then not all of
its possible sets of eigenvectors will also constitute a set of eigenvectors of B.
However, provided that by taking linear combinations one set of joint eigenvectorscan be found, the proof is still valid and the result still holds.
When extended to the case of Hermitian operators and continuous eigenfunc-
tions (sections 17.2 and 17.3 the conn ection between commuting matrices and
a set of common eigenvectors plays a fundamental role in the postulatory basis
284
8.14 DETERMINATION OF EIGENVALUES AND EIGENVECTORS
of quatum mechanics. It draws the distinction between commuting and non-
commuting observables and sets limits on how much information about a systemcan be known, even in principle, at any one time.
8.14 Determination of eigenvalues and eigenvectors
The next step is to show how the eigenvalues and eigenvectors of a given N×N
matrix Aare found. To do this we refer to (8.68) and as in (8.69) rewrite it as
Ax−λIx=(A−λI)x=0. (8.85)
The slight rearrangement used here is to write xasIx,w h e r e Iis the unit matrix
of order N. The point of doing this is immediate since (8.85) now has the form
of a homogeneous set of simultaneous equations, the theory of which will bedeveloped in section 8.18. What will be proved there is that the equation Bx=0
only has a non-trivial solution xif|B|= 0. Correspondingly, therefore, we must
have in the present case that
|A−λI|=0, (8.86)
if there are to be non-zero solutions xto (8.85).
Equation (8.86) is known as the characteristic equation forAa n di t sL H Sa s
thecharacteristic orsecular determinant ofA. The equation is a polynomial of
degree Nin the quantity λ.T h e Nroots of this equation λ
i,i=1,2,...,N , give
the eigenvalues of A. Corresponding to each λithere will be a column vector xi,
which is the ith eigenvector of Aand can be found by using (8.68).
It will be observed that when (8.86) is written out as a polynomial equation in
λ, the coefficient of −λN−1in the equation will be simply A11+A22+···+ANN
relative to the coefficient of λN. As discussed in section 8.8, the quantitysummationtextN
i=1Aii
is the traceofAand, from the ordinary theory of polynomial equations, will be
equal to the sum of the roots of (8.86):
Nsummationdisplay
i=1λi=T r A. (8.87)
This can be used as one check that a computation of the eigenvalues λihas been
done correctly. Unless equation (8.87) is satisfied by a computed set of eigenvalues,
they have not been calculated correctly. However, that equation (8.87) is satisfied is
a necessary, but not sufficient, condition for a correct computation. An alternativeproof of (8.87) is given in section 8.16.
285
MATRICES AND VECTOR SPACESIFind the eigenvalues and normalised eige nvectors of the real symmetric matrix
A=
/0/@11 3
11−3
3−3−3
/1A.
Using (8.86),//////1−λ13
11−λ−3
3−3−3−λ
//////=0.
Expanding out this determinant gives
(1−λ)[(1−λ)(−3−λ)−(−3)(−3)]+1[(−3)(3)−1(−3−λ)]
+3[1(−3)−(1−λ)(3)]=0,
which simplifies to give
(1−λ)(λ2+2λ−12) + ( λ−6) + 3(3 λ−6) = 0 ,
⇒ (λ−2)(λ−3)(λ+6 )=0 .
Hence the roots of the characteristic equation, which are the eigenvalues of A,a r e λ1=2 ,
λ2=3 , λ3=−6. We note that, as expected,
λ1+λ2+λ3=−1=1+1−3=A11+A22+A33=T r A.
For the first root, λ1= 2, a suitable eigenvector x1, with elements x1,x2,x3,m u s ts a t i s f y
Ax1=2x1or, equivalently,
x1+x2+3x3=2x1,
x1+x2−3x3=2x2, (8.88)
3x1−3x2−3x3=2x3.
These three equations are consistent (to ensure this was the purpose in finding the particular
values of λ)a n dy i e l d x3=0 , x1=x2=k,w h e r e kis any non-zero number. A suitable
eigenvector would thus be
x1=(kk 0)T.
If we apply the normalisation condition, we require k2+k2+02=1o r k=1/√
2. Hence
x1=
/1√
21√
20
/T
=1√
2(110 )T.
Repeating the last paragraph, but with the factor 2 on the RHS of (8.88) replaced
successively by λ2=3a n d λ3=−6, gives two further normalised eigenvectors
x2=1√
3(1−11)T, x3=1√
6(1−1−2)T.
J
In the above example, the three values of λare all different and Ais a
real symmetric matrix. Thus we expect, and it is easily checked, that the threeeigenvectors are mutually orthogonal, i.e.
parenleftbig
x
1parenrightbigTx2=parenleftbig
x1parenrightbigTx3=parenleftbig
x2parenrightbigTx3=0.
It will be apparent also that, as expected, the normalisation of the eigenvectors
has no effect on their orthogonality.
286
8.14 DETERMINATION OF EIGENVALUES AND EIGENVECTORS
8.14.1 Degenerate eigenvalues
We return now to the case of degenerate eigenvalues, i.e. those that have two or
more associated eigenvectors. We have shown already that it is always possibleto construct an orthogonal set of eigenvectors for a normal matrix, see subsec-tion 8.13.1, and the following example illustrates one method for constructingsuch a set.IConstruct an orthonormal set of eigenvectors for the matrix
A=
/0/@103
0−20
301
/1A.
We first determine the eigenvalues using |A−λI|=0 :
0=
//////1−λ 03
0−2−λ0
30 1 −λ
//////=−(1−λ)2(2 +λ) + 3(3)(2 + λ)
=( 4−λ)(λ+2 )2.
Thus λ1=4 , λ2=−2=λ3. The eigenvector x1=(x1x2x3)Tis found from/0/@103
0−20
301
/1A
/0/@x1
x2
x3
/1
A
=4
/0
/@
x1
x2
x3
/1
A
⇒ x1=1√
2
/0/@1
01
/1A.
A general column vector that is orthogonal to x1is
x=(ab−a)T, (8.89)
and it is easily shown that
Ax=
/0/@103
0−20
301
/1A
/0/@a
b
−a
/1A=−2
/0/@a
b
−a
/1A=−2x.
Thus xis a eigenvector of Awith associated eigenvalue −2. It is clear, however, that there
is an infinite set of eigenvectors xall possessing the required property; the geometrical
analogue is that there are an infinite number of corresponding vectors xlying in the
plane that has x1as its normal. We do require that the two remaining eigenvectors are
orthogonal to one another, but this still le aves an infinite number of possibilities. For x2,
therefore, let us choose a simple form of (8.89), suitably normalised, say,
x2=(010 )T.
The third eigenvector is then specified (to within an arbitrary multiplicative constant)
by the requirement that it must be orthogonal to x1and x2; thus x3may be found by
evaluating the vector product of x1and x2and normalising the result. This gives
x3=1√
2(−101 )T,
to complete the construction of an orthonormal set of eigenvectors.
J
287
MATRICES AND VECTOR SPACES
8.15 Change of basis and similarity transformations
Throughout this chapter we have considered the vector xas a geometrical quantity
that is independent of any basis (or coordinate system). If we introduce a basise
i,i=1,2,...,N , into our N-dimensional vector space then we may write
x=x1e1+x2e2+···+xNeN,
and represent xin this basis by the column matrix
x=(x1x2···xn)T,
having components xi. We now consider how these components change as a result
of a prescribed change of basis. Let us introduce a new basis e/prime
i,i=1,2,...,N ,
which is related to the old basis by
e/prime
j=Nsummationdisplay
i=1Sijei, (8.90)
the coefficient Sijbeing the ith component of e/prime
jwith respect to the old (unprimed)
basis. For an arbitrary vector xit follows that
x=Nsummationdisplay
i=1xiei=Nsummationdisplay
j=1x/prime
je/primej=Nsummationdisplay
j=1x/prime
jNsummationdisplay
i=1Sijei.
From this we derive the relationship between the components of xin the two
coordinate systems as
xi=Nsummationdisplay
j=1Sijx/prime
j,
w h i c hw ec a nw r i t ei nm a t r i xf o r ma s
x=Sx/prime(8.91)
where Sis the transformation matrix associated with the change of basis.
Furthermore, since the vectors e/prime
jare linearly independent, the matrix Sis
non-singular and so possesses an inverse S−1. Multiplying (8.91) on the left by
S−1we find
x/prime=S−1x, (8.92)
which relates the components of xin the new basis to those in the old basis.
Comparing (8.92) and (8.90) we note that the components of xtransform inversely
to the way in which the basis vectors eithemselves transform. This has to be so,
as the vector xitself must remain unchanged.
We may also find the transformation law for the components of a linear
operator under the same change of basis. Now, the operator equation y=Ax
288
8.15 CHANGE OF BASIS AND SIMILARITY TRANSFORMATIONS
(which is basis independent) can be written as a matrix equation in each of the
two bases as
y=Ax, y/prime=A/primex/prime. (8.93)
But, using (8.91), we may rewrite the first equation as
Sy/prime=ASx/prime⇒ y/prime=S−1ASx/prime.
Comparing this with the second equation in (8.93) we find that the components
of the linear operator Atransform as
A/prime=S−1AS. (8.94)
Equation (8.94) is an example of a similarity transformation – a transformation
that can be particularly useful in converting matrices into convenient forms forcomputation.
Given a square matrix A, we may interpret it as representing a linear operator
Ain a given basis e
i. From (8.94), however, we may also consider the matrix
A/prime=S−1AS, for any non-singular matrix S, as representing the same linear
operator Abut in a new basis e/prime
j, related to the old basis by
e/prime
j=summationdisplay
iSijei.
Therefore we would expect that any property of the matrix Athat represents
some (basis-independent) property of the linear operator Awill also be shared
by the matrix A/prime. We list these properties below.
(i) If A=Ithen A/prime=I, since, from (8.94),
A/prime=S−1IS=S−1S=I. (8.95)
(ii) The value of the determinant is unchanged:
|A/prime|=|S−1AS|=|S−1||A||S|=|A||S−1||S|=|A||S−1S|=|A|.(8.96)
(iii) The characteristic determinant and hence the eigenvalues of A/primeare the
same as those of A: from (8.86),
|A/prime−λI|=|S−1AS−λI|=|S−1(A−λI)S|
=|S−1||S||A−λI|=|A−λI|. (8.97)
(iv) The value of the trace is unchanged: from (8.87),
TrA/prime=summationdisplay
iA/prime
ii=summationdisplay
isummationdisplay
jsummationdisplay
k(S−1)ijAjkSki
=summationdisplay
isummationdisplay
jsummationdisplay
kSki(S−1)ijAjk=summationdisplay
jsummationdisplay
kδkjAjk=summationdisplay
jAjj
=T r A. (8.98)
289
MATRICES AND VECTOR SPACES
An important class of similarity transformations is that for which Sis a uni-
tary matrix; in this case A/prime=S−1AS=S†AS. Unitary transformation matrices
are particularly important, for the following reason. If the original basis eiis
orthonormal and the transformation matrix Sis unitary then
/angbracketlefte/prime
i|e/prime
j/angbracketright=angbracketleftBigsummationdisplay
kSkiekvextendsinglevextendsinglevextendsinglesummationdisplay
rSrjerangbracketrightBig
=summationdisplay
kS∗
kisummationdisplay
rSrj/angbracketleftek|er/angbracketright
=summationdisplay
kS∗
kisummationdisplay
rSrjδkr=summationdisplay
kS∗
kiSkj=(S†S)ij=δij,
showing that the new basis is also orthonormal.
Furthermore, in addition to the properties of general similarity transformations,
for unitary transformations the following hold.
(i) If Ais Hermitian (anti-Hermitian) then A/primeis Hermitian (anti-Hermitian),
i.e. if A†=±Athen
(A/prime)†=(S†AS)†=S†A†S=±S†AS=±A/prime. (8.99)
(ii) If Ais unitary (so that A†=A−1)t h e n A/primeis unitary, since
(A/prime)†A/prime=(S†AS)†(S†AS)=S†A†SS†AS=S†A†AS
=S†IS=I. (8.100)
8.16 Diagonalisation of matrices
Suppose that a linear operator Ais represented in some basis ei,i=1,2,...,N ,
by the matrix A. Consider a new basis xjgiven by
xj=Nsummationdisplay
i=1Sijei,
where the xjare chosen to be the eigenvectors of the linear operator A,i . e .
Axj=λjxj. (8.101)
In the new basis, Ais represented by the matrix A/prime=S−1AS, which has a
particularly simple form, as we shall see shortly. The element SijofSis the ith
component, in the old (unprimed) basis, of the jth eigenvector xjofA,i . e .t h e
columns of Sare the eigenvectors of the matrix A:
S=
↑↑ ↑
x1x2··· xN
↓↓ ↓
,
290
8.16 DIAGONALISATION OF MATRICES
That is Sij=(xj)i. Therefore A/primeis given by
(S−1AS)ij=summationdisplay
ksummationdisplay
l(S−1)ikAklSlj
=summationdisplay
ksummationdisplay
l(S−1)ikAkl(xj)l
=summationdisplay
k(S−1)ikλj(xj)k
=summationdisplay
kλj(S−1)ikSkj=λjδij.
So the matrix A/primeis diagonal with the eigenvalues of Aas the diagonal elements,
i.e.
A/prime=
λ10··· 0
0λ2...
......0
0··· 0λN
.
Therefore, given a matrix A,i fw ec o n s t r u c tt h em a t r i x Sthat has the eigen-
vectors of Aas its columns then the matrix A/prime=S−1ASis diagonal and has the
eigenvalues of Aas its diagonal elements. Since we require Sto be non-singular
(|S|/negationslash= 0), the Neigenvectors of Amust be linearly independent and form a basis
for the N-dimensional vector space. It may be shown that any matrix with distinct
eigenvalues can be diagonalised by this procedure. If, however, a general square
matrix has degenerate eigenvalues then it may, or may not, have Nlinearly
independent eigenvectors. If it does not then it cannot be diagonalised.
For normal matrices (which include Hermitian, anti-Hermitian and unitary
matrices) the Neigenvectors are indeed linearly independent. Moreover, when
normalised, these eigenvectors form an orthonormal s e t( o rc a nb em a d et od o
so). Therefore the matrix Swith these normalised eigenvectors as columns, i.e.
whose elements are Sij=(xj)i,h a st h ep r o p e r t y
(S†S)ij=summationdisplay
k(S†)ik(S)kj=summationdisplay
kS∗
kiSkj=summationdisplay
k(xi)∗
k(xj)k=(xi)†xj=δij.
Hence Sis unitary ( S−1=S†) and the original matrix Acan be diagonalised by
A/prime=S−1AS=S†AS.
Therefore, any normal matrix Acan be diagonalised by a similarity transformation
using a unitary transformation matrix S.
291
MATRICES AND VECTOR SPACESIDiagonalise the matrix
A=
/0/@103
0−20
301
/1A.
The matrix Ais symmetric and so may be diagonalised by a transformation of the form
A/prime=S†AS,w h e r e Shas the normalised eigenvectors of Aas its columns. We have already
found these eigenvectors in subsection 8. 14.1, and so we can write straightaway
S=1√
2
/0/@10−1
0√
20
10 1
/1A.
We note that although the eigenvalues of Aare degenerate, its three eigenvectors are
linearly independent and so Acan still be diagonalised. Thus, calculating S†ASwe obtain
S†AS=1
2
/0/@10 1
0√
20
−101
/1A
/0/@103
0−20
301
/1A
/0/@10−1
0√
20
10 1
/1A
=
/0/@40 0
0−20
00−2
/1A,
which is diagonal, as required, and has as its diagonal elements the eigenvalues of A.
J
If a matrix Ais diagonalised by the similarity transformation A/prime=S−1AS,s o
that A/prime=d i a g ( λ1,λ2...,λ N), then we have immediately
TrA/prime=T r A=Nsummationdisplay
i=1λi, (8.102)
|A/prime|=|A|=Nproductdisplay
i=1λi, (8.103)
since the eigenvalues of the matrix are unchanged by the transformation. More-
over, these results may be used to prove the rather useful trace formula
|exp A|=e x p ( T r A), (8.104)
where the exponential of a matrix is as defined in (8.38).IProve the trace formula (8.104).
At the outset, we note that for the similarity transformation A/prime=S−1AS, we have
(A/prime)n=(S−1AS)(S−1AS)···(S−1AS)=S−1AnS.
Thus, from (8.38), we obtain exp A/prime=S−1(exp A)S, from which it follows that |exp A/prime|=
292
8.17 QUADRATIC AND HERMITIAN FORMS
|exp A|. Moreover, by choosing the similarity transformation so that it diagonalises A,w e
have A/prime= diag( λ1,λ2,...,λ N), and so
|expA|=|expA/prime|=|exp[diag( λ1,λ2,...,λ N)]|=|diag(exp λ1,expλ2,...,expλN)|=NY
i=1expλi.
Rewriting the final product of exponentials of the eigenvalues as the exponential of the
sum of the eigenvalues, we find
|exp A|=NY
i=1expλi=e x p
/ NX
i=1λi
/!
=e x p ( T r A),
which gives the trace formula (8.104).
J
8.17 Quadratic and Hermitian forms
Let us now introduce the concept of quadratic forms (and their complex ana-
logues, Hermitian forms). A quadratic form Qis a scalar function of a real vector
xgiven by
Q(x)=/angbracketleftx|Ax/angbracketright, (8.105)
for some real linear operator A. In any given basis (coordinate system) we can
write (8.105) in matrix form as
Q(x)=xTAx, (8.106)
where Ais a real matrix. In fact, as will be explained below, we need only consider
t h ec a s ew h e r e Ais symmetric, i.e. A=AT. As an example in a three-dimensional
space,
Q=xTAx=parenleftBig
x1x2x3parenrightBig
11 3
11−3
3−3−3
x1
x2
x3
=x2
1+x2
2−3x2
3+2x1x2+6x1x3−6x2x3. (8.107)
It is reasonable to ask whether a quadratic form Q=xTMx,w h e r e Mis any
(possibly non-symmetric) real square matrix, is a more general definition. Thatthis not the case may be seen by expressing Min terms of a symmetric matrix
A=
1
2(M+MT) and an antisymmetric matrix B=1
2(M−MT) such that M=A+B.
We then have
Q=xTMx=xTAx+xTBx. (8.108)
However, Qis a scalar quantity and so
Q=QT=(xTAx)T+(xTBx)T=xTATx+xTBTx=xTAx−xTBx.
(8.109)
293
MATRICES AND VECTOR SPACES
Comparing (8.108) and (8.109) shows that xTBx= 0, and hence xTMx=xTAx,
i.e.Qis unchanged by considering only the symmetric part of M. Hence, with no
loss of generality, we may assume A=ATin (8.106).
From its definition (8.105), Qis clearly a basis- (i.e. coordinate-) independent
quantity. Let us therefore consider a new basis related to the old one by an
orthogonal transformation matrix S, the components in the two bases of any
vector xbeing related (as in (8.91)) by x=Sx/primeor, equivalently, by x/prime=S−1x=
STx. We then have
Q=xTAx=(x/prime)TSTASx/prime=(x/prime)TA/primex/prime,
where (as expected) the matrix describing the linear operator Ain the new
basis is given by A/prime=STAS(since ST=S−1). But, from the last section, if we
choose as Sthe matrix whose columns are the normalised eigenvectors of Athen
A/prime=STASis diagonal with the eigenvalues of Aas the diagonal elements. (Since
Ais symmetric, its normalised eigenvectors are orthogonal, or can be made so,
and hence Sis orthogonal with S−1=ST.)
In the new basis
Q=xTAx=(x/prime)TΛx/prime=λ1x/prime
12+λ2x/prime
22+···+λNx/prime
N2, (8.110)
where Λ = diag( λ1,λ2,...,λ N)a n dt h e λiare the eigenvalues of A. It should be
noted that Qcontains no cross-terms of the form x/prime
1x/prime2.IFind an orthogonal transformation that takes the quadratic form (8.107) into the form
λ1x/prime
12+λ2x/prime
22+λ3x/prime
32.
The required transformation matrix Shas the normalised eigenvectors of Aas its columns.
We have already found these in section 8.14, and so we can write immediately
S=1√
6
/0/@√
3√
21√
3−√
2−1
0√
2−2
/1A,
which is easily verified as being orthogonal. Since the eigenvalues of Aareλ=2 ,3 ,a n d
−6, the general result already proved shows that the transformation x=Sx/primewill carry
(8.107) into the form 2 x/prime
12+3x/prime
22−6x/prime
32.This may be verified most easily by writing out
the inverse transformation x/prime=S−1x=STxand substituting. The inverse equations are
x/prime
1=(x1+x2)/√
2,
x/prime
2=(x1−x2+x3)/√
3, (8.111)
x/prime
3=(x1−x2−2x3)/√
6.
If these are substituted into the form Q=2x/prime
12+3x/prime
22−6x/prime
32then the original expression
(8.107) is recovered.
J
In the definition of Qit was assumed that the components x1,x2,x3and the
matrix Awere real. It is clear that in this case the quadratic form Q≡xTAxis real
294
8.17 QUADRATIC AND HERMITIAN FORMS
also. Another, rather more general, expression that is also real is the Hermitian
form
H(x)≡x†Ax, (8.112)
where Ais Hermitian (i.e. A†=A) and the components of xmay now be complex.
It is straightforward to show that His real, since
H∗=(HT)∗=x†A†x=x†Ax=H.
With suitable generalisation, the properties of quadratic forms apply also to Her-
mitian forms, but to keep the presentation simple we will restrict our discussionto quadratic forms.
A special case of a quadratic (Hermitian) form is one for which Q=x
TAx
is greater than zero for all column matrices x. By choosing as the basis the
eigenvectors of Awe have Qin the form
Q=λ1x2
1+λ2x2
2+λ3x2
3.
The requirement that Q>0f o ra l l xmeans that all the eigenvalues λiofAmust
be positive. A symmetric (Hermitian) matrix Awith this property is called positive
definite .I f ,i n s t e a d , Q≥0f o ra l l xthen it is possible that some of the eigenvalues
are zero, and Ais called positive semi-definite .
8.17.1 The stationary properties of the eigenvectors
Consider a quadratic form, such as Q(x)=/angbracketleftx|Ax/angbracketrightgiven in (8.105), in a fixed
basis. As the vector xis varied, through changes in its three components x1,x2
andx3, the value of the quantity Qalso varies. Because of the homogeneous
form of Qwe may restrict any investigation of these variations to vectors of unit
length (since multiplying any vector xby any scalar ksimply multiplies the value
ofQby a factor k2).
Of particular interest are any vectors xthat make the value of the quadratic
form a maximum or minimum. A necessary, but not sufficient, condition for this
is that Qis stationary with respect to small variations ∆ xinx, whilst/angbracketleftx|x/angbracketrightis
maintained at a constant value (unity).
In the chosen basis the quadratic form is given by Q=xTAxand, using
Lagrange undetermined multipliers to incorporate the variational constraints, weare led to seek solutions of
∆[x
TAx−λ(xTx−1)] = 0 . (8.113)
This may be used directly, together with the fact that (∆ xT)Ax=xTA∆x,s i n c e A
is symmetric, to obtain
Ax=λx, (8.114)
295
MATRICES AND VECTOR SPACES
as the necessary condition that xmust satisfy. If (8.114) is satisfied for some
eigenvector xthen the value of Q(x)i sg i v e nb y
Q=xTAx=xTλx=λ. (8.115)
However, if xand yare eigenvectors corresponding to different eigenvalues then
they are (or can be chosen to be) orthogonal. Consequently the expression yTAx
is necessarily zero, since
yTAx=yTλx=λyTx=0. (8.116)
Summarising, those column matrices xof unit magnitude that make the
quadratic form Qstationary are eigenvectors of the matrix A, and the stationary
value of Qis then equal to the corresponding eigenvalue. It is straightforward
to see from the proof of (8.114) that, conversely, any eigenvector of Amakes Q
stationary.
Instead of maximising or minimising Q=xTAxsubject to the constraint
xTx= 1, an equivalent procedure is to extremise the function
λ(x)=xTAx
xTx.IShow that if λ(x)is stationary then xis an eigenvector of Aandλ(x)is equal to the
corresponding eigenvalue.
We require ∆ λ(x) = 0 with respect to small variations in x.N o w
∆λ=1
(xTx)2
/
(xTx)
/;
∆xTAx+xTA∆x
/
−xTAx
/;
∆xTx+xT∆x
//
=2∆xTAx
xTx−2
/xTAx
xTx
/∆xTx
xTx,
since xTA∆x=( ∆ xT)Axand xT∆x=( ∆ xT)x. Thus
∆λ=2
xTx∆xT[Ax−λ(x)x].
Hence, if ∆ λ=0t h e n Ax=λ(x)x,i . e . xis an eigenvector of Awith eigenvalue λ(x).
J
Thus the eigenvalues of a symmetric matrix Aare the values of the function
λ(x)=xTAx
xTx
at its stationary points. The eigenvectors of Alie along those directions in space
for which the quadratic form Q=xTAxhas stationary values, given a fixed
magnitude for the vector x. Similar results hold for Hermitian matrices.
296
8.18 SIMULTANEOUS LINEAR EQUATIONS
8.17.2 Quadratic surfaces
The results of the previous subsection may be turned round to state that the
surface given by
xTAx= constant = 1 (say) (8.117)
and called a quadratic surface , has stationary values of its radius (i.e. origin–
surface distance) in those directions that are along the eigenvectors of A.M o r e
specifically, in three dimensions the quadratic surface xTAx= 1 has its principal
axes along the three mutually perpendicular eigenvectors of A,a n dt h es q u a r e s
of the corresponding principal radii are given by λ−1
i,i=1,2,3. As well as
having this stationary property of the radius, a principal axis is characterised bythe fact that any section of the surface perpendicular to it has some degree ofsymmetry about it. If the eigenvalues corresponding to any two principal axes aredegenerate then the quadratic surface has rotational symmetry about the thirdprincipal axis and the choice of a pair of axes perpendicular to that axis is notuniquely defined.IFind the shape of the quadratic surface
x2
1+x2
2−3x2
3+2x1x2+6x1x3−6x2x3=1.
If, instead of expressing the quadratic surface in terms of x1,x2,x3, as in (8.107), we
were to use the new variables x/prime
1,x/prime
2,x/prime
3defined in (8.111), for which the coordinate axes
are along the three mutually perpendicular eigenvector directions (1 ,1,0), (1 ,−1,1) and
(1,−1,−2), then the equation of the surface would take the form (see (8.110))
x/prime
12
(1/√
2)2+x/prime
22
(1/√
3)2−x/prime
32
(1/√
6)2=1.
Thus, for example, a section of the quadratic surface in the plane x/prime
3=0 ,i . e . x1−x2−
2x3= 0, is an ellipse, with semi-axes 1 /√
2a n d1 /√
3. Similarly a section in the plane
x/prime
1=x1+x2= 0 is a hyperbola.
J
Clearly the simplest three-dimensional situation to visualise is that in which all
the eigenvalues are positive, since then the quadratic surface is an ellipsoid.
8.18 Simultaneous linear equations
In physical applications we often encounter sets of simultaneous linear equations.
In general we may have Mequations in Nunknowns x1,x2,...,x Nof the form
A11x1+A12x2+···+A1NxN=b1,
A21x1+A22x2+···+A2NxN=b2,
...
AM1x1+AM2x2+···+AMNxN=bM,(8.118)
297
MATRICES AND VECTOR SPACES
where the Aijandbihave known values. If all the biare zero then the system of
equations is called homogeneous , otherwise it is inhomogeneous . Depending on the
given values, this set of equations for the Nunknowns x1,x2,...,xNmay have
either a unique solution, no solution or infinitely many solutions. Matrix analysis
may be used to distinguish between the possibilities. The set of equations may be
expressed as a single matrix equation Ax=b, or, written out in full, as
A11 A12... A 1N
A21 A22... A 2N
............
AM1AM2... A MN
x1
x2
...
xN
=
b
1
b2
...
bM
.
8.18.1 The range and null space of a matrix
As we discussed in section 8.2, we may interpret the matrix equation Ax=bas
representing, in some basis, the linear transformation Ax=bof a vector xin an
N-dimensional vector space Vinto a vector bin some other (in general different)
M-dimensional vector space W.
In general the operator Awill map anyvector in Vinto some particular
subspace ofW, which may be the entire space. This subspace is called the range
ofA(orA) and its dimension is equal to the rankofA.M o r e o v e r ,i f A(and
hence A)i ssingular then there exists some subspace of Vthat is mapped onto
the zero vector 0inW; that is, any vector ythat lies in the subspace satisfies
Ay=0. This subspace is called the null space ofAand the dimension of this
null space is called the nullity ofA. We note that the matrix Amustbe singular
ifM/negationslash=Nandmaybe singular even if M=N.
The dimensions of the range and the null space of a matrix are related through
the fundamental relationship
rank A+ nullity A=N, (8.119)
where Nis the number of original unknowns x
1,x2,...,x N.IProve the relationship (8.119).
As discussed in section 8.11, if the columns of an M×Nmatrix Aare interpreted as the
components, in a given basis, of N(M-component) vectors v1,v2,...,vNthen rank Ais
equal to the number of linearly independent vectors in this set (this number is also equalto the dimension of the vector space spanned by these vectors). Writing (8.118) in termsof the vectors v
1,v2,...,vN, we have
x1v1+x2v2+···+xNvN=b. (8.120)
From this expression, we immediately deduce that the range of Ais merely the span of
the vectors v1,v2,...,vNand hence has dimension r=r a n k A.
298
8.18 SIMULTANEOUS LINEAR EQUATIONS
If a vector ylies in the null space of AthenAy=0,w h i c hw em a yw r i t ea s
y1v1+y2v2+···+yNvN=0. (8.121)
As just shown above, however, only r(≤N) of these vectors are linearly independent. By
renumbering, if necessary, we may assume that v1,v2,...,vrform a linearly independent
set; the remaining vectors, vr+1,vr+2,...,vN, can then be written as a linear superposition
ofv1,v2,...,vr.W ea r et h e r e f o r ef r e et oc h o o s et h e N−rcoefficients yr+1,yr+2,...,y N
arbitrarily and (8.121) will still be satisfied for some set of rcoefficients y1,y2,...,y r(which
are not all zero). The dimension of the null space is therefore N−r, and this completes
the proof of (8.119).
J
Equation (8.119) has far-reaching consequences for the existence of solutions
to sets of simultaneous linear equations such as (8.118). As mentioned previously,these equations may have no solution ,aunique solution orinfinitely many solutions .
We now discuss these three cases in turn.
No solution
The system of equations possesses no solution unless blies in the range of A;i n
this case (8.120) will be satisfied for some x
1,x2,...,x N. This in turn requires the
s e to fv e c t o r s b,v1,v2,...,vNto have the same span (see (8.8)) as v1,v2,...,vN.I n
terms of matrices, this is equivalent to the requirement that the matrix Aand the
augmented matrix
M=
A
11 A12... A 1Nb1
A21 A22... A 2Nb1
.........
AM1AM2... A MN bM
have the samerank r. If this condition is satisfied then bdoes lie in the range of
A, and the set of equations (8.118) will have either a unique solution or infinitely
many solutions. If, however, Aand Mhave different ranks then there will be no
solution.
A unique solution
Ifblies in the range of Aand if r=Nthen all the vectors v
1,v2,...,vNin (8.120)
are linearly independent and the equation has a unique solution x1,x2,...,x N.
Infinitely many solutions
Ifblies in the range of Aand if r<N then only rof the vectors v1,v2,...,vN
in (8.120) are linearly independent. We may therefore choose the coefficients of
n−rvectors in an arbitrary way, while still satisfying (8.120) for some set of
coefficients x1,x2,...,x N. There are therefore infinitely many solutions ,w h i c hs p a n
an (n−r) dimensional vector space. We may also consider this space of solutions
in terms of the null space of A:i fxis some vector satisfying Ax=bandyis
299
MATRICES AND VECTOR SPACES
anyvector in the null space of A(i.e.Ay=0)t h e n
A(x+y)=Ax+Ay=Ax+0=b,
and so x+yis also a solution. Since the null space is ( n−r)-dimensional, so too
is the space of solutions.
We may use the above results to investigate the special case of the solution of
ahomogeneous set of linear equations, for which b=0. Clearly the set always has
the trivial solution x1=x2=···=xn=0 ,a n di f r=Nthis will be the only
solution. If r<N , however, there are infinitely many solutions; they form the
null space of A, which has dimension n−r. In particular, we note that if M<N
(i.e. there are fewer equations than unknowns) then r<N automatically. Hence a
set of homogeneous linear equations with fewer equations than unknowns always
has infinitely many solutions.
8.18.2 Nsimultaneous linear equations in Nunknowns
A special case of (8.118) occurs when M=N. In this case the matrix Aissquare
and we have the same number of equations as unknowns. Since Ais square, the
condition r=Ncorresponds to |A|/negationslash= 0 and the matrix Aisnon-singular .T h e
caser<N corresponds to |A|=0 ,i nw h i c hc a s e Aissingular .
As mentioned above, the equations will have a solution provided blies in the
range of A. If this is true then the equations will possess a unique solution when
|A|/negationslash= 0 or infinitely many solutions when |A|= 0. There exist several methods
for obtaining the solution(s). Perhaps the most elementary method is Gaussian
elimination ; this method is discussed in subsection 28.3.1, where we also address
numerical subtleties such as equation interchange (pivoting). In this subsection,we will outline three further methods for solving a square set of simultaneouslinear equations.
Direct inversion
Since Ais square it will possess an inverse, provided |A|/negationslash= 0. Thus, if Ais
non-singular, we immediately obtain
x=A
−1b (8.122)
as the unique solution to the set of equations. However, if b=0,t h e nw es e e
immediately that the set of equations possesses only the trivial solution x=0.T h e
direct inversion method has the advantage that, once A−1has been calculated,
one may obtain the solutions xcorresponding to different vectors b1,b2, ...on
the RHS, with little further work.
300
8.18 SIMULTANEOUS LINEAR EQUATIONSIShow that the set of simultaneous equations
2x1+4x2+3x3=4,
x1−2x2−2x3=0, (8.123)
−3x1+3x2+2x3=−7,
has a unique solution, and find that solution.
The simultaneous equations can be represented by the matrix equation Ax=b,i . e ./0/@243
1−2−2
−33 2
/1A
/0/@x1
x2
x3
/1
A
=
/0
/@
4
0
−7
/1A.
As we have already shown that A−1exists and have calculated it, see (8.59), it follows that
x=A−1bor, more explicitly, that/0/@x1
x2
x3
/1
A
=1
11
/0/@21−2
41 37
−3−18−8
/1A
/0/@4
0
−7
/1A=
/0
/@
2
−3
4
/1A. (8.124)
Thus the unique solution is x1=2 , x2=−3,x3=4 .
J
LU decomposition
Although conceptually simple, finding the solution by calculating A−1can be
computationally demanding, especially when Nis large. In fact, as we shall now
show, it is not necessary to perform the full inversion of Ain order to solve the
simultaneous equations Ax=b. Rather, we can perform a decomposition of the
matrix into the product of a square lower triangular matrix Land a square upper
triangular matrix U, which are such that
A=LU, (8.125)
and then use the fact that triangular systems of equations can be solved very
simply.
We must begin, therefore, by finding the matrices Land Usuch that (8.125)
is satisfied. This may be achieved straightforwardly by writing out (8.125) incomponent form. For illustration, let us consider the 3 ×3 case. It is, in fact,
always possible, and convenient, to take the diagonal elements of Las unity, so
we have
A=
10 0
L
2110
L31L321
U11U12U13
0U22U23
00 U33
=
U11 U12 U13
L21U11 L21U12+U22 L21U13+U23
L31U11L31U12+L32U22L31U13+L32U23+U33
(8.126)
The nine unknown elements of Land Ucan now be determined by equating
301
MATRICES AND VECTOR SPACES
the nine elements of (8.126) to those of the 3 ×3m a t r i x A. This is done in the
particular order illustrated in the example below.
Once the matrices Land Uhave been determined, one can use the decomposition
to solve the set of equations Ax=bin the following way. From (8.125), we have
LUx=b,
but this can be written as twotriangular sets of equations
Ly=b and Ux=y,
where yis another column matrix to be determined. One may easily solve the first
triangular set of equations for y, which is then substituted into the second set.
The required solution xis then obtained readily from the second triangular set
of equations. We note that, as with direct inversion, once the LUdecomposition
has been determined, one can solve for various RHS column matrices b1,b2,...,
with little extra work.IUseLUdecomposition to solve the set of simultaneous equations (8.123).
We begin the determination of the matrices Land Uby equating the elements of the
matrix in (8.126) with those of the matrix
A=
/0/@243
1−2−2
−33 2
/1A.
This is performed in the following order:
1st row: U11=2 , U12=4 , U13=3
1st column: L21U11=1 , L31U11=−3⇒L21=1
2,L31=−3
2
2nd row: L21U12+U22=−2 L21U13+U23=−2⇒U22=−4,U23=−7
2
2nd column: L31U12+L32U22=3 ⇒L32=−9
4
3rd row: L31U13+L32U23+U33=2 ⇒U33=−11
8
Thus we may write the matrix Aas
A=LU=
/0/@10 0
1
210
−3
2−9
41
/1A
/0/@24 3
0−4−7
2
00−11
8
/1A.
We must now solve the set of equations Ly=b,w h i c hr e a d/0/@10 0
1
210
−3
2−9
41
/1A
/0/@y1
y2
y3
/1
A
=
/0
/@
4
0
−7
/1A.
Since this set of equations is triangular, we quickly find
y1=4,y 2=0−(1
2)(4) =−2,y 3=−7−(−3
2)(4)−(−9
4)(−2) =−11
2.
These values must then be substituted into the equations Ux=y,w h i c hr e a d/0/@24 3
0−4−7
2
00−11
8
/1A
/0/@x1
x2
x3
/1
A
=
/0
/@
4
−2
−11
2
/1A.
302
8.18 SIMULTANEOUS LINEAR EQUATIONS
This set of equations is also triangular, and we easily find the solution
x1=2,x 2=−3,x 3=4,
which agrees with the result found above by direct inversion.
J
We note, in passing, that one can calculate both the inverse and the determinant
ofAfrom its LUdecomposition. To find the inverse A−1, one solves the system
of equations Ax=brepeatedly for the Ndifferent RHS column matrices b=ei
(i=1,2,...,N ), where eiis the column matrix with its ith element equal to unity
and the others equal to zero. The solution xin each case gives the corresponding
column of A−1. Evaluation of the determinant |A|is much simpler. From (8.125),
we have
|A|=|LU|=|L||U|. (8.127)
Since Land Uare triangular, however, we see from (8.64) that their determinants
are equal to the products of their diagonal elements. Since Lii=1f o ra l l i,w e
thus find
|A|=U11U22···UNN=Nproductdisplay
i=1Uii.
As an illustration, in the above example we find |A|=( 2 ) (−4)(−11/8) = 11,
which, as it must, agrees with our earlier calculation (8.58).
Finally, we note that if the matrix Ais symmetric and positive semi-definite
then we can decompose it as
A=LL†, (8.128)
where Lis a lower triangular matrix whose diagonal elements are not,i ng e n e r a l ,
equal to unity. This is known as a Cholesky decomposition (in the special case
where Ais real, the decomposition becomes A=LLT). The reason that we cannot
set the diagonal elements of Lequal to unity in this case is that we require the
same number of independent elements in Las in A. The requirement that the
matrix be positive semi-definite is easily derived by considering the Hermitian
form (or quadratic form in the real case)
x†Ax=x†LL†x=(L†x)†(L†x).
Denoting the column matrix L†xbyy, we see that the last term on the RHS
isy†y, which must be greater than or equal to zero. Thus, we require x†Ax≥0
for any arbitrary column matrix x,a n ds o Amust be positive semi-definite (see
section 8.17).
We recall that the requirement that a matrix be positive semi-definite is equiv-
alent to demanding that all the eigenvalues of Aare positive or zero. If one of
the eigenvalues of Ais zero, however, then from (8.103) we have |A|=0a n ds o A
issingular . Thus, if Ais a non-singular matrix, it must be positive definite (rather
303
MATRICES AND VECTOR SPACES
than just positive semi-definite) in order to perform the Cholesky decomposition
(8.128). In fact, in this case, the inability to find a matrix Lthat satisfies (8.128)
implies that Acannot be positive definite.
The Cholesky decomposition can be applied in an analogous way to the LU
decomposition discussed above, but we shall not explore it further.
Cramer’s rule
An alternative method of solution is to use Cramer’s rule , which also provides
some insight into the nature of the solutions in the various cases. To illustratethis method let us consider a set of three equations in three unknowns,
A
11x1+A12x2+A13x3=b1,
A21x1+A22x2+A23x3=b2, (8.129)
A31x1+A32x2+A33x3=b3,
which may be represented by the matrix equation Ax=b. We wish either to find
the solution(s) xto these equations or to establish that there are no solutions.
From result (vi) of subsection 8.9.1, the determinant |A|is unchanged by adding
to its first column the combination
x2
x1×(second column of |A|)+x3
x1×(third column of |A|).
We thus obtain
|A|=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA
11A12A13
A21A22A23
A31A32A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA
11+(x2/x1)A12+(x3/x1)A13A12A13
A21+(x2/x1)A22+(x3/x1)A23A22A23
A31+(x2/x1)A32+(x3/x1)A33A32A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,
which, on substituting b
i/x1for the ith entry in the first column, yields
|A|=1
x1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb
1A12A13
b2A22A23
b3A32A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=1
x1∆1.
The determinant ∆ 1is known as a Cramer determinant . Similar manipulations of
the second and third columns of |A|yield x2andx3, and so the full set of results
reads
x1=∆1
|A|,x 2=∆2
|A|,x 3=∆3
|A|, (8.130)
where
∆1=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb
1A12A13
b2A22A23
b3A32A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,∆
2=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA
11b1A13
A21b2A23
A31b3A33vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,∆
3=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA
11A12b1
A21A22b2
A31A32b3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
It can be seen that each Cramer determinant ∆
iis simply|A|but with column i
replaced by the RHS of the original set of equations. If |A|/negationslash= 0 then (8.130) gives
304
8.18 SIMULTANEOUS LINEAR EQUATIONS
the unique solution. The proof given here appears to fail if any of the solutions
xiis zero, but it can be shown that result (8.130) is valid even in such a case.IUse Cramer’s rule to solve the set of simultaneous equations (8.123).
Let us again represent these simultaneous equations by the matrix equation Ax=b,i . e ./0/@243
1−2−2
−33 2
/1A
/0/@x1
x2
x3
/1
A
=
/0
/@
4
0
−7
/1A.
From (8.58), the determinant of Ais given by |A|= 11. Following the discussion given
above, the three Cramer determinants are
∆1=
//////443
0−2−2
−73 2
//////,∆2=
//////243
10−2
−3−72
//////,∆3=
//////244
1−20
−33−7
//////.
These may be evaluated using the properties of determinants listed in subsection 8.9.1
and we find ∆ 1= 22, ∆ 2=−33 and ∆ 3= 44. From (8.130) the solution to the equations
(8.123) is given by
x1=22
11=2,x 2=−33
11=−3,x 3=44
11=4,
which agrees with the solution found in the previous example.
J
At this point it is useful to consider each of the three equations (8.129) as rep-
resenting a plane in three-dimensional Cartesian coordinates. Using result (7.42)
of chapter 7, the sets of components of the vectors normal to the planes are(A
11,A12,A13), (A21,A22,A23)a n d( A31,A32,A33), and using (7.46) the perpendic-
ular distances of the planes from the origin are given by
di=biparenleftbig
A2
i1+A2
i2+A2
i3parenrightbig1/2fori=1,2,3.
Finding the solution(s) to the simultaneous equations above corresponds to finding
the point(s) of intersection of the planes.
If there is a unique solution the planes intersect at only a single point. This
happens if their normals are linearly independent vectors. Since the rows of A
represent the directions of these normals, this requirement is equivalent to |A|/negationslash=0 .
Ifb=(000 )T=0then all the planes pass through the origin and, since there
is only a single solution to the equations, the origin is that solution.
Let us now turn to the cases where |A|= 0. The simplest such case is that in
which all three planes are parallel; this implies that the normals are all paralleland so Ais of rank 1. Two possibilities exist:
(i) the planes are coincident, i.e. d
1=d2=d3,i nw h i c hc a s et h e r ei sa n
infinity of solutions;
(ii) the planes are not all coincident, i.e. d1/negationslash=d2and/or d1/negationslash=d3and/or
d2/negationslash=d3, in which case there are no solutions.
305
MATRICES AND VECTOR SPACES
(a) (b)
Figure 8.1 The two possible cases when Ais of rank 2. In both cases all the
normals lie in a horizontal plane but in ( a) the planes all intersect on a single
line (corresponding to an infinite number of solutions) whilst in ( b)t h e r ea r e
no common intersection points (no solutions).
It is apparent from (8.130) that case (i) occurs when all the Cramer determinants
are zero and case (ii) occurs when at least one Cramer determinant is non-zero.
The most complicated cases with |A|= 0 are those in which the normals to the
planes themselves lie in a plane but are not parallel. In this case Ahas rank 2.
Again two possibilities exist and these are shown in figure 8.1. Just as in therank-1 case, if all the Cramer determinants are zero then we get an infinity ofsolutions (this time on a line). Of course, in the special case in which b=0(and
the system of equations is homogeneous), the planes all pass through the originand so they must intersect on a line through it. If at least one of the Cramer
determinants is non-zero, we get no solution.
These rules may be summarised as follows.
(i)|A|/negationslash=0 , b/negationslash=0: The three planes intersect at a single point that is not the
origin, and so there is only one solution, given by both (8.122) and (8.130).
(ii)|A|/negationslash=0 , b=0: The three planes intersect at the origin only and there is
only the trivial solution, x=0 .
(iii)|A|=0 , b/negationslash=0, Cramer determinants all zero: There is an infinity of
solutions either on a line if Ais rank 2, i.e. the cofactors are not all zero,
or on a plane if Ais rank 1, i.e. the cofactors are all zero.
(iv)|A|=0 , b/negationslash=0, Cramer determinants not all zero: No solutions.
(v)|A|=0 , b=0: The three planes intersect on a line through the origin
giving an infinity of solutions.
8.18.3 Singular value decomposition
There exists a very powerful technique for dealing with a simultaneous set of
linear equations Ax=b, such as (8.118), which may be applied whether or not
306
8.18 SIMULTANEOUS LINEAR EQUATIONS
the number of simultaneous equations Mis equal to the number of unknowns N.
This technique is known as singular value decomposition (SVD) and is the method
of choice in analysing anyset of simultaneous linear equations.
We will consider the general case, in which Ais an M×N(complex) matrix.
Let us suppose we can write Aas the product §
A=USV†, (8.131)
where the matrices U,Sand Vhave the following properties.
(i) The square matrix Uhas dimensions M×Mand is unitary .
(ii) The matrix Shas dimensions M×N( t h es a m ed i m e n s i o n sa st h o s eo f A)
and is diagonal in the sense that Sij=0i f i/negationslash=j. We denote its diagonal
elements by sifori=1,2,...,p,w h e r e p= min( M,N); these elements are
termed the singular values ofA.
(iii) The square matrix Vhas dimensions N×Nand is unitary .
We must now determine the elements of these matrices in terms of the elements of
A. From the matrix A, we can construct two square matrices: A†Awith dimensions
N×Nand AA†with dimensions M×M. Both are clearly Hermitian . From (8.131),
and using the fact that Uand Vare unitary, we find
A†A=VS†U†USV†=VS†SV†(8.132)
AA†=USV†VS†U†=USS†U†, (8.133)
where S†Sand SS†are diagonal matrices with dimensions N×NandM×M
respectively. The first pelements of each diagonal matrix are s2
i,i=1,2,...,p,
where p= min( M,N), and the rest (where they exist) are zero.
These two equations imply that both V−1A†AVparenleftbig
=V−1A†A(V†)−1parenrightbig
and, by
a similar argument, U−1AA†U, must be diagonal. From our discussion of the
diagonalisation of Hermitian matrices in section 8.16, we see that the columns of
Vmust therefore be the normalised eigenvectors vi,i=1,2,...,N , of the matrix
A†Aand the columns of Umust be the normalised eigenvectors uj,j=1,2,...,M ,
of the matrix AA†. Moreover, the singular values simust satisfy s2
i=λi,w h e r e
theλiare the eigenvalues of the smaller of A†Aand AA†. Clearly, the λiare
also some of the eigenvalues of the larger of these two matrices, the remainingones being equal to zero. Since each matrix is Hermitian, the λ
iare real and the
singular values simay be taken as real and non-negative. Finally, to make the
decomposition (8.131) unique, it is customary to arrange the singular values in
decreasing order of their values, so that s1≥s2≥···≥sp.
§The proof that such a decomposition always exists is beyond the scope of this book. For a full
account of SVD one might consult, for example, Golub & Van Loan, Matrix Computations , second
edition (Johns Hopkins University Press).
307
MATRICES AND VECTOR SPACESIShow that, for i=1,2,...,p,Avi=siuiand A†ui=sivi,w h e r e p=m i n ( M,N).
Post-multiplying both sides of (8.131) by V, and using the fact that Vis unitary, we obtain
AV=US.
Since the columns of Vand Uconsist of the vectors viand ujrespectively and Shas only
diagonal non-zero elements, we find immediately that, for i=1,2,...,p,
Avi=siui. (8.134)
Moreover, we note that Avi=0f o r i=p+1,p+2,...,N .
Taking the Hermitian conjugate of both sides of (8.131) and post-multiplying by U,w e
obtain
A†U=VS†=VST,
where we have used the fact that Uis unitary and Sis real. We then see immediately that,
fori=1,2,...,p,
A†ui=sivi. (8.135)
We also note that A†ui=0f o r i=p+1,p+2,...,M . Results (8.134) and (8.135) are useful
for investigating the properties of the SVD.
J
The decomposition (8.131) has some advantageous features for the analysis of
sets of simultaneous linear equations. These are best illustrated by writing the
decomposition (8.131) in terms of the vectors uiand vias
A=psummationdisplay
i=1siui(vi)†,
where p= min( M,N). It may be, however, that some of the singular values si
arezero, as a result of degeneracies in the set of Mlinear equations Ax=b.
Let us suppose that there are rnon-zero singular values. Since our convention is
to arrange the singular values in order of decreasing size, the non-zero singular
values are si,i=1,2,...,r, and the zero singular values are sr+1,sr+2,...,s p.
Therefore we can write Aas
A=rsummationdisplay
i=1siui(vi)†. (8.136)
Let us consider the action of (8.136) on an arbitrary vector x.T h i si sg i v e nb y
Ax=rsummationdisplay
i=1siui(vi)†x.
Since ( vi)†xis just a number, we see immediately that the vectors ui,i=1,2,...,r,
must span the range of the matrix A; moreover, these vectors form an orthonor-
mal basis for the range. Further, since this subspace is r-dimensional, we have
rank A=r, i.e. the rank of Ais equal to the number of non-zero singular values.
The SVD is also useful in characterising the null space of A. From (8.119),
we already know that the null space must have dimension N−r;s oi f Ahasr
308
8.18 SIMULTANEOUS LINEAR EQUATIONS
non-zero singular values si,i=1,2,...,r, then from the worked example above
we have
Avi=0 f o r i=r+1,r+2,...,N.
Thus, the N−rvectors vi,i=r+1,r+2,...,N , form an orthonormal basis for
the null space of A.IFind the singular value decompostion of the matrix
A=
/0B/@22 2 2
17
101
10−17
10−1
10
3
59
5−3
5−9
5
/1CA. (8.137)
The matrix Ahas dimension 3 ×4( i . e . M=3 , N= 4), and so we may construct from
it the 3×3m a t r i x AA†and the 4×4m a t r i x A†A(in fact, since Ais real, the Hermitian
conjugates are just transposes). We begin by finding the eigenvalues λiand eigenvectors ui
of the smaller matrix AA†. This matrix is easily found to be given by
AA†=
/0/@16 0 0
029
512
5
012
536
5
/1A,
and its characteristic equation reads////
//
16−λ 00
029
5−λ12
5
012
536
5−λ
////
//
=( 1 6−λ)(36−13λ+λ2)=0 .
Thus, the eigenvalues are λ1= 16, λ2=9 , λ3= 4. Since the singular values of Aare given
bysi=√λiand the matrix Sin (8.131) has the same dimensions as A, we have
S=
/0/@4000
03000020
/1A, (8.138)
where we have arranged the singular values in order of decreasing size. Now the matrix U
has as its columns the normalised eigenvectors uiof the 3×3m a t r i x AA†. These normalised
eigenvectors correspond to the eigenvalues of AA†as follows:
λ1=1 6⇒ u1= ( 100 )T
λ2=9⇒ u2=( 03
54
5)T
λ3=4⇒ u3=( 0−4
53
5)T,
a n ds ow eo b t a i nt h em a t r i x
U=
/0/@10 0
03
5−4
5
04
53
5
/1A. (8.139)
The columns of the matrix Vin (8.131) are the normalised eigenvectors of the 4 ×4
matrix A†A, which is given by
A†A=1
4
/0B/@29 21 3 11
21 29 11 3
31 1 2 9 2 1
1 132 1 2 9
/1CA.
309
MATRICES AND VECTOR SPACES
We already know from the above discussion, however, that the non-zero eigenvalues of
this matrix are equal to those of AA†found above, and that the remaining eigenvalue is
zero. The corresponding normalised eigenvectors are easily found:
λ1=1 6⇒ v1=1
2( 1111 )T
λ2=9⇒ v2=1
2(1 1−1−1)T
λ3=4⇒ v3=1
2(−111−1)T
λ4=0⇒ v4=1
2(1−11−1)T
and so the matrix Vis given by
V=1
2
/0B/@11−11
11 1 −1
1−11 1
1−1−1−1
/1CA. (8.140)
Alternatively, we could have found the first three columns of Vby using the relation
(8.135) to obtain
vi=1
siA†uifori=1,2,3.
The fourth eigenvector could then be found using the Gram–Schmidt orthogonalisation
procedure. We note that if there were more than one eigenvector corresponding to a zeroeigenvalue then we would need to use this procedure to orthogonalise these eigenvectorsbefore constructing the matrix V.
Collecting our results together, we find the SVD of the matrix A:
A=USV
†=
/0/@10 0
03
5−4
5
04
53
5
/1A
/0/@4000
03000020
/1A
/0BBB/@1
21
21
21
2
1
21
2−1
2−1
2
−1
21
21
2−1
2
1
2−1
21
2−1
2
/1CCCA;
this can be verified by direct multiplication.
J
Let us now consider the use of SVD in solving a set of Msimultaneous linear
equations in Nunknowns, which we write again as Ax=b. Firstly, consider
the solution of a homogeneous set of equations, for which b=0. As mentioned
previously, if Ais square and non-singular (and so possesses no zero singular
values) then the equations have the unique trivial solution x=0.O t h e r w i s e , any
of the vectors vi,i=r+1,r+2,...,N , or any linear combination of them, will
be a solution.
In the inhomogeneous case, where bis not a zero vector, the set of equations
will possess solutions if blies in the range of A. To investigate these solutions, it
is convenient to introduce the N×Mmatrix S, which is constructed by taking
the transpose of Sin (8.131) and replacing each non-zero singular value sion the
diagonal by 1 /si. It is clear that, with this construction,
SS=I. (8.141)
We note, however, that the matrix Sisnotthe inverse of Ssince SS/negationslash=I.
310
8.18 SIMULTANEOUS LINEAR EQUATIONS
Nevertheless, using property (8.141) and the unitarity of the matrices Uand V,a
solution to the equations Ax=bis given by
x=VSU†b. (8.142)
We may, however, add to this solution anylinear combination of the p−rvectors
vi,i=r+1,r+2,...,p, that form an orthonormal basis for the null space of A;
thus, in general, there exists an infinity of solutions (although it is straightforward
to show that (8.142) is the solution vector of shortest length). The only way inwhich the solution (8.142) can be unique is if the rank requals N, so that the
matrix Adoes not possess a null space; this only occurs if Ais square and
non-singular.
Ifbdoes not lie in the range of Athen the set of equations Ax=bdoes
not have a solution. Nevertheless, the vector (8.142) provides the closest possible‘solution’ in a least-squares sense. In other words, although the vector (8.142)does not exactly solve Ax=b, it is the vector that minimises the residual
/epsilon1=|Ax−b|,
where here the vertical lines denote the absolute value of the quantity they
contain, not the determinant. This is proved as follows.
Suppose we were to add some arbitrary vector x
/primeto the vector xin (8.142).
This would result in the addition of the vector b/prime=Ax/primetoAx−b;b/primeis clearly in
the range of Asince any part of x/primebelonging to the null space of Acontributes
nothing to Ax/prime. We would then have
|Ax−b+b/prime|=|(USV†)(VSU†b)−b+b/prime|
=|(USSU†−I)b+b/prime|
=|U[(SS−I)U†b+U†b/prime]|
=|(SS−I)U†b+U†b/prime|; (8.143)
in the last line we have made use of the fact that the length of a vector is left
unchanged under the action of the unitary matrix U. Now, the diagonal square
matrix SS, with dimensions M×M, will have non-zero entries (actually all equal
to unity ) only for those values of jfor which sj/negationslash= 0. Thus, the jth component of
the vector ( SS−I)U†bwill only be non-zero when sj= 0. However, the jth element
of the vector U†b/primeis given by the scalar product ( uj)†b/prime, which is non-zero only if
sj/negationslash=0 ,s i n c e b/primelies in the range of A. Thus, as these two terms only contribute to
(8.143) for two disjoint sets of j-values, its minimum value, as x/primeis varied, occurs
when b/prime=0; this requires x/prime=0.
311
MATRICES AND VECTOR SPACESIFind the solution(s) to the set of simultaneous linear equations Ax=b,w h e r e Ais given
by (8.137) and b= ( 100 )T.
To solve the set of equations, we begin by calculating the vector given in (8.142),
x=VSU†b,
where Uand Vare given by (8.139) and (8.140) respectively and Sis obtained by taking
the transpose of Sin (8.138) and replacing all the non-zero singular values siby 1/si. Thus,
Sreads
S=
/0BBB/@1
400
01
30
001
2
000
/1CCCA.
Substituting the appropriate matrices into the expression for xwe find
x=1
8( 1111 )T. (8.144)
It is straightforward to show that this solves the set of equations Ax=bexactly, and
so the vector b=( 1 0 0 )Tmust lie in the range of A. This is, in fact, immediately
clear, since b=u1. The solution (8.144) is not, however, unique. There are three non-zero
singular values, but N= 4. Thus, the matrix Ahas a one-dimensional null space, which
is ‘spanned’ by v4, the fourth column of V, given in (8.140). The solutions to our set of
equations, consisting of the sum of the exact solution and anyvector in the null space of
A, therefore lie along the line
x=1
8( 1111 )T+α(1−11−1)T,
where the parameter αcan take any real value. We note that (8.144) is the point on this
line that is closest to the origin.
J
8.19 Exercises
8.1 Which of the following statements about linear vector spaces are true? Where a
statement is false, give a counter-example to demonstrate this.
(a) Non-singular N×Nmatrices form a vector space of dimension N2.
(b) Singular N×Nmatrices form a vector space of dimension N2.
(c) Complex numbers form a vector space of dimension 2.
(d) Polynomial functions of xform an infinite-dimensional vector space.
(e) Series{a0,a1,a2,...,a N}for which
PN
n=0|an|2=1f o r ma n N-dimensional
vector space.
(f) Absolutely convergent series form an infinite-dimensional vector space.
(g) Convergent series with terms of alternating sign form an infinite-dimensional
vector space.
8.2 Evaluate the determinants
(a)
//////ahg
hbf
gfc
//////, (b)
///
///
/
1023
01−21
3−34−2
−21−21
///
///
/
,
312
8.19 EXERCISES
and
(c)
//////
/
gc ge a +ge gb +ge
0bb b
ce e b +e
abb +fb +d
//////
/
.
8.3 Using the properties of determinants, solve with a minimum of calculation the
following equations for x:
(a)
///////xaa 1
axb 1
abx 1
abc 1
///////=0, (b)
//////x+2 x+4 x−3
x+3 xx +5
x−2x−1x+1
//////=0.
8.4 Consider the matrices
(a) B=
/0/@0−ii
i0−i
−ii 0
/1A,(b) C=1√
8
/0/@√
3−√
2−√
3
1√
6−1
20 2
/1A.
Are they (i) real, (ii) diagonal, (iii) symmetric, (iv) antisymmetric, (v) singular,
(vi) orthogonal, (vii) Hermitian, (viii) anti-Hermitian, (ix) unitary, (x) normal?
8.5 By considering the matrices
A=
/
10
00
/
, B=
/
00
34
/
show that AB=0doesnotimply that either AorBis the zero matrix but that
it does imply that at least one of them is singular.
8.6 (a) The basis vectors of the unit cell of a crystal, with the origin Oat one corner,
are denoted by e1,e2,e3.T h em a t r i x Ghas elements Gij,w h e r e Gij=ei·ej
and Hijare the elements of the matrix H≡G−1. Show that the vectors
fi=
P
jHijejare the reciprocal vectors and that Hij=fi·fj.
(b) If the vectors uandvare given by
u=
X
iuiei,v=
X
ivifi,
obtain expressions for |u|,|v|,a n du·v.
(c) If the basis vectors are each of length aand the angle between each pair is
π/3, write down Gand hence obtain H.
(d) Calculate (i) the length of the normal from Oonto the plane containing the
points p−1e1,q−1e2,r−1e3, and (ii) the angle between this normal and e1.
8.7 (a) Show that if Ais Hermitian and Uis unitary then U−1AUis Hermitian.
(b) Show that if Ais anti-Hermitian then iAis Hermitian.
(c) Prove that the product of two Hermitian matrices Aand Bis Hermitian if
and only if Aand Bcommute.
(d) Prove that if Sis a real antisymmetric matrix then A=(I−S)(I+S)−1is
orthogonal. If Ais given by
A=
/
cosθsinθ
−sinθcosθ
/
then find the matrix Sthat is needed to express Ain the above form.
(e) If Kis skew-hermitian, i.e. K†=−K, prove that V=(I+K)(I−K)−1is unitary.
313
MATRICES AND VECTOR SPACES
8.8 Aand Bare real non-zero 3 ×3 matrices and satisfy the equation
(AB)T+B−1A=0.
(a) Prove that if Bis orthogonal then Ais antisymmetric.
(b) Without assuming that Bis orthogonal, prove that Ais singular.
8.9 The commutator [X,Y] of two matrices is defined by the equation
[X,Y]=XY−YX.
Two anti-commuting matrices Aand Bsatisfy
A2=I, B2=I,[A,B]=2 iC.
(a) Prove that C2=Iand that [ B,C]=2 iA.
(b) Evaluate [[[ A,B],[B,C]],[A,B]].
8.10 The four matrices Sx,Sy,Szand Iare defined by
Sx=
/
01
10
/
, Sy=
/
0−i
i0
/
,
Sz=
/
10
0−1
/
,I=
/
10
01
/
,
where i2=−1. Show that S2
x=Iand SxSy=iSz, and obtain similar results
by permutting x,yandz.G i v e nt h a t vis a vector with Cartesian components
(vx,vy,vz), the matrix S(v) is defined as
S(v)=vxSx+vySy+vzSz.
Prove that, for general non-zero vectors aandb,
S(a)S(b)=a·bI+iS(a×b).
Without further calculation, deduce that S(a)a n d S(b) commute if and only if a
andbare parallel vectors.
8.11 A general triangle has angles α,βandγand corresponding opposite sides a,
bandc. Express the length of each side in terms of the lengths of the other
two sides and the relevant cosines, writing the relationships in matrix and vectorform using the vectors having components a, b, c and cos α,cosβ,cosγ. Invert the
matrix and hence deduce the cosine-law expressions involving α,βandγ.
8.12 Given a matrix
A=
/0/@1α0
β10
001
/1A,
where αandβare non-zero complex numbers, find its eigenvalues and eigenvec-
tors. Find the respective conditions for (a) the eigenvalues to be real and (b) theeigenvectors to be orthogonal. Show that the conditions are jointly satisfied ifand only if Ais Hermitian.
8.13 Using the Gram–Schmidt procedure:
(a) construct an orthonormal set of vectors from the following:
x
1= ( 0011 )T, x2=( 1 0−10 )T,
x3= ( 1202 )T, x4= ( 2111 )T;
314
8.19 EXERCISES
(b) find an orthonormal basis, within a four-dimensional Euclidean space, for
t h e s u b s p a c e s p a n n e d b y t h e t h r e e v e c t o r s ( 1200 )T,( 3−120 )T
a n d ( 0021 )T.
8.14 If a unitary matrix Uis written as A+iB,w h e r e Aand Bare Hermitian with
non-degenerate eigenvalues, show the following:
(a) Aand Bcommute;
(b) A2+B2=I;
(c) The eigenvectors of Aare also eigenvectors of B;
(d) The eigenvalues of Uhave unit modulus (as is necessary for any unitary
matrix).
8.15 Determine which of the matrices below are mutually commuting, and, for those
that are, demonstrate that they have a complete set of eigenfunctions in common:
A=
/
6−2
−29
/
, B=
/
18
8−11
/
,
C=
/
−9−10
−10 5
/
,D=
/
14 2
21 1
/
.
8.16 Find the eigenvalues and a set of eigenvectors of the matrix/0/@13−1
34−2
−1−22
/1A.
Verify that its eigenvectors are mutually orthogonal.
8.17 Find three real orthogonal column matrices, each of which is a simultaneous
eigenvector of
A=
/0/@001
010100
/1A and B=
/0/@011
101110
/1A.
8.18 Use the results of the first worked example in section 8.14 to evaluate, without
repeated matrix multiplication, the expression A6x,w h e r e x=( 2 4−1)Tand
Ais the matrix given in the example.
8.19 Given that Ais a real symmetric matrix with normalised eigenvectors eiobtain
the coefficients αiinvolved when column matrix x, which is the solution of
Ax−µx=v, (∗)
is expanded as x=
P
iαiei.H e r e µis a given constant and vis a given column
matrix.
(a) Solve (*) when
A=
/0/@210
120003
/1A,
µ=2a n d v= ( 123 )T.
(b) Would (*) have a solution if µ=1a n d( i ) v= ( 123 )T, (ii) v=
(2 2 3)T?
315
MATRICES AND VECTOR SPACES
8.20 Demonstrate that the matrix
A=
/0/@20 0
−644
3−10
/1A,
is defective, i.e. does not have three linearly independent eigenvectors, by showing
the following:
(a) its eigenvalues are degenerate and, in fact, all equal;
(b) any eigenvector has the form ( µ(3µ−2ν)ν)T.
(c) if two pairs of values, µ1,ν1andµ2,ν2, define two independent eigenvectors
v1and v2thenanythird similarly defined eigenvector v3c a nb ew r i t t e na sa
linear combination of v1and v2,i . e .
v3=av1+bv2
where
a=µ3ν2−µ2ν3
µ1ν2−µ2ν1and b=µ1ν3−µ3ν1
µ1ν2−µ2ν1.
Illustrate (c) using the example ( µ1,ν1)=( 1 ,1),(µ2,ν2)=( 1 ,2) and ( µ3,ν3)=
(0,1).
Show further that any matrix of the form/0/@200
6n−64−2n4−4n
3−3nn−12 n
/1A
is defective, with the same eigenvalues and eigenvectors as A.
8.21 By finding the eigenvectors of the Hermitian matrix
H=
/
10 3 i
−3i2
/
,
construct a unitary matrix Usuch that U†HU= Λ, where Λ is a real diagonal
matrix.
8.22 Use the stationary properties of quadratic forms to determine the maximum and
minimum values taken by the expression
Q=5x2+4y2+4z2+2xz+2xy
on the unit sphere x2+y2+z2= 1. For what values of x, y, z do they occur?
8.23 Given that the matrix
A=
/0/@2−10
−12−1
0−12
/1A
has two eigenvectors of the form (1 y1)T, use the stationary property of the
expression J(x)= xTAx/(xTx) to obtain the corresponding eigenvalues. Deduce
the third eigenvalue.
8.24 Find the lengths of the semi-axes of the ellipse
73x2+7 2xy+5 2y2= 100 ,
and determine its orientation.
8.25 The equation of a particular conic section is
Q≡8x2
1+8x2
2−6x1x2= 110 .
Determine the type of conic section this represents, the orientation of its principal
axes, and relevant lengths in the directions of these axes.
316
8.19 EXERCISES
8.26 Show that the quadratic surface
5x2+1 1y2+5z2−10yz+2xz−10xy=4
is an ellipsoid with semi-axes of lengths 2, 1 and 0 .5. Find the direction of its
longest axis.
8.27 Find the direction of the axis of symmetry of the quadratic surface
7x2+7y2+7z2−20yz−20xz+2 0xy=3.
8.28 Find the eigenvalues, and sufficient of the eigenvectors, of the following matrices
to be able to describe the quadratic surfaces associated with them.
(a)
/0/@51−1
151
−11 5
/1A,(b)
/0/@122
212221
/1A.(c)
/0/@12 1
24 2
−121
/1A.
8.29 (a) Rearrange the result A/prime=S−1ASof section 8.16 to express the original
matrix Ain terms of the unitary matrix Sand the diagonal matrix A/prime. Hence
show how to construct a matrix Athat has given eigenvalues and given
(orthogonal) column matrices as its eigenvectors.
(b) Find the matrix with eigenvectors (1 2 1)T,(1−11 )Tand (1 0 −1)T,
and corresponding eigenvalues λ,µandν.
(c) Try a particular case, say λ=3 , µ=−2a n d ν= 1, and verify by explicit
solution that the matrix so found does have these eigenvalues.
8.30 Find an orthogonal transformation that takes the quadratic form
Q≡−x2
1−2x2
2−x2
3+8x2x3+6x1x3+8x1x2
into the form
µ1y2
1+µ2y2
2−4y2
3,
and determine µ1andµ2(see section 8.17).
8.31 One method of determining the nullity (and hence the rank) of an M×Nmatrix
Ais as follows.
•Write down an augmented transpose of A, by adding on the right an N×N
unit matrix and thus producing an N×(M+N) array B.
•Subtract a suitable multiple of the first row of Bfrom each of the other lower
rows so as to make Bi1=0f o r i>1.
•Subtract a suitable multiple of the second row (or the uppermost row that
does not start with Mzero values) from each of the other lower rows so as to
make Bi2=0f o r i>2.
•Continue in this way until all remaining rows have zeroes in the first Mplaces.
The number of such rows is equal to the nullity of Aand the Nrightmost
entries of these rows are the components of vectors that span the null space.They can be made orthogonal if they are not so already.
Use this method to show that the nullity of
A=
/0BBB/@−13 2 7
31 0−61 7
−1−22−3
23−44
40−8−4
/1CCCA
is 2 and that an orthogonal base for the null space of Ais provided by any two
column matrices of the form (2 + αi−2αi1αi)Tfor which the αi(i=1,2)
are real and satisfy 6 α1α2+2 (α1+α2)+5=0 .
317
MATRICES AND VECTOR SPACES
8.32 Do the following sets of equations have non-zero solutions? If so, find them.
(a) 3 x+2y+z=0 , x−3y+2z=0 , 2 x+y+3z=0 .
(b) 2 x=b(y+z), x=2a(y−z), x=( 6a−b)y−(6a+b)z.
8.33 Solve the simultaneous equations
2x+3y+z=1 1,
x+y+z=6,
5x−y+1 0z=3 4.
8.34 Solve the following simultaneous equations for x1,x2and x3,u s i n gm a t r i x
methods:
x1+2x2+3x3=1,
3x1+4x2+5x3=2,
x1+3x2+4x3=3.
8.35 Show that the following equations have solutions only if η= 1 or 2, and find
them in these cases:
x+y+z=1,
x+2y+4z=η,
x+4y+1 0z=η2.
8.36 Find the condition(s) on αsuch that the simultaneous equations
x1+αx2=1,
x1−x2+3x3=−1,
2x1−2x2+αx3=−2
have (a) exactly one solution, (b) no solutions, or (c) an infinite number of
solutions; give all solutions where they exist.
8.37 Make an LUdecomposition of the matrix
A=
/0/@36 9
10 5
2−21 6
/1A
and hence solve Ax=b,w h e r e( i ) b= (21 9 28)T, (ii) b= (21 7 22)T.
8.38 Make an LUdecomposition of the matrix
A=
/0B/@2−31 3
14−3−3
53−1−1
3−6−31
/1CA.
Hence solve Ax=bfor (i) b=(−418−5)T, (ii) b=(−10 0−3−24)T.
Deduce that det A=−160 and confirm this by direct calculation.
8.39 Use the Cholesky separation method to determine whether the following matrices
are positive definite. For each that is, determine the corresponding lower diagonalmatrix L:
A=
/0/@21 3
13−1
3−11
/1A, B=
/0/@50√
3
030√
30 3
/1A.
318
8.20 HINTS AND ANSWERS
8.40 Find the equation satisfied by the squares of the singular values of the matrix
associated with the following over-determined set of equations:
2x+3y+z=0
x−y−z=1
2x+y=0
2y+z=−2.
Show that one of the singular values is close to zero. Determine the two larger
singular values by an appropriate iteration process and the smallest by indirectcalculation.
8.41 Find the SVD of/0/@0−1
11
−10
/1A,
showing that the singular values are√
3a n d1 .
8.42 Find the SVD form of the matrix
A=
/0B/@22 28−22
1−2−19
19−2−1
−61 2 6
/1CA.
Hence find the best solution xto the equation Ax=bwhen (i) b=( 6−
39 15 18)T, (ii) b=( 9−42 15 15)T, showing that (i) has an exact solution,
but that the best solution to (ii) has a residual of√
18.
8.43 Four experimental measurements of particular combinations of three physical
variables, x,yandz, gave the following inconsistent results:
13x+2 2y−13z=4,
10x−8y−10z=4 4,
10x−8y−10z=4 7,
9x−18y−9z=7 2.
Find the SVD best values for x,yandz. Identify the null space of Aand hence
obtain the general SVD solution.
8.20 Hints and answers
8.1 (a) False. ON,t h e N×Nnull matrix, is notnon-singular.
(b) False. Consider the sum of
/
10
00
/
and
/
0001
/
.
(c) True.
(d) True.
(e) False. Consider bn=an+anfor which
PN
n=0|bn|2=4/negationslash= 1, or note that there
is no zero vector with unit norm.
(f) True.(g) False. Consider the two series defined by
a
0=1
2,a n=2 (−1
2)nfor n≥1; bn=−(−1
2)nfor n≥0.
T h es e r i e st h a ti st h es u mo f {an}and{bn}does not have alternating signs
and so closure does not hold.
8.2 (a) abc+2fgh−af2−bg2−ch2,( b )0 ,( c ) ab(ab−cd).
319
MATRICES AND VECTOR SPACES
8.3 (a) x=a,borc;( b ) x=−1, equation is linear in x.
8.4 (a) iv, v, vii, x; (b) i, vi, ix, x.8.6 (b) (P
ijuiGijuj)1/2,(
P
ijviHijvj)1/2,
P
iuivi;
(c) H=1
a2
/0/@3/2−1/2−1/2
−1/23 /2−1/2
−1/2−1/23 /2
/1A.(d) (i) M−1, (ii) cos−1(p/Ma)w h e r e
M=a−1[3(p2+q2+r2)/2−qr−pr−pq]1/2.
8.7 (d) S=
/
0−tan(θ/2)
tan(θ/2) 0
/
.(e) Note that ( I+K)(I−K)= I−K2=
(I−K)(I+K).
8.8 (b) Note that |−A|=(−1)3|A|.
8.9 (b) 32 iA.
8.10 S(a)S(b)−S(b)S(a)=2 iS(a×b) and equals zero only if a×b=0.
8.11 a=bcosγ+ccosβ, and cyclic permutations; a2=b2+c2−2bccosα, and cyclic
permutations.
8.12 λ=1 ,(001 )T;
λ=1+( αβ)1/2,(α1/2β1/20)T;
λ=1−(αβ)1/2,(α1/2−β1/20)T;
(a)αβreal and >0; (b)|α|=|β|.
8.13 (a) 2−1/2( 0011 )T,6−1/2(2 0−11 )T,
39−1/2(−16−11 )T,1 3−1/2( 212 −2)T.
(b) 5−1/2( 1200 )T,(345)−1/2(14−71 00 )T,
(18285)−1/2(−56 28 98 69)T.
8.14 (a) Use UU†=U†U; (b) use UU†=I; (c) apply the result of subsection 8.13.5 to
give the eigenvalue for Uasλ+iµ; (d) apply result (b) to eigenvector uofUto
deduce that λ2+µ2=1.
8.15 Cdoes not commute with the others; A,Band Dhave (1−2)Tand (2 1)Tas
common eigenvectors.
8.16 λ=1 ,(113 )T;
λ=3±√
15, (5±√
15 7±2√
15−4∓√
15)T;
8.17 For A:( 1 0−1)T,(1α11)T,(1α21)T.
For B: ( 111 )T,(β1γ1−β1−γ1)T,(β2γ2−β2−γ2)T.
Theαi,βiandγiare arbitrary.
Simultaneous and orthogonal: (1 0 −1)T,( 111 )T,(1−21 )T.
8.18 Express xas a linear combination of the eigenvectors of Aand use the fact that
Anx=λnxfor an eigenvector; x=3x(1)−x(2);A6x=(−537 921 729)T.
8.19 αj=(v·ej∗)/(λj−µ), where λjis the eigenvalue corresponding to ej.
(a) x= ( 213 )T.
(b) Since µis equal to one of A’s eigenvalues λj, the equation only has a solution
ifv·ej∗= 0; (i) no solution; (ii) x= ( 113 /2)T.
8.20 (a) All eigenvalues equal 2; (c) a=−1,b=1 .
8.21 U= (10)−1/2(1,3i;3i,1), Λ = (1 ,0;0,11).
8.22 Maximum equal to 6 at ±(2,1,1)/√
6; minimum equal to 3 at ±(1,−1,−1)/√
3.
8.23 J=( 2y2−4y+4)/(y2+2) with stationary values at y=±√2 and corresponding
eigenvalues 2 ∓√2. From the trace property of A, the third eigenvalue equals 2.
8.24 The eigenvalues, after making the RHS unity, are 1 /4 and 1, corresponding to
semi-axis lengths of 2 and 1. The major axis makes an angle tan−1(−4/3) with
the positive x-axis.
8.25 Ellipse; θ=π/4,a=√
22;θ=3π/4,b=√
10.
320
8.20 HINTS AND ANSWERS
8.26 The eigenvector corresponding to the smallest eigenvalue is in the direction
(1,1,1)/√
3.
8.27 The direction of the eigenvector having the non-repeated eigenvalue is
(1,1,−1)/√
3.
8.28 (a) Eigenvalues 6, 6, 3; an ellipsoid with circular cross-section of radius r,
say, perpendicular to the direction (1 ,−1,1)/√3, and with semi-axis in that
direction of√2r.
(b) Eigenvalues 5, −1,−1; a hyperboloid of revolution about an axis in the
direction (1 ,1,1)/√3, the two halves of the hyperboloid being asymptotic to
that cone of semi-angle tan−1√5 that passes through the origin and also has
its axis in that direction.
(c) Eigenvalues 6, 0, 0; a pair of parallel planes, equidistant from the origin and
with their normals in the directions ±(1,2,1)/√6.
8.29 (a) A=SA/primeS†,w h e r e Sis the matrix whose columns are the eigenvectors of the
matrix Ato be constructed, and A/prime=d i a g( λ, µ, ν).
(b) A=(λ+2µ+3ν,2λ−2µ, λ+2µ−3ν;2λ−2µ,4λ+2µ,2λ−2µ;
λ+2µ−3ν,2λ−2µ, λ+2µ+3ν).
(c)1
3(1,5,−2;5,4,5;−2,5,1).
8.30 y1=(x1+x2+x3)/√
3,y2=(x1−2x2+x3)/√
6,y3=(−x1+x3)/√
2;
µ1=6 , µ2=−6.
8.31 The null space is spanned by (2 0 1 0)Tand (1−201 )T.
8.32 (a) No, |A|=−24/negationslash=0 ;y e s , x:y:z=4ab:4a+b:4a−b.
8.33 x=3 , y=1 , z=2.
8.34 x1=−3/2,x2=7/2,x3=−3/2.
8.35 η=1 , x=1+2 z,y=−3z;η=2 , x=2z,y=1−3z.
8.36 (a) α/negationslash=6,α/negationslash=1 ; x1=( 1−α)/(1 +α),x2=2/(1−α),x3=0 .
(b)α=1 .( c ) α=6 ; x1=1−6β, x 2=β, x 3=( 7β−2)/3f o ra n y β.
8.37 L=( 1,0,0;1
3,1,0;2
3,3,1),U =( 3,6,9;0,−2,2; 0,0,4).
(i)x=(−112 )T. (ii) x=(−322 )T.
8.38 L=( 1,0,0,0;1
2,1,0,0;5
2,21
11,1,0;3
2,−3
11,−12
7,1);
U=( 2,−3,1,3; 0,11
2−7
2,−9
2;0,0,35
11,1
11;0,0,0,−32
7).
(i)x=( 2−14−5)T. (ii) x=(−114−3)T.
8.39 Ais not positive definite as L33is calculated to be√
−6.
B=LLT, where the non-zero elements of Lare
L11=√
5,L31=
p
3/5,L22=√
3,L33=
p
12/5.
8.40 λ3−27λ2+121 λ−3=0 .Find the two larger roots for λusing the rearrangement
method described in subsection 28.1.1 and the smallest one using the property ofthe product of the roots. The singular values are 4.6190, 2.3748 and 0.1579.
8.41
A
†A=
/
21
12
/
,U=1√
6
/0/@−1√
3√
2
20√
2
−1−√
3√
2
/1A,V=
/
11
1−1
/
.
8.42 The singular values are 18√
6,−18,−12√
3.
(i)x= ( 112 )Twith all four equations exactly satisfied.
(ii)x=1
36(40 37 74)T, giving a residual column matrix ( −1223 )T.
8.43 The singular values are 12√
6,0,−18√
3 and the calculated best solution is x=
1.71,y=−1.94,z=−1.71. The null space is the line x=z,y= 0 and the general
SVD solution is x=1.71 + λ, y=−1.94,z=−1.71 + λ.
321
9
Normal modes
Any student of the physical sciences will encounter the subject of oscillations on
many occasions and in a wide variety of circumstances, for example the voltageand current oscillations in an electric circuit, the vibrations of a mechanicalstructure and the internal motions of molecules. The matrices studied in the
previous chapter provide a particularly simple way to approach what may appear,
at first glance, to be difficult physical problems.
We will consider only systems for which a position-dependent potential exists,
i.e., the potential energy of the system in any particular configuration dependsupon the coordinates of the configuration, which need not be be lengths however;the potential must notdepend upon the time derivatives (generalised velocities) of
these coordinates. So, for example, the potential −qv·Aused in the Lagrangian
description of a charged particle in an electromagnetic field is excluded. A
further restriction that we place is that the potential has a local minimum atthe equilibrium point; physically, this is a necessary and sufficient condition forstable equilibrium. By suitably defining the origin of the potential, we may takeits value at the equilibrium point as zero.
We denote the coordinates chosen to describe a configuration of the system
byq
i,i=1,2,...,N .T h e qineed not be distances; some could be angles, for
example. For convenience we can define the qiso that they are all zero at the
equilibrium point. The instantaneous veloc ities of various parts of the system will
depend upon the time derivatives of the qi, denoted by ˙qi. For small oscillations
the velocities will be linear in the ˙qiand consequently the total kinetic energy T
will be quadratic in them – and will include cross terms of the form ˙qi˙qjwith
i/negationslash=j. The general expression for Tcan be written as the quadratic form
T=summationdisplay
isummationdisplay
jaij˙qi˙qj=˙qTA˙q, (9.1)
where ˙qis the column vector ( ˙q1˙q2···˙qN)Tand the N×Nmatrix A
is real and may be chosen to be symmetric. Furthermore, A, like any matrix
322
9.1 TYPICAL OSCILLATORY SYSTEMS
corresponding to a kinetic energy, is positive definite (more strictly positive semi-
definite); that is, whatever real values the ˙qitake, the quadratic form (9.1) has a
value≥0.
Turning now to the potential energy, we may write its value for a configuration
qby means of a Taylor expansion about the origin q=0,
V(q)=V(0)+summationdisplay
i∂V(0)
∂qiqi+1
2summationdisplay
isummationdisplay
j∂2V(0)
∂qi∂qjqiqj+···.
However, we have chosen V(0) = 0 and, since the origin is an equilibrium point,
there is no force there and ∂V(0)/∂q i= 0. Consequently, to second order in the
qiwe also have a quadratic form, but in the coordinates rather than in their time
derivatives:
V=summationdisplay
isummationdisplay
jbijqiqj=qTBq, (9.2)
where Bis, or can be made, symmetric. In this case, and in general, the requirement
that the potential is a minimum means that the potential matrix B, like the kinetic
energy matrix A, is real and positive definite.
9.1 Typical oscillatory systems
We now introduce particular examples, although the results of this section are
general, given the above restrictions and the reader will find it easy to apply theresults to many other instances.
Consider first a uniform rod of mass Mand length l, attached by a light string
also of length lto a fixed point Pand executing small oscillations in a vertical
plane. We choose as coordinates the angles θ
1andθ2shown, with exaggerated
magnitude, in figure 9.1. In terms of these coordinates the centre of gravity of the
rod has, to first order in the θi, a velocity component in the x-direction equal to
l˙θ1+1
2l˙θ2a n di nt h e y-direction equal to zero. Adding in the rotational kinetic
energy of the rod about its centre of gravity we obtain, to second order in the ˙θi,
T≈1
2Ml2(˙θ2
1+1
4˙θ2
2+˙θ1˙θ2)+1
24Ml2˙θ2
2
=1
6Ml2parenleftbig
3˙θ2
1+3˙θ1˙θ2+˙θ2
2parenrightbig
=1
12Ml2˙qTparenleftbigg63
32parenrightbigg
˙q, (9.3)
where ˙qT=(˙θ1˙θ2).The potential energy is given by
V=Mlgbracketleftbig
(1−cosθ1)+1
2(1−cosθ2)bracketrightbig
(9.4)
≈1
4Mlg(2θ2
1+θ2
2)=1
12Mlg qTparenleftbigg60
03parenrightbigg
q, (9.5)
where gis the acceleration due to gravity and q=(θ1θ2)T; (9.5) is valid to
second order in the θi.
323
NORMAL MODES
P P P
l
lθ1θ1
θ1
θ2θ2θ2
(a) (b) (c)
Figure 9.1 A uniform rod of length lattached to the fixed point Pby a light
string of the same length: ( a) the general coordinate system; ( b) approximation
to the normal mode with lower frequency; ( c) approximation to the mode with
higher frequency.
With these expressions for TandVwe now apply the conservation of energy,
d
dt(T+V)=0 , (9.6)
assuming that there are no external forces other than gravity. In matrix form
(9.6) becomes
d
dt(˙qTA˙q+qTBq)=¨qTA˙q+˙qTA¨q+˙qTBq+qTB˙q=0,
which, using A=ATand B=BT,g i v e s
2˙qT(A¨q+Bq)=0 .
We will assume, although it is not clear that this gives the only possible solution,
that the above equation implies that the coefficient of each ˙qiis separately zero.
Hence
A¨q+Bq=0. (9.7)
For a rigorous derivation Lagrange’s equations should be used, as in chapter 22.
N o ww es e a r c hf o rs e t so fc o o r d i n a t e s qthatalloscillate with the same period,
i.e. the total motion repeats itself exactly after a finiteinterval. Solutions of this
form will satisfy
q=xcosωt; (9.8)
the relative values of the elements of xin such a solution will indicate how each
324
9.1 TYPICAL OSCILLATORY SYSTEMS
coordinate is involved in this special motion. In general there will be Nvalues
ofωif the matrices Aand BareN×Nand these values are known as normal
frequencies oreigenfrequencies .
Putting (9.8) into (9.7) yields
−ω2Ax+Bx=(B−ω2A)x=0. (9.9)
Our work in section 8.18 showed that this can have non-trivial solutions only if
|B−ω2A|=0. (9.10)
This is a form of characteristic equation for B, except that the unit matrix Ihas
been replaced by A. It has the more familiar form if a choice of coordinates is
made in which the kinetic energy Tis a simple sum of squared terms, i.e. it has
been diagonalised, and the scale of the new coordinates is then chosen to make
each diagonal element unity.
However, even in the present case, (9.10) can be solved to yield ω2
kfork=
1,2,...,N ,w h e r e Nis the order of Aand B. The values of ωkcan be used
with (9.9) to find the corresponding column vector xkand the initial (stationary)
physical configuration that, on release, will execute motion with period 2 π/ω k.
In equation (8.76) we showed that the eigenvectors of a real symmetric matrix
were, except in the case of degeneracy of the eigenvalues, mutually orthogonal.
In the present situation an analogous, but not identical, result holds. It is shown
in section 9.3 that if x1and x2are two eigenvectors satisfying (9.9) for different
values of ω2then they are orthogonal in the sense that
(x2)TAx1=0 a n d ( x2)TBx1=0.
The direct ‘scalar product’ ( x2)Tx1, formally equal to ( x2)TIx1,i sn o t ,i ng e n e r a l ,
equal to zero.
Returning to the suspended rod, we find from (9.10)
vextendsinglevextendsinglevextendsinglevextendsingleMlg
12parenleftbigg60
03parenrightbigg
−ω2Ml2
12parenleftbigg63
32parenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle=0.
Writing ω2l/g=λ, this becomes
vextendsinglevextendsinglevextendsinglevextendsingle6−6λ−3λ
−3λ3−2λvextendsinglevextendsinglevextendsinglevextendsingle=0⇒ λ2−10λ+6=0 ,
which has roots λ=5±√
19. Thus we find that the two normal frequencies are
given by ω1=( 0.641g/l)1/2andω2=( 9.359g/l)1/2. Putting the lower of the two
values for ω2,n a m e l y( 5 −√
19)g/l, into (9.9) shows that for this mode
x1:x2=3 ( 5−√
19) : 6(√
19−4) = 1 .923 : 2 .153.
This corresponds to the case where the rod and string are almost straight out, i.e.
they almost form a simple pendulum. Similarly it may be shown that the higher
325
NORMAL MODES
frequency corresponds to a solution where the string and rod are moving with
opposite phase and x1:x2=9.359 :−16.718. The two situations are shown in
figure 9.1.
In connection with quadratic forms it was shown in section 8.17 how to make
a change of coordinates such that the matrix for a particular form becomesdiagonal. In exercise 9.6 a method is developed for diagonalising simultaneouslytwo quadratic forms (though the transformation matrix may not be orthogonal).
If this process is carried out for Aand Bin a general system undergoing stable
oscillations, the kinetic and potential energies in the new variables η
itake the
forms
T=summationdisplay
iµi˙η2
i=˙ηTM˙η, M=d i a g( µ1,µ2,...,µ N), (9.11)
V=summationdisplay
iνiη2
i=ηTNη, N=d i a g( ν1,ν2...,ν N), (9.12)
and the equations of motion are the uncoupled equations
µi¨ηi+νiηi=0,i=1,2,...,N. (9.13)
Clearly a simple renormalisation of the ηican be made that reduces all the µi
in (9.11) to unity. When this is done the variables so formed are called normal
coordinates and equations (9.13) the normal equations .
When a system is executing one of these simple harmonic motions it is said to
be in a normal mode , and once started in such a mode it will repeat its motion
exactly after each interval of 2 π/ω i. Any arbitrary motion of the system may
be written as a superposition of the normal modes, and each component modewill execute harmonic motion with the corresponding eigenfrequency; however,unless by chance the eigenfrequencies are in integer relationship, the system will
never return to its initial configuration after any finite time interval.
As a second example we will consider a number of masses coupled together by
springs. For this type of situation the potential and kinetic energies are automat-
ically quadratic functions of the coordinates and their derivatives, provided theelastic limits of the springs are not exceeded, and the oscillations do not have tobe vanishingly small for the analysis to be valid.IFind the normal frequencies and modes of oscillation of three particles of masses m,µm,
mconnected in that order in a straight line by two equal light springs of force constant k.
(This arrangement could serve as a model for some linear molecules, e.g. CO2.)
The situation is shown in figure 9.2; the coordinates of the particles, x1,x2,x3,a r e
measured from their equilibrium positions, at which the springs are neither extended norcompressed.
The kinetic energy of the system is simply
T=
1
2m
/;˙x2
1+µ˙x2
2+˙x2
3
/
,
326
9.1 TYPICAL OSCILLATORY SYSTEMS
m m µm
x1 x2 x3k k
Figure 9.2 Three masses m,µmandmconnected by two equal light springs
of force constant k.
(a)
(b)
(c)
Figure 9.3 The normal modes of the masses and springs of a linear molecule
such as CO 2.(a)ω2=0 ;( b)ω2=k/m;(c)ω2=[ (µ+2 )/µ](k/m).
whilst the potential energy stored in the springs is
V=1
2k
/
(x2−x1)2+(x3−x2)2
/
.
The kinetic- and potential-energy symmetric matrices are thus
A=m
2
/0/@100
0µ0
001
/1A, B=k
2
/0/@1−10
−12−1
0−11
/1A.
From (9.10), to find the normal frequencies we have to solve |B−ω2A|=0.Thus, writing
mω2/k=λ, we have//////1−λ−10
−12−µλ−1
0−11−λ
//////=0,
which leads to λ=0 ,1o r1+2 /µ. The corresponding eigenvectors are respectively
x1=1√
3
/0/@1
11
/1A, x2=1√
2
/0/@1
0
−1
/1A, x3=1p
2+( 4 /µ2)
/0/@1
−2/µ
1
/1A.
The physical motions associated with these normal modes are illustrated in figure 9.3.
The first, with λ=ω=0a n da l lt h e xiequal, merely describes bodily translation of the
whole system, with no (i.e. zero-frequency) internal oscillations.
In the second solution the central particle remains stationary, x2= 0, whilst the other
two oscillate with equal amplitudes in antiphase with each other. This motion, which has
frequency ω=(k/m)1/2, is illustrated in figure 9.3( b).
The final and most complicated of the three modes has frequency ω={[(µ+
327
NORMAL MODES
2)/µ](k/m)}1/2, and involves a motion of the central particle which is in antiphase with
that of the two outer ones and which has an amplitude 2 /µt i m e sa sg r e a t .I nt h i sm o t i o n
(see figure 9.3( c)) the two springs are compressed and extended in turn. We also note
that in the second and third normal modes the centre of mass of the molecule remainsstationary.J
9.2 Symmetry and normal modes
It will have been noticed that the system in the above example has an obvious
symmetry under the interchange of coordinates 1 and 3: the matrices Aand B,
the equations of motion and the normal modes illustrated in figure 9.3 are all
unaltered by the interchange of x1and−x3. This reflects the more general result
that for each physical symmetry possessed by a system, there is at least onenormal mode with the same symmetry.
The general question of the relationship between the symmetries possessed by
a physical system and those of its normal modes will be taken up more formallyin chapter 25 where the representation theory of groups is considered. However,we can show here how an appreciation of a system’s symmetry properties will
sometimes allow its normal modes to be guessed (and then verified), something
that is particularly helpful if the number of coordinates involved is greater thantwo and the corresponding eigenvalue equation (9.10) is a cubic or higher-degreepolynomial equation.
Consider the problem of determining the normal modes of a system consist-
ing of four equal masses Mat the corners of a square of side 2 L, each pair
of masses being connected by a light spring of modulus kthat is unstretched
in the equilibrium situation. As shown in figure 9.4, we introduce Cartesiancoordinates x
n,yn, with n=1,2,3,4, for the positions of the masses and de-
note their displacements from their equilibrium positions Rnbyqn=xni+ynj.
Thus
rn=Rn+qnwith Rn=±Li±Lj.
The coordinates for the system are thus x1,y1,x2,...,y 4and the kinetic en-
ergy matrix Ais given trivially by MI8,w h e r e I8is the 8×8 identity ma-
trix.
The potential energy matrix Bis much more difficult to calculate and involves,
for each pair of values m, n, evaluating the quadratic approximation to the
expression
bmn=1
2kparenleftbig
|rm−rn|−|Rm−Rn|parenrightbig2.
Expressing each riin terms of qiandRiand remembering that |Rm−Rn|/greatermuch
328
9.2 SYMMETRY AND NORMAL MODES
M MM M
k k
kk kk
x1y1
x2y2
x3y3
x4y4
Figure 9.4 The arrangement of four equal masses and six equal springs
discussed in the text. The coordinate systems xn,ynforn=1,2,3,4 measure
the displacements of the masses from their equilibrium positions.
|qm−qn|,w eo b t a i n bmn(=bnm):
bmn=1
2kbracketleftbig
|(Rm−Rn)+(qm−qn)|−|Rm−Rn|bracketrightbig2
=1
2kbraceleftBigbracketleftbig
|Rm−Rn|2+2 (qm−qn)·(RM−Rn)+|qm−qn)|2bracketrightbig1/2−|Rm−Rn|bracerightBig2
=1
2k|Rm−Rn|2braceleftBiggbracketleftbigg
1+2(qm−qn)·(RM−Rn)
|Rm−Rn|2+···bracketrightbigg1/2
−1bracerightBigg2
≈1
2kbraceleftbigg(qm−qn)·(RM−Rn)
|Rm−Rn|bracerightbigg2
.
This final expression is readily interpretable as the potential energy stored in the
spring when it is extended by an amount equal to the component, along theequilibrium direction of the spring, of the relative displacement of its two ends.
Applying this result to each spring in turn gives the following expressions for
the elements of the potential matrix.
mn 2b
mn/k
12 ( x1−x2)2
13 ( y1−y3)2
141
2(−x1+x4+y1−y4)2
231
2(x2−x3+y2−y3)2
24 ( y2−y4)2
34 ( x3−x4)2.
329
NORMAL MODES
The potential matrix is thus constructed as
B=k
4
3−1−20 0 0 −11
−1 3000 −21−1
−2 031 −1−10 0
0013 −1−10−2
00−1−13 1 −20
0−2−1−1 1300
−1 100 −20 3 −1
1−10−20 0 −13
.
To solve the eigenvalue equation |B−λA|= 0 directly would mean solving
an eigth-degree polynomial equation. Fortunately, we can exploit intuition andthe symmetries of the system to obtain the eigenvectors and correspondingeigenvalues without such labour.
Firstly, we know that bodily translation of the whole system, without any
internal vibration, must be possible and that there will be two independentsolutions of this form, corresponding to translations in the x-a n d y- directions.
The eigenvector for the first of these (written in row form to save space) is
x
(1)= ( 10101010 )T.
Evaluation of Bx(1)gives
Bx(1)= ( 00000000 )T,
showing that x(1)is a solution of ( B−ω2A)x=0corresponding to the eigenvalue
ω2=0 ,w h a t e v e rf o r m Axmay take. Similarly,
x(2)= ( 01010101 )T
is a second eigenvector corresponding to the eigenvalue ω2=0 .
The next intuitive solution, again involving no internal vibrations, and, there-
fore, expected to correspond to ω2= 0, is pure rotation of the whole system
about its centre. In this mode each mass moves perpendicularly to the line joining
its position to the centre, and so the relevant eigenvector is
x(3)=1√
2( 111 −1−11−1−1)T.
It is easily verified that Bx(3)=0thus confirming both the eigenvector and the
corresponding eigenvalue. The three non-oscillatory normal modes are illustrated
in diagrams ( a)–(c) of figure 9.5.
We now come to solutions that do involve real internal oscillations, and,
because of the four-fold symmetry of the system, we expect one of them to be amode in which all the masses move along radial lines – the so-called ‘breathing
330
9.2 SYMMETRY AND NORMAL MODES
(a)ω2=0 ( b)ω2=0 ( c)ω2=0 (d)ω2=2k/M
(e)ω2=k/M (f)ω2=k/M (g)ω2=k/M (h)ω2=k/M
Figure 9.5 The displacements and frequencies of the eight normal modes of
the system shown in figure 9.4. Modes ( a), (b)a n d( c) are not true oscillations:
(a)a n d( b) are purely translational whilst ( c) is one of bodily rotation.
Mode ( d), the ‘breathing mode’, has the highest frequency and the remaining
four, ( e)– (h), of lower frequency, are degenerate.
mode’. Expressing this motion in coordinate form gives as the fourth eigenvector
x(4)=1√
2(−1111 −1−11−1)T.
Evaluation of Bx(4)yields
Bx(4)=k
4√
2(−8888 −8−88−8)T=2kx(4),
i.e. a multiple of x(4), confirming that it is indeed an eigenvector. Further, since
Ax(4)=Mx(4), it follows from ( B−ω2A)x=0that ω2=2k/Mf o rt h i sn o r m a l
mode. Diagram (d) of the figure illustrates the corresponding motions of the fourmasses.
As the next step in exploiting of the symmetry properties of the system we
note that, because of its reflection symmetry in the x-axis, the system is invariant
under the double interchange of y
1with−y3andy2with−y4. This leads us to
try an eigenvector of the form
x(5)=( 0 α0β0−α0−β)T.
Substituting this trial vector into ( B−ω2A)x=0gives, of course, eight simulta-
331
NORMAL MODES
neous equations for αandβ, but they are all equivalent to just two, namely
α+β=0,
5α+β=4Mω2
kα;
these have the solution α=−βandω2=k/M. The latter thus gives the frequency
of the mode with eigenvector
x(5)=( 0 1 0 −10−101 )T.
Note that, in this mode, when the spring joining masses 1 and 3 is most stretched,
the one joining masses 2 and 4 is at its most compressed. Similarly, based onreflection symmetry in the y-axis,
x
(6)=( 1 0−10−1010 )T
can be shown to be an eigenvector corresponding to the same frequency. These
two modes are sketched in diagrams ( e)a n d( f) of figure 9.5.
This accounts for six of the expected eight modes, and the other two could be
found by considering motions that are symmetric about both diagonals of the
square or are invariant under successive reflections in the x-a n d y- axes. However,
since Ais a multiple of the unit matrix, and since we know that ( x(j))TAx(i)=0i f
i/negationslash=j, we can find the two remaining eigenvectors more easily by requiring them
to be orthogonal to each of those found so far.
Let us take the next (seventh) eigenvector, x(7),t ob eg i v e nb y
x(7)=(abcdefgh )T.
Then orthogonality with each of the x(n)forn=1,2,...,6 yields six equations
satisfied by the unknowns a ,b ,...,h . As the reader may verify, they can be reduced
to the six simple equations
a+g=0,d+f=0,a+f=d+g,
b+h=0,c+e=0,b+c=e+h.
With six homogeneous equations for eight unknowns, effectively separated into
two groups of four, we may pick one in each group arbitrarily. Taking a=b=1
gives d=e=1a n d c=f=g=h=−1 as a solution. Substitution of
x(7)=( 1 1−111−1−1−1)T.
into the eigenvalue equation checks that it is an eigenvector and shows that the
corresponding eigenfrequency is given by ω2=k/M.
We now have the eigenvectors for seven of the eight normal modeshand the
eighth can be found by making it simultaneously orthogonal to each of the otherseven. It is left to the reader to show (or verify) that the final solution is
x
(8)=( 1−111−1−1−11 )T
332
9.3 RAYLEIGH–RITZ METHOD
and that this mode has the same frequency as three of the other modes. The
general topic of the degeneracy of normal modes is discussed in chapter 25. Themovements associated with the final two modes are shown in diagrams ( g)a n d
(h) of figure 9.5; this figure summarises all eight normal modes and frequencies.
Although this example has been lengthy to write out, we have seen that the
actual calculations are quite simple and provide the full solution to what isformally a matrix eigenvalue equation involving 8 ×8 matrices. It should be
noted that our exploitation of the intrinsic symmetries of the system played acrucial part in finding the correct eigenvectors for the various normal modes.
9.3 Rayleigh–Ritz method
We conclude this chapter with a discussion of the Rayleigh–Ritz method for
estimating the eigenfrequencies of an oscillating system. We recall from theintroduction to this chapter that for a system undergoing small oscillations thepotential and kinetic energy are given by
V=q
TBq and T=˙qTA˙q,
where the components of qare the coordinates chosen to represent the configura-
tion of the system and Aand Bare symmetric matrices (or may be chosen to be
such). We also recall from (9.9) that the normal modes xiand the eigenfrequencies
ωiare given by
(B−ω2
iA)xi=0. (9.14)
It may be shown that the eigenvectors xicorresponding to different normal modes
are linearly independent and so form a complete set. Thus, any coordinate vectorqcan be written q=summationtext
jcjxj. We now consider the value of the generalised
quadratic form
λ(x)=xTBx
xTAx=summationtext
m(xm)Tc∗
mBsummationtext
icixi
summationtext
j(xj)Tc∗
jAsummationtext
kckxk,
which, since both numerator and denominator are positive definite, is itself non-
negative. Equation (9.14) can be used to replace Bxi, with the result that
λ(x)=summationtext
m(xm)Tc∗
mAsummationtext
iω2
icixi
summationtext
j(xj)Tc∗
jAsummationtext
kckxk
=summationtext
m(xm)Tc∗
msummationtext
iω2
iciAxi
summationtext
j(xj)Tc∗
jAsummationtext
kckxk. (9.15)
Now the eigenvectors xiobtained by solving ( B−ω2A)x=0are not mutually
orthogonal unless either AorBis a multiple of the unit matrix. However, it may
333
NORMAL MODES
be shown that they do possess the desirable properties
(xj)TAxi=0 a n d ( xj)TBxi=0 i f i/negationslash=j. (9.16)
This result is proved as follows. From (9.14) it is clear that, for general iandj,
(xj)T(B−ω2
iA)xi=0. (9.17)
But, by taking the transpose of (9.14) with ireplaced by jand recalling that A
and Bare real and symmetric, we obtain
(xj)T(B−ω2
jA)=0.
Forming the scalar product of this with xiand subtracting the result from (9.17)
gives
(ω2
j−ω2
i)(xj)TAxi=0.
Thus, for i/negationslash=jand non-degenerate eigenvalues ω2
iand ω2
j, we have that
(xj)TAxi= 0, and substituting this into (9.17) immediately establishes the corre-
sponding result for ( xj)TBxi. Clearly, if either AorBis a multiple of the unit
matrix then the eigenvectors are mutually orthogonal in the normal sense. Theorthogonality relations (9.16) are re-derived and extended in exercise 9.6.
Using the first of the relationships (9.16) to simplify (9.15), we find that
λ(x)=summationtext
i|ci|2ω2
i(xi)TAxi
summationtext
k|ck|2(xk)TAxk. (9.18)
Now, if ω2
0is the lowest eigenfrequency then ω2
i≥ω2
0for all iand, further, since
(xi)TAxi≥0f o ra l l ithe numerator of (9.18) is ≥ω2
0summationtext
i|ci|2(xi)TAxi.H e n c e
λ(x)≡xTBx
xTAx≥ω2
0, (9.19)
for any xwhatsoever (whether xis an eigenvector or not). Thus we are able to
estimate the lowest eigenfrequency of the system by evaluating λfor a variety
of vectors x, the components of which, it will be recalled, give the ratios of the
coordinate amplitudes. This is sometimes a useful approach if many coordinates
are involved and direct solution for the eigenvalues is not possible.
An additional result is that the maximum eigenfrequency ω2
mmay also be
estimated. It is obvious that if we replace the statement ‘ ω2
i≥ω2
0for all i’b y
‘ω2
i≤ω2
mfor all i’, then λ(x)≤ω2
mfor any x. Thus λ(x) always lies between
the lowest and highest eigenfrequencies of the system. Furthermore, λ(x)h a sa
stationary value, equal to ω2
k,w h e n xis the kth eigenvector (see subsection 8.17.1).
334
9.4 EXERCISESIEstimate the eigenfrequencies of the oscillating rod of section 9.1.
Firstly we recall that
A=Ml2
12
/
63
32
/
and B=Mlg
12
/
60
03
/
.
Physical intuition suggests that the slower mode will have a configuration approximating
that of a simple pendulum (figure 9.1), in which θ1=θ2, and so we use this as a trial
vector.T a k i n g x=(θθ)T,
λ(x)=xTBx
xTAx=3Mlgθ2/4
7Ml2θ2/6=9g
14l=0.643g
l,
and we conclude from (9.19) that the lower (angular) frequency is ≤(0.643g/l)1/2.W e
have already seen on p. 325 that the true answer is (0 .641g/l)1/2and so we have come
very close to it.
Next we turn to the higher frequency. Here, a typical pattern of oscillation is not so
obvious but, rather preempting the answer, we try θ2=−2θ1; we then obtain λ=9g/l
and so conclude that the higher eigenfrequency ≥(9g/l)1/2. We have already seen that the
exact answer is (9 .359g/l)1/2and so again we have come close to it.
J
A simplified version of the Rayleigh–Ritz method may be used to estimate the
eigenvalues of a symmetric (or in general Hermitian) matrix B, the eigenvectors
of which will be mutually orthogonal. By repeating the calculations leading to(9.18), Abeing replaced by the unit matrix I, it is easily verified that if
λ(x)=x
TBx
xTx
is evaluated for anyvector xthen
λ1≤λ(x)≤λm,
where λ1,λ2...,λ mare the eigenvalues of Bin order of increasing size. A similar
result holds for Hermitian matrices.
9.4 Exercises
9.1 Three coupled pendulums swing perpendicularly to the horizontal line containing
their points of suspension, and the following equations of motion are satisfied:
−m¨x1=cmx 1+d(x1−x2),
−M¨x2=cMx 2+d(x2−x1)+d(x2−x3),
−m¨x3=cmx 3+d(x3−x2),
where x1,x2andx3are measured from the equilibrium points, m,Mandm
are the masses of the pendulum bobs and canddare positive constants. Find
the normal frequencies of the system and sketch the corresponding patterns ofoscillation. What happens as d→0o rd→∞?
9.2 A double pendulum, smoothly pivoted at A, consists of two light rigid rods, AB
andBC, each of length l, which are smoothly jointed at Band carry masses mand
αmatBandCrespectively. The pendulum makes s mall oscillations in one plane
335
NORMAL MODES
under gravity; at time t,ABandBCmake angles θ(t)a n d φ(t) respectively with
the downward vertical. Find quadratic e xpressions for the kinetic and potential
energies of the system and hence show that the normal modes have angularfrequencies given by
ω
2=g
l
h
1+α±
p
α(1 +α)
i
.
Forα=1/3, show that in one of the normal modes the mid-point of BCdoes
not move during the motion.
9.3 Continue the worked example modelling a linear molecule discussed at the end
o fs e c t i o n9 . 1 ,f o rt h ec a s ei nw h i c h µ=2 .
(a) Show that the eigenvectors derived there have the expected orthogonality
properties with respect to both Aand B.
(b) For the situation in which the atoms are released from rest with initial
displacements x1=2/epsilon1,x2=−/epsilon1andx3= 0, determine their subsequent
motions and maximum displacements.
9.4 Consider the circuit consisting of three equal capacitors and two different induc-
tors shown in the figure. For charges Qion the capacitors and currents Iithrough
L1 L2C C
C
I1 I2Q1 Q2
Q3
the components, write down Kirchhoff’s law for the total voltage change around
each of two complete circuit loops. Note that, to within an unimportant constant,the conservation of current implies that Q
3=Q1−Q2and hence express the loop
equations in the form given in (9.7), namely
A¨Q+BQ=0.
Use this to show that the normal frequencies of the circuit are given by
ω2=1
CL1L2
/
L1+L2±(L2
1+L2
2−L1L2)1/2
/
.
Obtain the same matrices and result by finding the total energy stored in the
various capacitors (typically Q2/(2C)) and in the inductors (typically LI2/2).
For the special case L1=L2=Ldetermine the relevant eigenvectors and so
describe the patterns of current flow in the circuit.
9.5 It is shown in physics and engineering textbooks that circuits containing capaci-
tors and inductors can be analysed by replacing a capacitor of capacitance Cby a
‘complex impedance’ 1 /(iωC) and an inductor of inductance Lby an impedance
iωL,w h e r e ωis the angular frequency of the currents flowing and i2=−1.
Use this approach and Kirchhoff’s circuit laws to analyse the circuit shown in
336
9.4 EXERCISES
the figure and obtain three linear equations governing the currents I1,I2andI3.
Show that the only possible frequencies of self-sustaining currents satisfy either
L L
C CCI1
I2 I3P Q
R STU
(a)ω2LC=1o r( b )3 ω2LC= 1. Find the corresponding current patterns and,
in each case, by identifying parts of the circuit in which no current flows, draw
an equivalent circuit that contains only one capacitor and one inductor.
9.6 The simultaneous reduction to diagonal form of two real symmetric quadratic forms.
Consider the two real symmetric quadratic forms uTAuand uTBu,w h e r e uT
stands for the row matrix ( xyz ), and denote by unthose column matrices
that satisfy
Bun=λnAun, (E9.1)
in which nis a label and the λnare real, non-zero and all different.
(a) By multiplying (E9.1) on the left by ( um)Tand the transpose of the corre-
sponding equation for umon the right by un, show that ( um)TAun=0f o r
n/negationslash=m.
(b) By noting that Aun=(λn)−1Bun, deduce that ( um)TBun=0f o r m/negationslash=n.
It can be shown that the unare linearly independent; the next step is to
construct a matrix Pwhose columns are the vectors un.
(c) Make a change of variables u=Pvsuch that uTAubecomes vTCv,a n d uTBu
becomes vTDv. Show that Cand Dare diagonal by showing that cij=0i f
i/negationslash=jand similarly for dij.
Thus u=Pvorv=P−1ureduces both quadratics to diagonal form.
To summarise, the method is as follows:
(a) find the λnthat allow (E9.1) a non-zero solution, by solving |B−λA|=0 ;
(b) for each λnconstruct un;
(c) construct the non-singular matrix Pwhose columns are the vectors un;
(d) make the change of variable u=Pv.
9.7 ( It is recommended that the reader does not attempt this question until exercise 9.6
has been studied .)
If, in the pendulum system studied in section 9.1, the string is replaced by a
second rod identical to the first then the expressions for the kinetic energy Tand
the potential energy Vbecome (to second order in the θi)
T≈Ml2
/;8
3˙θ2
1+2˙θ1˙θ2+2
3˙θ2
2
/
,
V≈Mgl
/;3
2θ2
1+1
2θ2
2
/
.
Determine the normal frequencies of the system and find new variables ξandη
that will reduce these two expressions to diagonal form, i.e. to
a1˙ξ2+a2˙η2and b1ξ2+b2η2.
337
NORMAL MODES
9.8 ( It is recommended that the reader does not attempt this question until exercise 9.6
has been studied .)
Find a real linear transformation that simultaneously reduces the quadratic
forms
3x2+5y2+5z2+2yz+6zx−2xy,
5x2+1 2y2+8yz+4zx
to diagonal form.
9.9 Three particles of mass mare attached to a light horizontal string having fixed
ends, the string being thus divided into four equal portions of length aeach
under a tension T. Show that for small transverse vibrations the amplitudes xi
of the normal modes satisfy Bx=(maω2/T)x,w h e r e Bis the matrix/0/@2−10
−12−1
0−12
/1A.
Estimate the lowest and highest eigenfrequencies using trial vectors (343 )T
and(3−43)T. Use also the exact vectors
/
1√
21
/T
and
/
1−√
21
/T
and compare the results.
9.10 Use the Rayleigh–Ritz method to estimate the lowest oscillation frequency of a
heavy chain of Nlinks, each of length a(=L/N), which hangs freely from one
end. (Try simple calculable configurations such as all links but one vertical, orall links collinear, etc.)
9.5 Hints and answers
9.1 See figure 9.6.
9.2 K.E. = (1 /2)ml2[(1 + α)˙θ2+α˙φ2+2α˙θ˙φ]; P.E. = (1 /2)mgl[(1 + α)θ2+αφ2]. For
α=1/3a n d ω=
p
2g/l,φ =−2θand the mid-point of BCremains vertically
below A.
9.3 (b) x1=/epsilon1(cosωt+c o s√2ωt),x2=−/epsilon1cos√2ωt,x3=/epsilon1(−cosωt+c o s√2ωt).
At various times the three displacements will reach 2 /epsilon1, /epsilon1,2/epsilon1respectively. For exam-
ple,x1c a nb ew r i t t e na s2 /epsilon1cos[(√
2−1)ωt/2]cos[(√
2+1) ωt/2], i.e. an oscillation
of angular frequency (√
2+1) ω/2 and modulated amplitude 2 /epsilon1cos[(√
2−1)ω/2];
the amplitude will reach 2 /epsilon1after a time ≈4π/[ω(√
2−1)].
9.4 Taking separate loops in the left-h and and right-hand sides of the diagram
the relevant matrices are A=(L1,0; 0,L2)a n d B=( 2C−1,−C−1;−C−1,2C−1).
Whatever the loop choice, ω2must satisfy L1L2C2ω4−2(L1+L2)Cω2+3=0 ,
which leads to the stated result. The energy stored in the central capacitor is(Q
1−Q2)2/(2C). IfL1=L2=Lthen one mode has ω2=(LC)−1and no current
flows through the central capacitor. The other mode has ω2=3 (LC)−1;i nt h i s
mode equal currents I(one clockwise, one anticlockwise) flow in the two loops
and therefore the current through the central capacitor is 2 I.
9.5 As the circuit loops contain no voltage sources the equations are homogeneous
and so for a non-trivial solution the determinant of coefficients must vanish.(a)I
1=0 , I2=−I3; no current in PQ; capacitance C/2 and inductance 2 L.
(b)I1=−2I2=−2I3; no current in TU; capacitance 3 C/2 and inductance 2 L.
9.6 (a) Obtain ( λ(n)−λ(m))(u(m))TAu(n)=0 ;( c ) cij=(PTAP)ij=(PT)ikAklPlj=
u(i)
kAklu(j)
l=(u(i))TAu(j)=0f o r i/negationslash=j.
338
9.5 HINTS AND ANSWERS
1 2 3
m m M
2kmkM kM(a)ω2=c+d
m
(b)ω2=c
(c)ω2=c+2d
M+d
m
Figure 9.6 The normal modes, as viewed from above, of the coupled pendu-
lums in example 9.1.
9.7 ω=( 2.634g/l)1/2or (0 .3661g/l)1/2;θ1=ξ+η,θ2=1.431ξ−2.097η.
9.8 λ=−1,2,4;x=2ξ−2η+2χ,y=ξ+η+χ,z=−3ξ+η−χ.
9.9 Estimated, 10 /17<M a ω2/T < 58/17; exact, 2 −√
2≤Maω2/T≤2+√
2.
9.10 The collinear case gives the best estimate, ω2≤6n2g/(4n3a)≈3g/(2l).
339
10
Vector calculus
In chapter 7 we discussed the algebra of vectors, and in chapter 8 we considered
how to transform one vector into another using a linear operator. In this chapterand the next we discuss the calculus of vectors, i.e. the differentiation and
integration both of vectors describing particular bodies, such as the velocity of
a particle, and of vector fields, in which a vector is defined as a function of thecoordinates throughout some volume (one-, two- or three-dimensional). Since theaim of this chapter is to develop methods for handling multi-dimensional physicalsituations, we will assume throughout that the functions with which we have todeal have sufficiently amenable mathematical properties, in particular that theyare continuous and differentiable.
10.1 Differentiation of vectors
L e tu sc o n s i d e rav e c t o r athat is a function of a scalar variable u.B yt h i s
we mean that with each value of uwe associate a vector a(u). For example, in
Cartesian coordinates a(u)=a
x(u)i+ay(u)j+az(u)k,w h e r e ax(u),ay(u)a n d az(u)
a r es c a l a rf u n c t i o n so f uand are the components of the vector a(u)i nt h e x-,y-
andz- directions respectively. We note that if a(u) is continuous at some point
u=u0then this implies that each of the Cartesian components ax(u),ay(u)a n d
az(u) is also continuous there.
Let us consider the derivative of the vector function a(u) with respect to u.
The derivative of a vector function is defined in a similar manner to the ordinaryderivative of a scalar function f(x) given in chapter 2. The small change in
the vector a(u) resulting from a small change ∆ uin the value of uis given by
∆a=a(u+∆u)−a(u) (see figure 10.1). The derivative of a(u)w i t hr e s p e c tt o uis
defined to be
da
du= lim
∆u→0a(u+∆u)−a(u)
∆u, (10.1)
340
10.1 DIFFERENTIATION OF VECTORS
a(u)a(u+∆u)∆a=a(u+∆u)−a(u)
Figure 10.1 A small change in a vector a(u) resulting from a small change
inu.
assuming that the limit exists, in which case a(u) is said to be differentiable at
that point. Note that da/duis also a vector, which is not, in general, parallel to
a(u). In Cartesian coordinates, the derivative of the vector a(u)=axi+ayj+azk
is given by
da
du=dax
dui+day
duj+daz
duk.
Perhaps the simplest application of the above is to finding the velocity and
acceleration of a particle in classical mechanics. If the time-dependent position
vector of the particle with respect to the origin in Cartesian coordinates is given
byr(t)=x(t)i+y(t)j+z(t)kthen the velocity of the particle is given by the vector
v(t)=dr
dt=dx
dti+dy
dtj+dz
dtk.
The direction of the velocity vector is along the tangent to the path r(t)a tt h e
instantaneous position of the particle, and its magnitude |v(t)|is equal to the
speed of the particle. The acceleration of the particle is given in a similar mannerby
a(t)=dv
dt=d2x
dt2i+d2y
dt2j+d2z
dt2k.IThe position vector of a particle at time tin Cartesian coordinates is given by r(t)=
2t2i+( 3t−2)j+( 3t2−1)k. Find the speed of the particle at t=1and the component of
its acceleration in the direction s=i+2j+k.
The velocity and acceleration of the particle are given by
v(t)=dr
dt=4ti+3j+6tk,
a(t)=dv
dt=4i+6k.
341
VECTOR CALCULUS
y
xφˆeφ
ˆeρ
ρij
Figure 10.2 Unit basis vectors for two-dimensional Cartesian and plane polar
coordinates.
The speed of the particle at t= 1 is simply
|v(1)|=
p
42+32+62=√
61.
The acceleration of the particle is constant (i.e. independent of t), and its component in
the direction sis given by
a·ˆs=(4i+6k)·(i+2j+k)√
12+22+12=5√
6
3.
J
Note that in the case discussed above i,jandkare fixed, time-independent
basis vectors. This may not be true of basis vectors in general; when we arenot using Cartesian coordinates the basis vectors themselves must also be dif-ferentiated. We discuss basis vectors for non-Cartesian coordinate systems indetail in section 10.10. Nevertheless, as a simple example, let us now considertwo-dimensional plane polar coordinates ρ, φ.
Referring to figure 10.2, imagine holding φfixed and moving radially outwards,
i.e. in the direction of increasing ρ. Let us denote the unit vector in this direction
byˆe
ρ. Similarly, imagine keeping ρfixed and moving around a circle of fixed radius
in the direction of increasing φ. Let us denote the unit vector tangent to the circle
byˆeφ. The two vectors ˆeρandˆeφare the basis vectors for this two-dimensional
coordinate system, just as iandjare basis vectors for two-dimensional Cartesian
coordinates. All these basis vectors are shown in figure 10.2.
An important difference between the two sets of basis vectors is that, while
iandjare constant in magnitude and direction , the vectors ˆeρandˆeφhave
constant magnitudes but their directions change as ρand φvary. Therefore,
when calculating the derivative of a vector written in polar coordinates we mustalso differentiate the basis vectors. One way of doing this is to express ˆe
ρandˆeφ
342
10.1 DIFFERENTIATION OF VECTORS
in terms of iandj. From figure 10.2, we see that
ˆeρ=c o s φi+s i n φj,
ˆeφ=−sinφi+c o s φj.
Since iandjare constant vectors, we find that the derivatives of the basis vectors
ˆeρandˆeφwith respect to tare given by
dˆeρ
dt=−sinφdφ
dti+c o s φdφ
dtj=˙φˆeφ, (10.2)
dˆeφ
dt=−cosφdφ
dti−sinφdφ
dtj=−˙φˆeρ, (10.3)
where the overdot is the conventional notation for differentiation with respect to
time.IThe position vector of a particle in plane polar coordinates is r(t)=ρ(t)ˆeρ.F i n de x p r e s -
sions for the velocity and acceleration of the particle in these coordinates.
Using result (10.4) below, the velocity of the particle is given by
v(t)=˙r(t)=˙ρˆeρ+ρ˙ˆeρ=˙ρˆeρ+ρ˙φˆeφ,
where we have used (10.2). In a similar way its acceleration is given by
a(t)=d
dt(˙ρˆeρ+ρ˙φˆeφ)
=¨ρˆeρ+˙ρ˙ˆeρ+ρ˙φ˙ˆeφ+ρ¨φˆeφ+˙ρ˙φˆeφ
=¨ρˆeρ+˙ρ(˙φˆeφ)+ρ˙φ(−˙φˆeρ)+ρ¨φˆeφ+˙ρ˙φˆeφ
=(¨ρ−ρ˙φ2)ˆeρ+(ρ¨φ+2˙ρ˙φ)ˆeφ.
J
Here we have used (10.2) and (10.3).
10.1.1 Differentiation of composite vector expressions
In composite vector expressions each of the vectors or scalars involved may be
a function of some scalar variable u, as we have seen. The derivatives of such
expressions are easily found using the definition (10.1) and the rules of ordinarydifferential calculus. They may be summarised by the following, in which we
assume that aandbare differentiable vector functions of a scalar uand that φ
is a differentiable scalar function of u:
d
du(φa)=φda
du+dφ
dua, (10.4)
d
du(a·b)=a·db
du+da
du·b, (10.5)
d
du(a×b)=a×db
du+da
du×b; (10.6)
343
VECTOR CALCULUS
the order of the factors in the terms on the RHS of (10.6) is, of course, just as
important as it is in the original vector product.IAp a r t i c l eo fm a s s mwith position vector rrelative to some origin Oexperiences a force
F, which produces a torque (moment) T=r×Fabout O. The angular momentum of the
particle about Ois given by L=r×mv,w h e r e vis the particle’s velocity. Show that the
rate of change of angular momentum is equal to the applied torque.
The rate of change of angular momentum is given by
dL
dt=d
dt(r×mv).
Using (10.6) we obtain
dL
dt=dr
dt×mv+r×d
dt(mv)
=v×mv+r×d
dt(mv)
=0+r×F=T,
where in the last line we use Newton’s second law, namely F=d(mv)/dt.
J
If a vector ais a function of a scalar variable sthat is itself a function of u,s o
thats=s(u), then the chain rule (see subsection 2.1.3) gives
da(s)
du=ds
duda
ds. (10.7)
The derivatives of more complicated vector expressions may be found by repeated
application of the above equations.
One further useful result can be derived by considering the derivative
d
du(a·a)=2a·da
du;
sincea·a=a2,w h e r e a=|a|,w es e et h a t
a·da
du=0 i f ais constant. (10.8)
In other words, if a vector a(u) has a constant magnitude as uvaries then it is
perpendicular to the vector da/du.
10.1.2 Differential of a vector
As a final note on the differentiation of vectors, we can also define the differential
of a vector, in a similar way to that of a scalar in ordinary differential calculus.
In the definition of the vector derivative (10.1), we used the notion of a small
change ∆ ain a vector a(u) resulting from a small change ∆ uin its argument. In
the limit ∆ u→0, the change in abecomes infinitesimally small, and we denote it
by the differential da. From (10.1) we see that the differential is given by
da=da
dudu. (10.9)
344
10.2 INTEGRATION OF VECTORS
Note that the differential of a vector is also a vector. As an example, the
infinitesimal change in the position vector of a particle in an infinitesimal time dtis
dr=dr
dtdt=vdt,
where vis the particle’s velocity.
10.2 Integration of vectors
The integration of a vector (or of an expression involving vectors that may itself
be either a vector or scalar) with respect to a scalar uc a nb er e g a r d e da st h e
inverse of differentiation. We must remember, however, that
(i) the integral has the same nature (vector or scalar) as the integrand,
(ii) the constant of integration for indefinite integrals must be of the same
nature as the integral.
For example, if a(u)=d[A(u)]/duthen the indefinite integral of a(u)i sg i v e nb y
integraldisplay
a(u)du=A(u)+b,
where bis a constant vector. The definite integral of a(u)f r o m u=u1tou=u2
is given by
integraldisplayu2
u1a(u)du=A(u2)−A(u1).IA small particle of mass morbits a much larger mass Mcentred at the origin O. According
to Newton’s law of gravitation, the position vector rof the small mass obeys the differential
equation
md2r
dt2=−GMm
r2ˆr.
Show that the vector r×dr/dtis a constant of the motion.
Forming the vector product of the differential equation with r,w eo b t a i n
r×d2r
dt2=−GM
r2r׈r.
Since randˆrare collinear, r׈r=0and therefore we have
r×d2r
dt2=0. (10.10)
However,
d
dt
/
r×dr
dt
/
=r×d2r
dt2+dr
dt×dr
dt=0,
345
VECTOR CALCULUS
xz
yC
Oˆbˆtˆn
P
r(u)
Figure 10.3 The unit tangent ˆt,n o r m a l ˆnand binormal ˆbto the space curve C
at a particular point P.
since the first term is zero by (10.10), and the second is zero because it is the vector product
of two parallel (in this case identical) vectors. Integrating, we obtain the required result
r×dr
dt=c, (10.11)
where cis a constant vector.
As a further point of interest we may note that in an infinitesimal time dtthe change
in the position vector of the small mass is drand the element of area swept out by the
position vector of the particle is simply dA=1
2|r×dr|. Dividing both sides of this equation
bydt, we conclude that
dA
dt=1
2
////r×dr
dt
////=|c|
2,
and that the physical interpretation of the above result (10.11) is that the position vector r
of the small mass sweeps out equal areas in equal times. This result is in fact valid formotion under any force that acts along the line joining the two particles.J
10.3 Space curves
In the previous section we mentioned that the velocity vector of a particle is a
tangent to the curve in space along which the particle moves. We now give a morecomplete discussion of curves in space and also of the geometrical interpretationof the vector derivative.
Ac u r v e Cin space can be described by the vector r(u) joining the origin Oof
a coordinate system to a point on the curve (see figure 10.3). As the parameter u
varies, the end-point of the vector moves along the curve. In Cartesian coordinates,
r(u)=x(u)i+y(u)j+z(u)k,
where x=x(u),y=y(u)a n d z=z(u)a r et h e parametric equations of the curve.
346
10.3 SPACE CURVES
This parametric representation can be very useful, particularly in mechanics when
the parameter may be the time t. We can, however, also represent a space curve
byy=f(x),z=g(x), which can be easily converted into the above parametric
form by setting u=x,s ot h a t
r(u)=ui+f(u)j+g(u)k.
Alternatively, a space curve can be represented in the form F(x, y, z)=0 ,
G(x, y, z) = 0, where each equation represents a surface and the curve is the
intersection of the two surfaces.
A curve may sometimes be described in parametric form by the vector r(s),
where the parameter sis the arc length along the curve measured from a fixed
point. Even when the curve is expressed in terms of some other parameter, it isstraightforward to find the arc length between any two points on the curve. Forthe curve described by r(u), let us consider an infinitesimal vector displacement
dr=dxi+dyj+dzk
along the curve. The square of the infinitesimal distance moved is then given by
(ds)
2=dr·dr=(dx)2+(dy)2+(dz)2,
f r o mw h i c hi tc a nb es h o w nt h a t
parenleftbiggds
duparenrightbigg2
=dr
du·dr
du.
Therefore, the arc length between two points on the curve r(u), given by u=u1
andu=u2,i s
s=integraldisplayu2
u1radicalbigg
dr
du·dr
dudu. (10.12)IAc u r v el y i n gi nt h e xy-plane is given by y=y(x),z=0. Using (10.12), show that the
arc length along the curve between x=aandx=bis given by s=
Rb
a
p
1+y/prime2dx,w h e r e
y/prime=dy/dx .
Let us first represent the curve in parametric form by setting u=x,s ot h a t
r(u)=ui+y(u)j.
Differentiating with respect to u, we find
dr
du=i+dy
duj,
from which we obtain
dr
du·dr
du=1+
/dy
du
/2
.
347
VECTOR CALCULUS
Therefore, remembering that u=x, from (10.12) the arc length between x=aandx=b
is given by
s=
Zb
a
r
dr
du·dr
dudu=
Zb
a
s
1+
/dy
dx
/2
dx.
This result was derived using more elementary methods in chapter 2.
J
If a curve Cis described by r(u) then, by considering figures 10.1 and 10.3, we
see that, at any given point on the curve, dr/duis a vector tangent to Cat that
point, in the direction of increasing u. In the special case where the parameter u
is the arc length salong the curve then dr/dsis aunittangent vector to Cand is
denoted by ˆt.
The rate at which the unit tangent ˆtchanges with respect to sis given by
dˆt/ds, and its magnitude is defined as the curvature κof the curve Cat a given
point,
κ=vextendsinglevextendsinglevextendsinglevextendsingledˆt
dsvextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingled
2ˆr
ds2vextendsinglevextendsinglevextendsinglevextendsingle.
We can also define the quantity ρ=1/κ, which is called the radius of curvature .
Since ˆtis of constant (unit) magnitude, it follows from (10.8) that it is perpen-
dicular to dˆt/ds. The unit vector in the direction perpendicular to ˆtis denoted
byˆnand is called the principal normal at the point. We therefore have
dˆt
ds=κˆn. (10.13)
The unit vector ˆb=ˆt׈n, which is perpendicular to the plane containing ˆt
andˆn, is called the binormal toC. The vectors ˆt,ˆnandˆbform a right-handed
rectangular cooordinate system (or triad) at any given point on C(see figure 10.3).
Asschanges so that the point of interest moves along C, the triad of vectors also
changes.
The rate at which ˆbchanges with respect to sis given by dˆb/dsand is a
measure of the torsion τof the curve at any given point. Since ˆbis of constant
magnitude, from (10.8) it is perpendicular to dˆb/ds. We may further show that
dˆb/dsis also perpendicular to ˆt, as follows. By definition ˆb·ˆt=0 ,w h i c ho n
differentiating yields
0=d
dsparenleftBig
ˆb·ˆtparenrightBig
=dˆb
ds·ˆt+ˆb·dˆt
ds
=dˆb
ds·ˆt+ˆb·κˆn
=dˆb
ds·ˆt,
where we have used the fact that ˆb·ˆn= 0. Hence, since dˆb/dsis perpendicular
to both ˆbandˆt, we must have dˆb/ds∝ˆn. The constant of proportionality is −τ,
348
10.3 SPACE CURVES
so we finally obtain
dˆb
ds=−τˆn. (10.14)
Taking the dot product of each side with ˆn, we see that the torsion of a curve is
given by
τ=−ˆn·dˆb
ds.
We may also define the quantity σ=1/τ, which is called the radius of torsion .
Finally, we consider the derivative dˆn/ds.S i n c e ˆn=ˆb׈twe have
dˆn
ds=dˆb
ds׈t+ˆb×dˆt
ds
=−τˆn׈t+ˆb×κˆn
=τˆb−κˆt. (10.15)
In summary, ˆt,ˆnandˆband their derivatives with respect to sare related to one
another by the relations (10.13), (10.14) and (10.15), the Frenet–Serret formulae ,
dˆt
ds=κˆn,dˆn
ds=τˆb−κˆt,dˆb
ds=−τˆn. (10.16)IShow that the acceleration of a particle travelling along a trajectory r(t)is given by
a(t)=dv
dtˆt+v2
ρˆn,
where vis the speed of the particle, ˆtis the unit tangent to the trajectory, ˆnis its principal
normal and ρis its radius of curvature.
The velocity of the particle is given by
v(t)=dr
dt=dr
dsds
dt=ds
dtˆt,
where ds/dt is the speed of the particle, which we denote by v,a n d ˆtis the unit vector
tangent to the trajectory. Writing the velocity as v=vˆt, and differentiating once more
with respect to time t,w eo b t a i n
a(t)=dv
dt=dv
dtˆt+vdˆt
dt;
but we note that
dˆt
dt=ds
dtdˆt
ds=vκˆn=v
ρˆn.
Therefore, we have
a(t)=dv
dtˆt+v2
ρˆn.
This shows that in addition to an acceleration dv/dt along the tangent to the particle’s
trajectory, there is also an acceleration v2/ρin the direction of the principal normal. The
latter is often called the centripetal acceleration.
J
349
VECTOR CALCULUS
Finally, we note that a curve r(u) representing the trajectory of a particle may
sometimes be given in terms of some parameter uthat is not necessarily equal to
the time tbut is functionally related to it in some way. In this case the velocity
of the particle is given by
v=dr
dt=dr
dudu
dt.
Differentiating again with respect to time gives the acceleration as
a=dv
dt=d
dtparenleftbiggdr
dudu
dtparenrightbigg
=d2r
du2parenleftbiggdu
dtparenrightbigg2
+dr
dud2u
dt2.
10.4 Vector functions of several arguments
The concept of the derivative of a vector is easily extended to cases where the
vectors (or scalars) are functions of more than one independent scalar variable,u
1,u2,...,u n. In this case, the results of subsection 10.1.1 are still valid, except
that the derivatives become partial derivatives ∂a/∂u idefined as in ordinary
differential calculus. For example, in Cartesian coordinates,
∂a
∂u=∂ax
∂ui+∂ay
∂uj+∂az
∂uk.
In particular, (10.7) generalises to the chain rule of partial differentiation discussed
in section 5.5. If a=a(u1,u2,...,u n) and each of the uiis also a function
ui(v1,v2,...,v n) of the variables vithen, generalising (5.17),
∂a
∂vi=∂a
∂u1∂u1
∂vi+∂a
∂u2∂u2
∂vi+···+∂a
∂un∂un
∂vi=nsummationdisplay
j=1∂a
∂uj∂uj
∂vi. (10.17)
A special case of this rule arises when ais an explicit function of some variable
v,a sw e l la so fs c a l a r s u1,u2,...,u nthat are themselves functions of v;t h e nw e
have
da
dv=∂a
∂v+nsummationdisplay
j=1∂a
∂uj∂uj
∂v. (10.18)
We may also extend the concept of the differential of a vector given in (10.9)
to vectors dependent on several variables u1,u2,...,u n:
da=∂a
∂u1du1+∂a
∂u2du2+···+∂a
∂undun=nsummationdisplay
j=1∂a
∂ujduj. (10.19)
As an example, the infinitesimal change in an electric field Ein moving from a
position rto a neighbouring one r+dris given by
dE=∂E
∂xdx+∂E
∂ydy+∂E
∂zdz. (10.20)
350
10.5 SURFACES
xyz
S
OT
P
r(u, v)∂r
∂u
∂r
∂v
v=c2u=c1
Figure 10.4 The tangent plane Tto a surface Sat a particular point P;
u=c1andv=c2are the coordinate curves, shown by dotted lines, that pass
through P. The broken line shows some particular parametric curve r=r(λ)
lying in the surface.
10.5 Surfaces
As u r f a c e Sin space can be described by the vector r(u, v) joining the origin Oof
a coordinate system to a point on the surface (see figure 10.4). As the parametersuandvvary, the end-point of the vector moves over the surface. This is very
similar to the parametric representation r(u) of a curve, discussed in section 10.3,
but with the important difference that we require twoparameters to describe a
surface, whereas we need only one to describe a curve.
In Cartesian coordinates the surface is given by
r(u, v)=x(u, v)i+y(u, v)j+z(u, v)k,
where x=x(u, v),y=y(u, v)a n d z=z(u, v) are the parametric equations of the
surface. We can also represent a surface by z=f(x, y)o rg(x, y, z)=0 .E i t h e r
of these representations can be converted into the parametric form in a similarmanner to that used for equations of curves. For example, if z=f(x, y)t h e nb y
setting u=xandv=ythe surface can be represented in parametric form by
r(u, v)=ui+vj+f(u, v)k.
Any curve r(λ), where λis a parameter, on the surface Scan be represented
by a pair of equations relating the parameters uandv, for example u=f(λ)
andv=g(λ). A parametric representation of the curve can easily be found by
straightforward substitution, i.e. r(λ)=r(u(λ),v(λ)). Using (10.17) for the case
where the vector is a function of a single variable λso that the LHS becomes a
351
VECTOR CALCULUS
total derivative, the tangent to the curve r(λ) at any point is given by
dr
dλ=∂r
∂udu
dλ+∂r
∂vdv
dλ. (10.21)
The two curves u=c o n s t a n ta n d v= constant passing through any point P
onSare called coordinate curves .F o rt h ec u r v e u= constant, for example, we
have du/dλ = 0, and so from (10.21) its tangent vector is in the direction ∂r/∂v.
Similarly, the tangent vector to the curve v= constant is in the direction ∂r/∂u.
If the surface is smooth then at any point PonSthe vectors ∂r/∂uand
∂r/∂vare linearly independent and define the tangent plane Tat the point P(see
figure 10.4). A vector normal to the surface at Pis given by
n=∂r
∂u×∂r
∂v. (10.22)
In the neighbourhood of P, an infinitesimal vector displacement dris written
dr=∂r
∂udu+∂r
∂vdv.
Theelement of area atP, an infinitesimal parallelogram whose sides are the
coordinate curves, has magnitude
dS=vextendsinglevextendsinglevextendsinglevextendsingle∂r
∂udu×∂r
∂vdvvextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingle∂r
∂u×∂r
∂vvextendsinglevextendsinglevextendsinglevextendsingledu dv=|n|du dv. (10.23)
Thus the total area of the surface is
A=integraldisplayintegraldisplay
Rvextendsinglevextendsinglevextendsinglevextendsingle∂r
∂u×∂r
∂vvextendsinglevextendsinglevextendsinglevextendsingledu dv=integraldisplayintegraldisplay
R|n|du dv, (10.24)
where Ris the region in the uv-plane corresponding to the range of parameter
values that define the surface.IFind the element of area on the surface of a sphere of radius a, and hence calculate the
total surface area of the sphere.
We can represent a point ron the surface of the sphere in terms of the two parameters θ
andφ:
r(θ,φ)=asinθcosφi+asinθsinφj+acosθk,
where θandφare the polar and azimuthal angles respectively. At any point P, vectors
tangent to the coordinate curves θ=c o n s t a n ta n d φ= constant are
∂r
∂θ=acosθcosφi+acosθsinφj−asinθk,
∂r
∂φ=−asinθsinφi+asinθcosφj.
352
10.6 SCALAR AND VECTOR FIELDS
An o r m a l nto the surface at this point is then given by
n=∂r
∂θ×∂r
∂φ=
///////ij k
acosθcosφa cosθsinφ−asinθ
−asinθsinφasinθcosφ 0
///////
=a2sinθ(sinθcosφi+s i n θsinφj+c o s θk),
which has a magnitude of a2sinθ. Therefore, the element of area at Pis, from (10.23),
dS=a2sinθd θd φ ,
and the total surface area of the sphere is given by
A=
Zπ
0dθ
Z2π
0dφ a2sinθ=4πa2.
This familiar result can, of course, be proved by much simpler methods!
J
10.6 Scalar and vector fields
We now turn to the case where a particular scalar or vector quantity is defined
not just at a point in space but continuously as a fieldthroughout some region
of space R(which is often the whole space). Although the concept of a field is
valid for spaces with an arbitrary number of dimensions, in the remainder of thischapter we will restrict our attention to the familiar three-dimensional case. Ascalar field φ(x, y, z) associates a scalar with each point in R, while a vector field
a(x, y, z) associates a vector with each point. In what follows, we will assume that
the variation in the scalar or vector field from point to point is both continuousand differentiable in R.
Simple examples of scalar fields include the pressure at each point in a fluid
and the electrostatic potential at each point in space in the presence of an electric
charge. Vector fields relating to the same physical systems are the velocity vector
in a fluid (giving the local speed and direction of the flow) and the electric field.
With the study of continuously varying scalar and vector fields there arises the
need to consider their derivatives and also the integration of field quantities alonglines, over surfaces and throughout volumes in the field. We defer the discussionof line, surface and volume integrals until the next chapter, and in the remainderof this chapter we concentrate on the definition of vector differential operatorsand their properties.
10.7 Vector operators
Certain differential operations may be performed on scalar and vector fields
and have wide-ranging applications in the physical sciences. The most important
operations are those of finding the gradient of a scalar field and the divergence
andcurlof a vector field. It is usual to define these operators from a strictly
353
VECTOR CALCULUS
mathematical point of view, as we do below. In the following chapter, however, we
will discuss their geometrical definitions, which rely on the concept of integratingvector quantities along lines and over surfaces.
Central to all these differential operations is the vector operator ∇,w h i c hi s
called del(or sometimes nabla) and in Cartesian coordinates is defined by
∇≡i∂
∂x+j∂
∂y+k∂
∂z. (10.25)
The form of this operator in non-Cartesian coordinate systems is discussed in
sections 10.9 and 10.10.
10.7.1 Gradient of a scalar field
Thegradient of a scalar field φ(x, y, z) is defined by
grad φ=∇φ=i∂φ
∂x+j∂φ
∂y+k∂φ
∂z. (10.26)
Clearly,∇φis a vector field whose x-,y-a n d z- components are the first partial
derivatives of φ(x, y, z) with respect to x,yandzrespectively. Also note that the
vector field ∇φshould not be confused with the vector operator φ∇, which has
components ( φ∂ / ∂ x ,φ∂ / ∂ y,φ∂ / ∂ z ).IFind the gradient of the scalar field φ=xy2z3.
From (10.26) the gradient of φis given by
∇φ=y2z3i+2xyz3j+3xy2z2k.
J
The gradient of a scalar field φhas some interesting geometrical properties.
Let us first consider the problem of calculating the rate of change of φin some
particular direction . For an infinitesimal vector displacement dr, forming its scalar
product with ∇φwe obtain
∇φ·dr=parenleftbigg
i∂φ
∂x+j∂φ
∂y+k∂φ
∂zparenrightbigg
·(idx+jdy+kdx),
=∂φ
∂xdx+∂φ
∂ydy+∂φ
∂zdz,
=dφ, (10.27)
which is the infinitesimal change in φin going from position rtor+dr.I n
particular, if rdepends on some parameter usuch that r(u) defines a space curve
354
10.7 VECTOR OPERATORS
φ=c o n s t a n t∇φ
a
PQ
dφ
dsin the direction aθ
Figure 10.5 Geometrical properties of ∇φ.PQgives the value of dφ/ds in
the direction a.
then the total derivative of φwith respect to ualong the curve is simply
dφ
du=∇φ·dr
du; (10.28)
in the particular case where the parameter uis the arc length salong the curve,
the total derivative of φwith respect to salong the curve is given by
dφ
ds=∇φ·ˆt, (10.29)
where ˆtis the unit tangent to the curve at the given point, as discussed in
section 10.3.
In general, the rate of change of φwith respect to the distance sin a particular
direction ais given by
dφ
ds=∇φ·ˆa (10.30)
and is called the directional derivative. Since ˆais a unit vector we have
dφ
ds=|∇φ|cosθ
where θis the angle between ˆaand∇φas shown in figure 10.5. Clearly ∇φlies
in the direction of the fastest increase in φ,a n d|∇φ|is the largest possible value
ofdφ/ds . Similarly, the largest rate of decrease of φisdφ/ds =−|∇φ|in the
direction of −∇φ.
355
VECTOR CALCULUSIFor the function φ=x2y+yzat the point (1,2,−1), find its rate of change with distance
in the direction a=i+2j+3k. At this same point, what is the greatest possible rate of
change with distance and in which direction does it occur?
The gradient of φis given by (10.26):
∇φ=2xyi+(x2+z)j+yk,
=4i+2kat the point (1 ,2,−1).
The unit vector in the direction of aisˆa=1√
14(i+2j+3k), so the rate of change of φ
with distance sin this direction is, using (10.30),
dφ
ds=∇φ·ˆa=1√
14(4 + 6) =10√
14.
From the above discussion, at the point (1 ,2,−1)dφ/ds will be greatest in the direction
of∇φ=4i+2kand has the value |∇φ|=√
20 in this direction.
J
We can extend the above analysis to find the rate of change of a vector
field (rather than a scalar field as above) in a particular direction. The scalardifferential operator ˆa·∇can be shown to give the rate of change with distance
in the direction ˆaof the quantity (vector or scalar) on which it acts. In Cartesian
coordinates it may be written as
ˆa·∇=a
x∂
∂x+ay∂
∂y+az∂
∂z. (10.31)
Thus we can write the infinitesimal change in an electric field in moving from r
tor+drgiven in (10.20) as dE=(dr·∇)E.
A second interesting geometrical property of ∇φmay be found by considering
the surface defined by φ(x, y, z)=c,w h e r e cis some constant. If ˆtis a unit
tangent to this surface at some point then clearly dφ/ds = 0 in this direction
and from (10.29) we have ∇φ·ˆt= 0. In other words, ∇φis a vector normal to
the surface φ(x, y, z)=cat every point , as shown in figure 10.5. If ˆnis a unit
normal to the surface in the direction of increasing φ(x, y, z), then the gradient is
sometimes written
∇φ≡∂φ
∂nˆn, (10.32)
where ∂φ/∂n≡| ∇φ|is the rate of change of φin the direction ˆnand is called
thenormal derivative .IFind expressions for the equations of the tangent plane and the line normal to the surface
φ(x, y, z)=cat the point Pwith coordinates x0,y0,z0. Use the results to find the equations
of the tangent plane and the line normal to the surface of the sphere φ=x2+y2+z2=a2
at the point (0,0,a).
A vector normal to the surface φ(x, y, z)=cat the point Pis simply∇φevaluated at that
point; we denote it by n0.I fr0is the position vector of the point Prelative to the origin,
356
10.7 VECTOR OPERATORS
xyz
ˆn0
(0,0,a)
O
a
φ=x2+y2+z2=a2z=a
Figure 10.6 The tangent plane and the normal to the surface of the sphere
φ=x2+y2+z2=a2at the point r0with coordinates (0 ,0,a).
andris the position vector of any point on the tangent plane, then the vector equation of
the tangent plane is, from (7.41),
(r−r0)·n0=0.
Similarly, if ris the position vector of any point on the straight line passing through P
(with position vector r0) in the direction of the normal n0then the vector equation of this
line is, from subsection 7.7.1,
(r−r0)×n0=0.
For the surface of the sphere φ=x2+y2+z2=a2,
∇φ=2xi+2yj+2zk
=2akat the point (0 ,0,a).
Therefore the equation of the tangent plane to the sphere at this point is
(r−r0)·2ak=0.
This gives 2 a(z−a)=0o r z=a, as expected. The equation of the line normal to the
sphere at the point (0 ,0,a)i s
(r−r0)×2ak=0,
which gives 2 ayi−2axj=0orx=y=0 ,i . e .t h e z-axis, as expected. The tangent plane
and normal to the surface of the sphere at this point are shown in figure 10.6.
J
Further properties of the gradient operation, which are analogous to those of
the ordinary derivative, are listed in subsection 10.8.1 and may be easily proved.
357
VECTOR CALCULUS
In addition to these, we note that the gradient operation also obeys the chain
rule as in ordinary differential calculus, i.e. if φandψare scalar fields in some
region Rthen
∇[φ(ψ)]=∂φ
∂ψ∇ψ.
10.7.2 Divergence of a vector field
Thedivergence of a vector field a(x, y, z) is defined by
diva=∇·a=∂ax
∂x+∂ay
∂y+∂az
∂z, (10.33)
where ax,ayandazare the x-,y-a n d z- components of a. Clearly,∇·ais a scalar
field. Any vector field afor which∇·a= 0 is said to be solenoidal .IFind the divergence of the vector field a=x2y2i+y2z2j+x2z2k.
From (10.33) the divergence of ais given by
∇·a=2xy2+2yz2+2x2z=2 (xy2+yz2+x2z).
J
We will discuss fully the geometric definition of divergence and its physical
meaning in the next chapter. For the moment, we merely note that the divergence
can be considered as a quantitative measure of how much a vector field diverges(spreads out) or converges at any given point. For example, if we consider thevector field v(x, y, z) describing the local velocity at any point in a fluid then ∇·v
is equal to the net rate of outflow of fluid per unit volume, evaluated at a point(by letting a small volume at that point tend to zero).
Now if some vector field ais itself derived from a scalar field via a=∇φthen
∇·ahas the form ∇·∇φo r ,a si ti su s u a l l yw r i t t e n , ∇
2φ,w h e r e∇2(del squared)
is the scalar differential operator
∇2≡∂2
∂x2+∂2
∂y2+∂2
∂z2. (10.34)
∇2φis called the Laplacian ofφand appears in several important partial differ-
ential equations of mathematical physics, discussed in chapters 18 and 19.IFind the Laplacian of the scalar field φ=xy2z3.
From (10.34) the Laplacian of φis given by
∇2φ=∂2φ
∂x2+∂2φ
∂y2+∂2φ
∂z2=2xz3+6xy2z.
J
358
10.7 VECTOR OPERATORS
10.7.3 Curl of a vector field
Thecurlof a vector field a(x, y, z) is defined by
curla=∇×a=parenleftbigg∂az
∂y−∂ay
∂zparenrightbigg
i+parenleftbigg∂ax
∂z−∂az
∂xparenrightbigg
j+parenleftbigg∂ay
∂x−∂ax
∂yparenrightbigg
k,
where ax,ayandazare the x-,y-a n d z- components of a. The RHS can be
written in a more memorable form as a determinant:
∇×a=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleij k
∂
∂x∂
∂y∂
∂z
axayazvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (10.35)
where it is understood that, on expanding the determinant, the partial derivatives
in the second row act on the components of ain the third row. Clearly, ∇×a
is itself a vector field. Any vector field afor which∇×a=0is said to be
irrotational .IFind the curl of the vector field a=x2y2z2i+y2z2j+x2z2k.
The curl of ais given by
∇φ=
//////
//
ij k
∂
∂x∂
∂y∂
∂z
x2y2z2y2z2x2z2
//////
//
=−2
/
y2zi+(xz2−x2y2z)j+x2yz2k
/
.
J
For a vector field v(x, y, z) describing the local velocity at any point in a fluid,
∇×vis a measure of the angular velocity of the fluid in the neighbourhood of
that point. If a small paddle wheel were placed at various points in the fluid thenit would tend to rotate in regions where ∇×v/negationslash=0, while it would not rotate in
regions where ∇×v=0.
Another insight into the physical interpretation of the curl operator is gained
by considering the vector field vdescribing the velocity at any point in a rigid
body rotating about some axis with angular velocity ω.I fris the position vector
of the point with respect to some origin on the axis of rotation then the velocityof the point is given by v=ω×r. Without any loss of generality, we may take
ωto lie along the z-axis of our coordinate system, so that ω=ωk.T h ev e l o c i t y
field is then v=−ωyi+ωxj. The curl of this vector field is easily found to be
∇×v=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleij k
∂
∂x∂
∂y∂
∂z
−ωy ωx 0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=2ωk=2ω. (10.36)
359
VECTOR CALCULUS
∇(φ+ψ)=∇φ+∇ψ
∇·(a+b)=∇·a+∇·b
∇×(a+b)=∇×a+∇×b
∇(φψ)=φ∇ψ+ψ∇φ
∇(a·b)=a×(∇×b)+b×(∇×a)+(a·∇)b+(b·∇)a
∇·(φa)=φ∇·a+a·∇φ
∇·(a×b)=b·(∇×a)−a·(∇×b)
∇×(φa)=∇φ×a+φ∇×a
∇×(a×b)=a(∇·b)−b(∇·a)+(b·∇)a−(a·∇)b
Table 10.1 Vector operators acting on sums and products. The operator ∇is
defined in (10.25); φandψare scalar fields, aandbare vector fields.
Therefore the curl of the velocity field is a vector equal to twice the angular
velocity vector of the rigid body about its axis of rotation. We give a fullgeometrical discussion of the curl of a vector in the next chapter.
10.8 Vector operator formulae
In the same way as for ordinary vectors (chapter 7), for vector operators certain
identities exist. In addition, we must consider various relations involving theaction of vector operators on sums and products of scalar and vector fields. Some
of these relations have been mentioned above, but we list all the most important
ones here for convenience. The validity of these relations may be easily verifiedby direct calculation (a quick method of deriving them using tensor notation isgiven in chapter 21).
Although some of the following vector relations are expressed in Cartesian
coordinates, it may be proved that they are all independent of the choice ofcoordinate system. This is to be expected since grad, div and curl all have cleargeometrical definitions, which are discussed more fully in the next chapter andwhich do not rely on any particular choice of coordinate system.
10.8.1 Vector operators acting on sums and products
Letφandψbe scalar fields and aandbbe vector fields. Assuming these fields
are differentiable, the action of grad, div and curl on various sums and productsof them is presented in table 10.1.
These relations can be proved by direct calculation.
360
10.8 VECTOR OPERATOR FORMULAEIShow that
∇×(φa)=∇φ×a+φ∇×a.
Thex-component of the LHS is
∂
∂y(φaz)−∂
∂z(φay)=φ∂az
∂y+∂φ
∂yaz−φ∂ay
∂z−∂φ
∂zay,
=φ
/∂az
∂y−∂ay
∂z
/
+
/∂φ
∂yaz−∂φ
∂zay
/
,
=φ(∇×a)x+(∇φ×a)x,
where, for example, ( ∇φ×a)xdenotes the x-component of the vector ∇φ×a. Incorporating
they-a n d z- components, which can be similarly found, we obtain the stated result.
J
Some useful special cases of the relations in table 10.1 are worth noting. If ris
the position vector relative to some origin and r=|r|,t h e n
∇φ(r)=dφ
drˆr,
∇·[φ(r)r]=3φ(r)+rdφ(r)
dr,
∇2φ(r)=d2φ(r)
dr2+2
rdφ(r)
dr,
∇×[φ(r)r]=0.
These results may be proved straightforwardly using Cartesian coordinates but
far more simply using spherical polar coordinates, which are discussed in subsec-tion 10.9.2. Particular cases of these results are
∇r=ˆr,∇·r=3,∇×r=0,
together with
∇parenleftbigg1
rparenrightbigg
=−ˆr
r2,
∇·parenleftbiggˆr
r2parenrightbigg
=−∇2parenleftbigg1
rparenrightbigg
=4πδ(r),
where δ(r) is the Dirac delta function, discussed in chapter 13. The last equation is
important in the solution of certain partial differential equations and is discussedfurther in chapter 18.
10.8.2 Combinations of grad, div and curl
We now consider the action of two vector operators in succession on a scalar or
vector field. We can immediately discard four of the nine obvious combinations ofgrad, div and curl, since they clearly do not make sense. If φis a scalar field and
361
VECTOR CALCULUS
ais a vector field, these four combinations are grad(grad φ), div(div a), curl(div a)
and grad(curl a). In each case the second (outer) vector operator is acting on the
wrong type of field, i.e. scalar instead of vector or vice versa. In grad(grad φ),
for example, grad acts on grad φ, which is a vector field, but we know that grad
only acts on scalar fields (although in fact we will see in chapter 21 that we can
form the outer product of the del operator with a vector to form a tensor, but
that need not concern us here).
Of the five valid combinations of grad, div and curl, two are identically zero,
namely
curl grad φ=∇×∇ φ=0, (10.37)
div curl a=∇·(∇×a)=0 . (10.38)
From (10.37), we see that if ais derived from the gradient of some scalar function
such that a=∇φthen it is necessarily irrotational ( ∇×a= 0). We also note
that if ais an irrotational vector field then another irrotational vector field is
a+∇φ+c,w h e r e φis any scalar field and cis a constant vector. This follows
since
∇×(a+∇φ+c)=∇×a+∇×∇ φ=0.
Similarly, from (10.38) we may infer that if bis the curl of some vector field a
such that b=∇×athenbis solenoidal ( ∇·b= 0). Obviously, if bis solenoidal
andcis any constant vector then b+cis also solenoidal.
The three remaining combinations of grad, div and curl are
div grad φ=∇·∇φ=∇2φ=∂2φ
∂x2+∂2φ
∂y2+∂2φ
∂z2, (10.39)
grad div a=∇(∇·a),
=parenleftbigg∂2ax
∂x2+∂2ay
∂x∂y+∂2az
∂x∂zparenrightbigg
i+parenleftbigg∂2ax
∂y∂x+∂2ay
∂y2+∂2az
∂y∂zparenrightbigg
j
+parenleftbigg∂2ax
∂z∂x+∂2ay
∂z∂y+∂2az
∂z2parenrightbigg
k, (10.40)
curl curl a=∇×(∇×a)=∇(∇·a)−∇2a, (10.41)
where (10.39) and (10.40) are expressed in Cartesian coordinates. In (10.41), the
term∇2ahas the linear differential operator ∇2acting on a vector (as opposed to
a scalar as in (10.39)), which of course consists of a sum of unit vectors multipliedby components. Two cases arise.
(i) If the unit vectors are constants (i.e. they are independent of the values of
the coordinates) then the differential operator gives a non-zero contribution
only when acting upon the coordinates, the unit vectors being merelymultipliers.
362
10.9 CYLINDRICAL AND SPHERICAL POLAR COORDINATES
(ii) If the unit vectors vary as the values of the coordinates change (i.e. are
not constant in direction throughout the whole space) then the derivativesof these vectors appear as contributions to ∇
2a.
Cartesian coordinates are an example of the first case in which each component
satisfies (∇2a)i=∇2ai. In this case (10.41) can be applied to each component
separately:
[∇×(∇×a)]i=[∇(∇·a)]i−∇2ai. (10.42)
However, cylindrical and spherical polar coordinates come in the second class.
For them (10.41) is still true, but the further step to (10.42) cannot be made.
More complicated vector operator relations may be proved using the relations
given above.IShow that
∇·(∇φ×∇ψ)=0 ,
where φandψare scalar fields.
From the previous section we have
∇·(a×b)=b·(∇×a)−a·(∇×b).
If we let a=∇φandb=∇ψthen we obtain
∇·(∇φ×∇ψ)=∇ψ·(∇×∇ φ)−∇φ·(∇×∇ ψ)=0 , (10.43)
since∇×∇ φ=0=∇×∇ ψ, from (10.37).
J
10.9 Cylindrical and spherical polar coordinates
The operators we have discussed in this chapter, i.e. grad, div, curl and ∇2,
have all been defined in terms of Cartesian coordinates, but for many physicalsituations other coordinate systems are more natural. For example, many systems,such as an isolated charge in space, have spherical symmetry and spherical polarcoordinates would be the obvious choice. For axisymmetric systems, such as fluidflow in a pipe, cylindrical polar coordinates are the natural choice. The physical
laws governing the behaviour of the systems are often expressed in terms of
the vector operators we have been discussing, and so it is necessary to be ableto express these operators in these other, non-Cartesian, coordinates. We firstconsider the two most common non-Cartesian coordinate systems, i.e. cylindricaland spherical polars, and go on to discuss general curvilinear coordinates in thenext section.
10.9.1 Cylindrical polar coordinates
As shown in figure 10.7, the position of a point in space Phaving Cartesian
coordinates x, y, z may be expressed in terms of cylindrical polar coordinates
363
VECTOR CALCULUS
ρ, φ, z,w h e r e
x=ρcosφ, y =ρsinφ, z =z, (10.44)
andρ≥0, 0≤φ<2πand−∞<z<∞. The position vector of Pmay therefore
be written
r=ρcosφi+ρsinφj+zk. (10.45)
If we take the partial derivatives of rwith respect to ρ,φandzrespectively then
we obtain the three vectors
eρ=∂r
∂ρ=c o s φi+s i n φj, (10.46)
eφ=∂r
∂φ=−ρsinφi+ρcosφj, (10.47)
ez=∂r
∂z=k. (10.48)
These vectors lie in the directions of increasing ρ,φandzrespectively but are
not all of unit length. Although eρ,eφandezform a useful set of basis vectors
in their own right (we will see in section 10.10 that such a basis is sometimes themostuseful), it is usual to work with the corresponding unitvectors, which are
obtained by dividing each vector by its modulus to give
ˆe
ρ=eρ=c o s φi+s i n φj, (10.49)
ˆeφ=1
ρeφ=−sinφi+c o s φj, (10.50)
ˆez=ez=k. (10.51)
These three unit vectors, like the Cartesian unit vectors i,jandk,f o r ma n
orthonormal triad at each point in space, i.e. the basis vectors are mutuallyorthogonal and of unit length (see figure 10.7). Unlike the fixed vectors i,jandk,
however, ˆe
ρandˆeφchange direction as Pmoves.
The expression for a general infinitesimal vector displacement drin the position
ofPis given, from (10.19), by
dr=∂r
∂ρdρ+∂r
∂φdφ+∂r
∂zdz
=dρeρ+dφeφ+dzez
=dρˆeρ+ρd φˆeφ+dzˆez. (10.52)
This expression illustrates an important difference between Cartesian and cylin-
drical polar coordinates (or non-Cartesian coordinates in general). In Cartesian
coordinates, the distance moved in going from xtox+dx, with yandzheld
constant, is simply ds=dx. However, in cylindrical polars, if φchanges by dφ,
with ρandzheld constant, then the distance moved is notdφ, but ds=ρd φ.
364
10.9 CYLINDRICAL AND SPHERICAL POLAR COORDINATES
xyz
z
ρr
ijk
OPˆez
ˆeφ
ˆeρ
φ
Figure 10.7 Cylindrical polar coordinates ρ, φ, z.
xyz
ρdzρd φ
ρd φφ dφdρ
Figure 10.8 The element of volume in cylindrical polar coordinates is given
byρd ρd φd z .
Factors, such as the ρinρd φ, that multiply the coordinate differentials (in the
orthonormal basis) to get distances are known as scale factors . From (10.52), the
scale factors for the ρ-,φ-a n d z- coordinates are therefore 1, ρand 1 respectively.
The magnitude dsof the displacement dris given in cylindrical polar coordinates
by
(ds)2=dr·dr=(dρ)2+ρ2(dφ)2+(dz)2,
where in the second equality we have used the fact that the basis vectors are
orthonormal. We can also find the volume element in a cylindrical polar system(see figure 10.8) by calculating the volume of the infinitesimal parallelepiped
365
VECTOR CALCULUS
∇Φ=∂Φ
∂ρˆeρ+1
ρ∂Φ
∂φˆeφ+∂Φ
∂zˆez
∇·a=1
ρ∂
∂ρ(ρaρ)+1
ρ∂aφ
∂φ+∂az
∂z
∇×a=1
ρ
////////ˆeρρˆeφˆez
∂
∂ρ∂
∂φ∂
∂z
aρρaφaz
////////
∇2Φ=1
ρ∂
∂ρ
/
ρ∂Φ
∂ρ
/
+1
ρ2∂2Φ
∂φ2+∂2Φ
∂z2
Table 10.2 Vector operators in cylindrical polar coordinates; Φ is a scalar
field and ais a vector field.
defined by the vectors dρˆeρ,ρd φˆeφanddzˆez;t h i si sg i v e nb y
dV=|dρˆeρ·(ρd φˆeφ×dzˆez)|=ρd ρd φd z,
which again uses the fact that the basis vectors are orthonormal. For a simple
coordinate system such as cylindrical polars the expressions for ( ds)2anddVare
obvious from the geometry.
We will now express the vector operators discussed in this chapter in terms of
cylindrical polar coordinates. Let us consider a scalar field Φ( ρ, φ, z), where we
use Φ for the scalar field to avoid confusion with the azimuthal angle φ,a n da
vector field a(ρ, φ, z). We must first write the vector field in terms of the basis
vectors of the cylindrical polar coordinate system, i.e.
a=aρˆeρ+aφˆeφ+azˆez,
where aρ,aφandazare the components of ain the ρ-,φ-a n d z- directions
respectively. The expressions for grad, div, curl and ∇2c a nt h e nb ec a l c u l a t e d
and are given in table 10.2. Since the derivations of these expressions are rathercomplicated we leave them until our discussion of general curvilinear coordinatesin the next section; the reader could well postpone examination of these formalproofs until some experience of using the expressions has been gained.IExpress the vector field a=yzi−yj+xz2kin cylindrical polar coordinates, and hence
calculate its divergence. Show that the same result is obtained by evaluating the divergencein Cartesian coordinates.
The basis vectors of the cylindrical polar coordinate system are given in (10.49)–(10.51).Solving these equations simultaneously for i,jandkwe obtain
i=c o s φˆe
ρ−sinφˆeφ
j=s i n φˆeρ+c o s φˆeφ
k=ˆez.
366
10.9 CYLINDRICAL AND SPHERICAL POLAR COORDINATES
xyz
r
ijk
OθPˆer
ˆeφ
ˆeθ
φ
Figure 10.9 Spherical polar coordinates r,θ,φ.
Substituting these relations and (10.44) into the expression for awe find
a=zρsinφ(cosφˆeρ−sinφˆeφ)−ρsinφ(sinφˆeρ+c o s φˆeφ)+z2ρcosφˆez
=(zρsinφcosφ−ρsin2φ)ˆeρ−(zρsin2φ+ρsinφcosφ)ˆeφ+z2ρcosφˆez.
Substituting into the expression for ∇·agiven in table 10.2,
∇·a=2zsinφcosφ−2si n2φ−2zsinφcosφ−cos2φ+s i n2φ+2zρcosφ
=2zρcosφ−1.
Alternatively, and much more quickly in this case, we can calculate the divergence
directly in Cartesian coordinates. We obtain
∇·a=∂ax
∂x+∂ay
∂y+∂az
∂z=2zx−1,
which on substituting x=ρcosφyields the same result as the calculation in cylindrical
polars.
J
Finally, we note that similar results can be obtained for (two-dimensional)
polar coordinates in a plane by omitting the z-dependence. For example, ( ds)2=
(dρ)2+ρ2(dφ)2, while the element of volume is replaced by the element of area
dA=ρd ρd φ .
10.9.2 Spherical polar coordinates
As shown in figure 10.9, the position of a point in space P, with Cartesian
coordinates x, y, z, may be expressed in terms of spherical polar coordinates
r,θ,φ,w h e r e
x=rsinθcosφ, y =rsinθsinφ, z =rcosθ, (10.53)
367
VECTOR CALCULUS
andr≥0, 0≤θ≤πand 0≤φ<2π. The position vector of Pmay therefore be
written as
r=rsinθcosφi+rsinθsinφj+rcosθk.
If, in a similar manner to that used in the previous section for cylindrical polars,
we find the partial derivatives of rwith respect to r,θandφrespectively and
divide each of the resulting vectors by its modulus then we obtain the unit basis
vectors
ˆer=s i n θcosφi+s i n θsinφj+c o s θk,
ˆeθ=c o s θcosφi+c o s θsinφj−sinθk,
ˆeφ=−sinφi+c o s φj.
These unit vectors are in the directions of increasing r,θandφrespectively
and are the orthonormal basis set for spherical polar coordinates, as shown infigure 10.9.
A general infinitesimal vector displacement in spherical polars is, from (10.19),
dr=drˆe
r+rd θˆeθ+rsinθd φˆeφ; (10.54)
thus the scale factors for the r-,θ-a n d φ- coordinates are 1, rand rsinθ
respectively. The magnitude dsof the displacement dris therefore given by
(ds)2=dr·dr=(dr)2+r2(dθ)2+r2sin2θ(dφ)2,
since the basis vectors form an orthonormal set. The element of volume in
spherical polar coordinates (see figure 10.10) is the volume of the infinitesimalparallelepiped defined by the vectors drˆe
r,rd θˆeθandrsinθd φˆeφand is given
by
dV=|drˆer·(rd θˆeθ×rsinθd φˆeφ)|=r2sinθd rd θd φ ,
where again we use the fact that the basis vectors are orthonormal. The expres-
sions for ( ds)2anddVin spherical polars can be obtained from the geometry of
this coordinate system.
We will now express the standard vector operators in spherical polar coordi-
nates, using the same techniques as for cylindrical polar coordinates. We considera scalar field Φ( r,θ,φ) and a vector field a(r,θ,φ). The latter may be written in
terms of the basis vectors of the spherical polar coordinate system as
a=a
rˆer+aθˆeθ+aφˆeφ,
where ar,aθandaφare the components of ain the r-,θ-a n d φ- directions
respectively. The expressions for grad, div, curl and ∇2are given in table 10.3.
The derivations of these results are given in the next section.
368
10.9 CYLINDRICAL AND SPHERICAL POLAR COORDINATES
∇Φ=∂Φ
∂rˆer+1
r∂Φ
∂θˆeθ+1
rsinθ∂Φ
∂φˆeφ
∇·a=1
r2∂
∂r(r2ar)+1
rsinθ∂
∂θ(sinθaθ)+1
rsinθ∂aφ
∂φ
∇×a=1
r2sinθ
////////ˆerrˆeθrsinθˆeφ
∂
∂r∂
∂θ∂
∂φ
arraθrsinθaφ
////////
∇2Φ=1
r2∂
∂r
/
r2∂Φ
∂r
/
+1
r2sinθ∂
∂θ
/
sinθ∂Φ
∂θ
/
+1
r2sin2θ∂2Φ
∂φ2
Table 10.3 Vector operators in spherical polar coordinates. Φ is a scalar field
andais a vector field.
xyz
rrd θ
φdφdφ
dr
rsinθ
rsinθd φrsinθd φ
θ
dθ
Figure 10.10 The element of volume in spherical polar coordinates is given
byr2sinθd rd θd φ .
As a final note we mention that in the expression for ∇2Φ given in table 10.3
we can rewrite the first term on the RHS as follows:
1
r2∂
∂rparenleftbigg
r2∂Φ
∂rparenrightbigg
=1
r∂2
∂r2(rΦ),
w h i c hc a no f t e nb eu s e f u li ns h o r t e n i n gc a l c u l a t i o n s .
369
VECTOR CALCULUS
10.10 General curvilinear coordinates
As indicated earlier, the contents of this section are more formal and technically
complicated than hitherto. The section could be omitted until the reader has had
some experience of using its results.
Cylindrical and spherical polars are just two examples of what are called
general curvilinear coordinates . In the general case, the position of a point P
having Cartesian coordinates x, y, z may be expressed in terms of the three
curvilinear coordinates u1,u2,u3,w h e r e
x=x(u1,u2,u3),y =y(u1,u2,u3),z =z(u1,u2,u3),
and similarly
u1=u1(x, y, z),u 2=u2(x, y, z),u 3=u3(x, y, z).
We assume that all these functions are continuous, differentiable and have a
single-valued inverse, except perhaps at or on certain isolated points or lines,
so that there is a one-to-one correspondence between the x, y, z andu1,u2,u3
systems. The u1-,u2-a n d u3- coordinate curves of a general curvilinear system
are analogous to the x-,y-a n d z- axes of Cartesian coordinates. The surfaces
u1=c1,u2=c2and u3=c3,w h e r e c1,c2,c3are constants, are called the
coordinate surfaces and each pair of these surfaces has its intersection in a curve
called a coordinate curve orline(see figure 10.11). If at each point in space the
three coordinate surfaces passing through the point meet at right angles then
the curvilinear coordinate system is called orthogonal . For example, in spherical
polars u1=r,u2=θ,u3=φand the three coordinate surfaces passing through
the point ( R,Θ,Φ) are the sphere r=R, the circular cone θ= Θ and the plane
φ= Φ, which intersect at right angles at that point. Therefore spherical polars
(and cylindrical polars) form an orthogonal coordinate system.
Ifr(u1,u2,u3) is the position vector of the point Pthene1=∂r/∂u1is a vector
tangent to the u1-curve at P(for which u2andu3are constants) in the direction
of increasing u1. Similarly, e2=∂r/∂u2ande3=∂r/∂u3are vectors tangent to
theu2-a n d u3-c u r v e sa t Pin the direction of increasing u2andu3respectively.
Denoting the lengths of these vectors by h1,h2andh3,t h eunitvectors in each of
these directions are given by
ˆe1=1
h1∂r
∂u1,
ˆe2=1
h2∂r
∂u2,
ˆe3=1
h3∂r
∂u3,
where h1=|∂r/∂u1|,h2=|∂r/∂u2|andh3=|∂r/∂u3|.
The quantities h1,h2,h3are called the scale factors of the curvilinear coordinate
370
10.10 GENERAL CURVILINEAR COORDINATES
z
xy ijk
OPu2=c2u1=c1
u3=c3u1u2u3
ˆ/epsilon11ˆ/epsilon12ˆ/epsilon13
ˆe1
ˆe2ˆe3
Figure 10.11 General curvilinear coordinates.
system. The element of distance associated with an infinitesimal change duiin
one of the coordinates is hidui. In the previous section we found that the scale
factors for cylindrical and spherical polar coordinates were
for cylindrical polars hρ=1 , hφ=ρ,hz=1 ,
for spherical polars hr=1 , hθ=r,hφ=rsinθ.
Although the vectors e1,e2,e3form a perfectly good basis for the curvilinear
coordinate system, it is usual to work with the corresponding unit vectors ˆe1,ˆe2,
ˆe3. For an orthogonal curvilinear coordinate system these unit vectors form an
orthonormal basis.
An infinitesimal vector displacement in general curvilinear coordinates is given
by, from (10.19),
dr=∂r
∂u1du1+∂r
∂u2du2+∂r
∂u3du3 (10.55)
=du1e1+du2e2+du3e3 (10.56)
=h1du1ˆe1+h2du2ˆe2+h3du3ˆe3. (10.57)
I nt h ec a s eo f orthogonal curvilinear coordinates, where the ˆeiare mutually
perpendicular, the element of arc length is given by
(ds)2=dr·dr=h2
1(du1)2+h2
2(du2)2+h2
3(du3)2. (10.58)
The volume element for the coordinate system is the volume of the infinitesimal
parallelepiped defined by the vectors ( ∂r/∂u i)dui=duiei=hiduiˆei,f o r i=1,2,3.
371
VECTOR CALCULUS
For orthogonal coordinates this is given by
dV=|du1e1·(du2e2×du3e3)|
=|h1ˆe1·(h2ˆe2×h3ˆe3)|du1du2du3
=h1h2h3du1du2du3.
Now, in addition to the set {ˆei},i=1,2,3, there exists another useful set of
three unit basis vectors at P.S i n c e∇u1is a vector normal to the surface u1=c1,
a unit vector in this direction is ˆ/epsilon11=∇u1/|∇u1|. Similarly, ˆ/epsilon12=∇u2/|∇u2|and
ˆ/epsilon13=∇u3/|∇u3|are unit vectors normal to the surfaces u2=c2andu3=c3
respectively.
Therefore at each point Pin a curvilinear coordinate system, there exist, in
general, two sets of unit vectors: {ˆei}, tangent to the coordinate curves, and {ˆ/epsilon1i},
normal to the coordinate surfaces. A vector acan be written in terms of either
set of unit vectors:
a=a1ˆe1+a2ˆe2+a3ˆe3=A1ˆ/epsilon11+A2ˆ/epsilon12+A3ˆ/epsilon13,
where a1,a2,a3andA1,A2,A3are the components of ain the two systems. It
may be shown that the two bases become identical if the coordinate system isorthogonal.
Instead of the unitvectors discussed above, we could instead work directly with
the two sets of vectors {e
i=∂r/∂u i}and{/epsilon1i=∇ui},w h i c ha r en o t ,i ng e n e r a l ,o f
unit length. We can then write a vector aas
a=α1e1+α2e2+α3e3=β1/epsilon11+β2/epsilon12+β3/epsilon13,
or more explicitly as
a=α1∂r
∂u1+α2∂r
∂u2+α3∂r
∂u3=β1∇u1+β2∇u2+β3∇u3,
where α1,α2,α3andβ1,β2,β3are called the contravariant andcovariant com-
ponents of arespectively. A more detailed discussion of these components, in
the context of tensor analysis, is given in chapter 21. The (in general) non-unit
bases{ei}and{/epsilon1i}are often the most natural bases in which to express vector
quantities.IShow that{ei}and{/epsilon1i}are reciprocal systems of vectors.
Let us consider the scalar product ei·/epsilon1j; using the Cartesian expressions for rand∇,w e
obtain
ei·/epsilon1j=∂r
∂ui·∇uj
=
/∂x
∂uii+∂y
∂uij+∂z
∂uik
/
·
/∂uj
∂xi+∂uj
∂yj+∂uj
∂zk
/
=∂x
∂ui∂uj
∂x+∂y
∂ui∂uj
∂y+∂z
∂ui∂uj
∂z=∂uj
∂ui,
372
10.10 GENERAL CURVILINEAR COORDINATES
in the last step we have used the chain rule for partial differentiation. Therefore ei·/epsilon1j=1
ifi=j,a n dei·/epsilon1j= 0 otherwise. Hence {ei}and{/epsilon1j}are reciprocal systems of vectors.
J
We now derive expressions for the standard vector operators in orthogonal
curvilinear coordinates. Despite the useful properties of the non-unit bases dis-cussed above, the remainder of our discussion in this section will be in terms ofthe unit basis vectors {ˆe
i}. The expressions for the vector operators in cylindrical
and spherical polar coordinates given in tables 10.2 and 10.3 respectively can be
found from those derived below by inserting the appropriate scale factors.
Gradient
The change dΦ in a scalar field Φ resulting from changes du1,d u2,d u3in the
coordinates u1,u2,u3is given by, from (5.5),
dΦ=∂Φ
∂u1du1+∂Φ
∂u2du2+∂Φ
∂u3du3.
For orthogonal curvilinear coordinates u1,u2,u3we find from (10.57), and com-
parison with (10.27), that we can write this as
dΦ=∇Φ·dr, (10.59)
where∇Φi sg i v e nb y
∇Φ=1
h1∂Φ
∂u1ˆe1+1
h2∂Φ
∂u2ˆe2+1
h3∂Φ
∂u3ˆe3. (10.60)
This implies that the del operator can be written
∇=ˆe1
h1∂
∂u1+ˆe2
h2∂
∂u2+ˆe3
h3∂
∂u3.IShow that for orthogonal curvilinear coordinates ∇ui=ˆei/hi. Hence show that the two
sets of vectors {ˆei}and{ˆ/epsilon1i}are identical in this case.
Letting Φ = uiin (10.60) we find immediately that ∇ui=ˆei/hi. Therefore |∇ui|=1/hi,a n d
soˆ/epsilon1i=∇ui/|∇ui|=hi∇ui=ˆei.
J
Divergence
In order to derive the expression for the divergence of a vector field in orthogonal
curvilinear coordinates, we must first write the vector field in terms of the basisvectors of the coordinate system:
a=a
1ˆe1+a2ˆe2+a3ˆe3.
The divergence is then given by
∇·a=1
h1h2h3bracketleftbigg∂
∂u1(h2h3a1)+∂
∂u2(h3h1a2)+∂
∂u3(h1h2a3)bracketrightbigg
.
(10.61)
373
VECTOR CALCULUSIProve the expression for ∇·ain orthogonal curvilinear coordinates.
Let us consider the sub-expression ∇·(a1ˆe1). Now ˆe1=ˆe2׈e3=h2∇u2×h3∇u3. Therefore
∇·(a1ˆe1)=∇·(a1h2h3∇u2×∇u3),
=∇(a1h2h3)·(∇u2×∇u3)+a1h2h3∇·(∇u2×∇u3).
However,∇·(∇u2×∇u3) = 0, from (10.43), so we obtain
∇·(a1ˆe1)=∇(a1h2h3)·
/ˆe2
h2׈e3
h3
/
=∇(a1h2h3)·ˆe1
h2h3;
letting Φ = a1h2h3in (10.60) and substituting into the above equation, we find
∇·(a1ˆe1)=1
h1h2h3∂
∂u1(a1h2h3).
Repeating the analysis for ∇·(a2ˆe2)a n d∇·(a3ˆe3), and adding the results we obtain (10.61),
as required.
J
Laplacian
In the expression for the divergence (10.61), let
a=∇Φ=1
h1∂Φ
∂u1ˆe1+1
h2∂Φ
∂u2ˆe2+1
h3∂Φ
∂u3ˆe3,
where we have used (10.60). We then obtain
∇2Φ=1
h1h2h3bracketleftbigg∂
∂u1parenleftbiggh2h3
h1∂Φ
∂u1parenrightbigg
+∂
∂u2parenleftbiggh3h1
h2∂Φ
∂u2parenrightbigg
+∂
∂u3parenleftbiggh1h2
h3∂Φ
∂u3parenrightbiggbracketrightbigg
,
which is the expression for the Laplacian in orthogonal curvilinear coordinates.
Curl
The curl of a vector field a=a1ˆe1+a2ˆe2+a3ˆe3in orthogonal curvilinear
coordinates is given by
∇×a=1
h1h2h3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleh
1ˆe1h2ˆe2h3ˆe3
∂
∂u1∂
∂u2∂
∂u3
h1a1h2a2h3a3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (10.62)IProve the expression for ∇×ain orthogonal curvilinear coordinates.
Let us consider the sub-expression ∇×(a1ˆe1). Since ˆe1=h1∇u1we have
∇×(a1ˆe1)=∇×(a1h1∇u1),
=∇(a1h1)×∇u1+a1h1∇×∇ u1.
But∇×∇ u1=0 ,s ow eo b t a i n
∇×(a1ˆe1)=∇(a1h1)׈e1
h1.
374
10.11 EXERCISES
∇Φ=1
h1∂Φ
∂u1ˆe1+1
h2∂Φ
∂u2ˆe2+1
h3∂Φ
∂u3ˆe3
∇·a=1
h1h2h3
/∂
∂u1(h2h3a1)+∂
∂u2(h3h1a2)+∂
∂u3(h1h2a3)
/
∇×a=1
h1h2h3
///////h1ˆe1h2ˆe2h3ˆe3
∂
∂u1∂
∂u2∂
∂u3
h1a1h2a2h3a3
///////
∇2Φ=1
h1h2h3
/∂
∂u1
/h2h3
h1∂Φ
∂u1
/
+∂
∂u2
/h3h1
h2∂Φ
∂u2
/
+∂
∂u3
/h1h2
h3∂Φ
∂u3
//
Table 10.4 Vector operators in orthogonal curvilinear coordinates u1,u2,u3.
Φ is a scalar field and ais a vector field.
Letting Φ = a1h1in (10.60) and substituting into the above equation, we find
∇×(a1ˆe1)=ˆe2
h3h1∂
∂u3(a1h1)−ˆe3
h1h2∂
∂u2(a1h1).
The corresponding analysis of ∇×(a2ˆe2) produces terms in ˆe3andˆe1, whilst that of
∇×(a3ˆe3) produces terms in ˆe1andˆe2. When the three results are added together, the
coefficients multiplying ˆe1,ˆe2andˆe3are the same as those obtained by writing out (10.62)
explicitly, thus proving the stated result.
J
The general expressions for the vector operators in orthogonal curvilinear
coordinates are shown for reference in table 10.4. The explicit results for cylindricaland spherical polar coordinates, given in tables 10.2 and 10.3 respectively, areobtained by substituting the appropriate set of scale factors in each case.
A discussion of the expressions for vector operators in tensor form, which
are valid even for non-orthogonal curvilinear coordinate systems, is given inchapter 21.
10.11 Exercises
10.1 Evaluate the integralZ/
a(˙b·a+b·˙a)+˙a(b·a)−2(˙a·a)b−˙b|a|2
/
dt
in which ˙a,˙bare the derivatives of a,bwith respect to t.
10.2 At time t= 0, the vectors EandBare given by E=E0andB=B0,w h e r et h e
fixed unit vectors E0andB0are orthogonal. The equations of motion are
dE
dt=E0+B×E0,
dB
dt=B0+E×B0.
FindEandBat a general time t, showing that after a long time the directions
ofEandBhave almost interchanged.
375
VECTOR CALCULUS
10.3 The general equation of motion of a (non-relativistic) particle of mass mand
charge qwhen it is placed in a region where there is a magnetic field Band an
electric field Eis
m¨r=q(E+˙r×B);
hereris the position of the particle at time tand˙r=dr/dtetc. Write this as
three separate equations in terms of the Cartesian components of the vectors
involved.
For the simple case of crossed uniform fields E=Ei,B=Bjin which the
particle starts from the origin at t=0w i t h ˙r=v0k, find the equations of motion
and show the following:
(a) if v0=E/Bthen the particle continues its initial motion;
(b) if v0= 0 then the particle follows the space curve given in terms of the
parameter ξby
x=mE
B2q(1−cosξ),y =0,z =mE
B2q(ξ−sinξ).
Interpret this curve geometrically and relate ξtot. Show that the total
distance travelled by the particle after time tis
2E
B
Zt
0
///
/
sinBqt/prime
2m
///
/
dt/prime.
10.4 Use vector methods to find the maximum angle to the horizontal at which a stone
may be thrown so as to ensure that it is always moving away from the thrower.
10.5 If two systems of coordinates with a common origin Oare rotating with respect
to each other, the measured accelerations differ in the two systems. Denotingbyrandr
/primeposition vectors in frames OXY Z andOX/primeY/primeZ/primerespectively, the
connection between the two is
¨r/prime=¨r+˙ω×r+2ω×˙r+ω×(ω×r),
where ωis the angular velocity vector of the rotation of OXY Z with respect to
OX/primeY/primeZ/prime(taken as fixed). The third term on the RHS is known as the Coriolis
acceleration, whilst the final term gives rise to a centrifugal force.
Consider the application of this result to the firing of a shell of mass mfrom
a stationary ship on the steadily rotating earth, working to the first order inω(= 7.3×10
−5rad s−1). If the shell is fired with velocity vat time t=0a n do n l y
reaches a height that is small compared to the radius of the earth, show that itsacceleration, as recorded on the ship, is given approximately by
¨r=g−2ω×(v+gt),
where mgis the weight of the shell measured on the ship’s deck.
The shell is fired at another stationary ship (a distance saway) and vis such
that the shell would have hit its target had there been no Coriolis effect.
(a) Show that without the Coriolis effect the time of flight of the shell would
have been τ=−2g·v/g
2.
(b) Show further that when the shell ac tually hits the sea it is off target by
approximately
2τ
g2[(g×ω)·v](gτ+v)−(ω×v)τ2−1
3(ω×g)τ3.
(c) Estimate the order of magnitude ∆ of this miss for a shell for which v= 300
ms−1, firing close to its maximum range ( vmakes an angle of π/4w i t ht h e
vertical) in a northerly direction, whilst the ship is stationed at latitude 45◦
North.
376
10.11 EXERCISES
10.6 Prove that for a space curve r=r(s), where sis the arc length measured along
the curve from a fixed point, the triple scalar product/dr
ds×d2r
ds2
/
·d3r
ds3
at any point on the curve has the value κ2τ,w h e r e κis the curvature and τthe
torsion at that point.
10.7 For the twisted space curve y3+2 7axz−81a2y= 0, given parametrically by
x=au(3−u2),y =3au2,z =au(3 +u2),
show the following:
(a) that ds/du =3√2a(1 +u2), where sis the distance along the curve measured
from the origin;
(b) that the length of the curve from the origin to the Cartesian point (2 a,3a,4a)
is 4√2a;
(c) that the radius of curvature at the point with parameter uis 3a(1 +u2)2;
(d) that the torsion τand curvature κat a general point are equal;
(e) that any of the Frenet–Serret formulae that you have not already used
directly are satisfied.
10.8 The shape of the curving slip road joining two motorways that cross at right
angles and are at vertical heights z=0a n d z=hcan be approximated by the
space curve
r=√
2h
πln cos
/zπ
2h
/
i+√
2h
πln sin
/zπ
2h
/
j+zk.
Show that the radius of curvature ρof the curve is (2 h/π)cosec ( zπ/h)a th e i g h t
zand that the torsion τ=−1/ρ. (To shorten the algebra, set z=2hθ/πand use
θas the parameter.)
10.9 In a magnetic field, field lines are curves to which the magnetic induction Bis
everywhere tangential. By evaluating dB/ds,w h e r e sis the distance measured
along a field line, prove that the radius of curvature at any point on a line isgiven by
ρ=B
3
|B×(B·∇)B|.
10.10 (a) Using the parameterization x=ucosφ,y=usinφ,z=ucotΩ, find the
sloping surface area of a right circular cone of semi-angle Ω whose base hasradius a.V e r i f yt h a ti ti se q u a lt o
1
2×perimeter of the base ×slope height.
(b) Using the same parameterization as in (a) for xandy, and an appropriate
choice for z, find the surface area between the planes z=0a n d z=Zof
the paraboloid of revolution z=α(x2+y2).
10.11 (a) Parameterising the hyperboloid
x2
a2+y2
b2−z2
c2=1
byx=acosθsecφ,y=bsinθsecφ,z=ctanφ, show that an area element
on its surface is
dS=s e c2φ
/
c2sec2φ
/;
b2cos2θ+a2sin2θ
/
+a2b2tan2φ
/1/2dθ dφ.
(b) Use this formula to show that the area of the curved surface x2+y2−z2=a2
between the planes z=0a n d z=2ais
πa2
/
6+1√
2sinh−12√
2
/
.
377
VECTOR CALCULUS
10.12 For the function
z(x, y)=(x2−y2)e−x2−y2,
find the location(s) at which the steepest gradient occurs. What are the magnitude
and direction of that gradient? (The algebra involved is easier if plane polarcoordinates are used.)
10.13 Verify by direct calculation that
∇·(a×b)=b·(∇×a)−a·(∇×b).
10.14 (a) Simplify
∇×a(∇·a)+ a×[∇×(∇×a)]+a×∇
2a.
(b) By explicitly writing out the terms in Cartesian coordinates prove that
[c·(b·∇)−b·(c·∇)]a=(∇×a)·(b×c).
(c) Prove that a×(∇×a)=∇(1
2a2)−(a·∇)a.
10.15 Evaluate the Laplacian of the function
ψ(x, y, z)=zx2
x2+y2+z2
(a) directly in Cartesian coordinates, and (b) after changing to a spherical polar
coordinate system. Verify that, as they must, the two methods give the sameresult.
10.16 Verify that (10.42) is valid for each component separately when ais the Cartesian
vector x
2yi+xyzj+z2yk, by showing that each side of the equation is equal to
zi+( 2x+2z)j+xk.
10.17 The (Maxwell) relationship between a time-independent magnetic field Band the
current density J(measured in S.I. units in A m−2) producing it,
∇×B=µ0J,
can be applied to a long cylinder of conducting ionised gas which, in cylindrical
polar coordinates, occupies the region ρ<a.
(a) Show that a uniform current density (0 ,C,0) and a magnetic field (0 ,0,B),
with Bconstant (= B0)f o r ρ>a andB=B(ρ)f o r ρ<a , are consistent
with this equation. Obtain expressions for CandB(ρ)i nt e r m so f B0anda,
given that Bis continuous at ρ=a.
(b) The magnetic field can be expressed as B=∇×A,w h e r e Ais known as the
vector potential. Show that a suitable Acan be found which has only one
non-vanishing component, Aφ(ρ), and obtain explicit expressions for Aφ(ρ)
for both ρ<a andρ>a.L i k e B, the vector potential is continuous at ρ=a.
(c) The gas pressure p(ρ) satisfies the hydrostatic equation ∇p=J×Band
vanishes at the outer wall of the cylinder. Find a general expression for p.
10.18 (a) For cylindrical polar coordinates ρ, φ, z evaluate the derivatives of the three
unit vectors with respect to each of the coordinates, showing that only ∂ˆeρ/∂φ
and∂ˆeφ/∂φare non-zero.
(i) Hence evaluate ∇2awhenais the vector ˆeρ, i.e. a vector of unit magnitude
everywhere directed radially outwards from the z-axis.
(ii) Note that it is trivially obvious that ∇×a=0and hence that equation
(10.41) requires that ∇(∇·a)=∇2a.
(iii) Evaluate ∇(∇·a) and show that the latter equation holds, but that
[∇(∇·a)]ρ/negationslash=∇2aρ.
378
10.11 EXERCISES
(b) Rework the same problem in Cartesian coordinates (where, as it happens,
the algebra is more complicated).
10.19 Maxwell’s equations for electromagnetism in free space (i.e. in the absence of
charges, currents and dielectric or magnetic media) can be written
(i)∇·B=0, (ii)∇·E=0,
(iii)∇×E+∂B
∂t=0,(iv)∇×B−1
c2∂E
∂t=0.
A vector Ais defined by B=∇×A,a n das c a l a r φbyE=−∇φ−∂A/∂t. Show
that if the condition
(v)∇·A+1
c2∂φ
∂t=0
is imposed (this is known as choosing the Lorenz gauge), then both Aandφ
satisfy the wave equations
(vi)∇2φ−1
c2∂2φ
∂t2=0,
(vii)∇2A−1
c2∂2A
∂t2=0.
The reader is invited to proceed as follows.
(a) Verify that the expressions for BandEin terms of Aandφare consistent
with (i) and (iii).
(b) Substitute for Ein (ii) and use the derivative with respect to time of (v) to
eliminate Afrom the resulting expression. Hence obtain (vi).
(c) Substitute for BandEin (iv) in terms of Aandφ. Then use the divergence
of (v) to simplify the resulting equation and so obtain (vii).
10.20 For a description in spherical polar coordinates with axial symmetry of the flow
of a very viscous fluid, the components of the velocity field uare given in terms
of the stream function ψby
ur=1
r2sinθ∂ψ
∂θ,u θ=−1
rsinθ∂ψ
∂r.
Find an explicit expression for the differential operator Edefined by
Eψ=−(rsinθ)(∇×u)φ.
The stream function satisfies the equation of motion E2ψ= 0 and, for the flow of
a fluid past a sphere, takes the form ψ(r,θ)=f(r)sin2θ. Show that f(r)s a t i s fi e s
the (ordinary) differential equation
r4f(4)−4r2f/prime/prime+8rf/prime−8f=0.
10.21 Paraboloidal coordinates u, v, φ are defined in terms of Cartesian coordinates by
x=uvcosφ, y =uvsinφ, z =1
2(u2−v2).
Identify the coordinate surfaces in the u, v, φ system. Verify that each coordinate
surface ( u= constant, say) intersects every coordinate surface on which one of
the other two coordinates ( v, say) is constant. Show further that the system of
coordinates is an orthogonal one and determine its scale factors. Prove that theu-component of ∇×ais given by
1
(u2+v2)1/2
/aφ
v+∂aφ
∂v
/
−1
uv∂av
∂φ.
379
VECTOR CALCULUS
10.22 Non-orthogonal curvilinear coordinates are difficult to work with and should be
avoided if at all possible, but the following example is provided to illustrate thecontent of section 10.10.
In a new coordinate system for the region of space in which the Cartesian
coordinate zsatisfies z≥0, the position of a point ris given by ( α
1,α2,R), where
α1andα2are respectively the cosines of the angles made by rwith the x-a n d y-
coordinate axes of a Cartesian system and R=|r|. The ranges are −1≤αi≤1,
0≤R<∞.
(a) Express rin terms of α1,α2,Rand the unit Cartesian vectors i,j,k.
(b) Obtain expressions for the vectors ei(=∂r/∂α1,...) and hence show that the
scale factors hiare given by
h1=R(1−α2
2)1/2
(1−α2
1−α2
2)1/2,h 2=R(1−α2
1)1/2
(1−α2
1−α2
2)1/2,h 3=1.
(c) Verify formally that the system is not an orthogonal one.
(d) Show that the volume element of the coordinate system is
dV=R2dα1dα2dR
(1−α2
1−α2
2)1/2,
and demonstrate that this is always less than or equal to the corresponding
expression for an orthogonal curvilinear system.
(e) Calculate the expression for ( ds)2for the system, and show that it differs
from that for the corresponding orthogonal system by
2α1α2R2
1−α2
1−α2
2dα1dα2.
10.23 Hyperbolic coordinates u, v, φ are defined in terms of Cartesian coordinates by
x=c o s h ucosvcosφ, y =c o s h ucosvsinφ, z =s i n h usinv.
Sketch the coordinate curves in the φ= 0 plane, showing that far from the origin
they become concentric circles and radial lines. In particular, identify the curvesu=0,v=0,v=π/2a n d v=π. Calculate the tangent vectors at a general
point, show that they are mutually orthogonal and deduce that the appropriatescale factors are
h
u=hv=( c o s h2u−cos2v)1/2,h φ=c o s h ucosv.
Find the most general function ψ(u)o fuonly that satisfies Laplace’s equation
∇2ψ=0 .
10.24 In a Cartesian system, AandBare the points (0 ,0,−1) and (0 ,0,1) respectively.
In a new coordinate system a general point Pis given by ( u1,u2,u3)w i t h
u1=1
2(r1+r2),u2=1
2(r1−r2),u3=φ;h e r e r1andr2are the distances APand
BPandφis the angle between the plane ABPandy=0 .
(a) Express zand the perpendicular distance ρfrom Pto the z-axis in terms of
u1,u2,u3.
(b) Evaluate ∂x/∂u i,∂y/∂u i,∂z/∂u i,f o r i=1,2,3.
(c) Find the Cartesian components of ˆujand hence show that the new coordi-
nates are mutually orthogonal. Evaluate the scale factors and the infinitesimalvolume element in the new coordinate system.
(d) Determine and sketch the forms of the surfaces u
i=c o n s t a n t .
(e) Find the most general function fofu1only that satisfies ∇2f=0 .
380
10.12 HINTS AND ANSWERS
10.12 Hints and answers
10.1 a×(a×b)+h.
10.2 Taking E0=iandB0=j,E=( 1+ t)i+(t2/2+t3/6)j−(t+t2/2)k,
B=(t2/2+t3/6)i+( 1+ t)j+(t+t2/2)k.
10.3 For crossed uniform fields ¨x+(Bq/m)2x=q(E−Bv0)/m,¨y=0 ,m˙z=qBx+mv0;
(b)ξ=Bqt/m ; the path is a cycloid in the plane y=0 ; ds=[ (dx/dt)2+
(dz/dt)2]1/2dt.
10.4 Prove that the vector equation of the stone is r=v0t+gt2/2. Impose the condition
r·˙r>0f o ra l l t,i . e .r·˙r=0h a sn or e a lr o o t sf o r t;8v2
0g2>9(v0·g)2. Maximum
angle is 70 .5◦.
10.5 g=¨r/prime−ω×(ω×r), where ¨r/primeis the shell’s acceleration measured by an observer
fixed in space. To first order in ω, the direction of gis radial, i.e. parallel to ¨r/prime.
(a) Note that sis orthogonal to g.
(b) If the actual time of flight is T,u s e( s+∆ )·g=0t os h o wt h a t
T≈τ(1 + 2 g−2(g×ω)·v+···).
In the Coriolis terms it is sufficient to put T≈τ.
(c) For this situation ( g×ω)·v=0a n d ω×v=0;τ≈43 s and ∆ = 10–15 m
to the East.
10.6 Differentiate ˆb=ˆt׈nwith respect to s; express the result in terms of the
derivatives of r; take the scalar product with d2r/ds2.
10.7 (a) Evaluate ( dr/du)·(dr/du).
(b) Integrate the previous result between u=0a n d u=1 .
(c)ˆt=[√2(1 + u2)]−1[(1−u2)i+2uj+( 1+ u2)k]. Use dˆt/ds=(dˆt/du)/(ds/du);
ρ−1=|dˆt/ds|.
(d)ˆn=( 1+ u2)−1[−2ui+(1−u2)j].ˆb=[√2(1+ u2)]−1[(u2−1)i−2uj+(1+ u2)k].
Usedˆb
ds=dˆb
du
/ds
duand show that this equals −[3a(1 +u2)2]−1ˆn.
(e) Show that dˆn/ds=τ(ˆb−ˆt)=−2[3√2a(1 +u2)3]−1[(1−u2)i+2uj].
10.8 ds/dθ =√
2h/(πsinθcosθ);ˆt=−sin2θi+c o s2θj+√
2si nθcosθk;
ˆb=c o s2θi−sin2θj+√
2si nθcosθk.
10.9 Note that dB=(dr·∇)Band that B=Bˆt,w i t h ˆt=dr/ds.O b t a i n( B·∇)B/B=
ˆt(dB/ds )+ˆn(B/ρ) and then take the vector product of ˆtwith this equation.
10.10 (a) dS=|(−ucosφcot Ω ,−usinφcot Ω ,u)|dφ du;S=πa2cosec Ω .
(b)z=αu2;dS=u(1 + 4 α2u2)1/2dφ du;S=(π/6)[(1 + 4 αZ)3/2−1].
10.11 (b) Put tan φ=2−1/2sinhψ.
10.12|∇z|2=4ρ2e−2ρ2[(1−ρ2)2cos22φ+ρ2sin22φ], which is extremal when φ=nπ/4
and 1−ρ2= 0. Maximum slope = 2 e−1atx=±1,y=±1, along azimuthal
directions x±y=±2a n d x±y=∓2.
10.14 (a) ( ∇·a)(∇×a); (b) terms of the form bxcx(∂ax/∂x) cancel; (c) for the x-
component, add and subtract ax(∂ax/∂x) and regroup.
10.15 (a) 2 z(x2+y2+z2)−3[(y2+z2)(y2+z2−3x2)−4x4]. (b) 2 r−1cosθ(1−5sin2θcos2φ);
both are equal to 2 zr−4(r2−5x2).
10.17 Use the formulae given in table 10.2.
(a)C=−B0/(µ0a);B(ρ)=B0ρ/a.
(b)B0ρ2/(3a)f o r ρ<a,a n d B0[ρ/2−a2/(6ρ)] for ρ>a.
(c) [ B2
0/(2µ0)][1−(ρ/a)2].
10.18 (a) ∂ˆeρ/∂φ=ˆeφ,∂ˆeφ/∂φ=−ˆeρ;( i )−ρ−2ˆeρ.( b )∇2a=−(x2+y2)−3/2(xi+yj).
381
VECTOR CALCULUS
10.20 E=∂2
∂r2+sinθ
r2∂
∂θ
/1
sinθ∂
∂θ
/
.
10.21 Two sets of paraboloids of revolution about the z-axis and the sheaf of planes
containing the z-axis. For constant u,−∞<z<u2/2; for constant v,−v2/2<
z<∞. The scale factors are hu=hv=(u2+v2)1/2,hφ=uv.
10.22 (c) e1·e2=R2α1α2/(1−α2
1−α2
2)/negationslash=0 .
10.23 The tangent vectors are as follows: for u= 0, the line joining (1 ,0,0) and
(−1,0,0); for v= 0, the line joining (1 ,0,0) and (∞,0,0). For v=π/2, the line
(0,0,z); for v=π, the line joining ( −1,0,0) and (−∞,0,0).ψ(u)=2t a n−1eu+c,
derived from ∂[cosh u(∂ψ/∂u )]/∂u=0 .
10.24 (a) z=u1u2,ρ=u2
1+u2
2−u2
1u22−1.
(b)u1(1−u2
2)c osu3/ρ,u1(1−u2
2)si nu3/ρ,u2;u2(1−u2
1)c osu3/ρ,u2(1−u2
1)si nu3/ρ,
u1;−ρsinu3,ρcosu3,0 .
(c) [( u2
1−u2
2)/(u2
1−1)]1/2,[ (u2
2−u2
1)/(u2
2−1)]1/2,ρ;|u2
1−u2
2|du1du2du3.
(d) Confocal ellipsoids, hyperboloids, half-planes containing the z-axis.
(e)Bln[(u1−1)/(u1+ 1)].
382
11
Line, surface and volume integrals
In the previous chapter we encountered continuously varying scalar and vector
fields and discussed the action of various differential operators on them. Inaddition to these differential operations, the need often arises to consider theintegration of field quantities along lines, over surfaces and throughout volumes.
In general the integrand may be scalar or vector in nature, but the evaluation
of such integrals involves their reduction to one or more scalar integrals, whichare then evaluated. In the case of surface and volume integrals this requires theevaluation of double and triple integrals (see chapter 6).
11.1 Line integrals
In this section we discuss lineorpath integrals , in which some quantity related
to the field is integrated between two given points in space, AandB, along a
prescribed curve Cthat joins them. In general, we may encounter line integrals
of the forms integraldisplay
Cφdr,integraldisplay
Ca·dr,integraldisplay
Ca×dr, (11.1)
where φis a scalar field and ais a vector field. The three integrals themselves are
respectively vector, scalar and vector in nature. As we will see below, in physical
applications line integrals of the second type are by far the most common.
The formal definition of a line integral closely follows that of ordinary integrals
and can be considered as the limit of a sum. We may divide the path Cjoining
the points AandBintoNsmall line elements ∆ rp,p=1,...,N .I f(xp,yp,zp)i s
any point on the line element ∆ rpthen the second type of line integral in (11.1),
for example, is defined as
integraldisplay
Ca·dr= lim
N→∞Nsummationdisplay
p=1a(xp,yp,zp)·∆rp,
where it is assumed that all |∆rp|→0a sN→∞.
383
LINE, SURFACE AND VOLUME INTEGRALS
Each of the line integrals in (11.1) is evaluated over some curve Cthat may be
either open ( AandBbeing distinct points) or closed (the curve Cforms a loop,
so that AandBare coincident). In the case where Cis closed, the line integral
is writtencontintegraltext
Cto indicate this. The curve may be given either parametrically by
r(u)=x(u)i+y(u)j+z(u)kor by means of simultaneous equations relating x, y, z
for the given path (in Cartesian coordinates). A full discussion of the differentrepresentations of space curves was given in section 10.3.
In general, the value of the line integral depends not only on the end-points
AandBbut also on the path Cjoining them. For a closed curve we must also
specify the direction around the loop in which the integral is taken. It is usuallytaken to be such that a person walking around the loop Cin this direction
always has the region Ron his/her left; this is equivalent to traversing Cin the
anticlockwise direction (as viewed from above).
11.1.1 Evaluating line integrals
The method of evaluating a line integral is to reduce it to a set of scalar integrals.
It is usual to work in Cartesian coordinates, in which case dr=dxi+dyj+dzk.
The first type of line integral in (11.1) then becomes simply
integraldisplay
Cφdr=iintegraldisplay
Cφ(x, y, z)dx+jintegraldisplay
Cφ(x, y, z)dy+kintegraldisplay
Cφ(x, y, z)dz.
The three integrals on the RHS are ordinary scalar integrals that can be evaluated
in the usual way once the path of integration Chas been specified. Note that in
the above we have used relations of the formintegraldisplay
φidx=iintegraldisplay
φd x ,
which is allowable since the Cartesian unit vectors are of constant magnitude
and direction and hence may be taken out of the integral. If we had been usinga different coordinate system, such as spherical polars, then, as we saw in thelast chapter, the unit basis vectors would not be constant. In that case the basisvectors could not be factorised out of the integral.
The second and third line integrals in (11.1) can also be reduced to a set of
scalar integrals by writing the vector field ain terms of its Cartesian components
asa=a
xi+ayj+azk,w h e r e ax,ay,azare each (in general) functions of x, y, z .
The second line integral in (11.1), for example, can then be written as
integraldisplay
Ca·dr=integraldisplay
C(axi+ayj+azk)·(dxi+dyj+dzk)
=integraldisplay
C(axdx+aydy+azdz)
=integraldisplay
Caxdx+integraldisplay
Caydy+integraldisplay
Cazdz. (11.2)
384
11.1 LINE INTEGRALS
A similar procedure may be followed for the third type of line integral in (11.1).
Line integrals have properties that are analogous to those of ordinary integrals.
In particular, the following are useful properties (which we illustrate using the
second form of line integral in (11.1) but which are valid for all three types).
(i) Reversing the path of integration changes the sign of the integral. If the
path Calong which the line integrals are evaluated has AandBas its
end-points then
integraldisplayB
Aa·dr=−integraldisplayA
Ba·dr.
This implies that if the path Cis a loop then integrating around the loop
in the opposite direction changes the sign of the integral.
(ii) If the path of integration is subdivided into smaller segments then the sum
of the separate line integrals along each segment is equal to the line integralalong the whole path. So, if Pis any point on the path of integration that
lies between the path’s end-points AandBthen
integraldisplay
B
Aa·dr=integraldisplayP
Aa·dr+integraldisplayB
Pa·dr.IEvaluate the line integral I=
R
Ca·dr,w h e r e a=(x+y)i+(y−x)j, along each of the
paths in the xy-plane shown in figure 11.1, namely
(i) the parabola y2=xfrom(1,1)to(4,2),
(ii) the curve x=2u2+u+1,y=1+ u2from(1,1)to(4,2),
(iii) the line y=1from(1,1)to(4,1), followed by the line x=4 from(4,1)
to(4,2).
Since each of the paths lies entirely in the xy-plane, we have dr=dxi+dyj.W ec a n
therefore write the line integral as
I=
Z
Ca·dr=
Z
C[(x+y)dx+(y−x)dy]. (11.3)
We must now evaluate this line integral along each of the prescribed paths.
Case (i) . Along the parabola y2=xwe have 2 yd y=dx. Substituting for xin (11.3)
and using just the limits on y,w eo b t a i n
I=
Z(4,2)
(1,1)[(x+y)dx+(y−x)dy]=
Z2
1[(y2+y)2y+(y−y2)]dy=1 11
3.
Note that we could just as easily have substituted for yand obtained an integral in x,
which would have given the same result.
Case (ii) . The second path is given in terms of a parameter u. We could eliminate u
between the two equations to obtain a relationship between xandydirectly and proceed
as above, but it is usually quicker to write the line integral in terms of the parameter u.
Along the curve x=2u2+u+1 , y=1+ u2we have dx=( 4u+1 )duanddy=2ud u.
385
LINE, SURFACE AND VOLUME INTEGRALS
y
x(i)
(ii)
(iii) (1,1)(4,2)
Figure 11.1 Different possible paths between the points (1, 1) and (4, 2).
Substituting for xandyin (11.3) and writing the correct limits on u,w eo b t a i n
I=
Z(4,2)
(1,1)[(x+y)dx+(y−x)dy]
=
Z1
0[(3u2+u+ 2)(4 u+1 )−(u2+u)2u]du=1 02
3.
Case (iii) . For the third path the line integral must be evaluated along the two line
segments separately and the results added together. First, along the line y= 1 we have
dy= 0. Substituting this into (11.3) and using just the limits on xf o rt h i ss e g m e n t ,w e
obtainZ(4,1)
(1,1)[(x+y)dx+(y−x)dy]=
Z4
1(x+1 )dx=1 01
2.
Next, along the line x= 4 we have dx= 0. Substituting this into (11.3) and using just the
limits on yfor this segment, we obtainZ(4,2)
(4,1)[(x+y)dx+(y−x)dy]=
Z2
1(y−4)dy=−21
2.
The value of the line integral along the whole path is just the sum of the values of the line
integrals along each segment, and is given by I=1 01
2−21
2=8 .
J
When calculating a line integral along some curve C, which is given in terms
ofx,yandz, we are sometimes faced with the problem that the curve Cis such
that x,yandzare not single-valued functions of one another over the entire
length of the curve. This is a particular problem for closed loops in the xy-plane
(and also for some open curves). In such cases the path may be subdivided intoshorter line segments along which one coordinate is a single-valued function of
the other two. The sum of the line integrals along these segments is then equal
to the line integral along the entire curve C. A better solution, however, is to
represent the curve in a parametric form r(u) that is valid for its entire length.
386
11.1 LINE INTEGRALSIEvaluate the line integral I=
H
Cxd y,w h e r e Cis the circle in the xy-plane defined by
x2+y2=a2,z=0.
Adopting the usual convention mentioned above, the circle Cis to be traversed in the
anticlockwise direction. Taking the circle as a whole means xis not a single-valued
function of y. We must therefore divide the path into two parts with x=+
p
a2−y2for
the semicircle lying to the right of x=0 ,a n d x=−
p
a2−y2for the semicircle lying to
the left of x= 0. The required line integral is then the sum of the integrals along the two
semicircles. Substituting for x,i ti sg i v e nb y
I=
I
Cxd y=
Za
−a
p
a2−y2dy+
Z−a
a
/
−
p
a2−y2
/
dy
=4
Za
0
p
a2−y2dy=πa2.
Alternatively, we can represent the entire circle parametrically, in terms of the azimuthal
angle φ,s ot h a t x=acosφandy=asinφwith φrunning from 0 to 2 π. The integral can
therefore be evaluated over the whole circle at once. Noting that dy=acosφd φ,w ec a n
rewrite the line integral completely in terms of the parameter φand obtain
I=
I
Cxd y=a2
Z2π
0cos2φd φ=πa2.
J
11.1.2 Physical examples of line integrals
There are many physical examples of line integrals, but perhaps the most common
is the expression for the total work done by a force Fwhen it moves its point
of application from a point Ato a point Balong a given curve C. We allow the
magnitude and direction of Fto vary along the curve. Let the force act at a point
rand consider a small displacement dralong the curve; then the small amount
of work done is dW=F·dr, as discussed in subsection 7.6.1 (note that dWcan
be either positive or negative). Therefore, the total work done in traversing thepath Cis
W
C=integraldisplay
CF·dr.
Naturally, other physical quantities can be expressed in such a way. For example,
the electrostatic potential energy gained by moving a charge qalong a path Cin
an electric field Eis−qintegraltext
CE·dr. We may also note that Amp `ere’s law concerning
the magnetic field Bassociated with a current-carrying wire can be written as
contintegraldisplay
CB·dr=µ0I,
where Iis the current enclosed by a closed path Ctraversed in a right-handed
sense with respect to the current direction.
Magnetostatics also provides a physical example of the third type of line
387
LINE, SURFACE AND VOLUME INTEGRALS
integral in (11.1). If a loop of wire Ccarrying a current Iis placed in a magnetic
fieldBthen the force dFon a small length drof the wire is given by dF=Idr×B,
and so the total (vector) force on the loop is
F=Icontintegraldisplay
Cdr×B.
11.1.3 Line integrals with respect to a scalar
In addition to those listed in (11.1), we can form other types of line integral,
which depend on a particular curve Cbut for which we integrate with respect
to a scalar du, rather than the vector differential dr. This distinction is somewhat
arbitrary, however, since we can always rewrite line integrals containing the vectordifferential dras a line integral with respect to some scalar parameter. If the path
Calong which the integral is taken is described parametrically by r(u)t h e n
dr=dr
dudu,
and the second type of line integral in (11.1), for example, can be written as
integraldisplay
Ca·dr=integraldisplay
Ca·dr
dudu.
A similar procedure can be followed for the other types of line integral in (11.1).
Commonly occurring special cases of line integrals with respect to a scalar are
integraldisplay
Cφd s ,integraldisplay
Cads,
where sis the arc length along the curve C. We can always represent Cparamet-
rically by r(u), and from section 10.3 we have
ds=radicalbigg
dr
du·dr
dudu.
The line integrals can therefore be expressed entirely in terms of the parameter u
and thence evaluated.IEvaluate the line integral I=
R
C(x−y)2ds,w h e r e Cis the semicircle of radius arunning
from A=(a,0)toB=(−a,0)and for which y≥0.
The semicircular path from AtoBcan be described in terms of the azimuthal angle φ
(measured from the x-axis) by
r(φ)=acosφi+asinφj,
where φruns from 0 to π. Therefore the element of arc length is given, from section 10.3,
by
ds=
s
dr
dφ·dr
dφdφ=a(cos2φ+s i n2φ)dφ=ad φ .
388
11.2 CONNECTIVITY OF REGIONS
(a)( b)( c)
Figure 11.2 ( a) A simply connected region; ( b) a doubly connected region;
(c) a triply connected region.
Since ( x−y)2=a2(1−sin 2φ), the line integral becomes
I=
Z
C(x−y)2ds=
Zπ
0a3(1−sin2φ)dφ=πa3.
J
As discussed in the previous chapter, the expression (10.58) for the square of
the element of arc length in three-dimensional orthogonal curvilinear coordinates
u1,u2,u3is
(ds)2=h2
1(du1)2+h2
2(du2)2+h2
3(du3)2,
where h1,h2,h3are the scale factors of the coordinate system. If a curve Cin
three dimensions is given parametrically by the equations ui=ui(λ)f o r i=1,2,3
then the element of arc length along the curve is
ds=radicalBigg
h2
1parenleftbiggdu1
dλparenrightbigg2
+h2
2parenleftbiggdu2
dλparenrightbigg2
+h2
3parenleftbiggdu3
dλparenrightbigg2
dλ.
11.2 Connectivity of regions
In physical systems it is usual to define a scalar or vector field in some region R.
In the next and some later sections we will need the concept of the connectivity
of such a region in both two and three dimensions.
We begin by discussing planar regions. A plane region Ris said to be simply
connected if every simple closed curve within Rcan be continuously shrunk to
a point without leaving the region (see figure 11.2( a)). If, however, the region
Rcontains a hole then there exist simple closed curves that cannot by shrunk
to a point without leaving R(see figure 11.2( b) ) .S u c har e g i o ni ss a i dt ob e
doubly connected, since its boundary has two distinct parts. Similarly, a regionwith n−1 holes is said to be n-fold connected ,o rmultiply connected (the region
in figure 11.2( c) is triply connected).
These ideas can be extended to regions that are not planar, such as general
389
LINE, SURFACE AND VOLUME INTEGRALS
y
d
c
a b xSR
TCUV
Figure 11.3 A simply connected region Rbounded by the curve C.
three-dimensional surfaces and volumes. The same criteria concerning the shrink-
ing of closed curves to a point also apply when deciding the connectivity of suchregions. In these cases, however, the curves must lie in the surface or volumein question. For example, the interior of a torus is not simply connected, sincethere exist closed curves in the interior that cannot be shrunk to a point without
leaving the torus. On the other hand, the region between two concentric spheres
of different radii is simply connected.
11.3 Green’s theorem in a plane
In subsection 11.1.1 we considered (amongst other things) the evaluation of line
integrals for which the path Cis closed and lies entirely in the xy-plane. Since
the path is closed it will enclose a region Rof the plane. We now discuss how to
express the line integral around the loop as a double integral over the enclosed
region R.
Suppose the functions P(x, y),Q(x, y) and their partial derivatives are single-
valued, finite and continuous inside and on the boundary Cof some simply
connected region Rin the xy-plane. Green’s theorem in a plane then states
contintegraldisplay
C(Pd x+Qd y)=integraldisplayintegraldisplay
Rparenleftbigg∂Q
∂x−∂P
∂yparenrightbigg
dx dy, (11.4)
and so relates the line integral around Cto a double integral over the enclosed
region R. This theorem may be proved straightforwardly in the following way.
Consider the simply connected region Rin figure 11.3, and let y=y1(x)a n d
y=y2(x) be the equations of the curves STU andSVU respectively. We then
390
11.3 GREEN’S THEOREM IN A PLANE
write
integraldisplayintegraldisplay
R∂P
∂ydx dy =integraldisplayb
adxintegraldisplayy2(x)
y1(x)dy∂P
∂y=integraldisplayb
adxbracketleftBig
P(x, y)bracketrightBigy=y2(x)
y=y1(x)
=integraldisplayb
abracketleftBig
P(x, y2(x))−P(x, y1(x))bracketrightBig
dx
=−integraldisplayb
aP(x, y1(x))dx−integraldisplaya
bP(x, y2(x))dx=−contintegraldisplay
CPd x .
If we now let x=x1(y)a n d x=x2(y) be the equations of the curves TSV and
TUV respectively, we can similarly show that
integraldisplayintegraldisplay
R∂Q
∂xdx dy =integraldisplayd
cdyintegraldisplayx2(y)
x1(y)dx∂Q
∂x=integraldisplayd
cdybracketleftBig
Q(x, y)bracketrightBigx=x2(y)
x=x1(y)
=integraldisplayd
cbracketleftBig
Q(x2(y),y)−Q(x1(y),y)bracketrightBig
dy
=integraldisplayc
dQ(x1,y)dy+integraldisplayd
cQ(x2,y)dy=contintegraldisplay
CQd y.
Subtracting these two results gives Green’s theorem in a plane.IShow that the area of a region Renclosed by a simple closed curve Cis given by A=
1
2
H
C(xd y−yd x)=
H
Cxd y=−
H
Cyd x. Hence calculate the area of the ellipse x=acosφ,
y=bsinφ.
In Green’s theorem (11.4) put P=−yandQ=x;t h e nI
C(xd y−yd x)=
ZZ
R(1 + 1) dx dy =2
ZZ
Rdx dy =2A.
Therefore the area of the region is A=1
2
H
C(xd y−yd x). Alternatively, we could put P=0
andQ=xand obtain A=
H
Cxd y, or put P=−yandQ= 0, which gives A=−
H
Cyd x.
The area of the ellipse x=acosφ,y=bsinφis given by
A=1
2
I
C(xd y−yd x)=1
2
Z2π
0ab(cos2φ+s i n2φ)dφ
=ab
2
Z2π
0dφ=πab.
J
It may further be shown that Green’s theorem in a plane is also valid for
multiply connected regions. In this case, the line integral must be taken overall the distinct boundaries of the region. Furthermore, each boundary must be
traversed in the positive direction, such that a person travelling along it in this
direction always has the region Ron their left. In order to apply Green’s theorem
to the region Rshown in figure 11.4, the line integrals must be taken over
391
LINE, SURFACE AND VOLUME INTEGRALS
y
xC1C2R
Figure 11.4 A doubly connected region Rbounded by the curves C1andC2.
both boundaries, C1andC2, in the directions indicated, and the results added
together.
We may also use Green’s theorem in a plane to investigate the path indepen-
dence (or not) of line integrals when the paths lie in the xy-plane. Let us consider
the line integral
I=integraldisplayB
A(Pd x+Qd y).
For the line integral from AtoBto be independent of the path taken, it must
have the same value along any two arbitrary paths C1andC2joining the points.
Moreover, if we consider as the path the closed loop Cformed by C1−C2then
the line integral around this loop must be zero. From Green’s theorem in a plane,(11.4), we see that a sufficient condition for I=0i st h a t
∂P
∂y=∂Q
∂x, (11.5)
throughout some simply connected region Rcontaining the loop, where we assume
that these partial derivatives are continuous in R.
It may be shown that (11.5) is also a necessary condition for I=0a n di s
equivalent to requiring Pd x+Qd yto be an exact differential of some function
φ(x, y) such that Pd x+Qd y=dφ. It follows thatintegraltextB
A(Pd x+Qd y)=φ(B)−φ(A)
and thatcontintegraltext
C(Pd x+Qd y) around any closed loop Cin the region Ris identically
zero. These results are special cases of the general results for paths in threedimensions, which are discussed in the next section.
392
11.4 CONSERVATIVE FIELDS AND POTENTIALSIEvaluate the line integral
I=
I
C[(exy+c o s xsiny)dx+(ex+s i n xcosy)dy],
around the ellipse x2/a2+y2/b2=1.
Clearly, it is not straightforward to calculate this line integral directly. However, if we let
P=exy+c o s xsiny and Q=ex+s i n xcosy,
then ∂P/∂y =ex+c o s xcosy=∂Q/∂x ,a n ds o Pd x+Qd yis an exact differential (it
is actually the differential of the function f(x, y)=exy+s i n xsiny). From the above
discussion, we therefore immediately conclude that I=0 .
J
11.4 Conservative fields and potentials
So far we have made the point that, in general, the value of a line integral
between two points AandBdepends on the path Ctaken from AtoB.I nt h e
previous section, however, we saw that, for paths in the xy-plane, line integrals
whose integrands have certain properties are independent of the path taken. We
now extend that discussion to the full three-dimensional case.
For line integrals of the formintegraltext
Ca·dr, there exists a class of vector fields for
which the line integral between two points is independent of the path taken. Such
vector fields are called conservative . A vector field athat has continuous partial
derivatives in a simply connected region Ris conservative if, and only if, any of
the following is true.
(i) The integralintegraltextB
Aa·dr,w h e r e AandBlie in the region R, is independent of
the path from AtoB. Hence the integralcontintegraltext
Ca·draround any closed loop
inRis zero.
(ii) There exists a single-valued function φof position such that a=∇φ.
(iii)∇×a=0.
(iv)a·dris an exact differential.
The validity or otherwise of any of these statements implies the same for the
other three, which we will now show.
First, let us assume that (i) above is true. If the line integral from AtoB
is independent of the path taken between the points then its value must be afunction only of the positions of AandB. We may therefore write
integraldisplay
B
Aa·dr=φ(B)−φ(A), (11.6)
which defines a single-valued scalar function of position φ. If the points AandB
are separated by an infinitesimal displacement drthen (11.6) becomes
a·dr=dφ,
393
LINE, SURFACE AND VOLUME INTEGRALS
which shows that we require a·drto be an exact differential: condition (iv). From
(10.27) we can write dφ=∇φ·dr, and so we have
(a−∇φ)·dr=0.
Since dris arbitrary, we find that a=∇φ; this immediately implies ∇×a=0,
condition (iii) (see (10.37)).
Alternatively, if we suppose that there exists a single-valued function of position
φsuch that a=∇φthen∇×a=0follows as before. The line integral around a
closed loop then becomes
contintegraldisplay
Ca·dr=contintegraldisplay
C∇φ·dr=contintegraldisplay
dφ.
Since we defined φto be single-valued, this integral is zero as required.
Now suppose ∇×a=0. From Stoke’s theorem, which is discussed in sec-
tion 11.9, we immediately obtaincontintegraltext
Ca·dr=0 ;t h e n a=∇φanda·dr=dφfollow
as above.
Finally, let us suppose a·dr=dφ. Then immediately we have a=∇φ,a n dt h e
other results follow as above.IEvaluate the line integral I=
RB
Aa·dr,w h e r e a=(xy2+z)i+(x2y+2 )j+xk,Ais the
point(c, c, h)andBis the point (2c, c/2,h), along the different paths
(i)C1, given by x=cu,y=c/u,z=h,
(ii)C2, given by 2y=3c−x,z=h.
Show that the vector field ais in fact conservative, and find φsuch that a=∇φ.
Expanding out the integrand, we have
I=
Z(2c, c/2,h)
(c, c, h)
/
(xy2+z)dx+(x2y+2 )dy+xd z
/
, (11.7)
which we must evaluate along each of the paths C1andC2.
(i) Along C1we have dx=cd u,dy=−(c/u2)du,dz= 0, and on substituting in (11.7)
and finding the limits on u,w eo b t a i n
I=
Z2
1c
/
h−2
u2
/
du=c(h−1).
(ii) Along C2we have 2 dy=−dx,dz= 0 and, on substituting in (11.7) and using the
limits on x,w eo b t a i n
I=
Z2c
c
/;1
2x3−9
4cx2+9
4c2x+h−1
/
dx=c(h−1).
Hence the line integral has the same value along paths C1andC2. Taking the curl of a,
we have
∇×a=( 0−0)i+( 1−1)j+( 2xy−2xy)k=0,
soais a conservative vector field, and the line integral between two points must be
394
11.5 SURFACE INTEGRALS
independent of the path taken. Since ais conservative, we can write a=∇φ. Therefore, φ
must satisfy
∂φ
∂x=xy2+z,
which implies that φ=1
2x2y2+zx+f(y,z) for some function f. Secondly, we require
∂φ
∂y=x2y+∂f
∂y=x2y+2,
which implies f=2y+g(z). Finally, since
∂φ
∂z=x+∂g
∂z=x,
we have g=c o n s t a n t= k. It can be seen that we have explicitly constructed the function
φ=1
2x2y2+zx+2y+k.
J
The quantity φthat figures so prominently in this section is called the scalar
potential function o ft h ec o n s e r v a t i v ev e c t o rfi e l d a(which satisfies ∇×a=0), and
is unique up to an arbitrary additive constant. Scalar potentials that are multi-valued functions of position (but in simple ways) are also of value in describingsome physical situations, the most obvious example being the scalar magnetic
potential associated with a current-carrying wire. When the integral of a field
quantity around a closed loop is considered, provided the loop does not enclosea net current, the potential is single-valued and all the above results still hold. Ifthe loop does enclose a net current, however, our analysis is no longer valid andextra care must be taken.
If, instead of being conservative, a vector field bsatisfies∇·b=0( i . e . b
is solenoidal) then it is both possible and useful, for example in the theory ofelectromagnetism, to define a vector potential field asuch that b=∇×a.I tm a y
be shown that such a vector field aalways exists. Further, if ais one such vector
field then a
/prime=a+∇ψ+c,w h e r e ψis any scalar function and cis any constant
vector, also satisfies the above relationship, i.e. b=∇×a/prime. This was discussed
more fully in subsection 10.8.2.
11.5 Surface integrals
As with line integrals, integrals over surfaces can involve vector and scalar fields
and, equally, can result in either a vector or a scalar. The simplest case involvesentirely scalars and is of the form
integraldisplay
Sφd S. (11.8)
As analogues of the line integrals listed in (11.1), we may also encounter surface
integrals involving vectors, namely
integraldisplay
SφdS,integraldisplay
Sa·dS,integraldisplay
Sa×dS. (11.9)
395
LINE, SURFACE AND VOLUME INTEGRALS
SS
V
CdS
dS
(a)( b)
Figure 11.5 ( a) A closed surface and ( b) an open surface. In each case a
normal to the surface is shown: dS=ˆndS.
All the above integrals are taken over some surface S, which may be either
open or closed, and are therefore, in general, double integrals. Following thenotation for line integrals, for surface integrals over a closed surfaceintegraltext
Sis replaced
bycontintegraltext
S.
The vector differential dSin (11.9) represents a vector area element of the
surface S. It may also be written dS=ˆndS,w h e r e ˆnis a unit normal to the
surface at the position of the element and dSis the scalar area of the element used
in (11.8). The convention for the direction of the normal ˆnto a surface depends
on whether the surface is open or closed. A closed surface, see figure 11.5( a),
does not have to be simply connected (for example, the surface of a torus is not),but it does have to enclose a volume V, which may be of infinite extent. The
direction of ˆnis taken to point outwards from the enclosed volume as shown.
An open surface, see figure 11.5( b), spans some perimeter curve C. The direction
ofˆnis then given by the right-hand sense with respect to the direction in which
the perimeter is traversed, i.e. follows the right-hand screw rule discussed in
section 7.6.2. An open surface does not have to be simply connected but for
our purposes it must be two-sided (a M ¨obius strip is an example of a one-sided
surface).
The formal definition of a surface integral is very similar to that of a line
integral. We divide the surface SintoNelements of area ∆ S
p,p=1,...,N ,e a c h
with a unit normal ˆnp.I f(xp,yp,zp) is any point in ∆ Spthen the second type of
surface integral in (11.9), for example, is defined as
integraldisplay
Sa·dS= lim
N→∞Nsummationdisplay
p=1a(xp,yp,zp)·ˆnp∆Sp,
where it is required that all ∆ Sp→0a sN→∞.
396
11.5 SURFACE INTEGRALS
xz
y
RdAα
SkdS
Figure 11.6 A surface S(or part thereof) projected onto a region Rin the
xy-plane; dSis the surface element at a point P.
11.5.1 Evaluating surface integrals
We now consider how to evaluate surface integrals over some general surface. This
involves writing the scalar area element dSin terms of the coordinate differentials
of our chosen coordinate system. In some particularly simple cases this is verystraightforward. For example, if Sis the surface of a sphere of radius a(or some
part thereof) then using spherical polar coordinates θ,φon the sphere we have
dS=a
2sinθd θd φ . For a general surface, however, it is not usually possible to
represent the surface in a simple way in any particular coordinate system. In suchcases, it is usual to work in Cartesian coordinates and consider the projections ofthe surface onto the coordinate planes.
Consider a surface (or part of a surface) Sas in figure 11.6. The surface Sis
projected onto a region Rof the xy-plane, so that an element of surface area dS
at point Pprojects onto the area element dA. From the figure, we see that dA=
|cosα|dS,w h e r e αis the angle between the unit vector kin the z-direction and
the unit normal ˆnto the surface at P. So, at any given point of S, we have simply
dS=dA
|cosα|=dA
|ˆn·k|.
Now, if the surface Sis given by the equation f(x, y, z) = 0 then, as shown in
subsection 10.7.1, the unit normal at any point of the surface is simply given by
ˆn=∇f/|∇f|evaluated at that point, cf. (10.32). The scalar element of surface
area then becomes
dS=dA
|ˆn·k|=|∇f|dA
∇f·k=|∇f|dA
∂f/∂z, (11.10)
397
LINE, SURFACE AND VOLUME INTEGRALS
where|∇f|and∂f/∂z are evaluated on the surface S. We can therefore express
any surface integral over Sas a double integral over the region Rin the xy-plane.IEvaluate the surface integral I=
R
Sa·dS,w h e r e a=xiandSis the surface of the
hemisphere x2+y2+z2=a2with z≥0.
The surface of the hemisphere is shown in figure 11.7. In this case dSmay be easily
expressed in spherical polar coordinates as dS=a2sinθd θd φ , and the unit normal to the
surface at any point is simply ˆr. On the surface of the hemisphere we have x=asinθcosφ
and so
a·dS=x(i·ˆr)dS=(asinθcosφ)(sinθcosφ)(a2sinθd θd φ ).
Therefore, inserting the correct limits on θandφ, we have
I=
Z
Sa·dS=a3
Zπ/2
0dθsin3θ
Z2π
0dφcos2φ=2πa3
3.
We could, however, follow the general prescription above and project the hemisphere S
onto the region Rin the xy-plane that is a circle of radius acentred at the origin. Writing
the equation of the surface of the hemisphere as f(x, y)=x2+y2+z2−a2=0a n du s i n g
(11.10), we have
I=
Z
Sa·dS=
Z
Sx(i·ˆr)dS=
Z
Rx(i·ˆr)|∇f|dA
∂f/∂z.
Now∇f=2xi+2yj+2zk=2r,s oo nt h es u r f a c e Swe have|∇f|=2|r|=2a.O n Swe
also have ∂f/∂z =2z=2
p
a2−x2−y2andi·ˆr=x/a. Therefore, the integral becomes
I=
ZZ
Rx2p
a2−x2−y2dx dy.
Although this integral may be evaluated directly, it is quicker to transform to plane polar
coordinates:
I=
ZZ
R/primeρ2cos2φp
a2−ρ2ρd ρd φ
=
Z2π
0cos2φd φ
Za
0ρ3dρp
a2−ρ2.
Making the substitution ρ=asinu, we finally obtain
I=
Z2π
0cos2φd φ
Zπ/2
0a3sin3ud u=2πa3
3.
J
In the above discussion we assumed that any line parallel to the z-axis intersects
Sonly once. If this is not the case, we must split up the surface into smaller
surfaces S1,S2etc. that are of this type. The surface integral over Sis then the
sum of the surface integrals over S1,S2and so on. This is always necessary for
closed surfaces.
We may also sometimes wish to project a surface S(or some part of it) onto
thezx-o ryz-plane, rather than the xy-plane. In such cases, the above analysis is
easily modified.
398
11.5 SURFACE INTEGRALS
dS
Sz
Ca
a
a
xy
dA=dx dy
Figure 11.7 The surface of the hemisphere x2+y2+z2=a2,z≥0.
11.5.2 Vector areas of surfaces
The vector area of a surface Sis defined simply as
S=integraldisplay
SdS,
where the surface integral may be evaluated as above.IFind the vector area of the surface of the hemisphere x2+y2+z2=a2with z≥0.
As in the previous example, dS=a2sinθd θd φ ˆrin spherical polar coordinates. Therefore
the vector area is given by
S=
ZZ
Sa2sinθˆrdθ dφ.
Now, since ˆrvaries over the surface S, it also must be integrated. This is most easily
achieved by writing ˆrin terms of the constant Cartesian basis vectors. On Swe have
ˆr=s i n θcosφi+s i n θsinφj+c o s θk,
so the expression for the vector area becomes
S=i
/
a2
Z2π
0cosφd φ
Zπ/2
0sin2θd θ
/!
+j
/
a2
Z2π
0sinφd φ
Zπ/2
0sin2θd θ
/!
+k
/
a2
Z2π
0dφ
Zπ/2
0sinθcosθd θ
/!
=0+0+πa2k=πa2k.
Note that the magnitude of Sis the projected area, of the hemisphere onto the xy-plane,
and not the surface area of the hemisphere.
J
399
LINE, SURFACE AND VOLUME INTEGRALS
dr
r
OC
Figure 11.8 The conical surface spanning the perimeter Cand having its
vertex at the origin.
The hemispherical shell discussed above is an example of an open surface. For
a closed surface, however, the vector area is always zero. This may be seen by
projecting the surface down onto each Cartesian coordinate plane in turn. For
each projection, every positive element of area on the upper surface is cancelledby the corresponding negative element on the lower surface. Therefore, eachcomponent of S=contintegraltext
SdSvanishes.
An important corollary of this result is that the vector area of an open surface
depends only on its perimeter, or boundary curve, C. This may be proved as
follows. If surfaces S1andS2have the same perimeter then S1−S2is a closed
surface, for which
contintegraldisplay
dS=integraldisplay
S1dS−integraldisplay
S2dS=0.
Hence S1=S2. Moreover, we may derive an expression for the vector area of
an open surface Ssolely in terms of a line integral around its perimeter C.
Since we may choose any surface with perimeter C, we will consider a cone
with its vertex at the origin (see figure 11.8). The vector area of the elementarytriangular region shown in the figure is dS=
1
2r×dr. Therefore, the vector area
of the cone, and hence of anyopen surface with perimeter C, is given by the line
integral
S=1
2contintegraldisplay
Cr×dr.
For a surface confined to the xy-plane, r=xi+yjand dr=dxi+dyj,
and we obtain for this special case that the area of the surface is given byA=
1
2contintegraltext
C(xd y−yd x), as we found in section 11.3.
400
11.5 SURFACE INTEGRALSIFind the vector area of the surface of the hemisphere x2+y2+z2=a2,z≥0,b y
evaluating the line integral S=1
2
H
Cr×draround its perimeter.
The perimeter Cof the hemisphere is the circle x2+y2=a2, on which we have
r=acosφi+asinφj,d r=−asinφd φi+acosφd φj.
Therefore the cross product r×dris given by
r×dr=
//
////
ij k
acosφa sinφ 0
−asinφd φ a cosφd φ 0
//
////
=a2(cos2φ+s i n2φ)dφk=a2dφk,
and the vector area becomes
S=1
2a2k
Z2π
0dφ=πa2k.
J
11.5.3 Physical examples of surface integrals
There are many examples of surface integrals in the physical sciences. Surface
integrals of the form (11.8) occur in computing the total electric charge on a
surface or the mass of a shell,integraltext
Sρ(r)dS, when the charge or mass density ρ(r)is
known. For surface integrals involving vectors, the second form in (11.9) is themost common. For a vector field a, the surface integralintegraltext
Sa·dSis called the flux
ofathrough S. Examples of physically important flux integrals are numerous.
For example, let us consider a surface Sin a fluid with density ρ(r)t h a th a sa
velocity field v(r). The mass of fluid crossing an element of surface area dSin
time dtisdM=ρv·dSdt. Therefore the nettotal mass flux of fluid crossing S
isM=integraltext
Sρ(r)v(r)·dS. As a another example, the electromagnetic flux of energy
out of a given volume Vbounded by a surface Siscontintegraltext
S(E×H)·dS.
The solid angle, to be defined below, subtended at a point Oby a surface (closed
or otherwise) can also be represented by an integral of this form, although it isnot strictly a flux integral (unless we imagine isotropic rays radiating from O).
The integral
Ω=integraldisplay
Sr·dS
r3=integraldisplay
Sˆr·dS
r2, (11.11)
gives the solid angle Ωsubtended at Oby a surface Sifris the position vector
measured from Oof an element of the surface. A little thought will show that
(11.11) takes account of all three relevant factors: the size of the element ofsurface, its inclination to the line joining the element to Oand the distance from
O. Such a general expression is often useful for computing solid angles when the
three-dimensional geometry is complicated. Note that (11.11) remains valid when
the surface Sis not convex and when a single ray from Oin certain directions
would cut Sin more than one place (but we exclude multiply connected regions).
401
LINE, SURFACE AND VOLUME INTEGRALS
In particular, when the surface is closed Ω = 0 if Ois outside Sand Ω = 4 πifO
is an interior point.
Surface integrals resulting in vectors occur less frequently. An example is
afforded, however, by the total resultant force experienced by a body immersed in
a stationary fluid in which the hydrostatic pressure is given by p(r). The pressure
is everywhere inwardly directed and the resultant force is F=−contintegraltext
SpdS,t a k e n
over the whole surface.
11.6 Volume integrals
Volume integrals are defined in an obvious way and are generally simpler than
line or surface integrals since the element of volume dVis a scalar quantity. We
may encounter volume integrals of the form
integraldisplay
Vφd V,integraldisplay
VadV. (11.12)
Clearly, the first form results in a scalar, whereas the second form yields a vector.
Two closely related physical examples, one of each kind, are provided by the totalmass of a fluid contained in a volume V,g i v e nb yintegraltext
Vρ(r)dV, and the total linear
momentum of that same fluid, given byintegraltext
Vρ(r)v(r)dV,w h e r e v(r)i st h ev e l o c i t y
field in the fluid. As a slightly more complicated example of a volume integral wemay consider the following.IFind an expression for the angular momentum of a solid body rotating with angular
velocity ωabout an axis through the origin.
Consider a small volume element dVsituated at position r; its linear momentum is ρd V˙r,
where ρ=ρ(r) is the density distribution, and its angular momentum about Oisr×ρ˙rdV.
Thus for the whole body the angular momentum Lis
L=
Z
V(r×˙r)ρd V.
Putting ˙r=ω×ryields
L=
Z
V[r×(ω×r)]ρd V=
Z
Vωr2ρd V−
Z
V(r·ω)rρd V.
J
The evaluation of the first type of volume integral in (11.12) has already been
considered in our discussion of multiple integrals in chapter 6. The evaluation of
the second type of volume integral follows directly since we can write
integraldisplay
VadV=iintegraldisplay
VaxdV+jintegraldisplay
VaydV+kintegraldisplay
VazdV, (11.13)
where ax,ay,azare the Cartesian components of a. Of course, we could have
written ain terms of the basis vectors of some other coordinate system (e.g.
spherical polars) but, since such basis vectors are not, in general, constant, they
402
11.6 VOLUME INTEGRALS
V
OS
rdS
Figure 11.9 A general volume Vcontaining the origin and bounded by the
closed surface S.
cannot be taken out of the integral sign as in (11.13) and must be included as
part of the integrand.
11.6.1 Volumes of three-dimensional regions
As discussed in chapter 6, the volume of a three-dimensional region Vis simply
V=integraltext
VdV, which may be evaluated directly once the limits of integration have
been found. However, the volume of the region obviously depends only on the
surface Sthat bounds it. We should therefore be able to express the volume V
in terms of a surface integral over S. This is indeed possible, and the appropriate
expression may derived as follows. Referring to figure 11.9, let us suppose thatthe origin Ois contained within V. The volume of the small shaded cone is
dV=
1
3r·dS; the total volume of the region is thus given by
V=1
3contintegraldisplay
Sr·dS.
It may be shown that this expression is still valid even when Ois not contained
inV. Although this surface integral form is available, in practice, in many cases
it is simpler to evaluate the volume integral directly.IFind the volume enclosed between a sphere of radius acentred on the origin and a circular
cone of half-angle αwith its vertex at the origin.
The element of vector area dSon the surface of the sphere is given in spherical polar
coordinates by a2sinθd θd φ ˆr. Now taking the axis of the cone to lie along the z-axis (from
which θis measured) the required volume is given by
V=1
3
I
Sr·dS=1
3
Z2π
0dφ
Zα
0a2sinθr·ˆrdθ
=1
3
Z2π
0dφ
Zα
0a3sinθd θ=2
3πa3(1−cosα).
J
403
LINE, SURFACE AND VOLUME INTEGRALS
11.7 Integral forms for grad,divandcurl
In the previous chapter we defined the vector operators grad, div and curl in purely
mathematical terms, which depended on the coordinate system in which they wereexpressed. An interesting application of line, surface and volume integrals is theexpression of grad, div and curl in coordinate-free, geometrical terms. If φis a
scalar field and ais a vector field then it may be shown that at any point P
∇φ= lim
V→0parenleftbigg1
Vcontintegraldisplay
SφdSparenrightbigg
(11.14)
∇·a= lim
V→0parenleftbigg1
Vcontintegraldisplay
Sa·dSparenrightbigg
(11.15)
∇×a= lim
V→0parenleftbigg1
Vcontintegraldisplay
SdS×aparenrightbigg
(11.16)
where Vis a small volume enclosing PandSis its bounding surface. Indeed,
we may consider these equations as the (geometrical) definitions of grad, div and
curl. An alternative, but equivalent, geometrical definition of ∇×aat a point P,
which is often easier to use than (11.16), is given by
(∇×a)·ˆn= lim
A→0parenleftbigg1
Acontintegraldisplay
Ca·drparenrightbigg
, (11.17)
where Cis a plane contour of area Aenclosing the point Pandˆnis the unit
normal to the enclosed planar area.
It may be shown, in any coordinate system , that all the above equations are
consistent with our definitions in the previous chapter although the difficulty of
proof depends on the chosen coordinate system. The most general coordinate
system encountered in that chapter was one with orthogonal curvilinear coordi-nates u
1,u2,u3, of which Cartesians, cylindrical polars and spherical polars are all
special cases. Although it may be shown that (11.14) leads to the usual expressionfor grad in curvilinear coordinates, the proof requires complicated manipulationsof the derivatives of the basis vectors with respect to the coordinates and is notpresented here. In Cartesian coordinates, however, the proof is quite simple.IShow that the geometrical definition of gradleads to the usual expression for ∇φin
Cartesian coordinates.
Consider the surface Sof a small rectangular volume element ∆ V=∆x∆y∆zthat has its
f a c e sp a r a l l e lt ot h e x,y,a n d zcoordinate surfaces and the point Pat one corner. We
must calculate the surface integral (11.14) over each of its six faces. Remembering that thenormal to the surface points outwards from the volume on each face, the two faces withx= constant have areas ∆ S=−i∆y∆zand ∆ S=i∆y∆zrespectively. Furthermore, over
each small surface element, we may take φto be constant, so that the net contribution to
404
11.7 INTEGRAL FORMS FOR grad, div AND curl
the surface integral fr om these two faces is then
[(φ+∆φ)−φ]∆y∆zi=
/
φ+∂φ
∂x∆x−φ
/
∆y∆zi
=∂φ
∂x∆x∆y∆zi.
The surface integral over the pairs of faces with y=c o n s t a n ta n d z= constant respectively
may be found in a similar way, and we obtainI
SφdS=
/∂φ
∂xi+∂φ
∂yj+∂φ
∂zk
/
∆x∆y∆z.
Therefore∇φat the point Pis given by
∇φ= lim
∆x,∆y,∆z→0
/1
∆x∆y∆z
/∂φ
∂xi+∂φ
∂yj+∂φ
∂zk
/
∆x∆y∆z
/
=∂φ
∂xi+∂φ
∂yj+∂φ
∂zk.
J
We now turn to (11.15) and (11.17). These geometrical definitions may be
shown straightforwardly to lead to the usual expressions for div and curl inorthogonal curvilinear coordinates.IBy considering the infinitesimal volume element dV=h1h2h3∆u1∆u2∆u3shown in fig-
ure 11.10, show that (11.15) leads to the usual expression for ∇·ain orthogonal curvilinear
coordinates.
Let us write the vector field in terms of its components with respect to the basis vectors
of the curvilinear coordinate system as a=a1ˆe1+a2ˆe2+a3ˆe3. We consider first the
contribution to the RHS of () from the two faces with u1=c o n s t a n t , i . e . PQ RS and the
face opposite it (see figure 11.10). Now, the volume element is formed from the orthogonalvectors h
1∆u1ˆe1,h2∆u2ˆe2andh3∆u3ˆe3, at the point Pand so for we have
∆S=h2h3∆u2∆u3ˆe3׈e2=−h2h3∆u2∆u3ˆe1.
Reasoning along the same lines as in the previous example, we conclude that the contri-
bution to the surface integral of a·dSover PQ RS and its opposite face taken together is
given by
∂
∂u1(a·∆S)∆u1=∂
∂u1(a1h2h3)∆u1∆u2∆u3.
The surface integrals over the pairs of faces with u2=c o n s t a n ta n d u3=c o n s t a n t
respectively may be found in a similar way, and we obtainI
Sa·dS=
/∂
∂u1(a1h2h3)+∂
∂u2(a2h3h1)+∂
∂u3(a3h1h2)
/
∆u1∆u2∆u3.
Therefore∇·aat the point Pis given by
∇·a= lim
∆u1,∆u2,∆u3→0
/1
h1h2h3∆u1∆u2∆u3
I
Sa·dS
/
=1
h1h2h3
/∂
∂u1(a1h2h3)+∂
∂u2(a2h3h1)+∂
∂u3(a3h1h2)
/
.
J
405
LINE, SURFACE AND VOLUME INTEGRALS
R
Q
PS
xyz
Th1∆u1ˆe1
h2∆u2ˆe2h3∆u3ˆe3
Figure 11.10 A general volume ∆ Vin orthogonal curvilinear coordinates
u1,u2,u3.PTgives the vector h1∆u1ˆe1,PSgives h2∆u2ˆe2and PQgives
h3∆u3ˆe3.IBy considering the infinitesimal planar surface element PQ RS in figure 11.10, show that
(11.17) leads to the usual expression for ∇×ain orthogonal curvilinear coordinates.
The planar surface PQ RS is defined by the orthogonal vectors h2∆u2ˆe2andh3∆u3ˆe3
at the point P. If we traverse the loop in the direction PSRQ then, by the right-hand
convention, the unit normal to the plane is ˆe1.W r i t i n g a=a1ˆe1+a2ˆe2+a3ˆe3, the line
integral around the loop in this direction is given byI
PSRQa·dr=a2h2∆u2+
/
a3h3+∂
∂u2(a3h3)∆u2
/
∆u3
−
/
a2h2+∂
∂u3(a2h2)∆u3
/
∆u2−a3h3∆u3
=
/∂
∂u2(a3h3)−∂
∂u3(a2h2)
/
∆u2∆u3.
Therefore from (11.17) the component of ∇×ain the direction ˆe1atPis given by
(∇×a)1= lim
∆u2,∆u3→0
/1
h2h3∆u2∆u3
I
PSRQa·dr
/
=1
h2h3
/∂
∂u2(h3a3)−∂
∂u3(h2a2)
/
.
The other two components are found by cyclically permuting the subscripts 1, 2, 3.
J
Finally, we note that we can also write the ∇2operator as a surface integral by
setting a=∇φin (11.15), to obtain
∇2φ=∇·∇φ= lim
V→0parenleftbigg1
Vcontintegraldisplay
S∇φ·dSparenrightbigg
.
406
11.8 DIVERGENCE THEOREM AND RELATED THEOREMS
11.8 Divergence theorem and related theorems
The divergence theorem relates the total flux of a vector field out of a closed
surface Sto the integral of the divergence of the vector field over the enclosed
volume V; it follows almost immediately from our geometrical definition of
divergence (11.15).
Imagine a volume V, in which a vector field ais continuous and differentiable,
to be divided up into a large number of small volumes Vi. Using (11.15), we have
for each small volume
(∇·a)Vi≈contintegraldisplay
Sia·dS,
where Siis the surface of the small volume Vi. Summing over iwe find that
contributions from surface elements interior to Scancel since each surface element
appears in two terms with opposite signs, the outward normals in the two termsbeing equal and opposite. Only contributions from surface elements that are alsoparts of Ssurvive. If each V
ii sa l l o w e dt ot e n dt oz e r ot h e nw eo b t a i nt h e
divergence theorem ,
integraldisplay
V∇·adV=contintegraldisplay
Sa·dS. (11.18)
We note that the divergence theorem holds for both simply and multiply con-
nected surfaces, provided that they are closed and enclose some non-zero volumeV. The divergence theorem may also be extended to tensor fields (see chapter 21).
The theorem finds most use as a tool in formal manipulations, but sometimes it
is of value in evaluating surface integrals of the formintegraltext
Sa·dSas volume integrals
or vice versa. For example, setting a=rwe immediately obtain
integraldisplay
V∇·rdV=integraldisplay
V3dV=3V=contintegraldisplay
Sr·dS,
which gives the expression for the volume of a region found in subsection 11.6.1.
The use of the divergence theorem is further illustrated in the following example.IEvaluate the surface integral I=
R
Sa·dS,w h e r e a=(y−x)i+x2zj+(z+x2)kandS
is the open surface of the hemisphere x2+y2+z2=a2,z≥0.
We could evaluate this surface integral directly, but the algebra is somewhat lengthy. We
will therefore evaluate it by use of the divergence theorem. Since the latter only holds
for closed surfaces enclosing a non-zero volume V, let us first consider the closed surface
S/prime=S+S1,w h e r e S1is the circular area in the xy-plane given by x2+y2≤a2,z=0 ; S/prime
then encloses a hemispherical volume V. By the divergence theorem we haveZ
V∇·adV=
I
S/primea·dS=
Z
Sa·dS+
Z
S1a·dS.
Now∇·a=−1+0+1=0 ,s ow ec a nw r i t eZ
Sa·dS=−
Z
S1a·dS.
407
LINE, SURFACE AND VOLUME INTEGRALS
Ry
C
xdxdydr
ˆnds
Figure 11.11 A closed curve Cin the xy-plane bounding a region R. Vectors
tangent and normal to the curve at a given point are also shown.
The surface integral over S1is easily evaluated. Remembering that the normal to the
surface points outward from the volume, a surface element on S1is simply dS=−kdx dy.
OnS1we also have a=(y−x)i+x2k,s ot h a t
I=−
Z
S1a·dS=
ZZ
Rx2dx dy,
where Ris the circular region in the xy-plane given by x2+y2≤a2. Transforming to plane
polar coordinates we have
I=
ZZ
R/primeρ2cos2φ ρ dρ dφ =
Z2π
0cos2φd φ
Za
0ρ3dρ=πa4
4.
J
It is also interesting to consider the two-dimensional version of the divergence
theorem. As an example, let us consider a two-dimensional planar region Rin
thexy-plane bounded by some closed curve C(see figure 11.11). At any point
on the curve the vector dr=dxi+dyjis a tangent to the curve and the vector
ˆnds=dyi−dxjis a normal pointing out of the region R. If the vector field ais
continuous and differentiable in Rthen the two-dimensional divergence theorem
in Cartesian coordinates gives
integraldisplayintegraldisplay
Rparenleftbigg∂ax
∂x+∂ay
∂yparenrightbigg
dx dy =contintegraldisplay
a·ˆnds=contintegraldisplay
C(axdy−aydx).
Letting P=−ayandQ=ax, we recover Green’s theorem in a plane, which was
discussed in section 11.3.
11.8.1 Green’s theorems
Consider two scalar functions φandψthat are continuous and differentiable in
some volume Vbounded by a surface S. Applying the divergence theorem to the
408
11.8 DIVERGENCE THEOREM AND RELATED THEOREMS
vector field φ∇ψwe obtain
contintegraldisplay
Sφ∇ψ·dS=integraldisplay
V∇·(φ∇ψ)dV
=integraldisplay
Vbracketleftbig
φ∇2ψ+(∇φ)·(∇ψ)bracketrightbig
dV. (11.19)
Reversing the roles of φandψin (11.19) and subtracting the two equations gives
contintegraldisplay
S(φ∇ψ−ψ∇φ)·dS=integraldisplay
V(φ∇2ψ−ψ∇2φ)dV. (11.20)
Equation (11.19) is usually known as Green’s first theorem and (11.20) as his
second. Green’s second theorem is useful in the development of the Green’sfunctions used in the solution of partial differential equations (see chapter 19).
11.8.2 Other related integral theorems
There exist two other integral theorems which are closely related to the divergence
theorem and which are of some use in physical applications. If φis a scalar field
andbis a vector field and both φandbsatisfy our usual differentiability
conditions in some volume Vbounded by a closed surface Sthen
integraldisplay
V∇φd V=contintegraldisplay
SφdS, (11.21)
integraldisplay
V∇×bdV=contintegraldisplay
SdS×b. (11.22)IUse the divergence theorem to prove (11.21).
In the divergence theorem (11.18) let a=φc,w h e r e cis a constant vector. We then haveZ
V∇·(φc)dV=
I
Sφc·dS.
Expanding out the integrand on the LHS we have
∇·(φc)=φ∇·c+c·∇φ=c·∇φ,
sincecis constant. Also, φc·dS=c·φdS,s ow eo b t a i nZ
Vc·(∇φ)dV=
I
Sc·φdS.
Since cis constant we may take it out of both integrals to give
c·
Z
V∇φd V=c·
I
SφdS,
and since cis arbitrary we obtain the stated result (11.21).
J
Equation (11.22) may be proved in a similar way by letting a=b×cin the
divergence theorem, where cis again a constant vector.
409
LINE, SURFACE AND VOLUME INTEGRALS
11.8.3 Physical applications of the divergence theorem
The divergence theorem is useful in deriving many of the most important partial
differential equations in physics (see chapter 18). The basic idea is to use the
divergence theorem to convert an integral form, often derived from observation,into an equivalent differential form (used in theoretical statements).IFor a compressible fluid with time-varying position-dependent density ρ(r,t)and velocity
field v(r,t), in which fluid is neither being created nor destroyed, show that
∂ρ
∂t+∇·(ρv)=0 .
For an arbitrary volume Vin the fluid, the conservation of mass tells us that the rate of
increase or decrease of the mass Mof fluid in the volume must equal the net rate at which
fluid is entering or leaving the volume, i.e.
dM
dt=−
I
Sρv·dS,
where Sis the surface bounding V. But the mass of fluid in Vis simply M=
R
Vρd V,s o
we have
d
dt
Z
Vρd V+
I
Sρv·dS=0.
Taking the derivative inside the first integral on the RHS and using the divergence theorem
to rewrite the second integral, we obtainZ
V∂ρ
∂tdV+
Z
V∇·(ρv)dV=
Z
V
/∂ρ
∂t+∇·(ρv)
/
dV=0.
Since the volume Vis arbitrary, the integrand (which is assumed continuous) must be
identically zero, so we obtain
∂ρ
∂t+∇·(ρv)=0 .
This is known as the continuity equation . It can also be applied to other systems, for
example those in which ρis the density of electric charge or the heat content, etc. In the
flow of an incompressible fluid ρ= constant and the continuity equation becomes simply
∇·v=0 .
J
In the previous example, we assumed that there were no sources or sinks in
the volume V, i.e. that there was no part of Vin which fluid was being created
or destroyed. We now consider the case where a finite number of point sources
and/or sinks are present in an incompressible fluid. Let us first consider thesimple case where a single source is located at the origin, out of which a quantityof fluid flows radially at a rate Q(m
3s−1). The velocity field is given by
v=Qr
4πr3=Qˆr
4πr2.
Now, for a sphere S1of radius rcentred on the source, the flux across S1is
contintegraldisplay
S1v·dS=|v|4πr2=Q.
410
11.8 DIVERGENCE THEOREM AND RELATED THEOREMS
Since vhas a singularity at the origin it is not differentiable there, i.e. ∇·vis not
defined there, but at all other points ∇·v= 0, as required for an incompressible
fluid. Therefore, from the divergence theorem, for any closed surface S2that does
not enclose the origin we have
contintegraldisplay
S2v·dS=integraldisplay
V∇·vdV=0.
Thus we see that the surface integralcontintegraltext
Sv·dShas value Qor zero depending on
whether or not Sencloses the source at the origin. In order that the divergence
theorem is valid for allsurfaces S, irrespective of whether they enclose the source,
we write
∇·v=Qδ(r),
where δ(r) is the three-dimensional Dirac delta function. The properties of this
function are discussed fully in chapter 13, but for the moment we note that it isdefined in such a way that
δ(r−a)=0 f o r r/negationslash=a,
integraldisplay
Vf(r)δ(r−a)dV=braceleftBigg
f(a)i falies in V
0o t h e r w i s e
for any well-behaved function f(r). Therefore, for any volume Vcontaining the
source at the origin, we have
integraldisplay
V∇·vdV=Qintegraldisplay
Vδ(r)dV=Q,
which is consistent withcontintegraltext
Sv·dS=Qfor a closed surface enclosing the source.
Hence, by introducing the Dirac delta function the divergence theorem can bemade valid even for non-differentiable point sources.
The generalisation to several sources and sinks is straightforward. For example,
if a source is located at r=aand a sink at r=bthen the velocity field is
v=(r−a)Q
4π|r−a|3−(r−b)Q
4π|r−b|3
and its divergence is given by
∇·v=Qδ(r−a)−Qδ(r−b).
Therefore, the integralcontintegraltext
Sv·dShas the value QifSencloses the source, −Qif
Sencloses the sink and 0 if Sencloses neither the source nor sink or encloses
them both. This analysis also applies to other physical systems – for example, in
electrostatics we can regard the sources and sinks as positive and negative pointcharges respectively and replace vby the electric field E.
411
LINE, SURFACE AND VOLUME INTEGRALS
11.9 Stokes’ theorem and related theorems
Stokes’ theorem is the ‘curl analogue’ of the divergence theorem and relates the
integral of the curl of a vector field over an open surface Sto the line integral of
the vector field around the perimeter Cbounding the surface.
Following the same lines as for the derivation of the divergence theorem, we
can divide the surface Sinto many small areas Siwith boundaries Ciand unit
normals ˆni. Using (11.17), we have for each small area
(∇×a)·ˆniSi≈contintegraldisplay
Cia·dr.
Summing over iwe find that on the RHS all parts of all interior boundaries
that are not part of Care included twice, being traversed in opposite directions
on each occasion and thus contributing nothing. Only contributions from line
elements that are also parts of Csurvive. If each Sii sa l l o w e dt ot e n dt oz e r o
then we obtain Stokes’ theorem,
integraldisplay
S(∇×a)·dS=contintegraldisplay
Ca·dr. (11.23)
We note that Stokes’ theorem holds for both simply and multiply connected open
surfaces, provided that they are two-sided. Stokes’ theorem may also be extendedto tensor fields (see chapter 21).
Just as the divergence theorem (11.18) can be used to relate volume and surface
integrals for certain types of integrand, Stokes’ theorem can be used in evaluatingsurface integrals of the formcontintegraltext
S(∇×a)·dSas line integrals or vice versa.IGiven the vector field a=yi−xj+zk, verify Stokes’ theorem for the hemispherical
surface x2+y2+z2=a2,z≥0.
Let us first evaluate the surface integralZ
S(∇×a)·dS
over the hemisphere. It is easily shown that ∇×a=−2k, and the surface element is
dS=a2sinθd θd φ ˆrin spherical polar coordinates. ThereforeZ
S(∇×a)·dS=
Z2π
0dφ
Zπ/2
0dθ
/;
−2a2sinθ
/ˆr·k
=−2a2
Z2π
0dφ
Zπ/2
0sinθ
/z
a
/
dθ
=−2a2
Z2π
0dφ
Zπ/2
0sinθcosθd θ=−2πa2.
We now evaluate the line integral around the perimeter curve Cof the surface, which
412
11.9 STOKES’ THEOREM AND RELATED THEOREMS
is the circle x2+y2=a2in the xy-plane. This is given byI
Ca·dr=
I
C(yi−xj+zk)·(dxi+dyj+dzk)
=
I
C(yd x−xd y).
Using plane polar coordinates, on Cwe have x=acosφ,y=asinφso that dx=
−asinφd φ,dy=acosφd φ, and the line integral becomesI
C(yd x−xd y)=−a2
Z2π
0(sin2φ+c o s2φ)dφ=−a2
Z2π
0dφ=−2πa2.
Since the surface and line integrals have the s ame value, we have verified Stokes’ theorem
in this case.
J
The two-dimensional version of Stokes’ theorem also yields Green’s theorem in
a plane. Consider the region Rin the xy-plane shown in figure 11.11, in which a
vector field ais defined. Since a=axi+ayj, we have∇×a=(∂ay/∂x−∂ax/∂y)k,
and Stokes’ theorem becomes
integraldisplayintegraldisplay
Rparenleftbigg∂ay
∂x−∂ax
∂yparenrightbigg
dx dy =contintegraldisplay
C(axdx+aydy).
Letting P=axandQ=aywe recover Green’s theorem in a plane, (11.4).
11.9.1 Related integral theorems
As for the divergence theorem, there exist two other integral theorems that are
closely related to Stokes’ theorem. If φis a scalar field and bis a vector field,
and both φandbsatisfy our usual differentiability conditions on some two-sided
open surface Sbounded by a closed perimeter curve C,t h e n
integraldisplay
SdS×∇φ=contintegraldisplay
Cφdr, (11.24)
integraldisplay
S(dS×∇)×b=contintegraldisplay
Cdr×b. (11.25)IUse Stokes’ theorem to prove (11.24).
In Stokes’ theorem, (11.23), let a=φc,w h e r e cis a constant vector. We then haveZ
S[∇×(φc)]·dS=
I
Cφc·dr. (11.26)
Expanding out the integrand on the LHS we have
∇×(φc)=∇φ×c+φ∇×c=∇φ×c,
sincecis constant, and the triple scalar product on the LHS of (11.26) can therefore be
written
[∇×(φc)]·dS=(∇φ×c)·dS=c·(dS×∇φ).
413
LINE, SURFACE AND VOLUME INTEGRALS
Substituting this into (11.26) and taking cout of both integrals because it is constant, we
find
c·
Z
SdS×∇φ=c·
I
Cφdr.
Since cis an arbitrary constant vector we therefore obtain the stated result (11.24).
J
Equation (11.25) may be proved in a similar way, by letting a=b×cin Stokes’
theorem, where cis again a constant vector. We also note that by setting b=r
in (11.25) we findintegraldisplay
S(dS×∇)×r=contintegraldisplay
Cdr×r.
Expanding out the integrand on the LHS we find
(dS×∇)×r=dS−dS(∇·r)=dS−3dS=−2dS.
Therefore, as we found in subsection 11.5.2, the vector area of an open surface S
is given by
S=integraldisplay
SdS=1
2contintegraldisplay
Cr×dr.
11.9.2 Physical applications of Stokes’ theorem
Like the divergence theorem, Stokes’ theorem is useful in converting integral
equations into differential equations.IFrom Amp `ere’s law derive Maxwell’s equation in the case where the currents are steady,
i.e.∇×B−µ0J=0.
Amp`ere’s rule for a distributed current with current density JisI
CB·dr=µ0
Z
SJ·dS,
for any circuit Cbounding a surface S. Using Stokes’ theorem, the LHS can be transformed
into
R
S(∇×B)·dS; henceZ
S(∇×B−µ0J)·dS=0
foranysurface S. This can only be so if ∇×B−µ0J=0, which is the required relation.
Similarly, from Faraday’s law of electromagnetic induction we can derive Maxwell’sequation∇×E=−∂B/∂t.J
In subsection 11.8.3 we discussed the flow of an incompressible fluid in the
presence of several sources and sinks. Let us now consider vortex flow in an
incompressible fluid with a velocity field
v=1
ρˆeφ,
in cylindrical polar coordinates ρ, φ, z . For this velocity field ∇×vequals zero
414
11.10 EXERCISES
everywhere except on the axis ρ=0 ,w h e r e vhas a singularity. Thereforecontintegraltext
Cv·dr
equals zero for any path Cthat does not enclose the vortex line on the axis and
2πifCdoes enclose the axis. In order for Stokes’ theorem to be valid for all
paths C, we therefore set
∇×v=2πδ(ρ),
where δ(ρ) is the Dirac delta function, to be discussed in subsection 13.1.3. Now,
since∇×v=0, except on the axis ρ= 0, there exists a scalar potential ψsuch
thatv=∇ψ. It may easily be shown that ψ=φ, the polar angle. Therefore, if C
does not enclose the axis thencontintegraldisplay
Cv·dr=contintegraldisplay
dφ=0,
and if Cdoes enclose the axis,
contintegraldisplay
Cv·dr=∆φ=2πn,
where nis the number of times we traverse C. Thus φis a multivalued potential.
A similar analysis is valid for other physical systems – for example, in magneto-
statics we may replace the vortex lines by current-carrying wires and the velocityfieldvby the magnetic field B.
11.10 Exercises
11.1 The vector field Fis defined by
F=2xzi+2yz2j+(x2+2y2z−1)k.
Calculate∇×Fand deduce that Fcan be written F=∇φ.Determine the form
ofφ.
11.2 The vector field Qis defined by
Q=
/
3x2(y+z)+y3+z3
/
i+
/
3y2(z+x)+z3+x3
/
j+
/
3z2(x+y)+x3+y3
/
k.
Show that Qis a conservative field, construct its potential function and hence
evaluate the integral J=
R
Q·dralong any line connecting the point Aat
(1,−1,1) to Bat (2,1,2).
11.3 Fis a vector field xy2i+2j+xk,a n d Lis a path parameterised by x=ct,y=c/t,
z=dfor the range 1 ≤t≤2. Evaluate (a)
R
LFdt,( b )
R
LFdyand (c)
R
LF·dr.
11.4 By making an appropriate choice for the functions P(x, y)a n d Q(x, y) that appear
in Green’s theorem in a plane, show that the integral of x−yover the upper half
of the unit circle centred on the origin has the value −2
3. Show the same result
by direct integration in Cartesian coordinates.
11.5 Determine the point of intersection P, in the first quadrant, of the two ellipses
x2
a2+y2
b2=1 a n dx2
b2+y2
a2=1.
Taking b<a, consider the contour Lthat bounds that area in the first quadrant
which is common to the two ellipses. Show that the parts of Lthat lie along the
coordinate axes contribute nothing to the line integral around Lofxd y−yd x,
and that this line integral can be written as the sum of two such integrals, I1
415
LINE, SURFACE AND VOLUME INTEGRALS
andI2, around closed contours. Using a parameterisation of each ellipse similar
to that employed in the example in section 11.3, evaluate these two integrals andhence find the total area common to the two ellipses.
11.6 By using parameterisations of the form x=acos
nθandy=asinnθfor suitable
values of n, find the area bounded by the curves
x2/5+y2/5=a2/5and x2/3+y2/3=a2/3
.
11.7 Evaluate the line integral
I=
I
C
/
y(4x2+y2)dx+x(2x2+3y2)dy
/
around the ellipse x2/a2+y2/b2=1 .
11.8 Criticise the following ‘proof’ that π=0 .
(a) Apply Green’s theorem in a plane to the functions P(x, y)=t a n−1(y/x)a n d
Q(x, y)=t a n−1(x/y), taking the region Rto be the unit circle centred on the
origin.
(b) The RHS of the equality so produced isZ Z
Ry−x
x2+y2dx dy
which, either by symmetry considerations or by changing to plane polar
coordinates, can be shown to have zero value.
(c) In the LHS of the equality set x=c o s θandy=s i n θ, yielding P(θ)=θ
andQ(θ)=π/2−θ. The line integral becomesZ2
0π
h/π
2−θ
/
cosθ−θsinθ
i
dθ,
which has value 2 π.
(d) Thus 2 π= 0 and the stated result follows.
11.9 A single-turn coil Cof arbitrary shape is placed in a magnetic field Band carries
a current I. Show that the couple acting upon the coil can be written as
M=I
Z
C(B·r)dr−I
Z
CB(r·dr).
For a planar rectangular coil of sides 2 aand 2 bplaced with its plane vertical
and at an angle φto a uniform horizontal field B, show that Mis, as expected,
4abBIcosφk.
11.10 Find the vector area Sof the curved surface of the hyperboloid of revolution
x2
a2−y2+z2
b2=1
which lies in the region z≥0a n d a≤x≤λa.
11.11 An axially symmetric solid body with its axis ABvertical is immersed in an
incompressible fluid of density ρ0. Use the following method to show that,
whatever the shape of the body, for ρ=ρ(z) in cylindrical polars the Archimedean
upthrust is, as expected, ρ0gV,w h e r e Vis the volume of the body.
Express the vertical component of the resultant force ( −
R
pdS,w h e r e pis the
pressure) on the body in terms of an integral; note that p=−ρ0gzand that for
an annular surface element of width dl,n·nzdl=−dr. Integrate by parts and
use the fact that ρ(zA)=ρ(zB)=0 .
416
11.10 EXERCISES
11.12 Show that the expression below is equal to the solid angle subtended by a
rectangular aperture of sides 2 aand 2 bat a point a distance cfrom the aperture
along the normal to its centre:
Ω=4
Zb
0ac
(y2+c2)(y2+c2+a2)1/2dy.
By setting y=(a2+c2)1/2tanφ, change this integral into the formZφ1
04accosφ
c2+a2sin2φdφ,
where tan φ1=b/(a2+c2)1/2, and hence show that
Ω=4t a n−1
/ab
c(a2+b2+c2)1/2
/
.
11.13 A vector field ais given by−zxr−3i−zyr−3j+(x2+y2)r−3k,w h e r e r2=x2+y2+z2.
Establish that the field is conservative (a) by showing ∇×a=0and (b) by
constructing its potential function φ.
11.14 A vector field ais given by ( z2+2xy)i+(x2+2yz)j+(y2+2zx)k. Show that
ais conservative and that the line integral
R
a·dralong any line joining (1 ,1,1)
and (1 ,2,2) has the value 11.
11.15 A force F(r) acts on a particle at r. In which of the following cases can Fbe
represented in terms of a potential? Where it can, find the potential.
(a)F=F0
/
i−j−2(x−y)
a2r
/
exp
/
−r2
a2
/
;
(b)F=F0
a
h
zk+(x2+y2−a2)
a2r
i
exp
/
−r2
a2
/
;
(c)F=F0
/
k+a(r×k)
r2
/
.
11.16 One of Maxwell’s electromagnetic equations states that all magnetic fields B
are solenoidal (i.e. ∇·B= 0). Determine whether each of the following vectors
could represent a real magnetic field; where it could, try to find a suitable vectorpotential A, i.e. such that B=∇×A. (Hint: seek a vector potential that is parallel
to∇×B.):
(a)B
0b
r3[(x−y)zi+(x−y)zj+x2−y2k] in Cartesians with r2=x2+y2+z2;
(b)B0b
r3[cosθcosφˆer−sinθcosφˆeθ+s i n2 θsinφˆeφ] in spherical polars;
(c)B0b2
/zr
(b2+z2)2ˆeρ+1
b2+z2ˆez
/
in cylindrical polars.
11.17 The vector field fhas components yi−xj+kandγis a curve given parametrically
by
r=(a−c+ccosθ)i+(b+csinθ)j+c2θk,0≤θ≤2π.
Describe the shape of the path γand show that the line integral
R
γf·drvanishes.
Does this result imply that fis a conservative field?
11.18 A vector field a=f(r)ris spherically symmetric and everywhere directed away
from the origin. Show that ais irrotational but that it is also solenoidal only if
f(r)i so ft h ef o r m Ar−3.
417
LINE, SURFACE AND VOLUME INTEGRALS
11.19 Evaluate the surface integral
R
r·dS,w h e r e ris the position vector, over that
part of the surface z=a2−x2−y2for which z≥0, by each of the following
methods:
(a) parameterize the surface as x=asinθcosφ,y=asinθsinφ,z=a2cos2θ,
and show that
r·dS=a4(2 sin3θcosθ+c o s3θsinθ)dθ dφ.
(b) apply the divergence theorem to the volume bounded by the surface and the
plane z=0 .
11.20 Obtain an expression for the value φPat a point Pof a scalar function φthat
satisfies∇2φ= 0 in terms of its value and normal derivative on a surface Sthat
encloses it, by proceeding as follows.
(a) In Green’s second theorem take ψat any particular point Qas 1/r,w h e r e r
is the distance of Qfrom P. Show that ∇2ψ= 0 except at r=0 .
(b) Apply the result to the doubly connected region bounded by Sand a small
sphere Σ of radius δcentred on P.
(c) Apply the divergence theorem to show that the surface integral over Σ
involving 1 /δvanishes, and prove that the term involving 1 /δ2has the value
4πφP.
(d) Conclude that
φP=−1
4π
Z
Sφ∂
∂n
/1
r
/
dS+1
4π
Z
S1
r∂φ
∂ndS.
This important result shows that the value at a point Pof a function φ
that satisfies ∇2φ= 0 everywhere within a closed surface Sthat encloses P
may be expressed entirely in terms of its value and normal derivative on S.
This matter is taken up more generally in connection with Green’s functionsin chapter 19 and in connection with functions of a complex variable insection 20.12.
11.21 Use result (11.21), together with an appropriately chosen scalar function φto
prove that the position vector ¯rof the centre of mass of an arbitrarily-shaped
body of volume Vand uniform density can be written
¯r=1
V
I
S1
2r2dS
.
11.22 A rigid body of volume Vand surface Srotates with angular velocity ω. Show
that
ω=−1
2V
I
Su×dS,
where u(x) is the velocity of the point xon the surface S.
11.23 Demonstrate the validity of the divergence theorem:
(a) by calculating the flux of the vector
F=αr
(r2+a2)3/2
through the spherical surface |r|=√
3a;
(b) by showing that
∇·F=3αa2
(r2+a2)5/2
418
11.10 EXERCISES
and evaluating the volume integral of ∇·Fover the interior of the sphere
|r|=√
3a.
(The substitution r=atanθwill prove useful in carrying out the integration.)
11.24 Prove equation (11.22) and, by taking b=zx2i+zy2j+(x2−y2)k, show that the
two integrals
I=
Z
x2dVand J=
Z
cos2θsin3θcos 2φd θd φ ,
both taken over the unit sphere, must have the same value. Evaluate both directly
to show that the common value is 4 π/15.
11.25 In a uniform, non-dielectric, conducting medium with unit relative permittivity,
charge density ρ, current density J, electric field Eand magnetic field B, Maxwell’s
electromagnetic equations take the form (with µ0/epsilon10=c−2)
(i)∇·B= 0, (ii) ∇·E=ρ//epsilon10,
(iii)∇×E+˙B=0,( i v )∇×B−(˙E/c2)=µ0J,
The density of stored energy in the medium is given by1
2(/epsilon10E2+µ−1
0B2). Show
that the rate of change of the total stored energy in a volume Vis equal to
−
Z
VJ·EdV−1
µ0
I
S(E×B)·dS,
where Sis the surface bounding V. (The first integral gives the ohmic heating
loss, whilst the second gives the electromagnetic energy flux out of the boundingsurface. The vector µ
−1
0(E×B) is known as the Poynting vector.)
11.26 A vector field Fis defined in cylindrical polar coordinates ρ, θ, z by
F=F0
/xcosλz
ai+ycosλz
aj+( s i n λz)k
/
≡ρ
a(cosλz)eρ+( s i n λz)k,
where i,jandkare the unit vectors along the Cartesian axes and eρis the unit
vector ( x/ρ)i+(y/ρ)j.
(a) Calculate, as a surface integral, the flux of Fthrough the closed surface
bounded by the cylinders ρ=aandρ=2aand the planes z=±aπ/2.
(b) Evaluate the same integral using the divergence theorem.
11.27 The vector field Fis given by
F=( 3x2yz+y3z+xe−x)i+( 3xy2z+x3z+yex)j+(x3y+y3x+xy2z2)k.
Calculate (a) directly and (b) by using Stokes’ theorem the value of the line
integral
R
LF·dr,w h e r e Lis the (three-dimensional) closed contour OABCDEO
defined by the successive vertices (0 ,0,0), (1 ,0,0), (1 ,0,1), (1 ,1,1), (1 ,1,0), (0 ,1,0),
(0,0,0).
11.28 A vector force field Fis defined in Cartesian coordinates by
F=F0
//y3
3a3+y
aexy/a2+1
/
i+
/xy2
a3+x+y
aexy/a2
/
j+z
aexy/a2k
/
.
Use Stokes’ theorem to calculateI
LF·dr,
where Lis the perimeter of the rectangle ABCD given by A=( 0,1,0),B=( 1,1,0),
C=( 1,3,0) and D=( 0,3,0).
419
LINE, SURFACE AND VOLUME INTEGRALS
11.11 Hints and answers
11.1 Show that ∇×F=0. The potential φF(r)=x2z+y2z2−z.
11.2 Show that one component of ∇×Qis zero and apply symmetry. The potential
φQ(r)=xy(x2+y2)+yz(y2+z2)+zx(z2+x2);J=φQ(B)−φQ(A) = 54.
11.3 (a) c3ln 2i+2j+( 3c/2)k;( b )(−3c4/8)i−cj−(c2ln 2)k;( c ) c4ln 2−c.
11.4 Take P=y2andQ=x2. Show that the line integral along the x-axis from
(−1,0) to (1 ,0) contributes nothing.
11.5 For P,x=y=ab/(a2+b2)1/2. Note that the integral along the straight line
joining Pto the origin is traversed in opposite directions in I1and I2.T h e
relevant limits are 0 ≤θ1≤tan−1(b/a)a n dt a n−1(a/b)≤θ2≤π/2. As required
by symmetry, I1=I2; the total common area is 4 abtan−1(b/a).
11.6 Use the result of the worked example in section 11.3 and the reduction formulae
derived in exercise 2.42. Bounded area = 33 πa2/128.
11.7 Show that, in the notation of section 11.3, ∂Q/∂x−∂P/∂y =2x2;I=πa3b/2.
11.8 The conditions for Green’s theorem are not met as PandQare not continuous
(or differentiable) at the origin.
11.9 M=I
R
Cr×(dr×B).
11.10 Since the vector area of a closed surface vanishes, S=−S1i+S2kwhere S1is
the area of the semicircular intersection with the plane x=λaandS2is the
area of the hyperbolic intersection with the plane z=0 ; S1=1
2πb2(λ2−1);
S2=ab[λ√(λ2−1)−cosh−1λ].
11.13 (b) φ=c+z/r.
11.14 The appropriate potential function is f(x, y, z)=z2x+x2y+y2z.
11.15 (a) Yes, F0(x−y)exp(−r2/a2); (b) yes,−F0[(x2+y2)/2a]exp(−r2/a2);
(c) no,∇×F/negationslash=0.
11.16 Only (c) has zero divergence. A possible vector potential is1
2B0b2ρ(b2+z2)−1ˆeφ;
to this could be added the gradient of anyscalar function.
11.17 A spiral of radius cwith its axis parallel to the z-direction and passing through
(a, b). The pitch of the spiral is 2 πc2. No, because (i) γis not a closed loop and
(ii) the line integral must be zero for every closed loop, not just for a particular
one. In fact ∇×f=−2k/negationslash=0shows that fis not conservative.
11.18∇×a=0;∇·a=3f(r)+rf/prime(r)=0i f f(r)=Ar−3.
11.19 (a) dS=( 2a3cosθsin2θcosφi+2a3cosθsin2θsinφj+a2cosθsinθk)dθ dφ.
(b)∇·r= 3; over the plane z=0 ,r·dS= 0; The necessarily common value is
3πa4/2.
11.20 (d) Remember that the outward normal to the region is the inward normal to Σ.11.21 Write ras∇(
1
2r2).
11.22 Use result (11.22) and the expression for ∇×(a×b) and note that ( ω·∇)x=ω.
11.23 T heansweris 3√
3πα/2i ne a c hc a s e .
11.24 Follow the method indicated in subsection 11.8.2, using an identity given in table
10.1. Use Cartesian coordinates for the LHS of equation (11.22) and sphericalpolars for the RHS. Employ (anti)symmetry and periodicity arguments to setseveral integrals to zero without explicit calculation.
11.25 Identify the expression for ∇·(E×B) and use the divergence theorem.
11.26 6 πF
0(a2+2a/λ)sin(λaπ/2).
11.27 (a) The successive contributions to the integral are 1 ,0,2+1
2e,−7
3,−1,−1
2.
(b)∇×F=2xyz2i−y2z2j+yexk. Show that the contour is equivalent to the
sum of two plane square contours in the planes z=0a n d x= 1, the latter being
traversed in the negative sense. Integral =1
6(3e−5).
11.28
Ra
0dx
R3a
ady F 0(y/a)2exy/a2=F0a(2e3−4).
420
12
Fourier series
We have already discussed, in chapter 4, how complicated functions may be
expressed as power series. However, this is not the only way in which a functionmay be represented as a series, and the subject of this chapter is the expressionof functions as a sum of sine and cosine terms. Such a representation is called aFourier series . Unlike Taylor series, a Fourier series can describe functions that are
not everywhere continuous and/or different iable. There are also other advantages
in using trigonometrical terms. They are easy to differentiate and integrate, their
moduli are easily taken and each term contains only one characteristic frequency.This last point is important because, as we shall see later, Fourier series are oftenused to represent the response of a system to a periodic input, and this responseoften depends directly on the frequency content of the input. Fourier series areused in a wide variety of such physical situations, including the vibrations of afinite string, the scattering of light by a diffraction grating and the transmission
of an input signal by an electronic circuit.
12.1 The Dirichlet conditions
We have already mentioned that Fourier series may be used to represent some
functions for which a Taylor series expansion is not possible. The particularconditions that a function f(x) must fulfil in order that it may be expanded as a
Fourier series are known as the Dirichlet conditions , and may be summarised by
the following four points:
(i) the function must be periodic;
(ii) it must be single-valued and continuous, except possibly at a finite number
of finite discontinuities;
(iii) it must have only a finite number of maxima and minima within one
period;
(iv) the integral over one period of |f(x)|must converge.
421
FOURIER SERIES
L Lf(x)
x
Figure 12.1 An example of a function that may be represented as a Fourier
series without modification.
If the above conditions are satisfied then the Fourier series converges to f(x)
at all points where f(x) is continuous. The convergence of the Fourier series
at points of discontinuity is discussed in section 12.4. The last three Dirichletconditions are almost always met in real applications, but not all functions areperiodic and hence do not fulfil the first condition. It may be possible, however,to represent a non-periodic function as a Fourier series by manipulation of the
function into a periodic form. This is discussed in section 12.5. An example of
a function that may, without modification, be represented as a Fourier series isshown in figure 12.1.
We have stated without proof that any function that satisfies the Dirichlet
conditions may be represented as a Fourier series. Let us now show why this is
a plausible statement. We require that any reasonable function (one that satisfiesthe Dirichlet conditions) can be expressed as a linear sum of sine and cosineterms. We first note that we cannot use just a sum of sine terms since sine, beingan odd function (i.e. a function for which f(−x)=−f(x)), cannot represent even
functions (i.e. functions for which f(−x)=f(x)). This is obvious when we try
to express a function f(x) that takes a non-zero value at x= 0. Clearly, since
sinnx= 0 for all values of n, we cannot represent f(x)a tx= 0 by a sine series.
Similarly odd functions cannot be represented by a cosine series since cosine isan even function. Nevertheless, it is possible to represent allodd functions by a
sine series and alleven functions by a cosine series. Now, since all functions may
b ew r i t t e na st h es u mo fa no d da n da ne v e np a r t ,
f(x)=
1
2[f(x)+f(−x)] +1
2[f(x)−f(−x)]
=feven(x)+fodd(x),
422
12.2 THE FOURIER COEFFICIENTS
we can write any function as the sum of a sine series and a cosine series.
All the terms of a Fourier series are mutually orthogonal, that is, the integrals,
over one period, of the product of any two terms have the following properties:
integraldisplayx0+L
x0sinparenleftbigg2πrx
Lparenrightbigg
cosparenleftbigg2πpx
Lparenrightbigg
dx=0 f o ra l l randp, (12.1)
integraldisplayx0+L
x0cosparenleftbigg2πrx
Lparenrightbigg
cosparenleftbigg2πpx
Lparenrightbigg
dx=
Lforr=p=0,
1
2Lforr=p>0,
0f o r r/negationslash=p,(12.2)
integraldisplayx0+L
x0sinparenleftbigg2πrx
Lparenrightbigg
sinparenleftbigg2πpx
Lparenrightbigg
dx=
0f o r r=p=0,
1
2Lforr=p>0,
0f o r r/negationslash=p,(12.3)
where randpare integers greater than or equal to zero; these formulae are easily
derived. A full discussion of why it is possible to expand a function as a sum ofmutually orthogonal functions is given in chapter 17.
The Fourier series expansion of the function f(x) is conventionally written
f(x)=a
0
2+∞summationdisplay
r=1bracketleftbigg
arcosparenleftbigg2πrx
Lparenrightbigg
+brsinparenleftbigg2πrx
Lparenrightbiggbracketrightbigg
, (12.4)
where a0,ar,brare constants called the Fourier coefficients . These coefficients are
analogous to those in a power series expansion and the determination of their
numerical values is the essential step in writing a function as a Fourier series.
This chapter continues with a discussion of how to find the Fourier coefficients
for particular functions. We then discuss simplifications to the general Fourier
series that may save considerable effort in calculations. This is followed by the
alternative representation of a function as a complex Fourier series, and weconclude with a discussion of Parseval’s theorem.
12.2 The Fourier coefficients
We have indicated that a series that satisfies the Dirichlet conditions may be
written in the form (12.4). We now consider how to find the Fourier coefficientsfor any particular function. For a periodic function f(x)o fp e r i o d Lwe will find
that the Fourier coefficients are given by
a
r=2
Lintegraldisplayx0+L
x0f(x)cosparenleftbigg2πrx
Lparenrightbigg
dx, (12.5)
br=2
Lintegraldisplayx0+L
x0f(x)s i nparenleftbigg2πrx
Lparenrightbigg
dx, (12.6)
where x0is arbitrary but is often taken as 0 or −L/2. The apparently arbitrary
factor1
2which appears in the a0term in (12.4) is included so that (12.5) may
423
FOURIER SERIES
apply for r= 0 as well as r>0. The relations (12.5) and (12.6) may be derived
as follows.
Suppose the Fourier series expansion of f(x) can be written as in (12.4),
f(x)=a0
2+∞summationdisplay
r=1bracketleftbigg
arcosparenleftbigg2πrx
Lparenrightbigg
+brsinparenleftbigg2πrx
Lparenrightbiggbracketrightbigg
.
Then, multiplying by cos(2 πpx/L ), integrating over one full period in xand
changing the order of the summation and integration, we get
integraldisplayx0+L
x0f(x)cosparenleftbigg2πpx
Lparenrightbigg
dx=a0
2integraldisplayx0+L
x0cosparenleftbigg2πpx
Lparenrightbigg
dx
+∞summationdisplay
r=1arintegraldisplayx0+L
x0cosparenleftbigg2πrx
Lparenrightbigg
cosparenleftbigg2πpx
Lparenrightbigg
dx
+∞summationdisplay
r=1brintegraldisplayx0+L
x0sinparenleftbigg2πrx
Lparenrightbigg
cosparenleftbigg2πpx
Lparenrightbigg
dx.
(12.7)
We can now find the Fourier coefficients by considering (12.7) as ptakes different
values. Using the orthogonality conditions (12.1)–(12.3) of the previous section,we find that when p= 0 (12.7) becomes
integraldisplay
x0+L
x0f(x)dx=a0
2L.
When p/negationslash= 0 the only non-vanishing term on the RHS of (12.7) occurs when
r=p,a n ds o
integraldisplayx0+L
x0f(x)cosparenleftbigg2πrx
Lparenrightbigg
dx=ar
2L.
The other Fourier coefficients brmay be found by repeating the above process
but multiplying by sin(2 πpx/L ) instead of cos(2 πpx/L ) (see exercise 12.2).IExpress the square-wave function illustrated in figure 12.2 as a Fourier series.
Physically this might represent the input to a el ectrical circuit that sw itches between a high
and a low state with time period T. The square wave may be represented by
f(t)=
/(
−1f o r−1
2T≤t<0,
+1 for 0≤t<1
2T.
In deriving the Fourier coefficients, we note firstly that the function is an odd function
and so the series will contain only sine terms (this simplification is discussed further in the
424
12.3 SYMMETRY CONSIDERATIONS
1
−10T
2−T
2 tf(t)
Figure 12.2 A square-wave function.
following section). To evaluate the coefficients in the sine series we use (12.6). Hence
br=2
T
ZT/2
−T/2f(t)sin
/2πrt
T
/
dt
=4
T
ZT/2
0sin
/2πrt
T
/
dt
=2
πr[1−(−1)r].
Thus the sine coefficients are zero if ris even and equal to 4 /(πr)i fris odd. Hence the
Fourier series for the square- wave function may be written as
f(t)=4
π
/
sinωt+sin3ωt
3+sin5ωt
5+···
/
, (12.8)
where ω=2π/Tis called the angular frequency .
J
12.3 Symmetry considerations
The example in the previous section employed the useful property that since the
function to be represented was odd, all the cosine terms of the Fourier series were
zero. It is often the case that the function we wish to express as a Fourier serieshas a particular symmetry, which we can exploit to reduce the calculational labourof evaluating Fourier coefficients. Functions that are symmetric or antisymmetricabout the origin (i.e. even and odd functions respectively) admit particularlyuseful simplifications. Functions that are odd in xhave no cosine terms (see
section 12.1) and all the a-coefficients are equal to zero. Similarly, functions that
are even in xhave no sine terms and all the b-coefficients are zero. Since the
Fourier series of odd or even functions contain only half the coefficients requiredfor a general periodic function, there is a considerable reduction in the algebraneeded to find a Fourier series.
The consequences of symmetry or antisymmetry of the function about the
quarter period (i.e. about L/4) are a little less obvious. Furthermore, the results
425
FOURIER SERIES
are not used as often as those above and the remainder of this section can be
omitted on a first reading without loss of continuity. The following argumentgives the required results.
Suppose that f(x) has even or odd symmetry about L/4, i.e. f(L/4−x)=
±f(x−L/4). For convenience, we make the substitution s=x−L/4 and hence
f(−s)=±f(s). We can now see that
b
r=2
Lintegraldisplayx0+L
x0f(s)s i nparenleftbigg2πrs
L+πr
2parenrightbigg
ds,
where the limits of integration have been left unaltered since fis, of course,
periodic in sas well as in x. If we use the expansion
sinparenleftbigg2πrs
L+πr
2parenrightbigg
=s i nparenleftbigg2πrs
Lparenrightbigg
cosparenleftBigπr
2parenrightBig
+c o sparenleftbigg2πrs
Lparenrightbigg
sinparenleftBigπr
2parenrightBig
,
we can immediately see that the trigonometrical part of the integrand is an odd
function of sifris even and an even function of sifris odd. Hence if f(s)i s
even and ris even then the integral is zero, and if f(s) is odd and ris odd then
the integral is zero. Similar results can be derived for the Fourier a-coefficients
and we conclude that
(i) if f(x) is even about L/4t h e n a2r+1=0a n d b2r=0 ,
(ii) if f(x) is odd about L/4t h e n a2r=0a n d b2r+1=0 .
All the above results follow automatically when the Fourier coefficients are
evaluated in any particular case, but prior knowledge of them will often enable
some coefficients to be set equal to zero on inspection and so substantially reducethe computational labour. As an example, the square-wave function shown infigure 12.2 is (i) an odd function of t,s ot h a ta l l a
r= 0, and (ii) even about the
point t=T/4, so that b2r= 0. Thus we can say immediately that only sine terms
of odd harmonics will be present and therefore will need to be calculated; this isconfirmed in the expansion (12.8).
12.4 Discontinuous functions
The Fourier series expansion usually works well for functions that are discon-
tinuous in the required range. However, the series itself does not produce adiscontinuous function and we state without proof that the value of the ex-
panded f(x) at a discontinuity will be half-way between the upper and lower
values. Expressing this more mathematically, at a point of finite discontinuity, x
d,
the Fourier series converges to
1
2lim
/epsilon1→0[f(xd+/epsilon1)+f(xd−/epsilon1)].
At a discontinuity, the Fourier series representation of the function will overshoot
its value. Although as more terms are included the overshoot moves in position
426
12.4 DISCONTINUOUS FUNCTIONS
(a)( b)
(c)( d)−1
−1−1
−11
11
1−T
2
−T
2−T
2
−T
2T
2
T
2T
2
T
2δ
Figure 12.3 The convergence of a Fourier series expansion of a square-wave
function, including ( a)o n et e r m ,( b) two terms, ( c) three terms and ( d)2 0
terms. The overshoot δis shown in ( d).
arbitrarily close to the discontinuity, it never disappears even in the limit of an
infinite number of terms. This behaviour is known as Gibbs’ phenomenon .Af u l l
discussion is not pursued here but suffice it to say that the size of the overshoot
is proportional to the magnitude of the discontinuity.IFind the value to which the Fourier series of the square-wave function discussed in sec-
tion 12.2 converges at t=0.
It can be seen that the function is discontinuous at t= 0 and, by the above rule, we expect
the series to converge to a value half-way between the upper and lower values, in otherwords to converge to zero in this case. Considering the Fourier series of this function,(12.8), we see that all the terms are zero and hence the Fourier series converges to zero asexpected. The Gibbs phenomenon for the square-wave function is shown in figure 12.3.J
427
FOURIER SERIES
(a)
(b)
(c)
(d)
0000
LLLL
2L2L2L
Figure 12.4 Possible periodic extensions of a function.12.5 Non-periodic functions
We have already mentioned that a Fourier representation may sometimes be used
for non-periodic functions. If we wish to find the Fourier series of a non-periodicfunction only within a fixed range then we may continue this function outside the
range so as to make it periodic. The Fourier series of this periodic function wouldthen correctly represent the non-periodic function in the desired range. Since we
are often at liberty to extend the function in a number of ways, we can sometimes
make it odd or even and so reduce the calculation required. Figure 12.4( b)s h o w s
the simplest extension to the function shown in figure 12.4( a). However, this
extension has no particular symmetry. Figures 12.4( c), (d) show extensions as odd
and even functions respectively with the benefit that only sine or cosine termsappear in the resulting Fourier series. We note that these last two extensions givea function of period 2 L.
In view of the result of section 12.4, it must be added that the continuation
must not be discontinuous at the end-points of the interval of interest; if it isthe series will not converge to the required value there. This requirement thatthe series converges appropriately may reduce the choice of continuations. Thisis discussed further at the end of the following example.IFind the Fourier series of f(x)=x2for0<x≤2.
We must first make the function periodic. We do this by extending the range of interest to
−2<x≤2i ns u c haw a yt h a t f(x)=f(−x) and then letting f(x+4k)=f(x), where kis
any integer. This is shown in figure 12.5. Now we have an even function of period 4. TheFourier series will faithfully represent f(x) in the range, −2<x≤2, although not outside
it. Firstly we note that since we have made the specified function even in xby extending
428
12.5 NON-PERIODIC FUNCTIONS
−22 0
Lxf(x)=x2
Figure 12.5 f(x)=x2,0<x≤2, with the range extended to give periodicity.
the range, all the coefficients brwill be zero. Now we apply (12.5) and (12.6) with L=4
to determine the remaining coefficients:
ar=2
4
Z2
−2x2cos
/2πrx
4
/
dx=4
4
Z2
0x2cos
/πrx
2
/
dx,
where the second equality holds because the function is even in x. Thus
ar=
/2
πrx2sin
/πrx
2
/
/2
0−4
πr
Z2
0xsin
/πrx
2
/
dx
=8
π2r2
h
xcos
/πrx
2
/i2
0−8
π2r2
Z2
0cos
/πrx
2
/
dx
=16
π2r2cosπr
=16
π2r2(−1)r.
Since this expression for arhasr2in its denominator, to evaluate a0we must return to the
original definition,
ar=2
4
Z2
−2f(x)cos
/πrx
2
/
dx.
From this we obtain
a0=2
4
Z2
−2x2dx=4
4
Z2
0x2dx=8
3.
The final expression for f(x)i st h e n
x2=4
3+1 6∞X
r=1(−1)r
π2r2cos
/πrx
2
/
for 0 <x≤2.
J
We note that in the above example we could have extended the range so as
to make the function odd. In other words we could have set f(x)=−f(−x)a n d
then made f(x) periodic in such a way that f(x+4 )= f(x). In this case the
resulting Fourier series would be a series of just sine terms. However, althoughthis will faithfully represent the function inside the required range, it does not
429
FOURIER SERIES
converge to the correct values of f(x)=±4a tx=±2; it converges, instead, to
zero, the average of the values at the two ends of the range.
12.6 Integration and differentiation
It is sometimes possible to find the Fourier series of a function by integration or
differentiation of another Fourier series. If the Fourier series of f(x)i si n t e g r a t e d
term by term then the resulting Fourier series converges to the integral of f(x).
Clearly, when integrating in such a way there is a constant of integration that must
be found. If f(x) is a continuous function of xfor all xandf(x) is also periodic
then the Fourier series that results from differentiating term by term converges tof
/prime(x), provided that f/prime(x) itself satisfies the Dirichlet conditions. These properties
of Fourier series may be useful in calculating complicated Fourier series, sincesimple Fourier series may easily be evaluated (or found from standard tables)and often the more complicated series can then be built up by integration and/ordifferentiation.IFind the Fourier series of f(x)=x3for0<x≤2.
In the example discussed in the previous section we found the Fourier series for f(x)=x2
in the required range. So, if we integrate this term by term, we obtain
x3
3=4
3x+3 2∞X
r=1(−1)r
π3r3sin
/πrx
2
/
+c,
where cis, so far, an arbitrary constant. We have not yet found the Fourier series for x3
because the term4
3xappears in the expansion. However, by now differentiating the same
initial expression for x2we obtain
2x=−8∞X
r=1(−1)r
πrsin
/πrx
2
/
.
We can now write the full Fourier expansion of x3as
x3=−16∞X
r=1(−1)r
πrsin
/πrx
2
/
+9 6∞X
r=1(−1)r
π3r3sin
/πrx
2
/
+c.
Finally, we can find the constant, c, by considering f(0). At x= 0, our Fourier expansion
gives x3=csince all the sine terms are zero, and hence c=0 .
J
12.7 Complex Fourier series
As a Fourier series expansion in general contains both sine and cosine parts, it
may be written more compactly using a complex exponential expansion. Thissimplification makes use of the property that exp( irx)=c o s rx+isinrx.T h e
430
12.7 COMPLEX FOURIER SERIES
complex Fourier series expansion is written
f(x)=∞summationdisplay
r=−∞crexpparenleftbigg2πirx
Lparenrightbigg
, (12.9)
where the Fourier coefficients are given by
cr=1
Lintegraldisplayx0+L
x0f(x)ex pparenleftbigg
−2πirx
Lparenrightbigg
dx. (12.10)
This relation can be derived, in a similar manner to that of section 12.2, by mul-
tiplying (12.9) by exp( −2πipx/L ) before integrating and using the orthogonality
relation
integraldisplayx0+L
x0expparenleftbigg
−2πipx
Lparenrightbigg
expparenleftbigg2πirx
Lparenrightbigg
dx=braceleftBigg
Lforr=p,
0f o r r/negationslash=p.
The complex Fourier coefficients in (12.9) have the following relations to the real
Fourier coefficients:
cr=1
2(ar−ibr),
c−r=1
2(ar+ibr).(12.11)
Note that if f(x)i sr e a lt h e n c−r=c∗
r, where the asterisk represents complex
conjugation.IFind a complex Fourier series for f(x)=xin the range −2<x< 2.
Using (12.10), for r/negationslash=0 ,
cr=1
4
Z2
−2xexp
/
−πirx
2
/
dx
=
/
−x
2πirexp
/
−πirx
2
//2
−2+
Z2
−21
2πirexp
/
−πirx
2
/
dx
=−1
πir[exp(−πir)+e x p ( πir)]+
/1
r2π2exp
/
−πirx
2
//2
−2
=2i
πrcosπr−2i
r2π2sinπr=2i
πr(−1)r. (12.12)
Forr= 0, we find c0= 0 and hence
x=∞X
r=−∞
r/negationslash=02i(−1)r
rπexp
/πirx
2
/
.
We note that the Fourier series derived for xin section 12.6 gives ar=0f o ra l l rand
br=−4(−1)r
πr,
and so, using (12.11), we confirm that crandc−rhave the forms derived above. It is also
apparent that the relationship c∗
r=c−rholds, as we expect, since f(x)i sr e a l .
J
431
FOURIER SERIES
12.8 Parseval’s theorem
Parseval’s theorem gives a useful way of relating the Fourier coefficients to the
function that they describe. Essentially a conservation law, it states that
1
Lintegraldisplayx0+L
x0|f(x)|2dx=∞summationdisplay
r=−∞|cr|2
=parenleftbig1
2a0parenrightbig2+1
2∞summationdisplay
r=1(a2
r+b2
r). (12.13)
In a more memorable form, this says that the sum of the moduli squared of
the complex Fourier coefficients is equal to the average value of |f(x)|2over one
period. Parseval’s theorem can be proved straightforwardly by writing f(x)a s
a Fourier series and evaluating the required integral, but the algebra is messy.
Therefore, we shall use an alternative method, for which the algebra is simple
and which in fact leads to a more general form of the theorem.
Let us consider two functions f(x)a n d g(x), which are (or can be made)
periodic with period Land which have Fourier series (expressed in complex
form)
f(x)=∞summationdisplay
r=−∞crexpparenleftbigg2πirx
Lparenrightbigg
,
g(x)=∞summationdisplay
r=−∞γrexpparenleftbigg2πirx
Lparenrightbigg
,
where crandγrare the complex Fourier coefficients of f(x)a n d g(x) respectively.
Thus
f(x)g∗(x)=∞summationdisplay
r=−∞crg∗(x)ex pparenleftbigg2πirx
Lparenrightbigg
.
Integrating this equation with respect to xover the interval ( x0,x0+L)a n d
dividing by L, we find
1
Lintegraldisplayx0+L
x0f(x)g∗(x)dx=∞summationdisplay
r=−∞cr1
Lintegraldisplayx0+L
x0g∗(x)ex pparenleftbigg2πirx
Lparenrightbigg
dx
=∞summationdisplay
r=−∞crbracketleftbigg1
Lintegraldisplayx0+L
x0g(x)ex pparenleftbigg−2πirx
Lparenrightbigg
dxbracketrightbigg∗
=∞summationdisplay
r=−∞crγ∗
r,
where the last equality uses (12.10). Finally, if we let g(x)=f(x) then we obtain
Parseval’s theorem (12.13). This proof can be performed in a similar manner
432
12.9 EXERCISES
using the sine and cosine form of the Fourier series, but the algebra is slightly
more complicated.
Parseval’s theorem is sometimes used to sum series. However, if one is presented
with a series to sum, it is not usually possible to decide which Fourier series
should be used to evaluate it. Instead, useful summations are sometimes found
serendipitously. The following example shows the evaluation of a sum by aFourier series method.IUsing Parseval’s theorem and the Fourier series for f(x)=x2found in section 12.5,
calculate the sum
P∞
r=1r−4.
Firstly we find the average value of [ f(x)]2over the interval −2<x≤2:
1
4
Z2
−2x4dx=16
5.
Now we evaluate the right-hand side of (12.13):/;1
2a0
/2+1
2∞X
1a2
r+1
2∞X
1b2
n=
/;4
3
/2+1
2∞X
r=1162
π4r4.
Equating the two expression we find
∞X
r=11
r4=π4
90.
J
12.9 Exercises
12.1 Prove the orthogonality relations stated in section 12.1.
12.2 Derive the Fourier coefficients brin a similar manner to the derivation of the ar
in section 12.2.
12.3 Which of the following functions of xcould be represented by a Fourier series
over the range indicated?
(a) tanh−1(x),−∞<x<∞.
(b) tan x, −∞<x<∞.
(c)|sinx|−1/2,−∞<x<∞.
(d) cos−1(sin2 x),−∞<x<∞.
(e)xsin(1/x)−π−1<x≤π−1, cyclically repeated.
12.4 By moving the origin of tto the centre of an interval in which f(t) = +1, i.e.
by changing to a new independent variable t/prime=t−1
4T, express the square-wave
function in the example in section 12.2 as a cosine series. Calculate the Fouriercoefficients involved (a) directly and (b) by changing the variable in result (12.8).
12.5 Find the Fourier series of the function f(x)=xin the range −π<x≤π. Hence
show that
1−1
3+1
5−1
7+···=π
4.
12.6 For the function
f(x)=1−x, 0≤x≤1,
find (a) the Fourier sine series and (b) the Fourier cosine series. Which would
be better for numerical evaluation? Relate your answer to the relevant periodiccontinuations.
433
FOURIER SERIES
12.7 For the continued functions used in exercise 12.6 and the derived corresponding
series, consider (i) their derivatives and (ii) their integrals. Do they give meaningfulequations? You will probably find it helpful to sketch all the functions involved.
12.8 The function y(x)=xsinxfor 0≤x≤πis to be represented by a Fourier series
of period 2 πthat is either even or odd. By sketching the function and considering
its derivative, determine which series will have the more rapid convergence. Findthe full expression for the better of these two series, showing that the convergence∼n
−3and that alternate terms are missing.
12.9 Find the Fourier coefficients in the expansion of f(x)=e x p xover the range
−1<x< 1. What value will the expansion have when x=2 ?
12.10 By integrating term by term the Fourier series found in the previous question
and using the Fourier series for f(x)= xfound in section 12.6, show thatR
expxd x=e x p x+c. Why is it not possible to show that d(expx)/dx=e x p x
by differentiating the Fourier series of f(x)=e x p xin a similar manner?
12.11 Consider the function f(x)=e x p (−x2) in the range 0 ≤x≤1. Show how it
should be continued to give as its Fourier series a series (the actual form is notwanted) (a) with only cosine terms, (b) with only sine terms, (c) with period 1and (d) with period 2.
Would there be any difference between the values of the last two series at (i)
x= 0, (ii) x=1 ?
12.12 Find, without calculation, which terms will be present in the Fourier series for
the periodic functions f(t), of period T, that are given in the range −T/2t oT/2
by:
(a)f(t)=2f o r0≤|t|<T/4,f=1f o r T/4≤|t|<T/2;
(b)f(t)=e x p [−(t−T/4)
2];
(c)f(t)=−1f o r−T/2≤t<−3T/8a n d3 T/8≤t<T / 2,f(t)=1f o r
−T/8≤t<−T/8; the graph of fis completed by two straight lines in the
remaining ranges so as to form a continuous function.
12.13 Consider the representation as a Fourier series of the displacement of a string
lying in the interval 0 ≤x≤Land fixed at its ends, when it is pulled aside by y0
at the point x=L/4. Sketch the continuations for the region outside the interval
that will
(a) produce a series of period L,
(b) produce a series that is antisymmetric about x=0 ,a n d
(c) produce a series that will contain only cosine terms.
(d) What are (i) the periods of the series in (b) and (c) and (ii) the value of the
‘a0-term’ in (c)?
(e) Show that a typical term of the series obtained in (b) is
32y0
3n2π2sinnπ
4sinnπx
L.
12.14 Show that the Fourier series for the function y(x)=|x|in the range −π≤x<π
is
y(x)=π
2−4
π∞X
m=0cos(2 m+1 )x
(2m+1 )2.
By integrating this equation term by term from 0 to x, find the function g(x)
whose Fourier series is
4
π∞X
m=0sin(2m+1 )x
(2m+1 )3.
434
12.9 EXERCISES
Deduce the value of the sum Sof the series
1−1
33+1
53−1
73+···.
12.15 Using the result of exercise 12.14, determine, as far as possible by inspection, the
form of the functions of which the following are the Fourier series:
(a)
cosθ+1
9cos 3θ+1
25cos 5θ+···;
(b)
sinθ+1
27sin 3θ+1
125sin 5θ+···;
(c)
L2
3−4L2
π2
/
cosπx
L−1
4cos2πx
L+1
9cos3πx
L−···
/
.
(You may find it helpful to first set x= 0 in the quoted result and so obtain
values for S0=
P(2m+1 )−2and other sums derivable from it.)
12.16 By finding a cosine Fourier series of period 2 for the function f(t)t h a tt a k e st h e
form f(t)=c o s h ( t−1) in the range 0 ≤t≤1, prove that
∞X
n=11
n2π2+1=1
e2−1.
Deduce values for the sums
P(n2π2+1 )−1over odd nand even nseparately.
12.17 Find the (real) Fourier series of period 2 for f(x)=c o s h xandg(x)=x2in the
range−1≤x≤1. By integrating the series for f(x) twice, prove that
∞X
n=1(−1)n+1
n2π2(n2π2+1 )=1
2
/1
sinh 1−5
6
/
.
12.18 Express the function f(x)=x2as a Fourier sine series in the range 0 <x≤2
and show that it converges to zero at x=±2.
12.19 Demonstrate explicitly for the square-w ave function discussed in section 12.2 that
Parseval’s theorem (12.13) is valid. You will need to use the relationship
∞X
m=01
(2m+1 )2=π2
8.
Show that a filter that transmits frequencies only up to 8 π/Twill still transmit
more than 90 per cent of the power in such a square-wave voltage signal.
12.20 Show that the Fourier series for |sinθ|in the range −π≤θ≤πis given by
|sinθ|=2
π−4
π∞X
m=1cos 2mθ
4m2−1.
By setting θ=0a n d θ=π/2, deduce values for
∞X
m=11
4m2−1and∞X
m=11
16m2−1.
435
FOURIER SERIES
12.21 Find the complex Fourier series for the periodic function of period 2 πdefined in
the range−π≤x≤πbyy(x)=c o s h x. By setting t= 0 prove that
∞X
n=1(−1)n
n2+1=1
2
/π
sinhπ−1
/
.
12.22 The repeating output from an electronic oscillator takes the form of a sine wave
f(t)=s i n tfor 0≤t≤π/2; it then drops instantaneously to zero and starts
again. The output is to be represented by a complex Fourier series of the form
∞X
n=−∞cne4nti.
Sketch the function and find an expression for cn.V e r i f yt h a t c−n=c∗
n.D e m o n -
strate that setting t=0a n d t=π/2 produces differing values for the sum
∞X
n=11
16n2−1.
Determine the correct value and check it using the quoted result of exercise 12.5.
12.23 Apply Parseval’s theorem to the series found in the previous exercise and so
derive a value for the sum of the series
17
(15)2+65
(63)2+145
(143)2+···+16n2+1
(16n2−1)2+···.
12.24 A string, anchored at x=±L/2, has a fundamental vibration frequency of 2 L/c,
where cis the speed of transverse waves on the string. It is pulled aside at its
centre point by a distance y0and released at time t= 0. Its subsequent motion
can be described by the series
y(x, t)=∞X
n=1ancosnπx
Lcosnπct
L.
Find a general expression for anand show that only odd harmonics of the
fundamental frequency are present in the sound generated by the released string.By applying Parseval’s theorem, find the sum Sof the seriesP∞
0(2m+1 )−4.
12.25 Show that Parseval’s theorem for two functions whose Fourier expansions have
cosine and sine coefficients an,bnandαn,βntakes the form
1
L
ZL
0f(x)g∗(x)dx=1
4a0α0+1
2∞X
n=1(anαn+bnβn).
(a) Demonstrate that for g(x)=s i n mxor cos mxthis reduces to the definition
of the Fourier coefficients.
(b) Explicitly verify the above result for the case in which f(x)=xandg(x)i s
the square-wave function, both in the interval −1≤x≤1.
12.26 An odd function f(x)o fp e r i o d2 πis to be approximated by a Fourier sine series
having only mterms. The error in this approximation is measured by the square
deviation
Em=
Zπ
−π
/"
f(x)−mX
n=1bnsinnx
/#2
dx.
By differentiating Emwith respect to the coefficients bn, find the values of bnthat
minimise Em.
436
12.10 HINTS AND ANSWERS
Sketch the graph of the function f(x), where
f(x)=
/
−x(π+x)f o r−π≤x<0,
x(x−π)f o r 0≤x<π .
f(x) is to be approximated by the first three terms of a Fourier sine series. What
coefficients minimise E3? What is the resulting value of E3?
12.10 Hints and answers
12.3 Only (c). In terms of the Dirichlet conditions (section 12.1), the others fail as
follows: (a) (i); (b) (ii); (d) (ii); (e) (iii).
12.4 an=[ ( 4 /(nπ)](−1)(n−1)/2fornodd and an=0f o r neven. In (b) use the expansion
of sin( A+B).
12.5 f(x)=2
P∞
1(−1)n+1n−1sinnx;s e t x=π/2.
12.6 (a)
P[(2/(nπ)]sin nπx,a l ln;( b )
P[(4/(n2π2)] cos nπxfor odd nonly. The cosine
series, with n−2convergence and alternate terms missing; the sine continuation
contains a discontinuity.
12.7 (i) Series (a) from exercise 12.6 does not converge and cannot represent the
function y(x)=−1. Series (b) reproduces the square-wave function of equation
(12.8).(ii) Series (a) gives the series for y(x)=−x−
1
2x2−1
2in the range −1≤x≤0
and for y(x)=x−1
2x2−1
2in the range 0 ≤x≤1. Series (b) gives the series for
y(x)=x+1
2x2+1
2in the range −1≤x≤0a n df o r y(x)=x−1
2x2+1
2in the
range 0≤x≤1.
12.8 The even continuation has a discontinuity in its derivative at x=π, whilst the
odd continuation does not; thus the sine series will have better convergence.b
1=π/2;b2m+1=0f o r m>0;b2m=−16m/[π(4m2−1)2].
12.9 f(x)=( s i n h1 )
/
1+2
P∞
1(−1)n(1 +n2π2)−1[cos(nπx)−nπsin(nπx)]
/
;
f(2) = f(0) = 1.
12.10 Combine the coefficients of the sin( nπx) terms from the Fourier series for xand
(part of)
R
expxd x; the partial series obtained by differentiating the sin( nπx)
terms does not converge, having coefficients of the form ( nπ)2/[1 + ( nπ)2].
12.11 See figure 12.6. (c) (i) (1 + e−1)/2, (ii) (1 + e−1)/2; (d) (i) (1 + e−4)/2, (ii) e−1.
(a) (b) (c) (d)00 0 0 11 12 4
Figure 12.6 Continuations of exp( −x2)i n0≤x≤1t og i v e :( a)c o s i n et e r m s
only; ( b) sine terms only; ( c)p e r i o d1 ;( d)p e r i o d2 .
12.12 (a) a0and odd cosines; (b) all, there is no symmetry about T/4 for the periodic
function; (c) Odd cosines.
12.13 (d) (i) The periods are both 2 L; (ii) y0/2.
12.14 g(x)=1
2x(π−x)f o r x≥0a n d=1
2x(π+x)f o r x≤0. Set x=π/2;S=π3/32.
437
FOURIER SERIES
12.15 So=π2/8. If Se=
P(2m)−2then Se=1
4(Se+So), yielding So−Se=π2/12 and
(Se+So)=π2/6.
(a) (π/4)(π/2−|θ|); (b) ( πθ/4)(π/2−|θ|/2) from integrating (a). (c) Even function;
average value L2/3;y(0) = 0; y(L)=L2; probably y(x)=x2. Compare with the
worked example in section 12.5.
12.16 cosh( t−1) = (sinh 1)[1+2
P∞
n=1(cosnπt)/(n2π2+1)]; set t= 0 to obtain the stated
result; set t= 1 to evaluate
P(−1)n/(n2π2+1) and add and subtract this quantity
from
P(n2π2+1 )−1.
P
odd=(e−1)/[4(e+ 1)];
P
even=( 3−e)/[4(e−1)].
12.17 cosh x=( s i n h1 ) [ 1+2
P∞
n=1(−1)n(cosnπx)/(n2π2+1)] and after integrating twice
this form must be recovered. Use x2=1
3+4
P(−1)n(cosnπx)/(n2π2)] to eliminate
the quadratic term arising from the constants of integration; there is no linearterm.
12.18 Consider f(x)=−x
2for−2<x≤0, to ensure a sine series;P
nbnsin(nπx/2), with bn=(−1)n+18/(nπ)f o r neven and ( −1)n+18/(nπ)−
32/(nπ)3fornodd.
12.19 C±(2m+1)=∓2i/[(2m+1 )π];
P|Cn|2=( 4/π2)×2×(π2/8); the values n=±1,
±3 contribute >90% of the total.
12.20 Write sin θcosnθas1
2[sin(n+1 )θ−sin(n−1)θ]; obtain
P∞
1(4m2−1)−1=1
2as
well as
P∞
1(−1)m(4m2−1)−1=1
2−π
4and add the two equations;1
2,1
2−π
8.
12.21 cn=(−1)n[sinh π+in(cosh π−1)]/[π(1 +n2)].
12.22 cn=( 2 /π)[(4ni−1)/(16n2−1)].The correct value is the mean of the two
incorrect ones, i.e. (4 −π)/8.Write (16 n2−1)−1in partial fractions and compare
with exercise 12.5.
12.23 ( π2−8)/16.
12.24 an=8y0/(n2π2) for odd n, an= 0 otherwise; S=π4/96.
12.25 (b) All anandαnare zero; bn=2 (−1)n+1/(nπ)a n d βn=4/(nπ). You will need
the result quoted in exercise 12.19.
12.26 Show that the minimising value bkis given by 0 =
Rπ
−πf(x)si nkx dx−
Pm
n=1bnπδkn
and hence that bkis equal to the normal Fourier coefficient; b1=−8/π,b2=0 ,
b3=−8/(27π);E3=( 6 4 /π)
P∞
2(2m+1 )−6.
438
13
Integral transforms
In the previous chapter we encountered the Fourier series representation of a
periodic function in a fixed interval as a superposition of sinusoidal functions. It isoften desirable, however, to obtain such a representation even for functions definedover an infinite interval and with no particular periodicity. Such a representationis called a Fourier transform and is one of a class of representations called integral
transforms .
We begin by considering Fourier transforms as a generalisation of Fourier
series. We then go on to discuss the properties of the Fourier transform and its
applications. In the second part of the chapter we present an analogous discussionof the closely related Laplace transform .
13.1 Fourier transforms
The Fourier transform provides a representation of functions defined over an
infinite interval and having no particular periodicity, in terms of a superpositionof sinusoidal functions. It may thus be considered as a generalisation of the
Fourier series representation of periodic functions. Since Fourier transforms are
often used to represent time-varying functions, we shall present much of ourdiscussion in terms of f(t), rather than f(x), although in some spatial examples
f(x) will be the more natural notation and we shall use it as appropriate. Our
only requirement on f(t) will be thatintegraltext
∞
−∞|f(t)|dtis finite.
In order to develop the transition from Fourier series to Fourier transforms, we
first recall that a function of period Tmay be represented as a complex Fourier
series, cf. (12.9),
f(t)=∞summationdisplay
r=−∞cre2πirt/T=∞summationdisplay
r=−∞creiωrt, (13.1)
where ωr=2πr/T. As the period Ttends to infinity, the ‘frequency quantum’
439
INTEGRAL TRANSFORMS
c(ω)e x p iωt
−100
12 r−2π
T2π
T4π
T ωr
Figure 13.1 The relationship between the Fourier terms for a function of
period Tand the Fourier integral (the area below the solid line) of the
function.
∆ω=2π/Tbecomes vanishingly small and the spectrum of allowed frequencies
ωrbecomes a continuum. Thus, the infinite sum of terms in the Fourier series
becomes an integral, and the coefficients crbecome functions of the continuous
variable ω, as follows.
We recall, cf. (12.10), that the coefficients crin (13.1) are given by
cr=1
TintegraldisplayT/2
−T/2f(t)e−2πirt/Tdt=∆ω
2πintegraldisplayT/2
−T/2f(t)e−iωrtdt, (13.2)
where we have written the integral in two alternative forms and, for convenience,
made one period run from −T/2t o+ T/2 rather than from 0 to T. Substituting
from (13.2) into (13.1) gives
f(t)=∞summationdisplay
r=−∞∆ω
2πintegraldisplayT/2
−T/2f(u)e−iωrudu eiωrt. (13.3)
At this stage ωris still a discrete function of requal to 2 πr/T.
The solid points in figure 13.1 are a plot of (say, the real part of) creiωrtas
a function of r(or equivalently of ωr) and it is clear that (2 π/T)creiωrtgives
the area of the rth broken-line rectangle. If Ttends to∞then ∆ ω(= 2π/T)
becomes infinitesimal, the width of the rectangles tends to zero and, from the
mathematical definition of an integral,
∞summationdisplay
r=−∞∆ω
2πg(ωr)eiωrt→1
2πintegraldisplay∞
−∞g(ω)eiωtdω.
In this particular case
g(ωr)=integraldisplayT/2
−T/2f(u)e−iωrudu,
440
13.1 FOURIER TRANSFORMS
and (13.3) becomes
f(t)=1
2πintegraldisplay∞
−∞dω eiωtintegraldisplay∞
−∞du f(u)e−iωu. (13.4)
This result is known as Fourier’s inversion theorem .
From it we may define the Fourier transform off(t)b y
tildewidef(ω)=1√
2πintegraldisplay∞
−∞f(t)e−iωtdt, (13.5)
and its inverse by
f(t)=1√
2πintegraldisplay∞
−∞tildewidef(ω)eiωtdω. (13.6)
Including the constant 1 /√
2πin the definition of tildewidef(ω) (whose mathematical
existence as T→∞is assumed here without proof) is clearly arbitrary, the only
requirement being that the product of the constants in (13.5) and (13.6) shouldequal 1 /(2π). Our definition is chosen to be as symmetric as possible.IFind the Fourier transform of the exponential decay function f(t)=0 fort<0and
f(t)=Ae−λtfort≥0(λ>0).
Using the definition (13.5) and separating the integral into two parts,ef(ω)=1√
2π
Z0
−∞(0)e−iωtdt+A√
2π
Z∞
0e−λte−iωtdt
=0+A√
2π
/
−e−(λ+iω)t
λ+iω
/∞
0
=A√
2π(λ+iω),
which is the required transform. It is clear that the multiplicative constant Adoes not
affect the form of the transform, merely its amplitude. This transform may be verified by
re-substitution of the above re sult into (13.6) to recover f(t), but evaluation of the integral
requires the use of complex-variable contour integration (chapter 20).
J
13.1.1 The uncertainty principle
An important function that appears in many areas of physical science, either
precisely or as an approximation to a physical situation, is the Gaussian or
normal distribution. Its Fourier transform is of importance both in itself and also
because, when interpreted statistically, it readily illustrates a form of uncertainty
principle .
441
INTEGRAL TRANSFORMSIFind the Fourier transform of the normalised Gaussian distribution
f(t)=1
τ√
2πexp
/
−t2
2τ2
/
,−∞<t<∞.
This Gaussian distribution is centred on t= 0 and has a root mean square deviation
∆t=τ. (Any reader who is unfamiliar with this interpretation of the distribution should
refer to chapter 26.)
Using the definition (13.5), the Fourier transform of f(t)i sg i v e nb yef(ω)=1√
2π
Z∞
−∞1
τ√
2πexp
/
−t2
2τ2
/
exp(−iωt)dt
=1√
2π
Z∞
−∞1
τ√
2πexp
/
−1
2τ2
/
t2+2τ2iωt+(τ2iω)2−(τ2iω)2
/
/
dt,
where the quantity −(τ2iω)2/(2τ2) has been both added and subtracted in the exponent
in order to allow the factors involving the variable of integration tto be expressed as a
complete square. Hence the expression can be writtenef(ω)=exp(−1
2τ2ω2)√
2π
/1
τ√
2π
Z∞
−∞exp
/
−(t+iτ2ω)2
2τ2
/
dt
/
.
The quantity inside the braces is the normalisation integral for the Gaussian and equals
unity, although to show this strictly needs results from complex variable theory (chapter 20).That it is equal to unity can be made plausible by changing the variable to s=t+iτ
2ω
and assuming that the imaginary parts introduced into the integration path and limits
(where the integrand goes rapidly to zero anyway) make no difference.
We are left with the result thatef(ω)=1√
2πexp
/−τ2ω2
2
/
, (13.7)
which is another Gaussian distribution, centred on zero and with a root mean square
deviation ∆ ω=1/τ. It is interesting to note, and an important property, that the Fourier
transform of a Gaussian is another Gaussian.
J
In the above example the root mean square deviation in twasτ, and so it is
seen that the deviations or ‘spreads’ in tand in ωare inversely related:
∆ω∆t=1,
independently of the value of τ. In physical terms, the narrower in time is, say, an
electrical impulse the greater the spread of frequency components it must contain.Similar physical statements are valid for other pairs of Fourier-related variables,such as spatial position and wave number. In an obvious notation, ∆ k∆x=1f o r
a Gaussian wave packet.
The uncertainty relations as usually expressed in quantum mechanics can be
related to this if the de Broglie and Einstein relationships for momentum and
energy are introduced; they are
p=/planckover2pi1kand E=/planckover2pi1ω.
Here/planckover2pi1is Planck’s constant hdivided by 2 π. In a quantum mechanics setting f(t)
442
13.1 FOURIER TRANSFORMS
is a wavefunction and the distribution of the wave intensity in time is given by
|f|2(also a Gaussian). Similarly, the intensity distribution in frequency is given
by|tildewidef|2. These two distributions have respective root mean square deviations of
τ/√
2a n d1 /(√
2τ), giving, after incorporation of the above relations,
∆E∆t=/planckover2pi1/2a n d∆ p∆x=/planckover2pi1/2.
The factors of 1 /2 that appear are specific to the Gaussian form, but any
distribution f(t) produces for the product ∆ E∆ta quantity λ/planckover2pi1in which λis
strictly positive (in fact the value 1 /2 for a Gaussian is the minimum possible).
13.1.2 Fraunhofer diffraction
We take our final example of the Fourier transform from the field of optics. The
pattern of transmitted light produced by a partially opaque (or phase-changing)object upon which a coherent beam of radiation falls is called a diffraction pattern
and, in particular, when the cross-section of the object is small compared with
the distance at which the light is observed the pattern is known as a Fraunhofer
diffraction pattern.
We will consider only the case in which the light is monochromatic with
wavelength λ. The direction of the incident beam of light can then be described
by the wave vector k; the magnitude of this vector is given by the wave number
k=2π/λof the light. The essential quantity in a Fraunhofer diffraction pattern
is the dependence of the observed amplitude (and hence intensity) on the angle θ
between the viewing direction k
/primeand the direction kof the incident beam. This
is entirely determined by the spatial distribution of the amplitude and phase ofthe light at the object, the transmitted intensity in a particular direction k
/primebeing
determined by the corresponding Fourier component of this spatial distribution.
As an example, we take as an object a simple two-dimensional screen of width
2Yon which light of wave number kis incident normally; see figure 13.2. We
suppose that at the position (0 ,y) the amplitude of the transmitted light is f(y)
per unit length in the y-direction ( f(y) may be complex). The function f(y)i s
called an aperture function . Both the screen and beam are assumed infinite in the
z-direction.
Denoting the unit vectors in the x-a n d y- directions by iandjrespectively,
the total light amplitude at a position r0=x0i+y0j, with x0>0, will be the
superposition of all the (Huyghens’) wavelets originating from the various parts
of the screen. For large r0(=|r0|), these can be treated as plane waves to give †
A(r0)=integraldisplayY
−Yf(y)ex p[ ik/prime·(r0−yj)]
|r0−yj|dy. (13.8)
†This is the approach first used by Fresnel. For simplicity we have omitted from the integral a
multiplicative inclination factor that depends on angle θand decreases as θincreases.
443
INTEGRAL TRANSFORMS
−YYy
xkk/prime
0θ
Figure 13.2 Diffraction grating of width 2 Ywith light of wavelength 2 π/k
being diffracted through an angle θ.
The factor exp[ ik/prime·(r0−yj)] represents the phase change undergone by the light
in travelling from the point yjon the screen to the point r0, and the denominator
represents the reduction in amplitude with distance. (Recall that the system isinfinite in the z-direction and so the ‘spreading’ is effectively in two dimensions
only.)
If the medium is the same on both sides of the screen then k
/prime=kcosθi+ksinθj,
and if r0/greatermuchYthen expression (13.8) can be approximated by
A(r0)=exp(ik/prime·r0)
r0integraldisplay∞
−∞f(y)ex p(−ikysinθ)dy. (13.9)
We have used that f(y)=0f o r|y|>Y, to extend the integral to infinite limits.
The intensity in the direction θis then given by
I(θ)=|A|2=2π
r02|tildewidef(q)|2, (13.10)
where q=ksinθ.IEvaluate I(θ)for an aperture consisting of two long slits each of width 2bwhose centres
are separated by a distance 2a,a>b; the slits are illuminated by light of wavelength λ.
The aperture function is plotted in figure 13.3. We first need to find
ef(q):ef(q)=1√
2π
Z−a+b
−a−be−iqxdx+1√
2π
Za+b
a−be−iqxdx
=1√
2π
/
−e−iqx
iq
/−a+b
−a−b+1√
2π
/
−e−iqx
iq
/a+b
a−b
=−1
iq√
2π
/
e−iq(−a+b)−e−iq(−a−b)+e−iq(a+b)−e−iq(a−b)
/
.
444
13.1 FOURIER TRANSFORMS
f(y)
1
a−b −a−b a+b −a+ba −a x
Figure 13.3 The aperture function f(y) for two wide slits.
After some manipulation we obtainef(q)=4c os qasinqb
q√
2π.
Now applying (13.10), and remembering that q=( 2πsinθ)/λ, we find
I(θ)=16cos2qasin2qb
q2r02,
where r0is the distance from the centre of the aperture.
J
13.1.3 The Dirac δ-function
Before going on to consider further properties of Fourier transforms we make a
digression to discuss the Dirac δ-function and its relation to Fourier transforms.
The δ-function is different from most functions encountered in the physical
sciences but we will see that a rigorous mathematical definition exists and theutility of the δ-function will be demonstrated throughout the remainder of this
chapter. It can be visualised as a very sharp narrow pulse (in space, time, density,etc.) which produces an integrated effect having a definite magnitude. The formal
properties of the δ-function may be summarised as follows.
The Dirac δ-function has the property that
δ(t)=0 f o r t/negationslash=0, (13.11)
but its fundamental defining property is
integraldisplay
f(t)δ(t−a)dt=f(a), (13.12)
provided the range of integration includes the point t=a; otherwise the integral
445
INTEGRAL TRANSFORMS
equals zero. This leads immediately to two further useful results:
integraldisplayb
−aδ(t)dt=1 f o ra l l a, b > 0 (13.13)
andintegraldisplay
δ(t−a)dt=1, (13.14)
provided the range of integration includes t=a.
Equation (13.12) can be used to derive further useful properties of the Dirac
δ-function:
δ(t)=δ(−t), (13.15)
δ(at)=1
|a|δ(t), (13.16)
tδ(t)=0 . (13.17)IProve that δ(bt)=δ(t)/|b|.
Let us first consider the case where b>0. It follows thatZ∞
−∞f(t)δ(bt)dt=
Z∞
−∞f
/t/prime
b
/
δ(t/prime)dt/prime
b=1
bf(0) =1
b
Z∞
−∞f(t)δ(t)dt,
where we have made the substitution t/prime=bt.B u t f(t) is arbitrary and so we immediately
see that δ(bt)=δ(t)/b=δ(t)/|b|forb>0.
Now consider the case where b=−c<0. It follows thatZ∞
−∞f(t)δ(bt)dt=
Z−∞
∞f
/t/prime
−c
/
δ(t/prime)
/dt/prime
−c
/
=
Z∞
−∞1
cf
/t/prime
−c
/
δ(t/prime)dt/prime
=1
cf(0) =1
|b|f(0) =1
|b|
Z∞
−∞f(t)δ(t)dt,
where we have made the substitution t/prime=bt=−ct.B u t f(t) is arbitrary and so
δ(bt)=1
|b|δ(t),
for all b, which establishes the result.
J
Furthermore, by considering an integral of the form
integraldisplay
f(t)δ(h(t))dt,
and making a change of variables to z=h(t), we may show that
δ(h(t)) =summationdisplay
iδ(t−ti)
|h/prime(ti)|, (13.18)
where the tiare those values of tfor which h(t)=0a n d h/prime(t) stands for dh/dt.
446
13.1 FOURIER TRANSFORMS
The derivative of the delta function, δ/prime(t), is defined by
integraldisplay∞
−∞f(t)δ/prime(t)dt=bracketleftBig
f(t)δ(t)bracketrightBig∞
−∞−integraldisplay∞
−∞f/prime(t)δ(t)dt
=−f/prime(0), (13.19)
and similarly for higher derivatives.
For many practical purposes, effects that are not strictly described by a δ-
function may be analysed as such, if they take place in an interval much shorter
than the response interval of the system on which they act. For example, the
idealised notion of an impulse of magnitude Japplied at time t0can be represented
by
j(t)=Jδ(t−t0). (13.20)
Many physical situations are described by a δ-function in space rather than in
time. Moreover, we often require the δ-function to be defined in more than one
dimension. For example, the charge density of a point charge qat a point r0may
be expressed as a three-dimensional δ-function
ρ(r)=qδ(r−r0)=qδ(x−x0)δ(y−y0)δ(z−z0), (13.21)
so that a discrete ‘quantum’ is expressed as if it were a continuous distribution.
From (13.21) we see that (as expected) the total charge enclosed in a volume V
is given by
integraldisplay
Vρ(r)dV=integraldisplay
Vqδ(r−r0)dV=braceleftBigg
qifr0lies in V,
0o t h e r w i s e .
Closely related to the Dirac δ-function is the Heaviside orunit step function
H(t), for which
H(t)=braceleftBigg
1f o r t>0,
0f o r t<0.(13.22)
This function is clearly discontinuous at t= 0 and it is usual to take H(0) = 1 /2.
The Heaviside function is related to the delta function by
H/prime(t)=δ(t). (13.23)
447
INTEGRAL TRANSFORMSIProve relation (13.23).
Considering the integralZ∞
−∞f(t)H/prime(t)dt=
/
f(t)H(t)
/∞
−∞−
Z∞
−∞f/prime(t)H(t)dt
=f(∞)−
Z∞
0f/prime(t)dt
=f(∞)−
/
f(t)
/∞
0=f(0),
and comparing it with (13.12) when a= 0 immediately shows that H/prime(t)=δ(t).
J
13.1.4 Relation of the δ-function to Fourier transforms
In the previous section we introduced the Dirac δ-function as a way of repre-
senting very sharp narrow pulses, but in no way related it to Fourier transforms.We now show that the δ-function can equally well be defined in a way that more
naturally relates it to the Fourier transform.
Referring back to the Fourier inversion theorem (13.4), we have
f(t)=1
2πintegraldisplay∞
−∞dω eiωtintegraldisplay∞
−∞du f(u)e−iωu
=integraldisplay∞
−∞du f(u)braceleftbigg1
2πintegraldisplay∞
−∞eiω(t−u)dωbracerightbigg
.
Comparison of this with (13.12) shows that we may write the δ-function as
δ(t−u)=1
2πintegraldisplay∞
−∞eiω(t−u)dω. (13.24)
Considered as a Fourier transform, this representation shows that a very
narrow time peak at t=uresults from the superposition of a complete spectrum
of harmonic waves, all frequencies having the same amplitude and all waves beingin phase at t=u. This suggests that the δ-function may also be represented as
the limit of the transform of a uniform distribution of unit height as the width
of this distribution becomes infinite.
Consider the rectangular distribution of frequencies shown in figure 13.4( a).
From (13.6), taking the inverse Fourier transform,
f
Ω(t)=1√
2πintegraldisplayΩ
−Ω1×eiωtdω
=2Ω√
2πsin Ωt
Ωt. (13.25)
This function is illustrated in figure 13.4( b) and it is apparent that, for large Ω, it
becomes very large at t= 0 and also very narrow about t= 0, as we qualitatively
448
13.1 FOURIER TRANSFORMS
ω
(a) (b)Ω −Ω t
π
Ω1
efΩfΩ(t)
2Ω
(2π)1/2
Figure 13.4 ( a) A Fourier transform showing a rectangular distribution of
frequencies between ±Ω; (b) the function of which it is the transform, which
is proportional to t−1sin Ωt.
expect and require. We also note that, in the limit Ω →∞,fΩ(t), as defined by
the inverse Fourier transform, tends to (2 π)1/2δ(t) by virtue of (13.24). Hence we
may conclude that the δ-function can also be represented by
δ(t) = lim
Ω→∞parenleftbiggsinΩt
πtparenrightbigg
. (13.26)
Several other function representations are equally valid, e.g. the limiting cases of
rectangular, triangular or Gaussian distributions; the only essential requirementsare a knowledge of the area under such a curve and that undefined operationssuch as dividing by zero are not inadvertently carried out on the δ-function whilst
some non-explicit representation is being employed.
We also note that the Fourier transform definition of the delta function, (13.24),
shows that the latter is real since
δ
∗(t)=1
2πintegraldisplay∞
−∞e−iωtdω=δ(−t)=δ(t).
Finally, the Fourier transform of a δ-function is simply
tildewideδ(ω)=1√
2πintegraldisplay∞
−∞δ(t)e−iωtdt=1√
2π. (13.27)
13.1.5 Properties of Fourier transforms
Having considered the Dirac δ-function, we now return to our discussion of the
properties of Fourier transforms. As we would expect, Fourier transforms have
many properties analogous to those of Fourier series in respect of the connectionbetween the transforms of related functions. Here we list these properties without
449
INTEGRAL TRANSFORMS
proof; they can be verified by working from the definition of the transform. As
previously, we denote the Fourier transform of f(t)b ytildewidef(ω)o r F[f(t)].
(i) Differentiation:Fbracketleftbig
f/prime(t)bracketrightbig
=iωtildewidef(ω). (13.28)
This may be extended to higher derivatives, so thatFbracketleftbig
f/prime/prime(t)bracketrightbig
=iω Fbracketleftbig
f/prime(t)bracketrightbig
=−ω2tildewidef(ω),
a n ds oo n .
(ii) Integration:Fbracketleftbiggintegraldisplayt
f(s)dsbracketrightbigg
=1
iωtildewidef(ω)+2πcδ(ω), (13.29)
where the term 2 πcδ(ω) represents the Fourier transform of the constant
of integration associated with the indefinite integral.
(iii) Scaling:F[f(at)]=1
atildewidefparenleftBigω
aparenrightBig
. (13.30)
(iv) Translation:F[f(t+a)]=eiaωtildewidef(ω). (13.31)
(v) Exponential multiplication:Fbracketleftbig
eαtf(t)bracketrightbig
=tildewidef(ω+iα), (13.32)
where αmay be real, imaginary or complex.IProve relation (13.28).
Calculating the Fourier transform of f/prime(t) directly, we obtainF
/
f/prime(t)
/
=1√
2π
Z∞
−∞f/prime(t)e−iωtdt
=1√
2π
/
e−iωtf(t)
/∞
−∞+1√
2π
Z∞
−∞iω e−iωtf(t)dt
=iω
ef(ω),
iff(t)→0a tt=±∞,a si tm u s ts i n c e
R∞
−∞|f(t)|dtis finite.
J
To illustrate a use and also a proof of (13.32), let us consider an amplitude-
modulated radio wave. Suppose a message to be broadcast is represented by f(t).
The message can be added electronically to a constant signal aof magnitude
such that a+f(t) is never negative, and then the sum can be used to modulate
450
13.1 FOURIER TRANSFORMS
the amplitude of a carrier signal of frequency ωc. Using a complex exponential
notation, the transmitted amplitude is now
g(t)=A[a+f(t)]eiωct. (13.33)
Ignoring in the present context the effect of the term Aaexp(iωct), which gives a
contribution to the transmitted spectrum only at ω=ωc,w eo b t a i nf o rt h en e w
spectrum
tildewideg(ω)=1√
2πAintegraldisplay∞
−∞f(t)eiωcte−iωtdt
=1√
2πAintegraldisplay∞
−∞f(t)e−i(ω−ωc)tdt
=Atildewidef(ω−ωc), (13.34)
which is simply a shift of the whole spectrum by the carrier frequency. The use
of different carrier frequencies enables signals to be separated.
13.1.6 Odd and even functions
Iff(t) is odd or even then we may derive alternative forms of Fourier’s inversion
theorem, which lead to the definition of different transform pairs. Let us firstconsider an odd function f(t)=−f(−t), whose Fourier transform is given by
tildewidef(ω)=1
√
2πintegraldisplay∞
−∞f(t)e−iωtdt
=1√
2πintegraldisplay∞
−∞f(t)(cos ωt−isinωt)dt
=−2i√
2πintegraldisplay∞
0f(t)s i nωtdt,
where in the last line we use the fact that f(t)a n ds i n ωtare odd, whereas cos ωt
is even.
We note that tildewidef(−ω)=−tildewidef(ω), i.e.tildewidef(ω) is an odd function of ω. Hence
f(t)=1√
2πintegraldisplay∞
−∞tildewidef(ω)eiωtdω=2i√
2πintegraldisplay∞
0tildewidef(ω)s i nωtdω
=2
πintegraldisplay∞
0dωsinωtbraceleftbiggintegraldisplay∞
0f(u)s i nωudubracerightbigg
.
Thus we may define the Fourier sine transform pair for odd functions:
tildewidefs(ω)=radicalbigg
2
πintegraldisplay∞
0f(t)s i n ωtdt, (13.35)
f(t)=radicalbigg
2
πintegraldisplay∞
0tildewidefs(ω)s i n ωtdω. (13.36)
451
INTEGRAL TRANSFORMS
g(y)
(a)
(b)
(c)
(d)
y 0
Figure 13.5 Resolution functions: ( a)i d e a l δ-function; ( b) typical unbiased
resolution; ( c)a n d( d) biases tending to shift observations to higher values
than the true one.
Note that although the Fourier sine transform pair was derived by considering
an odd function f(t) defined over all t, the definitions (13.35) and (13.36) only
require f(t)a n dtildewidefs(ω) to be defined for positive tandωrespectively. For an
even function, i.e. one for which f(t)=f(−t), we can define the Fourier cosine
transform pair in a similar way, but with sin ωtreplaced by cos ωt.
13.1.7 Convolution and deconvolution
It is apparent that any attempt to measure the value of a physical quantity is
limited, to some extent, by the finite resolution of the measuring apparatus used.
On the one hand, the physical quantity we wish to measure will be in general a
function of an independent variable, xsay, i.e. the true function to be measured
takes the form f(x). On the other hand, the apparatus we are using does not give
the true output value of the function; a resolution function g(y) is involved. By
this we mean that the probability that an output value y= 0 will be recorded
instead as being between yandy+dyis given by g(y)dy. Some possible resolution
functions of this sort are shown in figure 13.5. To obtain good results we wish
the resolution function to be as close to a δ-function as possible (case ( a)). A
typical piece of apparatus has a resolution function of finite width, although ifit is accurate the mean is centred on the true value (case ( b)). However, some
apparatus may show a bias that tends to shift observations to higher or lowervalues than the true ones (cases ( c)a n d( d)), thereby exhibiting systematic error.
Given that the true distribution is f(x) and the resolution function of our
452
13.1 FOURIER TRANSFORMS
−a −a aax y zf(x)
b−b2b 2bg(y) h(z) ∗ =
1
Figure 13.6 The convolution of two functions f(x)a n d g(y).
measuring apparatus is g(y), we wish to calculate what the observed distribution
h(z) will be. The symbols x,yandzall refer to the same physical variable (e.g.
length or angle), but are denoted differently because the variable appears in theanalysis in three different roles.
The probability that a true reading lying between xandx+dx, and so having
probability f(x)dxof being selected by the experiment, will be moved by the
instrumental resolution by an amount z−xinto a small interval of width dzis
g(z−x)dz. Hence the combined probability that the interval dxwill give rise to
an observation appearing in the interval dzisf(x)dx g(z−x)dz. Adding together
the contributions from all values of xt h a tc a nl e a dt oa no b s e r v a t i o ni nt h er a n g e
ztoz+dz, we find that the observed distribution is given by
h(z)=integraldisplay
∞
−∞f(x)g(z−x)dx. (13.37)
The integral in (13.37) is called the convolution of the functions fandgand is
often written f∗g. The convolution defined above is commutative ( f∗g=g∗f),
associative and distributive. The observed distribution is thus the convolution ofthe true distribution and the experimental resolution function. The result will bethat the observed distribution is broader and smoother than the true one and, if
g(y) has a bias, the maxima will normally be displaced from their true positions.
It is also obvious from (13.37) that if the resolution is the ideal δ-function,
g(y)=δ(y)t h e n h(z)=f(z) and the observed distribution is the true one.
It is interesting to note, and a very important property, that the convolution of
any function g(y) with a number of delta functions leaves a copy of g(y)a tt h e
position of each of the delta functions.IFind the convolution of the function f(x)=δ(x+a)+δ(x−a)with the function g(y)
plotted in figure 13.6.
Using the convolution integral (13.37)
h(z)=
Z∞
−∞f(x)g(z−x)dx=
Z∞
−∞[δ(x+a)+δ(x−a)]g(z−x)dx
=g(z+a)+g(z−a).
This convolution h(z) is plotted in figure 13.6.
J
453
INTEGRAL TRANSFORMS
Let us now consider the Fourier transform of the convolution (13.37); this is
given by
tildewideh(k)=1√
2πintegraldisplay∞
−∞dz e−ikzbraceleftbiggintegraldisplay∞
−∞f(x)g(z−x)dxbracerightbigg
=1√
2πintegraldisplay∞
−∞dx f(x)braceleftbiggintegraldisplay∞
−∞g(z−x)e−ikzdzbracerightbigg
.
If we let u=z−xin the second integral we have
tildewideh(k)=1√
2πintegraldisplay∞
−∞dx f(x)braceleftbiggintegraldisplay∞
−∞g(u)e−ik(u+x)dubracerightbigg
=1√
2πintegraldisplay∞
−∞f(x)e−ikxdxintegraldisplay∞
−∞g(u)e−ikudu
=1√
2π×√
2πtildewidef(k)×√
2πtildewideg(k)=√
2πtildewidef(k)tildewideg(k). (13.38)
Hence the Fourier transform of a convolution f∗gis equal to the product of the
separate Fourier transforms multiplied by√
2π; this result is called the convolution
theorem .
It may be proved similarly that the converse is also true, namely that the
Fourier transform of the product f(x)g(x)i sg i v e nb yF[f(x)g(x)]=1√
2πtildewidef(k)∗tildewideg(k). (13.39)IFind the Fourier transform of the function in figure 13.3 representing two wide slits by
considering the Fourier transforms of (i) two δ-functions, at x=±a, (ii) a rectangular
function of height 1and width 2bcentred on x=0.
(i) The Fourier transform of the two δ-functions is given byef(q)=1√
2π
Z∞
−∞δ(x−a)e−iqxdx+1√
2π
Z∞
−∞δ(x+a)e−iqxdx
=1√
2π
/;
e−iqa+eiqa
/
=2c os qa√
2π.
(ii) The Fourier transform of the broad slit iseg(q)=1√
2π
Zb
−be−iqxdx=1√
2π
/e−iqx
−iq
/b
−b
=−1
iq√
2π(e−iqb−eiqb)=2si nqb
q√
2π.
We have already seen that the convolution of these functions is the required function
representing two wide slits (see figure 13.6). So, using the convolution theorem, the Fourier
transform of the convolution is√
2πtimes the product of the individual transforms, i.e.
4cos qasinqb/(q√
2π). This is, of course, the same result as that obtained in the example
in subsection 13.1.2.
J
454
13.1 FOURIER TRANSFORMS
The inverse of convolution, called deconvolution , allows us to find a true
distribution f(x) given an observed distribution h(z) and a resolution function
g(y).IAn experimental quantity f(x)is measured using apparatus with a known resolution func-
tiong(y)to give an observed distribution h(z).H o wm a y f(x)be extracted from the mea-
sured distribution?
From the convolution theorem (13.38), the Fourier transform of the measured distributioniseh(k)=√
2π
ef(k)
eg(k),
from which we obtainef(k)=1√
2π
eh(k)eg(k).
Then on inverse Fourier transforming we find
f(x)=1√
2π
F−1
/"eh(k)eg(k)
/#
.
In words, to extract the true distribution, we divide the Fourier transform of the observed
distribution by that of the resolution function for each value of kand then take the inverse
Fourier transform of the function so generated.
J
This explicit method of extracting true distributions is straightforward for exact
functions but, in practice, because of experimental and statistical uncertainties inthe experimental data or because data over only a limited range are available, itis often not very precise, involving as it does three (numerical) transforms eachrequiring in principle an integral over an infinite range.
13.1.8 Correlation functions and energy spectra
Thecross-correlation of two functions fandgis defined by
C(z)=integraldisplay
∞
−∞f∗(x)g(x+z)dx. (13.40)
Despite the formal similarity between (13.40) and the definition of the convolution
in (13.37), the use and interpretation of the cross-correlation and of the convo-lution are very different; the cross-correlation provides a quantitative measure ofthe similarity of two functions fandgas one is displaced through a distance z
relative to the other. The cross-correlation is often notated as C=f⊗g, and, like
convolution, it is both associative and distributive. Unlike convolution, however,it isnotcommutative, in fact
[f⊗g](z)=[g⊗f]
∗(−z). (13.41)
455
INTEGRAL TRANSFORMSIProve the Wiener–Kinchin theorem,eC(k)=√
2π[
ef(k)]∗eg(k). (13.42)
Following a method similar to that for the convolution of fandg, let us consider the
Fourier transform of (13.40):eC(k)=1√
2π
Z∞
−∞dz e−ikz
/Z∞
−∞f∗(x)g(x+z)dx
/
=1√
2π
Z∞
−∞dx f∗(x)
/Z∞
−∞g(x+z)e−ikzdz
/
.
Making the substitution u=x+zin the second integral we obtaineC(k)=1√
2π
Z∞
−∞dx f∗(x)
/Z∞
−∞g(u)e−ik(u−x)du
/
=1√
2π
Z∞
−∞f∗(x)eikxdx
Z∞
−∞g(u)e−ikudu
=1√
2π×√
2π[
ef(k)]∗×√
2π
eg(k)=√
2π[
ef(k)]∗eg(k).
J
Thus the Fourier transform of the cross-correlation of fandgis equal to
the product of [ tildewidef(k)]∗andtildewideg(k) multiplied by√
2π.T h i sas t a t e m e n to ft h e
Wiener–Kinchin theorem . Similarly we can derive the converse theoremFbracketleftbig
f∗(x)g(x)bracketrightbig
=1√
2πtildewidef⊗tildewideg.
If we now consider the special case where gis taken to be equal to fin (13.40)
then, writing the LHS as a(z), we have
a(z)=integraldisplay∞
−∞f∗(x)f(x+z)dx; (13.43)
this is called the auto-correlation function off(x). Using the Wiener–Kinchin
theorem (13.42) we see that
a(z)=1√
2πintegraldisplay∞
−∞tildewidea(k)eikzdk
=1√
2πintegraldisplay∞
−∞√
2π[tildewidef(k)]∗tildewidef(k)eikzdk,
so that a(z) is the inverse Fourier transform of√
2π|tildewidef(k)|2, which is in turn called
theenergy spectrum off.
13.1.9 Parseval’s theorem
Using the results of the previous section we can immediately obtain Parseval’s
theorem . The most general form of this (also called the multiplication theorem )i s
456
13.1 FOURIER TRANSFORMS
obtained simply by noting from (13.42) that the cross-correlation (13.40) of two
functions fandgc a nb ew r i t t e na s
C(z)=integraldisplay∞
−∞f∗(x)g(x+z)dx=integraldisplay∞
−∞[tildewidef(k)]∗tildewideg(k)eikzdk. (13.44)
Then, setting z= 0 gives the multiplication theorem
integraldisplay∞
−∞f∗(x)g(x)dx=integraldisplay
[tildewidef(k)]∗tildewideg(k)dk. (13.45)
Specialising further, by letting g=f, we derive the most common form of
Parseval’s theorem,
integraldisplay∞
−∞|f(x)|2dx=integraldisplay∞
−∞|tildewidef(k)|2dk. (13.46)
When fis a physical amplitude these integrals relate to the total intensity involved
in some physical process. We have already met a form of Parseval’s theorem for
Fourier series in chapter 12; it is in fact a special case of (13.46).IThe displacement of a damped harmonic oscillator as a function of time is given by
f(t)=
/(
0 fort<0,
e−t/τsinω0tfort≥0.
Find the Fourier transform of this function and so give a physical interpretation of Parseval’s
theorem.
Using the usual definition for the Fourier transform we findef(ω)=
Z0
−∞0×e−iωtdt+
Z∞
0e−t/τsinω0te−iωtdt.
Writing sin ω0tas (eiω0t−e−iω0t)/2iwe obtainef(ω)=0+1
2i
Z∞
0
/
e−it(ω−ω0−i/τ)−e−it(ω+ω0−i/τ)
/
dt
=1
2
/1
ω+ω0−i/τ−1
ω−ω0−i/τ
/
,
which is the required Fourier transform. The physical interpretation of |
ef(ω)|2is the energy
content per unit frequency interval (i.e. the energy spectrum ) whilst|f(t)|2is proportional to
the sum of the kinetic and potential energies of the oscillator. Hence (to within a constant)Parseval’s theorem shows the equivalence of these two alternative specifications for the
total energy.J
13.1.10 Fourier transforms in higher dimensions
The concept of the Fourier transform can be extended naturally to more than
one dimension. For instance we may wish to find the spatial Fourier transform of
457
INTEGRAL TRANSFORMS
two- or three-dimensional functions of position. For example, in three dimensions
we can define the Fourier transform of f(x, y, z)a s
tildewidef(kx,ky,kz)=1
(2π)3/2integraldisplayintegraldisplayintegraldisplay
f(x, y, z)e−ikxxe−ikyye−ikzzdx dy dz, (13.47)
and its inverse by as
f(x, y, z)=1
(2π)3/2integraldisplayintegraldisplayintegraldisplay
tildewidef(kx,ky,kz)eikxxeikyyeikzzdkxdkydkz. (13.48)
Denoting the vector with components kx,ky,kzbykand that with components
x, y, z byr, we can write the Fourier transform pair (13.47), (13.48) as
tildewidef(k)=1
(2π)3/2integraldisplay
f(r)e−ik·rd3r, (13.49)
f(r)=1
(2π)3/2integraldisplay
tildewidef(k)eik·rd3k. (13.50)
From these relations we may deduce that the three-dimensional Dirac δ-function
c a nb ew r i t t e na s
δ(r)=1
(2π)3integraldisplay
eik·rd3k. (13.51)
Similar relations to (13.49), (13.50) and (13.51) exist for spaces of other dimen-
sionalities.IIn three-dimensional space a function f(r)possesses spherical symmetry, so that f(r)=
f(r). Find the Fourier transform of f(r)as a one-dimensional integral.
Let us choose spherical polar coordinates in which the vector kof the Fourier transform
lies along the polar axis ( θ= 0). This we can do since f(r) is spherically symmetric. We
then have
d3r=r2sinθd rd θd φ and k·r=krcosθ,
where k=|k|. The Fourier transform is then given byef(k)=1
(2π)3/2
Z
f(r)e−ik·rd3r
=1
(2π)3/2
Z∞
0dr
Zπ
0dθ
Z2π
0dφ f(r)r2sinθe−ikrcosθ
=1
(2π)3/2
Z∞
0dr2πf(r)r2
Zπ
0dθsinθe−ikrcosθ.
The integral over θmay be straightforwardly evaluated by noting that
d
dθ(e−ikrcosθ)=ikrsinθe−ikrcosθ.
Thereforeef(k)=1
(2π)3/2
Z∞
0dr2πf(r)r2
/e−ikrcosθ
ikr
/θ=π
θ=0
=1
(2π)3/2
Z∞
04πr2f(r)
/sinkr
kr
/
dr.
J
458
13.2 LAPLACE TRANSFORMS
A similar result may be obtained for two-dimensional Fourier transforms in
which f(r)=f(ρ), i.e. f(r) is independent of azimuthal angle φ. In this case, using
the integral representation of the Bessel function J0(x) given at the very end of
subsection 16.7.3, we find
tildewidef(k)=1
2πintegraldisplay∞
02πρf(ρ)J0(kρ)dρ. (13.52)
13.2 Laplace transforms
Often we are interested in functions f(t) for which the Fourier transform does
not exist because f/negationslash→0a s t→∞, and so the integral defining tildewidefdoes not
converge. For example, the function f(t)=tdoes not possess a Fourier transform.
Furthermore, often we are interested in a given function only for t>0, for example
when we are given the value at t= 0 in an initial-value problem. This leads us to
consider the Laplace transform, ¯f(s)o r L[f(t)],o ff(t), which is defined by
¯f(s)≡integraldisplay∞
0f(t)e−stdt, (13.53)
provided that the integral exists. We assume here that sis real, but complex values
would have to be considered in a more detailed study. In practice, for a givenfunction f(t) there will be some real number s
0such that the integral in (13.53)
exists for s>s 0but diverges for s≤s0.
Through (13.53) we define a linear transformation L[]that converts functions
of the variable tto functions of a new variable s:L[af1(t)+bf2(t)]=a L[f1(t)]+b L[f2(t)]=a¯f1(s)+b¯f2(s). (13.54)IFind the Laplace transforms of the functions (i) f(t)=1, (ii) f(t)=eat, (iii) f(t)=tn,
forn=0,1,2,....
(i) By direct application of the definition of a Laplace transform (13.53), we findL[1]=
Z∞
0e−stdt=
/−1
se−st
/∞
0=1
s,ifs>0,
where the restriction s>0 is required for the integral to exist.
(ii) Again using (13.53) directly, we find
¯f(s)=
Z∞
0eate−stdt=
Z∞
0e(a−s)tdt
=
/e(a−s)t
a−s
/∞
0=1
s−aifs>a .
459
INTEGRAL TRANSFORMS
(iii) Once again using the definition (13.53) we have
¯fn(s)=
Z∞
0tne−stdt.
Integrating by parts we find
¯fn(s)=
/−tne−st
s
/∞
0+n
s
Z∞
0tn−1e−stdt
=0+n
s¯fn−1(s),ifs>0.
We now have a recursion relation between successive transforms and by calculating
¯f0we can infer ¯f1,¯f2,e t c .S i n c e t0= 1, (i) above gives
¯f0=1
s,ifs>0, (13.55)
and
¯f1(s)=1
s2,¯f2(s)=2!
s3,. . . , ¯fn(s)=n!
sn+1ifs>0.
Thus, in each case (i)–(iii), direct application of the definition of the Laplace transform
(13.53) yields the required result.
J
Unlike that for the Fourier transform, the inversion of the Laplace transform
is not an easy operation to perform, since an explicit formula for f(t), given ¯f(s),
is not straightforwardly obtained from ( 13.53). The general method for obtaining
an inverse Laplace transform makes use of complex variable theory and is notdiscussed until chapter 20. However, progress can be made without having to findanexplicit inverse, since we can prepare from (13.53) a ‘dictionary’ of the Laplace
transforms of common functions and, when faced with an inversion to carry out,
hope to find the given transform (together with its parent function) in the listing.
Such a list is given in table 13.1.
When finding inverse Laplace transforms using table 13.1, it is useful to note
that for all practical purposes the inverse Laplace transform is unique †and linear
so thatL−1bracketleftbig
a¯f1(s)+b¯f2(s)bracketrightbig
=af1(t)+bf2(t). (13.56)
In many practical problems the method of partial fractions can be useful in
producing an expression from which the inverse Laplace transform can be found.IUsing table 13.1 find f(t)if
¯f(s)=s+3
s(s+1 ).
Using partial fractions ¯f(s) may be written
¯f(s)=3
s−2
s+1.
†This is not strictly true, since two functions can differ from one another at a finite number of
isolated points but have the sameLaplace transform.
460
13.2 LAPLACE TRANSFORMS
f(t) ¯f(s) s0
cc / s 0
ctncn!/sn+10
sinbt b/ (s2+b2)0
cosbt s/ (s2+b2)0
eat1/(s−a) a
tneatn!/(s−a)n+1a
sinhat a/ (s2−a2) |a|
coshat s/ (s2−a2) |a|
eatsinbt a/ [(s−a)2+b2] a
eatcosbt (s−a)/[(s−a)2+b2] a
t1/2 1
2(π/s3)1/20
t−1/2(π/s)1/20
δ(t−t0) e−st0 0
H(t−t0)=
/(
1f o r t≥t0
0f o r t<t 0e−st0/s 0
Table 13.1 Standard Laplace transforms. The transforms are valid for s>s 0.
Comparing this with the standard Laplace transforms in table 13.1, we find that the inverse
transform of 3 /sis 3 for s>0 and the inverse transform of 2 /(s+1 )i s2 e−tfors>−1,
and so
f(t)=3−2e−t,ifs>0.
J
13.2.1 Laplace transforms of derivatives and integrals
One of the main uses of Laplace transforms is in solving differential equations.
Differential equations are the subject of the next six chapters and we will returnto the application of Laplace transforms to their solution in chapter 15. Inthe meantime we will derive the required results, i.e. the Laplace transforms ofderivatives.
The Laplace transform of the first derivative of f(t)i sg i v e nb yLbracketleftbiggdf
dtbracketrightbigg
=integraldisplay∞
0df
dte−stdt
=bracketleftbig
f(t)e−stbracketrightbig∞
0+sintegraldisplay∞
0f(t)e−stdt
=−f(0) + s¯f(s),fors>0. (13.57)
The evaluation relies on integration by parts and higher-order derivatives may
be found in a similar manner.
461
INTEGRAL TRANSFORMSIFind the Laplace transform of d2f/dt2.
Using the definition of the Laplace transf orm and integrating by parts we obtainL
/d2f
dt2
/
=
Z∞
0d2f
dt2e−stdt
=
/df
dte−st
/∞
0+s
Z∞
0df
dte−stdt
=−df
dt(0) + s[s¯f(s)−f(0)],fors>0,
where (13.57) has been substituted for the integral. This can be written more neatly asL
/d2f
dt2
/
=s2¯f(s)−sf(0)−df
dt(0),fors>0.
J
In general the Laplace transform of the nth derivative is given byLbracketleftbiggdnf
dtnbracketrightbigg
=sn¯f−sn−1f(0)−sn−2df
dt(0)−···−dn−1f
dtn−1(0),fors>0.
(13.58)
We now turn to integration, which is much more straightforward. From the
definition (13.53),Lbracketleftbiggintegraldisplayt
0f(u)dubracketrightbigg
=integraldisplay∞
0dt e−stintegraldisplayt
0f(u)du
=bracketleftbigg
−1
se−stintegraldisplayt
0f(u)dubracketrightbigg∞
0+integraldisplay∞
01
se−stf(t)dt.
The first term on the RHS vanishes at both limits, and soLbracketleftbiggintegraldisplayt
0f(u)dubracketrightbigg
=1
s
L[f]. (13.59)
13.2.2 Other properties of Laplace transforms
From table 13.1 it will be apparent that multiplying a function f(t)b yeathas the
effect on its transform that sis replaced by s−a. This is easily proved generally:Lbracketleftbig
eatf(t)bracketrightbig
=integraldisplay∞
0f(t)eate−stdt
=integraldisplay∞
0f(t)e−(s−a)tdt
=¯f(s−a). (13.60)
As it were, multiplying f(t)b yeatmoves the origin of sby an amount a.
462
13.2 LAPLACE TRANSFORMS
We may now consider the effect of multiplying the Laplace transform ¯f(s)b y
e−bs(b>0). From the definition (13.53),
e−bs¯f(s)=integraldisplay∞
0e−s(t+b)f(t)dt
=integraldisplay∞
0e−szf(z−b)dz,
on putting t+b=z. Thus e−bs¯f(s) is the Laplace transform of a function g(t)
defined by
g(t)=braceleftBigg
0f o r 0 <t≤b,
f(t−b)f o r t>b .
In other words, the function fhas been translated to ‘later’ t(larger values of t)
by an amount b.
Further properties of Laplace transforms can be proved in similar ways and
are listed below.
(i) L[f(at)]=1
a¯fparenleftBigs
aparenrightBig
, (13.61)
(ii) L[tnf(t)]=(−1)ndn¯f(s)
dsn,forn=1,2,3,..., (13.62)
(iii) Lbracketleftbiggf(t)
tbracketrightbigg
=integraldisplay∞
s¯f(u)du, (13.63)
provided lim t→0[f(t)/t] exists.
Related results may be easily proved.IFind an expression for the Laplace transform of td2f/dt2.
From the definition of the Laplace transform we haveL
/
td2f
dt2
/
=
Z∞
0e−sttd2f
dt2dt
=−d
ds
Z∞
0e−std2f
dt2dt
=−d
ds[s2¯f(s)−sf(0)−f/prime(0)]
=−s2d¯f
ds−2s¯f+f(0).
J
Finally we mention the convolution theorem for Laplace transforms (which is
analogous to that for Fourier transforms discussed in subsection 13.1.7). If thefunctions fandghave Laplace transforms ¯f(s)a n d ¯g(s)t h e nLbracketleftbiggintegraldisplayt
0f(u)g(t−u)dubracketrightbigg
=¯f(s)¯g(s), (13.64)
463
INTEGRAL TRANSFORMS
t tt=u u=t
u u(a) (b)
Figure 13.7 Two representations of the Laplace transform convolution (see
text).
where the integral in the brackets on the LHS is the convolution offandg,
denoted by f∗g. As in the case of Fourier transforms, the convolution defined
above is commutative, i.e. f∗g=g∗f, and is associative and distributive. From
(13.64) we also see thatL−1bracketleftbig¯f(s)¯g(s)bracketrightbig
=integraldisplayt
0f(u)g(t−u)du=f∗g.IProve the convolution theorem (13.64) for Laplace transforms.
From the definition (13.64),
¯f(s)¯g(s)=
Z∞
0e−suf(u)du
Z∞
0e−svg(v)dv
=
Z∞
0du
Z∞
0dv e−s(u+v)f(u)g(v).
Now letting u+v=tchanges the limits on the integrals, with the result that
¯f(s)¯g(s)=
Z∞
0du f(u)
Z∞
udt g(t−u)e−st.
As shown in figure 13.7( a) the shaded area of integration may be considered as the sum
of vertical strips. However, we may instead integrate over this area by summing overhorizontal strips as shown in figure 13.7( b). Then the integral can be written as
¯f(s)¯g(s)=
Zt
0du f(u)
Z∞
0dt g(t−u)e−st
=
Z∞
0dt e−st
/Zt
0f(u)g(t−u)du
/
= L
/Zt
0f(u)g(t−u)du
/
.
J
464
13.3 CONCLUDING REMARKS
The properties of the Laplace transform derived in this section can sometimes
be useful in finding the Laplace transforms of particular functions.IFind the Laplace transform of f(t)=tsinbt.
Although we could calculate the Laplace transform directly, we can use (13.62) to give
¯f(s)=(−1)d
ds
L[sinbt]=−d
ds
/b
s2+b2
/
=2bs
(s2+b2)2,fors>0.
J
13.3 Concluding remarks
In this chapter we have discussed Fourier and Laplace transforms in some detail.
Both are examples of integral transforms , which can be considered in a more
general context.
A general integral transform of a function f(t)t a k e st h ef o r m
F(α)=integraldisplayb
aK(α, t)f(t)dt, (13.65)
where F(α)i st h et r a n s f o r mo f f(t) with respect to the kernel K(α, t), and αis
the transform variable. For example, in the Laplace transform case K(s, t)=e−st,
a=0 , b=∞.
Very often the inverse transform can also be written straightforwardly and
we obtain a transform pair similar to that encountered in Fourier transforms.Examples of such pairs are
(i) the Hankel transform
F(k)=integraldisplay
∞
0f(x)Jn(kx)xd x ,
f(x)=integraldisplay∞
0F(k)Jn(kx)kd k ,
where the Jnare Bessel functions of order n,a n d
(ii) the Mellin transform
F(z)=integraldisplay∞
0tz−1f(t)dt,
f(t)=1
2πiintegraldisplayi∞
−i∞t−zF(z)dz.
Although we do not have the space to discuss their general properties, the
reader should at least be aware of this wider class of integral transforms.
465
INTEGRAL TRANSFORMS
13.4 Exercises
13.1 Find the Fourier transform of the function f(t)=e x p (−|t|).
(a) By applying Fourier’s inversion theorem prove that
π
2exp(−|t|)=
Z∞
0cosωt
1+ω2dω.
(b) By making the substitution ω=t a n θ, demonstrate the validity of Parseval’s
theorem for this function.
13.2 Use the general definition and properties of Fourier transforms to show the
following.
(a) If f(x) is periodic with period athen˜f(k) = 0 unless ka=2πnfor integer n.
(b) The Fourier transform of tf(t)i sid˜f(ω)/dω.
(c) The Fourier transform of f(mt+c)i s
eiωc/m
m˜f
/ω
m
/
.
13.3 Find the Fourier transform of H(x−a)e−bx,w h e r e H(x) is the Heaviside function.
13.4 Prove that the Fourier transform of the function f(t) defined in the tf-plane by
straight-line segments joining ( −T,0) to (0 ,1) to ( T,0), with f(t) = 0 outside
|t|<T,i s
˜f(ω)=T√
2πsinc2
/ωT
2
/
,
where sinc xis defined as (sin x)/x.
Use the general properties of Fourier transforms to determine the transforms
of the following functions, graphically defined by straight-line segments and equalto zero outside the ranges specified:
(a) (0 ,0) to (0 .5,1) to (1 ,0) to (2 ,2) to (3 ,0) to (4 .5,3) to (6 ,0);
(b) (−2,0) to (−1,2) to (1 ,2) to (2 ,0);
(c) (0 ,0) to (0 ,1) to (1 ,2) to (1 ,0) to (2 ,−1) to (2 ,0).
13.5 By taking the Fourier transform of the equation
d
2φ
dx2−K2φ=f(x)
show that its solution φ(x) can be written as
φ(x)=−1√
2π
Z∞
−∞eikxef(k)
k2+K2dk,
where
ef(k) is the Fourier transform of f(x).
13.6 By differentiating the definition of the Fourier sine transform ˜fs(ω) of the function
f(t)=t−1/2with respect to ω, and then integrating the resulting expression by
parts, find an elementary differential equation satisfied by ˜fs(ω). Hence show that
this function is its own Fourier sine transform, i.e. ˜fs(ω)=Af(ω), where Ais a
constant. Show that it is also its own Fourier cosine transform. (Assume that thelimit as x→∞ofx
1/2sinαxcan be taken as zero.)
13.7 (a) Find the Fourier transform of the unit rectangular distribution
f(t)=
/(
1|t|<1
0 otherwise .
466
13.4 EXERCISES
(b) Determine the convolution of fwith itself and, without further integration,
deduce its transform.
(c) Deduce thatZ∞
−∞sin2ω
ω2dω=π,Z∞
−∞sin4ω
ω4dω=2π
3.
13.8 Calculate the Fraunhofer spectrum produced by a diffraction grating, uniformly
illuminated by light of wavelength 2 π/k, as follows. Consider a grating with 4 N
equal strips each of width aand alternately opaque and transparent. The aperture
function is then
f(y)=
/(
Afor (2 n+1 )a≤y≤(2n+2 )a,−N≤n<N ,
0 otherwise.
(a) Show, for diffraction at angle θto the normal to the grating, that the required
Fourier transform can be writtenef(q)=( 2 π)−1/2N−1X
r=−Nexp(−2iarq)
Z2a
aAexp(−iqu)du,
where q=ksinθ.
(b) Evaluate the integral and sum to show thatef(q)=( 2 π)−1/2exp(−iqa/2)Asin(2qaN)
qcos(qa/2),
and hence that the intensity distribution I(θ) in the spectrum is proportional
to
sin2(2qaN)
q2cos2(qa/2).
(c) For large values of N, the numerator in the above expression has very closely
spaced maxima and minima as a function of θand effectively takes its mean
value, 1 /2, giving a low-intensity background. Much more significant peaks
inI(θ) occur when θ= 0 or the cosine term in the denominator vanishes.
Show that the corresponding values of |
ef(q)|are
2aNA
(2π)1/2and4aNA
(2π)1/2(2m+1 )πwith mintegral .
Note that the constructive interference makes the maxima in I(θ)∝N2, not
N. Of course, observable maxima only occur for 0 ≤θ≤π/2.
13.9 By finding the complex Fourier series for its LHS show that either side of the
equation
∞X
n=−∞δ(t+nT)=1
T∞X
n=−∞e−2πnit/T
can represent a periodic train of impulses. By expressing the function f(t+nX),
in which Xis a constant, in terms of the Fourier transform ˜f(ω)o ff(t), show
that
∞X
n=−∞f(t+nX)=√
2π
X∞X
n=−∞˜f
/2nπ
X
/
e2πnit/X,
This result is known as the Poisson summation formula .
467
INTEGRAL TRANSFORMS
13.10 In many applications in which the frequency spectrum of an analogue signal is
required, the best that can be done is to sample the signal f(t) a finite number of
times at fixed intervals and then use a discrete Fourier transform Fkto estimate
discrete points on the (true) frequency spectrum ˜f(ω).
(a) By an argument that is essentially the converse of that given in section 13.1,
show that, if Nsamples fn, beginning at t= 0 and spaced τapart, are taken,
then˜f(2πk/(Nτ))≈Fkτwhere
Fk=1√
2πN−1X
n=0fne−2πnki/N.
(b) For the function f(t) defined by
f(t)=
/(
1f o r 0≤t<1
0 otherwise,
from which eight samples are drawn at intervals of τ=0.25, find a formula
for|Fk|and evaluate it for k=0,1,...,7.
(c) Find the exact frequency spectrum of f(t) and compare the actual and
estimated values of√
2π|˜f(ω)|atω=kπfork=0,1,...,7.Note the
relatively good agreement for k<4 and the lack of agreement for larger
values of k.
13.11 For a function f(t) that is non-zero only in the range |t|<T/2, the full frequency
spectrum ˜f(ω) can be constructed, in principle exactly, from values at discrete
sample points ω=n(2π/T) .P r o v et h i sa sf o l l o w s .
(a) Show that the coefficients of a complex Fourier series representation of f(t)
with period Tcan be written as
cn=√
2π
T˜f
/2πn
T
/
.
(b) Use this result to represent f(t) as an infinite sum in the defining integral for
˜f(ω), and hence show that
˜f(ω)=∞X
n=−∞˜f
/2πn
T
/
sinc
/
nπ−ωT
2
/
,
where sinc xis defined as (sin x)/x.
13.12 A signal obtained by sampling a function x(t) at regular intervals Tis passed
through an electronic filter, whose response g(t) to a unit δ-function input is
represented in a tg-plot by straight lines joining (0 ,0) to ( T,1/T)t o( 2 T,0) and
is zero for all other values of t. The output of the filter is the convolution of the
input,
P∞
−∞x(t)δ(t−nT), with g(t).
Using the convolution theorem, and the result given in exercise 13.4, show thatthe output of the filter can be written
y(t)=1
2π∞X
n=−∞x(nT)
Z∞
−∞sinc2
/ωT
2
/
e−iω[(n+1)T−t]dω.
13.13 (a) Find the Fourier transform of
f(γ,p,t)=
/(
e−γtsinpt t > 0
0 t<0,
where γ(>0) and pare constant parameters.
468
13.4 EXERCISES
(b) The current I(t) flowing through a certain system is related to the applied
voltage V(t) by the equation
I(t)=
Z∞
−∞K(t−u)V(u)du,
where
K(τ)=a1f(γ1,p1,τ)+a2f(γ2,p2,τ).
The function f(γ,p,t) is as given in (a) and all the ai,γi(>0) and piare fixed
parameters. By considering the Fourier transform of I(t), find the relationship
that must hold between a1anda2if the total net charge Qpassed through
the system (over a very long time) is to be zero for an arbitrary appliedvoltage.
13.14 Prove the equalityZ∞
0e−2atsin2at dt=1
π
Z∞
0a2
4a4+ω4dω.
13.15 A linear amplifier produces an output that is the convolution of its input and its
response function. The Fourier transform of the response function for a particularamplifier is
˜K(ω)=iω
√
2π(α+iω)2.
Determine the time variation of its output g(t) when its input is the Heaviside
step function. (Consider the Fourier transform of a decaying exponential function
and the result of exercise 13.2(b).)
13.16 In quantum mechanics, two equal-mass particles having momenta pj= /~kjand
energies Ej= /~ωjand represented by plane wavefunctions φj=e x p [ i(kj·rj−ωjt)],
j=1,2, interact through a potential V=V(|r1−r2|). In first-order perturbation
theory the probability of scattering to a state with momenta and energies p/prime
j,E/prime
j
is determined by the modulus squared of the quantity
M=
ZZ Z
ψ∗
fVψ idr1dr2dt.
The initial state ψiisφ1φ2and the final state ψfisφ/prime
1φ/prime2.
(a) By writing r1+r2=2Randr1−r2=rand assuming that dr1dr2=dRdr,
show that Mcan be written as the product of three one-dimensional integrals.
(b) From two of the integrals deduce energy and momentum conservation in the
form of δ-functions.
(c) Show that Mis proportional to the Fourier transform of V,i . e .
eV(k)w h e r e
2 /~k=(p2−p1)−(p/prime
2−p/prime
1).
13.17 For some ion–atom scattering processes, the potential Vof the previous example
may be approximated by V=|r1−r2|−1exp(−µ|r1−r2|). Show, using the result
of the worked example in subsection 13.1.10, that the probability that the ionwill scatter from, say, p
1top/prime
1is proportional to ( µ2+k2)−2where k=|k|andk
is as given in part (c) of exercise 13.16.
13.18 The equivalent duration and bandwidth, TeandBe, of a signal x(t) are defined
in terms of the latter and its Fourier transform ˜x(ω):
Te=1
x(0)
Z∞
−∞x(t)dt,
Be=1
˜x(0)
Z∞
−∞˜x(ω)dω,
469
INTEGRAL TRANSFORMS
where neither x(0) nor ˜x(0) is zero. Show that the product TeBe=2π(this is a
form of uncertainty principle), and find the equivalent bandwidth of the signal
x(t)=e x p (−|t|/T).
For this signal, determine the fraction of the total energy that lies in the frequency
range|ω|<B e/4. You will need the indefinite integral with respect to xof
(a2+x2)−2,w h i c hi s
x
2a2(a2+x2)+1
2a3tan−1x
a.
13.19 Calculate directly the auto-correlation function a(z) for the product of the expo-
nential decay distribution and the Heaviside step function
f(t)=1
λe−λtH(t).
Use the Fourier transform and energy spectrum of f(t) to deduce thatZ∞
−∞eiωz
λ2+ω2dω=π
λe−λ|z|.
13.20 Prove that the cross-correlation C(z) of the Gaussian and Lorentzian distributions
f(t)=1
τ√
2πexp
/
−t2
2τ2
/
,g (t)=
/a
π
/1
t2+a2,
has as its Fourier transform the function
1√
2πexp
/
−τ2ω2
2
/
exp(−a|ω|).
Hence show that
C(z)=1
τ√
2πexp
/a2−z2
2τ2
/
cos
/az
τ2
/
.
13.21 Prove the expressions given in table 13.1 for the Laplace transforms of t−1/2and
t1/2, by setting x2=tsin the resultZ∞
0exp(−x2)dx=1
2√π.
13.22 Find the functions y(t) whose Laplace transforms are the following,
(a) 1 /(s2−s−2),
(b) 2 s/[(s+1 ) ( s2+4 ) ] ,
(c)e−(γ+s)t0/[(s+γ)2+b2].
13.23 Use the properties of Laplace transforms to prove the following without evaluat-
ing any Laplace integrals explicitly:
(a) L
/
t5/2
/
=15
8√πs−7/2.
(b) L
/
(sinh at)/t
/
=1
2ln
/
(s+a)/(s−a)
/
,s >|a|.
(c) L[sinhatcosbt]=a(s2−a2+b2)[(s−a)2+b2]−1[(s+a)2+b2]−1.
13.24 Find the solution (the so-called impulse response orGreen’s function )o ft h e
equation
Tdx
dt+x=δ(t)
by proceeding as follows.
470
13.4 EXERCISES
(a) Show by substitution that
x(t)=A(1−e−t/T)H(t)
is a solution, for which x(0) = 0, of
Tdx
dt+x=AH(t), (*)
where H(t) is the Heaviside step function.
(b) Construct the solution when the RHS of (*) is replaced by AH(t−τ)w i t h
dx/dt =x=0f o r t<τ, and hence find the solution when the RHS is a
rectangular pulse of duration τ.
(c) By setting A=1/τand taking the limit when τ→0, show that the impulse
response is x(t)=T−1e−t/T.
(d) Obtain the same result much more directly by taking the Laplace transform
of each term in the original equation, solving the resulting algebraic equationand then using the entries in table 13.1.
13.25 (a) If f(t)=A+g(t), where Ais a constant and the indefinite integral of g(t)i s
bounded as its upper limit tends to ∞, show that
lim
s→0s¯f(s)=A.
(b) For t>0 the function y(t) obeys the differential equation
d2y
dt2+ady
dt+by=ccos2ωt,
where a,bandcare positive constants. Find ¯y(s)a n ds h o wt h a t s¯y(s)→c/2b
ass→0. Interpret the result in the t-domain.
13.26 By writing f(x) as an integral involving the δ-function δ(ξ−x) and taking the
Laplace transforms of both sides, show that the transform of the solution of the
equation
d4y
dx4−y=f(x)
for which yand its first three derivatives vanish at x= 0 can be written as
¯y(s)=
Z∞
0f(ξ)e−sξ
s4−1dξ.
Use the properties of Laplace transforms and the entries in table 13.1 to show
that
y(x)=1
2
Zx
0f(ξ)[sinh(x−ξ)−sin(x−ξ)]dξ.
13.27 The function fa(x) is defined as unity for 0 <x<a and zero otherwise. Find its
Laplace transform ¯fa(s) and deduce that the transform of xfa(x)i s
1
s2
/
1−(1 +as)e−sa
/
.
Write fa(x) in terms of Heaviside functions and hence obtain an explicit expres-
sion for
ga(x)=
Zx
0fa(y)fa(x−y)dy.
Use the expression to write ¯ga(s) in terms of the functions ¯fa(s), and ¯f2a(s)a n d
their derivatives, and hence show that ¯ga(s) is equal to the square of ¯fa(s), in
accordance with the convolution theorem.
471
INTEGRAL TRANSFORMS
13.28 (a) Show that the Laplace transform of f(t−a)H(t−a), where a≥0, ise−as¯f(s).
(b) If g(t) is a periodic function of period T, show that ¯g(s) can be written as
1
1−e−sT
ZT
0e−stg(t)dt.
(c) Sketch the periodic function defined in 0 ≤t≤Tby
g(t)=
/(
2t/T 0≤t<T/ 2
2(1−t/T)T/2≤t≤T,
and, using the result in (b), find its Laplace transform.
(d) Show, by sketching it, that
2
T[tH(t)+2∞X
n=1(−1)n(t−1
2nT)H(t−1
2nT)]
is another representation of g(t) and hence derive the relationship
tanh x=1+2∞X
n=1(−1)ne−2nx.
13.5 Hints and answers
13.1 (2 /π)1/2(1 +ω2)−1.
13.2 (a) Show ˜f(k)(1−e±ika)=0 .
13.3 (1 /√
2π)[(b−ik)/(b2+k2)]e−a(b+ik).
13.4 (a) [8 /(√
2πω2)][e−iω/2sin2(ω/4) +e−i2ωsin2(ω/2) +e−i9ω/2sin2(3ω/4)].
(b) Consider the superposition of a ‘triangle’ of height 2 with T=2a n dt w o
triangles, each of unit height with T= 1, displaced by ±1; [8 sin2(ω/2) (1 +
2cos ω)]/(√
2πω2).
(c) Consider the superposition of a triangle and its derivative.
[(1 + iω)e−iω/√
2π]sinc2(ω/2).
13.6 d˜fs(ω)/dω=−˜fs(ω)/(2ω).
13.7 (a) (2 /√
2π)(sinω/ω).
(b) 2−|t|for|t|<2, zero otherwise. Use convolution theorem; (4 /√
2π)(sin2ω/ω2).
(c) Apply Parseval’s theorem to fand to f∗f.
13.8 (c) Use l’H ˆopital’s rule to evaluate the expressions of the form 0 /0.
13.10 (b) |Fk|=c o s e c( kπ/8)/√
2πforkodd;|Fk|=0f o r keven, except |F0|=4/√
2π.
(c)√
2π˜f(ω)=e−iω/2[sin(ω/2)/(ω/2)]. Actual (estimated) values at ω=kπfor
k=0,1,...,7:
1( 1 ) ;0 .637 (0 .653); 0 (0); 0 .212 (0 .271); 0 (0); 0 .127 (0 .271); 0 (0); 0 .091 (0 .653).
13.11 (b) Recall that the infinite integral involved in defining ˜f(ω)o n l yh a san o n - z e r o
integrand in |t|<T/2.
13.12 The Fourier transform of g(t) is found by moving the time origin by Tand then
applying (13.31). It is (1 /√
2π)si nc2(ωT/2)e−iωT.
13.13 (a) (1 /√
2π){p/[(γ+iω)2+p2]}.
(b) Show that Q=√
2π˜I(0) and use the convolution theorem. The required
relationship is a1p1/(γ2
1+p2
1)+a2p2/(γ2
2+p2
2)=0 .
13.14 Set p=γ=ain part (a) of exercise 13.13 and then apply Parseval’s theorem.
13.15 ˜g(ω)=1 /[√
2π(α+iω)2], leading to g(t)=te−αt.
13.16 (b) The t-integral is
R
exp[i(E/prime
1+E/prime
2−E1−E2)]dt∝δ(E/prime
1+E/prime
2−(E1+E2));
similarly the R-integral yields δ(p/prime
1+p/prime
2−(p1+p2)).
472
13.5 HINTS AND ANSWERS
13.17
eV(k)∝[−2π/(ik)]
R
{exp[−(µ−ik)r]−exp[−(µ+ik)r]}dr.
13.18 By setting t=0a n d ω= 0 in the Fourier definitions, obtain two equations
connecting x(0) and ˜x(0).Be=π/T;˜x(ω), proportional to the Fourier cosine
transform of exp( −t/T), is equal to [2 T/√
2π](1+ ω2T2)−1. The energy spectrum
is proportional to |˜x(ω)|2. Fraction = 0.733.
13.19 Note that the lower limit in the calculation of a(z)i s0f o r z>0a n d|z|for
z<0. Auto-correlation a(z)=[ ( 1 /(2λ3)]exp(−λ|z|).
13.20 Use the result of exercise 13.18 to deduce that ˜g(ω)=( 1 /√
2π)e x p (−a|ω|). Apply
the Wiener–Kinchin theorem. Note that, because of the presence of |ω|,t h e
inverse transform giving C(z) is a cosine transform.
13.21 Prove the result for t1/2by integrating that for t−1/2by parts.
13.22 (a) y(t)=1
3(e2t−e−t).
(b)y(t)=1
5(4 sin2 t+2c o s2 t−2e−t).
(c) Note the factor e−st0and write y(t) as a function of ( t−t0);y(t)=
b−1e−γtsinb(t−t0)H(t−t0).
13.23 (a) Use (13.62) with n=2o n L
/√
t
/
; (b) use (13.63);
(c) consider L[exp(±at)cosbt]and use the translation property, subsection 13.2.2.
13.24 (b) Superimpose solutions with equal amplitudes but opposite signs. x(t)=
A(1−e−t/T)H(t)−A(1−e−(t−τ)/T)H(t−τ).
(c) Write e−(t−τ)/Tase−t/T[1 +τ/T+O(τ2)] and note that, with 0 <t<τ ,τ (1−
e−t/T)/τ→0a sτ→0.
(d) The algebraic equation is ¯x=( 1+ sT)−1.
13.25 (a) Note that |lim
R
g(t)e−stdt|≤|lim
R
g(t)dt|.
(b) (s2+as+b)¯y(s)={c(s2+2ω2)/[s(s2+4ω2)]}+(a+s)y(0) + y/prime(0).
For this damped system, at large t(corresponding to s→0) rates of change
are negligible and the equation reduces to by=ccos2ωt,w i t hc o s2ωthaving an
average value1
2.
13.26 Factorise ( s4−1)−1as1
2[(s2−1)−1−(s2+1 )−1].
13.27 s−1[1−exp(−sa)];ga(x)=xfor 0 <x<a ,ga(x)=2 a−xfora≤x≤2a,
ga(x) = 0 otherwise.
13.28 (a) Note that
R∞
TH(t)···=
R∞
0H(t−T)···and that H(t−T)g(t)=H(t−
T)g(t−T).
(c)¯g(s)=[ 2 /(Ts2)] tanh( sT/4).
(d) Use the result from (a) and L[tH(t)] =s−2;s e t sT=4x.
473
14
First-order ordinary differential
equations
Differential equations are the group of equations that contain derivatives. Chap-
ters 14–19 discuss a variety of differential equations, starting in this chapter andthe next with those ordinary differential equations (ODEs) that have closed-formsolutions. As its name suggests, an ODE contains only ordinary derivatives (andnot partial derivatives) and describes the relationship between these derivatives ofthedependent variable , usually called y, with respect to the independent variable ,
usually called x. The solution to such an ODE is therefore a function of xand
is written y(x). For an ODE to have a closed-form solution, it must be possible
to express y(x) in terms of the standard elementary functions such as exp x,l nx,
sinxetc. The solutions of some differential equations cannot, however, be written
in closed form, but only as an infinite series; these are discussed in chapter 16.
Ordinary differential equations may be separated conveniently into differ-
ent categories according to their general characteristics. The primary groupingadopted here is by the order of the equation. The order of an ODE is simply the
order of the highest derivative it contains. Thus equations containing dy/dx , but
no higher derivatives, are called first order, those containing d
2y/dx2are called
second order and so on. In this chapter we consider first-order equations, and inthe next, second- and higher-order equations.
Ordinary differential equations may be classified further according to degree .
The degree of an ODE is the power to which the highest-order derivative israised, after the equation has been rationalised to contain only integer powers ofderivatives. Hence the ODE
d
3y
dx3+xparenleftbiggdy
dxparenrightbigg3/2
+x2y=0,
is of third order and second degree, since after rationalisation it contains the term
(d3y/dx3)2.
Thegeneral solution to an ODE is the most general function y(x) that satisfies
the equation; it will contain constants of integration which may be determined by
474
14.1 GENERAL FORM OF SOLUTION
the application of some suitable boundary conditions . For example, we may be
told that for a certain first-order differential equation, the solution y(x)i se q u a lt o
zero when the parameter xis equal to unity; this allows us to determine the value
of the constant of integration. The general solutions tonth-order ODEs, which
are considered in detail in the next chapter, will contain n(essential) arbitrary
constants of integration and therefore we will need nboundary conditions if these
constants are to be determined (see section 14.1). When the boundary conditionshave been applied, and the constants found, we are left with a particular solution
to the ODE, which obeys the given boundary conditions. Some ODEs of degreegreater than unity also possess singular solutions , which are solutions that contain
no arbitrary constants and cannot be found from the general solution; singular
solutions are discussed in more detail in section 14.3. When any solution to an
ODE has been found, it is always possible to check its validity by substitutioninto the original equation and verification that any given boundary conditionsare met.
In this chapter, firstly we discuss various types of first-degree ODE, and then
go on to examine those higher-degree equations that can be solved in closed form.At the outset, however, we discuss the general form of the solutions of ODEs;
this discussion is relevant to both first- and higher-order ODEs.
14.1 General form of solution
It is helpful when considering the general form of the solution of an ODE to
consider the inverse process, namely that of obtaining an ODE from a givengroup of functions, each one of which is a solution of the ODE. Suppose the
members of the group can be written as
y=f(x, a
1,a2,...,a n), (14.1)
each member being specified by a different set of values of the parameters ai.F o r
example, consider the group of functions
y=a1sinx+a2cosx; (14.2)
here n=2 .
Since an ODE is required for which anyof the group is a solution, it clearly
must not contain any of the ai.A st h e r ea r e nof the aiin expression (14.1), we
must obtain n+ 1 equations involving them in order that, by elimination, we can
obtain one final equation without them.
Initially we have only (14.1), but if this is differentiated ntimes, a total of n+1
equations is obtained from which (in principle) all the aican be eliminated, to
give one ODE satisfied by all the group. As a result of the ndifferentiations,
dny/dxnwill be present in one of the n+ 1 equations and hence in the final
equation, which will therefore be of nth order.
475
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
In the case of (14.2), we have
dy
dx=a1cosx−a2sinx,
d2y
dx2=−a1sinx−a2cosx.
Here the elimination of a1anda2is trivial (because of the similarity of the forms
ofyandd2y/dx2), resulting in
d2y
dx2+y=0,
a second-order equation.
Thus, to summarise, a group of functions (14.1) with nparameters satisfies an
nth-order ODE in general (although in some degenerate cases an ODE of less
than nth order is obtained). The intuitive converse of this is that the general
solution of an nth-order ODE contains narbitrary parameters (constants); for
our purposes, this will be assumed to be valid although a totally general proof isdifficult.
As mentioned earlier, external factors affect a system described by an ODE,
by fixing the values of the dependent variables for particular values of theindependent ones. These externally imposed (or boundary ) conditions on the
solution are thus the means of determining the parameters and so of specifying
precisely which function is the required solution. It is apparent that the numberof boundary conditions should match the number of parameters and hence theorder of the equation, if a unique solution is to be obtained. Fewer independentboundary conditions than this will lead to a number of undetermined parametersin the solution, whilst an excess will usually mean that no acceptable solution ispossible.
For an nth-order equation the required nboundary conditions can take many
forms, for example the value of yatndifferent values of x, or the value of any
n−1o ft h e nderivatives dy/dx ,d
2y/dx2,...,dny/dxntogether with that of y,a l l
for the same value of x, or many intermediate combinations.
14.2 First-degree first-order equations
First-degree first-order ODEs contain only dy/dx equated to some function of x
andy, and can be written in either of two equivalent standard forms,
dy
dx=F(x, y),A(x, y)dx+B(x, y)dy=0,
where F(x, y)=−A(x, y)/B(x, y), and F(x, y),A(x, y)a n d B(x, y) are in general
functions of both xandy. Which of the two above forms is the more useful
for finding a solution depends on the type of equation being considered. There
476
14.2 FIRST-DEGREE FIRST-ORDER EQUATIONS
are several different types of first-degree first-order ODEs that are of interest in
the physical sciences. These equations and their respective solutions are discussedbelow.
14.2.1 Separable-variable equations
A separable-variable equation is one which may be written in the conventional
form
dy
dx=f(x)g(y), (14.3)
where f(x)a n d g(y) are functions of xandyrespectively, including cases in
which f(x)o rg(y) is simply a constant. Rearranging this equation so that the
terms depending on xand on yappear on opposite sides (i.e. are separated), and
integrating, we obtain
integraldisplaydy
g(y)=integraldisplay
f(x)dx.
Finding the solution y(x) that satisfies (14.3) then depends only on the ease with
which the integrals in the above equation can be evaluated. It is also worth
noting that ODEs that at first sight do not appear to be of the form (14.3) can
sometimes be made separable by an appropriate factorisation.ISolve
dy
dx=x+xy.
Since the RHS of this equation can be factorised to give x(1 +y), the equation becomes
separable and we obtainZdy
1+y=
Z
xd x .
Now integrating both sides separately, we find
ln(1 + y)=x2
2+c,
and so
1+y=e x p
/x2
2+c
/
=Aexp
/x2
2
/
,
where cand hence Ais an arbitrary constant.
J
Solution method. Factorise the equation so that it becomes separable. After rear-
ranging it so that the terms depending on xand those depending on yappear on
opposite sides, integrate directly. Remember the constant of integration, which canbe evaluated if further information is given.
477
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
14.2.2 Exact equations
Anexact first-degree first-order ODE is one of the form
A(x, y)dx+B(x, y)dy= 0 and for which∂A
∂y=∂B
∂x. (14.4)
In this case A(x, y)dx+B(x, y)dyis an exact differential, dU(x, y)s a y( s e e
section 5.3). In other words
Ad x+Bd y=dU=∂U
∂xdx+∂U
∂ydy,
from which we obtain
A(x, y)=∂U
∂x, (14.5)
B(x, y)=∂U
∂y. (14.6)
Since ∂2U/∂x∂y =∂2U/∂y∂x we therefore require
∂A
∂y=∂B
∂x. (14.7)
If (14.7) holds then (14.4) can be written dU(x, y) = 0, which has the solution
U(x, y)=c,w h e r e cis a constant and from (14.5) U(x, y)i sg i v e nb y
U(x, y)=integraldisplay
A(x, y)dx+F(y). (14.8)
The function F(y) can be found from (14.6) by differentiating (14.8) with respect
toyand equating to B(x, y).ISolve
xdy
dx+3x+y=0.
Rearranging into the form (14.4) we have
(3x+y)dx+xd y=0,
i.e.A(x, y)=3 x+yandB(x, y)=x.S i n c e ∂A/∂y =1= ∂B/∂x , the equation is exact, and
by (14.8) the solution is given by
U(x, y)=
Z
(3x+y)dx+F(y)=c1⇒3x2
2+yx+F(y)=c1.
Differentiating U(x, y) with respect to yand equating it to B(x, y)=xwe obtain dF/dy =0 ,
which integrates immediately to give F(y)=c2. Therefore, letting c=c1−c2,t h es o l u t i o n
to the original ODE is
3x2
2+xy=c.
J
478
14.2 FIRST-DEGREE FIRST-ORDER EQUATIONS
Solution method. Check that the equation is an exact differential using (14.7) then
solve using (14.8). Find the function F(y)by differentiating (14.8) with respect to
yand using (14.6).
14.2.3 Inexact equations: integrating factors
Equations that may be written in the form
A(x, y)dx+B(x, y)dy= 0 but for which∂A
∂y/negationslash=∂B
∂x(14.9)
are known as inexact equations. However, the differential Ad x+Bd ycan always
be made exact by multiplying by an integrating factor µ(x, y), which obeys
∂(µA)
∂y=∂(µB)
∂x. (14.10)
For an integrating factor that is a function of both xandy,i . e .µ=µ(x, y), there
exists no general method for finding it; in such cases it may sometimes be foundby inspection. If, however, an integrating factor exists that is a function of either
xoryalone then (14.10) can be solved to find it. For example, if we assume
that the integrating factor is a function of xalone, i.e. µ=µ(x), then (14.10)
reads
µ∂A
∂y=µ∂B
∂x+Bdµ
dx.
Rearranging this expression we find
dµ
µ=1
Bparenleftbigg∂A
∂y−∂B
∂xparenrightbigg
dx=f(x)dx,
where we require f(x) also to be a function of xonly; indeed this provides a
general method of determining whether the integrating factor µis a function of
xalone. This integrating factor is then given by
µ(x)=e x pbraceleftbiggintegraldisplay
f(x)dxbracerightbigg
where f(x)=1
Bparenleftbigg∂A
∂y−∂B
∂xparenrightbigg
. (14.11)
Similarly, if µ=µ(y)t h e n
µ(y)=e x pbraceleftbiggintegraldisplay
g(y)dybracerightbigg
where g(y)=1
Aparenleftbigg∂B
∂x−∂A
∂yparenrightbigg
. (14.12)
479
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
dy
dx=−2
y−3y
2x.
Rearranging into the form (14.9), we have
(4x+3y2)dx+2xy dy=0, (14.13)
i.e.A(x, y)=4 x+3y2andB(x, y)=2 xy.N o w
∂A
∂y=6y,∂B
∂x=2y,
so the ODE is not exact in its present form. However, we see that
1
B
/∂A
∂y−∂B
∂x
/
=2
x,
a function of xalone. Therefore an integrating factor exists that is also a function of x
alone and, ignoring the arbitrary constant of integration, is given by
µ(x)=e x p
/
2
Zdx
x
/
=e x p ( 2l n x)=x2.
Multiplying (14.13) through by µ(x)=x2we obtain
(4x3+3x2y2)dx+2x3yd y=4x3dx+( 3x2y2dx+2x3yd y)=0 .
By inspection this integrates immediately to give the solution x4+y2x3=c,w h e r e cis a
constant.
J
Solution method. Examine whether f(x)andg(y)are functions of only xory
respectively. If so, then the required integrating factor is a function of either xor
yonly, and is given by (14.11) or (14.12) respectively. If the integrating factor is
a function of both xandy, then sometimes it may be found by inspection or by
trial and error. In any case, the integrating factor µmust satisfy (14.10). Once the
equation has been made exact, solve by the method of subsection 14.2.2.
14.2.4 Linear equations
Linear first-order ODEs are a special case of inexact ODEs (discussed in the
previous subsection) and can be written in the conventional form
dy
dx+P(x)y=Q(x). (14.14)
Such equations can be made exact by multiplying through by an appropriate
integrating factor in a similar manner to that discussed above. In this case,
however, the integrating factor is always a function of xalone and may be
expressed in a particularly simple form. An integrating factor µ(x) must be such
that
µ(x)dy
dx+µ(x)P(x)y=d
dx[µ(x)y]=µ(x)Q(x), (14.15)
480
14.2 FIRST-DEGREE FIRST-ORDER EQUATIONS
which may then be integrated directly to give
µ(x)y=integraldisplay
µ(x)Q(x)dx. (14.16)
The required integrating factor µ(x) is determined by the first equality in (14.15),
i.e.
d
dx(µy)=µdy
dx+dµ
dxy=µdy
dx+µPy,
which immediately gives the simple relation
dµ
dx=µ(x)P(x)⇒ µ(x)=e x pbraceleftbiggintegraldisplay
P(x)dxbracerightbigg
. (14.17)ISolve
dy
dx+2xy=4x.
The integrating factor is given immediately by
µ(x)=e x p
/Z
2xd x
/
=e x p x2.
Multiplying through the ODE by µ(x)=e x p x2and integrating, we have
yexpx2=4
Z
xexpx2dx=2e x p x2+c.
The solution to the ODE is therefore given by y=2+ cexp(−x2).
J
Solution method. Rearrange the equation into the form (14.14) and multiply by the
integrating factor µ(x)given by (14.17). The left- and right-hand sides can then be
integrated directly, giving y from (14.16).
14.2.5 Homogeneous equations
Homogeneous equation are ODEs that may be written in the form
dy
dx=A(x, y)
B(x, y)=FparenleftBigy
xparenrightBig
, (14.18)
where A(x, y)a n d B(x, y) are homogeneous functions of the same degree. A
function f(x, y) is homogeneous of degree nif, for any λ,i to b e y s
f(λx, λy)=λnf(x, y).
For example, if A=x2y−xy2andB=x3+y3then we see that AandBare
both homogeneous functions of degree 3. In general, for functions of the form of
AandB, we see that for both to be homogeneous, and of the same degree, we
require the sum of the powers in xandyin each term of AandBto be the same
481
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
(in this example equal to 3). The RHS of a homogeneous ODE can be written as
a function of y/x. The equation may then be solved by making the substitution
y=vx,s ot h a t
dy
dx=v+xdv
dx=F(v).
This is now a separable equation and can be integrated directly to give
integraldisplaydv
F(v)−v=integraldisplaydx
x. (14.19)ISolve
dy
dx=y
x+t a n
/y
x
/
.
Substituting y=vxwe obtain
v+xdv
dx=v+t a n v.
Cancelling von both sides, rearranging and integrating givesZ
cotvd v=
Zdx
x=l nx+c1.
ButZ
cotvd v=
Zcosv
sinvdv=l n ( s i n v)+c2,
so the solution to the ODE is y=xsin−1Ax,w h e r e Ais a constant.
J
Solution method. Check to see whether the equation is homogeneous. If so, make
the substitution y=vx, separate variables as in (14.19) and then integrate directly.
Finally replace vbyy/xto obtain the solution.
14.2.6 Isobaric equations
An isobaric ODE is a generalisation of the homogeneous ODE discussed in the
previous section, and is of the form
dy
dx=A(x, y)
B(x, y), (14.20)
where the equation is dimensionally consistent if yanddyare each given a weight
mrelative to xanddx, i.e. if the substitution y=vxmmakes it separable.
482
14.2 FIRST-DEGREE FIRST-ORDER EQUATIONSISolve
dy
dx=−1
2yx
/
y2+2
x
/
.
Rearranging we have/
y2+2
x
/
dx+2yx dy =0.
Giving yanddythe weight mandxanddxthe weight 1, the sums of the powers in each
term on the LHS are 2 m+1 ,0a n d2 m+ 1 respectively. These are equal if 2 m+1=0 ,i . e .
ifm=−1
2. Substituting y=vxm=vx−1/2, with the result that dy=x−1/2dv−1
2vx−3/2dx,
we obtain
vd v+dx
x=0,
which is separable and may be integrated directly to give1
2v2+l nx=c. Replacing vby
y√xwe obtain the solution1
2y2x+l nx=c.
J
Solution method. Write the equation in the form Ad x+Bd y=0. Giving yand
dyeach a weight mandxanddxeach a weight 1, write down the sum of powers
in each term. Then, if a value of mthat makes all these sums equal can be found,
substitute y=vxminto the original equation to make it separable. Integrate the
separated equation directly, and then replace vbyyx−mto obtain the solution.
14.2.7 Bernoulli’s equation
Bernoulli’s equation has the form
dy
dx+P(x)y=Q(x)ynwhere n/negationslash=0o r1 . (14.21)
This equation is very similar in form to the linear equation (14.14), but is in fact
non-linear due to the extra ynfactor on the RHS. However, the equation can be
made linear by substituting v=y1−n, and correspondingly
dy
dx=parenleftbiggyn
1−nparenrightbiggdv
dx.
Substituting this into (14.21) and dividing through by yn, we find
dv
dx+( 1−n)P(x)v=( 1−n)Q(x),
which is a linear equation, and may be solved by the method described in
subsection 14.2.4.
483
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
dy
dx+y
x=2x3y4.
If we let v=y1−4=y−3then
dy
dx=−y4
3dv
dx.
Substituting this into the ODE and rearranging, we obtain
dv
dx−3v
x=−6x3,
which is linear and may be solved by multiplying through by the integrating factor (see
subsection 14.2.4)
exp
/
−3
Zdx
x
/
=e x p (−3lnx)=1
x3.
This yields the solution
v
x3=−6x+c.
Remembering that v=y−3,w eo b t a i n y−3=−6x4+cx3.
J
Solution method. Rearrange the equation into the form (14.21) and make the sub-
stitution v=y1−n. This leads to a linear equation in v, which can be solved by the
method of subsection 14.2.4. Then replace vbyy1−nto obtain the solution.
14.2.8 Miscellaneous equations
There are two further types of first-degree first-order equation that occur fairly
regularly but do not fall into any of the above categories. They may be reducedto one of the above equations, however, by a suitable change of variable.
Firstly, we consider
dy
dx=F(ax+by+c), (14.22)
where a,bandcare constants, i.e. xandyonlyappear on the RHS in the particular
combination ax+by+cand not in any other combination or by themselves. This
equation can be solved by making the substitution v=ax+by+c,i nw h i c hc a s e
dv
dx=a+bdy
dx=a+bF(v), (14.23)
which is separable and may be integrated directly.
484
14.2 FIRST-DEGREE FIRST-ORDER EQUATIONSISolve
dy
dx=(x+y+1 )2.
Making the substitution v=x+y+ 1, we obtain, as in (14.23),
dv
dx=v2+1,
which is separable and integrates directly to giveZdv
1+v2=
Z
dx⇒tan−1v=x+c1.
So the solution to the original ODE is tan−1(x+y+1 )= x+c1,w h e r e c1is a constant of
integration.
J
Solution method. In an equation such as (14.22), substitute v=ax+by+cto obtain
a separable equation that can be integrated directly. Then replace vbyax+by+c
to obtain the solution.
Secondly, we discuss
dy
dx=ax+by+c
ex+fy+g, (14.24)
where a,b,c,e,fandgare all constants. This equation may be solved by letting
x=X+αandy=Y+β,w h e r e αandβare constants found from
aα+bβ+c= 0 (14.25)
eα+fβ+g=0. (14.26)
Then (14.24) can be written as
dY
dX=aX+bY
eX+fY,
which is homogeneous and can be solved by the method of subsection 14.2.5.
Note, however, that if a/e=b/fthen (14.25) and (14.26) are not independent
and so cannot be solved uniquely for αandβ. However, in this case, (14.24)
reduces to an equation of the form (14.22), which was discussed above.ISolve
dy
dx=2x−5y+3
2x+4y−6.
Letx=X+αandy=Y+β,w h e r e αandβobey the relations
2α−5β+3=0
2α+4β−6=0 ,
which solve to give α=β= 1. Making these substitutions we find
dY
dX=2X−5Y
2X+4Y,
485
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
which is a homogeneous ODE and can be solved by substituting Y=vX(see subsec-
tion 14.2.5) to obtain
dv
dX=2−7v−4v2
X(2 + 4 v).
This equation is separable, and using partial fractions we findZ2+4 v
2−7v−4v2dv=−4
3
Zdv
4v−1−2
3
Zdv
v+2=
ZdX
X,
which integrates to give
lnX+1
3ln(4v−1) +2
3ln(v+2 )= c1,
or
X3(4v−1)(v+2 )2=e x p ( 3 c1).
Remembering that Y=vX,x=X+1a n d y=Y+ 1, the solution to the original ODE
is given by (4 y−x−3)(y+2x−3)2=c2,w h e r e c2=e x p ( 3 c1).
J
Solution method. If in (14.24) a/e/negationslash=b/fthen make the substitution x=X+α,
y=Y+β,w h e r e αandβare given by (14.25) and (14.26); the resulting equation
is homogeneous and can be solved as in subsection 14.2.5. Substitute v=Y/ X,
X=x−α, and Y=y−βto obtain the solution. If a/e=b/fthen (14.24) is of
the same form as (14.22) and may be solved accordingly.
14.3 Higher-degree first-order equations
First-order equations of degree higher than the first do not occur often in
the description of physical systems, since squared and higher powers of first-
order derivatives usually arise from resistive or driving mechanisms, when anacceleration or other higher-order derivative is also present. They do sometimesappear in connection with geometrical problems, however.
Higher-degree first-order equations can be written as F(x, y, dy/dx )=0 .T h e
most general standard form is
p
n+an−1(x, y)pn−1+···+a1(x, y)p+a0(x, y)=0 , (14.27)
where for ease of notation we write p=dy/dx . If the equation can be solved for
one of x,yorpthen either an explicit or a parametric solution can sometimes be
obtained. We discuss the main types of such equations below, including Clairaut’sequation, which is a special case of an equation explicitly soluble for y.
14.3.1 Equations soluble for p
Sometimes the LHS of (14.27) can be factorised into the form
(p−F
1)(p−F2)···(p−Fn)=0 , (14.28)
486
14.3 HIGHER-DEGREE FIRST-ORDER EQUATIONS
where Fi=Fi(x, y). We are then left with solving the nfirst-degree equations
p=Fi(x, y). Writing the solutions to these first-degree equations as Gi(x, y)=0 ,
the general solution to (14.28) is given by the product
G1(x, y)G2(x, y)···Gn(x, y)=0 . (14.29)ISolve
(x3+x2+x+1 )p2−(3x2+2x+1 )yp+2xy2=0. (14.30)
This equation may be factorised to give
[(x+1 )p−y][(x2+1 )p−2xy]=0 .
Taking each bracket in turn we have
(x+1 )dy
dx−y=0,
(x2+1 )dy
dx−2xy=0,
which have the solutions y−c(x+1 ) = 0 a n d y−c(x2+ 1) = 0 respectively (see
section 14.2 on first-degree first-order equations). Note that the arbitrary constants inthese two solutions can be taken to be the same, since only one is required for a first-orderequation. The general solution to (14.30) is then given by
[y−c(x+1 )]/
y−c(x2+1 )
/
=0.
J
Solution method. If the equation can be factorised into the form (14.28) then solve
the first-order ODE p−Fi=0in each factor and write the solution in the form
Gi(x, y)=0 . The solution to the original equation is then given by the product
(14.29).
14.3.2 Equations soluble for x
Equations that can be solved for x, i.e. such that they may be written in the form
x=F(y,p), (14.31)
can be reduced to first-degree first-order equations in pby differentiating both
sides with respect to y,s ot h a t
dx
dy=1
p=∂F
∂y+∂F
∂pdp
dy.
This results in an equation of the form G(y,p) = 0, which can be used together
with (14.31) to eliminate pand give the general solution. Note that often a singular
solution to the equation will be found at the same time (see the introduction tothis chapter).
487
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
6y2p2+3xp−y=0. (14.32)
This equation can be solved for xexplicitly to give 3 x=(y/p)−6y2p. Differentiating both
sides with respect to y, we find
3dx
dy=3
p=1
p−y
p2dp
dy−6y2dp
dy−12yp,
which factorises to give/;
1+6 yp2
/
/
2p+ydp
dy
/
=0. (14.33)
Setting the factor containing dp/dy equal to zero gives a first-degree first-order equation
inp, which may be solved to give py2=c. Substituting for pin (14.32) then yields the
general solution of (14.32):
y3=3cx+6c2. (14.34)
If we now consider the first factor in (14.33), we find 6 p2y=−1 as a possible solution.
Substituting for pin (14.32) we find the singular solution
8y3+3x2=0.
Note that the singular solution contains no arbitrary constants and cannot be found from
the general solution (14.34) by any choice of the constant c.
J
Solution method. Write the equation in the form (14.31) and differentiate both
sides with respect to y. Rearrange the resulting equation into the form G(y,p)=0,
which can be used together with the original ODE to eliminate pand so give the
general solution. If G(y,p)can be factorised then the factor containing dp/dy should
be used to eliminate pand give the general solution. Using the other factors in this
fashion will instead lead to singular solutions.
14.3.3 Equations soluble for y
Equations that can be solved for y, i.e. are such that they may be written in the
form
y=F(x, p), (14.35)
can be reduced to first-degree first-order equations in pby differentiating both
sides with respect to x,s ot h a t
dy
dx=p=∂F
∂x+∂F
∂pdp
dx.
This results in an equation of the form G(x, p) = 0, which can be used together
with (14.35) to eliminate pand give the general solution. An additional (singular)
solution to the equation is also often found.
488
14.3 HIGHER-DEGREE FIRST-ORDER EQUATIONSISolve
xp2+2xp−y=0. (14.36)
This equation can be solved for yexplicitly to give y=xp2+2xp. Differentiating both
sides with respect to x, we find
dy
dx=p=2xpdp
dx+p2+2xdp
dx+2p,
which after factorising gives
(p+1 )
/
p+2xdp
dx
/
=0. (14.37)
To obtain the general solution of (14.36), we consider the factor containing dp/dx .T h i s
first-degree first-order equation in phas the solution xp2=c(see subsection 14.3.1), which
we then use to eliminate pfrom (14.36). Thus we find that the general solution to (14.36)
is
(y−c)2=4cx. (14.38)
If instead, we set the other factor in (14.37) equal to zero, we obtain the very simple
solution p=−1. Substituting this into (14.36) then gives
x+y=0,
which is a singular solution to (14.36).
J
Solution method. Write the equation in the form (14.35) and differentiate both
sides with respect to x. Rearrange the resulting equation into the form G(x, p)=0,
which can be used together with the original ODE to eliminate pand so give the
general solution. If G(x, p)can be factorised then the factor containing dp/dx should
be used to eliminate pand give the general solution. Using the other factors in this
fashion will instead lead to singular solutions.
14.3.4 Clairaut’s equation
Finally, we consider Clairaut’s equation, which has the form
y=px+F(p) (14.39)
and is therefore a special case of equations soluble for y, as in (14.35). It may be
solved by a similar method to that given in subsection 14.3.3, but for Clairaut’sequation the form of the general solution is particularly simple. Differentiating(14.39) with respect to x, we find
dy
dx=p=p+xdp
dx+dF
dpdp
dx⇒dp
dxparenleftbiggdF
dp+xparenrightbigg
=0. (14.40)
Considering first the factor containing dp/dx , we find
dp
dx=d2y
dx2=0⇒ y=c1x+c2. (14.41)
489
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
Since p=dy/dx =c1, if we substitute (14.41) into (14.39) we find c1x+c2=
c1x+F(c1). Therefore the constant c2is given by F(c1), and the general solution
to (14.39) is
y=c1x+F(c1), (14.42)
i.e. the general solution to Clairaut’s equation can be obtained by replacing p
in the ODE by the arbitrary constant c1. Now, considering the second factor in
(14.40), we also have
dF
dp+x=0, (14.43)
which has the form G(x, p) = 0. This relation may be used to eliminate pfrom
(14.39) to give a singular solution.ISolve
y=px+p2. (14.44)
From (14.42) the general solution is y=cx+c2. But from (14.43) we also have 2 p+x=
0⇒p=−x/2. Substituting this into (14.44) we find the singular solution x2+4y=0 .
J
Solution method. Write the equation in the form (14.39), then the general solution
is given by replacing pby some constant c, as shown in (14.42). Using the relation
dF/dp +x=0to eliminate pfrom the original equation yields the singular solution.
14.4 Exercises
14.1 A radioactive isotope decays in such a way that the number of atoms present at
ag i v e nt i m e , N(t), obeys the equation
dN
dt=−λN.
If there are initially N0atoms present, find N(t)a tl a t e rt i m e s .
14.2 Solve the following equations by separation of the variables:
(a)y/prime−xy3=0 ;
(b)y/primetan−1x−y(1 +x2)−1=0 ;
(c)x2y/prime+xy2=4y2.
14.3 Show that the following equations are either exact or can be made exact, and
solve them:
(a)y(2x2y2+1 )y/prime+x(y4+1 )=0 ;
(b) 2 xy/prime+3x+y=0 ;
(c) (cos2x+ysin2x)y/prime+y2=0 .
14.4 Find the values of αandβthat make
F(x, y)=
/1
x2+2+α
y
/
dx+(xyβ+1 )dy
an exact differential. For these values solve F(x, y)=0 .
490
14.4 EXERCISES
14.5 By finding a suitable integrating factor, solve the following equations:
(a) (1−x2)y/prime+2xy=( 1−x2)3/2;
(b)y/prime−ycotx+c o s e c x=0 ;
(c) ( x+y3)y/prime=y(treat yas the independent variable).
14.6 By finding an appropriate integrating factor, solve
dy
dx=−2x2+y2+x
xy.
14.7 Find, in the form of an integral, the solution of the equation
αdy
dt+y=f(t)
for a general function f(t). Find the specific solutions for
(a)f(t)=H(t),
(b)f(t)=δ(t),
(c)f(t)=β−1e−t/βH(t)w i t h β<α .
For case (c), what happens if β→0?
14.8 An electric circuit contains a resistance Rand a capacitor Cin series, and a
battery supplying a time-varying electromotive force V(t). The charge qon the
capacitor therefore obeys the equation
Rdq
dt+q
C=V(t).
Assuming that initially there is no charge on the capacitor, and given that
V(t)=V0sinωt, find the charge on the capacitor as a function of time.
14.9 Using tangential-polar coordinates (see exercise 2.20), consider a particle of mass
mmoving under the influence of a force fdirected towards the origin O. By
resolving forces along the instantaneous tangent and normal and making use ofthe result of exercise 2.20 for the instantaneous radius of curvature, prove that
f=−mvdv
drand mv2=fpdr
dp.
Show further that h=mpvis a constant of the motion and that the law of force
can be deduced from
f=h2
p3dp
dr.
14.10 Use the result of the previous exercise to find the law of force, acting towards
the origin, under which a particle must move so as to describe the followingtrajectories:
(a) A circle of radius awhich passes through the origin;
(b) An equiangular spiral, which is defined by the property that the angle α
between the tangent and the radius vector is constant along the curve.
14.11 Solve
(y−x)dy
dx+2x+3y=0.
14.12 A mass mis accelerated by a time-varying force exp( −βt)v3,w h e r e vis its velocity.
It also experiences a resistive force ηv,w h e r e ηis a constant, owing to its motion
through the air. The equation of motion of the mass is therefore
mdv
dt=e x p (−βt)v3−ηv.
491
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
Find an expression for the velocity vof the mass as a function of time, given
that it has an initial velocity v0.
14.13 Using the results about Laplace transforms given in chapter 13 for df/dt and
tf(t), show, for a function y(t)t h a ts a t i s fi e s
tdy
dt+(t−1)y=0 ( * )
with y(0) finite, that ¯y(s)=C(1 +s)−2for some constant C.
Given that
y(t)=t+∞X
n=2antn,
determine Cand show that an=(−1)n/n!. Compare this result with that obtained
by integrating (*) directly.
14.14 Solve
dy
dx=1
x+2y+1.
14.15 Solve
dy
dx=−x+y
3x+3y−4.
14.16 If u=1+t a n y,c a l c u l a t e d(lnu)/dy; hence find the general solution of
dy
dx=t a n xcosy(cosy+s i n y).
14.17 Solve
x(1−2x2y)dy
dx+y=3x2y2,
given that y(1) = 1 /2.
14.18 A reflecting mirror is made in the shape of the surface of revolution generated by
revolving the curve y(x) about the x-axis. In order that light rays emitted from a
point source at the origin are reflected back parallel to the x-axis, the curve y(x)
must obey
y
x=2p
1−p2,
where p=dy/dx . By solving this equation for xfind the curve y(x).
14.19 Find the curve such that at each point on it the sum of the intercepts on the x-
andy-axes of the tangent to the curve (taking account of sign) is equal to 1.
14.20 Find a parametric solution of
x
/dy
dx
/2
+dy
dx−y=0
as follows.
(a) Write an equation for yin terms of p=dy/dx and show that
p=p2+( 2px+1 )dp
dx.
(b) Using pas the independent variable, arrange this as a linear first-order
equation for x.
492
14.4 EXERCISES
(c) Find an appropriate integrating factor to obtain
x=lnp−p+c
(1−p)2,
which, together with the expression for yobtained in (a), gives a parameter-
isation of the solution.
(d) Reverse the roles of xandyin steps (a) to (c), putting dx/dy =p−1,a n d
show that essentially the same parameterisation is obtained.
14.21 Using the substitutions u=x2andv=y2, reduce the equation
xy
/dy
dx
/2
−(x2+y2−1)dy
dx+xy=0
to Clairaut’s form. Hence show that the equation represents a family of conics
and the four sides of a square.
14.22 The action of the control mechanism on a particular system for an input f(t)i s
described, for t≥0, by the coupled first-order equations:
˙y+4z=f(t),
˙z−2z=˙y+1
2y.
Use Laplace transforms to find the response y(t) of the system to a unit step
input f(t)=H(t), given that y(0) = 1 and z(0) = 0.
Questions 23 to 31 are intended to give the reader practice in choosing an ap-
propriate method. The level of difficulty varies within the set; if necessary, the hints
may be consulted for an indication of the most appropriate approach .
14.23 Find the general solutions of the following:
(a)dy
dx+xy
a2+x2=x;( b )dy
dx=4y2
x2−y2.
14.24 Solve the following first-order equations for the boundary conditions given:
(a)y/prime−(y/x)=1 ,y (1) =−1;
(b)y/prime−ytanx=1,y(π/4) = 3;
(c)y/prime−y2/x2=1/4,y(1) = 1;
(d)y/prime−y2/x2=1/4,y(1) = 1 /2.
14.25 An electronic system has two inputs, to each of which a constant unit signal is
applied, but starting at different times. The equations governing the system thustake the form
˙x+2y=H(t),
˙y−2x=H(t−3).
Initially (at t=0 ) , x=1a n d y= 0; find x(t)a tl a t e rt i m e s .
14.26 Solve the differential equation
sinxdy
dx+2ycosx=1
subject to the boundary condition y(π/2) = 1.
14.27 Find the complete solution of/dy
dx
/2
−y
xdy
dx+A
x=0,
where Ais a positive constant.
493
FIRST-ORDER ORDINARY DIFFERENTIAL EQUATIONS
14.28 Find the solution of
(5x+y−7)dy
dx=3 (x+y+1 ).
14.29 Find the solution y=y(x)o f
xdy
dx+y−y2
x3/2=0,
subject to y(1) = 1.
14.30 Find the solution of
(2sin y−x)dy
dx=t a n y,
if (a) y(0) = 0, and (b) y(0) = π/2.
14.31 Find the family of solutions of
d2y
dx2+
/dy
dx
/2
+dy
dx=0
that satisfy y(0) = 0.
14.5 Hints and answers
14.1 N(t)=N0exp(−λt).
14.2 (a) y=±(c−x2)−1/2;( b ) y=ctan−1(x); (c) y=( l n x+4x−1−c)−1.
14.3 (a) exact, x2y4+x2+y2=c;( b )I F= x−1/2,x1/2(x+y)= c;( c )I F=
sec2x, y2tanx+y=c.
14.4 α=−1,β=−2; (1/√2) tan−1(x/√2)−(x/y)+y=c.
14.5 (a) IF = (1 −x2)−2,y=( 1−x2)(k+s i n−1x); (b) IF = cosec x, leading to
y=ksinx+c o s x; (c) exact equation is y−1(dx/dy )−xy−2=y, leading to
x=y(k+y2/2).
14.6 Integrating factor is x;3x4+2x3+3x2y2=c.
14.7 y(t)=e−t/α
Rtα−1et/prime/αf(t/prime)dt/prime;( a ) y(t)=1−e−t/α;( b ) y(t)=α−1e−t/α;( c ) y(t)=
(e−t/α−e−t/β)/(α−β). It becomes case (b).
14.8 q(t)=CV0[1 + ( ωCR)2]−1{sinωt+CRω[exp(−t/RC)−cosωt]}.
14.9 If the angle between the tangent and the radius vector is α, note that cos α=dr/ds
and sin α=p/r.
14.10 (a) r2=2ap, f∝ar−5;( b ) p=rsinα, f∝(sinα)−2r−3.
14.11 Homogeneous equation, put y=vxto obtain (1 −v)(v2+2v+2 )−1dv=x−1dx;
write 1−vas 2−(1 +v), and v2+2v+2a s1+( 1+ v)2;
A[x2+(x+y)2]=e x p
/
4t a n−1[(x+y)/x]
/
.
14.12 Bernoulli’s equation; set v=u−1/2to obtain m du/dt−2ηu=−2e x p (−βt);
v−2=2 (mβ+2η)−1[exp(−βt)−exp(2 ηt/m)] +v−2
0exp(2 ηt/m).
14.13 (1 + s)(d¯y/ds)+2¯y=0 . C=1 ; y(t)=te−t.
14.14 Follow subsection 14.2.8; k+y=l n ( x+2y+3 ) .
14.15 Equation is of the form of (14.22), set v=x+y;x+3y+2l n ( x+y−2) = A.
14.16 y=t a n−1(ksecx−1).
14.17 Equation is isobaric with weight y=−2; setting y=vx−2gives
v−1(1−v)−1(1−2v)dv=x−1dx;4xy(1−x2y)=1 .
14.18 Eliminate yto obtain, in turn, p(p2−1) = 2 x(dp/dx );p=±(1−Ax)−1/2;
A2y2=4 ( 1−Ax), i.e. a parabola.
14.19 The curve must satisfy y=( 1−p−1)−1(1−x+px), which has solution x=(p−1)−2,
leading to y=( 1±√x)2orx=( 1±√y)2; the singular solution p/prime= 0 gives
straight lines joining ( θ,0) and (0 ,1−θ)f o ra n y θ.
494
14.5 HINTS AND ANSWERS
14.20 (a) y=p2x+p; (d) the constants of integration will differ in the two cases.
14.21 v=qu+q/(q−1), where q=dv/du. General solution y2=cx2+c/(c−1),
hyperbolae for c>0 and ellipses for c<0. Singular solution y=±(x±1).
14.22 ¯y(s2+2s+2 )= s(¯f+1 )+( 2−2¯f);y(t)=−1+e−t(2cos t+3s i n t).
14.23 (a) Integrating factor is ( a2+x2)1/2,y=(a2+x2)/3+A(a2+x2)−1/2; (b) separable,
y=x(x2+Ax+4 )−1.
14.24 (a) y=xlnx−x;( b ) y=t a n x+√2se cx; (c) homogeneous, y=x(2−lnx)−1+
x/2;
(d) singular solution y=x/2.
14.25 Use Laplace transforms; ¯xs(s2+4 )= s+s2−2e−3s;
x(t)=1
2sin 2t+c o s2 t−1
2H(t−3) +1
2cos(2 t−6)H(t−3).
14.26 Integrating factor is sin x;y=( 1+c o s x)−1.
14.27 This is Clairaut’s equation with F(p)=A/p. General solution y=cx+A/c;
singular solution, y=2√
Ax.
14.28 Follow the second method demonstrated in subsection 14.2.8; x=X+2,y=
Y−3;
X(dv/dX )=( 3−2v−v2)/(5 +v); (x−y−5)3=A(3x+y−3).
14.29 Either Bernoulli’s equation with n= 2 or an isobaric equation with m=3/2;
y(x)=5 x3/2/(2 + 3 x5/2).
14.30 Treat yas the independent variable, giving the general solution xsiny=
−(cos2 y)/2+k.( a )y=s i n−1x;( b ) x=−cosycoty.
14.31 Show that p=(Cex−1)−1,w h e r e p=dy/dx ;y=l n [ C−e−x)/(C−1)] or
ln[D−(D−1)e−x]o rl n ( e−K+1−e−x)+K.
495
15
Higher-order ordinary differential
equations
Following on from the discussion of first-order ordinary differential equations
(ODEs) given in the previous chapter, we now examine equations of second andhigher order. Since a brief outline of the general properties of ODEs and theirsolutions was given at the beginning of the previous chapter, we will not repeat
it here. Instead, we will begin with a discussion of various types of higher-order
equation. This chapter is divided into three main parts. We first discuss linearequations with constant coefficients and then investigate linear equations withvariable coefficients. Finally, we discuss a few methods that may be of use insolving general linear or non-linear ODEs. Let us start by considering somegeneral points relating to alllinear ODEs.
Linear equations are of paramount importance in the description of physical
processes. Moreover, it is an empirical fact that, when put into mathematicalform, many natural processes appear as higher-order linear ODEs, most oftenas second-order equations. Although we could restrict our attention to thesesecond-order equations, the generalisation to nth-order equations requires little
extra work, and so we will consider this more general case.
A linear ODE of general order nhas the form
a
n(x)dny
dxn+an−1(x)dn−1y
dxn−1+···+a1(x)dy
dx+a0(x)y=f(x). (15.1)
Iff(x) = 0 then the equation is called homogeneous ; otherwise it is inhomogeneous .
The first-order linear equation studied in subsection 14.2.4 is a special case of(15.1). As discussed at the beginning of the previous chapter, the general solutionto (15.1) will contain narbitrary constants, which may be determined if nboundary
conditions are also provided.
In order to solve any equation of the form (15.1), we must first find the
general solution of the complementary equation , i.e. the equation formed by setting
496
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
f(x)=0 :
an(x)dny
dxn+an−1(x)dn−1y
dxn−1+···+a1(x)dy
dx+a0(x)y=0. (15.2)
To determine the general solution of (15.2), we must find nlinearly independent
functions that satisfy it. Once we have found these solutions, the general solutionis given by a linear superposition of these nfunctions. In other words, if the n
solutions of (15.2) are y
1(x),y2(x),...,y n(x), then the general solution is given by
the linear superposition
yc(x)=c1y1(x)+c2y2(x)+···+cnyn(x), (15.3)
where the cmare arbitrary constants that may be determined if nboundary
conditions are provided. The linear combination yc(x) is called the complementary
function of (15.1).
The question naturally arises how we establish that any nindividual solutions to
(15.2) are indeed linearly independent. For nfunctions to be linearly independent
over an interval, there must not exist anyset of constants c1,c2,...,c nsuch that
c1y1(x)+c2y2(x)+···+cnyn(x) = 0 (15.4)
over the interval in question, except for the trivial case c1=c2=···=cn=0 .
A statement equivalent to (15.4), which is perhaps more useful for the practical
determination of linear independence, can be found by repeatedly differentiating
(15.4), n−1 times in all, to obtain nsimultaneous equations for c1,c2,...,c n:
c1y1(x)+c2y2(x)+···+cnyn(x)=0
c1y1/prime(x)+c2y2/prime(x)+···+cnyn/prime(x)=0
...
c1y(n−1)
1(x)+c2y(n−1)
2+···+cny(n−1)
n(x)=0 ,(15.5)
where the primes denote differentiation with respect to x. Referring to the
discussion of simultaneous linear equations given in chapter 8, if the determinantof the coefficients of c
1,c2,...,c nis non-zero then the only solution to equations
(15.5) is the trivial solution c1=c2=···=cn= 0. In other words, the nfunctions
y1(x),y2(x),...,y n(x) are linearly independent over an interval if
W(y1,y2,...,y n)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingley
1 y2... y n
y1/primey2/prime...
.........
y(n−1)
1 ... ... y(n−1)
nvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle/negationslash= 0 (15.6)
over that interval; W(y
1,y2,...,y n) is called the Wronskian of the set of functions.
It should be noted, however, that the vanishing of the Wronskian does notguarantee that the functions are linearly dependent.
497
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
If the original equation (15.1) has f(x) = 0 (i.e. it is homogeneous) then of
course the complementary function yc(x) in (15.3) is already the general solution.
If, however, the equation has f(x)/negationslash= 0 (i.e. it is inhomogeneous) then yc(x)i so n l y
one part of the solution. The general solution of (15.1) is then given by
y(x)=yc(x)+yp(x), (15.7)
where yp(x)i st h e particular integral ,w h i c hc a nb e anyfunction that satisfies (15.1)
directly, provided it is linearly independent of yc(x). It should be emphasised for
practical purposes that anysuch function, no matter how simple (or complicated),
is equally valid in forming the general solution (15.7).
It is important to realise that the above method for finding the general solution
to an ODE by superposing particular solutions assumes crucially that the ODEis linear. For non-linear equations, discussed in section 15.3, this method cannotbe used, and indeed it is often impossible to find closed-form solutions to suchequations.
15.1 Linear equations with constant coefficients
If the a
min (15.1) are constants rather than functions of xthen we have
andny
dxn+an−1dn−1y
dxn−1+···+a1dy
dx+a0y=f(x). (15.8)
Equations of this sort are very common throughout the physical sciences and
engineering, and the method for their solution falls into two parts as discussed
in the previous section, i.e. finding the complementary function yc(x) and finding
the particular integral yp(x). If f(x) = 0 in (15.8) then we do not have to find
a particular integral, and the complementary function is by itself the generalsolution.
15.1.1 Finding the complementary function y
c(x)
The complementary function must satisfy
andny
dxn+an−1dn−1y
dxn−1+···+a1dy
dx+a0y= 0 (15.9)
and contain narbitrary constants (see equation (15.3)). The standard method
for finding yc(x) is to try a solution of the form y=Aeλx, substituting this into
(15.9). After dividing the resulting equation through by Aeλx, we are left with a
polynomial equation in λof order n;t h i si st h e auxiliary equation and reads
anλn+an−1λn−1+···+a1λ+a0=0. (15.10)
498
15.1 LINEAR EQUATIONS WITH CONSTANT COEFFICIENTS
In general the auxiliary equation has nroots, say λ1,λ2,...,λ n. In certain cases,
some of these roots may be repeated and some may be complex. The three maincases are as follows.
(i)All roots real and distinct. In this case the nsolutions to (15.9) are exp λ
mx
form=1t o n. It is easily shown by calculating the Wronskian (15.6)
of these functions that if all the λmare distinct then these solutions are
linearly independent. We can therefore linearly superpose them, as in(15.3), to form the complementary function
y
c(x)=c1eλ1x+c2eλ2x+···+cneλnx. (15.11)
(ii)Some roots complex. For the special (but usual) case that all the coefficients
amin (15.9) are real, if one of the roots of the auxiliary equation (15.10)
is complex, say α+iβ, then its complex conjugate α−iβis also a root. In
this case we can write
c1e(α+iβ)x+c2e(α−iβ)x=eαx(d1cosβx+d2sinβx)
=Aeαxbraceleftbiggsin
cosbracerightbigg
(βx+φ), (15.12)
where Aandφare arbitrary constants.
(iii)Some roots repeated. If, for example, λ1occurs ktimes ( k>1) as a root
of the auxiliary equation, then we have not found nlinearly independent
solutions of (15.9); formally the Wronskian (15.6) of these solutions, havingtwo or more identical columns, is equal to zero. We must therefore findk−1 further solutions that are linearly independent of those already found
and also of each other. By direct substitution into (15.9) we find that
xe
λ1x,x2eλ1x, ... , xk−1eλ1x
are also solutions, and by calculating the Wronskian it is easily shown that
they, together with the solutions already found, form a linearly independent
set of nfunctions. Therefore the complementary function is given by
yc(x)=(c1+c2x+···+ckxk−1)eλ1x+ck+1eλk+1x+ck+2eλk+2x+···+cneλnx.
(15.13)
If more than one root is repeated the above argument is easily extended.
For example, suppose as before that λ1is ak-fold root of the auxiliary
equation and, further, that λ2is an l-fold root (of course, k>1a n d l>1).
Then, from the above argument, the complementary function reads
yc(x)=(c1+c2x+···+ckxk−1)eλ1x
+(ck+1+ck+2x+···+ck+lxl−1)eλ2x
+ck+l+1eλk+l+1x+ck+l+2eλk+l+2x+···+cneλnx.(15.14)
499
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONSIFind the complementary function of the equation
d2y
dx2−2dy
dx+y=ex. (15.15)
Setting the RHS to zero, substituting y=Aeλxand dividing through by Aeλxwe obtain
the auxiliary equation
λ2−2λ+1=0 .
The root λ= 1 occurs twice and so, although exis a solution to (15.15), we must find
a further solution to the equation that is linearly independent of ex.F r o mt h ea b o v e
discussion, we deduce that xexis such a solution, so that the full complementary function
is given by the linear superposition
yc(x)=(c1+c2x)ex.
J
Solution method. Set the RHS of the ODE to zero (if it is not already so), and
substitute y=Aeλx. After dividing through the resulting equation by Aeλx, obtain
annth-order polynomial equation in λ(the auxiliary equation, see (15.10)). Solve
the auxiliary equation to find the nroots, λ1,λ2,...,λ n, say. If all these roots are
real and distinct then yc(x)is given by (15.11). If, however, some of the roots are
complex or repeated then yc(x)is given by (15.12) or (15.13), or the extension
(15.14) of the latter, respectively.
15.1.2 Finding the particular integral yp(x)
There is no generally applicable method for finding the particular integral yp(x)
but, for linear ODEs with constant coefficients and a simple RHS, yp(x) can often
be found by inspection or by assuming a parameterised form similar to f(x). The
latter method is sometimes called the method of undetermined coefficients .I ff(x)
contains only polynomial, exponential, or sine and cosine terms then, by assuminga trial function for y
p(x) of similar form but one which contains a number of
undetermined parameters and substituting this trial function into (15.9), theparameters can be found and y
p(x) deduced. Standard trial functions are as
follows.
(i) If f(x)=aerxthen try
yp(x)=berx.
(ii) If f(x)=a1sinrx+a2cosrx(a1ora2may be zero) then try
yp(x)=b1sinrx+b2cosrx.
(iii) If f(x)=a0+a1x+···+aNxN(some ammay be zero) then try
yp(x)=b0+b1x+···+bNxN.
500
15.1 LINEAR EQUATIONS WITH CONSTANT COEFFICIENTS
(iv) If f(x) is the sum or product of any of the above then try yp(x)a st h e
sum or product of the corresponding individual trial functions.
It should be noted that this method fails if any term in the assumed trial
function is also contained within the complementary function yc(x). In such a
case the trial function should be multiplied by the smallest integer power of x
such that it will then contain no term that already appears in the complementary
function. The undetermined coefficients in the trial function can now be foundby substitution into (15.8).
Three further methods that are useful in finding the particular integral y
p(x)
are Green’s functions, the variation of parameters, and making a change in thedependent variable based on knowledge of the complementary function. However,since these methods are also applicable to equations with variable coefficients, adiscussion of them is postponed until section 15.2.IFind a particular integral of the equation
d2y
dx2−2dy
dx+y=ex.
From the above discussion our first guess at a trial particular integral would be yp(x)=bex.
However, since the complementary function of this equation is yc(x)=( c1+c2x)ex(as
in the previous subsection), we see that exis already contained in it, as indeed is xex.
Multiplying our first guess by the lowest integer power of xsuch that the result does not
appear in yc(x), we therefore try yp(x)=bx2ex. Substituting this into the ODE, we find
thatb=1/2, so the particular integral is given by yp(x)=x2ex/2.
J
Solution method. If the RHS of an ODE contains only the functions mentioned at
the start of this subsection then the appropriate trial function should be substitutedinto it, thereby fixing the undetermined parameters. If, however, the RHS of the
equation is not of this form then one of the more general methods outlined in sub-
sections 15.2.3–15.2.5 should be used; perhaps the most straightforward of these isthe variation-of-parameters method.
15.1.3 Constructing the general solution y
c(x)+yp(x)
As stated earlier, the full solution to the ODE (15.8) is found by adding together
the complementary function and any particular integral. In order to illustrate
further the material discussed in the last two subsections, let us find the generalsolution to a new example, starting from the beginning.
501
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
d2y
dx2+4y=x2sin2x. (15.16)
First we set the RHS to zero and assume the trial solution y=Aeλx. Substituting this into
(15.16) leads to the auxiliary equation
λ2+4=0 ⇒ λ=±2i. (15.17)
Therefore the complementary function is given by
yc(x)=c1e2ix+c2e−2ix=d1cos 2x+d2sin2x. (15.18)
We must now turn our attention to the particular integral yp(x). Consulting the list of
standard trial functions in the previous subsection, we find that a first guess at a suitabletrial function for this case should be
(ax
2+bx+c)(dsin 2x+ecos 2x). (15.19)
However, we see that this trial function contains terms in sin 2 xand cos 2 x, both of which
already appear in the complementary function (15.18). We must therefore multiply (15.19)
by the smallest integer power of xthat ensures that none of the resulting terms appears
inyc(x). Since multiplying by xwill suffice, we finally assume the trial function
(ax3+bx2+cx)(dsin2x+ecos 2x). (15.20)
Substituting this into (15.16) to fix the constants appearing in (15.20), we find the particular
integral to be
yp(x)=−x3
12cos 2x+x2
16sin2x+x
32cos 2x. (15.21)
The general solution to (15.16) then reads
y(x)=yc(x)+yp(x)
=d1cos 2x+d2sin2x−x3
12cos 2x+x2
16sin2x+x
32cos 2x.
J
15.1.4 Linear recurrence relations
The discrete analogues of differential equations are called recurrence relations (or
sometimes difference equations ). Whereas a differential equation gives a prescrip-
tion, in terms of current values, for the new value of an dependent variable at a
point only infinitesimally far away, a recurrence relation describes how the nextin a sequence of values u
n, defined only at (non-negative) integer values of the
‘independent variable’ n,i st ob ec a l c u l a t e d .
In its most general form a recurrence relation expresses the way in which un+1
is to be calculated from all the preceding values u0,u1,... ,u n. Just as the most
general differential equations are intractable, so are the most general recurrence
relations, and we will limit ourselves to analogues of the types of differential
equations studied earlier in this chapter, namely those that are linear, haveconstant coefficients and possess simple functions on the RHS. Such equations
502
15.1 LINEAR EQUATIONS WITH CONSTANT COEFFICIENTS
occur over a broad range of engineering and statistical physics as well as in the
realms of finance, business planning and gambling! They form the basis of manynumerical methods, particularly those concerned with the numerical solution ofordinary and partial differential equations.
A general recurrence relation is exemplified by the formula
u
n+1=N−1summationdisplay
r=0arun−r+k, (15.22)
where Nand the arare fixed and kis a constant or a simple function of n.
Such an equation, involving terms of the series whose indices differ by up to N
(ranging from n−N+1 to n), is called an Nth-order recurrence relation. It is clear
that, given values for u0,u1,... ,u N−1, this is a definitive scheme for generating the
series and therefore has a unique solution.
Parallelling the nomenclature of differential equations, if the term not involving
anyunis absent, i.e. k= 0, then the recurrence relation is called homogeneous .
The parallel continues with the form of the general solution of (15.22). If vnis
the general solution of the homogeneous relation, and wnisanysolution of the
full relation, then
un=vn+wn
is the most general solution of the complete recurrence relation. This is straight-
forwardly verified as follows:
un+1=vn+1+wn+1
=N−1summationdisplay
r=0arvn−r+N−1summationdisplay
r=0arwn−r+k
=N−1summationdisplay
r=0ar(vn−r+wn−r)+k
=N−1summationdisplay
r=0arun−r+k.
Of course, if k=0t h e n wn=0f o ra l l nis a trivial particular solution and the
complementary solution, vn, is itself the most general solution.
First-order recurrence relations
First-order relations, for which N= 1, are exemplified by
un+1=aun+k, (15.23)
with u0specified. The solution to the homogeneous relation is immediate,
un=Can,
503
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
and, if kis a constant, the particular solution is equally straightforward: wn=K
for all n,p r o v i d e d Kis chosen to satisfy
K=aK+k,
i.e.K=k(1−a)−1. This will be sufficient unless a=1 ,i nw h i c hc a s e un=u0+nk
is obvious by inspection.
Thus the general solution of (15.23) is
un=braceleftBigg
Can+k/(1−a)a/negationslash=1,
u0+nk a =1.(15.24)
Ifu0is specified for the case of a/negationslash=1t h e n Cmust be chosen as C=u0−k/(1−a),
resulting in the equivalent form
un=u0an+k1−an
1−a. (15.25)
We now illustrate this method with a worked example.IA house-buyer borrows capital Bfrom a bank that charges a fixed annual rate of interest
R% .I ft h el o a ni st ob er e p a i do v e r Yyears, at what value should the fixed annual payments
P, made at the end of each year, be set? For a loan over 25 years at 6%, what percentage
of the first year’s payment goes towards paying off the capital?
Letundenote the outstanding debt at the end of year n,a n dw r i t e R/100 = r. Then the
relevant recurrence relation is
un+1=un(1 +r)−P
with u0=B. From (15.25) we have
un=B(1 +r)n−P1−(1 +r)n
1−(1 +r).
A st h el o a ni st ob er e p a i do v e r Yyears, uY= 0 and thus
P=Br(1 +r)Y
(1 +r)Y−1.
The first year’s interest is rBand so the fraction of the first year’s payment going
towards capital repayment is ( P−rB)/P, which, using the above expression for P,i se q u a l
to (1 + r)−Y. With the given figures, this is (only) 23%.
J
With only small modifications, the method just described can be adapted to
handle recurrence relations in which the constant kin (15.23) is replaced by kαn,
i.e. the relation is
un+1=aun+kαn. (15.26)
As for an inhomogeneous linear differential equation (see subsection 15.1.2), we
may try as a potential particular solution a form which resembles the term thatmakes the equation inhomogeneous. Here, the presence of the term kα
nindicates
504
15.1 LINEAR EQUATIONS WITH CONSTANT COEFFICIENTS
that a particular solution of the form un=Aαnshould be tried. Substituting this
into (15.26) gives
Aαn+1=aAαn+kαn,
from which it follows that A=k/(α−a) and that there is a particular solution
having the form un=kαn/(α−a), provided α/negationslash=a. For the special case α=a,t h e
reader can readily verify that a particular solution of the form un=Anαnis appro-
priate. This mirrors the corresponding situation for linear differential equations
when the RHS of the differential equation is contained in the complementary
function of its LHS.
In summary, the general solution to (15.26) is
un=braceleftBigg
C1an+kαn/(α−a)α/negationslash=a,
C2an+knαn−1α=a,(15.27)
with C1=u0−k/(α−a)a n d C2=u0.
Second-order recurrence relations
We consider next recurrence relations that involve un−1in the prescription for
un+1and treat the general case in which the intervening term, un, is also present.
A typical equation is thus
un+1=aun+bun−1+k. (15.28)
As previously, the general solution of this is un=vn+wn,w h e r e vnsatisfies
vn+1=avn+bvn−1 (15.29)
andwnisanyparticular solution of (15.28); the proof follows the same lines as
that given earlier.
We have already seen for a first-order recurrence relation that the solution to
the homogeneous equation is given by terms forming a geometric series, and weconsider a corresponding series of powers in the present case. Setting v
n=Aλnin
(15.29) for some λ, as yet undetermined, requires that λshould satisfy
Aλn+1=aAλn+bAλn−1.
Dividing through by Aλn−1(assumed non-zero) shows that λcould be either of
the roots, λ1andλ2,o f
λ2−aλ−b=0, (15.30)
which is known as the characteristic equation of the recurrence relation.
That there are two possible series of terms of the form Aλnis consistent with the
fact that two initial values (boundary conditions) have to be provided before the
series can be calculated by repeated use of (15.28). These two values are sufficientto determine the appropriate coefficient Afor each of the series. Since (15.29) is
505
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
both linear and homogeneous, and is satisfied by both vn=Aλn
1andvn=Bλn
2,i t s
general solution is
vn=Aλn
1+Bλn
2.
If the coefficients aandbare such that (15.30) has two equal roots, i.e. a2=−4b,
then, as in the analogous case of repeated roots for differential equations (seesubsection 15.1.1(iii)), the second term of the general solution is replaced by Bnλ
n
1
to give
vn=(A+Bn)λn
1.
Finding a particular solution is straightforward if kis a constant: a trivial but
adequate solution is wn=k(1−a−b)−1for all n. As with first-order equations,
particular solutions can be found for other simple forms of kby trying functions
similar to kitself. Thus particular solutions for the cases k=Cnandk=Dαn
can be found by trying wn=E+Fnandwn=Gαnrespectively.IFind the value of u16if the series unsatisfies
un+1+4un+3un−1=n
forn≥1,w i t h u0=1andu1=−1.
We first solve the characteristic equation,
λ2+4λ+3=0 ,
to obtain the roots λ=−1a n d λ=−3. Thus the complementary function is
vn=A(−1)n+B(−3)n.
In view of the form of the RHS of the original relation, we try
wn=E+Fn
as a particular solution and obtain
E+F(n+1 )+4 ( E+Fn)+3 [ E+F(n−1)] = n,
yielding F=1/8a n d E=1/32.
Thus the complete general solution is
un=A(−1)n+B(−3)n+n
8+1
32,
and now using the given values for u0andu1determines Aas 7/8a n d Bas 3/32. Thus
un=1
32[28(−1)n+3 (−3)n+4n+1].
Finally, substituting n=1 6g i v e s u16= 4035633, a value the reader may (or may not)
wish to verify by repeated application of the initial recurrence relation.
J
506
15.1 LINEAR EQUATIONS WITH CONSTANT COEFFICIENTS
Higher-order recurrence relations
It will be apparent that linear recurrence relations of order N>2 do not present
any additional difficulty in principle, though two obvious practical difficulties are
(i) that the characteristic equation is of order Nand in general will not have roots
that can be written in closed form and (ii) that a correspondingly large numberof given values is required to determine the Notherwise arbitrary constants
in the solution. The algebraic labour needed to solve the set of simultaneouslinear equations which determines them increases rapidly with N. We do not give
specific examples here, but some are included in the exercises at the end of the
chapter.
15.1.5 Laplace transform method
The method of Laplace transforms is very useful for solving linear ODEs with
constant coefficients. Taking the Laplace transform of such an equation trans-forms it into a purely algebraic equation in terms of the Laplace transform
of the required solution. Once the algebraic equation has been solved for thisLaplace transform, the general solution to the original ODE can be obtainedby performing an inverse Laplace transform. One advantage of this method isthat, for given boundary conditions, it provides the solution in just one step,
instead of having to find the complementary function and particular integral
separately.
In order to apply this method we need only two results from Laplace transform
theory (see section 13.2). First, the Laplace transform of a function f(x) is defined
by
¯f(s)≡integraldisplay
∞
0e−sxf(x)dx, (15.31)
from which we can derive a second useful relation. This concerns the Laplace
transform of derivatives of f(x):
f(n)(s)=sn¯f(s)−sn−1f(0)−sn−2f/prime(0)−···−sf(n−2)(0)−f(n−1)(0),
(15.32)
where the primes and superscripts in parentheses denote differentiation with
respect to x. Using these relations, along with the table 13.1, on p. 461, which
gives Laplace transforms of standard functions, we are in a position to solve alinear ODE with constant coefficients by this method.
507
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
d2y
dx2−3dy
dx+2y=2e−x, (15.33)
subject to the boundary conditions y(0) = 2 ,y/prime(0) = 1 .
Taking the Laplace transform of (15.33) and using the table of standard results we obtain
s2¯y(s)−sy(0)−y/prime(0)−3[s¯y(s)−y(0)]+2¯y(s)=2
s+1,
which reduces to
(s2−3s+2 )¯y(s)−2s+5=2
s+1. (15.34)
Solving this algebraic equation for ¯y(s), the Laplace transform of the required solution to
(15.33), we obtain
¯y(s)=2s2−3s−3
(s+1 ) ( s−1)(s−2)=1
3(s+1 )+2
s−1−1
3(s−2), (15.35)
where in the final step we have used partial fractions. Taking the inverse Laplace transform
of (15.35), again using table 13.1, we find the specific solution to (15.33) to be
y(x)=1
3e−x+2ex−1
3e2x.
J
Note that if the boundary conditions in a problem are given as symbols, rather
than just numbers, then the step involving partial fractions can often involvea considerable amount of algebra. The Laplace transform method is also veryconvenient for solving sets of simultaneous linear ODEs with constant coefficients.ITwo electrical circuits, both of negligible resistance, each consist of a coil having self-
inductance Land a capacitor having capacitance C. The mutual inductance of the two
circuits is M. There is no source of e.m.f. in either circuit. Initially the second capacitor
is given a charge CV0, the first capacitor being uncharged, and at time t=0as w i t c hi n
the second circuit is closed to complete the circuit. Find the subsequent current in the firstcircuit.
Subject to the initial conditions q1(0) = ˙q1(0) = ˙q2(0) = 0 and q2(0) = CV0=V0/G,s a y ,
we have to solve
L¨q1+M¨q2+Gq1=0,
M¨q1+L¨q2+Gq2=0.
On taking the Laplace transform of the above equations, we obtain
(Ls2+G)¯q1+Ms2¯q2=sMV 0C,
Ms2¯q1+(Ls2+G)¯q2=sLV 0C.
Eliminating ¯q2and rewriting as an equation for ¯q1, we find
¯q1(s)=MV 0s
[(L+M)s2+G][(L−M)s2+G]
=V0
2G
/(L+M)s
(L+M)s2+G−(L−M)s
(L−M)s2+G
/
.
508
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
Using table 13.1,
q1(t)=1
2V0C(cosω1t−cosω2t),
where ω2
1(L+M)=Gandω2
2(L−M)=G. Thus the current is given by
i1(t)=1
2V0C(ω2sinω2t−ω1sinω1t).
J
Solution method. Perform a Laplace transform, as defined in (15.31), on the entire
equation, using (15.32) to calculate the transform of the derivatives. Then solve theresulting algebraic equation for ¯y(s), the Laplace transform of the required solution
to the ODE. By using the method of partial fractions and consulting a table ofLaplace transforms of standard functions, calculate the inverse Laplace transform.The resulting function y(x)is the solution of the ODE that obeys the given boundary
conditions.
15.2 Linear equations with variable coefficients
There is no generally applicable method of solving equations with coefficients
that are functions of x. Nevertheless, there are certain cases in which a solution is
possible. Some of the methods discussed in this section are also useful in findingthe general solution or particular integral for equations with constant coefficientsthat have proved impenetrable by the techniques discussed above.
15.2.1 The Legendre and Euler linear equations
Legendre’s linear equation has the form
a
n(αx+β)ndny
dxn+···+a1(αx+β)dy
dx+a0y=f(x), (15.36)
where α,βand the anare constants and may be solved by making the substitution
αx+β=et. We then have
dy
dx=dt
dxdy
dt=α
αx+βdy
dt
d2y
dx2=d
dxdy
dx=α2
(αx+β)2parenleftbiggd2y
dt2−dy
dtparenrightbigg
and so on for higher derivatives. Therefore we can write the terms of (15.36) as
(αx+β)dy
dx=αdy
dt,
(αx+β)2d2y
dx2=α2d
dtparenleftbiggd
dt−1parenrightbigg
y,
...
(αx+β)ndny
dxn=αnd
dtparenleftbiggd
dt−1parenrightbigg
···parenleftbiggd
dt−n+1parenrightbigg
y.(15.37)
509
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
Substituting equations (15.37) into the original equation (15.36), the latter becomes
a linear ODE with constant coefficients, i.e.
anαnd
dtparenleftbiggd
dt−1parenrightbigg
···parenleftbiggd
dt−n+1parenrightbigg
y+···+a1αdy
dt+a0y=fparenleftbigget−β
αparenrightbigg
,
which can be solved by the methods of section 15.1.
A special case of Legendre’s linear equation, for which α=1a n d β=0 ,i s
Euler’s equation ,
anxndny
dxn+···+a1xdy
dx+a0y=f(x); (15.38)
it may be solved in a similar manner to the above by substituting x=et.I f
f(x) = 0 in (15.38), substituting y=xλleads to a simple algebraic equation in
λ, which can be solved to yield the solution to (15.38). In the event that the
algebraic equation for λhas repeated roots, extra care is needed. If λ1is ak-fold
root ( k>1) then the klinearly independent solutions corresponding to this root
arexλ1,xλ1lnx ,...,xλ1(lnx)k−1.ISolve
x2d2y
dx2+xdy
dx−4y= 0 (15.39)
by both of the methods discussed above.
First we make the substitution x=et, which, after cancelling et, gives an equation with
constant coefficients, i.e.
d
dt
/d
dt−1
/
y+dy
dt−4y=0⇒d2y
dt2−4y=0. (15.40)
Using the methods of section 15.1, the general solution of (15.40), and therefore of (15.39),
is given by
y=c1e2t+c2e−2t=c1x2+c2x−2.
Since the RHS of (15.39) is zero, we can reach the same solution by substituting y=xλ
into (15.39). This gives
λ(λ−1)xλ+λxλ−4xλ=0,
which reduces to
(λ2−4)xλ=0.
This has the solutions λ=±2, so we obtain again the general solution
y=c1x2+c2x−2.
J
Solution method. If the ODE is of the Legendre form (15.36) then substitute αx+
β=et. This results in an equation of the same order but with constant coefficients,
which can be solved by the methods of section 15.1. If the ODE is of the Euler
form (15.38) with a non-zero RHS then substitute x=et; this again leads to an
equation of the same order but with constant coefficients. If, however, f(x)=0 in
the Euler equation (15.38) then the equation may also be solved by substituting
510
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
y=xλ. This leads to an algebraic equation whose solution gives the allowed values
ofλ; the general solution is then the linear superposition of these functions.
15.2.2 Exact equations
Sometimes an ODE may be merely the derivative of another ODE of one order
lower. If this is the case then the ODE is called exact. The nth-order linear ODE
an(x)dny
dxn+···+a1(x)dy
dx+a0(x)y=f(x), (15.41)
is exact if the LHS can be written as a simple derivative, i.e. if
an(x)dny
dxn+···+a0(x)y=d
dxbracketleftbigg
bn−1(x)dn−1y
dxn−1+···+b0(x)ybracketrightbigg
. (15.42)
It may be shown that, for (15.42) to hold, we require
a0(x)−a/prime
1(x)+a/prime/prime
2(x)−···+(−1)na(n)
n(x)=0 , (15.43)
where the prime again denotes differentiation with respect to x. If (15.43) is
satisfied then straightforward integration leads to a new equation of one order
lower. If this simpler equation can be solved then a solution to the original
equation is obtained. Of course, if the above process leads to an equation that isitself exact then the analysis can be repeated to reduce the order still further.ISolve
(1−x2)d2y
dx2−3xdy
dx−y=1. (15.44)
Comparing with (15.41), we have a2=1−x2,a1=−3xanda0=−1. It is easily shown
thata0−a/prime
1+a/prime/prime
2= 0, so (15.44) is exact and can the refore be written in the form
d
dx
/
b1(x)dy
dx+b0(x)y
/
=1. (15.45)
Expanding the LHS of (15.45) we find
d
dx
/
b1dy
dx+b0y
/
=b1d2y
dx2+(b/prime
1+b0)dy
dx+b/prime
0y. (15.46)
Comparing (15.44) and (15.46) we find
b1=1−x2,b/prime
1+b0=−3x, b/prime
0=−1.
These relations integrate consistently to give b1=1−x2andb0=−x, so (15.44) can be
written asd
dx
/
(1−x2)dy
dx−xy
/
=1. (15.47)
Integrating (15.47) gives us directly the first-order linear ODE
dy
dx−
/x
1−x2
/
y=x+c1
1−x2,
which can be solved by the method of subsection 14.2.4 and has the solution
y=c1sin−1x+c2√
1−x2−1.
J
511
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
It is worth noting that, even if a higher-order ODE is not exact in its given form,
it may sometimes be made exact by multiplying through by some suitable function,anintegrating factor , cf. subsection 14.2.3. Unfortunately, no straightforward
method for finding an integrating factor exists and one often has to rely on
inspection or experience.ISolve
x(1−x2)d2y
dx2−3x2dy
dx−xy=x. (15.48)
It is easily shown that (15.48) is not exact, but we also see immediately that by multiplying
it through by 1 /xwe recover (15.44), which is exact and is solved above.
J
Another important point is that an ODE need not be linear to be exact,
although no simple rule such as (15.43) exists if it is not linear. Nevertheless, it isoften worth exploring the possibility that a non-linear equation is exact, since itcould then be reduced in order by one and may lead to a soluble equation. This
is discussed further in subsection 15.3.3.
Solution method. For a linear ODE of the form (15.41) check whether it is exact
using equation (15.43). If it is not then attempt to find an integrating factor which
when multiplying the equation makes it exact. Once the equation is exact write the
LHS as a derivative as in (15.42) and, by expanding this derivative and comparingwith the LHS of the ODE, determine the functions b
m(x)in (15.42). Integrate the
resulting equation to yield another ODE, of one order lower. This may be solved orsimplified further if the new ODE is itself exact or can be made so.
15.2.3 Partially known complementary function
Suppose we wish to solve the nth-order linear ODE
a
n(x)dny
dxn+···+a1(x)dy
dx+a0(x)y=f(x), (15.49)
a n dw eh a p p e nt ok n o wt h a t u(x) is a solution of (15.49) when the RHS is
set to zero, i.e. u(x) is one part of the complementary function. By making the
substitution y(x)=u(x)v(x), we can transform (15.49) into an equation of order
n−1i ndv/dx. This simpler equation may prove soluble.
In particular, if the original equation is of second order then we obtain
a first-order equation in dv/dx, which may be soluble using the methods of
section 14.2. In this way both the remaining term in the complementary function
and the particular integral are found. This method therefore provides a useful
way of calculating particular integrals for second-order equations with variable(or constant) coefficients.
512
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTSISolve
d2y
dx2+y=c o s e c x. (15.50)
We see that the RHS does not fall into any of the categories listed in subsection 15.1.2,
and so we are at an initial loss as to how to find the particular integral. However, thecomplementary function of (15.50) is
y
c(x)=c1sinx+c2cosx,
and so let us choose the solution u(x)=c o s x(we could equally well choose sin x)a n d
make the substitution y(x)=v(x)u(x)=v(x)c osxinto (15.50). This gives
cosxd2v
dx2−2si nxdv
dx=c o s e c x, (15.51)
which is a first-order linear ODE in dv/dx and may be solved by multiplying through by
a suitable integrating factor, as discussed in subsection 14.2.4. Writing (15.51) as
d2v
dx2−2tan xdv
dx=cosec x
cosx, (15.52)
we see that the required integrating factor is given by
exp
/
−2
Z
tanxd x
/
=e x p [2l n( c os x)]=c o s2x.
Multiplying both sides of (15.52) by the integrating factor cos2xwe obtain
d
dx
/
cos2xdv
dx
/
=c o t x,
which integrates to give
cos2xdv
dx=l n ( s i n x)+c1.
After rearranging and integrating again, this becomes
v=
Z
sec2xln(sin x)dx+c1
Z
sec2xd x
=t a n xln(sin x)−x+c1tanx+c2.
Therefore the general solution to (15.50) is given by y=uv=vcosx,i . e .
y=c1sinx+c2cosx+s i n xln(sin x)−xcosx,
which contains the full complementary function and the particular integral.
J
Solution method. Ifu(x)is a known solution of the nth-order equation (15.49) with
f(x)=0, then make the substitution y(x)=u(x)v(x)in (15.49). This leads to an
equation of order n−1indv/dx, which might be soluble.
513
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
15.2.4 Variation of parameters
The method of variation of parameters proves useful in finding particular integrals
for linear ODEs with variable (and constant) coefficients. However, it requiresknowledge of the entire complementary function, not just of one part of it as inthe previous subsection.
Suppose we wish to find a particular integral of the equation
a
n(x)dny
dxn+···+a1(x)dy
dx+a0(x)y=f(x), (15.53)
and the complementary function yc(x) (the general solution of (15.53) with
f(x)=0 )i s
yc(x)=c1y1(x)+c2y2(x)+···+cnyn(x),
where the functions ym(x) are known. We now assume that a particular integral of
(15.53) can be expressed in a form similar to that of the complementary function,but with the constants c
mreplaced by functions of x,i . e .w ea s s u m eap a r t i c u l a r
integral of the form
yp(x)=k1(x)y1(x)+k2(x)y2(x)+···+kn(x)yn(x). (15.54)
This will no longer satisfy the complementary equation (i.e. (15.53) with the RHS
set to zero) but might, with suitable choices of the functions ki(x), be made equal
tof(x), thus producing not a complementary function but a particular integral.
Since we have narbitrary functions k1(x),k2(x),...,k n(x), but only one restric-
tion on them (namely the ODE), we may impose a further n−1 constraints. We
can choose these constraints to be as convenient as possible, and the simplestchoice is given by
k
/prime
1(x)y1(x)+k/prime
2(x)y2(x)+···+k/prime
n(x)yn(x)=0
k/prime
1(x)y/prime
1(x)+k/prime
2(x)y/prime
2(x)+···+k/prime
n(x)y/prime
n(x)=0
... (15.55)
k/prime
1(x)y(n−2)
1(x)+k/prime
2(x)y(n−2)
2(x)+···+k/prime
n(x)y(n−2)
n(x)=0
k/prime
1(x)y(n−1)
1(x)+k/prime
2(x)y(n−1)
2(x)+···+k/prime
n(x)y(n−1)
n(x)=f(x)
an(x),
where the primes denote differentiation with respect to x. The last of these
equations is not a freely chosen constraint but must be satisfied given the previousn−1 constraints and the original ODE.
This choice of constraints is easily justified (although the algebra is quite
messy). Differentiating (15.54) with respect to x,w eo b t a i n
y
/prime
p=k1y/prime
1+k2y/prime
2+···+kny/prime
n+(k/prime
1y1+k/prime
2y2+···+k/prime
nyn),
where, for the moment, we drop the explicit x-dependence of these functions. Since
514
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
we are free to choose our constraints as we wish, let us define the expression in
parentheses to be zero, giving the first equation in (15.55). Differentiating againwe find
y
/prime/prime
p=k1y/prime/prime
1+k2y/prime/prime
2+···+kny/prime/prime
n+(k/prime
1y/prime
1+k/prime
2y/prime
2+···+k/prime
ny/prime
n).
Once more we can choose the expression in brackets to be zero, giving the second
equation in (15.55). We can repeat this procedure, choosing the corresponding
expression in each case to be zero. This yields the first n−1 equations in (15.55).
Themth derivative of ypform<n is then given by
y(m)
p=k1y(m)
1+k2y(m)
2+···+kny(m)
n.
Differentiating yponce more we find that its nth derivative is given by
y(n)
p=k1y(n)
1+k2y(n)
2+···+kny(n)
n+(k/prime
1y(n−1)
1+k/prime
2y(n−1)
2+···+k/prime
ny(n−1)
n).
Substituting the expressions for y(m)
p,m=0t o n, into the original ODE (15.53),
we obtain
nsummationdisplay
m=0am(k1y(m)
1+k2y(m)
2+···+kny(m)
n)+an(k/prime
1y(n−1)
1+k/prime
2y(n−1)
2+···+k/prime
ny(n−1)
n)=f(x).
i.e.
nsummationdisplay
m=0amnsummationdisplay
j=1kjy(m)
j+an(k1/primey(n−1)
1+k2/primey(n−1)
2+···+kn/primey(n−1)
n=f(x).
Rearranging the order of summations on the LHS, we find
nsummationdisplay
j=1kj(any(n)
j+···+a1y/prime
j+a0yj)+an(k/prime
1y(n−1)
1+k2/primey(n−1)
2+···+k/prime
ny(n−1)
n)=f(x).
(15.56)
But since the functions yjare solutions of the complementary equation of (15.53)
we have (for all j)
any(n)
j+···+a1y/prime
j+a0yj=0.
Therefore (15.56) becomes
an(k/prime
1y(n−1)
1+k2/primey(n−1)
2+···+k/prime
ny(n−1)
n)=f(x),
which is the final equation given in (15.55).
Considering (15.55) to be a set of simultaneous equations in the set of unknowns
k/prime
1(x),k/prime
2,...,k/prime
n(x), we see that the determinant of the coefficients of these functions
is equal to the Wronskian W(y1,y2,...,y n), which is non-zero since the solutions
ym(x) are linearly independent; see equation (15.6). Therefore (15.55) can be solved
for the functions k/prime
m(x), which in turn can be integrated, setting all constants of
515
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
integration equal to zero, to give km(x). The general solution to (15.53) is then
given by
y(x)=yc(x)+yp(x)=nsummationdisplay
m=1[cm+km(x)]ym(x).
Note that if the constants of integration are included in the km(x) then, as well
as finding the particular integral, we introduce an addition to the complementaryfunction.IUse the variation of parameters method to solve
d2y
dx2+y=c o s e c x, (15.57)
subject to the boundary conditions y(0) = y(π/2) = 0 .
The complementary function of (15.57) is again
yc(x)=c1sinx+c2cosx.
We therefore assume a particular integral of the form
yp(x)=k1(x)sinx+k2(x)cosx,
and impose the additional constraints of (15.55), i.e.
k/prime
1(x)sinx+k/prime
2(x)cosx=0,
k/prime
1(x)cosx−k/prime
2(x)sinx=c o s e c x.
Solving these equations for k/prime
1(x)a n d k/prime
2(x)g i v e s
k/prime
1(x)=c o s xcosec x=c o t x,
k/prime
2(x)=−sinxcosec x=−1.
Hence, ignoring the constants of integration, k1(x)a n d k2(x) are given by
k1(x) = ln(sin x),
k2(x)=−x.
The general solution to the ODE (15.57) is therefore
y(x)=[c1+l n ( s i n x)]sinx+(c2−x)cosx,
which is identical to the solution found in subsection 15.2.3. Applying the boundary
conditions y(0) = y(π/2) = 0 we find c1=c2=0a n ds o
y(x) = ln(sin x)sinx−xcosx.
J
Solution method. If the complementary function of (15.53) is known then assume
a particular integral of the same form but with the constants replaced by functions
ofx. Impose the constraints in (15.55) and solve the resulting system of equations
for the unknowns k/prime
1(x),k/prime
2,...,k/prime
n(x). Integrate these functions, setting constants of
integration equal to zero, to obtain k1(x),k2(x),...,k n(x)and hence the particular
integral.
516
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
15.2.5 Green’s functions
The Green’s function method of solving linear ODEs bears a striking resemblance
to the method of variation of parameters discussed in the previous subsection;it too requires knowledge of the entire complementary function in order to findthe particular integral and therefore the general solution. The Green’s functionapproach differs, however, since once the Green’s function for a particular LHS of
(15.1) and accompanying boundary conditions has been found, then the solution
foranyRHS (i.e. any f(x)) can be written down immediately, albeit in the form
of an integral.
Although the Green’s function method can be approached by considering the
superposition of eigenfunctions of the equation (see chapter 17) and is also
applicable to the solution of partial differential equations (see chapter 19), thissection adopts a more utilitarian approach based on the properties of the Diracdelta function (see subsection 13.1.3) and deals only with the use of Green’sfunctions in solving ODEs.
Let us again consider the equation
a
n(x)dny
dxn+···+a1(x)dy
dx+a0(x)y=f(x), (15.58)
but for the sake of brevity we now denote the LHS by Ly(x), i.e. as a linear
differential operator acting on y(x). Thus (15.58) now reads
Ly(x)=f(x). (15.59)
Let us suppose that a function G(x, z) exists (the Green’s function ) such that the
general solution to (15.59), which obeys some set of imposed boundary conditions
in the range a≤x≤b,i sg i v e nb y
y(x)=integraldisplayb
aG(x, z)f(z)dz, (15.60)
where zis the integration variable. If we apply the linear differential operator L
to both sides of (15.60) and use (15.59) then we obtain
Ly(x)=integraldisplayb
a[LG(x, z)]f(z)dz=f(x). (15.61)
Comparison of (15.61) with a standard property of the Dirac delta function (see
subsection 13.1.3), namely
f(x)=integraldisplayb
aδ(x−z)f(z)dz,
fora≤x≤b, shows that for (15.61) to hold for any arbitrary function f(x), we
require (for a≤x≤b)t h a t
LG(x, z)=δ(x−z), (15.62)
517
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
i.e. the Green’s function G(x, z)must satisfy the original ODE with the RHS set
equal to a delta function .G(x, z) may be thought of physically as the response of
a system to a unit impulse at x=z.
In addition to (15.62), we must impose two further sets of restrictions on
G(x, z). The first is the requirement that the general solution y(x) in (15.60) obeys
the boundary conditions. For homogeneous boundary conditions, in which y(x)
and/or its derivatives are required to be zeroat specified points, this is most
simply arranged by demanding that G(x, z) itself obeys the boundary conditions
when it is considered as a function of xalone; if, for example, we require
y(a)=y(b) = 0 then we should also demand G(a, z)=G(b, z)=0 .P r o b l e m s
having inhomogeneous boundary conditions are discussed at the end of thissubsection.
The second set of restrictions concerns the continuity or discontinuity of G(x, z)
and its derivatives at x=zand can be found by integrating (15.62) with respect
toxover the small interval [ z−/epsilon1, z+/epsilon1] and taking the limit as /epsilon1→0. We then
obtain
lim
/epsilon1→0nsummationdisplay
m=0integraldisplayz+/epsilon1
z−/epsilon1am(x)dmG(x, z)
dxmdx= lim
/epsilon1→0integraldisplayz+/epsilon1
z−/epsilon1δ(x−z)dx=1. (15.63)
Since dnG/dxnexists at x=zbut with value infinity, the ( n−1)th-order derivative
must have a finite discontinuity there, whereas all the lower-order derivatives,d
mG/dxmform<n−1, must be continuous at this point. Therefore the terms
containing these derivatives cannot co ntribute to the value of the integral on
the LHS of (15.63). Noting that, apart from an arbitrary additive constant,integraltext
(dmG/dxm)dx=dm−1G/dxm−1, and integrating the terms of (15.63) by parts we
find
lim
/epsilon1→0integraldisplayz+/epsilon1
z−/epsilon1am(x)dmG(x, z)
dxmdx= 0 (15.64)
form=0t o n−1. Thus, since only the term containing dnG/dxncontributes to
the integral in (15.63), we conclude, after performing an integration by parts, that
lim
/epsilon1→0bracketleftbigg
an(x)dn−1G(x, z)
dxn−1bracketrightbiggz+/epsilon1
z−/epsilon1=1. (15.65)
Thus we have the further nconstraints that G(x, z) and its derivatives up to order
n−2 are continuous at x=zbut that dn−1G/dxn−1has a discontinuity of 1 /an(z)
atx=z.
Thus the properties of the Green’s function G(x, z)f o ra n nth-order linear ODE
may be summarised by the following.
(i)G(x, z) obeys the original ODE but with f(x) on the RHS set equal to a
delta function δ(x−z).
518
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
(ii) When considered as a function of xalone G(x, z) obeys the specified
(homogeneous) boundary conditions on y(x).
(iii) The derivatives of G(x, z) with respect to xup to order n−2 are continuous
atx=z, but the ( n−1)th-order derivative has a discontinuity of 1 /an(z)
at this point.IUse Green’s functions to solve
d2y
dx2+y=c o s e c x, (15.66)
subject to the boundary conditions y(0) = y(π/2) = 0 .
From (15.62) we see that the Green’s function G(x, z)m u s ts a t i s f y
d2G(x, z)
dx2+G(x, z)=δ(x−z). (15.67)
Now it is clear that for x/negationslash=zthe RHS of (15.67) is zero, and we are left with the
task of finding the general solution to the homogeneous equation, i.e. the complementary
function. The complementary function of (15.67) consists of a linear superposition of sin x
and cos xandmustconsist of different superpositions on either side of x=z, since its
(n−1)th derivative (i.e. the first derivative in this case) is required to have a discontinuity
there. Therefore we assume the form of the Green’s function to be
G(x, z)=
/A(z)sinx+B(z)c osxforx<z ,
C(z)sinx+D(z)c osxforx>z .
Note that we have performed a similar (but not identical) operation to that used in the
variation of parameters method, i.e. we have replaced the constants in the complementaryfunction with func tions (this time of z).
We must now impose the relevant restrictions on G(x, z) in order to determine the
functions A(z),...,D (z). The first of these is that G(x, z) should itself obey the homogeneous
boundary conditions G(0,z)=G(π/2,z) = 0. This leads to the conclusion that B(z)=
C(z) = 0, so we now have
G(x, z)=
/A(z)sinxforx<z ,
D(z)c osxforx>z .
The second restriction is the continuity conditions given in equations (15.64), (15.65),
namely that, for this second-order equation, G(x, z) is continuous at x=zanddG/dx has
a discontinuity of 1 /a2(z) = 1 at this point. Applying these two constraints we have
D(z)c osz−A(z)sinz=0
−D(z)sinz−A(z)cosz=1.
Solving these equations for A(z)a n d D(z), we find
A(z)=−cosz, D (z)=−sinz.
Thus we have
G(x, z)=
/−coszsinxforx<z ,
−sinzcosxforx>z .
Therefore, from (15.60), the general solution to (15.66) that obeys the boundary conditions
519
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
y(0) = y(π/2) = 0 is given by
y(x)=
Zπ/2
0G(x, z)cosec zd z
=−cosx
Zx
0sinzcosec zd z−sinx
Zπ/2
xcoszcosec zd z
=−xcosx+s i n xln(sin x),
which agrees with the result obtained in the previous subsections.
J
As mentioned earlier, once a Green’s function has been obtained for a given
LHS and boundary conditions, it can be used to find a general solution for anyRHS; thus, the solution of d
2y/dx2+y=f(x), with y(0) = y(π/2) = 0, is given
immediately by
y(x)=integraldisplayπ/2
0G(x, z)f(z)dz
=−cosxintegraldisplayx
0sinzf(z)dz−sinxintegraldisplayπ/2
xcoszf(z)dz. (15.68)
As an example, the reader may wish to verify that if f(x)=s i n2 xthen (15.68)
gives y(x)=(−sin2x)/3, a solution easily verified by direct substitution. In
general, analytic integration of (15.68) for arbitrary f(x) will, prove intractable;
then the integrals must be evaluated numerically.
Another important point is that although the Green’s function method above
has provided a general solution, it is also useful for finding a particular integral
if the complementary function is known. This is easily seen since in (15.68) theconstant integration limits 0 and π/2 lead merely to constant values by which
the factors sin xand cos xare multiplied; thus the complementary function is
reconstructed. The rest of the general solution, i.e. the particular integral, comesfrom the variable integration limit x. Therefore by changingintegraltext
π/2
xto−integraltextx,a n ds o
dropping the constant integration limits, we can find just the particular integral.
For example, a particular integral of d2y/dx2+y=f(x) that satisfies the above
boundary conditions is given by
yp(x)=−cosxintegraldisplayx
sinzf(z)dz+s i n xintegraldisplayx
coszf(z)dz.
A very important point to realise about the Green’s function method is that a
particular G(x, z) applies to a given LHS of an ODE andthe imposed boundary
conditions, i.e. the same equation with different boundary conditions will have a
different Green’s function . To illustrate this point, let us consider again the same
ODE as solved above, but with different boundary conditions.
520
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTSIUse Green’s functions to solve
d2y
dx2+y=f(x), (15.69)
subject to the one-point boundary conditions y(0) = y/prime(0) = 0 .
We again require (15.67) to hold and so again we assume a Green’s function of the form
G(x, z)=
/A(z)sinx+B(z)c osxforx<z ,
C(z)sinx+D(z)c osxforx>z .
However, we now require G(x, z) to obey the boundary conditions G(0,z)=G/prime(0,z)=0 ,
which imply A(z)=B(z) = 0. Therefore we have
G(x, z)=
/0f o r x<z ,
C(z)sinx+D(z)c osxforx>z .
Applying the continuity conditions on G(x, z) as before now gives
C(z)sinz+D(z)c osz=0,
C(z)cosz−D(z)s i nz=1,
which are solved to give
C(z)=c o s z, D (z)=−sinz.
So finally the Green’s function is given by
G(x, z)=
/0f o r x<z ,
sin(x−z)f o r x>z ,
and the general solution to (15.69) that obeys the boundary conditions y(0) = y/prime(0) = 0 is
y(x)=
Z∞
0G(x, z)f(z)dz
=
Zx
0sin(x−z)f(z)dz.
J
Finally, we consider how to deal with inhomogeneous boundary conditions
such as y(a)=α,y(b)=βory(0) = y/prime(0) = γetc., where α, β, γ are non-
zero. The simplest method of solution in this case is to make a change of
variable such that the boundary conditions in the new variable, usay, are
homogeneous, i.e. u(a)=u(b)=0o r u(0) = u/prime(0) = 0 etc. For nth-order equations
we generally require nboundary conditions to fix the solution, but these n
boundary conditions can be of various types. For example we may have the n-
point boundary conditions y(xm)=ymform=1t o n, or the one-point boundary
conditions y(x0)=y/prime(x0)=···=y(n−1)(x0)=y0,o rs o m e t h i n gi nb e t w e e n .I na l l
cases a suitable change of variable is
u=y−h(x),
where h(x)i sa n( n−1)th-order polynomial that obeys the boundary conditions.
521
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
For example, if we consider the second-order case with boundary conditions
y(a)=α,y(b)=βthen a suitable change of variable is
u=y−(mx+c),
where y=mx+cis the straight line through the points ( a, α)a n d( b, β); this
is given by m=(α−β)/(a−b)a n d c=(βa−αb)/(a−b). Alternatively, if the
boundary conditions for our second-order equation are y(0) = y/prime(0) = γthen we
would make the same change of variable, but this time y=mx+cwould be the
straight line through (0 ,γ) with slope γ,i . e .m=c=γ.
Solution method. Require that the Green’s function G(x, z)obeys the original ODE,
but with the RHS set to a delta function δ(x−z). This is equivalent to assuming
thatG(x, z)is given by the complementary function of the original ODE, with the
constants replaced by functions of z; these functions are different for x<z andx>
z. Now require also that G(x, z)obeys the given homogeneous boundary conditions
and impose the continuity conditions given in (15.64) and (15.65). The generalsolution to the original ODE is then given by (15.60). For inhomogeneous boundary
conditions, make the change of dependent variable u=y−h(x),w h e r e h(x)is a
polynomial obeying the given boundary conditions.
15.2.6 Canonical form for second-order equations
In this section we specialise from nth-order linear ODEs with variable coefficients
to those of order 2. In particular we consider the equation
d
2y
dx2+a1(x)dy
dx+a0(x)y=f(x), (15.70)
which has been rearranged so that the coefficient of d2y/dx2is unity. By making
the substitution y(x)=u(x)v(x)w eo b t a i n
v/prime/prime+parenleftbigg2u/prime
u+a1parenrightbigg
v/prime+parenleftbiggu/prime/prime+a1u/prime+a0u
uparenrightbigg
v=f
u, (15.71)
where the prime denotes differentiation with respect to x. Since (15.71) would be
much simplified if there were no term in v/prime,l e tu sc h o o s e u(x) such that the first
factor in parentheses on the LHS of (15.71) is zero, i.e.
2u/prime
u+a1=0⇒ u(x)=e x pbraceleftbigg
−1
2integraldisplay
a1(z)dzbracerightbigg
. (15.72)
We then obtain an equation of the form
d2v
dx2+g(x)v=h(x), (15.73)
522
15.2 LINEAR EQUATIONS WITH VARIABLE COEFFICIENTS
where
g(x)=a0(x)−1
4[a1(x)]2−1
2a/prime
1(x)
h(x)=f(x)ex pbraceleftbigg
1
2integraldisplay
a1(z)dzbracerightbigg
.
Since (15.73) is of a simpler form than the original equation, (15.70), it may
prove easier to solve.ISolve
4x2d2y
dx2+4xdy
dx+(x2−1)y=0. (15.74)
Dividing (15.74) through by 4 x2, we see that it is of the form (15.70) with a1(x)=1 /x,
a0(x)=(x2−1)/4x2andf(x) = 0. Therefore, making the substitution
y=vu=vexp
/
−
Z1
2xdx
/
=Av√x,
we obtain
d2v
dx2+v
4=0. (15.75)
Equation (15.75) is easily solved to give
v=c1sin1
2x+c2cos1
2x,
so the solution of (15.74) is
y=v√x=c1sin1
2x+c2cos1
2x√x.
J
As an alternative to choosing u(x) such that the coefficient of v/primein (15.71) is
zero, we could choose a different u(x) such that the coefficient of vvanishes. For
this to be the case, we see from (15.71) that we would require
u/prime/prime+a1u/prime+a0u=0,
sou(x) would have to be a solution of the original ODE with the RHS set to
zero, i.e. part of the complementary function. If such a solution were known thenthe substitution y=uvwould yield an equation with no term in v, which could
be solved by two straightforward integrations. This is a special (second-order)
case of the method discussed in subsection 15.2.3.
Solution method. Write the equation in the form (15.70), then substitute y=uv,
where u(x)is given by (15.72). This leads to an equation of the form (15.73), in
which there is no term in dv/dx and which may be easier to solve. Alternatively,
if part of the complementary function is known then follow the method of subsec-tion 15.2.3.
523
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
15.3 General ordinary differential equations
In this section, we discuss miscellaneous methods for simplifying general ODEs.
These methods are applicable to both linear and non-linear equations and in
some cases may lead to a solution. More often than not, however, finding a
closed-form solution to a general non-linear ODE proves impossible.
15.3.1 Dependent variable absent
If an ODE does not contain the dependent variable yexplicitly, but only its
derivatives, then the change of variable p=dy/dx l e a d st oa ne q u a t i o no fo n e
order lower.ISolve
d2y
dx2+2dy
dx=4x (15.76)
This is transformed by the substitution p=dy/dx to the first-order equation
dp
dx+2p=4x. (15.77)
The solution to (15.77) is then found by the method of subsection 14.2.4 and reads
p=dy
dx=ae−2x+2x−1,
where ais a constant. Thus by direct integration the solution to the original equation,
(15.76), is
y(x)=c1e−2x+x2−x+c2.
J
An extension to the above method is appropriate if an ODE contains only
derivatives of ythat are of order mand greater. Then the substitution p=dmy/dxm
reduces the order of the ODE by m.
Solution method. If the ODE contains only derivatives of ythat are of order mand
greater then the substitution p=dmy/dxmreduces the order of the equation by m.
15.3.2 Independent variable absent
If an ODE does not contain the independent variable xexplicitly, except in d/dx,
d2/dx2etc., then as in the previous subsection we make the substitution p=dy/dx
524
15.3 GENERAL ORDINARY DIFFERENTIAL EQUATIONS
but also write
d2y
dx2=dp
dx=dy
dxdp
dy=pdp
dy
d3y
dx3=d
dxparenleftbigg
pdp
dyparenrightbigg
=dy
dxd
dyparenleftbigg
pdp
dyparenrightbigg
=p2d2p
dy2+pparenleftbiggdp
dyparenrightbigg2
, (15.78)
and so on for higher-order derivatives. This leads to an equation of one order
lower.ISolve
1+yd2y
dx2+
/dy
dx
/2
=0. (15.79)
Making the substitutions dy/dx =pandd2y/dx2=p(dp/dy ) we obtain the first-order
ODE
1+ypdp
dy+p2=0,
which is separable and may be solved as in subsection 14.2.1 to obtain
(1 +p2)y2=c1.
Using p=dy/dx we therefore have
p=dy
dx=±
s
c2
1−y2
y2,
which may be integrated to give the general solution of (15.79); after squaring this reads
(x+c2)2+y2=c2
1.
J
Solution method. If the ODE does not contain xexplicitly then substitute p=
dy/dx , along with the relations for higher derivatives given in (15.78), to obtain an
equation of one order lower, which may prove easier to solve.
15.3.3 Non-linear exact equations
As discussed in subsection 15.2.2, an exact ODE is one that can be obtained by
straightforward differentiation of an equation of one order lower. Moreover, thenotion of exact equations is useful for both linear and non-linear equations, sincean exact equation can be immediately integrated. It is possible, of course, thatthe resulting equation may itself be exact, so that the process can be repeated.
In the non-linear case, however, there is no simple relation (such as (15.43) for
the linear case) by which an equation can be shown to be exact. Nevertheless, ageneral procedure does exist and is illustrated in the following example.
525
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONSISolve
2yd3y
dx3+6dy
dxd2y
dx2=x. (15.80)
Directing our attention to the term on the LHS of (15.80) that contains the highest-order
derivative, i.e. 2 yd3y/dx3, we see that it can be obtained by differentiating 2 yd2y/dx2since
d
dx
/
2yd2y
dx2
/
=2yd3y
dx3+2dy
dxd2y
dx2. (15.81)
Rewriting the LHS of (15.80) using (15.81), we are left with 4( dy/dx )(d2y/dy2), which may
itself be written as a derivative, i.e.
4dy
dxd2y
dx2=d
dx
/"
2
/dy
dx
/2
/#
. (15.82)
Since, therefore, we can write the LHS of (15.80) as a sum of simple derivatives of other
functions, (15.80) is exact. Integrating (15.80) with respect to x, and using (15.81) and
(15.82), now gives
2yd2y
dx2+2
/dy
dx
/2
=
Z
xd x=x2
2+c1. (15.83)
Now we can repeat the process to find whether ( 15.83) is itself exact. Considering the term
on the LHS of (15.83) that contains the highest-order derivative, i.e. 2 yd2y/dx2, we note
that we obtain this by differentiating 2 yd y / d x , as follows:
d
dx
/
2ydy
dx
/
=2yd2y
dx2+2
/dy
dx
/2
.
The above expression already contains all the terms on the LHS of (15.83), so we can
integrate (15.83) to give
2ydy
dx=x3
6+c1x+c2.
Integrating once more we obtain the solution
y2=x4
24+c1x2
2+c2x+c3.
J
It is worth noting that both linear equations (as discussed in subsection 15.2.2)
and non-linear equations may sometimes be made exact by multiplying throughby an appropriate integrating factor. Although no general method exists forfinding such a factor, one may sometimes be found by inspection or inspiredguesswork.
Solution method. Rearrange the equation so that all the terms containing yor its
derivatives are on the LHS, then check to see whether the equation is exact byattempting to write the LHS as a simple derivative. If this is possible then theequation is exact and may be integrated directly to give an equation of one orderlower. If the new equation is itself exact the process can be repeated.
526
15.3 GENERAL ORDINARY DIFFERENTIAL EQUATIONS
15.3.4 Isobaric or homogeneous equations
It is straightforward to generalise the discussion of first-order isobaric equations
given in subsection 14.2.6 to equations of general order n.A n nth-order isobaric
equation is one in which every term can be made dimensionally consistent upongiving yanddyeach a weight m,a n d xanddxeach a weight 1. Then the nth
derivative of ywith respect to x, for example, would have dimensions minyand
−ninx. In the special case, with m= 1, for which the equation is dimensionally
consistent the equation is called homogeneous (not to be confused with linearequations with a zero RHS). If an equation is isobaric or homogeneous then the
change in dependent variable y=vx
m(y=vxin the homogeneous case) followed
by the change in independent variable x=etleads to an equation in which the
new independent variable tis absent except in the form d/dt.ISolve
x3d2y
dx2−(x2+xy)dy
dx+(y2+xy)=0 . (15.84)
Assigning yanddythe weight m,a n d xanddxthe weight 1, the weights of the five terms
on the LHS of (15.84) are, from left to right: m+1 , m+1 , 2 m,2m,m+1 . F o r t h e s e
weights all to be equal we require m= 1; thus (15.84) is a homogeneous equation. Since it
is homogeneous we now make the substitution y=vx, which, after dividing the resulting
equation through by x3,g i v e s
xd2v
dx2+( 1−v)dv
dx=0. (15.85)
Now substituting x=etinto (15.85) we obtain (after some working)
d2v
dt2−vdv
dt=0, (15.86)
which can be integrated directly to give
dv
dt=1
2v2+c1. (15.87)
Equation (15.87) is separable, and integrates to give
1
2t+d2=
Zdv
v2+d2
1
=1
d1tan−1
/v
d1
/
.
Rearranging and using x=etandy=vxwe finally obtain the solution to (15.84) as
y=d1xtan
/;1
2d1lnx+d1d2
/
.
J
Solution method. Assume that yanddyhave weight m, and xanddxweight 1,
and write down the combined weights of each term in the ODE. If these weights canbe made equal by assuming a particular value for mthen the equation is isobaric
(or homogeneous if m=1). Making the substitution y=vx
mfollowed by x=et
leads to an equation in which the new independent variable tis absent except in the
form d/dt.
527
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
15.3.5 Equations homogeneous in xoryalone
It will be seen that the intermediate equation (15.85) in the example of the
previous subsection was simplified by the substitution x=et, in that this led to
an equation in which the new independent variable toccurred only in the form
d/dt, see (15.86). A closer examination of (15.85) reveals that it is dimensionally
consistent in the independent variable xtaken alone ; this is equivalent to giving
the dependent variable and its differential a weight m= 0. For any equation that
is homogeneous in xalone, the substitution x=etwill lead to an equation that
does not contain the new independent variable texcept as d/dt. Note that the
Euler equation of subsection 15.2.1 is a special, linear example of an equation
homogeneous in xalone. Similarly, if an equation is homogeneous in yalone, then
substituting y=evleads to an equation in which the new dependent variable, v,
occurs only in the form d/dv.ISolve
x2d2y
dx2+xdy
dx+2
y3=0.
This equation is homogeneous in xalone, and on substituting x=etwe obtain
d2y
dt2+2
y3=0,
which does not contain the new independent variable texcept as d/dt. Such equations
may often be solved by the method of subsection 15.3.2, but in this case we can integratedirectly to obtain
dy
dt=
p
2(c1+1/y2).
This equation is separable, and we findZdyp
2(c1+1/y2)=t+c2.
By multiplying the numerator and denominator of the integrand on the LHS by y, we find
the solutionp
c1y2+1√
2c1=t+c2.
Remembering that t=l nxwe finally obtainp
c1y2+1√
2c1=l nx+c2.
J
Solution method. If the weight of xtaken alone is the same in every term in the
ODE then the substitution x=etleads to an equation in which the new independent
variable tis absent except in the form d/dt. If the weight of ytaken alone is the
same in every term then the substitution y=evleads to an equation in which the
new dependent variable vis absent except in the form d/dv.
528
15.4 EXERCISES
15.3.6 Equations having y=Aexas a solution
Finally, we note that if any general (linear or non-linear) nth-order ODE is
satisfied identically by assuming that
y=dy
dx=···=dny
dxn(15.88)
then y=Aexis a solution of that equation. This must be so because y=Aexis
a non-zero function that satisfies (15.88).IFind a solution of
(x2+x)dy
dxd2y
dx2−x2ydy
dx−x
/dy
dx
/2
=0. (15.89)
Setting y=dy/dx =d2y/dx2in (15.89), we obtain
(x2+x)y2−x2y2−xy2=0,
which is satisfied identically. Therefore y=Aexis a solution of (15.89); this is easily
verified by directly substituting y=Aexinto (15.89).
J
Solution method. If the equation is satisfied identically by assuming that y=
dy/dx =···=dny/dxnthen y=Aexis a solution.
15.4 Exercises
15.1 A simple harmonic oscillator, with natural frequency ω0, experiences an oscillating
driving force f(t)=c o s ωt. Therefore, its equation of motion is
d2x
dt2+ω2
0x=c o s ωt,
where xis its position. Given that at t= 0 we have x=dx/dt = 0, find the
function x(t). Describe the solution if ωis approximately, but not exactly, equal
toω0.
15.2 Find the roots of the auxiliary equation for the following. Hence solve them for
the boundary conditions stated.
(a)d2f
dt2+2df
dt+5f=0 w i t h f(0) = 1 ,f/prime(0) = 0 .
(b)d2f
dt2+2df
dt+5f=e−tcos 3twith f(0) = 0 ,f/prime(0) = 0 .
15.3 The theory of bent beams shows that at any point in the beam the ‘bending
moment’ is given by K/ρ,w h e r e Kis a constant (that depends upon the beam
material and cross-sectional shape) and ρis the radius of curvature at that point.
Consider a light beam of length Lwhose ends, x=0a n d x=L, are supported
at the same vertical height and which has a weight Wsuspended from its centre.
Verify that at any point x(0≤x≤L/2 for definiteness) the net magnitude of
the bending moments, (bending moment = force ×perpendicular distance) due
to the weight and support reactions, evaluated on either side of x,i sWx/2.
529
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
If the beam is only slightly bent, so that ( dy/dx )2/lessmuch1, where y=y(x)i st h e
downward displacement of the beam at x, show that the beam profile satisfies
the approximate equation
d2y
dx2=−Wx
2K.
By integrating this equation twice and using physically imposed conditions on
your solution at x=0a n d x=L/2, show that the downward displacement at
the centre of the beam is WL3/(48K).
15.4 Solve the differential equation
d2f
dt2+6df
dt+9f=e−t,
subject to the conditions f=0a n d df/dt =λatt=0 .
Find the equation satisfied by the positions of the turning points of f(t)a n d
hence, by drawing suitable sketch graphs, determine the number of turning pointsthe solution has in the range t>0i f( a ) λ=1/4, and (b) λ=−1/4.
15.5 The function f(t) satisfies the differential equation
d
2f
dt2+8df
dt+1 2f=1 2e−4t.
For the following sets of boundary conditions determine whether it has solutions,
and, if so, find them:
(a)f(0) = 0 ,f/prime(0) = 0 ,f(ln√2) = 0;
(b)f(0) = 0 ,f/prime(0) =−2,f(ln√2) = 0 .
15.6 Determine the values of αandβfor which the following functions are linearly
dependent:
y1(x)=xcoshx+s i n h x,
y2(x)=xsinhx+c o s h x,
y3(x)=(x+α)ex,
y4(x)=(x+β)e−x.
You will find it convenient to work with those linear combinations of the yi(x)
that can be written the most compactly.
15.7 A solution of the differential equation
d2y
dx2+2dy
dx+y=4e−x
takes the value 1 when x= 0 and the value e−1when x= 1. What is its value
when x=2 ?
15.8 The two functions x(t)a n d y(t) satisfy the simultaneous equations
dx
dt−2y=−sint,
dy
dt+2x=5c o s t.
Find explicit expressions for x(t)a n d y(t), given that x(0) = 3 and y(0) = 2.
Sketch the solution trajectory in the xy-plane for 0 ≤t<2π, showing that
the trajectory crosses itself at (0 ,1/2) and passes through the points (0 ,−3) and
(0,−1) in the negative x-direction.
530
15.4 EXERCISES
15.9 Find the general solutions of
(a)d3y
dx3−12dy
dx+1 6y=3 2x−8,
(b)d
dx
/1
ydy
dx
/
+( 2acoth2 ax)
/1
ydy
dx
/
=2a2,
where ais a constant.
15.10 Use the method of Laplace transforms to solve
(a)d2f
dt2+5df
dt+6f=0,f (0) = 1 ,f/prime(0) =−4,
(b)d2f
dt2+2df
dt+5f=0,f (0) = 1 ,f/prime(0) = 0 .
15.11 The quantities x(t),y(t) satisfy the simultaneous equations
¨x+2n˙x+n2x=0,
¨y+2n˙y+n2y=µ˙x,
where x(0) = y(0) = ˙y(0) = 0 and ˙x(0) = λ. Show that
y(t)=1
2µλt2
/;
1−1
3nt
/
exp(−nt).
15.12 Use Laplace transforms to solve, for t≥0, the differential equations
¨x+2x+y=c o s t,
¨y+2x+3y=2c o s t,
which describe a coupled system that starts from rest at the equilibrium position.
Show that the subsequent motion takes place along a straight line in the xy-plane.
Verify that the frequency at which the system is driven is equal to one of theresonance frequencies of the system; explain why there is noresonant behaviour
in the solution you have obtained.
15.13 Two unstable isotopes AandBand a stable isotope Chave the following decay
rates per atom present: A→B,3 s
−1;A→C,1 s−1;B→C,2 s−1. Initially
a quantity x0ofAis present and none of the other two types. Using Laplace
transforms, find the amount of Cpresent at a later time t.
15.14 For a lightly damped ( γ<ω 0) harmonic oscillator driven at its undamped
resonance frequency ω0, the displacement x(t)a tt i m e tsatisfies the equation
d2x
dt2+2γdx
dt+ω2
0x=Fsinω0t.
Use Laplace transforms to find the displacement at a general time if the oscillator
starts from rest at its equilibrium position.
(a) Show that ultimately the oscillation has amplitude F/(2ω0γ) with a phase
lag of π/2 relative to the driving force F.
(b) By differentiating the original equation, conclude that if x(t) is expanded as
a power series in tfor small tthen the first non-vanishing term is Fω0t3/6.
Confirm this conclusion by expanding your explicit solution.
15.15 The ‘golden mean’, which is said to describe the most aesthetically pleasing
proportions for the sides of a rectangle (e.g. the ideal picture frame), is givenby the limiting value of the ratio of successive terms of the Fibonacci series u
n,
which is generated by
un+2=un+1+un,
with u0=0a n d u1= 1. Find an expression for the general term of the series and
531
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
verify that the golden mean is equal to the larger root of the recurrence relation’s
characteristic equation.
15.16 In a particular scheme for modelling numerically one-dimensional fluid flow, the
successive values, un, of the solution are connected for n≥1 by the difference
equation
c(un+1−un−1)=d(un+1−2un+un−1),
where canddare positive constants. The boundary conditions are u0=0a n d
uM= 1. Find the solution to the equation and show that successive values of un
will have alternating signs if c>d.
15.17 The first few terms of a series un, starting with u0,a r e1 ,2,2,1,6,−3. The series
is generated by a recurrence relation of the form
un=Pun−2+Qun−4,
where PandQare constants. Find an expression for the general term of the
series and show that the series in fact consists of two other interleaved seriesgiven by
u
2m=2
3+1
34m,
u2m+1=7
3−1
34m,
form=0,1,2,... .
15.18 Find an explicit expression for the unsatisfying
un+1+5un+6un−1=2n,
given that u0=u1= 1. Deduce that 2n−26(−3)nis divisible by 5 for all integer
n.
15.19 Find the general expression for the unsatisfying
un+1=2un−2−un−1
with u0=u1=0a n d u2= 1, and show that they can be written in the form
un=1
5−2n/2
√5cos
/3πn
4−φ
/
,
where tan φ=2 .
15.20 Consider the seventh-order recurrence relation
un+7−un+6−un+5+un+4−un+3+un+2+un+1−un=0.
Find the most general form of its solution, and show that:
(a) if only the four initial values u0=0 ,u1=2 ,u3=6a n d u3= 12, are specified,
the relation has one solution which cycles repeatedly through this set of fournumbers.
(b) but if, in addition, it is required that u
4= 20, u5=3 0a n d u6= 42 then the
solution is unique, with un=n(n+1 ) .
15.21 Find the general solution of
x2d2y
dx2−xdy
dx+y=x,
given that y(1) = 1 and y(e)=2 e.
15.22 Find the general solution of
(x+1 )2d2y
dx2+3 (x+1 )dy
dx+y=x2.
532
15.4 EXERCISES
15.23 Prove that the general solution of
(x−2)d2y
dx2+3dy
dx+4y
x2=0
is given by
y(x)=1
(x−2)2
/
k
/2
3x−1
2
/
+cx2
/
.
15.24 Use the method of variation of parameters to find the general solutions of
(a)d2y
dx2−y=xn,( b )d2y
dx2−2dy
dx+y=2xex.
15.25 Use the intermediate result of exercise 15.24(a) to find the Green’s function which
satisfies
d2G(x, ξ)
dx2−G(x, ξ)=δ(x−ξ)w i t h G(0,ξ)=G(1,ξ)=0 .
15.26 (a) Given that y1(x)=1 /xis a solution of
F(x, y)=x(x+1 )d2y
dx2+( 2−x2)dy
dx−(2 +x)y=0,
find a second linearly independent solution,
(i) by setting y2(x)=y1(x)u(x),
(ii) by noting the sum of the coefficients in the equation.
(b) Hence, using the variation of parameters method, find the general solution
of
F(x, y)=(x+1 )2.
15.27 Show generally that if y1(x)a n d y2(x) are linearly independent solutions of
d2y
dx2+p(x)dy
dx+q(x)y=0,
with y1(0) = 0 and y2(1) = 0, then the Green’s function G(x, ξ) for the interval
0≤x, ξ≤1a n dw i t h G(0,ξ)=G(1,ξ)=0c a nb ew r i t t e ni nt h ef o r m
G(x, ξ)=
/(
y1(x)y2(ξ)/W(ξ)0<x<ξ
y2(x)y1(ξ)/W(ξ)ξ<x< 1,
where W(x)=W[y1(x),y2(x)] is the Wronskian of y1(x)a n d y2(x).
15.28 Use the result of the previous exercise to find the Green’s function G(x, ξ)t h a t
satisfies
d2G
dx2+3dG
dx+2G=δ(x−x),
in the interval 0 ≤x, ξ≤1w i t h G(0,ξ)=G(1,ξ) = 0. Hence obtain integral
expressions for the solution of
d2y
dx2+3dy
dx+2y=
/(
00 <x<x 0,
1x0<x< 1,
distinguishing between the cases (a) x<x 0,a n d( b ) x>x 0.
15.29 The equation of motion for a driven damped harmonic oscillator can be written
¨x+2˙x+( 1+ κ2)x=f(t),
with κ/negationslash=0 .I fi ts t a r t sf r o mr e s tw i t h x(0) = 0 and ˙x(0) = 0, find the corresponding
Green’s function G(t, τ) and verify that it can b e written as a function of t−τ
533
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
only. Find the explicit solution when the driving force is the unit step function,
i.e.f(t)=H(t). Confirm your solution by taking the Laplace transforms of both
it and the original equation.
15.30 Show that the Green’s function for the equation
d2y
dx2+y
4=f(x),
subject to the boundary conditions y(0) = y(π) = 0, is given by
G(x, z)=
/(
−2cos1
2xsin1
2z0≤z≤x,
−2sin1
2xcos1
2zx≤z≤π.
15.31 Find the Green’s function x=G(t, t0) that solves
d2x
dt2+αdx
dt=δ(t−t0)
under the initial conditions x=dx/dt =0a t t= 0. Hence solve
d2x
dt2+αdx
dt=f(t),
where f(t)=0f o r t<0.
Evaluate your answer explicitly for f(t)=Ae−at(t>0).
15.32 (a) By multiplying through by dy/dx , write down the solution to the equation
d2y
dx2+f(y)=0 ,
where f(y) can be any function.
(b) A mass m, initially at rest at the point x= 0, is accelerated by a force
f(x)=A(x0−x)
/
1+2l n
/
1−x
x0
//
.
Its equation of motion is md2x/dt2=f(x). Find xas a function of time and
show that ultimately the par ticle has travelled a distance x0.
15.33 Solve
2yd3y
dx3+2
/
y+3dy
dx
/d2y
dx2+2
/dy
dx
/2
=s i n x.
15.34 Find the general solution of the equation
xd3y
dx3+2d2y
dx2=Ax.
15.35 Express the equation
d2y
dx2+4xdy
dx+( 4x2+6 )y=e−x2sin2x
in canonical form and hence find its general solution.
15.36 Find the form of the solutions of the equation
dy
dxd3y
dx3−2
/d2y
dx2
/2
+
/dy
dx
/2
=0
which have y(0) =∞.
(You will need the result
Rzcosech ud u=−ln(cosech z+c o t h z).)
534
15.5 HINTS AND ANSWERS
15.37 Consider the equation
xpy/prime/prime+n+3−2p
n−1xp−1y/prime+
/p−2
n−1
/2
xp−2y=yn,
in which p/negationslash=2a n d n>−1 but n/negationslash= 1. For the boundary conditions y(1) = 0 and
y/prime(1) = λ, show that the solution is y(x)=v(x)x(p−2)/(n−1),w h e r e v(x)i sg i v e nb yZv(x)
0dz/
λ2+2zn+1/(n+1 )
/1/2=l nx.
15.5 Hints and answers
15.1 The function is ( ω2
0−ω2)−1(cosωt−cosω0t); for moderate t,x(t) is a sine wave
of linearly increasing amplitude ( tsinω0t)/(2ω0); for large tit shows beats of
maximum amplitude 2( ω2
0−ω2)−1.
15.2 m=−1±2i;( a ) f(t)=e−t(cos2 t+1
2sin2t); (b) f(t)=1
5e−t(cos2 t−cos 3t).
15.3 y=0a t x= 0. From symmetry, dy/dx =0a t x=L/2.
15.4 f(t)=1
4{e−t+[ ( 4 λ−2)t−1]e−3t}. For turning points, (4 λ+1 )+( 6−12λ)t=e2t.
(a) 1, (b) 2.
15.5 General solution f(t)= Ae−6t+Be−2t−3e−4t. (a) No solution, inconsistent
boundary conditions; (b) f(t)=2 e−6t+e−2t−3e−4t.
15.6 Set y5(x)=y1(x)+y2(x)a n d y6(x)=y1(x)−y2(x). Wronskian W(y3,y4,y5,y6)=
−16(α−1)(β+ 1). Thus linear dependence if α=1 ,o r β=−1, or both.
15.7 The auxiliary equation has repeated roots and the RHS is contained in the
complementary function. The solution is y(x)=(A+Bx)e−x+2x2e−x.y(2) = 5 e−2.
15.8 x=2s i n2 t+3c o s t, y=2c o s2 t−sint. The curve is symmetric about the y-axis
and crosses each axis four times. Its outer perimeter is heart-shaped.
15.9 (a) The auxiliary equation has roots 2, 2, −4; (A+Bx)exp2 x+Cexp(−4x)+2x+1;
(b) multiply through by sinh2 axand note thatR
cosech 2 ax dx =( 2a)−1ln(|tanh ax|);y=B(sinh2 ax)1/2(|tanh ax|)A.
15.10 (a) f(t)=2 e−3t−e−2t,( b ) f(t)=e−t(cos 2 t+1
2sin 2t); compare with exercise
15.2(a).
15.11 Use Laplace transforms; write s(s+n)−4as (s+n)−3−n(s+n)−4.
15.12 y=2x=2
3(cost−cos 2t), i.e. yis always a fixed multiple of x. There is no
resonance because the driving forces form the components, cos t,2c os t,o fa
vector that is a pure eigenvector corresponding to resonant frequency ω=2 ,a n d
contains no component of the eigenvector (1 ,−1) corresponding to ω=1 ,t h e
frequency of the forces.
15.13 L[C(t)]=x0(s+8 )/[s(s+2 ) ( s+ 4)], yielding
C(t)=x0[1 +1
2exp(−4t)−3
2exp(−2t)].
15.14 Write the numerator of the partial fraction with denominator ( s+γ)2+k2,w h e r e
k2=ω2
0−γ2,i nt h ef o r m A(s+γ)+B.
General solution is x(t)=(F/2ω0){γ−1[e−γtcoskt−cos(ω0t)] +k−1e−γtsinkt}.
(b) Since x=dx/dt =s i n ω0t=0a t t=0 , d2x/dt2= 0 also. Differentiating and
then setting t=0s h o w st h a t d3x/dt3has the initial value ω0F.
15.15 un=[ ( 1+√5)n−(1−√5)n]/(2n√5).
15.16 un=( 1−rn)/(1−rM)w h e r e r=(d+c)/(d−c). Ifc>d,t h e n r<−1.
15.17 P=5,Q=−4.un=3/2−5(−1)n/6+(−2)n/4+2n/12.
15.18 un= [35(−2)n−26(−3)n+2n]/10. Note that, with this recurrence relation and
these intial values, all unmust be integers.
15.19 The general solution is A+B2n/2exp(i3πn/4)+C2n/2exp(i5πn/4). The initial values
imply that A=1/5,B=(√5/10)exp[ i(π−φ)] and C=(√5/10)exp[ i(π+φ)].
535
HIGHER-ORDER ORDINARY DIFFERENTIAL EQUATIONS
15.20 The general solution is un=(A+Bn+Cn2)1n+(D+En)(−1)n+Fin+G(−i)n.
(a)B=C=E=0 , A=5 , D=−2,F=−3
2+5
2i,G=−3
2−5
2i;
(b)B=C= 1 and all other coefficients = 0.
15.21 This is Euler’s equation; setting x=e x p tproduces d2z/dt2−2dz/dt +z=e x p t
with complementary function ( A+Bt)e x p tand particular integral t2(expt)/2;
y(x)=x+[xlnx(1 + ln x)]/2.
15.22 This is Legendre’s linear equation with α=β= 1. Its reduced form is y/prime/prime+2y/prime+y=
(et−1)2,w h e r e y=y(t)a n d x+1= et.
A particular integral is y(t)=e2t/9−et/2 + 1 and the general solution is
y(x)=(x+1 )−1[A+Bln(x+1 ) ]+ x2/9−5x/18 + 11 /18.
15.23 After multiplication through by x2the coefficients are such that this is an
exact equation. The resulting first-order equation, in standard form, needs an
integrating factor ( x−2)2/x2.
15.24 (a) The complementary function is Aex+Be−x; writing the particular integral in
the form k1(x)ex+k2(x)e−xgives k/prime
1=xne−x/2a n d k/prime
2=−xnex/2. These lead to
the particular integral −(n!/2)
Pn
m=0[1 + (−1)n+m]xm/m!.
(b) Setting the particular integral equal to k1(x)ex+k2(x)xexgives the general
solution y=(A+Bx+x3/3)ex.
15.25 Given the boundary conditions, it is better to work with sinh xand sinh(1−x)
than with e±x;G(x, ξ)=−[sinh(1−ξ)sinh x]/sinh1 for x<ξ and−[sinh(1−
x)sinh ξ]/sinh1 for x>ξ.
15.26 (a) (i) (1 + x)u/prime/prime=( 2+ x)u/prime, (ii) follow subsection 15.3.6. Both give y2(x)=ex.
(b)y(x)=A/x+Bex−x/2−1.
15.27 Follow the method of subsection 15.2.5 but using general rather than specific
functions.
15.28 The relevant independent solutions are y1(x)=A(e−x−e−2x)a n d y2(x)=B(e−x−
e−2x+1) with Wronskian AB(e−1)e−3x.I fG1(x, ξ)=(e−1)−1(e−x−e−2x)(e2ξ−eξ+1)
and G2(x, ξ)=( e−1)−1(e−x−e−2x+1)(e2ξ−eξ)t h e n( a )f o r x<x 0,y(x)=R1
x0G1(x, ξ)dξ,a n d( b )f o r x>x 0,y(x)=
Rx
x0G2(x, ξ)dξ+
R1
xG1(x, ξ)dξ.
15.29 G(t, τ)=0f o r t<τ,a n d κ−1e−(t−τ)sin[κ(t−τ)] for t>τ. For a unit step input,
x(t)=( 1+ κ2)−1(1−e−tcosκt−κ−1e−tsinκt). Both transforms are equivalent to
s[(s+1 )2+κ2)]¯x=1.
15.30 With y=A(x)sin(x/2)+B(x)cos( x/2), obtain A/prime(z)=2 f(z)cos( z/2) and B/prime(z)=
−2f(z)sin(z/2) and hence identify G(x, z).
15.31 Use continuity and the step condition on ∂G/∂t att=t0to show that
G(t, t0)=α−1{1−exp[α(t0−t)]}for 0≤t0≤t;
x(t)=A(α−a)−1{a−1[1−exp(−at)]−α−1[1−exp(−αt)]}.
15.32 (a) B+x=
Rydz[A−2
Rzf(u)du]−1/2; (b) show that the force is proportional
to the derivative of ( x0−x)2ln[x0/(x0−x)];x=x0{1−exp[−At2/(2m)]}.
15.33 LHS of the equation is exact for two stages of integration and then needs an
integrating factor exp x;2yd2y/dx2+2yd y / d x +2 (dy/dx )2;2yd y / d x +y2=
d(y2)/dx+y2;y2=Aexp(−x)+Bx+C−(sinx−cosx)/2.
15.34 Set p=dy/dx ;y(x)=Ax3/18−Blnx+Cx+D.
15.35 Follow the method of subsection 15.2.6; u(x)=e−x2andv(x)s a t i s fi e s v/prime/prime+4v=
sin2x, for which a particular integral is ( −xcos 2x)/4. The general solution
y(x)=[Asin 2x+(B−1
4x)cos2 x]e−x2.
15.36 Set p=dy/dx and follow subsection 15.3.2 to obtain pd2p/dy2+1=( dp/dy )2
and then set q=dp/dy to obtain ( q2−1)1/2=Ap.The substitution sinh θ=Ap
gives finally that cosech ( Ay+B)+c o t h ( Ay+B)=e−x.
15.37 Equation is isobaric with yof weight m,w h e r e m+p−2=mn;v(x)s a t i s fi e s
x2v/prime/prime+xv/prime=vn.S e t x=etandv(x)=u(t), leading to u/prime/prime=unwith u(0) =
0,u/prime(0) = λ. Multiply both sides by u/primeto make the equation exact.
536
16
Series solutions of ordinary
differential equations
In the previous chapter the solution of both homogeneous and non-homogeneous
linear ordinary differential equations (ODEs) of order ≥2 were discussed. In par-
ticular we developed methods for solving some equations in which the coefficientswere not constant but functions of the independent variable x. In each case we
were able to write the solutions to such equations in terms of elementary func-tions, or as integrals. In general, however, the solutions of equations with variable
coefficients cannot be written in this way, and we must consider alternative
approaches.
In this chapter we discuss a method for obtaining solutions to linear ODEs
in the form of convergent series. Such series can be evaluated numerically, andthose occurring most commonly are named and tabulated. There is in fact nodistinct borderline between this and the previous chapter, since solutions in termsof elementary functions may equally well be written as convergent series (i.e. therelevant Taylor series). Indeed, it is partly because some series occur so frequently
that they are given special names such as sin x,c o sxor exp x.
Since we shall be concerned principally with second-order linear ODEs in this
chapter, we begin with a discussion of these equations, and obtain some general
results that will prove useful when we come to discuss series solutions.
16.1 Second-order linear ordinary differential equations
Any homogeneous second-order linear ODE can be written in the form
y
/prime/prime+p(x)y/prime+q(x)y=0, (16.1)
where y/prime=dy/dx andp(x)a n d q(x) are given functions of x. From the previous
chapter, we recall that the most general form of the solution to (16.1) is
y(x)=c1y1(x)+c2y2(x), (16.2)
537
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
where y1(x)a n d y2(x)a r elinearly independent solutions of (16.1), and c1andc2
are constants that are fixed by the boundary conditions (if supplied).
A full discussion of the linear independence of sets of functions was given
at the beginning of the previous chapter, but for just two functions y1andy2
to be linearly independent we simply require that y2is not a multiple of y1.
Equivalently, y1andy2must be such that the equation
c1y1(x)+c2y2(x)=0
isonlysatisfied for c1=c2= 0. Therefore the linear independence of y1(x)a n d
y2(x) can usually be deduced by inspection but in any case can always be verified
by the evaluation of the Wronskian of the two solutions,
W(x)=vextendsinglevextendsinglevextendsinglevextendsingley1y2
y/prime
1y/prime
2vextendsinglevextendsinglevextendsinglevextendsingle=y1y/prime
2−y2y/prime
1. (16.3)
IfW(x)/negationslash= 0 anywhere in a given interval then y1andy2are linearly independent
in that interval.
An alternative expression for W(x), of which we will make use later, may be
derived by differentiating (16.3) with respect to xto give
W/prime=y1y/prime/prime
2+y/prime
1y/prime
2−y2y/prime/prime
1−y/prime
2y/prime
1=y1y/prime/prime
2−y/prime/prime
1y2.
Since both y1andy2satisfy (16.1), we may substitute for y/prime/prime
1andy/prime/prime
2to obtain
W/prime=−y1(py/prime
2+qy2)+(py/prime
1+qy1)y2=−p(y1y/prime
2−y/prime
1y2)=−pW.
Integrating, we find
W(x)=Cexpbraceleftbigg
−integraldisplayx
p(u)dubracerightbigg
, (16.4)
where Cis a constant. We note further that in the special case p(x)≡0w eo b t a i n
W=c o n s t a n t .IThe functions y1=s i n xandy2=c o s xare both solutions of the equation y/prime/prime+y=
0. Evaluate the Wronskian of these two solutions, and hence show that they are linearly
independent.
The Wronskian of y1andy2is given by
W=y1y/prime
2−y2y/prime
1=−sin2x−cos2x=−1.
Since W/negationslash= 0 the two solutions are linearly independent. We also note that y/prime/prime+y=0i s
a special case of (16.1) with p(x) = 0. We therefore expect, from (16.4), that Wwill be a
constant, as is indeed the case.
J
From the previous chapter we recall that, once we have obtained the general
solution to the homogeneous second-order ODE (16.1) in the form (16.2), thegeneral solution to the inhomogeneous equation
y
/prime/prime+p(x)y/prime+q(x)y=f(x) (16.5)
538
16.1 SECOND-ORDER LINEAR ORDINARY DIFFERENTIAL EQUATIONS
can be written as the sum of the solution to the homogeneous equation yc(x)
(the complementary function) and anyfunction yp(x) (the particular integral) that
satisfies (16.5) and is linearly independent of yc(x). We have therefore
y(x)=c1y1(x)+c2y2(x)+yp(x). (16.6)
General methods for obtaining yp, that are applicable to equations with variable
coefficients, such as the variation of parameters or Green’s functions, were dis-
cussed in the previous chapter. An alternative description of the Green’s function
method for solving inhomogeneous equations is given in the next chapter. For thepresent, however, we will restrict our attention to the solutions of homogeneousODEs in the form of convergent series.
16.1.1 Ordinary and singular points of an ODE
So far we have implicitly assumed that y(x)i sa realfunction of a realvariable
x. However, this is not always the case, and in the remainder of this chapter we
broaden our discussion by generalising to a complex function y(z)o fa complex
variable z.
Let us therefore consider the second-order linear homogeneous ODE
y
/prime/prime+p(z)y/prime+q(z)=0 , (16.7)
where now y/prime=dy/dz ; this is a straightforward generalisation of (16.1). A full
discussion of complex functions and differentiation with respect to a complexvariable zis given in chapter 20, but for the purposes of the present chapter we
need not concern ourselves with many of the subtleties that exist. In particular,we may treat differentiation with respect to zin an way analogous to ordinary
differentiation with respect to a real variable x.
In (16.7), if at some point z=z
0the functions p(z)a n d q(z) are finite and can
be expressed as complex power series (see section 4.5)
p(z)=∞summationdisplay
n=0pn(z−z0)n,q (z)=∞summationdisplay
n=0qn(z−z0)n
then p(z)a n d q(z) are said to be analytic atz=z0, and this point is called an
ordinary point of the ODE. If, however, p(z)o rq(z), or both, diverge at z=z0
then it is called a singular point of the ODE.
Even if an ODE is singular at a given point z=z0, it may still possess a
non-singular (finite) solution at that point. In fact the necessary and sufficientcondition†for such a solution to exist is that ( z−z
0)p(z)a n d( z−z0)2q(z) are both
analytic at z=z0. Singular points that have this property are regular singular
†See, for example, Jeffreys and Jeffreys, Mathematical Methods of Physics ,3 r de d .( C a m b r i d g e
University Press, 1966), p. 479.
539
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
points, whereas any singular point not satisfying both these criteria is termed an
irregular oressential singularity.ILegendre’s equation has the form
(1−z2)y/prime/prime−2zy/prime+/lscript(/lscript+1 )y=0, (16.8)
where /lscriptis a constant. Show that z=0is an ordinary point and z=±1are regular singular
points of this equation.
Firstly, divide through by 1 −z2to put the equation into our standard form (16.7):
y/prime/prime−2z
1−z2y/prime+/lscript(/lscript+1 )
1−z2y=0.
Comparing this with (16.7), we identify p(z)a n d q(z)a s
p(z)=−2z
1−z2=−2z
(1 +z)(1−z),q (z)=/lscript(/lscript+1 )
1−z2=/lscript(/lscript+1 )
(1 +z)(1−z).
By inspection, p(z)a n d q(z) are analytic at z= 0, which is therefore an ordinary point,
but both diverge for z=±1, which are thus singular points. However, at z=1w es e e
that both ( z−1)p(z)a n d( z−1)2q(z) are analytic and hence z= 1 is a regular singular
point. Similarly, at z=−1 both ( z+1 )p(z)a n d( z+1 )2q(z) are analytic, and it too is a
regular singular point.
J
So far we have assumed that z0is finite. However, we may sometimes wish to
determine the nature of the point |z|→∞ . This may be achieved straightforwardly
by substituting w=1/zinto the equation and investigating the behaviour at
w=0 .IShow that Legendre’s equation has a regular singularity at |z|→∞ .
Letting w=1/z, the derivatives with respect to zbecome
dy
dz=dy
dwdw
dz=−1
z2dy
dw=−w2dy
dw,
d2y
dz2=dw
dzd
dw
/dy
dz
/
=−w2
/
−2wdy
dw−w2d2y
dw2
/
=w3
/
2dy
dw+wd2y
dw2
/
.
If we substitute these derivatives into Legendre’s equation (16.8) we obtain/
1−1
w2
/
w3
/
2dy
dw+wd2y
dw2
/
+21
ww2dy
dw+/lscript(/lscript+1 )y=0,
which simplifies to give
w2(w2−1)d2y
dw2+2w3dy
dw+/lscript(/lscript+1 )y=0.
Dividing through by w2(w2−1) to put the equation into standard form, and comparing
with (16.7), we identify p(w)a n d q(w)a s
p(w)=2w
w2−1,q (w)=/lscript(/lscript+1 )
w2(w2−1).
Atw=0 , p(w) is analytic but q(w) diverges, and so the point |z|→∞ is a singular point
of Legendre’s equation. However, since wpandw2qare both analytic at w=0 ,|z|→∞
is a regular singular point.
J
540
16.2 SERIES SOLUTIONS ABOUT AN ORDINARY POINT
Equation Regular singularities Essential singularities
Legendre∗
(1−z2)y/prime/prime−2zy/prime+/lscript(/lscript+1 )y=0 −1,1,∞ —
Chebyshev
(1−z2)y/prime/prime−zy/prime+n2y=0 −1,1,∞ —
Bessel
z2y/prime/prime+zy+(z2−ν2)y=0 0 ∞
Laguerre∗
zy/prime/prime+( 1−z)y/prime+αy=0 0 ∞
Simple harmonic oscillator
y/prime/prime+ω2y=0 — ∞
Hermite
y/prime/prime−2zy/prime+2αy=0 — ∞
Table 16.1 Important ODEs in the physical sciences and engineering. The
asterisks indicate that the corresponding associated equations (discussed in the
next chapter) have the same singular points.
Table 16.1 lists the singular points of several second-order linear ODEs that
play important roles in the analysis of many physics and engineering problems.In sections 16.6 and 16.7 we consider the the solution of Legendre’s and Bessel’sequations in terms of convergent series and discuss some useful properties ofthese solutions. The solutions of the remaining equations in table 16.1 may alsobe found in the form of convergent series, but a discussion of these solutions andtheir properties is left until the next chapter, where they are considered in the
context of Sturm–Liouville systems. We now discuss the methods by which series
solutions may be obtained.
16.2 Series solutions about an ordinary point
Ifz=z
0is an ordinary point of (16.7) then it may be shown that everysolution
y(z) of the equation is also analytic at z=z0. In our subsequent discussion
we will take z0as the origin, i.e. z0= 0. If this is not already the case, then a
substitution Z=z−z0will make it so. Since every solution is analytic, y(z)c a n
be represented by a power series of the form (see section 20.13)
y(z)=∞summationdisplay
n=0anzn. (16.9)
Moreover, it may be shown that such a power series converges for |z|<R,w h e r e
Ris the radius of convergence and is equal to the distance from z=0t ot h e
nearest singular point of the ODE (see chapter 20). At the radius of convergence,however, the series may or may not converge (as shown in section 4.5).
541
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
Since every solution of (16.7) is analytic at an ordinary point, it is always
possible to obtain two independent solutions (from which the general solution
(16.2) can be constructed) of the form (16.9). The derivatives of ywith respect to
zare given by
y/prime=∞summationdisplay
n=0nanzn−1=∞summationdisplay
n=0(n+1 )an+1zn, (16.10)
y/prime/prime=∞summationdisplay
n=0n(n−1)anzn−2=∞summationdisplay
n=0(n+2 ) ( n+1 )an+2zn. (16.11)
Note that, in each case, in the first equality the sum can still start at n=0s i n c e
the first term in (16.10) and the first two terms in (16.11) are automatically zero.The second equality in each case is obtained by shifting the summation index sothat the sum can be written in terms of coefficients of z
n. By substituting (16.9)–
(16.11) into the ODE (16.7), and requiring that the coefficients of each power of
zs u mt oz e r o ,w eo b t a i na recurrence relation expressing anas a function of the
previous ar(0≤r≤n−1).IFind the series solutions, about z=0,o f
y/prime/prime+y=0.
By inspection z= 0 is an ordinary point of the equation, and so we may obtain two
independent solutions by making the substitution y=
P∞
n=0anzn. Using (16.9) and (16.11)
we find
∞X
n=0(n+2 ) ( n+1 )an+2zn+∞X
n=0anzn=0,
which may be written as
∞X
n=0[(n+2 ) ( n+1 )an+2+an]zn=0.
For this equation to be satisfied we require that the coefficient of each power of zvanishes
separately , and so we obtain the two-term recurrence relation
an+2=−an
(n+2 ) ( n+1 )forn≥0.
Using this relation, we can calculate, say, the even coefficients a2,a4,a6a n ds oo n ,f o r
ag i v e n a0. Alternatively, starting with a1, we obtain the odd coefficients a3,a5etc. Two
independent solutions of the ODE can be obtained by setting either a0=0o r a1=0 .
Firstly if we set a1= 0 and choose a0= 1 then we obtain the solution
y1(z)=1−z2
2!+z4
4!−···=∞X
n=0(−1)n
(2n)!z2n.
Secondly, if we set a0= 0 and choose a1= 1 then we obtain a second, independent , solution
y2(z)=z−z3
3!+z5
5!−···=∞X
n=0(−1)n
(2n+1 ) !z2n+1.
542
16.2 SERIES SOLUTIONS ABOUT AN ORDINARY POINT
Recognising these two series as cos zand sin z, we can write the general solution as
y(z)=c1cosz+c2sinz,
where c1andc2are arbitrary constants that are fixed by boundary conditions (if supplied).
We note that both solutions converge for all z, as might be expected since the ODE
possesses no singular points (except |z|→∞ ).
J
Solving the above example was quite straightforward and the resulting series
were easily recognised and written in closed form (i.e. in terms of elementary
functions); this is not usually the case . Another simplifying feature of the previous
example was that we obtained a two-term recurrence relation relating an+2and
an, so that the odd- and even-numbered coefficients were independent of one
another. In general the recurrence relation expresses anas a function of any
number of the previous ar(0≤r≤n−1).IFind the series solutions, about z=0,o f
y/prime/prime−2
(1−z)2y=0.
By inspection z= 0 is an ordinary point, and therefore we may find two independent
solutions by substituting y=
P∞
n=0anzn. Using (16.10) and (16.11), and multiplying through
by (1−z)2, we find
(1−2z+z2)∞X
n=0n(n−1)anzn−2−2∞X
n=0anzn=0,
which leads to
∞X
n=0n(n−1)anzn−2−2∞X
n=0n(n−1)anzn−1+∞X
n=0n(n−1)anzn−2∞X
n=0anzn=0.
In order to write all these series in terms of the coefficients of zn, we must shift the
summation index in the first two sums, obtaining
∞X
n=0(n+2 ) ( n+1 )an+2zn−2∞X
n=0(n+1 )nan+1zn+∞X
n=0(n2−n−2)anzn=0,
which can be written as
∞X
n=0(n+1 ) [ ( n+2 )an+2−2nan+1+(n−2)an]zn=0.
By demanding that the coefficients of each power of zvanish separately, we obtain the
three-term recurrence relation
(n+2 )an+2−2nan+1+(n−2)an=0 f o r n≥0,
which determines anforn≥2i nt e r m so f a0anda1. Three-term (or more) recurrence
relations are a nuisance and, in general, can be difficult to solve. This particular recurrencerelation, however, has two straightf orward solutions. One solution is a
n=a0for all n,i n
which case (choosing a0= 1) we find
y1(z)=1+ z+z2+z3+···=1
1−z.
543
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
The other solution to the recurrence relation is a1=−2a0,a2=a0andan=0f o r n>2,
so that (again choosing a0=1 )w eo b t a i na polynomial solution to the ODE:
y2(z)=1−2z+z2=( 1−z)2.
The linear independence of y1andy2is obvious but can be checked by computing the
Wronskian
W=y1y/prime
2−y/prime
1y2=1
1−z[−2(1−z)]−1
(1−z)2(1−z)2=−3.
Since W/negationslash= 0 the two solutions y1andy2are indeed linearly independent. The general
solution of the ODE is therefore
y(z)=c1
1−z+c2(1−z)2.
We observe that y1(and hence the general solution) is singular at z=1 ,w h i c hi st h e
singular point of the ODE nearest to z= 0, but the polynomial solution y2is valid for all
finite z.
J
The above example illustrates the possibility that, in some cases, we may find
that the recurrence relation leads to an=0f o r n>N , for one or both of the
two solutions; we then obtain a polynomial solution to the equation. Polynomial
solutions are discussed more fully in section 16.5, but one obvious property ofsuch solutions is that they converge for all finite z. By contrast, as mentioned
above, for solutions in the form of an infinite series the circle of convergence
extends only as far as the singular point nearest to that about which the solutionis being obtained.
16.3 Series solutions about a regular singular point
From table 16.1 we see that several of the most important second-order linear
ODEs in physics and engineering have regular singular points in the finite complex
plane. We must extend our discussion, therefore, to obtaining series solutions toODEs about such points. In what follows we assume that the regular singularpoint about which the solution is required is at z= 0, since, as we have seen, if
this is not already the case then a substitution of the form Z=z−z
0will make
it so.
Ifz= 0 is a regular singular point of the equation
y/prime/prime+p(z)y/prime+q(z)y=0
then p(z)a n d q(z) are not analytic at z= 0, and in general we should not expect
to find a power series solution of the form (16.9). We must therefore extend the
method to include a more general form for the solution. In fact it may be shown
(Fuch’s theorem) that there exists at least one solution to the above equation, of
the form
y=zσ∞summationdisplay
n=0anzn, (16.12)
544
16.3 SERIES SOLUTIONS ABOUT A REGULAR SINGULAR POINT
where the exponent σis a number that may be real or complex and where a0/negationslash=0
(since, if it were otherwise, σcould be redefined as σ+1o r σ+2o r ···so as to
make a0/negationslash= 0). Such a series is called a generalised power series or Frobenius series .
As in the case of a simple power series solution, the radius of convergence of the
Frobenius series is, in general, equal to the distance to the nearest singularity of
the ODE.
Since z= 0 is a regular singularity of the ODE, it follows that zp(z)a n d z2q(z)
a r ea n a l y t i ca t z= 0, so that we may write
zp(z)≡s(z)=∞summationdisplay
n=0snzn
z2q(z)≡t(z)=∞summationdisplay
n=0tnzn,
where we have defined the analytic functions s(z)a n d t(z) for later convenience.
The original ODE therefore becomes
y/prime/prime+s(z)
zy/prime+t(z)
z2y=0.
Let us substitute the Frobenius series (16.12) into this equation. The derivatives
of (16.12) with respect to xare given by
y/prime=∞summationdisplay
n=0(n+σ)anzn+σ−1, (16.13)
y/prime/prime=∞summationdisplay
n=0(n+σ)(n+σ−1)anzn+σ−2, (16.14)
and we obtain
∞summationdisplay
n=0(n+σ)(n+σ−1)anzn+σ−2+s(z)∞summationdisplay
n=0(n+σ)anzn+σ−2+t(z)∞summationdisplay
n=0anzn+σ−2=0.
Dividing this equation through by zσ−2we find
∞summationdisplay
n=0[(n+σ)(n+σ−1) +s(z)(n+σ)+t(z)]anzn=0. (16.15)
Setting z= 0, all terms in the sum with n>0 vanish, implying that
[σ(σ−1) +s(0)σ+t(0)]a0=0,
which, since we require a0/negationslash= 0, yields the indicial equation
σ(σ−1) +s(0)σ+t(0) = 0 . (16.16)
This equation is a quadratic in σand in general has two roots, the nature of
which determines the forms of possible series solutions.
545
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
The two roots of the indicial equation σ1andσ2are called the indices of the
regular singular point. By substituting each of these roots into (16.15) in turn andrequiring that the coefficients of each power of zvanish separately, we obtain a
recurrence relation (for each root) expressing each a
nas a function of the previous
ar(0≤r≤n−1). Depending on the roots of the indicial equation σ1andσ2,
there are three possible general cases, which we now discuss.
16.3.1 Distinct roots not differing by an integer
If the roots of the indicial equation σ1andσ2differ by an amount that is not
an integer then the recurrence relations corresponding to each root lead to twolinearly independent solutions of the ODE,
y
1(z)=zσ1∞summationdisplay
n=0anzn,y 2(z)=zσ2∞summationdisplay
n=0bnzn.
The linear independence of these two solutions follows from the fact that y2/y1
is not a constant since σ1−σ2is not an integer. Because y1andy2are linearly
independent, we may use them to construct the general solution y=c1y1+c2y2.
We also note that this case includes complex conjugate roots where σ2=σ∗
1,
since σ1−σ2=σ1−σ∗
1=2iImσ1cannot be equal to a real integer.IFind the power series solutions about z=0of
4zy/prime/prime+2y/prime+y=0.
Dividing through by 4 zto put the equation into standard form, we obtain
y/prime/prime+1
2zy/prime+1
4zy=0, (16.17)
and on comparing with (16.7) we identify p(z)=1 /(2z)a n d q(z)=1 /(4z). Clearly z=0
is a singular point of (16.17), but since zp(z)=1 /2a n d z2q(z)=z/4 are finite there, it
is a regular singular point. We therefore substitute the Frobenius series y=zσ
P∞
n=0anzn
into (16.17). Using (16.13) and (16.14), we obtain
∞X
n=0(n+σ)(n+σ−1)anzn+σ−2+1
2z∞X
n=0(n+σ)anzn+σ−1+1
4z∞X
n=0anzn+σ=0,
which on dividing through by zσ−2gives
∞X
n=0
/
(n+σ)(n+σ−1) +1
2(n+σ)+1
4z
/
anzn=0. (16.18)
If we set z= 0 then all terms in the sum with n>0 vanish, and we obtain the indicial
equation
σ(σ−1) +1
2σ=0,
which has roots σ=1/2a n d σ= 0. Since these roots do not differ by an integer we expect
to find two independent solutions to (16.17), in the form of Frobenius series.
546
16.3 SERIES SOLUTIONS ABOUT A REGULAR SINGULAR POINT
Demanding that the coefficients of znvanish separately in (16.18), we obtain the
recurrence relation
(n+σ)(n+σ−1)an+1
2(n+σ)an+1
4an−1=0. (16.19)
If we choose the larger root, σ=1/2, of the indicial equation then (16.19) becomes
(4n2+2n)an+an−1=0⇒ an=−an−1
2n(2n+1 ).
Setting a0= 1 we find an=(−1)n/(2n+ 1)! and so the solution to (16.17) is
y1(z)=√z∞X
n=0(−1)n
(2n+1 ) !zn
=√z−(√z)3
3!+(√z)5
5!−···=s i n√z.
To obtain the second solution we set σ= 0 (the smaller root of the indicial equation) in
(16.19), which gives
(4n2−2n)an+an−1=0⇒ an=−an−1
2n(2n−1).
Setting a0=1n o wg i v e s an=(−1)n/(2n)!, and so the second (independent) solution to
(16.17) is
y2(z)=∞X
n=0(−1)n
(2n)!zn=1−(√z)2
2!+(√
4)4
4!−···=c o s√z.
We may check that y1(z)a n d y2(z) are indeed linearly independent by computing the
Wronskian
W=y1y/prime
2−y2y/prime
1
=s i n√z
/
−1
2√zsin√z
/
−cos√z
/1
2√zcos√z
/
=−1
2√z
/;
sin2√z+c o s2√z
/
=−1
2√z/negationslash=0.
Since W/negationslash= 0 the solutions y1(z)a n d y2(z) are linearly independent. Hence the general
solution to (16.17) is given by
y(z)=c1sin√z+c2cos√z.
J
16.3.2 Repeated root of the indicial equation
If the indicial equation has a repeated root, so that σ1=σ2=σ, then obviously
only one solution in the form of a Frobenius series (16.12) may be found asdescribed above, i.e.
y
1(z)=zσ∞summationdisplay
n=0anzn.
Methods for obtaining a second, linearly independent, solution are discussed in
section 16.4.
547
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
16.3.3 Distinct roots differing by an integer
Whatever the roots of the indicial equation, the recurrence relation corresponding
to the larger of the two always leads to a solution of the ODE. However, if theroots of the indicial equation differ by an integer then the recurrence relationcorresponding to the smaller root may or may not lead to a second linearlyindependent solution, depending on the ODE under consideration. Note that for
complex roots of the indicial equation, the ‘larger’ root is taken to be the one
with the larger real part.IFind the power series solutions about z=0of
z(z−1)y/prime/prime+3zy/prime+y=0. (16.20)
Dividing through by z(z−1) to put the equation into standard form, we obtain
y/prime/prime+3
(z−1)y/prime+1
z(z−1)y=0, (16.21)
and on comparing with (16.7) we identify p(z)=3 /(z−1) and q(z)=1 /[z(z−1)]. We
immediately see that z= 0 is a singular point of (16.21), but since zp(z)=3 z/(z−1) and
z2q(z)=z/(z−1) are finite there, it is a regular singular point and we expect to find at least
one solution in the form of a Frobenius series. We therefore substitute y=zσ
P∞
n=0anzn
into (16.21), and using (16.13) and (16.14), we obtain
∞X
n=0(n+σ)(n+σ−1)anzn+σ−2+3
z−1∞X
n=0(n+σ)anzn+σ−1
+1
z(z−1)∞X
n=0anzn+σ=0,
which on dividing through by zσ−2gives
∞X
n=0
/
(n+σ)(n+σ−1) +3z
z−1(n+σ)+z
z−1
/
anzn=0.
Although we could use this expression to find the indicial equation and recurrence relations,
the working is simpler if we now multiply through by z−1t og i v e
∞X
n=0[(z−1)(n+σ)(n+σ−1) + 3 z(n+σ)+z]anzn=0. (16.22)
If we set z= 0 then all terms in the sum with the exponent of zgreater than zero vanish,
and we obtain the indicial equation
σ(σ−1) = 0 ,
which has the roots σ=1a n d σ= 0. Since the roots differ by an integer (unity), it may not
be possible to find two linearly independent solutions of (16.21) in the form of Frobeniusseries. We are guaranteed, however, to find one such solution corresponding to the largerroot, σ=1 .
Demanding that the coefficients of z
nvanish separately in (16.22), we obtain the
recurrence relation
(n−1+σ)(n−2+σ)an−1−(n+σ)(n+σ−1)an+3 (n−1+σ)an−1+an−1=0,
548
16.4 OBTAINING A SECOND SOLUTION
which can be simplified to give
(n+σ−1)an=(n+σ)an−1. (16.23)
Substituting σ= 1 into this expression, we obtain
an=
/n+1
n
/
an−1,
and setting a0= 1 we find an=n+ 1; so one solution to (16.21) is
y1(z)=z∞X
n=0(n+1 )zn=z(1 + 2 z+3z2+···)
=z
(1−z)2. (16.24)
If we attempt to find a second solution (corresponding to the smaller root of the indicial
equation) by setting σ= 0 in (16.23), we find
an=
/n
n−1
/
an−1,
but we require a0/negationslash=0 ,s o a1is formally infinite and the method fails. We discuss how to
find a second linearly independent solution in the next section.
J
One particular case is also worth mentioning. If the point about which the
solution is required, i.e. z= 0, is in fact an ordinary point of the ODE rather than
a regular singular point, then substitution of the Frobenius series (16.12) leads toan indicial equation with roots σ=0a n d σ= 1. Although these roots differ by
an integer (unity), the recurrence relations corresponding to the two roots yieldtwo linearly independent power series solutions (one for each root), as expectedfrom section 16.2.
It is always worth investigating whether a series found as a solution to a
problem is summable in closed form or expressible in terms of known functions.Nevertheless, the reader should avoid gaining the impression that this is always
so or that, if one worked hard enough, a closed-form solution could always be
found without using the series method. As mentioned earlier, this is notthe case,
and very often an infinite series solution is the best one can do.
16.4 Obtaining a second solution
Whilst attempting to find a solution to an ODE in the form of a Frobenius series
about a regular singular point, we found in the previous section that when theindicial equation has a repeated root, or roots differing by an integer, we can (ingeneral) find only one solution of this form. In order to construct the general
solution to the ODE, however, we require two linearly independent solutions y
1
andy2. We now consider several methods for obtaining a second solution in this
case.
549
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
16.4.1 The Wronskian method
Ify1andy2are two linearly independent solutions of the standard equation
y/prime/prime+p(z)y/prime+q(z)y=0
then the Wronskian of these two solutions is given by W(z)= y1y/prime
2−y2y/prime
1.
Dividing the Wronskian by y2
1we obtain
W
y2
1=y/prime
2
y1−y/prime
1
y2
1y2=y/prime
2
y1+bracketleftbiggd
dzparenleftbigg1
y1parenrightbiggbracketrightbigg
y2=d
dzparenleftbiggy2
y1parenrightbigg
,
which integrates to give
y2(z)=y1(z)integraldisplayzW(u)
y2
1(u)du.
Now using the alternative expression for W(z) given in (16.4) with C=1( s i n c e
we are not concerned with this normalising factor), we find
y2(z)=y1(z)integraldisplayz1
y2
1(u)expbraceleftbigg
−integraldisplayu
p(v)dvbracerightbigg
du. (16.25)
Hence, given y1, we can in principle compute y2. Note that the lower limits of
integration have been omitted. If constant lower limits are included then theymerely lead to a constant times the first solution.IFind a second solution to (16.21) using the Wronskian method.
For the ODE (16.21) we have p(z)=3 /(z−1), and from (16.24) we see that one solution
to (16.21) is y1=z/(1−z)2. Substituting for pandy1in (16.25) we have
y2(z)=z
(1−z)2
Zz(1−u)4
u2exp
/
−
Zu3
v−1dv
/
du
=z
(1−z)2
Zz(1−u)4
u2exp[−3ln(u−1)]du
=z
(1−z)2
Zzu−1
u2du
=z
(1−z)2
/
lnz+1
z
/
.
By calculating the Wronskian of y1andy2it is easily shown that, as expected, the two
solutions are linearly independent. In fact, as the Wronskian has already been evaluated
asW(u)=e x p [−3ln(u−1)], i.e. W(z)=(z−1)−3, no calculation is needed.
J
An alternative (but equivalent) method of finding a second solution is simply to
assume that the second solution has the form y2(z)=u(z)y1(z) for some function
u(z) to be determined (this method was discussed more fully in subsection 15.2.3).
From (16.25), we see that the second solution derived from the Wronskian is
indeed of this form. Substituting y2(z)= u(z)y1(z) into the ODE leads to a
first-order ODE in which u/primeis the dependent variable; this may then be solved.
550
16.4 OBTAINING A SECOND SOLUTION
16.4.2 The derivative method
The derivative method of finding a second solution begins with the derivation of
a recurrence relation for the coefficients anin a Frobenius series solution, as in the
previous section. However, rather than putting σ=σ1in this recurrence relation
to evaluate the first series solution, we now keep σas a variable parameter. This
means that the computed anare functions of σand the computed solution is now
a function of zandσ:
y(z,σ)=zσ∞summationdisplay
n=0an(σ)zn. (16.26)
Of course, if we put σ=σ1in this, we obtain immediately the first series solution,
but for the moment we leave σas a parameter.
For brevity let us denote the differential operator on the LHS of our standard
ODE (16.7) by L,s ot h a t
L=d2
dz2+p(z)d
dz+q(z),
and examine the effect of Lon the series y(z,σ) in (16.26). It is clear that the
series Ly(z,σ) will contain only a term in zσ, since the recurrence relation defining
thean(σ) is such that these coefficients vanish for higher powers of z.B u tt h e
coefficient of zσis simply the LHS of the indicial equation. Therefore, if the roots
of the indicial equation are σ=σ1andσ=σ2then it follows that
Ly(z,σ)=a0(σ−σ1)(σ−σ2)zσ. (16.27)
Therefore, as in the previous section, we see that for y(z,σ)t ob eas o l u t i o no f
the ODE Ly=0 , σmust equal σ1orσ2. For simplicity we shall set a0= 1 in the
following discussion.
Let us first consider the case in which the two roots of the indicial equation
are equal, i.e. σ2=σ1. From (16.27) we then have
Ly(z,σ)=(σ−σ1)2zσ.
Differentiating this equation with respect to σwe obtain
∂
∂σ[Ly(z,σ)]=(σ−σ1)2zσlnz+2 (σ−σ1)zσ,
which equals zero if σ=σ1. But since ∂/∂σandLare operators that differentiate
with respect to different variables we can reverse their order, implying that
Lbracketleftbigg∂
∂σy(z,σ)bracketrightbigg
=0 a t σ=σ1.
Hence the function in square brackets, evaluated at σ=σ1and denoted by
bracketleftbigg∂
∂σy(z,σ)bracketrightbigg
σ=σ1, (16.28)
551
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
is also a solution of the original ODE Ly= 0, and is in fact the second linearly
independent solution for which we were looking.
The case in which the roots of the indicial equation differ by an integer is
slightly more complicated but can be treated in a similar way. In (16.27), since L
differentiates with respect to zwe may multiply (16.27) by any function of σ,s a y
σ−σ2, and take this function inside the operator Lon the LHS to obtain
L[(σ−σ2)y(z,σ)]=(σ−σ1)(σ−σ2)2zσ. (16.29)
Therefore the function
[(σ−σ2)y(z,σ)]σ=σ2
is also a solution of the ODE Ly= 0. However, it can be proved †that this
function is a simple multiple of the first solution y(z,σ1), showing that it is not
linearly independent and that we must find another solution. To do this we
differentiate (16.29) with respect to σand find
∂
∂σ{L[(σ−σ2)y(z,σ)]}=(σ−σ2)2zσ+2 (σ−σ1)(σ−σ2)zσ
+(σ−σ1)(σ−σ2)2zσlnz,
which is equal to zero if σ=σ2. As previously, since ∂/∂σ andLare operators
that differentiate with respect to different variables, we can reverse their order toobtain
Lbraceleftbigg∂
∂σ[(σ−σ2)y(z,σ)]bracerightbigg
=0 a t σ=σ2,
and so the function
braceleftbigg∂
∂σ[(σ−σ2)y(z,σ)]bracerightbigg
σ=σ2(16.30)
is also a solution of the original ODE Ly= 0, and is in fact the second linearly
independent solution.IFind a second solution to (16.21) using the derivative method.
From (16.23) the recurrence relation (with σas a parameter) is given by
(n+σ−1)an=(n+σ)an−1.
Setting a0= 1 we find that the cofficients have the particularly simple form an(σ)=
(σ+n)/σ. We therefore consider the function
y(z,σ)=zσ∞X
n=0an(σ)zn=zσ∞X
n=0σ+n
σzn.
†For a fuller discussion see, for example, Riley, Mathematical Methods for the Physical Sciences ,
(Cambridge University Press, 1974), pp. 158–9.
552
16.4 OBTAINING A SECOND SOLUTION
The smaller root of the indicial equation for (16.21) is σ2= 0, and so from (16.30) a
second, linearly independent, solution to the ODE is/∂
∂σ[σy(z,σ)]
/
σ=0=
/(
∂
∂σ
/"
zσ∞X
n=0(σ+n)zn
/# /)
σ=0.
The derivative with respect to σis given by
∂
∂σ
/"
zσ∞X
n=0(σ+n)zn
/#
=zσlnz∞X
n=0(σ+n)zn+zσ∞X
n=0zn,
which on setting σ= 0 gives the second solution
y2(z)=l n z∞X
n=0nzn+∞X
n=0zn
=z
(1−z)2lnz+1
1−z
=z
(1−z)2
/
lnz+1
z−1
/
.
This second solution is the same as that obtained by the Wronskian method in the previous
subsection except for the addition of some of the first solution.
J
16.4.3 Series form of the second solution
Using any of the methods discussed above, we can find the general form of the
second solution to the ODE. This form is most easily found, however, using thederivative method. Let us first consider the case where the two solutions of theindicial equation are equal. In this case a second solution is given by (16.28),which may be written as
y
2(z)=bracketleftbigg∂y(z,σ)
∂σbracketrightbigg
σ=σ1
=( l n z)zσ1∞summationdisplay
n=0an(σ1)zn+zσ1∞summationdisplay
n=1bracketleftbiggdan(σ)
dσbracketrightbigg
σ=σ1zn
=y1(z)l nz+zσ1∞summationdisplay
n=1bnzn,
where bn=[dan(σ)/dσ]σ=σ1.
In the case where the roots of the indicial equation differ by an integer (not
equal to zero), then from (16.30) a second solution is given by
y2(z)=braceleftbigg∂
∂σ[(σ−σ2)y(z,σ)]bracerightbigg
σ=σ2
=l nzbracketleftBigg
(σ−σ2)zσ∞summationdisplay
n=0an(σ)znbracketrightBigg
σ=σ2+zσ2∞summationdisplay
n=0bracketleftbiggd
dσ(σ−σ2)an(σ)bracketrightbigg
σ=σ2zn.
553
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
But, as we mentioned in the previous section, [(σ−σ2)y(z,σ)]atσ=σ2is just a
multiple of the first solution y(z,σ1). Therefore the second solution is of the form
y2(z)=cy1(z)l nz+zσ2∞summationdisplay
n=0bnzn,
where cis a constant. In some cases, however, cmight be zero and so the second
solution would not contain the term in ln zand could be written simply as a
Frobenius series. Clearly this corresponds to the case in which the substitution ofa Frobenius series into the original ODE yields two solutions automatically.
16.5 Polynomial solutions
We have seen that the evaluation of successive terms of a series solution to a
differential equation is carried out by means of a recurrence relation. The formof the relation for a
ndepends upon n, the previous values of ar(r<n)a n dt h e
parameters of the equation. It may happen, as a result of this, that for some
value of n=N+ 1 the computed value aN+1is zero and that all higher aralso
vanish. If this is so, and the corresponding solution of the indicial equation σ
is a positive integer or zero, then we are left with a finite polynomial of degreeN
/prime=N+σas a solution of the ODE:
y(z)=Nsummationdisplay
n=0anzn+σ. (16.31)
In many applications in theoretical physics (particularly in quantum mechanics)
the termination of a potentially infinite series after a finite number of termsis of crucial importance in establishing physically acceptable descriptions andproperties of systems. The condition under which such a termination occurs istherefore of considerable importance.IFind power series solutions about z=0of
y/prime/prime−2zy/prime+λy=0. (16.32)
For what values of λdoes the equation possess a polynomial solution? Find such a solution
forλ=4.
Clearly z= 0 is an ordinary point of (16.32) and so we look for solutions of the form
y=
P∞
n=0anzn. Substituting this into the ODE and multiplying through by z2we find
∞X
n=0[n(n−1)−2z2n+λz2]anzn=0.
By demanding that the coefficients of each power of zvanish separately we derive the
recurrence relation
n(n−1)an−2(n−2)an−2+λan−2=0,
554
16.6 LEGENDRE’S EQUATION
which may be rearranged to give
an=2(n−2)−λ
n(n−1)an−2forn≥2. (16.33)
The odd and even coefficients are therefore independent of one another, and two solutions
to (16.32) may be derived. We either set a1=0a n d a0=1t oo b t a i n
y1(z)=1−λz2
2!−λ(4−λ)z4
4!−λ(4−λ)(8−λ)z6
6!−··· (16.34)
or set a0=0a n d a1=1t oo b t a i n
y2(z)=z+( 2−λ)z3
3!+( 2−λ)(6−λ)z5
5!+( 2−λ)(6−λ)(10−λ)z7
7!+···.
Now from the recurrence relation (16.33) (or in this case from the expressions for y1
andy2themselves) we see that for the ODE to possess a polynomial solution we require
λ=2 (n−2) for n≥2o rm o r es i m p l y λ=2nforn≥0, i.e. λmust be an even positive
integer. If λ= 4 then from (16.34) the ODE has the polynomial solution
y1(z)=1−4z2
2!=1−2z2.
J
A simpler method of obtaining finite polynomial solutions is to assume a
solution of the form (16.31), where aN/negationslash= 0. Instead of starting with the lowest
power of z, as we have done up to now, this time we start by considering the
coefficient of the highest power zN; such a power now exists because of our
assumed form of solution.IBy assuming a polynomial solution find the values of λin (16.32) for which such a solution
exists.
We assume a polynomial solution to (16.32) of the form y=
PN
n=0anzn. Substituting this
form into (16.32) we find
NX
n=0
/
n(n−1)anzn−2−2zna nzn−1+λanzn
/
=0.
Now, instead of starting with the lowest power of z, we start with the highest. Thus,
demanding that the coefficient of zNvanishes, we require −2N+λ=0 ,i . e . λ=2N,a sw e
found in the previous example. By demanding that the coefficient of a general power of z
is zero, the same recurrence relation as above may be derived and the solutions found.
J
16.6 Legendre’s equation
In previous sections we have discussed methods for obtaining series solutions of
second-order linear ODEs. In this section and the next we apply some of thesemethods to finding the series solutions of the two most important equations listedin table 16.1, namely Legendre’s equation and Bessel’s equation. As mentioned
earlier, the remaining equations in table 16.1 may also be solved by the methods
discussed in this chapter. These equations, and the properties of their solutions,are discussed briefly in the next chapter.
555
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
We now consider Legendre’s equation
(1−z2)y/prime/prime−2zy/prime+/lscript(/lscript+1 )y=0, (16.35)
which occurs in numerous physical applications and particularly in problems with
axial symmetry when they are expressed in spherical polar coordinates. In normalusage the variable zin Legendre’s equation is the cosine of the polar angle in
spherical polars, and thus −1≤z≤1. The parameter /lscriptis a given real number,
and any solution of (16.35) is called a Legendre function .
In subsection 16.1.1, we showed that z= 0 is an ordinary point of (16.35), and
so we expect to find two linearly independent solutions of the form y=summationtext
∞
n=0anzn.
Substituting, we find
∞summationdisplay
n=0bracketleftbig
n(n−1)anzn−2−n(n−1)anzn−2nanzn+/lscript(/lscript+1 )anznbracketrightbig
=0,
which on collecting terms gives
∞summationdisplay
n=0{(n+2 ) ( n+1 )an+2−[n(n+1 )−/lscript(/lscript+1 ) ] an}zn=0.
The recurrence relation is therefore
an+2=[n(n+1 )−/lscript(/lscript+1 ) ]
(n+1 ) ( n+2 )an, (16.36)
forn=0,1,2,.... If we choose a0=1a n d a1= 0 then we obtain the solution
y1(z)=1−/lscript(/lscript+1 )z2
2!+(/lscript−2)/lscript(/lscript+1 ) ( /lscript+3 )z4
4!−···, (16.37)
whereas choosing a0=0a n d a1= 1 we find a second solution
y2(z)=z−(/lscript−1)(/lscript+2 )z3
3!+(/lscript−3)(/lscript−1)(/lscript+2 ) ( /lscript+4 )z5
5!−···.(16.38)
By applying the ratio test to these series (see subsection 4.3.2), we find that both
series converge for |z|<1, and so their radius of convergence is unity, which
(as expected) is the distance to the nearest singular point of the equation. Since(16.37) contains only even powers of zand (16.38) contains only odd powers,
these two solutions cannot be proportional to one another, and are thereforelinearly independent. Hence y=c
1y1+c2y2is the general solution to (16.35) for
|z|<1.
16.6.1 General solution for integer /lscript
Now, if /lscriptis an integer in Legendre’s equation (16.35), i.e. /lscript=0,1,2,..., then the
recurrence relation (16.36) gives
a/lscript+2=[/lscript(/lscript+1 )−/lscript(/lscript+1 ) ]
(/lscript+1 ) ( /lscript+2 )a/lscript=0,
556
16.6 LEGENDRE’S EQUATION
P0
P1P2
P3−1
−1−0.5 0.5 11
z
−22
Figure 16.1 The first four Legendre polynomials.
i.e. the series terminates and we obtain a polynomial solution of order /lscript.T h e s e
solutions (suitably normalised) are called Legendre polynomials of order /lscript;t h e y
are written P/lscript(z) and are valid for all finite z. It is conventional to normalise
P/lscript(z)i ns u c haw a yt h a t P/lscript(1) = 1, and as a consequence P/lscript(−1) = (−1)/lscript.T h e
first few Legendre polynomials are easily constructed and are given by
P0(z)=1 P1(z)=z
P2(z)=1
2(3z2−1) P3(z)=1
2(5z3−3z)
P4(z)=1
8(35z4−30z2+3 ) P5(z)=1
18(63z5−70z3+1 5z).
The first four Legendre polynomials are plotted in figure 16.1.
According to whether /lscriptis an even or odd integer respectively, either y1(z)
in (16.37) or y2(z) in (16.38) terminates to give a multiple of the corresponding
Legendre polynomial P/lscript(z). In either case, however, the other series does not
terminate and therefore converges only for |z|<1. According to whether /lscriptis
even or odd we define Legendre functions of the second kind asQ/lscript(z)=α/lscripty2(z)
orQ/lscript(z)=β/lscripty1(z) respectively, where the constants α/lscriptandβ/lscriptare conventionally
taken to have the values
α/lscript=(−1)/lscript/22/lscript[(/lscript/2)!]2
/lscript!for/lscripteven, (16.39)
β/lscript=(−1)(/lscript+1)/22/lscript−1{[(/lscript−1)/2]!}2
/lscript!for/lscriptodd. (16.40)
557
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
These normalisation factors are chosen so that the Q/lscript(z) obey the same recurrence
r e l a t i o n sa st h e P/lscript(z) (see subsection 16.6.2).
The general solution of Legendre’s equation for integer /lscriptis therefore
y(z)=c1P/lscript(z)+c2Q/lscript(z), (16.41)
where P/lscript(z) is a polynomial of order /lscript, and so converges for all z,a n d Q/lscript(z)i s
an infinite series that converges only for |z|<1.†
By using the Wronskian method, section 16.4, one may obtain closed forms for
theQ/lscript(z).IUse the Wronskian method to find a closed-form expression for Q0(z).
From (16.25) a second solution to Legendre’s equation (16.35), with /lscript=0 ,i s
y2(z)=P0(z)
Zz1
[P0(u)]2exp
/Zu2v
1−v2dv
/
du
=
Zz
exp
/
−ln(1−u2)
/
du
=
Zzdu
(1−u2)=1
2ln
/1+z
1−z
/
, (16.42)
where in the second line we have used the fact that P0(z)=1 .
All that remains is to adjust the normalisation of this solution so that it agrees with
(16.39). Expanding the logarithm in (16.42) as a Maclaurin series we obtain
y2(z)=z+z3
3+z5
5+···.
Comparing this with the expression for Q0(z), using (16.38) with /lscript= 0 and the normali-
sation (16.39), we find that y2(z) is already correctly normalised, and so
Q0(z)=1
2ln
/1+z
1−z
/
.
Of course, we might have recognised the series (16.38) for /lscript= 0, but to do so for larger /lscript
would prove progressively more difficult.
J
Using the above method for /lscript= 1, we find
Q1(z)=1
2zlnparenleftbigg1+z
1−zparenrightbigg
−1.
Closed forms for higher-order Q/lscript(z) may now be found using the recurrence
relation (16.55) derived in the next subsection.
†It is possible, in fact, to find a second solution in terms of an infinite series of negative powers of
zthat is finite for |z|>1.
558
16.6 LEGENDRE’S EQUATION
16.6.2 Properties of Legendre polynomials
As stated earlier, when encountered in physical problems the variable zin Leg-
endre’s equation is usually the cosine of the polar angle θin spherical polar
coordinates, and we then require the solution y(z) to be regular at z=±1, which
corresponds to θ=0o r θ=π. For this to occur we require the equation to have
a polynomial solution, and so /lscriptmust be an integer. Furthermore, we also require
the coefficient c2of the function Q/lscript(z) in (16.41) to be zero, since Q/lscript(z) is singular
atz=±1, with the result that the general solution is simply some multiple of the
relevant Legendre polynomial P/lscript(z). In this section we will study the properties
of the Legendre polynomials P/lscript(z) in some detail.
Rodrigues’ formula
As an aid to establishing further properties of the Legendre polynomials we now
develop Rodrigues’ representation of these functions. Rodrigues’ formula for theP
/lscript(z)i s
P/lscript(z)=1
2/lscript/lscript!d/lscript
dz/lscript(z2−1)/lscript. (16.43)
To prove that this is a representation we let u=(z2−1)/lscript,s ot h a t u/prime=2/lscriptz(z2−1)/lscript−1
and
(z2−1)u/prime−2/lscriptzu=0.
If we differentiate this expression /lscript+ 1 times using Leibnitz’ theorem, we obtain
bracketleftbig
(z2−1)u(/lscript+2)+2z(/lscript+1 )u(/lscript+1)+/lscript(/lscript+1 )u(/lscript)bracketrightbig
−2/lscriptbracketleftbig
zu(/lscript+1)+(/lscript+1 )u(/lscript)bracketrightbig
=0,
which reduces to
(z2−1)u(/lscript+2)+2zu(/lscript+1)−/lscript(/lscript+1 )u(/lscript)=0.
Changing the sign all through and comparing the resulting expression with
Legendre’s equation (16.35), we see that u(/lscript)satisfies the same equation as P/lscript(z),
and so
u(/lscript)(z)=c/lscriptP/lscript(z), (16.44)
for some constant c/lscriptthat depends on /lscript. To establish the value of c/lscriptwe note
that the only term in the expression for the /lscriptth derivative of ( z2−1)/lscriptthat
does not contain a factor z2−1, and therefore does not vanish at z=1 ,i s
(2z)/lscript/lscript!(z2−1)0. Putting z= 1 in (16.44) and recalling that P/lscript(1) = 1, therefore
shows that c/lscript=2/lscript/lscript!, thus completing the proof of Rodrigues’ formula (16.43).
559
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONSIUse Rodrigues’ formula to show that
I/lscript=
Z1
−1P/lscript(z)P/lscript(z)dz=2
2/lscript+1. (16.45)
The result is trivially obvious for /lscript= 0 and so we assume /lscript≥1. Then, by Rodrigues’
formula,
I/lscript=1
22/lscript(/lscript!)2
Z1
−1
/d/lscript(z2−1)/lscript
dz/lscript
//d/lscript(z2−1)/lscript
dz/lscript
/
dz.
Repeated integration by parts, with all boundary terms vanishing, reduces this to
I/lscript=(−1)/lscript
22/lscript(/lscript!)2
Z1
−1(z2−1)/lscriptd2/lscript
dz2/lscript(z2−1)/lscriptdz
=(2/lscript)!
22/lscript(/lscript!)2
Z1
−1(1−z2)/lscriptdz.
If we write
K/lscript=
Z1
−1(1−z2)/lscriptdz,
then integration by parts (taking a factor 1 as the second part) gives
K/lscript=
Z1
−12/lscriptz2(1−z2)/lscript−1dz.
Writing 2 /lscriptz2as 2/lscript−2/lscript(1−z2)w eo b t a i n
K/lscript=2/lscript
Z1
−1(1−z2)/lscript−1dz−2/lscript
Z1
−1(1−z2)/lscriptdz
=2/lscriptK/lscript−1−2/lscriptK/lscript
and hence the recurrence relation (2 /lscript+1 )K/lscript=2/lscriptK/lscript−1. We therefore find
K/lscript=2/lscript
2/lscript+12/lscript−2
2/lscript−1···2
3K0=2/lscript/lscript!2/lscript/lscript!
(2/lscript+1 ) !2=22/lscript+1(/lscript!)2
(2/lscript+1 ) !,
which, when substituted into the expression for I/lscript, establishes the required result.
J
Mutual orthogonality of Legendre polynomials
Another useful property of the P/lscript(z) is their mutual orthogonality, i.e. that
integraldisplay1
−1P/lscript(z)Pk(z)dz=0 i f /lscript/negationslash=k. (16.46)
More general considerations concerning the mutual orthogonality of solutions to
various classes of second-order linear ODEs are discussed in the next chapter,but for the moment we concentrate on the specific proof of (16.46).
Since the P
/lscript(z) satisfy Legendre’s equation we may write
bracketleftbig
(1−z2)P/prime
/lscriptbracketrightbig/prime+/lscript(/lscript+1 )P/lscript=0,
560
16.6 LEGENDRE’S EQUATION
where P/prime
/lscript=dP/lscript/dz. Multiplying through by Pkand integrating from z=−1t o
z=1 ,w eo b t a i n
integraldisplay1
−1Pkbracketleftbig
(1−z2)P/prime
/lscriptbracketrightbig/primedz+integraldisplay1
−1Pk/lscript(/lscript+1 )P/lscriptdz=0.
Integrating the first term by parts and noting that the boundary contribution
vanishes at both limits because of the factor 1 −z2, we find
−integraldisplay1
−1P/prime
k(1−z2)P/prime
/lscriptdz+integraldisplay1
−1Pk/lscript(/lscript+1 )P/lscriptdz=0.
Now, if we reverse the roles of /lscriptandkand subtract one expression from the
other, we conclude that
[k(k+1 )−/lscript(/lscript+1 ) ]integraldisplay1
−1PkP/lscriptdz=0,
and therefore since k/negationslash=/lscriptwe must have the result (16.46). As a particular case
we note that if we put k=0w eo b t a i n
integraldisplay1
−1P/lscript(z)dz=0 f o r /lscript/negationslash=0.
As will be discussed more fully in the next chapter, the mutual orthogonality of
theP/lscript(z) means that any reasonable function f(z) (i.e. one obeying the Dirichlet
conditions discussed at the start of chapter 12) can be expressed in the interval|z|<1 as an infinite sum of Legendre polynomials,
f(z)=∞summationdisplay
/lscript=0a/lscriptP/lscript(z), (16.47)
where the coefficients a/lscriptare given by
a/lscript=2/lscript+1
2integraldisplay1
−1f(z)P/lscript(z)dz. (16.48)IProve the expression (16.48) for the coefficients in the Legendre polynomial expansion
of a function f(z).
If we multiply (16.47) by Pm(z) and integrate from z=−1t oz= 1 then we obtainZ1
−1Pm(z)f(z)dz=∞X
/lscript=0a/lscript
Z1
−1Pm(z)P/lscript(z)dz
=am
Z1
−1Pm(z)Pm(z)dz=2am
2m+1,
where we have used the orthogonality property (16.46) and the normalisation property
(16.45).
J
561
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
Generating function for Legendre polynomials
A useful device for manipulating and studying sequences of functions or quantities
labelled by an integer variable (here, the Legendre polynomials P/lscript(z) labelled by
/lscript)i sagenerating function . The generating function has perhaps its greatest utility
in the area of probability theory (see chapter 26). However, it is also a greatconvenience in our present study.
The generating function for, say, a series of functions f
n(z)f o r n=0,1,2,...is
a function G(z,h), containing as well as za dummy variable h, such that
G(z,h)=∞summationdisplay
n=0fn(z)hn,
i.e.fn(z) is the coefficient of hnin the expansion of Gin powers of h. The utility
of the device lies in the fact that sometimes it is possible to find a closed form
forG(z,h).
For our study of Legendre polynomials let us consider the functions Pn(z)
defined by the equation
G(z,h)=( 1−2zh+h2)−1/2=∞summationdisplay
n=0Pn(z)hn. (16.49)
As we show below, the functions so defined are identical to the Legendre poly-
nomials and the function (1 −2zh+h2)−1/2is in fact the generating function for
them. In the process we will also deduce several useful relationships between thevarious polynomials and their derivatives.
In the following dP
n(z)/dzwill be denoted by P/prime
n. Firstly, we differentiate the
defining equation (16.49) with respect to zto get
h(1−2zh+h2)−3/2=summationdisplay
P/prime
nhn. (16.50)
Also, we differentiate (16.49) with respect to hto yield
(z−h)(1−2zh+h2)−3/2=summationdisplay
nPnhn−1; (16.51)
equation (16.50) can then be written using (16.49) as
hsummationdisplay
Pnhn=( 1−2zh+h2)summationdisplay
P/prime
nhn,
and thus equating coefficients of hn+1we obtain the recurrence relation
Pn=P/prime
n+1−2zP/prime
n+P/prime
n−1. (16.52)
Equations (16.50) and (16.51) can be combined as
(z−h)summationdisplay
P/prime
nhn=hsummationdisplay
nPnhn−1,
from which the coefficent of hnyields a second recurrence relation
zP/prime
n−P/prime
n−1=nPn; (16.53)
562
16.6 LEGENDRE’S EQUATION
eliminating P/prime
n−1between (16.52) and (16.53) then gives the further result
(n+1 )Pn=P/prime
n+1−zP/prime
n. (16.54)
If we now take the result (16.54) with nreplaced by n−1a n da d d ztimes
(16.53) to it then we obtain
(1−z2)P/prime
n=n(Pn−1−zPn);
finally, differentiating both sides with respect to zand using (16.53) again, we
find
(1−z2)P/prime/prime
n−2zP/prime
n=n[(P/prime
n−1−zP/prime
n)−Pn]
=n(−nPn−Pn)=−n(n+1 )Pn,
a n ds ot h e Pndefined by (16.49) do indeed satisfy Legendre’s equation.
It remains only to verify the normalisation. This is easily done at z=1 ,w h e n
Gbecomes
G(1,h)=[ ( 1−h)2]−1/2=1+ h+h2+···,
and we can see that all the Pnso defined have Pn(1) = 1 as required. Many other
useful recurrence relations can be derived from those found above.IProve the recurrence relation
(n+1 )Pn+1−(2n+1 )zPn+nPn−1=0. (16.55)
Substituting from (16.49) into (16.51) we find
(z−h)
X
Pnhn=( 1−2zh+h2)
X
nPnhn−1.
Equating coefficients of hnwe obtain
zPn−Pn−1=(n+1 )Pn+1−2znP n+(n−1)Pn−1,
which on rearrangment gives the stated result.
J
Another use of the generating function (16.49) is in representing the inverse
distance between two points in three-dimensional space in terms of Legendrepolynomials. If two points randr
/primeare at distances randr/primerespectively from the
origin, with r/prime<r,t h e n
1
|r−r/prime|=1
(r2+r/prime2−2rr/primecosθ)1/2
=1
r[1−2(r/prime/r)cosθ+(r/prime/r)2]1/2
=1
r∞summationdisplay
/lscript=0parenleftbiggr/prime
rparenrightbigg/lscript
P/lscript(cosθ), (16.56)
where θis the angle between the two position vectors randr/prime.I fr/prime>r,h o w e v e r ,
then randr/primemust be exchanged in (16.56) or the series would not converge.
563
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
To summarise the situation concerning Legendre polynomials, we now have
three possible starting points, which have been shown to be equivalent: thedefining equation (16.35) together with the condition P
n(1) = 1; Rodrigues’
formula (16.43); and the generating function (16.49). In addition we have proved
a variety of relationships and recurrence relations (not particularly memorable,
but collectively useful) and, as will be apparent from the work of chapter 18,have developed a powerful tool for use in axially symmetric situations in whichthe∇
2operator is involved and spherical polar coordinates are employed.
16.7 Bessel’s equation
Bessel’s equation arises from physical situations similar to those involving Legen-
dre’s equation but when cylindrical, rather than spherical, polar coordinates areemployed. It has the form
z
2y/prime/prime+zy/prime+(z2−ν2)y=0, (16.57)
where the parameter νis a given number, which we may take as ≥0w i t hn ol o s s
of generality. In Bessel’s equation, zis usually a multiple of a radial distance and
therefore ranges from 0 to ∞.
Writing (16.57) in our standard form we have
y/prime/prime+1
zy/prime+parenleftbigg
1−ν2
z2parenrightbigg
y=0. (16.58)
By inspection z= 0 is a regular singular point; hence we try a solution of the
form y=zσsummationtext∞
n=0anzn. Substituting this into (16.58) and multiplying the resulting
equation by z2−σ,w eo b t a i n
∞summationdisplay
n=0bracketleftbig
(σ+n)(σ+n−1) + ( σ+n)−ν2bracketrightbig
anzn+∞summationdisplay
n=0anzn+2=0,
which simplifies to
∞summationdisplay
n=0bracketleftbig
(σ+n)2−ν2bracketrightbig
anzn+∞summationdisplay
n=0anzn+2=0.
Considering coefficients of z0we obtain the indicial equation
σ2−ν2=0,
and so σ=±ν. For coefficients of higher powers of zwe find
bracketleftbig
(σ+1 )2−ν2bracketrightbig
a1=0, (16.59)bracketleftbig
(σ+n)2−ν2bracketrightbig
an+an−2=0 f o r n≥2. (16.60)
564
16.7 BESSEL’S EQUATION
Substituting σ=±νinto (16.59) and (16.60) we obtain the recurrence relations
(1±2ν)a1=0, (16.61)
n(n±2ν)an+an−2=0 f o r n≥2. (16.62)
We consider now the form of the general solution to Bessel’s equation (16.57) for
two cases, the case for which νis not an integer and that for which it is (including
zero).
16.7.1 General solution for non-integer ν
Ifνis a non-integer then in general the two roots of the indicial equation, σ1=ν
and σ2=−ν, will not differ by an integer, and we may obtain two linearly
independent solutions in the form of Frobenius series. Special considerations do
arise, however, when ν=m/2f o r m=1,3,5,...,a n d σ1−σ2=2ν=mis an
(odd positive) integer. When this happens, we may always obtain a solution inthe form of a Frobenius series corresponding to the larger root σ
1=ν=m/2,
as described above. For the smaller root σ2=−ν=−m/2, however, we must
determine whether a second Frobenius series solution is possible by examiningthe recurrence relation (16.62), which reads
n(n−m)a
n+an−2=0 f o r n≥2.
Since mis anoddpositive integer in this case, we can use this recurrence relation
(starting with a0/negationslash=0 )t oc a l c u l a t e a2,a4,a6,...in the knowledge that all these
terms will remain finite. It is possible in this case, therefore, to find a secondsolution in the form of a Frobenius series corresponding to the smaller root σ
2.
Thus, in general, for non-integer νwe have from (16.61) and (16.62)
an=−1
n(n±2ν)an−2forn=2,4,6,...,
=0 f o r n=1,3,5,....
Setting a0= 1 in each case, we obtain the two solutions
y±ν(z)=z±νbracketleftbigg
1−z2
2(2±2ν)+z4
2×4(2±2ν)(4±2ν)−···bracketrightbigg
.
It is customary, however, to set
a0=1
2±νΓ(1±ν),
where Γ( x)i st h e gamma function , described in the appendix; it may be regarded
as the generalisation of the factorial function to non-integer and/or negativearguments.†The two solutions of (16.57) are then written as J
ν(z)a n d J−ν(z),
†In particular, Γ( n+1 )= n!f o r n=0,1,2,...,and Γ( n) is infinite if nis any integer ≤0.
565
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
where
Jν(z)=1
Γ(ν+1 )parenleftBigz
2parenrightBigνbracketleftbigg
1−1
ν+1parenleftBigz
2parenrightBig2
+1
(ν+1 ) ( ν+2 )1
2!parenleftBigz
2parenrightBig4
−···bracketrightbigg
=∞summationdisplay
n=0(−1)n
n!Γ(ν+n+1 )parenleftBigz
2parenrightBigν+2n
; (16.63)
replacing νby−νgives J−ν(z). The functions Jν(z)a n d J−ν(z) are called Bessel
functions of the first kind, of order ν. Since the first term of each series is a
finite non-zero multiple of zνandz−νrespectively, if νis not an integer then
Jν(z)a n d J−ν(z) are linearly independent. This may be confirmed by calculating
the Wronskian of these two functions. Therefore, for non-integer νthe general
solution of Bessel’s equation (16.57) is
y(z)=c1Jν(z)+c2J−ν(z). (16.64)IFind the general solution of
z2y/prime/prime+zy/prime+(z2−1
4)y=0.
This is Bessel’s equation with ν=1/2, so from (16.64) the general solution is simply
y(z)=c1J1/2(z)+c2J−1/2(z).
However, Bessel functions of half-integral order can be expressed in terms of trigonometric
functions. To show this, we note from (16.63) that
J±1/2(z)=z±1/2∞X
n=0(−1)nz2n
22n±1/2n!Γ(1 + n±1
2).
Using the fact that Γ( x+1 )= xΓ(x)a n dΓ (1
2)=√πwe find that, for ν=1/2,
J1/2(z)=(1
2z)1/2
Γ(3
2)−(1
2z)5/2
1!Γ(5
2)+(1
2z)9/2
2!Γ(7
2)−···
=(1
2z)1/2
(1
2)√π−(1
2z)5/2
1!(3
2)(1
2)√π+(1
2z)9/2
2!(5
2)(3
2)(1
2)√π−···
=(1
2z)1/2
(1
2)√π
/
1−z2
3!+z4
5!−···
/
=(1
2z)1/2
(1
2)√πsinz
z=
r
2
πzsinz,
whereas for ν=−1/2w eo b t a i n
J−1/2(z)=(1
2z)−1/2
Γ(1
2)−(1
2z)3/2
1!Γ(3
2)+(1
2z)7/2
2!Γ(5
2)−···
=(1
2z)−1/2
√π
/
1−z2
2!+z4
4!−···
/
=
r
2
πzcosz.
Therefore the general solution we require is
y(z)=c1J1/2(z)+c2J−1/2(z)=c1
r
2
πzsinz+c2
r
2
πzcosz.
J
566
16.7 BESSEL’S EQUATION
Corresponding to the discussion in subsection 16.6.2 of the general solution
of Legendre’s equation, we note that when Bessel’s equation is encountered inphysical situations the argument zis usually some multiple of a radial distance
and so takes values in the range 0 ≤z≤∞. We often require that the solution
is regular at z= 0 but, from (16.63), we see immediately that J
−ν(z) is singular
at the origin (remember that we restricted νto be non-negative). In such cases,
the coefficient c2in (16.64) must be set to zero, and the solution is simply some
multiple of Jν(z).
16.7.2 General solution for integer ν
The definition of the Bessel function Jν(z) given in (16.63) is, of course, valid for
all values of νbut, as we shall see, in the case of integer νthe general solution of
Bessel’s equation cannot be written in the form (16.64). Firstly let us consider thecase ν= 0, so that the two solutions to the indicial equation are equal, and we
clearly obtain only one solution in the form of a Frobenius series. From (16.63),this is given by
J
0(z)=∞summationdisplay
n=0(−1)nz2n
22nn!Γ(1 + n)
=1−z2
22+z4
2242−z6
224262+···.
In general, however, if νis a positive integer then the solutions of the indicial
equation differ by an integer. For the larger root, σ1=ν, we may find a solution
Jν(z)f o r ν=1,2,3,..., in the form of a Frobenius series given by (16.63). Graphs
ofJ0(z),J1(z)a n d J2(z) are plotted in figure 16.2 for real z. For the smaller root
σ2=−ν, however, the recurrence relation (16.62) becomes
n(n−m)an+an−2=0 f o r n≥2,
where m=2νis now an evenpositive integer, i.e. m=2,4,6,.... Starting with
a0/negationslash= 0 we may then calculate a2,a4,a6,..., but we see that when n=mthe
coefficient anis formally infinite, and the method fails to produce a second
solution in the form of a Frobenius series.
In fact, by replacing νby−νin the definition of Jν(z) given in (16.63), it can
be shown that, for integer ν,
J−ν(z)=(−1)νJν(z)
and hence that Jν(z)a n d J−ν(z) are linearly dependent. So, in this case, we cannot
write the general solution to Bessel’s equation in the form (16.64). One therefore
defines the function
Yν(z)=Jν(z)cosνπ−J−ν(z)
sinνπ, (16.65)
567
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
J0
J1
J2
24 6 81 0z 0
−0.4−0.20.20.40.60.81
Figure 16.2 The first three integer-order Bessel functions.
which is called a Bessel’s function of the second kind of order ν.A sB e s s e l ’ se q u a -
tion is linear, Yν(z) is clearly a solution, since it is just the weighted sum of Bessel
functions of the first kind. Furthermore, for non-integer νit is clear that Yν(z)i s
linearly independent of Jν(z). It may also be shown that the Wronskian of Jν(z)
andYν(z) is non-zero for allvalues of ν. Hence Jν(z)a n d Yν(z) always constitute
a pair of independent solutions. The expression (16.65) becomes an indeterminate
form 0 /0w h e n νis an integer, however. This is so because for integer νwe have
cosνπ=(−1)νandJ−ν(z)=(−1)νJν(z). Nevertheless, this indeterminate form
can be evaluated using l’H ˆopital’s rule (see chapter 4). Thus for integer νwe set
Yν(z) = lim
µ→νbracketleftbiggJµ(z)cosµπ−J−µ(z)
sinµπbracketrightbigg
, (16.66)
which gives a linearly independent second solution for integer ν. Therefore, we
may write the general solution of Bessel’s equation, valid for allν,a s
y(z)=c1Jν(z)+c2Yν(z). (16.67)
As mentioned above for the case when νis not an integer, in physical situations
we often require the solution of Bessel’s equation to be regular at z=0 .B u t ,
from its definition (16.65) or (16.66), it is clear that Yν(z) is singular at the origin,
and so in such physical situations the coefficient c2in (16.67) must be set to zero;
the solution is then simply some multiple of Jν(z).
16.7.3 Properties of Bessel functions
Bessel functions of the first and second kind, Jν(z)a n d Yν(z), have various useful
properties that are worthy of further discussion.
568
16.7 BESSEL’S EQUATION
Recurrence relations
The recurrence relations enjoyed by Bessel functions of the first kind, Jν(z), can
be derived directly from the power series definition (16.63).IProve the recurrence relation
d
dz[zνJν(z)] =zνJν−1(z). (16.68)
From the power series definition (16.63) of Jν(z)w eo b t a i n
d
dz[zνJν(z)] =d
dz∞X
n=0(−1)nz2ν+2n
2ν+2nn!Γ(ν+n+1 )
=∞X
n=0(−1)nz2ν+2n−1
2ν+2n−1n!Γ(ν+n)
=zν∞X
n=0(−1)nz(ν−1)+2 n
2(ν−1)+2 nn!Γ((ν−1) +n+1 )=zνJν−1(z).
J
It may similarly be shown that
d
dz[z−νJν(z)] =−z−νJν+1(z). (16.69)
From (16.68) and (16.69) the remaining recurrence relations may be easily derived.
Expanding out the derivative on the LHS of (16.68) and dividing through by zν−1
we obtain the relation
zJ/prime
ν(z)+νJν(z)=zJν−1(z). (16.70)
Similarly, by expanding out the derivative on the LHS of (16.69), and multiplying
through by zν+1, we find
zJ/prime
ν(z)−νJν(z)=−zJν+1(z). (16.71)
Adding (16.70) and (16.71) and dividing through by zgives
Jν−1(z)−Jν+1(z)=2 J/prime
ν(z). (16.72)
Finally, subtracting (16.71) from (16.70) and dividing by zgives
Jν−1(z)+Jν+1(z)=2ν
zJν(z). (16.73)
569
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONSIGiven that J1/2(z)=( 2 /πz)1/2sinzand that J−1/2(z)=( 2 /πz)1/2cosz,e x p r e s s J3/2(z)
andJ−3/2(z)in terms of trigonometric functions.
From (16.71) we have
J3/2(z)=1
2zJ1/2(z)−J/prime
1/2(z)
=1
2z
/2
πz
/1/2
sinz−
/2
πz
/1/2
cosz+1
2z
/2
πz
/1/2
sinz
=
/2
πz
/1/2
/1
zsinz−cosz
/
.
Similarly, from (16.70), we have
J−3/2(z)=−1
2zJ−1/2(z)+J/prime
−1/2(z)
=−1
2z
/2
πz
/1/2
cosz−
/2
πz
/1/2
sinz−1
2z
/2
πz
/1/2
cosz
=
/2
πz
/1/2
/
−1
zcosz−sinz
/
.
We shall see that, by repeated use of these recurrence relations, all Bessel functions Jν(z)
of half-integer order may be expressed in terms of trigonometric functions. From their
definition (16.65), Bessel functions of the second kind, Yν(z), of half-integer order can be
similarly expressed.
J
Finally, we note that the relations (16.68) and (16.69) may be rewritten in
integral form as
integraldisplay
zνJν−1(z)dz=zνJν(z)
integraldisplay
z−νJν+1(z)dz=−z−νJν(z).
Ifνis an integer, the recurrence relations of this section may be proved using
the generating function for Bessel functions discussed below. It may be shownthat Bessel functions of the second kind, Y
ν(z), also satisfy the recurrence relations
derived above.
Mutual orthogonality of Bessel functions
Bessel functions of the first kind, Jν(z), possess an orthogonality relation analo-
gous to that of the Legendre polynomials discussed in subsection 16.6.2. A moregeneral discussion of the mutual orthogonality of solutions to second-order linearODEs (such as Bessel’s equation) is given in chapter 17.
By definition, the function J
ν(z) satisfies Bessel’s equation (16.57),
z2y/prime/prime+zy/prime+(z2−ν2)y=0.
570
16.7 BESSEL’S EQUATION
Let us instead consider the functions f(z)=Jν(λz)a n d g(z)=Jν(µz), which, as
will be proved below, respectively satisfy the equations
z2f/prime/prime+zf/prime+(λ2z2−ν2)f=0, (16.74)
z2g/prime/prime+zg/prime+(µ2z2−ν2)g=0. (16.75)IShow that f(z)=Jν(λz)satisfies (16.74).
Iff(z)=Jν(λz) and we write w=λz,t h e n
df
dz=λdJν(w)
dwandd2f
dz2=λ2d2Jν(w)
dw2.
When these expressions are substi tuted, the LHS of (16.74) becomes
z2λ2d2Jν(w)
dw2+zλdJν(w)
dw+(λ2z2−ν2)Jν(w)
=w2d2Jν(w)
dw2+wdJν(w)
dw+(w2−ν2)Jν(w).
But, from Bessel’s equation itself, this final expression is equal to zero, thus verifying that
f(z) does satisfy (16.74).
J
Now multiplying (16.75) by f(z) and (16.74) by g(z) and subtracting them gives
d
dz[z(fg/prime−gf/prime)] = ( λ2−µ2)zfg, (16.76)
w h e r ew eh a v eu s e dt h ef a c tt h a t
d
dz[z(fg/prime−gf/prime)] =z(fg/prime/prime−gf/prime/prime)+(fg/prime−gf/prime).
By integrating (16.76) over any given range z=atoz=bwe obtain
integraldisplayb
azf(z)g(z)dz=1
λ2−µ2bracketleftBig
zf(z)g/prime(z)−zg(z)f/prime(z)bracketrightBigb
a,
which, on setting f(z)=Jν(λz)a n d g(z)=Jν(µz), becomes
integraldisplayb
azJν(λz)Jν(µz)dz=1
λ2−µ2bracketleftBig
µzJ ν(λz)J/prime
ν(µz)−λzJ ν(µz)J/prime
ν(λz)bracketrightBigb
a.
(16.77)
Ifλ/negationslash=µ, and the interval [ a, b] is such that the expression on the RHS of (16.77)
equals zero then we obtain the orthogonality condition
integraldisplayb
azJν(λz)Jν(µz)dz=0. (16.78)
This happens, for example, if Jν(λz)a n d Jν(µz) vanish at z=aandz=b,o ri f
J/prime
ν(λz)a n d J/prime
ν(µz) vanish at z=aandz=b, or for many more general conditions.
Ifλ=µ, however, then the RHS of (16.77) takes the indeterminant form 0 /0.
This may be evaluated using l’H ˆopital’s rule, or alternatively we may calculate
the relevant integral directly.
571
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONSIEvaluate the integralZb
aJ2
ν(λz)zd z .
Ignoring the integration limits for the moment,Z
J2
ν(λz)zd z=1
λ2
Z
J2
ν(u)u du,
where u=λz. Integrating by parts yields
I=
Z
J2
ν(u)ud u=1
2u2J2
ν(u)−
Z
Jν(u)J/prime
ν(u)u2du.
Now Bessel’s equation (16.57) can be rearranged as
u2Jν(u)=ν2Jν(u)−uJ/prime
ν(u)−u2J/prime/prime
ν(u),
which, on substitution into the expression for I,g i v e s
I=1
2u2J2
ν(u)−
Z
J/prime
ν(u)[ν2Jν(u)−uJ/prime
ν(u)−u2J/prime/prime
ν(u)]du
=1
2u2J2
ν(u)−1
2ν2J2
ν(u)+1
2u2[J/prime
ν(u)]2+c.
Since u=λzthe required integral is given byZb
aJ2
ν(λz)zd z=1
2
//
z2−ν2
λ2
/
J2
ν(λz)+z2[J/prime
ν(λz)]2
/b
a, (16.79)
which gives the normalisation condition for Bessel functions of the first kind.
J
Since the Bessel functions Jν(z) possess the orthogonality property (16.78) we
may expand any reasonable function f(z) (i.e. one obeying the Dirichlet conditions
discussed in chapter 12) in the interval 0 ≤z≤aas a sum of Bessel functions of
a given order ν,
f(z)=∞summationdisplay
n=0cnJν(λnz), (16.80)
where the λnare chosen such that Jν(λna) = 0. The coefficients cnare then given
by
cn=2
a2J2
ν+1(λna)integraldisplaya
0f(z)Jν(λnz)zd z . (16.81)
572
16.7 BESSEL’S EQUATIONIProve the expression (16.81) for the coeffi cients in a Bessel function expansion of a
function f(z).
If we multiply (16.80) by zJν(λmz) and integrate from z=0t o z=athen we obtainZa
0zJν(λmz)f(z)dz=∞X
n=0cn
Za
0zJν(λmz)Jν(λnz)dz
=cm
Za
0J2
ν(λmz)zd z
=1
2cma2J/prime2
ν(λma)=1
2cma2J2
ν+1(λma),
where in the last two lines we have used (16.77), (16.79), the fact that Jν(λma)=0a n d
(16.71).
J
Generating function for Bessel functions
The Bessel functions Jν(z), where νis an integer, can be described by a gener-
ating function in a similar way to that discussed for Legendre polynomials insubsection 16.6.2. The generating function for Bessel functions of integer order isgiven by
G(z,h)=e x pbracketleftbiggz
2parenleftbigg
h−1
hparenrightbiggbracketrightbigg
=∞summationdisplay
n=−∞Jn(z)hn. (16.82)
By expanding the exponential as a power series, it is straightfoward to verify that
the functions Jn(z) defined by (16.82) are indeed Bessel functions of the first kind.
The generating function (16.82) is useful for finding, for Bessel functions of
integer order, properties which can often be extended to the non-integer case. Inparticular, the Bessel function recurrence relations may be derived.IUse the generating function (16.82) to prove, for integer ν, the recurrence relation (16.73),
i.e.
Jν−1(z)+Jν+1(z)=2ν
zJν(z).
Differentiating G(z,h) with respect to hwe obtain
∂G(z,h)
∂h=z
2
/
1+1
h2
/
G(z,h)=∞X
n=−∞nJn(z)hn−1,
which can be written using (16.82) again as
z
2
/
1+1
h2
/∞X
n=−∞Jn(z)hn=∞X
n=−∞nJn(z)hn−1.
Equating coefficients of hnwe obtain
z
2[Jn(z)+Jn+2(z)] = ( n+1 )Jn+1(z),
which on replacing nbyν−1 gives the required recurrence relation.
J
573
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
The generating function (16.82) is also useful in deriving the integral represen-
tation of Bessel functions of integer order.IShow that for integer nthe Bessel function Jn(z)is given by
Jn(z)=1
π
Zπ
0cos(nθ−zsinθ)dθ. (16.83)
By expanding out the cosine term in the integrand in (16.83) we obtain the integral
I=1
π
Zπ
0[cos(zsinθ)c osnθ+s i n ( zsinθ)sinnθ]dθ. (16.84)
Now, we may express cos( zsinθ) and sin( zsinθ) in terms of Bessel functions by setting
h=e x p iθin (16.82) to give
exp
hz
2(expiθ−exp(−iθ))
i
=e x p (izsinθ)=∞X
m=−∞Jm(z)exp imθ.
Using de Moivre’s theorem exp iθ=c o s θ+isinθwe then obtain
exp(izsinθ)=c o s ( zsinθ)+isin(zsinθ)=∞X
m=−∞Jm(z)(cos mθ+isinmθ).
Equating the real and imaginary parts of this expression we find
cos(zsinθ)=∞X
m=−∞Jm(z)cosmθ,
sin(zsinθ)=∞X
m=−∞Jm(z)sinmθ.
Substituting these expressions into (16.84) we find
I=1
π∞X
m=−∞
Zπ
0[Jm(z)c osmθcosnθ+Jm(z)si nmθsinnθ]dθ.
However, using the orthogonality of the trigonometric functions, see equations (12.1)–
(12.3), we obtain
I=1
ππ
2[Jn(z)+Jn(z)] =Jn(z),
which proves the integral representation (16.83).
J
Finally, we mention the special case of the integral representation (16.83) for
n=0 ,
J0(z)=1
πintegraldisplayπ
0cos(zsinθ)dθ=1
2πintegraldisplay2π
0cos(zsinθ)dθ,
since cos( zsinθ) repeats itself in the range θ=πtoθ=2π. However, sin( zsinθ)
changes sign in this range and so
1
2πintegraldisplay2π
0sin(zsinθ)dθ=0.
574
16.8 GENERAL REMARKS
Using de Moivre’s theorem, we can therefore write
J0(z)=1
2πintegraldisplay2π
0exp(izsinθ)dθ=1
2πintegraldisplay2π
0exp(izcosθ)dθ.
There are in fact many other integral representations of Bessel functions, which
can be derived from those given.
16.8 General remarks
As was our intention, in respect of infinite series solutions we have concentrated
to a very marked degree on Bessel’s equation and, in respect of finite polynomialsolutions, on Legendre’s equation. The techniques used are, however, applicable
to many equations other than these, but since the procedures are in all essentials
the same, we do not need to treat them explicitly. The solutions of the remainingequations in table 16.1 are discussed briefly in the next chapter in connectionwith Sturm–Liouville systems.
16.9 Exercises
16.1 Find two power series solutions about z= 0 of the differential equation
(1−z2)y/prime/prime−3zy/prime+λy=0.
Deduce that the value of λfor which the corresponding power series becomes an
Nth-degree polynomial UN(z)i sN(N+ 2). Construct U2(z)a n d U3(z).
16.2 Find solutions, as power series in z, of the equation
4zy/prime/prime+2 ( 1−z)y/prime−y=0.
Identify one of the solutions and verify it by direct substitution.
16.3 Find power series solutions in zof the differential equation
zy/prime/prime−2y/prime+9z5y=0.
Identify closed forms for the two series, calculate their Wronskian, and verify
that they are linearly independent. Compare the Wronskian with that calculatedfrom the differential equation.
16.4 Change the independent variable in the equation
d
2f
dz2+2 (z−a)df
dz+4f=0 ( * )
from ztox=z−α, and find two independent series solutions, expanded about
x= 0, of the resulting equation. Deduce that the general solution of (*) is
f(z,α)=A(z−α)e−(z−α)2+B∞X
m=0(−4)mm!
(2m)!(z−α)2m,
with AandBarbitrary constants.
16.5 (a) Verify that z= 1 is a regular singular point of Legendre’s equation and that
the indicial equation for a series solution in powers of ( z−1) has roots 0
and 3.
(b) Obtain the corresponding recurrence relation and show that σ= 0 does not
give a valid series solution.
575
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
(c) Determine the radius of convergence Rof the σ= 3 series and relate it to
the positions of the singularities of Legendre’s equation.
16.6 Verify that z= 0 is a regular singular point of the equation
z2y/prime/prime−3
2zy/prime+( 1+ z)y=0,
and that the indicial equation has roots 2 and 1 /2. Show that the general solution
is
y(z)=6 a0z2∞X
n=0(−1)n(n+1 ) 22nzn
(2n+3 ) !
+b0
/
z1/2+2z3/2−z1/2
4∞X
n=2(−1)n22nzn
n(n−1)(2n−3)!
/!
.
16.7 Use the derivative method to obtain as a second solution of Bessel’s equation for
the case when ν= 0 the following expression:
J0(z)lnz−∞X
n=1(−1)n
(n!)2
/ nX
r=11
r
/!/z
2
/2n
,
given that the first solution is J0(z) as specified by (16.63).
16.8 By initially writing y(x)a s x1/2f(x) and then making subsequent changes of
variable, reduce
d2y
dx2+λxy=0
to Bessel’s equation. Hence show that a solution that is finite at x=0i sa
multiple of x1/2J1/3(2
3√
λx3).
16.9 (a) Show that the indicial equation for
zy/prime/prime−2y/prime+yz=0
has roots that differ by an integer but that the two roots nevertheless generate
linearly independent solutions
y1(z)=3 a0∞X
n=1(−1)n+12nz2n+1
(2n+1 ) !,
y2(z)=a0∞X
n=0(−1)n+1(2n−1)z2n
(2n)!.
(b) Show that y1(z)i se q u a lt o3 a0(sinz−zcosz) by expanding the sinusoidal
functions. Then, using the Wronskian method, find an expression for y2(z)
in terms of sinusoids. (You will need to write z2as (z/sinz)(zsinz)a n d
integrate by parts to evaluate the integral involved.)
(c) Confirm that the two solutions are linearly independent by showing that
their Wronskian is equal to −z2, in accordance with (16.4).
16.10 Find series solutions of the equation y/prime/prime−2zy/prime−2y= 0. Identify one of the series
asy1(z)=e x p z2and verify this by direct substitution. By setting y2(z)=u(z)y1(z)
and solving the resulting equation for u(z), find an explicit form for y2(z)a n d
deduce thatZx
0e−v2dv=e−x2∞X
n=0n!
2(2n+1 ) !(2x)2n+1.
576
16.9 EXERCISES
16.11 (a) Identify and classify the singular points of the equation
z(1−z)d2y
dz2+( 1−z)dy
dz+λy=0,
and determine their indices.
(b) Find one series solution in powers of z. Give a formal expression for a
second linearly independent solution.
(c) Deduce the values of λfor which there is a polynomial solution PN(z)o f
degree N. Evaluate the first four polynomials, normalised in such a way that
PN(0) = 1 .
16.12 Find the general power series solution about z=0o ft h ee q u a t i o n
zd2y
dz2+( 2z−3)dy
dz+4
zy=0.
16.13 Find the radius of convergence of a series solution about the origin for the
equation ( z2+az+b)y/prime/prime+2y= 0 in the following cases:
(a)a=5 , b=6 ;( b ) a=5 , b=7 .
Show that if aandbare real and 4 b>a2then the radius of convergence is
always given by b1/2.
16.14 For the equation y/prime/prime+z−3y= 0, show that the origin becomes a regular singular
point if the independent variable is changed from ztox=1/z. Hence find a
series solution of the form y1(z)=
P∞
0anz−n. By setting y2(z)=u(z)y1(z)a n d
expanding the resulting expression for du/dz in powers of z−1, show that y2(z)
has the asymptotic form
y2(z)=c
/
z+l nz−1
2+O
/lnz
z
//
,
where cis an arbitrary constant.
16.15 Prove that the Laguerre equation
zd2y
dz2+( 1−z)dy
dz+λy=0
has polynomial solutions LN(z)i fλis a non-negative integer N, and determine
the recurrence relationship for the polynomial coefficients. Hence show that anexpression for L
N(z), normalised in such a way that LN(0) = N!, is
LN(z)=NX
n=0(−1)n(N!)2
(N−n)!(n!)2zn.
Evaluate L3(z) explicitly. [The Laguerre generating function is discussed in exer-
cise 17.9.]
16.16 (a) Use Leibniz’ theorem to show that the Rodrigues’ formula for the Laguerre
polynomials LN(z) of the previous question is
LN(z)=ezdN
dzN(zNe−z).
(b) Use the Rodrigue formulation to prove that
zL/prime
N(z)=LN+1(z)−(N+1−z)LN(z).
(c) Deduce the recurrence relation for the Laguerre polynomials, namely
LN+1(z)+(z−2N−1)LN(z)+N2LN−1(z)=0 .
577
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
16.17 Equation (16.32) was shown to have a polynomial solution provided that λ=2n
with nan integer≥0. The polynomials are known as Hermite polynomials Hn(x)
and are of importance in the quantum mechanical treatment of the harmonicoscillator problem. They may also be defined by
Φ(x, h)=e x p ( 2 xh−h
2)=∞X
n=01
n!Hn(x)hn.
Show that
∂2Φ
∂x2−2x∂Φ
∂x+2h∂Φ
∂h=0,
and hence that the Hn(x) satisfy (16.32). Use Φ to prove that
(a)H/prime
n(x)=2 nHn−1(x),
(b)Hn+1(x)−2xHn(x)+2nHn−1(x)=0 .
16.18 By writing Φ( x, h) of the previous exercise as a function of h−xrather than of
h, show that an alternative representation of the nth Hermite polynomial is
Hn(x)=(−1)n
/;
expx2
/dn
dxn[exp(−x2)].
(Note that Hn(x)=∂nΦ/∂hnath=0 . )
16.19 Obtain the recurrence relations for the solution of Legendre’s equation (16.35)
ininverse powers of z,i . e .s e t y(z)=
Panzσ−n,w i t h a0/negationslash= 0. Deduce that if /lscriptis
an integer then the series with σ=/lscriptwill terminate and hence converge for all z
whilst that with σ=−(/lscript+ 1) does not terminate and hence converges only for
|z|>1.
16.20 Carry through the following procedure as an alternative proof of result (16.45).
(a) Square both sides of (16.49), giving the generating-function definition of the
Legendre polynomials.
(b) Express the RHS as a sum of powers of h, obtaining expressions for the
coefficients.
(c) Integrate the RHS from −1 to 1 and use the orthogonality results (16.46).
(d) Similarly integrate the LHS and expand the result in powers of h.
(e) Compare coefficients.
16.21 A charge +2 qis situated at the origin and charges of −qare situated at distances
±afrom it along the polar axis. By relating it to the generating function for the
Legendre polynomials, show that the electrostatic potential Φ at a point ( r,θ,φ)
with r>a is given by
Φ(r,θ,φ)=2q
4π/epsilon10r∞X
s=1
/a
r
/2s
P2s(cosθ).
16.22 The origin is an ordinary point of the Chebyshev equation,
(1−z2)y/prime/prime−zy/prime+m2y=0,
which therefore has series solutions of the form zσ
P∞
0anznforσ=0a n d σ=1 .
(a) Find the recurrence relationships for the anin the two cases and show that
there exist polynomial solutions Tm(z):
(i) for σ=0 ,w h e n mis an even integer, the polynomial having1
2(m+2 )
terms;
(ii) for σ=1 ,w h e n mis an odd integer, the polynomial having1
2(m+1 )
terms.
578
16.10 HINTS AND ANSWERS
(b)Tm(z) is normalised so as to have Tm(1) = 1. Find explicit forms for Tm(z)
form=0,1,2,3.
(c) Show that the corresponding non-terminating series solutions Sm(z) have as
their first few terms
S0(z)=a0
/
z+1
3!z3+9
5!z5+···
/
,
S1(z)=a0
/
1−1
2!z2−3
4!z4−···
/
,
S2(z)=a0
/
z−3
3!z3−15
5!z5−···
/
,
S3(z)=a0
/
1−9
2!z2+45
4!z4+···
/
.
16.23 By choosing a suitable form for hin (16.82), show that further integral repe-
sentations of the Bessel functions of the first kind are given, for integral m,
by
J2m(z)=(−1)m
π
Z2π
0cos(zcosθ)c o s 2 mθ dθ m ≥1,
J2m+1(z)=(−1)m+1
π
Z2π
0cos(zcosθ)s i n ( 2 m+1 )θd θ m≥0.
16.24 Show from the definition given in (16. 66) that the Bessel function of the second
kind of order νcan be written as
Yν(z)=1
π
/∂Jµ(z)
∂µ−(−1)ν∂J−µ(z)
∂µ
/
µ=ν.
Using the explicit expression (16.63) for Jµ(z), show that ∂Jµ(z)/∂µcan be written
as
Jν(z)ln
/z
2
/
+g(ν,z),
and deduce that Yν(z) can be expressed as
Yν(z)=2
πJν(z)ln
/z
2
/
+h(ν,z),
h(ν,z), like g(ν,z), being a power series in z.
16.10 Hints and answers
16.1 Note that z= 0 is an ordinary point of the equation.
Forσ=0,an+2/an=[n(n+2)−λ]/[(n+1)(n+2)] and correspondingly for σ=1 ;
U2(z)=a0(1−4z2)a n d U3(z)=a0(z−2z3).
16.2 a0exp(z/2);b0z1/2
P∞
n=0(2z)nn!/(2n+1 ) ! .
16.3 σ=0a n d3 ; a6m/a0=(−1)m/(2m)! and a6m/a0=(−1)m/(2m+ 1)! respectively.
y1(z)=a0cosz3andy2(z)=a0sinz3.T h eW r o n s k i a ni s ±3a2
0z2/negationslash=0.
16.4 x= 0 is an ordinary point of the transformed equation and so σ=0a n d1 .
Forσ=1,an+2=−2an/(n+2 )a n ds o a2m/a0=(−1)m/m!.Forσ=0,an+2=
−2an/(n+1 )a n ds o a2m/a0=(−2)m/
Qm
r=1(2r−1).
16.5 (b) an+1/an=−[(σ+n)(σ+n−3) +/lscript(/lscript+1 ) ] /[2(σ+n)2−2]. For σ=0,a2=∞.
(c)R= 2, equal to the distance between z= 1 and the closest singularity at
z=−1.
579
SERIES SOLUTIONS OF ORDINARY DIFFERENTIAL EQUATIONS
16.8 x2f/prime/prime+xf/prime+(λx3−1
4)f= 0. Then, in turn, set x3/2=u,a n d2 λ1/2u/3=v;t h e n v
satisfies Bessel’s equation with ν=1/3.
16.9 (b) cos z+zsinz.
16.10 y2(z)=( e x p z2)
Rz
0exp(−x2)dx.
16.11 (a) Regular singular points at z= 0 (indices 0, 0) and at z= 1 (indices 0, 1).
(b)y1(z)=a0+a0
P∞
n=1(n!)−2zn
Qn−1
r=0(r2−λ).
y2(z)=y1(z)lnz+
P∞
n=1zn
/
(∂/∂σ)
nQn−1
r=0[(r+σ)2−λ]/(n+σ+1 )2
o/
σ=0.
(c)λ=N2; polynomials are 1, 1 −z,( 1−z)(1−3z), (1−z)(1−8z+1 0z2).
16.12 Repeated roots σ=2 .
y(z)=az2+∞X
n=1(n+1 ) (−2z)n+2
n!
na
4+b[lnz+g(n)]
o
,
where
g(n)=1
n+1−1
n−1
n−1−···−1
2−2.
16.13 (a) 2; (b)√
7.
16.14 Transformed equation is xy/prime/prime+2y/prime+y=0 ; an=(−1)n(n+1 )−1(n!)−2a0;du/dz =
A[y1(z)]−2.
16.15 an+1=−(N−n)an/(n+1 )2;L3(z)=6−18z+9z2−z3.
16.16 (b) Calculate LN+1(z), considering zN+1e−zaszzNe−z. Later write dN/dzN(zNe−z)
ase−zLN(z).(c) Use (b) to calculate L/prime
N+1(z), substituting for L/prime/prime
N(z)f r o mt h e
Laguerre equation. Substitute from (b) for the first derivatives, and finally change
n+1t o n.
16.17 Consider ∂Φ/∂x; (b) differentiate result (a) and then use (a) again to replace the
derivatives.
16.19 σ=/lscript;an+2=[ (/lscript−n)(/lscript−n−1)an]/[(n+2)(n−2/lscript+1)]. Note that ( n−2/lscript+1)/negationslash=0
forn≤/lscript+1a n d neven.
σ=−(/lscript+1 ) ; an+2=[ (/lscript+n+1 ) ( /lscript+n+2 )an]/[(n+2 ) ( n+2/lscript+3 ) ] .
16.20 At step (d)
1
hln1+h
1−h=∞X
h=0h2n
Z1
−1P2
n(x)dx.
16.21 Using the cosine law, the distances from the charges −qare of the form
r
/
1±2(a/r)cosθ+(a/r)2
/1/2.
16.22 (a) (i) an+2=[an(n2−m2)]/[(n+2 ) ( n+ 1)],
(ii)an+2={an[(n+1 )2−m2]}/[(n+3 ) ( n+ 2)]; (b) 1, z,2z2−1, 4z3−3z.
16.23 Set h=iexpiθand obtain an expression for cos( zcosθ).
16.24 Recall that J−ν(z)=(−1)νJν(z)f o ri n t e g e r ν.
580
17
Eigenfunction methods for
differential equations
In the previous three chapters we dealt with the solution of differential equations
of order nby two methods. In one method, we found nindependent solutions
of the equation and then combined them, weighted with coefficients determinedby the boundary conditions; in the other we found solutions in terms of serieswhose coefficients were related by (in general) an n-term recurrence relation and
thence fixed by the boundary conditions. For both approaches the linearity of theequation was an important or essential factor in the utility of the method, and
in this chapter our aim will be to exploit the superposition properties of linear
differential equations even further.
We will be concerned with the solution of equations of the inhomogeneous
form
Ly(x)=f(x), (17.1)
where f(x) is a prescribed or general function and the boundary conditions to
be satisfied by the solution y=y(x), for example at the limits x=aandx=b,
are given. The expression Ly(x) stands for a linear differential operator Lacting
upon the function y(x).
In general, unless f(x) is both known and simple, it will not be possible to find
particular integrals of (17.1), even if complementary functions can be found thatsatisfy Ly= 0. The idea is therefore to exploit the linearity of Lby building up
the required solution as a superposition , generally containing an infinite number
of terms, of some set of functions that each individually satisfy the boundaryconditions. Clearly this brings in a quite considerable complication but since,within reason, we may select the set of functions to suit ourselves, we can obtain
sizeable compensation for this complication. Indeed, if the set chosen is one
containing functions that, when acted upon by L, produce particularly simple
results then we can ‘show a profit’ on the operation. In particular, if the set
581
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
consists of those functions yifor which
Lyi(x)=λiyi(x), (17.2)
where λiis a constant, then a distinct advantage may be obtained from the
manoeuvre because all the differentiation will have disappeared from (17.1).
Equation (17.2) is clearly reminiscent of the equation satisfied by the eigenvec-
torsxiof a linear operator A,n a m e l y
Axi=λixi, (17.3)
where λiis a constant and is called the eigenvalue associated with xi. By analogy,
in the context of differential equations a function yi(x) satisfying (17.2) is called
aneigenfunction of the operator Landλiis then called the eigenvalue associated
with the eigenfunction yi(x).
Probably the most familiar equation of the form (17.2) is that which describes
a simple harmonic oscillator, i.e.
Ly≡−d2y
dt2=ω2y,where L≡−d2/dt2. (17.4)
In this case the eigenfunctions are given by yn(t)=Aneiωnt,w h e r e ωn=2πn/T,
Tis the period of oscillation, n=0,±1,±2,...and the Anare constants. The
eigenvalues are ω2
n=n2ω2
1=n2(2π/T)2. (Sometimes ωnis referred to as the
eigenvalue of this equation but we will avoid this confusing terminology here.)
Another equation of the form (17.2) is Legendre’s equation
Ly≡−(1−x2)d2y
dx2+2xdy
dx=/lscript(/lscript+1 )y, (17.5)
where
L=−(1−x2)d2
dx2+2xd
dx. (17.6)
We found the eigenfunctions of Lby a series method in chapter 16, and for
solutions to Legendre’s equation that are regular at x=±1 these are the
Legendre polynomials, given by
y/lscript(x)=P/lscript(x)=1
2/lscript/lscript!d/lscript
dx/lscript(x2−1)/lscript(17.7)
for/lscript=0,1,2,...; they have associated eigenvalues /lscript(/lscript+1). (Again, /lscriptis sometimes,
confusingly, referred to as the eigenvalue of this equation.)
We may discuss a somewhat wider class of differential equations by considering
a slightly more general form of (17.2), namely
Ly(x)=λρ(x)y(x), (17.8)
where ρ(x)i sa weight function . In many applications ρ(x) is unity for all x,i n
which case (17.2) is recovered; in general, though, it is a function determined by
582
17.1 SETS OF FUNCTIONS
the choice of coordinate system used in describing a particular physical situation.
The only requirement on ρ(x) is that it is real and does not change sign in the
range a≤x≤b, so that it can, without loss of generality, be taken to be non-
negative throughout. A function y(x) that satisfies (17.8) is called an eigenfunction
of the operator Lwith respect to the weight function ρ(x).
This chapter will not cover methods used to determine the eigenfunctions of
(17.2) or (17.8), since we have discussed these in previous chapters, but, rather,will use the properties of the eigenfunctions to solve inhomogeneous equationsof the form (17.1). We shall see later that the sets of eigenfunctions y
i(x)o f
a particular class of operators called Hermitian operators (the operators in the
simple harmonic oscillator equation and in Legendre’s equation are examples)
have particularly useful properties and these will be studied in detail. I turns
out that many of the interesting operators met with in the physical sciences areHermitian. Before continuing our discussion of the eigenfunctions of Hermitianoperators, however, we will consider the properties of general sets of functions.
17.1 Sets of functions
In chapter 8 we discussed the definition of a vector space but concentrated on
spaces of finite dimensionality. We consider now the infinite -dimensional space
of all reasonably well-behaved functions f(x),g(x),h(x),...on the interval
a≤x≤b. That these functions form a linear vector space can be verified since
the set is closed under
(i) addition, which is commutative and associative, i.e.
f(x)+g(x)=g(x)+f(x),
[f(x)+g(x)]+h(x)=f(x)+[g(x)+h(x)],
(ii) multiplication by a scalar, which is distributive and associative, i.e.
λ[f(x)+g(x)]=λf(x)+λg(x),
λ[µf(x)]=(λµ)f(x),
(λ+µ)f(x)=λf(x)+µf(x).
Furthermore, in such a space
(iii) there exists a ‘null vector’ 0 such that f(x)+0= f(x),
(iv) multiplication by unity leaves any function unchanged, i.e. 1 ×f(x)=f(x),
(v) each function has an associated negative function −f(x) that is such that
f(x)+[−f(x)] = 0.
By analogy with finite-dimensional vector spaces we now introduce a set
of linearly independent basis functions y
n(x),n=0,1,...,∞, such that any
583
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
‘reasonable’ function in the interval a≤x≤b(i.e. it obeys the Dirichlet conditions
discussed in chapter 12) can be expressed as the linear sum of these functions:
f(x)=∞summationdisplay
n=0cnyn(x).
Clearly if a different set of linearly independent basis functions zn(x) is chosen
then the function can be expressed in terms of the new basis,
f(x)=∞summationdisplay
n=0dnzn(x),
where the dnare a different set of coefficients. In each case, provided the basis
functions are linearly independent, the coefficients are unique.
We may also define an inner product on our function space by
/angbracketleftf|g/angbracketright=integraldisplayb
af∗(x)g(x)ρ(x)dx, (17.9)
where ρ(x) is the weight function, which we require to be real and non-negative
in the interval a≤x≤b. As mentioned above, ρ(x) is often unity for all x.T w o
functions are said to be orthogonal on the interval [ a, b]i f
/angbracketleftf|g/angbracketright=integraldisplayb
af∗(x)g(x)ρ(x)dx=0, (17.10)
and the normof a function is defined as
/bardblf/bardbl=/angbracketleftf|f/angbracketright1/2=bracketleftbiggintegraldisplayb
af∗(x)f(x)ρ(x)dxbracketrightbigg1/2
=bracketleftbiggintegraldisplayb
a|f(x)|2ρ(x)dxbracketrightbigg1/2
.(17.11)
An infinite-dimensional vector space of functions, for which an inner product
is defined, is called a Hilbert space . Using the concept of the inner product we
can choose a basis of linearly independent functions φn(x),n=0,1,2,...,t h a t
are orthonormal, i.e. such that
/angbracketleftφi|φj/angbracketright=integraldisplayb
aφ∗
i(x)φj(x)ρ(x)dx=δij. (17.12)
Ifyn(x),n=0,1,2,..., are a linearly independent, but not orthonormal, basis
for the Hilbert space then an orthonormal set of basis functions φnmay be
produced (in a similar manner to that used in the construction of a set of
orthogonal eigenvectors of an Hermitian matrix, see chapter 8) by the followingprocedure, in which each of the new functions ψ
nis to be normalised, giving
584
17.1 SETS OF FUNCTIONS
φn=ψn/angbracketleftψn|ψn/angbracketright−1/2, before proceeding to the construction of the next one:
ψ0=y0,
ψ1=y1−φ0/angbracketleftφ0|y1/angbracketright,
ψ2=y2−φ1/angbracketleftφ1|y2/angbracketright−φ0/angbracketleftφ0|y2/angbracketright,
...
ψn=yn−φn−1/angbracketleftφn−1|yn/angbracketright−···−φ0/angbracketleftφ0|yn/angbracketright
...
It is straightforward to check that each φn=ψn/angbracketleftψn|ψn/angbracketright−1/2is orthogonal to
all its predecessors φi,i=0,1,2,...,n−1. This method is called Gram–Schmidt
orthogonalisation . Clearly the functions ψnalso form an orthogonal set, but in
general they do not have unit norms.IStarting from the linearly independent functions yn(x)=xn,n=0,1,..., construct the
first three orthonormal functions over the range −1<x< 1.
The first unnormalised function ψ0is simply equal to the first of the original functions, i.e.
ψ0=1.
The normalisation is carried out by dividing by
/angbracketleftψ0|ψ0/angbracketright1/2=
/Z1
−11×1du
/1/2
=√
2,
with the result that the first normalised function φ0is given by
φ0=ψ0√
2=
q
1
2.
The second unnormalised function is found by applying the above Gram–Schmidt orthog-
onalisation procedure, i.e.
ψ1=y1−φ0/angbracketleftφ0|y1/angbracketright.
It can easily be shown that /angbracketleftφ0|y1/angbracketright=0 ,a n ds o ψ1=x. Normalising then gives
φ1=ψ1
/Z1
−1u×ud u
/−1/2
=
q
3
2x.
The third unnormalised function is similarly given by
ψ2=y2−φ1/angbracketleftφ1|y2/angbracketright−φ0/angbracketleftφ0|y2/angbracketright
=x2−0−1
3,
which, on normalising, gives
φ2=ψ2
/Z1
−1
/;
u2−1
3
/2du
/−1/2
=1
2
q
5
2(3x2−1).
By comparing the functions φ0,φ1andφ2, with the list in subsection 16.6.1, we see that
this procedure has generated (multiples of) the first three Legendre polynomials.
J
585
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
If a function is expressed in terms of an orthonormal basis φn(x)a s
f(x)=∞summationdisplay
n=0anφn(x) (17.13)
then the coefficients anare given by
an=/angbracketleftφn|f/angbracketright=integraldisplayb
aφ∗
n(x)f(x)ρ(x)dx. (17.14)
Note that this is true only if the basis is orthonormal.
17.1.1 Some useful inequalities
Since for a Hilbert space /angbracketleftf|f/angbracketright≥0, the inequalities discussed in subsection 8.1.3
hold. The proofs are not repeated here, but the relationships are listed forcompleteness.
(i) The Schwarz inequality states that
|/angbracketleftf|g/angbracketright|≤/angbracketleft f|f/angbracketright
1/2/angbracketleftg|g/angbracketright1/2, (17.15)
where the equality holds when f(x) is a scalar multiple of g(x), i.e. when
they are linearly dependent.
(ii) The triangle inequality states that
/bardblf+g/bardbl≤/bardbl f/bardbl+/bardblg/bardbl, (17.16)
where again equality holds when f(x) is a scalar multiple of g(x).
(iii) Bessel’s inequality requires the introduction of an orthonormal basis φn(x)
so that any function f(x) can be written as
f(x)=∞summationdisplay
n=0cnφn(x),
where cn=/angbracketleftφn|f/angbracketright. Bessel’s inequality then states that
/angbracketleftf|f/angbracketright≥summationdisplay
n|cn|2. (17.17)
The equality holds if the summation is over all the basis functions. If some
values of nare omitted from the sum then the inequality results (unless,
of course, the cnhappen to be zero for all values of nomitted, in which
case the equality remains).
586
17.2 ADJOINT AND HERMITIAN OPERATORS
17.2 Adjoint and Hermitian operators
Having discussed general sets of functions we now return to the discussion of
eigenfunctions of linear operators. The adjoint of an operator L, denoted by L†,
is defined by
integraldisplayb
af(x)∗[Lg(x)]ρ(x)dx=braceleftbiggintegraldisplayb
ag∗(x)bracketleftbig
L†f(x)bracketrightbig
ρ(x)dxbracerightbigg∗
, (17.18)
or, in inner product notation, /angbracketleftf|Lg/angbracketright=/angbracketleftg|L†f/angbracketright∗.A no p e r a t o ri st h e ns a i dt ob e
self-adjoint orHermitian ifL†=L,i . e .i f
integraldisplayb
af∗(x)[Lg(x)]ρ(x)dx=braceleftbiggintegraldisplayb
ag∗(x)[Lf(x)]ρ(x)dxbracerightbigg∗
, (17.19)
or, in inner product notation, /angbracketleftf|Lg/angbracketright=/angbracketleftg|Lf/angbracketright∗. From (17.19) we note that, when
applied to an Hermitian operator, the general property /angbracketleftb|a/angbracketright∗=/angbracketlefta|b/angbracketrighttakes the
form
/angbracketleftg|Lf/angbracketright∗=/angbracketleftLf|g/angbracketright⇒/angbracketleft Lf|g/angbracketright=/angbracketleftf|Lg/angbracketright=/angbracketleftf|L|g/angbracketright,
where the notation of the final equality emphasises that Lcan act on either forg
without changing the value of the inner product. A little careful study will revealthe similarity between the definition of an Hermitian operator and the definitionof an Hermitian matrix given in chapter 8. In general, however, an operator L
is Hermitian over an interval a≤x≤bonly if certain boundary conditions are
met by the functions fandgon which it acts.IFind the required boundary conditions for the linear operator L=d2/dt2to be Hermitian
over the interval t0tot0+T.
Substituting into the LHS of the definition of an Hermitian operator (17.19) and integrating
by parts givesZt0+T
t0f∗d2g
dt2dt=
/
f∗dg
dt
/t0+T
t0−
Zt0+T
t0df∗
dtdg
dtdt,
where we have taken the weight function ρ(x) to be unity. Integrating the second term on
the RHS by parts yieldsZt0+T
t0f∗d2g
dt2dt=
/
f∗dg
dt
/t0+T
t0+
/
−df∗
dtg
/t0+T
t0+
Zt0+T
t0gd2f∗
dt2dt.
Remembering that the operator is real and taking the complex conjugate outside the
integral givesZt0+T
t0f∗d2g
dt2dt=
/
f∗dg
dt
/t0+T
t0−
/df∗
dtg
/t0+T
t0+
/Zt0+T
t0g∗d2f
dt2dt
/∗
,
which, by comparison with (17.19), proves that Lis Hermitian provided/
f∗dg
dt
/t0+T
t0=
/df∗
dtg
/t0+T
t0.
J
587
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
We showed in chapter 8 that the eigenvalues of Hermitian matrices are real and
that their eigenvectors can be chosen to be orthogonal. Similarly, the eigenvaluesof Hermitian operators are real and their eigenfunctions can be chosen to beorthogonal (we will prove these properties in the following section). Hermitian
operators (or matrices) are often used in the formulation of quantum mechanics.
The eigenvalues then give the possible measured values of an observable quantitysuch as energy or angular momentum, and the physical requirement that suchquantities must be real is ensured by the reality of these eigenvalues. Furthermore,the infinite set of eigenfunctions of an Hermitian operator form a complete basisset, so that it is possible to expand in an eigenfunction series any function y(x)
obeying the appropriate conditions:
y(x)=
∞summationdisplay
n=0cnyn(x), (17.20)
where the choice of suitable values for the cnwill make the sum arbitrarily close
toy(x).†These useful properties provide the motivation for a detailed study of
Hermitian operators.
17.3 The properties of Hermitian operators
We now provide proofs of some of the useful properties of Hermitian operators.
Again much of the analysis is similar to that for Hermitian matrices in chapter 8,
although the present section stands alone. (Here, and throughout the remainderof this chapter, we will write out inner products in full. We note, however,that the inner product notation often provides a neat form in which to expressresults.)
17.3.1 Reality of the eigenvalues
Consider an Hermitian operator for which (17.8) is satisfied by at least two
eigenfunctions y
i(x)a n d yj(x), which have eigenvalues λiandλjrespectively, so
that
Lyi=λiρ(x)yi, (17.21)
Lyj=λjρ(x)yj, (17.22)
where ρ(x) is the weight function. Multiplying (17.21) by y∗
jand (17.22) by y∗
i
†The proof of the completeness of the eigenfunctions of an Hermitian operator is beyond the scope
of this book. The reader should refer to e.g. Courant and Hilbert, Methods of Mathematical Physics
(Interscience Publishers, 1953).
588
17.3 THE PROPERTIES OF HERMITIAN OPERATORS
and then integrating gives
integraldisplayb
ay∗
jLyidx=λiintegraldisplayb
ay∗
jyiρd x , (17.23)
integraldisplayb
ay∗
iLyjdx=λjintegraldisplayb
ay∗
iyjρd x . (17.24)
Remembering that we have required ρ(x) to be real, the complex conjugate of
(17.23) becomes
bracketleftbiggintegraldisplayb
ay∗
jLyidxbracketrightbigg∗
=λ∗
iintegraldisplayb
ay∗
iyjρd x , (17.25)
and using the definition of an Hermitian operator (17.19) it follows that the LHS
of (17.25) is equal to the LHS of (17.24). Thus
(λ∗
i−λj)integraldisplayb
ay∗
iyjρd x=0. (17.26)
Ifi=jthen λi=λ∗
i(sinceintegraltextb
ay∗
iyiρd x/negationslash= 0), which is a statement that the
eigenvalue λiis real.
17.3.2 Orthogonality of the eigenfunctions
From (17.26), it is immediately apparent that two eigenfunctions yiandyjthat
correspond to different eigenvalues, i.e. such that λi/negationslash=λj,s a t i s f y
integraldisplayb
ay∗
iyjρd x=0, (17.27)
which is a statement of the orthogonality of yiandyj. Because Lis linear, the
normalisation of the eigenfunctions yi(x) is arbitrary and we shall assume for
definiteness that they are normalised so thatintegraltextb
ay∗
iyiρd x= 1. Thus we can write
(17.27) in the form
integraldisplayb
ay∗
iyjρd x=δij, (17.28)
which is valid for all pairs of values i, j.
If one (or more) of the eigenvalues is degenerate, however, we have different
eigenfunctions corresponding to the same eigenvalue, and the proof of orthogo-nality is not so straightforward. Nevertheless, an orthogonal set of eigenfunctionsmay be constructed using the Gram–Schmidt orthogonalisation method mentioned
earlier in this chapter and used in chapter 8 to construct a set of orthogonal
eigenvectors of an Hermitian matrix. We repeat the analysis here for complete-ness.
589
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
Suppose, for the sake of our proof, that λ0isk-fold degenerate, i.e.
Lyi=λ0ρyifori=0,1,...,k−1, (17.29)
but that λ0is different from any of λk,λk+1, etc. Then any linear combination of
these yiis also an eigenfunction with eigenvalue λ0since
Lz≡Lk−1summationdisplay
i=0ciyi=k−1summationdisplay
i=0ciLyi=k−1summationdisplay
i=0ciλ0ρyi=λ0ρz. (17.30)
If the yidefined in (17.29) are not already mutually orthogonal then consider
the new eigenfunctions ziconstructed by the following procedure, in which each
of the new functions wiis to be normalised, to give zi, before proceeding to the
construction of the next one (the normalisation can be carried out by dividing
the eigenfunction wiby (integraltextb
aw∗
iwiρd x)1/2):
w0=y0,
w1=y1−parenleftbigg
z0integraldisplayb
az∗
0y1ρd xparenrightbigg
,
w2=y2−parenleftbigg
z1integraldisplayb
az∗
1y2ρd xparenrightbigg
−parenleftbigg
z0integraldisplayb
az∗
0y2ρd xparenrightbigg
,
...
wk−1=yk−1−parenleftbigg
zk−2integraldisplayb
az∗
k−2yk−1ρd xparenrightbigg
−···−parenleftbigg
z0integraldisplayb
az∗
0yk−1ρd xparenrightbigg
.
Each of the integrals is just a number and thus each new function zi=
wi(integraltextb
aw∗
iwiρd x)−1/2is, as can be shown from (17.30), an eigenvector of Lwith
eigenvalue λ0. It is straightforward to check that each ziis orthogonal to all its
predecessors. Thus, by this explicit construction we have shown that an orthog-onal set of eigenfunctions of an Hermitian operator Lcan be obtained. Clearly
the orthonormal set obtained, z
i, is not unique.
17.3.3 Construction of real eigenfunctions
Recall that the eigenfunction yisatisfies
Lyi=λiρyi (17.31)
and that the complex conjugate of this gives
Ly∗
i=λ∗
iρy∗
i=λiρy∗
i, (17.32)
where the last equality follows because the eigenvalues are real, i.e. λi=λ∗
i. Thus,
yiandy∗
iare eigenfunctions corresponding to the same eigenvalue and hence,
because of the linearity of L,a tl e a s to n eo f y∗
i+yiandi(y∗
i−yi) (which are both
590
17.4 STURM–LIOUVILLE EQUATIONS
real) is a non-zero eigenfunction corresponding to that eigenvalue. Therefore the
eigenfunctions can always be made real by taking suitable linear combinations.Such linear combinations will only be necessary in cases where a particular λis
degenerate, i.e. corresponds to more than one linearly independent eigenfunction.
17.4 Sturm–Liouville equations
One of the most important applications of our discussion of Hermitian operators
is to the study of Sturm–Liouville equations , which take the general form
p(x)d
2y
dx2+r(x)dy
dx+q(x)y+λρ(x)y=0,where r(x)=dp(x)
dx(17.33)
andp,qandrare real functions of x. (We note that sign conventions vary in this
expression for the general Sturm–Liouville equation; some authors use −λρ(x)y
on the LHS of (17.33).) A variational approach to the Sturm–Liouville equation,which is useful in estimating the eigenvalues λof the equation, is discussed
in chapter 22. For now, however, we concentrate on a demonstration that the
Sturm–Liouville equation can be solved by superposition methods.
It is clear that (17.33) can be written
Ly=λρ(x)ywhere L=−bracketleftbigg
p(x)d
2
dx2+r(x)d
dx+q(x)bracketrightbigg
. (17.34)
An example is Legendre’s equation (17.5), which is a Sturm–Liouville equation
with p(x)=1−x2,r(x)=−2x=p/prime(x),q(x)=0 , ρ(x) = 1 and eigenvalues
/lscript(/lscript+1 ) .
It will be seen that the general Sturm–Liouville equation (17.33) can be rewritten
(py/prime)/prime+qy+λρy=0, (17.35)
where primes denote differentiation with respect to x. Using (17.34) this may also
be written Ly=−(py/prime)/prime−qy=λρy. We will show in the next section that, under
certain boundary conditions on the solutions y(x), linear operators that can be
w r i t t e ni nt h i sf o r ma r e self-adjoint .
Whilst it is true that Sturm–Liouville equations represent only a small fraction
of the differential equations encountered in practice, as we shall demonstrate insubsection 17.4.2 anysecond-order differential equation of the form
p(x)y
/prime/prime+r(x)y/prime+q(x)y+λρ(x)y= 0 (17.36)
can be converted into Sturm–Liouville form by multiplying through by a suitable
factor; this is discussed in subsection 17.4.2.
591
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
17.4.1 Valid boundary conditions
For the linear operator of the Sturm–Liouville equation (17.34) to be Hermitian
over the range [ a, b] requires certain boundary conditions to be met, namely, that
any two eigenfunctions yiandyjof (17.34) must satisfy
bracketleftbig
y∗
ipy/prime
jbracketrightbig
x=a=bracketleftbig
y∗
ipy/prime
jbracketrightbig
x=bfor all i, j. (17.37)
Rearranging (17.37) we find that
bracketleftBig
y∗
ipy/prime
jbracketrightBigx=b
x=a=0, (17.38)
is an equivalent statement of the required boundary conditions. These boundary
conditions are in fact not too restrictive and are met, for instance, by the sets
y(a)=y(b)=0 ; y(a)=y/prime(b)=0 ; p(a)=p(b) = 0 and by many other sets. It
is important to note that in order to satisfy (17.37) and (17.38) one boundarycondition must be specified at each end of the range.IProve that the Sturm–Liouville operator is Hermitian over the range [a, b]and under the
boundary conditions (17.38).
Putting the Sturm–Liouville form Ly=−(py/prime)/prime−qyinto the definition (17.19) of an
Hermitian operator, the LHS may be written as a sum of two terms, i.e.
−
Zb
a
/
y∗
i(py/prime
j)/prime+y∗
iqyj
/
dx=−
Zb
ay∗
i(py/prime
j)/primedx−
Zb
ay∗
iqyjdx.
The first term may be integrated by parts to give
−
/
y∗
ipy/prime
j
/b
a+
Zb
a(y∗
i)/primepy/prime
jdx.
The first term is zero because of the boundary conditions, and thus, integrating by parts
again yields/
(y∗
i)/primepyj
/b
a−
Zb
a((y∗
i)/primep)/primeyjdx.
The first term is once again zero. Thus
−
Zb
a
/
y∗
i(py/prime
j)/prime+y∗
iqyj
/
dx=
Zb
a
/
−((y∗
i)/primep)/primeyj−y∗
iqyj
/
dx,
=
/
−
Zb
a
/
y∗
j(py/prime
i)/prime+y∗
jqyi
/
dx
/∗
,
which proves that the Sturm–Liouville operator is Hermitian over the prescribed interval.
J
17.4.2 Putting an equation into Sturm–Liouville form
The Sturm–Liouville equation (17.33) requires that r(x)=p/prime(x). However, any
equation of the form
p(x)y/prime/prime+r(x)y/prime+q(x)y+λρ(x)y=0, (17.39)
592
17.5 EXAMPLES OF STURM–LIOUVILLE EQUATIONS
can be put into self-adjoint form by multiplying through by the integrating factor
F(x)=e x pbraceleftbiggintegraldisplayxr(z)−p/prime(z)
p(z)dzbracerightbigg
. (17.40)
It is easily verified that (17.39) then takes the Sturm–Liouville form
[F(x)p(x)y/prime]/prime+F(x)q(x)y+λF(x)ρ(x)y=0, (17.41)
with a different, but still non-negative, weight function F(x)ρ(x).IPut the Hermite equation
y/prime/prime−2xy/prime+2αy=0
into Sturm–Liouville form.
Using (17.40), with p(z)=1 , p/prime(z)=0a n d r(z)=−2zgives the integrating factor
F(x)=e x p
/Zx
−2zd z
/
=e x p
/;
−x2
/
.
Thus, the Hermite equation becomes
e−x2y/prime/prime−2xe−x2y/prime+2αe−x2y=(e−x2y/prime)/prime+2αe−x2y=0,
which is clearly in Sturm–Liouville form with p(x)=e−x2,q(x)=0 , ρ(x)=e−x2and
λ=2α.
J
17.5 Examples of Sturm–Liouville equations
In order to illustrate the wide applicability of Sturm–Liouville theory, in this
section we present a short catalogue of some common equations of Sturm–Liouville form. Many of them have already been discussed in chapter 16. Inparticular the reader should note the orthogonality properties of the varioussolutions, which, in each case, follow because the differential operator is self-
adjoint. For completeness we also quote the associated generating functions.
17.5.1 Legendre’s equation
We have already met Legendre’s equation ,
(1−x
2)y/prime/prime−2xy/prime+/lscript(/lscript+1 )y=[ ( 1−x2)y/prime]/prime+/lscript(/lscript+1 )y= 0 (17.42)
and shown that it is a Sturm–Liouville equation with p(x)=1−x2,q(x)=0 ,
ρ(x) = 1 and eigenvalues /lscript(/lscript+1). In the previous chapter we found the solutions
of Legendre’s equation that are regular for all finite x. These are the Legendre
polynomials P/lscript(x), which are given by a Rodrigues’ formula:
P/lscript(x)=1
2/lscript/lscript!d/lscript
dx/lscript(x2−1)/lscript.
593
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
The orthogonality and normalisation of the functions in the interval −1≤x≤1
is expressed by
integraldisplay1
−1P/lscript(x)Pk(x)dx=2
2/lscript+1δ/lscriptk.
The generating function is
G(x, h)=( 1−2xh+h2)−1/2=∞summationdisplay
n=0Pn(x)hn.
Legendre’s equations appear in the analysis of physical situations involving the
operator∇2and axial symmetry, since the linear differential operator involved has
the form of the polar-angle part of ∇2, when the latter is expressed in spherical
polar coordinates. Examples include the solution of Laplace’s equation in axially
symmetric situations and the solution of the Schr ¨odinger equation for a quantum
mechanical system involving a central potential.
17.5.2 The associated Legendre equation
Very closely related to the Legendre equation is the associated Legendre equation
[(1−x2)y/prime]/prime+bracketleftbigg
/lscript(/lscript+1 )−m2
1−x2bracketrightbigg
y=0, (17.43)
which reduces to Legendre’s equation when m= 0. In physical applications
−/lscript≤m≤/lscriptand mis restricted to integer values. If y(x)i sas o l u t i o no f
Legendre’s equation then
w(x)=( 1−x2)|m|/2d|m|y
dx|m|
is a solution of the associated equation. The solutions of the associated Legendre
equation that are regular for all finite xare called the associated Legendre functions
and are therefore given by
Pm
/lscript(x)=( 1−x2)|m|/2d|m|P/lscript
dx|m|.
Note also that Pm
/lscript(x)=0f o r m>/lscript . Like the Legendre polynomials, the associated
Legendre functions Pm
/lscript(x) are orthogonal in the range −1≤x≤1. This property,
and their normalisation, is expressed by
integraldisplay1
−1Pm
/lscript(x)Pm
k(x)dx=2
2/lscript+1(/lscript+m)!
(/lscript−m)!δ/lscriptk.
They have the generating function
G(x, h)=(2m)!(1−x2)m/2
2mm!(1−2hx+h2)m+1/2=∞summationdisplay
n=0Pm
n+m(x)hn.
594
17.5 EXAMPLES OF STURM–LIOUVILLE EQUATIONS
The associated Legendre equation arises in physical situations in which there
is a dependence on azimuthal angle φof the form eimφor cos mφ.
17.5.3 Bessel’s equation
Physical situations that when described in spherical polar coordinates give rise
to Legendre and associated Legendre equations lead to Bessel’s equation when
cylindrical polar coordinates are used. Bessel’s equation has the form
x2y/prime/prime+xy/prime+(x2−n2)y=0, (17.44)
but on dividing by xand changing variables to ξ=x/a,†it takes on the
Sturm-Liouville form
(ξy/prime)/prime+a2ξy+−n2
ξy=0, (17.45)
where a prime now indicates differentiation with respect to ξ.
We met Bessel’s equation in chapter 16, where we saw that those of its solutions
that are regular for finite xare the Bessel functions, given by
Jn(x)=∞summationdisplay
r=0(−1)r(1
2x)n+2r
r!Γ(n+r+1 ), (17.46)
where Γ is the gamma function discussed in the Appendix. Their orthogonality
and normalisation over the range 0 ≤x<∞have been discussed in detail in
chapter 16. The generating function for the Bessel functions is
G(x, h)=e x pbracketleftbiggx
2parenleftbigg
h−1
hparenrightbiggbracketrightbigg
=∞summationdisplay
n=−∞Jn(x)hn. (17.47)
17.5.4 The simple harmonic equation
The most trivial of Sturm–Liouville equations is the simple harmonic motion
equation
y/prime/prime+ω2y=0, (17.48)
which has p(x)=1 , q(x)=0 , ρ(x) = 1 and eigenvalue ω2. We have already
met the solutions of this equation in the Fourier analysis of chapter 12, and the
properties of orthogonality and normalisation of the eigenfunctions given there
can now be seen in the wider context of general Sturm–Liouville equations.
†This change of scale is required to give the conventional normalisation, but is not needed for the
transformation into Sturm–Liouville form.
595
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
17.5.5 Hermite’s equation
The Hermite equation appears in the description of the wavefunction of a
harmonic oscillator and is given by
y/prime/prime−2xy/prime+2αy=0. (17.49)
We have already seen that it can be converted to Sturm–Liouville form by
multiplying by the integrating factor exp( −x2), which yields
e−x2y/prime/prime−2xe−x2y/prime+2αe−x2y=(e−x2y/prime)/prime+2αe−x2y=0. (17.50)
The solutions, the Hermite polynomials Hn(x), are given by a Rodrigues’
formula:
Hn(x)=(−1)nex2dn
dxnparenleftBig
e−x2parenrightBig
. (17.51)
Their orthogonality over the range −∞<x<∞and their normalisation are
summarised by
integraldisplay∞
−∞e−x2Hm(x)Hn(x)dx=2nn!√πδmn, (17.52)
and their generating function is
G(x, h)=e2hx−h2=∞summationdisplay
n=0Hn(x)
n!hn. (17.53)
17.5.6 Laguerre’s equation
The Laguerre equation appears in the description of the wavefunction of the
hydrogen atom and is given by
xy/prime/prime+( 1−x)y/prime+ny=0. (17.54)
It can be converted to Sturm–Liouville form by multiplying by the integrating
factor exp(−x), which yields
xe−xy/prime/prime+( 1−x)e−xy/prime+ne−xy=(xe−xy/prime)/prime+ne−xy=0. (17.55)
The solutions, the Laguerre polynomials Ln(x), are again given by a Rodrigues’
formula:
Ln(x)=exdn
dxnparenleftbig
xne−xparenrightbig
. (17.56)
Their orthogonality over the range 0 ≤x<∞and their normalisation are
expressed by
integraldisplay∞
0e−xLm(x)Ln(x)dx=(n!)2δmn, (17.57)
596
17.6 SUPERPOSITION OF EIGENFUNCTIONS: GREEN’S FUNCTIONS
and their generating function is
G(x, h)=e−xh/(1−h)
1−h=∞summationdisplay
n=0Ln(x)
n!hn. (17.58)
17.5.7 Chebyshev’s equation
The Chebyshev equation
(1−x2)y/prime/prime−xy/prime+n2y= 0 (17.59)
can be converted to an equation of Sturm–Liouville form by multiplying by the
integrating factor (1 −x2)−1/2. Simplifying, this yields
bracketleftBig
(1−x2)1/2y/primebracketrightBig/prime
+n2(1−x2)−1/2y=0. (17.60)
The solutions, the Chebyshev polynomials Tn(x), are once again given by a
Rodrigues’ formula:
Tn(x)=(−2)nn!(1−x2)1/2
(2n)!dn
dxn(1−x2)n−1/2. (17.61)
Their orthogonality over the range −1≤x≤1 and their normalisation are given
by
integraldisplay1
−1(1−x2)−1/2Tm(x)Tn(x)dx=
0f o r m/negationslash=n,
π/2f o r n=m/negationslash=0,
π forn=m=0,(17.62)
and their generating function is
G(x, h)=1−xh
1−2xh+h2=∞summationdisplay
n=0Tn(x)hn. (17.63)
17.6 Superposition of eigenfunctions: Green’s functions
We have already seen that if
Lyn(x)=λnρ(x)yn(x), (17.64)
where Lis an Hermitian operator, then the eigenvalues λnare real and the
eigenfunctions yn(x) are orthogonal (or can be made so). Let us assume that we
know the eigenfunctions yn(x)o fLthat individually satisfy (17.64) and some
imposed boundary conditions (for which Lis Hermitian).
Now let us suppose we wish to solve the inhomogeneous differential equation
Ly(x)=f(x), (17.65)
597
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
subject to the same boundary conditions. Since the eigenfunctions of Lform a
complete set, the full solution, y(x), to (17.65) may be written as a superposition
of eigenfunctions, i.e.
y(x)=∞summationdisplay
n=0cnyn(x), (17.66)
for some choice of the constants cn. Making full use of the linearity of L, we have
f(x)=Ly(x)=LparenleftBigg∞summationdisplay
n=0cnyn(x)parenrightBigg
=∞summationdisplay
n=0cnLyn(x)=∞summationdisplay
n=0cnλnρ(x)yn(x).
(17.67)
Multiplying the first and last terms of (17.67) by y∗
jand integrating, we obtain
integraldisplayb
ay∗
j(z)f(z)dz=∞summationdisplay
n=0integraldisplayb
acnλny∗
j(z)yn(z)ρ(z)dz, (17.68)
w h e r ew eh a v eu s e d zas the integration variable for later convenience. Finally,
using the orthogonality condition (17.28), we see that the integrals on the RHSare zero unless n=j,a n ds oo b t a i n
c
n=1
λnintegraltextb
ay∗
n(z)f(z)dz
integraltextb
ay∗n(z)yn(z)ρ(z)dz. (17.69)
Thus, if we can find all the eigenfunctions of a differential operator then (17.69)
can be used to find the weighting coefficients for the superposition, to give as the
full solution
y(x)=∞summationdisplay
n=01
λnintegraltextb
ay∗
n(z)f(z)dz
integraltextb
ay∗n(z)yn(z)ρ(z)dzyn(x). (17.70)
If the eigenfunctions have already been normalised, so that
integraldisplayb
ay∗
n(z)yn(z)ρ(z)dz=1 f o ra l l n,
and we assume that we may interchange the order of summation and integration,
then (17.70) can be written as
y(x)=integraldisplayb
abraceleftBigg∞summationdisplay
n=0bracketleftbigg1
λnyn(x)y∗
n(z)bracketrightbiggbracerightBigg
f(z)dz.
The quantity in braces, which is a function of xandzonly, is usually written
G(x, z), and is the Green’s function for the problem. With this notation,
y(x)=integraldisplayb
aG(x, z)f(z)dz, (17.71)
598
17.6 SUPERPOSITION OF EIGENFUNCTIONS: GREEN’S FUNCTIONS
where
G(x, z)=∞summationdisplay
n=01
λnyn(x)y∗
n(z). (17.72)
We note that G(x, z) is determined entirely by the boundary conditions and the
eigenfunctions yn, and hence by Litself, and that f(z) depends purely on the
RHS of the inhomogeneous equation (17.65). Thus, for a given Land boundary
conditions we can establish, once and for all, a function G(x, z) that will enable
us to solve the inhomogeneous equation for anyRHS. From (17.72) we also note
that
G(x, z)=G∗(z,x). (17.73)
We have already met the Green’s function in the solution of second-order dif-
ferential equations in chapter 15, as the function that satisfies the equationL[G(x, z)] = δ(x−z) (and the boundary conditions). The formulation given
above is an alternative, though equivalent, one.IFind an appropriate Green’s function for the equation
y/prime/prime+1
4y=f(x),
with boundary conditions y(0) = y(π)=0. Hence, solve for (i) f(x)=s i n 2 xand (ii)
f(x)=x/2.
One approach to solving this problem is to use the methods of chapter 15 and find
a complementary function and particular integral. However, in order to illustrate the
techniques developed in the present chapter we will use the superposition of eigenfunctions,which, as may easily be checked, produces the same solution.
The operator on the LHS of this equation is already self-adjoint under the given
boundary conditions, and so we seek its eigenfunctions. These satisfy the equation
y
/prime/prime+1
4y=λy.
This equation has the familiar solution
y(x)=Asin
/q
1
4−λ
/
x+Bcos
/q
1
4−λ
/
x.
Now, the boundary conditions require that B=0a n ds i n
/
q
1
4−λ
/
π=0 ,a n ds oq
1
4−λ=n,where n=0,±1,±2,....
Therefore, the independent eigenfunctions that satisfy the boundary conditions are
yn(x)=Ansinnx,
where nis any non-negative integer. The normalisation condition further requiresZπ
0A2
nsin2nx dx =1⇒ An=
/2
π
/1/2
.
599
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
Comparison with (17.72) shows that the appropriate Green’s function is therefore given
by
G(x, z)=2
π∞X
n=0sinnxsinnz
1
4−n2.
Case (i). Using (17.71), the solution with f(x)=s i n2 xis given by
y(x)=2
π
Zπ
0
/ ∞X
n=0sinnxsinnz
1
4−n2
/!
sin 2zd z=2
π∞X
n=0sinnx
1
4−n2
Zπ
0sinnzsin2zd z .
Now the integral is zero unless n= 2, in which case it isZπ
0sin22zd z=π
2.
Thus
y(x)=−2
πsin 2x
15/4π
2=−4
15sin 2x
is the full solution for f(x)=s i n2 x. This is, of course, exactly the solution found by using
the methods of chapter 15.
Case (ii). The solution with f(x)=x/2i sg i v e nb y
y(x)=
Zπ
0
/
2
π∞X
n=0sinnxsinnz
1
4−n2
/!
z
2dz=1
π∞X
n=0sinnx
1
4−n2
Zπ
0zsinnz dz.
The integral may be evaluated by integrating by parts, i.e.Zπ
0zsinnz dz=
/"
−zcosnz
n
/#π
0+
Zπ
0cosnz
ndz
=−πcosnπ
n+
/sinnz
n2
/π
0
=−π(−1)n
n.
Forn= 0 the integral is zero, and thus
y(x)=∞X
n=1(−1)n+1sinnx
n
/;1
4−n2
/,
is the full solution for f(x)=x/2. Using the methods of subsection 15.1.2 the solution
is found to be y(x)=2 x−2πsin(x/2), which may be shown to be equal to the above
solution by expanding 2 x−2πsin(x/2) as a Fourier sine series.
J
A useful relation between the eigenfunctions of Lis given by writing
f(x)=summationdisplay
nyn(x)integraldisplayb
ay∗
n(z)f(z)ρ(z)dz
=integraldisplayb
af(z)ρ(z)summationdisplay
nyn(x)y∗
n(z)dz,
and hence
ρ(z)summationdisplay
nyn(x)y∗
n(z)=δ(x−z). (17.74)
600
17.7 A USEFUL GENERALISATION
This is called the completeness orclosure property of the eigenfunctions. It defines
a complete set. If the spectrum of eigenvalues of Lis anywhere continuous then
the eigenfunction yn(x) must be treated as y(n, x) and an integration carried out
over n.
We also note that the RHS of (17.74) is a δ-function and so is only non-zero
when z=x; thus ρ(z) on the LHS can be replaced by ρ(x) if required, i.e.
ρ(z)summationdisplay
nyn(x)y∗
n(z)=ρ(x)summationdisplay
nyn(x)y∗
n(z). (17.75)
17.7 A useful generalisation
Sometimes we encounter inhomogeneous equations of a form slightly more gen-
eral than (17.1), given by
Ly(x)−λρ(x)y(x)=f(x) (17.76)
for some self-adjoint operator L, with ysubject to the appropriate boundary
conditions and λa given (i.e. fixed) constant. To solve this equation we expand
y(x)a n d f(x) in terms of the eigenfunctions yn(x) of the operator L, which satisfy
Lyn(x)=λnρ(x)yn(x).
Firstly, we expand f(x) as follows:
f(x)=∞summationdisplay
n=0yn(x)integraldisplayb
ay∗
n(z)f(z)ρ(z)dz
=integraldisplayb
aρ(z)∞summationdisplay
n=0yn(x)y∗
n(z)f(z)dz. (17.77)
Using (17.75) this becomes
f(x)=integraldisplayb
aρ(x)∞summationdisplay
n=0yn(x)y∗
n(z)f(z)dz
=ρ(x)∞summationdisplay
n=0yn(x)integraldisplayb
ay∗
n(z)f(z)dz. (17.78)
Next, we expand y(x)a sy=summationtext∞
n=0cnyn(x) and seek the coefficients cn. Substi-
tuting this and (17.78) in (17.76) we have
ρ(x)∞summationdisplay
n=0(λn−λ)cnyn(x)=ρ(x)∞summationdisplay
n=0yn(x)integraldisplayb
ay∗
n(z)f(z)dz,
601
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
from which we find that
cn=∞summationdisplay
n=0integraltextb
ay∗
n(z)f(z)dz
λn−λ.
Hence the solution of (17.76) is given by
y=∞summationdisplay
n=0cnyn(x)=∞summationdisplay
n=0yn(x)
λn−λintegraldisplayb
ay∗
n(z)f(z)dz=integraldisplayb
a∞summationdisplay
n=0yn(x)y∗
n(z)
λn−λf(z)dz.
From this we may identify the Green’s function
G(x, z)=∞summationdisplay
n=0yn(x)y∗
n(z)
λn−λ.
We note that if λ=λn,i . e .i f λequals one of the eigenvalues of L,t h e n G(x, z)
becomes infinite and this method runs into difficulty. No solution then existsunless the RHS of (17.76) satisfies the relation
integraldisplay
b
ay∗
n(x)f(x)dx=0.
If the spectrum of eigenvalues of the operator Lis anywhere continuous, the
orthogonality and closure relationships of the eigenfunctions become
integraldisplayb
ay∗
n(x)ym(x)ρ(x)dx=δ(n−m),
integraldisplay∞
0y∗
n(z)yn(x)ρ(x)dn=δ(x−z).
Repeating the above analysis we then find that the Green’s function is given by
G(x, z)=integraldisplay∞
0yn(x)y∗
n(z)
λn−λdn.
17.8 Exercises
17.1 By considering /angbracketlefth|h/angbracketright,w h e r e h=f+λgwith λreal, prove that, for two functions
fandg,
/angbracketleftf|f/angbracketright/angbracketleftg|g/angbracketright≥1
4[/angbracketleftf|g/angbracketright+/angbracketleftg|f/angbracketright]2.
The function y(x) is real and positive for all x. Its Fourier cosine transform ˜yc(k)
is defined by
˜yc(k)=
Z∞
−∞y(x)cos( kx)dx,
and it is given that ˜yc(0) = 1. Prove that
˜yc(2k)≥2[˜yc(k)]2−1.
602
17.8 EXERCISES
17.2 (a) Write the homogeneous Sturm-Liouville eigenvalue equation for which
y(a)=y(b)=0a s
L(y;λ)≡(py/prime)/prime+qy+λρy=0,
where p(x),q(x)a n d ρ(x) are continuously differentiable functions. Show that
ifz(x)a n d F(x)s a t i s f y L(z;λ)=F(x)w i t h z(a)=z(b)=0t h e nZb
ay(x)F(x)dx=0.
(b) Demonstrate the validity of result (a) by direct calculation for the case in
which p(x)=ρ(x)=1 , q(x)=0 , a=−1,b=1a n d z(x)=1−x2.
17.3 Consider the real eigenfunctions yn(x) of a Sturm–Liouville equation
(py/prime)/prime+qy+λρy=0,a≤x≤b
in which p(x),q(x)a n d ρ(x) are continuously differentiable real functions and
p(x) does not change sign in a≤x≤b.T a k e p(x) as positive throughout the
interval, if necessary by changing the signs of all eigenvalues. For a≤x1≤x2≤b,
establish the identity
(λn−λm)
Zx2
x1ρynymdx=
/
ynpy/prime
m−ympy/prime
n
/x2
x1.
Deduce that if λn>λ mthen yn(x) must change sign between two successive zeroes
ofym(x). (The reader may find it helpful to illustrate this result by sketching the
first few eigenfunctions of the system y/prime/prime+λy=0 ,w i t h y(0) = y(π) = 0, and the
Legendre polynomials Pn(z) given in subsection 16.6.1 for n=2,3,4,5.)
17.4 (a) Show that the equation
y/prime/prime+aδ(x)y+λy=0,
with y(±π)=0a n d areal, has a set of eigenvalues λsatisfying
tan(π√λ)=2√λ
a.
(b) Investigate the conditions under which negative eigenvalues, λ=−µ2with µ
real, are possible.
17.5 Express the hypergeometric equation
(x2−x)y/prime/prime+[ ( 1+ α+β)x−γ]y/prime+αβy=0
in Sturm–Liouville form, determining the conditions imposed on xand on the
parameters α,βandγby the boundary conditions and the allowed forms of
weight function.
17.6 (a) Find the solution of (1 −x2)y/prime/prime−2xy/prime+by=f(x) valid in the range −1≤x≤1
and finite at x= 0, in terms of Legendre polynomials.
(b) If b=1 4a n d f(x)=5 x3, find the explicit solution and verify it by direct
substitution.
17.7 Use the generating function for the Legendre polynomials Pn(x) to show thatZ1
0P2n+1(x)dx=(−1)n (2n)!
22n+1n!(n+1 ) !
and that, except for the case n=0 ,Z1
0P2n(x)dx=0.
603
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
17.8 The quantum mechanical wavefunction for a one-dimensional simple harmonic
oscillator in its nth energy level is of the form
ψ(x)=e x p (−x2/2)Hn(x),
where Hn(x)i st h e nth Hermite polynomial. The generating function for the
polynomials (17.53) is
G(x, h)=e2hx−h2=∞X
n=0Hn(x)
n!hn.
(a) Find Hi(x)f o r i=1,2,3,4.
(b) Evaluate by direct calculationZ∞
−∞e−x2Hp(x)Hq(x)dx,
(i) for p=2 , q= 3; (ii) for p=2 , q= 4; (iii) for p=q= 3. Check your
answers against equation (17.52). (You will find it convenient to useZ∞
−∞x2ne−x2dx=(2n)!√π
22nn!
for integer n≥0.)
17.9 The Laguerre polynomials, which are required for the quantum mechanical
description of the hydrogen atom, can be defined by the generating function(equation (17.58))
G(x, h)=e
−hx/(1−h)
1−h=∞X
n=0Ln(x)
n!hn.
By differentiating the equation separately with respect to xand h,a n dr e -
substituting for G(x, h), prove that LnandL/prime
n(=dLn(x)/dx) satisfy the recurrence
relations
L/prime
n−nL/prime
n−1+nLn−1=0,
Ln+1−(2n+1−x)Ln+n2Ln−1=0.
From these two equations and others derived from them, show that Ln(x)s a t i s fi e s
the Laguerre equation
xL/prime/prime
n+( 1−x)L/prime
n+nLn=0.
17.10 Starting from the linearly independent functions 1, x,x2,x3,..., in the range
0≤x<∞, find the first three orthogonal functions φ0,φ1andφ2, with respect
to the weight function ρ(x)=e−x. By comparing your answers with the Laguerre
polynomials generated by the recurrence relation derived in exercise 17.9, deducethe form of φ
3(x).
17.11 Consider the set of functions {f(x)}of the real variable x, defined in the interval
−∞<x<∞,t h a t→0a tl e a s ta sq u i c k l ya s x−1asx→±∞ . For unit weight
function, determine whether each of the following linear operators is Hermitianwhen acting upon {f(x)}:
(a)d
dx+x;( b )−id
dx+x2;( c ) ixd
dx;( d ) id3
dx3.
17.12 The Chebyshev polynomials Tn(x) can be written as
Tn(x)=c o s ( ncos−1x).
604
17.8 EXERCISES
(a) Verify that these functions do satisfy the Chebyshev equation.
(b) Use de Moivre’s theorem to show that an alternative expression is
Tn(x)=nX
reven(−1)r/2n!
(n−r)!r!xn−r(1−x2)r/2.
17.13 A particle moves in a parabolic potential in which its natural angular frequency
of oscillation is 1 /2. At time t= 0 it passes through the origin with velocity v
and is suddenly subjected to an addi tional acceleration of +1 for 0 ≤t≤π/2,
and then−1f o r π/2<t≤π. At the end of this period it is at the origin again.
Apply the results of the worked example in section 17.6 to show that
v=−8
π∞X
m=01
(4m+2 )2−1
4≈−0.81.
17.14 Find an eigenfunction expansion for the solution with boundary conditions
y(0) = y(π) = 0 of the inhomogeneous equation
d2y
dx2+κy=f(x),
where κis a constant and
f(x)=
/(
x, 0≤x≤π/2,
π−x, π/ 2<x≤π.
17.15 (a) Find those eigenfunctions yn(x) of the self-adjoint linear differential operator
d2/dx2that satisfy the boundary conditions yn(0) = yn(π) = 0, and hence
construct its Green’s function G(x, z).
(b) Construct the same Green’s function using the methods of subsection 15.2.5,
showing that it is
G(x, z)=
/(
x(z−π)/π,0≤x≤z,
z(x−π)/π, z≤x≤π.
(c) By expanding the function given in (b) in terms of the eigenfunctions yn(x),
verify that it is the same function as that derived in (a).
17.16 (a) The differential operator Lis defined by
Ly=−d
dx
/
exdy
dx
/
−exy
4.
Determine the eigenvalues λnof the problem
Lyn=λnexyn0<x< 1,
with boundary conditions
y(0) = 0 ,dy
dx+y
2=0 a t x=1.
(b) Find the corresponding unnormalised yn, and also a weight function ρ(x)w i t h
respect to which the ynare orthogonal. Hence, select a suitable normalisation
for the yn.
(c) By making an eigenfunction expansion, solve the equation
Ly=−ex/2,0<x< 1,
subject to the same boundary conditions as previously.
605
EIGENFUNCTION METHODS FOR DIFFERENTIAL EQUATIONS
17.17 Show that the linear operator
L≡1
4(1 +x2)2d2
dx2+1
2x(1 +x2)d
dx+a,
acting upon functions defined in −1≤x≤1 and vanishing at the endpoints of
the interval, is Hermitian with respect to the weight function (1 + x2)−1.
By making the change of variable x=t a n ( θ/2), find two even eigenfunctions,
f1(x)a n d f2(x), of the differential equation
Lu=λu.
17.18 By substituting x=e x p tfind the normalized eigenfunctions yn(x)a n dt h e
eigenvalues λnof the operator Ldefined by
Ly=x2y/prime/prime+2xy/prime+1
4y, 1≤x≤e,
with y(1) = y(e) = 0. Find, as a series
Panyn(x), the solution of Ly=x−1/2.
17.19 Express the solution of Poisson’s equation in electrostatics,
∇2φ(r)=−ρ(r)//epsilon10,
where ρis the non-zero charge density over a finite part of space, in the form of
an integral and hence identify the Green’s function for the ∇2operator.
17.20 In the quantum mechanical study of the scattering of a particle by a potential,
a Born-approximation solution can be obtained in terms of a function y(r)t h a t
satisfies an equation of the form
(−∇2−K2)y(r)=F(r).
Assuming that yk(r)=( 2 π)−3/2exp(ik·r) is a suitably normalised eigenfunction of
−∇2corresponding to eigenvalue −k2, find a suitable Green’s function GK(r,r/prime).
By taking the direction of the vector r−r/primeas the polar axis for a k-space
integration, show that GK(r,r/prime) can be reduced to
1
4π|r−r/prime|
Z∞
−∞wsinw
w2−w2
0dw,
where w0=K|r−r/prime|.
(This integral can be evaluated using a contour integration (chapter 20) to give(4π|r−r
/prime|)−1exp(iK|r−r/prime|).)
17.9 Hints and answers
17.1 Express the condition /angbracketlefth|h/angbracketright≥0 as a quadratic equation in λand then apply the
condition for no real roots, noting that /angbracketleftf|g/angbracketright+/angbracketleftg|f/angbracketrightis real. To put a limit onR
ycos2kx dx,s e tf=y1/2coskxandg=y1/2in the inequality.
17.2 (a) By twice integrating by parts the term containing p, show thatRb
ayL(z;λ)dx=
Rb
azL(y;λ)dx.
(b)y(x)=Acos(√λx)w i t h λ=n2π2/4, and F(x)=λ−2−λx2.
17.3 Follow an argument similar to that in subsection 17.3.1, but integrate from x1to
x2, rather than from atob.T a k e x1andx2as two successive zeroes of ym(x)a n d
note that, if the sign of ymisαthen the sign of y/prime
m(x1)i sαwhilst that of y/prime
m(x2)
is−α. Now assume that yn(x) does not change sign in the interval and has a
constant sign β; show that this leads to a contradiction between the signs of the
two sides of the identity.
17.4 (a) Different combinations of sinusoids are needed for negative and positive
ranges of x.( b ) µmust satisfy tanh µπ=2µ/a,w h i c hr e q u i r e s a>2/π.
606
17.9 HINTS AND ANSWERS
17.5 [ xγ(1−x)α+β−γ+1y/prime]/prime=αβxγ−1(1−x)α+β−γy;0≤x≤1,α+β>γ> 1.
17.6 (a) y=
PanPn(x)w i t h
an=n+1/2
b−n(n+1 )
Z1
−1f(z)Pn(z)dz;
(b) 5x3=2P3(x)+3P1(x), giving a1=1/4a n d a3= 1, leading to y=5 ( 2 x3−x)/4.
17.8 (a) 2 x,4x2−2, 8x3−12x,1 6x4−48x2+ 12; (b) (i) 0, (ii) 0, (iii) 48√π.
17.10 φ0(x)=1 ,φ1(x)=x−1,φ2(x)=(x2−4x+2 )/2;n!φn(x)=(−1)nLn(x);
φ3(x)=(x3−9x2+1 8x−6)/6.
17.11 (a) No,
R
gf∗/primedx/negationslash= 0; (b) yes; (c) no, i
R
f∗gdx/negationslash=0 ;( d )y e s .
17.14 The normalised eigenfunctions are (2 /π)1/2sinnx,w i t h nan integer.
y(x)=( 4 /π)
P
nodd[(−1)(n−1)/2sinnx]/[n2(κ−n2)].
17.15 (a) The normalised eigenfunctions are (2 /π)1/2sinnx,w i t h nan integer.
G(x, z)=(−2/π)
P∞
n=0[sin(nz)sin(nx)]/n2.
17.16 (a) λn=(n+1/2)2π2,n=0,1,2,... .
(b) Since yn(1)y/prime
m(1)/negationslash= 0, the Sturm–Liouville boundary conditions are not sat-
isfied and the appropriate weight function has to be justified by inspection. Thenormalised eigenfunctions are√2e
−x/2sin[(n+1/2)πx], with ρ(x)=ex.
(c)y(x)=(−2/π3)
P∞
n=0e−x/2sin[(n+1/2)πx]/(n+1/2)3.
17.17 In terms of θ,Lisd2/dθ2+aand has eigenfunctions u(θ)=c o s (√
a−λθ), where√
a−λ=2n+1 ;
f1(x)=( 1−x2)/(1 +x2);f2(x)=4 [ ( 1−x2)/(1 +x2)]3−3[(1−x2)/(1 +x2)].
17.18 yn(x)=√
2x−1/2sin(nπlnx)w i t h λn=−n2π2;
an=
/(
−(nπ)−2
Re
1√
2x−1sin(nπlnx)dx=−√
8(nπ)−3fornodd,
0f o r neven.
17.19 G(r,r/prime)=( 4 π|r−r/prime|)−1.
607
18
Partial differential equations:
general and particular solutions
In this chapter and the next the solution of differential equations of types
typically encountered in the physical sciences and engineering is extended tosituations involving more than one independent variable. A partial differentialequation (PDE) is an equation relating an unknown function (the dependentvariable) of two or more variables to its partial derivatives with respect tothose variables. The most commonly occurring independent variables are those
describing position and time, and so we will couch our discussion and examples
in notation appropriate to them.
As in other chapters we will focus our attention on the equations that arise
most often in physical situations. We will restrict our discussion, therefore, tolinear PDEs, i.e. those of first degree in the dependent variable. Furthermore, wewill discuss primarily second-order equations. The solution of first-order PDEs
will necessarily be involved in treating these, and some of the methods discussed
can be extended without difficulty to third- and higher-order equations. We shallalso see that many ideas developed for ordinary differential equations (ODEs)can be carried over directly into the study of PDEs.
In this chapter we will concentrate on general solutions of PDEs in terms
of arbitrary functions and the particular solutions that may be derived fromthem in the presence of boundary conditions. We also discuss the existence and
uniqueness of the solutions to PDEs under given boundary conditions.
In the next chapter the methods most commonly used in practice for obtaining
solutions to PDEs subject to given boundary conditions will be considered. Thesemethods include the separation of variables, integral transforms and Green’sfunctions. This division of material is rather arbitrary and really has been madeonly to emphasise the general usefulness of the latter methods. In particular, it
will be readily apparent that some of the results of the present chapter are in
fact solutions in the form of separated variables, but arrived at by a differentapproach.
608
18.1 IMPORTANT PARTIAL DIFFERENTIAL EQUATIONS
18.1 Important partial differential equations
Most of the important PDEs of physics are second-order and linear. In order to
gain familiarity with their general form, some of the more important ones willnow be briefly discussed. These equations apply to a wide variety of differentphysical systems.
Since, in general, the PDEs listed below describe three-dimensional situations,
the independent variables are randt,w h e r e ris the position vector and tis
time. The actual variables used to specify the position vector rare dictated by the
coordinate system in use. For example, in Cartesian coordinates the independentvariables of position are x,yandz, whereas in spherical polar coordinates they
arer,θandφ. The equations may be written in a coordinate-independent manner,
however, by the use of the Laplacian operator ∇
2.
18.1.1 The wave equation
The wave equation
∇2u=1
c2∂2u
∂t2(18.1)
describes as a function of position and time the displacement from equilibrium,
u(r,t), of a vibrating string or membrane or a vibrating solid, gas or liquid. The
equation also occurs in electromagnetism, where umay be a component of the
electric or magnetic field in an elecromagnetic wave or the current or voltagealong a transmission line. The quantity cis the speed of propagation of the waves.IFind the equation satisfied by small transverse displacements u(x, t)of a uniform string of
mass per unit length ρheld under a uniform tension T, assuming that the string is initially
located along the x-axis in a Cartesian coordinate system.
Figure 18.1 shows the forces acting on an elemental length ∆ sof the string. If the tension
Tin the string is uniform along its length then the net upward vertical force on the
element is
∆F=Tsinθ2−Tsinθ1.
Assuming that the angles θ1andθ2are both small, we may make the approximation
sinθ≈tanθ. Since at any point on the string the slope tan θ=∂u/∂x ,t h ef o r c ec a nb e
written
∆F=T
/∂u(x+∆x, t)
∂x−∂u(x, t)
∂x
/
≈T∂2u(x, t)
∂x2∆x,
where we have used the definition of the partial derivative to simplify the RHS.
This upward force may be equated, by Newton’s second law, to the product of the
mass of the element and its upward acceleration. The element has a mass ρ∆s,w h i c hi s
approximately equal to ρ∆xif the vibrations of the string are small, and so we have
ρ∆x∂2u(x, t)
∂t2=T∂2u(x, t)
∂x2∆x.
609
PDES: GENERAL AND PARTICULAR SOLUTIONS
u
x xT
T∆s
x+∆xθ1θ2
Figure 18.1 The forces acting on an element of a string under uniform
tension T.
Dividing both sides by ∆ xwe obtain, for the vibrations of the string, the one-dimensional
wave equation
∂2u
∂x2=1
c2∂2u
∂t2,
where c2=T/ρ.
J
The longitudinal vibrations of an elastic rod obey a very similar equation to
that derived in the above example, namely
∂2u
∂x2=ρ
E∂2u
∂t2;
here ρis the mass per unit volume and Eis Young’s modulus.
The wave equation can be generalised slightly. For example, in the case of the
vibrating string, there could also be an external upward vertical force f(x, t)p e r
unit length acting on the string at time t. The transverse vibrations would then
satisfy the equation
T∂2u
∂x2+f(x, t)=ρ∂2u
∂t2,
which is clearly of the form ‘upward force per unit length = mass per unit length
×upward acceleration’.
Similar examples, but involving two or three spatial dimensions rather than one,
are provided by the equation governing the transverse vibrations of a stretchedmembrane subject to an external vertical force density f(x, y, t),
Tparenleftbigg∂
2u
∂x2+∂2u
∂y2parenrightbigg
+f(x, y, t)=ρ(x, y)∂2u
∂t2,
where ρis the mass per unit area of the membrane and Tis the tension.
610
18.1 IMPORTANT PARTIAL DIFFERENTIAL EQUATIONS
18.1.2 The diffusion equation
The diffusion equation
κ∇2u=∂u
∂t(18.2)
describes the temperature uin a region containing no heat sources or sinks; it
also applies to the diffusion of a chemical that has a concentration u(r,t). The
constant κis called the diffusivity. The equation is clearly second-order in the
three spatial variables, but first order in time.IDerive the equation satisfied by the temperature u(r,t)at time tfor a material of uniform
thermal conductivity k, specific heat capacity sand density ρ. Express the equation in
Cartesian coordinates.
Let us consider an arbitrary volume Vlying within the solid and bounded by a surface S
(this may coincide with the surface of the solid if so desired). At any point in the solidthe rate of heat flow per unit area in any given direction ˆris proportional to minus the
component of the temperature gradient in that direction and so is given by ( −k∇u)·ˆr.T h e
total flux of heat outof the volume Vper unit time is given by
−dQ
dt=
ZZ
S(−k∇u)·ˆndS
=
ZZ Z
V∇·(−k∇u)dV, (18.3)
where Qis the total heat energy in Vat time tandˆnis the outward-pointing unit normal
toS; note that we have used the divergence theorem to convert the surface integral into
a volume integral.
W ec a na l s oe x p r e s s Qas a volume integral over V,
Q=
ZZZ
Vsρu dV,
and its rate of change is then given by
dQ
dt=
ZZ Z
Vsρ∂u
∂tdV, (18.4)
where we have taken the derivative with respect to time inside the integral (see section 5.12).
Comparing (18.3) and (18.4), and remembering that the volume Vis arbitrary, we obtain
the three-dimensional diffusion equation
κ∇2u=∂u
∂t,
where the diffusion coefficient κ=k/(sρ). To express this equation in Cartesian coordinates,
we simply write ∇2in terms of x,yandzto obtain
κ
/∂2u
∂x2+∂2u
∂y2+∂2u
∂z2
/
=∂u
∂t.
J
The diffusion equation just derived can be generalised to
k∇2u+f(r,t)=sρ∂u
∂t.
611
PDES: GENERAL AND PARTICULAR SOLUTIONS
The second term, f(r,t), represents a varying density of heat sources throughout
the material but is often not required in physical applications. In the most generalcase, k,sandρmay depend on position r, in which case the first term becomes
∇·(k∇u). However, in the simplest application the heat flow is one-dimensional
with no heat sources, and the equation becomes (in Cartesian coordinates)
∂
2u
∂x2=sρ
k∂u
∂t.
18.1.3 Laplace’s equation
Laplace’s equation,
∇2u=0, (18.5)
may be obtained by setting ∂u/∂t = 0 in the diffusion equation (18.2), and
describes (for example) the steady-state temperature distribution in a solid in
which there are no heat sources – i.e. the temperature distribution after a longtime has elapsed.
Laplace’s equation also describes the gravitational potential in a region con-
taining no matter or the electrostatic potential in a charge-free region. Further, it
applies to the flow of an incompressible fluid with no sources, sinks or vortices;in this case uis the velocity potential, from which the velocity is given by v=∇u.
18.1.4 Poisson’s equation
Poisson’s equation,
∇
2u=ρ(r), (18.6)
describes the same physical situations as Laplace’s equation, but in regions
containing matter, charges or sources of heat or fluid. The function ρ(r)i s
called the source density and in physical applications usually contains somemultiplicative physical constants. For example, if uis the electrostatic potential
in some region of space, in which case ρis the density of electric charge, then
∇
2u=−ρ(r)//epsilon10,w h e r e /epsilon10is the permittivity of free space. Alternatively, umight
represent the gravitational potential in some region where the matter density isgiven by ρ;t h e n∇
2u=4πGρ(r), where Gis the gravitational constant.
18.1.5 Schr ¨odinger’s equation
The Schr ¨odinger equation
−/planckover2pi12
2m∇2u+V(r)u=i/planckover2pi1∂u
∂t, (18.7)
612
18.2 GENERAL FORM OF SOLUTION
describes the quantum mechanical wavefunction u(r,t) of a non-relativistic particle
of mass m;/planckover2pi1is Planck’s constant divided by 2 π. Like the diffusion equation it is
second order in the three spatial variables and first order in time.
18.2 General form of solution
Before turning to the methods by which we may hope to solve PDEs such as
those listed in the previous section, it is instructive, as for ODEs in chapter 14, to
study how PDEs may be formed from a set of possible solutions. Such a studycan provide an indication of how equations obtained not from possible solutionsbut from physical arguments might be solved.
For definiteness let us suppose we have a set of functions involving two
independent variables xandy. Without further specification this is of course a
very wide set of functions, and we could not expect to find a useful equation thatthey all satisfy. However, let us consider a type of function u
i(x, y)i nw h i c h xand
yappear in a particular way, such that uican be written as a function (however
complicated) of a single variable p, itself a simple function of xandy.
Let us illustrate this by considering the three functions
u1(x, y)=x4+4 (x2y+y2+1 ),
u2(x, y)=s i n x2cos2y+c o s x2sin 2y,
u3(x, y)=x2+2y+2
3x2+6y+5.
These are all fairly complicated functions of xandyand a single differential
equation of which each one is a solution is not obvious. However, if we observe
that in fact each can be expressed as a function of the variable p=x2+2yalone
(with no other xoryinvolved) then a great simplification takes place. Written
in terms of pthe above equations become
u1(x, y)=(x2+2y)2+4= p2+4= f1(p),
u2(x, y)=s i n ( x2+2y)=s i n p=f2(p),
u3(x, y)=(x2+2y)+2
3(x2+2y)+5=p+2
3p+5=f3(p).
Let us now form, for each ui, the partial derivatives ∂ui/∂xand∂ui/∂y.I ne a c h
case these are (writing both the form for general pand the one appropriate to
o u rp a r t i c u l a rc a s e , p=x2+2y)
∂ui
∂x=dfi(p)
dp∂p
∂x=2xf/prime
i,
∂ui
∂y=dfi(p)
dp∂p
∂y=2f/prime
i,
fori= 1, 2, 3. All reference to the form of fican be eliminated from these
613
PDES: GENERAL AND PARTICULAR SOLUTIONS
equations by cross-multiplication, obtaining
∂p
∂y∂ui
∂x=∂p
∂x∂ui
∂y,
or, for our specific form, p=x2+2y,
∂ui
∂x=x∂ui
∂y. (18.8)
It is thus apparent that not only are the three functions u1,u2u3solutions of the
PDE (18.8) but so also is any arbitrary function f(p) of which the argument phas
the form x2+2y.
18.3 General and particular solutions
In the last section we found that the first-order PDE (18.8) has as a solution any
function of the variable x2+2y. This points the way for the solution of PDEs
of other orders, as follows. It is notgenerally true that an nth-order PDE can
always be considered as resulting from the elimination of narbitrary functions
from its solution (as opposed to the elimination of narbitrary constants for an
nth-order ODE, see section 14.1). However, given specific PDEs we can try to
solve them by seeking combinations of variables in terms of which the solutions
may be expressed as arbitrary functions. Where this is possible we may expect n
combinations to be involved in the solution.
Naturally, the exact functional form of the solution for any particular situation
must be determined by some set of boundary conditions. For instance, if the PDE
contains two independent variables xandythen for complete determination of
its solution the boundary conditions will take a form equivalent to specifyingu(x, y) along a suitable continuum of points in the xy-plane (usually along a line).
We now discuss the general and particular solutions of first- and second-
order PDEs. In order to simplify the algebra, we will restrict our discussionto equations containing just two independent variables xandy. Nevertheless,
the method presented below may be extended to equations containing severalindependent variables.
18.3.1 First-order equations
Although most of the PDEs encountered in physical contexts are second order
(i.e. they contain ∂
2u/∂x2or∂2u/∂x∂y , etc.), we now discuss first-order equations
to illustrate the general considerations involved in the form of the solution andin satisfying any boundary conditions on the solution.
The most general first-order linear PDE (containing two independent variables)
614
18.3 GENERAL AND PARTICULAR SOLUTIONS
is of the form
A(x, y)∂u
∂x+B(x, y)∂u
∂y+C(x, y)u=R(x, y), (18.9)
where A(x, y),B(x, y),C(x, y)a n d R(x, y) are given functions. Clearly, if either
A(x, y)o r B(x, y) is zero then the PDE may be solved straightforwardly as a
first-order linear ODE (as discussed in chapter 14), the only modification being
that the arbitrary constant of integration becomes an arbitrary function ofxory
respectively.IFind the general solution u(x, y)of
x∂u
∂x+3u=x2.
Dividing through by xwe obtain
∂u
∂x+3u
x=x,
which is a linear equation with integra ting factor (see subsection 14.2.4)
exp
/Z3
xdx
/
=e x p ( 3l n x)=x3.
Multiplying through by this factor we find
∂
∂x(x3u)=x4,
which, on integrating with respect to x,g i v e s
x3u=x5
5+f(y),
where f(y)i sa n arbitrary function ofy. Finally, dividing through by x3, we obtain the
solution
u(x, y)=x2
5+f(y)
x3.
J
When the PDE contains partial derivatives with respect to both independent
variables then, of course, we cannot employ the above procedure but must seekan alternative method. Let us for the moment restrict our attention to the specialcase in which C(x, y)=R(x, y) = 0 and, following the discussion of the previous
section, look for solutions of the form u(x, y)=f(p)w h e r e pis some, at present
unknown, combination of xandy. We then have
∂u
∂x=df(p)
dp∂p
∂x,
∂u
∂y=df(p)
dp∂p
∂y,
615
PDES: GENERAL AND PARTICULAR SOLUTIONS
which, when substituted into the PDE (18.9), give
bracketleftbigg
A(x, y)∂p
∂x+B(x, y)∂p
∂ybracketrightbiggdf(p)
dp=0.
This removes all reference to the actual form of the function f(p) since for
non-trivial pwe must have
A(x, y)∂p
∂x+B(x, y)∂p
∂y=0. (18.10)
Let us now consider the necessary condition for f(p) to remain constant as x
andyvary; this is that pitself remains constant. Thus for fto remain constant
implies that xandymust vary in such a way that
dp=∂p
∂xdx+∂p
∂ydy=0. (18.11)
The forms of (18.10) and (18.11) are very alike, and become the same if we
require that
dx
A(x, y)=dy
B(x, y). (18.12)
By integrating this expression the form of pcan be found.IFor
x∂u
∂x−2y∂u
∂y=0, (18.13)
find (i) the solution that takes the value 2y+1on the line x=1, and (ii) a solution that
has the value 4at the point (1,1).
If we seek a solution of the form u(x, y)=f(p), we deduce from (18.12) that u(x, y) will
be constant along lines of ( x, y)t h a ts a t i s f y
dx
x=dy
−2y,
w h i c ho ni n t e g r a t i n gg i v e s x=cy−1/2. Identifying the constant of integration cwith p1/2
(to avoid fractional powers), we conclude that p=x2y. Thus the general solution of the
PDE (18.13) is
u(x, y)=f(x2y),
where fis an arbitrary function.
We must now find the particular solutions that obey each of the imposed boundary
conditions. For boundary condition (i) a little thought shows that the particular solutionrequired is
u(x, y)=2 ( x
2y)+1=2 x2y+1. (18.14)
For boundary condition (ii) some ob viously acceptable solutions are
u(x, y)=x2y+3,
u(x, y)=4 x2y,
u(x, y)=4 .
616
18.3 GENERAL AND PARTICULAR SOLUTIONS
Each is a valid solution (the freedom of choice of form arises from the fact that u
is specified at only one point (1 ,1), and not along a continuum (say), as in boundary
condition (i)). All three are particular exam ples of the general solution, which may be
written, for example, as
u(x, y)=x2y+3+ g(x2y),
where g=g(x2y)=g(p) is an arbitrary function subject only to g(1) = 0. For this
example, the forms of gcorresponding to the particular solutions listed above are g(p)=0 ,
g(p)=3 p−3,g(p)=1−p.
J
As mentioned above, in order to find a solution of the form u(x, y)=f(p)w e
require that the original PDE contains no term in u, but only terms containing
its partial derivatives. If a term in uis present, so that C(x, y)/negationslash= 0 in (18.9),
then the procedure needs some modification, since we cannot simply divide outthe dependence on f(p) to obtain (18.10). In such cases we look instead for
a solution of the form u(x, y)=h(x, y)f(p). We illustrate this method in the
following example.IFind the general solution of
x∂u
∂x+2∂u
∂y−2u=0. (18.15)
We seek a solution of the form u(x, y)=h(x, y)f(p), with the consequence that
∂u
∂x=∂h
∂xf(p)+hdf(p)
dp∂p
∂x,
∂u
∂y=∂h
∂yf(p)+hdf(p)
dp∂p
∂y.
Substituting these expressions into the PDE (18.15) and rearranging, we obtain/
x∂h
∂x+2∂h
∂y−2h
/
f(p)+
/
x∂p
∂x+2∂p
∂y
/
hdf(p)
dp=0.
The first factor in parentheses is just the original PDE with ureplaced by h. Therefore, if
hisanysolution of the PDE, however simple , this term will vanish, to leave/
x∂p
∂x+2∂p
∂y
/
hdf(p)
dp=0,
from which, as in the previous case, we obtain
x∂p
∂x+2∂p
∂y=0.
From (18.11) and (18.12) we see that u(x, y) will be constant along lines of ( x, y)t h a t
satisfy
dx
x=dy
2,
which integrates to give x=cexp(y/2). Identifying the constant of integration cwith pwe
findp=xexp(−y/2). Thus the general solution of (18.15) is
u(x, y)=h(x, y)f(xexp(−1
2y)),
where f(p) is any arbitrary function of pandh(x, y) is any solution of (18.15).
617
PDES: GENERAL AND PARTICULAR SOLUTIONS
If we take, for example, h(x, y)=e x p y, which clearly satisfies (18.15), then the general
solution is
u(x, y)=( e x p y)f(xexp(−1
2y)).
Alternatively, h(x, y)=x2also satisfies (18.15) and so the general solution to the equation
c a na l s ob ew r i t t e n
u(x, y)=x2g(xexp(−1
2y)),
where gis an arbitrary function of p; clearly g(p)=f(p)/p2.
J
18.3.2 Inhomogeneous equations and problems
Let us discuss in a more general form the particular solutions of (18.13) found
in the second example of the previous subsection. It is clear that, so far as thisequation is concerned, if u(x, y) is a solution then so is any multiple of u(x, y)o r
any linear sum of separate solutions u
1(x, y)+u2(x, y). However, when it comes
to fitting the boundary conditions this is not so.
For example, although u(x, y) in (18.14) satisfies the PDE and the boundary
condition u(1,y)=2 y+ 1, the function u1(x, y)=4 u(x, y)=8 xy+ 4, whilst
satisfying the PDE, takes the value 8 y+4 on the line x= 1 and so does not satisfy
the required boundary condition. Likewise the function u2(x, y)=u(x, y)+f1(x2y),
for arbitrary f1, satisfies (18.13) but takes the value u2(1,y)=2 y+1+ f1(y)o n
the line x= 1, and so is not of the required form unless f1is identically zero.
Thus we see that when treating the superposition of solutions of PDEs two
considerations arise, one concerning the equation itself and the other connectedto the boundary conditions. The equation is said to be homogeneous if the fact
thatu(x, y) is a solution implies that λu(x, y), for any constant λ, is also a solution.
However, the problem is said to be homogeneous if, in addition, the boundary
conditions are such that if they are satisfied by u(x, y) then they are also satisfied
byλu(x, y). The last requirement itself is referred to as that of homogeneous
boundary conditions .
For example, the PDE (18.13) is homogeneous but the general first-order
equation (18.9) would not be homogeneous unless R(x, y) = 0. Furthermore,
the boundary condition (i) imposed on the solution of (18.13) in the previous
subsection is not homogeneous though, in this case, the boundary condition
u(x, y) = 0 on the line y=4x
−2
would be, since u(x, y)=λ(x2y−4) satisfies this condition for any λand, being a
function of x2y, satisfies (18.13).
The reason for discussing the homogeneity of PDEs and their boundary condi-
tions is that in linear PDEs there is a close parallel to the complementary-function
and particular-integral property of ODEs. The general solution of an inhomo-geneous problem can be written as the sum of anyparticular solution of the
618
18.3 GENERAL AND PARTICULAR SOLUTIONS
problem and the general solution of the corresponding homogeneous problem (as
for ODEs, we require that the particular solution is not already contained in thegeneral solution of the homogeneous problem). Thus, for example, the generalsolution of
∂u
∂x−x∂u
∂y+au=f(x, y), (18.16)
subject to, say, the boundary condition u(0,y)=g(y), is given by
u(x, y)=v(x, y)+w(x, y),
where v(x, y) is any solution (however simple) of (18.16) such that v(0,y)=g(y)
andw(x, y) is the general solution of
∂w
∂x−x∂w
∂y+aw=0, (18.17)
with w(0,y) = 0. If the boundary conditions are sufficiently specified then the only
possible solution of (18.17) will be w(x, y)≡0a n d v(x, y) will be the complete
solution by itself.
Alternatively, we may begin by finding the general solution of the inhomoge-
neous equation (18.16) without regard for any boundary conditions; it is just the
sum of the general solution to the homogeneous equation and a particular inte-
gral of (18.16), both without reference to the boundary conditions. The boundary
conditions can then be used to find the appropriate particular solution from thegeneral solution.
We will not discuss at length general methods of obtaining particular integrals
of PDEs but merely note that some of those methods available for ordinarydifferential equations can be suitably extended. †IFind the general solution of
y∂u
∂x−x∂u
∂y=3x. (18.18)
Hence find the most general particular solution (i) which satisfies u(x,0) = x2and (ii)
which has the value u(x, y)=2 at the point (1,0).
This equation is inhomogeneous, and so let us first find the general solution of (18.18)
without regard for any boundary conditions. We begin by looking for the solution of thecorresponding homogeneous equation ((18.18) with the RHS equal to zero) of the formu(x, y)=f(p). Following the same procedure as that used in the solution of (18.13) we
find that u(x, y) will be constant along lines of ( x, y)t h a ts a t i s f y
dx
y=dy
−x⇒x2
2+y2
2=c.
Identifying the constant of integration cwith p/2, we find that the general solution of the
†See for example Piaggio, Differential Equations (Bell, 1954), p. 175 et seq.
619
PDES: GENERAL AND PARTICULAR SOLUTIONS
homogeneous equation is u(x, y)=f(x2+y2) for arbitrary function f. Now by inspection
a particular integral of (18.18) is u(x, y)=−3y, and so the general solution to (18.18) is
u(x, y)=f(x2+y2)−3y.
Boundary condition (i) requires u(x,0) = f(x2)=x2,i . e .f(z)=z, and so the particular
solution in this case is
u(x, y)=x2+y2−3y.
Similarly, boundary condition (ii) requires u(1,0) = f(1) = 2. One possibility is f(z)=2 z,
and if we make this choice, then one way of writing the most general particular solutionis
u(x, y)=2 x
2+2y2−3y+g(x2+y2),
where gis any arbitrary function for which g(1) = 0. Alternatively, a simpler choice would
bef(z) = 2, leading to
u(x, y)=2−3y+g(x2+y2).
J
Although we have discussed the solution of inhomogeneous problems only
for first-order equations, the general considerations hold true for linear PDEs ofhigher order.
18.3.3 Second-order equations
As noted in section 18.1, second-order linear PDEs are of great importance in
describing the behaviour of many physical systems. As in our discussion of first-
order equations, for the moment we shall restrict our discussion to equations withjust two independent variables; extensions to a greater number of independentvariables are straightforward.
The most general second-order linear PDE (containing two independent vari-
ables) has the form
A∂
2u
∂x2+B∂2u
∂x∂y+C∂2u
∂y2+D∂u
∂x+E∂u
∂y+Fu=R(x, y), (18.19)
where A ,B,...,F andR(x, y) are given functions of xandy. Because of the nature
of the solutions to such equations, they are usually divided into three classes, adivision of which we will make further use in subsection 18.6.2. The equation
(18.19) is called hyperbolic ifB
2>4AC,parabolic ifB2=4ACandelliptic if
B2<4AC. Clearly, if A,BandCare functions of xandy(rather than just
constants) then the equation might be of different types in different parts of thexy-plane.
Equation (18.19) obviously represents a very large class of PDEs, and it is
usually impossible to find closed-form solutions to most of these equations.Therefore, for the moment we shall consider only homogeneous equations, with
R(x, y) = 0, and make the further (greatly simplifying) restriction that, throughout
the remainder of this section, A ,B,...,F are not functions of xandybut merely
constants.
620
18.3 GENERAL AND PARTICULAR SOLUTIONS
We now tackle the problem of solving some types of second-order PDE with
constant coefficients by seeking solutions that are arbitrary functions of particularcombinations of independent variables, just as we did for first-order equations.
Following the discussion of the previous section, we can hope to find such
solutions only if all the terms of the equation involve the same total number
of differentiations, i.e. all terms are of the same order, although the numberof differentiations with respect to the individual independent variables may bedifferent. This means that in (18.19) we require the constants D,EandFto be
identically zero (we have, of course, already assumed that R(x, y) is zero), so that
we are now considering only equations of the form
A∂
2u
∂x2+B∂2u
∂x∂y+C∂2u
∂y2=0, (18.20)
where A,BandCare constants. We note that both the one-dimensional wave
equation,
∂2u
∂x2−1
c2∂2u
∂t2=0,
and the two-dimensional Laplace equation,
∂2u
∂x2+∂2u
∂y2=0,
are of this form, but that the diffusion equation,
κ∂2u
∂x2−∂u
∂t=0,
is not, since it contains a first-order derivative.
Since all the terms in (18.20) involve two differentiations, by assuming a solution
of the form u(x, y)=f(p), where pis some unknown function of xandy(ort),
we may be able to obtain a common factor d2f(p)/dp2as the only appearance of
fon the LHS. Then, because of the zero RHS, all reference to the form of fcan
be cancelled out.
We can gain some guidance on suitable forms for the combination p=p(x, y)
by considering ∂u/∂x when uis given by u(x, y)=f(p), for then
∂u
∂x=df(p)
dp∂p
∂x.
Clearly differentiation of this equation with respect to x(ory) will not lead to a
single term on the RHS, containing fonly as d2f(p)/dp2, unless the factor ∂p/∂x
is a constant so that ∂2p/∂x2and∂2p/∂x∂y are necessarily zero. This shows that
pmust be a linear function of x. In an exactly similar way pmust also be a linear
function of y,i . e .p=ax+by.
If we assume a solution of (18.20) of the form u(x, y)=f(ax+by), and evaluate
621
PDES: GENERAL AND PARTICULAR SOLUTIONS
the terms ready for substitution into (18.20), we obtain
∂u
∂x=adf(p)
dp,∂u
∂y=bdf(p)
dp,
∂2u
∂x2=a2d2f(p)
dp2,∂2u
∂x∂y=abd2f(p)
dp2,∂2u
∂y2=b2d2f(p)
dp2,
w h i c ho ns u b s t i t u t i o ng i v e
parenleftbig
Aa2+Bab+Cb2parenrightbigd2f(p)
dp2=0. (18.21)
This is the form we have been seeking, since now a solution independent of
the form of fcan be obtained if we require that aandbsatisfy
Aa2+Bab+Cb2=0.
From this quadratic, two values for the ratio of the two constants aandbare
obtained,
b/a=[−B±(B2−4AC)1/2]/2C.
If we denote these two ratios by λ1andλ2thenanyfunctions of the two variables
p1=x+λ1y, p 2=x+λ2y
will be solutions of the original equation (18.20). The omission of the constant
factor afrom p1andp2is of no consequence since this can always be absorbed
into the particular form of any chosen function; only the relative weighting of x
andyinpis important.
Since p1andp2are in general different, we can thus write the general solution
of (18.20) as
u(x, y)=f(x+λ1y)+g(x+λ2y), (18.22)
where fandgare arbitrary functions.
Finally, we note that the alternative solution d2f(p)/dp2= 0 to (18.21) leads
only to the trivial solution u(x, y)=kx+ly+m, for which all second derivatives
are individually zero.IFind the general solution of the one-dimensional wave equation
∂2u
∂x2−1
c2∂2u
∂t2=0.
This equation is (18.20) with A=1 ,B=0a n d C=−1/c2, and so the values of λ1andλ2
are the solutions of
1−λ2
c2=0,
namely λ1=−candλ2=c. This means that arbitrary functions of the quantities
p1=x−ct, p 2=x+ct
622
18.3 GENERAL AND PARTICULAR SOLUTIONS
will be satisfactory solutions of the equation and that the general solution will be
u(x, t)=f(x−ct)+g(x+ct), (18.23)
where fandgare arbitrary functions. This solution is discussed further in section 18.4.
J
The method used to obtain the general solution of the wave equation may also
be applied straightforwardly to Laplace’s equation.IFind the general solution of the two-dimensional Laplace equation
∂2u
∂x2+∂2u
∂y2=0. (18.24)
Following the established procedure, we look for a solution that is a function f(p)o f
p=x+λy, where from (18.24) λsatisfies
1+λ2=0.
This requires that λ=±i, and satisfactory variables parep=x±iy. The general solution
required is therefore, in terms of arbitrary functions fandg,
u(x, y)=f(x+iy)+g(x−iy).
J
It will be apparent from the last two examples that the nature of the appropriate
linear combination of xandydepends upon whether B2>4ACorB2<4AC.
This is exactly the same criterion as determines whether the PDE is hyperbolic
or elliptic. Hence as a general result, hyperbolic and elliptic equations of the
form (18.20), given the restriction that the constants A,BandCare real, have as
solutions functions whose arguments have the form x+αyandx+iβyrespectively,
where αandβthemselves are real.
The one case not covered by this result is that in which B2=4AC,i . e .a
parabolic equation. In this case λ1andλ2are not different and only one suitable
combination of xandyresults, namely
u(x, y)=f(x−(B/2C)y).
To find the second part of the general solution we try, in analogy with the
corresponding situation for ordinary differential equations, a solution of the form
u(x, y)=h(x, y)g(x−(B/2C)y).
Substituting this into (18.20) and using A=B2/4Cresults in
parenleftbigg
A∂2h
∂x2+B∂2h
∂x∂y+C∂2h
∂y2parenrightbigg
g=0.
Therefore we require h(x, y) to be any solution of the original PDE. There are
several simple solutions of this equation, but as only one is required we take thesimplest non-trivial one, h(x, y)=x, to give the general solution of the parabolic
equation
u(x, y)=f(x−(B/2C)y)+xg(x−(B/2C)y). (18.25)
623
PDES: GENERAL AND PARTICULAR SOLUTIONS
We could, of course, have taken h(x, y)=y, but this only leads to a solution that
is already contained in (18.25).ISolve
∂2u
∂x2+2∂2u
∂x∂y+∂2u
∂y2=0,
subject to the boundary conditions u(0,y)=0 andu(x,1) = x2.
From our general result, functions of p=x+λywill be solutions provided
1+2 λ+λ2=0,
i.e.λ=−1 and the equation is parabolic. The general solution is therefore
u(x, y)=f(x−y)+xg(x−y).
The boundary condition u(0,y) = 0 implies f(p)≡0, whilst u(x,1) = x2yields
xg(x−1) = x2,
which gives g(p)=p+ 1, Therefore the particular solution required is
u(x, y)=x(p+1 )= x(x−y+1 ).
J
To reinforce the material discussed above we will now give alternative deriva-
tions of the general solutions (18.22) and (18.25) by expressing the original PDEin terms of new variables before solving it. The actual solution will then become
almost trivial; but, of course, it will be recognised that suitable new variables
could hardly have been guessed if it were not for the work already done. Thisdoes not detract from the validity of the derivation to be described, only fromthe likelihood that it would be discovered by inspection.
We start again with (18.20) and change to new variables
ζ=x+λ
1y, η =x+λ2y.
With this change of variables, we have from the chain rule that
∂
∂x=∂
∂ζ+∂
∂η,
∂
∂y=λ1∂
∂ζ+λ2∂
∂η.
Using these and the fact that
A+Bλi+Cλ2
i=0 f o r i=1,2,
equation (18.20) becomes
[2A+B(λ1+λ2)+2Cλ1λ2]∂2u
∂ζ∂η=0.
624
18.3 GENERAL AND PARTICULAR SOLUTIONS
Then, providing the factor in brackets does not vanish, for which the required
condition is easily shown to be B2/negationslash=4AC,w eo b t a i n
∂2u
∂ζ∂η=0,
which has the successive integrals
∂u
∂η=F(η),u(ζ,η)=f(η)+g(ζ).
This solution is just the same as (18.22),
u(x, y)=f(x+λ2y)+g(x+λ1y).
If the equation is parabolic (i.e. B2=4AC), we instead use the new variables
ζ=x+λy, η =x,
and recalling that λ=−(B/2C) we can reduce (18.20) to
A∂2u
∂η2=0.
Two straightforward integrations give as the general solution
u(ζ,η)=ηg(ζ)+f(ζ),
w h i c hi nt e r m so f xandyhas exactly the form of (18.25),
u(x, y)=xg(x+λy)+f(x+λy).
Finally, as hinted at in subsection 18.3.2 with reference to first-order linear
PDEs, some of the methods used to find particular integrals of linear ODEscan be suitably modified to find particular integrals of PDEs of higher order. Insimple cases, however, an appropriate solution may often be found by inspection.IFind the general solution of
∂2u
∂x2+∂2u
∂y2=6 (x+y).
Following our previous methods and results, the complementary function is
u(x, y)=f(x+iy)+g(x−iy),
and only a particular integral remains to be found. By inspection a particular integral of
the equation is u(x, y)=x3+y3, and so the general solution can be written
u(x, y)=f(x+iy)+g(x−iy)+x3+y3.
J
625
PDES: GENERAL AND PARTICULAR SOLUTIONS
18.4 The wave equation
We have already found that the general solution of the one-dimensional wave
equation is
u(x, t)=f(x−ct)+g(x+ct), (18.26)
where fandgare arbitrary functions. However, the equation is of such general
importance that further discussion will not be out of place.
Let us imagine that u(x, t)=f(x−ct) represents the displacement of a string at
time tand position x. It is clear that all positions xand times tfor which x−ct=
constant will have the same instantaneous displacement. But x−ct=c o n s t a n t
is exactly the relation between the time and position of an observer travelling
with speed calong the positive x-direction. Consequently this moving observer
sees a constant displacement of the string, whereas to a stationary observer, theinitial profile u(x,0) moves with speed calong the x-axis as if it were a rigid
system. Thus f(x−ct) represents a wave form of constant shape travelling along
the positive x-axis with speed c, the actual form of the wave depending upon
the function f. Similarly, the term g(x+ct) is a constant wave form travelling
with speed cin the negative x-direction. The general solution (18.23) represents
a superposition of these.
If the functions fand gare the same then the complete solution (18.23)
represents identical progressive waves going in opposite directions. This may
result in a wave pattern whose profile does not progress, described as a standing
wave. As a simple example, suppose both f(p)a n d g(p) have the form †
f(p)=g(p)=Acos(kp+/epsilon1).
Then (18.23) can be written as
u(x, t)=A[cos(kx−kct+/epsilon1)+c o s ( kx+kct+/epsilon1)]
=2Acos(kct)cos ( kx+/epsilon1).
The important thing to notice is that the shape of the wave pattern, given by the
factor in x, is the same at all times but that its amplitude 2 Acos(kct) depends
upon time. At some points xthat satisfy
cos(kx+/epsilon1)=0
there is no displacement at any time; such points are called nodes.
So far we have not imposed any boundary conditions on the solution (18.26).
The problem of finding a solution to the wave equation that satisfies given bound-ary conditions is normally treated using the method of separation of variables
†In the usual notation, kis the wave number (= 2 π/wavelength) and kc=ω, the angular frequency
of the wave.
626
18.4 THE WAVE EQUATION
discussed in the next chapter. Nevertheless, we now consider D’Alembert’s solution
u(x, t) of the wave equation subject to initial conditions (boundary conditions) in
the following general form:
initial displacement, u(x,0) = φ(x); initial velocity,∂u(x,0)
∂t=ψ(x).
The functions φ(x)a n d ψ(x) are given and describe the displacement and velocity
of each part of the string at the (arbitrary) time t=0 .
It is clear that what we need are the particular forms of the functions fandg
in (18.26) that lead to the required values at t= 0. This means that
φ(x)=u(x,0) = f(x−0) +g(x+0 ), (18.27)
ψ(x)=∂u(x,0)
∂t=−cf/prime(x−0) +cg/prime(x+0 ), (18.28)
where it should be noted that f/prime(x−0) stands for df(p)/dpevaluated, after the
differentiation, at p=x−c×0; likewise for g/prime(x+0 ) .
Looking on the above two left-hand sides as functions of p=x±ct, but
everywhere evaluated at t= 0, we may integrate (18.28) between an arbitrary
(and irrelevant) lower limit p0and an indefinite upper limit pto obtain
1
cintegraldisplayp
p0ψ(q)dq+K=−f(p)+g(p),
t h ec o n s t a n to fi n t e g r a t i o n Kdepending on p0. Comparing this equation with
(18.27), with xreplaced by p, we can establish the forms of the functions fand
gas
f(p)=φ(p)
2−1
2cintegraldisplayp
p0ψ(q)dq−K
2, (18.29)
g(p)=φ(p)
2+1
2cintegraldisplayp
p0ψ(q)dq+K
2. (18.30)
Adding (18.29) with p=x−ctto (18.30) with p=x+ctgives as the solution to
the original problem
u(x, t)=1
2[φ(x−ct)+φ(x+ct)]+1
2cintegraldisplayx+ct
x−ctψ(q)dq, (18.31)
in which we notice that all dependence on p0has disappeared.
Each of the terms in (18.31) has a fairly straightforward physical interpretation.
In each case the factor 1 /2 represents the fact that only half a displacement profile
that starts at any particular point on the string travels towards any other positionx, the other half travelling away from it. The first term
1
2φ(x−ct) arises from
the initial displacement at a distance ctto the left of x; this travels forward
arriving at xat time t. Similarly, the second contribution is due to the initial
displacement at a distance ctto the right of x. The interpretation of the final
627
PDES: GENERAL AND PARTICULAR SOLUTIONS
term is a little less obvious. It can be viewed as representing the accumulated
transverse displacement at position xdue to the passage past xof all parts of
the initial motion whose effects can reach xwithin a time t, both backward and
forward travelling.
The extension to the three-dimensional wave equation of solutions of the type
we have so far encountered presents no serious difficulty. In Cartesian coordinatesthe three-dimensional wave equation is
∂
2u
∂x2+∂2u
∂y2+∂2u
∂z2−1
c2∂2u
∂t2=0. (18.32)
In close analogy with the one-dimensional case we try solutions that are functions
of linear combinations of all four variables,
p=lx+my+nz+µt.
It is clear that a solution u(x, y, z, t )=f(p) will be acceptable provided that
parenleftbigg
l2+m2+n2−µ2
c2parenrightbiggd2f(p)
dp2=0.
Thus, as in the one-dimensional case, fcan be arbitrary provided that
l2+m2+n2=µ2/c2.
Using an obvious normalisation, we take µ=±candl,m,nas three numbers
such that
l2+m2+n2=1.
In other words ( l,m,n) are the Cartesian components of a unit vector ˆnthat
points along the direction of propagation of the wave. The quantity pcan be
written in terms of vectors as the scalar expression p=ˆn·r±ct, and the general
solution of (18.32) is then
u(x, y, z, t )=u(r,t)=f(ˆn·r−ct)+g(ˆn·r+ct), (18.33)
where ˆnisanyunit vector. It would perhaps be more transparent to write ˆn
explicitly as one of the arguments of u.
18.5 The diffusion equation
One important class of second-order PDEs, which we have not yet considered
in detail, is that in which the second derivative with respect to one variableappears, but only the first derivative with respect to another (usually time). Thisis exemplified by the one-dimensional diffusion equation
κ∂
2u(x, t)
∂x2=∂u
∂t, (18.34)
628
18.5 THE DIFFUSION EQUATION
in which κis a constant with the dimensions length2×time−1. The physical
constants that go to make up κin a particular case depend upon the nature of
the process (e.g. solute diffusion, heat flow, etc.) and the material being described.
With (18.34) we cannot hope to repeat successfully the method of subsection
18.3.3, since now u(x, t) is differentiated a different number of times on the two
sides of the equation; any attempted solution in the form u(x, t)=f(p) with
p=ax+btwill lead only to an equation in which the form of fcannot be
cancelled out. Clearly we must try other methods.
Solutions may be obtained by using the standard method of separation of
variables discussed in the next chapter. Alternatively, a simple solution is also
given if both sides of (18.34), as it stands, are separately set equal to a constant
α(say), so that
∂2u
∂x2=α
κ,∂u
∂t=α.
These equations have the general solutions
u(x, t)=α
2κx2+xg(t)+h(t)a n d u(x, t)=αt+m(x)
respectively, and may be made compatible with each other if g(t) is taken as
constant, g(t)=g(where gcould be zero), h(t)=αtandm(x)=(α/2κ)x2+gx.
An acceptable solution is thus
u(x, t)=α
2κx2+gx+αt+c o n s t a n t . (18.35)
Let us now return to seeking solutions of equations by combining the inde-
pendent variables in particular ways. Having seen that a linear combination ofxandtwill be of no value, we must search for other possible combinations. It
has been noted already that κhas the dimensions length
2×time−1and so the
combination of variables
η=x2
κt
will be dimensionless. Let us see if we can satisfy (18.34) with a solution of the
form u(x, t)=f(η). Evaluating the necessary derivatives we have
∂u
∂x=df(η)
dη∂η
∂x=2x
κtdf(η)
dη,
∂2u
∂x2=2
κtdf(η)
dη+parenleftbigg2x
κtparenrightbigg2d2f(η)
dη2,
∂u
∂t=−x2
κt2df(η)
dη.
Substituting these expressions into (18.34) we find that the new equation can be
629
PDES: GENERAL AND PARTICULAR SOLUTIONS
written entirely in terms of η,
4ηd2f(η)
dη2+( 2+ η)df(η)
dη=0.
This is a straightforward ODE, which can be solved (using a minimum of
explanation) as follows. Writing f/prime(η)=df(η)/dη, etc., we have
f/prime/prime(η)
f/prime(η)=−1
2η−1
4
⇒ ln[η1/2f/prime(η)] =−η
4+c
⇒ f/prime(η)=A
η1/2expparenleftBig−η
4parenrightBig
⇒ f(η)=Aintegraldisplayη
η0µ−1/2expparenleftBig−µ
4parenrightBig
dµ.
If we now write this in terms of a slightly different variable
ζ=η1/2
2=x
2(κt)1/2,
then dζ=1
4η−1/2dη, and the solution to (18.34) is given by
u(x, t)=f(η)=g(ζ)=Bintegraldisplayζ
ζ0exp(−ν2)dν. (18.36)
Here Bis a constant and it should be noticed that xandtappear on the RHS
only in the indefinite upper limit ζ, and then only in the combination xt−1/2.I f
ζ0is chosen as zero then u(x, t) is, to within a constant factor, †the error function
erf[x/2(κt)1/2], which is tabulated in many reference books. Only non-negative
values of xandtare to be considered here, so that ζ≥ζ0.
Let us try to determine what kind of (say) temperature distribution and flow
this represents. For definiteness we take ζ0= 0. Firstly, since u(x, t) in (18.36)
depends only upon the product xt−1/2, it is clear that all points xat times tsuch
that xt−1/2has the same value have the same temperature. Put another way, at
any specific time tthe region having a particular temperature has moved along
the positive x-axis a distance proportional to the square root of t.T h i si sat y p i c a l
diffusion process.
Notice that, on the one hand, at t= 0, the variable ζ→∞ andubecomes
quite independent of x(except perhaps at x= 0); the solution then represents a
uniform spatial temperature distribution. On the other hand, at x=0 , u(x, t)i s
identically zero for all t.
†Take B=2π−1/2to give the usual error function normalised such that erf( ∞) = 1. See the
Appendix.
630
18.5 THE DIFFUSION EQUATIONIAn infrared laser delivers a pulse of (heat) energy Eto a point Pon a large insulated
sheet of thickness b, thermal conductivity k, specific heat sand density ρ. The sheet is
initially at a uniform temperature. If u(r,t)is the excess temperature a time tlater, at a
point that is a distance r(/greatermuchb)from P, then show that a suitable expression for uis
u(r,t)=α
texp
/
−r2
2βt
/
, (18.37)
where αandβare constants. (Note that we use rinstead of ρto denote the radial coordinate
in plane polars so as to avoid confusion with the density.)
Further, (i) show that β=2k/(sρ); (ii) show that the excess heat energy in the sheet
is independent of t, and hence evaluate α; and (iii) show that the total heat flow past any
circle of radius risE.
The equation to be solved is the heat diffusion equation
k∇2u(r,t)=sρ∂u(r,t)
∂t.
Since we only require the solution for r/greatermuchbwe can treat the problem as two-dimensional
with obvious circular symmetry. Thus only the r-derivative term in the expression for ∇2u
is non-zero, giving
k
r∂
∂r
/
r∂u
∂r
/
=sρ∂u
∂t, (18.38)
where now u(r,t)=u(r,t).
(i) Substituting the given expression (18.37) into (18.38) we obtain
2kα
βt2
/r2
2βt−1
/
exp
/
−r2
2βt
/
=sρα
t2
/r2
2βt−1
/
exp
/
−r2
2βt
/
,
from which we find that (18.37) is a solution, provided β=2k/(sρ).
(ii) The excess heat in the system at any time tis
bρs
Z∞
0u(r,t)2πr dr=2πbρsα
Z∞
0r
texp
/
−r2
2βt
/
dr
=2πbρsαβ.
The excess heat is therefore independent of tand must be equal to the total heat
input E, implying that
α=E
2πbρsβ=E
4πbk.
(iii) The total heat flow past a circle of radius ris
−2πrbk
Z∞
0∂u(r,t)
∂rdt=−2πrbk
Z∞
0E
4πbkt
/−r
βt
/
exp
/
−r2
2βt
/
dt
=E
/
exp
/
−r2
2βt
//∞
0=Efor all r.
As we would expect, all the heat energy Edeposited by the laser will eventually flow past
a circle of any given radius r.
J
631
PDES: GENERAL AND PARTICULAR SOLUTIONS
18.6 Characteristics and the existence of solutions
So far in this chapter we have discussed how to find general solutions to various
types of first- and second-order linear PDE. Moreover, given a set of boundary
conditions we have shown how to find the particular solution (or class of solutions)that satisfies them. For first-order equations, for example, we found that if thevalue of u(x, y) is specified along some curve in the xy-plane then the solution to
the PDE is in general unique, but that if u(x, y) is specified at only a single point
then the solution is not unique: there exists a class of particular solutions all ofwhich satisfy the boundary condition. In this section and the next we make more
rigorous the notion of the types of boundary condition that cause a PDE to have
a unique solution, a class of solutions, or no solution at all.
18.6.1 First-order equations
Let us consider the general first-order PDE (18.9) but now write it as
A(x, y)∂u
∂x+B(x, y)∂u
∂y=F(x, y, u). (18.39)
Suppose we wish to solve this PDE subject to the boundary condition that
u(x, y)=φ(s) is specified along some curve Cin the xy-plane that is described
parametrically by the equations x=x(s)a n d y=y(s), where sis the arc length
along C. The variation of ualong Cis therefore given by
du
ds=∂u
∂xdx
ds+∂u
∂ydy
ds=dφ
ds. (18.40)
We may then solve the two (inhomogeneous) simultaneous linear equations
(18.39) and (18.40) for ∂u/∂x and∂u/∂y ,unless the determinant of the coefficients
vanishes (see section 8.18), i.e. unless
vextendsinglevextendsinglevextendsinglevextendsingledx/ds dy/ds
ABvextendsinglevextendsinglevextendsinglevextendsingle=0.
At each point in the xy-plane this equation determines a set of curves called
characteristic curves (or just characteristics ), which thus satisfy
Bdx
ds−Ady
ds=0,
or, multiplying through by ds/dx and dividing through by A,
dy
dx=B(x, y)
A(x, y). (18.41)
However, we have, already met (18.41) in subsection 18.3.1 on first-order PDEs,
where solutions of the form u(x, y)=f(p), where pis some combination of xandy,
632
18.6 CHARACTERISTICS AND THE EXISTENCE OF SOLUTIONS
were discussed. Comparing (18.41) with (18.12) we see that the characteristics are
merely those curves along which pis constant.
Since the partial derivatives ∂u/∂x and∂u/∂y may be evaluated provided the
boundary curve Cdoesnotlie along a characteristic, defining u(x, y)=φ(s)
along Cis sufficient to specify the solution to the original problem (equation
plus boundary conditions) near the curve C, in terms of a Taylor expansion
about C. Therefore the characteristics can be considered as the curves along
which information about the solution u(x, y) ‘propagates’. This is best understood
by using an example.IFind the general solution of
x∂u
∂x−2y∂u
∂y= 0 (18.42)
that takes the value 2y+1on the line x= 1 between y=0a n d y=1.
We solved this problem in subsection 18.3.1 for the case where u(x, y) takes the value
2y+1a l o ngt he entire linex= 1. We found then that the general solution to the equation
(ignoring boundary conditions) is of the form
u(x, y)=f(p)=f(x2y),
for some arbitrary function f. Hence the characteristics of (18.42) are given by x2y=c
where cis a constant; some of these curves are plotted in figure 18.2 for various values of
c. Furthermore, we found that the particular solution for which u(1,y)=2 y+1for all y
was given by
u(x, y)=2 x2y+1.
In the present case the value of x2yis fixed by the boundary conditions only between
y=0a n d y= 1. However, since the characteristics are curves along which x2y, and hence
f(x2y), remains constant, the solution is determined everywhere along any characteristic
that intersects the line segment denoting the boundary conditions. Thus u(x, y)=2 x2y+1
is the particular solution that holds in the shaded region in figure 18.2 (corresponding to0≤c≤1).
Outside this region, however, the solution is not precisely specified, and any function of
the form
u(x, y)=2 x
2y+1+ g(x2y)
will satisfy both the equation and the boundary condition, provided g(p)=0f o r
0≤p≤1.
J
In the above example the boundary curve was not itself a characteristic and
furthermore it crossed each characteristic once only . For a general boundary curve
Cthis may not be the case. Firstly, if Cis itself a characteristic (or is just a single
point) then information about the solution cannot ‘propagate’ away from C,a n d
so the solution remains unspecified everywhere except on C.
The second possibility is that C(although not a characteristic itself) crosses
some characteristics more than once, as in figure 18.3. In this case specifying the
value of u(x, y) along the curve PQdetermines the solution along all the character-
istics that intersect it. Therefore, also specifying u(x, y)a l o n g QRcanoverdetermine
the problem solution and generally results in there being no solution.
633
PDES: GENERAL AND PARTICULAR SOLUTIONS
112y
x−1
x=1c=1
y=c/x2
Figure 18.2 The characteristics of equation (18.42). The shaded region shows
where the solution to the equation is defined, given the imposed boundary
condition at x= 1 between y=0a n d y= 1, shown as a bold vertical line.
Py
QCR
x
Figure 18.3 A boundary curve Cthat crosses characteristics more than once.
18.6.2 Second-order equations
The concept of characteristics can be extended naturally to second- (and higher-)
order equations. In this case let us write the general second-order linear PDE
(18.19) as
A(x, y)∂2u
∂x2+B(x, y)∂2u
∂x∂y+C(x, y)∂2u
∂y2=Fparenleftbigg
x, y, u,∂u
∂x,∂u
∂yparenrightbigg
. (18.43)
634
18.6 CHARACTERISTICS AND THE EXISTENCE OF SOLUTIONS
Cy
xdxdy
ˆndsdr
Figure 18.4 A boundary curve Cand its tangent and unit normal at a given
point.
For second-order equations we might expect that relevant boundary conditions
would involve specifying u, or some of its first derivatives, or both, along a
suitable set of boundaries bordering or enclosing the region over which a solution
is sought. Three common types of boundary condition occur and are associatedwith the names of Dirichlet, Neumann and Cauchy. They are as follows.
(i)Dirichlet : The value of uis specified at each point of the boundary.
(ii)Neumann : The value of ∂u/∂n ,t h enormal derivative ofu,i ss p e c i fi e da t
each point of the boundary. Note that ∂u/∂n =∇u·ˆn,w h e r e ˆnis the
normal to the boundary at each point.
(iii)Cauchy :B o t h uand∂u/∂n are specified at each point of the boundary.
Let us consider for the moment the solution of (18.43) subject to the Cauchy
boundary conditions, i.e. uand∂u/∂n are specified along some boundary curve
Cin the xy-plane defined by the parametric equations x=x(s),y=y(s),sbeing
the arc length along C(see figure 18.4). Let us suppose that along Cwe have
u(x, y)=φ(s)a n d ∂u/∂n =ψ(s). At any point on Cthe vector dr=dxi+dyjis
a tangent to the curve and ˆnds=dyi−dxjis a vector normal to the curve. Thus
onCwe have
∂u
∂s≡∇u·dr
ds=∂u
∂xdx
ds+∂u
∂ydy
ds=dφ(s)
ds,
∂u
∂n≡∇u·ˆn=∂u
∂xdy
ds−∂u
∂ydx
ds=ψ(s).
These two equations may then be solved straightforwardly for the first partial
derivatives ∂u/∂x and∂u/∂y along C. Using the chain rule to write
d
ds=dx
ds∂
∂x+dy
ds∂
∂y,
635
PDES: GENERAL AND PARTICULAR SOLUTIONS
we may differentiate the two first derivatives ∂u/∂x and∂u/∂y along the boundary
to obtain the pair of equations
d
dsparenleftbigg∂u
∂xparenrightbigg
=dx
ds∂2u
∂x2+dy
ds∂2u
∂x∂y,
d
dsparenleftbigg∂u
∂yparenrightbigg
=dx
ds∂2u
∂x∂y+dy
ds∂2u
∂y2.
We may now solve these two equations, together with the original PDE (18.43),
for the second partial derivatives of u,except where the determinant of their
coefficients equals zero,vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleABC
dx
dsdy
ds0
0dx
dsdy
dsvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0.
Expanding out the determinant,
Aparenleftbiggdy
dsparenrightbigg2
−Bparenleftbiggdx
dsparenrightbiggparenleftbiggdy
dsparenrightbigg
+Cparenleftbiggdx
dsparenrightbigg2
=0.
Multiplying through by ( ds/dx)2we obtain
Aparenleftbiggdy
dxparenrightbigg2
−Bdy
dx+C=0, (18.44)
which is the ODE for the curves in the xy-plane along which the second partial
derivatives of ucannot be found.
As for the first-order case, the curves satisfying (18.44) are called characteristics
of the original PDE. These characteristics have tangents at each point given by(when A/negationslash=0 )
dy
dx=B±√
B2−4AC
2A. (18.45)
Clearly, when the original PDE is hyperbolic ( B2>4AC), equation (18.45)
defines two families of real curves in the xy-plane; when the equation is parabolic
(B2=4AC) it defines one family of real curves; and when the equation is elliptic
(B2<4AC) it defines two families of complex curves. Furthermore, when A,
BandCare constants, rather than functions of xandy, the equations of the
characteristics will be of the form x+λy= constant, which is reminiscent of the
form of solution discussed in subsection 18.3.3.
636
18.6 CHARACTERISTICS AND THE EXISTENCE OF SOLUTIONS
0ct
x+ct=c o n s t a n tx−ct=c o n s t a n t
x L
Figure 18.5 The characteristics for the one-dimensional wave equation. The
shaded region indicates the region over which the solution is determined by
specifying Cauchy boundary conditions at t= 0 on the line segment x=0t o
x=L.IFind the characteristics of the one-dimensional wave equation
∂2u
∂x2−1
c2∂2u
∂t2=0.
This is a hyperbolic equation with A=1 , B=0a n d C=−1/c2. Therefore from (18.44)
the characteristics are given by/dx
dt
/2
=c2,
and so the characteristics are the straight lines x−ct=c o n s t a n ta n d x+ct=c o n s t a n t .
J
The characteristics of second-order PDEs can be considered as the curves along
which partial information about the solution u(x, y) ‘propagates’. Consider a point
in the space that has the independent variables as its coordinates; if either or
both of the two characteristics which pass through the point does not intersectthe curve along which the boudary conditions are specified then the solution willnot be determined at that point.In particular, if the equation is hyperbolic, sothat we obtain two families of real characteristics in the xy-plane, then Cauchy
boundary conditions propagate partial information concerning the solution alongthe characteristics, belonging to each family, that intersect the boundary curve C.
The solution uis then specified in the region common to these two families of
characteristics. For instance, the characteristics of the hyperbolic one-dimensionalwave equation in the last example are shown in figure 18.5. By specifying Cauchy
637
PDES: GENERAL AND PARTICULAR SOLUTIONS
Equation type Boundary Conditions
hyperbolic open Cauchy
parabolic open Dirichlet or Neumannelliptic closed Dirichlet or Neumann
Table 18.1 The appropriate boundary conditions for different types of partialdifferential equation.
boundary conditions uand∂u/∂t on the line segment t=0 , x=0t o L,t h e
solution is specified in the shaded region.
As in the case of first-order PDEs, however, problems can arise. For example,
if for a hyperbolic equation the boundary curve intersects any characteristic
more than once then Cauchy conditions along Ccan overdetermine the problem,
resulting in there being no solution. In this case either the boundary curve C
must be altered, or the boundary conditions on the offending parts of Cmust be
relaxed to Dirichlet or Neumann conditions.
The general considerations involved in deciding which boundary conditions are
appropriate for a particular problem are complex, and we do not discuss themany further here. †We merely note that whether the various types of boundary
condition are appropriate (in that they give a solution that is unique, sometimesto within a constant, and is well defined) depends upon the type of second-orderequation under consideration and on whether the region of solution is boundedby a closed or an open curve (or a surface if there are more than two independent
variables). Note that part of a closed boundary may be at infinity if conditions
are imposed on uor∂u/∂n there.
It may be shown that the appropriate boundary-condition and equation-type
pairings are as given in table 18.1.
For example, Laplace’s equation ∇
2u= 0 is elliptic and thus requires either
Dirichlet or Neumann boundary conditions on a closed boundary which, as wehave already noted, may be at infinity if the behaviour of uis specified there
(most often uor∂u/∂n→0 at infinity).
18.7 Uniqueness of solutions
Although we have merely stated the appropriate boundary types and conditions
for which, in the general case, a PDE has a unique, well-defined solution, some-times to within an additive constant, it is often important to be able to prove
that a unique solution is obtained.
†For a discussion the reader is referred, for example, to Morse and Feshbach, Methods of Theoretical
Physics, Part I (McGraw-Hill, 1953) chapter 6.
638
18.7 UNIQUENESS OF SOLUTIONS
As an extremely important example let us consider Poisson’s equation in three
dimensions,
∇2u(r)=ρ(r), (18.46)
with either Dirichlet or Neumann conditions on a closed boundary appropriate
to such an elliptic equation; for brevity, in (18.46), we have absorbed any physicalconstants into ρ. We aim to show that, to within an unimportant constant, the
solution of (18.46) is unique if either the potential uor its normal derivative
∂u/∂n is specified on all surfaces bounding a given region of space (including, if
necessary, a hypothetical spherical surface of indefinitely large radius on which u
or∂u/∂n is prescribed to have an arbitrarily small value). Stated more formally
this is as follows.
Uniqueness theorem. Ifuis real and its first and second partial derivatives are
continuous in a region Vand on its boundary S, and∇
2u=ρinVand either
u=for∂u/∂n =gonS,w h e r e ρ,fandgare prescribed functions, then uis
unique (at least to within an additive constant).IProve the uniqueness theorem for Poisson’s equation.
Let us suppose on the contrary that two solutions u1(r)a n d u2(r) both satisfy the conditions
given above, and denote their difference by the function w=u1−u2. We then have
∇2w=∇2u1−∇2u2=ρ−ρ=0,
so that wsatisfies Laplace’s equation in V. Furthermore, since either u1=f=u2or
∂u1/∂n=g=∂u2/∂nonS, we must have either w=0o r ∂w/∂n =0o n S.
If we now use Green’s first theorem, (11.19), for the case where both scalar functions
are taken as wwe haveZ
V
/
w∇2w+(∇w)·(∇w)
/
dV=
Z
Sw∂w
∂ndS.
However, either condition, w=0o r ∂w/∂n = 0, makes the RHS vanish whilst the first
term on the LHS vanishes since ∇2w=0i n V. Thus we are left withZ
V|∇w|2dV=0.
Since|∇w|2can never be negative, this can only be satisfied if
∇w=0,
i.e. if w, and hence u1−u2, is a constant in V.
If Dirichlet conditions are given then u1≡u2on (some part of) Sand hence u1=u2
everywhere in V. For Neumann conditions, however, u1andu2can differ throughout V
by an arbitrary (but unimportant) constant.
J
The importance of this uniqueness theorem lies in the fact that if a solution to
Poisson’s (or Laplace’s) equation that fits the given set of Dirichlet or Neumann
conditions can be found by any means whatever, then that solution is the correct
one, since only one exists. This result is the mathematical justification for themethod of images , which is discussed more fully in the next chapter.
639
PDES: GENERAL AND PARTICULAR SOLUTIONS
We also note that often the same general method, used in the above example
for proving the uniqueness theorem for Poisson’s equation, can be employed toprove the uniqueness (or otherwise) of solutions to other equations and boundaryconditions.
18.8 Exercises
18.1 Determine whether the following can be written as functions of p=x2+2yonly,
and hence whether they are solutions of (18.8):
(a)x2(x2−4) + 4 y(x2−2) + 4( y2−1);
(b)x4+2x2y+y2;
(c) [ x4+4x2y+4y2+4 ]/[2x4+x2(8y+1 )+8 y2+2y].
18.2 Find partial differential equations satisfied by the following functions u(x, y)f o r
all arbitrary functions fand all arbitrary constants aandb:
(a)u(x, y)=f(x2−y2);
(b)u(x, y)=(x−a)2+(y−b)2;
(c)u(x, y)=ynf(y/x);
(d)u(x, y)=f(x+ay).
18.3 Solve the following partial differential equations for u(x, y) with the boundary
conditions given:
(a)x∂u
∂x+xy=u,u=2yon the line x=1 ;
(b) 1 + x∂u
∂y=xu,u(x,0) = x.
18.4 Find the most general solutions u(x, y) of the following equations consistent with
the boundary conditions stated:
(a)y∂u
∂x−x∂u
∂y=0 , u(x,0) = 1 + sin x;
(b)i∂u
∂x=3∂u
∂y,u=( 4+3 i)x2on the line x=y;
(c) sin xsiny∂u
∂x+c o s xcosy∂u
∂y=0 , u=c o s2 yonx+y=π/2;
(d)∂u
∂x+2x∂u
∂y=0 , u= 2 on the parabola y=x2.
18.5 Find solutions of
1
x∂u
∂x+1
y∂u
∂y=0
for which (a) u(0,y)=y,( b ) u(1,1) = 1.
18.6 Find the most general solutions u(x, y) of the following equations consistent with
the boundary conditions stated:
(a)y∂u
∂x−x∂u
∂y=3x,u=x2on the line y=0 ;
640
18.8 EXERCISES
(b)y∂u
∂x−x∂u
∂y=3x,u(1,0) = 2;
(c)y2∂u
∂x+x2∂u
∂y=x2y2(x3+y3), no boundary conditions.
18.7 Solve
sinx∂u
∂x+c o s x∂u
∂y=c o s x
subject to (a) u(π/2,y)=0 ,( b ) u(π/2,y)=y(y+1 ).
18.8 A function u(x, y)s a t i s fi e s
2∂u
∂x+3∂u
∂y=1 0,
and takes the value 3 on the line y=4x. Evaluate u(2,4).
18.9 If u(x, y)s a t i s fi e s
∂2u
∂x2−3∂2u
∂x∂y+2∂2u
∂y2=0
andu=−x2and∂u/∂y =0f o r y=0a n da l l x, find the value of u(0,1).
18.10 (a) Solve the previous question if the boundary condition is u=∂u/∂y =1
when y=0f o ra l l x.
(b) In which region of the xy-plane would ube determined if the boundary
condition were u=∂u/∂y = 1 when y=0f o ra l l x>0?
18.11 In those cases in which it is possible to do so, evaluate u(2,2), where u(x, y)i s
the solution of
2y∂u
∂x−x∂u
∂y=2xy(2y2−x2)
that satisfies the (separate) boundary conditions given below.
(a)u(x,1) = x2for all x.
(b)u(x,1) = x2forx≥0.
(c)u(x,1) = x2for 0≤x≤3.
(d)u(x,0) = xforx≥0.
(e)u(x,0) = xfor all x.
(f)u(1,√10) = 5 .
(g)u(√10,1) = 5 .
18.12 Solve
6∂2u
∂x2−5∂2u
∂x∂y+∂2u
∂y2=1 4,
subject to u=2x+1a n d ∂u/∂y =4−6x, both on the line y=0 .
18.13 By changing the independent variables in the previous question to
ξ=x+2y and η=x+3y,
show that it must be possible to write 14( x2+5xy+6y2)i nt h ef o r m
f1(x+2y)+f2(x+3y)−(x2+y2),
and determine the forms of f1(z)a n d f2(z).
18.14 Solve
∂2u
∂x∂y+3∂2u
∂y2=x(2y+3x).
18.15 Find the most general solution of ∂2u/∂x2+∂2u/∂y2=x2y2.
641
PDES: GENERAL AND PARTICULAR SOLUTIONS
18.16 An infinitely long string on which waves travel at speed chas an initial displace-
ment
y(x)=
/(
sin(πx/a),−a≤x≤a,
0,|x|>a .
It is released from rest at time t= 0, and its subsequent displacement is described
byy(x, t).
By expressing the initial displacement as one explicit function incorporating
Heaviside step functions, find an expression for y(x, t)a tag e n e r a lt i m e t>0. In
particular, determine the displacement as a function of time (a) at x=0 ,( b )a t
x=a,a n d( c )a t x=a/2.
18.17 The non-relativistic Schr ¨odinger equation (18.7) is similar to the diffusion equa-
tion in having different orders of derivatives in its various terms; this precludessolutions that are arbitrary functions of particular linear combinations of vari-ables. However, since exponential functions do not change their forms underdifferentiation, solutions in the form of exponential functions of combinations ofthe variables may still be possible.
Consider the Schr ¨odinger equation for the case of a constant potential, i.e. for
a free particle, and show that it has solutions of the form Aexp(lx+my+nz+λt)
where the only requirement is that
−/~2
2m
/;
l2+m2+n2
/
=i /~λ.
In particular, identify the equation and wavefunction obtained by taking λas
−iE/ /~,a n d l,mand nasipx/ /~,i py/ /~and ipz/ /~respectively, where Eis the
energy and pthe momentum of the particle; these identifications are essentially
the content of the de Broglie and Einstein relationships.
18.18 Like the Schr ¨odinger equation of the previous question, the equation describing
the transverse vibrations of a rod,
a4∂4u
∂x4+∂2u
∂t2=0,
has different orders of derivatives in its various terms. Show, however, that it has
solutions of exponential form u(x, t)=Aexp(λx+iωt) provided that the relation
a4λ4=ω2is satisfied.
Use a linear combination of such allowed solutions, expressed as the sum of
sinusoids and hyperbolic sinusoids of λx, to describe the transverse vibrations of
a rod of length Lclamped at both ends. At a clamped point both uand∂u/∂x
must vanish; show that this implies that cos( λL)cosh( λL) = 1, thus determining
the frequencies ωat which the rod can vibrate.
18.19 An incompressible fluid of density ρand negligible viscosity flows with velocity
valong a thin straight tube, perfectly light and flexible, of cross-section Aand
held under tension T. Assume that small transverse displacements uof the tube
are governed by
∂2u
∂t2+2v∂2u
∂x∂t+
/
v2−T
ρA
/∂2u
∂x2=0.
(a) Show that the general solution consists of a superposition of two waveforms
travelling with different speeds.
(b) The tube initially has a small transverse displacement u=acoskxand is
suddenly released from rest. Find its subsequent motion.
18.20 A sheet of material of thickness w, specific heat capacity cand thermal con-
ductivity kis isolated in a vacuum, but its two sides are exposed to fluxes of
642
18.8 EXERCISES
r a d i a n th e a to fs t r e n g t h s J1andJ2. Ignoring short-term transients, show that the
temperature difference between its two surfaces is steady at ( J2−J1)w/2k, whilst
their average temperature increases at a rate ( J2+J1)/cw.
18.21 In an electrical cable of resistance Rand capacitance Cper unit length, voltage
signals obey the equation ∂2V/∂x2=RC∂V/∂t . This has solutions of the form
given in (18.36) and also of the form V=Ax+D.
(a) Find a combination of these that re presents the situation after a steady
voltage V0is applied at x=0a tt i m e t=0 .
(b) Obtain a solution describing the propagation of the voltage signal resulting
from application of the signal V=V0for 0 <t<T ,V= 0 otherwise, to
the end x= 0 of an infinite cable.
(c) Show that for t/greatermuchTthe maximum signal occurs at a value of xproportional
tot1/2and has a magnitude proportional to t−1.
18.22 The daily and annual variations of temperature at the surface of the earth may
be represented by sine-wave oscillations with equal amplitudes and periods of1 day and 365 days respectively. Assume that for (angular) frequency ωthe
temperature at depth xin the earth is given by u(x, t)=Asin(ωt+µx)exp(−λx),
where λandµare constants.
(a) Use the diffusion equation to find the values of λandµ.
(b) Find the ratio of the depths below the surface at which the amplitudes have
dropped to 1 /20 of their surface values.
(c) At what time of year is the soil coldest at the greater of these depths,
assuming that the smoothed annual variation in temperature at the surfacehas a minimum on February 1st?
18.23 Consider each of the following situations in a qualitative way and determine
the equation type, the nature of the boundary curve and the type of boundaryconditions involved.
(a) a conducting bar given an initial temperature distribution and then thermally
isolated;
(b) two long conducting concentric cylinders on each of which the voltage
distribution is specified;
(c) two long conducting concentric cylinders on each of which the charge dis-
tribution is specified;
(d) a semi-infinite string the end of which is made to move in a prescribed way.
18.24 This example gives a formal demonstration that the type of a second-order PDE
(elliptic, parabolic or hyperbolic) cannot be changed by a new choi ce of independent
variable. The algebra is somewhat lengthy, but straightforward.
If a change of variable ξ=ξ(x, y),η=η(x, y) is made in (18.19), so that it
reads
A
/prime∂2u
∂ξ2+B/prime∂2u
∂ξ∂η+C/prime∂2u
∂η2+D/prime∂u
∂ξ+E/prime∂u
∂η+F/primeu=R/prime(ξ,η),
show that
B/prime2−4A/primeC/prime=(B2−4AC)
/∂(ξ,η)
∂(x, y)
/2
.
Hence deduce the conclusion stated above.
18.25 The Klein–Gordon equation (which is satisfied by the quantum-mechanical wave-
function Φ( r) of a relativistic spinless particle of non-zero mass m)i s
∇2Φ−m2Φ=0 .
643
PDES: GENERAL AND PARTICULAR SOLUTIONS
Show that the solution for the scalar field Φ( r)i na n yv o l u m e Vbounded by
as u r f a c e Sis unique if either Dirichlet or Neumann boundary conditions are
specified on S.
18.9 Hints and answers
18.1 (a) Yes, p2−4p−4; (b) no, ( p−y)2;( c )y e s ,( p2+4 )/(2p2+p).
18.2 (a) y(∂u/∂x )+x(∂u/∂y )=0 ;( b )( ∂u/∂x )2+(∂u/∂y )2=4u;( c ) x(∂u/∂x )+
y(∂u/∂y )= nu;( d )( ∂u/∂y )(∂2u/∂x2)=( ∂u/∂x )(∂2u/∂x∂y ), or with xand y
reversed.
18.3 Each equation is effectively an ordinary differential equation but with a function
of the non-integrated variable as the constant of integration;(a)u=xy(2−lnx); (b) u=x
−1(1−ey)+xey.
18.4 (a) p=x2+y2,u=s i n ( x2+y2)1/2+1 ;( b ) p=3x+iy,u=( 3x+iy)1/2/2;
(c)p=s i n xcosy,u=2s i n xcosy−1; (d) p=y−x2,u=y−x2+2 .
18.5 (a) ( y2−x2)1/2;( b )1+ f(y2−x2)w h e r e f(0) = 0.
18.6 (a) p=x2+y2, particular integral u=−3y,u=x2+y2−3y;
(b)u=x2+y2−3y+1+ g(x2+y2)w h e r e g(1) = 0;
(c) (x6+y6)/6+g(x3−y3).
18.7 u=y+f(y−ln(sin x)); (a) u=l n ( s i n x); (b) u=y+[y−ln(sin x)]2.
18.8 u=f(3x−2y)+2 ( x+y);f(p)=3+2 p;u=8x−2y+3a n d u(2,4) = 11.
18.9 General solution is u(x, y)=f(x+y)+g(x+y/2). Show that 2 p=−g/prime(p)/2,
and hence g(p)=k−2p2, whilst f(p)=p2−k, leading to u(x, y)=−x2+y2/2;
u(0,1) = 1 /2.
18.10 (a) u(x, y)=2 ( x+y)−2(x+y/2) + 1 = y+1 ; u(0,1) = 2; (b) in the sector
−π/4≤θ≤π/2+φ,w h e r et a n φ=1/2a n d θis measured from the positive
x-axis.
18.11 p=x2+2y2;u(x, y)=f(p)+x2y2/2.
(a)u(x, y)=( x2+2y2+x2y2−2)/2.u(2,2) = 13. The line y= 1 cuts each
characteristic in zero or two distinct points, but this causes no difficulty withthe given boundary conditions.
(b) As in (a).
(c) The solution is defined over the space between the ellipses p=2a n d p= 11;
(2,2) lies on p= 12, and so u(2,2) is undetermined.
(d)u(x, y)=(x
2+2y2)1/2+x2y2/2;u(2,2) = 8 +√12.
(e) The line y= 0, cuts each characteristic in two distinct points. No differ-
entiable form of f(p)g i v e s f(±a)=±arespectively, and so there is no
solution.
(f) The solution is only specified on p= 21, and so u(2,2) is undetermined.
(g) The solution is specified on p= 12, and so u(2,2) = 5 +1
2(4)(4) = 13 .
18.12 u(x, y)=f(x+2y)+g(x+3y)+x2+y2, leading to u=1+2 x+4y−6xy−8y2.
18.13 The equation becomes ∂2f/∂ξ∂η =−14, with solution f(ξ,η)=f(ξ)+g(η)−14ξη,
which can be compared with the answer from the previous question; f1(z)=1 0 z2
andf2(z)=5 z2.
18.14 u=f(y−3x)+g(x)+x2y2/2.
18.15 u(x, y)=f(x+iy)+g(x−iy)+( 1 /12)x4(y2−(1/15)x2). In the last term, xandy
may be interchanged.
18.16
y(x, t)=1
2sin[π(x−ct)/a][H(x−ct+a)−H(x−ct−a)]
+1
2sin[π(x+ct)/a][H(x+ct+a)−H(x+ct−a)].
644
18.9 HINTS AND ANSWERS
(a) zero at all times; (b)1
2sin(πct/a)f o r0≤t≤2a/c, and 0 otherwise;
(c) cos( πct/a)f o r0≤t≤a/2c,1
2cos(πct/a)f o r a/2c≤t≤3a/2c,a n d0
otherwise.
18.17 E=p2/(2m), the relationship between energy and momentum for a non-
relativistic particle; u(r,t)=Aexp[i(p·r−Et)/ /~], a plane wave of wave number
k=p/ /~and angular frequency ω=E/ /~travelling in the direction p/p.
18.18 λ=±ω1/2/aor±iω1/2/a;u(x, t)=e x p ( iωt)[Asinλx+Bcosλx+Csinhλx+
Dcoshλx], with C=−AandD=−B. The conditions at x=Land consistency
establish the quoted result.
18.19 (a) c=v±αwhere α2=T/ρA ;
(b)u(x, t)=acos[k(x−vt)] cos( kαt)−(va/α)sin[k(x−vt)] sin( kαt).
18.20 Use the first form of solution given in (18.35).
18.21 (a) V0
/
1−(2/√π)
R1
2x(CR/t)1/2exp(−ν2)dν
/
; (b) consider as V0applied at t=0
and continued and −V0att=Tand continued;
V(x, t)=2V0√π
Z1
2x[CR/(t−T)]1/2
1
2x(CR/t)1/2exp
/;
−ν2
/
dν;
(c) For t/greatermuchT, maximum at x=[ 2t/(CR)]1/2with value
V0Texp(−1/2)
(2π)1/2t.
18.22 (a) λ=−µ=[ω/(2κ)]1/2,w h e r e κis the diffusion constant; (b) xa= (365)1/2xd;
(c) only the annual variation is significant at this depth and has a phase µaxa=
ln 20 behind the surface. Thus the coldest day is 1 February + (365ln 20) /(2π)
days≈23 July.
18.23 (a) Parabolic, open, Dirichlet u(x,0) given, Neumann ∂u/∂x =0a t x=±L/2
for all t;
(b) elliptic, closed, Dirichlet;(c) elliptic, closed, Neumann ∂u/∂n =σ//epsilon1
0;
(d) hyperbolic, open, Cauchy.
18.24
A/prime=A
/∂ξ
∂x
/2
+B∂ξ
∂x∂ξ
∂y+C
/∂ξ
∂y
/2
,
B/prime=2A∂ξ
∂x∂η
∂x+B
/∂ξ
∂x∂η
∂y+∂ξ
∂y∂η
∂x
/
+2C∂ξ
∂y∂η
∂y,etc.
18.25 Follow an argument similar to that in section 18.7 and argue that the additional
term
R
m2|w|2dVmust be zero, and hence that w= 0 everywhere.
645
19
Partial differential equations:
separation of variables and other
methods
In the previous chapter we demonstrated the methods by which general solutions
of some partial differential equations (PDEs) may be obtained in terms ofarbitrary functions. In particular, solutions containing the independent variablesin definite combinations were sought, thus reducing the effective number of them.
In the present chapter we begin by taking the opposite approach, namely that
of trying to keep the independent variables as separate as possible, using the
method of separation of variables. We then consider integral transform methods
by which one of the independent variables may be eliminated, at least fromdifferential coefficients. Finally, we discuss the use of Green’s functions in solvinginhomogeneous problems.
19.1 Separation of variables: the general method
Suppose we seek a solution u(x, y, z, t ) to some PDE (expressed in Cartesian
coordinates). Let us attempt to obtain one that has the product form †
u(x, y, z, t )=X(x)Y(y)Z(z)T(t). (19.1)
A solution that has this form is said to be separable inx,y,zandt, and seeking
solutions of this form is called the method of separation of variables .
As simple examples we may observe that, of the functions
(i)xyz
2sinbt, (ii)xy+zt, (iii) ( x2+y2)zcosωt,
(i) is completely separable, (ii) is inseparable in that no single variable can be
separated out from it and written as a multiplicative factor, whilst (iii) is separableinzandtbut not in xandy.
†It should be noted that the conventional use here of upper-case (capital) letters to denote the
functions of the corresponding lower-case variable is intended to enable an easy correspondence
between a function and its argument to be made.
646
19.1 SEPARATION OF VARIABLES: THE GENERAL METHOD
When seeking PDE solutions of the form (19.1), we are requiring not that there
is no connection at all between the functions X,Y,ZandT(for example, certain
parameters may appear in two or more of them), but only that the Xdoes not
depend upon y,z,t,t h a t Ydoes not depend on x,z,ta n ds oo n .
For a general PDE it is likely that a separable solution is impossible, but
certainly some common and important equations do have useful solutions ofthis form and we will illustrate the method of solution by studying the three-dimensional wave equation
∇
2u(r)=1
c2∂2u(r)
∂t2. (19.2)
We will work in Cartesian coordinates for the present and assume a solution
of the form (19.1); the solutions in alternative coordinate systems, e.g. sphericalor cylindrical polars, are considered in section 19.3. Expressed in Cartesiancoordinates (19.2) takes the form
∂
2u
∂x2+∂2u
∂y2+∂2u
∂z2=1
c2∂2u
∂t2; (19.3)
substituting (19.1) gives
d2X
dx2YZT +Xd2Y
dy2ZT+XYd2Z
dz2T=1
c2XY Zd2T
dt2,
which can also be written as
X/prime/primeYZT +XY/prime/primeZT+XY Z/prime/primeT=1
c2XY ZT/prime/prime, (19.4)
w h e r ei ne a c hc a s et h ep r i m e sr e f e rt ot h e ordinary derivative with respect to the
independent variable upon which the function depends. This emphasises the factthat each of the functions X,Y,ZandThas only one independent variable and
thus its only derivative is its total derivative. For the same reason, in each termin (19.4) three of the four functions are unaltered by the partial differentiation
and behave exactly as constant multipliers.
If we now divide (19.4) throughout by u=XY ZT we obtain
X
/prime/prime
X+Y/prime/prime
Y+Z/prime/prime
Z=1
c2T/prime/prime
T. (19.5)
This form shows the particular characteristic that is the basis of the method of
separation of variables, namely that of the four terms the first is a function of x
only, the second of yonly, the third of zonly and the RHS a function of tonly
and yet there is an equation connecting them. This can only be so for all x,y,z
andtifeachof the terms does not in fact, despite appearances, depend upon the
corresponding independent variable but is equal to a constant , the four constants
being such that (19.5) is satisfied.
Since there is only one equation to be satisfied and four constants involved,
647
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
there is considerable freedom in the values they may take. For the purposes of
our illustrative example let us make the choice of −l2,−m2,−n2,f o rt h efi r s t
three constants. The constant associated with c−2T/prime/prime/Tmust then have the value
−µ2=−(l2+m2+n2).
Having recognised that each term of (19.5) is individually equal to a constant
(or parameter), we can now replace (19.5) by four separate ordinary differentialequations (ODEs),
X
/prime/prime
X=−l2,Y/prime/prime
Y=−m2,Z/prime/prime
Z=−n2,1
c2T/prime/prime
T=−µ2. (19.6)
The important point to notice is not the simplicity of the equations (19.6) (the
corresponding ones for a general PDE are usually far from simple) but that, bythe device of assuming a separable solution, a partial differential equation (19.3),
containing derivatives with respect to the four independent variables all in oneequation, has been reduced to four separate ordinary differential equations (19.6).
The ordinary equations are connected through four constant parameters that
satisfy an algebraic relation. These constants are called separation constants .
The general solutions of the equations (19.6) can be deduced straightforwardly
and are
X(x)=Aexp(ilx)+Bexp(−ilx)
Y(y)=Cexp(imy)+Dexp(−imy)
Z(z)=Eexp(inz)+Fexp(−inz)
T(t)=Gexp(icµt)+Hexp(−icµt),(19.7)
where A , B,...,H are constants, which may be determined if boundary condtions
are imposed on the solution. Depending on the geometry of the problem andany boundary conditions, it is sometimes more appropriate to write the solutions
(19.7) in the alternative form
X(x)=A
/primecoslx+B/primesinlx
Y(y)=C/primecosmy+D/primesinmy
Z(z)=E/primecosnz+F/primesinnz
T(t)=G/primecos(cµt)+H/primesin(cµt),(19.8)
for some different set of constants A/prime,B/prime,...,H/prime. Clearly the choice of how best
to represent the solution depends on the problem being considered.
As an example, suppose that we take as particular solutions the four functions
X(x)=e x p ( ilx),Y (y)=e x p ( imy),
Z(z)=e x p ( inz),T (t)=e x p (−icµt).
648
19.1 SEPARATION OF VARIABLES: THE GENERAL METHOD
This gives a particular solution of the original PDE (19.3)
u(x, y, z, t )=e x p ( ilx)ex p ( imy)ex p( inz)ex p (−icµt)
=e x p [ i(lx+my+nz−cµt)],
which is a special case of the solution (18.33) obtained in the previous chapter
and represents a plane wave of unit amplitude propagating in a direction givenby the vector with components l,m,n in a Cartesian coordinate system. In the
conventional notation of wave theory, l,mandnare the components of the
wave-number vector k, whose magnitude is given by k=2π/λ,w h e r e λis the
wavelength of the wave; cµis the angular frequency ωof the wave. This gives
the equation in the form
u(x, y, z, t )=e x p [ i(k
xx+kyy+kzz−ωt)]
=e x p [ i(k·r−ωt)],
and makes the exponent dimensionless.
The method of separation of variables can be applied to many commonly
occurring PDEs encountered in physical applications.IUse the method of separation of variables to obtain for the one-dimensional diffusion
equation
κ∂2u
∂x2=∂u
∂t, (19.9)
a solution that tends to zero as t→∞for all x.
Here we have only two independent variables xandtand we therefore assume a solution
of the form
u(x, t)=X(x)T(t).
Substituting this expression into (19.9) and dividing through by u=XT(and also by κ)
we obtain
X/prime/prime
X=T/prime
κT.
Now, arguing exactly as above that the LHS is a function of xonly and the RHS is a
function of tonly, we conclude that each side must equal a constant, which, anticipating
the result and noting the imposed boundary condition, we will take as −λ2. This gives us
two ordinary equations,
X/prime/prime+λ2X=0, (19.10)
T/prime+λ2κT=0, (19.11)
which have the solutions
X(x)=Acosλx+Bsinλx,
T(t)=Cexp(−λ2κt).
Combining these to give the assumed solution u=XTyields (absorbing the constant C
intoAandB)
u(x, t)=(Acosλx+Bsinλx)exp(−λ2κt). (19.12)
649
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
In order to satisfy the boundary condition u→0a s t→∞,λ2κmust be >0. Since κ
is real and >0, this implies that λis a real non-zero number and that the solution is
sinusoidal in xand is not a disguised hyperbolic function; this was our reason for choosing
the separation constant as −λ2.
J
As a final example we consider Laplace’s equation in Cartesian coordinates;
this may be treated in a similar manner.IUse the method of separation of variables to obtain a solution for the two-dimensional
Laplace equation,
∂2u
∂x2+∂2u
∂y2=0. (19.13)
If we assume a solution of the form u(x, y)=X(x)Y(y) then, following the above method,
and taking the separation constant as λ2, we find
X/prime/prime=λ2X, Y/prime/prime=−λ2Y.
Taking λ2as>0, the general solution becomes
u(x, y)=(Acoshλx+Bsinhλx)(Ccosλy+Dsinλy), (19.14)
An alternative form, in which the exponentials are written explicitly, may be useful for
other geometries or boundary conditions:
u(x, y)=[Aexpλx+Bexp(−λx)](Ccosλy+Dsinλy), (19.15)
with different constants AandB.
Ifλ2<0 then the roles of xandyinterchange. The particular combination of sinusoidal
and hyperbolic functions and the values of λallowed will be determined by the geometrical
properties of any specific problem, together with any prescribed or necessary boundaryconditions.J
We note here that a particular case of the solution (19.14) links up with the
‘combination’ result u(x, y)=f(x+iy) of the previous chapter (equations (18.24)
and following), namely that if A=B,a n d D=iCthen the solution is the same
asf(p)=ACexpλpwith p=x+iy.
19.2 Superposition of separated solutions
It will be noticed in the previous two examples that there is considerable freedom
in the values of the separation constant λ, the only essential requirement being
that λhas the samevalue in both parts of the solution, i.e. the part depending
onxand the part depending on y(ort). This is a general feature for solutions
in separated form, which, if the original PDE has nindependent variables, will
contain n−1 separation constants. All that is required in general is that we
associate the correct function of one independent variable with the appropriatefunctions of the others, the correct function being the one with the same values
of the separation constants.
If the original PDE is linear (as are the Laplace, Schr ¨odinger, diffusion and
wave equations) then mathematically acceptable solutions can be formed by
650
19.2 SUPERPOSITION OF SEPARATED SOLUTIONS
superposing solutions corresponding to different allowed values of the separation
constants. To take a two-variable example: if
uλ1(x, y)=Xλ1(x)Yλ1(y)
is a solution of a linear PDE obtained by giving the separation constant the value
λ1then the superposition
u(x, y)=a1Xλ1(x)Yλ1(y)+a2Xλ2(x)Yλ2(y)+···=summationdisplay
iaiXλi(x)Yλi(y),
(19.16)
is also a solution for any constants ai, provided that the λiare the allowed values
of the separation constant λgiven the imposed boundary conditions. Note that
if the boundary conditions allow any of the separation constants to be zero then
the form of the general solution is normally different and must be deduced byreturning to the separated ordinary differential equations. We will encounter thisbehaviour in section 19.3.
The value of the superposition approach is that a boundary condition, say that
u(x, y) takes a particular form f(x)w h e n y= 0, might be met by choosing the
constants a
isuch that
f(x)=summationdisplay
iaiXλi(x)Yλi(0).
In general, this will be possible provided that the functions Xλi(x) form a complete
set – as do the sinusoidal functions of Fourier series or the spherical harmonicsthat we shall discuss in subsection 19.3.2.IA semi-infinite rectangular metal plate occupies the region 0≤x≤∞and0≤y≤bin
thexy-plane. The temperature at the far end of the plate and along its two long sides is
fixed at 0◦C. If the temperature of the plate at x=0is also fixed and is given by f(y),fi n d
the steady-state temperature distribution u(x,y) of the plate. Hence find the temperaturedistribution if f(y)=u
0,w h e r e u0is a constant.
The physical situation is illustrated in figure 19.1. With the notation we have used several
times before, the two-dimensional heat diffu sion equation satisfied by the temperature
u(x, y, t)i s
κ
/∂2u
∂x2+∂2u
∂y2
/
=∂u
∂t,
with κ=k/(sρ). In this case, however, we are asked to find the steady-state temperature,
which corresponds to ∂u/∂t = 0, and so we are led to consider the (two-dimensional)
Laplace equation
∂2u
∂x2+∂2u
∂y2=0.
We saw that assuming a separable solution of the form u(x, y)=X(x)Y(y)l e dt o
solutions such as (19.14) or (19.15), or equivalent forms with xandyinterchanged. In
the current problem we have to satisfy the boundary conditions u(x,0) = 0 = u(x, b)a n d
so a solution that is sinusoidal in yseems appropriate. Furthermore, since we require
u(∞,y) = 0 it is best to write the x-dependence of the solution explicitly in terms of
651
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
y
b
0 xu=0
u=0u→0 u=f(y)
Figure 19.1 A semi-infinite metal plate whose edges are kept at fixed tem-
peratures.
exponentials rather than of hyperbolic functions. We therefore write the separable solution
in the form (19.15) as
u(x, y)=[Aexpλx+Bexp(−λx)](Ccosλy+Dsinλy).
Applying the boundary conditions, we see firstly that u(∞,y) = 0 implies A=0i fw e
take λ>0. Secondly, since u(x,0) = 0 we may set C= 0, which, if we absorb the constant
DintoB, leaves us with
u(x, y)=Bexp(−λx)sinλy.
But, using the condition u(x, b) = 0, we require sin λb= 0 and so the constant λis
constrained to equal nπ/b,w h e r e nis any positive integer.
Using the principle of superposition (19.16), the general solution satisfying the given
boundary conditions can therefore be written
u(x, y)=∞X
n=1Bnexp(−nπx/b )sin(nπy/b ), (19.17)
for some constants Bn. Notice that in the sum in (19.17) we have omitted negative values of
nsince they would lead to exponential terms that diverge as x→∞.T h e n= 0 term is also
omitted since it is identically zero. Using the remaining boundary condition u(0,y)=f(y)
we see that the constants Bnmust satisfy
f(y)=∞X
n=1Bnsin(nπy/b ). (19.18)
This is clearly a Fourier sine series expansion of f(y) (see chapter 12). For (19.18) to
hold, however, the continuation of f(y) outside the region 0 ≤y≤bmust be an odd
periodic function with period 2 b(see figure 19.2). We also see from figure 19.2 that if
the original function f(y) does not equal zero at either of y=0a n d y=bthen its
continuation has a discontinuity at the corresponding point(s); nevertheless, as discussedin chapter 12, the Fourier series will converge to the mid-points of these jumps and hencetend to zero in this case. If, however, the top and bottom edges of the plate were held notat 0
◦C but at some other non-zero temperature, then, in general, the final solution would
possess discontinuities at the corners x=0 , y=0a n d x=0 , y=b.
Bearing in mind these technicalities, the coefficients Bnin (19.18) are given by
Bn=2
b
Zb
0f(y)sin
/nπy
b
/
dy. (19.19)
652
19.2 SUPERPOSITION OF SEPARATED SOLUTIONS
f(y)
−b 0 b y
Figure 19.2 The continuation of f(y) for a Fourier sine series.
Therefore, if f(y)=u0(i.e. the temperature of the side at x= 0 is constant along its
length), (19.19) becomes
Bn=2
b
Zb
0u0sin
/nπy
b
/
dy
=
/
−2u0
bb
nπcos
/nπy
b
/
/b
0
=−2u0
nπ[(−1)n−1] =
/
4u0/nπ fornodd
0f o r neven.
Therefore the required solution is
u(x, y)=
X
nodd4u0
nπexp
/
−nπx
b
/
sin
/nπy
b
/
.
J
In the above example the boundary conditions meant that one term in each
part of the separable solution could be immediately discarded, making the prob-
lem much easier to solve. Sometimes, however, a little ingenuity is required in
writing the separable solution in such a way that certain parts can be neglectedimmediately.ISuppose that the semi-infinite rectangular metal plate in the previous example is replaced
by one that in the x-direction has finite length a. The temperature of the right-hand edge
is fixed at 0◦Cand all other boundary conditions remain as before. Find the steady-state
temperature in the plate.
As in the previous example, the boundary conditions u(x,0) = 0 = u(x, b) suggest a solution
that is sinusoidal in y. In this case, however, we require u=0o n x=a(rather than at
infinity) and so a solution in which the x-dependence is written in terms of hyperbolic
functions, such as (19.14), rather than exponentials is more appropriate. Moreover, sincethe constants in front of the hyperbolic functions are, at this stage, arbitrary, we maywrite the separable solution in the most convenient way that ensures that the conditionu(a, y) = 0 is straightforwardly satisfied. We therefore write
u(x, y)=[Acoshλ(a−x)+Bsinhλ(a−x)](Ccosλy+Dsinλy).
Now the condition u(a, y) = 0 is easily satisfied by setting A=0 .A sb e f o r et h e
conditions u(x,0) = 0 = u(x, b)i m p l y C=0a n d λ=nπ/bfor integer n. Superposing the
653
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
solutions for different nwe then obtain
u(x, y)=∞X
n=1Bnsinh[ nπ(a−x)/b] sin( nπy/b ), (19.20)
for some constants Bn. We have omitted negative values of nin the sum (19.20) since the
relevant terms are already included in those obtained for positive n. Again the n=0t e r m
is identically zero. Using the final boundary condition u(0,y)=f(y) as above we find that
the constants Bnmust satisfy
f(y)=∞X
n=1Bnsinh(nπa/b )sin(nπy/b ),
and, remembering the caveats discussed in the previous example, the Bnare therefore given
by
Bn=2
bsinh(nπa/b )
Zb
0f(y) sin(nπy/b )dy. (19.21)
For the case where f(y)=u0, following the working of the previous example gives
(19.21) as
Bn=4u0
nπsinh(nπa/b )fornodd,B n=0 f o r neven. (19.22)
The required solution is thus
u(x, y)=
X
nodd4u0
nπsinh(nπa/b )sinh[ nπ(a−x)/b]sin
/;
nπy/b
/
.
We note that, as required, in the limit a→∞ this solution tends to the solution of the
previous example.
J
Often the principle of superposition can be used to write the solution to
problems with more complicated boundary conditions as the sum of solutions toproblems that each satisfy only some part of the boundary condition but when
added togther satisfy all the conditions.IFind the steady-state temperature in the (finite) rectangular plate of the previous example,
subject to the boundary conditions u(x, b)=0,u(a, y)=0 andu(0,y)=f(y)as before, but
now in addition u(x,0) = g(x).
Figure 19.3( c) shows the imposed boundary conditions for the metal plate. Although we
could find a solution to this problem using the methods presented above, we can arrive at
the answer almost immediately by using the principle of superposition and the result of
the previous example.
Let us suppose the required solution u(x, y) is made up of two parts:
u(x, y)=v(x, y)+w(x, y),
where v(x, y) is the solution satisfying the boundary conditions shown in figure 19.3( a),
654
19.2 SUPERPOSITION OF SEPARATED SOLUTIONS
yy y
b
bb
f(y)f(y) 0
00
00
0
00
aa a
xx x
g(x)g(x)
(a)( b)
(c)
Figure 19.3 Superposition of boundary conditions for a metal plate.
whilst w(x, y) is the solution satisfying the boundary conditions in figure 19.3( b). It is clear
thatv(x, y) is simply given by the solution to the previous example,
v(x, y)=
X
noddBnsinh
/nπ(a−x)
b
/
sin
/nπy
b
/
,
where Bnis given by (19.21). Moreover, by symmetry, w(x, y)m u s tb eo ft h es a m ef o r ma s
v(x, y) but with xandainterchanged with yandbrespectively, and with f(y) in (19.21)
replaced by g(x). Therefore the required solution can be written down immediately without
further calculation as
u(x, y)=
X
noddBnsinh
/nπ(a−x)
b
/
sin
/nπy
b
/
+
X
noddCnsinh
/nπ(b−y)
a
/
sin
/nπx
a
/
,
theBnbeing given by (19.21) and Cnby
Cn=2
asinh(nπb/a )
Za
0g(x)sin(nπx/a )dx.
Clearly, this method may be extended to cases in which three or four sides of the plate
have non-zero boundary conditions.
J
As a final example of the usefulness of the principle of superposition we now
consider a problem that illustrates how to deal with inhomogeneous boundaryconditions by a suitable change of variables.
655
PDES: SEPARATION OF VARIABLES AND OTHER METHODSIA bar of length Lis initially at a temperature of 0◦C. One end of the bar (x=0 )is held
at0◦Cand the other is supplied with heat at a constant rate per unit area of H.F i n dt h e
temperature distribution within the bar after a time t.
With our usual notation, the heat diffusion equation satisfied by the temperature u(x, t)i s
κ∂2u
∂x2=∂u
∂t,
with κ=k/(sρ), where kis the thermal conductivity of the bar, sis its specific heat
capacity and ρits density.
The boundary conditions can be written as
u(x,0) = 0 ,u (0,t)=0 ,∂u(L, t)
∂x=H
k,
the last of which is inhomogeneous. In general, inhomogeneous boundary conditions can
cause difficulties and it is usual to attempt a transformation of the problem into anequivalent homogeneous one. To this end, let us assume that the solution to our problemtakes the form
u(x, t)=v(x, t)+w(x),
where the function w(x) is to be suitably determined. In terms of vandwthe problem
becomes
κ/∂2v
∂x2+d2w
dx2
/
=∂v
∂t,
v(x,0) +w(x)=0 ,
v(0,t)+w(0) = 0 ,
∂v(L, t)
∂x+dw(L)
dx=H
k.
There are several ways of choosing w(x) so as to make the new problem straightforward.
Using some physical insight, however, it is clear that ultimately (at t=∞), when all
transients have died away, the end x=Lwill attain a temperature u0such that ku0/L=H
and there will be a constant temperature gradient u(x,∞)=u0x/L. We therefore choose
w(x)=Hx
k.
Since the second derivative of w(x) is zero, vsatisfies the diffusion equation and the
boundary conditions on vare now
v(x,0) =−Hx
k,v (0,t)=0 ,∂v(L, t)
∂x=0,
which are homogeneous in x.
From (19.12) a separated solution for the one-dimensional diffusion equation is
v(x, t)=(Acosλx+Bsinλx)exp(−λ2κt),
corresponding to a separation constant −λ2.I fw er e s t r i c t λto be real then all these
solutions are transient ones decaying to zero as t→∞. These are just what is needed
for adding to w(x) to give the correct solution as t→∞. In order to satisfy v(0,t)=0 ,
however, we require A= 0. Furthermore, since
∂v
∂x=Bexp(−λ2κt)λcosλx,
656
19.2 SUPERPOSITION OF SEPARATED SOLUTIONS
L −Lf(x)
x 0
−HL/k
Figure 19.4 The appropriate continuation for a Fourier series containing
only sine terms.
in order to satisfy ∂v(L, t)/∂x=0w er e q u i r ec o s λL=0 ,a n ds o λis restricted to take the
values
λ=nπ
2L,
where nis an odd non-negative integer, i.e. n=1,3,5,....
Thus, to satisfy the boundary condition v(x,0) =−Hx/k, we must haveX
noddBnsin
/nπx
2L
/
=−Hx
k,
in the range x=0t o x=L. In this case we must be more careful about the continuation
of the function −Hx/k for which the Fourier sine series is needed. We want a series that
is odd in x(sine terms only) and continuous as x=0a n d x=L(no discontinuities, since
the series must converge at the end-points). This leads to a continuation of the functionas shown in figure 19.4, with a period of L
/prime=4L. Following the discussion of section 12.3,
since this continuation is odd about x= 0 and even about x=L/prime/4=Lit can indeed be
expressed as a Fourier sine series containing only odd-numbered terms.
The corresponding Fourier series coefficients are found to be
Bn=−8HL
kπ2(−1)(n−1)/2
n2fornodd,
and thus the final formula for u(x, t)i s
u(x, t)=Hx
k−8HL
kπ2
X
nodd(−1)(n−1)/2
n2sin
/nπx
2L
/
exp
/
−kn2π2t
4L2sρ
/
,
giving the temperature for all positions 0 ≤x≤Land for all times t≥0.
J
We note that in all the above examples the boundary conditions restricted the
separation constant(s) to an infinite number of discrete values, usually integers.
If, however, the boundary conditions allow the separation constant(s) λto take
acontinuum of values then the summation in (19.16) is replaced by an integral
over λ. This is discussed further in connection with integral transform methods
in section 19.4.
657
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
19.3 Separation of variables in polar coordinates
So far we have considered the solution of PDEs only in Cartesian coordinates,
but many systems in two and three dimensions are more naturally expressed
in some form of polar coordinates, in which full advantage can be taken ofany inherent symmetries. For example, the potential associated with an isolatedpoint charge has a very simple expression, q/(4π/epsilon1
0r), when polar coordinates are
used, but involves all three coordinates and square roots, when Cartesians areemployed. For these reasons we now turn to the separation of variables in planepolar, cylindrical polar and spherical polar coordinates.
Most of the PDEs we have considered so far have involved the operator ∇
2,e . g .
the wave equation, the diffusion equation, Schr ¨odinger’s equation and Poisson’s
equation (and of course Laplace’s equation). It is therefore appropriate that werecall the expressions for ∇
2when expressed in polar coordinate systems. From
chapter 10, in plane polars, cylindrical polars and spherical polars respectivelywe have
∇
2=1
ρ∂
∂ρparenleftbigg
ρ∂
∂ρparenrightbigg
+1
ρ2∂2
∂φ2, (19.23)
∇2=1
ρ∂
∂ρparenleftbigg
ρ∂
∂ρparenrightbigg
+1
ρ2∂2
∂φ2+∂2
∂z2, (19.24)
∇2=1
r2∂
∂rparenleftbigg
r2∂
∂rparenrightbigg
+1
r2sinθ∂
∂θparenleftbigg
sinθ∂
∂θparenrightbigg
+1
r2sin2θ∂2
∂φ2. (19.25)
Of course the first of these may be obtained from the second by taking zto be
identically zero.
19.3.1 Laplace’s equation in polar coordinates
The simplest of the equations containing ∇2is Laplace’s equation,
∇2u(r)=0 . (19.26)
Since it contains most of the essential features of the other more complicated
equations we will consider its solution first.
Laplace’s equation in plane polars
Suppose that we need to find a solution of (19.26) that has a prescribed behaviour
on the circle ρ=a(e.g. if we are finding the shape taken up by a circular drumskin
when its rim is slightly deformed from being planar). Then we may seek solutions
of (19.26) that are separable in ρandφ(measured from some arbitrary radius
asφ= 0) and hope to accommodate the boundary condition by examining the
solution for ρ=a.
658
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
Thus, writing u(ρ, φ)=P(ρ)Φ(φ) and using the expression (19.23), Laplace’s
equation (19.26) becomes
Φ
ρ∂
∂ρparenleftbigg
ρ∂P
∂ρparenrightbigg
+P
ρ2∂2Φ
∂φ2=0.
Now, employing the same device as previously, that of dividing through by
u=PΦ and multiplying through by ρ2, results in the separated equation
ρ
P∂
∂ρparenleftbigg
ρ∂P
∂ρparenrightbigg
+1
Φ∂2Φ
∂φ2=0.
Following our earlier argument, since the first term on the RHS is a function of
ρonly, whilst the second term depends only on φ, we obtain the two ordinary
equations
ρ
Pd
dρparenleftbigg
ρdP
dρparenrightbigg
=n2(19.27)
1
Φd2Φ
dφ2=−n2, (19.28)
where we have taken the separation constant to have the form n2for later
convenience; for the present nis a general (complex) number.
Let us first consider the case in which n/negationslash= 0. The second equation, (19.28), then
has the general solution
Φ(φ)=Aexp(inφ)+Bexp(−inφ). (19.29)
Equation (19.27), on the other hand, is the homogeneous equation
ρ2P/prime/prime+ρP/prime−n2P=0,
which must be solved either by trying a power solution in ρor by making the
substitution ρ=e x p tas described in subsection 15.2.1 and so reducing it to an
equation with constant coefficients. Carrying out this procedure we find
P(ρ)=Cρn+Dρ−n. (19.30)
Returning to the solution (19.29) of the azimuthal equation (19.28), we can
see that if Φ, and hence u, is to be single-valued and so not change when φ
increases by 2 πthen nmust be an integer. Mathematically, other values of nare
permissible, but for the description of real physical situations it is clear that thislimitation must be imposed. Having thus restricted the possible values of nin one
part of the solution, the same limitations must be carried over into the radial part(19.30). Thus we may write a particular solution of the two-dimensional Laplace
equation as
u(ρ, φ)=(Acosnφ+Bsinnφ)(Cρ
n+Dρ−n),
where A, B, C, D are arbitrary constants and nis any integer.
659
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
We have not yet, however, considered the solution when n=0 .I nt h i sc a s e ,
the solutions of the separated ordinary equations (19.28) and (19.27) respectivelyare easily shown to be
Φ(φ)=Aφ+B,
P(ρ)=Clnρ+D.
But, in order that u=PΦ is single-valued, we require A=0a n ds ot h es o l u t i o n
forn= 0 is simply (absorbing BintoCandD)
u(ρ, φ)=Clnρ+D.
Superposing the solutions for the different allowed values of n, we can write
the general solution to Laplace’s equation in plane polars as
u(ρ, φ)=(C
0lnρ+D0)+∞summationdisplay
n=1(Ancosnφ+Bnsinnφ)(Cnρn+Dnρ−n),
(19.31)
where ncan take only integer values. Negative values of nhave been omitted
from the sum since they are already included in the terms obtained for positive
n. We note that, since ln ρis singular at ρ= 0, whenever we solve Laplace’s
equation in a region containing the origin, C0must be identically zero.IA circular drumskin has a supporting rim at ρ=a. If the rim is twisted so that it
is displaced vertically by a small amount /epsilon1(sinφ+2 s i n2 φ),w h e r e φis the azimuthal
angle with respect to a given radius , find the resulting displacement u(ρ, φ)over the entire
drumskin.
The transverse displacement of a circular drumskin is usually described by the two-dimensional wave equation. In this case, however, there is no time dependence and sou(ρ, φ) solves the two-dimensional Laplace equation, subject to the imposed boundary
condition.
Referring to (19.31), since we wish to find a solution that is finite everywhere inside
ρ=a,w er e q u i r e C
0=0a n d Dn=0f o ra l l n>0. Now the boundary condition at the
rim requires
u(a, φ)=D0+∞X
n=1Cnan(Ancosnφ+Bnsinnφ)=/epsilon1(sinφ+2s i n2 φ).
Firstly we see that we require D0=0a n d An=0f o ra l l n. Furthermore, we must
have C1B1a=/epsilon1,C2B2a2=2/epsilon1andBn=0f o r n>2. Hence the appropriate shape for the
drumskin (valid over the whole skin, not just the rim) is
u(ρ, φ)=/epsilon1ρ
asinφ+2/epsilon1ρ2
a2sin2φ=/epsilon1ρ
a
/
sinφ+2ρ
asin2φ
/
.
J
660
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
Laplace’s equation in cylindrical polars
Passing to three dimensions, we now consider the solution of Laplace’s equation
in cylindrical polar coordinates,
1
ρ∂
∂ρparenleftbigg
ρ∂u
∂ρparenrightbigg
+1
ρ2∂2u
∂φ2+∂2u
∂z2=0. (19.32)
We note here that, even when considering a cylindrical physical system, if there
is no dependence of the physical variables on z(i.e. along the length of the
cylinder) then the problem may be treated using two-dimensional plane polars,
as discussed above.
For the more general case, however, we proceed as previously by trying a
solution of the form
u(ρ, φ, z)=P(ρ)Φ(φ)Z(z),
which on substitution into (19.32) and division through by u=PΦZgives
1
Pρd
dρparenleftbigg
ρdP
dρparenrightbigg
+1
Φρ2d2Φ
dφ2+1
Zd2Z
dz2=0.
The last term depends only on zand the first and second (taken together) only
onρandφ. Taking the separation constant to be k2, we find
1
Zd2Z
dz2=k2,
1
Pρd
dρparenleftbigg
ρdP
dρparenrightbigg
+1
Φρ2d2Φ
dφ2+k2=0.
The first of these equations has the straightforward solution
Z(z)=Eexp(−kz)+Fexpkz.
Multiplying the second equation through by ρ2,w eo b t a i n
ρ
Pd
dρparenleftbigg
ρdP
dρparenrightbigg
+1
Φd2Φ
dφ2+k2ρ2=0,
in which the second term depends only on Φ and the other terms only on ρ.
Taking the second separation constant to be m2, we find
1
Φd2Φ
dφ2=−m2, (19.33)
ρd
dρparenleftbigg
ρdP
dρparenrightbigg
+(k2ρ2−m2)P=0. (19.34)
The equation in the azimuthal angle φhas the very familiar solution
Φ(φ)=Ccosmφ+Dsinmφ.
661
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
As in the two-dimensional case, single-valuedness of urequires that mis an
integer. However, in the particular case m= 0 the solution is
Φ(φ)=Cφ+D.
This form is appropriate to a solution with axial symmetry ( C=0 )o ro n et h a ti s
multivalued, but manageably so, such as the magnetic scalar potential associatedwith a current I(in which case C=I/(2π)a n d Dis arbitrary).
Finally the ρ-equation (19.34) may be transformed into Bessel’s equation of
order mby writing µ=kρ. This has the solution
P(ρ)=AJ
m(kρ)+BYm(kρ).
The properties of these functions were investigated in chapter 16 and will not
be pursued here. We merely note that Ym(kρ) is singular at ρ=0 ,a n ds ow h e n
seeking solutions to Laplace’s equation in cylindrical coordinates within someregion containing the ρ= 0 axis, we require B=0 .
The complete separated-variable solution in cylindrical polars of Laplace’s
equation∇
2u= 0 is thus
u(ρ, φ, z)=[AJm(kρ)+BYm(kρ)][Ccosmφ+Dsinmφ][Eexp(−kz)+Fexpkz].
(19.35)
Of course we may use the principle of superposition to build up more general
solutions by adding together solutions of the form (19.35) for all allowed valuesof the separation constants kandm.IA semi-infinite solid cylinder of radius ahas its curved surface held at 0◦Cand its base
held at a temperature T0. Find the steady-state temperature distribution in the cylinder.
The physical situation is shown in figure 19.5. The steady-state temperature distribution
u(ρ, φ, z) must satisfy Laplace’s equation subject to the imposed boundary conditions. Let
us take the cylinder to have its base in the z= 0 plane and to extend along the positive
z-axis. From (19.35), in order that uis finite everywhere in the cylinder we immediately
require B=0a n d F= 0. Furthermore, since the boundary conditions, and hence the
temperature distribution, are axially symmetric we require m= 0, and so the general
solution must be a superposition of solutions of the form J0(kρ)exp(−kz) for all allowed
values of the separation constant k.
The boundary condition u(a, φ, z) = 0 restricts the allowed values of ksince we must
have J0(ka) = 0. The zeroes of Bessel functions are given in most books of mathematical
tables, and we find that, to two decimal places,
J0(x)=0 f o r x=2.40,5.52,8.65,. . . .
Writing the allowed values of kasknforn=1,2,3,...(so, for example, k1=2.40/a), the
required solution takes the form
u(ρ, φ, z)=∞X
n=1AnJ0(knρ)e x p (−knz).
662
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
z
xy au=0 u=0
u=T0
Figure 19.5 A uniform metal cylinder whose curved surface is kept at 0◦C
and whose base is held at a temperature T0.
By imposing the remaining boundary condition u(ρ, φ,0) = T0, the coefficients Ancan be
found in a similar way to Fourier coefficients but this time by exploiting the orthogonalityof the Bessel functions, as discussed in chapter 16. From this boundary condition werequire
u(ρ, φ,0) =
∞X
n=1AnJ0(knρ)=T0.
If we multiply this expression by ρJ0(krρ) and integrate from ρ=0t o ρ=a, and use the
orthogonality of the Bessel functions J0(knρ), then the coefficients are given by (16.81) as
An=2T0
a2J2
1(kna)
Za
0J0(knρ)ρd ρ . (19.36)
The integral on the RHS can be evaluated using the recurrence relation (16.68) of
chapter 16,
d
dz[zJ1(z)] =zJ0(z),
which on setting z=knρyields
1
knd
dρ[knρJ1(knρ)] =knρJ0(knρ).
Therefore the integral in (19.36) is given byZa
0J0(knρ)ρd ρ=
/1
knρJ1(knρ)
/a
0=1
knaJ1(kna),
663
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
and the coefficients Anmay be expressed as
An=2T0
a2J2
1(kna)
/aJ1(kna)
kn
/
=2T0
knaJ1(kna).
The steady-state temperature in the cylinder is then given by
u(ρ, φ, z)=∞X
n=12T0
knaJ1(kna)J0(knρ)e x p (−knz).
J
We note that if, in the above example, the base of the cylinder were not kept at
a uniform temperature T0, but instead had some fixed temperature distribution
T(ρ, φ), then the solution of the problem would become more complicated. In
such a case, the required temperature distribution u(ρ, φ, z) is in general notaxially
symmetric, and so the separation constant mis not restricted to be zero but may
take any integer value. The solution will then take the form
u(ρ, φ, z)=∞summationdisplay
m=0∞summationdisplay
n=1Jm(knmρ)(Cnmcosmφ+Dnmsinmφ)ex p (−knmz),
where the separation constants knmare such that Jm(knma)=0 ,i . e . knmais the nth
zero of the mth-order Bessel function. At the base of the cylinder we would then
require
u(ρ, φ,0) =∞summationdisplay
m=0∞summationdisplay
n=1Jm(knmρ)(Cnmcosmφ+Dnmsinmφ)=T(ρ, φ).
(19.37)
The coefficients Cnmcould be found by multiplying (19.37) by Jq(krqρ)cosqφ,
integrating with respect to ρandφover the base of the cylinder and exploiting
the orthogonality of the Bessel functions and of the trigonometric functions. The
Dnmcould be found in a similar way by multiplying (19.37) by Jq(krqρ)s i nqφ.
Laplace’s equation in spherical polars
We now come to an eequation that is very widely applicable in physical science,
namely∇2u= 0 in spherical polar coordinates:
1
r2∂
∂rparenleftbigg
r2∂u
∂rparenrightbigg
+1
r2sinθ∂
∂θparenleftbigg
sinθ∂u
∂θparenrightbigg
+1
r2sin2θ∂2u
∂φ2=0. (19.38)
Our method of procedure will be as before; we try a solution of the form
u(r,θ,φ)=R(r)Θ(θ)Φ(φ).
Substituting this in (19.38), dividing through by u=RΘΦ and multiplying by r2,
we obtain
1
Rd
drparenleftbigg
r2dR
drparenrightbigg
+1
Θs i n θd
dθparenleftbigg
sinθdΘ
dθparenrightbigg
+1
Φsin2θd2Φ
dφ2=0. (19.39)
664
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
The first term depends only on rand the second and third terms (taken together)
only on θandφ. Thus (19.39) is equivalent to the two equations
1
Rd
drparenleftbigg
r2dR
drparenrightbigg
=λ, (19.40)
1
Θs i n θd
dθparenleftbigg
sinθdΘ
dθparenrightbigg
+1
Φsin2θd2Φ
dφ2=−λ. (19.41)
Equation (19.40) is a homogeneous equation,
r2d2R
dr2+2rdR
dr−λR=0,
which can be reduced by the substitution r=e x p t(and writing R(r)=S(t)) to
d2S
dt2+dS
dt−λS=0.
This has the straightforward solution
S(t)=Aexpλ1t+Bexpλ2t,
and so the solution to the radial equation is
R(r)=Arλ1+Brλ2,
where λ1+λ2=−1a n d λ1λ2=−λ. We can thus take λ1andλ2as given by /lscript
and−(/lscript+1 ) ; λthen has the form /lscript(/lscript+ 1). (It should be noted that at this stage
nothing has been either assumed or proved about whether /lscriptis an integer.)
Hence we have obtained some information about the first factor in the
separated-variable solution, which will now have the form
u(r,θ,φ)=bracketleftbig
Ar/lscript+Br−(/lscript+1)bracketrightbig
Θ(θ)Φ(φ), (19.42)
where Θ and Φ must satisfy (19.41) with λ=/lscript(/lscript+1 ) .
The next step is to take (19.41) further. Multiplying through by sin2θand
substituting for λ, it too takes a separated form:
bracketleftbiggsinθ
Θd
dθparenleftbigg
sinθdΘ
dθparenrightbigg
+/lscript(/lscript+1 )s i n2θbracketrightbigg
+1
Φd2Φ
dφ2=0. (19.43)
Taking the separation constant as m2, the equation in the azimuthal angle φ
has the same solution as in cylindrical polars, namely
Φ(φ)=Ccosmφ+Dsinmφ.
As before, single-valuedness of urequires that mis an integer; for m= 0 we again
have Φ( φ)=Cφ+D.
665
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
Having settled the form of Φ( φ), we are left only with the equation satisfied by
Θ(θ), which is
sinθ
Θd
dθparenleftbigg
sinθdΘ
dθparenrightbigg
+/lscript(/lscript+1 )s i n2θ=m2. (19.44)
A change of independent variable from θtoµ=c o s θwill reduce this to a
form for which solutions are known, and of which some study has been made inchapter 16. Putting
µ=c o s θ,dµ
dθ=−sinθ,d
dθ=−(1−µ2)1/2d
dµ,
the equation for M(µ)≡Θ(θ)r e a d s
d
dµbracketleftbigg
(1−µ2)dM
dµbracketrightbigg
+bracketleftbigg
/lscript(/lscript+1 )−m2
1−µ2bracketrightbigg
M=0. (19.45)
This equation is the associated Legendre equation , which was mentioned in sub-
section 17.5.2 in the context of Sturm–Liouville equations.
We recall that for the case m= 0, (19.45) reduces to Legendre’s equation, which
was studied at length in chapter 16, and has the solution
M(µ)=EP/lscript(µ)+FQ/lscript(µ). (19.46)
We have not solved (19.45) explicitly for general m, but the solutions were given
in subsection 17.5.2 and are the associated Legendre functions Pm
/lscript(µ)a n d Qm
/lscript(µ),
where
Pm
/lscript(µ)=( 1−µ2)|m|/2d|m|
dµ|m|P/lscript(µ), (19.47)
and similarly for Qm
/lscript(µ). We then have
M(µ)=EPm
/lscript(µ)+FQm
/lscript(µ); (19.48)
here mmust be an integer, 0 ≤|m|≤/lscript. We note that if we require solutions to
Laplace’s equation that are finite when µ=c o s θ=±1 (i.e. on the polar axis
where θ=0,π), then we must have F= 0 in (19.46) and (19.48) since Qm
/lscript(µ)
diverges at µ=±1.
It will be remembered that one of the important conditions for obtaining
finite polynomial solutions of Legendre’s equation is that /lscriptis an integer ≥0.
This condition therefore applies also to the solutions (19.46) and (19.48) and isreflected back into the radial part of the general solution given in (19.42).
Now that the solutions of each of the three ordinary differential equations
governing R, Θ and Φ have been obtained, we may assemble a complete separated-
666
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
variable solution of Laplace’s equation in spherical polars. It is
u(r,θ,φ)=(Ar/lscript+Br−(/lscript+1))(Ccosmφ+Dsinmφ)[EPm
/lscript(cosθ)+FQm
/lscript(cosθ)],
(19.49)
where the three bracketted factors are connected only through the integer pa-
rameters /lscriptandm,0≤|m|≤/lscript. As before, a general solution may be obtained
by superposing solutions of this form for the allowed values of the separationconstants /lscriptandm. As mentioned above, if the solution is required to be finite on
the polar axis then F=0f o ra l l /lscriptandm.IAn uncharged conducting sphere of radius ais placed at the origin in an initially uniform
electrostatic field E. Show that it behaves as an electric dipole.
The uniform field, taken in the direction of the polar axis, has a electrostatic potential
u=−Ez=−Ercosθ,
where uis arbitrarily taken as zero at z=0 .T h i ss a t i s fi e sL a p l a c e ’ se q u a t i o n ∇2u=0 ,a s
must the potential vwhen the sphere is present; for large rthe asymptotic form of vmust
still be−Ercosθ.
Since the problem is clearly axially symmetric we have immediately that m=0 ,a n d
since we require vto be finite on the polar axis we must have F= 0 in (19.49). Therefore
the solution must be of the form
v(r,θ,φ)=∞X
/lscript=0(A/lscriptr/lscript+B/lscriptr−(/lscript+1))P/lscript(cosθ).
Now the cos θ-dependence of vfor large rindicates that the ( θ,φ)-dependence of v(r,θ,φ)
is given by P0
1(cosθ)=c o s θ. Thus the r-dependence of vmust also correspond to an
/lscript= 1 solution, and the most general such solution (outside the sphere, i.e. for r≥a)i s
v(r,θ,φ)=(A1r+B1r−2)P1(cosθ).
The asymptotic form of vfor large rimmediately gives A1=−Eand so yields the solution
v(r,θ,φ)=
/
−Er+B1
r2
/
cosθ.
Since the sphere is conducting, it is an equipotential region and so vmust not depend on
θforr=a. This can only be the case if B1/a2=Ea, thus fixing B1. The final solution is
therefore
v(r,θ,φ)=−Er
/
1−a3
r3
/
cosθ.
Since a dipole of moment pgives rise to a potential p/(4π/epsilon10r2), this result shows that the
sphere behaves as a dipole of moment 4 π/epsilon10a3E, because of the charge distribution induced
on its surface; see figure 19.6.
J
Often the boundary conditions are not so easily met, and it is necessary to
use the mutual orthogonality of the associated Legendre functions (and thetrigonometric functions) to obtain the coefficients in the general solution.
667
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
+
+
+
+
+
+
+
+++
++
−−−−
−
−−
−−−−−
θ
a
Figure 19.6 Induced charge and field lines associated with a conducting
sphere placed in an initially uniform electrostatic field.IA hollow split conducting sphere of radius ais placed at the origin. If one half of its
surface is charged to a potential v0and the other half is kept at zero potential, find the
potential vinside and outside the sphere.
Let us choose the top hemisphere to be charged to v0and the bottom hemisphere to be
at zero potential, with the plane in which the two hemispheres meet perpendicular to thepolar axis; this is shown in figure 19.7. The boundary condition then becomes
v(a, θ, φ)=
/
v0for 0 <θ<π / 2( 0 <cosθ<1),
0f o r π/2<θ<π (−1<cosθ<0).(19.50)
The problem is clearly axially symmetric and so we may set m= 0. Also, we require the
solution to be finite on the polar axis and so it cannot contain Q/lscript(cosθ). Therefore the
general form of the solution to (19.38) is
v(r,θ,φ)=∞X
/lscript=0(A/lscriptr/lscript+B/lscriptr−(/lscript+1))P/lscript(cosθ). (19.51)
Inside the sphere (for r<a) we require the solution to be finite at the origin and so
B/lscript=0f o ra l l /lscriptin (19.51). Imposing the boundary condition at r=awe must then have
v(a, θ, φ)=∞X
/lscript=0A/lscripta/lscriptP/lscript(cosθ),
where v(a, θ, φ) is also given by (19.50). Exploiting the mutual orthogonality of the Legendre
polynomials, the coefficients in the Legendre polynomial expansion are given by (16.48)as (writing µ=c o s θ)
A
/lscripta/lscript=2/lscript+1
2
Z1
−1v(a, θ, φ)P/lscript(µ)dµ
=2/lscript+1
2v0
Z1
0P/lscript(µ)dµ,
668
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
xz
ya
−aφθ
v=0v=v0
r
Figure 19.7 A hollow split conducting sphere with its top half charged to a
potential v0and its bottom half at zero potential.
where in the last line we have used (19.50). The integrals of the Legendre polynomials are
easily evaluated (see exercise 17.7) and we find
A0=v0
2,A 1=3v0
4a,A 2=0,A 3=−7v0
16a3,···,
so that the required solution inside the sphere is
v(r,θ,φ)=v0
2
/
1+3r
2aP1(cosθ)−7r3
8a3P3(cosθ)+···
/
.
Outside the sphere (for r>a ) we require the solution to be bounded as rtends to
infinity and so in (19.51) we must have A/lscript=0f o ra l l /lscript. In this case, by imposing the
boundary condition at r=awe require
v(a, θ, φ)=∞X
/lscript=0B/lscripta−(/lscript+1)P/lscript(cosθ),
where v(a, θ, φ) is given by (19.50). Following the above argument the coefficients in the
expansion are given by
B/lscripta−(/lscript+1)=2/lscript+1
2v0
Z1
0P/lscript(µ)dµ,
so that the required solution outside the sphere is
v(r,θ,φ)=v0a
2r
/
1+3a
2rP1(cosθ)−7a3
8r3P3(cosθ)+···
/
.
J
669
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
In the above example, on the equator of the sphere (i.e. at r=aandθ=π/2)
the potential is given by
v(a, π/2,φ)=v0/2,
i.e. mid-way between the potentials of the top and bottom hemispheres. This is
so because a Legendre polynomial expansion of a function behaves in the sameway as a Fourier series expansion, in that it converges to the average of the twovalues at any discontinuities present in the original function.
If the potential on the surface of the sphere had been given as a function of θ
andφ, then we would have had to consider a double series summed over /lscriptand
m(for−/lscript≤m≤/lscript), since, in general, the solution would not have been axially
symmetric.
19.3.2 Spherical harmonics
When obtaining solutions in spherical polar coordinates of ∇
2u= 0, we found
that, for solutions that are finite on the polar axis, the angular part of the solutionwas given by
Θ(θ)Φ(φ)=P
m
/lscript(cosθ)(Ccosmφ+Dsinmφ).
This general form is sufficiently common that particular functions of θandφ
called spherical harmonics are defined and tabulated. The spherical harmonics
Ym
/lscript(θ,φ) are defined for m≥0b y
Ym
/lscript(θ,φ)=(−1)mbracketleftbigg2/lscript+1
4π(/lscript−m)!
(/lscript+m)!bracketrightbigg1/2
Pm
/lscript(cosθ)ex p ( imφ). (19.52)
For values of m<0t h er e l a t i o n
Y−|m|
/lscript(θ,φ)=(−1)|m|bracketleftBig
Y|m|
/lscript(θ,φ)bracketrightBig∗
defines the spherical harmonic, the asterisk denoting complex conjugation. Since
they contain as their θ-dependent part the solution Pm
/lscriptto the associated Legendre
equation, which is a Sturm–Liouville equation (see chapter 17), the Ym
/lscriptare
mutually orthogonal when integrated from −1t o+ 1o v e r d(cosθ). Their mutual
orthogonality with respect to φ(0≤φ≤2π) is even more obvious. The numerical
factor in (19.52) is chosen to make the Ym
/lscriptan orthonormal set, that is
integraldisplay1
−1integraldisplay2π
0bracketleftbig
Ym
/lscript(θ,φ)bracketrightbig∗Ym/prime
/lscript/prime(θ,φ)dφ d(cosθ)=δ/lscript/lscript/primeδmm/prime.
In addition, the spherical harmonics form a complete set in that any reasonable
function (i.e. one that is likely to be met in a physical situation) of θandφcan
670
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
be expanded as a sum of such functions,
f(θ,φ)=∞summationdisplay
/lscript=0/lscriptsummationdisplay
m=−/lscripta/lscriptmYm
/lscript(θ,φ), (19.53)
the constants a/lscriptmbeing given by
a/lscriptm=integraldisplay1
−1integraldisplay2π
0bracketleftbig
Ym
/lscript(θ,φ)bracketrightbig∗f(θ,φ)dφ d(cosθ). (19.54)
This is in exact analogy with a Fourier series and is a particular example of the
general property of Sturm–Liouville solutions.
The first few spherical harmonics Ym
/lscript(θ,φ)≡Ym
/lscriptare as follows:
Y0
0=radicalBig
1
4π,Y0
1=radicalBig
3
4πcosθ,
Y±1
1=∓radicalBig
3
8πsinθexp(±iφ),Y0
2=radicalBig
5
16π(3 cos2θ−1),
Y±1
2=∓radicalBig
15
8πsinθcosθexp(±iφ),Y±2
2=radicalBig
15
32πsin2θexp(±2iφ).
19.3.3 Other equations in polar coordinates
The development of the solutions of ∇2u= 0 carried out in the previous subsection
can be employed to solve other equations in which the ∇2operator appears. Since
we have discussed the general method in some depth already, only an outline ofthe solutions will be given here.
Let us first consider the wave equation
∇
2u=1
c2∂2u
∂t2, (19.55)
and look for a separated solution of the form u=F(r)T(t), so that initially we
are separating only the spatial and time dependences. Substituting this form into
(19.55) and taking the separation constant as k2we obtain
∇2F+k2F=0,d2T
dt2+k2c2T=0. (19.56)
The second equation has the simple solution
T(t)=Aexp(iωt)+Bexp(−iωt), (19.57)
where ω=kc; this may also be expressed in terms of sines and cosines, of course.
The first equation in (19.56) is referred to as Helmholtz’s equation ; we discuss it
below.
We may treat the diffusion equation
κ∇2u=∂u
∂t
671
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
in a similar way. Separating the spatial and time dependences by assuming a
solution of the form u=F(r)T(t), and taking the separation constant as k2,w e
find
∇2F+k2F=0,dT
dt+k2κT=0.
Just as in the case of the wave equation, the spatial part of the solution satisfies
Helmholtz’s equation. It only remains to consider the time dependence, whichhas the simple solution
T(t)=Aexp(−k
2κt).
Helmholtz’s equation is clearly of central importance in the solutions of the
wave and diffusion equations. It can be solved in polar coordinates in much the
same way as Laplace’s equation, and indeed reduces to Laplace’s equation when
k= 0. Therefore, we will merely sketch the method of its solution in each of the
three polar coordinate systems.
Helmholtz’s equation in plane polars
In two-dimensional plane polar cooordinates Helmholtz’s equation takes the form
1
ρ∂
∂ρparenleftbigg
ρ∂F
∂ρparenrightbigg
+1
ρ2∂2F
∂φ2+k2F=0.
If we try a separated solution of the form F(r)= P(ρ)Φ(φ), and take the
separation constant as m2, we find
d2Φ
dφ2+m2φ=0,
d2P
dρ2+1
ρdP
dρ+parenleftbigg
k2−m2
ρ2parenrightbigg
P=0.
As for Laplace’s equation, the angular part has the familiar solution (if m/negationslash=0 )
Φ(φ)=Acosmφ+Bsinmφ,
or an equivalent form in terms of complex exponentials. The radial equation
differs from that found in the solution of Laplace’s equation, but by making thesubstitution µ=kρit is easily transformed into Bessel’s equation of order m
(discussed in chapter 16), and has the solution
P(ρ)=CJ
m(kρ)+DYm(kρ),
where Ymis a Bessel function of the second kind, which is infinite at the origin
and is not to be confused with a spherical harmonic (these are written with asuperscript as well as a subscript).
Putting the two parts of the solution together we have
F(ρ, φ)=[Acosmφ+Bsinmφ][CJ
m(kρ)+DYm(kρ)]. (19.58)
672
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
Clearly, for solutions of Helmholtz’s equation that are required to be finite at the
origin, we must set D=0 .IFind the four lowest frequency modes of oscillation of a circular drumskin of radius a
whose circumference is held fixed in a plane.
The transverse displacement u(r,t) of the drumskin satisfies the two-dimensional wave
equation
∇2u=1
c2∂2u
∂t2,
with c2=T/σ,w h e r e Tis the tension of the drumskin and σis its mass per unit area.
From (19.57) and (19.58) a separated solution of this equation, in plane polar coordinates,that is finite at the origin is
u(ρ, φ, t)=J
m(kρ)(Acosmφ+Bsinmφ)e x p(±iωt),
where ω=kc. Since we require the solution to be single-valued we must have mas an
integer. Furthermore, if the drumskin is clamped at its outer edge ρ=athen we also
require u(a, φ, t) = 0. Thus we need
Jm(ka)=0 ,
which in turn restricts the allowed values of k. The zeroes of Bessel functions can be
obtained from most books of tables, and the first few are
J0(x)=0 f o r x≈2.40,5.52,8.65,...,
J1(x)=0 f o r x≈3.83,7.02,10.17,...,
J2(x)=0 f o r x≈5.14,8.42,11.62....
The smallest value of xfor which any of the Bessel functions is zero is x≈2.40, which
occurs for J0(x). Thus the lowest-frequency mode has k=2.40/aand angular frequency
ω=2.40c/a.S i n c e m= 0 for this mode, the shape of the drumskin is
u∝J0
/
2.40ρ
a
/
;
this is illustrated in figure 19.8.
Continuing in the same way the next three modes are given by
ω=3.83c
a,u∝J1
/
3.83ρ
a
/
cosφ, J 1
/
3.83ρ
a
/
sinφ;
ω=5.14c
a,u∝J2
/
5.14ρ
a
/
cos 2φ, J 2
/
5.14ρ
a
/
sin 2φ;
ω=5.52c
a,u∝J0
/
5.52ρ
a
/
.
These modes are also shown in figure 19.8. We note that the second and third frequencies
have twocorresponding modes of oscillation; these frequencies are therefore two-fold
degenerate.
J
Helmholtz’s equation in cylindrical polars
Generalising the above method to three-dimensional cylindrical polars is straight-
forward, and following a similar procedure to that used for Laplace’s equation
673
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
ω=2.40c/a ω=3.83c/a
ω=5.14c/a ω=5.52c/aa
Figure 19.8 For a circular drumskin of radius a, the modes of oscillation
with the four lowest frequencies. The dotted lines indicate the nodes, wherethe displacement of the drumskin is always zero.
we find the separated solution of Helmholtz’s equation takes the form
F(ρ, φ, z)=bracketleftBig
AJmparenleftBig√
k2−α2ρparenrightBig
+BYmparenleftBig√
k2−α2ρparenrightBigbracketrightBig
×(Ccosmφ+Dsinmφ)[Eexp(iαz)+Fexp(−iαz)],
where αandmare separation constants. We note that the angular part of the
solution is the same as for Laplace’s equation in cylindrical polars.
Helmholtz’s equation in spherical polars
In spherical polars, we find again that the angular parts of the solution Θ( θ)Φ(φ)
are identical to those of Laplace’s equation in this coordinate system, i.e. they are
the spherical harmonics Ym
/lscript(θ,φ), and so we shall not discuss them further.
The radial equation in this case is given by
r2R/prime/prime+2rR/prime+[k2r2−/lscript(/lscript+1 ) ] R=0, (19.59)
which has an additional term k2r2Rcompared with the radial equation for the
Laplace solution. The equation (19.59) looks very much like Bessel’s equation and
can in fact be reduced to it by writing R(r)=r−1/2S(r). The function S(r)t h e n
satisfies
r2S/prime/prime+rS/prime+bracketleftBig
k2r2−parenleftbig
/lscript+1
2parenrightbig2bracketrightBig
S=0,
which, after changing the variable to µ=kr,i sB e s s e l ’ se q u a t i o no fo r d e r /lscript+1
2
and has as its solutions S(µ)=J/lscript+1/2(µ)a n d Y/lscript+1/2(µ). The separated solution to
674
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
Helmholtz’s equation in spherical polars is thus
F(r,θ,φ)=r−1/2[AJ/lscript+1/2(kr)+BY/lscript+1/2(kr)](Ccosmφ+Dsinmφ)
×[EPm
/lscript(cosθ)+FQm
/lscript(cosθ)]. (19.60)
For solutions that are finite at the origin we require B= 0, and for solutions that
are finite on the polar axis we require F=0 .
It is worth mentioning that the solutions proportional to r−1/2J/lscript+1/2(kr)w h e n
suitably normalised are called spherical Bessel functions and are denoted by j/lscript(kr):
j/lscript(µ)=radicalbiggπ
2µJ/lscript+1/2(µ).
They are trigonometric functions of µ(as discussed in chapter 16), and for /lscript=0
and/lscript=1a r eg i v e nb y
j0(µ)=s i n µ,
j1(µ)=sinµ
µ−cosµ.
The second, linearly-independent, solution of (19.59), n/lscript(µ), is derived from
Y/lscript+1/2(µ) in a similar way.
As mentioned at the beginning of this subsection, the separated solution of
the wave equation in spherical polars is the product of the time-dependent part(19.57) and a spatial part (19.60). It will be noticed that, although this solutioncorresponds to a solution of definite frequency ω=kc, the zeroes of the radial
function j
/lscript(kr) are not equally spaced in r, except for the case /lscript= 0 involving
j0(kr), and so there is no precise wavelength associated with the solution.
To conclude this subsection, let us mention briefly the Schr ¨odinger equation
for the electron in a hydrogen atom, the nucleus of which is taken at the originand is assumed massive compared with the electron. Under these circumstancesthe Schr ¨odinger equation is
−/planckover2pi1
2
2m∇2u−e2
4π/epsilon10u
r=i/planckover2pi1∂u
∂t.
For a ‘stationary-state’ solution, for which the energy is a constant Eand the time-
dependent factor Tinuis given by T(t)=Aexp(−iEt//planckover2pi1), the above equation
is similar to, but not quite the same as, the Helmholtz equation. †However, as
with the wave equation, the angular parts of the solution are identical to thosefor Laplace’s equation and are expressed in terms of spherical harmonics.
The important point to note is that for anyequation involving ∇
2,p r o v i d e d θ
andφdo not appear in the equation other than as part of ∇2, a separated-variable
†For the solution by series of the r-equation in this case the reader may consult, e.g., Schiff, Quantum
Mechanics (McGraw-Hill, 1955) p. 82.
675
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
solution in spherical polars will always lead to spherical harmonic solutions. This
is the case for the Schr ¨odinger equation describing an atomic electron whenever
the potential is central, i.e. whenever V(r)i si nf a c t V(r).
19.3.4 Solution by expansion
It is sometimes possible to use the uniqueness theorem discussed in the last
chapter, together with the results of the last few subsections, in which Laplace’s
equation (and other equations) were considered in polar coordinates, to obtainsolutions of such equations appropriate to particular physical situations.
We will illustrate the method for Laplace’s equation in spherical polars and first
assume that the required solution of ∇
2u= 0 can be written as a superposition
in the normal way:
u(r,θ,φ)=∞summationdisplay
/lscript=0/lscriptsummationdisplay
m=−/lscript(Ar/lscript+Br−(/lscript+1))Pm
/lscript(cosθ)(Ccosmφ+Dsinmφ).
(19.61)
Here, all the constants A, B, C, D may depend upon /lscriptand m, and we have
assumed that the required solution is finite on the polar axis. As usual, boundaryconditions of a physical nature will then fix or eliminate some of the constants;for example, ufinite at the origin implies all B= 0, or axial symmetry implies
that only m= 0 terms are present.
The essence of the method is then to find the remaining constants by determin-
inguat values of r,θ,φ for which it can be evaluated by other means , e.g. by direct
calculation on an axis of symmetry. Once the remaining constants have been fixed
by these special considerations to have particular values, the uniqueness theoremcan be invoked to establish that they must have these values in general.ICalculate the gravitational potential at a general point in space due to a uniform ring of
matter of radius aand total mass M.
Everywhere except on the ring the potential u(r) satisfies the Laplace equation, and so if
we use polar coordinates with the normal to the ring as polar axis, as in figure 19.9, asolution of the form (19.61) can be assumed.
We expect the potential u(r,θ,φ) to tend to zero as r→∞, and also to be finite at r=0 .
At first sight this might seem to imply that all AandB, and hence u, must be identically
zero, an unacceptable result. In fact, what it means is that different expressions must applyto different regions of space. On the ring itself we no longer have ∇
2u= 0 and so it is not
surprising that the form of the expression for uchanges there. Let us therefore take two
separate regions.
In the region r>a
(i) we must have u→0a sr→∞, implying that all A=0 ,a n d
(ii) the system is axially symmetric and so only m= 0 terms appear.
With these restrictions we can write as a trial form
u(r,θ,φ)=∞X
/lscript=0B/lscriptr−(/lscript+1)P0
/lscript(cosθ). (19.62)
676
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
yP
aθ
O−arz
x
Figure 19.9 The polar axis Ozis taken as normal to the plane of the ring of
matter and passing through its centre.
The constants B/lscriptare still to be determined; this we do by calculating directly the potential
where this can be done simply – in this case, on the polar axis.
Considering a point Pon the polar axis at a distance z(>a) from the plane of the ring
(taken as θ=π/2), all parts of the ring are at a distance ( z2+a2)1/2from it. The potential
atPis thus straightforwardly
u(z,0,φ)=−GM
(z2+a2)1/2, (19.63)
where Gis the gravitational constant. This must be the same as (19.62) for the particular
values r=z,θ=0 ,a n d φundefined. Since P0
/lscript(cosθ)=P/lscript(cosθ)w i t h P/lscript(1) = 1, putting
r=zin (19.62) gives
u(z,0,φ)=∞X
/lscript=0B/lscript
z/lscript+1. (19.64)
However, expanding (19.63) for z>a (as it applies to this region of space) we obtain
u(z,0,φ)=−GM
z
/
1−1
2
/a
z
/2
+3
8
/a
z
/4
−···
/
,
which on comparison with (19.64) gives †
B0=−GM,
B2/lscript=−GMa2/lscript(−1)/lscript(2/lscript−1)!!
2/lscript/lscript!for/lscript≥1, (19.65)
B2/lscript+1=0.
We now conclude the argument by saying that if a solution for a general point ( r,θ,φ)
exists at all, which of course we very much expect on physical grounds, then it must be(19.62) with the B
/lscriptgiven by (19.65). This is so because thus defined it is a function with
no arbitrary constants and which satisfies all the boundary conditions, and the uniqueness
†(2/lscript−1)!! = 1×3×···×(2/lscript−1).
677
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
theorem states that there is only one such function. The expression for the potential in the
region r>a is therefore
u(r,θ,φ)=−GM
r
/"
1+∞X
/lscript=1(−1)/lscript(2/lscript−1)!!
2/lscript/lscript!
/a
r
/2/lscript
P2/lscript(cosθ)
/#
.
The expression for r<a can be found in a similar way. The finiteness of uatr=0a n d
the axial symmetry give
u(r,θ,φ)=∞X
/lscript=0A/lscriptr/lscriptP0
/lscript(cosθ).
Comparing this expression for r=z,θ= 0 with the z<a expansion of (19.63), which is
valid for any z, establishes A2/lscript+1=0 , A0=−GM/a and
A2/lscript=−GM
a2/lscript+1(−1)/lscript(2/lscript−1)!!
2/lscript/lscript!,
so that the final expression valid, and convergent, for r<a is thus
u(r,θ,φ)=−GM
a
/"
1+∞X
/lscript=1(−1)/lscript(2/lscript−1)!!
2/lscript/lscript!
/r
a
/2/lscript
P2/lscript(cosθ)
/#
.
It is easy to check that the solution obtained has the expected physical value for large r
and for r= 0 and is continuous at r=a.
J
19.3.5 Separation of variables for inhomogeneous equations
So far our discussion of the method of separation of variables has been limited
to the solution of homogeneous equations such as the Laplace equation and thewave equation. The solutions of inhomogeneous PDEs are usually obtained usingthe Green’s function methods to be discussed below in section 19.5. However, as afinal illustration of the usefulness of the separation of variables, we now consider
its application to the solution of inhomogeneous equations.
Because of the added complexity in dealing with inhomogeneous equations, we
shall restrict our discussion to the solution of Poisson’s equation,
∇
2u=ρ(r), (19.66)
in spherical polar coordinates, although the general method can accommodate
other coordinate systems and equations. In physical problems the RHS of (19.66)usually contains some multiplicative constant(s). If uis the electrostatic potential
in some region of space in which ρis the density of electric charge then ∇
2u=
−ρ(r)//epsilon10. Alternatively, umight represent the gravitational potential in some
region where the matter density is given by ρ,s ot h a t∇2u=4πGρ(r).
We will simplify our discussion by assuming that the required solution uis
finite on the polar axis and also that the system possesses axial symmetry about
that axis – in which case ρdoes not depend on the azimuthal angle φ.T h ek e y
to the method is then to assume a separated form for both the solution uandthe
density term ρ.
678
19.3 SEPARATION OF VARIABLES IN POLAR COORDINATES
From the discussion of Laplace’s equation, for systems with axial symmetry
only m= 0 terms appear, and so the angular part of the solution can be
expressed in terms of Legendre polynomials P/lscript(cosθ). Since these functions form
an orthogonal set let us expand both uandρin terms of them:
u=∞summationdisplay
/lscript=0R/lscript(r)P/lscript(cosθ), (19.67)
ρ=∞summationdisplay
/lscript=0F/lscript(r)P/lscript(cosθ), (19.68)
where the coefficients R/lscript(r)a n d F/lscript(r) in the Legendre polynomial expansions
are functions of r. Since in any particular problem ρis given, we can find the
coefficients F/lscript(r) in the expansion in the usual way (see subsection 16.6.2). It then
only remains to find the coefficients R/lscript(r) in the expansion of the solution u.
Writing∇2in spherical polars and substituting (19.67) and (19.68) into (19.66)
we obtain
∞summationdisplay
/lscript=0bracketleftbiggP/lscript(cosθ)
r2d
drparenleftbigg
r2dR/lscript
drparenrightbigg
+R/lscript
r2sinθd
dθparenleftbigg
sinθdP/lscript(cosθ)
dθparenrightbiggbracketrightbigg
=∞summationdisplay
/lscript=0F/lscript(r)P/lscript(cosθ).
(19.69)
However, if, in equation (19.44) of our discussion of the angular part of the
solution to Laplace’s equation, we set m= 0 we conclude that
1
sinθd
dθparenleftbigg
sinθdP/lscript(cosθ)
dθparenrightbigg
=−/lscript(/lscript+1 )P/lscript(cosθ).
Substituting this into (19.69), we find that the LHS is greatly simplified and we
obtain
∞summationdisplay
/lscript=0bracketleftbigg1
r2d
drparenleftbigg
r2dR/lscript
drparenrightbigg
−/lscript(/lscript+1 )R/lscript
r2bracketrightbigg
P/lscript(cosθ)=∞summationdisplay
/lscript=0F/lscript(r)P/lscript(cosθ).
This relation is most easily satisfied by equating terms on both sides for each
value of /lscriptseparately, so that for /lscript=0,1,2,...we have
1
r2d
drparenleftbigg
r2dR/lscript
drparenrightbigg
−/lscript(/lscript+1 )R/lscript
r2=F/lscript(r). (19.70)
This is an ODE in which F/lscript(r) is given, and it can therefore be solved for
R/lscript(r). The solution to Poisson’s equation, u, is then obtained by making the
superposition (19.67).
679
PDES: SEPARATION OF VARIABLES AND OTHER METHODSIIn a certain system, the electric charge density ρis distributed as follows:
ρ=
/
Arcosθfor0≤r<a ,
0 forr≥a.
Find the electrostatic potential inside and outside the charge distribution, given that both
the potential and its radial derivative are continuous everywhere.
The electrostatic potential usatisfies
∇2u=
/
−(A//epsilon10)rcosθfor 0≤r<a ,
0f o r r≥a.
Forr<a the RHS can be written −(A//epsilon10)rP1(cosθ), and the coefficients in (19.68) are
simply F1(r)=−(Ar//epsilon1 0)a n d F/lscript(r)=0f o r /lscript/negationslash= 1. Therefore we need only calculate R1(r),
which satisfies (19.70) for /lscript=1 :
1
r2d
dr
/
r2dR1
dr
/
−2R1
r2=−Ar
/epsilon10.
This can be rearranged to give
r2R/prime/prime
1+2rR/prime
1−2R1=−Ar3
/epsilon10,
where the prime denotes differentiation with respect to r. The LHS is homogeneous and
the equation can be reduced by the substitution r=e x p t, and writing R1(r)=S(t), to
¨S+˙S−2S=−A
/epsilon10exp 3 t, (19.71)
where the dots indicate differentiation with respect to t.
This is an inhomogeneous second-order ODE with constant coefficients and can be
straightforwardly solved by the methods of subsection 15.2.1 to give
S(t)=c1expt+c2exp(−2t)−A
10/epsilon10exp 3 t.
Recalling that r=e x p twe find
R1(r)=c1r+c2r−2−A
10/epsilon10r3.
Since we are interested in the region r<a we must have c2= 0 for the solution to remain
finite. Thus inside the charge distribution the electrostatic potential has the form
u1(r,θ,φ)=
/
c1r−A
10/epsilon10r3
/
P1(cosθ). (19.72)
Outside the charge distribution (for r≥a), however, the electrostatic potential obeys
Laplace’s equation, ∇2u= 0, and so given the symmetry of the problem and the requirement
thatu→∞asr→∞the solution must take the form
u2(r,θ,φ)=∞X
/lscript=0B/lscript
r/lscript+1P/lscript(cosθ). (19.73)
We can now use the boundary conditions at r=ato fix the constants in (19.72) and
(19.73). The requirement of continuity of the potential and its radial derivative at r=a
imply that
u1(a, θ, φ)=u2(a, θ, φ),
∂u1
∂r(a, θ, φ)=∂u2
∂r(a, θ, φ).
680
19.4 INTEGRAL TRANSFORM METHODS
Clearly B/lscript=0f o r /lscript/negationslash= 1; carrying out the necessary differentiations and setting r=ain
(19.72) and (19.73) we obtain the simultaneous equations
c1a−A
10/epsilon10a3=B1
a2,
c1−3A
10/epsilon10a2=−2B1
a3
which may be solved to give c1=Aa2/(6/epsilon10)a n d B1=Aa5/(15/epsilon10). Since P1(cosθ)=c o s θ,
the electrostatic potentials inside and outside the charge distribution are given respectivelyby
u
1(r,θ,φ)=A
/epsilon10
/a2r
6−r3
10
/
cosθ, u 2(r,θ,φ)=Aa5
15/epsilon10cosθ
r2.
J
19.4 Integral transform methods
In the method of separation of variables our aim was to keep the independent
variables in a PDE as separate as possible. We now discuss the use of integraltransforms in solving PDEs, a method by which one of the independent variablescan be eliminated from the differential coefficients. It will be assumed that thereader is familiar with Laplace and Fourier transforms and their properties, asdiscussed in chapter 13.
The method consists simply of transforming the PDE into one containing
derivatives with respect to a smaller number of variables. Thus, if the originalequation has just two independent variables, it may be possible to reduce thePDE into a soluble ODE. The solution obtained can then (where possible) betransformed back to give the solution of the original PDE. As we shall see,
boundary conditions can usually be incorporated in a natural way.
Which sort of transform to use, and the choice of the variable(s) with respect
to which the transform is to be taken, is a matter of experience; we illustrate thisin the example below. In practice, transforms can be taken with respect to eachvariable in turn, and the transformation that affords the greatest simplificationcan be pursued further.IA semi-infinite tube of constant cross-section contains initially pure water. At time t=0,
one end of the tube is put into contact with a salt solution and maintained at a concentrationu
0. Find the total amount of salt that has diffused into the tube after time t, if the diffusion
constant is κ.
The concentration u(x, t)a tt i m e tand distance xfrom the end of the tube satisfies the
diffusion equation
κ∂2u
∂x2=∂u
∂t, (19.74)
which has to be solved subject to the boundary conditions u(0,t)=u0for all tand
u(x,0) = 0 for all x>0.
Since we are interested only in t>0, the use of the Laplace transform is suggested.
681
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
Furthermore, it will be recalled from chapter 13 that one of the major virtues of Laplace
transformations is the possibility they afford of replacing derivatives of functions by simplemultiplication by a scalar. If the derivative with respect to time were so removed, equation(19.74), would contain only differentiation with respect to a single variable. Let us thereforetake the Laplace transform of (19.74) with respect to t:Z∞
0κ∂2u
∂x2exp(−st)dt=
Z∞
0∂u
∂texp(−st)dt.
On the LHS the (double) differentiation is with respect to x, whereas the integration is
with respect to the independent variable t. Therefore the derivative can be taken outside
the integral. Denoting the Laplace transform of u(x, t)b y¯u(x, s) and using result (13.57)
to rewrite the transform of the derivative on the RHS (or by integrating directly by parts),we obtain
κ∂
2¯u
∂x2=s¯u(x, s)−u(x,0).
But from the boundary condition u(x,0) = 0 the last term on the RHS vanishes, and the
solution is immediate:
¯u(x, s)=Aexp
/
rs
κx
/
+Bexp
/
−
rs
κx
/
,
where the constants AandBmay depend on s.
We require u(x, t)→0a s x→∞ and so we must also have ¯u(∞,s) = 0; consequently
we require that A= 0. The value of Bis determined by the need for u(0,t)=u0and hence
that
¯u(0,s)=
Z∞
0u0exp(−st)dt=u0
s.
We thus conclude that the appropriate expression for the Laplace transform of u(x, t)i s
¯u(x, s)=u0
sexp
/
−
rs
κx
/
. (19.75)
To obtain u(x, t) from this result requires the inversion of this transform – a task that is
generally difficult and requires a contour integration. This is discussed in chapter 20, butfor completeness we note that the solution is
u(x, t)=u
0
/
1−erf
/x√
4κt
//
,
where erf( x) is the error function discussed in the Appendix. (The more complete sets of
mathematical tables list this inverse Laplace transform.)
In the present problem, however, an alternative method is available. Let w(t)b et h e
amount of salt that has diffused into the tube in time t;t h e n
w(t)=
Z∞
0u(x, t)dx,
and its transform is given by
¯w(s)=
Z∞
0dtexp(−st)
Z∞
0u(x, t)dx
=
Z∞
0dx
Z∞
0u(x, t)e x p(−st)dt
=
Z∞
0¯u(x, s)dx.
682
19.4 INTEGRAL TRANSFORM METHODS
Substituting for ¯u(x, s) from (19.75) into the last integral and integrating, we obtain
¯w(s)=u0κ1/2s−3/2.
This expression is much simpler to invert, and referring to the table of standard Laplace
transforms (table 13.1) we find
w(t)=2 ( κ/π)1/2u0t1/2,
which is thus the required expression for the amount of diffused salt at time t.
J
The above example shows that in some circumstances the use of a Laplace
transformation can greatly simplify the solution of a PDE. However, it will havebeen observed that (as with ODEs) the easy elimination of some derivatives isusually paid for by the introduction of a difficult inverse transformation. This
problem, although still present, is less severe for Fourier transformations.IAn infinite metal bar has an initial temperature distribution f(x)along its length. Find
the temperature distribution at a later time t.
We are interested in values of xfrom−∞to∞, which suggests Fourier transformation
with respect to x. Assuming that the solution obeys the boundary conditions u(x, t)→0
and∂u/∂x→0a s|x|→∞ , we may Fourier-transform the one-dimensional diffusion
equation (19.74) to obtain
κ√
2π
Z∞
−∞∂2u(x, t)
∂x2exp(−ikx)dx=1√
2π∂
∂t
Z∞
−∞u(x, t)exp(−ikx)dx,
where on the RHS we have taken the partial derivative with respect to toutside the
integral. Denoting the Fourier transform of u(x, t)b y
eu(k,t), and using equation (13.28) to
rewrite the Fourier transform of the second derivative on the LHS, we then have
−κk2eu(k,t)=∂
eu(k,t)
∂t.
This first-order equation has the simple solutioneu(k,t)=
eu(k,0)exp(−κk2t),
where the initial conditions giveeu(k,0) =1√
2π
Z∞
−∞u(x,0)exp(−ikx)dx
=1√
2π
Z∞
−∞f(x)exp(−ikx)dx=
ef(k).
Thus we may write the Fourier transform of the solution aseu(k,t)=
ef(k)exp(−κk2t)=√
2π
ef(k)
eG(k,t), (19.76)
where we have defined the function
eG(k,t)=(√
2π)−1exp(−κk2t). Since
eu(k,t)c a nb e
written as the product of two Fourier transforms, we can use the convolution theorem,subsection 13.1.7, to write the solution as
u(x, t)=
Z∞
−∞G(x−x/prime,t)f(x/prime)dx/prime,
683
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
where G(x, t) is the Green’s function for this problem (see subsection 15.2.5). This function
is the inverse Fourier transform of
eG(k,t) and is thus given by
G(x, t)=1
2π
Z∞
−∞exp(−κk2t)exp( ikx)dk
=1
2π
Z∞
−∞exp
/
−κt
/
k2−ix
κtk
//
dk.
Completing the square in the integrand we find
G(x, t)=1
2πexp
/
−x2
4κt
/Z∞
−∞exp
/"
−κt
/
k−ix
2κt
/2
/#
dk
=1
2πexp
/
−x2
4κt
/Z∞
−∞exp
/
−κtk/prime2
/
dk/prime
=1√
4πκtexp
/
−x2
4κt
/
,
where in the second line we have made the substitution k/prime=k−ix/(2κt) ,a n di nt h el a s t
line we have used the standard result for the integral of a Gaussian, given in subsection6.4.2. (Strictly speaking the change of variable from ktok
/primeshifts the path of integration
off the real axis, since k/primeis complex for real k, and so results in a complex integral, as will
be discussed in chapter 20. Nevertheless, in this case the path of integration can be shiftedback to the real axis without affecting the value of the integral.)
Thus the temperature in the bar at a later time tis given by
u(x, t)=1
√
4πκt
Z∞
−∞exp
/
−(x−x/prime)2
4κt
/
f(x/prime)dx/prime, (19.77)
which may be evaluated (numerically if necessary) when the form of f(x)i sg i v e n .
J
As we might expect from our discussion of Green’s functions in chapter 15,
we see from (19.77) that, if the initial temperature distribution is f(x)=δ(x−a),
i.e. a ‘point’ source at x=a, then the temperature distribution at later times is
simply given by
u(x, t)=G(x−a, t)=1√
4πκtexpbracketleftbigg
−(x−a)2
4κtbracketrightbigg
.
The temperature at several later times is illustrated in figure 19.10, which shows
that the heat diffuses out from its initial position; the width of the Gaussianincreases as√
t, a dependence on time which is characteristic of diffusion processes.
The reader may have noticed that in both examples using integral transforms
the solutions have been obtained in closed form – albeit in one case in the formof an integral. This differs from the infinite series solutions usually obtained viathe separation of variables. It should be noted that this behaviour is a result ofthe infinite range in xrather than of the transform method itself. In fact the
method of separation of variables would yield the same solutions, since in the
infinite-range case the separation constant is not restricted to take on an infinite
set of discrete values but may have any real value, with the result that the sumover λbecomes an integral, as mentioned at the end of section 19.2.
684
19.4 INTEGRAL TRANSFORM METHODS
xu
x=at1
t2
t3
Figure 19.10 Diffusion of heat from a point source in a metal bar: the curves
show the temperature uat position xfor various different times t1<t2<t3.
The area under the curves remains constant, since the total heat energy is
conserved.IAn infinite metal bar has an initial temperature distribution f(x)along its length. Find
the temperature distribution at a later time tusing the method of separation of variables.
This is the same problem as in the previous example, but we now seek a solution by
separating variables. From (19.12) a separated solution for the one-dimensional diffusionequation is
u(x, t)=[Aexp(iλx)+Bexp(−iλx)]exp(−κλ
2t),
where−λ2is the separation constant. Since the bar is infinite we do not require the
solution to take a given form at any finite value of x(for instance at x=0 )a n ds ot h e r e
is no restriction on λother than its being real. Therefore instead of the superposition of
such solutions in the form of a sum over allowed values of λwe have an integral over
allλ,
u(x, t)=1√
2π
Z∞
−∞A(λ)exp(−κλ2t)exp( iλx)dλ, (19.78)
where in taking λfrom−∞to∞we need include only one of the complex exponentials;
we have taken a factor 1 /√
2πout of A(λ) for convenience. We can see from (19.78)
that the expression for u(x, t) has the form of an inverse Fourier transform (where λis
the transform variable). Therefore, Fourier-transforming both sides and using the Fourierinversion theorem, we findeu(λ, t)=A(λ)exp(−κλ2t).
Now the initial boundary condition requires
u(x,0) =1√
2π
Z∞
−∞A(λ)exp( iλx)dλ=f(x),
685
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
from which, using the Fourier inversion theorem once more, we see that A(λ)=
ef(λ).
Therefore we haveeu(λ, t)=
ef(λ)exp(−κλ2t),
which is identical to (19.76) in the previous example (but with kreplaced by λ), and hence
leads to the same result.
J
19.5 Inhomogeneous problems – Green’s functions
In chapters 15 and 17 we encountered Green’s functions and found them a useful
tool for solving inhomogeneous linear ODEs. We now discuss their usefulness in
solving inhomogeneous linear PDEs.
For the sake of brevity we shall again denote a linear PDE by
Lu(r)=ρ(r), (19.79)
where Lis a linear partial differential operator. For example, in Laplace’s equation
we have L=∇2, whereas for Helmholtz’s equation L=∇2+k2. Note that we have
not specified the dimensionality of the problem, and (19.79) may, for example,represent Poisson’s equation in two or three (or more) dimensions. The readerwill also notice that for the sake of simplicity we have not included any timedependence in (19.79). Nevertheless, the following discussion can be generalisedto include it.
As we discussed in subsection 18.3.2, a problem is inhomogeneous if the fact
that u(r) is a solution does notimply that any constant multiple λu(r)i sa l s oa
solution. This inhomogeneity may derive from either the PDE itself or from theboundary conditions imposed on the solution.
In our discussion of Green’s function solutions of inhomogeneous ODEs (see
subsection 15.2.5) we dealt with inhomogeneous boundary conditions by making asuitable change of variable such that in the new variable the boundary conditionswere homogeneous. In an analogous way, as illustrated in the final exampleof section 19.2, it is usually possible to make a change of variables in PDEs totransform between inhomogeneity of the boundary conditions and inhomogeneityof the equation. Therefore let us assume for the moment that the boundary
conditions imposed on the solution u(r) of (19.79) are homogeneous. This most
commonly means that if we seek a solution to (19.79) in some region Vthen
on the surface Sthat bounds Vthe solution obeys the conditions u(r)=0o r
∂u/∂n =0 ,w h e r e ∂u/∂n is the normal derivative of uat the surface S.
We shall discuss the extension of the Green’s function method to the direct so-
lution of problems with inhomogeneous boundary conditions in subsection 19.5.2,
but we first highlight how the Green’s function approach to solving ODEs canbe simply extended to PDEs for homogeneous boundary conditions.
686
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
19.5.1 Similarities with Green’s functions for ODEs
As in the discussion of ODEs in chapter 15, we may consider the Green’s
function for a system described by a PDE as the response of the system to a ‘unitimpulse’ or ‘point source’. Thus if we seek a solution to (19.79) that satisfies somehomogeneous boundary conditions on u(r) then the Green’s function G(r,r
0)f o r
the problem is a solution of
LG(r,r0)=δ(r−r0), (19.80)
where r0lies in V. The Green’s function G(r,r0) must also satisfy the imposed
(homogeneous) boundary conditions.
It is understood that in (19.80) the Loperator expresses differentiation with
respect to ras opposed to r0.A l s o , δ(r−r0) is the Dirac delta function (see
chapter 13) of dimension appropriate for the problem; it may be thought of asrepresenting a unit-strength point source at r=r
0.
Following an analogous argument to that given in subsection 15.2.5 for ODEs,
if the boundary conditions on u(r) are homogeneous then a solution to (19.79)
that satisfies the imposed boundary conditions is given by
u(r)=integraldisplay
G(r,r0)ρ(r0)dV(r0), (19.81)
where the integral on r0is over some appropriate ‘volume’. In two or more
dimensions, however, the task of finding directly a solution to (19.80) that satisfiesthe imposed boundary conditions on Scan be a difficult one, and we return to
this in the next subsection.
An alternative approach is to follow a similar argument to that presented in
chapter 17 for ODEs and so to construct the Green’s function for (19.79) as a
superposition of eigenfunctions of the operator L,p r o v i d e d Lis Hermitian. By
analogy with an ordinary differential operator, a partial differential operator isHermitian if it satisfies
integraldisplay
Vv∗(r)Lw(r)dV=bracketleftbiggintegraldisplay
Vw∗(r)Lv(r)dVbracketrightbigg∗
,
where the asterisk denotes complex conjugation and vandware arbitrary func-
tions obeying the imposed (homogeneous) boundary condition on the solution ofLu(r)=0 .
The eigenfunctions u
n(r),n=0,1,2,...,o fLsatisfy
Lun(r)=λnun(r),
where λnare the corresponding eigenvalues, which are all real for an Hermitian
operator L. Furthermore, each eigenfunction must obey any imposed (homo-
geneous) boundary conditions. Using an argument analogous to that given in
687
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
chapter 17, the Green’s function for the problem is given by
G(r,r0)=∞summationdisplay
n=0un(r)u∗
n(r0)
λn. (19.82)
From (19.82) we see immediately that the Green’s function (irrespective of how
it is found) enjoys the property
G(r,r0)=G∗(r0,r).
Thus, if the Green’s function is real then it is symmetric in its two arguments.
Once the Green’s function has been obtained, the solution to (19.79) is again
given by (19.81). For PDEs this approach can become very cumbersome, however,and so we shall not pursue it further here.
19.5.2 General boundary-value problems
As mentioned above, often inhomogeneous boundary conditions can be dealt
with by making an appropriate change of variables, such that the boundaryconditions in the new variables are homogeneous although the equation itself isgenerally inhomogeneous. In this section, however, we extend the use of Green’sfunctions to problems with inhomogeneous boundary conditions (and equations).
This provides a more consistent and intuitive approach to the solution of such
boundary-value problems .
For definiteness we shall consider Poisson’s equation
∇
2u(r)=ρ(r), (19.83)
but the material of this section may be extended to other linear PDEs of the form
(19.79). Clearly, Poisson’s equation reduces to Laplace’s equation for ρ(r)=0a n d
so our discussion is equally applicable to this case.
We wish to solve (19.83) in some region Vbounded by a surface S, which may
consist of several disconnected parts. As stated above, we shall allow the possibilitythat the boundary conditions on the solution u(r) may be inhomogeneous on S,
although as we shall see this method reduces to those discussed above in thespecial case that the boundary conditions are in fact homogeneous.
The two common types of inhomogeneous boundary condition for Poisson’s
equation are (as discussed in subsection 18.6.2):
(i) Dirichlet conditions, in which u(r)i ss p e c i fi e do n S,a n d
(ii) Neumann conditions, in which ∂u/∂n is specified on S.
In general, specifying bothDirichlet andNeumann conditions on Soverdetermines
the problem and leads to there being no solution.
The specification of the surface Srequires some further comment, since S
may have several disconnected parts. If we wish to solve Poisson’s equation
688
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
ˆnˆn
ˆnV
VS
S1
S2
(a) (b)
Figure 19.11 Surfaces used for solving Poisson’s equation in different
regions V.
inside some closed surface Sthen the situation is straightforward and is shown
in figure 19.11( a). If, however, we wish to solve Poisson’s equation in the gap
between two closed surfaces (for example in the gap between two concentricconducting cylinders) then the volume Vis bounded by a surface Sthat has two
disconnected parts S
1andS2, as shown in figure 19.11( b); the direction of the
normal to the surface is always taken as pointing outof the volume V. A similar
situation arises when we wish to solve Poisson’s equation outside some closed
surface S1. In this case the volume Vis infinite but is treated formally by taking
the surface S2as a large sphere of radius Rand letting Rtend to infinity.
In order to solve (19.83) subject to either Dirichlet or Neumann boundary
conditions on S, we first remind ourselves of Green’s second theorem, equation
(11.20), which states that for two scalar functions φ(r)a n d ψ(r) defined in some
volume Vbounded by a surface S
integraldisplay
V(φ∇2ψ−ψ∇2φ)dV=integraldisplay
S(φ∇ψ−ψ∇φ)·ˆndS, (19.84)
where on the RHS it is common to write, for example, ∇ψ·ˆndSas (∂ψ/∂n )dS.
The expression ∂ψ/∂n stands for∇ψ·ˆn, the rate of change of ψin the direction
of the unit outward normal ˆnto the surface S.
The Green’s function for Poisson’s equation (19.83) must satisfy
∇2G(r,r0)=δ(r−r0), (19.85)
where r0lies in V. (As mentioned above, we may think of G(r,r0)a st h es o l u t i o n
to Poisson’s equation for a unit-strength point source located at r=r0.) Let us
for the moment impose no boundary conditions on G(r,r0).
If we now let φ=u(r)a n d ψ=G(r,r0) in Green’s theorem (19.84) then we
689
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
obtain
integraldisplay
Vbracketleftbig
u(r)∇2G(r,r0)−G(r,r0)∇2u(r)bracketrightbig
dV(r)
=integraldisplay
Sbracketleftbigg
u(r)∂G(r,r0)
∂n−G(r,r0)∂u(r)
∂nbracketrightbigg
dS(r),
where we have made explicit that the volume and surface integrals are with
respect to r. Using (19.83) and (19.85) the LHS can be simplified to give
integraldisplay
V[u(r)δ(r−r0)−G(r,r0)ρ(r)]dV(r)
=integraldisplay
Sbracketleftbigg
u(r)∂G(r,r0)
∂n−G(r,r0)∂u(r)
∂nbracketrightbigg
dS(r),(19.86)
Since r0lies within the volume V,
integraldisplay
Vu(r)δ(r−r0)dV(r)=u(r0),
and thus rearranging (19.86) the solution to Poisson’s equation (19.83) can be
written
u(r0)=integraldisplay
VG(r,r0)ρ(r)dV(r)+integraldisplay
Sbracketleftbigg
u(r)∂G(r,r0)
∂n−G(r,r0)∂u(r)
∂nbracketrightbigg
dS(r).
(19.87)
Clearly, we can interchange the roles of randr0in (19.87) if we wish. (Remember
also that, for a real Green’s function, G(r,r0)=G(r0,r).)
Equation (19.87) is central to the extension of the Green’s function method
to problems with inhomogeneous boundary conditions, and we next discuss itsapplication to both Dirichlet and Neumann boundary-value problems. But, beforedoing so, we also note that if the boundary condition on Sis in fact homogeneous,
so that u(r)=0o r ∂u(r)/∂n=0o n S, then demanding that the Green’s function
G(r,r
0) also obeys the same boundary condition causes the surface integral in
(19.87) to vanish, and we are left with the familiar form of solution given in
(19.81). The extension of (19.87) to a PDE other than Poisson’s equation isdiscussed in exercise 19.30.
19.5.3 Dirichlet problems
In a Dirichlet problem we require the solution u(r) of Poisson’s equation (19.83)
to take specific values on some surface Sthat bounds V,i . e .w er e q u i r et h a t
u(r)=f(r)o nSwhere fis a given function.
If we seek a Green’s function G(r,r
0) for this problem it must clearly satisfy
(19.85), but we are free to choose the boundary conditions satisfied by G(r,r0)i n
690
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
such a way as to make the solution (19.87) as simple as possible. From (19.87),
we see that by choosing
G(r,r0)=0 f o r ronS (19.88)
the second term in the surface integral vanishes. Since u(r)=f(r)o n S, (19.87)
then becomes
u(r0)=integraldisplay
VG(r,r0)ρ(r)dV(r)+integraldisplay
Sf(r)∂G(r,r0)
∂ndS(r). (19.89)
Thus we wish to find the Dirichlet Green’s function that
(i) satisfies (19.85) and hence is singular at r=r0,a n d
(ii) obeys the boundary condition G(r,r0)=0f o r ronS.
In general, it is, difficult to obtain this function directly, and so it is useful to
separate these two requirements. We therefore look for a solution of the form
G(r,r0)=F(r,r0)+H(r,r0),
where F(r,r0) satisfies (19.85) and has the required singular character at r=r0but
does not necessarily obey the boundary condition on S, whilst H(r,r0) satisfies
the corresponding homogeneous equation (i.e. Laplace’s equation) inside Vbut
is adjusted in such a way that the sum G(r,r0) equals zero on S. The Green’s
function G(r,r0) is still a solution of (19.85) since
∇2G(r,r0)=∇2F(r,r0)+∇2H(r,r0)=∇2F(r,r0)+0= δ(r−r0).
The function F(r,r0) is called the fundamental solution and will clearly take
different forms depending on the dimensionality of the problem. Let us first
consider the fundamental solution to (19.85) in three dimensions.IFind the fundamental solution to Poisson’s equation in three dimensions that tends to zero
as|r|→∞ .
We wish to solve
∇2F(r,r0)=δ(r−r0) (19.90)
in three dimensions, subject to the boundary condition F(r,r0)→0a s|r|→∞ .S i n c et h e
problem is spherically symmetric about r0, let us consider a large sphere Sof radius R
centred on r0, and integrate (19.90) over the enclosed volume V. We then obtainZ
V∇2F(r,r0)dV=
Z
Vδ(r−r0)dV=1, (19.91)
since Vencloses the point r0. However, using the divergence theorem,Z
V∇2F(r,r0)dV=
Z
S∇F(r,r0)·ˆndS, (19.92)
where ˆnis the unit normal to the large sphere Sat any point.
691
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
Since the problem is spherically symmetric about r0, we expect that
F(r,r0)=F(|r−r0|)=F(r),
i.e.Fhas the same value everywhere on S. Thus, evaluating the surface integral in (19.92)
and equating it to unity from (19.91), we have †
4πr2dF
dr
////
r=R=1.
Integrating this expression we obtain
F(r)=−1
4πr+c o n s t a n t ,
but, since we require F(r,r0)→0a s|r|→∞ , the constant must be zero. The fundamental
solution in three dimensions is consequently given by
F(r,r0)=−1
4π|r−r0|. (19.93)
This is clearly also the full Green’s function for Poisson’s equation subject to the boundary
condition u(r)→0a s|r|→∞ .
J
Using (19.93) we can write down the solution of Poisson’s equation to find,
for example, the electrostatic potential u(r) due to some distribution of electric
charge ρ(r). The electrostatic potential satisfies
∇2u(r)=−ρ
/epsilon10,
where u(r)→0a s|r|→∞ . Since the boundary condition on the surface at
infinity is homogeneous the surface integral in (19.89) vanishes, and using (19.93)we recover the familiar solution
u(r
0)=integraldisplayρ(r)
4π/epsilon10|r−r0|dV(r), (19.94)
where the volume integral is over all space.
We can develop an analogous theory in two dimensions. As before the funda-
mental solution satisfies
∇2F(r,r0)=δ(r−r0), (19.95)
where δ(r−r0) is now the two-dimensional delta function. Following an analogous
method to that used in the previous example, we find the fundamental solutionin two dimensions to be given by
F(r,r
0)=1
2πln|r−r0|+c o n s t a n t . (19.96)
†A vertical bar to the right of an expression is a common alternative notation to enclosing the
expression in square brackets; as usual, the subscript shows the value of the variable at which the
expression is to be evaluated.
692
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
From the form of the solution we see that in two dimensions we cannot apply
the condition F(r,r0)→0a s|r|→∞ , and in this case the constant does not
necessarily vanish.
We now return to the task of constructing the full Dirichlet Green’s function. To
do so we wish to add to the fundamental solution a solution of the homogeneous
equation (in this case Laplace’s equation) such that G(r,r0)=0o n S,a sr e q u i r e d
by (19.89) and its attendant conditions. The appropriate Green’s function isconstructed by adding to the fundamental solution ‘copies’ of itself that represent‘image’ sources at different locations outside V. Hence this approach is called the
method of images .
In summary, if we wish to solve Poisson’s equation in some region Vsubject to
Dirichlet boundary conditions on its surface Sthen the procedure and argument
are as follows.
(i) To the single source δ(r−r
0) inside Vadd image sources outside V
Nsummationdisplay
n=1qnδ(r−rn) with rnoutside V,
where the positions rnand the strengths qnof the image sources are to be
determined as described in step (iii) below.
(ii) Since all the image sources lie outside V, the fundamental solution cor-
responding to each source satisfies Laplace’s equation inside V. Thus we
may add the fundamental solutions F(r,rn) corresponding to each image
source to that corresponding to the single source inside V, obtaining the
Green’s function
G(r,r0)=F(r,r0)+Nsummationdisplay
n=1qnF(r,rn).
(iii) Now adjust the positions rnand strengths qnof the image sources so
that the required boundary conditions are satisfied on S. For a Dirichlet
Green’s function we require G(r,r0)=0f o r ronS.
(iv) The solution to Poisson’s equation subject to the Dirichlet boundary
condition u(r)=f(r)o n Sis then given by (19.89).
In general it is very difficult to find the correct positions and strengths for the
images, i.e. to make them such that the boundary conditions on Sare satisfied.
Nevertheless, it is possible to do so for certain problems that have simple geometry.In particular, for problems in which the boundary Sconsists of straight lines (in
two dimensions) or planes (in three dimensions), positions of the image points
can be deduced simply by imagining the boundary lines or planes to be mirrorsin which the single source in V(atr
0) is reflected.
693
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
yz
xVr0
r1+
−
Figure 19.12 The arrangement of images for solving Laplace’s equation in
the half-space z>0.ISolve Laplace’s equation ∇2u=0in three dimensions in the half-space z>0, given that
u(r)=f(r)on the plane z=0.
The surface Sbounding Vconsists of the xy-plane and the surface at infinity. Therefore,
the Dirichlet Green’s function for this problem must satisfy G(r,r0)=0o n z=0a n d
G(r,r0)→0a s|r|→∞ . Thus it is clear in this case that we require one image source at a
position r1that is the reflection of r0in the plane z= 0, as shown in figure 19.12 (so that
r1lies in z<0, outside the region in which we wish to obtain a solution). It is also clear
that the strength of this image should be −1.
Therefore by adding the fundamental solutions corresponding to the original source
and its image we obtain the Green’s function
G(r,r0)=−1
4π|r−r0|+1
4π|r−r1|, (19.97)
where r1is the reflection of r0in the plane z=0 ,i . e .i f r0=(x0,y0,z0)t h e n r1=(x0,y0,−z0).
Clearly G(r,r0)→0a s|r|→∞ as required. Also G(r,r0)=0o n z= 0, and so (19.97) is
the desired Dirichlet Green’s function.
The solution to Laplace’s equation is then given by (19.89) with ρ(r)=0 ,
u(r0)=
Z
Sf(r)∂G(r,r0)
∂ndS(r). (19.98)
Clearly the surface at infinity makes no contribution to this integral. The outward-pointing
unit vector normal to the xy-plane is simply ˆn=−k(where kis the unit vector in the
z-direction), and so
∂G(r,r0)
∂n=−∂G(r,r0)
∂z=−k·∇G(r,r0).
We may evaluate this normal derivative by writing the Green’s function (19.97) explicitly
in terms of x,yandz(and x0,y0andz0) and calculating the partial derivative with respect
694
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
tozdirectly. It is usually quicker, however, to use the fact that †
∇|r−r0|=r−r0
|r−r0|; (19.99)
thus
∇G(r,r0)=r−r0
4π|r−r0|3−r−r1
4π|r−r1|3.
Since r0=(x0,y0,z0)a n d r1=(x0,y0,−z0) the normal derivative is given by
−∂G(r,r0)
∂z=−k·∇G(r,r0)
=−z−z0
4π|r−r0|3+z+z0
4π|r−r1|3.
Therefore on the surface z= 0, writing out the dependence on x,yandzexplicitly, we
have
−∂G(r,r0)
∂z
////
z=0=2z0
4π[(x−x0)2+(y−y0)2+z2
0]3/2.
Inserting this expression into (19.98) we obtain the solution
u(x0,y0,z0)=z0
2π
Z∞
−∞
Z∞
−∞f(x, y)
[(x−x0)2+(y−y0)2+z2
0]3/2dx dy.
J
An analogous procedure may be applied in two-dimensional problems. For
example, in solving Poisson’s equation in two dimensions in the half-space x>0
we again require just one image charge, of strength q1=−1, at a position r1that
is the reflection of r0in the line x= 0. Since we require G(r,r0)=0w h e n rlies
onx= 0, the constant in (19.96) must equal zero, and so the Dirichlet Green’s
function is
G(r,r0)=1
2πparenleftbig
ln|r−r0|−ln|r−r1|parenrightbig
.
Clearly G(r,r0) tends to zero as |r|→∞ . If, however, we wish to solve the two-
dimensional Poisson equation in the quarter space x>0,y>0, then more image
points are required.
†Since|r−r0|2=(r−r0)·(r−r0) we have∇|r−r0|2=2 (r−r0),from which we obtain
∇(|r−r0|2)1/2=1
22(r−r0)
(|r−r0|2)1/2=r−r0
|r−r0|.
Note that this result holds in two andthree dimensions.
695
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
−λ−λ
+λ+λ
x0
y0r0
r1 r2r3
C xy
V
Figure 19.13 The arrangement of images for finding the force on a line
charge situated in the (two-dimensional) quarter-space x>0,y>0, when the
planes x=0a n d y= 0 are earthed.IA line charge in the z-direction of charge density λis placed at some position r0in the
quarter-space x>0,y>0. Calculate the force per unit length on the line charge due to
the presence of thin earthed plates along x=0andy=0.
Here we wish to solve Poisson’s equation
∇2u=−λ
/epsilon10δ(r−r0)
in the quarter space x>0,y>0. It is clear that we require three image line charges
with positions and strengths as shown in figure 19.13 (all of which lie outside the regionin which we seek a solution). The boundary condition that the electrostatic potential uis
zero on x=0a n d y= 0 (shown as the ‘curve’ Cin figure 19.13) is then automatically
satisfied, and so this system of image charges is directly equivalent to the original situationof a single line charge in the presence of the earthed plates along x=0a n d y= 0. Thus
the electrostatic potential is simply equal to the Dirichlet Green’s function
u(r)=G(r,r
0)=−λ
2π/epsilon10
/;
ln|r−r0|−ln|r−r1|+l n|r−r2|−ln|r−r3|
/
,
which equals zero on Cand on the ‘surface’ at infinity.
The force on the line charge at r0, therefore, is simply that due to the three line charges
atr1,r2andr3. The elecrostatic potential due to a line charge at ri,i= 1, 2 or 3, is given
by the fundamental solution
ui(r)=∓λ
2π/epsilon10ln|r−ri|+c,
the upper or lower sign being taken according to whether the line charge is positive or
negative respectively. Therefore the force per unit length on the line charge at r0, due to
the one at ri,i sg i v e nb y
−λ∇ui(r)
////
r=r0=±λ2
2π/epsilon10r0−ri
|r0−ri|2.
696
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
Adding the contributions from the three image charges shown in figure 19.13, the total
force experienced by the line charge at r0is
F=λ2
2π/epsilon10
/
−r0−r1
|r0−r1|2+r0−r2
|r0−r2|2−r0−r3
|r0−r3|2
/
,
where, from the figure, r0−r1=2y0j,r0−r2=2x0i+2y0jandr0−r3=2x0i. Thus, in
terms of x0andy0, the total force on the line charge due to the charge induced on the
plates is given by
F=λ2
2π/epsilon10
/
−1
2y0j+2x0i+2y0j
4x2
0+4y2
0−1
2x0i
/
=−λ2
4π/epsilon10(x2
0+y2
0)
/y2
0
x0i+x2
0
y0j
/
.
J
Further generalisations are possible. For instance, solving Poisson’s equation in
the two-dimensional strip −∞<x<∞,0<y<b requires an infinite series of
image points.
So far we have considered problems in which the boundary Sconsists of
straight lines (in two dimensions) or planes (in three dimensions), in which simple
reflection of the source at r0in these boundaries fixes the positions of the image
points. For more complicated (curved) boundaries this is no longer possible, andfinding the appropriate position(s) and strength(s) of the image source(s) requiresfurther work.IUse the method of images to find the Dirichlet Green’s function for solving Poisson’s
equation outside a sphere of radius acentred at the origin.
We need to find a solution of Poisson’s equation valid outside the sphere of radius a.
Since an image point r1cannot lie in this region, it must be located within the sphere. The
Green’s function for this problem is therefore
G(r,r0)=−1
4π|r−r0|−q
4π|r−r1|,
where|r0|>a,|r1|<aandqis the strength of the image which we have yet to determine.
Clearly, G(r,r0)→0 on the surface at infinity.
By symmetry we expect the image point r1to lie on the same radial line as the original
source, r0, as shown in figure 19.14, and so r1=kr0where k<1. However, for a Dirichlet
Green’s function we require G(r−r0)=0o n|r|=a, and the form of the Green’s function
suggests that we need
|r−r0|∝|r−r1|for all|r|=a. (19.100)
Referring to figure 19.14, if this relationship is to hold over the whole surface of the
sphere, then it must certainly hold for the points AandB. We thus require
|r0|−a
a−|r1|=|r0|+a
a+|r1|,
which reduces to |r1|=a2/|r0|. Therefore the image point must be located at the position
r1=a2
|r0|2r0.
697
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
yz
a
−aBAV
x−a
|r0|r1r0 +1
Figure 19.14 The arrangement of images for solving Poisson’s equation
outside a sphere of radius acentred at the origin. For a charge +1 at r0,t h e
image point r1is given by ( a/|r0|)2r0and the strength of the image charge is
−a/|r0|.
It may now be checked that, for this location of the image point, (19.100) is satisfied
over the whole sphere. Using the geometrical result
|r−r1|2=|r|2−2a2
|r0|2r·r0+a4
|r0|2
=a2
|r0|2
/;
|r0|2−2r·r0+a2
/
for|r|=a, (19.101)
we see that, on the surface of the sphere,
|r−r1|=a
|r0||r−r0|for|r|=a. (19.102)
Therefore, in order that G=0a t|r|=a, the strength of the image charge must be
−a/|r0|. Consequently, the Dirichlet Green’s function for the exterior of the sphere is
G(r,r0)=−1
4π|r−r0|+a/|r0|
4π|r−(a2/|r0|2)r0|.
For a less formal treatment of the same problem see exercise 19.24.
J
If we seek solutions to Poisson’s equation in the interior of a sphere then the
above analysis still holds, but randr0are now inside the sphere and the image
r1lies outside it.
For two-dimensional Dirichlet problems outside the circle |r|=a, we are led
by arguments similar to those employed previously to use the same image pointas in the three-dimensional case, namely
r
1=a2
|r0|2r0. (19.103)
698
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
As illustrated below, however, it is usually necessary to take the image strength
as−1 in two-dimensional problems.ISolve Laplace’s equation in the two-dimensional region |r|≤a, subject to the boundary
condition u=f(φ)on|r|=a.
In this case we wish to find the Dirichlet Green’s function in the interior of a disc of
radius a, so the image charge must lie outside the disc. Taking the strength of the image
to be−1, we have
G(r,r0)=1
2πln|r−r0|−1
2πln|r−r1|+c,
where r1=(a2/|r0|2)r0lies outside the disc, and cis a constant that includes the strength
of the image charge and does not necessarily equal zero.
Since we require G(r,r0)=0w h e n |r|=a, the value of the constant cis determined,
and the Dirichlet Green’s function for this problem is given by
G(r,r0)=1
2π
/
ln|r−r0|−ln
///
/
r−a2
|r0|2r0
///
/
−ln|r0|
a
/
. (19.104)
Using plane polar coordinates, the solution to the boundary-value problem can be written
as a line integral around the circle ρ=a:
u(r0)=
Z
Cf(r)∂G(r,r0)
∂ndl
=
Z2π
0f(r)∂G(r,r0)
∂ρ
////
ρ=aad φ . (19.105)
The normal derivative of the Green’s function (19.104) is given by
∂G(r,r0)
∂ρ=r
|r|·∇G(r,r0)
=r
2π|r|·
/r−r0
|r−r0|2−r−r1
|r−r1|2
/
. (19.106)
Using the fact that r1=(a2/|r0|2)r0and the geometrical result (19.102), we find that
∂G(r,r0)
∂ρ
////
ρ=a=a2−|r0|2
2πa|r−r0|2.
In plane polar coordinates, r=ρcosφi+ρsinφjandr0=ρ0cosφ0i+ρ0sinφ0j,a n d
so
∂G(r,r0)
∂ρ
////
ρ=a=
/1
2πa
/a2−ρ2
0
a2+ρ2
0−2aρ0cos(φ−φ0).
On substituting into (19.105), we obtain
u(ρ0,φ0)=1
2π
Z2π
0(a2−ρ2
0)f(φ)dφ
a2+ρ2
0−2aρ0cos(φ−φ0), (19.107)
which is the solution to the problem.
J
699
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
19.5.4 Neumann problems
In a Neumann problem we require the normal derivative of the solution of
Poisson’s equation to take on specific values on some surface Sthat bounds V,
i.e. we require ∂u(r)/∂n=f(r)o nS,w h e r e fis a given function. As we shall see,
much of our discussion of Dirichlet problems can be immediately taken over intothe solution of Neumann problems.
As we proved in section 18.7 of the previous chapter, specifying Neumann
boundary conditions determines the relevant solution of Poisson’s equation towithin an (unimportant) additive constant. Unlike Dirichlet conditions, Neumannconditions impose a self-consistency requirement. In order for a solution uto exist,
it is necessary that the following consistency condition holds:
integraldisplay
Sfd S=integraldisplay
S∇u·ˆndS=integraldisplay
V∇2ud V=integraldisplay
Vρd V, (19.108)
where we have used the divergence theorem to convert the surface integral into
a volume integral. As a physical example, the integral of the normal componentof an electric field over a surface bounding a given volume cannot be chosenarbitrarily when the charge inside the volume has already been specified (Gauss’stheorem).
Let us again consider (19.87), which is central to our discussion of Green’s
functions in inhomogeneous problems. It reads
u(r
0)=integraldisplay
VG(r,r0)ρ(r)dV(r)+integraldisplay
Sbracketleftbigg
u(r)∂G(r,r0)
∂n−G(r,r0)∂u(r)
∂nbracketrightbigg
dS(r).
As always, the Green’s function must obey
∇2G(r,r0)=δ(r−r0),
where r0lies in V. In the solution of Dirichlet problems in the previous subsection,
we chose the Green’s function to obey the boundary condition G(r,r0)=0o n S
and, in a similar way, we might wish to choose ∂G(r,r0)/∂n= 0 in the solution of
Neumann problems. However, in general this is notpermitted since the Green’s
function must obey the consistency condition
integraldisplay
S∂G(r,r0)
∂ndS=integraldisplay
S∇G(r,r0)·ˆndS=integraldisplay
V∇2G(r,r0)dV=1.
The simplest permitted boundary condition is therefore
∂G(r,r0)
∂n=1
AforronS,
where Ais the area of the surface S; this defines a Neumann Green’s function .
If we require ∂u(r)/∂n=f(r)o n S, the solution to Poisson’s equation is given
700
19.5 INHOMOGENEOUS PROBLEMS – GREEN’S FUNCTIONS
by
u(r0)=integraldisplay
VG(r,r0)ρ(r)dV(r)+1
Aintegraldisplay
Su(r)dS(r)−integraldisplay
SG(r,r0)f(r)dS(r)
=integraldisplay
VG(r,r0)ρ(r)dV(r)+/angbracketleftu(r)/angbracketrightS−integraldisplay
SG(r,r0)f(r)dS(r), (19.109)
where/angbracketleftu(r)/angbracketrightSis the average of uover the surface Sand is a freely specifiable
constant. For Neumann problems in which the volume Vis bounded by a surface
Sat infinity, we do not need the /angbracketleftu(r)/angbracketrightSterm. For example, if we wish to solve
a Neumann problem outside the unit sphere centred at the origin then r>a
is the region Vthroughout which we require the solution; this region may be
considered as being bounded by two disconnected surfaces, the surface of thesphere and a surface at infinity. By requiring that u(r)→0a s|r|→∞ ,t h et e r m
/angbracketleftu(r)/angbracketright
Sbecomes zero.
As mentioned above, much of our discussion of Dirichlet problems can be
taken over into the solution of Neumann problems. In particular, we may use the
method of images to find the appropriate Neumann Green’s function.ISolve Laplace’s equation in the two-dimensional region |r|≤asubject to the boundary
condition ∂u/∂n =f(φ)on|r|=a,w i t h
R2π
0f(φ)dφ=0as required by the consistency
condition (19.108).
Let us assume, as in Dirichlet problems with this geometry, that a single image charge isplaced outside the circle at
r
1=a2
|r0|2r0,
where r0is the position of the source inside the circle (see equation (19.103)). Then, from
(19.102), we have the useful geometrical result
|r−r1|=a
|r0||r−r0|for|r|=a. (19.110)
Leaving the strength qof the image as a parameter, the Green’s function has the form
G(r,r0)=1
2π
/;
ln|r−r0|+qln|r−r1|+c
/
. (19.111)
Using plane polar coordinates, the radial (i.e. normal) derivative of this function is given
by
∂G(r,r0)
∂ρ=r
|r|·∇G(r,r0)
=r
2π|r|·
/r−r0
|r−r0|2+q(r−r1)
|r−r1|2
/
.
Using (19.110), on the circumference of the circle ρ=athe radial derivative is
∂G(r,r0)
∂ρ
////
ρ=a=1
2π|r|
/|r|2−r·r0
|r−r0|2+q|r|2−q(a2/|r0|2)r·r0
(a2/|r0|2)|r−r0|2
/
=1
2πa1
|r−r0|2
/
|r|2+q|r0|2−(1 +q)r·r0
/
,
701
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
where we have set |r|2=a2in the second term on the RHS, but not in the first. If we take
q= 1, the radial derivative simplifies to
∂G(r,r0)
∂ρ
////
ρ=a=1
2πa,
or 1/Lwhere Lis the length of the circumference, and so (19.111) with q=1i st h e
required Neumann Green’s function.
Since ρ(r) = 0, the solution to our boundary-value problem is now given by (19.109) as
u(r0)=/angbracketleftu(r)/angbracketrightC−
Z
CG(r,r0)f(r)dl(r),
where the integral is around the circumference of the circle C. In plane polar coordinates
r=ρcosφi+ρsinφjandr0=ρ0cosφ0i+ρ0sinφ0j, and again using (19.110) we find
that on Cthe Green’s function is given by
G(r,r0)|ρ=a=1
2π
/
ln|r−r0|+l n
/a
|r0||r−r0|
/
+c
/
=1
2π
/
ln|r−r0|2+l na
|r0|+c
/
=1
2π
/
ln
/
a2+ρ2
0−2aρ0cos(φ−φ0)
/
+l na
ρ0+c
/
. (19.112)
Since dl=ad φonC, the solution to the problem is given by
u(ρ0,φ0)=/angbracketleftu/angbracketrightC−a
2π
Z2π
0f(φ)ln[a2+ρ2
0−2aρ0cos(φ−φ0)]dφ.
The contributions of the final two terms terms in the Green’s function (19.112) vanish
because
R2π
0f(φ)dφ= 0. The average value of uaround the circumference, /angbracketleftu/angbracketrightC, is a freely
specifiable constant as we would expect for a Neumann problem. This result should becompared with the result (19.107) for the corresponding Dirichlet problem, but it shouldbe remembered that in the one case f(φ) is a potential, and in the other the gradient of a
potential.J
19.6 Exercises
19.1 Solve the following first-order partial differential equations by separating the
variables:
(a)∂u
∂x−x∂u
∂y=0 ; ( b ) x∂u
∂x−2y∂u
∂y=0.
19.2 A conducting cube has as its six faces the planes x=±a,y=±aandz=±a,
and contains no internal heat sources. Verify that the temperature distribution
u(x, y, z, t )=Acosπx
asinπz
aexp
/
−2κπ2t
a2
/
obeys the appropriate diffusion equation. Across which faces is there heat flow?
What is the direction and rate of heat flow at the point (3 a/4,a /4,a)a tt i m e
t=a2/(κπ2)?
19.3 The wave equation describing the transverse vibrations of a stretched membrane
under tension Tand having a uniform surface density ρis
T
/∂2u
∂x2+∂2u
∂y2
/
=ρ∂2u
∂t2.
702
19.6 EXERCISES
Find a separable solution appropriate to a membrane stretched on a frame of
length aand width b, showing that the natural angular frequencies of such a
membrane are
ω2=π2T
ρ
/n2
a2+m2
b2
/
,
where nandmare any positive integers.
19.4 Schr ¨odinger’s equation for a non-relativistic particle in a constant potential region
can be taken as
−
/~2
2m
/∂2u
∂x2+∂2u
∂y2+∂2u
∂z2
/
=i /~∂u
∂t.
(a) Find a solution, separable in the four independent variables, that can be
written in the form of a plane wave,
ψ(x, y, z, t )=Aexp[i(k·r−ωt)].
Using the relationships associated with de Broglie ( p= /~k) and Einstein
(E= /~ω), show that the separation constants must be such that
p2
x+p2
y+p2
z=2mE.
(b) Obtain a different separable solution describing a particle confined to a box
of side a(ψmust vanish at the walls of the box). Show that the energy of
the particle can only take the quantised values
E=
/~2π2
2ma2(n2
x+n2
y+n2
z),
where nx,ny,nzare integers.
19.5 Denoting the three terms of ∇2in spherical polars by ∇2
r,∇2
θ,∇2
φin an obvious
way, evaluate ∇2
ru, etc. for the two functions given below and verify that, in each
case, although the individual terms are not necessarily zero their sum ∇2uis zero.
Identify the corresponding values of /lscriptandm.
(a)u(r,θ,φ)=
/
Ar2+B
r3
/3c os2θ−1
2.
(b)u(r,θ,φ)=
/
Ar+B
r2
/
sinθexpiφ.
19.6 Prove that the expression given in equation (19.47) for the associated Legendre
function Pm
/lscript(µ) satisfies the appropriate equation, (19.45), as follows.
(a) Evaluate dPm
/lscript(µ)/dµandd2Pm
/lscript(µ)/dµ2using the forms given in (19.47) and
substitute them into (19.45).
(b) Differentiate Legendre’s equation mtimes using Leibniz’ theorem.
(c) Show that the equations obtained in (a) and (b) are multiples of each other,
and hence that the validity of (b) implies that of (a).
19.7 Use the expressions at the end of subsection 19.3.2 to verify for /lscript=0,1,2t h a t
/lscriptX
m=−/lscript|Ym
/lscript(θ,φ)|2=2/lscript+1
4π
and so is independent of the values of θandφ.T h i si st r u ef o ra n y /lscript, but
a general proof is more involved. This result helps to reconcile intuition withthe apparently arbitrary choice of polar axis in a general quantum mechanicalsystem.
703
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
19.8 Express the function
f(θ,φ)=s i n θ[sin2(θ/2)cos φ+icos2(θ/2)sin φ]+s i n2(θ/2)
as a sum of spherical harmonics.
19.9 Continue the analysis of exercise 10.20, concerned with the flow of a very viscous
fluid past a sphere, to find the full expression for the stream function ψ(r,θ). At
the surface of the sphere r=athe velocity field u=0, whilst far from the sphere
ψ/similarequal(Ur2sin2θ)/2.
Show that f(r) can be expressed as a superposition of powers of r,a n d
determine which powers give acceptable solutions. Hence show that
ψ(r,θ)=U
4
/
2r2−3ar+a3
r
/
sin2θ.
19.10 The motion of a very viscous fluid in the two-dimensional (wedge) region −α<
φ<α can be described in ( ρ, φ) coordinates by the (biharmonic) equation
∇2∇2ψ≡∇4ψ=0,
together with the boundary conditions ∂ψ/∂φ =0a t φ=±α, which represents
the fact that there is no radial fluid velocity close to either of the boundingwalls because of the viscosity, and ∂ψ/∂ρ =±ρatφ=±α, which imposes the
condition that azimuthal flow increases linearly with ralong any radial line.
Assuming a solution in separated-variable form, show that the full expression forψis
ψ(ρ, φ)=ρ
2
2sin 2φ−2φcos 2α
sin 2α−2αcos 2α.
19.11 A circular disk of radius ais such a way that its perimeter ρ=ais maintained
with a temperature distribution A+Bcos2φ,w h e r e ρandφare plane polar
coordinates and AandBare constants. Find the temperature T(ρ, φ) everywhere
in the region ρ<a .
19.12 (a) Find the form of the solution of Laplace’s equation in plane polar coordinates
ρ, φthat takes the value +1 for 0 <φ<π and the value −1f o r−π<φ< 0,
when ρ=a.
(b) For a point ( x, y) on or inside the circle x2+y2=a2, identify the angles α
andβdefined by
α=t a n−1y
a+xand β=t a n−1y
a−x.
Show that u(x, y)=( 2 /π)(α+β) is a solution of Laplace’s equation that
satisfies the boundary conditions given in (a).
(c) Deduce a Fourier series expansion for the function
tan−1sinφ
1+c o s φ+t a n−1sinφ
1−cosφ.
19.13 The free transverse vibrations of a thick rod satisfy the equation
a4∂4u
∂x4+∂2u
∂t2=0.
Obtain a solution in separated-variable form and, for a rod clamped at one end,
x= 0, and free at the other, x=L, show that the angular frequency of vibration
ωsatisfies
cosh
/ω1/2L
a
/
=−sec
/ω1/2L
a
/
.
704
19.6 EXERCISES
(At a clamped end both uand∂u/∂x vanish, whilst at a free end, where there is
no bending moment, ∂2u/∂x2and∂3u/∂x3are both zero.)
19.14 A membrane is stretched between two concentric rings of radii aandb(b>a).
If the smaller ring is transversely distorted from the planar configuration by anamount c|φ|,−π≤φ≤π, show that the membrane then has a shape given by
u(ρ, φ)=cπ
2ln(b/ρ)
ln(b/a)−4c
π
X
moddam
m2(b2m−a2m)
/b2m
ρm−ρm
/
cosmφ.
19.15 A string of length L, fixed at its two ends, is plucked at its mid-point by an
amount Aand then released. Prove that the subsequent displacement is given by
u(x, t)=∞X
n=08A
π2(2n+1 )2sin
/(2n+1 )πx
L
/
cos
/(2n+1 )πct
L
/
,
where, in the usual notation, c2=T/ρ.
Find the total kinetic energy of the string when it passes through its unplucked
position, by calculating it in each mode (each n) and summing, using the result
∞X
01
(2n+1 )2=π2
8.
Confirm that the total energy is equal to the work done in plucking the string
initially.
19.16 Prove that the potential for ρ<a associated with a vertical split cylinder of
radius a, the two halves of which (cos φ>0a n dc o s φ<0) are maintained at
equal and opposite potentials ±V,i sg i v e nb y
u(ρ, φ)=4V
π∞X
n=0(−1)n
2n+1
/ρ
a
/2n+1
cos(2 n+1 )φ.
19.17 A conducting spherical shell of radius ais cut round its equator and the two
halves connected to voltages of + Vand−V. Show that an expression for the
potential at the point ( r,θ,φ) anywhere inside the two hemispheres is
u(r,θ,φ)=V∞X
n=0(−1)n(2n)!(4n+3 )
22n+1n!(n+1 ) !
/r
a
/2n+1
P2n+1(cosθ).
(This is the spherical polar analogue of the previous question.)
19.18 A slice of biological material of thickness Lis placed into a solution of a
radioactive isotope of constant concentration C0at time t=0 .F o ral a t e rt i m e t
find the concentration of radioactive ions at a depth xinside one of its surfaces
if the diffusion constant is κ.
19.19 Two identical copper bars are each of length a. Initially, one is at 0◦Ca n dt h e
other at 100◦C; they are then joined together end to end and thermally isolated.
Obtain in the form of a Fourier series an expression u(x, t) for the temperature
at any point a distance xfrom the join at a later time t.( B e a ri nm i n dt h eh e a t
flow conditions at the free ends of the bars.)
Taking a=0.5m estimate the time it takes for one of the free ends to
attain a temperature of 55◦C. The thermal conductivity of copper is 3 .8×
102Jm−1K−1s−1, and its specific heat capacity is 3 .4×106Jm−3K−1.
19.20 A sphere of radius aand thermal conductivity k1is surrounded by an infinite
medium of conductivity k2in which, far away, the temperature tends to T∞.
A distribution of heat sources q(θ) embedded in the sphere’s surface establish
steady temperature fields T1(r,θ) inside the sphere and T2(r,θ) outside it. It can
705
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
be shown, by considering the heat flow through a small volume that includes
part of the sphere’s surface, that
k1∂T1
∂r−k2∂T2
∂r=q(θ)o n r=a.
Given that
q(θ)=1
a∞X
n=0qnPn(cosθ),
find complete expressions for T1(r,θ)a n d T2(r,θ). What is the temperature at
the centre of the sphere?
19.21 Using result (19.77) from the worked example in the text, find the general
expression for the temperature u(x, t) in the bar, given that the temperature
distribution at time t=0i s u(x,0) = exp(−x2/a2).
19.22 (a) Show that the gravitational potential due to a uniform disc of radius aand
mass M, centred at the origin, is given for r<a by
2GM
a
/
1−r
aP1(cosθ)+1
2
/r
a
/2
P2(cosθ)−1
8
/r
a
/4
P4(cosθ)+···
/
,
and for r>a by
GM
r
/
1−1
4
/a
r
/2
P2(cosθ)+1
8
/a
r
/4
P4(cosθ)−···
/
,
where the polar axis is normal to the plane of the disc.
(b) Reconcile the presence of a term P1(cosθ), which is odd under θ→π−θ,
with the symmetry with respect to the plane of the disc of the physicalsystem.
(c) Deduce that the gravitational field near an infinite sheet of matter of constant
density ρper unit area is 2 πGρ.
19.23 In the region −∞<x,y<∞and−t≤z≤t, a charge-density wave ρ(r)=
Acosqx,i nt h e x-direction, is represented by
ρ(r)=e
iqx
√
2π
Z∞
−∞˜ρ(α)eiαzdα.
The resulting potential is represented by
V(r)=eiqx
√
2π
Z∞
−∞˜V(α)eiαzdα.
Determine the relationship between ˜V(α)a n d ˜ρ(α), and hence show that the
potential at the point ( x,0,0) is
A
π
Z∞
−∞sinkt
k(k2+q2)dk.
19.24 Point charges qand−qa/b(with a<b) are placed respectively at a point P,a
distance bfrom the origin O, and a point Qbetween OandP,ad i s t a n c e a2/b
from O. Show, by considering similar triangles QOSandSOP,w h e r e Sis any
point on the surface of the sphere centred at Oand of radius a, that the net
potential anywhere on the sphere due to the two charges is zero.
Use this result (backed up by the uniqueness theorem) to find the force with
which a point charge qplaced a distance bfrom the centre of a spherical
conductor of radius a(<b) is attracted to the sphere (i) if the sphere is earthed,
and (ii) if the sphere is uncharged and insulated.
706
19.6 EXERCISES
19.25 Find the Green’s function G(r,r0) in the half-space z>0 for the solution of
∇2Φ = 0 with Φ specified in cylindrical polar coordinates ( ρ, φ, z) on the plane
z=0b y
Φ(ρ, φ, z)=
/(
1f o r ρ≤1,
1/ρforρ>1.
Determine the variation of Φ(0 ,0,z)a l o n gt h e z-axis.
19.26 Electrostatic charge is distributed in a sphere of radius Rcentred on the origin.
Determine the form of the resultant potential φ(r) at distances much greater than
R, as follows.
(a) express in the form of an integral over all space the solution of
∇2φ=−ρ(r)
/epsilon10;
(b) show that, for r/greatermuchr/prime,
|r−r/prime|=r−r·r/prime
r+O
/1
r
/
.
(c) use results (a) and (b) to show that φ(r)h a st h ef o r m
φ(r)=M
r+d·r
r3+O
/1
r3
/
;
Find expressions for Mandd, and identify them physically.
19.27 Find, in the form of an infinite series the Green’s function of the ∇2operator for
the Dirichlet problem in the region −∞<x<∞,−∞<y<∞,−c≤z≤c.
19.28 Find the Green’s function for the three-dimensional Neumann problem
∇2φ=0 f o r z>0a n d∂φ
∂z=f(x, y)o n z=0.
Determine φ(x, y, z)i f
f(x, y)=
/(
δ(y)f o r|x|<a ,
0f o r|x|≥a.
19.29 (a) By applying the divergence theorem to the volume integralZ
V
/
φ(∇2−m2)ψ−ψ(∇2−m2)φ
/
dV
obtain a Green’s function expression, as the sum of a volume integral and a
surface integral, for φ(r/prime)t h a ts a t i s fi e s
∇2φ−m2φ=ρ
inVand takes the specified form φ=fonS, the boundary of V.T h e
Green’s function G(r,r/prime)t ob eu s e ds a t i s fi e s
∇2G−m2G=δ(r−r/prime)
and vanishes when ris on S.
(b) When Vis all space, G(r,r/prime) can be written as G(t)=g(t)/twhere t=|r−r/prime|
andg(t) is bounded as t→∞. Find the form of G(t).
(c) Find φ(r) in the half space x>0i fρ(r)=δ(r−r1)a n d φ= 0 both on x=0
and as r→∞.
707
PDES: SEPARATION OF VARIABLES AND OTHER METHODS
19.30 Consider the PDE Lu(r)=ρ(r), for which the differential operator Lis given by
L=∇·[p(r)∇]+q(r),
where p(r)a n d q(r) are functions of position. By proving the generalised form of
Green’s theorem,Z
V(φLψ−ψLφ)dV=
I
Sp(φ∇ψ−ψ∇φ)·ˆndS,
show that the solution of the PDE is given by
u(r0)=
Z
VG(r,r0)ρ(r)dV(r)+
I
Sp(r)
/
u(r)∂G(r,r0)
∂n−G(r,r0)∂u(r)
∂n
/
dS(r),
where G(r,r0) is the Green’s function satisfying LG(r,r0)=δ(r−r0).
19.7 Hints and answers
19.1 (a) Cexp[λ(x2+2y)]; (b) C(x2y)λ.
19.2 There is heat flow only across z=±a. It is into the cube at a rate of κAe−2/√2.
19.3 u(x, y, t)=s i n ( nπx/a )sin(mπy/b )(Asinωt+Bcosωt).
19.4 (a) −
/~2
2mX/prime/prime
X=p2
x
2m,etc.,i /~T/prime
T=E;
(b) As in (a), but with solutions X=Asin(pxx/ /~), etc. with pxa/ /~=nxπ.
19.5 (a) 6 u/r2,−6u/r2,0 ,/lscript=2 , m=0 ;
(b) 2u/r2,( c o t2θ−1)u/r2;−u/(r2sin2θ),/lscript=1 , m=1 .
19.8 The first term can contain only /lscript=1,2a n d m=±1, the second only /lscript=0,1,2
andm=0 ; f(θ,φ)=(π)1/2[Y0
0−3−1/2Y0
1−(2/3)1/2Y1
1−(2/15)1/2Y−1
2].
19.9 Solutions of the form r/lscriptgive/lscriptas−1,1,2,4. Because of the asymptotic form of
ψ,a nr4term cannot be present. The coefficients of the three remaining terms are
determined by the two boundary conditions u=0on the sphere and the form of
ψfor large r.
19.10 If ψ(ρ, φ)=R(ρ)Φ(φ), show that Φ(4)+4 Φ/prime/prime= 0 and hence that Φ = A+Bφ+
Ccos 2φ+Dsin 2φ.
19.11 Express cos2φin terms of cos 2 φ;T(ρ, φ)=A+B/2+(Bρ2/2a2)c os2 φ.
19.12 (a) u(ρ, φ)=( 4 /π)
P
noddn−1(ρ/a)nsinnφ.
(b)∇2α=0 ,a n d∇2β= 0 separately. On ρ=a, α+β+π/2=π.
(c) Equate the two forms (uniqueness theorem) and then set ρ=a.
The Fourier series is 2
P
noddn−1sinnφ.
19.13 ( Acosmx+Bsinmx+Ccoshmx+Dsinhmx)cos( ωt+/epsilon1), with m4a4=ω2.
19.15 En=1 6ρA2c2/[(2n+1 )2π2L];E=2ρc2A2/L=
RA
0[2Tv/(1
2L)]dv.
19.17 You will need the result from exercise 17.7.19.18 Write C(x, t)=C
0+
P∞
1Ansin(nπx/L )fn(t)w h e r e fn(t)→0a st→∞;
An=−4C0/(nπ)a n d fn(t)=e x p [−(κn2π2/L2)t]f o r nodd, and An=0f o r neven.
19.19 Since there is no heat flow at x=±a, use a series of period 4 a,u(x,0) = 100 for
0<x≤2a,u(x,0) = 0 for−2a≤x<0.
u(x, t)=5 0+200
π∞X
n=01
2n+1sin
/(2n+1 )πx
2a
/
exp
/
−k(2n+1 )2π2t
4a2s
/
.
Taking only the n= 0 term gives t≈2300 s.
19.20 T1(r,θ)=
P∞
1bm(r/a)mPm(cosθ)+q0/k2+T∞,
T2(r,θ)=
P∞
1bm(a/r)m+1Pm(cosθ)+aq0/(k2r)+T∞,
where in both cases bm=qm/[mk1+(m+1 )k2];T(0,θ)=q0/k2+T∞.
708
19.7 HINTS AND ANSWERS
19.21 u(x, t)=[a/(a2+4κt)1/2]e x p [−x2/(a2+4κt)].
19.22 (a) u(r=z,0) = 2 MGa−2[(a2+z2)1/2−z]. (b) For θ>π / 2, the factor in the
square brackets is ( a2+z2)1/2+z.( c )F i n d ∂u/∂r atθ=0f o r r<a,a n dl e t
a→∞.
19.23 Fourier-transform Poisson’s equation to show that ˜ρ(α)=/epsilon10(α2+q2)˜V(α).
19.24 (i) q2ab/[4π/epsilon10(b2−a2)2]; (ii) [ q2ab/(4π/epsilon10)][(b2−a2)−2−b−4]. Obtain (ii) from (i)
by adding a further image charge + qa/batO, to give a net zero electrostatic
flux from the sphere while maintaining its equipotential property.
19.25 Follow the worked example that includes result (19.98). For part of the explicit
integration, substitute ρ=ztanα.
Φ(0,0,z)=z(1 +z2)1/2−z2+( 1+ z2)1/2−1
z(1 +z2)1/2.
19.26 (a) See equation (19.94); (c) M=( 4 π/epsilon10)−1
R
ρ(r/prime)dV/prime= total charge on the
sphere. d=( 4π/epsilon10)−1
R
ρ(r/prime)r/primedV/prime= dipole moment of the sphere.
19.27
G(r,r0)=1
4π∞X
n=2(−1)n
/"
1p
(x−x0)2+(y−y0)2+(z+(−1)nz0−nc)2
+1p
(x−x0)2+(y−y0)2+(z+(−1)nz0+nc)2
/#
.
19.28
G(r,r0)=−1
4π
/"
1p
(x−x0)2+(y−y0)2+(z−z0)2
+1p
(x−x0)2+(y−y0)2+(z+z0)2
/#
.
φ(x, y, z)=1
2π
/
sinh−1a+xp
y2+z2+s i n h−1a−xp
y2+z2
/!
.
19.29 (a) As given in equation (19.89), but with r0replaced by r/prime.
(b) Move the origin to r/primeand integrate the defining Green’s equation to obtain
4πt2dG
dt−m2
Zt
0G(t/prime)4πt/prime2dt/prime=1,
leading to G(t)=[−1/(4πt)]e−mt.
(c)φ(r)=[−1/(4π)](p−1e−mp−q−1e−mq), where p=|r−r1|andq=|r−r2|with
r1=(x1,y1,z1)a n d r2=(−x1,y1,z1).
709
20
Complex variables
Throughout this book references have been made to results derived from the
theory of complex variables. This theory thus becomes an integral part of themathematics appropriate to physical applications. The difficulty with it, from thepoint of view of a book such as the present one, is that although it has many
practical applications its underlying basis has a distinctly pure mathematics
flavour.
Thus, to adopt a comprehensive rigorous approach would involve a large
amount of groundwork in analysis, for example formulating precise definitionsof continuity and differentiability, developing the theory of sets and making adetailed study of boundedness. Instead, we will be selective and pursue only thoseparts of the formal theory that are needed to establish the results used elsewherein this book and some others of general utility.
In this spirit, the proofs that have been adopted for some of the standard
results of complex variable theory have been chosen with an eye to simplicity
rather than sophistication. This means that in some cases the imposed conditionsare more stringent than would be strictly necessary if more sophisticated proofswere used; where this happens the less restrictive results are usually stated aswell. The reader who is interested in a fuller treatment should consult one of themany excellent textbooks on this fascinating subject. †
One further concession to ‘hand-waving’ has been made in the interests of
keeping the treatment to a moderate length. In several places phrases such as ‘can
be made as small as we like’ are used, rather than a careful treatment in terms
of ‘given /epsilon1>0, there exists a δ>0 such that’. In the authors’ experience, some
students are more at ease with the former type of statement despite its lack of
†For example, Knopp, Theory of Functions, Part I (Dover, 1945); Phillips, Functions of a Complex
Variable (Oliver and Boyd, 1954); Titchmarsh, The Theory of Functions (Oxford, 1952).
710
20.1 FUNCTIONS OF A COMPLEX VARIABLE
precision whilst others, those who would contemplate only the latter, are usually
well able to supply it for themselves.
20.1 Functions of a complex variable
The quantity f(z) is said to be a function of the complex variable zif to every value
ofzin a certain domain R(a region of the Argand diagram) there corresponds
one or more values of f(z). Stated like this f(z) could be any function consisting
of a real and an imaginary part, each of which is, in general, itself a function of x
andy. If we denote the real and imaginary parts of f(z)b yuandvrespectively,
then
f(z)=u(x, y)+iv(x, y).
In this chapter, however, we will be primarily concerned with functions that are
single-valued, so that for each value of zthere corresponds just one value of f(z),
and differentiable in a particular sense, which we now discuss.
A function f(z) that is single-valued in some domain Risdifferentiable at the
point zinRif the derivative
f/prime(z) = lim
∆z→0bracketleftbiggf(z+∆z)−f(z)
∆zbracketrightbigg
(20.1)
exists and is unique, in that its value does not depend upon the direction in the
Argand diagram from which ∆ ztends to zero.IShow that the function f(z)=x2−y2+i2xyis differentiable for all values of z.
Considering the definition (20.1), and taking ∆ z=∆x+i∆y, we have
f(z+∆z)−f(z)
∆z
=(x+∆x)2−(y+∆y)2+2i(x+∆x)(y+∆y)−x2+y2−2ixy
∆x+i∆y
=2x∆x+( ∆x)2−2y∆y−(∆y)2+2i(x∆y+y∆x+∆x∆y)
∆x+i∆y
=2x+i2y+(∆x)2−(∆y)2+2i∆x∆y
∆x+i∆y.
Now, in whatever way ∆ xand ∆ yare allowed to tend to zero (e.g. taking ∆ y=0a n d
letting ∆ x→0 or vice versa), the last term on the right will tend to zero and the unique
limit 2 x+i2ywill be obtained. Since zwas arbitrary, f(z)w i t h u=x2−y2andv=2xy
is differentiable at all points in the (finite) complex plane.
J
We note that the above working can be considerably reduced by recognising
that, since z=x+iy, we can write f(z)a s
f(z)=x2−y2+2ixy=(x+iy)2=z2.
711
COMPLEX VARIABLES
We then find that
f/prime(z) = lim
∆z→0bracketleftbigg(z+∆z)2−z2
∆zbracketrightbigg
= lim
∆z→0bracketleftbigg(∆z)2+2z∆z
∆zbracketrightbigg
=parenleftBig
lim
∆z→0∆zparenrightBig
+2z=2z,
from which we see immediately that the limit both exists and is independent of
the way in which ∆ z→0. Thus we have verified that f(z)=z2is differentiable
for all (finite) z. We also note that the derivative is analogous to that found for
real variables.
Although the definition of a differentiable function clearly includes a wide
class of functions, the concept of differentiability is restrictive and, indeed, some
functions are not differentiable at any point in the complex plane.IShow that the function f(z)=2 y+ixis not differentiable anywhere in the complex plane.
In this case f(z) cannot be written simply in terms of z, and so we must consider the
limit (20.1) in terms of xandyexplicitly. Following the same procedure as in the previous
example we find
f(z+∆z)−f(z)
∆z=2y+2 ∆ y+ix+i∆x−2y−ix
∆x+i∆y
=2∆y+i∆x
∆x+i∆y.
In this case the limit will clearly depend on the direction from which ∆ z→0. Suppose
∆z→0 along a line through zof slope m,s ot h a t∆ y=m∆x,t h e n
lim
∆z→0
/f(z+∆z)−f(z)
∆z
/
= lim
∆x,∆y→0
/2∆y+i∆x
∆x+i∆y
/
=2m+i
1+im.
This limit is dependent on mand hence on the direction from which ∆ z→0. Since this
conclusion is independent of the value of z, and hence true for all z,f(z)=2 y+ixis
nowhere differentiable.
J
A function that is single-valued and differentiable at all points of a domain R
is said to be analytic (orregular )i nR. A function may be analytic in a domain
except at a finite number of points (or an infinite number if the domain is
infinite); in this case it is said to be analytic except at these points, which are
called the singularities off(z). (In our treatment we will not consider cases in
which an infinite number of singularities occur in a finite domain.)
712
20.2 THE CAUCHY–RIEMANN RELATIONSIShow that the function f(z)=1 /(1−z)is analytic everywhere except at z=1.
Since f(z) is given explicitly as a function of z, evaluation of the limit (20.1) is somewhat
easier. We find
f/prime(z) = lim
∆z→0
/f(z+∆z)−f(z)
∆z
/
= lim
∆z→0
/1
∆z
/1
1−z−∆z−1
1−z
//
= lim
∆z→0
/1
(1−z−∆z)(1−z)
/
=1
(1−z)2,
independently of the way in which ∆ z→0, provided z/negationslash= 1. Hence f(z)i sa n a l y t i c
everywhere except at the singularity z=1 .
J
20.2 The Cauchy–Riemann relations
From examining the previous examples, it is apparent that for a function f(z)
to be differentiable and hence analytic there must be some particular connectionbetween its real and imaginary parts uand v. We next establish what this
connection must be, by considering a general function.
If the limit
L= lim
∆z→0bracketleftbiggf(z+∆z)−f(z)
∆zbracketrightbigg
(20.2)
is to exist and be unique, in the way required for differentiability, then any two
specific ways of letting ∆ z→0 must produce the same limit. In particular, moving
parallel to the real axis and moving parallel to the imaginary axis must do so.
This is certainly a necessary condition, although it may not be sufficient.
If we let f(z)=u(x, y)+iv(x, y)a n d∆ z=∆x+i∆ythen we have
f(z+∆z)=u(x+∆x, y+∆y)+iv(x+∆x, y+∆y),
and the limit (20.2) is given by
L= lim
∆x,∆y→0bracketleftbiggu(x+∆x, y+∆y)+iv(x+∆x, y+∆y)−u(x, y)−iv(x, y)
∆x+i∆ybracketrightbigg
.
If we first suppose that ∆ zis purely real, so that ∆ y=0 ,w eo b t a i n
L= lim
∆x→0bracketleftbiggu(x+∆x, y)−u(x, y)
∆x+iv(x+∆x, y)−v(x, y)
∆xbracketrightbigg
=∂u
∂x+i∂v
∂x,
(20.3)
provided each limit exists at the point z. Similarly, if ∆ zis taken as purely
imaginary, so that ∆ x= 0, we find
L= lim
∆y→0bracketleftbiggu(x, y+∆y)−u(x, y)
i∆y+iv(x, y+∆y)−v(x, y)
i∆ybracketrightbigg
=1
i∂u
∂y+∂v
∂y.
(20.4)
713
COMPLEX VARIABLES
Forfto be differentiable at the point z, expressions (20.3) and (20.4) must
be identical. It follows from equating real and imaginary parts that necessary
conditions for this are
∂u
∂x=∂v
∂yand∂v
∂x=−∂u
∂y. (20.5)
These two equations are known as the Cauchy–Riemann relations .
We can now see why for the earlier examples (i) f(z)=x2−y2+i2xymight
be differentiable and (ii) f(z)=2 y+ixcould not be.
(i)u=x2−y2,v=2xy:
∂u
∂x=2x=∂v
∂yand∂v
∂x=2y=−∂u
∂y,
(ii)u=2y,v=x:
∂u
∂x=0=∂v
∂ybut∂v
∂x=1/negationslash=−2=−∂u
∂y.
It is apparent that for f(z) to be analytic something more than the existence
of the partial derivatives of uandvwith respect to xandyis required; this
something is that they satisfy the Cauchy–Riemann relations.
We may enquire also as to the sufficient conditions for f(z) to be analytic in
R.I tc a nb es h o w n †that a sufficient condition is that the four partial derivatives
exist, are continuous and satisfy the Cauchy–Riemann relations. It is the addi-
tional requirement of continuity that makes the difference between the necessaryconditions and the sufficient conditions.IIn which domain(s) of the complex plane is f(z)=|x|−i|y|an analytic function?
Writing f=u+ivit is clear that both ∂u/∂y and∂v/∂x a r ez e r oi na l lf o u rq u a d r a n t s
and hence that the second Cauchy–Riemann relation in (20.5) is satisfied everywhere.
Turning to the first Cauchy–Riemann relation, in the first quadrant ( x>0,y>0) we
have f(z)=x−iyso that
∂u
∂x=1,∂v
∂y=−1,
which clearly violates the first relation in (20.5). Thus f(z) is not analytic in the first
quadrant.
Following a similiar argument for the other quadrants, we find
∂u
∂x=−1o r + 1 f o r x<0a n d x>0 respectively,
∂v
∂y=−1o r + 1 f o r y>0a n d y<0 respectively.
Therefore ∂u/∂x and∂v/∂y are equal, and hence f(z) is analytic, only in the second and
fourth quadrants.
J
†See for example any of the references given earlier.
714
20.2 THE CAUCHY–RIEMANN RELATIONS
Since xandyare related to zand its complex conjugate z∗by
x=1
2(z+z∗)a n d y=1
2i(z−z∗), (20.6)
we may formally regard any function f=u+ivas a function of zandz∗,r a t h e r
than xandy. If we do this and examine ∂f/∂z∗we obtain
∂f
∂z∗=∂f
∂x∂x
∂z∗+∂f
∂y∂y
∂z∗
=parenleftbigg∂u
∂x+i∂v
∂xparenrightbiggparenleftbigg1
2parenrightbigg
+parenleftbigg∂u
∂y+i∂v
∂yparenrightbiggparenleftbigg
−1
2iparenrightbigg
=1
2parenleftbigg∂u
∂x−∂v
∂yparenrightbigg
+i
2parenleftbigg∂v
∂x+∂u
∂yparenrightbigg
. (20.7)
Now, if fis analytic then the Cauchy–Riemann relations (20.5) must be satisfied,
and these immediately give that ∂f/∂z∗is identically zero. Thus we conclude that
iffis analytic then fcannot be a function of z∗and any expression representing
an analytic function of zcan contain xandyonly in the combination x+iy,not
in the combination x−iy.
We conclude this section by discussing some properties of analytic functions
that are of great practical importance in theoretical physics. These can be obtainedsimply from the requirement that the Cauchy–Riemann relations must be satisfiedby the real and imaginary parts of an analytic function.
The most important of these results can be obtained by differentiating the
first Cauchy–Riemann relation with respect to one independent variable, and thesecond with respect to the other independent variable, to obtain the two chains
of equalities:
∂
∂xparenleftbigg∂u
∂xparenrightbigg
=∂
∂xparenleftbigg∂v
∂yparenrightbigg
=∂
∂yparenleftbigg∂v
∂xparenrightbigg
=−∂
∂yparenleftbigg∂u
∂yparenrightbigg
;
∂
∂xparenleftbigg∂v
∂xparenrightbigg
=−∂
∂xparenleftbigg∂u
∂yparenrightbigg
=−∂
∂yparenleftbigg∂u
∂xparenrightbigg
=−∂
∂yparenleftbigg∂v
∂yparenrightbigg
.
Thus both uandvareseparately solutions of Laplace’s equation in two dimen-
sions, i.e.
∂2u
∂x2+∂2u
∂y2=0 a n d∂2v
∂x2+∂2v
∂y2=0. (20.8)
We shall make use of this result in section 20.9.
A further useful result concerns the two families of curves u(x, y) = constant
andv(x, y) = constant, where uandvare the real and imaginary parts of any
analytic function f=u+iv. As discussed in chapter 10, the vector normal to the
curve u(x, y) = constant is given by
∇u=∂u
∂xi+∂u
∂yj, (20.9)
715
COMPLEX VARIABLES
where iandjare the unit vectors along the x-a n d y- axes respectively. A similar
expression exists for ∇v, the normal to the curve v(x, y) = constant. Taking the
scalar product of these two normal vectors we obtain
∇u·∇v=∂u
∂x∂v
∂x+∂u
∂y∂v
∂y
=−∂u
∂x∂u
∂y+∂u
∂y∂u
∂x=0,
where in the last line we have used the Cauchy–Riemann relations to rewrite the
partial derivatives of vas partial derivatives of u. Since the scalar product of the
normal vectors is zero, they must be orthogonal and the curves u(x, y) = constant
andv(x, y) = constant must therefore intersect at right angles .IUse the Cauchy–Riemann relations to show that, for any analytic function f=u+iv,t h e
relation|∇u|=|∇v|must hold.
From (20.9) we have
|∇u|2=∇u·∇u=
/∂u
∂x
/2
+
/∂u
∂y
/2
.
Using the Cauchy–Riemann relations to write the partial derivatives of uin terms of those
ofv,w eo b t a i n
|∇u|2=
/∂v
∂y
/2
+
/∂v
∂x
/2
=|∇v|2,
from which the result |∇u|=|∇v|follows immediately.
J
20.3 Power series in a complex variable
The theory of power series in a real variable was considered in chapter 4, which
also contained a brief discussion of the natural extension of this theory to a seriessuch as
f(z)=∞summationdisplay
n=0anzn, (20.10)
where zis a complex variable and the anare in general complex. We now consider
complex power series in more detail.
Expression (20.10) is a power series about the origin and may be used for
general discussion since a power series about any other point z0can be obtained
by a change of variable from ztoz−z0.I fzwere written in its modulus and
argument form z=rexpiθ, expression (20.10) would become
f(z)=∞summationdisplay
n=0anrnexp(inθ). (20.11)
716
20.3 POWER SERIES IN A COMPLEX VARIABLE
This series is absolutely convergent if
∞summationdisplay
n=0|an|rn, (20.12)
which is a series of positive real terms, is convergent. Thus tests for the absolute
convergence of real series can be used in the present context, and of these themost appropriate form is based on the Cauchy root test. The radius of convergence
Ris defined by
1
R= lim
n→∞|an|1/n; (20.13)
the series (20.10) is absolutely convergent if |z|<Rand divergent if |z|>R.I f
|z|=Rno particular conclusion may be drawn, and this case must be considered
separately, as discussed in subsection 4.5.1.
A circle of radius Rcentred on the origin is called the circle of convergence
of the seriessummationtextanzn.T h ec a s e s R=0a n d R=∞correspond respectively to
convergence at the origin only and convergence everywhere. For Rfinite the
convergence occurs in a restricted part of the z-plane (the Argand diagram). For
a power series about a general point z0, the circle of convergence is of course
centred on that point.IFind the parts of the z-plane for which the following series are convergent:
(i)∞X
n=0zn
n!, (ii)∞X
n=0n!zn,(iii)∞X
n=1zn
n.
(i) Since ( n!)1/nbehaves like nasn→∞ we find lim(1 /n!)1/n= 0. Hence R=∞and the
series is convergent for all z. (ii) Correspondingly, lim( n!)1/n=∞. Thus R=0a n dt h e
series converges only at z= 0. (iii) As n→∞,(n)1/nhas a lower limit of 1 and hence
lim(1 /n)1/n=1/1 = 1. Thus the series is absolutely convergent if |z|<1.
J
Case (iii) in the above example provides a good illustration of the fact that on its
circle of convergence a power series may or may not converge. For this particular
series the circle of convergence is |z|= 1, so let us consider the convergence of
the series at two different points on this circle. Taking z= 1, the series becomes
∞summationdisplay
n=11
n=1+1
2+1
3+1
4+···,
which is easily shown to diverge (by, for example, grouping terms, as discussed in
subsection 4.3.2). Taking z=−1, however, the series is given by
∞summationdisplay
n=1(−1)n
n=−1+1
2−1
3+1
4−···,
717
COMPLEX VARIABLES
which is an alternating series whose terms decrease in magnitude and which
therefore converges.
The ratio test discussed in subsection 4.3.2 may also be employed to investi-
gate the absolute convergence of a complex power series. A series is absolutely
convergent if
lim
n→∞|an+1||z|n+1
|an||z|n= lim
n→∞|an+1||z|
|an|<1 (20.14)
and hence the radius of convergence Ro ft h es e r i e si sg i v e nb y
1
R= lim
n→∞|an+1|
|an|.
For instance, in case (i) of the previous example, we have
1
R= lim
n→∞n!
(n+1 ) != lim
n→∞1
n+1=0.
Thus the series is absolutely convergent for all (finite) z, confirming the previous
result.
Before turning to particular power series, we conclude this section by stating
the important result †thatthe power seriessummationtext∞
0anznhas a sum that is an analytic
function of zinside its circle of convergence.
As a corollary to the above theorem, it may further be shown that if f(z)=summationtextanznthen, inside the circle of convergence of the series,
f/prime(z)=∞summationdisplay
n=0nanzn−1.
Repeated application of this result demonstrates that any power series can be
differentiated any number of times inside its circle of convergence.
20.4 Some elementary functions
In the example at the end of the previous section it was shown that the function
expzdefined by
expz=∞summationdisplay
n=0zn
n!(20.15)
is convergent for all zof finite modulus and is thus, by the discussion of
the previous section, an analytic function over the whole z-plane.‡Like its
†For a proof see, for example, Riley, Mathematical Methods for the Physical Sciences (CUP, 1974),
p. 446.
‡Functions that are analytic in the whole z-plane are usually called integral orentire functions.
718
20.4 SOME ELEMENTARY FUNCTIONS
real-variable counterpart it is called the exponential function ; also like its real
counterpart it is equal to its own derivative.
The multiplication of two exponential functions results in a further exponential
function, in accordance with the corresponding result for real variables.IShow that expz1expz2=e x p ( z1+z2).
From the series expansion (20.15) of exp z1and a similar expansion for exp z2,i ti sc l e a r
that the coefficient of zr
1zs
2in the corresponding series expansion of exp z1expz2is simply
1/(r!s!).
But, from (20.15) we also have
exp(z1+z2)=∞X
n=0(z1+z2)n
n!.
In order to find the coefficient of zr
1zs
2in this expansion, we clearly have to consider the
term in which n=r+s,n a m e l y
(z1+z2)r+s
(r+s)!=1
(r+s)!
/;r+sC0zr+s
1+···+r+sCszr
1zs
2+···+r+sCr+szr+s
2
/
.
The coefficient of zr
1zs
2in this is given by
r+sCs1
(r+s)!=(r+s)!
s!r!1
(r+s)!=1
r!s!.
Thus, since the corresponding coefficients on the two sides are equal and all the series
involved are absolutely convergent for all z, we can conclude that exp z1expz2=e x p ( z1+
z2).
J
As an extension of (20.15) we may also define the complex exponent of a real
number a>0 by the equation
az=e x p ( zlna), (20.16)
where ln ais the natural logarithm of a. The particular case a=eand the fact
that ln e= 1 enable us to write exp zinterchangeably with ez.I fzis real then the
definition agrees with the familiar one.
The result for z=iy,
expiy=c o s y+isiny, (20.17)
has been met already in equation (3.23). Its immediate extension is
expz=( e x p x)(cos y+isiny). (20.18)
Aszvaries over the complex plane the modulus of exp ztakes all real positive
values, except that of 0. However, two values of zthat differ by 2 πni,f o ra n y
integer n, produce the same value of exp z, as given by (20.18), and so exp zis
periodic with period 2 πi. If we denote exp zbytthen the strip −π<y≤πin
thez-plane corresponds to the whole of the t-plane, except for the point t=0 .
719
COMPLEX VARIABLES
The sine, cosine, sinh and cosh functions of a complex variable are defined from
the exponential function exactly as are those for real variables. The functionsderived from them (e.g. tan and tanh), the identities they satisfy, and theirderivative properties, are also just as for real variables. In view of this we will
not give them further attention here.
The inverse function of exp zis given by w, the solution of
expw=z. (20.19)
This inverse function was discussed in chapter 3, but we mention it again here
for completeness. By virtue of the discussion following (20.18), wis not uniquely
defined and is indeterminate to the extent of any integer multiple of 2 πi.I fw e
express zas
z=rexpiθ,
where ris the (real) modulus of z,a n d θis its argument ( −π<θ≤π), then
multiplying zby exp(2 inπ), where nis an integer, will result in the same complex
number z. Thus we may write
z=rexp[i(θ+2nπ)],
where nis an integer. If we denote win (20.19) by
w=L n z=l nr+i(θ+2nπ), (20.20)
where ln ris the natural logarithm (to base e) of the real positive quantity r,t h e n
Lnzis an infinitely multivalued function of z.I t sprincipal value , denoted by ln z,
is obtained by taking n= 0 so that its argument lies in the range −πtoπ. Thus
lnz=l nr+iθ, with−π<θ≤π. (20.21)
Now that the logarithm of a complex variable has been defined, definition
(20.16) of a general power can be extended to cases other than those in which a
is real and positive. If t(/negationslash=0 )a n d zare both complex then the zth power of tis
defined by
t
z=e x p ( zLnt). (20.22)
Since Ln tis multivalued, so is this definition. Its principal value is obtained by
giving Ln tits principal value, ln t.
Ift(/negationslash= 0) is complex but zis real and equal to 1 /n, then (20.22) provides a
definition of the nth root of t. Because of the multivaluedness of Ln t, there will
be more than one nth root of any given t.
720
20.5 MULTIVALUED FUNCTIONS AND BRANCH CUTSIShow that there are exactly ndistinct nth roots of t.
From (20.22) the nth roots of tare given by
t1/n=e x p
/1
nLnt
/
.
On the RHS let us write tas follows:
t=rexp[i(θ+2kπ)],
where kis an integer. We then obtain
t1/n=e x p
/1
nlnr+i(θ+2kπ)
n
/
=r1/nexp
/
i(θ+2kπ)
n
/
,
where k=0,1,...,n−1; for other values of kwe simply recover the roots already found.
Thus thasndistinct nth roots.
J
20.5 Multivalued functions and branch cuts
In the definition of an analytic function, one of the conditions imposed was
that the function is single-valued. However, as shown in the previous section, thelogarithmic function, a complex power and a complex root are all multivalued.
Nevertheless, it happens that the properties of analytic functions can still be
applied to these and other multivalued functions of a complex variable providedthat suitable care is taken. This care amounts to identifying the branch points of
the multivalued function f(z) in question. If zis varied in such a way that its
path in the Argand diagram forms a closed curve that encloses a branch point,then, in general, f(z) will not return to its original value.
For definiteness let us consider the multivalued function f(z)=z
1/2and express
zasz=rexpiθ. From figure 20.1( a), it is clear that, as the point ztraverses any
closed contour Cthat does not enclose the origin, θwill return to its original
value after one complete circuit. However, for any closed contour C/primethat does
enclose the origin, after one circuit θ→θ+2π(see figure 20.1( b)). Thus, for the
function f(z)=z1/2, after one circuit
r1/2exp(iθ/2)→r1/2exp[i(θ+2π)/2] =−r1/2exp(iθ/2).
In other words, the value of f(z) changes around any closed loop enclosing the
origin; in this case f(z)→− f(z). Thus z= 0 is a branch point of the function
f(z)=z1/2.
We note in this case that if any closed contour enclosing the origin is traversed
twicethen f(z)=z1/2returns to its original value. The number of loops around
a branch point required for any given function f(z) to return to its original value
721
COMPLEX VARIABLES
y y y
x x xC
θ θr r
(a)( b)( c)C/prime
Figure 20.1 ( a) A closed contour not enclosing the origin; ( b) a closed
contour enclosing the origin; ( c) a possible branch cut for f(z)=z1/2.
depends on the function in question, and for some functions (e.g. Ln z,w h i c ha l s o
has a branch point at the origin) the original value is never recovered.
In order that f(z) may be treated as single-valued we may define a branch cut
in the Argand diagram. A branch cut is a line (or curve) in the complex plane
and may be regarded as an artificial barrier that we must not cross. Branch cuts
are positioned in such a way that we are prevented from making a completecircuit around any one branch point, and so the function in question remainssingle-valued.
For the function f(z)=z
1/2, we may take as a branch cut any curve starting
at the origin z= 0 and extending out to |z|=∞in any direction, since all
such curves would equally well prevent us from making a closed loop around the
branch point at the origin. It is usual, however, to take the cut along the real or
imaginary axis. For example, in figure 20.1( c), we take the cut as the positive real
axis. By agreeing not to cross this cut, we restrict θto lie in the range 0 ≤θ<2π,
a n ds ok e e p f(z) single-valued.
These ideas are easily extended to functions with more than one branch point.IFind the branch points of f(z)=√
z2+1, and hence sketch suitable arrangements of
branch cuts.
We begin by writing f(z)a s
f(z)=
p
z2+1=
p
(z−i)(z+i).
As shown above the function g(z)=z1/2has a branch point at z= 0. Thus we might
expect f(z) to have branch points at values of zthat make the expression under the square
root equal to zero, i.e. at z=iandz=−i.
As shown in figure 20.2( a), we use the notation
z−i=r1expiθ1 and z+i=r2expiθ2.
722
20.6 SINGULARITIES AND ZEROES OF COMPLEX FUNCTIONS
(a)( b)( c)−i −i −ii i iy y y
x x xr1
r2θ1
θ2z
Figure 20.2 ( a) Coordinates used in the analysis of the branch points of
f(z)=( z2+1 )1/2;(b) one possible arrangement of branch cuts; ( c) another
possible branch cut, which is finite.
We can therefore write f(z)a s
f(z)=√r1r2exp(iθ1/2)exp( iθ2/2) =√r1r2exp
/
i(θ1+θ2)/2
/
.
Let us now consider how f(z) changes as we make one complete circuit around various
closed loops Cin the Argand diagram. If Cencloses
(i) neither branch point, then θ1→θ1,θ2→θ2and so f(z)→f(z);
(ii)z=ibut not z=−i,t h e n θ1→θ1+2π,θ2→θ2and so f(z)→−f(z);
(iii)z=−ibut not z=i,t h e n θ1→θ1,θ2→θ2+2πand so f(z)→−f(z);
(iv) both branch points, then θ1→θ1+2π,θ2→θ2+2πand so f(z)→f(z).
Thus, as expected, f(z) changes value around loops containing either z=iorz=−i
(but not both). We must therefore choose branch cuts that prevent us from making acomplete loop around either branch point; one suitable choice is shown in figure 20.2( b).
For this f(z), however, we have noted that after traversing a loop containing bothbranch
points the function returns to its original value. Thus we may choose an alternative, finite,
branch cut that allows this possibility but still prevents us from making a complete looparound just one of the points. A suitable cut is shown in figure 20.2( c).J
20.6 Singularities and zeroes of complex functions
A singular point of a complex function f(z) is any point in the Argand diagram
at which f(z) fails to be analytic. We have already met one sort of singularity,
the branch point, and in this section we shall consider other types of singularityas well as discuss the zeroes of complex functions.
Iff(z) has a singular point at z=z
0but is analytic at all points in some
neighbourhood containing z0but no other singularities then z=z0is called an
isolated singularity . (Clearly branch points are not isolated singularities.)
723
COMPLEX VARIABLES
The most important type of isolated singularity is the pole.I ff(z)h a st h ef o r m
f(z)=g(z)
(z−z0)n, (20.23)
where nis a positive integer, g(z) is analytic at all points in some neighbourhood
containing z=z0andg(z0)/negationslash=0 ,t h e n f(z)h a sa pole of order natz=z0.A n
alternative (though equivalent) definition is that
lim
z→z0[(z−z0)nf(z)]=a, (20.24)
where ais a finite, non-zero complex number. (If the limit equals zero then z=z0
is a pole of order less than n,o rf(z) is analytic there; if the limit is infinite then
the pole is of order greater than n.) It may also be shown that if f(z) has a pole
atz=z0,t h e n|f(z)|→∞ asz→z0from any direction in the Argand diagram. †
If no finite value of ncan be found such that (20.24) is satisfied then z=z0is
called an essential singularity .IFind the singularities of the functions
(i)f(z)=1
1−z−1
1+z, (ii)f(z)=t a n h z.
(i) If we write f(z)a s
f(z)=1
1−z−1
1+z=2z
(1−z)(1 + z),
we see immediately from either (20.23) or (20.24) that f(z) has poles of order 1 (or simple
poles)a tz=1a n d z=−1. (ii) In this case we write
f(z)=t a n h z=sinhz
coshz=expz−exp(−z)
expz+e x p (−z).
Thus f(z) has a singularity when exp z=−exp(−z) or, equivalently, when
expz=e x p [ i(2n+1 )π]e x p (−z),
where nis any integer. Equating the arguments of the exponentials we find z=(n+1
2)πi,
for integer n.
Furthermore, using l’H ˆopital’s rule (see chapter 4) we have
lim
z→(n+1
2)πi
/(
[z−(n+1
2)πi]si nh z
coshz
/)
= lim
z→(n+1
2)πi
/(
[z−(n+1
2)πi]cosh z+s i n h z
sinhz
/)
=1.
Therefore, from (20.24), each singularity is a simple pole.
J
Another type of singularity exists at points for which the value of f(z)t a k e s
an indeterminate form such as 0 /0 but lim z→z0f(z) exists and is independent
†Although perhaps intuitively obvious this result really requires formal demonstration by analysis.
724
20.7 COMPLEX POTENTIALS
of the direction from which z0is approached. Such points are called removable
singularities .IShow that f(z)=( s i n z)/zhas a removable singularity at z=0.
It is clear that f(z) takes the indeterminate form 0 /0a t z= 0. However, by expanding
sinzas a power series in z, we find
f(z)=1
z
/
z−z3
3!+z5
5!−···
/
=1−z2
3!+z4
5!−···.
Thus lim z→0f(z) = 1 independently of the way in which z→0, and so f(z) has a removable
singularity at z=0 .
J
An expression common in mathematics, but which we have so far avoided
using explicitly in this chapter, is ‘ ztends to infinity’. For a real variable such
as|z|orR, ‘tending to infinity’ has a reasonably well-defined meaning. For a
complex variable needing a two-dimensional plane to represent it, the meaning is
not intrinsically well defined. However, it is convenient to have a unique meaningand this is provided by the following definition : the behaviour of f(z)at infinity
is given by that of f(1/ξ)a tξ=0 ,w h e r e ξ=1/z.IFind the behaviour at infinity of (i) f(z)=a+bz−2, (ii) f(z)=z(1 + z2)and (iii)
f(z)=e x p z.
(i)f(z)=a+bz−2: on putting z=1/ξ,f(1/ξ)=a+bξ2, which is analytic at ξ= 0; thus f
is analytic at z=∞. (ii)f(z)=z(1 +z2):f(1/ξ)=1 /ξ+1/ξ3; thus fhas a pole of order
3a tz=∞. (iii) f(z)=e x p z:f(1/ξ)=
P∞
0(n!)−1ξ−n; thus fhas an essential singularity
atz=∞.
J
We conclude this section by briefly mentioning the zeroes of a complex function.
As the name suggests, if f(z0)=0t h e n z=z0is called a zero of the function
f(z). Zeroes are classified in a similar way to poles, in that if
f(z)=(z−z0)ng(z),
where nis a positive integer and g(z0)/negationslash=0 ,t h e n z=z0is called a zero of order
noff(z). Ifn=1t h e n z=z0is called a simple zero . It may further be shown
that if z=z0is a zero of order noff(z) then it is also a pole of order nof the
function 1 /f(z).
We will return in section 20.13 to the classification of zeroes and poles in terms
of their series expansions.
20.7 Complex potentials
Towards the end of section 20.2 it was shown that the real and the imaginary
parts of an analytic function of zare separately solutions of Laplace’s equation
in two dimensions. Analytic functions thus offer a possible way of solving some
725
COMPLEX VARIABLES
y
x
Figure 20.3 The equipotentials (broken) and field lines (solid) for a line
charge perpendicular to the z-plane.
two-dimensional physical problems describable by a potential satisfying ∇2φ=0 .
The general method is known as that of complex potentials .
We found also that if f=u+ivis an analytic function of zthen any curve
u= constant intersects any curve v= constant at right angles. In the context of
solutions of Laplace’s equation, this result implies that the real and imaginary
parts of f(z) have an additional connection between them, for if the set of
contours on which one of them is a constant represents the equipotentials of asystem then the contours on which the other is constant, being orthogonal toeach of the first set, must represent the corresponding field lines or stream lines,
depending on the context. The analytic function fis the complex potential. It
is conventional to use φandψ(rather than uandv) to denote the real and
imaginary parts of a complex potential, so that f=φ+iψ.
As an example consider the function
f(z)=−q
2π/epsilon10lnz, (20.25)
in connection with the physical situation of a line charge of strength qper unit
length passing through the origin, perpendicular to the z-plane (figure 20.3). Its
real and imaginary parts are
φ=−q
2π/epsilon10ln|z|,ψ =−q
2π/epsilon10argz. (20.26)
The contours in the z-plane of φ= constant are concentric circles and of ψ=
constant are radial lines. As expected these are orthogonal sets, but in additionthey are respectively the equipotentials and electric field lines appropriate to the
726
20.7 COMPLEX POTENTIALS
field produced by the line charge (the minus sign is needed in (20.25) because the
value of φmust decrease with increasing distance from the origin).
Suppose we make the choice that the real part φof the analytic function f
gives the conventional potential function; ψcould equally well be selected. Then
we may consider how the direction and magnitude of the field are related to f.IShow that for any complex (electrostatic) potential f(z)the strength of the electric field
is given by E=|f/prime(z)|and that its direction makes an angle of π−arg[f/prime(z)]with the
x-axis.
Because φ= constant is an equipotential, the field has components
Ex=−∂φ
∂xand Ey=−∂φ
∂y. (20.27)
Since fis analytic, (i) we may use the Cauchy–Riemann relations (20.5) to change the
second of these, obtaining
Ex=−∂φ
∂xand Ey=∂ψ
∂x; (20.28)
(ii) the direction of differentiation at a point is immaterial and so
df
dz=∂f
∂x=∂φ
∂x+i∂ψ
∂x=−Ex+iEy. (20.29)
From these it can be seen that the field at a point is given in magnitude by E=|f/prime(z)|
and that it makes an angle with the x-axis given by π−arg[f/prime(z)].
J
It will be apparent from the above that much of physical interest can be
calculated by working directly in terms of fandz. In particular, the electric field
vector Emay be represented, using (20.29) above, by the quantity
E=Ex+iEy=−[f/prime(z)]∗.
Complex potentials can be used in two-dimensional fluid mechanics problems
in a similar way. If the flow is stationary (i.e. the velocity of the fluid does notdepend on time) and irrotational, and the fluid is both incompressible and non-
viscous, then the velocity of the fluid can be described by V=∇φ,w h e r e φis the
velocity potential and satisfies ∇
2φ= 0. If, for a complex potential f=φ+iψ,
the real part φis taken to represent the velocity potential then the curves ψ=
constant will be the streamlines of the flow. In a direct parallel with the electricfield, the velocity may be represented in terms of the complex potential by
V=V
x+iVy=[f/prime(z)]∗,
the difference of a minus sign reflecting the same difference between the definitions
ofEandV. The speed of the flow is equal to |f/prime(z)|. Points where f/prime(z) = 0, and
so the velocity is zero, are called stagnation points of the flow.
Analogously to the electrostatic case, a line source of fluid at z=z0, perpendic-
ular to the z-plane (i.e. a point from which fluid is emerging at a constant rate)
727
COMPLEX VARIABLES
is described by the complex potential
f(z)=kln(z−z0),
where kis the strength of the source. A sink is similarly represented, but with k
replaced by −k. Other simple examples are as follows.
(i) The flow of a fluid at a constant speed V0and at an angle αto the x-axis
is described by f(z)=V0(expiα)z.
(ii) Vortex flow, in which fluid flows azimuthally in an anticlockwise direction
around some point z0, the speed of the flow being inversely proportional
to the distance from z0, is described by f(z)=−ikln(z−z0), where kis
the strength of the vortex. For a clockwise vortex kis replaced by −k.IVerify that the complex potential
f(z)=V0
/
z+a2
z
/
,
is appropriate to a circular cylinder of radius aplaced so that it is perpendicular to a
uniform fluid flow of speed V0parallel to the x-axis.
Firstly, since f(z)i sa n a l y t i ce x c e p ta t z= 0, both its real and imaginary parts satisfy
Laplace’s equation in the region exterior to the cylinder. Also f(z)→V0zasz→∞,s o
that Re f(z)→V0x, which is appropriate to a uniform flow of speed V0in the x-direction
far from the cylinder.
Writing z=rexpiθand using de Moivre’s theorem we have
f(z)=V0
/
rexpiθ+a2
rexp(−iθ)
/
=V0
/
r+a2
r
/
cosθ+iV0
/
r−a2
r
/
sinθ.
Thus we see that the streamlines of the flow described by f(z) are given by
ψ=V0
/
r−a2
r
/
sinθ=c o n s t a n t .
In particular, ψ=0o n r=a, independently of the value of θ,a n ds o r=amust be a
streamline. Since there can be no flow of fluid across streamlines, r=amust correspond
to a boundary along which the fluid flows tangentially. Thus f(z) is a solution of Laplace’s
equation that satisfies all the physical boundary conditions of the problem, and so it is theappropriate complex potential.J
By a similar argument the complex potential f(z)=−E(z−a2/z) (note the
minus signs) is appropriate to a conducting circular cylinder of radius aplaced
perpendicular to a uniform electric field Ein the x-direction.
The real and imaginary parts of a complex potential f=φ+iψhave another
interesting relationship in the context of Laplace’s equation in electrostatics or
fluid mechanics. Let us choose φas the conventional potential, so that ψrepresents
the stream function (or electric field, depending on the application), and consider
728
20.7 COMPLEX POTENTIALS
PQy
x
ˆn
Figure 20.4 A curve joining the points PandQ. Also shown is ˆn, the unit
vector normal to the curve.
the difference in the values of ψat any two points PandQconnected by some
path C, as shown in figure 20.4. This difference is given by
ψ(Q)−ψ(P)=integraldisplayQ
Pdψ=integraldisplayQ
Pparenleftbigg∂ψ
∂xdx+∂ψ
∂ydyparenrightbigg
,
which, on using the Cauchy–Riemann relations, becomes
ψ(Q)−ψ(P)=integraldisplayQ
Pparenleftbigg
−∂φ
∂ydx+∂φ
∂xdyparenrightbigg
=integraldisplayQ
P∇φ·ˆnds=integraldisplayQ
P∂φ
∂nds,
where ˆnis the vector unit normal to the path Candsis the arc length along the
path; the last equality is written in terms of the normal derivative ∂φ/∂n≡∇φ·ˆn.
Now suppose that in an electrostatics application, the path Cis the surface of
a conductor; then
∂φ
∂n=−σ
/epsilon10,
where σis the surface charge density per unit length normal to the xy-plane.
Therefore−/epsilon10[ψ(Q)−ψ(P)] is equal to the charge per unit length normal to the
xy-plane on the surface of the conductor between the points PandQ. Similarly,
in fluid mechanics applications, if the density of the fluid is ρand its velocity V
then
ρ[ψ(Q)−ψ(P)] =ρintegraldisplayQ
P∇φ·ˆnds=ρintegraldisplayQ
PV·ˆnds
is equal to the mass flux between PandQper unit length perpendicular to the
xy-plane.
729
COMPLEX VARIABLESIA conducting circular cylinder of radius ais placed with its centre line passing through
the origin and perpendicular to a uniform electric field Ein the x-direction. Find the charge
per unit length induced on the half of the cylinder that lies in the region x<0.
As mentioned after the previous example, the appropriate complex potential for this
problem is f(z)=−E(z−a2/z). Writing z=rexpiθthis becomes
f(z)=−E
/
rexpiθ−a2
rexp(−iθ)
/
=−E
/
r−a2
r
/
cosθ−iE
/
r+a2
r
/
sinθ,
so that on r=athe imaginary part of fis given by
ψ=−2Easinθ.
Therefore the induced charge qper unit length on the left half of the cylinder, between
θ=π/2a n d θ=3π/2, is given by
q=2/epsilon10Ea[sin(3 π/2)−sin(π/2)] =−4/epsilon10Ea.
J
20.8 Conformal transformations
We now turn our attention to the subject of transformations, by which we mean
a change of coordinates from the complex variable z=x+iyto another, say
w=r+is, by means of a prescribed formula:
w=g(z)=r(x, y)+is(x, y).
Under such a transformation, or mapping , the Argand diagram for the z-variable
is transformed into one for the w-variable, although the complete z-plane might
be mapped onto only a part of the w-plane, or onto the whole of the w-plane, or
onto some or all of the w-plane covered more than once.
We shall consider only those mappings for which wandzare related by a
function w=g(z) and its inverse z=h(w) that are analytic, except possibly
at a few isolated points; such mappings are called conformal . Their important
properties are that, except at points at which g/prime(z), and hence h/prime(z), is zero or
infinite:
(i) continuous lines in the z-plane transform into continuous lines in the
w-plane;
(ii) the angle between two intersecting curves in the z-plane equals the angle
between the corresponding curves in the w-plane;
(iii) the magnification, as between the z-a n d w-plane, of a small line element in
the neighbourhood of any particular point is independent of the directionof the element;
(iv) any analytic function of ztransforms to an analytic function of wand
vice versa.
730
20.8 CONFORMAL TRANSFORMATIONS
y
θ1 θ2C1
C2z1
z2
z0
w=g(z)s
rC/prime
1
C/prime
2
φ1 φ2w0w1
w2
x
Figure 20.5 Two curves C1andC2in the z-plane, which are mapped onto
C/prime
1andC/prime
2in the w-plane.
Result (i) is immediate, and results (ii) and (iii) can be justified by the following
argument. Let two curves C1andC2pass through the point z0in the z-plane and
z1andz2be two points on their respective tangents at z0, each a distance ρfrom
z0. The same prescription with wreplacing zdescribes the transformed situation;
however, the transformed tangents may not be straight lines and the distancesofw
1andw2from w0have not yet been shown to be equal. This situation is
illustrated in figure 20.5.
In the z-plane z1andz2are given by
z1−z0=ρexpiθ1and z2−z0=ρexpiθ2.
The corresponding descriptions in the w-plane are
w1−w0=ρ1expiφ1and w2−w0=ρ2expiφ2.
The angles θiandφiare clear from figure 20.5. Now since w=g(z), where gis
analytic, we have
lim
z1→z0parenleftbiggw1−w0
z1−z0parenrightbigg
= lim
z2→z0parenleftbiggw2−w0
z2−z0parenrightbigg
=dg
dzvextendsinglevextendsinglevextendsinglevextendsingle
z=z0,
which may be written as
lim
ρ→0braceleftbiggρ1
ρexp[i(φ1−θ1)]bracerightbigg
= lim
ρ→0braceleftbiggρ2
ρexp[i(φ2−θ2)]bracerightbigg
=g/prime(z0). (20.30)
Comparing magnitudes and phases (i.e. arguments) in the equalities (20.30)
gives the stated results (ii) and (iii) and adds quantitative information to them,
731
COMPLEX VARIABLES
namely that for smallline elements
ρ1
ρ≈ρ2
ρ≈|g/prime(z0)|, (20.31)
φ1−θ1≈φ2−θ2≈argg/prime(z0). (20.32)
For strict comparison with result (ii), (20.32) must be written as θ1−θ2=φ1−φ2,
with an ordinary equality sign, since the angles are only defined in the limit
ρ→0 when (20.32) becomes a true identity. We also see from (20.31) that
the linear magnification factor is |g/prime(z0)|; similarly, small areas are magnified by
|g/prime(z0)|2.
Since in the neighbourhoods of corresponding points in a transformation angles
are preserved and magnifications are independent of direction, it follows that smallplane figures are transformed into figures of the same shape, but, in general, onesthat are magnified and rotated (though not distorted). However, we also notethat at any point where g
/prime(z) = 0, the angle arg g/prime(z) through which line elements
are rotated is undefined; these are called critical points of the transformation.
The final result (iv) is perhaps the most important property of conformal
transformations. If f(z) is an analytic function of zandz=h(w) is also analytic,
then F(w)=f(h(w)) is analytic in w. Its importance lies in the further conclusions
it allows us to draw from the fact that, since fis analytic, the real and imaginary
parts of f=φ+iψare necessarily solutions of
∂2φ
∂x2+∂2φ
∂y2=0 a n d∂2ψ
∂x2+∂2ψ
∂y2=0. (20.33)
Since the transformation property ensures that F=Φ+ iΨ is also analytic, we
can conclude that its real and imaginary parts must themselves satisfy Laplace’sequation in the w-plane,
∂
2Φ
∂r2+∂2Φ
∂s2=0 a n d∂2Ψ
∂r2+∂2Ψ
∂s2=0. (20.34)
Further, suppose that (say) Re f(z)=φis constant over a boundary Cin the
z-plane; then Re F(w) = Φ is constant over Cin the z-plane. But this is the same
as saying that Re F(w) is constant over the boundary C/primein the w-plane, C/primebeing
the curve into which Cis transformed by the conformal transformation w=g(z).
This is discussed further in the next section.
Examples of useful conformal transformations are numerous. For instance,
w=z+b,w=( e x p iφ)zandw=azcorrespond respectively to a translation by
b, a rotation through an angle φand a stretching (or contraction) in the radial
direction (for areal). These three examples can be combined into the general
linear transformation w=az+b, where in general aandbare complex. Another
example is the inversion mapping w=1/z, which maps the interior of the unit
circle to the exterior and vice versa. Other, more complicated, examples also exist.
732
20.8 CONFORMAL TRANSFORMATIONS
y
x Q R S TPi
w=g(z)R/prime
P/prime
S/primeQ/prime
T/primes
r
Figure 20.6 Transforming the upper half of the z-plane into the interior of
the unit circle in the w-plane, in such a way that z=iis mapped onto w=0
and the points x=±∞are mapped onto w=1 .IShow that, if the point z0lies in the upper half of the z-plane then the transformation
w=( e x p iφ)z−z0
z−z∗
0
maps the upper half of the z-plane into the interior of the unit circle in the w-plane. Hence
find a similar transformation that maps the point z=ionto w=0and the points x=±∞
onto w=1.
Taking the modulus of w, we have
|w|=
////(expiφ)z−z0
z−z∗
0
////=
////z−z0
z−z∗
0
////.
However, since the complex conjugate z∗
0is the reflection of z0in the real axis, if zandz0
both lie in the upper half of the z-plane then |z−z0|≤|z−z∗
0|; thus|w|≤1 as required.
We also note that (i) the equality holds only when zlies on the real axis, and so this axis
is mapped onto the boundary of the unit circle in the w-plane; (ii) the point z0is mapped
onto w= 0, the origin of the w-plane.
By fixing the images of two points in the z-plane, the constants z0andφc a na l s ob e
fixed. Since we require the point z=ito be mapped onto w= 0, we have immediately
z0=i. By further requiring z=±∞to be mapped onto w=1 ,w efi n d1= w=e x p iφ
and so φ= 0. The required transformation is therefore
w=z−i
z+i,
and is illustrated in figure 20.6.
J
We conclude this section by mentioning the rather curious Schwarz–Christoffel
transformation. †Suppose, as shown in figure 20.7, that we are interested in a
(finite) number of points x1,x2,...,x non the real axis in the z-plane. Then by
means of the transformation
w=braceleftbigg
Aintegraldisplayz
0(ξ−x1)(φ1/π)−1(ξ−x2)(φ2/π)−1···(ξ−xn)(φn/π)−1dξbracerightbigg
+B,(20.35)
†Strictly speaking the use of this transformation requires an understanding of complex integrals,
which are discussed in section 20.10 below.
733
COMPLEX VARIABLES
x1x2x3 x4x5w1
w2
w3w4w5
φ1
φ2 φ3φ4φ5y
xs
rw=g(z)
Figure 20.7 Transforming the upper half of the z-plane into the interior
of a polygon in the w-plane, in such a way that the points x1,x2,...,x nare
mapped onto the vertices w1,w2,...,w nof the polygon with interior angles
φ1,φ2,...,φ n.
we may map the upper half of the z-plane onto the interior of a closed polygon in
thew-plane having nvertices w1,w2,...,w n(which are the images of x1,x2,...,x n)
with corresponding interior angles φ1,φ2,...,φ n, as shown in figure 20.7. The
real axis in the z-plane is transformed into the boundary of the polygon itself.
The constants Aand Bare complex in general and determine the position,
size and orientation of the polygon. It is clear from (20.35) that dw/dz =0a t
x=x1,x2,...,x n, and so the transformation is not conformal at these points.
There are various subtleties associated with the use of the Schwarz–Christoffel
transformation. For example, if one of the points on the real axis in the z-plane
(usually xn) is taken at infinity then the corresponding factor in (20.35) (i.e. the
one involving xn) is not present. In this case, the point(s) x=±∞are considered
as one point, since they transform to a single vertex of the polygon in the w-plane.
We can also map the upper half of the z-plane onto an infinite openpolygon
by considering it as the limiting case of some closed polygon.IFind a transformation that maps the upper half of the z-plane onto the triangular region
shown in figure 20.8 in such a way that the points x1=−1andx2=1are mapped
onto the points w=−aandw=arespectively, and the point x3=±∞is mapped onto
w=ib. Hence find a transformation that maps the upper half of the z-plane into the region
−a<r<a ,s>0of the w-plane, as shown in figure 20.9.
Let us denote the angles at w1andw2in the w-plane by φ1=φ2=φ,w h e r e φ=t a n−1(b/a).
Since x3is taken at infinity we may omit the corresponding factor in (20.35) to obtain
w=
/
A
Zz
0(ξ+1 )(φ/π)−1(ξ−1)(φ/π)−1dξ
/
+B
=
/
A
Zz
0(ξ2−1)(φ/π)−1dξ
/
+B. (20.36)
The required transformation may then be found by fixing the constants AandBas
follows. Since the point z= 0 lies on the line segment x1x2it will be mapped onto the line
734
20.9 APPLICATIONS OF CONFORMAL TRANSFORMATIONS
φ1 φ2φ3
x1 x2
−11y
xw1 w2w3 ib
−aas
rw=g(z)
Figure 20.8 Transforming the upper half of the z-plane into the interior of
a triangle in the w-plane.
φ1 φ2 x1 x2
−11y
xw1 w2w3 w3
−aas
rw=g(z)
Figure 20.9 Transforming the upper half of the z-plane into the interior of
the region−a<r<a ,s>0i nt h e w-plane.
segment w1w2in the w-plane, and by symmetry must be mapped onto the point w=0 .
Thus setting z=0a n d w= 0 in (20.36) we obtain B= 0. An expression for Acan be
found in the form of an integral by setting (for example) z=1a n d w=ain (20.36).
We may consider the region in the w-plane in figure 20.9 to be the limiting case of the
triangular region in figure 20.8 with the vertex w3at infinity. Thus we may use the above,
but with the angles at w1andw2set to φ=π/2. From (20.36), we obtain
w=A
Zz
0dξp
ξ2−1=iAsin−1z.
By setting z=1a n d w=a, we find iA=2a/π, so the required transformation is
w=2a
πsin−1z.
J
20.9 Applications of conformal transformations
In the previous section it was shown that, under a conformal transformation
w=g(z)f r o m z=x+iyto a new variable w=r+is, if a solution of Laplace’s
equation in some region Rof the xy-plane can be found as the real or imaginary
735
COMPLEX VARIABLES
part of an analytic function †ofzthen the same expression put in terms of rand
swill be a solution of Laplace’s equation in the corresponding region R/primeof the
w-plane, and vice versa. In addition, if the solution is constant over the boundary
Cof the region Rin the xy-plane then the solution in the w-plane will take the
same constant value over the corresponding curve C/primethat bounds R/prime.
Thus, from any two-dimensional solution of Laplace’s equation for a particular
geometry, further solutions for other geometries can be obtained by makingconformal transformations. From the physical point of view the given geometryis usually complicated and so the solution is sought by transforming to a simplerone. However, working from simpler to more complicated situations can provideuseful experience, and make it more likely that the reverse procedure can be
tackled successfully.IFind the complex electrostatic potential associated with an infinite charged conducting
plate y=0, and thus obtain those associated with
(i) a semi-infinite charged conducting plate ( r>0,s=0),
(ii) the inside of a right-angled charged conducting wedge ( r> 0,s=0 and
r=0,s>0).
Figure 20.10( a) shows the equipotentials (broken lines) and field lines (solid) for the
infinite charged conducting plane y= 0. Suppose that we elect to make the real part of
the complex potential coincide with the conventional electrostatic potential. If the plate isc h a r g e dt oap o t e n t i a l Vthen clearly
φ(x, y)=V−ky, (20.37)
where kis related to the charge density σbyk=σ//epsilon1
0, since physically the electric field E
has components (0 ,σ/ /epsilon1 0)a n d E=−∇φ.
Thus what is needed is an analytic function of zof which the real part is V−ky.T h i s
can be obtained by inspection, but we may proceed formally and use the Cauchy–Riemannrelations to obtain the imaginary part ψ(x, y) thus:
∂ψ
∂y=∂φ
∂x=0 a n d∂ψ
∂x=−∂φ
∂y=k.
Hence ψ=kx+cand, absorbing cintoV, the required complex potential is
f(z)=V−ky+ikx=V+ikz. (20.38)
(i) Now consider the transformation
w=g(z)=z2. (20.39)
This satisfies the criteria for a conformal mapping (except at z= 0) and carries the upper
half of the z-plane into the entire w-plane; the equipotential plane y=0g o e si n t ot h e
half-plane r>0,s=0 .
By the general results proved, f(z) when expressed in terms of randswill give a
complex potential of which the real part will be constant on the half-plane in question;
†In fact, the original solution in the xy-plane need not be given explicitly as the real or imaginary
part of an analytic function. Any solution of ∇2φ= 0 in the xy-plane is carried over into another
solution of ∇2φ= 0 in the new variables by a conformal transformation, and vice versa.
736
20.9 APPLICATIONS OF CONFORMAL TRANSFORMATIONS
y s s
r r x
(a)z-plane (b)w-plane ( c)w-plane
Figure 20.10 ( a) The equipotential lines (broken) and field lines (solid) for
an infinite charged conducting plane at y=0 ,w h e r e z=x+iy;(b), (c) after
the transformations w=z2,w=z1/2of the situation shown in ( a).
we deduce that
F(w)=f(z)=V+ikz=V+ikw1/2(20.40)
is the required potential. Expressed in terms of r,sandρ=(r2+s2)1/2,w1/2is given by
w1/2=ρ1/2
/"/ρ+r
2ρ
/1/2
+i
/ρ−r
2ρ
/1/2
/#
(20.41)
and, in particular, the electrostatic potential is given by
Φ(r,s)=R e F(w)=V−k√
2
/
(r2+s2)1/2−r
/1/2. (20.42)
The corresponding equipotentials and field lines are shown in figure 20.10( b). Using results
(20.27)–(20.29), the magnitude of the electric field is
|E|=|F/prime(w)|=|1
2ikw−1/2|=1
2k(r2+s2)−1/4.
(ii) A transformation ‘converse’ to that used in (i),
w=g(z)=z1/2,
has the effect of mapping the upper half of the z-plane into the first quadrant of the
w-plane and the conducting plane y= 0 into the wedge r>0,s=0a n d r=0 , s>0.
The complex potential now becomes
F(w)=V+ikw2
=V+ik[(r2−s2)+2irs], (20.43)
showing that the electrostatic potential is V−2krsand the electric field has components
E=( 2ks,2kr). (20.44)
Figure 20.10( c) indicates the approximate equipotentials and field lines. (Note that, in
both transformations, g/prime(z)i se i t h e r0o r ∞at the origin and so neither transformation
is conformal there. Consequently there is no violation of result (ii), given at the start ofsection 20.8, concerning the angles between intersecting lines.)J
737
COMPLEX VARIABLES
φ=0φ=0
Φ=0 Φ=0y
xπ/az0w=zαs
rw0
w∗
0 (a) (b)
Figure 20.11 ( a) An infinite conducting wedge with interior angle π/αand a
line charge at z=z0;(b) after the transformation w=zα,w i t ha na d d i t i o n a l
image charge placed at w=w∗
0.
Themethod of images discussed in section 19.5 can also be used in conjunction
with conformal transformations to solve Laplace’s equation in two dimensions.IA wedge of angle π/αwith its vertex at z=0is formed by two semi-infinite conducting
plates, as shown in figure 20.11(a). A line charge of strength qper unit length is positioned
atz=z0, perpendicular to the z-plane. By considering the transformation w=zα,fi n dt h e
complex electrostatic potential for this situation.
Let us consider the action of the transformation w=zαon the lines defining the positions
of the conducting plates. The plate that lies along the positive x-axis is mapped onto the
positive r-axis in the w-plane, whereas the plate that lies along the direction exp( iπ/α)i s
mapped into the negative r-axis, as shown in figure 20.11( b). Similarly the line charge at
z0is mapped onto the point w0=zα
0.
From figure 20.11( b), we see that in the w-plane the problem can be solved by introducing
a second line charge of opposite sign at the point w∗
0, so that the potential Φ = 0 along
ther-axis. The complex potential for such an arrangement is simply
F(w)=−q
2π/epsilon10ln(w−w0)+q
2π/epsilon10ln(w−w∗
0).
Substituting w=zαinto the above shows that the required complex potential in the
original z-plane is
f(z)=q
2π/epsilon10ln
/zα−z∗α
0
zα−zα
0
/
.
J
20.10 Complex integrals
Corresponding to integration with respect to a real variable, it is possible to
define integration with respect to a complex variable between two complex limits.Since the z-plane is two-dimensional there is clearly greater freedom and hence
ambiguity in what is meant by a complex integral. If a complex function f(z)i s
single-valued and continuous in some region Rin the complex plane, then we can
define the complex integral of f(z) between two points AandBalong some curve
738
20.10 COMPLEX INTEGRALS
AB
C1C2
C3xy
Figure 20.12 Alternative paths for the integral of a function f(z) between A
andB.
inR; its value will depend, in general, upon the path taken between AandB
(see figure 20.12). However, we will find that for some paths that are differentbut bear a particular relationship to each other, the value of the integral does not
depend upon which of the paths is adopted.
Let a particular path Cbe described by a continuous (real) parameter t
(α≤t≤β) that gives successive positions on Cby means of the equations
x=x(t),y =y(t), (20.45)
with t=αandt=βcorresponding to the points AandBrespectively. Then the
integral along path Cof a continuous function f(z) is written
integraldisplay
Cf(z)dz (20.46)
and can be given explicitly as a sum of real integrals as follows:
integraldisplay
Cf(z)dz=integraldisplay
C(u+iv)(dx+idy)
=integraldisplay
Cud x−integraldisplay
Cvd y+iintegraldisplay
Cud y+iintegraldisplay
Cvd x
=integraldisplayβ
αudx
dtdt−integraldisplayβ
αvdy
dtdt+iintegraldisplayβ
αudy
dtdt+iintegraldisplayβ
αvdx
dtdt.
(20.47)
The question of when such an integral exists will not be pursued, except to state
that a sufficient condition is that dx/dt anddy/dt are continuous.
739
COMPLEX VARIABLES
(a)( b)( c)y y y
x x x RRR R
t t
−R −Rs=1iR
t=0C3b C3aC1 C2
Figure 20.13 Different paths for an integral of f(z)=z−1. See the text for
details.IEvaluate the complex integral of f(z)=z−1along the circle |z|=R, starting and finishing
atz=R.
The path C1is parameterised as follows (figure 20.13( a)):
z(t)=Rcost+iRsint, 0≤t≤2π,
whilst f(z)i sg i v e nb y
f(z)=1
x+iy=x−iy
x2+y2.
Thus the real and imaginary parts of f(z)a r e
u=x
x2+y2=Rcost
R2and v=−y
x2+y2=−Rsint
R2.
Hence, using expression (20.47),Z
C11
zdz=
Z2π
0cost
R(−Rsint)dt−
Z2π
0
/−sint
R
/
Rcostd t
+i
Z2π
0cost
RRcostd t+i
Z2π
0
/−sint
R
/
(−Rsint)dt (20.48)
=0+0+ iπ+iπ=2πi.
J
With a bit of experience, the reader may be able to evaluate integrals like
the LHS of (20.48) directly without having to write them as four separate real
integrals. In the present case,
integraldisplay
C1dz
z=integraldisplay2π
0−Rsint+iRcost
Rcost+iRsintdt=integraldisplay2π
0id t=2πi. (20.49)
This very important result will be used many times later, and the following should
be carefully noted: (i) its value, (ii) that this value is independent of R.
In the above example the contour was closed, and so it began and ended at
740
20.10 COMPLEX INTEGRALS
the same point in the Argand diagram. We can evaluate complex integrals along
open paths in a similar way.IEvaluate the complex integral of f(z)=z−1along
(i) the contour C2consisting of the semicircle |z|=Rin the half-plane y≥0,
(see figure 20.13(b)),
(ii) the contour C3made up of the two straight lines C3aand C3b(see
figure 20.13(c)).
(i) This is just as in the previous example, except that now 0 ≤t≤π. With this change we
have from (20.48) or (20.49) thatZ
C2dz
z=πi. (20.50)
(ii) The straight lines that make up the countour C3may be parameterised as follows:
C3a,z =( 1−t)R+itR for 0≤t≤1;
C3b,z =−sR+i(1−s)Rfor 0≤s≤1.
With these parameterisations the required integrals may be writtenZ
C3dz
z=
Z1
0−R+iR
R+t(−R+iR)dt+
Z1
0−R−iR
iR+s(−R−iR)ds. (20.51)
If we could take over from real-variable theory that, for real t,
R
(a+bt)−1dt=b−1ln(a+bt)
even if aandbare complex, then these integrals could be evaluated immediately. However,
to do this would be presuming to some extent what we wish to show, and so the evaluationmust be made in terms of entirely real integrals. For example, the first isZ1
0−R+iR
R(1−t)+itRdt=
Z1
0(−1+i)(1−t−it)
(1−t)2+t2dt
=
Z1
02t−1
1−2t+2t2dt+i
Z1
01
1−2t+2t2dt
=1
2
/
ln(1−2t+2t2)
/1
0+i
2
/"
2tan−1
/
t−1
2
1
2
/!/#1
0
=0+i
2
hπ
2−
/
−π
2
/i
=πi
2.
The second integral on the right of (20.51) can also be shown to have the value πi/2. ThusZ
C3dz
z=πi.
J
Considering the results of the last two examples, which have common inte-
grands and limits, some interesting observations are possible. Firstly, the twointegrals from z=Rtoz=−R,a l o n g C
2andC3respectively, have the same
value even though the paths taken are different. It also follows that if we took aclosed path C
4,g i v e nb y C2from Rto−RandC3traversed backwards from −R
toR, then the integral round C4ofz−1would be zero (both parts contributing
equal and opposite amounts). This is to be compared with result (20.49), in whichclosed path C
1, beginning and ending at the same place as C4, yields a value 2 πi.
741
COMPLEX VARIABLES
It is not true, however, that the integrals along the paths C2andC3are equal
for any function f(z), or, indeed, that their values are independent of Rin general.IEvaluate the complex integral of f(z)=R e zalong the paths C1,C2andC3shown in
figure 20.13.
(i) If we take f(z)=R e zand the contour C1thenZ
C1Rezd z=
Z2π
0Rcost(−Rsint+iRcost)dt=iπR2.
(ii) Using C2as the contour,Z
C2Rezd z=
Zπ
0Rcost(−Rsint+iRcost)dt=1
2iπR2.
(iii) Finally the integral along C3=C3a+C3bis given byZ
C3Rezd z=
Z1
0(1−t)R(−R+iR)dt+
Z1
0(−sR)(−R−iR)ds
=1
2R2(−1+i)+1
2R2(1 +i)=iR2.
J
The results of this section demonstrate that the value of an integral between
the same two points may depend upon the path that is taken between them but,at the same time, suggest that under some circumstances it is independent of the
path. The general situation is summarised in the result of the next section, namely
Cauchy’s theorem, which is the cornerstone of the integral calculus of complexvariables.
Before discussing Cauchy’s theorem, however, we note an important result
concerning complex integrals that will be of some use later. Let us consider theintegral of a function f(z) along some path C.I fMis an upper bound on the
value of|f(z)|on the path, i.e. |f(z)|≤MonC,a n d Lis the length of the path C,
thenvextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
Cf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle≤integraldisplay
c|f(z)||dz|≤Mintegraldisplay
Cdl=ML. (20.52)
It is straightforward to verify that this result does indeed hold for the complex
integrals considered earlier in this section.
20.11 Cauchy’s theorem
Cauchy’s theorem states that if f(z) is an analytic function, and f/prime(z) is continuous
at each point within and on a closed contour C,t h e n
contintegraldisplay
Cf(z)dz=0. (20.53)
In this statement and from now on we denote an integral around a closed contour
bycontintegraltext
C.
742
20.11 CAUCHY’S THEOREM
To prove this theorem we will need the two-dimensional form of the divergence
theorem, known as Green’s theorem in a plane (see section 11.3). This says thatifpandqare two functions with continuous first derivatives within and on a
closed contour C(bounding a domain R)i nt h e xy-plane, then
integraldisplayintegraldisplay
Rparenleftbigg∂p
∂x+∂q
∂yparenrightbigg
dxdy=contintegraldisplay
C(pd y−qd x). (20.54)
With f(z)=u+ivanddz=dx+id y, this can be applied to
I=contintegraldisplay
Cf(z)dz=contintegraldisplay
C(ud x−vd y)+icontintegraldisplay
C(vd x+ud y)
to give
I=integraldisplayintegraldisplay
Rbracketleftbigg∂(−u)
∂y+∂(−v)
∂xbracketrightbigg
dx dy+iintegraldisplayintegraldisplay
Rbracketleftbigg∂(−v)
∂y+∂u
∂xbracketrightbigg
dx dy. (20.55)
Now, recalling that f(z) is analytic and therefore that the Cauchy–Riemann
relations (20.5) apply, we see that each integrand is identically zero and thus Iis
also zero; this proves Cauchy’s theorem.
In fact the conditions of the above proof are more stringent than are needed.
The continuity of f/prime(z) is not necessary for the proof of Cauchy’s theorem,
analyticity of f(z) within and on Cbeing sufficient. However, the proof then
becomes more complicated and is too long to be given here. †
The connection between Cauchy’s theorem and the zero value of the integral
ofz−1around the composite path C4discussed towards the end of the previous
section is apparent: the function z−1is analytic in the two regions of the z-plane
enclosed by contours ( C2andC3a)a n d( C2andC3b).ISuppose two points AandBin the complex plane are joined by two different paths C1
andC2. Show that if f(z)is an analytic function on each path and in the region enclosed
by the two paths then the integral of f(z)is the same along C1andC2.
The situation is shown in figure 20.14. Since f(z)i sa n a l y t i ci n Rit follows from Cauchy’s
theorem that we haveZ
C1f(z)dz−
Z
C2f(z)dz=
I
C1−C2f(z)dz=0,
since C1−C2forms a closed contour enclosing R. Thus we immediately obtainZ
C1f(z)dz=
Z
C2f(z)dz,
and so the values of the integrals along C1andC2are equal.
J
An important application of Cauchy’s theorem is in proving that in some cases it
†The reader may refer to almost any book devoted to complex variables and the theory of functions.
743
COMPLEX VARIABLES
AB
Ry
xC1
C2
Figure 20.14 Two paths C1andC2enclosing a region R.
is possible to deform a closed contour Cinto another contour γ,i ns u c haw a yt h a t
the integrals of a function f(z) around each of the contours have the same value.IConsider two closed contours Candγin the Argand diagram, γbeing sufficiently small
that it lies completely within C. Show that if the function f(z)is analytic in the region
between the two contours thenI
Cf(z)dz=
I
γf(z)dz. (20.56)
To prove this result we consider a contour as shown in figure 20.15. The two close parallel
lines C1andC2joinγandC, which are ‘cut’ to accommodate them. The new contour Γ
so formed consists of C,C1,γandC2.
Within the area bounded by Γ the function f(z) is analytic and therefore, by Cauchy’s
theorem (20.53),I
Γf(z)dz=0. (20.57)
Now the parts C1andC2of Γ are traversed in opposite directions, and in the limit lie on
top of each other, and so their contributions to (20.57) cancel. ThusI
Cf(z)dz+
I
γf(z)dz=0. (20.58)
The sense of the integral round γis opposite to the conventional (anticlockwise) one, and
so by traversing γin the usual sense, we establish the result (20.56).
J
A sort of converse of Cauchy’s theorem is known as Morera’s theorem ,w h i c h
states that if f(z) is a continuous function of zin a closed domain Rbounded by
ac u r v e Cand, further,contintegraltext
Cf(z)dz=0 ,t h e n f(z) is analytic in R.
744
20.12 CAUCHY’S INTEGRAL FORMULA
y
xγ C1
C2C
Figure 20.15 The contour used to prove the result (20.56).
20.12 Cauchy’s integral formula
Another very important theorem in the theory of complex variables is Cauchy’s
integral formula , which states that if f(z) is analytic within and on a closed
contour Candz0is a point within Cthen
f(z0)=1
2πicontintegraldisplay
Cf(z)
z−z0dz. (20.59)
This formula is saying that the value of an analytic function anywhere inside
a closed contour is uniquely determined by its values on the contour †and that
the specific expression (20.59) can be given for the value at the interior point.
We may prove Cauchy’s integral formula by using (20.56) and taking γto be a
circle centred on the point z=z0, of small enough radius ρthat it all lies inside
C. Then, since f(z) is analytic inside C, the integrand f(z)/(z−z0) is analytic
in the space between Candγ. Thus, from (20.56), the integral around γhas the
same value as that around C.
We then use the fact that any point zonγis given by z=z0+ρexpiθ(and
sodz=iρexpiθ dθ). Thus the value of the integral around γis given by
I=contintegraldisplay
γf(z)
z−z0dz=integraldisplay2π
0f(z0+ρexpiθ)
ρexpiθiρexpiθ dθ
=iintegraldisplay2π
0f(z0+ρexpiθ)dθ.
†The similarity between this and the uniqueness theorem for Dirichlet boundary conditions (see
chapter 18) is apparent.
745
COMPLEX VARIABLES
If the radius of the circle γis now shrunk to zero, i.e. ρ→0, then I→2πif(z0),
thus establishing the result (20.59).
An extension to Cauchy’s integral formula can be made, yielding an integral
expression for f/prime(z0):
f/prime(z0)=1
2πiintegraldisplay
Cf(z)
(z−z0)2dz, (20.60)
under the same conditions as previously stated.IProve Cauchy’s integral formula for f/prime(z0)given in (20.60).
To show this, we use the definition of a derivative and (20.59) itself to evaluate
f/prime(z0) = lim
h→0f(z0+h)−f(z0)
h
= lim
h→0
/1
2πi
I
Cf(z)
h
/1
z−z0−h−1
z−z0
/
dz
/
= lim
h→0
/1
2πi
I
Cf(z)
(z−z0−h)(z−z0)dz
/
=1
2πi
I
Cf(z)
(z−z0)2dz,
which establishes the result (20.60).
J
Further, it may be proved by induction that the nth derivative of f(z) is also
given by a Cauchy integral,
f(n)(z0)=n!
2πicontintegraldisplay
Cf(z)dz
(z−z0)n+1. (20.61)
Thus, if the value of the analytic function is known on Cthen not only may the
value of the function at any interior point be calculated, but also the values ofallits derivatives.
The observant reader will notice that (20.61) may also be obtained by the
formal device of differentiating under the integral sign with respect to z
0in
Cauchy’s integral formula (20.59),
f(n)(z0)=1
2πicontintegraldisplay
C∂n
∂zn
0bracketleftbiggf(z)
(z−z0)bracketrightbigg
dz
=n!
2πicontintegraldisplay
Cf(z)dz
(z−z0)n+1.
746
20.13 TAYLOR AND LAURENT SERIESISuppose that f(z)is analytic inside and on a circle Cof radius Rcentred on the point
z=z0.I f|f(z)|≤Mon the circle, where Mis some constant, show that
|f(n)(z0)|≤Mn!
Rn. (20.62)
From (20.61) we have
|f(n)(z0)|=n!
2π
////
I
Cf(z)dz
(z−z0)n+1
////
and using (20.52) this becomes
|f(n)(z0)|≤n!
2πM
Rn+12πR=Mn!
Rn.
This result is known as Cauchy’s inequality .
J
We may use Cauchy’s inequality to prove Liouville’s theorem , which states that
iff(z) is analytic and bounded for all zthen fis a constant. Setting n=1i n
(20.62) and letting R→∞we find|f/prime(z0)|= 0 and hence f/prime(z0) = 0. Since f(z)i s
analytic for all zwe may take z0as any point in the z-plane and thus f/prime(z)=0
for all z; this implies f(z) = constant. Liouville’s theorem may be used in turn to
prove the fundamental theorem of algebra (see exercise 20.12).
20.13 Taylor and Laurent series
Following on from (20.61), we may establish Taylor’s theorem f o rf u n c t i o n so fa
complex variable. If f(z) is analytic inside and on a circle Cof radius Rcentred
on the point z=z0,a n d zis a point inside C,t h e n
f(z)=∞summationdisplay
n=0an(z−z0)n, (20.63)
where anis given by f(n)(z0)/n!. The Taylor expansion is valid inside the region
of analyticity and, for any particular z0,c a nb es h o w nt ob eu n i q u e .
To prove Taylor’s theorem (20.63), we note that, since f(z) is analytic inside
and on C, we may use Cauchy’s formula to write f(z)a s
f(z)=1
2πicontintegraldisplay
Cf(ξ)
ξ−zdξ, (20.64)
where ξlies on C. Now we may expand the factor ( ξ−z)−1as a geometric series
in (z−z0)/(ξ−z0),
1
ξ−z=1
ξ−z0∞summationdisplay
n=0parenleftbiggz−z0
ξ−z0parenrightbiggn
,
747
COMPLEX VARIABLES
so (20.64) becomes
f(z)=1
2πicontintegraldisplay
Cf(ξ)
ξ−z0∞summationdisplay
n=0parenleftbiggz−z0
ξ−z0parenrightbiggn
dξ
=1
2πi∞summationdisplay
n=0(z−z0)ncontintegraldisplay
Cf(ξ)
(ξ−z0)n+1dξ
=1
2πi∞summationdisplay
n=0(z−z0)n2πif(n)(z0)
n!, (20.65)
where we have used Cauchy’s integral formula (20.61) for the derivatives of
f(z). Cancelling the factors of 2 πi, we thus establish the result (20.63) with
an=f(n)(z0)/n!.IShow that if f(z)andg(z)are analytic in some region R, and f(z)=g(z)within some
subregion SofR,t h e n f(z)=g(z)throughout R.
It is simpler to consider the (analytic) function h(z)=f(z)−g(z), and to show that because
h(z)=0i n Sit follows that h(z) = 0 throughout R.
If we choose a point z=z0inSthen we can expand h(z) in a Taylor series about z0,
h(z)=h(z0)+h/prime(z0)(z−z0)+1
2h/prime/prime(z0)(z−z0)2+···,
which will converge inside some circle Cthat extends at least as far as the nearest part of
the boundary of R,s i n c e h(z)i sa n a l y t i ci n R. But since z0lies in S, we have
h(z0)=h/prime(z0)=h/prime/prime(z0)=···=0,
and so h(z)=0i n s i d e C. We may now expand about a new point, which can lie anywhere
within C, and repeat the process. By continuing this procedure we may show that h(z)=0
throughout R.
This result is called the identity theorem and, in fact, the equality of f(z)a n d g(z)
throughout Rfollows from their equality along as little as some curve in R,o re v e na ta
countably infinite number of points in R.
J
So far we have assumed that f(z) is analytic inside and on the (circular)
contour C.I f ,h o w e v e r , f(z) has a singularity inside Cat the point z=z0,t h e ni t
cannot be expanded in a Taylor series. Nevertheless, suppose that f(z) has a pole
of order patz=z0but is analytic at every other point inside and on C.T h e n
the function g(z)=( z−z0)pf(z) is analytic at z=z0, and so may be expanded
as a Taylor series about z=z0,
g(z)=∞summationdisplay
n=0bn(z−z0)n. (20.66)
Thus, for all zinside C,f(z) will have a power series representation of the form
f(z)=a−p
(z−z0)p+···+a−1
z−z0+a0+a1(z−z0)+a2(z−z0)2+···,
(20.67)
with a−p/negationslash= 0. Such a series, which is an extension of the Taylor expansion, is
748
20.13 TAYLOR AND LAURENT SERIES
xy
R
z0C1C2
Figure 20.16 The region of convergence Rfor a Laurent series of f(z) about
a point z=z0where f(z) has a singularity.
called a Laurent series . By comparing the coefficients in (20.66) and (20.67), we
see that an=bn+p. Now, the coefficients bnin the Taylor expansion of g(z)a r e
seen from (20.65) to be given by
bn=g(n)(z0)
n!=1
2πicontintegraldisplayg(z)
(z−z0)n+1dz,
and so for the coefficients anin (20.67) we have
an=1
2πicontintegraldisplayg(z)
(z−z0)n+1+pdz=1
2πicontintegraldisplayf(z)
(z−z0)n+1dz,
an expression that is valid for both positive and negative n.
The terms in the Laurent series with n≥0 are collectively called the analytic
part, whilst the remainder of the series, consisting of terms in inverse powers of
z−z0, is called the principal part . Depending on the nature of the point z=z0,
the principal part may contain an infinite number of terms, so that
f(z)=+∞summationdisplay
n=−∞an(z−z0)n. (20.68)
In this case we would expect the principal part to converge only for |(z−z0)−1|
less than some constant, i.e. outside some circle centred on z0. However, the
analytic part will converge inside some (different) circle also centred on z0.I ft h e
latter circle has the greater radius then the Laurent series will converge in theregion Rbetween the two circles (see figure 20.16); otherwise it does not converge
at all.
In fact, it may be shown that any function f(z) that is analytic in a region
Rbetween two such circles C
1andC2centred on z=z0can be expressed as
749
COMPLEX VARIABLES
a Laurent series about z0that converges in R. We note that, depending on the
nature of the point z=z0, the inner circle may be a point (when the principal
part contains only a finite number of terms) and the outer circle may have aninfinite radius.
We may use the Laurent series of a function f(z) about any point z=z
0to
classify the nature of that point. If f(z) is actually analytic at z=z0then in
(20.68) all anforn<0 must be zero. It may happen that not only are all an
zero for n<0 but a0,a1,...,am−1are all zero as well. In this case the first
non-vanishing term in (20.68) is am(z−z0)mwith m>0, and f(z)i st h e ns a i dt o
have a zero of order matz=z0.
Iff(z) is not analytic at z=z0then two cases arise, as discussed above ( pis
here taken as positive):
(i) It is possible to find an integer psuch that a−p/negationslash= 0 but a−p−k=0f o ra l l
integers k>0;
(ii) it is not possible to find such a lowest value of −p.
In case (i), f(z) is of the form (20.67) and is described as having a pole of order
patz=z0; the value of a−1(not a−p) is called the residue off(z) at the pole
z=z0, and will play an important part in later applications.
For case (ii), in which the negatively decreasing powers of z−z0do not
terminate, f(z) is said to have an essential singularity . These definitions should be
compared with those given in section 20.6.IFind the Laurent series of
f(z)=1
z(z−2)3
about the singularities z=0andz=2(separately). Hence verify that z=0is a pole of
order 1andz=2is a pole of order 3, and find the residue of f(z)at each pole.
To obtain the Laurent series about z=0 ,w es i m p l yw r i t e
f(z)=−1
8z(1−z/2)3
=−1
8z
/
1+(−3)
/
−z
2
/
+(−3)(−4)
2!
/
−z
2
/2
+(−3)(−4)(−5)
3!
/
−z
2
/3
+···
/
=−1
8z−3
16−3z
16−5z2
32−···.
Since the lowest power of zis−1, the point z= 0 is a pole of order 1. The residue of f(z)
atz= 0 is simply the coefficient of z−1in the Laurent expansion about that point and is
equal to−1/8.
The Laurent series about z= 2 is most easily found by letting z=2+ ξ(orz−2=ξ)
750
20.13 TAYLOR AND LAURENT SERIES
and substituting into the expression for f(z)t oo b t a i n
f(z)=1
(2 +ξ)ξ3=1
2ξ3(1 +ξ/2)
=1
2ξ3
/"
1−
/ξ
2
/
+
/ξ
2
/2
−
/ξ
2
/3
+
/ξ
2
/4
−···
/#
=1
2ξ3−1
4ξ2+1
8ξ−1
16+ξ
32−···
=1
2(z−2)3−1
4(z−2)2+1
8(z−2)−1
16+z−2
32−···.
From this series we see that z= 2 is a pole of order 3 and that the residue of f(z)a tz=2
is 1/8.
J
As we shall see in the next few sections, finding the residue of a function
at a singularity is of crucial importance in the evaluation of complex integrals.Specifically, formulae exist for calculating the residue of a function at a particular(singular) point z=z
0without having to expand the function explicitly as a
Laurent series about z0and identify the coefficient of ( z−z0)−1. The type of
formula generally depends on the nature of the singularity at which the residue
is required.ISuppose that f(z)has a pole of order mat the point z=z0. By considering the Laurent
series of f(z)about z0, derive a general expression for the residue R(z0)off(z)atz=z0.
Hence evaluate the residue of the function
f(z)=expiz
(z2+1 )2
at the point z=i.
Iff(z)h a sap o l eo fo r d e r matz=z0then its Laurent series about this point has the
form
f(z)=a−m
(z−z0)m+···+a−1
(z−z0)+a0+a1(z−z0)+a2(z−z0)2+···,
which, on multiplying both sides of the equation by ( z−z0)m,g i v e s
(z−z0)mf(z)=a−m+a−m+1(z−z0)+···+a−1(z−z0)m−1+···.
Differentiating both sides m−1t i m e s ,w eo b t a i n
dm−1
dzm−1[(z−z0)mf(z)] = ( m−1)!a−1+∞X
n=1bn(z−z0)n,
for some coefficients bn. In the limit z→z0, however, the terms in the sum disappear and
after rearranging we obtain the formula
R(z0)=a−1= lim
z→z0
/1
(m−1)!dm−1
dzm−1[(z−z0)mf(z)]
/
, (20.69)
which gives the value of the residue of f(z) at the point z=z0.
If we now consider the function
f(z)=expiz
(z2+1 )2=expiz
(z+i)2(z−i)2,
751
COMPLEX VARIABLES
we see immediately that it has poles of order 2 ( double poles) at z=iandz=−i.T o
calculate the residue at (for example) z=i, we may apply the formula (20.69) with m=2 .
Performing the required differentiation we obtain
d
dz[(z−i)2f(z)] =d
dz
/expiz
(z+i)2
/
=1
(z+i)4[(z+i)2iexpiz−2(exp iz)(z+i)].
Setting z=iwe find the residue is given by
R(i)=1
1!1
16
/;
−4ie−1−4ie−1
/
=−i
2e.
J
An important special case of (20.69) occurs when f(z)h a sa simple pole (a pole
of order 1) at z=z0. Then the residue at z0is given by
R(z0) = lim
z→z0[(z−z0)f(z)]. (20.70)
Iff(z) has a simple pole at z=z0and, as is often the case, has the form
g(z)/h(z), where g(z) is analytic and non-zero at z0andh(z0) = 0, then (20.70)
becomes
R(z0) = lim
z→z0(z−z0)g(z)
h(z)=g(z0) lim
z→z0(z−z0)
h(z)
=g(z0) lim
z→z01
h/prime(z)=g(z0)
h/prime(z0), (20.71)
where we have used l’H ˆopital’s rule. This result often provides the simplest way
of determining the residue at a simple pole.
20.14 Residue theorem
Having seen from Cauchy’s theorem that the value of an integral round a closed
contour Cis zero if the integrand is analytic inside the contour, it is natural to
ask what value it takes when the integrand is not analytic inside C.T h ea n s w e r
to this is contained in the residue theorem, which we now discuss.
Suppose the function f(z) has a pole of order mat the point z=z0,a n ds o
can be written as a Laurent series about z0of the form
f(z)=∞summationdisplay
n=−man(z−z0)n. (20.72)
Now consider the integral Ioff(z) around a closed contour Cthat encloses
z=z0, but no other singular points. Using Cauchy’s theorem this integral has
the same value as the integral around a circle γof radius ρcentred on z=z0,
since f(z) is analytic in the region between Cand γ. On the circle we have
752
20.14 RESIDUE THEOREM
z=z0+ρexpiθ(and dz=iρexpiθ dθ), and so
I=contintegraldisplay
γf(z)dz
=∞summationdisplay
n=−mancontintegraldisplay
(z−z0)ndz
=∞summationdisplay
n=−manintegraldisplay2π
0iρn+1exp[i(n+1 )θ]dθ.
For every term in the series with n/negationslash=−1, we have
integraldisplay2π
0iρn+1exp[i(n+1 )θ]dθ=bracketleftbiggiρn+1exp[i(n+1 )θ]
i(n+1 )bracketrightbigg2π
0=0,
but for the n=−1t e r mw eo b t a i n
integraldisplay2π
0id θ=2πi.
Therefore only the term in ( z−z0)−1contributes to the value of the integral
around γ(and therefore C), and Itakes the value
I=contintegraldisplay
Cf(z)dz=2πia−1. (20.73)
Thus the integral around any closed contour containing a single pole of general
order m(or, by extension, an essential singularity) is equal to 2 πitimes the residue
off(z)a tz=z0.
If we extend the above argument to the case where f(z) is continuous within
and on a closed contour Cand analytic, except for a finite number of poles,
within C, then we arrive at the residue theorem
contintegraldisplay
Cf(z)dz=2πisummationdisplay
jRj, (20.74)
wheresummationtext
jRjis the sum of the residues of f(z) at its poles within C.
The method of proof is indicated by figure 20.17, in which ( a) shows the original
contour Creferred to in (20.74) and ( b) shows a contour C/primegiving the same value
to the integral, because fis analytic between CandC/prime. Now the contribution to
theC/primeintegral from the polygon (a triangle for the case illustrated) joining the
small circles is zero, since fis also analytic inside C/prime. Hence the whole value of
the integral comes from the circles and, by result (20.73), each of these contributes2πitimes the residue at the pole it encloses. All the circles are traversed in their
positive sense if Cis thus traversed and so the residue theorem follows. Formally,
Cauchy’s theorem (20.53) is a particular case of (20.74) in which Cencloses no
poles.
Finally we mention another important result, which we will use later. Suppose
753
COMPLEX VARIABLES
CCC/prime
(a) (b)
Figure 20.17 The contours used to prove the residue theorem: ( a) the original
contour; ( b) the contracted contour encircling each of the poles.
that f(z) has a simple pole at z=z0and so may be expanded as the Laurent
series
f(z)=φ(z)+a−1(z−z0)−1,
where φ(z) is analytic within some neighbourhood surrounding z0. We wish to
find an expression for the integral Ioff(z)a l o n ga n opencontour C,w h i c hi s
the arc of a circle of radius ρcentred on z=z0given by
|z−z0|=ρ, θ 1≤arg(z−z0)≤θ2, (20.75)
where ρis chosen small enough that no singularity of f, other than z=z0, lies
within the circle. Then Iis given by
I=integraldisplay
Cf(z)dz=integraldisplay
Cφ(z)dz+a−1integraldisplay
C(z−z0)−1dz.
If the radius of the arc Cis now allowed to tend to zero then the first integral
tends to zero, since the path becomes of zero length and φis analytic and
therefore continuous along it. On C,z=ρeiθand hence the required expression
forIis
I= lim
ρ→0integraldisplay
Cf(z)dz= lim
ρ→0parenleftbigg
a−1integraldisplayθ2
θ11
ρeiθiρeiθdθparenrightbigg
=ia−1(θ2−θ1).(20.76)
We note that result (20.73) is a special case of (20.76) in which θ2is equal to
θ1+2π.
20.15 Location of zeroes
An important use of the residue theorem is to locate the zeroes of functions
of a complex variable. The location of such zeroes has a particular applicationin electrical network and general oscillation theory, since the complex zeroes of
754
20.15 LOCATION OF ZEROES
certain functions give the system parameters (usually frequencies) at which system
instabilities occur. As the basis of a method for locating these zeroes we nextprove three important theorems.
(i) If f(z) has poles as its only singularities inside a closed contour Cand is
not zero at any point on Cthen
contintegraldisplay
Cf/prime(z)
f(z)dz=2πisummationdisplay
j(Nj−Pj). (20.77)
Here Njis the order of the jth zero of f(z) enclosed by C. Similarly Pjis the
order of the jth pole of f(z) inside C.
To prove this we note that, at each position zj,f(z) can be written as
f(z)=(z−zj)mjφ(z), (20.78)
where φ(z) is analytic and non-zero at z=zjandmjis positive for a zero and
negative for a pole. Then the integrand f/prime(z)/f(z)t a k e st h ef o r m
f/prime(z)
f(z)=mj
z−zj+φ/prime(z)
φ(z). (20.79)
Since φ(zj)/negationslash= 0, the second term on the right is analytic; thus the integrand
has a simple pole at z=zj, with residue mj. For zeroes mj=Njand for poles
mj=−Pj, and thus by the residue theorem (20.77) follows.
(ii) If f(z) is analytic inside Cand not zero at any point on it then
2πsummationdisplay
jNj=∆ C[argf(z)], (20.80)
where ∆ C[x] denotes the variation in xaround the contour C.
Since fis analytic there are no Pj; further, since
f/prime(z)
f(z)=d
dz[Lnf(z)], (20.81)
equation (20.77) can be written
2πisummationdisplay
Nj=contintegraldisplay
Cf/prime(z)
f(z)dz=∆ C[Lnf(z)]. (20.82)
However,
∆C[Lnf(z)] = ∆ C[ln|f(z)|]+i∆C[argf(z)], (20.83)
and, since Cis a closed contour, ln |f(z)|must return to its original value and
so the real term on the RHS is zero. Comparison of (20.82) and (20.83) then
establishes (20.80), which is known as the principle of the argument .
(iii) If f(z)a n d g(z) are analytic within and on a closed contour Cand
|g(z)|<|f(z)|onCthen f(z)a n d f(z)+g(z) have the same number of zeroes
inside C;t h i si s Rouch ´e’s theorem .
755
COMPLEX VARIABLES
With the conditions given, neither f(z)n o r f(z)+g(z) can have a zero on C.
So, applying theorem (ii) with an obvious notation,
2πsummationtext
jNj(f+g)=∆ C[arg(f+g)]
=∆ C[argf]+∆ C[arg(1 + g/f)]
=2πsummationtext
kNk(f)+∆ C[arg(1 + g/f)]. (20.84)
Further, since |g|<|f|onC,1+ g/falways lies within a unit circle centred
onz= 1; thus its argument always lies in the range −π/2<arg(1 + g/f)<π /2
and cannot change by any multiple of 2 π. It must therefore return to its
original value when zreturns to its starting point having traversed C.H e n c e
the second term on the right of (20.84) is zero and the theorem is estab-
lished.
The importance of Rouch ´e’s theorem is that for some functions, in particular
polynomials, only the behaviour of a single term in the function need be con-sidered if the contour is chosen appropriately. For example, for a polynomial,treated as f(z)+g(z), only the properties of its largest- (smallest-) power, taken as
f(z), need be investigated, if a circular contour is chosen with radius Rsufficiently
large (small) that, on the contour, the magnitude of the largest (smallest) power
term is greater than the sum of the magnitudes of all other terms. Further, if the
zeroes of f(z)+g(z)=summationtext
N
0bnznare considered as the roots of f(z)+g(z)=0 ,
w r i t t e ni nt h ef o r m
1+g(z)
f(z)=0, (20.85)
then it is apparent that no roots can lie outside (inside) |z|=Rand also that
f(z)=bNzN(orb0)h a s N(or 0) zeroes inside |z|=R;f+gconsequently has
the same number of zeroes inside the same circle.
Aw e a kf o r mo ft h e maximum-modulus theorem may also be deduced. This
states that if f(z) is analytic within and on a simple closed contour Cthen|f(z)|
attains its maximum value on the boundary of C.
Let|f(z)|≤MonCwith equality at at least one point of C. Now suppose
that there is a point z=ainside Csuch that|f(a)|>M. Then the function
h(z)≡f(a) is such that |h(z)|>|−f(z)|onC, and thus h(z)a n d h(z)−f(z) have
the same number of zeroes inside C.B u t h(z)(≡f(a)) has no zeroes inside C
and by Rouch ´e’s theorem this would imply that f(a)−f(z) has no zeroes in C.
However, f(a)−f(z) clearly has a zero at z=a, and so we have a contradiction;
the assumption of a point z=ainside Csuch that|f(a)|>Mmust be invalid.
This establishes the theorem.
The stronger form of the maximum-modulus theorem, which we do not prove,
states in addition that the maximum value of f(z) is not attained at any interior
point except for the case where f(z) is a constant.
756
20.15 LOCATION OF ZEROES
Yy
XR
x O
Figure 20.18 A contour for locating the zeroes of a polynomial that occur
in the first quadrant of the Argand diagram.IShow that the four zeroes of h(z)=z4+z+1occur one in each quadrant of the Argand
diagram and that all four lie between the circles |z|=2/3and|z|=3/2.
Putting z=xandz=iyshows that no zeroes occur on the real or imaginary axes. They
must therefore occur in conjugate pairs (as is shown by taking the complex conjugate ofh(z)=0 ) .
Now take Cas the contour OXY O shown in figure 20.18 and consider the changes
∆[arg h] in the argument of h(z)a sztraverses C.
(i)OX:a r g his everywhere zero, since his real, and thus ∆
OX[argh]=0 .
(ii)XY:z=Rexpiθa n ds oa r g hchanges by an amount
∆XY[argh]=∆ XY[argz4]+∆ XY[arg(1 + z−3+z−4)]
=∆ XY[argR4e4iθ]+∆ XY
/
arg[1 + O( R−3)]
/
=2π+O ( R−3). (20.86)
(iii)YO:z=iyand so arg h=y/(y4+ 1), which starts at O( R−3) and finishes at 0 as y
goes from large Rto 0. It never reaches π/2 because y4+1 = 0 has no real positive
root. Thus ∆ YO[argh]=0 .
Hence for the complete contour ∆ C[argh]=0+2 π+0+O ( R−3) and, if Ris allowed
to tend to infinity, we deduce from (20.80) that h(z) has one zero in the first quadrant.
Furthermore, since the roots occur in conjugate pairs, a second root must lie in the fourthquadrant and the other pair in the second and third quadrants.
To show that the zeroes lie within a given annulus in the z-plane we must apply Rouch ´e’s
theorem, as follows.
(i) With Cas|z|=3/2,f=z
4,g=z+1 .N o w|f|=8 1/16 on Cand|g|≤1+|z|<
5/2<81/16. Thus since z4=0h a sf o u rr o o t si n s i d e |z|=3/2, so also does
z4+z+1=0 .
(ii) With Cas|z|=2/3,f=1 , g=z4+z.N o w f=1o n Cand|g|≤|z4|+|z|=
16/81 + 2 /3=7 0 /81<1. Thus since f= 0 has no roots inside |z|=2/3, neither
does 1 + z+z4=0 .
Hence the four zeroes of h(z)=z4+z+ 1 occur one in each quadrant and all lie between
the circles|z|=2/3a n d|z|=3/2.
J
757
COMPLEX VARIABLES
A further technique useful in locating function zeroes is explained in exer-
cise 20.16.
20.16 Integrals of sinusoidal functions
The remainder of this chapter is devoted to methods of applying contour inte-
gration and the residue theorem to various types of definite integral. In each case
not much preamble is given since, for this material, the simplest explanation is
felt to be via a series of worked examples that can be used as models.
Suppose that an integral of the form
integraldisplay2π
0F(cosθ,sinθ)dθ (20.87)
is to be evaluated. It can be made into a contour integral around the unit circle
Cby writing z=e x p iθand hence
cosθ=1
2(z+z−1),sinθ=−1
2i(z−z−1),d θ =−iz−1dz. (20.88)
This contour integral can then be evaluated using the residue theorem, provided
the transformed integrand has only a finite number of poles inside the unit circlea n dn o n eo ni t .IEvaluate
I=
Z2π
0cos 2θ
a2+b2−2abcosθdθ, b > a > 0. (20.89)
By de Moivre’s theorem (section 3.4),
cosnθ=1
2(zn+z−n). (20.90)
Using n= 2 in (20.90) and straightforward substitution for the other functions of θin
(20.89) gives
I=i
2ab
I
Cz4+1
z2(z−a/b)(z−b/a)dz.
Thus there are two poles inside C, a double pole at z= 0 and a simple pole at z=a/b
(recall that b>a).
We could find the residue of the integrand at z= 0 by expanding the integrand as a
Laurent series in zand identifying the coefficient of z−1. Alternatively, we may use the
formula (20.69) with m= 2. Denoting the integrand by f(z) we have
d
dz[z2f(z)] =d
dz
/z4+1
(z−a/b)(z−b/a)
/
=(z−a/b)(z−b/a)4z3−(z4+1 ) [ ( z−a/b)+(z−b/a)]
(z−a/b)2(z−b/a)2.
Setting z= 0 and applying (20.69), we find
R(0) =a
b+b
a.
758
20.17 SOME INFINITE INTEGRALS
y
xR −R OΓ
Figure 20.19 A semicircular contour in the upper half-plane.
For the simple pole at z=a/b, equation (20.70) gives the residue as
R(a/b) = lim
z→(a/b)
/
(z−a/b)f(z)
/
=(a/b)4+1
(a/b)2(a/b−b/a)
=−a4+b4
ab(b2−a2).
Therefore by the residue theorem
I=2πi×i
2ab
/a2+b2
ab−a4+b4
ab(b2−a2)
/
=2πa2
b2(b2−a2).
J
20.17 Some infinite integrals
Suppose we wish to evaluate an integral of the form
integraldisplay∞
−∞f(x)dx,
where f(z) has the following properties.
(i)f(z) is analytic in the upper half-plane, Im z≥0, except for a finite number
of poles, none of which are on the real axis.
(ii) on a semicircle Γ of radius R(figure 20.19), Rtimes the maximum of |f|
on Γ tends to zero as R→∞ (a sufficient condition is that zf(z)→0a s
|z|→∞ ).
(iii)integraltext0
−∞f(x)dxandintegraltext∞
0f(x)dxboth exist.
The required integral is then given by
integraldisplay∞
−∞f(x)dx=2πi×(sum of the residues at poles with Im z>0).
(20.91)
759
COMPLEX VARIABLES
Sincevextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
Γf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle≤2πR×(maximum of |f|on Γ) ,
condition (ii) ensures that the integral along Γ tends to zero as R→∞, after
which (20.91) is obvious from the residue theorem.IEvaluate
I=
Z∞
0dx
(x2+a2)4,where ais real .
The complex function ( z2+a2)−4has poles of order 4 at z=±aiof which only z=ai
is in the upper half-plane. Conditions (ii) and (iii) are clearly satisfied. For higher-orderpoles, formula (20.69) for evaluating residues can be tiresome to apply. So ,instead, we putz=ai+ξand expand for small ξto obtain†
1
(z2+a2)4=1
(2aiξ+ξ2)4=1
(2aiξ)4
/
1−iξ
2a
/−4
.
The coefficient of ξ−1is
1
(2a)4(−4)(−5)(−6)
3!
/−i
2a
/3
=−5i
32a7,
and hence by the residue theoremZ∞
−∞dx
(x2+a2)4=10π
32a7,
and so I=5π/(32a7).
J
Condition (i) of the previous method required there to be no poles of the
integrand on the real axis, but in fact simple poles on the real axis can beaccommodated by indenting the contour as shown in figure 20.20. The indentationat the pole z=z
0is in the form of a semicircle γof radius ρin the upper half-
plane, thus excluding the pole from the interior of the contour.
What is then obtained from a contour integration, apart from the contributions
for Γ and γ, is called the principal value of the integral, defined as ρ→0b y :
PintegraldisplayR
−Rf(x)dx≡integraldisplayz0−ρ
−Rf(x)dx+integraldisplayR
z0+ρf(x)dx.
The remainder of the calculation goes through as before, but the contribution
from the semicircle γmust be included. Result (20.76) of section 20.14 shows that
since only a simple pole is involved its contribution is
−ia−1π, (20.92)
where a−1is the residue at the pole and the minus sign arises because γis
traversed in the clockwise (negative) sense.
†This illustrates another useful technique for determining residues.
760
20.17 SOME INFINITE INTEGRALS
y
xR −R OΓ
γ
Figure 20.20 An indented contour used when the integrand has a simple
pole on the real axis.
We defer giving an example of an indented contour until we have established
Jordan’s lemma ; we will then work through an example illustrating both. Jordan’s
lemma enables infinite integrals involving sinusoidal functions to be evaluated.
Jordan’s lemma. For a function f(z)of a complex variable z,i f
(i)f(z)is analytic in the upper half-plane except for a finite number of poles
inImz>0,
(ii)the maximum of |f(z)|→0as|z|→∞ in the upper half-plane,
(iii)m>0,
then
IΓ=integraldisplay
Γeimzf(z)dz→0asR→∞ , (20.93)
where Γis the same semicircular contour as in figure 20.19.
Notice that this condition (ii) is less stringent than the earlier condition (ii)
(see the start of this section), since we now only require M(R)→0 and not
RM(R)→0, where Mis the maximum †of|f(z)|on|z|=R.
The proof of the lemma is straightforward, once it has been observed that, for
0≤θ≤π/2,
1≥sinθ
θ≥2
π. (20.94)
Then, since on Γ we have |exp(imz)|=|exp(−mRsinθ)|,
IΓ≤integraldisplay
Γ|eimzf(z)||dz|≤MRintegraldisplayπ
0e−mRsinθdθ=2MRintegraldisplayπ/2
0e−mRsinθdθ.
†More strictly the least upper bound.
761
COMPLEX VARIABLES
Thus, using (20.94),
IΓ<2MRintegraldisplayπ/2
0e−mR(2θ/π)dθ=πM
mparenleftbig
1−e−mRparenrightbig
<πM
m;
hence,as R→∞,IΓtends to zero since Mdoes.IFind the principal value ofZ∞
−∞cosmx
x−adx, forareal, m>0.
Consider the function ( z−a)−1exp(imz); although it has no poles in the upper half-plane
it does have a simple pole at z=a, and further |(z−a)−1|→0a s|z|→∞ . We will use a
contour like that shown in figure 20.20 and apply the residue theorem. Symbolically,Za−ρ
−R+
Z
γ+
ZR
a+ρ+
Z
Γ=0. (20.95)
Now as R→∞ andρ→0 we have
R
Γ→0, by Jordan’s lemma, and from (20.91) and
(20.92) we obtain
P
Z∞
−∞eimx
x−adx−iπa−1=0, (20.96)
where a−1is the residue of ( z−a)−1exp(imz)a tz=a, which is exp( ima). Then taking the
real and imaginary parts of (20.96) gives
P
Z∞
−∞cosmx
x−adx=−πsinma, as required ,
P
Z∞
−∞sinmx
x−adx=πcosma, as a bonus.
J
20.18 Integrals of multivalued functions
We have discussed briefly some of the properties and difficulties associated with
certain multivalued functions such as z1/2or Ln z. It was mentioned that one
method of managing such functions is by means of a ‘cut plane’. A similartechnique can be used with advantage to evaluate some kinds of infinite integral
involving real functions for which the corresponding complex functions are multi-
valued. A typical contour employed for functions with a single branch pointlocated at the origin is shown in figure 20.21. Here Γ is a large circle of radius R
andγa small one of radius ρ, both centred on the origin. Eventually we will let
R→∞andρ→0.
The success of the method is due to the fact that because the integrand is
multivalued, its values along the two lines ABandCDjoining z=ρtoz=R
arenotequal and opposite although both are related to the corresponding real
integral. Again an example gives the best explanation.
762
20.18 INTEGRALS OF MULTIVALUED FUNCTIONS
y
xA B
C DΓ
γ
Figure 20.21 A typical cut-plane contour for use with multivalued functions
that have a single branch point located at the origin.IEvaluate
I=
Z∞
0dx
(x+a)3x1/2,a > 0.
We consider the integrand f(z)=( z+a)−3z−1/2and note that |zf(z)|→0o nt h et w o
circles as ρ→0a n d R→∞. Thus the two circles make no contribution to the contour
integral.
The only pole of the integrand inside the contour is at z=−a(and is of order 3).
To determine its residue we put z=−a+ξand expand (noting that ( −a)1/2equals
a1/2exp(iπ/2) = ia1/2):
1
(z+a)3z1/2=1
ξ3ia1/2(1−ξ/a)1/2
=1
iξ3a1/2
/
1+1
2ξ
a+3
8ξ2
a2+···
/
.
The residue is thus −3i/(8a5/2).
The residue theorem (20.74) now givesZ
AB+
Z
Γ+
Z
DC+
Z
γ=2πi
/−3i
8a5/2
/
.
We have seen that
R
Γand
R
γvanish and if we denote zbyxalong the line ABthen it has
the value z=xexp2 πialong the line DC(note that exp2 πimust not be set equal to 1
until after the substitution for zhas been made in
R
DC). Substituting these expressions,Z∞
0dx
(x+a)3x1/2+
Z0
∞dx
[xexp2 πi+a]3x1/2exp(1
22πi)=3π
4a5/2.
763
COMPLEX VARIABLES
Thus/
1−1
expπi
/Z∞
0dx
(x+a)3x1/2=3π
4a5/2,
and
I=1
2×3π
4a5/2.
J
20.19 Summation of series
Sometimes a real infinite series may be summed if a suitable complex function
can be found that has poles on the real axis at positions corresponding to thevalues of the dummy variable in the summation, and whose residues at thesepoles are equal to the values of the terms of the series there.IBy consideringI
Cπcotπz
(a+z)2dz,
where ais not an integer and Cis a circle of large radius, evaluate
∞X
n=−∞1
(a+n)2.
The integrand has (i) simple poles at z= integer n,f o r−∞<n<∞, (ii) a double pole at
z=−a.
(i) To find the residue of cot πz, put z=n+ξfor small ξ:
cotπz=cos(nπ+ξπ)
sin(nπ+ξπ)≈cosnπ
(cosnπ)ξπ=1
ξπ.
The residue of the integrand at z=nis thus π(a+n)−2π−1.
(ii) Putting z=−a+ξfor small ξand determining the coefficient of ξ−1,†
πcotπz
(a+z)2=π
ξ2cot(−aπ+ξπ)
=π
ξ2
/
cot(−aπ)+ξ
/d
dz(cotπz)
/
z=−a+···
/
,
so that the residue at the double pole z=−ais
π[−πcosec2πz]z=−a=−π2cosec2πa.
Collecting together these results to express the residue theorem gives
I=
I
Cπcotπz
(a+z)2dz=2πi
/"NX
n=−N1
(a+n)2−π2cosec2πa
/#
, (20.97)
where Nequals the integer part of R. But as the radius RofCtends to∞,c o tπz→∓i
(depending on whether Im zis greater or less than zero respectively). Thus
I<k
Zdz
(a+z)2,
†This again illustrates one of the techniques for determining residues.
764
20.20 INVERSE LAPLACE TRANSFORM
which tends to 0 as R→∞. Thus I→0a sR(and hence N)→∞and (20.97) establishes
the result
∞X
n=−∞1
(a+n)2=π2
sin2πa.
J
Series with alternating signs in the terms, i.e. ( −1)n, can also be attempted
in this way but using cosec πzinstead of cot πz, since the former has residue
(−1)nπ−1atz=n(see exercise 20.30).
20.20 Inverse Laplace transform
As a final example of contour integration we mention a method whereby the
process of Laplace transformation, discussed in chapter 13, can be inverted.
It will be recalled that the Laplace transform ¯f(s) of a function f(x),x≥0, is
given by
¯f(s)=integraldisplay∞
0e−sxf(x)dx, Res>s 0. (20.98)
In chapter 13, functions f(x) were deduced from the transforms by means of a
prepared dictionary. However, an explicit formula for an unknown inverse may
be written in the form of an integral. It is known as the Bromwich integral and is
given by
f(x)=1
2πiintegraldisplayλ+i∞
λ−i∞esx¯f(s)ds, λ > 0, (20.99)
where sis treated as a complex variable and the integration is along the line L
indicated in figure 20.22. The position of the line is dictated by the requirementsthatλis positive and that all singularities of ¯f(s) lie to the left of the line.
That (20.99) really is the unique inverse of (20.98) is difficult to show for general
functions and transforms, but the following verification should at least make itplausible:
f(x)=1
2πiintegraldisplayλ+i∞
λ−i∞ds esxintegraldisplay∞
0e−suf(u)du, Re(s)>0,i.e.λ>0,
=1
2πiintegraldisplay∞
0du f(u)integraldisplayλ+i∞
λ−i∞es(x−u)ds
=1
2πiintegraldisplay∞
0du f(u)integraldisplay∞
−∞eλ(x−u)eip(x−u)i dp, putting s=λ+ip,
=1
2πintegraldisplay∞
0f(u)eλ(x−u)2πδ(x−u)du
=braceleftBigg
f(x)x≥0,
0 x<0.(20.100)
765
COMPLEX VARIABLES
Ims
Res
L
λ
Figure 20.22 The integration path of the inverse Laplace transform is along
the infinite line L.T h eq u a n t i t y λmust be positive and large enough for all
poles of the integrand to lie to the left of L.
Our main interest here is in the use of contour integration. To employ it to
evaluate the line integral in (20.99), the path Lmust be made into a closed
contour in such a way that the contribution from the completion either vanishes
or is simply calculable.
A typical completion is shown in figure 20.23( a) and would be appropriate if
¯f(s) had a finite number of poles. For more complicated cases, in which ¯f(s)h a s
an infinite sequence of poles but all to the left of Las in figure 20.23( b), a sequence
of circular-arc completions that pass between the poles must be used and f(x)i s
obtained as a series. If ¯f(s) is a multivalued function then a cut plane is needed
and a contour such as that shown in figure 20.23( c) might be appropriate.
We consider here only the simple case in which the contour in figure 20.23( a)
is used; we refer the reader to the exercises at the end of the chapter for others.Ideally, we would like the contribution to the integral from the circular arc Γ to
tend to zero as its radius R→∞. Using a modified version of Jordan’s lemma,
it may be shown that this is indeed the case if there exist constants M>0a n d
α>0 such that on Γ
|¯f(s)|≤M
Rα.
Moreover, this condition always holds when ¯f(s)h a st h ef o r m
¯f(s)=P(s)
Q(s),
where P(s)a n d Q(s) are polynomials and the degree of Q(s) is greater than that
ofP(s).
766
20.20 INVERSE LAPLACE TRANSFORM
ΓΓ ΓR R R
L L L
(a)( b)( c)
Figure 20.23 Some contour completions for the integration path Lof the
inverse Laplace transform. For details of when each is appropriate see the
main text.
When the contribution from the part-circle Γ tends to zero as R→∞,w e
have from the residue theorem that the inverse Laplace transform (20.99) is givensimply by
f(t)=summationdisplayparenleftbig
residues of ¯f(s)e
sxat all polesparenrightbig
. (20.101)IFind the function f(x)whose Laplace transform is
¯f(s)=s
s2−k2,
where kis a constant.
It is clear that ¯f(s) is of the form required for the integral over the circular arc Γ to tend
to zero as R→∞, and so we may use the result (20.101). Now
¯f(s)esx=sesx
(s−k)(s+k)
and thus has simple poles at s=kands=−k. Using (20.70) the residues at each pole can
be easily calculated as
R(k)=kekx
2kand R(−k)=ke−kx
2k.
Thus the inverse Laplace transform is given by
f(x)=1
2
/;
ekx+e−kx
/
=c o s h kx.
This result may be checked by computing the forward transform of cosh kx.
J
Sometimes a little more care is required when deciding in which half-plane to
close the contour C.
767
COMPLEX VARIABLESIFind the function f(x)whose Laplace transform is
¯f(s)=1
s(e−as−e−bs),
where aandbare fixed and positive, with b>a.
From (20.99) we have the integral
f(x)=1
2πi
Zλ+i∞
λ−i∞e(x−a)s−e(x−b)s
sds. (20.102)
Now, despite appearances to the contrary, the integrand has no poles, as may be confirmed
by expanding the exponentials as Taylor series about s= 0. Depending on the value of x,
several cases arise.
(i) For x<a both exponentials in the integrand will tend to zero as Re s→∞. Thus
we may close Lwith a circular arc Γ in the righthalf-plane ( λcan be as small as desired),
and we observe that s×integrand tends to zero everywhere on Γ as R→∞.W i t hn o
poles enclosed and no contribution from Γ, the integral along Lmust also be zero. Thus
f(x)=0 f o r x<a . (20.103)
(ii) For x>b the exponentials in the integrand will tend to zero as Re s→−∞ ,a n ds o
we may close Lin the left half-plane, as in figure 20.23( a). Again the integral around Γ
vanishes for infinite Rand so, by the residue theorem,
f(x)=0 f o r x>b . (20.104)
(iii) For a<x<b the two parts of the integrand behave in different ways and have to
be treated separately:
I1−I2≡1
2πi
Z
Le(x−a)s
sds−1
2πi
Z
Le(x−b)s
sds.
The integrand of I1then vanishes in the far left-han d half-plane, but does now have a
(simple) pole at s=0 .C l o s i n g Lin the left half-plane, and using the residue theorem, we
obtain
I1= residue at s=0o f s−1e(x−a)s=1. (20.105)
The integrand of I2, however, vanishes in the far right-hand half-plane (and also has
a simple pole at s= 0) and is evaluated by a circular-arc completion in that half-plane.
Such a contour encloses no poles and leads to I2=0 .
Thus, collecting together results (20.103)–(20.105) we obtain
f(x)=
/8/>/</>/:0f o r x<a,
1f o r a<x<b ,
0f o r x>b,
as shown in figure 20.24.
J
20.21 Exercises
20.1 Find an analytic function of z=x+iywhose imaginary part is
(ycosy+xsiny)e x p x.
768
20.21 EXERCISES
a b1f(x)
x
Figure 20.24 The result of the Laplace inversion of ¯f(s)=s−1(e−as−e−bs)
with b>a.
20.2 Find a function f(z), analytic in a suitable part of the Argand diagram, for which
Ref=sin2x
cosh 2 y−cos 2x.
W h e r ea r et h es i n g u l a r i t i e so f f(z)?
20.3 Find the radii of convergence of the following Taylor series:
(a)∞X
n=2zn
lnn,(b)∞X
n=1n!zn
nn,
(c)∞X
n=1znnlnn,(d)∞X
n=1
/n+p
n
/n2
zn,with preal.
20.4 Find the Taylor series expansion about the origin of the function f(z) defined by
f(z)=∞X
r=1(−1)r+1sin
/pz
r
/
where pis a constant. Hence verify that f(z) is a convergent series for all z.
20.5 Determine the types of singularities (if any) possessed by the following functions
atz=0a n d z=∞:
(a) (z−2)−1,(b) (1 + z3)/z2, (c) sinh(1 /z),
(d)ez/z3, (e)z1/2/(1 +z2)1/2.
20.6 Identify the zeroes, poles and essential singularities of the following functions:
(a) tan z, (b) [( z−2)/z2] sin[1 /(1−z)],(c) exp(1 /z),
(d) tan(1 /z),(e)z2/3.
20.7 Find the real and imaginary parts of the functions (i) z2, (ii)ez, and (iii) cosh πz.
By considering the values taken by these parts on the boundaries of the region0≤x,y≤1, determine the solution of Laplace’s equation in that region that
satisfies the boundary conditions
φ(x,0) = 0 ,φ (0,y)=0 ,
φ(x,1) = x, φ (1,y)=y+s i n πy.
769
COMPLEX VARIABLES
20.8 For the function
f(z)=l n
/z+c
z−c
/
where cis real, show that the real part uoffis constant on a circle of radius
ccosech ucentred on the point z=ccothu. Use this result to show that the
electrical capacitance per unit length of two parallel cylinders of radii a, placed
with their axes 2 dapart, is proportional to [cosh−1(d/a)]−1.
20.9 Find a complex potential in the z-plane appropriate to a physical situation in
which the half-plane x>0,y= 0 has zero potential and the half-plane x<0,
y= 0 has potential V.
By making the transformation w=a(z+z−1)/2, with areal and positive, find
the electrostatic potential associated with the half-plane r>a ,s=0a n dt h e
half-plane r<−a,s= 0 at potentials 0 and Vrespectively.
20.10 By considering in turn the transformations
z=1
2c(w+w−1),w =e x p ζ,
where z=x+iy,w=rexpiθ,ζ=ξ+iηandcis a real positive constant, show
thatz=ccoshζmaps the strip ξ≥0, 0≤η≤2π, onto the whole z-plane. Which
curves in the z-plane correspond to the lines ξ=c o n s t a n ta n d η=c o n s t a n t ?
Identify those corresponding respectively to ξ=0 , η=0a n d η=2π.
The electric potential φof a charged conducting strip −c≤x≤c,y=0 ,
satisfies
φ∼−kln(x2+y2)1/2for large ( x2+y2)1/2,
with φconstant on the strip. Show that φ=R e [−kcosh−1(z/c)] and that the
magnitude of the electric field near the strip is k(c2−x2)−1/2.
20.11 Show that the transformation
w=
Zz
01
(ζ3−ζ)1/2dζ
transforms the upper half-plane into the interior of a square that has one corner
at the origin of the w-plane and sides of length L,w h e r e
L=
Zπ/2
0cosec1/2θd θ .
20.12 The fundamental theorem of algebra states that a complex polynomial pn(z)o f
degree nhas precisely ncomplex roots. By applying Liouville’s theorem (see the
end of section 20.12) to f(z)=1 /pn(z) prove that pn(z) has at least one complex
root. Factor out that root to obtain pn−1(z) and, by repeating the process, prove
the above theorem.
20.13 Show that, if ais a positive real constant, the function exp( iaz2) is analytic and
→0a s|z|→∞ for 0 <argz≤π/4. By applying Cauchy’s theorem to a suitable
contour prove thatZ∞
0cos(ax2)dx=
rπ
8a.
20.14 For the equation 8 z3+z+1=0 :
(a) show that all three roots lie between the circles |z|=3/8a n d|z|=5/8;
(b) find the approximate location of the real root, and hence deduce that the
complex ones lie in the first and fourth quadrants and have moduli greaterthan 0 .5.
770
20.21 EXERCISES
20.15 (a) Prove that z8+3z3+7z+ 5 has two zeroes in the first quadrant.
(b) Find in which quadrants the zeroes of 2 z3+7z2+1 0z+ 6 lie. Try to locate
them.
20.16 The following is a method of determining the number of zeroes of an nth-degree
polynomial f(z) inside the contour Cgiven by|z|=R:
(a) put z=R(1 +it)/(1−it)w i t h t=t a n ( θ/2) in−∞≤ t≤∞;
(b) obtain f(z)a s
A(t)+iB(t)
(1−it)n(1 +it)n
(1 +it)n;
(c) show that arg f(z)=t a n−1(B/A)+ntan−1t;
(d) show that ∆ C[argf(z)] = ∆ C[tan−1(B/A)] +nπ;
(e) using inspection or a sketch graph, determine ∆ C[tan−1(B/A)] by finding the
discontinuities in B/Aand evaluating tan−1(B/A)a tt=±∞.
Use this method, together with the results of the worked example in section 20.15,
to show that the zeroes of z4+z+ 1 in the second and third quadrants have
|z|<1.
20.17 By considering the real part ofZ−izn−1dz
1−a(z+z−1)+a2,
where z=e x p iθandnis a non-negative integer, evaluateZπ
0cosnθ
1−2acosθ+a2dθ,
forareal and >1.
20.18 Prove that if f(z) has a simple pole at z0then 1 /f(z) has residue 1 /f/prime(z0)t h e r e .
Hence evaluateZπ
−πsinθ
a−sinθdθ,
where ais real and >1.
20.19 The equation of an ellipse in plane polar coordinates r,θ, with one of its foci at
the origin, is
l
r=1−/epsilon1cosθ,
where lis a length (that of the latus rectum) and /epsilon1(0</epsilon1< 1) is the eccentricity
of the ellipse. Express the area of the ellipse as an integral around the unit circlein the complex plane, and show that the only singularity of the integrand inside
the circle is a double pole at z
0=/epsilon1−1−(/epsilon1−2−1)1/2.
By setting z=z0+ξand expanding the integrand in powers of ξ, find the
residue at z0and hence show that the area is equal to πl2(1−/epsilon12)−3/2. (In terms
of the semi-axes aandbof the ellipse, l=b2/aand/epsilon12=(a2−b2)/a2.)
20.20 Prove that, for α>0, the integralZ∞
0tsinαt
1+t2dt
has the value ( π/2)exp(−α).
20.21 Prove thatZ∞
0cosmx
4x4+5x2+1dx=π
6
/;
4e−m/2−e−m
/
form>0.
771
COMPLEX VARIABLES
20.22 Show that the principal value of the integralZ∞
−∞cos(x/a)
x2−a2dx
is−(π/a)sin1.
20.23 (a) Prove that the integral of [exp( iπz2)]cosec πzaround the parallelogram with
corners±1/2±Rexp(iπ/4) has the value 2 i.
(b) Show that the parts of the contour parallel to the real axis give no contribu-
tion when R→∞.
(c) Evaluate the integrals along the other two sides by putting z/prime=rexp(iπ/4)
and working in terms of z/prime+1
2andz/prime−1
2. Hence by letting R→∞show thatZ∞
−∞e−πr2dr=1.
20.24 By applying the residue theorem around a wedge-shaped contour of angle 2 π/n,
with one side along the real axis, prove that the integralZ∞
0dx
1+xn,
where nis real and≥2, has the value ( π/n)cosec( π/n).
20.25 Using a suitable cut plane, prove that if αis real and 0 <α< 1t h e nZ∞
0x−α
1+xdx
has the value πcosec πα.
20.26 Show thatZ∞
0lnx
x3/4(1 +x)dx=−√
2π2.
20.27 By integrating a suitable function around a large semicircle in the upper half
plane and a small semicircle centred on the origin, determine the value of
I=
Z∞
0(lnx)2
1+x2dx
and deduce, as a by-product of your calculation, thatZ∞
0lnx
1+x2dx=0.
20.28 Prove that
∞X
−∞1
n2+3
4n+1
8=4π.
Carry out the summation numerically, say between −4 and 4, and note how
much of the sum comes from values near the poles of the contour integration.
20.29 (a) Determine the residues at all the poles of the function
f(z)=πcotπz
a2+z2,
where ais a positive real constant.
(b) By evaluating, in two different ways, the integral Ioff(z) along the straight
line joining −∞− ia/2a n d+∞−ia/2, show that
∞X
n=11
a2+n2=πcothπa
2a−1
2a2.
(c) Deduce the value of
P∞
1n−2.
772
20.22 HINTS AND ANSWERS
20.30 By considering the integral of/sinαz
αz
/2π
sinπz,α <π
2
around a circle of large radius, prove that
∞X
m=1(−1)m−1sin2mα
(mα)2=1
2.
20.31 Use the Bromwich inversion, and contours such as those shown in figure 20.23( a),
to find the functions of which the following are the Laplace transforms:
(a)s(s2+b2)−1;
(b)n!(s−a)−(n+1),w i t h na positive integer and s>a;
(c)a(s2−a2)−1,w i t h s>|a|. (Change variable to t=s−|a|.)
Compare your answers with those given in table 13.1.
20.32 Find the function f(t) whose Laplace transform is
¯f(s)=e−s−1+s
s2.
20.33 A function f(t) has the Laplace transform
F(s)=1
2iln
/s+i
s−i
/
the complex logarithm being defined by a finite branch cut running along the
imaginary axis from −itoi.
(a) Convince yourself that, for t>0,f(t) can be expressed as a closed contour
integral that encloses only the branch cut.
(b) Calculate F(s) on either side of the branch cut, evaluate the integral and
hence determine f(t).
(c) Confirm that the derivative with respect to sof the Laplace transform
integral of your answer is the same as that given by dF/ds .
20.34 Use the contour in figure 20.23( c) to show that the function with Laplace
transform s−1/2is (πx)−1/2. (For an integrand of the form r−1/2exp(−rx) change
variable to t=r1/2.)
20.22 Hints and answers
20.1 ∂u/∂y =−(expx)(ycosy+xsiny+s i n y);zexpz.
20.2 f=( s i n 2 x−isinh2 y)/(cosh2 y−cos 2x); the special case of zreal shows that
f(z)=c o t z; poles at z=nπ.
20.3 (a) 1; (b) 1; (c) 1; (d) e−p.
20.4 The series is given by
a2n+1=(−1)n+1p2n+1
(2n+1 ) !∞X
r=1(−1)r
r2n+1,
for integer n≥0,a2n=0 ; R−1= lim [ p×1×(2n+1 )−1] = 0, and so the series
is convergent by the root test.
20.5 (a) analytic, analytic; (b) double pole, single pole; (c) essential singularity, ana-
lytic; (d) triple pole, essential singularity; (e) branch point, branch point.
773
COMPLEX VARIABLES
20.6 (a) Zeroes at z=nπ, simple poles at z=nπ+π/2, essential singularity at z=∞;
(b) zeroes at z=∞,2a n d1−(nπ)−1, double pole at z= 0, essential singularity
atz=1 ;
(c) zero at∞, essential singularity at z=0 ;
(d) zeroes at z=∞and ( nπ)−1, simple poles at z=(nπ+π/2)−1, essential
singularity at z=0 ;
(e) zero and branch point at the origin, essential singularity at z=∞.
20.7 (i) x2−y2,2xy; (ii) excosy,exsiny; (iii) cosh πxcosπy,sinhπxsinπy;
φ(x, y)=xy+( s i n h πxsinπy)/sinhπ.
20.8 Set ccothu1=−d,ccothu2=+d,|ccosech u|=aand note that the capacitance
is proportional to ( u2−u1)−1.
20.9 f(z)=−i(V/π)lnz;−i(V/π)ln
/
(z/a)±[(z/a)2−1]1/2
/
.
20.10 ξ= constant, ellipses x2(a+1)−2+y2(a−1)−2=c2/(4a2);η= constant, hyperbolae
x2(cosα)−2−y2(sinα)−2=c2. The curves are the cuts −c≤x≤c,y=0 ;a n d
|x|≥c,y= 0. The curves for η=2πare the same as those for η=0 .
20.13 Use a contour bounding the sector 0 ≤argz≤π/4 to establish the relationship
between the required integral and that with exp( −au2) as the integrand.
20.14 (a) |z|=3/8,|8z3+z|≤51/64<1;|z|=5/8,|8z3|= 125 /64>104/64≥|z+1|;
(b) write as 8( z−γ)(z−α−iβ)(z−α+iβ)=0 , γ<0, and then the zero coefficient
ofz2shows α>0. Show−3/8>γ>−1/2 and use−8γ(α2+β2)=1 .
20.15 (a) For a quarter-circular contour enclosing the first quadrant, the change in the
argument of the function is 0 + 8( π/2) + 0 (since y8+ 5 = 0 has no real roots);
(b) one negative real zero; a conjugate pair in the second and third quadrants,−
3
2,−1±i.
20.16 A=3−12t2+t4,B=−2t−2t3,∆C[tan−1(B/A)] = 0, ∆ C[argf(z)] = 4 π; hence
there are two zeroes inside |z|=1 .
20.17 Pole at z=1/a;πa−n(a2−1)−1.
20.18 The only pole inside the unit circle is at z=ia−i(a2−1)1/2; the residue is given
by−(i/2)(a2−1)−1/2; the integral has value 2 π[a(a2−1)−1/2−1].
20.19 The integrand is 2 l2z(2z−/epsilon1z2−/epsilon1)−2; residue = (4 /epsilon1)−1(/epsilon1−2−1)−3/2.
20.20 Follow the first example in section 20.17 and use Jordan’s lemma, pole at z=i.
20.21 Factorise the denominator, showing that the relevant simple poles are at i/2a n d i.
20.22 Use Jordan’s lemma and a semicircular contour indented at z=±a.
20.23 (a) The only pole is at the origin with residue π−1; (b) each is O[exp( −πR2−
πR/√
2)]; (c) the sum of the integrals is 2 i
RR
−Rexp(−πr2)dr.
20.24 The residue at the only pole inside the contour, z=e x p ( iπ/n)i s−n−1exp(iπ/n).
The values of the integrals along the two radii differ by a factor −exp(2 πi/n).
20.25 Use a contour like that shown in figure 20.21.20.26 See the previous example.20.27 Note that ρln
nρ→0a s ρ→0f o ra l l n.W h e n zis on the negative real axis,
(lnz)2contains three terms; one of the corresponding integrals is a standard
form. The residue at z=iisiπ2/8;I=π3/8.
20.28 EvaluateZπcotπz/;1
2+z
//;1
4+z
/dz
around a large circle centred on the origin; residue at z=−1/2 is 0; residue at
z=−1/4i s4 πcot(−π/4).
20.29 (a) ( a2+n2)−1atz=n(integer);−πcoth( πa)/(2a)a tz=±ia.
(b) Complete the contour separately in the upper half-plane (including all the
poles on the real axis and the one at z=ia), and in the lower half plane
(including only the pole at z=−ia). Equate the two expressions for I.
(c) Take the limit as a→0, using l’H ˆopital’s rule to give π2/6.
774
20.22 HINTS AND ANSWERS
20.30 The behaviour of the integrand for large |z|is|z|−2exp[(2 α−π)|z|]. The residue
atz=±m,f o re a c hi n t e g e r m,i ss i n2(mα)(−1)m/(mα)2. The contour contributes
nothing. Required summation = [total sum −(m= 0 term)]/2.
20.31 Poles at (a) ±ib;( b ) t=s−a=0 ,o fo r d e r n+1 ,( c ) t=0a n d t=−2|a|.S e e
table 13.1.
20.32 Note that ¯f(s) has no pole at s=0 .F o r t<0 close the Bromwich contour in the
right half-plane, and for t>1 in the left half-plane. For 0 <t< 1 the integrand
has to be split into separate terms containing e−sands−1 and the completions
made in the right and left half-planes respectively. The last of these completedcontours now contains a second-order pole at s=0 . f(t)=1−tfor 0 <t< 1
but is 0 otherwise.
20.33 (a) Note that F(s) has no singularities in Re s<0 and apply Cauchy’s theorem
to reshape the Bromwich contour.
(b) The real parts of F(s)d i ff e rb y πon either side of the branch cut;
f(t)=s i n t/t.
(c) Both are −1/(1 +s
2).
20.34
R
Γand
R
γtend to 0 as R→∞andρ→0. Put s=rexpiπands=rexp(−iπ)o n
the two sides of the cut and use
R∞
0exp(−t2x)dt=1
2(π/x)1/2. There are no poles
inside the contour.
775
21
Tensors
It may seem obvious that the quantitative description of physical processes cannot
depend on the coordinate system in which they are represented. However, we mayturn this argument around: since physical results must indeed be independent of
the choice of coordinate system, what does this imply about the nature of the
quantities involved in the description of physical processes? The study of theseimplications and of the classification of physical quantities by means of themforms the content of the present chapter.
Although the concepts presented here may be applied, with little modifi-
cation, to more abstract spaces (most notably the four-dimensional space–time ofspecial or general relativity), we shall restrict our attention to our familiar three-dimensional Euclidean space. This removes the need to discuss the properties ofdifferentiable manifolds and their tangent and dual spaces. The reader who isinterested in these more technical aspects of tensor calculus in general spaces,and in particular their application to general relativity, should consult one of the
many excellent textbooks on the subject. †
Before the presentation of the main development of the subject, we begin by
introducing the summation convention, which will prove very useful in writingtensor equations in a more compact form. We then review the effects of a change
of basis in a vector space; such spaces were discussed in chapter 8. This is
followed by an investigation of the rotation of Cartesian coordinate systems, andfinally we broaden our discussion to include more general coordinate systems andtransformations.
†For example, D’Inverno, Introducing Einstein’s Relativity (Oxford, 1992); Foster and Nightingale,
A Short Course in General Relativity (Springer-Verlag, 1994); Schutz, A First Course in General
Relativity (Cambridge, 1990).
776
21.1 SOME NOTATION
21.1 Some notation
Before proceeding further, we introduce the summation convention for subscripts,
since its use looms large in the work of this chapter. The convention is that
anylower-case alphabetic subscript that appears exactly twice in any term of an
expression is understood to be summed over all the values that a subscript inthat position can take (unless the contrary is specifically stated). The subscriptedquantities may appear in the numerator and/or the denominator of a term in anexpression. This naturally implies that any such pair of repeated subscripts mustoccur only in subscript positions that have the same range of values. Sometimesthe ranges of values have to be specified but usually they are apparent from the
context.
The following simple examples illustrate what is meant (in the three-dimensional
case):
(i)a
ixistands for a1x1+a2x2+a3x3;
(ii)aijbjkstands for ai1b1k+ai2b2k+ai3b3k;
(iii)aijbjkckstands forsummationtext3
j=1summationtext3
k=1aijbjkck;
(iv)∂vi
∂xistands for∂v1
∂x1+∂v2
∂x2+∂v3
∂x3;
(v)∂2φ
∂xi∂xistands for∂2φ
∂x2
1+∂2φ
∂x2
2+∂2φ
∂x2
3.
Subscripts that are summed over are called dummy subscripts and the others
free subscripts . It is worth remarking that when introducing a dummy subscript
into an expression, care should be taken not to use one that is already present,either as a free or as a dummy subscript. For example, a
ijbjkcklcannot, and must
not, be replaced by aijbjjcjlor by ailblkckl, but could be replaced by aimbmkckl
or by aimbmncnl. Naturally, free subscripts must not be changed at all unless the
working calls for it.
Furthermore, as we have done throughout this book, we will make frequent
use of the Kronecker delta δij, which is defined by
δij=braceleftBigg
1i f i=j,
0o t h e r w i s e .
When the summation convention has been adopted, the main use of δijis to
replace one subscript by another in certain expressions. Examples might include
bjδij=bi,
and
aijδjk=aijδkj=aik. (21.1)
777
TENSORS
In the second of these the dummy index shared by both terms on the left-hand
side (namely j) has been replaced by the free index carried by the Kronecker delta
(namely k), and the delta symbol has disappeared. In matrix language, (21.1) can
be written as AI=A,w h e r e Ais the matrix with elements aijand Iis the unit
matrix having the same dimensions as A.
In some expressions we may use the Kronecker delta to replace indices in a
number of different ways, e.g.
aijbjkδki=aijbjior akjbjk,
where the two expressions on the RHS are totally equivalent to one another.
21.2 Change of basis
In chapter 8 some attention was given to the subject of changing the basis set (or
coordinate system) in a vector space and it was shown that, under such a change,different types of quantity behave in different ways. These results are given insection 8.15, but are summarised below for convenience, using the summationconvention. Although throughout this section we will remind the reader that weare using this convention, it will simply be assumed in the remainder of the
chapter.
If we introduce a set of basis vectors e
1,e2,e3into our familiar three-dimensional
(vector) space, then we can describe any vector xin terms of its components
x1,x2,x3with respect to this basis:
x=x1e1+x2e2+x3e3=xiei,
where we have used the summation convention to write the sum in a more
compact form. If we now introduce a new basis e/prime
1,e/prime
2,e/prime
3related to the old one by
e/prime
j=Sijei(sum over i), (21.2)
where the coefficient Sijis the ith component of the vector e/prime
jwith respect to the
unprimed basis, then we may write xwith respect to the new basis as
x=x/prime
1e/prime1+x/prime
2e/prime2+x/prime
3e/prime3=x/prime
ie/primei(sum over i).
If we denote the matrix with elements SijbyS, then the components x/prime
iandxi
in the two bases are related by
x/prime
i=(S−1)ijxj (sum over j),
where, using the summation convention, there is an implicit sum over jfrom
j=1t o j= 3. In the special case where the transformation is a rotation of the
coordinate axes, the transformation matrix Sis orthogonal and we have
x/prime
i=(ST)ijxj=Sjixj(sum over j). (21.3)
778
21.3 CARTESIAN TENSORS
Scalars behave differently under transformations, however, since they remain
unchanged. For example, the value of the scalar product of two vectors x·y
(which is just a number) is unaffected by the transformation from the unprimedto the primed basis. Different again is the behaviour of linear operators. If a
linear operator Ais represented by some matrix Ain a given coordinate system
then in the new (primed) coordinate system it is represented by a new matrixA
/prime=S−1AS.
In this chapter we develop a general formulation to describe and classify these
different types of behaviour under a change of basis (or coordinate transfor-mation). In the development, the generic name tensor is introduced, and certain
scalars, vectors and linear operators are described respectively as tensors of ze-
roth, first and second order (the order –o rrank– corresponds to the number of
subscripts needed to specify a particular element of the tensor). Tensors of thirdand fourth order will also occupy some of our attention.
21.3 Cartesian tensors
We begin our discussion of tensors by considering a particular class of coordinate
transformation – namely rotations – and we shall confine our attention strictly
to the rotation of Cartesian coordinate systems. Our object is to study the prop-erties of various types of mathematical quantities, and their associated physicalinterpretations, when they are described in terms of Cartesian coordinates andthe axes of the coordinate system are rigidly rotated from a basis e
1,e2,e3(lying
along the Ox1,Ox2andOx3axes) to a new one e/prime
1,e/prime
2,e/prime
3(lying along the Ox/prime
1,
Ox/prime
2andOx/prime
3axes).
Since we shall be more interested in how the components of a vector or linear
operator are changed by a rotation of the axes than in the relationship betweenthe two sets of basis vectors e
iande/prime
i, let us define the transformation matrix L
as the inverse of the matrix Sin (21.2). Thus, from (21.2), the components of a
position vector x, in the old and new bases respectively, are related by
x/prime
i=Lijxj. (21.4)
Because we are considering only rigid rotations of the coordinate axes, the
transformation matrix Lwill be orthogonal, i.e. such that L−1=LT. Therefore
the inverse transformation is given by
xi=Ljix/prime
j. (21.5)
The orthogonality of Lalso implies relations among the elements of Lthat
express the fact that LLT=LTL=I. In subscript notation they are given by
LikLjk=δij and LkiLkj=δij. (21.6)
Furthermore, in terms of the basis vectors of the primed and unprimed Cartesian
779
TENSORS
O x1x2
x/prime
1x/prime
2
θ
θ
θ
Figure 21.1 Rotation of Cartesian axes by an angle θabout the x3-axis. The
three angles marked θand the parallels (broken lines) to the primed axes
show how the first two equations of (21.7) are constructed.
coordinate systems, the transformation matrix is given by
Lij=e/prime
i·ej.
We note that the product of two rotations is also a rotation. For example,
suppose that x/prime
i=Lijxjandx/prime/prime
i=Mijx/prime
j; then the composite rotation is described
by
x/prime/prime
i=Mijx/prime
j=MijLjkxk=(ML)ikxk,
corresponding to the matrix ML.IFind the transformation matrix Lcorresponding to a rotation of the coordinate axes
through an angle θabout the e3-axis (or x3-axis), as shown in figure 21.1.
Taking xas a position vector – the most obvious choice – we see from the figure that
the components of xwith respect to the new (primed) basis are given in terms of the
components in the old (unprimed) basis by
x/prime
1=x1cosθ+x2sinθ,
x/prime
2=−x1sinθ+x2cosθ, (21.7)
x/prime
3=x3.
The (orthogonal) transformation matrix is thus
L=
/0/@cosθsinθ0
−sinθcosθ0
00 1
/1A.
The inverse equations are
x1=x/prime
1cosθ−x/prime
2sinθ,
x2=x/prime
1sinθ+x/prime
2cosθ, (21.8)
x3=x/prime
3,
in line with (21.5).
J
780
21.4 FIRST- AND ZERO-ORDER CARTESIAN TENSORS
21.4 First- and zero-order Cartesian tensors
Using the above example as a guide, we may consider any set of three quantities
vi, which are directly or indirectly functions of the coordinates xiand possibly
involve some constants, and ask how their values are changed by any rotation ofthe Cartesian axes. The specific question to be answered is whether the specificforms v
/prime
iin the new variables can be obtained from the old ones viusing (21.4),
v/prime
i=Lijvj. (21.9)
If so, the viare said to form the components of a vector orfirst-order Cartesian
tensor v. The first-order tensor vdoes not change under the rotation of the
coordinate axes; nevertheless, since the basis set does change, from e1,e2,e3to
e/prime
1,e/prime
2,v/prime
3, the components of vmust also change. The changes must be such that
vgiven by
v=viei=v/prime
ie/primei(21.10)
is unchanged. By definition, the position coordinates are themselves the compo-
nents of such a tensor.
Since the transformation (21.9) is orthogonal, the components of any such
first-order Cartesian tensor obey a relation that is the inverse of (21.9),
vi=Ljiv/prime
j. (21.11)
We now consider explicit examples. In order to keep the equations to reasonable
proportions, the examples will be restricted to the x1x2-plane, i.e. there are
no components in the x3-direction. Three-dimensional cases are no different in
principle – but much longer to write out.IWhich of the following pairs (v1,v2)form the components of a first-order Cartesian tensor
in two dimensions?:
(i)(x2,−x1), (ii)(x2,x1), (iii)(x2
1,x22).
We shall consider the rotation discussed in the previous example, and to save space we
denote cos θbycand sin θbys.
(i) Here v1=x2andv2=−x1, referred to the old axes. In terms of the new coordinates
they will be v/prime
1=x/prime
2andv/prime
2=−x/prime
1,i . e .
v/prime
1=x/prime
2=−sx1+cx2
v/prime
2=−x/prime
1=−cx1−sx2.(21.12)
Now if we start again and evaluate v/prime
1andv/prime
2as given by (21.9) we find that
v/prime
1=L11v1+L12v2=cx2+s(−x1)
v/prime
2=L21v1+L22v2=−s(x2)+c(−x1).(21.13)
The expressions for v/prime
1andv/prime
2in (21.12) and (21.13) are the same whatever the values
ofθ(i.e. for allrotations) and thus by de finition (21.9) the pair ( x2,−x1)isa first-order
Cartesian tensor.
781
TENSORS
(ii) Here v1=x2andv2=x1. Following the same procedure,
v/prime
1=x/prime
2=−sx1+cx2
v/prime
2=x/prime
1=cx1+sx2.
But, by (21.9), for a Cartesian tensor we must have
v/prime
1=cv1+sv2=cx2+sx1
v/prime
2=(−s)v1+cv2=−sx2+cx1.
These two sets of expressions do not agree and thus the pair ( x2,x1) is not a first-order
Cartesian tensor.
(iii)v1=x2
1andv2=x2
2. As in (ii) above, considering the first component alone is
sufficient to show that this pair is nota first-order tensor. Evaluating v/prime
1directly gives
v/prime
1=x/prime
12=c2x2
1+2csx1x2+s2x2
2,
whilst (21.9) requires that
v/prime
1=cv1+sv2=cx2
1+sx2
2,
which is quite different.
J
There are many physical examples of first-order tensors (i.e. vectors) that will be
familiar to the reader. As a straightforward one, we may take the set of Cartesian
components of the momentum of a particle of mass m,(m˙x1,m˙x2,m˙x3). This set
transforms in all essentials as ( x1,x2,x3), since the other operations involved,
multiplication by a number and differentiation with respect to time, are quiteunaffected by any orthogonal transformation of the axes. Similarly, accelerationand force are represented by the components of first-order tensors.
Other more complicated vectors involving the position coordinates more than
once, such as the angular momentum of a particle of mass m,n a m e l y J=
x×p=m(x×˙x), are also first-order tensors. That this is so is less obvious in
component form than for the earlier examples, but may be verified by writing
out the components of Jexplicitly or by appealing to the quotient law to be
discussed in section 21.7 and using the Cartesian tensor /epsilon1
ijkfrom section 21.8.
Having considered the effects of rotatio ns on vector-like sets of quantities we
may consider quantities that are unchanged by a rotation of axes. In our previousnomenclature these have been called scalars but we may also describe them as
tensors of zero order . They contain only one element (formally, the number of
subscripts needed to identify a particular element is zero); the most obvious non-trivial example associated with a rotation of axes is the square of the distance of
a point from the origin, r
2=x2
1+x2
2+x2
3. In the new coordinate system it will
have the form r/prime2=x/prime
12+x/prime
22+x/prime
32, which for any rotation has the same value as
x2
1+x2
2+x2
3.
782
21.4 FIRST- AND ZERO-ORDER CARTESIAN TENSORS
In fact any scalar product of two first-order tensors (vectors) is a zero-order
tensor (scalar), as might be expected since it can be written in a coordinate-freeway as u·v.IBy considering the components of the vectors uandvwith respect to two Cartesian
coordinate systems (related by a rotation), show that the scalar product u·vis invariant
under rotation.
In the original (unprimed) system the scalar product is given in terms of components byu
ivi(summed over i), and in the rotated (primed) system by
u/prime
iv/prime
i=LijujLikvk=LijLikujvk=δjkujvk=ujvj,
where we have used the orthogonality relation (21.6). Since the resulting expression in the
rotated system is the same as that in the original system, the scalar product is indeedinvariant under rotations.J
The above result leads directly to the identification of many physically im-
portant quantities as zero-order tensors. Perhaps the most immediate of these isenergy, either as potential energy or as an energy density (e.g. F·dr,eE·dr,D·E,
B·H,µ·B), but others, such as the angle between two directed quantities, are
important.
As mentioned in the first paragraph of this chapter, in most analyses of physical
situations it is a scalar quantity (such as energy) that is to be determined. Such
quantities are invariant under a rotation of axes and so it is possible to work with
the most convenient set of axes and still have confidence in the results.
Complementing the way in which a zero-order tensor was obtained from two
first-order tensors, so a first-order tensor can be obtained from a zero-order
tensor. We show this by taking a specific example, that of the electric field
E=−∇φ; this is derived from a scalar, the electrostatic potential φand has
components
E
i=−∂φ
∂xi. (21.14)
Clearly, Eisa first-order tensor, but we may prove this more formally by
considering the behaviour of its components (21.14) under a rotation of thecoordinate axes, since the components of the electric field E
/prime
iare then given
by
E/prime
i=parenleftbigg
−∂φ
∂xiparenrightbigg/prime
=−∂φ/prime
∂x/prime
i=−∂xj
∂x/prime
i∂φ
∂xj=LijEj, (21.15)
where (21.5) has been used to evaluate ∂xj/∂x/prime
i. Now (21.15) is in the form
(21.9), thus confirming that the components of the electric field do behave as thecomponents of a first-order tensor.
783
TENSORSIIfviare the components of a first-order tensor, show that ∇·v=∂vi/∂x iis a zero-order
tensor.
In the rotated coordinate system ∇·vis given by/∂vi
∂xi
//prime
=∂v/prime
i
∂x/primei=∂xj
∂x/prime
i∂
∂xj(Likvk)=LijLik∂vk
∂xj,
since the elements Lijare not functions of position. Using the orthogonality relation (21.6)
we then find
∂v/prime
i
∂x/primei=LijLik∂vk
∂xj=δjk∂vk
∂xj=∂vj
∂xj.
Hence ∂vi/∂x iis invariant under rotation of the axes and is thus a zero-order tensor; this
was to be expected since it can be written in a coordinate-free way as ∇·v.
J
21.5 Second- and higher-order Cartesian tensors
Following on from scalars with no subscripts and vectors with one subscript,
we turn to sets of quantities that require two subscripts to identify a particularelement of the set. Let these quantities by denoted by T
ij.
Taking (21.9) as a guide we define a second-order Cartesian tensor as follows:
theTijform the components of such a tensor if, under the same conditions as
for (21.9),
T/prime
ij=LikLjlTkl (21.16)
and
Tij=LkiLljT/prime
kl. (21.17)
At the same time we may define a Cartesian tensor of general order as follows.
The set of expressions Tij···kform the components of a Cartesian tensor if, for all
rotations of the axes of coordinates given by (21.4) and (21.5), subject to (21.6),the expressions using the new coordinates, T
/prime
ij···kare given by
T/prime
ij···k=LipLjq···LkrTpq···r (21.18)
and
Tij···k=LpiLqj···LrkT/prime
pq···r. (21.19)
It is apparent that in three dimensions, an Nth-order Cartesian tensor has 3N
components.
Since a second-order tensor has two subscripts, it is natural to display its
components in matrix form. The notation [ Tij] is used, as well as T, to denote
the matrix having Tijas the element in the ith row and jth column.†
We may think of a second-order tensor Tas a geometrical entity in a similar
way to that in which we viewed linear operators (which transform one vector into
†We can also denote the column matrix containing the elements viof a vector by [ vi].
784
21.5 SECOND- AND HIGHER-ORDER CARTESIAN TENSORS
another, without reference to any coordinate system) and consider the matrix
containing its components as a representation of the tensor with respect to aparticular coordinate system. Moreover, the matrix T=[T
ij], containing the
components of a second-order tensor, behaves in the same way under orthogonal
transformations T/prime=LTLTas a linear operator.
However, not all linear operators are second-order tensors. More specifically, we
require the two subscripts in a second-order tensor to refer to the same coordinatesystem, so that only linear operators that transform a vector to another vector inthe same vector space could be second-order tensors. Thus, although the elementsL
ijof the transformation matrix are written with two subscripts, they cannot
be the components of a tensor since the two subscripts each refer to a different
coordinate system.
As examples of sets of quantities that are readily shown to be second-order
tensors we consider the following.
(i)The outer product of two vectors .L e t uiandvi,i=1,2,3, be the components
of two vectors uandv, and consider the set of quantities Tijdefined by
Tij=uivj. (21.20)
The set Tijare called the components of the the outer product ofuandv. Under
rotations the components Tijbecome
T/prime
ij=u/prime
iv/prime
j=LikukLjlvl=LikLjlukvl=LikLjlTkl, (21.21)
which shows that they do transform as the components of a second-order tensor.
Use has been made in (21.21) of the fact that uiandviare the components of
first-order tensors.
The outer product of two vectors is often denoted, without reference to any
coordinate system, as
T=u⊗v. (21.22)
(This is not to be confused with the vector product of two vectors, which is itself
a vector and is discussed in chapter 7.) The expression (21.22) gives the basis towhich the components T
ijof the second-order tensor refer. Since u=uieiand
v=viei, we may write the tensor Tas
T=uiei⊗vjej=uivjei⊗ej=Tijei⊗ej. (21.23)
Moreover, as for the case of first-order tensors (see equation(21.10)) we note
that the quantities T/prime
ijare the components of the sametensor T, but referred to
a different coordinate system, i.e.
T=Tijei⊗ej=T/prime
ije/primei⊗e/prime
j.
These concepts can be extended to higher-order tensors.
(ii)The gradient of a vector. Suppose virepresents the components of a vector;
785
TENSORS
let us consider the quantities generated by forming the derivatives of each vi,
i=1,2,3, with respect to each xj,j=1,2,3, i.e.
Tij=∂vi
∂xj.
These nine quantities form the components of a second-order tensor, as can be
seen from the fact that
T/prime
ij=∂v/prime
i
∂x/primej=∂(Likvk)
∂xl∂xl
∂x/prime
j=Lik∂vk
∂xlLjl=LikLjlTkl.
In coordinate-free language the tensor Tmay be written as T=∇vand hence
gives meaning to the concept of the gradient of a vector, a quantity that was not
discussed in the chapter on vector calculus (chapter 10).
A test of whether any given set of quantities forms the components of a second-
order tensor can always be made by direct substitution of the x/prime
iin terms of the
xi, followed by comparison with the right-hand side of (21.16). This procedure is
extremely laborious, however, and it is almost always better to try to recognisethe set as being expressible in one of the forms just considered, or to make
alternative tests based on the quotient law of section 21.7 below.IShow that the Tijgiven by
T=[Tij]=
/
x2
2−x1x2
−x1x2 x2
1
/
(21.24)
are the components of a second-order tensor.
Again we consider a rotation θabout the e3-axis. Carrying out the direct evaluation first
we obtain, using (21.7),
T/prime
11=x/prime
22=s2x2
1−2scx1x2+c2x2
2,
T/prime
12=−x/prime
1x/prime2=scx2
1+(s2−c2)x1x2−scx2
2,
T/prime
21=−x/prime
1x/prime2=scx2
1+(s2−c2)x1x2−scx2
2,
T/prime
22=x/prime
12=c2x2
1+2scx1x2+s2x2
2.
Now, evaluating the right-hand side of (21.16),
T/prime
11=ccx2
2+cs(−x1x2)+sc(−x1x2)+ssx2
1,
T/prime
12=c(−s)x2
2+cc(−x1x2)+s(−s)(−x1x2)+scx2
1,
T/prime
21=(−s)cx2
2+(−s)s(−x1x2)+cc(−x1x2)+csx2
1,
T/prime
22=(−s)(−s)x2
2+(−s)c(−x1x2)+c(−s)(−x1x2)+ccx2
1.
After reorganisation, the corresponding expressions are seen to be the same, showing, as
required, that the Tijare the components of a second-order tensor.
The same result could be inferred much more easily, however, by noting that the Tij
are in fact the components of the outer product of the vector ( x2,−x1) with itself. That
(x2,−x1) is indeed a vector was established by (21.12) and (21.13).
J
Physical examples involving second-order tensors will be discussed in the later
sections of this chapter, but we might note here that, for example, the magnetic
786
21.6 THE ALGEBRA OF TENSORS
susceptibility and electrical conductivity of materials are described by second-
order tensors.
21.6 The algebra of tensors
Because of the similarity of first- and second-order tensors to column vectors and
matrices, it would be expected that similar types of algebraic operation can be
carried out with them and so provide ways of constructing new tensors from oldones. In the remainder of this chapter, instead of referring to the T
ij(say) as the
components of a second-order tensor T, we may sometimes simply refer to Tij
as the tensor. It should always be remembered, however, that the Tijare in fact
just the components of Tin a given coordinate system and that T/prime
ijrefers to the
components of the sametensor Tin a different coordinate system.
The addition and subtraction of tensors follows an obvious definition; namely
that if Vij···kandWij···kare (the components of) tensors of the same order, then
their sum and difference, Sij···kandDij···krespectively, are given by
Sij···k=Vij···k+Wij···k,
Dij···k=Vij···k−Wij···k,
for each set of values i ,j,...,k .T h a t Sij···kandDij···kare the components of
tensors follows immediately from the linearity of a rotation of coordinates.
It is equally straightforward to show that if the Tij···kare the components of
a tensor, then so is the set of quantities formed by interchanging the order of (apair of) indices, e.g. T
ji···k.
IfTji···kis found to be identical with Tij···kthen Tij···kis said to be symmetric
with respect to its first two subscripts (or simply ‘symmetric’, for second-ordertensors). If, however, T
ji···k=−Tij···kfor every element then it is an antisymmetric
tensor. An arbitrary tensor is neither symmetric nor antisymmetric but can alwaysbe written as the sum of a symmetric tensor S
ij···kand an antisymmetric tensor
Aij···k:
Tij···k=1
2(Tij···k+Tji···k)+1
2(Tij···k−Tji···k)
=Sij···k+Aij···k.
Of course these properties are valid for any pair of subscripts.
In (21.20) in the previous section we had an example of a kind of ‘multiplication’
of two tensors, thereby producing a tensor of higher order – in that case two
first-order tensors were multiplied to give a second-order tensor. Inspection of
(21.21) shows that there is nothing particular about the orders of the tensorsinvolved and it follows as a general result that the outer product of an Nth-order
tensor with an Mth-order tensor will produce an ( M+N)th-order tensor.
An operation that produces the opposite effect – namely, generates a tensor
787
TENSORS
of smaller rather than larger order – is known as contraction and consists of
making two of the subscripts equal and summing over all values of the equalisedsubscripts.IShow that the process of contraction of a tensor produces another tensor, but with an
order reduced by 2.
LetTij···l···m···kbe the components of an Nth-order tensor, then
T/prime
ij···l···m···k=LipLjq···Llr···Lms···Lkn/| /{z /}
NfactorsTpq···r···s···n.
Thus if, for example, we make the two subscripts landmequal and sum over all values
of these subscripts, we obtain
T/prime
ij···l···l···k=LipLjq···Llr···Lls···LknTpq···r···s···n
=LipLjq···δrs···LknTpq···r···s···n
=LipLjq···Lkn/| /{z /}
(N−2) factorsTpq···r···r···n,
showing that Tij···l···l···kare the components of a (different) Cartesian tensor of order
N−2.
J
For a second-rank tensor, the process of contraction is the same as taking the
trace of the corresponding matrix. The trace Tiiitself is thus a zero-order tensor
(or scalar) and hence invariant under rotations, as was noted in chapter 8.
The process of taking the scalar product of two vectors can be recast into tensor
language as forming the outer product Tij=uivjof two first-order tensors uand
vand then contracting the second-order tensor Tso formed, to give Tii=uivi,a
scalar (invariant under a rotation of axes).
As yet another example of a familiar operation that is a particular case of a
contraction, we may note that the multiplication of a column vector [ ui]b ya
matrix [ Bij] to produce another column vector [ vi],
Bijuj=vi,
can be looked upon as the contraction Tijjof the third-order tensor Tijkformed
from the outer product of Bijanduk.
21.7 The quotient law
The previous paragraph appears to give a heavy-handed way of describing a
familiar operation, but it leads us to ask whether it has a converse. To put thequestion in more general terms: if we know that BandCare tensors and also
that
A
pq···k···mBij···k···n=Cpq···mij···n, (21.25)
788
21.7 THE QUOTIENT LAW
does this imply that the Apq···k···malso form the components of a tensor A?H e r e
A,BandCare respectively of Mth,Nth and ( M+N−2)th order and it should be
noted that the subscript kthat has been contracted may be any of the subscripts
inAandBindependently.
Thequotient law for tensors states that if (21.25) holds in all rotated coordinate
frames then the Apq···k···mdo indeed form the components of a tensor A.T op r o v e
it for general MandNis no more difficult regarding the ideas involved than to
show it for specific MandN, but this does involve the introduction of a large
number of subscript symbols. We will therefore take the case M=N= 2, but
it will be readily apparent that the principle of the proof holds for general M
andN.
We thus start with (say)
ApkBik=Cpi, (21.26)
where BikandCpiare arbitrary second-order tensors. Under a rotation of coor-
dinates the set Apk(tensor or not) transforms into a new set of quantities that
we will denote by A/prime
pk. We thus obtain in succession the following steps, using
(21.16), (21.17) and (21.6):
A/prime
pkB/prime
ik=C/prime
pi (transforming (21.26)),
=LpqLijCqj (since Cis a tensor),
=LpqLijAqlBjl (from (21.26)),
=LpqLijAqlLmjLnlB/prime
mn(since Bis a tensor),
=LpqLnlAqlB/prime
in (since LijLmj=δim).
Now kon the left and non the right are dummy subscripts and thus we may
write
(A/prime
pk−LpqLklAql)B/prime
ik=0. (21.27)
Since Bik, and hence B/prime
ik, is an arbitrary tensor, we must have
A/prime
pk=LpqLklAql,
showing that the A/prime
pkare given by the general formula (21.18) and hence that
theApkare the components of a second-order tensor. By following an analogous
argument, the same result (21.27) and deduction could be obtained if (21.26) were
replaced by
ApkBki=Cpi,
i.e. the contraction being now with respect to a different pair of indices.
Use of the quotient law to test whether a given set of quantities is a tensor is
generally much more convenient than making a direct substitution. A particularway in which it is applied is by contracting the given set of quantities, having
789
TENSORS
Nsubscripts, with an arbitrary Nth-order tensor (i.e. one having independently
variable components) and determining whether the result is a scalar.IUse the quotient law to show that the elements of T, equation (21.24), are the components
of a second-order tensor.
The outer product xixjis a second-order tensor. Contracting this with the Tijgiven in
(21.24) we obtain
Tijxixj=x2
2x21−x1x2x1x2−x1x2x2x1+x2
1x22=0,
which is clearly invariant (a zeroth-order tensor). Hence by the quotient theorem Tijmust
also be a tensor.
J
21.8 The tensors δijand/epsilon1ijk
In many places throughout this book we have encountered and used the two-
subscript quantity δijdefined by
δij=braceleftBigg
1i f i=j,
0o t h e r w i s e .
Let us now also introduce the three-subscript Levi–Civita symbol /epsilon1ijk, the value
of which is given by
/epsilon1ijk=
+1 if i, j, kis an even permutation of 1 ,2,3,
−1i f i, j, kis an odd permutation of 1 ,2,3,
0o t h e r w i s e .
We will now show that δ
ijand/epsilon1ijkare respectively the components of a second-
and a third-order Cartesian tensor. Notice that the coordinates xido not appear
explicitly in the components of these tensors, their components consisting entirely
of 0 and 1.
In passing, we also note that /epsilon1ijkis totally antisymmetric, i.e. it changes sign
under the interchange of any pair of subscripts. In fact /epsilon1ijk, or any scalar multiple
of it, is the onlythree-subscript quantity with this property.
Treating δijfirst, the proof that it is a second-order tensor is straightforward
since, if, from (21.16), we consider the equation
δ/prime
kl=LkiLljδij=LkiLli=δkl,
we see that the transformation generates the same expression (a pattern of
0’s and 1’s) as does the definition of δ/prime
ijin the transformed coordinates. Thus
δijtransforms according to the appropriate tensor transformation law and is
therefore a second-order tensor.
Turning now to /epsilon1ijk, we have to consider the quantity
/epsilon1/prime
lmn=LliLmjLnk/epsilon1ijk.
790
21.8 THE TENSORS δijAND /epsilon1ijk
Let us begin, however, by noting that we may use the Levi–Civita symbol to
write an expression for the determinant of a 3 ×3m a t r i x A,
|A|/epsilon1lmn=AliAmjAnk/epsilon1ijk, (21.28)
which may be shown to be equivalent to the Laplace expansion (see chapter 8). †
Indeed many of the properties of determinants discussed in chapter 8 can be
proved very efficiently using this expression (see exercise 21.9).IEvaluate the determinant of the matrix
A=
/0/@21−3
34 01−21
/1A.
Setting l=1 , m=2a n d n= 3 in (21.28) we find
|A|=/epsilon1ijkA1iA2jA3k
= (2)(4)(1)−(2)(0)(−2)−(1)(3)(1) + ( −3)(3)(−2)
+ (1)(0)(1)−(−3)(4)(1) = 35 ,
which may be verified using the Laplace expansion method.
J
We can now show that the /epsilon1ijkare in fact the components of a third-order tensor.
Using (21.28) with the general matrix Areplaced by the specific transformation
matrix L, we can rewrite the RHS of (21.8)in terms of |L|
/epsilon1/prime
lmn=LliLmjLnk/epsilon1ijk=|L|/epsilon1lmn.
Since Lis orthogonal its determinant has the value unity, and so /epsilon1/prime
lmn=/epsilon1lmn.
Thus we see that /epsilon1/prime
lmnhas exactly the properties of /epsilon1ijkbut with i, j, kreplaced by
l,m,n, i.e. it is the same as the expression /epsilon1ijkwritten using the new coordinates.
This shows that /epsilon1ijkis a third-order Cartesian tensor.
In addition to providing a convenient notation for the determinant of a matrix,
δijand /epsilon1ijkcan be used to write many of the familiar expressions of vector
algebra and calculus as contracted tensors. For example, provided we are using
right-handed Cartesian coordinates, the vector product a=b×chas as its
ith component ai=/epsilon1ijkbjck; this should be contrasted with the outer product
T=b⊗c, which is a second-order tensor having the components Tij=bicj.
†This may be readily extended to an N×Nmatrix A,i . e .
|A|/epsilon1i1i2···iN=Ai1j1Ai2j2···AiNjN/epsilon1j1j2···jN,
where /epsilon1i1i2···iNequals 1 if i1i2···iNis an even permutation of 1 ,2,...,N and equals−1i fi ti sa n
odd permutation; otherwise it equals zero.
791
TENSORSIWrite the following as contracted Cartesian tensors: a·b;∇2φ;∇×v;∇(∇·v);∇×(∇×v);
(a×b)·c.
The corresponding (contracted) tensor expressions are readily seen to be as follows:
a·b=aibi=δijaibj,
∇2φ=∂2φ
∂xi∂xi=δij∂2φ
∂xi∂xj,
(∇×v)i=/epsilon1ijk∂vk
∂xj,
[∇(∇·v)]i=∂
∂xi
/∂vj
∂xj
/
=δjk∂2vj
∂xi∂xk,
[∇×(∇×v)]i=/epsilon1ijk∂
∂xj
/
/epsilon1klm∂vm
∂xl
/
=/epsilon1ijk/epsilon1klm∂2vm
∂xj∂xl,
(a×b)·c=δijci/epsilon1jklakbl=/epsilon1iklciakbl.
J
An important relationship between the /epsilon1-a n d δ- tensors is expressed by the
identity
/epsilon1ijk/epsilon1klm=δilδjm−δimδjl. (21.29)
To establish the validity of this identity between two fourth-order tensors (the
LHS is a once-contracted sixth-order tensor) we consider the various possible
cases.
The RHS of (21.29) has the values
+1 if i=landj=m/negationslash=i, (21.30)
−1i fi=mandj=l/negationslash=i, (21.31)
0 for any other set of subscript values i, j, l, m . (21.32)
In each product on the LHS khas the same value in both factors and for a
non-zero contribution none of i, l, j, m can have the same value as k. Since there
are only three values, 1, 2 and 3, that any of the subscripts may take, the onlynon-zero possibilities are i=landj=mor vice versa but not all four subscripts
equal (since then each /epsilon1factor is zero, as it would be if i=jorl=m). This
reproduces (21.32) for the LHS of (21.29) and also the conditions (21.30) and
(21.31). The values in (21.30) and (21.31) are also reproduced in the LHS of(21.29) since
(i) if i=landj=m,/epsilon1
ijk=/epsilon1lmk=/epsilon1klmand, whether /epsilon1ijkis +1 or−1, the
product of the two factors is +1; and
(ii) if i=mandj=l,/epsilon1ijk=/epsilon1mlk=−/epsilon1klmand thus the product /epsilon1ijk/epsilon1klm(no
summation) has the value −1.
This concludes the establishment of identity (21.29).
792
21.9 ISOTROPIC TENSORS
A useful application of (21.29) is in obtaining alternative expressions for vector
quantities that arise from the vector product of a vector product.IObtain an alternative expression for ∇×(∇×v).
As shown in the previous example, ∇×(∇×v) can be expressed in tensor form as
[∇×(∇×v)]i=/epsilon1ijk/epsilon1klm∂2vm
∂xj∂xl
=(δilδjm−δimδjl)∂2vm
∂xj∂xl
=∂
∂xi
/∂vj
∂xj
/
−∂2vi
∂xj∂xj
=[∇(∇·v)]i−∇2vi,
where in the second line we have used the identity (21.29). This result has already
been mentioned in chapter 10 and the reader is referred there for a discussion of itsapplicability.J
By examining the various possibilities, it is straightforward to verify that, more
generally,
/epsilon1ijk/epsilon1pqr=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleδ
ipδiqδir
δjpδjqδjr
δkpδkqδkrvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(21.33)
and it is easily seen that (21.29) is a special case of this result. From (21.33) we
can derive alternative forms of (21.29), for example,
/epsilon1
ijk/epsilon1ilm=δjlδkm−δjmδkl. (21.34)
The pattern of subscripts in these identities is most easily remembered by noting
that the subscripts on the first δon the RHS are those that immediately follow
(cyclically, if necessary) the common subscript, here i,i ne a c h /epsilon1-term on the LHS;
the remaining combinations of j,k,l,m as subscripts in the other δ-terms on the
RHS can then be filled in automatically.
Contracting (21.34) by setting j=l(say) we obtain, since δkk= 3 when using
the summation convention,
/epsilon1ijk/epsilon1ijm=3δkm−δkm=2δkm,
and by contracting once more, setting k=m, we further find that
/epsilon1ijk/epsilon1ijk=6. (21.35)
21.9 Isotropic tensors
It will have been noticed that, unlike most of the tensors discussed (except for
scalars), δijand/epsilon1ijkhave the property that all their components have values
that are the same whatever rotation of axes is made, i.e. the component values
793
TENSORS
are independent of the transformation Lij. Specifically, δ11has the value 1 in
all coordinate frames, whereas for a general second-order tensor Tall we know
is that if T11=f11(x1,x2,x3)t h e n T/prime
11=f11(x/prime
1,x/prime2,x/prime3). Tensors with the former
property are called isotropic (orinvariant ) tensors.
It is important to know the most general form that an isotropic tensor can take,
since the description of the physical properties, e.g. the conductivity, magneticsusceptibility or tensile strength, of an isotropic medium (i.e. a medium havingthe same properties whichever way it is orientated) involves an isotropic tensor.In the previous section it was shown that δ
ijand/epsilon1ijkare second- and third-order
isotropic tensors; we will now show that, to within a scalar multiple, they are theonly such isotropic tensors.
Let us begin with isotropic second-order tensors. Suppose T
ijis an isotropic
tensor; then, by definition, for anyrotation of the axes we must have that
Tij=T/prime
ij=LikLjlTkl (21.36)
for each of the nine components.
First consider a rotation of the axes by 2 π/3 about the (1 ,1,1) direction; this
takes Ox1,Ox2,Ox3intoOx/prime
2,Ox/prime
3,Ox/prime
1respectively. For this rotation L13=1 ,
L21=1 , L32= 1 and all other Lij= 0. This requires that T11=T/prime
11=T33.
Similarly T12=T/prime
12=T31. Continuing in this way, we find:
(a)T11=T22=T33;
(b)T12=T23=T31;
(c)T21=T32=T13.
Next, consider a rotation of the axes (from their original position) by π/2
about the Ox3-axis. In this case L12=−1,L21=1 ,L33= 1 and all other Lij=0 .
Amongst other relationships, we must have from (21.36) that:
T13=(−1)×1×T23;
T23=1×1×T13.
Hence T13=T23= 0 and therefore, by parts (b) and (c) above, each element Tij=
0 except for T11,T22andT33, which are all the same. This shows that Tij=λδij.IShow that λ/epsilon1ijkis the only isotropic third-order Cartesian tensor.
The general line of attack is as above and so only a minimum of explanation will be given.
Tijk=T/prime
ijk=LilLjmLknTlmn (in all, there are 27 elements) .
Rotate about the (1 ,1,1) direction: this is equivalent to making subscript permutations
1→2→3→1. We find
(a)T111=T222=T333,
(b)T112=T223=T331 (and two similar sets),
(c)T123=T231=T312 (and a set involving odd permutations of 1 ,2,3).
794
21.10 IMPROPER ROTATIONS AND PSEUDOTENSORS
Rotate by π/2 about the Ox3-axis: L12=−1,L21=1 , L33= 1, the other Lij=0 .
(d)T111=(−1)×(−1)×(−1)×T222=−T222,
(e)T112=(−1)×(−1)×1×T221,
(f)T221=1×1×(−1)×T112,
(g)T123=(−1)×1×1×T213.
Relations (a) and (d) show that elements with all subscripts the same are zero. Relations
(e), (f) and (b) show that all elements with repe ated subscripts are zero. Relations (g) and
(c) show that T123=T231=T312=−T213=−T321=−T132.
In total, Tijkdiffers from /epsilon1ijkby at most a scalar factor, but since /epsilon1ijk(and hence λ/epsilon1ijk)
has already been shown to be an isotropic tensor, Tijkmust be the most general third-order
isotropic Cartesian tensor.
J
Using exactly the same procedures as those employed for δijand/epsilon1ijk,i tm a yb e
shown that the only isotropic first-order tensor is the trivial one with all elementszero.
21.10 Improper rotations and pseudotensors
So far we have considered rigid rotations of the coordinate axes described by
an orthogonal matrix Lwith|L|= +1, (21.4). Strictly speaking such transfor-
mations are called proper rotations . We now broaden our discussion to include
transformations that are still described by an orthogonal matrix Lbut for which
|L|=−1; these are called improper rotations .
This kind of transformation can always be considered as an inversion of the
coordinate axes through the origin represented by the equation
x
/prime
i=−xi, (21.37)
combined with a proper rotation. The transformation may be looked upon
alternatively as one that changes an initially right-handed coordinate system intoa left-handed one; any prior or subsequent proper rotation will not change thisstate of affairs. The most obvious example of a transformation with |L|=−1i s
the matrix corresponding to (21.37) itself; in this case L
ij=−δij.
As we have emphasised in earlier chapters, any real physical vector vmay be
considered as a geometrical object (i.e. an arrow in space), which can be referred
to independently of any coordinate system and whose direction and magnitude
cannot be altered merely by describing it in terms of a different coordinate system.Thus the components of vtransform as v
/prime
i=Lijvjunder allrotations (proper and
improper).
We can define another type of object, however, whose components may also
be labelled by a single subscript but which transforms as v/prime
i=Lijvjunder proper
rotations and as v/prime
i=−Lijvj(note the minus sign) under improper rotations. In
this case, the viare not strictly the components of a true first-order Cartesian
tensor but instead are said to form the components of a first-order Cartesianpseudotensor orpseudovector .
795
TENSORS
OOpv
x1x2x3
x/prime
1
x/prime2
x/prime
3v/prime
p/prime
Figure 21.2 The behaviour of a vector vand a pseudovector punder a
reflection through the origin of the coordinate system x1,x2,x3giving the new
system x/prime
1,x/prime2,x/prime3.
It is important to realise that a pseudovector (as its name suggests) is not a
geometrical object in the usual sense. In particular, it should notbe considered
as a real physical arrow in space, since its direction is reversed by an improper
transformation of the coordinate axes (such as an inversion through the origin).This is illustrated in figure 21.2, in which the pseudovector pis shown as a broken
line to indicate that it is not a real physical vector.
Corresponding to vectors and pseudovectors, zeroth-order objects may be
divided into scalars and pseudoscalars – the latter being invariant under rotationbut changing sign on reflection.
We may also extend the notion of scalars and pseudoscalars, vectors and pseu-
dovectors, to objects with two or more subscripts. For two subcripts, as definedpreviously, any quantity with components that transform as T
/prime
ij=LikLjlTklun-
derallrotations (proper and improper) is called a second-order Cartesian tensor.
If, however, T/prime
ij=LikLjlTklunder proper rotations but T/prime
ij=−LikLjlTklunder
improper ones (which include reflections), then the Tijare the components of
a second-order Cartesian pseudotensor. In general the components of Cartesian
pseudotensors of arbitary order transform as
T/prime
ij···k=|L|LilLjm···LknTlm···n, (21.38)
where|L|is the determinant of the transformation matrix.
For example, from (21.28) we have that
|L|/epsilon1ijk=LilLjmLkn/epsilon1lmn,
796
21.10 IMPROPER ROTATIONS AND PSEUDOTENSORS
but since|L|=±1 we may rewrite this as
/epsilon1ijk=|L|LilLjmLkn/epsilon1lmn.
From this expression, we see that although /epsilon1ijkbehaves as a tensor under proper
rotations, as discussed in section 21.8, it should properly be regarded as a third-order Cartesian pseudo tensor.IIfbjandckare the components of vectors, show that the quantities ai=/epsilon1ijkbjckform
the components of a pseudovector.
In a new coordinate system we have
a/prime
i=/epsilon1/prime
ijkb/primejc/primek
=|L|LilLjmLkn/epsilon1lmnLjpbpLkqcq
=|L|Lil/epsilon1lmnδmpδnqbpcq
=|L|Lil/epsilon1lmnbmcn
=|L|Lilal,
from which we see immediately that the quantities aiform the components of a pseu-
dovector.
J
The above example is worth some further comment. If we denote the vec-
tors with components bjandckbybandcrespectively then, as mentioned in
section 21.8, the quantities ai=/epsilon1ijkbjckare the components of the real vector
a=b×c,provided that we are using a right-handed Cartesian coordinate system .
However, in a coordinate system that is left-handed the quantitites a/prime
i=/epsilon1/prime
ijkb/prime
jc/prime
k
arenotthe components of the physical vector a=b×c, which has, instead, the
components −a/prime
i. It is therefore important to note the handedness of a coordinate
system before attempting to write in component form the vector relation a=b×c
(which is true without reference to any coordinate system).
It is worth noting that, although pseudotensors can be useful mathematical
objects, the description of the real physical world must usually be in terms oftensors (i.e. scalars, vectors, etc.). †For example, the temperature or density of a
gas must be a scalar quantity (rather than a pseudoscalar), since its value doesnot change when the coordinate system used to describe it is inverted through
the origin. Similarly, velocity, magnetic field strength or angular momentum can
only be described by a vector, and not by a pseudovector.
At this point, it may be useful to make a brief comment on the distinction
between active andpassive transformations of a physical system, as this difference
often causes confusion. In this chapter, we are concerned solely with passive trans-
†In fact the quantum-mechanical description of elementary particles, such as electrons, protons and
neutrons, requires the introduction of a new kind of mathematical object called a spinor ,w h i c hi s
not a scalar, vector, or more general tensor. The study of spinors, however, falls beyond the scope
of this book.
797
TENSORS
formations, for which the physical system of interest is left unaltered, and only
the coordinate system used to describe it is changed. In an active transformation,however, the system itself is altered.
As an example, let us consider a particle of mass mthat is located at a position
xrelative to the origin Oand hence has velocity ˙x. The angular momentum of
the particle about Ois thus J=m(x×˙x). If we merely invert the Cartesian
coordinates used to describe this system through O, neither the magnitude nor
direction of any these vectors will be changed, since they may be consideredsimply as arrows in space that are independent of the coordinates used to de-scribe them. If, however, we perform the analogous active transformation on
the system, by inverting the position vector of the particle through O,t h e ni t
is clear that the direction of particle’s velocity will also be reversed, since itis simply the time derivative of the position vector, but that the direction ofits angular momentum vector remains unaltered. This suggests that vectors canbe divided into two categories, as follows: polar vectors (such as position and
velocity) which reverse direction under an active inversion of the physical sys-tem through the origin and axial vectors (such as angular momentum), which
remain unchanged. It should be emphasised that at no point in this discus-
sion have we used the concept of a pseudovector to describe a real physicalquantity.†
21.11 Dual tensors
Although pseudotensors are not themselves appropriate for the description of
physical phenomena, they are sometimes needed; for example, we may use thepseudotensor /epsilon1
ijkto associate with every antisymmetric second-order tensor Aij
(in three dimensions) a pseudovector pigiven by
pi=1
2/epsilon1ijkAjk; (21.39)
piis called the dualofAij. Thus if we denote the antisymmetric tensor Aby the
matrix
A=[Aij]=
0 A12−A31
−A12 0 A23
A31−A23 0
then the components of its dual pseudovector are ( p1,p2,p3)=(A23,A31,A12).
†The scalar product of a polar vector and an axial vector is a pseudoscalar. It was the experimental
detection of the dependence of the angular distribution of electrons of (polar vector) momentum
peemitted by polarised nuclei of (axial vector) spin JNupon the pseudoscalar quantity JN·pethat
established the existence of the non-conservation of parity in β-decay.
798
21.12 PHYSICAL APPLICATIONS OF TENSORSIUsing (21.39) show that Aij=/epsilon1ijkpk.
By contracting both sides of (21.39) with /epsilon1ijk, we find
/epsilon1ijkpk=1
2/epsilon1ijk/epsilon1klmAlm.
Using the identity (21.29) then gives
/epsilon1ijkpk=1
2(δilδjm−δimδjl)Alm
=1
2(Aij−Aji)=1
2(Aij+Aij)=Aij,
where in the last line we use the fact that Aij=−Aji.
J
By a simple extension, we may associate a dual pseudoscalar swith every
totally antisymmetric third-rank tensor Aijk,i.e. one that is antisymmetric with
respect to the interchange of every possible pair of subscripts; sis given by
s=1
3!/epsilon1ijkAijk. (21.40)
Since Aijkis a totally antisymmetric three-subscript quantity, we expect it to
equal some multiple of /epsilon1ijk(since this is the only such quantity). In fact Aijk=s/epsilon1ijk,
as can be proved by substituting this expression into (21.40) and using (21.35).
21.12 Physical applications of tensors
In this section some physical applications of tensors will be given. First-order
tensors are familiar as vectors and so we will concentrate on second-order tensors,starting with an example taken from mechanics.
Consider a collection of rigidly connected point particles of which the αth,
which has mass m
(α)and is positioned at r(α)with respect to an origin O,i s
typical. Suppose that the rigid assembly is rotating about an axis through Owith
angular velocity ω.
The angular momentum Jabout Oof the assembly is given by
J=summationdisplay
αparenleftbig
r(α)×p(α)parenrightbig
.
Butp(α)=m(α)˙r(α)and˙r(α)=ω×r(α),f o ra n y α, and so in subscript form the
components of Jare given by
Ji=summationdisplay
αm(α)/epsilon1ijkx(α)
j˙x(α)
k
=summationdisplay
αm(α)/epsilon1ijkx(α)
j/epsilon1klmωlx(α)
m
=summationdisplay
αm(α)(δilδjm−δimδjl)x(α)
jx(α)
mωl
=summationdisplay
αm(α)bracketleftBigparenleftbig
r(α)parenrightbig2δil−x(α)
ix(α)
lbracketrightBig
ωl≡Iilωl, (21.41)
where Iilis a symmetric second-order Cartesian tensor (by the quotient rule, see
799
TENSORS
section 21.7, since Jand ωare vectors). The tensor is called the inertia tensor atO
of the assembly and depends only on the distribution of masses in the assemblyand not upon the direction or magnitude of ω.
A more realistic situation obtains if a continuous rigid body is considered. In
this case, m
(α)must be replaced everywhere by ρ(r)dx dy dz and all summations
by integrations over the volume of the body. Written out in full in Cartesians,the inertia tensor for a continuous body would have the form
I=[I
ij]=
integraltext
(y2+z2)ρd V−integraltext
xyρ dV −integraltext
xzρ dV
−integraltext
xyρ dVintegraltext
(z2+x2)ρd V−integraltext
yzρdV
−integraltext
xzρ dV −integraltext
yzρdVintegraltext
(x2+y2)ρd V
,
where ρ=ρ(x, y, z) is the mass distribution and dVstands for dx dy dz ;t h e
integrals are to be taken over the whole body. The diagonal elements of thistensor are called the moments of inertia and the off-diagonal elements without the
minus signs are known as the products of inertia .IShow that the kinetic energy of the rotating system is given by T=1
2Ijlωjωl.
By an argument parallel to that already made for J, the kinetic energy is given by
T=1
2
X
αm(α)
/;˙r(α)·˙r(α)
/
=1
2
X
αm(α)/epsilon1ijkωjx(α)
k/epsilon1ilmωlx(α)
m
=1
2
X
αm(α)(δjlδkm−δjmδkl)x(α)
kx(α)
mωjωl
=1
2
X
αm(α)
h
δjl
/;
r(α)
/2−x(α)
jx(α)
l
i
ωjωl
=1
2Ijlωjωl.
Alternatively, since Jj=Ijlωlwe may write the kinetic energy of the rotating system as
T=1
2Jjωj.
J
The above example shows that the kinetic energy of the rotating body can be
expressed as a scalar obtained by twice contracting ωwith the inertia tensor. It
also shows that the moment of inertia of the body about a line given by the unitvector ˆnisI
jlˆnjˆnl(orˆnTIˆnin matrix form).
Since I(≡Ijl) is a real symmetric second-order tensor, it has associated with it
three mutually perpendicular directions that are its principal axes and have the
following properties (proved in chapter 8):
(i) with each axis is associated a principal moment of inertia λµ,µ=1,2,3;
(ii) when the rotation of the body is about one of these axes, the angular
velocity and the angular momentum are parallel and given by
J=Iω=λµω,
i.e.ωis an eigenvector of Iwith eigenvalue λµ;
800
21.12 PHYSICAL APPLICATIONS OF TENSORS
(iii) referred to these axes as coordinate axes, the inertia tensor is diagonal
with diagonal entries λ1,λ2,λ3.
Two further examples of physical quantities represented by second-order tensors
are magnetic susceptibility and electrical conductivity. In the first case we have(in standard notation)
M
i=χijHj, (21.42)
and in the second case
ji=σijEj. (21.43)
Here Mis the magnetic moment per unit volume and jthe current density
(current per unit perpendicular area). In both cases we have on the left-hand sidea vector and on the right-hand side the contraction of a set of quantities withanother vector. Each set of quantities must therefore form the components of a
second-order tensor.
For isotropic media M∝Handj∝E, but for anisotropic materials such as
crystals the susceptibility and conductivity may be different along different crystalaxes, making χ
ijandσijgeneral second-order tensors, although they are usually
symmetric.IThe electrical conductivity σin a crystal is measured by an observer to have components
as shown
[σij]=
/0/@1√
20√
231
01 1
/1A. (21.44)
Show that there is one direction in the crystal along which no current can flow. Does the
current flow equally easily in the two perpendicular directions?
The current density in the crystal is given by ji=σijEj,w h e r e σij,r e l a t i v et ot h e
observer’s coordinate system, is given by (21.44). Since [ σij] is a symmetric matrix, it
possess three mutually perpendicular eigenvectors (or principal axes) with respect to
which the conductivity tensor is diagonal, with diagonal entries λ1,λ2,λ3, the eigenvalues
of [σij].
As discussed in chapter 8, the eigenvalues of [ σij] are given by |σ−λI|= 0. Thus we
require//////1−λ√
20√
23−λ1
01 1 −λ
//////=0,
from which we find
(1−λ)[(3−λ)(1−λ)−1]−2(1−λ)=0 .
This simplifies to give λ=0,1,4 so that, with respect to its principal axes, the conductivity
tensor has components σ/prime
ijgiven by
[σ/prime
ij]=
/0/@400
010000
/1A.
Since j/prime
i=σ/prime
ijE/prime
j, we see immediately that along one of the principal axes there is no
current flow and along the two perpendicular directions the current flows are not equal.J
801
TENSORS
We can extend the idea of a second-order tensor that relates two vectors to a
situation where two physical second-order tensors are related by a fourth-ordertensor. The most common occurrence of such relationships is in the theory ofelasticity. This is not the place to give a detailed account of elasticity theory,
but suffice it to say that the local deformation of an elastic body at any interior
point Pcan be described by a second-order symmetric tensor e
ijcalled the strain
tensor .I ti sg i v e nb y
eij=1
2parenleftbigg∂ui
∂xj+∂uj
∂xiparenrightbigg
,
where uis the displacement vector describing the strain of a small volume element
whose unstrained position relative to the origin is x. Similarly we can describe
the stress in the body at Pby the second-order symmetric stress tensor pij;t h e
quantity pijis the xj-component of the stress vector acting across a plane through
Pwhose normal lies in the xi-direction. A generalisation of Hooke’s law then
relates the stress and strain tensors by
pij=cijklekl (21.45)
where cijklis a fourth-order Cartesian tensor.IAssuming that the most general fourth-order isotropic tensor is
cijkl=λδijδkl+ηδikδjl+νδilδjk, (21.46)
find the form of (21.45) for an isotropic medium having Young’s modulus Eand Poisson’s
ratio σ.
For an isotropic medium we must have an isotropic tensor for cijkl, and so we assume the
form (21.46). Substituting this into (21.45) yields
pij=λδijekk+ηeij+νeji.
Buteijis symmetric, and if we write η+ν=2µ, then this takes the form
pij=λekkδij+2µeij,
in which λandµare known as Lam´e constants . It will be noted that if eij=0f o r i/negationslash=j
t h e nt h es a m ei st r u eo f pij, i.e. the principal axes of the stress and strain tensors coincide.
Now consider a simple tension in the x1-direction, i.e. p11=Sbut all other pij=0 .
Then denoting ekk(summed over k)b yθwe have, in addition to eij=0f o r i/negationslash=j,t h et h r e e
equations
S=λθ+2µe11,
0=λθ+2µe22,
0=λθ+2µe33.
Adding them gives
S=θ(3λ+2µ).
Substituting for θfrom this into the first of the three, and recalling that Young’s modulus
is defined by S=Ee11,g i v e s Eas
E=µ(3λ+2µ)
λ+µ. (21.47)
802
21.13 INTEGRAL THEOREMS FOR TENSORS
Further, Poisson’s ratio is defined as σ=−e22/e11(or−e33/e11) and is thus
σ=
/1
e11
/λθ
2µ=
/1
e11
//λ
2µ
/Ee11
3λ+2µ=λ
2(λ+µ). (21.48)
Solving (21.47) and (21.48) for λandµgives finally
pij=σE
(1 +σ)(1−2σ)ekkδij+E
(1 +σ)eij.
J
21.13 Integral theorems for tensors
In chapter 11, we discussed various integral theorems involving vector and scalar
fields. Most notably, we considered the divergence theorem, which states that, forany vector field a, integraldisplay
V∇·adV=contintegraldisplay
Sa·ˆndS, (21.49)
where Sis the surface enclosing the volume Vandˆnis the outward-pointing unit
normal to Sat each point.
Writing (21.49) in subscript notation, we have
integraldisplay
V∂ak
∂xkdV=contintegraldisplay
SakˆnkdS. (21.50)
Although we shall not prove it rigorously, (21.50) can be extended in an obvious
manner to relate integrals of tensor fields , rather than just vector fields, over
volumes and surfaces, with the result
integraldisplay
V∂Tij···k···m
∂xkdV=contintegraldisplay
STij···k···mˆnkdS.
This form of the divergence theorem for general tensors can be very useful in
vector calculus manipulations.IA vector field asatisfies∇·a=0inside some volume Vanda·ˆn=0on the bound-
ary surface S. By considering the divergence theorem applied to Tij=xiaj, show thatR
VadV=0.
Applying the divergence theorem to Tij=xiajwe findZ
V∂Tij
∂xjdV=
Z
V∂(xiaj)
∂xjdV=
I
SxiajˆnjdS=0,
since ajˆnj= 0. By expanding the volume integral we obtainZ
V∂(xiaj)
∂xjdV=
Z
V∂xi
∂xjajdV+
Z
Vxi∂aj
∂xjdV
=
Z
VδijajdV
=
Z
VaidV=0,
where in going from the first to the second line we used the fact that ∂xi/∂x j=δijand
∂aj/∂x j=0 .
J
803
TENSORS
The other integral theorems discussed in chapter 11 can be extended in a
similar way. For example, written in tensor notation Stokes’ theorem states that,for a vector field a
i,integraldisplay
S/epsilon1ijk∂ak
∂xjˆnidS=contintegraldisplay
Cakdxk.
For a general tensor field this has the straightforward extension
integraldisplay
S/epsilon1ijk∂Tlm···k···n
∂xjˆnidS=contintegraldisplay
CTlm···k···ndxk.
21.14 Non-Cartesian coordinates
So far we have restricted our attention to the study of tensors when they are
described in terms of Cartesian coordinates and the axes of coordinates are rigidly
rotated, sometimes together with an inversion of axes through the origin. In theremainder of this chapter we shall extend the concepts discussed in the previoussections by considering arbitrary coordinate transformations from one generalcoordinate system to another. Although this generalisation brings with it severalcomplications, we shall find that many of the properties of Cartesian tensorsare still valid for more general tensors. Before considering general coordinate
transformations, however, we begin by reminding ourselves of some properties of
general curvilinear coordinates, as discussed in chapter 10.
The position of an arbitrary point Pin space may be expressed in terms of the
three curvilinear coordinates u
1,u2,u3. We saw in chapter 10 that if r(u1,u2,u3)i s
the position vector of the point Pthen at Pthere exist two sets of basis vectors
ei=∂r
∂uiand /epsilon1i=∇ui, (21.51)
where i=1,2,3. In general, the vectors in each set neither are of unit length nor
form an orthogonal basis. However, the sets eiand /epsilon1iare reciprocal systems of
vectors and so
ei·/epsilon1j=δij. (21.52)
In the context of general tensor analysis, it is more usual to denote the second
set of vectors /epsilon1iin (21.51) by ei, the index being placed as a superscript to
distinguish it from the (different) vector ei, which is a member of the first set in
(21.51). Although this positioning of the index may seem odd (not least becauseof the possibility of confusion with powers) it forms part of a slight modificationto the summation convention that we will adopt for the remainder of this chapter.
This is as follows: any lower-case alphabetic index that appears exactly twice in
any term of an expression, once as a subscript and once as a superscript ,i st ob e
summed over all the values that an index in that position can take (unless the
804
21.14 NON-CARTESIAN COORDINATES
contrary is specifically stated). All other aspects of the summation convention
remain unchanged.
With the introduction of superscripts, the reciprocity relation (21.52) should be
rewritten so that both sides of (21.53) have one subscript and one superscript, i.e.
as
ei·ej=δj
i. (21.53)
The alternative form of the Kronecker delta is defined in a similar way to
previously, i.e. it equals unity if i=jand is zero otherwise.
For similar reasons it is usual to denote the curvilinear coordinates themselves
byu1,u2,u3, with the index raised, so that
ei=∂r
∂uiand ei=∇ui. (21.54)
From the first equality we see that we may consider a superscript that appears in
the denominator of a partial derivative as a subscript.
Given the two bases eiandei,w em a yw r i t eag e n e r a lv e c t o r aequally well in
terms of either basis as follows:
a=a1e1+a2e2+a3e3=aiei;
a=a1e1+a2e2+a3e3=aiei.
The aiare called the contravariant components of the vector aand the ai
thecovariant components, the position of the index (either as a subscript or
superscript) serving to distinguish between them. Similarly, we may call the eithe
covariant basis vectors and the eithe contravariant ones.IShow that the contravariant and covariant components of a vector aare given by ai=a·ei
andai=a·eirespectively.
For the contravariant components, we find
a·ei=ajej·ei=ajδi
j=ai,
where we have used the reciprocity relation (21.53). Similarly, for the covariant components,
a·ei=ajej·ei=ajδj
i=ai.
J
The reason that the notion of contravariant and covariant components of
a vector (and the resulting superscript notation) was not introduced earlier isthat for Cartesian coordinate systems the two sets of basis vectors e
iandeiare
identical and, hence, so are the components of a vector with respect to eitherbasis. Thus, for Cartesian coordinates, we may speak simply of the componentsof the vector and there is no need to differentiate between contravariance and
covariance, or to introduce superscripts to make a distinction between them.
If we consider the components of higher-order tensors in non-Cartesian co-
ordinates, there are even more possibilities. As an example, let us consider a
805
TENSORS
second-order tensor T. Using the outer product notation in (21.23), we may write
Tin three different ways:
T=Tijei⊗ej=Ti
jei⊗ej=Tijei⊗ej,
where Tij,Ti
jandTijare called the contravariant, mixed andcovariant com-
ponents of Trespectively. It is important to remember that these three sets of
quantities form the components of the sametensor Tbut refer to different (tensor)
bases made up from the basis vectors of the coordinate system. Again, if we are
using Cartesian coordinates then all three sets of components are identical.
We may generalise the above equation to higher-order tensors; then com-
ponents carrying only superscripts or only subscripts are referred to as thecontravariant and covariant components respectively and all others are calledmixed components.
21.15 The metric tensor
Any particular curvilinear coordinate system is completely characterised at each
point in space by the nine quantities
g
ij=ei·ej, (21.55)
which, as we will show, are the covariant components of a symmetric second-order
tensor gcalled the metric tensor .
Since an infinitesimal vector displacement can be written as dr=duiei, we find
that the square of the infinitesimal arc length ( ds)2c a nb ew r i t t e ni nt e r m so ft h e
metric tensor as
(ds)2=dr·dr=duiei·dujej=gijduiduj. (21.56)
It may further be shown that the volume element dVis given by
dV=√gd u1du2du3, (21.57)
where gis the determinant of the matrix [ gij], which has the covariant components
of the metric tensor as its elements.
If we compare equations (21.56) and (21.57) with the analogous ones in section
10.10 then we see that in the special case where the coordinate system is orthogonal(so that e
i·ej=0f o r i/negationslash=j) the metric tensor can be written in terms of the
coordinate-system scale factors hi,i=1,2,3a s
gij=braceleftBigg
h2
ii=j,
0i/negationslash=j.
Its determinant is then given by g=h2
1h22h23.
806
21.15 THE METRIC TENSORICalculate the elements gijof the metric tensor for cylindrical polar coordinates. Hence
find the square of the infinitesimal arc length (ds)2and the volume dVfor this coordinate
system.
As discussed in section 10.9, in cylindrical polar coordinates ( u1,u2,u3)=( ρ, φ, z)a n ds o
the position vector rof any point Pmay be written
r=ρcosφi+ρsinφj+zk.
From this we obtain the (covariant) basis vectors:
e1=∂r
∂ρ=c o s φi+s i n φj;
e2=∂r
∂φ=−ρsinφi+ρcosφj;
e3=∂r
∂z=k. (21.58)
Thus the components of the metric tensor [ gij]=[ei·ej] are found to be
G=[gij]=
/0/@100
0ρ20
001
/1A, (21.59)
from which we see that, as expected for an orthogonal coordinate system, the metric tensor
is diagonal, the diagonal elements being equal to the squares of the scale factors of thecoordinate system.
From (21.56), the square of the infinitesimal arc length in this coordinate system is given
by
(ds)
2=gijduiduj=(dρ)2+ρ2(dφ)2+(dz)2,
and, using (21.57), the volume element is found to be
dV=√gd u1du2du3=ρd ρd φd z.
These expressions are identical to those derived in section 10.9.
J
We may also express the scalar product of two vectors in terms of the metric
tensor:
a·b=aiei·bjej=gijaibj, (21.60)
where we have used the contravariant components of the two vectors. Similarly,
using the covariant components, we can write the same scalar product as
a·b=aiei·bjej=gijaibj, (21.61)
where we have defined the nine quantities gij=ei·ej. As we shall show, they form
the contravariant components of the metric tensor gand are, in general, different
from the quantities gij. Finally, we could express the scalar product in terms of
the contravariant components of one vector and the covariant components of the
other,
a·b=aiei·bjej=aibjδi
j=aibi, (21.62)
807
TENSORS
where we have used the reciprocity relation (21.53). Similarly, we could write
a·b=aiei·bjej=aibjδj
i=aibi. (21.63)
By comparing the four alternative expressions (21.60)–(21.63) for the scalar
product of two vectors we can deduce one of the most useful properties of
the quantities gijand gij.S i n c e gijaibj=aibiholds for any arbitrary vector
components ai, it follows that
gijbj=bi,
which illustrates the fact that the covariant components gijof the metric tensor
can be used to lower an index . In other words, it provides a means of obtaining
the covariant components of a vector from its contravariant components. By asimilar argument, we have
g
ijbj=bi,
so that the contravariant components gijcan be used to perform the reverse
operation of raising an index .
It is straightforward to show that the contravariant and covariant basis vectors,
eiandeirespectively, are related in the same way as other vectors, i.e. by
ei=gijejand ei=gijej.
We also note that, since eiandeiare reciprocal systems of vectors in three-
dimensional space (see chapter 7), we may write
ei=ej×ek
ei·(ej×ek),
for the combination of subscripts i, j, k=1,2,3 and its cyclic permutations. A
similar expression holds for eiin terms of the ei-basis. Moreover, it may be shown
that the triple scalar product |e1·(e2×e3)|=√g.IShow that the matrix [gij]is the inverse of the matrix [gij]. Hence calculate the con-
travariant components gijof the metric tensor in cylindrical polar coordinates.
Using the index-lowering and index-raising properties of gijandgijon an arbitrary vector
a, we find
δi
kak=ai=gijaj=gijgjkak.
But, since ais arbitrary, we must have
gijgjk=δi
k. (21.64)
Denoting the matrix [ gij]b y Gand [ gij]b yˆG, equation (21.64) can be written in matrix
form as ˆGG=I,w h e r e Iis the unit matrix. Hence GandˆGare inverse matrices of each
other.
808
21.16 GENERAL COORDINATE TRANSFORMATIONS AND TENSORS
Thus, by inverting the matrix Gin (21.59), we find that the elements gijare given in
cylindrical polar coordinates by
ˆG=[gij]=
/0/@100
01 /ρ20
001
/1A.
J
So far we have not considered the components of the metric tensor gi
jwith one
subscript and one superscript. By analogy with (21.55), these mixed componentsare given by
g
i
j=ei·ej=δj
i,
and so the components of gi
jare identical to those of δi
j.W em a yt h e r e f o r e
consider the δi
jto be the mixed components of the metric tensor g.
21.16 General coordinate transformations and tensors
We now discuss the concept of general transformations from one coordinate
system, u1,u2,u3, to another, u/prime1,u/prime2,u/prime3. We can describe the coordinate transform
using the three equations
u/primei=u/primei(u1,u2,u3),
fori=1,2,3, in which the new coordinates u/primeican be arbitrary functions of the old
ones uirather than just represent linear orthogonal transformations (rotations)
of the coordinate axes. We shall assume also that the transformation can beinverted, so that we can write the old coordinates in terms of the new ones as
u
i=ui(u/prime1,u/prime2,u/prime3),
As an example, we may consider the transformation from spherical polar to
Cartesian coordinates, given by
x=rsinθcosφ,
y=rsinθsinφ,
z=rcosθ,
which is clearly not a linear transformation.
The two sets of basis vectors in the new coordinate system, u/prime1,u/prime2,u/prime3, are given
as in (21.54) by
e/prime
i=∂r
∂u/primeiand e/primei=∇u/primei. (21.65)
Considering the first set, we have from the chain rule that
∂r
∂uj=∂u/primei
∂uj∂r
∂u/primei,
809
TENSORS
so that the basis vectors in the old and new coordinate systems are related by
ej=∂u/primei
∂uje/prime
i. (21.66)
Now, since we can write any arbitrary vector ain terms of either basis as
a=a/primeie/prime
i=ajej=aj∂u/primei
∂uje/prime
i,
it follows that the contravariant components of a vector must transform as
a/primei=∂u/primei
∂ujaj. (21.67)
In fact, we use this relation as the defining property for a set of quantities aito
form the contravariant components of a vector.IFind an expression analogous to (21.66) relating the basis vectors eiande/primeiin the two
coordinate systems. Hence deduce the way i n which the covariant components of a vector
change under a coordinate transformation.
If we consider the second set of basis vectors in (21.65), e/primei=∇u/primei, we have from the chain
rule that
∂uj
∂x=∂uj
∂u/primei∂u/primei
∂x
and similarly for ∂uj/∂yand∂uj/∂z. So the basis vectors in the old and new coordinate
systems are related by
ej=∂uj
∂u/primeie/primei. (21.68)
For any arbitrary vector a,
a=a/prime
ie/primei=ajej=aj∂uj
∂u/primeie/primei
and so the covariant components of a vector must transform as
a/prime
i=∂uj
∂u/primeiaj. (21.69)
Analogously to the contravariant case (21.67), we take this result as the defining property
of the covariant components of a vector.
J
We may compare the transformation laws (21.67) and (21.69) with those for
a first-order Cartesian tensor under a rigid rotation of axes. Let us consider
a rotation of Cartesian axes xithrough an angle θabout the 3-axis to a new
setx/primei,i=1,2,3, as given by (21.7) and the inverse transformation (21.8). It is
straightforward to show that
∂xj
∂x/primei=∂x/primei
∂xj=Lij,
810
21.16 GENERAL COORDINATE TRANSFORMATIONS AND TENSORS
where the elements Lijare given by
L=
cosθsinθ0
−sinθcosθ0
00 1
.
Thus (21.67) and (21.69) agree with our earlier definition in the special case of a
rigid rotation of Cartesian axes.
Following on from (21.67) and (21.69), we proceed in a similar way to de-
fine general tensors of higher rank. For example, the contravariant, mixed and
covariant components, respectively, of a second-order tensor must transform as
follows:
contravariant components, T/primeij=∂u/primei
∂uk∂u/primej
∂ulTkl;
mixed components, T/primei
j=∂u/primei
∂uk∂ul
∂u/primejTk
l;
covariant components, T/prime
ij=∂uk
∂u/primei∂ul
∂u/primejTkl.
It is important to remember that these quantities form the components of the
sametensor Tbut refer to different tensor bases made up from the basis vectors
of the different coordinate systems. For example, in terms of the contravariant
components we may write
T=Tijei⊗ej=T/primeije/prime
i⊗e/prime
j.
We can clearly go on to define tensors of higher order, with arbitrary numbers
of covariant (subscript) and contravariant (superscript) indices, by demandingthat their components transform as follows:
T
/primeij···k
lm···n=∂u/primei
∂ua∂u/primej
∂ub···∂u/primek
∂uc∂ud
∂u/primel∂ue
∂u/primem···∂uf
∂u/primenTab···c
de···f.(21.70)
Using the revised summation convention described in section 21.14, the algebra
of general tensors is completely analogous to that of the Cartesian tensorsdiscussed earlier. For example, as with Cartesian coordinates, the Kroneckerdelta is a tensor provided it is written as the mixed tensor δ
i
jsince
δ/primei
j=∂u/primei
∂uk∂ul
∂u/primejδk
l=∂u/primei
∂uk∂uk
∂u/primej=∂u/primei
∂u/primej=δi
j,
where we have used the chain rule to justify the third equality. This also shows
that δi
jis isotropic. As discussed at the end of section 21.15, the δi
jcan be
considered as the mixed components of the metric tensor g.
811
TENSORSIShow that the quantities gij=ei·ejform the covariant com ponents of a second-order
tensor.
In the new (primed) coordinate system we have
g/prime
ij=e/prime
i·e/prime
j,
but using (21.66) for the inverse transformation, we have
e/prime
i=∂uk
∂u/primeiek,
and similarly for e/prime
j. Thus we may write
g/prime
ij=∂uk
∂u/primei∂ul
∂u/primejek·el=∂uk
∂u/primei∂ul
∂u/primejgkl,
which shows that the gijare indeed the covariant components of a second-order tensor
(the metric tensor g).
J
A similar argument to that used in the above example shows that the quantities
gijform the contravariant components of a second-order tensor which transforms
according to
g/primeij=∂u/primei
∂uk∂u/primej
∂ulgkl.
In the previous section we discussed the use of the components gijandgijin
the raising and lowering of indices in contravariant and covariant vectors. This
can be extended to tensors of arbitrary rank. In general, contraction of a tensorwith g
ijwill convert the contracted index fro m being contravariant (superscript)
to covariant (subscript), i.e. it is lowered. This can be repeated for as many indicesare required. For example,
T
ij=gikTk
j=gikgjlTkl. (21.71)
Similarly contraction with gijraises an index, i.e.
Tij=gikTj
k=gikgjlTkl. (21.72)
That (21.71) and (21.72) are mutually consistent may be shown by using the fact
thatgikgkj=δi
j.
21.17 Relative tensors
In section 21.10 we introduced the concept of pseudotensors in the context of the
rotation (proper or improper) of a set of Cartesian axes. Generalising to arbitrarycoordinate transformations leads to the notion of a relative tensor .
For an arbitrary coordinate transformation from one general coordinate system
812
21.17 RELATIVE TENSORS
uito another u/primei, we may define the Jacobian of the transformation (see chapter 6)
as the determinant of the transformation matrix [ ∂u/primei/∂uj]: this is usually denoted
by
J=vextendsinglevextendsinglevextendsinglevextendsingle∂u/prime
∂uvextendsinglevextendsinglevextendsinglevextendsingle.
Alternatively, we may interchange the primed and unprimed coordinates to
obtain|∂u/∂u/prime|=1/J: unfortunately this also is often called the Jacobian of the
transformation.
Using the Jacobian J, we define a relative tensor of weight was one whose
components transform as follows:
T/primeij···k
lm···n=∂u/primei
∂ua∂u/primej
∂ub···∂u/primek
∂uc∂ud
∂u/primel∂ue
∂u/primem···∂uf
∂u/primenTab···c
de···fvextendsinglevextendsinglevextendsinglevextendsingle∂u
∂u/primevextendsinglevextendsinglevextendsinglevextendsinglew
.
(21.73)
Comparing this expression with (21.70), we see that a true (or absolute )g e n e r a l
tensor may be considered as a relative tensor of weight w=0 .I f w=−1, on the
other hand, the relative tensor is known as a general pseudotensor ,a n di f w=1
as atensor density .
It is worth comparing (21.73) with the definition (21.38) of a Cartesian pseu-
dotensor. For the latter, we are concerned only with its behaviour under a rotation(proper or improper) of Cartesian axes, for which the Jacobian J=±1. Thus,
general relative tensors of weight w=−1a n d w= 1 would both satisfy the
definition (21.38) of a Cartesian pseudotensor.IIf the gijare the covariant components of the metric tensor, show that the determinant g
of the matrix [gij]is a relative scalar of weight w=2.
The components gijtransform as
g/prime
ij=∂uk
∂u/primei∂ul
∂u/primejgkl.
Defining the matrices U=[∂ui/∂u/primej],G=[gij]a n d G/prime=[g/prime
ij], we may write this expression
as
G/prime=UTGU.
Taking the determinant of both sides, we obtain
g/prime=|U|2g=
////∂u
∂u/prime
////2
g,
which shows that gis a relative scalar of weight w=2 .
J
From the discussion in section 21.8, it can be seen that /epsilon1ijkis a covariant
relative tensor of weight −1. We may also define the contravariant tensor /epsilon1ijk,
which is numerically equal to /epsilon1ijkbut is a relative tensor of weight +1.
If two relative tensors have weights w1andw2respectively then, from (21.73),
813
TENSORS
the outer product of the two tensors, or any contraction of them, is a relative
tensor of weight w1+w2. As a special case, we may use /epsilon1ijkand/epsilon1ijkto construct
pseudovectors from antisymmetric tensors and vice versa, in an analogous wayto that discussed in section 21.11.
For example, if the A
ijare the contravariant components of an antisymmetric
tensor ( w=0 )t h e n
pi=1
2/epsilon1ijkAjk
are the covariant components of a pseudovector ( w=−1), since /epsilon1ijkhas weight
w=−1. Similarly, we may show that
Aij=/epsilon1ijkpk.
21.18 Derivatives of basis vectors and Christoffel symbols
In Cartesian coordinates, the basis vectors eiare constant and so their derivatives
with respect to the coordinates vanish. In a general coordinate system, however,
the basis vectors eiandeiare functions of the coordinates. Therefore, in order
that we may differentiate general tensors we must consider the derivatives of thebasis vectors.
First consider the derivative ∂e
i/∂uj. Since this is itself a vector, it can be
written as a linear combination of the basis vectors ek,k=1,2,3. If we introduce
the symbol Γk
ijto denote the coefficients in this combination, we have
∂ei
∂uj=Γk
ijek. (21.74)
The coefficient Γk
ijis the kth component of the vector ∂ei/∂uj.U s i n gt h er e c i -
procity relation ei·ej=δi
j, these 27 numbers are given (at each point in space)
by
Γk
ij=ek·∂ei
∂uj. (21.75)
Furthermore, by differentiating the reciprocity relation ei·ej=δi
jwith respect
to the coordinates, and using (21.75), it is straightforward to show that thederivatives of the contravariant basis vectors are given by
∂e
i
∂uj=−Γi
kjek. (21.76)
The symbol Γk
ijis called a Christoffel symbol (of the second kind), but, despite
appearances to the contrary, these quantities do notform the components of a
third-order tensor. It is clear from (21.75) that in Cartesian coordinates Γk
ij=0
for all values of the indices i,jandk.
814
21.18 DERIVATIVES OF BASIS VECTORS AND CHRISTOFFEL SYMBOLSIUsing (21.75), deduce the way in which the quantities Γk
ijtransform under a general
coordinate transformation, and hence show that they do not form the components of athird-order tensor.
In a new coordinate system
Γ/primek
ij=e/primek·∂e/prime
i
∂u/primej,
but from (21.68) and (21.66) respectively we have, on reversing primed and unprimed
variables,
e/primek=∂u/primek
∂unenand e/prime
i=∂ul
∂u/primeiel.
Therefore in the new coordinate system the quantities Γ/primek
ijare given by
Γ/primek
ij=∂u/primek
∂unen·∂
∂u/primej
/∂ul
∂u/primeiel
/
=∂u/primek
∂unen·
/∂2ul
∂u/primej∂u/primeiel+∂ul
∂u/primei∂el
∂u/primej
/
=∂u/primek
∂un∂2ul
∂u/primej∂u/primeien·el+∂u/primek
∂un∂ul
∂u/primei∂um
∂u/primejen·∂el
∂um
=∂u/primek
∂ul∂2ul
∂u/primej∂u/primei+∂u/primek
∂un∂ul
∂u/primei∂um
∂u/primejΓn
lm, (21.77)
where in the last line we have used (21.75) and the reciprocity relation en·el=δn
l.F r o m
(21.77), because of the presence of the first term on the right-hand side, we concludeimmediately that the Γ
k
ijdo not form the components of a third-order tensor.
J
In a given coordinate system, in principle, we may calculate the Γk
ijusing
(21.75). In practice, however, it is often quicker to use an alternative expression,
which we now derive, for the Christoffel symbol in terms of the metric tensor gij
and its derivatives with respect to the coordinates.
Firstly we note that the Christoffel symbol Γk
ijis symmetric with respect to
the interchange of its two subscripts iandj. This is easily shown: since
∂ei
∂uj=∂2r
∂uj∂ui=∂2r
∂ui∂uj=∂ej
∂ui,
it follows from (21.74) that Γk
ijek=Γk
jiek. Taking the scalar product with eland
using the reciprocity relation ek·el=δl
kgives immediately that
Γl
ij=Γl
ji.
To obtain an expression for Γk
ijwe then use gij=ei·ejand consider the
derivative
∂gij
∂uk=∂ei
∂uk·ej+ei·∂ej
∂uk
=Γl
ikel·ej+ei·Γl
jkel
=Γl
ikglj+Γl
jkgil, (21.78)
815
TENSORS
where we have used the definition (21.74). By cyclically permuting the free indices
i, j, kin (21.78), we obtain two further equivalent relations,
∂gjk
∂ui=Γl
jiglk+Γl
kigjl (21.79)
and
∂gki
∂uj=Γl
kjgli+Γl
ijgkl. (21.80)
If we now add (21.79) and (21.80) together and subtract (21.78) from the result,
we find
∂gjk
∂ui+∂gki
∂uj−∂gij
∂uk=Γl
jiglk+Γl
kigjl+Γl
kjgli+Γl
ijgkl−Γl
ikglj−Γl
jkgil
=2 Γl
ijgkl,
where we have used the symmetry properties of both Γl
ijandgij. Contracting
both sides with gmkleads to the required expression for the Christoffel symbol in
terms of the metric tensor and its derivatives, namely
Γm
ij=1
2gmkparenleftbigg∂gjk
∂ui+∂gki
∂uj−∂gij
∂ukparenrightbigg
. (21.81)ICalculate the Christoffel symbols Γm
ijfor cylindrical polar coordinates.
We may use either (21.74) or (21.81) to calculate the Γm
ijfor this simple coordinate system.
In cylindrical polar coordinates ( u1,u2,u3)=( ρ, φ, z), the basis vectors eiare given by
(21.58). It is straightforward to show that the only derivatives of these vectors with respectto the coordinates that are non-zero are
∂e
ρ
∂φ=1
ρeφ,∂eφ
∂ρ=1
ρeφ,∂eφ
∂φ=−ρeρ.
Thus, from (21.74), we have immediately that
Γ2
12=Γ2
21=1
ρand Γ1
22=−ρ. (21.82)
Alternatively, using (21.81) and the fact that g11=1 , g22=ρ2,g33= 1 and the other
components are zero, we see that the only three non-zero Christoffel symbols are indeedΓ
2
12=Γ2
21and Γ1
22.T h e s ea r eg i v e nb y
Γ2
12=Γ2
21=1
2g22∂g22
∂u1=1
2ρ2∂
∂ρ(ρ2)=1
ρ,
Γ1
22=−1
2g11∂g22
∂u1=−1
2∂
∂ρ(ρ2)=−ρ,
which agree with the expressions found directly from (21.74) and given in (21.82 ).
J
816
21.19 COVARIANT DIFFERENTIATION
21.19 Covariant differentiation
For Cartesian tensors we noted that the derivative of a scalar is a (covariant)
vector. This is also true for general tensors, as may be shown by considering thedifferential of a scalar
dφ=∂φ
∂uidui.
Since the duiare the components of a contravariant vector and dφis a scalar,
we have by the quotient law, discussed in section 21.7, that the quantities ∂φ/∂ui
must form the components of a covariant vector.
As a second example, in Cartesian coordinates, if the viare the contravariant
components of a vector then the quantites ∂vi/∂xjform the components of a
second-order tensor. It is straightforward, however, to show that (in contrast to
what happens in Cartesian coordinates) the differentiation of the components ofa general tensor, other than a scalar, with respect to the coordinates does notin
general result in the components of another tensor.IShow that, in general coordinates, the quantities ∂vi/∂ujdo not form the components of
a tensor.
We may show this directly by considering/∂vi
∂uj
//prime
=∂v/primei
∂u/primej=∂uk
∂u/primej∂v/primei
∂uk
=∂uk
∂u/primej∂
∂uk
/
∂u/primei
∂ulvl
/!
=∂uk
∂u/primej∂u/primei
∂ul∂vl
∂uk+∂uk
∂u/primej∂2u/primei
∂uk∂ulvl. (21.83)
The presence of the second term on the right-hand side of (21.83) shows that the ∂vi/∂xj
do not form the components of a second-order tensor. This term arises because the
‘transformation matrix’ [ ∂u/primei/∂uj] changes as the position in space at which it is evaluated
is changed. This is not true in Cartesian coordinates, for which the second term vanishes,and∂v
i/∂xjis a second-order tensor.
J
We may, however, use the Christoffel symbols discussed in the previous section
to define a new covariant derivative of the components of a tensor that does
result in the components of another tensor.
Let us first consider the derivative of a vector vwith respect to the coordinates.
Writing the vector in terms of its contravariant components v=viei, we find
∂v
∂uj=∂vi
∂ujei+vi∂ei
∂uj, (21.84)
where the second term arises because, in general, the basis vectors eiare not
constant (this term vanishes in Cartesian coordinates). Using (21.74) we may
817
TENSORS
write
∂v
∂uj=∂vi
∂ujei+viΓk
ijek.
Since iandkare dummy indices in the last term on the right-hand side, we may
interchange them to obtain
∂v
∂uj=∂vi
∂ujei+vkΓi
kjei=parenleftbigg∂vi
∂uj+vkΓi
kjparenrightbigg
ei. (21.85)
The reason for the interchanging the dummy indices, as shown in (21.85), is that
we may now factor out ei. The quantity in parentheses is called the covariant
derivative , for which the standard notation is
vi
;j≡∂vi
∂uj+Γi
kjvk, (21.86)
the semicolon subscript denoting covariant differentiation. A similar short-hand
notation also exists for the partial derivatives, a comma being used for these
instead of a semicolon; for example, ∂vi/∂ujis denoted by vi
,j. In Cartesian
coordinates all the Γi
kjare zero, and so the covariant derivative reduces to the
simple partial derivative ∂vi/∂uj.
Using the short-hand semicolon notation, the derivative of a vector may be
written in the very compact form
∂v
∂uj=vi
;jei
and, by the quotient rule (section 21.7), it is clear that the vi
;jare the (mixed)
components of a second-order tensor. This may also be verified directly, usingthe transformation properties of ∂v
i/∂ujand Γi
kjgiven in (21.83) and (21.77)
respectively.
In general, we may regard the vi
;jas the mixed components of a second-
order tensor called the covariant derivative of vand denoted by ∇v. In Cartesian
coordinates, the components of this tensor are just ∂vi/∂xj.ICalculate vi
;iin cylindrical polar coordinates.
Contracting (21.86) we obtain
vi
;i=∂vi
∂ui+Γi
kivk.
Now from (21.82) we have
Γi
1i=Γ1
11+Γ2
12+Γ3
13=1/ρ,
Γi
2i=Γ1
21+Γ2
22+Γ3
23=0,
Γi
3i=Γ1
31+Γ2
32+Γ3
33=0,
818
21.19 COVARIANT DIFFERENTIATION
and so
vi
;i=∂vρ
∂ρ+∂vφ
∂φ+∂vz
∂z+1
ρvρ
=1
ρ∂
∂ρ(ρvρ)+∂vφ
∂φ+∂vz
∂z.
This result is identical to the expression for the divergence of a vector field in cylindrical
polar coordinates given in section 10.9. This is discussed further in section 21.20.
J
So far we have considered only the covariant derivative of the contravariant
components viof a vector. The corresponding result for the covariant components
vimay be found in a similar way, by considering the derivative of v=vieiand
using (21.76) to obtain
vi;j=∂vi
∂uj−Γk
ijvk. (21.87)
Comparing the expressions (21.86) and (21.87) for the covariant derivative
of the contravariant and covariant components of a vector respectively, we seethat there are some similarities and some differences. It may help to rememberthat the index with respect to which the covariant derivative is taken ( jin this
case), is also the last subscript on the Christoffel symbol; the remaining indicescan then be arranged in only one way without raising or lowering them. It
remains to remember the sign difference, i.e. that for a covariant index (subscript)
the Christoffel symbol carries a minus sign, whereas for a contravariant index(superscript) the sign is positive.
Following a similar procedure to that which led to equation (21.86), we may
obtain expressions for the covariant derivatives of higher-order tensors.IBy considering the derivative of the second-order tensor Twith respect to the coordinate
uk, find an expression for the covariant derivative Tij
;kof its contravariant components.
Expressing Tin terms of its contravariant components, we have
∂T
∂uk=∂
∂uk(Tijei⊗ej)
=∂Tij
∂ukei⊗ej+Tij∂ei
∂uk⊗ej+Tijei⊗∂ej
∂uk.
Using (21.74), we can rewrite the derivatives of the basis vectors in terms of Christoffel
symbols to obtain
∂T
∂uk=∂Tij
∂ukei⊗ej+TijΓl
ikel⊗ej+Tijei⊗Γl
jkel.
Interchanging the dummy indices iandlin the second term and jandlin the third term
on the right-hand s ide, this becomes
∂T
∂uk=
/∂Tij
∂uk+Γi
lkTlj+Γj
lkTil
/
ei⊗ej,
819
TENSORS
where the expression in brackets is the required covariant derivative
Tij
;k=∂Tij
∂uk+Γi
lkTlj+Γj
lkTil. (21.88)
Using (21.88), the derivative of the tensor Twith respect to ukcan now be written in terms
of its contravariant components as
∂T
∂uk=Tij
;kei⊗ej.
J
Results similar to (21.88) may be obtained for the the covariant derivatives of
the mixed and covariant components of a second-order tensor. Collecting these
results together, we have
Tij
;k=Tij
,k+Γi
lkTlj+Γj
lkTil,
Ti
j;k=Ti
j,k+Γi
lkTl
j−Γl
jkTi
l,
Tij;k=Tij, k−Γl
ikTlj−Γl
jkTil,
where we have used the comma notation for partial derivatives. The position of
the indices in these expressions is very systematic: for each contravariant index(superscript) on the LHS we add a term on the RHS containing a Christoffelsymbol with a plus sign, and for every covariant index (subscript) we add acorresponding term with a minus sign. This is extended straightforwardly totensors with an arbitrary number of contravariant and covariant indices.
We note that the quantities T
ij
;k,Ti
j;kandTij;kare the components of the
samethird-order tensor ∇Twith respect to different tensor bases, i.e.
∇T=Tij
;kei⊗ej⊗ek=Ti
j;kei⊗ej⊗ek=Tij;kei⊗ej⊗ek.
We conclude this section by considering briefly the covariant derivative of a
scalar. The covariant derivative differs from the simple partial derivative withrespect to the coordinates only because the basis vectors of the coordinate
system change with position in space (hence for Cartesian coordinates there is no
difference). However, a scalar φdoes not depend on the basis vectors at all and
so its covariant derivative must be the same as its partial derivative, i.e.
φ
;j=∂φ
∂uj=φ,j. (21.89)
21.20 Vector operators in tensor form
In section 10.10 we used vector calculus methods to find expressions for vector
differential operators, such as grad, div, curl and the Laplacian, in general orthog-
onalcurvilinear coordinates, taking cylindrical and spherical polars as particular
examples. In this section we use the framework of general tensors that we have
developed to obtain, in tensor form, expressions for these operators that are validinallcoordinate systems, whether orthogonal or not.
820
21.20 VECTOR OPERATORS IN TENSOR FORM
In order to compare the results obtained here with those given in section
10.10 for orthogonal coordinates, it is necessary to remember that here we areworking with the (in general) non-unit basis vectors e
i=∂r/∂uiorei=∇ui.
Thus the components of a vector v=vieiare not the same as the components ˆvi
appropriate to the corresponding unit basis ˆei. In fact, if the scale factors of the
coordinate system are hi,i=1,2,3, then vi=ˆvi/hi(no summation over i).
As mentioned in section 21.15, for an orthogonal coordinate system with scale
factors hiwe have
gij=braceleftBigg
h2
iifi=j,
0o t h e r w i s eand gij=braceleftBigg
1/h2
iifi=j,
0o t h e r w i s e ,
a n ds ot h ed e t e r m i n a n t gof the matrix [ gij]i sg i v e nb y g=h2
1h22h23.
Gradient
The gradient of a scalar φis given by
∇φ=φ;iei=∂φ
∂uiei, (21.90)
since the covariant derivative of a scalar is the same as its partial derivative.
Divergence
Replacing the partial derivatives that occur in Cartesian coordinates with covari-
ant derivatives, the divergence of a vector field vin a general coordinate system
is given by
∇·v=vi
;i=∂vi
∂ui+Γi
kivk.
Using the expression (21.81) for the Christoffel symbol in terms of the metric
tensor, we find
Γi
ki=1
2gilparenleftbigg∂gil
∂uk+∂gkl
∂ui−∂gki
∂ulparenrightbigg
=1
2gil∂gil
∂uk. (21.91)
The last two terms have cancelled because
gil∂gkl
∂ui=gli∂gki
∂ul=gil∂gki
∂ul,
where in the first equality we have interchanged the dummy indices iandl,a n d
in the second equality have used the symmetry of the metric tensor.
We may simplify (21.91) still further by using a result concerning the derivative
of the determinant of a matrix whose elements are functions of the coordinates.
821
TENSORSISuppose A=[aij],B=[bij]and that B=A−1. By considering the determinant a=|A|,
show that
∂a
∂uk=abji∂aij
∂uk.
If we denote the cofactor of the element aijby ∆ijthen the elements of the inverse matrix
are given by (see chapter 8)
bij=1
a∆ji. (21.92)
However, the determinant of Ais given by
a=
X
jaij∆ij,
in which we have fixed iand written the sum over jexplicitly, for clarity. Partially
differentiating both sides with respect to aij, we then obtain
∂a
∂aij=∆ij, (21.93)
since aijdoes not occur in any of the cofactors ∆ij.
Now, if the aijdepend on the coordinates then so will the determinant aand, by the
chain rule, we have
∂a
∂uk=∂a
∂aij∂aij
∂uk=∆ij∂aij
∂uk=abji∂aij
∂uk, (21.94)
in which we have used (21.92) and (21.93).
J
Applying the result (21.94) to the determinant gof the metric tensor, and
remembering both that gikgkj=δi
jand that gijis symmetric, we obtain
∂g
∂uk=ggij∂gij
∂uk. (21.95)
Substituting (21.95) into (21.91) we find that the expression for the Christoffel
symbol can be much simplified to give
Γi
ki=1
2g∂g
∂uk=1√g∂√g
∂uk.
Thus finally we obtain the expression for the divergence of a vector field in a
general coordinate system as
∇·v=vi
;i=1√g∂
∂uj(√gvj). (21.96)
Laplacian
If we replace vby∇φin∇·vthen we obtain the Laplacian ∇2φ. From (21.90),
we have
viei=v=∇φ=∂φ
∂uiei,
822
21.20 VECTOR OPERATORS IN TENSOR FORM
and so the covariant components of vare given by vi=∂φ/∂ui. In (21.96),
however, we require the contravariant components vi. These may be obtained by
raising the index using the metric tensor, to give
vj=gjkvk=gjk∂φ
∂uk.
Substituting this into (21.96) we obtain
∇2φ=1√g∂
∂ujparenleftbigg√ggjk∂φ
∂ukparenrightbigg
. (21.97)IUse (21.97) to find the expression for ∇2φin an orthogonal coordinate system with scale
factors hi,i=1,2,3.
For an orthogonal coordinate system√g=h1h2h3andgij=1/h2
iifi=jandgij=0
otherwise. Therefore, from (21.97) we have
∇2φ=1
h1h2h3∂
∂uj
/
h1h2h3
h2
j∂φ
∂uj
/!
,
which agrees with the results of section 10.10.
J
Curl
The special vector form of the curl of a vector field exists only in three dimensions.
We therefore consider a more general form valid in higher-dimensional spaces aswell. In a general space the operation curl vis defined by
(curlv)
ij=vi;j−vj;i,
which is an antisymmetric covariant tensor.
In fact the difference of derivatives can be simplified, since
vi;j−vj;i=∂vi
∂uj−Γl
ijvl−∂vj
∂ui+Γl
jivl
=∂vi
∂uj−∂vj
∂ui,
where the Christoffel symbols have cancelled because of their symmetry properties.
Thus curl vcan be written in terms of partial derivatives as
(curlv)ij=∂vi
∂uj−∂vj
∂ui.
Generalising slightly the discussion of section 21.17, in three dimensions we may
associate with this antisymmetric second-order tensor a vector with contravariant
823
TENSORS
components,
(∇×v)i=−1
2√g/epsilon1ijk(curlv)jk
=−1
2√g/epsilon1ijkparenleftbigg∂vj
∂uk−∂vk
∂ujparenrightbigg
=1√g/epsilon1ijk∂vk
∂uj;
this is the analogue of the expression in Cartesian coordinates discussed in
section 21.8.
21.21 Absolute derivatives along curves
In section 21.19 we discussed how to differentiate a general tensor with respect
to the coordinates and introduced the covariant derivative. In this section weconsider the slightly different problem of calculating the derivative of a tensor
along a curve r(t) that is parameterised by some variable t.
Let us begin by considering the derivative of a vector valong the curve. If we
introduce an arbitrary coordinate system u
iwith basis vectors ei,i=1,2,3, then
we may write v=vieia n ds oo b t a i n
dv
dt=dvi
dtei+videi
dt
=dvi
dtei+vi∂ei
∂ukduk
dt;
here the chain rule has been used to rewrite the last term on the right-hand side.
Using (21.74) to write the derivatives of the basis vectors in terms of Christoffelsymbols, we obtain
dv
dt=dvi
dtei+Γj
ikviduk
dtej.
Interchanging the dummy indices iandjin the last term, we may factor out the
basis vector, and we find
dv
dt=parenleftbiggdvi
dt+Γi
jkvjduk
dtparenrightbigg
ei.
The term in parentheses is called the absolute (orintrinsic ) derivative of the
components vialong the curve r(t)and is usually denoted by
δvi
δt≡dvi
dt+Γi
jkvjduk
dt=vi
;kduk
dt.
With this notation, we may write
dv
dt=δvi
δtei=vi
;kduk
dtei. (21.98)
824
21.22 GEODESICS
Using the same method, the absolute derivative of the covariant components
viof a vector is given by
δvi
δt≡vi;kduk
dt.
Similarly, the absolute derivatives of the contravariant, mixed and covariant
components of a second-order tensor Tare
δTij
δt≡Tij
;kduk
dt,
δTi
j
δt≡Ti
j;kduk
dt,
δTij
δt≡Tij;kduk
dt.
The derivative of Talong the curve r(t) may then be written in terms of, for
example, its contravariant components as
dT
dt=δTij
δtei⊗ej=Tij
;kduk
dtei⊗ej.
21.22 Geodesics
As an example of the use of the absolute derivative, we conclude this chapter
with a brief discussion of geodesics. A geodesic in real three-dimensional spaceis a straight line, which has two equivalent defining properties. Firstly, it is thecurve of shortest length between two points and, secondly, it is the curve whosetangent vector always points in the same direction (along the line). Althoughin this chapter we have considered explicitly only our familiar three-dimensional
space, much of the mathematical formalism developed can be generalised to more
abstract spaces of higher dimensionality in which the familiar ideas of Euclideangeometry are no longer valid. It is often of interest to find geodesic curves insuch spaces by using the defining properties of straight lines in Euclidean space.
We shall not consider these more complicated spaces explicitly but will de-
termine the equation that a geodesic in Euclidean three-dimensional space (i.e.a straight line) must satisfy, deriving it in a sufficiently general way that our
method may be applied with little modification to finding the equations satisfied
by geodesics in more abstract spaces.
Let us consider a curve r(s), parameterised by the arc length sfrom some point
on the curve, and choose as our defining property for a geodesic that its tangentvector t=dr/dsalways points in the same direction everywhere on the curve, i.e.
dt
ds=0. (21.99)
(We could alternatively exploit the property that the distance between two points
825
TENSORS
is a minimum along a geodesic and use the calculus of variations (see chapter 22);
this would lead to the same final result (21.100).)
If we now introduce an arbitrary coordinate system uiwith basis vectors ei,
i=1,2,3, then we may write t=tieiand from (21.98) we find
dt
ds=ti
;kduk
dsei=0.
Writing out the covariant derivative, we obtain
parenleftbiggdti
ds+Γi
jktjduk
dsparenrightbigg
ei=0.
But, since tj=duj/ds, it follows that the equation satisfied by a geodesic is
d2ui
ds2+Γi
jkduj
dsduk
ds=0. (21.100)IFind the equations satisfied by a geodesic (straight line) in cylindrical polar coordinates.
From (21.82), the only non-zero Christoffel symbols are Γ1
22=−ρand Γ2
12=Γ2
21=1/ρ.
Thus the required geodesic equations are
d2u1
ds2+Γ1
22du2
dsdu2
ds=0⇒d2ρ
ds2−ρ
/dφ
ds
/2
=0,
d2u2
ds2+2 Γ2
12du1
dsdu2
ds=0⇒d2φ
ds2+2
ρdρ
dsdφ
ds=0,
d2u3
ds2=0⇒d2z
ds2=0.
J
21.23 Exercises
21.1 (a) Show that for any general, but fixed, φ,
(u1,u2)=(x1cosφ−x2sinφ, x 1sinφ+x2cosφ)
are the components of a first-order tensor in two dimensions.
(b) Show that/
x2
2x1x2
x1x2x21
/
is not a (Cartesian) tensor of order 2. To establish that a single element does
not transform correctly is sufficient.
21.2 The components of two vectors AandBand a second-order tensor Tare given
in one coordinate system by
A=
/0/@1
00
/1A,B=
/0/@0
10
/1A,T=
/0/@2√30√340
00 2
/1A.
826
21.23 EXERCISES
In a second coordinate system, obtained from the first by rotation, the components
ofAandBare
A/prime=1
2
/0/@√3
01
/1A,B/prime=1
2
/0/@−1
0√3
/1A.
Find the components of Tin this new coordinate system and hence evaluate,
with a minimum of calculation,
TijTji,T kiTjkTij,T ikTmnTniTkm.
21.3 In section 21.3 the transformation matrix for a rotation of the coordinate axes
was derived, and this approach is used in the rest of the chapter. An alternative
view is that of taking the coordinate axes as fixed and rotating the componentsof the system; this is equivalent to reversing the signs of all rotation angles.
Using this alternative view, determine the matrices representing (a) a positive
rotation of π/4 about the x-axis, and (b) a rotation of −π/4 about the y-axis.
Determine the initial vector rwhich, when subjected to (a) followed by (b),
finishes at (3 ,2,1).
21.4 Show how to decompose the tensor T
ijinto three tensors,
Tij=Uij+Vij+Sij,
where Uijis symmetric and has zero trace, Vijis isotropic and Sijhas only three
independent components.
21.5 Use the quotient law discussed in section 21.7 to show that the array/0/@y2+z2−x2−2xy −2xz
−2yx x2+z2−y2−2yz
−2zx −2zy x2+y2−z2
/1A
forms a second-order tensor.
21.6 Use tensor methods to establish the following vector identities:
(a) (u×v)×w=(u·w)v−(v·w)u;
(b) curl( φu)=φcurlu+ (grad φ)×u;
(c) div ( u×v)=v·curlu−u·curlv;
(d) curl( u×v)=(v·grad)u−(u·grad)v+udivv−vdivu;
(e) grad1
2(u·u)=u×curlu+(u·grad)u.
21.7 Use result (e) of the previous question and the general divergence theorem for
tensors to show thatZ
S
/
A(A·dS)−1
2A2dS
/
=
Z
V[AdivA−A×curlA]dV.
21.8 A column matrix ahas components ax,ay,azand Ais the matrix with elements
Aij=−/epsilon1ijkak.
(a) What is the relationship between column matrices band cifAb=c?
(b) Find the eigenvalues of Aand show that ais one of its eigenvectors. Explain
why this must be so.
21.9 Equation (21.28),
|A|/epsilon1lmn=AliAmjAnk/epsilon1ijk,
is a more general form of the expression (8.47) for the determinant of a 3 ×3
matrix A. The latter could have been written as
|A|=/epsilon1ijkAi1Aj2Ak3,
827
TENSORS
whilst the former removes the explicit mention of 1 ,2,3 at the expense of an
additional Levi–Civita symbol. As stated in the footnote on p. 791, (21.28) canbe readily extended to cover a general N×Nmatrix.
Use the form given in (21.28) to prove properties (i), (iii), (v), (vi) and (vii)
of determinants stated in subsection 8.9.1. Property (iv) is obvious by inspection.For definiteness take N= 3, but convince yourself that your methods of proof
would be valid for any positive integer N.
21.10 A symmetric second-order Cartesian tensor is defined by
T
ij=δij−3xixj.
Evaluate the following surface integrals, each taken over the surface of the unit
sphere:
(a)
Z
TijdS;( b )
Z
TikTkjdS;( c )
Z
xiTjkdS.
21.11 Given a non-zero vector v, find the value that should be assigned to αto make
Pij=αvivjand Qij=δij−αvivj
into parallel and orthogonal projection tensors respectively, i.e. tensors that satisfy
respectively Pijvj=vi,Pijuj=0a n d Qijvj=0 , Qijuj=ui, for any vector uthat
is orthogonal to v,
Show, in particular, that Qijis unique, i.e. that if another tensor Tijhas the
same properties as Qijthen ( Qij−Tij)wj=0f o r anyvector w.
21.12 In four dimensions define second-order antisymmetric tensors FijandQijand a
first-order tensor Sias follows:
(a)F23=H1,Q23=B1and their cyclic permutations;
(b)Fi4=−Di,Qi4=Eifori=1,2,3;
(c)S4=ρ,Si=Jifori=1,2,3.
Then, taking x4astand the other symbols to have their usual meanings in
electromagnetic theory, show that the equations
P
j∂Fij/∂x j=Siand∂Qjk/∂x i+
∂Qki/∂x j+∂Qij/∂x k= 0 reproduce Maxwell’s equations. Here i, j, kis any set of
three subscripts selected from 1 ,2,3,4, but chosen in such a way that they are
all different.
21.13 In a certain crystal the unit cell can be taken as six identical atoms lying at the
corners of a regular octahedron. Convince yourself that these atoms can also beconsidered as lying at the centres of the faces of a cube and hence that the crystalhas cubic symmetry. Use this result to prove that the conductivity tensor for the
crystal, σ
ij,m u s tb ei s o t r o p i c .
21.14 Assuming that the current density jand the electric field Eappearing in equation
(21.43) are first-order Cartesian tensors, show explicitly that the electrical con-ductivity tensor σ
ijtransforms according to the law appropriate to a second-order
tensor.
The rate Wat which energy is dissipated per unit volume, as a result of the
current flow, is given by E·j. Determine the limits between which Wmust lie for
a given value of |E|as the direction of Eis varied.
21.15 In a certain system of units the electromagnetic stress tensor Mijis given by
Mij=EiEj+BiBj−1
2δij(EkEk+BkBk),
where the electric and magnetic fields, EandB, are first-order tensors. Show that
Mijis a second-order tensor.
Consider a situation in which |E|=|B|but the directions of EandBare
not parallel. Show that E±Bare principal axes of the stress tensor and find
828
21.23 EXERCISES
the corresponding principal values. Determine the third principal axis and its
corresponding principal value.
21.16 A rigid body consists of four particles of masses m,2m,3m,4m, respectively
situated at the points ( a, a, a), (a,−a,−a), (−a, a,−a), (−a,−a, a) and connected
together by a light framework.
(a) Find the inertia tensor at the origin and show that the principal moments of
inertia are 20 ma2,a n d( 2 0±2√
5)ma2.
(b) Find the principal axes and verify that they are orthogonal.
21.17 A rigid body consists of eight particles, each of mass m, held together by light
rods. In a certain coordinate frame the particles are at
±a(3,1,−1),±a(1,−1,3),a(1,3,−1),a(−1,1,3).
Show that, when the body rotates about an axis through the origin, if the angular
velocity and angular momentum vectors are parallel then their ratio must be40ma
2,6 4ma2or 72 ma2.
21.18 The paramagnetic tensor χijof a body placed in a magnetic field, in which its
energy density is −1
2µ0M·Hwith Mi=
P
jχijHj,i s/0/@2k00
03 kk
0 k3k
/1A.
Assuming depolarizing effects are negligible, find how the body will orientate
itself if the field is horizontal, in the following circumstances:
(a) the body can rotate freely;
(b) the body is suspended with the (1,0,0) axis vertical;(c) the body is suspended with the (0,1,0) axis vertical.
21.19 A block of wood contains a number of thin soft iron nails (of constant permeabil-
ity). A unit magnetic field directed eastwards induces a magnetic moment in theblock having components (3 ,1,−2) and similar fields directed northwards and
vertically upwards induce moments (1 ,3,−2) and (−2,2,2) respectively. Show
that all the nails lie in parallel planes.
21.20 For tin the conductivity tensor is diagonal, with entries a, a,andbwhen referred
to its crystal axes. A single crystal is grown in the shape of a long wire of length L
and radius r, the axis of the wire making polar angle θwith respect to the crystal’s
3-axis. Show that the resistance of the wire is L(πr
2ab)−1
/;
acos2θ+bsin2θ
/
.
21.21 By considering an isotropic body subjected to a uniform hydrostatic pressure
(no shearing stress), show that the bulk modulus k, defined by the ratio of the
pressure to the fractional decrease in volume, is given by k=E/[3(1−2σ)] where
Eis Young’s modulus and σPoisson’s ratio.
21.22 For an isotropic elastic medium under dynamic stress, at time tthe displacement
uiand the stress tensor pijsatisfy
pij=cijkl
/∂uk
∂xl+∂ul
∂xk
/
and∂pij
∂xj=ρ∂2ui
∂t2,
where cijklis the isotropic tensor given in equation (21.46) and ρis a constant.
Show that both ∇·uand∇×usatisfy wave equations and find the corresponding
wave speeds.
829
TENSORS
21.23 A fourth-order tensor Tijklhas the properties
Tjikl=−Tijkl,T ijlk=−Tijkl.
Prove that for any such tensor there exists a second-order tensor Kmnsuch that
Tijkl=/epsilon1ijm/epsilon1klnKmn
and give an explicit expression for Kmn. Consider two (separate) special cases, as
follows.
(a) Given that Tijklis isotropic and Tijji= 1, show that Tijklis uniquely deter-
mined and express it in terms of Kronecker deltas.
(b) If now Tijklhas the additional property
Tklij=−Tijkl,
show that Tijklhas only three linearly independent components and find an
expression for Tijklin terms of the vector
Vi=−1
4/epsilon1jklTijkl.
21.24 Working in cylindrical polar coordinates ρ, φ, z, parameterise the straight line
(geodesic) joining (1 ,0,0) to (1 ,π/2,1) in terms of s, the distance along the line.
Show by substitution that the geodesic equations derived at the end of section21.22 are satisfied.
21.25 In a general coordinate system u
i,i=1,2,3, in three-dimensional Euclidean
space, a volume element is given by
dV=|e1du1·(e2du2×e3du3)|.
Show that an alternative form for this expression, written in terms of the deter-
minant gof the metric tensor, is given by
dV=√gd u1du2du3.
Show that under a general coordinate transformation to a new coordinate system
u/primeithe volume element dVremains unchanged, i.e. show that it is a scalar quantity.
21.26 By writing down the expression for the square of the infinitesimal arc length ( ds)2
in spherical polar coordinates, find the components gijof the metric tensor in this
coordinate system. Hence, using (21.96), find the expression for the divergenceof a vector field vin spherical polars. Calculate the Christoffel symbols (of the
second kind) Γ
i
jkin this coordinate system.
21.27 Find an expression for the second covariant derivative vi;jk≡(vi;j);kof a vector
vi(see(21.86)). By interchanging the order of differentiation and then subtracting
the two expressions, we define the components Rl
ijkof the Riemann tensor as
vi;jk−vi;kj≡Rl
ijkvl.
Show that in a general coordinate system uithese components are given by
Rl
ijk=∂Γl
ik
∂uj−∂Γl
ij
∂uk+Γm
ikΓl
mj−Γm
ijΓl
mk.
By first considering Cartesian coordinates, show that all the components Rl
ijk≡0
foranycoordinate system in three-dimensional Euclidean space.
In such a space, therefore, we may change the order of the covariant derivativeswithout changing the resulting expression.
830
21.24 HINTS AND ANSWERS
21.28 A curve r(t) is parameterised by a scalar variable t. Show that the length of the
curve between two points, AandB,i sg i v e nb y
L=
ZB
A
r
gijdui
dtduj
dtdt.
Using the calculus of variations (see chapter 22), show that the curve r(t)t h a t
minimises Lsatisfies the equation
d2ui
dt2+Γi
jkduj
dtduk
dt=¨s
˙sdui
dt,
where sis the arc length along the curve, ˙s=ds/dt and¨s=d2s/dt2. Hence, show
that if the parameter tis of the form t=as+b,w h e r e aandbare constants,
then we recover the equation for a geodesic (21.100).
(A parameter which, like t, is the sum of a linearly transformation of sand a
translation is called an affine parameter.)
21.29 We may define Christoffel symbols of the first kind by
Γijk=gilΓl
jk.
Show that these are given by
Γijk=1
2
/∂gik
∂uj+∂gjk
∂ui−∂gij
∂uk
/
.
By permuting indices, verify that
∂gij
∂uk=Γ ijk+Γ jik.
Using the fact that Γl
jk=Γl
kj, show that
gij;k≡0,
i.e. that the covariant derivative of the metric tensor is identically zero in all
coordinate systems.
21.24 Hints and answers
21.1 (a) u/prime
1=x1cos(φ−θ)−x2sin(φ−θ), etc.;
(b)u/prime
11=s2x2
1−2scx1x2+c2x2
2/negationslash=c2x2
2+csx1x2+scx1x2+s2x2
1.
21.2 Determine entries for the third column of Lby requiring that it is orthogonal
and has determinant +1. T=1
2(√3,−1,0;0,0,−2;1√3,0). They are all scalars
with values 30, 134, 642.
21.3 (a) (1 /√2)(√2,0,0; 0,1,−1; 0,1,1). (b) (1 /√2)(1,0,−1;0,√2,0; 1,0,1).
r=( 2√2,−1+√2,−1−√2)T.
21.4 If T0is Tr Tijthen Uij=1
2(Tij+Tji)−1
3T0δij,Vij=1
3T0δij,Sij=1
2(Tij−Tji).
21.5 Twice contract the array with the outer product of ( x, y, z)w i t hi t s e l ft oo b t a i n
the expression −(x2+y2+z2)2, which is an invariant and therefore a scalar.
21.6 (a) /epsilon1ijk/epsilon1jlmulvmwkand use (21.29); (b) /epsilon1ijk∂(φuk)/∂x j;( c ) ∂(/epsilon1ijkujvk)/∂x i;
(d)/epsilon1ijk/epsilon1klm∂(ulvm)/∂x jand use (21.29); (e) start with u×curluand obtain
/epsilon1ijkuj/epsilon1klm∂um
∂xl=···=uj
/∂uj
∂xi
/
−uj
/∂ui
∂xj
/
.
21.7 Write Aj(∂Ai/∂x j)a s∂(AiAj)/∂x j−Ai(∂Aj/∂x j).
21.8 (a) c=a×b.( b )0 ,±i|a|.Aa=0asince a×a=0.
831
TENSORS
21.9 (i) Write out the expression for |AT|, contract both sides of the equation with /epsilon1lmn
and pick out the expression for |A|o nt h eR H S .N o t et h a t /epsilon1lmn/epsilon1lmnis a numerical
scalar.(iii) Each non-zero term on the RHS contains any particular row index once andonly once. The same can be said for the Levi–Civita symbol on the LHS. Thusinterchanging two rows is equivalent to interchanging two of the subscripts of/epsilon1
lmnand thereby reversing its sign. Consequently, the magnitude of |A|remains
the same but its sign is changed.(v) If, say, A
pi=λApj, for some particular pair of values iandjand all pthen,
in the (multiple-) summation on the RHS, each Ankappears multiplied by (no
summation over iandj)
/epsilon1ijkAliAmj+/epsilon1jikAljAmi=/epsilon1ijkλAljAmj+/epsilon1jikAljλAmj=0,
since /epsilon1ijk=−/epsilon1jik. Consequently, grouped in this way all terms are zero and
|A|=0 .
(vi) Replace AmjbyAmj+λAljand note that λAliAljAnk/epsilon1ijk= 0 by virtue of
result (v).(vii) If C=AB,
|C|/epsilon1
lmn=AlxBxiAmyByjAnzBzk/epsilon1ijk.
Contract this with /epsilon1lmnand show that the RHS is equal to /epsilon1xyz|AT|/epsilon1xyz|B|.I tt h e n
follows from result (i) that |C|=|A||B|.
21.10 Note that
R
xidS=
R
(xi)3dS=0a n dt h a t
R
(xi)2dS=4π/3. (a) 0 (the two
contributions cancel when i=j); (b) 8 πδij; (c) 0 for all sets of i, j, k,w h e t h e ro r
not some or all are equal.
21.11 α=|v|−2. Note that the most general vector has components wi=λvi+µu(1)
i+νu(2)
i,
where both u(1)andu(2)are orthogonal to v.
21.12∇×H=J+˙D;∇·D=ρ;∇×E+˙B=0;∇·B=0 .
21.13 Construct the orthogonal transformation matrix Sfor the symmetry operation
of (say) a rotation of 2 π/3 about a body diagonal and, setting L=S−1=ST,
construct σ/prime=LσLTand require σ/prime=σ. Repeat the procedure for (say) a rotation
ofπ/2 about the x3-axis. These together show that σ11=σ22=σ33and that
all other σij= 0. Further symmetry requirements do not provide any additional
constraints.
21.14 W=EiσijEjhas to be maximised or minimised subject to EiEibeing held
constant. Extreme values are W±=λ±|E|2,w h e r e λ±are the maximum and
minimum eigenvalues of the matrix σij.
21.15 The transformation of δijhas to be included; the principal values are ±E·B.
The third axis is in the direction ±B×Ewith principal value −|E|2.
21.16 (b) xT
1=( 2−10 ) ,xT
2=( 1 2√
5),xT
3=( 1 2−√
5).
21.17 The principal moments give the required ratios.21.18 The principal susceptibilit ies and (unnormalised) axes are λ=4 ,±(0,1,1);
λ=2 ,±(c
i,1,−1) with c1c2=−2, leading to:
(a) lowest energy when (0 ,1,1) axis is parallel to the field;
(b) permitted values o f orientation are (0 ,n2,n3), hence as in (a);
(c) permitted values o f orientation are ( n1,0,n3), subject to n2
1+n2
3=1 .
The energy = −1
2µ0kH2V(2n2
1+3n2
3), which is minimised when (0 ,0,1) is parallel
to the field.
21.19 The principal permeability, in direction (1 ,1,2), has value 0. Thus all the nails lie
in planes to which this is the normal.
21.20 ji=σikEkgives Isinθcosφ=aπr2E1,Isinθsinφ=aπr2E2,Icosθ=bπr2E3.
Also V/L=E1sinθcosφ+E2sinθsinφ+E3cosθ. The current must flow along
the wire; Eis not parallel to the wire.
832
21.24 HINTS AND ANSWERS
21.21 Take p11=p22=p33=−p,a n d pij=eij=0f o r i/negationslash=j, leading to −p=
(λ+2µ/3)eii. The fractional volume change is eii;λandµare as defined in (21.45)
and the worked example that follows it.
21.22 Show that pij=2λδij∇·u+(η+ν)(∂ui/∂x j+∂uj/∂x i). Form the sum of the
derivatives
P
j(∂/∂x j) for this equation, substitute for ∂pij/∂x jand then formP
i(∂/∂x i) of the result. The wave speed for ∇·uis [2( λ+η+ν)/ρ]1/2.Show that
ρ∂2(∇×u)k
∂t2=/epsilon1kji∂2pil
∂xj∂xl
and then use the previous expression for pijand the identity /epsilon1kji∂2/∂x j∂xi=0 .
The wave speed for ∇×uis [(η+ν)/ρ]1/2.
21.23 Consider Qpq=/epsilon1pij/epsilon1qklTijkland show that Kmn=Qmn/4 has the required property.
(a) Argue from the isotropy of Tijkland/epsilon1ijkfor that of Kmnand hence that it
must be a multiple of δmn. Show that the multiplier is uniquely determined and
thatTijkl=(δilδjk−δikδjl)/6.
(b) By relabelling dummy subscripts and using the stated antisymmetry property,show that K
nm=−Kmn. Show that −2Vi=/epsilon1minKmnand hence that Kmn=/epsilon1imnVi.
Tijkl=/epsilon1kliVj−/epsilon1kljVi.
21.24 ρ=( 1−2s/√
3+2 s2/3)1/2,φ=t a n−1[s/(√
3−s)],z=s/√
3.
21.25 Use |e1·(e2×e3)|=√g.
Recall that√g/prime=|∂u/∂u/prime|√ganddu/prime1du/prime2du/prime3=|∂u/prime/∂u|du1du2du3.
21.26 g=r4sin2θ; recall that, for each i,vi=ˆvi/hi,e . g . v3=vφ/(rsinθ).
Γ1
22=−r;Γ1
33=−rsin2θ;Γ2
12=r−1;Γ2
32=−sinθcosθ;Γ3
13=r−1;Γ3
23=
cotθ.
21.27 ( vi;j);k=(vi;j),k−Γl
ikvl;j−Γl
jkvi;landvi;j=vi, j−Γm
ijvm. If all components of a
tensor equal zero in one coordinate system then they are zero in all coordinatesystems.
21.28 Using ˙s=p
gij˙ui˙uj, the Euler–Lagrange equation is
d
dt
/gik˙ui
˙s
/
−1
2˙s∂gij
∂uk˙ui˙uj=0.
Calculate the t-derivative, write
∂gik
∂uj=1
2
/∂gik
∂uj+∂gjk
∂ui
/
and multiply through by glk.I ft=as+bthen¨s=0 .
833
22
Calculus of variations
In chapters 2 and 5 we discussed how to find stationary values of functions of a
single variable f(x), of several variables f(x ,y,... ) and of constrained variables,
where x ,y,... are subject to the nconstraints gi(x ,y,... )=0 , i=1,2,...,n.I na l l
these cases the forms of the functions fandgiwere known, and the problem was
one of finding the appropriate values of the variables x,yetc.
We now turn to a different kind of problem in which we are interested in
bringing about a particular condition for a given expression (usually maximising
or minimising it) by varying the functions on which the expression depends. For
instance, we might want to know in what shape a fixed length of rope shouldbe arranged so as to enclose the largest possible area, or in what shape it willhang when suspended under gravity from two fixed points. In each case we areconcerned with a general maximisation or minimisation criterion by which thefunction y(x) that satisfies the given problem may be found.
The calculus of variations provides a method for finding the function y(x).
The problem must first be expressed in a mathematical form, and the formmost commonly applicable to such problems is an integral . In each of the above
questions, the quantity that has to be maximised or minimised by an appropriatechoice of the function y(x) may be expressed as an integral involving y(x)a n d
the variables describing the geometry of the situation.
In our example of the rope hanging from two fixed points, we need to find
the shape function y(x) that minimises the gravitational potential energy of the
rope. Each elementary piece of the rope has a gravitational potential energyproportional both to its vertical height above an arbitrary zero level and to thelength of the piece. Therefore the total potential energy is given by an integralfor the whole rope of such elementary contributions. The particular function y(x)
for which the value of this integral is a minimum will give the shape assumed by
the hanging rope.
So in general we are led by this type of question to study the value of an
834
22.1 THE EULER–LAGRANGE EQUATION
y
x a b
Figure 22.1 Possible paths for the integral (22.1). The solid line is the curve
along which the integral is assumed stationary. The broken curves represent
small variations from this path.
integral whose integrand has a specified form in terms of a certain function
and its derivatives, and to study how that value changes when the form ofthe function is varied. Specifically, we aim to find the function that makes theintegral stationary , i.e. the function that makes the value of the integral a local
maximum or minimum. Note that, unless stated otherwise, y
/primeis used to denote
dy/dx throughout this chapter. We also assume that all the functions we need to
deal with are sufficiently smooth and differentiable.
22.1 The Euler–Lagrange equation
Let us consider the integral
I=integraldisplayb
aF(y,y/prime,x)dx, (22.1)
where a,band the form of the function Fare fixed by given considerations,
e.g. the physics of the problem, but the curve y(x) is to be chosen so as to
make stationary the value of I, which is clearly a function (or more accurately a
functional ) of this curve, i.e. I=I[y(x)]. Referring to figure 22.1, we wish to find
the function y(x) (given, say, by the solid line) such that first-order small changes
in it (for example the two broken lines) will make only second-order changes in
the value of I.
Writing this in a more mathematical form, let us suppose that y(x)i st h e
function required to make Istationary and consider making the replacement
y(x)→y(x)+αη(x), (22.2)
where the parameter αis small and η(x) is an arbitrary function with sufficiently
amenable mathematical properties. For the value of Ito be stationary with respect
835
CALCULUS OF VARIATIONS
to these variations, we require
dI
dαvextendsinglevextendsinglevextendsinglevextendsingle
α=0=0 f o ra l l η(x). (22.3)
Substituting (22.2) into (22.1) and expanding as a Taylor series in αwe obtain
I(y,α)=integraldisplayb
aF(y+αη, y/prime+αη/prime,x)dx
=integraldisplayb
aF(y,y/prime,x)dx+integraldisplayb
aparenleftbigg∂F
∂yαη+∂F
∂y/primeαη/primeparenrightbigg
dx+O ( α2).
With this form for I(y,α) the condition (22.3) implies that for all η(x)w er e q u i r e
δI=integraldisplayb
aparenleftbigg∂F
∂yη+∂F
∂y/primeη/primeparenrightbigg
dx=0,
where δIdenotes the first-order variation in the value of Idue to the variation
(22.2) in the function y(x). Integrating the second term by parts this becomes
bracketleftbigg
η∂F
∂y/primebracketrightbiggb
a+integraldisplayb
abracketleftbigg∂F
∂y−d
dxparenleftbigg∂F
∂y/primeparenrightbiggbracketrightbigg
η(x)dx=0. (22.4)
In order to simplify the result we will assume, for the moment, that the end-points
are fixed, i.e. not only aandbare given but also y(a)a n d y(b). This restriction
means that we require η(a)=η(b) = 0, in which case the first term on the LHS of
(22.4) equals zero at both end-points. Since (22.4) must be satisfied for arbitrary
η(x), it is easy to see that we require
∂F
∂y=d
dxparenleftbigg∂F
∂y/primeparenrightbigg
. (22.5)
This is known as the Euler–Lagrange (EL) equation, and is a differential equation
fory(x),since the function Fis known.
22.2 Special cases
In certain special cases a first integral of the EL equation can be obtained for a
general form of F.
22.2.1 Fdoes not contain yexplicitly
In this case ∂F/∂y = 0, and (22.5) can be integrated immediately giving
∂F
∂y/prime=c o n s t a n t . (22.6)
836
22.2 SPECIAL CASES
A(a, y(a))dxdydsB(b, y(b))
Figure 22.2 An arbitrary path between two fixed points.IShow that the shortest curve joining two points is a straight line.
Let the two points be labelled Aand Band have coordinates ( a, y(a)) and ( b, y(b))
respectively (see figure 22.2). Whatever the shape of the curve joining AtoB, the length
of an element of path dsis given by
ds=
/
(dx)2+(dy)2
/1/2=( 1+ y/prime2)1/2dx,
and hence the total path length along the curve is given by
L=
Zb
a(1 +y/prime2)1/2dx. (22.7)
We must now apply the results of the previous section to determine that path which makes
Lstationary (clearly a minimum in this case). Since the integral does not contain y(or
indeed x) explicitly, we may use (22.6) to obtain
k=∂F
∂y/prime=y/prime
(1 +y/prime2)1/2.
where kis a constant. This is easily rearranged and integrated to give
y=k
(1−k2)1/2x+c,
which, as expected, is the equation of a straight line in the form y=mx+c,w i t h
m=k/(1−k2)1/2. The value of m(ork) can be found by demanding that the straight line
passes through the points AandBand is given by m=[y(b)−y(a)]/(b−a). Substituting
the equation of the straight line into (22.7) we find that, again as expected, the total pathlength is given by
L
2=[y(b)−y(a)]2+(b−a)2.
J
837
CALCULUS OF VARIATIONS
dxdydsy
x
Figure 22.3 A convex closed curve that is symmetrical about the x-axis.
22.2.2 Fdoes not contain xexplicitly
In this case, multiplying the EL equation (22.5) by y/primeand using
d
dxparenleftbigg
y/prime∂F
∂y/primeparenrightbigg
=y/primed
dxparenleftbigg∂F
∂y/primeparenrightbigg
+y/prime/prime∂F
∂y/prime
we obtain
y/prime∂F
∂y+y/prime/prime∂F
∂y/prime=d
dxparenleftbigg
y/prime∂F
∂y/primeparenrightbigg
.
But since Fis a function of yandy/primeonly, and not explicitly of x,t h eL H So f
this equation is just the total derivative of F,n a m e l y dF/dx . Hence, integrating
we obtain
F−y/prime∂F
∂y/prime=c o n s t a n t . (22.8)IFind the closed convex curve of length lthat encloses the greatest possible area.
Without any loss of generality we can assume that the curve passes through the origin,
and can further suppose that it is symmetric with respect to the x-axis; this assumption
is not essential. Using the distance salong the curve, measured from the origin, as the
independent variable and yas the dependent one, we have the boundary conditions
y(0) = y(l/2) = 0. The element of area shown in figure 22.3 is then given by
dA=yd x=y
/
(ds)2−(dy)2
/1/2,
a n dt h et o t a la r e ab y
A=2
Zl/2
0y(1−y/prime2)1/2ds; (22.9)
herey/primestands for dy/ds rather than dy/dx . Since the integrand does not contain sexplicitly,
838
22.2 SPECIAL CASES
we can use (22.8) to obtain a first integral of the EL equation for y,n a m e l y
y(1−y/prime2)1/2+yy/prime2(1−y/prime2)−1/2=k,
where kis a constant. On rearranging this gives
ky/prime=±(k2−y2)1/2,
which, using y(0) = 0, integrates to
y/k=s i n ( s/k). (22.10)
The other end-point, y(l/2) = 0, fixes the value of kasl/2πto yield
y=l
2πsin2πs
l.
From this we obtain dy=c o s ( 2 πs/l)dsand since ( ds)2=(dx)2+(dy)2we find also that
dx=±sin(2πs/l)ds. This in turn can be integrated and, using x(0) = 0, gives xin terms
ofsas
x−l
2π=−l
2πcos2πs
l.
We thus obtain the expected result that xandylie on the circle of radius l/(2π)g i v e nb y/
x−l
2π
/2
+y2=l2
4π2.
Substituting the solution (22.10) into the expression for the total area (22.9), it is easily
verified that A=l2/(4π). A much quicker derivation of this result is possible using plane
polar coordinates.
J
The previous two examples have been carried out in some detail, even though
the answers are more easily obtained in other ways, expressly so that the method
is transparent and the way in which it works can be filled in mentally at almostevery step. The next example, however, does not have such an intuitively obvioussolution.ITwo rings, each of radius a, are placed parallel with their centres 2bapart and on a
common normal. An axially symmetric soap film is formed between them but does not coverthe ends of the rings (see figure 22.4). Find the shape assumed by the film.
Creating the soap film requires an energy γper unit area (numerically equal to the surface
tension of the soap solution). So the stable shape of the soap film, i.e. the one thatminimises the energy, will also be the one that minimises the surface area (neglectinggravitational effects).
It is obvious that any convex surface, shaped such as that shown as the broken line in
figure 22.4( a) cannot be a minimum but it is not clear whether some shape intermediate
between the cylinder shown by solid lines in ( a), with area 4 πab(or twice this for the
double surface of the film), and the form shown in ( b), with area approximately 2 πa
2, will
produce a lower total area than both of these extremes. If there is such a shape (e.g. that
in figure 22.4( c)), then it will be that which best compromises between two requirements,
the need to minimise the ring-to-ring distance measured on the film surface ( a)a n dt h e
need to minimise the average waist measurement of the surface ( b).
We take cylindrical polar coordinates as in figure 22.4( c) and let the radius of the soap
film at height zbeρ(z)w i t h ρ(±b)=a. Counting only one side of the film, the element of
839
CALCULUS OF VARIATIONS
(a)( b)( c)b
−bz
ρ
a
Figure 22.4 Possible soap films between two parallel circular rings.
surface area between zandz+dzis
dS=2πρ
/
(dz)2+(dρ)2
/1/2,
so the total surface area is given by
S=2π
Zb
−bρ(1 +ρ/prime2)1/2dz. (22.11)
Since the integrand does not contain zexplicitly, we can use (22.8) to obtain an equation
forρthat minimises S,i . e .
ρ(1 +ρ/prime2)1/2−ρρ/prime2(1 +ρ/prime2)−1/2=k,
where kis a constant. Multiplying through by (1 + ρ/prime2)1/2, rearranging to find an explicit
expression for ρ/primeand integrating we find
cosh−1ρ
k=z
k+c.
where cis the constant of integration. Using the boundary conditions ρ(±b)=a,w e
require c=0a n d ksuch that a/k=c o s h b/k(ifb/ais too large, no such kcan be found).
Thus the curve that minimises the surface area is
ρ/k=c o s h ( z/k),
and in profile the soap film is a catenary (see section 22.4) with the minimum distance
from the axis equal to k.
J
22.3 Some extensions
It is quite possible to relax many of the restrictions we have imposed so far. For
example, we can allow end-points that are constrained to lie on given curves rather
than being fixed, or we can consider problems with several dependent and/or
independent variables or higher-order derivatives of the dependent variable. Eachof these extensions is now discussed.
840
22.3 SOME EXTENSIONS
22.3.1 Several dependent variables
Here we have F=F(y1,y/prime
1,y2,y/prime
2,...,y n,y/prime
n,x)w h e r ee a c h yi=yi(x). The analysis
in this case proceeds as before, leading to nseparate but simultaneous equations
for the yi(x),
∂F
∂yi=d
dxparenleftbigg∂F
∂y/prime
iparenrightbigg
,i =1,2,...,n . (22.12)
22.3.2 Several independent variables
With nindependent variables, we need to extremise multiple integrals of the form
I=integraldisplayintegraldisplay
···integraldisplay
Fparenleftbigg
y,∂y
∂x1,∂y
∂x2,...,∂y
∂xn,x1,x2,...,x nparenrightbigg
dx1dx2···dxn.
Using the same kind of analysis as before, we find that the extremising function
y=y(x1,x2,...,x n) must satisfy
∂F
∂y=nsummationdisplay
i=1∂
∂xiparenleftbigg∂F
∂yxiparenrightbigg
, (22.13)
where yxistands for ∂y/∂x i.
22.3.3 Higher-order derivatives
If in (22.1) F=F(y,y/prime,y/prime/prime,...,y(n),x) then using the same method as before
and performing repeated integration by parts, it can be shown that the requiredextremising function y(x) satisfies
∂F
∂y−d
dxparenleftbigg∂F
∂y/primeparenrightbigg
+d2
dx2parenleftbigg∂F
∂y/prime/primeparenrightbigg
−···+(−1)ndn
dxnparenleftbigg∂F
∂y(n)parenrightbigg
=0,(22.14)
provided that y=y/prime=···=y(n−1)= 0 at both end-points. If y, or any of its
derivatives, is not zero at the end-points then a corresponding contribution orcontributions will appear on the RHS of (22.14).
22.3.4 Variable end-points
We now discuss the very important generalisation to variable end-points. Suppose,
as before, we wish to find the function y(x) that extremises the integral
I=integraldisplay
b
aF(y,y/prime,x)dx,
but this time we demand only that the lower end-point is fixed, while we allow
y(b) to be arbitrary. Repeating the analysis of section 22.1, we find from (22.4)
841
CALCULUS OF VARIATIONS
∆x∆y
y(x)
h(x, y)=0y(x)+η(x)
b
Figure 22.5 Variation of the end-point balong the curve h(x, y)=0 .
that we require
bracketleftbigg
η∂F
∂y/primebracketrightbiggb
a+integraldisplayb
abracketleftbigg∂F
∂y−d
dxparenleftbigg∂F
∂y/primeparenrightbiggbracketrightbigg
η(x)dx=0. (22.15)
Obviously the EL equation (22.5) must still hold for the second term on the LHS
to vanish. Also, since the lower end-point is fixed, i.e. η(a)=0 ,t h efi r s tt e r mo n
the LHS automatically vanishes at the lower limit. However, in order that it alsovanishes at the upper limit, we require in addition that
∂F
∂y/primevextendsinglevextendsinglevextendsinglevextendsingle
x=b=0. (22.16)
Clearly if both end-points may vary then ∂F/∂y/primemust vanish at both ends.
An interesting and more general case is where the lower end-point is again
fixed at x=a, but the upper end-point is free to lie anywhere on the curve
h(x, y) = 0. Now in this case, the variation in the value of Idue to the arbitrary
variation (22.2) is given to first order by
δI=bracketleftbigg∂F
∂y/primeηbracketrightbiggb
a+integraldisplayb
aparenleftbigg∂F
∂y−d
dx∂F
∂y/primeparenrightbigg
ηd x+F(b)∆x, (22.17)
where ∆ xis the displacement in the x-direction of the upper end-point, as
indicated in figure 22.5, and F(b) is the value of Fatx=b. In order for (22.17)
to be valid, we of course require the displacement ∆ xto be small.
From the figure we see that ∆ y=η(b)+y/prime(b)∆x. Since the upper end-point
must lie on h(x, y) = 0 we also require that, at x=b,
∂h
∂x∆x+∂h
∂y∆y=0,
which on substituting our expression for ∆ yand rearranging becomes
parenleftbigg∂h
∂x+y/prime∂h
∂yparenrightbigg
∆x+∂h
∂yη=0. (22.18)
842
22.3 SOME EXTENSIONS
A
yx=x0
Bx
Figure 22.6 A frictionless wire along which a small bead slides. We seek the
shape of the wire that allows the bead to travel from the origin Oto the line
x=x0in the least possible time.
Now, from (22.17) the condition δI= 0 requires, besides the EL equation, that
atx=b, the other two contributions cancel, i.e.
F∆x+∂F
∂y/primeη=0. (22.19)
Eliminating ∆ xandηbetween (22.18) and (22.19) leads to the condition that at
the end-point
parenleftbigg
F−y/prime∂F
∂y/primeparenrightbigg∂h
∂y−∂F
∂y/prime∂h
∂x=0. (22.20)
In the special case where the end-point is free to lie anywhere on the vertical line
x=b, we have ∂h/∂x =1a n d ∂h/∂y = 0. Substituting these values into (22.20),
we recover the end-point condition given in (22.16).IA frictionless wire in a vertical plane connects two points AandB,Abeing higher than B.
Let the position of Abe fixed at the origin of an xy-coordinate system, but allow Bto lie
anywhere on the vertical line x=x0(see figure 22.6). Find the shape of the wire such that
ab e a dp l a c e do ni ta t Awill slide under gravity to Bin the shortest possible time.
This is a variant of the famous brachistochrone (shortest time) problem, which is often
used to illustrate the calculus of variations. Conservation of energy tells us that the particle
speed is given by
v=ds
dt=
p
2gy,
where sis the path length along the wire and gis the acceleration due to gravity. Since
the element of path length is ds=( 1+ y/prime2)1/2dx, the total time taken to travel to the line
x=x0is given by
t=
Zx=x0
x=0ds
v=1√2g
Zx0
0
s
1+y/prime2
ydx.
Because the integrand does not contain xexplicitly, we can use (22.8) with the specific
form F=
p
1+y/prime2/√yto find a first integral; on simplification this yieldsh
y(1 +y/prime2)
i1/2
=k,
843
CALCULUS OF VARIATIONS
where kis a constant. Letting a=k2and solving for y/primewe find
y/prime=dy
dx=
ra−y
y,
which on substituting y=asin2θintegrates to give
x=a
2(2θ−sin2θ)+c.
Thus the parametric equations of the curve are given by
x=b(φ−sinφ)+c, y =b(1−cosφ),
where b=a/2a n d φ=2θ; they define a cycloid, the curve traced out by a point on
the rim of a wheel of radius brolling along the x-axis. We must now use the end-point
conditions to determine the constants bandc. Since the curve passes through the origin,
we see immediately that c=0 .N o ws i n c e y(x0) is arbitrary, i.e. the upper end-point can
lie anywhere on the curve x=x0, the condition (22.20) reduces to (22.16), so that we also
require
∂F
∂y/prime
////
x=x0=y/primep
y(1 +y/prime2)
/////
x=x0=0,
which implies that y/prime=0a t x=x0. In words, that the tangent to the cycloid at Bmust
b ep a r a l l e lt ot h e x-axis; this requires πb=x0.
J
22.4 Constrained variation
Just as the problem of finding thestationary values of a function f(x, y)s u b j e c tt o
the constraint g(x, y) = constant is solved by means of Lagrange’s undetermined
multipliers (see chapter 5), so the corresponding problem in the calculus ofvariations is solved by an analogous method.
Suppose that we wish to find the stationary values of
I=integraldisplay
b
aF(y,y/prime,x)dx,
subject to the constraint that the value of
J=integraldisplayb
aG(y,y/prime,x)dx
is held constant. Following the method of Lagrange undetermined multipliers let
us define a new functional
K=I+λJ=integraldisplayb
a(F+λG)dx,
and find its unconstrained stationary values. Repeating the analysis of section 22.1
we find that we require
∂F
∂y−d
dxparenleftbigg∂F
∂y/primeparenrightbigg
+λbracketleftbigg∂G
∂y−d
dxparenleftbigg∂G
∂y/primeparenrightbiggbracketrightbigg
=0,
844
22.4 CONSTRAINED VARIATION
−ay
O a
x
Figure 22.7 A uniform rope with fixed end-points suspended under gravity.
which, together with the original constraint J= constant, will yield the required
solution y(x).
This method is easily generalised to cases with more than one constraint by the
introduction of more Lagrange multipliers. If we wish to find the stationary valuesof an integral Isubject to the multiple constraints that the values of the integrals
J
ibe held constant for i=1,2,...,n, then we simply find the unconstrained
stationary values of the new integral
K=I+nsummationdisplay
1λiJi.IFind the shape assumed by a uniform rope when suspended by its ends from two points
at equal heights.
We will solve this problem using x(see figure 22.7) as the independent variable. Let
the rope of length 2 Lbe suspended between the points x=±a,y=0( L>a )a n d
have uniform linear density ρ. We then need to find the stationary value of the rope’s
gravitational potential energy,
I=−ρg
Z
yd s=−ρg
Za
−ay(1 +y/prime2)1/2dx,
with respect to small changes in the form of the rope but subject to the constraint that
the total length of the rope remains constant, i.e.
J=
Z
ds=
Za
−a(1 +y/prime2)1/2dx=2L.
We thus define a new integral (omitting the factor −1f r o m Ifor brevity)
K=I+λJ=
Za
−a(ρgy+λ)(1 + y/prime2)1/2dx
and find its stationary values. Since the integrand does not contain the independent
variable xexplicitly, we can use (22.8) to find the first integral:
(ρgy+λ)
/
1+y/prime2
/1/2
−(ρgy+λ)
/
1+y/prime2
/−1/2
y/prime2=k,
845
CALCULUS OF VARIATIONS
where kis a constant; this reduces to
y/prime2=
/ρgy+λ
k
/2
−1.
Making the substitution ρgy+λ=kcoshz, this can be integrated easily to give
k
ρgcosh−1
/ρgy+λ
k
/
=x+c,
where ci st h ec o n s t a n to fi n t e g r a t i o n .
We now have three unknowns, λ,kandc, that must be evaluated using the two end
conditions y(±a) = 0 and the constraint J=2L. The end conditions give
coshρg(a+c)
k=λ
k=c o s hρg(−a+c)
k,
and since a/negationslash= 0, these imply c=0a n d λ/k=c o s h ( ρga/k ). Putting c=0i n t ot h e
constraint, in which y/prime= sinh( ρgx/k ), we obtain
2L=
Za
−a
h
1+s i n h2
/ρgx
k
/ i1/2
dx
=2k
ρgsinh
/ρga
k
/
.
Collecting together the values for the constants, the form adopted by the rope is therefore
y(x)=k
ρg
h
cosh
/ρgx
k
/
−cosh
/ρga
k
/ i
,
where kis the solution of sinh( ρga/k )=ρgL/k . This curve is known as a catenary.
J
22.5 Physical variational principles
Many results in both classical and quantum physics can be expressed as varia-
tional principles, and it is often when expressed in this form that their physical
meaning is most clearly understood. Moreover, once a physical phenomenon hasbeen written as a variational principle, we can use all the results derived in thischapter to investigate its behaviour. It is usually possible to identify conservedquantities, or symmetries of the system of interest, that otherwise might be foundonly with considerable effort. From the wide range of physical variational princi-ples we will select two examples from familiar areas of classical physics, namely
geometric optics and mechanics.
22.5.1 Fermat’s principle in optics
Fermat’s principle in geometrical optics states that a ray of light travelling in a
region of variable refractive index follows a path such that the total optical pathlength (physical length ×refractive index) is stationary.
846
22.5 PHYSICAL VARIATIONAL PRINCIPLES
θ1θ2
n1n2
AB
xy
Figure 22.8 Path of a light ray at the plane interface between media with
refractive indices n1andn2,w h e r e n2<n1.IFrom Fermat’s principle deduce Snell’s law of refraction at an interface.
Let the interface be at y= constant (see figure 22.8) and let it separate two regions with
refractive indices n1andn2respectively. On a ray the element of physical path length is
ds=( 1+ y/prime2)1/2dx, and so for a ray that passes through the points AandB,t h et o t a l
optical path length is
P=
ZB
An(y)(1 + y/prime2)1/2dx.
Since the integrand does not contain the independent variable xexplicitly, we use (22.8)
to obtain a first integral, which, a fter some rearrangement, reads
n(y)
/
1+y/prime2
/−1/2
=k,
where kis a constant. Recalling that y/primeis the tangent of the angle φbetween the
instantaneous direction of the ray and the x-axis, this general result, which is not dependent
on the configuration presently under consideration, can be put in the form
ncosφ=c o n s t a n t
along a ray, even though nandφvary individually.
For our particular configuration nis constant in each medium and therefore so is
y/prime. Thus the rays travel in straight lines in each medium (as anticipated in figure 22.8,
but not assumed in our analysis), and since kis constant along the whole path we have
n1cosφ1=n2cosφ2, or in terms of the conventional angles in the figure
n1sinθ1=n2sinθ2.
J
22.5.2 Hamilton’s principle in mechanics
Consider a mechanical system whose configuration can be uniquely defined by a
number of coordinates qi(usually distances and angles) together with time tand
which experiences only forces derivable from a potential. Hamilton’s principle
847
CALCULUS OF VARIATIONS
y
Odx lx
Figure 22.9 Transverse displacement on a taut string that is fixed at two
points a distance lapart.
states that in moving from one configuration at time t0to another at time t1the
motion of such a system is such as to make
L=integraldisplayt1
t0L(q1,q2...,q n,˙q1,˙q2,...,˙qn,t)dt (22.21)
stationary. The Lagrangian Lis defined, in terms of the kinetic energy Tand
the potential energy V(with respect to some reference situation), by L=T−V.
Here Vis a function of the qionly, not of the ˙qi. Applying the EL equation to L
we obtain Lagrange’s equations ,
∂L
∂qi=d
dtparenleftbigg∂L
∂˙qiparenrightbigg
,i =1,2,...,n .IUsing Hamilton’s principle derive the wave e quation for small transverse oscillations of a
taut string.
In this example we are in fact considering a generalisation of (22.21) to a case involvingone isolated independent coordinate t,t o g e t h e rw i t ha continuum in which the q
ibecome
the continuous variable x. The expressions for TandVtherefore become integrals over x
rather than sums over the label i.
Ifρandτare the local density and tension of the string, both of which may depend on
x, then, referring to figure 22.9, the kinetic and potential energies of the string are given
by
T=
Zl
0ρ
2
/∂y
∂t
/2
dx, V =
Zl
0τ
2
/∂y
∂x
/2
dx
and (22.21) becomes
L=1
2
Zt1
t0dt
Zl
0
/"
ρ
/∂y
∂t
/2
−τ
/∂y
∂x
/2
/#
dx.
848
22.6 GENERAL EIGENVALUE PROBLEMS
Using (22.13) and the fact that ydoes not appear explicitly, we obtain
∂
∂t
/
ρ∂y
∂t
/
−∂
∂x
/
τ∂y
∂x
/
=0.
If, in addition, ρandτdo not depend on xortthen
∂2y
∂x2=1
c2∂2y
∂t2,
where c2=τ/ρ. This is the wave equation for small transverse oscillations of a taut
uniform string.
J
22.6 General eigenvalue problems
We have seen in this chapter that the problem of finding a curve that makes the
value of a given integral stationary when the integral is taken along the curveresults, in each case, in a differential equation for the curve. It is not a great
extension to ask whether this may be used to solve differential equations, by
setting up a suitable variational problem and then seeking ways other than theEuler equation of finding or estimating stationary solutions.
We shall be concerned with differential equations of the form Ly=λρ(x)y,
where the differential operator Lis self-adjoint, so that L=L
†(with appropriate
boundary conditions on the solution y)a n d ρ(x) is some weight function, as
discussed in chapter 17. In particular, we will concentrate on the Sturm–Liouvilleequation as an explicit example, but much of what follows can be applied toother equations of this type.
We have already discussed the solution of equations of the Sturm–Liouville
type in chapter 17 and the same notation will be used here. In this section,however, we will adopt a variational approach to estimating the eigenvalues ofsuch equations.
Suppose we search for stationary values of the integral
I=integraldisplay
b
abracketleftBig
p(x)y/prime2(x)−q(x)y2(x)bracketrightBig
dx, (22.22)
with y(a)=y(b)=0a n d pandqany sufficiently smooth and differentiable
functions of x. However, in addition we impose a normalisation condition
J=integraldisplayb
aρ(x)y2(x)dx=c o n s t a n t . (22.23)
Here ρ(x) is a positive weight function defined in the interval a≤x≤b, but
which may in particular cases be a constant.
Then, as in section 22.4, we use undetermined Lagrange multipliers, †and
†We use−λ, rather than λ, so that the final equation (22.24) appears in the conventional Sturm–
Liouville form.
849
CALCULUS OF VARIATIONS
consider K=I−λJgiven by
K=integraldisplayb
abracketleftBig
py/prime2−(q+λρ)y2bracketrightBig
dx.
On application of the EL equation (22.5) this yields
d
dxparenleftbigg
pdy
dxparenrightbigg
+qy+λρy=0, (22.24)
which is exactly the Sturm–Liouville equation (17.35), with eigenvalue λ. Now,
since both IandJare quadratic in yand its derivative, finding stationary values
ofKis equivalent to finding stationary values of I/J. This may also be shown
by considering the functional Λ = I/J, for which
δΛ=( δI/J)−(I/J2)δJ
=(δI−ΛδJ)/J
=δK/J.
Hence, extremising Λ is equivalent to extremising K. Thus we have the important
result that finding functions ythat make I/Jstationary is equivalent to finding
functions ythat are solutions of the Sturm–Liouville equation; the resulting value
ofI/Jequals the corresponding eigenvalue of the equation.
Of course this does not tell us how to find such a function yand, naturally, to
have to do this by solving (22.24) directly defeats the purpose of the exercise. Wewill see in the next section how some progress can be made. It is worth recallingthat the functions p(x),q(x)a n d ρ(x) can have many different forms, and so
(22.24) represents quite a wide variety of equations.
We now recall some properties of the solutions of the Sturm–Liouville equation.
The eigenvalues λ
iof (22.24) are real and will be assumed non-degenerate (for
simplicity). We also assume that the corresponding eigenfunctions have been made
real, so that normalised eigenfunctions yi(x) satisfy the orthogonality relation (as
in (17.27))
integraldisplayb
ayiyjρd x=δij. (22.25)
Further, we take the boundary condition in the form
bracketleftBig
yipy/prime
jbracketrightBigx=b
x=a= 0; (22.26)
this can be satisfied by y(a)=y(b) = 0, but also by many other sets of boundary
conditions.
850
22.7 ESTIMATION OF EIGENVALUES AND EIGENFUNCTIONSIShow thatZb
a
/;
y/prime
jpy/prime
i−yjqyi
/
dx=λiδij. (22.27)
Letyibe an eigenfunction of (22.24), corresponding to a particular eigenvalue λi,s ot h a t/;
py/prime
i
//prime+(q+λiρ)yi=0.
Multiplying this through by yjand integrating from atob(the first term by parts) we
obtain/
yj
/;
py/prime
i
/
/b
a−
Zb
ay/prime
j(py/prime
i)dx+
Zb
ayj(q+λiρ)yidx=0. (22.28)
The first term vanishes by virtue of (22.26), and on rearranging the other terms and using
(22.25), we find the result (22.27).
J
We see at once that, if the function y(x) minimises I/J, i.e. satisfies the Sturm–
Liouville equation, then putting yi=yj=yin (22.25) and (22.27) yields Jand
Irespectively on the left-hand sides; thus, as mentioned above, the minimised
value of I/Jis just the eigenvalue λ, introduced originally as the undetermined
multiplier.IFor a function ysatisfying the Sturm–Liouville equat ion verify that, provided (22.26) is
satisfied, λ=I/J.
Firstly, we multiply (22.24) through by yto give
y(py/prime)/prime+qy2+λρy2=0.
Now integrating this expression by parts we have/
ypy/prime
/b
a−
Zb
a
/
py/prime2−qy2
/
dx+λ
Zb
aρy2dx=0.
The first term on the LHS is zero, the second is simply −Ia n dt h et h i r di s λJ. Thus
λ=I/J.
J
22.7 Estimation of eigenvalues and eigenfunctions
Since the eigenvalues λiof the Sturm–Liouville equation are the stationary values
ofI/J(see above), it follows that any evaluation of I/Jmust yield a value that lies
between the lowest and highest eigenvalues of the corresponding Sturm–Liouville
equation, i.e.
λmin≤I
J≤λmax,
where, depending on the equation under consideration, either λmin=−∞and
851
CALCULUS OF VARIATIONS
λmaxis finite, or λmax=∞andλminis finite. Notice that here we have departed
from direct consideration of the minimising problem and made a statement abouta calculation in which no actual minimisation is necessary.
Thus, as an example, for an equation with a finite lowest eigenvalue λ
0any
evaluation of I/Jprovides an upper bound on λ0. Further, we will now show that
the estimate λobtained is a better estimate of λ0than the estimated (guessed)
function yis of y0, the true eigenfunction corresponding to λ0. The sense in which
‘better’ is used here will be clear from the final result.
Firstly, we expand the estimated or trial function yin terms of the complete
setyi:
y=y0+c1y1+c2y2+···,
where, if a good trial function has been guessed, the ciwill be small. Using (22.25)
we have immediately that J=1+summationtext
i|ci|2. The other required integral is
I=integraldisplayb
abracketleftBigg
pparenleftbigg
y/prime
0+summationdisplay
iciy/prime
iparenrightbigg2
−qparenleftbigg
y0+summationdisplay
iciyiparenrightbigg2bracketrightBigg
dx.
On multiplying out the squared terms, all the cross terms vanish because of
(22.27) to leave
λ=I
J=λ0+summationtext
i|ci|2λi
1+summationtext
j|cj|2
=λ0+summationdisplay
i|ci|2(λi−λ0)+O ( c4).
Hence λdiffers from λ0by a term second order in the ci, even though ydiffered
from y0by a term first order in the ci; this is what we aimed to show. We notice
incidentally that, since λ0<λ ifor all i,λis shown to be necessarily ≥λ0, with
equality only if all ci=0 ,i . e .i f y≡y0.
The method can be extended to the second and higher eigenvalues by imposing,
in addition to the original constraints and boundary conditions, a restrictionof the trial functions to only those that are orthogonal to the eigenfunctionscorresponding to lower eigenvalues. (Of course, this requires complete or nearlycomplete knowledge of these latter eigenfunctions.) An example is given at theend of the chapter (exercise 22.26).
We now illustrate the method we have discussed by considering a simple
example, one for which, as on previous occasions, the answer is obvious.
852
22.7 ESTIMATION OF EIGENVALUES AND EIGENFUNCTIONS
(a)(b)(c)
(d)
0.20.2
0.40.4
0.60.6
0.80.8
11
xy(x)
Figure 22.10 Trial solutions used to estimate the lowest eigenvalue λof
−y/prime/prime=λywith y(0) = y/prime(1) = 0. They are: ( a)y=s i n ( πx/2), the exact result;
(b)y=2x−x2;(c)y=x3−3x2+3x;(d)y=s i n2(πx/2).IEstimate the lowest eigenvalue of the equation
−d2y
dx2=λy, 0≤x≤1, (22.29)
with boundary conditions
y(0) = 0 ,y/prime(1) = 0 . (22.30)
We need to find the lowest value λ0ofλfor which (22.29) has a solution y(x) that satisfies
(22.30). The exact answer is of course y=Asin(xπ/2) and λ0=π2/4≈2.47.
Firstly we note that the Sturm–Liouville equation reduces to (22.29) if we take p(x)=1 ,
q(x)=0a n d ρ(x) = 1 and that the boundary conditions satisfy (22.26). Thus we are able
to apply the previous theory.
We will use three trial functions so that the effect on the estimate of λ0of making better
or worse ‘guesses’ can be seen. One further preliminary remark is relevant, namely that theestimate is independent of any constant multiplicative factor in the function used. Thisis easily verified by looking at the form of I/J. We normalise each trial function so that
y(1) = 1, purely in order to facilitate comparison of the various function shapes.
Figure 22.10 illustrates the trial functions used, curve ( a) being the exact solution
y= sin( πx/2). The other curves are ( b)y(x)=2 x−x
2,(c)y(x)=x3−3x2+3x,a n d( d)
y(x)=s i n2(πx/2). The choice of trial function is governed by the following considerations:
(i) the boundary conditions (22.30) mustbe satisfied.
(ii) a ‘good’ trial function ought to mimic the correct solution as far as possible, but
it may not be easy to guess even the general shape of the correct solution in somecases.
(iii) the evaluation of I/Jshould be as simple as possible.
853
CALCULUS OF VARIATIONS
It is easily verified that functions ( b), (c)a n d( d) all satisfy (22.30) but, so far as mimicking
the correct solution is concerned, we would expect from the figure that ( b)w o u l db e
superior to the other two. All three evaluations are straightforward, using (22.22) and(22.23):
λ
b=
R1
0(2−2x)2dxR1
0(2x−x2)2dx=4/3
8/15=2.50
λc=
R1
0(3x2−6x+3 )2dxR1
0(x3−3x2+3x)2dx=9/5
9/14=2.80
λd=
R1
0(π2/4)sin2(πx)dxR1
0sin4(πx/2)dx=π2/8
3/8=3.29.
We expected all evaluations to yield estimates greater than the lowest eigenvalue, 2.47,
and this is indeed so. From these trials alone we are able to say (only) that λ0≤2.50.
As expected, the best approximation ( b) to the true eigenfunction yields the lowest, and
therefore the best, upper bound on λ0.
J
We may generalise the work of this section to other differential equations of
the form Ly=λρy,w h e r e L=L†. In particular, one finds
λmin≤I
J≤λmax,
where IandJare now given by
I=integraldisplayb
ay∗(Ly)dx and J=integraldisplayb
aρy∗yd x . (22.31)
It is straightforward to show that, for the special case of the Sturm–Liouville
equation, for which
Ly=−(py/prime)/prime−qy,
the expression for Iin (22.31) leads to (22.22).
22.8 Adjustment of parameters
Instead of trying to estimate λ0by selecting a large number of different trial
functions, we may also use trial functions that include one or more parameters
which themselves may be adjusted to give the lowest value to λ=I/Jand
hence the best estimate of λ0. The justification for this method comes from the
knowledge that no matter what form of function is chosen, nor what values areassigned to the parameters, provided the boundary conditions are satisfied λcan
never be less than the required λ
0.
To illustrate this method an example from quantum mechanics will be used.
The time-independent Schr ¨odinger equation is formally written as the eigenvalue
equation Hψ=Eψ,w h e r e His a linear operator, ψthe wavefunction describing
a quantum mechanical system and Ethe energy of the system. The energy
854
22.8 ADJUSTMENT OF PARAMETERS
operator His called the Hamiltonian and for a particle of mass mmoving in a
one-dimensional harmonic oscillator potential is given by
H=−/planckover2pi12
2md2
dx2+kx2
2, (22.32)
where/planckover2pi1is Planck’s constant divided by 2 π.IEstimate the ground-state energy of a quantum harmonic oscillator.
Using (22.32) in Hψ=Eψ, the Schr ¨odinger equation is
−
/~2
2md2ψ
dx2+kx2
2ψ=Eψ,−∞<x<∞. (22.33)
The boundary conditions are that ψshould vanish as x→±∞ . Equation (22.33) is a form
of the Sturm–Liouville equation in which p= /~2/(2m),q=−kx2/2,ρ=1a n d λ=E;i t
can be solved by the methods developed previously, e.g. by writing the eigenfunction ψas
a power series in x.
However, our purpose here is to illustrate variational methods and so we take as a trial
wavefunction ψ=e x p (−αx2), where αis a positive parameter whose value we will choose
later. This function certainly →0a sx→±∞ and is convenient for calculations. Whether
it approximates the true wave function is unknown, but if it does not our estimate willstill be valid, although the upper bound will be a poor one.
With y=e x p (−αx
2) and therefore y/prime=−2αxexp(−αx2), the required estimate is
E=λ=
R∞
−∞[( /~2/2m)4α2x2+(k/2)x2]e−2αx2dxR∞
−∞e−2αx2dx=
/~2α
2m+k
8α. (22.34)
This evaluation is easily carried out using the reduction formula
In=n−1
4αIn−2,for integrals of the form In=
Z∞
−∞xne−2αx2dx.
(22.35)
So, we have obtained the estimate (22.34), involving the parameter α, for the oscillator’s
ground-state energy, i.e. the lowest eigenvalue of H. In line with our previous discussion
we now minimise λwith respect to α. Putting dλ/dα = 0 (clearly a minimum), yields
α=(km)1/2/(2 /~), which in turn gives as the minimum value for λ
E=
/~
2
/k
m
/1/2
=
/~ω
2, (22.36)
where we have put ( k/m)1/2equal to the classical angular frequency ω.
The method thus leads to the conclusion that the ground-state energy E0is≤1
2
/~ω.
In fact, as is well known, the equality sign holds,1
2
/~ωbeing just the zero-point energy
of a quantum mechanical oscillator. Our estimate gives the exact value because ψ(x)=
exp(−αx2) is the correct functional form for the ground state wavefunction and the
particular value of αthat we have found is the that needed to make ψan eigenfunction of
Hwith eigenvalue ≤1
2
/~ω.
J
An alternative but equivalent approach to this is developed in the exercises
that follow, as is an extension of this particular problem to estimating the second-lowest eigenvalue (see exercise 22.26).
855
CALCULUS OF VARIATIONS
22.9 Exercises
22.1 A surface of revolution, whose equation in cylindrical polar coordinates is ρ=
ρ(z), is bounded by the circles ρ=a,z=±c(a>c). Show that the function
that makes the surface integral I=
R
ρ−1/2dSstationary with respect to small
variations is given by ρ(z)=k+z2/(4k), where k=[a±(a2−c2)1/2]/2.
22.2 Show that the lowest value of the integralZB
A(1 +y/prime2)1/2
ydx,
where Ais (−1,1) and Bis (1,1), is 2ln(1+√
2). Assume that the Euler–Lagrange
equation gives a minimising curve.
22.3 The refractive index nof a medium is a function only of the distance rfrom a
fixed point O. Prove that the equation of a light ray, assumed to lie in a plane
through O, travelling in the medium satisfies (in plane polar coordinates)
1
r2
/dr
dφ
/2
=r2
a2n2(r)
n2(a)−1,
where ais the distance of the ray from Oat the point at which dr/dφ =0 .
Ifn=[ 1+( α2/r2)]1/2and the ray starts and ends far from O, find its deviation
(the angle through which the ray is tu rned) if its minimum distance from Oisa.
22.4 The Lagrangian for a π-meson is given by
L(x,t)=1
2(˙φ2−|∇φ|2−µ2φ2),
where µis the meson mass and φ(x,t) is its wavefunction. Assuming Hamilton’s
principle find the wave equation satisfied by φ.
22.5 (a) For a system described in terms of coordinates qiandt, show that if tdoes
not appear explicitly in the expressions for x,yandz(x=x(qi,t), etc.) then
the kinetic energy Tis a homogeneous quadratic function of the ˙qi(it may
also involve the qi). Deduce that
P
i˙qi(∂T/∂˙qi)=2 T.
(b) Assuming that the forces acting on th e system are derivable from a potential
V, show, by expressing dT/dt in terms of qiand˙qi,t h a t d(T+V)/dt=0 .
22.6 For a system specified by the coordinates qandt, show that the equation of
motion is unchanged if the Lagrangian L(q,˙q,t) is replaced by
L1=L+dφ(q,t)
dt,
where φis an arbitrary function. Deduce that the equation of motion of a particle
that moves in one dimension subject to a force −dV(x)/dx(xbeing measured
from a point O) is unchanged if Ois forced to move with a constant velocity v
(xstill being measured from O).
22.7 In cylindrical polar coordinates, the curve ( ρ(θ),θ,α ρ(θ)) lies on the surface of
the cone z=αρ. Show that geodesics (curves of minimum length joining two
points) on the cone satisfy
ρ4=c2[β2ρ/prime2+ρ2],
where cis an arbitrary constant, but βhas to have a particular value. Determine
the form of ρ(θ) and hence find the equation of the shortest path on the
cone between the points ( R,−θ0,α R)a n d( R,θ0,α R). (You will find it useful to
determine the form of the derivative of cos−1(u−1).)
22.8 Derive the differential equations for the polar coordinates r,θof a particle of
unit mass moving in a field of potential V(r). Find the form of Vif the path of
the particle is given by r=asinθ.
856
22.9 EXERCISES
22.9 You are provided with a line of length πa/2 and negligible mass and some lead
shot of total mass M. Use a variational method to determine how the lead shot
must be distributed along the line if the loaded line is to hang in a circular arc ofradius awhen its ends are attached to two points at the same height. (Measure
the distance salong the line from its centre.)
22.10 Extend the result of subsection 22.2.2 to the case of several dependent variables
y
i(x), showing that, if xdoes not appear explicitly in the integrand, then a first
integral of the Euler–Lagrange equations is
F−nX
i=1y/prime
i∂F
∂y/prime
i=c o n s t a n t .
22.11 A general result is that light travels through a variable medium by a path which
minimises the travel time (this is an alternative formulation of Fermat’s principle).With respect to a particular cylindrical polar coordinate system ( ρ, φ, z) the speed
of light v(ρ, φ) is independent of z. If the path of the light is parameterised as
ρ=ρ(z),φ=φ(z), use the result of the previous exercise to show that
v
2(ρ/prime2+ρ2φ/prime2+1 )
is constant along the path.
For the particular case when v=v(ρ)=b(a2+ρ2)1/2, show that the two Euler–
Lagrange equations have a common solution in which the light travels along ahelical path given by φ=Az+B,ρ=C, provided that Ahas a particular value.
22.12 Light travels in a vertical xz-plane through a slab of material which lies between
the planes z=z
0andz=2z0and in which the speed of light v(z)=c0z/z0.U s i n g
the alternative formulation of Fermat’s principle given in the previous question,show that the ray paths are arcs of circles.
Deduce that, if a ray enters the material at (0 ,z
0) at an angle to the vertical,
π/2−θ,o fm o r et h a n3 0◦, it does not reach the far side of the slab.
22.13 A dam of capacity V(less than πb2h/2) is to be constructed on level ground next
to a long straight wall which runs from ( −b,0) to ( b,0 ) .T h i si st ob ea c h i e v e db y
joining the ends of a new wall, of height h, to those of the existing wall. Show
that, in order to minimise the length Lof new wall to be built, it should form
part of a circle, and that Lis then given byZb
−bdx
(1−λ2x2)1/2,
where λis found from
V
hb2=sin−1µ
µ2−(1−µ2)1/2
µ,
andµ=λb.
22.14 The Schwarzchild metric for the static field of a non-rotating spherically sym-
metric black hole of mass Mis
(ds)2=c2
/
1−2GM
c2r
/
(dt)2−dr2
1−2GM/(c2r)−r2(dθ)2−r2sin2θ(dφ)2.
Considering only motion confined to the plane θ=π/2, and assuming that the
path of a small test particle is such as to make
R
dsstationary, find two first
integrals of the equations of motion. From their Newtonian limits, in whichGM/r ,˙r
2andr2˙φ2are all/lessmuchc2, identify the constants of integration.
22.15 In the brachistochrone problem of subsection 22.3.4 show that if the upper end-
point can lie anywhere on the curve h(x, y) = 0 then the curve of quickest descent
y(x) meets h(x, y) = 0 at right angles.
857
CALCULUS OF VARIATIONS
22.16 Use result (22.27) to evaluate
J=
Z1
−1(1−x2)P/prime
m(x)P/prime
n(x)dx,
where Pm(x) is a Legendre polynomial of order m.
22.17 Determine the minimum value that the integral
J=
Z1
0[x4(y/prime/prime)2+4x2(y/prime)2]dx,
can have, given that yis not singular at x=0a n dt h a t y(1) = y/prime(1) = 1.
Assume that the Euler–Lagrange equation does give the lower limit, and verifyretrospectively that your solution makes the first term on the LHS of equation(22.15) vanish.
22.18 Show that y
/prime/prime−xy+λx2y=0h a sas o l u t i o nf o rw h i c h y(0) = y(1) = 0 and
λ≤147/4.
22.19 Find an appropriate but simple trial function and use it to estimate the lowest
eigenvalue λ0of Stokes’ equation
d2y
dx2+λxy=0,y (0) = y(π)=0 .
Explain why your estimate must be strictly greater than λ0.
22.20 Estimate the lowest eigenvalue λ0of the equation
d2y
dx2−x2y+λy=0,y (−1) = y(1) = 0 ,
using a quadratic trial function.
22.21 A drumskin is stretched across a fixed circular rim of radius a. Small transverse
vibrations of the skin have an amplitude z(ρ, φ, t) that satisfies
∇2z=1
c2∂2z
∂t2
in plane polar coordinates. For a normal mode independent of azimuth, z=
Z(ρ)cosωt, find the differential equation satisfied by Z(ρ). By using a trial
function of the form aν−ρν, obtain an estimate for the lowest normal mode
frequency. (The exact answer is (5 .78)1/2c/a.)
22.22 (a) Recast the problem of finding the lowest eigenvalue λ0of the equation
(1 +x2)d2y
dx2+2xdy
dx+λy=0,y (±1) = 0 ,
in variational form, and derive an approximation λ1toλ0by using the trial
function y1(x)=1−x2.
(b) Show that an improved estimate λ2is obtained by using y2(x)=c o s ( πx/2).
(c) Prove that the estimate λ(γ) obtained by taking y1(x)+γy2(x)a st h et r i a l
function is
λ(γ)=64/15 + 16 γ/π+(π2/3+1 /2)γ2
16/15 + 64 γ/π3+γ2.
Investigate λ(γ) numerically as γis varied, or, more simply, show that
λ(1) = 3 .183, a significant improvement on both λ1andλ2.
22.23 For the boundary conditions given below, obtain a functional Λ( y) whose sta-
tionary values give the eigenvalues of the equation
(1 +x)d2y
dx2+( 2+ x)dy
dx+λy=0,y (0) = 0 ,y/prime(2) = 0 .
858
22.9 EXERCISES
Derive an approximation to the lowest eigenvalue λ0using the trial function
y(x)=xe−x/2. For what value(s) of γwould
y(x)=xe−x/2+βsinγx
be a suitable trial function for attemp ting to obtain an improved estimate of λ0?
22.24 The upper and lower surfaces of a film of liquid with surface energy per unit
area (surface tension) equal to γand with density ρhave equations z=p(x)a n d
z=q(x) respectively. The film has a given volume V(per unit depth in the y-
direction) and lies in the region −L<x<L ,w i t h p(0) = q(0) = p(L)=q(L)=0 .
The total energy (per unit depth) of the film consists of its surface energy and its
gravitational energy, and is expressed by
E=1
2
ZL
−L(p2−q2)dx+γ
ZL
−L
h
(1 +p/prime2)1/2+( 1+ q/prime2)1/2
i
dx.
(a) Express Vin terms of pandq.
(b) Show that, if the total energy is minimised, pandqmust satisfy
p/prime2
(1 +p/prime2)1/2−q/prime2
(1 +q/prime2)1/2=c o n s t a n t .
(c) As an approximate solution, consider the equations
p=a(L−|x|),q =b(L−|x|),
where aandbare sufficiently small that a3andb3can be neglected compared
to unity. Find the values of aandbthat minimise E.
22.25 This is an alternative approach to the example in section 22.8. Using the notation
of that section, the expectation value of the energy of the state ψis given byR
ψ∗Hψdv . Denote the eigenfunctions of Hbyψi,s ot h a t Hψ i=Eiψi, and, since
His self-adjoint (Hermitian),
R
ψ∗
jψidv=δij.
(a) By writing any function ψas
Pcjψjand following an argument similar to
that in section 22.7, show that
E=
R
ψ∗HψdvR
ψ∗ψd v≥E0,
the energy of the lowest state. (This is the Rayleigh–Ritz principle.)
(b) Using the same trial function as in section 22.8, ψ=e x p (−αx2),show that
the same result is obtained.
22.26 This is an extension to section 22.8 and the previous question. With the ground-
state (i.e. the lowest-energy) wavefunction as exp( −αx2), take as a trial function
the orthogonal wave function x2n+1exp(−αx2), using the integer nas a variable
parameter. Use either Sturm–Liouville theory or the Rayleigh–Ritz principle toshow that the energy of the second lowest state of a quantum harmonic oscillatoris≤3/~ω/2.
22.27 The Hamiltonian Hfor the hydrogen atom is
−
/~2
2m∇2−q
4π/epsilon10r.
For a spherically symmetric state, as may be assumed for the ground state, the
only relevant part of ∇2is that involving differentiation with respect to r.
(a) Define the integrals Jnby
Jn=
Z∞
0rne−2βrdr
859
CALCULUS OF VARIATIONS
and show that, for a trial wavefunction of the form exp( −βr)w i t h β>0,R
ψ∗Hψdv and
R
ψ∗ψd v(see exercise 22.25(a)) can be expressed as aJ1−bJ2
andcJ2respectively, where a, b, c are factors which you should determine.
(b) Show that the estimate of Eis minimised when β=mq2/(4π/epsilon10
/~2).
(c) Hence find an upper limit for the ground-state energy of the hydrogen atom.
In fact, exp( −βr) is the correct form for the wavefunction and the limit gives
the actual value.
22.28 A particle of mass mmoves in a one-dimensional potential well of the form
V(x)=−µ
/~2α2
msech2αx,
where µandαare positive constants. As in exercise 22.27, the expectation value
/angbracketleftE/angbracketrightof the energy of the system is
R
ψ∗Hψdx , where the self-adjoint operator
His given by −( /~2/2m)d2/dx2+V(x). Using trial wavefunctions of the form
y=Asech βx, show the following:
(a) for µ= 1 there is an exact eigenfunction of H, with a corresponding /angbracketleftE/angbracketrightof
half of the maximum depth of the well;
(b) for µ= 6 the ‘binding energy’ of the ground state is at least 10 /~2α2/(3m).
(You will find it useful to note that for u,v≥0, sech usech v≥sech ( u+v).)
22.29 The Sturm–Liouville equation can be extended to two independent variables, x
andz, with little modification. In equation (22.22) y/prime2is replaced by ( ∇y)2and
the integrals of the various functions of y(x, z) become two-dimensional, i.e. the
infinitesimal is dx dz.
The vibrations of a trampoline 4 units long and 1 unit wide satisfy the equation
∇2y+k2y=0.
By taking the simplest possible permissible polynomial as a trial function, show
that the lowest mode of vibration has k2≤10.63 and, by direct solution, that the
actual value is 10.49.
22.10 Hints and answers
22.2 The minimising curve is x2+y2=2 .
22.3 I=
R
n(r)[r2+(dr/dφ )2]1/2dφ. Take axes such that φ= 0 when r=∞.
Ifβ=(π−deviation angle) /2t h e n β=φatr=a, and the equation reduces to
β
(a2+α2)1/2=
Z∞
−∞dr
r(r2−a2)1/2,
which can be evaluated by putting r=a(y+y−1)/2, or successively r=acoshψ,
y=e x p ψto yield a deviation π[(a2+α2)1/2−a]/a.
22.4∇2φ−∂2φ/∂t2=µ2φ.
22.5 (a) ∂x/∂t =0a n ds o ˙x=
P
i˙qi∂x/∂q i; (b) useX
i˙qid
dt
/∂T
∂˙qi
/
=d
dt(2T)−
X
i¨qi∂T
∂˙qi.
22.6 φ(x, t)=m(vx+v2t/2).
22.7 Use result (22.8); β2=1+ α2.P u t ρ=ucto obtain dθ/du =β/[u(u2−1)1/2].
Remember that cos−1is a multivalued function; ρ(θ)=[Rcos(θ0/β)]/[cos(θ/β)].
22.8 r2˙θ=k,¨r−r˙θ2+dV/dr =0 , V(r)=−k2a2/(2r4) + constant.
22.9−λy/prime(1−y/prime2)−1/2=2gP(s),y=y(s),P(s)=
Rs
0ρ(s/prime)ds/prime.T h es o l u t i o n y=
860
22.10 HINTS AND ANSWERS
−acos(s/a)a n d2 P(πa/4) = Mgive λ=−gM.T h er e q u i r e d ρ(s)i s
[M/(2a)]sec2(s/a).
22.10 Note that dF/dx contains partial contributions from all yi(x)a n da l l y/prime
i(x) but
no∂F/∂x term.
22.11 A=1/a.
22.12 Circle is ( x−z0tanθ)2+z2=z2
0sec2θ. Consider the value of zwhen dz/dx =0 .
22.13 Circle is λ2x2+[λy+( 1−λ2b2)1/2]2=1 .U s et h ef a c tt h a t
R
yd x=V/hto
determine the condition on λ.
22.14 Denoting ( ds)2/(dt)2byf2, the Euler–Lagrange equation for φgives r2˙φ=Af
where Acorresponds to the angular momentum of the particle. Use the result
of exercise 22.10 to obtain c2−(2GM/r )=Bf, where to first order in small
quantities
cB=c2−GM
r+1
2(˙r2+r2˙φ2),
which reads ‘total energy = rest mass + gravitational energy + radial and
azimuthal kinetic energy’.
22.16 Note that Legendre’s equation is a Sturm–Liouville equation with p(±1) = 0
andρ(x) = 1. For normalised eigenfunctions take ym(x)=[ ( 2 m+1 )/2]1/2Pm(x);
J={[2m(m+1 ) ] /(2m+1 )}δmn.
22.17 Convert the equation to the usual form, by writing y/prime(x)=u(x), and obtain
x2u/prime/prime+4xu/prime−4u= 0 with general solution Ax−4+Bx. Integrating a second time
and using the boundary conditions gives y(x)=( 1+ x2)/2a n d J=1 ; η(1) = 0,
since y/prime(1) is fixed, and ∂F/∂u/prime=2x4u/prime=0a t x=0 .
22.18 The equation is of SL form with p=1 , q=−xand weight function x2.T r y
y=x(1−x). The integrals have values 7 /20 and 1 /105.
22.19 Using y=s i n xas a trial function shows that λ0≤2/π. The estimate must be
>λ0since the trial function does not satisfy the original equation.
22.20 Using y=1−x2as a trial function shows that λ0≤37/14.
22.21 Z/prime/prime+ρ−1Z/prime+(ω/c)2Z=0 ,w i t h Z(a)=0a n d Z/prime(0) = 0, an SL equation with
p=ρ,q= 0 and weight function ρ/c2.E s t i m a t eo f ω2=[c2ν/(2a2)][0.5−2(ν+
2)−1+( 2ν+2 )−1]−1, which minimises to c2(2 +√
2)2/(2a2)=5 .83c2/a2when
ν=√
2.
22.22 (a) Follow the method of section 22.7 with p=1+ x2,q=0a n d ρ=1 ; λ1=4.
(b)λ2=π2/3+1 /2≈3.79.
(c)λ(γ) has a minimum value 3 .1768 at γ=1.1976.
22.23 Note that the original equation is not self-adjoint; it needs an integrating factor
ofex.Λ (y)=[
R2
0(1 +x)exy/prime2dx]/[
R2
0exy2dx;λ0≤3/8. Since y/prime(2) must equal 0,
γ=(π/2)(n+1
2)f o rs o m ei n t e g e r n.
22.24 (a) V=
RL
−L(p−q)dx.( c )U s e V=(a−b)L2to eliminate bfrom the expression
forE; now the minimisation is with respect to aalone. The values for aandb
are±V/(2L2)−Vρg/(6γ).
22.25 The estimate is /~2α/(2m)+k/(8α) and the minimum occurs at the value of αthat
makes the two terms equal.
22.26 E1≤( /~ω/2)(8n2+1 2n+3 )/(4n+ 1), which has a minimum value 3 /~ω/2w h e n
integer n=0 .
22.27 (a) a=4π /~2β/m−q2//epsilon10,b=2π /~2β2/m,c=4π;( c )−mq4/[2(4π/epsilon10
/~)2].
22.28 (a) Hy=λyrequires that β=α.
(b)
R
ψ∗[−( /~2/2m)d2/dx2]ψd x= /~2β2/(6m)a n d
R
ψ∗Vψd x≤−6 /~2α2β/[m(α+
β)]. The sum of the two integrals is minimised when β=2α, leading to the
stated upper limit for /angbracketleftE/angbracketright.
22.29 The SL equation has p=1 , q=0 ,a n d ρ=1 .
Useu(x, y)=x(4−x)y(1−y) as a trial function.
Numerator = 1088 /90, denominator = 512 /450. Direct solution k2=1 7π2/16.
861
23
Integral equations
It is not unusual in the analysis of a physical system to encounter an equation
in which an unknown but required function y(x), say, appears under an integral
sign. Such an equation is called an integral equation , and in this chapter we discuss
several methods for solving the more straightforward examples of such equations.
Before embarking on our discussion of methods for solving various integral
equations, we begin with a warning that many of the integral equations met inpractice cannot be solved by the elementary methods presented here but mustinstead be solved numerically, usually on a computer. Nevertheless, the regularoccurrence of several simple types of integral equation that may be solved
analytically is sufficient reason to explore these equations more fully.
We shall begin this chapter by discussing how a differential equation can be
transformed into an integral equation and by considering the most commontypes of linear integral equation. After introducing the operator notation andconsidering the existence of solutions for various types of equation, we go onto discuss elementary methods of obtaining closed-form solutions of simpleintegral equations. We then consider the solution of integral equations in terms of
infinite series and conclude by discussing the properties of integral equations with
Hermitian kernels, i.e. those in which the integrands have particular symmetryproperties.
23.1 Obtaining an integral equation from a differential equation
Integral equations occur in many situations, partly because we may always rewrite
a differential equation as an integral equation. It is sometimes advantageous tomake this transformation, since questions concerning the existence of a solu-
tion are more easily answered for integral equations (see section 23.3), and,
furthermore, an integral equation can incorporate automatically any boundaryconditions on the solution.
862
23.2 TYPES OF INTEGRAL EQUATION
We shall illustrate the principles involved by considering the differential equa-
tion
y/prime/prime(x)=f(x, y), (23.1)
where f(x, y) can be any function of xandybut not of y/prime(x). Equation (23.1)
thus represents a large class of linear and non-linear second-order differential
equations.
We can convert (23.1) into the corresponding integral equation by first inte-
grating with respect to xto obtain
y/prime(x)=integraldisplayx
0f(z,y(z))dz+c1.
Integrating once more, we find
y(x)=integraldisplayx
0duintegraldisplayu
0f(z,y(z))dz+c1x+c2.
Provided we do not change the region in the uz-plane over which the double
integral is taken, we can reverse the order of the two integrations. Changing the
integration limits appropriately, we find
y(x)=integraldisplayx
0f(z,y(z))dzintegraldisplayx
zdu+c1x+c2 (23.2)
=integraldisplayx
0(x−z)f(z,y(z))dz+c1x+c2; (23.3)
this is a non-linear (for general f(x, y))Volterra integral equation.
It is straightforward to incorporate any boundary conditions on the solution
y(x) by fixing the constants c1andc2in (23.3). For example, we might have the
one-point boundary condition y(0) = aandy/prime(0) = b, for which it is clear that
we must set c1=bandc2=a.
23.2 Types of integral equation
From (23.3), we can see that even a relatively simple differential equation such
as (23.1) can lead to a corresponding integral equation that is non-linear. In this
chapter, however, we will restrict our attention to linear integral equations, which
have the general form
g(x)y(x)=f(x)+λintegraldisplayb
aK(x, z)y(z)dz. (23.4)
In (23.4), y(x) is the unknown function, while the functions f(x),g(x)a n d K(x, z)
are assumed known. K(x, z) is called the kernel of the integral equation. The
integration limits aandbare also assumed known, and may be constants or
functions of x,a n d λis a known constant or parameter.
863
INTEGRAL EQUATIONS
In fact, we shall be concerned with various special cases of (23.4), which are
known by particular names. Firstly, if g(x) = 0 then the unknown function y(x)
appears only under the integral sign, and (23.4) is called a linear integral equationof the first kind . Alternatively, if g(x) = 1, so that y(x) appears twice, once inside
the integral and once outside, then (23.4) is called a linear integral equation of
the second kind . In either case, if f(x) = 0 the equation is called homogeneous ,
otherwise inhomogeneous .
We can distinguish further between different types of integral equation by the
form of the integration limits aandb. If these limits are fixed constants then the
equation is called a Fredholm equation. If, however, the upper limit b=x(i.e. it
is variable) then the equation is called a Volterra equation; such an equation is
analogous to one with fixed limits but for which the kernel K(x, z)=0f o r z>x.
Finally, we note that any equation for which either (or both) of the integrationlimits is infinite, or for which K(x, z) becomes infinite in the range of integration,
is called a singular integral equation.
23.3 Operator notation and the existence of solutions
There is a close correspondence between linear integral equations and the matrix
equations discussed in chapter 8. However, the former involve linear, integral rela-
tions between functions in an infinite-dimensional function space (see chapter 17),whereas the latter specify linear relations among vectors in a finite-dimensionalvector space.
Since we are restricting our attention to linear integral equations, it will be
convenient to introduce the linear integral operator K, whose action on an
arbitrary function yis given by
Ky=integraldisplay
b
aK(x, z)y(z)dz. (23.5)
This is analogous to the introduction in chapters 16 and 17 of the notation Lto
describe a linear differential operator. Furthermore, we may define the Hermitianconjugate K
†by
K†y=integraldisplayb
aK∗(z,x)y(z)dz,
where the asterisk denotes complex conjugation and we have reversed the order
of the arguments in the kernel.
It is clear from (23.5) that Kis indeed linear. Moreover, since Koperates on
the infinite-dimensional space of (reasonable) functions, we may make an obviousanalogy with matrix equations and consider the action of Kon a function fas
that of a matrix on a column vector (both of infinite dimension).
When written in operator form, the integral equations discussed in the pre-
vious section resemble equations familiar from linear algebra. For example, the
864
23.4 CLOSED-FORM SOLUTIONS
inhomogeneous Fredholm equation of the first kind may be written as
0=f+λKy,
which has the unique solution y=−K−1f/λ, provided that f/negationslash= 0 and the inverse
operator K−1exists.
Similarly, we may write the corresponding Fredholm equation of the second
kind as
y=f+λKy. (23.6)
In the homogeneous case, where f= 0, this reduces to y=λKy,w h i c hi s
reminiscent of an eigenvalue problem in linear algebra (except that λappears on
the other side of the equation) and, similarly, only has solutions for at most acountably infinite set of eigenvalues λ
i. The corresponding solutions yiare called
the eigenfunctions.
In the inhomogeneous case ( f/negationslash= 0), the solution to (23.6) can be written
symbolically as
y=( 1−λK)−1f,
again provided that the inverse operator exists. It may be shown that, in general,
(23.6) does possess a unique solution if λ/negationslash=λi,i . e .w h e n λdoes not equal one of
the eigenvalues of the corresponding homogeneous equation.
When λdoes equal one of these eigenvalues, (23.6) may have either many
solutions or no solution at all, depending on the form of f. If the function fis
orthogonal to everyeigenfunction of the equation
g=λ∗K†g (23.7)
that belongs to the eigenvalue λ∗,i . e .
/angbracketleftg|f/angbracketright=integraldisplayb
ag∗(x)f(x)dx=0
for every function gobeying (23.7), then it can be shown that (23.6) has many
solutions. Otherwise the equation has no solution. These statements are discussed
further in section 23.7, for the special case of integral equations with Hermitiankernels, i.e. those for which K=K
†.
23.4 Closed-form solutions
In certain very special cases, it may be possible to obtain a closed-form solution
of an integral equation. The reader should realise, however, when faced with an
integral equation, that in general it will not be soluble by the simple methods
presented in this section but must instead be solved using (numerical) iterativemethods, such as those outlined in section 23.5.
865
INTEGRAL EQUATIONS
23.4.1 Separable kernels
The most straightforward integral equations to solve are Fredholm equations
withseparable (ordegenerate ) kernels. A kernel is separable if it has the form
K(x, z)=nsummationdisplay
i=1φi(x)ψi(z), (23.8)
where φi(x)a r e ψi(z) are respectively functions of xonly and of zonly and the
number of terms in the sum, n, is finite.
Let us consider the solution of the (inhomogeneous) Fredholm equation of the
second kind,
y(x)=f(x)+λintegraldisplayb
aK(x, z)y(z)dz, (23.9)
which has a separable kernel of the form (23.8). Writing the kernel in its separated
form, the functions φi(x) may be taken outside the integral over zto obtain
y(x)=f(x)+λnsummationdisplay
i=1φi(x)integraldisplayb
aψi(z)y(z)dz.
Since the integration limits aandbare constant for a Fredholm equation, the
integral over zin each term of the sum is just a constant. Denoting these constants
by
ci=integraldisplayb
aψi(z)y(z)dz, (23.10)
the solution to (23.9) is found to be
y(x)=f(x)+λnsummationdisplay
i=1ciφi(x), (23.11)
where the constants cican be evalutated by substituting (23.11) into (23.10).ISolve the integral equation
y(x)=x+λ
Z1
0(xz+z2)y(z)dz. (23.12)
The kernel for this equation is K(x, z)=xz+z2, which is clearly separable, and using the
notation in (23.8) we have φ1(x)=x,φ2(x)=1 , ψ1(z)=zandψ2(z)=z2. From (23.11)
the solution to (23.12) has the form
y(x)=x+λ(c1x+c2),
where the constants c1andc2are given by (23.10) as
c1=
Z1
0z[z+λ(c1z+c2)]dz=1
3+1
3λc1+1
2λc2,
c2=
Z1
0z2[z+λ(c1z+c2)]dz=1
4+1
4λc1+1
3λc2.
866
23.4 CLOSED-FORM SOLUTIONS
These two simultaneous linear equations may be straightforwardly solved for c1andc2to
give
c1=24 + λ
72−48λ−λ2and c2=18
72−48λ−λ2,
so that the solution to (23.12) is
y(x)=(72−24λ)x+1 8λ
72−48λ−λ2.
J
In the above example, we see that (23.12) has a (finite) unique solution provided
that λis not equal to either root of the quadratic in the denominator of y(x).
The roots of this quadratic are in fact the eigenvalues of the corresponding
homogeneous equation, as mentioned in the previous section. In general, if theseparable kernel contains nterms, as in (23.8), there will be nsuch eigenvalues,
although they may not all be different.
Kernels consisting of trigonometric (or hyperbolic) functions of sums or differ-
ences of xandzare also often separable.IFind the eigenvalues and corresponding ei genfunctions of the homogeneous Fredholm
equation
y(x)=λ
Zπ
0sin(x+z)y(z)dz. (23.13)
The kernel of this integral equation can be written in separated form as
K(x, z)=s i n ( x+z)=s i n xcosz+c o s xsinz,
so, comparing with (23.8), we have φ1(x)=s i n x,φ2(x)=c o s x,ψ1(z)=c o s zand
ψ2(z)=s i n z.
Thus, from (23.11), the solution to (23.13) has the form
y(x)=λ(c1sinx+c2cosx),
where the constants c1andc2are given by
c1=λ
Zπ
0cosz(c1sinz+c2cosz)dz=λπ
2c2, (23.14)
c2=λ
Zπ
0sinz(c1sinz+c2cosz)dz=λπ
2c1. (23.15)
Combining these two equations we find c1=(λπ/2)2c1, and, assuming that c1/negationslash=0 ,t h i s
gives λ=±2/π, the two eigenvalues of the integral equation (23.13).
By substituting each of the eigenvalues back into (23.14) and (23.15), we find that
the eigenfunctions corresponding to the eigenvalues λ1=2/πandλ2=−2/πare given
respectively by
y1(x)=A(sinx+c o s x)a n d y2(x)=B(sinx−cosx), (23.16)
where AandBare arbitrary constants.
J
867
INTEGRAL EQUATIONS
23.4.2 Integral transform methods
If the kernel of an integral equation can be written as a function of the difference
x−zof its two arguments, then it is called a displacement kernel. An integral
equation having such a kernel, and which also has the integration limits −∞to
∞, may be solved by the use of Fourier transforms (chapter 13).
If we consider the following integral equation with a displacement kernel,
y(x)=f(x)+λintegraldisplay∞
−∞K(x−z)y(z)dz, (23.17)
the integral over zclearly takes the form of a convolution (see chapter 13).
Therefore, Fourier-transforming (23.17) and using the convolution theorem, we
obtain
˜y(k)=˜f(k)+√
2πλ˜K(k)˜y(k),
which may be rearranged to give
˜y(k)=˜f(k)
1−√
2πλ˜K(k). (23.18)
Taking the inverse Fourier transform, the solution to (23.17) is given by
y(x)=1√
2πintegraldisplay∞
−∞˜f(k)ex p( ikx)
1−√
2πλ˜K(k)dk.
If we can perform this inverse Fourier transformation then the solution can be
found explicitly; otherwise it must be left in the form of an integral.IFind the Fourier transform of the function
g(x)=
/(
1if|x|≤a,
0if|x|>a.
Hence find an explicit expression for the solution of the integral equation
y(x)=f(x)+λ
Z∞
−∞sin(x−z)
x−zy(z)dz. (23.19)
Find the solution for the special case f(x)=( s i n x)/x.
The Fourier transform of g(x) is given directly by
˜g(k)=1√
2π
Za
−aexp(−ikx)dx=
/1√
2πexp(−ikx)
(−ik)
/a
−a=
r
2
πsinka
k.
(23.20)
The kernel of the integral equation (23.19) is K(x−z)=[ s i n ( x−z)]/(x−z). Using
(23.20), it is straightforward to show that the Fourier transform of the kernel is
˜K(k)=
/(p
π/2i f|k|≤1,
0i f|k|>1.(23.21)
868
23.4 CLOSED-FORM SOLUTIONS
Thus, using (23.18), we find the Fourier transform of the solution to be
˜y(k)=
/(˜f(k)/(1−πλ)i f|k|≤1,
˜f(k)i f |k|>1.(23.22)
Inverse Fourier-transforming, and writing the result in a slightly more convenient form,
the solution to (23.19) is given by
y(x)=f(x)+
/1
1−πλ−1
/1√
2π
Z1
−1˜f(k)e x p ( ikx)dk
=f(x)+πλ
1−πλ1√
2π
Z1
−1˜f(k)e x p ( ikx)dk. (23.23)
It is clear from (23.22) that when λ=1/π, which is the only eigenvalue of the
corresponding homogeneous equation to (23.19), the solution becomes infinite, as wewould expect.
For the special case f(x)=( s i n x)/x, the Fourier transform ˜f(k) is identical to that in
(23.21), and the solution (23.23) becomes
y(x)=sinx
x+
/πλ
1−πλ
/1√
2π
Z1
−1
rπ
2exp(ikx)dk
=sinx
x+
/πλ
1−πλ
/1
2
/exp(ikx)
ix
/k=1
k=−1
=sinx
x+
/πλ
1−πλ
/sinx
x=
/1
1−πλ
/sinx
x.
J
If, instead, the integral equation (23.17) had integration limits 0 and x(so
making it a Volterra equation) then its solution could be found, in a similar way,by using the convolution theorem for Laplace transforms (see chapter 13). Wewould find
¯y(s)=¯f(s)
1−λ¯K(s),
where sis the Laplace transform variable. Often one may use the dictionary of
Laplace transforms given in table 13.1 to invert this equation and find the solution
y(x). In general, however, the evaluation of inverse Laplace transform integrals
is difficult, since (in principle) it requires a contour integration; see chapter 20.
As a final example of the use of Fourier transforms in solving integral equations,
we mention equations that have integration limits −∞and∞and a kernel of the
form
K(x, z)=e x p (−ixz).
Consider, for example, the inhomogeneous Fredholm equation
y(x)=f(x)+λintegraldisplay∞
−∞exp(−ixz)y(z)dz. (23.24)
The integral over zis clearly just (a multiple of) the Fourier transform of y(z),
869
INTEGRAL EQUATIONS
so we can write
y(x)=f(x)+√
2πλ˜y(x). (23.25)
If we now take the Fourier transform of (23.25) but continue to denote the
independent variable by x(i.e. rather than k, for example), we obtain
˜y(x)=˜f(x)+√
2πλy(−x). (23.26)
Substituting (23.26) into (23.25) we find
y(x)=f(x)+√
2πλbracketleftBig
˜f(x)+√
2πλy(−x)bracketrightBig
,
but on making the change x→−xand substituting back in for y(−x), this gives
y(x)=f(x)+√
2πλ˜f(x)+2πλ2bracketleftBig
f(−x)+√
2πλ˜f(−x)+2πλ2y(x)bracketrightBig
.
Thus the solution to (23.24) is given by
y(x)=1
1−(2π)2λ4bracketleftBig
f(x)+( 2 π)1/2λ˜f(x)+2πλ2f(−x)+( 2 π)3/2λ3˜f(−x)bracketrightBig
.
(23.27)
Clearly, (23.24) possesses a unique solution provided λ/negationslash=±1/√
2πor±i/√
2π;
these are easily shown to be the eigenvalues of the corresponding homogeneousequation (for which f(x)≡0).ISolve the integral equation
y(x)=e x p
/
−x2
2
/
+λ
Z∞
−∞exp(−ixz)y(z)dz, (23.28)
where λis a real constant. Show that the solution is unique unless λhas one of two particular
values. Does a solution exist for either of these two values of λ?
Following the argument given above, the solution to (23.28) is given by (23.27) with
f(x)=e x p (−x2/2). In order to write the solution exp licitly, however, we must calculate
the Fourier transform of f(x). Using equation (13.7), we find ˜f(k)=e x p (−k2/2), from
which we note that f(x) has the special property that its functional form is identical to
that of its Fourier transform. Thus, the solution to (23.28) is given by
y(x)=1
1−(2π)2λ4
/
1+( 2 π)1/2λ+2πλ2+( 2π)3/2λ3
/
exp
/
−x2
2
/
.
(23.29)
Since λis restricted to be real, the solution to (23.28) will be unique unless λ=±1/√
2π,
at which points (23.29) becomes infinite. In order to find whether solutions exist for eitherof these values of λwe must return to equations (23.25) and (23.26).
Let us first consider the case λ=+ 1 /√
2π. Putting this value into (23.25) and (23.26),
we obtain
y(x)=f(x)+˜y(x), (23.30)
˜y(x)=˜f(x)+y(−x). (23.31)
870
23.4 CLOSED-FORM SOLUTIONS
Substituting (23.31) into (23.30) we find
y(x)=f(x)+˜f(x)+y(−x),
but on changing xto−xand substituting back in for y(−x), this gives
y(x)=f(x)+˜f(x)+f(−x)+˜f(−x)+y(x).
Thus, in order for a solution to exist, we require that the function f(x)o b e y s
f(x)+˜f(x)+f(−x)+˜f(−x)=0 .
This is satisfied if f(x)=−˜f(x), i.e. if the functional form of f(x)i sm i n u st h ef o r mo fi t s
Fourier transform. We may repeat this analysis for the case λ=−1/√
2π, and, in a similar
way, we find that this time we require f(x)=˜f(x).
In our case f(x)=e x p (−x2/2), for which, as we mentioned above, f(x)=˜f(x).
Therefore, (23.28) possesses no solution when λ=+ 1 /√
2πbut has many solutions when
λ=−1/√
2π.
J
A similar approach to the above may be taken to solve equations with kernels
of the form K(x, y)=c o s xyor sin xy, either by considering the integral over yin
each case as the real or imaginary part of the corresponding Fourier transformor by using Fourier cosine or sine transforms directly.
23.4.3 Differentiation
A closed-form solution to a Volterra equation may sometimes be obtained by
differentiating the equation to obtain the corresponding differential equation,which may be easier to solve.ISolve the integral equation
y(x)=x−
Zx
0xz2y(z)dz. (23.32)
Dividing through by x,w eo b t a i n
y(x)
x=1−
Zx
0z2y(z)dz,
which may be differentiated with respect to xto give
d
dx
/y(x)
x
/
=−x2y(x)=−x3
/y(x)
x
/
.
This equation may be integrated straightforwardly, and we find
ln
/y(x)
x
/
=−x4
4+c,
where cis a constant of integration. Thus the solution to (23.32) has the form
y(x)=Axexp
/
−x4
4
/
, (23.33)
where Ais an arbitrary constant.
Since the original integral equation (23.32) contains no arbitrary constants, neither
should its solution. We may calculate the value of the constant, A, by substituting the
solution (23.33) back into (23.32), from which we find A=1 .
J
871
INTEGRAL EQUATIONS
23.5 Neumann series
As mentioned above, most integral equations met in practice will not be of the
simple forms discussed in the last section and so, in general, it is not possible tofind closed-form solutions. In such cases, we might try to obtain a solution in theform of an infinite series, as we did for differential equations (see chapter 16).
Let us consider the equation
y(x)=f(x)+λintegraldisplay
b
aK(x, z)y(z)dz, (23.34)
where either both integration limits are constants (for a Fredholm equation) or
the upper limit is variable (for a Volterra equation). Clearly, if λwere small then
a crude (but reasonable) approximation to the solution would be
y(x)≈y0(x)=f(x),
where y0(x) stands for our ‘zeroth-order’ approximation to the solution (and is
not to be confused with an eigenfunction).
Substituting this crude guess under the integral sign in the original equation,
we obtain what should be a better approximation:
y1(x)=f(x)+λintegraldisplayb
aK(x, z)y0(z)dz=f(x)+λintegraldisplayb
aK(x, z)f(z)dz,
which is first order in λ. Repeating the procedure once more results in the
second-order approximation
y2(x)=f(x)+λintegraldisplayb
aK(x, z)y1(z)dz
=f(x)+λintegraldisplayb
aK(x, z1)f(z1)dz1+λ2integraldisplayb
adz1integraldisplayb
aK(x, z1)K(z1,z2)f(z2)dz2.
It is clear that we may continue this process to obtain progressively higher-order
approximations to the solution. Introducing the functions
K1(x, z)=K(x, z),
K2(x, z)=integraldisplayb
aK(x, z1)K(z1,z)dz1,
K3(x, z)=integraldisplayb
adz1integraldisplayb
aK(x, z1)K(z1,z2)K(z2,z)dz2,
a n ds oo n ,w h i c ho b e yt h er e c u r r e n c er e l a t i o n
Kn(x, z)=integraldisplayb
aK(x, z1)Kn−1(z1,z)dz1,
872
23.5 NEUMANN SERIES
we may write the nth-order approximation as
yn(x)=f(x)+nsummationdisplay
m=1λmintegraldisplayb
aKm(x, z)f(z)dz. (23.35)
The solution to the original integral equation is then given by y(x)=
limn→∞yn(x),provided the infinite series converges . Using (23.35), this solution
may be written as
y(x)=f(x)+λintegraldisplayb
aR(x, z;λ)f(z)dz, (23.36)
where the resolvent kernel R(x, z;λ)i sg i v e nb y
R(x, z;λ)=∞summationdisplay
m=0λmKm+1(x, z). (23.37)
Clearly, the resolvent kernel, and hence the series solution, will converge
provided λis sufficiently small. In fact, it may be shown that the series converges
in some domain of |λ|provided the original kernel K(x, z) is bounded in such a
way that
|λ|2integraldisplayb
adxintegraldisplayb
a|K(x, z)|2dz <1. (23.38)IUse the Neumann series method to solve the integral equation
y(x)=x+λ
Z1
0xzy(z)dz. (23.39)
Following the method outlined above, we begin with the crude approximation y(x)≈
y0(x)=x. Substituting this under the integral sign in (23.39), we obtain the next approxi-
mation
y1(x)=x+λ
Z1
0xzy0(z)dz=x+λ
Z1
0xz2dz=x+λx
3,
Repeating the procedure once more, we obtain
y2(x)=x+λ
Z1
0xzy1(z)dz
=x+λ
Z1
0xz
/
z+λz
3
/
dz=x+
/λ
3+λ2
9
/
x.
For this simple example, it is easy to see that by continuing this process the solution to
(23.39) is obtained as
y(x)=x+
/"
λ
3+
/λ
3
/2
+
/λ
3
/3
+···
/#
x.
Clearly the expression in brackets is an infinite geometric series with first term λ/3a n d
873
INTEGRAL EQUATIONS
common ratio λ/3. Thus, provided|λ|<3, this infinite series converges to the value
λ/(3−λ), and the solution to (23.39) is
y(x)=x+λx
3−λ=3x
3−λ. (23.40)
Finally, we note that the requirement that |λ|<3 may also be derived very easily from
the condition (23.38).
J
23.6 Fredholm theory
In the previous section, we found that a solution to the integral equation (23.34)
can be obtained as a Neumann series of the form (23.36), where the resolventkernel R(x, z;λ) is written as an infinite power series in λ. This solution is valid
provided the infinite series converges.
A related, but more elegant, approach to the solution of integral equations
using infinite series was found by Fredholm. We will not reproduce Fredholm’s
analysis here, but merely state the results we need. Essentially, Fredholm theory
provides a formula for the resolvent kernel R(x, z;λ) in (23.36) in terms of the
ratio of two infinite series:
R(x, z;λ)=D(x, z;λ)
d(λ). (23.41)
The numerator and denominator in (23.41) are given by
D(x, z;λ)=∞summationdisplay
n=0(−1)n
n!Dn(x, z)λn, (23.42)
d(λ)=∞summationdisplay
n=0(−1)n
n!dnλn, (23.43)
where the functions Dn(x, z) and the constants dnare found from recurrence
relations as follows. We start with
D0(x, z)=K(x, z)a n d d0=1, (23.44)
where K(x, z) is the kernel of the original integral equation (23.34). The higher-
order coefficients of λin (23.43) and (23.42) are then obtained from the two
recurrence relations
dn=integraldisplayb
aDn−1(x, x)dx, (23.45)
Dn(x, z)=K(x, z)dn−nintegraldisplayb
aK(x, z1)Dn−1(z1,z)dz1. (23.46)
Although the formulae for the resolvent kernel appear complicated, they are
often simple to apply. Moreover, for the Fredholm solution the power series(23.42) and (23.43) are both guaranteed to converge for all values of λ, unlike
874
23.7 SCHMIDT–HILBERT THEORY
Neumann series, which converge only if the condition (23.38) is satisfied. Thus the
Fredholm method leads to a unique, non-singular solution, provided that d(λ)/negationslash=0 .
In fact, as we might suspect, the solutions of d(λ) = 0 give the eigenvalues of the
homogeneous equation corresponding to (23.34), i.e. with f(x)≡0.IUse Fredholm theory to solve the integral equation (23.39).
Using (23.36) and (23.41), the solution to (23.39) can be written in the form
y(x)=x+λ
Z1
0R(x, z;λ)zd z=x+λ
Z1
0D(x, z;λ)
d(λ)zd z . (23.47)
In order to find the form of the resolvent kernel R(x, z;λ), we begin by setting
D0(x, z)=K(x, z)=xz and d0=1
and use the recurrence relations (23.45) and (23.46) to obtain
d1=
Z1
0D0(x, x)dx=
Z1
0x2dx=1
3,
D1(x, z)=xz
3−
Z1
0xz2
1zd z1=xz
3−xz
/z3
1
3
/1
0=0.
Applying the recurrence relations again we find that dn=0a n d Dn(x, z)=0f o r n>1.
Thus, from (23.42) and (23.43), the numerator and denominator of the resolvent respectively
are given by
D(x, z;λ)=xz and d(λ)=1−λ
3.
Substituting these expressions into (23.47), we find that the solution to (23.39) is given
by
y(x)=x+λ
Z1
0xz2
1−λ/3dz
=x+λ
/x
1−λ/3z3
3
/1
0=x+λx
3−λ=3x
3−λ,
which, as expected, is the same as the solution (23.40) found by constructing a Neumann
series.
J
23.7 Schmidt–Hilbert theory
The Schmidt–Hilbert (SH) theory of integral equations may be considered as
analogous to the Sturm–Liouville (SL) theory of differential equations, discussed
in chapter 17, and is concerned with the properties of integral equations withHermitian kernels. An Hermitian kernel enjoys the property
K(x, z)=K
∗(z,x), (23.48)
and it is clear that a special case of (23.48) occurs for a real kernel that is also
symmetric with respect to its two arguments.
875
INTEGRAL EQUATIONS
Let us begin by considering the homogeneous integral equation
y=λKy,
where the integral operator Khas an Hermitian kernel. As discussed in sec-
tion 23.3, in general, this equation will have solutions only for λ=λi,w h e r et h e λi
are the eigenvalues of the integral equation, the corresponding solutions yibeing
the eigenfunctions of the equation.
By following similar arguments to those presented in chapter 17 for SL theory,
it may be shown that the eigenvalues λiof an Hermitian kernel are real and
that the corresponding eigenfunctions yibelonging to different eigenvalues are
orthogonal and form a complete set. If the eigenfunctions are suitably normalised,
we have
/angbracketleftyi|yj/angbracketright=integraldisplayb
ay∗
i(x)yj(x)dx=δij. (23.49)
If an eigenvalue is degenerate then the eigenfunctions corresponding to that
eigenvalue can be made orthogonal by the Gram–Schmidt procedure, in a similarway to that discussed in chapter 17 in the context of SL theory.
Like SL theory, SH theory does not provide a method of obtaining the eigen-
values and eigenfunctions of any particular homogeneous integral equation withan Hermitian kernel; for this we have to turn to the methods discussed in the
previous sections of this chapter. Rather, SH theory is concerned with the gen-
eral properties of the solutions to such equations. Where SH theory becomesapplicable, however, is in the solution of inhomogeneous integral equations withHermitian kernels for which the eigenvalues and eigenfunctions of the corre-sponding homogeneous equation are already known.
Let us consider the inhomogeneous equation
y=f+λKy, (23.50)
where K=K
†and for which we know the eigenvalues λiand normalised
eigenfunctions yiof the corresponding homogeneous problem. The function f
may or may not be expressible solely in terms of the eigenfunctions yi,a n dt o
accommodate this situation we write the unknown solution yasy=f+summationtext
iaiyi,
where the aiare expansion coefficients to be determined.
Substituting this into (23.50), we obtain
f+summationdisplay
iaiyi=f+λsummationdisplay
iaiyi
λi+λKf, (23.51)
where we have used the fact that yi=λiKyi. Forming the inner product of both
876
23.7 SCHMIDT–HILBERT THEORY
sides of (23.51) with yj, we find
summationdisplay
iai/angbracketleftyj|yi/angbracketright=λsummationdisplay
iai
λi/angbracketleftyj|yi/angbracketright+λ/angbracketleftyj|Kf/angbracketright. (23.52)
Since the eigenfunctions are orthonormal and Kis an Hermitian operator,
we have that both /angbracketleftyj|yi/angbracketright=δijand/angbracketleftyj|Kf/angbracketright=/angbracketleftKyj|f/angbracketright=λ−1
j/angbracketleftyj|f/angbracketright. Thus the
coefficients ajare given by
aj=λλ−1
j/angbracketleftyj|f/angbracketright
1−λλ−1
j=λ/angbracketleftyj|f/angbracketright
λj−λ, (23.53)
and the solution is
y=f+summationdisplay
iaiyi=f+λsummationdisplay
i/angbracketleftyi|f/angbracketright
λi−λyi. (23.54)
This also shows, incidentally, that a formal representation for the resolvent kernel
is
R(x, z;λ)=summationdisplay
iyi(x)y∗
i(z)
λi−λ. (23.55)
Iffcanbe expressed as a linear superposition of the yi,i . e .f=summationtext
ibiyi,t h e n
bi=/angbracketleftyi|f/angbracketrightand the solution can be written more briefly as
y=summationdisplay
ibi
1−λλ−1
iyi. (23.56)
We see from (23.54) that the inhomogeneous equation (23.50) has a unique
solution provided λ/negationslash=λi,i . e .w h e n λis not equal to one of the eigenvalues of
the corresponding homogeneous equation. However, if λdoes equal one of the
eigenvalues λjthen, in general, the coefficients ajbecome singular and no (finite)
solution exists.
Returning to (23.53) we notice that even if λ=λja non-singular solution to
the integral equation is still possible provided that the function fis orthogonal
to every eigenfunction corresponding to the eigenvalue λj,i . e .
/angbracketleftyj|f/angbracketright=integraldisplayb
ay∗
j(x)f(x)dx=0.
The following worked example illustrates the case in which fcan be expressed in
terms of the yi. One in which it cannot is considered in exercise 23.14.
877
INTEGRAL EQUATIONSIUse Schmidt–Hilbert theory to solve the integral equation
y(x) = sin( x+α)+λ
Zπ
0sin(x+z)y(z)dz. (23.57)
It is clear that the kernel K(x, z)=s i n ( x+z) is real and symmetric in xandzand is
thus Hermitian. In order to solve this inhomogeneous equation using SH theory, however,we must first find the eigenvalues and eigenfunctions of the corresponding homogeneousequation.
In fact, we have considered the solution of the corresponding homogeneous equation
(23.13) already in subsection 23.4.1, where we found that it has two eigenvalues λ
1=2/π
andλ2=−2/π, with eigenfunctions given by (23.16). The normalised eigenfunctions are
y1(x)=1√π(sinx+c o s x)a n d y2(x)=1√π(sinx−cosx) (23.58)
and are easily shown to obey the orthonormality condition (23.49).
Using (23.54), the solution to the inhomogeneous equation (23.57) has the form
y(x)=a1y1(x)+a2y2(x), (23.59)
where the coefficients a1anda2are given by (23.53) with f(x) = sin( x+α). Therefore,
using (23.58),
a1=1
1−πλ/2
Zπ
01√π(sinz+c o s z)si n(z+α)dz=√π
2−πλ(cosα+s i n α),
a2=1
1+πλ/2
Zπ
01√π(sinz−cosz)si n(z+α)dz=√π
2+πλ(cosα−sinα).
Substituting these expressions for a1anda2into (23.59) and simplifying, we find that
the solution to (23.57) is given by
y(x)=1
1−(πλ/2)2
/
sin(x+α)+(πλ/2)cos( x−α)
/
.
J
23.8 Exercises
23.1 Solve the integral equationZ∞
0cos(xv)y(v)dv=e x p (−x2/2),
for the function y=y(x)f o r x>0. Note that for x<0,y(x) can be chosen as
is most convenient.
23.2 SolveZ∞
0f(t)exp(−st)dt=a
a2+s2.
23.3 Use the fact that its kernel is separable to solve for y(x) the integral equation
y(x)=Acos(x+a)+λ
Zπ
0sin(x+z)y(z)dz.
(This equation is an inhomogeneous extension of the homogeneous Fredholm
equation (23.13), and is similar to equation (23.57).)
878
23.8 EXERCISES
23.4 Convert
f(x)=e x p x+
Zx
0(x−y)f(y)dy
into a differential equation, and hence show that its solution is
(α+βx)exp x+γexp(−x),
where α, β, γ are constants that should be determined.
23.5 Solve for φ(x) the integral equation
φ(x)=f(x)+λ
Z1
0
//x
y
/n
+
/y
x
/n
/
φ(y)dy,
where f(x) is bounded for 0 <x< 1a n d−1
2<n<1
2, expressing your answer
in terms of the quantities Fm=
R1
0f(y)ymdy.
(a) Give the explicit solution when λ=1 .
(b) For what values of λare there no solutions unless F±ntake particular values?
What are these values?
23.6 (a) Consider the inhomogeneous integral equation
f(x)=g(x)+λ
Zb
aK(x, y)f(y)dy;
its kernel K(x, y) is real, symmetric and continuous in a≤x≤b,a≤y≤b.
Ifλis one of the eigenvalues λiof the homogeneous equation
fi(x)=λi
Zb
aK(x, y)fi(y)dy,
prove that the inhomogeneous equation can only a have non-trivial solution
ifg(x) is orthogonal to the corresponding eigenfunction fi(x).
(b) Show that the only values of λfor which
f(x)=λ
Z1
0xy(x+y)f(y)dy
has a non-trivial solution are the roots of the equation
λ2+ 120 λ−240 = 0 .
(c) Solve
f(x)=µx2+
Z1
02xy(x+y)f(y)dy.
23.7 (a) If the kernel of the integral equation
ψ(x)=λ
Zb
aK(x, y)ψ(y)dy
has the form
K(x, y)=∞X
n=0hn(x)gn(y),
where the hn(x) form a complete orthonormal set of functions over the
interval [ a, b], show that the eigenvalues λiare given by
|M−λ−1I|=0,
879
INTEGRAL EQUATIONS
where Mis the matrix with elements
Mkj=
Zb
agk(u)hj(u)du.
If the corresponding solutions are ψ(i)(x)=
P∞
n=0a(i)
nhn(x), find an expression
fora(i)
n.
(b) Obtain the eigenvalues and eigenfunctions over the interval [0 ,2π]i f
K(x, y)=∞X
n=11
ncosnxcosny.
23.8 By taking its Laplace transform, and that of xne−ax, obtain the explicit solution
of
f(x)=e−x
/
x+
Zx
0(x−u)euf(u)du
/
.
Verify your answer by substitution.
23.9 For f(t)=exp(−t2
2), use the relationships of the Fourier transforms of f/prime(t)a n d
tf(t)t ot h a to f f(t) itself to find a simple differential equation satisfied by ˜f(ω),
the Fourier transform of f(t) and hence determine ˜f(ω) to within a constant. Use
this result to solve the integral equationZ∞
−∞e−t(t−2x)/2h(t)dt=e3x2/8
forh(t).
23.10 Show that the equation
f(x)=x−1/3+λ
Z∞
0f(y)e x p(−xy)dy
has a solution of the form Axα+Bxβ. Determine the values of αandβand show
that those of AandBare
1
1−λ2Γ(1
3)Γ(2
3)andλΓ(2
3)
1−λ2Γ(1
3)Γ(2
3),
where Γ( z) is the gamma function, discussed in the appendix.
23.11 At an international ‘peace’ conference a large number of delegates are seated
around a circular table with each delegation sitting near its allies and diametricallyopposite the delegation most bitterly opposed to it. The position of a delegate isdenoted by θ,w i t h0≤θ≤2π.T h ef u r y f(θ)f e l tb yt h ed e l e g a t ea t θis the sum
of his own natural hostility h(θ) and the influences on him of each of the other
delegates; a delegate at position φcontributes an amount K(θ−φ)f(φ). Thus
f(θ)=h(θ)+Z2π
0K(θ−φ)f(φ)dφ.
Show that if K(ψ)t a k e st h ef o r m K(ψ)=k0+k1cosψthen
f(θ)=h(θ)+p+qcosθ+rsinθ
and evaluate p,qandr. A positive value for k1implies that delegates tend to
placate their opponents but upset their allies, whilst negative values imply thatthey calm their allies but infuriate their opponents. A walkout will occur if f(θ)
exceeds a certain threshold value for some θ. Is this more likely to happen for
positive or for negative values of k
1?
880
23.8 EXERCISES
23.12 By considering functions of the form h(x)=
Rx
0(x−y)f(y)dy, show that the
solution f(x)o ft h ei n t e g r a le q u a t i o n
f(x)=x+1
2
Z1
0|x−y|f(y)dy
satisfies the equation f/prime/prime(x)=f(x).
By examining the special cases x=0a n d x= 1, show that
f(x)=2
(e+3 ) ( e+1 )[(e+2 )ex−ee−x].
23.13 The operator Mis defined by
Mf(x)≡
Z∞
−∞K(x, y)f(y)dy,
where K(x, y) = 1 inside the square |x|<a ,|y|<a, and is equal to 0 elsewhere.
Consider the possible eigenvalues of Mand the eigenfunctions that correspond
to them; show that the only possible eigenvalues are 0 and 2 aand determine the
corresponding eigenfunctions. Hence find the general solution of
f(x)=g(x)+λ
Z∞
−∞K(x, y)f(y)dy.
23.14 For the integral equation
y(x)=x−3+λ
Zb
ax2z2y(z)dz,
show that the resolvent kernel is 5 x2z2/[5−λ(b5−a5)] and hence solve the
equation. For what range of λis the solution valid?
23.15 Use Fredholm theory to show that, for the kernel
K(x, z)=(x+z)e x p( x−z)
over the interval [0 ,1], the resolvent kernel is
R(x, z;λ)=exp(x−z)[(x+z)−λ(1
2x+1
2z−xz−1
3)]
1−λ−1
12λ2,
and hence solve
y(x)=x2+2
Z1
0(x+z)e x p ( x−z)y(z)dz,
expressing your answer in terms of In,w h e r e In=
R1
0unexp(−u)du.
23.16 (a) Determine the eigenvalues λ±of the kernel K(x, z)=(xz)1/2(x1/2+z1/2)a n d
show that the corresponding eigenfunctions have the forms
y±(x)=A±(√2x1/2±√3x),
where A2
±=5/(10±4√6).
(b) Use Schmidt–Hilbert theory to solve
y(x)=1+5
2
Z1
0K(x, z)y(z)dz.
(c) As may be apparent, the algebra involved in the formal method used in (b)
is long and error-prone, and it is in fact much more straightforward to usea trial function 1 + αx
1/2+βx. Check your answer by doing so.
881
INTEGRAL EQUATIONS
23.9 Hints and answers
23.1 Define y(−x)=y(x) and use the cosine Fourier transform inversion theorem;
y(x)=( 2 /π)1/2exp(−x2/2).
23.2 Use the Laplace transform; f(t)=s i n at.
23.3 Set y(x)=c1sinx+c2cosx;y(x)=A[cos(x+a)+(λπ/2)sin( x−a)]/[1−(λ2π2/4)].
23.4 f/prime/prime(x)−f(x)=e x p x;α=3/4,β=1/2,γ=1/4.
23.5 (a) φ(x)=f(x)−( 1+2 n)Fnxn−(1−2n)F−nx−n. (b) There are no solutions for
λ=[ 1±(1−4n2)−1/2]−1unless F±n=0o r F−n/Fn=±[(1−2n)/(1 + 2 n)]1/2.
23.6 (b) Set f(x)=a1x2+a2xand obtain a1=(λ/4)a1+(λ/3)a2,a2=(λ/5)a1+(λ/4)a2;
(c) set f(x)=(µ+a1)x2+a2x;f(x)=−6µx(5x+4 ) .
23.7 (a) a(i)
n=
Rb
ahn(x)ψ(x)dx; (b) use (1 /√π)cosnxand (1 /√π)si nnx;Mis diagonal;
eigenvalues λk=k/πwith ψ(k)(x)=( 1 /√π)c oskx.
23.8 Writing p(x)= exf(x)a n d q(x)= x, the integrand can be expressed as a
convolution. Show that ¯p(s)=¯q(s)/[1−¯q(s)], leading to f(x)=( 1−e−2x)/2.
23.9 d˜f/dω =−ω˜f, leading to ˜f(ω)=Ae−ω2/2. Rearrange the integral as a convolution
and deduce that ˜h(ω)= Be−3ω2/2;h(t)= Ce−t2/6, where re-substitution and
Gaussian normalisation show that C=√
6π/2.
23.10 Recall or prove that the Laplace transform of x−n,w h e r e n<1 but is not
necessarily an integer, is Γ(1 −n)sn−1. For a possible solution α=−1/3a n d
β=−2/3, or vice versa.
23.11 p=k0H/(1−2πk0),q=k1Hc/(1−1
2k1), and r=k1Hs/(1−1
2k1),
where H=
R2π
0h(z)dz,Hc=
R2π
0h(z)coszd z,a n d Hs=
R2π
0h(z)sinzd z. Positive
values of k1(≈2) are most likely to cause a conference breakdown.
23.12 Write
R1
0|x−y|f(y)dyas
Rx
0(x−y)f(y)dy+
R1
x(y−x)f(y)dy.
23.13 For eigenvalue 0 : f(x)=0f o r|x|<aorf(x) is such that
Ra
−af(y)dy=0 .F o r
eigenvalue 2 a:f(x)=µS(x, a)w i t h µac o n s t a n ta n d S(x, a)≡[H(a+x)−H(x−
a)], where H(z) is the Heaviside step function. Take f(x)=g(x)+cGS(x, a),
where G=
Ra
−ag(z)dz. Show that c=λ/(1−2aλ).
23.14 y(x)=x−3+[ 5x2λln(b/a)]/[5−λ(b5−a5)];|λ|<5/|b5−a5|.
23.15 y(x)=x2−(3I3x+I2)e x p x.
23.16 (a) 5√6/(2√6±5). (b)/angbracketlefty±|K1/angbracketright=A±[(31/60)√2±(19/45)√3]. [1−(5/2)/λ±]−1=
∓2√6/5. For (b) and (c) y(x)=1−4
3x1/2−3
2x.
882
24
Group theory
For systems that have some degree of symmetry, full exploitation of that symmetry
is desirable. Significant physical results can sometimes be deduced simply by astudy of the symmetry properties of the system under investigation. Consequentlyit becomes important, for such a system, to identify all those operations (rotations,reflections, inversions) that carry the system into a physically indistinguishablecopy of itself.
The study of the properties of the complete set of such operations forms
one application of group theory . Though this is the aspect of most interest to
the physical scientist, group theory itself is a much larger subject and of greatimportance in its own right. Consequently we leave until the next chapter any
direct applications of group theoretical results and concentrate on building up
the general mathematical properties of groups.
24.1 Groups
As an example of symmetry properties, let us consider the sets of operations,
such as rotations, reflections, and inversions, that transform physical objects, forexample molecules, into physically indistinguishable copies of themselves, so thatonly the labelling of identical components of the system (the atoms) changes in
the process. For differently shaped molecules there are different sets of operations,
but in each case it is a well-defined set, and with a little practice all members ofeach set can be identified.
As simple examples, consider ( a) the hydrogen molecule, and ( b) the ammonia
molecule illustrated in figure 24.1. The hydrogen molecule consists of two atoms
H of hydrogen and is carried into itself by any of the following operations:
(i) any rotation about its long axis;
(ii) rotation through πabout an axis perpendicular to the long axis and
passing through the point Mthat lies midway between the atoms;
883
GROUP THEORY
(a)( b)H
HH
HHN
M
Figure 24.1 ( a) The hydrogen molecule, and ( b) the ammonia molecule.
(iii) inversion through the point M;
(iv) reflection in the plane that passes through Mand has its normal parallel
to the long axis.
These operations collectively form the set of symmetry operations for the hydro-
gen molecule.
The somewhat more complex ammonia molecule consists of a tetrahedron with
an equilateral triangular base at the three corners of which lie hydrogen atomsH, whilst a nitrogen atom N is sited at the fourth vertex of the tetrahedron. Theset of symmetry operations on this molecule is limited to rotations of π/3a n d
2π/3 about the axis joining the centroid of the equilateral triangle to the nitrogen
atom, and reflections in the three planes containing that axis and each of thehydrogen atoms in turn. However, if the nitrogen atom could be replaced by a
fourth hydrogen atom, and all interatomic distances equalised in the process, the
number of symmetry operations would be greatly increased.
Once allthe possible operations in any particular set have been identified, it
must follow that the result of applying two such operations in succession will beidentical to that obtained by the sole application of some third (usually different)operation in the set – for if it were not, a new member of the set would havebeen found, contradicting the assumption that all members have been identified.
Such observations introduce two of the main considerations relevant to decid-
ing whether a set of objects, here the rotation, reflection and inversion operations,qualifies as a group in the mathematically tightly defined sense. These two consid-
erations are (i) whether there is some law for combining two members of the set,and (ii) whether the result of the combination is also a member of the set. Theobvious rule of combination has to be that the second operation is carried outon the system that results from application of the first operation, and we have
already seen that the second requirement is satisfied by the inclusion of all such
operations in the set. However, for a set to qualify as a group, more than thesetwo conditions have to be satisfied, as will now be made clear.
884
24.1 GROUPS
24.1.1 Definition of a group
Ag r o u p Gis a set of elements {X,Y,...}, together with a rule for combining
them that associates with each ordered pair X,Ya ‘product’ or combination law
X•Yfor which the following conditions must be satisfied.
(i) For everypair of elements X, Y that belongs to G, the product X•Yalso
belongs to G. (This is known as the closure property of the group.)
(ii) For all triples X, Y , Z theassociative law holds; in symbols,
X•(Y•Z)=(X•Y)•Z. (24.1)
(iii) There exists a unique element I,b e l o n g i n gt o G, with the property that
I•X=X=X•I (24.2)
forallXbelonging to G. This element Iis known as the identity element
of the group.
(iv) For every element XofG, there exists an element X−1,a l s ob e l o n g i n gt o
G, such that
X−1•X=I=X•X−1. (24.3)
X−1is called the inverse ofX.
An alternative notation in common use is to write the elements of a group G
as the set{G1,G2,...}or, more briefly, as {Gi}, a typical element being denoted
byGi.
It should be noticed that, as given, the nature of the operation •is not stated. It
should also be noticed that the more general term element , rather than operation ,
has been used in this definition. We will see that the general definition of agroup allows as elements not only sets of operations on an object but also sets of
numbers, of functions and of other objects, provided that the interpretation of •
is appropriately defined.
In one of the simplest examples of a group, namely the group of all integers
under addition, the operation •is taken to be ordinary addition. In this group the
role of the identity Iis played by the integer 0, and the inverse of an integer Xis
−X. That requirements (i) and (ii) are satisfied by the integers under addition is
trivially obvious. A second simple group, under ordinary multiplication, is formedby the two numbers 1 and −1; in this group, closure is obvious, 1 is the identity
element, and each element is its own inverse.
It will be apparent from these two examples that the number of elements in a
group can be either finite or infinite. In the former case the group is called a finite
group and the number of elements it contains is called the order of the group,
which we will denote by g; an alternative notation is |G|but has obvious dangers
885
GROUP THEORY
if matrices are involved. In the notation in which G={G1,G2,...,G n}the order
of the group is clearly n.
As we have noted, for the integers under addition zero is the identity. For
the group of rotations and reflections, the operation of doing nothing, i.e. thenull operation, plays this role. This latter identification may seem artificial, butit is an operation, albeit trivial, which does leave the system in a physicallyindistinguishable state, and needs to be included. One might add that without itthe set of operations would not form a group and none of the powerful resultswe will derive later in this and the next chapter could be justifiably applied to
give deductions of physical significance.
In the examples of rotations and reflections mentioned earlier, •has been taken
to mean that the left-hand operation is carried out on the system that resultsfrom application of the right-hand operation. Thus
Z=X•Y (24.4)
means that the effect on the system of carrying out Zi st h es a m ea sw o u l d
be obtained by first carrying out Yand then carrying out X. The order of the
operations should be noted; it is arbitrary in the first instance but, once chosen,
must be adhered to. The choice we have made is dictated by the fact that mostof our applications involve the effect of rotations and reflections on functions ofspace coordinates, and it is usual, and our practice in the rest of this book, towrite operators acting on functions to the left of the functions.
It will be apparent that for the above-mentioned group, integers under ordinary
addition, it is true that
Y•X=X•Y (24.5)
for all pairs of integers X,Y. If any two particular elements of a group satisfy
(24.5), they are said to commute under the operation •; if all pairs of elements in
a group satisfy (24.5), then the group is said to be Abelian .T h es e to fa l li n t e g e r s
forms an infinite Abelian group under (ordinary) addition.
As we show below, requirements (iii) and (iv) of the definition of a group
are over-demanding (but self-consistent), since in each of equations (24.2) and(24.3) the second equality can be deduced from the first by using the associativityrequired by (24.1). The mathematical steps in the following arguments are allvery simple, but care has to be taken to make sure that nothing that has notyet been proved is used to justify a step. For this reason, and to act as a modelin logical deduction, a reference in Roman numerals to the previous result,
or to the group definition used, is given over each equality sign. Such explicit
detailed referencing soon becomes tiresome, but it should always be available ifneeded.
886
24.1 GROUPSIUsing only the first equalities in (24.2) and (24.3), deduce the second ones.
Consider the expression X−1•(X•X−1);
X−1•(X•X−1)(ii)=(X−1•X)•X−1(iv)=I•X−1
(iii)=X−1. (24.6)
ButX−1belongs to G, and so from (iv) there is an element UinGsuch that
U•X−1=I. (v)
Form the product of Uwith the first and last expressions in (24.6) to give
U•(X−1•(X•X−1)) =U•X−1(v)=I. (24.7)
Transforming the left-hand side of this equation gives
U•(X−1•(X•X−1))(ii)=(U•X−1)•(X•X−1)
(v)=I•(X•X−1)
(iii)=X•X−1. (24.8)
Comparing (24.7), (24.8) shows that
X•X−1=I, (iv)/prime
i.e. the second equality in group definition (iv). Similarly
X•I(iv)=X•(X−1•X)(ii)=(X•X−1)•X
(iv)/prime
=I•X
(iii)=X. (iii/prime)
i.e. the second equality in group definition (iii).
J
The uniqueness of the identity element Ic a na l s ob ed e m o n s t r a t e dr a t h e rt h a n
assumed. Suppose that I/prime,b e l o n g i n gt o G, also has the property
I/prime•X=X=X•I/primefor all Xbelonging to G.
Take XasI,t h e n
I/prime•I=I. (24.9)
Further, from (iii/prime),
X=X•I for all Xbelonging to G,
887
GROUP THEORY
and setting X=I/primegives
I/prime=I/prime•I. (24.10)
It then follows from (24.9), (24.10) that I=I/prime, showing that in any particular
group the identity element is unique.
In a similar way it can be shown that the inverse of any particular element
is unique. If UandVare two postulated inverses of an element XofG,b y
considering the product
U•(X•V)=(U•X)•V,
it can be shown that U=V. The proof is left to the reader.
Given the uniqueness of the inverse of any particular group element, it follows
that
(U•V•···•Y•Z)•(Z−1•Y−1•···•V−1•U−1)
=(U•V•···•Y)•(Z•Z−1)•(Y−1•···•V−1•U−1)
=(U•V•···•Y)•(Y−1•···•V−1•U−1)
...
=I,
where use has been made of the associativity and of the two equations Z•Z−1=I
andI•X=X. Thus the inverse of a product is the product of the inverses in
reverse order, i.e.
(U•V•···•Y•Z)−1=(Z−1•Y−1•···•V−1•U−1). (24.11)
Further elementary results that can be obtained by arguments similar to those
above are as follows.
(i) Given any pair of elements X, Y belonging to G, there exist unique
elements U, V,a l s ob e l o n g i n gt o G, such that
X•U=Y and V•X=Y.
Clearly U=X−1•Y,a n d V=Y•X−1, and they can be shown to be
unique. This result is sometimes called the division axiom .
(ii) The cancellation law can be stated as follows. If
X•Y=X•Z
for some Xbelonging to G,t h e n Y=Z.Similarly,
Y•X=Z•X
implies the same conclusion.
888
24.1 GROUPS
L M
K
Figure 24.2 Reflections in the three perpendicular bisectors of the sides of
an equilateral triangle take the triangle into itself.
(iii) Forming the product of each element of Gwith a fixed element XofG
simply permutes the elements of G; this is often written symbolically as
G•X=G. If this were not so, and X•YandX•Zwere not different
even though YandZwere, application of the cancellation law would lead
to a contradiction. This result is called the permutation law .
In any finite group of order g, any element Xwhen combined with itself to
form successively X2=X•X,X3=X•X2,...will, after at most g−1s u c h
combinations, produce the group identity I. Of course X2,X3,...are some of
the original elements of the group, and not new ones. If the actual number of
combinations needed is m−1, i.e. Xm=I,t h e n mis called the order of the element
XinG. The order of the identity of a group is always 1, and that of any other
element of a group that is its own inverse is always 2.IDetermine the order of the group of (two-dimensional) rotations and reflections that take
a plane equilateral triangle into itself and the order of each of the elements. The group is
usually known as 3m(to physicists and crystallographers) or C3v(to chemists).
There are two (clockwise) rotations, by 2 π/3a n d4 π/3, about an axis perpendicular to
the plane of the triangle. In addition, reflections in the perpendicular bisectors of the threesides (see figure 24.2) have the defining property. To these must be added the identityoperation. Thus in total there are six distinct operations and so g= 6 for this group.
To reproduce the identity operation either of the rotations has to be applied three times,whilst any of the reflections has to be applied just twice in order to recover the originalsituation. Thus each rotation element of the group has order 3, and each reflection elementhas order 2.J
A so-called cyclic group is one for which all members of the group can be
generated from just one element X(say). Thus a cyclic group of order gcan be
written as
G=braceleftbig
I,X,X2,X3,...,Xg−1bracerightbig
.
889
GROUP THEORY
It is clear that cyclic groups are always Abelian and that each element, apart
from the identity, has order g, the order of the group itself.
24.1.2 Further examples of groups
In this section we consider some sets of objects, each set together with a law of
combination, and investigate whether they qualify as groups and, if not, why not.
We have already seen that the integers form a group under ordinary addition,
but it is immediately apparent that (even if zero is excluded) they do notdo
so under ordinary multiplication. Unity must be the identity of the set, but therequisite inverse of any integer n,n a m e l y1 /n, does not belong to the set of
integers for any nother than unity.
Other infinite sets of quantities that do form groups are the sets of all real
numbers, or of all complex numbers, under addition, and of the same two setsexcluding 0 under multiplication. All these groups are Abelian.
Although subtraction and division are normally considered the obvious coun-
terparts of the operations of (ordinary) addition and multiplication, they are notacceptable operations for use within groups since the associative law, (24.1), doesnot hold. Explicitly,
X−(Y−Z)/negationslash=(X−Y)−Z,
X÷(Y÷Z)/negationslash=(X÷Y)÷Z.
From within the field of all non-zero complex numbers we can select just those
that have unit modulus, i.e. are of the form e
iθwhere 0≤θ<2π,t of o r ma
group under multiplication, as can easily be verified:
eiθ1×eiθ2=ei(θ1+θ2)(closure) ,
ei0= 1 (identity) ,
ei(2π−θ)×eiθ=ei2π≡ei0= 1 (inverse) .
Closely related to the above group is the set of 2 ×2 rotation matrices that take
the form
M(θ)=parenleftbiggcosθ−sinθ
sinθcosθparenrightbigg
where, as before, 0 ≤θ<2π. These form a group when the law of combination
is that of matrix multiplication. The reader can easily verify that
M(θ)M(φ)= M(θ+φ) (closure),
M(0) = I2 (identity),
M(2π−θ)= M−1(θ) (inverse).
Here I2is the unit 2 ×2 matrix.
890
24.2 FINITE GROUPS
24.2 Finite groups
Whilst many properties of physical systems (e.g. angular momentum) are related
to the properties of infinite, and, in particular, continuous groups, the symmetryproperties of crystals and molecules are more intimately connected with those offinite groups. We therefore concentrate in this section on finite sets of objects thatcan be combined in a way satisfying the group postulates.
Although it is clear that the set of all integers does not form a group under
ordinary multiplication, restricted sets can do so if the operation involved ismultiplication (mod N) for suitable values of N; this operation will be explained
below.
As a simple example of a group with only four members, consider the set S
defined as follows:
S={1,3,5,7}under multiplication (mod 8) .
To find the product (mod 8) of any two elements, we multiply them together in
the ordinary way, and then divide the answer by 8, treating the remainder after
doing so as the product of the two elements. For example, 5 ×7 = 35, which on
dividing by 8 gives a remainder of 3. Clearly, since Y×Z=Z×Y, the full set
of different products is
1×1=1 ,1×3=3 ,1×5=5 ,1×7=7 ,
3×3=1 ,3×5=7 ,3×7=5 ,
5×5=1 ,5×7=3 ,
7×7=1 .
The first thing to notice is that each multiplication produces a member of the
original set, i.e. the set is closed. Obviously the element 1 takes the role of theidentity, i.e. 1 ×Y=Yfor all members Yof the set. Further, for each element Y
of the set there is an element Z(equal to Y, as it happens, in this case) such that
Y×Z= 1, i.e. each element has an inverse. These observations, together with the
associativity of multiplication (mod 8), show that the set Sis an Abelian group
of order 4.
It is convenient to present the results of combining any two elements of a
group in the form of multiplication tables – akin to those which used to appear inelementary arithmetic books before electronic calculators were invented! Written
in this much more compact form the above example is expressed by table 24.1.
Although the order of the two elements being combined does not matter herebecause the group is Abelian, we adopt the convention that if the product in ageneral multiplication table is written X•Ythen Xis taken from the left-hand
column and Yis taken from the top row. Thus the bold ‘ 7’ in the table is the
result of 3×5, rather than of 5 ×3.
Whilst it would make no difference to the basic information content in a table
891
GROUP THEORY
1357
11357
331 75
55713
77531
Table 24.1 The table of products for the elements of the group S={1,3,5,7}
under multiplication (mod 8).
to present the rows and columns with their headings in random orders, it is
usual to list the elements in the same order in both the vertical and horizontalheadings in any one table. The actual order of the elements in the common list,whilst arbitrary, is normally chosen to make the table have as much symmetry as
possible. This is initially a matter of convenience, but, as we shall see later, some
of the more subtle properties of groups are revealed by putting next to each otherelements of the group that are alike in certain ways.
Some simple general properties of group multiplication tables can be deduced
immediately from the fact that each row or column constitutes the elements ofthe group.
(i) Each element appears once and only once in each row or column of the
table; this must be so since G•X=G(the permutation law) holds.
(ii) The inverse of any element Ycan be found by looking along the row
in which Yappears in the left-hand column (the Yth row), and noting
the element Zat the head of the column (the Zth column) in which
the identity appears as the table entry. An immediate corollary is that
whenever the identity appears on the leading diagonal, it indicates thatthe corresponding header element is of order 2 (unless it happens to bethe identity itself).
(iii) For any Abelian group the multiplication table is symmetric about the
leading diagonal.
To get used to the ideas involved in using group multiplication tables, we now
consider two more sets of integers under multiplication (mod N):
S
/prime={1,5,7,11}under multiplication (mod 24), and
S/prime/prime={1,2,3,4}under multiplication (mod 5) .
These have group multiplication tables 24.2( a)a n d( b) respectively, as the reader
should verify.
If tables 24.1 and 24.2( a) for the groups SandS/primeare compared, it will be seen
that they have essentially the same struc ture, i.e if the elements are written as
{I,A,B,C}in both cases, then the two tables are each equivalent to table 24.3.
ForS,I=1 ,A=3 ,B=5 ,C= 7 and the law of combination is multiplication
892
24.2 FINITE GROUPS
(a)157 1 1
1157 1 1
551 1 1 7
771 11 5
111 1 751(b)1234
11234
22413
33142
44321
Table 24.2 ( a) The multiplication table for the group S/prime={1,5,7,11}under
multiplication (mod 24). ( b) The multiplication table for the group S/prime/prime=
{1,2,3,4}under multiplication (mod 5).
IA B C
IIA B C
AAICB
BBCI A
CCBAI
Table 24.3 The common structure exemplified by tables 24.1 and 24.2( a).
1 i−1−i
1 1 i−1−i
i i−1−i1
−1−1−i1 i
−i−i1 i−1
Table 24.4 The group table for the set {1,i ,−1,−i}under ordinary multipli-
cation of complex numbers.
(mod 8), whilst for S/prime,I=1 , A=5 , B=7 , C= 11 and the law of combination
is multiplication (mod 24). However, the really important point is that the twogroups SandS
/primehave equivalent group multiplication tables – they are said to
beisomorphic , a matter to which we will return more formally in section 24.5.IDetermine the behaviour of the set of four elements
{1,i ,−1,−i}
under the ordinary multiplication of complex numbers. Show that they form a group and
determine whether the group is isomorphic to either of the groups S(itself isomorphic to
S/prime) andS/prime/primedefined above.
That the elements form a group under the associative operation of complex multiplication
is immediate; there is an identity (1), each possible product generates a member of the setand each element has an inverse (1 ,−i,−1,i, respectively). The group table has the form
shown in table 24.4.
We now ask whether this table can be made to look like table 24.3, which is the
standardised form of the tables for SandS
/prime. Since the identity element of the group (1)
will have to be represented by I, and ‘1’ only appears on the leading diagonal twice whereas
893
GROUP THEORY
1 i−1−i
1 1 i−1−i
i i−1−i1
−1−1−i1 i
−i−i1 i−11243
11243
22431
44312
33124
Table 24.5 A comparison between tables 24.4 and 24.2( b), the latter with its
columns reordered.
IA B C
IIA B C
AABC I
BBCI A
CCIA B
Table 24.6 The common structure exemplified by tables 24.4 and 24.2( b), the
latter with its columns reordered.
Iappears on the leading diagonal four times in table 24.3, it is clear that no amount of
relabelling (or, equivalently, no allocation of the symbols A, B, C ,a m o n g s t i,−1,−i)c a n
bring table 24.4 into the form of table 24.3. We conclude that the group {1,i ,−1,−i}is
not isomorphic to SorS/prime. An alternative way of stating the observation is to say that
the group contains only one element of order 2 whilst a group corresponding to table 24.3contains three such elements. However, if the rows and columns of table 24.2( b) – in which
the identity does appear twice on the diagonal and which therefore has the potential to beequivalent to table 24.4 – are rearranged by making the heading order 1 ,2,4,3 then the
two tables can be compared in the forms sh own in table 24.5. They can thus be seen to
have the same structure, namely that shown in table 24.6.
We therefore conclude that the group of four elements {1,i ,−1,−i}under ordinary mul-
tiplication of complex numbers is isomorphic to the group {1,2,3,4}under multiplication
(mod 5).J
What we have done does not prove it, but the two tables 24.3 and 24.6 are in
fact the only possible tables for a group of order 4, i.e. a group containing exactlyfour elements.
24.3 Non-Abelian groups
So far, all the groups for which we have constructed multiplication tables have
been based on some form of arithmetic multiplication, a commutative operation,with the result that the groups have been Abelian and the tables symmetricabout the leading diagonal. We now turn to examples of groups in which some
non-commutation occurs. It should be noted, in passing, that non-commutation
cannot occur throughout a group, as the identity always commutes with any
element in its group.
894
24.3 NON-ABELIAN GROUPS
As the first example we consider again as elements of a group the two-
dimensional operations which transform an equilateral triangle into itself (seethe end of subsection 24.1.1). It has already been shown that there are sixsuch operations; the null operation, two rotations (by 2 π/3a n d4 π/3 about
an axis perpendicular to the plane of the triangle) and three reflections in the
perpendicular bisectors of the three sides. To abbreviate we will denote theseoperations by symbols as follows.
(i)Iis the null operation.
(ii)RandR
/primeare (clockwise) rotations by 2 π/3a n d4 π/3 respectively.
(iii)K,L,Mare reflections in the three lines indicated in figure 24.2.
Some products of the operations of the form X•Y(where it will be recalled
that the symbol •means that the second operation Xis carried out on the system
resulting from the application of the first operation Y) are easily calculated:
R•R=R/prime,R/prime•R/prime=R, R•R/prime=I=R/prime•R(24.12)
K•K=L•L=M•M=I.
Others, such as K•M, are more difficult, but can be found by a little thought,
or by making a model triangle or drawing a sequence of diagrams such as thosefollowing.
xx x
xK•M = = =K R/prime
showing that K•M=R/prime. In the same way,
x x x xM•K = = =M R
shows that M•K=R,a n d
x x xxR•L = = =R K
shows that R•L=K.
Proceeding in this way we can build up the complete multiplication table
(table 24.7). In fact, it is not necessary to draw any more diagrams, as allremaining products can be deduced algebraically from the three found above and
895
GROUP THEORY
IR R/primeKL M
I IR R/primeKL M
R RR/primeIM KL
R/primeR/primeIRL M K
K KL MI RR/prime
L LMKR/primeIR
M MK L RR/primeI
Table 24.7 The group table for the two-dimensional symmetry operations onan equilateral triangle.
the more self-evident results given in (24.12). A number of things may be noticed
about this table.
(i) It is notsymmetric about the leading diagonal, indicating that some pairs
of elements in the group do not commute.
(ii) There is some symmetry within the 3 ×3 blocks that form the four quarters
of the table. This occurs because we have elected to put similar operations
close to each other when choosing the order of table headings – the tworotations (or three if Iis viewed as a rotation by 0 π/3) are next to each
other, and the three reflections also occupy adjacent columns and rows.We will return to this later.
That two groups of the same order may be isomorphic carries over to non-
Abelian groups. The next two examples are each concerned with sets of sixobjects; they will be shown to form groups that, although very different in naturefrom the rotation–reflection group just considered, are isomorphic to it.
We consider first the set Mof six orthogonal 2 ×2 matrices given by
I=parenleftbigg10
01parenrightbigg
A=parenleftBigg
−
1
2√
3
2
−√
3
2−1
2parenrightBigg
B=parenleftBigg
−1
2−√
3
2√
3
2−1
2parenrightBigg
C=parenleftbigg−10
01parenrightbigg
D=parenleftBigg1
2−√
3
2
−√
3
2−1
2parenrightBigg
E=parenleftBigg1
2√
3
2√
3
2−1
2parenrightBigg(24.13)
the combination law being that of ordinary matrix multiplication. Here we use
italic, rather than the sans serif used for matrices elsewhere, to emphasise thatthe matrices are group elements.
Although it is tedious to do so, it can be checked that the product of any
two of these matrices, in either order, is also in the set. However, the result isgenerally different in the two cases, as matrix multiplication is non-commutative.
The matrix Iclearly acts as the identity element of the set, and during the checking
for closure it is found that the inverse of each matrix is contained in the set, I,
C,DandEbeing their own inverses. The group table is shown in table 24.8.
896
24.3 NON-ABELIAN GROUPS
IA B C D E
IIA B C D E
AAB I ECD
BBIA D EC
CCDEI AB
DDECB I A
EECDABI
Table 24.8 The group table, under matrix multiplication, for the set Mof
six orthogonal 2 ×2 matrices given by (24.13).
The similarity to table 24.7 is striking. If {R,R/prime,K,L ,M}of that table are
replaced by {A, B, C, D, E }respectively, the two tables are identical, without even
the need to reshuffle the rows and columns. The two groups, one of reflectionsand rotations of an equilateral triangle, the other of matrices, are isomorphic.
Our second example of a group isomorphic to the same rotation–reflection
group is provided by a set of functions of an undetermined variable x.T h e
functions are as follows:
f
1(x)=x, f 2(x)=1 /(1−x),f 3(x)=(x−1)/x,
f4(x)=1 /x, f 5(x)=1−x, f 6(x)=x/(x−1),
and the law of combination is
fi(x)•fj(x)=fi(fj(x)),
i.e. the function on the right acts as the argument of the function on the left to
produce a new function of x. It should be emphasised that it is the functions
that are the elements of the group. The variable xis the ‘system’ on which they
act, and plays much the same role as the triangle does in our first example of a
non-Abelian group.
To show an explicit example, we calculate the product f6•f3. The product
will be the function of xobtained by evaluating y/(y−1), when yis set equal to
(x−1)/x. Explicitly
f6(f3)=(x−1)/x
(x−1)/x−1=1−x=f5(x).
Thus f6•f3=f5. Further examples are
f2•f2=1
1−1/(1−x)=x−1
x=f3,
and
f6•f6=x/(x−1)
x/(x−1)−1=x=f1. (24.14)
897
GROUP THEORY
The multiplication table for this set of six functions has all the necessary proper-
ties to show that they form a group. Further, if the symbols f1,f2,f3,f4,f5,f6are
replaced by I,A,B,C,D,E respectively the table becomes identical to table 24.8.
This justifies our earlier claim that this group of functions, with argument sub-
stitution as the law of combination, is isomorphic to the group of reflections and
rotations of an equilateral triangle.
24.4 Permutation groups
The operation of rearranging ndistinct objects amongst themselves is called a
permutation of degree n, and since many symmetry operations on physical systems
can be viewed in that light, the properties of permutations are of interest. For
example, the symmetry operations on an equilateral triangle, to which we have
already given much attention, can be considered as the six possible rearrangementsof the marked corners of the triangle amongst three fixed points in space, muchas in the diagrams used to compute table 24.7. In the same way, the symmetryoperations on a cube can be viewed as a rearrangement of its corners amongsteight points in space, albeit with many constraints, or, with fewer complications,as a rearrangement of its body diagonals in space. The details will be left until
we review the possible finite groups more systematically.
The notations and conventions used in the literature to describe permutations
are very varied and can easily lead to confusion. We will try to avoid this by usingletters a ,b ,c ,... (rather than numbers) for the objects that are rearranged by a
permutation and by adopting, before long, a ‘cycle notation’ for the permutationsthemselves. It is worth emphasising that it is the permutations ,i . e .t h ea c t so f
rearranging, and not the objects themselves (represented by letters) that form
the elements of permutation groups. The complete group of all permutations of
degree nis usually denoted by S
nor Σ n. The number of possible permutations of
degree nisn!, and so this is the order of Sn.
Suppose the ordered set of six distinct objects {abcdef}is rearranged by
some process into {befadc}; then we can represent this mathematically as
θ{abcdef}={befadc},
where θis a permutation of degree 6. The permutation θcan be denoted by
[ 256143 ] ,s i n c et h efi r s to b j e c t , a, is replaced by the second, b, the second
object, b, is replaced by the fifth, e, the third by the sixth, f,e t c .T h ee q u a t i o n
can then be written more explicitly as
θ{abcdef}=[ 256143 ] {abcdef}={befadc}.
Ifφis a second permutation, also of degree 6, then the obvious interpretation of
the product φ•θof the two permutations is
φ•θ{abcdef}=φ(θ{abcdef}).
898
24.4 PERMUTATION GROUPS
Suppose that φis the permutation [4 5 3 6 2 1]; then
φ•θ{abcdef}=[ 453621 ] [ 256143 ] {abcdef}
=[ 453621 ] {befadc}
={adfceb}
=[ 146352 ] {abcdef}.
Written in terms of the permutation notation this result is
[ 453621 ] [ 256143 ]=[ 146352 ] .
A concept that is very useful for working with permutations is that of decom-
position into cycles. The cycle notation is most easily explained by example. For
the permutation θgiven above:
the 1st object, a, has been replaced by the 2nd, b;
the 2nd object, b, has been replaced by the 5th, e;
the 5th object, e, has been replaced by the 4th, d;
the 4th object, d, has been replaced by the 1st, a.
This brings us back to the beginning of a closed cycle, which is conveniently
represented by the notation (1 2 5 4), in which the successive replacementp o s i t i o n sa r ee n c l o s e d ,i ns e q u e n c e ,i np a r e n t h e s e s .T h u s( 1254 )m e a n s2 n d
→1st, 5th→2nd, 4th→5th, 1st→4th. It should be noted that the object
initially in the first listed position replaces that in the final position indicated inthe bracket – here ‘ a’ is put into the fourth position by the permutation. Clearly
the cycle (5 4 1 2), or any other which involved the same numbers in the samerelative order, would have exactly the same meaning and effect. The remainingtwo objects, candf, are interchanged by θor, more formally, are rearranged
according to a cycle of length 2, a transposition , represented by (3 6). Thus the
complete representation (specification) of θis
θ=( 1254 ) ( 36 ) .
The positions of objects that are unaltered by a permutation are either placed by
themselves in a pair of parentheses or omitted altogether. The former is recom-mended as it helps to indicate how many objects are involved – important whenthe object in the last position is unchanged, or the permutation is the identity,which leaves all objects unaltered in position! Thus the identity permutation of
degree 6 is
I= (1)(2)(3)(4)(5)(6) ,
though in practice it is often shortened to (1).
It will be clear that the cycle representation is unique, to within the internal
899
GROUP THEORY
absolute ordering of the numbers in each bracket as already noted, and that
each number appears once and only once in the representation of any particularpermutation.
Theorder of any permutation of degree nwithin the group S
ncan be read off
from the cyclic representation and is given by the lowest common multiple (LCM)
of the lengths of the cycles. Thus Ihas order 1, as it must, and the permutation
θdiscussed above has order 4 (the LCM of 4 and 2).
Expressed in cycle notation our second permutation φis (3)(1 4 6)(2 5), and
the product φ•θis calculated as
(3)(1 4 6)(2 5) •(1 2 5 4)(3 6) {abcdef}= (3)(1 4 6)(2 5) {befadc}
={adfceb}
= (1)(5)(2 4 3 6) {abcdef}.
i.e. expressed as a relationship amongst elements of the group of permutations of
degree 6 (not yet proved as a group, but reasonably anticipated), this result reads
(3)(1 4 6)(2 5) •(1 2 5 4)(3 6) = (1)(5)(2 4 3 6) .
We note, for practice, that φhas order 6 (the LCM of 1, 3, and 2) and that the
product φ•θhas order 4.
The number of elements in the group Snof all permutations of degree nis
n! and clearly increases very rapidly as nincreases. Fortunately, to illustrate the
essential features of permutation groups it is sufficient to consider the case n=3 ,
which involves only six elements. They are as follows (with labelling which thereader will by now recognise as anticipatory):
I= (1)(2)(3) A=( 123 ) B=( 132 )
C= (1)(2 3) D= (3)(1 2) E= (2)(1 3)
It will be noted that AandBhave order 3, whilst C,DandEhave order 2. As
perhaps anticipated, their combination products are exactly those correspondingto table 24.8, I,C,DandEbeing their own inverses. For example, putting in all
steps explicitly,
D•C{abc}= (3)(1 2)•(1)(2 3){abc}
= (3)(12){acb}
={cab}
=( 321 ){abc}
=( 132 ){abc}
=B{abc}.
In brief, the six permutations belonging to S
3form yet another non-Abelian group
isomorphic to the rotation–reflection symmetry group of an equilateral triangle.
900
24.5 MAPPINGS BETWEEN GROUPS
24.5 Mappings between groups
Now that we have available a range of groups that can be used as examples,
we return to the study of more general group properties. From here on, when
there is no ambiguity we will write the product of two elements, X•Y, simply
asXY, omitting the explicit combination symbol. We will also continue to use
‘multiplication’ as a loose generic name for the combination process betweenelements of a group.
IfGandG
/primeare two groups, we can study the effect of a mapping
Φ:G→G/prime
ofGontoG/prime.I fXis an element of Gwe denote its image inG/primeunder the mapping
Φb y X/prime=Φ ( X).
A technical term that we have already used is isomorphic . We will now define
it formally. Two groups G={X,Y,...}andG/prime={X/prime,Y/prime,...}are said to be
isomorphic if there is a one-to-one correspondence
X↔X/prime,Y↔Y/prime,···
between their elements such that
XY=Z implies X/primeY/prime=Z/prime
and vice versa.
In other words, isomorphic groups have the same (multiplication) structure,
although they may differ in the nature of their elements, combination law andnotation. Clearly if groups GandG
/primeare isomorphic, and GandG/prime/primeare isomorphic,
then it follows that G/primeandG/prime/primeare isomorphic. We have already seen an example of
four groups (of functions of x, of orthogonal matrices, of permutations and of the
symmetries of an equilateral triangle) that are isomorphic, all having table 24.8
as their multiplication table.
Although our main interest is in isomorphic relationships between groups, the
wider question of mappings of one set of elements onto another is of someimportance, and we start with the more general notion of a homomorphism.
LetGandG
/primebe two groups and Φa mapping of G→G/prime. If for every pair of
elements XandYinG
(XY)/prime=X/primeY/prime
thenΦis called a homomorphism, and G/primeis said to be a homomorphic image of G.
The essential defining relationship, expressed by ( XY)/prime=X/primeY/prime, is that the
same result is obtained whether the product of two elements is formed first andthe image then taken or the images are taken first and the product then formed.
Three immediate consequences of the above definition are proved as follows.
901
GROUP THEORY
(i) If Iis the identity of Gthen IX=Xfor all XinG.C o n s e q u e n t l y
X/prime=(IX)/prime=I/primeX/prime,
for all X/primeinG/prime. Thus I/primeis the identity in G/prime. In words, the identity element
ofGmaps into the identity element of G/prime.
(ii) Further,
I/prime=(XX−1)/prime=X/prime(X−1)/prime.
That is, ( X−1)/prime=(X/prime)−1.In words, the image of an inverse is the same
element in G/primeas the inverse of the image.
(iii) If element XinGis of order m,i . e .I=Xm,t h e n
I/prime=(Xm)/prime=(XXm−1)/prime=X/prime(Xm−1)/prime=···=X/primeX/prime···X/prime
bracehtipupleftbracehtipdownrightbracehtipdownleftbracehtipupright
mfactors.
In words, the image of an element has the same order as the element.
What distinguishes an isomorphism from the more general homomorphism are
the requirements that in an isomorphism:
(I) different elements in Gmust map into different elements in G/prime(whereas in
a homomorphism several elements in Gmay have the same image in G/prime),
that is, x/prime=y/primemust imply x=y;
(II) any element in G/primemust be the image of some element in G.
An immediate consequence of (I) and result (iii) for homomorphisms is that
groups that are isomorphic each have the same number of elements of any given
order.
For a general homomorphism, the set of elements of Gwhose image in G/prime
isI/primeis called the kernel of the homomorphism; this is discussed further in the
next section. In an isomorphism the kernel consists of the identity Ialone. To
illustrate both this point and the general notion of a homomorphism, considera mapping between the additive group of real numbers /Rfracturand the multiplicative
group of complex numbers with unit modulus, U(1). Suppose that the mapping
/Rfractur→ U(1) is
Φ:x→e
ix;
then this is a homomorphism since
(x+y)/prime→ei(x+y)=eixeiy=x/primey/prime.
However, it is not an isomorphism because many (an infinite number) of the
elements of /Rfracturhave the same image in U(1). For example, π,3π,5π,... in/Rfracturall
have the image −1i nU(1) and, furthermore, all elements of /Rfracturof the form 2 πn,
where nis an integer, map onto the identity element in U(1). The latter set forms
the kernel of the homomorphism.
902
24.6 SUBGROUPS
(a)IA B C D E
IIAB CDE
AABI ECD
BBIA DEC
CCDEI AB
DDECB I A
EECDABI(b)IA B C
IIA BC
AAI CB
BBCI A
CCBAI
Table 24.9 Reproduction of ( a) table 24.8 and ( b) table 24.3 with the relevant
subgroups shown in bold.
For the sake of completeness, we add that a homomorphism for which (I) above
holds is said to be a monomorphism (or an isomorphism into), whilst a homomor-
phism for which (II) holds is called an epimorphism (or an isomorphism onto). If,
in either case, the other requirement is met as well then the monomorphism orepimorphism is also an isomorphism.
Finally, if the initial and final groups are the same, G=G
/prime, then the isomorphism
G→G/primeis termed an automorphism .
24.6 Subgroups
More detailed inspection of tables 24.8 and 24.3 shows that not only do the
complete tables have the properties associated with a group multiplication table(see section 24.2) but so do the upper left corners of each table taken on theirown. The relevant parts are shown in bold in the tables 24.9( a)a n d( b).
This observation immediately prompts the notion of a subgroup . A subgroup
of a group Gcan be formally defined as any non-empty subset H={H
i}of
G, the elements of which themselves behave as a group under the same rule of
combination as applies in Gitself. As for all groups, the order of the subgroup is
equal to the number of elements it contains; we will denote it by hor|H|.
All groups Gcontain two trivial subgroups:
(i)Gitself;
(ii) the set Iconsisting of the identity element alone.
All other subgroups are termed proper subgroups . In a group with multiplication
table 24.8 the elements {I,A,B}form a proper subgroup, as do {I,A}in a group
with table 24.3 as its group table.
Some groups have no proper subgroups. For example, the so-called cyclic
groups , mentioned at the end of subsection 24.1.1, have no subgroups other
than the whole group or the identity alone. Tables 24.10( a)a n d( b) show the
multiplication tables for two of these groups. Table 24.6 is also the group tablefor a cyclic group, that of order 4.
903
GROUP THEORY
(a)IA B
IIA B
AAB I
BBIA(b)IA B C D
IIA B C D
AABCDI
BBCDI A
CCDI AB
DDIABC
Table 24.10 The group tables of two cyclic groups, of orders 3 and 5. They
have no proper subgroups.
It will be clear that for a cyclic group Grepeated combination of any element
with itself generates all other elements of G, before finally reproducing itself. So,
for example, in table 24.10( b), starting with (say) D, repeated combination with
itself produces, in turn, C,B,A,Iand finally Dagain. As noted earlier, in any
cyclic group Gevery element, apart from the identity, is of order g, the order of
the group itself.
The two tables shown are for groups of orders 3 and 5. It will be proved in
subsection 24.7.2 that the order of any group is a multiple of the order of any ofits subgroups (Lagrange’s theorem), i.e. in our general notation, gis a multiple
ofh. It thus follows that a group of order p,w h e r e pis any prime, must be cyclic
and cannot have any proper subgroups. The groups for which tables 24.10( a)a n d
(b) are the group tables are two such examples. Groups of non-prime order may
(table 24.3) or may not (table 24.6) have proper subgroups.
As we have seen, repeated multiplication of an element X(not the identity)
by itself will generate a subgroup {X,X
2,X3,...}. The subgroup will clearly be
Abelian, and if Xis of order m,i . e . Xm=I, the subgroup will have mdistinct
members. If mis less than g– though, in view of Lagrange’s theorem, mmust
be a factor of g– the subgroup will be a proper subgroup. We can deduce, in
passing, that the order of any element of a group is an exact divisor of the orderof the group.
Some obvious properties of the subgroups of a group G, which can be listed
without formal proof, are as follows.
(i) The identity element of Gbelongs to every subgroup H.
(ii) If element Xbelongs to a subgroup H, so does X
−1.
(iii) The set of elements in Gthat belong to every subgroup of Gthemselves
form a subgroup, though it may consist of the identity alone.
Properties of subgroups that need more explicit proof are given in the follow-
ing sections, though some need the development of new concepts before they
can be established. However, we can begin with a theorem, applicable to allhomomorphisms, not just isomorphisms, that requires no new concepts.
Let Φ : G→G
/primebe a homomorphism of GintoG/prime;t h e n
904
24.7 SUBDIVIDING A GROUP
(i) the set of elements H/primeinG/primethat are images of the elements of Gforms a
subgroup of G/prime;
(ii) the set of elements KinGthat are mapped onto the identity I/primeinG/primeforms
a subgroup of G.
As indicated in the previous section, the subgroup Kis called the kernel of the
homomorphism.
To prove (i), suppose ZandWbelong to H/prime, with Z=X/primeandW=Y/prime,w h e r e
XandYbelong to G.T h e n
ZW=X/primeY/prime=(XY)/prime
and therefore belongs to H/prime,a n d
Z−1=(X/prime)−1=(X−1)/prime
and therefore belongs to H/prime. These two results, together with the fact that I/prime
belongs to H/prime, are enough to establish result (i).
To prove (ii), suppose XandYbelong to K;t h e n
(XY)/prime=X/primeY/prime=I/primeI/prime=I/prime(closure),
I/prime=(XX−1)/prime=X/prime(X−1)/prime=I/prime(X−1)/prime=(X−1)/prime
and therefore X−1belongs to K. These two results, together with the fact that I
belongs to K, are enough to establish (ii). An illustration of this result is provided
by the mapping Φ of /Rfractur→ U(1) considered in the previous section. Its kernel
consists of the set of real numbers of the form 2 πnwhere nis an integer; they
form a subgroup of R, the additive group of real numbers.
In fact the kernel Kof a homomorphism is a normal subgroup of G.T h e
defining property of such a subgroup is that for every element XinGand every
element Yin the subgroup, XY X−1belongs to the subgroup. This property is
easily verified for the kernel K,s i n c e
(XY X−1)/prime=X/primeY/prime(X−1)/prime=X/primeI/prime(X−1)/prime=X/prime(X−1)/prime=I/prime.
Anticipating the discussion of subsection 24.7.2, the cosets of a normal subgroup
themselves form a group (see exercise 24.16).
24.7 Subdividing a group
We have already noted, when looking at the (arbitrary) order of headings in a
group table, that some choices appear to make the table more orderly than do
others. In the following subsections we will identify ways in which the elements
of a group can be divided up into sets with the property that the members of anyone set are more like the other members of the set, in some particular regard,
905
GROUP THEORY
than they are like any element that does not belong to the set. We will find that
these divisions will be such that the group is partitioned , i.e. the elements will be
divided into sets in such a way that each element of the group belongs to one,and only one, such set.
We note in passing that the subgroups of a group do notform such a partition,
not least because the identity element is in every subgroup, rather than being inprecisely one. In other words, despite the nomenclature, a group is not simply theaggregate of its proper subgroups.
24.7.1 Equivalence relations and classes
We now specify in a more mathematical manner what it means for two elements
of a group to be ‘more like’ one another than like a third element, as mentionedin section 24.2. Our introduction will apply to any set, whether a group or not,
but our main interest will ultimately be in two particular applications to groups.
We start with the formal definition of an equivalence relation.
Anequivalence relation on a set Sis a relationship X∼Y,b e t w e e nt w o
elements XandYbelonging to S, in which the definition of the symbol ∼must
satisfy the requirements of
(i) reflexivity, X∼X;
(ii) symmetry, X∼Yimplies Y∼X;
(iii) transitivity, X∼YandY∼Zimply X∼Z.
Any particular two elements either satisfy or do not satisfy the relationship.
The general notion of an equivalence relation is very straightforward, and the
requirements on ∼seem undemanding; but not all relationships qualify. As an
example within the topic of groups, if ∼meant ‘has the same order as’ then
clearly all the requirements would be satisfied. However, if ∼meant ‘commutes
with’ then it would not be an equivalence relation, since although Acommutes
with I,a n d Icommutes with C, this does not necessarily imply that Acommutes
with C, as is obvious from table 24.8.
It may be shown that an equivalence relation on Sdivides up Sintoclasses C
i
such that:
(i)XandYbelong to the same class if, and only if, X∼Y;
(ii) every element WofSbelongs to exactly one class.
This may be shown as follows. Let Xbelong to S, and define the subset SXof
Sto be the set of all elements UofSsuch that X∼U. Clearly by reflexivity
Xbelongs to SX. Suppose first that X∼Y, and let Zbe any element of SY.
Then Y∼Z, and hence by transitivity X∼Z, which means that Zbelongs to
SX. Conversely, since the symmetry law gives Y∼X,i fZbelongs to SXthen
906
24.7 SUBDIVIDING A GROUP
this implies that Zbelongs to SY. These two results together mean that the two
subsets SXandSYhave the same members and hence are equal.
Now suppose that SXequals SY.S i n c e Ybelongs to SYit also belongs to SX
and hence X∼Y. This completes the proof of (i), once the distinct subsets of
typeSXare identified as the classes Ci. Statement (ii) is an immediate corollary,
the class in question being identified as SW.
The most important property of an equivalence relation is as follows.
Two different subsets SXandSYcan have no element in common, and the collection
of all the classes Ciis a ‘partition’ of S, i.e. every element in Sbelongs to one, and
only one, of the classes.
To prove this, suppose SXandSYhave an element Zin common; then X∼Z
andY∼Zand so by the symmetry and transitivity laws X∼Y.B yt h ea b o v e
theorem this implies SXequals SY. But this contradicts the fact that SXandSY
are different subsets. Hence SXandSYcan have no element in common.
Finally, if the elements of Sare used in turn to define subsets and hence classes
inS, every element Uis in the subset SUthat is either a class already found or
constitutes a new one. It follows that the classes exhaust S,i . e .e v e r ye l e m e n ti s
in some class.
Having established the general properties of equivalence relations, we now turn
to two specific examples of such relationships, in which the general set Shas the
more specialised properties of a group Gand the equivalence relation ∼is chosen
in such a way that the relatively transparent general results for equivalencerelations can be used to derive powerful, but less obvious, results about theproperties of groups.
24.7.2 Congruence and cosets
As the first application of equivalence relations we now prove Lagrange’s theorem
which is stated as follows.
IfGis a finite group of order gandHis a subgroup of Gof order h
then gis a multiple of h.
We take as the definition of ∼that, given XandYbelonging to G,X∼Yif
X
−1Ybelongs to H. This is the same as saying that Y=XH ifor some element
Hibelonging to H; technically XandYare said to be left-congruent with respect
toH.
This defines an equivalence relation, since it has the following properties.
(i) Reflexivity: X∼X,s i n c e X−1X=IandIbelongs to any subgroup.
(ii) Symmetry: X∼Yimplies that X−1Ybelongs to Hand so, therefore, does
its inverse, since His a group. But ( X−1Y)−1=Y−1Xand, as this belongs
toH, it follows that Y∼X.
907
GROUP THEORY
(iii) Transitivity: X∼YandY∼Zimply that X−1YandY−1Zbelong to H
and so, therefore, does their product ( X−1Y)(Y−1Z)=X−1Z,f r o mw h i c h
it follows that X∼Z.
With∼proved as an equivalence relation, we can immediately deduce that it
divides Ginto disjoint (non-overlapping) classes. For this particular equivalence
relation the classes are called the left cosets ofH. Thus each element of Gis in
one and only one left coset of H. The left coset containing any particular Xis
usually written XH, and denotes the set of elements of the form XH i(one of
which is Xitself since Hcontains the identity element); it must contain hdifferent
elements, since if it did not, and two elements were equal,
XH i=XH j,
we could deduce that Hi=Hjand that Hcontained fewer than helements.
From our general results about equivalence relations it now follows that the
left cosets of Hare a ‘partition’ of Ginto a number of sets each containing h
members. Since there are gmembers of Gand each must be in just one of the
sets, it follows that gis a multiple of h. This concludes the proof of Lagrange’s
theorem.
The number of left cosets of HinGis known as the index ofHinGand is
written [ G:H]; numerically the index = g/h. For the record we note that, for
the trivial subgroup I, which contains only the identity element, [ G:I]=gand
that, for a subgroup Jof subgroup H,[G:H][H:J]=[G:J].
The validity of Lagrange’s theorem was established above using the far-reaching
properties of equivalence relations. However, for this specific purpose there is amore direct and self-contained proof, which we now give.
LetXbe some particular element of a finite group Gof order g,a n dHbe a
subgroup of Gof order h, with typical element Y
i. Consider the set of elements
XH≡{XY1,XY 2,...,XY h}.
This set contains hdistinct elements, since if any two were equal, i.e. XYi=XYj
with i/negationslash=j, this would contradict the cancellation law. As we have already seen,
the set is called a left coset of H.
We now prove three simple results.
•Two cosets are either disjoint or identical. Suppose cosets X1HandX2Hhave
an element in common, i.e. X1Y1=X2Y2for some Y1,Y2inH.T h e n X1=
X2Y2Y−1
1, and since Y1andY2both belong to Hso does Y2Y−1
1; thus X1
belongs to the left coset X2H. Similarly X2belongs to the left coset X1H.
Consequently, either the two cosets are identical or it was wrong to assumethat they have an element in common.
908
24.7 SUBDIVIDING A GROUP
•Two cosets X1HandX2Hare identical if, and only if, X−1
2X1belongs to H.If
X−1
2X1belongs to Hthen X1=X2Yifor some i,a n d
X1H=X2YiH=X2H,
since by the permutation law YiH=H. Thus the two cosets are identical.
Conversely, suppose X1H=X2H.T h e n X−1
2X1H=H. But one element of
H(on the left of the equation) is I; thus X−1
2X1must also be an element of H
(on the right). This proves the stated result.
•Every element of Gis in some left coset XH.This follows trivially since H
contains I, and so the element Xiis in the coset XiH.
The final step in establishing Lagrange’s theorem is, as previously, to note that
each coset contains helements, that the cosets are disjoint and that every one of
thegelements in Gappears in one and only one distinct coset. It follows that
g=khfor some integer k.
As noted earlier, Lagrange’s theorem justifies our statement that any group of
order p,w h e r e pis prime, must be cyclic and cannot have any proper subgroups:
since any subgroup must have an order that divides p, this can only be 1 or p,
corresponding to the two trivial subgroups Iand the whole group.
It may be helpful to see an example worked through explicitly, and we again
use the same six-element group.IFind the left cosets of the proper subgroup Hof the group Gthat has table 24.8 as its
multiplication table.
The subgroup consists of the set of elements H={I,A,B}. We note in passing that it has
order 3, which, as required by Lagrange’s theorem, is a divisor of 6, the order of G.A si n
all cases, Hitself provides the first (left) coset, formally the coset
IH={II,IA,IB}={I,A,B}.
We continue by choosing an element not already selected, Csay, and form
CH={CI,CA,CB}={C,D,E}.
These two cosets of Hexhaust G, and are therefore the only cosets, the index of HinG
being equal to 2.
This completes the example, but it is useful to demonstrate that it would not have
mattered if we had taken D, say, instead of Ito form a first coset
DH={DI, DA, DB}={D, E,C},
and then, from previously unselected elements, picked B, say:
BH={BI,BA,BB}={B,I,A}.
The same two cosets would have resulted.
J
It will be noticed that the cosets are the same groupings of the elements
ofGwhich we earlier noted as being the choice of adjacent column and row
headings that give the multiplication table its ‘neatest’ appearance. Furthermore,
909
GROUP THEORY
ifHis anormal subgroup of Gthen its (left) cosets themselves form a group (see
exercise 24.16).
24.7.3 Conjugates and classes
Our second example of an equivalence relation is concerned with those elements
XandYof a group Gthat can be connected by a transformation of the form
Y=G−1
iXG i,w h e r e Giis an (appropriate) element of G. Thus X∼Yif there
exists an element GiofGsuch that Y=G−1
iXG i. Different pairs of elements X
andYwill, in general, require different group elements Gi. Elements connected
in this way are said to be conjugates .
We first need to establish that this does indeed define an equivalence relation,
as follows.
(i) Reflexivity: X∼X,s i n c e X=I−1XIandIbelongs to the group.
(ii) Symmetry: X∼Yimplies Y=G−1
iXG iand therefore X=(G−1
i)−1YG−1
i.
Since Gibelongs to Gso does G−1
i, and it follows that Y∼X.
(iii) Transitivity: X∼YandY∼Zimply Y=G−1
iXG iandZ=G−1
jYG j
and therefore Z=G−1
jG−1
iXG iGj=(GiGj)−1X(GiGj). Since GiandGj
belong to Gso does GiGj, from which it follows that X∼Z.
These results establish conjugacy as an equivalence relation and hence show
that it divides Ginto classes, two elements being in the same class if, and only if,
they are conjugate.
Immediate corollaries are:
(i) If Zis in the class containing Ithen
Z=G−1
iIGi=G−1
iGi=I.
Thus, since any conjugate of Ican be shown to be I, the identity must be
in a class by itself.
(ii) If Xis in a class by itself then
Y=G−1
iXG i
must imply that Y=X.B u t
X=GiG−1
iXG iG−1
i
for any Gi,a n ds o
X=Gi(G−1
iXG i)G−1
i=GiYG−1
i=GiXG−1
i,
i.e.XG i=GiXfor all Gi.
Thus commutation with all elements of the group is a necessary (and
sufficient) condition for any particular group element to be in a class byitself. In an Abelian group each element is in a class by itself.
910
24.7 SUBDIVIDING A GROUP
(iii) In any group Gthe set Sof elements in classes by themselves is an Abelian
subgroup (known as the centre ofG). We have shown that Ibelongs to S,
and so if, further, XG i=GiXandYG i=GiYfor all Gibelonging to G
then:
(a) (XY)Gi=XG iY=Gi(XY),i.e. the closure of S,a n d
(b)XG i=GiXimplies X−1Gi=GiX−1, i.e. the inverse of Xbelongs
toS.
Hence Sis a group, and clearly Abelian.
Yet again for illustration purposes, we use the six-element group that has
table 24.8 as its group table.IFind the conjugacy classes of the group Ghaving table 24.8 as its multiplication table.
As always, Iis in a class by itself, and we need consider it no further.
Consider next the results of forming X−1AX,a sXruns through the elements of G.
I−1AI A−1AA B−1AB C−1AC D−1AD E−1AE
=IA =IA =AI =CE =DC =ED
=A =A =A =B =B =B
Only AandBare generated. It is clear that {A, B}is one of the conjugacy classes of G.
This can be verified by forming all elements X−1BX; again only AandBappear.
We now need to pick an element not in the two classes already found. Suppose we
pick C.J u s ta sf o r A, we compute X−1CX,a sXruns through the elements of G.T h e
calculations can be done directly using the table and give the following:
X :IA B C D E
X−1CX :CEDCED
Thus C,DandEbelong to the same class. The group is now exhausted, and so the three
conjugacy classes are
{I},{A, B},{C,D,E}.
J
In the case of this small and simple, but non-Abelian, group, only the identity
is in a class by itself (i.e. only Icommutes with all other elements). It is also the
only member of the centre of the group.
Other areas from which examples of conjugacy classes can be taken include
permutations and rotations. Two permutations are in the same class if theircycle specifications have the same structure. For example, in S
5the permutations
(1 3 5)(2)(4) and (2 5 3)(1)(4) are in the same class as each other but in a differentclass from that which contains (1 5)(2 4)(3).
In the case of the continuous rotation group, rotations by the same angle θ
about any two axes labelled iandjare in the same class, because the group
contains a rotation that takes the first axis into the second. Without going intomathematical details, a rotation about axis ican be represented by the operator
R
i(θ), and the two rotations are connected by a relationship of the form
Rj(θ)=φ−1
ijRi(θ)φij,
911
GROUP THEORY
in which φijis the member of the full continuous rotation group that takes axis
iinto axis j.
24.8 Exercises
24.1 For each of the following sets, determine whether they form a group under the op-
eration indicated (where it is relevant you may assume that matrix multiplicationis associative):
(a) the integers (mod 10) under addition;
(b) the integers (mod 10) under multiplication;
(c) the integers 1,2,3, 4,5,6 under multiplication (mod 7);(d) the integers 1,2,3, 4,5 under multiplication (mod 6);
(e) all matrices of the form/
aa−b
0 b
/
where aandbare integers (mod 5), and a/negationslash=0/negationslash=b, under matrix multiplica-
tion;
(f) those elements of the set in (e) that are of order 1 or 2 (taken together);
(g) all matrices of the form/0/@100
a10
bc 1
/1A where a,b,care integers,
under matrix multiplication.
24.2 Which of the following relationships between XandYare equivalence relations?
Give a proof of your conclusions in each case:
(a)XandYare integers and X−Yis odd;
(b)XandYare integers and X−Yis even;
(c)XandYare people and have the same postcode;
(d)XandYare people and have a parent in common;
(e)XandYare people and have the same mother;
(f)XandYaren×nmatrices satisfying Y=PXQ,w h e r e PandQare elements
of a group Gofn×nmatrices.
24.3 Define a binary operation •on the set of real numbers by
x•y=x+y+rxy,
where ris a non-zero real number. Show that the operation •is associative.
Prove that x•y=−r−1if, and only if, x=−r−1ory=−r−1. Hence prove that
the set of all real numbers excluding −r−1forms a group under the operation •.
24.4 Prove that the relationship X∼Y, defined by X∼YifYcan be expressed in
the form
Y=aX+b
cX+d,
with a,b,canddas integers, is an equivalence relation on the set of real numbers
/Rfractur. Identify the class that contains the real number 1.
24.5 The following is a ‘proof’ that reflexivity is an unnecessary axiom for an equiva-
lence relation.
912
24.8 EXERCISES
Because of symmetry X∼Yimplies Y∼X. Then by transitivity X∼Yand
Y∼Ximply X∼X. Thus symmetry and transitivity imply reflexivity, which
therefore need not be separately required.
Demonstrate the flaw in this proof using the set consisting of all real numbers plus
the number i. Show by investigating the following specific cases that, whether or
not reflexivity actually holds, it cannot be deduced from symmetry and transitivityalone.
(a)X∼YifX+Yis real.
(b)X∼YifXYis real.
24.6 Prove that the set Mof matrices
A=
/
ab
0c
/
,
where a, b, c are integers (mod 5) and a/negationslash=0/negationslash=c, forms a non-Abelian group
under matrix multiplication.
Show that the subset containing elements of Mthat are of order 1 or 2 does
not form a proper subgroup of M
(a) using Lagrange’s theorem,
(b) by direct demonstration that the set is not closed.
24.7 Sis the set of all 2 ×2 matrices of the form
A=
/
wx
yz
/
where wz−xy=1 .
Show that Sis a group under matrix multiplication. Which element(s) have order
2? Prove that an element Ahas order 3 if w+z+1=0 .
24.8 Show that, under matrix multiplication, matrices of the form
M(a0,a)=
/
a0+a1i−a2+a3i
a2+a3ia 0−a1i
/
,
where a0and the components of column matrix a=(a1a2a3)Tare real num-
bers satisfying a2
0+|a|2= 1, form a group. Deduce that, under the transformation
z→Mz,w h e r e zis any column matrix, |z|2is invariant.
24.9 If Ais a group in which every element other than the identity, I,h a so r d e r2 ,
prove that Ais Abelian. Hence show that if XandYare distinct elements of A,
neither being equal to the identity, then the set {I,X,Y ,XY }forms a subgroup
ofA.
Deduce that if Bis a group of order 2 p,w i t h pa prime greater than 2, then B
must contain an element of order p.
24.10 The group of rotations (excluding reflections and inversions) in three dimensions
that take a cube into itself is known as the group 432 (or Oin the usual chemical
notation). Show by each of the following methods that this group has 24 elements.
(a) Identify the distinct relevant axes and count the number of qualifying rota-
tions about each.
(b) The orientation of the cube is determined if the directions of two of its body
diagonals are given. Consider the number of distinct ways in which one bodydiagonal can be chosen to be ‘vertical’ and a second diagonal made to liealong a particular direction.
24.11 Identify the eight symmetry operations on a square. Show that they form a group
(known to crystallographers as 4 mmor to chemists as C
4v) having one element
of order 1, five of order 2 and two of order 4. Find its proper subgroups and thecorresponding cosets.
913
GROUP THEORY
24.12 If AandBare two groups then their direct product, A×B,i sd e fi n e dt ob e
the set of ordered pairs ( X,Y), with Xan element of A,Yan element of B
and multiplication given by ( X,Y)(X/prime,Y/prime)=(XX/prime,YY/prime).Prove that A×Bis a
group.
Denote the cyclic group of order nbyCnand the symmetry group of a regular
n-sided figure (an n-gon) by Dn– thus D3is the symmetry group of an equilateral
triangle, as discussed in the text.
(a) By considering the orders of each of their elements, show (i) that C2×C3is
isomorphic to C6, and (ii) that C2×D3is isomorphic to D6.
(b) Are any of D4,C8,C2×C4,C2×C2×C2isomorphic?
24.13 Find the group Ggenerated under matrix multiplication by the matrices
A=
/
01
10
/
, B=
/
0i
i0
/
.
Determine its proper subgroups, and verify for each of them that its cosets
exhaust G.
24.14 Show that if pis prime then the set of rational number pairs ( a, b), excluding
(0,0), with multiplication defined by
(a, b)•(c, d)=(e, f),where ( a+b√p)(c+d√p)=e+f√p,
forms an Abelian group. Show further that the mapping ( a, b)→(a,−b)i sa n
automorphism.
24.15 (a) Denote by Anthe subset of the permutation group Snthat contains all the
even permutations. Show that Anis a subgroup of Sn.
(b) List the elements of S3in cycle notation and identify the subgroup A3.
(c) For each element XofS3,l e tp(X)=1i f Xbelongs to A3andp(X)=−1i fi t
does not. Denote by C2the multiplicative cyclic group of order 2. Determine
the images of each of the elements of S3for the following four mappings:
Φ1:S3→C2 X→p(X)
Φ2:S3→C2 X→−p(X)
Φ3:S3→A3 X→X2
Φ4:S3→S3 X→X3
(d) For each mapping, determine whether the kernel Kis a subgroup of S3and,
if so, whether the mapping is a homomorphism.
24.16 For the group Gwith multiplication table 24.8 and proper subgroup H={I,A,B},
denote the coset {I,A,B}byC1and the coset {C,D,E}byC2.F o r mt h es e to f
all possible products of a member of C1with itself, and denote this by C1C1.
Similarly compute C2C2,C1C2andC2C1. Show that each product coset is equal to
C1or toC2and that a 2 ×2 multiplication table can be formed demonstrating
thatC1andC2are themselves the elements of a group of order 2. A subgroup
likeHwhose cosets themselves form a group is a normal subgroup .
24.17 The group of all non-singular n×nmatrices is known as the general linear
group GL(n) and that with only real elements as GL(n,R). IfR∗denotes the
multiplicative group of non-zero real numbers, prove that the mapping Φ :GL(n,R)→R
∗, defined by Φ( M)=d e t M, is a homomorphism.
Show that the kernel Kof Φ is a subgroup of GL(n,R). Determine its cosets
and show that they themselves form a group.
24.18 The group of reflection–rotation symmetries of a square is known as D4;l e t
Xbe one of its elements. Consider a mapping Φ : D4→S4, the permutation
group on four objects, defined by Φ( X) = the permutation induced by Xon
914
24.9 HINTS AND ANSWERS
the set{x, y, d, d/prime},w h e r e xand yare the two principal axes and dand d/prime
the two principal diagonals, of the square. For example, if Ris a rotation by
π/2, Φ( R) = (12)(34). Show that D4is mapped onto a subgroup of S4and, by
constructing the multiplication tables for D4and the subgroup, prove that the
mapping is a homomorphism.
24.19 Given that matrix Mis a member of the multiplicative group GL(3,R), determine,
for each of the following additional constraints on M(applied separately), whether
the subset satisfying the constraint is a subgroup of GL(3,R):
(a) MT=M;
(b) MTM=I;
(c)|M|=1 ;
(d)Mij=0f o r j>iandMii/negationslash=0 .
24.20 In the quaternion group Qthe elements form the set
{1,−1,i ,−i, j,−j,k,−k},
with i2=j2=k2=−1,ij=kand its cyclic permutations, and ji=−kand
its cyclic permutations. Find the proper subgroups of Qand the corresponding
cosets. Show that the subgroup of order 2 is a normal subgroup, but that theother subgroups are not. Show that Qcannot be isomorphic to the group 4 mm
(C
4v) considered in exercise 24.11.
24.21 Show that D4, the group of symmetries of a square, has two isomorphic subgroups
of order 4. Show further that there exists a two-to-one homomorphism from thequaternion group Qof exercise 24.20 onto one (and hence either) of these two
subgroups, and determine its kernel.
24.22 Show that the matrices
M(θ,x,y)=
/0/@cosθ−sinθx
sinθcosθy
00 1
/1A,
where 0≤θ<2π,−∞<x<∞,−∞<y<∞, form a group under matrix
multiplication.
Show thatthose Mfor which θ= 0 form a subgroup and identify its cosets.
Show that the cosets themselves form a group.
24.23 Find (a) all the proper subgroups and (b) all the conjugacy classes of the symmetry
group of a regular pentagon.
24.9 Hints and answers
24.1†(a) Yes. (b) no, no inverse for 2. (c) yes. (d) no, 2 ×3 is not in the set. (e) yes.
(f) yes, they form a subgroup of order 4, [1 ,0; 0,1] [4,0;0,4] [1,2; 0,4] [4,3;0,1].
(g) yes.
24.2 (a) No, not reflexive; (b) yes, partition of integers into odd and even; (c) yes. (d)
no, not transitive, X→Y→ZifY’s parents both re-marry and XandZare
children of the two second marriages; (e) yes; (f) yes.
24.3 x•(y•z)=x+y+z+r(xy+xz+yz)+r2xyz=(x•y)•z.Show that assuming
x•y=−r−1leads to ( rx+1 ) ( ry+1 )=0 .The inverse of xisx−1=−x/(1 +rx);
show that this is not equal to −r−1.
†Where matrix elements are given as a list, the convention used is [row 1; row 2; ...], individual
entries in each row being separated by commas.
915
GROUP THEORY
24.4 The relevant sets of values for [ a, b, c, d ]a r e[ 1 ,0,0,1],[−d, b, c,−a]a n d[ a/primea+
b/primec, a/primeb+b/primed, c/primea+d/primec, c/primeb+d/primed] for reflexivity, symmetry and transitivity respec-
tively; the rational numbers.
24.5 (a) Consider both X=iandX/negationslash=i. Here, i/negationslash∼i. (b) In this case i∼i, but the
conclusion cannot be deduced from the other axioms. In both cases iis in a class
by itself and no Y, as used in the false proof, can be found.
24.6†Matrices [1 ,3;0,1] and [2 ,3;0,1] do not commute, so the group is non-
Abelian. (a) 12 elements in the set {[1,0;0,1] [4,0; 0,4] [1,b;0,4] [4,b;0,1] with
barbitrary}. The full group has order 4 ×4×5 = 80, which is not divisible by
12. (b) [1 ,0;0,4][1,3;0,4] = [1 ,3; 0,1], which has order >2.
24.7†Use|AB|=|A||B|=1×1 = 1 to prove closure. The inverse has w↔z,
x↔−x,y↔−y, giving|A−1|= 1, i.e. it is in the set. The only element of order
2i s−I;A2can be simplified to [ −(w+1 ),−x;−y,−(z+ 1)].
24.8 Note that if each matrix is written in the form N=(n1,−n∗
2;n2,n∗
1)w i t h|n1|2+
|n2|2=1t h e n NQ=P,w h e r e p1=n1q1−n∗
2q2and p2=n2q1+n∗
1q2with
|p1|2+|p2|2= 1. The inverse of M(a0,a)i sM(a0,−a). Show that M∗TM=I.
24.9 If XY=Z, show that Y=XZandX=ZY,t h e nf o r m YX. Note that the
elements of Bcan only have orders 1, 2 or p. Suppose they all have order 1 or
2; then using the earlier result, whilst noting that 4 does not divide 2 p,l e a d st o
a contradiction.
24.10 (a) Identity = 1, three rotations of πabout face normals, six rotations of ±π/2
about face normals, six rotations of πabout edge diagonals, eight rotations of
±2π/3 about body diagonals. (b) The ‘vertical’ diagonal can be chosen in 4 ×2
ways (either end of each diagonal can be ‘up’). There are then three equivalentrotational positions about the vertical and thus 4 ×2×3 possibilities altogether.
24.11 Using the notation indicated in figure 24.3, Rbeing a rotation of π/2 about an
axis perpendicular to the square, we have: Ihas order 1; R
2,m1,m2,m3,m4have
order 2; R,R3have order 4.
m1(π)
m2(π)
m3(π) m4(π)
Figure 24.3 The notation for exercise 24.11.
Subgroup{I,R,R2,R3}has cosets{I,R,R2,R3},{m1,m2,m3,m4};
subgroup{I,R2}has cosets{I,R2},{R,R3},{m1,m2},{m3,m4};
subgroup{I,m1}has cosets{I,m1},{R,m 3},{R2,m2},{R3,m4};
subgroup{I,m2}has cosets{I,m2},{R,m 4},{R2,m1},{R3,m3};
subgroup{I,m3}has cosets{I,m3},{R,m 2},{R2,m4},{R3,m1};
subgroup{I,m4}has cosets{I,m4},{R,m 1},{R2,m3},{R3,m2}.
24.12 (a) (i) Each has one element of order 1, one element of order 2 two elements of
order 3 and two elements of order 6.(ii) Each has one element of order 1, seven of order 2, two of order 3 andtwo elements of order 6.
916
24.9 HINTS AND ANSWERS
(b) No. C8contains elements of order 8; none of the others could. Every element
ofC2×C2×C2is of order 1 or 2; the remaining two groups must each contain
an element of order 4. D4has one, five and two elements of order 1, 2 and
4 respectively; C2×C4has correspondingly one, three and four elements.
24.13 G={I,A,B,B2,B3,AB,AB2,AB3}. The proper subgroups are as follows:
{I,A};{I,B2},{I,AB2},{I,B,B2,B3},{I,B2,AB,AB3}.
24.14 ( a, b)−1=(a2−pb2)−1(a,−b), which has rational entries with a2/negationslash=pb2since pis
prime.
24.15 (b) A3={(1),(123) ,(132)}.
(d) For Φ 1,K={(1),(123) ,(132)}is a subgroup.
For Φ 2,K={(23),(13),(12)}is not a subgroup because it has no identity element.
For Φ 3,K={(1),(23),(13),(12)}is not a subgroup because it is not closed.
For Φ 4,K={(1),(123) ,(132)}is a subgroup.
Only Φ 1is a homomorphism; Φ 4fails because, for example, [(23)(13)]/prime/negationslash=
(23)/prime(13)/prime.
24.16 C1C1=C2C2=C1,C1C2=C2C1=C2.
24.17 Recall that, for any pair of matrices Pand Q,|PQ|=|P||Q|.Kis the set of all
matrices with unit determinant. The cosets of Kare the sets of matrices whose
determinants are equal; Kitself is the identity in the group of cosets.
24.18 I,R2→(1);R,R3→(12)(34); mx,my→(34); md,md/prime→(12). The multiplication
table for the subgroup is that given in table 24.3.
24.19 (a) No, because the set is not closed. (b) yes. (c) yes. (d) yes.24.20 The subgroup {1,−1}has cosets C
1={1,−1},Ci={i,−i},Cj={j,−j},
Ck={k,−k}. The subgroup {1,i ,−1,−i}has cosets Di={1,i ,−1,−i},D/prime
i=
{j,−j,k,−k}; corresponding pairs of cosets Dj,Dj/primeandDk,Dk/primeare obtained
from subgroups {1,j,−1,−j}and{1,k,−1,−k}respectively. They can be written
down by cyclically permuting i, j, kinDi,D/prime
i. The cosets of {1,−1}form a group
withC1as the identity and CiCj=Cketc. The cosets of {1,i ,−1,−i}do not form
a group since, for example, the product DiD/prime
iinvolves all elements of Q.I ti s
sufficient to notice that 4 mmhas six elements of order 2, whilst Qhas only two.
24.21 Each subgroup contains the identity, a rotation by π, and two reflections. The
homomorphism is ±1→I,±i→R2,±j→mx,±k→mywith kernel {1,−1}.
24.22 Closure is shown by M(θ,x,y)M(φ, x/prime,y/prime)=M(θ+φ, X, Y ),
where X=x+x/primecosθ−y/primesinθandY=y+y/primecosθ+x/primesinθ.
The inverse is given by M(θ, x,y)−1=M(−θ,−xcosθ−ysinθ, xsinθ−ycosθ).
All members of any coset Cθhave the same value for θ.
Cθ1×Cθ2=Cθ1+θ2(mod 2 π). The inverse coset is C−1
θ=C2π−θ.
24.23 There are 10 elements: I, rotations Ri(i=1,4) and reflections mj(j=1,5).
(a) Five proper subgroups of order 2, {I,m j}and one of order 5, {I,R,R2,R3,R4}.
(b) Four conjugacy classes, {I},{R,R4},{R2,R3},{m1,m2,m3,m4,m5}.
917
25
Representation theory
As indicated at the start of the previous chapter, significant conclusions can
often be drawn about a physical system simply from the study of its symmetry
properties. That chapter was devoted to setting up a formal mathematical basis,group theory, with which to describe and classify such properties; the currentchapter shows how to implement the consequences of the resulting classificationsand obtain concrete physical conclusions about the system under study. Theconnection between the two chapters is akin to that between working withcoordinate-free vectors, each denoted by a single symbol, and working with a
coordinate system in which the same vectors are expressed in terms of components.
The ‘coordinate systems’ that we will choose will be ones that are expressed in
terms of matrices; it will be clear that ordinary numbers would not be sufficient,as they make no provision for any non-commutation amongst the elements
of a group. Thus, in this chapter the group elements will be represented by
matrices that have the same commutation relations as the members of the group,whatever the group’s original nature (symmetry operations, functional forms,matrices, permutations, etc.). For some abstract groups it is difficult to give awritten description of the elements and their properties without recourse to suchrepresentations. Most of our applications will be concerned with representationsof the groups that consist of the symmetry operations on molecules containing
two or more identical atoms.
Firstly, in section 25.1, we use an elementary example to demonstrate the kind
of conclusions that can be reached by arguing purely on symmetry grounds. Then
in sections 25.2–25.10 we develop the formal side of representation theory and
establish general procedures and results. Finally, these are used in section 25.11to tackle a variety of problems drawn from across the physical sciences.
918
25.1 DIPOLE MOMENTS OF MOLECULES
(a)H C l (b)C O 2 (c)O3AB
A/primeB/prime
Figure 25.1 Three molecules, ( a)h y d r o g e nc h l o r i d e ,( b) carbon dioxide and
(c) ozone, for which symmetry considerations impose varying degrees of
constraint on their possible electric dipole moments.
25.1 Dipole moments of molecules
Some simple consequences of symmetry can be demonstrated by considering
whether a permanent electric dipole moment can exist in any particular molecule;three simple molecules, hydrogen chloride, carbon dioxide and ozone, are illus-
trated in figure 25.1. Even if a molecule is electrically neutral, an electric dipole
moment will exist in it if the centres of gravity of the positive charges (due toprotons in the atomic nuclei) and of the negative charges (due to the electrons)do not coincide.
For hydrogen chloride there is no reason why they should coincide; indeed, the
normal picture of the binding mechanism in this molecule is that the electron fromthe hydrogen atom moves its average position from that of its proton nucleus tosomewhere between the hydrogen and chlorine nuclei. There is no compensating
movement of positive charge, and a net dipole moment is to be expected – and
is found experimentally.
For the linear molecule carbon dioxide it seems obvious that it cannot have
a dipole moment, because of its symmetry. Putting this rather more rigorously,we note that any rotation about the long axis of the molecule leaves it totallyunchanged; consequently, any component of a permanent electric dipole perpen-dicular to that axis must be zero (a non-zero component would rotate althoughno physical change had taken place in the molecule). That only leaves the pos-
sibility of a component parallel to the axis. However, a rotation of πradians
about the axis AA
/primeshown in figure 25.1( b) carries the molecule into itself, as
does a reflection in a plane through the carbon atom and perpendicular to themolecular axis (i.e. one with its normal parallel to the axis). In both cases the twooxygen atoms change places but, as they are identical, the molecule is indistin-guishable from the original. Either ‘symmetry operation’ would reverse the signof any dipole component directed parallel to the molecular axis; this can only be
compatible with the indistinguishability of the original and final systems if the
parallel component is zero. Thus on symmetry grounds carbon dioxide cannothave a permanent electric dipole moment.
Finally, for ozone, which is angular rather than linear, symmetry does not
919
REPRESENTATION THEORY
place such tight constraints. A dipole-moment component parallel to the axis
BB/prime(figure 25.1( c)) is possible, since there is no symmetry operation that reverses
the component in that direction and at the same time carries the molecule intoan indistinguishable copy of itself. However, a dipole moment perpendicular to
BB
/primeis not possible, since a rotation of πabout BB/primewould both reverse any
such component and carry the ozone molecule into itself – two contradictoryconclusions unless the component is zero.
In summary, symmetry requirements appear in the form that some or all
components of permanent electric dipoles in molecules are forbidden; they do
not show that the other components do exist, only that they may. The greaterthe symmetry of the molecule, the tighter the restrictions on potentially non-zerocomponents of its dipole moment.
In section 23.11 other, more complicated, physical situations will be analysed
using results derived from representation theory. In anticipation of these results,and since it may help the reader to understand where the developments in thenext nine sections are leading, we make here a broad, powerful, but rather formal,statement as follows.
If a physical system is such that after the application of particular rotations or
reflections (or a combination of the two) the final system is indistinguishable fromthe original system then its behaviour, and hence the functions that describe itsbehaviour, must have the corresponding property of invariance when subjected tothe same rotations and reflections.
25.2 Choosing an appropriate formalism
As mentioned in the introduction to this chapter, the elements of a finite group
Gcan be represented by matrices; this is done in the following way. A suitable
column matrix u, known as a basis vector ,†is chosen and is written in terms of
its components u
i,t h ebasis functions ,a s u=(u1u2···un)T.T h e uimay be of
a variety of natures, e.g. numbers, coordinates, functions or even a set of labels,though for any one basis vector they will all be of the same kind.
Once chosen, the basis vector can be used to generate an n-dimensional rep-
resentation of the group as follows. An element Xof the group is selected and
its effect on each basis function u
iis determined. If the action of Xonu1is to
produce u/prime
1, etc. then the set of equations
u/prime
i=Xui (25.1)
generates a new column matrix u/prime=(u/prime
1u/prime2···u/prime
n)T. Having established uand u/prime
†This usage of the term basis vector is not exactly the same as that introduced in subsection 8.1.1.
920
25.2 CHOOSING AN APPROPRIATE FORMALISM
we can determine the n×nmatrix, M(X) say, that connects them by
u/prime=M(X)u. (25.2)
It may seem natural to use the matrix M(X) so generated as the representative
matrix of the element X; in fact, because because we have already chosen the
convention whereby Z=XYimplies that the effect of applying element Zis the
same as that of first applying Yand then applying Xto the result, one further
step has to be taken. So that the representative matrices D(X) may follow the
same convention, i.e.
D(Z)=D(X)D(Y),
and at the same time respect the normal rules of matrix multiplication, it is
necessary to take the transpose ofM(X) as the representative matrix D(X).
Explicitly,
D(X)=MT(X) (25.3)
and (25.2) becomes
u/prime=DT(X)u. (25.4)
Thus the procedure for determining the matrix D(X) that represents the group
element Xin a representation based on basis vector uis summarised by equations
(25.1)–(25.4). †
This procedure is then repeated for each element Xof the group, and the
resulting set of n×nmatrices D={D(X)}is said to be the n-dimensional
representation of Ghaving uas its basis. The need to take the transpose of each
matrix M(X) is not of any fundamental significance, since the only thing that
really matters is whether the matrices D(X) have the appropriate multiplication
properties – and, as defined, they do.
In cases in which the basis functions are labels, the actions of the group
elements are such as to cause rearrangements of the labels. Correspondingly thematrices D(X) contain only ‘1’s and ‘0’s as entries; each row and each column
contains a single ‘1’.
†An alternative procedure in which a row vector is used as the basis vector is possible. Defining
equations of the form uTX=uTD(X) are used, and no additional transpositions are needed to
define the representative matrices. However, row-matrix equations are cumbersome to write out
and in all other parts of this book we have conventionally written operators (here the group
element) to the left of the object on which they operate (here the basis vector).
921
REPRESENTATION THEORYIFor the group S3of permutations on three objects, which has group multiplication ta-
ble 24.8 on p. 897, with (in cycle notation)
I= (1)(2)(3) ,A=( 123 ) ,B =( 132
C= (1)(23) ,D = (3)(12) ,E= (2)(13) ,
use as the components of a basis vector the ordered letter triplets
u1={PQR},u 2={QRP},u 3={RPQ},
u4={PRQ},u 5={QPR},u 6={RQP}.
Generate a six-dimensional representation D={D(X)}of the group and confirm that the
representative matrices multiply according to table 24.8, e.g.
D(C)D(B)=D(E).
It is immediate that the identity permutation I= (1)(2)(3) leaves all uiunchanged, i.e.
u/prime
i=uifor all i. The representative matrix D(I) is thus I6,t h e6×6 unit matrix.
We next take Xas the permutation A= (12 3) and, using (25.1), let it act on each of
the components of the basis vector:
u/prime
1=Au1=( 123 ){PQR}={QRP}=u2
u/prime
2=Au2=( 123 ){QRP}={RPQ}=u3
......
u/prime
6=Au6=( 123 ){RQP}={QPR}=u5.
The matrix M(A) has to be such that u/prime=M(A)u(here dots replace zeroes to aid
readability):
u/prime=
/0BBBBB/@u2
u3
u1
u6
u4
u5
/1CCCCCA=
/0BBBBB/@·1····
·· 1···
1·····
····· 1
··· 1··
···· 1·
/1CCCCCA
/0BBBBB/@u1
u2
u3
u4
u5
u6
/1CCCCCA≡M(A)u.
D(A)i st h e ne q u a lt o MT(A).
The other D(X) are determined in a similar way. In general, if
Xui=uj,
then[M(X)]ij= 1, leading to [D(X)]ji=1a n d [D(X)]jk=0f o r k/negationslash=i. For example,
Cu3= (1)(23){RPQ}={RQP}=u6
implies that [D(C)]63=1a n d [D(C)]6k=0f o r k=1,2,4,5,6. When calculated in full
D(C)=
/0BBBBB/@··· 1··
···· 1·
····· 1
1·····
·1····
·· 1···
/1CCCCCA, D(B)=
/0BBBBB/@·1····
·· 1···
1·····
····· 1
··· 1··
···· 1·
/1CCCCCA,
922
25.2 CHOOSING AN APPROPRIATE FORMALISM
(a) (b) (c)P P P
Q Q Q R R R1 11
22 2 33 3
Figure 25.2 Diagram ( a) shows the definition of the basis vector, ( b)s h o w s
the effect of applying a clockwise rotation of 2 π/3a n d( c) shows the effect of
applying a reflection in the mirror axis through Q.
D(E)=
/0BBBBB/@····· 1
··· 1··
···· 1·
·1····
·· 1···
1·····
/1CCCCCA,
from which it can be verified that D(C)D(B)=D(E).
J
Whilst a representation obtained in this way necessarily has the same dimension
as the order of the group it represents, there are, in general, square matrices ofboth smaller and larger dimensions that can be used to represent the group,though their existence may be less obvious.
One possibility that arises when the group elements are symmetry opera-
tions on an object whose position and orientation can be referred to a spacecoordinate system is called the natural representation . In it the representative
matrices D(X) describe, in terms of a fixed coordinate system, what happens
to a coordinate system that moves with the object when Xis applied. There
is usually some redundancy of the coordinates used in this type of represen-
tation, since interparticle distances are fixed and fewer than 3 Ncoordinates,
where Nis the number of identical particles, are needed to specify uniquely
the object’s position and orientation. Subsection 25.11.1 gives an example thatillustrates both the advantages and disadvantages of the natural representation.We continue here with an example of a natural representation that has no suchredundancy.IUse the fact that the group considered in the previous worked example is isomorphic to
the group of two-dimensional symmetry operati ons on an equilateral triangle to generate a
three-dimensional representation of the group.
Label the triangle’s corners as 1, 2, 3 and three fixed points in space as P, Q, R, so thatinitially corner 1 lies at point P, 2 lies at point Q, and 3 at point R. We take P, Q, R asthe components of the basis vector.
In figure 25.2, ( a) shows the initial configuration and also, formally, the result of applying
the identity Ito the triangle; it is therefore described by the basis vector, (P Q R)
T.
923
REPRESENTATION THEORY
Diagram ( b) shows the the effect of a clockwise rotation by 2 π/3, corresponding to
element Ain the previous example; the new column matrix is (Q R P)T.
Diagram ( c) shows the effect of a typical mirror reflection – the one that leaves the
corner at point Q unchanged (element Din table 24.8 and the previous example); the new
column matrix is now (R Q P)T.
In similar fashion it can be concluded that the column matrix corresponding to element
B,r o t a t i o nb y4 π/3, is (R P Q)T, and that the other two reflections CandEresult in
column matrices (P R Q)Tand (Q P R)Trespectively. The forms of the representative
matrices Mnat(X), (25.2), are now determined by equations such as, for element E,/0/@Q
P
R
/1A=
/0
/@
010
100001
/1A
/0/@P
QR
/1A
implying that
Dnat(E)=
/0/@010
100001
/1AT
=
/0
/@
010
100001
/1A.
In this way the complete repr esentation is obtained as
Dnat(I)=
/0/@100
010
001
/1A,Dnat(A)=
/0/@001
100
010
/1A,Dnat(B)=
/0/@010
001
100
/1A,
Dnat(C)=
/0/@100
001010
/1A,Dnat(D)=
/0/@001
010100
/1A,Dnat(E)=
/0/@010
100001
/1A.
It should be emphasised that although the group contains six elements this representation
is three-dimensional.
J
We will concentrate on matrix representations of finite groups, particularly
rotation and reflection groups (the so-called crystal point groups). The general
ideas carry over to infinite groups, such as the continuous rotation groups, but ina book such as this, which aims to cover many areas of applicable mathematics,some topics can only be mentioned and not explored. We now give the formaldefinition of a representation.
Definition. A representation D={D(X)}of a group Gis an assignment of a non-
singular square n×nmatrix D(X)to each element Xbelonging to G, such that
(i)D(I)=I
n,the unit n×nmatrix ,
(ii)D(X)D(Y)=D(XY)for any two elements XandYbelonging to G,i . e .t h e
matrices multiply in the same way as the group elements they represent.
As mentioned previously, a representation by n×nmatrices is said to be an
n-dimensional representation ofG. The dimension nis not to be confused with
g, the order of the group, which gives the number of matrices needed in the
representation, though they might not all be different.
A consequence of the two defining conditions for a representation is that the
924
25.2 CHOOSING AN APPROPRIATE FORMALISM
matrix associated with the inverse of Xis the inverse of the matrix associated
with X. This follows immediately from setting Y=X−1in (ii):
D(X)D(X−1)=D(XX−1)=D(I)=In;
hence
D(X−1)=[D(X)]−1.
As an example, the four-element Abelian group that consists of the set {1,i ,−1,−i}
under ordinary multiplication has a two-dimensional representation based on thecolumn matrix (1i)
T:
D(1) =parenleftbigg10
01parenrightbigg
, D(i)=parenleftbigg0−1
10parenrightbigg
,
D(−1) =parenleftbigg−10
0−1parenrightbigg
,D(−i)=parenleftbigg01
−10parenrightbigg
.
The reader should check that D(i)D(−i)= D(1), D(i)D(i)= D(−1) etc., i.e. that
the matrices do have exactly the same multiplication properties as the elementsof the group. Having done so, the reader may also wonder why anybody wouldbother with the representative matrices, when the original elements are so muchsimpler to handle! As we will see later, once some general properties of matrixrepresentations have been established, the analysis of large groups, both Abelian
and non-Abelian, can be reduced to routine, almost cookbook, procedures.
Ann-dimensional representation of Gis a homomorphism of Ginto the set of
invertible n×nmatrices (i.e. n×nmatrices that have inverses or, equivalently,
have non-zero determinants); this set is usually known as the general linear groupand denoted by GL( n). In general the same matrix may represent more than one
element of G; if, however, all the matrices representing the elements of Gare
different then the representation is said to be faithful , and the homomorphism
becomes an isomorphism onto a subgroup of GL( n).
A trivial but important representation is D(X)= I
nfor all elements XofG.
Clearly both of the defining relationships are satisfied, and there is no restriction
on the value of n. However, such a representation is not a faithful one.
To sum up, in the context of a rotation–reflection group, the transposes of
the set of n×nmatrices D(X)t h a tm a k eu par e p r e s e n t a t i o n Dmay be thought
of as describing what happens to an n-component basis vector of coordinates,
(xy ···)T, or of functions, (Ψ 1Ψ2···)T,t h eΨ ithemselves being functions
of coordinates, when the group operation Xis carried out on each of the
coordinates or functions. For example, to return to the symmetry operationson an equilateral triangle, the clockwise rotation by 2 π/3,R, carries the three-
925
REPRESENTATION THEORY
dimensional basis vector ( xyz )Tinto the column matrix
−1
2x+√
3
2y
−√
3
2x−1
2y
z
whilst the two-dimensional basis vector of functions ( r23z2−r2)Tis unaltered,
as neither rnorzis changed by the rotation. The fact that zis unchanged by
any of the operations of the group shows that the components x,y,zactually
divide (i.e. are ‘reducible’, to anticipate a more formal description) into two sets:one comprises z, which is unchanged by any of the operations, and the other
comprises x,y, which change as a pair into linear combinations of themselves.
This is an important observation to which we return in section 25.4.
25.3 Equivalent representations
IfDis an n-dimensional representation of a group G,a n d Qis any fixed invert-
iblen×nmatrix (|Q|/negationslash= 0), then the set of matrices defined by the similarity
transformation
D
Q(X)=Q−1D(X)Q (25.5)
also forms a representation DQofG,s a i dt ob e equivalent toD.W ec a ns e ef r o ma
comparison with the definition in section 25.2 that they do form a representation:
(i)DQ(I)=Q−1D(I)Q=Q−1InQ=In,
(ii)DQ(X)DQ(Y)=Q−1D(X)QQ−1D(Y)Q=Q−1D(X)D(Y)Q
=Q−1D(XY)Q=DQ(XY).
Since we can always transform between equivalent representations using a non-
singular matrix Q, we will consider such representations to be one and the same.
Despite the similarity of words and mani pulations to those of subsection 24.7.1,
that two representations are equivalent does not constitute an ‘equivalence re-lation’ – for example, the reflexive property does not hold for a general fixedmatrix Q. However, if Qwere not fixed, but simply restricted to belonging to
a set of matrices that themselves form a group, then (25.5) would constitute anequivalence relation.
The general invertible matrix Qthat appears in the definition (25.5) of equiv-
alent matrices describes changes arising from a change in the coordinate system
(i.e. in the set of basis functions). As before, suppose that the effect of an opera-
tionXon the basis functions is expressed by the action of M(X)( w h i c hi se q u a l
toD
T(X)) on the corresponding basis vector:
u/prime=M(X)u=DT(X)u. (25.6)
926
25.3 EQUIVALENT REPRESENTATIONS
A change of basis would be given by uQ=Quand u/prime
Q=Qu/prime, and we may write
u/prime
Q=Qu/prime=QM(X)u=QDT(X)Q−1uQ. (25.7)
This is of the same form as (25.6), i.e.
u/prime
Q=DT
QT(X)uQ, (25.8)
where DQT(X)=( QT)−1D(X)QTis related to D(X) by a similarity transforma-
tion. Thus DQT(X) represents the same linear transformation as D(X), but with
respect to a new basis vector uQ; this supports our contention that representa-
tions connected by similarity transformations should be considered as the same
representation.IFor the four-element Abelian group consisting of the set {1,i ,−1,−i}under ordinary
multiplication, discussed near the end of section 25.2, change the basis vector from u=
(1i)TtouQ=( 3−i2i−5)T. Find the real transformation matrix Q. Show that the
transformed representative matrix for element i,DQT(i), is given by
DQT(i)=
/
17−29
10−17
/
and verify that DT
QT(i)uQ=iuQ.
Firstly, we solve the matrix equation/
3−i
2i−5
/
=
/
ab
cd
//
1
i
/
,
with a, b, c, d real. This gives Qand hence Q−1as
Q=
/
3−1
−52
/
, Q−1=
/
21
53
/
.
Following (25.7) we now find the transpose of DQT(i)a s
QDT(i)Q−1=
/
3−1
−52
//
01
−10
//
21
53
/
=
/
17 10
−29−17
/
and hence DQT(i) is as stated. Finally,
DT
QT(i)uQ=
/
17 10
−29−17
//
3−i
2i−5
/
=
/
1+3 i
−2−5i
/
=i
/
3−i
2i−5
/
=iuQ,
as required.
J
Although we will not prove it, it can be shown that any finite representation
of a finite group of linear transformations that preserve spatial length (or, inquantum mechanics, preserve the magnitude of a wavefunction) is equivalent to
927
REPRESENTATION THEORY
a representation in which all the matrices are unitary (see chapter 8) and so from
now on we will consider only unitary representations .
25.4 Reducibility of a representation
We have seen already that it is possible to have more than one representation
of any particular group. For example, the group {1,i ,−1,−i}under ordinary
multiplication has been shown to have a set of 2 ×2 matrices, and a set of four
unitn×nmatrices In, as two of its possible representations.
Consider two or more representations, D(1),D(2), ..., D(N), which may be
of different dimensions, of a group G. Now combine the matrices D(1)(X),
D(2)(X), ..., D(N)(X) that correspond to element XofGinto a larger block-
diagonal matrix:
D(2)(X)
D(N)(X). . .D(X) =
00
D(1)(X)
(25.9)
Then D={D(X)}is the matrix representation of the group obtained by combining
the basis vectors of D(1),D(2), ..., D(N)into one larger basis vector. If, knowingly or
unknowingly, we had started with this larger basis vector and found the matricesof the representation Dto have the form shown in (25.9), or to have a form
that can be transformed into this by a similarity transformation (25.5) (using,of course, the samematrix Qfor each of the matrices D(X)) then we would say
that Disreducible and that each matrix D(X) can be written as the direct sum of
smaller representations:
D(X)=D
(1)(X)⊕D(2)(X)⊕···⊕D(N)(X).
It may be that some or all of the matrices D(1)(X),D(2)(X), ..., D(N)themselves
can be further reduced – i.e. written in block diagonal form. For example,suppose that the representation D
(1), say, has a basis vector ( xyz )T; then, for
the symmetry group of an equilateral triangle, whilst xandyare mixed together
for at least one of the operations X,zis never changed. In this case the 3 ×3
representative matrix D(1)(X) can itself be written in block diagonal form as a
928
25.4 REDUCIBILITY OF A REPRESENTATION
2×2m a t r i xa n da1 ×1 matrix. The direct-sum matrix D(X) can now be written
a b
cd
1
D(2)(X)
D(N)(X). . .D (X) =
00
(25.10)
but the first two blocks can be reduced no further.
When all the other representations D(2)(X),...have been similarly treated,
w h a tr e m a i n si ss a i dt ob e irreducible and has the characteristic of being block
diagonal, with blocks that individually cannot be reduced further. The blocks are
known as the irreducible representations of G, often abbreviated to the irreps of
G, and we denote them by ˆD(i). They form the building blocks of representation
theory, and it is their properties that are used to analyse any given physicalsituation which is invariant under the operations that form the elements of G.
Any representation can be written as a linear combination of irreps.
If, however, the initial choice uof basis vector for the representation Dis
arbitrary, as it is in general, then it is unlikely that the matrices D(X) will
assume obviously block diagonal forms (it should be noted, though, that sincethe matrices are square, even a matrix with non-zero entries only in the extremetop right and bottom left positions is technically block diagonal). In general, itwill be possible to reduce them to block diagonal matrices with more than oneblock; this reduction corresponds to a transformation Qto a new basis vector
u
Q, as described in section 25.3.
In any particular representation D, each constituent irrep ˆD(i)may appear any
number of times, or not at all, subject to the obvious restriction that the sum ofall the irrep dimensions must add up to the dimension of Ditself. Let us say that
ˆD(i)appears mitimes. The general expansion of Dis then written
D=m1ˆD(1)⊕m2ˆD(2)⊕···⊕mNˆD(N), (25.11)
where if Gis finite so is N.
This is such an important result that we shall now restate the situation in
somewhat different language. When the set of matrices that forms a representation
929
REPRESENTATION THEORY
of a particular group of symmetry operations has been brought to irreducible
form, the implications are as follows.
(i) Those components of the basis vector that correspond to rows in the
representation matrices with a single-entry block, i.e. a 1 ×1b l o c k ,a r e
unchanged by the operations of the group. Such a coordinate or functionis said to transform according to a one-dimensional irrep of G.I nt h e
example given in (25.10), that the entry on the third row forms a 1 ×1
block implies that the third entry in the basis vector ( xyz ···)
T,
namely z, is invariant under the two-dimensional symmetry operations on
an equilateral triangle in the xy-plane.
(ii) If, in any of the gmatrices of the representation, the largest-sized block
located on the row or column corresponding to a particular coordinate(or function) in the basis vector is n×n, then that coordinate (or function)
is mixed by the symmetry operations with n−1 others and is said to
transform according to an n-dimensional irrep of G. Thus in the matrix
(25.10), xis the first entry in the complete basis vector; the first row of
the matrix contains two non-zero entries, as does the first column, and soxis part of a two-component basis vector whose components are mixed
by the symmetry operations of G. The other component is y.
The result (25.11) may also be formulated in terms of the more abstract notion
of vector spaces (chapter 8). The set of gmatrices that forms an n-dimensional
representation Dof the group Gcan be thought of as acting on column matrices
corresponding to vectors in an n-dimensional vector space Vspanned by the basis
functions of the representation. If there exists a proper subspace WofV,s u c h
that if a vector whose column matrix is wbelongs to Wthen the vector whose
column matrix is D(X)walso belongs to W,f o ra l l Xbelonging to G,t h e ni t
follows that Dis reducible. We say that the subspace Wis invariant under the
actions of the elements of G. With Dunitary, the orthogonal complement W
⊥of
W,i . e .t h ev e c t o rs p a c e Vremaining when the subspace Whas been removed, is
also invariant, and all the matrices D(X) split into two blocks acting separately
onWandW⊥.B o t h WandW⊥may contain further invariant subspaces and
be split still further.
As a concrete example of this approach, consider in plane polar coordinates
ρ, φ, the effect of rotations about the polar axis on the infinite-dimensional vector
space Vof all functions of φthat satisfy the Dirichlet conditions for expansion
as a Fourier series (see section 12.1). We take as our basis functions the set{sinmφ,cosmφ}for integer values m=0,1,2,...; this is an infinite-dimensional
representation ( n=∞) and, since a rotation about the polar axis can be through
any angle α(0≤α<2π), the group Gis a subgroup of the continuous rotation
group and has its order gformally equal to infinity.
930
25.4 REDUCIBILITY OF A REPRESENTATION
Now, for some k, consider a vector win the space Wkspanned by{sinkφ,coskφ},
say w=asinkφ+bcoskφ. Under a rotation by αabout the polar axis, asinkφ
becomes asink(φ+α), which can be written as acoskαsinkφ+asinkαcoskφ,i . e
as a linear combination of sin kφand cos kφ; similarly cos kφbecomes another
linear combination of the same two functions. Thus w/prime=D(α)walso belongs
toWkfor any αand we can conclude that Wkis an invariant irreducible two-
dimensional subspace of V. It follows that D(α) is reducible and that, since the
result holds for every k, in its reduced form D(α) has an infinite series of identical
2×2 blocks on its leading diagonal; each block will have the form
parenleftbiggcosα−sinα
sinαcosαparenrightbigg
.
We note that the particular case k= 0 is special, in that then sin kφ=0a n d
coskφ=1 ,f o ra l l φ; consequently the first 2 ×2 block in D(α) is reducible further
and becomes two single-entry blocks.
A second illustration of the connection between the behaviour of vector spaces
under the actions of the elements of a group and the form of the matrix repre-sentation of the group is provided by the vector space spanned by the sphericalharmonics Y
/lscriptm(θ,φ). This contains subspaces, corresponding to the different
values of /lscript, that are invariant under the actions of the elements of the full three-
dimensional rotation group; the corresponding matrices are block-diagonal, andthose entries that correspond to the part of the basis containing Y
/lscriptm(θ,φ)f o r ma
(2/lscript+1 )×(2/lscript+ 1) block.
To illustrate further the irreps of a group, we return again to the group Gof
two-dimensional rotation and reflection symmetries of an equilateral triangle, orequivalently the permutation group S
3; this may be shown, using the methods of
section 25.7 below, to have three irreps. Firstly, we have already seen that the setMof six orthogonal 2 ×2 matrices given in section (24.3), equation (24.13), is
isomorphic to G. These matrices therefore form not only a representation of G,
but a faithful one. It should be noticed that, although Gcontains six elements,
the matrices are only 2 ×2. However, they contain no invariant 1 ×1 sub-block
(which for 2 ×2 matrices would require them all to be diagonal) and neither can
allthe matrices be made block diagonal by the samesimilarity transformation;
they therefore form a two-dimensional irrep of G.
Secondly, as previously noted, every group has one (unfaithful) irrep in which
every element is represented by the 1 ×1m a t r i x I
1, or, more simply, 1.
Thirdly an (unfaithful) irrep of Gis given by assignment of the one-dimensional
set of six ‘matrices’ {1,1,1,−1,−1,−1}to the symmetry operations {I,R,R/prime,K,
L, M}respectively, or to the group elements {I,A,B,C,D,E }respectively; see
section 24.3. In terms of the permutation group S3, 1 corresponds to even
permutations and −1 to odd permutations, ‘odd’ or ‘even’ referring to the number
of simple pair interchanges to which a permutation is equivalent. That these
931
REPRESENTATION THEORY
assignments are in accord with the group multiplication table 24.8 should be
checked.
Thus the three irreps of the group G(i.e. the group 3 morC3vorS3), are, using
the conventional notation A 1,A2, E (see section 25.8), as follows:
Element
IABCDE
A11111 1 1
Irrep A 2111 −1−1−1
E MIMAMBMCMDME(25.12)
where
MI=parenleftBigg
10
01parenrightBigg
, MA=parenleftBigg
−1
2√
3
2
−√
3
2−1
2parenrightBigg
, MB=parenleftBigg
−1
2−√
3
2√
3
2−1
2parenrightBigg
,
MC=parenleftBigg
−10
01parenrightBigg
,MD=parenleftBigg
1
2−√
3
2
−√
3
2−1
2parenrightBigg
,ME=parenleftBigg
1
2√
3
2√
3
2−1
2parenrightBigg
.
25.5 The orthogonality theorem for irreducible representations
We come now to the central theorem of representation theory, a theorem that
justifies the relatively routine application of certain procedures to determinethe restrictions that are inherent in physical systems that have some degree ofrotational or reflection symmetry. The development of the theorem is long andquite complex when presented in its entirety, and the reader will have to referelsewhere for the proof. †
The theorem states that, in a certain sense, the irreps of a group Gare as
orthogonal as possible, as follows. If, for each irrep, the elements in any one
p o s i t i o ni ne a c ho ft h e gm a t r i c e sa r eu s e dt om a k eu p g-component column
matrices then
(i) any two such column matrices coming from different irreps are orthogonal;
(ii) any two such column matrices coming from different positions in the
matrices of the same irrep are orthogonal.
This orthogonality is in addition to the irreps’ being in the form of orthogo-
nal (unitary) matrices and thus each comprising mutually orthogonal rows andcolumns.
†See, e.g., Groups, Representation and Physics , H.F. Jones (Institute of Physics), Group Theory in
Quantum Mechanics , J. F. Cornwell (Academic Press), or Linear Representations of Finite Groups ,
J. P. Sore (Springer-Verlag).
932
25.5 ORTHOGONALITY THEOREM FOR IRREDUCIBLE REPRESENTATIONS
More mathematically, if we denote the entry in the ith row and jth column of a
matrix D(X)b y[D(X)]ij,a n d ˆD(λ)andˆD(µ)are two irreps of Ghaving dimensions
nλandnµrespectively, then
summationdisplay
XbracketleftBig
ˆD(λ)(X)bracketrightBig∗
ijbracketleftBig
ˆD(µ)(X)bracketrightBig
kl=g
nλδikδjlδλµ. (25.13)
This rather forbidding-looking equation needs some further explanation.
Firstly, the asterisk indicates that the complex conjugate should be taken if
necessary, though all our representations so far have involved only real matrixelements. Each Kronecker delta function on the right-hand side has the value 1if its two subscripts are equal and has the value 0 otherwise. Thus the right-handside is only non-zero if i=k,j=landλ=µ, all at the same time.
Secondly, the summation over the group elements Xmeans that gcontributions
have to be added together, each contribution being a product of entries drawn
from the representative matrices in the two irreps ˆD
(λ)={ˆD(λ)(X)}andˆD(µ)=
{ˆD(µ)(X)}.T h e gcontributions arise as Xruns over the gelements of G.
Thus, putting these remarks together, the summation will produce zero if either
(i) the matrix elements are not taken from exactly the same position in every
matrix, including cases in which it is not possible to do so because the
irreps ˆD(λ)andˆD(µ)have different dimensions, or
(ii) even if ˆD(λ)andˆD(µ)do have the same dimensions and the matrix elements
are from the same positions in every matrix, they are different irreps, i.e.λ/negationslash=µ.
Some numerical illustrations based on the irreps A
1,A2and E of the group 3 m
(orC3vorS3) will probably provide the clearest explanation (see (25.12)).
(a) Take i=j=k=l= 1, with ˆD(λ)=A 1andˆD(µ)=A 2. Equation (25.13)
then reads
1(1) + 1(1) + 1(1) + 1( −1) + 1(−1) + 1(−1) = 0 ,
as expected, since λ/negationslash=µ.
(b) Take ( i, j)a s( 1 ,2) and ( k,l)a s( 2 ,2), corresponding to different matrix
positions within the same irrep ˆD(λ)=ˆD(µ)= E. Substituting in (25.13)
gives
0(1) +parenleftBig
−√
3
2parenrightBigparenleftbig
−1
2parenrightbig
+parenleftBig√
3
2parenrightBigparenleftbig
−1
2parenrightbig
+0 ( 1 )+parenleftBig
−√
3
2parenrightBigparenleftbig
−1
2parenrightbig
+parenleftBig√
3
2parenrightBigparenleftbig
−1
2parenrightbig
=0.
(c) Take ( i, j)a s( 1 ,2), and ( k,l)a s( 1 ,2), corresponding to the same matrix
positions within the same irrep ˆD(λ)=ˆD(µ)= E. Substituting in (25.13)
gives
0(0)+parenleftBig
−√
3
2parenrightBigparenleftBig
−√
3
2parenrightBig
+parenleftBig√
3
2parenrightBigparenleftBig√
3
2parenrightBig
+0(0)+parenleftBig
−√
3
2parenrightBigparenleftBig
−√
3
2parenrightBig
+parenleftBig√
3
2parenrightBigparenleftBig√
3
2parenrightBig
=6
2.
933
REPRESENTATION THEORY
(d) No explicit calculation is needed to see that if i=j=k=l= 1, with
ˆD(λ)=ˆD(µ)=A 1(or A 2), then each term in the sum is either 12or (−1)2
and the total is 6, as predicted by the right-hand side of (25.13) since g=6
andnλ=1 .
25.6 Characters
The actual matrices of general representations and irreps are cumbersome to
work with, and they are not unique since there is always the freedom to change
the coordinate system, i.e. the components of the basis vector (see section 25.3),and hence the entries in the matrices. However, one thing that does not changefor a matrix under such an equivalence (similarity) transformation – i.e. undera change of basis – is the trace of the matrix. This was shown in chapter 8,but is repeated here. The trace of a matrix Ais the sum of its diagonal ele-
ments,
TrA=
nsummationdisplay
i=1Aii
or, using the summation convention (section 21.1), simply Aii. Under a similarity
transformation, again using the summation convention,
[DQ(X)]ii=[Q−1]ij[D(X)]jk[Q]ki
=[D(X)]jk[Q]ki[Q−1]ij
=[D(X)]jk[I]kj
=[D(X)]jj,
showing that the traces of equivalent matrices are equal.
This fact can be used to greatly simplify work with representations, though with
some partial loss of the information content of the full matrices. For example,
using trace values alone it is not possible to distinguish between the two groupsknown as 4 mmand¯42m,o ra s C
4vandD2drespectively, even though the two
groups are not isomorphic. To make use of these simplifications we now definethe characters of a representation.
Definition. Thecharacters χ(D)of a representation Dof a group Gare defined as
the set of traces of the matrices D(X), one for each element XofG.
At this stage there will be gcharacters, but, as we noted in subsection 24.7.3,
elements A,BofGin the same conjugacy class are connected by equations of
the form B=X
−1AX. It follows that their matrix representations are connected
by corresponding equations of the form D(B)=D(X−1)D(A)D(X) ,a n ds ob yt h e
argument just given their representations will have equal traces and hence equalcharacters. Thus elements in the same conjugacy class have the same characters ,
934
25.6 CHARACTERS
3mIA ,BC ,D,E
A111 1 z;z2;x2+y2
A211 −1 Rz
E2−10 (x, y); (xz, yz); (Rx,Ry); (x2−y2,2xy)
Table 25.1 The character table for the irreps of group 3 m(C3vorS3). The
right-hand column lists some common functions that transform according to
the irrep against which each is shown (see text).
though, in general, these will vary from one representation to another. However,
it might also happen that two or more conjugacy classes have the same characters
in a representation – indeed, in the trivial irrep A 1, see (25.12), every element
inevitably has the character 1.
For the irrep A 2of the group 3 m, the classes {I},{A, B}and{C,D,E}have
characters 1, 1 and −1, respectively, whilst they have characters 2, −1a n d0
respectively in irrep E.
We are thus able to draw up a character table for the group 3 mas shown
in table 25.1. This table holds in compact form most of the important infor-
mation on the behaviour of functions under the two-dimensional rotational andreflection symmetries of an equilateral triangle, i.e. under the elements of group3m. The entry under Ifor any irrep gives the dimension of the irrep, since it
is equal to the trace of the unit matrix whose dimension is equal to that ofthe irrep. In other words, for the λth irrep χ
(λ)(I)=nλ,w h e r e nλis its dimen-
sion.
In the extreme right-hand column we list some common functions of Cartesian
coordinates that transform, under the group 3 m, according to the irrep on whose
line they are listed. Thus, as we have seen, z,z2,a n d x2+y2are all unchanged
by the group operations (though xandyindividually are affected) and so are
listed against the one-dimensional irrep A 1. Each of the pairs ( x, y), (xz, yz), and
(x2−y2,2xy), however, is mixed as a pair by some of the operations, and so these
pairs are listed against the two-dimensional irrep E: each pair forms a basis forthis irrep.
The quantities R
x,Ryand Rzrefer to rotations about the indicated axes;
they transform in the same way as the corresponding components of angularmomentum J, and their behaviour can be established by examining how the
components of J=r×ptransform under the operations of the group. To do
this explicitly is beyond the scope of this book. However, it can be noted that
R
z, being listed opposite the one-dimensional A 2, is unchanged by Iand by the
rotations AandBbut changes sign under the mirror reflections C,D,a n d E,a s
would be expected.
935
REPRESENTATION THEORY
25.6.1 Orthogonality property of characters
Some of the most important properties of characters can be deduced from the
orthogonality theorem (25.13),
summationdisplay
XbracketleftBig
ˆD(λ)(X)bracketrightBig∗
ijbracketleftBig
ˆD(µ)(X)bracketrightBig
kl=g
nλδikδjlδλµ.
If we set j=iandl=k,s ot h a tb o t hf a c t o r si na n yp a r t i c u l a rt e r mi nt h e
summation refer to diagonal elements of the representative matrices, and then
sum both sides over iandk,w eo b t a i n
summationdisplay
Xnλsummationdisplay
i=1nµsummationdisplay
k=1bracketleftBig
ˆD(λ)(X)bracketrightBig∗
iibracketleftBig
ˆD(µ)(X)bracketrightBig
kk=g
nλnλsummationdisplay
i=1nµsummationdisplay
k=1δikδikδλµ.
Expressed in term of characters, this reads
summationdisplay
Xbracketleftbig
χ(λ)(X)bracketrightbig∗χ(µ)(X)=g
nλnλsummationdisplay
i=1δ2
iiδλµ=g
nλnλsummationdisplay
i=11×δλµ=gδλµ.
(25.14)
In words, the ( g-component) ‘vectors’ formed from the characters of the various
irreps of a group are mutually orthogonal, but each one has a squared magnitude(the sum of the squares of its components) equal to the order of the group.
Since, as noted in the previous subsection, group elements in the same class
have the same characters, (25.14) can be written as a sum over classes rather than
elements. If c
idenotes the number of elements in class CiandXiany element of
Ci,t h e n
summationdisplay
icibracketleftbig
χ(λ)(Xi)bracketrightbig∗χ(µ)(Xi)=gδλµ. (25.15)
Although we do not prove it here, there also exists a ‘completeness’ relation for
characters. It makes a statement about the products of characters for a fixed pair
of group elements, X1andX2, when the products are summed over all possible
irreps of the group. This is the converse of the summation process defined by(25.14). The completeness relation states that
summationdisplay
λbracketleftbig
χ(λ)(X1)bracketrightbig∗χ(λ)(X2)=g
c1δC1C2, (25.16)
where element X1belongs to conjugacy class C1andX2belongs to C2. Thus the
sum is zero unless X1andX2belong to the same class. For table 25.1 we can
verify that these results are valid.
(i) For ˆD(λ)=ˆD(µ)=A 1or A 2, (25.15) reads
1(1) + 2(1) + 3(1) = 6 ,
936
25.7 COUNTING IRREPS USING CHARACTERS
whilst for ˆD(λ)=ˆD(µ)=E ,i tg i v e s
1(22) + 2(1) + 3(0) = 6 .
(ii) For ˆD(λ)=A 2andˆD(µ)= E, say, (25.15) reads
1(1)(2) + 2(1)( −1) + 3(−1)(0) = 0 .
(iii) For X1=AandX2=D, say, (25.16) reads
1(1) + 1(−1) + (−1)(0) = 0 ,
whilst for X1=CandX2=E, both of which belong to class C3for which
c3=3 ,
1(1) + (−1)(−1) + (0)(0) = 2 =6
3.
25.7 Counting irreps using characters
The expression of a general representation D={D(X)}in terms of irreps, as
given in (25.11), can be simplified by going from the full matrix form to that ofcharacters. Thus
D(X)=m
1ˆD(1)(X)⊕m2ˆD(2)(X)⊕···⊕mNˆD(N)(X)
becomes, on taking the trace of both sides,
χ(X)=Nsummationdisplay
λ=1mλχ(λ)(X). (25.17)
Given the characters of the irreps of the group Gto which the elements Xbelong,
and the characters of the representation D={D(X)},t h e gequations (25.17)
can be solved as simultaneous equations in the mλ, either by inspection or by
multiplying both sides bybracketleftbig
χ(µ)(X)bracketrightbig∗and summing over X, making use of (25.14)
and (25.15), to obtain
mµ=1
gsummationdisplay
Xbracketleftbig
χ(µ)(X)bracketrightbig∗χ(X)=1
gsummationdisplay
icibracketleftbig
χ(µ)(Xi)bracketrightbig∗χ(Xi). (25.18)
That an unambiguous formula can be given for each mλ, once the character
set(the set of characters of each of the group elements or, equivalently, of
each of the conjugacy classes) of Dis known, shows that, for any particular
group, two representations with the same characters are equivalent. This stronglysuggests something that can be shown, namely, the number of irreps = the number
of conjugacy classes. The argument is as follows. Equation (25.17) is a set of
simultaneous equations for Nunknowns, the m
λ, some of which may be zero. The
value of Nis equal to the number of irreps of G.T h e r ea r e gdifferent values of
X, but the number of different equations is only equal to the number of distinct
937
REPRESENTATION THEORY
conjugacy classes, since any two elements of Gin the same class have the same
character set and therefore generate the same equation. For a unique solutionto simultaneous equations in Nunknowns, exactly Nindependent equations are
needed. Thus Nis also the number of classes, establishing the stated result.IDetermine the irreps contained in the representation of the group 3min the vector space
spanned by the functions x2,y2,xy.
We first note that although these functions are not orthogonal they form a basis set for a
representation, since they are linearly independent quadratic forms in xandyand any other
quadratic form can be written (uniquely) in terms of them. We must establish how they
transform under the symmetry operations of group 3 m. We need to do so only for a repre-
sentative element of each conjugacy class, and naturally we take the simplest in each case.
T h efi r s tc l a s sc o n t a i n so n l y I(as always) and clearly D(I)i st h e3×3 unit matrix.
The second class contains the rotations, AandB, and we choose to find D(A). Since,
under A,
x→−1
2x+√
3
2y and y→−√
3
2x−1
2y,
it follows that
x2→1
4x2−√
3
2xy+3
4y2,y2→3
4x2+√
3
2xy+1
4y2(25.19)
and
xy→√
3
4x2−1
2xy−√
3
4y2. (25.20)
Hence D(A) can be deduced and is given below.
The third and final class contains the reflections, C,DandE;o ft h e s e Cis much the
easiest to deal with. Under C,x→−xandy→y,c a u s i n g xyto change sign but leaving
x2andy2unaltered. The three matrices needed are thus
D(I)=I3,D(C)=
/0/@10 0
01 0
00−1
/1A,D(A)=
/0BB/@1
43
4−√
3
2
3
41
4√
3
2√
3
4−√
3
4−1
2
/1CCA;
their traces are respectively 3, 1 and 0.
It should be noticed that much more work has been done here than is necessary, since
the traces can be computed immediately from the effects of the symmetry operations on thebasis functions. All that is needed is the weight of each basis function in the transformedexpression for that function; these are clearly 1, 1, 1 for I,a n d
1
4,1
4,−1
2forA, from (25.19)
and (25.20), and 1, 1, −1f o r C, from the observations made just above the displayed
matrices. The traces are then the sums of these weights. The off-diagonal elements of thematrices need not be found, nor need the matrices be written out.
From (25.17) we now need to find a superposition of the characters of the irreps that
gives representation Din the bottom line of table 25.2.
By inspection it is obvious that D=A
1⊕E, but we can use (25.18) formally:
mA1=1
6[1(1)(3) + 2(1)(0) + 3(1)(1)] = 1 ,
mA2=1
6[1(1)(3) + 2(1)(0) + 3( −1)(1)] = 0 ,
mE=1
6[1(2)(3) + 2( −1)(0) + 3(0)(1)] = 1 .
Thus A 1and E appear once each in the reduction of D,a n dA 2not at all. Table 25.1
gives the further information, not needed here, that it is the combination x2+y2that
transforms as a one-dimensional irrep and the pair ( x2−y2,2xy)t h a tf o r m sab a s i so f
the two-dimensional irrep, E.
J
938
25.7 COUNTING IRREPS USING CHARACTERS
Classes
Irrep IA BC D E
A1 11 1
A2 11 −1
E 2−10
D 30 1
Table 25.2 The characters of the irreps of the group 3 mand of the represen-
tation D, which must be a superposition of some of them.
25.7.1 Summation rules for irreps
The first summation rule for irreps is a simple restatement of (25.14), with µset
equal to λ;i tt h e nr e a d s
summationdisplay
Xbracketleftbig
χ(λ)(X)bracketrightbig∗χ(λ)(X)=g.
In words, the sum of the squares (modulus squared if necessary) of the characters
of an irrep taken over all elements of the group adds up to the order of thegroup. For group 3 m(table 25.1), this takes the following explicit forms:
for A
1, 1(12)+2 ( 12)+3 ( 12)=6 ;
for A 2, 1(12)+2 ( 12)+3 (−1)2=6 ;
for E , 1(22)+2 (−1)2+3 ( 02)=6 .
We next prove a theorem that is concerned not with a summation within an irrep
but with a summation over irreps.
Theorem. Ifnµis the dimension of the µth irrep of a group Gthen
summationdisplay
µn2
µ=g,
where gis the order of the group.
Proof. Define a representation of the group in the following way. Rearrange
the rows of the multiplication table of the group so that whilst the elements ina particular order head the columns, their inverses in the same order head therows. In this arrangement of the g×gtable, the leading diagonal is entirely
occupied by the identity element. Then, for each element Xof the group, take as
representative matrix the multiplication-table array obtained by replacing Xby
1 and all other element symbols by 0. The matrices D
reg(X) so obtained form the
regular representation ofG;t h e ya r ee a c h g×g, have a single non-zero entry ‘1’
in each row and column and (as will be verified by a little experimentation) have
939
REPRESENTATION THEORY
(a)IA B
IIA B
AAB I
BBIA(b)IA B
IIA B
BBIA
AAB I
Table 25.3 ( a) The multiplication table of the cyclic group of order 3, and
(b) its reordering used to generate the regular representation of the group.
the same multiplication structure as the group Gitself, i.e. they form a faithful
representation of G.
Although not part of the proof, a simple example may help to make these
ideas more transparent. Consider the cyclic group of order 3. Its multiplicationtable is shown in table 25.3( a) (a repeat of table 24.10( a) of the previous chapter),
whilst table 25.3( b) shows the same table reordered so that the columns are
still labelled in the order I,A,Bbut the rows are now labelled in the order
I
−1=I, A−1=B, B−1=A. The three matrices of the regular representation are
then
Dreg(I)=
100
010001
,D
reg(A)=
010
001100
,D
reg(B)=
001
100010
.
An alternative, more mathematical, definition of the regular representation of a
group is
bracketleftbig
D
reg(Gk)bracketrightbig
ij=braceleftBigg
1i f GkGj=Gi,
0o t h e r w i s e .
We now return to the proof. With the construction given, the regular representa-
t i o nh a sc h a r a c t e r sa sf o l l o w s :
χreg(I)=g, χreg(X)=0 i f X/negationslash=I.
We now apply (25.18) to Dregto obtain for the number mµof times that the irrep
ˆD(µ)appears in Dreg(see 25.11))
mµ=1
gsummationdisplay
Xbracketleftbig
χ(µ)(X)bracketrightbig∗χreg(X)=1
gbracketleftbig
χ(µ)(I)bracketrightbig∗χreg(I)=1
gnµg=nµ.
Thus an irrep ˆD(µ)of dimension nµappears nµtimes in Dreg, and so by counting
the total number of basis functions, or by considering χreg(I), we can conclude
940
25.7 COUNTING IRREPS USING CHARACTERS
that
summationdisplay
µn2
µ=g. (25.21)
This completes the proof.
As before, our standard demonstration group 3 mprovides an illustration. In
this case we have seen already that there are two one-dimensional irreps and onetwo-dimensional irrep. This is in accord with (25.21) since
1
2+12+22=6,which is the order gof the group.
Another straightforward application of the relation (25.21), to the group with
multiplication table 25.3( a), yields immediate results. Since g=3 ,n o n eo fi t s
irreps can have dimension 2 or more, as 22= 4 is too large for (25.21) to be
satisfied. Thus all irreps must be one-dimensional and there must be three ofthem (consistent with the fact that each element is in a class of its own, and thatthere are therefore three classes). The three irreps are the sets of 1 ×1 matrices
(numbers)
A
1={1,1,1}A2={1,ω,ω2}A∗
2={1,ω2,ω},
where ω=e x p ( 2 πi/3); since the matrices are 1 ×1, the same set of nine numbers
would be, of course, the entries in the character table for the irreps of the group.
The fact that the numbers in each irrep are all cube roots of unity is discussed
below. As will be noticed, two of these irreps are complex – an unusual occurrencein most applications – and form a complex conjugate pair of one-dimensionalirreps. In practice, they function much as a two-dimensional irrep, but this is tobe ignored for formal purposes such as theorems.
A further property of characters can be derived from the fact that all elements
in a conjugacy class have the same order. Suppose that the element Xhas order
m,i . e .X
m=I. This implies for a representation Dof dimension nthat
[D(X)]m=In. (25.22)
Representations equivalent to Dare generated as before by using similarity
transformations of the form
DQ(X)=Q−1D(X)Q.
In particular, if we choose the columns of Qto be the eigenvectors of D(X) then,
as discussed in chapter 8,
DQ(X)=
λ
10···0
0λ2...
......0
0···0 λn
941
REPRESENTATION THEORY
where the λiare the eigenvalues of D(X). Therefore, from (25.22), we have that
λm
10···0
0λm
2...
......0
0···0 λm
n
=
10 ···0
01...
......0
0···01
.
Hence all the eigenvalues λ
iaremth roots of unity, and so χ(X), the trace of
D(X), is the sum of nof these. In view of the implications of Lagrange’s theorem
(section 24.6 and subsection 24.7.1), the only values of mallowed are the divisors
of the order gof the group.
25.8 Construction of a character table
In order to decompose representations into irreps on a routine basis using
characters, it is necessary to have available a character table for the group inquestion. Such a table gives, for each irrep µof the group, the character χ
(µ)(X)
of the class to which group element Xbelongs. To construct such a table the
following properties of a group, established earlier in this chapter, may be used:
(i) the number of classes equals the number of irreps;
(ii) the ‘vector’ formed by the characters from a given irrep is orthogonal to
the ‘vector’ formed by the characters from a different irrep;
(iii)summationtext
µn2
µ=g,w h e r e nµis the dimension of the µth irrep and gis the order
of the group;
(iv) the identity irrep (one-dimensional with all characters equal to 1) is present
for every group;
(v)summationtext
Xvextendsinglevextendsingleχ(µ)(X)vextendsinglevextendsingle2=g.
(vi)χ(µ)(X)i st h es u mo f nµmth roots of unity, where mis the order of X.IConstruct the character table for the group 4mm(orC4v) using the properties of classes,
irreps and characters so far established.
The group 4 mmis the group of two-dimensional symmetries of a square, namely rotations
of 0, π/2,πand 3 π/2 and reflections in the mirror planes parallel to the coordinate axes
and along the main diagonals. These are illustrated in figure 25.3. For this group there areeight elements:
•the identity, I;
•rotations by π/2a n d3 π/2,RandR
/prime;
•ar o t a t i o nb y π,Q;
•four mirror reflections mx,my,mdandmd/prime.
Requirements (i) to (iv) at the start of this section put tight constraints on the possible
character sets, as the following argument shows.
The group is non-Abelian (clearly Rm x/negationslash=mxR), and so there are fewer than eight
classes, and hence fewer than eight irreps. But requirement (iii), with g= 8, then implies
942
25.8 CONSTRUCTION OF A CHARACTER TABLE
mx
mymd m/prime
d
Figure 25.3 The mirror planes associated with 4 mm,t h eg r o u po ft w o -
dimensional symmetries of a square.
that at least one irrep has dimension 2 or gr eater. However, there can be no irrep with
dimension 3 or greater, since 32>8, nor can there be more than one two-dimensional
irrep, since 22+22= 8 would rule out a contribution to the sum in (iii) of 12from the
identity irrep, and this must be present. Thus the only possibility is one two-dimensionalirrep and, to make the sum in (iii) correct, four one-dimensional irreps.
Therefore using (i) we can now deduce that there are five classes. This same conclusion
can be reached by evaluating X
−1YXfor every pair of elements in G, as in the description
of conjugacy classes given in the previous chapter. However, it is tedious to do so andcertainly much longer than the above. The five classes are I,Q,{R,R
/prime},{mx,my},{md,md/prime}.
It is straightforward to show that only IandQcommute with every element of the
group, so they are the only elements in classes of their own. Each other class must haveat least 2 members, but, as there are three classes to accommodate 8 −2=6e l e m e n t s ,
there must be exactly 2 in each class. This does not pair up the remaining 6 elements, butdoes say that the five classes have 1, 1, 2, 2, and 2 elements. Of course, if we had startedby dividing the group into classes, we would know the number of elements in each classdirectly.
We cannot entirely ignore the group structure (though it sometimes happens that the
results are independent of the group structure – for example, all non-Abelian groups oforder 8 have the same character table!); thus we need to note in the present case thatm
2
i=Ifori=x, y, d ord/primeand, as can be proved directly, Rm i=miR/primefor the same four
values of label i. We also recall that for any pair of elements XandY,D(XY)=D(X)D(Y).
We may conclude the following for the one-dimensional irreps.
(a) In view of result (vi), χ(mi)=D(mi)=±1.
(b) Since R4=I, result (vi) requires that χ(R) is one of 1, i,−1,−i.B u t ,s i n c e
D(R)D(mi)=D(mi)D(R/prime), and the D(mi) are just numbers, D(R)=D(R/prime). Further
D(R)D(R)=D(R)D(R/prime)=D(RR/prime)=D(I)=1 ,
and so D(R)=±1= D(R/prime).
(c)D(Q)=D(RR)=D(R)D(R)=1 .
If we add this to the fact that the characters of the identity irrep A 1are all unity then we
can fill in those entries in character table 25.4 shown in bold.
Suppose now that the three missing entries in a one-dimensional irrep are p,qandr,
where each can only be ±1. Then, allowing for the numbers in each class, orthogonality
943
REPRESENTATION THEORY
4mm IQR ,R/primemx,mymd,md/prime
A1 11 1 1 1
A2 11 1−1−1
B1 11−11 −1
B2 11−1−11
E 2−20 0 0
Table 25.4 The character table deduced for the group 4 mm. For an explana-
tion of the entries in bold see the text.
with the characters of A 1requires that
1(1)(1) + 1(1)(1) + 2(1)( p) + 2(1)( q) + 2(1)( r)=0 .
The only possibility is that two of p,q,a n d requal−1 and the other equals +1. This
can be achieved in three different ways, corresponding to the need to find three furtherdifferent one-dimensional irreps. Thus the first four lines of entries in character table 25.4can be completed. The final line can be completed by requiring it to be orthogonal to theother four. Property (v) has not been used here though it could have replaced part of theargument given.J
25.9 Group nomenclature
The nomenclature of published character tables, as we have said before, is erratic
and sometimes unfortunate; for example, often Eis used to represent, not only
a two-dimensional irrep, but also the identity operation, where we have used I.
Thus the symbol Emight appear in both the column and row headings of a
table, though with quite different meanings in the two cases. In this book we useroman capitals to denote irreps.
One-dimensional irreps are regularly denoted by A and B, B being used if a
rotation about the principal axis of 2 π/nhas character −1. Here nis the highest
integer such that a rotation of 2 π/nis a symmetry operation of the system, and
the principal axis is the one about which this occurs. For the group of operations
on a square, n= 4, the axis is the perpendicular to the square and the rotation
in question is R. The names for the group, 4 mmandC
4v,d e r i v ef r o mt h ef a c t
that here nis equal to 4. Similarly, for the operations on an equilateral triangle,
n= 3 and the group names are 3 mandC3v, but because the rotation by 2 π/3h a s
character +1 in all its one-dimensional irreps (see table 25.1), only A appears inthe irrep list.
Two-dimensional irreps are denoted by E, as we have already noted, and three-
dimensional irreps by T, although in many cases the symbols are modified by
primes and other alphabetic labels to denote variations in behaviour from one
irrep to another in respect of mirror reflections and parity inversions. In the studyof molecules, alternative names based on molecular angular momentum properties
944
25.10 PRODUCT REPRESENTATIONS
are common. It is beyond the scope of this book to list all these variations, or to
give a large selection of character tables; our aim is to demonstrate and justifythe use of those found in the literature specifically dedicated to crystal physics ormolecular chemistry.
Variations in notation are not restricted to the naming of groups and their
irreps, but extend to the symbols used to identify a typical element, and henceall members, of a conjugacy class in a group. In physics these are usually of thetypes n
z,¯nzormx. The first of these denotes a rotation of 2 π/nabout the z-axis,
and the second the same thing followed by parity inversion (all vectors rgo to
−r), whilst the third indicates a mirror reflection in a plane, in this case the plane
x=0 .
Typical chemistry symbols for classes are NC n,NC2
n,NCx
n,NSn,σv,σxy.H e r e
the first symbol N, where it appears, shows that there are Nelements in the
class (a useful feature). The subscript nhas the same meaning as in the physics
notation, but σrather than mis used for a mirror reflection, subscripts v,dorhor
superscripts xy,xzoryzdenoting the various orientations of the relevant mirror
planes. Symmetries involving parity inversions are denoted by S; thus Snis the
chemistry analogue of ¯n. None of what is said in this and the previous paragraph
should be taken as definitive, but merely as a warning of common variations innomenclature and as an initial guide to corresponding entities. Before using anyset of group character tables, the reader should ensure that he or she understands
the precise notation being employed.
25.10 Product representations
In quantum mechanical investigations we are often faced with the calculation of
what are called matrix elements. These normally take the form of integrals over allspace of the product of two or more functions whose analytic forms depend on themicroscopic properties (usually angular momentum and its components) of the
electrons or nuclei involved. For ‘bonding’ calculations involving ‘overlap integrals’
there are usually two functions involved, whilst for transition probabilities a thirdfunction, giving the spatial variation of the interaction Hamiltonian, also appearsunder the integral sign.
If the environment of the microscopic system under investigation has some
symmetry properties, then sometimes these can be used to establish, without
detailed evaluation, that the multiple integral must have zero value. We nowexpress the essential content of these ideas in group theoretical language.
Suppose we are given an integral of the form
J=integraldisplay
Ψφd τ or J=integraldisplay
Ψξφdτ
to be evaluated over all space in a situation in which the physical system is
945
REPRESENTATION THEORY
invariant under a particular group Gof symmetry operations. For the integral to
be non-zero the integrand must be invariant under each of these operations. Ingroup theoretical language, the integrand must transform as the identity, the one-
dimensional representation A
1ofG; more accurately, some non-vanishing part of
the integrand must do so.
An alternative way of saying this is that if under the symmetry operations
ofGthe integrand transforms according to a representation Dand Ddoes not
contain A 1amongst its irreps then the integral Jis necessarily zero. It should be
noted that the converse is not true; Jm a yb ez e r oe v e ni fA 1is present, since the
integral, whilst showing the required invariance, may still have the value zero.
It is evident that we need to establish how to find the irreps that go to make
up a representation of a double or triple product when we already know theirreps according to which the factors in the product transform. The method isestablished by the following theorem.
Theorem. For each element of a group the character in a product representation is
the product of the corresponding characters in the separate representations.
Proof. Suppose that {u
i}and{vj}are two sets of basis functions, that transform
under the operations of a group Gaccording to representations D(λ)and D(µ)
respectively. Denote by uand vthe corresponding basis vectors and let Xbe an
element of the group. Then the functions generated from uiandvjby the action
ofXare calculated as follows, using (25.1) and (25.4):
u/prime
i=Xui=bracketleftBigparenleftbig
D(λ)(X)parenrightbigTubracketrightBig
i=bracketleftbig
D(λ)(X)bracketrightbig
iiui+summationdisplay
l/negationslash=ibracketleftBigparenleftbig
D(λ)(X)parenrightbigTbracketrightBig
ilul,
v/prime
j=Xvj=bracketleftBigparenleftbig
D(µ)(X)parenrightbigTvbracketrightBig
j=bracketleftbig
D(µ)(X)bracketrightbig
jjvj+summationdisplay
m/negationslash=jbracketleftBigparenleftbig
D(µ)(X)parenrightbigTbracketrightBig
jmvm.
Here[D(X)]ijis just a single element of the matrix D(X)a n d[ D(X)]kk=[DT(X)]kk
is simply a diagonal element from the matrix – the repeated subscript does not
indicate summation. Now, if we take as basis functions for a product represen-tation D
prod(X) the products wk=uivj(where the nλnµvarious possible pairs of
values i,jare labelled by k), we have also that
w/prime
k=Xw k=Xuivj=(Xui)(Xvj)
=bracketleftbig
D(λ)(X)bracketrightbig
iibracketleftbig
D(µ)(X)bracketrightbig
jjuivj+ terms not involving the product uivj.
This is to be compared with
w/prime
k=Xw k=bracketleftBigparenleftbig
Dprod(X)parenrightbigTwbracketrightBig
k=bracketleftbig
Dprod(X)bracketrightbig
kkwk+summationdisplay
n/negationslash=kbracketleftBigparenleftbig
Dprod(X)parenrightbigTbracketrightBig
knwn,
where Dprod(X) is the product representation matrix for element Xof the group.
946
25.11 PHYSICAL APPLICATIONS OF GROUP THEORY
The comparison shows that
bracketleftbig
Dprod(X)bracketrightbig
kk=bracketleftbig
D(λ)(X)bracketrightbig
iibracketleftbig
D(µ)(X)bracketrightbig
jj.
It follows that
χprod(X)=nλnµsummationdisplay
k=1bracketleftbig
Dprod(X)bracketrightbig
kk
=nλsummationdisplay
i=1nµsummationdisplay
j=1bracketleftbig
D(λ)(X)bracketrightbig
iibracketleftbig
D(µ)(X)bracketrightbig
jj
=braceleftBiggnλsummationdisplay
i=1bracketleftbig
D(λ)(X)bracketrightbig
iibracerightBiggbraceleftBiggnµsummationdisplay
j=1bracketleftbig
D(µ)(X)bracketrightbig
jjbracerightBigg
=χ(λ)(X)χ(µ)(X). (25.23)
This proves the theorem, and a similar argument leads to the corresponding result
for integrands in the form of a product of three or more factors.
An immediate corollary is that an integral whose integrand is the product of
two functions transforming according to two different irreps is necessarily zero .T o
see this, we use (25.18) to determine whether irrep A 1appears in the product
character set χprod(X):
mA1=1
gsummationdisplay
Xbracketleftbig
χ(A1)(X)bracketrightbig∗χprod(X)=1
gsummationdisplay
Xχprod(X)=1
gsummationdisplay
Xχ(λ)(X)χ(µ)(X).
We have used the fact that χ(A1)(X)=1f o ra l l Xbut now note that, by virtue of
(25.14), the expression on the right of this equation is equal to zero unless λ=µ.
Any complications due to non-real characters have been ignored – in practice,
they are handled automatically as it is usually Ψ∗φ, rather than Ψ φ, that appears
in integrands, though many functions are real in any case, and nearly all charactersare.
Equation (25.23) is a general result for integrands but, specifically in the context
of chemical bonding, it implies that for the possibility of bonding to exist, the
two quantum wavefunctions must transform according to the same irrep. This is
discussed further in the next section.
25.11 Physical applications of group theory
As we indicated at the start of chapter 24 and discussed in a little more detail at
the beginning of the present chapter, some physical systems possess symmetries
that allow the results of the present chapter to be used in their analysis. We
consider now some of the more common sorts of problem in which these resultsfind ready application.
947
REPRESENTATION THEORY
1
2
34xy
Figure 25.4 A molecule consisting of four atoms of iodine and one of
manganese.
25.11.1 Bonding in molecules
We have just seen that whether chemical bonding can take place in a molecule
is strongly dependent upon whether the wavefunctions of the two atoms forming
a bond transform according to the same irrep. Thus it is sometimes useful to beable to find a wavefunction that does transform according to a particular irrepof a group of transformations. This can be done if the characters of the irrep areknown and a sensible starting point can be guessed. We state without proof thatstarting from any n-dimensional basis vector Ψ ≡(Ψ
1Ψ2···Ψn)Twhere{Ψi}is
a set of wavefunctions, the new vector Ψ(λ)≡(Ψ(λ)
1Ψ(λ)
2···Ψ(λ)
n)Tgenerated by
Ψ(λ)
i=summationdisplay
Xχ(λ)∗(X)XΨi (25.24)
will transform according to the λth irrep. If the randomly chosen Ψ happens not
to contain any component that transforms in the desired way then the Ψ(λ)so
generated is found to be a zero vector and it is necessary to select a new startingvector. An illustration of the use of this ‘projection operator’ is given in the next
example.IConsider a molecule made up of four iodine atoms lying at the corners of a square in the
xy-plane, with a manganese atom at its centre, as shown in figure 25.4. Investigate whether
the molecular orbital given by the superposition of p-state (angular momentum l=1)
atomic orbitals
Ψ1=Ψ y(r−R1)+Ψ x(r−R2)−Ψy(r−R3)−Ψx(r−R4)
can bond to the d-state atomic orbitals of the mangane se atom described by either (a)
φ1=( 3z2−r2)f(r)or (b) φ2=(x2−y2)f(r),w h e r e f(r)is a function of rand so is
unchanged by any of the symmetry operations of th e molecule. Such linear combinations of
atomic orbitals are known as ring orbitals.
We have eight basis functions, the atomic orbitals Ψ x(N)a n dΨ y(N), where N=1,2,3,4
and indicates the position of an iodine atom. Since the wavefunctions are those of p-states
they have the forms xf(r)o ryf(r) and lie in the directions of the x-a n d y-a x e ss h o w ni n
the figure. Since ris not changed by any of the symmetry operations, f(r) can be treated as
a constant. The symmetry group of the system is 4 mm, whose character table is table 25.4.
Case (a). The manganese atomic orbital φ1=( 3z2−r2)f(r), lying at the centre of the
948
25.11 PHYSICAL APPLICATIONS OF GROUP THEORY
molecule, is not affected by any of the symmetry operations since zandrare unchanged
by them. It clearly transforms according to the identity irrep A 1. We therefore need to
know which combination of the iodine orbitals Ψ x(N)a n dΨ y(N), if any, also transforms
according to A 1.
We use the projection operator (25.24). If we choose Ψ x(1) as the arbitrary one-
dimensional starting vector, we unfortunately obtain zero (as the reader may wish toverify), but Ψ
y(1) does generate a new non-zero one-dimensional vector transforming
according to A 1. The results of acting on Ψ y(1) with the various symmetry elements X
can be written down by inspection (see the discussion in section 25.2). So, for example, theΨ
y(1) orbital centred on iodine atom 1 and aligned along the positive y-axis, is changed
by the anticlockwise rotation of π/2 produced by R/primeinto an orbital centred on atom 4
and aligned along the negative x-axis; thus R/primeΨy(1) =−Ψx(4). The complete set of group
actions on Ψ y(1) is:
I,Ψy(1); Q,−Ψy(3); R,Ψx(2); R/prime,−Ψx(4);
mx,Ψy(1); my,−Ψy(3); md,Ψx(2); md/prime,−Ψx(4).
Now χ(A1)(X)=1f o ra l l X, so (25.24) states that the sum of the above results for XΨy(1),
all with weight 1, gives a vector (in this case of a one-dimensional irrep, just a wave-function) that transforms according to A
1and is therefore capable of forming a chemical
bond with the manganese wavefunction φ1.I ti s
Ψ(A1)=2 [ Ψ y(1)−Ψy(3) + Ψ x(2)−Ψx(4)],
though, of course, the factor 2 is irrelevant. This is precisely the ring orbital Ψ 1given in
the problem, but here it is generated rather than guessed beforehand.
Case (b). The atomic orbital φ2=(x2−y2)f(r) behaves as follows under the action of
typical conjugacy class members:
I, φ 2;Q, φ 2;R,(y2−x2)f(r)=−φ2;mx,φ2;md,−φ2.
From this we see that φ2transforms as a one-dimensional irrep, but, from table 25.4, that
irrep is B 1not A 1(the irrep according to which Ψ 1transforms, as already shown). Thus
φ2and Ψ 1cannot form a bond.
J
The original question did not ask for the the ring orbital to which φ2may
bond, but it can be generated easily by using the values of XΨy(1) calculated in
case (a) but now weighting them according to the characters of B1:
Ψ(B1)=Ψ y(1)−Ψy(3) + (−1)Ψ x(2)−(−1)Ψ x(4)
+Ψ y(1)−Ψy(3) + (−1)Ψ x(2)−(−1)Ψ x(4)
=2 [ Ψ y(1)−Ψx(2)−Ψy(3) + Ψ x(4)].
Now we will find the other irreps of 4 mmpresent in the space spanned by
the basis functions Ψ x(N)a n dΨ y(N); at the same time this will illustrate the
important point that since we are working with characters we are only interestedin the diagonal elements of the representative matrices. This means (section 25.2)that if we work in the natural representation D
natwe need consider only those
functions that transform, wholly or partially, into themselves. Since we have no
need to write out the matrices explicitly, their size (8 ×8) is no drawback. All the
irreps spanned by the basis functions Ψ x(N)a n dΨ y(N)c a nb ed e t e r m i n e db y
considering the actions of the group elements upon them, as follows.
949
REPRESENTATION THEORY
(i) Under Iall eight basis functions are unchanged, and χ(I)=8 .
(ii) The rotations R,R/primeandQchange the value of Nin every case and so
all diagonal elements of the natural representation are zero and χ(R)=
χ(Q)=0 .
(iii)mxtakes xinto−xandyintoyand, for N= 1 and 3, leaves Nunchanged,
with the consequences (remember the forms of Ψ x(N)a n dΨ y(N)) that
Ψx(1)→−Ψx(1),Ψx(3)→−Ψx(3),
Ψy(1)→Ψy(1),Ψy(3)→Ψy(3).
Thus χ(mx) has four non-zero contributions, −1,−1, 1 and 1, together
with four zero contributions. The total is thus zero.
(iv)mdandmd/primeleave no atom unchanged and so χ(md)=0 .
The character set of the natural representation is thus 8, 0, 0, 0, 0, which, either
by inspection or by applying formula (25.18), shows that
Dnat=A 1⊕A2⊕B1⊕B2⊕2E,
i.e. that all possible irreps are present. We have constructed previously the
combinations of Ψ x(N)a n dΨ y(N) that transform according to A 1and B 1.
The others can be found in the same way.
25.11.2 Matrix elements in quantum mechanics
In section 25.10 we outlined the procedure for determining whether a matrix
element that involves the product of three factors as an integrand is necessarily
zero. We now illustrate this with a specific worked example.IDetermine whether a ‘dipole’ matrix element of the form
J=
Z
Ψd1xΨd2dτ,
where Ψd1andΨd2ared-state wavefunctions of the forms xyf(r)and(x2−y2)g(r)respec-
tively, can be non-zero (i) in a molecule with symmetry C3v(or3m), such as ammonia, and
(ii) in a molecule with symmetry C4v(or4mm), such as the MnI 4molecule considered in
the previous example.
We will need to make reference to the character tables of the two groups. The table forC
3vis table 25.1 (section 25.6); that for C4vis reproduced as table 25.5 from table 25.4 but
with the addition of another column showing how some common functions transform.
We make use of (25.23), extended to the product of three functions. No attention need
be paid to f(r)a n d g(r) as they are unaffected by the group operations.
Case (a). From the character table 25.1 for C3v, we see that each of xy,xandx2−y2
forms part of a basis set transforming according to the two-dimensional irrep E. Thus we
may fill in the array of characters (using chemical notation for the classes, except thatwe continue to use Irather than E) as shown in table 25.6. The last line is obtained by
950
25.11 PHYSICAL APPLICATIONS OF GROUP THEORY
4mm IQR ,R/primemx,mymd,md/prime
A1 11 1 1 1 z;z2;x2+y2
A2 11 1 −1−1 Rz
B1 11−11 −1 x2−y2
B2 11−1−11 xy
E 2−20 0 0 (x, y); (xz, yz); (Rx,Ry)
Table 25.5 The character table for the irreps of group 4 mm(orC4v). The
right-hand column lists some common func tions, or, for the two-dimensional
irrep E, pairs of functions, that transform according to the irrep against which
they are shown.
Function Irrep Classes
I2C33σv
xy E2−10
x E2−10
x2−y2E2−10
product 8 −10
Table 25.6 The character sets, for the group C3v(or 3mm), of three functions
and of their product x2y(x2−y2).
Function Irrep Classes
IC 22C62σv2σd
xy B2 11−1−11
x E2 −20 0 0
x2−y2B1 11−11−1
product 2 −20 0 0
Table 25.7 The character sets, for the group C4v(or 4mm), of three functions,
and of their product x2y(x2−y2).
multiplying together the corresponding characters for each of the three elements. Now, by
inspection, or by applying (25.18), i.e.
mA1=1
6[1(1)(8) + 2(1)( −1) + 3(1)(0)] = 1 ,
we see that irrep A 1does appear in the reduced representation of the product, and so J
is not necessarily zero.
Case (b). From table 25.5 we find that, under the group C4v,xyandx2−y2transform
as irreps B 2and B 1respectively and that xis part of a basis set transforming as E. Thus
the calculation table takes the form of table 25.7 (again, chemical notation for the classeshas been used).
Here inspection is sufficient, as the product is exactly that of irrep E and irrep A
1is
certainly not present. Thus Jis necessarily zero and the dipole matrix element vanishes.
J
951
REPRESENTATION THEORY
x1x2x3
y1 y2y3
Figure 25.5 An equilateral array of masses and springs.
25.11.3 Degeneracy of normal modes
As our final area for illustrating the usefulness of group theoretical results we
consider the normal modes of a vibrating system (see chapter 9). This analysishas far-reaching applications in physics, chemistry and engineering. For a givensystem, normal modes that are related by some symmetry operation have the samefrequency of vibration; the modes are said to be degenerate .I tc a nb es h o w nt h a t
such modes span a vector space that transforms according to some irrep of thegroup Gof symmetry operations of the system. Moreover, the degeneracy of
the modes equals the dimension of the irrep. As an illustration, we consider the
following example.IInvestigate the possible vibrational modes o f the equilateral triangular arrangement of
equal masses and springs shown in figure 25. 5. Demonstrate that two are degenerate.
Clearly the symmetry group is that of the symme try operations on an equilateral triangle,
namely 3 m(orC3v), whose character table is table 25.1. As on a previous occasion, it is
most convenient to use the natural representation Dnatof this group (it almost always
saves having to write out matrices explicitly) acting on the six-dimensional vector space(x
1,y1,x2,y2,x3,y3). In this example the natural and regular representations coincide, but
this is not usually the case.
We note that in table 25.1 the second class contains the rotations A(byπ/3) and B(by
2π/3), also known as RandR/prime. This class is known as 3 zin crystallographic notation, or
C3in chemical notation, as explained in section 25.9. The third class contains C,D,E,t h e
three mirror reflections.
Clearly χ(I) = 6. Since all position labels are changed by a rotation, χ(3z)=0 .F o rt h e
mirror reflections the simplest representative class member to choose is the reflection myin
the plane containing the y3-axis, since then only label 3 is unchanged; under my,x3→−x3
andy3→y3, leading to the conclusion that χ(my) = 0. Thus the character set is 6, 0, 0.
Using (25.18) and the character table 25.1 shows that
Dnat=A 1⊕A2⊕2E.
952
25.11 PHYSICAL APPLICATIONS OF GROUP THEORY
However, we have so far allowed xi,yito be completely general, and we must now identify
and remove those irreps that do not correspond to vibrations. These will be the irrepscorresponding to bodily translations of the triangle and to its rotation without relativemotion of the three masses.
Bodily translations are linear motions of the centre of mass, which has coordinates
x=(x
1+x2+x3)/3a n d y=(y1+y2+y3)/3).
Table 25.1 shows that such a coordinate pair ( x, y) transforms according to the two-
dimensional irrep E; this accounts for one of the two such irreps found in the naturalrepresentation.
It can be shown that, as stated in table 25.1, planar bodily rotations of the triangle
– rotations about the z-axis, denoted by R
z– transform as irrep A 2. Thus, when the
linear motions of the centre of mass, and pure rotation about it, are removed from ourreduced representation, we are left with E ⊕A
1. These must be the irreps corresponding
to the internal vibrations of the triangle. – one doubly degenerate mode and one non-degenerate mode. The physical interpretation of this is that two of the normal modes of thesystem have the same frequency and one normal mode has a different frequency (barringaccidental coincidences for other reasons). It may be noted that in quantum mechanicsthe energy quantum of a normal mode is proportional to its frequency.J
In general, group theory does not tell us what the frequencies are, since it is
entirely concerned with the symmetry of the system and not with the values of
masses and spring constants. However, using this type of reasoning, the results
from representation theory can be used to predict the degeneracies of atomicenergy levels and, given a perturbation whose Hamiltonian (energy operator) hassome degree of symmetry, the extent to which the perturbation will resolve thedegeneracy. Some of these ideas are explored a little further in the next sectionand in the exercises.
25.11.4 Breaking of degeneracies
If a physical system has a high degree of symmetry, invariant under a group Gof
reflections and rotations, say, then, as implied above, it will normally be the casethat some of its eigenvalues (of energy, frequency, angular momentum etc.) are
degenerate. However, if a perturbation that is invariant only under the operations
of the elements of a smaller symmetry group (a subgroup of G)is added, some of
the original degeneracies may be broken. The results derived from representationtheory can be used to decide the extent of the degeneracy-breaking.
The normal procedure is to use an N-dimensional basis vector, consisting of
theNdegenerate eigenfunctions, to generate an N-dimensional representation of
the symmetry group of the perturbation. This representation is then decomposedinto irreps. In general, eigenfunctions that transform according to different irrepsno longer share the same frequency of vibration.
We illustrate this with the following example.
953
REPRESENTATION THEORY
M MM
Figure 25.6 A circular drumskin loaded with three symmetrically placed
masses.IA circular drumskin has three equal masses placed on it at the vertices of an equilateral
triangle, as shown in figure 25.6. Determine which degenerate normal modes of the drumskincan be split in frequency by this perturbation.
When no masses are present the normal modes of the drum-skin are either non-degenerateor two-fold degenerate (see chapter 19). The degenerate eigenfunctions Ψ of the nth normal
mode have the forms
J
n(kr)(cos nθ)e±iωtor Jn(kr)(sinnθ)e±iωt.
Therefore, as explained above, we need to consider the two-dimensional vector space
spanned by Ψ 1=s i n nθand Ψ 2=c o s nθ. This will generate a two-dimensional representa-
tion of the group 3 m(orC3v), the symmetry group of the perturbation. Taking the easiest
element from each of the three classes (identity, rotations, and reflections) of group 3 m,
we have
IΨ1=Ψ 1,IΨ2=Ψ 2,
AΨ1=s i n
/
n
/;
θ−2
3π
//
=
/;
cos2
3nπ
/
Ψ1−
/;
sin2
3nπ
/
Ψ2,
AΨ2=c o s
/
n
/;
θ−2
3π
//
=
/;
cos2
3nπ
/
Ψ2+
/;
sin2
3nπ
/
Ψ1,
CΨ1= sin[ n(π−θ)] =−(cosnπ)Ψ1,
CΨ2=c o s [ n(π−θ)] = (cos nπ)Ψ2.
The three representative matrices are therefore
D(I)=I2,D(A)=
/
cos2
3nπ−sin2
3nπ
sin2
3nπ cos2
3nπ
/!
,D(C)=
/
−cosnπ 0
0c o s nπ
/!
.
The characters of this representation are χ(I)=2 , χ(A)=2 c o s ( 2 nπ/3) and χ(C)=0 .
Using (25.18) and table 25.1, we find that
mA1=1
6
/;
2+4c o s2
3nπ
/
=mA2
mE=1
6
/;
4−4c os2
3nπ
/
.
Thus
D=
/(
A1⊕A2ifn=3,6,9,. . .,
E otherwise .
Hence the normal modes n=3,6,9,. . .each transform under the operations of 3 m
954
25.12 EXERCISES
as the sum of two one-dimensional irreps and, using the reasoning given in the previous
example, are therefore split in frequency by the perturbation. For other values of nthe
representation is irreducible and so the degeneracy cannot be split.
J
25.12 Exercises
25.1 A group Gh a sf o u re l e m e n t s I,X,Y andZ, which satisfy X2=Y2=Z2=
XY Z =I. Show that Gis Abelian and hence deduce the form of its character
table.
Show that the matrices
D(I)=
/
10
01
/
, D(X)=
/
−10
0−1
/
,
D(Y)=
/
−1−p
01
/
, D(Z)=
/
1p
0−1
/
,
where pis a real number, form a representation DofG. Find its characters and
decompose it into irreps.
25.2 Using a square whose corners lie at coordinates ( ±1,±1), form a natural rep-
resentation of the dihedral group D4. Find the characters of the representation,
and, using the information (and class order) in table 25.4 (p. 944), express therepresentation in terms of irreps.
Now form a representation in terms of eight 2 ×2 orthogonal matrices, by
considering the effect of each of the elements of D
4on a general vector ( x, y).
Confirm that this representation is one of the irreps found using the naturalrepresentation.
25.3 The quaternion group Q(see exercise 24.20) has eight elements {±1,±i,±j,±k}
obeying the relations
i
2=j2=k2=−1,i j=k=−ji.
Determine the conjugacy classes of Qand deduce the dimensions of its irreps.
Show that Qis homomorphic to the four-element group V, which is generated by
two distinct elements aandbwith a2=b2=(ab)2=I. Find the one-dimensional
irreps of Vand use these to help determine the full character table for Q.
25.4 (a) By considering the possible forms of its cycle notation, determine the number
of elements in each conjugacy class of the permutation group S4and show
thatS4has five irreps. Give the logical reasoning that shows they must consist
of two three-dimensional, one two-dimensional, and two one-dimensionalirreps.
(b) By considering the odd and even permutations in the group S
4establish the
characters for one of the one-dimensional irreps.
(c) Form a natural matrix representation of 4 ×4 matrices based on a set of
objects{a, b, c, d}, which may or may not be equal to each other, and, by
selecting one example from each conjugacy class, show that this natural rep-resentation has characters 4, 2, 1, 0, 0. The one-dimensional vector subspace
spanned by sets of the form {a, a, a, a}is invariant under the permutation
group and hence transforms according to the invariant irrep A
1.T h er e m a i n -
ing three-dimensional subspace is irreducible; use this and the charactersdeduced above to establish the characters for one of the three-dimensionalirreps, T
1.
(d) Complete the character table using orthogonality properties, and check the
summation rule for each irrep. You should obtain table 25.8.
955
REPRESENTATION THEORY
Typical element and class size
Irrep (1) (12) (123) (1234) (12)(34)
16 8 6 3
A1 11 1 1 1
A2 1−11 −11
E 20 −10 2
T1 31 0 −1−1
T2 3−10 1 −1
Table 25.8 The character table for the permutation group S4.
25.5 In exercise 24.10, the group of pure rotations taking a cube into itself was found
to have 24 elements. The group is isomorphic to the permutation group S4,
considered in the previous question, and hence has the same character table, oncecorresponding classes have been established. By counting the number of elementsin each class make the correspondences below (the final two cannot be decidedpurely by counting, and should be taken as given).
Permutation Symbol Action
class type (physics)
(1) I none
(123) 3 rotations about a body diagonal(12)(34) 2
z rotation of πabout the normal to a face
(1234) 4 z rotations of ±π/2 about the normal to a face
(12) 2 d rotation of πabout an axis through the
centres of opposite edges
Reformulate the character table 25.8 in terms of the elements of the rotation
symmetry group (432 or O) of a cube and use it when answering exercises 25.7
and 25.8.
25.6 Consider a regular hexagon orientated so that two of its vertices lie on the x-axis.
Find matrix representations of a rotation Rthrough π/6 and a reflection myin
they-axis by determining their effects on vectors lying in the xy-plane . Show
that a reflection mxin the x-axis can be written as mx=myR3and that the (12)
elements of the symmetry group of the hexagon are given by RnorRnmy.
Using the representations of Randmyas generators, find a two-dimensional
representation of the symmetry group, C6, of the regular hexagon. Is it a faithful
representation?
25.7 In a certain crystalline compound, a thorium atom lies at the centre of a regular
octahedron of six sulphur atoms at positions ( ±a,0,0), (0 ,±a,0), (0 ,0,±a). These
can be considered as being positioned at the centres of the faces of a cube ofside 2 a. The sulphur atoms produce at the site of the thorium atom an electric
field that has the same symmetry group as a cube (432 or O).
The five degenerate d-electron orbitals of the thorium atom can be expressed,
relative to any arbitrary polar axis, as
(3cos
2θ−1)f(r),e±iφsinθcosθf(r),e±2iφsin2θf(r).
A rotation about that polar axis by an angle φ/primeeffectively changes φtoφ−φ/prime.
Use this to show that the character of the rotation in a representation based onthe orbital wavefunctions is given by
1+2c o s φ
/prime+2c o s2 φ/prime
956
25.12 EXERCISES
and hence that the characters of the representation, in the order of the symbols
given in exercise 25.5, is 5, −1, 1,−1, 1. Deduce that the five-fold degenerate
level is split into two levels, a doublet and a triplet.
25.8 Sulphur hexafluoride is a molecule with the same structure as the crystalline
compound in exercise 25.7, except that a sulphur atom is now the central atom.The following are the forms of some of the electronic orbitals of the sulphuratom, together with the irreps according to which they transform under thesymmetry group 432 (or O).
Ψ
s=f(r)A 1
Ψp1=zf(r)T 1
Ψd1=( 3z2−r2)f(r)E
Ψd2=(x2−y2)f(r)E
Ψd3=xyf(r), T2
The function xtransforms according to the irrep T 1. Use the above data to
determine whether dipole matrix elements of the form J=
R
φ1xφ2dτcan be
non-zero for the following pairs of orbitals φ1,φ2in a sulphur hexafluoride
molecule: (a) Ψ d1,Ψs;( b )Ψ d1,Ψp1;( c )Ψ d2,Ψd1;( d )Ψ s,Ψd3;( e )Ψ p1,Ψs.
25.9 The hydrogen atoms in a methane molecule CH 4form a perfect tetrahedron
with the carbon atom at its centre. The molecule is most conveniently describedmathematically by placing the hydrogen atoms at the points (1 ,1,1), (1 ,−1,−1),
(−1,1,−1) and (−1,−1,1 ) .T h es y m m e t r yg r o u pt ow h i c hi tb e l o n g s ,t h et e t r a h e -
dral group ( ¯43morT
d) has classes typified by I,3 ,2 z,mdand¯4z, where the first
three are as in exercise 25.5, mdis a reflection in the mirror plane x−y=0a n d
¯4zis a rotation of π/2 about the z-axis followed by an inversion in the origin. A
reflection in a mirror plane can be considered as a rotation of πabout an axis
perpendicular to the plane, followed by an inversion in the origin.
T h ec h a r a c t e rt a b l ef o rt h eg r o u p ¯43mis very similar to that for the group
432, and has the form shown in table 25.9.
Typical element and class size Functions transforming
Irreps I32 z¯4z md according to irrep
18 3 6 6
A1 11 1 1 1 x2+y2+z2
A2 11 1 −1−1
E 2−12 0 0 (x2−y2,3z2−r2)
T1 30−11 −1 (Rx,Ry,Rz)
T2 30−1−11 (x, y, z); (xy, yz, zx )
Table 25.9 The character table for group ¯43m.
By following the steps given below, determine how many different internal vibra-
tion frequencies the CH 4molecule has.
(a) Consider a representation based on the 12 coordinates xi,yi,zifori=
1,2,3,4. For those hydrogen atoms that transform into themselves, a rota-
tion through an angle θabout an axis parallel to one of the coordinate axes
gives rise in the natural representation to the diagonal elements 1 for thecorresponding coordinate and 2cos θfor the two orthogonal coordinates. If
the rotation is followed by an inversion then these entries are multiplied by−1. Atoms not transforming into themselves give a zero diagonal contribu-
tion. Show that the characters of the natural representation are 12, 0, 0, 0, 2
957
REPRESENTATION THEORY
and hence that its expression in terms of irreps is
A1⊕E⊕T1⊕2T2.
(b) The irreps of the bodily translational and rotational motions are included in
this expression and need to be identified and removed. Show that when thisis done it can be concluded that there are three different internal vibrationfrequencies in the CH
4molecule. State their degeneracies and check that
they are consistent with the expected number of normal coordinates neededto describe the internal motions of the molecule.
25.10 (a) The set of even permutations of four objects (a proper subgroup of S
4)
is known as the alternating group A4. List its twelve members using cycle
notation.
(b) Assume that all permutations with the same cycle structure belong to the
same conjugacy class. Show that this leads to a contradiction and hencedemonstrates that even if two permutations have the same cycle structurethey do not necessarily belong to the same class.
(c) By evaluating the products p
1= (123)(4)•(12)(34)•(132)(4) and p2=
(132)(4)•(12)(34)•(123)(4) deduce that the three elements of A4with structure
of the form (12)(34) belong to the same class.
(d) By evaluating products of the form (1 α)(βγ)•(123)(4)•(1α)(βγ), where α, β, γ
are various combinations of 2, 3, 4, show that the class to which (123)(4)belongs contains at least four members. Show the same for (124)(3).
(e) By combining results (b), (c) and (d) deduce that A
4has exactly four classes,
and determine the dimensions of its irreps.
(f) Using the orthogonality properties of characters and noting that elements of
the form (124)(3) have order 3, find the character table for A4.
25.11 Use the results of exercise 24.23 to find the character table for the dihedral group
D5, the symmetry group of a regular pentagon.
25.12 Demonstrate that equation (25.24) does indeed generate a set of vectors trans-
forming according to an irrep λ, by sketching and superposing drawings of an
equilateral triangle of springs and masses, based on that shown in figure 25.7.
(a) (b) (c)A A A BB BC CC
30◦30◦
Figure 25.7 The three normal vibration modes of the equilateral array. Mode
(a) is known as the ‘breathing mode’. Modes ( b)a n d( c) transform according
to irrep E and have equal vibrational frequencies.
(a) Make an initial sketch showing an arbitrary small mass displacement from,
say, vertex C. Draw the results of operating on the initial sketch with each
of the symmetry elements of the group 3 m(C3v).
(b) Superimpose the results, weighting them according to the characters of irrep
A1(table 25.1 in section 25.6) and verify that the resultant is a symmetrical
arrangement in which all three masses move symmetrically towards (or awayfrom) the centroid of the triangle. The mode is illustrated in figure 25.7( a).
958
25.13 HINTS AND ANSWERS
(c) Start again, now considering a displacement δofCparallel to the x-axis.
Form a similar superposition of sketches weighted according to the charactersof irrep E (note that the reflections are not needed). The resultant containssome bodily displacement of the triangle, since this also transforms accordingto E. Show that the displacement of the centre of mass is ¯x=δ,¯y=0 .
Subtract this out and verify that the remainder is of the form shown infigure 25.7( c).
(d) Using an initial displacement parallel to the y-axis, and an analogous proce-
dure, generate the remaining normal mode, degenerate with that in ( c)a n d
shown in figure 25.7( b).
25.13 Further investigation of the crystalline compound considered in exercise 25.7
shows that the octahedron is not quite perfect but is elongated along the (1 ,1,1)
direction with the sulphur atoms at positions ±(a+δ,δ,δ),±(δ,a+δ,δ),±(δ,δ,a+
δ), where δ/lessmucha. This structure is invariant under the (crystallographic) symmetry
group 32 with three two-fold axes along directions typified by (1 ,−1,0). The
latter axes, which are perpendicular to the (1 ,1,1) direction, are axes of two-
fold symmetry for the perfect octahedron. The group 32 is really the three-dimensional version of the group 3 mand has the same character table as table 25.1
(section 25.6). Use this to show that, when the distortion of the octahedron isincluded, the doublet found in exercise 25.7 is unsplit but the triplet breaks upinto a singlet and a doublet.
25.13 Hints and answers
25.1 There are four classes and hence four one-dimensional irreps, which must have
e n t r i e sa sf o l l o w s :1 ,1 ,1 ,1 ; 1 ,1 , −1,−1; 1,−1, 1,−1; 1,−1,−1, 1. The
characters of Dare 2,−2, 0, 0 and so the irreps present are the last two of these.
25.2 The characters are 4, 0, 0, 0, 2, and the irreps present are A 1+B 2+E .T h e
characters of the classes are 2, −2, 0, 0, 0, showing that the representation is the
irrep E.
25.3 There are five classes {1},{−1},{±i},{±j},{±k}; there are four one-dimensional
irreps and one two-dimensional irrep. Show that ab=ba. The homomorphism
is±1→I,±i→a,±j→b,±k→ab.Vis Abelian and hence has four
one-dimensional irreps.
In the class order given above, the characters for Qare as follows: ˆD(1),1,1,1,1,1;
ˆD(2),1,1,1,−1,−1;ˆD(3),1,1,−1,1,−1;ˆD(4),1,1,−1,−1,1;ˆD(5),2,−2,0,0,0.
25.4 (a) One element of type (1)(2)(3)(4), six of type (12)(3)(4), eight of type (123)(4),
six of type (1234), three of type (12)(34). Five classes implies five irreps. SincePn2
imust equal 24, at least one ni≥3. Assuming ni≥4 leads to a contradiction,
and so n5(say) equals 3. The inequalities 12+3 ( 22)<15<4(22)i m p l yt h a t
a second niequals 3.
P3
1n2
i= 6 has only one integer solution. (b) D(2)[(12)] =
D(2)[(1234)] =−1. (c) Characters for T 1are (4−1), (2−1), (1−1), (0−1), (0−1),
i.e. 3, 1, 0,−1,−1.
25.6 The matrix representations are R=1
2[1,−√
3;√
3,1]; my=[−1,0;0,1]
As examples, R4=1
2[−1,√
3;−√
3,−1] and R2my=1
2[1,−√
3;−√
3,−1].
The representation is faithful.
25.7 The five basis functions of the representation are multiplied by 1, e−iφ/prime,e+iφ/prime,
e−2iφ/prime,e+2iφ/primeas a result of the rotation. The character is the sum of these for
rotations of 0, 2 π/3,π,π/2,π;Drep=E+T 2.
25.8 (a) No; (b) yes; (c) no; (d) no; (e) yes.
959
REPRESENTATION THEORY
25.9 (b) The bodily translation has irrep T 2and the rotation has irrep T 1. The irreps
of the internal vibrations are A 1,E ,T 2, with respective degeneracies 1, 2, 3,
making six internal coordinates (12 in total minus three translational minus threerotational).
25.10 (a) The identity, eight elements of the form (124)(3) and three elements of the
form (12)(34).
(b) The assumption implies that there are three irreps, of which one must be the
identity irrep. However, 1 + n
2
2+n2
3= 12 has no integer solutions.
(c)p1= (13)(24) ,p2= (14)(23).
(d) For example, (123)(4) generates (134)(2) ,(142)(3) ,(243)(1).
(e) There are three one-dimensional irreps and one three-dimensional irrep.
(f) The four sets of characters are: 1 ,1,1,1; 1,ω,ω2,1; 1,ω2,ω,1; 3,0,0,−1.
Here ω=e x p ( 2 πi/3) and 1 + ω+ω2=0 .
25.11 There are four classes and hence four irreps, which can only be the identity
irrep, one other one-dimensional irrep, and two two-dimensional irreps. In theclass order {I},{R,R
4},{R2,R3},{mi}the second one-dimensional irrep must
(because of orthogonality) have characters 1, 1, 1, −1. The summation rules and
orthogonality require the other two character sets to be 2 ,(−1+√5)/2,(−1−√5)/2,0a n d2 ,(−1−√5)/2,(−1+√5)/2,0. Note that Rhas order 5 and that,
e.g., (−1+√5)/2=e x p ( 2 πi/5) + exp(8 πi/5).
25.12 (c) ¯x=1
3[2δ+(−1)(−1
2δ)+(−1)(−1
2δ)],¯y=1
3[0 + (−1)(−√
3
2δ)+(−1)(√
3
2δ)].
25.13 The doublet irrep E (characters 2, −1, 0) appears in both 432 and 32 and so
is unsplit. The triplet T 2(characters 3, 0, 1) splits under 32 into doublet E
(characters 2, −1, 0) and singlet A 1(characters 1, 1, 1).
960
26
Probability
All scientists will know the importance of experiment and observation and,
equally, be aware that the results of some experiments depend to a degree onchance. For example, in an experiment to measure the heights of a random sampleof people, we would not be in the least surprised if all the heights were found to
be different; but, if the experiment were repeated often enough, we would expect
to find some sort of regularity in the results. Statistics, which is the subject of thenext chapter, is concerned with the analysis of real experimental data of this sort.First, however, we discuss probability. To a pure mathematician, probability is anentirely theoretical subject based on axioms. Although this axiomatic approach isimportant, and we discuss it briefly, an approach to probability more in keepingwith its eventual applications in statistics is adopted here.
We first discuss the terminology required, with particular reference to the
convenient graphical representation of experimental results as Venn diagrams.The concepts of random variables and distributions of random variables are thenintroduced. It is here that the connection with statistics is made; we assert thatthe results of many experiments are random variables and that those results havesome sort of regularity, which is represented by a distribution. Precise definitionsof a random variable and a distribution are then given, as are the defining
equations for some important distributions. We also derive some useful quantities
associated with these distributions.
26.1 Venn diagrams
We call a single performance of an experiment a trialand each possible result
anoutcome .T h e sample space Sof the experiment is then the set of all possible
outcomes of an individual trial. For example, if we throw a six-sided die then there
are six possible outcomes that together form the sample space of the experiment.At this stage we are not concerned with how likely a particular outcome might
961
PROBABILITY
ii iiii
ivA B
S
Figure 26.1 A Venn diagram.
be (we will return to the probability of an outcome in due course) but rather
will concentrate on the classification of possible outcomes. It is clear that some
sample spaces are finite (e.g. the outcomes of throwing a die) whilst others are
infinite (e.g. the outcomes of measuring people’s heights). Most often, one is notinterested in individual outcomes but in whether an outcome belongs to a givensubset A(say) of the sample space S; these subsets are called events . For example,
we might be interested in whether a person is taller or shorter than 180 cm, inwhich case we divide the sample space into just two events: namely, that theoutcome (height measured) is (i) greater than 180 cm or (ii) less than 180 cm.
A common graphical representation of the outcomes of an experiment is the
Venn diagram . A Venn diagram usually consists of a rectangle, the interior of
which represents the sample space, together with one or more closed curves insideit. The interior of each closed curve then represents an event. Figure 26.1 showsa typical Venn diagram representing a sample space Sand two events Aand
B. Every possible outcome is assigned to an appropriate region; in this example
there are four regions to consider (marked i to iv in figure 26.1):
(i) outcomes that belong to event Abut not to event B;
(ii) outcomes that belong to event Bbut not to event A;
(iii) outcomes that belong to both event Aand event B;
(iv) outcomes that belong to neither event Anor event B.IA six-sided die is thrown. Let event Abe ‘the number obtained is divisible by 2’ and event
Bbe ‘the number obtained is divisible by 3’. Draw a Venn diagram to represent these events.
It is clear that the outcomes 2, 4, 6 belong to event Aand that the outcomes 3, 6 belong
to event B. Of these, 6 belongs to both AandB. The remaining outcomes, 1, 5, belong to
neither AnorB. The appropriate Venn diagram is shown in figure 26.2.
J
In the above example, one outcome, 6, is divisible by both 2 and 3 and so
belongs to both AandB. This outcome is placed in region iii of figure 26.1, which
is called the intersection ofAandBand is denoted by A∩B(see figure 26.3( a)).
If no events lie in the region of intersection then AandBare said to be mutually
exclusive ordisjoint . In this case, often the Venn diagram is drawn so that the
closed curves representing the events AandBdo not overlap, so as to make
962
26.1 VENN DIAGRAMS
A B
S12 34
56
Figure 26.2 The Venn diagram for the outcomes of the die-throwing trials
described in the worked example.
AAA
ABBB
SSS
S(a)( b)
(c)( d)¯A
Figure 26.3 Venn diagrams: the shaded regions show ( a)A∩B,t h ei n t e r -
section of two events AandB,(b)A∪B, the union of events AandB,(c)
the complement ¯Aof an event A,(d)A−B, those outcomes in Athat do not
belong to B.
graphically explicit the fact that AandBare disjoint. It is not necessary, however,
to draw the diagram in this way, since we may simply assign zero outcomes to
the shaded region in figure 26.3(a). An event that contains no outcomes is calledtheempty event and denoted by ∅. The event comprising all the elements that
belong to either AorB, or to both, is called the union ofAandBand is denoted
byA∪B(see figure 26.3( b)). In the previous example, A∪B={2,3,4,6}.
It is sometimes convenient to talk about those outcomes that do notbelong to
a particular event. The set of outcomes that do not belong to Ais called the
complement ofAand is denoted by ¯A(see figure 26.3( c)); this can also be written
as¯A=S−A. It is clear that A∪¯A=SandA∩¯A=∅.
The above notation can be extended in an obvious way, so that A−Bdenotes
the outcomes in Athat do not belong to B. It is clear from figure 26.3( d)t h a t
A−Bc a na l s ob ew r i t t e na s A∩¯B. Finally, when allthe outcomes in event B
(say) also belong to event A, but Amay contain, in addition, outcomes that do
963
PROBABILITY
AB
C
S12
34 5
678
Figure 26.4 The general Venn diagram for three events is divided into eight
regions.
not belong to B,t h e n Bis called a subset ofA, a situation that is denoted by
B⊂A; alternatively, one may write A⊃B, which states that Acontains B.I nt h i s
case, the closed curve representing the event Bis often drawn lying completely
within the closed curve representing the event A.
The operations ∪and∩are extended straightforwardly to more than two
events. If there exist nevents A1,A2,...,A n, in some sample space S, then the
event consisting of all those outcomes that belong to one or more of the Aiis the
union ofA1,A2,...,A nand is denoted by
A1∪A2∪···∪An. (26.1)
Similarly, the event consisting of all the outcomes that belong to every one of the
Aiis called the intersection ofA1,A2,...,A nand is denoted by
A1∩A2∩···∩An. (26.2)
If, for anypair of values i, jwith i/negationslash=j,
Ai∩Aj=∅ (26.3)
then the events AiandAjare said to be mutually exclusive ordisjoint .
Consider three events A,BandCwith a Venn diagram such as is shown in
figure 26.4. It will be clear that, in general, the diagram will be divided into eight
regions and they will be of four different types. Three regions correspond to a
single event; three regions are each the intersection of exactly two events; oneregion is the three-fold intersection of all three events; and finally one regioncorresponds to none of the events. Let us now consider the numbers of differentregions in a general n-event Venn diagram.
For one-event Venn diagrams there are two regions, for the two-event case
there are four regions and, as we have just seen, for the three-event case there areeight. In the general n-event case there are 2
nregions, as is clear from the fact
that any particular region Rlies either inside or outside the closed curve of any
particular event. With two choices (inside or outside) for each of nclosed curves,
there are 2ndifferent possible combinations with which to characterise R.O n c e n
964
26.1 VENN DIAGRAMS
gets beyond three it becomes impossible to draw a simple two-dimensional Venn
diagram, but this does not change the results.
The 2nregions will break down into n+1 types, with the numbers of each type
as follows†
no events,nC0=1 ;
one event but no intersections,nC1=n;
two-fold intersections,nC2=1
2n(n−1);
three-fold intersections,nC3=1
3!n(n−1)(n−2);
...
ann-fold intersection,nCn=1 .
That this makes a total of 2ncan be checked by considering the binomial
expansion
2n=( 1+1 )n=1+ n+1
2n(n−1) +···+1.
Using Venn diagrams, it is straightforward to show that the operations ∩and
∪obey the following algebraic laws:
commutativity, A∩B=B∩A, A∪B=B∪A;
associativity, ( A∩B)∩C=A∩(B∩C),(A∪B)∪C=A∪(B∪C);
distributivity, A∩(B∪C)=(A∩B)∪(A∩C),
A∪(B∩C)=(A∪B)∩(A∪C);
idempotency, A∩A=A, A∪A=A.IShow that (i) A∪(A∩B)=A∩(A∪B)=A, (ii)(A−B)∪(A∩B)=A.
(i) Using the distributivity and idempotency laws above, we see that
A∪(A∩B)=(A∪A)∩(A∪B)=A∩(A∪B).
By sketching a Venn diagram it is immediately clear that both expressions are equal to
A. Nevertheless, we here proceed in a more formal manner in order to deduce this result
algebraically. Let us begin by writing
X=A∪(A∩B)=A∩(A∪B), (26.4)
from which we want to deduce a simpler expression for the event X. Using the first equality
in (26.4) and the algebraic laws for ∩and∪,w em a yw r i t e
A∩X=A∩[A∪(A∩B)]
=(A∩A)∪[A∩(A∩B)]
=A∪(A∩B)=X.
†The symbolsnCi,f o r i=0,1,2,...,n, are a convenient notation for combinations; they and their
properties are discussed in chapter 1.
965
PROBABILITY
Since A∩X=Xwe must have X⊂A. Now, using the second equality in (26.4) in a
similar way, we find
A∪X=A∪[A∩(A∪B)]
=(A∪A)∩[A∪(A∪B)]
=A∩(A∪B)=X,
from which we deduce that A⊂X. Thus, since X⊂AandA⊂X, we must conclude that
X=A.
(ii) Since we do not know how to deal with compound expressions containing a minus
sign, we begin by writing A−B=A∩¯Bas mentioned above. Then, using the distributivity
law, we obtain
(A−B)∪(A∩B)=(A∩¯B)∪(A∩B)
=A∩(¯B∪B)
=A∩S=A.
In fact, this result, like the first one, can be proved trivially by drawing a Venn diagram.
J
Further useful results may be derived from Venn diagrams. In particular, it is
simple to show that the following rules hold:
(i) if A⊂Bthen¯A⊃¯B;
(ii)A∪B=¯A∩¯B;
(iii)A∩B=¯A∪¯B.
Statements (ii) and (iii) are known jointly as de Morgan’s laws and are sometimes
useful in simplifying logical expressions.IThere exist two events AandBsuch that
(X∪A)∪(X∪¯A)=B.
Find an expression for the event Xin terms of AandB.
We begin by taking the complement of both sides of the above expression: applying de
Morgan’s laws we obtain
¯B=(X∪A)∩(X∪¯A).
We may then use the algebraic laws obeyed by ∩and∪to yield
¯B=X∪(A∩¯A)=X∪∅=X.
Thus, we find that X=¯B.
J
26.2 Probability
In the previous section we discussed Venn diagrams, which are graphical repre-
sentations of the possible outcomes of experiments. We did not, however, giveany indication of how likely each outcome or event might be when any particular
experiment is performed. Most experiments show some regularity. By this we
mean that the relative frequency of an event is approximately the same on eachoccasion that a set of trials is performed. For example, if we throw a die N
966
26.2 PROBABILITY
times then we expect that a six will occur approximately N/6 times (assuming,
of course, that the die is not biased). The regularity of outcomes allows us todefine the probability ,P r (A), as the expected relative frequency of event Ain a
large number of trials. More quantitatively, if an experiment has a total of n
S
outcomes in the sample space S,a n d nAof these outcomes correspond to the
event A, then the probability that event Awill occur is
Pr(A)=nA
nS. (26.5)
26.2.1 Basic theorems
From (26.5) we may deduce the following properties of the probability Pr( A).
(i) For any event Ain a sample space S,
0≤Pr(A)≤1. (26.6)
If Pr( A)=1t h e n Ais a certainty; if Pr( A)=0t h e n Ais an impossibility.
(ii) For the entire sample space Swe have
Pr(S)=nS
nS=1, (26.7)
which simply states that we are certain to obtain one of the possible
outcomes.
(iii) If AandBare two events in Sthen, from the Venn diagrams in figure 26.3,
we see that
nA∪B=nA+nB−nA∩B, (26.8)
the final subtraction arising because the outcomes in the intersection of
AandBare counted twice when the outcomes of Aare added to those
ofB. Dividing both sides of (26.8) by nS, we obtain the addition rule for
probabilities
Pr(A∪B)=P r ( A)+P r ( B)−Pr(A∩B). (26.9)
However, if Aand Baremutually exclusive events ( A∩B=∅)t h e n
Pr(A∩B) = 0 and we obtain the special case
Pr(A∪B)=P r ( A)+P r ( B). (26.10)
(iv) If ¯Ais the complement of Athen¯AandAare mutually exclusive events.
Thus, from (26.7) and (26.10) we have
1=P r ( S)=P r ( A∪¯A)=P r ( A)+P r ( ¯A),
from which we obtain the complement law
Pr(¯A)=1−Pr(A). (26.11)
967
PROBABILITY
This is particularly useful for problems in which evaluating the probability
of the complement is easier than evaluating the probability of the eventitself.ICalculate the probability of drawing an ace or a spade from a pack of cards.
LetAbe the event that an ace is drawn and Bthe event that a spade is drawn. It
immediately follows that Pr( A)=4
52=1
13and Pr( B)=13
52=1
4. The intersection of Aand
Bconsists of only the ace of spades and so Pr( A∩B)=1
52. Thus, from (26.9)
Pr(A∪B)=1
13+1
4−1
52=4
13.
In this case it is just as simple to recognise that there are 16 cards in the pack that satisfy
the required condition (13 spades plus three other aces) and so the probability is16
52.
J
The above theorems can easily be extended to a greater number of events. For
example, if A1,A2,...,A nare mutually exclusive events then (26.10) becomes
Pr(A1∪A2∪···∪An)=P r ( A1)+P r ( A2)+···+P r ( An).
(26.12)
Furthermore, if A1,A2,...,A n(whether mutually exclusive or not) exhaust S,i . e .
are such that A1∪A2∪···∪An=S,t h e n
Pr(A1∪A2∪···∪An)=P r ( S)=1 . (26.13)IA biased six-sided die has probabilities1
2p,p,p,p,p,2pof showing 1, 2, 3, 4, 5, 6
respectively. Calculate p.
Given that the individual events are mutually exclusive, (26.12) can be applied to give
Pr(1∪2∪3∪4∪5∪6) =1
2p+p+p+p+p+2p=13
2p.
The union of all possible outcomes on the LHS of this equation is clearly the sample
space, S,a n ds o
Pr(S)=13
2p.
Now using (26.7),
13
2p=P r ( S)=1⇒ p=2
13.
J
When the possible outcomes of a trial correspond to more than two events,
and those events are notmutually exclusive, the calculation of the probability of
the union of a number of events is more complicated, and the generalisation ofthe addition law (26.9) requires further work. Let us begin by considering theunion of three events A
1,A2andA3, which need not be mutually exclusive. We
first define the event B=A2∪A3and, using the addition law (26.9), we obtain
Pr(A1∪A2∪A3)=P r ( A1∪B)=P r ( A1)+P r ( B)−Pr(A1∩B).
(26.14)
968
26.2 PROBABILITY
However, we may write Pr( A1∩B)a s
Pr(A1∩B)=P r [ A1∩(A2∪A3)]
=P r [ ( A1∩A2)∪(A1∩A3)]
=P r ( A1∩A2)+P r ( A1∩A3)−Pr(A1∩A2∩A3).
Substituting this expression, and that for Pr( B) obtained from (26.9), into (26.14)
we obtain the probability addition law for three general events,
Pr(A1∪A2∪A3)=P r ( A1)+P r ( A2)+P r ( A3)−Pr(A2∩A3)−Pr(A1∩A3)
−Pr(A1∩A2)+P r ( A1∩A2∩A3). (26.15)ICalculate the probability of drawing from a pack of cards one that is an ace or is a spade
or shows an even number ( 2 ,4 ,6 ,8 ,1 0 ) .
If, as previously, Ais the event that an ace is drawn, Pr( A)=4
52. Similarly the event B,
that a spade is drawn, has Pr( B)=13
52. The further possibility C, that the card is even (but
not a picture card) has Pr( C)=20
52. The two-fold intersections have probabilities
Pr(A∩B)=1
52,Pr(A∩C)=0 ,Pr(B∩C)=5
52.
There is no three-fold intersection as events AandCare mutually exclusive. Hence
Pr(A∪B∪C)=1
52[(4 + 13 + 20) −( 1+0+5 )+( 0 ) ]=31
52.
The reader should identify the 31 cards involved.
J
When the probabilities are combined to calculate the probability for the union
of the ngeneral events, the result, which may be proved by induction upon n(see
the answer to exercise 26.4), is
Pr(A1∪A2∪···∪An)=summationdisplay
iPr(Ai)−summationdisplay
i,jPr(Ai∩Aj)+summationdisplay
i,j,kPr(Ai∩Aj∩Ak)
−···+(−1)n+1Pr(A1∩A2∩···∩An). (26.16)
Each summation runs over all possible sets of subscripts, except those in which
any two subscripts in a set are the same. The number of terms in the summationof probabilities of m-fold intersections of the nevents is given by
nCm(as discussed
in section 26.1). Equation (26.9) is a special case of (26.16) in which n=2a n d
only the first two terms on the RHS survive. We now illustrate this result with aworked example that has n= 4 and includes a four-fold intersection.
969
PROBABILITYIFind the probability of drawing from a pack a card that has at least one of the following
properties:
A,i ti sa na c e ;
B, it is a spade;
C, it is a black honour card (ace, king, queen, jack or 10);
D, it is a black ace.
Measuring all probabilities in units of1
52, the single-event probabilities are
Pr(A)=4 , Pr(B)=1 3 , Pr(C)=1 0 , Pr(D)=2 .
The two-fold intersection probabilities, measured in the same units, are
Pr(A∩B)=1 , Pr(A∩C)=2 , Pr(A∩D)=2 ,
Pr(B∩C)=5 , Pr(B∩D)=1 , Pr(C∩D)=2 .
The three-fold intersections have probabilities
Pr(A∩B∩C)=1 ,Pr(A∩B∩D)=1 ,Pr(A∩C∩D)=2 ,Pr(B∩C∩D)=1 .
Finally, the four-fold intersection, requiring all four conditions to hold, is satisfied only by
the ace of spades, and hence (again in units of1
52)
Pr(A∩B∩C∩D)=1 .
Substituting in (26.16) gives
P=1
52[( 4+1 3+1 0+2 ) −( 1+2+2+5+1+2 )+( 1+1+2+1 ) −(1)]=20
52.
J
We conclude this section on basic theorems by deriving a useful general
expression for the probability Pr( A∩B)t h a tt w oe v e n t s AandBboth occur in
the case where A(say) is the union of a set of nmutually exclusive events Ai.I n
this case
A∩B=(A1∩B)∪···∪(An∩B),
where the events Ai∩Bare also mutually exclusive. Thus, from the addition law
(26.12) for mutually exclusive events, we find
Pr(A∩B)=summationdisplay
iPr(Ai∩B). (26.17)
Moreover, in the special case where the events Aiexhaust the sample space S,w e
have A∩B=S∩B=B, and we obtain the total probability law
Pr(B)=summationdisplay
iPr(Ai∩B). (26.18)
26.2.2 Conditional probability
So far we have defined only probabilities of the form ‘what is the probability that
event Ahappens?’. In this section we turn to conditional probability , the probability
that a particular event occurs given the occurrence of another, possibly related,
event. For example, we may wish to know the probability of event B,d r a w i n ga n
970
26.2 PROBABILITY
ace from a pack of cards from which one has already been removed, given that
event A, the card already removed was itself an ace, has occurred.
We denote this probability by Pr( B|A) and may obtain a formula for it by
considering the total probability Pr( A∩B)=P r ( B∩A) that both AandBwill
occur. This may be written in two ways, i.e.
Pr(A∩B)=P r ( A)P r(B|A)
=P r ( B)P r (A|B).
From this we obtain
Pr(A|B)=Pr(A∩B)
Pr(B)(26.19)
and
Pr(B|A)=Pr(B∩A)
Pr(A). (26.20)
In terms of Venn diagrams, we may think of Pr( B|A) as the probability of Bin
the reduced sample space defined by A. Thus, if two events AandBare mutually
exclusive then
Pr(A|B)=0=P r ( B|A). (26.21)
When an experiment consists of drawing objects at random from a given set
of objects, it is termed sampling a population . We need to distinguish between
two different ways in which such a sampling experiment may be performed. After
an object has been drawn at random from the set it may either be put asideor returned to the set before the next object is randomly drawn. The former istermed ‘sampling without replacement’, the latter ‘sampling with replacement’.IFind the probability of drawing two aces at random from a pack of cards (i) when the
first card drawn is replaced at random into the pack before the second card is drawn, and(ii) when the first card is put aside after being drawn.
LetAbe the event that the first card is an ace, and Bthe event that the second card is an
ace. Now
Pr(A∩B)=P r ( A)P r(B|A),
and for both (i) and (ii) we know that Pr( A)=4
52=1
13.
(i) If the first card is replaced in the pack before the next is drawn then Pr( B|A)=
Pr(B)=4
52=1
13,s i n c e AandBare independent events. We then have
Pr(A∩B)=P r ( A)P r(B)=1
13×1
13=1
169.
(ii) If the first card is put aside and the second then drawn, AandBare not independent
and Pr( B|A)=3
51, with the result that
Pr(A∩B)=P r ( A)P r(B|A)=1
13×3
51=1
221.
J
971
PROBABILITY
Two events AandBarestatistically independent if Pr( A|B)=P r ( A) (or equiva-
lently if Pr( B|A)=P r ( B)). In words, the probability of Agiven Bis then the same
as the probability of Aregardless of whether Boccurs. For example, if we throw
a coin and a die at the same time, we would normally expect that the probability
of throwing a six was independent of whether a head was thrown. If AandBare
statistically independent then it follows that
Pr(A∩B)=P r ( A)P r(B). (26.22)
In fact, on the basis of intuition and experience, (26.22) may be regarded as the
definition of the statistical independence of two events.
The idea of statistical independence is easily extended to an arbitrary number
of events A1,A2,...,A n. The events are said to be (mutually) independent if
Pr(Ai∩Aj)=P r ( Ai)P r (Aj),
Pr(Ai∩Aj∩Ak)=P r ( Ai)P r (Aj)P r (Ak),
...
Pr(A1∩A2∩···∩An)=P r ( A1)P r(A2)···Pr(An),
for all combinations of indices i,jandkfor which no two indices are the same.
Even if all nevents are not mutually independent, any two events for which
Pr(Ai∩Aj)=P r ( Ai)P r(Aj) are said to be pairwise independent .
We now derive two results that often prove useful when working with condi-
tional probabilities. Let us suppose that an event Ais the union of nmutually
exclusive events Ai.I fBis some other event then from (26.17) we have
Pr(A∩B)=summationdisplay
iPr(Ai∩B).
Dividing both sides of this equation by Pr( B), and using (26.19), we obtain
Pr(A|B)=summationdisplay
iPr(Ai|B), (26.23)
which is the addition law for conditional probabilities .
Furthermore, if the set of mutually exclusive events Aiexhausts the sample
space Sthen, from the total probability law (26.18), the probability Pr( B) of some
event BinScan be written as
Pr(B)=summationdisplay
iPr(Ai)P r (B|Ai). (26.24)IA collection of traffic islands connected by a system of one-way roads is shown in fig-
ure 26.5. At any given island a car driver chooses a direction at random from those available.
What is the probability that a driver starting at Owill arrive at B?
In order to leave Othe driver must pass through one of A1,A2,A3orA4, which thus
form a complete set of mutually exclusive events. Since at each island (including O)t h e
driver chooses a direction at random from those available, we have that Pr( Ai)=1
4for
972
26.2 PROBABILITY
O
BA1
A2A3A4
Figure 26.5 A collection of traffic islands connected by one-way roads.
i=1,2,3,4. From figure 26.5, we see also that
Pr(B|A1)=1
3,Pr(B|A2)=1
3,Pr(B|A3)=0 ,Pr(B|A4)=2
4=1
2.
Thus, using the total probability law (26.24), we find that the probability of arriving at B
is given by
Pr(B)=
X
iPr(Ai)P r(B|Ai)=1
4
/;1
3+1
3+0+1
2
/
=7
24.
J
Finally, we note that the concept of conditional probability may be straightfor-
wardly extended to several compound events. For example, in the case of three
events A, B, C ,w em a yw r i t eP r ( A∩B∩C) in several ways, e.g.
Pr(A∩B∩C)=P r ( C)P r(A∩B|C)
=P r ( B∩C)P r(A|B∩C)
=P r ( C)P r(B|C)P r(A|B∩C).ISuppose{Ai}is a set of mutually exclusive events that exhausts the sample space S.I fB
andCare two other events in S, show that
Pr(B|C)=
X
iPr(Ai|C)P r(B|Ai∩C).
Using (26.19) and (26.17), we may write
Pr(C)P r(B|C)=P r ( B∩C)=
X
iPr(Ai∩B∩C). (26.25)
Each term in the sum on the RHS can be expanded as an appropriate product of
conditional probabilities,
Pr(Ai∩B∩C)=P r ( C)P r(Ai|C)P r(B|Ai∩C).
Substituting this form into (26.25) and dividing through by Pr( C) gives the required
result.
J
973
PROBABILITY
26.2.3 Bayes’ theorem
In the previous section we saw that the probability that both an event Aand a
related event Bwill occur can be written either as Pr( A)P r (B|A)o rP r ( B)P r(A|B).
Hence
Pr(A)P r (B|A)=P r ( B)P r (A|B),
from which we obtain Bayes’ theorem ,
Pr(A|B)=Pr(A)
Pr(B)Pr(B|A). (26.26)
This theorem clearly shows that Pr( B|A)/negationslash=P r ( A|B), unless Pr( A)=P r ( B). It is
sometimes useful to rewrite Pr( B), if it is not known directly, as
Pr(B)=P r ( A)P r(B|A)+P r ( ¯A)P r (B|¯A)
so that Bayes’ theorem becomes
Pr(A|B)=Pr(A)P r(B|A)
Pr(A)P r(B|A)+P r ( ¯A)P r (B|¯A). (26.27)ISuppose that the blood test for some disease is reliable in the following sense: for people
who are infected with the disease the test produces a positive result in 99.99% of cases; forpeople not infected a positive test result is obtained in only 0.02% of cases. Furthermore,assume that in the general population one person in 10000 people is infected. A person is
selected at random and found to test positive for the disease. What is the probability thatthe individual is actually infected?
LetAbe the event that the individual is infected and Bbe the event that the individual
tests positive for the disease. Using Bayes’ theorem the probability that a person who tests
positive is actually infected is
Pr(A|B)=Pr(A)P r(B|A)
Pr(A)P r(B|A)+Pr(¯A)Pr(B|¯A).
Now Pr( A)=1 /10000 = 1−Pr(¯A), and we are told that Pr( B|A) = 9999 /10000 and
Pr(B|¯A)=2 /10000. Thus we obtain
Pr(A|B)=1/10000×9999/10000
(1/10000×9999/10000) + (9999 /10000×2/10000)=1
3.
Thus, there is only a one in three chance that a person chosen at random, who tests
positive for the disease, is actually infected.
At a first glance, this answer may seem a little surprising, but the reason for the counter-
intuitive result is that the probability tha t a randomly selected person is not infected is
9999/10000, which is very high. Thus, the 0.02% chance of a match for an uninfected
person becomes significant.
J
We note that (26.27) may be written in a more general form if Sis not simply
974
26.3 PERMUTATIONS AND COMBINATIONS
divided into Aand¯Abut, rather, into anyset of mutually exclusive events Aithat
exhaust S. Using the total probability law (26.24), we may then write
Pr(B)=summationdisplay
iPr(Ai)P r (B|Ai),
so that Bayes’ theorem takes the form
Pr(A|B)=Pr(A)P r (B|A)summationtext
iPr(Ai)P r (B|Ai), (26.28)
where the event Aneed not coincide with any of the Ai.
As a final point, we comment that sometimes we are concerned only with the
relative probabilities of two events AandC(say), given the occurrence of some
other event B. From (26.26) we then obtain a different form of Bayes’ theorem,
Pr(A|B)
Pr(C|B)=Pr(A)P r (B|A)
Pr(C)P r (B|C), (26.29)
which does not contain Pr( B)a ta l l .
26.3 Permutations and combinations
In equation (26.5) we defined the probability of an event Ain a sample space S
as
Pr(A)=nA
nS,
where nAis the number of outcomes belonging to event AandnSis the total
number of possible outcomes. It is therefore necessary to be able to count thenumber of possible outcomes in various common situations.
26.3.1 Permutations
Let us first consider a set of nobjects that are all different. We may ask in
how many ways these nobjects may be arranged, i.e. how many permutations of
these objects exist. This is straightforward to deduce, as follows: the object in thefirst position may be chosen in ndifferent ways, that in the second position in
n−1 ways, and so on until the final object is positioned. The number of possible
arrangements is therefore
n(n−1)(n−2)···(1) = n! (26.30)
Generalising (26.30) slightly, let us suppose we choose only k(<n)o b j e c t s
from n. The number of possible permutations of these kobjects selected from n
is given by
n(n−1)(n−2)···(n−k+1 )bracehtipupleft
bracehtipdownrightbracehtipdownleft bracehtipupright
kfactors=n!
(n−k)!≡nPk. (26.31)
975
PROBABILITY
In calculating the number of permutations of the various objects we have so
far assumed that the objects are sampled without replacement – i.e. once an object
has been drawn from the set it is put aside. As mentioned previously, however,we may instead replace each object before the next is chosen. The number of
permutations of kobjects from nwith replacement may be calculated very easily
since the first object can be chosen in ndifferent ways, as can the second, the
third, etc. Therefore the number of permutations is simply n
k. This may also be
viewed as the number of permutations of kobjects from nwhere repetitions are
allowed, i.e. each object may be used as often as one likes.IFind the probability that in a group of kpeople, at least two have the same birthday
(ignoring 29February).
It is simplest to begin by calculating the probability that no two people share a birthday,
as follows. Firstly, we imagine each of the kpeople in turn pointing to their birthday on
a year planner. Thus, we are sampling the 365 days of the year ‘with replacement’ and
so the total number of possible outcomes is (365)k. Now, (for the moment) we assume
that no two people share a birthday and imagine the process being repeated, but as eachperson points out their birthday it is crossed off the planner. In this case, we are samplingthe days of the year ‘without replacement’, and so the possible number of outcomes forwhich all the birthdays are different is
365Pk=365!
(365−k)!.
Hence the probability that all the birthdays are different is
p=365!
(365−k)! 365k.
Now using the complement rule (26.11), the probability qthat two or more people have
the same birthday is simply
q=1−p=1−365!
(365−k)! 365k.
This expression may be conveniently evalutated using Stirling’s approximation for n!w h e n
nis large, namely
n!∼√
2πn
/n
e
/n
,
to give
q≈1−e−k
/365
365−k
/365−k+0.5
.
It is interesting to note that if k= 23 the probability is a little greater than a half that
at least two people have the same birthday, and if k= 50 the probability rises to 0.970.
This can prove a good bet at a party of non-mathematicians!
J
So far we have assumed that all nobjects are different (or distinguishable ). Let
us now consider nobjects of which n1are identical and of type 1, n2are identical
and of type 2, ...,n mare identical and of type m(clearly n=n1+n2+···+nm).
From (26.30) the number of permutations of these nobjects is again n!. However,
976
26.3 PERMUTATIONS AND COMBINATIONS
the number of distinguishable permutations is only
n!
n1!n2!···nm!, (26.32)
since the ith group of identical objects can be rearranged in ni! ways without
changing the distinguishable permutation.IA set of snooker balls consists of a white, a yellow, a green, a brown, a blue, a pink, a
black and 15reds. How many distinguishable permutations of the balls are there?
In total there are 22 balls, the 15 reds being indistinguishable. Thus from (26.32) the
number of distinguishable permutations is
22!
(1!)(1!)(1!)(1!)(1!)(1!)(15!)=22!
15!= 859541760 .
J
26.3.2 Combinations
We now consider the number of combinations of various objects when their order
is immaterial. Assuming all the objects to be distinguishable, from (26.31) we seethat the number of permutations of kobjects chosen from nis
nPk=n!/(n−k)!.
Now, since we are no longer concerned with the order of the chosen objects, which
can be internally arranged in k! different ways, the number of combinations of k
objects from nis
n!
(n−k)!k!≡nCk≡parenleftbiggn
kparenrightbigg
for 0≤k≤n, (26.33)
where, as noted in chapter 1,nCkis called the binomial coefficient since it also
appears in the binomial expansion for positive integer n,n a m e l y
(a+b)n=nsummationdisplay
k=0nCkakbn−k. (26.34)IA hand of 13 playing cards is dealt from a well-shuffled deck of 52. What is the probability
that the hand contains two aces?
Since the order of the cards in the hand is immaterial, the total number of distinct handsis simply equal to the number of combinations of 13 objects drawn from 52, i.e.
52C13.
However, the number of hands containing two aces is equal to the number of ways,4C2,
in which the two aces can be drawn from the four available, multiplied by the number ofways,
48C11, in which the remaining 11 cards in the hand can be drawn from the 48 cards
that are not aces. Thus the required probability is given by
4C248C11
52C13=4!
2!2!48!
11!37!13!39!
52!
=(3)(4)
2(12)(13)(38)(39)
(49)(50)(51)(52)=0.213
J
977
PROBABILITY
Another useful result that may be derived using the binomial coefficients is the
number of ways in which ndistinguishable objects can be divided into mpiles,
with niobjects in the ith pile, i=1,2,...,m (the ordering of objects within each
pile being unimportant). This may be straightforwardly calculated as follows. We
may choose the n1objects in the first pile from the original nobjects innCn1ways.
Then2objects in the second pile can then be chosen from the n−n1remaining
objects inn−n1Cn2ways, etc. We may continue in this fashion until we reach the
(m−1)th pile, which may be formed inn−n1−···−nm−2Cnm−1ways. The remaining
objects then form the mth pile and so can only be ‘chosen’ in one way. Thus the
total number of ways of dividing the original nobjects into mpiles is given by
the product
N=nCn1n−n1Cn2···n−n1−···−nm−2Cnm−1
=n!
n1!(n−n1)!(n−n1)!
n2!(n−n1−n2)!···(n−n1−n2−···−nm−2)!
nm−1!(n−n1−n2−···−nm−2−nm−1)!
=n!
n1!(n−n1)!(n−n1)!
n2!(n−n1−n2)!···(n−n1−n2−···−nm−2)!
nm−1!nm!
=n!
n1!n2!···nm!. (26.35)
These numbers are called multinomial coefficients since (26.35) is the coefficient of
xn1
1xn2
2···xnmmin the multinomial expansion of ( x1+x2+···+xm)n, i.e. for positive
integer n
(x1+x2+···+xm)n=summationdisplay
n1,n2,... ,nm
n1+n2+···+nm=nn!
n1!n2!···nm!xn1
1xn2
2···xnmm.
For the case m=2 ,n1=k,n2=n−k, (26.35) reduces to the binomial coefficient
nCk. Furthermore, we note that the multinomial coefficient (26.35) is identical to
the expression (26.32) for the number of distinguishable permutations of nobjects,
niof which are identical and of type i(fori=1,2,...,m andn1+n2+···+nm=n).
A few moments’ thought should convince the reader that the two expressions(26.35) and (26.32) must be identical.IIn the card game of bridge, each of four players is dealt 13 cards from a full pack of 52.
What is the probability that each player is dealt an ace?
From (26.35), the total number of distinct bridge dealings is 52! /(13!13!13!13!). However,
the number of ways in which the four aces can be distributed with one in each hand is
4!/(1!1!1!1!) = 4!; the remaining 48 cards can then be dealt out in 48! /(12!12!12!12!)
ways. Thus the probability that each player receives an ace is
4!48!
(12!)4(13!)4
52!=24(13)4
(49)(50)(51)(52)=0.105.
J
As in the case of permutations we might ask how many combinations of k
objects can be chosen from nwith replacement (repetition). To calculate this, we
978
26.3 PERMUTATIONS AND COMBINATIONS
may imagine the n(distinguishable) objects set out on a table. Each combination
ofkobjects can then be made by pointing to kof the no b j e c t si nt u r n( w i t h
repetitions allowed). These kequivalent selections distributed amongst ndifferent
but re-choosable objects are strictly analogous to the placing of kindistinguishable
‘balls’ in ndifferent boxes with no restriction on the number of balls in each box.
A particular selection in the case k=7 , n= 5 may be symbolised as
xxx||x|xx|x.
This denotes three balls in the first box, none in the second, one in the third,
two in the fourth and one in the fifth. We therefore need only to consider the
number of (distinguishable) ways in which kcrosses and n−1 vertical lines can
be arranged, i.e. the number of permutations of k+n−1 objects of which kare
identical crosses and n−1 are identical lines. This is given by (26.33) as
(k+n−1)!
k!(n−1)!=n+k−1Ck. (26.36)
We note that this expression also occurs in the binomial expansion for negative
integer powers. If nis a positive integer, it is straightforward to show that (see
chapter 1)
(a+b)−n=∞summationdisplay
k=0(−1)kn+k−1Cka−n−kbk,
where ai st a k e nt ob el a r g e rt h a n b.IA system contains a number Nof (non-interacting) particles, each of which can be in
any of the quantum states of the system. The structure of the set of quantum states is suchthat there exist Renergy levels with corresponding energies E
iand degeneracies gi(i.e. the
ith energy level contains giquantum states). Find the numbers of distinct ways in which
the particles can be distributed among the quantum states of the system such that the ith
energy level contains niparticles, for i=1,2,...,R , in the cases where the particles are
(i) distinguishable with no restriction on the number in each state;
(ii) indistinguishable with no restriction on the number in each state;
(iii) indistinguishable with a maxim um of one particle in each state;
(iv) distinguishable with a maximum of one particle in each state.
It is easiest to solve this problem in two stages. Let us first consider distributing the N
particles among the Renergy levels, without regard for the individual degenerate quantum
states that comprise each level. If the particles are distinguishable then the number of
distinct arrangements with niparticles in the ith level, i=1,2,...,R , is given by (26.35) as
N!
n1!n2!···nR!.
If, however, the particles are indistinguishable then clearly there exists only one distinct
arrangement having niparticles in the ith level, i=1,2,...,R . Now let us suppose there
exist wiways in which the niparticles in the ith energy level can be distributed among
thegidegenerate states. Thus it follows that the number of distinct ways in which the N
979
PROBABILITY
particles can be distributed among all Rquantum states of the system, with niparticles in
theith level, is given by
W{ni}=
/8/>/>
/>
/>/</>/>/>
/>/:
N!
n1!n2!···nR!RY
i=1wifor distinguishable particles ,
RY
i=1wi for indistinguishable particles .(26.37)
It therefore remains only for us to find the appropriate expression for wiin each of the
cases (i)–(iv) above.
Case (i). If there is no restriction on the number of particles in each quantum state,
then in the ith energy level each particle can reside in any of the gidegenerate quantum
states. Thus, if the particles are distinguishable then the number of distinct arrangementsis simply w
i=gni
i. Thus, from (26.37),
W{ni}=N!
n1!n2!···nR!RY
i=1gni
i=N!RY
i=1gni
i
ni!.
Such a system of particles (for example atoms or molecules in a classical gas) is said to
obey Maxwell–Boltzmann statistics.
Case (ii). If the particles are indistinguishable and there is no restriction on the number
in each state then, from (26.36), the number of distinct arrangements of the niparticles
among the gistates in the ith energy level is
wi=(ni+gi−1)!
ni!(gi−1)!.
Substituting this expression in (26.37), we obtain
W{ni}=RY
i=1(ni+gi−1)!
ni!(gi−1)!.
Such a system of particles (for example a gas of photons) is said to obey Bose–Einstein
statistics.
Case (iii). If a maximum of one particle can reside in each of the gidegenerate quantum
states in the ith energy level then the number of particles in each state is either 0 or 1.
Since the particles are indistinguishable, wiis equal to the number of distinct arrangements
in which nistates are occupied and gi−nistates are unoccupied; this is given by
wi=giCni=gi!
ni!(gi−ni)!.
Thus, from (26.37), we have
W{ni}=RY
i=1gi!
ni!(gi−ni)!.
Such a system is said to obey Fermi–Dirac statistics, and an example is provided by an
electron gas.
Case (iv). Again, the number of particles in each state is either 0 or 1. If the particles
are distinguishable, however, each arrangement identified in case (iii) can be reordered inn
i! different ways, so that
wi=giPni=gi!
(gi−ni)!.
980
26.4 RANDOM VARIABLES AND DISTRIBUTIONS
Substituting this expression into (26.37) gives
W{ni}=N!RY
i=1gi!
ni!(gi−ni)!.
Such a system of particles has the names of no famous scientists attached to it, since it
appears that it never occurs in nature.
J
26.4 Random variables and distributions
Suppose an experiment has an outcome sample space S. A real variable Xthat
is defined for all possible outcomes in S(so that a real number – not necessarily
unique – is assigned to each possible outcome) is called a random variable (RV).
The outcome of the experiment may already be a real number and hence a randomvariable, e.g. the number of heads obtained in 10 throws of a coin, or the sum of
the values if two dice are thrown. However, more arbitrary assignments are possi-
ble, e.g. the assignment of a ‘quality’ rating to each successive item produced by amanufacturing process. Furthermore, assuming that a probability can be assignedto all possible outcomes in a sample space S, it is possible to assign a probability
distribution to any random variable. Random variables may be divided into two
classes, discrete and continuous, and we now examine each of these in turn.
26.4.1 Discrete random variables
A random variable Xthat takes only discrete values x
1,x2,...,x n, with proba-
bilities p1,p2,...,p n, is called a discrete random variable. The number of values
nfor which Xhas a non-zero probability is finite or at most countably infinite.
As mentioned above, an example of a discrete random variable is the number ofheads obtained in 10 throws of a coin. If Xis a discrete random variable, we can
define a probability function (PF) f(x) that assigns probabilities to all the distinct
values that Xcan take, such that
f(x)=P r ( X=x)=braceleftBigg
p
iifx=xi,
0o t h e r w i s e .(26.38)
A typical PF (see figure 26.6) thus consists of spikes, at valid values ofX, whose
height at xcorresponds to the probability that X=x. Since the probabilities
must sum to unity, we require
nsummationdisplay
i=1f(xi)=1 . (26.39)
We may also define the cumulative probability function (CPF) of X,F(x), whose
value gives the probability that X≤x,s ot h a t
F(x)=P r ( X≤x)=summationdisplay
xi≤xf(xi). (26.40)
981
PROBABILITY
xf(x) F(x)
2p
p
1
2p
11
12 2 3 3 4 4 5 5 6 6
(a) (b)
Figure 26.6 ( a) A typical probability function for a discrete distribution, that
for the biased die discussed earlier. Since the probabilities must sum to unitywe require p=2/13. (b) The cumulative probability function for the same
discrete distribution. (Note that a different scale has been used for ( b).)
Hence F(x) is a step function that has upward jumps of piatx=xi,i=
1,2,...,n, and is constant between possible values of X. We may also calculate
the probability that Xlies between two limits, l1andl2(l1<l2); this is given by
Pr(l1<X≤l2)=summationdisplay
l1<xi≤l2f(xi)=F(l2)−F(l1), (26.41)
i.e. it is the sum of all the probabilities for which xilies within the relevant interval.IA bag contains seven red balls and three white balls. Three balls are drawn at random
and not replaced. Find the probability function for the number of red balls drawn.
LetXbe the number of red balls drawn. Then
Pr(X=0 )= f(0) =3
10×2
9×1
8=1
120,
Pr(X=1 )= f(1) =3
10×2
9×7
8×3=7
40,
Pr(X=2 )= f(2) =3
10×7
9×6
8×3=21
40,
Pr(X=3 )= f(3) =7
10×6
9×5
8=7
24.
It should be noted that
P3
i=0f(i) = 1, as expected.
J
26.4.2 Continuous random variables
A random variable Xis said to have a continuous distribution if Xis defined for a
continuous range of values between given limits (often −∞to∞). An example of
a continuous random variable is the height of a person drawn from a population,which can take anyvalue (within limits!). We can define the probability density
function (PDF) f(x) of a continuous random variable Xsuch that
Pr(x<X≤x+dx)=f(x)dx,
982
26.4 RANDOM VARIABLES AND DISTRIBUTIONS
l1 l2 abxf(x)
Figure 26.7 The probability density function for a continuous random vari-
ableXthat can take values only between the limits l1andl2. The shaded area
under the curve gives Pr( a<X≤b), whereas the total area under the curve,
between the limits l1andl2, is equal to unity.
i.e.f(x)dxis the probability that Xlies in the interval x<X≤x+dx. Clearly
f(x) must be a real function that is everywhere ≥0. If Xcan only take values
between the limits l1andl2then in order for the sum of the probabilities of all
possible outcomes to be equal to unity, we require
integraldisplayl2
l1f(x)dx=1.
Often Xcan take any value between −∞and∞and so
integraldisplay∞
−∞f(x)dx=1.
The probability that Xlies in the interval a<X≤bis then given by
Pr(a<X≤b)=integraldisplayb
af(x)dx, (26.42)
i.e. Pr( a<X≤b) is equal to the area under the curve of f(x) between these
limits (see figure 26.7).
We may also define the cumulative probability function F(x) for a continuous
random variable by
F(x)=P r ( X≤x)=integraldisplayx
l1f(u)du, (26.43)
where uis a (dummy) integration variable. We can then write
Pr(a<X≤b)=F(b)−F(a).
From (26.43) it is clear that f(x)=dF(x)/dx.
983
PROBABILITYIA random variable Xhas a PDF f(x)given by Ae−xin the interval 0<x<∞and zero
elsewhere. Find the value of the constant Aand hence calculate the probability that Xlies
in the interval 1<X≤2.
We require the integral of f(x) between 0 and ∞to equal unity. Evaluating this integral,
we findZ∞
0Ae−xdx=
/
−Ae−x
/∞
0=A,
and hence A= 1. From (26.42), we then obtain
Pr(1<X≤2) =
Z2
1f(x)dx=
Z2
1e−xdx=−e−2−(−e−1)=0 .23.
J
It is worth mentioning here that a discrete RV can in fact be treated as
continuous and assigned a corresponding probability density function. If Xis a
discrete RV that takes only the values x1,x2,...,x nwith probabilities p1,p2,...,p n
then we may describe Xas a continuous RV with PDF
f(x)=nsummationdisplay
i=1piδ(x−xi), (26.44)
where δ(x) is the Dirac delta function discussed in subsection 13.1.3. From (26.42)
and the fundamental property of the delta function (13.12), we see that
Pr(a<X≤b)=integraldisplayb
af(x)dx,
=nsummationdisplay
i=1piintegraldisplayb
aδ(x−xi)dx=summationdisplay
ipi,
where the final sum extends over those values of ifor which a<x i≤b.
26.4.3 Sets of random variables
It is common in practice to consider two or more random variables simultane-
ously. For example, one might be interested in both the height and weight ofa person drawn at random from a population. In the general case, these vari-ables may depend on one another and are described by joint probability density
functions . These are discussed fully in section 26.11, and we simply note here
that, if we have (say) two random variables XandY, then by analogy with the
single-variable case we define their joint probability density function f(x, y)i n
such a way that, if XandYare discrete RVs,
Pr(X=x
i,Y=yj)=f(xi,yj),
or, if XandYare continuous RVs,
Pr(x<X≤x+dx, y < Y≤y+dy)=f(x, y)dx dy.
984
26.5 PROPERTIES OF DISTRIBUTIONS
In many circumstances, however, random variables do not depend on one
another, i.e. they are independent . As an example, for a person drawn at random
from a population, we might expect height and IQ to be independent randomvariables. Let us suppose that XandYare two random variables with probability
density functions g(x)a n d h(y) respectively. In mathematical terms, XandYare
independent RVs if their joint probability density function is given by f(x, y)=
g(x)h(y). Thus, for independent RVs, if XandYare both discrete then
Pr(X=x
i,Y=yj)=g(xi)h(yj)
or, if XandYare both continuous, then
Pr(x<X≤x+dx, y < Y≤y+dy)=g(x)h(y)dx dy.
The important point in each case is that the RHS is simply the product of the
individual probability density functions (compare with the expression for Pr( A∪B)
in (26.22) for statistically independent events AandB). By a simple extension,
one may also consider the case where one of the random variables is discrete and
the other continuous. The above discussion may also be trivially extended to any
number of independent RVs Xi,i=1,2,...,N .IThe independent random variables XandYhave the PDFs g(x)=e−xandh(y)=2 e−2y
respectively. Calculate the probability that Xlies in the interval 1<X≤2andYlies in
the interval 0<Y≤1.
Since XandYare independent RVs, the required probability is given by
Pr(1<X≤2,0<Y≤1) =
Z2
1g(x)dx
Z1
0h(y)dy
=
Z2
1e−xdx
Z1
02e−2ydy
=
/
−e−x
/2
1×
/
−e−2y
/1
0=0.23×0.86 = 0 .20.
J
26.5 Properties of distributions
For a single random variable X, the probability density function f(x) contains
all possible information about how the variable is distributed. However, for the
purposes of comparison, it is conventional and useful to characterise f(x)b y
certain of its properties. Most of these standard properties are defined in termsofaverages orexpectation values . In the most general case, the expectation value
E[g(X)] of any function g(X) of the random variable Xis defined as
E[g(X)] =braceleftBiggsummationtext
ig(xi)f(xi) for a discrete distribution,integraltext
g(x)f(x)dxfor a continuous distribution,(26.45)
985
PROBABILITY
where the sum or integral is over all allowed values of X. It is assumed that
the series is absolutely convergent or that the integral exists, as the case may be.From its definition it is straightforward to show that the expectation value hasthe following properties:
(i) if ais a constant then E[a]=a;
(ii) if ais a constant then E[ag(X)] =aE[g(X)];
(iii) if g(X)=s(X)+t(X)t h e n E[g(X)] =E[s(X)] +E[t(X)].
It should be noted that the expectation value is not a function of Xbut is
instead a number that depends on the form of the probability density functionf(x) and the function g(x). Most of the standard quantities used to characterise
f(x) are simply the expectation values of various functions of the random variable
X. We now consider these standard quantities.
26.5.1 Mean
The property most commonly used to characterise a probability distribution is
itsmean, which is defined simply as the expectation value E[X] of the variable X
itself. Thus, the mean is given by
E[X]=braceleftBiggsummationtext
ixif(xi) for a discrete distribution,integraltext
xf(x)dxfor a continuous distribution.(26.46)
The alternative notations µand/angbracketleftx/angbracketrightare also commonly used to denote the mean.
If in (26.46) the series is not absolutely convergent, or the integral does not exist,
we say that the distribution does not have a mean, but this is very rare in physicalapplications.IThe probability of finding a 1selectron in a hydrogen atom in a given infinitesimal volume
dVisψ∗ψd V, where the quantum mechanical wavefunction ψis given by
ψ=Ae−r/a0.
Find the value of the real constant Aand thereby deduce the mean distance of the electron
from the origin.
Let us consider the random variable R= ‘distance of the electron from the origin’. Since
the 1s orbital has no θ-o rφ-dependence (it is spherically symmetric), we may consider
the infinitesimal volume element dVas the spherical shell with inner radius rand outer
radius r+dr. Thus, dV=4πr2drand the PDF of Ris simply
Pr(r<R≤r+dr)≡f(r)dr=4πr2A2e−2r/a0dr.
The value of Ais found by requiring the total probability (i.e. the probability that the
electron is somewhere ) to be unity. Since Rmust lie between zero and infinity, we require
that
A2
Z∞
0e−2r/a04πr2dr=1.
986
26.5 PROPERTIES OF DISTRIBUTIONS
Integrating by parts we find A=1/(πa3
0)1/2. Now, using the definition of the mean (26.46),
we find
E[R]=
Z∞
0rf(r)dr=4
a3
0
Z∞
0r3e−2r/a0dr.
The expression on the RHS may be integrated by parts and takes the value 3 a4
0/8;
consequently we find that E[R]=3 a0/2.
J
26.5.2 Mode and median
Although the mean discussed in the last section is the most common measure
of the ‘average’ of a distribution, two other measures, which do not rely on theconcept of expectation values, are frequently encountered.
Themodeof a distribution is the value of the random variable Xat which the
probability (density) function f(x) has its greatest value. If there is more than one
value of Xfor which this is true then each value may equally be called the mode
of the distribution.
Themedian Mof a distribution is the value of the random variable Xat which
the cumulative probability function F(x) takes the value
1
2,i . e .F(M)=1
2. Related
to the median are the lower and upper quartiles QlandQuof the PDF, which
are defined such that
F(Ql)=1
4,F (Qu)=3
4.
Thus the median and lower and upper quartiles divide the PDF into four regions
each containing one quarter of the probability. Smaller subdivisions are alsopossible, e.g. the nth percentile, P
n, of a PDF is defined by F(Pn)=n/100.IFind the mode of the PDF for the distance from the origin of the electron whose wave-
function was given in the previous example.
We found in the previous example that the PDF for the electron’s distance from the originwas given by
f(r)=4r
2
a3
0e−2r/a0. (26.47)
Differentiating f(r) with respect to r,w eo b t a i n
df
dr=8r
a3
0
/
1−r
a0
/
e−2r/a0.
Thus f(r) has turning points at r=0a n d r=a0,w h e r e df/dr = 0. It is straightforward
to show that r= 0 is a minimum and r=a0is a maximum. Moreover, it is also clear that
r=a0is a global maximum (as opposed to just a local one). Thus the mode of f(r) occurs
atr=a0.
J
987
PROBABILITY
26.5.3 Variance and standard deviation
Thevariance of a distribution, V[X], also written σ2, is defined by
V[X]=Ebracketleftbig
(X−µ)2bracketrightbig
=braceleftBiggsummationtext
j(xj−µ)2f(xj) for a discrete distribution,integraltext
(x−µ)2f(x)dxfor a continuous distribution.
(26.48)
Here µhas been written for the expectation value E[X]o fX. As in the case of
the mean, unless the series and the integral in (26.48) converge the distributiondoes not have a variance. From the definition (26.48) we may easily derive thefollowing useful properties of V[X]. Ifaandbare constants then
(i)V[a]=0 ,
(ii)V[aX+b]=a
2V[X].
The variance of a distribution is always positive; its positive square root is
known as the standard deviation of the distribution and is often denoted by σ.
Roughly speaking, σmeasures the spread (about x=µ) of the values that Xcan
assume.IFind the standard deviation of the PDF for the distance from the origin of the electron
whose wavefunction was discussed in the previous two examples.
Inserting the expression (26.47) for the PDF f(r) into (26.48), the variance of the random
variable Ris given by
V[R]=
Z∞
0(r−µ)24r2
a3
0e−2r/a0dr=4
a3
0
Z∞
0(r4−2r3µ+r2µ2)e−2r/a0dr,
where the mean µ=E[R]=3 a0/2. Integrating each term in the integrand by parts we
obtain
V[R]=3 a2
0−3µa0+µ2=3a2
0
4.
Thus the standard deviation of the distribution is σ=√
3a0/2.
J
We may also use the definition (26.48) to derive the Bienaym ´e–Chebyshev
inequality , which provides a useful upper limit on the probability that random
variable Xtakes values outside a given range centred on the mean. Let us consider
the case of a continuous random variable, for which
Pr(|X−µ|≥c)=integraldisplay
|x−µ|≥cf(x)dx,
where the integral on the RHS extends over all values of xsatisfying the inequality
988
26.5 PROPERTIES OF DISTRIBUTIONS
|x−µ|≥c. From (26.48), we find that
σ2≥integraldisplay
|x−µ|≥c(x−µ)2f(x)dx≥c2integraldisplay
|x−µ|≥cf(x)dx. (26.49)
The first inequality holds because both ( x−µ)2andf(x) are non-negative for
allx, and the second inequality holds because ( x−µ)2≥c2over the range of
integration. However, the RHS of (26.49) is simply equal to c2Pr(|X−µ|≥c),
and thus we obtain the required inequality
Pr(|X−µ|≥c)≤σ2
c2.
A similar derivation may be carried through for the case of a discrete random
variable. Thus, for anydistribution f(x) that possesses a variance we have, for
example,
Pr(|X−µ|≥2σ)≤1
4and Pr(|X−µ|≥3σ)≤1
9.
26.5.4 Moments
The mean (or expectation) of Xis sometimes called the first moment ofX,s i n c e
it is defined as the sum or integral of the probability density function multiplied
by the first power of x. By a simple extension the kth moment of a distribution
is defined by
µk≡E[Xk]=braceleftBiggsummationtext
jxk
jf(xj) for a discrete distribution,integraltext
xkf(x)dxfor a continuous distribution.(26.50)
For notational convenience, we have introduced the symbol µkto denote E[Xk],
thekth moment of the distribution. Clearly, the mean of the distribution is then
denoted by µ1, often abbreviated simply to µ, as in the previous subsection, as
this rarely causes confusion.
A useful result that relates the second moment, the mean and the variance of
a distribution is proved using the properties of the expectation operator:
V[X]=Ebracketleftbig
(X−µ)2bracketrightbig
=Ebracketleftbig
X2−2µX+µ2bracketrightbig
=Ebracketleftbig
X2bracketrightbig
−2µE[X]+µ2
=Ebracketleftbig
X2bracketrightbig
−2µ2+µ2
=Ebracketleftbig
X2bracketrightbig
−µ2. (26.51)
In alternative notations, this result can be written
/angbracketleft(x−µ)2/angbracketright=/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2or σ2=µ2−µ2
1.
989
PROBABILITYIA biased die has probabilities p/2,p,p,p,p,2pof showing 1, 2, 3, 4, 5, 6 respectively. Find
(i) the mean, (ii) the second moment and (iii) th e variance of this probability distribution.
By demanding that the sum of the probabilities equals unity we require p=2/13. Now,
using the definition of the mean (26.46) for a discrete distribution,
E[X]=
X
jxjf(xj)=1×1
2p+2×p+3×p+4×p+5×p+6×2p
=53
2p=53
2×2
13=53
13.
Similarly, using the definition of the second moment (26.50),
E[X2]=
X
jx2
jf(xj)=12×1
2p+22p+32p+42p+52p+62×2p
=253
2p=253
13.
Finally, using the definition of the variance (26.48), with µ=5 3/13, we obtain
V[X]=
X
j(xj−µ)2f(xj)
=( 1−µ)21
2p+( 2−µ)2p+( 3−µ)2p+( 4−µ)2p+( 5−µ)2p+( 6−µ)22p
=
/3120
169
/
p=480
169.
It is easy to verify that V[X]=E
/
X2
/
−(E[X])2.
J
In practice, to calculate the moments of a distribution it is often simpler to use
the moment generating function discussed in subsection 26.7.2. This is particularlytrue for higher-order moments, where direct evaluation of the sum or integral in
(26.50) can be somewhat laborious.
26.5.5 Central moments
The variance V[X] is sometimes called the second central moment of the distribu-
tion, since it is defined as the sum or integral of the probability density functionmultiplied by the second power of x−µ. The origin of the term ‘central’ is that by
subtracting µfrom xbefore squaring we are considering the moment about the
mean of the distribution, rather than about x= 0. Thus the kthcentral moment
of a distribution is defined as
ν
k≡Ebracketleftbig
(X−µ)kbracketrightbig
=braceleftBiggsummationtext
j(xj−µ)kf(xj) for a discrete distribution,integraltext
(x−µ)kf(x)dxfor a continuous distribution.(26.52)
It is convenient to introduce the notation νkfor the kth central moment. Thus
V[X]≡ν2and we may write (26.51) as ν2=µ2−µ2
1. Clearly, the first central
990
26.5 PROPERTIES OF DISTRIBUTIONS
moment of a distribution is always zero since, for example in the continuous case,
ν1=integraldisplay
(x−µ)f(x)dx=integraldisplay
xf(x)dx−µintegraldisplay
f(x)dx=µ−(µ×1) = 0 .
We note that the notation µkand νkfor the moments and central moments
respectively is not universal. Indeed, in some books their meanings are reversed.
We can write the kth central moment of a distribution in terms of its kth and
lower-order moments by expanding ( X−µ)kin powers of X. We have already
noted that ν2=µ2−µ2
1, and similar expressions may be obtained for higher-order
central moments. For example,
ν3=Ebracketleftbig
(X−µ1)3bracketrightbig
=Ebracketleftbig
X3−3µ1X2+3µ2
1X−µ3
1bracketrightbig
=µ3−3µ1µ2+3µ2
1µ1−µ3
1
=µ3−3µ1µ2+2µ3
1. (26.53)
In general, it is straightforward to show that
νk=µk−kC1µk−1µ1+···+(−1)rkCrµk−rµr
1+···+(−1)k−1(kCk−1−1)µk
1.
(26.54)
Once again, direct evaluation of the sum or integral in (26.52) can be rather
tedious for higher moments, and it is usually quicker to use the moment generatingfunction (see subsection 26.7.2), from which the central moments can be easilyevaluated as well.IThe PDF for a Gaussian distribution (see subsection 26.9.1) with mean µand variance
σ2is given by
f(x)=1
σ√
2πexp
/
−(x−µ)2
2σ2
/
.
Obtain an expression for the kth central moment of this distribution.
As an illustration, we will perform this calculation by evaluating the integral in (26.52)
directly. Thus, the kth central moment of f(x)i sg i v e nb y
νk=
Z∞
−∞(x−µ)kf(x)dx
=1
σ√
2π
Z∞
−∞(x−µ)kexp
/
−(x−µ)2
2σ2
/
dx
=1
σ√
2π
Z∞
−∞ykexp
/
−y2
2σ2
/
dy, (26.55)
where in the last line we have made the substitution y=x−µ. It is clear that if kis
odd then the integrand is an odd function of yand hence the integral equals zero. Thus,
νk=0i f kis odd. When kis even, we could calculate νkby integrating by parts to obtain
a reduction formula, but it is more elegant to consider instead the standard integral (seesubsection 6.4.2)
I=
Z∞
−∞exp(−αy2)dy=π1/2α−1/2,
991
PROBABILITY
and differentiate it repeatedly with respect to α(see section 5.12). Thus, we obtain
dI
dα=−
Z∞
−∞y2exp(−αy2)dy=−1
2π1/2α−3/2
d2I
dα2=
Z∞
−∞y4exp(−αy2)dy=(1
2)(3
2)π1/2α−5/2
...
dnI
dαn=(−1)n
Z∞
−∞y2nexp(−αy2)dy=(−1)n(1
2)(3
2)···(1
2(2n−1))π1/2α−(2n+1)/2.
Setting α=1/(2σ2) and substituting the above result into (26.55), we find (for keven)
νk=(1
2)(3
2)···(1
2(k−1))(2σ2)k/2= (1)(3) ···(k−1)σk.
J
One may also characterise a probability distribution f(x) using the closely
related normalised and dimensionless central moments
γk≡νk
νk/2
2=νk
σk.
From this set, γ3andγ4are more commonly called, respectively, the skewness
andkurtosis of the distribution. The skewness γ3of a distribution is zero if it is
symmetrical about its mean. If the distribution is skewed to values of xsmaller
than the mean then γ3<0. Similarly γ3>0 if the distribution is skewed to higher
values of x.
From the above example, we see that the kurtosis of the Gaussian distribution
(subsection 26.9.1) is given by
γ4=ν4
ν2
2=3σ4
σ4=3.
It is therefore common practice to define the excess kurtosis of a distribution
asγ4−3. A positive value of the excess kurtosis implies a relatively narrower
peak and wider wings than the Gaussian distribution with the same mean andvariance. A negative excess kurtosis implies a wider peak and shorter wings.
Finally, we note here that one can also describe a probability density function
f(x) in terms of its cumulants , which are again related to the central moments.
However, we defer the discussion of cumulants until subsection 26.7.4, since their
definition is most easily understood in terms of generating functions.
26.6 Functions of random variables
Suppose Xis some random variable for which the probability density function
f(x) is known. In many cases, we may be more interested in the related random
variable Y=Y(X), where Y(X) is some function of X. What is the probability
992
26.6 FUNCTIONS OF RANDOM VARIABLES
density function g(y) for the new random variable Y? We now discuss how to
obtain this function.
26.6.1 Discrete random variables
IfXis a discrete RV that takes only the values xi,i=1,2,...,n,t h e n Ymust
also be discrete and takes the values yi=Y(xi), although some of these values
may be identical. The probability function for Yis given by
g(y)=braceleftBiggsummationtext
jf(xj)i f y=yi,
0o t h e r w i s e ,(26.56)
where the sum extends over those values of jfor which yi=Y(xj). The simplest
case arises when the function Y(X) possesses a single-valued inverse X(Y). In this
case, only one x-value corresponds to each y-value, and we obtain a closed-form
expression for g(y)g i v e nb y
g(y)=braceleftBigg
f(x(y)) if y=yi,
0o t h e r w i s e .
IfY(X) does not possess a single-valued inverse then the situation is more
complicated and it may not be possible to obtain a closed-form expression forg(y). Nevertheless, whatever the form of Y(X), one can always use (26.56) to
obtain the numerical values of the probability function g(y)a ty=y
i.
26.6.2 Continuous random variables
IfXis a continuous RV, then so too is the new random variable Y=Y(X). The
probability that Ylies in the range ytoy+dyis given by
g(y)dy=integraldisplay
dSf(x)dx, (26.57)
where dScorresponds to all values of xfor which Ylies in the range ytoy+dy.
Once again the simplest case occurs when Y(X) possesses a single-valued inverse
X(Y). In this case, we may write
g(y)dy=vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplayx(y+dy)
x(y)f(x/prime)dx/primevextendsinglevextendsinglevextendsinglevextendsingle=integraldisplayx(y)+|dx
dy|dy
x(y)f(x/prime)dx/prime,
from which we obtain
g(y)=f(x(y))vextendsinglevextendsinglevextendsinglevextendsingledx
dyvextendsinglevextendsinglevextendsinglevextendsingle. (26.58)
993
PROBABILITY
lighthouse
beam L
0θ
coastliney
Figure 26.8 The illumination of a coastline by the beam from a lighthouse.IA lighthouse is situated at a distance Lfrom a straight coastline, opposite a point O, and
sends out a narrow continuous beam of light simultaneously in opposite directions. The beamrotates with constant angular velocity. If the random variable Yis the distance along the
coastline, measured from O, of the spot that the light beam illuminates, find its probability
density function.
The situation is illustrated in figure 26.8. Since the light beam rotates at a constant angularvelocity, θis distributed uniformly between −π/2a n d π/2, and so f(θ)=1 /π.N o w
y=Ltanθ, which possesses the single-valued inverse θ=t a n
−1(y/L), provided that θlies
between−π/2a n d π/2. Since dy/dθ =Lsec2θ=L(1 + tan2θ)=L[1 + ( x/L)2], from
(26.58) we find
g(y)=1
π
////dθ
dy
////=1
πL[1 + ( y/L)2]for−∞<y<∞.
A distribution of this form is called a Cauchy distribution and is discussed in subsec-
tion 26.9.5.
J
IfY(X) does not possess a single-valued inverse then we encounter complica-
tions, since there exist several intervals in the X-domain for which Ylies between
yandy+dy. This is illustrated in figure 26.9, which shows a function Y(X)
such that X(Y) is a double-valued function of Y. Thus the range ytoy+dy
corresponds to X’s being either in the range x1tox1+dx1or in the range x2to
x2+dx2. In general, it may not be possible to obtain an expression for g(y)i n
closed form, although the distribution may always be obtained numerically using(26.57). However, a closed-form expression may be obtained in the case wherethere exist single-valued functions x
1(y)a n d x2(y) giving the two values of xthat
correspond to any given value of y.I nt h i sc a s e ,
g(y)dy=vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
x1(y+dy)
x1(y)f(x)dxvextendsinglevextendsinglevextendsinglevextendsingle+vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
x2(y+dy)
x2(y)f(x)dxvextendsinglevextendsinglevextendsinglevextendsingle,
from which we obtain
g(y)=f(x
1(y))vextendsinglevextendsinglevextendsinglevextendsingledx1
dyvextendsinglevextendsinglevextendsinglevextendsingle+f(x2(y))vextendsinglevextendsinglevextendsinglevextendsingledx2
dyvextendsinglevextendsinglevextendsinglevextendsingle. (26.59)
994
26.6 FUNCTIONS OF RANDOM VARIABLES
yy+dy
dx1 dx2XY
Figure 26.9 Illustration of a function Y(X) such that its inverse X(Y)i sa
double-valued function of Y. The range ytoy+dycorresponds to Xbeing
either in the range x1tox1+dx1or in the range x2tox2+dx2.
This result may be generalised straightforwardly to the case where the range yto
y+dycorresponds to more than two x-intervals.IThe random variable Xis Gaussian distributed (see subsection 26.9.1) with mean µand
variance σ2. Find the PDF of the new variable Y=(X−µ)2/σ2.
It is clear that X(Y) is a double-valued function of Y. However, in this case, it is
straightforward to obtain single-valued functions giving the two values of xthat correspond
to a given value of y;t h e s ea r e x1=µ−σ√yandx2=µ+σ√y,w h e r e√yis taken to
mean the positive square root.
The PDF of Xis given by
f(x)=1
σ√
2πexp
/
−(x−µ)2
2σ2
/
.
Since dx1/dy=−σ/(2√y)a n d dx2/dy=σ/(2√y), from (26.59) we obtain
g(y)=1
σ√
2πexp(−1
2y)
////−σ
2√y
////+1
σ√
2πexp(−1
2y)
////σ
2√y
////
=1
2√π(1
2y)−1/2exp(−1
2y).
As we shall see in subsection 26.9.3, this is the gamma distribution γ(1
2,1
2).
J
26.6.3 Functions of several random variables
We may extend our discussion further, to the case in which the new random
variable is a function of several other random variables. For definiteness, let us
consider the random variable Z=Z(X,Y), which is a function of two other
RVs XandY. Given that these variables are described by the joint probability
density function f(x, y), we wish to find the probability density function p(z)o f
the variable Z.
995
PROBABILITY
IfXandYare both discrete RVs then
p(z)=summationdisplay
i,jf(xi,yj), (26.60)
where the sum extends over all values of iandjfor which Z(xi,yj)=z. Similarly,
ifXandYare both continuous RVs then p(z) is found by requiring that
p(z)dz=integraldisplayintegraldisplay
dSf(x, y)dx dy, (26.61)
where dSis the infinitesimal area in the xy-plane lying between the curves
Z(x, y)=zandZ(x, y)=z+dz.ISuppose XandYare independent continuous random variables in the range −∞to∞,
with PDFs g(x)andh(y)respectively. Obtain expressions for the PDFs of Z=X+Yand
W=XY.
Since XandYare independent RVs, their joint PDF is simply f(x, y)=g(x)h(y). Thus,
from (26.61), the PDF of the sum Z=X+Yis given by
p(z)dz=
Z∞
−∞dx g(x)
Zz+dz−x
z−xdy h(y)
=
/Z∞
−∞g(x)h(z−x)dx
/
dz.
Thus p(z)i st h e convolution of the PDFs of gandh(i.e.p=g∗h, see subsection 13.1.7).
In a similar way, the PDF of the product W=XYis given by
q(w)dw=
Z∞
−∞dx g(x)
Z(w+dw)/|x|
w/|x|dy h(y)
=
/Z∞
−∞g(x)h(w/x)dx
|x|
/
dw
J
The prescription (26.61) is readily generalised to functions of nrandom variables
Z=Z(X1,X2,...,X n), in which case the infinitesimal ‘volume’ element dSis the
region in x1x2···xn-space between the (hyper)surfaces Z(x1,x2,...,x n)=zand
Z(x1,x2,...,x n)=z+dz. In practice, however, the integral is difficult to evaluate,
since one is faced with the complicated geometrical problem of determining thelimits of integration. Fortunately, an alternative (and powerful) technique existsfor evaluating integrals of this kind. One eliminates the geometrical problem byintegrating over allvalues of the variables x
iwithout restriction, while shifting
the constraint on the variables to the integrand. This is readily achieved bymultiplying the integrand by a function that equals unity in the infinitesimal
region dSand zero elsewhere. From the discussion of the Dirac delta function in
subsection 13.1.3, we see that δ(Z(x
1,x2,...,x n)−z)dzsatisfies these requirements,
and so in the most general case we have
p(z)=integraldisplayintegraldisplay
···integraldisplay
f(x1,x2,...,x n)δ(Z(x1,x2,...,x n)−z)dx1dx2...d x n,
(26.62)
996
26.6 FUNCTIONS OF RANDOM VARIABLES
where the range of integration is over all possible values of the variables xi.T h i s
integral is most readily evaluated by substituting in (26.62) the Fourier integralrepresentation of the Dirac delta function discussed in subsection 13.1.4, namely
δ(Z(x
1,x2,...,x n)−z)=1
2πintegraldisplay∞
−∞eik(Z(x1,x2,...,x n)−z)dk. (26.63)
This is best illustrated by considering a specific example.IA general one-dimensional random walk consists of nindependent steps, each of which
can be of a different length and in either direction along the x-axis. If g(x)is the PDF for
the (positive or negative) displacement Xalong the x-axis achieved in a single step, obtain
an expression for the PDF of the total displacement Safter nsteps.
The total displacement Sis simply the algebraic sum of the displacements Xiachieved in
each of the nsteps, so that
S=X1+X2+···+Xn.
Since the random variables Xiare independent and have the same PDF g(x), their joint
PDF is simply g(x1)g(x2)···g(xn). Substituting this into (26.62), together with (26.63), we
obtain
p(s)=
Z∞
−∞
Z∞
−∞···
Z∞
−∞g(x1)g(x2)···g(xn)1
2π
Z∞
−∞eik[(x1+x2+···+xn)−s]dk dx 1dx2···dxn
=1
2π
Z∞
−∞dk e−iks
/Z∞
−∞g(x)eikxdx
/n
. (26.64)
It is convenient to define the characteristic function C(k) of the variable Xas
C(k)=
Z∞
−∞g(x)eikxdx,
which is simply related to the Fourier transform of g(x). Then (26.64) may be written as
p(s)=1
2π
Z∞
−∞e−iks[C(k)]ndk.
Thus p(s) can be found by evaluating two Fourier in tegrals. Characteristic functions will
be discussed in more detail in subsection 26.7.3.
J
26.6.4 Expectation values and variances
In some cases, one is interested only in the expectation value or the variance
of the new variable Zrather than in its full probability density function. For
definiteness, let us consider the random variable Z=Z(X,Y), which is a function
of two RVs XandYwith a known joint distribution f(x, y); the results we will
obtain are readily generalised to more (or fewer) variables.
It is clear that E[Z]a n d V[Z] can be obtained, in principle, by first using the
methods discussed above to obtain p(z) and then evaluating the appropriate sums
or integrals. The intermediate step of calculating p(z) is not necessary, however,
since it is straightforward to obtain expressions for E[Z]a n d V[Z] in terms of
997
PROBABILITY
the variables XandY. For example, if XandYare continuous RVs then the
expectation value of Zis given by
E[Z]=integraldisplay
zp(z)dz=integraldisplayintegraldisplay
Z(x, y)f(x, y)dx dy. (26.65)
An analogous result exists for discrete random variables.
Integrals of the form (26.65) are often difficult to evaluate. Nevertheless, we
may use (26.65) to derive an important general result concerning expectationvalues. If XandYareanytwo random variables and aandbare arbitrary
constants then by letting Z=aX+bYwe find
E[aX+bY]=aE[X]+bE[Y].
Furthermore, we may use this result to obtain an approximate expression for the
expectation value E[Z(X,Y)] of any arbitrary function of XandY. Letting µ
X=
E[X]a n d µY=E[Y], then, provided Z(X,Y) can be reasonably approximated
by the linear terms of its Taylor expansion about the point ( µX,µY), we have
Z(X,Y)≈Z(µX,µY)+parenleftbigg∂Z
∂Xparenrightbigg
(X−µX)+parenleftbigg∂Z
∂Yparenrightbigg
(Y−µY),
(26.66)
where the partial derivatives are evaluated at X=µXandY=µY.T a k i n gt h e
expectation value of both sides, we find
E[Z(X,Y)]≈Z(µX,µY)+parenleftbigg∂Z
∂Xparenrightbigg
(E[X]−µX)+parenleftbigg∂Z
∂Yparenrightbigg
(E[Y]−µY)=Z(µX,µY),
which gives the approximate result E[Z(X,Y)]≈Z(µX,µY).
By analogy with (26.65), the variance of Z=Z(X,Y)i sg i v e nb y
V[Z]=integraldisplay
(z−µZ)2p(z)dz=integraldisplayintegraldisplay
[Z(x, y)−µZ]2f(x, y)dx dy,
(26.67)
where µZ=E[Z]. We may use this expression to derive a useful general result. If
XandYare two independent random variables, so that f(x, y)=g(x)h(y), and
a,bandcare constants then by setting Z=aX+bY+cin (26.67) we obtain
V[aX+bY+c]=a2V[X]+b2V[Y]. (26.68)
From (26.68) we also obtain the important special case
V[X+Y]=V[X−Y]=V[X]+V[Y].
Provided XandYare indeed independent random variables, we may obtain
an approximate expression for V[Z(X,Y)], for any arbitrary function Z(X,Y),
in a similar manner to that used in approximating E[Z(X,Y)] above. Taking the
998
26.7 GENERATING FUNCTIONS
variance of both sides of (26.66), and using (26.68), we find
V[Z(X,Y)]≈parenleftbigg∂Z
∂Xparenrightbigg2
V[X]+parenleftbigg∂Z
∂Yparenrightbigg2
V[Y], (26.69)
the partial derivatives being evaluated at X=µXandY=µY.
26.7 Generating functions
As we saw in chapter 16, when dealing with particular sets of functions fn,
each member of the set being characterised by a different non-negative integern, it is sometimes possible to summarise the whole set by a single function of a
dummy variable (say t), called a generating function. The relationship between
the generating function and the nth member f
nof the set is that if the generating
function is expanded as a power series in tthen fnis the coefficient of tn.F o r
example, in the expansion of the generating function G(z,t)=( 1−2zt+t2)−1/2,
the coefficient of tnis the nth Legendre polynomial Pn(z), i.e.
G(z,t)=( 1−2zt+t2)−1/2=∞summationdisplay
n=0Pn(z)tn.
We found that many useful properties of, and relationships between, the members
of a set of functions could be established using the generating function and otherfunctions obtained from it, e.g. its derivatives.
Similar ideas can be used in the area of probability theory, and two types of
generating function can be usefully defined, one more generally applicable than
the other. The more restricted of the two, applicable only to discrete integral
distributions, is called a probability generating function; this is discussed in thenext section. The second type, a moment generating function, can be used withboth discrete and continuous distributio ns and is considered in subsection 26.7.2.
From the moment generating function, we may also construct the closely re-lated characteristic and cumulant generating functions; these are discussed insubsections 26.7.3 and 26.7.4 respectively.
26.7.1 Probability generating functions
As already indicated, probability generating functions are restricted in applicabil-
ity to integer distributions, of which the most common (the binomial, the Poissonand the geometric) are considered in this and later subsections. In such distribu-tions a random variable may take only non-negative integer values. The actualpossible values may be finite or infinite in number, but, for formal purposes,
all integers, 0 ,1,2,...are considered possible. If only a finite number of integer
values can occur in any particular case then those that cannot occur are includedbut are assigned zero probability.
999
PROBABILITY
If, as previously, the probability that the random variable Xtakes the value xn
isf(xn), then
summationdisplay
nf(xn)=1 .
In the present case, however, only non-negative integer values of xnare possible,
and we can, without ambiguity, write the probability that Xtakes the value nas
fn, with
∞summationdisplay
n=0fn=1. (26.70)
We may now define the probability generating function ΦX(t)b y
ΦX(t)≡∞summationdisplay
n=0fntn. (26.71)
It is immediately apparent that Φ X(t)=E[tX] and that, by virtue of (26.70),
ΦX(1) = 1.
Probably the simplest example of a probability generating function (PGF) is
provided by the random variable Xdefined by
X=braceleftBigg
1 if the outcome of a single trial is a ‘success’,
0 if the trial ends in ‘failure’.
If the probability of success is pand that of failure q(= 1−p)t h e n
ΦX(t)=qt0+pt1+0+0+ ···=q+pt. (26.72)
This type of random variable is discussed much more fully in subsection 26.8.1.
In a similar but slightly more complicated way, a Poisson-distributed integervariable with mean λ(see subsection 26.8.4) has a PGF
Φ
X(t)=∞summationdisplay
n=0e−λλn
n!tn=e−λeλt. (26.73)
We note that, as required, Φ X(1) = 1 in both cases.
Useful results will be obtained from this kind of approach only if the summation
(26.71) can be carried out explicitly in particular cases and the functions derivedfrom Φ
X(t) can be shown to be related to meaningful parameters. Two such
relationships can be obtained by differentiating (26.71) with respect to t. Taking
the first derivative we find
dΦX(t)
dt=∞summationdisplay
n=0nfntn−1⇒Φ/prime
X(1) =∞summationdisplay
n=0nfn=E[X], (26.74)
1000
26.7 GENERATING FUNCTIONS
and differentiating once more we obtain
d2ΦX(t)
dt2=∞summationdisplay
n=0n(n−1)fntn−2⇒Φ/prime/prime
X(1) =∞summationdisplay
n=0n(n−1)fn=E[X(X−1)].
(26.75)
Equation (26.74) shows that Φ/prime
X(1) gives the mean of X. Using both (26.75) and
(26.51) allows us to write
Φ/prime/prime
X(1) + Φ/primeX(1)−bracketleftbig
Φ/prime
X(1)bracketrightbig2=E[X(X−1)] + E[X]−(E[X])2
=Ebracketleftbig
X2bracketrightbig
−E[X]+E[X]−(E[X])2
=Ebracketleftbig
X2bracketrightbig
−(E[X])2
=V[X], (26.76)
and so express the variance of Xin terms of the derivatives of its probability
generating function.IA random variable Xis given by the number of trials needed to obtain a first success
when the chance of success at each trial is constant and equal to p. Find the probability
generating function for Xand use it to determine the mean and variance of X.
Clearly, at least one trial is needed, and so f0=0 .I f n(≥1) trials are needed for the first
success, the first n−1 trials must have resulted in failure. Thus
Pr(X=n)=qn−1p, n≥1, (26.77)
where q=1−pis the probability of failure in each individual trial.
The corresponding probability generating function is thus
ΦX(t)=∞X
n=0fntn=∞X
n=1(qn−1p)tn
=p
q∞X
n=1(qt)n=p
q×qt
1−qt=pt
1−qt, (26.78)
where we have used the result for the sum of a geometric series, given in chapter 4, to
obtain a closed-form expression for Φ X(t). Again, as must be the case, Φ X(1) = 1.
To find the mean and variance of Xwe need to evaluate Φ/prime
X(1) and Φ/prime/primeX(1). Differentiating
(26.78) gives
Φ/prime
X(t)=p
(1−qt)2⇒Φ/prime
X(1) =p
p2=1
p,
Φ/prime/prime
X(t)=2pq
(1−qt)3⇒Φ/prime/prime
X(1) =2pq
p3=2q
p2.
Thus, using (26.74) and (26.76),
E[X]=Φ/prime
X(1) =1
p,
V[X]=Φ/prime/prime
X(1) + Φ/primeX(1)−[Φ/prime
X(1)]2
=2q
p2+1
p−1
p2=q
p2.
A distribution with probabilities of the general form (26.77) is known as a geometric
distribution and is discussed in subsection 26.8.2. This form of distribution is common in
‘waiting time’ problems (subsection 26.9.3).
J
1001
PROBABILITY
n
r=n
r
Figure 26.10 The pairs of values of nandrused in the evaluation of Φ X+Y(t).
Sums of random variables
We now turn to considering the sum of two or more independent random
variables, say XandY, and denote by S2the random variable
S2=X+Y.
If Φ S2(t)i st h eP G Ff o r S2, the coefficient of tnin its expansion is given by the
probability that X+Y=nand is thus equal to the sum of the probabilities that
X=randY=n−rfor all values of rin 0≤r≤n. Since such outcomes for
different values of rare mutually exclusive, we have
Pr(X+Y=n)=∞summationdisplay
r=0Pr(X=r)P r (Y=n−r). (26.79)
Multiplying both sides of (26.79) by tnand summing over all values of nenables
us to express this relationship in terms of probability generating functions asfollows:
Φ
X+Y(t)=∞summationdisplay
n=0Pr(X+Y=n)tn=∞summationdisplay
n=0nsummationdisplay
r=0Pr(X=r)trPr(Y=n−r)tn−r
=∞summationdisplay
r=0∞summationdisplay
n=rPr(X=r)trPr(Y=n−r)tn−r.
The change in summation order is justified by reference to figure 26.10, which
illustrates that the summations are over exactly the same pairs of values of nand
r, but with the first (inner) summation over the points in a column rather than
over the points in a row. Now, setting n=r+sgives the final result,
ΦX+Y(t)=∞summationdisplay
r=0Pr(X=r)tr∞summationdisplay
s=0Pr(Y=s)ts
=Φ X(t)ΦY(t), (26.80)
1002
26.7 GENERATING FUNCTIONS
i.e. the PGF of the sum of two independent random variables is equal to the
product of their individual PGFs. The same result can be deduced in a less formalway by noting that if XandYare independent then
Ebracketleftbig
t
X+Ybracketrightbig
=Ebracketleftbig
tXbracketrightbig
Ebracketleftbig
tYbracketrightbig
.
Clearly result (26.80) can be extended to more than two random variables by
writing S3=S2+Zetc., to give
Φ(
Pn
i=1Xi)(t)=nproductdisplay
i=1ΦXi(t), (26.81)
and, further, if all the Xihave the same probability distribution,
Φ(
Pn
i=1Xi)(t)=[ΦX(t)]n. (26.82)
This latter result has immediate application in the deduction of the PGF for the
binomial distribution from that for a single trial, equation (26.72).
Variable-length sums of random variables
As a final result in the theory of probability generating functions we show how to
calculate the PGF for a sum of Nrandom variables, all with the same probability
distribution, when the value of Nis itself a random variable but one with a
known probability distribution. In symbols, we wish to find the distribution
of
SN=X1+X2+···+XN, (26.83)
where Nis a random variable with Pr( N=n)= hnand PGF χN(t)=summationtexthntn.
The probability ξkthatSN=kis given by a sum of conditional probabilities,
namely†
ξk=∞summationdisplay
n=0Pr(N=n)P r(X0+X1+X2+···+Xn=k)
=∞summationdisplay
n=0hn×coefficient of tkin [Φ X(t)]n.
Multiplying both sides of this equation by tkand summing over all k,w eo b t a i n
†Formally X0= 0 has to be included, since Pr( N= 0) may be non-zero.
1003
PROBABILITY
an expression for the PGF Ξ S(t)o fSN:
ΞS(t)=∞summationdisplay
k=0ξktk=∞summationdisplay
k=0tk∞summationdisplay
n=0hn×coefficient of tkin [Φ X(t)]n
=∞summationdisplay
n=0hn∞summationdisplay
k=0tk×coefficient of tkin [Φ X(t)]n
=∞summationdisplay
n=0hn[ΦX(t)]n
=χN(ΦX(t)). (26.84)
In words, the PGF of the sum SNis given by the compound function χN(ΦX(t))
obtained by substituting Φ X(t)f o r tin the PGF for the number of terms Nin
the sum. We illustrate this with the following example.IThe probability distribution for the number of eggs in a clutch is Poisson distributed with
mean λ, and the probability that each egg will hatch is p(and is independent of the size of
the clutch). Use the results stated in (26.72) and (26.73) to show that the PGF (and hencethe probability distribution) for the numbe r of chicks that hatch corresponds to a Poisson
distribution having mean λp.
The number of chicks that hatch is given by a sum of the form (26.83) in which Xi=1i f
theith chick hatches and Xi= 0 if it does not. As given by (26.72), Φ X(t) is thus (1−p)+pt.
The value of Nis given by a Poisson distribution with mean λ; thus, from (26.73), in the
terminology of our previous discussion,
χN(t)=e−λeλt.
We now substitute these forms into (26.84) to obtain
ΞS(t)=e x p (−λ)exp[ λΦX(t)]
=e x p (−λ)exp{λ[(1−p)+pt]}
=e x p (−λp)exp( λpt).
But this is exactly the PGF of a Poisson distribution with mean λp.
That this implies that the probability is Poisson distributed is intuitively obvious since,
in the expansion of the PGF as a power series in t, every coefficient will be precisely
that implied by such a distribution. A solution of the same problem by direct calculation
appears in the answer to exercise 26.29.
J
26.7.2 Moment generating functions
As we saw in section 26.5 a probability function is often expressed in terms of
its moments. This leads naturally to the second type of generating function, amoment generating function . For a random variable X, and a real number t,t h e
moment generating function (MGF) is defined by
M
X(t)=Ebracketleftbig
etXbracketrightbig
=braceleftBiggsummationtext
ietxif(xi) for a discrete distribution,integraltext
etxf(x)dxfor a continuous distribution.(26.85)
1004
26.7 GENERATING FUNCTIONS
The MGF will exist for all values of tprovided that Xis bounded and always
exists at the point t=0w h e r e M(0) = E(1) = 1.
It will be apparent that the PGF and the MGF for a random variable X
are closely related. The former is the expectation of tXwhilst the latter is the
expectation of etX:
ΦX(t)=Ebracketleftbig
tXbracketrightbig
,M X(t)=Ebracketleftbig
etXbracketrightbig
.
The MGF can thus be obtained from the PGF by replacing tbyet,a n dv i c e
versa. The MGF has more general applicability, however, since it can be used
with both continuous and discrete distributions whilst the PGF is restricted to
non-negative integer distributions.
As its name suggests, the MGF is particularly useful for obtaining the moments
of a distribution, as is easily seen by noting that
Ebracketleftbig
etXbracketrightbig
=Ebracketleftbigg
1+tX+t2X2
2!+···bracketrightbigg
=1+ E[X]t+Ebracketleftbig
X2bracketrightbigt2
2!+···.
Assuming that the MGF exists for all taround the point t= 0, we can deduce
that the moments of a distribution are given in terms of its MGF by
E[Xn]=dnMX(t)
dtnvextendsinglevextendsinglevextendsinglevextendsingle
t=0. (26.86)
Similarly, by substitution in (26.51), the variance of the distribution is given by
V[X]=M/prime/prime
X(0)−bracketleftbig
M/prime
X(0)bracketrightbig2, (26.87)
where the prime denotes differentiation with respect to t.IThe MGF for the Gaussian distribution (see the end of subsection 26.9.1) is given by
MX(t)=e x p
/;
µt+1
2σ2t2
/
.
Find the expectation and variance of this distribution.
Using (26.86),
M/prime
X(t)=
/;
µ+σ2t
/
exp
/;
µt+1
2σ2t2
/
⇒ E[X]=M/prime
X(0) = µ,
M/prime/prime
X(t)=
/
σ2+(µ+σ2t)2
/
exp
/;
µt+1
2σ2t2
/
⇒ M/prime/prime
X(0) = σ2+µ2.
Thus, using (26.87),
V[X]=σ2+µ2−µ2=σ2.
That the mean is found to be µand the variance σ2justifies the use of these symbols in
the Gaussian distribution.
J
The moment generating function has several useful properties that follow from
its definition and can be employed in simplifying calculations.
1005
PROBABILITY
Scaling and shifting
IfY=aX+b,w h e r e aandbare arbitrary constants, then
MY(t)=Ebracketleftbig
etYbracketrightbig
=Ebracketleftbig
et(aX+b)bracketrightbig
=ebtEbracketleftbig
eatXbracketrightbig
=ebtMX(at). (26.88)
This result is often useful for obtaining the central moments of a distribution. If the
MFG of XisMX(t) then the variable Y=X−µhas the MGF MY(t)=e−µtMX(t),
which clearly generates the central moments of X,i . e .
E[(X−µ)n]=E[Yn]=M(n)
Y(0) =parenleftbiggdn
dtn[e−µtMX(t)]parenrightbigg
t=0.
Sums of random variables
IfX1,X2,...,X Nare independent random variables and SN=X1+X2+···+XN
then
MSN(t)=Ebracketleftbig
etSNbracketrightbig
=Ebracketleftbig
et(X1+X2+···+XN)bracketrightbig
=EbracketleftBiggNproductdisplay
i=1etXibracketrightBigg
.
Since the Xiareindependent ,
MSN(t)=Nproductdisplay
i=1Ebracketleftbig
etXibracketrightbig
=Nproductdisplay
i=1MXi(t). (26.89)
In words, the MGF of the sum of Nindependent random variables is the product
of their individual MGFs. By combining (26.89) with (26.88), we obtain the moregeneral result that the MGF of S
N=c1X1+c2X2+···+cNXN(where the ciare
constants) is given by
MSN(t)=Nproductdisplay
i=1MXi(cit). (26.90)
Variable-length sums of random variables
Let us consider the sum of Nindependent random variables Xi(i=1,2,...,N ), all
with the same probability distribution, and let us suppose that Nis itself a random
variable with a known distribution. Following the notation of section 26.7.1,
SN=X1+X2+···+XN,
where Nis a random variable with Pr( N=n)=hnand probability generating
function χN(t)=summationtexthntn. For definiteness, let us assume that the Xiare continuous
RVs (an analogous discussion can be given in the discrete case). Thus, the
1006
26.7 GENERATING FUNCTIONS
probability that value of SNlies in the interval stos+dsis given by†
Pr(s<S N≤s+ds)=∞summationdisplay
n=0Pr(N=n)P r (s<X 0+X1+X2···+Xn≤s+ds).
Let us denote Pr( s<S N≤s+ds)b yfN(s)dsand Pr( s<X 0+X1+X2···+Xn≤
s+ds)b yfn(s)ds. Thus, the kth moment of the PDF fN(s)i sg i v e nb y
µk=integraldisplay
skfN(s)ds=integraldisplay
sk∞summationdisplay
n=0Pr(N=n)fn(s)ds
=∞summationdisplay
n=0Pr(N=n)integraldisplay
skfn(s)ds
=∞summationdisplay
n=0hn×(k!×coefficient of tkin [MX(t)]n)
Thus the MGF of SNis given by
MSN(t)=∞summationdisplay
k=0µk
k!tk=∞summationdisplay
n=0hn∞summationdisplay
k=0tk×coefficient of tkin [MX(t)]n
=∞summationdisplay
n=0hn[MX(t)]n
=χN(MX(t)).
In words, the MGF of the sum SNis given by the compound function χN(MX(t))
obtained by substituting MX(t)f o r tin the PGF for the number of terms Nin
the sum.
Uniqueness
If the MGF of the random variable X1is identical to that for X2then the
probability distributions of X1andX2are identical. This is intuitively reasonable
although a rigorous proof is complicated, ‡and beyond the scope of this book.
26.7.3 Characteristic function
Thecharacteristic function (CF) of a random variable Xis defined as
CX(t)=Ebracketleftbig
eitXbracketrightbig
=braceleftBiggsummationtext
jeitxjf(xj) for a discrete distribution,integraltext
eitxf(x)dxfor a continuous distribution(26.91)
so that CX(t)=MX(it), where MX(t)i st h eM G Fo f X. Clearly, the characteristic
†As in the previous section, X0has to be formally included, since Pr( N= 0) may be non-zero.
‡See, for example, Moran, An Introduction to Probability Theory (Oxford Science Publications).
1007
PROBABILITY
function and the MGF are very closely related and can be used interchangeably.
Because of the formal similarity between the definitions of CX(t)a n d MX(t), the
characteristic function possesses analogous properties to those listed in the previ-ous section for the MGF, with only minor modifications. Indeed, by substituting it
fortin any of the relations obeyed by the MGF and noting that C
X(t)=MX(it),
we obtain the corresponding relationship for the characteristic function. Thus, forexample, the moments of Xare given in terms of the derivatives of C
X(t)b y
E[Xn]=(−i)nC(n)
X(0).
Similarly, if Y=aX+bthen CY(t)=eibtCX(t).
Whether to describe a random variable by its characteristic function or by its
MGF is partly a matter of personal preference. However, the use of the CF doeshave some advantages. Most importantly, the replacement of the exponential e
tX
in the definition of the MGF by the complex oscillatory function eitXin the CF
means that in the latter we avoid any difficulties associated with convergence of
the relevant sum or integral. Furthermore, when Xis a continous RV, we see
from (26.91) that CX(t) is related to the Fourier transform of the PDF f(x). As
a consequence of Fourier’s inversion theorem, we may obtain f(x)f r o m CX(t)b y
performing the inverse transform
f(x)=1
2πintegraldisplay∞
−∞CX(t)e−itxdt.
26.7.4 Cumulant generating function
As mentioned at the end of subsection 26.5.5, we may also describe a probability
density function f(x) in terms of its cumulants . These quantities may be expressed
in terms of the moments of the distribution and are important in sampling theory,which we discuss in the next chapter. The cumulants of a distribution are bestdefined in terms of its cumulant generating function (CGF), given by K
X(t)=
lnMX(t)w h e r e MX(t) is the MGF of the distribution. If KX(t)i se x p a n d e da sa
power series in tthen the kth cumulant κkoff(x) is the coefficient of tk/k!:
KX(t)=l n MX(t)≡κ1t+κ2t2
2!+κ3t3
3!+···. (26.92)
Since MX(0) = 1, KX(t) contains no constant term.IFind all the cumulants of the Gaussian distribution discussed in the previous example.
The moment generating function for the Gaussian distribution is MX(t)=e x p
/;
µt+1
2σ2t2
/
.
Thus, the cumulant generating function has the simple form
KX(t)=l n MX(t)=µt+1
2σ2t2.
Comparing this expression with (26.92), we find that κ1=µ,κ2=σ2and all other
cumulants are equal to zero.
J
1008
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
We may obtain expressions for the cumulants of a distribution in terms of its
moments by differentiating (26.92) with respect to tto give
dKX
dt=1
MXdMX
dt.
Expanding each term as power series in tand cross-multiplying, we obtain
parenleftbigg
κ1+κ2t+κ3t2
2!+···parenrightbiggparenleftbigg
1+µ1t+µ2t2
2!+···parenrightbigg
=parenleftbigg
µ1+µ2t+µ3t2
2!+···parenrightbigg
,
and, on equating coefficients of like powers of ton each side, we find
µ1=κ1,
µ2=κ2+κ1µ1,
µ3=κ3+2κ2µ1+κ1µ2,
µ4=κ4+3κ3µ1+3κ2µ2+κ1µ3,
...
µk=κk+k−1C1κk−1µ1+···+k−1Crκk−rµr+···+κ1µk−1.
Solving these equations for the κk, we obtain (for the first four cumulants)
κ1=µ1,
κ2=µ2−µ2
1=ν2,
κ3=µ3−3µ2µ1+2µ3
1=ν3,
κ4=µ4−4µ3µ1+1 2µ2µ2
1−3µ2
2−6µ4
1=ν4−3ν2
2. (26.93)
Higher-order cumulants may be calculated in the same way but become increas-
ingly lengthy to write out in full.
The principal property of cumulants is their additivity, which may be proved
by combining (26.92) with (26.90). If X1,X2,...,XNare independent random
variables and KXi(t)f o r i=1,2,...,N is the CGF for Xithen the CGF of
SN=c1X1+c2X2+···+cNXN(where the ciare constants) is given by
KSN(t)=Nsummationdisplay
i=1KXi(cit).
Cumulants also have the useful property that, under a change of origin X→
X+athe first cumulant undergoes the change κ1→κ1+abut all higher-order
cumulants remain unchanged. Under a change of scale X→bX, cumulant κr
undergoes the change κr→brκr.
26.8 Important discrete distributions
Having discussed the some general properties of distributions, we now consider
the more important discrete distributions encountered in physical applications.
1009
PROBABILITY
Distribution Probability law f(x)M G F E[X] V[X]
binomialnCxpxqn−x(pet+q)nnp npq
negative binomialr+x−1Cxprqx
/p
1−qet
/rrq
prq
p2
geometric qx−1ppet
1−qet1
pq
p2
hypergeometric(Np)!(Nq)!n!(N−n)!
x!(Np−x)!(n−x)!(Nq−n+x)!N!npN−n
N−1npq
Poissonλx
x!e−λeλ(et−1)λλ
Table 26.1 Some important discrete probability distributions.
These are discussed in detail below, and summarised for convenience in table 26.1;
we refer the reader to the relevant section below for an explanation of the symbolsused.
26.8.1 The binomial distribution
Perhaps the most important discrete probability distribution is the binomial dis-
tribution . This distribution describes processes that consist of a number of inde-
pendent identical trials with two possible outcomes, AandB=¯A.W em a yc a l l
these outcomes ‘success’ and ‘failure’ respectively. If the probability of a success
is Pr( A)=p, then the probability of a failure is Pr( B)=q=1−p.I fw ep e r f o r m
ntrials then the discrete random variable
X= number of times Aoccurs
can take the values 0 ,1,2,...,n; its distribution amongst these values is described
by the binomial distribution .
We now calculate the probability that in ntrials we obtain xsuccesses (and so
n−xfailures). One way of obtaining such a result is to have xsuccesses followed
byn−xfailures. Since the trials are assumed independent, the probability of this is
pp···pbracehtipupleft
bracehtipdownrightbracehtipdownleftbracehtipupright
xtimes×qq···qbracehtipupleftbracehtipdownrightbracehtipdownleftbracehtipupright
n−xtimes=pxqn−x.
This is, however, just one permutation of xsuccesses and n−xfailures. The total
number of permutations of nobjects, of which xare identical and of type 1 and
n−xare identical and of type 2, is given by (26.33) as
n!
x!(n−x)!≡nCx.
1010
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
1 11 1
2 22 2
33
33
444 4
55 5
5 6 6 7 78 8 9 910 100 0
00000 00.1 0.1
0.1 0.10.2 0.2
0.2 0.20.3 0.3
0.3 0.30.4 0.4
0.4 0.4
x xx x
f(x) f(x)f(x) f(x)
n=5 , p=0.6 n=5 , p=0.167
n= 10, p=0.6 n= 10, p=0.167
Figure 26.11 Some typical binomial distributions with various combinations
of parameters nandp.
Therefore, the total probability of obtaining xsuccesses from ntrials is
f(x)=P r ( X=x)=nCxpxqn−x=nCxpx(1−p)n−x, (26.94)
which is the binomial probability distribution formula . When a random variable
Xfollows the binomial distribution for ntrials, with a probability of success p,
we write X∼Bin(n, p). Then the random variable Xi so f t e nr e f e r r e dt oa sa
binomial variate . Some typical binomial distributions are shown in figure 26.11.IIf a single six-sided die is rolled five times, what is the probability that a six is thrown
exactly three times?
Here the number of ‘trials’ n= 5, and we are interested in the random variable
X= number of sixes thrown.
Since the probability of a ‘success’ is p=1
6, the probability of obtaining exactly three sixes
in five throws is given by (26.94) as
Pr(X=3 )=5!
3!(5−3)!
/1
6
/3
/5
6
/(5−3)
=0.032.
J
In evaluating binomial probabilities a useful result is the binomial recurrence
formula
Pr(X=x+1 )=p
qparenleftbiggn−x
x+1parenrightbigg
Pr(X=x), (26.95)
1011
PROBABILITY
which enables successive probabilities Pr( X=x+k),k=1,2,..., to be calculated
once Pr( X=x) is known; it is often quicker to use than (26.94).IThe random variable Xis distributed as X∼Bin(3 ,1
2). Evaluate the probability function
f(x)using the binomial recurrence formula.
The probability Pr( X= 0) may be calculated using (26.94) and is
Pr(X=0 )=3C0
/;1
2
/0
/;1
2
/3=1
8.
The ratio p/q=1
2/1
2= 1 in this case and so, using the binomial recurrence formula
(26.95), we find
Pr(X=1 )=1×3−0
0+1×1
8=3
8,
Pr(X=2 )=1×3−1
1+1×3
8=3
8,
Pr(X=3 )=1×3−2
2+1×3
8=1
8,
results which may be verified by direct application of (26.94).
J
We note that, as required, the binomial distribution satifies
nsummationdisplay
x=0f(x)=nsummationdisplay
x=0nCxpxqn−x=(p+q)n=1.
Furthermore, from the definitions of E[X]a n d V[X] for a discrete distribution,
we may show that for the binomial distribution E[X]=npandV[X]=npq.T h e
direct summations involved are, however, rather cumbersome and these resultsare obtained much more simply using the moment generating function.
The moment generating function for the binomial distribution
To find the MGF for the binomial distribution we consider the binomial random
variable Xto be the sum of the random variables X
i,i=1,2,...,n,w h i c ha r e
defined by
Xi=braceleftBigg
1 if a ‘success’ occurs on the ith trial,
0 if a ‘failure’ occurs on the ith trial.
Thus
Mi(t)=Ebracketleftbig
etXibracketrightbig
=e0t×Pr(Xi=0 )+ e1t×Pr(Xi=1 )
=1×q+et×p
=pet+q.
From (26.89), it follows that the MGF for the binomial distribution is given by
M(t)=nproductdisplay
i=1Mi(t)=(pet+q)n. (26.96)
1012
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
We can now use the moment generating function to derive the mean and
variance of the binomial distribution. From (26.96)
M/prime(t)=npet(pet+q)n−1,
and from (26.86)
E[X]=M/prime(0) = np(p+q)n−1=np,
where the last equality follows from p+q=1 .
Differentiating with respect to tonce more gives
M/prime/prime(t)=et(n−1)np2(pet+q)n−2+etnp(pet+q)n−1,
and from (26.86)
E[X2]=M/prime/prime(0) = n2p2−np2+np.
Thus, using (26.87)
V[X]=M/prime/prime(0)−bracketleftbig
M/prime(0)bracketrightbig2=n2p2−np2+np−n2p2=np(1−p)=npq.
Multiple binomial distributions
Suppose Xand Yare two independent random variables, both of which are
described by binomial distributions with a common probability of success p, but
with (in general) different numbers of trials n1andn2,s ot h a t X∼Bin(n1,p)
andY∼Bin(n2,p). Now consider the random variable Z=X+Y.W ec o u l d
calculate the probability distribution of Zdirectly using (26.60), but it is much
easier to use the MGF (26.96).
Since XandYare independent random variables, the MGF MZ(t) of the new
variable Z=X+Yis given simply by the product of the individual MGFs
MX(t)a n d MY(t). Thus, we obtain
MZ(t)=MX(t)MY(t)=(pet+q)n1(pet+q)n1=(pet+q)n1+n2,
which we recognise as the MGF of Z∼Bin(n1+n2,p). Hence Zis also described
by a binomial distribution.
This result may be extended to any number of binomial distributions. If Xi,
i=1,2,...,N , is distributed as Xi∼Bin(ni,p)t h e n Z=X1+X2+···+XNis
distributed as Z∼Bin(n1+n2+···+nN,p), as would be expected since the result
ofsummationtext
initrials cannot depend on how they are split up. A similar proof is also
possible using either the probability or cumulant generating functions.
Unfortunately, no equivalent simple result exists for the probability distribution
of the difference Z=X−Yof two binomially distributed variables.
1013
PROBABILITY
26.8.2 The geometric and negative binomial distributions
A special case of the binomial distribution occurs when instead of the number of
successes we consider the discrete random variable
X= number of trials required to obtain the first success .
The probability that xtrials are required in order to obtain the first success, is
simply the probability of obtaining x−1 failures followed by one success. If the
probability of a success on each trial is p,t h e nf o r x>0
f(x)=P r ( X=x)=( 1−p)x−1p=qx−1p,
where q=1−p. This distribution is sometimes called the geometric distribution .
The probability generating function for this distribution is given in (26.78). Byreplacing tbye
tin (26.78) we immediately obtain the MGF of the geometric
distribution
M(t)=pet
1−qet,
from which its mean and variance are found to be
E[X]=1
p,V [X]=q
p2.
Another distribution closely related to the binomial is the negative binomial
distribution. This describes the probability distribution of the random variable
X= number of failures before the rth success .
One way of obtaining xfailures before the rth success is to have r−1s u c c e s s e s
followed by xfailures followed by the rth success, for which the probability is
pp···pbracehtipupleftbracehtipdownrightbracehtipdownleftbracehtipupright
r−1t i m e s×qq···qbracehtipupleftbracehtipdownrightbracehtipdownleftbracehtipupright
xtimes×p=prqx.
However, the first r+x−1 factors constitute just one permutation of r−1
successes and xfailures. The total number of permutations of these r+x−1
objects, of which r−1 are identical and of type 1 and xare identical and of type
2, isr+x−1Cx. Therefore, the total probability of obtaining xfailures before the
rth success is
f(x)=P r ( X=x)=r+x−1Cxprqx,
which is called the negative binomial distribution (see the related discussion on
p. 979). It is straightforward to show that the MGF of this distribution is
M(t)=parenleftbiggp
1−qetparenrightbiggr
,
1014
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
and that its mean and variance are given by
E[X]=rq
pand V[X]=rq
p2.
26.8.3 The hypergeometric distribution
In subsection 26.8.1 we saw that the probability of obtaining xsuccesses in n
independent trials was given by the binomial distribution. Suppose that these n
‘trials’ actually consist of drawing at random nballs, from a set of Nsuch balls
of which Mare red and the rest white. Let us consider the random variable
X= number of red balls drawn.
On the one hand, if the balls are drawn with replacement then the trials are
independent and the probability of drawing a red ball is p=M/N each time.
Therefore, the probability of drawing xred balls in ntrials is given by the
binomial distribution as
Pr(X=x)=n!
x!(n−x)!px(1−p)n−x.
On the other hand, if the balls are drawn without replacement the trials are not
independent and the probability of drawing a red ball depends on how many redballs have already been drawn. We can, however, still derive a general formulafor the probability of drawing xred balls in ntrials, as follows.
The number of ways of drawing xred balls from Mis
MCx, and the number
of ways of drawing n−xwhite balls from N−MisN−MCn−x. Therefore, the
total number of ways to obtain xred balls in ntrials isMCxN−MCn−x. However,
the total number of ways of drawing nobjects from Nis simplyNCn. Hence the
probability of obtaining xred balls in ntrials is
Pr(X=x)=MCxN−MCn−x
NCn
=M!
x!(M−x)!(N−M)!
(n−x)!(N−M−n+x)!n!(N−n)!
N!,(26.97)
=(Np)!(Nq)!n!(N−n)!
x!(Np−x)!(n−x)!(Nq−n+x)!N!, (26.98)
where in the last line p=M/N andq=1−p.T h i si sc a l l e dt h e hypergeometric
distribution .
By performing the relevant summations directly, it may be shown that the
hypergeometric distribution has mean
E[X]=nM
N=np
and variance
V[X]=nM(N−M)(N−n)
N2(N−1)=N−n
N−1npq.
1015
PROBABILITYIIn the UK National Lottery each participant chooses six different numbers between 1
and49. In each weekly draw six numbered winning balls are subsequently drawn. Find the
probabilities that a participant chooses 0, 1,2,3, 4,5,6 winning numbers correctly.
The probabilities are given by a hypergeometric distribution with N(the total number of
balls) = 49, M(the number of winning balls drawn) = 6, and n(the number of numbers
chosen by each participant) = 6. Thus, substituting in (26.97), we find
Pr(0) =6C043C6
49C6=1
2.29,Pr(1) =6C143C5
49C6=1
2.42,
Pr(2) =6C243C4
49C6=1
7.55,Pr(3) =6C343C3
49C6=1
56.6,
Pr(4) =6C443C2
49C6=1
1032,Pr(5) =6C543C1
49C6=1
54200,
Pr(6) =6C643C0
49C6=1
13.98×106.
It can easily be seen that
6X
i=0Pr(i)=0 .44 + 0 .41 + 0 .13 + 0 .02 + O(10−3)=1 ,
as expected.
J
Note that if the number of trials (balls drawn) is small compared with N,M
andN−Mthen not replacing the balls is of little consequence, and we may
approximate the hypergeometric distribution by the binomial distribution (withp=M/N); this is much easier to evaluate.
26.8.4 The Poisson distribution
We have seen that the binomial distribution describes the number of successful
outcomes in a certain number of trials n. The Poisson distribution also describes
the probability of obtaining a given number of successes but for situationsin which the number of ‘trials’ cannot be enumerated; rather it describes thesituation in which discrete events occur in a continuum. Typical examples ofdiscrete random variables Xdescribed by a Poisson distribution are the number
of telephone calls received by a switchboard in a given interval, or the numberof stars above a certain brightness in a particular area of the sky. Given a mean
rate of occurrence λof these events in the relevant interval or area, the Poisson
distribution gives the probability Pr( X=x) that exactly xevents will occur.
We may derive the form of the Poisson distribution as the limit of the binomial
distribution when the number of trials n→∞ and the probability of ‘success’
p→0, in such a way that np=λremains finite. Thus, in our example of a
telephone switchboard, suppose we wish to find the probability that exactly x
calls are received during some time interval, given that the mean number of calls
1016
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
in such an interval is λ. Let us begin by dividing the time interval into a large
number, n, of equal shorter intervals, in each of which the probability of receiving
ac a l li s p.A sw el e t n→∞ then p→0, but since we require the mean number
of calls in the interval to equal λ, we must have np=λ. The probability of x
successes in ntrials is given by the binomial formula as
Pr(X=x)=n!
x!(n−x)!px(1−p)n−x. (26.99)
Now as n→∞, with xfinite, the ratio of the n-dependent factorials in (26.99)
behaves asymptotically as a power of n,i . e .
lim
n→∞n!
(n−x)!= lim
n→∞n(n−1)(n−2)···(n−x+1 )∼nx.
Also
lim
n→∞lim
p→0(1−p)n−x= lim
p→0(1−p)λ/p
(1−p)x=e−λ
1,
Thus, using λ=np, (26.99) tends to the Poisson distribution
f(x)=P r ( X=x)=e−λλx
x!, (26.100)
which gives the probability of obtaining exactly xcalls in the given time interval.
As we shall show below, λis the mean of the distribution. Events following a
Poisson distribution are usually said to occur randomly in time.
Alternatively we may derive the Poisson distribution directly, without consid-
ering a limit of the binomial distribution. Let us again consider our exampleof a telephone switchboard. Suppose that the probability that xcalls have been
received in a time interval tisP
x(t). If the average number of calls received in a
unit time is λthen in a further small time interval ∆ tthe probability of receiving
ac a l li s λ∆t,p r o v i d e d∆ tis short enough that the probability of receiving two or
more calls in this small interval is negligible. Similarly the probability of receivingno call during the same small interval is simply 1 −λ∆t.
Thus, for x>0, the probability of receiving exactly xcalls in the total interval
t+∆tis given by
P
x(t+∆t)=Px(t)(1−λ∆t)+Px−1(t)λ∆t.
Rearranging the equation, dividing through by ∆ tand letting ∆ t→0, we obtain
the differential recurrence equation
dPx(t)
dt=λPx−1(t)−λPx(t). (26.101)
Forx= 0 (i.e. no calls received), however, (26.101) simplifies to
dP0(t)
dt=−λP0(t),
1017
PROBABILITY
which may be integrated to give P0(t)=P0(0)e−λt. But since the probability P0(0)
of receiving no calls in a zero time interval must equal unity, we have P0(t)=e−λt.
This expression for P0(t) may then be substituted back into (26.101) with x=1
to obtain a differential equation for P1(t) that has the solution P1(t)=λte−λt.
We may repeat this process to obtain expressions for P2(t),P3(t),...,P x(t), and we
find
Px(t)=(λt)x
x!e−λt. (26.102)
By setting t= 1 in (26.102), we again obtain the Poisson distribution (26.100) for
obtaining exactly xcalls in a unit time interval.
If a discrete random variable is described by a Poisson distribution of mean λ
then we write X∼Po(λ). As it must be, the sum of the probabilities is unity:
∞summationdisplay
x=0Pr(X=x)=e−λ∞summationdisplay
x=0λx
x!=e−λeλ=1.
From (26.100) we may also derive the Poisson recurrence formula ,
Pr(X=x+1 )=λ
x+1Pr(X=x)f o r x=0,1,2,...,
(26.103)
which enables successive probabilities to be calculated easily once one is known.IA person receives on average one e-mail message per half-hour interval. Assuming that
the e-mails are received randomly in time, find the probabilities that in any particular hour
0,1,2,3,4,5messages are received.
LetX= number of e-mails received per hour. Clearly the mean number of e-mails per
hour is two, and so Xfollows a Poisson distribution with λ=2 ,i . e .
Pr(X=x)=2x
x!e−2.
Thus Pr( X=0 )= e−2=0.135, Pr( X=1 )=2 e−2=0.271, Pr( X=2 )=22e−2/2! = 0 .271,
Pr(X=3 )=23e−2/3! = 0 .180, Pr( X=4 )=24e−2/4! = 0 .090, Pr( X=5 )=25e−2/5! =
0.036. These results may also be calculated using the recurrence formula (26.103).
J
The above example illustrates the point that a Poisson distribution typically
rises and then falls. It either has a maximum when xis equal to the integer part
ofλor, if λhappens to be an integer, has equal maximal values at x=λ−1a n d
x=λ. The Poisson distribution always has a long ‘tail’ towards higher values of X
but the higher the value of the mean the more symmetric the distribution becomes.Typical Poisson distributions are shown in figure 26.12. Using the definitions ofmean and variance, we may show that, for the Poisson distribution, E[X]=λand
V[X]=λ. Nevertheless, as in the case of the binomial distribution, performing
the relevant summations directly is rather tiresome, and these results are muchmore easily proved using the MGF.
1018
26.8 IMPORTANT DISCRETE DISTRIBUTIONS
1
112
22 3
33 4
44 5
55 6
67
78910110
000
000.1 0.1
0.10.20.2 0.2
0.30.3 0.3
x
xxf(x)
f(x)f(x)
λ=1 λ=2
λ=5
Figure 26.12 Three Poisson distributions for different values of the parame-
terλ.
The moment generating function for the Poisson distribution
The MGF of the Poisson distribution is given by
MX(t)=Ebracketleftbig
etXbracketrightbig
=∞summationdisplay
x=0etxe−λλx
x!=e−λ∞summationdisplay
x=0(λet)x
x!=e−λeλet=eλ(et−1)
(26.104)
from which we obtain
M/prime
X(t)=λeteλ(et−1),
M/prime/prime
X(t)=(λ2e2t+λet)eλ(et−1).
Thus, the mean and variance of the Poisson distribution are given by
E[X]=M/prime
X(0) = λ and V[X]=M/prime/prime
X(0)−[M/prime
X(0)]2=λ.
The Poisson approximation to the binomial distribution
Earlier we derived the Poisson distribution as the limit of the binomial distribution
when n→∞andp→0i ns u c haw a yt h a t np=λremains finite, where λis the
1019
PROBABILITY
mean of the Poisson distribution. It is not surprising, therefore, that the Poisson
distribution is a very good approximation to the binomial distribution for largen(≥50, say) and small p(≤0.1, say). Moreover, it is easier to calculate as it
involves fewer factorials.IIn a large batch of light bulbs, the probability that a bulb is defective is 0.5%.F o ra
sample of 200bulbs taken at random, find the approximate probabilities that 0,1and2of
the bulbs respectively are defective.
Let the random variable X= number of defective bulbs in a sample. This is distributed
asX∼Bin(200, 0.005), implying that λ=np=1.0. Since nis large and psmall, we may
approximate the distribution as X∼Po(1), giving
Pr(X=x)≈e−11x
x!,
from which we find Pr( X=0 )≈0.37, Pr( X=1 )≈0.37, Pr( X=2 )≈0.18. For comparison,
it may be noted that the exact values calculated from the binomial distribution are identical
to those found here to two decimal places.
J
Multiple Poisson distributions
Mirroring our discussion of multiple binomial distributions in subsection 26.8.1,
let us suppose XandYare two independent random variables, both of which
are described by Poisson distributions with (in general) different means, so thatX∼Po(λ
1)a n d Y∼Po(λ2). Now consider the random variable Z=X+Y.W e
may calculate the probability distribution of Zdirectly using (26.60), but we may
derive the result much more easily by using the moment generating function (or
indeed the probability or cumulant generating functions).
Since XandYare independent RVs, the MGF for Zis simply the product of
the individual MGFs for XandY. Thus, from (26.104),
MZ(t)=MX(t)MY(t)=eλ1(et−1)eλ2(et−1)=e(λ1+λ2)(et−1),
which we recognise as the MGF of Z∼Po(λ1+λ2). Hence Zis also Poisson
distributed and has mean λ1+λ2. Unfortunately, no such simple result holds for
thedifference Z=X−Yof two independent Poisson variates. A closed-form
expression for the PDF of this Zdoes exist, but it is a rather complicated
combination of exponentials and a modified Bessel function. †ITwo types of e-mail arrive independently and at random: external e-mails at a mean rate
of one every five minutes and internal e-mails at a rate of two every five minutes. Calculatethe probability of receiving two or more e-mails in any two-minute interval.
Let
X= number of external e-mails per two-minute interval,
Y= number of internal e-mails per two-minute interval.
†For a derivation see, for example, Hobson & Lasenby, Monthly Notices of the Royal Astronomical
Society ,298, 905 (1998).
1020
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
Distribution Probability law f(x)M G F E[X] V[X]
Gaussian1
σ√
2πexp
/
−(x−µ)2
2σ2
/
exp(µt+1
2σ2t2) µσ2
exponential λe−λx
/λ
λ−t
/1
λ1
λ2
gammaλ
Γ(r)(λx)r−1e−λx
/λ
λ−t
/rr
λr
λ2
chi-squared1
2n/2Γ(n/2)x(n/2)−1e−x/2
/1
1−2t
/n/2
n 2n
uniform1
b−aebt−eat
(b−a)ta+b
2(b−a)2
12
Table 26.2 Some important continuous probability distributions.
Since we expect on average one external e-mail and two internal e-mails every five minutes
we have X∼Po(0.4) and Y∼Po(0.8). Letting Z=X+Ywe have Z∼Po(0.4+0 .8) =
Po(1.2). Now
Pr(Z≥2) = 1−Pr(Z<2) = 1−Pr(Z=0 )−Pr(Z=1 )
and
Pr(Z=0 )= e−1.2=0.301,
Pr(Z=1 )= e−1.21.2
1=0.361.
Hence Pr( Z≥2) = 1−0.301−0.361 = 0 .338.
J
The above result can be extended, of course, to any number of Poisson processes,
so that if Xi=P o ( λi),i=1,2,...,n then the random variable Z=X1+X2+
···+Xnis distributed as Z∼Po(λ1+λ2+···+λn).
26.9 Important continuous distributions
Having discussed the most commonly encountered discrete probability distri-
butions, we now consider some of the more important continuous probability
distributions. These are summarised for convenience in table 26.2; we refer thereader to the relevant subsection below for an explanation of the symbols used.
26.9.1 The Gaussian distribution
By far the most important continuous probability distribution is the Gaussian
ornormal distribution. The reason for its importance is that a great many
random variables of interest, in all areas of the physical sciences and beyond, are
described either exactly or approximately by a Gaussian distribution. Moreover,
the Gaussian distribution can be used to approximate other, more complicated,probability distributions.
1021
PROBABILITY
−6−4−2 2346810 120.10.20.30.4
σ=1
σ=2
σ=3µ=3
Figure 26.13 The Gaussian or normal distribution for mean µ=3a n d
various values of the standard deviation σ.
The probability density function for a Gaussian distribution of a random
variable X, with mean E[X]=µand variance V[X]=σ2,t a k e st h ef o r m
f(x)=1
σ√
2πexpbracketleftbigg
−1
2parenleftBigx−µ
σparenrightBig2bracketrightbigg
. (26.105)
The factor 1 /√
2πarises from the normalisation of the distribution,
integraldisplay∞
−∞f(x)dx=1 ;
the evaluation of this integral is discussed in subsection 6.4.2. The Gaussian
distribution is symmetric about the point x=µand has the characteristic ‘bell’
shape shown in figure 26.13. The width of the curve is described by the standarddeviation σ:i fσis large then the curve is broad, and if σis small then the curve
is narrow (see the figure). At x=µ±σ,f(x) falls to e
−1/2≈0.61 of its peak
value; these points are points of inflection, where d2f/dx2= 0. When a random
variable Xfollows a Gaussian distribution with mean µand variance σ2,w ew r i t e
X∼N(µ, σ2).
The effects of changing µandσare only to shift the curve along the x-axis or
to broaden or narrow it, respectively. Thus all Gaussians are equivalent in thata change of origin and scale can reduce them to a standard form. We thereforeconsider the random variable Z=(X−µ)/σ, for which the PDF takes the form
φ(z)=1
√
2πexpparenleftbigg
−z2
2parenrightbigg
, (26.106)
which is called the standard Gaussian distribution and has mean µ=0a n d
variance σ2= 1. The random variable Zis called the standard variable .
1022
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
y
−4 −12−2 −2 0 a1
2 4 az zφ(z)
Φ(z)
Φ(a)Φ(a)
0.20.40.60.8
0.10.20.30.4
Figure 26.14 On the left, the standard Gaussian distribution φ(z); the shaded
area gives Pr( Z<a )=Φ ( a). On the right, the cumulative probability function
Φ(z) for a standard Gaussian distribution φ(z).
From (26.105) we can define the cumulative probability function for a Gaussian
distribution as
F(x)=P r ( X<x )=1
σ√
2πintegraldisplayx
−∞expbracketleftbigg
−1
2parenleftBigu−µ
σparenrightBig2bracketrightbigg
du,
(26.107)
where uis a (dummy) integration variable. Unfortunately, this (indefinite) integral
cannot be evaluated analytically. It is therefore standard practice to tabulate val-ues of the cumulative probability function for the standard Gaussian distribution(see figure 26.14), i.e.
Φ(z)=P r ( Z<z )=1
√
2πintegraldisplayz
−∞expparenleftbigg
−u2
2parenrightbigg
du. (26.108)
It is usual only to tabulate Φ( z)f o r z>0, since it can be seen easily, from
figure 26.14 and the symmetry of the Gaussian distribution, that Φ( −z)=1−Φ(z);
see table 26.3. Using such a table it is then straightforward to evaluate the
probability that Zlies in a given range of z-values. For example, for aandb
constant,
Pr(Z<a )=Φ ( a),
Pr(Z>a )=1−Φ(a),
Pr(a<Z≤b)=Φ ( b)−Φ(a).
Remembering that Z=(X−µ)/σand comparing (26.107) and (26.108), we see
that
F(x)=ΦparenleftBigx−µ
σparenrightBig
,
and so we may also calculate the probability that the original random variable
1023
PROBABILITY
Φ(z).00 .01 .02 .03 .04 .05 .06 .07 .08 .09
0.0 .5000 .5040 .5080 .5120 .5160 .5199 .5239 .5279 .5319 .5359
0.1 .5398 .5438 .5478 .5517 .5557 .5596 .5636 .5675 .5714 .5753
0.2 .5793 .5832 .5871 .5910 .5948 .5987 .6026 .6064 .6103 .6141
0.3 .6179 .6217 .6255 .6293 .6331 .6368 .6406 .6443 .6480 .6517
0.4 .6554 .6591 .6628 .6664 .6700 .6736 .6772 .6808 .6844 .6879
0.5 .6915 .6950 .6985 .7019 .7054 .7088 .7123 .7157 .7190 .7224
0.6 .7257 .7291 .7324 .7357 .7389 .7422 .7454 .7486 .7517 .7549
0.7 .7580 .7611 .7642 .7673 .7704 .7734 .7764 .7794 .7823 .7852
0.8 .7881 .7910 .7939 .7967 .7995 .8023 .8051 .8078 .8106 .8133
0.9 .8159 .8186 .8212 .8238 .8264 .8289 .8315 .8340 .8365 .8389
1.0 .8413 .8438 .8461 .8485 .8508 .8531 .8554 .8577 .8599 .8621
1.1 .8643 .8665 .8686 .8708 .8729 .8749 .8770 .8790 .8810 .8830
1.2 .8849 .8869 .8888 .8907 .8925 .8944 .8962 .8980 .8997 .9015
1.3 .9032 .9049 .9066 .9082 .9099 .9115 .9131 .9147 .9162 .9177
1.4 .9192 .9207 .9222 .9236 .9251 .9265 .9279 .9292 .9306 .9319
1.5 .9332 .9345 .9357 .9370 .9382 .9394 .9406 .9418 .9429 .9441
1.6 .9452 .9463 .9474 .9484 .9495 .9505 .9515 .9525 .9535 .9545
1.7 .9554 .9564 .9573 .9582 .9591 .9599 .9608 .9616 .9625 .9633
1.8 .9641 .9649 .9656 .9664 .9671 .9678 .9686 .9693 .9699 .9706
1.9 .9713 .9719 .9726 .9732 .9738 .9744 .9750 .9756 .9761 .9767
2.0 .9772 .9778 .9783 .9788 .9793 .9798 .9803 .9808 .9812 .9817
2.1 .9821 .9826 .9830 .9834 .9838 .9842 .9846 .9850 .9854 .9857
2.2 .9861 .9864 .9868 .9871 .9875 .9878 .9881 .9884 .9887 .9890
2.3 .9893 .9896 .9898 .9901 .9904 .9906 .9909 .9911 .9913 .9916
2.4 .9918 .9920 .9922 .9925 .9927 .9929 .9931 .9932 .9934 .9936
2.5 .9938 .9940 .9941 .9943 .9945 .9946 .9948 .9949 .9951 .9952
2.6 .9953 .9955 .9956 .9957 .9959 .9960 .9961 .9962 .9963 .9964
2.7 .9965 .9966 .9967 .9968 .9969 .9970 .9971 .9972 .9973 .9974
2.8 .9974 .9975 .9976 .9977 .9977 .9978 .9979 .9979 .9980 .9981
2.9 .9981 .9982 .9982 .9983 .9984 .9984 .9985 .9985 .9986 .9986
3.0 .9987 .9987 .9987 .9988 .9988 .9989 .9989 .9989 .9990 .9990
3.1 .9990 .9991 .9991 .9991 .9992 .9992 .9992 .9992 .9993 .9993
3.2 .9993 .9993 .9994 .9994 .9994 .9994 .9994 .9995 .9995 .9995
3.3 .9995 .9995 .9995 .9996 .9996 .9996 .9996 .9996 .9996 .9997
3.4 .9997 .9997 .9997 .9997 .9997 .9997 .9997 .9997 .9997 .9998
Table 26.3 The cumulative probability function Φ( z) for the standard Gaus-
sian distribution, as given by (26.108). The units and the first decimal place
ofzare specified in the column under Φ( z) and the second decimal place is
specified by the column headings. Thus, for example, Φ(1 .23) = 0 .8907.
1024
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
Xlies in a given x-range. For example,
Pr(a<X≤b)=1
σ√
2πintegraldisplayb
aexpbracketleftbigg
−1
2parenleftBigu−µ
σparenrightBig2bracketrightbigg
du (26.109)
=F(b)−F(a) (26.110)
=Φparenleftbiggb−µ
σparenrightbigg
−ΦparenleftBiga−µ
σparenrightBig
. (26.111)IIfXis described by a Gaussian distribution of mean µand variance σ2,c a l c u l a t et h e
probabilities that Xlies within 1σ,2σand3σof the mean.
From (26.111)
Pr(µ−nσ < X≤µ+nσ)=Φ ( n)−Φ(−n)=Φ ( n)−[1−Φ(n)],
and so from table 26.3
Pr(µ−σ<X≤µ+σ)=2 Φ ( 1 )−1=0 .6826≈68.3%,
Pr(µ−2σ<X≤µ+2σ)=2 Φ ( 2 )−1=0 .9544≈95.4%,
Pr(µ−3σ<X≤µ+3σ)=2 Φ ( 3 )−1=0 .9974≈99.7%.
Thus we expect Xto be distributed in such a way that about two thirds of the values will
lie between µ−σandµ+σ, 95% will lie within 2 σof the mean and 99 .7% will lie within
3σof the mean. These limits are called the one-, t wo- and three-sigma limits respectively;
it is particularly important to note that they are independent of the actual values of the
mean and variance.
J
There are many other ways in which the Gaussian distribution may be used.
We now illustrate some of the uses in more complicated examples.ISawmill Aproduces boards whose lengths are Gaussian distributed with mean 209.4 cm
and standard deviation 5.0 cm . A board is accepted if it is longer than 200 cm but is
rejected otherwise. Show that 3%of boards are rejected.
Sawmill Bproduces boards of the same standard deviation but of mean length 210.1 cm .
Find the proportion of boards rejected if they are drawn at random from the outputs of A
andBin the ratio 3:1.
LetX= length of boards from A,s ot h a t X∼N(209.4,(5.0)2)a n d
Pr(X<200) = Φ
/200−µ
σ
/
=Φ
/200−209.4
5.0
/
=Φ (−1.88).
But, since Φ( −z)=1−Φ(z) we have, using table 26.3,
Pr(X<200) = 1−Φ(1.88) = 1−0.9699 = 0 .0301,
i.e. 3.0% of boards are rejected.
Now let Y= length of boards from B,s ot h a t Y∼N(210.1,(5.0)2)a n d
Pr(Y<200) = Φ
/200−210.1
5.0
/
=Φ (−2.02)
=1−Φ(2.02)
=1−0.9783 = 0 .0217.
1025
PROBABILITY
Therefore, when taken alone, only 2 .2% of boards from Bare rejected. If, however, boards
are drawn at random from AandBin the ratio 3 : 1 then the proportion rejected is
1
4(3×0.030 + 1×0.022) = 0 .028 = 2 .8%.
J
We may sometimes work backwards to derive the mean and standard deviation
of a population that is known to be Gaussian distributed.IThe time taken for a computer ‘packet’ to travel from Cambridge UK to Cambridge MA
is Gaussian distributed. 6.8% of the packets take over 200 ms to make the journey, and
3.0% take under 140 ms . Find the mean and standard deviation of the distribution.
LetX= journey time in ms; we are told that X∼N(µ, σ2)w h e r e µandσare unknown.
Since 6.8% of journey times are longer than 200 ms,
Pr(X>200) = 1−Φ
/200−µ
σ
/
=0.068,
from which we find
Φ
/200−µ
σ
/
=1−0.068 = 0 .932.
Using table 26.3, we have therefore
200−µ
σ=1.49. (26.112)
Also, 3 .0% of journey times are under 140 ms, so
Pr(X<140) = Φ
/140−µ
σ
/
=0.030.
Now using Φ( −z)=1−Φ(z)g i v e s
Φ
/µ−140
σ
/
=1−0.030 = 0 .970.
Using table 26.3 again, we find
µ−140
σ=1.88. (26.113)
Solving the simultaneous equations (26.112) and (26.113) gives µ= 173 .5,σ=1 7.8.
J
The moment generating function for the Gaussian distribution
Using the definition of the MGF (26.85),
MX(t)=Ebracketleftbig
etXbracketrightbig
=integraldisplay∞
−∞1
σ√
2πexpbracketleftbigg
tx−(x−µ)2
2σ2bracketrightbigg
dx
=cexpparenleftbig
µt+1
2σ2t2parenrightbig
,
where the final equality is established by completing the square in the argument
of the exponential and writing
c=integraldisplay∞
−∞1
σ√
2πexpbraceleftbigg
−[x−(µ+σ2t)]2
2σ2bracerightbigg
dx.
1026
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
However, the final integral is simply the normalisation integral for the Gaussian
distribution, and so c=1a n dt h eM G Fi sg i v e nb y
MX(t)=e x pparenleftbig
µt+1
2σ2t2parenrightbig
. (26.114)
We showed in subsection 26.7.2 that this MGF leads to E[X]=µandV[X]=σ2,
as required.
Gaussian approximation to the binomial distribution
We may consider the Gaussian distribution as the limit of the binomial distribu-
tion when the number of trials n→∞but the probability of a success premains
finite, so that np→∞ also. (This contrasts with the Poisson distribution, which
corresponds to the limit n→∞ andp→0 with np=λremaining finite.) In
other words, a Gaussian distribution results when an experiment with a finite
probability of success is repeated a large number of times. We now show howthis Gaussian limit arises.
The binomial probability function gives the probability of xsuccesses in ntrials
as
f(x)=n!
x!(n−x)!px(1−p)n−x.
Taking the limit as n→∞ (and x→∞) we may approximate the factorials by
Stirling’s approximation
n!∼√
2πnparenleftBign
eparenrightBign
to obtain
f(x)≈1√
2πnparenleftBigx
nparenrightBig−x−1/2parenleftBign−x
nparenrightBig−n+x−1/2
px(1−p)n−x
=1√
2πnexpbracketleftBig
−parenleftbig
x+1
2parenrightbig
lnx
n−parenleftbig
n−x+1
2parenrightbig
lnn−x
n
+xlnp+(n−x)l n ( 1−p)bracketrightBig
.
By expanding the argument of the exponential in terms of y=x−np,w h e r e
1/lessmuchy/lessmuchnpand keeping only the dominant terms, it can be shown that
f(x)≈1√
2πn1√p(1−p)expbracketleftbigg
−1
2(x−np)2
np(1−p)bracketrightbigg
,
which is of Gaussian form with µ=npandσ=√np(1−p).
Thus we see that the value of the Gaussian probability density function f(x)i s
a good approximation to the probability of obtaining xsuccesses in ntrials. This
approximation is actually very good even for relatively small n. For example, if
n=1 0a n d p=0.6 then the Gaussian approximation to the binomial distribution
is (26.105) with µ=1 0×0.6=6a n d σ=√10×0.6(1−0.6) = 1 .549. The
1027
PROBABILITY
xf (x) (binomial) f(x) (Gaussian)
0 0.0001 0.0001
1 0.0016 0.00142 0.0106 0.0092
3 0.0425 0.0395
4 0.1115 0.11195 0.2007 0.2091
6 0.2508 0.2575
7 0.2150 0.20918 0.1209 0.1119
9 0.0403 0.0395
10 0.0060 0.0092
Table 26.4 Comparison of the binomial distribution for n=1 0a n d p=0.6
with its Gaussian approximation.
probability functions f(x) for the binomial and associated Gaussian distributions
for these parameters are given in table 26.4, and it can be seen that the Gaussianapproximation is a good one.
Strictly speaking, however, since the Gaussian distribution is continuous and
the binomial distribution is discrete, we should use the integral of f(x)f o rt h e
Gaussian distribution in the calculation of approximate binomial probabilities.More specifically, we should apply a continuity correction so that the discrete
integer xin the binomial distribution becomes the interval [ x−0.5,x+0.5] in
the Gaussian distribution. Explicitly,
Pr(X=x)≈1
σ√
2πintegraldisplayx+0.5
x−0.5expbracketleftbigg
−1
2parenleftBigu−µ
σparenrightBig2bracketrightbigg
du.
The Gaussian approximation is particularly useful for estimating the binomial
probability that Xlies between the (integer) values x1andx2,
Pr(x1<X≤x2)≈1
σ√
2πintegraldisplayx2+0.5
x1−0.5expbracketleftbigg
−1
2parenleftBigu−µ
σparenrightBig2bracketrightbigg
du.IA manufacturer makes computer chips of which 10%are defective. For a random sample
of200chips, find the approximate probability that more than 15are defective.
We first define the random variable
X= number of defective chips in the sample ,
which has a binomial distribution X∼Bin(200,0.1). Therefore, t he mean and variance of
this distribution are
E[X] = 200×0.1 = 20 and V[X] = 200×0.1×(1−0.1) = 18 ,
and we may approximate the binomial distribution with a Gaussian distribution such that
1028
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
X∼N(20,18). The standard variable is
Z=X−20√
18,
and so, using X=1 5.5 to allow for the continuity correction,
Pr(X>15.5) = Pr
/
Z>15.5−20√
18
/
=P r ( Z>−1.06)
=P r ( Z<1.06) = 0 .86.
J
Gaussian approximation to the Poisson distribution
We first met the Poisson distribution as the limit of the binomial distribution for
n→∞ andp→0 ,t a k e ni ns u c haw a yt h a t np=λremains finite. Further, in
the previous subsection, we considered the Gaussian distribution as the limit ofthe binomial distribution when n→∞butpremains finite, so that np→∞also.
It should come as no surprise, therefore, that the Gaussian distribution can also
be used to approximate the Poisson distribution when the mean λbecomes large.
The probability function for the Poisson distribution is
f(x)=e
−λλx
x!,
which, on taking the logarithm of both sides, gives
lnf(x)=−λ+xlnλ−lnx!. (26.115)
Stirling’s approximation for large xgives
x!≈√
2πxparenleftBigx
eparenrightBigx
implying that
lnx!≈ln√
2πx+xlnx−x,
which, on substituting into (26.115), yields
lnf(x)≈−λ+xlnλ−(xlnx−x)−ln√
2πx.
Since we expect the Poisson distribution to peak around x=λ, we substitute
/epsilon1=x−λto obtain
lnf(x)≈−λ+(λ+/epsilon1)braceleftBig
lnλ−lnbracketleftBig
λparenleftBig
1+/epsilon1
λparenrightBigbracketrightBigbracerightBig
+(λ+/epsilon1)−lnradicalbig
2π(λ+/epsilon1).
Using the expansion ln(1 + z)=z−z2/2+···, we find
lnf(x)≈/epsilon1−(λ+/epsilon1)parenleftbigg/epsilon1
λ−/epsilon12
2λ2parenrightbigg
−ln√
2πλ−parenleftbigg/epsilon1
λ−/epsilon12
2λ2parenrightbigg
≈−/epsilon12
2λ−ln√
2πλ,
1029
PROBABILITY
when only the dominant terms are retained, after using the fact that /epsilon1is of the
order of the standard deviation of x,i . e .o fo r d e r λ1/2. On exponentiating this
result we obtain
f(x)≈1√
2πλexpbracketleftbigg
−(x−λ)2
2λbracketrightbigg
,
which is the Gaussian distribution with µ=λandσ2=λ.
The larger the value of λ, the better is the Gaussian approximation to the
Poisson distribution; the approximation is reasonable even for λ= 5, but λ≥10
is safer. As in the case of the Gaussian approximation to the binomial distribution,a continuity correction is necessary since the Poisson distribution is discrete.IE-mail messages are received by an author at an average rate of one per hour. Find the
probability that in a day the author receives 24messages or more.
We first define the random variable
X= number of messages received in a day .
Thus E[X]=1×24 = 24, and so X∼Po(24). Since λ>10 we may approximate the
Poisson distribution by X∼N(24,24). Now the standard variable is
Z=X−24√
24,
and, using the continuity correction, we find
Pr(X>23.5) = P
/
Z>23.5−24√
24
/
=P r ( Z>−0.102) = Pr( Z<0.102) = 0 .54.
J
In fact, almost all probability distributions tend towards a Gaussian when the
numbers involved become large – that this should happen is required by thecentral limit theorem, which we discuss in section 26.10.
Multiple Gaussian distributions
Suppose Xand Yareindependent Gaussian-distributed random variables, so
thatX∼N(µ
1,σ2
1)a n d Y∼N(µ2,σ2
2). Let us now consider the random variable
Z=X+Y. The PDF for this random variable may be found directly using
(26.61), but it is easier to use the MGF. From (26.114), the MGFs of XandY
are
MX(t)=e x pparenleftbig
µ1t+1
2σ2
1t2parenrightbig
,M Y(t)=e x pparenleftbig
µ2t+1
2σ2
2t2parenrightbig
.
Using (26.89), since XandYare independent RVs, the MGF of Z=X+Yis
simply the product of MX(t)a n d MY(t). Thus, we have
MZ(t)=MX(t)MY(t)=e x pparenleftbig
µ1t+1
2σ2
1t2parenrightbig
expparenleftbig
µ2t+1
2σ2
2t2parenrightbig
=e x pbracketleftbig
(µ1+µ2)t+1
2(σ2
1+σ2
2)t2bracketrightbig
,
1030
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
which we recognise as the MGF for a Gaussian with mean µ1+µ2and variance
σ2
1+σ2
2. Thus, Zis also Gaussian distributed: Z∼N(µ1+µ2,σ2
1+σ2
2).
A similar calculation may be performed to calculate the PDF of the random
variable W=X−Y. If we introduce the variable ˜Y=−Ythen W=X+˜Y,
where ˜Y∼N(−µ1,σ2
1). Thus, using the result above, we find W∼N(µ1−
µ2,σ2
1+σ2
2).IAn executive travels home from her office every evening. Her journey consists of a train
ride, followed by a bicycle ride. The time spent on the train is Gaussian distributed withmean 52minutes and standard deviation 1.8minutes, while the time for the bicycle journey
is Gaussian distributed with mean 8minutes and standard deviation 2.6minutes. Assuming
these two factors are independent, estimate the percentage of occasions on which the whole
journey takes more than 65minutes.
We first define the random variables
X=t i m es p e n to nt r a i n ,Y = time spent on bicycle,
so that X∼N(52,(1 .8)2)a n d Y∼N(8,(2.6)2). Since XandYare independent, the total
journey time T=X+Yis distributed as
T∼N(52 + 8 ,(1.8)2+( 2.6)2)=N(60,(3.16)2).
The standard variable is thus
Z=T−60
3.16,
and the required probability is given by
Pr(T>65) = Pr
/
Z>65−60
3.16
/
=P r ( Z>1.58) = 1−0.943 = 0 .057.
Thus the total journey time exceeds 65 minutes on 5 .7% of occasions.
J
The above results may be extended. For example, if the random variables
Xi,i=1,2,...,n, are distributed as Xi∼N(µi,σ2
i) then the random variable
Z=summationtext
iciXi(where the ciare constants) is distributed as Z∼N(summationtext
iciµi,summationtext
ic2
iσ2
i).
26.9.2 The log-normal distribution
If the random variable Xfollows a Gaussian distribution then the variable
Y=eXis described by a log-normal distribution. Clearly, if Xcan take values
in the range −∞to∞,t h e n Ywill lie between 0 and ∞. The probability density
function for Yis found using the result (26.58). It is
g(y)=f(x(y))vextendsinglevextendsinglevextendsinglevextendsingledx
dyvextendsinglevextendsinglevextendsinglevextendsingle=1
σ√
2π1
yexpbracketleftbigg
−(lny−µ)2
2σ2bracketrightbigg
.
We note that µand σ2are not the mean and variance of the log-normal
distribution, but rather the parameters of t he corresponding Gaussian distribution
forX. The mean and variance of Y, however, can be found straightforwardly
1031
PROBABILITY
000.20.40.60.8
11
2 3 4yg(y)
µ=0 , σ=0
µ=0 , σ=0.5
µ=0 , σ=1.5
µ=1 , σ=1
Figure 26.15 The PDF g(y) for the log-normal distribution for various values
of the parameters µandσ.
using the MGF of X,w h i c hr e a d s MX(t)=E[etX]=e x p ( µt+1
2σ2t2). Thus, the
mean of Yis given by
E[Y]=E[eX]=MX(1) = exp( µ+1
2σ2),
and the variance of Yreads
V[Y]=E[Y2]−(E[Y])2=E[e2X]−(E[eX])2
=MX(2)−[MX(1)]2=e x p ( 2 µ+σ2)[exp( σ2)−1].
In figure 26.15, we plot some examples of the log-normal distribution for various
values of the parameters µandσ2.
26.9.3 The exponential and gamma distributions
The exponential distribution with positive parameter λis given by
f(x)=braceleftBigg
λe−λxforx>0,
0f o r x≤0(26.116)
and satisfiesintegraltext∞
−∞f(x)dx= 1 as required. The exponential distribution occurs nat-
urally if we consider the distribution of the length of intervals between successive
events in a Poisson process or, equivalently, the distribution of the interval (i.e.the waiting time) before the first event. If the average number of events per unitinterval is λthen on average there are λxevents in interval x, so that from the
Poisson distribution the probability that there will be no events in this interval isgiven by
Pr(no events in interval x)=e
−λx.
1032
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
The probability that an event occurs in the next infinitestimal interval [ x, x+dx]
is given by λd x,s ot h a t
Pr(the first event occurs in interval [ x, x+dx]) =e−λxλd x .
Hence the required probability density function is given by
f(x)=λe−λx.
The expectation and variance of the exponential distribution can be evaluated as
1/λand (1 /λ)2respectively. The MGF is given by
M(t)=λ
λ−t. (26.117)
We may generalise the above discussion to obtain the PDF for the interval
between every rth event in a Poisson process or, equivalently, the interval (waiting
time) before the rth event. We begin by using the Poisson distribution to give
Pr(r−1 events occur in interval x)=e−λx(λx)r−1
(r−1)!,
from which we obtain
Pr(rth event occurs in the interval [ x, x+dx]) =e−λx(λx)r−1
(r−1)!λd x .
Thus the required PDF is
f(x)=λ
(r−1)!(λx)r−1e−λx, (26.118)
which is known as the gamma distribution of order rwith parameter λ. Although
our derivation applies only when ris a positive integer, the gamma distribution is
defined for all positive rby replacing ( r−1)! by Γ( r) in (26.118); see the appendix
for a discussion of the gamma function Γ( x). If a random variable Xis described
by a gamma distribution of order rwith parameter λ,w ew r i t e X∼γ(λ, r);
we note that the exponential distribution is the special case γ(λ,1). The gamma
distribution γ(λ, r) is plotted in figure 26.16 for λ=1a n d r=1,2,5,10. For
large r, the gamma distribution tends to the Gaussian distribution whose mean
and variance are specified by (26.120) below.
The MGF for the gamma distribution is obtained from that for the exponential
distribution, by noting that we may consider the interval between every rth event
in a Poisson process as the sum of rintervals between successive events. Thus the
rth-order gamma variate is the sum of rindependent exponentially distributed
random variables. From (26.117) and (26.90), the MGF of the gamma distribution
is therefore given by
M(t)=parenleftbiggλ
λ−tparenrightbiggr
, (26.119)
1033
PROBABILITY
0
024 68 1 0 12 14 16 18 200.20.40.60.81
r=1
r=2
r=5
r=1 0
xf(x)
Figure 26.16 The PDF f(x) for the gamma distributions γ(λ, r)w i t h λ=1
andr=1,2,5,10.
from which the mean and variance are found to be
E[X]=r
λ,V [X]=r
λ2. (26.120)
We may also use the above MGF to prove another useful theorem regarding
multiple gamma distributions. If Xi∼γ(λ, ri),i=1,2,...,n, are independent
gamma variates then the random variable Y=X1+X2+···+Xnhas MGF
M(t)=nproductdisplay
i=1parenleftbiggλ
λ−tparenrightbiggri
=parenleftbiggλ
λ−tparenrightbiggr1+r2+···+rn
. (26.121)
Thus Yis also a gamma variate, distributed as Y∼γ(λ, r1+r2+···+rn).
26.9.4 The chi-squared distribution
In subsection 26.6.2, we showed that if Xis Gaussian distributed with mean µand
variance σ2, such that X∼N(µ, σ2), then the random variable Y=(x−µ)2/σ2
is distributed as the gamma distribution Y∼γ(1
2,1
2). Let us now consider n
independent Gaussian random variables Xi∼N(µi,σ2
i),i=1,2,...,n, and define
the new variable
χ2
n=nsummationdisplay
i=1(Xi−µi)2
σ2
i. (26.122)
1034
26.9 IMPORTANT CONTINUOUS DISTRIBUTIONS
Using the result (26.121) for multiple gamma distributions, χ2
nmust be distributed
as the gamma variate χ2
n∼γ(1
2,1
2n), which from (26.118) has the PDF
f(χ2
n)=1
2
Γ(1
2n)(1
2χ2
n)(n/2)−1exp(−1
2χ2
n)
=1
2n/2Γ(1
2n)(χ2
n)(n/2)−1exp(−1
2χ2
n). (26.123)
This is known as the chi-squared distribution of order nand has numerous
applications in statistics (see chapter 27). Setting λ=1
2andr=1
2nin (26.120),
we find that
E[χ2
n]=n, V [χ2
n]=2 n.
An important generalisation occurs when the nGaussian variables Xiarenot
linearly independent but are instead required to satisfy a linear constraint of theform
c
1X1+c2X2+···+cnXn=0, (26.124)
in which the constants ciare not all zero. In this case, it may be shown (see
exercise 26.40) that the variable χ2
ndefined in (26.122) is still described by a chi-
squared distribution, but one of order n−1. Indeed, this result may be trivially
extended to show that if the nGaussian variables Xisatisfy mlinear constraints
of the form (26.124) then the variable χ2
ndefined in (26.122) is described by a
chi-squared distribution of order n−m.
26.9.5 The Cauchy and Breit–Wigner distributions
A random variable X(in the range −∞to∞)t h a to b e y st h e Cauchy distribution
is described by the PDF
f(x)=1
π1
1+x2.
This is a special case of the Breit–Wigner distribution
f(x)=1
π1
2Γ
1
4Γ2+(x−x0)2,
which is encountered in the study of nuclear and particle physics. In figure 26.17,
we plot some examples of the Breit–Wigner distribution for several values of theparameters x
0and Γ.
We see from the figure that the peak (or mode) of the distribution occurs
atx=x0. It is also straightforward to show that the parameter Γ is equal to
the width of the peak at half the maximum height. Although the Breit–Wignerdistribution is symmetric about its peak, it does not formally possess a mean since
1035
PROBABILITY
000.20.40.60.8
−4−22 4xf(x)
x0=0 ,
x0=0 ,x0=2 ,
Γ=1 Γ=1
Γ=3
Figure 26.17 The PDF f(x) for the Breit–Wigner distribution for different
values of the parameters x0and Γ.
the integralsintegraltext0
−∞xf(x)dxandintegraltext∞
0xf(x)dxboth diverge. Similar divergences occur
for all higher moments of the distribution.
26.9.6 The uniform distribution
Finally we mention the very simple, but common, uniform distribution ,w h i c h
describes a continuous random variable that has a constant PDF over its allowedrange of values. If the limits on Xareaandbthen
f(x)=braceleftBigg
1/(b−a)f o r a≤x≤b,
0o t h e r w i s e .
The MGF of the uniform distribution is found to be
M(t)=e
bt−eat
(b−a)t,
and its mean and variance are given by
E[X]=a+b
2,V [X]=(b−a)2
12.
26.10 The central limit theorem
In subsection 26.9.1 we discussed approximating the binomial and Poisson distri-
butions by the Gaussian distribution when the number of trials is large. We now
discuss why the Gaussian distribution is so common and therefore so important.Thecentral limit theorem may be stated as follows.
1036
26.10 THE CENTRAL LIMIT THEOREM
Central limit theorem. Suppose that Xi,i=1,2,...,n,a r eindependent random
variables, each of which is described by a probability density function fi(x)(these
may all be different) with a mean µiand a variance σ2
i. The random variable Z=parenleftbigsummationtext
iXiparenrightbig
/n, i.e. the ‘mean’ of the Xi, has the following properties:
(i)its expectation value is given by E[Z]=parenleftbigsummationtext
iµiparenrightbig
/n;
(ii)its variance is given by V[Z]=parenleftbigsummationtext
iσ2
iparenrightbig
/n2;
(iii)asn→∞ the probability function of Ztends to a Gaussian with corre-
sponding mean and variance.
We note that for the theorem to hold, the probability density functions fi(x)
must possess formal means and variances. Thus, for example, if each Xiwere
described by a Cauchy distribution then the theorem would not apply.
Properties (i) and (ii) of the theorem are easily proved, as follows. Firstly
E[Z]=1
n(E[X1]+E[X2]+···+E[Xn]) =1
n(µ1+µ2+···+µn)=summationtext
iµi
n,
a result which does notrequire that the Xiareindependent random variables. If
µi=µfor all ithen this becomes
E[Z]=nµ
n=µ.
Secondly, if the Xiareindependent, it follows from an obvious extension of
(26.68) that
V[Z]=Vbracketleftbigg1
n(X1+X2+···+Xn)bracketrightbigg
=1
n2(V[X1]+V[X2]+···+V[Xn])=summationtext
iσ2
i
n2.
Let us now consider property (iii), which is the reason for the ubiquity of
the Gaussian distribution and is most easily proved by considering the momentgenerating function M
Z(t)o fZ. From (26.90), this MGF is given by
MZ(t)=nproductdisplay
i=1MXiparenleftbiggt
nparenrightbigg
,
where MXi(t)i st h eM G Fo f fi(x). Now
MXiparenleftbiggt
nparenrightbigg
=1+t
nE[Xi]+1
2t2
n2E[X2
i]+···
=1+ µit
n+1
2(σ2
i+µ2
i)t2
n2+···,
and as nbecomes large
MXiparenleftbiggt
nparenrightbigg
≈expparenleftbiggµit
n+1
2σ2
it2
n2parenrightbigg
,
1037
PROBABILITY
as may be verified by expanding the exponential up to terms including ( t/n)2.
Therefore
MZ(t)≈nproductdisplay
i=1expparenleftbiggµit
n+1
2σ2
it2
n2parenrightbigg
=e x pparenleftbiggsummationtext
iµi
nt+1
2summationtext
iσ2
i
n2t2parenrightbigg
.
Comparing this with the form of the MGF for a Gaussian distribution, (26.114),
we can see that the probability density function g(z)o fZtends to a Gaussian dis-
tribution with meansummationtext
iµi/nand variancesummationtext
iσ2
i/n2. In particular, if we consider
Zto be the mean of nindependent measurements of the samerandom variable X
(so that Xi=Xfori=1,2,...,n) then, as n→∞,Zhas a Gaussian distribution
with mean µand variance σ2/n.
We may use the central limit theorem to derive an analogous result to (iii)
above for the product W=X1X2···Xnof the nindependent random variables
Xi. Provided the Xionly take values between zero and infinity, we may write
lnW=l nX1+l nX2+···+l nXn,
which is simply the sum of nnew random variables ln Xi. Thus, provided these
new variables each possess a formal mean and variance, the PDF of ln Wwill
tend to a Gaussian in the limit n→∞, and so the product Wwill be described
by a log-normal distribution (see subsection 26.9.2).
26.11 Joint distributions
As mentioned briefly in subsection 26.4.3, it is common in the physical sciences to
consider simultaneously two or more random variables that are not independent,
in general, and are thus described by joint probability density functions . We will
return to the subject of the interdependence of random variables after firstpresenting some of the general ways of characterising joint distributions. Wewill concentrate mainly on bivariate distributions, i.e. distributions of only two
random variables, though the results may be extended readily to multivariatedistributions. The subject of multivariate distributions is large and a detailed
study is beyond the scope of this book; the interested reader should therefore
consult one of the many specialised texts. However, we do discuss the multinomialand multivariate Gaussian distributions, in section 26.15.
The first thing to note when dealing with bivariate distributions is that the
distinction between discrete and continuous distributions may not be as clear asfor the single variable case; the random variables can both be discrete, or bothcontinuous, or one discrete and the other continuous. In general, for the randomvariables XandY, the joint distribution will take an infinite number of values
unless both XandYhave only a finite number of values. In this chapter we
will consider only the cases where XandYare either both discrete or both
continuous random variables.
1038
26.11 JOINT DISTRIBUTIONS
26.11.1 Discrete bivariate distributions
In direct analogy with the one-variable (univariate) case, if Xis a discrete random
variable that takes the values {xi}andYone that takes the values {yj}then the
probability function of the joint distribution is defined as
f(x, y)=braceleftBigg
Pr(X=xi,Y=yj)f o r x=xi,y=yj,
0o t h e r w i s e .
We may therefore think of f(x, y) as a set of spikes at valid points in the xy-plane,
whose heights represent the probability of obtaining X=xiandY=yj.T h e
normalisation of f(x, y) implies
summationdisplay
isummationdisplay
jf(xi,yj)=1 , (26.125)
where the sums over iandjtake all valid pairs of values. We can also define the
cumulative probability function
F(x, y)=summationdisplay
xi≤xsummationdisplay
yj≤yf(xi,yj), (26.126)
from which it follows that the probability that Xlies in the range [ a1,a2]a n d Y
lies in the range [ b1,b2]i sg i v e nb y
Pr(a1<X≤a2,b1<Y≤b2)=F(a2,b2)−F(a1,b2)−F(a2,b1)+F(a1,b1).
Finally, we define XandYto beindependent if we can write their joint distribution
in the form
f(x, y)=fX(x)fY(y), (26.127)
i.e. as the product of two univariate distributions.
26.11.2 Continuous bivariate distributions
In the case where both XandYare continuous random variables, the PDF of
the joint distribution is defined by
f(x, y)dx dy =P r ( x<X≤x+dx, y < Y≤y+dy),
(26.128)
sof(x, y)dx dy is the probability that xlies in the range [ x, x+dx]a n d ylies in
the range [ y,y+dy]. It is clear that the two-dimensional function f(x, y) must be
everywhere non-negative and that normalisation requires
integraldisplay∞
−∞integraldisplay∞
−∞f(x, y)dx dy =1.
1039
PROBABILITY
It follows further that
Pr(a1<X≤a2,b1<Y≤b2)=integraldisplayb2
b1integraldisplaya2
a1f(x, y)dx dy.
(26.129)
We can also define the cumulative probability function by
F(x, y)=P r ( X≤x, Y≤y)=integraldisplayx
−∞integraldisplayy
−∞f(u, v)du dv,
from which we see that (as for the discrete case),
Pr(a1<X≤a2,b1<Y≤b2)=F(a2,b2)−F(a1,b2)−F(a2,b1)+F(a1,b1).
Finally we note that the definition of independence (26.127) for discrete bivariate
distributions also applies to continuous bivariate distributions.IA flat table is ruled with parallel straight lines a distance Dapart, and a thin needle of
length l<D is tossed onto the table at random. What is the probability that the needle
will cross a line?
Letθbe the angle that the needle makes with the lines, and let xbe the distance from
the centre of the needle to the nearest line. Since the needle is tossed ‘at random’ ontothe table, the angle θis uniformly distributed in the interval [0 ,π], and the distance x
is uniformly distributed in the interval [0 ,D/2]. Assuming that θandxare independent,
their joint distribution is just the product of their individual distributions, and is given by
f(θ,x)=1
π1
D/2=2
πD.
The needle will cross a line if the distance xof its centre from that line is less than1
2lsinθ.
Thus the required probability is
2
πD
Zπ
0
Z1
2lsinθ
0dx dθ =2
πDl
2
Zπ
0sinθd θ=2l
πD.
This gives an experimental (but cumbersome) method of determining π.
J
26.11.3 Marginal and conditional distributions
Given a bivariate distribution f(x, y), we may only be interested in the proba-
bility function for Xirrespective of the value of Y(or vice versa). This marginal
distribution of Xis obtained by summing or integrating, as appropriate, the
joint probability distribution over all allowed values of Y. Thus, the marginal
distribution of X(for example) is given by
fX(x)=braceleftBiggsummationtext
jf(x, yj) for a discrete distribution,integraltext
f(x, y)dyfor a continuous distribution.(26.130)
It is clear that an analogous definition exists for the marginal distribution of Y.
Alternatively, one might be interested in the probability function of Xgiven
1040
26.12 PROPERTIES OF JOINT DISTRIBUTIONS
thatYtakes some specific value of Y=y0,i . e .P r ( X=x|Y=y0). This conditional
distribution of Xis given by
g(x)=f(x, y0)
fY(y0),
where fY(y) is the marginal distribution of Y. The division by fY(y0) is necessary
in order that g(x) is properly normalised.
26.12 Properties of joint distributions
The probability density function f(x, y) contains all the information on the joint
probability distribution of two random variables XandY. In a similar manner
to that presented for univariate distributions, however, it is conventional tocharacterise f(x, y) by certain of its properties, which we now discuss. Once
again, most of these properties are based on the concept of expectation values,which are defined for joint distributions in an analogous way to those for single-
variable distributions (26.46). Thus, the expectation value of any function g(X,Y)
of the random variables XandYis given by
E[g(X,Y)] =braceleftBiggsummationtext
isummationtext
jg(xi,yj)f(xi,yj) for the discrete case,integraltext∞
−∞integraltext∞
−∞g(x, y)f(x, y)dx dy for the continuous case.
26.12.1 Means
The means of XandYare defined respectively as the expectation values of the
variables XandY. Thus, the mean of Xis given by
E[X]=µX=braceleftBiggsummationtext
isummationtext
jxif(xi,yj) for the discrete case,integraltext∞
−∞integraltext∞
−∞xf(x, y)dx dy for the continuous case.(26.131)
E[Y] is obtained in a similar manner.IShow that if XandYare independent random variables then E[XY]=E[X]E[Y].
Let us consider the case where XandYare continuous random variables. Since Xand
Yare independent f(x, y)=fX(x)fY(y), so that
E[XY]=
Z∞
−∞
Z∞
−∞xyf X(x)fY(y)dx dy=
Z∞
−∞xfX(x)dx
Z∞
−∞yfY(y)dy=E[X]E[Y].
An analogous proof exists for the discrete case.
J
1041
PROBABILITY
26.12.2 Variances
The definitions of the variances of Xand Yare analogous to those for the
single-variable case (26.48), i.e. the variance of Xis given by
V[X]=σ2
X=braceleftBiggsummationtext
isummationtext
j(xi−µX)2f(xi,yj) for the discrete case,integraltext∞
−∞integraltext∞
−∞(x−µX)2f(x, y)dx dy for the continuous case.(26.132)
Equivalent definitions exist for the variance of Y.
26.12.3 Covariance and correlation
Means and variances of joint distributions provide useful information about
their marginal distributions, but we have not yet given any indication of how to
measure the relationship between the two random variables. Of course, it may
be that the two random variables are independent, but often this is not so. Forexample, if we measure the heights and weights of a sample of people we wouldnot be surprised to find a tendency for tall people to be heavier than short peopleand vice versa. We will show in this section that two functions, the covariance
and the correlation , can be defined for a bivariate distribution and that these are
useful in characterising the relationship between the two random variables.
Thecovariance of two random variables XandYis defined by
Cov[X,Y]=E[(X−µ
X)(Y−µY)], (26.133)
where µXandµYare the expectation values of XandYrespectively. Clearly
related to the covariance is the correlation of the two random variables, defined
by
Corr[ X,Y]=Cov[X,Y]
σXσY, (26.134)
where σXandσYare the standard deviations of XandYrespectively. It can be
shown that the correlation function lies between −1 and +1. If the value assumed
is negative, XandYare said to be negatively correlated , if it is positive they are
said to be positively correlated a n di fi ti sz e r ot h e ya r es a i dt ob e uncorrelated .
We will now justify the use of these terms.
One particularly useful consequence of its definition is that the covariance
of two independent variables, Xand Y, is zero. It immediately follows from
(26.134) that their correlation is also zero, and this justifies the use of the term‘uncorrelated’ for two such variables. To show this extremely important property
1042
26.12 PROPERTIES OF JOINT DISTRIBUTIONS
we first note that
Cov[X,Y]=E[(X−µX)(Y−µY)]
=E[XY−µXY−µYX+µXµY]
=E[XY]−µXE[Y]−µYE[X]+µXµY
=E[XY]−µXµY. (26.135)
Now, if XandYare independent then E[XY]=E[X]E[Y]=µXµYand so
Cov[X,Y] = 0. It is important to note that the converse of this result is not
necessarily true; two variables dependent on each other can still be uncorrelated.
In other words, it is possible (and not uncommon) for two variables XandY
to be described by a joint distribution f(x, y)t h a t cannot be factorised into a
product of the form g(x)h(y), but for which Corr[ X,Y] = 0. Indeed, from the
definition (26.133), we see that for any joint distribution f(x, y) that is symmetric
inxabout µX(or similarly in y) we have Corr[ X,Y]=0 .
We have already asserted that if the correlation of two random variables is
positive (negative) they are said to be positively (negatively) correlated. We havealso stated that the correlation lies between −1 and +1. The terminology suggests
that if the two RVs are identical (i.e. X=Y) then they are completely correlated
and that their correlation should be +1. Likewise, if X=−Ythen the functions
are completely anticorrelated and their correlation should be −1. Values of the
correlation function between these extremes show the existence of some degreeof correlation. In fact it is not necessary that X=Yfor Corr[ X,Y] = 1; it is
sufficient that Yis a linear function of X,i . e .Y=aX+b(with apositive). If a
is negative then Corr[ X,Y]=−1. To show this we first note that µ
Y=aµX+b.
Now
Y=aX+b=aX+µY−aµX⇒ Y−µY=a(X−µX),
and so using the definition of the covariance (26.133)
Cov[X,Y]=aE[(X−µX)2]=aσ2
X.
It follows from the properties of the variance (subsection 26.5.3) that σY=|a|σX
and so, using the definition (26.134) of the correlation,
Corr[ X,Y]=aσ2
X
|a|σ2
X=a
|a|,
which is the stated result.
It should be noted that, even if the possibilities of XandYbeing non-zero are
mutually exclusive, Corr[ X,Y] need not have value ±1.
1043
PROBABILITYIA biased die gives probabilities1
2p,p,p,p,p,2pof throwing 1, 2, 3, 4, 5, 6 respectively.
If the random variable Xis the number shown on the die and the random variable Yis
defined as X2, calculate the covariance and correlation of XandY.
We have already calculated in subsections 26.2.1 and 26.5.4 that p=2
13,E[X]=53
13,
E
/
X2
/
=253
13andV[X]=480
169. Using (26.135)
Cov[X,Y]=C o v [ X,X2]=E[X3]−E[X]E[X2].
Now E[X3]i sg i v e nb y
E[X3]=13×1
2p+( 23+33+43+53)p+63×2p
=1313
2p= 101 ,
and the covariance of XandYis given by
Cov[X,Y] = 101−53
13×253
13=3660
169.
The correlation is defined by Corr[ X,Y]=C o v [ X,Y]/σXσY. The standard deviation of
Ymay be calculated from the definition of the variance. Letting µY=E[X2]=253
13gives
σ2
Y=p
2
/;
12−µY
/2+p
/;
22−µY
/2+p
/;
32−µY
/2+p
/;
42−µY
/2
+p
/;
52−µY
/2+2p
/;
62−µY
/2
=187356
169p=28824
169.
We deduce that
Corr[ X,Y]=3660
169
r
169
28824
r
169
480≈0.984.
Thus the random variables XandYdisplay a strong degree of positive correlation, as we
would expect.
J
We note that the covariance of XandYoccurs in various expressions. For
example, if XandYarenotindependent then
V[X+Y]=Ebracketleftbig
(X+Y)2bracketrightbig
−(E[X+Y])2
=Ebracketleftbig
X2bracketrightbig
+2E[XY]+Ebracketleftbig
Y2bracketrightbig
−{(E[X])2+2E[X]E[Y]+(E[Y])2}
=V[X]+V[Y]+2 ( E[XY]−E[X]E[Y])
=V[X]+V[Y]+2C o v [ X,Y].
More generally, we find (for a,bandcconstant)
V[aX+bY+c]=a2V[X]+b2V[Y]+2abCov[X,Y].
(26.136)
1044
26.12 PROPERTIES OF JOINT DISTRIBUTIONS
Note that if XandYare in fact independent then Cov[ X,Y]=0a n dw er e c o v e r
the expression (26.68) in subsection 26.6.4.
We may use (26.136) to obtain an approximate expression for V[f(X,Y)]
for any arbitrary function f, even when the random variables Xand Yare
correlated. Approximating f(X,Y) by the linear terms of its Taylor expansion
about the point ( µX,µY), we have
f(X,Y)≈f(µX,µY)+parenleftbigg∂f
∂Xparenrightbigg
(X−µX)+parenleftbigg∂f
∂Yparenrightbigg
(Y−µY),
(26.137)
where the partial derivatives are evaluated at X=µXandY=µY.T a k i n gt h e
variance of both sides, and using (26.136), we find
V[f(X,Y)]≈parenleftbigg∂f
∂Xparenrightbigg2
V[X]+parenleftbigg∂f
∂Yparenrightbigg2
V[Y]+2parenleftbigg∂f
∂Xparenrightbiggparenleftbigg∂f
∂Yparenrightbigg
Cov[X,Y].
(26.138)
Clearly, if Cov[ X,Y] = 0, we recover the result (26.69) derived in subsection 26.6.4.
We note that (26.138) is exact if f(X,Y) is linear in XandY.
For several variables Xi,i=1,2,...,n, we can define the symmetric (positive
definite) covariance matrix whose elements are
Vij=C o v [ Xi,Xj], (26.139)
and the symmetric (positive definite) correlation matrix
ρij=C o r r [ Xi,Xj].
The diagonal elements of the covariance matrix are the variances of the variables,
whilst those of the correlation matrix are unity. For several variables, (26.138)generalises to
V[f(X
1,X2,...,X n)]≈summationdisplay
iparenleftbigg∂f
∂Xiparenrightbigg2
V[Xi]+summationdisplay
isummationdisplay
j/negationslash=iparenleftbigg∂f
∂Xiparenrightbiggparenleftbigg∂f
∂Xjparenrightbigg
Cov[Xi,Xj],
where the partial derivatives are evaluated at Xi=µXi.
1045
PROBABILITYIA card is drawn at random from a normal 52-card pack and its identity noted. The card
is replaced, the pack shuffled and the process repeated. Random variables W, X, Y, Z are
defined as follows:
W=2 if the drawn card is a heart; W=0otherwise.
X=4 if the drawn card is an ace, king, or queen; X=2if the card is
aj a c ko rt e n ; X=0otherwise.
Y=1 if the drawn card is red; Y=0otherwise.
Z=2 if the drawn card is black and an ace, king or queen; Z=0
otherwise.
Establish the correlation matrix for W, X, Y, Z .
The means of the variables are given by
µW=2×1
4=1
2,µ X=
/;
4×3
13
/
+
/;
2×2
13
/
=16
13,
µY=1×1
2=1
2,µ Z=2×6
52=3
13.
The variances, calculated from σ2
U=V[U]=E
/
U2
/
−(E[U])2,w h e r e U=W,X,Yor
Z,a r e
σ2
W=
/;
4×1
4
/
−
/;1
2
/2=3
4,σ2
X=
/;
16×3
13
/
+
/;
4×2
13
/
−
/;16
13
/2=472
169,
σ2
Y=
/;
1×1
2
/
−
/;1
2
/2=1
4,σ2
Z=
/;
4×6
52
/
−
/;3
13
/2=69
169.
The covariances are found by first calculating E[WX] etc. and then forming E[WX]−µWµX
etc.
E[WX]=2(4)
/;3
52
/
+2(2)
/;2
52
/
=8
13,Cov[W,X]=8
13−1
2
/;16
13
/
=0,
E[WY] = 2(1)
/;1
4
/
=1
2, Cov[W,Y]=1
2−1
2
/;1
2
/
=1
4,
E[WZ]=0 , Cov[W,Z]=0−1
2
/;3
13
/
=−3
26,
E[XY] = 4(1)
/;6
52
/
+2 ( 1 )
/;4
52
/
=8
13,Cov[X,Y]=8
13−16
13
/;1
2
/
=0,
E[XZ] = 4(2)
/;6
52
/
=12
13, Cov[X,Z]=12
13−16
13
/;3
13
/
=108
169,
E[YZ]=0 , Cov[Y,Z]=0−1
2
/;3
13
/
=−3
26.
The correlations Corr[ W,X] and Corr[ X,Y] are clearly zero; the remainder are given by
Corr[ W,Y]=1
4
/;3
4×1
4
/−1/2=0.577,
Corr[ W,Z]=−3
26
/;3
4×69
169
/−1/2=−0.209,
Corr[ X,Z]=108
169
/;472
169×69
169
/−1/2=0.598,
Corr[ Y,Z]=−3
26
/;1
4×69
169
/−1/2=−0.361.
Finally, then, we can write down the correlation matrix:
ρ=
/0B/@10 0 .58−0.21
010 0 .60
0.58 0 1 −0.36
−0.21 0 .60−0.36 1
/1CA.
1046
26.13 GENERATING FUNCTIONS FOR JOINT DISTRIBUTIONS
As would be expected, Xis uncorrelated with either WorY, colour and face-value being
two independent characteristics. Positive correlations are to be expected between Wand
Yand between XandZ; both correlations are fairly strong. Moderate anticorrelations
exist between Zand both WandY, reflecting the fact that it is impossible for WandY
to be positive if Zis positive.
J
Finally, let us suppose that the random variables Xi,i=1,2,...,n, are related
to a second set of random variables Yk=Yk(X1,X2,...,X n),k=1,2,...,m .B y
expanding each Ykas a Taylor series as in (26.137) and inserting the resulting
expressions into the definition of the covariance (26.133), we find that the elements
of the covariance matrix for the Ykvariables are given by
Cov[Yk,Yl]≈summationdisplay
isummationdisplay
jparenleftbigg∂Yk
∂Xiparenrightbiggparenleftbigg∂Yl
∂Xjparenrightbigg
Cov[Xi,Xj].
(26.140)
It is straightforward to show that this relation is exact if the Ykare linear
combinations of the Xi. Equation (26.140) can then be written in matrix form as
VY=SVXST, (26.141)
where VYand VXare the covariance matrices of the YkandXivariables re-
spectively and Sis the rectangular m×nmatrix with elements Ski=∂Yk/∂X i.
26.13 Generating functions for joint distributions
It is straightforward to generalise the discussion of generating function in section
26.7 to joint distributions. For a multivariate distribution f(X1,X2,...,X n)o f
non-negative integer random variables Xi,i=1,2,...,n, we define the probability
generating function to be
Φ(t1,t2,...,t n)=E[tX1
1tX2
2···tXnn].
As in the single-variable case, we may also define the closely related moment
generating function, which has wider applicability since it is not restricted tonon-negative integer random variables but can be used with any set of discreteor continuous random variables X
i(i=1,2,...,n). The MGF of the multivariate
distribution f(X1,X2,...,X n) is defined as
M(t1,t2,...,t n)=E[et1X1et2X2···etnXn]=E[et1X1+t2X2+···+tnXn]
(26.142)
and may be used to evaluate (joint) moments of f(X1,X2,...,X n). By performing
a derivation analogous to that presented for the single-variable case in subsection26.7.2, it can be shown that
E[X
m1
1Xm2
2···Xmn
n]=∂m1+m2+···+mnM(0,0,...,0)
∂tm1
1∂tm2
2···∂tmnn. (26.143)
1047
PROBABILITY
Finally we note that, by analogy with the single-variable case, the characteristic
function and the cumulant generating function of a multivariate distribution aredefined respectively as
C(t
1,t2,...,t n)=M(it1,i t2,...,i t n)a n d K(t1,t2,...,t n)=l n M(t1,t2,...,t n).ISuppose that the random variables Xi,i=1,2,...,n,a r ed e s c r i b e db yt h eP D F
f(x)=f(x1,x2,...,x n)=Nexp(−1
2xTAx),
where the column vector x=(x1x2··· xn)T,Ais an n×nsymmetric matrix and N
is a normalisation constant such thatZ
∞f(x)dnx≡
Z∞
−∞
Z∞
−∞···
Z∞
−∞f(x1,x2,...,x n)dx1dx2···dxn=1.
Find the MGF of f(x).
From (26.142), the MGF is given by
M(t1,t2,...,t n)=N
Z
∞exp(−1
2xTAx+tTx)dnx, (26.144)
where the column vector t=(t1t2··· tn)T. In order to evaluate this multiple integral,
we begin by noting that
xTAx−2tTx=(x−A−1t)TA(x−A−1t)−tTA−1t,
which is the matrix equivalent of ‘completing the square’. Using this expression in (26.144)
and making the substitution y=x−A−1t,w eo b t a i n
M(t1,t2,...,t n)=cexp(1
2tTA−1t), (26.145)
where the constant cis given by
c=N
Z
∞exp(−1
2yTAy)dny.
From the normalisation condition for N,w es e et h a t c= 1, as indeed it must be in order
thatM(0,0,...,0) = 1.
J
26.14 Transformation of variables in joint distributions
Suppose the random variables Xi,i=1,2,...,n, are described by the multivariate
PDF f(x1,x2...,x n). If we wish to consider random variables Yj,j=1,2,...,m ,
related to the XibyYj=Yj(X1,X2,...,X m) then we may calculate g(y1,y2,...,y m),
the PDF for the Yj, in a similar way to that in the univariate case by demanding
that
|f(x1,x2...,x n)dx1dx2···dxn|=|g(y1,y2,...,y m)dy1dy2···dym|.
From the discussion of changing the variables in multiple integrals given in
chapter 6 it follows that, in the special case where n=m,
g(y1,y2,...,y m)=f(x1,x2...,x n)|J|,
1048
26.15 IMPORTANT JOINT DISTRIBUTIONS
where
J≡∂(x1,x2...,x n)
∂(y1,y2,...,y n)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
1
∂y1...∂xn
∂y1.........
∂x1
∂yn...∂xn
∂ynvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,
is the Jacobian of the x
iwith respect to the yj.ISuppose that the random variables Xi,i=1,2,...,n, are independent and Gaussian dis-
tributed with means µiand variances σ2
irespectively. Find the PDF for the new variables
Zi=(Xi−µi)/σi,i=1,2,...,n. By considering an elemental spherical shell in Z-space,
find the PDF of the chi-squared random variable χ2
n=
Pn
i=1Z2
i.
Since the Xiare independent random variables,
f(x1,x2,...,x n)=f(x1)f(x2)···f(xn)=1
(2π)n/2σ1σ2···σnexp
/"
−nX
i=1(xi−µi)2
2σ2
i
/#
.
To derive the PDF for the variables Zi,w er e q u i r e
|f(x1,x2,...,x n)dx1dx2···dxn|=|g(z1,z2,...,z n)dz1dz2···dzn|,
and, noting that dzi=dxi/σi,w eo b t a i n
g(z1,z2,...,z n)=1
(2π)n/2exp
/
−1
2nX
i=1z2
i
/!
.
Let us now consider the random variable χ2
n=
Pn
i=1Z2
i, which we may regard as the
square of the distance from the origin in the n-dimensional Z-space. We now require that
g(z1,z2,...,z n)dz1dz2···dzn=h(χ2
n)dχ2
n.
If we consider the infinitesimal volume dV=dz1dz2···dznto be that enclosed by the
n-dimensional spherical shell of radius χnand thickness dχnthen we may write dV=
Aχn−1
ndχn, for some constant A. We thus obtain
h(χ2
n)dχ2
n∝exp(−1
2χ2
n)χn−1
ndχn∝exp(−1
2χ2
n)χn−2
ndχ2n,
where we have used the fact that dχ2
n=2χndχn. Thus we see that the PDF for χ2
nis given
by
h(χ2
n)=Bexp(−1
2χ2
n)χn−2
n,
for some constant B. This constant may be determined from the normalisation conditionZ∞
0h(χ2
n)dχ2
n=1
and is found to be B=[ 2n/2Γ(1
2n)]−1. This is the nth-order chi-squared distribution
discussed in subsection 26.9.4.
J
26.15 Important joint distributions
In this section we will examine two important multivariate distributions, the
multinomial distribution , which is an extension of the binomial distribution, and
themultivariate Gaussian distribution .
1049
PROBABILITY
26.15.1 The multinomial distribution
The binomial distribution describes the probability of obtaining x‘successes’ from
nindependent trials, where each trial has only two possible outcomes. This may
be generalised to the case where each trial has kpossible outcomes with respective
probabilities p1,p2,...,pk. If we consider the random variables Xi,i=1,2,...,n,
to be the number of outcomes of type iinntrials then we may calculate their
joint probability function
f(x1,x2,...,x k)=P r ( X1=x1,X2=x2, ..., X k=xk),
w h e r ew em u s th a v esummationtextk
i=1xi=n.I n ntrials the probability of obtaining x1
outcomes of type 1, followed by x2outcomes of type 2 etc. is given by
px1
1px2
2···pxk
k.
However, the number of distinguishable permutations of this result is
n!
x1!x2!···xk!,
and thus
f(x1,x2,...,x k)=n!
x1!x2!···xk!px1
1px2
2···pxk
k. (26.146)
This is the multinomial probability distribution .
Ifk= 2 then the multinomial distribution reduces to the familiar binomial
distribution. Although in this form the binomial distribution appears to be afunction of two random variables, it must be remembered that, in fact, sincep
2=1−p1andx2=n−x1, the distribution of X1is entirely determined by the
parameters pandn.T h a t X1has abinomial distribution is shown by remembering
that it represents the number of objects of a particular type obtained fromsampling with replacement, which led to the original definition of the binomialdistribution. In fact, any of the random variables X
ihas a binomial distribution,
i.e. the marginal distribution of each Xiis binomial with parameters nandpi.I t
immediately follows that
E[Xi]=npiand V[Xi]2=npi(1−pi). (26.147)IAt a village f ˆete patrons were invited, for a 10pentry fee, to pick without looking six
tickets from a drum containing equal large numbers of red, blue and green tickets. If fiveor more of the tickets were of the same colour a prize of 100 p was awarded. A consolation
award of 40pwas made if two tickets of each colour were picked. Was a good time had by
all?
In this case, all types of outcome (red, blue and green) have the same probabilities. Theprobability of obtaining any given combination of tickets is given by the multinomialdistribution with n=6 , k=3a n d p
i=1
3,i=1,2,3.
1050
26.15 IMPORTANT JOINT DISTRIBUTIONS
(i) The probability of picking six tickets of the same colour is given by
Pr (six of the same colour) = 3 ×6!
6!0!0!
/1
3
/6
/1
3
/0
/1
3
/0
=1
243.
The factor of 3 is present because there are three different colours.
(ii) The probability of picking five tickets of one colour and one ticket of another
colour is
Pr(five of one colour; one of another) = 3 ×2×6!
5!1!0!
/1
3
/5
/1
3
/1
/1
3
/0
=4
81.
The factors of 3 and 2 are included because there are three ways to choose the
colour of the five matching tickets, and then two ways to choose the colour of theremaining ticket.
(iii) Finally, the probability of picking two tickets of each colour is
Pr (two of each colour) =6!
2!2!2!
/1
3
/2
/1
3
/2
/1
3
/2
=10
81.
Thus the expected return to any patron was, in pence,
100
/1
243+4
81
/
+
/
40×10
81
/
=1 0.29.
A good time was had by all but the stallholder!
J
26.15.2 The multivariate Gaussian distribution
A particularly interesting multivariate distribution is provided by the generalisa-
tion of the Gaussian distribution to multiple random variables Xi,i=1,2,...,n.
If the expectation value of XiisE(Xi)=µithen the general form of the PDF is
given by
f(x1,x2,...,x n)=Nexpbracketleftbigg
−1
2summationdisplay
isummationdisplay
jaij(xi−µi)(xj−µj)bracketrightbigg
,
where aij=ajiandNis a normalisation constant that we give below. If we write
the column vectors x=(x1x2··· xn)Tandµ=(µ1µ2··· µn)T,a n d
denote the matrix with elements aijbyAthen
f(x)=f(x1,x2,...,x n)=Nexpbracketleftbig
−1
2(x−µ)TA(x−µ)bracketrightbig
,
where Ais symmetric. Using the same method as that used to derive (26.145) it
is straightforward to show that the MGF of f(x)i sg i v e nb y
M(t1,t2,...,t n)=e x pparenleftbig
µTt+1
2tTA−1tparenrightbig
,
where the column matrix t=(t1t2··· tn)T. From the MGF, we find that
E[XiXj]=∂2M(0,0,...,0)
∂ti∂tj=µiµj+(A−1)ij,
1051
PROBABILITY
and thus, using (26.135), we obtain
Cov[Xi,Xj]=E[(Xi−µi)(Xj−µj)] = ( A−1)ij.
Hence Ais equal to the inverse of the covariance matrix Vof the Xi, see (26.139).
Thus, with the correct normalisation, f(x)i sg i v e nb y
f(x)=1
(2π)n/2(det V)1/2expbracketleftbig
−1
2(x−µ)TV−1(x−µ)bracketrightbig
.
(26.148)IEvaluate the integral
I=
Z
∞exp
/
−1
2(x−µ)TV−1(x−µ)
/
dnx,
where Vis a symmetric matrix, and hence verify the normalisation in (26.148).
We begin by making the substitution y=x−µto obtain
I=
Z
∞exp(−1
2yTV−1y)dny.
Since Vis a symmetric matrix, it may be diagonalised by an orthogonal transformation to
the new set of variables y/prime=STy,w h e r e Sis the orthogonal matrix with the normalised
eigenvectors of Vas its columns (see section 8.16). In this new basis, the matrix Vbecomes
V/prime=STVS=d i a g ( λ1,λ2,...,λ n),
where the λiare the eigenvalues of V. Also, since Sis orthogonal, det S=±1, and so
dny=|detS|dny/prime=dny/prime.
Thus we can write Ias
I=
Z∞
−∞
Z∞
−∞···
Z∞
−∞exp
/
−nX
i=1y/prime
i2
2λi
/!
dy/prime
1dy/prime
2···dy/prime
n
=nY
i=1
Z∞
−∞exp
/
−y/prime
i2
2λi
/!
dy/prime
i=( 2π)n/2(λ1λ2···λn)1/2, (26.149)
where we have used the standard integral
R∞
−∞exp(−αy2)dy=(π/α)1/2(see subsection
6.4.2). From section 8.16, however, we note that the product of eigenvalues in (26.149) ise q u a lt od e t V. Thus we finally obtain
I=( 2π)
n/2(det V)1/2,
and hence the normalisation in (26.148) ensures that f(x) integrates to unity.
J
The above example illustrates some importants points concerning the multi-
variate Gaussian distribution. In particular, we note that the Y/prime
iareindependent
Gaussian variables with mean zero and variance λi. Thus, given a general set of
nGaussian variables xwith means µand covariance matrix V, one can always
perform the above transformation to obtain a new set of variables y/prime,w h i c ha r e
linear combinations of the old ones and are distributed as independent Gaussians
with zero mean and variances λi.
This result is extremely useful in proving many of the properties of the mul-
1052
26.16 EXERCISES
tivariate Gaussian. For example, let us consider the quadratic form (multiplied
by 2) appearing in the exponent of (26.148) and write it as χ2
n,i . e .
χ2
n=(x−µ)TV−1(x−µ). (26.150)
From (26.149), we see that we may also write it as
χ2
n=nsummationdisplay
i=1y/prime
i2
λi,
which is the sum of nindependent Gaussian variables with mean zero and unit
variance. Thus, as our notation implies, the quantity χ2
nis distributed as a chi-
squared variable of order n. As illustrated in exercise 26.40, if the variables Xiare
required to satisfy mlinear constraints of the formsummationtextn
i=1ciXi=0t h e n χ2
ndefined
in (26.150) is distributed as a chi-squared variable of order n−m.
26.16 Exercises
26.1 By shading Venn diagrams, determine which of the following are valid rela-
tionships between events. For those that are, prove them using de Morgan’slaws.
(a)
(¯X∪Y)=X∩¯Y.
(b)¯X∪¯Y=(X∪Y).
(c) ( X∪Y)∩Z=(X∪Z)∩Y.
(d)X∪(Y∩Z)=(X∪¯Y)∩¯Z.
(e)X∪(Y∩Z)=(X∪¯Y)∪¯Z.
26.2 Given that events X,YandZsatisfy
(X∩Y)∪(Z∩X)∪(¯X∪¯Y)=(Z∪¯Y)∪{[(¯Z∪¯X)∪(¯X∩Z)]∩Y},
prove that X⊇Yand either Y∩Z=∅orY⊇Z.
26.3 AandBeach have two unbiased four-faced dice, the four faces being numbered
1, 2, 3, 4. Without looking, Btries to guess the sum xof the numbers on the
bottom faces of A’s two dice after they have been thrown onto a table. If the
guess is correct Breceives x2euros, but if not he loses xeuros.
Determine B’s expected gain per throw of A’s dice when he adopts each of the
following strategies:
(a) he selects xat random in the range 2 ≤x≤8;
(b) he throws his own two dice and guesses xto be whatever they indicate;
(c) he takes your advice and always chooses the same value for x. Which number
would you advise?
26.4 Use the method of induction to prove e quation (26.16), the probability addition
law for the union of ngeneral events.
26.5 Two duellists, AandB, take alternate shots at each other, and the duel is over
when a shot (fatal or otherwise!) hits its target. Each shot fired by Ahas a
probability αof hitting B, and each shot fired by Bhas a probability βof hitting
A. Calculate the probabilities P1andP2, defined as follows, that Awill win such
a duel: P1,Afires the first shot; P2,Bfires the first shot.
If they agree to fire simultaneously, rath er than alternately, what is the proba-
bility P3thatAwill win? Verify that your results satisfy the intuitive inequality
P1≥P3≥P2.
1053
PROBABILITY
26.6 X1,X2,...,X nare independent identically distributed random variables drawn
from a uniform distribution on [0 ,1]. The random variables AandBare defined
by
A=m i n ( X1,X2,...,X n),B = max( X1,X2,...,X n).
For any fixed ksuch that 0 ≤k≤1
2, find the probability pnthat both
A≤k and B≥1−k.
Check your general formula by considering directly the cases (a) k=0 ,( b ) k=1
2,
(c)n=1a n d( d ) n=2 .
26.7 A tennis tournament is arranged on a straight knockout basis for 2nplayers and
for each round, except the final, opponents for those still in the competition aredrawn at random. The quality of the field is so even that in any match it isequally likely that either player will win. Two of the players have surnames thatbegin with ‘ Q’. Find the probabilities that they play each other
(a) in the final,
(b) at some stage in the tournament.
26.8 (a) Gamblers AandBeach roll a fair six-faced die, and Bwins if his score is
strictly greater than A’s. Show that the odds are 7 to 5 in A’s favour.
(b) Calculate the probabilities of scoring a total Tfrom two rolls of a fair die
forT=2,3,...,12. Gamblers CandDeach roll a fair die twice and score
respective totals T
CandTD,Dwinning if TD>T C. Realising that the odds
are not equal, Dinsists that Cshould increase her stake for each game. C
agrees to stake £1.10 per game, as compared to D’s£1.00 stake. Who will
show a profit?
26.9 An electronics assembly firm buys its microchips from three different suppliers;
half of them are bought from firm X, whilst firms YandZsupply 30% and
20% respectively. The suppliers use different quality-control procedures and thepercentages of defective chips are 2%, 4% and 4% for X,YandZrespectively.
The probabilities that a defective chip will fail two or more assembly-line testsare 40%, 60% and 80% respectively, whilst all defective chips have a 10% chanceof escaping detection. An assembler finds a chip that fails only one test. What isthe probability that it came from supplier X?
26.10 As every student of probability theory will know, Bayesylvania is awash with
natives, not all of whom can be trusted to tell the truth, and lost and apparentlysomewhat deaf travellers who ask the sam e question several times in an attempt
to get directions to the nearest village.
One such traveller finds himself at a T-junction in an area populated by the
Asciis and Bisciis in the ratio 11 to 5. As is well known, the Biscii always lie butthe Ascii tell the truth three quarters o f the time, giving independent answers to
all questions, even to immediately repeated ones.
(a) The traveller asks one particular native twice whether he should go to the
left or to the right to reach the local village. Each time he is told ‘left’. Shouldhe take this advice, and, if he does, what are his chances of reaching thevillage?
(b) The traveller then asks the same native the same question a third time and
for a third time receives the answer ‘left’. What should the traveller do now?Have his chances of finding the village been altered by asking the thirdquestion?
26.11 A boy is selected at random from amongs t the children belonging to families with
nchildren. It is known that he has at least two sisters. Show that the probability
1054
26.16 EXERCISES
that he has k−1b r o t h e r si s
(n−1)!
(2n−1−n)(k−1)!(n−k)!,
for 1≤k≤n−2 and zero for other values of k.
26.12 Villages A,B,CandDare connected by overhead telephone lines joining AB,
AC,BC,BDandCD. As a result of severe gales, there is a probability p(the
same for each link) that any particular link is broken.
(a) Show that the probability that a call can be made from AtoBis
1−2p2+p3.
(b) Show that the probability that a call can be made from DtoAis
1−2p2−2p3+5p4−2p5.
26.13 A set of 2 N+ 1 rods consists of one of each integer length 1 ,2,... ,2N,2N+1 .
Three, of lengths a,bandc, are selected, of which ais the longest. By considering
the possible values of bandc, determine the number of ways in which a non-
degenerate triangle (i.e. one of non-zero area) can be formed (i) if ais even,
and (ii) if ais odd. Combine these results appropriately to determine the total
number of non-degenerate triangles that can be formed with the 2 N+ 1 rods,
and hence show that the probability that such a triangle can be formed from arandom selection (without replacement) of three rods is
(N−1)(4N+1 )
2(4N2−1).
26.14 A certain marksman never misses his target, which consists of a disc of unit
radius with centre O. The probability that any given shot will hit the target
within a distance tofOist2for 0≤t≤1. The marksman fires nindependendent
shots at the target, and the random variable Yis the radius of the smallest circle
with centre Othat encloses all the shots. Determine the PDF for Yand hence
find the expected area of the circle.
The shot that is furthest from Ois now rejected and the corresponding circle
determined for the remaining n−1 shots. Show that its expected area is
n−1
n+1π.
26.15 The duration of a telephone call made from a public call-box is a random variable
T. The probability density function of Tis
f(t)=
/8/>/</>/:0 t<0,
1
20≤t<1,
ke−2tt≥1,
where kis a constant. To pay for the call, 20 pence has to be inserted at the
beginning, and a further 20 pence after each subsequent half-minute. Determineby how much the average cost of a call exceeds the cost of a call of averagelength charged at 40 pence per minute.
26.16 Kittens from different litters do not get on with each other and fighting breaks out
whenever two kittens from different litters are present together. A cage initiallycontains xkittens from one litter and yfrom another. To quell the fighting,
kittens are removed at random, one at a time, until peace is restored. Show, byinduction, that the expected number of kittens finally remaining is
N(x, y)=x
y+1+y
x+1.
1055
PROBABILITY
26.17 ( A more difficult question. )
If the scores in a cup football match are equal at the end of the normal
period of play, a ‘penalty shoot-out’ is held in which each side takes up to fiveshots (from the penalty spot) alternately, the shoot-out being stopped if oneside acquires an unassailable lead (i.e. has a lead greater than its opponentshave shots remaining). If the scores are still level after the shoot-out a ‘suddendeath’ competition takes place. In sudden death each side takes one shot and thecompetition is over if one side scores and the other does not; if both score, orboth fail to score, a further shot is taken by each side, and so on. Team 1, whichtakes the first penalty, has a probability p
1, which is independent of the player
involved, of scoring and a probability q1(= 1−p1) of missing; p2andq2are
defined likewise.
Define Pr( i:x, y) as the probability that team ihas scored xgoals after y
attempts, and let f(M) be the probability that the shoot-out terminates after a
totalofMshots.
(a) Prove that the probability tha t ‘sudden death’ will be needed is
f(11+) =5X
r=0(5Cr)2(p1p2)r(q1q2)5−r.
(b) Give reasoned arguments (preferably without first looking at the expressions
involved) which show that
f(M=2N)=2N−6X
r=0
/
p2Pr(1: r,N)P r(2:5−N+r,N−1)
+q2Pr(1:6−N+r,N)P r(2: r,N−1)
/
forN=3,4,5a n d
f(M=2N+1 )=2N−5X
r=0
/
p1Pr(1:5−N+r,N)P r(2: r,N)
+q1Pr(1: r,N)P r( 2:5−N+r,N)
/
forN=3,4.
(c) Give an explicit expression for Pr( i:x, y) and hence show that if the teams
are so well matched that p1=p2=1/2t h e n
f(2N)=2N−6X
r=0
/1
22N
/N!(N−1)!6
r!(N−r)!(6−N+r)!(2N−6−r)!,
f(2N+1 )=2N−5X
r=0
/1
22N
/(N!)2
r!(N−r)!(5−N+r)!(2N−5−r)!.
(d) Evaluate these expressions to show that, expressing f(M) in units of 2−8,w e
have
M 6 7 8 9 10 11+
f(M) 8 24 42 56 63 63
Give a simple explanation of why f(10) = f(11+).
26.18 A particle is confined to the one-dimensional space 0 ≤x≤aand classically
it can be in any small interval dxwith equal probability. However, quantum
mechanics gives the result that the probability distribution is proportional tosin
2(nπx/a ), where nis an integer. Find the variance in the particle’s position
in both the classical and quantum mechanical pictures and show that, althoughthey differ, the latter tends to the former in the limit of large n, in agreement
with the correspondence principle of physics.
1056
26.16 EXERCISES
26.19 A continuous random variable Xhas a probability density function f(x); the
corresponding cumulative probability function is F(x). Show that the random
variable Y=F(X) is uniformly distributed between 0 and 1.
26.20 For a non-negative integer random variable X, in addition to the probability
generating function Φ X(t) defined in equation (26.71) it is possible to define the
probability generating function
ΨX(t)=∞X
n=0gntn,
where gnis the probability that X>n .
(a) Prove that Φ Xand Ψ Xare related by
ΨX(t)=1−ΦX(t)
1−t.
(b) Show that E[X]i sg i v e nb yΨ X(1) and that the variance of Xcan be
expressed as 2Ψ/prime
X(1) + Ψ X(1)−[ΨX(1)]2.
(c) For a particular random variable X, the probability that X>n is equal to
αn+1with 0 <α< 1. Use the results in (b) to show that V[X]=α(1−α)−2.
26.21 (a) In two sets of binomial trials Tandtthe probabilities that a trial has a
successful outcome are Pandprespectively, with corresponding probabilites
of failure of Q=1−Pandq=1−p. One ‘game’ consists of a trial T
followed, if Tis successful, by a trial tand then a further trial T.T h et w o
trials continue to alternate until one of the Ttrials fails, at which point the
game ends. The score Sfor the game is the total number of successes in the
t-trials. Find the PGF for Sand use it to show that
E[S]=Pp
Q,V [S]=Pp(1−Pq)
Q2.
(b) Two normal unbiased six-faced dice AandBare rolled alternately starting
with A;i fAshows a 6 the experiment ends. If Bshows an odd number no
points are scored, if it shows a 2 or a 4 then one point is scored, whilst ifit records a 6 then two points are awarded. Find the average and standarddeviation of the score for the experiment and show that the latter is thegreater.
26.22 Use the formula obtained in subsection 26.8.2 for the moment generating function
of the negative binomial distribution to determine the CGF K
n(t) for the number
of trials needed to record nsuccesses. Evaluate the first four cumulants and use
them to confirm the stated results for the mean and variance and to show thatthe distribution has skewness and kurtosis given respectively by
2−p
√n(1−p)and 3 +6−6p+p2
√n(1−p).
26.23 A point Pis chosen at random on the circle x2+y2= 1. The random variable
Xdenotes the distance of Pfrom (1 ,0). Find the mean and variance of Xand
the probability that Xis greater than its mean.
26.24 As assistant to a celebrated and imperious newspaper proprietor, you are given
the job of running a lottery in which each of his five million readers will havean equal independent chance pof winning a million pounds ; you have the job of
choosing p. However, if nobody wins it will be bad for publicity whilst if more
than two readers do so, the prize cost will more than offset the profit from extracirculation – in either case you will be sacked! Show that, however you choosep, there is more than a 40% chance you will soon be clearing your desk.
1057
PROBABILITY
26.25 The number of errors needing correction on each page of a set of proofs follows
a Poisson distribution of mean µ. The cost of the first correction on any page is
αand that of each subsequent correction on the same page is β. Prove that the
average cost of correcting a page is
α+β(µ−1)−(α−β)e−µ.
26.26 In the game of Blackball, at each turn Muggins draws a ball at random from a
bag containing five white balls, three red balls and two black balls; after being
recorded, the ball is replaced in the bag. A white ball earns him $1 whilst a redball gets him $2; in either case he also has the option of leaving with his currentwinnings or of taking a further turn on the same basis. If he draws a black ballthe game ends and he loses all he may have gained previously. Find an expressionfor Muggins’ expected return if he adopts the strategy to drawing up to nballs
if he has not been eliminated by then.
Show that, as the entry fee to play is $3, Muggins should be dissuaded from
playing Blackball, but if that cannot be done what value of nwould you advise
him to adopt?
26.27 Show that for large rthe value at the maximum of the PDF for the gamma
distribution of order rwith parameter λis approximately λ/√
2π(r−1).
26.28 A husband and wife decide that their family will be complete when it includes
two boys and two girls – but that this would then be enough! The probabilitythat a new baby will be a girl is p. Ignoring the possibility of identical twins,
show that the expected size of their family is
2
/1
pq−1−pq
/
,
where q=1−p.
26.29 The probability distribution for the number of eggs in a clutch is Po( λ), and the
probability that each egg will hatch is p(independently of the size of the clutch).
Show by direct calculation that the probability distribution for the number ofchicks that hatch is Po( λp) and so justify the assumptions made in the worked
example at the end of subsection 26.7.1.
26.30 A shopper buys 36 items at random in a supermarket where, because of the sales
tax imposed, the final digit (the number of pence) in the price is uniformly and
randomly distributed from 0 to 9. Instead of adding up the bill exactly she rounds
each item to the nearest 10 pence, rounding up or down with equal probabilityif the price ends in a ‘5’. Should she suspect a mistake if the cashier asks her for23 pence more than she estimated?
26.31 Under EU legislation on harmonisation, all kippers are to weigh 0.2000 kg and
vendors who sell underweight kippers must be fined by their government. Theweight of a kipper is normally distributed with a mean of 0.2000 kg and astandard deviation of 0.0100 kg. They are packed in cartons of 100 and largequantities of them are sold.
Every day a carton is to be selected at random from each vendor and tested
according to one of the following schemes, which have been approved for thepurpose.
(a) The entire carton is weighed and the vendor is fined 2500 euros if the average
weight of a kipper is less than 0.1975 kg.
(b) Twenty-five kippers are selected at random from the carton; the vendor is
fined 100 euros if the average weight of a kipper is less than 0.1980 kg.
(c) Kippers are removed one at a time, at random, until one has been found
that weighs morethan 0.2000 kg; the vendor is fined n(n−1) euros, where n
is the number of kippers removed.
1058
26.16 EXERCISES
Which scheme should the Chancellor of the Exchequer be urging his government
to adopt?
26.32 In a certain parliament the government consists of 75 New Socialites and the
opposition consists of 25 Preservatives. Preservatives never change their mind, al-ways voting against government policy without a second thought; New Socialitesvote randomly, but with probability pthat they will vote for their party leader’s
policies.
Following a decision by the New Socialites’ leader to drop certain manifesto
commitments, Nof his party decide to vote consistently with the opposition. The
leader’s advisors reluctantly admit that an election must be called if Nis such
that, at any vote on government policy, the chance of a simple majority in favour
would be less than 80%. Given that p=0.8, estimate the lowest value of Nthat
would precipitate an election.
26.33 A practical-class demonstrator sends his 12 students to the storeroom to collect
apparatus for an experiment, but forgets to tell each which type of componentto bring. There are three types, A,BandC,h e l di nt h es t o r e s( i nl a r g en u m b e r s )
in the proportions 20%, 30% and 50% respectively, and each student picks acomponent at random. In order to set up one experiment, one unit each of Aand
Band two units of Care needed. Find an expression for the probability Pr( N)
that at least Nexperiments can be set up.
(a) Evaluate Pr(3).
(b) Show that Pr(2) can be written in the form
Pr(2) = (0 .5)
126X
i=212Ci(0.4)i8−iX
j=212−iCj(0.6)j.
(c) By considering the conditions under which no experiments can be set up,
show that Pr(1) = 0 .9145.
26.34 The random variables XandYtake integer values ≥1 such that 2 x+y≤2a,
where ais an integer greater than 1. The joint probability within this region is
given by
Pr(X=x, Y=y)=c(2x+y),
where cis a constant, and it is zero elsewhere.
Show that the marginal probability Pr( X=x)i s
Pr(X=x)=6(a−x)(2x+2a+1 )
a(a−1)(8a+5 ),
and obtain expressions for Pr( Y=y), (a) when yis even and (b) when yis odd.
Show further that
E[Y]=6a2+4a+1
8a+5.
(You will need the results about series involving the natural numbers given in
subsection 4.2.5.)
26.35 The continuous random variables XandYhave a joint PDF proportional to
xy(x−y)2with 0≤x≤1a n d0≤y≤1. Find the marginal distributions
forXand Yand show that they are negatively correlated with correlation
coefficient−2
3.
1059
PROBABILITY
26.36 A discrete random variable Xtakes integer values n=0,1,... ,N with probabil-
itiespn. A second random variable Yis defined as Y=(X−µ)2,w h e r e µis the
expectation value of X. Prove that the covariance of XandYis given by
Cov[X,Y]=NX
n=0n3pn−3µNX
n=0n2pn+2µ3.
Now suppose that Xtakes all its possible values with equal probability and hence
demonstrate that two random variables can be uncorrelated even though one isdefined in terms of the other.
26.37 Two continuous random variables XandYhave a joint probability distribution
f(x, y)=A(x
2+y2),
where Ais a constant and 0 ≤x≤a,0≤y≤a. Show that XandYare negatively
correlated with correlation coefficient −15/73. By sketching a rough contour
map of f(x, y) and marking off the regions of positive and negative correlation,
convince yourself that this (perhaps counter-intuitive) result is plausible.
26.38 A continuous random variable Xis uniformly distributed over the interval [ −c, c].
As a m p l eo f2 n+ 1 values of Xis selected at random and the random variable
Zis defined as the median of that sample. Show that Zis distributed over [ −c, c]
with probability density function
fn(z)=(2n+1 ) !
(n!)2(2c)2n+1(c2−z2)n.
Find the variance of Z.
26.39 Show that, as the number of trials nbecomes large but npi=λi,i=1,2,...,k−1,
remains finite, the multinomial p robability distribution (26.146),
Mn(x1,x2,...,x k)=n!
x1!x2!···xk!px1
1px2
2···pxk
k,
can be approximated by a multiple Poisson distribution (with k−1 factors)
M/prime
n(x1,x2,...,x k−1)=k−1Y
i=1e−λiλxi
i
xi!.
(Write
Pk−1
ipi=δand express all terms involving subscript kin terms of nand
δ, either exactly or approximately. You will need to use n!≈n/epsilon1[(n−/epsilon1)!] and
(1−a/n)n≈e−afor large n.)
(a) Verify that the terms of M/prime
nwhen summed over all values of x1,x2,...,x k−1
a d du pt ou n i t y .
(b) If k=7a n d λi=9f o ra l l i=1,2,...,6, estimate, using the appropriate
Gaussian approximation, the chance that at least three of x1,x2,...,x 6will
be 15 or greater.
26.40 The variables Xi,i=1,2,...,n, are distributed as a multivariate Gaussian, with
means µiand a covariance matrix V.I ft h e Xiare required to satisfy the linear
constraint
Pn
i=1ciXi=0 ,w h e r et h e ciare constants (and not all equal to zero),
show that the variable
χ2
n=(x−µ)TV−1(x−µ)
follows a chi-squared distribution of order n−1.
1060
26.17 HINTS AND ANSWERS
26.17 Hints and answers
26.1 (a) Yes, (b) no, (c) no, (d) no, (e) yes.
26.2 Reduce the equality to X∩(Y∪Z)=Y.
26.3 Show that if px/16 is the probability that the total will be xthen the corrsponding
gain is [ px(x2+x)−16x]/16. (a) A loss of 2.5 euros; (b) a gain of27
64euros; (c)
a gain of 2.5 euros, provided he takes your advice and guesses ‘5’ each time.
26.4 Let Bbe the union of events A1,A2,... ,A nand apply (26.9) with AasAn+1.
Evaluate Pr( B∩An+1) by applying the assumed result to the set of nevents Ci=
Ai∩An+1fori=1,2,...,n and noting that Ci∩Cj∩···∩Cm=Ai∩Aj∩···∩Am∩An+1.
26.5 P1=α(α+β−αβ)−1;P2=α(1−β)(α+β−αβ)−1;P3=α(α+β)−1.
26.6 Find simple expressions for the separate probabilities that A≥kandB≤1−k,
and also for the two conditions at the same time. Then applying (26.11) andidentities typified by Pr( C)=P r ( CandD)+Pr( Cand¯D), show that p
n=1−2(1−
k)n+( 1−2k)n. (a) 0, (b) 1 −2−(n−1),( c )0 ,( d )2 k2.
26.7 If pris the probability that before the rth round both players are still in the
tournament (and have not met each other), show that
pr+1=1
42n+1−r−2
2n+1−r−1prand hence that pr=
/1
2
/r−12n+1−r−1
2n−1.
(a) The probability that they meet in the final is pn=2−(n−1)(2n−1)−1.
(b) The probability that they meet at some stage in the tournament is given by
the sum
Pn
r=1pr(2n+1−r−1)−1=2−(n−1).
26.8 (b) Pr( TD>T C)=0 .5{1−[146/(36)2]}=0.4437; C’s expected return is equal to
£2.10(1−0.4437)≈£1.17 for a £1.10 stake.
26.9 The relative probabilities are X:Y:Z= 50 : 36 : 8 (in units of 10−4);25
47.
26.10 (a) Show that the probability that an Ascii gives the same answer twice in
succession to the same question is 5 /8 and that if he gives the same answer
twice the probability that he is telling the truth is 9 /10. Conclude that the
probability that the native questioned is an Ascii is 55 /95 and that the
probability that the traveller is being correctly directed is 99 /190. As this is
more than1
2, he should go left.
(b) For the same answer given three times the corresponding fractions are 28 /64,
27/28 and 308 /628. The chance that the traveller is being told the truth has
d r o p p e dt o2 9 7 /628, and, as this is less than one half, he should go ‘ right’
with a 331 /628 chance of success. This is a (very) slight improvement on his
previous situation.
26.11 Take Ajas the event that a family consists of jboys and n−jgirls, and Bas
the event that the boy has at least two sisters. Apply Bayes’ theorem.
26.12 If q=1−p, the probability is q3+3pq2+p2q, the separate terms corresponding
to zero, one and a particular set of double breaks; (b) similarly, the probabilityisq
5+5q4p+( 1 0−2)p2q3+2p3q2.
26.13 (i) For aeven, the number of ways is 1 + 3 + 5 + ···+(a−3), and (ii) for aodd
i ti s2+4+6+ ···+(a−3). Combine the results for a=2manda=2m+1 ,
with mrunning from 2 to N, to show that the total number of non-degenerate
triangles is given by N(4N+1 ) ( N−1)/6. The number of possible selections of a
set of three rods is (2 N+ 1)(2 N)(2N−1)/6.
26.14 The CPF for Yisy2nand the PDF is the derivative of this, namely 2 ny2n−1.
This leads to an expected area equal to nπ/(n+ 1). The same PDF gives the
distribution of the rejected shot and, for a given y, the remaining n−1 shots,
all lying within yofO, have a CPF of ( z2/y2)n−1. Show, from the corresponding
PDF, that the expected area is then ( n−1)πy2/nand that when this is averaged
over ythe stated result is obtained.
1061
PROBABILITY
26.15 Show that k=e2and that the average duration of a call is 1 minute. Let pn
be the probability that the call ends during the interval 0 .5(n−1)≤t<0.5n
and cn=2 0 nbe the corresponding cost. Prove that p1=p2=1
4and that
pn=1
2e2(e−1)e−nforn≥3. It follows that the average cost is
E[C]=30
2+2 0e2(e−1)
2∞X
n=3ne−n.
The arithmetico-geome tric series has sum (3 e−1−2e−2)/(e−1)2and the total
charge is 5( e+1 )/(e−1) = 10 .82 pence more than the 40 pence a uniform rate
would cost.
26.16 Establish that N(x, y+1 )=[ ( y+1 )N(x, y)+xN(x−1,y+1 ) ] /(x+y+1 ) .
26.17 (a) The scores must be equal, at reach, after five attempts each.
(b)Mcan only be even if team 2 gets too far ahead (or drops too far behind)
to be caught (or catch up), with conditional probability p2(orq2). Conversely M
can only be odd as a result of a final action by team 1.(c) Pr( i:x, y)=
yCxpx
iqy−x
i.
(d) if the match is still alive at the tenth kick, team 2 is just as likely to lose it asto take it into sudden death.
26.18 a
2/12;a2/12−a2/(2π2n2).
26.19 Show that dY /dX =fand use g(y)=f(x)|dx/dy|.
26.20 (a) Note that gn−gn−1=−fnand that g0=1−f0.
(b) Show that Φ/prime
X(1) = Ψ X(1) and relate Φ/prime/primeX(1) to Ψ/primeX(1).
(c) Ψ X(t)=α/(1−αt).
26.21 (a) Use result (26.84) to show that the PGF for SisQ/(1−Pq−Ppt). Then use
equations (26.74) and (26.76).
(b) The PGF for the score is 6 /(21−10t−5t2) and the average score is 10 /3.
The variance is 145 /9 and the standard deviation is 4 .01.
26.22 Kn(t)=nlnp+nt+n
P∞
r=1r−1(1−p)retr. This gives the first four cumulants as
n/p,n(1−p)/p2,n(1−p)(2−p)/p3andn(1−p)(6−6p+p2)/p4.
26.23 Mean = 4 /π.V a r i a n c e=2 −(16/π2). Probability that Xexceeds its mean
=1−(2/π)sin−1(2/π)=0 .561.
26.24 Write x=5×106p. Show that ( x+1
2x2)e−xhas a maximum value of 0 .587
whatever the value of x, and hence of p.
26.25 Consider 0, 1 and ≥2 errors on a page separately.
26.26 Show that the expected return is/4
5
/nnX
r=0nCr
/3
8
/r
/5
8
/n−r
=11
8n
/4
5
/n
.
This is maximal when n=( l n 5 /4)−1=4.48;n=4a n d n= 5 both give an
expected return of $2.2528, i.e. less than the entry fee, but are the best that canbe advised.
26.27 Show that the maximum occurs at x=(r−1)/λand then use Stirling’s approxi-
mation to find the maximum value.
26.28 Show that the probability that the ‘trials’ end with the nth child, ( n≥4), is given
by (
n−1C1pqn−2)p+(n−1C1qpn−2)q. The expectation value for nis then given by
the sum
P∞
n=4n(n−1)(p2qn−2+q2pn−2). By twice differentiating the result for the
sum of a geometric series, prove that
P∞
n=2n(n−1)rn−2=2/(1−r)3.U s et h i s
result, after explicitly removing the first two terms, to show that E[n]i sa sg i v e n .
26.29 Pr( kchicks hatching) =
P∞
n=kPo(n, λ)B i n ( n, p).
26.30 Show that the variance of the distribution that has probabilities of 1 /20 for i=−5
andi= 5, and probabilities of 1 /10 for i=−4,−3,...,4i s1 7 /2. Conclude that
1062
26.17 HINTS AND ANSWERS
23 pence is only 1 .3×the standard deviation expected for the total bill and that
a bigger discrepancy would occur about 20% of the time.
26.31 There is not much to choose between the schemes. In (a) the critical value of
the standard variable is −2.5 and the average fine would be 15.5 euros. For
(b) the corresponding figures are −1.0 and 15.9 euros. Scheme (c) is governed
by a geometric distribution with p=q=1
2, and leads to an expected fine
of
P∞
n=1n(n−1)(1
2)n. The sum can be evaluated by differentiating the resultP∞
n=1pn=p/(1−p) with respect to p, and gives the expected fine as 16 euros.
26.32 By making a Gaussian approximation to the binomial distribution, establish that
Nmust be such that
25−75q−Np=0.841
p
(75−N)pq.
With p=0.8a n d q=0.2, this has solution N=9.1.
26.33 (a) [12!(0 .5)6(0.3)3(0.2)3]/(6!3!3!) = 0 .063.
26.34 Show that Pr( X=x)=c(a−x)(2x+2a+ 1) and use the fact that
Pa−1
x=1Pr(X=
x) = 1 to prove that c=6/[a(a−1)(8a+5)]. When evaluating Pr( Y=y)c o n s i d e r
carefully the value of the upper limit in the summation over x.
(a) Pr( Y=y)=3(2a−y)(2a+y+2 )
2a(a−1)(8a+5 ),
(b) Pr( Y=y)=3(2a−y−1)(2a+y+1 )
2a(a−1)(8a+5 ).
Express the expectation value as a summation over mfrom 1 to a−1, combining
the terms involving y=2m−1a n d y=2m.
26.35 You will need to establish the normalisation constant for the distribution (36),
the common mean value (3 /5), and the common standard deviation (3 /10). The
marginal distributions are f(x)=3 x(6x2−8x+ 3) and the same function of y.
The covariance has the value −3/50, yielding a correlation of −2/3.
26.36 E[XY]=
PN
n=0n3pn−2µ
PN
n=0n2pn+µ3.
Setpn=1/(N+1 )f o ra l l nand use the results for series involving the natural
numbers given in subsection 4.2.5 to show that Cov[ X,Y]=0 .
26.37 A=3/(24a4);µX=µY=5a/8;σ2
X=σ2
Y=7 3 a2/960; E[XY]=3 a2/8;
Cov[X,Y]=−a2/64.
26.38 This is the multinomial distribution for nRVs in each of the intervals [ −c, z],
[z+dz, c] and one RV in the interval [ z,z+dz]. The corresponding basic
probabilities are
[(c±z)/(2c)]nanddz/(2c). Use the fact that fnandfn+1are normalised to deduce
the value of
R
z2fn(z)dz. The variance is c2/(2n+3 ) .
26.39 (b) With the continuity correction Pr( xi≥15) = 0 .0334. The probability that at
least three are 15 or greater is 7 .5×10−4.
26.40 Perform successive transformations of variables y/prime=ST(x−µ)a n d zi=y/prime
i/√λi,
where the columns of Sare eigenvectors of Vand the λiare eigenvalues of V.
Then bothχ
2
n=
Pn
i=1z2
iandziare independent Gaussian variables with mean zero and
unit variance, which are required to satisfy the linear constraint
Pn
i=1c/prime
izi=0f o r
some constants c/prime
i. Now require that
f(z1,z2,...,z n)dz1dz2···dzn=h(χ2
n)dχ2
n,
where dz1dz2···dznis the infinitesimal volume enclosed by the intersection of
then-dimensional spherical shell of radius χ2
nand thickness dχ2
nwith the ( n−1)-
dimensional hyperplane
Pn
i=1c/prime
izi=0 .
1063
27
Statistics
In this chapter, we turn to the study of statistics, which is concerned with
the analysis of experimental data. In a book of this nature we cannot hopeto do justice to such a large subject; indeed, many would argue that statisticsbelongs to the realm of experimental science rather than in a mathematicstextbook. Nevertheless, physical scientists and engineers are regularly called uponto perform a statistical analysis of their data and to present their results in astatistical context. Therefore, we will concentrate on this aspect of a much more
extensive subject. †
27.1 Experiments, samples and populations
We may regard the product of any experiment as a set of Nmeasurements of some
quantity xor set of quantities x ,y,...,z . This set of measurements constitutes the
data. Each measurement (or data item ) consists accordingly of a single number x
i
or a set of numbers ( xi,yi,...,,z i), where i=1,...,,N . For the moment, we will
assume that each data item is a single number, although our discussion can beextended to the more general case.
As a result of inaccuracies in the measurement process, or because of intrinsic
variability in the quantity xbeing measured, one would expect the Nmeasured
values x
1,x2,...,x Nto be different each time the experiment is performed. We may
therefore consider the xias a set of Nrandom variables. In the most general case,
†There are, in fact, two separate schools of thought concerning statistics: the frequentist approach
and the Bayesian approach. Indeed, which of these approaches is the more fundamental is still a
matter of heated debate. Here we shall concentrate primarily on the more traditional frequentistapproach (despite the preference of some of the authors for the Bayesian viewpoint!). For a fuller
discussion of the frequentist approach one could refer to, for example, Stuart & Ord, Kendall’s
Advanced Theory of Statistics Vol. I (Edward Arnold) or Kenney & Keeping, Mathematics of
Statistics (Van Nostrand). For a discussion of the Bayesian approach one might consult, for
example, Sivia, Data Analysis: A Bayesian Tutorial (OUP).
1064
27.2 SAMPLE STATISTICS
these random variables will be described by some N-dimensional joint probability
density function P(x1,x2,...,x N).†In other words, an experiment consisting of N
measurements is considered as a single random sample from the joint distribution
(orpopulation )P(x), where xdenotes a point in the N-dimensional data space
having coordinates ( x1,x2,...,x N).
The situation is simplified considerably if the sample values xiareindependent .
In this case, the N-dimensional joint distribution P(x) factorises into the product
ofNone-dimensional distributions,
P(x)=P(x1)P(x2)···P(xN). (27.1)
In the general case, each of the one-dimensional distributions P(xi)m a yb e
different. A typical example of this occurs when Nindependent measurements
are made of some quantity xbut the accuracy of the measuring procedure varies
between measurements.
It is often the case, however, that each sample value xiis drawn independently
from the samepopulation. In this case, P(x) is of the form (27.1), but, in addition,
P(xi) has the same form for each value of i. The measurements x1,x2,...,x N
a r et h e ns a i dt of o r ma random sample of size Nfrom the one-dimensional
population P(x). This is the most common situation met in practice and, unless
stated otherwise, we will assume from now on that this is the case.
27.2 Sample statistics
Suppose we have a set of Nmeasurements x1,x2,...,x N. Any function of these
measurements (that contains no unknown parameters) is called a sample statistic ,
or often simply a statistic . Sample statistics provide a means of characterising the
data. Although the resulting characteris ation is inevitably incomplete, it is useful
to be able to describe a set of data in terms of a few pertinent numbers. We nowdiscuss the most commonly used sample statistics.
27.2.1 Averages
The simplest number used to characterise a sample is the mean,w h i c hf o r N
values x
i,i=1,2,...,N , is defined by
¯x=1
NNsummationdisplay
i=1xi. (27.2)
†In this chapter, we will adopt the common convention that P(x) denotes the particular probability
density function that applies to its argument, x. This obviates the need to use a different letter
for the PDF of each new variable. For example, if XandYare random variables with different
PDFs, then properly one should denote these distributions by f(x)a n d g(y), say. In our shorthand
notation, these PDFs are denoted by P(x)a n d P(y), where it is understood that the functional
form of the PDF may be different in each case.
1065
STATISTICS
188.7 204.7 193.2 169.0
168.1 189.8 166.3 200.0
Table 27.1 Experimental data giving eight measurements of the round trip
time in milliseconds for a computer ‘packet’ to travel from Cambridge UK to
Cambridge MA.
In words, the sample mean is the sum of the sample values divided by the number
of values in the sample.ITable 27.1 gives eight values for the round trip time in milliseconds for a computer ‘packet’
to travel from Cambridge UK to Cambridge MA. Find the sample mean.
Using (27.2) the sample mean in milliseconds is given by
¯x=1
8(188.7 + 204 .7 + 193 .2 + 169 .0 + 168 .1 + 189 .8 + 166 .3 + 200 .0)
=1479.8
8= 184 .975.
Since the sample values in table 27.1 are quoted to an accuracy of one decimal place, it is
usual to quote the mean to the same accuracy, i.e. as ¯x= 185 .0.
J
Strictly speaking the mean given by (27.2) is the arithmetic mean and this is by
far the most common definition used for a mean. Other definitions of the meanare possible, though less common, and include
(i) the geometric mean ,
¯x
g=parenleftBiggNproductdisplay
i=1xiparenrightBigg1/N
, (27.3)
(ii) the harmonic mean ,
¯xh=NsummationtextN
i=11/xi, (27.4)
(iii) the root mean square ,
¯xrms=parenleftBiggsummationtextN
i=1x2
i
NparenrightBigg1/2
. (27.5)
It should be noted that, ¯x,¯xhand¯xrmswould remain well defined even if some
sample values were negative, but the value of ¯xgcould then become complex.
The geometric mean should not be used in such cases.
1066
27.2 SAMPLE STATISTICSICalculate ¯xg,¯xhand¯xrmsfor the sample given in table 27.1.
The geometric mean is given by (27.3) to be
¯xg= (188 .7×204.7×···×200.0)1/8= 184 .4.
The harmonic mean is given by (27.4) to be
¯xh=8
(1/188.7) + (1 /204.7) +···+( 1/200.0)= 183 .9.
Finally, the root mean square is given by (27.5) to be
¯xrms=
/1
8(188.72+ 204 .72+···+ 200 .02)
/1/2= 185 .5.
J
Two other measures of the ‘average’ of a sample are its modeandmedian .T h e
mode is simply the most commonly occurring value in the sample. A sample maypossess several modes, however, and it can thus be misleading in such cases touse the mode as a measure of the average of the sample. The median of a sampleis the halfway point when the sample values x
i(i=1,2,...,N ) are arranged in
ascending (or descending) order. Clearly, this depends on whether the size of
the sample, N, is odd or even. If Nis odd then the median is simply equal to
x(N+1)/2,w h e r e a si f Nis even the median of the sample is usually taken to be
1
2(xN/2+x(N/2)+1).IFind the mode and median of the sample given in table 27.1.
From the table we see that each sample value occurs exactly once, and so any value may
be called the mode of the sample.
To find the sample median, we first arrange the sample values in ascending order and
obtain
166.3, 168.1, 169.0, 188.7, 189.8, 193.2, 200.0, 204.7 .
Since the number of sample values N= 8, which is even, the median of the sample is
1
2(x4+x5)=1
2(188.7 + 189 .8) = 189 .25.
J
27.2.2 Variance and standard deviation
The variance and standard deviation both give a measure of the spread of values
in a sample about the sample mean ¯x.T h e sample variance is defined by
s2=1
NNsummationdisplay
i=1(xi−¯x)2, (27.6)
and the sample standard deviation is the positive square root of the sample
variance, i.e.
s=radicaltpradicalvertexradicalvertexradicalbt1
NNsummationdisplay
i=1(xi−¯x)2. (27.7)
1067
STATISTICSIFind the sample variance and sample standard deviation of the data given in table 27.1.
We have already found that the sample mean is 185.0 to one decimal place. However,
when the mean is to be used in the subsequent calculation of the sample variance it isbetter to use the most accurate value available. In this case the exact value is 184.975, andso using (27.6),
s
2=1
8
/
(188.7−184.975)2+···+ (200 .0−184.975)2
/
=1608.36
8= 201 .0,
where once again we have quoted the result to one decimal place. The sample standard
deviation is then given by s=√
201.0=1 4 .2. As it happens, in this case the difference
between the true mean and the rounded value is very small compared to the variation
of the individual readings about the mean and using the rounded value makes negligibledifference; however, this would not be so if the difference were comparable to the samplestandard deviation.J
Using the definition (27.7), it is clear that in order to calculate the standard
deviation of a sample we must first calculate the sample mean. This requirementcan be avoided, however, by using an alternative form for s
2. From (27.6), we see
that
s2=1
NNsummationdisplay
i=1(xi−¯x)2
=1
NNsummationdisplay
i=1x2
i−1
NNsummationdisplay
i=12xi¯x+1
NNsummationdisplay
i=1¯x2
=x2−2¯x2+¯x2=x2−¯x2
We may therefore write the sample variance s2as
s2=x2−¯x2=1
NNsummationdisplay
i=1x2
i−parenleftBigg
1
NNsummationdisplay
i=1xiparenrightBigg2
, (27.8)
from which the sample standard deviation is found by taking the positive square
root. Thus, by evaluating the quantitiessummationtextN
i=1xiandsummationtextN
i=1x2
ifor our sample, we
can calculate the sample mean and sample standard deviation at the same time.ICalculate
PN
i=1xiand
PN
i=1x2
ifor the data given in table 27.1 and hence find the mean
and standard deviation of the sample.
From table 27.1, we obtain
NX
i=1xi= 188 .7 + 204 .7+···+ 200 .0 = 1479 .8,
NX
i=1x2
i= (188 .7)2+ (204 .7)2+···+ (200 .0)2= 275334 .36.
1068
27.2 SAMPLE STATISTICS
Since N= 8, we find as before (quoting the final results to one decimal place)
¯x=1479.8
8= 185 .0,s =
s
275334 .36
8−
/1479.8
8
/2
=1 4.2.
J
27.2.3 Moments and central moments
By analogy with our discussion of probability distributions in section 26.5, the
sample mean and variance may also be described respectively as the first momentand second central moment of the sample. In general, for a sample x
i,i=
1,2,...,N , we define the rth moment mrandrth central moment nras
mr=1
NNsummationdisplay
i=1xr
i, (27.9)
nr=1
NNsummationdisplay
i=1(xi−m1)r. (27.10)
Thus the sample mean ¯xand variance s2m a ya l s ob ew r i t t e na s m1andn2
respectively. As is common practice, we have introduced a notation in which a
sample statistic is denoted by the Roman letter corresponding to whichever Greekletter is used to describe the corresponding population statistic. Thus, we use m
r
andnrto denote the moment and central moment of a sample, since in section
26.5, we denoted the rth moment and central moment of a population by µrand
νrrespectively.
This notation is particularly useful, since the rth central moment of a sample,
mr, may be expressed in terms of the rth- and lower-order sample moments nrin
a way exactly analogous to that derived in subsection 26.5.5 for the correspondingpopulation statistics. For example, as discussed in the previous section, the samplevariance is given by s
2=x2−¯x2but this may also be written as n2=m2−m2
1,
which is to be compared with the corresponding relation ν2=µ2−µ2
1derived
in subsection 26.5.3 for population statistics. This correspondence also holds for
higher-order central moments of the sample. For example,
n3=1
NNsummationdisplay
i=1(xi−m1)3
=1
NNsummationdisplay
i=1(x3
i−3m1x2
i+3m2
1xi−m3
1)
=m3−3m1m2+3m2
1m1−m3
1
=m3−3m1m2+2m3
1, (27.11)
which may be compared with equation (26.53) in the previous chapter.
1069
STATISTICS
Mirroring our discussion of the normalised central moments γrof a population
in subsection 26.5.5, we may also describe a sample in terms of the dimensionlessquantities
g
k=nk
nk/2
2=nk
sk;
g3andg4are called the sample skewness and kurtosis. Likwise, it is common to
define the excess kurtosis of a sample by g4−3.
27.2.4 Covariance and correlation
So far we have assumed that each data item of the sample consists of a single
number. Now let us suppose that each item of data consists of a pair of numbers,so that the sample is given by ( x
i,yi)(i=1,2,...,N ).
We may calculate the sample means, ¯xand¯y, and sample variances, s2
xand
s2
y,o ft h e xiandyivalues individually but these statistics do not provide any
measure of the relationship between the xiandyi. By analogy with our discussion
in subsection 26.12.3 we measure any interdependence between the xiandyiin
terms of the sample covariance , which is given by
Vxy=1
NNsummationdisplay
i=1(xi−¯x)(yi−¯y)
=(x−¯x)(y−¯y)
=xy−¯x¯y. (27.12)
Writing out the last expression in full, we obtain the form most useful for
calculations, which reads
Vxy=1
NparenleftBiggNsummationdisplay
i=1xiyiparenrightBigg
−1
N2parenleftBiggNsummationdisplay
i=1xiparenrightBiggparenleftBiggNsummationdisplay
i=1yiparenrightBigg
.
We may also define the closely related sample correlation by
rxy=Vxy
sxsy,
which can take values between −1 and +1. If the xiandyiare independent then
Vxy=0= rxy, and from (27.12) we see that xy=¯x¯y. It should also be noted
that the value of rxyis not altered by shifts in the origin or by changes in the
scale of the xioryi. In other words, if x/prime=ax+bandy/prime=cy+d,w h e r e a,
b,c,dare constants, then rx/primey/prime=rxy. Figure 27.1 shows scatter plots for several
two-dimensional random samples xi,yiof size N= 1000, each with a different
value of rxy.
1070
27.2 SAMPLE STATISTICS
rxy=0.0 rxy=0.1 rxy=0.5
rxy=−0.7 rxy=−0.9 rxy=0.99xy
Figure 27.1 Scatter plots for two-dimensional data samples of size N= 1000,
with various values of the correlation r. No scales are plotted, since the value
ofris unaffected by shifts of origin or changes of scale in xandy.ITen UK citizens are selected at random and their heights and weights are found to be as
follows (to the nearest cmorkgrespectively):
Person ABCDEFGH IJ
Height (cm) 194 168 177 180 171 190 151 169 175 182Weight (kg) 75 53 72 80 75 75 57 67 46 68
Calculate the sample correlation between the heights and weights.
In order to find the sample correlation, we begin by calculating the following sums (where
xiare the heights and yiare the weights)X
ixi= 1757 ,
X
iyi= 668 ,X
ix2
i= 310041 ,
X
iy2
i= 45746 ,
X
ixiyi= 118029 .
T h es a m p l ec o n s i s t so f N= 10 pairs of numbers, so the means of the xiand of the yiare
given by ¯x= 175 .7a n d ¯y=6 6 .8. Also, xy= 11802 .9. Similarly, the standard deviations
of the xiandyiare calculated, using (27.8), as
sx=
s
310041
10−
/1757
10
/2
=1 1.6,
sy=
s
45746
10−
/668
10
/2
=1 0.6.
1071
STATISTICS
Thus the sample correlation is given by
rxy=xy−¯x¯y
sxsy=11 802 .9−(175.7)(66 .8)
(11.6)(10 .6)=0.54.
Thus there is a moderate positive correlation between the heights and weights of the
people measured.
J
It is straightforward to generalise the above discussion to data samples of
arbitrary dimension, the only complication being one of notation. We chooseto denote the ith data item from an n-dimensional sample as ( x
(1)
i,x(2)i,...,x(n)
i),
where the bracketted superscript runs from 1 to nand labels the elements within
a given data item whereas the subscript iruns from 1 to Nand labels the data
items within the sample. In this n-dimensional case, we can define the sample
covariance matrix whose elements are
Vkl=x(k)x(l)−x(k)x(l)
and the sample correlation matrix with elements
rkl=Vkl
sksl.
Both these matrices are clearly symmetric but are notnecessarily positive definite.
27.3 Estimators and sampling distributions
In general, the population P(x) from which a sample x1,x2,...,x Nis drawn
isunknown .T h e central aim of statistics is to use the sample values xito infer
certain properties of the unknown population P(x), such as its mean, variance and
higher moments. To keep our discussion in general terms, let us denote the variousproperties of the population by a
1,a2,..., or collectively by a. Moreover, we make
the dependence of the population on the values of these quantities explicit by
writing the population as P(x|a). For the moment, we are assuming that the
sample values xiare independent and drawn from the same (one-dimensional)
population P(x|a), in which case
P(x|a)=P(x1|a)P(x2|a)···P(xN|a).
Suppose, we wish to estimate the value of one of the quantities a1,a2,...,w h i c h
we will denote simply by a. Since the sample values xiare our only source of
information, any estimate of amust be some function of the xi, i.e. some sample
statistic. Such a statistic is called an estimator ofaand is usually denoted by ˆa(x),
where xdenotes the sample elements x1,x2,...,x N.
Since an estimator ˆais a function of the sample values of the random variables
x1,x2,...,x N, it too must be a random variable. In other words, if a number of
random samples, each of the same size N, are taken from the (one-dimensional)
1072
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
population P(x|a) then the value of the estimator ˆawill vary from one sample to
the next and in general will not be equal to the true value a. This variation of
the estimator is described by its sampling distribution P(ˆa|a). From section 26.14,
this is given by
P(ˆa|a)dˆa=P(x|a)dNx,
where dNxis the infinitesimal ‘volume’ in x-space lying between the ‘surfaces’
ˆa(x)=ˆaandˆa(x)=ˆa+dˆa. The form of the sampling distribution generally
depends upon the estimator under consideration and upon the form of thepopulation from which the sample was drawn, including, as indicated, the truevalues of the quantities a. It is also usually dependent on the sample size N.IThe sample values x1,x2,...,x Nare drawn independently from a Gaussian distribution
with mean µand variance σ. Suppose that we choose the sample mean ¯xas our estimator
ˆµof the population mean. Find the sampling distributions of this estimator.
The sample mean ¯xis given by
¯x=1
N(x1+x2+···+xN),
where the xiare independent random variables distributed as xi∼N(µ, σ2). From our
discussion of multiple Gaussian distributions on page 1030, we see immediately that ¯xwill
also be Gaussian distributed as N(µ, σ2/N). In other words, the sampling distribution of
¯xis given by
P(¯x|µ, σ)=1p
2πσ2/Nexp
/
−(¯x−µ)2
2σ2/N
/
. (27.13)
Note that the variance of this distribution is σ2/N.
J
27.3.1 Consistency, bias and efficiency of estimators
For any particular quantity a, we may in fact define any number of different
estimators, each of which will have its own sampling distribution. The qualityof a given estimator ˆamay be assessed by investigating certain properties of its
sampling distribution P(ˆa|a). In particular, an estimator ˆais usually judged on
the three criteria of consistency ,biasandefficiency , each of which we now discuss.
Consistency
An estimator ˆaisconsistent if its value tends to the true value ain the large-sample
limit, i.e.
lim
N→∞ˆa=a.
Consistency is usually a minimum requirement for a useful estimator. An equiv-
alent statement of consistency is that in the limit of large Nthe sampling
1073
STATISTICS
distribution P(ˆa|a) of the estimator must satisfy
lim
N→∞P(ˆa|a)→δ(ˆa−a).
Bias
The expectation value of an estimator ˆais given by
E[ˆa]=integraldisplay
ˆaP(ˆa|a)dˆa=integraldisplay
ˆa(x)P(x|a)dNx, (27.14)
where the second integral extends over all possible values that can be taken by
the sample elements x1,x2,...,x N. This expression gives the expected mean value
ofˆafrom an infinite number of samples, each of size N.T h ebiasof an estimator
ˆais then defined as
b(a)=E[ˆa]−a. (27.15)
We note that the bias bdoes not depend on the measured sample values
x1,x2,...,x N. In general, though, it will depend on the sample size N, the func-
tional form of the estimator ˆaand, as indicated, on the true properties aof
the population, including the true value of aitself. If b=0t h e n ˆais called an
unbiased estimator of a.IAn estimator ˆais biased in such a way that E[ˆa]=a+b(a),w h e r et h eb i a s b(a)is given
by(b1−1)a+b2andb1andb2are known constants. Construct an unbiased estimator of a.
Let us first write E[ˆa] is the clearer form
E[ˆa]=a+(b1−1)a+b2=b1a+b2.
The task of constructing an unbiased estimator is now trivial, and an appropriate choice
isˆa/prime=(ˆa−b2)/b1, which (as required) has the expectation value
E[ˆa/prime]=E[ˆa]−b2
b1=a.
J
Efficiency
The variance of an estimator is given by
V[ˆa]=integraldisplay
(ˆa−E[ˆa])2P(ˆa|a)dˆa=integraldisplay
(ˆa(x)−E[ˆa])2P(x|a)dNx
(27.16)
and describes the spread of values ˆaabout E[ˆa] that would result from a large
number of samples, each of size N. An estimator with a smaller variance is said
to be more efficient than one with a larger variance. As we show in the next
section, for any given quantity aof the population there exists a theoretical lower
1074
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
limiton the variance of anyestimator ˆa. This result is known as Fisher’s inequality
(or the Cram´er–Rao inequality )a n dr e a d s
V[ˆa]≥parenleftbigg
1+∂b
∂aparenrightbigg2slashbigg
Ebracketleftbigg
−∂2lnP
∂a2bracketrightbigg
, (27.17)
where Pstands for the population P(x|a)a n d bis the bias of the estimator.
Denoting the quantity on the RHS of (27.17) by Vmin,t h e efficiency eof an
estimator is defined as
e=Vmin/V[ˆa].
An estimator for which e= 1 is called a minimum-variance orefficient estimator.
Otherwise, if e<1,ˆais called an inefficient estimator.
It should be noted that, in general, there is no unique ‘optimal’ estimator ˆafor
a particular property a. To some extent, there is always a trade-off between bias
and efficiency. One must often weigh the relative merits of an unbiased, inefficientestimator against another that is more efficient but slightly biased. Nevertheless, a
common choice is the best unbiased estimator (BUE), which is simply the unbiased
estimator ˆahaving the smallest variance V[ˆa].
Finally, we note that some qualities of estimators are related. For example,
suppose ˆais an unbiased estimator, so that E[ˆa]=aandV[ˆa]→0a s N→
∞. Using the Bienaym ´e–Chebyshev inequality discussed in subsection 26.5.3, it
follows immediately that ˆais also a consistent estimator. Nevertheless, it does not
follow that a consistent estimator is unbiased.IThe sample values x1,x2,...,x Nare drawn independently from a Gaussian distribution
with mean µand variance σ. Show that the sample mean ¯xis a consistent, unbiased,
minimum-variance estimator of µ.
We found earlier that the sampling distribution of ¯xis given by
P(¯x|µ, σ)=1p
2πσ2/Nexp
/
−(¯x−µ)2
2σ2/N
/
,
from which we see immediately that E[¯x]=µandV[¯x]=σ2/N. Thus ¯xis an unbiased
estimator of µ. Moreover, since it is also true that V[¯x]→0a sN→∞,¯xis a consistent
estimator of µ.
In order to determine whether ¯xis a minimum-variance estimator of µ,w em u s tu s e
Fisher’s inequality (27.17). Since the sample values xiare independent and drawn from a
Gaussian of mean µand standard deviation σ, we have
lnP(x|µ, σ)=−1
2NX
i=1
/
ln(2πσ2)+(xi−µ)2
σ2
/
,
and, on differentiating twice with respect to µ, we find
∂2lnP
∂µ2=−N
σ2.
This is independent of the xiand so its expectation value is also equal to −N/σ2.W i t h b
1075
STATISTICS
set equal to zero in (27.17), Fisher’s inequality thus states that, for anyunbiased estimator
ˆµof the population mean,
V[ˆµ]≥σ2
N.
Since V[¯x]=σ2/N, the sample mean ¯xis a minimum-variance estimator of µ.
J
27.3.2 Fisher’s inequality
As mentioned above, Fisher’s inequality provides a lower limit on the variance of
anyestimator ˆaof the quantity a;i tr e a d s
V[ˆa]≥parenleftbigg
1+∂b
∂aparenrightbigg2slashbigg
Ebracketleftbigg
−∂2lnP
∂a2bracketrightbigg
, (27.18)
where Pstands for the population P(x|a)a n d bis the bias of the estimator.
We now present a proof of this inequality. Since the derivation is somewhatcomplicated, and many of the details are unimportant, this section can be omittedon a first reading. Nevertheless, some aspects of the proof will be useful whenthe efficiency of maximum-likelihood estimators is discussed in section 27.5.IProve Fisher’s inequality (27.18).
The normalisation of P(x|a)i sg i v e nb yZ
P(x|a)dNx=1, (27.19)
where dNx=dx1dx2···dxNand the integral extends over all the allowed values of the
sample items xi. Differentiating (27.19) with respect to the parameter a,w eo b t a i nZ∂P
∂adNx=
Z∂lnP
∂aPdNx=0. (27.20)
We note that the second integral is simply the expectation value of ∂lnP/∂a,w h e r et h e
average is taken over all possible samples xi,i=1,2,...,N .F u r t h e r ,b ye q u a t i n gt w o
expressions for ∂E[ˆa]/∂a, obtained by differentiating (27.15) and (27.14) with respect to a
we obtain, dropping the functional dependencies, a second relationship,
1+∂b
∂a=
Z
ˆa∂P
∂adNx=
Z
ˆa∂lnP
∂aPdNx. (27.21)
Now, multiplying (27.20) by α(a), where α(a)i sanyfunction of a, and subtracting the
result from (27.21), we obtainZ
[ˆa−α(a)]∂lnP
∂aPdNx=1+∂b
∂a.
At this point we must invoke the Schwarz inequality proved in subsection 8.1.3. The proof
is trivially extended to multiple integrals and shows that for two real functions, g(x)a n d
h(x),/Z
g2(x)dNx
//Z
h2(x)dNx
/
≥
/Z
g(x)h(x)dNx
/2
. (27.22)
1076
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
If we now let g=[ˆa−α(a)]√
Pandh=(∂lnP/∂a)√
P, we find/Z
[ˆa−α(a)]2PdNx
/
/"Z
/∂lnP
∂a
/2
PdNx
/#
≥
/
1+∂b
∂a
/2
.
On the LHS, the factor in braces represents the expected spread of ˆa-values around the
point α(a). The minimum value that this integral may take occurs when α(a)=E[ˆa].
Making this substitution, we recognise the integral as the variance V[ˆa], and so obtain the
result
V[ˆa]≥
/
1+∂b
∂a
/2
/"Z
/∂lnP
∂a
/2
PdNx
/#−1
. (27.23)
We note that the factor in brackets is the expectation value of ( ∂lnP/∂a)2.
Fisher’s inequality is, in fact, often quoted in the form (27.23). We may recover the form
(27.18) by noting that on differentiating (27.20) with respect to awe obtainZ
/∂2lnP
∂a2P+∂lnP
∂a∂P
∂a
/
dNx=0.
Writing ∂P/∂a as (∂lnP/∂a)Pand rearranging we find thatZ
/∂lnP
∂a
/2
PdNx=−
Z∂2lnP
∂a2PdNx.
Substituting this result in (27.23) gives
V[ˆa]≥−
/
1+∂b
∂a
/2
/Z∂2lnP
∂a2PdNx
/−1
.
Since the factor in brackets is the expectation value of ∂2lnP/∂a2, we have recovered
result (27.18).
J
27.3.3 Standard errors on estimators
For a given sample x1,x2,...,x N, we may calculate the value of an estimator ˆa(x)
for the quantity a. It is also necessary, however, to give some measure of the
statistical uncertainty in this estimate. One way of characterising this uncertainty
is with the standard deviation of the sampling distribution P(ˆa|a), which is given
simply by
σˆa=(V[ˆa])1/2. (27.24)
If the estimator ˆa(x) were calculated for a large number of samples, each of size
N, then the standard deviation of the resulting ˆavalues would be given by (27.24).
Consequently, σˆais called the standard error on our estimate.
In general, however, the standard error σˆadepends on the true values of some
or all of the quantities aand they may be unknown. When this occurs, one must
substitute estimated values of any unknown quantities into the expression for σˆa
in order to obtain an estimated standard error ˆσˆa. One then quotes the result as
a=ˆa±ˆσˆa.
1077
STATISTICSITen independent sample values xi,i=1,2,...,10, are drawn at random from a Gaussian
distribution with standard deviation σ=1. The sample values are as follows (to two decimal
places):
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Estimate the population mean µ, quoting the standard error on your result.
We have shown in the final worked example of subsection 27.3.1 that, in this case, ¯xis
a consistent, unbiased, minimum-variance estimator of µand has variance V[¯x]=σ2/N.
Thus, our estimate of the population mean with its associated standard error is
ˆµ=¯x±σ√
N=1.11±0.32.
If the true value of σhad not been known, we would have needed to use an estimated
value ˆσin the expression for the standard error. Useful basic estimators of σare discussed
in subsection 27.4.2.
J
It should be noted that the above approach is most meaningful for unbiased
estimators. In this case, E[ˆa]=aand so σˆadescribes the spread of ˆa-values about
the true value a. For a biased estimator, however, the spread about the true value
ais given by the root mean square error /epsilon1ˆa, which is defined by
/epsilon12
ˆa=E[(ˆa−a)2]
=E[(ˆa−E[ˆa])2]+(E[ˆa]−a)2
=V[ˆa]+b(a)2.
We see that /epsilon12
ˆais the sum of the variance of ˆaand the square of the bias and so
can be interpreted as the sum of squares of statistical and systematic errors. For
a biased estimator, it is often more appropriate to quote the result as
a=ˆa±/epsilon1ˆa.
As above, it may be necessary to use estimated values ˆain the expression for the
root mean square error and thus to quote only an estimate ˆ/epsilon1ˆaoftheerror .
27.3.4 Confidence limits on estimators
An alternative (and often equivalent) way of quoting a statistical error is with a
confidence interval . Let us assume that, other than the quantity of interest a,t h e
quantities ahave known fixed values. Thus we denote the sampling distribution
ofˆabyP(ˆa|a). For any particular value of a, one can determine the two values
ˆaα(a)a n d ˆaβ(a) such that
Pr(ˆa<ˆaα(a)) =integraldisplayˆaα(a)
−∞P(ˆa|a)dˆa=α, (27.25)
Pr(ˆa>ˆaβ(a)) =integraldisplay∞
ˆaβ(a)P(ˆa|a)dˆa=β. (27.26)
1078
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
ˆaP(ˆa|a)
ˆaα(a) ˆaβ(a)α β
Figure 27.2 The sampling distribution P(ˆa|a)o fs o m ee s t i m a t o r ˆafor a given
value of a. The shaded regions indicate the two probabilities Pr( ˆa<ˆaα(a)) =α
and Pr( ˆa>ˆaβ(a)) =β.
This is illustrated in figure 27.2. Thus, for any particular value of a, the probability
that the estimator ˆalies within the limits ˆaα(a)a n d ˆaβ(a)i sg i v e nb y
Pr(ˆaα(a)<ˆa<ˆaβ(a)) =integraldisplayˆaβ(a)
ˆaα(a)P(ˆa|a)dˆa=1−α−β.
Now, let us suppose that from our sample x1,x2,...,x N, we actually obtain the
value ˆaobsfor our estimator. If ˆais a good estimator of athen we would expect
ˆaα(a)a n d ˆaβ(a) to be monotonically increasing functions of a(i.e.ˆaαandˆaβboth
change in the samesense as awhen the latter is varied). Assuming this to be the
case, we can uniquely define the two numbers a−anda+by the relationships
ˆaα(a+)=ˆaobs and ˆaβ(a−)=ˆaobs.
From (27.25) and (27.26) it follows that
Pr(a+<a)=α and Pr( a−>a)=β,
which when taken together imply
Pr(a−<a<a +)=1−α−β. (27.27)
Thus, from our estimate ˆaobs, we have determined two values a−anda+such that
this interval contains the true value of awith probability 1 −α−β. It should be
emphasised that a−anda+are random variables. If a large number of samples,
each of size N, were analysed then the interval [ a−,a+] would contain the true
value aon a fraction 1 −α−βof occasions.
The interval [ a−,a+] is called a confidence interval onaat the confidence
level1−α−β. The values a−and a+themselves are called respectively the
lower confidence limit and the upper confidence limit at this confidence level. In
practice, the confidence level is often quoted as a percentage. A convenient way
1079
STATISTICS
ˆaP(ˆa|a−) P(ˆa|a+)
ˆaobsα β
Figure 27.3 An illustration of how the observed value of the estimator, ˆaobs,
and the given values αandβdetermine the two confidence limits a−anda+,
which are such that ˆaα(a+)=ˆaobs=ˆaβ(a−).
of presenting our results isintegraldisplayˆaobs
−∞P(ˆa|a+)dˆa=α, (27.28)
integraldisplay∞
ˆaobsP(ˆa|a−)dˆa=β. (27.29)
The confidence limits may then be found by solving these equations for a−and
a+either analytically or numerically.
Occasionally one might not combine the results (27.28) and (27.29) but use
either one or the other to provide a one-sided confidence interval on a. Whenever
the results are combined to provide a two-sided confidence interval, the interval
isnotspecified uniquely by the confidence level 1 −α−β. In other words, there
are generally an infinite number of intervals [ a−,a+] for which (27.27) holds.
To specify a unique interval, one often chooses α=β, resulting in the central
confidence interval ona. All cases can be covered by calculating the quantities
c=ˆa−a−andd=a+−ˆaand quoting the result of an estimate as
a=ˆa+d
−c.
We have so far assumed that the quantities aother than the quantity of interest
aare known in advance. If this is not the case then, in principle, the construction
of confidence limits is considerably more complicated. This is discussed briefly insubsection 27.3.6.
27.3.5 Confidence limits for a Gaussian sampling distribution
An important special case occurs when the sampling distribution is Gaussian; if
the mean is aand the standard deviation is σ
ˆathen
P(ˆa|a, σˆa)=1radicalBig
2πσ2
ˆaexpbracketleftbigg
−(ˆa−a)2
2σ2
ˆabracketrightbigg
. (27.30)
1080
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
For almost any (consistent) estimator ˆa, the sampling distribution will tend to
this form in the large-sample limit N→∞, as a consequence of the central limit
theorem. For a sampling distribution of the form (27.30), the above procedurefor determining confidence intervals becomes straightforward. Suppose, from our
sample, we obtain the value ˆa
obsfor our estimator. In this case, equations (27.28)
and (27.29) become
Φparenleftbiggˆaobs−a+
σˆaparenrightbigg
=α,
1−Φparenleftbiggˆaobs−a−
σˆaparenrightbigg
=β,
where Φ( z) is the cumulative probability function for the standard Gaussian distri-
bution, discussed in subsection 26.9.1. Solving these equations for a−anda+gives
a−=ˆaobs−σˆaΦ−1(1−β), (27.31)
a+=ˆaobs+σˆaΦ−1(1−α); (27.32)
we have used the fact that Φ−1(α)=−Φ−1(1−α) to make the equations symmetric.
The value of the inverse function Φ−1(z) can be read off directly from table 26.3,
given in subsection 26.9.1. For the normally-used central confidence interval onehasα=β. In this case, we see that quoting a result using the standard error, as
a=ˆa±σ
ˆa, (27.33)
is equivalent to taking Φ−1(1−α) = 1. From table 26.3, we find α=1−0.8413 =
0.1587, and so this corresponds to a confidence level of 1 −2(0.1587)≈0.683.
Thus, the standard error limits give the 68.3% central confidence interval.ITen independent sample values xi(i=1,2,...,10)are drawn at random from a Gaussian
distribution with standard deviation σ=1. The sample values are as follows (to two decimal
places):
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Find the 90% central confidence interval on the population mean µ.
Our estimator ˆµis the sample mean ¯x. As shown towards the end of section 27.3, the
sampling distribution of ¯xis Gaussian with mean E[¯x] and variance V[¯x]=σ2/N.S i n c e
σ= 1 in this case, the standard error is given by σˆx=σ/√
N=0.32. Moreover, in
subsection 27.3.3, we found the mean of the above sample to be ¯x=1.11.
For the 90% central confidence interval, we require α=β=0.05. From table 26.3, we
find
Φ−1(1−α)=Φ−1(0.95) = 1 .65,
and using (27.31) and (27.32) we obtain
a−=¯x−1.65σ¯x=1.11−(1.65)(0 .32) = 0 .58,
a+=¯x+1.65σ¯x=1.11 + (1 .65)(0 .32) = 1 .64.
Thus, the 90% central confidence interval on µis [0.58,1.64]. For comparison, the true
value used to create the sample was µ=1 .
J
1081
STATISTICS
In the case where the standard error σˆain (27.33) is not known in advance,
one must use a value ˆσˆaestimated from the sample. In principle, this complicates
somewhat the construction of confidence intervals, since properly one shouldconsider the two-dimensional joint sampling distribution P(ˆa,ˆσ
ˆa|a). Nevertheless,
in practice, provided ˆσˆais a fairly good estimate of σˆathe above procedure may
be applied with reasonable accuracy. In the special case where the sample valuesx
iare drawn from a Gaussian distribution with unknown µandσ,i ti si nf a c t
possible to obtain exact confidence intervals on the mean µ, for a sample of any
sizeN, using Student’s t-distribution. This is discussed in section 27.7.5.
27.3.6 Estimation of several quantities simultaneously
Suppose one uses a sample x1,x2,...,x Nto calculate the values of several es-
timators ˆa1,ˆa2,...,ˆaM(collectively denoted by ˆa)o ft h eq u a n t i t i e s a1,a2,...,a M
(collectively denoted by a) that describe the population from which the sample was
drawn. The joint sampling distribution of these estimators is an M-dimensional
PDF P(ˆa|a)g i v e nb y
P(ˆa|a)dMˆa=P(x|a)dNx.ISample values x1,x2,...,x Nare drawn independently from a Gaussian distribution with
mean µand standard deviation σ. Suppose we choose the sample mean ¯xand sample stan-
dard deviation srespectively as estimators ˆµandˆσ. Find the joint sampling distribution of
these estimators.
Since each data value xiin the sample is assumed to be independent of the others, the
joint probability distribution of sample values is given by
P(x|µ, σ)=( 2 πσ2)−N/2exp
/
−
P
i(xi−µ)2
2σ2
/
.
We may rewrite the sum in the exponent as follows:X
i(xi−µ)2=
X
i(xi−¯x+¯x−µ)2
=
X
i(xi−¯x)2+2 (¯x−µ)
X
i(xi−¯x)+
X
i(¯x−µ)2
=Ns2+N(¯x−µ)2,
where in the last line we have used the fact that
P
i(xi−¯x) = 0. Hence, for given values
ofµandσ, the sampling distribution is in fact a function only of the sample mean ¯xand
the standard deviation s. Thus the sampling distribution of ¯xandsmust satisfy
P(¯x, s|µ, σ)d¯xd s=( 2πσ2)−N/2exp
/
−N[(¯x−µ)2+s2]
2σ2
/
dV, (27.34)
where dV=dx1dx2···dxNis an element of volume in the sample space which yields
simultaneously values of ¯xandsthat lie within the region bounded by [ ¯x,¯x+d¯x]a n d
[s, s+ds]. Thus our only remaining task is to express dVin terms of ¯xandsand their
differentials.
1082
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
LetSbe the point in sample space representing the sample ( x1,x2,...,x N). For given
values of ¯xands, we require the sample values to satisfy both the conditionX
ixi=N¯x,
which defines an ( N−1)-dimensional hyperplane in the sample space, and the conditionX
i(xi−¯x)2=Ns2,
which defines an ( N−1)-dimensional hypersphere. Thus Sis constrained to lie in the
intersection of these two hypersurfaces, which is itself an ( N−2)-dimensional hypersphere.
Now, the volume of an ( N−2)-dimensional hypersphere is proportional to sN−1. It follows
from this that the volume dVbetween two concentric ( N−2)-dimensional hyperspheres
of radius√
Nsand√
N(s+ds)a n dt w o( N−1)-dimensional hyperplanes corresponding
to¯xand¯x+d¯xis
dV=AsN−2ds d¯x,
where Ais some constant. Thus, substituting this expression for dVinto (27.34), we find
P(¯x, s|µ, σ)=C1exp
/
−N(¯x−µ)2
2σ2
/
C2sN−2exp
/
−Ns2
2σ2
/
=P(¯x|µ, σ)P(s|σ),
(27.35)
where C1andC2are constants. We have written P(¯x, s|µ, σ) in this form to illustrate that
it separates naturally into two parts, one depending only on ¯xand the other only on s.
Thus, ¯xandsareindependent variables. Separate normalisations of the two factors in
(27.35) require
C1=
/N
2πσ2
/1/2
and C2=2
/N
2σ2
/(N−1)/21
Γ
/;1
2(N−1)
/,
where the calculation of C2requires the use of the gamma function discussed in the
Appendix.
J
Themarginal sampling distribution of any one of the estimators ˆaiis given
simply by
P(ˆai|a)=integraldisplay
···integraldisplay
P(ˆa|a)dˆa1···dˆai−1dˆai+1···dˆaM,
and the expectation value E[ˆai] and variance V[ˆai]o fˆaiare again given by (27.14)
and (27.16) respectively. By analogy with the one-dimensional case, the standarderror σ
ˆaion the estimator ˆaiis given by the positive square root of V[ˆai]. With
several estimators, however, it is usual to quote their full covariance matrix. This
M×Mmatrix has elements
Vij=C o v [ ˆai,ˆaj]=integraldisplay
(ˆai−E[ˆai])(ˆaj−E[ˆaj])P(ˆa|a)dMˆa
=integraldisplay
(ˆai−E[ˆai])(ˆaj−E[ˆaj])P(x|a)dNx.
Fisher’s inequality can be generalised to the multi-dimensional case. Adapting
the proof given in subsection 27.3.2, one may show that, in the case where the
1083
STATISTICS
estimators are efficient and have zero bias, the elements of the inverse of the
covariance matrix are given by
(V−1)ij=Ebracketleftbigg
−∂2lnP
∂ai∂ajbracketrightbigg
, (27.36)
where Pdenotes the population P(x|a) from which the sample is drawn. The
quantity on the RHS of (27.36) is the element Fijof the so-called Fisher matrix
Fof the estimators.ICalculate the covariance matrix of the estimators ¯xandsin the previous example.
As shown in (27.35), the joint sampling distribution P(¯x, s|µ, σ) factorises, and so the
estimators ¯xandsare independent. Thus, we conclude immediately that
Cov[¯x, s]=0 .
Since we have already shown in the worked example at the end of subsection (27.3.1) that
V[¯x]=σ2/N, it only remains to calculate V[s]. From (27.35), we find
E[sr]=C2
Z∞
0sN−2+rexp
/
−Ns2
2σ2
/
ds=
/2
N
/r/2Γ
/;1
2(N−1+r)
/
Γ
/;1
2(N−1)
/σr,
where we have evaluated the integral using the definition of the gamma function given in
the Appendix. Thus, the expectation value of the sample standard deviation is
E[s]=
/2
N
/1/2Γ
/;1
2N
/
Γ
/;1
2(N−1)
/σ, (27.37)
and its variance is given by
V[s]=E[s2]−(E[s])2=σ2
N
/8/</:N−1−2
/"
Γ
/;1
2N
/
Γ
/;1
2(N−1)
/
/#2
/9/=/;;
We note, in passing, that (27.37) shows that sis abiased estimator of σ.
J
The idea of a confidence interval can also be extended to the case where several
quantities are estimated simultaneously but then the practical construction of an
interval is considerably more complicated. The general approach is to constructanM-dimensional confidence region Rina-space. By analogy with the one-
dimensional case, for a given confidence level of (say) 1 −α, one first constructs
a region ˆRinˆa-space, such that
integraldisplayintegraldisplay
ˆRP(ˆa|a)dMˆa=1−α.
A common choice for such a region is that bounded by the ‘surface’ P(ˆa|a)=
constant. By considering all possible values aand the values of ˆalying within
the region ˆR, one can construct a 2 M-dimensional region in the combined space
(ˆa,a). Suppose now that, from our sample x, the values of the estimators are
ˆai,obs,i=1,2,...,M . The intersection of the M‘hyperplanes’ ˆai=ˆai,obswith
the 2 M-dimensional region will determine an M-dimensional region which, when
1084
27.3 ESTIMATORS AND SAMPLING DISTRIBUTIONS
a1a2
ˆa1ˆa2
atrue atrue
ˆaobs ˆaobs(a) (b)
Figure 27.4 (a) The ellipse Q(ˆa,a)=cinˆa-space. (b) The ellipse Q(a,ˆaobs)=c
ina-space that corresponds to a confidence region Rat the level 1 −α,w h e n
csatisfies (27.39).
projected onto a-space, will determine a confidence limit Rat the confidence
level 1−α. It is usually the case that this confidence region has to be evaluated
numerically.
The above procedure is clearly rather complicated in general and a simpler
approximate method that uses the likelihood function is discussed in subsec-tion 27.5.5. As a consequence of the central limit theorem, however, in thelarge-sample limit, N→∞, the joint sampling distribution P(ˆa|a) will tend, in
general, towards the multivariate Gaussian
P(ˆa|a)=1
(2π)M/2|V|1/2expbracketleftbig
−1
2Q(ˆa,a)bracketrightbig
, (27.38)
where Vis the covariance matrix of the estimators and the quadratic form Qis
given by
Q(ˆa,a)=(ˆa−a)TV−1(ˆa−a).
Moreover, in the limit of large N, the inverse covariance matrix tends to the
Fisher matrix Fgiven in (27.36), i.e. V−1→F.
For the Gaussian sampling distribution (27.38), the process of obtaining confi-
dence intervals is greatly simplified. The surfaces of constant P(ˆa|a) correspond
to surfaces of constant Q(ˆa,a), which have the shape of M-dimensional ellipsoids
inˆa-space, centred on the true values a. In particular, let us suppose that the
ellipsoid Q(ˆa,a)=c(where cis some constant) contains a fraction 1 −αsay of
the total probability. Now suppose that, from our sample x, we obtain the values
ˆaobsfor our estimators. Because of the obvious symmetry of the quadratic form
Qwith respect to aandˆa, it is clear that the ellipsoid Q(a,ˆaobs)=cina-space
that is centred on ˆaobsshould contain the true values awith probability 1 −α.
Thus Q(a,ˆaobs)=cdefines our required confidence region Rat this confidence
level. This is illustrated in figure 27.4 for the two-dimensional case.
1085
STATISTICS
It remains only to determine the constant ccorresponding to the confidence
level 1−α. As discussed in subsection 26.15.2, the quantity Q(ˆa,a) is distributed
as a χ2variable of order M. Thus, the confidence region corresponding to the
confidence level 1 −αis given by Q(a,ˆaobs)=c, where the constant csatisfies
integraldisplayc
0P(χ2
M)d(χ2
M)=1−α, (27.39)
andP(χ2
M) is the chi-squared PDF of order M, discussed in subsection 26.9.4. This
integral may be evaluated numerically to determine the constant c. Alternatively,
some reference books tabulate the values of ccorresponding to given confidence
levels and various values of M.
27.4 Some basic estimators
In many cases, one does not know the functional form of the population from
which a sample is drawn. Nevertheless, in a case where the sample valuesx
1,x2,...,x Nare each drawn independently from a one-dimensional population
P(x), it is possible to construct some basic estimators for the moments and central
moments of P(x). In this section, we investigate the estimating properties of the
common sample statistics presented in section 27.2. In fact, expectation values
and variances of these sample statistics can be calculated without prior knowledge
of the functional form of the population; they depend only on the sample size N
and certain moments and central moments of P(x).
27.4.1 Population mean µ
Let us suppose that the parent population P(x) has mean µand variance σ2.A n
obvious estimator ˆµof the population mean is the sample mean ¯x.P r o v i d e d µ
andσ2are both finite, we may apply the central limit theorem directly to obtain
exact expressions, valid for samples of any size N, for the expectation value and
variance of ¯x. From parts (i) and (ii) of the central limit theorem, discussed in
section 26.10, we immediately obtain
E[¯x]=µ, V [¯x]=σ2
N. (27.40)
Thus we see that ¯xis an unbiased estimator of µ. Moreover, we note that the
standard error in ¯xisσ/√
N, and so the sampling distribution of ¯xbecomes more
tightly centred around µas the sample size Nincreases. Indeed, since V[¯x]→0
asN→∞,¯xis also a consistent estimator of µ.
In the limit of large N, we may in fact obtain an approximate form for the
full sampling distribution of ¯x. Part (iii) of the central limit theorem (see section
26.10) tells us immediately that, for large N, the sampling distribution of ¯xis
1086
27.4 SOME BASIC ESTIMATORS
given approximately by the Gaussian form
P(¯x|µ, σ)≈1radicalbig
2πσ2/Nexpbracketleftbigg
−(¯x−µ)2
2σ2/Nbracketrightbigg
.
Note that this does notdepend on the form of the original parent population.
If, however, the parent population is in fact Gaussian then this result is exact
for samples of anysizeN(as is immediately apparent from our discussion of
multiple Gaussian distributions in subsection 26.9.1).
27.4.2 Population variance σ2
An estimator for the population variance σ2is not so straightforward to define
as one for the mean. Complications arise because, in many cases, the true mean
of the population µis not known. Nevertheless, let us begin by considering the
case where in fact µis known. In this event, a useful estimator is
hatwideσ2=1
NNsummationdisplay
i=1(xi−µ)2=parenleftBigg
1
NNsummationdisplay
i=1x2
iparenrightBigg
−µ2. (27.41)IShow that
bσ2is an unbiased and consistent estimator of the population variance σ2.
The expectation value of
bσ2is given by
E[
bσ2]=1
NE
/"NX
i=1x2
i
/#
−µ2=E[x2
i]−µ2=µ2−µ2=σ2,
from which we see that the estimator is unbiased. The variance of the estimator is
V[
bσ2]=1
N2V
/"NX
i=1x2
i
/#
+V[µ2]=1
NV[x2
i]=1
N(µ4−µ2
2),
in which we have used that fact that V[µ2]=0a n d V[x2
i]=E[x4
i]−(E[x2
i])2=µ4−µ2
2,
where µris the rth population moment. Since
bσ2is unbiased and V[
bσ2]→0a s N→∞,
showing that it is also a consistent estimator of σ2, the result is established.
J
If the true mean of the population is unknown, however, a natural alternative
is to replace µby¯xin (27.41), so that our estimator is simply the sample variance
s2given by
s2=1
NNsummationdisplay
i=1x2
i−parenleftBigg
1
NNsummationdisplay
i=1xiparenrightBigg2
.
In order to determine the properties of this estimator, we must calculate E[s2]
andV[s2]. This task is straightforward but lengthy. However, for the investigation
of the properties of a central moment of the sample, there exists a useful trick
that simplifies the calculation. We can assume, with no loss of generality, that
1087
STATISTICS
the mean µ1of the population from which the sample is drawn is equal to zero.
With this assumption, the population central moments, νr,a r ei d e n t i c a lt ot h e
corresponding moments µr, and we may perform our calculation in terms of the
latter. At the end, however, we replace µrbyνrin the final result and so obtain
a general expression that is valid even in cases where µ1/negationslash=0 .ICalculate E[s2]andV[s2]for a sample of size N.
The expectation value of the sample variance s2for a sample of size Nis given by
E[s2]=1
NE
/"X
ix2
i
/#
−1
N2E
/2/4
/ X
ixi
/!2
/3/5
=1
NNE[x2
i]−1
N2E
/2/6/4
X
ix2
i+
X
i,j
j/negationslash=ixixj
/3/7/5. (27.42)
The number of terms in the double summation in (27.42) is N(N−1), so we find
E[s2]=E[x2
i]−1
N2(NE[x2
i]+N(N−1)E[xixj]).
Now, since the sample elements xiandxjare independent, E[xixj]=E[xi]E[xj]=0 ,
assuming the mean µ1of the parent population to be zero. Denoting the rth moment of
the population by µr, we thus obtain
E[s2]=µ2−µ2
N=N−1
Nµ2=N−1
Nσ2, (27.43)
where in the last line we have used the fact that the population mean is zero, and so
µ2=ν2=σ2. However, the final result is also valid in the case where µ1/negationslash=0 .
Using the above method, we can also find the variance of s2, although the algebra is
rather heavy going. The variance of s2is given by
V[s2]=E[s4]−(E[s2])2, (27.44)
where E[s2] is given by (27.43). We therefore need only consider how to calculate E[s4],
where s4is given by
s4=
/"P
ix2
i
N−
/
P
ixi
N
/2
/#2
=(
P
ix2
i)2
N2−2(
P
ix2
i)(
P
ixi)2
N3+(
P
ixi)4
N4. (27.45)
We will consider in turn each of the three terms on the RHS. In the first term, the sum
(
P
ix2
i)2can be written as/ X
ix2
i
/!2
=
X
ix4i+
X
i,j
j/negationslash=ix2
ix2j,
where the first sum contains Nterms and the second contains N(N−1) terms. Since the
sample elements xiandxjare assumed independent, we have E[x2
ix2j]=E[x2
i]E[x2
j]=µ2
2,
1088
27.4 SOME BASIC ESTIMATORS
and so
E
/2/4
/ X
ix2
i
/!2
/3/5=Nµ4+N(N−1)µ2
2.
Turning to the second term on the RHS of (27.45),/ X
ix2
i
/!/ X
ixi
/!2
=
X
ix4
i+
X
i,j
j/negationslash=ix3
ixj+
X
i,j
j/negationslash=ix2
ix2j+
X
i,j,k
k/negationslash=j/negationslash=ix2
ixjxk.
Since the mean of the population has been assumed to equal zero, the expectation values
of the second and fourth sums on the RHS vanish. The first and third sums contain N
andN(N−1) terms respectively, and so
E
/2/4
/ X
ix2
i
/!/ X
ixi
/!2
/3/5=Nµ4+N(N−1)µ2
2.
Finally, we consider the third term on the RHS of (27.45), and write/ X
ixi
/!4
=
X
ix4
i+
X
i,j
j/negationslash=ix3
ixj+
X
i,j
j/negationslash=ix2
ix2j+
X
i,j,k
k/negationslash=j/negationslash=ix2
ixjxk+
X
i,j,k,l
l/negationslash=k/negationslash=j/negationslash=ixixjxkxl.
The expectation values of the second, fourth and fifth sums are zero, and the first and third
sums contain Nand 3 N(N−1) terms respectively (for the third sum, there are N(N−1)/2
ways of choosing iandj, and the multinomial coefficient of x2
ix2jis 4!/(2!2!) = 6). Thus
E
/2/4
/ X
ixi
/!4
/3/5=Nµ4+3N(N−1)µ2
2.
Collecting together terms, we therefore obtain
E[s4]=(N−1)2
N3µ4+(N−1)(N2−2N+3 )
N3µ2
2, (27.46)
which, together with the result (27.43), may be substituted into (27.44) to obtain finally
V[s2]=(N−1)2
N3µ4−(N−1)(N−3)
N3µ2
2
=N−1
N3[(N−1)ν4−(N−3)ν2
2], (27.47)
where in the last line we have used again the fact that, since the population mean is zero,
µr=νr. However, result (27.47) holds even when the population mean is not zero.
J
From (27.43), we see that s2is abiased estimator of σ2, although the bias
becomes negligible for large N. However, it immediately follows that an unbiased
estimator of σ2is given simply by
hatwideσ2=N
N−1s2, (27.48)
where the multiplicative factor N/(N−1) is often called Bessel’s correction . Thus
1089
STATISTICS
in terms of the sample values xi,i=1,2,...,N , an unbiased estimator of the
population variance σ2is given by
hatwideσ2=1
N−1Nsummationdisplay
i=1(xi−¯x)2. (27.49)
Using (27.47), we find that the variance of the estimatorhatwideσ2is
V[hatwideσ2]=parenleftbiggN
N−1parenrightbigg2
V[s2]=1
Nparenleftbigg
ν4−N−3
N−1ν2
2parenrightbigg
,
where νris the rth central moment of the parent population. We note that,
since E[hatwideσ2]=σ2andV[hatwideσ2]→0a s N→∞, the statistichatwideσ2is also a consistent
estimator of the population variance.
27.4.3 Population standard deviation σ
The standard deviation σof a population is defined as the positive square root of
the population variance σ2(as, indeed, our notation suggests). Thus, it is common
practice to take the positive square root of the variance estimator as our estimatorforσ.T h u s ,w et a k e
ˆσ=parenleftBighatwideσ
2parenrightBig1/2
, (27.50)
wherehatwideσ2is given by either (27.41) or (27.48), depending on whether the population
mean µis known or unknown. Because of the square root in the definition of
ˆσ, it is not possible in either case to obtain an exact expression for E[ˆσ]a n d
V[ˆσ]. Indeed, although in each case the estimator is the positive square root of
an unbiased estimator of σ2,i ti snotitself an unbiased estimator of σ. However,
the bias does becomes negligible for large N.IObtain approximate expressions for E[ˆσ]andV[ˆσ]for a sample of size Nin the case
where the population mean µis unknown.
As the population mean is unknown, from (27.50) and (27.48) our estimator is given by
ˆσ=
/N
N−1
/1/2
s,
where sis the sample standard deviation. The expectation value of this estimator is given
by
E[ˆσ]=
/N
N−1
/1/2
E[(s2)1/2]≈
/N
N−1
/1/2
(E[s2])1/2=σ.
An approximate expression for the variance of ˆσmay be found using (27.47) and is given
1090
27.4 SOME BASIC ESTIMATORS
by
V[ˆσ]=N
N−1V[(s2)1/2]≈N
N−1
/d
d(s2)(s2)1/2
/2
s2=E[s2]V[s2]
≈N
N−1
/1
4s2
/
s2=E[s2]V[s2].
Using the expressions (27.43) and (27.47) for E[s2]a n d V[s2] respectively, we obtain
V[ˆσ]≈1
4Nν2
/
ν4−N−3
N−1ν2
2
/
.
J
27.4.4 Population moments µr
We may straightforwardly generalise our discussion of estimation of the popu-
lation mean µ(=µ1) in section 27.4.1 to the estimation of the rth population
moment µr. An obvious choice of estimator is the rth sample moment mr.T h e
expectation value of mris given by
E[mr]=1
NNsummationdisplay
i=1E[xr
i]=Nµr
N=µr,
and so it is an unbiased estimator of µr.
The variance of mrmay be found in a similar manner, although the calculation
is a little more complicated. We find that
V[mr]=E[(mr−µr)2]
=1
N2E
parenleftBiggsummationdisplay
ixr
i−NµrparenrightBigg2
=1
N2E
summationdisplay
ix2r
i+summationdisplay
isummationdisplay
j/negationslash=ixr
ixrj−2Nµrsummationdisplay
ixr
i+N2µ2
r
=1
Nµ2r−µ2
r+1
N2summationdisplay
isummationdisplay
j/negationslash=iE[xr
ixrj]. (27.51)
However, since the sample values xiare assumed to be independent, we have
E[xr
ixrj]=E[xr
i]E[xr
j]=µ2
r. (27.52)
The number of terms in the sum on the RHS of (27.51) is N(N−1), and so we
find
V[mr]=1
Nµ2r−µ2
r+N−1
Nµ2
r=µ2r−µ2
r
N. (27.53)
Since E[mr]=µrandV[mr]→0a sN→∞,t h e rth sample moment mris also a
consistent estimator of µr.
1091
STATISTICSIFind the covariance of the sample moments mrandmsfor a sample of size N.
We obtain the covariance of the sample moments mrandmsin a similar manner to that
used above to obtain the variance of mr. From the definition of covariance, we have
Cov[mr,ms]=E[(mr−µr)(ms−µs)]
=1
N2E
/" / X
ixr
i−Nµr
/!/ X
jxs
j−Nµs
/!/#
=1
N2E
/2/4
X
ixr+s
i+
X
i
X
j/negationslash=ixr
ixsj−Nµr
X
jxs
j−Nµs
X
ixr
i+N2µrµs
/3/5
Assuming the xito be independent, we may again use result (27.52) to obtain
Cov[mr,ms]=1
N2[Nµr+s+N(N−1)µrµs−N2µrµs−N2µsµr+N2µrµs]
=1
Nµr+s+N−1
Nµrµs−µrµs
=µr+s−µrµs
N.
We note that by setting r=s, we recover the expression (27.53) for V[mr].
J
27.4.5 Population central moments νr
We may generalise the discussion of estimators for the second central moment ν2
(or equivalently σ2) given in subsection 27.4.2 to the estimation of the rth central
moment νr. In particular, we saw in that subsection that our choice of estimator
forν2depended on whether the population mean µ1is known; the same is true
for the estimation of νr.
Let us first consider the case in which µ1is known. From (26.54), we may write
νras
νr=µr−rC1µr−1µ1+···+(−1)krCkµr−kµk
1+···+(−1)r−1(rCr−1−1)µr
1.
Ifµ1is known, a suitable estimator is obviously
ˆνr=mr−rC1mr−1µ1+···+(−1)krCkmr−kµk
1+···+(−1)r−1(rCr−1−1)µr
1,
where mris the rth sample moment. Since µ1and the binomial coefficients are
(known) constants, it is immediately clear that E[ˆνr]=νr,a n ds o ˆνris an unbiased
estimator of νr.I ti sa l s op o s s i b l et oo b t a i na ne x p r e s s i o nf o r V[ˆνr], though the
calculation is somewhat lengthy.
In the case where the population mean µ1isnotknown, the situation is more
complicated. We saw in subsection 27.4.2 that the second sample moment n2(or
s2)i snotan unbiased estimator of ν2(orσ2). Similarly, the rthcentral moment of
a sample, nr, is not an unbiased estimator of the rth population central moment
νr. However, in all cases the bias becomes negligible in the limit of large N.
1092
27.4 SOME BASIC ESTIMATORS
As we also found in the same subsection, it is rather complicated to calculate
the expectation and variance of n2; this complication increases considerably for
general r. Nevertheless, we have derived already in this chapter exact expressions
for the expectation value of the first few sample central moments, which are valid
for samples of any size N. From (27.40), (27.43) and (27.46), we find
E[n1]=0 ,
E[n2]=N−1
Nν2, (27.54)
E[n2
2]=N−1
N3[(N−1)ν4+(N2−2N+3 )ν2
2].
By similar arguments it can be shown that
E[n3]=(N−1)(N−2)
N2ν3, (27.55)
E[n4]=N−1
N3[(N2−3N+3 )ν4+3 ( 2 N−3)ν2
2]. (27.56)
From (27.54) and (27.55), we see that unbiased estimators of ν2andν3are
ˆν2=N
N−1n2, (27.57)
ˆν3=N2
(N−1)(N−2)n3, (27.58)
where (27.57) simply re-establishes our earlier result thathatwideσ2=Ns2/(N−1) is an
unbiased estimator of σ2.
Unfortunately, the pattern that appears to be emerging in (27.57) and (27.58)
isnotcontinued for higher r, as is seen immediately from (27.56). Nevertheless,
in the limit of large N, the bias becomes negligible, and often one simply takes
ˆνr=nr. For large N,i tm a yb es h o w nt h a t
E[nr]≈νr
V[nr]≈1
N(ν2r−ν2
r+r2ν2ν2
r−1−2rνr−1νr+1)
Cov[nr,ns]≈1
N(νr+s−νrνs+rsν2νr−1νs−1−rνr−1νs+1−sνs−1νr+1)
27.4.6 Population covariance Cov [x, y]and correlation Corr [x, y]
So far we have assumed that each of our Nindependent samples consists of
a single number xi. Let us now extend our discussion to a situation in which
each sample consists of two numbers xi,yi, which we may consider as being
drawn randomly from a two-dimensional population P(x, y). In particular, we
now consider estimators for the population covariance Cov[ x, y] and for the
correlation Corr[ x, y].
1093
STATISTICS
When µxandµyareknown , an appropriate estimator of the population covari-
ance is
hatwidestCov[x, y]=xy−µxµy=parenleftBigg
1
NNsummationdisplay
i=1xiyiparenrightBigg
−µxµy. (27.59)
This estimator is unbiased since
EbracketleftBig
hatwidestCov[x, y]bracketrightBig
=1
NEbracketleftBiggNsummationdisplay
i=1xiyibracketrightBigg
−µxµy=E[xiyi]−µxµy=C o v [ x, y].
Alternatively, if µxandµyareunknown , it is natural to replace µxandµyin
(27.59) by the sample means ¯xand¯yrespectively, in which case we recover the
sample covariance Vxy=xy−¯x¯ydiscussed in subsection 27.2.4. This estimator
is biased but an unbiased estimator of the population covariance is obtained byforming
hatwidestCov[x, y]=N
N−1Vxy. (27.60)ICalculate the expectation value of the sample covariance Vxyfor a sample of size N.
The sample covariance is given by
Vxy=
/
1
N
X
ixiyi
/!
−
/
1
N
X
ixi
/!/
1
N
X
jyj
/!
.
Thus its expectation value is given by
E[Vxy]=1
NE
/"X
ixiyi
/#
−1
N2E
/"/ X
ixi
/!/ X
jxj
/! /#
=E[xiyi]−1
N2E
/2/6/4
X
ixiyi+
X
i,j
j/negationslash=ixiyj
/3/7/5
Since the number of terms in the double sum on the RHS is N(N−1), we have
E[Vxy]=E[xiyi]−1
N2(NE[xiyi]+N(N−1)E[xiyj])
=E[xiyi]−1
N2(NE[xiyi]+N(N−1)E[xi]E[yj])
=E[xiyi]−1
N
/;
E[xiyi]+(N−1)µxµy
/
=N−1
NCov[x, y],
where we have used the fact that, since the samples are independent, E[xiyj]=E[xi]E[yj].
J
1094
27.4 SOME BASIC ESTIMATORS
It is possible to obtain expressions for the variances of the estimators (27.59)
and (27.60) but these quantities depend upon higher moments of the populationP(x, y) and are extremely lengthy to calculate.
Whether the means µ
xandµyare known or unknown, an estimator of the
population correlation Corr[ x, y]i sg i v e nb y
/hatwideCorr[ x, y]=hatwidestCov[x, y]
ˆσxˆσy, (27.61)
wherehatwidestCov[x, y],ˆσxandˆσyare the appropriate estimators of the population co-
variance and standard deviations. Although this estimator is only asymptoticallyunbiased, i.e. for large N, it is widely used because of its simplicity. Once again
the variance of the estimator depends on the higher moments of P(x, y)a n di s
difficult to calculate.
In the case in which the means µ
xandµyare unknown, a suitable (but biased)
estimator is
/hatwideCorr[ x, y]=N
N−1Vxy
sxsy=N
N−1rxy, (27.62)
where sxandsyare the sample standard deviations of the xiandyirespectively
andrxyis the sample correlation. In the special case when the parent population
P(x, y) is Gaussian, it may be shown that, if ρ= Corr[ x, y],
E[rxy]=ρ−ρ(1−ρ2)
2N+O ( N−2), (27.63)
V[rxy]=1
N(1−ρ2)2+O ( N−2), (27.64)
from which the expectation value and variance of the estimator /hatwideCorr[ x, y]m a y
be found immediately.
We note finally that our discussion may be extended, without significant al-
teration, to the general case in which each data item consists of nnumbers
xi,yi,...,z i.
27.4.7 A worked example
We conclude our discussion of basic estimators by reconsidering the set of
experimental data given in subsection 27.2.4.
1095
STATISTICSITen UK citizens are selected at random and their heights and weights are found to be as
follows (to the nearest cmorkgrespectively):
Person ABCDEFGH IJ
Height (cm) 194 168 177 180 171 190 151 169 175 182Weight (kg) 75 53 72 80 75 75 57 67 46 68
Estimate the means, µ
xandµy, and standard deviations, σxandσy, of the two-dimensional
joint population from which the sample was drawn, quoting the standard error on the esti-mate in each case. Estimate also the correlation Corr[ x, y]of the population, and quote the
standard error on the estimate under the assumption that the population is a multivariate
Gaussian.
In subsection 27.2.4, we calculated various sample statistics for these data. In particular,we found that for our sample of size N= 10,
¯x= 175 .7,¯y=6 6.8,
s
x=1 1.6,s y=1 0.6,r xy=0.54.
Let us begin by estimating the means µxandµy. As discussed in subsection 27.4.1, the
sample mean is an unbiased, consistent estimator of the population mean. Moreover, the
standard error on ¯x(say) is σx/√
N. In this case, however, we do not know the true value
ofσxand we must estimate it using
bσx=
p
N/(N−1)sx. Thus, our estimates of µxand
µy, with associated standard errors, are
ˆµx=¯x±sx√
N−1= 175 .7±3.9,
ˆµy=¯y±sy√
N−1=6 6.8±3.5.
We now turn to estimating σxandσy. As just mentioned, our estimate of σx(say)
is
bσx=
p
N/(N−1)sx. Its variance (see the final line of subsection 27.4.3) is given
approximately by
V[ˆσ]≈1
4Nν2
/
ν4−N−3
N−1ν2
2
/
.
Since we do not know the true values of the population central moments ν2andν4,w e
must use their estimated values in this expression. We may take ˆν2=
bσ2
x=(ˆσ)2,w h i c hw e
have already calculated. It still remains, however, to estimate ν4.A si m p l i e dn e a rt h ee n d
of subsection 27.4.5, it is acceptable to take ˆν4=n4. Thus for the xiandyivalues, we have
(ˆν4)x=1
NNX
i=1(xi−¯x)4= 53411 .6
(ˆν4)y=1
NNX
i=1(yi−¯y)4= 27732 .5
Substituting these values into (27.50), we obtain
ˆσx=
/N
N−1
/1/2
sx±(ˆV[ˆσx])1/2=1 2.2±6.7, (27.65)
ˆσy=
/N
N−1
/1/2
sy±(ˆV[ˆσy])1/2=1 1.2±3.6. (27.66)
1096
27.5 MAXIMUM-LIKELIHOOD METHOD
Finally, we estimate the population correlation Corr[ x, y], which we shall denote by ρ.
From (27.62), we have
ˆρ=N
N−1rxy=0.60.
Under the assumption that the sample was drawn from a two-dimensional Gaussian
population P(x, y), the variance of our estimator is given by (27.64). Since we do not know
the true value of ρ, we must use our estimate ˆρ. Thus, we find that the standard error ∆ ρ
in our estimate is given approximately by
∆ρ≈10
9
/1
10
/
[1−(0.60)2]2=0.05.
J
27.5 Maximum-likelihood method
The population from which the sample x1,x2,...,x Nis drawn is, in general,
unknown . In the previous section, we assumed that the sample values were inde-
pendent and drawn from a one-dimensional population P(x), and we considered
basic estimators of the moments and central moments of P(x). We did not,h o w -
ever, assume a particular functional form for P(x). We now discuss the process
ofdata modelling , in which a specific form is assumed for the population.
In the most general case, it will not be known whether the sample values are
independent, and so let us consider the full joint population P(x), where xis the
point in the N-dimensional data space with coordinates x1,x2,...,x N.W et h e n
adopt the hypothesis Hthat the probability distribution of the sample values has
some particular functional form L(x;a), dependent on the values of some set of
parameters ai,i=1,2,...,m . Thus, we have
P(x|a,H)=L(x;a),
where we make explicit the conditioning on both the assumed functional form and
on the parameter values; L(x;a) is called the likelihood function . Hypotheses of
this type form the basis of data modelling andparameter estimation . One proposes a
particular model for the underlying population and then attempts to estimate fromthe sample values x
1,x2,...,x Nthe values of the parameters adefining this model.IA company measures the duration (in minutes) of the Nintervals xi,i=1,2,...,N
between successive telephone calls received by its switchboard. Suppose that the samplevalues x
iare drawn independently from the distribution P(x|τ)=( 1 /τ)e x p (−x/τ),w h e r e τ
is the mean interval between calls. C alculate the likelihood function L(x;τ).
Since the sample values are independent and drawn from the stated distribution, the
likelihood is given by
L(x;τ)=P(xi|τ)P(x2|τ)···P(xN|τ)
=1
τexp
/
−x1
τ
/1
τexp
/
−x2
τ
/
···1
τexp
/
−xN
τ
/
=1
τNexp
/
−1
τ(x1+x2+···+xN)
/
. (27.67)
which is to be considered as a function of τ, given that the sample values xiare fixed.
J
1097
STATISTICS
0024 681 0 12 14 16 18 200.51
N=5L(x;τ)
τ0024 681 0 12 14 16 18 200.51
N=1 0L(x;τ)
τ
0024 681 0 12 14 16 18 200.51
N=2 0L(x;τ)
τ0024 681 0 12 14 16 18 200.51
N=5 0L(x;τ)
τ
Figure 27.5 Examples of the likelihood function (27.67) for samples of dif-
ferent size N. In each case, the true value of the parameter is τ=4a n dt h e
sample values xiare indicated by the short vertical lines. For the purposes
of illustration, in each case the likelihood function is normalised so that its
maximum value is unity.
The likelihood function (27.67) depends on just a single parameter τ.P l o t so f
the likelihood function, considered as a function of τ, are shown in figure 27.5 for
samples of different size N. The true value of the parameter τused to generate the
sample values was 4. In each case, the sample values xiare indicated by the short
vertical lines. For the purposes of illustration, the likelihood function in each
case has been scaled so that its maximum value is unity (this is, in fact, commonpractice). We see that when the sample size is small, the likelihood function is verybroad. As Nincreases, however, the likelihood becomes narrower (it is inversely
proportional to√
N) and tends to a Gaussian-like shape, with its peak centred
on 4, the true value of τ. We discuss these properties of the likelihood function
in more detail in subsection 27.5.6.
27.5.1 The maximum-likelihood estimator
Since the likelihood function L(x;a) gives the probability density associated with
any particular set of values of the parameters a, our best estimate ˆaof these
parameters is given by the values of afor which L(x;a) is a maximum. This is
called the maximum-likelihood estimator (or ML estimator).
In general, the likelihood function can have a complicated shape when con-
1098
27.5 MAXIMUM-LIKELIHOOD METHOD
a aa a
L(x;a) L(x;a)L(x;a) L(x;a)
ˆa ˆaˆa ˆa(a)( b)
(c)( d)
Figure 27.6 Typical shapes of one-dimensional likelihood functions L(x;a)
encountered in practice, where, for illustration purposes, it is assumed that
the parameter ais restricted to the range zero to infinity. The ML estimator
in the various cases occurs at: ( a) the only stationary point; ( b) one of several
stationary points; ( c) an end-point of the allowed parameter range that is not
a stationary point (although stationary points do exist); ( d) an end-point of
the allowed parameter range in whic h no stationary point exists.
sidered as a function of a, particularly when the dimensionality of the space of
parameters a1,a2,...,a Mis large. It may be that the values of some parameters
are either known or assumed in advance, in which case the effective dimension-ality of the likelihood function is reduced accordingly. However, even when thelikelihood depends on just a single parameter a(either intrinsically or as the
result of assuming particular values for the remaining parameters), its form maybe complicated when the sample size Nis small. Frequently occurring shapes of
one-dimensional likelihood functions are illustrated in figure 27.6, where we have
assumed, for definiteness, that the allowed range of the parameter ais zero to
infinity. In each case, the ML estimate ˆais also indicated. Of course, the ‘shape’ of
higher-dimensional likelihood functions may be considerably more complicated.
In many simple cases, however, the likelihood function L(x;a) has a single
maximum that occurs at a stationary point (the likelihood function is then termedunimodal ). In this case, the ML estimators of the parameters a
i,i=1,2,...,M ,
may be found without evaluating the full likelihood function L(x;a). Instead, one
simply solves the Msimultaneous equations
∂L
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=ˆa=0 f o r i=1,2,...,M. (27.68)
1099
STATISTICS
Since ln zis a monotonically increasing function of z(and therefore has the
same stationary points), it is often more convenient, in fact, to maximise thelog-likelihood function ,l nL(x;a), with respect to the a
i. Thus, one may, as an
alternative, solve the equations
∂lnL
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=ˆa=0 f o r i=1,2,...,M. (27.69)
Clearly, (27.68) and (27.69) will lead to the same ML estimates ˆaof the parameters.
In either case, it is, of course, prudent to check that the point a=ˆais a local
maximum.IFind the ML estimate of the parameter τin the previous example, in terms of the measured
values xi,i=1,2,...,N .
From (27.67), the log-likelihood function in this case is given by
lnL(x;τ)=NX
i=1ln
/1
τe−xi/τ
/
=−NX
i=1
/
lnτ+xi
τ
/
. (27.70)
Differentiating with respect to the parameter τand setting the result equal to zero, we find
∂lnL
∂τ=−NX
i=1
/1
τ−xi
τ2
/
=0.
Thus the ML estimate of the parameter τis given by
ˆτ=1
NNX
i=1xi, (27.71)
which is simply the sample mean of the Nmeasured intervals.
J
In the previous example we assumed that the sample values xiwere drawn
independently from the sameparent distribution. The ML method is more flexible
than this restriction might seem to imply and it can equally well be applied to thecommon case in which the samples x
iare independent but each is drawn from a
different distribution.IIn an experiment, Nindependent measurements xiof some quantity are made. Suppose
that the random measurement error on the ith sample value is Gaussian distributed with
mean zero and known standard deviation σi. Calculate the ML estimate of the true value
µof the quantity being measured.
As the measurements are independent, the likelihood factorises:
L(x;µ,{σk})=NY
i=1P(xi|µ, σ i),
where{σk}denotes collectively the set of known standard deviations σ1,σ2,...,σ N.T h e
individual distributions are given by
P(xi|µ, σ i)=1p
2πσ2
iexp
/
−(xi−µ)2
2σ2
i
/
.
1100
27.5 MAXIMUM-LIKELIHOOD METHOD
and so the full log-likelihood function is given by
lnL(x;µ,{σk})=−1
2NX
i=1
/
ln(2πσ2
i)+(xi−µ)2
σ2
i
/
.
Differentiating this expression with respect to µand setting the result equal to zero, we
find
∂lnL
∂µ=NX
i=1xi−µ
σ2
i=0,
from which we obtain the ML estimator
ˆµ=
PN
i=1(xi/σ2
i)PN
i=1(1/σ2
i). (27.72)
This estimator is commonly used when averaging data with different statistical weights
wi=1/σ2
i. We note that when all the variances σ2
ihave the same value the estimator
reduces to the sample mean of the data xi.
J
There is, in fact, no requirement in the ML method that the sample values
be independent. As an illustration, we shall generalise the above example to a
case in which the measurements xiare not all independent. This would occur, for
example, if these measurements were based at least in part on the same data.IIn an experiment Nmeasurements xiof some quantity are made. Suppose that the random
measurement errors on the samples are drawn from a joint Gaussian distribution with meanzero and known covariance matrix V. Calculate the ML estimate of the true value µof the
quantity being measured.
From (26.148), the likelihood in this case is given by
L(x;µ,V)=1
(2π)N/2|V|1/2exp
/
−1
2(x−µ1)TV−1(x−µ1)
/
,
where xis the column vector with components x1,x2,...,x Nand 1is the column vector
with all components equal to unity. Thus, the log-likelihood function is given by
lnL(x;µ,V)=−1
2
/
Nln(2π)+l n|V|+(x−µ1)TV−1(x−µ1)
/
.
Differentiating with respect to µand setting the result equal to zero gives
∂lnL
∂µ=1TV−1(x−µ1)=0 .
Thus, the ML estimator is given by
ˆµ=1TV−1x
1TV−11=
P
i,j(V−1)ijxjP
i,j(V−1)ij.
In the case of uncorrelated errors in measurement, ( V−1)ij=δij/σ2
iand our estimator
reduces to that given in (27.72).
J
In all the examples considered so far, the likelihood function has been effectively
one-dimensional, either instrinsically or under the assumption that the values ofall but one of the parameters are known in advance. As the following example
1101
STATISTICS
involving two parameters shows, the application of the ML method to the
estimation of several parameters simultaneously is straightforward.IIn an experiment Nmeasurements xiof some quantity are made. Suppose the random error
on each sample value is drawn independently from a Gaussian distribution of mean zero butunknown standard deviation σ(which is the same for each measurement). Calculate the ML
estimates of the true value µof the quantity being measured and the standard deviation σ
of the random errors.
In this case the log-likelihood function is given by
lnL(x;µ, σ)=−1
2NX
i=1
/
ln(2πσ2)+(xi−µ)2
σ2
/
.
Taking partial derivatives of ln Lwith respect to µandσand setting the results equal to
zero at the joint estimate ˆµ,ˆσ,w eo b t a i n
NX
i=1xi−ˆµ
ˆσ2=0, (27.73)
NX
i=1(xi−ˆµ)2
ˆσ3−NX
i=11
ˆσ=0. (27.74)
In principle, one should solve these two equations simultaneously for ˆµandˆσ, but in this
case we notice that the first is solved immediately by
ˆµ=1
NNX
i=1xi=¯x,
where ¯xis the sample mean. Substituting this result into the second equation, we find
ˆσ=
vuut1
NNX
i=1(xi−¯x)2=s,
where sis the sample standard deviation. As shown in subsection 27.4.3, sis a biased
estimator of σ. The reason why the ML method may produce a biased estimator is
discussed in the next subsection.
J
27.5.2 Transformation invariance and bias of ML estimators
An extremely useful property of ML estimators is that they are invariant to
parameter transformations. Suppose that, instead of estimating some parameter
aof the assumed population, we wish to estimate some function α(a)o ft h e
parameter. The ML estimator ˆα(a) is given by the value assumed by the function
α(a) at the maximum point of the likelihood, which is simply equal to α(ˆa). Thus,
we have the very convenient property
ˆα(a)=α(ˆa).
We do not have to worry about the distinction between the two cases, estimating
aand estimating a function of a.T h i si s nottrue, in general, for other estimation
procedures.
1102
27.5 MAXIMUM-LIKELIHOOD METHODIA company measures the duration (in minutes) of the Nintervals xi,i=1,2,...,N ,
between successive telephone calls received by its switchboard. Suppose that the samplevalues x
iare drawn independently from the distribution P(x|τ)=( 1 /τ)exp(−x/τ).F i n dt h e
ML estimate of the parameter λ=1/τ.
This is the same problem as that considered at the start of section 27.5.1. In terms of the
new parameter λ, the log-likelihood function is given by
lnL(x;λ)=NX
i=1ln(λe−λxi)=NX
i=1(lnλ−λxi).
Differentiating with respect to λand setting the result equal to zero, we have
∂lnL
∂λ=NX
i=1
/1
λ−xi
/
=0.
Thus, the ML estimator of the parameter λis given by
ˆλ=
/
1
NNX
i=1xi
/!−1
=¯x−1. (27.75)
Referring back to (27.71), we see that, as expected, the ML estimators of λandτare
related by ˆλ=1/ˆτ.
J
Although this invariance property is useful it also means that, in general, ML
estimators may be biased . In particular, one must be aware of the fact that even
ifˆais an unbiased ML estimator of ait does notfollow that the estimator ˆα(a)i s
also unbiased. In the limit of large N, however, the bias of ML estimators always
tends to zero. As an illustration, it is straightforward to show (see exercise 27.8)that the ML estimators ˆτandˆλin the above example have expectation values
E[ˆτ]=τand E[ˆλ]=N
N−1λ. (27.76)
In fact, since ˆτ=¯xand the sample values are independent, the first result follows
immediately from (27.40). Thus, ˆτis unbiased, but ˆλ=1/ˆτis biased, albeit that
the bias tends to zero for large N.
27.5.3 Efficiency of ML estimators
We showed in subsection 27.3.2 that Fisher’s inequality puts a lower limit on the
variance V[ˆa] of any estimator of the parameter a. Under our hypothesis Hon
p. 1097, the functional form of the population is given by the likelihood function,
i.e.P(x|a,H)=L(x;a). Thus, if this hypothesis is correct, we may replace Pby
Lin Fisher’s inequality (27.18), which then reads
V[ˆa]≥parenleftbigg
1+∂b
∂aparenrightbigg2slashbigg
Ebracketleftbigg
−∂2lnL
∂a2bracketrightbigg
,
where bis the bias in the estimator ˆa. We usually denote the RHS by Vmin.
1103
STATISTICS
An important property of ML estimators is that ifthere exists an efficient
estimator ˆaeff,i . e .o n ef o rw h i c h V[ˆaeff]=Vmin,t h e ni t mustbe the ML estimator
or some function thereof. This is easily shown by replacing PbyLin the proof
of Fisher’s inequality given in subsection 27.3.2. In particular, we note that the
equality in (27.22) holds only if h(x)=cg(x), where cis a constant. Thus, if an
efficient estimator ˆaeffexists, this is equivalent to demanding that
∂lnL
∂a=c[ˆaeff−α(a)].
Now, the ML estimator ˆaMLis given by
∂lnL
∂avextendsinglevextendsinglevextendsinglevextendsingle
a=ˆaML=0⇒ c[ˆaeff−α(ˆaML)] = 0 ,
which, in turn, implies that ˆaeffmust be some function of ˆaML.IShow that the ML estimator ˆτgiven in (27.71) is an efficient estimator of the parameter τ.
As shown in (27.70), the log-likelihood function in this case is
lnL(x;τ)=−NX
i=1
/
lnτ+xi
τ
/
.
Differentiating twice with respect to τ, we find
∂2lnL
∂τ2=NX
i=1
/1
τ2−2xi
τ3
/
=N
τ2
/
1−2
τNNX
i=1xi
/!
, (27.77)
and so the expectation value of this expression is
E
/∂2lnL
∂τ2
/
=N
τ2
/
1−2
τE[xi]
/
=−N
τ2,
where we have used the fact that E[x]=τ. Setting b= 0 in (27.18), we thus find that for
anyunbiased estimator of τ,
V[ˆτ]≥τ2
N.
From (27.76), we see that the ML estimator ˆτ=
P
ixi/Nis unbiased. Moreover, using
the fact that V[x]=τ2, it follows immediately from (27.40) that V[ˆτ]=τ2/N. Thus ˆτis a
minimum-variance estimator of τ.
J
27.5.4 Standard errors and confidence limits on ML estimators
The ML method provides a procedure for obtaining a particular set of estimators
ˆaMLfor the parameters aof the assumed population P(x|a). As for any other set
of estimators, the associated standard errors, covariances and confidence intervalscan be found as described in subsections 27.3.3 and 27.3.4.
1104
27.5 MAXIMUM-LIKELIHOOD METHOD
0
00.10.20.30.4
24 681 0 12 14P(ˆτ|τ)
ˆτ
Figure 27.7 The sampling distribution P(ˆτ|τ) for the estimator ˆτfor the case
τ=4a n d N= 10.IA company measures the duration (in minutes) of the 10intervals xi,i=1,2,...,10,
between successive telephone calls made to its switchboard to be as follows:
0.43 0 .24 3 .03 1 .93 1 .16 8 .65 5 .33 6 .06 5 .62 5 .22.
Supposing that the sample value s are drawn independently from the probability distribution
P(x|τ)=( 1 /τ)e x p (−x/τ), find the ML estimate of the mean τand quote an estimate of
the standard error on your result.
As shown in (27.71) the (unbiased) ML estimator ˆτin this case is simply the sample mean
¯x=3.77. Also, as shown in subsection 27.5.3, ˆτis a minimum-variance estimator with
V[ˆτ]=τ2/N. Thus, the standard error in ˆτis simply
σˆτ=τ√
N. (27.78)
Since we do not know the true value of τ, however, we must instead quote an estimate ˆσˆτ
of the standard error, obtained by substituting our estimate ˆτforτin (27.78). Thus, we
quote our final result as
τ=ˆτ±ˆτ√
N=3.77±1.19. (27.79)
For comparison, the true value used to create the sample was τ=4 .
J
For the particular problem considered in the above example, it is in fact possible
to derive the full sampling distribution of the ML estimator ˆτusing characteristic
functions, and it is given by
P(ˆτ|τ)=NN
(N−1)!ˆτN−1
τNexpparenleftbigg
−Nˆτ
τparenrightbigg
, (27.80)
where Nis the size of the sample. This function is plotted in figure 27.7 for the
case τ=4a n d N= 10, which pertains to the above example. Knowledge of the
analytic form of the sampling distribution allows one to place confidence limits
on the estimate ˆτobtained, as discussed in subsection 27.3.4.
1105
STATISTICSIUsing the sample values in the above example, obtain the 68% central confidence interval
on the value of τ.
For the sample values given, our observed value of the ML estimator is ˆτobs=3.77. Thus,
from (27.28) and (27.29), the 68% central confidence interval [ τ−,τ+] on the value of τis
found by solving the equationsZˆτobs
−∞P(ˆτ|τ+)dˆτ=0.16,Z∞
ˆτobsP(ˆτ|τ−)dˆτ=0.16,
where P(ˆτ|τ) is given by (27.80) with N= 10. The above integrals can be evaluated
analytically but the calculations are rather cumbersome. It is much simpler to evaluatethem by numerical integration, from which we find [ τ
−,τ+]=[ 2 .86,5.46]. Alternatively,
we could quote the estimate and its 68% confidence interval as
τ=3.77+1.69
−0.91.
Thus we see that the 68% central confidence interval is not symmetric about the estimated
value, and differs from the standard error calculated above. This is a result of the (non-Gaussian) shape of the sampling distribution P(ˆτ|τ), apparent in figure 27.7.J
In many problems, however, it is not possible to derive the full sampling
distribution of an ML estimator ˆain order to obtain its confidence intervals.
Indeed, one may not even be able to obtain an analytic formula for its standarderror σ
ˆa. This is particularly true when one is estimating several parameter ˆa
simultaneously, since the joint sampling distribution will be, in general, very
complicated. Nevertheless, as we discuss below, the likelihood function L(x;a)
itselfcan be used very simply to obtain standard errors and confidence intervals.
The justification for this has its roots in the Bayesian approach to statistics, as
opposed to the more traditional frequentist approach we have adopted here. We
now give a brief discussion of the Bayesian viewpoint on parameter estimation.
27.5.5 The Bayesian interpretation of the likelihood function
As stated at the beginning of section 27.5, the likelihood function L(x;a)i s
defined by
P(x|a,H)=L(x;a),
where Hdenotes our hypothesis of an assumed functional form. Now, using
Bayes’ theorem (see subsection 26.2.3), we may write
P(a|x,H)=P(x|a,H)P(a|H)
P(x|H), (27.81)
which provides us with an expression for the probability distribution P(a|x,H)
of the parameters a, given the (fixed) data xand our hypothesis H, in terms of
1106
27.5 MAXIMUM-LIKELIHOOD METHOD
other quantities that we may assign. The various terms in (27.81) have special
formal names, as follows.
•The quantity P(a|H)o nt h eR H Si st h e priorprobability, which represents our
state of knowledge of the parameter values (given the hypothesis H)before we
have analysed the data.
•This probability is modified by the experimental data xthrough the likelihood
P(x|a,H).
•When appropriately normalised by the evidence P(x|H), this yields the posterior
probability P(a|x,H), which is the quantity of interest.
•The posterior encodes allour inferences about the values of the parameters a.
Strictly speaking, from a Bayesian viewpoint, this entire function ,P(a|x,H), is
the ‘answer’ to a parameter estimation problem.
Given a particular hypothesis, the (normalising) evidence factor P(x|H)i s
unimportant, since it does not depend explicitly upon the parameter values a.
Thus, it is often omitted and one considers only the proportionality relation
P(a|x,H)∝P(x|a,H)P(a|H). (27.82)
If necessary, the posterior distribution can be normalised empirically, by requiring
that it integrates to unity, i.e.integraltext
P(a|x,H)dma= 1, where the integral extends over
all values of the parameters a1,a2,...,a m.
The prior P(a|H) in (27.82) should reflect our entire knowledge concerning the
values of the parameters a,before the analysis of the current data x. For example,
there may be some physical reason to require some or all of the parameters to
lie in a given range. If we are largely ignorant of the values of the parameters,we often indicate this by choosing a uniform (or very broad) prior,
P(a|H) = constant ,
in which case the posterior distribution is simply proportional to the likelihood.
In this case, we thus have
P(a|x,H)∝L(x;a). (27.83)
In other words, if we assume a uniform prior then we can identify the posterior
distribution (up to a normalising factor) with L(x;a), considered as a function of
the parameters a.
Thus, a Bayesian statistician considers the ML estimates ˆa
MLof the parameters
to be the values that maximise the posterior P(a|x,H) under the assumption of
a uniform prior. More importantly, however, a Bayesian would notcalculate the
standard error or confidence interval on this estimate using the (classical) methodemployed in subsection 27.3.4. Instead, a far more straightforward approach is
1107
STATISTICS
adopted. Let us assume, for the moment, that one is estimating just a single
parameter a. Using (27.83), we may determine the values a−anda+such that
Pr(a<a−|x,H)=integraldisplaya−
∞L(x;a)da=α,
Pr(a>a +|x,H)=integraldisplay∞
a+L(x;a)da=β.
where it is assumed that the likelihood has been normalised in such a way thatintegraltext
L(x;a)da= 1. Combining these equations gives
Pr(a−≤a<a +|x,H)=integraldisplaya+
a−L(x;a)da=1−α−β, (27.84)
and [ a−,a+]i st h e Bayesian confidence interval on the value of aat the confidence
level 1−α−β. As in the case of classical confidence intervals, one often quotes
the central confidence interval, for which α=β. Another common choice (where
possible) is to use the two values a−anda+satisfying (27.84), for which L(x;a−)=
L(x;a+).
It should be understood that a frequentist would consider the Bayesian confi-
dence interval as an approximation to the (classical) confidence interval discussed
in subsection 27.3.4. Conversely, a Bayesian would consider the confidence inter-val defined in (27.84) to be the more meaningful. In fact, the difference betweenthe Bayesian and classical confidence intervals is rather subtle. The classical con-fidence interval is defined in such a way that if one took a large number of
samples each of size Nand constructed the confidence interval in each case then
the proportion of cases in which the true value of awould be contained within the
interval is 1 −α−β. For the Bayesian confidence interval, one does not rely on the
frequentist concept of a large number of repeated samples. Instead, its meaning isthat, given the single sample x(and our hypothesis Hfor the functional form of
the population), the probability that alies within the interval [ a
−,a+]i s1−α−β.
By adopting the Bayesian viewpoint, the likelihood function L(x;a) may also
be used to obtain an approximation ˆσˆato the standard error in the ML estimator;
the approximation is given by
ˆσˆa=parenleftbigg
−∂2lnL
∂a2vextendsinglevextendsinglevextendsinglevextendsingle
a=ˆaparenrightbigg−1/2
. (27.85)
Clearly, if L(x;a) were a Gaussian centred on a=ˆathenˆσˆawould be its standard
deviation. Indeed, in this case, the resulting ‘one-sigma’ limits would constitute a
68.3% Bayesian central confidence interval. Even when L(x;a) is not Gaussian,
however, (27.85) is often used as a measure of the standard error.
1108
27.5 MAXIMUM-LIKELIHOOD METHOD
0
00.10.20.30.4
24 681 0 12 14L(x;τ)
τ
Figure 27.8 The likelihood function L(x;τ) (normalised to unit area) for the
sample values given in the worked example in subsection 27.5.4, and indicated
here by short vertical lines.IFor the sample data given in section 27.5.4, use the likelihood function to estimate the
standard error ˆσˆτin the ML estimator ˆτand obtain the Bayesian 68% central confidence
interval on τ.
We showed in (27.67) that the likelihood function in this case is given by
L(x;τ)=1
τNexp[−1
τ(x1+x2+···+xN)].
where xi,i=1,2,...,N , denotes the sample value and N= 10. This likelihood function is
plotted in figure 27.8, after normalising (numerically) to unit area. The short vertical linesin the figure indicate the sample values. We see that the likelihood function peaks at theML estimate ˆτ=3.77 that we found in subsection 27.5.4. Also, from (27.77), we have
∂
2lnL
∂τ2=N
τ2
/
1−2
τNNX
i=1xi
/!
,
Remembering that ˆτ=
P
ixi/N, our estimate of the standard error in ˆτis
ˆσˆτ=
/
−∂2lnL
∂τ2
////
τ=ˆτ
/−1/2
=ˆτ√
N=1.19,
which is precisely the estimate of the standard error we obtained in subsection 27.5.4.
It should be noted, however, that in general we would not expect the two estimates ofstandard error made by the different methods to be identical.
In order to calculate the Bayesian 68% central confidence interval, we must determine
the values a
−anda+that satisfy (27.84) with α=β=0.16. In this case, the calculation
can be performed analytically but is somewhat tedious. It is trivial, however, to determine
a−anda+numerically and we find the confidence interval to be [3 .16,6.20]. Thus we can
quote our result with 68% central confidence limits as
τ=3.77+2.43
−0.61.
By comparing this result with that given towards the end of subsection 27.5.4, we see that,
as we might expect, the Bayesian and classical confidence intervals differ somewhat.
J
1109
STATISTICS
The above discussion is generalised straightforwardly to the estimation of
several parameters a1,a2,...,a Msimultaneously. The elements of the covariance
matrix of the ML estimators can be approximated by
ˆVij=hatwidestCov[ˆai,ˆaj]=parenleftbigg
−∂2lnL
∂ai∂ajvextendsinglevextendsinglevextendsinglevextendsingle
a=ˆaparenrightbigg−1
. (27.86)
From (27.36), we see that (at least for unbiased estimators) the expectation value
of (27.86) is equal to the element Fijof the Fisher matrix.
The construction of a multi-dimensional Bayesian confidence region is also
straightforward. For a given confidence level 1 −α(say), it is most common
to construct the confidence region as the M-dimensional region Rina-space,
bounded by the ‘surface’ L(x;a) = constant, for which
integraldisplay
RL(x;a)dMa=1−α,
where it is assumed that L(x;a) is normalised to unit volume. Moreover, we
see from (27.83) that (assuming a uniform prior probability) we may obtain themarginal posterior distribution for any parameter a
isimply by integrating the
likelihood function L(x;a) over the other parameters:
P(ai|x,H)=integraldisplay
···integraldisplay
L(x;a)da1···dai−1dai+1···daM.
Here the integral extends over all possible values of the parameters, and again
is it assumed that the likelihood function is normalised in such a way thatintegraltext
L(x;a)dMa= 1. This marginal distribution can then be used as above to
determine Bayesian confidence intervals on each aiseparately.ITen independent sample values xi,i=1,2,...,10, are drawn at random from a Gaussian
distribution with unknown mean µand standard deviation σ. The samples values are as
follows (to two decimal places):
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Find the Bayesian 95% central confidence intervals on µandσseparately.
The likelihood function in this case is
L(x;µ, σ)=( 2 πσ2)−N/2exp
/"
−1
2σ2NX
i=1(xi−µ)2
/#
. (27.87)
Assuming uniform priors on µandσ(over their natural ranges of −∞ → ∞ and 0→∞
respectively), we may identify this likelihood function with the posterior probability, as in(27.83). Thus, the marginal posterior distribution on µis given by
P(µ|x,H)∝
Z∞
01
σNexp
/"
−1
2σ2NX
i=1(xi−µ)2
/#
dσ.
1110
27.5 MAXIMUM-LIKELIHOOD METHOD
By substituting σ=1/u(so that dσ=−du/u2) and integrating by parts either ( N−2)/2
or (N−3)/2 times, we find
P(µ|x,H)∝
/
N(¯x−µ)2+Ns2
/−(N−1)/2,
where we have used the fact that
P
i(xi−µ)2=N(¯x−µ)2+Ns2,¯xbeing the sample
mean and s2the sample variance. We may now find the 95% central confidence interval
by finding the values µ−andµ+for whichZµ−
−∞P(µ|x,H)dµ=0.025 and
Z∞
µ+P(µ|x,H)dµ=0.025.
The normalisation of the posterior distribution and the values µ−and µ+are easily
obtained by numerical integration. Substituting in the appropriate values N= 10, ¯x=1.11
ands=1.01, we find the required confidence interval to be [0 .29,1.97].
To obtain a confidence interval on σ, we must first obtain the corresponding marginal
posterior distribution. From (27.87), again using the fact that
P
i(xi−µ)2=N(¯x−µ)2+Ns2,
this is given by
P(σ|x,H)∝1
σNexp
/
−Ns2
2σ2
/Z∞
−∞exp
/
−N(¯x−µ)2
2σ2
/
dµ.
Noting that the integral of a one-dimensional Gaussian is proportional to σ, we conclude
that
P(σ|x,H)∝1
σN−1exp
/
−Ns2
2σ2
/
.
The 95% central confidence interval on σcan then be found in an analogous manner to
that on µ, by solving numerically the equationsZσ−
0P(σ|x,H)dσ=0.025 and
Z∞
σ+P(σ|x,H)dσ=0.025.
We find the required interval to be [0 .76,2.16].
J
27.5.6 Behaviour of ML estimators for large N
As mentioned in subsection 27.3.6, in the large-sample limit N→∞, the sampling
distribution of a set of (consistent) estimators ˆa, whether ML or not, will tend,
in general, to a multivariate Gaussian centred on the true values a.T h i si sa
direct consequence of the central limit theorem. Similarly, in the limit N→∞the
likelihood function L(x;a)alsotends towards a multivariate Gaussian but one
centred on the ML estimate(s) ˆa. Thus ML estimators are always asymptotically
consistent . This limiting process was illustrated for the one-dimensional case by
figure 27.5.
Thus, as Nbecomes large, the likelihood function tends to the form
L(x;a)=Lmaxexpbracketleftbig
−1
2Q(a,ˆa)bracketrightbig
,
where Qdenotes the quadratic form
Q(a,ˆa)=(a−ˆa)TV−1(a−ˆa)
1111
STATISTICS
and the matrix V−1is given by
parenleftbig
V−1parenrightbig
ij=−∂2lnL
∂ai∂ajvextendsinglevextendsinglevextendsinglevextendsingle
a=ˆa.
Moreover, in the limit of large N, this matrix tends to the Fisher matrix given in
(27.36), i.e. V−1→F. Hence ML estimators are asymptotically minimum-variance .
Comparison of the above results with those in subsection 27.3.6 shows that
the large-sample limit of the likelihood function L(x;a) has the same form as the
large-sample limit of the joint estimator sampling distribution P(ˆa|a). The only
difference is that P(ˆa|a) is centred in ˆa-space on the true values ˆa=awhereas
L(x;a) is centred in a-space on the ML estimates a=ˆa. From figure 27.4 and its
accompanying discussion, we therefore conclude that, in the large-sample limit,
the Bayesian and classical confidence limits on the parameters coincide .
27.5.7 Extended maximum-likelihood method
It is sometimes the case that the number of data items Nin our sample is itself a
random variable. Such experiments are typically those in which data are collectedfor a certain period of time during which events occur at random in some way,as opposed to those in which a prearranged number of data items are collected.In particular, let us consider the case where the sample values x
1,x2,...,x Nare
drawn independently from some distribution P(x|a) and the sample size Nis a
random variable described by a Poisson distribution with mean λ,i . e .N∼Po(λ).
The likelihood function in this case is given by
L(x;λ,a)=λN
N!e−λNproductdisplay
i=1P(xi|a), (27.88)
a n di so f t e nc a l l e dt h e extended likelihood function . The function L(x;λ,a)c a n
be used as before to estimate parameter values or obtain confidence intervals.Two distinct cases arise in the use of the extended likelihood function, dependingon whether the Poisson parameter λis a function of the parameters aor is an
independent parameter.
Let us first consider the case in which λis a function of the parameters a.F r o m
(27.88), we can write the extended log-likelihood function as
lnL=Nlnλ(a)−λ(a)+
Nsummationdisplay
i=1lnP(xi|a)=−λ(a)+Nsummationdisplay
i=1ln[λ(a)P(xi|a)].
where we have ignored terms not depending on a. The ML estimates ˆaof the
parameters can then be found in the usual way, and the ML estimate of the
Poisson parameter is simply ˆλ=λ(ˆa). The errors on our estimators ˆawill be, in
general, smaller than those obtained in the usual likelihood approach, since ourestimate includes information from the value of Nas well as the sample values x
i.
1112
27.6 THE METHOD OF LEAST SQUARES
The other possibility is that λis an independent parameter and not a function
of the parameters a. In this case, the extended log-likelihood function is
lnL=Nlnλ−λ+Nsummationdisplay
i=1lnP(xi|a), (27.89)
where we have omitted terms not depending on λora. Differentiating with
respect to λand setting the result equal to zero, we find that the ML estimate of
λis simply
ˆλ=N.
By differentiating (27.89) with respect to the parameters aiand setting the results
equal to zero, we obtain the usual ML estimates ˆaiof their values. In this case,
however, the errors in our estimates will be larger, in general, than those in thestandard likelihood approach, since they must include the effect of statisticaluncertainty in the parameter λ.
27.6 The method of least squares
The method of least squares is, in fact, just a special case of the method of
maximum-likelihood. Nevertheless, it is so widely used as a method of parameterestimation that it has acquired a special name of its own. At the outset, let ussuppose that a data sample consists of a set of pairs ( x
i,yi),i=1,2,...,N .F o r
example, these data might correspond to the temperature yimeasured at various
points xialong some metal rod.
For the moment, we will suppose that the xiare known exactly, whereas there
exists a measurement error (or noise)nion each of the values yi. Moreover, let
us assume that the true value of yat any position xis given by some function
y=f(x;a) that depends on the Munknown parameters a.T h e n
yi=f(xi;a)+ni.
Our aim is to estimate the values of the parameters afrom the data sample.
Bearing in mind the central limit theorem, let us suppose that the niare
drawn from a Gaussian distribution with zero mean and no systematic bias. In
the most general case the measurement errors nimight notbe independent but
described by an N-dimensional multivariate Gaussian with non-trivial covariance
matrix N, whose elements Nij=C o v [ ni,nj] we assume to be known. Under these
assumptions it follows from (26.148), that the likelihood function is
L(x,y;a)=1
(2π)N/2|N|1/2expbracketleftbig
−1
2χ2(a)bracketrightbig
,
1113
STATISTICS
where the quantity denoted by χ2is given by the quadratic form
χ2(a)=Nsummationdisplay
i,j=1[yi−f(xi;a)](N−1)ij[yj−f(xj;a)] = ( y−f)TN−1(y−f).
(27.90)
In the last equality, we have rewritten the expression in matrix notation by
defining the column vector fwith elements fi=f(xi;a). We note that in the
(common) special case in which the measurement errors niareindependent ,t h e i r
covariance matrix takes the diagonal form N=d i a g ( σ2
1,σ2
2,...,σ2
N), where σiis
the standard deviation of the measurement error ni. In this case, the expression
(27.90) for χ2reduces to
χ2(a)=Nsummationdisplay
i=1bracketleftbiggyi−f(xi;a)
σibracketrightbigg2
.
The least squares (LS) estimators ˆaLSof the parameter values are defined as
those that minimise the value of χ2(a); they are usually determined by solving
theMequations
∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=ˆaLS=0 f o r i=1,2,...,M. (27.91)
Clearly, if the measurement errors niare indeed Gaussian distributed, as assumed
above, then the LS and ML estimators of the parameters acoincide. Because
of its relative simplicity, the method of least squares is often applied to cases in
which the niare not Gaussian distributed. The resulting estimators ˆaLSarenotthe
ML estimators, and the best that can be said in justification is that the method isan obviously sensible procedure for parameter estimation that has stood the testof time.
Finally, we note that the method of least squares is easily extended to the case
in which each measurement y
idepends on several variables, which we denote
byxi. For example, yimight represent the temperature measured at the (three-
dimensional) position xiin a room. In this case, the data is modelled by a
function y=f(xi;a), and the remainder of the above discussion carries through
unchanged.
27.6.1 Linear least squares
We have so far made no restriction on the form of the function f(x;a). It so
happens, however, that, for a model in which f(x;a)i sa linear function of the
parameters a1,a2,...,a M, one can always obtain analytic expressions for the LS
estimators ˆaLSand their variances. The general form of this kind of model is
f(x;a)=Msummationdisplay
i=1aihi(x), (27.92)
1114
27.6 THE METHOD OF LEAST SQUARES
where h1(x),h2(x),...,h M(x) are some set of linearly independent fixed functions
ofx,o f t e nc a l l e dt h e basis functions . Note that the functions hi(x) themselves may
be highly non-linear functions of x. The ‘linear’ nature of the model (27.92) refers
only to its dependence on the parameters ai. Furthermore, in this case, it may
be shown that the LS estimators ˆaihave zero bias and are minimum-variance,
irrespective of the probability density function from which the measurement errorsn
iare drawn.
In order to obtain analytic expressions for the LS estimators ˆaLS,i ti sc o n v e n i e n t
to write (27.92) in the form
f(x;a)=Msummationdisplay
j=1Rijaj, (27.93)
where Rij=hj(xi) is an element of the response matrix Rof the experiment. The
expression for χ2given in (27.90) can then be written, in matrix notation, as
χ2(a)=( y−Ra)TN−1(y−Ra). (27.94)
The LS estimates of the parameters aare now found, as shown in (27.91), by
differentiating (27.94) with respect to the aiand setting the resulting expressions
equal to zero. Denoting by ∇χ2the vector with elements ∂χ2/∂a i, we find
∇χ2=−2RTN−1(y−Ra). (27.95)
This can be verified by writing out the expression (27.94) in component form and
differentiating directly.IVerify result (27.95) by formulating the calculation in component form.
To make the derivation less cumbersome, let us adopt the summation convention discussed
in section 21.1, in which it is understood that any subscript that appears exactly twice in
any term of an expression is to be summed over all the values that a subscript in thatposition can take. Thus, writing (27.94) in component form, we have
χ
2(a)=(yi−Rikak)(N−1)ij(yj−Rjlal).
Differentiating with respect to apgives
∂χ2
∂ap=−Rikδkp(N−1)ij(yj−Rjlal)+(yi−Rikak)(N−1)ij(−Rjlδlp)
=−Rip(N−1)ij(yj−Rjlal)−(yi−Rikak)(N−1)ijRjp, (27.96)
where δijis the Kronecker delta symbol discussed in section 21.1. By swapping the indices
iandjin the second term on the RHS of (27.96) and using the fact that the matrix N−1
is symmetric, we obtain
∂χ2
∂ap=−2Rip(N−1)ij(yj−Rjkak)
=−2(RT)pi(N−1)ij(yj−Rjkak). (27.97)
If we denote the vector with components ∂χ2/∂a p,p=1,2,...,M ,b y∇χ2and write the
RHS of (27.97) in matrix notation, we recover the result (27.95).
J
1115
STATISTICS
Setting the expression (27.95) equal to zero at a=ˆa, we find
−2RTN−1y+2RTN−1Rˆa=0.
Provided the matrix RTN−1Ris not singular, we may solve this equation for ˆato
obtain
ˆa=(RTN−1R)−1RTN−1y≡Sy, (27.98)
thus defining the M×Nmatrix S. It follows that the LS estimates ˆai,i=1,2,...,M ,
are linear functions of the original measurements yj,j=1,2,...,N .M o r e o v e r ,
using the error propagation formula (26.141) derived in subsection 26.12.3, wefind that the covariance matrix of the estimators ˆa
iis given by
V≡Cov[ˆai,ˆaj]=SNST=(RTN−1R)−1. (27.99)
The two equations (27.98) and (27.99) contain the complete method of least
squares. In particular, we note that, if one calculates the LS estimates using(27.98) then one has already obtained their covariance matrix (27.99).IProve result (27.99).
Using the definition of Sgiven in (27.98), the covariance matrix (27.99) becomes
V=SNST
=[ ( RTN−1R)−1RTN−1]N[(RTN−1R)−1RTN−1]T.
Using the result ( AB···C)T=CT···BTATfor the transpose of a product of matrices and
noting that, for any non-singular matrix, ( A−1)T=(AT)−1we find
V=(RTN−1R)−1RTN−1N(NT)−1R[(RTN−1R)T]−1
=(RTN−1R)−1RTN−1R(RTN−1R)−1
=(RTN−1R)−1,
where we have also used the fact that Nis symmetric and so NT=N.
J
It is worth noting that one may also write the elements of the (inverse)
covariance matrix as
(V−1)ij=1
2parenleftbigg∂2χ2
∂ai∂ajparenrightbigg
a=ˆa,
which is the same as the Fisher matrix (27.36) in cases where the measurement
errors are Gaussian distributed (and so the log-likelihood is ln L=−χ2/2). This
p r o v e s ,a tl e a s tf o rt h i sc a s e ,o u re a r l i e rs t a t e m e n tt h a tt h eL Se s t i m a t o r sa r eminimum-variance. In fact, since f(x;a) is linear in the parameters a,o n ec a n
write χ
2exactly as
χ2(a)=χ2(ˆa)+1
2Msummationdisplay
i,j=1parenleftbigg∂2χ2
∂ai∂ajparenrightbigg
a=ˆa(ai−ˆai)(aj−ˆaj),
which is quadratic in the parameters ai. Hence the likelihood function L∝
1116
27.6 THE METHOD OF LEAST SQUARES
0
01
12
23
34
4 5567y
x
Figure 27.9 A set of data points with error bars indicating the uncertainty
σ=0.5o nt h e y-values. The straight line is y=ˆmx+ˆc,w h e r e ˆmandˆcare
the least squares estimates of the slope and intercept.
exp(−χ2/2) is Gaussian. From the discussions of sections 27.3.6 and 27.5.6, it
follows that the ‘surfaces’ χ2(a)=c,w h e r e cis a constant, bound ellipsoidal
confidence regions for the parameters ai. The relationship between the value of
the constant cand the confidence level is given by (27.39).IAn experiment produces the following data sample pairs (xi,yi):
xi:1.85 2 .72 2 .81 3 .06 3 .42 3 .76 4 .31 4 .47 4 .64 4 .99
yi:2.26 3 .10 3 .80 4 .11 4 .74 4 .31 5 .24 4 .03 5 .69 6 .57
where the xi-values are known exactly but each yi-value is measured only to an accuracy
ofσ=0.5. Assuming the underlying model for the data to be a straight line y=mx+c,
find the LS estimates of the slope mand intercept cand quote the standard error on each
estimate.
The data are plotted in figure 27.9, together wi th error bars indicating the uncertainty in
theyi-values. Our model of the data is a straight line, and so we have
f(x;c, m)=c+mx.
In the language of (27.92), our basis functions are h1(x)=1a n d h2(x)=xand our model
parameters are a1=canda2=m. From (27.93) the elements of the response matrix are
Rij=hj(xi), so that
R=
/0BBB/@1x1
1x2
......
1xN
/1CCCA, (27.100)
where xiare the data values and N= 10 in our case. Further, since the standard deviation
on each measurement error is σ, we have N=σ2I,w h e r e Iis the N×Nidentity matrix.
Because of this simple form for N, the expression (27.98) for the LS estimates reduces to
ˆa=σ2(RTR)−11
σ2RTy=(RTR)−1RTy. (27.101)
Note that we cannot expand the inverse in the last line, since Ritself is not square and
1117
STATISTICS
hence does not possess an inverse. Inserting the form for Rin (27.100) into the expression
(27.101), we find/
ˆc
ˆm
/
=
/P
i1
P
ixiP
ixi
P
ix2
i
/−1
/P
iyiP
ixiyi
/
=1
N(x2−¯x2)
/
x2−¯x
−¯x1
//
N¯y
Nxy
/
.
We thus obtain the LS estimates
ˆm=xy−¯x¯y
x2−¯x2and ˆc=x2¯y−¯xxy
x2−¯x2=¯y−ˆm¯x, (27.102)
where the last expression for ˆcshows that the best-fit line passes through the ‘centre
of mass’ ( ¯x,¯y) of the data sample. To find the standard errors on our results, we must
calculate the covariance matrix of the estimators. This is given by (27.99), which in ourcase reduces to
V=σ
2(RTR)−1=σ2
N(x2−¯x2)
/
x2−¯x
−¯x1
/
. (27.103)
The standard error on each estimator is simply the positive square root of the corresponding
diagonal element, i.e. σˆc=√V11andσˆm=√V22, and the covariance of the estimators ˆm
andˆcis given by Cov[ ˆc,ˆm]=V12=V21. Inserting the data sample averages and moments
into (27.102) and (27.103), we find
c=ˆc±σˆc=0.40±0.62 and m=ˆm±σˆm=1.11±0.17.
The ‘best-fit’ straight line y=ˆmx+ˆcis plotted in figure 27.9. For comparison, the true
values used to create the data were m=1a n d c=1 .
J
The extension to the fitting the data to a higher-order polynomial, such as
f(x;a)=a1+a2x+a3x2, is obvious. Nevertheless, as the order of the polynomial
increases the matrix inversions become rather complicated. Indeed, even when the
matrices are inverted numerically, the inversion is prone to numerical instabilities.
A better approach is to replace the basis functions hm(x)=xm,m=1,2,...,M ,
with a set of polynomials that are ‘orthogonal over the data’, i.e. such that
Nsummationdisplay
i=1hl(xi)hm(xi)=0 f o r l/negationslash=m.
Such a set of polynomial basis functions can always be found by using the Gram–
Schmidt orthogonalisation procedure presented in section 17.1. The details of thisapproach are beyond the scope of our discussion but we note that, in this case,
the matrix R
TRis diagonal and may be inverted easily.
27.6.2 Non-linear least squares
If the function f(x;a)i snotlinear in the parameters athen, in general, it is
not possible to obtain an explicit expression for the LS estimates ˆa.I n s t e a d ,o n e
must use an iterative (numerical) procedure, which we now outline. In practice,
1118
27.7 HYPOTHESIS TESTING
however, such problems are best solved using one of the many commercially
available software packages.
One begins by making a first guess a0for the values of the parameters. At this
point in parameter space, the components of the gradient ∇χ2will, not be equal
to zero, in general (unless one makes a very lucky guess!). Thus, for at least some
values of i, we have
∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=a0/negationslash=0.
Our aim is to find a small increment δain the values of the parameters, such that
∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=a0+δa=0 f o ra l l i. (27.104)
If our first guess a0were sufficiently close to the true (local) minimum of χ2,
we could find the required increment δaby expanding the LHS of (27.104) as a
Taylor series about a=a0, keeping only the zeroth-order and first-order terms:
∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=a0+δa≈∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=a0+Msummationdisplay
j=1∂2χ2
∂ai∂ajvextendsinglevextendsinglevextendsinglevextendsingle
a=a0δaj. (27.105)
Setting this expression to zero, we find that the increments δajmay be found by
solving the set of Mlinear equations
Msummationdisplay
j=1∂2χ2
∂ai∂ajvextendsinglevextendsinglevextendsinglevextendsingle
a=a0δaj=−∂χ2
∂aivextendsinglevextendsinglevextendsinglevextendsingle
a=a0.
It most cases, however, our first guess a0will not be sufficiently close to the true
minimum for (27.105) to be an accurate approximation, and consequently (27.104)will not be satisfied. In this case, a
1=a0+δais (hopefully) an improved guess
at the parameter values; the whole process is then repeated until convergence isachieved.
It is worth noting that, when one is estimating several parameters a,t h e
function χ
2(a)m a yb e verycomplicated. In particular, it may possess numerous
local extrema. The procedure outlined above will converge to the local extremum
‘nearest’ to the first guess a0. Since, in fact,we are interested only in the local
minimum that has the absolute lowest value of χ2(a), it is clear that a large part
of solving the problem is to make a ‘good’ first guess.
27.7 Hypothesis testing
So far we have concentrated on using a data sample to obtain a number or a set
of numbers. These numbers may be estimated values for the moments or central
moments of the population from which the sample was drawn or, more generally,the values of some parameters ain an assumed model for the data. Sometimes,
1119
STATISTICS
however, one wishes to use the data to give a ‘yes’ or ‘no’ answer to a particular
question. For example, one might wish to know whether some assumed modeldoes, in fact, provide a good fit to the data, or whether two parameters have thesame value.
27.7.1 Simple and composite hypotheses
In order to use data to answer questions of this sort, the question must be
posed precisely. This is done by first asserting that some hypothesis is true.
The hypothesis under consideration is traditionally called the null hypothesis
and is denoted by H
0. In particular, this usually specifies some form P(x|H0)
for the probability density function from which the data xare drawn. If the
hypothesis determines the PDF uniquely, then it is said to be a simple hypothesis .
If, however, the hypothesis determines the functional form of the PDF but not thevalues of certain parameters aon which it depends then it is called a composite
hypothesis .
One decides whether to accept orreject the null hypothesis H
0by performing
some statistical test , as described below in subsection 27.7.2. In fact, formally
one uses a statistical test to decide between the null hypothesis H0and the
alternative hypothesis H1. We define the latter to be the complement H0of the
null hypothesis within some restricted hypothesis space known (or assumed) in
advance . Hence, rejection of H0implies acceptance of H1, and vice versa.
As an example, let us consider the case in which a sample xis drawn from a
Gaussian distribution with a known variance σ2but with an unknown mean µ.
If one adopts the null hypothesis H0that µ=0w h i c hw ew r i t ea s H0:µ=0 ,
then the corresponding alternative hypothesis must be H1:µ/negationslash= 0. Note that,
in this case, H0is a simple hypothesis whereas H1is a composite hypothesis.
If, however, one adopted the null hypothesis H0:µ<0 then the alternative
hypothesis would be H1:µ≥0, so that both H0andH1would be composite
hypotheses. Very occasionally both H0andH1will be simple hypotheses. In our
illustration, this would occur, for example, if one knew in advance that the meanµof the Gaussian distribution were equal to either zero or unity. In this case, if
one adopted the null hypothesis H
0:µ= 0 then the alternative hypothesis would
beH1:µ=1 .
27.7.2 Statistical tests
In our discussion of hypothesis testing we will restrict our attention to cases in
which the null hypothesis H0issimple (see above). We begin by constructing a
test statistic t(x) from the data sample. Although, in general, the test statistic need
not be just a (scalar) number, and could be a multi-dimensional (vector) quantity,we will restrict our attention to the former case. Like any statistic, t(x) will be a
1120
27.7 HYPOTHESIS TESTING
tP(t|H0)
tcritα
t
tcritP(t|H1)
β
Figure 27.10 The sampling distributions P(t|H0)a n d P(t|H1) of a test statistic
t. The shaded areas indicate the (one-tailed) regions for which Pr( t>t crit|H0)=
αand Pr( t<t crit|H1)=βrespectively.
random variable. Moreover, given the simple null hypothesis H0concerning the
PDF from which the sample was drawn, we may determine (in principle) thesampling distribution P(t|H
0) of the test statistic. A typical example of such a
sampling distribution is shown in figure 27.10. One defines for tarejection region
containing some fraction αof the total probability. For example, the (one-tailed)
rejection region could consist of values of tgreater than some value tcrit,f o r
which
Pr(t>t crit|H0)=integraldisplay∞
tcritP(t|H0)dt=α; (27.106)
this is indicated by the shaded region in the upper half of figure 27.10. Equally,
a (one-tailed) rejection region could consist of values of tless than some value
tcrit. Alternatively, one could define a (two-tailed) rejection region by two values
t1andt2such that Pr( t1<t<t 2|H0)=α. In all cases, if the observed value of t
lies in the rejection region then H0isrejected atsignificance level α;o t h e r w i s e H0
isaccepted at this same level.
It is clear that there is a probability αof rejecting the null hypothesis H0
even if it is true. This is called an error of the first kind .C o n v e r s e l y ,a n error
of the second kind occurs when the hypothesis H0i sa c c e p t e de v e nt h o u g hi ti s
1121
STATISTICS
false (in which case H1is true). The probability β( s a y )t h a ts u c ha ne r r o rw i l l
occur is, in general, difficult to calculate, since the alternative hypothesis H1is
often composite. Nevertheless, in the case where H1is a simple hypothesis, it is
straightforward (in principle) to calculate β. Denoting the corresponding sampling
distribution of tbyP(t|H1), the probability βis the integral of P(t|H1)o v e rt h e
complement of the rejection region, called the acceptance region . For example, in
the case corresponding to (27.106) this probability is given by
β=P r ( t<t crit|H1)=integraldisplaytcrit
−∞P(t|H1)dt.
This is illustrated in figure 27.10. The quantity 1 −βis called the power of the
statistical test to reject the wrong hypothesis.
27.7.3 The Neyman–Pearson test
In the case where H0andH1are both simple hypotheses, the Neyman–Pearson
lemma (which we shall not prove) allows one to determine the ‘best’ rejection
region and test statistic to use.
We consider first the choice of rejection region. Even in the general case, in
which the test statistic tis a multi-dimensional (vector) quantity, the Neyman–
Pearson lemma states that, for a given significance level α, the rejection region for
H0giving the highest power for the test is the region of t-space for which
P(t|H0)
P(t|H1)>c , (27.107)
where cis some constant determined by the required significance level.
In the case where the test statistic tis a simple scalar quantity, the Neyman–
Pearson lemma is also useful in deciding which such statistic is the ‘best’ inthe sense of having the maximum power for a given significance level α.F r o m
(27.107), we can see that the best statistic is given by the likelihood ratio
t(x)=P(x|H
0)
P(x|H1). (27.108)
and that the corresponding rejection region for H0is given by t<t crit.I nf a c t ,
it is clear that any statistic u=f(t) will be equally good, provided that f(t)i sa
monotonically increasing function of t. The rejection region is then u<f (tcrit).
Alternatively, one may use any test statistic v=g(t)w h e r e g(t) is a monotonically
decreasing function of t; in this case the rejection region becomes v>g(tcrit). To
construct such statistics, however, one must know P(x|H0)a n d P(x|H1) explicitly,
and such cases are rare.
1122
27.7 HYPOTHESIS TESTINGITen independent sample values xi,i=1,2,...,10, are drawn at random from a Gaussian
distribution with standard deviation σ=1. The mean µof the distribution is known to
equal either zero or unity. The sample values are as follows:
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Test the null hypothesis H0:µ=0at the 10% significance level.
The restricted nature of the hypothesis space means that our null and alternative hypotheses
areH0:µ=0a n d H1:µ= 1 respectively. Since H0andH1are both simple hypotheses,
the best test statistic is given by the likelihood ratio (27.108). Thus, denoting the meansbyµ
0andµ1, we have
t(x)=exp
/
−1
2
P
i(xi−µ0)2
/
exp
/
−1
2
P
i(xi−µ1)2
/=exp
/
−1
2
P
i(x2
i−2µ0xi+µ2
0)
/
exp
/
−1
2
P
i(x2
i−2µ1xi+µ2
1)
/
=e x p
/
(µ0−µ1)
P
ixi−1
2N(µ2
0−µ2
1)
/
.
Inserting the values, µ0=0 ,a n d µ1= 1, yields t=e x p (−N¯x+1
2N), where ¯xis the
sample mean. Since −lntis a monotonically decreasing function of t, however, we may
equivalently use as our test statistic
v=−1
Nlnt+1
2=¯x,
where we have divided by the sample size Nand added1
2for convenience. Thus we
may take the sample mean as our test statistic. From (27.13), we know that the samplingdistribution of the sample mean under our null hypothesis H
0is the Gaussian distribution
N(µ0,σ2/N), where µ0=0 , σ2=1a n d N= 10. Thus ¯x∼N(0,0.1).
Since ¯xis a monotonically decreasing function of t, our best rejection region for a given
significance αis¯x>¯xcrit,w h e r e ¯xcritdepends on α. Thus, in our case, ¯xcritis given by
α=1−Φ
/¯xcrit−µ0
σ
/
=1−Φ(10¯xcrit),
where Φ( z) is the cumulative distribution function for the standard Gaussian. For a 10%
significance level we have α=0.1 and, from table 26.3 in subsection 26.9.1, we find
¯xcrit=0.128. Thus the rejection region on ¯xis
¯x>0.128.
From the sample, we deduce that ¯x=1.11, and so we can clearly reject the null hypothesis
H0:µ= 0 at the 10% significance level It can, in fact, be rejected at a much higher
significance level. As revealed on p. 1081, the data was generated using µ=1 .
J
27.7.4 The generalised likelihood-ratio test
If the null hypothesis H0or the alternative hypothesis H1(or both) is composite
then the corresponding distributions P(x|H0)a n d P(x|H1) are not uniquely de-
termined, in general, and so we cannot use the Neyman–Pearson lemma to obtain
the ‘best’ test statistic t. Nevertheless, in many cases, there still exists a general
procedure for constructing a test statistic twhich has useful properties and which
1123
STATISTICS
reduces to the Neyman–Pearson statistic (27.108) in the special case where H0
andH1are both simple hypotheses.
Consider the quite general, and commonly occurring, case in which the
data sample xis drawn from a population P(x|a) with a known (or as-
sumed) functional form but depends on the unknown values of some parametersa
1,a2,...,a M. Moreover, suppose we wish to test the null hypothesis H0that
the parameter values alie in some subspace Sof the full parameter space
A. In other words, on the basis of the sample xit is desired to test the
null hypothesis H0:(a1,a2,...,a Mlies in S) against the alternative hypothesis
H1:(a1,a2,...,a Mlies in S), where SisA−S.
Since the functional form of the population is known, we may write down the
likelihood function L(x;a) for the sample. Ordinarily, the likelihood will have
a maximum as the parameters aare varied over the entire parameter space A.
This is the usual maximum-likelihood estimate of the parameter values, which
we denote by ˆa. If, however, the parameter values are allowed to vary only over
the subspace Sthen the likelihood function will be maximised at the point ˆaS,
which may or may not coincide with the global maximum ˆa. Now, let us take as
our test statistic the generalised likelihood ratio
t(x)=L(x;ˆaS)
L(x;ˆa), (27.109)
where L(x;ˆaS) is the maximum value of the likelihood function in the subspace
SandL(x;ˆa) is its maximum value in the entire parameter space A.I ti sc l e a r
thattis a function of the sample values only and must lie between 0 and 1.
We will concentrate on the special case where H0is the simple hypothesis
H0:a=a0. The subspace Sthen consists of only the single point a0. Thus
(27.109) becomes
t(x)=L(x;ˆa0)
L(x;ˆa), (27.110)
and the sampling distribution P(t|H0) can be determined (in principle). As in the
previous subsection, the best rejection region for a given significance αis simply
t<t crit, where the value tcritdepends on α. Moreover, as before, an equivalent
procedure is to use as a test statistic u=f(t), where f(t) is any monotonically
increasing function of t; the corresponding rejection region is then u<f (tcrit).
Similarly, one may use a test statistic v=g(t), where g(t) is any monotonically
decreasing function of t; the rejection region then becomes v>g(tcrit). Finally,
we note that if H1is also a simple hypothesis H1:a=a1, then (27.110) reduces
to the Neyman–Pearson test statistic (27.108).
1124
27.7 HYPOTHESIS TESTINGITen independent sample values xi(i=1,2,...,10)are drawn at random from a Gaussian
distribution with standard deviation σ=1. The sample values are as follows:
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Test the null hypothesis H0:µ=0at the 10% significance level.
We must test the (simple) null hypothesis H0:µ= 0 against the (composite) alternative
hypothesis H1:µ/negationslash= 0. Thus, the subspace Sis the single point µ=0 ,w h e r e a s Ais the
entire µ-axis. The likelihood function is
L(x;µ)=1
(2π)N/2exp
/
−1
2
P
i(xi−µ)2
/
,
which has its global maximum at µ=¯x. The test statistic tis then given by
t(x)=L(x;0)
L(x;¯x)=exp
/
−1
2
P
ix2
i
/
exp
/
−1
2
P
i(xi−¯x)2
/=e x p
/;
−1
2N¯x2
/
.
It is in fact more convenient to consider the test statistic
v=−2lnt=N¯x2.
Since−2l ntis a monotonically decreasing function of t, the rejection region now becomes
v>v crit,w h e r eZ∞
vcritP(v|H0)dv=α, (27.111)
αbeing the significance level of the test. Thus it only remains to determine the sampling
distribution P(v|H0). Under the null hypothesis H0, we expect ¯xto be Gaussian distributed,
with mean zero and variance 1 /N. Thus, from subsection 26.9.4, vwill follow a chi-squared
distribution of order 1. Substituting the appropriate form for P(v|H0) in (27.111) and
setting α=0.1, we find by numerical integration (or from tables of the cumulative chi-
squared distribution) that vcrit=N¯x2
crit=2.71. Since N= 10, the rejection region on ¯xat
the 10% significance level is thus
¯x<−0.52 and ¯x>0.52.
As noted before, for this sample ¯x=1.11, and so we may reject the null hypothesis
H0:µ= 0 at the 10% significance level.
J
The above example illustrates the general situation that if the maximum-
likelihood estimates ˆaof the parameters fall in or near the subspace Sthen the
sample will be considered consistent with H0and the value of twill be near
unity. If ˆais distant from Sthen the sample will not be in accord with H0and
ordinarily twill have a small (positive) value.
It is clear that in order to prescribe the rejection region for t, or for a related
statistic uorv, it is necessary to know the sampling distribution P(t|H0). IfH0
is simple then one can in principle determine P(t|H0), although this may prove
difficult in practice. Moreover, if H0is composite, then it may not be possible
to obtain P(t|H0), even in principle. Nevertheless, a useful approximate form for
P(t|H0) exists in the large-sample limit. Consider the null hypothesis
H0:(a1=a0
1,a2=a0
2,...,a R=a0
R),where R≤M
1125
STATISTICS
and the a0
iare fixed numbers. (In fact, we may fix the values of any subset
containing Rof the Mparameters.) If H0is true then it follows from our
discussion in subsection 27.5.6 (although we shall not prove it) that, when thesample size Nis large, the quantity −2l ntfollows approximately a chi-squared
distribution of order R.
27.7.5 Student’s t-test
Student’s t-test is just a special case of the generalised likelihood ratio test applied
to a sample x
1,x2,...,x Ndrawn independently from a Gaussian distribution for
which boththe mean µand variance σ2are unknown, and for which one wishes
to distinguish between the hypotheses
H0:µ=µ0,0<σ2<∞,and H1:µ/negationslash=µ0,0<σ2<∞,
where µ0is a given number. Here, the parameter space Ais the half-plane
−∞<µ<∞,0<σ2<∞, whereas the subspace Scharacterised by the null
hypothesis H0is the line µ=µ0,0<σ2<∞.
The likelihood function for this situation is given by
L(x;µ, σ2)=1
(2πσ2)N/2expbracketleftbigg
−summationtext
i(xi−µ)2
2σ2bracketrightbigg
.
On the one hand, as shown in subsection 27.5.1, the values of µandσ2that
maximise LinAareµ=¯xandσ2=s2,w h e r e ¯xis the sample mean and s2is
the sample variance. On the other hand, to maximise Lin the subspace Swe set
µ=µ0, and the only remaining parameter is σ2; the value of σ2that maximises
Lis then easily found to be
hatwideσ2=1
NNsummationdisplay
i=1(xi−µ0)2.
To retain, in due course, the standard notation for Student’s t-test, in this section
we will denote the generalised likelihood ratio by λ(rather than t); it is thus
given by
λ(x)=L(x;µ0,hatwideσ2)
L(x;¯x, s2)
=[(2π/N)summationtext
i(xi−µ0)2]−N/2exp(−N/2)
[(2π/N)summationtext
i(xi−¯x)2]−N/2exp(−N/2)=bracketleftbiggsummationtext
i(xi−¯x)2
summationtext
i(xi−µ0)2bracketrightbiggN/2
.(27.112)
Normally, our next step would be to find the sampling distribution of λunder
the assumption that H0were true. It is more conventional, however, to work in
terms of a related test statistic t, which was first devised by William Gossett, who
wrote under the pen name of ‘Student’.
1126
27.7 HYPOTHESIS TESTING
The sum of squares in the denominator of (27.112) may be put into the form
summationtext
i(xi−µ0)2=N(¯x−µ0)2+summationtext
i(xi−¯x)2.
Thus, on dividing the numerator and denominator in (27.112) bysummationtext
i(xi−¯x)2and
rearranging, the generalised likelihood ratio λcan be written
λ=parenleftbigg
1+t2
N−1parenrightbigg−N/2
,
where we have defined the new variable
t=¯x−µ0
s/√
N−1. (27.113)
Since t2is a monotonically decreasing function of λ, the corresponding rejection
region is t2>c,w h e r e cis a positive constant depending on the required
significance level α. It is conventional, however, to use titself as our test statistic,
in which case our rejection region becomes two-tailed and is given by
t<−tcrit and t>t crit, (27.114)
where tcritis the positive square root of the constant c.
The definition (27.113) and the rejection region (27.114) form the basis of
Student’s t-test. It only remains to determine the sampling distribution P(t|H0).
At the outset, it is worth noting that if we write the expression (27.113) for t
in terms of the standard estimator ˆσ=radicalbig
Ns2/(N−1) of the standard deviation
then we obtain
t=¯x−µ0
ˆσ/√
N. (27.115)
If, in fact, we knew the true value of σand used it in this expression for tthen
it is clear from our discussion in section 27.3 that twould follow a Gaussian
distribution with mean 0 and variance 1, i.e. t∼N(0,1). When σis not known,
however, we have to use our estimate ˆσin (27.115), with the result that tis
no longer distributed as the standard Gaussian. As one might expect fromthe central limit theorem, however, the distribution of tdoes tend towards the
standard Gaussian for large values of N.
As noted earlier, the exact distribution of t, valid for any value of N, was first
discovered by William Gossett. From (27.35), if the hypothesis H
0is true then the
joint sampling distribution of ¯xandsis given by
P(¯x, s|H0)=CsN−2expparenleftbigg
−Ns2
2σ2parenrightbigg
expbracketleftbigg
−N(¯x−µ)2
2σ2bracketrightbigg
,
(27.116)
where Cis a normalisation constant. We can use this result to obtain the joint
sampling distribution of sandtby demanding that
P(¯x, s|H0)d¯xd s=P(t, s|H0)dt ds.
1127
STATISTICS
Using (27.113) to substitute for ¯x−µin (27.116), and noting that d¯x=
(s/√
N−1)dt, we find
P(¯x, s|H0)d¯xd s=AsN−1expbracketleftbigg
−Ns2
2σ2parenleftbigg
1+t2
N−1parenrightbiggbracketrightbigg
dt ds,
where Ais another normalisation constant. In order to obtain the sampling
distribution of talone, we must integrate P(t, s|H0)w i t hr e s p e c tt o sover its
a l l o w e dr a n g e ,f r o m0t o ∞. Thus, the required distribution of talone is given by
P(t|H0)=integraldisplay∞
0P(t, s|H0)ds=Aintegraldisplay∞
0sN−1expbracketleftbigg
−Ns2
2σ2parenleftbigg
1+t2
N−1parenrightbiggbracketrightbigg
ds.
(27.117)
To perform this integration, we make the change of variable y=s{1+[t2/(N−
1)]}1/2, which on substitution into (27.117) yields
P(t|H0)=Aparenleftbigg
1+t2
N−1parenrightbigg−N/2integraldisplay∞
0yN−1expparenleftbigg
−Ny2
2σ2parenrightbigg
dy.
Since the integral over ydoes not depend on t, it is simply a constant. We thus
find that that the sampling distribution of the variable tis
P(t|H0)=1√(N−1)πΓparenleftbig1
2Nparenrightbig
Γparenleftbig1
2(N−1)parenrightbigparenleftbigg
1+t2
N−1parenrightbigg−N/2
,
(27.118)
w h e r ew eh a v eu s e dt h ec o n d i t i o nintegraltext∞
−∞P(t|H0)dt= 1 to determine the normali-
sation constant (see exercise 27.18).
The distribution (27.118) is called Student’s t-distribution with N−1degrees of
freedom . A plot of Student’s t-distribution is shown in figure 27.11 for various
values of N. For comparison, we also plot the standard Gaussian distribution,
to which the t-distribution tends for large N. As is clear from the figure, the
t-distribution is symmetric about t= 0. In table 27.2 we list some critical points
of the cumulative probability function Cn(t)o ft h e t-distribution, which is defined
by
Cn(t)=integraldisplayt
−∞P(t/prime|H0)dt/prime,
where n=N−1 is the number of degrees of freedom. Clearly, Cn(t) is analogous
to the cumulative probability function Φ( z) of the Gaussian distribution, discussed
in subsection 26.9.1. For comparison purposes, we also list the critical points ofΦ(z), which corresponds to the t-distribution for N=∞.
1128
27.7 HYPOTHESIS TESTING
0
00.10.20.30.40.5
−4−3−2−11 2 3 4tP(t|H0)
N=2N=3N=5N=1 0
Figure 27.11 Student’s t-distribution for various values of N. The broken
curve shows the standard Gaussian distribution for comparison.ITen independent sample values xi,i=1,2,...,10, are drawn at random from a Gaussian
distribution with unknown mean µand unknown standard deviation σ. The sample values
are as follows:
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Test the null hypothesis, H0:µ=0at the 10% significance level.
For our null hypothesis µ0=0 .S i n c ef o rt h i ss a m p l e ¯x=1.11,s=1.01 and N= 10, it
follows from (27.113) that
t=¯x
s/√
N−1=3.33.
The rejection region for tis given by (27.114) where tcritis such that
CN−1(tcrit)=1−α/2,
andαis the required significance of the test. In our case α=0.1a n d N= 10, and from
table 27.2 we find tcrit=1.83. Thus our rejection region for H0at the 10% significance
level is
t<−1.83 and t>1.83.
For our sample t=3.30 and so we can clearly reject the null hypothesis H0:µ=0a tt h i s
level.
J
It is worth noting the connection between the t-test and the classical confidence
interval on the mean µ. The central confidence interval on µat the confidence
level 1−α, is the set of values for which
−tcrit<¯x−µ
s/√
N−1<tcrit,
1129
STATISTICS
Cn(t)0.5 0.6 0.7 0.8 0.9 0.950 0.975 0.990 0.995 0.999
n=1 0.00 0.33 0.73 1.38 3.08 6.31 12.7 31.8 63.7 318.3
2 0.00 0.29 0.62 1.06 1.89 2.92 4.30 6.97 9.93 22.3
3 0.00 0.28 0.58 0.98 1.64 2.35 3.18 4.54 5.84 10.2
4 0.00 0.27 0.57 0.94 1.53 2.13 2.78 3.75 4.60 7.17
5 0.00 0.27 0.56 0.92 1.48 2.02 2.57 3.37 4.03 5.89
6 0.00 0.27 0.55 0.91 1.44 1.94 2.45 3.14 3.71 5.21
7 0.00 0.26 0.55 0.90 1.42 1.90 2.37 3.00 3.50 4.79
8 0.00 0.26 0.55 0.89 1.40 1.86 2.31 2.90 3.36 4.50
9 0.00 0.26 0.54 0.88 1.38 1.83 2.26 2.82 3.25 4.30
10 0.00 0.26 0.54 0.88 1.37 1.81 2.23 2.76 3.17 4.14
11 0.00 0.26 0.54 0.88 1.36 1.80 2.20 2.72 3.11 4.03
12 0.00 0.26 0.54 0.87 1.36 1.78 2.18 2.68 3.06 3.93
13 0.00 0.26 0.54 0.87 1.35 1.77 2.16 2.65 3.01 3.85
14 0.00 0.26 0.54 0.87 1.35 1.76 2.15 2.62 2.98 3.79
15 0.00 0.26 0.54 0.87 1.34 1.75 2.13 2.60 2.95 3.73
16 0.00 0.26 0.54 0.87 1.34 1.75 2.12 2.58 2.92 3.69
17 0.00 0.26 0.53 0.86 1.33 1.74 2.11 2.57 2.90 3.65
18 0.00 0.26 0.53 0.86 1.33 1.73 2.10 2.55 2.88 3.61
19 0.00 0.26 0.53 0.86 1.33 1.73 2.09 2.54 2.86 3.58
20 0.00 0.26 0.53 0.86 1.33 1.73 2.09 2.53 2.85 3.55
25 0.00 0.26 0.53 0.86 1.32 1.71 2.06 2.49 2.79 3.46
30 0.00 0.26 0.53 0.85 1.31 1.70 2.04 2.46 2.75 3.39
40 0.00 0.26 0.53 0.85 1.30 1.68 2.02 2.42 2.70 3.31
50 0.00 0.26 0.53 0.85 1.30 1.68 2.01 2.40 2.68 3.26
100 0.00 0.25 0.53 0.85 1.29 1.66 1.98 2.37 2.63 3.17
200 0.00 0.25 0.53 0.84 1.29 1.65 1.97 2.35 2.60 3.13
∞ 0.00 0.25 0.52 0.84 1.28 1.65 1.96 2.33 2.58 3.09
Table 27.2 The confidence limits tof the cumulative probability function
Cn(t) for Student’s t-distribution with ndegrees of freedom. For example,
C5(0.92) = 0 .8. The n=∞row is also the corresponding result for the
standard Gaussian distribution.
where tcritsatisfies CN−1(tcrit)=α/2. Thus the required confidence interval is
¯x−tcrits√
N−1<µ< ¯x+tcrits√
N−1.
Hence, in the above example, the 90% classical central confidence interval on µ
is
0.49<µ< 1.73.
Thet-distribution may also be used to compare different samples from Gaussian
distributions. In particular, let us consider the case where we have two independent
1130
27.7 HYPOTHESIS TESTING
samples of sizes N1andN2, drawn respectively from Gaussian distributions with
a common variance σ2but with possibly different means µ1andµ2.O n et h eb a s i s
of the samples, one wishes to distinguish between the hypotheses
H0:µ1=µ2,0<σ2<∞ and H1:µ1/negationslash=µ2,0<σ2<∞.
In other words, we wish to test the null hypothesis that the samples are drawn
from populations having the same mean. Suppose that the measured samplemeans and standard deviations are ¯x
1,¯x2ands1,s2respectively. In an analogous
way to that presented above, one may show that the generalised likelihood ratiocan be written as
λ=parenleftbigg
1+t
2
N1+N2−2parenrightbigg−(N1+N2)/2
.
In this case, the variable tis given by
t=¯w−ω
ˆσparenleftbiggN1N2
N1+N2parenrightbigg1/2
, (27.119)
where ¯w=¯x1−¯x2,ω=µ1−µ2and
ˆσ=bracketleftbiggN1s2
1+N2s2
2
N1+N2−2bracketrightbigg1/2
.
It is straightforward (albeit with complicated algebra) to show that the variable t
in (27.119) follows Student’s t-distribution with N1+N2−2 degrees of freedom,
and so we may use an appropriate form of Student’s t-test to investigate the null
hypothesis H0:µ1=µ2(or equivalently H0:ω= 0). As above, the t-test can be
used to place a confidence interval on ω=µ1−µ2.ISuppose that two classes of students take the same mathematics examination and the
following percentage marks are obtained:
C l a s s 1 :6 66 23 45 57 78 05 56 06 94 75 0
C l a s s 2 :6 49 07 65 68 17 27 0
Assuming that the two sets of examinations marks are drawn from Gaussian distributions
with a common variance, test the hypothesis H0:µ1=µ2at the 5% significance level. Use
your result to obtain the 95% classical central confidence interval on ω=µ1−µ2.
We begin by calculating the mean and standard deviation of each sample. The number of
values in each sample is N1=1 1a n d N2= 7 respectively, and we find
¯x1=5 9.5,s1=1 2.8a n d ¯x2=7 2.7,s2=1 0.3,
leading to ¯w=¯x1−¯x2=−13.2a n d ˆσ=1 2 .6. Setting ω= 0 in (27.119), we thus find
t=−2.17.
The rejection region for H0is given by (27.114), where tcritsatisfies
CN1+N2−2(tcrit)=1−α/2, (27.120)
1131
STATISTICS
where αis the required significance level of the test. In our case we set α=0.05, and from
table 27.2 with n= 16 we find that tcrit=2.12. The rejection region is therefore
t<−2.12 and t>2.12.
Since t=−2.17 for our samples, we can reject the null hypothesis H0:µ1=µ2, although
only by a small margin. (Indeed, it is easily shown that one cannot reject H0at the 2%
significance level). The 95% central confidence interval on ω=µ1−µ2is given by
¯w−ˆσtcrit
/N1+N2
N1N2
/1/2
<ω< ¯w+ˆσtcrit
/N1+N2
N1N2
/1/2
,
where tcritis given by (27.120). Thus, we find
−26.1<ω<−0.28,
which, as expected, does not (quite) contain ω=0 .
J
In order to apply Student’s t-test in the above example, we had to make the
assumption that the samples were drawn from Gaussian distributions possessing acommon variance, which is clearly unjustified ap r i o r i . We can, however, perform
another test on the data to investigate whether the additional hypothesis σ
2
1=σ2
2
is reasonable; this test is discussed in the next subsection. If this additional test
shows that the hypothesis σ2
1=σ2
2may be accepted (at some suitable significance
level), then we may indeed use the analysis in the above example to infer that
the null hypothesis H0:µ1=µ2may be rejected at the 5% significance level.
If, however, we find that the additional hypothesis σ2
1=σ2
2must be rejected,
then we can only infer from the above example that the hypothesis that the twosamples were drawn from the same Gaussian distribution may be rejected at the5% significance level.
Throughout the above discussion, we have assumed that samples are drawn
from a Gaussian distribution. Although this is true for many random variables,
in practice it is usually impossible to know ap r i o r i whether this is case. It can
be shown, however, that Student’s t-test remains reasonably accurate even if the
sampled distribution(s) differ considerably from a Gaussian. Indeed, for sampleddistributions that differ only slightly from a Gaussian form, the accuracy ofthe test is remarkably good. Nevertheless, when applying the t-test, it is always
important to remember that the assumption of a Gaussian parent population is
central to the method.
27.7.6 Fisher’s F-test
Having concentrated on tests for the mean µof a Gaussian distribution, we
now consider tests for its standard deviation σ. Before discussing Fisher’s F-test
for comparing the standard deviations of two samples, we begin by considering
the case when an independent sample x
1,x2,...,x Nis drawn from a Gaussian
distribution with unknown µandσ, and we wish to distinguish between the two
1132
27.7 HYPOTHESIS TESTING
0
00.050.10
10 20 30 40λ(u)
uλcrit
ab
Figure 27.12 The sampling distribution P(u|H0)f o r N= 10; this is a chi-
squared distribution for N−1 degrees of freedom.
hypotheses
H0:σ2=σ2
0,−∞<µ<∞ and H1:σ2/negationslash=σ2
0,−∞<µ<∞,
where σ2
0is a given number. Here, the parameter space Ais the half-plane
−∞<µ<∞,0<σ2<∞, whereas the subspace Scharacterised by the null
hypothesis H0is the line σ2=σ2
0,−∞<µ<∞.
The likelihood function for this situation is given by
L(x;µ, σ2)=1
(2πσ2)N/2expbracketleftbigg
−summationtext
i(xi−µ)2
2σ2bracketrightbigg
.
The maximum of LinAoccurs at µ=¯xandσ2=s2, whereas the maximum of
LinSis at µ=¯xandσ2=σ2
0. Thus, the generalised likelihood ratio is given by
λ(x)=L(x;¯x, σ2
0)
L(x;¯x, s2)=parenleftBigu
NparenrightBigN/2
expbracketleftbig
−1
2(u−N)bracketrightbig
,
where we have introduced the variable
u=Ns2
σ2
0=summationtext
i(xi−¯x)2
σ2
0. (27.121)
An example of this distribution is plotted in figure 27.12 for N= 10. From
the figure, we see that the rejection region λ<λ critcorresponds to a two-tailed
rejection region on ugiven by
0<u<a and b<u<∞,
where aandbare such that λcrit(a)=λcrit(b), as shown in figure 27.12. In practice,
1133
STATISTICS
however, it is difficult to determine aandbfor a given significance level α,s oa
slightly different rejection region, which we now describe, is usually adopted.
The sampling distribution P(u|H0) may be found straightforwardly from the
sampling distribution of sgiven in (27.35). Let us first determine P(s2|H0)b y
demanding that
P(s|H0)ds=P(s2|H0)d(s2),
from which we find
P(s2|H0)=P(s|H0)
2s=(N/2σ2
0)(N−1)/2
Γparenleftbig1
2(N−1)parenrightbig(s2)(N−3)/2expparenleftbigg
−Ns2
2σ2
0parenrightbigg
.
(27.122)
Thus, the sampling distribution of u=Ns2/σ2
0is given by
P(u|H0)=1
2(N−1)/2Γparenleftbig1
2(N−1)parenrightbigu(N−3)/2expparenleftbig
−1
2uparenrightbig
.
We note, in passing, that the distribution of uis precisely that of an ( N−1) th-
order chi-squared variable (see subsection 26.9.4), i.e. u∼χ2
N−1. Although it does
not give quite the best test, one then takes the rejection region to be
0<u<a and b<u<∞,
with aandbchosen such that the two tails have equal areas ; the advantage of
this choice is that tabulations of the chi-squared distribution make the size of this
region relatively easy to estimate. Thus, for a given significance level α, we have
integraldisplaya
0P(u|H0)du=α/2a n dintegraldisplay∞
bP(u|H0)du=α/2.ITen independent sample values xi,i=1,2,...,10, are drawn at random from a Gaussian
distribution with unknown mean µand standard deviation σ. The sample values are as
follows:
2.22 2 .56 1 .07 0 .24 0 .18 0 .95 0 .73−0.79 2 .09 1 .81
Test the null hypothesis H0:σ2=2at the 10% significance level.
For our null hypothesis σ2
0=2 .S i n c ef o rt h i ss a m p l e s=1.01 and N= 10, from (27.121)
we have u=5.10. For α=0.1 we find, either numerically or using tables, that a=3.30
andb=1 6.92. Thus, our rejection region is
0<u< 3.33 and 16 .92<u<∞.
The value u=5.10 from our sample does not lie in the rejection region, and so we cannot
reject the null hypothesis H0:σ2=2 .
J
1134
27.7 HYPOTHESIS TESTING
We now turn to Fisher’s F-test. Let us suppose that two independent samples
of sizes N1and N2are drawn from Gaussian distributions with means and
variances µ1,σ2
1andµ2,σ2
2respectively, and we wish to distinguish between the
two hypotheses
H0:σ2
1=σ2
2and H1:σ2
1/negationslash=σ2
2.
In this case, the generalised likelihood ratio is found to be
λ=(N1+N2)(N1+N2)/2
NN1/2
1NN2/2
2bracketleftbig
F(N1−1)/(N2−1)bracketrightbigN1/2
bracketleftbig
1+F(N1−1)/(N2−1)bracketrightbig(N1+N2)/2,
where Fis given by the variance ratio
F=N1s2
1/(N1−1)
N2s2
2/(N2−1)≡u2
v2(27.123)
ands1ands2are the standard deviations of the two samples. On plotting λas a
function of F, it is apparent that the rejection region λ<λ critcorresponds to a
two-tailed test on F. Nevertheless, as will shall see below, by defining the fraction
(27.123) appropriately, it is customary to make a one-tailed test on F.
The distribution of Fmay be obtained in a reasonably straightforward manner
by making use of the distribution of the sample variance s2given in (27.122).
Under our null hypothesis H0, the two Gaussian distributions share a common
variance, which we denote by σ2. Changing the variable in (27.122) from s2tou2
we find that u2has the sampling distribution
P(u2|H0)=parenleftbiggN−1
2σ2parenrightbigg(N−1)/21
Γparenleftbig1
2(N−1)parenrightbig(u2)(N−3)/2expbracketleftbigg
−(N−1)u2
2σ2bracketrightbigg
.
Since u2andv2are independent, their joint distribution is simply the product of
their individual distributions and is given by
P(u2|H0)P(v2|H0)=A(u2)(N1−3)/2(v2)(N2−3)/2expbracketleftbigg
−(N1−1)u2+(N2−1)v2
2σ2bracketrightbigg
,
where the constant Ais given by
A=(N1−1)(N1−1)/2(N2−1)(N2−1)/2
2(N1+N2−2)/2σ(N1+N2−2)Γparenleftbig1
2(N1−1)parenrightbig
Γparenleftbig1
2(N2−1)parenrightbig.
(27.124)
Now, for fixed vwe have u2=Fv2andd(u2)=v2dF. Thus, the joint sampling
1135
STATISTICS
distribution P(v2,F|H0) is obtained by requiring that
P(v2,F|H0)d(v2)dF=P(u2|H0)P(v2|H0)d(u2)d(v2).
(27.125)
In order to find the distribution of Falone, we now integrate P(v2,F|H0) with
respect to v2from 0 to∞,f r o mw h i c hw eo b t a i n
P(F|H0)=
parenleftbiggN1−1
N2−1parenrightbigg(N1−1)/21
Bparenleftbig1
2(N1−1),1
2(N2−1)parenrightbigF(N1−3)/2parenleftbigg
1+N1−1
N2−1Fparenrightbigg−(N1+N2−2)/2
,
(27.126)
where Bparenleftbig1
2(N1−1),1
2(N2−1)parenrightbig
is the beta function defined in the Appendix.
P(F|H0) is called the F-distribution (or occasionally the Fisher distribution ) with
(N1−1,N2−1) degrees of freedom..IEvaluate the integral
R∞
0P(v2,F|H0)d(v2)to obtain result (27.126).
From (27.125), we have
P(F|H0)=AF(N1−3)/2
Z∞
0(v2)(N1+N2−4)/2exp
/
−[(N1−1)F+(N2−1)]v2
2σ2
/
d(v2).
Making the substitution x=[ (N1−1)F+(N2−1)]v2/(2σ2), we obtain
P(F|H0)=A
/2σ2
(N1−1)F+(N2−1)
/(N1+N2−2)/2
F(N1−3)/2
Z∞
0x(N1+N2−4)/2e−xdx
=A
/2σ2
(N1−1)F+(N2−1)
/(N1+N2−2)/2
F(N1−3)/2Γ
/;1
2(N1+N2−2)
/
,
where in the last line we have used the definition of the gamma function given in the
Appendix. Using the further result (A11), which expresses the beta function in terms ofthe gamma function, and the expression for Agiven in (27.124), we see that P(F|H
0)i s
indeed given by (27.126).
J
As it does not matter whether the ratio Fgiven in (27.123) is defined as u2/v2
or as v2/u2, it is conventional to put the larger sample variance on the top, so
thatFis always greater than or equal to unity. A large value of Findicates that
the sample variances u2andv2are very different whereas a value of Fclose to
unity means that they are very similar. Therefore, for a given significance α,i ti s
1136
27.7 HYPOTHESIS TESTING
Cn1,n2(F)n1= 1 2345678
n2=1 161 200 216 225 230 234 237 239
2 18.5 19.0 19.2 19.2 19.3 19.3 19.4 19.4
3 10.1 9.55 9.28 9.12 9.01 8.94 8.89 8.85
4 7.71 6.94 6.59 6.39 6.26 6.16 6.09 6.04
5 6.61 5.79 5.41 5.19 5.05 4.95 4.88 4.82
6 5.99 5.14 4.76 4.53 4.39 4.28 4.21 4.15
7 5.59 4.74 4.35 4.12 3.97 3.87 3.79 3.73
8 5.32 4.46 4.07 3.84 3.69 3.58 3.50 3.44
9 5.12 4.26 3.86 3.63 3.48 3.37 3.29 3.23
10 4.96 4.10 3.71 3.48 3.33 3.22 3.14 3.07
20 4.35 3.49 3.10 2.87 2.71 2.60 2.51 2.45
30 4.17 3.32 2.92 2.69 2.53 2.42 2.33 2.27
40 4.08 3.23 2.84 2.61 2.45 2.34 2.25 2.18
50 4.03 3.18 2.79 2.56 2.40 2.29 2.20 2.13
100 3.94 3.09 2.70 2.46 2.31 2.19 2.10 2.03
∞ 3.84 3.00 2.60 2.37 2.21 2.10 2.01 1.94
n1= 9 10 20 30 40 50 100 ∞
n2=1 241 242 248 250 251 252 253 254
2 19.4 19.4 19.4 19.5 19.5 19.5 19.5 19.5
3 8.81 8.79 8.66 8.62 8.59 8.58 8.55 8.53
4 6.00 5.96 5.80 5.75 5.72 5.70 5.66 5.63
5 4.77 4.74 4.56 4.50 4.46 4.44 4.41 4.37
6 4.10 4.06 3.87 3.81 3.77 3.75 3.71 3.67
7 3.68 3.64 3.44 3.38 3.34 3.32 3.27 3.23
8 3.39 3.35 3.15 3.08 3.04 3.02 2.97 2.93
9 3.18 3.14 2.94 2.86 2.83 2.80 2.76 2.71
10 3.02 2.98 2.77 2.70 2.66 2.64 2.59 2.54
20 2.39 2.35 2.12 2.04 1.99 1.97 1.91 1.84
30 2.21 2.16 1.93 2.69 1.79 1.76 1.70 1.62
40 2.12 2.08 1.84 1.74 1.69 1.66 1.59 1.51
50 2.07 2.03 1.78 1.69 1.63 1.60 1.52 1.44
100 1.97 1.93 1.68 1.57 1.52 1.48 1.39 1.28
∞ 1.88 1.83 1.57 1.46 1.39 1.35 1.24 1.00
Table 27.3 Values of Ffor which the cumulative probability function Cn1,n2(F)
of the F-distribution with ( n1,n2) degrees of freedom has the value 0.95. For
example, for n1=1 0a n d n2=6 , Cn1,n2(4.06) = 0 .95.
customary to define the rejection region on FasF>F crit,w h e r e
Cn1,n2(Fcrit)=integraldisplayFcrit
1P(F|H0)dF=α,
and n1=N1−1a n d n2=N2−1 are the numbers of degrees of freedom.
Table 27.3 lists values of Fcritcorresponding to the 5% significance level (i.e.
α=0.05) for various values of n1andn2.
1137
STATISTICSISuppose that two classes of students take the same mathematics examination and the
following percentage marks are obtained:
C l a s s 1 :6 66 23 45 57 78 05 56 06 94 75 0
C l a s s 2 :6 49 07 65 68 17 27 0
Assuming that the two sets of examinations marks are drawn from Gaussian distributions,
test the hypothesis H0:σ2
1=σ2
2at the 5% significance level.
The variances of the two samples are s2
1=( 1 2 .8)2ands2
2=( 1 0 .3)2and the sample sizes
areN1=1 1a n d N2= 7. Thus, we have
u2=N1s2
1
N1−1= 180 .2a n d v2=N2s2
2
N2−1= 123 .8,
where we have taken u2to be the larger value. Thus, F=u2/v2=1.46 to two decimal
places. Since the first sample contains eleven values and the second contains seven values,
we take n1=1 0a n d n2= 6. Consulting table 27.3, we see that, at the 5% significance
level, Fcrit=4.06. Since our value lies comfortably below this, we conclude that there is
no statistical evidence for rejecting the hypothesis that the two samples were drawn fromGaussian distributions with a common variance.J
It is also common to define the variable z=1
2lnF, the distribution of which
can be found straightfowardly from (27.126). This is a useful change of variablesince it can be shown that, for large values of n
1and n2, the variable zis
distributed approximately as a Gaussian with mean1
2(n−1
2−n−1
1) and variance
1
2(n−1
2+n−1
1).
27.7.7 Goodness of fit in least squares problems
We conclude our discussion of hypothesis testing with an example of a goodness
of fit test. In section 27.6, we discussed the use of the method of least squares inestimating the best-fit values of a set of parameters ain a given model y=f(x;a)
for a data set( x
i,yi),i=1,2,...,N . We have not addressed, however, the question
of whether the best-fit model y=f(x;ˆa) does, in fact, provide a good fit of the
data. In other words, we have not considered thus far how to verify that thefunctional form fof our assumed model is indeed correct. In the language of
hypothesis testing, we wish to distinguish between the two hypotheses
H
0: model is correct and H1: model is incorrect .
Given the vague nature of the alternative hypothesis H1, we clearly cannot use
the generalised likelihood-ratio test. Nevertheless, it is still possible to test the
null hypothesis H0at a given significance level α.
The least squares estimates of the parameters ˆa1,ˆa2,...,ˆaM, as discussed in
section 27.6, are those values that minimise the quantity
χ2(a)=Nsummationdisplay
i,j=1[yi−f(xi;a)](N−1)ij[yj−f(xj;a)] = ( y−f)TN−1(y−f).
1138
27.7 HYPOTHESIS TESTING
In the last equality, we rewrote the expression in matrix notation by defining the
column vector fwith elements fi=f(xi;a). The value χ2(ˆa) at this minimum can
be used as a statistic to test the null hypothesis H0, as follows. The Nquantities
yi−f(xi;a) are Gaussian distributed. However, provided the function f(xj;a)i s
linear in the parameters a, the equations (27.98) that determine the least squares
estimate ˆaconstitute a set of Mlinear constraints on these Nquantities. Thus,
as discussed in subsection 26.15.2, the sampling distribution of the quantity χ2(ˆa)
will be a chi-squared distribution with N−Mdegrees of freedom (d.o.f), which has
the expectation value and variance
E[χ2(ˆa)] =N−M and V[χ2(ˆa)] = 2( N−M).
Thus we would expect the value of χ2(ˆa) to lie typically in the range ( N−M)±√2(N−M). A value lying outside this range may suggest that the assumed model
for the data is incorrect. A very small value of χ2(ˆa) is usually an indication that
the model has too many free parameters and has ‘over-fitted’ the data. Morecommonly, the assumed model is simply incorrect, and this usually results in avalue of χ
2(ˆa) that is larger than expected.
One can choose to perform either a one-tailed or a two-tailed test on the
value of χ2(ˆa). It is usual, for a given significance level α, to define the one-tailed
rejection region to be χ2(ˆa)>c,w h e r et h ec o n s t a n t csatisfies
integraldisplay∞
cP(χ2
n)dχ2
n=α (27.127)
andP(χ2
n) is the PDF of the chi-squared distributiom with n=N−Mdegrees of
freedom (see subsection 26.9.4).IAn experiment produces the following data sample pairs (xi,yi):
xi:1.85 2 .72 2 .81 3 .06 3 .42 3 .76 4 .31 4 .47 4 .64 4 .99
yi:2.26 3 .10 3 .80 4 .11 4 .74 4 .31 5 .24 4 .03 5 .69 6 .57
where the xi-values are known exactly but each yi-value is measured only to an accuracy
ofσ=0.5. At the one-tailed 5% significance level, tes t the null hypothesis H0that the
underlying model for the data is a straight line y=mx+c.
These data are the same as those investigated in section 27.6 and plotted in figure 27.9. As
shown previously, the least squares estimates of the slope mand intercept care given by
ˆm=1.11 and ˆc=0.4. (27.128)
Since the error on each yi-value is drawn independently from a Gaussian distribution with
standard deviation σ, we have
χ2(a)=NX
i=1
/yi−f(xi;a)
σ
/2
=NX
i=1
hyi−mxi−c
σ
i2
. (27.129)
Inserting the values (27.128) into (27.129), we obtain χ2(ˆm,ˆc)=1 1 .5. In our case, the
number of data points is N= 10 and the number of fitted parameters is M= 2. Thus, the
1139
STATISTICS
number of degrees of freedom is n=N−M= 8. Setting n=8a n d α=0.05 in (27.127)
we find, either numerically or from tables, that c=1 5.51. Hence our rejection region is
χ2(ˆm,ˆc)>15.51.
Since we found that χ2(ˆm,ˆc)=1 1 .5, we cannot reject the null hypothesis that the underlying
model for the data is a straight line y=mx+c.
J
As mentioned above, our analysis is only valid if the function f(x;a) is linear
in the parameters a. Nevertheless, it is so convenient that it is sometimes applied
in non-linear cases, provided the non-linearity is not too severe.
27.8 Exercises
27.1 A group of students uses a pendulum experiment to measure g, the acceleration
of free fall, and obtain the following values (in m s−2): 9.80, 9.84, 9.72, 9.74,
9.87, 9.77, 9.28, 9.86, 9.81, 9.79, 9.82. What would you give as the best value andstandard error for gas measured by the group?
27.2 Measurements of a certain quantity gave the following values: 296, 316, 307, 278,
312, 317, 314, 307, 313, 306, 320, 309. Within what limits would you say there isa 50% chance that the correct value lies?
27.3 The following are the values obtained by a class of 14 students when measuring
a physical quantity x: 53.8, 53.1, 56.9, 54.7, 58.2, 54.1, 56.4, 54.8, 57.3, 51.0, 55.1,
55.0, 54.2, 56.6.
(a) Display these results as a histogram and state what you would give as the
best value for x.
(b) Without calculation estimate how much reliance could be placed upon your
answer to (a).
(c) Databooks give the value of xas 53.6 with negligible error. Are the data
obtained by the students in conflict with this?
27.4 Two physical quantities xandyare connected by the equation
y
1/2=x
ax1/2+b,
and measured pairs of values for xandyare as follows:
x:1 01 21 6 2 0
y: 409 196 114 94.
Determine the best values for aandbby graphical means and (either by hand
or by using a built-in calculator routine) by a least squares fit to an appropriatestraight line,
27.5 Measured quantities xandyare known to be connected by the formula
y=ax
x2+b,
where aandbare constants. Pairs of values obtained experimentally are
x: 2.0 3.0 4.0 5.0 6.0
y: 0.32 0.29 0.25 0.21 0.18.
Use these data to make best estimates of the values of ythat would be obtained
for (a) x=7.0, and (b) x=−3.5. As measured by fractional error, which estimate
is likely to be the more accurate?
1140
27.8 EXERCISES
27.6 Prove that the sample mean is the best linear unbiased estimator of the population
mean µas follows.
(a) If the real numbers a1,a2,...,a nsatisfy the constraint
Pn
i=1ai=C,w h e r e C
is a given constant, show that
Pn
i=1a2
iis minimised by ai=C/nfor all i.
(b) Consider the linear estimator ˆµ=
Pn
i=1aixi. Impose the conditions (i) that it
isunbiased , and (ii) that it is as efficient as possible.
27.7 A population contains individuals of ktypes in equal proportions. A quantity X
has mean µiamongst individuals of type i, and variance σ2which has the same
value for all types. In order to estimate the mean of Xover the whole population,
two schemes are considered; each involves a total sample size of nk. In the first
the sample is drawn randomly from the whole population, whilst in the second(stratified sampling )nindividuals are randomly selected from each of the ktypes.
Show that in both cases the estimate has expectation
µ=1
kkX
i=1µi,
but that the variance of the first scheme exceeds that of the second by an amount
1
k2nkX
i=1(µi−µ)2.
27.8 Carry through the following proofs of statements made in subsections 27.5.2 and
27.5.3 about the ML estimators ˆτandˆλ.
(a) Find the expectation values of the ML estimators ˆτandˆλgiven respectively
in (27.71) and (27.75). Hence verify equations (27.76) which show that, eventhough an ML estimator is unbiased, it does not follow that functions of itare also unbiased.
(b) Show that E[ˆτ
2]=(N+1)τ2/Nand hence prove that ˆτis a minimum-variance
estimator of τ.
27.9 An experiment consists of a large, but unknown, number n(/greatermuch1) of trials in
each of which the probability of success pis the same, but also unkown. In the
ith trial, i=1,2,...,N , the total number of successes is xi(/greatermuch1). Determine the
log-likelihood function.Using Stirling’s approximation to ln( n−x), show that
dln(n−x)
dn≈1
2(n−x)+l n ( n−x),
and hence evaluate ∂(nCx)/∂n.
By finding the (coupled) equations determining the ML estimators ˆpandˆn,s h o w
that, to order N−1, they must satisfy the simultaneous ‘arithmetic’ and ‘geometric’
mean constraints
ˆnˆp=1
NNX
i=1xiand (1−ˆp)N=NY
i=1
/
1−xi
ˆn
/
.
1141
STATISTICS
27.10 This exercise is intended to illustrate the dangers of applying formalised estimator
techniques to distributions that are not well behaved in a statistical sense.
The following are five sets of 10 values, all drawn from the same Cauchy
distribution with parameter a.
(i) 4 .81−1.24 1 .30−0.23 2 .98
−1.13−8.32 2 .62−0.79−2.85
(ii) 0 .07 1 .54 0 .38−2.76−8.82
1.86−4.75 4 .81 1 .14−0.66
(iii) 0 .72 4 .57 0 .86−3.86 0 .30
−2.00 2 .65−17.44−2.26−8.83
(iv)−0.15 202 .76−0.21−0.58−0.14
0.36 0 .44 3 .36−2.96 5 .51
(v) 0 .24−3.33−1.30 3 .05 3 .99
1.59−7.76 0 .91 2 .80−6.46
Ignoring the fact that the Cauchy distribution does not have a finite variance (or
even a formal mean), show that ˆa,t h eM Le s t i m a t o ro f a, has to satisfy
s(ˆa)=10X
i=11
1+x2
i/ˆa2=5.(∗)
Using a programmable calculator, spreadsheet or computer, find the value of
ˆathat satisfies (*) for each of the data sets and compare it with the value
a=1.6 used to generate the data. Form an opinion regarding the variance of the
estimator.
Show further that if it is assumed that (E[ˆa])2=E[ˆa2]t h e n E[ˆa]=ν1/2
2,w h e r e
ν2is the second (central) moment of the distribution, which for the Cauchy
distribution is infinite!
27.11 According to a particular theory, two dimensionless quantities XandYhave
equal values. Nine measurements of Xgave values of 22, 11, 19, 19, 14, 27, 8,
24 and 18, whilst seven measured values of Ywere 11, 14, 17, 14, 19, 16 and
14. Assuming that the measurements of both quantities are Gaussian distributedwith a common variance, are they consistent with the theory? An alternativetheory predicts that Y
2=π2X; is the data consistent with this proposal?
27.12 On a certain (testing) steeplechase course there are 12 fences to be jumped and
any horse that falls is not allowed to continue in the race. In a season of racing
a total of 500 horses started the course and the following numbers fell at eachfence:
F e n c e :123456789 1 0 1 1 1 2
F a l l s : 6 27 54 92 93 32 53 01 71 91 11 51 2
Use this data to determine the overall probability of a horse falling at a fence,
and test the hypothesis that it is the same for all horses and fences as follows.
(a) draw up a table of the expected number of falls at each fence on the basis
of the hypothesis;
(b) consider for each fence ithe standardised variable
z
i=estimated falls −actual falls
standard deviation of estimated falls
and use it in an appropriate χ2test;
(c) show that the data indicates that the odds against all fences being equally
testing are about 40 to 1. Identify the fences that are significantly easier orharder than the average.
1142
27.8 EXERCISES
27.13 A similar technique to that employed in exercise 27.12 can be used to test
correlations between characteristics of sampled data. To illustrate this considerthe following problem.
During an investigation into possible links between mathematics and classical
music, pupils at a school were asked whether they had preferences (a) betweenmathematics and english, and (b) between classical and pop music. The resultsare given below.
Classical None Pop
Mathematics 23 13 14None 17 17 36English 30 10 40
By computing tables of expected numbers, based on the assumption that no
correlations exist, and calculating the relevant values of χ
2, determine whether
there is any evidence for
(a) a link between academic and musical tastes, and
(b) a claim that pupils either had preferences in both areas or had no preference.
You will need to consider the appropriate value for the number of degrees of
freedom to use when applying the χ2test.
27.14 Three candidates X,YandZwere standing for election to a vacant seat on their
college’s Student Committee. The members of the electorate (current first-yearstudents, consisting of 150 men and 105 women) were each allowed to crossout the name of the candidate they least wished to be elected, the other twocandidates then being credited with one vote each. the following data are known.
(a)Xreceived 100 votes from men, whilst Yreceived 65 votes from women.
(b)Zreceived five more votes from men than Xreceived from women.
(c) The total votes cast for XandYwere equal.
Analyse this data in such a way that a χ
2test can be used to determine whether
voting was other than random (i) amongst men, and (ii) amongst women.
27.15 A particle detector consisting of a shielded scintillator is being tested by placing
it near a particle source of controlled intensity (by the use of absorbers). It mightregister counts even in the absence of particles from the source because of thecosmic ray background.
The number of counts nregistered in a fixed time interval as a function of the
source strength sis given in the following table:
Source strength: s0123456
Counts n: 6 11 20 42 44 62 61
At any given source strength the number of counts is expected to be Poisson
distributed with mean
n=a+bs,
where aandbare constants. Analyse the data for a fit to this relationship and
obtain the best values for aandbtogether with their standard errors.
(a) How well is the cosmic ray background determined?
(b) What is the value of the correlation coefficient between aandb?I st h i s
consistent with what would happen if the cosmic ray background were
imagined to be negligible?
(c) Does the data fit the expected relationship well? Is there any evidence that
t h er e p o r t e dd a t a‘ i st o og o o dafi t ’ ?
1143
STATISTICS
27.16 The function y(x) is known to be a quadratic function of x. The following table
gives the measured values and uncorrelated standard errors of ymeasured at
various values of x(in which there is negligible error):
x 1234 5
y(x)3 .5±0.52 .0±0.53 .0±0.56 .5±1.01 0 .5±1.0
Construct the response matrix Rusing as basis functions 1 ,x ,x2. Calculate the
matrix RTN−1Rand show that its inverse, the covariance matrix V,h a st h ef o r m
V=1
9184
/0/@12592−9708 1580
−9708 8413 −1461
1580−1461 269
/1A.
Use this matrix to find the best values, and their uncertainties, for the coefficients
of the quadratic form for y(x).
27.17 The following are the values and standard errors of a physical quantity f(θ)
measured at various values of θ(in which there is negligible error):
θ 0 π/6 π/4 π/3
f(θ)3 .72±0.21 .98±0.1−0.06±0.1−2.05±0.1
θπ / 22 π/33 π/4 π
f(θ)−2.83±0.21 .15±0.13 .99±0.29 .71±0.4
Theory suggests that fshould be of the form a1+a2cosθ+a3cos 2θ. Show that
the normal equations for the coefficients aiare
481.3a1+ 158 .4a2−43.8a3= 284 .7,
158.4a1+ 218 .8a2+6 2.1a3=−31.1,
−43.8a1+6 2.1a2+ 131 .3a3= 368 .4.
(a) If you have matrix inversion routines available on a computer, determine the
best values and variances for the coefficients aiand the correlation between
the coefficients a1anda2.
(b) If you have only a calculator available, solve for the values using Gauss–
Seidel iteration starting from the approximate solution a1=2,a2=−2,a3=
4.
27.18 Prove that the expression given for the Student’s t-distribution in equation (27.118)
is correctly normalised.
27.19 Verify that the F-distribution P(F) given explicitly in equation (27.126) is symme-
tric between the two data samples, i.e. that it retains the same form but with N1
andN2interchanged, if Fis replaced by F/prime=F−1. Symbolically, if P/prime(F/prime)i st h e
distribution of F/primeandP(F)=η(F,N 1,N2), then P/prime(F/prime)=η(F/prime,N2,N1).
27.20 It is claimed that the two following sets of values were obtained (a) by ran-
domly drawing from a normal distribution that is N(0,1) and then (b) randomly
assigning each reading to one of the two sets A and B.
Set A−0.314 0 .603−0.551−0.537−0.160−1.635 0 .719
0.610 0 .482−1.757 0 .058
Set B−0.691 1 .515−1.642−1.736 1 .224 1 .423 1 .165
Make tests, including t-a n d F-tests, to establish whether there is any evidence
that either claims is, or both claims are, false.
1144
27.9 HINTS AND ANSWERS
27.9 Hints and answers
27.1 Note that the reading of 9.28 m s−2is clearly in error and should not be used in
the calculation; 9 .80±0.02 m s−2.
27.2 The reading of 278 should probably be rejected. The other readings do not look
as though they are Gaussian distributed and the best estimate is probably givenby the inter-quartile range of the remaining 11 readings, i.e. 307 – 316.
27.3 (a) 55.1. (b) Note that two thirds of the readings lie within ±2 of the mean and
that 14 readings are being used. This gives a standard error in the mean ≈0.6.
(c) Student’s thas a value of about 2.5 for 13 d.o.f. (degrees of freedom), and
therefore it is likely at the 3% significance level that the data are in conflict withthe accepted value.
27.4 Plot either xy
−1/2versus x1/2or (x/y)1/2versus x−1/2;a=1.20,b=−3.29.
27.5 Plot or calculate a least squares fit of either x2versus x/yorxyversus y/x
to obtain a≈1.19 and b≈3.4. (a) 0.16; (b) −0.27. Estimate (b) is the more
accurate because, using the fact that y(−x)=−y(x), it is effectively obtained by
interpolation rather than extrapolation.
27.6 (a) Use Lagrange multipliers. (b) Write xiasµ+zi,w h e r e/angbracketleftz2
i/angbracketright=σ2for all i,a n d
use the result of (a) to evaluate E[(ˆµ−µ)2].
27.7 Recall that, because of the equal proportions of each type, the expected numbers
of each type in the first scheme is n. Show that the variance of the estimator for
the second scheme is σ2/(kn). When calculating that for the first scheme, recall
that¯x2
i=µ2
i+σ2and note that µ2
ican be written as ( µi−µ+µ)2.
27.8 (a) Note that
R
P(x|τ)
Q
j/negationslash=idxj=τ−1exp(−xi/τ).
With E[ˆλ]=
R
N(
Pxi)−1λNexp(−λ
Pxi)dx, find the first-order differential
equation involving dE[ˆλ]/dλ. The relevant integrating factor is λ−N.
(b) Denoting τ−1
R
xrexp(−x/τ)dxbyJr, show that E[ˆτ2]=N−1[J2+(N−1)J2
1].
27.9 The log-likelihood function is
lnL=NX
i=1lnnCxi+NX
i=1xilnp+
/
Nn−NX
i=1xi
/!
ln(1−p);
∂(nCx)
∂n≈ln
/n
n−x
/
−x
2n(n−x).
Ignore the second term on the RHS of the above to obtain
NX
i=1ln
/n
n−xi
/
+Nln(1−p)=0 .
27.10 Remember that aappears in the normalisation constant of the Cauchy distribu-
tion.(i) 1.85; (ii) 1.66; (iii) 2.46; (iv) 0.68; (v) 2.44. Although the estimates have thecorrect order of magnitude, there is clearly a very large (perhaps infinite) sam-pling variance. Even when all 50 samples are combined the estimated value of1.84 is still 0.24 different from that used to generate the data.
27.11 ¯X=1 8 .0±2.2,¯Y=1 5 .0±1.1.ˆσ=4.92 giving t=1.21 for 14 d.o.f., and
is significant only at the 75% level. Thus there is no significant disagreementbetween the data and the theory. For the second theory, only the mean values canbe tested as Y
2will not be Gaussian distributed. ¯Y2−π2¯X=4 7±38 and is not
significantly different from zero. Again the data is consistent with the proposedtheory.
1145
STATISTICS
27.12 Whilst the distribution of falls at each fence is formally binomial, it can be
approximated for the purposes of the question by a Poisson distribution. Thetotal number of falls is 377. The total number of attempted jumps = 3202.Overall probability of a fall is 0.1177.(a) 58.9, 51.6, 42.7, 37.0, 33.6, 29.7, 26.7, 23.2, 21.2, 19.0, 17.7, 15.9. (b) χ
2=2 1.2
for 11 d.o.f. (c) The χ2-value is close to the 97.5% confidence limit. Fence 2 is
much harder than the rest, fence 10 is easier, and fences 4 and 8 are somewhateasier than the average.
27.13 Consider how many entries may be chosen freely in the table if all row and
column totals are to match the observed values. It should be clear that for anm×ntable the number of degrees of freedom is ( m−1)(n−1).
(a) In order to make the fractions expre ssing each preference or lack of prefer-
ence correct, the expected distribution, if there were no correlation, is
Classical None Pop
Mathematics 17.5 10 22.5
None 24.5 14 31.5
English 28 16 36
This gives a χ
2of 12.3 for 4 d.o.f., making it less than 2% likely, that no
correlation exists.
(b) The expected distribution, if there were no correlation, is
Music preference No music preference
Academic preference 104 26No academic preference 56 14
This gives a χ
2of 1.2 for one d.o.f and no evidence for the claim.
27.14 The votes are not statistically independent and, once established, must be con-
verted to a table of deletions before any χ2test is applied. This table is
NotXNotYNotZ
Men 50 35 65Women 25 40 40
The relevant values of χ
2are (i) 9.0 and (ii) 4.3, both for 2 d.o.f., suggesting that
the voting by men was almost certainly not random but that by women mayhave been.
27.15 As the distribution at each value of sis Poisson, the best estimate of the
measurement error is the square root of the number of counts, i.e.√
n(s). Linear
regression gives a=4.3±2.1a n d b=1 0.06±0.94.
(a) The cosmic ray background must be present, since n(0)/negationslash= 0 but its value of
about 4 is uncertain to within a factor of 2.
(b) The correlation coefficient between aandbis−0.63. Yes; if awere reduced
towards zero then bwould have to be increased to compensate.
(c) Yes, χ2=4.9 for 5 d.o.f., which is almost exactly the ‘expected’ value, neither
t o og o o dn o rt o ob a d .
27.16 The matrix RTN−1Rhas entries 14, 33, 97; 33, 97, 333; 97, 333, 1273, whilst
RTN−1b,w h e r e bis the data vector, has entries 51, 144.5, 520.5.
y(x)=( 6 .73±1.17)−(4.34±0.96)x+( 1.03±0.17)x2.
27.17 a1=2.02±0.06,a2=−2.99±0.09,a3=4.90±0.10;r12=−0.60.
1146
27.9 HINTS AND ANSWERS
27.18 Make the substitution t=√
N−1t a n θto reduce the integral of the t-dependent
part of (27.118) to 2√
N−1
Rπ/2
0cosN−2θd θ. Relate this to the beta function
B
/;1
2,1
2(N−1)
/
, and hence to the gamma functions, using the relationships given
in the Appendix. Note that Γ
/;1
2
/
=√π.
27.19 Note that |dF|=|dF/prime/F/prime2|and write
1+N1−1
(N2−1)F/primeas
/N1−1
(N2−1)F/prime
//
1+(N2−1)F/prime
N1−1
/
.
27.20 (a) The mean and variance of the whole sample are −0.068 and 1 .180, which are
obviously compatible with N(0,1) without the need for statistical tests.
(b) The means and variances of the two sets are: A, −0.226 and 0 .815; B, 0 .180
and 2 .554. The value of tfor the difference between the two means is 0 .692; for
16 degrees of freedom, this or a greater value of tcan be expected in marginally
more than half all cases. The value of Fis 3.13. For n1=6a n d n2= 10, this
value is very close to the 95% confidence limit of 3 .22. Thus it is rather unlikely
that the allocation between the two groups was made at random – set B hassignificantly more readings that are more than one standard deviation from themean for a N(0,1) distribution than it should.
1147
28
Numerical methods
It happens frequently that the end product of a calculation or piece of analysis
is one or more algebraic or differential equations, or an integral that cannot beevaluated in closed form or in terms of tabulated or pre-programmed functions.From the point of view of the physical scientist or engineer, who needs numericalvalues for prediction or comparison with experiment, the calculation or analysisis thus incomplete.
With the ready availability of standard packages on powerful computers for
the numerical solution of equations, both algebraic and differential, and for theevaluation of integrals, in principle there is no need for the investigator to doother than turn to them. However, it should be a part of every engineer’s orscientist’s competence to have some understanding of the kinds of procedure thatare being put into practice within those packages. The present chapter indicates
(at a simple level) some of the ways in which analytically intractable problems
can be tackled using numerical methods.
In the restricted space available in a book of this nature it is clearly not
possible to give anything like a full discussion, even of the elementary points thatwill be made in this chapter. The limited objective adopted is that of explainingand illustrating by simple examples some of the basic principles involved. Inmany cases, the examples used can be solved in closed form anyway, but this
‘obviousness’ of the answers should not detract from their illustrative usefulness,
and it is hoped that their transparency will help the reader to appreciate some ofthe inner workings of the methods described.
The student who proposes to study complicated sets of equations or make
repeated use of the same procedures by, for example, writing computer programsto carry out the computations, will find it essential to acquire a good under-
standing of topics hardly mentioned here. Amongst these are the sensitivity of
the adopted procedures to errors introduced by the limited accuracy with whicha numerical value can be stored in a computer (rounding errors) and to the
1148
28.1 ALGEBRAIC AND TRANSCENDENTAL EQUATIONS
errors introduced as a result of approximations made in setting up the numerical
procedures (truncation errors). For this scale of application, books specificallydevoted to numerical analysis, data analysis and computer programming shouldbe consulted.
So far as is possible, the method of presentation here is that of indicating
and discussing in a qualitative way the main steps in the procedure, and thenof following this with an elementary worked example. The examples have beenrestricted in complexity to a level at which they can be carried out with a pocketcalculator. Naturally it will not be possible for the student to check all thenumerical values presented unless he or she has a programmable calculator orcomputer readily available, and even then it might be tedious to do so. However,
it is advisable to check the initial step and at least one step in the middle of
each repetitive calculation given in the text, so that how the symbolic equationsare used with actual numbers is understood. Clearly the intermediate step shouldbe chosen to be at a point in the calculation at which the changes are stillsufficiently large that they can be detected by whatever calculating device isused.
Where alternative methods for solving the same type of problem are discussed,
for example in finding the roots of a polynomial equation, we have usually
taken the same example to illustrate each method. This could give the mistakenimpression that the methods are very restricted in applicability, but it is felt bythe authors that using the same examples repeatedly has sufficient advantages, interms of illustrating the relative characteristics of competing methods, to justify
doing so. Once the principles are clear, little is to be gained by using new exampleseach time and, in fact, having some prior knowledge of the ‘correct answer’ should
allow the reader to judge the efficiency and dangers of particular methods as the
successive steps are followed through.
One other point remains to be mentioned. Here, in contrast with every other
chapter of this book, the value of a large selection of exercises is not clear cut.The reader with sufficient computing resources to tackle them can easily devisealgebraic or differential equations to be solved, or functions to be integrated
(which perhaps have arisen in other contexts). Further, the solutions of these
problems will be self-checking, for the most part. Consequently, although anumber of exercises are included, no attempt has been made to test the full rangeof ideas treated in this chapter.
28.1 Algebraic and transcendental equations
The problem of finding the real roots of an equation of the form f(x)=0 ,w h e r e
f(x) is an algebraic or transcendental function of x, is one that can sometimes
be treated numerically even if explicit solutions in closed form are not feasible.
1149
NUMERICAL METHODS
0.20.40.60.81.01.21.41.61.8
−4−202468101214
xf(x)
f(x)=x5−2x2−3
Figure 28.1 A graph of the function f(x)=x5−2x2−3f o r xin the range
0≤x≤1.9.
Examples of the types of equation mentioned are the quartic equation
ax4+bx+c=0,
and the transcendental equation
x−3t a nh x=0.
The latter type is characterised by the fact that it contains in effect a polynomial
of infinite order on the left-hand side.
We will discuss four methods that, in various circumstances, can be used to
obtain the real roots of equations of the above types. In all cases we will take as
the specific equation to be solved the fifth-order polynomial equation
f(x)≡x5−2x2−3=0 . (28.1)
The reasons for using the same equation each time were discussed in the intro-
duction to this chapter.
For future reference and so that the reader may follow some of the calculations
leading to the evaluation of the real root of (28.1), a graph of f(x) in the range
0≤x≤1.9 is shown in figure 28.1.
Equation (28.1) is one for which no solution can be found in closed form, that
is in the form x=awhere adoes not explicitly contain x. The general scheme to
be employed will be an iterative one in which successive approximations to a real
root of (28.1) will be obtained, each approximation, it is to be hoped, being better
than the preceding one; certainly, we require that the approximations convergeand that they have as their limit the sought-for root. Let us denote the required
1150
28.1 ALGEBRAIC AND TRANSCENDENTAL EQUATIONS
root by ξand the values of successive approximations by x1,x2,...,xn,....T h e n
for any particular method to be successful,
lim
n→∞xn=ξwhere f(ξ)=0 . (28.2)
However, success as defined here is not the only criterion. Since, in practice,
only a finite number of iterations will be possible, it is important that the values
ofxnbe close to that of ξfor all n>N ,w h e r e Nis a relatively low number;
exactly how low it is naturally depends on the computing resources available andthe accuracy required in the final answer.
So that the reader may assess the progress of the calculations that follow, we
record that to nine significant figures the real root of (28.1) has the value
ξ=1.495 106 40 . (28.3)
We now consider in turn four methods for determining the value of this root.
28.1.1 Rearrangement of the equation
If equation (28.1), f(x) = 0, can be recast into the form
x=φ(x) (28.4)
where φ(x)i saslowly varying function of xthen an iteration scheme
x
n+1=φ(xn) (28.5)
will often produce a fair approximation to the root ξafter a few iterations, as
follows. Clearly ξ=φ(ξ)s i n c e f(ξ) = 0; thus when xnis close to ξthe next
approximation, xn+1, will differ little from xn, the actual size of the difference giving
an order-of-magnitude indication of the inaccuracy in xn+1(when compared
with ξ).
In the present case the equation can be written
x=( 2x2+3 )1/5. (28.6)
Because of the presence of the one-fifth power, the RHS is rather insensitive
to the value of xused to compute it, and so the form (28.6) fits the general
requirements for the method to work satisfactorily. It remains only to choose astarting approximation. It is easy to see from figure 28.1 that the value x=1.5
would be a good starting point but, so that the behaviour of the procedure at
values some way from the actual root can be studied, we will make a poorerchoice, x
1=1.7.
With this starting value and the general recurrence relationship
xn+1=( 2x2
n+3 )1/5, (28.7)
1151
NUMERICAL METHODS
nx n f(xn)
1 1.7 5.42
2 1.544 18 1.013 1.506 86 2 .28×10
−1
4 1.497 92 5 .37×10−2
5 1.495 78 1 .28×10−2
6 1.495 27 3 .11×10−3
7 1.495 14 7 .34×10−4
8 1.495 12 1 .76×10−4
Table 28.1 Successive approximations to the root of (28.1) using the rear-
rangement method.
nA n f(An) Bnf(Bn) xn f(xn)
11 . 0 −4.0000 1.7 5.4186 1.2973 −2.6916
2 1.2973 −2.6916 1.7 5.4186 1.4310 −1.0957
3 1.4310 −1.0957 1.7 5.4186 1.4762 −0.3482
4 1.4762 −0.3482 1.7 5.4186 1.4897 −0.1016
5 1.4897 −0.1016 1.7 5.4186 1.4936 −0.0289
6 1.4936 −0.0289 1.7 5.4186 1.4947 −0.0082
Table 28.2 Successive approximations to the root of (28.1) using linear
interpolation.
successive values can be found. These are recorded in table 28.1. Although not
strictly necessary, the value of f(xn)≡x5
n−2x2
n−3 is also shown at each stage.
It will be seen that x7and all later xnagree with the precise answer (28.3)
to within one part in 104. However, f(xn)a n d xn−ξare both reduced by a
factor of only about 4 for each iteration; thus a large number of iterations wouldbe needed to produce a very accurate answer. The factor 4 is of course specificto this particular problem and would be different for a different equation. The
successive values of x
nare shown in graph ( a) of figure 28.2.
28.1.2 Linear interpolation
In this approach two values A1and B1ofxare chosen with A1<B 1and
such that f(A1)a n d f(B1) have opposite signs. The chord joining the two points
(A1,f(A1)) and ( B1,f(B1)) is then notionally constructed, as illustrated in graph
(b) of figure 28.2, and the value x1at which the chord cuts the x-axis is determined
by the interpolation formula
xn=Anf(Bn)−Bnf(An)
f(Bn)−f(An), (28.8)
1152
28.1 ALGEBRAIC AND TRANSCENDENTAL EQUATIONS
1.01.0 1.0
1.01 .21.2 1.2
1.21 .41.4 1.4
1.4 1.61.6 1.6
1.6
−4−4 −4
−4−2−2 −2
−222 2
244 4
466 6
6x1x1
x2x2x3
x3
x4 x1
x1x2 x2x3
x3ξ
ξ
ξξ(a) (b)
(c) (d)
Figure 28.2 Graphical illustrations o f the iteration methods discussed in
the text: ( a) rearrangement; ( b) linear interpolation; ( c) binary chopping;
(d) Newton–Raphson.
with n=1 .N e x t f(x1) is evaluated and the process repeated after replacing by x1
either A1orB1, according to whether f(x1) has the same sign as f(A1)o rf(B1)
respectively. In figure 28.2( b),A1is the one replaced.
As can be seen in the particular example that we are considering, with this
method there is a tendency, if the curvature of f(x)i so fc o n s t a n ts i g nn e a r
the root, for one of the two ends of the successive chords to remain un-changed.
Starting with the initial values A
1=1a n d B1=1.7, the results of the first
five iterations using (28.8) are given in table 28.2 and indicated in graph ( b)o f
figure 28.2. As with the rearrangement method, the improvement in accuracy,as measured by f(x
n)a n d xn−ξ, is a fairly constant factor at each iteration
(approximately 3 in this case), and for our particular example there is little tochoose between the two. Both tend to their limiting value of ξmonotonically,
from either higher or lower values, and this makes it difficult to estimate limits
within which ξcan safely be presumed to lie. The next method to be described
gives at any stage a range of values within which ξisknown to lie.
1153
NUMERICAL METHODS
nA n f(An) Bn f(Bn) xn f(xn)
1 1.0000 −4.0000 1.7000 5.4186 1.3500 −2.1610
2 1.3500 −2.1610 1.7000 5.4186 1.5250 0 .5968
3 1.3500 −2.1610 1.5250 0.5968 1.4375 −0.9946
4 1.4375 −0.9946 1.5250 0.5968 1.4813 −0.2573
5 1.4813 −0.2573 1.5250 0.5968 1.5031 0 .1544
6 1.4813 −0.2573 1.5031 0.1544 1.4922 −0.0552
7 1.4922 −0.0552 1.5031 0.1544 1.4977 0 .0487
8 1.4922 −0.0552 1.4977 0.0487 1.4949 −0.0085
Table 28.3 Successive approximations to the root of (28.1) using binary
chopping.
28.1.3 Binary chopping
Again two values of x,A1andB1, that straddle the root are chosen, such that
A1<B 1andf(A1)a n d f(B1) have opposite signs. The interval between them is
then halved by forming
xn=1
2(An+Bn), (28.9)
with n=1 ,a n d f(x1) is evaluated. It should be noted that x1is determined
solely by A1andB1, and not by the values of f(A1)a n d f(B1) as in the linear
interpolation method. Now x1is used to replace either A1orB1, depending on
which of f(A1)o rf(B1) has the same sign as f(x1), i.e. if f(A1)a n d f(x1) have the
same sign then x1replaces A1. The process is then repeated to obtain x2,x3etc.
This has been carried through in table 28.3 for our standard equation (28.1)
and is illustrated in figure 28.2( c). The entries have been rounded to four places
of decimals. It is suggested that the reader follows through the sequential replace-ments of the A
nandBnin the table and correlates the first few of these with
graph ( c) of figure 28.2.
Clearly the accuracy with which ξis known in this approach increases by only
a factor of 2 at each step, but this accuracy is predictable at the outset of the
calculation and (unless f(x) has very violent behaviour near x=ξ)ar a n g eo f x
in which ξlies can be safely stated at any stage. At the stage reached in the last
line of table 28.3 it may be stated that 1 .4949 <ξ< 1.4977. Thus binary chopping
gives a simple approximation method (it involves less multiplication than linearinterpolation, for example) that is predictable and relatively safe, although itsconvergence is slow.
28.1.4 Newton–Raphson method
The Newton–Raphson (NR) procedure is somewhat similar to the interpolation
method but, as will be seen, has one distinct advantage over the latter. Instead
1154
28.1 ALGEBRAIC AND TRANSCENDENTAL EQUATIONS
nx n f(xn)
1 1.7 5.42
2 1.545 01 1.033 1.498 87 7 .20×10
−2
4 1.495 13 4 .49×10−4
5 1.495 106 40 2 .6×10−8
6 1.495 106 40 –
Table 28.4 Successive approximations to the root of (28.1) using the Newton–
Raphson method.
of (notionally) constructing the chord between two points on the curve of f(x)
against x, the tangent to the curve is notionally constructed at each successive
value of xnand the next value xn+1taken as the point at which the tangent cuts
the axis f(x) = 0. This is illustrated in graph ( d) of figure 28.2.
If the nth value is xn, the tangent to the curve of f(x) at that point has slope
f/prime(xn) and passes through the point x=xn,y=f(xn). Its equation is thus
y(x)=(x−xn)f/prime(xn)+f(xn). (28.10)
The value of xat which y= 0 is then taken as xn+1; thus the condition y(xn+1)=0
yields from (28.10) the iteration scheme
xn+1=xn−f(xn)
f/prime(xn). (28.11)
This is the Newton–Raphson iteration formula . Clearly,if xnis close to ξthen xn+1
is close to xn, as it should be. It is also apparent that if any of the xncomes close
to a stationary point of f,s ot h a t f/prime(xn) is close to zero, the scheme is not going
to work well.
For our standard example, (28.11) becomes
xn+1=xn−x5
n−2x2
n−3
5x4n−4xn=4x5
n−2x2
n+3
5x4n−4xn. (28.12)
Again taking a starting value of x1=1.7 we obtain in succession the entries
in table 28.4. The different values are given to an increasing number of decimal
places as the calculation proceeds; f(xn) is also recorded.
It is apparent that this method is unlike the previous ones in that the increase
in accuracy of the answer is not constant throughout the iterations but improvesdramatically as the required root is approached. Away from the root the behaviour
of the series is less satisfactory and from its geometrical interpretation it can be
seen that if, for example, there were a maximum or minimum near the root thenthe series could oscillate between values on either side of it (instead of ‘homing
1155
NUMERICAL METHODS
in’ on the root). The reason for the good convergence near the root is discussed
in the next section.
Of the four methods mentioned, no single one is ideal and, in practice, some
mixture of them is usually to be preferred. The particular combination of methodsselected will depend a great deal on how easily the progress of the calculationmay be monitored, but some combination of the first three methods mentioned,followed by the NR scheme if great accuracy were required, would be suitable
for most situations.
28.2 Convergence of iteration schemes
For iteration schemes in which x
n+1can be expressed as a differentiable function
ofxn, e.g. the rearrangement or NR methods of the previous section, a partial
analysis of the conditions necessary for a successful scheme can be made as
follows.
Suppose the general iteration formula is expressed as
xn+1=F(xn) (28.13)
((28.7) and (28.12) are examples). Then the sequence of values x1,x2,...,x n,...is
required to converge to the value ξthat satisfies both
f(ξ)=0 a n d ξ=F(ξ). (28.14)
If the error in the solution at the nth stage is /epsilon1n,i . e .xn=ξ+/epsilon1n,t h e n
ξ+/epsilon1n+1=xn+1=F(xn)=F(ξ+/epsilon1n). (28.15)
For the iteration process to converge, a decreasing error is required, i.e. |/epsilon1n+1|<
|/epsilon1n|. To see what this implies about F, we expand the right-hand term of (28.15)
by means of a Taylor series and use (28.14) to replace (28.15) by
ξ+/epsilon1n+1=ξ+/epsilon1nF/prime(ξ)+1
2/epsilon12
nF/prime/prime(ξ)+···. (28.16)
This shows that, for small /epsilon1n,
/epsilon1n+1≈F/prime(ξ)/epsilon1n
and that a necessary (but not sufficient) condition for convergence is that
|F/prime(ξ)|<1. (28.17)
It should be noticed that this is a condition on F/prime(ξ) and not on f/prime(ξ), which
may have any finite value. Figure 28.3 illustrates in a graphical way how theconvergence proceeds for the case 0 <F
/prime(ξ)<1.
1156
28.2 CONVERGENCE OF ITERATION SCHEMES
xn xn+1xn+2y=x
y=F(x)
ξ
xy
Figure 28.3 Illustration of the convergence of the iteration scheme xn+1=
F(xn)w h e n0 <F/prime(ξ)<1, where ξ=F(ξ). The line y=xmakes an angle
π/4 with the axes. The broken line makes an angle tan−1F/prime(ξ)w i t ht h e x-axis.
Equation (28.16) suggests that if F(x) can be chosen so that F/prime(ξ) = 0 then the
ratio|/epsilon1n+1//epsilon1n|could be made very small, of order /epsilon1ni nf a c t .T og oe v e nf u r t h e r ,
if it can be arranged that the first few derivatives of Fvanish at x=ξthen the
convergence, once xnhas become close to ξ, could be very rapid indeed. If the
firstN−1 derivatives of Fvanish at x=ξ,i . e .
F/prime(ξ)=F/prime/prime(ξ)=···=F(N−1)(ξ) = 0 (28.18)
and consequently
/epsilon1n+1=O ( /epsilon1N
n), (28.19)
then the scheme is said to have Nth-order convergence .
This is the explanation of the significant difference in convergence between the
NR scheme and the others discussed (judged by reference to (28.19), so that thedifferentiability of the function Fis not a prerequisite). The NR procedure has
second-order convergence, as is shown by the following analysis. Since
F(x)=x−f(x)
f/prime(x),
F/prime(x)=1−f/prime(x)
f/prime(x)+f(x)f/prime/prime(x)
[f/prime(x)]2=f(x)f/prime/prime(x)
[f/prime(x)]2.
Now, provided f/prime(ξ)/negationslash= 0, it follows that F/prime(ξ)=0b e c a u s e f(x)=0a t x=ξ.
1157
NUMERICAL METHODS
nx n+1 /epsilon1n
18 . 5 4 . 5
2 5.191 1.193 4.137 1 .4×10
−1
4 4.002257 2 .3×10−3
5 4.000000637 6 .4×10−7
64 —
Table 28.5 Successive approximations to√
16 using the iteration scheme
(28.20).IThe following is an iteration scheme for finding the square root of X:
xn+1=1
2
/
xn+X
xn
/
. (28.20)
Show that it has second-order convergence and illustrate its efficiency by finding, say,√
16
starting with a very poor guess√
16 = 1 .
If this scheme does converge to ξthen ξwill satisfy
ξ=1
2
/
ξ+X
ξ
/
⇒ ξ2=X,
as required. The iteration function Fis given by
F(x)=1
2
/
x+X
x
/
,
and so, since ξ2=X,
F/prime(ξ)=1
2
/
1−X
x2
/
x=ξ=0,
whilst
F/prime/prime(ξ)=
/X
x3
/
x=ξ=1
ξ/negationslash=0.
Thus the procedure has second-order, but not third-order, convergence.
We now show the procedure in action. Table 28.5 gives successive values of xnand of
/epsilon1n, the difference between xnand the true value, 4.
As we can see the scheme is crude initially, but once xngets close to ξ, it homes in on
the true value extremely rapidly.
J
28.3 Simultaneous linear equations
As we saw in chapter 8, many situations in physical science can be described
approximately or exactly by a set of Nsimultaneous linear equations in N
1158
28.3 SIMULTANEOUS LINEAR EQUATIONS
variables (unknowns), xi,i=1,2,...,N . The equations take the general form
A11x1+A12x2+···+A1NxN=b1,
A21x1+A22x2+···+A2NxN=b2, (28.21)
...
AN1x1+AN2x2+···+ANNxN=bN,
where the Aijare constants and form the elements of a square matrix A.T h e bi
are given and form a column matrix b.I fAis non-singular then (28.21) can be
solved for the xiusing the inverse of A, according to the formula
x=A−1b.
This approach was discussed at length in chapter 8 and will not be considered
further here.
28.3.1 Gaussian elimination
We follow instead a continuation of one of the earliest techniques acquired by a
student of algebra, namely the solving of simultaneous equations (initially only
two in number) by the successive elimination of all the variables but one. This
(known as Gaussian elimination ) is achieved by using, at each stage, one of the
equations to obtain an explicit expression for one of the remaining xiin terms
of the others and then substituting for that xiin all other remaining equations.
Eventually a single linear equation in just one of the unknowns is obtained. Thisis then solved and the result re-substituted in previously derived equations (inreverse order) to establish values for all the x
i.
This method is probably very familiar to the reader and so a specific example
to illustrate this alone seems unnecessary. Instead, we will show how a calculationalong such lines might be arranged so that the errors due to the inherent lack ofprecision in any calculating equipment do not become excessive. This can happen
if the value of Nis large and particularly (and we will merely state this) if the
elements A
11,A22,...,A NNon the leading diagonal of the matrix in (28.21) are
small compared with the off-diagonal elements.
The process to be described is known as Gaussian elimination with interchange .
The only, but essential, difference from straightforward elimination is that before
each variable xiis eliminated, the equations are reordered to put the largest (in
modulus) remaining coefficient of xion the leading diagonal.
We will take as an illustration a straightforward three-variable example, which
can in fact be solved perfectly well without any interchange since, with simple
numbers and only two eliminations to perform, rounding errors do not havea chance to build up. However, the important thing is that the reader should
1159
NUMERICAL METHODS
appreciate how this would apply in (say) a computer program for a 1000-variable
case, perhaps with unforseeable zeroes or very small numbers appearing on theleading diagonal.ISolve the simultaneous equations
(a) x1+6x2−4x3=8,
(b) 3 x1−20x2+x3=1 2 ,
(c)−x1+3x2+5x3=3.(28.22)
Firstly, we interchange rows (a) and (b) to bring the term 3 x1onto the leading diagonal. In
the following, we label the important equations (I), (II), (III), and the others alphabetically.
(I) 3 x1−20x2+x3=1 2 ,
(d) x1+6x2−4x3=8,
(e)−x1+3x2+5x3=3.
For ( j) = (d) and (e), replace row (j) by
row ( j)−aj1
3×row (I) ,
where aj1is the coefficient of x1in row ( j), to give the two equations
(II)
/;
6+20
3
/
x2+
/;
−4−1
3
/
x3=8−12
3,
(f)
/;
3−20
3
/
x2+
/;
5+1
3
/
x3=3 +12
3.
Now|6+20
3|>|3−20
3|and so no interchange is needed before the next elimination. To
eliminate x2, replace row (f) by
row (f)−
/;
−11
3
/
38
3×row (II) .
This gives
(III)
/16
3+11
38×(−13)
3
/
x3=7+11
38×4.
Collecting together and tidying up the final equations, we have
(I) 3 x1−20x2+x3=1 2 ,
(II) 38 x2−13x3=1 2 ,
(III) x3=2.
Starting with (III) and working backwards it is now a simple matter to obtain
x1=1 0,x 2=1,x 3=2.
J
28.3.2 Gauss–Seidel iteration
In the example considered in the previous subsection an explicit way of solving
a set of simultaneous equations was given, the accuracy obtainable being limited
only by the rounding errors in the calculating facilities available, and the calcula-
tion was planned to minimise these. However, in some situations it may be thatonly an approximate solution is needed. If, for a large number of variables, this is
1160
28.3 SIMULTANEOUS LINEAR EQUATIONS
the case then an iterative method may produce a satisfactory degree of precision
with less calculation. Such a method, known as Gauss–Seidel iteration ,i sb a s e d
upon the following analysis.
The problem is again that of finding the components of the column matrix x
that satisfies
Ax=b (28.23)
when Aand bare a given matrix and column matrix respectively.
The steps of the Gauss–Seidel scheme are as follows.
(i) Rearrange the equations (usually by simple division on both sides of each
equation) so that all diagonal elements of the new matrix Care unity, i.e.
(28.23) becomes
Cx=d, (28.24)
where C=I−F,a n d Fhas zeroes as its diagonal elements.
(ii) Step (i) produces
Fx+d=Ix=x, (28.25)
and this forms the basis of an iteration scheme
xn+1=Fxn+d, (28.26)
where xnis the nth approximation to the required solution vector ξ.
(iii) To improve the convergence, the matrix F, which has zeroes on its leading
diagonal, can be written as the sum of two matrices Land Uthat have
non-zero elements only below and above the leading diagonal respectively:
Lij=braceleftBigg
Fijifi>j ,
0o t h e r w i s e ,
(28.27)
Uij=braceleftBigg
Fijifi<j ,
0o t h e r w i s e .
This allows the latest values of the components of xto be used at each
stage and an improved form of (28.26) to be obtained,
xn+1=Lxn+1+Uxn+d. (28.28)
To see why this is possible we note, for example, that when calculating,
say, the fourth component of xn+1, its first three components are already
known, and, because of the structure of L, these are the only ones needed
to evaluate the fourth component of Lxn+1.
1161
NUMERICAL METHODS
nx 1 x2 x3
12 2 2
24 0 . 1 1 . 3 43 12.76 1.381 2.323
4 9.008 0.867 1.881
5 10.321 1.042 2.0396 9.902 0.987 1.988
7 10.029 1.004 2.004
Table 28.6 Successive approximations to the solution of simultaneous equa-tions (28.29) using the Gauss–Seidel iteration method.IObtain an approximate solution to the simultaneous equations
x1+6x2−4x3=8,
3x1−20x2+x3=1 2 ,
−x1+3x2+5x3=3.(28.29)
These are the same equations as were solved in subsection 28.3.1.
Divide the equations by 1, −20 and 5 respectively to give
x1+6x2−4x3=8,
−0.15x1+x2−0.05x3=−0.6,
−0.2x1+0.6x2+x3=0.6.
Thus, set out in matrix form, (28.28) is in this case/0/@x1
x2
x3
/1
A
n+1=
/0/@00 0
0.1 500
0.2−0.60
/1A
/0/@x1
x2
x3
/1
A
n+1
+
/0/@0−64
000 .05
00 0
/1A
/0/@x1
x2
x3
/1
A
n+
/0
/@
8
−0.6
0.6
/1A.
Suppose initially ( n= 1) we guess each component to have the value 2. Then the successive
sets of values of the three quantities generated by this scheme are as shown in table 28.6.Even with the rather poor initial guess, a close approximation to the exact result x
1= 10,
x2=1 , x3= 2 is obtained in only a few iterations.
J
28.3.3 Tridiagonal matrices
Although for the solution of most matrix equations Ax=bthe number of
operations needed increases rapidly with the size N×Nof the matrix (roughly as
N3), for one particularly simple kind of matrix the computing required increases
only linearly with N. This type often occurs in physical situations in which objects
in an ordered set interact only with their nearest neighbours and is one in whichonly the leading diagonal and the diagonals immediately above and below it
1162
28.3 SIMULTANEOUS LINEAR EQUATIONS
contain non-zero entries. Such matrices are known as tridiagonal matrices. They
may also be used in numerical approximations to the solutions of certain typesof differential equation.
A typical matrix equation involving a tridiagonal matrix is thus
00b1c1
b2c2
b3c3
bNa2
a3
aN–1
aN. . .x1
x2
x3
. . .
xN–1
xNy1
y2
y3
. . .
yN–1
yN=
bN–1cN–1. . . . . .(28.30)
So as to keep the entries in the matrix as free from subscripts as possible, we
have used a,bandcto indicate subdiagonal, leading diagonal and superdiagonal
elements respectively. As a consequence we have had to change the notation forthe column matrix on the right-hand side from bto (say) y.
In such an equation the first and last rows involve x
1andxNrespectively, and
so the solution could be found by letting x1be unknown and then solving in turn
each row of the equation in terms of x1, and finally determining x1by requiring
the next-to-last line to generate for xNan equation compatible with that given by
the last line. However, if the matrix is large then this becomes a very cumbersome
operation, and a simpler method is to assume a form of solution
xi−1=θi−1xi+φi−1. (28.31)
Since the ith line of the matrix equation is
aixi−1+bixi+cixi+1=yi,
we must have, by substituting for xi−1,t h a t
(aiθi−1+bi)xi+cixi+1=yi−aiφi−1.
This is also in the form of (28.31), but with ireplaced by i+1. Thus the recurrence
formulae for θiandφiare
θi=−ci
aiθi−1+bi,φ i=yi−aiφi−1
aiθi−1+bi, (28.32)
provided the denominator does not vanish for any i. From the first of the matrix
equations it follows that θ1=−c1/b1andφ1=y1/b1. The equations may now
be solved for the xiin two stages without carrying through an unknown quantity.
First, all the θiandφiare generated using (28.32) and the values of θ1andφ1
and then, as a second stage, (28.31) is used to evaluate the xi, starting with xN
(=φN) and working backwards.
1163
NUMERICAL METHODSISolve the following tridiagonal matrix equation, in which only non-zero elements are
shown./0BBBBB/@12
−12 1
2−12
311
342
−22
/1CCCCCA
/0BBBBB/@x1
x2
x3
x4
x5
x6
/1CCCCCA=
/0BBBBB/@4
3
−3
10
7
−2
/1CCCCCA. (28.33)
The solution is set out in table 28.7, in which the arrows indicate the general flow of the
calculation. First, the columns of ai,bi,ciandyiare filled in from the original equation
(28.33) and then the recurrence relations ( 28.32) are used to fill in the successive rows
starting from the top; on each row we work from left to right as far as and including the φi
column. Finally, the bottom entry in the the xicolumn is set equal to the bottom entry in the
completed φicolumn and the rest of the xicolumn completed by using (28.31) and working
up from the bottom. Thus the solution is x1=2 ;x2=1 ;x3=3 ;x4=−1;x5=2 ;x6=1 .
J
ai bici aiθi−1+bi θi yi aiφi−1 φi xi
↓01 2 → 1−2 40 4 2↑
↓− 121 → 4−1/4 3−4 7/4 1↑
↓2−12→− 3/2 4/3 −3 7/2 13/3 3↑
↓31 1 → 5−1/510 13 −3/5−1↑
↓34 2 → 17/5−10/17 7−9/5 44/17 2↑
↓− 220 → 54/17 0 −2−88/17 1→ 1↑
Table 28.7 The solution of tridiagonal matrix equation (28.33). The arrows
indicate the general flow of the calculation, as described in the text.
28.4 Numerical integration
As noted at the start of this chapter, with modern computers and computer
packages – some of which will present solutions in algebraic form, where thatis possible – the inability to find a closed-form expression for an integral no
longer presents a problem. But, just as for the solution of algebraic equations, it
is extremely important that scientists and engineers should have some idea of theprocedures on which such packages are based. In this section we discuss some ofthe more elementary methods used to evaluate integrals numerically and at thesame time indicate the basis of more sophisticated procedures.
The standard integral evaluation has the form
I=integraldisplay
b
af(x)dx, (28.34)
where the integrand f(x) may be given in analytic or tabulated form, but for the
cases under consideration no closed-form expression for Ican be obtained. All
1164
28.4 NUMERICAL INTEGRATION
xi xi xi xi+1/2 xi+1 xi+1 xi+1 xi−1hh h hf(x) f(x)(a) (b) (c)
fififi+1
fi+1 fi+1
fi−1
Figure 28.4 ( a) Definition of nomenclature. ( b) The approximation in using
the trapezium rule; f(x) is indicated by the broken curve. ( c) Simpson’s rule
approximation; f(x) is indicated by the broken curve. The solid curve is part
of the approximating parabola.
numerical evaluations of Iare based on regarding Ias the area under the curve
off(x) between the limits x=aandx=band attempting to estimate that area.
The simplest methods of doing this involve dividing up the interval a≤x≤b
into Nequal sections, each of length h=(b−a)/N. The dividing points are
labelled xiwith x0=a,xN=b,irunning from 0 to N. The point xiis a distance
ihfrom a. The central value of xin a strip ( x=xi+h/2) is denoted for brevity
byxi+1/2, and for the same reason f(xi) is written as fi. This nomenclature is
indicated graphically in figure 28.4( a).
So that we may compare later estimates of the area under the curve with the
true value, we next obtain an exact expression for I, even though we cannot
evaluate it. To do this we need to consider only one strip, say that between xi
andxi+1. For this strip the area is, using Taylor’s expansion,
integraldisplayh/2
−h/2f(xi+1/2+y)dy=integraldisplayh/2
−h/2∞summationdisplay
n=0f(n)(xi+1/2)yn
n!dy
=∞summationdisplay
n=0f(n)
i+1/2integraldisplayh/2
−h/2yn
n!dy
=∞summationdisplay
nevenf(n)
i+1/22
(n+1 ) !parenleftbiggh
2parenrightbiggn+1
. (28.35)
It should be noticed that, in this exact expression, only the even derivatives
offsurvive the integration and all derivatives are evaluated at xi+1/2. Clearly
1165
NUMERICAL METHODS
other exact expressions are possible, e.g. the integral of f(xi+y) over the range
0≤y≤h, but we will find (28.35) the most useful for our purposes.
We now turn to practical ways of approximating I, given the values of fi,o ra
means to calculate them, for i=0,1,...,N .
28.4.1 Trapezium rule
In this simple case the area shown in figure 28.4( a) is approximated as shown in
figure 28.4( b), i.e. by a trapezium. The area Aiof the trapezium is
Ai=1
2(fi+fi+1)h, (28.36)
and if such contributions from all strips are added together then the estimate of
the total, and hence of I,i s
I(estim.) =N−1summationdisplay
i=0Ai=h
2(f0+2f1+2f2+···+2fN−1+fN). (28.37)
This provides a very simple expression for estimating integral (28.34); its accuracy
is limited only by the extent to which hcan be made very small (and hence N
very large) without making the calculation excessively long. Clearly the estimateprovided is only exact if f(x) is a linear function of x.
The error made in calculating the area of the strip when the trapezium rule is
used may be estimated as follows. The values used are f
iandfi+1, as in (28.36).
These can be expressed accurately in terms of fi+1/2and its derivatives by the
Taylor series
fi+1/2±1/2=fi+1/2±h
2f/prime
i+1/2+1
2!parenleftbiggh
2parenrightbigg2
f/prime/prime
i+1/2±1
3!parenleftbiggh
2parenrightbigg3
f(3)
i+1/2+···.
Thus
Ai(estim.) =1
2h(fi+fi+1),
=hbracketleftBigg
fi+1/2+1
2!parenleftbiggh
2parenrightbigg2
f/prime/prime
i+1/2+O ( h4)bracketrightBigg
,
whilst, from the first few terms of the exact result (28.35),
Ai(exact) = hfi+1/2+2
3!parenleftbiggh
2parenrightbigg3
f/prime/prime
i+1/2+O ( h5).
Thus the error ∆ Ai=Ai(estim.)−Ai(exact) is given by
∆Ai=parenleftbig1
8−1
24parenrightbig
h3f/prime/prime
i+1/2+O ( h5)
≈1
12h3f/prime/prime
i+1/2.
1166
28.4 NUMERICAL INTEGRATION
The total error in I(estim.) is thus given approximately by
∆I(estim.)≈1
12nh3/angbracketleftf/prime/prime/angbracketright=1
12(b−a)h2/angbracketleftf/prime/prime/angbracketright, (28.38)
where/angbracketleftf/prime/prime/angbracketrightrepresents an average value for the second derivative of fover the
interval atob.IUse the trapezium rule with h=0.5to evaluate
I=
Z2
0(x2−3x+4 )dx,
and, by evaluating the integral exactly, examine how well (28.38) estimates the error.
With h=0.5, we will need five values of f(x)=x2−3x+ 4 for use in formula (28.37).
They are f(0) = 4, f(0.5) = 2 .75,f(1) = 2, f(1.5) = 1 .75 and f(2) = 2. Putting these into
(28.37) gives
I(estim.) =0.5
2(4 + 2×2.75 + 2×2+2×1.75 + 2) = 4 .75.
The exact value is
I(exact) =
/x3
3−3x2
2+4x
/2
0=42
3.
The difference between the estimate of the integral and the exact answer is 1 /12. Equation
(28.38) estimates this error as 2 ×0.25×/angbracketleftf/prime/prime/angbracketright/12. Our (deliberately chosen!) integrand is
one for which /angbracketleftf/prime/prime/angbracketrightcan be evaluated trivially. Because f(x) is a quadratic function of x,
its second derivative is constant, and equal to 2 in this case. Thus /angbracketleftf/prime/prime/angbracketrighthas value 2 and
(28.38) estimates the error as 1 /12; that the estimate is exactly right should be no surprise
since the Taylor expansion for a quadratic polynomial about any point always terminatesafter three terms and so no higher-order terms in hhave been ignored in (28.38).J
28.4.2 Simpson’s rule
Whereas the trapezium rule makes a linear interpolation of f, Simpson’s rule
effectively mimics the local variation of f(x) using parabolas. The strips are
treated two at a time (figure 28.4( c)) and therefore their number, N, should be
made even.
In the neighbourhood of xi,f o riodd, it is supposed that f(x) can be adequately
represented by a quadratic form,
f(xi+y)=fi+ay+by2. (28.39)
In particular, applying this to y=±hyields two expressions involving b,
fi+1=f(xi+h)=fi+ah+bh2,
fi−1=f(xi−h)=fi−ah+bh2;
thus
bh2=1
2(fi+1+fi−1−2fi).
1167
NUMERICAL METHODS
Now, in the representation (28.39), the area of the double strip from xi−1to
xi+1is given by
Ai(estim.) =integraldisplayh
−h(fi+ay+by2)dy=2hfi+2
3bh3.
Substituting for bh2then yields for the estimated area
Ai(estim.) = 2 hfi+2
3h×1
2(fi+1+fi−1−2fi)
=1
3h(4fi+fi+1+fi−1),
an expression involving only given quantities. It should be noted that the values
of neither bnoraneed be calculated.
For the full integral
I(estim.) =1
3h(f0+fN+4summationdisplay
moddfm+2summationdisplay
mevenfm). (28.40)
It can be shown, by following the same procedure as in the trapezium rule case,
that the error in the estimated area is approximately
∆I(estim.)≈(b−a)
180h4/angbracketleftf(4)/angbracketright.
28.4.3 Gaussian integration
In the cases considered in the previous two subsections, the function fwas
mimicked by linear and quadratic functions. These yield exact answers if f
itself is a linear or quadratic function (respectively) of x. This process could
be continued by increasing the order of the polynomial mimicking-function soas to increase the accuracy with which more complicated functions fcould be
numerically integrated; but the same effect can be achieved with less effort by
not insisting upon equally spaced points x
i.
The detailed analysis of such methods of numerical integration, in which the
integration points are not equally spaced and the weightings given to the values
at each point do not fall into a few simple groups, is too long to be given here.The reader is referred to books devoted specifically to the theory of numericalanalysis, where details of the integration points and weights for many schemeswill be found. †
We will content ourselves here with describing Gaussian integration, which
is based upon the orthogonality properties, in the interval −1≤x≤1, of the
Legendre polynomials P
/lscript(x), discussed in subsection 16.6.2. In order to use these
properties, the integral between limits aandbin (28.34) has to be changed to
†The points and weights may be found in, e.g. Abramowitz and Stegun, Handbook of Mathematical
Functions (Dover, 1965).
1168
28.4 NUMERICAL INTEGRATION
one between the limits −1 and +1. This is easily done with a change of variable
from xtozgiven by
z=2x−b−a
b−a,
so that Ibecomes
I=b−a
2integraldisplay1
−1g(z)dz, (28.41)
in which g(z)≡f(x).
Thenintegration points xifor an n-point Gaussian integration are given by
the zeroes of Pn(x), i.e. the xiare such that Pn(xi) = 0. The integrand g(x)i s
mimicked by the ( n−1)th-degree polynomial
G(x)=nsummationdisplay
i=1Pn(x)
(x−xi)P/primen(xi)g(xi),
which coincides with g(x) at each of the points xi,i=1,2,...,n. To see this it
should be noted that
lim
x→xkPn(x)
(x−xi)P/primen(xi)=δik.
It then follows, to the extent that g(x) is well reproduced by G(x), that
integraldisplay1
−1g(x)dx≈nsummationdisplay
i=1g(xi)
P/primen(xi)integraldisplay1
−1Pn(x)
x−xidx. (28.42)
The expression
w(xi)≡1
P/primen(xi)integraldisplay1
−1Pn(x)
x−xidx
can be shown, using the properties of Legendre polynomials, to be equal to
wi=2
(1−x2
i)|P/primen(xi)|2,
and is thus the weighting to be attached to the factor g(xi) in the sum (28.42),
which becomes
integraldisplay1
−1g(x)dx≈nsummationdisplay
i=1wig(xi). (28.43)
In fact, because of the particular properties of Legendre polynomials, it can be
shown that (28.43) integrates exactly any polynomial of degree up to 2 n−1. The
error in the approximate equality is of the order of the 2 nth derivative of gand
so, provided g(x) is a reasonably smooth function, the approximation is a good
one.
1169
NUMERICAL METHODS
As an example, for a three-point integration, the three xiare the zeroes of
P3(x)=1
2(5x3−3x), namely 0 and ±0.774 60, and the corresponding weights are
2
1×parenleftbig
−3
2parenrightbig2=8
9and2
(1−0.6)×parenleftbig6
2parenrightbig2=5
9.
For other forms of integrand, formulae based on other sets of orthogonal
functions give better results. For example, integrals over finite ranges involvingfactors of the form (1 −x
2)±1/2in the integrand are best treated using formulae
based on Chebyshev polynomials, whilst infinite integrals containing e−x(0≤
x<∞)o re−x2(−∞<x<∞) are best handled using schemes based on Laguerre
or Hermite polynomials respectively.IUsing a three-point formula in each case, evaluate the integral
I=
Z1
01
1+x2dx,
(i) using the trapezium rule, (ii) using Simpson’s rule, (iii) using Gaussian integration.
Also evaluate the integral analytically and compare the results.
(i) Using the trapezium rule, we obtain
I=1
2×1
2
/
f(0) + 2 f
/;1
2
/
+f(1)
/
=1
4
/
1+8
5+1
2
/
=0.7750.
(ii) Using Simpson’s rule, we obtain
I=1
3×1
2
/
f(0) + 4 f
/;1
2
/
+f(1)
/
=1
6
/
1+16
5+1
2
/
=0.7833.
(iii) Using Gaussian integration, we obtain
I=1−0
2
Z1
−1dz
1+1
4(z+1 )2
=1
2
n
0.55556 [f(−0.77460) + f(0.77460) ]+0.88889 f(0)
o
=1
2
n
0.55556 [0.987458 + 0 .559503 ]+0.88889×0.8
o
=0.78527 .
(iv) Exact evaluation gives
I=
Z1
0dx
1+x2=
/
tan−1x
/1
0=π
4=0.78540 .
In practice, a compromise has to be struck between the accuracy of the result achieved
and the calculational labour that goes into obtaining it.
J
28.4.4 Monte Carlo methods
Surprising as it may at first seem, random numbers may be used to carry out
numerical integration. The random element comes in principally when selecting
1170
28.4 NUMERICAL INTEGRATION
the points at which the integrand is evaluated, and naturally does not extend to
the actual values of the integrand!
For the most part we will continue to use as our model one-dimensional
integrals between finite limits, as typified by equation (28.34). Extensions to cover
infinite or multidimensional integrals will be indicated briefly at the end of thesection. It should be noted here, however, that Monte Carlo methods – the namehas become attached to methods based on randomly generated numbers – inmany ways come into their own when used on multidimensional integrals overregions with complicated boundaries.
It goes without saying that in order to use random numbers for calculational
purposes a supply of them must be available. There was a time when theywere provided in book form as a two-dimensional array of random digits inthe range 0 to 9, and the user could generate a random number of any desiredlength by selecting the positions in the table of its successive digits in anypredetermined and systematic way. Nowadays all computers and nearly all pocket
calculators offer a function which supplies a sequence of decimal numbers ξthat,
for all practical purposes, are randomly and uniformly chosen in the range0≤ξ<1. The maximum number of significant figures available in each random
number depends on the precision of the generating device. We will defer thedetails of how these numbers are produced to a later subsection, where it willalso be shown how random numbers distributed in a prescribed way can begenerated.
All integrals of the general form shown in equation (28.34) can, by a suitable
change of variable, be brought to the form
θ=integraldisplay
1
0f(x)dx, (28.44)
and we will use this as our standard model.
All approaches to integral evaluation based on random numbers proceed by
estimating a quantity whose expectation value is equal to the sought-for value θ.
The estimator tmust be unbiased, i.e. we must have E[t]=θ, and the method
must provide some measure of the likely error in the result. The latter will appeargenerally as the variance of the estimate, with its usual statistical interpretation,and not as a band in which the true answer is known to lie with certainty.
The various approaches really differ from each other only in the degree of
sophistication employed to keep the variance of the estimate of θsmall. The
overall efficiency of any particular method has to take into account not only thevariance of the estimate but also the computing and book-keeping effort requiredto achieve it.
We do not have the space to describe even the more elementary methods in
full detail, but the main thrust of each approach should be apparent to the readerfrom the brief descriptions that follow.
1171
NUMERICAL METHODS
Crude Monte Carlo
The most straightforward application is one in which the random numbers areused to pick sample points at which f(x) is evaluated. These values are then
averaged:
t=1
nnsummationdisplay
i=1f(ξi). (28.45)
Stratified sampling
Here the range of xis broken up into ksubranges,
0=α0<α1<···<α k=1,
and crude Monte Carlo evaluation is carried out in each subrange. The estimate
E[t] is then calculated as
E[t]=ksummationdisplay
j=1njsummationdisplay
i=1αj−αj−1
njfparenleftbig
αj−1+ξij(αj−αj−1)parenrightbig
. (28.46)
This is an unbiased estimator of θwith variance
σ2
t=ksummationdisplay
j=1αj−αj−1
njintegraldisplayαj
αj−1[f(x)]2dx−ksummationdisplay
j=11
njbracketleftBiggintegraldisplayαj
αj−1f(x)dxbracketrightBigg2
.
This variance can be made less than that for crude Monte Carlo, whilst using
the same total number of random numbers, n=summationtextnj, if the differences between
the average values of f(x) in the various subranges are significantly greater than
the variations in fwithin each subrange. It is easier administratively to make all
subranges equal in length but better, if it can be managed, to make them suchthat the variations in fare approximately equal in all the individual subranges.
Importance sampling
Although we cannot integrate f(x) analytically – we would not be using Monte
Carlo methods if we could – if we can find another function g(x)t h a t canbe
integrated analytically and mimics the shape of fthen the variance in the estimate
ofθcan be reduced significantly compared with that resulting from the use of
crude Monte Carlo evaluation.
Firstly, if necessary the function gmust be renormalised, so that G(x)=integraltext
x
0g(y)dyhas the property G(1) = 1. Clearly, it also has the property G(0) = 0.
Then, since
θ=integraldisplay1
0f(x)
g(x)dG(x),
it follows that finding the expectation value of f(η)/g(η) using a random number
η, distributed in such a way that ξ=G(η) is uniformly distributed on (0 ,1), is
equivalent to estimating θ. This involves being able to find the inverse function
1172
28.4 NUMERICAL INTEGRATION
ofG; a discussion of how to do this is given in a later subsection. If g(η) mimics
f(η) well, f(η)/g(η) will be nearly constant and the estimation will have a very
small variance. Further, any error in inverting the relationship between ηandξ
will not be important since f(η)/g(η) will be largely independent of the value
ofη.
As an example, consider the function f(x)=[ t a n−1(x)]1/2, which is not analyti-
cally integrable over the range (0 ,1) but is well mimicked by the easily-integrated
function g(x)=x1/2(1−x2/6). The ratio of the two varies from 1 .00 to 1 .06 as x
varies from 0 to 1. The integral of gover this range is 0 .619 048, and so it has to
be renormalised by the factor 1 .615 38. The value of the integral of f(x)f r o m0
to 1 can then be estimated by averaging the value of
[tan−1(η)]1/2
1.615 38 η1/2(1−1
6η2)
for random variables ηwhich are such that G(η) is uniformly distributed on
(0,1). Using batches of as few as 10 random numbers gave a value 0 .630 for θ,
with standard deviation 0 .003. The corresponding result for crude Monte Carlo,
using the same random numbers, was 0 .634±0.065. The increase in precision is
obvious, though the additional labour involved would not be justified for a single
application.
Control variates
The control-variate method is similar to, but not the same as, importance sam-pling. Again, an analytically integrable function that mimics f(x) in shape has
to be found. The function, known as the control variate, is first scaled so as tomatch fas closely as possible in magnitude and then its integral is found in
closed form. If we denote the scaled control variate by h(x) then the estimate of
θis computed as
t=integraldisplay
1
0[f(x)−h(x)]dx+integraldisplay1
0h(x)dx. (28.47)
The first integral in (28.47) is evaluated using (crude) Monte Carlo, whilst the
second is known analytically. Although the first integral should have been ren-dered small by the choice of h(x), it is its variance that matters. The method relies
on the result (see equation (26.136))
V[t−t
/prime]=V[t]+V[t/prime]−2C o v [ t, t/prime]
a n do nt h ef a c tt h a ti f testimates θwhilst t/primeestimates θ/primeusing the same random
numbers then the covariance of tandt/primecan be larger than the variance of t/prime,a n d
indeed will be so if the integrands producing θandθ/primeare highly correlated.
To evaluate the same integral as was estimated previously using importance
sampling, we take as h(x) the function g(x) used there, before it was renormalised.
Again using batches of 10 random numbers, the estimated value for θwas found
1173
NUMERICAL METHODS
to be 0 .629±0.004, a result almost identical to that obtained using importance
sampling, in both value and precision. Since we knew already that f(x)a n d g(x)
diverge monotonically by about 6% as xvaries over the range (0 ,1), we could
have made a small improvement to our control variate by scaling it by 1 .03 before
using it in equation (28.47).
Antithetic variates
As a final example of a method that improves on crude Monte Carlo, and one thatis particularly useful when monotonic functions are to be integrated, we mentionthe use of antithetic variates. This method relies on finding two estimates tand
t
/primeofθthat are strongly anticorrelated (i.e. Cov[ t, t/prime] is large and negative) and
using the result
V[1
2(t+t/prime)] =1
4V[t]+1
4V[t/prime]+1
2Cov[t, t/prime].
For example, the use of1
2[f(ξ)+f(1−ξ)] instead of f(ξ) involves only twice
as many evaluations of f, and no more random variables, but generally gives
an improvement in precision significantly greater than this. For the integral of
f(x)=[ t a n−1(x)]1/2, using as previously a batch of 10 random variables, an
estimate of 0 .623±0.018 was found. This to be compared with the crude Monte
Carlo result, 0 .634±0.065, obtained using the same number of random variables.
For a fuller discussion of these methods, and of theoretical estimates of their
efficiencies, the reader is referred to more specialist treatments. For practicalimplementation schemes, a book dedicated to scientific computing should beconsulted.†
Hit or miss method
We now come to the approach that, in spirit, is closest to the activities that gave
Monte Carlo methods their name. In this approach, one or more straightforwardyes/no decisions are made on the basis of numbers drawn at random – the endresult of each trial is either a hit or a miss! In this section we are concernedwith numerical integration, but the general Monte Carlo approach, in whichone estimates a physical quantity that is hard or impossible to calculate directlyby simulating the physical processes that determine it, is widespread in modern
science. For example, the calculation of the efficiencies of detector arrays in
experiments to study elementary particle interactions are nearly always carriedout in this way. Indeed, in a normal experiment, far more simulated interactionsare generated in computers than ever actually occur when the experiment istaking real data.
As was noted in chapter 2, the process of evaluating a one-dimensional integralintegraltext
b
af(x)dxcan be regarded as that of finding the area between the curve y=f(x)
†e.g.,Numerical Recipes ,W .H .P r e s s et al. (Cambridge University Press).
1174
28.4 NUMERICAL INTEGRATION
xy=f(x)
y=c
x=ax =b
Figure 28.5 A simple rectangular figure enclosing the area (shown shaded)
which is equal to
Rb
af(x)dx.
and the x-axis in the range a≤x≤b. It may not be possible to do this
analytically but if, as shown in figure 28.5, we can enclose the curve in a simplefigure whose area can be found trivially then the ratio of the required area (shown
shaded) to that of the bounding figure, c(b−a), is the same as the probability
that a randomly selected point inside the boundary will lie below the line.
In order to accommodate cases in which f(x) can be negative in part of the
x-range, we treat a slightly more general case. Suppose that, for a≤x≤b,f(x)
is bounded and known to lie in the range A≤f(x)≤B, then the transformation
z=x−a
b−a
will reduce the integralintegraltextb
af(x)dxto the form
A(b−a)+(B−A)(b−a)integraldisplay1
0h(z)dz, (28.48)
where
h(z)=1
B−A[f((b−a)z+a)−A].
In this form zlies in the range 0 ≤z≤1a n d h(z) lies in the range 0 ≤h(z)≤1,
i.e. both are suitable for simulation using the standard random-number generator.It should be noted that for an efficient estimation the bounds AandBshould
be drawn as tightly as possible – preferably, but not necessarily, they should be
equal to the minimum and maximum values of fin the range. The reason for
this is that random numbers corresponding to values which f(x) cannot reach
add nothing to the estimation but do increase its variance.
It only remains to estimate the final integral on the RHS of equation (28.48).
This we do by selecting pairs of random numbers ξ
1andξ2and testing whether
1175
NUMERICAL METHODS
h(ξ1)>ξ2. The fraction of times that this inequality is satisfied estimates the value
of the integral (without the scaling factors ( B−A)(b−a)) since the expectation
value of this fraction is the ratio of the area below the curve y=h(z)t ot h ea r e a
of a unit square.
To illustrate the evaluation of multiple integrals using Monte Carlo techniques,
consider the relatively elementary problem of finding the volume of an irregularsolid bounded by planes, say an octahedron. In order to keep the descriptionbrief, but at the same time illustrate the general principles involved, let us supposethat the octahedron has two vertices on each of the three Cartesian axes, one on
either side of the origin for each axis. Denote those on the x-axis by x
1(<0) and
x2(>0), and similarly for the y-a n d z-axes. Then the whole of the octahedron
can be enclosed by the rectangular parallelepiped
x1≤x≤x2,y 1≤y≤y2,z 1≤z≤z2.
Any point in the octahedron lies inside or on the parallelepiped, but any point
in the parallelepiped may or may not lie inside the octahedron.
The equation of the plane containing the three vertex points ( xi,0,0),(0,yj,0)
and (0 ,0,zk)i s
x
xi+y
yj+z
zk=1 f o r i, j, k=1,2, (28.49)
and the condition that any general point ( x, y, z) lies on the same side of the
plane as the origin is that
x
xi+y
yj+z
zk−1≤0. (28.50)
For the point to be inside or on the octahedron, equation (28.50) must therefore
be satisfied for all eight of the sets of i, jandkgiven in (28.49).
Thus an estimate of the volume of the octahedron can be made by generating
random numbers ξfrom the usual uniform distribution and then using them in
sets of three, according to the following scheme.
With integer mlabelling the mth set of three random numbers, calculate
x=x1+ξ3m−2(x2−x1),
y=y1+ξ3m−1(y2−y1),
z=z1+ξ3m(z2−z1).
Define a variable nmas 1 if (28.50) is satisfied for all eight combinations of i, j, k
values and as 0 otherwise. The volume Vcan then be estimated using 3 Mrandom
numbers from the formula
V
(x2−x1)(y2−y1)(z2−z1)=1
MMsummationdisplay
m=1nm.
1176
28.4 NUMERICAL INTEGRATION
It will be seen that, by replacing each nmin the summation by f(x, y, z)nm,t h i s
procedure could be extended to estimating the integral of the function fover
the volume of the solid. The method has special value if fis too complicated to
have analytic integrals with respect to x, yandzor if the limits of any of these
integrals are determined by anything other than the simplest combinations of the
other variables. If large values of fare known to be concentrated in particular
regions of the integration volume then some form of stratified sampling shouldbe used.
It will be apparent that this general method can be extended to integrals
of general functions, bounded but not necessarily continuous, over volumes withcomplicated bounding surfaces and, if appropriate, in more than three dimensions.
Random number generation
Earlier in this subsection we showed how to evaluate integrals using sequences of
numbers that we took to be distributed uniformly on the interval 0 ≤ξ<1. In
reality the sequence of numbers is not truly random, since each is generated in amechanistic way from its predecessor and eventually the sequence repeats itself.However, the cycle is so long that in practice this is unlikely to be a problem,and the reproducibility of the sequence can even be turned to advantage whenchecking the accuracy of the rest of a calculational program. Much research has
gone into the best ways to produce such ‘pseudo-random’ sequences of numbers.
We do not have space to pursue them here and will limit ourselves to one recipethat works well in practice.
Given any particular starting (integer) value x
0, the following algorithm will
generate a full cycle of mvalues for ξi, uniformly distributed on 0 ≤ξi<1,
before repeats appear:
xi=axi−1+c(mod m); ξi=xi
m.
Here cis an odd integer and ahas the form a=4k+1w i t h kan integer. For
practical reasons, in computers and calculators mis taken as a (fairly high) power
of 2, typically 32.
The uniform distribution can be used to generate random numbers ydistributed
according to a more general probability distribution f(y) on the range a≤y≤b
if the inverse of the indefinite integral of fcan be found, either analytically or by
means of a look-up table. In other words, if
F(y)=integraldisplayy
af(t)dt,
for which F(a)=0a n d F(b)=1t h e n F(y) is uniformly distributed on (0 ,1). This
approach is not limited to finite aandb;acould be−∞andbcould be∞.
The procedure is thus to select a random number ξfrom a uniform distribution
1177
NUMERICAL METHODS
on (0 ,1) and then take as the random number ythe value of F−1(ξ). We now
illustrate this with a worked example.IFind an explicit formula that will generate a random number ydistributed on (−∞,∞)
according to the Cauchy distribution
f(y)dy=
/a
π
/dy
a2+y2,
given a random number ξuniformly distributed on (0,1).
The first task is to determine the indefinite integral
F(y)=
Zy
−∞
/a
π
/dt
a2+t2=1
πtan−1y
a+1
2.
Now, if yis distributed as we wish then F(y) is uniformly distributed on (0 ,1). This follows
from the fact that the derivative of F(y)i sf(y). We therefore set F(y)e q u a lt o ξand
obtain
ξ=1
πtan−1y
a+1
2,
yielding
y=atan[π(ξ−1
2)].
This explicit formula shows how to change a random number ξdrawn from a population
uniformly distributed on (0 ,1) into a random number ydistributed according to the
Cauchy distribution.
J
Look-up tables operate as described below for cumulative distributions F(y)
that are non-invertible, i.e. F−1(y) cannot be expressed in closed form. They
are especially useful if many random numbers are needed but great samplingaccuracy is not essential. The method for an N-entry table can be summarised as
follows.
Define w
mbyF(wm)=m/Nform=1,2,...,N , and store a table of
y(m)=1
2(wm+wm−1).
As each random number yis needed, calculate kas the integral part of Nξand
take yas given by y(k).
Normally, such a look-up table would have to be used for generating random
numbers with a Gaussian distribution, as the cumulative integral of a Gaussian isnon-invertible. It would be in essence table 26.3, with the roles of argument andvalue interchanged. In this particular case an alternative, based on the centrallimit theorem, can be considered.
With ξ
igenerated in the usual way, i.e. uniformly distibuted on the interval
0≤ξ<1, the random variable
y=nsummationdisplay
i=1ξi−1
2n (28.51)
is normally distributed with mean 0 and variance n/12 when nis large. This
1178
28.5 FINITE DIFFERENCES
approach does produce a continuous spectrum of possible values for y, but needs
many values of ξifor each value of yand is a very poor approximation if the
wings of the Gaussian distribution have to be sampled accurately. For nearly allpractical purposes a Gaussian look-up table is to be preferred.
28.5 Finite differences
It will have been noticed that earlier sections included several equations linking
sequential values of f
iand the derivatives of fevaluated at one of the xi.I n
this section, by way of preparation for the numerical treatment of differentialequations, we establish these relationships in a more systematic way.
Again we consider a set of values f
iof a function f(x)e v a l u a t e da te q u a l l y
spaced points xi, their separation being h. As before, the basis for our discussion
will be a Taylor series expansion, but on this occasion about the point xi:
fi±1=fi±hf/prime
i+h2
2!f/prime/prime
i±h3
3!f(3)
i+···. (28.52)
In this section, and subsequently, we denote the nth derivative evaluated at xi
byf(n)
i.
From (28.52), three different expressions that approximate f(1)
ican be derived.
The first of these, obtained by subtracting the ±equations, is
f(1)
i≡parenleftbiggdf
dxparenrightbigg
xi=fi+1−fi−1
2h−h2
3!f(3)
i−···. (28.53)
The quantity ( fi+1−fi−1)/(2h) is known as the central difference approximation
tof(1)
iand can be seen from (28.53) to be in error by approximately ( h2/6)f(3)
i.
An alternative approximation, obtained from (28.52+) alone, is given by
f(1)
i≡parenleftbiggdf
dxparenrightbigg
xi=fi+1−fi
h−h
2!f(2)
i−···. (28.54)
Theforward difference approximation, ( fi+1−fi)/h, is clearly a poorer approxima-
tion, since it is in error by approximately ( h/2)f(2)
i, as compared with ( h2/6)f(3)
i.
Similarly, the backward difference ( fi−fi−1)/hobtained from (28.52 −) is not as
good as the central difference; the sign of the error is reversed in this case.
This type of differencing approximation can be continued to the higher deriva-
tives of fin an obvious manner. By adding the two equations (28.52 ±), a central
difference approximation to f(2)
ican be obtained:
f(2)
i≡parenleftbiggd2f
dx2parenrightbigg
≈fi+1−2fi+fi−1
h2. (28.55)
The error in this approximation (also known as the second difference of f)i s
e a s i l ys h o w nt ob ea b o u t( h2/12)f(4)
i.
Of course, if the function f(x) is a sufficiently simple polynomial in x,a l l
1179
NUMERICAL METHODS
derivatives beyond a particular one will vanish and there is no error in taking
the differences to obtain the derivatives.IThe following is copied from the tabulation of a second-degree polynomial f(x)at values
ofxfrom1to12inclusive,
2,2,?,8,14,22,32,46,?,74,92,112.
The entries marked ?were illegible and in addition one er ror was made in transcription.
Complete and correct the table. Would your procedure have worked if the copying errorhad been in f(6)?
Write out the entries again in row (a) below, and where possible calculate first differences
in row (b) and second differences in row (c). Denote the jth entry in row ( n)b y( n)j.
(a) 2 2 ? 8 14 22 32 46 ? 74 92 112
(b) 0 ? ? 6 8 10 14 ? ? 18 20(c) ? ? ? 2 2 4 ? ? ? 2
Because the polynomial is second-degree the second differences (c)
j, which are proportional
tod2f/dx2, should be constant, and clearly the constant should be 2. That is, (c) 6should
equal 2 and (b) 7should equal 12 (not 14). Since all the (c) j= 2, we can conclude that
(b)2=2 ,( b ) 3=4 ,( b ) 8= 14, and (b) 9= 16. Working these changes back to row (a) shows
that (a) 3=4 ,( a ) 8= 44 (not 46), and (a) 9= 58.
The entries therefore should read
(a) 2,2,4,8,14,22,32,44,58,74,92,112,
where the amended entries are shown in bold type.
It is easily verified that if the error were in f(6) no two computable entries in row (c)
would be equal, and it would not be clear what the correct common entry should be.Nevertheless, trial and error might arrive at a self-consistent scheme.J
28.6 Differential equations
For the remaining sections of this chapter our attention will be on the solution
of differential equations by numerical methods. Some of the general difficultiesof applying numerical methods to differentia l equations will be all too apparent.
Initially we consider only the simplest kind of equation – one of first order,
typically represented by
dy
dx=f(x, y), (28.56)
where yis taken as the dependent variable and xthe independent one. If this
equation can be solved analytically then that is the best course to adopt. Butsometimes it is not possible to do so and a numerical approach becomes theonly one available. In fact, most of the examples that we will use can be solved
easily by an explicit integration, but, for the purposes of illustration, this is an
advantage rather than the reverse since useful comparisons can then be madebetween the numerically derived solution and the exact one.
1180
28.6 DIFFERENTIAL EQUATIONS
x h y(exact)
0.01 0.1 0.5 1.0 1.5 2 3
0 (1) (1) (1) (1) (1) (1) (1) (1)
0.5 0.605 0.590 0.500 0 −0.500−1−20.607
1.0 0.366 0.349 0.250 0 0 .250 1 4 0.368
1.5 0.221 0.206 0.125 0 −0.125−1−80.223
2.0 0.134 0.122 0.063 0 0 .063 1 16 0.135
2.5 0.081 0.072 0.032 0 −0.032−1−32 0.082
3.0 0.049 0.042 0.016 0 0 .016 1 64 0.050
Table 28.8 The solution yof differential equation (28.57) using the Euler
forward difference method for various values of h. The exact solution is also
shown.
28.6.1 Difference equations
Consider the differential equation
dy
dx=−y, y (0) = 1 , (28.57)
and the possibility of solving it numerically by approximating dy/dx by a finite
difference along the lines indicated in section 28.5. We start with the forwarddifference
parenleftbiggdy
dxparenrightbigg
xi≈yi+1−yi
h, (28.58)
where we use the notation of section 28.5 but with freplaced by y.I nt h i s
particular case, it leads to the recurrence relation
yi+1=yi+hparenleftbiggdy
dxparenrightbigg
i=yi−hyi=( 1−h)yi. (28.59)
Thus, since y0=y(0) = 1 is given, y1=y(0 +h)=y(h) can be calculated, and so
on (this is the Euler method). Table 28.8 shows the values of y(x) obtained if this
is done using various values of hand for selected values of x. The exact solution,
y(x)=e x p (−x), is also shown.
It is clear that to maintain anything like a reasonable accuracy only very small
steps hcan be used. Indeed, if his taken to be too large, not only is the accuracy
bad but, as can be seen, for h>1 the calculated solution oscillates (when it
should be monotonic) and for h>2 it diverges. Equation (28.59) is of the form
yi+1=λyi, and a necessary condition for non-divergence is |λ|<1, i.e. 0 <h< 2,
though in no way does this ensure accuracy.
Part of this difficulty arises from the poor approximation (28.58); its right-
hand side is a closer approximation to dy/dx evaluated at x=xi+h/2t h a nt o
dy/dx atx=xi. This is the result of using a forward difference rather than the
1181
NUMERICAL METHODS
xy (estim.) y(exact)
−0.5 (1.648) –
0 (1.000) (1.000)
0.5 0.648 0.607
1.0 0.352 0.368
1.5 0.296 0.223
2.0 0.056 0.135
2.5 0.240 0.082
3.0−0.184 0.050
Table 28.9 The solution of differential equation (28.57) using the Milne
central difference method with h=0.5 and accurate starting values.
more accurate, but of course still approximate, central difference. A more accurate
method based on central differences ( Milne’s method ) gives the recurrence relation
yi+1=yi−1+2hparenleftbiggdy
dxparenrightbigg
i(28.60)
in general and, in this particular case,
yi+1=yi−1−2hyi. (28.61)
An additional difficulty now arises, since two initial values of yare needed.
The second must be estimated by other means (e.g. by using a Taylor series,as discussed later) but for illustration purposes we will take the accurate value,y(−h)=e x p h, as the value of y
−1.I fhis taken as, say, 0.5 and (28.61) applied
repeatedly then the results shown in table 28.9 are obtained.
Although some improvement in the early values of the calculated y(x)i s
noticeable, as compared with the corresponding ( h=0.5) column of table 28.8,
this scheme soon runs into difficulties, as is obvious from the last two rows of thetable.
Some part of this poor performance is not really attributable to the approxi-
mations made in estimating dy/dx but to the form of the equation itself and
hence of its solution. Anyrounding error occurring in the evaluation effectively
introduces into ysome contamination by the solution of
dy
dx=+y.
This equation has the solution y(x)=e x p xand so grows without limit; ultimately
it will dominate the sought-for solution and thus render the calculations totally
inaccurate.
We have only illustrated, rather than analysed, some of the difficulties associated
with simple finite-difference iteration schemes for first-order differential equations,
1182
28.6 DIFFERENTIAL EQUATIONS
but they may be summarised as (i) insufficiently precise approximations to the
derivatives and (ii) inherent instability due to rounding errors.
28.6.2 Taylor series solutions
Since a Taylor series expansion is exact if all its terms are included, and the limits
of convergence are not exceeded, we may seek to use one to evaluate y1,y2etc.
for an equation
dy
dx=f(x, y), (28.62)
when the initial value y(x0)=y0is given.
The Taylor series is
y(x+h)=y(x)+hy/prime(x)+h2
2!y/prime/prime(x)+h3
3!y(3)(x)+···. (28.63)
In the present notation, at the point x=xithis is written
yi+1=yi+hy(1)
i+h2
2!y(2)
i+h3
3!y(3)
i+···. (28.64)
But, for the required solution y(x), we know that
y(1)
i≡parenleftbiggdy
dxparenrightbigg
xi=f(xi,yi) (28.65)
and the value of the second derivative at x=xi,y=yican be obtained from it:
y(2)
i=∂f
∂x+∂f
∂ydy
dx=∂f
∂x+f∂f
∂y. (28.66)
This process can be continued for the third and higher derivatives, all of which
are to be evaluated at ( xi,yi).
Having obtained expressions for the derivatives y(n)
iin (28.63), two alternative
ways of proceeding are open:
(i) equation (28.64) is used to evaluate yi+1and then the whole process is
repeated to obtain yi+2a n ds oo n ;
(ii) equation (28.64) is applied several times but using a different value of h
each time, and so the corresponding values of y(x+h) are obtained.
It is clear that, on the one hand, approach (i) does not require so many terms of
(28.63) to be kept but, on the other hand, the yi(n) have to be recalculated at each
step. With approach (ii), fairly accurate results for ymay be obtained for values
ofxclose to the given starting value, but for large values of ha large number
of terms of (28.63) must be kept. As an example of approach (ii) we solve thefollowing problem.
1183
NUMERICAL METHODS
xy (estim.) y(exact)
0 1.0000 1.0000
0.1 1.2346 1.23460.2 1.5619 1.5625
0.3 2.0331 2.0408
0.4 2.7254 2.77780.5 3.7500 4.0000
Table 28.10 The solution of differential equation (28.67) using a Taylor series.IFind the numerical solution of the equation
dy
dx=2y3/2,y (0) = 1 , (28.67)
forx=0.1to0.5in steps of 0.1. Compare it with the exact solution obtained analytically.
Since the right-hand side of the equation does not contain xexplicitly, (28.66) is greatly
simplified and the calculation becomes a repeated application of
y(n+1)
i=∂y(n)
∂ydy
dx=f∂y(n)
∂y.
The necessary derivatives and their values at x=0 ,w h e r e y= 1, are given below.
y(0) = 1 1
y/prime=2y3/22
y/prime/prime=( 3/2)(2y1/2)(2y3/2)=6 y26
y(3)=( 1 2 y)2y3/2=2 4y5/224
y(4)=( 6 0 y3/2)2y3/2= 120 y3120
y(5)= (360 y2)2y3/2= 720 y7/2720
Thus the Taylor expansion of the solution about the origin (in fact a Maclaurin series) is
y(x)=1+2 x+6
2!x2+24
3!x3+120
4!x4+720
5!x5+···.
Hence, y(estim.) = 1 + 2 x+3x2+4x3+5x4+6x5. Values calculated from this are given
in table 28.10. Comparison with the exact values shows that using the first six terms givesa value that is correct to one part in 100, up to x=0.3.J
28.6.3 Prediction and correction
An improvement in the accuracy obtainable using difference methods is possible
if steps are taken, sometimes retrospectively, to allow for inaccuracies in approx-
imating derivatives by differences. We will describe only the simplest schemes ofthis kind and begin with a prediction method, usually called the Adams method .
1184
28.6 DIFFERENTIAL EQUATIONS
The forward difference estimate of yi+1,n a m e l y
yi+1=yi+hparenleftbiggdy
dxparenrightbigg
i=yi+hf(xi,yi), (28.68)
would give exact results if ywere a linear function of xin the range xi≤x≤xi+h.
The idea behind the Adams method is to allow some relaxation of this andsuppose that ycan be adequately approximated by a parabola over the interval
x
i−1≤x≤xi+1. In the same interval dy/dx can then be approximated by a linear
function:
f(x, y)=dy
dx≈a+b(x−xi)f o r xi−h≤x≤xi+h.
The values of aandbare fixed by the calculated values of fatxi−1andxi,w h i c h
we may denote by fi−1andfi:
a=fi,b =fi−fi−1
h.
Thus
yi+1−yi≈integraldisplayxi+h
xibracketleftbigg
fi+(fi−fi−1)
h(x−xi)bracketrightbigg
dx,
which yields
yi+1=yi+hfi+1
2h(fi−fi−1). (28.69)
The last term of this expression is seen to be a correction to result (28.68). That
it is, in some sense, the second-order correction
1
2h2y(2)
i−1/2
to a first-order formula is apparent.
Such a procedure requires, in addition to a value for y0, a value for either y1or
y−1,s ot h a t f1orf−1can be used to initiate the iteration. This has to be obtained
by other methods, e.g. a Taylor series expansion.
Improvements to simple difference formulae can also be obtained by using
correction methods. Here a rough prediction of the value yi+1is first made and
then this is used in a better formula, not originally usable since it in turn requires
a value of yi+1for its evaluation. The value of yi+1is then recalculated using this
better formula.
Such a scheme based on the forward difference formula might be as follows:
(i) predict yi+1using yi+1=yi+hfi;
(ii) calculate fi+1using this value;
(iii) recalculate yi+1using yi+1=yi+h(fi+fi+1)/2. Here ( fi+fi+1)/2h a s
replaced the fiused in (i), since it better represents the average value of
dy/dx in the interval xi≤x≤xi+h.
1185
NUMERICAL METHODS
Steps (ii) and (iii) can be iterated to improve further the approximation to the
average value of dy/dx , but this will not compensate for the omission of higher-
order derivatives in the forward difference formula.
Many more complex schemes of prediction and correction, in most cases
combining the two in the same process, have been devised, but the reader isreferred to more specialist texts for discussions of them. However, because itoffers some clear advantages, one group of methods will be set out explicitly inthe next subsection. This is the general class of schemes known as Runge–Kuttamethods.
28.6.4 Runge–Kutta methods
The Runge–Kutta method of integrating
dy
dx=f(x, y) (28.70)
is a step-by-step process of obtaining an approximation for yi+1by starting from
the value of yi. Among its advantages are that no functions other than fare used,
no subsidiary differentiation is needed and no additional starting values need be
calculated.
To be set against these advantages is the fact that fis evaluated using somewhat
complicated arguments and that this has to be done several times for each increase
in the value of i. However, once a procedure has been established, for example
on a computer, the method usually gives good results.
The basis of the method is to simulate the (accurate) Taylor series for y(xi+h),
not by calculating all the higher derivatives of yat the point xibut by taking
a particular combination of the values of the first derivative of yevaluated at
a number of carefully chosen points. Equation (28.70) is used to evaluate thesed e r i v a t i v e s .T h ea c c u r a c yc a nb em a d et ob eu pt ow h a t e v e rp o w e ro f his desired
but, naturally, the greater the accuracy the more complex the calculation and, inany case, rounding errors cannot ultimately be avoided.
The setting up of the calculational scheme may be illustrated by considering
the particular case in which second-order accuracy in his required. To second
order, the Taylor expansion is
y
i+1=yi+hfi+h2
2parenleftbiggdf
dxparenrightbigg
xi, (28.71)
whereparenleftbiggdf
dxparenrightbigg
xi=parenleftbigg∂f
∂x+f∂f
∂yparenrightbigg
xi≡∂fi
∂x+fi∂fi
∂y,
the last step being merely the definition of an abbreviated notation.
1186
28.6 DIFFERENTIAL EQUATIONS
We assume that this can be simulated by a form
yi+1=yi+α1hfi+α2hf(xi+β1h, y i+β2hfi), (28.72)
which in effect uses a weighted mean of the value of dy/dx atxiand its value at
some point yet to be determined. The object is to choose values of α1,α2,β1and
β2such that (28.72) coincides with (28.71) up to the coefficient of h2.
Expanding the function fin the last term of (28.72) in a Taylor series of its
own we obtain
f(xi+β1h, y i+β2hfi)=f(xi,yi)+β1h∂fi
∂x+β2hfi∂fi
∂y+O ( h2).
Putting this result into (28.72) and rearranging in powers of hwe obtain
yi+1=yi+(α1+α2)hfi+α2h2parenleftbigg
β1∂fi
∂x+β2fi∂fi
∂yparenrightbigg
. (28.73)
Comparing this with (28.71) shows that there is in fact some freedom remaining
in the choice of the α’s and β’s. In terms of an arbitrary α1(/negationslash=1 )
α2=1−α1,β 1=β2=1
2(1−α1).
One possible choice is α1=0.5, giving α2=0.5,β1=β2= 1. In this case the
procedure (equation (28.72)) can be summarised by
yi+1=yi+1
2(a1+a2), (28.74)
where
a1=hf(xi,yi),
a2=hf(xi+h, y i+a1).
Similar schemes giving higher-order accuracy in hcan be devised. Two such
schemes, given without derivation, are
(i) to order h3,
yi+1=yi+1
6(b1+4b2+b3), (28.75)
where
b1=hf(xi,yi),
b2=hf(xi+1
2h, y i+1
2b1),
b3=hf(xi+h, y i+2b2−b1),
1187
NUMERICAL METHODS
(ii) to order h4,
yi+1=yi+1
6(c1+2c2+2c3+c4), (28.76)
where
c1=hf(xi,yi),
c2=hf(xi+1
2h, y i+1
2c1),
c3=hf(xi+1
2h, y i+1
2c2),
c4=hf(xi+h, y i+c3).
28.6.5 Isoclines
The final method to be described for first-order differential equations is not so
much numerical as graphical, but since it is sometimes useful it is included here.The method, known as that of isoclines , involves sketching for a number of
values of a parameter cthose curves (the isoclines) in the xy-plane along which
f(x, y)=c, i.e. those curves along which dy/dx is a constant of known value. It
should be noted that they are not generally straight lines. Since a straight lineof slope dy/dx at and through any particular point is a tangent to the curve
y=y(x) at that point, small elements of straight lines, with slopes appropriate
to the isoclines they cut, effectively form the curve y=y(x).
Figure 28.6 illustrates in outline the method as applied to the solution of
dy
dx=−2xy. (28.77)
The thinner curves (rectangular hyperbolae) are a selection of the isoclines along
which−2xyis constant and equal to the corresponding value of c.T h es m a l l
cross lines on each curve show the slopes (= c) that solutions of (28.77) must
have if they cross the curve. The thick line is the solution for which y=1a t
x= 0; it takes the slope dictated by the value of con each isocline it crosses. The
analytic solution with these properties is y(x)=e x p (−x2).
28.7 Higher-order equations
So far the discussion of numerical solutions of differential equations has been
in terms of one dependent and one independent variable related by a first-orderequation. It is straightforward to carry out an extension to the case of severaldependent variables y
[r]governed by Rfirst-order equations
dy[r]
dx=f[r](x, y[1],y[2],...,y [R]),r =1,2,...,R.
We have enclosed the label rin brackets so that there is no confusion between,
say, the second dependent variable y[2]and the value y2of a variable yat the
1188
28.7 HIGHER-ORDER EQUATIONS
0.20.2
0.40.4
0.60.6
0.80.8
1.01.0
cy
y
x−1.0
−0.8
−0.6
−0.4
−0.2
−0.1
Figure 28.6 The isocline method. The cross lines on each isocline show the
slopes that solutions of dy/dx =−2xymust have at the points where they
cross the isoclines. The heavy line is the solution with y(0) = 1, namely
exp(−x2).
second calculational point x2. The integration of these equations by the methods
discussed in the previous section presents no particular difficulty, provided that
all the equations are advanced through each particular step before any of them
is taken through the following step.
Higher-order equations in one dependent and one independent variable can be
reduced to a set of simultaneous equations, provided that they can be written inthe form
d
Ry
dxR=f(x, y, y/prime,...,y(R−1)), (28.78)
where Ris the order of the equation. To do this, a new set of variables p[r]is
defined by
p[r]=dry
dxr,r =1,2,...,R−1. (28.79)
Equation (28.78) is then equivalent to the set of simultaneous first-order equations
dy
dx=p[1],
dp[r]
dx=p[r+1],r =1,2,...,R−2, (28.80)
dp[R−1]
dx=f(x, y, p [1],...,p [R−1]).
These can then be treated in the way indicated in the previous paragraph. The
extension to more than one dependent variable is straightforward.
1189
NUMERICAL METHODS
In practical problems it often happens that boundary conditions applicable
to a higher-order equation consist not of the values of the function and all itsderivatives at one particular point but of, say, the values of the function at twoseparate end-points. In these cases a solution cannot be found using an explicit
step-by-step ‘marching’ scheme, in which the solutions at successive values of the
independent variable are calculated using solution values previously found. Othermethods have to be tried.
One obvious method is to treat the problem as a ‘marching one’, but to use a
number of (intelligently guessed) initial values for the derivatives at the startingpoint. The aim is then to find, by interpolation or some other form of iteration,those starting values for the derivatives that will produce the given value of the
function at the finishing point.
In some cases the problem can be reduced by a differencing scheme to a matrix
equation. Such a case is that of a second-order equation for y(x) with constant
coefficients and given values of yat the two end-points. Consider the second-order
equation
y
/prime/prime+2ky/prime+µy=f(x), (28.81)
with the boundary conditions
y(0) = A, y (1) = B.
If (28.81) is replaced by a central difference equation,
yi+1−2yi+yi−1
h2+2kyi+1−yi−1
2h+µyi=f(xi),
we obtain from it the recurrence relation
(1 +kh)yi+1+(µh2−2)yi+( 1−kh)yi−1=h2f(xi).
Forh=1/(N−1) this is in exactly the form of the N×Ntridiagonal matrix
equation (28.30), with
b1=bN=1,c 1=aN=0,
ai=1−kh, b i=µh2−2,c i=1+ kh, i =2,3,...,N−1,
andy1replaced by A,yNbyBandyibyh2f(xi)f o r i=2,3,...,N−1. The
solutions can be obtained as in (28.31) and (28.32).
28.8 Partial differential equations
The extension of previous methods to partial differential equations, thus involving
two or more independent variables, proceeds in a more or less obvious way. Ratherthan an interval divided into equal steps by the points at which solutions to the
1190
28.8 PARTIAL DIFFERENTIAL EQUATIONS
equations are to be found, a mesh of points in two or more dimensions has to be
set up and all the variables given an increased number of subscripts.
Considerations of the stability, accuracy and feasibility of particular calcula-
tional schemes in principle are the same as for the one-dimensional case, but in
practice are too complicated to be discussed here.
Rather than note generalities that we are unable to pursue in any quantitative
way, we will conclude this chapter by indicating in outline how two familiarpartial differential equations of physical science can be set up for numericalsolution. The first of these is Laplace’s equation in two dimensions,
∂
2φ
∂x2+∂2φ
∂y2= 0 (28.82)
the value of φbeing given on the perimeter of a closed domain.
A grid with spacings ∆ xand ∆ yin the two directions is first chosen, so that,
for example, xistands for the point x0+i∆xandφi,jfor the value φ(xi,yj). Next,
using a second central difference formula, (28.82) is turned into
φi+1,j−2φi,j+φi−1,j
(∆x)2+φi,j+1−2φi,j+φi,j−1
(∆y)2=0,
(28.83)
fori=0,1,...,N andj=0,1,...,M .I f( ∆ x)2=λ(∆y)2then this becomes the
recurrence relationship
φi+1,j+φi−1,j+λ(φi,j+1+φi,j−1)=2 ( 1+ λ)φi,j. (28.84)
The boundary conditions in their simplest form (i.e. for a rectangular domain)
mean that
φ0,j,φ N,j,φ i,0,φ i,M (28.85)
have predetermined values. Non-rectangular boundaries can be accommodated,
either by more complex boundary-value prescriptions or by using non-Cartesiancoordinates.
To find a set of values satisfying (28.84), an initial guess at a complete set of
values for the φ
i,jis made, subject to the requirement that the quantities listed
in (28.85) have the given fixed values; those values that are not on the boundaryare then adjusted iteratively in order to try to bring about condition (28.84)everywhere. Clearly one scheme is to set λ= 1 and recalculate each φ
i,jas the
mean of the four current values at neighbouring grid-points, using (28.84) directly,and then to iterate this recalculation until no value of φchanges significantly
after a complete cycle through all values of iandj. This procedure is the simplest
of such ‘relaxation’ methods; for a slightly more sophisticated scheme see exercise
28.22 at the end of the chapter. The reader is referred to specialist books forfuller accounts of how this approach can be made faster and more accurate.
1191
NUMERICAL METHODS
Our final example is based upon the one-dimensional diffusion equation for
the temperature φof a system,
∂φ
∂t=κ∂2φ
∂x2. (28.86)
Ifφi,jstands for φ(x0+i∆x, t0+j∆t) then a forward difference representation
of the time derivative and a central difference representation for the spatialderivative lead to the following relationship:
φ
i,j+1−φi,j
∆t=κφi+1,j−2φi,j+φi−1,j
(∆x)2. (28.87)
This allows the construction of an explicit scheme for generating the temperature
distribution at later times, given that it is known at some earlier time:
φi,j+1=α(φi+1,j+φi−1,j)+( 1−2α)φi,j, (28.88)
where α=κ∆t/(∆x)2.
Although this scheme is explicit it is not a good one, because of the asymmetric
way in which the differences are formed. However, the effect of this can be
minimised if we study and correct for the errors introduced, in the following way.
Taylor’s series for the time variable gives
φi,j+1=φi,j+∆t∂φi,j
∂t+(∆t)2
2!∂2φi,j
∂t2+···, (28.89)
using the same notation as previously. Thus the first correction term to the
left-hand side of (28.87) is
−∆t
2∂2φi,j
∂t2. (28.90)
The first term omitted on the right-hand side of the same equation is, by a similar
argument,
−κ2(∆x)2
4!∂4φi,j
∂x4. (28.91)
But, using the fact that φsatisfies (28.86) we obtain
∂2φ
∂t2=∂
∂tparenleftbigg
κ∂2φ
∂x2parenrightbigg
=κ∂2
∂x2parenleftbigg∂φ
∂tparenrightbigg
=κ2∂4φ
∂x4, (28.92)
and so, to this accuracy, the two errors (28.90) and (28.91) can be made to cancel
ifαis chosen such that
−κ2∆t
2=−2κ(∆x)2
4!,i.e.α=1
6.
1192
28.9 EXERCISES
28.9 Exercises
28.1 Use an iteration procedure to find to four significant figures the root of the
equation 40 x=e x p x.
28.2 Using the Newton–Raphson procedure find, correct to three decimal places, the
root nearest to 7 of the equation 4 x3+2x2−200x−50 = 0.
28.3 (a) Show that if a polynomial equation g(x)≡xm−f(x)=0 ,where f(x)i sa
polynomial of degree less than mand for which f(0)/negationslash= 0, is solved using
a rearrangement iteration scheme xn+1=[f(xn)]1/m, then, in general, the
scheme will have only first-order convergence.
(b) By considering the cubic equation
x3−ax2+2abx−(b3+ab2)=0
for arbitrary non-zero values of aandb, demonstrate that, in special cases,
a rearrangement scheme can give second- (or higher-) order convergence.
28.4 The square root of a number Nis to be determined by means of the iteration
scheme
xn+1=xn
/
1−
/;
N−x2
n
/
f(N)
/
.
Determine how to choose f(N) so that the process has second-order convergence.
Given that√
7≈2.65, calculate√
7 as accurately as a single application of the
formula will allow.
28.5 Solve the following set of simultaneous equations using Gaussian elimination
(including interchange where it is formally desirable),
x1+3x2+4x3+2x4=0,
2x1+1 0x2−5x3+x4=6,
4x2+3x3+3x4=2 0,
−3x1+6x2+1 2x3−4x4=1 6.
28.6 The following table of values of a polynomial p(x) of low degree contains an
error. Identify and correct the erroneous value and extend the table up to x=1.2.
xp (x) xp (x)
0.0 0.000 0.5 0.165
0.1 0.011 0.6 0.2160.2 0.040 0.7 0.2450.3 0.081 0.8 0.2560.4 0.128 0.9 0.243
28.7 Simultaneous linear equations that result in tridiagonal matrices can be treated
sometimes as three-term recurrence relations and their solution found in a similar
manner to that described in chapter 15. Consider the tridiagonal simultaneousequations
x
i−1+4xi+xi+1=3 (δi+1,0−δi−1,0),i=0,±1,±2,... .
Prove that for i>0 the equations have a general solution of the form xi=
αpi+βqi,w h e r e pandqare the roots of a certain quadratic equation. Show that
a similar result holds for i<0. In each case express x0in terms of the arbitrary
constants α, β, . . . .
Now impose the condition that xiis bounded as i→±∞ and obtain a unique
solution.
1193
NUMERICAL METHODS
28.8 A possible rule for obtaining an approximation to an integral is the mid-point
rule,g i v e nb yZx0+∆x
x0f(x)dx=∆xf(x0+1
2∆x)+O ( ∆ x3).
Writing hfor ∆ x, and evaluating all derivates at the mid-point of the interval
(x, x+∆x), use a Taylor series expansion to find, up to O( h5), the coefficients of
the higher-order errors in both the trapezium and midpoint rules. Hence find alinear combination of these two rules that gives O( h
5)a c c u r a c yf o re a c hs t e p∆ x.
28.9 Given a random number ηuniformly distributed on (0 ,1), determine the function
ξ=ξ(η) that would generate a random number ξdistributed as
(a) 2 ξon 0≤ξ<1.
(b)3
2√ξon 0≤ξ<1.
(c)π
4acosπξ
2aon−a≤ξ<a.
(d)1
2exp(−|ξ|)o n−∞ <ξ<∞.
28.10 A, Band Care three circles of unit radius with centres in the xy-plane at
(1,2),(2.5,1.5) and (2 ,3) respectively. Devise a hit or miss Monte Carlo calculation
to determine the size of the area that lies outside Cbut inside AandB, as well
as inside the square centred on (2 ,2.5) that has sides of length 2 parallel to the
coordinate axes. You should choose your sampling region so as to make the
estimation as efficient as possible. Take the random number distribution to be
uniform on (0 ,1) and determine the inequalities that have to be tested using the
random numbers chosen.
28.11 Use a Taylor series to solve the equation
dy
dx+xy=0,y (0) = 1 ,
evaluating y(x)f o r x=0.0 to 0.5 in steps of 0.1.
28.12 Consider the application of the predictor–corrector method described near the
end of subsection 28.6.3 to the equation
dy
dx=x+y.
Show, by comparison with a Taylor series expansion, that the expression obtained
foryi+1in terms of xiandyiby applying the three steps indicated (without any
repeat of the last two) is correct to O( h2). Using steps of h=0.1 compute the
value of y(0.3) and compare it with the value obtained by solving the equation
analytically.
28.13 A more refined form of the Adams predictor–corrector method for solving the
first-order differential equation
dy
dx=f(x, y)
is known as the Adams–Moulton–Bashforth scheme. At any stage (say the nth)
in an Nth-order scheme the values of xandyat the previous Nsolution points
are first used to predict the value of yn+1. This approximate value of yat the
next solution point xn+1, denoted by ¯yn+1, is then used together with those at the
previous N−1 solution points to make a more refined ( corrected ) estimation of
y(xn+1). The calculational procedure for a third-order scheme is summarised by
the two equations
¯yn+1=yn+h(a1fn+a2fn−1+a3fn−2)( p r e d i c t o r ) ,
yn+1=yn+h(b1f(xn+1,¯yn+1)+b2fn+b3fn−1) (corrector) .
1194
28.9 EXERCISES
(a) Find Taylor series expansions for fn−1andfn−2in terms of the function
fn=f(xn,yn) and its derivatives at xn.
(b) Substitute them into the predictor equation and, by making that expression
for¯yn+1coincide with the true Taylor series for yn+1up to order h3, establish
simultaneous equations that determine the values of a1,a2anda3.
(c) Find the Taylor series for fn+1and substitute it and that for fn−1into the
corrector equation. Make the corrected prediction for yn+1coincide with the
true Taylor series by choosing the weights b1,b2andb3appropriately.
(d) The values of the numerical solution of the differential equation
dy
dx=2(1 + x)y+x3/2
2x(1 +x)
at three values of xare given in the following table:
x0.1 0.2 0.3
y(x) 0.030628 0.084107 0.150328.
Use the above predictor–corrector scheme to find the value of y(0.4) and
compare your answer with the accurate value, 0.225577.
28.14 If dy/dx =f(x, y) then show that
d2f
dx2=∂2f
∂x2+2f∂2f
∂x∂y+f2∂2f
∂y2+∂f
∂x∂f
∂y+f
/df
dy
/2
.
Hence verify, by substitution and the subsequent expansion of arguments in
Taylor series of their own, that the scheme given in (28.75) coincides with theTaylor expansion (28.64), i.e.
y
i+1=yi+hy(1)
i+h2
2!y(2)
i+h3
3!y(3)
i+···.
up to terms in h3.
28.15 To solve the ordinary differential equation
du
dt=f(u, t)
forf=f(t), the explicit two-step finite difference scheme
un+1=αun+βun−1+h(µfn+νfn−1)
may be used. Here, in the usual notation, his the time step, tn=nh,un=u(tn)
andfn=f(un,tn);α,β,µ,a n d νare constants.
(a) A particular scheme has α=1 ,β=0,µ=3/2a n d ν=−1/2. By considering
Taylor expansions about t=tnfor both un+jandfn+j, show that this scheme
gives errors of order h3.
(b) Find the values of α,β,µ,a n d νthat will give the greatest accuracy.
28.16 Set up a finite difference scheme to solve the ordinary differential equation
xd2φ
dx2+dφ
dx=0
in the range 1 ≤x≤4 and subject to the boundary conditions φ(1) = 2 and
dφ/dx =2a t x=4 .U s i n g Nequal increments, ∆ x,i nx, obtain the general
difference equation and state how the boundary conditions are incorporatedinto the scheme. Setting ∆ xequal to the (crude) value 1, obtain the relevant
simultaneous equations and so obtain rough estimates for φ(2),φ(3) and φ(4).
Finally, solve the original equation analytically and compare your numerical
estimates with the accurate values.
1195
NUMERICAL METHODS
28.17 Write a computer program that would solve, for a range of values of λ,t h e
differential equation
dy
dx=1p
x2+λy2,y (0) = 1 ,
using a third-order Runge–Kutta scheme. Consider the difficulties that might
arise when λ<0.
28.18 Use the isocline approach to sketch the family of curves that satisfies the non-
linear first-order differential equation
dy
dx=ap
x2+y2.
28.19 For some problems, numerical or algebraic experimentation may suggest the
form of the complete solution. Consider the problem of numerically integratingthe first-order wave equation
∂u
∂t+A∂u
∂x=0,
in which Ais a positive constant. A finite difference scheme for this partial
differential equation is
u(p, n+1 )−u(p, n)
∆t+Au(p, n)−u(p−1,n)
∆x=0,
where x=p∆xandt=n∆t,w i t h pany integer and na non-negative integer.
The initial values are u(0,0) = 1 and u(p,0) = 0 for p/negationslash=0 .
(a) Carry the difference equation forward in time for two or three steps and
attempt to identify the pattern of solution. Establish the criterion for themethod to be numerically stable.
(b) Suggest a general form for u(p, n), expressing it in generator function form,
i.e. ‘as u(p, n) is the coefficient of s
pin the expansion of G(n, s)’.
(c) Using your form of solution (or that given in the answers!), obtain an
explicit general expression for u(p, n) and verify it by direct substitution into
the difference equation.
(d) An analytic solution of the original PDE indicates that an initial distur-
bance propagates undistorted. Under what circumstances would the differ-
ence scheme reproduce that behaviour?
28.20 In the previous question the difference scheme for solving
∂u
∂t+∂u
∂x=0,
in which Ahas been set equal to unity, was one-sided in both space ( x)a n d
time ( t). A more accurate procedure (known as the Lax–Wendroff scheme) is
u(p, n+1 )−u(p, n)
∆t+u(p+1,n)−u(p−1,n)
2∆x
=∆t
2
/u(p+1,n)−2u(p, n)+u(p−1,n)
(∆x)2
/
.
(a) Establish the orders of accuracy of the two finite difference approximations
on the LHS of the equation.
(b) Establish the accuracy with which the expression in the brackets approxi-
mates ∂2u/∂x2.
(c) Show that the RHS of the equation is such as to make the whole difference
scheme accurate to second order in both space and time.
1196
28.9 EXERCISES
28.21 Laplace’s equation
∂2V
∂x2+∂2V
∂y2=0,
is to be solved for the region and boundary conditions shown in figure 28.7.
40 40 40 40 40 40 40
20 20 20V=8 0
V=0∞ −∞
Figure 28.7 Region, boundary values and initial guessed solution for exer-
cise 28.21.
Starting from the given initial guess for the potential values Vand using the
simplest possible form of relaxation, obtain a better approximation to the actualsolution. Do not aim to be more accurate than ±0.5 units and so terminate the
process when subsequent changes would be no greater than this.
28.22 Consider the solution φ(x, y) of Laplace’s equation in two dimensions using a
relaxation method on a square grid with common spacing h.A si nt h em a i nt e x t ,
denote φ(x
0+ih, y 0+jh)b yφi,j. Further, define φm,n
i,jby
φm,n
i,j≡∂m+nφ
∂xm∂yn
evaluated at ( x0+ih, y 0+jh).
(a) Show that
φ4,0
i,j+2φ2,2
i,j+φ0,4
i,j=0.
(b) Working up to terms of order h5, find Taylor series expansions, expressed in
terms of the φm,n
i,j,f o r
S±,0=φi+1,j+φi−1,j
S0,±=φi,j+1+φi,j−1.
(c) Find a corresponding expansion, to the same order of accuracy, for φi±1,j+1+
φi±1,j−1and hence show that
S±,±=φi+1,j+1+φi+1,j−1+φi−1,j+1+φi−1,j−1
has the form
4φ0,0
i,j+2h2(φ2,0
i,j+φ0,2
i,j)+h4
6(φ4,0
i,j+6φ2,2
i,j+φ0,4
i,j).
1197
NUMERICAL METHODS
(d) Evaluate the expression 4( S±,0+S0,±)+S±,±and hence deduce that a possible
relaxation scheme, good to the fifth order in h, is to recalculate each φi,jas
the weighted mean of the current values of its four nearest neighbours (eachwith weight
1
5) and its four next-nearest neighbours (each with weight1
20).
28.23 The Schr ¨odinger equation for a quantum mechanical particle of mass mmoving
in a one-dimensional harmonic oscillator potential V(x)=kx2/2i s
−
/~2
2md2ψ
dx2+kx2ψ
2=Eψ.
For physically acceptable solutions the wavefunction ψ(x) must be finite at x=0 ,
tend to zero as x→± ∞ and be normalised so that
R
|ψ|2dx=1 .I np r a c t i c e
these constraints mean that only certain (quantised) values of E, the energy of
the particle, are allowed. The allowed values fall into two groups, those for whichthe corresponding y(0) = 0 and those for which the corresponding y(0)/negationslash=0 .
Show that if the unit of length is taken as [/~2/(mk)]1/4and the unit of energy
as /~(k/m)1/2then the Schr ¨odinger equation takes the form
d2ψ
dy2+( 2E/prime−y2)ψ=0.
Devise an outline computerised scheme, using Runge–Kutta integration, that will
enable you to:
(a) determine the three lowest allowed values of E;
(b) tabulate the normalised wavefunction corresponding to the lowest allowed
energy.
You should consider explicitly:
(i) the variables to use in the numerical integration;
(ii) how starting values near y= 0 are to be chosen;
(iii) how the condition on ψasy→±∞ is to be implemented;
(iv) how the required values of Eare to be extracted from the results of the
integration;
(v) how the normalisation is to be carried out.
28.10 Hints and answers
28.1 5.370.
28.2 6.951 after two iterations.28.3 (a) ξ/negationslash=0a n d f
/prime(ξ)/negationslash= 0 in general; (b) ξ=b, but f/prime(b) = 0 whilst f(b)/negationslash=0 .
28.4 f(N)=(N−3x2)−1=−(2N)−1; 2.6457411, accurate value 2.6457513.
28.5 Interchange is formally needed for the first two steps, though in this case no error
will result if it is not carried out; x1=−12,x2=2,x3=−1,x4=5.
28.6 p(0.5) = 0 .175,p(1.0) = 0 .200,p(1.1) = 0 .121,p(1.2) = 0 .000.
28.7 The quadratic equation is z2+4z+1=0 ; α+β−3=x0=α/prime+β/prime+3 .
With p=−2+√3a n d q=−2−√3,βmust be zero for i>0a n d α/primemust be
zero for i<0;xi=3 (−2+√3)ifori>0,xi=0f o r i=0,xi=−3(−2−√3)i
fori<0.
28.8 Iexact=hf+h3f/24;IT=hf+h3f/8;IM=hf. Thus Iexact=1
3IT+2
3IM+O (h5).
28.9 Listed below are the relevant indefinite integrals F(y) of the distributions together
with the functions ξ=ξ(η):
(a)y2;ξ=√η.
(b)y3/2;ξ=η2/3.
1198
28.10 HINTS AND ANSWERS
aa
2a2a
−a
−a−2a
−2axy
Figure 28.8 Typical solutions y=y(x), shown by solid lines, of dy/dx =
a(x2+y2)−1/2. The short arrows give the direction that the tangent to any
solution must have at that point.
(c)1
2[sin(πy/2a)+1 ] ; ξ=( 2a/π)sin−1(2η−1).
(d)1
2exp(y)f o r y≤0;1
2(1−exp(−y)) for y>0.ξ=l n ( 2 η)f o r0 <η≤1
2;
ξ=−ln[2(1−η)] for frac12<η< 1.
28.10 Show that the corners of the required area (with three curved sides and one
straight one) are at (1 .5,1.5),(1.866,1.5),(2,2) and (1 .669,2.056). For a pair of
random numbers ( ξ1,ξ2), take x=1.5+αξ1andy=1.5+βξ2. The best values
forαandβare 0.5 and 0.556 respectively. Test the inequalities
(αξ1+0.5)2+(βξ2−0.5)2≤1,
(αξ1−1)2+(βξ2)2≤1,
(αξ1−0.5)2+(βξ2−1.5)2≥1.
If all three conditions are satisfied for nout of Npairs, the area can be estimated
asnαβ/N .
28.11 1−x2/2+x4/8−x6/48; 1.0000, 0.9950, 0.9802, 0.9560, 0.9231, 0.8825; exact
solution y=e x p (−x2/2).
28.12 yi+1=yi+h(xi+yi)+1
2h2(1 + xi+yi). The numerical estimate is 0.04923; the
analytic solution y(x)=ex−1−xgives 0.04986.
28.13 (b) a1=2 3/12,a2=−4/3,a3=5/12.
(c)b1=5/12,b2=2/3,b3=−1/12.
(d)¯y(0.4) = 0 .224582 ,y(0.4) = 0 .225527 after correction.
28.15 (a) The error is5
12h3un+O ( h4).
(b)α=−4,β=5,µ=4 ,a n d ν=2
1199
NUMERICAL METHODS
40 40
2041.5 41.5 46.5 46.5 48
16.5 16.5V=8 0
V=0∞ −∞
Figure 28.9 The solution to exercise 28.21.
28.16 With xj=1+ j∆x,N∆x=3a n d φj=φ(xj),
[2 + (2 j+1 ) ∆ x]φj+1−(4 + 4 j∆x)φj+[ 2+( 2 j−1)∆x]φj−1=0
forj=1,2,...,N−2. In addition φ0=2a n d φN=φN−1+2 ∆ x. Analytically
φ(x)=2+8l n x. Estimated (accurate) values are 6.67 (7.55), 9.47 (10.79), 11.47
(13.09).
28.18 See figure 28.8.28.19 (a) Setting A∆t=c∆xgives, for example, u(0,2) = (1−c)
2,u(1,2) = 2 c(1−c),
u(2,2) = c2. For stability 0 <c< 1.
(b)G(n, s)=[ ( 1−c)+cs]nfor 0≤p≤n.
(c) [ n!(1−c)n−pcp]/[p!(n−p)!].
(d) When c= 1 and the difference equation becomes u(p, n+1 )= u(p−1,n).
28.20 (a) First order in time and second order in space; (b) O(∆ x)2;( c )s h o wt h a t
∂2u/∂t2=∂2u/∂x2and write the second-order correction to ∂u/∂t in the form
1
2(∆t)2∂2u/∂x2.
28.21 See figure 28.9.
28.22 (a) Write Laplace’s equation as φ2,0
i,j+φ0,2
i,j= 0 and differentiate twice with respect
toxandyseparately. Then add the resulting equations.
(b)S±,0=2φ0,0
i,j+h2φ2,0
i,j+1
12h4φ4,0
i,jand corresponding results for S0,±.
(c) Note that φ0,0
i±1,j,φ0,2
i±1,jandφ0,4
i±1,jhave themselves to be expanded as Taylor
series in x, the number of terms to be retained being determined by the power
of ∆y=hthat multiplies them. The roles of xandycould be reversed.
(d) Use Laplace’s equation and result (a) to show that the given expression has
the value 20 φi,j.
1200
Appendix
Gamma, beta and error functions
In several places in this book we have made mention of the gamma, beta and error
functions. These convenient functions appear in a number of contexts and herewe gather together some of their properties. This appendix should be regardedmerely as a reference containing some useful relations with a minimum of formalproofs.
A1.1 The gamma function
Thegamma function Γ(n) is defined by
Γ(n)=integraldisplay
∞
0xn−1e−xdx, (A1)
which converges for n>0. Replacing nbyn+1 in (A1) and integrating the RHS
by parts, we find
Γ(n+1 )=integraldisplay∞
0xne−xdx
=bracketleftbig
−xne−xbracketrightbig∞
0+integraldisplay∞
0nxn−1e−xdx
=nintegraldisplay∞
0xn−1e−xdx,
from which we obtain the important result
Γ(n+1 )= nΓ(n). (A2)
From (A1), we see that Γ(1) = 1, and so, if nis a positive integer,
Γ(n+1 )= n!. (A3)
1201
GAMMA, BETA AND ERROR FUNCTIONS
Γ( )
1 22
344
−1−2
−2−3−4
−4
−66
nn
Figure A1.1 The gamma function Γ( n).
In fact, equation (A3) serves as a definition of the factorial function even for
non-integer n. For negative nthe factorial function is defined by
n!=(n+m)!
(n+m)(n+m−1)···(n+1 ), (A4)
where mis any positive integer that makes n+m>0. Different choices of m
(>−n) do not lead to different values for n!. A plot of the gamma function is
given in figure A1.1, where it can be seen that the function is infinite for negativeinteger values of n, in accordance with (A4).
By letting x=y
2in (A1), we immediately obtain another useful representation
of the gamma function given by
Γ(n)=2integraldisplay∞
0y2n−1e−y2dy. (A5)
Setting n=1
2we find the result
Γparenleftbig1
2parenrightbig
=2integraldisplay∞
0e−y2dy=integraldisplay∞
−∞e−y2dy=√π,
where have used the standard integral discussed in section 6.4.2. From this result,
Γ(n) for half-integral ncan be found using (A2). Some immediately derivable
factorial values of half integers are
parenleftbig
−3
2parenrightbig
!=−2√π,parenleftbig
−1
2parenrightbig
!=√π,parenleftbig1
2parenrightbig
!=1
2√π,parenleftbig3
2parenrightbig
!=3
4√π.
1202
A1.2 THE BETA FUNCTION
It can also be shown that the gamma function is given by
Γ(n+1 )=√
2πn nne−nparenleftbigg
1+1
12n+1
288n2−139
51 840 n3+...parenrightbigg
=n!,(A6)
which is known as Stirling’s asymptotic series . For large nthe first term dominates
and so
n!≈√
2πn nne−n; (A7)
this is known as Stirling’s approximation . This approximation is particularly useful
in statistical thermodynamics, when arrangements of a large number of particlesa r et ob ec o n s i d e r e d .IProve Stirling’s approximation n!≈√
2πn nne−nfor large n.
From (A1), the extended definition of the factorial function (which is valid for n>−1) is
given by
n!=
Z∞
0xne−xdx=
Z∞
0enlnx−xdx. (A8)
If we let x=n+y,t h e n
lnx=l nn+l n
/
1+y
n
/
=l nn+y
n−y2
2n2+y3
3n3−···.
Substituting this result into (A8), we obtain
n!=
Z∞
−nexp
/
n
/
lnn+y
n−y2
2n2+···
/
−n−y
/
dy.
Thus, when nis sufficiently large, we may approximate n!b y
n!≈enlnn−n
Z∞
−∞e−y2/(2n)dy=enlnn−n√
2πn=√
2πn nne−n,
which is Stirling’s approximation (A7).
J
A1.2 The beta function
Thebeta function is defined by
B(m, n)=integraldisplay1
0xm−1(1−x)n−1dx, (A9)
which converges for m>0,n>0. By letting x=1−yin (A9) it is easy to show
thatB(m, n)=B(n, m). Other useful representations of the beta function may be
obtained by suitable changes of variable. For example, putting x=( 1+ y)−1in
(A9), we find that
B(m, n)=integraldisplay∞
0yn−1dy
(1 +y)m+n.
1203
GAMMA, BETA AND ERROR FUNCTIONS
Alternatively, if we let x=s i n2θin (A9), we obtain immediately
B(m, n)=2integraldisplayπ/2
0sin2m−1θcos2n−1θd θ . (A10)
The beta function may also be written in terms of the gamma function as
B(m, n)=Γ(m)Γ(n)
Γ(m+n). (A11)IProve the result (A11).
Using (A5), we have
Γ(n)Γ(m)=4
Z∞
0x2n−1e−x2dx
Z∞
0y2m−1e−y2dy
=4
Z∞
0
Z∞
0x2n−1y2m−1e−(x2+y2)dx dy.
Changing variables to plane polar coordinates ( ρ, φ)g i v e nb y x=ρcosφ,y=ρsinφ,w e
obtain
Γ(n)Γ(m)=4
Zπ/2
0
Z∞
0ρ2(m+n−1)e−ρ2sin2m−1θcos2n−1θ ρ dρ dθ
=4
Zπ/2
0sin2m−1θcos2n−1θd θ
Z∞
0ρ2(m+n)−1e−ρ2dρ
=B(m, n)Γ(m+n),
where in the last line we have used the results (A5) and (A10).
J
A1.3 The error function
Finally we mention the error function , which is encountered in probability theory
and in the solutions of some partial differential equations, and which is definedby
erf(x)=2
√πintegraldisplayx
0e−u2du=1−2√πintegraldisplay∞
xe−u2du. (A12)
From this definition we can easily see that
erf(0) = 0 ,erf(∞)=1 ,erf(−x)=−erf(x).
By making the substitution y=√
2uin (A12), we find
erf(x)=1√
2πintegraldisplay√
2x
0e−y2/2dy.
The cumulative probability function Φ( x) for the standard Gaussian distribution
(discussed in section 26.9.1) may be written in terms of the error function as
1204
A1.3 THE ERROR FUNCTION
follows
Φ(x)=1√
2πintegraldisplayx
−∞e−y2/2dy
=1
2+1√
2πintegraldisplayx
0e−y2/2dy
=1
2+e r fparenleftbiggx√
2parenrightbigg
.
It is also sometimes useful to define the complementary error function
erfc(x)=1−erf(x)=2√πintegraldisplay∞
xe−u2du. (A13)
1205
Index
F-distribution (Fisher), 1132–1138
critical points table, 1137logarithmic form, 1138
t-test,seeStudent’s t-test
correlation, chi-squared test , 1143
Cram ´er-Rao (Fisher’s) inequality , 1075, 1076
Fisher’s inequality , 1075, 1076maximum-likelihood, method of
extended, 1112
A, B, one-dimensional irreps, 932, 944, 951
Abelian groups, 886absolute convergence of series, 127, 717
absolute derivative, 824–826
acceleration vector, 341Adams method, 1184Adams–Moulton–Bashforth, predictor-corrector
scheme, 1194
addition rule for probabilities, 967, 972
adjoint, seeHermitian conjugate
adjoint operators, 587–591adjustment of parameters, 854–855algebra of
complex numbers, 88–89
functions in a vector space, 583
matrices, 256–257power series, 137series, 134tensors, 787–790vectors, 217–218
in a vector space, 247
in component form, 222
algebraic equations, numerical methods for, see
numerical methods for algebraic equations
alternating group, 958
alternating series test, 133
ammonia molecule, symmetries of, 884Amp`ere’s rule (law), 387, 414
amplitude modulation of radio waves, 450analytic (regular) functions, 712angle between two vectors, 225
angular frequency, 626n
in Fourier series, 425
angular momentum, 782, 798
and irreps, 935of particle system, 799–801of particles, 344of solid body, 402, 800
vector representation, 241
angular velocity, vector representation, 227, 241,
359
anti-Hermitian matrices, 276
eigenvalues, 281–283
imaginary nature, 282–283
eigenvectors, 281–283
orthogonality, 282–283
anticommutativity of vector/cross product, 226
antisymmetric functions, 422
and Fourier series, 425–426and Fourier transforms, 451
antisymmetric matrices, 275
general properties, seeanti-Hermitian
matrices
antisymmetric tensors, 787, 790antithetic variates, in Monte Carlo methods,
1174
aperture function, 443approximately equal ≈, definition, 135
arbitrary parameters for ODE, 475
arc length of
plane curves, 74–75space curves, 347
arccosech, arccosh, arccoth, arcsech, arcsinh,
arctanh, seehyperbolic functions, inverses
Archimedean upthrust, 402, 416area element in
Cartesian coordinates, 191plane polars, 205
area of
circle, 72
ellipse, 72, 210
1206
INDEX
parallelogram, 227
region, using multiple integrals, 194–196surfaces, 352
as vector, 399–401, 414
area, maximal enclosure, 838arg, argument of a complex number, 90Argand diagram, 87, 711
argument, principle of the, 755
arithmetic series, 120arithmetico-geometric series, 121arrays, seematrices
associated Legendre equation, 594–595, 666, 670,
703
associated Legendre functions P
m
/lscript(x), 666
generating function, 594
orthogonality, 594Rodrigues’ formula, 594
associated Legendre functions P
m
/lscript(x), 703
associative law for
addition
in a vector space of finite dimensionality,
247
in a vector space of infinite dimensionality,583of complex numbers, 89of matrices, 256of vectors, 217
convolution, 453, 464
group operations, 885
linear operators, 254multiplication
of a matrix by a scalar, 256o fav e c t o rb yas c a l a r ,2 1 8of complex numbers, 91of matrices, 258
multiplication by a scalar
in a vector space of finite dimensionality,
247
in a vector space of infinite dimensionality,583
atomic orbitals, 957
d-states, 948, 950, 956
p-states, 948
s-states, 986
auto-correlation functions, 456automorphism, 903auxiliary equation, 499
repeated roots, 499
average value, seemean value
axial vectors, 798
backward differences, 1179
basis functions
for linear least squares estimation, 1115in a vector space of infinite dimensionality,
583–584
of a representation, 920
change in, 926–927, 929, 934
basis vectors, 221–222, 248–249, 778, 920
derivatives, 814–816Christoffel symbol Γ
k
ij, 814
for particular irrep, 948–950, 958
linear dependence and independence, 221
non-orthogonal, 250
orthonormal, 249–250
required properties, 221
Bayes’ theorem, 974–975Bernoulli equation, 483
Bessel correction to variance estimate, 1090
Bessel equation, 541, 564–568, 595, 674Bessel functions J
ν(z)
zeroes of, 662, 673
Bessel functions Jν(z), 662, 672
generating function, 573, 595
integral relationships, 572
integral representation, 574–575orthogonality, 570–573
recurrence relations, 569–570
second kind, Y
ν(z), 568
series, 565, 566, 595
ν= 0, 567
ν=±1/2, 566
spherical, 675
Bessel inequality, 251, 586best unbiased estimator, 1075
beta function, 1203
bias, of estimator, 1074bilinear transformation, general, 113
binary chopping, 1154
binomial coefficient
nCk, 27–30
elementary properties, 26
identities, 27
negative n,2 9
non-integral n,2 9
binomial coefficientnCk, 977–979
in Leibniz’ theorem, 50
binomial distribution Bin( n, p), 1010–1013
and Gaussian distribution, 1027and Poisson distribution, 1016, 1019
mean and variance, 1013
MGF, 1012recurrence formula, 1011
binomial expansion, 143
binormal to space curves, 348birthdays, different, 976
bivariate distributions, 1038–1049
conditional, 1040
continuous, 1039
correlation, 1042–1049
and independence, 1042
matrix, 1045–1049
positive/negative, 1042
covariance, 1042–1049
matrix, 1045
expectation (mean), 1041independent (uncorrelated), 1039
marginal, 1040
variance, 1042
Boltzmann distribution
from constraints on total energy, 174–176
1207
INDEX
bonding in molecules, 945, 947–950
Born approximation, 152, 606Bose–Einstein statistics, 980boundary conditions
and characteristics, 633and Laplace equation, 699, 701
for Green’s functions, 518, 520–522
inhomogeneous, 521
for ODE, 474, 476, 507for PDE, 614, 618–620for Sturm–Liouville equations, 592homogeneous and inhomogeneous, 618, 656,
686, 688
superposition solutions, 651–657types, 635–638
brachistochrone problem, 843Bragg formula, 241branch cut, 722branch points, 721Bromwich integral, 765bulk modulus, 829
calculus of residues, seezeroes of a function of
a complex variable andcontour integration
calculus of variations
constrained variation, 844–846estimation of ODE eigenvalues, 849
Euler–Lagrange equation, 835–836
Fermat’s principle, 846Hamilton’s principle, 847higher-order derivatives, 841several dependent variables, 841several independent variables, 841soap films, 839–840variable end-points, 841–844
calculus, elementary, 42–77cancellation law in a group, 888
canonical form, for second-order ODE, 522
card drawing, seeprobability
carrier frequency of radio waves, 451Cartesian coordinates, 221–222Cartesian tensors, 779–804
algebra, 787–790contraction, 788definition, 784first-order, 781–784
from scalar, 783
general order, 784–803
integral theorems, 803–804isotropic, 793–795physical applications, 783, 788–790, 799–803second-order, 784–803, 817symmetry and antisymmetry, 787tensor fields, 803zero-order, 781–784
from vector, 784
Cartesian tensors, particular
conductivity, 801inertia, 800strain, 802stress, 802
susceptibility, 801
catenary, 840, 846Cauchy
boundary conditions, 635distribution, 994inequality, 747integrals, 745–747product, 134root test, 132, 717theorem, 742
Cauchy–Riemann relations, 713–716, 727, 729,
743
in terms of zandz
∗, 715
central differences, 1179central limit theorem, 1036–1038, 1178central moments, seemoments, central
centre of a group, 911centre of mass, 198
of hemisphere, 198of semicircular lamina, 200
centroid, 198
of plane area, 198of plane curve, 200of triangle, 220–221
CF,seecomplementary function
chain rule for functions of
one real variable, 47–48several real variables, 160–161
change of basis, seesimilarity transformations
change of variables
and coordinate systems, 161–163in multiple integrals, 202–210
evaluation of Gaussian integral, 205–207
general properties, 209–210
in RVD, 992–999
character tables, 935
4mmorC
4v, 944, 948, 950
S4or 432 or O, 956, 957
¯43morTd, 957
3mor 32 or C3vorS3, 935, 939, 950, 952, 958,
959
construction of, 942–944
characteristic equation, 285
normal mode form, 325of recurrence relation, 505
characteristic function, seemoment generating
functions (MGF)
characteristics
and boundary curves, 633
multiple intersections, 633, 638
and the existence of solutions, 632–638first-order equations, 632–633second-order equations, 636
and equation type, 636
characters, 934–938, 942–944
and conjugacy classes, 934, 937character tables
4mmorC
4vorD4, 955
A4, 958
1208
INDEX
D5, 958
quaternion, 955
counting irreps, 937–938
definition, 934of product representation, 946–947
orthogonality properties, 936, 944
summation rules, 939
charge (point), Dirac δ-function respresentation,
447
charged particle in electromagnetic fields, 376
Chebyshev equation, 541, 597
polynomial solutions, 578
Chebyshev polynomials T
n(x)
generating function, 597
orthogonality, 597Rodrigues’ formula, 597
chi-squared ( χ
2) distribution
and likelihood-ratio test, 1125
chi-squared ( χ2) distribution, 1034
and goodness of fit, 1139
and likelihood-ratio test, 1134and multiple estimators, 1086
test for correlation, 1143
Cholesky separation, 318Christoffel symbol Γ
k
ij, 814–816
from metric tensor, 815, 822
circle of convergence, 717–718circle, area of, 72
circuits, electrical
transients, 491
Clairaut equation, 489
classes and equivalence relations, 906
closure of a group, 885closure property of eigenfunctions of an
Hermitian operator, 601
cofactor of a matrix element, 264
column matrix, 255
column vector, 255combinations (probability), 975–981
common ratio in geometric series, 120
commutation law for group elements, 886commutative law for
addition
in a vector space of finite dimensionality,247
in a vector space of infinite dimensionality,
583
of complex numbers, 89
of matrices, 256of vectors, 217
complex scalar/dot product, 226
convolution, 453, 464inner product, 249
multiplication
o fav e c t o rb yas c a l a r ,2 1 8of complex numbers, 91
scalar/dot product, 224
commutator, of two matrices, 314comparison test, 128
complement, 963probability for, 967
complementary equation, 496complementary error function, 1205complementary function (CF), 497
for ODE, 498partially known, 512repeated roots of auxiliary equation, 499
completeness of
basis vectors, 248eigenfunctions of an Hermitian operator, 588,
601
eigenvectors of a normal matrix, 280spherical harmonics Y
m
/lscript(θ, φ), 670
completing the square
as a means of integration, 67–68for quadratic equations, 35to evaluate Gaussian integral, 442, 684
complex conjugate
z
∗, of complex number, 92–94, 715
of a matrix, 261–263of scalar/dot product, 226properties of, 93
complex exponential function, 95, 719complex Fourier series, 430–431complex integrals, 738–742, see also zeroes of a
function of a complex variable andcontour
integration
Cauchy’s, 745–747Cauchy’s theorem, 742definition, 739
Jordan’s lemma, 761–762
Morera’s theorem, 744ofz
−1, 740
principal value, 760residue theorem, 752–754
complex logarithms, 102–103, 720
principal value of, 103, 720
complex numbers, 86–117
addition and subtraction of, 88–89applications to differentiation and integration,
104
argument of, 90associativity of
addition and subtraction, 89
multiplication, 91
commutativity of
addition and subtraction, 89multiplication, 91
complex conjugate of, seecomplex conjugate
components of, 87de Moivre’s theorem, seede Moivre’s theorem
division of, 94–95, 97–98
properties, 95
from roots of polynomial equations, 86–87imaginary part of, 86–87modulus of, 90multiplication of, 91–92, 97–98
as rotation in the Argand diagram, 91–92
notation, 87polar representation of, 95–98
1209
INDEX
real part of, 86–87
trigonometric representation of, 96
complex potentials, 725–730
and fluid flow, 727–728equipotentials and field lines, 726for circular and elliptic cylinders, 728, 730for parallel cylinders, 770for plates, 736–738, 770for strip, 770for wedges, 737under conformal transformations, 730–738
complex power series, 136
complex powers, 102–103complex variables, seefunctions of a complex
variable andpower series in a complex
variable andcomplex integrals
components
of a complex number, 87of a vector, 221–222
in a non-orthogonal basis, 238uniqueness, 248
conditional (constrained) variation, 844–846conditional convergence, 127conditional distributions, 1040conditional probability, seeprobability,
conditional
cone
surface area of, 75–76volume of, 76
confidence interval, 1079
confidence region, 1084
conformal transformations (mappings), 730–738
applications, 735–738examples, 732–735properties, 730–732Schwarz–Christoffel transformation, 733–735
congruence, 907conic sections, 15
eccentricity, 16parametric forms, 17standard forms, 16
conjugacy classes, 910–912
in a class by itself, 910
conjugate roots of polynomial equations, 102connectivity of regions, 389conservative fields, 393–395
necessary and sufficient conditions, 393–395potential (function), 395
consistency, of estimator, 1073constant coefficients in ODE, 498–509
auxiliary equation, 499
constants of integration, 63, 474constrained variation, 844–846
constraints, stationary values under, see
Lagrange undetermined multiplers
continuity correction for discrete RV, 1028continuity equation, 410contour integration, 758–768
infinite integrals, 759–764inverse Laplace transforms, 765–768residue theorem, 752–754, 758–768
sinusoidal functions, 758–759summing series, 764–765
contraction of tensors, 788contradiction, proof by, 32–34contravariant
basis vectors, 810
derivative, 814
components of tensor, 805–806
definition, 810
control variates, in Monte Carlo methods, 1173convergence of infinite series, 717–718
absolute, 127, 717complex power series, 136conditional, 127necessary condition, 128power series, 135
under various manipulations, seepower
series, manipulation
ratio test, 718rearrangement of terms, 127tests for convergence, 128–134
alternating series test, 133comparison test, 128grouping terms, 132integral test, 131quotient test, 130ratio comparison test, 130ratio test (D’Alembert), 129, 135root test (Cauchy), 132, 717
convergence of numerical iteration schemes,
1156–1158
convolution
Fourier tranforms, seeFourier transforms,
convolution
Laplace tranforms, seeLaplace transforms,
convolution
convolution theorem
Fourier transforms, 454Laplace transforms, 463
Coordinate geometry, 15–18
straight line, 15
coordinate systems, seeCartesian, curvilinear,
cylindrical polar, plane polar andspherical
polar coordinates
coordinate transformations
and integrals, seechange of variables
and matrices, seesimilarity transformations
general, 809–814
relative tensors, 812
tensor transformations, 811weight, 813
orthogonal, 781
coplanar vectors, 229correlation functions, 455–457
auto-correlation, 456cross-correlation, 455energy spectrum, 456Parseval’s theorem, 457Wiener–Kinchin theorem, 456
1210
INDEX
correlation matrix, of sample, 1072
correlation of bivariate distributions, 1042–1049correlation of sample data, 1072correspondence principle in quantum mechanics,
1056
cosets and congruence, 907cosh, hyperbolic cosine, 105, 720, see also
hyperbolic functions
cosine
in terms of exponential functions, 105Maclaurin series for, 143orthogonality relations, 423
counting irreps, seecharacters, counting irreps
coupled pendulums, 335, 337covariance matrix, of linear least squares
estimators, 1116
covariance matrix, of sample, 1072covariance of bivariate distributions, 1042–1049covariance of sample data, 1072covariant
basis vector, 810
derivative, 814
components of tensor, 805–806
definition, 810
derivative, 817
of scalar, 820semi-colon notation, 818
differentiation, 817–820
CPF,seeprobability functions, cumulative
Cramer determinant, 304–305Cramer’s rule, 304–305cross product, seevector product
cross-correlation functions, 455crystal lattice, 151crystal point groups, 924cube roots of unity, 101cube, rotational symmetries of, 956curl of a vector field, 359
as a determinant, 359
as integral, 404, 406
curl curl, 362in curvilinear coordinates, 374in cylindrical polars, 366in spherical polars, 368Stoke’s theorem, 412–415tensor form, 823
current-carrying wire, magnetic potential, 662Curvature, 53–56
circle of, 54of a function, 53radius of, 54
curvature of space curves, 348curves, seeplane curves andspace curves
curvilinear coordinates, 370–375
basis vectors, 370length and volume elements, 371scale factors, 370surfaces and curves, 370tensors, 804–826vector operators, 373–375cut plane, 762
cycle notation for permutations, 899cyclic groups, 903, 940cyclic relation for partial derivatives, 160cycloid, 376, 844cylinders, conducting, seecomplex potentials,
for circular and elliptic cylinders
cylindrical polar coordinates, 363–367
area element, 366basis vectors, 364Laplace equation, 661–664length element, 366vector operators, 363–367volume element, 366
δ-function (Dirac), seeDirac δ-function
δ
ij(δj
i), Kronecker delta, tensor, seeKronecker
delta, δij(δj
i), tensor
D’Alembert’s ratio test, 129, 718
in convergence of power series, 135
D’Alembert’s solution to wave equation, 627damped harmonic oscillators, 243
and Parseval’s theorem, 457
data modelling, maximum-likelihood, 1097de Broglie relation, 442, 642, 703de Moivre’s theorem, 98, 758
applications, 98–102
finding the nth roots of unity, 100–101
solving polynomial equations, 101–102trigonometric identities, 98–100
deconvolution, seeFourier transforms,
deconvolution
defective matrices, 283, 316degeneracy
breaking of, 953–955of normal modes, 952
degenerate eigenvalues, 280, 287degenerate kernel, seekernel of integral
equations, separable
degree
of polynomial equation, 2
degree of ODE, 474del∇,seegradient operator (grad)
del squared ∇
2(Laplacian), 358, 609
as integral, 406in curvilinear coordinates, 374in cylindrical polar coordinates, 366in polar coordinates, 658
in spherical polar coordinates, 368, 675
tensor form, 822
delta function (Dirac), seeDirac δ-function
dependent random variables, 1038–1047derivative, see also differentiation
absolute, 824–826covariant, 817Fourier transform of, 450Laplace transform of, 461normal, 356of a function of a complex variable, 711
1211
INDEX
of a vector, 340
of basis vectors, 342–343of composite vector expressions, 343–344of function of a function, 47–48of hyperbolic functions, 109–112of products, 45–47, 49–51of quotients, 48of simple functions, 45ordinary, first, second and nth, 43–44
partial, seepartial differentiation
total, 157
derivative method for second series solution of
ODE, 551–554
determinant form
and/epsilon1
ijk, 791
for curl, 359
determinants, 264–268
adding rows or columns, 267and singular matrices, 268as product of eigenvalues, 292evaluation
using /epsilon1
ijk, 791
using Laplace expansion, 264–265
identical rows or columns, 267in terms of cofactors, 264–265
interchanging two rows or two columns, 267
Jacobian representation, 204, 208, 209notation, 264of Hermitian conjugate matrices, 267of order three, in components, 265of transpose matrices, 266product rule, 267properties, 266–268, 827relationship with rank, 272–273removing factors, 267secular, 285
diagonal matrices, 273diagonalisation of matrices, 290–293
normal matrices, 291–292properties of eigenvalues, 292–293simultaneous, 337–338
diamond, unit cell, 239die throwing, seeprobability
difference method for summation of series,
122–123
difference schemes for differential equations,
1180–1183, 1190–1192
difference, finite, seefinite differences
differentiable
function of a complex variable, 711–713function of a real variable, 43
differential
definition, 44
exact and inexact, 158–159
of vectors, 344, 350total, 157
differential equations, seeordinary differential
equations andpartial differential equations
differential equations, particular
Bernoulli, 483Bessel, 541, 564–568
Chebyshev, 541Clairaut, 489diffusion, 611, 628–631, 649, 656–657, 1192Euler, 510Euler–Lagrange, 835–836Helmholtz, 671–676Hermite, 541Lagrange, 848Laguerre, 541Laplace, 612, 623, 650, 651, 1191Legendre, 540, 541, 555–564Legendre linear, 509–511Poisson, 612, 678–681Schr¨odinger, 675, 703
Schr¨odinger, 612, 854
simple harmonic oscillator, 541Sturm–Liouville, 849wave, 609–610, 622, 626–628, 647, 671, 849
differential operators, seelinear differential
operator
differentiation, see also derivative
as gradient, 43as rate of change, 42chain rule, 47–48covariant, 817–820from first principles, 42–45implicit, 48–49logarithmic, 49notation, 44of Fourier series, 430of integrals, 181–182of power series, 138partial, seepartial differentiation
product rule, 45–47, 49–51quotient rule, 48theorems, 56–58using complex numbers, 104
diffraction, seeFraunhofer diffraction
diffusion equation, 611, 621, 628–631
combination of variables, 629–631integral transforms, 681–683numerical methods, 1192separation of variables, 649simple solution, 629superposition, 656–657
diffusion of solute, 611, 629, 681–683dihedral group, 955, 958dimension of irrep, 930dimensionality of vector space, 248dipole matrix elements, 211, 950, 957dipole moments of molecules, 919–920Dirac δ-function, 361, 411, 445–449
and convolution, 453and Green’s functions, 517, 518as limit of various distributions, 449as sum of harmonic waves, 448definition, 445Fourier transform of, 449impulses, 447
1212
INDEX
point charges, 447
properties, 445reality of, 449relation to Fourier transforms, 448–449relation to Heaviside (unit step) function, 447three-dimensional, 447, 458
direct product, of groups, 914direct sum⊕, 928
direction cosines, 225Dirichlet boundary conditions, 635, 745n
Green’s functions, 688, 690–699
method of images, 693–699
Dirichlet conditions, for Fourier series, 421–422disc, moment of inertia, 211discontinuous functions and Fourier series,
426–428
discrete Fourier transforms, 468disjoint events, seemutually exclusive events
displacement kernel, seekernel of integral
equations, displacement
distance from a
line to a line, 235–236line to a plane, 236–237point to a line, 233–234point to a plane, 234–235
distributive law for
addition of matrix products, 259convolution, 453, 464inner product, 249linear operators, 254multiplication
of a matrix by a scalar, 256of a vector by a complex scalar, 226o fav e c t o rb yas c a l a r ,2 1 8
multiplication by a scalar
in a vector space of finite dimensionality,247in a vector space of infinite dimensionality,
583
scalar/dot product, 224vector/cross product, 226
div,seedivergence of vector fields
divergence of vector fields, 358
as integral, 404–405
in curvilinear coordinates, 373
in cylindrical polars, 366in spherical polars, 368tensor form, 821–822
divergence theorem
for tensors, 803for vectors, 407–408physical applications, 410–411related theorems, 409
division axiom in a group, 888division of complex numbers, 94–95dot product, seescalar product
double integrals, seemultiple integrals
drumskin, seemembrane
dual tensors, 798–799dummy variable, 62/epsilon1
ijk, Levi-Civita symbol, tensor, 790–795
and determinant, 791identities, 792–793isotropic, 794vector products, 791weight, 813
e
x,seeexponential function
E, two-dimensional irrep, 932, 944, 950eccentricity, of conic sections, 16
efficiency, of estimator, 1074
eigenequation for differential operators, 581
more general form, 582–583, 601–602
eigenfrequencies, 325
estimation using Rayleigh–Ritz method,
333–335
eigenfunctions
completeness for an Hermitian operator, 588construction of a real set for an Hermitian
operator, 590–591
definition, 582of integral equations, 876of Legendre equation, 582of simple harmonic oscillators, 582orthogonality for Hermitian operators,
589–590
eigenvalues, 277–287
characteristic equation, 285definition, 277degenerate, 287determination, 285–287estimation for ODE, 849estimation using Rayleigh–Ritz method,
333–335
notation, 278of a general square matrix, 283of a representative matrix, 942
of a unitary matrix, 283
of an Hermitian operator
reality, 588–589
of anti-Hermitian matrices, see
anti-Hermitian matrices
of Fredholm equations, 867of Hermitian matrices, seeHermitian matrices
of integral equations, 867, 875of linear differential operators
adjustment of parameters, 854–855definition, 582
error in estimate of, 852
estimation, 849–855higher eigenvalues, 852, 859Legendre equation, 582simple harmonic oscillator, 582
of linear operators, 277of normal matrices, 278–281under similarity transformation, 292–293
eigenvectors, 277–287
characteristic equation, 285definition, 277determination, 285–287normalisation condition, 278
1213
INDEX
notation, 278
of a general square matrix, 283of a unitary matrix, 283of anti-Hermitian matrices, see
anti-Hermitian matrices
of commuting matrices, 283of Hermitian matrices, seeHermitian matrices
of linear operators, 277of normal matrices, 278–281stationary properties for quadratic/Hermitian
forms, 295–296
Einstein relation, 442, 642, 703elastic deformations, 802–803
electromagnetic fields
flux, 401Maxwell’s equations, 379, 414, 828
electrostatic fields and potentials
charged split sphere, 668conducting cylinder in uniform field, 730conducting sphere in uniform field, 667from charge density, 680, 692from complex potential, 727infinite charged plate, 694, 736infinite wedge with line charge, 738ininite charged wedge, 736of line charges, 696, 726semi-infinite charged plate, 736sphere with point charge, 698
ellipse
area of, 72, 210, 391as section of quadratic surface, 297
ellipsoid, volume of, 210
elliptic PDE, 620, 623
empty event ∅, 963
end-points for variations
contributions from, 841fixed, 836variable, 841–844
energy levels of
particle in a box, 703simple harmonic oscillator, 604
energy spectrum and Fourier transforms, 456,
457
envelopes, 176–178
equations of, 177to a family of curves, 176
epimorphism, 903equilateral triangle, symmetries of, 889, 894,
923–924, 952
equivalence relations, 906–908, 910
and classes, 906congruence, 907–909examples, 912
equivalence transformations, seesimilarity
transformations
equivalent representations, 926–928, 941error function, erf, 630, 682, 1204error terms
in Fourier series, 436–437in Taylor series, 142–143errors, first and second kind, 1122
essential singularity, 724, 750estimation of eigenvalues
linear differential operator, 851–854Rayleigh–Ritz method, 333–335
estimator, maximum-likelihood, 1098
estimators (statistics), 1072
best unbiased, 1075bias, 1074central confidence interval, 1080confidence interval, 1079confidence limits, 1079confidence region, 1084consistency, 1073efficiency, 1074
minimum-variance, 1075
standard error, 1077
Euler equation
differential, 510, 528trigonometric, 96
Euler method, numerical, 1181Euler–Lagrange equation, 835–836
special cases, 836–840
even functions, seesymmetric functions
events, 962
complement of, 963
empty∅, 963
intersection of ∩, 962
mutually exclusive, 971statistically independent, 971union of∪, 963
exact differentials, 158–159exact equations, 478, 511–512
condition for, 478non-linear, 525
expectation values, seeprobability distributions,
mean
exponential distribution, 1032–1033
from Poisson, 1032MGF, 1033
exponential function
Maclaurin series for, 143of a complex variable, 95, 719relation with hyperbolic functions, 105
Fabry–P ´erot interferometer, 149
factorial function, general, 1202factorisation, of a polynomial equation, 7faithful representation, 925, 940Fermat’s principle, 846, 857Fermi–Dirac statistics, 980
Fibonacci series, 531
field lines and complex potentials, 726fields
conservative, 393–395scalar, 353tensor, 803vector, 353
fields, electrostatic, seeelectrostatic fields and
potentials
1214
INDEX
fields, gravitational, seegravitational fields and
potentials
finite differences, 1179–1180
central, 1179for differential equations, 1180–1183forward and backward, 1179from Taylor series, 1179, 1186schemes for differential equations, 1190–1192
finite groups, 885first law of thermodynamics, 179first-order differential equations, seeordinary
differential equations
Fisher distribution, seeF-distribution (Fisher)
fluids
Archimedean upthrust, 402, 416complex velocity potential, 727continuity equation, 410cylinder in uniform flow, 728flow, 727–728flux, 401, 729irrotational flow, 359
sources and sinks, 410–411, 727
stagnation points, 727velocity potential, 415, 612vortex flow, 414, 728
forward differences, 1179Fourier cosine transforms, 452Fourier series, 421–438
and separation of variables, 652–655, 657coefficients, 423–425, 431complex, 430–431differentiation, 430Dirichlet conditions, 421–422discontinuous functions, 426–428error term, 436–437examples
square-wave, 424–425x, 430, 431
x
2, 428–429
x3, 430
integration, 430non-periodic functions, 428–430orthogonality of terms, 423
complex case, 431
Parseval’s theorem, 432–433raison d’ ˆetre, 421
standard form, 423summation of series, 433symmetry considerations, 425–426
Fourier sine transforms, 451Fourier transforms, 439–459
as generalisation of Fourier series, 439–441
convolution, 452–455
and the Dirac δ-function, 453
associativity, commutativity, distributivity,
453
definition, 453resolution function, 452
convolution theorem, 454correlation functions, 455–457cosine transforms, 452
deconvolution, 455definition, 441discrete, 468evaluation using convolution theorem, 454for integral equations, 868–871for PDE, 683–686Fourier-related (conjugate) variables, 442in higher dimensions, 457–459inverse, definition, 441odd and even functions, 451Parseval’s theorem, 456–457properties: differentiation, exponential
multiplication, integration, scaling,
translation, 450
relation to Dirac δ-function, 448–449
sine transforms, 451
Fourier transforms, examples
convolution, 454damped harmonic oscillator, 457Dirac δ-function, 449
exponential decay function, 441
Gaussian (normal) distribution, 441rectangular distribution, 448spherically symmetric functions, 458two narrow slits, 454two wide slits, 444, 454
Fourier’s inversion theorem, 441Fraunhofer diffraction, 443–445
diffraction grating, 467two narrow slits, 454two wide slits, 444, 454
Fredholm integral equations, 864
eigenvalues, 867operator form, 865with separable kernel, 866–867
Fredholm theory, 874–875Frenet–Serret formulae, 349Frobenius series, 545Fuch’s theorem, 544function of a matrix, 260functional, 835functions of a complex variable, 711–725,
747–752
analyticity, 712behaviour at infinity, 725branch points, 721Cauchy’s integrals, 745–747Cauchy–Riemann relations, 713–716conformal transformations, 730–738derivative, 711differentiation, 711–716identity theorem, 748Laplace equation, 715, 725Laurent expansion, 749–752multivalued and branch cuts, 721–723, 766particular functions, 718–721poles, 724power series, 716–718real and imaginary parts, 711, 716
1215
INDEX
singularities, 712, 723–725
Taylor expansion, 747–748zeroes, 725, 754–758
functions of one real variable
decomposition into even and odd functions,
422
differentiation of, 42–51Fourier series, seeFourier series
integration of, 60–73limits, seelimits
maxima and minima of, 51–53stationary values of, 51–53
Taylor series, seeTaylor series
functions of several real variables
chain rule, 160–161differentiation of, 154–182integration of, seemultiple integrals,
evaluation
maxima and minima, 165–170
points of inflection, 165–170rates of change, 156–158saddle points, 165–170stationary values, 165–170Taylor series, 163–165
fundamental solution, 691–693fundamental theorem of
algebra, 86, 88, 770calculus, 62–63complex numbers, seede Moivre’s theorem
gamma function
as general factorial function, 1202definition and properties, 1201
Gauss’s theorem, 700
Gauss–Seidel iteration, 1160–1162Gaussian (normal) distribution N(µ, σ
2),
1021–1031
and Binomial distribution, 1027and central limit theorem, 1037
and Poisson distribution, 1029–1030
continuity correction, 1028CPF, 1023
tabulation, 1024
Fourier transform, 441integration with infinite limits, 205integration with infinite limits, 207mean and variance, 1022–1026MGF, 1027, 1030multiple, 1030–1031multivariate, 1051
sigma limits, 1025
standard variable, 1022
Gaussian (normal) distribution N(µ, σ
2),
cumulative probability function, 1178random number generation, 1178
Gaussian elimination with interchange,
1159–1160
Gaussian integration, 1168–1170general tensors
algebra, 787–790contraction, 788
contravariant, 810covariant, 810dual, 798–799metric, 806–809physical applications, 806–809, 825–826pseudotensors, 813tensor densities, 813
generalised likelihood ratio, 1124generating functions
associated Legendre polynomials, 594Bessel functions, 573, 595Chebyshev polynomials, 597Hermite polynomials, 578, 596Laguerre polynomials, 597
Legendre polynomials, 562–564, 594
generating functions, probability, 999–1009, see
alsomoment generating functions and
probability generating functions
geodesics, 825–826, 831, 856geometric distribution, 1001
geometric series, 120
Gibbs’ free energy, 181Gibbs’ phenonmenon, 427gradient of a function of
one variable, 43several real variables, 156–158
gradient of scalar, 354–358
tensor form, 821
gradient of vector, 785, 818gradient operator (grad), 354
as integral, 404in curvilinear coordinates, 373in cylindrical polars, 366in spherical polars, 368tensor form, 821
Gram–Schmidt orthogonalisation of
eigenfunctions of Hermitian operators,
589–590
eigenvectors of
Hermitian matrices, 282normal matrices, 280
functions in a Hilbert space, 584–586
gravitational fields and potentials
Laplace equation, 612Newton’s law, 345Poisson equation, 612, 678uniform disc, 706uniform ring, 676
Green’s functions, 597–601, 686–702
and boundary conditions, 518, 520and Dirac δ-function, 517
and partial differential operators, 687and Wronskian, 533diffusion equation, 684Dirichlet problems, 690–699for ODE, 188, 517–522Neumann problems, 700–702particular integrals from, 520Poisson’s equation, 689
1216
INDEX
Green’s theorems
applications, 639, 689, 743in a plane, 390–393, 413in three dimensions, 408
ground-state energy
harmonic oscillator, 855
hydrogen atom, 860
group multiplication tables, 892
order five, 904order four, 892, 894, 903order six, 897, 903order three, 904
grouping terms as a test for convergence, 132groups
Abelian, 886
associative law, 885cancellation law, 888centre, 911closure, 885cyclic, 903definition, 885–888direct product, 914division axiom, 888elements, 885
order, 889
finite, 885identity element, 885–888inverse, 885, 888isomorphic, 893mappings between, 901–903
homomorphic, 901–903image, 901isomorphic, 901
nomenclature, 944–945
non-Abelian, 894–898order, 885, 923, 924, 936, 939, 942permutation law, 889subgroups, seesubgroups
groups, examples
±1 under multiplication, 885
alternating, 958complex numbers e
iθ, 890
functions, 897
general linear, 914integers under addition, 885integers under multiplication (mod N),
891–893
matrices, 896
permutations, 898–900quaternion, 915rotation matrices, 890symmetries of an equilateral triangle, 889
H
n(x),seeHermite polynomials
Hamilton’s principle, 847Hamiltonian, 855
Hankel transforms, 465
harmonic oscillators
damped, 243, 457ground-state energy, 855Schr¨odinger equation, 855
simple, seesimple harmonic oscillator
heat flow
diffusion equation, 611, 629, 656in bar, 656–657, 683, 705in thin sheet, 631
Heaviside function, 447
relation to Dirac δ-function, 447
Heisenberg’s uncertainty principle, 441–443
Helmholtz equation, 671–676
cylindrical polars, 673plane polars, 672–673spherical polars, 674–676
Helmholtz potential, 180hemisphere, centre of mass and centroid, 198Hermite equation, 541, 593, 596Hermite polynomials H
n(x)
generating function, 578, 596orthogonality, 596Rodrigues’ formula, 596
Hermitian conjugate, 261–263
and inner product, 263product rule, 262
Hermitian forms, 293–297
positive definite and semi-definite, 295stationary properties of eigenvectors, 295–296
Hermitian kernel, seekernel of integral
equations, Hermitian
Hermitian matrices, 276
eigenvalues, 281–283
reality, 281–282
eigenvectors, 281–283
orthogonality, 282
Hermitian operators, 587–591
boundary condition for simple harmonic
oscillators, 587–588
eigenfunctions
completeness, 588orthogonality, 589–590
eigenvalues
reality, 588–589
Green’s functions, 597–601importance of, 583, 588in Sturm–Liouville equations, 591–592properties, 588–591superposition methods, 597–601
higher-order differential equations, seeordinary
differential equations
Hilbert spaces, 584–586hit or miss, in Monte Carlo methods, 1174homogeneous
boundary conditions, seeboundary
conditions, homogeneous and
inhomogeneous
differential equations, 496dimensionally, 481, 527–528simultaneous linear equations, 298
homomorphism, 901–903
kernel of, 902representation as, 925
1217
INDEX
Hooke’s law, 802
hydrogen atom, 604
s-states, 986
electron wavefunction, 211ground-state energy, 860
hydrogen molecule, symmetries of, 883
hyperbola, as section of quadratic surface, 297
hyperbolic functions, 105–112, 720
calculus of, 109–112definitions, 105, 720graphs, 105identities, 107in equations, 108inverses, 108–109
graphs, 109
trigonometric analogies, 105–107
hyperbolic PDE, 620, 623
hypergeometric distribution, 1015–1016
mean and variance, 1015
hypergeometric equation, 603hypothesis testing, 1119–1140
errors, first and second kind, 1122generalised likelihood ratio, 1124generalised likelihood ratio test, 1123goodness of fit, 1138Neyman–Pearson test, 1122null, 1120
power, 1122
rejection region, 1121simple or composite, 1120statistical tests, 1120test statistic, 1120
i,j,k(unit vectors), 223
i,s q u a r er o o to f −1, 87
identity element of a group, 885–888
uniqueness, 885, 887
identity matrices, 259, 260identity operator, 254images, method of, seemethod of images
imaginary part/term of a complex number,
86–87
importance sampling, in Monte Carlo methods,
1172
improper
integrals, 71rotations, 795–797
impulses, δ-function respresentation, 447
independent random variables, 998, 1042index of a subgroup, 908indices, of regular singular points, 546
indicial equation, 545
distinct roots with non-integral difference,
546–547
repeated roots, 547, 551, 553roots differ by integer, 548–549, 552
induction, proof by, 31–32
inequalities
amongst integrals, 73Bessel, 251, 586Schwarz, 251, 586
triangle, 251, 586
inertia, see also moments of inertia
moments and products, 800tensor, 800
inexact differentials, 158–159inexact equation, 479infinite integrals, 71
contour integration, 759–764
infinite series, seeseries
inflection
general points of, 53
stationary points of, 51–53
inhomogeneous
boundary conditions, seeboundary
conditions, homogeneous and
inhomogeneous
differential equations, 496simultaneous linear equations, 298
inner product in a vector space, see also scalar
product
of finite dimensionality, 249–250
and Hermitian conjugate, 263commutativity, 249distributivity over addition, 249
of infinite dimensionality, 584
integral equations
eigenfunctions, 876eigenvalues, 867, 875Fredholm, 864
from differential equations, 862–863
homogeneous, 864linear, 863
first kind, 864second kind, 864
nomenclature, 863singular, 864Volterra, 864
integral equations, methods for
differentiation, 871Fredholm theory, 874–875integral transforms, 868–871
Fourier, 868–871Laplace, 869
Neumann series, 872–874Schmidt–Hilbert theory, 875–878separable (degenerate) kernels, 866–867
integral test for convergence of series, 131integral transforms, see also Fourier transforms
andLaplace transforms
general form, 465Hankel transforms, 465Mellin transforms, 465
integrals, see also integration
complex, seecomplex integrals
definite, 60double, seemultiple integrals
Fourier transform of, 450improper, 71indefinite, 63
1218
INDEX
inequalities, 73, 586
infinite, 71Laplace transform of, 462limits
containing variables, 191fixed, 60variable, 62
line,seeline integrals
multiple, seemultiple integrals
non-zero, 946–947properties, 61triple, seemultiple integrals
undefined, 60
integrals of vectors, seevectors, calculus of,
integration
integrand, 60integrating factor (IF), 512
first-order ODE, 479–481
integration, see also integrals
applications, 73–77
finding the length of a curve, 74–75mean value of a function, 73–74surfaces of revolution, 75–76volumes of revolution, 76–77
as area under a curve, 60as the inverse of differentiation, 62–63formal definition, 60from first principles, 60–61in plane polar coordinates, 71–72logarithmic, 65multiple, seemultiple integrals
multivalued functions, 762–764, 766of Fourier series, 430
of functions of several real variables, see
multiple integrals
of hyperbolic functions, 109–112of power series, 138of simple functions, 63–64of singular functions, 71of sinusoidal functions, 64–65
integration constant, 63integration, methods for
by inspection, 63–64by parts, 68–70by substitution, 66–68
tsubstitution, 66–67
completing the square, 68ing the square, 67
change of variables, seechange of variables
contour, seecontour integration
Gaussian, 1168–1170
numerical, 1164–1170
partial fractions, 65–66reduction formulae, 70trigonometrical expansions, 64–65using complex numbers, 104
intersection ∩, probability for, seeprobability,
for intersection
intrinsic derivative, seeabsolute derivative
invariant tensors, seeisotropic tensorsinverse hyperbolic functions, 108–109
inverse integral transforms
Fourier, 441Laplace, 460, 765–768
uniqueness, 460
inverse matrices, 268–271
elements, 269in solution of simultaneous linear equations,
300–301
product rule, 271
properties, 270–271
inverse of a linear operator, 254inverse of a product in a group, 888inverse of element in a group
uniqueness, 885, 888
inversion theorem, Fourier’s, 441inversions as
improper rotations, 795symmetry operations, 883–884
irregular singular points, 540
irreps, 929
n-dimensional, 931, 944
counting, 937–938dimension n
λ, 939
direct sum⊕, 928
identity A 1, 942, 946
n-dimensional, 930
number in a representation, 929, 937one-dimensional, 930, 931, 935, 941, 944
orthogonality theorem, 932–934
projection operators for, 949reduction to, 938summation rules for, 939–941
irrotational vectors, 359isobaric ODE, 482
non-linear, 527–528
isoclines, method of, 1188, 1196isomorphic groups, 893–898, 900, 901
isomorphism (mapping), 902
isotope decay, 490, 531isotropic (invariant) tensors, 793–795, 802iteration schemes
convergence of, 1156–1158for algebraic equations, 1150–1158for differential equations, 1185for integral equations, 872–875Gauss–Seidel, 1160–1162
order of convergence, 1157
J
ν(z),seeBessel functions
j,s q u a r er o o to f −1, 87
j/lscript(z),seespherical Bessel functions
Jacobians
analogy with derivatives, 210and change of variables, 209–210definition in
three dimensions, 208
two dimensions, 204
general properties, 209–210in terms of a determinant, 204, 208, 209
1219
INDEX
joint distributions, seebivariate distributions
andmultivariate distributions
Jordan’s lemma, 761–762
kernel of a homomorphism, 902, 905
kernel of an integral transform, 465
kernel of integral equations
displacement, 868Hermitian, 875
of form exp( −ixz), 869–871
of linear integral equations, 863
resolvent, 873, 874
separable (degenerate), 866–867
kinetic energy of oscillating system, 322
Klein-Gordon equation, 643
Kronecker delta δ
ijand orthogonality, 249
Kronecker delta, δij(δj
i), tensor, 777, 790–795,
805, 811
identities, 792–793
isotropic, 794
vector products, 791
L’Hˆopital’s rule, 145–147
Lagrange equations, 848
and energy conservation, 856
Lagrange undetermined multipliers, 170–176
and ODE eigenvalue estimation, 851
application to stationary properties of the
eigenvectors of quadratic/Hermitian forms,
295
for functions of more than two variables,
172–176
in deriving the Boltzmann distribution,
174–176
integral constraints, 844
with several constraints, 172–176
Lagrange’s identity, 230
Lagrange’s theorem, 907–908
and the order of a subgroup, 904and the order of an element, 904
Lagrangian, 848, 856
Laguerre equation, 541, 596–597
Laguerre polynomials L
n(x)
generating function, 597
orthogonality, 596
Rodrigues’ formula, 577, 596
Lam´e constants, 802
lamina: mass, centre of mass and centroid,
196–198
Laplace equation, 612
expansion methods, 676–678
in three dimensions
cylindrical polars, 661–664
spherical polars, 664–671
in two dimensions, 621, 623, 650, 651
and analytic functions, 715
and conformal transformations, 735–738
numerical method for, 1191, 1197
plane polars, 658–660
separated variables, 650uniqueness of solution, 676
with specified boundary values, 699, 701
Laplace expansion, 264–265Laplace transforms, 459–465, 765
convolution
associativity, commutativity, distibutivity,464definition, 463
convolution theorem, 463definition, 459for ODE with constant coefficients, 507–509for PDE, 681–683inverse, 460, 765–768
uniqueness, 460
properties: translation, exponential
multiplication, etc., 462
table for common functions, 461
Laplace transforms, examples
constant, 459derivatives, 461exponential function, 459integrals, 462polynomial, 459
Laplacian, seedel squared ∇
2(Laplacian)
Laurent expansion, 749–752
analytic and principal parts, 749region of convergence, 749
least squares, method of, 1113–1119
basis functions, 1115linear, 1114non-linear, 1118response matrix, 1115
Legendre equation, 540, 541, 555–564, 582,
593–594
as an example of a Sturm–Liouville equation,
591
associated, seeassociated Legendre equation
general series solution, 556
Legendre functions, 556
of second kind, 557
Legendre functions P
/lscript(x)
associated Legendre functions, 703
Legendre linear equation, 509Legendre polynomials P
/lscript(x)
orthogonality, 668
Legendre polynomials P/lscript(x), 557
associated Legendre functions, 666generating function, 562–564, 594graph of, 557in Gaussian integration, 1168normalisation, 557, 559orthogonality, 560, 594recurrence relation, 562Rodrigues’ formula, 559, 594
Leibnitz’ rule for differentiation of integrals, 181Leibniz’ theorem, 49–51length of
a vector, 222–223plane curves, 74–75, 347space curves, 347
1220
INDEX
tensor form, 831
Levi-Civita symbol, see/epsilon1ijk, Levi-Civita symbol,
tensor
likelihood function, 1097limits, 144–147
definition, 144L’Hˆopital’s rule, 145–147
of functions containing exponents, 145
of integrals, 60
containing variables, 191
of products, 144of quotients, 144–147of sums, 144
line charge, electrostatic potential, 726, 738
line integrals
and Cauchy integrals, 745–747and Stokes’ theorem, 412–415of scalars, 383–393
of vectors, 383–395
physical examples, 387round closed loop, 392
line, vector equation of, 230–231linear dependence and independence
definition in a vector space, 247
of basis vectors, 221
relationship with rank, 272
linear differential operator L, 517, 551, 581
adjoint L
†, 587
eigenfunctions, seeeigenfunctions
eigenvalues, seeeigenvalues, of linear
differential operators
for Sturm-Liouville equation, 591–593Hermitian (self-adjoint), 583, 587–591
linear equations, differential
first-order ODE, 480general ODE, 496–523ODE with constant coefficients, 498–509ODE with variable coefficients, 509–523
linear equations, simultaneous, seesimultaneous
linear equations
linear independence of functions, 497
Wronskian test, 497, 538
linear integral operator K, 864
and Schmidt–Hilbert theory, 875–877Hermitian conjugate, 864inverse, 865
linear interpolation for algebraic equations,
1152–1153
linear least squares, method of, 1114linear molecules
normal modes of, 326–328
symmetries of, 919
linear operators, 252–254
associativity, 254distributivity over addition, 254eigenvalues and eigenvectors, 277in a particular basis, 253
inverse, 254
non-commutativity, 254particular: identity, null/zero,
singular/non-singular, 254
properties, 254
linear vector spaces, seevector spaces
Liouville’s theorem, 747Ln of a complex number, 102–103, 720ln
Maclaurin series for, 143of a complex number, 102–103, 720
log-likelihood function, 1100longitudinal vibrations in a rod, 610lottery (UK), and hypergeometric distribution,
1016
lower triangular matrices, 274
Maclaurin series, 141
standard expressions, 143
Madelung constant, 151
magnetic dipole, 224
magnitude of a vector, 222–223
in terms of scalar/dot product, 225
mappings between groups, seegroups, mappings
between
marginal distributions, 1040mass of non-uniform bodies, 196matrices, 246–312
as a vector space, 257
as arrays of numbers, 254as representation of a linear operator, 254column, 255elements, 254
minors and cofactors, 264
identity/unit, 259row, 255zero/null, 259
matrices, algebra of, 255
Cholesky separation , 318addition, 256–257and normal modes, seenormal modes
change of basis, 288–290
diagonalisation, seediagonalisation of
matrices
multiplication, 257–259
and common eigenvalues, 283commutator, 314non-commutativity, 259
multiplication by a scalar, 256–257
numerical methods, seenumerical methods
for simultaneous linear equations
similarity transformations, seesimilarity
transformations
simultaneous linear equations, see
simultaneous linear equations
subtraction, 256
matrices, derived
adjoint, 261–263
complex conjugate, 261–263Hermitian conjugate, 261–263inverse, seeinverse matrices
transpose, 255
1221
INDEX
matrices, properties of
anti-Hermitian, seeanti-Hermitian matrices
antisymmetric/skew-symmetric, 275determinant, seedeterminants
diagonal, 273eigenvalues, seeeigenvalues
eigenvectors, seeeigenvectors
Hermitian, seeHermitian matrices
normal, seenormal matrices
nullity, 298order, 254orthogonal, 275–276rank, 272square, 254symmetric, 275trace/spur, 263triangular, 274tridiagonal, 1162–1164, 1190unitary, seeunitary matrices
matrix elements in quantum mechanics
as integrals, 945dipole, 950–951, 957
maxima and minima (local) of a function of
constrained variables, seeLagrange
undetermined multipliers
one real variable, 51–53
sufficient conditions, 52
several real variables, 165–170
sufficient conditions, 167, 170
maximum modulus theorem, 756maximum-likelihood, method of, 1097–1113
bias, 1102data modelling, 1097estimator, 1098log-likelihood function, 1100parameter estimation, 1097transformation invariance, 1102
Maxwell’s
electromagnetic equations, 379, 414, 828thermodynamic relations, 179–181
Maxwell–Boltzmann statistics, 980mean µ
from MGF, 1005from PGF, 1000of RVD, 986–987of sample, 1066of sample: geometric, harmonic, root mean
square, 1066
mean value of a function of
one variable, 73–74several variables, 202
mean value theorem, 57–58median of RVD, 987membrane
deformed rim, 658–660normal modes, 673, 954transverse vibrations, 610, 673, 702, 954
method of images, 639, 693–699, 738
disc (section of cylinder), 699, 701infinite plate, 694intersecting plates in two dimensions, 696
sphere, 697–698, 706
metric tensor, 806–809, 812
and Christoffel symbols, 815and scale factors, 806, 821covariant derivative of, 831determinant, 806, 813
derivative of, 822
length element, 806raising/lowering index, 808, 812scalar product, 807volume element, 806, 830
MGF, seemoment generating functions
Milne’s method, 1182minimum-variance estimator, 1075minor of a matrix element, 264mixed, components of tensor, 806, 811, 818ML estimator, 1098ML estimators, 1098
bias, 1102confidence limits, 1104efficiency, 1103transformation invariance, 1102
mod N, multiplication, 891
mode of RVD, 987modulo, seemod N, multiplication
modulus
of a complex number, 90of a vector, seemagnitude of a vector
molecules
bonding in, 945, 947–950dipole moments of, 919–920symmetries of, 919
moment generating functions (MGF), 1004–1009
and central limit theorem, 1037–1038and PGF, 1005mean and variance, 1005particular distributions
binomial, 1012exponential, 1033Gaussian, 1005, 1027Poisson, 1019
properties, 1005
moments
central, 990of RVD, 989
moments of inertia
and inertia tensor, 800definition, 201of disc, 211of rectangular lamina, 201of right circular cylinder, 212of sphere, 208perpendicular axes theorem, 212
moments, vector representation of, 227momentum as first-order tensor, 782monomorphism, 903Monte Carlo methods
antithetic variates, 1174control variates, 1173
1222
INDEX
crude, 1172
hit or miss, 1174importance sampling, 1172multiple integrals, 1176random number generation, 1177stratified sampling, 1172
Monte Carlo methods, of integration, 1170–1177
Morera’s theorem, 744multinomial distribution, 1050–1051
and multiple Poisson distribution, 1060
multiple angles, trigonometric formulae, 10multiple integrals
application in finding
area and volume, 194–196mass, centre of mass and centroid, 196–198
mean value of a function of several
variables, 202moments of inertia, 201
change of variables
double integrals, 203–207general properties, 209–210
triple integrals, 207–209
definitions of
double integrals, 190–191triple integrals, 193
evaluation, 191–193notation, 191, 192, 194order of integration, 191–192, 194
caveats, 193
multiplication tables for groups, seegroup
multiplication tables
multiplication theorem, seeParseval’s theorem
multivalued functions, 721–723
integration of, 762–764
multivariate distributions, 1038, 1049–1053
change of variables, 1048–1049
Gaussian, 1051multinomial, 1050–1051
mutually exclusive events, 962, 971
n
/lscript(z),seespherical Bessel functions
nabla∇,seegradient operator (grad)
natural logarithm, seelnandLn
natural numbers, in series, 31, 124–125natural representations, 923, 952Necessary and sufficient conditions, 34–35negative
function, 583
vector, 247
Neumann boundary conditions, 635
Green’s functions, 688, 700–702method of images, 700–702self-consistency, 700
Neumann series, 872–874Newton–Raphson (NR) method, 1154–1156
order of convergence, 1157
Neyman–Pearson test, 1122
nodes of oscillation, 626non-Abelian groups, 894–898
functions, 897matrices, 896
permutations, 898–900rotations–reflections, 894
non-Cartesian coordinates, seecurvilinear,
cylindrical polar, plane polar andspherical
polar coordinates
non-linear differential equations, seeordinary
differential equations, non-linear
non-linear least squares, method of, 1118norm of
function, 584vector, 249
normal
to a plane, 232to coordinate surface, 372to surface, 352, 356, 396
normal derivative, 356normal distribution, seeGaussian (normal)
distribution
normal matrices, 277
eigenvectors
completeness, 280orthogonality, 280–281
eigenvectors and eigenvalues, 278–281
normal modes, 322–335
characteristic equation, 325coupled pendulums, 335, 337definition, 326degeneracy, 952–955frequencies of, 325linear molecular system, 326–328membrane, 673, 954normal coordinates, 326normal equations, 326rod–string system, 323–326symmetries of, 328
normalisation of
eigenfunctions, 589eigenvectors, 278functions, 585vectors, 223
null (zero)
matrix, 259, 260operator, 254space, of a matrix, 298vector, 218, 247, 583
null operation, as identity element of group, 886nullity, of a matrix, 298numerical methods for algebraic equations,
1149–1156
binary chopping, 1154convergence of iteration schemes, 1156–1158linear interpolation, 1152–1153Newton–Raphson, 1154–1156rearrangement methods, 1151–1152
numerical methods for integration, 1164–1170
Monte Carlo, 1170Gaussian integration, 1168–1170midpoint rule, 1194nomenclature, 1165
1223
INDEX
Simpson’s rule, 1167
trapezium rule, 1166–1167
numerical methods for ordinary differential
equations, 1180–1190
accuracy and convergence, 1181Adams method, 1184difference schemes, 1181–1183Euler method, 1181first-order equations, 1181–1188higher-order equations, 1188–1190isoclines, 1188Milne’s method, 1182prediction and correction, 1184–1186reduction to matrix form, 1190Runge–Kutta methods, 1186–1188
Taylor series methods, 1183–1184
numerical methods for partial differential
equations, 1190–1192
diffusion equation, 1192Laplace’s equation, 1191minimising error, 1192
numerical methods for simultaneous linear
equations, 1158–1164
Gauss–Seidel iteration, 1160–1162Gaussian elimination with interchange,
1159–1160
matrix form, 1158–1164tridiagonal matrices, 1162–1164
O(x), order of, 135
observables in quantum mechanics, 282, 588odd functions, seeantisymmetric functions
ODE, seeordinary differential equations
operators
Hermitian, seeHermitian operators
linear, seelinear operators andlinear
differential operator andlinear integral
operator
order of
approximation in Taylor series, 140nconvergence of iteration schemes, 1157group, 885group element, 889ODE, 474permutation, 900subgroup, 903
and Lagrange’s theorem, 907
tensor, 779
ordinary differential equations (ODE), see also
differential equations, particular
boundary conditions, 474, 476, 507complementary function, 497degree, 474dimensionally homogeneous, 481exact, 478, 511–512
first-order, 474–490
first-order higher-degree, 486–490
soluble for p, 486
soluble for x, 487
soluble for y, 488general form of solution, 474–476
higher-order, 496–529homogeneous, 496inexact, 479isobaric, 482, 527–528linear, 480, 496–523non-linear, 524–529
xabsent, 524
yabsent, 524
exact, 525isobaric (homogeneous), 527–528
order, 474
ordinary point, seeordinary points of ODE
particular integral (solution), 475, 498,
500–501
singular point, seesingular points of ODE
singular solution, 475, 487, 488, 490
ordinary differential equations, methods for
canonical form for second-order equations,
522
eigenfunctions, 581–602equations containing linear forms, 484–486equations with constant coefficients, 498–509Green’s functions, 517–522integrating factors, 479–481Laplace transforms, 507–509numerical, 1180–1190partially known CF, 512separable variables, 477series solutions, 537–558, 564–568undetermined coefficients, 500variation of parameters, 514–516
ordinary points of ODE, 539, 541–544
indicial equation, 549
orthogonal lines, condition for, 12orthogonal matrices, 275–276, 778, 779
general properties, seeunitary matrices
orthogonal systems of coordinates, 370orthogonal transformations, 781
orthogonalisation (Gram–Schmidt) of
eigenfunctions of an Hermitian operator,
589–590
eigenvectors of a normal matrix, 280functions in a Hilbert space, 584–586
orthogonality of
eigenfunctions of an Hermitian operator,
589–590
eigenvectors of a normal matrix, 280–281eigenvectors of an Hermitian matrix, 282functions, 584terms in Fourier series, 423, 431vectors, 223, 249
orthogonality properties of characters, 936, 944orthogonality theorem for irreps, 932–934orthonormal
basis functions, 584basis vectors, 249–250
under unitary transformation, 290
oscillations, seenormal modes
outcome, of trial, 961
1224
INDEX
outer product of two vectors, 785
P/lscript(x),seeLegendre polynomials
Pm
/lscript(x),seeassociated Legendre functions
Pappus’ theorems, 198–200parabolic PDE, 620, 623parallel axis theorem, 242parallel vectors, 227parallelepiped, volume of, 229–230parallelogram equality, 252parallelogram, area of, 227, 228parameter estimation (statistics), 1072–1097,
1140
Bessel correction, 1090error in mean, 1140mean, 1086variance, 1087–1090
parameter estimation, maximum-likelihood, 1097parameters, variation of, 514–516parametric equations
of cycloid, 376, 844of space curves, 346
of surfaces, 351
parity inversion, 944Parseval’s theorem
conservation of energy, 457for Fourier series, 432–433for Fourier transforms, 456–457
partial derivative, seepartial differentiation
partial differential equations (PDE), 608–640,
646–702, see also differential equations,
particular
arbitrary functions, 613–618boundary conditions, 614, 632–640, 656characteristics, 632–638
and equation type, 636
equation types, 620, 643first-order, 614–620general solution, 614–625homogeneous, 618inhomogeneous equation and problem,
618–620, 678–681, 686–702
particular solutions (integrals), 618–625
second-order, 620–631
partial differential equations (PDE), methods for
change of variables, 624–625, 629–631constant coefficients, 620
general solution, 622
integral transform methods, 681–686method of images, seemethod of images
numerical, 1190–1192separation of variables, seeseparation of
variables
superposition methods, 650–657with no undifferentiated term, 617–618
partial differentiation, 154–182
as gradient of a function of several real
variables, 154–155
chain rule, 160–161change of variables, 161–163definitions, 154–156
properties, 160
cyclic relation, 160reciprocity relation, 160
Partial fractions, 18–25
and degree of numerator, 21complex roots, 22repeated roots, 23
partial fractions
as a means of integration, 65–66in inverse Laplace transforms, 460, 508
partial sum, 118
particular integrals (PI), 475, see also ordinary
differential equation, methods for and
partial differential equations, methods for
partition of a
group, 906
set, 907
parts, integration by, 68–70path integrals, seeline integrals
PDE, seepartial differential equations
PDF, seeprobability functions, density functions
penalty shoot-out, 1056pendulums, coupled, 335, 337periodic function representation, seeFourier
series
permutation groups S
n, 898–900
cycle notation, 899
permutation law in a group, 889permutations, 975–981
degree, 898
distinguishable, 977order of, 900symbol
nPk, 975
perpendicular axes theorem, 212perpendicular vectors, 223, 249PF,seeprobability functions
PGF, seeprobability generating functions
PI,seeparticular integrals
plane curves, length of, 74–75
in Cartesian coordinates, 74in plane polar coordinates, 75
plane polar coordinates, 71, 342
arc length, 75, 367area element, 205, 367basis vectors, 342velocity and acceleration, 343
plane waves, 628, 649planes
and simultaneous linear equations, 305–306vector equation of, 231–232
plates, conducting, see also complex potentials,
for plates
l i n ec h a r g en e a r ,6 9 6point charge near, 694
point charges, δ-function respresentation, 447
point groups, 924points of inflection of a function of
one real variable, 51–53several real variables, 165–170
1225
INDEX
Poisson distribution Po( λ), 1016–1021
and Gaussian distribution, 1029–1030as limit of Binomial distribution, 1016, 1019mean and variance, 1018MGF, 1019multiple, 1020–1021
recurrence formula, 1018
Poisson equation, 606, 612, 678–681
fundamental solution, 691–693Green’s functions, 688–702uniqueness, 638–640
Poisson summation formula, 467Poisson’s ratio, 802polar coordinates, seeplane polar andcylindrical
polar andspherical polar coordinates
polar representation of complex numbers, 95–98polar vectors, 798pole, of a function of a complex variable
contours containing, 758–768order, 724, 750residue, 750–752
polynomial equations, 1–10
conjugate roots, 102factorisation, 7
multiplicities of roots, 4
number of roots, 86, 88, 770properties of roots, 9real roots, 1solution of using de Moivre’s theorem,
101–102
polynomial solutions of ODE, 544, 554–555
populations, sampling of, 1065
positive definite and semi-definite
quadratic/Hermitian forms, 295
positive semi-definite norm, 249potential energy of
ion in a crystal lattice, 151magnetic dipoles
vector representation, 224
oscillating system, 323
potential function
and conservative fields, 395complex, 725–730electrostatic, seeelectrostatic fields and
potentials
gravitational, seegravitational fields and
potentials
vector, 395
power series
and differential equations, seeseries solutions
of differential equations
interval of convergence, 135Maclaurin, seeMaclaurin series
manipulation: difference, differentiation,
integration, product, substitution, sum,
137–138
Taylor, seeTaylor series
power series in a complex variable, 136, 716–718
analyticity, 718circle and radius of convergence, 136, 717–718convergence tests, 717, 718
form, 716
power, in hypothesis testing, 1122powers, complex, 102–103, 719prediction and correction methods, 1184–1186,
1194
prime, non-existence of largest, 34principal axes of
Cartesian tensors, 800–802conductivity tensors, 801inertia tensors, 800quadratic surfaces, 297rotation symmetry, 944
principal normals of space curves, 348principal value of
complex integrals, 760complex logarithms, 103, 720
principle of the argument, 755probability, 966–1053
axioms, 967
conditional, 970–975
Bayes’ theorem, 974–975combining, 972
definition, 967for intersection ∩, 962
for union∪, 963, 967–970
probability distributions, 981, see also individual
distributions
bivariate, seebivariate distributions
change of variables, 992–999generating functions, seemoment generating
functions andprobability generating
functions
mean µ, 986–987
mean of functions, 987mode, median and quartiles, 987
moments, 989–992
multivariate, seemultivariate distributions
standard deviation σ, 988
variance σ
2, 988
probability functions (PF), 981
cumulative (CPF), 981, 983density functions (PDF), 982
probability generating functions (PGF),
999–1004
and MGF, 1005binomial, 1003
definition, 1000
geometric, 1001mean and variance, 1000–1001Poisson, 1000sums of RV, 1003trials, 1000variable sums of RV, 1003–1004
product rule for differentiation, 45–47, 49–51products of inertia, 800projection operators for irreps, 949, 958projection tensors, 828proper rotations, 795proper subgroups, 903
1226
INDEX
pseudoscalars, 796, 799
pseudotensors, 795–799, 813pseudovectors, 795–799
quadratic equations
properties of roots, 10roots of, 2
quadratic equations, complex roots of, 86–87
quadratic forms, 293–297
positive definite and semi-definite, 295quadratic surfaces, 297removing cross terms, 294stationary properties of eigenvectors, 295–296
quartiles, of RVD, 987
quaternion group, 915, 955
quotient law for tensors, 788–790quotient rule, for differentiation, 48quotient test for series, 130
radius of convergence, 136, 717
radius of curvature of space curves, 348radius of torsion of space curves, 349random number generation, 1177random numbers
non-uniform distribution, 1194
random variable distributions, seeprobability
distributions
random variables (RV), 961, 981–985
continuous, 982–985dependent, 1038–1047discrete, 981–982independent, 998, 1042sums of, 1002–1004
range, of a matrix, 298
rank of matrices, 272
and determinants, 272–273and linear dependence, 272
rank of tensors, seeorder of, tensors
rate of change of a function of
one real variable, 42several real variables, 156–158
ratio comparison test, 130ratio test (D’Alembert), 129, 718
in convergence of power series, 135
ratio theorem, 219
and centroid of a triangle, 220–221
Rayleigh–Ritz method, 333–335, 859real part/term of a complex number, 86–87real roots, of a polynomial equation, 1
rearrangement methods for algebraic equations,
1151–1152
reciprocal vectors, 237–238, 372, 804, 808
reciprocity relation for partial derivatives, 160
rectangular distribution, 1036
Fourier transform of, 448
recurrence relations, 502–507
characteristic equation, 505coefficients, 542, 543, 1163
first-order, 503
functions, 562, 569–570higher-order, 507
second-order, 505
reducible representations, 926, 928reduction formulae for integrals, 70reflections
and improper rotations, 795as symmetry operations, 883–884
reflexivity, and equivalence relations, 906regular functions, seeanalytic functions
regular representations, 939, 952regular singular points, 540, 544–546relative velocities, 222remainder term in Taylor series, 141repeated roots of auxiliary equation, 499representations, 918
definition, 924dimension of, 920, 924equivalent, 926–928faithful, 925, 940generation of, 920–926, 954irreducible, seeirreps
natural, 923, 952
product, 945–947
reducible, 926, 928regular, 939, 952
counting irreps, 940
unitary, 928
representative matrices, 921
block-diagonal, 928eigenvalues, 942inverse, 925number needed, and order of group, 924of identity, 924
residue
at a pole, 750–752theorem, 752–754
resolution function, 452resolvent kernel, 873, 874response matrix, for linear least squares, 1115rhomboid, volume of, 241Riemann tensor, 830Riemann theorem for conditional convergence,
127
Riemann zeta series, 131, 132right hand screw rule, 226Rodrigues’ formula for
associated Legendre functions, 594Chebyshev polynomials, 597
Hermite polynomials, 596
Laguerre polynomials, 577, 596Legendre polynomials, 559, 594
Rolle’s theorem, 56root test (Cauchy), 132, 717roots
of a polynomial equation, properties, 9
roots of unity, 100–101roots, of a polynomial equation, 2rope, suspended at its ends, 845rotation groups (continuous), invariant
subspaces, 930
1227
INDEX
rotation matrices as a group, 890
rotation of a vector, seecurl
rotations
as symmetry operations, 883–884axes and orthogonal matrices, 779, 780, 810improper, 795–797
invariance under, 783
product of, 780proper, 795
Rouch ´e’s theorem, 755–757
row matrix, 255Runge–Kutta methods, 1186–1188RV,seerandom variables
RVD (random variable distributions), see
probability distributions
saddle points, 165
sufficient conditions, 167, 170
sampling
correlation, 1070covariance, 1070space, 961statistics, 1065–1072with/without replacement, 971
scalar fields, 353
derivative along a space curve, 355gradient, 354–358line integrals, 383–393rate of change, 355
scalar product, 223–226
and inner product, 249and metric tensor, 807and perpendicular vectors, 223, 249for vectors with complex components, 225in Cartesian coordinates, 225
invariance, 779, 788
scalar triple product, 228–230
cyclic permutation of, 229in Cartesian coordinates, 229
determinant form, 229
interchange of dot and cross, 229
scalars, 216–217
invariance, 779zero-order tensors, 782
scale factors, 365, 368, 370
and metric tensor, 806, 821
scattering in quantum mechanics, 469Schmidt–Hilbert theory, 875–878Schr¨odinger equation, 612
constant potential, 703hydrogen atom, 675numerical solution, 1198variational approach, 854
Schwarz inequality, 251, 586Schwarz–Christoffel transformation, 733–735
second differences, 1179
second-order differential equations, seeordinary
differential equations andpartial differential
equations
secular determinant, 285self-adjoint operators, seeHermitian operators
semicircle, angle in, 18semicircular lamina, centre of mass, 200separable
kernel in integral equations, 866–867variables in ODE, 477
separation constants, 648, 650separation of variables, for PDE, 646–681
diffusion equation, 649, 655–657, 671, 685expansion methods, 676–678general method, 646–650Helmholtz equation, 671–675inhomogeneous boundary conditions, 655–657inhomogeneous equations, 678–681Laplace equation, 650–655, 658–671, 676
polar coordinates, 658–681
separation constants, 648, 650superposition methods, 650–657wave equation, 647–649, 671, 673
series, 118–144
convergence of, seeconvergence of infinite
series
differentiation of, 134finite and infinite, 119integration of, 134multiplication by a scalar, 134multiplication of (Cauchy product), 134notation, 119operations, 134summation, seesummation of series
series, particular
arithmetic, 120arithmetico-geometric, 121
Fourier, seeFourier series
geometric, 120Maclaurin, 141, 143power, seepower series
powers of natural numbers, 124–125Riemann zeta, 131, 132Taylor, seeTaylor series
series solutions of differential equations,
537–558, 564–568
about ordinary points, 541–544about regular singular points, 544–546
Frobenius series, 545
convergence, 541indicial equation, 545linear independence, 546polynomial solutions, 544, 554–555recurrence relation, 542, 543second solution, 542, 549–554
derivative method, 551–554
Wronskian method, 550, 558
shortest path, 837
and geodesics, 825, 831
similarity transformations, 288–290, 778–779,
934
properties of matrix under, 289–290unitary transformations, 290
simple harmonic oscillator, 582, 595
1228
INDEX
energy levels of, 604
equation, 541
simple poles, 724Simpson’s rule, 1167simultaneous linear equations, 297–312
and intersection of planes, 305–306homogeneous and inhomogeneous, 298singular value decomposition, 306–312solution using
Cramer’s rule, 304–305inverse matrix, 300–301numerical methods, seenumerical methods
for simultaneous linear equations
sine
in terms of exponential functions, 105Maclaurin series for, 143orthogonality relations, 423
singular and non-singular
integral equations, 864linear operators, 254matrices, 268
singular integrals, seeimproper, integrals
singular points (singularities), 712, 723–725
essential, 724, 750removable, 725
singular points of ODE, 539
irregular, 540particular equations, 541regular, 540
singular solution of ODE, 475, 487, 488, 490singular value decomposition
and simultaneous linear equations, 306–312singular values, 307
sinh, hyperbolic sine, 105, 720, see also
hyperbolic functions
skew-symmetric matrices, 275Snell’s law, 847soap films, 839–840solenoidal vectors, 358, 395solid angle
as surface integral, 401subtended by rectangle, 417
solid: mass, centre of mass and centroid,
196–198
source density, 612space curves, 346–350
arc length, 347binormal, 348curvature, 348Frenet–Serret formulae, 349parametric equations, 346principal normal, 348radius of curvature, 348radius of torsion, 349tangent vector, 348torsion, 348
spaces, seevector spaces
span of a set of vectors, 247sphere, vector equation of, 232spherical Bessel functions j
/lscript(z), 675spherical harmonics Ym
/lscript(θ, φ), 670–671
spherical polar coordinates, 367–369
area element, 368basis vectors, 368length element, 368vector operators, 367–369volume element, 208, 368
spur of a matrix, 263–264spur, of a matrix, seetrace, of a matrix
square matrices, 254square, symmetries of, 942square-wave, Fourier series for, 424–425stagnation points of fluid flow, 727standard deviation σ, 988
of sample, 1067
standing waves, 626stationary values
of functions of
one real variable, 51–53several real variables, 165–170
of integrals, 835under constraints, seeLagrange undetermined
multipliers
statistical tests, and hypothesis testing , 1120statistics, 961, 1064–1140
describing data, 1065–1072estimating parameters, 1072–1097, 1140
Stirling’s
approximation, 1027, 1203asymptotic series, 1203
Stokes’ equation, 858Stokes’ theorem, 394, 412–415
for tensors, 804physical applications, 414related theorems, 413
strain tensor, 802stratified sampling, in Monte Carlo methods,
1172
streamlines and complex potentials, 727stress tensor, 802stress waves, 829string
loaded, 857plucked, 705transverse vibrations of, 609, 848
Student’s t-distribution
comparison of means, 1131critical points table, 1130normalisation, 1128plots, 1129
Student’s t-test, 1126–1132
Student’s t-distribution
one/two-tailed confidence limits, 1130
Sturm–Liouville equations, 591–597
boundary conditions, 592examples, 593–597
associated Legendre equation, 594–595Bessel equation, 595Chebyshev equation, 597Hermite equation, 596
1229
INDEX
hypergeometric equation, 603
Laguerre equation, 596–597Legendre equation, 593–594simple harmonic oscillator, 595
manipulation to self-adjoint form, 592–593two independent variables, 860variational approach, 849–854weight function, 849
Sturm-Liouville equations
zeroes of eigenfunctions, 603
subgroups, 903–905
index, 908normal, 905order, 903
Lagrange’s theorem, 907
proper, 903trivial, 903
submatrices, 272–273subscripts and superscripts, 777
contra- and covariant, 805covariant derivative, 818dummy, 777free, 777partial derivative, 818summation convention, 777, 804
substitution, integration by, 66–68summation convention, 777, 804summation of series, 119–127
arithmetic, 120arithmetico-geometric, 121contour integration method, 764–765difference method, 122–123Fourier series method, 433geometric, 120powers of natural numbers, 124–125transformation methods, 125–127
differentiation, 125integration, 125substitution, 126
superposition methods
for ODE, 581, 597–601for PDE, 650–657
surface integrals
and divergence theorem, 407Archimedean upthrust, 402, 416of scalars, vectors, 395–402physical examples, 401
surfaces, 351–353
area of, 352
cone, 75–76solid, and Pappus’ theorem, 198–200sphere, 352
coordinate curves, 352normal to, 352, 356of revolution, 75–76parametric equations, 351quadratic, 297tangent plane, 352
symmetric functions, 422
and Fourier series, 425–426and Fourier transforms, 451
symmetric matrices, 275
general properties, seeHermitian matrices
symmetric tensors, 787symmetry operations
on molecules, 883
order of application, 886
symmetry, and equivalence relations, 906
tsubstitution, 66–67
tan
−1x, Maclaurin series for, 143
tangent planes to surfaces, 352tangent vectors to space curves, 348tanh, hyperbolic tangent, seehyperbolic
functions
Taylor series, 139–144
and finite differences, 1179, 1186
and Taylor’s theorem, 139–142, 747approximation errors, 142–143
in numerical methods, 1156, 1166
as solution of ODE, 1183–1184for functions of a complex variable, 747–748for functions of several real variables, 163–165remainder term, 141required properties, 139standard forms, 139
tensors, seeCartesian tensors andCartesian
tensors, particular andgeneral tensors
test statistic, 1120tetrahedral group, 957tetrahedron
mass of, 197volume of, 195
thermodynamics
first law of, 179Maxwell’s relations, 179–181
top-hat function, seerectangular distribution
torque, vector representation of, 227torsion of space curves, 348total derivative, 157total differential, 157trace of a matrix, 263–264
and second-order tensors, 788as sum of eigenvalues, 285, 292
invariance under similarity transformations,
289, 934
trace formula, 292
transcendental equations, 1150transformation matrix, 288, 294transformations
active and passive, 797
conformal, 730–738coordinate, seecoordinate transformations
similarity, seesimilarity transformations
transforms, integral, seeintegral transforms and
Fourier transforms andLaplace transforms
transients
in diffusion equation, 656in electric circuits, 491
transitivity, and equivalence relations, 906
1230
INDEX
transpose of a matrix, 255, 260–261
product rule, 261
transverse vibrations
membrane, 610, 673, 702
rod, 704
string, 609
trapezium rule, 1166–1167trial functions
for eigenvalue estimation, 852
for particular integrals of ODE, 500
trials, 961triangle inequality, 251, 586triangle, centroid of, 220–221triangular matrices, 274tridiagonal matrices, 1162–1164, 1190, 1193
trignometric identities, 15
trigonometric identities, 10triple integrals, seemultiple integrals
triple scalar product, seescalar triple product
triple vector product, seevector triple product
uncertainty principle (Heisenberg), 441–443
undetermined coefficients, method of, 500undetermined multipliers, seeLagrange
undetermined multipliers
uniform distribution, 1036union∪, probability for, seeprobability for
union
uniqueness theorem
Laplace equation, 676Poisson equation, 638–640
unit step function, seeHeaviside function
unit vectors, 223
unitary
matrices, 276
eigenvalues and eigenvectors, 283
representations, 928transformations, 290
upper triangular matrices, 274
variable end-points, seeend-points for
variations, variable
variable, dummy, 62
variables, separation of, seeseparation of
variables
variance σ
2, 988
from MGF, 1005
from PGF, 1001of dependent RV, 1044of sample, 1067
variation of parameters, 514–516variation, constrained, 844–846
variational principles, physical, 846–849
Fermat, 846Hamilton, 847
variations, calculus of, seecalculus of variations
vector operators, 353–375
acting on sums and products, 360–361
combinations of, 361–363
curl, 359, 374del∇, 354
del squared ∇
2, 358
divergence (div), 358geometrical definitions, 404–406gradient operator (grad), 354–358, 373identities, 362, 827Laplacian, 358, 374non-Cartesian, 363–375tensor forms, 820–824
curl, 823divergence, 821–822gradient, 821Laplacian, 822
vector product, 226–228
anticommutativity, 226definition, 226determinant form, 228in Cartesian coordinates, 228non-associativity, 226
vector spaces, 247–252, 955
action of group on, 930associativity of addition, 247basis vectors, 248–249commutativity of addition, 247complex, 247defining properties, 247dimensionality, 248inequalities: Bessel, Schwarz, triangle, 251–252invariant, 930, 955matrices as an example, 257of infinite dimensionality, 583–586
associativity of addition, 583basis functions, 583–584commutativity of addition, 583defining properties, 583Hilbert spaces, 584–586inequalities: Bessel, Schwarz, triangle, 586
parallelogram equality, 252real, 247span of a set of vectors in, 247
vector triple product, 230
identities, 230non-associativity, 230
vectors
as first-order tensors, 781as geometrical objects, 246base, 342column, 255compared with scalars, 216–217component form, 221–222examples of, 216graphical representation of, 216–217irrotational, 359magnitude of, 222–223non-Cartesian, 342, 364, 368notation, 216polar and axial, 798solenoidal, 358, 395span of, 247
vectors, algebra of, 216–238
1231
INDEX
addition and subtraction, 217–218
in component form, 222
angle between, 225associativity of addition and subtraction, 217commutativity of addition and subtraction,
217
multiplication by a complex scalar, 226multiplication by a scalar, 218multiplication of, seescalar product and
vector product
outer product, 785
vectors, applications
centroid of a triangle, 220–221equation of a line, 230–231equation of a plane, 231–232equation of a sphere, 232finding distance from a
line to a line, 235–236line to a plane, 236–237point to a line, 233–234point to a plane, 234–235
intersection of two planes, 232
vectors, calculus of, 340–375
differentiation, 340–345, 350integration, 345–346line integrals, 383–395surface integrals, 395–402volume integrals, 402–403
vectors, derived quantities
curl, 359
derivative, 340differential, 344, 350divergence (div), 358reciprocal, 237–238, 372, 804, 808vector fields, 353
curl, 412divergence, 358flux, 401rate of change, 356
vectors, physical
acceleration, 341angular momentum, 241angular velocity, 227, 241, 359area, 399–401, 414area of parallelogram, 227, 228force, 216, 217, 224moment/torque of a force, 227velocity, 341
velocity vectors, 341Venn diagrams, 961–966vibrations
internal, seenormal modes
longitudinal, in a rod, 610transverse
membrane, 610, 673, 702, 858, 860rod, 704string, 609, 848
Volterra integral equation, 863, 864
differentiation methods, 871Laplace transform methods, 869volume elements
curvilinear coordinates, 371
cylindrical polars, 366
spherical polars, 208, 368
volume integrals, 402–403
and divergence theorem, 407
volume of
cone, 76ellipsoid, 210
parallelepiped, 229
rhomboid, 241tetrahedron, 195
volumes
as surface integrals, 403, 407
in many dimensions, 213of regions, using multiple integrals, 194–196
volumes of revolution, 76–77
and surface area & centroid, 198–200
wave equation, 609–610, 621, 849
boundary conditions, 626–628
characteristics, 637
from Maxwell’s equations, 379in one dimension, 622, 626–628
in three dimensions, 628, 647, 671
standing waves, 626
wave number, 443, 626nwave packet, 442
wave vector, k, 443
wavefunction of electron in hydrogen atom, 211wedge product, seevector product
weight
of relative tensor, 813of variable, 483
weight function, 582–583, 849
Wiener–Kinchin theorem, 456
work done
by force, 387
vector representation, 224
Wronskian
and Green’s functions, 533for second solution of ODE, 550, 558
from ODE, 538
test for linear independence, 497, 538
X-ray scattering, 241
Y
m
/lscript(θ, φ),seespherical harmonics
Yν(z), Bessel functions of second kind, 568
Young’s modulus, 610, 802
z, as a complex number, 87
z∗, as complex conjugate, 92–94
zero (null)
matrix, 259, 260
operator, 254
vector, 218, 247, 583
zero-order tensors, 781–784
zeroes of a function of a complex variable, 725
location of, 754–758, 771
1232
INDEX
order, 725, 750
principle of the argument, 755
Rouch ´e’s theorem, 755, 757
zeroes of Sturm-Liouville eigenfunctions, 603
zeroes, of a polynomial, 2
zeta series (Riemann), 131, 132
z-plane, seeArgand diagram
1233