Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books / PDF Originals
Chow T.L. Mathematical methods for physicists.. a concise introduction (CUP, 2000)(569s)
PDF · 569 pages · 3.3 MB
Open PDF file
Published textbook by Tai L. Chow of California State University, Stanislaus, written for a two-semester intermediate course in mathematical physics. The table of contents covers vector and tensor analysis, ordinary differential equations, matrix algebra, Fourier series and integrals, linear vector spaces, complex variables, special functions, calculus of variations, Laplace transforms, partial differential equations and integral equations. It is a book by someone else, kept in the archive's collection of downloaded physics books.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
Mathematical Methods
for Physicists:
A concise introduction
CAMBRIDGE UNIVERSITY PRESSTAI L. CHOW
Mathematical Methods for Ph/C121sicists
A concise introduction
This text is designed for an intermediate-level, two-semester undergraduate course
in mathematical physics. It provides an accessible account of most of the current,
important mathematical tools required in physics these days. It is assumed that
the reader has an adequate preparation in general physics and calculus.
The book bridges the gap between an introductory physics course and more
advanced courses in classical mechanics, electricity and magnetism, quantum
mechanics, and thermal and statistical physics. The text contains a large numberof worked examples to illustrate the mathematical techniques developed and to
show their relevance to physics.
The book is designed primarily for undergraduate physics majors, but could
also be used by students in other subjects, such as engineering, astronomy and
mathematics.
TAI L . CHOW was born and raised in China. He received a BS degree in physics
from the National Taiwan University, a Masters degree in physics from CaseWestern Reserve University, and a PhD in physics from the University of
Rochester. Since 1969, Dr Chow has been in the Department of Physics at
California State University, Stanislaus, and served as department chairman for17 years, until 1992. He served as Visiting Professor of Physics at University of
California (at Davis and Berkeley) during his sabbatical years. He also worked as
Summer Faculty Research Fellow at Stanford University and at NASA. Dr Chow
has published more than 35 articles in physics journals and is the author of two
textbooks and a solutions manual.
PUBLISHED BY CAMBRIDGE UNIVERSITY PRESS (VIRTUAL PUBLISHING) FOR AND ON BEHALF OF THE PRESS SYNDICATE OF THE UNIVERSITY OF CAMBRIDGE The Pitt Building, Trumpington Street, Cambridge CB2 IRP 40 West 20th Street, New York, NY 10011-4211, USA 477 Williamstown Road, Port Melbourne, VIC 3207, Australia http://www.cambridge.org © Cambridge University Press 2000 This edition © Cambridge University Press (Virtual Publishing) 2003 First published in printed format 2000 A catalogue record for the original printed book is available from the British Library and from the Library of Congress Original ISBN 0 521 65227 8 hardback Original ISBN 0 521 65544 7 paperback ISBN 0 511 01022 2 virtual (netLibrary Edition)
Mathematical Methods for Physicists
A concise introduction
TAI L . CHOW
/C67alifornia State University
/C67ontents
Preface xv
/C49 /C86ector and tensor anal/C121sis /C49
Vectors and scalars 1
Direction angles and direction cosines 3
Vector algebra 4
Equality of vectors 4Vector addition 4
Multiplication by a scalar 4
The scalar product 5
The vector (cross or outer) product 7
The triple scalar product /C65
/C66/C6710
The triple vector product 11Change of coordinate system 11
The linear vector space /C86
n13
Vector di/C128erentiation 15Space curves 16
Motion in a plane 17
A vector treatment of classical orbit theory 18
Vector di/C128erential of a scalar field and the gradient 20
Conservative vector field 21The vector di/C128erential operator /C114 22
Vector di/C128erentiation of a vector field 22
The divergence of a vector 22
The operator /C114
2, the Laplacian 24
The curl of a vector 24
Formulas involving /C114 27
Orthogonal curvilinear coordinates 27
v
Special orthogonal coordinate systems 32
Cylindrical coordinates
/C26; /C30; z32
Spherical coordinates ( r; ;/C30 34
Vector integration and integral theorems 35
Gauss’ theorem (the divergence theorem) 37
Continuity equation 39Stokes’ theorem 40
Green’s theorem 43
Green’s theorem in the plane 44
Helmholtz’s theorem 44Some useful integral relations 45Tensor analysis 47Contravariant and covariant vectors 48Tensors of second rank 48Basic operations with tensors 49Quotient law 50The line element and metric tensor 51Associated tensors 53Geodesics in a Riemannian space 53Covariant di/C128erentiation 55
Problems 57
/C50 /C79rdinar/C121 di/C128erential equations /C54/C50
First-order di/C128erential equations 63
Separable variables 63
Exact equations 67Integrating factors 69Bernoulli’s equation 72
Second-order equations with constant coecients 72
Nature of the solution of linear equations 73General solutions of the second-order equations 74Finding the complementary function 74Finding the particular integral 77Particular integral and the operator D
d=dx78
Rules for /C68operators 79
The Euler linear equation 83Solutions in power series 85
Ordinary and singular points of a di/C128erential equation 86Frobenius and Fuchs theorem 86
Simultaneous equations 93The gamma and beta functions 94
Problems 96CONTENTS
vi
/C51 Matri/C120 algebra /C49/C48/C48
Definition of a matrix 100
Four basic algebra operations for matrices 102
Equality of matrices 102Addition of matrices 102
Multiplication of a matrix by a number 103
Matrix multiplication 103
The commutator 107
Powers of a matrix 107
Functions of matrices 107
Transpose of a matrix 108Symmetric and skew-symmetric matrices 109
The matrix representation of a vector product 110
The inverse of a matrix 111
A method for finding ~A
ÿ1112
Systems of linear equations and the inverse of a matrix 113Complex conjugate of a matrix 114
Hermitian conjugation 114
Hermitian/anti-hermitian matrix 114
Orthogonal matrix (real) 115
Unitary matrix 116
Rotation matrices 117Trace of a matrix 121
Orthogonal and unitary transformations 121
Similarity transformation 122
The matrix eigenvalue problem 124
Determination of eigenvalues and eigenvectors 124
Eigenvalues and eigenvectors of hermitian matrices 128
Diagonalization of a matrix 129
Eigenvectors of commuting matrices 133
Cayley–Hamilton theorem 134Moment of inertia matrix 135
Normal modes of vibrations 136
Direct product of matrices 139
Problems 140
/C52 Fourier series and integrals /C49/C52/C52
Periodic functions 144
Fourier series; Euler–Fourier formulas 146
Gibb’s phenomena 150
Convergence of Fourier series and Dirichlet conditions 150CONTENTS
vii
Half-range Fourier series 151
Change of interval 152Parseval’s identity 153
Alternative forms of Fourier series 155
Integration and di/C128erentiation of a Fourier series 157
Vibrating strings 157
The equation of motion of transverse vibration 157Solution of the wave equation 158
/C82L/C67 circuit 160
Orthogonal functions 162
Multiple Fourier series 163Fourier integrals and Fourier transforms 164
Fourier sine and cosine transforms 172
Heisenberg’s uncertainty principle 173
Wave packets and group velocity 174
Heat conduction 179
Heat conduction equation 179
Fourier transforms for functions of several variables 182
The Fourier integral and the delta function 183
Parseval’s identity for Fourier integrals 186
The convolution theorem for Fourier transforms 188
Calculations of Fourier transforms 190The delta function and Green’s function method 192
Problems 195
/C53 Linear /C118ector spaces /C49/C57/C57
Euclidean n-space /C69
n199
General linear vector spaces 201
Subspaces 203
Linear combination 204Linear independence, bases, and dimensionality 204
Inner product spaces (unitary spaces) 206
The Gram–Schmidt orthogonalization process 209
The Cauchy–Schwarz inequality 210
Dual vectors and dual spaces 211
Linear operators 212
Matrix representation of operators 214
The algebra of linear operators 215
Eigenvalues and eigenvectors of an operator 217
Some special operators 217
The inverse of an operator 218CONTENTS
viii
The adjoint operators 219
Hermitian operators 220Unitary operators 221
The projection operators 222
Change of basis 224Commuting operators 225
Function spaces 226
Problems 230
/C54 Functions of a comple/C120 /C118ariable /C50/C51/C51
Complex numbers 233
Basic operations with complex numbers 234
Polar form of complex number 234
De Moivre’s theorem and roots of complex numbers 237
Functions of a complex variable 238Mapping 239
Branch lines and Riemann surfaces 240
The di/C128erential calculus of functions of a complex variable 241
Limits and continuity 241Derivatives and analytic functions 243The Cauchy–Riemann conditions 244
Harmonic functions 247
Singular points 248
Elementary functions of z249
The exponential functions e
z(or exp( z)249
Trigonometric and hyperbolic functions 251The logarithmic functions /C119lnz252
Hyperbolic functions 253
Complex integration 254
Line integrals in the complex plane 254
Cauchy’s integral theorem 257
Cauchy’s integral formulas 260
Cauchy’s integral formulas for higher derivatives 262
Series representations of analytic functions 265
Complex sequences 265
Complex series 266
Ratio test 268
Uniform covergence and the Weierstrass M/C45test 268
Power series and Taylor series 269Taylor series of elementary functions 272Laurent series 274CONTENTS
ix
Integration by the method of residues 279
Residues 279
The residue theorem 282
Evaluation of real definite integrals 283
Improper integrals of the rational functionZ1
ÿ1f
xdx 283
Integrals of the rational functions of sin and cos Z2
0/C71
sin;cosd286
Fourier integrals of the formZ1
ÿ1f
xsinmx
cosmx/C26/C27
dx 288
Problems 292
/C55 /C83pecial functions of mathematical ph/C121sics /C50/C57/C54
Legendre’s equation 296
Rodrigues’ formula for Pn
x299
The generating function for Pn
x301
Orthogonality of Legendre polynomials 304
The associated Legendre functions 307
Orthogonality of associated Legendre functions 309
Hermite’s equation 311
Rodrigues’ formula for Hermite polynomials Hn
x313
Recurrence relations for Hermite polynomials 313Generating function for the H
n
x314
The orthogonal Hermite functions 314
Laguerre’s equation 316
The generating function for the Laguerre polynomials Ln
x317
Rodrigues’ formula for the Laguerre polynomials Ln
x318
The orthogonal Laugerre functions 319
The associated Laguerre polynomials Lm
n
x320
Generating function for the associated Laguerre polynomials 320
Associated Laguerre function of integral order 321
Bessel’s equation 321
Bessel functions of the second kind Yn
x325
Hanging flexible chain 328Generating function for /C74
n
x330
Bessel’s integral representation 331Recurrence formulas for /C74
n
x332
Approximations to the Bessel functions 335Orthogonality of Bessel functions 336
Spherical Bessel functions 338CONTENTS
x
Sturm–Liouville systems 340
Problems 343
/C56 /C84he calculus of /C118ariations /C51/C52/C55
The Euler–Lagrange equation 348
Variational problems with constraints 353Hamilton’s principle and Lagrange’s equation of motion 355
Rayleigh–Ritz method 359
Hamilton’s principle and canonical equations of motion 361
The modified Hamilton’s principle and the Hamilton–Jacobi equation 364
Variational problems with several independent variables 367
Problems 369
/C57 /C84he Laplace transformation /C51/C55/C50
Definition of the Lapace transform 372
Existence of Laplace transforms 373
Laplace transforms of some elementary functions 375
Shifting (or translation) theorems 378
The first shifting theorem 378The second shifting theorem 379
The unit step function 380
Laplace transform of a periodic function 381
Laplace transforms of derivatives 382
Laplace transforms of functions defined by integrals 383
A note on integral transformations 384
Problems 385
/C49/C48 Partial di/C128erential equations /C51/C56/C55
Linear second-order partial di/C128erential equations 388
Solutions of Laplace’s equation: separation of variables 392
Solutions of the wave equation: separation of variables 402
Solution of Poisson’s equation. Green’s functions 404
Laplace transform solutions of boundary-value problems 409
Problems 410
/C49/C49 /C83imple linear integral equations /C52/C49/C51
Classification of linear integral equations 413
Some methods of solution 414
Separable kernel 414Neumann series solutions 416CONTENTS
xi
Transformation of an integral equation into a di/C128erential equation 419
Laplace transform solution 420Fourier transform solution 421
The Schmidt–Hilbert method of solution 421Relation between di/C128erential and integral equations 425
Use of integral equations 426
Abel’s integral equation 426Classical simple harmonic oscillator 427
Quantum simple harmonic oscillator 427
Problems 428
/C49/C50 /C69lements of group theor/C121 /C52/C51/C48
Definition of a group (group axioms) 430
Cyclic groups 433
Group multiplication table 434
Isomorphic groups 435
Group of permutations and Cayley’s theorem 438Subgroups and cosets 439
Conjugate classes and invariant subgroups 440
Group representations 442
Some special groups 444
The symmetry group D
2;D3446
One-dimensional unitary group U
1449
Orthogonal groups SO
2andSO
3450
TheSU
ngroups 452
Homogeneous Lorentz group 454
Problems 457
/C49/C51 Numerical methods /C52/C53/C57
Interpolation 459
Finding roots of equations 460
Graphical methods 460
Method of linear interpolation (method of false position) 461
Newton’s method 464
Numerical integration 466
The rectangular rule 466The trapezoidal rule 467
Simpson’s rule 469
Numerical solutions of di/C128erential equations 469
Euler’s method 470
The three-term Taylor series method 472CONTENTS
xii
The Runge–Kutta method 473
Equations of higher order. System of equations 476
Least-squares fit 477
Problems 478
/C49/C52 Introduction to probabilit/C121 theor/C121 /C52/C56/C49
A definition of probability 481Sample space 482
Methods of counting 484
Permutations 484
Combinations 485
Fundamental probability theorems 486
Random variables and probability distributions 489
Random variables 489Probability distributions 489
Expectation and variance 490
Special probability distributions 491
The binomial distribution 491
The Poisson distribution 495
The Gaussian (or normal) distribution 497
Continuous distributions 500
The Gaussian (or normal) distribution 502The Maxwell–Boltzmann distribution 503
Problems 503
/C65ppendi/C120 /C49 Preliminaries (review of fundamental concepts) /C53/C48/C54
Inequalities 507
Functions 508
Limits 510
Infinite series 511
Tests for convergence 513Alternating series test 516
Absolute and conditional convergence 517
Series of functions and uniform convergence 520
Weistrass Mtest 521
Abel’s test 522Theorem on power series 524
Taylor’s expansion 524
Higher derivatives and Leibnitz’s formula for nth derivative of
a product 528
Some important properties of definite integrals 529CONTENTS
xiii
Some useful methods of integration 531
Reduction formula 533Di/C128erentiation of integrals 534
Homogeneous functions 535
Taylor series for functions of two independent variables 535
Lagrange multiplier 536
/C65ppendi/C120 /C50 /C68eterminants /C53/C51/C56
Determinants, minors, and cofactors 540Expansion of determinants 541
Properties of determinants 542Derivative of a determinant 547
/C65ppendi/C120 /C51 /C84able of function /C70
x
1
2pZx
0eÿt2=2dt/C53/C52/C56
/C70urther reading 549
Index 551CONTENTS
xiv
Preface
This book evolved from a set of lecture notes for a course on ‘Introduction to
Mathematical Physics’, that I have given at California State University, Stanislaus
(CSUS) for many years. Physics majors at CSUS take introductory mathematical
physics before the physics core courses, so that they may acquire the expected
level of mathematical competency for the core course. It is assumed that the
student has an adequate preparation in general physics and a good understanding
of the mathematical manipulations of calculus. For the student who is in need of a
review of calculus, however, Appendix 1 and Appendix 2 are included.
This book is not encyclopedic in character, nor does it give in a highly mathe-
matical rigorous account. Our emphasis in the text is to provide an accessibleworking knowledge of some of the current important mathematical tools required
in physics.
The student will find that a generous amount of detail has been given mathe-
matical manipulations, and that ‘it-may-be-shown-thats’ have been kept to a
minimum. However, to ensure that the student does not lose sight of the develop-
ment underway, some of the more lengthy and tedious algebraic manipulations
have been omitted when possible.
Each chapter contains a number of physics examples to illustrate the mathe-
matical techniques just developed and to show their relevance to physics. They
supplement or amplify the material in the text, and are arranged in the order in
which the material is covered in the chapter. No e/C128ort has been made to trace theorigins of the homework problems and examples in the book. A solution manual
for instructors is available from the publishers upon adoption.
Many individuals have been very helpful in the preparation of this text. I wish
to thank my colleagues in the physics department at CSUS.
Any suggestions for improvement of this text will be greatly appreciated.
Turlock/C44 /C67alifornia
TAI L . CHOW
2000
xv
1
/C86ector and tensor analysis
/C86ectors and scalars
Vector methods have become standard tools for the physicists. In this chapter we
discuss the properties of the vectors and vector fields that occur in classical
physics. We will do so in a way, and in a notation, that leads to the formation
of abstract linear vector spaces in Chapter 5.
A physical quantity that is completely specified, in appropriate units, by a single
number (called its magnitude) such as volume, mass, and temperature is called a
scalar. Scalar quantities are treated as ordinary real numbers. They obey all the
regular rules of algebraic addition, subtraction, multiplication, division, and soon.
There are also physical quantities which require a magnitude and a direction for
their complete specification. These are called vectors iftheir combination with
each other is commutative (that is the order of addition may be changed withouta/C128ecting the result). Thus not all quantities possessing magnitude and direction
are vectors. Angular displacement, for example, may be characterised by magni-
tude and direction but is not a vector, for the addition of two or more angular
displacements is not, in general, commutative (Fig. 1.1).
In print, we shall denote vectors by boldface letters (such as /C65) and use ordin-
ary italic letters (such as A) for their magnitudes; in writing, vectors are usually
represented by a letter with an arrow above it such as /C126A. A given vector /C65(or/C126A)
can be written as
/C65A^A;
1:1
where Ais the magnitude of vector /C65and so it has unit and dimension, and ^Ais a
dimensionless unit vector with a unity magnitude having the direction of /C65. Thus
^A/C65=A.
1
A vector quantity may be represented graphically by an arrow-tipped line seg-
ment. The length of the arrow represents the magnitude of the vector, and the
direction of the arrow is that of the vector, as shown in Fig. 1.2. Alternatively, avector can be specified by its components (projections along the coordinate axes)
and the unit vectors along the coordinate axes (Fig. 1.3):
/C65A
1^e1A2^e2A^e3X3
i1Ai^ei;
1:2
where ^ei(i1;2;3) are unit vectors along the rectangular axes xi
x1x;x2y;
x3z; they are normally written as ^i;^j;^kin general physics textbooks. The
component triplet ( A1;A2;A3) is also often used as an alternate designation for
vector /C65:
/C65
A1;A2;A3:
1:2a
This algebraic notation of a vector can be extended (or generalized) to spaces of
dimension greater than three, where an ordered n-tuple of real numbers,
(A1;A2;...;An), represents a vector. Even though we cannot construct physical
vectors for n/C623, we can retain the geometrical language for these n-dimensional
generalizations. Such abstract ‘‘vectors’’ will be the subject of Chapter 5.
2VECTOR AND TENSOR ANALYSIS
Figure 1.1. Rotation of a parallelpiped about coordinate axes.
Figure 1.2. Graphical representation of vector /C65.
/C68irection angles and direction cosines
We can express the unit vector ^Ain terms of the unit coordinate vectors ^ei.F r o m
Eq. (1.2), /C65A1^e1A2^e2A^e3, we have
/C65AA1
A^e1A2
A^e2A3
A^e3
A^A:
Now A1=Acos;A2=Acos/C12,a n d A3=Acos/C13are the direction cosines of
the vector /C65, and ,/C12,a n d /C13are the direction angles (Fig. 1.4). Thus we can write
/C65A
cos^e1cos/C12^e2cos/C13^e3A^A;
it follows that
^A
cos^e1cos/C12^e2cos/C13^e3
cos;cos/C12;cos/C13:
1:3
3DIRECTION ANGLES AND DIRECTION COSINES
Figure 1.3. A vector /C65in Cartesian coordinates.
Figure 1.4. Direction angles of vector /C65.
/C86ector algebra
Equality of vectors
Two vectors, say /C65and/C66, are equal if, and only if, their respective components
are equal:
/C65/C66or
A1;A2;A3
B1;B2;B3
is equivalent to the three equations
A1B1;A2B2;A3B3:
Geometrically, equal vectors are parallel and have the same length, but do not
necessarily have the same position.
/C86ector addition
The addition of two vectors is defined by the equation
/C65/C66
A1;A2;A3
B1;B2;B3
A1B1;A2B2;A3B3:
That is, the sum of two vectors is a vector whose components are sums of thecomponents of the two given vectors.
We can add two non-parallel vectors by graphical method as shown in Fig. 1.5.
To add vector /C66to vector /C65, shift /C66parallel to itself until its tail is at the head of
/C65. The vector sum /C65/C66is a vector /C67drawn from the tail of /C65to the head of /C66.
The order in which the vectors are added does not a/C128ect the result.
/C77ultiplication by a scalar
Ifcis scalar then
c/C65
cA
1;cA2;cA3:
Geometrically, the vector c/C65is parallel to /C65and is ctimes the length of /C65. When
cÿ1, the vector ÿ/C65is one whose direction is the reverse of that of /C65, but both
4VECTOR AND TENSOR ANALYSIS
Figure 1.5. Addition of two vectors.
have the same length. Thus, subtraction of vector /C66from vector /C65is equivalent to
adding ÿ/C66to/C65:
/C65ÿ/C66/C65
ÿ /C66:
We see that vector addition has the following properties:
(a)/C65/C66/C66/C65 (commutativity);
(b) (/C65/C66/C67/C65
/C66/C67 (associativity);
(c)/C65/C48/C48/C65/C65;
(d)/C65
ÿ /C65/C48:
We now turn to vector multiplication. Note that division by a vector is not
defined: expressions such as k=/C65or/C66=/C65are meaningless.
There are several ways of multiplying two vectors, each of which has a special
meaning; two types are defined.
/C84he scalar product
The scalar (dot or inner) product of two vectors /C65and/C66is a real number defined
(in geometrical language) as the product of their magnitude and the cosine of the
(smaller) angle between them (Figure 1.6):
/C65/C66ABcos
0:
1:4
It is clear from the definition (1.4) that the scalar product is commutative:
/C65/C66/C66/C65;
1:5
and the product of a vector with itself gives the square of the dot product of thevector:
/C65/C65A
2:
1:6
If/C65/C660 and neither /C65nor/C66is a null (zero) vector, then /C65is perpendicular to /C66.
5THE SCALAR PRODUCT
Figure 1.6. The scalar product of two vectors.
We can get a simple geometric interpretation of the dot product from an
inspection of Fig. 1.6:
BcosAprojection of /C66onto /C65multiplied by the magnitude of /C65;
AcosBprojection of /C65onto /C66multiplied by the magnitude of /C66:
If only the components of /C65and/C66are known, then it would not be practical to
calculate /C65/C66from definition (1.4). But, in this case, we can calculate /C65/C66in
terms of the components:
/C65/C66
A1^e1A2^e2A3^e3
B1^e1B2^e2B3^e3;
1:7
the right hand side has nine terms, all involving the product ^ei^ej. Fortunately,
the angle between each pair of unit vectors is 90 8, and from (1.4) and (1.6) we find
that
^ei^ejij; i;j1;2;3;
1:8
where ijis the Kronecker delta symbol
ij0;ifi6j;
1;ifij:(
1:9
After we use (1.8) to simplify the resulting nine terms on the right-side of (7), we
obtain
/C65/C66A1B1A2B2A3B3X3
i1AiBi:
1:10
The law of cosines for plane triangles can be easily proved with the application
of the scalar product: refer to Fig. 1.7, where /C67is the resultant vector of /C65and/C66.
Taking the dot product of /C67with itself, we obtain
C2/C67/C67
/C65/C66
/C65/C66
A2B22/C65/C66A2B22ABcos;
which is the law of cosines.
6VECTOR AND TENSOR ANALYSIS
Figure 1.7. Law of cosines.
A simple application of the scalar product in physics is the work /C87done by a
constant force F:/C87Fr, where ris the displacement vector of the object
moved by F.
/C84he /C118ector (cross or outer) product
The vector product of two vectors /C65and/C66is a vector and is written as
/C67/C65/C66:
1:11
As shown in Fig. 1.8, the two vectors /C65and/C66form two sides of a parallelogram.
We define /C67to be perpendicular to the plane of this parallelogram with its
magnitude equal to the area of the parallelogram. And we choose the direction
of/C67along the thumb of the right hand when the fingers rotate from /C65to/C66(angle
of rotation less than 180 8).
/C67/C65/C66ABsin^eC
0:
1:12
From the definition of the vector product and following the right hand rule, we
can see immediately that
/C65/C66ÿ/C66/C65:
1:13
Hence the vector product is not commutative. If /C65and/C66are parallel, then it
follows from Eq. (1.12) that
/C65/C660:
1:14
In particular
/C65/C650:
1:14a
In vector components, we have
/C65/C66
A1^e1A2^e2A3^e3
B1^e1B2^e2B3^e3:
1:15
7THE VECTOR (CROSS OR OUTER) PRODUCT
Figure 1.8. The right hand rule for vector product.
Using the following relations
^ei^ei0;i1;2;3;
^e1^e2^e3;^e2^e3^e1;^e3^e1^e2;
1:16
Eq. (1.15) becomes
/C65/C66
A2B3ÿA3B2^e1
A3B1ÿA1B3^e2
A1B2ÿA2B1^e3:
1:15a
This can be written as an easily remembered determinant of third order:
/C65/C66^e1^e2^e3
A1A2A3
B1B2B3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
1:17
The expansion of a determinant of third order can be obtained by diagonal multi-
plication by repeating on the right the first two columns of the determinant and
adding the signed products of the elements on the various diagonals in the result-ing array:
The non-commutativity of the vector product of two vectors now appears as a
consequence of the fact that interchanging two rows of a determinant changes itssign, and the vanishing of the vector product of two vectors in the same direction
appears as a consequence of the fact that a determinant vanishes if one of its rows
is a multiple of another.
The determinant is a basic tool used in physics and engineering. The reader is
assumed to be familiar with this subject. Those who are in need of review should
read Appendix II.
The vector resulting from the vector product of two vectors is called an axial
vector, while ordinary vectors are sometimes called polar vectors. Thus, in Eq.(1.11), /C67is a pseudovector, while /C65and/C66are axial vectors. On an inversion of
coordinates, polar vectors change sign but an axial vector does not change sign.
A simple application of the vector product in physics is the torque /C115of a force F
about a point O:/C115Fr, where ris the vector from Oto the initial point of the
force F(Fig. 1.9).
We can write the nine equations implied by Eq. (1.16) in terms of permutation
symbols /C34
ijk:
^ei^ej/C34ijk^ek;
1:16a
8VECTOR AND TENSOR ANALYSIS
a1a2a3
b1b2b3
c1c2cc2
435a
1a2
b1b2
c1c2
ÿÿÿ
------
------
------ÿ
ÿ
!ÿ
ÿ
!ÿ
ÿ
!
where /C34ijkis defined by
/C34ijk1
ÿ1
0if
i;j;kis an even permutation of
1;2;3;
if
i;j;kis an odd permutation of
1;2;3;
otherwise
for example ;if 2 or more indices are equal :8
<
:
1:18
It follows immediately that
/C34ijk/C34kij/C34jkiÿ/C34jikÿ/C34kjiÿ/C34ikj:
There is a very useful identity relating the /C34ijkand the Kronecker delta symbol:
X3
k1/C34mnk/C34ijkminjÿmjni;
1:19
X
j;k/C34mjk/C34njk2mn;X
i;j;k/C342
ijk6:
1:19a
Using permutation symbols, we can now write the vector product /C65/C66as
/C65/C66X3
i1Ai^ei/C32!
X3
j1Bj^ej/C32!
X3
i;jAiBj^ei^ejÿ
X3
i;j;kAiBj/C34ijkÿ^ek:
Thus the kth component of /C65/C66is
/C65/C66kX
i;jAiBj/C34ijkX
i;j/C34kijAiBj:
Ifk1, we obtain the usual geometrical result:
/C65/C661X
i;j/C341ijAiBj/C34123A2B3/C34132A3B2A2B3ÿA3B2:
9THE VECTOR (CROSS OR OUTER) PRODUCT
Figure 1.9. The torque of a force about a point O.
/C84he triple scalar product /C65 E(/C66/C67)
We now briefly discuss the scalar /C65
/C66/C67. This scalar represents the volume of
the parallelepiped formed by the coterminous sides /C65,/C66,/C67, since
/C65
/C66/C67ABC sincos/C104Svolume ;
Sbeing the area of the parallelogram with sides /C66and/C67, and hthe height of the
parallelogram (Fig. 1.10).
Now
/C65
/C66/C67A1^e1A2^e2A3^e3
^e1^e2^e3
B1B2B3
C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
A
1
B2C3ÿB3C2A2
B3C1ÿB1C3A3
B1C2ÿB2C1
so that
/C65
/C66/C67A1A2A3
B1B2B3
C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
1:20
The exchange of two rows (or two columns) changes the sign of the determinant
but does not change its absolute value. Using this property, we find
/C65
/C66/C67A
1A2A3
B1B2B3
C1C2C3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ÿC
1C2C3
B1B2B3
A1A2A3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C67
/C65/C66;
that is, the dot and the cross may be interchanged in the triple scalar product.
/C65
/C66/C67
/C65/C66/C67
1:21
10VECTOR AND TENSOR ANALYSIS
Figure 1.10. The triple scalar product of three vectors /C65,/C66,/C67.
In fact, as long as the three vectors appear in cyclic order, /C65!/C66!/C67!/C65, then
the dot and cross may be inserted between any pairs:
/C65
/C66/C67/C66
/C67/C65/C67
/C65/C66:
It should be noted that the scalar resulting from the triple scalar product changes
sign on an inversion of coordinates. For this reason, the triple scalar product is
sometimes called a pseudoscalar.
/C84he triple /C118ector product
The triple product /C65
/C66/C67) is a vector, since it is the vector product of two
vectors: /C65and/C66/C67. This vector is perpendicular to /C66/C67and so it lies in the
plane of /C66and/C67.I f/C66is not parallel to /C67,/C65
/C66/C67x/C66y/C67. Now dot both
sides with /C65and we obtain x
/C65/C66y
/C65/C670, since /C65/C65
/C66/C67 0.
Thus
x=
/C65/C67ÿ y=
/C65/C66
is a scalar
and so
/C65
/C66/C67x/C66y/C67/C66
/C65/C67ÿ/C67
/C65/C66:
We now show that 1. To do this, let us consider the special case when /C66/C65.
Dot the last equation with /C67:
/C67/C65
/C65/C67
/C65/C672ÿ/C652/C672;
or, by an interchange of dot and cross
ÿ
/C65/C672
/C65/C672ÿ/C652/C672:
In terms of the angles between the vectors and their magnitudes the last equationbecomes
ÿA
2C2sin2
A2C2cos2ÿA2C2ÿ A2C2sin2;
hence 1. And so
/C65
/C66/C67/C66
/C65/C67ÿ/C67
/C65/C66:
1:22
/C67hange of coordinate s/C121stem
Vector equations are independent of the coordinate system we happen to use. But
the components of a vector quantity are di/C128erent in di/C128erent coordinate systems.
We now make a brief study of how to represent a vector in di/C128erent coordinate
systems. As the rectangular Cartesian coordinate system is the basic type of
coordinate system, we shall limit our discussion to it. Other coordinate systems
11THE TRIPLE VECTOR PRODUCT
will be introduced later. Consider the vector /C65expressed in terms of the unit
coordinate vectors
^e1;^e2;^e3:
/C65A1^e1A2^e2A^e3X3
i1Ai^ei:
Relative to a new system
^e0
1;^e0
2;^e0
3that has a di/C128erent orientation from that of
the old system
^e1;^e2;^e3, vector /C65is expressed as
/C65A0
1^e0
1A0
2^e0
2A0^e0
3X3
i1A0
i^e0
i:
Note that the dot product /C65^e0
1is equal to A0
1, the projection of /C65on the direction
of^e0
1;/C65^e0
2is equal to A0
2, and /C65^e0
3is equal to A0
3. Thus we may write
A0
1
^e1^e0
1A1
^e2^e0
1A2
^e3^e0
1A3;
A0
2
^e1^e0
2A1
^e2^e0
2A2
^e3^e0
2A3;
A0
3
^e1^e0
3A1
^e2^e0
3A2
^e3^e0
3A3:9
>>=
>>;
1:23
The dot products
^ei^e0
jare the direction cosines of the axes of the new coordi-
nate system relative to the old system: ^e0
i^ejcos
x0
i;xj; they are often called the
coecients of transformation. In matrix notation, we can write the above system
of equations as
A0
1
A0
2
A0
30
B@1
CA^e1^e0
1^e2^e0
1^e3^e0
1
^e1^e0
2^e2^e0
2^e3^e0
2
^e1^e0
3^e2^e0
3^e3^e0
30
B@1
CAA1
A2
A30
B@1
CA:
The 3 3 matrix in the above equation is called the rotation (or transformation)
matrix, and is an orthogonal matrix. One advantage of using a matrix is that
successive transformations can be handled easily by means of matrix multiplica-
tion. Let us digress for a quick review of some basic matrix algebra. A full account
of matrix method is given in Chapter 3.
A matrix is an ordered array of scalars that obeys prescribed rules of addition
and multiplication. A particular matrix element is specified by its row numberfollowed by its column number. Thus a
ijis the matrix element in the ith row and
jth column. Alternative ways of representing matrix ~Aare /C91aij/C93 or the entire array
~Aa11a12:::a1n
a21a22:::a2n
::: ::: ::: :::
am1am2:::amn0
BBBB@1
CCCCA:
12VECTOR AND TENSOR ANALYSIS
~Ais an nmmatrix. A vector is represented in matrix form by writing its
components as either a row or column array, such as
~B
b11b12b13or ~Cc11
c21
c310
B@1
CA;
where b11bx;b12by;b13bz, and c11cx;c21cy;c31cz.
The multiplication of a matrix ~Aand a matrix ~Bis defined only when the
number of columns of ~Ais equal to the number of rows of ~B, and is performed
in the same way as the multiplication of two determinants: if ~C/C61~A~B, then
cijX
kaikbk/C108:
We illustrate the multiplication rule for the case of the 3 3 matrix ~Amultiplied
by the 3 3 matrix ~B:
If we denote the direction cosines ^e0
i^ejbyij, then Eq. (1.23) can be written as
A0
iX3
j1^e0
i^ejAjX3
j1ijAj:
1:23a
It can be shown (Problem 1.9) that the quantities ijsatisfy the following relations
X3
i1ijikjk
j;k1;2;3:
1:24
Any linear transformation, such as Eq. (1.23a), that has the properties required by
Eq. (1.24) is called an orthogonal transformation, and Eq. (1.24) is known as the
orthogonal condition.
/C84he linear /C118ector space /C86n
We have found that it is very convenient to use vector components, in particular,the unit coordinate vectors ^e
i(i1, 2, 3). The three unit vectors ^eiare orthogonal
and normal, or, as we shall say, orthonormal. This orthonormal propertyis conveniently written as Eq. (1.8). But there is nothing special about these
13THE LINEAR VECTOR SPACE /C86n
.
orthonormal unit vectors ^ei. If we refer the components of the vectors to a
di/C128erent system of rectangular coordinates, we need to introduce another set of
three orthonormal unit vectors ^f1;^f2, and ^f3:
^fi^fjij
i;j1;2;3:
1:8a
For any vector /C65we now write
/C65X3
i1ci^fi;and ci^fi/C65:
We see that we can define a large number of di/C128erent coordinate systems. But
the physically significant quantities are the vectors themselves and certain func-tions of these, which are independent of the coordinate system used. The ortho-
normal condition (1.8) or (1.8a) is convenient in practice. If we also admit oblique
Cartesian coordinates then the ^f
ineed neither be normal nor orthogonal; they
could be any three non-coplanar vectors, and any vector /C65can still be written as a
linear superposition of the ^fi
/C65c1^f1c2^f2c3^f3:
1:25
Starting with the vectors ^fi, we can find linear combinations of them by the
algebraic operations of vector addition and multiplication of vectors by scalars,and then the collection of all such vectors makes up the three-dimensional linear
space often called /C86
3(V for vector) or R3(/C82for real) or /C693(Efor Euclidean). The
vectors ^f1;^f2;^f3are called the base vectors or bases of the vector space /C863. Any set
of vectors, such as the ^fi, which can serve as the bases or base vectors of /C863is
called complete, and we say it spans the linear vector space. The base vectors arealso linearly independent because no relation of the form
c
1^f1c2^f2c3^f30
1:26
exists between them, unless c1c2c30.
The notion of a vector space is much more general than the real vector space
/C863. Extending the concept of /C863, it is convenient to call an ordered set of n
matrices, or functions, or operators, a ‘vector’ (or an n-vector) in the n-dimen-
sional space /C86n. Chapter 5 will provide justification for doing this. Taking a cue
from /C863, vector addition in /C86nis defined to be
x1;...;xn
y1;...;yn
x1y1;...;xnyn
1:27
and multiplication by scalars is defined by
x1;...;xn
x1;...;xn;
1:28
14VECTOR AND TENSOR ANALYSIS
where is real. With these two algebraic operations of vector addition and multi-
plication by scalars, we call /C86na vector space. In addition to this algebraic
structure, /C86nhas geometric structure derived from the length defined to be
Xn
j1x2
j/C32!1=2
x2
1 x2nq
1:29
The dot product of two n-vectors can be defined by
x1;...;xn
y1;...;ynXn
j1xjyj:
1:30
In/C86n, vectors are not directed line segments as in /C863; they may be an ordered set
ofnoperators, matrices, or functions. We do not want to become sidetracked
from our main goal of this chapter, so we end our discussion of vector space here.
/C86ector di/C128erentiation
Up to this point we have been concerned mainly with vector algebra. A vector
may be a function of one or more scalars and vectors. We have encountered, for
example, many important vectors in mechanics that are functions of time and
position variables. We now turn to the study of the calculus of vectors.
Physicists like the concept of field and use it to represent a physical quantity
that is a function of position in a given region. Temperature is a scalar field,
because its value depends upon location: to each point ( x,y,z) is associated a
temperature T
x;y;z. The function T
x;y;zis a scalar field, whose value is a
real number depending only on the point in space but not on the particular choiceof the coordinate system. A vector field, on the other hand, associates with each
point a vector (that is, we associate three numbers at each point), such as the wind
velocity or the strength of the electric or magnetic field. When described in a
rotated system, for example, the three components of the vector associated with
one and the same point will change in numerical value. Physically and geo-
metrically important concepts in connection with scalar and vector fields are
the gradient, divergence, curl, and the corresponding integral theorems.
The basic concepts of calculus, such as continuity and di/C128erentiability, can be
naturally extended to vector calculus. Consider a vector /C65, whose components are
functions of a single variable u. If the vector /C65represents position or velocity, for
example, then the parameter uis usually time t, but it can be any quantity that
determines the components of /C65. If we introduce a Cartesian coordinate system,
the vector function /C65(u) may be written as
/C65
uA
1
u^e1A2
u^e2A3
u^e3:
1:31
15VECTOR DIFFERENTIATION
/C65(u) is said to be continuous at uu0if it is defined in some neighborhood of
u0and
lim
u!u0A
uA
u0:
1:32
Note that /C65(u) is continuous at u0if and only if its three components are con-
tinuous at u0.
/C65(u) is said to be di/C128erentiable at a point uif the limit
d/C65
u
dulim
u!0/C65
uuÿ/C65
u
u
1:33
exists. The vector /C650
ud/C65
u=duis called the derivative of /C65(u); and to di/C128er-
entiate a vector function we di/C128erentiate each component separately:
/C650
uA0
1
u^e1A0
2
u^e2A0
3
u^e3:
1:33a
Note that the unit coordinate vectors are fixed in space. Higher derivatives of /C65(u)
can be similarly defined.
If/C65is a vector depending on more than one scalar variable, say u,/C118for
example, we write /C65/C65
u;/C118. Then
d/C65
/C64/C65=/C64udu
/C64/C65=/C64/C118d/C118
1:34
is the di/C128erential of /C65, and
/C64/C65
/C64ulim
u!0/C65
uu;/C118ÿ/C65
u;/C118
/C64u
1:34a
and similarly for /C64/C65=/C64/C118.
Derivatives of products obey rules similar to those for scalar functions.
However, when cross products are involved the order may be important.
/C83pace cur/C118es
As an application of vector di/C128erentiation, let us consider some basic facts about
curves in space. If /C65(u) is the position vector r(u) joining the origin of a coordinate
system and any point P
x1;x2;x3in space as shown in Fig. 1.11, then Eq. (1.31)
becomes
r
ux1
u^e1x2
u^e2x3
u^e3:
1:35
Asuchanges, the terminal point Pofrdescribes a curve /C67in space. Eq. (1.35) is
called a parametric representation of the curve /C67, and uis the parameter of this
representation. Then
r
ur
uuÿr
u
u
16VECTOR AND TENSOR ANALYSIS
is a vector in the direction of r, and its limit (if it exists) dr=duis a vector in the
direction of the tangent to the curve at
x1;x2;x3.I fuis the arc length smeasured
from some fixed point on the curve /C67, then dr=ds^Tis a unit tangent vector to
the curve /C67. The rate at which ^Tchanges with respect to sis a measure of the
curvature of /C67and is given by d^T/ds. The direction of d^T/dsat any given point on
/C67is normal to the curve at that point: ^T^T1,d
^T^T=ds0, from this we
get^Td^T=ds0, so they are normal to each other. If ^/C78is a unit vector in this
normal direction (called the principal normal to the curve), then d^T=ds/C20^/C78,
and/C20is called the curvature of /C67at the specified point. The quantity /C261=/C20is
called the radius of curvature. In physics, we often study the motion of particles
along curves, so the above results may be of value.
In mechanics, the parameter uis time t, then dr=dt/C118is the velocity
of the particle which is tangent to the curve at the specific point. Now wecan write
/C118dr
dtdr
dsds
dt/C118^T
where /C118is the magnitude of /C118, called the speed. Similarly, ad/C118=dtis the accel-
eration of the particle.
Motion in a plane
Consider a particle Pmoving in a plane along a curve /C67(Fig. 1.12). Now rr^er,
where ^eris a unit vector in the direction of r. Hence
/C118dr
dtdr
dt^errd^er
dt:
17MOTION IN A PLANE
Figure 1.11. Parametric representation of a curve.
Now d^er=dtis perpendicular to ^er. Also jd^er=dtjd=dt; we can easily verify this
by di/C128erentiating ^ercos^e1sin^e2:Hence
/C118dr
dtdr
dt^errd
dt^e;
^eis a unit vector perpendicular to ^er.
Di/C128erentiating again we obtain
ad/C118
dtd2r
dt2^erdr
dtd^er
dtdr
dtd
dt^erd2
dt2^erd
dt^e
d2r
dt2^er2dr
dtd
dt^erd2
dt2^eÿrd
dt2
^er5d^e
dtÿd
dt^er
:
Thus
ad2r
dt2ÿrd
dt2"#
^er1
rd
dtr2d
dt
^e:
/C65 /C118ector treatment of classical orbit theor/C121
To illustrate the power and use of vector methods, we now employ them to work
out the Keplerian orbits. We first prove Kepler’s second law which can be stated
as: angular momentum is constant in a central force field. A central force is a force
whose line of action passes through a single point or center and whose magnitude
depends only on the distance from the center. Gravity and electrostatic forces are
central forces. A general discussion on central force can be found in, for example,
Chapter 6 of /C67lassical Mechanics , Tai L. Chow, John Wiley, New York, 1995.
Di/C128erentiating the angular momentum Lrpwith respect to time, we
obtain
dL=dtdr=dtprdp=dt:
18VECTOR AND TENSOR ANALYSIS
Figure 1.12. Motion in a plane.
The first vector product vanishes because pmdr=dtsodr=dtandpare parallel.
The second vector product is simply rFby Newton’s second law, and hence
vanishes for all forces directed along the position vector r, that is, for all central
forces. Thus the angular momentum Lis a constant vector in central force
motion. This implies that the position vector r, and therefore the entire orbit,
lies in a fixed plane in three-dimensional space. This result is essentially Kepler’s
second law, which is often stated in terms of the conservation of area velocity,
jLj=2m.
We now consider the inverse-square central force of gravitational and electro-
statics. Newton’s second law then gives
md/C118=dtÿ
k=r2^n;
1:36
where ^nr=ris a unit vector in the r-direction, and k/C71m 1m2for the gravita-
tional force, and k/C1131/C1132for the electrostatic force in cgs units. First we note that
/C118dr=dtdr=dt^nrd^n=dt:
Then Lbecomes
Lr
m/C118mr2^n
d^n=dt:
1:37
Now consider
d
dt
/C118Ld/C118
dtLÿk
mr2
^nLÿk
mr2^nmr2
^nd^n=dt
ÿk^n
d^n=dt^nÿ
d^n=dt
^n^n:
Since ^n^n1, it follows by di/C128erentiation that ^nd^n=dt0. Thus we obtain
d
dt
/C118Lkd^n=dt;
integration gives
/C118Lk^n/C67;
1:38
where /C67is a constant vector. It lies along, and fixes the position of, the major axis
of the orbit as we shall see after we complete the derivation of the orbit. To find
the orbit, we form the scalar quantity
L2L
rm/C118mr
/C118Lmr
kCcos;
1:39
where is the angle measured from /C67(which we may take to be the x-axis) to r.
Solving for r, we obtain
rL2=km
1C=
kcosA
1/C34cos:
1:40
Eq. (1.40) is a conic section with one focus at the origin, where /C34represents the
eccentricity of the conic section; depending on its values, the conic section may be
19A VECTOR TREATMENT OF CLASSICAL ORBIT THEORY
a circle, an ellipse, a parabola, or a hyperbola. The eccentricity can be easily
determined in terms of the constants of motion:
/C34C
k1
kj
/C118Lÿk^nj
1
kj/C118Lj2k2ÿ2k^n
/C118L1=2
Now j/C118Lj2/C1182L2because /C118is perpendicular to L. Using Eq. (1.39), we obtain
/C341
k/C1182L2k2ÿ2kL2
mr"#1=2
12L2
mk21
2m/C1182ÿk
r"#1=2
12L2/C69
mk2"#1=2
;
where Eis the constant energy of the system.
/C86ector di/C128erentiation of a scalar field and the gradient
Given a scalar field in a certain region of space given by a scalar function
/C30
x1;x2;x3that is defined and di/C128erentiable at each point with respect to the
position coordinates
x1;x2;x3, the total di/C128erential corresponding to an infini-
tesimal change dr
dx1;dx2;dx3is
d/C30/C64/C30
/C64x1dx1/C64/C30
/C64x2dx2/C64/C30
/C64x3dx3:
1:41
We can express d/C30as a scalar product of two vectors:
d/C30/C64/C30
/C64x1dx1/C64/C30
/C64x2dx2/C64/C30
/C64x3dx3/C114 /C30
dr;
1:42
where
/C114/C30/C64/C30
/C64x1^e1/C64/C30
/C64x2^e2/C64/C30
/C64x3^e3
1:43
is a vector field (or a vector point function). By this we mean to each pointr
x
1;x2;x3in space we associate a vector /C114/C30as specified by its three compo-
nents ( /C64/C30=/C64 x1; /C64/C30=/C64 x2;/C64 /C30 = /C64 x3):/C114/C30is called the gradient of/C30and is often written
as grad /C30.
There is a simple geometric interpretation of /C114/C30. Note that /C30
x1;x2;x3c,
where cis a constant, represents a surface. Let rx1^e1x2^e2x3^e3be the
position vector to a point P
x1;x2;x3on the surface. If we move along the
surface to a nearby point /C81
rdr, then drdx1^e1dx2^e2dx3^e3lies in the
tangent plane to the surface at P. But as long as we move along the surface /C30has a
constant value and d/C300. Consequently from (1.41),
dr/C114/C300:
1:44
20VECTOR AND TENSOR ANALYSIS
Eq. (1.44) states that /C114/C30is perpendicular to drand therefore to the surface (Fig.
1.13). Let us return to
d/C30
/C114 /C30dr:
The vector /C114/C30is fixed at any point P, so that d/C30, the change in /C30, will depend to a
great extent on dr. Consequently d/C30will be a maximum when dris parallel to /C114/C30,
since dr/C114/C30jdrjj/C114/C30jcos, and cos is a maximum for 0. Thus /C114/C30is in
the direction of maximum increase of /C30
x1;x2;x3. The component of /C114/C30in the
direction of a unit vector ^uis given by /C114/C30^uand is called the directional deri-
vative of /C30in the direction ^u. Physically, this is the rate of change of /C30at
(x1;x2;x3in the direction ^u.
/C67onser/C118ati/C118e /C118ector field
By definition, a vector field is said to be conservative if the line integral of the
vector along any closed path vanishes. Thus, if Fis a conservative vector field
(say, a conservative force field in mechanics), then
I
Fds0;
1:45
where dsis an element of the path. A (necessary and sucient) condition for F
to be conservative is that Fcan be expressed as the gradient of a scalar, say
/C30:Fÿgrad /C30:
Zb
aFdsÿZb
agrad /C30dsÿZb
ad/C30/C30
aÿ/C30
b:
it is obvious that the line integral depends solely on the value of the scalar /C30at the
initial and final points, andH
FdsÿH
grad /C30ds0.
21CONSERVATIVE VECTOR FIELD
Figure 1.13. Gradient of a scalar.
/C84he /C118ector di/C128erential operator /C114
We denoted the operation that changes a scalar field to a vector field in Eq. (1.43)
by the symbol /C114(del or nabla):
/C114/C64
/C64x1^e1/C64
/C64x2^e2/C64
/C64x3^e3;
1:46
which is called a gradient operator. We often write /C114/C30as grad /C30, and the vector
field/C114/C30
ris called the gradient of the scalar field /C30
r. Notice that the operator
/C114contains both partial di/C128erential operators and a direction: it is a vector di/C128er-
ential operator. This important operator possesses properties analogous to those
of ordinary vectors. It will help us in the future to keep in mind that /C114acts both
as a di/C128erential operator and as a vector.
/C86ector di/C128erentiation of a /C118ector field
Vector di/C128erential operations on vector fields are more complicated because of the
vector nature of both the operator and the field on which it operates. As we know
there are two types of products involving two vectors, namely the scalar andvector products; vector di/C128erential operations on vector fields can also be sepa-
rated into two types called the curl and the divergence.
/C84he divergence of a vector
If/C86
x
1;x2;x3/C861^e1/C862^e2/C863^e3is a di/C128erentiable vector field (that is, it is
defined and di/C128erentiable at each point ( x1;x2;x3) in a certain region of space),
the divergence of /C86, written /C114/C86or div /C86, is defined by the scalar product
/C114/C86/C64
/C64x1^e1/C64
/C64x2^e2/C64
/C64x3^e3
/C861^e1/C862^e2/C863^e3
/C64/C861
/C64x1/C64/C862
/C64x2/C64/C863
/C64x3:
1:47
The result is a scalar field. Note the analogy with /C65/C66A1B1A2B2A3B3,
but also note that /C114/C866/C86/C114(bear in mind that /C114is an operator). /C86/C114is a
scalar di/C128erential operator:
/C86/C114 /C861/C64
/C64x1/C862/C64
/C64x2/C863/C64
/C64x3:
What is the physical significance of the divergence/C63 Or why do we call the scalar
product /C114/C86the divergence of /C86/C63 To answer these questions, we consider, as an
example, the steady motion of a fluid of density /C26
x1;x2;x3, and the velocity field
is given by /C118
x1;x2;x3/C1181
x1;x2;x3e1/C1182
x1;x2;x3e2/C1183
x1;x2;x3e3.W e
22VECTOR AND TENSOR ANALYSIS
now concentrate on the flow passing through a small parallelepiped AB/C67/C68E/C70/C71/C72
of dimensions dx1dx2dx3(Fig. 1.14). The x1andx3components of the velocity /C118
contribute nothing to the flow through the face AB/C67/C68 . The mass of fluid entering
AB/C67/C68 per unit time is given by /C26/C1182dx1dx3and the amount leaving the face E/C70/C71/C72
per unit time is
/C26/C1182/C64
/C26/C1182
/C64x2dx2
dx1dx3:
So the loss of mass per unit time is /C64
/C26/C1182=/C64x2dx1dx2dx3. Adding the net rate of
flow out all three pairs of surfaces of our parallelepiped, the total mass loss per
unit time is
/C64
/C64x1
/C26/C1181/C64
/C64x2
/C26/C1182/C64
/C64x3
/C26/C1183
dx1dx2dx3/C114
/C26/C118dx1dx2dx3:
So the mass loss per unit time per unit volume is /C114
/C26/C118. Hence the name
divergence.
The divergence of any vector /C86is defined as /C114/C86. We now calculate /C114
f/C86,
where fis a scalar:
/C114
f/C86/C64
/C64x1
f/C861/C64
/C64x2
f/C862/C64
/C64x3
f/C863
f/C64/C861
/C64x1/C64/C862
/C64x2/C64/C863
/C64x3
/C861/C64f
/C64x1/C862/C64f
/C64x2/C863/C64f
/C64x3
or
/C114
f/C86f/C114/C86/C86/C114f:
1:48
It is easy to remember this result if we remember that /C114acts both as a di/C128erential
operator and a vector. Thus, when operating on f/C86, we first keep ffixed and let /C114
23VECTOR DIFFERENTIATION OF A VECTOR FIELD
Figure 1.14. Steady flow of a fluid.
operate on /C86, and then we keep /C86fixed and let /C114operate on f
/C114 fis nonsense),
and as /C114fand/C86are vectors we complete their multiplication by taking their dot
product.
A vector /C86is said to be solenoidal if its divergence is zero: /C114/C860.
/C84he operator /C1142/C44 the /C76aplacian
The divergence of a vector field is defined by the scalar product of the operator /C114
with the vector field. What is the scalar product of /C114with itself /C63
/C1142/C114/C114/C64
/C64x1^e1/C64
/C64x2^e2/C64
/C64x3^e3
/C64
/C64x1^e1/C64
/C64x2^e2/C64
/C64x3^e3
/C642
/C64x2
1/C642
/C64x22/C642
/C64x23:
This important quantity
/C1142/C642
/C64x21/C642
/C64x22/C642
/C64x23
1:49
is a scalar di/C128erential operator which is called the Laplacian, after a French
mathematician of the eighteenth century named Laplace. Now, what is the diver-
gence of a gradient/C63
Since the Laplacian is a scalar di/C128erential operator, it does not change the
vector character of the field on which it operates. Thus /C1142/C30
ris a scalar field
if/C30
ris a scalar field, and /C1142/C114/C30
ris a vector field because the gradient /C114/C30
r
is a vector field.
The equation /C1142/C300 is called Laplace’s equation.
/C84he curl of a vector
If/C86
x1;x2;x3is a di/C128erentiable vector field, then the curl or rotation of /C86,
written /C114/C86(or curl /C86or rot /C86), is defined by the vector product
curl/C86/C114 /C86^e1 ^e2 ^e3
/C64
/C64x1/C64
/C64x2/C64
/C64x3
/C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
^e
1/C64/C863
/C64x2ÿ/C64/C862
/C64x3
^e2/C64/C861
/C64x3ÿ/C64/C863
/C64x1
^e3/C64/C862
/C64x1ÿ/C64/C861
/C64x2
X
i;j;k/C34ijk^ei/C64/C86k
/C64xj:
1:50
24VECTOR AND TENSOR ANALYSIS
The result is a vector field. In the expansion of the determinant the operators
/C64=/C64ximust precede /C86i;P
ijkstands forP
iP
jP
k; and /C34ijkare the permutation
symbols: an even permutation of ijkwill not change the value of the resulting
permutation symbol, but an odd permutation gives an opposite sign. That is,
/C34ijk/C34jki/C34kijÿ/C34jikÿ/C34kjiÿ/C34ikj;and
/C34ijk0 if two or more indices are equal :
A vector /C86is said to be irrotational if its curl is zero: /C114/C86
r0. From this
definition we see that the gradient of any scalar field /C30
ris irrotational. The proof
is simple:
/C114
/C114 /C30^e1 ^e2 ^e3
/C64
/C64x1/C64
/C64x2/C64
/C64x3
/C64
/C64x1/C64
/C64x2/C64
/C64x3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C30
x
1;x2;x30
1:51
because there are two identical rows in the determinant. Or, in terms of the
permutation symbols, we can write /C114
/C114 /C30as
/C114
/C114 /C30X
ijk/C34ijk^ei/C64
/C64xj/C64
/C64xk/C30
x1;x2;x3:
Now /C34ijkis antisymmetric in j,k, but /C642=/C64xj/C64xkis symmetric, hence each term in
the sum is always cancelled by another term:
/C34ijk/C64
/C64xj/C64
/C64xk/C34ikj/C64
/C64xk/C64
/C64xj0;
and consequently /C114
/C114 /C300. Thus, for a conservative vector field F,w eh a v e
curlFcurl (grad /C300.
We learned above that a vector /C86is solenoidal (or divergence-free) if its diver-
gence is zero. From this we see that the curl of any vector field /C86(r) must be
solenoidal:
/C114
/C114 /C86X
i/C64
/C64xi
/C114 /C86iX
i/C64
/C64xiX
j;k/C34ijk/C64
/C64xj/C86k/C32!
0;
1:52
because /C34ijkis antisymmetric in i,j.
If/C30
ris a scalar field and /C86(r) is a vector field, then
/C114
/C30/C86/C30
/C114 /C86
/C114 /C30/C86:
1:53
25VECTOR DIFFERENTIATION OF A VECTOR FIELD
We first write
/C114
/C30/C86^e1 ^e2 ^e3
/C64
/C64x1/C64
/C64x2/C64
/C64x3
/C30/C861/C30/C862/C30/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
then notice that
/C64
/C64x1
/C30/C862/C30/C64/C862
/C64x1/C64/C30
/C64x1/C862;
so we can expand the determinant in the above equation as a sum of two deter-
minants:
/C114
/C30/C86/C30^e1 ^e2 ^e3
/C64
/C64x1/C64
/C64x2/C64
/C64x3
/C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12^e
1 ^e2 ^e3
/C64/C30
/C64x1/C64/C30
/C64x2/C64/C30
/C64x3
/C861/C862/C863/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
/C30
/C114 /C86
/C114 /C30/C86:
Alternatively, we can simplify the proof with the help of the permutation symbols
/C34
ijk:
/C114
/C30/C86X
i;j;k/C34ij k^ei/C64
/C64xj
/C30/C86k
/C30X
i;j;k/C34ij k^ei/C64/C86k
/C64xjX
i;j;k/C34ijk^ei/C64/C30
/C64xj/C86k
/C30
/C114 /C86
/C114 /C30/C86:
A vector field that has non-vanishing curl is called a vortex field, and the curl of
the field vector is a measure of the vorticity of the vector field.
The physical significance of the curl of a vector is not quite as transparent as
that of the divergence. The following example from fluid flow will help us to
develop a better feeling. Fig. 1.15 shows that as the component /C1182of the velocity
/C118of the fluid increases with x3, the fluid curls about the x1-axis in a negative sense
(rule of the right-hand screw), where /C64/C1182=/C64x3is considered positive. Similarly, a
positive curling about the x1-axis would result from /C1183if/C64/C1183=/C64x2were positive.
Therefore, the total x1component of the curl of /C118is
curl/C1181/C64/C1183=
/C64x2ÿ/C64/C1182=/C64x3;
which is the same as the x1component of Eq. (1.50).
26VECTOR AND TENSOR ANALYSIS
Formulas in/C118ol/C118ing /C114
We now list some important formulas involving the vector di/C128erential operator /C114,
some of which are recapitulation. In these formulas, /C65and/C66are di/C128erentiable
vector field functions, and fand gare di/C128erentiable scalar field functions of
position
x1;x2;x3:
(1)/C114
f/C103f/C114/C103/C103/C114f;
(2)/C114
f/C65f/C114/C65/C114f/C65;
(3)/C114
f/C65f/C114/C65/C114f/C65;
(4)/C114
/C114 f0;
(5)/C114
/C114 /C650;
(6)/C114
/C65/C66
/C114 /C65/C66ÿ
/C114 /C66/C65;
(7)/C114
/C65/C66
/C66/C114 /C65ÿ/C66
/C114 /C65/C65
/C114 /C66ÿ
/C65/C114 /C66;
(8)/C114
/C114 /C65 /C114
/C114 /C65ÿ/C1142/C65;
(9)/C114
/C65/C66/C65
/C114 /C66/C66
/C114 /C65
/C65/C114 /C66
/C66/C114 /C65;
(10)
/C65/C114 r/C65;
(11)/C114r3;
(12)/C114r0;
(13)/C114
rÿ3r0;
(14) dF
dr/C114 F/C64F
/C64tdt
Fa di/C128erentiable vector field quantity);
(15) d’dr/C114’/C64’
/C64tdt(’a di/C128erentiable scalar field quantity).
/C79rthogonal cur/C118ilinear coordinates
Up to this point all calculations have been performed in rectangular Cartesian
coordinates. Many calculations in physics can be greatly simplified by using,
instead of the familiar rectangular Cartesian coordinate system, another kind of
27FORMULAS INVOLVING /C114
Figure 1.15. Curl of a fluid flow.
system which takes advantage of the relations of symmetry involved in the parti-
cular problem under consideration. For example, if we are dealing with sphere, wewill find it expedient to describe the position of a point in sphere by the spherical
coordinates ( r; ;/C30. Spherical coordinates are a special case of the orthogonal
curvilinear coordinate system. Let us now proceed to discuss these more generalcoordinate systems in order to obtain expressions for the gradient, divergence,
curl, and Laplacian. Let the new coordinates u
1;u2;u3be defined by specifying the
Cartesian coordinates ( x1;x2;x3) as functions of ( u1;u2;u3:
x1f
u1;u2;u3;x2/C103
u1;u2;u3;x3/C104
u1;u2;u3;
1:54
where f,g,hare assumed to be continuous, di/C128erentiable. A point P(Fig. 1.16) in
space can then be defined not only by the rectangular coordinates ( x1;x2;x3) but
also by curvilinear coordinates ( u1;u2;u3).
Ifu2andu3are constant as u1varies, P(or its position vector r) describes a curve
which we call the u1coordinate curve. Similarly, we can define the u2andu3coordi-
nate curves through P. We adopt the convention that the new coordinate system is a
right handed system, like the old one. In the new system drtakes the form:
dr/C64r
/C64u1du1/C64r
/C64u2du2/C64r
/C64u3du3:
The vector /C64r=/C64u1is tangent to the u1coordinate curve at P.I f^u1is a unit vector
atPin this direction, then ^u1/C64r=/C64u1=j/C64r=/C64u1j, so we can write /C64r=/C64u1/C1041^u1,
where /C1041j/C64r=/C64u1j. Similarly we can write /C64r=/C64u2/C1042^u2and /C64r=/C64u3/C1043^u3,
where /C1042j/C64r=/C64u2jand/C1043j/C64r=/C64u3j, respectively. Then drcan be written
dr/C1041du1^u1/C1042du2^u2/C1043du3^u3:
1:55
28VECTOR AND TENSOR ANALYSIS
Figure 1.16. Curvilinear coordinates.
The quantities /C1041;/C1042;/C1043are sometimes called scale factors. The unit vectors ^u1,^u2,
^u3are in the direction of increasing u1;u2;u3, respectively.
If^u1,^u2,^u3are mutually perpendicular at any point P, the curvilinear coordi-
nates are called orthogonal. In such a case the element of arc length dsis given by
ds2drdr/C1042
1du21/C10422du22/C10423du23:
1:56
Along a u1curve, u2andu3are constants so that dr/C1041du1^u1. Then the
di/C128erential of arc length ds1along u1atPis/C1041du1. Similarly the di/C128erential arc
lengths along u2andu3atPareds2/C1042du2,ds3/C1043du3respectively.
The volume of the parallelepiped is given by
d/C86j
/C1041du1^u1
/C1042du2^u2
/C1043du3^u3j /C1041/C1042/C1043du1du2du3
since j^u1^u2^u3j1. Alternatively d/C86can be written as
d/C86/C64r
/C64u1/C64r
/C64u2/C64r
/C64u3/C12/C12/C12/C12/C12/C12/C12/C12du
1du2du3/C64
x1;x2;x3
/C64
u1;u2;u3/C12/C12/C12/C12/C12/C12/C12/C12du
1du2du3;
1:57
where
/C74/C64
x1;x2;x3
/C64
u1;u2;u3/C64x1
/C64u1/C64x1
/C64u2/C64x1
/C64u3
/C64x2
/C64u1/C64x2
/C64u2/C64x2
/C64u3
/C64x3
/C64u1/C64x3
/C64u2/C64x3
/C64u3/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
is called the Jacobian of the transformation.
We assume that the Jacobian /C7460 so that the transformation (1.54) is one to
one in the neighborhood of a point.
We are now ready to express the gradient, divergence, and curl in terms of
u
1;u2,a n d u3.I f/C30is a scalar function of u1;u2,a n d u3, then the gradient takes the
form
/C114/C30grad /C301
/C1041/C64/C30
/C64u1^u11
/C1042/C64/C30
/C64u2^u21
/C1043/C64/C30
/C64u3^u3:
1:58
To derive this, let
/C114/C30f1^u1f2^u2f3^u3;
1:59
where f1;f2;f3are to be determined. Since
dr/C64r
/C64u1du1/C64r
/C64u2du2/C64r
/C64u3du3
/C1041du1^u1/C1042du2^u2/C1043du3^u3;
29ORTHOGONAL CURVILINEAR COORDINATES
we have
d/C30/C114/C30dr/C1041f1du1/C1042f2du2/C1043f3du3:
But
d/C30/C64/C30
/C64u1du1/C64/C30
/C64u2du2/C64/C30
/C64u3du3;
and on equating the two equations, we find
fi1
/C104i/C64/C30
/C64ui;i1;2;3:
Substituting these into Eq. (1.57), we obtain the result Eq. (1.58).
From Eq. (1.58) we see that the operator /C114takes the form
/C114^u1
/C1041/C64
/C64u1^u2
/C1042/C64
/C64u2^u3
/C1043/C64
/C64u3:
1:60
Because we will need them later, we now proceed to prove the following two
relations:
(a)j/C114uij/C104ÿ1
i;i1, 2, 3.
(b)^u1/C1042/C1043/C114u2/C114u3with similar equations for ^u2and ^u3. (1.61)
Proof: ( a) Let /C30u1in Eq. (1.51), we then obtain /C114u1^u1=/C1041and so
j/C114u1jj ^u1j/C104ÿ1
1/C104ÿ1
1;since j^u1j1:
Similarly by letting /C30u2andu3, we obtain the relations for i2 and 3.
(b) From ( a) we have
/C114u1^u1=/C1041;/C114u2^u2=/C1042;and /C114u3^u3=/C1043:
Then
/C114u2/C114u3^u2^u3
/C1042/C1043^u1
/C1042/C1043and ^u1/C1042/C1043/C114u2/C114u3:
Similarly
^u2/C1043/C1041/C114u3/C114u1and ^u3/C1041/C1042/C114u1/C114u2:
We are now ready to express the divergence in terms of curvilinear coordinates.
If/C65A1^u1A2^u2A3^u3is a vector function of orthogonal curvilinear coordi-
nates u1,u2, and u3, the divergence will take the form
/C114/C65div/C651
/C1041/C1042/C1043/C64
/C64u1
/C1042/C1043A1/C64
/C64u2
/C1043/C1041A2/C64
/C64u3
/C1041/C1042A3
:
1:62
To derive (1.62), we first write /C114/C65as
/C114/C65/C114
A1^u1/C114
A2^u2/C114
A3^u3;
1:63
30VECTOR AND TENSOR ANALYSIS
then, because ^u1/C1041/C1042/C114u2/C114u3, we express /C114
A1^u1)a s
/C114
A1^u1/C114
A1/C1042/C1043/C114u2/C114u3
^u1/C1042/C1043/C114u2/C114u3
/C114
A1/C1042/C1043/C114u2/C114u3A1/C1042/C1043/C114
/C114 u2/C114u3;
where in the last step we have used the vector identity: /C114
/C30/C65
/C114/C30/C65/C30
/C114 /C65.N o w /C114ui^ui=/C104i;i1, 2, 3, so /C114
A1^u1) can be rewritten
as
/C114
A1^u1/C114
A1/C1042/C1043^u2
/C1042^u3
/C10430/C114
A1/C1042/C1043^u1
/C1042/C1043:
The gradient /C114
A1/C1042/C1043is given by Eq. (1.58), and we have
/C114
A1^u1^u1
/C1041/C64
/C64u1
A1/C1042/C1043^u2
/C1042/C64
/C64u2
A1/C1042/C1043^u3
/C1043/C64
/C64u3
A1/C1042/C1043
^u1
/C1042/C1043
1
/C1041/C1042/C1043/C64
/C64u1
A1/C1042/C1043:
Similarly, we have
/C114
A2^u21
/C1041/C1042/C1043/C64
/C64u2
A2/C1043/C1041;and /C114
A3^u31
/C1041/C1042/C1043/C64
/C64u3
A3/C1042/C1041:
Substituting these into Eq. (1.63), we obtain the result, Eq. (1.62).
In the same manner we can derive a formula for curl /C65. We first write it as
/C114/C65/C114
A1^u1A2^u2A3^u3
and then evaluate /C114Ai^ui.
Now ^ui/C104i/C114ui;i1, 2, 3, and we express /C114
A1^u1as
/C114
A1^u1/C114
A1/C1041/C114u1
/C114
A1/C1041/C114 u1A1/C1041/C114/C114 u1
/C114
A1/C1041^u1
/C10410
^u1
/C1041/C64
/C64u1A1/C1041
^u2
/C1042/C64
/C64u2A2/C1042
^u3
/C1043/C64
/C64u3A3/C1043
^u1
/C1041
^u2
/C1043/C1041/C64
/C64u3A1/C1041
ÿ^u3
/C1041/C1042/C64
/C64u2
A1/C1041;
31ORTHOGONAL CURVILINEAR COORDINATES
with similar expressions for /C114
A2^u2and/C114
A3^u3. Adding these together,
we get /C114/C65in orthogonal curvilinear coordinates:
/C114/C65^u1
/C1042/C1043/C64
/C64u2A3/C1043
ÿ/C64
/C64u3A2/C1042
^u2
/C1043/C1041/C64
/C64u3A1/C1041
ÿ/C64
/C64u1A3/C1043
^u3
/C1041/C1042/C64
/C64u1A2/C1042
ÿ/C64
/C64u2
A1/C1041
:
1:64
This can be written in determinant form:
/C114/C651
/C1041/C1042/C1043/C1041^u1/C1042^u2/C1043^u3
/C64
/C64u1/C64
/C64u2/C64
/C64u3
A1/C1041A2/C1042A3/C1043/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
1:65
We now express the Laplacian in orthogonal curvilinear coordinates. From
Eqs. (1.58) and (1.62) we have
/C114/C30grad /C30
1
/C1041/C64/C30
/C64u1^u11
/C1042/C64/C30
/C64u2^u1
/C1043/C64/C30
/C64u3^u3;
/C114/C65div/C651
/C1041/C1042/C1043/C64
/C64u1
/C1042/C1043A1/C64
/C64u2
/C1043/C1041A2/C64
/C64u3
/C1041/C1042A3
:
If/C65/C114/C30, then Ai
1=/C104i/C64/C30=/C64 ui,i1, 2, 3; and
/C114/C65/C114/C114 /C30/C1142/C30
1
/C1041/C1042/C1043/C64
/C64u1/C1042/C1043
/C1041/C64/C30
/C64u1
/C64
/C64u2/C1043/C1041
/C1042/C64/C30
/C64u2
/C64
/C64u3/C1041/C1042
/C1043/C64/C30
/C64u3
:
1:66
/C83pecial orthogonal coordinate s/C121stems
There are at least nine special orthogonal coordinates systems, the most common
and useful ones are the cylindrical and spherical coordinates; we introduce these
two coordinates in this section.
/C67ylindrical coordinates
/C26; /C30;z
u1/C26;u2/C30;u3z;and ^u1e/C26;^u2e/C30^u3ez:
From Fig. 1.17 we see that
x1/C26cos/C30;x2/C26sin/C30;x3z
32VECTOR AND TENSOR ANALYSIS
where
/C260;0/C302;ÿ1 <z<1:
The square of the element of arc length is given by
ds2/C1042
1
d/C262/C10422
d/C302/C10423
dz2:
To find the scale factors /C104i, we notice that ds2drdrwhere
r/C26cos/C30e1/C26sin/C30e2ze3:
Thus
ds2drdr
d/C262/C262
d/C302
dz2:
Equating the two ds2, we find the scale factors:
/C1041/C104/C261;/C1042/C104/C30/C26;/C1043/C104z1:
1:67
From Eqs. (1.58), (1.62), (1.64), and (1.66) we find the gradient, divergence, curl,
and Laplacian in cylindrical coordinates:
/C114/C64
/C64/C26e/C261
/C26/C64
/C64/C30e/C30/C64
/C64zez;
1:68
where
/C26; /C30;zis a scalar function;
/C114/C651
/C26/C64
/C64/C26
/C26A/C26/C64A/C30
/C64/C30/C64
/C64z
/C26Az
;
1:69
33SPECIAL ORTHOGONAL COORDINATE SYSTEMS
Figure 1.17. Cylindrical coordinates.
where
/C65A/C26e/C26A/C30e/C30Azez;
/C114/C651
/C26e/C26/C26e/C30ez
/C64
/C64/C26/C64
/C64/C30/C64
/C64z
A/C26/C26A/C30Az/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
1:70
and
/C114
21
/C26/C64
/C64/C26/C26/C64
/C64/C26
1
/C262/C642
/C64/C302/C642
/C64z2:
1:71
Spherical coordinates
r; ;/C30
u1r;u2;u3/C30;^u1er;^u2e;^u3e/C30
From Fig. 1.18 we see that
x1rsincos/C30;x2rsinsin/C30;x3rcos:
Now
ds2/C1042
1
dr2/C10422
d2/C10423
d/C302
but
rrsincos/C30^e1rsinsin/C30^e2rcos^e3;
34VECTOR AND TENSOR ANALYSIS
Figure 1.18. Spherical coordinates.
so
ds2drdr
dr2r2
d2r2sin2
d/C302:
Equating the two ds2, we find the scale factors: /C1041/C104r1,/C1042/C104r,
/C1043/C104/C30rsin. We then find, from Eqs. (1.58), (1.62), (1.64), and (1.66), the
gradient, divergence, curl, and the Laplacian in spherical coordinates:
/C114^er/C64
/C64r^e1
r/C64
/C64^e/C301
rsin/C64
/C64/C30;
1:72
/C114/C651
r2sinsin/C64
/C64r
r2Arr/C64
/C64
sinAr/C64A/C30
/C64/C30
;
1:73
/C114/C651
r2sin^err^ersin^e/C30
/C64
/C64r/C64
/C64/C64
/C64/C30
ArrArrsinA/C30/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
1:74
/C114
21
r2sinsin/C64
/C64rr2/C64
/C64r
/C64
/C64sin/C64
/C64
1
sin/C642
/C64/C302"#
:
1:75
/C86ector integration and integral theorems
Having discussed vector di/C128erentiation, we now turn to a discussion of vector
integration. After defining the concepts of line, surface, and volume integrals of
vector fields, we then proceed to the important integral theorems of Gauss,
Stokes, and Green.
The integration of a vector, which is a function of a single scalar u, can proceed
as ordinary scalar integration. Given a vector
/C65
uA1
u^e1A2
u^e2A3
u^e3;
then
Z
/C65
udu^e1Z
A1
udu^e2Z
A2
udu^e3Z
A3
udu/C66;
where /C66is a constant of integration, a constant vector. Now consider the integral
of the scalar product of a vector /C65
x1;x2;x3) and drbetween the limit
P1
x1;x2;x3) and P2
x1;x2;x3:
35VECTOR INTEGRATION AND INTEGRAL THEOREMS
ZP2
P1/C65drZP2
P1
A1^e1A2^e2A3^e3
dx1^e1dx2^e2dx3^e3
ZP2
P1A1
x1;x2;x3dx1ZP2
P1A2
x1;x2;x3dx2
ZP2
P1A3
x1;x2;x3dx3:
Each integral on the right hand side requires for its execution more than a knowl-
edge of the limits. In fact, the three integrals on the right hand side are not
completely defined because in the first integral, for example, we do not theknow value of x
2andx3inA1:
I1ZP2
P1A1
x1;x2;x3dx1:
1:76
What is needed is a statement such as
x2f
x1;x3/C103
x1
1:77
that specifies x2,x3for each value of x1. The integrand now reduces to
A1
x1;x2;x3A1
x1;f
x1;/C103
x1 B1
x1so that the integral I1becomes
well defined. But its value depends on the constraints in Eq. (1.77). The con-straints specify paths on the x
1x2andx3x1planes connecting the starting point
P1to the end point P2. The x1integration in (1.76) is carried out along these
paths. It is a path-dependent integral and is called a line integral (or a path
integral). It is very helpful to keep in mind that: when the number of integration
variables is less than the number of variables in the integrand/C44 the integral is not yet
completely de/C174ned and it is path/C45dependent . However, if the scalar product /C65dris
equal to an exact di/C128erential, /C65drd’/C114’dr, the integration depends only
upon the limits and is therefore path-independent:
ZP2
P1/C65drZP2
P1d’’2ÿ’1:
A vector field /C65which has above (path-independent) property is termed conser-
vative. It is clear that the line integral above is zero along any close path, and the
curl of a conservative vector field is zero
/C114 /C65/C114
/C114 ’0. A typical
example of a conservative vector field in mechanics is a conservative force.
The surface integral of a vector function /C65
x1;x2;x3over the surface Sis an
important quantity; it is defined to be
Z
S/C65da;
36VECTOR AND TENSOR ANALYSIS
where the surface integral symbolR
sstands for a double integral over a certain
surface S, and dais an element of area of the surface (Fig. 1.19), a vector quantity.
We attribute to daa magnitude daand also a direction corresponding the normal,
^n, to the surface at the point in question, thus
da^nda:
The normal ^nto a surface may be taken to lie in either of two possible directions.
But if dais part of a closed surface, the sign of ^nrelative to dais so chosen that it
points outward away from the interior. In rectangular coordinates we may write
da^e1da1^e2da2^e3da3^e1dx2dx3^e2dx3dx1^e3dx1dx2:
If a surface integral is to be evaluated over a closed surface S, the integral is
written as
I
S/C65da:
Note that this is di/C128erent from a closed-path line integral. When the path of
integration is closed, the line integral is write it as
I
ÿ^/C65ds;
where ÿspecifies the closed path, and dsis an element of length along the given
path. By convention, dsis taken positive along the direction in which the path is
traversed. Here we are only considering simple closed curves. A simple closed
curve does not intersect itself anywhere.
Gauss’ theorem /C40the divergence theorem/C41
This theorem relates the surface integral of a given vector function and the volumeintegral of the divergence of that vector. It was introduced by Joseph Louis
Lagrange and was first used in the modern sense by George Green. Gauss’
37VECTOR INTEGRATION AND INTEGRAL THEOREMS
Figure 1.19. Surface integral over a surface S.
name is associated with this theorem because of his extensive work on general
problems of double and triple integrals.
If a continuous, di/C128erentiable vector field /C65is defined in a simply connected
region of volume /C86bounded by a closed surface S, then the theorem states that
Z
/C86/C114/C65d/C86I
S/C65da;
1:78
where d/C86dx1dx2dx3. A simple connected region /C86has the property that every
simple closed curve within it can be continuously shrunk to a point withoutleaving the region. To prove this, we first write
Z
/C86/C114/C65d/C86Z
/C86X3
i1/C64Ai
/C64xid/C86;
then integrate the right hand side with respect to x1while keeping x2x3constant,
thus summing up the contribution from a rod of cross section dx2dx3(Fig. 1.20).
The rod intersects the surface Sat the points Pand/C81and thus defines two
elements of area daPandda/C81:
Z
/C86/C64A1
/C64x1d/C86I
Sdx2dx3Z/C81
P/C64A1
/C64x1dx1I
Sdx2dx3Z/C81
PdA1;
where we have used the relation dA1
/C64A1=/C64x1dx1along the rod. The last
integration on the right hand side can be performed at once and we have
Z
/C86/C64A1
/C64x1d/C86I
SA1
/C81ÿA1
Pdx2dx3;
where A1
/C81denotes the value of A1evaluated at the coordinates of the point /C81,
and similarly for A1
P.
The component of the surface element dawhich lies in the x1-direction is
da1dx2dx3at the point /C81, and da1ÿdx2dx3at the point P. The minus sign
38VECTOR AND TENSOR ANALYSIS
Figure 1.20. A square tube of cross section dx2dx3.
arises since the x1component of daatPis in the direction of negative x1.W ec a n
now rewrite the above integral as
Z
/C86/C64A1
/C64x1d/C86Z
S/C81A1
/C81da1Z
SPA1
Pda1;
where S/C81denotes that portion of the surface for which the x1component of the
outward normal to the surface element da1is in the positive x1-direction, and SP
denotes that portion of the surface for which da1is in the negative direction. The
two surface integrals then combine to yield the surface integral over the entire
surface S(if the surface is suciently concave, there may be several such as right
hand and left hand portions of the surfaces):
Z
/C86/C64A1
/C64x1d/C86I
SA1da1:
Similarly we can evaluate the x2andx3components. Summing all these together,
we have Gauss’ theorem:
Z
/C86X
i/C64Ai
/C64xid/C86I
SX
iAidai orZ
/C86/C114/C65d/C86I
S/C65da:
We have proved Gauss’ theorem for a simply connected region (a volume
bounded by a single surface), but we can extend the proof to a multiply connectedregion (a region bounded by several surfaces, such as a hollow ball). For inter-
ested readers, we recommend the book Electromagnetic /C70ields , Roald K.
Wangsness, John Wiley, New York, 1986.
/C67ontinuity equation
Consider a fluid of density /C26
rwhich moves with velocity /C118(r) in a certain region.
If there are no sources or sinks, the following continuity equation must be satis-fied:
/C64/C26
r=/C64t/C114 /C106
r0;
1:79
where /C106is the current
/C106
r/C26
r/C118
r
1:79a
and Eq. (1.79) is called the continuity equation for a conserved current.
To derive this important equation, let us consider an arbitrary surface Senclos-
ing a volume /C86of the fluid. At any time the mass of fluid within /C86isMR
/C86/C26d/C86
and the time rate of mass increase (due to mass flowing into /C86)i s
/C64M
/C64t/C64
/C64tZ
/C86/C26d/C86Z
/C86/C64/C26
/C64td/C86;
39VECTOR INTEGRATION AND INTEGRAL THEOREMS
while the mass of fluid leaving /C86per unit time isZ
S/C26/C118^ndsZ
/C86/C114
/C26/C118d/C86;
where Gauss’ theorem is used in changing the surface integral to volume integral.
Since there is neither a source nor a sink, mass conservation requires an exact
balance between these e/C128ects:Z
/C86/C64/C26
/C64td/C86ÿZ
/C86/C114
/C26/C118d/C86;orZ
/C86/C64/C26
/C64t/C114
/C26/C118
d/C860:
Also since /C86is arbitrary, mass conservation requires that the continuity equation
/C64/C26
/C64t/C114
/C26/C118/C64/C26
/C64t/C114/C1060
must be satisfied everywhere in the region.
Sto/C107es’ theorem
This theorem relates the line integral of a vector function and the surface integralof the curl of that vector. It was first discovered by Lord Kelvin in 1850 and
rediscovered by George Gabriel Stokes four years later.
If a continuous, di/C128erentiable vector field /C65is defined a three-dimensional
region /C86, and Sis a regular open surface embedded in /C86bounded by a simple
closed curve ÿ, the theorem states thatZ
S/C114/C65daI
ÿ/C65dl;
1:80
where the line integral is to be taken completely around the curve ÿanddlis an
element of line (Fig. 1.21).
40VECTOR AND TENSOR ANALYSIS
Figure 1.21. Relation between daanddlin defining curl.
The surface S, bounded by a simple closed curve, is an open surface; and the
normal to an open surface can point in two opposite directions. We adopt the
usual convention, namely the right hand rule: when the fingers of the right hand
follow the direction of dl, the thumb points in the dadirection, as shown in Fig.
1.21.
Note that Eq. (1.80) does not specify the shape of the surface Sother than that
it be bounded by ÿ; thus there are many possibilities in choosing the surface. But
Stokes’ theorem enables us to reduce the evaluation of surface integrals which
depend upon the shape of the surface to the calculation of a line integral which
depends only on the values of /C65along the common perimeter.
To prove the theorem, we first expand the left hand side of Eq. (1.80); with the
aid of Eq. (1.50), it becomes
Z
S/C114/C65daZ
S/C64A1
/C64x3da2ÿ/C64A1
/C64x2da3
Z
S/C64A2
/C64x1da3ÿ/C64A2
/C64x3da1
Z
S/C64A3
/C64x2da1ÿ/C64A3
/C64x1da2
;
1:81
where we have grouped the terms by components of /C65. We next subdivide the
surface Sinto a large number of small strips, and integrate the first integral on the
right hand side of Eq. (1.81), denoted by I1, over one such a strip of width dx1,
which is parallel to the x2x3plane and a distance x1from it, as shown in Fig. 1.21.
Then, by integrating over x1, we sum up the contributions from all of the strips.
Fig. 1.21 also shows the projections of the strip on the x1x3andx1x2planes that
will help us to visualize the orientation of the surface. The element area dais
shown at an intermediate stage of the integration, when the direction angles havevalues such that and /C13are less than 90 8and /C12is greater than 90 8. Thus,
da
2ÿdx1dx3andda3dx1dx2and we can write
I1ÿZ
stripsdx1Z/C81
P/C64A1
/C64x2dx2/C64A1
/C64x3dx3
:
1:82
Note that dx2anddx3in the parentheses are not independent because x2andx3
are related by the equation for the surface Sand the value of x1involved. Since
the second integral in Eq. (1.82) is being evaluated on the strip from Pto/C81for
which x1const., dx10 and we can add
/C64A1=/C64x1dx10 to the integrand to
make it dA1:
/C64A1
/C64x1dx1/C64A1
/C64x2dx2/C64A1
/C64x3dx3dA1:
And Eq. (1.82) becomes
I1ÿZ
stripsdx1Z/C81
PdA1Z
stripsA1
PÿA1
/C81 dx1:
41VECTOR INTEGRATION AND INTEGRAL THEOREMS
Next we consider the line integral of /C65around the lines bounding each of the
small strips. If we trace each one of these lines in the same sense as we trace the
path ÿ, then we will trace all of the interior lines twice (once in each direction) and
all of the contributions to the line integral from the interior lines will cancel,leaving only the result from the boundary line ÿ. Thus, the sum of all of the
line integrals around the small strips will equal the line integral ÿofA
1:
Z
S/C64A1
/C64x3da2ÿ/C64A1
/C64x2da3
I
ÿA1d/C1081:
1:83
Similarly, the last two integrals of Eq. (1.81) can be shown to have the respectivevalues
I
ÿA2d/C1082andI
ÿA3d/C1083:
Substituting these results and Eq. (1.83) into Eq. (1.81) we obtain Stokes’
theorem:
Z
S/C114/C65daI
ÿ
A1d/C1081A2d/C1082A3d/C1083I
ÿ/C65dl:
Stokes’ theorem in Eq. (1.80) is valid whether or not the closed curve ÿlies in a
plane, because in general the surface Sis not a planar surface. Stokes’ theorem
holds for any surface bounded by ÿ.
In fluid dynamics, the curl of the velocity field /C118
ris called its vorticity (for
example, the whirls that one creates in a cup of co/C128ee on stirring it). If the velocityfield is derivable from a potential
/C118
rÿ /C114 /C30
r
it must be irrotational (see Eq. (1.51)). For this reason, an irrotational flow is alsocalled a potential flow, which describes a steady flow of the fluid, free of vortices
and eddies.
One of Maxwell’s equations of electromagnetism (Ampe /C193re’s law) states that
/C114/C66/C22
0/C106;
where /C66is the magnetic induction, /C106is the current density (per unit area), and /C220is
the permeability of free space. From this equation, current densities may bevisualized as vortices of /C66. Applying Stokes’ theorem, we can rewrite Ampe /C193re’s
law as
I
ÿ/C66dr/C220Z
S/C106da/C220I;
it states that the circulation of the magnetic induction is proportional to the total
current Ipassing through the surface Senclosed by ÿ.
42VECTOR AND TENSOR ANALYSIS
Green’s theorem
Green’s theorem is an important corollary of the divergence theorem, and it
has many applications in many branches of physics. Recall that the divergence
theorem Eq. (1.78) states that
Z
/C86/C114/C65d/C86I
S/C65da:
Let/C65/C32/C66, where /C32is a scalar function and /C66a vector function, then /C114/C65
becomes
/C114/C65/C114
/C32/C66/C32/C114/C66/C66/C114/C32:
Substituting these into the divergence theorem, we have
I
S/C32/C66daZ
/C86
/C32/C114/C66/C66/C114/C32d/C86:
1:84
If/C66represents an irrotational vector field, we can express it as a gradient of a
scalar function, say, ’:
/C66/C114’:
Then Eq. (1.84) becomes
I
S/C32/C66daZ
/C86/C32/C114
/C114 ’
/C114 ’
/C114 /C32d/C86:
1:85
Now
/C66da
/C114 ’^nda:
The quantity
/C114’^nrepresents the rate of change of /C30in the direction of the
outward normal; it is called the normal derivative and is written as
/C114’^n/C64’=/C64 n:
Substituting this and the identity /C114
/C114 ’/C1142’into Eq. (1.85), we have
I
S/C32/C64’
/C64ndaZ
/C86/C32/C1142’/C114’/C114/C32d/C86:
1:86
Eq. (1.86) is known as Green’s theorem in the first form.
Now let us interchange ’and/C32, then Eq. (1.86) becomes
I
S’/C64/C32
/C64ndaZ
/C86’/C1142/C32/C114’/C114/C32d/C86:
Subtracting this from Eq. (1.85):
I
S/C32/C64’
/C64nÿ’/C64/C32
/C64n
daZ
/C86/C32/C1142’ÿ’/C1142/C32ÿ
d/C86:
1:87
43VECTOR INTEGRATION AND INTEGRAL THEOREMS
This important result is known as the second form of Green’s theorem, and has
many applications.
Green’s theorem in the plane
Consider the two-dimensional vector field /C65M
x1;x2^e1/C78
x1;x2^e2. From
Stokes’ theorem
I
ÿ/C65drZ
S/C114/C65daZ
S/C64/C78
/C64x1ÿ/C64M
/C64x2
dx1dx2;
1:88
which is often called Green’s theorem in the plane.
SinceH
ÿ/C65drH
ÿ
Mdx 1/C78dx 2, Green’s theorem in the plane can be writ-
ten as
I
ÿMdx 1/C78dx 2Z
S/C64/C78
/C64x1ÿ/C64M
/C64x2
dx1dx2:
1:88a
As an illustrative example, let us apply Green’s theorem in the plane to show
that the area bounded by a simple closed curve ÿis given by
1
2I
ÿx1dx2ÿx2dx1:
Into Green’s theorem in the plane, let us put Mÿx2;/C78x1, giving
I
ÿx1dx2ÿx2dx1Z
S/C64
/C64x1x1ÿ/C64
/C64x2
ÿx2
dx1dx22Z
Sdx1dx22A;
where Ais the required area. Thus A1
2H
ÿx1dx2ÿx2dx1.
/C72elmholt/C122/C39s theorem
The divergence and curl of a vector field play very important roles in physics. We
learned in previous sections that a divergence-free field is solenoidal and a curl-
free field is irrotational. We may classify vector fields in accordance with their
being solenoidal and/or irrotational. A vector field /C86is:
(1) Solenoidal and irrotational if /C114/C860a n d /C114/C860. A static electric
field in a charge-free region is a good example.
(2) Solenoidal if /C114/C860 but /C114/C8660. A steady magnetic field in a current-
carrying conductor meets these conditions.
(3) Irrotational if /C114/C860 but /C114/C860. A static electric field in a charged
region is an irrotational field.
The most general vector field, such as an electric field in a charged medium with
a time-varying magnetic field, is neither solenoidal nor irrotational, but can be
44VECTOR AND TENSOR ANALYSIS
considered as the sum of a solenoidal field and an irrotational field. This is made
clear by Helmholtz’s theorem, which can be stated as (C. W. Wong: Introduction
to Mathematical Physics , Oxford University Press, Oxford 1991; p. 53):
A vector field is uniquely determined by its divergence and curl in
a region of space, and its normal component over the boundary
of the region. In particular, if both divergence and curl arespecified everywhere and if they both disappear at infinity
suciently rapidly, then the vector field can be written as a
unique sum of an irrotational part and a solenoidal part.
In other words, we may write
/C86
rÿ /C114 /C30
r/C114 /C65
r;
1:89
where ÿ/C114/C30is the irrotational part and /C114/C65is the solenoidal part, and /C30(r) and
/C65
rare called the scalar and the vector potential, respectively, of /C86
r). If both /C65
and/C30can be determined, the theorem is verified. How, then, can we determine /C65
and/C30/C63 If the vector field /C86
ris such that
/C114/C86
r/C26;and /C114/C86
r/C118;
then we have
/C114/C86
r/C26ÿ /C114
/C114 /C30/C114
/C114 /C65
or
/C114
2/C30ÿ/C26;
which is known as Poisson’s equation. Next, we have
/C114/C86
r/C118 /C114 ÿ/C114 /C30/C114 /C65
r
or
/C1142/C65/C118;
or in component, we have
/C1142Ai/C118i;i1;2;3
where these are also Poisson’s equations. Thus, both /C65and/C30can be determined
by solving Poisson’s equations.
/C83ome useful integral relations
These relations are closely related to the general integral theorems that we have
proved in preceding sections.
(1) The line integral along a curve /C67between two points aandbis given by
Zb
a/C114/C30
dl/C30
bÿ/C30
a:
1:90
45SOME USEFUL INTEGRAL RELATIONS
Proof:
Zb
a/C114/C30
dlZb
a/C64/C30
/C64x^i/C64/C30
/C64y^j/C64/C30
/C64z^k
dx^idy^jdz^k
Zb
a/C64/C30
/C64xdx/C64/C30
/C64ydy/C64/C30
/C64zdz
Zb
a/C64/C30
/C64xdx
dt/C64/C30
/C64ydy
dt/C64/C30
/C64zdz
dt
dt
Zb
ad/C30
dt
dt/C30
bÿ/C30
a:
2I
S/C64’
/C64ndaZ
/C86/C1142’d/C86:
1:91
Proof: Set /C321 in Eq. (1.87), then /C64/C32=/C64 n0/C1142/C32and Eq. (1.87) reduces to
Eq. (1.91).
3Z
/C86/C114’d/C86I
S’^nda:
1:92
Proof: In Gauss’ theorem (1.78), let /C65’/C67, where /C67is constant vector. Then
we have
Z
/C86/C114
’/C67d/C86Z
S’/C67^nda:
Since
/C114
’/C67/C114 ’/C67/C67/C114’and ’/C67^n/C67
’^n;
we have
Z
/C86/C67/C114’d/C86Z
S/C67
’^nda:
Taking /C67outside the integrals,
/C67Z
/C86/C114’d/C86/C67Z
S
’^nda
and since /C67is an arbitrary constant vector, we have
Z
/C86/C114’d/C86I
S’^nda:
4Z
/C86/C114/C66d/C86Z
S^n/C66da
1:93
46VECTOR AND TENSOR ANALYSIS
Proof: In Gauss’ theorem (1.78), let /C65/C66/C67where /C67is a constant vector. We
then have
Z
/C86/C114
/C66/C67d/C86Z
S
/C66/C67^nda:
Since /C114
/C66/C67/C67
/C114 /C66and
/C66/C67^n/C66
/C67^n
/C67^n/C66
/C67
^n/C66;
Z
/C86/C67
/C114 /C66d/C86Z
S/C67
^n/C66da:
Taking /C67outside the integrals
/C67Z
/C86
/C114 /C66d/C86/C67Z
S
^n/C66da
and since /C67is an arbitrary constant vector, we have
Z
/C86/C114/C66d/C86Z
S^n/C66da:
/C84ensor anal/C121sis
Tensors are a natural generalization of vectors. The beginnings of tensor analysis
can be traced back more than a century to Gauss’ works on curved surfaces.
Today tensor analysis finds applications in theoretical physics (for example, gen-
eral theory of relativity, mechanics, and electromagnetic theory) and to certainareas of engineering (for example, aerodynamics and fluid mechanics). The gen-
eral theory of relativity uses tensor calculus of curved space-time, and engineers
mainly use tensor calculus of Euclidean space. Only general tensors are
considered in this section. The general definition of a tensor is given, followed
by a concise discussion of tensor algebra and tensor calculus (covariant di/C128eren-
tiation).
Tensors are defined by means of their properties of transformation under
coordinate transformation. Let us consider the transformation from one coordi-
nate system
x
1;x2;...;x/C78to another
x01;x02;...;x0/C78in an /C78-dimensional
space /C86/C78. Note that in writing x/C22, the index /C22is a superscript and should not
be mistaken for an exponent. In three-dimensional space we use subscripts. Wenow use superscripts in order that we may maintain a ‘balancing’ of the indices in
all the general equations. The meaning of ‘balancing’ will become clear a little
later. When we transform the coordinates, their di/C128erentials transform according
to the relation
dx
/C22/C64x/C22
/C64x0/C23dx0/C23:
1:94
47TENSOR ANALYSIS
Here we have used Einstein’s summation convention: repeated indexes which
appear once in the lower and once in the upper position are automaticallysummed over. Thus,
X
/C78
/C221A/C22A/C22A/C22A/C22:
It is important to remember that indexes repeated in the lower part or upper part
alone are not summed over. An index which is repeated and over which summa-
tion is implied is called a dummy index. Clearly, a dummy index can be replaced
by any other index that does not appear in the same term.
/C67ontra/C118ariant and co/C118ariant /C118ectors
A set of /C78quantities A/C22
/C221;2;...;/C78which, under a coordinate change,
transform like the coordinate di/C128erentials, are called the components of a contra-
variant vector or a contravariant tensor of the first rank or first order:
A/C22/C64x/C22
/C64x0/C23A0/C23:
1:95
This relation can easily be inverted to express A0/C23in terms of A/C22. We shall leave
this as homework for the reader (Problem 1.32).
If/C78quantities A/C22
/C221;2;...;/C78in a coordinate system
x1;x2;...;x/C78are
related to /C78other quantities A0
/C23
/C231;2;...;/C78in another coordinate system
x01;x02;...;x0/C78by the transformation equations
A/C22/C64x0/C23
/C64x/C22A/C23
1:96
they are called components of a covariant vector or covariant tensor of the firstrank or first order.
One can show easily that velocity and acceleration are contravariant vectors
and that the gradient of a scalar field is a covariant vector (Problem 1.33).
Instead of speaking of a tensor whose components are A
/C22orA/C22we shall simply
refer to the tensor A/C22orA/C22.
/C84ensors of second ran/C107
From two contravariant vectors A/C22andB/C23we may form the /C782quantities A/C22B/C23.
This is known as the outer product of tensors. These /C782quantities form the
components of a contravariant tensor of the second rank: any aggregate of /C782
quantities T/C22/C23which, under a coordinate change, transform like the product of
48VECTOR AND TENSOR ANALYSIS
two contravariant vectors
T/C22/C23/C64x/C22
/C64x0/C64x/C23
/C64x0/C12T0
/C12;
1:97
is a contravariant tensor of rank two. We may also form a covariant tensor of
rank two from two covariant vectors, which transforms according to the formula
T/C22/C23/C64x0
/C64x/C22/C64x0/C12
/C64x/C23T0
/C12:
1:98
Similarly, we can form a mixed tensor T/C22
/C23of order two that transforms as
follows:
T/C22
/C23/C64x/C22
/C64x0/C64x0/C12
/C64x/C23T0
/C12:
1:99
We may continue this process and multiply more than two vectors together,
taking care that their indexes are all di/C128erent. In this way we can construct tensors
of higher rank. The total number of free indexes of a tensor is its rank (or order).
In a Cartesian coordinate system, the distinction between the contravariant and
the covariant tensors vanishes. This can be illustrated with the velocity andgradient vectors. Velocity and acceleration are contravariant vectors, they are
represented in terms of components in the directions of coordinate increase; the
gradient vector is a covariant vector and it is represented in terms of components
in the directions orthogonal to the constant coordinate surfaces. In a Cartesian
coordinate system, the coordinate direction x
/C22coincides with the direction ortho-
gonal to the constant- x/C22surface, hence the distinction between the covariant and
the contravariant vectors vanishes. In fact, this is the essential di/C128erence between
contravariant and covariant tensors: a covariant tensor is represented by com-
ponents in directions orthogonal to like constant coordinate surface, and acontravariant tensor is represented by components in the directions of coordinate
increase.
If two tensors have the same contravariant rank and the same covariant rank,
we say that they are of the same type.
/C66asic operations /C119ith tensors
(1) Equality: Two tensors are said to be equal if and only if they have the same
covariant rank and the same contravariant rank, and every component of
one is equal to the corresponding component of the other:
A
/C12
/C22B/C12
/C22:
49BASIC OPERATIONS WITH TENSORS
(2) Addition (subtraction): The sum (di/C128erence) of two or more tensors of the
same type and rank is also a tensor of the same type and rank. Addition of
tensors is commutative and associative.
(3) Outer product of tensors: The product of two tensors is a tensor whose rank
is the sum of the ranks of the given two tensors. This product involves
ordinary multiplication of the components of the tensor and it is called
the outer product. For example, A/C22/C23B/C12
C/C22/C23/C12is the outer product
ofA/C22/C23andB/C12
.
(4) Contraction: If a covariant and a contravariant index of a mixed tensor are
set equal, a summation over the equal indices is to be taken according to thesummation convention. The resulting tensor is a tensor of rank two less thanthat of the original tensor. This process is called contraction. For example, if
we start with a fourth-order tensor T
/C22
/C23/C26, one way of contracting it is to set
/C26, which gives the second rank tensor T/C22
/C23/C26/C26. We could contract it again
to get the scalar T/C22
/C22/C26/C26.
(5) Inner product of tensors: The inner product of two tensors is produced by
contracting the outer product of the tensors. For example, given two tensors
A/C12
andB/C22
/C23, the outer product is A/C12
B/C22
/C23. Setting /C22, we obtain the
inner product A/C12
/C22B/C22
/C23.
(6) Symmetric and antisymmetric tensors: A tensor is called symmetric with
respect to two contravariant or two covariant indices if its componentsremain unchanged upon interchange of the indices:
A
/C12A/C12;A/C12A/C12:
A tensor is called anti-symmetric with respect to two contravariant or two
covariant indices if its components change sign upon interchange of the
indices:
A/C12ÿA/C12;A/C12ÿA/C12:
Symmetry and anti-symmetry can be defined only for similar indices, not
when one index is up and the other is down.
/C81uotient la/C119
A quantity /C81...
/C22...with various up and down indexes may or may not be a tensor.
We can test whether it is a tensor or not by using the quotient law, which can be
stated as follows:
Suppose it is not known whether a quantity Xis a tensor or not.
If an inner product of Xwith an arbitrary tensor is a tensor, then
Xis also a tensor.
50VECTOR AND TENSOR ANALYSIS
As an example, let XP/C22/C23;Abe an arbitrary contravariant vector, and AP/C22/C23
be a tensor, say /C81/C22/C23:AP/C22/C23/C81/C22/C23, then
AP/C22/C23/C64x0
/C64x/C22/C64x0/C12
/C64x/C23A0/C13P0
/C13/C12:
But
A0/C13/C64x0/C13
/C64xA
and so
AP/C22/C23/C64x0
/C64x/C22/C64x0/C12
/C64x/C23/C64x0/C13
/C64xA0P0
/C13/C12:
This equation must hold for all values of A, hence we have, after canceling the
arbitrary A,
P/C22/C23/C64x0
/C64x/C22/C64x0/C12
/C64x/C23/C64x0/C13
/C64xP0
/C13/C12;
which shows that P/C22/C23is a tensor (contravariant tensor of rank 3).
/C84he line element and metric tensor
So far covariant and contravariant tensors have nothing to do each other except
that their product is an invariant:
A0
/C22B0/C22/C64x
/C64x0/C22/C64x0/C22
/C64x/C12AA/C12/C64x
/C64x/C12AA/C12
/C12AA/C12AA:
A space in which covariant and contravariant tensors exist separately is called
ane. Physical quantities are independent of the particular choice of the modeof description (that is, independent of the possible choice of contravariance orcovariance). Such a space is called a metric space. In a metric space, contravariant
and covariant tensors can be converted into each other with the help of the metric
tensor /C103
/C22/C23. That is, in metric spaces there exists the concept of a tensor that may be
described by covariant indices, or by contravariant indices. These two descrip-tions are now equivalent.
To introduce the metric tensor /C103
/C22/C23, let us consider the line element in /C86/C78.I n
rectangular coordinates the line element (the di/C128erential of arc length) dsis given by
ds2dx2dy2dz2
dx12
dx22
dx32;
there are no cross terms dxidxj. In curvilinear coordinates ds2cannot be repre-
sented as a sum of squares of the coordinate di/C128erentials. As an example, in
spherical coordinates we have
ds2dr2r2d2r2sin2d/C302
which can be in a quadratic form, with x1r;x2;x3/C30.
51THE LINE ELEMENT AND METRIC TENSOR
A generalization to /C86/C78is immediate. We define the line element dsin/C86/C78to be
given by the following quadratic form, called the metric form, or metric
ds2X3
/C221X3
/C231/C103/C22/C23dx/C22dx/C23/C103/C22/C23dx/C22dx/C23:
1:100
For the special cases of rectangular coordinates and spherical coordinates, we
have
~/C103
/C103/C22/C23100
010
0010
B@1
CA; ~/C103
/C103/C22/C2310 0
0r20
00 r2sin20
B@1
CA:
1:101
In an /C78-dimensional orthogonal coordinate system /C103/C22/C230 for /C226/C23. And in a
Cartesian coordinate system /C103/C22/C221a n d /C103/C22/C230 for /C226/C23. In the general case of
Riemannian space, the /C103/C22/C23are functions of the coordinates x/C22
/C221;2;...;/C78.
Since the inner product of /C103/C22/C23and the contravariant tensor dx/C22dx/C23is a scalar
(ds2, the square of line element), then according to the quotient law /C103/C22/C23is a
covariant tensor. This can be demonstrated directly:
ds2/C103/C12dxdx/C12/C1030
/C12dx0dx0/C12:
Now dx0
/C64x0=/C64x/C22dx/C22;so that
/C1030
/C12/C64x0
/C64x/C22/C64x0/C12
/C64x/C23dx/C22dx/C23/C103/C22/C23dx/C22dx/C23
or
/C1030
/C12/C64x0
/C64x/C22/C64x0/C12
/C64x/C23ÿ/C103/C22/C23/C32!
dx/C22dx/C230:
The above equation is identically zero for arbitrary dx/C22, so we have
/C103/C22/C23/C64x0
/C64x/C22/C64x0/C12
/C64x/C23/C1030
/C12;
1:102
which shows that /C103/C22/C23is a covariant tensor of rank two. It is called the metric
tensor or the fundamental tensor.
Now contravariant and covariant tensors can be converted into each other with
the help of the metric tensor. For example, we can get the covariant vector (tensor
of rank one) A/C22from the contravariant vector A/C23:
A/C22/C103/C22/C23A/C23:
1:103
Since we expect that the determinant of /C103/C22/C23does not vanish, the above equations
can be solved for A/C23in terms of the A/C22. Let the result be
A/C23/C103/C23/C22A/C22:
1:104
52VECTOR AND TENSOR ANALYSIS
By combining Eqs. (1.103) and (1.104) we get
A/C22/C103/C22/C23/C103/C23A:
Since the equation must hold for any arbitrary A/C22,w eh a v e
/C103/C22/C23/C103/C23/C22;
1:105
where /C22is Kronecker’s delta symbol. Thus, /C103/C22/C23is the inverse of /C103/C22/C23and vice
versa; /C103/C22/C23is often called the conjugate or reciprocal tensor of /C103/C22/C23. But remember
that /C103/C22/C23and/C103/C22/C23are the contravariant and covariant components of the same
tensor, that is the metric tensor. Notice that the matrix ( /C103/C22/C23) is just the inverse of
the matrix ( /C103/C22/C23).
We can use /C103/C22/C23to lower any upper index occurring in a tensor, and use /C103/C22/C23to
raise any lower index. It is necessary to remember the position from which the
index was lowered or raised, because when we bring the index back to its original
site, we do not want to interchange the order of indexes, in general T/C22/C236T/C23/C22.
Thus, for example
A/C112
/C113/C103r/C112Ar/C113;A/C112/C113/C103r/C112/C103s/C113Ars;A/C112
rs/C103r/C113A/C112/C113
s:
/C65ssociated tensors
All tensors obtained from a given tensor by forming an inner product with themetric tensor are called associated tensors of the given tensor. For example, A
andAare associated tensors:
A/C103/C12A/C12;A/C103/C12A/C12:
/C71eodesics in a /C82iemannian space
In a Euclidean space, the shortest path between two points is a straight linejoining the two points. In a Riemannian space, the shortest path between two
points, called the geodesic, may be a curved path. To find the geodesic, let us
consider a space curve in a Riemannian space given by x
/C22f/C22
tand compute
the distance between two points of the curve, which is given by the formula
sZ/C81
P
/C103/C22dxdx/C22q
Zt2
t1/C103
/C22d_xd_x/C22q
dt;
1:106
where d_xdx=dt, and t(a parameter) varies from point to point of the geo-
desic curve described by the relations which we are seeking. A geodesic joining
53ASSOCIATED TENSORS
two points Pand/C81has a stationary value compared with any other neighboring
path that connects Pand/C81. Thus, to find the geodesic we extremalize (1.106), and
this leads to the di/C128erential equation of the geodesic (Problem 1.37)
d
dt/C64/C70
/C64_x
ÿ/C64/C70
/C64x0;
1:107
where /C70
/C103/C12_x_x/C12q
;and _xdx=dt. Now
/C64/C70
/C64x/C131
2/C103/C12_x_x/C12ÿ1=2/C64/C103/C12
/C64x/C13_x_x/C12;/C64/C70
/C64_x/C1312/C103
/C12_x_x/C12ÿ1=2
2/C103/C13_x
and
ds=dt
/C103/C12_x_x/C12q
:
Substituting these into (1.107) we obtain
d
dt/C103/C13_x_sÿ1ÿ
ÿ1
2/C64/C103/C12
/C64x/C13_x_x/C12_sÿ10;_sds
dt
or
/C103/C13/C127x/C64/C103/C13
/C64x/C12_x_x/C12ÿ12/C64/C103/C12
/C64x/C13_x_x/C12/C103/C13_x/C127s_sÿ1:
We can simplify this equation by writing
/C64/C103/C13
/C64x/C12_x_x/C121
2/C64/C103/C13
/C64x/C12/C64/C103/C12/C13
/C64x
_x_x/C12;
then we have
/C103/C13/C127x/C12; /C13 _x_x/C12/C103/C13_x/C127s_sÿ1:
We can further simplify this equation by taking arc length as the parameter t, then
_s1;/C127s0 and we have
/C103/C13d2x
ds2/C12; /C13dx
dsdx/C12
ds0:
1:108
where the functions
/C12; /C13ÿ/C12;/C131
2/C64/C103/C13
/C64x/C12/C64/C103/C12/C13
/C64xÿ/C64/C103/C12
/C64x/C13
1:109
are called the Christo/C128el symbols of the first kind.
Multiplying (1.108) by /C103/C26/C13, we obtain
d2x/C26
ds2/C26
/C12()
dx
dsdx/C12
ds0;
1:110
54VECTOR AND TENSOR ANALYSIS
where the functions
/C26
/C12()
ÿ/C26
/C12/C103/C26/C13/C12; /C13
1:111
are the Christo/C128el symbol of the second kind.
Eq. (1.110) is, of course, a set of /C78coupled di/C128erential equations; they are the
equations of the geodesic. In Euclidean spaces, geodesics are straight lines. In a
Euclidean space, /C103/C12are independent of the coordinates x/C22, so that the Christo/C128el
symbols identically vanish, and Eq. (1.110) reduces to
d2x/C26
ds20
with the solution
x/C26a/C26sb/C26;
where a/C26andb/C26are constants independent of s. This solution is clearly a straight
line.
The Christo/C128el symbols are not tensors. Using the defining Eqs. (1.109) and the
transformation of the metric tensor, we can find the transformation laws of the
Christo/C128el symbol. We now give the result, without the mathematical details:
ÿ/C22/C23;ÿ/C12;/C13/C64x
/C64x/C22/C64x/C12
/C64x/C23/C64x/C13
/C64x/C103/C12/C64x
/C64x/C642x/C12
/C64x/C22/C64x/C23:
1:112
The Christo/C128el symbols are not tensors because of the presence of the second term
on the right hand side.
/C67o/C118ariant di/C128erentiation
We have seen that a covariant vector is transformed according to the formula
A/C22/C64x/C23
/C64x/C22A/C23;
where the coecients are functions of the coordinates, and so vectors at di/C128erent
points transform di/C128erently. Because of this fact, dA/C22is not a vector, since it is the
di/C128erence of vectors located at two (infinitesimally separated) points. We can
verify this directly:
/C64A/C22
/C64x/C13/C64A/C23
/C64x/C12/C64x/C23
/C64x/C22/C64x/C12
/C64x/C13A/C23/C642x/C23
/C64x/C22/C64x/C13;
1:113
55COVARIANT DIFFERENTIATION
which shows that /C64A=/C64x/C12are not the components of a tensor because of the
second term on the right hand side. The same also applies to the di/C128erential of
a contravariant vector. But we can construct a tensor by the following device.
From Eq. (1.111) we have
ÿ
/C22/C13ÿ/C26
/C28/C64x
/C64x/C22/C64x/C28
/C64x/C13/C64x
/C64x/C26/C642x
/C64x/C22/C64x/C13/C64x
/C64x:
1:114
Multiplying (1.114) by Aand subtracting from (1.113), we obtain
/C64A/C22
/C64x/C13ÿAÿ
/C22/C13/C64A
/C64x/C12ÿA/C26ÿ/C26
/C12/C64x
/C64x/C22/C64x/C12
/C64x/C13:
1:115
If we define
A;/C12/C64A
/C64x/C12ÿA/C26ÿ/C26
/C12;
1:116
then (1.115) can be rewritten as
A/C22;/C13A;/C12/C64x
/C64x/C22/C64x/C12
/C64x/C13;
which shows that A;/C12is a covariant tensor of rank 2. This tensor is called the
covariant derivative of Awith respect to x/C12. The semicolon denotes covariant
di/C128erentiation. In a Cartesian coordinate system, the Christo/C128el symbols vanish,
and so covariant di/C128erentiation reduces to ordinary di/C128erentiation.
The contravariant derivative is found by raising the index which denotes di/C128er-
entiation:
A/C22;/C103A/C22
;:
1:117
We can similarly determine the covariant derivative of a tensor of arbitrary
rank. In doing so we find the following simple rule helps greatly:
To obtain the covariant derivative of the tensor T
with respect to
x/C22/C44 we add to the ordinary derivative /C64T
=/C64x/C22for each covariant
index /C23
T
/C23:a term ÿÿ
/C22/C23T
:/C44 and for each contravariant index
/C23
T/C23
a term ÿ
/C23/C22T
....
Thus,
T/C22/C23;/C64T/C22/C23
/C64xÿÿ/C12
/C22T/C12/C23ÿÿ/C12
/C23T/C22/C12;
T/C22
/C23;/C64T/C22
/C23
/C64xÿÿ/C12
/C23T/C22
/C12ÿ/C22
/C12T/C12
/C23:
The covariant derivatives of both the metric tensor and the Kronnecker delta
are identically zero (Problem 1.38).
56VECTOR AND TENSOR ANALYSIS
Problems
1.1. Given the vector /C65
2;2;ÿ1and/C66
6;ÿ3;2, determine:
(a)6/C65ÿ3/C66,(b)A2B2,(c)/C65/C66,(d) the angle between /C65and/C66,(e) the
direction cosines of /C65,(f) the component of /C66in the direction of /C65.
1.2. Find a unit vector perpendicular to the plane of /C65
2;ÿ6;ÿ3and
/C66
4;3;ÿ1.
1.3. Prove that:
(a) the median to the base of an isosceles triangle is perpendicular to the
base; ( b) an angle inscribed in a semicircle is a right angle.
1.4. Given two vectors /C65
2;1;ÿ1,/C66
1;ÿ1;2find: ( a)/C65/C66, and ( b)a
unit vector perpendicular to the plane containing vectors /C65and/C66.
1.5. Prove: ( a) the law of sines for plane triangles, and ( b) Eq. (1.16a).
1.6. Evaluate
2^e1ÿ3^e2
^e1^e2ÿ^e3
3^e1ÿ^e3.
1.7. ( a) Prove that a necessary and sucient condition for the vectors /C65,/C66and
/C67to be coplanar is that /C65
/C66/C670:
(b) Find an equation for the plane determined by the three points
P1
2;ÿ1;1,P2
3;2;ÿ1andP3
ÿ1;3;2.
1.8. ( a) Find the transformation matrix for a rotation of new coordinate system
through an angle /C30about the x3
z-axis.
(b) Express the vector /C653^e12^e2^e3in terms of the triad ^e0
1^e0
2^e0
3where
thex0
1x0
2axes are rotated 45 8about the x3-axis (the x3-a n d x0
3-axes
coinciding).
1.9. Consider the linear transformation A0
iP3
j1^e0
i^ejAjP3j1ijAj. Show,
using the fact that the magnitude of the vector is the same in both systems,
that
X3
i1ijikjk
j;k1;2;3:
1.10. A curve /C67is defined by the parametric equation
r
ux1
u^e1x2
u^e2x3
u^e3;
where uis the arc length of C measured from a fixed point on /C67,a n d ris the
position vector of any point on /C67; show that:
(a)dr=duis a unit vector tangent to /C67;
(b) the radius of curvature of the curve /C67is given by
/C26d2x1
du2/C32!2
d2x2
du2/C32!2
d2x3
du2/C32!22
435ÿ1=2
:
57PROBLEMS
1.11. ( a) Show that the acceleration aof a particle which travels along a space
curve with velocity /C118is given by
ad/C118
dt^T/C1182
/C26^/C78;
where ^T,^/C78, and /C26are as defined in the text.
(b) Consider a particle Pmoving on a circular path of radius rwith constant
angular speed /C33d=dt(Fig. 1.22). Show that the acceleration aof the
particle is given by
aÿ/C332r:
1.12. A particle moves along the curve x12t2;x2t2ÿ4t;x33tÿ5, where t
is the time. Find the components of the particle’s velocity and acceleration
at time t1 in the direction ^e1ÿ3^e22^e3.
1.13. ( a) Find a unit vector normal to the surface x2
1x22ÿx31 at the point
P(1,1,1).
(b) Find the directional derivative of /C30x21x2x34x1x23at (1, ÿ2;ÿ1) in
the direction 2 ^e1ÿ^e2ÿ2^e3.
1.14. Consider the ellipse given by r1r2const :(Fig. 1.23). Show that r1andr2
make equal angles with the tangent to the ellipse.
1.15. Find the angle between the surfaces x2
1x22x239a n d x3x21x22ÿ3.
at the point (2, ÿ1, 2).
1.16. ( a)I ffandgare di/C128erentiable scalar functions, show that
/C114
f/C103f/C114/C103/C103/C114f:
(b) Find /C114rifr
x21x22x231=2.
(c) Show that /C114rnnrnÿ2r.
1.17. Show that:
(a)/C114
r=r30. Thus the divergence of an inverse-square force is zero.
58VECTOR AND TENSOR ANALYSIS
Figure 1.22. Motion on a circle.
(b)I ffis a di/C128erentiable function and /C65is a di/C128erentiable vector function,
then
/C114
f/C65
/C114 f/C65f
/C114 /C65:
1.18. ( a) What is the divergence of a gradient/C63
(b) Show that /C1142
1=r0.
(c) Show that r
/C114 r6
r/C114r.
1.19 Given /C114/C690;/C114/C720;/C114/C69ÿ/C64H=/C64t;/C114/C72/C64/C69=/C64t, show that
/C69and/C72satisfy the wave equation /C1142u/C642u=/C64t2.
The given equations are related to the source-free Maxwell’s equations of
electromagnetic theory, /C69and/C72are the electric field and magnetic field
intensities.
1.20. ( a) Find constants a,b,csuch that
/C65
x12x2ax3^e1
bx1ÿ3x2ÿx3^e2
4x1cx22x3^e3
is irrotational.
(b) Show that /C65can be expressed as the gradient of a scalar function.
1.21. Show that a cylindrical coordinate system is orthogonal.1.22. Find the volume element d/C86in: (a) cylindrical and (b) spherical coordinates.
Hint: The volume element in orthogonal curvilinear coordinates is
d/C86/C104
1/C1042/C1043du1du2du3/C64
x1;x2;x3
/C64
u1;u2;u3/C12/C12/C12/C12/C12/C12/C12/C12du
1du2du3:
1.23. Evaluate the integralR
1;2
0;1
x2ÿydx
y2xdyalong
(a) a straight line from (0, 1) to (1, 2);
(b) the parabola xt;yt21;
(c) straight lines from (0, 1) to (1, 1) and then from (1, 1) to (1, 2).
1.24. Evaluate the integralR
1;1
0;0
x2y2dxalong (see Fig. 1.24):
(a) the straight line yx,
(b) the circle arc of radius 1 ( xÿ12y21.
59PROBLEMS
Figure 1.23.
1.25. Evaluate the surface integralR
S/C65daR
S/C65^nda, where /C65x1x2^e1ÿ
x2
1^e2
x1x2^e3,Sis that portion of the plane 2 x12x2x36
included in the first octant.
1.26. Verify Gauss’ theorem for /C65
2x1ÿx3^e1x21x2^e2ÿx1x23^e3taken over
the region bounded by x10;x11;x20;x21;x30;x31.
1.28 Show that the electrostatic field intensity /C69
rof a point charge /C81at the
origin has an inverse-square dependence on r.
1.28. Show, by using Stokes’ theorem, that the gradient of a scalar field is irrota-
tional:
/C114
/C114 /C30
r 0:
1.29. Verify Stokes’ theorem for /C65
2x1ÿx2^e1ÿx2x23^e2ÿx22x3^e3, where Sis
the upper half surface of the sphere x2
1x22x231 and ÿis its boundary
(a circle in the x1x2plane of radius 1 with its center at the origin).
1.30. Find the area of the ellipse x1acos;x2bsin.
1.31. Show thatR
Sr^nda0, where Sis a closed surface which encloses a
volume /C86.
1.33. Starting with Eq. (1.95), express A0/C23in terms of A/C22.
1.33. Show that velocity and acceleration are contravariant vectors and that the
gradient of a scalar field is a covariant vector.
1.34. The Cartesian components of the acceleration vector are
axd2x
dt2; ayd2y
dt2; azd2z
dt2:
Find the component of the acceleration vector in the spherical polar co-
ordinates.
1.35. Show that the property of symmetry (or anti-symmetry) with respect to
indexes of a tensor is invariant under coordinate transformation.
1.36. A covariant tensor has components xy;2yÿz2;xzin rectangular coordi-
nates, find its covariant components in spherical coordinates.
60VECTOR AND TENSOR ANALYSIS
Figure 1.24. Paths for a path integral.
1.37. Prove that a necessary condition that IRt/C81
tP/C70
t;x;_xdtbe an extremum
(maximum or minimum) is that
d
dt/C64/C70
/C64_x
ÿ/C64/C70
/C64x0:
1.38. Show that the covariant derivatives of: ( a) the metric tensor, and ( b) the
Kronecker delta are identically zero.
61PROBLEMS
2
Ordinary di/C128erential e/C113uations
Physicists have a variety of reasons for studying di/C128erential equations: almost all
the elementary and numerous of the advanced parts of theoretical physics are
posed mathematically in terms of di/C128erential equations. We devote three chapters
to di/C128erential equations. This chapter will be limited to ordinary di/C128erential
equations that are reducible to a linear form. Partial di/C128erential equations
and special functions of mathematical physics will be dealt with in Chapters 10
and 7.
A di/C128erential equation is an equation that contains derivatives of an
unknown function which expresses the relationship we seek. If there is only one
independent variable and, as a consequence, total derivatives like dx=dt, the
equation is called an ordinary di/C128erential equation (ODE). A partial di/C128erential
equation (PDE) contains several independent variables and hence partial deriva-
tives.
Theorder of a di/C128erential equation is the order of the highest derivative appear-
ing in the equation; its degree is the power of the derivative of highest order after
the equation has been rationalized, that is, after fractional powers of all deriva-tives have been removed. Thus the equation
d2y
dx23dy
dx2y0
is of second order and first degree, and
d3y
dx3
1
dy=dx3q
is of third order and second degree, since it contains the term ( d3y=dx32after it is
rationalized.
62
A di/C128erential equation is said to be linear if each term in it is such that the
dependent variable or its derivatives occur only once, and only to the first power.
Thus
d3y
dx3ydy
dx0
is not linear, but
x3d3y
dx3exsinxdy
dxylnx
is linear. If in a linear di/C128erential equation there are no terms independent of y,
the dependent variable, the equation is also said to be homogeneous ; this would
have been true for the last equation above if the ‘ln x’ term on the right hand side
had been replaced by zero.
A very important property of linear homogeneous equations is that, if we know
two solutions y1andy2, we can construct others as linear combinations of them.
This is known as the principle of superposition and will be proved later when wedeal with such equations.
Sometimes di/C128erential equations look unfamiliar. A trivial change of variables
can reduce a seemingly impossible equation into one whose type is readily recog-nizable.
Many di/C128erential equations are very dicult to solve. There are only a rela-
tively small number of types of di/C128erential equation that can be solved in closed
form. We start with equations of first order. A first-order di/C128erential equation can
always be solved, although the solution may not always be expressible in terms of
familiar functions. A solution (or integral) of a di/C128erential equation is the relation
between the variables, not involving di/C128erential coecients, which satisfies thedi/C128erential equation. The solution of a di/C128erential equation of order nin general
involves narbitrary constants.
First-order di/C128erential equations
A di/C128erential equation of the general form
dy
dxÿf
x;y
/C103
x;y;or/C103
x;ydyf
x;ydx0
2:1
is clearly a first-order di/C128erential equation.
Separable variables
Iff
x;yand/C103
x;yare reducible to P
xand/C81
y, respectively, then we have
/C81
ydyP
xdx0:
2:2
Its solution is found at once by integrating.
63FIRST-ORDER DIFFERENTIAL EQUATIONS
The reader may notice that dy=dxhas been treated as if it were a ratio of dyand
dx, that can be manipulated independently. Mathematicians may be unhappy
about this treatment. But, if necessary, we can justify it by considering dyand
dxto represent small finite changes yandx, before we have actually reached the
limit where each becomes infinitesimal.
Example 2.1
Consider the di/C128erential equation
dy=dxÿy2ex:
We can rewrite it in the following form ÿdy=y2exdxwhich can be integrated
separately giving the solution
1=yexc;
where cis an integration constant.
Sometimes when the variables are not separable a di/C128erential equation may be
reduced to one in which they are separable by a change of variable. The general
form of di/C128erential equation amenable to this approach is
dy=dxf
axby;
2:3
where fis an arbitrary function and aandbare constants. If we let /C119axby,
then bdy=dxd/C119=dxÿa, and the di/C128erential equation becomes
d/C119=dxÿabf
/C119
from which we obtain
d/C119
abf
/C119dx
in which the variables are separated.
Example 2.2
Solve the equation
dy=dx8x4y
2xyÿ12:
Solution: Letw2xy, then dy=dxdw=dxÿ2, and the di/C128erential equation
becomes
d/C119=dx24/C119
/C119ÿ12
or
d/C119=4/C119
/C119ÿ12ÿ2dx:
The variables are separated and the equation can be solved.
64ORDINARY DIFFERENTIAL EQUATIONS
A homogeneous di/C128erential equation which has the general form
dy=dxf
y=x
2:4
may also be reduced, by a change of variable, to one with separable variables.
This can be illustrated by the following example:
Example 2.3
Solve the equation
dy
dxy2xy
x2:
Solution: The right hand side can be rewritten as
y=x2
y=x, and hence is a
function of the single variable
/C118y=x:
We thus use /C118both for simplifying the right hand side of our equation, and also
for rewriting dy=dxin terms of /C118and x. Now
dy
dxd
dx
x/C118/C118xd/C118
dx
and our equation becomes
/C118xd/C118
dx/C1182/C118
from which we have
d/C118
/C1182dx
x:
Integration gives
ÿ1
/C118lnxcorxAeÿx=y;
where cand A
eÿcare constants.
Sometimes a nearly homogeneous di/C128erential equation can be reduced to
homogeneous form which can then be solved by variable separation. This canbe illustrated by the by the following:
Example 2.4
Solve the equation
dy=dx
yxÿ5=
yÿ3xÿ1:
65FIRST-ORDER DIFFERENTIAL EQUATIONS
Solution: Our equation would be homogeneous if it were not for the constants
ÿ5 and ÿ1 in the numerator and denominator respectively. But we can eliminate
them by a change of variable:
x0x;y0y/C12;
where and /C12are constants specially chosen in order to make our equation
homogeneous:
dy0=dx0
y0x0=y0ÿ3x0:
Note that dy0=dx0dy=dx. Trivial algebra yields ÿ1;/C12ÿ4. Now let
/C118y0=x0, then
dy0
dx0d
dx0
x0/C118/C118x0d/C118
dx0
and our equation becomes
/C118x0d/C118
dx0/C1181
/C118ÿ3;or/C118ÿ3
ÿ/C11824/C1181d/C118dx0
x0
in which the variables are separated and the equation can be solved by integra-
tion.
Example 2.5
Fall of a skydiver.
Solution: Assuming the parachute opens at the beginning of the fall, there are
two forces acting on the parachute: the downward force of gravity mg, and the
upward force of air resistance kv2. If we choose a coordinate system that has y0
at the earth’s surface and increases upward, then the equation of motion of the
falling diver, according to Newton’s second law, is
md/C118=dtÿm/C103k/C1182;
where mis the mass, gthe gravitational acceleration, and ka positive constant. In
general the air resistance is very complicated, but the power-law approximation is
useful in many instances in which the velocity does not vary appreciably.
Experiments show that for a subsonic velocity up to 300 m/s, the air resistance
is approximately proportional to /C1182.
The equation of motion is separable:
md/C118
m/C103ÿk/C1182dt
66ORDINARY DIFFERENTIAL EQUATIONS
or, to make the integration easier
d/C118
/C1182ÿ
m/C103=kÿk
mdt:
Now
1
/C1182ÿ
m/C103=k1
/C118/C118t
/C118ÿ/C118t1
2/C118t1
/C118ÿ/C118tÿ1
/C118/C118t
;
where /C1182
tm/C103=k. Thus
1
2/C118td/C118
/C118ÿ/C118tÿd/C118
/C118ÿ/C118t
ÿk
mdt:
Integrating yields
1
2/C118tln/C118ÿ/C118t
/C118/C118t
ÿk
mtc;
where cis an integration constant.
Solving for /C118we finally obtain
/C118
t/C118t1Bexp
ÿ2/C103t=/C118t
1ÿBexp
ÿ2/C103t=/C118t;
where Bexp
2/C118tC.
It is easy to see that as t!1 , exp( ÿ2/C103t=/C118t!0, and so /C118!/C118t; that is, if he
falls from a sucient height, the diver will eventually reach a constant velocity
given by /C118t, the terminal velocity. To determine the constants of integration, we
need to know the value of k, which is about 30 kg/m for the earth’s atmosphere
and a standard parachute.
Exact equations
We may integrate Eq. (2.1) directly if its left hand side is the di/C128erential duof
some function u
x;y, in which case the solution is of the form
u
x;yC
2:5
and Eq. (2.1) is said to be exact. A convenient test to see if Eq. (2.1) is exact isdoes
/C64/C103
x;y
/C64x/C64f
x;y
/C64y:
2:6
To see this, let us go back to Eq. (2.5) and we have
du
x;y 0:
On performing the di/C128erentiation we obtain
/C64u
/C64xdx/C64u
/C64ydy0:
2:7
67FIRST-ORDER DIFFERENTIAL EQUATIONS
It is a general property of partial derivatives of any well-behaved function that
the order of di/C128erentiation is immaterial. Thus we have
/C64
/C64y/C64u
/C64x
/C64
/C64x/C64u
/C64y
:
2:8
Now if our di/C128erential equation (2.1) is of the form of Eq. (2.7), we must be able
to identify
f
x;y/C64u=/C64xand /C103
x;y/C64u=/C64y:
2:9
Then it follows from Eq. (2.8) that
/C64/C103
x;y
/C64x/C64f
x;y
/C64y;
which is Eq. (2.6).
Example 2.6
Show that the equation xdy=dx
xy0 is exact and find its general solu-
tion.
Solution: We first write the equation in standard form
xydxxdy0:
Applying the test of Eq. (2.6) we notice that
/C64f
/C64y/C64
/C64y
xy1 and/C64/C103
/C64x/C64x
/C64x1:
Therefore the equation is exact, and the solution is of the form indicated by Eq.
(2.7). From Eq. (2.9) we have
/C64u=/C64xxy;/C64u=/C64yx;
from which it follows that
u
x;yx2=2xy/C104
y;u
x;yxyk
x;
where /C104
yandk
xarise from integrating u
x;ywith respect to xandy, respec-
tively. For consistency, we require that
/C104
y0a n d k
xx2=2:
Thus the required solution is
x2=2xyc:
It is interesting to consider a di/C128erential equation of the type
/C103
x;ydy
dxf
x;yk
x;
2:10
68ORDINARY DIFFERENTIAL EQUATIONS
where the left hand side is an exact di/C128erential
d=dxu
x;y, and k
xon the
right hand side is a function of xonly. Then the solution of the di/C128erential
equation can be written as
u
x;yZ
k
xdx:
2:11
Alternatively Eq. (2.10) can be rewritten as
/C103
x;ydy
dxf
x;yÿk
x 0:
2:10a
Since the left hand side of Eq. (2.10) is exact, we have
/C64/C103=/C64x/C64f=/C64y:
Then Eq. (2.10a) is exact as well. To see why, let us apply the test for exactness for
Eq. (2.10a) which requires
/C64
/C64x/C103
x;y /C64
/C64yf
x;yÿk
x /C64
/C64yf
x;y:
Thus Eq. (2.10a) satisfies the necessary requirement for being exact. We can thus
write its solution as
U
x;yc;
where
/C64U
/C64y/C103
x;yand/C64U
/C64xf
x;yÿk
x:
Of course, the solution U
x;ycmust agree with Eq. (2.11).
Integrating factors
If a di/C128erential equation in the form of Eq. (2.1) is not already exact, it sometimescan be made so by multiplying by a suitable factor, called an integrating factor.
Although an integrating factor always exists for each equation in the form of Eq.
(2.1), it may be dicult to find it. However, if the equation is linear, that is, if can
be written
dy
dxf
xy/C103
x
2:12
an integrating factor of the form
expZ
f
xdx
2:13
is always available. It is easy to verify this. Suppose that R
xis the integrating
factor we are looking for. Multiplying Eq. (2.12) by /C82, we have
Rdy
dxRf
xyR/C103
x;orRdyRf
xydxR/C103
xdx:
69FIRST-ORDER DIFFERENTIAL EQUATIONS
The right hand side is already integrable; the condition that the left hand side of
Eq. (2.12) be exact gives
/C64
/C64yRf
xy/C64R
/C64x;
which yields
dR=dxRf
x;ordR=Rf
xdx;
and integrating gives
lnRZ
f
xdx
from which we obtain the integrating factor /C82we were looking for
RexpZ
f
xdx
:
It is now possible to write the general solution of Eq. (2.12). On applying the
integrating factor, Eq. (2.12) becomes
d
ye/C70
dx/C103
xe/C70;
where /C70
xR
f
xdx. The solution is clearly given by
yeÿ/C70Z
e/C70/C103
xdxC
:
Example 2.7Show that the equation xdy=dx2yx
20 is not exact; then find a suitable
integrating factor that makes the equation exact. What is the solution of thisequation/C63
Solution: We first write the equation in the standard form
2yx
2dxxdy0;
then we notice that
/C64
/C64y
2yx22a n d/C64
/C64xx1;
which indicates that our equation is not exact. To find the required integrating
factor that makes our equation exact, we rewrite our equation in the form of Eq.
(2.12):
dy
dx2y
xÿx
70ORDINARY DIFFERENTIAL EQUATIONS
from which we find f
x1=x, and so the required integrating factor is
expZ
1=xdx
exp
lnxx:
Applying this to our equation gives
x2dy
dx2xyx30o rd
dxx2yx4=4ÿ
0
which integrates to
x2yx4=4c;
or
ycÿx4
4x2:
Example 2.8
/C82Lcircuits: A typical /C82Lcircuit is shown in Fig. 2.1. Find the current I
tin the
circuit as a function of time t.
Solution: We need first to establish the di/C128erential equation for the current
flowing in the circuit. The resistance /C82and the inductance Lare both constant.
The voltage drop across the resistance is I/C82, and the voltage drop across the
inductance is LdI=dt. Kirchho/C128 ’s second law for circuits then gives
LdI
t
dtRI
t/C69
t;
which is in the form of Eq. (2.12), but with tas the independent variable instead of
xand Ias the dependent variable instead of y. Thus we immediately have the
general solution
I
t1
LeÿRt=LZ
eRt=L/C69
tdtkeÿRt=L;
71FIRST-ORDER DIFFERENTIAL EQUATIONS
Figure 2.1. /C82Lcircuit.
where kis a constant of integration (in electric circuits, /C67is used for capacitance).
Given Ethis equation can be solved for I
t. If the voltage Eis constant, we
obtain
I
t1
LeÿRt=L/C69L
ReÿRt=L
keÿRt=L/C69
RkeÿRt=L:
Regardless of the value of k, we see that
I
t!/C69=Rast!1 :
Setting t0 in the solution, we find
kI
0ÿ/C69=R:
Bernoulli’s equation
Bernoulli’s equation is a non-linear first-order equation that occurs occasionally
in physical problems:
dy
dxf
xy/C103
xyn;
2:14
where nis not necessarily integer.
This equation can be made linear by the substitution /C119yawith suitably
chosen. We find this can be achieved if 1ÿn:
/C119y1ÿnory/C1191=
1ÿn:
This converts Bernoulli’s equation into
d/C119
dx
1ÿnf
x/C119
1ÿn/C103
x;
which can be made exact using the integrating factor exp
R
1ÿnf
xdx.
/C83econd-order equations /C119ith constant coe/C129cients
The general form of the nth-order linear di/C128erential equation with constant coef-
ficients is
dny
dxn/C1121dnÿ1y
dxnÿ1 /C112nÿ1dy
dx/C112ny
Dn/C1121Dnÿ1 /C112nÿ1D/C112nyf
x;
where /C1121;/C1122;...are constants, f
xis some function of x, and Dd=dx.I f
f
x0, the equation is called homogeneous; otherwise it is called a non-homo-
geneous equation. It is important to note that the symbol /C68is meaningless unless
applied to a function of xand is therefore not a mathematical quantity in the
usual sense. /C68is an operator.
72ORDINARY DIFFERENTIAL EQUATIONS
Many of the di/C128erential equations of this type which arise in physical problems
are of second order and we shall consider in detail the solution of the equation
d2y
dt2ady
dtby
D2aDbyf
t;
2:15
where aandbare constants, and tis the independent variable. As an example, the
equation of motion for a mass on a spring is of the form Eq. (2.15), with a
representing the friction, cbeing the constant of proportionality in Hooke’s law
for the spring, and f
tsome time-dependent external force acting on the mass.
Eq. (2.15) can also apply to an electric circuit consisting of an inductor, a resistor,
a capacitor and a varying external voltage.
The solution of Eq. (2.15) involves first finding the solution of the equation with
f
treplaced by zero, that is,
d2y
dt2ady
dtby
D2aDby0;
2:16
this is called the reduced or homogeneous equation corresponding to Eq. (2.15).
Nature of the solution of linear equations
We now establish some results for linear equations in general. For simplicity, weconsider the second-order reduced equation (2.16). If y
1andy2are independent
solutions of (2.16) and Aand Bare any constants, then
D
Ay1By2ADy 1BDy 2;D2
Ay1By2AD2y1BD2y2
and hence
D2aDb
Ay1By2A
D2aDby1B
D2aDby20:
Thus yAy1By2is a solution of Eq. (2.16), and since it contains two arbitrary
constants, it is the general solution. A necessary and sucient condition for two
solutions y1andy2to be linearly independent is that the Wronskian determinant
of these functions does not vanish:
y1y2
dy1
dtdy2
dt/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C1260:
Similarly, if y
1;y2;...;ynarenlinearly independent solutions of the nth-order
linear equations, then the general solution is
yA1y1A2y2 Anyn;
where A1;A2;...;Anare arbitrary constants. This is known as the superposition
principle.
73SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS
General solutions of the second-order equations
Suppose that we can find one solution, y/C112
tsay, of Eq. (2.15):
D2aDby/C112
tf
t:
2:15a
Then on defining
yc
ty
tÿy/C112
t
we find by subtracting Eq. (2.15a) from Eq. (2.15) that
D2aDbyc
t0:
That is, yc
tsatisfies the corresponding homogeneous equation (2.16), and it is
known as the complementary function yc
tof non-homogeneous equation (2.15).
while the solution y/C112
tis called a particular integral of Eq. (2.15). Thus, the
general solution of Eq. (2.15) is given by
y
tAyc
tBy/C112
t:
2:17
Finding the complementary function
Clearly the complementary function is independent of f
t, and hence has nothing
to do with the behavior of the system in response to the external applied influence.
What it does represent is the free motion of the system. Thus, for example, even
without external forces applied, a spring can oscillate, because of any initial
displacement and/or velocity. Similarly, had a capacitor already been charged
att0, the circuit would subsequently display current oscillations even if there
is no applied voltage.
In order to solve Eq. (2.16) for yc
t, we first consider the linear first-order
equation
ady
dtby0:
Separating the variables and integrating, we obtain
yAeÿbt=a;
where Ais an arbitrary constant of integration. This solution suggests that Eq.
(2.16) might be satisfied by an expression of the type
ye/C112t;
where pis a constant. Putting this into Eq. (2.16), we have
e/C112t
/C1122a/C112b0:
Therefore ye/C112tis a solution of Eq. (2.16) if
/C1122a/C112b0:
74ORDINARY DIFFERENTIAL EQUATIONS
This is called the auxiliary (or characteristic) equation of Eq. (2.16). Solving it
gives
/C1121ÿa
a2ÿ4bp
2; /C1122ÿaÿa2ÿ4bp
2:
2:18
We now distinguish between the cases in which the roots are real and distinct,
complex or coincident.
(i) Real and distinct roots ( a2ÿ4b/C620
In this case, we have two independent solutions y1e/C1121t;y2e/C1122tand the general
solution of Eq. (2.16) is a linear combination of these two:
yAe/C1121tBe/C1122t;
2:19
where Aand Bare constants.
Example 2.9
Solve the equation
D2ÿ2Dÿ3y0, given that y1 and y0dy=dx2
when t0.
Solution: The auxiliary equation is p2ÿ2pÿ30, from which we find pÿ1
orp3. Hence the general solution is
yAeÿtBe3t:
The constants Aand Bcan be determined by the boundary conditions at t0.
Since y1 when t0, we have
1AB:
Now
y0ÿAeÿt3Be3t
and since y02 when t0, we have 2 ÿA3B. Hence
A1=4;B3=4
and the solution is
4yeÿt3e3t:
(ii) Complex roots
a2ÿ4b<0
If the roots /C1121,/C1122of the auxiliary equation are imaginary, the solution given by
Eq. (2.18) is still correct. In order to give the solutions in terms of real quantities,we can use the Euler relations to express the exponentials. If we let
rÿa=2;is
a
2ÿ4bp
=2, then
e/C1121terteistertcosstisinst;
e/C1122terteistertcosstÿisinst
75SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS
and the general solution can be written as
yAe/C1121tBe/C1122t
ert
ABcossti
AÿBsinst
ertA0cosstB0sinst
2:20
with A0AB;B0i
AÿB:
The solution (2.20) may be expressed in a slightly di/C128erent and often more
useful form by writing B0=A0tan. Then
y
A2
0B201=2ert
coscosstsinsinstCertcos
stÿ;
2:20a
where /C67andare arbitrary constants.
Example 2.10
Solve the equation
D24D13y0, given that y1 and y02 when t0.
Solution: The auxiliary equation is p24p130, and hence pÿ23i.
The general solution is therefore, from Eq. (2.20),
yeÿ2t
A0cos 3 tB0sin 3t:
Since y/C108when t0, we have A01. Now
y0ÿ2eÿ2t
A0cos 3 tB0sin 3t3eÿ2t
ÿA0sin 3tB0cos 3 t
and since y02 when t0, we have 2 ÿ2A03B0. Hence B04=3, and the
solution is
3yeÿ2t
3 cos 3 t4 sin 3 t:
(iii) Coincident roots
When a24b, the auxiliary equation yields only one value for p, namely
/C112ÿa=2, and hence the solution yAet. This is not the general solution
as it does not contain the necessary two arbitrary constants. In order to obtain thegeneral solution we proceed as follows. Assume that y/C118e
t, where vis a func-
tion of tto be determined. Then
y0/C1180et/C118et;y00/C11800et2/C1180et2/C118et:
Substituting for y;y0, and y00in the di/C128erential equation we have
et/C118002/C11802/C118a
/C1180/C118b/C1180
and hence
/C11800/C1180
a2/C118
2ab0:
76ORDINARY DIFFERENTIAL EQUATIONS
Now
2ab0;and a20
so that
/C118000:
Hence, integrating gives
/C118AtB;
where Aand Bare arbitrary constants, and the general solution of Eq. (2.16) is
y
AtBet
2:21
Example 2.11
Solve the equation ( D2ÿ4D4y0 given that y1a n d Dy3 when t0:
Solution: The auxiliary equation is p2ÿ4p4
pÿ220 which has one
root p2. The general solution is therefore, from Eq. (2.21)
y
AtBe2t:
Since y1 when t0, we have B1. Now
y02
AtBe2tAe2t
and since Dy3 when t0,
32BA:
Hence A1 and the solution is
y
t1e2t:
Finding the particular integral
The particular integral is a solution of Eq. (2.15) that takes the term f
ton the
right hand side into account. The complementary function is transient in nature,so from a physical point of view, the particular integral will usually dominate the
response of the system at large times.
The method of determining the particular integral is to guess a suitable func-
tional form containing arbitrary constants, and then to choose the constants to
ensure it is indeed the solution. If our guess is incorrect, then no values of these
constants will satisfy the di/C128erential equation, and so we have to try a di/C128erent
form. Clearly this procedure could take a long time; fortunately, there are some
guiding rules on what to try for the common examples of f(t):
77SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS
(1)f
ta polynomial in t.
Iff
tis a polynomial in twith highest power tn, then the trial particular
integral is also a polynomial in t, with terms up to the same power. Note
that the trial particular integral is a power series in t, even if f
tcontains
only a single terms Atn.
(2)f
tAekt.
The trial particular integral is yBekt.
(3)f
tAsinktorAcoskt.
The trial particular integral is yAsinktCcoskt. That is, even though
f
tcontains only a sine or cosine term, we need both sine and cosine terms
for the particular integral.
(4)f
tAetsin/C12torAetcos/C12t.
The trial particular integral is yet
Bsin/C12tCcos/C12t.
(5)f
tis a polynomial of order nint, multiplied by ekt.
The trial particular integral is a polynomial in twith coecients to be
determined, multiplied by ekt.
(6)f
tis a polynomial of order nint, multiplied by sin kt.
The trial particular integral is y/C6n
j0
BjsinktCjcoskttj. Can we try
y
BsinktCcoskt/C6nj0Djtj/C63 The answer is no. Do you know why/C63
If the trial particular integral or part of it is identical to one of the terms of the
complementary function, then the trial particular integral must be multiplied by
an extra power of t. Therefore, we need to find the complementary function before
we try to work out the particular integral. What do we mean by ‘identical inform’/C63 It means that the ratio of their t-dependences is a constant. Thus ÿ2e
ÿt
andAeÿtare identical in form, but eÿtandeÿ2tare not.
Particular integral and the operator /C68
d=dx
We now describe an alternative method that can be used for finding particular
integrals. As compared with the method described in previous section, it involves
less guesswork as to what the form of the solution is, and the constants multi-plying the functional forms of the answer are obtained automatically. It does,
however, require a fair amount of practice to ensure that you are familiar with
how to use it.
The technique involves using the di/C128erential operator Dd
=dt, which is an
interesting and simple example of a linear operator without a matrix representa-
tion. It is obvious that /C68obeys the relevant laws of operator algebra: suppose f
and gare functions of t, and ais a constant, then
(i)D
f/C103DfD/C103 (distributive);
(ii)DafaDf (commutative);
(iii)D
nDmfDnmf (index law).
78ORDINARY DIFFERENTIAL EQUATIONS
We can form a polynomial function of /C68and write
/C70
Da0Dna1Dnÿ1 anÿ1Dan
so that
/C70
Df
ta0Dnfa1Dnÿ1f anÿ1Dfanf
and we can interpret Dÿ1as follows
Dÿ1Df
tf
t
andZ
Dfdtf:
Hence Dÿ1indicates the operation of integration (the inverse of di/C128erentiation).
Similarly Dÿmfmeans ‘integrate f
tmtimes’.
These properties of the linear operator /C68can be used to find the particular
integral of Eq. (2.15):
d2y
dt2ady
dtbyD2aDbÿ
yf
t
from which we obtain
y1
D2aDbf
t1
/C70
Df
t;
2:22
where
/C70
DD2aDb:
The trouble with Eq. (2.22) is that it contains an expression involving /C68s in the
denominator. It requires a fair amount of practice to use Eq. (2.22) to express yin
terms of conventional functions. For this, there are several rules to help us.
Rules for /C68operators
Given a power series of /C68
/C71
Da0a1D anDn
and since Dnetnet, it follows that
/C71
Det
a0a1D anDn et/C71
et:
Thus we have
Rule (a): /C71
Det/C71
etprovided /C71
is convergent.
When /C71
Dis the expansion of 1 =/C70
Dthis rule gives
1
/C70
Det1
/C70
etprovided /C70
6 0:
79SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS
Now let us operate /C71
Don a product function et/C86
t:
/C71
Det/C86
t /C71
Det/C86
tet/C71
D/C86
t
et/C71
/C71
D/C86
tet/C71
D/C86
t:
That is, we have
Rule (b): /C71
Det/C86
t et/C71
D/C86
t:
Thus, for example
D2ett2et
D2t2:
Rule (c): /C71
D2sinkt/C71
ÿk2sinkt:
Thus, for example
1
D2
sin 3tÿ1
9sin 3t:
Example 2.12 /C68amped oscillations (Fig. 2.2)
Suppose we have a spring of natural length L(that is, in its unstretched state). If
we hang a ball of mass mfrom it and leave the system in equilibrium, the spring
stretches an amount d, so that the ball is now Ldfrom the suspension point.
We measure the vertical displacement of the ball from this static equilibrium
point. Thus, Ldisy0, and yis chosen to be positive in the downward
direction, and negative upward. If we pull down on the ball and then release it,
it oscillates up and down about the equilibrium position. To analyze the oscilla-
tion of the ball, we need to know the forces acting on it:
80ORDINARY DIFFERENTIAL EQUATIONS
Figure 2.2. Damped spring system.
(1) the downward force of gravity, mg:
(2) the restoring force kywhich always opposes the motion (Hooke’s law),
where kis the spring constant of the spring. If the ball is pulled down a
distance yfrom its static equilibrium position, this force is ÿk
dy.
Thus, the total net force acting on the ball is
m/C103ÿk
dym/C103ÿkdÿky:
In static equilibrium, y0 and all forces balances. Hence
kdm/C103
and the net force acting on the spring is just ÿky; and the equation of motion of
the ball is given by Newton’s second law of motion:
md2y
dt2ÿky;
which describes free oscillation of the ball. If the ball is connected to a dashpot
(Fig. 2.2), a damping force will come into play. Experiment shows that the damp-
ing force is given by ÿbdy=dt, where the constant bis called the damping constant.
The equation of motion of the ball now is
md2y
dt2ÿkyÿbdy
dtory00b
my0k
my0:
The auxiliary equation is
/C1122b
m/C112k
m0
with roots
/C1121ÿb
2m1
2m
b2ÿ4km/C112
; /C1122ÿb
2mÿ1
2mb
2ÿ4km/C112
:
We now have three cases, resulting in quite di/C128erent motions of the oscillator.
/C67ase 1 b2ÿ4km/C620 (overdamping)
The solution is of the form
y
tc1e/C1121tc2e/C1122t:
Now, both band kare positive, so
1
2m
b2ÿ4km/C112
<b
2m
and accordingly
/C1121ÿb
2m1
2m
b2ÿ4km/C112
<0:
81SECOND-ORDER EQUATIONS WITH CONSTANT COEFFICIENTS
Obviously /C1122<0 also. Thus, y
t!0a s t!1 . This means that the oscillation
dies out with time and eventually the mass will assume the static equilibrium
position.
/C67ase 2 b2ÿ4km0 (critical damping)
The solution is of the form
y
teÿbt=2m
c1c2t:
As both bandmare positive, y
t!0a st!1 as in case 1. But c1andc2play a
significant role here. Since eÿbt=2m60 for finite t,y
tcan be zero only when
c1c2t0, and this happens when
tÿc1=c2:
If the number on the right is positive, the mass passes through the equilibrium
position y0 at that time. If the number on the right is negative, the mass never
passes through the equilibrium position.
It is interesting to note that c1y
0, that is, c1measures the initial position.
Next, we note that
y0
0c2ÿbc1=2m;orc2y0
0by
0=2m:
/C67ase 3 b2ÿ4km<0 (underdamping)
The auxiliary equation now has complex roots
/C1121ÿb
2mi
2m
4kmÿb2/C112
;/C1122ÿb
2mÿi
2m4kmÿb
2/C112
and the solution is of the form
y
teÿbt=2mc1cos4kmÿb
2/C112t
2mc2sin4kmÿb
2/C112t
2m /C104/C105
;
which can be rewritten as
y
tceÿbt=2mcos
/C33tÿ;
where
c
c2
1c22q
;tanÿ1c2
c1
;and /C33
4kmÿb2/C112
=2m:
As in case 2, eÿbt=2m!0a s t!1 , and the oscillation gradually dies down to
zero with increasing time. As the oscillator dies down, it oscillates with a fre-
quency /C33=2. But the oscillation is not periodic.
82ORDINARY DIFFERENTIAL EQUATIONS
/C84he /C69uler linear equation
The linear equation with variable coecients
xndny
dxn/C1121xnÿ1dnÿ1y
dxnÿ1 /C112nÿ1xdy
dx/C112nyf
x;
2:23
in which the derivative of the jth order is multiplied by xjand by a constant, is
known as the Euler or Cauchy equation. It can be reduced, by the substitution
xet, to a linear equation with constant coecients with tas the independent
variable. Now if xet, then dx=dtx,a n d
dy
dxdy
dtdt
dx1
xdy
dt;orxdy
dxdy
dt
and
d2y
dx2d
dxdy
dx
d
dt1
xdy
dtdt
dx1
xd
dt1
xdy
dt
or
xd2y
dx21
xd2y
dt2dy
dtd
dt1
x
1
xd2y
dx2ÿ1
xdy
dt
and hence
x2d2y
dx2d2y
dt2ÿdy
dtd
dtdy
dtÿ1
y:
Similarly
x3d3y
dx3d
dtd
dtÿ1d
dtÿ2
y;
and
xndny
dxnd
dtd
dtÿ1d
dtÿ2
d
dtÿn1
y:
Substituting for xj
djy=dxjin Eq. (2.23) the equation transforms into
dny
dtn/C1131dnÿ1y
dtnÿ1 /C113nÿ1dy
dt/C113nyf
et
in which /C1131,/C1132;...;/C113nare constants.
Example 2.13
Solve the equation
x2d2y
dx26xdy
dx6y1
x2:
83THE EULER LINEAR EQUATION
Solution: Putxet, then
xdy
dxdy
dt;x2d2y
dx2d2y
dt2ÿdy
dt:
Substituting these in the equation gives
d2y
dt25dy
dt6yet:
The auxiliary equation /C11225/C1126
/C1122
/C11230 has two roots: /C1121ÿ2,
/C11223. So the complementary function is of the form ycAeÿ2tBeÿ3tand the
particular integral is
y/C1121
D2
D3eÿ2tteÿ2t:
The general solution is
yAeÿ2tBeÿ3tteÿ2t:
The Euler equation is a special case of the general linear second-order equation
D2y/C112
xDy/C113
xyf
x;
where /C112
x,/C113
x, and f
xare given functions of x. In general this type of
equation can be solved by series approximation methods which will be introduced
in next section, but in some instances we may solve it by means of a variable
substitution, as shown by the following example:
D2y
4xÿxÿ1Dy4x2y0;
where
/C112
x
4xÿxÿ1;/C113
x4x2;and f
x0:
If we let
xz1=2
the above equation is transformed into the following equation with constant
coecients:
D2y2Dyy0;
which has the solution
y
ABzeÿz:
Thus the general solution of the original equation is y
ABx2eÿx2:
84ORDINARY DIFFERENTIAL EQUATIONS
/C83olutions in po/C119er series
In many problems in physics and engineering, the di/C128erential equations are of
such a form that it is not possible to express the solution in terms of elementary
functions such as exponential, sine, cosine, etc.; but solutions can be obtained as
convergent infinite series. What is the basis of this method/C63 To see it, let us
consider the following simple second-order linear di/C128erential equation
d2y
dx2y0:
Now assuming the solution is given by ya0a1xa2x2 , we further
assume the series is convergent and di/C128erentiable term by term for suciently
small x. Then
dy=dxa12a2x3a3x2
and
d2y=dx22a223a3x34a4x2 :
Substituting the series for yandd2y=dx2in the given di/C128erential equation and
collecting like powers of xyields the identity
2a2a0
23a3a1x
34a4a2x2 0:
Since if a power series is identically zero all of its coecients are zero, equating to
zero the term independent of xand coecients of x,x2;...;gives
2a2a00; 45a5a30;
23a3a10; 56a6a40;
34a4a20;
and it follows that
a2ÿa0
2;a3ÿa1
23ÿa1
3/C33;a4ÿa2
34ÿa0
4/C33
a5ÿa3
45a1
5/C33;a6ÿa4
56ÿa0
6/C33;...:
The required solution is
ya01ÿx2
2/C33x4
4/C33ÿx6
6/C33ÿ/C32!
a1xÿx3
3/C33x5
5/C33ÿ/C32!
;
you should recognize this as equivalent to the usual solutionya
0cosxa1sinx,a0anda1being arbitrary constants.
85SOLUTIONS IN POWER SERIES
Ordinary and singular points of a di/C128erential equation
We shall concentrate on the linear second-order di/C128erential equation of the form
d2y
dx2P
xdy
dx/C81
xy0
2:24
which plays a very important part in physical problems, and introduce certain
definitions and state (without proofs) some important results applicable to equa-
tions of this type. With some small modifications, these are applicable to linear
equation of any order. If both the functions Pand/C81can be expanded in Taylor
series in the neighborhood of x, then Eq. (2.24) is said to possess an ordinary
point at x. But when either of the functions Por/C81does not possess a Taylor
series in the neighborhood of x, Eq. (2.24) is said to have a singular point at
x.I f
P
x=
xÿand /C81/C22
x=
xÿ2
and
xand/C22
xcan be expanded in Taylor series near x. In such cases,
xis a singular point but the singularity is said to be regular.
Frobenius and Fuchs theorem
Frobenius and Fuchs showed that:
(1) If P
xand/C81
xare regular at x, then the di/C128erential equation (2.24)
possesses two distinct solutions of the form
yX1
0a
xÿ
a060:
2:25
(2) If P
xand/C81
xare singular at x, but
xÿP
xand
xÿ2/C81
x
are regular at x, then there is at least one solution of the di/C128erential
equation (2.24) of the form
yX1
0a
xÿ/C26
a060;
2:26
where /C26is some constant, which is valid for jxÿj</C12whenever the Taylor
series for
xand/C22
xare valid for these values of x.
(3) If P
xand/C81
xare irregular singular at x(that is,
xand/C22
xare
singular at x, then regular solutions of the di/C128erential equation (2.24)
may not exist.
86ORDINARY DIFFERENTIAL EQUATIONS
The proofs of these results are beyond the scope of the book, but they can be
found, for example, in E. L. Ince’s Ordinary /C68i/C128erential E/C113uations , Dover
Publications Inc., New York, 1944.
The first step in finding a solution of a second-order di/C128erential equation
relative to a regular singular point xis to determine possible values for the
index /C26in the solution (2.26). This is done by substituting series (2.26) and its
appropriate di/C128erential coecients into the di/C128erential equation and equating to
zero the resulting coecient of the lowest power of xÿ.This leads to a quadratic
equation, called the indicial equation, from which suitable values of /C26can be
found. In the simplest case, these values of /C26will give two di/C128erent series solutions
and the general solution of the di/C128erential equation is then given by a linear
combination of the separate solutions. The complete procedure is shown in
Example 2.14 below.
Example 2.14
Find the general solution of the equation
4xd2y
dx22dy
dxy0:
Solution: The origin is a regular singular point and, writing
y/C61
0ax/C26
a060we have
dy=dxX1
0a
/C26x/C26ÿ1;d2y=dx2X1
0a
/C26
/C26ÿ1x/C26ÿ2:
Before substituting in the di/C128erential equation, it is convenient to rewrite it in the
form
4xd2y
dx22dy
dx()
yfg0:
When ax/C26is substituted for y, each term in the first bracket yields a multiple of
x/C26ÿ1, while the second bracket gives a multiple of x/C26and, in this form, the
di/C128erential equation is said to be arranged according to weight, the weights of the
bracketed terms di/C128ering by unity. When the assumed series and its di/C128erential
coecients are substituted in the di/C128erential equation, the term containing the
lowest power of xis obtained by writing ya0x/C26in the first bracket. Since the
coecient of the lowest power of xmust be zero and, since a060, this gives the
indicial equation
4/C26
/C26ÿ12/C262/C26
2/C26ÿ10;
its roots are /C260,/C261=2.
87SOLUTIONS IN POWER SERIES
The term in x/C26is obtained by writing ya1x/C261in first bracket and
yax/C26in the second. Equating to zero the coecient of the term obtained in
this way we have
f4
/C261
/C262
/C261ga1a0;
giving, with replaced by n,
an1ÿ1
2
/C26n1
2/C262n1an:
This relation is true for n1;2;3;...and is called the recurrence relation for the
coecients. Using the first root /C260 of the indicial equation, the recurrence
relation gives
an11
2
n1
2n1an
and hence
a1ÿa0
2;a2ÿa1
12a0
4/C33;a3ÿa2
30ÿa0
6/C33;...:
Thus one solution of the di/C128erential equation is the series
a01ÿx
2/C33x2
4/C33ÿx3
6/C33ÿ/C32!
:
With the second root /C261=2, the recurrence relation becomes
an1ÿ1
2n3
2n2an:
Replacing a0(which is arbitrary) by b0, this gives
a1ÿb0
32ÿb0
3/C33;a2ÿa1
54b0
5/C33;a3ÿa2
76ÿb0
7/C33;...:
and a second solution is
b0x1=21ÿx
3/C33x2
5/C33ÿx3
7/C33ÿ/C32!
:
The general solution of the equation is a linear combination of these two solu-
tions.
Many physical problems require solutions which are valid for large values of
the independent variable x. By using the transformation x1=t, the di/C128erential
equation can be transformed into a linear equation in the new variable tand the
solutions required will be those valid for small t.
In Example 2.14 the indicial equation has two distinct roots. But there are two
other possibilities: ( a) the indicial equation has a double root; ( b) the roots of the
88ORDINARY DIFFERENTIAL EQUATIONS
indicial equation di/C128er by an integer. We now take a general look at these cases.
For this purpose, let us consider the following di/C128erential equation which is highlyimportant in mathematical physics:
x
2y00x/C103
xy0/C104
xy0;
2:27
where the functions /C103
xand/C104
xare analytic at x0. Since the coecients are
not analyic at x0, the solution is of the form
y
xxrX1
m0amxm
a060:
2:28
We first expand /C103
xand/C104
xin power series,
/C103
x/C1030/C1031x/C1032x2 /C104
x/C1040/C1041x/C1042x2 :
Then di/C128erentiating Eq. (2.28) term by term, we find
y0
xX1
m0
mramxmrÿ1;y00
xX1
m0
mr
mrÿ1amxmrÿ2:
By inserting all these into Eq. (2.27) we obtain
xrr
rÿ1a0
/C1030/C1031x xr
ra0
/C1040/C1041x xr
a0a1x 0:
Equating the sum of the coecients of each power of xto zero, as before, yields a
system of equations involving the unknown coecients am. The smallest power is
xr, and the corresponding equation is
r
rÿ1/C1030r/C1040a00:
Since by assumption a060, we obtain
r
rÿ1/C1030r/C10400o r r2
/C1030ÿ1r/C10400:
2:29
This is the indicial equation of the di/C128erential equation (2.27). We shall see thatour series method will yield a fundamental system of solutions; one of the solu-
tions will always be of the form (2.28), but for the form of other solution there will
be three di/C128erent possibilities corresponding to the following cases.
Case 1 The roots of the indicial equation are distinct and do not di/C128er by an
integer.
Case 2 The indicial equation has a double root.
Case 3 The roots of the indicial equation di/C128er by an integer.
We now discuss these cases separately.
89SOLUTIONS IN POWER SERIES
/C67ase 1 /C68istinct roots not di/C128ering by an integer
This is the simplest case. Let r1andr2be the roots of the indicial equation (2.29).
If we insert rr1into the recurrence relation and determine the coecients
a1,a2;...successively, as before, then we obtain a solution
y1
xxr1
a0a1xa2x2 :
Similarly, by inserting the second root rr2into the recurrence relation, we will
obtain a second solution
y2
xxr2
a0/C42a1/C42xa2/C42x2 :
Linear independence of y1andy2follows from the fact that y1=y2is not constant
because r1ÿr2is not an integer.
/C67ase 2 /C68ouble rootsThe indicial equation (2.29) has a double root rif, and only if,
/C103
0ÿ12ÿ4/C10400, and then r
1ÿ/C1030=2. We may determine a first solution
y1
xxr
a0a1xa2x2 r1ÿ/C1030
2
2:30
as before. To find another solution we may apply the method of variation of
parameters, that is, we replace constant cin the solution cy1
xby a function
u
xto be determined, such that
y2
xu
xy1
x
2:31
is a solution of Eq. (2.27). Inserting y2and the derivatives
y0
2u0y1uy0
1 y00
2u00y12u0y0
1uy00
1
into the di/C128erential equation (2.27) we obtain
x2
u00y12u0y0
1uy00
1x/C103
u0y1uy0
1/C104uy 10
or
x2y1u002x2y0
1u0x/C103y 1u0
x2y00
1x/C103y0
1/C104y1u0:
Since y1is a solution of Eq. (2.27), the quantity inside the bracket vanishes; and
the last equation reduces to
x2y1u002x2y0
1u0x/C103y 1u00:
Dividing by x2y1and inserting the power series for gwe obtain
u002y0
1
y1/C1030
x
u00:
Here and in the following the dots designate terms which are constants or involvepositive powers of x. Now from Eq. (2.30) it follows that
y0
1
y1xrÿ1ra0
r1a1x
xra0a1x 1
xra0
r1a1x
a0a1xr
x :
90ORDINARY DIFFERENTIAL EQUATIONS
Hence the last equation can be written
u002r/C1030
x
u00:
2:32
Since r
1ÿ/C1030=2 the term
2r/C1030=xequals 1 =x, and by dividing by u0we
thus have
u00
u0ÿ1
x :
By integration we obtain
lnu0ÿlnx oru01
xe
...:
Expanding the exponential function in powers of xand integrating once more, we
see that the expression for uwill be of the form
ulnxk1xk2x2 :
By inserting this into Eq. (2.31) we find that the second solution is of the form
y2
xy1
xlnxxrX1
m1Amxm:
2:33
/C67ase 3 /C82oots di/C128ering by an integer
If the roots r1andr2of the indicial equation (2.29) di/C128er by an integer, say, r1r
andr2rÿ/C112, where pis a positive integer, then we may always determine one
solution as before, namely, the solution corresponding to r1:
y1
xxr1
a0a1xa2x2 :
To determine a second solution y2, we may proceed as in Case 2. The first steps
are literally the same and yield Eq. (2.32). We determine 2 r/C1030in Eq. (2.32).
Then from the indicial equation (2.29), we find ÿ
r1r2/C1030ÿ1. In our case,
r1rand r2rÿ/C112, therefore, /C1030ÿ1/C112ÿ2r. Hence in Eq. (2.32) we have
2r/C1030/C1121, and we thus obtain
u00
u0ÿ/C1121
x
:
Integrating, we find
lnu0ÿ
/C1121lnx oru0xÿ
/C1121e
...;
where the dots stand for some series of positive powers of x. By expanding the
exponential function as before we obtain a series of the form
u01
x/C1121k1
x/C112k/C112
xk/C1121k/C1122x :
91SOLUTIONS IN POWER SERIES
Integrating, we have
uÿ1
/C112x/C112ÿ k/C112lnxk/C1121x :
2:34
Multiplying this expression by the series
y1
xxr1
a0a1xa2x2
and remembering that r1ÿ/C112r2we see that y2uy1is of the form
y2
xk/C112y1
xlnxxr2X1
m0amxm:
2:35
While for a double root of Eq. (2.29) the second solution always contains a
logarithmic term, the coecient k/C112may be zero and so the logarithmic term may
be missing, as shown by the following example.
Example 2.15
Solve the di/C128erential equation
x2y00xy0
x2ÿ1
4y0:
Solution: Substituting Eq. (2.28) and its derivatives into this equation, we obtain
X1
m0
mr
mrÿ1
mrÿ14amxmrP1
m0amxmr20:
By equating the coecient of xrto zero we get the indicial equation
r
rÿ1rÿ140o r r214:
The roots r112andr2ÿ12di/C128er by an integer. By equating the sum of the
coecients of xsrto zero we find
r1r
rÿ1ÿ14a10
s1:
2:36a
sr
srÿ1srÿ1
4asasÿ20
s2;3;...:
2:36b
For rr112, Eq. (2.36a) yields a10, and the indicial equation (2.36b)
becomes
s1sasasÿ20:
From this and a10 we obtain a30,a50, etc. Solving the indicial equation
forasand setting s2/C112, we get
a2/C112ÿa2/C112ÿ2
2/C112
2/C1121
/C1121;2;...:
92ORDINARY DIFFERENTIAL EQUATIONS
Hence the non-zero coecients are
a2ÿa0
3/C33;a4ÿa2
45a0
5/C33;a6ÿa0
7/C33;etc:;
and the solution y1is
y1
xa0xpX1
m0
ÿ1mx2m
2m1/C33a0xÿ1=2X1
m0
ÿ1mx2m1
2m1/C33a0sinxxp:
2:37
From Eq. (2.35) we see that a second independent solution is of the form
y2
xky1
xlnxxÿ1=2X1
m0amxm:
Substituting this and the derivatives into the di/C128erential equation, we see that the
three expressions involving ln xand the expressions ky1andÿky1drop out.
Simplifying the remaining equation, we thus obtain
2kxy0
1X1
m0m
mÿ1amxmÿ1=2X1
m0amxm3=20:
From Eq. (2.37) we find 2 kxy0ÿka0x1=2 . Since there is no further term
involving x1=2anda060, we must have k0. The sum of the coecients of the
power xsÿ1=2is
s
sÿ1asasÿ2
s2;3;...:
Equating this to zero and solving for as, we have
asÿasÿ2=s
sÿ1
s2;3;...;
from which we obtain
a2ÿa0
2/C33;a4ÿa2
43a0
4/C33;a6ÿa0
6/C33;etc:;
a3ÿa1
3/C33;a5ÿa3
54a1
5/C33;a7ÿa1
7/C33;etc:
We may take a10, because the odd powers would yield a1y1=a0. Then
y2
xa0xÿ1=2X1
m0
ÿ1mx2m
2m/C33a0cosxxp:
/C83imultaneous equations
In some physics and engineering problems we may face simultaneous dif-
ferential equations in two or more dependent variables. The general solution
93SIMULTANEOUS EQUATIONS
of simultaneous equations may be found by solving for each dependent variable
separately, as shown by the following example
Dx2y3x0
3xDyÿ2y0)
Dd=dt
which can be rewritten as
D3x2y0;
3x
Dÿ2y0:)
We then operate on the first equation with ( Dÿ2) and multiply the second by a
factor 2:
Dÿ2
D3x2
Dÿ2y0;
6x2
Dÿ2y0:)
Subtracting the first from the second leads to
D2Dÿ6xÿ6x
D2Dÿ12x0;
which can easily be solved and its solution is of the form
x
tAe3tBeÿ4t:
Now inserting x
tback into the original equation to find ygives:
y
tÿ 3Ae3t1
2Beÿ4t:
/C84he gamma and beta functions
The factorial notation n/C33n
nÿ1
nÿ2 321 has proved useful in
writing down the coecients in some of the series solutions of the di/C128erential
equations. However, this notation is meaningless when nis not a positive integer.
A useful extension is provided by the gamma (or Euler) function, which is definedby the integral
ÿ
Z
1
0eÿxxÿ1dx
/C620
2:38
and it follows immediately that
ÿ
1Z1
0eÿxdx ÿ eÿx1
01:
2:39
Integration by parts gives
ÿ
1Z1
0eÿxxdx ÿ eÿxx10Z1
0eÿxxÿ1dxÿ
:
2:40
94ORDINARY DIFFERENTIAL EQUATIONS
When n, a positive integer, repeated application of Eq. (2.40) and use of Eq.
(2.39) gives
ÿ
n1nÿ
nn
nÿ1ÿ
nÿ1 ...n
nÿ1 32ÿ
1
n
nÿ1 321n/C33:
Thus the gamma function is a generalization of the factorial function. Eq. (2.40)
enables the values of the gamma function for any positive value of to be
calculated: thus
ÿ
7
2
52ÿ
52
52
32ÿ
32
52
32
12ÿ
12:
Write uxpin Eq. (2.38) and we then obtain
ÿ
2Z1
0u2ÿ1eÿu2du;
so that
ÿ
1
22Z1
0eÿu2dup:
The function ÿ
has been tabulated for values of between 0 and 1.
When <0 we can define ÿ
with the help of Eq. (2.40) and write
ÿ
ÿ
1=:
Thus
ÿ
ÿ3
2ÿ23ÿ
ÿ12ÿ23
ÿ21ÿ
1243p:
When !0;R1
0eÿxxÿ1dxdiverges so that ÿ
0is not defined.
Another function which will be useful later is the beta function which is defined
by
B
/C112;/C113Z1
0t/C112ÿ1
1ÿt/C113ÿ1dt
/C112;/C113/C620:
2:41
Substituting t/C118=
1/C118, this can be written in the alternative form
B
/C112;/C113Z1
0/C118/C112ÿ1
1/C118ÿ/C112ÿ/C113d/C118:
2:42
By writing t01ÿtwe deduce that B
/C112;/C113B
/C113;/C112.
The beta function can be expressed in terms of gamma functions as follows:
B
/C112;/C113ÿ
/C112ÿ
/C113
ÿ
/C112/C113:
2:43
To prove this, write xat(a/C620) in the integral (2.38) defining ÿ
, and it is
straightforward to show that
ÿ
aZ1
0eÿattÿ1dt
2:44
95THE GAMMA AND BETA FUNCTIONS
and, with /C112/C113,a1/C118, this can be written
ÿ
/C112/C113
1/C118ÿ/C112ÿ/C113Z1
0eÿ
1/C118tt/C112/C113ÿ1dt:
Multiplying by /C118/C112ÿ1and integrating with respect to /C118between 0 and 1,
ÿ
/C112/C113Z1
0/C118/C112ÿ1
1/C118ÿ/C112ÿ/C113d/C118Z1
0/C118/C112ÿ1d/C118Z1
0eÿ
1/C118tt/C112/C1131dt:
Then interchanging the order of integration in the double integral on the right and
using Eq. (2.42),
ÿ
/C112/C113B
/C112;/C113Z1
0eÿtt/C112/C113ÿ1dtZ1
0eÿ/C118t/C118/C112ÿ1d/C118
Z1
0eÿtt/C112/C113ÿ1ÿ
/C112
t/C112dt;using Eq :
2:44
ÿ
/C112Z1
0eÿtt/C113ÿ1dtÿ
/C112ÿ
/C113:
Example 2.15
Evaluate the integralZ1
03ÿ4x2dx:
Solution: We first notice that 3 eln 3, so we can rewrite the integral as
Z1
03ÿ4x2dxZ1
0
eln 3
ÿ4x2dxZ1
0eÿ
4l n3x2dx:
Now let (4 ln 3) x2z, then the integral becomes
Z1
0eÿzdz1=2
4l n3p/C32!
1
24l n3pZ
1
0zÿ1=2eÿzdzÿ
1
2
2
4l n3p p
2
4l n3p :
Problems
2.1 Solve the following equations:
(a)xdy=dxy21;
(b)dy=dx
xy2.
2.2 Melting of a sphere of ice: Assume that a sphere of ice melts at a rate
proportional to its surface area. Find an expression for the volume at any
time t.
2.3 Show that
3x2ycosxdx
sinxÿ4y3dy0 is an exact di/C128erential
equation and find its general solution.
96ORDINARY DIFFERENTIAL EQUATIONS
2.4 /C82/C67circuits: A typical /C82/C67circuit is shown in Fig. 2.3. Find current flow I
t
in the circuit, assuming /C69
t/C690.
Hint: the voltage drop across the capacitor is given /C81/C47/C67, with /C81(t) the
charge on the capacitor at time t.
2.5 Find a constant such that
xyis an integrating factor of the equation
4x22xy6ydx
2x29y3xdy0:
What is the solution of this equation/C63
2.6 Solve dy=dxyy3x:
2.7 Solve:
(a) the equation
D2ÿDÿ12y0 with the boundary conditions y0,
Dy3 when t0;
(b) the equation
D22D3y0 with the boundary conditions y2,
Dy0 when t0;
(c) the equation
D2ÿ2D1y0 with the boundary conditions y5,
Dy3 when t0.
2.8 Find the particular integral of
D22Dÿ1y3t3.
2.9 Find the particular integral of
2D25D73e2t.
2.10 Find the particular integral of
3D2Dÿ5ycos 3 t:
2.11 Simple harmonic motion of a pendulum (Fig. 2.4): Suspend a ball of mass m
at the end of a massless rod of length Land set it in motion swinging back
and forth in a vertical plane. Show that the equation of motion of the ball is
d2
dt2/C103
Lsin0;
where gis the local gravitational acceleration. Solve this pendulum equation
for small displacements by replacing sin by.
2.12 Forced oscillations with damping: If we allow an external driving force /C70
t
in addition to damping (Example 2.12), the motion of the oscillator is
governed by
y00b
my0k
my/C70
t;
97PROBLEMS
Figure 2.3. /C82/C67circuit.
a constant coecient non-homogeneous equation. Solve this equation for
/C70
tAcos
/C33t:
2.13 Solve the equation
r2d2R
dr22rdR
drÿn
n1R0
nconstant :
2.14 The first-order non-linear equation
dy
dxy2/C81
xyR
x0
is known as Riccati’s equation. Show that, by use of a change of dependentvariable
y1
zdz
dx;
Riccati’s equation transforms into a second-order linear di/C128erential equa-tion
d2z
dx2/C81
xdz
dxR
xz0:
Sometimes Riccati’s equation is written as
dy
dxP
xy2/C81
xyR
x0:
Then the transformation becomes
yÿ1
P
xdz
dx
and the second-order equation takes the form
d2z
dx2/C811
PdP
dxdz
dxPRz0:
98ORDINARY DIFFERENTIAL EQUATIONS
Figure 2.4. Simple pendulum.
2.15 Solve the equation 4 x2y004xy0
x2ÿ1y0 by using Frobenius’
method, where y0dy=dx,a n d y00d2y=dx2.
2.16 Find a series solution, valid for large values of x, of the equation
1ÿx2y00ÿ2xy02y0:
2.17 Show that a series solution of Airy’s equation y00ÿxy0i s
ya01x3
23x6
2356/C32!
b0xx4
34x7
3467/C32!
:
2.18 Show that Weber’s equation y00
n1
2ÿ14x2y0 is reduced by the sub-
stitution yeÿx2=4/C118to the equation d2/C118=dx2ÿx
d/C118=dxn/C1180. Show
that two solutions of this latter equation are
/C11811ÿn
2/C33x2n
nÿ2
4/C33x4ÿn
nÿ2
nÿ4
6/C33x6ÿ ;
/C1182xÿ
nÿ1
3/C33x3
nÿ1
nÿ3
5/C33x5ÿ
nÿ1
nÿ3
nÿ5
7/C33x7ÿ :
2.19 Solve the following simultaneous equations
Dxyt3
Dyÿxt)
Dd=dt:
2.20 Evaluate the integrals:
(a)Z1
0x3eÿxdx:
(b)Z1
0x6eÿ2xdx
hint: let y2x:
(c)Z1
0ypeÿy2dy
hint: let y2x:
(d)Z1
0dx
ÿlnxp
hint : let ÿlnxu:
2.21 ( a) Prove that B
/C112;/C1132Z=2
0sin2/C112ÿ1cos2/C113ÿ1d.
(b) Evaluate the integralZ1
0x4
1ÿx3dx:
2.22 Show that n/C33/C25
2np
nneÿn. This is known as Stirling’s factorial approxima-
tion or asymptotic formula for n/C33.
99PROBLEMS
3
Matrix algebra
As vector methods have become standard tools for physicists, so too matrix
methods are becoming very useful tools in sciences and engineering. Matrices
occur in physics in at least two ways: in handling the eigenvalue problems in
classical and quantum mechanics, and in the solutions of systems of linear equa-
tions. In this chapter, we introduce matrices and related concepts, and define some
basic matrix algebra. In Chapter 5 we will discuss various operations with
matrices in dealing with transformations of vectors in vector spaces and the
operation of linear operators on vector spaces.
/C68efinition of a matri/C120
A matrix consists of a rectangular block or ordered array of numbers that obeys
prescribed rules of addition and multiplication. The numbers may be real or
complex. The array is usually enclosed within curved brackets. Thus
12 4
2ÿ17
is a matrix consisting of 2 rows and 3 columns, and it is called a 2 3( 2b y3 )
matrix. An mnmatrix consists of mrows and ncolumns, which is usually
expressed in a double sux notation:
~Aa11a12a13 a1n
a21a22a23 ... a2n
............
am1am2am3 ...amn0
BBBB@1
CCCCA:
3:1
Each number a
ijis called an element of the matrix, where the first subscript i
denotes the row, while the second subscript jindicates the column. Thus, a23
100
refers to the element in the second row and third column. The element aijshould
be distinguished from the element aji.
It should be pointed out that a matrix has no single numerical value; therefore it
must be carefully distinguished from a determinant.
We will denote a matrix by a letter with a tilde over it, such as ~Ain (3.1).
Sometimes we write ( aij)o r( aijmn, if we wish to express explicitly the particular
form of element contained in ~A.
Although we have defined a matrix here with reference to numbers, it is easy to
extend the definition to a matrix whose elements are functions fi
x; for a 2 3
matrix, for example, we have
f1
xf2
xf3
x
f4
xf5
xf6
x
:
A matrix having only one row is called a row matrix or a row vector, while a
matrix having only one column is called a column matrix or a column vector. An
ordinary vector /C65A1^e1A2^e2A3^e3can be represented either by a row
matrix or by a column matrix.
If the numbers of rows mand columns nare equal, the matrix is called a square
matrix of order n.
In a square matrix of order n, the elements a11;a22;...;annform what is called
the principal (or leading) diagonal, that is, the diagonal from the top left hand
corner to the bottom right hand corner. The diagonal from the top right hand
corner to the bottom left hand corner is sometimes termed the trailing diagonal.
Only a square matrix possesses a principal diagonal and a trailing diagonal.
The sum of all elements down the principal diagonal is called the trace, or spur,
of the matrix. We write
Tr ~AXn
i1aii:
If all elements of the principal diagonal of a square matrix are unity while all
other elements are zero, then it is called a unit matrix (for a reason to be explained
later) and is denoted by ~I. Thus the unit matrix of order 3 is
~I100
010
0010
B@1
CA:
A square matrix in which all elements other than those along the principal
diagonal are zero is called a diagonal matrix.
A matrix with all elements zero is known as the null (or zero) matrix and is
denoted by the symbol ~0, since it is not an ordinary number, but an array of zeros.
101DEFINITION OF A MATRI/C88
Four basic algebra operations for matrices
Equality of matrices
Two matrices ~A
ajkand ~B
bjkare equal if and only if ~Aand ~Bhave the
same order (equal numbers of rows and columns) and corresponding elements are
equal, that is
ajkbjkfor all jandk:
Then we write
~A~B:
/C65ddition of matrices
Addition of matrices is defined only for matrices of the same order. If ~A
ajk
and ~B
bjkhave the same order, the sum of ~Aand ~Bis a matrix of the same
order
~C~A~B
with elements
cjkajkbjk:
3:2
We see that ~Cis obtained by adding corresponding elements of ~Aand ~B.
Example 3.1
If
~A214
302
; ~B35 121 ÿ3
hen
~C~A~B214
302/C32!
35 121 ÿ3/C32!
231541
32012ÿ3/C32!
56 551 ÿ1/C32!
:
From the definitions we see that matrix addition obeys the commutative and
associative laws, that is, for any matrices ~A,~B,~Cof the same order
~A~B~B~A; ~A
~B~C
~A~B ~C:
3:3
Similarly, if ~A
a
jkand ~B
bjk) have the same order, we define the di/C128er-
ence of ~Aand ~Bas
~D~Aÿ~B
102MATRI/C88 ALGEBRA
with elements
djkajkÿbjk:
3:4
/C77ultiplication of a matrix by a number
If~A
ajkandcis a number (or scalar), then we define the product of ~Aandcas
c~A~Ac
cajk;
3:5
we see that c~Ais the matrix obtained by multiplying each element of ~Abyc.
We see from the definition that for any matrices and any numbers,
c
~A~Bc~Ac~B;
ck~Ac~Ak~A;c
k~Ack~A:
3:6
Example 3.2
7abc
def
7a7b7c
7d7e7f
:
Formulas (3.3) and (3.6) express the properties which are characteristic for a
vector space. This gives vector spaces of matrices. We will discuss this further in
Chapter 5.
/C77atrix multiplication
The matrix product ~A~Bof the matrices ~Aand ~Bis defined if and only if the
number of columns in ~Ais equal to the number of rows in ~B. Such matrices
are sometimes called ‘conformable’. If ~A
ajkis an nsmatrix and
~B
bjkis an smmatrix, then ~Aand ~Bare conformable and their matrix
product, written ~C~A~B,i sa n nmmatrix formed according to the rule
cikXs
j1aijbjk;i1;2;...;nk 1;2;...;m:
3:7
Consequently, to determine the ijth element of matrix ~C, the corresponding terms
of the ith row of ~Aandjth column of ~Bare multiplied and the resulting products
added to form cij.
Example 3.3Let
~A21 4
ÿ302
; ~B35
2ÿ1
420
B@1
CA
103FOUR BASIC ALGEBRA OPERATIONS FOR MATRICES
then
~A~B2312442 51
ÿ 142
ÿ330224
ÿ350
ÿ 122/C32!
24 17
ÿ1ÿ11/C32!
:
The reader should master matrix multiplication, since it is used throughout the
rest of the book.
In general, matrix multiplication is not commutative: ~A~B6~B~A. In fact, ~B~Ais
often not defined for non-square matrices, as shown in the following example.
Example 3.4
If
~A12
34/C32!
; ~B37/C32!
then
~A~B12
3437
1327
3347
1737
:
But
~B~A371234
is not defined.
Matrix multiplication is associative and distributive:
~A~B~C~A
~B~C;
~A~B~C~A~C~B~C:
To prove the associative law, we start with the matrix product ~A~B, then multi-
ply this product from the right by ~C:
~A~BX
kaikbkj;
~A~B~CX
jX
kaikbkj/C32!
cjs"#
X
kaikX
jbkjcjs/C32!
~A
~B~C:
Products of matrices di/C128er from products of ordinary numbers in many
remarkable ways. For example, ~A~B0 does not imply ~A0o r ~B0. Even
more bizarre is the case where ~A20, ~A60; an example of which is
~A0100
:
104MATRI/C88 ALGEBRA
When you first run into Eq. (3.7), the rule for matrix multiplication, you might
ask how anyone would arrive at it. It is suggested by the use of matrices in
connection with linear transformations. For simplicity, we consider a very simple
case: three coordinates systems in the plane denoted by the x1x2-system, the y1y2-
system, and the z1z2-system. We assume that these systems are related by the
following linear transformations
x1a11y1a12y2;x2a21y1a22y2;
3:8
y1b11z1b12z2;y2b21z1b22z2:
3:9
Clearly, the x1x2-coordinates can be obtained directly from the z1z2-coordinates
by a single linear transformation
x1c11z1c12z2;x2c21z1c22z2;
3:10
whose coecients can be found by inserting (3.9) into (3.8),
x1a11
b11z1b12z2a12
b21z1b22z2;
x2a21
b11z1b12z2a22
b21z1b22z2:
Comparing this with (3.10), we find
c11a11b11a12b21;c12a11b12a12b22;
c21a21b11a22b21;c22a21b12a22b22;
or briefly
cjkX2
i1ajibik;j;k1;2;
3:11
which is in the form of (3.7).
Now we rewrite the transformations (3.8), (3.9) and (3.10) in matrix form:
~X~A~Y; ~Y~B~Z;and ~X~C~Z;
where
~Xx1
x2/C32!
; ~Yy1
y2/C32!
; ~Zz1
z2/C32!
;
~Aa11a12
a21a22/C32!
; ~Bb11b12
b21b22/C32!
;C:::
c11c12
c21c22/C32!
:
We then see that ~C~A~B, and the elements of ~Care given by (3.11).
Example 3.5
Rotations in three-dimensional space: An example of the use of matrix multi-
plication is provided by the representation of rotations in three-dimensional
105FOUR BASIC ALGEBRA OPERATIONS FOR MATRICES
space. In Fig. 3.1, the primed coordinates are obtained from the unprimed coor-
dinates by a rotation through an angle about the x3-axis. We see that x0
1is the
sum of the projection of x1onto the x0
1-axis and the projection of x2onto the x0
1-
axis:
x0
1x1cosx2cos
=2ÿx1cosx2sin;
similarly
x0
2x1cos
=2x2cosÿx1sinx2cos
and
x0
3x3:
We can put these in matrix form
X0RX;
where
X0x0
1
x0
2
x0
30
B@1
CA;Xx1
x2
x30
B@1
CA;Rcossin0
ÿsincos0
00 10
B@1
CA:
106MATRI/C88 ALGEBRA
Figure 3.1. Coordinate changes by rotation.
/C84he commutator
Even if matrices ~Aand ~Bare both square matrices of order n, the products ~A~B
and ~B~A, although both square matrices of order n, are in general quite di/C128erent,
since their individual elements are formed di/C128erently. For example,
12
131012
3446
but10121213
1238
:
The di/C128erence between the two products ~A~Band ~B~Ais known as the commu-
tator of ~Aand ~Band is denoted by
~A;~B ~A~Bÿ~B~A:
3:12
It is obvious that
~B;~Aÿ ~A;~B:
3:13
If two square matrices ~Aand ~Bare very carefully chosen, it is possible to make the
product identical. That is ~A~B~B~A. Two such matrices are said to commute with
each other. Commuting matrices play an important role in quantum mechanics.
If~Acommutes with ~Band ~Bcommutes with ~C, it does not necessarily follow
that ~Acommutes with ~C.
Po/C119ers of a matri/C120
Ifnis a positive integer and ~Ais a square matrix, then ~A
2~A~A,~A3~A~A~A, and
in general, ~An~A~A ~A(ntimes). In particular, ~A0~I.
Functions of matrices
As we define and study various functions of a variable in algebra, it is possible to
define and evaluate functions of matrices. We shall briefly discuss the following
functions of matrices in this section: integral powers and exponential.
A simple example of integral powers of a matrix is polynomials such as
f
~A ~A23~A5:
Note that a matrix can be multiplied by itself if and only if it is a square matrix.Thus ~Ahere is a square matrix and we denote the product ~A~Aas ~A
2. More fancy
examples can be obtained by taking series, such as
~SX1
k0ak~Ak;
where akare scalar coecients. Of course, the sum has no meaning if it does not
converge. The convergence of the matrix series means every matrix element of the
107THE COMMUTATOR
infinite sum of matrices converges to a limit. We will not discuss the general
theory of convergence of matrix functions. Another very common series is definedby
e~AX1
n0~An
n/C33:
/C84ranspose of a matri/C120
Consider an mnmatrix ~A, if the rows and columns are systematically changed
to columns to rows, without changing the order in which they occur, the newmatrix is called the transpose of matrix ~A. It is denoted by ~A
T:
~Aa11a12a13 ... a1n
a21a22a23 ... a2n
............
am1am2am3 ...amn0
BBBB@1
CCCCA; ~A
Ta11a21a31 ...am1
a12a22a32 ...am2
............
an1a2na3n ...amn0
BBBB@1
CCCCA:
Thus the transpose matrix has nrows and mcolumns. If ~Ais written as ( a
jk), then
~ATmay be written as
akj).
~A
ajk; ~AT
akj:
3:14
The transpose of a row matrix is a column matrix, and vice versa.
Example 3.6
~A123
456
; ~AT14
25
360
B@1
CA; ~B
123 ; ~BT1
2
30
B@1
CA:
It is obvious that
~ATT~A,a n d
~A~BT~AT~BT. It is also easy to prove
that the transpose of the product is the product of the transposes in reverse:
~A~BT~BT~AT:
3:15
Proof:
~A~BT
ij
~A~Bjiby definition
X
kAjkBki
X
kBT
ikATkj
~BT~ATij
108MATRI/C88 ALGEBRA
so that
ABTBTATq:e:d:
Because of (3.15), even if ~A~ATand ~B~BT,
~A~BT6~A~Bunless the matrices
commute.
/C83/C121mmetric and s/C107e/C119-s/C121mmetric matrices
A square matrix ~A
ajkis said to be symmetric if all its elements satisfy the
equations
akjajk;
3:16
that is, ~Aand its transpose are equal ~A~AT. For example,
~A157
53 ÿ4
7ÿ400
B@1
CA
is a third-order symmetric matrix: the elements of the ith row equal the elements
ofith column, for all i.
On the other hand, if the elements of ~Asatisfy the equations
akjÿajk;
3:17
then ~Ais said to be skew-symmetric, or antisymmetric. Thus, for a skew-sym-
metric ~A, its transpose equals minus ÿ~A:~ATÿ ~A.
Since the elements ajjalong the principal diagonal satisfy the equations
ajjÿajj, it is evident that they must all vanish. For example,
~A0ÿ25
20 1
ÿ5ÿ100
B@1
CA
is a skew-symmetric matrix.
Any real square matrix ~Amay be expressed as the sum of a symmetric matrix ~R
and a skew-symmetric matrix ~S, where
~R1
2
~A~ATand ~S12
~Aÿ~AT:
3:18
Example 3.7
The matrix
~A23
5ÿ1
may be written in the form ~A~R~S, where
~R1
2
~A~AT24
4ÿ1
~S1
2
~Aÿ~AT0ÿ1
10
:
109SYMMETRIC AND SKEW-SYMMETRIC MATRICES
The product of two symmetric matrices need not be symmetric. This is
because of (3.15): even if ~A~ATand ~B~BT,
~A~BT6~A~Bunless the matrices
commute.
A square matrix whose elements above or below the principal diagonal are all
zero is called a triangular matrix. The following two matrices are triangular
matrices:
100
230
5020
B@1
CA;16 ÿ1
02 3
00 40
B@1
CA:
A square matrix ~Ais said to be singular if det ~A0, and non-singular if
det ~A60, where det ~Ais the determinant of the matrix ~A.
/C84he matri/C120 representation of a /C118ector product
The scalar product defined in ordinary vector theory has its counterpart in matrix
theory. Consider two vectors /C65
A1;A2;A3and/C66
B1;B2;B3the counter-
part of the scalar product is given by
~A~BT
A1A2A3B1
B2
B30
B@1
CAA1B1A2B2A3B3:
Note that ~B~ATis the transpose of ~A~BT,a n d ,b e i n ga1 1 matrix, the transpose
equals itself. Thus a scalar product may be written in these two equivalent forms.
Similarly, the vector product used in ordinary vector theory must be replaced
by something more in keeping with the definition of matrix multiplication. Note
that the vector product
/C65/C66
A2B3ÿA3B2^e1
A3B1ÿA1B3^e2
A1B2ÿA2B1^e3
can be represented by the column matrix
A2B3ÿA3B2
A3B1ÿA1B3
A1B2ÿA2B10
B@1
CA:
This can be split into the product of two matrices
A2B3ÿA3B2
A3B1ÿA1B3
A1B2ÿA2B10
B@1
CA0ÿA2A2
A3 0ÿA1
ÿA2A1 00
B@1
CAB1
B2
B30
B@1
CA
110MATRI/C88 ALGEBRA
or
A2B3ÿA3B2
A3B1ÿA1B3
A1B2ÿA2B10
B@1
CA0ÿB2B2
B3 0ÿB1
ÿB2B1 00
B@1
CAA1
A2
A30
B@1
CA:
Thus the vector product may be represented as the product of a skew-symmetric
matrix and a column matrix. However, this definition only holds for 3 3
matrices.
Similarly, curl Amay be represented in terms of a skew-symmetric matrix
operator, given in Cartesian coordinates by
/C114A0 ÿ/C64=/C64x3/C64=/C64x2
/C64=/C64x3 0 ÿ/C64=/C64x1
ÿ/C64=/C64x2/C64=/C64x1 00
B@1
CAA1
A2
A30
B@1
CA:
In a similar way, we can investigate the triple scalar product and the triple vector
product.
/C84he in/C118erse of a matri/C120
If for a given square matrix ~Athere exists a matrix ~Bsuch that ~A~B~B~A~I,
where ~Iis a unit matrix, then ~Bis called an inverse of matrix ~A.
Example 3.8The matrix
~B35
12
is an inverse of
~A2ÿ5
ÿ13
;
since
~A~B2ÿ5
ÿ1335
12
1001
~I
and
~B~A35122ÿ5
ÿ13
1001
~I:
An invertible matrix has a unique inverse. That is, if ~Band ~Care both inverses
of the matrix ~A, then ~B~C. The proof is simple. Since ~Bis an inverse of ~A,
111THE INVERSE OF A MATRI/C88
~B~A~I. Multiplying both sides on the right by ~Cgives
~/C66~A~C~I~C~C. On the
other hand, ( ~B~A~C~B
~A~C ~B~I~B, so that ~B~C. As a consequence of this
result, we can now speak of theinverse of an invertible matrix. If ~Ais invertible,
then its inverse will be denoted by ~Aÿ1. Thus
~A~Aÿ1~Aÿ1~A~I:
3:19
It is obvious that the inverse of the inverse is the given matrix, that is,
~Aÿ1ÿ1~A:
3:20
It is easy to prove that the inverse of the product is the product of the inverse in
reverse order, that is,
~A~Bÿ1~Bÿ1~Aÿ1:
3:21
To prove (3.21), we start with ~A~Aÿ1~I, with ~Areplaced by ~A~B, that is,
~A~B
~A~Bÿ1~I:
By premultiplying this by ~Aÿ1we get
~B
~A~Bÿ1~Aÿ1:
If we premultiply this by ~Bÿ1, the result follows.
/C65 method for finding ~Aÿ1
The positive power for a square matrix ~Ais defined as ~An~A~A ~A(nfactors)
and ~A0~I, where nis a positive integer. If, in addition, ~Ais invertible, we define
~Aÿn
~Aÿ1n~Aÿ1~Aÿ1 ~Aÿ1
nfactors :
We are now in position to construct the inverse of an invertible matrix ~A:
~Aa11a12 ...a1n
a21a22 ...a2n
.........
an1an2 ...ann0
BBBB@1
CCCCA:
Thea
jkare known. Now let
~Aÿ1a0
11a0
12 ...a0
1n
a0
21a0
22 ...a0
2n
.........
a0
n1a0
n2 ...a0
nn0
BBBB@1
CCCCA:
112MATRI/C88 ALGEBRA
Thea0
jkare required to construct ~Aÿ1. Since ~A~Aÿ1~I, we have
a11a0
11a12a0
12 a1na0
1n1;
a21a0
21a22a0
22 a2na0
2n0;
...
an1a0
n1an2a0
n2 anna0
nn0:
3:22
The solution to the above set of linear algebraic equations (3.22) may be facili-
tated by applying Cramer’s rule. Thus
a0
jkcofactor akj
det ~A:
3:23
From (3.23) it is clear that ~Aÿ1exists if and only if matrix ~Ais non-singular (that
is, det ~A60).
/C83/C121stems of linear equations and the in/C118erse of a matri/C120
As an immediate application, let us apply the concept of an inverse matrix to a
system of nlinear equations in nunknowns
x1;...;xn:
a11x1a12x2 a1nxnb1;
a21x2a22x2 a2nxnb2;
...
an1xnan2xn annxnbn;
in matrix form we have
~A~X~B;
3:24
where
~Aa11a12 ...a1n
a21a22 ...a2n
.........
an1an2 ...ann0
BBBB@1
CCCCA; ~Xx
1
x2
...
xn0
BBBB@1
CCCCA; ~Bb
1
b2
...
bn0
BBBB@1
CCCCA:
We can prove that the above linear system possesses a unique solution given by
~X~A
ÿ1~B:
3:25
The proof is simple. If ~Ais non-singular it has a unique inverse ~Aÿ1. Now pre-
multiplying (3.24) by ~Aÿ1we obtain
~Aÿ1
~A~X ~Aÿ1~B;
113SYSTEMS OF LINEAR EQUATIONS
but
~Aÿ1
~A~X
~Aÿ1~A~X~X
so that
~X~Aÿ1~Bis a solution to
3:24;~A~X~B:
/C67omple/C120 con/C106ugate of a matri/C120
If~A
ajkis an arbitrary matrix whose elements may be complex numbers, the
complex conjugate matrix, denoted by ~A/C42, is also a matrix of the same order,
every element of which is the complex conjugate of the corresponding element of
~A, that is,
A/C42jka/C42jk:
3:26
/C72ermitian con/C106ugation
If~A
ajkis an arbitrary matrix whose elements may be complex numbers, when
the two operations of transposition and complex conjugation are carried out on
~A, the resulting matrix is called the hermitian conjugate (or hermitian adjoint) of
the original matrix ~Aand will be denoted by ~A/C121. We frequently call ~A/C121A-dagger.
The order of the two operations is immaterial:
~A/C121
~AT/C42
~A/C42T:
3:27
In terms of the elements, we have
~A/C121jka/C42kj:
3:27a
It is clear that if ~Ais a matrix of order mn, then ~A/C121is a matrix of order nm.
We can prove that, as in the case of the transpose of a product, the adjoint of the
product is the product of the adjoints in reverse:
~A~B/C121~B/C121~A/C121:
3:28
/C72ermitian/C47anti-hermitian matri/C120
A matrix ~Athat obeys
~A/C121~A
3:29
is called a hermitian matrix. It is very clear the following matrices are hermitian:
1ÿi
i2
;45 2i63i
5ÿ2i 5 ÿ1ÿ2i
6ÿ3iÿ12i 60
B@1
CA;where i
ÿ1p
:
114MATRI/C88 ALGEBRA
Evidently all the elements along the principal diagonal of a hermitian matrix must
be real.
A hermitian matrix is also defined as a matrix whose transpose equals its
complex conjugate:
~AT~A/C42
that is ;akja/C42jk:
3:29a
These two definitions are the same. First note that the elements in the principaldiagonal of a hermitian matrix are always real. Furthermore, any real symmetric
matrix is hermitian, so a real hermitian matrix is a symmetric matrix.
The product of two hermitian matrices is not generally hermitian unless they
commute. This is because of property (3.28): even if ~A
/C121~Aand ~B/C121~B,
~A~B/C1216~A~Bunless the matrices commute.
A matrix ~Athat obeys
~A/C121ÿ ~A
3:30
is called an anti-hermitian (or skew-hermitian) matrix. All the elements along the
principal diagonal must be pure imaginary. An example is
6i 52i63i
ÿ52iÿ8iÿ1ÿ2i
ÿ63i1ÿ2i 00
B@1
CA:
We summarize the three operations on matrices discussed above in Table 3.1.
/C79rthogonal matri/C120 (real)
A matrix ~A
ajkmnsatisfying the relations
~A~AT~In;
3:31a
~AT~A~Im
3:31b
is called an orthogonal matrix. It can be shown that if ~Ais a finite matrix satisfy-
ing both relations (3.31a) and (3.31b), then ~Amust be square, and we have
~A~AT~AT~A~I:
3:32
115ORTHOGONAL MATRI/C88 (REAL)
Table 3.1. Operations on matrices
Operation Matrix element ~A ~B If~B~A
Transposition ~B~ATbijaji mnn m Symmetrica
Complex conjugation ~B~A/C42bija/C42ij mnm n Real
Hermitian conjugation ~B~AT/C42bija/C42ji mnn m Hermitian
aFor square matrices only.
But if ~Ais an infinite matrix, then ~Ais orthogonal if and only if both (3.31a) and
(3.31b) are simultaneously satisfied.
Now taking the determinant of both sides of Eq. (3.32), we have (det ~A21,
or det ~A1. This shows that ~Ais non-singular, and so ~Aÿ1exists.
Premultiplying (3.32) by ~Aÿ1we have
~Aÿ1~AT:
3:33
This is often used as an alternative way of defining an orthogonal matrix.
The elements of an orthogonal matrix are not all independent. To find the
conditions between them, let us first equate the ijth element of both sides of
~A~AT~I; we find that
Xn
k1aikajkij:
3:34a
Similarly, equating the ijth element of both sides of ~AT~A~I, we obtain
Xn
k1akiakjij:
3:34b
Note that either (3.34a) and (3.34b) gives 2 n
n1relations. Thus, for a real
orthogonal matrix of order n, there are only n2ÿn
n1=2n
nÿ1=2 di/C128er-
ent elements.
/C85nitar/C121 matri/C120
A matrix ~U
ujkmnsatisfying the relations
~U~U/C121~In;
3:35a
~U/C121~U~Im
3:35b
is called a unitary matrix. If ~Uis a finite matrix satisfying both (3.35a) and
(3.35b), then ~Umust be a square matrix, and we have
~U~U/C121 ~U/C121~U~I:
3:36
This is the complex generalization of the real orthogonal matrix. The elements of
a unitary matrix may be complex, for example
1
2p1i
i1
is unitary. From the definition (3.35), a real unitary matrix is orthogonal.
Taking the determinant of both sides of (3.36) and noting that
det ~U/C121
det ~U)/C42, we have
det ~U
det ~U/C421o r jdet ~Uj1:
3:37
116MATRI/C88 ALGEBRA
This shows that the determinant of a unitary matrix can be a complex number of
unit magnitude, that is, a number of the form ei, where is a real number. It also
shows that a unitary matrix is non-singular and possesses an inverse.
Premultiplying (3.35a) by ~Uÿ1, we get
~U/C121~Uÿ1:
3:38
This is often used as an alternative way of defining a unitary matrix.
Just as in the case of an orthogonal matrix that is a special (real) case of a
unitary matrix, the elements of a unitary matrix satisfy the following conditions:
Xn
k1uikujk/C42ij;Xn
k1ukiukj/C42ij:
3:39
The product of two unitary matrices is unitary. The reason is as follows. If ~U1
and ~U2are two unitary matrices, then
~U1~U2
~U1~U2/C121~U1~U2
~U/C121
2~U/C121
1 ~U1~U/C121
1~I;
3:40
which shows that U1U2is unitary.
/C82otation matrices
Let us revisit Example 3.5. Our discussion will illustrate the power and usefulnessof matrix methods. We will also see that rotation matrices are orthogonal
matrices. Consider a point Pwith Cartesian coordinates
x
1;x2;x3(see Fig.
3.2). We rotate the coordinate axes about the x3-axis through an angle and
create a new coordinate system, the primed system. The point Pnow has the
coordinates
x0
1;x0
2;x0
3in the primed system. Thus the position vector rof
point Pcan be written as
rX3
i1xi^eiX3
i1x0
i^e0
i:
3:41
117ROTATION MATRICES
Figure 3.2. Coordinate change by rotation.
Taking the dot product of Eq. (3.41) with ^e0
1and using the orthonormal relation
^e0
i^e0
jij(where ijis the Kronecker delta symbol), we obtain x0
1r^e0
1.
Similarly, we have x0
2r^e0
2andx0
3r^e0
3. Combining these results we have
x0
iX3
j1^e0
i^ejxjX3
j1ijxj; i1;2;3:
3:42
The quantities ij^e0
i^ejare called the coecients of transformation. They are
the direction cosines of the primed coordinate axes relative to the unprimed ones
ij^e0
i^ejcos
x0
i;xj; i;j1;2;3:
3:42a
Eq. (3.42) can be written conveniently in the following matrix form
x0
1
x0
2
x0
30
B@1
CA111213
212223
3132330
B@1
CAx1
x2
x30
B@1
CA
3:43a
or
~X0~
~X;
3:43b
where ~X0and ~Xare the column matrices, ~
is called a transformation (or
rotation) matrix; it acts as a linear operator which transforms the vector /C88into
the vector /C880. Strictly speaking, we should describe the matrix ~
as the matrix
representation of the linear operator ^. The concept of linear operator is more
general than that of matrix.
Not all of the nine quantities ijare independent; six relations exist among the
ij, hence only three of them are independent. These six relations are found by
using the fact that the magnitude of the vector must be the same in both systems:
X3
i1
x0
i2X3
i1x2
i:
3:44
With the help of Eq. (3.42), the left hand side of the last equation becomes
X3
i1X3
j1ijxj/C32!X3
k1ikxk/C32!
X3
i1X3
j1X3
k1ijikxjxk;
which, by rearranging the summations, can be rewritten as
X3
k1X3
j1X3
i1ijik/C32!
xjxk:
This last expression will reduce to the right hand side of Eq. (3.43) if and only if
X3
i1ijikjk; j;k1;2;3:
3:45
118MATRI/C88 ALGEBRA
Eq. (3.45) gives six relations among the ij, and is known as the orthogonal
condition.
If the primed coordinates system is generated by a rotation about the x3-axis
through an angle as shown in Fig. 3.2. Then from Example 3.5, we have
x0
1x1cosx2sin;x0
2ÿx1sinx2cos;x0
3x3:
3:46
Thus
11cos; 12sin; 130;
21ÿsin; 22cos; 230;
310; 320; 331:
We can also obtain these elements from Eq. (3.42a). It is obvious that only three
of them are independent, and it is easy to check that they satisfy the condition
given in Eq. (3.45). Now the rotation matrix takes the simple form
~
cossin0
ÿsincos0
00 10
B@1
CA
3:47
and its transpose is
~T
cosÿsin0
sin cos0
00 10
B@1
CA:
Now take the product
~T
~
cossin0
ÿsincos0
00 10
B@1
CAcosÿsin0
sin cos0
00 10
B@1
CA100
0100010
B@1
CA~I;
which shows that the rotation matrix is an orthogonal matrix. In fact, rotation
matrices are orthogonal matrices, not limited to ~
of Eq. (3.47). The proof of
this is easy. Since coordinate transformations are reversible by interchanging oldand new indices, we must have
~
ÿ1ÿ
ij^eold
i^enewj^enewj^eoldiji ~Tÿ
ij:
Hence rotation matrices are orthogonal matrices. It is obvious that the inverse of
an orthogonal matrix is equal to its transpose.
A rotation matrix such as given in Eq. (3.47) is a continuous function of its
argument . So its determinant is also a continuous function of and, in fact, it is
equal to 1 for any . There are matrices of coordinate changes with a determinant
ofÿ1. These correspond to inversion of the coordinate axes about the origin and
119ROTATION MATRICES
change the handedness of the coordinate system. Examples of such parity trans-
formations are
~P1ÿ100
010
0010
B@1
CA; ~P3ÿ100
0ÿ10
00 ÿ10
B@1
CA; ~P2
iI:
They change the signs of an odd number of coordinates of a fixed point rin space
(Fig. 3.3).
What is the advantage of using matrices in describing rotation in space/C63 One of
the advantages is that successive transformations 1 ;2;...;mof the coordinate
axes about the origin are described by successive matrix multiplications as far
as their e/C128ects on the coordinates of a fixed point are concerned:
If~X
1~1~X;~X
2~2~X
1;...;then
~X
m~m~X
mÿ1
~m~mÿ1 ~1~X~R~X
where
~R~m~mÿ1 ~1
is the resultant (or net) rotation matrix for the msuccessive transformations taken
place in the specified manner.
Example 3.9
Consider a rotation of the x1-,x2-axes about the x3-axis by an angle . If this
rotation is followed by a back-rotation of the same angle in the opposite direction,
120MATRI/C88 ALGEBRA
Figure 3.3. Parity transformations of the coordinate system.
that is, by ÿ, we recover the original coordinate system. Thus
~R
ÿ~R
100
010
0010
B@1
CA~Rÿ1
~R
:
Hence
~Rÿ1
~R
ÿcosÿsin0
sin cos0
00 10
B@1
CA~RT
;
which shows that a rotation matrix is an orthogonal matrix.
We would like to make one remark on rotation in space. In the above discus-
sion, we have considered the vector to be fixed and rotated the coordinate axes.
The rotation matrix can be thought of as an operator that, acting on the unprimed
system, transforms it into the primed system. This view is often called the passive
view of rotation. We could equally well keep the coordinate axes fixed and rotatethe vector through an equal angle, but in the opposite direction. Then the rotation
matrix would be thought of as an operator acting on the vector, say /C88, and
changing it into /C88
0. This procedure is called the active view of the rotation.
/C84race of a matri/C120
Recall that the trace of a square matrix ~Ais defined as the sum of all the principal
diagonal elements:
Tr ~AX
kakk:
It can be proved that the trace of the product of a finite number of matrices isinvariant under any cyclic permutation of the matrices. We leave this as home
work.
/C79rthogonal and unitar/C121 transformations
Eq. (3.42) is a linear transformation and it is called an orthogonal transformation,
because the rotation matrix is an orthogonal matrix. One of the properties of anorthogonal transformation is that it preserves the length of a vector. A more
useful linear transformation in physics is the unitary transformation:
~Y~U~X
3:48
in which ~Xand ~Yare column matrices (vectors) of order n1 and ~Uis a unitary
matrix of order nn. One of the properties of a unitary transformation is that it
121TRACE OF A MATRI/C88
preserves the norm of a vector. To see this, premultiplying Eq. (3.48) by
~Y/C121
~X/C121~U/C121) and using the condition ~U/C121~U~I, we obtain
~Y/C121~Y~X/C121~U/C121~U~X~X/C121~X
3:49a
or
Xn
k1yk/C42ykXn
k1xk/C42xk:
3:49b
This shows that the norm of a vector remains invariant under a unitary transfor-
mation. If the matrix ~Uof transformation happens to be real, then ~Uis also an
orthogonal matrix and the transformation (3.48) is an orthogonal transformation,and Eqs. (3.49) reduce to
~Y
T~Y~XT~X;
3:50a
Xn
k1y2
kXn
k1x2k;
3:50b
as we expected.
/C83imilarit/C121 transformation
We now consider a di/C128erent linear transformation, the similarity transformation
that, we shall see later, is very useful in diagonalization of a matrix. To get the
idea about similarity transformations, we consider vectors rand/C82in a particular
basis, the coordinate system Ox1x2x3, which are connected by a square matrix ~A:
/C82~Ar:
3:51a
Now rotating the coordinate system about the origin Owe obtain a new system
Ox0
1x0
2x0
3(a new basis). The vectors rand /C82have not been a/C128ected by this
rotation. Their components, however, will have di/C128erent values in the new system,and we now have
/C82
0~A0r0:
3:51b
The matrix ~A0in the new (primed) system is called similar to the matrix ~Ain the
old (unprimed) system, since they perform same function. Then what is the rela-tionship between matrices ~Aand ~A
0/C63 This information is given in the form
of coordinate transformation. We learned in the previous section that the com-ponents of a vector in the primed and unprimed systems are connected by a
matrix equation similar to Eq. (3.43). Thus we have
r~Sr
0and /C82~S/C820;
122MATRI/C88 ALGEBRA
where ~Sis a non-singular matrix, the transition matrix from the new coordinate
system to the old system. With these, Eq. (3.51a) becomes
~S/C820~A~Sr0
or
/C820~Sÿ1~A~Sr0:
Combining this with Eq. (3.51) gives
~A0~Sÿ1~A~S;
3:52
where ~A0and ~Aare similar matrices. Eq. (3.52) is called a similarity transforma-
tion.
Generalization of this idea to n-dimensional vectors is straightforward. In this
case, we take rand/C82as two n-dimensional vectors in a particular basis, having
their coordinates connected by the matrix ~A(annsquare matrix) through Eq.
(3.51a). In another basis they are connected by Eq. (3.51b). The relationship
between ~Aand ~A0is given by Eq. (3.52). The transformation of ~Ainto ~Sÿ1~A~S
is called a similarity transformation.
All identities involving vectors and matrices will remain invariant under a
similarity transformation since this arises only in connection with a change inbasis. That this is so can be seen in the following two simple examples.
Example 3.10
Given the matrix equation ~A~B~C, and the matrices ~A,~B,~Csubjected to the
same similarity transformation, show that the matrix equation is invariant.
Solution: Since the three matrices are all subjected to the same similarity trans-
formation, we have
~A
0~S~A~Sÿ1; ~B0~S~B~Sÿ1; ~C0~S~C~Sÿ1
and it follows that
~A0~B0
~S~A~Sÿ1
~S~B~Sÿ1 ~S~A~I~B~Sÿ1~S~A~B~Sÿ1~S~C~Sÿ1~C0:
Example 3.11
Show that the relation ~A/C82~Bris invariant under a similarity transformation.
Solution: Since matrices ~Aand ~Bare subjected to the same similarity transfor-
mation, we have
~A0~S~A~Sÿ1; ~B0~S~B~Sÿ1
we also have
/C820~S/C82; r0~Sr:
123SIMILARITY TRANSFORMATION
Then
~A0/C820
~S~A~Sÿ1
S/C82 ~S~A/C82and ~B0r0
~S~B~Sÿ1
Sr ~S~Br
thus
~A0/C820~B0r0:
We shall see in the following section that similarity transformations are very
useful in diagonalization of a matrix, and that two similar matrices have the same
eigenvalues.
/C84he matri/C120 eigen/C118alue problem
As we saw in preceding sections, a linear transformation generally carries a vector
/C88
x1;x2;...;xninto a vector /C89
y1;y2;...;yn:However, there may exist
certain non-zero vectors for which ~A/C88is just /C88multiplied by a constant
~A/C88/C88:
3:53
That is, the transformation represented by the matrix (operator) ~Ajust multiplies
the vector /C88by a number . Such a vector is called an eigenvector of the matrix ~A,
and is called an eigenvalue (German: eigenwert ) or characteristic value of the
matrix ~A. The eigenvector is said to ‘belong’ (or correspond) to the eigenvalue.
And the set of the eigenvalues of a matrix (an operator) is called its eigenvaluespectrum.
The problem of finding the eigenvalues and eigenvectors of a matrix is called an
eigenvalue problem. We encounter problems of this type in all branches of
physics, classical or quantum. Various methods for the approximate determina-
tion of eigenvalues have been developed, but here we only discuss the
fundamental ideas and concepts that are important for the topics discussed inthis book.
There are two parts to every eigenvalue problem. First, we compute the eigen-
value , given the matrix ~A. Then, we compute an eigenvector Xfor each
previously computed eigenvalue .
/C68etermination of eigenvalues and eigenvectors
We shall now demonstrate that any square matrix of order nhas at least 1 and at
most ndistinct (real or complex) eigenvalues. To this purpose, let us rewrite the
system of Eq. (3.53) as
~Aÿ~IX0:
3:54
This matrix equation really consists of nhomogeneous linear equations in the n
unknown elements x
iofX:
124MATRI/C88 ALGEBRA
a11ÿ
x1a12x2 a1nxn0
a21x1a22ÿ
x2 a2nxn0
...
an1x1an2x2 annÿ
xn09
>>>>>=
>>>>>;
3:55
In order to have a non-zero solution, we recall that the determinant of the coe-
cients must be zero; that is,
det
~Aÿ~Ia
11ÿ a12 a1n
a21 a22ÿ a2n
.........
an1 an2 annÿ/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C120:
3:56
The expansion of the determinant gives an nth order polynomial equation in ,
and we write this as
c
0nc1nÿ1c2nÿ2 cnÿ1cn0;
3:57
where the coecients ciare functions of the elements ajkof ~A. Eq. (3.56) or (3.57)
is called the characteristic equation corresponding to the matrix ~A. We have thus
obtained a very important result: the eigenvalues of a square matrix ~Aare the
roots of the corresponding characteristic equation (3.56) or (3.57).
Some of the coecients cican be readily determined; by an inspection of Eq.
(3.56) we find
c0
ÿ 1n;c1
ÿ 1nÿ1
a11a22 ann;cndet ~A:
3:58
Now let us rewrite the characteristic polynomial in terms of its nroots
1;2;...;n
c0nc1nÿ1c2nÿ2 cnÿ1cn1ÿ
2ÿ
nÿ
;
then we see that
c1
ÿ 1nÿ1
12 n;cn12n:
3:59
Comparing this with Eq. (3.58), we obtain the following two important results on
the eigenvalues of a matrix:
(1) The sum of the eigenvalues equals the trace (spur) of the matrix:
12 na11a22 annTr ~A:
3:60
(2) The product of the eigenvalues equals the determinant of the matrix:
12ndet ~A:
3:61
125THE MATRI/C88 EIGENVALUE PROBLEM
Once the eigenvalues have been found, corresponding eigenvectors can be
found from the system (3.55). Since the system is homogeneous, if Xis an
eigenvector of ~A, then kX, where kis any constant (not zero), is also an eigen-
vector of ~Acorresponding to the same eigenvalue. It is very easy to show this.
Since ~AXX, multiplying by an arbitrary constant kwill give k~AXkX.
Now k~A~Ak (every matrix commutes with a scalar), so we have
~A
kX
kX;showing that kXis also an eigenvector of ~Awith the same
eigenvalue . But kXis linearly dependent on X, and if we were to count all
such eigenvectors separately, we would have an infinite number of them. Such
eigenvectors are therefore not counted separately.
A matrix of order ndoes not necessarily have nlinearly independent
eigenvectors; some of them may be repeated. (This will happen when the char-acteristic polynomial has two or more identical roots.) If an eigenvalue occurs m
times, mis called the multiplicity of the eigenvalue. The matrix has at most m
linearly independent eigenvectors all corresponding to the same eigenvalue. Such
linearly independent eigenvectors having the same eigenvalue are said to be degen-
erate eigenvectors; in this case, m-fold degenerate. We will deal only with those
matrices that have nlinearly independent eigenvectors and they are diagonalizable
matrices.
Example 3.12
Find ( a) the eigenvalues and ( b) the eigenvectors of the matrix
~A54
12
:
Solution: (a) The eigenvalues: The characteristic equation is
det
~Aÿ~I5ÿ 4
12 ÿ/C12/C12/C12/C12/C12/C12/C12/C12
2ÿ760
which has two roots
16 and 21:
(b) The eigenvectors: For 1the system (3.55) assumes the form
ÿx14x20;
x1ÿ4x20:
Thus x14x2, and
X14
1
126MATRI/C88 ALGEBRA
is an eigenvector of ~Acorresponding to 16. In the same way we find the
eigenvector corresponding to 21:
X21
ÿ1
:
Example 3.13
If~Ais a non-singular matrix, show that the eigenvalues of ~Aÿ1are the reciprocals
of those of ~Aand every eigenvector of ~Ais also an eigenvector of ~Aÿ1.
Solution: Letbe an eigenvalue of ~Acorresponding to the eigenvector X,s o
that
~AXX:
Since ~Aÿ1exists, multiply the above equation from the left by ~Aÿ1
~Aÿ1~AX~Aÿ1X/C41X~Aÿ1X:
Since ~Ais non-singular, must be non-zero. Now dividing the above equation by
, we have
~Aÿ1X
1=X:
Since this is true for every value of ~A, the results follows.
Example 3.14Show that all the eigenvalues of a unitary matrix have unit magnitude.
Solution: Let ~Ube a unitary matrix and Xan eigenvector of ~Uwith the eigen-
value , so that
~UXX:
Taking the hermitian conjugate of both sides, we have
X
/C121~U/C121/C42X/C121:
Multiplying the first equation from the left by the second equation, we obtain
X/C121~U/C121~UX/C42X/C121X:
Since ~Uis unitary, ~U/C121~U/C61 ~I, so that the last equation reduces to
X/C121X
jj2ÿ10:
Now X/C121Xis the square of the norm of Xand hence cannot vanish unless Xis a
null vector and so we must have jj21o rjj1;proving the desired result.
127THE MATRI/C88 EIGENVALUE PROBLEM
Example 3.15
Show that similar matrices have the same characteristic polynomial and hence the
same eigenvalues. (Another way of stating this is to say that the eigenvalues of a
matrix are invariant under similarity transformations.)
Solution: Let ~Aand ~Bbe similar matrices. Thus there exists a third matrix ~S
such that ~B~Sÿ1~A~S. Substituting this into the characteristic polynomial of
matrix ~Bwhich is j~Bÿ~Ij, we obtain
j~BÿIjj ~Sÿ1~A~Sÿ~Ijj ~Sÿ1
~Aÿ~I~Sj:
Using the properties of determinants, we have
j~Sÿ1
~Aÿ~I~Sjj ~Sÿ1jj~Aÿ~Ijj~Sj:
Then it follows that
j~Bÿ~Ijj ~Sÿ1
~Aÿ~I~Sjj ~Sÿ1jj~Aÿ~Ijj~Sjj ~Aÿ~Ij;
which shows that the characteristic polynomials of ~Aand ~Bare the same; their
eigenvalues will also be identical.
/C69igen/C118alues and eigen/C118ectors of hermitian matrices
In quantum mechanics complex variables are unavoidable because of the form of
the Schro /C200dinger equation. And all quantum observables are represented by her-
mitian operators. So physicists are almost always dealing with adjoint matrices,hermitian matrices, and unitary matrices. Why are physicists interested in hermi-
tian matrices/C63 Because they have the following properties: (1) the eigenvalues of a
hermitian matrix are real, and (2) its eigenvectors corresponding to distinct eigen-values are orthogonal, so they can be used as basis vectors. We now proceed to
prove these important properties.
(1) the eigenvalues of a hermitian matrix are real.
Let ~Hbe a hermitian matrix and Xa non-trivial eigenvector corresponding to the
eigenvalue , so that
~HXX:
3:62
Taking the hermitian conjugate and note that ~H
/C121 ~H, we have
X/C121~H/C42X/C121:
3:63
Multiplying (3.62) from the left by X/C121, and (3.63) from the right by X/C121, and then
subtracting, we get
ÿ/C42X/C121X0:
3:64
Now, since X/C121Xcannot be zero, it follows that /C42, or that is real.
128MATRI/C88 ALGEBRA
(2) The eigenvectors corresponding to distinct eigenvalues are orthogonal.
LetX1andX2be eigenvectors of ~Hcorresponding to the distinct eigenvalues 1
and 2, respectively, so that
~HX 11X1;
3:65
~HX 22X2:
3:66
Taking the hermitian conjugate of (3.66) and noting that /C42, we have
X/C121
2~H2X/C121
2:
3:67
Multiplying (3.65) from the left by X/C121
2and (3.67) from the right by X1, then
subtracting, we obtain
1ÿ2X/C121
2X10:
3:68
Since 12, it follows that X/C121
2X10 or that X1andX2are orthogonal.
IfXis an eigenvector of ~H, any multiple of X,X, is also an eigenvector of ~H.
Thus we can normalize the eigenvector Xwith a properly chosen scalar . This
means that the eigenvectors of ~Hcorresponding to distinct eigenvalues are ortho-
normal. Just as the three orthogonal unit coordinate vectors ^e1;^e2;and ^e3form
the basis of a three-dimensional vector space, the orthonormal eigenvectors of ~H
may serve as a basis for a function space.
/C68iagonali/C122ation of a matri/C120
Let ~A
aijbe a square matrix of order n, which has nlinearly independent
eigenvectors Xiwith the corresponding eigenvalues i:~AXiiXi. If we denote
the eigenvectors Xiby column vectors with elements x1i;x2i;...;xni, then the
eigenvalue equation can be written in matrix form:
a11a12a1n
a21a22a2n
.........
an1an2ann0
BBBBB@1
CCCCCAx
1i
x2i
...
xni0
BBBBB@1
CCCCCA
ix1i
x2i
...
xni0
BBBBB@1
CCCCCA:
3:69
From the above matrix equation we obtain
X
n
k1ajkxkiixji:
3:69b
Now we want to diagonalize ~A. To this purpose, we can follow these steps. We
first form a matrix ~Sof order nnwhose columns are the vector Xi, that is,
129DIAGONALI/C90ATION OF A MATRI/C88
~Sx11x1ix1n
x21x2ix2n
.........
xn1xnixnn0
BBBBB@1
CCCCCA;
~S
ijxij:
3:70
Since the vectors Xiare linear independent, ~Sis non-singular and ~Sÿ1exists. We
then form a matrix ~Sÿ1~A~S; this is a diagonal matrix whose diagonal elements are
the eigenvalues of ~A.
To show this, we first define a diagonal matrix ~Bwhose diagonal elements are i
i1;2;...;n:
~B1
2
...
n0
BBBBB@1
CCCCCA;
3:71
and we then demonstrate that
~S
ÿ1~A~S~B:
3:72a
Eq. (3.72a) can be rewritten by multiplying it from the left by ~Sas
~A~S~S~B:
3:72b
Consider the left hand side first. Taking the jith element, we obtain
~A~SjiXn
k1
~Ajk
~SkiXn
k1ajkxki:
3:73a
Similarly, the jith element of the right hand side is
~S~BjiXn
k1
~Sjk
~BkiXn
k1xjkikiixji:
3:73b
Eqs. (3.73a) and (3.73b) clearly show the validity of Eq. (3.72a).
It is important to note that the matrix ~Sthat is able to diagonalize matrix ~Ais
not unique. This is because we could arrange the eigenvectors X1;X2;...;Xnin
any order to construct ~S.
We summarize the procedure for diagonalizing a diagonalizable nnmatrix ~A:
Step 1. Find nlinearly independent eigenvectors of ~A;X1;X2;...;Xn.
Step 2. Form the matrix ~Shaving X1;X2;...;Xnas its column vectors.
Step 3. Find the inverse of ~S,~Sÿ1.
Step 4. The matrix ~Sÿ1~A~Swill then be diagonal with 1;2;...;nas its succes-
sive diagonal elements, where iis the eigenvalue corresponding to Xi.
130MATRI/C88 ALGEBRA
Example 3.16
Find a matrix ~Sthat diagonalizes
~A3ÿ20
ÿ23 0
00 50
B@1
CA:
Solution: We have first to find the eigenvalues and the corresponding eigen-
vectors of matrix ~A. The characteristic equation of ~Ais
3ÿÿ20
ÿ23 ÿ 0
00 5 ÿ/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
ÿ1
ÿ5
20;
so that the eigenvalues of ~Aare1 and 5.
By definition
~Xx1
x2
x30
B@1
CA
is an eigenvector of ~Acorresponding to if and only if ~Xis a non-trivial solution
of (~Iÿ~A~X0, that is, of
ÿ32 0
2 ÿ30
00 ÿ50
B@1
CAx1
x2
x30
B@1
CA0
000
B@1
CA:
If5 the above equation becomes
220
2200000
B@1
CAx
1
x2
x30
B@1
CA0
000
B@1
CAor2x
12x20x3
2x12x20x3
0x10x20x30
B@1
CA0
000
B@1
CA:
Solving this system yields
x
1ÿs;x2s;x3t;
where sand tare arbitrary values. Thus the eigenvectors of ~Acorresponding to
5 are the non-zero vectors of the form
~Xÿs
s
t0
B@1
CAÿs
s
00
B@1
CA0
0
t0
B@1
CAsÿ1
1
00
B@1
CAt0
0
10
B@1
CA:
131DIAGONALI/C90ATION OF A MATRI/C88
Since
ÿ1
1
00
B@1
CAand0
0
10
B@1
CA
are linearly independent, they are the eigenvectors corresponding to 5.
For 1, we have
ÿ22 0
2ÿ20
00 ÿ40
B@1
CAx1
x2
x30
B@1
CA0
000
B@1
CAorÿ2x
12x20x3
2x1ÿ2x20x3
0x10x2ÿ4x30
B@1
CA0
000
B@1
CA:
Solving this system yields
x
1t;x2t;x30;
where tis arbitrary. Thus the eigenvectors corresponding to 1 are non-zero
vectors of the form
~Xt
t
00
B@1
CAt1
1
00
B@1
CA:
It is easy to check that the three eigenvectors
~X1ÿ1
1
00
B@1
CA; ~X20
0
10
B@1
CA; ~X31
1
00
B@1
CA;
are linearly independent. We now form the matrix ~Sthat has ~X1,~X2,a n d ~X3as its
column vectors:
~Sÿ101
101
0100
B@1
CA:
The matrix ~Sÿ1~A~Sis diagonal:
~Sÿ1~A~Sÿ1=21 =20
00 1
1=21 =200
B@1
CA3ÿ20
ÿ23 0
00 50
B@1
CAÿ101
101
0100
B@1
CA500
050
0010
B@1
CA:
There is no preferred order for the columns of ~S. If had we written
~Sÿ110
1100010
B@1
CA
132MATRI/C88 ALGEBRA
then we would have obtained (verify)
~Sÿ1~A~S500
010
0010
B@1
CA:
Example 3.17
Show that the matrix
~Aÿ32
ÿ21
is not diagonalizable.
Solution: The characteristic equation of ~Ais
3ÿ2
2 ÿ1/C12/C12/C12/C12/C12/C12/C12/C12
120:
Thus ÿ1 the only eigenvalue of ~A; the eigenvectors corresponding to ÿ1
are the solutions of
3ÿ2
2 ÿ1x1
x2/C32!
0
0
/C412ÿ2
2ÿ2x1
x2/C32!
00
from which we have
2x
1ÿ2x20;
2x1ÿ2x20:
The solutions to this system are x1t;x2t; hence the eigenvectors are of the
form
tt
t11
:
Adoes not have two linearly independent eigenvectors, and is therefore not
diagonalizable.
/C69igen/C118ectors of commuting matrices
There is a theorem on eigenvectors of commuting matrices that is of great impor-
tance in matrix algebra as well as in quantum mechanics. This theorem states that:
Two commuting matrices possess a common set of eigenvectors.
133EIGENVECTORS OF COMMUTING MATRICES
We now proceed to prove it. Let ~Aand ~Bbe two square matrices, each of order
n, which commute with each other, that is,
~A~Bÿ~B~A ~A;~B0:
First, let be an eigenvalue of ~Awith multiplicity 1, corresponding to the eigen-
vector X, so that
~AXX:
3:74
Multiplying both sides from the left by ~B
~B~AX~BX:
Because ~B~A~A~B, we have
~A
~BX
~BX:
Now ~Bis an nnmatrix and Xis an n1 vector; hence ~BXis also an n1
vector. The above equation shows that ~BXis also an eigenvector of ~Awith the
eigenvalue . Now Xis a non-degenerate eigenvector of ~A, any other vector which
is an eigenvector of ~Awith the same eigenvalue as that of Xmust be multiple of X.
Accordingly
~BX/C22X;
where /C22is a scalar. Thus we have proved that:
If two matrices commute, every non-degenerate eigenvector of
one is also an eigenvector of the other, and vice versa.
Next, let be an eigenvalue of ~Awith multiplicity k.S o ~Ahasklinearly inde-
pendent eigenvectors, say X1;X2;...;Xk, each corresponding to :
~AXiXi;1ik:
Multiplying both sides from the left by ~B, we obtain
~A
~BXi
~BXi;
which shows again that ~BXis also an eigenvector of ~Awith the same eigenvalue .
/C67a/C121le/C121/C177/C72amilton theorem
The Cayley–Hamilton theorem is useful in evaluating the inverse of a square
matrix. We now introduce it here. As given by Eq. (3.57), the characteristic
equation associated with a square matrix ~Aof order nmay be written as a poly-
nomial
f
Xn
i0cinÿi0;
134MATRI/C88 ALGEBRA
where are the eigenvalues given by the characteristic determinant (3.56). If we
replace inf
by the matrix ~Aso that
f
~AXn
i0ci~Anÿi:
The Cayley–Hamilton theorem says that
f
~A0o rXn
i0ci~Anÿi0;
3:75
that is, the matrix ~Asatisfies its characteristic equation.
We now formally multiply Eq. (3.75) by ~Aÿ1so that we obtain
~Aÿ1f
~Ac0~Anÿ1c1~Anÿ2 cnÿ1~Icn~Aÿ10:
Solving for ~Aÿ1gives
~Aÿ1ÿ1
cnXnÿ1
i0ci~Anÿ1ÿi"#
;
3:76
we can use this to find ~Aÿ1(Problem 3.28).
Moment of inertia matri/C120
We shall see that physically diagonalization amounts to a simplification of the
problem by a better choice of variable or coordinate system. As an illustrative
example, we consider the moment of inertia matrix ~Iof a rotating rigid body (see
Fig. 3.4). A rigid body can be considered to be a many-particle system, with the
135MOMENT OF INERTIA MATRI/C88
Figure 3.4. A rotating rigid body.
distance between any particle pair constant at all times. Then its angular momen-
tum about the origin Oof the coordinate system is
LX
mr/C118X
mr
xr
where the subscript refers to mass malocated at r
x1;x2;x3, andxthe
angular velocity of the rigid body.
Expanding the vector triple product by using the vector identity
/C65
/C66/C67/C66
/C65/C67ÿ/C67
/C65/C66;
we obtain
LX
m/C98r2
xÿr
rx/C99:
In terms of the components of the vectors randx, the ith component of Liis
LiX
m/C33iX3
k1x2
;kÿx;iX3
j1x;j/C33j"#
X
j/C33jX
mijX
kx2;kÿx;ix;j"#
X
jIij/C33j
or
~L~I~/C33:
Both ~Land ~/C33are three-dimensional column vectors, while ~Iis a 3 3 matrix and
is called the moment inertia matrix.
In general, the angular momentum vector Lof a rigid body is not always
parallel to its angular velocity xand ~Iis not a diagonal matrix. But we can orient
the coordinate axes in space so that all the non-diagonal elements Iij
i6j
vanish. Such special directions are called the principal axes of inertia. If the
angular velocity is along one of these principal axes, the angular momentum
and the angular velocity will be parallel.
In many simple cases, especially when symmetry is present, the principal axes of
inertia can be found by inspection.
Normal modes of /C118ibrations
Another good illustrative example of the application of matrix methods in classi-cal physics is the longitudinal vibrations of a classical model of a carbon dioxide
molecule that has the chemical structure O–C–O. In particular, it provides a good
example of the eigenvalues and eigenvectors of an asymmetric real matrix.
136MATRI/C88 ALGEBRA
We can regard a carbon dioxide molecule as equivalent to a set of three par-
ticles jointed by elastic springs (Fig. 3.5). Clearly the system will vibrate in some
manner in response to an external force. For simplicity we shall consider only
longitudinal vibrations, and the interactions of the oxygen molecules with oneanother will be neglected, so we consider only nearest neighbor interaction. The
Lagrangian function Lfor the system is
L
1
2m
_x2
1_x231
2M_x2
2ÿ1
2k
x2ÿx12ÿ12k
x3ÿx22;
substituting this into Lagrange’s equations
d
dt/C64L
/C64_xi
ÿ/C64L
/C64xi0
i1;2;3;
we find the equations of motion to be
/C127x1ÿk
m
x1ÿx2ÿk
mx1k
mx2;
/C127x2ÿk
M
x2ÿx1ÿk
M
x2ÿx3k
Mx1ÿ2k
Mx2k
Mx3;
/C127x3k
mx2ÿk
mx3;
where the dots denote time derivatives. If we define
~Xx1
x2
x30
B@1
CA; ~Aÿk
mk
m0
ÿk
Mÿ2k
Mk
M
0k
mÿk
m0
BBBBBBB@1
CCCCCCCA
and, furthermore, if we define the derivative of a matrix to be the matrix obtained
by di/C128erentiating each matrix element, then the above system of di/C128erential equa-
tions can be written as
/C127~X~A~X:
137NORMAL MODES OF VIBRATIONS
Figure 3.5. A linear symmetrical carbon dioxide molecule.
This matrix equation is reminiscent of the single di/C128erential equation /C127xax, with
aa constant. The latter always has an exponential solution. This suggests that we
try
~X~Ce/C33t;
where /C33is to be determined and
~CC1
C2
C30
B@1
CA
is an as yet unknown constant matrix. Substituting this into the above matrix
equation, we obtain a matrix-eigenvalue equation
~A~C/C332~C
or
ÿk
mk
m0
ÿk
Mÿ2k
Mk
M
0k
mÿk
m0
BBBBBBBBB@1
CCCCCCCCCAC
1
C2
C30
B@1
CA/C332C1
C2
C30
B@1
CA:
3:77
Thus the possible values of /C33are the square roots of the eigenvalues of the
asymmetric matrix ~Awith the corresponding solutions being the eigenvectors of
the matrix ~A. The secular equation is
ÿk
mÿ/C332 k
m0
ÿk
Mÿ2k
Mÿ/C332 k
M
0k
mÿk
mÿ/C332/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C120:
This leads to
/C33
2ÿ/C332k
m
ÿ/C332k
m2k
M
0:
The eigenvalues are
/C3320;k
m;andk
m2k
M;
138MATRI/C88 ALGEBRA
all real. The corresponding eigenvectors are determined by substituting the eigen-
values back into Eq. (3.77) one eigenvalue at a time:
(1) Setting /C3320 in Eq. (3.77) we find that C1C2C3. Thus this mode is
not an oscillation at all, but is a pure translation of the system as a whole, no
relative motion of the masses (Fig. 3.6( a)).
(2) Setting /C332k=min Eq. (3.77), we find C20 and C3ÿC1. Thus the
center mass Mis stationary while the outer masses vibrate in opposite
directions with the same amplitude (Fig. 3.6( b)).
(3) Setting /C332k=m2k=Min Eq. (3.77), we find C1C3, and
C2ÿ2C1
m=M. In this mode the two outer masses vibrate in
unison and the center mass vibrates oppositely with di/C128erent amplitude
(Fig. 3.6( c)).
/C68irect product of matrices
Sometimes the direct product of matrices is useful. Given an mmmatrix ~A
and an nnmatrix ~B, the direct product of ~Aand ~Bis an mnmnmatrix,
defined by
~C~A/C10~Ba11~Ba 12~B a1m~B
a21~Ba 22~B a2m~B
.........
am1~Ba m2~Bamm~B0
BBBBB@1
CCCCCA:
For example, if
~Aa
11a12
a21a22/C32!
; ~Bb11b12
b21b22/C32!
;
then
139DIRECT PRODUCT OF MATRICES
Figure 3.6. Longitudinal vibrations of a carbon dioxide molecule.
~A/C10~Ba11~Ba 12~B
a21~Ba 22~B/C32!
a11b11a11b12a12b11a12b12
a11b21a11b22a12b21a12b22
a21b11a21b12a22b11a22b12
a21b21a21b22a22b21a22b220
BBBB@1
CCCCA:
Problems
3.1 For the pairs ~Aand ~Bgiven below, find ~A~B,~A~B,a n d ~A
2:
~A12
34
; ~B5678
:
3.2 Show that an n-rowed diagonal matrix ~D
~Dk0 0
0k 0
.........
k0
BBBB@1
CCCCA
commutes with any n-rowed square matrix ~A:~A~D~D~Ak~A.
3.3 If ~A,~B,a n d ~Care any matrices such that the addition ~B~Cand the
products ~A~Band ~A~Care defined, show that ~A(~B~C ~A~B/C43~A~C. That
is, that matrix multiplication is distributive.
3.4 Given
~A010
101
0100
B@1
CA; ~B100
010
0010
B@1
CA; ~C10 0
00 0
00 ÿ10
B@1
CA;
show that /C91 ~A;~B0, and /C91 ~B;~C0, but that ~Adoes not commute with ~C.
3.5 Prove that ( ~A~B
T~AT~BT.
3.6 Given
~A2ÿ3
04
; ~Bÿ52
21
;and ~C01 ÿ2
30 4
/C58
(a) Find 2 ~Aÿ4~B,2 ( ~Aÿ2~B)
(b) Find ~AT;~BT;
~BTT
(c) Find ~CT;
~CTT
(d)I s ~A~Cdefined/C63
(e)I s ~C~CTdefined/C63
(f)I s ~A~ATsymmetric/C63
140MATRI/C88 ALGEBRA
(g)I s ~Aÿ~ATantisymmetric/C63
3.7 Show that the matrix
~A140
250
3600
B@1
CA
is not invertible.
3.8 Show that if ~Aand ~Bare invertible matrices of the same order, then ~A~Bis
invertible.
3.9 Given
~A123
2531080
B@1
CA;
find ~A
ÿ1and check the answer by direct multiplication.
3.10 Prove that if ~Ais a non-singular matrix, then det( ~Aÿ11=det
~A).
3.11 If ~Ais an invertible nnmatrix, show that ~AX0 has only the trivial
solution.
3.12 Show, by computing a matrix inverse, that the solution to the following
system is x14,x21:
x1ÿx23;
x1x25:
3.13 Solve the system ~AX~Bif
~A100
020
0010
B@1
CA; ~B1
2
30
B@1
CA:
3.14 Given matrix ~A, find A/C42, AT, and A/C121, where
~A23i1ÿi 5i ÿ3
1i6ÿi13iÿ1ÿ2i
5ÿ6i30 ÿ40
B@1
CA:
3.15 Show that:
(a) The matrix ~A~A/C121, where ~Ais any matrix, is hermitian.
(b)
~A~B/C121~B/C121~A/C121:
(c)I f ~A;~Bare hermitian, then ~A~B~B~Ais hermitian.
(d)I f ~Aand ~Bare hermitian, then i
~A~Bÿ~B~Ais hermitian.
3.16 Obtain the most general orthogonal matrix of order 2.
/C91Hint: use relations (3.34a) and (3.34b)./C93
3.17. Obtain the most general unitary matrix of order 2.
141PROBLEMS
3.18 If ~A~B0, show that one of these matrices must have zero determinant.
3.19 Given the Pauli spin matrices (which are very important in quantum
mechanics)
101
10
; 20ÿi
i0
; 3100ÿ1
;
(note that the subscripts x;y, and zare sometimes used instead of 1, 2, and
3). Show that
(a) they are hermitian,
(b)
2
i~I;i1;2;3
(c) as a result of ( a) and ( b) they are also unitary, and
(d)/C911;22I3et cycl .
Find the inverses of 1;2;3:
3.20 Use a rotation matrix to show that
sin
12sin1cos2sin2cos1:
3.21 Show that: Tr ~A~BTr ~B~Aand Tr ~A~B~CTr ~B~C~ATr ~C~A~B:
3.22 Show that: ( a) the trace and ( b) the commutation relation between two
matrices are invariant under similarity transformations.
3.23 Determine the eigenvalues and eigenvectors of the matrix
~Aab
ÿba
:
Given
~A57 ÿ5
04 ÿ1
28 ÿ30
B@1
CA;
find a matrix ~Sthat diagonalizes ~A, and show that ~Sÿ1~A~Sis diagonal.
3.25 If ~Aand ~Bare square matrices of the same order, then
det( ~A~Bdet
~Adet
~B:Verify this theorem if
~A2ÿ1
32
; ~B72
ÿ34
:
3.26 Find a common set of eigenvectors for the two matrices
~Aÿ1
6p
2p
6p
03p
2p
3p
ÿ20
B@1
CA; ~B10
6p
ÿ
2p
6p
93p
ÿ2p
3p
110
B@1
CA:
3.27 Show that two hermitian matrices can be made diagonal if and only if they
commute.
142MATRI/C88 ALGEBRA
3.28 Show the validity of the Cayley–Hamilton theorem by applying it to the
matrix
~A54
12
;
then use the Cayley–Hamilton theorem to find the inverse of the matrix ~A.
3.29 Given
~A01
10
; ~B0ÿi
i0
;
find the direct product of these matrices, and show that it does not com-
mute.
143PROBLEMS
4
/C70ourier series and integrals
Fourier series are infinite series of sines and cosines which are capable of repre-
senting almost any periodic function whether continuous or not. Periodic func-
tions that occur in physics and engineering problems are often very complicated
and it is desirable to represent them in terms of simple periodic functions.
Therefore the study of Fourier series is a matter of great practical importance
for physicists and engineers.
The first part of this chapter deals with Fourier series. Basic concepts, facts, and
techniques in connection with Fourier series will be introduced and developed,
along with illustrative examples. They are followed by Fourier integrals and
Fourier transforms.
Periodic functions
If function f
xis defined for all xand there is some positive constant Psuch that
f
xPf
x
4:1
then we say that f
xis periodic with a period P(Fig. 4.1). From Eq. (4.1) we also
144Figure 4.1. A general periodic function.
have, for all xand any integer n,
f
xnPf
x:
That is, every periodic function has arbitrarily large periods and contains arbi-
trarily large numbers in its domain. We call Pthe fundamental (or least) period,
or simply the period.
A periodic function need not be defined for all values of its independent vari-
able. For example, tan xis undefined for the values x
=2n. But tan xis a
periodic function in its domain of definition, with as its fundamental period:
tan(xtanx.
Example 4.1
(a) The period of sin xis 2, since sin( x2, sin(x4;sin
x6;...are all
equal to sin x, but 2 is the least value of P. And, as shown in Fig. 4.2, the period
of sin nxis 2=n, where nis a positive integer.
(b) A constant function has any positive number as a period. Since f
x
c(const.) is defined for all real x, then, for every positive number P,
f
xPcf
x. Hence Pis a period of f. Furthermore, fhas no fundamental
period.
(c)
f
x/C75for 2 nx
2n1
ÿ/C75for
2n1x<
2n2(
n0;1;2;3;...
is periodic of period 2 (Fig. 4.3).
145PERIODIC FUNCTIONS
Figure 4.2. Sine functions.
Figure 4.3. A square wave function.
Fourier series/C59 /C69uler/C177Fourier formulas
If the general periodic function f
xis defined in an interval ÿx, the
Fourier series of f
xin /C91ÿ; /C93 is defined to be a trigonometric series of the form
f
x1
2a0a1cosxa2cos 2x ancosnx
b1sinxb2sin 2x bnsinnx ;
4:2
where the numbers a0;a1;a2;...;b1;b2;b3;...are called the Fourier coecients of
f
xinÿ; . If this expansion is possible, then our power to solve physical
problems is greatly increased, since the sine and cosine terms in the series can be
handled individually without diculty. Joseph Fourier (1768–1830), a French
mathematician, undertook the systematic study of such expansions. In 1807 he
submitted a paper (on heat conduction) to the Academy of Sciences in Paris and
claimed that every function defined on the closed interval ÿ; could be repre-
sented in the form of a series given by Eq. (4.2); he also provided integral formulasfor the coecients a
nandbn. These integral formulas had been obtained earlier by
Clairaut in 1757 and by Euler in 1777. However, Fourier opened a new avenue byclaiming that these integral formulas are well defined even for very arbitrary
functions and that the resulting coecients are identical for di/C128erent functions
that are defined within the interval. Fourier’s paper was rejected by the Academy
on the grounds that it lacked mathematical rigor, because he did not examine thequestion of the convergence of the series.
The trigonometric series (4.2) is the only series which corresponds to f
x.
Questions concerning its convergence and, if it does, the conditions under
which it converges to f
xare many and dicult. These problems were partially
answered by Peter Gustave Lejeune Dirichlet (German mathematician, 1805–1859) and will be discussed briefly later.
Now let us assume that the series exists, converges, and may be integrated term
by term. Multiplying both sides by cos mx, then integrating the result from ÿto
,w eh a v e
Z
ÿf
xcosmx dx a0
2Z
ÿcosmx dx X1
n1anZ
ÿcosnxcosmx dx
X1
n1bnZ
ÿsinnxcosmx dx :
4:3
Now, using the following important properties of sines and cosines:
Z
ÿcosmx dx Z
ÿsinmx dx 0i f m1;2;3;...;
Z
ÿcosmxcosnx dx Z
ÿsinmxsinnx dx 0i f n6m;
ifnm;(
146FOURIER SERIES AND INTEGRALS
Z
ÿsinmxcosnx dx 0;for all m;n/C620;
we find that all terms on the right hand side of Eq. (4.3) except one vanish:
an1
Z
ÿf
xcosnx dx ;nintegers ;
4:4a
the expression for a0can be obtained from the general expression for anby setting
n0.
Similarly, if Eq. (4.2) is multiplied through by sin mxand the result is integrated
fromÿto, all terms vanish save that involving the square of sin nx, and so we
have
bn1
Z
ÿf
xsinnx dx :
4:4b
Eqs. (4.4a) and (4.4b) are known as the Euler–Fourier formulas.
From the definition of a definite integral it follows that, if f
xis single-valued
and continuous within the interval ÿ; or merely piecewise continuous (con-
tinuous except at a finite numbers of finite jumps in the interval), the integrals in
Eqs. (4.4) exist and we may compute the Fourier coecients of f
xby Eqs. (4.4).
If there exists a finite discontinuity in f
xat the point x0(Fig. 4.1), the coe-
cients a0;an;bnare determined by integrating first to xx0and then from x0to,
as
an1
Zx0
ÿf
xcosnx dx Z
x0f
xcosnx dx
;
4:5a
bn1
Zx0
ÿf
xsinnx dx Z
x0f
xsinnx dx
:
4:5b
This procedure may be extended to any finite number of discontinuities.
Example 4.2
Find the Fourier series which represents the function
f
xÿkÿ<x<0
k0<x<and f
x2f
x;/C26
in the interval ÿx.
147FOURIER SERIES; EULER–FOURIER FORMULAS
Solution: The Fourier coecients are readily calculated:
an1
Z0
ÿ
ÿkcosnx dx Z
0kcosnx dx
1
ÿksinnx
n/C12/C12/C12/C120
ÿksinnx
n/C12/C12/C12/C12
0
0"
bn1
Z0
ÿ
ÿksinnx dx Z
0ksinnx dx
1
kcosnx
n/C12/C12/C12/C120
ÿÿkcosnx
n/C12/C12/C12/C12
0
2k
n
1ÿcosn"
Now cos nÿ1 for odd n, and cos n1 for even n. Thus
b14k=; b20;b34k=3;b40;b54k=5;...
and the corresponding. Fourier series is
4k
sinx1
3sin 3x15sin 5x
:
For the special case k=2, the Fourier series becomes
2 sinx23sin 3x25sin 5x :
The first two terms are shown in Fig. 4.4, the solid curve is their sum. We will see
that as more and more terms in the Fourier series expansion are included, the sum
more and more nearly approaches the shape of f
x. This will be further demon-
strated by next example.
Example 4.3
Find the Fourier series that represents the function defined by
f
t0; ÿ<t<0
sint; 0<t</C26
in the interval ÿ<t< :
148FOURIER SERIES AND INTEGRALS
Solution:
an1
Z0
ÿ0cosnt dtZ
0sintcosnt dt
ÿ1
2cos
1ÿnt
1ÿncos
1nt
1n
0cosn1
1ÿn2; n61/C12/C12/C12/C12/C12;
a
11
Z
0sintcostd t1
sin2t
2/C12/C12/C12/C12
00;
bn1
Z0
ÿ0sinnt dtZ
0sintsinnt dt
1
2sin
1ÿnt
1ÿnÿsin
1nt
1n
00
b11
Z
0sin2td t1
t
2ÿsin 2t
40
1
2:
Accordingly the Fourier expansion of f
tin /C91ÿ; /C93 may be written
f
t1
sint
2ÿ2
cos 2t
3cos 4t
15cos 6t
35cos 8t
63
:
The first three partial sums Sn
n1;2;3) are shown in Fig. 4.5: S11=;
S21=sint=2, and S31=sin
t=2ÿ2 cos
2t=3:
149FOURIER SERIES; EULER–FOURIER FORMULAS
Figure 4.4. The first two partial sums.
/C71ibb/C39s phenomena
From Figs. 4.4 and 4.5, two features of the Fourier expansion should be noted:
(a) at the points of the discontinuity, the series yields the mean value;
(b) in the region immediately adjacent to the points of discontinuity, the expan-
sion overshoots the original function. This e/C128ect is known as the /C71ibb/C39s
phenomena and occurs in all order of approximation.
/C67on/C118ergence of Fourier series and /C68irichlet conditions
The serious question of the convergence of Fourier series still remains: if we
determine the Fourier coecients an;bnof a given function f
xfrom Eq. (4.4)
and form the Fourier series given on the right hand side of Eq. (4.2), will itconverge toward f
x/C63 This question was partially answered by Dirichlet.
Here is a restatement of the results of his study, which is often calledDirichlet’s theorem:
(1) If f
xis defined and single-valued except at a finite number of point in
ÿ; ,
(2) if f
xis periodic outside ÿ; with period 2 (that is, f
x2f
x,
and
(3) if f
xandf
0
xare piecewise continuous in ÿ; ,
150FOURIER SERIES AND INTEGRALS
Figure 4.5. The first three partial sums of the series.
then the series on the right hand side of Eq. (4.2), with coecients anandbngiven
by Eqs. (4.4), converges to
(i)f
x,i fxis a point of continuity, or
(ii)1
2f
x0f
xÿ0,i fxis a point of discontinuity as shown in Fig. 4.6,
where f
x0andf
xÿ0are the right and left hand limits of f
xatxand
represent lim /C34!0f
x/C34and lim /C34!0f
xÿ/C34respectively, where /C34/C620.
The proof of Dirichlet’s theorem is quite technical and is omitted in this treat-
ment. The reader should remember that the Dirichlet conditions (1), (2), and (3)
imposed on f
xare sucient but not necessary. That is, if the above conditions
are satisfied the convergence is guaranteed; but if they are not satisfied, the seriesmay or may not converge. The Dirichlet conditions are generally satisfied in
practice.
/C72alf-range Fourier series
Unnecessary work in determining Fourier coecients of a function can beavoided if the function is odd or even. A function f
xis called odd if
f
ÿxÿ f
xand even if f
xf
ÿxf
x. It is easy to show that in the
Fourier series corresponding to an odd function f
o
x, only sine terms can be
present in the series expansion in the interval ÿ<x<, for
an1
Z
ÿfo
xcosnx dx 1
Z0
ÿfo
xcosnx dx Z
0fo
xcosnx dx
1
ÿZ
0fo
xcosnx dx Z
0fo
xcosnx dx
0 n0;1;2;...;
4:6a
151HALF-RANGE FOURIER SERIES
Figure 4.6. A piecewise continuous function.
but
bn1
Z0
ÿfo
xsinnx dx Z
0fo
xsinnx dx
2
Z
0fo
xsinnx dx n 1;2;3;...:
4:6b
Here we have made use of the fact that cos( ÿnxcosnxand sin
ÿnx
ÿsinnx. Accordingly, the Fourier series becomes
fo
xb1sinxb2sin 2x :
Similarly, in the Fourier series corresponding to an even function fe
x, only
cosine terms (and possibly a constant) can be present. Because in this case,
fe
xsinnxis an odd function and accordingly bn0 and the anare given by
an2
Z
0fe
xcosnx dx n 0;1;2;...:
4:7
Note that the Fourier coecients anandbn, Eqs. (4.6) and (4.7) are computed in
the interval (0, ) which is halfof the interval ( ÿ; ). Thus, the Fourier sine or
cosine series in this case is often called a half-range Fourier series.
Any arbitrary function (neither even nor odd) can be expressed as a combina-
tion of fe
xandfo
xas
f
x1
2f
xf
ÿx 12f
xÿf
ÿx fe
xfo
x:
When a half-range series corresponding to a given function is desired, the
function is generally defined in the interval (0, ) and then the function is specified
as odd or even, so that it is clearly defined in the other half of the interval
ÿ;0.
/C67hange of inter/C118al
A Fourier expansion is not restricted to such intervals as ÿ<x< and
0<x<. In many problems the period of the function to be expanded may
be some other interval, say 2 L. How then can the Fourier series developed
above be applied to the representation of periodic functions of arbitrary period/C63
The problem is not a dicult one, for basically all that is involved is to change the
variable. Let
z
Lx
4:8a
then
f
zf
x=L/C70
x:
4:8b
Thus, if f
zis expanded in the interval ÿ<z<, the coecients being deter-
mined by expressions of the form of Eqs. (4.4a) and (4.4b), the coecients for the
152FOURIER SERIES AND INTEGRALS
expansion of /C70
xin the interval ÿL<x<Lmay be obtained merely by sub-
stituting Eqs. (4.8) into these expressions. We have then
an1
LZL
ÿL/C70
xcosn
Lxd x n 0;1;2;3;...;
4:9a
bn1
LZL
ÿL/C70
xsinn
Lxd x; n1;2;3;...:
4:9b
The possibility of having expanding functions in which the period is other than
2increases the usefulness of Fourier expansion. As an example, consider the
value of L, it is obvious that the larger the value of L, the larger the basic period
of the function being expanded. As L!1 , the function would not be periodic at
all. We will see later that in such cases the Fourier series becomes a Fourier
integral.
Parse/C118al/C39s identit/C121
Parseval’s identity states that:
1
2LZL
ÿLf
x2dxa0
22
1
2X1
n1
a2
nb2n;
4:10
ifanandbnare coecients of the Fourier series of f
xand if f
xsatisfies the
Dirichlet conditions.
It is easy to prove this identity. Assuming that the Fourier series corresponding
tof
xconverges to f
x
f
xa0
2X1
n1ancosnx
Lbnsinnx
L
:
Multiplying by f
xand integrating term by term from ÿLtoL, we obtain
ZL
ÿLf
x2dxa0
2ZL
ÿLf
xdx
X1
n1anZL
ÿLf
xcosnx
LdxbnZL
ÿLf
xsinnx
Ldx/C26/C27
a20
2LLX1
n1a2
nb2nÿ
;
4:11
where we have used the results
ZL
ÿLf
xcosnx
LdxLan;ZL
ÿLf
xsinnx
LdxLbn;ZL
ÿLf
xdxLa0:
The required result follows on dividing both sides of Eq. (4.11) by L.
153PARSEVAL’S IDENTITY
Parseval’s identity shows a relation between the average of the square of f
x
and the coecients in the Fourier series for f
x:
the average of ff
xg2isRL
ÿLf
x2dx=2L;
the average of ( a0=2is (a0=22;
the average of ( ancosnxisa2
n=2;
the average of ( bnsinnxisb2
n=2.
Example 4.4
Expand f
xx;0<x<2, in a half-range cosine series, then write Parseval’s
identity corresponding to this Fourier cosine series.
Solution: We first extend the definition of f
xto that of the even function of
period 4 shown in Fig. 4.7. Then 2 L4;L2. Thus bn0 and
an2
LZL
0f
xcosnx
Ldx2
2Z2
0f
xcosnx
2dx
x2
nsinnx
2
ÿ1ÿ4
n22cosnx
2 2
0
ÿ4
n22cosnÿ1
ifn60:
Ifn0,
a0ZL
0xdx2:
Then
f
x1X1
n14
n22cosnÿ1
cosnx
2:
We now write Parseval’s identity. We first compute the average of f
x2:
the average of f
x21
2Z2
ÿ2f
xfg2dx12Z2
ÿ2x2dx8
3;
154FOURIER SERIES AND INTEGRALS
Figure 4.7.
then the average
a2
0
2X1
n1a2nb2nÿ
22
2X1
n116
n44cosnÿ1
2:
Parseval’s identity now becomes
8
3264
41
141
341
54
;
or
1
141
341
544
96
which shows that we can use Parseval’s identity to find the sum of an infinite
series. With the help of the above result, we can find the sum Sof the following
series:
1
141
241
341
441
n4 :
S1
141
241
341
441
141
341
54
1
241
441
64
1
141
341
54
1
241
141
241
341
44
4
96S
16
from which we find S4=90.
/C65lternati/C118e forms of Fourier series
Up to this point the Fourier series of a function has been written as an infinite
series of sines and cosines, Eq. (4.2):
f
xa0
2X1
n1ancosnx
Lbnsinnx
L
:
This can be converted into other forms. In this section, we just discuss two alter-
native forms. Let us first write, with =L
ancosnxbnsinnx
a2nb2nqan
a2nb2n/C112 cosnxbn
a2nb2n/C112 sinnx/C32!
:
155ALTERNATIVE FORMS OF FOURIER SERIES
Now let (see Fig. 4.8)
cosnan
a2nb2n/C112 ;sinnbn
a2nb2n/C112 ;sontanÿ1bn
an
;
Cn
a2nb2nq
;C01
2a0;
then we have the trigonometric identity
ancosnxbnsinnxCncosnxÿn
;
and accordingly the Fourier series becomes
f
xC0X1
n1Cncosnxÿn
:
4:12
In this new form, the Fourier series represents a periodic function as a sum of
sinusoidal components having di/C128erent frequencies. The sinusoidal component of
frequency nis called the nth harmonic of the periodic function. The first har-
monic is commonly called the fundamental component. The angles nand the
coecients Cnare known as the phase angle and amplitude.
Using Euler’s identities eicosisinwhere i2ÿ1, the Fourier series
forf
xcan be converted into complex form
f
xX1
nÿ1cneinx=L;
4:13a
where
cnan/C7ibn
1
2LZL
ÿLf
xeÿinx=Ldx;forn/C620:
4:13b
Eq. (4.13a) is obtained on the understanding that the Dirichlet conditions are
satisfied and that f
xis continuous at x.I ff
xis discontinuous at x, the left
hand side of Eq. (4.13a) should be replaced by f
x0f
xÿ0=2.
The exponential form (4.13a) can be considered as a basic form in its own right:
it is not obtained by transformation from the trigonometric form, rather it is
156FOURIER SERIES AND INTEGRALS
Figure 4.8.
constructed directly from the given function. Furthermore, in the complex repre-
sentation defined by Eqs. (4.13a) and (4.13b), a certain symmetry between theexpressions for a function and for its Fourier coecients is evident. In fact the
expressions (4.13a) and (4.13b) are of essentially the same structure, as the follow-
ing correlation reveals:
x/C24L;f
x/C24c
nc
n;einx=L/C24eÿinx=L;X1
nÿ1
/C241
2LZL
ÿL
dx:
This duality is worthy of note, and as our development proceeds to the Fourierintegral, it will become more striking and fundamental.
Integration and di/C128erentiation of a Fourier series
The Fourier series of a function f
xmay always be integrated term-by-term to
give a new series which converges to the integral of f
x.I ff
xis a continuous
function of xfor all x, and is periodic (of period 2 ) outside the interval
ÿ<x<, then term-by-term di/C128erentiation of the Fourier series of f
x
leads to the Fourier series of f
0
x, provided f0
xsatisfies Dirichlet’s conditions.
/C86ibrating strings
/C84he equation of motion of transverse vibration
There are numerous applications of Fourier series to solutions of boundary value
problems. Here we consider one of them, namely vibrating strings. Let a string of
length Lbe held fixed between two points (0, 0) and ( L, 0) on the x-axis, and then
given a transverse displacement parallel to the y-axis. Its subsequent motion, with
no external forces acting on it, is to be considered; this is described by finding thedisplacement yas a function of xandt(if we consider only vibration in one plane,
and take the xyplane as the plane of vibration). We will assume that /C26, the mass
per unit length is uniform over the entire length of the string, and that the string is
perfectly flexible, so that it can transmit tension but not bending or shearing
forces.
As the string is drawn aside from its position of rest along the x-axis, the
resulting increase in length causes an increase in tension, denoted by P. This
tension at any point along the string is always in the direction of the tangent tothe string at that point. As shown in Fig. 4.9, a force P
xAacts at the left hand
side of an element ds, and a force P
xdxAacts at the right hand side, where A
is the cross-sectional area of the string. If is the inclination to the horizontal,
then
/C70
x/C129APcos
dÿAPcos; /C70y/C129APsin
dÿAPsin:
157INTEGRATION AND DIFFERENTIATION OF A FOURIER SERIES
We limit the displacement to small values, so that we may set
cos1ÿ2=2;sin/C129/C129tandy=dx;
then
/C70yAPdy
dx
xdxÿdy
dx
x
APd2y
dx2dx:
Using Newton’s second law, the equation of motion of transverse vibration of the
element becomes
/C26Adx/C642y
/C64t2AP/C642y
/C64x2dx; or/C642y
/C64x21
/C1182/C642y
/C64t2;/C118
P=/C26/C112
:
Thus the transverse displacement of the string satisfies the partial di/C128erential wave
equation
/C642y
/C64x21
/C1182/C642y
/C64t2;0<x<L;t/C620
4:14
with the following boundary conditions: y
0;ty
L;t0;/C64y=/C64t0;
y
x;0f
x; where f
xdescribes the initial shape (position) of the string,
and/C118is the velocity of propagation of the wave along the string.
Solution of the /C119ave equation
To solve this boundary value problem, let us try the method of separation vari-
ables:
y
x;tX
xT
t:
4:15
Substituting this into Eq. (4.14) yields
1=X
d2X=dx2
1=/C1182T
d2T=dt2.
158FOURIER SERIES AND INTEGRALS
Figure 4.9. A vibrating string.
Since the left hand side is a function of xonly and the right hand side is a function
of time only, they must be equal to a common separation constant, which we will
callÿ2. Then we have
d2X=dx2ÿ2X;X
0X
L0
4:16a
and
d2T=dt2ÿ2/C1182Td T =dt0a t t0:
4:16b
Both of these equations are typical eigenvalue problems: we have a di/C128erentialequation containing a parameter , and we seek solutions satisfying certain
boundary conditions. If there are special values of for which non-trivial solu-
tions exist, we call these eigenvalues, and the corresponding solutions eigensolu-tions or eigenfunctions.
The general solution of Eq. (4.16a) can be written as
X
xA
1sin
xB1cos
x:
Applying the boundary conditions
X
00/C41B10;
and
X
L0/C41A1sin
L0
A10 is the trivial solution X0 (so y0); hence we must have sin
L0,
that is,
Ln;n1;2;...;
and we obtain a series of eigenvalues
nn=L;n1;2;...
and the corresponding eigenfunctions
Xn
xsin
n=Lx;n1;2;...:
To solve Eq. (4.16b) for T
twe must use one of the values nfound above. The
general solution is of the form
T
tA2cos
n/C118tB2sin
n/C118t:
The boundary condition leads to B20.
The general solution of Eq. (4.14) is hence a linear superposition of the solu-
tions of the form
y
x;tX1
n1Ansin
nx=Lcos
n/C118t=L;
4:17
159VIBRATING STRINGS
theAnare as yet undetermined constants. To find An, we use the boundary
condition y
x;tf
xatt0, so that Eq. (4.17) reduces to
f
xX1
n1Ansin
nx=L:
Do you recognize the infinite series on the right hand side/C63 It is a Fourier sine
series. To find An, multiply both sides by sin( mx=L) and then integrate with
respect to xfrom 0 to Land we obtain
Am2
LZL
0f
xsin
mx=Ldx; m1;2;...
where we have used the relation
ZL
0sin
mx=Lsin
nx=LdxL
2mn:
Eq. (4.17) now gives
y
x;tX1
n12
LZL
0f
xsinnx
Ldx
sinnx
Lcosn/C118t
L:
4:18
The terms in this series represent the natural modes of vibration. The frequency
of the nth normal mode fnis obtained from the term involving cos
n/C118t=Land is
given by
2fnn/C118=Lorfnn/C118=2L:
All frequencies are integer multiples of the lowest frequency f1.W ec a l l f1the
fundamental frequency or first harmonic, and f2andf3the second and third
harmonics (or first and second overtones) and so on.
/C82L/C67 circuit
Another good example of application of Fourier series is an /C82L/C67 circuit driven by
a variable voltage /C69
twhich is periodic but not necessarily sinusoidal (see Fig.
4.10). We want to find the current I
tflowing in the circuit at time t.
According to Kirchho/C128 ’s second law for circuits, the impressed voltage /C69
t
equals the sum of the voltage drops across the circuit components. That is,
LdI
dtRI/C81
C/C69
t;
where /C81is the total charge in the capacitor /C67. But Id/C81=dt, thus di/C128erentiating
the above di/C128erential equation once we obtain
Ld2I
dt2RdI
dt1
CId/C69
dt:
160FOURIER SERIES AND INTEGRALS
Under steady-state conditions the current I
tis also periodic, with the same
period Pas for /C69
t. Let us assume that both /C69
tandI
tpossess Fourier
expansions and let us write them in their complex forms:
/C69
tX1
nÿ1/C69nein/C33t; I
tX1
nÿ1cnein/C33t
/C332=P:
Furthermore, we assume that the series can be di/C128erentiated term by term. Thus
d/C69
dtX1
nÿ1in/C33/C69nein/C33t;dI
dtX1
nÿ1in/C33cnein/C33t;d2I
dt2X1
nÿ1
ÿn2/C332cnein/C33t:
Substituting these into the last (second-order) di/C128erential equation and equating
the coecients with the same exponential eint, we obtain
ÿn2/C332Lin/C33R1=Cÿ
cnin/C33/C69n:
Solving for cn
cnin/C33=L
1=CL
2ÿn2/C332iR=L
n/C33/C69n:
Note that 1/ L/C67is the natural frequency of the circuit and /C82/C47L is the attenuation
factor of the circuit. The Fourier coecients for /C69
tare given by
/C69n1
PZP=2
ÿP=2/C69
teÿin/C33tdt:
The current I
tin the circuit is given by
I
tX1
nÿ1cnein/C33t:
161/C82L/C67 CIRCUIT
Figure 4.10. The /C82L/C67 circuit.
/C79rthogonal functions
Many of the properties of Fourier series considered above depend on orthogonal
properties of sine and cosine functions
ZL
0sinmx
Lsinnx
Ldx0;ZL
0cosmx
Lcosnx
Ldx0
m6n:
In this section we seek to generalize this orthogonal property. To do so we firstrecall some elementary properties of real vectors in three-dimensional space.
Two vectors /C65and/C66are called orthogonal if /C65/C660. Although not geome-
trically or physically obvious, we generalize these ideas to think of a function, say
A
x, as being an infinite-dimensional vector (a vector with an infinity of compo-
nents), the value of each component being specified by substituting a particularvalue of xtaken from some interval ( a,b), and two functions, A
xandB
xare
orthogonal in ( a,b)i f
Z
b
aA
xB
xdx0:
4:19
The left-side of Eq. (4.19) is called the scalar product of A
xandB
xand
denoted by, in the Dirac bracket notation, hA
xjB
xi. The first factor in the
bracket notation is referred to as the bra and the second factor as the ket, sotogether they comprise the bracket.
A vector /C65is called a unit vector or normalized vector if its magnitude is unity:
/C65/C65A
21. Extending this concept, we say that the function A
xis normal
or normalized in ( a,b)i f
hA
xjA
xi Zb
aA
xA
xdx1:
4:20
If we have a set of functions ’i
x;i1;2;3;...;having the properties
’m
x hj ’n
xi Zb
a’m
x’n
xdxmn;
4:20a
where nmis the Kronecker delta symbol, we then call such a set of functions an
orthonormal set in ( a,b). For example, the set of functions ’m
x
2=1=2sin
mx;m1;2;3;...is an orthonormal set in the interval 0 x.
Just as in three-dimensional vector space, any vector /C65can be expanded in the
form /C65A1^e1A2^e2A3^e3, we can consider a set of orthonormal functions ’i
as base vectors and expand a function f
xin terms of them, that is,
f
xX1
n1cn’n
xaxb;
4:21
162FOURIER SERIES AND INTEGRALS
the series on the right hand side is called an orthonormal series; such series are
generalizations of Fourier series. Assuming that the series on the right convergestof
x, we can then multiply both sides by ’
m
xand integrate both sides from a
tobto obtain
cmhf
xj’m
xi Zb
af
x’m
xdx;
4:21a
cmcan be called the generalized Fourier coecients.
Multiple Fourier series
A Fourier expansion of a function of two or three variables is often very useful inmany applications. Let us consider the case of a function of two variables, say
f
x;y. For example, we can expand f
x;yinto a double Fourier sine series
f
x;yX
1
m1X1
n1Bmnsinmx
L1sinny
L2;
4:22
where
Bmn4
L1L2ZL1
0ZL2
0f
x;ysinmx
L1sinny
L2dxdy :
4:22a
Similar expansions can be made for cosine series and for series having both sines
and cosines.
To obtain the coecients Bmn, let us rewrite f
x;yas
f
x;yX1
m1Cmsinmx
L1;
4:23
where
CmX1
n1Bmnsinny
L2:
4:23a
Now we can consider Eq. (4.23) as a Fourier series in which yis kept constant
so that the Fourier coecients Cmare given by
Cm2
L1ZL1
0f
x;ysinmx
L1dx:
4:24
On noting that Cmis a function of y, we see that Eq. (4.23a) can be considered as a
Fourier series for which the coecients Bmnare given by
Bmn2
L2ZL2
0Cmsinny
L2dy:
163MULTIPLE FOURIER SERIES
Substituting Eq. (4.24) for Cminto the above equation, we see that Bmnis given by
Eq. (4.22a).
Similar results can be obtained for cosine series or for series containing both
sines and cosines. Furthermore, these ideas can be generalized to triple Fourier
series, etc. They are very useful in solving, for example, wave propagation and
heat conduction problems in two or three dimensions. Because they lie outside of
the scope of this book, we have to omit these interesting applications.
Fourier integrals and Fourier transforms
The properties of Fourier series that we have thus far developed are adequate forhandling the expansion of any periodic function that satisfies the Dirichlet con-
ditions. But many problems in physics and engineering do not involve periodic
functions, and it is therefore desirable to generalize the Fourier series method toinclude non-periodic functions. A non-periodic function can be considered as a
limit of a given periodic function whose period becomes infinite, as shown in
Examples 4.5 and 4.6.
Example 4.5
Consider the periodic functions f
L
x
fL
x0 when ÿL=2<x<ÿ1
1 when ÿ1<x<1
0 when 1 <x<L=28
><
>:;
164FOURIER SERIES AND INTEGRALS
Figure 4.11. Square wave function:
aL4;
bL8;
cL!1 .
which has period L/C622. Fig. 4.11( a) shows the function when L4. If Lis
increased to 8, the function looks like the one shown in Fig. 4.11( b). As
L!1 we obtain a non-periodic function f
x, as shown in Fig. 4.11( c):
f
x1ÿ1<x<1
0 otherwise/C26
:
Example 4.6
Consider the periodic function /C103L
x(Fig. 4.12( a)):
/C103L
xeÿjxjwhen ÿL=2<x<L=2:
AsL!1 we obtain a non-periodic function /C103
x/C58/C103
xlimL!1/C103L
x(Fig.
4.12( b)).
By investigating the limit that is approached by a Fourier series as the period of
the given function becomes infinite, a suitable representation for non-periodic
functions can perhaps be obtained. To this end, let us write the Fourier series
representing a periodic function f
xin complex form:
f
xX1
nÿ1cnei/C33x;
4:25
cn1
2LZL
ÿLf
xeÿi/C33xdx
4:26
where /C33denotes n=L
/C33n
L;npositive or negative :
4:27
The transition L!1 is a little tricky since cnapparently approaches zero, but
these coecients should not approach zero. We can ask for help from Eq. (4.27),
from which we have
/C33
=Ln;
165FOURIER INTEGRALS AND FOURIER TRANSFORMS
Figure 4.12. Sawtooth wave functions:
aÿL=2<x<L=2;
bL!1 .
and the ‘adjacent’ values of /C33are obtained by setting n1, which corresponds
to
L=/C331:
Then we can multiply each term of the Fourier series by
L=/C33and obtain
f
xX1
nÿ1L
cn
ei/C33x/C33;
where
L
cn1
2ZL
ÿLf
xeÿi/C33xdx:
The troublesome factor 1 =Lhas disappeared. Switching completely to the /C33
notation and writing
L=cncL
/C33, we obtain
cL
/C331
2ZL
ÿLf
xeÿi/C33xdx
and
f
xX1
L/C33=ÿ1cL
/C33ei/C33x/C33:
In the limit as L!1 , the /C33s are distributed continuously instead of discretely,
/C33!d/C33and this sum is exactly the definition of an integral. Thus the last
equations become
c
/C33lim
L!1cL
/C331
2Z1
ÿ1f
xeÿi/C33xdx
4:28
and
f
xZ1
ÿ1c
/C33ei/C33xd/C33:
4:29
This set of formulas is known as the Fourier transformation, in somewhat di/C128er-
ent form. It is easy to put them in a symmetrical form by defining
/C103
/C33
2p
c
ÿ/C33;
then Eqs. (4.28) and (4.29) take the symmetrical form
/C103
/C331
2pZ1
ÿ1f
x0eÿi/C33x0dx0;
4:30
f
x1 2pZ
1
ÿ1/C103
/C33ei/C33xd/C33:
4:31
166FOURIER SERIES AND INTEGRALS
The function /C103
/C33is called the Fourier transform of f
xand is written
/C103
/C33/C70ff
xg. Eq. (4.31) is the inverse Fourier transform of /C103
/C33and is written
f
x/C70ÿ1f/C103
/C33g; sometimes it is also called the Fourier integral representation
off
x. The exponential function eÿi/C33xis sometimes called the kernel of trans-
formation.
It is clear that /C103
/C33is defined only if f
xsatisfies certain restrictions. For
instance, f
xshould be integrable in some finite region. In practice, this means
thatf
xhas, at worst, jump discontinuities or mild infinite discontinuities. Also,
the integral should converge at infinity. This would require that f
x!0a s
x! 1 .
A very common sucient condition is the requirement that f
xis absolutely
integrable. That is, the integral
Z1
ÿ1f
xjj dx
exists. Since jf
xeÿi/C33xjjf
xj, it follows that the integral for /C103
/C33is absolutely
convergent; therefore it is convergent.
It is obvious that /C103
/C33is, in general, a complex function of the real variable /C33.
So if f
xis real, then
/C103
ÿ/C33/C103/C42
/C33:
There are two immediate corollaries to this property:
(1)f
xis even, /C103
/C33is real;
(2) if f
xis odd, /C103
/C33is purely imaginary.
Other, less symmetrical forms of the Fourier integral can be obtained by working
directly with the sine and cosine series, instead of with the exponential functions.
Example 4.7
Consider the Gaussian probability function f
x/C78eÿx2, where /C78andare
constant. Find its Fourier transform /C103
/C33, then graph f
xand/C103
/C33.
Solution: Its Fourier transform is given by
/C103
/C331
2pZ1
ÿ1f
xeÿi/C33xdx/C78 2pZ
1
ÿ1eÿx2eÿi/C33xdx:
This integral can be simplified by a change of variable. First, we note that
ÿx2ÿi/C33xÿ
x pi/C33=2 p2ÿ/C332=4;
and then make the change of variable x pi/C33=2 puto obtain
/C103
/C33 /C78
2p eÿ/C332=4Z1
ÿ1eÿu2du/C782peÿ/C332=4:
167FOURIER INTEGRALS AND FOURIER TRANSFORMS
It is easy to see that /C103
/C33is also a Gaussian probability function with a peak at the
origin, monotonically decreasing as /C33! 1 . Furthermore, for large ,f
xis
sharply peaked but /C103
/C33is flattened, and vice versa as shown in Fig. 4.13. It is
interesting to note that this is a general feature of Fourier transforms. We shall see
later that in quantum mechanical applications it is related to the Heisenberg
uncertainty principle.
The original function f
xcan be retrieved from Eq. (4.31) which takes the
form
1
2pZ1
ÿ1/C103
/C33ei/C33xd/C331 2p /C78
2pZ1
ÿ1eÿ/C332=4ei/C33xd/C33
1
2p/C78
2pZ1
ÿ1eÿ0/C332eÿi/C33x0d/C33
in which we have set 01=4, and x0ÿx. The last integral can be evaluated
by the same technique, and we finally find
1
2pZ1
ÿ1/C103
/C33ei/C33xd/C331 2p /C78
2pZ1
ÿ1eÿ0/C332eÿi/C33x0d/C33
/C782p
2p
eÿx2
/C78eÿx2f
x:
Example 4.8
Given the box function which can represent a single pulse
f
x1jxja
0xjj/C62a/C26
find the Fourier transform of f
x,/C103
/C33; then graph f
xand/C103
/C33fora3.
168FOURIER SERIES AND INTEGRALS
Figure 4.13. Gaussian probability function:
alarge ;
bsmall .
Solution: The Fourier transform of f
xis, as shown in Fig. 4.14,
/C103
/C331
2pZ1
ÿ1f
x0eÿi/C33x0dx01 2pZ
a
ÿa
1eÿi/C33x0dx01 2p eÿi/C33x0
ÿi/C33a
ÿa/C12/C12/C12/C12/C12
2
/C114
sin/C33a
/C33;/C3360:
For/C330, we obtain /C103
/C33
2=/C112
a.
The Fourier integral representation of f
xis
f
x1
2pZ1
ÿ1/C103
/C33ei/C33xd/C331
2Z1
ÿ12 sin /C33a
/C33ei/C33xd/C33:
Now
Z1
ÿ1sin/C33a
/C33ei/C33xd/C33Z1
ÿ1sin/C33acos/C33x
/C33d/C33iZ1
ÿ1sin/C33asin/C33x
/C33d/C33:
The integrand in the second integral is odd and so the integral is zero. Thus we
have
f
x1
2pZ1
ÿ1/C103
/C33ei/C33xd/C331
Z1
ÿ1sin/C33acos/C33x
/C33d/C332
Z1
0sin/C33acos/C33x
/C33d/C33;
the last step follows since the integrand is an even function of /C33.
It is very dicult to evaluate the last integral. But a known property of f
xwill
help us. We know that f
xis equal to 1 for jxja, and equal to 0 for jxj/C62a.
Thus we can write
2
Z1
0sin/C33acos/C33x
/C33d/C331jxja
0jxj/C62a/C26
169FOURIER INTEGRALS AND FOURIER TRANSFORMS
Figure 4.14. The box function.
Just as in Fourier series expansion, we also expect to observe Gibb’s
phenomenon in the case of Fourier integrals. Approximations to the Fourier
integral are obtained by replacing 1by:
Z
0sin/C33cos/C33x
/C33d/C33;
where we have set a1. Fig. 4.15 shows oscillations near the points of disconti-
nuity of f
x. We might expect these oscillations to disappear as !1 , but they
are just shifted closer to the points x1.
Example 4.9
Consider now a harmonic wave of frequency /C330,ei/C330t, which is chopped to a life-
time of 2 Tseconds (Fig. 4.16( a)):
f
tei/C330tÿTtT
0 jtj/C620:(
The chopping process will introduce many new frequencies in varying amounts,given by the Fourier transform. Then we have, according to Eq. (4.30),
/C103
/C33
2
ÿ1=2ZT
ÿTei/C330teÿi/C33tdt
2ÿ1=2ZT
ÿTei
/C330ÿ/C33tdt
2ÿ1=2ei
/C330ÿ/C33t
i
/C330ÿ/C33/C12/C12/C12/C12T
ÿT
2=1=2Tsin
/C330ÿ/C33T
/C330ÿ/C33T:
This function is plotted schematically in Fig. 4.16( b). (Note that
limx!0
sinx=x1.) The most striking aspect of this graph is that, although
the principal contribution comes from the frequencies in the neighborhood of
/C330, an infinite number of frequencies are presented. Nature provides an example
of this kind of chopping in the emission of photons during electronic and nucleartransitions in atoms. The light emitted from an atom consists of regular vibrationsthat last for a finite time of the order of 10
ÿ9s or longer. When light is examined
by a spectroscope (which measures the wavelengths and, hence, the frequencies)
we find that there is an irreducible minimum frequency spread for each spectrum
line. This is known as the natural line width of the radiation.
The relative percentage of frequencies, other than the basic one, present
depends on the shape of the pulse, and the spread of frequencies depends on
170FOURIER SERIES AND INTEGRALS
Figure 4.15. The Gibb’s phenomenon.
the time Tof the duration of the pulse. As Tbecomes larger the central peak
becomes higher and the width /C33
2=Tbecomes smaller. Considering only
the spread of frequencies in the central peak we have
/C332=T;orT1:
Multiplying by the Planck constant hand replacing Tbyt, we have the relation
t/C69/C104:
4:32
A wave train that lasts a finite time also has a finite extension in space. Thus the
radiation emitted by an atom in 10ÿ9s has an extension equal to 3 10810ÿ9
310ÿ1m. A Fourier analysis of this pulse in the space domain will yield a graph
identical to Fig. 4.11( b), with the wave numbers clustered around
k0
2= 0/C330=/C118. If the wave train is of length 2 a, the spread in wave number
will be given by ak2, as shown below. This time we are chopping an infinite
plane wave front with a shutter such that the length of the packet is 2 a, where
2a2/C118T, and 2 Tis the time interval that the shutter is open. Thus
/C32
xeik0x;ÿaxa
0; jxj/C62a:(
Then
/C30
k
2ÿ1=2Z1
ÿ1/C32
xeÿikxdx
2ÿ1=2Za
ÿa/C32
xeÿikxdx
2=1=2asin
k0ÿka
k0ÿka:
This function is plotted in Fig. 4.17: it is identical to Fig. 4.16( b), but here it is the
wave vector (or the momentum) that takes on a spread of values around k0. The
breadth of the central peak is k2=a,o rak2.
171FOURIER INTEGRALS AND FOURIER TRANSFORMS
Figure 4.16. ( a) A chopped harmonic wave ei/C330tthat lasts a finite time 2 T.
b
Fourier transform of e
i/C330t;jtj<T, and 0 otherwise.
Fourier sine and cosine transforms
Iff
xis an odd function, the Fourier transforms reduce to
/C103
/C33
2
/C114Z1
0f
x0sin/C33x0dx0; f
x
2
/C114Z1
0/C103
/C33sin/C33xd/C33:
4:33a
Similarly, if f
xis an even function, then we have Fourier cosine transforma-
tions:
/C103
/C33
2
/C114Z1
0f
x0cos/C33x0dx0; f
x
2
/C114Z1
0/C103
/C33cos/C33xd/C33:
4:33b
To demonstrate these results, we first expand the exponential function on the
right hand side of Eq. (4.30)
/C103
/C331
2pZ1
ÿ1f
x0eÿi/C33x0dx0
1 2pZ
1
ÿ1f
x0cos/C33x0dx0ÿi 2pZ
1
ÿ1f
x0sin/C33x0dx0:
Iff
xis even, then f
xcos/C33xis even and f
xsin/C33xis odd. Thus the second
integral on the right hand side of the last equation is zero and we have
/C103
/C331 2pZ
1
ÿ1f
x0cos/C33x0dx0
2
/C114Z1
0f
x0cos/C33x0dx0;
/C103
/C33is an even function, since /C103
ÿ/C33/C103
/C33. Next from Eq. (4.31) we have
f
x1
2pZ1
ÿ1/C103
/C33ei/C33xd/C33
1
2pZ1
ÿ1/C103
/C33cos/C33xd/C33i 2pZ
1
ÿ1/C103
/C33sin/C33xd/C33:
172FOURIER SERIES AND INTEGRALS
Figure 4.17. Fourier transform of eikx;jxja:
Since /C103
/C33is even, so /C103
/C33sin/C33xis odd and the second integral on the right hand
side of the last equation is zero, and we have
f
x1
2pZ1
ÿ1/C103
/C33cos/C33xd/C33
2
/C114Z1
0/C103
/C33cos/C33xd/C33:
Similarly, we can prove Fourier sine transforms by replacing the cosine by the
sine.
/C72eisenberg/C39s uncertaint/C121 principle
We have demonstrated in above examples that if f
xis sharply peaked, then /C103
/C33
is flattened, and vice versa. This is a general feature in the theory of Fourier
transforms and has important consequences for all instances of wave propaga-
tion. In electronics we understand now why we use a wide-band amplification in
order to reproduce a sharp pulse without distortion.
In quantum mechanical applications this general feature of the theory of
Fourier transforms is related to the Heisenberg uncertainty principle. We sawin Example 4.9 that the spread of the Fourier transform in kspace ( k) times
its spread in coordinate space ( a) is equal to 2
ak/C1292. This result is of
special importance because of the connection between values of kand momentum
/C112/C58/C112pk(where pis the Planck constant hdivided by 2 ). A particle localized in
space must be represented by a superposition of waves with di/C128erent momenta.
As a result, the position and momentum of a particle cannot be measured simul/C45
taneously with infinite precision; the product of ‘uncertainty in the position deter-
mination’ and ‘uncertainty in the momentum determination’ is governed by the
relation x/C112/C129/C104
apk/C1292p/C104,o r x/C112/C129/C104;xa. This statement is
called Heisenberg’s uncertainty principle. If position is known better, knowledge
of the momentum must be unavoidably reduced proportionally, and vice versa. A
complete knowledge of one, say k(and so p), is possible only when there is
complete ignorance of the other. We can see this in physical terms. A wavewith a unique value of kis infinitely long. A particle represented by an infinitely
long wave (a free particle) cannot have a definite position, since the particle can beanywhere along its length. Hence the position uncertainty is infinite in order that
the uncertainty in kis zero.
Equation (4.32) represents Heisenberg’s uncertainty principle in a di/C128erent
form. It states that we cannot know with infinite precision the exact energy of aquantum system at every moment in time. In order to measure the energy of a
quantum system with good accuracy, one must carry out such a measurement for
a suciently long time. In other words, if the dynamical state exists only for a
time of order t, then the energy of the state cannot be defined to a precision
better than /C104=t.
173HEISENBERG’S UNCERTAINTY PRINCIPLE
We should not look upon the uncertainty principle as being merely an unfor-
tunate limitation on our ability to know nature with infinite precision. We can use
it to our advantage. For example, when combining the time–energy uncertainty
relation with Einstein’s mass–energy relation ( /C69mc2) we obtain the relation
mt/C129/C104=c2. This result is very useful in our quest to understand the universe, in
particular, the origin of matter.
/C87a/C118e pac/C107ets and group /C118elocit/C121
Energy (that is, a signal or information) is transmitted by groups of waves, not asingle wave. Phase velocity may be greater than the speed of light c, ‘group
velocity’ is always less than c. The wave groups with which energy is transmitted
from place to place are called wave packets. Let us first consider a simple casewhere we have two waves ’
1and’2: each has the same amplitude but di/C128ers
slightly in frequency and wavelength,
’1
x;tAcos
/C33tÿkx;
’2
x;tAcos
/C33/C33tÿ
kkx;
where /C33/C28/C33and k/C28k. Each represents a pure sinusoidal wave extending to
infinite along the x-axis. Together they give a resultant wave
’’1’2
Acos
/C33tÿkxcos
/C33/C33tÿ
kkx fg :
Using the trigonometrical identity
cosAcosB2 cosAB
2cosAÿB
2;
we can rewrite ’as
’2 cos2/C33tÿ2kx/C33tÿkx
2cosÿ/C33tkx
2
2 cos1
2
/C33tÿkxcos
/C33tÿkx:
This represents an oscillation of the original frequency /C33, but with a modulated
amplitude as shown in Fig. 4.18. A given segment of the wave system, such as AB,
can be regarded as a ‘wave packet’ and moves with a velocity /C118/C103(not yet deter-
mined). This segment contains a large number of oscillations of the primary wave
that moves with the velocity /C118. And the velocity /C118/C103with which the modulated
amplitude propagates is called the group velocity and can be determined by therequirement that the phase of the modulated amplitude be constant. Thus
/C118
/C103dx=dt/C33=k!d/C33=dk:
174FOURIER SERIES AND INTEGRALS
The modulation of the wave is repeated indefinitely in the case of superposition of
two almost equal waves. We now use the Fourier technique to demonstrate thatany isolated packet of oscillatory disturbance of frequency /C33can be described in
terms of a combination of infinite trains of frequencies distributed around /C33. Let
us first superpose a system of nwaves
/C32
x;tX
n
j1Ajei
kjxÿ/C33jt;
where Ajdenotes the amplitudes of the individual waves. As napproaches infinity,
the frequencies become continuously distributed. Thus we can replace the sum-
mation with an integration, and obtain
/C32
x;tZ1
ÿ1A
kei
kxÿ/C33tdk;
4:34
the amplitude A
kis often called the distribution function of the wave. For
/C32
x;tto represent a wave packet traveling with a characteristic group velocity,
it is necessary that the range of propagation vectors included in the superposition
be fairly small. Thus, we assume that the amplitude A
k6 0 only for a small
range of values about a particular k0ofk:
A
k6 0;k0ÿ/C34<k<k0/C34; /C34 /C28k0:
The behavior in time of the wave packet is determined by the way in which the
angular frequency /C33depends upon the wave number k/C58/C33/C33
k, known as the
law of dispersion. If /C33varies slowly with k, then /C33
kcan be expanded in a power
series about k0:
/C33
k/C33
k0d/C33
dk/C12/C12/C12/C12
0
kÿk0 /C330/C330
kÿk0O
kÿk02/C104/C105
;
where
/C330/C33
k0;and /C330d/C33
dk/C12/C12/C12/C12
0
175WAVE PACKETS AND GROUP VELOCITY
Figure 4.18. Superposition of two waves.
and the subscript zero means ‘evaluated’ at kk0. Now the argument of the
exponential in Eq. (4.34) can be rewritten as
/C33tÿkx
/C330tÿk0x/C330
kÿk0tÿ
kÿk0x
/C330tÿk0x
kÿk0
/C330tÿx
and Eq. (4.34) becomes
/C32
x;texpi
k0xÿ/C330tZk0/C34
k0ÿ/C34A
kexpi
kÿk0
xÿ/C330tdk:
4:35
If we take kÿk0as the new integration variable yand assume A
kto be a slowly
varying function of kin the integration interval 2 /C34, then Eq. (4.35) becomes
/C32
x;t/C129expi
k0xÿ/C330tZk0/C34
k0ÿ/C34A
k0yexpi
xÿ/C330tydy:
Integration, transformation, and the approximation A
k0y/C129A
k0lead to
the result
/C32
x;tB
x;texpi
k0xÿ/C330t
4:36
with
B
x;t2A
k0sink
xÿ/C330t
xÿ/C330t:
4:37
As the argument of the sine contains the small quantity k;B
x;tvaries slowly
depending on time tand coordinate x. Therefore, we can regard B
x;tas the
small amplitude of an approximately monochromatic wave and k0xÿ/C330tas its
phase. If we multiply the numerator and denominator on the right hand side of
Eq. (4.37) by kand let
zk
xÿ/C330t
then B
x;tbecomes
B
x;t2A
k0ksinz
z
and we see that the variation in amplitude is determined by the factor sin ( z=z.
This has the properties
lim
z!0sinz
z1 for z0
and
sinz
z0 for z;2;...:
176FOURIER SERIES AND INTEGRALS
If we further increase the absolute value of z, the function sin ( z=zruns alter-
nately through maxima and minima, the function values of which are small com-
pared with the principal maximum at z0, and quickly converges to zero.
Therefore, we can conclude that superposition generates a wave packet whoseamplitude is non-zero only in a finite region, and is described by sin ( z=z(see Fig.
4.19).
The modulating factor sin ( z=zof the amplitude assumes the maximum value 1
asz!0. Recall that zk
xÿ/C33
0t), thus for z0, we have
xÿ/C330t0;
which means that the maximum of the amplitude is a plane propagating with
velocity
dx
dt/C330d/C33
dk/C12/C12/C12/C12
0;
that is, /C330is the group velocity, the velocity of the whole wave packet.
The concept of a wave packet also plays an important role in quantum
mechanics. The idea of associating a wave-like property with the electron and
other material particles was first proposed by Louis Victor de Broglie (1892–1987)
in 1925. His work was motivated by the mystery of the Bohr orbits. After
Rutherford’s successful -particle scattering experiments, a planetary-type
nuclear atom, with electrons orbiting around the nucleus, was in favor withmost physicists. But, according to classical electromagnetic theory, a charge
undergoing continuous centripetal acceleration emits electromagnetic radiation
177WAVE PACKETS AND GROUP VELOCITY
Figure 4.19. A wave packet.
continuously and so the electron would lose energy continuously and it would
spiral into the nucleus after just a fraction of a second. This does not occur.Furthermore, atoms do not radiate unless excited, and when radiation does
occur its spectrum consists of discrete frequencies rather than the continuum of
frequencies predicted by the classical electromagnetic theory. In 1913 Niels Bohr
(1885–1962) proposed a theory which successfully explained the radiation spectra
of the hydrogen atom. According to Bohr’s postulates, an atom can exist in
certain allowed stationary states without radiation. Only when an electron
makes a transition between two allowed stationary states, does it emit or absorb
radiation. The possible stationary states are those in which the angular momen-
tum of the electron about the nucleus is quantized, that is, m/C118rnp, where vis
the speed of the electron in the nth orbit and r is its radius. Bohr didn’t clearly
describe this quantum condition. De Broglie attempted to explain it by fitting astanding wave around the circumference of each orbit. Thus de Broglie proposed
thatn2r, where is the wavelength associated with the nth orbit. Combining
this with Bohr’s quantum condition we immediately obtain
/C104
m/C118/C104
/C112:
De Broglie proposed that any material particle of total energy Eand momentum p
is accompanied by a wave whose wavelength is given by /C104=/C112and whose
frequency is given by the Planck formula /C23/C69=/C104. Today we call these waves
de Broglie waves or matter waves. The physical nature of these matter waves was
not clearly described by de Broglie, we shall not ask what these matter waves are –
this is addressed in most textbooks on quantum mechanics. Let us ask just onequestion: what is the (phase) velocity of such a matter wave/C63 If we denote this
velocity by u, then
u/C23/C69
/C1121
/C112
/C1122c2m2
0c4q
c
1
m0c=/C1122q
c2
/C118/C112m0/C118
1ÿ/C1182=c2/C112/C32!
;
which shows that for a particle with m0/C620 the wave velocity uis always greater
than c, the speed of light in a vacuum. Instead of individual waves, de Broglie
suggested that we can think of particles inside a wave packet, synthesized from a
number of individual waves of di/C128erent frequencies, with the entire packet travel-
ing with the particle velocity /C118.
De Broglie’s matter wave idea is one of the cornerstones of quantum
mechanics.
178FOURIER SERIES AND INTEGRALS
/C72eat conduction
We now consider an application of Fourier integrals in classical physics. A semi-
infinite thin bar ( x0), whose surface is insulated, has an initial temperature
equal to f
x. The temperature of the end x0 is suddenly dropped to and
maintained at zero. The problem is to find the temperature T
x;tat any point
xat time t. First we have to set up the boundary value problem for heat conduc-
tion, and then seek the general solution that will give the temperature T
x;tat
any point xat time t.
/C72ead conduction equation
To establish the equation for heat conduction in a conducting medium we need
first to find the heat flux (the amount of heat per unit area per unit time) across a
surface. Suppose we have a flat sheet of thickness n, which has temperature Ton
one side and TTon the other side (Fig. 4.20). The heat flux which flows from
the side of high temperature to the side of low temperature is directly proportionalto the di/C128erence in temperature Tand inversely proportional to the thickness
n. That is, the heat flux from I to II is equal to
ÿ/C75T
n;
where /C75, the constant of proportionality, is called the thermal conductivity of the
conducting medium. The minus sign is due to the fact that if T/C620 the heat
actually flows from II to I. In the limit of n!0, the heat flux across from II to I
can be written
ÿ/C75/C64T
/C64nÿ/C75/C114T:
The quantity /C64T=/C64nis called the gradient of Twhich in vector form is /C114T.
We are now ready to derive the equation for heat conduction. Let /C86be an
arbitrary volume lying within the solid and bounded by surface S. The total
179HEAT CONDUCTION
n
Figure 4.20. Heat flux through a thin sheet.
amount of heat entering Sper unit time is
ZZ
S/C75/C114T
^ndS;
where ^nis an outward unit vector normal to element surface area dS. Using the
divergence theorem, this can be written as
ZZ
S/C75/C114T
^ndSZZZ
/C86/C114/C75/C114T
d/C86:
4:38
Now the heat contained in /C86is given by
ZZZ
/C86c/C26Td/C86;
where cand/C26are respectively the specific heat capacity and density of the solid.
Then the time rate of increase of heat is
/C64
/C64tZZZ
/C86c/C26Td/C86ZZZ
/C86c/C26/C64T
/C64td/C86:
4:39
Equating the right hand sides of Eqs. (4.38) and (4.39) yields
ZZZ
/C86c/C26/C64T
/C64tÿ/C114 /C75/C114T
d/C860:
Since /C86is arbitrary, the integrand (assumed continuous) must be identically zero:
c/C26/C64T
/C64t/C114 /C75/C114T
or if /C75,c,/C26are constants
/C64T
/C64tk/C114/C114 Tk/C1142T;
4:40
where k/C75=c/C26. This is the required equation for heat conduction and was first
developed by Fourier in 1822. For the semiinfinite thin bar, the boundary condi-
tions are
T
x;0f
x;T
0;t0; jT
x;tj<M;
4:41
where the last condition means that the temperature must be bounded for physicalreasons.
A solution of Eq. (4.40) can be obtained by separation of variables, that is by
letting
TX
xH
t:
Then
XH
0kX00H or X00=XH0=kH:
180FOURIER SERIES AND INTEGRALS
Each side must be a constant which we call ÿ2. (If we use 2, the resulting
solution does not satisfy the boundedness condition for real values of .) Then
X002X0;H02kH0
with the solutions
X
xA1cosxB1sinx;H
tC1eÿk2t:
A solution to Eq. (4.40) is thus given by
T
x;tC1eÿk2t
A1cosxB1sinx
eÿk2t
AcosxBsinx:
From the second of the boundary conditions (4.41) we find A0 and so T
x;t
reduces to
T
x;tBeÿk2tsinx:
Since there is no restriction on the value of , we can replace Bby a function B
and integrate over from 0 to 1and still have a solution:
T
x;tZ1
0B
eÿk2tsinxd:
4:42
Using the first of boundary conditions (4.41) we find
f
xZ1
0B
sinxd:
Then by the Fourier sine transform we find
B
2
Z1
0f
xsinxdx2
Z1
0f
usinudu
and the temperature distribution along the semiinfinite thin bar is
T
x;t2
Z1
0Z1
0f
ueÿk2tsinusinxddu:
4:43
Using the relation
sinusinx1
2cos
uÿxÿcos
ux;
Eq. (4.43) can be rewritten
T
x;t1
Z1
0Z1
0f
ueÿk2tcos
uÿxÿcos
uxddu
1
Z1
0f
uZ1
0eÿk2tcos
uÿxdÿZ1
0eÿk2tcos
uxd
du:
181HEAT CONDUCTION
Using the integral
Z1
0eÿ2cos/C12d1
2
/C114
eÿ/C122=4;
we find
T
x;t1
2
ktpZ1
0f
ueÿ
uÿx2=4ktduÿZ1
0f
ueÿ
ux2=4ktdu
:
Letting
uÿx=2ktp
/C119in the first integral and
ux=2 ktp
/C119in the second
integral, we obtain
T
x;t
1pZ1
ÿx=2
ktpeÿ/C1192f
2/C119
ktp
xd/C119ÿZ1
x=2
ktpeÿ/C1192f
2/C119 ktp
ÿxd/C119"#
:
Fourier transforms for functions of se/C118eral /C118ariables
We can extend the development of Fourier transforms to a function of several
variables, such as f
x;y;z. If we first decompose the function into a Fourier
integral with respect to x, we obtain
f
x;y;z 1
2pZ1
ÿ1/C13
/C33x;y;zei/C33xxd/C33x;
where /C13is the Fourier transform. Similarly, we can decompose the function with
respect to yandzto obtain
f
x;y;z1
22=3Z1
ÿ1/C103
/C33x;/C33y;/C33zei
/C33xx/C33yy/C33zzd/C33xd/C33yd/C33z;
with
/C103
/C33x;/C33y;/C33z1
22=3Z1
ÿ1f
x;y;zeÿi
/C33xx/C33yy/C33zzdxdydz :
We can regard /C33x;/C33y;/C33zas the components of a vector /C33whose magnitude is
/C33
/C332x/C332y/C332zq
;
then we express the above results in terms of the vector /C33:
f
r1
22=3Z1
ÿ1/C103
xeixrdx;
4:44
/C103
x1
22=3Z1
ÿ1f
reÿ
ixrdr:
4:45
182FOURIER SERIES AND INTEGRALS
/C84he Fourier integral and the delta function
The delta function is a very useful tool in physics, but it is not a function in the
usual mathematical sense. The need for this strange ‘function’ arises naturally
from the Fourier integrals. Let us go back to Eqs. (4.30) and (4.31) and substitute
/C103
/C33intof
x; we then have
f
x1
2Z1
ÿ1d/C33Z1
ÿ1dx0f
x0ei/C33
xÿx0:
Interchanging the order of integration gives
f
xZ1
ÿ1dx0f
x01
2Z1
ÿ1d/C33ei/C33
xÿx0:
4:46
If the above equation holds for any function f
x, then this tells us something
remarkable about the integral
1
2Z1
ÿ1d/C33ei/C33
xÿx0
considered as a function of x0. It vanishes everywhere except at x0x, and its
integral with respect to x0over any interval including xis unity. That is, we may
think of this function as having an infinitely high, infinitely narrow peak atxx
0. Such a strange function is called Dirac’s delta function (first introduced
by Paul A. M. Dirac):
xÿx01
2Z1
ÿ1d/C33ei/C33
xÿx0:
4:47
Equation (4.46) then becomes
f
xZ1
ÿ1f
x0
xÿx0dx0:
4:48
Equation (4.47) is an integral representation of the delta function. We summarize
its properties below:
xÿx00;ifx06x;
4:49a
Zb
a
xÿx0dx00;ifx/C62borx<a
1;ifa<x<b;/C26
4:49b
f
xZ1
ÿ1f
x0
xÿx0dx0:
4:49c
183THE FOURIER INTEGRAL AND THE DELTA FUNCTION
It is often convenient to place the origin at the singular point, in which case the
delta function may be written as
x1
2Z1
ÿ1d/C33ei/C33x:
4:50
To examine the behavior of the function for both small and large x, we use an
alternative representation of this function obtained by integrating as follows:
x1
2lim
a!1Za
ÿaei/C33xd/C33lim
a!11
2eiaxÿeÿiax
ix
lim
a!1sinax
x;
4:51
where ais positive and real. We see immediately that
ÿx
x. To examine
its behavior for small x, we consider the limit as xgoes to zero:
lim
x!0sinax
xa
lim
x!0sinax
axa
:
Thus,
0lima!1
a=!1 , or the amplitude becomes infinite at the singu-
larity. For large jxj, we see that sin
ax=xoscillates with period 2 =a, and its
amplitude falls o/C128 as 1 =jxj. But in the limit as agoes to infinity, the period
becomes infinitesimally narrow so that the function approaches zero everywhere
except for the infinite spike of infinitesimal width at the singularity. What is the
integral of Eq. (4.51) over all space/C63
Z1
ÿ1lim
a!1sinax
xdxlim
a!12
Z1
0sinax
xdx2
21:
Thus, the delta function may be thought of as a spike function which has unit area
but a non-zero amplitude at the point of singularity, where the amplitude becomes
infinite. No ordinary mathematical function with these properties exists. How do
we end up with such an improper function/C63 It occurs because the change of order
of integration in Eq. (4.46) is not permissible. In spite of this, the Dirac deltafunction is a most convenient function to use symbolically. For in applications the
delta function always occurs under an integral sign. Carrying out this integration,
using the formal properties of the delta function, is really equivalent to inverting
the order of integration once more, thus getting back to a mathematically correct
expression. Thus, using Eq. (4.49) we have
Z
1
ÿ1f
x
xÿx0dxf
x0;
but, on substituting Eq. (4.47) for the delta function, the integral on the left hand
side becomes
Z1
ÿ1f
x1
2pZ1
ÿ1d/C33ei/C33
xÿx0/C26/C27
dx
184FOURIER SERIES AND INTEGRALS
or, using the property
ÿx
x,
Z1
ÿ1f
x1
2pZ1
ÿ1d/C33eÿi/C33
xÿx0/C26/C27
dx
andchanging the order of integration , we have
Z1
ÿ1f
x1 2pZ
1
ÿ1d/C33eÿi/C33x/C26/C27
ei/C33x0dx:
Comparing this expression with Eqs. (4.30) and (4.31), we see at once that this
double integral is equal to f
x0, the correct mathematical expression.
It is important to keep in mind that the delta function cannot be the end result of a
calculation and has meaning only so long as a subse/C113uent integration over its argu/C45
ment is carried out .
We can easily verify the following most frequently required properties of the
delta function:
Ifa<b
Zb
af
x
xÿx0dxf
x0;ifa<x0<b
0; ifx0<aorx0<b/C26
;
4:52a
ÿx
x;
4:52b
0
xÿ 0
ÿx;0
xd
x=dx;
4:52c
x
x0;
4:52d
axaÿ1
x;a/C620;
4:52e
x2ÿa2
2aÿ1
xÿa
xa ;a/C620;
4:52f
Z
aÿx
xÿbdx
aÿb;
4:52g
f
x
xÿaf
a
xÿa:
4:52h
Each of the first six of these listed properties can be established by multiplying
both sides by a continuous, di/C128erentiable function f
xand then integrating over
x. For example, multiplying x0
xbyf
xand integrating over xgives
Z
f
xx0
xdxÿZ
xd
dxxf
xdx
ÿZ
xf
xxf0
x/C2/C3
dxÿZ
f
x
xdx:
Thus x
xhas the same e/C128ect when it is a factor in an integrand as has ÿ
x.
185THE FOURIER INTEGRAL AND THE DELTA FUNCTION
Parse/C118al/C39s identit/C121 for Fourier integrals
We arrived earlier at Parseval’s identity for Fourier series. An analogy exists for
Fourier integrals. If /C103
and/C71
are Fourier transforms of f
xand/C70
x
respectively, we can show that
Z1
ÿ1f
x/C70/C42
xdx1
2Z1
ÿ1/C103
/C71/C42
d;
4:54
where /C70/C42
xis the complex conjugate of /C70
x. In particular, if /C70
xf
xand
hence /C71
/C103
, then we have
Z1
ÿ1f
xjj2dxZ1
ÿ1/C103
jj d:
4:54
Equation (4.53), or the more general Eq. (4.54), is known as the Parseval’s iden-tity for Fourier integrals. Its proof is straightforward:
Z
1
ÿ1f
x/C70/C42
xdxZ1
ÿ11
2pZ1
ÿ1/C103
eÿixd
1 2pZ
1
ÿ1/C71/C42
0ei0xd0
dx
Z1
ÿ1dZ1
ÿ1d0/C103
/C71/C42
01
2Z1
ÿ1eix
ÿ0dx
Z1
ÿ1d/C103
Z1
ÿ1d0/C71/C42
0
0ÿZ1
ÿ1/C103
/C71/C42
d:
Parseval’s identity is very useful in understanding the physical interpretation of
the transform function /C103
when the physical significance of f
xis known. The
following example will show this.
Example 4.10
Consider the following function, as shown in Fig. 4.21, which might represent thecurrent in an antenna, or the electric field in a radiated wave, or displacement of a
damped harmonic oscillator:
f
t0 t<0
e
ÿt=Tsin/C330tt /C620:/C26
186FOURIER SERIES AND INTEGRALS
Its Fourier transform /C103
/C33is
/C103
/C331
2pZ1
ÿ1f
teÿi/C33tdt
1 2pZ
1
ÿ1eÿt=Teÿi/C33tsin/C330tdt
1
2 2p 1
/C33/C330ÿi=Tÿ1
/C33ÿ/C330ÿi=T
:
Iff
tis a radiated electric field, the radiated power is proportional to jf
tj2
and the total energy radiated is proportional toR1
0f
tjj2dt. This is equal toR1
0/C103
/C33jj2d/C33by Parseval’s identity. Then j/C103
/C33j2must be the energy radiated
per unit frequency interval.
Parseval’s identity can be used to evaluate some definite integrals. As an exam-
ple, let us revisit Example 4.8, where the given function is
f
x1xjj<a
0xjj/C62a(
and its Fourier transform is
/C103
/C33
2
/C114
sin/C33a
/C33:
By Parseval’s identity, we have
Z1
ÿ1f
xfg2dxZ1
ÿ1/C103
/C33fg2d/C33:
187PARSEVAL’S IDENTITY FOR FOURIER INTEGRALS
Figure 4.21. A damped sine wave.
This is equivalent to
Za
ÿa
12dxZ1
ÿ12
sin2/C33a
/C332d/C33;
from which we find
Z1
02
sin2/C33a
/C332d/C33a
2:
/C84he con/C118olution theorem for Fourier transforms
The convolution of the functions f
xandH
x, denoted by f/C3H, is defined by
f/C3HZ1
ÿ1f
uH
xÿudu:
4:55
If/C103
/C33and/C71
/C33are Fourier transforms of f
xandH
xrespectively, we can
show that
1
2Z1
ÿ1/C103
/C33/C71
/C33ei/C33xd/C33Z1
ÿ1f
uH
xÿudu:
4:56
This is known as the convolution theorem for Fourier transforms. It means that
the Fourier transform of the product /C103
/C33/C71
/C33, the left hand side of Eq. (55), is
the convolution of the original function.
The proof is not dicult. We have, by definition of the Fourier transform,
/C103
/C331
2pZ1
ÿ1f
xeÿi/C33xdx; /C71
/C331 2pZ
1
ÿ1H
x0eÿi/C33x0dx0:
Then
/C103
/C33/C71
/C331
2Z1
ÿ1Z1
ÿ1f
xH
x0eÿi/C33
xx0dxdx0:
4:57
Letxx0uin the double integral of Eq. (4.57) and we wish to transform from
(x,x0)t o( x;u). We thus have
dxdx0/C64
x;x0
/C64
x;ududx ;
where the Jacobian of the transformation is
/C64
x;x0
/C64
x;u/C64x
/C64x/C64x
/C64u
/C64x0
/C64x/C64x0
/C64u/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C1210
01/C12/C12/C12/C12/C12/C12/C12/C121:
188FOURIER SERIES AND INTEGRALS
Thus Eq. (4.57) becomes
/C103
/C33/C71
/C331
2Z1
ÿ1Z1
ÿ1f
xH
uÿxeÿi/C33udxdu
1
2Z1
ÿ1eÿi/C33uZ1
ÿ1f
xH
uÿxdu
dx
/C70Z1
ÿ1f
xH
uÿxdu/C26/C27
/C70f/C3Hfg :
4:58
From this we have equivalently
f/C3H/C70ÿ1/C103
/C33/C71
/C33 fg
1=2Z1
ÿ1ei/C33x/C103
/C33/C71
/C33;
which is Eq. (4.56).
Equation (4.58) can be rewritten as
/C70ffg/C70Hfg/C70f/C3Hfg
/C103/C70ffg;/C71/C70Hfg ;
which states that the /C70ourier transform of the convolution of f(x) and /C72(x) is e/C113ual
to the product of the /C70ourier transforms of f(x) and /C72(x) . This statement is often
taken as the convolution theorem.
The convolution obeys the commutative, associative and distributive laws of
algebra that is, if we have functions f1;f2;f3then
f1/C3f2f2/C3f1 commutative;
f1/C3
f2/C3f3
f1/C3f2/C3f3 associative;
f1/C3
f2f3f1/C3f2f1/C3f3distributive :9
>>=
>>;
4:59
It is not dicult to prove these relations. For example, to prove the commutative
law, we first have
f1/C3f2Z1
ÿ1f1
uf2
xÿudu:
Now let xÿu/C118, then
f1/C3f2Z1
ÿ1f1
uf2
xÿudu
Z1
ÿ1f1
xÿ/C118f2
/C118d/C118f2/C3f1:
Example 4.11Solve the integral equation y
xf
xR
1
ÿ1y
ur
xÿudu, where f
xand
r
xare given, and the Fourier transforms of y
x;f
xandr
xexist.
189THE CONVOLUTION THEOREM FOR FOURIER TRANSFORMS
Solution: Let us denote the Fourier transforms of y
x;f
xand r
xby
/C89
/C33;/C70
/C33;and /C82
/C33respectively. Taking the Fourier transform of both sides
of the given integral equation, we have by the convolution theorem
Y
/C33/C70
/C33Y
/C33R
/C33orY
/C33/C70
/C33
1ÿR
/C33:
/C67alculations of Fourier transforms
Fourier transforms can often be used to transform a di/C128erential equation which is
dicult to solve into a simpler equation that can be solved relatively easy. In
order to use the transform methods to solve first- and second-order di/C128erential
equations, the transforms of first- and second-order derivatives are needed. By
taking the Fourier transform with respect to the variable x, we can show that
a/C70/C64u
/C64x
i/C70
u;
b/C70/C642u
/C64x2/C32!
ÿ2/C70
u;
c/C70/C64u
/C64t
/C64
/C64t/C70
u:9
>>>>>>>>>=
>>>>>>>>>;
4:60
Proof: (a) By definition we have
/C70
/C64u
/C64x
Z1
ÿ1/C64u
/C64xeÿixdx;
where the factor 1 =
2p
has been dropped. Using integration by parts, we obtain
/C70/C64u
/C64x
Z1
ÿ1/C64u
/C64xeÿixdx
ueÿix/C12/C12/C12/C121
ÿ1iZ1
ÿ1ueÿixdx
i/C70
u:
(b) Let u/C64/C118=/C64xin (a), then
/C70/C642/C118
/C64x2/C32!
i/C70/C64/C118
/C64x
i2/C70
/C118:
Now if we formally replace /C118byuwe have
/C70/C642u
/C64x2/C32!
ÿ2/C70
u;
190FOURIER SERIES AND INTEGRALS
provided that uand/C64u=/C64x!0a sx! 1 . In general, we can show that
/C70/C64nu
/C64xn
in/C70
u
ifu;/C64u=/C64x;...;/C64nÿ1u=/C64xnÿ1! 1 asx! 1 .
(c) By definition
/C70/C64u
/C64t
Z1
ÿ1/C64u
/C64teÿixdx/C64
/C64tZ1
ÿ1ueÿixdx/C64
/C64t/C70
u:
Example 4.12
Solve the inhomogeneous di/C128erential equation
d2
dx2/C112d
dx/C113/C32!
f
xR
x;ÿ1 x1 ;
where pand/C113are constants.
Solution: We transform both sides
/C70d2f
dx2/C112df
dx/C113f()
i2/C112
i/C113/C70f
xfg
/C70R
xfg :
If we denote the Fourier transforms of f
xandR
xby/C103
and/C71
, respec-
tively,
/C70f
xfg /C103
;/C70R
xfg /C71
;
we have
ÿ2i/C112/C113/C103
/C71
;or/C103
/C71
=
ÿ2i/C112/C113
and hence
f
x1
2pZ1
ÿ1eix/C103
d
1 2pZ
1
ÿ1eix /C71
ÿ2i/C112/C113d:
We will not gain anything if we do not know how to evaluate this complex
integral. This is not a dicult problem in the theory of functions of complex
variables (see Chapter 7).
191CALCULATIONS OF FOURIER TRANSFORMS
/C84he delta function and the /C71reen/C39s function method
The Green’s function method is a very useful technique in the solution of partial
di/C128erential equations. It is usually used when boundary conditions, rather than
initial conditions, are specified. To appreciate its usefulness, let us consider the
inhomogeneous di/C128erential equation
L
xf
xÿf
xR
x
4:61
over a domain /C68, with Lan arbitrary di/C128erential operator, and a given con-
stant. Suppose we can expand f
xandR
xin eigenfunctions unof the operator
L
Lunnun:
f
xX
ncnun
x;R
xX
ndnun
x:
Substituting these into Eq. (4.61) we obtain
X
ncn
nÿun
xX
ndnun
x:
Since the eigenfunctions un
xare linearly independent, we must have
cn
nÿdnorcndn=
nÿ:
Moreover,
dnZ
Dun/C42R
xdx:
Now we may write cnas
cn1
nÿZ
Dun/C42R
xdx;
therefore
f
xX
nun
nÿZ
Dun/C42
x0R
x0dx0:
This expression may be written in the form
f
xZ
D/C71
x;x0R
x0dx0;
4:62
where /C71
x;x0is given by
/C71
x;x0X
nun
xun/C42
x0
nÿ
4:63
and is called the Green’s function. Some authors prefer to write /C71
x;x0;to
emphasize the dependence of /C71onas well as on xandx0.
192FOURIER SERIES AND INTEGRALS
What is the di/C128erential equation obeyed by /C71
x;x0/C63 Suppose f
x0in Eq.
(4.62) is taken to be
x0ÿx0), then we obtain
f
xZ
D/C71
x;x0
x0ÿx0dx/C71
x;x0:
Therefore /C71
x;x0is the solution of
L/C71
x;x0ÿ/C71
x;x0
xÿx0;
4:64
subject to the appropriate boundary conditions. Eq. (4.64) shows clearly that the
/C71reen/C39s function is the solution of the problem for a unit point /C96source/C39
R
x
xÿx0.
Example 4.13
Find the solution to the di/C128erential equation
d2u
dx2ÿk2uf
x
4:65
on the interval 0 x/C108, with u
0u
/C1080, for a general function f
x.
Solution: We first solve the di/C128erential equation which /C71
x;x0obeys:
d2/C71
x;x0
dx2ÿk2/C71
x;x0
xÿx0:
4:66
Forxequal to anything but x0(that is, for x<x0orx/C62x0),
xÿx00a n d
we have
d2/C71<
x;x0
dx2ÿk2/C71<
x;x00
x<x0;
d2/C71/C62
x;x0
dx2ÿk2/C71/C62
x;x00
x/C62x0:
Therefore, for x<x0
/C71<AekxBeÿkx:
By the boundary condition u
00 we find AB0, and /C71<reduces to
/C71<A
ekxÿeÿkx;
4:67a
similarly, for x/C62x0
/C71/C62CekxDeÿkx:
193THE DELTA FUNCTION AND THE GREEN’S FUNCTION METHOD
By the boundary condition u
/C1080 we find Cek/C108Deÿk/C1080, and /C71/C62can be
rewritten as
/C71/C62C0ek
xÿ/C108ÿeÿk
xÿ/C108;
4:67b
where C0Cek/C108.
How do we determine the constants AandC0/C63 First, continuity of /C71atxx0
gives
A
ekxÿeÿkxC0
ek
xÿ/C108ÿeÿk
xÿ/C108:
4:68
A second constraint is obtained by integrating Eq. (4.61) from x0ÿ/C34tox0/C34,
where /C34is infinitesimal:
Zx0/C34
x0ÿ/C34d2/C71
dx2ÿk2/C71"#
dxZx0/C34
x0ÿ/C34
xÿx0dx1:
4:69
But
Zx0/C34
x0ÿ/C34k2/C71dxk2
/C71/C62ÿ/C71<0;
where the last step is required by the continuity of /C71. Accordingly, Eq. (4.64)
reduces to
Zx0/C34
x0ÿ/C34d2/C71
dx2dxd/C71/C62
dxÿd/C71<
dx1:
4:70
Now
d/C71<
dx xx0/C12/C12/C12/C12/C12Ak
ekx0eÿkx0
and
d/C71/C62
dx xx0/C12/C12/C12/C12/C12C
0kek
x0ÿ/C108eÿk
x0ÿ/C108:
Substituting these into Eq. (4.70) yields
C0k
ek
x0ÿ/C108eÿk
x0ÿ/C108ÿAk
ekx0eÿkx01:
4:71
We can solve Eqs. (4.68) and (4.71) for the constants AandC0. After some
algebraic manipulation, the solution is
A1
2ksinhk
x0ÿ/C108
sinhk/C108;C01
2ksinhkx0
sinhk/C108
194FOURIER SERIES AND INTEGRALS
and the Green’s function is
/C71
x;x01
ksinhkxsinhk
x0ÿ/C108
sinhk/C108;
4:72
which can be combined with f
xto obtain u
x:
u
xZ/C108
0/C71
x;x0f
x0dx0:
Problems
4.1 ( a) Find the period of the function f
xcos
x=3cos
x=4.
(b) Show that, if the function f
tcos/C331tcos/C332tis periodic with a
period T, then the ratio /C331=/C332must be a rational number.
4.2 Show that if f
xPf
x, then
ZaP=2
aÿP=2f
xdxZP=2
ÿP=2f
xdx;ZPx
Pf
xdxZx
0f
xdx:
4.3 ( a) Using the result of Example 4.2, prove that
1ÿ1
315ÿ17ÿ
4:
(b) Using the result of Example 4.3, prove that
1
13ÿ1
351
57ÿÿ2
4:
4.4 Find the Fourier series which represents the function f
xjxjin the inter-
valÿx.
4.5 Find the Fourier series which represents the function f
xxin the interval
ÿx.
4.6 Find the Fourier series which represents the function f
xx2in the inter-
valÿx.
4.7 Represent f
xx;0<x<2, as: ( a) in a half-range sine series, ( b) a half-
range cosine series.
4.8 Represent f
xsinx,0<x<, as a Fourier cosine series.
4.9 ( a) Show that the function f
xof period 2 which is equal to xon
ÿ1;1
can be represented by the following Fourier series
ÿi
eixÿeÿixÿ12e
2ix12e
ÿ2ix13e
3ixÿ13e
ÿ3ix
:
(b) Write Parseval’s identity corresponding to the Fourier series of ( a).
(c) Determine from ( b) the sum Sof the series 1 1
419P1
n11=n2.
195PROBLEMS
4.10 Find the exponential form of the Fourier series of the function whose defi-
nition in one period is f
xeÿx;ÿ1<x<1.
4.11 ( a) Show that the set of functions
1;sinx
L;cosx
L;sin2x
L;cos2x
L;sin3x
L;cos3x
L;...
form an orthogonal set in the interval
ÿL;L.
(b) Determine the corresponding normalizing constants for the set in ( a)s o
that the set is orthonormal in
ÿL;L.
4.12 Express f
x;yxyas a Fourier series for 0 x1;0y2.
4.13 Steady-state heat conduction in a rectangular plate: Consider steady-state
heat conduction in a flat plate having temperature values prescribed on the
sides (Fig. 4.22). The boundary value problem modeling this is:
/C642u
/C642x2/C642u
/C642y20; 0<x< ; 0<y</C12 ;
u
x;0u
x;/C120; 0<x< ;
u
0;y0;u
;yT; 0<y</C12 :
Determine the temperature at any point of the plate.
4.14 Derive and solve the following eigenvalue problem which occurs in the
theory of a vibrating square membrane whose sides, of length L, are kept
fixed:
/C642/C119
/C64x2/C642/C119
/C64y2/C1190;
/C119
0;y/C119
L;y0
0yL;
/C119
x;0/C119
x;L0
0yL:
196FOURIER SERIES AND INTEGRALS
Figure 4.22. Flat plate with prescribed temperature.
4.15 Show that the Fourier integral can be written in the form
f
x1
Z1
0d/C33Z1
ÿ1f
x0cos/C33
xÿx0dx0:
4.16 Starting with the form obtained in Problem 4.15, show that the Fourier
integral can be written in the form
f
xZ1
0A
/C33cos/C33xB
/C33sin/C33x fg d/C33;
where
A
/C331
Z1
ÿ1f
xcos/C33xd x; B
/C331
Z1
ÿ1f
xsin/C33xd x:
4.17 ( a) Find the Fourier transform of
f
x1ÿx2jxj<1
0 jxj/C621:(
(b) Evaluate
Z1
0xcosxÿsinx
x3cosx
2dx:
4.18 ( a) Find the Fourier cosine transform of f
xeÿmx;m/C620.
(b) Use the result in ( a) to show that
Z1
0cos/C112x
x22dx
2eÿ/C112
/C112/C620;/C62 0:
4.19 Solve the integral equation
Z1
0f
xsinxd x1ÿ 01
0 /C621/C26
:
4.20 Find a bounded solution to Laplace’s equation /C1142u
x;y0 for the half-
plane y/C620i futakes on the value of f(x) on the x-axis:
/C642u
/C64x2/C642u
/C64y20; u
x;0f
x; u
x;y jj <M:
4.21 Show that the following two functions are valid representations of the delta
function, where /C34is positive and real:
a
x1plim
/C34!01/C34peÿx2=/C34
b
x1
lim
/C34!0/C34
x2/C342:
197PROBLEMS
4.22 Verify the following properties of the delta function:
(a)
x
ÿx,
(b)x
x0,
(c)0
ÿxÿ 0
x,
(d)x0
xÿ
x,
(e)c
cx
x;c/C620.
4.23 Solve the integral equation for y
x
Z1
ÿ1y
udu
xÿu2a21
x2b20<a<b:
4.24 Use Fourier transforms to solve the boundary value problem
/C64u
/C64tk/C642u
/C64x2; u
x;0f
x; u
x;t jj <M;
where ÿ1 <x<1;t/C620.
4.25 Obtain a solution to the equation of a driven harmonic oscillator
/C127x
t2/C12_x
t/C332
0x
t0R
t;
where /C12and/C330are positive and real constants.
198FOURIER SERIES AND INTEGRALS
5
Linear vector spaces
Linear vector space is to quantum mechanics what calculus is to classical
mechanics. In this chapter the essential ideas of linear vector spaces will be dis-
cussed. The reader is already familiar with vector calculus in three-dimensional
Euclidean space /C693(Chapter 1). We therefore present our discussion as a general-
ization of elementary vector calculus. The presentation will be, however, slightlyabstract and more formal than the discussion of vectors in Chapter 1. Any reader
who is not already familiar with this sort of discussion should be patient with the
first few sections. You will then be amply repaid by finding the rest of this chapter
relatively easy reading.
/C69uclidean n-space E
n
In the study of vector analysis in /C693, an ordered triple of numbers ( a1,a2,a3) has
two di/C128erent geometric interpretations. It represents a point in space, with a1,a2,
a3being its coordinates; it also represents a vector, with a1,a2, and a3being its
components along the three coordinate axes (Fig. 5.1). This idea of using triples of
numbers to locate points in three-dimensional space was first introduced in the
mid-seventeenth century. By the latter part of the nineteenth century physicists
and mathematicians began to use the quadruples of numbers ( a1,a2,a3,a4)a s
points in four-dimensional space, quintuples ( a1,a2,a3,a4,a5) as points in five-
dimensional space etc. We now extend this to n-dimensional space /C69n, where nis a
positive integer. Although our geometric visualization doesn’t extend beyond
three-dimensional space, we can extend many familiar ideas beyond three-dimen-
sional space by working with analytic or numerical properties of points and
vectors rather than their geometric properties.
For two- or three-dimensional space, we use the terms ‘ordered pair’ and
‘ordered triple.’ When n/C623, we use the term ‘ordered- n-tuplet’ for a sequence
199
ofnnumbers, real or complex, ( a1,a2,a3;...;an); they will be viewed either as a
generalized point or a generalized vector in a n-dimensional space /C69n.
Two vectors u
u1;u2;...;unand/C118
/C1181;/C1182;...;/C118nin/C69nare called equal if
ui/C118i;i1;2;...;n
5:1
The sum u/C118is defined by
u/C118
u1/C1181;u2/C1182;...;un/C118n
5:2
and if kis any scalar, the scalar multiple kuis defined by
ku
ku1;ku2;...;kun:
5:3
Ifu
u1;u2;...;unis any vector in /C69n, its negative is given by
ÿu
ÿ u1;ÿu2;...;ÿun
5:4
and the subtraction of vectors in /C69ncan be considered as addition: /C118ÿu
/C118
ÿ u. The null (zero) vector in /C69nis defined to be the vector 0
0;0;...;0.
The addition and scalar multiplication of vectors in /C69nhave the following
arithmetic properties:
u/C118/C118u;
5:5a
u
/C118/C119
u/C118/C119;
5:5b
u/C48/C48uu;
5:5c
a
bu
abu;
5:5d
a
u/C118aua/C118;
5:5e
abuaubu;
5:5f
where u,/C118,/C119are vectors in /C69nand aand bare scalars.
200LINEAR VECTOR SPACES
Figure 5.1. A space point Pwhose position vector is /C65.
We usually define the inner product of two vectors in /C693in terms of lengths of
the vectors and the angle between the vectors: /C65/C66ABcos; /C128
/C65;/C66.W e
do not define the inner product in /C69nin the same manner. However, the inner
product in /C693has a second equivalent expression in terms of components:
/C65/C66A1B1A2B2A3B3. We choose to define a similar formula for the gen-
eral case. We made this choice because of the further generalization that will be
outlined in the next section. Thus, for any two vectors u
u1;u2;...;unand
/C118
/C1181;/C1182;...;/C118nin/C69n, the inner (or dot) product u/C118is defined by
u/C118u1/C42/C1181u2/C42/C1182 un/C42/C118n
5:6
where the asterisk denotes complex conjugation. uis often called the prefactor
and/C118the post-factor. The inner product is linear with respect to the post-factor,
and anti-linear with respect to the prefactor:
u
a/C118b/C119au/C118bu/C119;
aub/C118/C119a/C42
u/C118b/C42
u/C119:
We expect the inner product for the general case also to have the following three
main features:
u/C118
/C118u/C42
5:7a
u
a/C118b/C119au/C118bu/C119
5:7b
uu0
0;if and only if u0:
5:7c
Many of the familiar ideas from /C692and/C693have been carried over, so it is
common to refer to /C69nwith the operations of addition, scalar multiplication, and
with the inner product that we have defined here as Euclidean n-space.
/C71eneral linear /C118ector spaces
We now generalize the concept of vector space still further: a set of ‘objects’ (or
elements) obeying a set of axioms, which will be chosen by abstracting the most
important properties of vectors in /C69n, forms a linear vector space /C86nwith the
objects called vectors. Before introducing the requisite axioms, we first adapt a
notation for our general vectors: general vectors are designated by the symbol ji,
which we call, following Dirac, ket vectors; the conjugates of ket vectors aredenoted by the symbol hj, the bra vectors. However, for simplicity, we shall
refer in the future to the ket vectors jisimply as vectors, and to the hjsa s
conjugate vectors. We now proceed to define two basic operations on these
vectors: addition and multiplication by scalars.
By addition we mean a rule for forming the sum, denoted j/C32
1ij/C322i, for
any pair of vectors j/C321iandj/C322i.
By scalar multiplication we mean a rule for associating with each scalar k
and each vector j/C32ia new vector kj/C32i.
201GENERAL LINEAR VECTOR SPACES
We now proceed to generalize the concept of a vector space. An arbitrary set of
nobjects j1i;j2i;j3i;...;j/C30i;...;j’iform a linear vector /C86nif these objects, called
vectors, meet the following axioms or properties:
A.1 If /C30jiand ’jiare objects in /C86nandkis a scalar, then /C30ji’jiandk/C30jiare
in/C86n, a feature called closure.
A.2 /C30ji’ji’ji/C30ji; that is, addition is commutative.
A.3 ( /C30ji’ji /C32ji/C30ji
’ji/C32ji); that is, addition is associative.
A.4k
/C30ji’ji k/C30jik’ji; that is, scalar multiplication is distributive in the
vectors.
A.5
k/C30jik/C30ji/C30ji; that is, scalar multiplication is distributive in the
scalars.
A.6k
/C30ji k/C30ji; that is, scalar multiplication is associative.
A.7 There exists a null vector 0 jiin/C86nsuch that /C30ji0ji/C30jifor all /C30jiin/C86n.
A.8 For every vector /C30jiin/C86n, there exists an inverse under addition, ÿ/C30ji such
that /C30jiÿ /C30ji 0ji.
The set of numbers a;b;...used in scalar multiplication of vectors is called the
field over which the vector field is defined. If the field consists of real numbers, we
have a real vector field; if they are complex, we have a complex field. /C78ote that the
vectors themselves are neither real nor complex/C44 the nature of the vectors is notspeci/C174ed. /C86ectors can be any kinds of objects/C59 all that is re/C113uired is that the vector
space axioms be satis/C174ed . Thus we purposely do not use the symbol /C86to denote
the vectors as the first step to turn the reader away from the limited concept of the
vector as a directed line segment. Instead, we use Dirac’s ket and bra symbols, ji
and jh, to denote generic vectors.
The familiar three-dimensional space of position vectors /C69
3is an example of
a vector space over the field of real numbers. Let us now examine two simpleexamples.
Example 5.1
Let/C86be any plane through the origin in /C69
3. We wish to show that the points in
the plane /C86form a vector space under the addition and scalar multiplication
operations for vector in /C693.
Solution: Since E3itself is a vector space under the addition and scalar multi-
plication operations, thus Axioms A.2, A.3, A.4, A.5, and A.6 hold for all points
inE3and consequently for all points in the plane /C86. We therefore need only show
that Axioms A.1, A.7, and A.8 are satisfied.
Now the plane /C86, passing through the origin, has an equation of the form
ax1bx2cx30:
202LINEAR VECTOR SPACES
Hence, if u
u1;u2;u3and/C118
/C1181;/C1182;/C1183are points in /C86, then we have
au1bu2cu30a n d a/C1181b/C1182c/C11830:
Addition gives
a
u1/C1181b
u2/C1182c
u3/C11830;
which shows that the point u/C118also lies in the plane /C86. This proves that Axiom
A.1 is satisfied. Multiplying au1bu2cu30 through by ÿ1 gives
a
ÿu1b
ÿu2c
ÿu30;
that is, the point ÿu
ÿ u1;ÿu2;ÿu3lies in /C86. This establishes Axiom A.8.
The verification of Axiom A.7 is left as an exercise.
Example 5.2
Let/C86be the set of all mnmatrices with real elements. We know how to add
matrices and multiply matrices by scalars. The corresponding rules obey closure,associativity and distributive requirements. The null matrix has all zeros in it, and
the inverse under matrix addition is the matrix with all elements negated. Thus the
set of all mnmatrices, together with the operations of matrix addition and
scalar multiplication, is a vector space. We shall denote this vector space by thesymbol M
mn.
/C83ubspaces
Consider a vector space /C86.I f/C87is a subset of /C86and forms a vector space under
the addition and scalar multiplication, then /C87is called a subspace of /C86. For
example, lines and planes passing through the origin form vector spaces andthey are subspaces of /C69
3.
Example 5.3We can show that the set of all 2 2 matrices having zero on the main diagonal is
a subspace of the vector space M
22of all 2 2 matrices.
Solution: To prove this, let
~X0x12
x21 0/C32!
~Y0y12
y21 0/C32!
be two matrices in /C87and kany scalar. Then
k~X0x12
kx21 0/C32!
and ~X~Y0 x12y12
x21y21 0/C32!
and thus they lie in /C87. We leave the verification of other axioms as exercises.
203SUBSPACES
Linear combination
A vector /C87jiis a linear combination of the vectors /C1181ji;/C1182ji;...;/C118rjiif it can be
expressed in the form
/C87jik1j/C1181ik2j/C1182i krj/C118ri;
where k1;k2;...;krare scalars. For example, it is easy to show that the vector
/C87ji
9;2;7in/C693is a linear combination of /C1181ji
1;2;ÿ1and
/C1182ji
6;4;2. To see this, let us write
9;2;7k1
1;2;ÿ1k2
6;4;2
or
9;2;7
k16k2;2k14k2;ÿk12k2:
Equating corresponding components gives
k16k29;2k14k22;ÿk12k27:
Solving this system yields k1ÿ3 and k22 so that
/C87jiÿ3/C1181ji2/C1182ji:
Linear independence/C44 bases/C44 and dimensionalit/C121
Consider a set of vectors 1 ji;2ji;...;rji;...njiin a linear vector space /C86. If every
vector in /C86is expressible as a linear combination of 1 ji;2ji;...;rji;...;nji, then
we say that these vectors span the vector space /C86, and they are called the base
vectors orbasis of the vector space /C86. For example, the three unit vectors
e1
1;0;0;e2
0;1;0, and e3
0;0;1span /C693because every vector in /C693
is expressible as a linear combination of e1,e2, and e3. But the following three
vectors in /C693do not span /C693/C581ji
1;1;2;2ji
1;0;1, and 3 ji
2;1;3.
Base vectors are very useful in a variety of problems since it is often possible to
study a vector space by first studying the vectors in a base set, then extending the
results to the rest of the vector space. Therefore it is desirable to keep the spanning
set as small as possible. Finding the spanning sets for a vector space depends upon
the notion of linear independence.
We say that a finite set of nvectors 1 ji;2ji;...;rji;...;nji, none of which is a
null vector, is linearly independent if no set of non-zero numbers akexists such
that
Xn
k1akkijj0i:
5:8
In other words, the set of vectors is linearly independent if it is impossible to
construct the null vector from a linear combination of the vectors except when all
204LINEAR VECTOR SPACES
the coecients vanish. For example, non-zero vectors 1 jiand 2jiof/C692that lie
along the same coordinate axis, say x1, are not linearly independent, since we can
write one as a multiple of the other: 1 jia2ji, where ais a scalar which may be
positive or negative. That is, 1 jiand 2jidepend on each other and so they are not
linearly independent. Now let us move the term a2jito the left hand side and the
result is the null vector: 1 jiÿa2ji0ji. Thus, for these two vectors 1 jiand 2jiin
/C692, we can find two non-zero numbers (1, ÿasuch that Eq. (5.8) is satisfied, and
so they are not linearly independent.
On the other hand, the nvectors 1 ji;2ji;...;rji;...;njiare linearly dependent if
it is possible to find scalars a1;a2;...;an, at least two of which are non-zero, such
that Eq. (5.8) is satisfied. Let us say a960. Then we could express 9 jiin terms of
the other vectors
9jiXn
i1;69ÿai
a9iji:
That is, the nvectors in the set are linearly dependent if any one of them can be
expressed as a linear combination of the remaining nÿ1 vectors.
Example 5.4
The set of three vectors 1 ji
2;ÿ1;0;3,2ji
1;2;5;ÿ1;3ji
7;ÿ1;5;8is
linearly dependent, since 3 1 ji 2ji ÿ 3ji 0ji.
Example 5.5The set of three unit vectors e
1ji
1;0;0;e2ji
0;1;0, and e3ji
0;0;1in
/C693is linearly independent. To see this, let us start with Eq. (5.8) which now takes
the form
a1e1jia2e2jia3e3ji0ji
or
a1
1;0;0a2
0;1;0a3
0;0;1
0;0;0
from which we obtain
a1;a2;a3
0;0;0;
the set of three unit vectors e1ji;e2ji, and e3jiis therefore linearly independent.
Example 5.6
The set Sof the following four matrices
1ji10
00
;2ji0100
;j3i0010
;4ji0001
;
205LINEAR INDEPENDENCE, BASES, AND DIMENSIONALITY
is a basis for the vector space M22of 22 matrices. To see that Sspans M22, note
that a typical 2 2 vector (matrix) can be written as
ab
cd/C32!
a10
00/C32!
b0100/C32!
c0010/C32!
d0001/C32!
a1jib2jic3jid4ji:
To see that Sis linearly independent, assume that
a1jib2jic3jid4ji0ji;
that is,
a10
00
b0100
c0010
d0001
0000
;
from which we find abcd0 so that Sis linearly independent.
We now come to the dimensionality of a vector space. We think of space
around us as three-dimensional. How do we extend the notion of dimension to
a linear vector space/C63 Recall that the three-dimensional Euclidean space /C69
3is
spanned by the three base vectors: e1
1;0;0,e2
0;1;0,e3
0;0;1.
Similarly, the dimension nof a vector space /C86is defined to be the number nof
linearly independent base vectors that span the vector space /C86. The vector space
will be denoted by /C86n
Rif the field is real and by /C86n
Cif the field is complex.
For example, as shown in Example 5.6, 2 2 matrices form a four-dimensional
vector space whose base vectors are
1ji10
00
;2ji0100
;3ji0010
;4ji0001
;
since any arbitrary 2 2 matrix can be written in terms of these:
ab
cd
a1jib2jic3jid4ji:
If the scalars a;b;c;dare real, we have a real four-dimensional space, if they are
complex we have a complex four-dimensional space.
Inner product spaces (unitar/C121 spaces)
In this section the structure of the vector space will be greatly enriched by the
addition of a numerical function, the inner product (or scalar product). Linear
vector spaces in which an inner product is defined are called inner-product spaces
(or unitary spaces). The study of inner-product spaces enables us to make a real
juncture with physics.
206LINEAR VECTOR SPACES
In our earlier discussion, the inner product of two vectors in /C69nwas defined by
Eq. (5.6), a generalization of the inner product of two vectors in /C693. In a general
linear vector space, an inner product is defined axiomatically analogously with the
inner product on /C69n. Thus given two vectors Ujiand /C87ji
UjiXn
i1uiiji;/C87jiXn
i1/C119iiji;
5:9
where Ujiand /C87ji are expressed in terms of the nbase vectors iji, the inner
product, denoted by the symbol U/C87jih , is defined to be
hU/C87jiXn
i1Xn
j1ui/C42/C119jijjih:
5:10
Uhjis often called the pre-factor and /C87jithe post-factor. The inner product obeys
the following rules (or axioms):
B.1 U/C87ji/C87Uji/C42 h h (skew-symmetry);
B.2 UjUhi 0;0 if and only if Uji0ji (positive semidefiniteness);
B.3 UXji/C87ji UXjiU/C87jih h h (additivity);
B.4 aU /C87jia/C42U/C87ji;Ub /C87ji bU/C87jih h h h (homogeneity);
where aandbare scalars and the asterisk (/C42) denotes complex conjugation. Note
that Axiom B.1 is di/C128erent from the one for the inner product on /C693: the inner
product on a general linear vector space depends on the order of the two factors
for a complex vector space. In a real vector space /C693, the complex conjugation in
Axioms B.1 and B.4 adds nothing and may be ignored. In either case, real orcomplex, Axiom B.1 implies that UUjih is real, so the inequality in Axiom B.2
makes sense.
The inner product is linear with respect to the post-factor:
Uhja/C87bXiaUhj/C87ibUhjXi;
and anti-linear with respect to the prefactor,
aUbX /C87jia/C42U/C87jib/C42X/C87ji: h h h
Two vectors are said to be orthogonal if their inner product vanishes. And we
will refer to the quantity hUUji
1=2kUkas the norm or length of the vector. A
normalized vector, having unit norm, is a unit vector. Any given non-zero vector
may be normalized by dividing it by its length. An orthonormal basis is a set ofbasis vectors that are all of unit norm and pair-wise orthogonal. It is very handy
to have an orthonormal set of vectors as a basis for a vector space, so for hijjiin
Eq. (5.10) we shall assume
ihjji
ij1 for ij
0 for i6j(
;
207INNER PRODUCT SPACES (UNITARY SPACES)
then Eq. (5.10) reduces to
Uhj/C87iX
iX
jui/C42/C119jijX
iui/C42X
j/C119jij/C32!
X
iui/C42/C119i:
5:11
Note that Axiom B.2 implies that if a vector Ujiis orthogonal to every vector
of the vector space, then Uji0: since U0ijh for all jibelongs to the vector
space, so we have in particular UUji 0 h .
We will show shortly that we may construct an orthonormal basis from an
arbitrary basis using a technique known as the Gram–Schmidt orthogonalization
process.
Example 5.7
LetjUi
3ÿ4ij1i
5ÿ6ij2iandj/C87i
1ÿij1i
2ÿ3ij2ibe two vec-
tors expanded in terms of an orthonormal basis j1iandj2i. Then we have, using
Eq. (5.10):
UhjUi
34i
3ÿ4i
56i
5ÿ6i86;
/C87hj/C87i
1i
1ÿi
23i
2ÿ3i15;
Uhj/C87i
34i
1ÿi
56i
2ÿ3i35ÿ2i/C87hjUi/C42:
Example 5.8If~Aand ~Bare two matrices, where
~Aa
11a12
a21a22/C32!
; ~Bb11b12
b21b22/C32!
;
then the following formula defines an inner product on M22:
~A/C10/C12/C12~B/C11
a11b11a12b12a21b21a22b22:
To see this, let us first expand ~Aand ~Bin terms of the following base vectors
1ji 10
00
;2ji 0100
;3ji 0010
;4ji 0001
;
~Aa
111jia122jia213jia224ji; ~Bb111jib122jib213jib224ji:
The result follows easily from the defining formula (5.10).
Example 5.9
Consider the vector jUi, in a certain orthonormal basis, with components
Uji1i
3p
i/C32!
;i
ÿ1p
:
208LINEAR VECTOR SPACES
We now expand it in a new orthonormal basis je1i;je2iwith components
e1ji1
2p1
1
;e2ji1
2p1
ÿ1
:
To do this, let us write
Ujiu1e1jiu2e2ji
and determine u1andu2. To determine u1, we take the inner product of both sides
with he1j:
u1e1hjUi12p11
1i
3p
i/C32!
1
2p
1
3p
2i;
likewise,
u21
2p
1ÿ
3p
:
As a check on the calculation, let us compute the norm squared of the vector and
see if it equals j1ij2j
3p
ij26. We find
u1jj2u2jj21
2
132
3p
413ÿ23p
6:
/C84he /C71ram/C177/C83chmidt orthogonali/C122ation process
We now take up the Gram–Schmidt orthogonalization method for converting a
linearly independent basis into an orthonormal one. The basic idea can be clearly
illustrated in the following steps. Let j1i;j2i;...;jii;...be a linearly independent
basis. To get an orthonormal basis out of these, we do the following:
Step 1. Rescale the first vector by its own length, so it becomes a unit vector.
This will be the first basis vector.
e
1ji1ji
1jijj;
where 1 jijj
1j1hi/C112
. Clearly
e1je1 hi 1j1hi
1jijj1:
Step 2. To construct the second member of the orthonormal basis, we subtract
from the second vector j2iits projection along the first, leaving behind
only the part perpendicular to the first.
IIji2jiÿe1jie1j2hi :
209THE GRAM–SCHMIDT ORTHOGONALI/C90ATION PROCESS
Clearly
e1jIIhi e1j2hi ÿe1je1hi e1j2hi 0;i:e:;
IIj?je1i:
Dividing jIIiby its norm (length), we now have the second basis vector
and it is orthogonal to the first base vector je1iand of unit length.
Step 3. To construct the third member of the orthonormal basis, consider
jIIIij3iÿje1ihe1jIIIiÿje2i2jIIIi
which is orthogonal to both je1iandje2i. Dividing by its norm we get
je3i.
Continuing in this way, we will obtain an orthonormal basis je1i;je2i;...;jeni.
/C84he /C67auch/C121/C177/C83ch/C119ar/C122 inequalit/C121
If/C65and/C66are non-zero vectors in /C693, then the dot product gives /C65/C66ABcos,
where is the angle between the vectors. If we square both sides and use the fact
that cos21, we obtain the inequality
/C65/C662A2B2or /C65/C66jj AB:
This is known as the Cauchy–Schwarz inequality. There is an inequality corre-
sponding to the Cauchy–Schwarz inequality in any inner-product space that
obeys Axioms B.1–B.4, which can be stated as
Uhj/C87i jj jUj/C87jj;Ujj
UhjUi/C112
etc:;
5:13
where jUiandj/C87iare two non-zero vectors in an inner-product space.
This can be proved as follows. We first note that, for any scalar , the following
inequality holds
0U/C87 hj U/C87i jj2U/C87 hj U/C87i
UhjUi/C87hj UiUhj/C87i/C87hj /C87i
Ujj2/C42/C86hjUiUhj/C87ijj2/C87jj2:
Now let hUj/C87i/C42=jhUj/C87ij, with real. This is possible if j/C87i6 0, but if
hUj/C87i0, then Cauchy–Schwarz inequality is trivial. Making this substitution
in the above, we have
0Ujj22Uhj/C87i jj 2/C87jj2:
This is a quadratic expression in the real variable with real coecients.
Therefore, the discriminant must be less than or equal to zero:
4Uhj/C87i jj2ÿ4Ujj2/C87jj20
210LINEAR VECTOR SPACES
or
Uhj/C87i jj Ujj/C87jj;
which is the Cauchy–Schwarz inequality.
From the Cauchy–Schwarz inequality follows another important inequality,
known as the triangle inequality,
U/C87 jj Ujj /C87jj:
5:14
The proof of this is very straightforward. For any pair of vectors, we have
U/C87 jj2U/C87 hj U/C87iUjj2/C87jj2Uhj/C87i/C87hjUi
Ujj2/C87jj22Uhj/C87i jj
Ujj2/C87jj22Ujj/C87jj
Ujj2/C87jj2
from which it follows that
U/C87 jj Ujj/C87jj:
If/C86denotes the vector space of real continuous functions on the interval
axb, and fand gare any real continuous functions, then the following is
an inner product on /C86:
fhj/C103iZb
af
x/C103
xdx:
The Cauchy–Schwarz inequality now gives
Zb
af
x/C103
xdx 2
Zb
af2
xdxZb
a/C1032
xdx
or in Dirac notation
fhj/C103ijj2fjj2/C103jj2:
/C68ual /C118ectors and dual spaces
We begin with a technical point regarding the inner product huj/C118i. If we set
j/C118ij/C119i/C12jzi;
then
huj/C118ihuj/C119i/C12hujzi
is a linear function of and/C12. However, if we set
juij/C119i/C12jzi;
211DUAL VECTORS AND DUAL SPACES
then
huj/C118ih/C118jui/C42/C42h/C118j/C119i/C42/C12/C42h/C118jzi/C42/C42h/C119j/C118i/C12/C42hzj/C118i
is no longer a linear function of and /C12. To remove this asymmetry, we can
introduce, besides the ket vectors ji, bra vectors hjwhich form a di/C128erent vector
space. We will assume that there is a one-to-one correspondence between ket
vectors ji, and bra vectors hj. Thus there are two vector spaces, the space of
kets and a dual space of bras. A pair of vectors in which each is in correspondence
with the other will be called a pair of dual vectors. Thus, for example, h/C118jis the
dual vector of j/C118i. Note they always carry the same identification label.
We now define the multiplication of ket vectors by bra vectors by requiring
hujj/C118ihuj/C118i:
Setting
uhj/C119hj/C42zhj/C12/C42;
we have
huj/C118i/C42h/C119j/C118i/C12/C42hzj/C118i;
the same result we obtained above, and we see that /C119hj/C42zhj/C12/C42 is the dual
vector of j/C119i/C12jzi.
From the above discussion, it is obvious that inner products are really defined
only between bras and kets and hence from elements of two distinct but relatedvector spaces. There is a basis of vectors jiifor expanding kets and a similar basis
hijfor expanding bras. The basis ket jtiis represented in the basis we are using by
a column vector with all zeros except for a 1 in the ith row, while the basis hijis a
row vector with all zeros except for a 1 in the ith column.
Linear operators
A useful concept in the study of linear vector spaces is that of a linear transfor-
mation, from which the concept of a linear operator emerges naturally. It is
instructive first to review the concept of transformation or mapping. Given vector
spaces /C86and /C87and function T
~,i fT
~associates each vector in /C86with a unique
vector in /C87, we say T
~maps /C86 into /C87, and write T
~:/C86!/C87.I fT
~associates the
vector j/C119iin/C87with the vector j/C118iin/C86, we say that j/C119iis the image ofj/C118iunder T
~and write j/C119iT
~j/C118i. Further, T
~is a linear transformation if:
(a)T
~
juij/C118i T
~juiT
~j/C118ifor all vectors juiandj/C118iin/C86.
(b)T
~
kj/C118i kT
~j/C118ifor all vectors j/C118iin/C86and all scalars k.
We can illustrate this with a very simple example. If j/C118i
x;yis a vector in
/C692, then T
~
j/C118i
x;xy;xÿydefines a function (a transformation) that maps
212LINEAR VECTOR SPACES
/C692into /C693. In particular, if j/C118i
1;1, then the image of j/C118iunder T
~isT
~
j/C118i
1;2;0. It is easy to see that the transformation is linear. If
jui
x1;y1andj/C118i
x2;y2, then
juij/C118i
x1x2;y1y2;
so that
T
~uji/C118ji
x1x2;
x1x2
y1y2;
x1x2ÿ
y1y2
x1;x1y1;x1ÿy1
x2;x2y2;x2ÿy2
T
~uji
T
~/C118ji
and if kis a scalar, then
T
~kuji
kx1;kx1ky1;kx1ÿky1
kx 1;x1y1;x1ÿy1
kT
~uji
:
Thus T
~is a linear transformation.
IfT
~maps the vector space onto itself ( T
~:/C86!/C86), then it is called a linear
operator on /C86.I n/C693a rotation of the entire space about a fixed axis is an example
of an operation that maps the space onto itself. We saw in Chapter 3 that rotation
can be represented by a matrix with elements ij
i;j1;2;3;i fx1;x2;x3are the
components of an arbitrary vector in /C693before the transformation and x0
1;x0
2;x0
3
the components of the transformed vector, then
x0
111x112x213x3;
x0
221x122x223x3;
x0
331x132x233x3:9
>>=
>>;
5:15
In matrix form we have
~x0~
~x;
5:16
where is the angle of rotation, and
~x0x0
1
x0
2
x0
30
B@1
CA; ~xx1
x2
x30
B@1
CA;and ~
111213
212223
3132330
B@1
CA:
In particular, if the rotation is carried out about x3-axis, ~
has the following
form:
~
cosÿsin0
sincos0
00 10
B@1
CA:
213LINEAR OPERATORS
Eq. (5.16) determines the vector /C1200if the vector /C120is given, and ~
is the operator
(matrix representation of the rotation operator) which turns /C120into /C1200.
Loosely speaking, an operator is any mathematical entity which operates on
any vector in /C86and turns it into another vector in /C86. Abstractly, an operator L
~is
a mapping that assigns to a vector j/C118iin a linear vector space /C86another vector jui
in/C86:juiL
~j/C118i. The set of vectors j/C118ifor which the mapping is defined, that is,
the set of vectors j/C118ifor which L
~j/C118ihas meaning, is called the domain of L
~. The
set of vectors juiin the domain expressible as juiL
~j/C118iis called the range of the
operator. An operator L
~is linear if the mapping is such that for any vectors
jui;j/C119iin the domain of L
~and for arbitrary scalars ,/C12, the vector
jui/C12j/C119iis in the domain of L
~and
L
~
jui/C12j/C119i L
~jui/C12L
~j/C119i:
A linear operator is bounded if its domain is the entire space /C86and if there exists a
single constant /C67such that
jL
~j/C118ij<Cjj/C118ij
for all j/C118iin/C86. We shall consider linear bounded operators only.
Matri/C120 representation of operators
Linear bounded operators may be represented by matrix. The matrix will have a
finite or an infinite number of rows according to whether the dimension of /C86is
finite or infinite. To show this, let j1i;j2i;...be an orthonormal basis in /C86; then
every vector j’iin/C86may be written in the form
j’i1j1i2j2i :
Since L
~jiis also in /C86, we may write
L
~j’i/C121j1i/C122j2i :
But
L
~j’i1L
~j1i2L
~j2i ;
so
/C121j1i/C122j2i 1L
~j1i2L
~j2i :
Taking the inner product of both sides with h1jwe obtain
/C121h1jL
~j1i1h1jL
~j2i2/C13111/C13122 ;
214LINEAR VECTOR SPACES
Similarly
/C122h2jL
~j1i1h2jL
~j2i2/C13211/C13222 ;
/C123h3jL
~j1i1h3jL
~j2i2/C13311/C13322 :
In general, we have
/C12iX
j/C13ijj;
where
/C13ijihjL
~jji:
5:17
Consequently, in terms of the vectors j1i;j2i;...as a basis, operator L
~is repre-
sented by the matrix whose elements are /C13ij.
A matrix representing L
~can be found by using any basis, not necessarily an
orthonormal one. Of course, a change in the basis changes the matrix representing
L
~.
/C84he algebra of linear operators
LetA
~andB
~be two operators defined in a linear vector space /C86of vectors ji. The
equation A
~B
~will be understood in the sense that
A
~jiB
~ji for all ji2 /C86:
We define the addition and multiplication of linear operators as
C
~A
~B
~and D
~A
~B
~
if for any ji
C
~ji
A
~B
~jiA
~jiB
~ji;
D
~ji
A
~B
~j i A
~
B
~ji :
Note that A
~B
~andA
~B
~are themselves linear operators.
Example 5.10
(a)
A
~B
~
uji/C12/C118ji A
~
B
~uji /C12
B
~/C118j i
A
~B
~uji/C12
A
~B
~/C118ji;
(b)C
~
A
~B
~/C118ji C
~
A
~/C118ji B
~/C118ji C
~A
~jiC
~B
~ji,
which shows that
C
~
A
~B
~C
~A
~C
~B
~:
215THE ALGEBRA OF LINEAR OPERATORS
In general A
~B
~6B
~A
~. The di/C128erence A
~B
~ÿB
~A
~is called the commutator of A
~andB
~and is denoted by the symbol A
~;B
~/C93:
A
~;B
~A
~B
~ÿB
~A
~:
5:18
An operator whose commutator vanishes is called a commuting operator.
The operator equation
B
~A
~A
~
is equivalent to the vector equation
B
~jiA
~jifor any ji:
And the vector equation
A
~jiji
is equivalent to the operator equation
A
~/C69
~
where /C69
~is the identity (or unit) operator:
/C69
~jiji for any ji:
It is obvious that the equation A
~is meaningless.
Example 5.11
To illustrate the non-commuting nature of operators, let A
~x;B
~d=dx. Then
A
~B
~f
xxd
dxf
x;
and
B
~A
~f
xd
dxxf
xdx
dx
fxdf
dxfxdf
dx
/C69
~A
~B
~f:
Thus,
A
~B
~ÿB
~A
~f
xÿ /C69
~f
x
or
x;d
dx
xd
dxÿd
dxxÿ/C69
~:
Having defined the product of two operators, we can also define an operator
raised to a certain power. For example
A
~mjiA
~A
~ A
~ji:
216LINEAR VECTOR SPACES
/C124/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C123/C122/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C130/C125
mfactor
By combining the operations of addition and multiplication, functions of opera-
tors can be formed. We can also define functions of operators by their powerseries expansions. For example, eA
~formally means
eA
~1A
~1
2/C33A
~21
3/C33A
~3 :
A function of a linear operator is a linear operator.
Given an operator A
~that acts on vector ji, we can define the action of the same
operator on vector hj. We shall use the convention of operating on hjfrom the
right. Then the action of A
~on a vector hjis defined by requiring that for any jui
andh/C118j, we have
fuhjA
~g/C118jiuhj fA
~/C118ji guhjA
~/C118ji:
We may write j/C118ij/C118iand the corresponding bra as h/C118j. However, it is
important to note that h/C118jA/C42h/C118j.
/C69igen/C118alues and eigen/C118ectors of an operator
The result of operating on a vector with an operator A
~is, in general, a di/C128erent
vector. But there may be some vector j/C118iwith the property that operating with A
~on it yields the same vector j/C118imultiplied by a scalar, say :
A
~j/C118ij/C118i:
This is called the eigenvalue equation for the operator A
~, and the vector j/C118iis
called an eigenvector of A
~belonging to the eigenvalue . A linear operator has, in
general, several eigenvalues and eigenvectors, which can be distinguished by asubscript
A
~j/C118kikj/C118ki:
The set fkgof all the eigenvalues taken together constitutes the spectrum of the
operator. The eigenvalues may be discrete, continuous, or partly discrete andpartly continuous. In general, an eigenvector belongs to only one eigenvalue. If
several linearly independent eigenvectors belong to the same eigenvalue, the eigen-
value is said to be degenerate, and the degree of degeneracy is given by the number
of linearly independent eigenvectors.
/C83ome special operators
Certain operators with rather special properties play very important roles inphysics. We now consider some of them below.
217EIGENVALUES AND EIGENVECTORS OF AN OPERATOR
/C84he inverse of an operator
The operator X
~satisfying X
~A
~/C69
~is called the left inverse of A
~and we denote it
byA
~ÿ1
L. Thus, A
~ÿ1
LA
~/C69
~. Similarly, the right inverse of A
~is defined by the
equation
A
~A
~ÿ1
R/C69
~:
In general, A
~ÿ1
LorA
~ÿ1
R, or both, may not be unique and even may not exist at all.
However, if both A
~ÿ1
LandA
~ÿ1
Rexist, then they are unique and equal to each
other:
A
~ÿ1
LA
~ÿ1
RA
~ÿ1;
and
A
~A
~ÿ1A
~ÿ1A
~/C69
~:
5:19
A
~ÿ1is called the operator inverse to A
~. Obviously, an operator is the inverse of
another if the corresponding matrices are.
An operator for which an inverse exists is said to be non-singular, whereas one
for which no inverse exists is singular. A necessary and sucient condition for an
operator A
~to be non-singular is that corresponding to each vector jui, there
should be a unique vector j/C118isuch that juiA
~j/C118i:
The inverse of a linear operator is a linear operator. The proof is simple: let
ju1iA
~j/C1181i;ju2iA
~j/C1182i:
Then
j/C1181iA
~ÿ1ju1i;j/C1182iA
~ÿ1ju2i
so that
c1j/C1181ic1A
~ÿ1ju1i;c2j/C1182ic2A
~ÿ1ju2i:
Thus,
A
~ÿ1c1ju1ic2ju2i A
~ÿ1c1A
~j/C1181ic2A
~j/C1182i
A
~ÿ1A
~c1j/C1181ic2j/C1182i
c1j/C1181ic2j/C1182i
or
A
~ÿ1c1ju1ic2ju2i c1A
~ÿ1ju1ic2A
~ÿ1ju2i:
218LINEAR VECTOR SPACES
The inverse of a product of operators is the product of the inverse in the reverse
order
A
~B
~ÿ1B
~ÿ1A
~ÿ1:
5:20
The proof is straightforward: we have
A
~B
~
A
~B
~ÿ1/C69
~:
Multiplying successively from the left by A
~ÿ1andB
~ÿ1, we obtain
A
~B
~ÿ1B
~ÿ1A
~ÿ1;
which is identical to Eq. (5.20).
/C84he ad/C106oint operators
Assuming that /C86is an inner-product space, then the operator X
~satisfying the
relation
uhjX
~/C118ji/C118hjA
~uji/C42 for any jui;j/C118i2/C86
is called the adjoint operator of A
~and is denoted by A
~. Thus
uhjA
~/C118ji/C118hjA
~uji/C42 for any jui;j/C118i2/C86:
5:21
We first note that hjA
~is a dual vector of A
~ji. Next, it is obvious that
A
~A
~:
5:22:
To see this, let A
~B
~, then ( A
~becomes B
~, and from Eq. (5.21) we find
/C118hjB
~ujiuhjB
~/C118ji/C42;for any jui;j/C118i2/C86:
But
uhjB
~/C118ji/C42uhjA
~/C118ji/C42/C118hjA
~uji:
Thus
/C118hjB
~ujiuhjB
~/C118ji/C42/C118hjA
~uji
from which we find
A
~A
~:
219SOME SPECIAL OPERATORS
It is also easy to show that
A
~B
~B
~A
~:
5:23
For any jui;j/C118i,h/C118jB
~andB
~j/C118iis a pair of dual vectors; hujA
~andA
~juiis also a
pair of dual vectors. Thus we have
/C118hjB
~A
~uji f /C118hjB
~gfA
~uji g f uhjA
~gfB
~/C118j ig/C42
uhjA
~B
~/C118ji/C42/C118hj
A
~B
~uji
and therefore
A
~B
~B
~A
~:
/C72ermitian operators
An operator H
~that is equal to its adjoint, that is, that obeys the relation
H
~H
~
5:24
is called Hermitian or self-adjoint. And H
~is anti-Hermitian if
H
~ÿH
~:
Hermitian operators have the following important properties:
(1) The eigenvalues are real: Let H
~be the Hermitian operator and let j/C118ibe an
eigenvector belonging to the eigenvalue :
H
~j/C118ij/C118i:
By definition, we have
/C118hjA
~/C118ji/C118hjA
~/C118ji/C42;
that is,
/C42ÿ/C118/C118jih 0:
Since h/C118j/C118i6 0, we have
/C42:
(2) Eigenvectors belonging to di/C128erent eigenvalues are orthogonal: Let juiand
j/C118ibe eigenvectors of H
~belonging to the eigenvalues and/C12respectively:
H
~juijui;H
~j/C118i/C12j/C118i:
220LINEAR VECTOR SPACES
Then
uhjH
~/C118ji/C118hjH
~uji/C42:
That is,
ÿ/C12h/C118jui0
since /C42:
But6/C12, so that
h/C118jui0:
(3) The set of all eigenvectors of a Hermitian operator forms a complete set:
The eigenvectors are orthogonal, and since we can normalize them, this
means that the eigenvectors form an orthonormal set and serve as a basis
for the vector space.
/C85nitary operators
A linear operator U
~is unitary if it preserves the Hermitian character of an
operator under a similarity transformation:
U
~A
~U
~ÿ1U
~A
~U
~ÿ1;
where
A
~A
~:
But, according to Eq. (5.23)
U
~A
~U
~ÿ1
U
~ÿ1A
~U
~;
thus, we have
U
~ÿ1A
~U
~U
~A
~U
~ÿ1:
Multiplying from the left by U
~and from the right by U
~, we obtain
U
~
U
~ÿ1A
~U
~U
~U
~U
~A
~;
this reduces to
A
~
U
~U
~
U
~U
~A
~;
since
U
~
U
~ÿ1
U
~ÿ1U
~/C69
~:
Thus
U
~U
~/C69
~
221SOME SPECIAL OPERATORS
or
U
~U
~ÿ1:
5:25
We often use Eq. (5.25) for the definition of the unitary operator.
Unitary operators have the remarkable property that transformation by a uni-
tary operator preserves the inner product of the vectors. This is easy to see: under
the operation U
~, a vector j/C118iis transformed into the vector j/C1180iU
~j/C118i. Thus, if
two vectors j/C118iandjuiare transformed by the same unitary operator U
~, then
u0/C10/C12/C12/C1180/C11
hU
~ujU
~/C118iuhjU
~U
~/C118iu/C118jih ;
that is, the inner product is preserved. In particular, it leaves the norm of a vector
unchanged. Thus, a unitary transformation in a linear vector space is analogous
to a rotation in the physical space (which also preserves the lengths of vectors and
the inner products).
Corresponding to every unitary operator U
~, we can define a Hermitian opera-
torH
~and vice versa by
U
~ei/C34H
~;
5:26
where /C34is a parameter. Obviously
U
~e
i/C34H
~
=eÿi/C34H
~U
~ÿ1:
A unitary operator possesses the following properties:
(1) The eigenvalues are unimodular; that is, if U
~j/C118ij/C118i, then jj1.
(2) Eigenvectors belonging to di/C128erent eigenvalues are orthogonal.(3) The product of unitary operators is unitary.
/C84he pro/C106ection operators
A symbol of the type of juih/C118jis quite useful: it has all the properties of a linear
operator, multiplied from the right by a ket ji, it gives juiwhose magnitude is
h/C118ji; and multiplied from the left by a bra hjit gives h/C118jwhose magnitude is hjui.
The linearity of juih/C118jresults from the linear properties of the inner product. We
also have
fuji/C118hj g
/C118jiuhj:
The operator P
~jjjijhjis a very particular example of projection operator. To
see its e/C128ect on an arbitrary vector jui, let us expand jui:
222LINEAR VECTOR SPACES
uji Xn
j1ujjji;ujjhjui:
5:27
We may write the above as
ujiXn
j1jjijhj/C32!
uji;
which is true for all jui. Thus the object in the brackets must be identified with the
identity operator:
I
~Xn
j1jjijhjXn
j1P
~j:
5:28
Now we will see that the e/C128ect of this particular projection operator on juiis to
produce a new vector whose direction is along the basis vector jjiand whose
magnitude is hjjui:
P
~jujijjijhjuijjiuj:
We see that whatever juiis,P
~jjuiis a multiple of jjiwith a coecient ujwhich is
the component of juialong jji. Eq. (5.28) says that the sum of the projections of a
vector along all the ndirections equals the vector itself.
When P
~jjjijhjacts on jji, it reproduces that vector. On the other hand, since
the other basis vectors are orthogonal to jji, a projection operation on any one of
them gives zero (null vector). The basis vectors are therefore eigenvectors of P
~k
with the property
P
~kjjikjjji;
j;k1;...;n:
In this orthonormal basis the projection operators have the matrix form
P
~1100
000
000
............0
BBBB@1
CCCCA;P
~2000
010
000
............0
BBBB@1
CCCCA;P
~/C78000
000
000
............
10
BBBBB@1
CCCCCA:
Projection operators can also act on bras in the same way:
uhjP
~juhjjijhjuj/C42jhj:
223SOME SPECIAL OPERATORS
/C67hange of basis
The choice of base vectors (basis) is largely arbitrary and di/C128erent representations
are physically equally acceptable. How do we change from one orthonormal set of
base vectors j’1i;j’2i;...;j/C34nito another such set j/C241i;j/C242i;...;j/C24ni/C63 In other
words, how do we generate the orthonomal set j/C241i;j/C242i;...;j/C24nifrom the old
setj’1i;j’2i;...;j’n/C63 This task can be accomplished by a unitary transformation:
j/C24iiU
~j’ii
i1;2;...;n:
5:29
Then given a vector XjiPn
i1ai’iji, it will be transformed into jX0i:
jX0iU
~jXiU
~Xn
i1ai’ijiXn
i1U
~ai’ijiXn
i1ai/C24iji:
We can see that the operator U
~possesses an inverse U
~ÿ1which is defined by the
equation
j’iiU
~ÿ1j/C24ii
i1;2;...;n:
The operator U
~is unitary; for, if XjiPni1ai’ijiand YjiPni1bi’iji, then
XjYhi Xn
i;j1ai/C42bj’ij’j/C10/C11
Xn
i1ai/C42bi; hUXjUYiXn
i;j1ai/C42bj/C24ij/C24j/C10/C11
Xn
i1ai/C42bi:
Hence
Uÿ1U:
The inner product of two vectors is independent of the choice of basis which
spans the vector space, since unitary transformations leave all inner products
invariant. In quantum mechanics inner products give physically observable quan-
tities, such as expectation values, probabilities, etc.
It is also clear that the matrix representation of an operator is di/C128erent in a
di/C128erent basis. To find the e/C128ect of a change of basis on the matrix representation
of an operator, let us consider the transformation of the vector jXiintojYiby the
operator A
~:
jYiA
~j
Xi:
5:30
Referred to the basis j’1i;j’2i;...;j’i;jXiand jYiare given by
jXiPn
i1ai’iijandjYiPni1bij’ii, and the equation jYiA
~jXibecomes
Xn
i1bi’ijiA
~Xn
j1aj’j/C12/C12/C11
:
224LINEAR VECTOR SPACES
Multiplying both sides from the left by the bra vector h’ijwe find
biXn
j1aj’ihjA
~’j/C12/C12/C11
Xn
j1ajAij:
5:31
Referred to the basis j/C241i;j/C242i;...;j/C24nithe same vectors jXiand jYiare
jXiPn
i1a0
i/C24iij, and jYiPni1b0
ij/C24ii, and Eqs. (5.31) are replaced by
b0
iXn
j1a0
j/C24ihjA
~/C24j/C12/C12/C11
Xn
j1a0
jA0
ij;
where A0
ijh/C24ijA
~/C24ji/C12/C12, which is related to A
ijby the following relation:
A0
ij/C24ihjA
~/C24j/C12/C12/C11
U’ihj A
~U’j/C12/C12/C11
’
ihjU/C42A
~U’j/C12/C12/C11
U/C42A
~Uij
or using the rule for matrix multiplication
A0
ij/C24ihjA
~/C24j/C12/C12/C11
U/C42A
~UijXn
r1Xn
s1Uir/C42ArsUsj:
5:32
From Eqs. (5.32) we can find the matrix representation of an operator with
respect to a new basis.
If the operator A
~transforms vector jXiinto vector jYiwhich is vector jXiitself
multiplied by a scalar /C58jYijXi, then Eq. (5.30) becomes an eigenvalue
equation:
A
~jXijXi:
/C67ommuting operators
In general, operators do not commute. But commuting operators do exist andthey are of importance in quantum mechanics. As Hermitian operators play a
dominant role in quantum mechanics, and the eigenvalues and the eigenvectors of
a Hermitian operator are real and form a complete set, respectively, we shall
concentrate on Hermitian operators. It is straightforward to prove that
Two commuting Hermitian operators possess a complete ortho-normal set of common eigenvectors, and vice versa.
IfA
~andA
~j/C118ij/C118iare two commuting Hermitian operators, and if
A
~j/C118ij/C118i;
5:33
then we have to show that
B
~j/C118i/C12j/C118i:
5:34
225COMMUTING OPERATORS
Multiplying Eq. (5.33) from the left by B
~, we obtain
B
~
A
~j/C118i
B
~j/C118i;
which using the fact A
~B
~B
~A
~, can be rewritten as
A
~
B
~j/C118i
B
~j/C118i:
Thus, B
~j/C118iis an eigenvector of A
~belonging to eigenvalue .I fis non-degen-
erate, then B
~j/C118ishould be linearly dependent on j/C118i, so that
a
B
~j/C118i bj/C118i0;with a60 and b60:
It follows that
B
~j/C118iÿ
b=aj/C118i/C12j/C118i:
IfAis degenerate, then the matter becomes a little complicated. We now state
the results without proof. There are three possibilities:
(1) The degenerate eigenvectors (that is, the linearly independent eigenvectors
belonging to a degenerate eigenvalue) of A
~are degenerate eigenvectors of B
~also.
(2) The degenerate eigenvectors of A
~belong to di/C128erent eigenvalues of B
~. In this
case, we say that the degeneracy is removed by the Hermitian operator B
~.
(3) Every degenerate eigenvector of A
~is not an eigenvector of B
~. But there are
linear combinations of the degenerate eigenvectors, as many in number as
the degrees of degeneracy, which are degenerate eigenvectors of A
~but
are non-degenerate eigenvectors of B
~. Of course, the degeneracy is removed
byB
~.
Function spaces
We have seen that functions can be elements of a vector space. We now return to
this theme for a more detailed analysis. Consider the set of all functions that arecontinuous on some interval. Two such functions can be added together to con-
struct a third function /C104
x:
/C104
xf
x/C103
x;axb;
where the plus symbol has the usual operational meaning of ‘add the value of fat
the point xto the value of gat the same point.’
A function f
xcan also be multiplied by a number kto give the function /C112
x:
/C112
xkf
x;axb:
226LINEAR VECTOR SPACES
The centred dot, the multiplication symbol, is again understood in the conven-
tional meaning of ‘multiply by kthe value of f
xat the point x.’
It is evident that the following conditions are satisfied:
(a) By adding two continuous functions, we obtain a continuous function.
(b) The multiplication by a scalar of a continuous function yields again a con-
tinuous function.
(c) The function that is identically zero for axbis continuous, and its
addition to any other function does not alter this function.
(d) For any function f
xthere exists a function
ÿ1f
x, which satisfies
f
x
ÿ 1f
x 0:
Comparing these statements with the axioms for linear vector spaces (Axioms
A.1–A.8), we see clearly that the set of all continuous functions defined on someinterval forms a linear vector space; this is called a function space. We shall
consider the entire set of values of a function f
xas representing a vector jfi
of this abstract vector space /C70(/C70stands for function space). In other words, we
shall treat the number f
xat the point xas the component with ‘index x’o fa n
abstract vector jfi. This is quite similar to what we did in the case of finite-
dimensional spaces when we associated a component a
iof a vector with each
value of the index i. The only di/C128erence is that this index assumed a discrete set
of values 1, 2, etc., up to /C78(for/C78-dimensional space), whereas the argument xof
a function f
xis a continuous variable. In other words, the function f
xhas an
infinite number of components, namely the values it takes in the continuum of
points labeled by the real variable x. However, two questions may be raised.
The first question concerns the orthonormal basis. The components of a vector
are defined with respect to some basis and we do not know which basis has been(or could be) chosen in the function space. Unfortunately, we have to postpone
the answer to this question. Let us merely note that, once a basis has been chosen,
we work only with the components of a vector. Therefore, provided we do not
change to other basis vectors, we need not be concerned about the particular basis
that has been chosen.
The second question is how to define an inner product in an infinite-dimen-
sional vector space. Suppose the function f
xdescribes the displacement of a
string clamped at x0 and xL. We divide the interval of length Linto /C78equal
parts and measure the displacements f
x
ifiat/C78point xi;i1;2;...;/C78.A t
fixed /C78, the functions are elements of a finite /C78-dimensional vector space. An
inner product is defined by the expression
fhj/C103iX/C78
i1fi/C103i:
227FUNCTION SPACES
For a vibrating string, the space is real and there is no need to conjugate anything.
To improve the description, we can increase the number /C78. However, as
/C78!1 by increasing the number of points without limit, the inner product
diverges as we subdivide further and further. The way out of this is to modify
the definition by a positive prefactor L=/C78which does not violate any of the
axioms for the inner product. But now
fhj/C103ilim
!0X/C78
i1fi/C103i!ZL
0f
x/C103
xdx;
by the usual definition of an integral. Thus the inner product of two functions is
the integral of their product. Two functions are orthogonal if this inner product
vanishes, and a function is normalized if the integral of its square equals unity.
Thus we can speak of an orthonormal set of functions in a function space just as
in finite dimensions. The following is an example of such a set of functions defined
in the interval 0 xLand vanishing at the end points:
emji!m
x
2
L/C114
sinmx
L; m1;2;...;1;
emhjeni2
LZL
0sinmx
Lsinnx
Ldxmn:
For the details, see ‘Vibrating strings’ of Chapter 4.
In quantum mechanics we often deal with complex functions and our definition
of the inner product must then modified. We define the inner product of f
xand
/C103
xas
fhj/C103iZL
0f/C42
x/C103
xdx;
where f/C42 is the complex conjugate of f. An orthonormal set for this case is
m
x1
2peimx; m0;1;2;...;
which spans the space of all functions of period 2 with finite norm. A linear
vector space with a complex-type inner product is called a /C72ilbert space .
Where and how did we get the orthonormal functions/C63 In general, by solving
the eigenvalue equation of some Hermitian operator. We give a simple example
here. Consider the derivative operator Dd
=dx/C58
Df
xdf
x=dx;Djfidjfi=dx:
However, /C68is not Hermitian, because it does not meet the condition:
ZL
0f/C42
xd/C103
x
dxdxZL
0/C103/C42
xdf
x
dxdx/C42:
228LINEAR VECTOR SPACES
Here is why:
ZL
0/C103/C42
xdf
x
dxdx/C42ZL
0/C103
xdf/C42
x
dxdx
/C103f/C42L
0ÿZL
0f/C42
xd/C103
x
dxdx:/C12/C12/C12/C12/C12
It is easy to see that hermiticity of /C68is lost on two counts. First we have the term
coming from the end points. Second the integral has the wrong sign. We can fix
both of these by doing the following:
(a) Use operator ÿiD. The extra iwill change sign under conjugation and kill
the minus sign in front of the integral.
(b) Restrict the functions to those that are periodic: f
0f
L.
Thus, ÿiDis a Hermitian operator on period functions. Now we have
ÿidf
x
dxf
x;
where is the eigenvalue. Simple integration gives
f
xAeix:
Now the periodicity requirement gives
eiLei01
from which it follows that
2m=L; m0;1;2;
and the normalization condition gives
A1
Lp:
Hence the set of orthonormal eigenvectors is given by
fm
x1Lpe2imx=L:
In quantum mechanics the eigenvalue equation is the Schro /C200dinger equation and
the Hermitian operator is the Hamiltonian operator. Quantum mechanically, a
system with ndegrees of freedom which is classically specified by ngeneralized
coordinates /C1131;...;/C1132;/C113nis specified at a fixed instant of time by a wave function
/C32
/C1131;/C1132;...;/C113nwhose norm is unity, that is,
/C32hj/C32iZ
/C32
/C1131;/C1132;...;/C113n jj2d/C1131;d/C1132;...;d/C113n1;
229FUNCTION SPACES
the integration being over the accessible values of the coordinates /C1131;/C1132;...;/C113n.
The set of all such wave functions with unit norm spans a Hilbert space /C72. Every
possible state of the system is represented by a function in this Hilbert space, and
conversely, every vector in this Hilbert space represents a possible state of the
system. In addition to depending on the coordinates /C1131;/C1132;...;/C113n, the wave func-
tion depends also on the time t, but the dependence on the /C113s and on tare
essentially di/C128erent. The Hilbert space H is formed with respect to the spatial
coordinates /C1131;/C1132;...;/C113nonly, for example, the inner product is formed with
respect to the /C113s only, and one wave function /C32
/C1131;/C1132;...;/C113n) states its complete
spatial dependence. On the other hand the states of the system at di/C128erent instants
of time t1;t2;... are given by the di/C128erent wave functions
/C321
/C1131;/C1132;...;/C113n;/C322
/C1131;/C1132;...;/C113n...of the Hilbert space.
Problems
5.1 Prove the three main properties of the dot product given by Eq. (5.7).
5.2 Show that the points on a line /C86passing through the origin in /C693form a
linear vector space under the addition and scalar multiplication operationsfor vectors in /C69
3.
Hint: The points of /C86satisfy parametric equations of the form
x1at;x2bt;x3ct; ÿ1 <t<1:
5.3 Do all Hermitian 2 2 matrices form a vector space under addition/C63 Is there
any requirement on the scalars that multiply them/C63
5.4 Let /C86be the set of all points ( x1;x2)i n/C692that lie in the first quadrant; that
is, such that x10 and x20. Show that the set /C86fails to be a vector space
under the operations of addition and scalar multiplication.Hint: Consider u/C61(1, 1) which lies in /C86. Now form the scalar multiplication
ÿ1u
ÿ 1;ÿ1; where is this point located/C63
5.5 Show that the set /C87of all 2 2 matrices having zeros on the principal
diagonal is a subspace of the vector space M
22of all 2 2 matrices.
5.6 Show that j/C87i
4;ÿ1;8is not a linear combination of jUi
1;2;ÿ1
andj/C86i
6;4;2.
5.7 Show that the following three vectors in /C693cannot serve as base vectors of /C693:
1ji
1;1;2;2ji
1;0;1;and 3 ji
2;1;3:
5.8 Determine which of the following lie in the space spanned by jficos2x
andj/C103isin2x:(a) cos 2 x;
b3x2;
c1;
dsinx.
5.9 Determine whether the three vectors
1ji
1;ÿ2;3;2ji
5;6;ÿ1;3ji
3;2;1
are linearly dependent or independent.
230LINEAR VECTOR SPACES
5.10 Given the following three vectors from the vector space of real 2 2
matrices:
1ji01
00
;2ji1101
;3jiÿ2ÿ1
0ÿ2
;
determine whether they are linearly dependent or independent.
5.11 If S1ji;2ji;...;nji fg is a basis for a vector space /C86, show that every set
with more than nvectors is linearly dependent.
5.12 Show that any two bases for a finite-dimensional vector space have the same
number of vectors.
5.13 Consider the vector space /C69
3with the Euclidean inner product. Apply the
Gram–Schmidt process to transform the basis
j1i
1;1;1;j2i
0;1;1;j3i
0;0;1
into an orthonormal basis.
5.14 Consider the two linearly independent vectors of Example 5.10:
jUi
3ÿ4ij1i
5ÿ6ij2i;
j/C87i
1ÿij1i
2ÿ3ij2i;
where j1iandj2iare an orthonormal basis. Apply the Gram–Schmidt pro-
cess to transform the two vectors into an orthonormal basis.
5.15 Show that the eigenvalue of the square of an operator is the square of the
eigenvalue of the operator.
5.16 Show that if, for a given A
~, both operators A
~ÿ1
LandA
~ÿ1
Rexist, then
A
~ÿ1
LA
~ÿ1
RA
~ÿ1:
5.17 Show that if a unitary operator U
~can be written in the form U
~1ie/C70
~,
where eis a real infinitesimally small number, then the operator /C70
~is
Hermitian.
5.18 Show that the di/C128erential operator
/C112
~p
id
dx
is linear and Hermitian in the space of all di/C128erentiable wave functions /C30
x
that, say, vanish at both ends of an interval ( a,b).
5.19 The translation operator T
ais defined to be such that T
a/C30
x
/C30
xa. Show that:
(a)T
amay be expressed in terms of the operator
/C112
~p
id
dx;
231PROBLEMS
(b)T
ais unitary.
5.21 Verify that:
a2
LZL
0sinmx
Lsinnx
Ldxmn:
b1
2pZ2
0ei
mÿndxmn:
232LINEAR VECTOR SPACES
6
/C70unctions of a complex variable
The theory of functions of a complex variable is a basic part of mathematical
analysis. It provides some of the very useful mathematical tools for physicists and
engineers. In this chapter a brief introduction to complex variables is presented
which is intended to acquaint the reader with at least the rudiments of this
important subject.
/C67omple/C120 numbers
The number system as we know it today is a result of gradual development. The
natural numbers (positive integers 1, 2, ...) were first used in counting. Negative
integers and zero (that is, 0, ÿ1;ÿ2;...) then arose to permit solutions of equa-
tions such as x32. In order to solve equations such as bxafor all integers
aandbwhere b60, rational numbers (or fractions) were introduced. Irrational
numbers are numbers which cannot be expressed as a/C47b,w i t h aandbintegers and
b60, such as
2p
1:41423 ;3:14159
Rational and irrational numbers are all real numbers. However, the real num-
ber system is still incomplete. For example, there is no real number xwhich
satisfies the algebraic equation x210/C58x ÿ1p
. The problem is that we
do not know what to make of
ÿ1p
because there is no real number whose square
isÿ1. Euler introduced the symbol i
ÿ1p
in 1777 years later Gauss used the
notation aibto denote a complex number, where aandbare real numbers.
Today, i ÿ1p
is called the unit imaginary number.
In terms of i, the answer to equation x
210i sxi. It is postulated that i
will behave like a real number in all manipulations involving addition and multi-
plication.
We now introduce a general complex number, in Cartesian form
zxiy
6:1
233
and refer to xandyas its real and imaginary parts and denote them by the
symbols Re zand Im z, respectively. Thus if zÿ32i, then Re zÿ3 and
Imz2.
A number with just y60 is called a pure imaginary number.
The complex conjugate, or briefly conjugate, of the complex number zxiy
is
z/C42xÿiy
6:2
and is called ‘ z-star’. Sometimes we write it zand call it ‘ z-bar’. Complex con-
jugation can be viewed as the process of replacing ibyÿiwithin the complex
number.
Basic operations /C119ith complex numbers
Two complex numbers z1x1iy1andz2x2iy2are equal if and only if
x1x2andy1y2.
In performing operations with complex numbers we can proceed as in the
algebra of real numbers, replacing i2byÿ1 when it occurs. Given two complex
numbers z1andz2where z1aib;z2cid, the basic rules obeyed by com-
plex numbers are the following:
(1) Addition:
z1z2
aib
cid
aci
bd:
(2) Subtraction:
z1ÿz2
aibÿ
cid
aÿci
bÿd:
(3) Multiplication:
z1z2
aib
cid
acÿbdi
adÿbc:
(4) Division:
z1
z2aib
cid
aib
cÿid
cid
cÿidacbd
c2d2ibcÿad
c2d2:
Polar form of complex numbers
All real numbers can be visualized as points on a straight line (the x-axis). A
complex number, containing two real numbers, can be represented by a point in a
two-dimensional xyplane, known as the zplane or the complex plane (also
known as the Gauss plane or Argand diagram). The complex variablezxiyand its complex conjugation z/C42 are labeled in Fig. 6.1.
234FUNCTIONS OF A COMPLE/C88 VARIABLE
The complex variable can also be represented by the plane polar coordinates
(r;):
zr
cosisin:
With the help of Euler’s formula
eicosisin;
we can rewrite the last equation in polar form:
zr
cosisinrei;r
x2y2q
zz/C42p
:
6:3
ris called the modulus or absolute value of z, denoted by jzjor mod z; and is
called the phase or argument of zand it is denoted by arg z. For any complex
number z60 there corresponds only one value of in 02. The absolute
value of zhas the following properties. If z1;z2;...;zmare complex numbers, then
we have:
(1)jz1z2zmjjz1jjz2jjzmj:
(2)z1
z2/C12/C12/C12/C12/C12/C12/C12/C12jz1j
jz2j;z260.
(3)jz1z2 zmjjz1jjz2jj zmj.
(4)jz1z2jjz1jÿjz2j.
Complex numbers zreiwith r1 have jzj1 and are called unimodular.
235COMPLE/C88 NUMBERS
Figure 6.1. The complex plane.
We may imagine them as lying on a circle of unit radius in the complex plane.
Special points on this circle are
0
1
=2
i
ÿ1
ÿ=2
ÿi:
The reader should know these points at all times.
Sometimes it is easier to use the polar form in manipulations. For example, to
multiply two complex numbers, we multiply their moduli and add their phases; todivide, we divide by the modulus and subtract the phase of the denominator:
zz
1
rei
r1ei1rr1ei
1;z
z1rei
r1ei1r
r1ei
ÿ1:
On the other hand to add two complex numbers we have to go back to theCartesian forms, add the components and revert to the polar form.
If we view a complex number zas a vector, then the multiplication of zbye
i
(where is real) can be interpreted as a rotation of zcounterclockwise through
angle ; and we can consider eias an operator which acts on zto produce this
rotation. Similarly, the multiplication of two complex numbers represents a rota-tion and a change of length: z
1r1ei1;z2r2ei2,z1z2r1r2ei
12; the new
complex number has length r1r2and phase 12.
Example 6.1Find
1i
8.
Solution: We first write zin polar form: z1ir
cosisin, from which
we find r
2p
;=4. Then
z
2p
cos=4isin=4
2p
ei=4:
Thus
1i8
2p
ei=4816e2i16:
Example 6.2
Show that
1
3p
i
1ÿ
3p
i/C32!10
ÿ1
2i
3p
2:
236FUNCTIONS OF A COMPLE/C88 VARIABLE
1i
3p
1ÿi3p/C32!
10
2ei=3
2eÿi=3/C32!10
e2i=310
e20i=3
e6ie2i=31c o s
2=3isin
2=3 ÿ1
2i
3p
2:
/C68e /C77oivre’s theorem and roots of complex numbers
Ifz1r1ei1andz2r2ei2, then
z1z2r1r2ei
12r1r2cos
12isin
12:
A generalization of this leads to
z1z2znr1r2rnei
12 n
r1r2rncos
12 nisin
12 n;
ifz1z2 znzthis becomes
zn
reinrncos
nisin
n;
from which it follows that
cosisinncos
nisin
n;
6:4
a result known as De Moivre’s theorem. Thus we now have a general rule for
calculating the nth power of a complex number z. We first write zin polar form
zr
cosisin, then
znrn
cosisinnrncosnisinn:
6:5
The general rule for calculating the nth root of a complex number can now be
derived without diculty. A number wis called an nth root of a complex number
zif/C119nz, and we write /C119z1=n.I fzr
cosisin, then the complex num-
ber
/C1190rnpcos
nisin
n
is definitely the nth root of zbecause /C119n
0z. But the numbers
/C119krnpcos2k
nisin2k
n
; k1;2;...;
nÿ1;
are also nth roots of zbecause /C119n
kz. Thus the general rule for calculating the nth
root of a complex number is
/C119rnpcos2k
nisin2k
n
; k0;1;2;...;
nÿ1:
6:6
237COMPLE/C88 NUMBERS
It is customary to call the number corresponding to k0 (that is, /C1190) the princi-
pal root of z.
Thenth roots of a complex number zare always located at the vertices of a
regular polygon of nsides inscribed in a circle of radiusrnpabout the origin.
Example 6.3
Find the cube roots of 8.
Solution: In this case z8i0r
cosisin;r2 and the principal argu-
ment 0. Formula (6.6) then yields
83p
2 cos2k
3isin2k
3
; k0;1;2:
These roots are plotted in Fig. 6.2:
2
k0;08;
ÿ1i
3p
k1;120 8;
ÿ1ÿi
3p
k2;240 8:
Functions of a comple/C120 /C118ariable
Complex numbers zxiybecome variables if xory(or both) vary. Then
functions of a complex variable may be formed. If to each value which a complex
variable zcan assume there corresponds one or more values of a complex variable
w, we say that wis a function of zand write /C119f
zor/C119/C103
z, etc. The
variable zis sometimes called an independent variable, and then wis a dependent
238FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.2. The cube roots of 8.
variable. If only one value of wcorresponds to each value of z, we say that wis a
single-valued function of zor that f
zis single-valued; and if more than one value
ofwcorresponds to each value of z,wis then a multiple-valued function of z. For
example, /C119z2is a single-valued function of z, but /C119 zpis a double-valued
function of z. In this chapter, whenever we speak of a function we shall mean a
single-valued function, unless otherwise stated.
Mapping
Note that wis also a complex variable and so can be written in the form
/C119ui/C118f
xiy;
6:7
where uand/C118are real. By equating real and imaginary parts this is seen to be
equivalent to
uu
x;y;/C118/C118
x;y:
6:8
If/C119f
zis a single-valued function of z, then to each point of the complex z
plane, there corresponds a point in the complex wplane. If f
zis multiple-valued,
a point in the zplane is mapped in general into more than one point. The
following two examples show the idea of mapping clearly.
Example 6.4
Map /C119z2r2e2i:
Solution: This is single-valued function. The mapping is unique, but not
one-to-one. It is a two-to-one mapping, since zandÿzgive the same square.
For example as shown in Fig. 6.3, zÿ2iandz2ÿiare mapped to the
same point w3ÿ4i;a n d z1ÿ3iandÿ13iare mapped into the same
point wÿ8ÿ6i.
The line joining the points P
ÿ2;1and/C81
1;ÿ3in the z-plane is mapped by
/C119z2into a curve joining the image points P0
3;ÿ4and/C810
ÿ8;ÿ6. It is not
239MAPPING
Figure 6.3. The mapping function /C119z2.
very dicult to determine the equation of this curve. We first need the equation of
the line joining Pand/C81in the zplane. The parametric equations of the line
joining Pand/C81are given by
xÿ
ÿ 2
1ÿ
ÿ 2yÿ1
ÿ3ÿ1tor x3tÿ2;y1ÿ4t:
The equation of the line P/C81is then given by z3tÿ2i
1ÿ4t. The curve in
thewplane into which the line P/C81is mapped has the equation
/C119z23tÿ2i
1ÿ4t23ÿ4tÿ7t2i
ÿ422tÿ24t2;
from which we obtain
u3ÿ4tÿ7t2; /C118ÿ422tÿ24t2:
By assigning various values to the parameter t, this curve may be graphed.
Sometimes it is convenient to superimpose the zandwplanes. Then the images
of various points are located on the same plane and the function /C119f
zmay be
said to transform the complex plane to itself (or a part of itself).
Example 6.5
Map /C119f
z zp;zrei:
Solution: There are two square roots: f1
reirpei=2;f2ÿf1rpei
2=2.
The function is double-valued, and the mapping is one-to-two. This is shown in
Fig. 6.4, where for simplicity we have used the same complex plane for both zand
wf
z.
/C66ranch lines and /C82iemann surfaces
We now take a close look at the function /C119 zpof Example 6.5. Suppose we allow z
to make a complete counterclockwise motion around the origin starting from point
240FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.4. The mapping function /C119 zp:
A, as shown in Fig. 6.5. At A,1and/C119rpei=2. After a complete circuit back
toA;12and/C119rpei
2=2ÿrpei=2. However, by making a second
complete circuit back to A,14, and so /C119rpei
4=2rpei=2; that is,
we obtain the same value of wwith which we started.
We can describe the above by stating that if 0 <2we are on one branch of
the multiple-valued function zp, while if 2 <4we are on the other branch
of the function. It is clear that each branch of the function is single-valued. In
order to keep the function single-valued, we set up an artificial barrier such as OB
(the wavy line) which we agree not to cross. This artificial barrier is called abranch line or branch cut, point Ois called a branch point. Any other line
from Ocan be used for a branch line.
Riemann (George Friedrich Bernhard Riemann, 1826–1866) suggested another
way to achieve the purpose of the branch line described above. Imagine the zplane
consists of two sheets superimposed on each other. We now cut the two sheets alongOBand join the lower edge of the bottom sheet to the upper edge of the top sheet.
Then on starting in the bottom sheet and making one complete circuit about Owe
arrive in the top sheet. We must now imagine the other cut edges to be joined
together (independent of the first join and actually disregarding its existence) so
that by continuing the circuit we go from the top sheet back to the bottom sheet.
The collection of two sheets is called a Riemann surface corresponding to the
function zp. Each sheet corresponds to a branch of the function and on each
sheet the function is singled-valued. The concept of Riemann surfaces has the
advantage that the various values of multiple-valued functions are obtained in acontinuous fashion.
/C84he di/C128erential calculus of functions of a comple/C120 /C118ariable
/C76imits and continuity
The definitions of limits and continuity for functions of a complex variable are
similar to those for a real variable. We say that f
zhas limit /C119
0aszapproaches
241Figure 6.5. Branch cut for the function /C119 zp.DIFFERENTIAL CALCULUS
z0, which is written as
lim
z!z0f
z/C1190;
6:9
if
(a)f
zis defined and single-valued in a neighborhood of zz0, with the
possible exception of the point z0itself; and
(b) given any positive number /C34(however small), there exists a positive number
such that f
zÿ/C1190 jj </C34whenever 0 <zÿz0 jj <.
The limit must be independent of the manner in which zapproaches z0.
Example 6.6
(a)I ff
zz2, prove that lim z!z0;f
zz2
0
(b) Find lim z!z0f
zif
f
zz2z6z0
0 zz0:(
Solution: (a) We must show that given any /C34/C620 we can find (depending in
general on /C34) such that jz2ÿz2
0j</C34whenever 0 <jzÿz0j<.
Now if 1, then 0 <jzÿz0j<implies that
zÿz0 jj zz0 jj <zz0 jj zÿz02z0 jj ;
z2ÿz2
0/C12/C12/C12/C12<
zÿz
0 jj 2z0jj <12z0jj
:
Taking as 1 or /C34=
12jz0j, whichever is smaller, we then have jz2ÿz2
0j</C34
whenever 0 <jzÿz0j<, and the required result is proved.
(b) There is no di/C128erence between this problem and that in part (a), since in
both cases we exclude zz0from consideration. Hence lim z!z0f
zz20. Note
that the limit of f
zasz!z0has nothing to do with the value of f
zatz0.
A function f
zis said to be continuous at z0if, given any /C34/C620, there exists a
/C620 such that f
zÿf
z0 jj </C34whenever 0 <zÿz0 jj <. This implies three
conditions that must be met in order that f
zbe continuous at zz0:
(1) lim z!z0f
z/C1190must exist;
(2)f
z0must exist, that is, f
zis defined at z0;
(3)/C1190f
z0.
For example, complex polynomials, 01z12z2nzn(where imay be
complex), are continuous everywhere. Quotients of polynomials are continuous
whenever the denominator does not vanish. The following example provides
further illustration.
242FUNCTIONS OF A COMPLE/C88 VARIABLE
A function f
zis said to be continuous in a region /C82of the zplane if it is
continuous at all points of /C82.
Points in the zplane where f
zfails to be continuous are called discontinuities
off
z, and f
zis said to be discontinuous at these points. If lim z!z0f
zexists
but is not equal to f
z0, we call the point z0a removable discontinuity, since by
redefining f
z0to be the same as lim z!z0f
zthe function becomes continuous.
To examine the continuity of f
zatz1 , we let z1=/C119and examine the
continuity of f
1=/C119at/C1190.
/C68erivatives and analytic functions
Given a continuous, single-valued function of a complex variable f
zin some
region /C82of the zplane, the derivative f0
z
df=dzat some fixed point z0in/C82is
defined as
f0
z0lim
z!0f
z0zÿf
z0
z;
6:10
provided the limit exists independently of the manner in which z!0. Here
zzÿz0,a n d zis any point of some neighborhood of z0.I ff0
zexists at z0
and every point zin some neighborhood of z0, then f
zis said to be analytic at z0.
And f
zis analytic in a region /C82of the complex zplane if it is analytic at every
point in /C82.
In order to be analytic, f
zmust be single-valued and continuous. It is
straightforward to see this. In view of Eq. (6.10), whenever f0
z0exists, then
lim
z!0f
z0zÿf
z0 lim
z!0f
z0zÿf
z0
zlim
z!0z0
that is,
lim
z!0f
zf
z0:
Thus fis necessarily continuous at any point z0where its derivative exists. But the
converse is not necessarily true, as the following example shows.
Example 6.7
The function f
zz/C42 is continuous at z0, but dz/C42=dzdoes not exist anywhere. By
definition,
dz/C42
dzlim
z!0
zz/C42ÿz/C42
zlim
x;y!0
xiyxiy/C42ÿ
xiy/C42
xiy
lim
x;y!0xÿiyxÿiyÿ
xÿiy
xiylim
x;y!0xÿiy
xiy:
243DIFFERENTIAL CALCULUS
Ify0, the required limit is lim x!0x=x1. On the other hand, if
x0, the required limit is ÿ1. Then since the limit depends on the manner
in which z!0, the derivative does not exist and so f
zz/C42 is non-analytic
everywhere.
Example 6.8
Given f
z2z2ÿ1, find f0
zatz01ÿi.
Solution:
f0
z0f0
1ÿi lim
z!1ÿi
2z2ÿ1ÿ2
1ÿi2ÿ1
zÿ
1ÿi
lim
z!1ÿi2zÿ
1ÿiz
1ÿi
zÿ
1ÿi
lim
z!1ÿi2z
1ÿi 4
1ÿi:
The rules for di/C128erentiating sums, products, and quotients are, in general, the
same for complex functions as for real-valued functions. That is, if f0
z0and
/C1030
z0exist, then:
(1)
f/C1030
z0f0
z0/C1030
z0;
(2)
f/C1030
z0f0
z0/C103
z0f
z0/C1030
z0;
(3)f
/C1030
z0/C103
z0f0
z0ÿf
z0/C1030
z0
/C103
z02; if/C1030
z06 0:
/C84he /C67auchy/C177Riemann conditions
We call f
zanalytic at z0,i ff0
zexists for all zin some neighborhood of z0;
andf
zis analytic in a region /C82if it is analytic at every point of /C82. Cauchy and
Riemann provided us with a simple but extremely important test for the analyti-
city of f
z. To deduce the Cauchy–Riemann conditions for the analyticity of
f
z, let us return to Eq. (6.10):
f0
z0lim
z!0f
z0zÿf
z0
z:
If we write f
zu
x;yi/C118
x;y, this becomes
f0
z lim
x;y!0u
xx;yyÿu
x;yi
same for /C118
xiy:
There are of course an infinite number of ways to approach a point zon a two-
dimensional surface. Let us consider two possible approaches – along xand along
244FUNCTIONS OF A COMPLE/C88 VARIABLE
y. Suppose we first take the xroute, so yis fixed as we change x, that is, y0
and x!0, and we have
f0
z lim
x!0u
xx;yÿu
x;y
xi/C118
xx;yÿ/C118
x;y
x
/C64u
/C64xi/C64/C118
/C64x:
We next take the yroute, and we have
f0
z lim
y!0u
x;yyÿu
x;y
iyi/C118
x;yyÿ/C118
x;y
iy
ÿi/C64u
/C64y/C64/C118
/C64y:
Now f
zcannot possibly be analytic unless the two derivatives are identical.
Thus a necessary condition for f
zto be analytic is
/C64u
/C64xi/C64/C118
/C64xÿi/C64u
/C64y/C64/C118
/C64y;
from which we obtain
/C64u
/C64x/C64/C118
/C64yand/C64u
/C64yÿ/C64/C118
/C64x:
6:11
These are the Cauchy–Riemann conditions, named after the French mathemati-
cian A. L. Cauchy (1789–1857) who discovered them, and the German mathema-
tician Riemann who made them fundamental in his development of the theory ofanalytic functions. Thus if the function f
zu
x;yi/C118
x;yis analytic in a
region /C82, then u
x;yand/C118
x;ysatisfy the Cauchy–Riemann conditions at all
points of /C82.
Example 6.9Iff
zz
2x2ÿy22ixy, then f0
zexists for all z/C58f0
z2z,a n d
/C64u
/C64x2x/C64/C118
/C64y;and/C64u
/C64yÿ2yÿ/C64/C118
/C64x:
Thus, the Cauchy–Riemann equations (6.11) hold in this example at all points z.
We can also find examples in which u
x;yand/C118
x;ysatisfy the Cauchy–
Riemann conditions (6.11) at zz0, but f0
z0doesn’t exist. One such example
is the following:
f
zu
x;yi/C118
x;yz5=jzj4ifz60
0i f z0(
:
The reader can show that u
x;yand/C118
x;ysatisfy the Cauchy–Riemann condi-
tions (6.11) at z0, but that f0
0does not exist. Thus f
zis not analytic at
z0. The proof is straightforward, but very tedious.
245DIFFERENTIAL CALCULUS
However, the Cauchy–Riemann conditions do imply analyticity provided an
additional hypothesis is added:
Given f
zu
x;yi/C118
x;y,i fu
x;yand/C118
x;yare contin-
uous with continuous first partial derivatives and satisfy the
Cauchy–Riemann conditions (11) at all points in a region /C82,
then f
zis analytic in /C82.
To prove this, we need the following result from the calculus of real-valued
functions of two variables: If /C104
x;y;/C64/C104=/C64x, and /C64/C104=/C64yare continuous in some
region /C82about
x0;y0, then there exists a function H
x;ysuch that
H
x;y!0a s
x;y!
0;0and
/C104
x0x;y0yÿ/C104
x0;y0/C64/C104
x0;y0
/C64xx/C64/C104
x0;y0
/C64yy
H
x;y
x2
y2q
:
Let us return to
lim
z!0f
z0zÿf
z0
z;
where z0is any point in region /C82and zxiy. Now we can write
f
z0zÿf
z0u
x0x;y0yÿu
x0;y0
i/C118
x0x;y0yÿ/C118
x0;y0
/C64u
x0y0
/C64xx/C64u
x0y0
/C64yyH
x;y
x
2
y2q
i/C64/C118
x0y0
/C64xx/C64/C118
x0y0
/C64yy
/C71
x;y
x
2
y2q
;
where H
x;y!0 and /C71
x;y!0a s
x;y!
0;0.
Using the Cauchy–Riemann conditions and some algebraic manipulation we
obtain
f
z0zÿf
z0/C64u
x0;y0
/C64xi/C64/C118
x0;y0
/C64x
xiy
H
x;yi/C71
x;y
x2
y2q
246FUNCTIONS OF A COMPLE/C88 VARIABLE
and
f
z0zÿf
z0
z/C64u
x0;y0
/C64xi/C64/C118
x0;y0
/C64x
H
xyi/C71
xy
x2
y2q
xiy:
But
x2
y2q
xiy/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C121:
Thus, as z!0, we have
x;y!
0;0and
lim
z!0f
z0zÿf
z0
z/C64u
x0;y0
/C64xi/C64/C118
x0;y0
/C64x;
which shows that the limit and so f0
z0exist. Since f
zis di/C128erentiable at all
points in region /C82,f
zis analytic at z0which is any point in /C82.
The Cauchy–Riemann equations turn out to be both necessary and sucient
conditions that f
zu
x;yi/C118
x;ybe analytic. Analytic functions are also
called regular or holomorphic functions. If f
zis analytic everywhere in the finite
zcomplex plane, it is called an entire function. A function f
zis said to be
singular at zz0, if it is not di/C128erentiable there; the point z0is called a singular
point of f
z.
/C72armonic functions
Iff
zu
x;yi/C118
x;yis analytic in some region of the zplane, then at every
point of the region the Cauchy–Riemann conditions are satisfied:
/C64u
/C64x/C64/C118
/C64y;and/C64u
/C64yÿ/C64/C118
/C64x;
and therefore
/C642u
/C64x2/C642/C118
/C64x/C64y;and/C642u
/C64y2ÿ/C642/C118
/C64y/C64x;
provided these second derivatives exist. In fact, one can show that if f
zis
analytic in some region /C82, all its derivatives exist and are continuous in /C82.
Equating the two cross terms, we obtain
/C642u
/C64x2/C642u
/C64y20
6:12a
throughout the region /C82.
247DIFFERENTIAL CALCULUS
Similarly, by di/C128erentiating the first of the Cauch–Riemann equations with
respect to y, the second with respect to x, and subtracting we obtain
/C642/C118
/C64x2/C642/C118
/C64y20:
6:12b
Eqs. (6.12a) and (6.12b) are Laplace’s partial di/C128erential equations in two inde-
pendent variables xandy. Any function that has continuous partial derivatives of
second order and that satisfies Laplace’s equation is called a harmonic function.
We have shown that if f
zu
x;yi/C118
x;yis analytic, then both uand/C118are
harmonic functions. They are called conjugate harmonic functions. This is a
di/C128erent use of the word conjugate from that employed in determining z/C42.
Given one of two conjugate harmonic functions, the Cauchy–Riemann equa-
tions (6.11) can be used to find the other.
Singular points
A point at which f
zfails to be analytic is called a singular point or a singularity
off
z; the Cauchy–Riemann conditions break down at a singularity. Various
types of singular points exist.
(1) Isolated singular points: The point zz0is called an isolated singular point
off
zif we can find /C620 such that the circle jzÿz0jencloses no
singular point other than z0. If no such can be found, we call z0a non-
isolated singularity.
(2) Poles: If we can find a positive integer nsuch that
limz!z0
zÿz0nf
zA60, then zz0is called a pole of order n.I f
n1,z0is called a simple pole. As an example, f
z1=
zÿ2has a
simple pole at z2. But f
z1=
zÿ23has a pole of order 3 at z2.
(3) Branch point: A function has a branch point at z0if, upon encircling z0and
returning to the starting point, the function does not return to the startingvalue. Thus the function is multiple-valued. An example is f
z zp, which
has a branch point at z0.
(4) Removable singularities: The singular point z
0is called a removable singu-
larity of f
zif lim z!z0f
zexists. For example, the singular point at z0
off
zsin
z=zis a removable singularity, since lim z!0sin
z=z1.
(5) Essential singularities: A function has an essential singularity at a point z0if
it has poles of arbitrarily high order which cannot be eliminated by multi-
plication by
zÿz0n, which for any finite choice of n. An example is the
function f
ze1=
zÿ2, which has an essential singularity at z2.
(6) Singularities at infinity: The singularity of f
zatz1 is the same type as
that of f
1=/C119at/C1190. For example, f
zz2has a pole of order 2 at
z1 , since f
1=/C119/C119ÿ2has a pole of order 2 at /C1190.
248FUNCTIONS OF A COMPLE/C88 VARIABLE
/C69lementar/C121 functions of z
/C84he exponential function ez/C40orexp(z))
The exponential function is of fundamental importance, not only for its own sake,
but also as a basis for defining all the other elementary functions. In its definition
we seek to preserve as many of the characteristic properties of the real exponential
function exas possible. Specifically, we desire that:
(a)ezis single-valued and analytic.
(b)dez=dzez.
(c)ezreduces to exwhen Im z0:
Recall that if we approach the point zalong the x-axis (that is, y0;
x!0), the derivative of an analytic function f0
zcan be written in the form
f0
zdf
dz/C64u
/C64xi/C64/C118
/C64x:
If we let
ezui/C118;
then to satisfy ( b) we must have
/C64u
/C64xi/C64/C118
/C64xui/C118:
Equating real and imaginary parts gives
/C64u
/C64xu;
6:13
/C64/C118
/C64x/C118:
6:14
Eq. (6.13) will be satisfied if we write
uex/C30
y;
6:15
where /C30
yis any function of y. Moreover, since ezis to be analytic, uand/C118must
satisfy the Cauchy–Riemann equations (6.11). Then using the second of Eqs.(6.11), Eq. (6.14) becomes
ÿ/C64u
/C64y/C118:
249ELEMENTARY FUNCTIONS OF z
Di/C128erentiating this with respect to y, we obtain
/C642u
/C64y2ÿ/C64/C118
/C64y
ÿ/C64u
/C64x
with the aid of the f irst of Eqs :
6:11:
Finally, using Eq. (6.13), this becomes
/C642u
/C64y2ÿu;
which, on substituting Eq. (6.15), becomes
ex/C3000
yÿ ex/C30
yor/C3000
yÿ /C30
y:
This is a simple linear di/C128erential equation whose solution is of the form
/C30
yAcosyBsiny:
Then
uex/C30
y
ex
AcosyBsiny
and
/C118ÿ/C64u
/C64yÿex
ÿAsinyBcosy:
Therefore
ezui/C118ex
AcosyBsinyi
AsinyÿBcosy:
If this is to reduce to exwhen y0, according to ( c), we must have
exex
AÿiB
from which we find
A1 and B0:
Finally we find
ezexiyex
cosyisiny:
6:16
This expression meets our requirements ( a), (b), and ( c); hence we adopt it as the
definition of ez. It is analytic at each point in the entire zplane, so it is an entire
function. Moreover, it satisfies the relation
ez1ez2ez1z2:
6:17
It is important to note that the right hand side of Eq. (6.16) is in standard polar
form with the modulus of ezgiven by exand an argument by y:
mod ezjezjexand arg ezy:
250FUNCTIONS OF A COMPLE/C88 VARIABLE
From Eq. (6.16) we obtain the Euler formula: eiycosyisiny. Now let
y2, and since cos 2 1 and sin 2 0, the Euler formula gives
e2i1:
Similarly,
eiÿ1;ei=2i:
Combining this with Eq. (6.17), we find
ez2ieze2iez;
which shows that ezis periodic with the imaginary period 2 i. Thus
ez2nIez
n0;1;2;...:
6:18
Because of the periodicity all the values that /C119f
zezcan assume are already
assumed in the strip ÿ<y. This infinite strip is called the fundamental
region of ez.
/C84rigonometric and hyperbolic functions
From the Euler formula we obtain
cosx1
2
eixeÿix;sinx1
2i
eixÿeÿix
xreal:
This suggests the following definitions for complex z:
cosz12
e
izeÿiz;sinz1
2i
eizÿeÿiz:
6:19
The other trigonometric functions are defined in the usual way:
tanzsinz
cosz;cotzcosz
sinz;secz1
cosz;cosec z1
sinz;
whenever the denominators are not zero.
From these definitions it is easy to establish the validity of such familiar for-
mulas as:
sin
ÿzÿ sinz;cos
ÿzcosz;and cos2zsin2z1;
cos
z1z2cosz1cosz2/C7sinz1sinz2;sin
z1z2sinz1cosz2cosz1sinz2
dcosz
dzÿsinz;dsinz
dzcosz:
Since ezis analytic for all z, the same is true for the function sin zand cos z. The
functions tan zand sec zare analytic except at the points where cos zis zero,
and cot zand cosec zare analytic except at the points where sin zis zero. The
251ELEMENTARY FUNCTIONS OF z
functions cos zand sec zare even, and the other functions are odd. Since the
exponential function is periodic, the trigonometric functions are also periodic,
and we have
cos
z2ncosz;sin
z2nsinz;
tan
z2ntanz;cot
z2ncotz;
where n0;1;...:
Another important property also carries over: sin zand cos zhave the same
zeros as the corresponding real-valued functions:
sinz0 if and only if zn
ninteger ;
cosz0 if and only if z
2n1=2
ninteger :
We can also write these functions in the form u
x;yi/C118
x;y. As an example,
we give the details for cos z. From Eq. (6.19) we have
cosz1
2
eizeÿiz12
e
i
xiyeÿi
xiy12
e
ÿyeixeyeÿix
12e
ÿy
cosxisinxey
cosxÿisinx
cosxeyeÿy
2ÿisinxeyÿeÿy
2
or, using the definitions of the hyperbolic functions of real variables
coszcos
xiycosxcosh yÿisinxsinhy;
similarly,
sinzsin
xiysinxcoshyicosxsinhy:
In particular, taking x0 in these last two formulas, we find
cos
iycosh y;sin
iyisinhy:
There is a big di/C128erence between the complex and real sine and cosine func-
tions. The real functions are bounded between ÿ1a n d 1, but the
complex functions can take on arbitrarily large values. For example, if yis real,
then cos iy1
2
eÿyey!1 asy!1 ory!ÿ 1 .
/C84he logarithmic function wlnz
The real natural logarithm ylnxis defined as the inverse of the exponential
function eyx. For the complex logarithm, we take the same approach and
define /C119lnzwhich is taken to mean that
e/C119z
6:20
for each z60.
252FUNCTIONS OF A COMPLE/C88 VARIABLE
Setting /C119ui/C118andzreijzjeiwe have
e/C119eui/C118euei/C118rei:
It follows that
eurzjjorulnrlnzjj
and
/C118argz:
Therefore
/C119lnzlnrilnzjj iargz:
Since the argument of zis determined only in multiples of 2 , the complex
natural logarithm is infinitely many-valued. If we let 1be the principal argument
ofz, that is, the particular argument of zwhich lies in the interval 0 <2,
then we can rewrite the last equation in the form
lnzlnzjji
2n n0;1;2;...:
6:21
For any particular value of n, a unique branch of the function is determined, and
the logarithm becomes e/C128ectively single-valued. If n0, the resulting branch of
the logarithmic function is called the principal value. Any particular branch of the
logarithmic function is analytic, for we have by di/C128erentiating the definitive rela-
tionze/C119,
dz=d/C119e/C119zord/C119=dzd
lnz=dz1=z:
For a particular value of nthe derivative of ln zthus exists for all z60.
For the real logarithm, ylnxmakes sense when x/C620. Now we can take a
natural logarithm of a negative number, as shown in the following example.
Example 6.10
lnÿ4lnjÿ4jiarg
ÿ4ln 4i
2n; its principal value is ln 4 i
,
a complex number. This explains why the logarithm of a negative number makesno sense in real variable.
/C72yperbolic functions
We conclude this section on ‘‘elementary functions’’ by mentioning briefly thehyperbolic functions; they are defined at points where the denominator does not
vanish:
sinhz
1
2
ezÿeÿz;coshz12
e
zeÿz;
tanh zsinhz=coshz;cothzcoshz=sinhz;
sech z1=cosh z;cosech z1=sinhz:
253ELEMENTARY FUNCTIONS OF z
Since ezandeÿzare entire functions, sinh zand cosh zare also entire functions.
The singularities of tanh zand sech zoccur at the zeros of cosh z, and the
singularities of coth zand cosech zoccur at the zeros of sinh z.
As with the trigonometric functions, basic identities and derivative formulas
carry over in the same form to the complex hyperbolic functions (just replace x
byz). Hence we shall not list them here.
/C67omple/C120 integration
Complex integration is very important. For example, in applications we often en-
counter real integrals which cannot be evaluated by the usual methods, but we canget help and relief from complex integration. In theory, the method of complex
integration yields proofs of some basic properties of analytic functions, which
would be very dicult to prove without using complex integration.
The most fundamental result in complex integration is Cauchy’s integral theo-
rem, from which the important Cauchy integral formula follows. These will be the
subject of this section.
/C76ine integrals in the complex plane
As in real integrals, the indefinite integralRf
zdzstands for any function whose
derivative is f
z. The definite integral of real calculus is now replaced by integrals
of a complex function along a curve. Why/C63 To see this, we can express zin terms
of a real parameter t:z
tx
tiy
t, where, say, atb. Now as tvaries
from atob, the point ( x;y) describes a curve in the plane. We say this curve is
smooth if there exists a tangent vector at all points on the curve; this means that
dx/C47dt anddy/C47dt are continuous and do not vanish simultaneously for a<t<b.
Let/C67be such a smooth curve in the complex zplane (Fig. 6.6), and we shall
assume that /C67has a finite length (mathematicians call /C67a rectifiable curve). Let
f
zbe continuous at all points of /C67. Subdivide /C67intonparts by means of points
z
1;z2;...;znÿ1, chosen arbitrarily, and let az0;bzn. On each arc joining zkÿ1
tozk(k1;2;...;n) choose a point /C119k(possibly /C119kzkÿ1or/C119kzk) and form
the sum
SnXn
k1f
/C119kzk zkzkÿzkÿ1:
Now let the number of subdivisions nincrease in such a way that the largest of the
chord lengths jzkjapproaches zero. Then the sum Snapproaches a limit. If this
limit exists and has the same value no matter how the zjs and /C119js are chosen, then
254FUNCTIONS OF A COMPLE/C88 VARIABLE
this limit is called the integral of f
zalong /C67and is denoted by
Z
Cf
zdzorZb
af
zdz:
6:22
This is often called a contour integral (with contour /C67) or a line integral of f
z.
Some authors reserve the name contour integral for the special case in which /C67is
a closed curve (so end aand end bcoincide), and denote it by the symbolH
f
zdz.
We now state, without proof, a basic theorem regarding the existence of the
contour integral: If /C67 is piecewise smooth and f(z) is continuous on /C67/C44 thenR
Cf
zdzexists .
Iff
zu
x;yi/C118
x;y, the complex line integral can be expressed in terms
of real line integrals as
Z
Cf
zdzZ
C
ui/C118
dxidyZ
C
udxÿ/C118dyiZ
C
/C118dxudy;
6:23
where curve /C67may be open or closed but the direction of integration must be
specified in either case. Reversing the direction of integration results in the change
of sign of the integral. Complex integrals are, therefore, reducible to curvilinear
real integrals and possess the following properties:
(1)R
Cf
z/C103
zdzR
Cf
zdzR
C/C103
zdz;
(2)R
Ckf
zdzkR
Cf
zdz,kany constant (real or complex);
(3)Rb
af
zdzÿRa
bf
zdz;
(4)Rb
af
zdzRm
af
zdRb
mf
zdz;
(5)jR
Cf
zdzjML, where Mmaxjf
zjon/C67, and Lis the length of /C67.
Property (5) is very useful, because in working with complex line integrals it is
often necessary to establish bounds on their absolute values. We now give a brief
255COMPLE/C88 INTEGRATION
Figure 6.6. Complex line integral.
proof. Let us go back to the definition:
Z
Cf
zdzlim
n!1Xn
k1f
/C119kzk:
Now
Xn
k1f
/C119kzk/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12X
n
k1f
/C119k jj zkjj MXn
k1zkjj ML;
where we have used the fact that jf
zj Mfor all points zon/C67and thatPjzkjrepresents the sum of all the chord lengths joining zkÿ1andzk, and
that this sum is not greater than the length Lof/C67. Now taking the limit of
both sides, and property (5) follows. It is possible to show, more generally, that
Z
Cf
zdz/C12/C12/C12/C12/C12/C12/C12/C12Z
Cf
zjj dzjj:
6:24
Example 6.11
Evaluate the integralR
C
z/C422dz, where /C67is a straight line joining the points z0
andz12i.
Solution: Since
z/C422
xÿiy2x2ÿy2ÿ2xyi;
we have
Z
C
z/C422dzZ
C
x2ÿy2dx2xydyiZ
Cÿ2xydx
x2ÿy2dy:
But the Cartesian equation of /C67isy2x, and the above integral therefore
becomes
Z
C
z/C422dzZ1
05x2dxiZ1
0
ÿ10x2dx5=3ÿi10=3:
Example 6.12
Evaluate the integral
Z
Cdz
zÿz0n1;
where /C67is a circle of radius rand center at z0, and nis an integer.
256FUNCTIONS OF A COMPLE/C88 VARIABLE
Solution: For convenience, let zÿz0rei, where ranges from 0 to 2 asz
ranges around the circle (Fig. 6.7). Then dzrieid, and the integral becomes
Z2
0rieid
rn1ei
n1i
rnZ2
0eÿind:
Ifn0, this reduces to
iZ2
0d2i
and if n60, we have
i
rnZ2
0
cosnÿisinnd0:
This is an important and useful result to which we will refer later.
/C67auchy’s integral theorem
Cauchy’s integral theorem has various theoretical and practical consequences. It
states that if f
zis analytic in a simply-connected region (domain) and on its
boundary /C67, then
I
Cf
zdz0:
6:25
What do we mean by a simply-connected region/C63 A region /C82(mathematicians
prefer the term ‘domain’) is called simply-connected if any simple closed curve
which lies in /C82can be shrunk to a point without leaving /C82. That is, a simply-
connected region has no hole in it (Fig. 6.7( a)); this is not true for a multiply-
connected region. The multiply-connected regions of Fig. 6.7( b) and ( c)h a v e
respectively one and three holes in them.
257COMPLE/C88 INTEGRATION
Figure 6.7. Simply-connected and doubly-connected regions.
Although a rigorous proof of Cauchy’s integral theorem is quite demanding
and beyond the scope of this book, we shall sketch the main ideas. Note that the
integral can be expressed in terms of two-dimensional vector fields /C65and/C66:
I
Cf
zdzI
C
udxÿ/C118dyiZ
C
/C118dxudy
I
C/C65
rdriI
C/C66
rdr;
where
/C65
ru^e1ÿ/C118^e2;/C66
r/C118^e1u^e2:
Applying Stokes’ theorem, we obtain
I
Cf
zdzZZ
Rda
/C114 /C65i/C114 /C66
ZZ
Rdxdy ÿ/C64/C118
/C64x/C64u
/C64y
i/C64u
/C64xÿ/C64/C118
/C64y
;
where /C82is the region enclosed by /C67. Since f
xsatisfies the Cauchy–Riemann
conditions, both the real and the imaginary parts of the integral are zero, thusproving Cauchy’s integral theorem.
Cauchy’s theorem is also valid for multiply-connected regions. For simplicity
we consider a doubly-connected region (Fig. 6.8). f
zis analytic in and on the
boundary of the region /C82between two simple closed curves C
1andC2. Construct
a cross-cut AF. Then the region bounded by AB/C68EA/C70/C71/C72/C70A is simply-connected
so by Cauchy’s theorem
I
Cf
zdzI
ABD/C69A/C70/C71H/C70Af
zdz0
or
Z
ABD/C69Af
zdzZ
A/C70f
zdzZ
/C70/C71H/C70f
zdzZ
/C70Af
zdz0:
258FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.8. Proof of Cauchy’s theorem for a doubly-connected region.
ButR
A/C70f
zdzÿR
/C70Af
zdz, therefore this becomes
Z
ABD/C69Af
zdzyZ
/C70/C71H/C70f
zdzy0
or
I
Cf
zdzI
C1f
zdzI
C2f
zdz0;
6:26
where both C1andC2are traversed in the positive direction (in the sense that an
observer walking on the boundary always has the region /C82on his left). Note that
curves C1andC2are in opposite directions.
If we reverse the direction of C2(now C2is also counterclockwise, that is, both
C1andC2are in the same direction.), we have
I
C1f
zdzÿI
C2f
zdz0o rI
C2f
zdzI
C1f
zdz:
Because of Cauchy’s theorem, an integration contour can be moved across any
region of the complex plane over which the integrand is analytic without changing
the value of the integral. It cannot be moved across a hole (the shaded area) or a
singularity (the dot), but it can be made to collapse around one, as shown in Fig.
6.9. As a result, an integration contour /C67enclosing nholes or singularities can be
replaced by nseparated closed contours Ci, each enclosing a hole or a singularity:
I
Cf
zdzXn
k1I
Cif
zdz
which is a generalization of Eq. (6.26) to multiply-connected regions.
There is a converse of the Cauchy’s theorem, known as Morera’s theorem. We
now state it without proof:
Morera/C39s theorem:
If f(z) is continuous in a simply/C45connected region /C82 and the /C67auchy/C39s theorem isvalid around every simple closed curve /C67 in /C82/C44 then f
zis analytic in /C82.
259COMPLE/C88 INTEGRATION
Figure 6.9. Collapsing a contour around a hole and a singularity.
Example 6.13
EvaluateH
Cdz=
zÿawhere /C67is any simple closed curve and zais (a) outside
/C67,(b) inside /C67.
Solution: (a)I fais outside /C67, then f
z1=
zÿais analytic everywhere inside
and on /C67. Hence by Cauchy’s theoremI
Cdz=
zÿa0:
(b)I fais inside /C67and ÿis a circle of radius 2with center at zaso that ÿis
inside C (Fig. 6.10). Then by Eq. (6.26) we haveI
Cdz=
zÿaI
ÿdz=
zÿa:
Now on ÿ,jzÿaj/C34,o rzÿa/C34ei, then dzi/C34eid, and
I
ÿdz
zÿaZ2
0i/C34eid
/C34eiiZ2
0d2i:
/C67auchy’s integral formulas
One of the most important consequences of Cauchy’s integral theorem is what isknown as Cauchy’s integral formula. It may be stated as follows.
If f(z) is analytic in a simply/C45connected region /C82/C44 and z
0is any
point in the interior of /C82 which is enclosed by a simple closed curve
/C67/C44 then
f
z01
2iI
Cf
z
zÿz0dz;
6:27
the integration around /C67 being taken in the positive sense (counter/C45clockwise).
260FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.10.
To prove this, let ÿbe a small circle with center at z0and radius r(Fig. 6.11),
then by Eq. (6.26) we have
I
Cf
z
zÿz0dzI
ÿf
z
zÿz0dz:
Now jzÿz0jrorzÿz0rei;0<2. Then dzireidand the integral
on the right becomes
I
ÿf
z
zÿz0dzZ2
0f
z0reiirei
reidiZ2
0f
z0reid:
Taking the limit of both sides and making use of the continuity of f
z,w eh a v e
I
Cf
z
zÿz0dzlim
r!0Z2
0f
z0reid
iZ2
0lim
r!0f
z0reidiZ2
0f
z0d2if
z0;
from which we obtain
f
z01
2iI
Cf
z
zÿz0dzq:e:d:
Cauchy’s integral formula is also true for multiply-connected regions, but we shall
leave its proof as an exercise.
It is useful to write Cauchy’s integral formula (6.27) in the form
f
z1
2iI
Cf
z0dz0
z0ÿz
to emphasize the fact that zcan be any point inside the close curve /C67.
Cauchy’s integral formula is very useful in evaluating integrals, as shown in the
following example.
261COMPLE/C88 INTEGRATION
Figure 6.11. Cauchy’s integral formula.
Example 6.14
Evaluate the integralH
Cezdz=
z21,i f/C67is a circle of unit radius with center at
(a)ziand ( b)zÿi.
Solution: (a) We first rewrite the integral in the form
I
Cez
zidz
zÿi;
then we see that f
zez=
ziand z0i. Moreover, the function f
zis
analytic everywhere within and on the given circle of unit radius around zi.
By Cauchy’s integral formula we have
I
Cez
zidz
zÿi2if
i2iei
2i
cos 1isin 1:
(b) We find z0ÿiandf
zez=
zÿi. Cauchy’s integral formula gives
I
Cez
zÿidz
ziÿ
cos 1ÿisin 1:
/C67auchy’s integral formula for higher derivatives
Using Cauchy’s integral formula, we can show that an analytic function f
zhas
derivatives of all orders given by the following formula:
f
n
z0n/C33
2iI
Cf
zdz
zÿz0n1;
6:28
where /C67is any simple closed curve around z0andf
zis analytic on and inside /C67.
Note that this formula implies that each derivative of f
zis itself analytic, since it
possesses a derivative.
We now prove the formula (6.28) by induction on n. That is, we first prove the
formula for n1:
f0
z01
2iI
Cf
zdz
zÿz02:
As shown in Fig. 6.12, both z0andz0/C104lie in /C82, and
f0
z0lim
/C104!0f
z0/C104ÿf
z0
/C104:
Using Cauchy’s integral formula we obtain
262FUNCTIONS OF A COMPLE/C88 VARIABLE
f0
z0lim
/C104!0f
z0/C104ÿf
z0
/C104
lim
/C104!01
2i/C104I
C1
zÿ
z0/C104ÿ1
zÿz0/C26/C27
f
zdz:
Now
1
/C1041
zÿ
z0/C104ÿ1
zÿz0
1
zÿz02/C104
zÿz0ÿ/C104
zÿz02:
Thus,
f0
z01
2iI
Cf
z
zÿz02dz1
2ilim
/C104!0/C104I
Cf
z
zÿz0ÿ/C104
zÿz02dz:
The proof follows if the limit on the right hand side approaches zero as /C104!0. To
show this, let us draw a small circle ÿof radius centered at z0(Fig. 6.12), then
1
2ilim
/C104!0/C104I
Cf
z
zÿz0ÿ/C104
zÿz02dz1
2ilim
/C104!0/C104I
ÿf
z
zÿz0ÿ/C104
zÿz02dz:
Now choose hso small (in absolute value) that z0/C104lies in ÿandj/C104j< =2,
and the equation for ÿisjzÿz0j. Thus, we have jzÿz0ÿ/C104j
jzÿz0jÿj/C104j/C62ÿ=2=2. Next, as f
zis analytic in /C82, we can find a positive
number Msuch that jf
zj M. And the length of ÿis 2. Thus,
/C104
2iI
ÿf
zdz
zÿz0ÿ/C104
zÿz02/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
/C104jj
2M
2
=2
22/C104jjM
2!0a s /C104!0;
proving the formula for f0
z0.
263COMPLE/C88 INTEGRATION
Figure 6.12.
Forn2, we begin with
f0
z0/C104ÿf0
z0
/C1041
2i/C104I
C1
zÿz0/C1042ÿ1
zÿz02()
f
zdz
2/C33
2iI
Cf
z
zÿz03dz/C104
2iI
C3
zÿz0ÿ2/C104
zÿz0ÿ/C1042
zÿz03f
zdz:
The result follows on taking the limit as /C104!0 if the last term approaches zero.
The proof is similar to that for the case n1, for using the fact that the integral
around /C67equals the integral around ÿ, we have
/C104
2iI
ÿ3
zÿz0ÿ2/C104
zÿz0ÿ/C1042
zÿz03f
zdz/C104jj
2M
2
=2234/C104jjM
4;/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
assuming Mexists such that j3
zÿz
0ÿ2/C104f
zj<M.
In a similar manner we can establish the results for n3;4;.... We leave it to
the reader to complete the proof by establishing the formula for f
n1
z0, assum-
ing that f
n
z0is true.
Sometimes Cauchy’s integral formula for higher derivatives can be used to
evaluate integrals, as illustrated by the following example.
Example 6.15
Evaluate
I
Ce2z
z14dz;
where /C67is any simple closed path not passing through ÿ1. Consider two cases:
(a)/C67does not enclose ÿ1. Then e2z=
z14is analytic on and inside /C67, and the
integral is zero by Cauchy’s integral theorem.
(b)/C67encloses ÿ1. Now Cauchy’s integral formula for higher derivatives
applies.
Solution: Letf
ze2z, then
f
3
ÿ13/C33
2iI
Ce2z
z14dz:
Now f
3
ÿ18eÿ2, hence
I
Ce2z
z14dz2i
3/C33f
3
ÿ18
3eÿ2i:
264FUNCTIONS OF A COMPLE/C88 VARIABLE
/C83eries representations of anal/C121tic functions
We now turn to a very important notion: series representations of analytic func-
tions. As a prelude we must discuss the notion of convergence of complex series.
Most of the definitions and theorems relating to infinite series of real terms can be
applied with little or no change to series whose terms are complex.
/C67omplex sequences
A complex sequence is an ordered list which assigns to each positive integer na
complex number zn:
z1;z2;...;zn;...:
The numbers znare called the terms of the sequence. For example, both
i;i2;...;in;...or 1i;
1i=2;
1i=4;
1i=8;...are complex sequences.
The nth term of the second sequence is (1 i=2nÿ1. A sequence
z1;z2;...;zn;...is said to be convergent with the limit l(or simply to converge
to the number l) if, given /C34/C620, we can find a positive integer /C78such that
jznÿ/C108j</C34for each n/C78(Fig. 6.13). Then we write
lim
n!1zn/C108:
In words, or geometrically, this means that each term znwith n/C62/C78(that is,
z/C78;z/C781;z/C782;...lies in the open circular region of radius /C34with center at l.
In general, /C78depends on the choice of /C34. Here is an illustrative example.
Example 6.17
Using the definition, show that lim n!1
1z=n1 for all z.
265SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
Figure 6.13. Convergent complex sequence.
Solution: Given any number /C34/C620, we must find /C78such that
1z
nÿ1/C12/C12/C12/C12/C12/C12</C34 ; for all n/C62/C78
from which we find
z=njj </C34
or
zjj=n</C34 ifn/C62zjj=/C34/C78:
Setting znxniyn, we may consider a complex sequence z1;z2;...;znin
terms of real sequences, the sequence of the real parts and the sequence of the
imaginary parts: x1;x2;...;xn, and y1;y2;...;yn. If the sequence of the real parts
converges to the number A, and the sequence of the imaginary parts converges to
the number B, then the complex sequence z1;z2;...;znconverges to the limit
AiB, as illustrated by the following example.
Example 6.18Consider the complex sequence whose nth term is
z
nn2ÿ2n3
3n2ÿ4i2nÿ1
2n1:
Setting znxniyn, we find
xnn2ÿ2n3
3n2ÿ41ÿ
2=n
3=n2
3ÿ4=n2and yn2nÿ1
2n12ÿ1=n
21=n:
Asn!1 ;xn!1=3 and yn!1, thus, zn!1=3i.
/C67omplex series
We are interested in complex series whose terms are complex functions
f1
zf2
zf3
z fn
z :
6:29
The sum of the first nterms is
Sn
zf1
zf2
zf3
z fn
z;
which is called the nth partial sum of the series (6.29). The sum of the remaining
terms after the nth term is called the remainder of the series.
We can now associate with the series (6.29) the sequence of its partial sums
S1;S2;...:If this sequence of partial sums is convergent, then the series converges;
and if the sequence diverges, then the series diverges. We can put this in a formal
way. The series (6.29) is said to converge to the sum S
zin a region /C82if for any
266FUNCTIONS OF A COMPLE/C88 VARIABLE
/C34/C620 there exists an integer /C78depending in general on /C34and on the particular
value of zunder consideration such that
Sn
zÿS
z jj </C34 for all n/C62/C78
and we write
lim
n!1Sn
zS
z:
The di/C128erence Sn
zÿS
zis just the remainder after nterms, Rn
z; thus the
definition of convergence requires that jRn
zj ! 0a sn!1 .
If the absolute values of the terms in (6.29) form a convergent series
f1
zjj f2
zjj f3
zjj fn
zjj
then series (6.29) is said to be absolutely convergent. If series (6.29) converges but
is not absolutely convergent, it is said to be conditionally convergent. The terms
of an absolutely convergent series can be rearranged in any manner whatsoever
without a/C128ecting the sum of the series whereas rearranging the terms of a con-
ditionally convergent series may alter the sum of the series or even cause the series
to diverge.
As with complex sequences, questions about complex series can also be reduced
to questions about real series, the series of the real part and the series of theimaginary part. From the definition of convergence it is not dicult to prove
the following theorem:
A necessary and su/C129cient condition that the series of complex
terms
f
1
zf2
zf3
z fn
z
should convergence is that the series of the real parts and the seriesof the imaginary parts of these terms should each converge.
Moreover/C44 if
X
1
n1RefnandX1
n1Imfn
converge to the respective functions /C82(z) and I(z)/C44 then the
given series converges to R
zI
z/C44 and the series
f1
zf2
zf3
z fn
z converges to R
ziI
z.
Of all the tests for the convergence of infinite series, the most useful is probably
the familiar ratio test , which applies to real series as well as complex series.
267SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
Ratio test
/C71iven the series f1
zf2
zf3
z fn
z /C44 the series converges abso/C45
lutely if
0<r
zjj lim
n!1fn1
z
fn
z/C12/C12/C12/C12/C12/C12/C12/C12<1
6:30
and diverges if jr
zj/C621. /C87hen jr
zj 1/C44 the ratio test provides no information
about the convergence or divergence of the series .
Example 6.19
Consider the complex series
X
nSnX1
n02ÿnieÿn
X1
n02ÿniX1
n0eÿn:
The ratio tests on the real and imaginary parts show that both converge:
lim
n!12ÿ
n1
2ÿn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
1
2, which is positive and less than 1;
lim
n!1eÿ
n1
eÿn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
1
e, which is also positive and less than 1.
One can prove that the full series converges to
X1
n1Sn1
1ÿ1=2i1
1ÿeÿ1:
/C85niform convergence and the /C87eierstrass M-test
To establish conditions, under which series can legitimately be integrated or
di/C128erentiated term by term, the concept of uniform convergence is required:
A series of functions is said to converge uniformly to the functionS(z) in a region /C82/C44 either open or closed/C44 if corresponding to an
arbitrary /C34<0there exists an integral /C78/C44 depending on /C34but not
on z/C44 such that for every value of z in /C82
S
zÿS
n
z jj </C34 f/C111r a/C108/C108 n /C62/C78:
One of the tests for uniform convergence is the Weierstrass M-test (a sucient
test).
268FUNCTIONS OF A COMPLE/C88 VARIABLE
If a se/C113uence of positive constants fMngexists such that
jfn
zj Mnfor all positive integers n and for all values of z in
a given region /C82/C44 and if the series
M1M2 Mn
is convergent/C44 then the series
f1
zf2
zf3
z fn
z
converges uniformly in /C82.
As an illustrative example, we use it to test for uniform convergence of the
series
X1
n1unX1
n1zn
n
n1p
in the region jzj1. Now
junjjzjn
n
n1p 1
n3=2
ifjzj1. Calling Mn1=n3=2, we see thatPMnconverges, as it is a pseries with
/C1123=2. Hence by Wierstrass M-test the given series converges uniformly (and
absolutely) in the indicated region jzj1.
Po/C119er series and /C84aylor series
Power series are one of the most important tools of complex analysis, as power
series with non-zero radii of convergence represent analytic functions. As an
example, the power series
SX1
n0anzn
6:31
clearly defines an analytic function as long as the series converge. We will only be
interested in absolute convergence. Thus we have
lim
n!1an1zn1
anzn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12<1o r zjj<Rlim
n!1anjj
an1jj;
where /C82is the radius of convergence since the series converges for all zlying
strictly inside a circle of radius /C82centered at the origin. Similarly, the series
SX1
n0an
zÿz0n
converges within a circle of radius /C82centered at z0.
269SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
Notice that the Eq. (6.31) is just a Taylor series at the origin of a function with
fn
0ann/C33. Every choice we make for the infinite variables andefines a new
function with its own set of derivatives at the origin. Of course we can go beyond
the origin, and expand a function in a Taylor series centered at zz0. Thus in the
complex analysis there is a Taylor expansion for every analytic function. This isthe question addressed by Taylor/C39s theorem (named after the English mathemati-
cian Brook Taylor, 1685–1731):
If f(z) is analytic throughout a region /C82 bounded by a simple
closed curve /C67/C44 and if z and a are both interior to /C67/C44 then f(z)
can be expanded in a Taylor series centered at zafor
jzÿaj<R:
f
zf
af
0
a
zÿaf00
a
zÿa2
2/C33
fn
a
zÿanÿ1
n/C33Rn;
6:32
where the remainder Rnis given by
Rn
z
zÿan1
2iI
Cf
/C119d/C119
/C119ÿan
/C119ÿz:
Proof: To prove this, we first rewrite Cauchy’s integral formula as
f
z1
2iI
Cf
/C119d/C119
/C119ÿz1
2iI
Cf
/C119
/C119ÿa1
1ÿ
zÿa=
/C119ÿa
d/C119:
6:33
For later use we note that since wis on /C67while zis inside /C67,
zÿa
/C119ÿa/C12/C12/C12/C12/C12/C12<1:
From the geometric progression
1/C113/C113
2 /C113n1ÿ/C113n1
1ÿ/C1131
1ÿ/C113ÿ/C113n1
1ÿ/C113
we obtain the relation
1
1ÿ/C1131/C113 /C113n/C113n1
1ÿ/C113:
270FUNCTIONS OF A COMPLE/C88 VARIABLE
By setting /C113
zÿa=
/C119ÿawe find
1
1ÿ
zÿa=
/C119ÿa1zÿa
/C119ÿazÿa
/C119ÿa2
zÿa
/C119ÿan
zÿa=
/C119ÿan1
/C119ÿz=
/C119ÿa:
We insert this into Eq. (6.33). Since zandaare constant, we may take the powers
of (zÿa) out from under the integral sign, and then Eq. (6.33) takes the form
f
z1
2iI
Cf
/C119d/C119
/C119ÿazÿa
2iI
Cf
/C119d/C119
/C119ÿa2
zÿan
2iI
Cf
/C119d/C119
/C119ÿan1Rn
z:
Using Eq. (6.28), we may write this expansion in the form
f
zf
azÿa
1/C33f0
a
zÿa2
2/C33f00
a
zÿan
n/C33fn
aRn
z;
where
Rn
z
zÿan1
2iI
Cf
/C119d/C119
/C119ÿan
/C119ÿz:
Clearly, the expansion will converge and represent f
zif and only if
limn!1Rn
z0. This is easy to prove. Note that wis on /C67while zis inside
/C67,s ow eh a v e j/C119ÿzj/C620. Now f
zis analytic inside /C67and on /C67, so it follows
that the absolute value of f
/C119=
/C119ÿzis bounded, say,
f
/C119
/C119ÿz/C12/C12/C12/C12/C12/C12/C12/C12<M
for all won/C67. Let rbe the radius of /C67, then j/C119ÿajrfor all won/C67, and /C67has
the length 2 r. Hence we obtain
R
njjjzÿajn
2I
Cf
/C119d/C119
/C119ÿan
/C119ÿz/C12/C12/C12/C12/C12/C12/C12/C12<zÿajjn
2M1
rn2r
Mrzÿa
r/C12/C12/C12/C12/C12/C12
n
!0a s n!1 :
Thus
f
zf
azÿa
1/C33f0
a
zÿa2
2/C33f00
a
zÿan
n/C33fn
a
is a valid representation of f
zat all points in the interior of any circle with its
center at aand within which f
zis analytic. This is called the Taylor series of f
z
with center at a. And the particular case where a0 is called the Maclaurin series
off
z/C91Colin Maclaurin 1698–1746, Scots mathematician/C93.
271SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
The Taylor series of f
zconverges to f
zonly within a circular region around
the point za, the circle of convergence; and it diverges everywhere outside this
circle.
/C84aylor series of elementary functions
Taylor series of analytic functions are quite similar to the familiar Taylor series of
real functions. Replacing the real variable in the latter series by a complex vari-
able we may ‘continue’ real functions analytically to the complex domain. The
following is a list of Taylor series of elementary functions: in the case of multiple-
valued functions, the principal branch is used.
ezX1
n0zn
n/C331zz2
2/C33 ; jzj<1;
sinzX1
n0
ÿ1nz2n1
2n1/C33zÿz3
3/C33z5
5/C33ÿ ; jzj<1;
coszX1
n0
ÿ1nz2n
2n/C331ÿz2
2/C33z4
4/C33ÿ ; jzj<1;
sinhzX1
n0z2n1
2n1/C33zz3
3/C33z5
5/C33 ; jzj<1;
cosh zX1
n0z2n
2n/C331z2
2/C33z4
4/C33 ; jzj<1;
ln
1zX1
n0
ÿ1n1zn
nzÿz2
2z3
3ÿ ; jzj<1:
Example 6.20Expand (1 ÿz
ÿ1about a.
Solution:
1
1ÿz1
1ÿaÿ
zÿa1
1ÿa1
1ÿ
zÿa=
1ÿa1
1ÿaX1
n0zÿa
1ÿan
:
We have established two surprising properties of complex analytic functions:
(1)They have derivatives of all order .
(2)They can always be represented by Taylor series .
This is not true in general for real functions; there are real functions which have
derivatives of all orders but cannot be represented by a power series.
272FUNCTIONS OF A COMPLE/C88 VARIABLE
Example 6.21
Expand ln( az) about a.
Solution: Suppose we know the Maclaurin series, then
ln
1zln
1azÿaln
1a1zÿa
1a
ln
1aln 1zÿa
1a
ln
1azÿa
1a
ÿ1
2zÿa
1a2
13zÿa
1a3
ÿ :
Example 6.22
Letf
zln
1z, and consider that branch which has the value zero when
z0.
(a) Expand f
zin a Taylor series about z0, and determine the region of
convergence.
(b) Expand ln/C91(1 z=
1ÿz)/C93 in a Taylor series about z0.
Solution: (a)
f
zln
1z f
00
f0
z
1zÿ1f0
01
f00
zÿ
1zÿ2f00
0ÿ 1
fF
z2
1zÿ3fF
02/C33
......
f
n1
z
ÿ 1nn/C33
1n
n1f
n1
0
ÿ 1nn/C33:
Then
f
zln
1zf
0f0
0zf00
0
2/C33z2fF
0
3/C33z3
zÿz2
2z3
3ÿz4
4ÿ :
Thenth term is un
ÿ 1nÿ1zn=n. The ratio test gives
lim
n!1un1
un/C12/C12/C12/C12/C12/C12/C12/C12lim
n!1nz
n1/C12/C12/C12/C12/C12/C12/C12/C12zjj
and the series converges for jzj<1.
273SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
(b)l n
1z=
1ÿz ln
1zÿln
1ÿz. Next, replacing zbyÿzin
Taylor’s expansion for ln
1z, we have
ln
1ÿzÿ zÿz2
2ÿz3
3ÿz4
4ÿ :
Then by subtraction, we obtain
ln1z
1ÿz2zz3
3z5
5/C32!
X1
n02z2n1
2n1:
/C76aurent series
In many applications it is necessary to expand a function f
zaround points
where or in the neighborhood of which the function is not analytic. The Taylor
series is not applicable in such cases. A new type of series known as the Laurentseries is required. The following is a representation which is valid in an annular
ring bounded by two concentric circles of C
1andC2such that f
zis single-valued
and analytic in the annulus and at each point of C1andC2, see Fig. 6.14. The
function f
zmay have singular points outside C1and inside C2. Hermann
Laurent (1841–1908, French mathematician) proved that, at any point in theannular ring bounded by the circles, f
zcan be represented by the series
f
zX
1
nÿ1an
zÿan
6:34
where
an1
2iI
Cf
/C119d/C119
/C119ÿan1;n0;1;2;...;
6:35
274FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.14. Laurent theorem.
each integral being taken in the counterclockwise sense around curve /C67lying in
the annular ring and encircling its inner boundary (that is, /C67is any concentric
circle between C1andC2).
To prove this, let zbe an arbitrary point of the annular ring. Then by Cauchy’s
integral formula we have
f
z1
2iI
C1f
/C119d/C119
/C119ÿz1
2iI
C2f
/C119d/C119
/C119ÿz;
where C2is traversed in the counterclockwise direction and C2is traversed in the
clockwise direction, in order that the entire integration is in the positive direction.
Reversing the sign of the integral around C2and also changing the direction of
integration from clockwise to counterclockwise, we obtain
f
z1
2iI
C1f
/C119d/C119
/C119ÿzÿ1
2iI
C2f
/C119d/C119
/C119ÿz:
Now
1=
/C119ÿz1=
/C119ÿaf1=1ÿ
zÿa=
/C119ÿag;
ÿ1=
/C119ÿz1=
zÿ/C1191=
zÿaf1=1ÿ
/C119ÿa=
zÿag:
Substituting these into f
zwe obtain:
f
z1
2iI
C1f
/C119d/C119
/C119ÿzÿ1
2iI
C2f
/C119d/C119
/C119ÿz
1
2iI
C1f
/C119
/C119ÿa1
1ÿ
zÿa=
/C119ÿa
d/C119
1
2iI
C2f
/C119
zÿa1
1ÿ
/C119ÿa=
zÿa
d/C119:
Now in each of these integrals we apply the identity
1
1ÿ/C1131/C113/C1132 /C113nÿ1/C113n
1ÿ/C113
275SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
to the last factor. Then
f
z1
2iI
C1f
/C119
/C119ÿa1zÿa
/C119ÿazÿa
/C119ÿanÿ1
zÿan=
/C119ÿan
1ÿ
zÿa=
/C119ÿa
d/C119
1
2iI
C2f
/C119
zÿa1/C119ÿa
zÿa/C119ÿa
zÿanÿ1
/C119ÿan=
zÿan
1ÿ
/C119ÿa=
zÿa
d/C119
1
2iI
C1f
/C119d/C119
/C119ÿazÿa
2iI
C2f
/C119d/C119
/C119ÿa2
zÿanÿ1
2iI
C2f
/C119d/C119
/C119ÿanRn1
1
2i
zÿaI
C2f
/C119d/C1191
2i
zÿa2I
C1
/C119ÿaf
/C119d/C119
1
2i
zÿanI
C1
/C119ÿanÿ1f
/C119d/C119Rn2;
where
Rn1
zÿan
2iI
C1f
/C119d/C119
/C119ÿan
/C119ÿz;
Rn21
2i
zÿanI
C2
/C119ÿanf
/C119d/C119
zÿ/C119:
The theorem will be established if we can show that lim n!1Rn20
and lim n!1Rn10. The proof of lim n!1Rn10 has already been given in the
derivation of the Taylor series. To prove the second limit, we note that for values
ofwonC2
j/C119ÿajr1;jzÿaj/C26say;jzÿ/C119jj
zÿaÿ
/C119ÿaj /C26ÿr1;
and
j
f
/C119j M;
where Mis the maximum of jf
/C119jonC2. Thus
Rn2/C12/C12/C12/C12 1
2i
zÿanI
C2
/C119ÿanf
/C119d/C119
zÿ/C119/C12/C12/C12/C12/C12/C12/C12/C12
1
2ijj zÿajjnI
C2/C119ÿajjnf
/C119jj d/C119jj
zÿ/C119jj
or
Rn2/C12/C12/C12/C12rn
1M
2/C26n
/C26ÿr1I
C2d/C119jjM
2r1
/C26n2r1
/C26ÿr1:
276FUNCTIONS OF A COMPLE/C88 VARIABLE
Since r1=/C26 < 1, the last expression approaches zero as n!1 . Hence
limn!1Rn20 and we have
f
z1
2iI
C1f
/C119d/C119
/C119ÿa1
2iI
C1f
/C119d/C119
/C119ÿa2"#
zÿa
1
2iI
C1f
/C119d/C119
/C119ÿa3"#
zÿa2
1
2iI
C2f
/C119d/C1191
zÿa1
2iI
C2
/C119ÿaf
/C119d/C1191
zÿa2 :
Since f
zis analytic throughout the region between C1andC2, the paths of
integration C1andC2can be replaced by any other curve /C67within this region
and enclosing C2. And the resulting integrals are precisely the coecients angiven
by Eq. (6.35). This proves the Laurent theorem.
It should be noted that the coecients of the positive powers ( zÿa) in the
Laurent expansion, while identical in form with the integrals of Eq. (6.28), cannot
be replaced by the derivative expressions
fn
a
n/C33
as they were in the derivation of Taylor series, since f
zis not analytic through-
out the entire interior of C2(or/C67), and hence Cauchy’s generalized integral
formula cannot be applied.
In many instances the Laurent expansion of a function is not found through the
use of the formula (6.34), but rather by algebraic manipulations suggested by the
nature of the function. In particular, in dealing with quotients of polynomials it isoften advantageous to express them in terms of partial fractions and then expand
the various denominators in series of the appropriate form through the use of the
binomial expansion, which we assume the reader is familiar with:
st
nsnnsnÿ1tn
nÿ1
2/C33snÿ2t2n
nÿ1
nÿ2
3/C33snÿ3t3 :
This expansion is valid for all values of n if jsj/C62jtj:Ifjsjjtjthe expansion is
valid only if n is a non/C45negative integer.
That such procedures are correct follows from the fact that the Laurent expan/C45
sion of a function over a given annular ring is uni/C113ue . That is, if an expansion of the
Laurent type is found by any process, it must be the Laurent expansion.
Example 6.23
Find the Laurent expansion of the function f
z
7zÿ2=
z1z
zÿ2in
the annulus 1 <jz1j<3.
277SERIES REPRESENTATIONS OF ANALYTIC FUNCTIONS
Solution: We first apply the method of partial fractions to f
zand obtain
f
zÿ3
z11
z2
zÿ2:
Now the center of the given annulus is zÿ1, so the series we are seeking must
be one involving powers of z1. This means that we have to modify the second
and third terms in the partial fraction representation of f
z:
f
zÿ3
z11
z1ÿ12
z1ÿ3;
but the series for
z1ÿ3ÿ1converges only where jz1j/C623, whereas we
require an expansion valid for jz1j<3. Hence we rewrite the third term in
the other order:
f
zÿ3
z11
z1ÿ12
ÿ3
z1
ÿ3
z1ÿ1
z1ÿ1ÿ12ÿ3
z1ÿ1
z1ÿ2ÿ2
z1ÿ1ÿ2
3ÿ29
z1
ÿ2
27
z12ÿ ; 1<jz1j<3:
Example 6.24
Given the following two functions:
ae3z
z1ÿ3;
b
z2sin1
z2;
find Laurent series about the singularity for each of the functions, name the
singularity, and give the region of convergence.
Solution: (a)zÿ1 is a triple pole (pole of order 3). Let z1u, then
zuÿ1 and
e3z
z13e3
uÿ1
u3eÿ3e3u
u3eÿ3
u313u
3u2
2/C33
3u3
3/C33
3u4
4/C33/C32!
eÿ3 1
z133
z129
2
z19
227
z1
8/C32!
:
The series converges for all values of z6 ÿ1.
(b)zÿ2 is an essential singularity. Let z2u, then zuÿ2, and
278FUNCTIONS OF A COMPLE/C88 VARIABLE
z2sin1
z2usin1
uu1uÿ1
3/C33u31
5/C33u5
1ÿ1
6
z221
120
z24ÿ :
The series converges for all values of z6 ÿ2.
Integration b/C121 the method of residues
We now turn to integration by the method of residues which is useful in evaluat-
ing both real and complex integrals. We first discuss briefly the theory of residues,
then apply it to evaluate certain types of real definite integrals occurring in physics
and engineering.
Residues
Iff
zis single-valued and analytic in a neighborhood of a point za, then, by
Cauchy’s integral theorem,
I
Cf
zdz0
for any contour in that neighborhood. But if f
zhas a pole or an isolated
essential singularity at zaand lies in the interior of /C67, then the above integral
will, in general, be di/C128erent from zero. In this case we may represent f
zby a
Laurent series:
f
zX1
nÿ1an
zÿana0a1
zÿaa2
zÿa2aÿ1
zÿaaÿ2
zÿa2 ;
where
an1
2iI
Cf
z
zÿan1dz; n0;1;2;...:
The sum of all the terms containing negative powers, namely
aÿ1=
zÿaaÿ2=
zÿa2 ;is called the principal part of f
zatza.I n
the special case nÿ1, we have
aÿ11
2iI
Cf
zdz
or
I
Cf
zdz2iaÿ1;
6:36
279INTEGRATION BY THE METHOD OF RESIDUES
the integration being taken in the counterclockwise sense around a simple closed
curve /C67that lies in the region 0 <jzÿaj<Dand contains the point za, where
/C68is the distance from ato the nearest singular point of f
z. The coecient aÿ1is
called the residue of f
zatza, and we shall use the notation
aÿ1Res
zaf
z:
6:37
We have seen that Laurent expansions can be obtained by various methods,
without using the integral formulas for the coecients. Hence, we may determine
the residue by one of those methods and then use the formula (6.36) to evaluate
contour integrals. To illustrate this, let us consider the following simple example.
Example 6.25
Integrate the function f
zzÿ4sinzaround the unit circle /C67in the counter-
clockwise sense.
Solution: Using
sinzX1
n0
ÿ1nz2n1
2n1/C33zÿz3
3/C33z5
5/C33ÿ ;
we obtain the Laurent series
f
zsinz
z41
z3ÿ1
3/C33zz
5/C33ÿz3
7/C33ÿ :
We see that f
zhas a pole of third order at z0, the corresponding residue is
aÿ1ÿ1=3/C33, and from Eq. (6.36) it follows that
Isinz
z4dz2iaÿ1ÿi
3:
There is a simple standard method for determining the residue in the case of a
pole. If f
zhas a simple pole at a point za, the corresponding Laurent series is
of the form
f
zX1
nÿ1an
zÿana0a1
zÿaa2
zÿa2aÿ1
zÿa;
where aÿ160. Multiplying both sides by zÿa,w eh a v e
zÿaf
z
zÿaa0a1
zÿa aÿ1
and from this we have
Res
zaf
zaÿ1lim
z!a
zÿaf
z:
6:38
280FUNCTIONS OF A COMPLE/C88 VARIABLE
Another useful formula is obtained as follows. If f
zcan be put in the form
f
z/C112
z
/C113
z;
where /C112
zand/C113
zare analytic at za;/C112
z6 0, and /C113
z0a tza(that is,
/C113
zhas a simple zero at za. Consequently, /C113
zcan be expanded in a Taylor
series of the form
/C113
z
zÿa/C1130
a
zÿa2
2/C33/C11300
a :
Hence
Res
zaf
zlim
z!a
zÿa/C112
z
/C113
zlim
z!a
zÿa/C112
z
zÿa/C1130
a
zÿa/C11300
a=2 /C112
a
/C1130
a:
6:39
Example 6.26
The function f
z
4ÿ3z=
z2ÿzis analytic except at z0 and z1 where
it has simple poles. Find the residues at these poles.
Solution: We have p
z4ÿ3z;/C113
zz2ÿz. Then from Eq. (6.39) we obtain
Res
z0f
z4ÿ3z
2zÿ1
z0ÿ4; Res
z1f
z4ÿ3z
2zÿ1
z11:
We now consider poles of higher orders. If f
zhas a pole of order m/C621a ta
point za, the corresponding Laurent series is of the form
f
za0a1
zÿaa2
zÿa2aÿ1
zÿaaÿ2
zÿa2aÿm
zÿam;
where aÿm60 and the series converges in some neighborhood of za, except at
the point itself. By multiplying both sides by
zÿamwe obtain
zÿamf
zaÿmaÿm1
zÿaaÿm2
zÿa2 aÿm
mÿ1
zÿa
mÿ1
zÿama0a1
zÿa :
This represents the Taylor series about zaof the analytic function on the left
hand side. Di/C128erentiating both sides ( mÿ1) times with respect to z,w eh a v e
dmÿ1
dzmÿ1
zÿamf
z
mÿ1/C33aÿ1m
mÿ12a0
zÿa :
281INTEGRATION BY THE METHOD OF RESIDUES
Thus on letting z!a
lim
z!admÿ1
dzmÿ1
zÿamf
z
mÿ1/C33aÿ1;
that is,
Res
zaf
z1
mÿ1/C33lim
z!admÿ1
dzmÿ1
zÿamf
z ()
:
6:40
Of course, in the case of a rational function f
zthe residues can also be
determined from the representation of f
zin terms of partial fractions.
/C84he residue theorem
So far we have employed the residue method to evaluate contour integrals whose
integrands have only a single singularity inside the contour of integration. Now
consider a simple closed curve /C67containing in its interior a number of isolated
singularities of a function f
z. If around each singular point we draw a circle so
small that it encloses no other singular points (Fig. 6.15), these small circles,
together with the curve /C67, form the boundary of a multiply-connected region in
which f
zis everywhere analytic and to which Cauchy’s theorem can therefore be
applied. This gives
1
2iI
Cf
zdzI
C1f
zdzI
Cmf
zdz
0:
If we reverse the direction of integration around each of the circles and change thesign of each integral to compensate, this can be written
1
2iI
Cf
zdz1
2iI
C1f
zdz1
2iI
C2f
zdz1
2iI
Cmf
zdz;
282FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.15. Residue theorem.
where all the integrals are now to be taken in the counterclockwise sense. But the
integrals on the right are, by definition, just the residues of f
zat the various
isolated singularities within /C67. Hence we have established an important theorem,
the residue theorem:
Iff
zis analytic inside a simple closed curve /C67 and on /C67 ,except
at a /C174nite number of singular points a1/C44a2;...;amin the interior of
C/C44 then
I
Cf
zdz2iXm
j1Res
zajf
z2i
r1r2 rm;
6:41
where rjis the residue of f
zat the singular point aj.
Example 6.27
The function f
z
4ÿ3z=
z2ÿzhas simple poles at z0 and z1; the
residues are ÿ4 and 1, respectively (cf. Example 6.26). ThereforeI
C4ÿ3z
z2ÿzdz2i
ÿ41ÿ 6i
for every simple closed curve /C67which encloses the points 0 and 1, andI
C4ÿ3z
z2ÿzdz2i
ÿ4ÿ 8i
for any simple closed curve /C67for which z0 lies inside /C67andz1 lies outside,
the integrations being taken in the counterclockwise sense.
/C69/C118aluation of real definite integrals
The residue theorem yields a simple and elegant method for evaluating certainclasses of complicated real definite integrals. One serious restriction is that the
contour must be closed. But many integrals of practical interest involve integra-
tion over open curves. Their paths of integration must be closed before the residue
theorem can be applied. So our ability to evaluate such an integral depends
crucially on how the contour is closed, since it requires knowledge of the addi-
tional contributions from the added parts of the closed contour. A number oftechniques are known for closing open contours. The following types are most
common in practice.
Improper integrals of the rational functionZ
1
ÿ1f
xdx
The improper integral has the meaning
Z1
ÿ1f
xdxlim
a!1Z0
af
xdxlim
b!1Zb
0f
xdx:
6:42
283EVALUATION OF REAL DEFINITE INTEGRALS
If both limits exist, we may couple the two independent passages to ÿ1 and1,
and write
Z1
ÿ1f
xdxlim
r!1Zr
ÿrf
xdx:
6:43
We assume that the function f
xis a real rational function whose denominator
is di/C128erent from zero for all real xand is of degree at least two units higher than
the degree of the numerator. Then the limits in (6.42) exist and we can start from
(6.43). We consider the corresponding contour integral
I
Cf
zdz;
along a contour /C67consisting of the line along the x-axis from ÿrtorand the
semicircle ÿabove (or below) the x-axis having this line as its diameter (Fig. 6.16).
Then let r!1 .I ff
xis an even function this can be used to evaluate
Z1
0f
xdx:
Let us see why this works. Since f
xis rational, f
zhas finitely many poles in the
upper half-plane, and if we choose rlarge enough, /C67encloses all these poles. Then
by the residue theorem we have
I
Cf
zdzZ
ÿf
zdzZr
ÿrf
xdx2iX
Resf
z:
This gives
Zr
ÿrf
xdx2iX
Resf
zÿZ
ÿf
zdz:
We next prove thatR
ÿf
zdz!0i fr!1 . To this end, we set zrei, then ÿ
is represented by rconst, and as zranges along ÿ;ranges from 0 to . Since
284FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.16. Path of the contour integral.
the degree of the denominator of f
zis at least 2 units higher than the degree of
the numerator, we have
f
zjj <k=zjj2
zjjr/C62r0
for suciently large constants kandr. By applying (6.24) we thus obtain
Z
ÿf
zdz/C12/C12/C12/C12/C12/C12/C12/C12<k
r2rk
r:
Hence, as r!1 , the value of the integral over ÿapproaches zero, and we obtain
Z1
ÿ1f
xdx2iX
Resf
z:
6:44
Example 6.28
Using (6.44), show that
Z1
0dx
1x4
2
2p:
Solution: f
z1=
1z4has four simple poles at the points
z1ei=4;z2e3i=4;z3eÿ3i=4;z4eÿi=4:
The first two poles, z1andz2, lie in the upper half-plane (Fig. 6.17) and we find,
using L’Hospital’s rule
Res
zz1f
z1
1z40
zz11
4z3
zz11
4eÿ3i=4ÿ14e
i=4;
Res
zz2f
z1
1z40
zz21
4z3
zz214e
ÿ9i=414e
ÿi=4;
285EVALUATION OF REAL DEFINITE INTEGRALS
Figure 6.17.
then
Z1
ÿ1dx
1x42i
4
ÿei=4eÿi=4sin
4
2p
and so
Z1
0dx
1x41
2Z1
ÿ1dx
1x4
2
2p:
Example 6.29
Show that
Z1
ÿ1x2dx
x212
x22x27
50:
Solution: The poles of
f
zz2
z212
z22z2
enclosed by the contour of Fig. 6.17 are ziof order 2 and zÿ1iof order 1.
The residue at ziis
lim
z!id
dz
zÿi2 z2
zi21
zÿi2
z22z2"#
9iÿ12
100:
The residue at zÿ1iis
lim
z!ÿ1i
z1ÿiz2
z212
z1ÿi
z1i3ÿ4i
25:
Therefore
Z1
ÿ1x2dx
x212
x22x22i9iÿ12
1003ÿ4i
25
7
50:
Integrals of the rational functions of sinandcosZ2
0/C71
sin;cosd
/C71
sin;cosis a real rational function of sin and cos finite on the interval
02. Let zei, then
dzieid;orddz=iz;sin
zÿzÿ1=2i;cos
zzÿ1=2
286FUNCTIONS OF A COMPLE/C88 VARIABLE
and the given integrand becomes a rational function of z, say, f
z.A s ranges
from 0 to 2 , the variable zranges once around the unit circle jzj1 in the
counterclockwise sense. The given integral takes the formI
Cf
zdz
iz;
the integration being taken in the counterclockwise sense around the unit circle.
Example 6.30
Evaluate
Z2
0d
3ÿ2c o s sin:
Solution: Letzei, then dzieid,o rddz=iz, and
sinzÿzÿ1
2i;coszzÿ1
2;
then
Z2
0d
3ÿ2 cos sinI
C2dz
1ÿ2iz26izÿ1ÿ2i;
where /C67is the circle of unit radius with its center at the origin (Fig. 6.18).
We need to find the poles of
1
1ÿ2iz26izÿ1ÿ2i/C58
zÿ6i
6i2ÿ4
1ÿ2i
ÿ1ÿ2iq
2
1ÿ2i
2ÿi;
2ÿi=5;
287EVALUATION OF REAL DEFINITE INTEGRALS
Figure 6.18.
only (2 ÿi=5 lies inside /C67, and residue at this pole is
lim
z!
2ÿi=5zÿ
2ÿi=52
1ÿ2iz26izÿ1ÿ2i
lim
z!
2ÿi=52
2
1ÿ2iz6i1
2iby L’Hospital’s rule :
Then
Z2
0d
3ÿ2c o s sinI
C2dz
1ÿ2iz26izÿ1ÿ2i2i
1=2i:
Fourier integrals of the formZ1
ÿ1f
xsinmx
cosmx/C26/C27
dx
Iff
xis a rational function satisfying the assumptions stated in connection with
improper integrals of rational functions, then the above integrals may be evalu-
ated in a similar way. Here we consider the corresponding integral
I
Cf
zeimzdz
over the contour /C67as that in improper integrals of rational functions (Fig. 6.16),
and obtain the formula
Z1
ÿ1f
xeimxdx2iX
Resf
zeimz
m/C620;
6:45
where the sum consists of the residues of f
zeimzat its poles in the upper half-
plane. Equating the real and imaginary parts on each side of Eq. (6.45), we obtain
Z1
ÿ1f
xcosmxdx ÿ2X
Im Res f
zeimz;
6:46
Z1
ÿ1f
xsinmxdx 2X
Re Res f
zeimz:
6:47
To establish Eq. (6.45) we should now prove that the value of the integral over
the semicircle ÿin Fig. 6.16 approaches zero as r!1 . This can be done as
follows. Since ÿlies in the upper half-plane y0a n d m/C620, it follows that
jeimzjjeimxjeÿmyjj eÿmy1
y0;m/C620:
From this we obtain
jf
zeimzjf
zjj jeimzjf
zjj
y0;m/C620;
which reduces our present problem to that of an improper integral of a rational
function of this section, since f
xis a rational function satisfying the assumptions
288FUNCTIONS OF A COMPLE/C88 VARIABLE
stated in connection these improper integrals. Continuing as before, we see that
the value of the integral under consideration approaches zero as rapproaches 1,
and Eq. (6.45) is established.
Example 6.31
Show that
Z1
ÿ1cosmx
k2x2dx
keÿkm;Z1
ÿ1sinmx
k2x2dx0
m/C620;k/C620:
Solution: The function f
zeimz=
k2z2has a simple pole at zikwhich
lies in the upper half-plane. The residue of f
zatzikis
Res
zikeimz
k2z2eimz
2z
zikeÿmk
2ik:
Therefore
Z1
ÿ1eimx
k2x2dx2ieÿmk
2ik
keÿmk
and this yields the above results.
Other types of real improper integrals
These are definite integrals
ZB
Af
xdx
whose integrand becomes infinite at a point ain the interval of integration,
limx!af
xjj 1 . This means that
ZB
Af
xdxlim
/C34!0Zaÿ/C34
Af
xdxlim
/C17!0Z
a/C17f
xdx;
where both /C34and/C17approach zero independently and through positive values. It
may happen that neither of these limits exists when /C34; /C17!0 independently, but
lim
/C34!0Zaÿ/C34
Af
xdxZB
a/C34f
xdx
exists; this is called Cauchy’s principal value of the integral and is often written
pr:v:ZB
Af
xdx:
289EVALUATION OF REAL DEFINITE INTEGRALS
To evaluate improper integrals whose integrands have poles on the real axis, we
can use a path which avoids these singularities by following small semicircles with
centers at the singular points. We now illustrate the procedure with a simple
example.
Example 6.32
Show that
Z1
0sinx
xdx
2:
Solution: The function sin
z=zdoes not behave suitably at infinity. So we con-
sider eiz=z, which has a simple pole at z0, and integrate around the contour /C67
orAB/C68E/C70/C71A (Fig. 6.19). Since eiz=zis analytic inside and on /C67, it follows from
Cauchy’s integral theorem that
I
Ceiz
zdz0
or
Zÿ/C34
ÿReix
xdxZ
C2eiz
zdzZR
/C34eix
xdxZ
C1eiz
zdz0:
6:48
We now prove that the value of the integral over large semicircle C1approaches
zero as /C82approaches infinity. Setting zRei, we have dziReid;dz=zid
and therefore
Z
C1eiz
zdz/C12/C12/C12/C12/C12/C12/C12/C12Z
0eizid/C12/C12/C12/C12/C12/C12/C12/C12Z
0eiz/C12/C12/C12/C12d:
In the integrand on the right,
eiz/C12/C12/C12/C12je
iR
cosisinjjeiRcosjjeÿRsinjeÿRsin:
290FUNCTIONS OF A COMPLE/C88 VARIABLE
Figure 6.19.
By inserting this and using sin( ÿsinwe obtain
Z
0eiz/C12/C12/C12/C12dZ
0eÿRsind2Z=2
0eÿRsind
2Z/C34
0eÿRsindZ=2
/C34eÿRsind"#
;
where /C34has any value between 0 and =2. The absolute values of the integrands in
the first and the last integrals on the right are at most equal to 1 and eÿRsin/C34,
respectively, because the integrands are monotone decreasing functions of in the
interval of integration. Consequently, the whole expression on the right is smaller
than
2Z/C34
0deÿRsinZ=2
/C34d
2/C34eÿRsin
2ÿ/C34
<2/C34eÿRsin/C34:
Altogether
Z
C1eiz
zdz/C12/C12/C12/C12/C12/C12/C12/C12<2/C34eÿRsin/C34:
We first take /C34arbitrarily small. Then, having fixed /C34, the last term can be made as
small as we please by choosing /C82suciently large. Hence the value of the integral
along C1approaches 0 as R!1 .
We next prove that the value of the integral over the small semicircle C2
approaches zero as /C34!0. Let z/C34i, then
Z
C2eiz
zdzÿlim
/C34!0Z0
exp
i/C34ei
/C34eii/C34eidÿlim
/C34!0Z0
iexp
i/C34eidi
and Eq. (6.48) reduces to
Zÿ/C34
ÿReix
xdxiZR
/C34eix
xdx0:
Replacing xbyÿxin the first integral and combining with the last integral, we
find
ZR
/C34eixÿeÿix
xdxi0:
Thus we have
2iZR
/C34sinx
xdxi:
291EVALUATION OF REAL DEFINITE INTEGRALS
Taking the limits R!1 and/C34!0
Z1
0sinx
xdx
2:
Problems
6.1. Given three complex numbers z1aib,z2cid, and z3/C103i/C104,
show that:
(a)z1z2z2z1 commutative law of addition;
(b)z1
z2z3
z1z2z3 associative law of addition;
(c)z1z2z2z1 commutative law of multiplication;
(d)z1
z2z3
z1z2z3 associative law of multiplication.
6.2. Given
z134i
3ÿ4i;z212i
1ÿ3i2
find their polar forms, complex conjugates, moduli, product, the quotient
z1=z2:
6.3. The absolute value or modulus of a complex number zxiyis defined as
zjj
zz/C42p
x2y2q
:
Ifz1;z2;...;zmare complex numbers, show that the following hold:
(a)jz1z2jjz1jjz2jorjz1z2zmjjz1jjz2jjzmj:
(b)jz1=z2jjz1j=jz2jifz260:
(c)jz1z2jjz1jjz2j:
(d)jz1z2jjz1jÿjz2jorjz1ÿz2jjz1jÿjz2j.
6.4 Find all roots of ( a)
ÿ325p
, and ( b)
1i3p
, and locate them in the complex
plane.
6.5 Show, using De Moivre’s theorem, that:
(a) cos 5 16 cos5ÿ20 cos35c o s ;
(b) sin 5 5 cos4sinÿ10 cos2sin3sin5.
6.6 Given zrei, interpret zei, where is real geometrically.
6.7 Solve the quadratic equation az2bzc0;a60.
6.8 A point Pmoves in a counterclockwise direction around a circle of radius 1
with center at the origin in the zplane. If the mapping function is /C119z2,
show that when Pmakes one complete revolution the image P0ofPin the w
plane makes three complete revolutions in a counterclockwise direction on a
circle of radius 1 with center at the origin.
6.9 Show that f
zlnzhas a branch point at z0.
6.10 Let /C119f
z
z211=2, show that:
292FUNCTIONS OF A COMPLE/C88 VARIABLE
(a)f
zhas branch points at zI.
(b) a complete circuit around both branch points produces no change in the
branches of f
z.
6.11 Apply the definition of limits to prove that:
lim
z!1z2ÿ1
zÿ12:
6.12. Prove that:
(a)f
zz2is continuous at zz0,a n d
(b)f
zz2;z6z0
0;zz0(
is discontinuous at zz0, where z060.
6.13 Given f
zz/C42, show that f0
idoes not exist.
6.14 Using the definition, find the derivative of f
zz3ÿ2zat the point where:
(a)zz0, and (b) zÿ1.
6.15. Show that fis an analytic function of zif it does not depend on
z/C42/C58f
z;z/C42f
z. In other words, f
x;yf
xiy, that is, xand y
enter fonly in the combination x/C43iy .
6.16. (a) Show that uy3ÿ3x2yis harmonic.
(b) Find /C118such that f
zui/C118is analytic.
6.17 ( a)I ff
zu
x;yi/C118
x;yis analytic in some region /C82of the zplane,
show that the one-parameter families of curves u
x;yC1and
/C118
x;yC2are orthogonal families.
(b) Illustrate ( a) by using f
zz2.
6.18 For each of the following functions locate and name the singularities in the
finite zplane:
(a)f
zz
z244;(b)f
zsin zp
zp ;(c)f
zP1
n01
znn/C33:
6.19 ( a) Locate and name all the singularities of
f
zz8z42
zÿ13
3z22:
(b) Determine where f
zis analytic.
6.20 ( a) Given ezex
cosyisiny, show that
d=dzezez.
(b) Show that ez1ez2ez1z2.
(Hint: set z1x1iy1andz2x2iy2and apply the addition formulas
for the sine and cosine.)
6.21 Show that: ( a)l nezz2ni,(b)l nz1=z2lnz1ÿlnz22ni.
6.22 Find the values of: (a) ln i,(b)l n ( 1 ÿi).
6.23 EvaluateR
Cz/C42dzfrom z0t oz42ialong the curve /C67given by:
(a)zt2it;
(b) the line from z0t oz2iand then the line from z2itoz42i.
293PROBLEMS
6.24 EvaluateH
Cdz=
zÿan;n2;3;4;...where zais inside the simple
closed curve /C67.
6.25 If f
zis analytic in a simply-connected region /C82, and aandzare any two
points in /C82, show that the integral
Zz
af
zdz
is independent of the path in /C82joining aandz.
6.26 Let f
zbe continuous in a simply-connected region /C82and let aandzbe
points in /C82. Prove that /C70
zRz
af
z0dz0is analytic in /C82, and /C700
zf
z.
6.27 Evaluate
(a)I
Csinz2cosz2
zÿ1
zÿ2dz
(b)I
Ce2z
z14dz,
where /C67is the circle jzj1.
6.28 Evaluate
I
C2 sinz2
zÿ14dz;
where /C67is any simple closed path not passing through 1.
6.29 Show that the complex sequence
zn1
nÿn2ÿ1
ni
diverges.
6.30 Find the region of convergence of the seriesP1
n1
z2n1=
n134n.
6.31 Find the Maclaurin series of f
z1=
1z2.
6.32 Find the Taylor series of f
zsinzabout z=4, and determine its circle
of convergence. (Hint: sin zsina
zÿa:
6.33 Find the Laurent series about the indicated singularity for each of the
following functions. Name the singularity in each case and give the region
of convergence of each series.
(a)
zÿ3sin1
z2;zÿ2;
(b)z
z1
z2;zÿ2;
(c)1
z
zÿ32;z3:
6.34 Expand f
z1=
z1
z3in a Laurent series valid for:
(a)1<jzj<3, ( b)jzj/C623, ( c)0<jz1j<2.
294FUNCTIONS OF A COMPLE/C88 VARIABLE
6.35 Evaluate
Z1
ÿ1x2dx
x2a2
x2b2; a/C620;b/C620:
6.36 Evaluate
aZ2
0d
1ÿ2/C112cos/C1122;
where pis a fixed number in the interval 0 </C112<1;
bZ2
0d
5ÿ3 sin 2:
6.37 Evaluate
Z1
ÿ1xsinx
x22x5dx:
6.38 Show that:
aZ1
0sinx2dxZ1
0cosx2dx1
2
2/C114
;
bZ1
0x/C112ÿ1
1xdx
sin/C112; 0</C112<1:
295PROBLEMS
7
Special functions of
mathematical physics
The functions discussed in this chapter arise as solutions of second-order di/C128er-
ential equations which appear in special, rather than in general, physical pro-
blems. So these functions are usually known as the special functions of
mathematical physics. We start with Legendre’s equation (Adrien MarieLegendre, 1752–1833, French mathematician).
Legendre/C39s equation
Legendre’s di/C128erential equation
1ÿx
2d2y
dx2ÿ2xdy
dx/C23
/C231y0;
7:1
where vis a positive constant, is of great importance in classical and quantum
physics. The reader will see this equation in the study of central force motion inquantum mechanics. In general, Legendre’s equation appears in problems inclassical mechanics, electromagnetic theory, heat, and quantum mechanics, with
spherical symmetry.
Dividing Eq. (7.1) by 1 ÿx
2, we obtain the standard form
d2y
dx2ÿ2x
1ÿx2dy
dx/C23
/C231
1ÿx2y0:
We see that the coecients of the resulting equation are analytic at x0, so the
origin is an ordinary point and we may write the series solution in the form
yX1
m0amxm:
7:2
296
Substituting this and its derivatives into Eq. (7.1) and denoting the constant
/C23
/C231bykwe obtain
1ÿx2X1
m2m
mÿ1amxmÿ2ÿ2xX1
m1mamxmÿ1kX1
m0amxm0:
By writing the first term as two separate series we have
X1
m2m
mÿ1amxmÿ2ÿX1
m2m
mÿ1amxmÿ2X1
m1mamxmkX1
m0amxm0;
which can be written as:
21a232a3x43a4x2
s2
s1as2xs
ÿ21a2x2ÿ ÿ
s
sÿ1asxsÿ
ÿ21a1xÿ22a2x2ÿ ÿ 2sasxsÿ
ka0ka1x ka2x2 kasxs 0:
Since this must be an identity in xif Eq. (7.2) is to be a solution of Eq. (7.1), the
sum of the coecients of each power of xmust be zero; remembering that
k/C23
/C231we thus have
2a2/C23
/C231a00;
7:3a
6a3 ÿ 2/C118
/C1181a10;
7:3b
and in general, when s2;3;...;
s2
s1as2 ÿs
sÿ1ÿ2s/C23
/C231as0:
4:4
The expression in square brackets /C91 .../C93 can be written
/C23ÿs
/C23s1:
We thus obtain from Eq. (7.4)
as2ÿ
/C23ÿs
/C23s1
s2
s1as
s0;1;...:
7:5
This is a recursion formula, giving each coecient in terms of the one two places
before it in the series, except for a0anda1, which are left as arbitrary constants.
297LEGENDRE’S EQUATION
We find successively
a2ÿ/C23
/C231
2/C33a0; a3ÿ
/C23ÿ1
/C232
3/C33a1;
a4ÿ
/C23ÿ2
/C233
43a2; a5ÿ
/C23ÿ3
/C234
3/C33a3;
/C23ÿ2/C23
/C231
/C233
4/C33a0;
/C23ÿ3
/C23ÿ1
/C232
/C234
5/C33a1;
etc. By inserting these values for the coecients into Eq. (7.2) we obtain
y
xa0y1
xa1y2
x;
7:6
where
y1
x1ÿ/C23
/C231
2/C33x2
/C23ÿ2/C23
/C231
/C233
4/C33x4ÿ
7:7a
and
y2
xx
/C23ÿ1
/C232
3/C33x3
/C23ÿ2
/C23ÿ1
/C232
/C234
5/C33x5ÿ :
7:7b
These series converge for jxj<1. Since Eq. (7.7a) contains even powers of x, and
Eq. (7.7b) contains odd powers of x, the ratio y1=y2is not a constant, and y1and
y2are linearly independent solutions. Hence Eq. (7.6) is a general solution of Eq.
(7.1) on the interval ÿ1<x<1.
In many applications the parameter /C23in Legendre’s equation is a positive
integer n. Then the right hand side of Eq. (7.5) is zero when snand, therefore,
an20 and an40;...:Hence, if nis even, y1
xreduces to a polynomial of
degree n.I fnis odd, the same is true with respect to y2
x. These polynomials,
multiplied by some constants, are called Legendre polynomials. Since they are of
great practical importance, we will consider them in some detail. For this purpose
we rewrite Eq. (7.5) in the form
asÿ
s2
s1
nÿs
ns1as2
7:8
and then express all the non-vanishing coecients in terms of the coecient anof
the highest power of xof the polynomial. The coecient anis then arbitrary. It is
customary to choose an1 when n0 and
an
2n/C33
2n
n/C332135
2nÿ1
n/C33; n1;2;...;
7:9
298SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
the reason being that for this choice of anall those polynomials will have the value
1 when x1. We then obtain from Eqs. (7.8) and (7.9)
anÿ2ÿn
nÿ1
2
2nÿ1anÿn
nÿ1
2n/C33
2
2nÿ12n
n/C332
ÿn
nÿ12n
2nÿ1
2nÿ2/C33/C33
2
2nÿ12nn
nÿ1/C33n
nÿ1
nÿ2/C33;
that is,
anÿ2ÿ
2nÿ2/C33
2n
nÿ1/C33
nÿ2/C33:
Similarly,
anÿ4ÿ
nÿ2
nÿ3
4
2nÿ3anÿ2
2nÿ4/C33
2n2/C33
nÿ2/C33
nÿ4/C33
etc., and in general
anÿ2m
ÿ 1m
2nÿ2m/C33
2nm/C33
nÿm/C33
nÿ2m/C33:
7:10
The resulting solution of Legendre’s equation is called the Legendre polynomial
of degree nand is denoted by Pn
x; from Eq. (7.10) we obtain
Pn
xXM
m0
ÿ1m
2nÿ2m/C33
2nm/C33
nÿm/C33
nÿ2m/C33xnÿ2m
2n/C33
2n
n/C332xnÿ
2nÿ2/C33
2n1/C33
nÿ1/C33
nÿ2/C33xnÿ2ÿ ;
7:11
where Mn=2o r
nÿ1=2, whichever is an integer. In particular (Fig. 7.1)
P0
x1;P1
xx;P2
x1
2
3x2ÿ1;P3
x12
5x3ÿ3x;
P4
x18
35x4ÿ30x23;P5
x18
63x5ÿ70x315x:
Rodrigues’ formula for Pn
x
The Legendre polynomials Pn
xare given by the formula
Pn
x1
2nn/C33dn
dxn
x2ÿ1n:
7:12
We shall establish this result by actually carrying out the indicated di/C128erentia-
tions, using the Leibnitz rule for nth derivative of a product, which we state below
without proof:
299LEGENDRE’S EQUATION
If we write DnuasunandDn/C118as/C118n, then
u/C118nu/C118nnC1u1/C118nÿ1nCrur/C118nÿr un/C118;
where Dd=dxandnCris the binomial coecient and is equal to n/C33=r/C33
nÿr/C33.
We first notice that Eq. (7.12) holds for n0, 1. Then, write
z
x2ÿ1n=2nn/C33
so that
x2ÿ1Dz2nxz:
7:13
Di/C128erentiating Eq. (7.13)
n1times by the Leibnitz rule, we get
1ÿx2Dn2zÿ2xDn1zn
n1Dnz0:
Writing yDnz, we then have:
(i)yis a polynomial.
(ii) The coecient of xnin
x2ÿ1nis
ÿ1n=2nCn=2(neven) or 0 ( nodd).
Therefore the lowest power of xiny
xisx0(neven) or x1(nodd). It
follows that
yn
00
nodd
and
yn
01
2nn/C33
ÿ1n=2nCn=2n/C33
ÿ1n=2n/C33
2n
n=2/C332
neven:
300SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Figure 7.1. Legendre polynomials.
By Eq. (7.11) it follows that
yn
0Pn
0
alln:
(iii)
1ÿx2D2yÿ2xDyn
n1y0, which is Legendre’s equation.
Hence Eq. (7.12) is true for all n.
/C84he generating function for Pn
x
One can prove that the polynomials Pn
xare the coecients of znin the expan-
sion of the function
x;z
1ÿ2xzz2ÿ1=2, with jzj<1; that is,
x;z
1ÿ2xzz2ÿ1=2X1
n0Pn
xzn; zjj<1:
7:14
x;zis called the generating function for Legendre polynomials Pn
x. We shall
be concerned only with the case in which
xcos
ÿ<
and then
z2ÿ2xz1
zÿei
zÿei:
The expansion (7.14) is therefore possible when jzj<1. To prove expansion (7.14)
we have
lhs11
2z
2xÿ113
222/C33z2
2xÿz2
13
2nÿ1
2nn/C33zn
2xÿzn :
The coecient of znin this power series is
13
2nÿ1
2nn/C33
2nxn13
2nÿ3
2nÿ1
nÿ1/C33ÿ
nÿ1
2xnÿ2 Pn
x
by Eq. (7.11). We can use Eq. (7.14) to find successive polynomials explicitly.
Thus, di/C128erentiating Eq. (7.14) with respect to zso that
xÿz
1ÿ2xzz2ÿ3=2X1
n1nznÿ1Pn
x
and using Eq. (7.14) again gives
xÿzP0
xX1
n1Pn
xzn"#
1ÿ2xzz2X1
n1nznÿ1Pn
x:
7:15
301LEGENDRE’S EQUATION
Then expanding coecients of znin Eq. (7.15) leads to the recurrence relation
2n1xPn
x
n1Pn1
xnPnÿ1
x:
7:16
This gives P4;P5;P6, etc. very quickly in terms of P0;P1, and P3.
Recurrence relations are very useful in simplifying work, helping in proofs
or derivations. We list four more recurrence relations below without proofs or
derivations:
xP0
n
xÿP0
nÿ1
xnPn
x;
7:16a
P0
n
xÿxP0
nÿ1
xnPnÿ1
x;
7:16b
1ÿx2P0
n
xnPnÿ1
xÿnxP n
x;
7:16c
2n1Pn
xP0
n1
xÿP0
nÿ1
x:
7:16d
With the help of the recurrence formulas (7.16) and (7.16b), it is straight-
forward to establish the other three. Omitting the full details, which are left for
the reader, these relations can be obtained as follows:
(i) di/C128erentiation of Eq. (7.16) with respect to xand the use of Eq. (7.16b) to
eliminate P0
n1
xleads to relation (7.16a);
(ii) the addition of Eqs. (7.16a) and (7.16b) immediately yields relation
(7.16d);
(iii) the elimination of P0
nÿ1
xbetween Eqs. (7.16b) and (7.16a) gives relation
(7.16c).
Example 7.1The physical significance of expansion (7.14) is apparent in this simple example:
find the potential /C86of a point charge at point Pdue to a charge /C113at/C81.
Solution: Suppose the origin is at O(Fig. 7.2). Then
/C86
P/C113
R/C113
/C262ÿ2r/C26cosr2ÿ1=2:
302SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Figure 7.2.
Thus, if r</C26
/C86P/C113
/C26
1ÿ2zcosz2ÿ1=2; zr=/C26;
which gives
/C86P/C113
/C26X1
n0r
/C26n
Pn
cos
r</C26:
Similarly, when r/C62/C26, we get
/C86P/C113
/C26X1
n0/C26
rn1
Pn
cos:
There are many problems in which it is essential that the Legendre polynomials
be expressed in terms of , the colatitude angle of the spherical coordinate system.
This can be done by replacing xby cos . But this will lead to expressions that are
quite inconvenient because of the powers of cos they contain. Fortunately, using
the generating function provided by Eq. (7.14), we can derive more useful forms in
which cosines of multiples of take the place of powers of cos . To do this, let us
substitute
xcos
eieÿi=2
into the generating function, which gives
1ÿz
eieÿiz2ÿ1=2
1ÿzei
1ÿzeÿiÿ1=2X1
n0Pn
coszn:
Now by the binomial theorem, we have
1ÿzeiÿ1=2X1
n0anzneni;
1ÿzeÿiÿ1=2X1
n0anzneÿni;
where
an135
2nÿ1
246
2n; n1; a01:
7:17
To find the coecient of znin the product of these two series, we need to form the
Cauchy product of these two series. What is a Cauchy product of two series/C63 We
state it below for the reader who is in need of a review:
The Cauchy product of two infinite series,P1
n0un
x
andP1
n0/C118n
x, is defined as the sum over n
X1
n0sn
xX1
n0Xn
k0uk
x/C118nÿk
x;
303LEGENDRE’S EQUATION
where sn
xis given by
sn
xXn
k0uk
x/C118nÿk
xu0
x/C118n
x un
x/C1180
x:
Now the Cauchy product for our two series is given by
X1
n0Xn
k0anÿkznÿke
nÿki
akzkeÿki
X1
n0
znXn
k0akanÿke
nÿ2ki
:
7:18
In the inner sum, which is the sum of interest to us, it is straightforward to prove
that, for n1, the terms corresponding to kjandknÿjare identical except
that the exponents on eare of opposite sign. Hence these terms can be paired, and
we have for the coecient of zn,
Pn
cosa0an
enieÿnia1anÿ1
e
nÿ2ieÿ
nÿ2i
2a0ancosna1anÿ1cos
nÿ2 :
7:19
Ifnis odd, the number of terms is even and each has a place in one of the pairs. In
this case, the last term in the sum is
a
nÿ1=2a
n1=2cos:
Ifnis even, the number of terms is odd and the middle term is unpaired. In this
case, the series (7.19) for Pn
cosends with the constant term
an=2an=2:
Using Eq. (7.17) to compute values of the an, we find from the unit coecient of z0
in Eqs. (7.18) and (7.19), whether nis odd or even, the specific expressions
P0
cos1; P1
coscos;P2
cos
3 cos 2 1=4
P3
cos
5 cos 3 3 cos =8
P4
cos
35 cos 4 20 cos 2 9=64
P5
cos
63 cos 5 35 cos 3 30 cos =128
P6
cos
231 cos 6 126 cos 4 105 cos 2 50=5129
>>>>>>>>=
>>>>>>>>;:
7:20
Orthogonality of /C76egendre polynomials
The set of Legendre polynomials fP
n
xgis orthogonal for ÿ1x1. In
particular we can show that
Z1
ÿ1Pn
xPm
xdx2=
2n1ifmn
0i f m6n:/C26
7:21
304SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
(i)m6n: Let us rewrite the Legendre equation (7.1) for Pm
xin the form
d
dx
1ÿx2P0
m
x/C2/C3
m
m1Pm
x0
7:22
and the one for Pn
x
d
dx
1ÿx2P0
n
x/C2/C3
n
n1Pn
x0:
7:23
We then multiply Eq. (7.22) by Pn
xand Eq. (7.23) by Pm
x, and subtract to get
Pmd
dx
1ÿx2P0
n/C2/C3
ÿPnd
dx
1ÿx2P0
m/C2/C3
n
n1ÿm
m1PmPn0:
The first two terms in the last equation can be written as
d
dx
1ÿx2
PmP0
nÿPnP0
m/C2/C3
:
Combining this with the last equation we have
d
dx
1ÿx2
PmP0
nÿPnP0
m/C2/C3
n
n1ÿm
m1PmPn0:
Integrating the above equation between ÿ1 and 1 we obtain
1ÿx2
PmP0
nÿPnP0
mj1
ÿ1n
n1ÿm
m1Z1
ÿ1Pm
xPn
xdx0:
The integrated term is zero because (1 ÿx20a tx1, and Pm
xandPn
x
are finite. The bracket in front of the integral is not zero since m6n. Therefore
the integral must be zero and we have
Z1
ÿ1Pm
xPn
xdx0; m6n:
(ii)mn: We now use the recurrence relation (7.16a), namely
nPn
xxP0
n
xÿP0
nÿ1
x:
Multiplying this recurrence relation by Pn
xand integrating between ÿ1 and 1,
we obtain
nZ1
ÿ1Pn
x2dxZ1
ÿ1xPn
xP0
n
xdxÿZ1
ÿ1Pn
xP0
nÿ1
xdx:
7:24
The second integral on the right hand side is zero. (Why/C63) To evaluate the first
integral on the right hand side, we integrate by parts
Z1
ÿ1xPn
xP0
n
xdxx
2Pn
x2j1
ÿ1ÿ1
2Z1
ÿ1Pn
x2dx1ÿ12Z
1
ÿ1Pn
x2dx:
305LEGENDRE’S EQUATION
Substituting these into Eq. (7.24) we obtain
nZ1
ÿ1Pn
x2dx1ÿ1
2Z1
ÿ1Pn
x2dx;
which can be simplified to
Z1
ÿ1Pn
x2dx2
2n1:
Alternatively, we can use generating function
1
1ÿ2xzz2p X1
n0Pn
xzn:
We have on squaring both sides of this:
1
1ÿ2xzz2X1
m0X1
n0Pm
xPn
xzmn:
Then by integrating from ÿ1 to 1 we have
Z1
ÿ1dx
1ÿ2xzz2X1
m0X1
n0Z1
ÿ1Pm
xPn
xdx/C26/C27
zmn:
Now
Z1
ÿ1dx
1ÿ2xzz2ÿ1
2zZ1
ÿ1d
1ÿ2xzz2
1ÿ2xzz2ÿ1
2zln
1ÿ2xzz2j1
ÿ1
and
Z1
ÿ1Pm
xPn
xdx0; m6n:
Thus, we have
ÿ1
2zln
1ÿ2xzz2j1ÿ1X1
n0Z1
ÿ1P2
n
xdx/C26/C27
z2n
or
1
zln1z
1ÿz
X1
n0Z1
ÿ1P2
n
xdx/C26/C27
z2n;
that is,
X1
n02z2n
2n1X1
n0Z1
ÿ1P2n
xdx/C26/C27
z2n:
Equating coecients of z2nwe have as requiredR1
ÿ1P2n
xdx2=
2n1.
306SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Since the Legendre polynomials form a complete orthogonal set on ( ÿ1, 1), we
can expand functions in Legendre series just as we expanded functions in Fourier
series:
f
xX1
i0ciPi
x:
The coecients cican be found by a method parallel to the one we used in finding
the formulas for the coecients in a Fourier series. We shall not pursue this line
further.
There is a second solution of Legendre’s equation. However, this solution is
usually only required in practical applications in which jxj/C621 and we shall only
briefly discuss it for such values of x. Now solutions of Legendre’s equation
relative to the regular singular point at infinity can be investigated by writingx
2t. With this substitution,
dy
dxdy
dtdt
dx2t1=2dy
dtandd2y
dx2d
dxdy
dx
2dy
dx4td2y
dt2;
and Legendre’s equation becomes, after some simplifications,
t
1ÿtd2y
dt21
2ÿ32t
dy
dt/C23
/C231
4y0:
This is the hypergeometric equation with ÿ/C23=2;/C12
1/C23=2, and /C131
2:
x
1ÿxd2y
dx2/C13ÿ
/C121xdy
dxÿ/C12y0;
we shall not seek its solutions. The second solution of Legendre’s equation is
commonly denoted by /C81/C23
xand is called the Legendre function of the second
kind of order /C23. Thus the general solution of Legendre’s equation (7.1) can be
written
yAP/C23
xB/C81 /C23
x;
AandBbeing arbitrary constants. P/C23
xis called the Legendre function of the
first kind of order /C23and it reduces to the Legendre polynomial Pn
xwhen /C23is an
integer n.
/C84he associated Legendre functions
These are the functions of integral order which are solutions of the associatedLegendre equation
1ÿx
2y00ÿ2xy0n
n1ÿm2
1ÿx2()
y0
7:25
with m2n2.
307THE ASSOCIATED LEGENDRE FUNCTIONS
We could solve Eq. (7.25) by series; but it is more useful to know how the
solutions are related to Legendre polynomials, so we shall proceed in the follow-
ing way. We write
y
1ÿx2m=2u
x
and substitute into Eq. (7.25) whence we get, after a little simplification,
1ÿx2u00ÿ2
m1xu0n
n1ÿm
m1u0:
7:26
Form0, this is a Legendre equation with solution Pn
x. Now we di/C128erentiate
Eq. (7.26) and get
1ÿx2
u000ÿ2
m11x
u00n
n1ÿ
m1
m2u00:
7:27
Note that Eq. (7.27) is just Eq. (7.26) with u0in place of u, and ( m1) in place of
m. Thus, if Pn
xis a solution of Eq. (7.26) with m0,P0
n
xis a solution of Eq.
(7.26) with m1,P00
n
xis a solution with m2, and in general for integral
m;0mn;
dm=dxmPn
xis a solution of Eq. (7.26). Then
y
1ÿx2m=2dm
dxmPn
x
7:28
is a solution of the associated Legendre equation (7.25). The functions in Eq.(7.28) are called associated Legendre functions and are denoted by
P
m
n
x
1ÿx2m=2dm
dxmPn
x:
7:29
Some authors include a factor ( ÿ1min the definition of Pm
n
x:
A negative value of min Eq. (7.25) does not change m2, so a solution of Eq.
(7.25) for positive mis also a solution for the corresponding negative m. Thus
many references define Pm
n
xforÿnmnas equal to Pjmj
n
x.
When we write xcos, Eq. (7.25) becomes
1
sind
dsindy
d
n
n1ÿm2
sin2()
y0
7:30
and Eq. (7.29) becomes
Pm
n
cossinmdm
d
cosmPn
cos fg :
In particular
Dÿ1meansZx
1Pn
xdx:
308SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Orthogonality of associated /C76egendre functions
As in the case of Legendre polynomials, the associated Legendre functions Pm
n
x
are orthogonal for ÿ1x1 and in particular
Z1
ÿ1Ps
m
xPs
n
xdx
ns/C33
nÿs/C33mn:
7:31
To prove this, let us write for simplicity
MPm
s
x;and /C78Psn
x
and from Eq. (7.25), the associated Legendre equation, we have
d
dx
1ÿx2dM
dx/C26/C27
m
m1ÿs2
1ÿx2()
M0
7:32
and
d
dx
1ÿx2d/C78
dx/C26/C27
n
n1ÿs2
1ÿx2()
/C780:
7:33
Multiplying Eq. (7.32) by /C78, Eq. (7.33) by Mand subtracting, we get
Md
dx
1ÿx2d/C78
dx/C26/C27
ÿ/C78d
dx
1ÿx2dM
dx/C26/C27
fm
m1ÿn
n1gM/C78:
Integration between ÿ1 and 1 gives
mÿn
mnÿ1Z1
ÿ1M/C78dx Z1
ÿ1
Md
dx
1ÿx2d/C78
dx/C26/C27
ÿ/C78d
dx
1ÿx2dM
dx/C26/C27
dx:
7:34
Integration by parts gives
Z1
ÿ1Md
dxf
1ÿx2/C780gdxM/C780
1ÿx21
ÿ1ÿZ1
ÿ1
1ÿx2M0/C780dx
ÿZ1
ÿ1
1ÿx2M0/C780dx:
309THE ASSOCIATED LEGENDRE FUNCTIONS
Then integrating by parts once more, we obtain
Z1
ÿ1Md
dxf
1ÿx2/C780gdxÿZ1
ÿ1
1ÿx2M0/C780dx
ÿ M/C78
1ÿx21
ÿ1Z1
ÿ1/C78d
dx
1ÿx2M0/C8/C9
dx
Z1
ÿ1/C78d
dxf
1ÿx2M0gdx:
Substituting this in Eq. (7.34) we get
mÿn
mnÿ1Z1
ÿ1M/C78dx 0:
Ifmn, we have
Z1
ÿ1M/C78dx Z1
ÿ1Ps
m
xPsm
xdx0
m6n:
Ifmn, let us write
Ps
n
x
1ÿx2s=2ds
dxsPn
x
1ÿx2s=2
2nn/C33dsn
dxsnf
x2ÿ1ng:
Hence
Z1
ÿ1Ps
n
xPsn
xdx1
22n
n/C332Z1
ÿ1
1ÿx2sDnsf
x2ÿ1ngDns
f
x2ÿ1ngdx;
Dkdk=dxk:
Integration by parts gives
1
22n
n/C332
1ÿx2sDnsf
x2ÿ1ngDnsÿ1f
x2ÿ1ng1
ÿ1
ÿ1
22n
n/C332Z1
ÿ1D
1ÿx2sDns
x2ÿ1n/C8/C9 /C2/C3
Dnsÿ1
x2ÿ1n/C8/C9
dx:
The first term vanishes at both limits and we have
Z1
ÿ1fPs
n
xg2dxÿ1
22n
n/C332Z1
ÿ1D
1ÿx2sDnsf
x2ÿ1ngDnsÿ1f
x2ÿ1ngdx:
7:35
We can continue to integrate Eq. (7.35) by parts and the first term continues to
vanish since D/C112
1ÿx2sDnsf
x2ÿ1ngcontains the factor
1ÿx2when /C112<s
310SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
andDnsÿ/C112f
x2ÿ1ngcontains it when /C112s. After integrating
nstimes we
find
Z1
ÿ1fPs
n
xg2dx
ÿ1ns
22n
n/C332Z1
ÿ1Dns
1ÿx2sDnsf
x2ÿ1ng
x2ÿ1ndx:
7:36
But Dnsf
x2ÿ1ngis a polynomial of degree ( nÿs) so that
(1ÿx2sDnsf
x2ÿ1ngis of degree nÿ22sns. Hence the first factor
in the integrand is a polynomial of degree zero. We can find this constant by
examining the following:
Dns
x2n2n
2nÿ1
2nÿ2
nÿ1xnÿs:
Hence the highest power in
1ÿx2sDnsf
x2ÿ1ngis the term
ÿ1s2n
2nÿ1
nÿs1xns;
so that
Dns
1ÿx2sDnsf
x2ÿ1ng
ÿ 1s
2n/C33
ns/C33
nÿs/C33:
Now Eq. (7.36) gives, by writing xcos,
Z1
ÿ1Ps
nf
xg2dx
ÿ1n
22n
n/C332Z1
ÿ1
2n/C33
ns/C33
nÿs/C33
x2ÿ1ndx
2
2n1
ns/C33
nÿs/C33
7:37
/C72ermite/C39s equation
Hermite’s equation is
y00ÿ2xy02/C23y0;
7:38
where y0dy=dx. The reader will see this equation in quantum mechanics (when
solving the Schro /C200dinger equation for a linear harmonic potential function).
The origin x0 is an ordinary point and we may write the solution in the form
ya0a1xa2x2X1
j0ajxj:
7:39
Di/C128erentiating the series term by term, we have
y0X1
j0jajxjÿ1; y00X1
j0
j1
j2aj2xj:
311HERMITE’S EQUATION
Substituting these into Eq. (7.38) we obtain
X1
j0
j1
j2aj22
/C23ÿjaj/C2/C3
xj0:
For a power series to vanish the coecient of each power of xmust be zero; this
gives
j1
j2aj22
/C23ÿjaj0;
from which we obtain the recurrence relations
aj22
jÿ/C23
j1
j2aj:
7:40
We obtain polynomial solutions of Eq. (7.38) when /C23n, a positive integer. Then
Eq. (7.40) gives
an2an4 0:
For even n, Eq. (7.40) gives
a2
ÿ 12n
2/C33a0;a4
ÿ 1222
nÿ2n
4/C33a0;a6
ÿ 1323
nÿ4
nÿ2n
6/C33a0
and generally
an
ÿ 1n=22n=2n
nÿ242
n/C33a0:
This solution is called a Hermite polynomial of degree nand is written Hn
x.I f
we choose
a0
ÿ1n=22n=2n/C33
n
nÿ242
ÿ1n=2n/C33
n=2/C33
we can write
Hn
x
2xnÿn
nÿ1
1/C33
2xnÿ2n
nÿ1
nÿ2
nÿ3
2/C33
2xnÿ4 :
7:41
When nis odd the polynomial solution of Eq. (7.38) can still be written as Eq.
(7.41) if we write
a1
ÿ1
nÿ1=22n/C33
n=2ÿ1=2/C33:
In particular,
H0
x1;H1
x2x;H3
x4x2ÿ2;H3
x8x2ÿ12x;
H4
x16x4ÿ48x212;H5
x32x5ÿ160x3120x;...:
312SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Rodrigues’ formula for /C72ermite polynomials /C72n
x
The Hermite polynomials are also given by the formula
Hn
x
ÿ 1nex2dn
dxn
eÿx2:
7:42
To prove this formula, let us write /C113eÿx2. Then
D/C1132x/C1130;Dd
dx:
Di/C128erentiate this ( n1) times by the Leibnitz’ rule giving
Dn2/C1132xDn1/C1132
n1Dn/C1130:
Writing y
ÿ 1nDn/C113gives
D2y2xDy 2
n1y0
7:43
substitute uex2ythen
Duex2f2xyDyg
and
D2uex2fD2y4xDy 4x2y2yg:
Hence by Eq. (7.43) we get
D2uÿ2xDu 2nu0;
which indicates that
u
ÿ 1nex2Dn
eÿx2
is a polynomial solution of Hermite’s equation (7.38).
Recurrence relations for /C72ermite polynomials
Rodrigues’ formula gives on di/C128erentiation
H0
n
x
ÿ 1n2xex2Dn
eÿx2
ÿ 1nex2Dn1
eÿx2:
that is,
H0
n
x2xHn
xÿHn1
x:
7:44
Eq. (7.44) gives on di/C128erentiation
H00
n
x2Hn
x2xH0
n
xÿH0
n1
x:
313HERMITE’S EQUATION
Now Hn
xsatisfies Hermite’s equation
H00
n
xÿ2xH0
n
x2nHn
x0:
Eliminating H00
n
xfrom the last two equations, we obtain
2xH0
n
xÿ2nHn
x2Hn
x2xH0
n
xÿH0
n1
x
which reduces to
H0
n1
x2
n1Hn
x:
7:45
Replacing nbyn1 in Eq. (7.44), we have
H0
n1
x2xHn1
xÿHn2
x:
Combining this with Eq. (7.45) we obtain
Hn2
x2xHn1
xÿ2
n1Hn
x:
7:46
This will quickly give the higher polynomials.
Generating function for the /C72n
x
By using Rodrigues’ formula we can also find a generating formula for the Hn
x.
This is
x;te2txÿt2efx2ÿ
tÿx2gX1
n0Hn
x
n/C33tn:
7:47
Di/C128erentiating Eq. (7.47) ntimes with respect to twe get
ex2/C64n
/C64tneÿ
tÿx2ex2
ÿ1n/C64n
/C64xneÿ
tÿx2X1
k0Hnk
xtk
k/C33:
Putt0 in the last equation and we obtain Rodrigues’ formula
Hn
x
ÿ 1nex2dn
dxn
eÿx2:
/C84he orthogonal /C72ermite functions
These are defined by
/C70n
xeÿx2=2Hn
x;
7:48
from which we have
D/C70n
xÿ x/C70n
xeÿx2=2H0
n
x;
D2/C70n
xeÿx2=2H00
n
xÿ2xeÿx2=2H0
n
xx2eÿx2=2Hn
xÿ/C70n
x
eÿx2=2H00
n
xÿ2xH0
n
x x2/C70n
xÿ/C70n
x;
314SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
butH00
n
xÿ2xH0
n
xÿ 2nHn
x, so we can rewrite the last equation as
D2/C70n
xeÿx2=2ÿ2nH0
n
x x2/C70n
xÿ/C70n
x
ÿ2n/C70n
xx2/C70n
xÿ/C70n
x;
which gives
D2/C70n
xÿx2/C70n
x
2n1/C70n
x0:
7:49
We can now show that the set f/C70n
xgis orthogonal in the infinite range
ÿ1 <x<1. Multiplying Eq. (7.49) by /C70m
xwe have
/C70m
xD2/C70n
xÿx2/C70n
x/C70m
x
2n1/C70n
x/C70m
x0:
Interchanging mandngives
/C70n
xD2/C70m
xÿx2/C70m
x/C70n
x
2m1/C70m
x/C70n
x0:
Subtracting the last two equations from the previous one and then integrating
from ÿ1 to1,w eh a v e
In;mZ1
ÿ1/C70n
x/C70m
xdx1
2
nÿmZ1
ÿ1
/C7000
n/C70mÿ/C7000
m/C70ndx:
The integration by parts gives
2
nÿmIn;m/C700
n/C70mÿ/C700
m/C70n/C2/C31
ÿ1ÿZ1
ÿ1
/C700
n/C700
mÿ/C700
m/C700
ndx:
Since the right hand side vanishes at both limits and if m6m, we have
In;mZ1
ÿ1/C70n
x/C70m
xdx0:
7:50
When nmwe can proceed as follows
In;nZ1
ÿ1eÿx2Hn
xHn
xdxZ1
ÿ1ex2Dn
eÿx2Dm
eÿx2dx:
Integration by parts, that is,R
ud/C118u/C118ÿR
/C118du with ueÿx2Dn
eÿx2
and/C118Dnÿ1
eÿx2, gives
In;nÿZ1
ÿ12xex2Dn
eÿx2ex2Dn1
eÿx2Dnÿ1
eÿx2dx:
By using Eq. (7.43) which is true for y
ÿ 1nDn/C113
ÿ 1nDn
eÿx2we obtain
In;nZ1
ÿ12nex2Dnÿ1
eÿx2Dnÿ1
eÿx2dx2nInÿ1;nÿ1:
315HERMITE’S EQUATION
Since
I0;0Z1
ÿ1eÿx2dxÿ
1=2p;
we find that
In;nZ1
ÿ1eÿx2Hn
xHn
xdx2nn/C33p:
7:51
We can also use the generating function for the Hermite polynomials:
e2txÿt2X1
n0Hn
xtn
n/C33;e2sxÿs2X1
m0Hm
xsm
m/C33:
Multiplying these, we have
e2txÿt22sxÿs2X1
m0X1
n0Hm
xHn
xsmtn
m/C33n/C33:
Multiplying by eÿx2and integrating from ÿ1 to1gives
Z1
ÿ1eÿ
xst2ÿ2stdxX1
m0X1
n0smtn
m/C33n/C33Z1
ÿ1eÿx2Hm
xHn
xdx:
Now the left hand side is equal to
e2stZ1
ÿ1eÿ
xst2dxe2stZ1
ÿ1eÿu2due2stppX
1
m02msmtm
m/C33:
By equating coecients the required result follows.
It follows that the functions
1=2nn/C33np1=2eÿx2Hn
xform an orthonormal set.
We shall assume it is complete.
Laguerre/C39s equation
Laguerre’s equation is
xD2y
1ÿxDy/C23y0:
7:52
This equation and its solutions (Laguerre functions) are of interest in quantum
mechanics (e.g., the hydrogen problem). The origin x0 is a regular singular
point and so we write
y
xX1
k0akxk/C26:
7:53
By substitution, Eq. (7.52) becomes
X1
k0
k/C262akxk/C26ÿ1
/C23ÿk/C26akxk0
7:54
316SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
from which we find that the indicial equation is /C2620. And then (7.54) reduces to
X1
k0k2akxkÿ1
/C23ÿkakxk0:
Changing kÿ1t ok0in the first term, then renaming k0k, we obtain
X1
k0f
k12ak1
/C23ÿkakgxk0;
whence the recurrence relations are
ak1kÿ/C23
k12ak:
7:55
When /C23is a positive integer n, the recurrence relations give ak1ak2 0,
and
a1ÿn
12a0; a2ÿ
nÿ1
22a1
ÿ12
nÿ1n
122a0;
a3ÿ
nÿ2
32a2
ÿ13
nÿ2
nÿ1n
1232a0;etc:
In general
ak
ÿ 1k
nÿk1
nÿk2
nÿ1n
k/C332a0:
7:56
We usually choose a0
ÿ 1n/C33, then the polynomial solution of Eq. (7.52) is given
by
Ln
x
ÿ 1nxnÿn2
1/C33xnÿ1n2
nÿ12
2/C33xnÿ2ÿ
ÿ 1nn/C33()
:
7:57
This is called the Laguerre polynomial of degree n. We list the first four Laguerre
polynomials below:
L0
x1;L1
x1ÿx;L2
x2ÿ4xx2;L3
x6ÿ18x9x2ÿx3:
/C84he generating function for the /C76aguerre polynomials Ln
x
This is given by
x;zeÿxz=
1ÿz
1ÿzX1
n0Ln
x
n/C33zn:
7:58
317LAGUERRE’S EQUATION
By writing the series for the exponential and collecting powers of z, you can verify
the first few terms of the series. And it is also straightforward to show that
x/C642
/C64x2
1ÿx/C64
/C64xz/C64
/C64z0:
Substituting the right hand side of Eq. (7.58), that is,
x;zP1
n0Ln
x=n/C33zn,
into the last equation we see that the functions Ln
xsatisfy Laguerre’s equation.
Thus we identify
x;zas the generating function for the Laguerre polynomials.
Now multiplying Eq. (7.58) by zÿnÿ1and integrating around the origin, we
obtain
Ln
xn/C33
2iIeÿxz=
1ÿz
1ÿzzn1dz;
7:59
which is an integral representation of Ln
x.
By di/C128erentiating the generating function in Eq. (7.58) with respect to xandz,
we obtain the recurrence relations
Ln1
x
2n1ÿxLn
xÿn2Lnÿ1
x;
nLnÿ1
xnL0
nÿ1
xÿL0
n
x:)
7:60
Rodrigues’ formula for the /C76aguerre polynomials Ln
x
The Laguerre polynomials are also given by Rodrigues’ formula
Ln
xexdn
dxn
xneÿx:
7:61
To prove this formula, let us go back to the integral representation of Ln
x, Eq.
(7.59). With the transformation
xz
1ÿzsÿxorzsÿx
s;
Eq. (7.59) becomes
Ln
xn/C33ex
2iIsneÿn
sÿxn1ds;
the new contour enclosing the point sxin the splane. By Cauchy’s integral
formula (for derivatives) this reduces to
Ln
xexdn
dxn
xneÿx;
which is Rodrigues’ formula.
Alternatively, we can di/C128erentiate Eq. (7.58) ntimes with respect to zand
afterwards put z/C610, and thus obtain
exlim
z!0/C64n
/C64zn
1ÿzÿ1expÿx
1ÿz /C104/C105
Ln
x:
318SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
But
lim
z!0/C64n
/C64zn
1ÿzÿ1expÿx
1ÿz /C104/C105
dn
dxnxneÿx
;
hence
Ln
xexdn
dxn
xneÿx:
/C84he orthogonal /C76aguerre functions
The Laguerre polynomials, Ln
x, do not by themselves form an orthogonal set.
But the functions eÿx=2Ln
xare orthogonal in the interval (0, 1). For any two
Laguerre polynomials Lm
xandLn
xwe have, from Laguerre’s equation,
xL00
m
1ÿxL0
mmLm0;
xL00
n
1ÿxL0
nmLn0:
Multiplying these equations by Ln
xandLm
xrespectively and subtracting, we
find
xLnL00
mÿLmL00
n
1ÿxLnL0
mÿLmL0
n
nÿmLmLn
or
d
dxLnL0
mÿLmL0
n1ÿx
xLnL0
mÿLmL0
n
nÿmLmLn
x:
Then multiplying by the integrating factor
expZ
1ÿx=xdxexp
lnxÿxxeÿx;
we have
d
dxfxeÿxLnL0
mÿLmL0
ng
nÿmeÿxLmLn:
Integrating from 0 to 1gives
nÿmZ1
0eÿxLm
xLn
xdxxeÿxLnL0
mÿLmL0
nj1
00:
Thus if m6n
Z1
0eÿxLm
xLn
xdx0
m6n;
7:62
which proves the required result.
Alternatively, we can use Rodrigues’ formula (7.61). If mis a positive integer,
Z1
0eÿxxmLm
xdxZ1
0xmdn
dxn
xneÿxdx
ÿ 1mm/C33Z1
0dnÿm
dxnÿm
xneÿxdx;
7:63
319LAGUERRE’S EQUATION
the last step resulting from integrating by parts mtimes. The integral on the right
hand side is zero when n/C62mand, since Ln
xis a polynomial of degree minx,i t
follows that
Z1
0eÿxLm
xLn
xdx0
m6n;
which is Eq. (7.62). The reader can also apply Eq. (7.63) to show that
Z1
0eÿxLn
x fg2dx
n/C332:
7:64
Hence the functions feÿx=2Ln
x=n/C33gform an orthonormal system.
/C84he associated Laguerre pol/C121nomials Lm
n
x
Di/C128erentiating Laguerre’s equation (7.52) mtimes by the Leibnitz theorem we
obtain
xDm2y
m1ÿxDm1y
nÿmDmy0
/C23n
and writing zDmywe obtain
xD2z
m1ÿxDz
nÿmz0:
7:65
This is Laguerre’s associated equation and it clearly possesses a polynomial solu-
tion
zDmLn
xLm
n
x
mn;
7:66
called the associated Laguerre polynomial of degree ( nÿm). Using Rodrigues’
formula for Laguerre polynomial Ln
x, Eq. (7.61), we obtain
Lmn
xdm
dxmLn
xdm
dxmexdn
dxn
xneÿx/C26/C27
:
7:67
This result is very useful in establishing further properties of the associated
Laguerre polynomials. The first few polynomials are listed below:
L0
0
x1;L01
x1ÿx;L11
xÿ 1;
L02
x2ÿ4xx2;L12
xÿ 42x;L22
x2:
Generating function for the associated /C76aguerre polynomials
The Laguerre polynomial Ln
xcan be generated by the function
1
1ÿtexpÿxt
1ÿt
X1
n0Ln
xtn
n/C33:
320SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Di/C128erentiating this ktimes with respect to x, it is seen at once that
ÿ1k
1ÿtÿ1t
1ÿt k
expÿxt
1ÿt
X1
kLk
x
/C33t:
7:68
/C65ssociated /C76aguerre function of integral order
A function of great importance in quantum mechanics is the associated Laguerre
function that is defined as
/C71m
n
xeÿx=2x
mÿ1=2Lmn
x
mn:
7:69
It is significant largely because j/C71m
n
xj ! 0a sx!1 . It satisfies the di/C128erential
equation
x2D2u2xDu nÿmÿ1
2
xÿx2
4ÿm2ÿ1
4"#
u0:
7:70
If we substitute ueÿx=2x
mÿ1=2zin this equation, it reduces to Laguerre’s asso-
ciated equation (7.65). Thus u/C71mnsatisfies Eq. (7.70). You will meet this equa-
tion in quantum mechanics in the study of the hydrogen atom.
Certain integrals involving /C71mnare often used in quantum mechanics and they
are of the form
In;mZ1
0eÿxxkÿ1Lkn
xLkm
xx/C112dx;
where pis also an integer. We will not consider these here and instead refer the
interested reader to the following book: The Mathematics of Physics and
/C67hemistry , by Henry Margenau and George M. Murphy; D. Van Nostrand Co.
Inc., New York, 1956.
/C66essel/C39s equation
The di/C128erential equation
x2y00xy0
x2ÿ2y0
7:71
in which is a real and positive constant, is known as Bessel’s equation and its
solutions are called Bessel functions. These functions were used by Bessel
(Friedrich Wilhelm Bessel, 1784–1864, German mathematician and astronomer)
extensively in a problem of dynamical astronomy. The importance of this
equation and its solutions (Bessel functions) lies in the fact that they occur fre-
quently in the boundary-value problems of mathematical physics and engineering
321BESSEL’S EQUATION
involving cylindrical symmetry (so Bessel functions are sometimes called cylind-
rical functions), and many others. There are whole books on Bessel functions.
The origin is a regular singular point, and all other values of x are ordinary
points. At the origin we seek a series solution of the form
y
xX1
m0amxm/C26
a060:
7:72
Substituting this and its derivatives into Bessel’s equation (7.71), we have
X1
m0
m/C26
m/C26ÿ1amxm/C26X1
m0
m/C26amxm/C26
X1
m0amxm/C262ÿ2X1
m0amxm/C260:
This will be an identity if and only if the coecient of every power of xis zero. By
equating the sum of the coecients of xk/C26to zero we find
/C26
/C26ÿ1a0/C26a0ÿ2a00
k0;
7:73a
/C26ÿ1/C26a1
/C261a1ÿ2a10
k1;
7:73b
k/C26
k/C26ÿ1ak
k/C26akakÿ2ÿ2ak0
k2;3;...:
7:73c
From Eq. (7.73a) we obtain the indicial equation
/C26
/C26ÿ1/C26ÿ2
/C26
/C26ÿ0:
The roots are /C26. We first determine a solution corresponding to the positive
root. For /C26, Eq. (7.73b) yields a10, and Eq. (7.73c) takes the form
k2kakakÿ20; or akÿ1
k
k2akÿ2;
7:74
which is a recurrence formula: since a10 and 0, it follows that
a30;a50;...;successively. If we set k2min Eq. (7.74), the recurrence
formula becomes
a2mÿ1
22m
ma2mÿ2; m1;2;...
7:75
and we can determine the coecients a2;a4, successively. We can rewrite a2min
terms of a0:
a2m
ÿ1m
22mm/C33
m
2
1a0:
322SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Now a2mis the coecient of x2min the series (7.72) for y. Hence it would be
convenient if a2mcontained the factor 22min its denominator instead of just 22m.
To achieve this, we write
a2m
ÿ1m
22mm/C33
m
2
1
2a0:
Furthermore, the factors
m
2
1
suggest a factorial. In fact, if were an integer, a factorial could be created by
multiplying numerator by /C33. However, since is not necessarily an integer, we
must use not /C33but its generalization ÿ
1for this purpose. Then, except for
the values
ÿ1;ÿ2;ÿ3;...
for which ÿ
1is not defined, we can write
a2m
ÿ1m
22mm/C33
m
2
1ÿ
12ÿ
1a0:
Since the gamma function satisfies the recurrence relation zÿ
zÿ
z1, the
expression for a2mbecomes finally
a2m
ÿ1m
22mm/C33ÿ
m12ÿ
1a0:
Since a0is arbitrary, and since we are looking only for particular solutions, we
choose
a01
2ÿ
1;
so that
a2m
ÿ1m
22mm/C33ÿ
m1; a2m10
and the series for yis, from Eq. (7.72),
y
xx 1
2ÿ
1ÿx2
22ÿ
2x4
242/C33ÿ
3ÿ"#
X1
m0
ÿ1m
22mm/C33ÿ
m1x2m:
7:76
The function defined by this infinite series is known as the Bessel function of the
first kind of order and is denoted by the symbol /C74
x. Since Bessel’s equation
323BESSEL’S EQUATION
of order has no finite singular points except the origin, the ratio test will show
that the series for /C74
xconverges for all values of xif0.
When n, an integer, solution (7.76) becomes, for n0
/C74n
xxnX1
m0
ÿ1mx2m
22mnm/C33
nm/C33:
7:76a
The graphs of /C740
x;/C741
x, and /C742
xare shown in Fig. 7.3. Their resemblance to
the graphs of cos xand sin xis interesting (Problem 7.16 illustrates this for the
first few terms). Fig. 7.3 also illustrates the important fact that for every value of
the equation /C74
x0 has infinitely many real roots.
With the second root /C26ÿof the indicial equation, the recurrence relation
takes the form (from Eq. (7.73c))
akÿ1
k
kÿ2akÿ2:
7:77
Ifis not an integer, this leads to an independent second solution that can be
written
/C74ÿ
xX1
m0
ÿ1m
m/C33ÿ
ÿm1
x=2ÿ2m
7:78
and the complete solution of Bessel’s equation is then
y
xA/C74
xB/C74ÿ
x;
7:79
where AandBare arbitrary constants.
When is a positive integer n, it can be shown that the formal expression for
/C74ÿn
xis equal to ( ÿ1n/C74n
x.S o/C74n
xand/C74ÿn
xare linearly dependent and Eq.
(7.79) cannot be a general solution. In fact, if is a positive integer, the recurrence
324SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Figure 7.3. Bessel functions of the first kind.
relation (7.77) breaks down when 2 kand a second solution has to be found
by other methods. There is a diculty also when 0, in which case the two
roots of the indicial equation are equal; the second solution must also found by
other methods. These will be discussed in next section.
The results of Problem 7.16 are a special case of an important general theorem
which states that /C74
xis expressible in finite terms by means of algebraic and
trigonometrical functions of xwhenever is half of an odd integer. Further
examples are
/C743=2
x2
x1=2sinx
xÿcosx
;
/C74ÿ5=2
x2
x1=23 sinx
x3
x2ÿ1
cosx/C26/C27
:
The functions /C74
n1=2
xand/C74ÿ
n1=2
x, where nis a positive integer or zero, are
called spherical Bessel functions; they have important applications in problems of
wave motion in which spherical polar coordinates are appropriate.
Bessel functions of the second /C107ind /C89n
x
For integer n;/C74n
xand/C74ÿn
xare linearly dependent and do not form a
fundamental system. We shall now obtain a second independent solution, startingwith the case n0. In this case Bessel’s equation may be written
xy
00y0xy0;
7:80
the indicial equation (7.73a) now, with 0, has the double root /C260. Then we
see from Eq. (7.33) that the desired solution must be of the form
y2
x/C740
xlnxX1
m1Amxm:
7:81
Next we substitute y2and its derivatives
y0
2/C740
0lnx/C740
xX1
m1mA mxmÿ1;
y00
2/C7400
0lnx2/C740
0
xÿ/C740
x2X1
m1m
mÿ1Amxmÿ2
into Eq. (7.80). Then the logarithmic terms disappear because /C740is a solution of
Eq. (7.80), the other two terms containing /C740cancel, and we find
2/C740
0X1
m1m
mÿ1Amxmÿ1X1
m1mA mxmÿ1X1
m1Amxm10:
325BESSEL’S EQUATION
From Eq. (7.76a) we obtain /C740
0as
/C740
0
xX1
m1
ÿ1m2mx2mÿ1
22m
m/C332X1
m1
ÿ1mx2mÿ1
22mÿ1m/C33
mÿ1/C33:
By inserting this series we have
X1
m1
ÿ1mx2mÿ1
22mÿ2m/C33
mÿ1/C33X1
m1m2Amxmÿ1X1
m1Amxm10:
We first show that Amwith odd subscripts are all zero. The coecient of the
power x0isA1and so A10. By equating the sum of the coecients of the power
x2sto zero we obtain
2s12A2s1A2sÿ10; s1;2;...:
Since A10, we thus obtain A30;A50;...;successively. We now equate the
sum of the coecients of x2s1to zero. For s0 this gives
ÿ14A20o r A21=4:
For the other values of swe obtain
ÿ1s1
2s
s1/C33s/C33
2s22A2s2A2s0:
Fors1 this yields
1=816A4A20o r A4ÿ3=128
and in general
A2m
ÿ1mÿ1
2m
m/C33211
2131
m
; m1;2;...:
7:82
Using the short notation
/C104m11
2131
m
and inserting Eq. (7.82) and A1A3 0 into Eq. (7.81) we obtain the result
y2
x/C740
xlnxX1
m1
ÿ1mÿ1/C104m
22m
m/C332x2m
/C740
xlnx14x
2ÿ3
128x4ÿ :
7:83
Since /C740andy2are linearly independent functions, they form a fundamental
system of Eq. (7.80). Of course, another fundamental system is obtained by
replacing y2by an independent particular solution of the form a
y2b/C740,
where a
60and bare constants. It is customary to choose a2=and
326SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
b/C13ÿln 2, where /C130:577 215 664 90 ...is the so-called Euler constant, which
is defined as the limit of
11
21
sÿlns
assapproaches infinity. The standard particular solution thus obtained is known
as the Bessel function of the second kind of order zero or Neumann’s function of
order zero and is denoted by Y0
x:
Y0
x2
/C740
xlnx
2/C13
X1
m1
ÿ1mÿ1/C104m
22m
m/C332x2m
:
7:84
If1;2;...;a second solution can be obtained by similar manipulations,
starting from Eq. (7.35). It turns out that in this case also the solution contains a
logarithmic term. So the second solution is unbounded near the origin and is
useful in applications only for x60.
Note that the second solution is defined di/C128erently, depending on whether the
order is integral or not. To provide uniformity of formalism and numerical
tabulation, it is desirable to adopt a form of the second solution that is valid forall values of the order. The common choice for the standard second solution
defined for all is given by the formula
Y
x/C74
xcosÿ/C74ÿ
x
sin;Yn
xlim
!nY
x:
7:85
This function is known as the Bessel function of the second kind of order .I ti s
also known as Neumann’s function of order and is denoted by /C78
x(Carl
Neumann 1832–1925, German mathematician and physicist). In G. N. Watson’sA Treatise on the Theory of Bessel /C70unctions (2nd ed. Cambridge University Press,
Cambridge, 1944), it was called Weber’s function and the notation Y
xwas
used. It can be shown that
Yÿn
x
ÿ 1nYn
x:
We plot the first three Yn
xin Fig. 7.4.
A general solution of Bessel’s equation for all values of can now be written:
y
xc1/C74
xc2Y
x:
In some applications it is convenient to use solutions of Bessel’s equation that
are complex for all values of x, so the following solutions were introduced
H
1
x/C74
xiY
x;
H
2
x/C74
xÿiY
x:9
=
;
7:86
327BESSEL’S EQUATION
These linearly independent functions are known as Bessel functions of the third
kind of order or first and second Hankel functions of order (Hermann
Hankel, 1839–1873, German mathematician).
To illustrate how Bessel functions enter into the analysis of physical problems,
we consider one example in classical physics: small oscillations of a hanging chain,
which was first considered as early as 1732 by Daniel Bernoulli.
/C72anging /C175exible chain
Fig. 7.5 shows a uniform heavy flexible chain of length lhanging vertically under
its own weight. The x-axis is the position of stable equilibrium of the chain and its
lowest end is at x0. We consider the problem of small oscillations in the vertical
xyplane caused by small displacements from the stable equilibrium position. This
is essentially the problem of the vibrating string which we discussed in Chapter 4,
with two important di/C128erences: here, instead of being constant, the tension Tat a
given point of the chain is equal to the weight of the chain below that point, andnow one end of the chain is free, whereas before both ends were fixed. The
analysis of Chapter 4 generally holds. To derive an equation for y, consider an
element dx, then Newton’s second law gives
T/C64y
/C64x
2ÿT/C64y
/C64x
1/C26dx/C642y
/C64t2
or
/C26dx/C642y
/C64t2/C64
/C64xT/C64y
/C64x
dx;
328SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Figure 7.4. Bessel functions of the second kind.
from which we obtain
/C26/C642y
/C64t2/C64
/C64xT/C64y
/C64x
:
Now T/C26/C103x. Substituting this into the above equation for y, we obtain
/C642y
/C64t2/C103/C64y
/C64x/C103x/C642y
/C64x2;
where yis a function of two variables xandt. The first step in the solution is to
separate the variables. Let us attempt a solution of the form y
x;tu
xf
t.
Substitution of this into the partial di/C128erential equation yields two equations:
f00
t/C332f
t0;xu00
xu0
x
/C332=/C103u
x0;
where /C332is the separation constant. The di/C128erential equation for f
tis ready for
integration and the result is f
tcos
/C33tÿ, with a phase constant. The
di/C128erential equation for u
xis not in a recognizable form yet. To solve it, first
change variables by putting
x/C103z2=4;/C119
zu
x;
then the di/C128erential equation for u
xbecomes Bessel’s equation of order zero:
z/C11900
z/C1190
z/C332z/C119
z0:
Its general solution is
/C119
zA/C740
/C33zBY0
/C33z
or
u
xA/C7402/C33x
/C103/C114
BY02/C33x
/C103/C114
:
329BESSEL’S EQUATION
Figure 7.5. A flexible chain.
Since Y0
2/C33
x=/C103/C112
!ÿ 1 asx!0, we are forced by physics to choose B0
and then
y
x;tA/C7402/C33x
/C103/C114
cos
/C33tÿ:
The upper end of the chain at x/C108is fixed, requiring that
/C7402/C33
‘
/C103/C115/C32!
0:
The frequencies of the normal vibrations of the chain are given by
2/C33n
‘
/C103/C115
n;
where nare the roots of /C740. Some values of /C740
xand/C741
xare tabulated at the
end of this chapter.
Generating function for /C74n
x
The function
x;te
x=2
tÿtÿ1X1
nÿ1/C74n
xtn
7:87
is called the generating function for Bessel functions of the first kind of integral
order. It is very useful in obtaining properties of /C74n
xfor integral values of n
which can then often be proved for all values of n.
To prove Eq. (7.87), let us consider the exponential functions ext=2andeÿxt=2.
The Laurent expansions for these two exponential functions about t0 are
ext=2X1
k0
xt=2k
k/C33;eÿxt=2X1
m0
ÿxt=2k
m/C33:
Multiplying them together, we get
ex
tÿtÿ1=2X1
k0X1
m0
ÿ1m
k/C33m/C33x
2km
tkÿm:
7:88
It is easy to recognize that the coecient of the t0term which is made up of those
terms with kmis just /C740
x:
X1
k0
ÿ1k
22k
k/C332x2k/C740
x:
330SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Similarly, the coecient of the term tnwhich is made up of those terms for
which kÿmnis just /C74n
x:
X1
k0
ÿ1k
kn/C33k/C3322knx2kn/C74n
x:
This shows clearly that the coecients in the Laurent expansion (7.88) of the
generating function are just the Bessel functions of integral order. Thus we
have proved Eq. (7.87).
Bessel’s integral representation
With the help of the generating function, we can express /C74n
xin terms of a
definite integral with a parameter. To do this, let teiin the generating func-
tion, then
ex
tÿtÿ1=2ex
eiÿeÿi=2eixsin
cos
xsinisin
xcos:
Substituting this into Eq. (7.87) we obtain
cos
xsinisin
xcosX1
nÿ1/C74n
x
cosisinn
X1
ÿ1/C74n
xcosniX1
ÿ1/C74n
xsinn:
Since /C74ÿn
x
ÿ 1n/C74n
x;cosncos
ÿn, and sin nÿsin
ÿn, we have,
upon equating the real and imaginary parts of the above equation,
cos
xsin/C740
x2X1
n1/C742n
xcos 2n;
sin
xsin2X1
n1/C742nÿ1
xsin
2nÿ1:
It is interesting to note that these are the Fourier cosine and sine series of
cos
xsinand sin
xsin. Multiplying the first equation by cos kand integrat-
ing from 0 to , we obtain
1
Z
0coskcos
xsind/C74k
x;ifk0;2;4;...
0; ifk1;3;5;...(
:
331BESSEL’S EQUATION
Now multiplying the second equation by sin kand integrating from 0 to ,w e
obtain
1
Z
0sinksin
xsind/C74k
x;ifk1;3;5;...
0; ifk0;2;4;...(
:
Adding these two together we obtain Bessel’s integral representation
/C74n
x1
Z
0cos
nÿxsind;npositive integer :
7:89
Recurrence formulas for /C74n
x
Bessel functions of the first kind, /C74n
x, are the most useful, because they are
bounded near the origin. And there exist some useful recurrence formulas between
Bessel functions of di/C128erent orders and their derivatives.
1/C74n1
x2n
x/C74n
xÿ/C74nÿ1
x:
7:90
Proof: Di/C128erentiating both sides of the generating function with respect to t,w e
obtain
ex
tÿtÿ1=2x
211
t2
X1
nÿ1n/C74n
xtnÿ1
or
x
211
t2X1
nÿ1/C74n
xtnX1
nÿ1n/C74n
xtnÿ1:
This can be rewritten as
x
2X1
nÿ1/C74n
xtnx
2X1
nÿ1/C74n
xtnÿ2X1
nÿ1n/C74n
xtnÿ1
or
x
2X1
nÿ1/C74n
xtnx
2X1
nÿ1/C74n2
xtnX1
nÿ1
n1/C74n1
xtn:
Equating coecients of tnon both sides, we obtain
x
2/C74n
xx
2/C74n2
x
n1/C74n
x:
Replacing nbynÿ1, we obtain the required result.
2x/C740
n
xn/C74n
xÿx/C74n1
x:
7:91
332SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Proof:
/C74n
xX1
k0
ÿ1k
k/C33ÿ
nk12n2kxn2k:
Di/C128erentiating both sides once, we obtain
/C740
n
xX1
k0
n2k
ÿ1k
k/C33ÿ
nk12n2kxn2kÿ1;
from which we have
x/C740
n
xn/C74n
xxX1
k1
ÿ1k
kÿ1/C33ÿ
nk12n2kÿ1xn2kÿ1:
Letting km1 in the sum on the right hand side, we obtain
x/C740
n
xn/C74n
xÿxX1
m0
ÿ1m
m/C33ÿ
nm22n2m1xn2m1
n/C74n
xÿx/C74n1
x:
3x/C740
n
xÿ n/C74n
xx/C74nÿ1
x:
7:92
Proof: Di/C128erentiating both sides of the following equation with respect to x
xn/C74n
xX1
k0
ÿ1k
k/C33ÿ
nk12n2kx2n2k;
we have
d
dxfxn/C74n
xg xn/C740
n
xnxnÿ1/C74n
x;
d
dxX1
k0
ÿ1kx2n2k
2n2kk/C33ÿ
nk1X1
k0
ÿ1kx2n2kÿ1
2n2kÿ1k/C33ÿ
nk
xnX1
k0
ÿ1kx
nÿ12k
2
nÿ12kk/C33ÿ
nÿ1k1
xn/C74nÿ1
x:
Equating these two results, we have
xn/C740
n
xnxnÿ1/C74n
xxn/C74nÿ1
x:
333BESSEL’S EQUATION
Canceling out the common factor xnÿ1, we obtained the required result (7.92).
4/C740
n
x/C74nÿ1
xÿ/C74n1
x=2:
7:93
Proof: Adding (7.91) and (7.92) and dividing by 2 x, we obtain the required
result (7.93).
If we subtract (7.91) from (7.92), /C740
n
xis eliminated and we obtain
x/C74n1
xx/C74nÿ1
x2n/C74n
x
which is Eq. (7.90).
These recurrence formulas (or important identities) are very useful. Here are
some illustrative examples.
Example 7.2
Show that /C740
0
x/C74ÿ1
xÿ /C741
x.
Solution: From Eq. (7.93), we have
/C740
0
x/C74ÿ1
xÿ/C741
x=2;
then using the fact that /C74ÿn
x
ÿ 1n/C74n
x, we obtain the required results.
Example 7.3
Show that
/C743
x8
x2ÿ1
/C741
xÿ4
x/C740
x:
Solution: Letting n4 in (7.90), we have
/C743
x4
x/C742
xÿ/C741
x:
Similarly, for /C742
xwe have
/C742
x2
x/C741
xÿ/C740
x:
Substituting this into the expression for /C743
x, we obtain the required result.
Example 7.4FindR
t
0x/C740
xdx.
334SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Solution: Taking derivative of the quantity x/C741
xwith respect to x, we obtain
d
dxfx/C741
xg /C741
xx/C740
1
x:
Then using Eq. (7.92) with n1,x/C740
1
xÿ /C741
xx/C740
x, we find
d
dxfx/C741
xg /C741
xx/C740
1
xx/C740
x;
thus,
Zt
0x/C740
xdxx/C741
xjt
0t/C741
t:
/C65pproximations to the Bessel functions
For very large or very small values of xwe might be able to make some approxi-
mations to the Bessel functions of the first kind /C74n
x. By a rough argument, we
can see that the Bessel functions behave something like a damped cosine function
when the value of xis very large. To see this, let us go back to Bessel’s equation
(7.71)
x2y00xy0
x2ÿ2y0
and rewrite it as
y001
xy01ÿ2
x2/C32!
y0:
Ifxis very large, let us drop the term 2=x2and then the di/C128erential equation
reduces to
y001
xy0y0:
Letuyx1=2, then u0y0x1=21
2xÿ1=2y, and u00y00x1=2xÿ1=2y0ÿ14xÿ3=2y.
From u00we have
y001
xy0xÿ1=2u001
4x2y:
Adding yon both sides, we obtain
y001
xy0y0xÿ1=2u001
4x2yy;
xÿ1=2u001
4x2yy0
335BESSEL’S EQUATION
or
u001
4x21
x1=2yu001
4x21
u0;
the solution of which is
uAcosxBsinx:
Thus the approximate solution to Bessel’s equation for very large values of xis
yxÿ1=2
AcosxBsinxCxÿ1=2cos
x/C12:
A more rigorous argument leads to the following asymptotic formula
/C74n
x/C252
x1=2
cosxÿ
4ÿn
2
:
7:94
For very small values of x(that is, near 0), by examining the solution itself and
dropping all terms after the first, we find
/C74n
x/C25xn
2nÿ
n1:
7:95
Orthogonality of Bessel functions
Bessel functions enjoy a property which is called orthogonality and is of general
importance in mathematical physics. If and/C22are two di/C128erent constants, we
can show that under certain conditions
Z1
0x/C74n
x/C74n
/C22xdx0:
Let us see what these conditions are. First, we can show that
Z1
0x/C74n
x/C74n
/C22xdx/C22/C74n
/C740
n
/C22ÿ/C74n
/C22/C740
n
2ÿ/C222:
7:96
To show this, let us go back to Bessel’s equation (7.71) and change the indepen-dent variable to x, where is a constant, then the resulting equation is
x
2y00xy0
2x2ÿn2y0
and its general solution is /C74n
x. Now suppose we have two such equations, one
fory1with constant , and one for y2with constant /C22:
x2y00
1xy0
1
2x2ÿn2y10;x2y00
2xy0
2
/C222x2ÿn2y20:
Now multiplying the first equation by y2, the second by y1and subtracting, we get
x2y2y00
1ÿy1y00
2xy2y0
1ÿy1y0
2
/C222ÿ2x2y1y2:
336SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
Dividing by xwe obtain
xd
dxy2y0
1ÿy1y0
2y2y0
1ÿy1y0
2
/C222ÿ2xy1y2
or
d
dxfxy2y0
1ÿy1y0
2g
/C222ÿ2xy1y2
and then integration gives
/C222ÿ2Z
xy1y2dxxy2y0
1ÿy1y0
2;
where we have omitted the constant of integration. Now y1/C74n
x;y2/C74n
x,
and if 6/C22we then have
Z
x/C74n
x/C74n
/C22xdxx/C74n
/C22x/C740
n
xÿ/C22/C74n
x/C740
n
/C22x
/C222ÿ2:
Thus
Z1
0x/C74n
x/C74n
/C22xdx/C22/C74n
/C740
n
/C22ÿ/C74n
/C22/C740
n
2ÿ/C222q:e:d:
Now letting /C22!and using L’Hospital’s rule, we obtain
Z1
0x/C742
n
xdxlim
/C22!/C740
n
/C22/C740
n
ÿ/C74n
/C740
n
/C22ÿ/C22/C74n
/C7400
n
/C22
2/C22
/C740
n2
ÿ/C74n
/C740
n
ÿ/C74n
/C7400
n
2:
But
2/C7400
n
/C740
n
2ÿn2/C74n
0:
Solving for /C7400
n
and substituting, we obtain
Z1
0x/C742
n
xdx1
2/C740
n2
1ÿn2
2/C32!
/C742
n
x"#
:
7:97
Furthermore, if and /C22are any two di/C128erent roots of the equation
R/C74n
xSx/C740
n
x0, where /C82andSare constant, we then have
R/C74n
S/C740
n
0;R/C74n
/C22S/C22/C740
n
/C220;
from these two equations we find, if R60;S60,
/C22/C74n
/C740
n
/C22ÿ/C74n
/C22/C740
n
0
337BESSEL’S EQUATION
and then from Eq. (7.96) we obtain
Z1
0x/C74n
x/C74n
/C22xdx0:
7:98
Thus, the two functionsxp/C74n
xandxp/C74
n
/C22xare orthogonal in (0, 1). We can
also say that the two functions /C74n
xand/C74n
/C22xare orthogonal with respect to
the weighted function x.
Eq. (7.98) is also easily proved if R0 and S60, or R60 but S0. In this
case, and/C22can be any two di/C128erent roots of /C74n
x0o r/C740
n
x0.
/C83pherical /C66essel functions
In physics we often meet the following equation
d
drr2dR
dr
k2r2ÿ/C108
/C1081R0;
/C1080;1;2;...:
7:99
In fact, this is the radial equation of the wave and the Helmholtz partial di/C128er-
ential equation in the spherical coordinate system (see Problem 7.22). If we let
xkrandy
xR
r, then Eq. (7.99) becomes
x2y002xy0x2ÿ/C108
/C1081y0
l0;1;2;...;
7:100
where y0dy=dx. This equation almost matches Bessel’s equation (7.71). Let us
make the further substitution
y
x/C119
x=xp;
then we obtain
x2/C11900x/C1190x2ÿ
/C1081
2/C1190
/C1080;1;2;...:
7:101
The reader should recognize this equation as Bessel’s equation of order /C10812.I t
follows that the solutions of Eq. (7.100) can be written in the form
y
xA/C74/C1081=2
xxp B/C74ÿ/C108ÿ1=2
xxp :
This leads us to define spherical Bessel functions j
/C108
xC/C74/C108/C69
x=xp. The factor
/C67is usually chosen to be
=2/C112
for a reason to be explained later:
j
/C108
x
=2x/C112
/C74/C108/C69
x:
7:102
Similarly, we can define
n/C108
x =2x/C112
/C78
/C108/C69
x:
338SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
We can express j/C108
xin terms of j0
x. To do this, let us go back to /C74n
xand we
find that
d
dxfxÿn/C74n
xg ÿ xÿn/C74n1
x;or/C74n1
xÿ xnd
dxfxÿn/C74n
xg:
The proof is simple and straightforward:
d
dxfxÿn/C74n
xg d
dxX1
k0
ÿ1kx2k
2n2kk/C33ÿ
nk1
xÿnX1
k0
ÿ1kxn2kÿ1
2n2kÿ1
kÿ1/C33ÿ
nk1
xÿnX1
k0
ÿ1k1xn2k1
2n2k1k/C33ÿ
nk2gÿxÿn/C74n1
x:
Now if we set n/C1081
2and divide by x/C1083=2, we obtain
/C74/C1083=2
x
x/C1083=2ÿ1
xd
dx/C74/C1081=2
x
x/C1081=2
orj/C1081
x
x/C1081ÿ1
xd
dxj/C108
x
x/C108
:
Starting with /C1080 and applying this formula ltimes, we obtain
j/C108
xx/C108ÿ1
xd
dx/C108
j0
x
/C1081;2;3;...:
7:103
Once j0
xhas been chosen, all j/C108
xare uniquely determined by Eq. (7.103).
Now let us go back to Eq. (7.102) and see why we chose the constant factor /C67to
be
=2/C112
. If we set /C1080 in Eq. (7.101), the resulting equation is
xy002y0xy0:
Solving this equation by the power series method, the reader will find that func-
tions sin ( x=xand cos ( x=xare among the solutions. It is customary to define
j0
xsin
x=x:
Now by using Eq. (7.76), we find
/C741=2
xX1
k0
ÿ1k
x=21=22k
k/C33ÿ
k3=2
x=21=2
1=2p 1ÿx2
3/C33x4
5/C33ÿ/C32!
x=21=2
1=2psinx
x
2
x/C114
sinx:
Comparing this with j0
xshows that j0
x
=2x/C112
/C741=2
x, and this explains the
factor
=2/C112
chosen earlier.
339SPHERICAL BESSEL FUNCTIONS
/C83turm/C177Liou/C118ille s/C121stems
A boundary-value problem having the form
d
dxr
xdy
dx
/C113
x/C112
xy0; axb
7:104
and satisfying boundary conditions of the form
k1y
ak2y0
a0; /C1081y
b/C1082y0
b0
7:104a
is called a Sturm–Liouville boundary-value problem; Eq. (7.104) is known as the
Sturm–Liouville equation. Legendre’s equation, Bessel’s equation and many other
important equations can be written in the form of (7.104).
Legendre’s equation (7.1) can be written as
1ÿx2y00y0; /C23
/C231;
we can then see it is a Sturm–Liouville equation with r1ÿx2;/C1130 and /C1121.
Then, how do Bessel functions fit into the Sturm–Liouville framework/C63 /C74
s
satisfies the Bessel equation (7.71)
s2/C127/C74ns_/C74n
s2ÿn2/C74n0; _/C74nd/C74n=ds:
7:71a
We assume nis a positive integer and setting sx, with a non-zero constant,
we have
ds
dx; _/C74nd/C74n
dxdx
ds1
d/C74n
dx;/C127/C74nd
dx1
d/C74n
dxdx
ds1
2d2/C74n
dx2
and Eq. (7.71a) becomes
x2/C7400
n
xx/C740
n
x
2x2ÿn2/C74n
x0;/C740
nd/C74n=dx
or
x/C7400
n
x/C740
n
x
2xÿn2=x/C74n
x0;
which can be written as
x/C740
n
x0ÿn2
x2x/C32!
/C74n
x0:
It is easy to see that for each fixed nthis is a Sturm–Liouville equation (7.104),
with r
xx,/C113
xÿ n2=x;/C112
xx, and with the parameter now written as
2.
For the Sturm–Liouville system (7.104) and (7.104a), a non-trivial solution
exists in general only for a particular set of values of the parameter . These
values are called the eigenvalues of the system. If r
xand/C113
xare real, the
eigenvalues are real. The corresponding solutions are called eigenfunctions of
340SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
the system. In general there is one eigenfunction to each eigenvalue. This is the
non-degenerate case. In the degenerate case, more than one eigenfunction maycorrespond to the same eigenvalue. The eigenfunctions form an orthogonal set
with respect to the density function /C112
xwhich is generally 0.Thus by suitable
normalization the set of functions can be made an orthonormal set with respect to/C112
xinaxb. We now proceed to prove these two general claims.
Propert/C121 /C49Ifr
xand/C113
xare real, the eigenvalues of a Sturm–Liouville
system are real.
We start with the Sturm–Liouville equation (7.104) and the boundary condi-
tions (7.104a):
d
dxr
xdy
dx
/C113
x/C112
xy0; axb;
k1y
ak2y0
a0;/C1081y
b/C1082y0
b0;
and assume that r
x;/C113
x;/C112
x;k1;k2;/C1081,a n d /C1082are all real, but andymay be
complex. Now take the complex conjugates
d
dxr
xdy
dx
/C113
x /C112
xy0;
7:105
k1y
ak2y0
a0;/C1081y
b/C1082y0
b0;
7:105a
where yand are the complex conjugates of yand, respectively.
Multiplying (7.104) by y, (7.105) by y, and subtracting, we obtain after
simplifying
d
dxr
x
yy0ÿyy0/C2/C3
ÿ/C112
xyy:
Integrating from atob, and using the boundary conditions (7.104a) and (7.105a),
we then obtain
ÿZb
a/C112
xy0/C12/C12/C12/C122dxr
x
yy0ÿyy0jb
a0:
Since /C112
x0i n axb, the integral on the left is positive and therefore
, that is, is real.
Propert/C121 /C50
The eigenfunctions corresponding to two di/C128erent eigenvalues
are orthogonal with respect to /C112
xinaxb.
341STURM–LIOUVILLE SYSTEMS
Ify1andy2are eigenfunctions corresponding to the two di/C128erent eigenvalues
1;2, respectively,
d
dxr
xdy1
dx
/C113
x1/C112
xy10; axb;
7:106
k1y1
ak2y0
1
a0;/C1081y1
b/C1082y0
1
b0;
7:106a
d
dxr
xdy2
dx
/C113
x2/C112
xy20; axb;
7:107
k1y2
ak2y0
2
a0; /C1081y2
b/C1082y0
2
b0:
7:107a
Multiplying (7.106) by y2and (7.107) by y1, then subtracting, we obtain
d
dxr
x
y1y0
2ÿy2y0
1/C2/C3
ÿ/C112
xy1y2:
Integrating from atob, and using (7.106a) and (7.107a), we obtain
1ÿ2Zb
a/C112
xy1y2dxr
x
y1y0
2ÿy2y0
1jb
a0:
Since 162we have the required result; that is,
Zb
a/C112
xy1y2dx0:
We can normalize these eigenfunctions to make them an orthonormal set, and
so we can expand a given function in a series of these orthonormal eigenfunctions.
We have shown that Legendre’s equation is a Sturm–Liouville equation with
r
x1ÿx;/C1130 and /C1121. Since r0 when x1, no boundary conditions
are needed to form a Sturm–Liouville problem on the interval ÿ1x1. The
numbers nn
n1are eigenvalues with n0;1;2;3;.... The corresponding
eigenfunctions are ynPn
x. Property 2 tells us that
Z1
ÿ1Pn
xPm
xdx0 n6m:
For Bessel functions we saw that
x/C740
n
x0ÿn2
x2x/C32!
/C74n
x0
is a Sturm–Liouville equation (7.104), with r
xx;/C113
xÿ n2=x;/C112
xx, and
with the parameter now written as 2. Typically, we want to solve this equation
342SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
on an interval 0 xbsubject to
/C74n
b0:
which limits the selection of . Property 2 then tells us that
Zb
0x/C74n
kx/C74n
/C108xdx0; k6/C108:
Problems
7.1 Using Eq. (7.11), show that Pn
ÿx
ÿ 1nPn
xand Pn0
ÿx
ÿ1n1P0
n
x:
7.2 Find P0
x;P1
x;P2
x;P3
x, and P4
xfrom Rodrigues’ formula (7.12).
Compare your results with Eq. (7.11).
7.3 Establish the recurrence formula (7.16b) by manipulating Rodrigues’
formula.
7.4 Prove that P0
5
x9P4
x5P2
xP0
x.
Hint: Use the recurrence relation (7.16d).
7.5 Let Pand/C81be two points in space (Fig. 7.6). Using Eq. (7.14), show that
1
r1
r2
1r22ÿ2r1r2cosq
1
r2P0P1
cosr1
r2P2
cosr1
r22
"#
:
7.6 What is Pn
1/C63What is Pn
ÿ1/C63
7.7 Obtain the associated Legendre functions:
aP1
2
x;
bP23
x;
cP32
x:
7.8 Verify that P2
3
xis a solution of Legendre’s associated equation (7.25) for
m2,n3.
7.9 Verify the orthogonality conditions (7.31) for the functions P1
2
xandP13
x.
343PROBLEMS
Figure 7.6.
7.10 Verify Eq. (7.37) for the function P1
2
x:
7.11 Show that
dnÿm
dxnÿm
x2ÿ1n
nÿm/C33
nm/C33
x2ÿ1mdnm
dxnm
x2ÿ1m
Hint: Write
x2ÿ1n
xÿ1n
x1nand find the derivatives by
Leibnitz’s rule.
7.12 Use the generating function for the Hermite polynomials to find:
(a)H0
x; (b) H1
x;(c)H2
x;( d ) H3
x.
7.13 Verify that the generating function satisfies the identity
/C642
/C64x2ÿ2x/C64
/C64x2t/C64
/C64t0:
Show that the functions Hn
xin Eq. (7.47) satisfy Eq. (7.38).
7.14 Given the di/C128erential equation y00
/C34ÿx2y0, find the possible values
of/C34(eigenvalues) such that the solution y
xof the given di/C128erential equa-
tion tends to zero as x! 1 . For these values of /C34, find the eigenfunctions
y
x.
7.15 In Eq. (7.58), write the series for the exponential and collect powers of zto
verify the first few terms of the series. Verify the identity
x/C642
/C64x2
1ÿx/C64
/C64xz/C64
/C64z0:
Substituting the series (7.58) into this identity, show that the functions Ln
x
in Eq. (7.58) satisfy Laguerre’s equation.
7.16 Show that
/C740
x1ÿx2
22
1/C332x4
24
2/C332ÿx6
26
3/C332ÿ ;
/C741
xx
2ÿx3
231/C332/C33x5
252/C333/C33ÿx7
273/C334/C33ÿ :
7.17 Show that
/C741=2
x2
x1=2
sinx;/C74ÿ1=2
x2
x1=2
cosx:
7.18 If nis a positive integer, show that the formal expression for /C74ÿn
xgives
/C74ÿn
x
ÿ 1n/C74n
x.
7.19 Find the general solution to the modified Bessel’s equation
x2y00xy0
x2s2ÿ2y0
which di/C128ers from Bessel’s equation only in that sxtakes the place of x.
344SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
(Hint: Reduce the given equation to Bessel’s equation first.)
7.20 The lengthening simple pendulum: Consider a small mass msuspended by a
string of length l. If its length is increased at a steady rate ras it swings back
and forth freely in a vertical plane, find the equation of motion and the
solution for small oscillations.
7.21 Evaluate the integrals:
aZ
xn/C74nÿ1
xdx;
bZ
xÿn/C74n1
xdx;
cZ
xÿ1/C741
xdx:
7.22 In quantum mechanics, the three-dimensional Schro /C200dinger equation is
ip/C64/C32
r;t
/C64tÿp2
2m/C1142/C32
r;t/C86/C32
r;t; i
ÿ1p
;p/C104=2:
(a) When the potential /C86is independent of time, we can write /C32
r;t
u
rT
t. Show that in this case the Schro /C200dinger equation reduces to
ÿp2
2m/C1142u
r/C86u
r/C69u
r;
a time-independent equation along with T
teÿi/C69t=p, where Eis a
separation constant.
(b) Show that, in spherical coordinates, the time-independent Schro /C200dinger
equation takes the form
ÿp2
2m1
r2/C64
/C64rr2/C64u
/C64r
1
r2sin/C64
/C64sin/C64u
/C64
1
r2sin2/C642u
/C64/C30"#
/C86
ru/C69u;
then use separation of variables, u
r; ;/C30R
rY
; /C30, to split it into
two equations, with as a new separation constant:
ÿp2
2m1
r2d
drr2dR
dr
/C86
r2
R/C69R;
ÿp2
2m1
sin/C64
/C64sin/C64Y
/C64
ÿp2
2m1
sin2/C642Y
/C64/C302Y:
It is straightforward to see that the radial equation is in the form of Eq.
(7.99). Continuing the separation process by putting Y
; /C30/C2
,
the angular equation can be separated further into two equations, with/C12as separation constant:
ÿp2
2m1
d2
d/C302/C12;
ÿp2
2msind
dsind/C2
d
ÿsin2/C2/C12/C20:
345PROBLEMS
The first equation is ready for integration. Do you recognize the
second equation in as Legendre’s equation/C63 (Compare it with Eq.
(7.30).) If you are unsure, try to simplify it by putting /C132m=p;
/C22
2m/C12=p1=2, and you will obtain
sind
dsind/C2
d
/C13sin2ÿ/C222/C20
or
1
sind
dsind/C2
d
/C13ÿ/C222
sin2
/C20;
which more closely resembles Eq. (7.30).
7.23 Consider the di/C128erential equation
y00R
xy0/C81
xP
xy0:
Show that it can be put into the form of the Sturm–Liouville equation
(7.104) with
r
xeR
R
xdx;/C113
x/C81
xeR
R
xdx;and /C112
xP
xeR
R
xdx:
7.24. ( a) Show that the system y00y0;y
00;y
10 is a Sturm–
Liouville system.
(b) Find the eigenvalues and eigenfunctions of the system.
(c) Prove that the eigenfunctions are orthogonal on the interval 0 x1.
(d) Find the corresponding set of normalized eigenfunctions, and expand
the function f
x1 in a series of these orthonormal functions.
346SPECIAL FUNCTIONS OF MATHEMATICAL PHYSICS
8
The calculus of variations
The calculus of variations, in its present form, provides a powerful method for the
treatment of variational principles in physics and has become increasingly impor-
tant in the development of modern physics. It is originated as a study of certain
extremum (maximum and minimum) problems not treatable by elementary
calculus. To see this more precisely let us consider the following integral whose
integrand is a function of x,y, and of the first derivative y0
xdy=dx:
IZx2
x1fy
x;y0
x;x/C8/C9
dx;
8:1
where the semicolon in fseparates the independent variable xfrom the dependent
variable y
xand its derivative y0
x. For what function y
xis the value of the
integral Ia maximum or a minimum/C63 This is the basic problem of the calculus of
variations.
The quantity fdepends on the functional form of the dependent variable y
x
and is called the functional which is considered as given, the limits of integrationare also given. It is also understood that yy
1atxx1,yy2atxx2.I n
contrast with the simple extreme-value problem of di/C128erential calculus, the func-tiony
xis not known here, but is to be varied until an extreme value of the
integral Iis found. By this we mean that if y
xis a curve which gives to Ia
minimum value, then any neighboring curve will make Iincrease.
We can make the definition of a neighboring curve clear by giving y
xa
parametric representation:
y
/C34;xy
0;x/C34/C17
x;
8:2
where /C17
xis an arbitrary function which has a continuous first derivative and /C34is
a small arbitrary parameter. In order for the curve (8.2) to pass through
x
1;y1
and
x2;y2, we require that /C17
x1/C17
x20 (see Fig. 8.1). Now the integral I
347
also becomes a function of the parameter /C34
I
/C34Zx2
x1ffy
/C34;x;y0
/C34;x;xgdx:
8:3
We then require that y
xy
0;xmakes the integral Ian extreme, that is, the
integral I
/C34has an extreme value for /C340:
I
/C34Zx2
x1ffy
/C34;x;y0
/C34;x;xgdxextremum for /C340:
This gives us a very simple method of determining the extreme value of the
integral I. The necessary condition is
dI
d/C34/C12/C12/C12/C12
/C3400
8:4
for all functions /C17
x. The sucient conditions are quite involved and we shall
not pursue them. The interested reader is referred to mathematical texts on the
calculus of variations.
The problem of the extreme-value of an integral occurs very often in geometry
and physics. The simplest example is provided by the problem of determining the
shortest curve (or distance) between two given points. In a plane, this is the
straight line. But if the two given points lie on a given arbitrary surface, then
the analytic equation of this curve, which is called a geodesic, is found by solution
of the above extreme-value problem.
/C84he /C69uler/C177Lagrange equation
In order to find the required curve y
xwe carry out the indicated di/C128erentiation
in the extremum condition (8.4):
348THE CALCULUS OF VARIATIONS
Figure 8.1.
/C64I
/C64/C34/C64
/C64/C34Zx2
x1ffy
/C34;x;y0
/C34;x;xgdx
Zx2
x1/C64f
/C64y/C64y
/C64/C34/C64f
/C64y0/C64y0
/C64/C34
dx;
8:5
where we have employed the fact that the limits of integration are fixed, so the
di/C128erential operation a/C128ects only the integrand. From Eq. (8.2) we have
/C64y
/C64/C34/C17
xand/C64y0
/C64/C34d/C17
dx:
Substituting these into Eq. (8.5) we obtain
/C64I
/C64/C34Zx2
x1/C64f
/C64y/C17
x/C64f
/C64y0d/C17
dx
dx:
8:6
Using integration by parts, the second term on the right hand side becomes
Zx2
x1/C64f
/C64y0d/C17
dxdx/C64f
/C64y0/C17
x/C12/C12/C12/C12x2
x1ÿZx2
x1d
dx/C64f
/C64y0
/C17
xdx:
The integrated term on the right hand side vanishes because /C17
x1/C17
x20
and Eq. (8.6) becomes
/C64I
/C64/C34Zx2
x1/C64f
/C64y/C64y
/C64/C34ÿd
dx/C64f
/C64y0/C64y
/C64/C34
dx
Zx2
x1/C64f
/C64yÿd
dx/C64f
/C64y0
/C17
xdx:
8:7
Note that /C64f=/C64yand /C64f=/C64y0are still functions of /C34. However, when
/C340;y
/C34;xy
xand the dependence on /C34disappears.
Then
/C64I=/C64/C34j/C340vanishes, and since /C17
xis an arbitrary function, the inte-
grand in Eq. (8.7) must vanish for /C340:
d
dx/C64f
/C64y0ÿ/C64f
/C64y0:
8:8
Eq. (8.8) is known as the Euler–Lagrange equation; it is a necessary but not
sucient condition that the integral Ihave an extreme value. Thus, the solution
of the Euler–Lagrange equation may not yield the minimizing curve. Ordinarilywe must verify whether or not this solution yields the curve that actually mini-
mizes the integral, but frequently physical or geometrical considerations enable us
to tell whether the curve so obtained makes the integral a minimum or a max-
imum. The Euler–Lagrange equation can be written in the form (Problem 8.2)
d
dxfÿy0/C64f
/C64y0
ÿ/C64f
/C64x0:
8:8a
349THE EULER–LAGRANGE EQUATION
This is often called the second form of the Euler–Lagrange equation. If fdoes not
involve xexplicitly, it can be integrated to yield
fÿy0/C64f
/C64y0c;
8:8b
where cis an integration constant.
The Euler–Lagrange equation can be extended to the case in which fis a
functional of several dependent variables:
ffy 1
x;y0
1
x;y2
x;y0
2
x;...;x/C8/C9
:
Then, in analogy with Eq. (8.2), we now have
yi
/C34;xyi
0;x/C34/C17i
x; i1;2;...;n:
The development proceeds in an exactly analogous manner, with the result
/C64I
/C64/C34Zx2
x1/C64f
/C64yiÿd
dx/C64f
/C64yi0
/C17i
xdx:
Since the individual variations, that is, the /C17i
x, are all independent, the vanish-
ing of the above equation when evaluated at /C340 requires the separate vanishing
of each expression in the brackets:
d
dx/C64f
/C64y0
iÿ/C64f
/C64yi0; i1;2;...;n:
8:9
Example 8.1
The brachistochrone problem: Historically, the brachistochrone problem was the
first to be treated by the method of the calculus of variations (first solved by
Johann Bernoulli in 1696). As shown in Fig. 8.2, a particle is constrained to
move in a gravitational field starting at rest from some point P1to some lower
350THE CALCULUS OF VARIATIONS
Figure 8.2
point P2. Find the shape of the path such that the particle goes from P1toP2in
the least time. (The word brachistochrone was derived from the Greek brachistos
(shortest) and chronos (time).)
Solution: If O andPare not very far apart, the gravitational field is constant, and
if we ignore the possibility of friction, then the total energy of the particle is
conserved:
0m/C103y 11
2mds
dt2
m/C103
y1ÿy;
where the left hand side is the sum of the kinetic energy and the potential energy
of the particle at point P1, and the right hand side refers to point P
x;y. Solving
fords=dt:
ds=dt
2/C103y/C112
:
Thus the time required for the particle to move from P1toP2is
tZP2
P1dtZP2
P1ds2/C103yp :
The line element dscan be expressed as
ds
dx2dy2q
1y
02q
dx; y0dy=dx;
thus, we have
tZP2
P1dtZP2
P1ds2/C103yp 12/C103pZ
x2
0
1y02/C112
yp dx:
We now apply the Euler–Lagrange equation to find the shape of the path for
the particle to go from P1toP2in the least time. The constant does not a/C128ect the
final equation and the functional fmay be identified as
f
1y02q
=yp;
which does not involve xexplicitly. Using Problem 8.2(b), we find
fÿy0/C64f
/C64y0
1y02/C112
yp ÿy0 y0
1y02/C112 yp"#
c;
which simplifies to
1y02qyp1=c:
351THE EULER–LAGRANGE EQUATION
Letting 1 =capand solving for y0gives
y0dy
dxaÿy
y/C114
;
and solving for dxand integrating we obtain
Z
dxZy
aÿy/C114
dy:
We then let
yasin2a
2
1ÿcos 2
which leads to
x2aZ
sin2daZ
1ÿcos 2 da2
2ÿsin 2 k:
Thus the parametric equation of the path is given by
xb
1ÿcos/C30; yb
/C30ÿsin/C30k;
where ba=2;/C302. The path passes through the origin so we have k0 and
xb
1ÿcos/C30; yb
/C30ÿsin/C30:
The constant bis determined from the condition that the particle passes through
P
2
x2;y2:
The required path is a cycloid and is the path of a fixed point P0on a circle of
radius bas it rolls along the x-axis (Fig. 8.3).
A line that represents the shortest path between any two points on some surface
is called a geodesic. On a flat surface, the geodesic is a straight line. It is easy to
show that, on a sphere, the geodesic is a great circle; we leave this as an exercise
for the reader (Problem 8.3).
352THE CALCULUS OF VARIATIONS
Figure 8.3.
/C86ariational problems /C119ith constraints
In certain problems we seek a minimum or maximum value of the integral (8.1)
IZx2
x1fy
x;y0
x;x/C8/C9
dx
8:1
subject to the condition that another integral
/C74Zx2
x1/C103y
x;y0
x;x/C8/C9
dx
8:10
has a known constant value. A simple problem of this sort is the problem of
determining the curve of a given perimeter which encloses the largest area, or
finding the shape of a chain of fixed length which minimizes the potential energy.
In this case we can use the method of Lagrange multipliers which is based on
the following theorem:
The problem of the stationary value of /C70(x/C44 y) subject to the con/C45
dition /C71
x;yc/C111nst .is e/C113uivalent to the problem of stationary
values/C44 without constraint/C44 of /C70/C71for some constant /C44 pro/C45
vided either /C64/C71=/C64xor/C64/C71=/C64ydoes not vanish at the critical point.
The constant is called a Lagrange multiplier and the method is known as the
method of Lagrange multipliers. To see the ideas behind this theorem, let us
assume that /C71
x;y0 defines yas a unique function of x, say, y/C103
x, having
a continuous derivative /C1030
x. Then
/C70
x;y/C70x;/C103
x
and its maximum or minimum can be found by setting the derivative with respect
toxequal to zero:
/C64/C70
/C64x/C64/C70
/C64ydy
dx0o r /C70x/C70y/C1030
x0:
8:11
We also have
/C71x;/C103
x 0;
from which we find
/C64/C71
/C64x/C64/C71
/C64ydy
dx0o r /C71x/C71y/C1030
x0:
8:12
Eliminating /C1030
xbetween Eq. (8.11) and Eq. (8.12) we obtain
/C70xÿ/C70y=/C71yÿ
/C71x0;
8:13
353VARIATIONAL PROBLEMS WITH CONSTRAINTS
provided /C71y/C64/C71=/C64y60. Defining ÿ/C70y=/C71yor
/C70y/C71y/C64/C70
/C64y/C64/C71
/C64y0;
8:14
Eq. (8.13) becomes
/C70x/C71x/C64/C70
/C64x/C64/C71
/C64x0:
8:15
If we define
H
x;y/C70
x;y/C71
x;y;
then Eqs. (8.14) and (8.15) become
/C64H
x;y=/C64x0; H
x;y=/C64y0;
and this is the basic idea behind the method of Lagrange multipliers.
It is natural to attempt to solve the problem Iminimum subject to the con-
dition /C74constant by the method of Lagrange multipliers. We construct the
integral
I/C74Zx2
x1/C70
y;y0;x/C71
y;y0;xdx
and consider its free extremum. This implies that the function y
xthat makes the
value of the integral an extremum must satisfy the equation
d
dx/C64
/C70/C71
/C64y0/C64
/C70/C71
/C64y0
8:16
or
d
dx/C64/C70
/C64y0
ÿ/C64/C70
/C64y
d
dx/C64/C71
/C64y0
ÿ/C64/C71
/C64y
0:
8:16a
Example 8.2
Isoperimetric problem: Find that curve /C67having the given perimeter lthat
encloses the largest area.
Solution: The area bounded by /C67can be expressed as
A1
2Z
C
xdyÿydx12Z
C
xy0ÿydx
and the length of the curve /C67is
sZ
C
1y02q
dx/C108:
354THE CALCULUS OF VARIATIONS
Then the function /C72is
HZ
C1
2
xy0ÿy
1y02/C112
dx
and the Euler–Lagrange equation gives
d
dx1
2xy0
1y02/C112/C32!
1
20
or
y0
1y02/C112 ÿxc1:
Solving for y0, we get
y0dy
dxxÿc1
2ÿ
xÿc12q ;
which on integrating gives
yÿc2
2ÿ
xÿc12q
or
xÿc12
yÿc222;a circle :
/C72amilton/C39s principle and Lagrange/C39s equation of motion
One of the most important applications of the calculus of variations is in classicalmechanics. In this case, the functional fin Eq. (8.1) is taken to be the Lagrangian
Lof a dynamical system. For a conservative system, the Lagrangian Lis defined
as the di/C128erence of kinetic and potential energies of the system:
LTÿ/C86;
where time tis the independent variable and the generalized coordinates /C113
i
tare
the dependent variables. What do we mean by generalized coordinates/C63 Any
convenient set of parameters or quantities that can be used to specify the config-
uration (or state) of the system can be assumed to be generalized coordinates;
therefore they need not be geometrical quantities, such as distances or angles. In
suitable circumstances, for example, they could be electric currents.
Eq. (8.1) now takes the form that is known as the action (or the action integral)
IZt2
t1L/C113 i
t;_/C113i
t;t
dt; _/C113d/C113=dt
8:17
355HAMILTON’S PRINCIPLE
and Eq. (8.4) becomes
I/C64I
/C64/C34/C12/C12/C12/C12
/C340d/C34Zt2
t1L/C113 i
t; _/C113i
t;t
dt0;
8:18
where /C113i
t, and hence _/C113i
t, is to be varied subject to /C113i
t1/C113i
t20.
Equation (8.18) is a mathematical statement of Hamilton’s principle of classical
mechanics. In this variational approach to mechanics, the Lagrangian Lis given,
and/C113i
ttaken on the prescribed values at t1andt2, but may be arbitrarily varied
for values of tbetween t1andt2.
In words, Hamilton’s principle states that for a conservative dynamical system,
the motion of the system from its position in configuration space at time t1to its
position at time t2follows a path for which the action integral (8.17) has a
stationary value. The resulting Euler–Lagrange equations are known as theLagrange equations of motion:
d
dt/C64L
/C64_/C113iÿ/C64L
/C64/C113i0:
8:19
These Lagrange equations can be derived from Newton’s equations of motion
(that is, the second law written in di/C128erential equation form) and Newton’s equa-
tions can be derived from Lagrange’s equations. Thus they are ‘equivalent.’
However, Hamilton’s principle can be applied to a wide range of physical phe-
nomena, particularly those involving fields, with which Newton’s equations are
not usually associated. Therefore, Hamilton’s principle is considered to be more
fundamental than Newton’s equations and is often introduced as a basic postulate
from which various formulations of classical dynamics are derived.
Example 8.3
Electric oscillations: As an illustration of the generality of Lagrangian dynamics,
we consider its application to an L/C67circuit (inductive–capacitive circuit) as shown
in Fig. 8.4. At some instant of time the charge on the capacitor /C67is/C81
tand the
current flowing through the inductor is I
t _/C81
t. The voltage drop around the
356THE CALCULUS OF VARIATIONS
Figure 8.4. L/C67circuit.
circuit is, according to Kirchho/C128 ’s law
LdI
dt1
CZ
I
tdt0
or in terms of /C81
L/C127/C811
C/C810:
This equation is of exactly the same form as that for a simple mechanical oscil-
lator:
m/C127xkx0:
If the electric circuit also contains a resistor /C82, Kirchho/C128 ’s law then gives
L/C127/C81R_/C811
C/C810;
which is of exactly the same form as that for a damped oscillator
m/C127xb_xkx0;
where bis the damping constant.
By comparing the corresponding terms in these equations, an analogy between
mechanical and electric quantities can be established:
x displacement /C81 charge (generalized coordinate)
_x velocity _/C81Ielectric current
m mass L inductance
1=kk spring constant /C67 capacitance
b damping constant /C82 electric resistance
1
2m_x2kinetic energy12L_/C812energy stored in inductance
1
2mx2potential energy12/C812=Cenergy stored in capacitance
If we recognize in the beginning that the charge /C81in the circuit plays the role of a
generalized coordinate, and T1
2L_/C812and/C8612/C812=C, then the Langrangian L
of the system is
LTÿ/C861
2L_/C812ÿ12/C812=C
and the Lagrange equation gives
L/C127/C811
C/C810;
the same equation as given by Kirchho/C128 ’s law.
357HAMILTON’S PRINCIPLE
Example 8.4
A bead of mass mslides freely on a frictionless wire of radius bthat rotates in a
horizontal plane about a point on the circular wire with a constant angular
velocity /C33. Show that the bead oscillates as a pendulum of length /C108/C103=/C332.
Solution: The circular wire rotates in the xyplane about the point O, as shown
in Fig. 8.5. The rotation is in the counterclockwise direction, /C67is the center of the
circular wire, and the angles and /C30are as indicated. The wire rotates with an
angular velocity /C33,s o/C30/C33t. Now the coordinates xandyof the bead are given
by
xbcos/C33tbcos
/C33t;
ybsin/C33tbsin
/C33t;
and the generalized coordinate is . The potential energy of the bead (in a hor-
izontal plane) can be taken to be zero, while its kinetic energy is
T1
2m
_x2_y212mb2/C332
_/C3322/C33
_/C33cos;
which is also the Lagrangian of the bead. Inserting this into Lagrange’s equation
d
d/C64L
/C64_
ÿ/C64L
/C640
we obtain, after some simplifications,
/C127/C332sin0:
Comparing this equation with Lagrange’s equation for a simple pendulum of
length l
/C127
/C103=/C108sin0
358THE CALCULUS OF VARIATIONS
Figure 8.5.
(Fig. 8.6) we see that the bead oscillates about the line OAlike a pendulum of
length /C108/C103=/C332.
/C82a/C121leigh/C177/C82it/C122 method
Hamilton’s principle views the motion of a dynamical system as a whole and
involves a search for the path in configuration space that yields a stationary
value for the action integral (8.17):
IZt2
t1L/C113 i
t;_/C113i
t;t
dt0;
8:18
with /C113i
t1/C113i
t20. Ordinarily it is used as a variational method to obtain
Lagrange’s and Hamilton’s equations of motion, so we do not often think of it as
a computational tool. But in other areas of physics variational formulationsare used in a much more active way. For example, the variational method for
determining the approximate ground-state energies in quantum mechanics is
very well known. We now use the Rayleigh–Ritz method to illustrate that
Hamilton’s principle can be used as computational device in classical
mechanics. The Rayleigh–Ritz method is a procedure for obtaining approximate
solutions of problems expressed in variational form directly from the variational
equation.
The Lagrangian is a function of the generalized coordinates /C113s and their time
derivatives _/C113s. The basic idea of the approximation method is to guess a solution
for the /C113s that depends on time and a number of parameters. The parameters are
then adjusted so that Hamilton’s principle is satisfied. The Rayleigh–Ritz method
takes a special form for the trial solution. A complete set of functions ff
i
tgis
chosen and the solution is assumed to be a linear combination of a finite number
of these functions. The coecients in this linear combination are the parameters
that are chosen to satisfy Hamilton’s principle (8.18). Since the variations of the /C113s
359RAYLEIGH–RIT/C90 METHOD
Figure 8.6.
must vanish at the endpoints of the integral, the variations of the parameter must
be so chosen that this condition is satisfied.
To summarize, suppose a given system can be described by the action integral
IZt2
t1L/C113 i
t;_/C113i
t;t
dt; _/C113d/C113=dt:
The Rayleigh–Ritz method requires the selection of a trial solution, ideally in theform
/C113X
n
i1aifi
t;
8:20
which satisfies the appropriate conditions at both the initial and final times, and
where as are undetermined constant coecients and the fs are arbitrarily chosen
functions. This trial solution is substituted into the action integral Iand integra-
tion is performed so that we obtain an expression for the integral Iin terms of the
coecients. The integral Iis then made ‘stationary’ with respect to the assumed
solution by requiring that
/C64I
/C64ai0
8:21
after which the resulting set of nsimultaneous equations is solved for the values of
the coecients ai. To illustrate this method, we apply it to two simple examples.
Example 8.5A simple harmonic oscillator consists of a mass Mattached to a spring of force
constant k. As a trial function we take the displacement xas a function tin the
form
x
tX
1
n1Ansinn/C33t:
For the boundary conditions we have x0;t0, and x0;t2=/C33. Then the
potential energy and the kinetic energy are given by, respectively,
/C861
2kx212kP1
n1P1
m1AnAmsinn/C33tsinm/C33t;
T1
2M_x212M/C332P1
n1P1
m1AnAmnmcosn/C33tcosm/C33t:
The action Ihas the form
IZ2=/C33
0LdtZ2=/C33
0
Tÿ/C86dt
2/C33X1
n1
kA2
nÿMn2A2n/C332:
360THE CALCULUS OF VARIATIONS
In order to satisfy Hamilton’s principle we must choose the values of Anso as to
make Ian extremum:
dI
dAn
kÿn2/C332MAn0:
The solution that meets the physics of the problem is
A10;/C332k=M; or /C282=/C33
1=22M=k
1=2;
An0;for n2;3;etc:
Example 8.6
As a second example, we consider a bead of mass Msliding freely along a wire
shaped in the form of a parabola along the vertical axis and of the form yax2.
In this case, we have
LTÿ/C861
2M
_x2_y2ÿM/C103y 12M
14a2x2_x2ÿM/C103y :
We assume
xAsin/C33t
to be an approximate value for the displacement x, and then the action integral
becomes
IZ2=/C33
0LdtZ2=/C33
0
Tÿ/C86dtA2/C332
1a2A2
2ÿ/C103a()
M
/C33:
The extremum condition, dI=dA0, gives an approximate /C33:
/C332/C103ap
1a2A2;
and the approximate period is
/C282
1a2A22/C103ap :
The Rayleigh–Ritz method discussed in this section is a special case of the
general Rayleigh–Ritz methods that are designed for finding approximate solu-
tions of boundary-value problems by use of varitional principles, for example, the
eigenvalues and eigenfunctions of the Sturm–Liouville systems.
/C72amilton/C39s principle and canonical equations of motion
Newton first formulated classical mechanics in the seventeenth century and it is
known as Newtonian mechanics. The essential physics involved in Newtonian
361HAMILTON’S PRINCIPLE
mechanics is contained in Newton’s three laws of motion, with the second law
serving as the equation of motion. Classical mechanics has since been reformu-lated in a few di/C128erent forms: the Lagrange, the Hamilton, and the Hamilton–
Jacobi formalisms, to name just a few.
The essential physics of Lagrangian dynamics is contained in the Lagrange
function Lof the dynamical system and Lagrange’s equations (the equations of
motion). The Lagrangian Lis defined in terms of independent generalized coor-
dinates /C113
iand the corresponding generalized velocity _/C113i. In Hamiltonian
dynamics, we describe the state of a system by Hamilton’s function (or theHamiltonian) /C72defined in terms of the generalized coordinates /C113
iand the corre-
sponding generalized momenta /C112i, and the equations of motion are given by
Hamilton’s equations or canonical equations
_/C113i/C64H
/C64/C112i; _/C112iÿ/C64H
/C64/C113i; i1;2;...;n:
8:22
Hamilton’s equations of motion can be derived from Hamilton’s principle.
Before doing so, we have to define the generalized momentum and the
Hamiltonian. The generalized momentum /C112icorresponding to /C113iis defined as
/C112i/C64L
/C64/C113i
8:23
and the Hamiltonian of the system is defined by
HX
i/C112i_/C113iÿL:
8:24
Even though _/C113iexplicitly appears in the defining expression (8.24), /C72is a function
of the generalized coordinates /C113i, the generalized momenta /C112i, and the time t,
because the defining expression (8.23) can be solved explicitly for the _/C113isi n
terms of /C112i;/C113i, and t. The /C113sa n d ps are now treated the same: HH
/C113i;/C112i;t.
Just as with the configuration space spanned by the nindependent /C113s, we can
imagine a space of 2 ndimensions spanned by the 2 nvariables
/C1131;/C1132;...;/C113n;/C1121;/C1122;...;/C112n. Such a space is called phase space, and is particularly
useful in both statistical mechanics and the study of non-linear oscillations. Theevolution of a representative point in this space is determined by Hamilton’sequations.
We are ready to deduce Hamilton’s equation from Hamilton’s principle. The
original Hamilton’s principle refers to paths in configuration space, so in order to
extend the principle to phase space, we must modify it such that the integrand of
the action Iis a function of both the generalized coordinates and momenta and
their derivatives. The action Ican then be evaluated over the paths of the system
362THE CALCULUS OF VARIATIONS
point in phase space. To do this, first we solve Eq. (8.24) for L
LX
i/C112i_/C113iÿH
and then substitute Linto Eq. (8.18) and we obtain
IZt2
t1X
i/C112i_/C113iÿH
/C112;/C113;t
dt0;
8:25
where /C113I
tis still varied subject to /C113i
t1/C113i
t20, but /C112iis varied without
such end-point restrictions.
Carrying out the variation, we obtain
Zt2
t1X
i/C112i_/C113i_/C113i/C112iÿ/C64H
/C64/C113i/C113iÿ/C64H
/C64/C112i/C112i
dt0;
8:26
where the _/C113s are related to the /C113s by the relation
_/C113id
dt/C113i:
8:27
Now we integrate the term /C112i_/C113idtby parts. Using Eq. (8.27) and the endpoint
conditions on /C113i, we find that
Zt2
t1X
i/C112i_/C113idtZt2
t1X
i/C112id
dt/C113idt
Zt2
t1X
id
dt/C112i/C113idtÿZt2
t1X
i_/C112i/C113idt
/C112i/C113i/C12/C12/C12/C12t2
t1ÿZt2
t1X
i_/C112i/C113idt
ÿZt2
t1X
i_/C112i/C113idt:
Substituting this back into Eq. (8.26), we obtain
Zt2
t1X
i_/C113iÿ/C64H
/C64/C112i
/C112iÿ _/C112i/C64H
/C64/C113i
/C113i
dt0:
8:28
Since we view Hamilton’s principle as a variational principle in phase space, both
the/C113s and the /C112s are arbitrary, the coecients of /C113iand /C112iin Eq. (8.28) must
vanish separately, which results in the 2 nHamilton’s equations (8.22).
Example 8.7
Obtain Hamilton’s equations of motion for a one-dimensional harmonic oscilla-
tor.
363HAMILTON’S PRINCIPLE
Solution: We have
T1
2m_x2; /C8612/C75x2;
/C112/C64L
/C64_x/C64T
/C64_xm_x; _x/C112
m:
Hence
H/C112_xÿLT/C861
2m/C11221
2/C75x2:
Hamilton’s equations
_x/C64H
/C64/C112; _/C112ÿ/C64H
/C64x
then read
_x/C112
m; _/C112ÿ/C75x:
Using the first equation, the second can be written
d
dt
m_xÿ /C75x orm/C127x/C75x0
which is the familiar equation of the harmonic oscillator.
/C84he modified /C72amilton/C39s principle and the /C72amilton/C177/C74acobi equation
The Hamilton–Jacobi equation is the cornerstone of a general method of integrat-
ing equations of motion. Before the advent of modern quantum theory, Bohr’s
atomic theory was treated in terms of Hamilton–Jacobi theory. It also plays an
important role in optics as well as in canonical perturbation theory. In classical
mechanics books, the Hamilton–Jacobi equation is often obtained via canonical
transformations. We want to show that the Hamilton–Jacobi equation can also beobtained directly from Hamilton’s principle, or, a modified Hamilton’s principle.
In formulating Hamilton’s principle, we have considered the action
IZ
t2
t1L/C113 i
t;_/C113i
t;t
dt; _/C113d/C113=dt;
taken along a path between two given positions /C113i
t1and/C113i
t2which the dyna-
mical system occupies at given instants t1andt2. In varying the action, we com-
pare the values of the action for neighboring paths with fixed ends, that is, with
/C113i
t1/C113i
t20. Only one of these paths corresponds to the true dynamical
path for which the action has its extremum value.
We now consider another aspect of the concept of action, by regarding Ias a
quantity characterizing the motion along the true path, and comparing the value
364THE CALCULUS OF VARIATIONS
ofIfor paths having a common beginning at /C113i
t1, but passing through di/C128erent
points at time t2. In other words we consider the action Ifor the true path as a
function of the coordinates at the upper limit of integration:
II
/C113i;t;
where /C113iare the coordinates of the final position of the system, and tis the instant
when this position is reached.
If/C113i
t2are the coordinates of the final position of the system reached at time t2,
the coordinates of a point near the point /C113i
t2can be written as /C113i
t1/C113i,
where /C113iis a small quantity. The action for the trajectory bringing the system
to the point /C113i
t1/C113idi/C128ers from the action for the trajectory bringing the
system to the point /C113i
t2by the quantity
IZt2
t1/C64L
/C64/C113i/C113i/C64L
/C64_/C113i_/C113i
dt;
8:29
where /C113iis the di/C128erence between the values of /C113itaken for both paths at the same
instant t; similarly, _/C113iis the di/C128erence between the values of _/C113iat the instant t.
We now integrate the second term on the right hand side of Eq. (8.25) by parts:
Zt2
t1/C64L
/C64_/C113i_/C113idt/C64L
/C64_/C113i/C113iÿZt2
t1d
dt/C64L
/C64_/C113i
/C113idt
/C112i/C113iÿZt2
t1d
dt/C64L
/C64_/C113i
/C113idt;
8:30
where we have used the fact that the starting points of both paths coincide, hence
/C113i
t10; the quantity /C113i
t2is now written as just /C113i. Substituting Eq. (8.30)
into Eq. (8.29), we obtain
IX
i/C112i/C113iZt2
t1X
i/C64L
/C64/C113iÿd
dt/C64L
/C64_/C113i
/C113idt:
8:31
Since the true path satisfies Lagrange’s equations of motion, the integrand and,consequently, the integral itself vanish. We have thus obtained the following value
for the increment of the action Idue to the change in the coordinates of the final
position of the system by /C113
i(at a constant time of motion):
IX
i/C112i/C113i;
8:32
from which it follows that
/C64I
/C64/C113i/C112i;
8:33
that is, the partial derivatives of the action with respect to the generalized co-ordinates equal the corresponding generalized momenta.
365THE MODIFIED HAMILTON’S PRINCIPLE
The action Imay similarly be regarded as an explicit function of time, by
considering paths starting from a given point /C113i
1at a given instant t1, ending
at a given point /C113i
2at various times t2t:
II
/C113i;t:
Then the total time derivative of Iis
dI
dt/C64I
/C64tX
i/C64I
/C64/C113i_/C113i/C64I
/C64tX
i/C112i_/C113i:
8:34
From the definition of the action, we have dI=dtL. Substituting this into Eq.
(8.34), we obtain
/C64I
/C64tLÿX
i/C112i_/C113iÿH
or
/C64I
/C64tH
/C113i;/C112i;t0:
8:35
Replacing the momenta /C112iin the Hamiltonian Hby/C64I=/C64/C113ias given by Eq. (8.33),
we obtain the Hamilton–Jacobi equation
H
/C113i;/C64I=/C64/C113i;t/C64I
/C64t0:
8:36
For a conservative system with stationary constraints, the time is not contained
explicitly in Hamiltonian /C72,a n d H/C69(the total energy of the system).
Consequently, according to Eq. (8.35), the dependence of action Ion time tis
expressed by the term ÿ/C69t. Therefore, the action breaks up into two terms, one of
which depends only on /C113i, and the other only on t:
I
/C113i;tI/C111
/C113iÿ/C69t:
8:37
The function I/C111
/C113iis sometimes called the contracted action, and the Hamilton–
Jacobi equation (8.36) reduces to
H
/C113i;/C64I/C111=/C64/C113i/C69:
8:38
Example 8.8
To illustrate the method of Hamilton–Jacobi, let us consider the motion of an
electron of charge ÿerevolving about an atomic nucleus of charge Ze(Fig. 8.7).
As the mass Mof the nucleus is much greater than the mass mof the electron, we
may consider the nucleus to remain stationary without making any very appreci-
able error. This is a central force motion and so its motion lies entirely in one
plane (see /C67lassical Mechanics , by Tai L. Chow, John Wiley, 1995). Employing
366THE CALCULUS OF VARIATIONS
polar coordinates rand in the plane of motion to specify the position of the
electron relative to the nucleus, the kinetic and potential energies are, respectively,
T1
2m
_r2r2_2; /C86ÿZe2
r:
Then
LTÿ/C8612m
_r
2r2_2Ze2
r
and
/C112r/C64L
/C64_rm_r/C112 /C112/C64L
/C64_mr2_:
The Hamiltonian /C72is
H1
2m/C1122
r/C1122
r2/C32!
ÿZe2
r:
Replacing /C112rand/C112in the Hamiltonian by /C64I=/C64rand /C64I=/C64, respectively, we
obtain, by Eq. (8.36), the Hamilton–Jacobi equation
1
2m/C64I
/C64r2
1
r2/C64I
/C642"#
ÿZe2
r/C64I
/C64t0:
/C86ariational problems /C119ith se/C118eral independent /C118ariables
The functional fin Eq. (8.1) contains only one independent variable, but very
often fmay contain several independent variables. Let us now extend the theory
to this case of several independent variables:
IZZZ
/C86ffu;ux;uy;uz;x;y;zdxdydz ;
8:39
where /C86is assumed to be a bounded volume in space with prescribed values of
u
x;y;zat its boundary S;ux/C64u=/C64x, and so on. Now, the variational problem
367VARIATIONAL PROBLEMS
Figure 8.7.
is to find the function u
x;y;zfor which Iis stationary with respect to small
changes in the functional form u
x;y;z.
Generalizing Eq. (8.2), we now let
u
x;y;z;/C34u
x;y;z;0/C34/C17
x;y;z;
8:40
where /C17
x;y;zis an arbitrary well-behaved (that is, di/C128erentiable) function which
vanishes at the boundary S. Then we have, from Eq. (8.40),
ux
x;y;z;/C34ux
x;y;z;0/C34/C17x;
and similar expressions for uy;uz;a n d
/C64I
/C64/C34/C12/C12/C12/C12
/C340ZZZ
/C86/C64f
/C64u/C17/C64f
/C64ux/C17x/C64f
/C64uy/C17y/C64f
/C64uz/C17z
dxdydz 0:
We next integrate each of the terms
/C64f=/C64ui/C17iusing ‘integration by parts’ and the
integrated terms vanish at the boundary as required. After some simplifications,
we finally obtain
ZZZ
/C86/C64f
/C64uÿ/C64
/C64x/C64f
/C64uxÿ/C64
/C64y/C64f
/C64uyÿ/C64
/C64z/C64f
/C64uz/C26/C27
/C17
x;y;zdxdydz 0:
Again, since /C17
x;y;zis arbitrary, the term in the braces may be set equal to zero,
and we obtain the Euler–Lagrange equation:
/C64f
/C64uÿ/C64
/C64x/C64f
/C64uxÿ/C64
/C64y/C64f
/C64uyÿ/C64
/C64z/C64f
/C64uz0:
8:41
Note that in Eq. (8.41) /C64=/C64xis a partial derivative, in that yandzare constant.
But/C64=/C64xis also a total derivative in that it acts on implicit xdependence and on
explicit xdependence:
/C64
/C64x/C64f
/C64ux/C642f
/C64x/C64ux/C642f
/C64u/C64uxux/C642f
/C64u2x/C642f
/C64uy/C64uxuxy/C642f
/C64uz/C64uxuxz:
8:42
Example 8.9
The Schro /C200dinger wave equation. The equations of motion of classical mechanics
are the Euler–Lagrange di/C128erential equations of Hamilton’s principle. Similarly,the Schro /C200dinger equation, the basic equation of quantum mechanics, is also a
Euler–Lagrange di/C128erential equation of a variational principle the form of which
is, in the case of a system of /C78particles, the following
Z
Ld/C280;
8:43
368THE CALCULUS OF VARIATIONS
with
LX/C78
i1p2
2mi/C64/C32/C42
/C64xi/C64/C32
/C64xi/C64/C32/C42
/C64yi/C64/C32
/C64yi/C64/C32/C42
/C64zi/C64/C32
/C64zi
/C86/C32/C42/C32
8:44
and the constraint
Z
/C32/C42/C32d/C281;
8:45
where miis the mass of particle I,/C86is the potential energy of the system, and d/C28is
a volume element of the 3 /C78-dimensional space.
Condition (8.45) can be taken into consideration by introducing a Lagrangian
multiplier ÿ/C69:
Z
Lÿ/C69/C32/C42/C32d/C280:
8:46
Performing the variation we obtain the Schro /C200dinger equation for a system of /C78
particles
X/C78
i1p2
2mi/C1142
i/C32
/C69ÿ/C86/C320;
8:47
where /C1142
iis the Laplace operator relating to particle i. Can you see that Eis the
energy parameter of the system/C63 If we use the Hamiltonian operator ^H, Eq. (8.47)
can be written as
^H/C32/C69/C32:
8:48
From this we obtain for E
/C69Z
/C32/C42H/C32d/C28
Z
/C32/C42/C32d/C28:
8:49
Through partial integration we obtain
Z
Ld/C28Z
/C32/C42H/C32d/C28
and thus the variational principle can be formulated in another way:
R
/C32/C42
Hÿ/C69/C32d/C280.
Problems
8.1 As a simple practice of using varied paths and the extremum condition, we
consider the simple function y
xxand the neighboring paths
369PROBLEMS
y
/C34;xx/C34sinx. Draw these paths in the xyplane between the limits
x0a n d x2for/C340 for two di/C128erent non-vanishing values of /C34. If the
integral I
/C34is given by
I
/C34Z2
0
dy=dx2dx;
show that the value of I
/C34is always greater than I
0, no matter what value
of/C34(positive or negative) is chosen. This is just condition (8.4).
8.2 ( a) Show that the Euler–Lagrange equation can be written in the form
d
dxfÿy0/C64f
/C64y0
ÿ/C64f
/C64x0:
This is often called the second form of the Euler–Lagrange equation.
(b)I ffdoes not involve xexplicitly, show that the Euler–Lagrange equation
can be integrated to yield
fÿy0/C64f
/C64y0c;
where cis an integration constant.
8.3 As shown in Fig. 8.8, a curve /C67joining points
x1;y1and
x2;y2is
revolved about the x-axis. Find the shape of the curve such that the surface
thus generated is a minimum.
8.4 A geodesic is a line that represents the shortest distance between two points.
Find the geodesic on the surface of a sphere.
8.5 Show that the geodesic on the surface of a right circular cylinder is a helix.
8.6 Find the shape of a heavy chain which minimizes the potential energy while
the length of the chain is constant.
8.7 A wedge of mass Mand angle slides freely on a horizontal plane. A
particle of mass mmoves freely on the wedge. Determine the motion of
the particle as well as that of the wedge (Fig. 8.9).
370THE CALCULUS OF VARIATIONS
Figure 8.8.
8.8 Use the Rayleigh–Ritz method to analyze the forced oscillations of a har-
monic oscillation:
m/C127xkx/C700sin/C33t:
8.9 A particle of mass mis attracted to a fixed point Oby an inverse square
force /C70rÿk=r2(Fig. 8.10). Find the canonical equations of motion.
8.10 Set up the Hamilton–Jacobi equation for the simple harmonic oscillator.
371PROBLEMS
Figure 8.9.
Figure 8.10.
9
The Laplace transformation
The Laplace transformation method is generally useful for obtaining solutions of
linear di/C128erential equations (both ordinary and partial). It enables us to reduce a
di/C128erential equation to an algebraic equation, thus avoiding going to the trouble
of finding the general solution and then evaluating the arbitrary constants. This
procedure or technique can be extended to systems of equations and to integral
equations, and it often yields results more readily than other techniques. In this
chapter we shall first define the Laplace transformation, then evaluate the trans-
formation for some elementary functions, and finally apply it to solve some simple
physical problems.
/C68efinition of the Lapace transform
The Laplace transform Lf
xof a function f
xis defined by the integral
Lf
x Z1
0eÿ/C112xf
xdx/C70
/C112;
9:1
whenever this integral exists. The integral in Eq. (9.1) is a function of the para-
meter pand we denote it by /C70
/C112. The function /C70
/C112is called the Laplace trans-
form of f
x. We may also look upon Eq. (9.1) as a definition of a Laplace
transform operator Lwhich tranforms f
xin to /C70
/C112. The operator Lis linear,
since from Eq. (9.1) we have
Lc1f
xc2/C103
x Z1
0eÿ/C112xfc1f
xc2/C103
xgdx
c1Z1
0eÿ/C112xf
xdxc2Z1
0eÿ/C112x/C103
xdx
c1Lf
x c2L/C103
x;
372
where c1andc2are arbitrary constants and /C103
xis an arbitrary function defined
forx/C620.
The inverse Laplace transform of /C70
/C112is a function f
xsuch that
Lf
x /C70
/C112. We denote the operation of taking an inverse Laplace transform
byLÿ1:
Lÿ1/C70
/C112 f
x:
9:2
That is, we operate algebraically with the operators LandLÿ1, bringing them
from one side of an equation to the other side just as we would in writing axb
implies xaÿ1b. To illustrate the calculation of a Laplace transform, let us
consider the following simple example.
Example 9.1
Find Leax, where ais a constant.
Solution: The transform is
LeaxZ1
0eÿ/C112xeaxdxZ1
0eÿ
/C112ÿaxdx:
For/C112a, the exponent on eis positive or zero and the integral diverges. For
/C112/C62a, the integral converges:
LeaxZ1
0eÿ/C112xeaxdxZ1
0eÿ
/C112ÿaxdxeÿ
/C112ÿax
ÿ
/C112ÿa/C12/C12/C12/C121
01
/C112ÿa:
This example enables us to investigate the existence of Eq. (9.1) for a general
function f
x.
/C69/C120istence of Laplace transforms
We can prove that:
(1) if f
xis piecewise continuous on every finite interval 0 xX, and
(2) if we can find constants Mand asuch that jf
xj MeaxforxX,
then Lf
xexists for /C112/C62a. A function f
xwhich satisfies condition (2) is said
to be of exponential order asx!1 ; this is mathematician’s jargon/C33
These are sucient conditions on f
xunder which we can guarantee the
existence of Lf
x. Under these conditions the integral converges for /C112/C62a:
ZX
0f
xeÿ/C112xdx/C12/C12/C12/C12/C12/C12/C12/C12Z
X
0f
xjj eÿ/C112xdxZX
0Meaxeÿ/C112xdx
MZ1
0eÿ
/C112ÿaxdxM
/C112ÿa:
373E/C88ISTENCE OF LAPLACE TRANSFORMS
This establishes not only the convergence but the absolute convergence of the
integral defining Lf
x. Note that M=
/C112ÿatends to zero as /C112!1 . This
shows that
lim
/C112!1/C70
/C1120
9:3
for all functions /C70
/C112Lf
xsuch that f
xsatisfies the foregoing conditions
(1) and (2). It follows that if lim /C112!1/C70
/C1126 0,/C70
/C112cannot be the Laplace trans-
form of any function f
x.
It is obvious that functions of exponential order play a dominant role in the use
of Laplace transforms. One simple way of determining whether or not a specifiedfunction is of exponential order is the following one: if a constant bexists such
that
lim
x!1eÿbxf
xjj/C104/C105
9:4
exists, the function f
xis of exponential order (of the order of eÿbx. To see this,
let the value of the above limit be /C7560. Then, when xis large enough, jeÿbxf
xj
can be made as close to /C75as possible, so certainly
jeÿbxf
xj<2/C75:
Thus, for suciently large x,
jf
xj<2/C75ebx
or
jf
xj<Mebx;with M2/C75:
On the other hand, if
lim
x!1eÿcxf
xjj 1
9:5
for every fixed c, the function f
xis not of exponential order. To see this, let us
assume that bexists such that
jf
xj<Mebxfor xX
from which it follows that
jeÿ2bxf
xj<Meÿbx:
Then the choice of c2bwould give us jeÿcxf
xj<Meÿbx, and eÿcxf
x!0a s
x!1 which contradicts Eq. (9.5).
Example 9.2
Show that x3is of exponential order as x!1 .
374THE LAPLACE TRANSFORMATION
Solution: We have to check whether or not
lim
x!1eÿbxx3
lim
x!1x3
ebx
exists. Now if b/C620, then L’Hospital’s rule gives
lim
x!1eÿbxx3
lim
x!1x3
ebxlim
x!13x2
bebxlim
x!16x
b2ebxlim
x!16
b3ebx0:
Therefore x3is of exponential order as x!1 .
Laplace transforms of some elementar/C121 functions
Using the definition (9.1) we now obtain the transforms of polynomials, expo-
nential and trigonometric functions.
(1)f
x1 for x/C620.
By definition, we have
L1Z1
0eÿ/C112xdx1
/C112; /C112/C620:
(2)f
xxn, where nis a positive integer.
By definition, we have
LxnZ1
0eÿ/C112xxndx:
Using integration by parts:
Z
u/C1180dxu/C118ÿZ
/C118u0dx
with
uxn;d/C118/C1180dxeÿ/C112xdxÿ
1=/C112d
eÿ/C112x;/C118ÿ
1=/C112eÿ/C112x;
we obtain
Z1
0eÿ/C112xxndxÿxneÿ/C112x
/C112 1
0n/C112Z
1
0eÿ/C112xxnÿ1dx:
For/C112/C620 and n/C620, the first term on the right hand side of the above equation is
zero, and so we have
Z1
0eÿ/C112xxndxn/C112Z
1
0eÿ/C112xxnÿ1dx
375LAPLACE TRANSFORMS OF ELEMENTARY FUNCTIONS
or
Lxnn
/C112Lxnÿ1
from which we may obtain for n/C621
Lxnÿ1nÿ1
/C112Lxnÿ2:
Iteration of this process yields
Lxnn
nÿ1
nÿ2 21
/C112nLx0:
By (1) above we have
Lx0L11=/C112:
Hence we finally have
Lxnn/C33
/C112n1; /C112/C620:
(3)f
xeax, where ais a real constant.
LeaxZ1
0eÿ/C112xeaxdx1
/C112ÿa;
where /C112/C62afor convegence. (For details, see Example 9.1.)
(4)f
xsinax, where ais a real constant.
LsinaxZ1
0eÿ/C112xsinaxdx :
UsingZ
u/C1180dxu/C118ÿZ
/C118u0dx with ueÿ/C112x;d/C118ÿd
cosax=a;
and
Z
emxsinnxdx emx
msinnxÿncosnx
n2m2
(you can obtain this simply by using integration by parts twice) we obtain
LsinaxZ1
0eÿ/C112xsinaxdx eÿ/C112x
ÿ/C112sinaxÿacosax
/C1122a2 1
0:
Since pis positive, eÿ/C112x!0a s x!1 , but sin axand cos axare bounded as
x!1 , so we obtain
Lsinax0ÿ1
0ÿa
/C1122a2a
/C1122a2;/C112/C620:
376THE LAPLACE TRANSFORMATION
(5)f
xcosax, where ais a real constant.
Using the result
Z
emxcosnxdx emx
mcosnxnsinmx
n2m2;
we obtain
LcosaxZ1
0eÿ/C112xcosaxdx /C112
/C1122a2; /C112/C620:
(6)f
xsinhax, where ais a real constant.
Using the linearity property of the Laplace transform operator L, we obtain
Lcosh axLeaxeÿax
2
1
2Leax12Le
ÿax
12 1
/C112ÿa1
/C112a
/C112
/C1122ÿa2:
(7)f
xxk, where k/C62ÿ1.
By definition we have
LxkZ1
0eÿ/C112xxkdx:
Let/C112xu, then dx/C112ÿ1du;xkuk=/C112k, and so
LxkZ1
0eÿ/C112xxkdx1
/C112k1Z1
0ukeÿuduÿ
k1
/C112k1:
Note that the integral defining the gamma function converges if and only if
k/C62ÿ1.
The following example illustrates the calculation of inverse Laplace transforms
which is equally important in solving di/C128erential equations.
Example 9.3
Find
aLÿ15
/C1122
;
bLÿ11
/C112s
;s/C620:
Solution:
aLÿ15
/C1122
5Lÿ11
/C1122
:
377LAPLACE TRANSFORMS OF ELEMENTARY FUNCTIONS
Recall Leax1=
/C112ÿa, hence Lÿ11=
/C112ÿa eax. It follows that
Lÿ15
/C1122
5Lÿ11
/C1122
5eÿ2x:
(b) Recall
LxkZ1
0eÿ/C112xxkdx1
/C112k1Z1
0ukeÿuduÿ
k1
/C112k1:
From this we have
Lxk
ÿ
k1"#
1
/C112k1;
hence
Lÿ11
/C112k1
xk
ÿ
k1:
If we now let k1s, then
Lÿ11
/C112s
xsÿ1
ÿ
s:
/C83hifting (or translation) theorems
In practical applications, we often meet functions multiplied by exponential fac-
tors. If we know the Laplace transform of a function, then multiplying it by an
exponential factor does not require a new computation as shown by the following
theorem.
/C84he /C174rst shifting theorem
IfLf
x/C70
/C112;/C112/C62b;t/C104en L eatf
x /C70
/C112ÿa;/C112/C62ab.
Note that /C70
/C112ÿadenotes the function /C70
/C112‘shifted’ a units to the right.
Hence the theorem is called the shifting theorem.
The proof is simple and straightforward. By definition (9.1) we have
Lf
x Z1
0eÿ/C112xf
xdx/C70
/C112:
Then
Leaxf
x Z1
0eÿ/C112xfeaxf
xgdxZ1
0eÿ
/C112ÿaxf
xdx/C70
/C112ÿa:
The following examples illustrate the use of this theorem.
378THE LAPLACE TRANSFORMATION
Example 9.4
Show that:
aLeÿaxxn n/C33
/C112an1; /C112/C62ÿa;
bLeÿaxsinbx b
/C112a2b2; /C112/C62ÿa:
Solution: (a) Recall
Lxnn/C33=/C112n1; /C112/C620;
the shifting theorem then gives
Leÿaxxnn/C33
/C112an1; /C112/C62ÿa:
(b) Since
Lsinaxa
/C1122a2;
it follows from the shifting theorem that
Leÿaxsinbxb
/C112a2b2; /C112/C62ÿa:
Because of the relationship between Laplace transforms and inverse Laplace
transforms, any theorem involving Laplace transforms will have a corresponding
theorem involving inverse Lapace transforms. Thus
If Lÿ1/C70
/C112 f
x;t/C104en Lÿ1/C70
/C112ÿa eaxf
x:
/C84he second shifting theorem
This second shifting theorem involves the shifting xvariable and states that
/C71iven Lf
x /C70
/C112,where f
x0forx<0;and if /C103
xf
xÿa,
then
L/C103
x eÿa/C112Lf
x:
To prove this theorem, let us start with
/C70
/C112Lf
x Z1
0eÿ/C112xf
xdx
from which it follows that
eÿa/C112/C70
/C112eÿa/C112Lf
x Z1
0eÿ/C112
xaf
xdx:
379SHIFTING (OR TRANSLATION) THEOREMS
Letuxa, then
eÿa/C112/C70
/C112Z1
0eÿ/C112
xaf
xdxZ1
0eÿ/C112uf
uÿadu
Za
0eÿ/C112u0duZ1
aeÿ/C112uf
uÿadu
Z1
0eÿ/C112u/C103
uduL/C103
u :
Example 9.5
Show that given
f
xxfor x0
0f o r x<0;/C26
and if
/C103
x0; for x<5
xÿ5;for x5/C26
then
L/C103
x eÿ5/C112=/C1122:
Solution: We first notice that
/C103
xf
xÿ5:
Then the second shifting theorem gives
L/C103
x eÿ5/C112Lxeÿ5/C112=/C1122:
/C84he unit step function
It is often possible to express various discontinuous functions in terms of the unitstep function, which is defined as
U
xÿa0x<a
1xa:/C26
Sometimes it is convenient to state the second shifting theorem in terms of the
unit step function:
Iff
x0forx<0and Lf
x /C70
/C112,then
LU
xÿaf
xÿa e
ÿa/C112/C70
/C112:
380THE LAPLACE TRANSFORMATION
The proof is straightforward:
LU
xÿaf
xÿa Z1
0eÿ/C112xU
xÿaf
xÿdx
Za
0eÿ/C112x0dxZ1
aeÿ/C112xf
xÿadx:
Letxÿau, then
LU
xÿaf
xÿa Z1
aeÿ/C112xf
xÿadx
Z1
aeÿ/C112
uaf
udueÿa/C112Z1
aeÿ/C112uf
udueÿa/C112/C70
/C112:
The corresponding theorem involving inverse Laplace transforms can be stated as
Iff
x0forx<0and Lÿ1/C70
/C112 f
x/C44 then
Lÿ1eÿa/C112/C70
/C112 U
xÿaf
xÿa:
Laplace transform of a periodic function
Iff
xis a periodic function of period P/C620, that is, if f
xPf
x, then
Lf
x 1
1ÿeÿ/C112PZP
0eÿ/C112xf
xdx:
To prove this, we assume that the Laplace transform of f
xexists:
Lf
x Z1
0eÿ/C112xf
xdxZP
0eÿ/C112xf
xdxZ2P
Peÿ/C112xf
xdx
Z3P
2Peÿ/C112xf
xdx :
On the right hand side, let xuPin the second integral, xu2Pin the
third integral, and so on, we then have
Lf
x ZP
0eÿ/C112xf
xdxZP
0eÿ/C112
uPf
uPdu
ZP
0eÿ/C112
u2Pf
u2Pdu :
381LAPLACE TRANSFORM OF A PERIODIC FUNCTION
But f
uPf
u;f
u2Pf
u;etc:Also, let us replace the dummy
variable ubyx, then the above equation becomes
Lf
x ZP
0eÿ/C112xf
xdxZP
0eÿ/C112
xPf
xdxZP
0eÿ/C112
x2Pf
xdx
ZP
0eÿ/C112xf
xdxeÿ/C112PZP
0eÿ/C112xf
xdxeÿ2/C112PZP
0eÿ/C112xf
xdx
1eÿ/C112Peÿ2/C112P ZP
0eÿ/C112xf
xdx
1
1ÿeÿ/C112PZP
0eÿ/C112xf
xdx:
Laplace transforms of deri/C118ati/C118es
Iff
xis a continuous for x0, and f0
xis piecewise continuous in every finite
interval 0 xk, and if jf
xj Mebx(that is, f
xis of exponential order),
then
Lf0
x /C112L f
x ÿ f
0;/C112/C62b:
We may employ integration by parts to prove this result:Z
ud/C118u/C118ÿZ
/C118du with ueÿ/C112x;and d/C118f0
xdx;
Lf0
x Z1
0eÿ/C112xf0
xdxeÿ/C112xf
x1
0ÿZ1
0
ÿ/C112eÿ/C112xf
xdx:
Since jf
xj Mebxfor suciently large x, then jf
xeÿ/C112xjMe
bÿ/C112for su-
ciently large x.I f /C112/C62b, then Me
bÿ/C112!0a s x!1 ; and eÿ/C112xf
x!0a s
x!1 . Next, f
xis continuous at x0, and so eÿ/C112xf
x!f
0asx!0.
Thus, the desired result follows:
Lf0
x /C112Lf
x ÿf
0;/C112/C62b:
This result can be extended as follows:
Iff
xis such that f
nÿ1
xis continuous and f
n
xpiecewise continuous in
every interval 0 xkand furthermore, if f
x;f0
x;...;f
n
xare of
exponential order for 0 /C62k, then
Lf
n
x /C112nLf
x ÿ/C112nÿ1f
0ÿ/C112nÿ2f0
0ÿÿ f
nÿ1
0:
Example 9.6
Solve the initial value problem:
y00y0;y
0y0
00;and f
t0f o r t<0 but f
t1f o r t0:
382THE LAPLACE TRANSFORMATION
Solution: Note that y0dy=dt. We know how to solve this simple di/C128erential
equation, but as an illustration we now solve it using Laplace transforms. Taking
both sides of the equation we obtain
Ly00LyL1;
LfL1:
Now
Ly00/C112Ly0ÿy0
0/C112f/C112Lyÿy
0g ÿ y0
0
/C1122Lyÿ/C112y
0ÿy0
0
/C1122Ly
and
L11=/C112:
The transformed equation then becomes
/C1122Ly Ly 1=/C112
or
Ly 1
/C112
/C112211
/C112ÿ/C112
/C11221;
therefore
yLÿ11
/C112
ÿLÿ1 /C112
/C11221
:
We find from Eqs. (9.6) and (9.10) that
Lÿ11
/C112
1 and Lÿ1/C112
/C11221
cost:
Thus, the solution of the initial problem is
y1ÿcostfor t0; y0 for t<0:
Laplace transforms of functions defined b/C121 integrals
If/C103
xRx
0f
udu,and if Lf
x /C70
/C112,thenL/C103
x /C70
/C112=/C112.
Similarly, if Lÿ1/C70
/C112 f
x, then Lÿ1/C70
/C112=/C112 /C103
x:
It is easy to prove this. If /C103
xRx
0f
udu, then /C103
00;/C1030
xf
x. Taking
Laplace transform, we obtain
L/C1030
x Lf
x
383FUNCTIONS DEFINED BY INTEGRALS
but
L/C1030
x /C112L/C103
x ÿ/C103
0/C112L/C103
x
and so
/C112L/C103
x Lf
x;orL/C103
x 1
/C112Lf
x /C70
/C112
/C112:
From this we have
Lÿ1/C70
/C112=/C112 /C103
x:
Example 9.7
If/C103
xRu
0sinau du , then
L/C103
x LZu
0sinau du
1
/C112Lsinau a
/C112
/C1122a2:
/C65 note on integral transformations
A Laplace transform is one of the integral transformations. The integral trans-
formation Tf
xof a function f
xis defined by the integral equation
Tf
x Zb
af
x/C75
/C112;xdx/C70
/C112;
9:6
where /C75
/C112;x, a known function of pand x, is called the kernel of the transfor-
mation. In the application of integral transformations to the solution of bound-
ary-value problems, we have so far made use of five di/C128erent kernels:
Laplace transform: /C75
/C112;xeÿ/C112x,a n d a0;b1 :
Lf
x Z1
0eÿ/C112xf
xdx
/C112:
Fourier sine and cosine transforms: /C75
/C112;xsin/C112xor cos px, and
a0;b1 :
/C70f
x Z1
0f
x/C26sin
/C112x
cos
/C112xdx/C70
/C112:
Complex Fourier transform: /C75
/C112;xei/C112x;andaÿ 1 ,bÿ 1 :
/C70f
x Z1
ÿ1ei/C112xf
xdx/C70
/C112:
384THE LAPLACE TRANSFORMATION
Hankel transform: /C75
/C112;xx/C74n
/C112x;a0;b1 , where /C74n
/C112xis the
Bessel function of the first kind of order n:
Hf
x Z1
0f
xx/C74n
xdx/C70
/C112:
Mellin transform: /C75
/C112;xx/C112ÿ1,a n d a0;b1 :
Mf
x Z1
0f
xx/C112ÿ1dx/C70
/C112:
The Laplace transform has been the subject of this chapter, and the Fouier
transform was treated in Chapter 4. It is beyond the scope of this book to include
Hankel and Mellin transformations.
Problems
9.1 Show that:
(a)et2is not of exponential order as x!1 .
(b) sin et2is of exponential order as x!1 .
9.2 Show that:
(a)Lsinhaxa
/C1122ÿa2; /C112/C620:
(b)L3x4ÿ2x3=2672
/C1125ÿ3p
2/C1125=26
/C112.
(c)Lsinxcosx1=
/C11224:
(d)I f
f
xx;0<x<4
5;x/C624;/C26
then
Lf
x1
/C1122eÿ4/C112
/C112ÿeÿ4/C112
/C1122:
9.4 Show that LU
xÿa eÿa/C112=/C112;/C112/C620:
9.5 Find the Laplace transform of H
x, where
H
xx;
5;0<x<4
x/C624:/C26
9.5 Let f
xbe the rectified sine wave of period P2:
f
xsinx;0<x<
0; x<2:/C26
Find the Laplace transform of f
x.
385PROBLEMS
9.6 Find
Lÿ1 15
/C11224/C11213
:
9.7 Prove that if f0
xis continuous and f00
xis piecewise continuous in every
finite interval 0 xkand if f
xandf0
xare of exponential order for
x/C62k, then
LfF
x /C1122Lf
x ÿ/C112f
0ÿf0
0:
(Hint: Use (9.19) with f0
xin place of f
xandf00
xin place of f0
x.)
9.8 Solve the initial problem y00
t/C122y
tAsin/C33t;y
01;y0
00.
9.9 Solve the initial problem yF
tÿy0
tsintsubject to
y
02; y
00; y00
01:
9.10 Solve the linear simultaneous di/C128erential equation with constant coecients
y002yÿx0;
x002xÿy0;
subject to x
02;y
00, and x0
0y0
00, where xand yare the
dependent variables and tis the independent variable.
9.11 Find
LZ1
0cosau du
:
9.12. Prove that if Lf
x /C70
/C112then
Lf
ax 1
a/C70/C112a
:
Similarly if L
ÿ1/C70
/C112 f
xthen
Lÿ1/C70/C112a/C104/C105
af
ax:
386THE LAPLACE TRANSFORMATION
10
Partial di/C128erential e/C113uations
We have met some partial di/C128erential equations in previous chapters. In this
chapter we will study some elementary methods of solving partial di/C128erential
equations which occur frequently in physics and in engineering. In general, the
solution of partial di/C128erential equations presents a much more dicult problem
than the solution of ordinary di/C128erential equations. A complete discussion of the
general theory of partial di/C128erential equations is well beyond the scope of this
book. We therefore limit ourselves to a few solvable partial di/C128erential equations
that are of physical interest.
Any equation that contains an unknown function of two or more variables and
its partial derivatives with respect to these variables is called a partial di/C128erentialequation, the order of the equation being equal to the order of the highest partial
derivatives present. For example, the equations
3y
2/C64u
/C64x/C64u
/C64y2u;/C642u
/C64x/C64y2xÿy
are typical partial di/C128erential equations of the first and second orders, respec-
tively, xandybeing independent variables and u
x;ythe function to be found.
These two equations are linear, because both uand its derivatives occur only to
the first order and products of uand its derivatives are absent. We shall not
consider non-linear partial di/C128erential equations.
We have seen that the general solution of an ordinary di/C128erential equation
contains arbitrary constants equal in number to the order of the equation. But
the general solution of a partial di/C128erential equation contains arbitrary functions
(equal in number to the order of the equation). After the particular choice of the
arbitrary functions is made, the general solution becomes a particular solution.
The problem of finding the solution of a given di/C128erential equation subject to
given initial conditions is called a boundary-value problem or an initial-value
387
problem. We have seen already that such problems often lead to eigenvalue
problems.
Linear second-order partial di/C128erential equations
Many physical processes can be described to some degree of accuracy by linear
second-order partial di/C128erential equations. For simplicity, we shall restrict our
discussion to the second-order linear partial di/C128erential equation in two indepen-
dent variables, which has the general form
A/C642u
/C64x2B/C642u
/C64x/C64yC/C642u
/C64y2D/C64u
/C64x/C69/C64u
/C64y/C70u/C71;
10:1
where A;B;C;...;/C71may be dependent on variables xandy.
If/C71is a zero function, then Eq. (10.1) is called homogeneous; otherwise it is
said to be non-homogeneous. If u1;u2;...;unare solutions of a linear homoge-
neous partial di/C128erential equation, then c1u1c2u2 cnunis also a solution,
where c1;c2;...are constants. This is known as the superposition principle; it does
not apply to non-linear equations. The general solution of a linear non-homo-geneous partial di/C128erential equation is obtained by adding a particular solution
of the non-homogeneous equation to the general solution of the homogeneous
equation.
The homogeneous form of Eq. (10.1) resembles the equation of a general conic:
ax
2bxycy2dxeyf0:
We thus say that Eq. (10.1) is of
elliptic
hyperbolic
parabolic9
>=
>;type whenB2ÿ4AC<0
B2ÿ4AC/C620
B2ÿ4AC08
><
>::
For example, according to this classification the two-dimensional Laplace equation
/C642u
/C64x2/C642u
/C64y20
is of elliptic type ( AC1;BD/C69/C70/C710, and the equation
/C642u
/C64x2ÿ2/C642u
/C64y20
is a real constant
is of hyperbolic type. Similarly, the equation
/C642u
/C64x2ÿ/C64u
/C64y0
is a real constant
is of parabolic type.
388PARTIAL DIFFERENTIAL EQUATIONS
We now list some important linear second-order partial di/C128erential equations
that are of physical interest and we have seen already:
(1) Laplace’s equation:
/C1142u0;
10:2
where /C1142is the Laplacian operator. The function umay be the electrostatic
potential in a charge-free region. It may be the gravitational potential in a region
containing no matter or the velocity potential for an incompressible fluid with no
sources or sinks.
(2) Poisson’s equation:
/C1142u/C26
x;y;z;
10:3
where the function /C26
x;y;zis called the source density. For example, if urepre-
sents the electrostatic potential in a region containing charges, then /C26is propor-
tional to the electrical charge density. Similarly, for the gravitational potential
case, /C26is proportional to the mass density in the region.
(3) Wave equation:
/C1142u1
/C1182/C642u
/C64t2;
10:4
transverse vibrations of a string, longitudinal vibrations of a beam, or propaga-tion of an electromagnetic wave all obey this same type of equation. For a vibrat-
ing string, urepresents the displacement from equilibrium of the string; for a
vibrating beam, uis the longitudinal displacement from the equilibrium.
Similarly, for an electromagnetic wave, umay be a component of electric field
/C69or magnetic field /C72.
(4) Heat conduction equation:
/C64u
/C64t/C1142u;
10:5
where uis the temperature in a solid at time t. The constant is called the
di/C128usivity and is related to the thermal conductivity, the specific heat capacity,
and the mass density of the object. Eq. (10.5) can also be used as a di/C128usion
equation: uis then the concentration of a di/C128using substance.
It is obvious that Eqs. (10.2)–(10.5) all are homogeneous linear equations with
constant coecients.
Example 10.1
Laplace’s equation: arises in almost all branches of analysis. A simple example
can be found from the motion of an incompressible fluid. Its velocity /C118
x;y;z;t
and the fluid density /C26
x;y;z;tmust satisfy the equation of continuity:
/C64/C26
/C64t/C114
/C26/C1180:
389LINEAR SECOND-ORDER PDEs
If/C26is constant we then have
/C114/C1180:
If, furthermore, the motion is irrotational, the velocity vector can be expressed as
the gradient of a scalar function /C86:
/C118ÿ /C114 /C86;
and the equation of continuity becomes Laplace’s equation:
/C114/C118 /C114
ÿ/C114 /C860;or/C1142/C860:
The scalar function /C86is called the velocity potential.
Example 10.2
Poisson’s equation: The electrostatic field provides a good example of Poisson’s
equation. The electric force between any two charges /C113and/C1130in a homogeneous
isotropic medium is given by Coulomb’s law
FC/C113/C1130
r2^r;
where ris the distance between the charges, and ^ris a unit vector in the direction
of the force. The constant /C67determines the system of units, which is not of
interest to us; thus we leave /C67as it is.
An electric field /C69is said to exist in a region if a stationary charge /C1130in that
region experiences a force F:
/C69lim
/C1130!0
F=/C1130:
The lim /C1130!0guarantees that the test charge /C1130will not alter the charge distribution
that existed prior to the introduction of the test charge /C1130. From this definition
and Coulomb’s law we find that the electric field at a point rdistant from a point
charge is given by
/C69C/C113
r2^r:
Taking the curl on both sides we get
/C114 /C690;
which shows that the electrostatic field is a conservative field. Hence a potential
function /C30exists such that
/C69ÿ /C114 /C30:
Taking the divergence of both sides
/C114
/C114 /C30 ÿ/C114 /C69
390PARTIAL DIFFERENTIAL EQUATIONS
or
/C1142/C30ÿ /C114 /C69:
/C114/C69is given by Gauss’ law. To see this, consider a volume /C28containing a total
charge /C113. Let dsbe an element of the surface Swhich bounds the volume /C28. Then
ZZ
S/C69dsC/C113ZZ
S^rds
r2:
The quantity ^rdsis the projection of the element area dson a plane perpendi-
cular to r. This projected area divided by r2is the solid angle subtended by ds,
which is written d/C10. Thus, we have
ZZ
S/C69dsC/C113ZZ
S^rds
r2C/C113ZZ
Sd/C104C/C113:
If we write /C113as
/C113ZZZ
/C28/C26d/C86;
where /C26is the charge density, then
ZZ
S/C69ds4CZZZ
/C28/C26d/C86:
But (by the divergence theorem)
ZZ
S/C69dsZZZ
/C28/C114/C69d/C86:
Substituting this into the previous equation, we obtain
ZZZ
/C28/C114/C69d/C864CZZZ
/C28/C26d/C86
or
ZZZ
/C28/C114/C69ÿ4C/C26
d/C860:
This equation must be valid for all volumes, that is, for any choice of the volume
/C28. Thus, we have Gauss’ law in di/C128erential form:
/C114/C694C/C26:
Substituting this into the equation /C1142/C30ÿ /C114 /C69, we get
/C1142/C30ÿ4C/C26;
which is Poisson’s equation. In the Gaussian system of units, C1; in the SI
system of units, C1=4/C340, where the constant /C340is known as the permittivity of
free space. If we use SI units, then
/C1142/C30ÿ/C26=/C34 0:
391LINEAR SECOND-ORDER PDEs
In the particular case of zero charge density it reduces to Laplace’s equation,
/C1142/C300:
In the following sections, we shall consider a number of problems to illustrate
some useful methods of solving linear partial di/C128erential equations. There are
many methods by which homogeneous linear equations with constant coecients
can be solved. The following are commonly used in the applications.
(1) General solutions: In this method we first find the general solution and then
that particular solution which satisfies the boundary conditions. It is always
satisfying from the point of view of a mathematician to be able to find general
solutions of partial di/C128erential equations; however, general solutions are dicult
to find and such solutions are sometimes of little value when given boundary
conditions are to be imposed on the solution. To overcome this diculty it is
best to find a less general type of solution which is satisfied by the type of
boundary conditions to be imposed. This is the method of separation of variables.
(2) Separation of variables: The method of separation of variables makes use of
the principle of superposition in building up a linear combination of individualsolutions to form a solution satisfying the boundary conditions. The basic
approach of this method in attempting to solve a di/C128erential equation (in, say,
two dependent variables xandy) is to write the dependent variable u
x;yas a
product of functions of the separate variables u
x;yX
xY
y. In many cases
the partial di/C128erential equation reduces to ordinary di/C128erential equations for X
and/C89.
(3) Laplace transform method: We first obtain the Laplace transform of the
partial di/C128erential equation and the associated boundary conditions with respect
to one of the independent variables, and then solve the resulting equation for the
Laplace transform of the required solution which can be found by taking the
inverse Laplace transform.
/C83olutions of Laplace/C39s equation/C58 separation of /C118ariables
(1) Laplace’s equation in two dimensions
x;y: If the potential /C30is a function of
only two rectangular coordinates, Laplace’s equation reads
/C642/C30
/C64x2/C642/C30
/C64y20:
It is possible to obtain the general solution to this equation by means of a trans-
formation to a new set of independent variables:
/C24xiy;/C17xÿiy;
392PARTIAL DIFFERENTIAL EQUATIONS
where Iis the unit imaginary number. In terms of these we have
/C64
/C64x/C64
/C64/C24/C64/C24
/C64x/C64
/C64/C17/C64/C17
/C64x/C64
/C64/C24/C64
/C64/C17;
/C642
/C64x2/C64
/C64x/C64
/C64/C24/C64
/C64/C17
/C64
/C64/C24/C64
/C64/C24/C64
/C64/C17/C64/C24
/C64x/C64
/C64/C17/C64
/C64/C24/C64
/C64/C17/C64/C17
/C64x
/C642
/C64/C2422/C64
/C64/C24/C64
/C64/C17/C642
/C64/C172:
Similarly, we have
/C642
/C64y2ÿ/C642
/C64/C2422/C64
/C64/C24/C64
/C64/C17ÿ/C642
/C64/C172
and Laplace’s equation now reads
/C1142/C304/C642/C30
/C64/C24/C64/C170:
Clearly, a very general solution to this equation is
/C30f1
/C24f2
/C17f1
xiyf2
xÿiy;
where f1andf2are arbitrary functions which are twice di/C128erentiable. However, it
is a somewhat dicult matter to choose the functions f1andf2such that the
equation is, for example, satisfied inside a square region defined by the lines
x0;xa;y0;yband such that /C30takes prescribed values on the boundary
of this region. For many problems the method of separation of variables is moresatisfactory. Let us apply this method to Laplace’s equation in three dimensions.
(2) Laplace’s equation in three dimensions ( x;y;z): Now we have
/C642/C30
/C64x2/C642/C30
/C64y2/C642/C30
/C64z20:
10:6
We make the assumption, justifiable by its success, that /C30
x;y;zmay be written
as the product
/C30
x;y;zX
xY
yZ
z:
Substitution of this into Eq. (10.6) yields, after division by /C30;
1
Xd2X
dx21
Yd2Y
dy2ÿ1
Zd2Z
dz2:
10:7
393SOLUTIONS OF LAPLACE’S EQUATION
The left hand side of Eq. (10.7) is a function of xandy, while the right hand side is
a function of zalone. If Eq. (10.7) is to have a solution at all, each side of the
equation must be equal to the same constant, say k2
3. Then Eq. (10.7) leads to
d2Z
dz2k23Z0;
10:8
1
Xd2X
dx2ÿ1
Yd2Y
dy2k2
3:
10:9
The left hand side of Eq. (10.9) is a function of xonly, while the right hand side is
a function of yonly. Thus, each side of the equation must be equal to a constant,
sayk2
1. Therefore
d2X
dx2k21X0;
10:10
d2Y
dy2k2
2Y0;
10:11
where
k22k21ÿk23:
The solution of Eq. (10.10) is of the form
X
xa
k1ek1x;k160;ÿ1 <k1<1
or
X
xa
k1ek1xa0
k1eÿk1x;k160;0<k1<1:
10:12
Similarly, the solutions of Eqs. (10.11) and (10.8) are of the forms
Y
yb
k2ek2yb0
k2eÿk2y;k260;0<k2<1;
10:13
Z
zc
k3ek3zc0
k3eÿk3z;k360;0<k3<1:
10:14
Hence
/C30a
k1ek1xa0
k1eÿk1xb
k2ek2yb0
k2eÿk2yc
k3ek3zc0
k3eÿk3z;
and the general solution of Eq. (10.6) is obtained by integrating the above equa-
tion over all the permissible values of the ki
i1;2;3.
In the special case when ki0
i1;2;3, Eqs. (10.8), (10.10), and (10.11)
have solutions of the form
Xi
xiaixibi;
where x1x, and X1Xetc.
394PARTIAL DIFFERENTIAL EQUATIONS
Let us now apply the above result to a simple problem in electrostatics: that of
finding the potential /C30at a point Pa distance hfrom a uniformly charged infinite
plane in a dielectric of permittivity /C34. Let be the charge per unit area of the
plane, and take the origin of the coordinates in the plane and the x-axis perpen-
dicular to the plane. It is evident that /C30is a function of xonly. There are two types
of solutions, namely:
/C30
xa
k1ek1xa0
k1eÿk1x;
/C30
xa1xb1;
the boundary conditions will eliminate the unwanted one. The first boundary
condition is that the plane is an equipotential, that is, /C30
0constant, and the
second condition is that /C69ÿ/C64/C30=/C64 x=2/C34. Clearly, only the second type of
solution satisfies both the boundary conditions. Hence b1/C30
0;a1ÿ=2/C34,
and the solution is
/C30
xÿ
2/C34x/C30
0:
(3) Laplace’s equation in cylindrical coordinates
/C26; ’;z: The cylindrical co-
ordinates are shown in Fig. 10.1, where
x/C26cos’
y/C26sin’
zz9
>=
>;or/C262x2y2
’tanÿ1
y=x
zz:8
><
>:
Laplace’s equation now reads
/C1142/C30
/C26; ’;z1
/C26/C64
/C64/C26/C26/C64/C30
/C64/C26
1
/C262/C642/C30
/C64’2/C642/C30
/C642z20:
10:15
395SOLUTIONS OF LAPLACE’S EQUATION
Figure 10.1. Cylindrical coordinates.
We assume that
/C30
/C26; ’;zR
/C26
’Z
z:
10:16
Substitution into Eq. (10.15) yields, after division by /C30,
1
/C26Rd
d/C26/C26dR
d/C26
1
/C262d2
d’2ÿ1
Zd2Z
dz2:
10:17
Clearly, both sides of Eq. (10.17) must be equal to a constant, say ÿk2. Then
1
Zd2Z
dz2k2ord2Z
dz2ÿk2Z0
10:18
and
1
/C26Rd
d/C26/C26dR
d/C26
1
/C262d2
d’2ÿk2
or
/C26
Rd
d/C26/C26dR
d/C26
k2/C262ÿ1
d2
d’2:
Both sides of this last equation must be equal to a constant, say 2. Hence
d2
d’220;
10:19
1
Rd
d/C26/C26dR
d/C26
k2ÿ2
/C262/C32!
R0:
10:20
Equation (10.18) has for solutions
Z
zc
kekzc0
keÿkz;k60;0<k<1;
c1zc2; k0;(
10:21
where candc0are arbitrary functions of kandc1andc2are arbitrary constants.
Equation (10.19) has solutions of the form
’a
ei’;60;ÿ1 << 1;
b’b0; 0:(
That the potential must be single-valued requires that
’
’2n, where
nis an integer. It follows from this that must be an integer or zero and that
b0. Then the solution
’becomes
’a
ei’a0
eÿi’;60;integer ;
b0; 0:(
10:22
396PARTIAL DIFFERENTIAL EQUATIONS
In the special case k0, Eq. (10.20) has solutions of the form
R
/C26d
/C26d0
/C26ÿ;60;
fln/C26/C103; 0:/C26
10:23
When k60, a simple change of variable can put Eq. (10.20) in the form of
Bessel’s equation. Let xk/C26, then dxkd/C26and Eq. (10.20) becomes
d2R
dx21
xdR
dx1ÿ2
x2/C32!
R0;
10:24
the well-known Bessel’s equation (Eq. (7.71)). As shown in Chapter 7, R
xcan be
written as
R
xA/C74
xB/C74ÿ
x;
10:25
where AandBare constants, and /C74
xis the Bessel function of the first kind.
When is not an integer, /C74and/C74ÿare independent. But when is an integer,
/C74ÿ
x
ÿ 1n/C74
x, thus /C74and/C74ÿare linearly dependent, and Eq. (10.25)
cannot be a general solution. In this case the general solution is given by
R
xA1/C74
xB1Y
x;
10:26
where A1andB2are constants; Y
xis the Bessel function of the second kind of
order or Neumann’s function of order /C78
x.
The general solution of Eq. (10.20) when k60 is therefore
R
/C26/C112
/C74
k/C26/C113
Y
k/C26;
10:27
where pand/C113are arbitrary functions of . Then these functions are also solu-
tions:
H
1
k/C26/C74
k/C26iY
k/C26;H
2
k/C26/C74
k/C26ÿiY
k/C26:
These are the Hankel functions of the first and second kinds of order , respec-
tively.
The functions /C74;Y(or/C78), and H
1
,a n d H
2
which satisfy Eq. (10.20) are
known as cylindrical functions of integral order and are denoted by Z
k/C26,
which is not the same as Z
z. The solution of Laplace’s equation (10.15) can now
be written
/C30
/C26; ’;z
c1zb
fln/C26/C103; k0; 0;
c1zbd
/C26d0
/C26ÿa
ei’a0
eÿi’;
k0; 60;
c
kekzc0
keÿkzZ0
k/C26; k60; 0;
c
kekzc0
keÿkzZ
k/C26a
ei’a0
eÿi’;k60; 60:8
>>>>>><
>>>>>>:
397SOLUTIONS OF LAPLACE’S EQUATION
Let us now apply the solutions of Laplace’s equation in cylindrical coordinates
to an infinitely long cylindrical conductor with radius land charge per unit length
. We want to find the potential at a point Pa distance /C26/C62/C108from the axis of the
cylindrical. Take the origin of the coordinates on the axis of the cylinder that is
taken to be the z-axis. The surface of the cylinder is an equipotential:
/C30
/C108const :forr/C108and all ’andz:
The secondary boundary condition is that
/C69ÿ/C64/C30=/C64/C26 =2/C108/C34forr/C108and all ’andz:
Of the four types of solutions to Laplace’s equation in cylindrical coordinates
listed above only the first can satisfy these two boundary conditions. Thus
/C30
/C26b
fln/C26/C103ÿ
2/C34ln/C26
/C108/C30
a:
(4) Laplace’s equation in spherical coordinates
r; ;’: The spherical coordinates
are shown in Fig. 10.2, where
xrsincos’;
yrsinsin’;
zrcos’:
Laplace’s equation now reads
/C1142/C30
r; ;’1
r/C64
/C64rr2/C64/C30
/C64r
1
r2sin/C64
/C64sin/C64/C30
/C64
1
r2sin2/C642/C30
/C64’20:
10:28
398PARTIAL DIFFERENTIAL EQUATIONS
Figure 10.2. Spherical coordinates.
Again, assume that
/C30
r; ;’R
r/C2
’:
10:29
Substituting into Eq. (10.28) and dividing by /C30we obtain
sin2
Rd
drr2dR
dr
sin
/C2d
dsind/C2
d
ÿ1
d2
d’2:
For a solution, both sides of this last equation must be equal to a constant, say
m2. Then we have two equations
d2
d’2m20;
10:30
sin2
Rd
drr2dR
dr
sin
/C2d
dsind/C2
d
m2;
the last equation can be rewritten as
1
/C2sind
dsind/C2
d
ÿm2
sin2ÿ1
Rd
drr2dR
dr
:
Again, both sides of the last equation must be equal to a constant, say ÿ/C12. This
yields two equations
1
Rd
drr2dR
dr
/C12;
10:31
1
/C2sind
dsind/C2
d
ÿm2
sin2ÿ/C12:
By a simple substitution: xcos, we can put the last equation in a more familiar
form:
d
dx
1ÿx2dP
dx
/C12ÿm2
1ÿx2/C32!
P0
10:32
or
1ÿx2d2P
dx2ÿ2xdP
dx/C12ÿm2
1ÿx2"#
P0;
10:32a
where we have set P
x/C2
.
You may have already noticed that Eq. (10.32) is very similar to Eq. (10.25),
the associated Legendre equation. Let us take a close look at this resemblance.
In Eq. (10.32), the points x1 are regular singular points of the equation. Let
us first study the behavior of the solution near point x1; it is convenient to
399SOLUTIONS OF LAPLACE’S EQUATION
bring this regular singular point to the origin, so we make the substitution
u1ÿx;U
uP
x. Then Eq. (10.32) becomes
d
duu
2ÿudU
du
/C12ÿm2
u
2ÿu"#
U0:
When we solve this equation by a power series: UP1
n0anun/C26, we find that the
indicial equation leads to the values m=2 for /C26. For the point xÿ1, we make
the substitution /C1181x, and then solve the resulting di/C128erential equation by the
power series method; we find that the indicial equation leads to the same values
m=2 for /C26.
Let us first consider the value m=2;m0. The above considerations lead us
to assume
P
x
1ÿxm=2
1xm=2y
x
1ÿx2m=2y
x; m0
as the solution of Eq. (10.32). Substituting this into Eq. (10.32) we find
1ÿx2d2y
dx2ÿ2
m1xdy
dx/C12ÿm
m1 y0:
Solving this equation by a power series
y
xX1
n0cnxn;
we find that the indicial equation is
ÿ10. Thus the solution can be written
y
xX
nevencnxnX
noddcnxn:
The recursion formula is
cn2
nm
nm1ÿ/C12
n1
n2cn:
Now consider the convergence of the series. By the ratio test,
Rncnxn
cnÿ2xnÿ2/C12/C12/C12/C12/C12/C12/C12/C12
nm
nm1ÿ/C12
n1
n2/C12/C12/C12/C12/C12/C12/C12/C12xjj
2:
The series converges for jxj<1, whatever the finite value of /C12may be. For
jxj1, the ratio test is inconclusive. However, the integral test yields
Z
M
tm
tm1ÿ/C12
t1
t2dtZ
M
tm
tm1
t1
t2dtÿZ
M/C12
t1
t2dt
and since
Z
M
tm
tm1
t1
t2dt!1 asM!1 ;
400PARTIAL DIFFERENTIAL EQUATIONS
the series diverges for jxj1. A solution which converges for all xcan be
obtained if either the even or odd series is terminated at the term in xj. This
may be done by setting /C12equal to
/C12
jm
jm1/C108
/C1081:
On substituting this into Eq. (10.32a), the resulting equation is
1ÿx2d2P
dx2ÿ2xdP
dx/C108
/C1081ÿm2
1ÿx2"#
P0;
which is identical to Eq. (7.25). Special solutions were studied there: they were
written in the form Pm
/C108
xand are known as the associated Legendre functions of
the first kind of degree land order m, where land m, take on the values
/C1080;1;2;...;andm0;1;2;...;/C108. The general solution of Eq. (10.32) for
m0 is therefore
P
x/C2
a/C108Pm/C108
x:
10:33
The second solution of Eq. (10.32) is given by the associated Legendre function of
the second kind of degree land order m:/C81m
/C108
x. However, only the associated
Legendre function of the first kind remains finite over the range ÿ1x1 (or
02.
Equation (10.31) for R
rbecomes
d
drr2dR
dr
ÿ/C108
/C1081R0:
10:31a
When /C10860, its solution is
R
rb
/C108r/C108b0
/C108rÿ/C108ÿ1;
10:34
and when /C1080, its solution is
R
rcrÿ1d:
10:35
The solution of Eq. (10.30) is
f
meim’f0
/C108eÿim’;m60;positive integer ;
/C103; m0:(
10:36
The solution of Laplace’s equation (10.28) is therefore given by
/C30
r; ;’br/C108b0rÿ/C108ÿ1Pm/C108
cosfeim’f0eÿim’;/C10860;m60;
br/C108b0rÿ/C108ÿ1P/C108
cos; /C10860;m0;
crÿ1dP0
cos; /C1080;m0;8
><
>:
10:37
where P/C108P0
/C108.
401SOLUTIONS OF LAPLACE’S EQUATION
We now illustrate the usefulness of the above result for an electrostatic problem
having spherical symmetry. Consider a conducting spherical shell of radius aand
charge per unit area. The problem is to find the potential /C30
r; ;’at a point Pa
distance r/C62afrom the center of shell. Take the origin of coordinates to be at the
center of the shell. As the surface of the shell is an equipotential, we have the first
boundary condition
/C30
rconstant /C30
aforraand all and’:
10:38
The second boundary condition is that
/C30!0 for r!1 and all and’:
10:39
Of the three types of solutions (10.37) only the last can satisfy the boundaryconditions. Thus
/C30
r; ;’
cr
ÿ1dP0
cos:
10:40
Now P0
cos1, and from Eq. (10.38) we have
/C30
acaÿ1d:
But the boundary condition (10.39) requires that d0. Thus /C30
acaÿ1,o r
ca/C30
a, and Eq. (10.40) reduces to
/C30
ra/C30
a
r:
10:41
Now
/C30
a=a/C69
a/C81=4a2/C34;
where /C34is the permittivity of the dielectric in which the shell is embedded,
/C814a2. Thus /C30
aa=/C34, and Eq. (10.41) becomes
/C30
ra2
/C34r:
10:42
/C83olutions of the /C119a/C118e equation/C58 separation of /C118ariables
We now use the method of separation of variables to solve the wave equation
/C642u
x;t
/C64x2/C118ÿ2/C642u
x;t
/C64t2;
10:43
subject to the following boundary conditions:
u
0;tu
/C108;t0;t0;
10:44
u
x;0f
t;0x/C108;
10:45
402PARTIAL DIFFERENTIAL EQUATIONS
and
/C64u
x;t
/C64t/C12/C12/C12/C12
t0/C103
x;0x/C108;
10:46
where fandgare given functions.
Assuming that the solution of Eq. (10.43) may be written as a product
u
x;tX
xT
t;
10:47
then substituting into Eq. (10.43) and dividing by XTwe obtain
1
Xd2X
dx21
/C1182Td2T
dt2:
Both sides of this last equation must be equal to a constant, say ÿb2=/C1182. Then we
have two equations
1
Xd2X
dx2ÿb2
/C1182;
10:48
1
Td2T
dt2ÿb2:
10:49
The solutions of these equations are periodic, and it is more convenient to write
them in terms of trigonometric functions
X
xAsinbx
/C118Bcosbx
/C118; T
tCsinbtDcosbt;
10:50
where A;B;C, and /C68are arbitrary constants, to be fixed by the boundary condi-
tions. Equation (10.47) then becomes
u
x;t Asinbx
/C118Bcosbx
/C118
CsinbtDcosbt:
10:51
The boundary condition u
0;t0
t/C620gives
0B
CsinbtDcosbt
for all t, which implies
B0:
10:52
Next, from the boundary condition u
/C108;t0
t/C620we have
0Asinb/C108
/C118
CsinbtDcosbt:
Note that B0 would make u0. However, the last equation can be satisfied
for all twhen
sinb/C108
/C1180;
403SOLUTIONS OF LAPLACE’S EQUATION
which implies
bn/C118
/C108;n1;2;3;...:
10:53
Note that ncannot be equal to zero, because it would make b0, which in turn
would make u0.
Substituting Eq. (10.53) into Eq. (10.51) we have
un
x;tsinnx
/C108Cnsinn/C118t
/C108Dncosn/C118t
/C108
; n1;2;3;...:
10:54
We see that there is an infinite set of discrete values of band that to each value of
bthere corresponds a particular solution. Any linear combination of these parti-
cular solutions is also a solution:
un
x;tX1
n1sinnx
/C108Cnsinn/C118t
/C108Dncosn/C118t
/C108
:
10:55
The constants CnandDnare fixed by the boundary conditions (10.45) and
(10.46).
Application of boundary condition (10.45) yields
f
xX1
n1Dnsinnx
/C108:
10:56
Similarly, application of boundary condition (10.46) gives
/C103
x/C118
/C108X1
n1nCnsinnx
/C108:
10:57
The coecients CnandDnmay then be determined by the Fourier series method:
Dn2
/C108Z/C108
0f
xsinnx
/C108dx; Cn2
n/C118Z/C108
0/C103
xsinnx
/C108dx:
10:58
We can use the method of separation of variable to solve the heat conduction
equation. We shall leave this as a home work problem.
In the following sections, we shall consider two more methods for the solution
of linear partial di/C128erential equations: the method of Green’s functions, and the
method of the Laplace transformation which was used in Chapter 9 for the
solution of ordinary linear di/C128erential equations with constant coecients.
/C83olution of Poisson/C39s equation/C46 /C71reen/C39s functions
The Green’s function approach to boundary-value problems is a very powerfultechnique. The field at a point caused by a source can be considered to be the total
e/C128ect due to each ‘‘unit’’ (or elementary portion) of the source. If /C71
x;x
0is the
404PARTIAL DIFFERENTIAL EQUATIONS
field at a point xdue to a unit point source at x0, then the total field at xdue to a
distributed source /C26
x0is the integral of /C71/C26over the range of x0occupied by the
source. The function /C71
x;x0is the well-known Green’s function. We now apply
this technique to solve Poisson’s equation for electric potential /C30(Example 10.2)
/C1142/C30
rÿ1
/C34/C26
r;
10:59
where /C26is the charge density and /C34the permittivity of the medium, both are given.
By definition, Green’s function /C71
r;r0is the solution of
/C1142/C71
r;r0
rÿr0;
10:60
where
rÿr0is the Dirac delta function.
Now, multiplying Eq. (10.60) by /C30and Eq. (10.59) by /C71, and then subtracting,
we find
/C30
r/C1142/C71
r;r0ÿ/C71
r;r0/C1142/C30
r/C30
r
rÿr01
/C34/C71
r;r0/C26
r;
and on interchanging rand r0,
/C30
r0/C11402/C71
r0;rÿ/C71
r0;r/C11402/C30
r0/C30
r0
r0ÿr1
/C34/C71
r0;r/C26
r0
or
/C30
r0
r0ÿr/C30
r0/C11402/C71
r0;rÿ/C71
r0;r/C11402/C30
r0ÿ1
/C34/C71
r0;r/C26
r0;
10:61
the prime on /C114indicates that di/C128erentiation is with respect to the primed co-
ordinates. Integrating this last equation over all r0within and on the surface S0
which encloses all sources (charges) yields
/C30
rÿ1
/C34Z
/C71
r;r0/C26
r0dr0
Z
/C30
r0/C11402/C71
r;r0ÿ/C71
r;r0/C11402/C30
r0dr0;
10:62
where we have used the property of the delta function
Z1
ÿ1f
r0
rÿr0dr0f
r:
We now use Green’s theorem
ZZZ
f/C11402/C32ÿ/C32/C11402fd/C280ZZ
f/C1140/C32ÿ/C32/C1140fd/C83
405SOLUTIONS OF POISSON’S EQUATION
to transform the second term on the right hand side of Eq. (10.62) and obtain
/C30
rÿ1
/C34Z
/C71
r;r0/C26
r0dr0
Z
/C30
r0/C1140/C71
r;r0ÿ/C71
r;r0/C1140/C30
r0 d/C830
10:63
or
/C30
rÿ1
/C34Z
/C71
r;r0/C26
r0dr0
Z
/C30
r0/C64
/C64n0/C71
r;r0ÿ/C71
r;r0/C64
/C64n0/C30
r0
d/C830;
10:64
where n0is the outward normal to dS0. The Green’s function /C71
r;r0can be found
from Eq. (10.60) subject to the appropriate boundary conditions.
If the potential /C30vanishes on the surface S0or/C64/C30=/C64n0vanishes, Eq. (10.64)
reduces to
/C30
rÿ1
/C34Z
/C71
r;r0/C26
r0dr0:
10:65
On the other hand, if the surface S0encloses no charge, then Poisson’s equation
reduces to Laplace’s equation and Eq. (10.64) reduces to
/C30
rZ
/C30
r0/C64
/C64n0/C71
r;r0ÿ/C71
r;r0/C64
/C64n0/C30
r0
d/C830:
10:66
The potential at a field point rdue to a point charge /C113located at the point r0is
/C30
r1
4/C34/C113
rÿr0jj:
Now
/C1142 1
rÿr0jj
ÿ4
rÿr0
(the proof is left as an exercise for the reader) and it follows that the Green’s
function /C71
r;r0in this case is equal
/C71
r;r01
4/C341
rÿr0jj:
If the medium is bounded, the Green’s function can be obtained by direct solutionof Eq. (10.60) subject to the appropriate boundary conditions.
To illustrate the procedure of the Green’s function technique, let us consider a
simple example that can easily be solved by other methods. Consider two
grounded parallel conducting plates of infinite extent: if the electric charge density
/C26between the two plates is given, find the electric potential distribution /C30between
406PARTIAL DIFFERENTIAL EQUATIONS
the plates. The electric potential distribution /C30is described by solving Poisson’s
equation
/C1142/C30ÿ/C26=/C34
subject to the boundary conditions
(1)/C30
00;
(2)/C30
10:
We take the coordinates shown in Fig. 10.3. Poisson’s equation reduces to the
simple form
d2/C30
dx2ÿ/C26
/C34:
10:67
Instead of using the general result (10.64), it is more convenient to proceed
directly. Multiplying Eq. (10.67) by /C71
x;x0and integrating, we obtain
Z1
0/C71d2/C30
dx2dxÿZ1
0/C26
x/C71
/C34dx:
10:68
Then using integration by parts gives
Z1
0/C71d2/C30
dx2dx/C71
x;x0d/C30
x
dx/C12/C12/C12/C121
0ÿZ1
0d/C71
dxd/C30
dxdx
and using integration by parts again on the right hand side, we obtain
ÿZ1
0/C71d2/C30
dx2dxÿ/C71
x;x0d/C30
x
dx1
0d/C71
dx/C3010
ÿZ1
0/C30d2/C71
dx2dx/C12/C12/C12/C12/C12"#/C12/C12/C12/C12/C12
/C71
0;x
0d/C30
0
dxÿ/C71
1;x0d/C30
1
dxÿZ1
0/C30d2/C71
dx2dx:
407SOLUTIONS OF POISSON’S EQUATION
Figure 10.3.
Substituting this into Eq. (10.68) we obtain
/C71
0;x0d/C30
0
dxÿ/C71
1;x0d/C30
1
dxÿZ1
0/C30d2/C71
dx2dxZ1
0/C71
x;x0/C26
x
/C34dx
or
Z1
0/C30d2/C71
dx2dx/C71
1;x0d/C30
1
dxÿ/C71
0;x0d/C30
0
dxÿZ1
0/C71
x;x0/C26
x
/C34dx:
10:69
We must now choose a Green’s function which satisfies the following equation
and the boundary conditions:
d2/C71
dx2ÿ
xÿx0;/C71
0;x0/C71
1;x00:
10:70
Combining these with Eq. (10.69) we find the solution to be
/C30
x0Z1
01
/C34/C26
x/C71
x;x0dx:
10:71
It remains to find /C71
x;x0. By integration, we obtain from Eq. (10.70)
d/C71
dxÿZ
xÿx0dxaÿU
xÿx0a;
where Uis the unit step function and ais an integration constant to be determined
later. Integrating once we get
/C71
x;x0ÿZ
U
xÿx0dxaxbÿ
xÿx0U
xÿx0axb:
Imposing the boundary conditions on this general solution yields two equations:
/C71
0;x0x0U
ÿx0a0b00b0;
/C71
1;x0ÿ
1ÿx0U
1ÿx0ab0:
From these we find
a
1ÿx0U
1ÿx0;b0
and the Green’s function is
/C71
x;x0ÿ
xÿx0U
xÿx0
1ÿx0x:
10:72
This gives the response at x0due to a unit source at x. Interchanging xandx0in
Eqs. (10.70) and (10.71) we find the solution of Eq. (10.67) to be
/C30
xZ1
01
/C34/C26
x0/C71
x0;xdx0Z1
01
/C34/C26
x0ÿ
x0ÿxU
x0ÿx
1ÿxx0dx0:
10:73
408PARTIAL DIFFERENTIAL EQUATIONS
Note that the Green’s function in the last equation can be written in the form
/C71
x;x0
1ÿxxx <x0
1ÿxx0x/C62x0(
:
Laplace transform solutions of boundar/C121-/C118alue problems
Laplace and Fourier transforms are useful in solving a variety of partial di/C128er-
ential equations, the choice of the appropriate transforms depends on the type of
boundary conditions imposed on the problem. To illustrate the use of the Lapace
transforms in solving boundary-value problems, we solve the following equation:
/C64u
/C64t2/C642u
/C64x2;
10:74
u
0;tu
3;t0; u
x;010 sin 2 xÿ6 sin 4 x:
10:75
Taking the Laplace transform of Eq. (10.74) with respect to tgives
L/C64u
/C64t
2L/C642u
/C64x2"#
:
Now
L/C64u
/C64t
/C112L u
ÿ u
x;0
and
L/C642u
/C64x2"#
Z1
0eÿ/C112t/C642u
/C64x2dt/C642
/C64x2Z1
0eÿ/C112tu
x;tdt/C642
/C64x2Lu:
Here /C642=/C64x2andR1
0dtare interchangeable because xandtare independent.
For convenience, let
UU
x;/C112Lu
x;t Z1
0eÿ/C112tu
x;tdt:
We then have
/C112Uÿu
x;02d2U
dx2;
from which we obtain, on using the given condition (10.75),
d2U
dx2ÿ1
2/C112U3 sin 4 xÿ5 sin 2 x:
10:76
409BOUNDARY-VALUE PROBLEMS
Now think of this as a di/C128erential equation in terms of x,w i t h pas a parameter.
Then taking the Laplace transform of the given conditions u
0;tu
3;t0,
we have
Lu
0;t 0;Lu
3;t 0
or
U
0;/C1120;U
3;/C1120:
These are the boundary conditions on U
x;/C112. Solving Eq. (10.76) subject to these
conditions we find
U
x;/C1125 sin 2 x
/C112162ÿ3 sin 4 x
/C112642:
The solution to Eq. (10.74) can now be obtained by taking the inverse Laplace
transform
u
x;tLÿ1U
x;/C112 5eÿ162tsin 2xÿ3eÿ642sin 4x:
The Fourier transform method was used in Chapter 4 for the solution of
ordinary linear ordinary di/C128erential equations with constant coecients. It can
be extended to solve a variety of partial di/C128erential equations. However, we shall
not discuss this here. Also, there are other methods for the solution of linear
partial di/C128erential equations. In general, it is a dicult task to solve partial
di/C128erential equations analytically, and very often a numerical method is the
best way of obtaining a solution that satisfies given boundary conditions.
Problems
10.1 ( a) Show that y
x;t/C70
2x5t/C71
2xÿ5tis a general solution of
4/C642y
/C64t225/C642y
/C64x2:
(b) Find a particular solution satisfying the conditions
y
0;ty
;t0;y
x;0sin 2x;y0
x;00:
10.2. State the nature of each of the following equations (that is, whether elliptic,
parabolic, or hyperbolic)
a/C642y
/C64t2/C642y
/C64x20;
bx/C642u
/C64x2y/C642u
/C64y23y2/C64u
/C64x:
10.3 The electromagnetic wave equation: Classical electromagnetic theory was
worked out experimentally in bits and pieces by Coulomb, Oersted, Ampere,
Faraday and many others, but the man who put it all together and built it
into the compact and consistent theory it is today was James Clerk Maxwell.
410PARTIAL DIFFERENTIAL EQUATIONS
His work led to the understanding of electromagnetic radiation, of which
light is a special case.
Given the four Maxwell equations
/C114/C69/C26=/C34 0;
Gauss’ law ;
/C114 /C66/C220/C106/C340/C64/C69=/C64t
Ampere’s law ;
/C114/C660
Gauss’ law ;
/C114 /C69ÿ/C64/C66=/C64t
Faraday’s law ;
where /C66is the magnetic induction, /C106/C26/C118is the current density, and /C220is the
permeability of the medium, show that:
(a) the electric field and the magnetic induction can be expressed as
/C69ÿ /C114 /C30ÿ/C64/C65=/C64t; /C66/C114 /C65;
where /C65is called the vector potential, and /C30the scalar potential. It
should be noted that /C69and /C66are invariant under the following trans-
formations:
/C650/C65/C114/C31; /C300/C30ÿ/C64/C30=/C64 t
in which /C31is an arbitrary real function. That is, both ( /C650;/C30, and
(/C650;/C300) yield the same /C69and/C66. Any condition which, for computational
convenience, restricts the form of /C65and/C30is said to define a gauge. Thus
the above transformation is called a gauge transformation and /C31is
called a gauge parameter.
(b) If we impose the so-called Lorentz gauge condition on /C65and/C30:
/C114/C65/C220/C340
/C64/C30=/C64 t0;
then both /C65and/C30satisfy the following wave equations:
/C1142/C65ÿ/C220/C340/C642/C65
/C64t2ÿ/C220/C106;
/C1142/C30ÿ/C220/C340/C642/C30
/C64t2ÿ/C26=/C34 0:
10.4 Given Gauss’ lawRR
S/C69ds/C113=/C34, find the electric field produced by a
charged plane of infinite extension is given by /C69=/C34, where is the charge
per unit area of the plane.
10.5 Consider an infinitely long uncharged conducting cylinder of radius lplaced
in an originally uniform electric field /C690directed at right angles to the axis of
the cylinder. Find the potential at a point /C26
/C62/C108from the axis of the cylin-
der. The boundary conditions are:
411PROBLEMS
/C30
/C26; ’ÿ/C690/C26cos’ÿ/C690xfor /C26!1 ;
0f o r /C26/C108;/C26
where the x-axis has been taken in the direction of the uniform field /C690.
10.6 Obtain the solution of the heat conduction equation
/C642u
x;t
/C64x21
/C64u
x;t
/C64t
which satisfies the boundary conditions
(1)u
0;tu
/C108;t0;t0;(2)u
x;0f
x;0x, where f
xis a
given function and lis a constant.
10.7 If a battery is connected to the plates as shown in Fig. 10.4, and if the charge
density distribution between the two plates is still given by /C26
x, find the
potential distribution between the plates.
10.8 Find the Green’s function that satisfies the equation
d2/C71
dx2
xÿx0
and the boundary conditions /C710 when x0 and /C71remains bounded
when xapproaches infinity. (This Green’s function is the potential due to a
surface charge ÿ/C34per unit area on a plane of infinite extent located at xx0
in a dielectric medium of permittivity /C34when a grounded conducting plane
of infinite extent is located at x0.)
10.9 Solve by Laplace transforms the boundary-value problem
/C642u
/C64x21
/C75/C64u
/C64tfor x/C620;t/C620;
given that uu0(a constant) on x0 for t/C620, and u0 for x/C620;t0.
412PARTIAL DIFFERENTIAL EQUATIONS
Figure 10.4.
11
Simple linear integral e/C113uations
In previous chapters we have met equations in which the unknown functions
appear under an integral sign. Such equations are called integral equations.
Fourier and Laplace transforms are important integral equations, In Chapter 4,
by introducing the method of Green’s function we were led in a natural way to
reformulate the problem in terms of integral equations. Integral equations have
become one of the very useful and sometimes indispensable mathematical tools of
theoretical physics and engineering.
/C67lassification of linear integral equations
In this chapter we shall confine our attention to linear integral equations. Linearintegral equations can be divided into two major groups:
(1) If the unknown function occurs only under the integral sign, the integral
equation is said to be of the first kind. Integral equations having the
unknown function both inside and outside the integral sign are of the second
kind.
(2) If the limits of integration are constants, the equation is called a Fredholm
integral equation. If one limit is variable, it is a Volterra equation.
These four kinds of linear integral equations can be written as follows:
f
xZ
b
a/C75
x;tu
tdt Fredholm equation of the first kind;
11:1
u
xf
xZb
a/C75
x;tu
tdtFredholm equation of the second kind;
11:2
413
f
xZx
a/C75
x;tu
tdt Volterra equation of the first kind;
11:3
u
xf
xZx
a/C75
x;tu
tdtVolterra equation of the second kind :
11:4
In each case u
tis the unknown function, /C75
x;tandf
xare assumed to be
known. /C75
x;tis called the kernel or nucleus of the integral equation. is a
parameter, which often plays the role of an eigenvalue. The equation is said to
be homogeneous if f
x0.
If one or both of the limits of integration are infinite, or the kernel /C75
x;t
becomes infinite in the range of integration, the equation is said to be singular;
special techniques are required for its solution.
The general linear integral equation may be written as
/C104
xu
xf
xZb
a/C75
x;tu
tdt:
11:5
If/C104
x0, we have a Fredholm equation of the first kind; if /C104
x1, we have a
Fredholm equation of the second kind. We have a Volterra equation when theupper limit is x.
It is beyond the scope of this book to present the purely mathematical general
theory of these various types of equations. After a general discussion of a fewmethods of solution, we will illustrate them with some simple examples. We will
then show with a few examples from physical problems how to convert di/C128erential
equations into integral equations.
/C83ome methods of solution
Separable /C107ernel
When the two variables xand twhich appear in the kernel /C75
x;tare separable,
the problem of solving a Fredholm equation can be reduced to that of solving a
system of algebraic equations, a much easier task. When the kernel /C75
x;tcan be
written as
/C75
x;tX
n
i1/C103i
x/C104i
t;
11:6
where /C103
xis a function of xonly and /C104
ta function of tonly, it is said to be
degenerate. Putting Eq. (11.6) into Eq. (11.2), we obtain
u
xf
xXn
i1Zb
a/C103i
x/C104i
tu
tdt:
414SIMPLE LINEAR INTEGRAL EQUATIONS
Note that /C103
xis a constant as far as the tintegration is concerned, hence it may
be taken outside the integral sign and we have
u
xf
xXn
i1/C103i
xZb
a/C104i
tu
tdt:
11:7
Now
Zb
a/C104i
tu
tdtCi
const ::
11:8
Substituting this into Eq. (11.7) and solving for u
t, we obtain
u
tf
xCXn
i1/C103i
x:
11:9
The value of Cimay now be obtained by substituting Eq. (11.9) into Eq. (11.8).
The solution is only valid for certain values of , and we call these the eigenvalues
of the integral equation. The homogeneous equation has non-trivial solutions
only if is one of these eigenvalues; these solutions are called eigenfunctions of
the kernel (operator) /C75.
Example 11.1As an example of this method, we consider the following equation:
u
xxZ
1
0
xt2x2tu
tdt:
11:10
This is a Fredholm equation of the second kind, with f
xxand
/C75
x;txt2x2t. If we define
Z1
0t2u
tdt;/C12 Z1
0tu
tdt;
11:11
then Eq. (11.10) becomes
u
xx
x/C12x2:
11:12
To determine Aand B, we put Eq. (11.12) back into Eq. (11.11) and obtain
1
41415/C12; /C12 131314/C12:
11:13
Solving this for and/C12we find
60
240ÿ120ÿ2;/C12 80
240ÿ120ÿ2;
and the final solution is
u
t
240ÿ60x80x2
240ÿ120ÿ2:
415SOME METHODS OF SOLUTION
The solution blows up when 117:96 or 2:04. These are the eigenvalues of
the integral equation.
Fredholm found that if: (1) f
xis continuous, (2) /C75
x;tis piecewise contin-
uous, (3) the integralsRR
/C752
x;tdxdt;R
f2
tdtexist, and (4) the integralsRR
/C752
x;tdtandR
/C752
t;xdtare bounded, then the following theorems apply:
(a) Either the inhomogeneous equation
u
xf
xZb
a/C75
x;tu
tdt
has a unique solution for any function f
x
is not an eigenvalue), or the
homogeneous equation
u
xZb
a/C75
x;tu
tdt
has at least one non-trivial solution corresponding to a particular value of .
In this case, is an eigenvalue and the solution is an eigenfunction.
(b)I fis an eigenvalue, then is also an eigenvalue of the transposed equation
u
xZb
a/C75
t;xu
tdt;
and, if is not an eigenvalue, then is also not an eigenvalue of the
transposed equation
u
xf
xZb
a/C75
t;xu
tdt:
(c)I fis an eigenvalue, the inhomogeneous equation has a solution if, and only
if,
Zb
au
xf
xdx0
for every function f
x.
We refer the readers who are interested in the proof of these theorems to the
book by R. Courant and D. Hilbert ( Methods of Mathematical Physics , Vol. 1,
Wiley, 1961).
Neumann series solutions
This method is due largely to Neumann, Liouville, and Volterra. In this method
we solve the Fredholm equation (11.2)
u
xf
xZb
a/C75
x;tu
tdt
416SIMPLE LINEAR INTEGRAL EQUATIONS
by iteration or successive approximations, and begin with the approximation
u
x/C25u0
x/C25f
x:
This approximation is equivalent to saying that the constant or the integral is
small. We then put this crude choice into the integral equation (11.2) under the
integral sign to obtain a second approximation:
u1
xf
xZb
a/C75
x;tf
tdt
and the process is then repeated and we obtain
u2
xf
xZb
a/C75
x;tf
tdt2Zb
aZb
a/C75
x;t/C75
t;t0f
t0dt0dt:
We can continue iterating this process, and the resulting series is known as theNeumann series, or Neumann solution:
u
xf
xZ
b
a/C75
x;tf
tdt2Zb
aZb
a/C75
x;t/C75
t;t0f
t0dt0dt :
This series can be written formally as
un
xXn
i1i’i
x;
11:14
where
’0
xu0
xf
x;
’1
xZb
a/C75
x;t1f
t1dt1;
’2
xZb
aZb
a/C75
x;t1/C75
t1;t2f
t2dt1dt2;
...
’n
xZb
aZb
aZb
a/C75
x;t1/C75
t1;t2/C75
tnÿ1;tnf
tndt1dt2dtn:9
>>>>>>>>>>>>>>=
>>>>>>>>>>>>>>;
11:15
The series (11.14) will converge for suciently small , when the kernel /C75
x;tis
bounded. This can be checked with the Cauchy ratio test (Problem 11.4).
Example 11.2
Use the Neumann method to solve the integral equation
u
xf
x
1
2Z1
ÿ1/C75
x;tu
tdt;
11:16
417SOME METHODS OF SOLUTION
where
f
xx; /C75
x;ttÿx:
Solution: We begin with
u0
xf
xx:
Then
u1
xx1
2Z1
ÿ1
tÿxtdtx13:
Putting u
1
xinto Eq. (11.16) under the integral sign, we obtain
u2
xx12Z
1
ÿ1
tÿxt13
dtx13ÿx
3:
Repeating this process of substituting back into Eq. (11.16) once more, we obtain
u3
xx13ÿx
3ÿ1
32:
We can improve the approximation by iterating the process, and the convergence
of the resulting series (solution) can be checked out with the ratio test.
The Neumann method is also applicable to the Volterra equation, as shown by
the following example.
Example 11.3
Use the Neumann method to solve the Volterra equation
u
x1Zx
0u
tdt:
Solution: We begin with the zeroth approximation u0
x1. Then
u1
x1Zx
0u0
tdt1Zx
0dt1x:
This gives
u2
x1Zx
0u1
tdt1Zx
0
1tdt1x1
22x2;
similarly,
u3
x1Zx
01t12
2t2
dt1t12
2t21
3/C333x3:
418SIMPLE LINEAR INTEGRAL EQUATIONS
By induction
un
xXn
k11
k/C33kxk:
When n!1 ,un
xapproaches
u
xex:
/C84ransformation of an integral equation into a di/C128erential equation
Sometimes the Volterra integral equation can be transformed into an ordinary
di/C128erential equation which may be easier to solve than the original integral equa-
tion, as shown by the following example.
Example 11.4
Consider the Volterra integral equation u
x2x4Rx
0
tÿxu
tdt. Before we
transform it into a di/C128erential equation, let us recall the following very usefulformula: if
I
Z
b
a
f
x;dx;
where aandbare continuous and at least once di/C128erentiable functions of , then
dI
df
b;db
dÿf
a;da
dZb
a/C64f
x;
/C64dx:
With the help of this formula, we obtain
d
dxu
x24
tÿxu
t fgtxÿZx
0u
tdt
2ÿ4Zx
0u
tdt:
Di/C128erentiating again we obtain
d2u
x
dx2ÿ4u
x:
This is a di/C128erentiation equation equivalent to the original integral equation, butits solution is much easier to find:
u
xAcos 2 xBsin 2x;
where Aand Bare integration constants. To determine their values, we put the
solution back into the original integral equation under the integral sign, and then
419SOME METHODS OF SOLUTION
integration gives A0 and B1. Thus the solution of the original integral
equation is
u
xsin 2x:
/C76aplace transform solution
The Volterra integral equation can sometime be solved with the help of the
Laplace transformation and the convolution theorem. Before we consider the
Laplace transform solution, let us review the convolution theorem. If f1
xand
f2
xare two arbitrary functions, we define their convolution ( faltung in German)
to be
/C103
xZ1
ÿ1f1
yf2
xÿydy:
Its Laplace transform is
L/C103
x Lf1
xLf2
x:
We now consider the Volterra equation
u
xf
xZx
0/C75
x;tu
tdt
f
xZx
0/C103
xÿtu
tdt;
11:17
where /C75
xÿt/C103
xÿt, a so-called displacement kernel. Taking the Laplace
transformation and using the convolution theorem, we obtain
LZx
0/C103
xÿtu
tdt
L/C103
xÿt Lu
t /C71
/C112U
/C112;
where U
/C112Lu
t R1
0eÿ/C112tu
tdt, and similarly for /C71
/C112. Thus, taking the
Laplace transformation of Eq. (11.17), we obtain
U
/C112/C70
/C112/C71
/C112U
/C112
or
U
/C112/C70
/C112
1ÿ/C71
/C112:
Inverting this we obtain u
t:
u
tLÿ1 /C70
/C112
1ÿ/C71
/C112
:
420SIMPLE LINEAR INTEGRAL EQUATIONS
Fourier transform solution
If the kernel is a displacement kernel and if the limits are ÿ1and1, we can use
Fourier transforms. Consider a Fredholm equation of the second kind
u
xf
xZ1
ÿ1/C75
xÿtu
tdt:
11:18
Taking Fourier transforms (indicated by overbars)
1
2pZ1
ÿ1dxf
xeÿi/C112xf
/C112;etc:;
and using the convolution theorem
Z1
ÿ1f
t/C103
xÿtdtZ1
ÿ1f
y/C103
yeÿiyxdy;
we obtain the transform of our integral equation (11.18):
u
/C112 f
/C112/C75
/C112u
/C112:
Solving for u
/C112we obtain
u
/C112f
/C112
1ÿ/C75
/C112:
If we can invert this equation, we can solve the original integral equation:
u
x1 2pZ
1
ÿ1f
teÿixt
1ÿ 2p
/C75
t:
11:19
/C84he /C83chmidt/C177/C72ilbert method of solution
In many physical problems, the kernel may be symmetric. In such cases, the
integral equation may be solved by a method quite di/C128erent from any of thosein the preceding section. This method, devised by Schmidt and Hilbert, is based
on considering the eigenfunctions and eigenvalues of the homogeneous integral
equation.
A kernel /C75
x;tis said to be symmetric if /C75
x;t/C75
t;xand Hermitian if
/C75
x;y/C75/C42
t;x. We shall limit our discussion to such kernels.
(a) The homogeneous Fredholm equation
u
xZ
b
a/C75
x;tu
tdt:
421THE SCHMIDT–HILBERT METHOD OF SOLUTION
A Hermitian kernel has at least one eigenvalue and it may have an infinite num-
ber. The proof will be omitted and we refer interested readers to the book byCourant and Hibert mentioned earlier (Chapter 3).
The eigenvalues of a Hermitian kernel are real, and eigenfunctions belonging to
di/C128erent eigenvalues are orthogonal; two functions f
xand/C103
xare said to be
orthogonal if
Z
f/C42
x/C103
xdx0:
To prove the reality of the eigenvalue, we multiply the homogeneous Fredholm
equation by u/C42
x, then integrating with respect to x, we obtain
Z
b
au/C42
xu
xdxZb
aZb
a/C75
x;tu/C42
xu
tdtdx:
11:20
Now, multiplying the complex conjugate of the Fredholm equation by u
xand
then integrating with respect to x,w eg e t
Zb
au/C42
xu
xdx/C42Zb
aZb
a/C75/C42
x;tu/C42
tu
xdtdx:
Interchanging xandton the right hand side of the last equation and remembering
that the kernel is Hermitian /C75/C42
t;x/C75
x;t, we obtain
Zb
au/C42
xu
xdx/C42Zb
aZb
a/C75
x;tu
tu/C42
xdtdx:
Comparing this equation with Eq. (11.2), we see that /C42, that is, is real.
We now prove the orthogonality. Let i,jbe two di/C128erent eigenvalues and
ui
x;uj
x, the corresponding eigenfunctions. Then we have
ui
xiZb
a/C75
x;tui
tdt; uj
xjZb
a/C75
x;tuj
tdt:
Now multiplying the first equation by uj
x, the second by iui
x, and then
integrating with respect to x, we obtain
jZb
aui
xuj
xdxijZb
aZb
a/C75
x;tui
tuj
xdtdx;
iZb
aui
xuj
xdxijZb
aZb
a/C75
x;tuj
tui
xdtdx:
11:21
Now we interchange xandton the right hand side of the last integral and because
of the symmetry of the kernel, we have
iZb
aui
xuj
xdxijZb
aZb
a/C75
x;tui
tuj
xdtdx:
11:22
422SIMPLE LINEAR INTEGRAL EQUATIONS
Subtracting Eq. (11.21) from Eq. (11.22), we obtain
iÿjZb
aui
xuj
xdx0:
11:23
Since i6j, it follows that
Zb
aui
xuj
xdx0:
11:24
Such functions may always be nomalized. We will assume that this has been done
and so the solutions of the homogeneous Fredholm equation form a complete
orthonomal set:
Zb
aui
xuj
xdxij:
11:25
Arbitrary functions of x, including the kernel for fixed t, may be expanded in
terms of the eigenfunctions
/C75
x;tX
Ciui
x:
11:26
Now substituting Eq. (11.26) into the original Fredholm equation, we have
uj
tjZb
a/C75
t;xuj
xdxjZb
a/C75
x;tuj
xdx
jX
iZb
aCiui
xuj
xdxjX
iCiijjCj
or
Ciui
t=i
and for our homogeneous Fredholm equation of the second kind the kernel maybe expressed in terms of the eigenfunctions and eigenvalues as
/C75
x;tX
1
n1un
xun
t
n:
11:27
The Schmidt–Hilbert theory does not solve the homogeneous integral equation;
its main function is to establish the properties of the eigenvalues (reality) and
eigenfunctions (orthogonality and completeness). The solutions of the homoge-
neous integral equation come from the preceding section on methods of solution.
(b) Solution of the inhomogeneous equation
u
xf
xZb
a/C75
x;tu
tdt:
11:28
We assume that we have found the eigenfunctions of the homogeneous equation
by the methods of the preceding section, and we denote them by ui
x. We may
423THE SCHMIDT–HILBERT METHOD OF SOLUTION
now expand both u
xandf
xin terms of ui
x, which forms an orthonormal
complete set.
u
xX1
n1nun
x;f
xX1
n1/C12nun
x:
11:29
Substituting Eq. (11.29) into Eq. (11.28), we obtain
Xn
n1nun
xXn
n1/C12nun
xZb
a/C75
x;tXn
n1nun
tdt
Xn
n1/C12nun
xX1
n1num
x
mZb
aum
tun
tdt;
Xn
n1/C12nun
xX1
n1num
x
mnm;
from which it follows that
Xn
n1nun
xXn
n1/C12nun
xX1
n1nun
x
n:
11:30
Multiplying by ui
xand then integrating with respect to xfrom atob, we obtain
n/C12nn=n;
11:31
which can be solved for nin terms of /C12n:
nn
nÿ/C12n;
11:32
where /C12nis given by
/C12nZb
af
tun
tdt:
11:33
Finally, our solution is given by
u
xf
xX1
n1nun
x
n
f
xX1
n1/C12n
nÿun
x;
11:34
where /C12nis given by Eq. (11.33), and i6.
When for the inhomogeneous equation is equal to one of the eigenvalues, k,
of the kernel, our solution (11.31) blows up. Let us return to Eq. (11.31) and see
what happens to k:
k/C12kkk=k/C12kk:
424SIMPLE LINEAR INTEGRAL EQUATIONS
Clearly, /C12k0, and kis no longer determined by /C12k. But we have, according to
Eq. (11.33),
Zb
af
tuk
tdt/C12k0;
11:35
that is, f
xis orthogonal to the eigenfunction uk
x. Thus if k, the inho-
mogeneous equation has a solution only if f
xis orthogonal to the correspond-
ing eigenfunction uk
x. The general solution of the equation is then
u
xf
xkuk
xkX1
n100Rb
af
tun
tdt
nÿkun
x;
11:36
where the prime on the summation sign means that the term nkis to be omitted
from the sum. In Eq. (11.36) the kremains as an undetermined constant.
/C82elation bet/C119een di/C128erential and integral equations
We have shown how an integral equation can be transformed into a di/C128erential
equation that may be easier to solve than the original integral equation. We now
show how to transform a di/C128erential equation into an integral equation. After we
become familiar with the relation between di/C128erential and integral equations, we
may state the physical problem in either form at will. Let us consider a linear
second-order di/C128erential equation
x00A
tx0B
tx/C103
t;
11:37
with the initial condition
x
ax0;x0
ax0
0:
Integrating Eq. (11.37), we obtain
x0ÿZt
aAx0dtÿZt
aBxdt Zt
a/C103dtC1:
The initial conditions require that C1x0
0. We next integrate the first integral on
the right hand side by parts and obtain
x0ÿAxÿZt
a
BÿA0xdtZt
a/C103dtA
ax0x0
0:
Integrating again, we get
xÿZt
aAxdt ÿZt
aZt
aB
yÿA0
y/C2/C3
x
ydydt
Zt
aZt
a/C103
ydydtA
ax0x0
0/C2/C3
tÿax0:
425DIFFERENTIAL AND INTEGRAL EQUATIONS
Then using the relation
Zt
aZt
af
ydydtZt
a
tÿyf
ydy;
we can rewrite the last equation as
x
tÿZt
aA
y
tÿyB
yÿA0
y/C8/C9 /C2/C3
x
ydy
Zt
a
tÿy/C103
ydyA
ax0x0
0/C2/C3
tÿax0;
11:38
which can be put into the form of a Volterra equation of the second kind
x
tf
tZt
a/C75
t;yx
ydy;
11:39
with
/C75
t;y
yÿtB
yÿA0
y ÿA
y;
11:39a
f
tZt
0
tÿy/C103
ydyA
ax0x0
0
tÿax0:
11:39b
/C85se of integral equations
We have learned how linear integral equations of the more common types may be
solved. We now show some uses of integral equations in physics; that is, we are
going to state some physical problems in integral equation form. In 1823, Abel
made one of the earliest applications of integral equations to a physical problem.Let us take a brief look at this old problem in mechanics.
/C65bel’s integral equation
Consider a particle of mass mfalling along a smooth curve in a vertical plane, the
yzplane, under the influence of gravity, which acts in the negative zdirection.
Conservation of energy gives
1
2m
_z2_y2m/C103z/C69;
where _zdz=dt;and _ydy=dt:If the shape of the curve is given by y/C70
z,w e
can write _y
d/C70=dz_z. Substituting this into the energy conservation equation
and solving for _z, we obtain
_z
2/C69=mÿ2/C103z/C112
1
d/C70=dz2q
/C69=m/C103ÿz/C112
u
z;
11:40
426SIMPLE LINEAR INTEGRAL EQUATIONS
where
u
z
1
d/C70=dz2=2/C103q
:
If_z0a n d zz0att0, then /C69=m/C103z0and Eq. (11.40) becomes
_zz0ÿzp /C14
u
z:
Solving for time t, we obtain
tÿZz0
zu
zz0ÿzp dzZz
z0u
zz
0ÿzp dz;
where zis the height the particle reaches at time t.
/C67lassical simple harmonic oscillator
Consider a linear oscillator
/C127x/C332x0;with x
00; _x
01:
We can transform this di/C128erential equation into an integral equation.
Comparing with Eq. (11.37), we have
A
t0;B
t/C332;and /C103
t0:
Substituting these into Eq. (11.38) (or (11.39), (11.39a), and (11.39b)), we obtain
the integral equation
x
tt/C332Zt
0
yÿtx
ydy;
which is equivalent to the original di/C128erential equation plus the initial conditions.
/C81uantum simple harmonic oscillator
The Schro /C200dinger equation for the energy eigenstates of the one-dimensional
simple harmonic oscillator is
ÿp2
2md2/C32
dx21
2m/C332x2/C32/C69/C32:
11:41
Changing to the dimensionless variable y
m/C33=p/C112
x, Eq. (11.41) reduces to a
simpler form:
d2/C32
dy2
2ÿy2/C320;
11:42
where 2/C69=p/C33/C112
. Taking the Fourier transform of Eq. (11.42), we obtain
d2/C103
k
dk2
2ÿk2/C103
k0;
11:43
427U S EO FI N T E G R A LE Q U A T I O N S
where
/C103
k1
2pZ1
ÿ1/C32
yeikydy
11:44
and we also assume that /C32and/C320vanish as y! 1 .
Eq. (11.43) is formally identical to Eq. (11.42). Since quantities such as the total
probability and the expectation value of the potential energy must be remain finite
for finite E, we should expect /C103
k;d/C103
k=dk!0a sk! 1 . Thus gand/C32di/C128er
at most by a normalization constant
/C103
kc/C32
k:
It follows that /C32satisfies the integral equation
c/C32
k1
2pZ1
ÿ1/C32
yeikydy:
11:45
The constant cmay be determined by substituting c/C32on the right hand side:
c2/C32
k1
2Z1
ÿ1Z1
ÿ1/C32
zeizyeikydzdy
Z1
ÿ1/C32
z
zkdz
/C32
ÿk:
Recall that /C32may be simultaneously chosen to be a parity eigenstate
/C32
ÿx /C32
x. We see that eigenstates of even parity require c21, or
c1; and for eigenstates of odd parity we have c2ÿ1, or ci.
We shall leave the solution of Eq. (11.45), which can be approached in several
ways, as an exercise for the reader.
Problems
11.1 Solve the following integral equations:
(a)u
x1
2ÿxZ1
0u
tdt;
(b)u
xZ1
0u
tdt;
(c)u
xxZ1
0u
tdt:
11.2 Solve the Fredholm equation of the second kind
f
xu
xZb
a/C75
x;tu
tdt;
where f
xcosh x;/C75
x;txt.
428SIMPLE LINEAR INTEGRAL EQUATIONS
11.3 The homogeneous Fredholm equation
u
xZ=2
0sinxsintu
tdt
only has a solution for a particular value of . Find the value of and the
solution corresponding to this value of .
11.4 Solve homogeneous Fredholm equation u
xR1
ÿ1
txu
tdt. Find the
values of and the corresponding solutions.
11.5 Check the convergence of the Neumann series (11.14) by the Cauchy ratio
test.
11.6 Transform the following di/C128erential equations into integral equations:
adx
dtÿx0 with x1 when t0;
bd2x
dt2dx
dtx1 with x0;dx
dt1 when t0:
11.7 By using the Laplace transformation and the convolution theorem solve the
equation
u
xxZx
0sin
xÿtu
tdt:
11.8 Given the Fredholm integral equation
eÿx2Z1
ÿ1eÿ
xÿt2u
tdt;
apply the Fouurier convolution technique to solve it for u
t.
11.9 Find the solution of the Fredholm equation
u
xxZ1
0
xtu
tdt
by the Schmidt–Hilbert method for not equal to an eigenvalue. Show that
there are no solutions when is an eigenvalue.
429PROBLEMS
12
Elements of group theory
Group theory did not find a use in physics until the advent of modern quantum
mechanics in 1925. In recent years group theory has been applied to many
branches of physics and physical chemistry, notably to problems of molecules,
atoms and atomic nuclei. Mostly recently, group theory has been being applied in
the search for a pattern of ‘family’ relationships between elementary particles.
Mathematicians are generally more interested in the abstract theory of groups,
but the representation theory of groups of direct use in a large variety of physical
problems is more useful to physicists. In this chapter, we shall give an elementary
introduction to the theory of groups, which will be needed for understanding the
representation theory.
/C68efinition of a group (group a/C120ioms)
A group is a set of distinct elements for which a law of ‘combination’ is welldefined. Hence, before we give ‘group’ a formal definition, we must first define
what kind of ‘elements’ do we mean. Any collection of objects, quantities or
operators form a set, and each individual object, quantity or operator is calledan element of the set.
A group is a set of elements A/C44 B/C44 /C67 ;...;finite or infinite in number, with a rule
for combining any two of them to form a ‘product’, subject to the following four
conditions:
(1) The product of any two group elements must be a group element; that is, if
AandBare members of the group, then so is the product AB.
(2) The law of composition of the group elements is associative; that is, if A,B,
and/C67are members of the group, then
ABCA
BC.
(3) There exists a unit group element E, called the identity, such that
/C69AA/C69Afor every member of the group.
430
(4) Every element has a unique inverse, Aÿ1, such that AAÿ1Aÿ1A/C69.
The use of the word ‘product’ in the above definition requires comment. The
law of combination is commonly referred as ‘multiplication’, and so the result of a
combination of elements is referred to as a ‘product’. However, the law of com-
bination may be ordinary addition as in the group consisting of the set of all
integers (positive, negative, and zero). Here ABAB, ‘zero’ is the identity, and
Aÿ1
ÿ A. The word ‘product’ is meant to symbolize a broad meaning of
‘multiplication’ in group theory, as will become clearer from the examples below.
A group with a finite number of elements is called a finite group; and the
number of elements (in a finite group) is the order of the group.
A group containing an infinite number of elements is called an infinite group.
An infinite group may be either discrete or continuous. If the number of theelements in an infinite group is denumerably infinite, the group is discrete; if
the number of elements is non-denumerably infinite, the group is continuous.
A group is called Abelian (or commutative) if for every pair of elements A,Bin
the group, ABBA. In general, groups are not Abelian and so it is necessary to
preserve carefully the order of the factors in a group ‘product’.
A subgroup is any subset of the elements of a group that by themselves satisfy
the group axioms with the same law of combination.
Now let us consider some examples of groups.
Example 12.1
The real numbers 1 and ÿ1 form a group of order two, under multiplication. The
identity element is 1; and the inverse is 1 =x, where xstands for 1 or ÿ1.
Example 12.2The set of all integers (positive, negative, and zero) forms a discrete infinite group
under addition. The identity element is zero; the inverse of each element is its
negative. The group axioms are satisfied:
(1) is satisfied because the sum of any two integers (including any integer with
itself) is always another integer.
(2) is satisfied because the associative law of addition A
BC
ABCis true for integers.
(3) is satisfied because the addition of 0 to any integer does not alter it.(4) is satisfied because the addition of the inverse of an integer to the integer
itself always gives 0, the identity element of our group: A
ÿ A0.
Obviously, the group is Abelian since ABBA. We denote this group by
S
1.
431DEFINITION OF A GROUP (GROUP A/C88IOMS)
The same set of all integers does not form a group under multiplication. Why/C63
Because the inverses of integers are not integers and so they are not members of
the set.
Example 12.3
The set of all rational numbers ( /C112=/C113, with /C11360) forms a continuous infinite group
under addition. It is an Abelian group, and we denote it by S2. The identity
element is 0; and the inverse of a given element is its negative.
Example 12.4
The set of all complex numbers
zxiyforms an infinite group under
addition. It is an Abelian group and we denote it by S3. The identity element
is 0; and the inverse of a given element is its negative (that is, ÿzis the inverse
ofz).
The set of elements in S1is a subset of elements in S2, and the set of elements in
S2is a subset of elements in S3. Furthermore, each of these sets forms a group
under addition, thus S1is a subgroup of S2,a n d S2a subgroup of S3. Obviously
S1is also a subgroup of S3.
Example 12.5The three matrices
~A10
01
; ~B01
ÿ1ÿ1
; ~Cÿ1ÿ1
10
form an Abelian group of order three under matrix multiplication. The identity
element is the unit matrix, /C69~A. The inverse of a given matrix is the inverse
matrix of the given matrix:
~A
ÿ110
01
~A; ~Bÿ1ÿ1ÿ1
10
~C; ~Cÿ101
ÿ1ÿ1
~B:
It is straightforward to check that all the four group axioms are satisfied. We
leave this to the reader.
Example 12.6
The three permutation operations on three objects a;b;c
123;231;312
form an Abelian group of order three with sequential performance as the law ofcombination.
The operation /C911 2 3/C93 means we put the object afirst, object bsecond, and object
cthird. And two elements are multiplied by performing first the operation on the
432ELEMENTS OF GROUP THEORY
right, then the operation on the left. For example
231312abc231cababc:
Thus two operations performed sequentially are equivalent to the operation
/C911 2 3/C93:
231312123:
similarly
312231abc312bcaabc;
that is,
312231123:
This law of combination is commutative. What is the identity element of thisgroup/C63 And the inverse of a given element/C63 We leave the reader to answer these
questions. The group illustrated by this example is known as a cyclic group of
order 3, C
3.
It can be shown that the set of all permutations of three objects
123;231;312;132;321;213
forms a non-Abelian group of order six denoted by S3. It is called the symmetric
group of three objects. Note that C3is a subgroup of S3.
/C67/C121clic groups
We now revisit the cyclic groups. The elements of a cyclic group can be expressed
as power of a single element A, say, as A;A2;A3;...;A/C112ÿ1;A/C112/C69;pis the smal-
lest integer for which A/C112/C69and is the order of the group. The inverse of Akis
A/C112ÿk, that is, an element of the set. It is straightforward to check that all group
axioms are satisfied. We leave this to the reader. It is obvious that cyclic groups
are Abelian since AkAAAk
k</C112.
Example 12.7The complex numbers 1, i;ÿ1;ÿiform a cyclic group of order 3. In this case,
Aiand/C1123:i
n,n0;1;2;3. These group elements may be interpreted as
successive 90 8rotations in the complex plane
0; =2; ;and 3 =2. Con-
sequently, they can be represented by four 2 2 matrices. We shall come back
to this later.
Example 12.8
We now consider a second example of cyclic groups: the group of rotations of an
equilateral triangle in its plane about an axis passing through its center that brings
433CYCLIC GROUPS
it onto itself. This group contains three elements (see Fig. 12.1):
/C69
08/C58 the identity; triangle is left alone;
A
120 8/C58the triangle is rotated through 120 8counterclockwise, which
sends Pto/C81,/C81to/C82, and /C82toP;
B
240 8/C58the triangle is rotated through 240 8counterclockwise, which
sends Pto/C82,/C82to/C81, and /C81toP;
C
360 8/C58the triangle is rotated through 360 8counterclockwise, which
sends Pback to P,/C81back to /C81and/C82back to /C82.
Notice that C/C69. Thus there are only three elements represented by E,A, and B.
This set forms a group of order three under addition. The reader can check that
all four group axioms are satisfied. It is also obvious that operation Bis equiva-
lent to performing operation Atwice (240 8120 8120 8, and the operation /C67
corresponds to performing Athree times. Thus the elements of the group may be
expressed as the power of the single element AasE,A,A2,A3
/C69): that is, it is a
cyclic group of order three, and is generated by the element A.
The cyclic group considered in Example 12.8 is a special case of groups of
transformations (rotations, reflection, translations, permutations, etc.), thegroups of particular interest to physicists. A transformation that leaves a physical
system invariant is called a symmetry transformation of the system. The set of all
symmetry transformations of a system is a group, as illustrated by this example.
/C71roup multiplication table
A group of order nhasn
2products. Once the products of all ordered pairs of
elements are specified the structure of a group is uniquely determined. It is some-
times convenient to arrange these products in a square array called a group multi-
plication table. Such a table is indicated schematically in Table 12.1. The element
that appears at the intersection of the row labeled Aand the column labeled Bis
the product AB, (in the table A2means AA, etc). It should be noted that all the
434ELEMENTS OF GROUP THEORY
Figure 12.1.
elements in each row or column of the group multiplication must be distinct: that
is, each element appears once and only once in each row or column. This can be
proved easily: if the same element appeared twice in a given row, the row labeledAsay, then there would be two distinct elements /C67and/C68such that ACAD.I f
we multiply the equation by A
ÿ1on the left, then we would have Aÿ1AC
Aÿ1AD,o r /C69C/C69D. This cannot be true unless CD, in contradiction to
our hypothesis that /C67and/C68are distinct. Similarly, we can prove that all the
elements in any column must be distinct.
As a simple practice, consider the group C3of Example 12.6 and label the
elements as follows
123!/C69;231!X;312!Y:
If we label the columns of the table with the elements E,X,/C89and the rows with
their respective inverses, E,Xÿ1,Yÿ1, the group multiplication table then takes
the form shown in Table 12.2.
Isomorphic groups
Two groups are isomorphic to each other if the elements of one group can be put inone-to-one correspondence with the elements of the other so that the corresponding
elements multiply in the same way. Thus if the elements A;B;C;...of the group /C71
435ISOMORPHIC GROUPS
Table 12.1. Group multiplication table
EA B /C67 ...
EE A B /C67 ...
AA A2AB A/C67 ...
BB B A B2B/C67 ...
/C67 /C67 /C67A /C67B C2...
...............
Table 12.2.
EX /C89
EE X /C89
Xÿ1Xÿ1E Xÿ1Y
Yÿ1Yÿ1Yÿ1XE
correspond respectively to the elements A0;B0;C0;...of/C710, then the equation
ABCimplies that A0B0C0, etc., and vice versa. Two isomorphic groups
have the same multiplication tables except for the labels attached to the group
elements. Obviously, two isomorphic groups must have the same order.
Groups that are isomorphic and so have the same multiplication table are the
same or identical, from an abstract point of view. That is why the concept of
isomorphism is a key concept to physicists. Diverse groups of operators that act
on diverse sets of objects have the same multiplication table; there is only one
abstract group. This is where the value and beauty of the group theoretical
method lie; the same abstract algebraic results may be applied in making predic-
tions about a wide variety physical objects.
The isomorphism of groups is a special instance of homomorphism, which
allows many-to-one correspondence.
Example 12.9
Consider the groups of Problems 12.2 and 12.4. The group /C71of Problem 12.2
consists of the four elements /C691;Ai;Bÿ1;Cÿiwith ordinary multi/C45
plication as the rule of combination. The group multiplication table has the form
shown in Table 12.3. The group /C710of Problem 12.4 consists of the following four
elements, with matrix multiplication as the rule of combination
/C69010
01
;A001
ÿ10
;B0ÿ100ÿ1
;C
00ÿ1
10
:
It is straightforward to check that the group multiplication table of group /C710has
the form of Table 12.4. Comparing Tables 12.3 and 12.4 we can see that they have
precisely the same structure. The two groups are therefore isomorphic.
Example 12.10
We stated earlier that diverse groups of operators that act on diverse sets of
objects have the same multiplication table; there is only one abstract group. To
illustrate this, we consider, for simplicity, an abstract group of order two, /C712: that
436ELEMENTS OF GROUP THEORY
Table 12.3.
1 iÿ1iE A B /C67
11 iÿ1 iE E A B /C67
ii ÿ1ÿi 1o r AA B /C67 E
ÿ1 ÿ1ÿi 1 iB B /C67 E A
ÿi ÿi 1 iÿ1 /C67/C67 E A B
is, we make no a priori assumption about the significance of the two elements of
our group. One of them must be the identity E, and we call the other X. Thus we
have
/C692/C69;/C69XX/C69/C69:
Since each element appears once and only once in each row and column, the
group multiplication table takes the form:
We next consider some groups of operators that are isomorphic to /C712. First,
consider the following two transformations of three-dimensional space into itself:
(1) the transformation /C690, which leaves each point in its place, and
(2) the transformation /C82, which maps the point
x;y;zinto the point
ÿx;ÿy;ÿz. Evidently, R2RR(the transformation Rfollowed by R)
will bring each point back to its original position. Thus we have
/C69
02/C690,R/C690/C690RR/C690R;R2/C690; and the group multiplication
table has the same form as /C712: that is, the group formed by the set of the two
operations /C690andRis isomorphic to /C712.
We now associate with the two operations /C690andRtwo operators ^O/C690and ^OR,
which act on real- or complex-valued functions of the spatial coordinates
x;y;z,
/C32
x;y;z, with the following e/C128ects:
^O/C690/C32
x;y;z/C32
x;y;z; ^OR/C32
x;y;z/C32
ÿx;ÿy;ÿz:
From these we see that
^O/C6902 ^O/C690; ^O/C690^OR ^OR^O/C690 ^OR;
^OR2 ^OR:
437ISOMORPHIC GROUPS
Table 12.4.
/C690A0B0C0
/C690/C690A0B0C0
A0A0B0C0/C690
B0B0C0/C690A0
C0C0/C690/C690B0
/C69X
/C69/C69 X
XX /C69
Obviously these two operators form a group that is isomorphic to /C712. These two
groups (formed by the elements /C690,R, and the elements ^O/C690and ^OR, respectively)
are the two representations of the abstract group /C712. These two simple examples
cannot illustrate the value and beauty of the group theoretical method, but they
do serve to illustrate the key concept of isomorphism.
/C71roup of permutations and /C67a/C121le/C121/C39s theorem
In Example 12.6 we examined briefly the group of permutations of three objects.
We now come back to the general case of nobjects (1 ;2;...;n) placed in nboxes
(or places) labeled 1,2;...;n. This group, denoted by Sn, is called the sym-
metric group on nobjects. It is of order n/C33 How do we know/C63 The first object may
be put in any of nboxes, and the second object may then be put in any of nÿ1
boxes, and so forth: n
nÿ1
nÿ2 321n/C33:
We now define, following common practice, a permutation symbol P
P123 n
123 n/C32!
;
12:1
which shifts the object in box 1 to box 1, the object in box 2 to box 2, and so
forth, where 12nis some arrangement of the numbers 1 ;2;3;...;n. The old
notation in Example 12.6 can now be written as
231123
231
:
Fornobjects there are n/C33permutations or arrangements, each of which may be
written in the form (12.1). Taking a specific example of three objects, we have
P1123
123
;P2123231
;P
3123132
;
P
4123
213
;P5123321
;P
6123312
:
For the product of two permutations P
iPj
i;j1;2;...;6, we first perform the
one on the right, Pj, and then the one on the left, Pi. Thus
P3P6123
132123312
123213
P
4:
To the reader who has diculty seeing this result, let us explain. Consider the first
column. We first perform P6, so that 1 is replaced by 3, we then perform P3and 3
438ELEMENTS OF GROUP THEORY
is replaced by 2. So by the combined action 1 is replaced by 2 and we have the first
column
1
2
:
We leave the other two columns to be completed by the reader.
Each element of a group has an inverse. Thus, for each permutation Pithere is
Pÿ1
i, the inverse of Pi. We can use the property PiPÿ1
iP1to find Pÿ1
i. Let us find
Pÿ1
6:
Pÿ1
6312
123
123231
P
2:
It is straightforward to check that
P6Pÿ1
6P6P2123312123231
123123
P
1:
The reader can verify that our group S3is generated by the elements P2andP3,
while P1serves as the identity. This means that the other three distinct elements
can be expressed as distinct multiplicative combinations of P2andP3:
P4P2
2P3;P5P2P3;P6P22:
The symmetric group Snplays an important role in the study of finite groups.
Every finite group of order nis isomorphic to a subgroup of the permutation
group Sn. This is known as Cayley’s theorem. For a proof of this theorem the
interested reader is referred to an advanced text on group theory.
In physics, these permutation groups are of considerable importance in the
quantum mechanics of identical particles, where, if we interchange any two or
more these particles, the resulting configuration is indistinguishable from the
original one. Various quantities must be invariant under interchange or permuta-
tion of the particles. Details of the consequences of this invariant property may befound in most first-year graduate textbooks on quantum mechanics that cover the
application of group theory to quantum mechanics.
/C83ubgroups and cosets
A subset of a group /C71, which is itself a group, is called a subgroup of /C71. This idea
was introduced earlier. And we also saw that C
3, a cyclic group of order 3, is a
subgroup of S3, a symmetric group of order 6. We note that the order of C3is a
factor of the order of S3. In fact, we will show that, in general,
the order of a subgroup is a factor of the order of the full group(that is, the group from which the subgroup is derived).
439SUBGROUPS AND COSETS
This can be proved as follows. Let /C71be a group of order nwith elements /C1031
/C69,
/C1032;...;/C103n:Let/C72, of order m, be a subgroup of /C71with elements /C1041
/C69,
/C1042;...;/C104m. Now form the set /C103/C104k
0km, where gis any element of /C71not
in/C72. This collection of elements is called the left-coset of /C72with respect to g(the
left-coset, because gis at the left of /C104k).
If such an element gdoes not exist, then H/C71, and the theorem holds trivially.
Ifgdoes exist, than the elements /C103/C104kare all di/C128erent. Otherwise, we would have
/C103/C104k/C103/C104‘,o r /C104k/C104‘, which contradicts our assumption that /C72is a group.
Moreover, the elements /C103/C104kare not elements of /C72. Otherwise, /C103/C104k/C104j, and we
have
/C103/C104j=/C104k:
This implies that gis an element of /C72, which contradicts our assumption that g
does not belong to /C72.
This left-coset of /C72does not form a group because it does not contain the
identity element ( /C1031/C1041/C69. If it did form a group, it would require for some /C104j
such that /C103/C104j/C69or, equivalently, /C103/C104ÿ1
j. This requires gto be an element of /C72.
Again this is contrary to assumption that gdoes not belong to /C72.
Now every element gin/C71but not in /C72belongs to some coset g/C72. Thus /C71is a
union of /C72and a number of non-overlapping cosets, each having mdi/C128erent
elements. The order of /C71is therefore divisible by m. This proves that the order
of a subgroup is a factor of the order of the full group. The ratio n/C47mis the index
of/C72in/C71.
It is straightforward to prove that a group of order p, where pis a prime
number, has no subgroup. It could be a cyclic group generated by an element a
of period p.
/C67on/C106ugate classes and in/C118ariant subgroups
Another way of dividing a group into subsets is to use the concept of classes. Let
a,b, and ube any three elements of a group, and if
buÿ1au;
bis said to be the transform of aby the element u;aandbare conjugate (or
equivalent) to each other. It is straightforward to prove that conjugate has thefollowing three properties:
(1) Every element is conjugate with itself (reflexivity). Allowing uto be the
identity element E, then we have a/C69
ÿ1a/C69:
(2) If ais conjugate to b, then bis conjugate to a(symmetry). If auÿ1bu, then
buauÿ1
uÿ1ÿ1a
uÿ1, where uÿ1is an element of /C71ifuis.
440ELEMENTS OF GROUP THEORY
(3) If ais conjugate with both bandc, then bandcare conjugate with each
other (transitivity). If auÿ1buand b/C118ÿ1c/C118, then auÿ1/C118ÿ1c/C118u
/C118uÿ1c
/C118u, where uand/C118belong to /C71so that /C118uis also an element of /C71.
We now divide our group up into subsets, such that all elements in any subset
are conjugate to each other. These subsets are called classes of our group.
Example 12.11
The symmetric group S3has the following six distinct elements:
P1/C69;P2;P3;P4P2
2P3;P5P2P3;P6P22;
which can be separated into three conjugate classes:
fP1g;fP2;P6g;fP3;P4;P5g:
We now state some simple facts about classes without proofs:
(a) The identity element always forms a class by itself.
(b) Each element of an Abelian group forms a class by itself.
(c) All elements of a class have the same period.
Starting from a subgroup /C72of a group /C71, we can form a set of elements u/C104ÿ1u
for each ubelong to /C71. This set of elements can be seen to be itself a group. It is a
subgroup of /C71and is isomorphic to /C72. It is said to be a conjugate subgroup to /C72
in/C71. It may happen, for some subgroup /C72, that for all ubelonging to /C71, the sets
/C72andu/C104uÿ1are identical. /C72is then an invariant or self-conjugate subgroup of /C71.
Example 12.12
Let us revisit S3of Example 12.11, taking it as our group /C71S3. Consider the
subgroup HC3fP1;P2;P6g:The following relation holds
P2P1
P2
P60
BB@1
CCAPÿ1
2P1P2P1
P2
P60
BB@1
CCAPÿ1
2Pÿ1
1
P1P2P1
P2
P60
BB@1
CCA
P1P2ÿ1
P2
1P2P1
P2
P60
BB@1
CCA
P
2
1P2ÿ1P1
P6
P20
BB@1
CCA:
Hence HC
3fP1;P2;P6gis an invariant subgroup of S3.
441CONJUGATE CLASSES AND INVARIANT SUBGROUPS
/C71roup representations
In previous sections we have seen some examples of groups which are isomorphic
with matrix groups. Physicists have found that the representation of group
elements by matrices is a very powerful technique. It is beyond the scope of
this text to make a full study of the representation of groups; in this section we
shall make a brief study of this important subject of the matrix representations of
groups.
If to every element of a group /C71,/C1031;/C1032;/C1033;...;we can associate a non-singular
square matrix D
/C1031;D
/C1032;D
/C1033;...;in such a way that
/C103i/C103j/C103k implies D
/C103iD
/C103jD
/C103k;
12:2
then these matrices themselves form a group /C710, which is either isomorphic or
homomorphic to /C71. The set of such non-singular square matrices is called a
representation of group /C71. If the matrices are nn, we have an n-dimensional
representation; that is, the order of the matrix is the dimension (or order) of therepresentation D
n. One trivial example of such a representation is the unit matrix
associated with every element of the group. As shown in Example 12.9, the four
matrices of Problem 12.4 form a two-dimensional representation of the group /C71
of Problem 12.2.
If there is one-to-one correspondence between each element of /C71and the matrix
representation group /C710, the two groups are isomorphic, and the representation is
said to be faithful (or true). If one matrix /C68represents more than one group
element of /C71, the group /C71is homomorphic to the matrix representation group
/C710and the representation is said to be unfaithful.
Now suppose a representation of a group /C71has been found which consists of
matrices DD
/C1031;D
/C1032;D
/C1033;...;D
/C103/C112, each matrix being of dimension
n. We can form another representation D0by a similarity transformation
D0
/C103Sÿ1D
/C103S;
12:3
Sbeing a non-singular matrix, then
D0
/C103iD0
/C103jSÿ1D
/C103iSSÿ1D
/C103jS
Sÿ1D
/C103iD
/C103jS
Sÿ1D
/C103i/C103jS
D0
/C103i/C103j:
In general, representations related in this way by a similarity transformation are
regarded as being equivalent. However, the forms of the individual matrices in the
two equivalent representations will be quite di/C128erent. With this freedom in the
442ELEMENTS OF GROUP THEORY
choice of the forms of the matrices it is important to look for some quantity that is
an invariant for a given transformation. This is found in considering the traces ofthe matrices of the representation group because the trace of a matrix is
invariant under a similarity transformation. It is often possible to bring, by a
similarity transformation, each matrix in the representation group into a diagonal
form
S
ÿ1DSD
10
0D
2/C32!
;
12:4
where D
1is of order m;m<nandD
2is of order nÿm. Under these conditions,
the original representation is said to be reducible to D
1andD
2. We may write
this result as
DD
1/C8D
2
12:5
and say that /C68has been decomposed into the two smaller representation D
1and
D
2;/C68is often called the direct sum of D
1andD
2.
A representation D
/C103is called irreducible if it is not of the form (12.4)
and cannot be put into this form by a similarity transformation. Irreducible
representations are the simplest representations, all others may be built up from
them, that is, they play the role of ‘building blocks’ for the study of group
representation.
In general, a given group has many representations, and it is always possible to
find a unitary representation – one whose matrices are unitary. Unitary matricescan be diagonalized, and the eigenvalues can serve for the description or classi-
fication of quantum states. Hence unitary representations play an especially
important role in quantum mechanics.
The task of finding all the irreducible representations of a group is usually very
laborious. Fortunately, for most physical applications, it is sucient to know only
the traces of the matrices forming the representation, for the trace of a matrix is
invariant under a similarity transformation. Thus, the trace can be used to iden-
tify or characterize our representation, and so it is called the character in group
theory. A further simplification is provided by the fact that the character of every
element in a class is identical, since elements in the same class are related to each
other by a similarity transformation. If we know all the characters of one elementfrom every class of the group, we have all of the information concerning the group
that is usually needed. Hence characters play an important part in the theory of
group representations. However, this topic and others related to whether a given
representation of a group can be reduced to one of smaller dimensions are beyond
the scope of this book. There are several important theorems of representation
theory, which we now state without proof.
443GROUP REPRESENTATIONS
(1) A matrix that commutes with all matrices of an irreducible representation of
a group is a multiple of the unit matrix (perhaps null). That is, if matrix A
commutes with D
/C103which is irreducible,
D
/C103AAD
/C103
for all gin our group, then Ais a multiple of the unit matrix.
(2) A representation of a group is irreducible if and only if the only matrices to
commute with all matrices are multiple of the unit matrix.
Both theorems (1) and (2) are corollaries of Schur’s lemma.
(3) Schur’s lemma: Let D
1andD
2be two irreducible representations of (a
group /C71) dimensionality nandn0, if there exists a matrix Asuch that
AD
1
/C103D
2
/C103A for all /C103in the group /C71
then for n6n0,A0; for nn0, either A0o rAis a non-singular matrix
andD
1andD
2are equivalent representations under the similarity trans-
formation generated by A.
(4) Orthogonality theorem: If /C71is a group of order handD
1andD
2are any
two inequivalent irreducible (unitary) representations, of dimensions d1and
d2, respectively, then
X
/C103D
i
/C12
/C103/C42D
j
/C13
/C103/C104
d1ij/C13/C12;
where D
i
/C103is a matrix, and D
i
/C12
/C103is a typical matrix element. The sum
runs over all gin/C71.
/C83ome special groups
Many physical systems possess symmetry properties that always lead to certain
quantity being invariant. For example, translational symmetry (or spatial homo-
geneity) leads to the conservation of linear momentum for a closed system, and
rotational symmetry (or isotropy of space) leads to the conservation of angular
momentum. Group theory is most appropriate for the study of symmetry. In this
section we consider the geometrical symmetries. This provides more illustrations
of the group concepts and leads to some special groups.
Let us first review some symmetry operations. A plane of symmetry is a plane in
the system such that each point on one side of the plane is the mirror image of acorresponding point on the other side. If the system takes up an identical position on
rotation through a certain angle about an axis, that axis is called an axis of sym-
metry. A center of inversion is a point such that the system is invariant under the
operation r!ÿ r, where ris the position vector of any point in the system referred to
the inversion center. If the system takes up an identical position after a rotationfollowed by an inversion, the system possesses a rotation–inversion center.
444ELEMENTS OF GROUP THEORY
Some symmetry operations are equivalent. As shown in Fig. 12.2, a two-fold
inversion axis is equivalent to a mirror plane perpendicular to the axis.
There are two di/C128erent ways of looking at a rotation, as shown in Fig. 12.3.
According to the so-called active view, the system (the body) undergoes a rotation
through an angle , say, in the clockwise direction about the x3-axis. In the passive
view, this is equivalent to a rotation of the coordinate system through the sameangle but in the counterclockwise sense. The relation between the new and old
coordinates of any point in the body is the same in both cases:
x
0
1x1cosx2sin;
x0
2ÿx1sinx2cos;
x0
3x3;9
>>=
>>;
12:6
where the prime quantities represent the new coordinates.
A general rotation, reflection, or inversion can be represented by a linear
transformation of the form
x0
111x112x213x3;
x0
221x122x223x3;
x0
331x132x233x3:9
>>=
>>;
12:7
445SOME SPECIAL GROUPS
Figure 12.2.
Figure 12.3. ( a) Active view of rotation; ( b) passive view of rotation.
Equation (12.7) can be written in matrix form
~x0~~x
12:8
with
~111213
212223
3132330
B@1
CA; ~xx1
x2
x30
B@1
CA; ~x0x0
1
x0
2
x0
30
B@1
CA:
The matrix ~is an orthogonal matrix and the value of its determinant is 1. The
‘ÿ1’ value corresponds to an operation involving an odd number of reflections.
For Eq. (12.6) the matrix ~has the form
~cossin0
ÿsincos0
00 10
B@1
CA:
12:6a
For a rotation, an inversion about an axis, or a reflection in a plane through the
origin, the distance of a point from the origin remains unchanged:
r2x2
1x22x23x0
12x0
22x0
32:
12:9
/C84he symmetry group /C682;/C683
Let us now examine two simple examples of symmetry and groups. The first one is
on twofold symmetry axes. Our system consists of six particles: two identicalparticles Alocated at aon the x-axis, two particles, Batbon the y-axis,
and two particles /C67atcon the z-axis. These particles could be the atoms of a
molecule or part of a crystal. Each axis is a twofold symmetry axis. Clearly, the
identity or unit operator (no rotation) will leave the system unchanged. What
rotations can be carried out that will leave our system invariant/C63 A certain com-
bination of rotations of radians about the three coordinate axes will do it. The
orthogonal matrices that represent rotations about the three coordinate axes canbe set up in a similar manner as was done for Eq. (12.6a), and they are
~
100
0ÿ10
00 ÿ10
B@1
CA; ~/C12
ÿ10 0
01 0
00 ÿ10
B@1
CA; ~/C13
ÿ10 0
0ÿ10
00 10
B@1
CA;
where ~is the rotational matrix about the x-axis, and ~/C12and ~/C13are the rotational
matrices about y-a n d z-axes, respectively. Of course, the identity operator is a
unit matrix
446ELEMENTS OF GROUP THEORY
~/C69100
010
0010
B@1
CA:
These four elements form an Abelian group with the group multiplication table
shown in Table 12.5. It is easy to check this group table by matrix multiplication.
Or you can check it by analyzing the operations themselves, a tedious task. This
demonstrates the power of mathematics: when the system becomes too complexfor a direct physical interpretation, the usefulness of mathematics shows.
This symmetry group is usually labeled D
2, a dihedral group with a twofold
symmetry axis. A dihedral group Dnwith an n-fold symmetry axis has naxes with
an angular separation of 2 =nradians and is very useful in crystallographic study.
We next consider an example of threefold symmetry axes. To this end, let us re-
visit Example 12.8. Rotations of the triangle of 0 8, 120 8, 240 8, and 360 8leave the
triangle invariant. Rotation of the triangle of 0 8means no rotation, the triangle is
left unchanged; this is represented by a unit matrix (the identity element). Theother two orthogonal rotational matrices can be set up easily:
~AR
z
120 8ÿ1=2ÿ
3p
=2
3p
=2ÿ1=20
@1A;
~BR
z
240 8ÿ1=2
3p
=2
ÿ3p
=2ÿ1=20
@1A;
and
~/C69R
z
010
01
:
We notice that ~CRz
360 8 ~/C69. The set of the three elements
~/C69;~A;~Bforms a
cyclic group C3with the group multiplication table shown in Table 12.6. The z-
447SOME SPECIAL GROUPS
Table 12.5.
~/C69 ~ ~/C12 ~/C13
~/C69 ~/C69 ~ ~/C12 ~/C13
~ ~ ~/C69 ~/C13 ~/C12
~/C12 ~/C12 ~/C13 ~/C69 ~
~/C13 ~/C13 ~/C12 ~ ~/C69
axis is a threefold symmetry axis. There are three additional axes of symmetry in
thexyplane: each corner and the geometric center Odefining an axis; each of
these is a twofold symmetry axis (Fig. 12.4). Now let us consider reflection opera-
tions. The following successive operations will bring the equilateral angle onto
itself (that is, be invariant):
~/C69the identity; triangle is left unchanged;
~Atriangle is rotated through 120 8clockwise;
~Btriangle is rotated through 240 8clockwise;
~Ctriangle is reflected about axis O/C82(or the y-axis);
~Dtriangle is reflected about axis O/C81;
~/C70triangle is reflected about axis OP.
Now the reflection about axis O/C82is just a rotation of 180 8about axis O/C82, thus
~CROR
180 8ÿ10
01
:
Next, we notice that reflection about axis O/C81is equivalent to a rotation of 240 8
about the z-axis followed by a reflection of the x-axis (Fig. 12.5):
~DRO/C81
180 8 ~C~Bÿ10
01ÿ1=2
3p
=2
ÿ3p
=2ÿ1=2/C32!
1=2ÿ3p
=2
ÿ3p
=2ÿ1=2/C32!
:
448ELEMENTS OF GROUP THEORY
Table 12.6.
~/C69 ~A ~B
~/C69 ~/C69 ~A ~B
~A ~A ~B ~/C69
B ~B ~/C69 ~A
Figure 12.4.
Similarly, reflection about axis OPis equivalent to a rotation of 180 8followed by
a reflection of the x-axis:
~/C70ROP
180 8 ~C~A1=2
3p
=23p
=2ÿ1=2/C32!
:
The group multiplication table is shown in Table 12.7. We have constructed a six-
element non-Abelian group and a 2 2 irreducible matrix representation of it.
Our group is known as D
3in crystallography, the dihedral group with a threefold
axis of symmetry.
One-dimensional unitrary group U
1
We now consider groups with an infinite number of elements. The group element
will contain one or more parameters that vary continuously over some range so
they are also known as continuous groups. In Example 12.7, we saw that the
complex numbers (1 ;i;ÿ1;ÿiform a cyclic group of order 3. These group ele-
ments may be interpreted as successive 90 8rotations in the complex plane
0; =2; ;3=2, and so they may be written as ei’with ’0,=2,,3=2. If
’is allowed to vary continuously over the range 0;2, then we will have, instead
of a four-member cyclic group, a continuous group with multiplication for the
composition rule. It is straightforward to check that the four group axioms are all
449SOME SPECIAL GROUPS
Table 12.7.
~/C69 ~A ~B ~C ~D ~/C70
~/C69 ~/C69 ~A ~B ~C ~D ~/C70
~A ~A ~B ~/C69 ~D ~/C70 ~C
~B ~B ~/C69 ~A ~/C70 ~C ~D
~C ~C ~/C70 ~D ~/C69 ~B ~A
~D ~D ~C ~/C70 ~A ~/C69 ~B
~/C70 ~/C70 ~D ~C ~B ~A ~/C69
Figure 12.5.
met. In quantum mechanics, ei’is a complex phase factor of a wave function,
which we denote by U
’). Obviously, U
0is an identity element. Next,
U
’U
’0ei
’’0U
’’0;
andU
’’0is an element of the group. There is an inverse: Uÿ1
’U
ÿ’,
since
U
’U
ÿ’U
ÿ’U
’U
0/C69
for any ’. The associative law is satisfied:
U
’1U
’2U
’3ei
’1’2ei’3ei
’1’2’3ei’1ei
’2’3
U
’1U
’2U
’3:
This group is a one-dimensional unitary group; it is called U
1. Each element is
characterized by a continuous parameter ’,0’2;’can take on an infinite
number of values. Moreover, the elements are di/C128erentiable:
dUU
’d’ÿU
’ei
’d’ÿei’
ei’
1id’ÿei’iei’d’iUd’
or
dU=d’iU:
Infinite groups whose elements are di/C128erentiable functions of their parameters
are called Lie groups. The di/C128erentiability of group elements allows us to develop
the concept of the generator. Furthermore, instead of studying the whole group,
we can study the group elements in the neighborhood of the identity element.
Thus Lie groups are of particular interest. Let us take a brief look at a few more
Lie groups.
Orthogonal groups SO
2andSO
3
The rotations in an n-dimensional Euclidean space form a group, called O
n. The
group elements can be represented by nnorthogonal matrices, each with
n
nÿ1=2 independent elements (Problem 12.12). If the determinant of Ois set
to be 1 (rotation only, no reflection), then the group is often labeled SO
n. The
label O
nis also often used.
The elements of SO
2are familiar; they are the rotations in a plane, say the xy
plane:
x0
y0/C32!
~Rx
y/C32!
cossin
ÿsincos x
y
:
450ELEMENTS OF GROUP THEORY
This group has one parameter: the angle . As we stated earlier, groups enter
physics because we can carry out transformations on physical systems and the
physical systems often are invariant under the transformations. Here x2y2is
left invariant.
We now introduce the concept of a generator and show that rotations of SO
2
are generated by a special 2 2 matrix ~2, where
~20ÿi
i0
:
Using the Euler identity, eicosisin, we can express the 2 2 rotation
matrices R
in exponential form:
~R
cossin
ÿsincos
~I2cosi~2sinei~2;
where ~I2is a 2 2 unit matrix. From the exponential form we see that multi-
plication is equivalent to addition of the arguments. The rotations close to theidentity element have small angles /C1290:We call ~
2the generator of rotations for
SO
2.
It has been shown that any element gof a Lie group can be written in the form
/C103
1;2;...;nexpX
i1ii/C70i/C32!
:
For nparameters there are nof the quantities /C70i, and they are called the
generators of the Lie group.
Note that we can get ~2from the rotation matrix ~R
by di/C128erentiation at the
identity of SO
2, that is, /C1290:This suggests that we may find the generators of
other groups in a similar manner.
Forn3 there are three independent parameters, and the set of 3 3 ortho-
gonal matrices with determinant 1 also forms a group, the SO
3, its general
member may be expressed in terms of the Euler angle rotation
R
; /C12; /C13 Rz0
0;0;Ry
0;/C12 ;0Rz
0;0;/C13;
where Rzis a rotation about the z-axis by an angle /C13,Rya rotation about the y-
axis by an angle /C12, and Rz0a rotation about the z0-axis (the new z-axis) by an
angle . This sequence can perform a general rotation. The separate rotations can
be written as
~Ry
/C12cos/C120ÿsin/C12
01 0
sin/C120 cos /C120
B@1
CA; ~Rz
/C13cos/C13 sin/C130
ÿsin/C13cos/C130
0 /C111 10
B@1
CA;
451SOME SPECIAL GROUPS
~Rx
10 0
0c o s sin
0ÿsincos0
B@1
CA:
TheSO
3rotations leave x2y2z2invariant.
The rotations Rz
/C13form a group, called the group Rz, which is an Abelian sub-
group of SO
3. To find the generator of this group, let us take the following
di/C128erentiation
ÿid ~Rz
/C13=d/C13/C130/C12/C120ÿi0
i00
0000
B@1
CA~Sz;
where the insertion of iis to make ~SzHermitian. The rotation Rz
/C13through an
infinitesimal angle /C13can be written in terms of ~Sz:
Rz
/C13 ~I3dRz
/C13
d/C13/C12/C12/C12/C12
/C130/C13O
/C132 ~I3i/C13~Sz:
A finite rotation R
/C13may be constructed from successive infinitesimal rotations
Rz
/C131/C132
~I3i/C131~Sz
~I3i/C132~Sz:
Now let
/C13/C13=/C78for/C78rotations, with /C78!1 , then
Rz
/C13 lim
/C78!1~I3
i/C13=/C78~Sz/C2/C3 /C78exp
i~Sz;
which identifies ~Szas the generator of the rotation group Rz. Similarly, we can
find the generators of the subgroups of rotations about the x-axis and the y-axis.
/C84heSU
ngroups
Thennunitary matrices ~Ualso form a group, the U
ngroup. If there is the
additional restriction that the determinant of the matrices be 1, we have the
special unitary or unitary unimodular group, SU
n. Each nnunitary matrix
hasn2ÿ1 independent parameters (Problem 12.14).
Forn2w eh a v e SU
2and possible ways to parameterize the matrix Uare
~Uab
ÿb/C42a/C42
;
where a,bare arbitrary complex numbers and jaj2jbj21. These parameters
are often called the Cayley–Klein parameters, and were first introduced by Cayley
and Klein in connection with problems of rotation in classical mechanics.
Now let us write our unitary matrix in exponential form:
~Uei~H;
452ELEMENTS OF GROUP THEORY
where ~His a Hermitian matrix. It is easy to show that ei~His unitary:
ei~H
ei~Heÿi~Hei~Hei
~Hÿ~H1:
This implies that any nnunitary matrix can be written in exponential form with
a particularly selected set of n2Hermitian nnmatrices, ~Hj
~Uexp iXn2
j1j~Hj/C32!
;
where the jare real parameters. The n2~Hjare the generators of the group U
n.
To specialize to SU
nwe need to meet the restriction det ~U1. To impose this
restriction we need to use the identity
dete~AeTr~A
for any square matrix ~A. The proof is left as homework (Problem 12.15). Thus the
condition det ~U1 requires Tr ~H0 for every ~H. Accordingly, the generators
ofSU
nare any set of nntraceless Hermitian matrices.
Forn2,SU
nreduces to SU
2, which describes rotations in two-dimen-
sional complex space. The determinant is 1. There are three continuous para-
meters (22ÿ13). We have expressed these as Cayley–Klein parameters. The
orthogonal group SO
3, determinant 1, describes rotations in ordinary three-
dimensional space and leaves x2y2z2invariant. There are also three inde-
pendent parameters. The rotation interpretations and the equality of numbers of
independent parameters suggest these two groups may be isomorphic or homo-
morphic. The correspondence between these groups has been proved to be two-to-one. Thus SU
2andSO
3are isomorphic. It is beyond the scope of this book to
reproduce the proof here.
TheSU
2group has found various applications in particle physics. For exam-
ple, we can think of the proton ( p) and neutron ( n) as two states of the same
particle, a nucleon /C78, and use the electric charge as a label. It is also useful to
imagine a particle space, called the strong isospin space, where the nucleon statepoints in some direction, as shown in Fig. 12.6. If (or assuming that) the theory
that describes nucleon interactions is invariant under rotations in strong isospin
space, then we may try to put the proton and the neutron as states of a spin-likedoublet, or SU
2doublet. Other hadrons (strong-interacting particles) can also
be classified as states in SU
2multiplets. Physicists do not have a deep under-
standing of why the Standard Model (of Elementary Particles) has an SU
2
internal symmetry.
Forn3 there are eight independent parameters
3
2ÿ18), and we have
SU
3, which is very useful in describing the color symmetry.
453SOME SPECIAL GROUPS
/C72omogeneous /C76orent/C122 group
Before we describe the homogeneous Lorentz group, we need to know the
Lorentz transformation. This will bring us back to the origin of special theory
of relativity. In classical mechanics, time is absolute and the Galilean trans-
formation (the principle of Newtonian relativity) asserts that all inertial frames
are equivalent for describing the laws of classical mechanics. But physicists in
the nineteenth century found that electromagnetic theory did not seem to obey
the principle of Newtonian relativity. Classical electromagnetic theory is sum-
marized in Maxwell’s equations, and one of the consequences of Maxwell’sequations is that the speed of light (electromagnetic waves) is independent of
the motion of the source. However, under the Galilean transformation, in a
frame of reference moving uniformly with respect to the light source the light
wave is no longer spherical and the speed of light is also di/C128erent. Hence, for
electromagnetic phenomena, inertial frames are not equivalent and Maxwell’s
equations are not invariant under Galilean transformation. A number of
experiments were proposed to resolve this conflict. After the Michelson–
Morley experiment failed to detect ether, physicists finally accepted that
Maxwell’s equations are correct and have the same form in all inertial frames.There had to be some transformation other than the Galilean transformation
that would make both electromagnetic theory and classical mechanical invar-
iant.
This desired new transformation is the Lorentz transformation, worked out by
H. Lorentz. But it was not until 1905 that Einstein realized its full implications
and took the epoch-making step involved. In his paper, ‘On the Electrodynamics
of Moving Bodies’ ( The Principle of /C82elativity , Dover, New York, 1952), he
developed the Special Theory of Relativity from two fundamental postulates,which are rephrased as follows:
454ELEMENTS OF GROUP THEORY
Figure 12.6. The strong isospin space.
(1) The laws of physics are the same in all inertial frame. No preferred inertial
frame exists.
(2) The speed of light in free space is the same in all inertial frames and is inde-
pendent of the motion of the source (the emitting body).
These postulates are often called Einstein’s principle of relativity, and they radi-
cally revised our concepts of space and time. Newton’s laws of motion abolish theconcept of absolute space, because according to the laws of motion there is no
absolute standard of rest. The non-existence of absolute rest means that we can-
not give an event an absolute position in space. This in turn means that space is
not absolute. This disturbed Newton, who insisted that there must be some abso-
lute standard of rest for motion, remote stars or the ether system. Absolute space
was finally abolished in its Maxwellian role as the ether. Then absolute time was
abolished by Einstein’s special relativity. We can see this by sending a pulse of
light from one place to another. Since the speed of light is just the distance it has
traveled divided by the time it has taken, in Newtonian theory, di/C128erent observers
would measure di/C128erent speeds for the light because time is absolute. Now in
relativity, all observers agree on the speed of light, but they do not agree on thedistance the light has traveled. So they cannot agree on the time it has taken. That
is, time is no longer absolute.
We now come to the Lorentz transformation, and suggest that the reader to
consult books on special relativity for its derivation. For two inertial frames with
their corresponding axes parallel and the relative velocity /C118along the x
1
xaxis,
the Lorentz transformation has the form:
x0
1/C13
x1i/C12x4;
x0
2x2;
x0
3x3;
x0
4/C13
x4ÿi/C12x1;
where x4ict;/C12/C118=c,a n d /C131=
1ÿ/C122/C112
. We will drop the two directions
perpendicular to the motion in the following discussion.
For an infinitesimal relative velocity /C118, the Lorentz transformation reduces to
x0
1x1i/C12x4;
x0
4x4ÿi/C12x1;
where /C12/C118=c;/C131=
1ÿ
/C122q
/C251:In matrix form we have
x0
1
x0
4/C32!
1 i/C12
ÿi/C12 1/C32!
x1
x4/C32!
:
455SOME SPECIAL GROUPS
We can express the transformation matrix in exponential form:
1 i/C12
ÿi/C12 1/C32!
10
01
/C120i
ÿi0
~I/C12~;
where
~I10
01
; ~0i
ÿi0
:
Note that ~is the negative of the Pauli spin matrix ~2. Now we have
x0
1
x0
4/C32!
~I/C12~x1
x4/C32!
:
We can generate a finite transformation by repeating the infinitesimal transforma-
tion/C78times with /C78/C12:
x0
1
x0
4
~I~
/C78/C78x1
x4
:
In the limit as /C78!1 ,
lim
/C78!1~I~
/C78/C78
e~:
Now we can expand the exponential in a Maclaurin series:
e~~I~
~2=2/C33
~3=3/C33
and, noting that ~21 and
sinh3=3/C335=5/C337=7/C33 ;
cosh 12=2/C334=4/C336=6/C33 ;
we finally obtain
e~~Icosh ~sinh:
Our finite Lorentz transformation then takes the form:
x0
1
x0
2
cosh isinh
ÿisinhcosh x1
x2
;
and ~is the generator of the representations of our Lorentz transformation. The
transformation
cosh isinh
ÿisinhcosh
456ELEMENTS OF GROUP THEORY
can be interpreted as the rotation matrix in the complex x4x1plane (Problem
12.16).
It is straightforward to generalize the above discussion to the general case
where the relative velocity is in an arbitrary direction. The transformation matrix
will be a 4 4 matrix, instead of a 2 2 matrix one. For this general case, we have
to take x2- and x3-axes into consideration.
Problems
12.1. Show that
(a) the unit element (the identity) in a group is unique, and
(b) the inverse of each group element is unique.
12.2. Show that the set of complex numbers 1 ;i;ÿ1, and ÿiform a group of
order four under multiplication.
12.3. Show that the set of all rational numbers, the set of all real numbers,
and the set of all complex numbers form infinite Abelian groups underaddition.
12.4. Show that the four matrices
~A10
01
; ~B01
ÿ10
; ~Cÿ100ÿ1
; ~D0ÿ1
10
form an Abelian group of order four under multiplication.
12.5. Show that the set of all permutations of three objects
123;231;312;132;321;213
forms a non-Abelian group of order six, with sequential performance as
the law of combination.
12.6. Given two elements AandBsubject to the relations A
2B2/C69(the
identity), show that:(a)AB6BA, and
(b) the set of six elements /C69;A;B;A
2;AB;BAform a group.
12.7. Show that the set of elements 1 ;A;A2;...;Anÿ1,An1, where Ae2i=n
forms a cyclic group of order nunder multiplication.
12.8. Consider the rotations of a line about the z-axis through the angles
=2; ;3=2;and 2 in the xyplane. This is a finite set of four elements,
the four operations of rotating through =2; ;3=2, and 2 . Show that
this set of elements forms a group of order four under addition.
12.9. Construct the group multiplication table for the group of Problem 12.2.
12.10. Consider the possible rearrangement of two objects. The operation /C69/C112
leaves each object in its place, and the operation I/C112interchanges the two
objects. Show that the two operations form a group that is isomorphic to
/C712.
457PROBLEMS
Next, we associate with the two operations two operators ^O/C69/C112and ^OI/C112,
which act on the real or complex function f
x1;y1;z1;x2;y2;z2with the
following e/C128ects:
^O/C69/C112ff; ^OI/C112f
x1;y1;z1;x2;y2;z2f
x2;y2;z2;x1;y1;z1:
Show that the two operators form a group that is isomorphic to /C712.
12.11. Verify that the multiplication table of S3has the form:
12.12. Show that an nnorthogonal matrix has n
nÿ1=2 independent
elements.
12.13. Show that the 2 2 matrix 2can be obtained from the rotation matrix
R
by di/C128erentiation at the identity of SO
2, that is, 0.
12.14. Show that an nnunitary matrix has n2ÿ1 independent parameters.
12.15. Show that det e~AeTr ~Awhere ~Ais any square matrix.
12.16. Show that the Lorentz transformation
x0
1/C13
x1i/C12x4;
x0
2x2;
x0
3x3;
x0
4/C13
x4ÿi/C12x1
corresponds to an imaginary rotation in the x4x1plane. (A detailed dis-
cussion of this can be found in the book /C67lassical Mechanics , by Tai L.
Chow, John Wiley, 1995.)
458ELEMENTS OF GROUP THEORY
P1P2P3P4P5P6
P1 P1P2P3P4P5P6
P2 P2P1P6P5P6P4
P3 P3P4P5P6P2P1
P4 P4P5P3P1P6P2
P5 P5P3P4P2P1P6
P6 P6P2P1P3P4P5
13
/C78umerical methods
Very few of the mathematical problems which arise in physical sciences and
engineering can be solved analytically. Therefore, a simple, perhaps crude, tech-
nique giving the desired values within specified limits of tolerance is often to be
preferred. We do not give a full coverage of numerical analysis in this chapter; but
some methods for numerically carrying out the processes of interpolation, finding
roots of equations, integration, and solving ordinary di/C128erential equations will be
presented.
Interpolation
In the eighteenth century Euler was probably the first person to use the interpola-
tion technique to construct planetary elliptical orbits from a set of observed
positions of the planets. We discuss here one of the most common interpolation
techniques: the polynomial interpolation. Suppose we have a set of observed or
measured data
x0;y0,
x1;y1;...;
xn;yn, how do we represent them by a
smooth curve of the form yf
x/C63 For analytical convenience, this smooth
curve is usually assumed to be polynomial:
f
xa0a1x1a2x2 anxn
13:1
and we use the given points to evaluate the coecients a0;a1;...;an:
f
x0a0a1x0a2x2
0 anxn0y0;
f
x1a0a1x1a2x2
1 anxn1y1;
...
f
xna0a1xna2x2
n anxnnyn:9
>>>>>=
>>>>>;
13:2
This provides n1 equations to solve for the n1 coecients a
0;a1;...;an:
However, straightforward evaluation of coecients in the way outlined above
459
is rather tedious, as shown in Problem 13.1, hence many shortcuts have been
devised, though we will not discuss these here because of limited space.
Finding roots of equations
A solution of an equation f
x0 is sometimes called a root, where f
xis a real
continuous function. If f
xis suciently complicated that a direct solution may
not be possible, we can seek approximate solutions. In this section we will sketch
some simple methods for determining the approximate solutions of algebraic and
transcendental equations. A polynomial equation is an algebraic equation. An
equation that is not reducible to an algebraic equation is called transcendental.
Thus, tan xÿx0a n d ex2c o s x0 are transcendental equations.
Graphical methods
The approximate solution of the equation
f
x0
13:3
can be found by graphing the function yf
xand reading from the graph the
values of xfor which y0. The graphing procedure can often be simplified by
first rewriting Eq. (13.3) in the form
/C103
x/C104
x
13:4
and then graphing y/C103
xandy/C104
x. The xvalues of the intersection points
of the two curves gives the approximate values of the roots of Eq. (13.4). As anexample, consider the equation
f
xx
3ÿ146:25xÿ682:50;
we can graph
yx3ÿ146:25xÿ682:5
to find its roots. But it is simpler to graph the two curves
yx3
a cubic
and
y146:25x682:5
a straight line :
See Fig. 13.1.
There is one drawback of graphical methods: that is, they require plotting
curves on a large scale to obtain a high degree of accuracy. To avoid this, methods
of successive approximations (or simple iterative methods) have been devised, and
we shall sketch a couple of these in the following sections.
460NUMERICAL METHODS
Method of linear interpolation (method of false position)
Make an initial guess of the root of Eq. (13.3), say x0, located between x1andx2,
and in the interval ( x1;x2) the graph of yf
xhas the appearance as shown in
Fig. 13.2. The straight line connecting P1andP2cuts the x-axis at point x3;which
is usually closer to x0than either x1orx2. From similar triangles
x3ÿx1
ÿf
x1x2ÿx1
f
x2;
and solving for x3we get
x3x1f
x2ÿx2f
x1
f
x2ÿf
x1:
Now the straight line connecting the points P3andP2intersects the x-axis at point
x4, which is a closer approximation to x0than x3. By repeating this process we
obtain a sequence of values x3;x4;...;xnthat generally converges to the root of
the equation.
The iterative method described above can be simplified if we rewrite Eq. (13.3)
in the form of Eq. (13.4). If the roots of
/C103
xc
13:5
can be determined for every real c, then we can start the iterative process as
follows. Let x1be an approximate value of the root x0of Eq. (13.3) (and, of
461METHOD OF LINEAR INTERPOLATION
Figure 13.1.
course, also equation 13.4). Now setting xx1on the right hand side of Eq.
(13.4) we obtain the equation
/C103
x/C104
x1;
13:6
which by hypothesis we can solve. If the solution is x2, we set xx2on the right
hand side of Eq. (13.4) and obtain
/C103
x/C104
x2:
13:7
By repeating this process, we obtain the nth approximation
/C103
x/C104
xnÿ1:
13:8
From geometric considerations or interpretation of this procedure, we can see
that the sequence x1;x2;...;xnconverges to the root x0 if, in the interval
2jx1ÿx0jcentered at x0, the following conditions are met:
1j/C1030
xj/C62j/C1040
xj;and
2The derivatives are bounded :)
13:9
Example 13.1
Find the approximate values of the real roots of the transcendental equation
exÿ4x0:
Solution: Letg
xxand h
xex=4;so the original equation can be rewrit-
ten as
xex=4:
462NUMERICAL METHODS
Figure 13.2.
According to Eq. (13.8) we have
xn1exn=4;n1;2;3;...:
13:10
There are two roots (see Fig. 13.3), with one around x0:3:If we take it as x1,
then we have, from Eq. (13.10)
x2ex1=40:3374 ;
x3ex2=40:3503 ;
x4ex3=40:3540 ;
x5ex4=40:3565 ;
x6ex5=40:3571 ;
x7ex6=40:3573 :
The computations can be terminated at this point if only three-decimal-place
accuracy is required.
The second root lies between 2 and 3. As the slope of y4xis less than that of
yex, the first condition of Eq. (13.9) cannot be met, so we rewrite the original
equation in the form
ex4x;orxlog 4x
463Figure 13.3.METHOD OF LINEAR INTERPOLATION
and take /C103
xx,/C104
xlog 4 x. We now have
xn1log 4 xn;n1;2;...:
If we take x12:1, then
x2log 4 x12:12823 ;
x3log 4 x22:14158 ;
x4log 4 x32:14783 ;
x5log 4 x42:15075 ;
x6log 4 x52:15211 ;
x7log 4 x62:15303 ;
x8log 4 x72:15316 ;
and we see that the value of the root correct to three decimal places is 2.153.
Ne/C119ton’s method
In Newton’s method, the successive terms in the sequence of approximate values
x1;x2;...;xnthat converges to the root is obtained by the intersection with the x-
axis of the tangent line to the curve yf
x. Fig. 13.4 shows a portion of the
graph of f
xclose to one of its roots, x0. We start with x1, an initial guess of the
value of the root x0. Now the equation of the tangent line to yf
xatP1is
yÿf
x1f0
x1
xÿx1:
13:11
This tangent line intersects the x-axis at x2that is a better approximation to the
root than x1. To find x2, we set y0 in Eq. (13.11) and find
x2x1ÿf
x1=f0
x1
464NUMERICAL METHODS
Figure 13.4.
provided f0
x16 0. The equation of the tangent line at P2is
yÿf
x2f0
x2
xÿx2
and it intersects the x-axis at x3:
x3x2ÿf
x2=f0
x2:
This process is continued until we reach the desired level of accuracy. Thus, in
general
xn1xnÿf
xn
f0
xn;n1;2;...:
13:12
Newton’s method may fail if the function has a point of inflection, or other bad
behavior, near the root. To illustrate Newton’s method, let us consider the follow-ing trivial example.
Example 13.2
Solve, by Newton’s method, x
3ÿ20.
Solution: Here we have yx3ÿ2. If we take x11:5 (note that 1 <21=3<3,
then Eq. (13.12) gives
x21:296296296 ;
x31:260932225 ;
x41:259921861 ;
x51:25992105
x61:25992105)
repetition :
Thus, to eight-decimal-place accuracy, 21=31:25992105.
When applying Newton’s method, it is often convenient to replace f0
xnby
f
xnÿf
xn
;
with small. Usually 0:001 will give good accuracy. Eq. (13.12) then reads
xn1xnÿf
xn
f
xnÿf
xn;n1;2;...:
13:13
Example 13.3Solve the equation x
2ÿ20.
465METHOD OF LINEAR INTERPOLATION
Solution: Here f
xx2ÿ2. Take x11a n d 0:001, then Eq. (13.13) gives
x21:499750125 ;
x31:416680519 ;
x41:414216580 ;
x51:414213563 ;
x61:414113562
x71:414113562)
x6x7:
Numerical integration
Very often definite integrations cannot be done in closed form. When this happens
we need some simple and useful techniques for approximating definite integrals.
In this section we discuss three such simple and useful methods.
/C84he rectangular rule
The reader is familiar with the interpretation of the definite integralRb
af
xdxas
the area under the curve yf
xbetween the limits xaandxb:
Zb
af
xdxXn
i1f
i
xiÿxiÿ1;
where xiÿ1ixi;ax0<x1<x2<<xnb:We can obtain a good
approximation to this definite integral by simply evaluating such an area underthe curve yf
x. We can divide the interval axbintonsubintervals of
length /C104
bÿa=n, and in each subinterval, the function f
iis replaced by a
466NUMERICAL METHODS
Figure 13.5.
straight line connecting the values at each head or end of the subinterval (or at the
center point of the interval), as shown in Fig. 13.5. If we choose the head,
ixiÿ1, then we have
Zb
af
xdx/C25/C104
y0y1 ynÿ1;
13:14
where y0f
x0;y1f
x1;...;ynÿ1f
xnÿ1. This method is called the rec-
tangular rule.
It will be shown later that the error decreases as n2. Thus, as nincreases, the
error decreases rapidly.
/C84he trape/C122oidal rule
The trapezoidal rule evaluates the small area of a subinterval slightly di/C128erently.The area of a trapezoid as shown in Fig. 13.6 is given by
1
2/C104
Y1Y2:
Thus, applied to Fig. 13.5, we have the approximation
Zb
af
xdx/C25
bÿa
n
1
2y0y1y2 ynÿ112yn:
13:15
What are the upper and lower limits on the error of this method/C63 Let us first
calculate the error for a single subinterval of length /C104
bÿa=n. Writing
xi/C104zand/C34i
zfor the error, we have
Zz
xif
xdx/C104
2yiyz/C34i
z;
467NUMERICAL INTEGRATION
Figure 13.6.
where yif
xi;yzf
z.O r
/C34i
zZz
xif
xdxÿ/C104
2f
xiÿf
z
Zz
xif
xdxÿzÿxi
2f
xiÿf
z:
Di/C128erentiating with respect to z:
/C340
i
zf
zÿf
xif
z=2ÿ
zÿxif0
z=2:
Di/C128erentiating once again,
/C3400
i
zÿ
zÿxif00
z=2:
IfmiandMiare, respectively, the minimum and the maximum values of f00
zin
the subinterval /C91 xi;z/C93, we can write
zÿxi
2miÿ/C3400
i
zzÿxi
2Mi:
Anti-di/C128erentiation gives
zÿxi2
4miÿ/C340
i
z
zÿxi2
4Mi
Anti-di/C128erentiation once more gives
zÿxi3
12miÿ/C34i
z
zÿxi3
12Mi:
or, since zÿxi/C104,
/C1043
12miÿ/C34i/C1043
12Mi:
Ifmand Mare, respectively, the minimum and the maximum of f00
zin the
interval /C91 a;b/C93 then
/C1043
12mÿ/C34i/C1043
12M for all i:
Adding the errors for all subintervals, we obtain
/C1043
12nmÿ/C34/C1043
12nM
or, since /C104
bÿa=n;
bÿa3
12n2mÿ/C34
bÿa3
12n2M:
13:16
468NUMERICAL METHODS
Thus, the error decreases rapidly as nincreases, at least for twice-di/C128erentiable
functions.
Simpson’s rule
Simpson’s rule provides a more accurate and useful formula for approximating a
definite integral. The interval axbis subdivided into an even number of
subintervals. A parabola is fitted to points a,a/C104,a2/C104; another to a2/C104,
a3/C104,a4/C104; and so on. The area under a parabola, as shown in Fig. 13.7, is
(Problem 13.8)
/C104
3
y14y2y3:
Thus, applied to Fig. 13.5, we have the approximation
Zb
af
xdx/C25/C1043
y
04y12y24y32y4 2ynÿ24ynÿ1yn;
13:17
with neven and /C104
bÿa=n.
The analysis of errors for Simpson’s rule is fairly involved. It has been shown
that the error is proportional to /C1044(or inversely proportional to n4).
There are other methods of approximating integrals, but they are not so simple
as the above three. The method called Gaussian quadrature is very fast but more
involved to implement. Many textbooks on numerical analysis cover this method.
Numerical solutions of di/C128erential equations
We noted in Chapter 2 that the methods available for the exact solution of di/C128er-ential equations apply only to a few, principally linear, types of di/C128erential equa-
tions. Many equations which arise in physical science and in engineering are not
solvable by such methods and we are therefore forced to find ways of obtaining
approximate solutions of these di/C128erential equations. The basic idea of approxi-
mate solutions is to specify a small increment hand to obtain approximate values
of a solution yy
xatx
0,x0/C104,x02/C104;...:
469NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS
Figure 13.7.
The first-order ordinary di/C128erential equation
dy
dxf
x;y;
13:18
with the initial condition yy0when xx0, has the solution
yÿy0Zx
x0f
t;y
tdt:
13:19
This integral equation cannot be evaluated because the value of yunder the
integral sign is unknown. We now consider three simple methods of obtaining
approximate solutions: Euler’s method, Taylor series method, and the Runge–
Kutta method.
Euler’s method
Euler proposed the following crude approach to finding the approximate solution.He began at the initial point
x
0;y0and extended the solution to the right to the
point x1x0/C104, where /C104is a small quantity. In order to use Eq. (13.19) to
obtain the approximation to y
x1, he had to choose an approximation to fon
the interval x0;x1. The simplest of all approximations is to use f
t;y
t
f
x0;y0. With this choice, Eq. (13.19) gives
y
x1y0Zx1
x0f
x0;y0dty0f
x0;y0
x1ÿx0:
Letting y1y
x1, we have
y1y0f
x0;y0
x1ÿx0:
13:20
From y1;y0
1f
x1;y1can be computed. To extend the approximate solution
further to the right to the point x2x1/C104, we use the approximation:
f
t;y
t y0
1f
x1;y1. Then we obtain
y2y
x2y1Zx2
x1f
x1;y1dty1f
x1;y1
x2ÿx1:
Continuing in this way, we approximate y3,y4, and so on.
There is a simple geometrical interpretation of Euler’s method. We first note
that f
x0;y0y0
x0, and that the equation of the tangent line at the point
x0;y0to the actual solution curve (or the integral curve) yy
xis
yÿy0Zx
x0f
t;y
tdtf
x0;y0
xÿx0:
Comparing this with Eq. (13.20), we see that
x1;y1lies on the tangent line to the
actual solution curve at ( x0;y0). Thus, to move from point ( x0;y0) to point ( x1;y1)
we proceed along this tangent line. Similarly, to move to point ( x2;y2we proceed
parallel to the tangent line to the solution curve at ( x1;y1, as shown in Fig. 13.8.
470NUMERICAL METHODS
The merit of Euler’s method is its simplicity, but the successive use of the
tangent line at the approximate values y1;y2;...can accumulate errors. The accu-
racy of the approximate vale can be quite poor, as shown by the following simple
example.
Example 13.4
Use Euler’s method to approximate solution to
y0x2y;y
13 on interval 1;2:
Solution: Using h0:1, we obtain Table 13.1. Note that the use of a smaller
step-size hwill improve the accuracy.
Euler’s method can be improved upon by taking the gradient of the integral
curve as the means of obtaining the slopes at x0andx0/C104, that is, by using the
471NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS
Table 13.1.
xy (Euler) y(actual)
1.0 3 3
1.1 3.4 3.431371.2 3.861 3.931221.3 4.3911 4.50887
1.4 4.99921 5.1745
1.5 5.69513 5.939771.6 6.48964 6.81695
1.7 7.39461 7.82002
1.8 8.42307 8.964331.9 9.58938 10.2668
2.0 10.9093 11.7463
Figure 13.8.
approximate value obtained for y1, we obtain an improved value, denoted by
y11:
y11y01
2ff
x0;y0f
x0/C104;y1g:
13:21
This process can be repeated until there is agreement to a required degree of
accuracy between successive approximations.
/C84he three-term /C84aylor series method
The rationale for this method lies in the three-term Taylor expansion. Let ybe the
solution of the first-order ordinary equation (13.18) for the initial condition
yy0when xx0and suppose that it can be expanded as a Taylor series in
the neighborhood of x0.I fyy1when xx0/C104, then, for suciently small
values of /C104,w eh a v e
y1y0/C104dy
dx
0/C1042
2/C33d2y
dx2/C32!
0/C1043
3/C33d3y
dx3/C32!
0 :
13:22
Now
dy
dxf
x;y;
d2y
dx2/C64f
/C64xdy
dx/C64f
/C64y/C64f
/C64xf/C64f
/C64y;
and
d3y
dx3/C64
/C64xf/C64
/C64y/C64f
/C64xf/C64f
/C64y
/C642f
/C64x2/C64f
/C64y/C64f
/C64y2f/C642f
/C64x/C64yf/C64f
/C64y2
f2/C642f
/C64y2:
Equation (13.22) can be rewritten as
y1y0/C104f
x0;y0/C1042
2/C64f
x0;y0
/C64xf
x0;y0/C64f
x0;y0
/C64y
;
where we have dropped the /C1043term. We now use this equation as an iterative
equation:
yn1yn/C104f
xn;yn/C1042
2/C64f
xn;yn
/C64xf
xn;yn/C64f
xn;yn
/C64y
:
13:23
That is, we compute y1y
x0/C104from y0,y2y
x1/C104from y1by replacing x
byx1, and so on. The error in this method is proportional /C1043. A good approxima-
472NUMERICAL METHODS
tion can be obtained for ynby summing a number of terms of the Taylor’s
expansion. To illustrate this method, let us consider a very simple example.
Example 13.5
Find the approximate values of y1through y10for the di/C128erential equation
y0xy, with the initial condition x01:0 and yÿ2:0.
Solution: Now f
x;yxy;/C64f=/C64x/C64f=/C64y1 and Eq. (13.23) reduces to
yn1yn/C104
xnyn/C1042
2
1xnyn:
Using this simple formula with /C1040:1 we obtain the results shown in Table 13.2.
/C84he Runge/C177/C75utta method
In practice, the Taylor series converges slowly and the accuracy involved is not
very high. Thus we often resort to other methods of solution such as the Runge–
Kutta method, which replaces the Taylor series, Eq. (13.23), with the following
formula:
yn1yn/C104
6
k14k2k3;
13:24
where
k1f
xn;yn;
13:24a
k2f
xn/C104=2;yn/C104k1=2;
13:24b
k3f
xn/C104;y02/C104k2ÿ/C104k1:
13:24c
This approximation is equivalent to Simpson’s rule for the approximate
integration of f
x;y, and it has an error proportional to /C1044. A beauty of the
473NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS
Table 13.2.
n xn yn yn1
0 1.0 ÿ2.0 ÿ2.1
1 1.1 ÿ2.1 ÿ2.2
2 1.2 ÿ2.2 ÿ2.3
3 1.3 ÿ2.3 ÿ2.4
4 1.4 ÿ2.4 ÿ2.5
5 1.5 ÿ2.5 ÿ2.6
Runge–Kutta method is that we do not need to compute partial derivatives, but it
becomes rather complicated if pursued for more than two or three steps.
The accuracy of the Runge–Kutta method can be improved with the following
formula:
yn1yn/C104
6
k12k22k3k4;
13:25
where
k1f
xn;yn;
13:25a
k2f
xn/C104=2;yn/C104k1=2;
13:25b
k3f
xn/C104;y0/C104k2=2;
13:25c
k4f
xn/C104;yn/C104k3:
13:25d
With this formula the error in yn1is of order /C1045.
You may wonder how these formulas are established. To this end, let us go
back to Eq. (13.22), the three-term Taylor series, and rewrite it in the form
y1y0/C104f0
1=2/C1042
A0f0B0
1=6/C1043
C02f0D0f2
0/C690
A0B0f0B2
0O
/C1044;
13:26
where
A/C64f
/C64x;B/C64f
/C64y;C/C642f
/C64x2;D/C642f
/C64x/C64y;/C69/C642f
/C64y2
and the subscript 0 denotes the values of these quantities at
x0;y0.
Now let us expand k1;k2,a n d k3in the Runge–Kutta formula (13.24) in powers
ofhin a similar manner:
k1/C104f
x0;y0;
k2f
x0/C104=2;y0k1/C104=2;
f01
2/C104
A0f0B018/C104
2
C02f0D0f2
0/C690O
/C1043:
Thus
2k2ÿk1f0/C104
A0f0B0
and
d
d/C104
2k2ÿk1
/C1040f0;d2
d/C1042
2k2ÿk1/C32!
/C10402
A0f0B0:
474NUMERICAL METHODS
Then
k3f
x0/C104;y02/C104k2ÿ/C104k1
f0/C104
A0f0B0
1=2/C1042fC02f0D0f2
0/C6902B0
A0f0B0g
O
/C1043:
and
1=6
k14k2k3/C104f0
1=2/C1042
A0f0B0
1=6/C1043
C02f0D0f2
0/C690A0B0f0B2
0O
/C1044:
Comparing this with Eq. (13.26), we see that it agrees with the Taylor series
expansion (up to the term in /C1043) and the formula is established. Formula (13.25)
can be established in a similar manner by taking one more term of the Taylor
series.
Example 13.6
Using the Runge–Kutta method and /C1040:1, solve
y0xÿy2=10;x00;y01:
Solution: With h0:1,h40:0001 and we may use the Runge–Kutta third-
order approximation.
First step: x00;y00;f0ÿ0:1;
k1ÿ0:1,y0/C104k1=20:995;
k2ÿ0:049;2k2ÿk10:002;k30;
y1y0/C104
6
k14k2k10:9951 :
Second step: x1x0/C1040:1,y10:9951, f10:001,
k10:001,y1/C104k1=20:9952 ;
k20:051, 2 k2ÿk10:101,k30:099,
y2y1/C1046
k
14k2k11:0002 :
Third step: x2x1/C1040:2;y21:0002, f20:1,
k10:1,y2/C104k1=21:0052,
k20:149;2k2ÿk10:198;k30:196;
y3y2/C1046
k
14k2k11:0151 :
475NUMERICAL SOLUTIONS OF DIFFERENTIAL EQUATIONS
Equations of higher order/C46 System of equations
The methods in the previous sections can be extended to obtain numerical solu-
tions of equations of higher order. An nth-order di/C128erential equation is equivalent
tonfirst-order di/C128erential equations in n1 variables. Thus, for instance, the
second-order equation
y00f
x;y;y0;
13:27
with initial conditions
y
x0y0;y0
x0y0
0;
13:28
can be written as a system of two equations of first order by setting
y0u;
13:29
then Eqs. (13.27) and (13.28) become
u0f
x;y;u;
13:30
y
x0y0;u
x0u0:
13:31
The two first-order equations (13.29) and (13.30) with the initial conditions
(13.31) are completely equivalent to the original second-order equation (13.27)
with the initial conditions (13.28). And the methods in the previous sections for
determining approximate solutions can be extended to solve this system of two
first-order equations. For example, the equation
y00ÿy2;
with initial conditions
y
0ÿ 1;y0
01;
is equivalent to the system
y0xu;u01y;
with
y
0ÿ 1;u
01:
These two first-order equations can be solved with Taylor’s method (Problem
13.12).
The simple methods outlined above all have the disadvantage that the error in
approximating to values of yis to a certain extent cumulative and may become
large unless some form of checking process is included. For this reason, methodsof solution involving finite di/C128erence are devised, most of them being variations of
the Adams–Bashforth method that contains a self-checking process. This method
476NUMERICAL METHODS
is quite involved and because of limited space we shall not cover it here, but it is
discussed in any standard textbook on numerical analysis.
Least-squares fit
We now look at the problem of fitting of experimental data. In some experimentalsituations there may be underlying theory that suggests the kind of function to be
used in fitting the data. Often there may be no theory on which to rely in selecting
a function to represent the data. In such circumstances a polynomial is often used.We saw earlier that the m1 coecients in the polynomial
ya
0a1x amxm
can always be determined so that a given set of m1 points ( xi;yi), where the xs
may be unequal, lies on the curve described by the polynomial. However, whenthe number of points is large, the degree mof the polynomial is high, and an
attempt to fit the data by using a polynomial is very laborious. Furthermore, theexperimental data may contain experimental errors, and so it may be more
sensible to represent the data approximately by some function yf
xthat
contains a few unknown parameters. These parameters can then be determinedso that the curve yf
xfits the data. How do we determine these unknown
parameters/C63
Let us represent a set of experimental data ( x
i;yi), where i1;2;...;n, by some
function yf
xthat contains rparameters a1;a2;...;ar. We then take the
deviations (or residuals)
dif
xiÿyi
13:32
and form the weighted sum of squares of the deviations
SXn
i1/C119i
di2Xn
i1/C119if
xiÿyi2;
13:33
where the weights /C119iexpress our confidence in the accuracy of the experimental
data. If the points are equally weighted, the /C119s can all be set to 1.
It is clear that the quantity Sis a function of as:SS
a1;a2;...;ar:We can
now determine these parameters so that Sis a minimum:
/C64S
/C64a10;/C64S
/C64a20;...;/C64S
/C64ar0:
13:34
The set of requations (13.34) is called the normal equations and serves to
determine the runknown asi nyf
x. This particular method of determining
the unknown as is known as the method of least squares.
477LEAST-SQUARES FIT
We now illustrate the construction of the normal equations with the simplest
case: yf
xis a linear function:
ya1a2x:
13:35
The deviations diare given by
di
a1a2xÿyi
and so, assuming /C119i1
SXn
i1d2
i
a1a2x1ÿy12
a1a2x2ÿy22
a1a2xrÿyr2:
We now find the partial derivatives of S with respect to a1anda2and set these to
zero:
/C64S=/C64a12
a1a2x1ÿy12
a1a2x2ÿy2 2
a1a2xnÿyn0;
/C64S=/C64a22x1
a1a2x1ÿy12x2
a1a2x2ÿy2 2xn
a1a2xnÿyn
0:
Dividing out the factor 2 and collecting the coecients of a1anda2, we obtain
na1Xn
i1xi/C32!
a2Xn
i1y1;
13:36
Xn
i1xi/C32!
a1Xn
i1x2
i/C32!
a2Xn
i1xiyi:
13:37
These equations can be solved for a1anda2.
Problems
13.1. Given six points
ÿ1;0,
ÿ0:8;2,
ÿ0:6;1,
ÿ0:4;ÿ1,
ÿ0:2;0;and
0;ÿ4, determine a smooth function yf
xsuch that yif
xi:
13.2. Find an approximate value of the real root of
xÿtanx0
near x3=2:
13.3. Find the angle subtended at the center of a circle by an arc whose length is
double the length of the chord.
13.4. Use Newton’s method to solve
ex2ÿx33xÿ40;
with x00 and /C1040:001:
478NUMERICAL METHODS
13.5. Use Newton’s method to find a solution of
sin
x321=x;
with x01a n d /C1040:001.
13.6. Approximate the following integrals using the rectangular rule, the
trapezoidal rule, and Simpson’s rule, with n2;4;10;20;50:
(a)Z=2
0eÿx2sin
x21dx ;
(b)Z
2p
0sin
x23xÿ2
x4dx ;
(c)Z1
0dx
2ÿsin2x/C112 :
13.7 Show that the area under a parabola, as shown in Fig. 13.7, is given by
A/C104
3
y14y2y3:
13.8. Using the improved Euler’s method, find the value of ywhen x0:2o n
the integral curve of the equation y0x2ÿ2ythrough the point x0,
y1.
13.9. Using Taylor’s method, find correct to four places of decimals values of y
corresponding to x0:2 and xÿ0:2 for the solution of the di/C128erential
equation
dy=dxxÿy2=10;
with the initial condition y1 when x0.
13.10. Using the Runge–Kutta method and /C1040:1, solve
y0x2ÿsin
y2;x01 and y04:7:
13.11. Using the Runge–Kutta method and /C1040:1, solve
y0yeÿx2;x01 and y03:
13.12. Using Taylor’s method, obtain the solution of the system
y0xu;u01y
withy
0ÿ 1;u
01:
.
13.13. Find to four places of decimals the solution between x0a n d x0:5o f
the equations
y01
2
yu;u012
y2ÿu2;
with yu1 when x0.
479PROBLEMS
13.14. Find to three places of decimals a solution of the equation
y002xy0ÿ4y0;
with yy01 when x0:
13.15. Use Eqs. (13.36) and (13.37) to calculate the coecients in ya1a2xto
fit the following data:
x;y
1;1:7;
2;1:8;
3;2:3;
4;3:2:
480NUMERICAL METHODS
14
Introduction to probability theory
The theory of probability is so useful that it is required in almost every branch of
science. In physics, it is of basic importance in quantum mechanics, kinetic theory,
and thermal and statistical physics to name just a few topics. In this chapter the
reader is introduced to some of the fundamental ideas that make probability
theory so useful. We begin with a review of the definitions of probability, a
brief discussion of the fundamental laws of probability, and methods of counting
(some facts about permutations and combinations), probability distributions are
then treated.
A notion that will be used very often in our discussion is ‘equally likely’. This
cannot be defined in terms of anything simpler, but can be explained and illu-strated with simple examples. For example, heads and tails are equally likely
results in a spin of a fair coin; the ace of spades and the ace of hearts are equally
likely to be drawn from a shu/C130ed deck of 52 cards. Many more examples can be
given to illustrate the concept of ‘equally likely’.
/C65 definition of probabilit/C121
Now a question that arises naturally is that of how shall we measure the
probability that a particular case (or outcome) in an experiment (such as the
throw of dice or the draw of cards) out of many equally likely cases that will
occur. Let us flip a coin twice, and ask the question: what is the probability of itcoming down heads at least once. There are four equally likely results in flipping a
coin twice: /C72/C72,/C72T,T/C72/C44 TT , where /C72stands for head and Tfor tail. Three of the
four results are favorable to at least one head showing, so the probability ofgetting one head is 3/4. In the example of drawn cards, what is the probability
of drawing the ace of spades/C63 Obviously there is one chance out of 52, and the
probability, accordingly, is 1/52. On the other hand, the probability of drawing an
481
ace is four times as great ÿ4=52, for there are four aces, equally likely. Reasoning
in this way, we are led to give the notion of probability the following definition:
If there are /C78mutually exclusive, collective exhaustive, and
equally likely outcomes of an experiment, and nof these are
favorable to an event A, then the probability /C112
Aof an event
Aisn=/C78:/C112n=/C78,o r
/C112
Anumber of outcomes favorable to A
total number of results:
14:1
We have made no attempt to predict the result, just to measure it. The definition
of probability given here is often called a posteriori probability.
The terms exclusive and exhaustive need some attention. Two events are said to
be mutually exclusive if they cannot both occur together in a single trial; and the
term collective exhaustive means that all possible outcomes or results are enum-
erated in the /C78outcomes.
If an event is certain not to occur its probability is zero, and if an event is
certain to occur, then its probability is 1. Now if pis the probability that an event
will occur, then the probability that it will fail to occur is 1 ÿ/C112, and we denote it
by/C113:
/C1131ÿ/C112:
14:2
Ifpis the probability that an event will occur in an experiment, and if the
experiment is repeated Mtimes, then the expected number of times the event will
occur is Mp. For suciently large M,Mpis expected to be close to the actual
number of times the event will occur. For example, the probability of a headappearing when tossing a coin is 1/2, the expected number of times heads appear
is 41=2 or 2. Actually, heads will not always appear twice when a coin is tossed
four times. But if it is tossed 50 times, the number of heads that appear will, on theaverage, be close to 25
501=225). Note that closeness is computed on a
percentage basis: 20 is 20/C37 of 25 away from 25 while 1 is 50/C37 of 2 away from 2.
/C83ample space
The equally likely cases associated with an experiment represent the possible out-comes. For example, the 36 equally likely cases associated with the throw of a pairof dice are the 36 ways the dice may fall, and if 3 coins are tossed, there are 8
equally likely cases corresponding to the 8 possible outcomes. A list or set that
consists of all possible outcomes of an experiment is called a sample space and
each individual outcome is called a sample point (a point of the sample space).
The outcomes composing the sample space are required to be mutually exclusive.
As an example, when tossing a die the outcomes ‘an even number shows’ and
482INTRODUCTION TO PROBABILITY THEORY
‘number 4 shows’ cannot be in the same sample space. Often there will be more
than one sample space that can describe the outcome of an experiment but there isusually only one that will provide the most information. In a throw of a fair die,
one sample space is the set of all possible outcomes /C1231, 2, 3, 4, 5, 6/C125, and another
could be /C123even/C125 or /C123odd/C125.
A finite sample space is one that has only a finite number of points. The points
of the sample space are weighted according to their probabilities. To see this, let
the points have the probabilities
/C112
1;/C1122;...;/C112/C78
with
/C1121/C1122 /C112/C781:
Suppose the first nsample points are favorable to another event A. Then the
probability of Ais defined to be
/C112
A/C1121/C1122 /C112n:
Thus the points of the sample space are weighted according to their probabilities.
If each point has the sample probability 1/ n, then /C112
Abecomes
/C112
A1
/C781
/C781
/C78n
/C78
and this definition is consistent with that given by Eq. (14.1).
A sample space with constant probability is called uniform. Non-uniform sam-
ple spaces are more common. As an example, let us toss four coins and count the
number of heads. An appropriate sample space is composed of the outcomes
0 heads ;1 head ;2 heads ;3 heads ;4 heads ;
with respective probabilities, or weights
1=16;4=16;6=16;4=16;1=16:
The four coins can fall in 2 22224, or 16 ways. They give no heads (all
land tails) in only one outcome, and hence the required probability is 1/16. There
are four ways to obtain 1 head: a head on the first coin or on the second coin, and
so on. This gives 4/16. Similarly we can obtain the probabilities for the other
cases.
We can also use this simple example to illustrate the use of sample space. What
is the probability of getting at least two heads/C63 Note that the last three samplepoints are favorable to this event, hence the required probability is given by
6
164
161
1611
16:
483SAMPLE SPACE
Methods of counting
In many applications the total number of elements in a sample space or in an
event needs to be counted. A fundamental principle of counting is this: if one
thing can be done in ndi/C128erent ways and another thing can be done in mdi/C128erent
ways, then both things can be done together or in succession in mndi/C128erent ways.
As an example, in the example of throwing a pair of dice cited above, there are 36
equally like outcomes: the first die can fall in six ways, and for each of these the
second die can also fall in six ways. The total number of ways is
6666666636
and these are equally likely.
Enumeration of outcomes can become a lengthy process, or it can become a
practical impossibility. For example, the throw of four dice generates a samplespace with 6
41296 elements. Some systematic methods for the counting are
desirable. Permutation and combination formulas are often very useful.
Permutations
A permutation is a particular ordered selection. Suppose there are nobjects and r
of these objects are arranged into rnumbered spaces. Since there are nways of
choosing the first object, and after this is done there are nÿ1 ways of choosing
the second object, ...;and finally nÿ
rÿ1ways of choosing the rth object, it
follows by the fundamental principle of counting that the number of di/C128erentarrangements or permutations is given by
nPrn
nÿ1
nÿ2
nÿr1:
14:3
where the product on the right-hand side has rfactors. We call nPrthe number of
permutations of nobjects taken rat a time. When rn, we have
nPnn
nÿ1
nÿ21n/C33:
We can rewrite nPrin terms of factorials:
nPrn
nÿ1
nÿ2
nÿr1
n
nÿ1
nÿ2
nÿr1
nÿr21
nÿr21
n/C33
nÿr/C33:
When rn,w eh a v e nPnn/C33=
nÿn/C33n/C33=0/C33. This reduces to n/C33 if we have
0/C331 and mathematicians actually take this as the definition of 0/C33.
Suppose the nobjects are not all di/C128erent. Instead, there are n1objects of one
kind (that is, indistinguishable from each other), n2that is of a second kind ;...;nk
484INTRODUCTION TO PROBABILITY THEORY
of a kth kind so that n1n2 nkn. A natural question is that of how
many distinguishable arrangements are there of these nobjects. Assuming that
there are /C78di/C128erent arrangements, and each distinguishable arrangement appears
n1/C33,n2/C33;...times, where n1/C33is the number of ways of arranging the n1objects,
similarly for n2/C33;...;nk/C33:Then multiplying /C78byn1/C33n2/C33;...;nk/C33we obtain the
number of ways of arranging the nobjects if they were all distinguishable, that
is,nPnn/C33:
/C78n1/C33n2/C33nk/C33n/C33 or /C78n/C33=
n1/C33n2/C33...nk/C33:
/C78is often written as nPn1n2:::nk, and then we have
nPn1n2:::nkn/C33
n1/C33n2/C33nk/C33:
14:4
For example, given six coins: one penny, two nickels and three dimes, the number
of permutations of these six coins is
6P1236/C33=1/C332/C333/C3360:
/C67ombinations
A permutation is a particular ordered selection. Thus 123 is a di/C128erent permuta-tion from 231. In many problems we are interested only in selecting objects with-
out regard to order. Such selections are called combinations. Thus 123 and 231
are now the same combination. The notation for a combination is
nCrwhich
means the number of ways in which robjects can be selected from nobjects
without regard to order (also called the combination of nobjects taken rat a
time). Among the nPrpermutations there are r/C33 that give the same combination.
Thus, the total number of permutations of ndi/C128erent objects selected rat a time is
r/C33nCrnPrn/C33
nÿr/C33:
Hence, it follows that
nCrn/C33
r/C33
nÿr/C33:
14:5
It is straightforward to show that
nCrn/C33
r/C33
nÿr/C33n/C33
nÿ
nÿr/C33
nÿr/C33nCnÿr:
nCris often written as
nCrn
r
:
485METHODS OF COUNTING
The numbers (14.5) are often called binomial coecients because they arise in
the binomial expansion
xynxnn
1
xnÿ1yn2
x
nÿ2y2nn
y
n:
When nis very large a direct evaluation of n/C33 is impractical. In such cases we use
Stirling’s approximate formula
n/C33/C25
2np
nneÿn:
The ratio of the left hand side to the right hand side approaches 1 as n!1 . For
this reason the right hand side is often called an asymptotic expansion of the left
hand side.
Fundamental probabilit/C121 theorems
So far we have calculated probabilities by directly making use of the definitions; itis doable but it is not always easy. Some important properties of probabilities will
help us to cut short our computation works. These important properties are often
described in the form of theorems. To present these important theorems, let us
consider an experiment, involving two events Aand B, with /C78equally likely
outcomes and let
n
1number of outcomes in which Aoccurs ;but not B;
n2number of outcomes in which Boccurs ;but not A;
n3number of outcomes in which both AandBoccur ;
n4number of outcomes in which neither AnorBoccurs :
This covers all possibilities, hence n1n2n3n4/C78:
The probabilities of AandBoccurring are respectively given by
P
An1n3
/C78; P
Bn2n3
/C78;
14:6
the probability of either AorB(or both) occurring is
P
ABn1n2n3
/C78;
14:7
and the probability of both AandBoccurring successively is
P
ABn3
/C78:
14:8
Let us rewrite P
ABas
P
ABn3
/C78n1n3
/C78n3
n1n3:
486INTRODUCTION TO PROBABILITY THEORY
Now
n1n3=/C78isP
Aby definition. After Ahas occurred, the only possible
cases are the
n1n3cases favorable to A. Of these, there are n3cases favorable
toB, the quotient n3=
n1n3represents the probability of Bwhen it is known
that Aoccurred, PA
B. Thus we have
P
ABP
APA
B:
14:9
This is often known as the theorem of joint (or compound) probability. In words,
the joint probability (or the compound probability) of AandBis the product of
the probability that Awill occur times the probability that Bwill occur if Adoes.
PA
Bis called the conditional probability of Bgiven A(that is, given that Ahas
occurred).
To illustrate the theorem of joint probability (14.9), we consider the probability
of drawing two kings in succession from a shu/C130ed deck of 52 playing cards. The
probability of drawing a king on the first draw is 4/52. After the first king has
been drawn, the probability of drawing another king from the remaining 51 cards
is 3/51, so that the probability of two kings is
4
523
511
221:
If the events Aand Bare independent, that is, the information that Ahas
occurred does not influence the probability of B, then PA
BP
Band the
joint probability takes the form
P
ABP
AP
B;for independent events :
14:10
As a simple example, let us toss a coin and a die, and let Abe the event ‘head
shows’ and Bis the event ‘4 shows.’ These events are independent, and hence the
probability that 4 and a head both show is
P
ABP
AP
B
1=2
1=61=12:
Theorem (14.10) can be easily extended to any number of independent events
A;B;C;...:
Besides the theorem of joint probability, there is a second fundamental relation-
ship, known as the theorem of total probability. To present this theorem, let us go
back to Eq. (14.4) and rewrite it in a slightly di/C128erent form
P
ABn1n2n3
/C78
n1n22n3ÿn3
/C78
n1n3
n2n3ÿn3
/C78
n1n3
/C78n2n3
/C78ÿn3
/C78P
AP
BÿP
AB;
P
ABP
AP
BÿP
AB:
14:11
487FUNDAMENTAL PROBABILITY THEOREMS
This theorem can be represented diagrammatically by the intersecting points sets
Aand Bshown in Fig. 14.1. To illustrate this theorem, consider the simple
example of tossing two dice and find the probability that at least one die gives2. The probability that both give 2 is 1/36. The probability that the first die gives 2is 1/6, and similarly for the second die. So the probability that at least one gives 2
is
P
AB1=61=6ÿ1=3611=36:
For mutually exclusive events, that is, for events A,Bwhich cannot both occur,
P
AB0 and the theorem of total probability becomes
P
ABP
AP
B; for mutually exclusive events :
4:12
For example, in the toss of a die, ‘4 shows’ (event A) and ‘5 shows’ (event B)a r e
mutually exclusive, the probability of getting either 4 or 5 is
P
ABP
AP
B1=61=61=3:
The theorems of total and joint probability for uniform sample spaces estab-
lished above are also valid for arbitrary sample spaces. Let us consider a finite
sample space, its events /C69
iare so numbered that /C691;/C692;...;/C69jare favorable to
A;/C69j1;...;/C69kare favorable to both AandB, and /C69k1;...;/C69mare favorable to B
only. If the associated probabilities are /C112i, then Eq. (14.11) is equivalent to the
identity
/C1121 /C112m
/C1121 /C112j/C112j1 /C112k
/C112j1 /C112k/C112k1 /C112mÿ
/C112j1 /C112m:
The sums within the three parentheses on the right hand side represent, respec-tively, P
AP
B;andP
ABby definition. Similarly, we have
P
AB/C112
j1 /C112k
/C1121 /C112k/C112j1
/C1121 /C112k/C112k
/C1121 /C112k
P
APA
B;
which is Eq. (14.9).
488INTRODUCTION TO PROBABILITY THEORY
Figure 14.1.
/C82andom /C118ariables and probabilit/C121 distributions
As demonstrated above, simple probabilities can be computed from elementary
considerations. We need more ecient ways to deal with probabilities of whole
classes of events. For this purpose we now introduce the concepts of random
variables and a probability distribution.
Random variables
A process such as spinning a coin or tossing a die is called random since it isimpossible to predict the final outcome from the initial state. The outcomes of a
random process are certain numerically valued variables that are often called
random variables. For example, suppose that three dimes are tossed at the
same time and we ask how many heads appear. The answer will be 0, 1, 2, or 3
heads, and the sample space Shas 8 elements:
SfTTT ;HTT ;THT ;TTH ;HHT ;HTH ;THH ;HHH g:
The random variable Xin this case is the number of heads obtained and it
assumes the values
0;1;1;1;2;2;2;3:
For instance, X1 corresponds to each of the three outcomes:
HTT ;THT ;TTH . That is, the random variable Xcan be thought of as a function
of the number of heads appear.
A random variable that takes on a finite or countable infinite number of values
(that is it has as many values as the natural numbers 1 ;2;3;...) is called a discrete
random variable while one that takes on a non-countable infinite number ofvalues is called a non-discrete or continuous random variable.
Probability distributions
A random variable, as illustrated by the simple example of tossing three dimes atthe same time, is a numerical-valued function defined on a sample space. In
symbols,
X
s
ixi i1;2;...;n;
14:13
where siare the elements of the sample space and xiare the values of the random
variable X. The set of numbers xican be finite or infinite.
In terms of a random variable we will write P
Xxias the probability that
the random variable Xtakes the value xi, and P
X<xias the probability that
the random variable takes values less than xi, and so on. For simplicity, we often
write P
Xxias/C112i. The pairs
xi;/C112ifori1;2;3;...define the probability
489RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS
distribution or probability function for the random variable X. Evidently any
probability distribution /C112ifor a discrete random variable must satisfy the follow-
ing conditions:
(i)0/C112i1;
(ii) the sum of all the probabilities must be unity (certainty),P
i/C112i1:
Expectation and variance
The expectation or expected value or mean of a random variable is defined in
terms of a weighted average of outcomes, where the weighting is equal to the
probability /C112iwith which xioccurs. That is, if Xis a random variable that can take
the values x1;x2;...;with probabilities /C1121;/C1122;...;then the expectation or
expected value /C69
Xis defined by
/C69
X/C1121x1/C1122x2X
i/C112ixi:
14:14
Some authors prefer to use the symbol /C22for the expectation value /C69
X. For the
three dimes tossed at the same time, we have
xi01 2 3
/C112i1=83=83=81=8
and
/C69
X1
8038138218332:
We often want to know how much the individual outcomes are scattered away
from the mean. A quantity measure of the spread is the di/C128erence Xÿ/C69
Xand
this is called the deviation or residual. But the expectation value of the deviations
is always zero:
/C69
Xÿ/C69
X X
i
xiÿ/C69
X/C112iX
ixi/C112iÿ/C69
XX
i/C112i
/C69
Xÿ/C69
X10:
This should not be particularly surprising; some of the deviations are positive, and
some are negative, and so the mean of the deviations is zero. This means that the
mean of the deviations is not very useful as a measure of spread. We get around
the problem of handling the negative deviations by squaring each deviation,
thereby obtaining a quantity that is always positive. Its expectation value is calledthe variance of the set of observations and is denoted by
2
2/C69
Xÿ/C69
X2/C69
Xÿ/C222:
14:15
490INTRODUCTION TO PROBABILITY THEORY
The square root of the variance, , is known as the standard deviation, and it is
always positive.
We now state some basic rules for expected values. The proofs can be found in
any standard textbook on probability and statistics. In the following cis a con-
stant, Xand/C89are random variables, and /C104
Xis a function of X:
(1)/C69
cXc/C69
X;
(2)/C69
XY/C69
X/C69
Y;
(3)/C69
XY/C69
X/C69
Y(provided Xand/C89are independent);
(4)/C69
/C104
X P
i/C104
xi/C112i(for a finite distribution).
/C83pecial probabilit/C121 distributions
We now consider some special probability distributions in which we will use all
the things we have learned so far about probability.
/C84he binomial distribution
Before we discuss the binomial distribution, let us introduce a term, the Bernoullitrials. Consider an experiment such as spinning a coin or throw a die repeatedly.
Each spin or toss is called a trial. In any single trial there will be a probability p
associated with a particular event (or outcome). If pis constant throughout (that
is, does not change from one trial to the next), such trials are then said to be
independent and are known as Bernoulli trials.
Now suppose that we have nindependent events of some kind (such as tossing a
coin or die), each of which has a probability pof success and probability of
/C113
1ÿ/C112of failure. What is the probability that exactly mof the events will
succeed/C63 If we select mevents from n, the probability that these mwill succeed and
all the rest
nÿmwill fail is /C112
m/C113nÿm. We have considered only one particular
group or combination of mevents. How many combinations of mevents can be
chosen from n/C63 It is the number of combinations of nthings taken mat a time:
nCm. Thus the probability that exactly mevents will succeed from a group of nis
f
mP
Xm nCm/C112m/C113
nÿm
n/C33
m/C33
nÿm/C33/C112m/C113
nÿm:
14:16
This discrete probability function (14.16) is called the binomial distribution for X,
the random variable of the number of successes in the ntrials. It gives the prob-
ability of exactly msuccesses in nindependent trials with constant probability p.
Since many statistical studies involve repeated trials, the binomial distribution has
great practical importance.
491SPECIAL PROBABILITY DISTRIBUTIONS
Why is the discrete probability function (14.16) called the binomial distribu-
tion/C63 Since for m0;1;2;...;nit corresponds to successive terms in the binomial
expansion
/C113/C112n/C113nnC1/C113nÿ1/C112nC2/C113nÿ2/C1122 /C112nXn
m0nCm/C112m/C113nÿm:
To illustrate the use of the binomial distribution (14.16), let us find the prob-
ability that a one will appear exactly 4 times if a die is thrown 10 times. Here
n10,m4,/C1121=6, and /C113
1ÿ/C1125=6. Hence the probability is
f
4P
X410/C33
4/C336/C331
6456
6
0:0543:
A few examples of binomial distributions, computed from Eq. (14.16), are
shown in Figs. 14.2, and 14.3 by means of histograms.
One of the key requirements for a probability distribution is that
Xn
m/C111f
mXn
m/C111nCm/C112m/C113nÿm1:
14:17
To show that this is in fact the case, we note that
Xn
m/C111nCm/C112m/C113nÿm
492INTRODUCTION TO PROBABILITY THEORY
Figure 14.2. The distribution is symmetric about m10:
is exactly equal to the binomial expansion of
/C113/C112n. But here /C113/C1121, so
/C113/C112n1 and our proof is established.
The mean (or average) number of successes, m, is given by
mXn
m0mnCm/C112m
1ÿ/C112nÿm:
14:18
The sum ranges from m0t onbecause in every one of the sets of trials the same
number of successes between 0 and nmust occur. It is similar to Eq. (14.17); the
di/C128erence is that the sum in Eq. (14.18) contains an extra factor n. But we can
convert it into the form of the sum in Eq. (14.17). Di/C128erentiating both sides of Eq.
(14.17) with respect to p, which is legitimate as the equation is true for all p
between 0 and 1, gives
X
nCmm/C112mÿ1
1ÿ/C112nÿmÿ
nÿm/C112m
1ÿ/C112nÿmÿ10;
where we have dropped the limits on the sum, remembering that mranges from 0
ton. The last equation can be rewritten as
X
mnCm/C112mÿ1
1ÿ/C112nÿmX
nÿmnCm/C112m
1ÿ/C112nÿmÿ1
nX
nCm/C112m
1ÿ/C112nÿmÿ1ÿX
mnCm/C112m
1ÿ/C112nÿmÿ1
or
X
mnCm/C112mÿ1
1ÿ/C112nÿm/C112m
1ÿ/C112nÿmÿ1nX
nCm/C112m
1ÿ/C112nÿmÿ1:
493SPECIAL PROBABILITY DISTRIBUTIONS
Figure 14.3. The distribution favors smaller value of m.
Now multiplying both sides by /C112
1ÿ/C112we get
X
mnCm
1ÿ/C112/C112m
1ÿ/C112nÿm/C112m1
1ÿ/C112nÿmn/C112X
nCm/C112m
1ÿ/C112nÿm:
Combining the two terms on the left hand side, and using Eq. (14.17) in the right
hand side we have
X
mnCm/C112m
1ÿ/C112nÿmX
mf
mn/C112:
14:19
Note that the left hand side is just our original expression for m, Eq. (14.18). Thus
we conclude that
mn/C112
14:20
for the binomial distribution.
The variance 2is given by
2X
mÿm2f
mX
mÿn/C1122f
m;
14:21
here we again drop the summation limits for convenience. To evaluate this sumwe first rewrite Eq. (14.21) as
2X
m2ÿ2mn/C112n2/C1122f
m
X
m2f
mÿ2n/C112X
mf
mn2/C1122X
f
m:
This reduces to, with the help of Eqs. (14.17) and (14.19),
2X
m2f
mÿ
n/C1122:
14:22
To evaluate the first term on the right hand side, we first di/C128erentiate Eq. (14.19):
X
mnCmm/C112mÿ1
1ÿ/C112nÿmÿ
nÿm/C112m
1ÿ/C112nÿmÿ1/C112;
then multiplying by /C112
1ÿ/C112and rearranging terms as before
X
m2
nCm/C112m
1ÿ/C112nÿmÿn/C112X
mnCm/C112m
1ÿ/C112nÿmn/C112
1ÿ/C112:
By using Eq. (14.19) we can simplify the second term on the left hand side andobtain
X
m
2
nCm/C112m
1ÿ/C112nÿm
n/C1122n/C112
1ÿ/C112
or
X
m2f
mn/C112
1ÿ/C112n/C112:
Inserting this result back into Eq. (14.22), we obtain
2n/C112
1ÿ/C112n/C112ÿ
n/C1122n/C112
1ÿ/C112n/C112/C113;
14:23
494INTRODUCTION TO PROBABILITY THEORY
and the standard deviation ;
n/C112/C113p:
14:24
Two di/C128erent limits of the binomial distribution for large nare of practical
importance: (1) n!1 and/C112!0 in such a way that the product n/C112remains
constant; (2) both nandpnare large. The first case will result a new distribution,
the Poisson distribution, and the second cases gives us the Gaussian (or Laplace)
distribution.
/C84he Poisson distribution
Now n/C112;so/C112=n. The binomial distribution (14.16) then becomes
f
mP
Xmn/C33
m/C33
nÿm/C33
nm
1ÿ
nnÿm
n
nÿ1
nÿ2
nÿm1
m/C33nmm1ÿ
nnÿm
1ÿ1
n
1ÿ2n
1ÿmÿ1
nm
m/C331ÿ
nnÿm
:
14:25
Now as n!1 ,
1ÿ1n
1ÿ2n
1ÿmÿ1
n
!1;
while
1ÿ
nnÿm
1ÿ
nn
1ÿ
nÿm
!eÿÿ
1
eÿ;
where we have made use of the result
lim
n!11
nn
e:
It follows that Eq. (14.25) becomes
f
mP
Xmmeÿ
m/C33:
14:26
This is known as the Poisson distribution. Note thatP1
m0P
Xm1;as it
should.
495SPECIAL PROBABILITY DISTRIBUTIONS
The Poisson distribution has the mean
/C69
XX1
m0mmeÿ
m/C33X1
m1meÿ
mÿ1/C33X1
m0meÿ
m/C33
eÿX1
m0m
m/C33eÿe;
14:27
where we have made use of the result
X1
m0m
m/C33e:
The variance 2of the Poisson distribution is
2Var
X/C69
Xÿ/C69
X2/C69
X2ÿ/C69
X2
X1
m0m2meÿ
m/C33ÿ2eÿX1
m1mm
mÿ1/C33ÿ2
eÿd
deÿ
ÿ2:
14:28
To illustrate the use of the Poisson distribution, let us consider a simple exam-
ple. Suppose the probability that an individual su/C128ers a bad reaction from a flu
injection is 0.001; what is the probability that out of 2000 individuals ( a) exactly 3,
(b) more than 2 individuals will su/C128er a bad reaction/C63 Now Xdenotes the number
of individuals who su/C128er a bad reaction and it is binomially distributed. However,
we can use the Poisson approximation, because the bad reactions are assumed to
be rare events. Thus
P
Xmmeÿ
m/C33;with m/C112
2000
0:0012/C58
(a)P
X323eÿ2
3/C330:18;
bP
X/C6221ÿP
X0P
X1P
X2
1ÿ20eÿ2
0/C3321eÿ2
1/C3322eÿ2
2/C33"#
1ÿ5eÿ20:323:
An exact evaluation of the probabilities using the binomial distribution would
require much more labor.
496INTRODUCTION TO PROBABILITY THEORY
The Poisson distribution is very important in nuclear physics. Suppose that we
have nradioactive nuclei and the probability for any one of these to decay in a
given interval of time Tisp, then the probability that mnuclei will decay in the
interval Tis given by the binomial distribution. However, nmay be a very large
number (such as 1023), and pmay be the order of 10ÿ20, and it is impractical to
evaluate the binomial distribution with numbers of these magnitudes.
Fortunately, the Poisson distribution can come to our rescue.
The Poisson distribution has its own significance beyond its connection with the
binomial distribution and it can be derived mathematically from elementary con-
siderations. In general, the Poisson distribution applies when a very large number
of experiments is carried out, but the probability of success in each is very small,so that the expected number of successes is a finite number.
/C84he Gaussian /C40or normal/C41 distribution
The second limit of the binomial distribution that is of interest to us results when
both nandpnare large. Clearly, we assume that m,n,a n d nÿmare large enough
to permit the use of Stirling’s formula ( n/C33/C25
2np
n
neÿn). Replacing m/C33,n/C33, and
(nÿm)/C33 by their approximations and after simplification, we obtain
P
Xm/C129n/C112
mmn/C113
nÿmnÿmn
2m
nÿm/C114
:
14:29
The binomial distribution has the mean value np(see Eq. (14.20). Now let
denote the deviation of mfrom np; that is, mÿn/C112. Then nÿmn/C113ÿ;
and Eq. (14.29) becomes
P
Xm1
2n/C112/C1131=n/C112
1ÿ=n/C112
/C112 1
n/C112ÿ
n/C112
1ÿ
n/C113ÿ
n/C113
or
P
XmA1
n/C112ÿ
n/C112
1ÿ
n/C113ÿ
n/C113ÿ
;
where
A
2n/C112/C113 1
n/C112
1ÿ
n/C113/C115
:
Then
logP
XmA
/C129 ÿ
n/C112log 1 =n/C112
ÿ
n/C113ÿlog
1ÿ=n/C113:
497SPECIAL PROBABILITY DISTRIBUTIONS
Assuming jj<n/C112/C113, so that =n/C112jj <1 and =n/C113jj <1, this permits us to write the
two convergent series
log 1
n/C112
n/C112ÿ2
2n2/C11223
3n3/C1123ÿ ;
log 1 ÿ
n/C113
ÿ
n/C113ÿ2
2n2/C1132ÿ3
3n3/C1133ÿ :
Hence
logP
XmA
/C129 ÿ2
2n/C112/C113ÿ3
/C1122ÿ/C1132
23n2/C1122/C1132ÿ4
/C1123/C1133
34n3/C1123/C1133ÿ :
Now, if jjis so small in comparison with np/C113that we ignore all but the first term
on the right hand side of this expansion and Acan be replaced by
2n/C112/C1131=2, then
we get the approximation formula
P
Xm12n/C112/C113p eÿ2=2n/C112/C113:
14:30
When n/C112/C113p;Eq. (14.30) becomes
f
mP
Xm1
2p
eÿ2=22:
14:31
This is called the Guassian, or normal, distribution. It is a very good approxima-
tion even for quite small values of n.
The Gaussian distribution is a symmetrical bell-shaped distribution about its
mean /C22, and is a measure of the width of the distribution. Fig. 14.4 gives a
comparison of the binomial distribution and the Gaussian approximation.
The Gaussian distribution also has a significance far beyond its connection with
the binomial distribution. It can be derived mathematically from elementary con-
siderations, and is found to agree empirically with random errors that actually
498INTRODUCTION TO PROBABILITY THEORY
Figure 14.4.
occur in experiments. Everyone believes in the Gaussian distribution: mathe-
maticians think that physicists have verified it experimentally and physiciststhink that mathematicians have proved it theoretically.
One of the main uses of the Gaussian distribution is to compute the probability
X
m2
mm1f
m
that the number of successes is between the given limits m1andm2. Eq. (14.31)
shows that the above sum may be approximated by a sum
X 1
2p
eÿ2=22
14:32
over appropriate values of . Since mÿn/C112, the di/C128erence between successive
values of is 1, and hence if we let z=, the di/C128erence between successive
values of zisz1=. Thus Eq. (14.32) becomes the sum over z,
X 1 2peÿz2=2z:
14:33
Asz!0, the expression (14.33) approaches an integral, which may be evalu-
ated in terms of the function
zZz
01 2peÿz2=2dz1 2pZ
z
0eÿz2=2dz:
14:34
The function
zis related to the extensively tabulated error function, erf( z):
erf
z2pZz
0eÿz2dz;and
z1
2erfz
2p
:
These considerations lead to the following important theorem, which we state
without proof: If mis the number of successes in nindependent trials with con-
stant probability p, the probability of the inequality
z1mÿn/C112n/C112/C113p z2
14:35
approaches the limit
1
2pZz2
z1eÿz2=2dz
z2ÿ
z1
14:36
asn!1 . This theorem is known as Laplace–de Moivre limit theorem.
To illustrate the use of the result (14.36), let us consider the simple example of a
die tossed 600 times, and ask what the probability is that the number of ones will
499SPECIAL PROBABILITY DISTRIBUTIONS
be between 80 and 110. Now n600, /C1121=6,/C1131ÿ/C1125=6, and mvaries
from 80 to 110. Hence
z180ÿ100
100
5=6/C112 ÿ2:19 and z1110ÿ100 100
5=6/C112 1:09:
The tabulated error function gives
z
2
1:090:362;
and
z1
ÿ2:19ÿ
2:19ÿ 0:486;
where we have made use of the fact that
ÿzÿ
z;you can check this with
Eq. (14.34). So the required probability is approximately given by
0:362ÿ
ÿ 0:4860:848:
/C67ontinuous distributions
So far we have discussed several discrete probability distributions: since measure-
ments are generally made only to a certain number of significant figures, the
variables that arise as the result of an experiment are discrete. However, discretevariables can be approximated by continuous ones within the experimental error.
Also, in some applications a discrete random variable is inappropriate. We now
give a brief discussion of continuous variables that will be denoted by x. We shall
see that continuous variables are easier to handle analytically.
Suppose we want to choose a point randomly on the interval 0 x1, how
shall we measure the probabilities associated with that event/C63 Let us dividethis interval
0;1into a number of subintervals, each of length x0:1 (Fig.
14.5), the point xis then equally likely to be in any of these subintervals. The
probability that 0 :3<x<0:6, for example, is 0.3, as there are three favorable
cases. The probability that 0 :32<x<0:64 is found to be 0 :64ÿ0:320:32
when the interval is divided into 100 parts, and so on. From these we see thatthe probability for xto be in a given subinterval of (0, 1) is the length of that
subinterval. Thus
P
a<x<bbÿa; 0ab1:
14:37
500INTRODUCTION TO PROBABILITY THEORY
Figure 14.5.
The variable xis said to be uniformly distributed on the interval 0 x1.
Expression (14.37) can be rewritten as
P
a<x<bZb
adxZb
a1dx:
For a continuous variable it is customary to speak of the probability density,
which in the above case is unity. More generally, a variable may be distributed
with an arbitrary density f
x. Then the expression
f
zdz
measures approximately the probability that xis on the interval
z<x<zdz:
And the probability that xis on a given interval ( a;b)i s
P
a<x<bZb
af
xdx
14:38
as shown in Fig. 14.6.
The function f
xis called the probability density function and has the proper-
ties:
(1)f
x0;
ÿ1 <x<1;
(2)Z1
ÿ1f
xdx1;a real-valued random variable must lie between 1.
The function
/C70
xP
XxZx
ÿ1f
udu
14:39
defines the probability that the continuous random variable Xis in the interval
(ÿ1;x, and is called the cumulative distributive function. If f
xis continuous,
then Eq. (14.39) gives
/C700
xf
x
and we may speak of a probability di/C128erential d/C70
xf
xdx:
501CONTINUOUS DISTRIBUTIONS
Figure 14.6.
By analogy with those for discrete random variables the expected value or mean
and the variance of a continuous random variable Xwith probability
density function f
xare defined, respectively, to be:
/C69
X/C22Z1
ÿ1xf
xdx;
14:40
Var
X2/C69
Xÿ/C222Z1
ÿ1
xÿ/C222f
xdx:
14:41
/C84he Gaussian /C40or normal/C41 distribution
One of the most important examples of a continuous probability distribution is
the Gaussian (or normal) distribution. The density function for this distribution is
given by
f
x1
2peÿ
xÿ/C222=22; ÿ1 <x<1;
14:42
where /C22andare the mean and standard deviation, respectively. The correspond-
ing distribution function is
/C70
xP
Xx1
2pZ
x
ÿ1eÿ
uÿ/C222=22du:
14:43
The standard normal distribution has mean zero
/C220and standard devia-
tion ( 1)
f
z1 2peÿz2=2:
14:44
Any normal distribution can be ‘standardized’ by considering the substitution
z
xÿ/C22=in Eqs. (14.42) and (14.43). A graph of the density function (14.44),
known as the standard normal curve, is shown in Fig. 14.7. We have also indi-
cated the areas within 1, 2 and 3 standard deviations of the mean (that is between
zÿ1 and 1,ÿ2 and 2,ÿ3a n d 3):
P
ÿ1Z11
2pZ1
ÿ1eÿz2=2dz0:6827 ;
P
ÿ2Z21 2pZ
2
ÿ2eÿz2=2dz0:9545 ;
P
ÿ3Z31 2pZ
3
ÿ3eÿz2=2dz0:9973:
502INTRODUCTION TO PROBABILITY THEORY
The above three definite integrals can be evaluated by making numerical approx-
imations. A short table of the values of the integral
/C70
x1
2pZx
0eÿt2dt1
21
2pZx
ÿxeÿt2dt
is included in Appendix 3. A more complete table can be found in Tables of
/C78ormal Probability /C70unctions , National Bureau of Standards, Washington, DC,
1953.
/C84he /C77ax/C119ell/C177Bolt/C122mann distribution
Another continuous distribution that is very important in physics is the Maxwell–
Boltzmann distribution
f
x4aa
/C114
x2eÿax2;0x<1;a/C620;
14:45
where am=2kT,mis the mass, Tis the temperature (K), kis the Boltzmann
constant, and xis the speed of a gas molecule.
Problems
14.1 If a pair of dice is rolled what is the probability that a total of 8 shows/C63
14.2 Four coins are tossed, and we are interested in the number of heads. What
is the probability that there is an odd number of heads/C63 What is the prob-ability that the third coin will land heads/C63
503PROBLEMS
Figure 14.7.
14.3 Two coins are tossed. A reliable witness tells us ‘at least 1 coin showed
heads.’ What e/C128ect does this have on the uniform sample space/C63
14.4 The tossing of two coins can be described by the following sample space:
Event no heads one head two head
Probability 1/4 1/2 1/4
What happens to this sample space if we know at least one coin showed
heads but have no other specific information/C63
14.5 Two dice are rolled. What are the elements of the sample space/C63 What is the
probability that a total of 8 shows/C63 What is the probability that at least one
5 shows/C63
14.6 A vessel contains 30 black balls and 20 white balls. Find the probability of
drawing a white ball and a black ball in succession from the vessel.
14.7 Find the number of di/C128erent arrangements or permutations consisting of
three letters each which can be formed from the seven letters A/C44 B/C44 /C67/C44 /C68/C44 E/C44
/C70/C44 /C71.
14.8 It is required to sit five boys and four girls in a row so that the girls occupy
the even seats. How many such arrangements are possible/C63
14.9 A balanced coin is tossed five times. What is the probability of obtaining
three heads and two tails/C63
14.10 How many di/C128erent five-card hands can be dealt from a shu/C130ed deck of 52
cards/C63 What is the probability that a hand dealt at random consists of fivespades/C63
14.11 ( a) Find the constant term in the expansion of ( x
21=x12:
(b) Evaluate 50/C33.
14.12 A box contains six apples of which two are spoiled. Apples are selected at
random without replacement until a spoiled one is found. Find the
probability distribution of the number of apples drawn from the box,
and present this distribution graphically.
14.13 A fair coin is tossed six times. What is the probability of getting exactly two
heads/C63
14.14 Suppose three dice are rolled simultaneously. What is the probability that
two 5 sappear with the third face showing a di/C128erent number/C63
14.15 Verify thatP1
m0P
Xm1 for the Poisson distribution.
14.16 Certain processors are known to have a failure rate of 1.2/C37. There are
shipped in batches of 150. What is the probability that a batch has exactly
one defective processor/C63 What is the probability that it has two/C63
14.17 A Geiger counter is used to count the arrival of radioactive particles. Find:
(a) the probability that in time tno particles will be counted;
(b) the probability of exactly one count in time t.
504INTRODUCTION TO PROBABILITY THEORY
14.18 Given the density function f
x
f
xkx20<x<3
0 otherwise/C58(
(a) find the constant k;
(b) compute P
1<x<2;
(c) find the distribution function and use it to find P
1<x
2:
505PROBLEMS
Appendix 1
Preliminaries (review of
fundamental concepts)
This appendix is for those readers who need a review; a number of fundamental
concepts or theorem will be reviewed without giving proofs or attempting to
achieve completeness.
We assume that the reader is already familiar with the classes of real
numbers used in analysis. The set of positive integers (also known as natural
numbers) 1, 2, ...;nadmits the operations of addition without restriction, that
is, they can be added (and therefore multiplied) together to give other positiveintegers. The set of integers 0,1;2;...;nadmits the operations of addition
and subtraction among themselves. /C82ational numbers are numbers of the form
/C112=/C113, where pand/C113are integers and /C11360. Examples of rational numbers are 2/3,
ÿ10=7. This set admits the further property of division among its members. The
set of irrational numbers includes all numbers which cannot be expressed as the
quotient of two integers. Examples of irrational numbers are
2p
;11
3p
;and any
number of the form
a=bn/C112
, where aand bare integers which are perfect nth
powers.
The set of real numbers contains all the rationals and irrationals. The important
property of the set of real numbers fxgis that it can be put into (1:1) cor-
respondence with the set of points fPgof a line as indicated in Fig. A.1.
The basic rules governing the combinations of real numbers are:
commutative law: abba;abba;
associative law: a
bc
abc;a
bc
abc;
distributive law: a
bcabac;
index law amanamn;am=anamÿn
a60;
where a;b;c, are algebraic symbols for the real numbers.
Problem A1.1
Prove that
2p
is an irrational number.
506
(Hint: Assume the contrary, that is, assume that
2p
/C112=/C113, where pand /C113are
positive integers having no common integer factor.)
Inequalities
Ifxandyare real numbers, x/C62ymeans that xis greater than y;a n d x<ymeans
that xis less than y. Similarly, xyimplies that xis either greater than or equal
toy. The following basic rules governing the operations with inequalities:
(1) Multiplication by a constant: If x/C62y, then ax/C62ayifais a positive num-
ber, and ax<ayifais a negative number.
(2) Addition of inequalities: If x;y;u;/C118are real numbers, and if x/C62y, and
u/C62/C118, than xu/C62y/C118.
(3) Subtraction of inequalities: If x/C62y,a n d u/C62/C118, we cannot deduce that
xÿu/C62
yÿ/C118. Why/C63 It is evident that
xÿuÿ
yÿ/C118
xÿyÿ
uÿ/C118is not necessarily positive.
(4) Multiplication of inequalities: If x/C62y, and u/C62/C118, and x;y;u;/C118areallposi-
tive, then xu/C62y/C118. When some of the numbers are negative, then the result is
not necessarily true.
(5) Division of inequalities: x/C62yandu/C62/C118do not imply x=u/C62y=/C118.
When we wish to consider the numerical value of the variable xwithout regard
to its sign, we write jxjand read this as ‘absolute or mod x’. Thus the inequality
jxjais equivalent to ax a.
Problem A1.2
Find the values of xwhich satisfy the following inequalities:
(a)x3ÿ7x221xÿ27/C620,
(b)j7ÿ3xj<2,
(c)5
5xÿ1/C622
2x1. (Warning: cross multiplying is not permitted.)
Problem A1.3
Ifa1;a2;...;anand b1;b2;...;bnare any real numbers, prove Schwarz’s
inequality:
a1b1a2b2 anbn2
a2
1a22 ann
b21b22 bnn:
507INEQUALITIES
Figure A1.1.
Problem A1.4
Show that
1
214181
2nÿ11 for all positive integers n/C621:
Ifx1;x2;...;xnarenpositive numbers, their arithmetic mean is defined by
A1nX
n
k1xkx1x2 xn
n
and their geometric mean by
/C71n/C89n
k1xk/C115
x1x2xnnp;
wherePand/C81are the summation and product signs. The harmonic mean /C72is
sometimes useful and it is defined by
1
H1
nXn
k11
xk1n1
x11
x21
xn
:
There is a basic inequality among the three means: A/C71H, the equality sign
occurring when x1x2 xn.
Problem A1.5
Ifx1andx2are two positive numbers, show that A/C71H.
Functions
We assume that the reader is familiar with the concept of functions and theprocess of graphing functions.
A polynomial of degree nis a function of the form
f
x/C112
n
xa0xna1xnÿ1a2xnÿ2 an
ajconstant ;a060:
A polynomial can be di/C128erentiated and integrated. Although we have written
ajconstant, they might still be functions of some other variable independent
ofx. For example,
tÿ3x3sintx2
tp
xt
is a polynomial function of x(of degree 3) and each of the as is a function of a
certain variable t:a0tÿ3;a1sint;a2t1=2;a3t.
The polynomial equation f
x0 has exactly nroots provided we count repe-
titions. For example, x3ÿ3x23xÿ10 can be written
xÿ130 so that
the three roots are 1, 1, 1. Note that here we have used the binomial theorem
axnannanÿ1xn
nÿ1
2/C33anÿ2x2 xn:
508APPENDI/C88 1 PRELIMINARIES
A rational function is of the form f
x/C112n
x=/C113n
x, where /C112n
xand/C113n
x
are polynomials.
A transcendental function is any function which is not algebraic, for example,
the trigonometric functions sin x, cos x, etc., the exponential functions ex, the
logarithmic functions log x, and the hyperbolic functions sinh x, cosh x, etc.
The exponential functions obey the index law. The logarithmic functions are
inverses of the exponential functions, that is, if axythen xlogay, where ais
called the base of the logarithm. If ae, which is often called the natural base of
logarithms, we denote logexby ln x, called the natural logarithm of x. The funda-
mental rules obeyed by logarithms are
ln
mnlnmlnn;ln
m=nlnmÿlnn;and ln m/C112/C112lnm:
The hyperbolic functions are defined in terms of exponential functions as
follows
sinhxexÿeÿx
2; coshxexeÿx
2;
tanhxsinhx
coshxexÿeÿx
exex; cothx1
tanhxexeÿx
exÿeÿx;
sechx1
coshx2
exeÿx; cosech x1
sinhx2
exÿeÿx:
Rough graphs of these six functions are given in Fig. A1.2.
Some fundamental relationships among these functions are as follows:
cosh2xÿsinh2x1;sech2xtanh2x1;coth2xÿcosech2x1;
sinh
xysinhxcoshycoshxsinhy;
cosh
xycoshxcoshysinhxsinhy;
tanh
xytanhxtanhy
1tanhxtanhy:
509FUNCTIONS
Figure A1.2. Hyperbolic functions.
Problem A1.6
Using the rules of exponents, prove that ln
mnlnmlnn:
Problem A1.7
Prove that:
asin2x1
2
1ÿcos 2x;cos2x12
1cos 2x, and ( b)Acosx
Bsinx
A2B2p
sin
x, where tan A=B
Problem A1.8
Prove that:
acosh2xÿsinh2x1, and ( b)2xtanh2x1.
Limits
We are sometimes required to find the limit of a function f
xasxapproaches
some particular value :
lim
x!f
x/C108:
This means that if jxÿjis small enough, jf
xÿ/C108jcan be made as small as we
please. A more precise analytic description of lim x!f
x/C108is the following:
For any /C34/C620 (however small) we can always find a number /C17
(which, in general, depends upon /C34) such that f
xÿ/C108 jj </C34
whenever xÿjj </C17.
As an example, consider the limit of the simple function f
x2ÿ1=
xÿ1as
x!2. Then
lim
x!2f
x1
for if we are given a number, say /C3410ÿ3, we can always find a number /C17which is
such that
2ÿ1
xÿ1
ÿ1<10ÿ3
A1:1
provided jxÿ2j</C17. In this case (A1.1) will be true if 1 =
xÿ1/C621ÿ10ÿ3
0:999. This requires xÿ1<
0:999ÿ1,o rxÿ2<
0:999ÿ1ÿ1. Thus we need
only take /C17
0:999ÿ1ÿ1.
The function f
xis said to be continuous at if lim x!f
x/C108.I ff
xis con-
tinuous at each point ofan interval such as axbora<xb, etc., it is said to be
continuous in the interval (for example, a polynomial is continuous at all x).
The definition implies that lim x!ÿ0f
xlimx!0f
xf
at all points
of the interval ( a;b), but this is clearly inapplicable at the endpoints aandb.A t
these points we define continuity by
lim
x!a0f
xf
aand lim
x!bÿ0f
xf
b:
510APPENDI/C88 1 PRELIMINARIES
A finite discontinuity may occur at x. This will arise when
limx!ÿ0f
x/C1081, lim x!ÿ0f
x/C1082, and /C10816/C1082.
It is obvious that a continuous function will be bounded in any finite interval.
This means that we can find numbers mandMindependent of xand such that
mf
xMforaxb. Furthermore, we expect to find x0;x1such that
f
x0mandf
x1M.
The order of magnitude of a function is indicated in terms of its variable. Thus,
ifxis very small, and if f
xa1xa2x2a3x3 (akconstant), its magni-
tude is governed by the term in xand we write f
xO
x. When a10, we
write f
xO
x2, etc. When f
xO
xn, then lim x!0ff
x=xngis finite and/
or lim x!0ff
x=xnÿ1g0.
A function f
xis said to be di/C128erentiable or to possess a derivative at the point
xif lim /C104!0f
x/C104ÿf
x=/C104exists. We write this limit in various forms
df=dx;f0orDf, where Dd
=dx. Most of the functions in physics can be
successively di/C128erentiated a number of times. These successive derivatives are
written as f0
x;f00
x;...;fn
x;...;orDf;D2f;...;Dnf;...:
Problem A1.9
Iff
xx2, prove that: ( a) lim x!2f
x4, and
bf
xis continuous at x2.
Infinite series
Infinite series involve the notion of sequence in a simple way. For example,
2p
is
irrational and can only be expressed as a non-recurring decimal 1 :414 ...:We can
approximate to its value by a sequence of rationals, 1, 1.4, 1.41, 1.414, ...sayfang
which is a countable set limit of anwhose values approach indefinitely close to2p
.
Because of this we say the limit of a
nasntends to infinity exists and equals2p
,
and write lim
n!1an2p
.
In general, a sequence u
1;u2;...;fungis a function defined on the set of natural
numbers. The sequence is said to have the limit lor to converge to l, if given any
/C34/C620 there exists a number /C78/C620 such that junÿ/C108j</C34for all n/C62/C78, and in such
case we write lim n!1un/C108.
Consider now the sums of the sequence fung
snXn
r1uru1u2u3 ;
A:2
where ur/C620 for all r.I fn!1 , then (A.2) is an infinite series of positive terms.
We see that the behavior of this series is determined by the behavior of the
sequence fungas it converges or diverges. If lim n!1sns(finite) we say that
(A.2) is convergent and has the sum s. When sn!1 asn!1 , we say that (A.2)
is divergent.
511INFINITE SERIES
Example A1.1.
Show that the series
X1
n11
2n1
21
221
23
is convergent and has sum s1.
Solution: Let
sn121
221
231
2n;
then
12s
n1
221
231
2n1:
Subtraction gives
1ÿ12
s
n12ÿ1
2n1121ÿ1
2n
; or sn1ÿ1
2n:
Then since lim n!1snlimn!1
1ÿ1=2n1, the series is convergent and has
the sum s1.
Example A1.2.
Show that the seriesP1
n1
ÿ1nÿ11ÿ11ÿ1 is divergent.
Solution: Here sn0 or 1 according as nis even or odd. Hence lim n!1sndoes
not exist and so the series is divergent.
Example A1.3.
Show that the geometric seriesP1
n1arnÿ1aarar2 ;where aandrare
constants, ( a) converges to sa=
1ÿrifjrj<1;and ( b) diverges if jrj/C621.
Solution: Let
snaarar2 arnÿ1:
Then
rsn arar2 arnÿ1arn:
Subtraction gives
1ÿrsnaÿarnor sna
1ÿrn
1ÿr:
512APPENDI/C88 1 PRELIMINARIES
(a)I fjrj<1;
lim
n!1snlim
n!1a
1ÿrn
1ÿra
1ÿr:
(b)I fjrj/C621,
lim
n!1snlim
n!1a
1ÿrn
1ÿr
does not exist.
Example A1.4.
Show that the pseriesP1
n11=n/C112converges if /C112/C621 and diverges if /C1121.
Solution: Using f
n1=npwe have f
x1=xpso that if p61,
Z1
1dx
x/C112lim
M!1ZM
1xÿ/C112dxlim
M!1x1ÿ/C112
1ÿ/C112/C12/C12/C12/C12M
1lim
M!1M1ÿ/C112
1ÿ/C112ÿ1
1ÿ/C112"#
:
Now if /C112/C621 this limit exists and the corresponding series converges. But if /C112<1
the limit does not exist and the series diverges.
If/C1121 then
Z1
1dx
xlim
M!1ZM
1dx
xlim
M!1lnx/C12/C12/C12/C12M
1lim
M!1lnM;
which does not exist and so the corresponding series for /C1121 diverges.
This shows that 1 1
213 diverges even though the nth term approaches
zero.
/C84ests for convergence
There are several important tests for convergence of series of positive terms.
Before using these simple tests, we can often weed out some very badly divergent
series with the following preliminary test:
If the terms of an infinite series do not tend to zero (that is, iflim
n!1an60, the series diverges. If lim n!1an0, we must
test further :
Four of the common tests are given below:
/C67omparison test
Ifun/C118n(alln), thenP1
n1unconverges whenP1n1/C118nconverges. If un/C118n(all
n), thenP1
n1undiverges whenP1n1/C118ndiverges.
513INFINITE SERIES
Since the behavior ofP1
n1unis una/C128ected by removing a finite number of
terms from the series, this test is true if un/C118norun/C118nfor all n/C62/C78. Note that
n/C62/C78means from some term onward. Often, /C781.
Example A1.5
(a) Since 1 =
2n11=2nandP1=2nconverges,P1=
2n1also converges.
(b) Since 1 =lnn/C621=nandP1
n21=ndiverges,P1n21=lnnalso diverges.
/C81uotient test
Ifun1=un/C118n1=/C118n(alln), thenP1
n1unconverges whenP1n1/C118nconverges. And
ifun1=un/C118n1=/C118n(alln), thenP1
n1undiverges whenP1n1/C118ndiverges.
We can write
unun
unÿ1unÿ1
unÿ2u2
u1u1/C118n
/C118nÿ1/C118nÿ1
/C118nÿ2/C1182
/C1181/C1181
so that un/C118nu1which proves the quotient test by using the comparison test.
A similar argument shows that if un1=un/C118n1=/C118n(alln), thenP1n1un
diverges whenP1n1/C118ndiverges.
Example A1.6
Consider the series
X1
n14n2ÿn3
n32n:
For large n,
4n2ÿn3=
n32nis approximately 4 =n. Taking
un
4n2ÿn3=
n32nand /C118n1=n, we have lim n!1un=/C118n1. Now
sinceP/C118nP1=ndiverges,Punalso diverges.
/C68/C39Alembert/C39s ratio test:P1
n1unconverges when un1=un<1 (all n/C78) and diverges when un1=un/C621.
Write /C118nxnÿ1in the quotient test so thatP1
n1/C118nis the geometric series with
common ratio /C118n1=/C118nx. Then the quotient test proves thatP1
n1unconverges
when x<1 and diverges when x/C621:
Sometimes the ratio test is stated in the following form: if lim n!1un1=un/C26,
thenP1n1unconverges when /C26<1 and diverges when /C26/C621.
Example A1.7
Consider the series
11
2/C331
3/C331
n/C33 :
514APPENDI/C88 1 PRELIMINARIES
Using the ratio test, we have
un1
un1
n1/C33/C41
n/C33n/C33
n1/C331
n1<1;
so the series converges.
Integral test.
Iff
xis positive, continuous and monotonic decreasing and is such that
f
nunforn/C62/C78, thenPunconverges or diverges according as
Z1
/C78f
xdxlim
M!1ZM
/C78f
xdx
converges or diverges. We often have /C781 in practice.
To prove this test, we will use the following property of definite integrals:
If in axb;f
x/C103
x, thenZb
af
xdxZb
a/C103
xdx.
Now from the monotonicity of f
x, we have
un1f
n1f
xf
nun; n1;2;3;...:
Integrating from xntoxn1 and using the above quoted property of
definite integrals we obtain
un1Zn1
nf
xdxun; n1;2;3;...:
Summing from n1t oMÿ1,
u1u2 uMZM
1f
xdxu1u2 uMÿ1:
A1:3
Iff
xis strictly decreasing, the equality sign in (A1.3) can be omitted.
If lim M!1RM
1f
xdxexists and is equal to s, we see from the left hand inequal-
ity in (A1.3) that u1u2 uMis monotonically increasing and bounded
above by s, so thatPunconverges. If lim M!1RM
1f
xdxis unbounded, we see
from the right hand inequality in (A1.3) thatPundiverges.
Geometrically, u1u2 uMis the total area of the rectangles shown
shaded in Fig. A1.3, while u1u2 uMÿ1is the total area of the rectangles
which are shaded and non-shaded. The area under the curve yf
xfrom x1
toxMis intermediate in value between the two areas given above, thus illus-
trating the result (A1.3).
515INFINITE SERIES
Example A1.8P1
n11=n2converges since lim M!1RM
1dx=x2limM!1
1ÿ1=Mexists.
Problem A1.10
Find the limit of the sequence 0.3, 0.33, 0 :333;...;and justify your conclusion.
/C65lternating series test
An alternating series is one whose successive terms are alternately positive and
negative u1ÿu2u3ÿu4 :It converges if the following two conditions are
satisfied:
(a)jun1junjforn1; (b) lim
n!1un0
or lim
n!1unjj0
:
The sum of the series to 2 Mis
S2M
u1ÿu2
u3ÿu4
u2Mÿ1ÿu2M
u1ÿ
u2ÿu3ÿ
u4ÿu5ÿÿ
u2Mÿ2ÿu2Mÿ1ÿu2M:
Since the quantities in parentheses are non-negative, we have
S2M0; S2S4S6 S2Mu1:
Therefore fS2Mgis a bounded monotonic increasing sequence and thus has the
limit S.
Also S2M1S2Mu2M1. Since lim M!1S2MSand lim M!1u2M10
(for, by hypothesis, lim n!1un0), it follows that lim M!1S2M1
limM!1S2MlimM!1u2M1S0S. Thus the partial sums of the series
approach the limit Sand the series converges.
516APPENDI/C88 1 PRELIMINARIES
Figure A1.3.
Problem A1.11
Show that the error made in stopping after 2Mterms is less than or equal to
u2M1.
Example A1.9For the series
1ÿ1
213ÿ14X
1
n1
ÿ1nÿ1
n;
we have un
ÿ 1n1=n;unjj1=n;un1jj 1=
n1. Then for n1;
un1jj unjj. Also we have lim n!1unjj0. Hence the series converges.
/C65bsolute and conditional convergence
The seriesPunis called absolutely convergent ifPunjjconverges. IfPuncon-
verges butPunjjdiverges, thenPunis said to be conditionally convergent.
It is easy to show that ifPunjjconverges, thenPunconverges (in words, an
absolutely convergent series is convergent). To this purpose, let
SMu1u2 uMTMu1jju2jj uMjj;
then
SMTM
u1u1jj
u2u2jj
uMuMjj
2u1jj2u2jj 2uMjj:
SincePunjjconverges and since ununjj0, for n1;2;3;...;it follows that
SMTM is a bounded monotonic increasing sequence, and so
limM!1
SMTMexists. Also lim M!1TMexists (since the series is absolutely
convergent by hypothesis),
lim
M!1SMlim
M!1
SMTMÿTM lim
M!1
SMTMÿlim
M!1TM
must also exist and so the seriesPunconverges.
The terms of an absolutely convergent series can be rearranged in any order,
and all such rearranged series will converge to the same sum. We refer the reader
to text-books on advanced calculus for proof.
Problem A1.12
Prove that the series
1ÿ1
221
32ÿ1
421
52ÿ
converges.
517INFINITE SERIES
How do we test for absolute convergence/C63 The simplest test is the ratio test,
which we now review, along with three others – Raabe’s test, the nth root test, and
Gauss’ test.
/C82atio test
Let lim n!1jun1=unjL. Then the seriesPun:
(a) converges (absolutely) if L<1;
(b) diverges if L/C621;
(c) the test fails if L1.
Let us consider first the positive-termPun, that is, each term is positive. We
must now prove that if lim n!1un1=unL<1, then necessarilyPunconverges.
By hypothesis, we can choose an integer /C78so large that for all
n/C78;
un1=un<r, where L<r<1. Then
u/C781<ru/C78;u/C782<ru/C781<r2u/C78;u/C783<ru/C782<r3u/C78;etc:
By addition
u/C781u/C782 <u/C78
rr2r3
and so the given series converges by the comparison test, since 0 <r<1.
When the series has terms with mixed signs, we consider u1jju2jju3jj ,
then by the above proof and because an absolutely convergent series isconvergent, it follows that if lim
n!1un1=un jj L<1, thenPunconverges
absolutely.
Similarly we can prove that if lim n!1un1=un jj L/C621, the seriesPun
diverges.
Example A1.10
Consider the seriesP1
n1
ÿ1nÿ12n=n2. Here un
ÿ 1nÿ12n=n2. Then
limn!1un1=un jj limn!12n2=
n122. Since L2/C621, the series diverges.
When the ratio test fails, the following three tests are often very helpful.
/C82aabe/C39s test
Let lim n!1n
1ÿun1=un jj ‘, then the seriesPun:
(a) converges absolutely if ‘<1;
(b) diverges if ‘/C621.
The test fails if ‘1.
The n throot test
Let lim n!1
unjjn/C112
R, then the seriesPun:
518APPENDI/C88 1 PRELIMINARIES
(a) converges absolutely if R<1;
(b) diverges if R/C621.
The test fails if R1.
/C71auss/C39 test
If
un1
un/C12/C12/C12/C12/C12/C12/C12/C121ÿ/C71
ncn
n2;
where jcnj<Pfor all n/C62/C78, then the seriesPun:
(a) converges (absolutely) if /C71/C621;
(b) diverges or converges conditionally if /C711.
Example A1.11
Consider the series 1 2rr22r3r42r5 . The ratio test gives
un1
un/C12/C12/C12/C12/C12/C12/C12/C122rjj;nodd
rjj=2;neven;/C26
which indicates that the ratio test is not applicable. We now try the nth root test:
u
njjn/C112
2rnjjn/C112
2np
rjj;nodd
rnjjn/C112
rjj; neven(
and so lim n!1
unjjn/C112
rjj. Thus if jrj<1 the series converges, and if jrj/C621 the
series diverges.
Example A1.12
Consider the series
1
32
14
362
147
3692
147
3nÿ2
369
3n
:
The ratio test is not applicable, since
lim
n!1un1
un/C12/C12/C12/C12/C12/C12/C12/C12lim
n!1
3n1
3n3/C12/C12/C12/C12/C12/C12/C12/C122
1:
But Raabe’s test gives
lim
n!1n1ÿun1
un/C12/C12/C12/C12/C12/C12/C12/C12
lim
n!1n1ÿ3n1
3n32()
4
3/C621;
and so the series converges.
519INFINITE SERIES
Problem A1.13
Test for convergence the series
1
22
13
242
135
2462
135
2nÿ1
135
2n
:
Hint: Neither the ratio test nor Raabe’s test is applicable (show this). Try Gauss’
test.
/C83eries of functions and uniform con/C118ergence
The series considered so far had the feature that undepended just on n. Thus the
series, if convergent, is represented by just a number. We now consider series
whose terms are functions of x;unun
x. There are many such series of func-
tions. The reader should be familiar with the power series in which the nth term is
a constant times xn:
S
xX1
n0anxn:
A1:4
We can think of all previous cases as power series restricted to x1. In later
sections we shall see Fourier series whose terms involve sines and cosines, and
other series in which the terms may be polynomials or other functions. In this
section we consider power series in x.
The convergence or divergence of a series of functions depends, in general, on
the values of x. With xin place, the partial sum Eq. (A1.2) now becomes a
function of the variable x:
sn
xu1u2
x un
x:
A1:5
as does the series sum. If we define S
xas the limit of the partial sum
S
xlim
n!1sn
xX1
n0un
x;
A1:6
then the series is said to be convergent in the interval /C91 a,b/C93 (that is, axb), if
for each /C34/C620 and each xin /C91a,b/C93 we can find /C78/C620 such that
S
xÿsn
x jj </C34 ; for all n/C78:
A1:7
If/C78depends only on /C34and not on x, the series is called uniformly convergent in
the interval /C91 a,b/C93. This says that for our series to be uniformly convergent, it must
be possible to find a finite /C78so that the remainder of the series after /C78terms,P1
i/C781ui
x, will be less than an arbitrarily small /C34for all xin the given interval.
The domain of convergence (absolute or uniform) of a series is the set of values
ofxfor which the series of functions converges (absolutely or uniformly).
520APPENDI/C88 1 PRELIMINARIES
We deal with power series in xexactly as before. For example, we can use the
ratio test, which now depends on x, to investigate convergence or divergence of a
series:
r
xlim
n!1un1
un/C12/C12/C12/C12/C12/C12/C12/C12lim
n!1/C12/C12/C12/C12an1xn1
anxn/C12/C12/C12/C12xjjlim
n!1an1
an/C12/C12/C12/C12/C12/C12/C12/C12xjjr; rlim
n!1an1
an/C12/C12/C12/C12/C12/C12/C12/C12;
thus the series converges (absolutely) if jxjr<1o r
jxj<R
1
rlim
n!1an
an1/C12/C12/C12/C12/C12/C12/C12/C12
and the domain of convergence is given by R/C58ÿR<x<R. Of course, we need to
modify the above discussion somewhat if the power series does not contain every
power of x.
Example A1.13
For what value of xdoes the seriesP
1
n1xnÿ1=n3nconverge/C63
Solution: Now unxnÿ1=n3n,a n d x60 (if x0 the series converges). We
have
lim
n!1un1
un/C12/C12/C12/C12/C12/C12/C12/C12lim
n!1n
3
n1xjj1
3xjj:
Then the series converges if jxj<3, and diverges if jxj/C623. Ifjxj3, that is,
x3, the test fails.
Ifx3, the series becomesP1
n11=3nwhich diverges. If xÿ3, the series
becomesP1n1
ÿ1nÿ1=3nwhich converges. Then the interval of convergence is
ÿ3x<3. The series diverges outside this interval. Furthermore, the series
converges absolutely for ÿ3<x<3 and converges conditionally at xÿ3.
As for uniform convergence, the most commonly encountered test is the
Weierstrass Mtest:
/C87eierstrass Mtest
If a sequence of positive constants M1;M2;M3;...;can be found such that: ( a)
Mnjun
xjfor all xin some interval /C91 a,b/C93, and ( b)PMnconverges, thenPun
xis uniformly and absolutely convergent in /C91 a,b/C93.
The proof of this common test is direct and simple. SincePMnconverges,
some number /C78exists such that for n/C78,
X1
i/C781Mi</C34 :
521SERIES OF FUNCTIONS AND UNIFORM CONVERGENCE
This follows from the definition of convergence. Then, with Mnjun
xjfor all x
in /C91a,b/C93,
X1
i/C781ui
xjj </C34 :
Hence
S
xÿsn
x jj /C12/C12/C12/C12X1
i/C781ui
x/C12/C12/C12/C12</C34 ; for all n/C78
and by definitionPu
n
xis uniformly convergent in /C91 a,b/C93. Furthermore, since we
have specified absolute values in the statement of the Weierstrass Mtest, the seriesPun
xis also seen to be absolutely convergent.
It should be noted that the Weierstrass Mtest only provides a sucient con-
dition for uniform convergence. A series may be uniformly convergent even when
theMtest is not applicable. The Weierstrass Mtest might mislead the reader to
believe that a uniformly convergent series must be also absolutely convergent, andconversely. In fact, the uniform convergence and absolute convergence are inde-
pendent properties. Neither implies the other.
A somewhat more delicate test for uniform convergence that is especially useful
in analyzing power series is Abel’s test. We now state it without proof.
/C65bel’s test
If
au
n
xanfn
x, andPanA, convergent, and ( b) the functions fn
xare
monotonic fn1
xfn
xand bounded, 0 fn
xMfor all xin /C91a,b/C93, thenPun
xconverges uniformly in /C91 a,b/C93.
Example A1.14
Use the Weierstrass Mtest to investigate the uniform convergence of
aX1
n1cosnx
n4;
bX1
n1xn
n3=2;
cX1
n1sinnx
n:
Solution:
(a) cos
nx=n4/C12/C12/C12/C121=n4Mn. Then sincePMnconverges ( pseries with
/C1124/C621), the series is uniformly and absolutely convergent for all xby
theMtest.
(b) By the ratio test, the series converges in the interval ÿ1x1 (or jxj1).
For all xinjxj1;xn=n3=2/C12/C12/C12/C12/C12/C12xjj
n=n3=21=n3=2. Choosing Mn1=n3=2,
we see thatPMnconverges. So the given series converges uniformly for
jxj1 by the Mtest.
522APPENDI/C88 1 PRELIMINARIES
(c) sin
nx=n=n jj 1=nMn. However,PMndoes not converge. The Mtest
cannot be used in this case and we cannot conclude anything about the
uniform convergence by this test.
A uniformly convergent infinite series of functions has many of the properties
possessed by the sum of finite series of functions. The following three are parti-cularly useful. We state them without proofs.
(1) If the individual terms u
n
xare continuous in /C91 a,b/C93 and ifPun
xcon-
verges uniformly to the sum S
xin /C91a,b/C93, then S
xis continuous in /C91 a,b/C93.
Briefly, this states that a uniformly convergent series of continuous func-tions is a continuous function.
(2) If the individual terms u
n
xare continuous in /C91 a,b/C93 and ifPun
xcon-
verges uniformly to the sum S
xin /C91a,b/C93, then
Zb
aS
xdxX1
n1Zb
aun
xdx
or
Zb
aX1
n1un
xdxX1
n1Zb
aun
xdx:
Briefly, a uniform convergent series of continuous functions can be inte-grated term by term.
(3) If the individual terms u
n
xare continuous and have continuous derivatives
in /C91a,b/C93 and ifPun
xconverges uniformly to the sum S
xwhilePdun
x=dxis uniformly convergent in /C91 a,b/C93, then the derivative of the
series sum S
xequals the sum of the individual term derivatives,
d
dxS
xX1
n1d
dxun
xord
dxX1
n1un
x()
X1
n1d
dxun
x:
Term-by-term integration of a uniformly convergent series requires only con-
tinuity of the individual terms. This condition is almost always met in physicalapplications. Term-by-term integration may also be valid in the absence of uni-
form convergence. On the other hand term-by-term di/C128erentiation of a series is
often not valid because more restrictive conditions must be satisfied.
Problem A1.14
Show that the series
sinx
13sin 2x
23sinnx
n3
is uniformly convergent for ÿx.
523SERIES OF FUNCTIONS AND UNIFORM CONVERGENCE
/C84heorems on po/C119er series
When we are working with power series and the functions they represent, it is very
useful to know the following theorems which we will state without proof. We will
see that, within their interval of convergence, power series can be handled much
like polynomials.
(1) A power series converges uniformly and absolutely in any interval which lies
entirely within its interval of convergence.
(2) A power series can be di/C128erentiated or integrated term by term over any
interval lying entirely within the interval of convergence. Also, the sum of a
convergent power series is continuous in any interval lying entirely within its
interval of convergence.
(3) Two power series can be added or subtracted term by term for each value of
xcommon to their intervals of convergence.
(4) Two power series, for example,P1
n0anxnandP1n0bnxn, can be multiplied
to obtainP1
n0cnxn;where cna0bna1bnÿ1a2bnÿ2 anb0, the
result being, valid for each xwithin the common interval of convergence.
(5) If the power seriesP1n0anxnis divided by the power seriesP1n0bnxn,
where b060, the quotient can be written as a power series which converges
for suciently small values of x.
/C84a/C121lor/C39s e/C120pansion
It is very useful in most applied work to find power series that represent the given
functions. We now review one method of obtaining such series, the Taylor expan-
sion. We assume that our function f
xhas a continuous nth derivative in the
interval /C91 a,b/C93 and that there is a Taylor series for f
xof the form
f
xa0a1
xÿa2
xÿ2a3
xÿ3 an
xÿn ;
A1:8
where lies in the interval /C91 a,b/C93. Di/C128erentiating, we have
f0
xa12a2
xÿ3a3
xÿ2 nan
xÿnÿ1 ;
f00
x2a232a3
xÿ43a4
xÿa2 n
nÿ1an
xÿnÿ2 ;
...
f
n
xn
nÿ1
nÿ21anterms containing powers of
xÿ:
We now put xin each of the above derivatives and obtain
f
a0;f0
a1;f00
2a2;fF
3/C33a3;;f
n
n/C33an;
524APPENDI/C88 1 PRELIMINARIES
where f0
means that f
xhas been di/C128erentiated and then we have put x;
and by f00
we mean that we have found f00
xand then put x, and so on.
Substituting these into (A1.8) we obtain
f
xf
f0
xÿ1
2/C33f00
xÿ21
n/C33f
n
xÿn :
A1:9
This is the Taylor series for f
xabout x. The Maclaurin series for f
xis the
Taylor series about the origin. Putting 0 in (A1.9), we obtain the Maclaurin
series for f
x:
f
xf
0f0
0x1
2/C33f00
0x21
3/C33fF
0x31
n/C33f
n
0xn :
A1:10
Example A1.15
Find the Maclaurin series expansion of the exponential function ex.
Solution: Here f
xex. Di/C128erentiating, we have f
n
01 for all
n;n1;2;3.... Then, by Eq. (A1.10), we have
ex1x1
2/C33x21
3/C33x3 X1
n0xn
n/C33; ÿ1 <x<1:
The following series are frequently employed in practice:
1sinxxÿx3
3/C33x5
5/C33ÿx7
7/C33
ÿ 1nÿ1x2nÿ1
2nÿ1/C33 ; ÿ1 <x<1:
2cosx1ÿx2
2/C33x4
4/C33ÿx6
6/C33
ÿ 1nÿ1x2nÿ2
2nÿ2/C33 ; ÿ1 <x<1:
3ex1xx2
2/C33x3
3/C33xnÿ1
nÿ1/C33 ; ÿ1 <x<1:
4lnj
1xjxÿx2
2x3
3ÿx4
4
ÿ 1nÿ1xn
n ; ÿ1<x1:
51
2ln1x
1ÿx/C12/C12/C12/C12/C12/C12/C12/C12xx3
3x5
5x7
7x2nÿ1
2nÿ1 ; ÿ1<x<1:
6tanÿ1xxÿx3
3x5
5ÿx7
7
ÿ 1nÿ1x2nÿ1
2nÿ1 ; ÿ1x1:
7
1x/C1121/C112x/C112
/C112ÿ1
2/C33x2/C112
/C112ÿ1
/C112ÿn1
n/C33xn :
525TAYLOR’S E/C88PANSION
This is the binomial series: ( a)I fpis a positive integer or zero, the series termi-
nates. ( b)I f/C112/C620 but is not an integer, the series converges absolutely for
ÿ1x1. (c)I f ÿ1</C112<0, the series converges for ÿ1<x1. (d)I f
/C112ÿ1, the series converges for ÿ1<x<1.
Problem A1.16
Obtain the Maclaurin series for sin x(the Taylor series for sin xabout x0).
Problem A1.17Use series methods to obtain the approximate value ofR
1
0
1ÿeÿx=xdx.
We can find the power series of functions other than the most common ones
listed above by the successive di/C128erentiation process given by Eq. (A1.9). Thereare simpler ways to obtain series expansions. We give several useful methods here.
(a) For example to find the series for
x1sinx, we can multiply the series for
sinxby
x1and collect terms:
x1sinx
x1xÿx3
3/C33x5
5/C33ÿ/C32!
xx2ÿx3
3/C33ÿx4
3/C33 :
To find the expansion for excosx, we can multiply the series for exby the series
for cos x:
excosx1xx2
2/C33x3
3/C33/C32!
1ÿx2
2/C33x4
4/C33/C32!
1xx2
2/C33x3
3/C33x4
4/C33
ÿx2
2/C33ÿx3
3/C33ÿx4
2/C332/C33
x4
4/C33
1xÿx3
3ÿx4
6:
Note that in the first example we obtained the desired series by multiplication of
a known series by a polynomial; and in the second example we obtained thedesired series by multiplication of two series.
(b) In some cases, we can find the series by division of two series. For
example, to find the series for tan x, we can divide the series for sinx by the series
for cos x:
526APPENDI/C88 1 PRELIMINARIES
tanxsinx
cosx
1ÿx2
2x4
4/C33/C30
xÿx3
3/C33x5
5/C33
x1
3x32
15x5:
The last step is by long division
x13x
32
15x5
1ÿx2
2/C33x4
4/C33
xÿx3
3/C33x5
5/C33/C115
xÿx3
2/C33x5
4/C33
x3
3ÿx5
30
x3
3ÿx5
6
2x5
15;etc:
Problem A1.18
Find the series expansion for 1 =
1xby long division. Note that the series can
be found by using the binomial series: 1 =
1x
1xÿ1.
(c) In some cases, we can obtain a series expansion by substitution of a poly-
nomial or a series for the variable in another series. As an example, let us find the
series for eÿx2. We can replace xin the series for exbyÿx2and obtain
eÿx21ÿx2
ÿx22
2/C33
ÿx3
3/C33 1ÿx2x4
2/C33ÿx6
3/C33:
Similarly, to find the series for sinxp=xpwe replace xin the series for sin xbyxpand obtain
sinxp
xp 1ÿx
3/C33x2
5/C33; x/C620:
Problem A1.19
Find the series expansion for etanx.
Problem A1.20
Assuming the power series for exholds for complex numbers, show that
eixcosxisinx:
527TAYLOR’S E/C88PANSION
(d) Find the series for tanÿ1x(arc tan x). We can find the series by the succes-
sive di/C128erentiation process. But it is very tedious to find successive derivatives of
tanÿ1x. We can take advantage of the following integration
Zx
0dt
1t2tanÿ1t/C12/C12/C12/C12x
0tanÿ1x:
We now first write out
1t2ÿ1as a binomial series and then integrate term by
term:
Zx
0dt
1t2Zx
01ÿt2t4ÿt6ÿ
dttÿt3
3t5
5ÿt7
7/C12/C12/C12/C12x
0:
Thus, we have
tanÿ1xxÿx3
3x5
5ÿx7
7 :
(e) Find the series for ln xabout x1. We want a series of powers ( xÿ1)
rather than powers of x. We first write
lnxln1
xÿ1
and then use the series ln
1xwith xreplaced by ( xÿ1):
lnxln1
xÿ1
xÿ1ÿ1
2
xÿ1213
xÿ1
3ÿ14
xÿ1
4:
Problem A1.21
Expand cos xabout x3=2.
/C72igher deri/C118ati/C118es and Leibnit/C122/C39s formula for nth deri/C118ati/C118e of a product
Higher derivatives of a function yf
xwith respect to xare written as
d2y
dx2d
dxdy
dx
;d3y
dx3d
dxd2y
dx2/C32!
; ...;dny
dxnd
dydnÿ1y
dxnÿ1/C32!
:
These are sometimes abbreviated to either
f00
x;fF
x;...;f
n
xorD2y;D3y;...;Dny
where Dd=dx.
When higher derivatives of a product of two functions f
xand /C103
xare
required, we can proceed as follows:
D
f/C103fD/C103/C103Df
528APPENDI/C88 1 PRELIMINARIES
and
D2
f/C103D
fD/C103/C103DffD2/C1032DfD/C103D2/C103:
Similarly we obtain
D3
f/C103fD3/C1033DfD2/C1033D2fD/C103/C103D3f;
D4
f/C103fD4/C1034DfD3/C1036D2fD2/C1034D3D/C103/C103D4/C103;
and so on. By inspection of these results the following formula (due to Leibnitz)
may be written down for nth derivative of the product fg:
Dn
f/C103f
Dn/C103n
Df
Dnÿ1/C103n
nÿ1
2/C33
D2f
Dnÿ2/C103
n/C33
k/C33
nÿk/C33
Dkf
Dnÿkg
Dnf/C103:
Example A1.16
Iff1ÿx2;/C103D2y, where yis a function of x, say u
x, then
Dnf
1ÿx2D2yg
1ÿx2Dn2yÿ2nxDn1yÿn
nÿ1Dny:
Leibnitz’s formula may also be applied to a di/C128erential equation. For example,
ysatisfies the di/C128erential equation
D2yx2ysinx:
Then di/C128erentiating each term ntimes we obtain
Dn2y
x2Dny2nxDnÿ1yn
nÿ1Dnÿ2ysinn
2x
;
where we have used Leibnitz’s formula for the product term x2y.
Problem A1.22Using Leibnitz’s formula, show that
D
n
x2sinxf x2ÿn
nÿ1gsin
xn=2ÿ2nxcos
xn=2:
/C83ome important properties of definite integrals
Integration is an operation inverse to that of di/C128erentiation; and it is a device forcalculating the ‘area under a curve’. The latter method regards the integral as the
limit of a sum and is due to Riemann. We now list some useful properties of
definite integrals.
(1) If in axb;mf
xM, where mandMare constants, then
m
bÿaZ
b
af
xdM
bÿa:
529PROPERTIES OF DEFINITE INTEGRALS
Divide the interval /C91 a,b/C93 into nsubintervals by means of the points
x1;x2;...;xnÿ1chosen arbitrarily. Let /C17kbe any point in the subinterval
xkÿ1/C17kxk, then we have
mxkf
/C17kxkMxk;k1;2;...;n;
where xkxkÿxkÿ1. Summing from k1t o nand using the fact that
Xn
k1xk
x1ÿa
x2ÿx1
bÿxnÿ1bÿa;
it follows that
m
bÿaXn
rk1f
/C17kxkM
bÿa:
Taking the limit as n!1 and each xk!0 we have the required result.
(2) If in axb;f
x/C103
x, then
Zb
af
xdxZb
a/C103
xdx:
(3)Zb
af
xdx/C12/C12/C12/C12/C12/C12/C12/C12Z
b
af
xjj dx ifa<b:
From the inequality
abc jj ajjbjjcjj ;
where jajis the absolute value of a real number a, we have
Xn
k1f
/C17kxk/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12X
n
k1f
/C17kxk jj Xn
k1f
/C17kjj xk:
Taking the limit as n!1 and each xk!0 we have the required result.
(4) The mean value theorem: If f
xis continuous in /C91 a,b/C93, we can find a point /C17
in (a,b) such that
Zb
af
xdx
bÿaf
/C17:
Since f
xis continuous in /C91 a,b/C93, we can find constants mandMsuch that
mf
xM. Then by (1) we have
m1
bÿaZb
af
xdxM:
Since f
xis continuous it takes on all values between mandM; in parti-
cular there must be a value /C17such that
f
/C17Zb
af
xdx=
bÿa;a</C17< b:
The required result follows on multiplying by bÿa.
530APPENDI/C88 1 PRELIMINARIES
/C83ome useful methods of integration
(1) Changing variables: We use a simple example to illustrate this common
procedure. Consider the integral
IZ1
0eÿax2dx;
which is equal to
=a1=2=2:To show this let us write
IZ1
0eÿax2dxZ1
0eÿay2dy:
Then
I2Z1
0eÿax2dxZ1
0eÿay2dyZ1
0Z1
0eÿa
x2y2dxdy :
We now rewrite the integral in plane polar coordinates
r;/C58
x2y2r2;dxdy rdrd. Then
I2Z1
0Z=2
0eÿar2rddr
2Z1
0eÿar2rdr
2
ÿeÿar2
2a/C12/C12/C12/C121
0
4a
and
IZ1
0eÿax2dx
=a1=2=2:
(2) Integration by parts: Since
d
dxu/C118
ud/C118
dx/C118du
dx;
where uf
xand/C118/C103
x, it follows that
Z
ud/C118
dx
dxu/C118ÿZ
/C118du
dx
dx:
This can be a useful formula in evaluating integrals.
Example A1.17
Evaluate IR
tanÿ1xdx
Solution: Since tanÿ1x can be easily di/C128erentiated, we write
IR
tanÿ1xdxR
1tanÿ1xdxand let utanÿ1x;dv=dx1. Then
Ixtanÿ1xÿZxdx
1x2xtanÿ1xÿ1
2log
1x2c:
531SOME USEFUL METHODS OF INTEGRATION
Example A1.18
Show that
Z1
ÿ1x2eÿax2dx1=2
2a3=2:
Solution: Let us first consider the integral
IZ1
0eÿax2dx
‘Integration-by-parts’ gives
IZc
beÿax2dxeÿax2x/C12/C12/C12/C12c
b2Zc
bax2eÿax2dx;
from which we obtain
Zc
bx2eÿax2dx1
2aZc
beÿax2dxÿeÿax2x/C12/C12/C12/C12c
b
:
We let limits bandcbecome ÿ1 and1, and thus obtain the desired result.
Problem A1.23
Evaluate IZ
xexdx(constant).
(3) Partial fractions: Any rational function P
x=/C81
x, where P
xand/C81
x
are polynomials, with the degree of P
xless than that of /C81
x, can be written as
the sum of rational functions having the form A=
axbk,
AxB=
ax2bxck, where k1;2;3;...which can be integrated in
terms of elementary functions.
Example A1.19
3xÿ2
4xÿ3
2x53A
4xÿ3B
2x53C
2x52D
2x5;
5x2ÿx2
x22x42
xÿ1AxB
x22x42CxD
x22x4/C69
xÿ1:
Solution: The coecients A,B,/C67etc., can be determined by clearing the frac-
tions and equating coecients of like powers of xon both sides of the equation.
Problem A1.24
Evaluate
IZ6ÿx
xÿ3
2x5dx:
532APPENDI/C88 1 PRELIMINARIES
(4) Rational functions of sin xand cos xcan always be integrated in terms of
elementary functions by substitution tan
x=2u, as shown in the following
example.
Example A1.20
Evaluate
IZdx
53c o s x:
Solution: Let tan
x=2u, then
sin
x=2u=
1u2/C112
;cos
x=21=1u
2/C112
and
cosxcos2
x=2ÿsin2
x=21ÿu2
1u2;
also
du1
2sec2
x=2dx ordx2 cos2
x=22du=
1u2:
Thus
IZdu
u241
2tanÿ1
u=2c12tanÿ112tanx=2 c:
/C82eduction formulas
Consider an integral of the formR
xneÿxdx. Since this depends upon nlet us call it
In. Then using integration by parts we have
InÿxneÿxnZ
xnÿ1eÿxdxxneÿxnInÿ1:
The above equation gives Inin terms of Inÿ1
Inÿ2;Inÿ3, etc.) and is therefore called
a reduction formula.
Problem A1.25
Evaluate InZ=2
0sinnxdxZ=2
0sinxsinnÿ1xdx:
533REDUCTION FORMULAS
/C68i/C128erentiation of integrals
(1) Indefinite integrals: We first consider di/C128erentiation of indefinite integrals. If
f
x;is an integrable function of xand is a variable parameter, and if
Z
f
x;dx/C71
x;;
A1:11
then we have
/C64/C71
x;=/C64xf
x;:
A1:12
Furthermore, if f
x;is such that
/C642/C71
x;
/C64x/C64/C642/C71
x;
/C64/C64x;
then we obtain
/C64
/C64x/C64/C71
x;
/C64
/C64
/C64/C64/C71
x;
/C64x
/C64f
x;
/C64
and integrating gives
Z/C64f
x;
/C64dx/C64/C71
x;
/C64;
A1:13
which is valid provided /C64f
x;=/C64is continuous in xas well as .
(2) Definite integrals: We now extend the above procedure to definite integrals:
I
Zb
af
x;dx;
A1:14
where f
x;is an integrable function of xin the interval axb,a n d aandb
are in general continuous and di/C128erentiable (at least once) functions of . We now
have a relation similar to Eq. (A1.11):
I
Zb
af
x;dx/C71
b;ÿ/C71
a;
A1:15
and, from Eq. (A1.13),
Zb
a/C64f
x;
/C64dx/C64/C71
b;
/C64ÿ/C64/C71
a;
/C64:
A1:16
Di/C128erentiating (A1.15) totally
dI
d/C64/C71
b;
/C64bdb
d/C64/C71
b;
/C64ÿ/C64/C71
a;
/C64ada
dÿ/C64/C71
a;
/C64/C58
534APPENDI/C88 1 PRELIMINARIES
which becomes, with the help of Eqs. (A1.12) and (A1.16),
dI
dZb
a/C64f
x;
/C64dxf
b;db
dÿf
a;da
d;
A1:17
which is known as Leibnitz’s rule for di/C128erentiating a definite integral. If aandb,
the limits of integration, do not depend on , then Eq. (A1.17) reduces to
dI
dd
dZb
af
x;dxZb
a/C64f
x;
/C64dx:
Problem A1.26
IfI
Z2
0sin
x=xdx, find dI=d.
/C72omogeneous functions
A homogeneous function f
x1;x2;...;xnof the kth degree is defined by the
relation
f
x1;x2;...;xnkf
x1;x2;...;xn:
For example, x33x2yÿy3is homogeneous of the third degree in the variables x
andy.
Iff
x1;x2;...;xn) is homogeneous of degree kthen it is straightforward to
show that
Xn
j1xj/C64f
/C64xjkf:
This is known as Euler’s theorem on homogeneous functions.
Problem A1.27
Show that Euler’s theorem on homogeneous functions is true.
/C84a/C121lor series for functions of t/C119o independent /C118ariables
The ideas involved in Taylor series for functions of one variable can be general-ized. For example, consider a function of two variables ( x;y). If all the nth partial
derivatives of f
x;yare continuous in a closed region and if the
n1)st partial
535HOMOGENEOUS FUNCTIONS
derivatives exist in the open region, then we can expand the function f
x;yabout
xx0;yy0in the form
f
x0/C104;y0kf
x0;y0 /C104/C64
/C64xk/C64
/C64y
f
x0;y0
1
2/C33/C104/C64
/C64xk/C64
/C64y2
f
x0;y0
1
n/C33/C104/C64
/C64xk/C64
/C64yn
f
x0;y0Rn;
where /C104xxÿx0;kyyÿy0;Rn, the remainder after nterms, is given
by
Rn1
n1/C33/C104/C64
/C64xk/C64
/C64yn1
f
x0/C104;y0k;0<< 1;
and where we use the operator notation
/C104/C64
/C64xk/C64
/C64y
f
x0;y0/C104fx
x0;y0kfy
x0;y0;
/C104/C64
/C64xk/C64
/C64y2
f
x0;y0 /C1042/C642
/C64x22/C104k/C642
/C64x/C64yk2/C642
/C64y2/C32!
f
x0;y0;
etc., when we expand
/C104/C64
/C64xk/C64
/C64yn
formally by the binomial theorem.
When lim n!1Rn0 for all
x;y) in a region, the infinite series expansion is
called a Taylor series in two variables. Extensions can be made to three or more
variables.
Lagrange multiplier
For functions of one variable such as f
xto have a stationary value (maximum
or minimum) at xa, we have f0
a0. If fn
a<0 it is a relative maximum
while if f
a/C620 it is a relative minimum.
Similarly f
x;yhas a relative maximum or minimum at xa;ybif
fx
a;b0,fy
a;b0. Thus possible points at which f
x;yhas a relative max-
imum or minimum are obtained by solving simultaneously the equations
/C64f=/C64x0;/C64 f=/C64y0:
536APPENDI/C88 1 PRELIMINARIES
Sometimes we wish to find the relative maxima or minima of f
x;y0 subject to
some constraint condition /C30
x;y0. To do this we first form the function
/C103
x;yf
x;yf
x;yand then set
/C64/C103=/C64x0;/C64 /C103=/C64y0:
The constant is called a Lagrange multiplier and the method is known as the
method of undetermined multipliers.
537LAGRANGE MULTIPLIER
Appendix 2
/C68eterminants
The determinant is a tool used in many branches of mathematics, science, and
engineering. The reader is assumed to be familiar with this subject. However, for
those who are in need of review, we prepared this appendix, in which the deter-
minant is defined and its properties developed. In Chapters 1 and 3, the reader will
see the determinant’s use in proving certain properties of vector and matrix
operations.
The concept of a determinant is already familiar to us from elementary algebra,
where, in solving systems of simultaneous linear equation, we find it convenient to
use determinants. For example, consider the system of two simultaneous linear
equations
a11x1a12x2b1;
a21x1a22x2b2;)
A2:1
in two unknowns x1;x2where aij
i;j1;2are constants. These two equations
represent two lines in the x1x2plane. To solve the system (A2.1), multiplying the
first equation by a22, the second by ÿa12and then adding, we find
x1b1a22ÿb2a12
a11a22ÿa21a12:
A2:2a
Next, by multiplying the first equation by ÿa21, the second by a11and adding, we
find
x2b2a11ÿb1a21
a11a22ÿa21a12:
A2:2b
We may write the solutions (A2.2) of the system (A2.1) in the determinant form
x1D1
D;x2D2
D;
A2:3
538
where
D1b1a12
b2a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;D
2a11b1
a21b2/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;Da
11a12
a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
A2:4
are called determinants of second order or order 2. The numbers enclosed between
vertical bars are called the elements of the determinant. The elements in ahorizontal line form a row and the elements in a vertical line form a column
of the determinant. It is obvious that in Eq. (A2.3) D60.
Note that the elements of determinant /C68are arranged in the same order as they
occur as coecients in Eqs. (A1.1). The numerator D
1forx1is constructed from
/C68by replacing its first column with the coecients b1andb2on the right-hand
side of (A2.1). Similarly, the numerator for x2is formed by replacing the second
column of /C68byb1;b2. This procedure is often called Cramer’s rule.
Comparing Eqs. (A2.3) and (A2.4) with Eq. (A2.2), we see that the determinant
is computed by summing the products on the rightward arrows and subtractingthe products on the leftward arrows:
a
11a12
a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12a
11a22ÿa12a21;etc:
ÿ
This idea is easily extended. For example, consider the system of three linear
equations
a11x1a12x2a13x3b1;
a21x1a22x2a23x3b2;
a31x1a32x2a33x3b3;9
>>=
>>;
A2:5
in three unknowns x1;x2;x3. To solve for x1, we multiply the equations by
a22a33ÿa32a23;ÿ
a12a33ÿa32a13;a12a23ÿa22a13;
respectively, and then add, finding
x1b1a22a33ÿb1a23a32b2a13a32ÿb2a12a33b3a12a23ÿb3a13a22
a11a22a33ÿa11a32a23a21a32a13ÿa21a12a33a31a12a23ÿa31a22a13;
which can be written in determinant form
x1D1=D;
A2:6
539APPENDI/C88 2 DETERMINANTS
where
Da11a12a13
a21a22a23
a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;D
1b1a12a13
b2a22a23
b3a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
A2:7
Again, the elements of /C68are arranged in the same order as they appear as
coecients in Eqs. (A2.5), and D
1is obtained by Cramer’s rule. In the same
manner we can find solutions for x2;x3. Moreover, the expansion of a determi-
nant of third order can be obtained by diagonal multiplication by repeating on the
right the first two columns of the determinant and adding the signed products of
the elements on the various diagonals in the resulting array:
This method of writing out determinants is correct only for second- and third-
order determinants.
Problem A2.1
Solve the following system of three linear equations using Cramer’s rule:
2x1ÿx22x32;
x110x2ÿ3x35;
ÿx1x2x3ÿ3:
Problem A2.2
Evaluate the following determinants
a12
43/C12/C12/C12/C12/C12/C12/C12/C12;
b51 8
1 536
1 042/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
ccos ÿsin
sin cos /C12/C12/C12/C12/C12/C12/C12/C12:
/C68eterminants/C44 minors/C44 and cofactors
We are now in a position to define an nth-order determinant. A determinant of
order nis a square array of n
2quantities enclosed between vertical bars,
540APPENDI/C88 2 DETERMINANTS
a11a12a13
a21a22a23
a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12a
11
a21
a31a12
a22
a32
ÿ
ÿ
ÿ
Da11a12 a1n
a21a22 a2n
.........
an1an2 ann/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
A2:8
By deleting the ith row and the kth column from the determinant /C68we obtain
an (nÿ1)st order determinant (a square array of nÿ1 rows and nÿ1 columns
between vertical bars), which is called the minor of the element a
ik(which belongs
to the deleted row and column) and is denoted by Mik. The minor Mikmultiplied
by
ÿikis called the cofactor of aikand is denoted by Cik:
Cik
ÿ 1ikMik:
A2:9
For example, in the determinant
a11a12a13
a21a22a23
a31a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
we have
C
11
ÿ 111M11a22a23
a32a33/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12; C
32
ÿ 132M32ÿa11a13
a21a23/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;etc:
It is very convenient to get the proper sign (plus or minus) for the cofactor
ÿ1
ikby thinking of a checkerboard of plus and minus signs like this
ÿ ÿ
ÿÿ
ÿ ÿ etc:
ÿÿ
etc:...
ÿ
ÿ/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
thus, for the element a
23we can see that the checkerboard sign is minus.
/C69/C120pansion of determinants
Now we can see how to find the value of a determinant: multiply each of one row
(or one column) by its cofactor and then add the results, that is,
541E/C88PANSION OF DETERMINANTS
Dai1Ci1ai2Ci2 ainCin
Xn
k1aikCik
i1;2;...;orn
A2:10a
cofactor expansion along the ith row
or
Da1kC1ka2kC2k ankCnk
Xn
i1aikCik
k1;2;...;orn:
A2:10b
cofactor expansion along the kth column
We see that /C68is defined in terms of ndeterminants of order nÿ1, each of
which, in turn, is defined in terms of nÿ1 determinants of order nÿ2, and so on;
we finally arrive at second-order determinants, in which the cofactors of the
elements are single elements of /C68. The method of evaluating a determinant just
described is one form of Laplace’s development of a determinant.
Problem A2.3
For a second-order determinant
Da11a12
a21a22/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12
show that the Laplace’s development yields the same value of /C68no matter which
row or column we choose.
Problem A2.4
Let
D130
264
ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
Evaluate /C68, first by the first-row expansion, then by the first-column expansion.
Do you get the same value of /C68/C63
Properties of determinants
In this section we develop some of the fundamental properties of the determinant
function. In most cases, the proofs are brief.
542APPENDI/C88 2 DETERMINANTS
(1) If all elements of a row (or a column) of a determinant are zero, the value of
the determinant is zero.
Proof: Let the elements of the kth row of the determinant /C68be zero. If we
expand /C68in terms of the ith row, then
Dai1Ci1ai2Ci2 ainCin:
Since the elements ai1;ai2;...;ainare zero, D0. Similarly, if all the elements in
one column are zero, expanding in terms of that column shows that the determi-
nant is zero.
(2) If all the elements of one row (or one column) of a determinant are multi-
plied by the same factor k, the value of the new determinant is ktimes the value of
the original determinant. That is, if a determinant Bis obtained from determinant
/C68by multiplying the elements of a row (or a column) of /C68by the same factor k,
then BkD.
Proof: Suppose Bis obtained from /C68by multiplying its ith row by k. Hence the
ith row of Biskaij, where j1;2;...;n, and all other elements of Bare the same
as the corresponding elements of A. Now expand Bin terms of the ith row:
Bkai1Ci1kai2Ci2 kainCin
k
ai1Ci1ai2Ci2 ainCin
kD:
The proof for columns is similar.
Note that property (1) can be considered as a special case of property (2) with
k0.
Example A2.1
If
D12 3
01 1
4ÿ10/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12and B16 3
03 1
4ÿ30/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
then we see that the second column of Bis three times the second column of /C68.
Evaluating the determinants, we find that the value of /C68isÿ3, and the value of B
isÿ9 which is three times the value of /C68, illustrating property (2).
Property (2) can be used for simplifying a given determinant, as shown in the
following example.
543PROPERTIES OF DETERMINANTS
Example A2.2
130
264
ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C122130
132
ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C1223110
112
ÿ102/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12232110
111
ÿ101/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ÿ12:
(3) The value of a determinant is not altered if its rows are written as columns,
in the same order.
Proof: Since the same value is obtained whether we expand a determinant by
any row or any column, thus we have property (3). The following example will
illustrate this property.
Example A2.3
D10 2
ÿ11 0
2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12110
ÿ13/C12/C12/C12/C12/C12/C12/C12/C12ÿ0ÿ10
23/C12/C12/C12/C12/C12/C12/C12/C122ÿ11
2ÿ1/C12/C12/C12/C12/C12/C12/C12/C121:
Now interchanging the rows and the columns, then evaluating the value of the
resulting determinant, we find
1ÿ12
01 ÿ1
203/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C1211ÿ1
03/C12/C12/C12/C12/C12/C12/C12/C12ÿ
ÿ 10ÿ1
23/C12/C12/C12/C12/C12/C12/C12/C12201
20/C12/C12/C12/C12/C12/C12/C12/C121;
illustrating property (3).
(4) If any two rows (or two columns) of a determinant are interchanged, the
resulting determinant is the negative of the original determinant.
Proof: The proof is by induction. It is easy to see that it holds for 2 2 deter-
minants. Assuming the result holds for nndeterminants, we shall show that it
also holds for
n1
n1determinants, thereby proving by induction that it
holds in general.
LetBbe an
n1
n1determinant obtained from /C68by interchanging
two rows. Expanding B in terms of a row that is not one of those interchanged,
such as the kth row, we have
BX
n
j1
ÿ1jkbkjM0
kj;
where M0
kjis the minor of bkj. Each bkjis identical to the corresponding akj(the
elements of /C68). Each M0
kjis obtained from the corresponding Mkj(ofakj)b y
544APPENDI/C88 2 DETERMINANTS
interchanging two rows. Thus bkjakj, and M0
kjÿMkj. Hence
BÿXn
j1
ÿ1jkbkjMkjÿD:
The proof for columns is similar.
Example A2.4
Consider
D10 2
ÿ11 0
2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C121:
Now interchanging the first two rows, we have
Bÿ11 0
10 2
2ÿ13/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ÿ1
illustrating property (4).
(5) If corresponding elements of two rows (or two columns) of a determinant
are proportional, the value of the determinant is zero.
Proof: Let the elements of the ith and jth rows of /C68be proportional, say,
a
ikcajk;k1;2;...;n.I fc0, then /C680. For c60, then by property (2),
/C68cB, where the ith and jth rows of Bare identical. Interchanging these two
rows, Bgoes over to ÿB(by property (4)). But the rows are identical, the new
determinant is still B. Thus BÿB;B0, and /C680.
Example A2.5
B11 2
ÿ1ÿ10
22 8/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C120;D36 ÿ4
1ÿ13
ÿ6ÿ12 8/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C120:
InBthe first and second columns are identical, and in /C68the first and the third
rows are proportional.
(6) If each element of a row of a determinant is a binomial, then the determi-
nant can be written as the sum of two determinants, for example,
4x232
x 43
3xÿ121/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C124x32
x43
3x21/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12232
043
ÿ121/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
545PROPERTIES OF DETERMINANTS
Proof: Expanding the determinant by the row whose terms are binomials, we
will see property (6) immediately.
(7) If we add to the elements of a row (or column) any constant multiple of the
corresponding elements in any other row (or column), the value of the determi-
nant is unaltered.
Proof: Applying property (6) to the determinant that results from the given
addition, we obtain a sum of two determinants: one is the original determinant
and the other contains two proportional rows. Then by property (4), the seconddeterminant is zero, and the proof is complete.
It is advisable to simplify a determinant before evaluating it. This may be done
with the help of properties (7) and (2), as shown in the following example.
Example A2.6
Evaluate
D12 4 2 1 9 3
2ÿ37ÿ1 194
ÿ23 50 ÿ171
ÿ3 177 63 234/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
To simplify this, we want the first elements of the second, third and last rows all to
be zero. To achieve this, add the second row to the third, and add three times thefirst to the last, subtract twice the first row from the second; then develop the
resulting determinant by the first column:
D12 42 1 9 3
0ÿ85 ÿ43 8
0ÿ2 ÿ12 3
0 249 126 513/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ÿ85ÿ43 8
ÿ2ÿ12 3
249 126 513/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
We can simplify the resulting determinant further. Add three times the first row to
the last row:
Dÿ85ÿ43 8
ÿ2ÿ12 3
ÿ6ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
Subtract twice the second column from the first, and then develop the resulting
determinant by the first column:
D1ÿ43 8
0ÿ12 3
0ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12ÿ12 3
ÿ3 537/C12/C12/C12/C12/C12/C12/C12/C12ÿ537ÿ23
ÿ 3ÿ 468 :
546APPENDI/C88 2 DETERMINANTS
By applying the product rule of di/C128erentiation we obtain the following
theorem.
/C68eri/C118ati/C118e of a determinant
If the elements of a determinant are di/C128erentiable functions of a variable, then the
derivative of the determinant may be written as a sum of individual determinants,
for example,
d
dxabc
ef /C103
/C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12a
0b0c0
ef/C103
/C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12abc
e
0f0/C1030
/C104mn/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12abc
ef /C103
/C104
0m0n0/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;
where a;b;...;m;nare di/C128erentiable functions of x, and the primes denote deri-
vatives with respect to x.
Problem A2.5
Show, without computation, that the following determinants are equal to zero:
0 aÿb
ÿa 0 c
bÿc 0/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12;02 ÿ3
ÿ204
3ÿ40/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12/C12:
Problem A2.6
Find the equation of a plane which passes through the three points (0, 0, 0),
(1, 2, 5), and (2, ÿ1, 0).
547DERIVATIVE OF A DETERMINANT
Appendix 3
Table of /C42 /C70
x1
2pZx
0eÿt2=2dt:
548x 0.0 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09
0.0 0.0000 0.0040 0.0080 0.0120 0.0160 0.0199 0.0239 0.0279 0.0319 0.0359
0.1 0.0398 0.0438 0.0478 0.0517 0.0557 0.0596 0.0636 0.0675 0.0714 0.07530.2 0.0793 0.0832 0.0871 0.0910 0.0948 0.0987 0.1026 0.1064 0.1103 0.11410.3 0.1179 0.1217 0.1255 0.1293 0.1331 0.1368 0.1406 0.1443 0.1480 0.15170.4 0.1554 0.1591 0.1628 0.1664 0.1700 0.1736 0.1772 0.1808 0.1844 0.18790.5 0.1915 0.1950 0.1985 0.2019 0.2054 0.2088 0.2123 0.2157 0.2190 0.2224
0.6 0.2257 0.2291 0.2324 0.2357 0.2389 0.2422 0.2454 0.2486 0.2517 0.2549
0.7 0.2580 0.2611 0.2642 0.2673 0.2704 0.2734 0.2764 0.2794 0.2823 0.28520.8 0.2881 0.2910 0.2939 0.2967 0.2995 0.3023 0.3051 0.3078 0.3106 0.31330.9 0.3159 0.3186 0.3212 0.3238 0.3264 0.3289 0.3315 0.3340 0.3365 0.3389
1.0 0.3413 0.3438 0.3461 0.3485 0.3508 0.3531 0.3554 0.3577 0.3599 0.3621
1.1 0.3643 0.3665 0.3686 0.3708 0.3729 0.3749 0.3770 0.3790 0.3810 0.3830
1.2 0.3849 0.3869 0.3888 0.3907 0.3925 0.3944 0.3962 0.3980 0.3997 0.40151.3 0.4032 0.4049 0.4066 0.4082 0.4099 0.4115 0.4131 0.4147 0.4162 0.41771.4 0.4192 0.4207 0.4222 0.4236 0.4251 0.4265 0.4279 0.4292 0.4306 0.43191.5 0.4332 0.4345 0.4357 0.4370 0.4382 0.4394 0.4406 0.4418 0.4429 0.4441
1.6 0.4452 0.4463 0.4474 0.4484 0.4495 0.4505 0.4515 0.4525 0.4535 0.4545
1.7 0.4554 0.4564 0.4573 0.4582 0.4591 0.4599 0.4608 0.4616 0.4625 0.4633
1.8 0.4641 0.4649 0.4656 0.4664 0.4671 0.4678 0.4686 0.4693 0.4699 0.4706
1.9 0.4713 0.4719 0.4726 0.4732 0.4738 0.4744 0.4750 0.4756 0.4761 0.47672.0 0.4472 0.4778 0.4783 0.4788 0.04793 0.4798 0.4803 0.4808 0.4812 0.4817
2.1 0.4821 0.4826 0.4830 0.4834 0.4838 0.4842 0.4846 0.4850 0.4854 0.4857
2.2 0.4861 0.4864 0.4868 0.4871 0.4875 0.4878 0.4881 0.4884 0.4887 0.48902.3 0.4893 0.4896 0.4898 0.4901 0.4904 0.4906 0.4909 0.4911 0.4913 0.49162.4 0.4918 0.4920 0.4922 0.4925 0.4927 0.4929 0.4931 0.4932 0.4934 0.49362.5 0.4938 0.4940 0.4941 0.4943 0.4945 0.4946 0.4948 0.4949 0.4951 0.4952
2.6 0.4953 0.4955 0.4956 0.4957 0.4959 0.4960 0.4961 0.4962 0.4963 0.4964
2.7 0.4965 0.4966 0.4967 0.4968 0.4969 0.4970 0.4971 0.4972 0.4973 0.49742.8 0.4974 0.4975 0.4976 0.4977 0.4977 0.4978 0.4979 0.4979 0.4980 0.49812.9 0.4981 0.4982 0.4982 0.4983 0.4984 0.4984 0.4985 0.4986 0.4986 0.49863.0 0.4987 0.4987 0.4987 0.4988 0.4988 0.4989 0.4989 0.4989 0.4990 0.4990
x 0.0 0.2 0.4 0.6 0.8
1.0 0.3413447 0.3849303 0.4192433 0.4452007 0.4640697
2.0 0.4772499 0.4860966 0.4918025 0.4953388 0.4974449
3.0 0.4986501 0.4993129 0.4998409 0.4999277 0.49992774.0 0.4999683 0.4999867 0.4999946 0.4999979 0.4999992
/C42 This table is reproduced, by permission, from the Biometrica Tables for Statisticians , vol. 1, 1954,
edited by E. S. Pearson and H. O. Hartley and published by the Cambridge University Press for the
Biometrica Trustees.
/C70urther reading
Anton, Howard, Elementary Linear Algebra , 3rd ed., John Wiley, New York, 1982.
Arfken, G. B., Weber, H. J., Mathematical Methods for Physicists , 4th ed., Academic Press,
New York, 1995.
Boas, Mary L., Mathematical Methods in the Physical Sciences , 2nd ed., John Wiley, New
York, 1983.
Butkov, Eugene, Mathematical Physics , Addison-Wesley, Reading (MA), 1968.
Byon, F. W., Fuller, R. W., Mathematics of /C67lassical and /C81uantum Physics , Addison-
Wesley, Reading (MA), 1968.
Churchill, R. V., Brown, J. W., Verhey, R. F., /C67omplex /C86ariables /C38 Applications , 3rd ed.,
McGraw-Hill, New York, 1976.
Harper, Charles, Introduction to Mathematical Physics , Prentice Hall, Englewood Cli/C128s,
NJ, 1976.
Kreyszig, E., Advanced Engineering Mathematics , 3rd ed., John Wiley, New York, 1972.
Joshi, A. W., Matrices and Tensor in Physics , John Wiley, New York, 1975.
Joshi, A. W., Elements of /C71roup Theory for Physicists , John Wiley, New York, 1982.
Lass, Harry, /C86ector and Tensor Analysis , McGraw-Hill, New York, 1950.
Margenus, Henry, Murphy, George M., The Mathematics of Physics and /C67hemistry ,D .
Van Nostrand, New York, 1956.
Mathews, Fon, Walker, R. L., Mathematical Methods of Physics , W. A. Benjamin, New
York, 1965.
Spiegel, M. R., Advanced Mathematics for Engineers and Scientists , Schaum’s Outline
Series, McGraw-Hill, New York, 1971.
Spiegel, M. R., Theory and Problems of /C86ector Analysis , Schaum’s Outline Series,
McGraw-Hill, New York, 1959.
Wallace, P. R., Mathematical Analysis of Physical Problems , Dover, New York, 1984.
Wong, Chun Wa, Introduction to Mathematical Physics/C44 Methods and /C67oncepts , Oxford,
New York, 1991.
Wylie, C., Advanced Engineering Mathematics , 2nd ed., McGraw-Hill, New York, 1960.
549
Index
Abel’s integral equation, 426
Abelian group, 431
adjoint operator, 212
analytic functions, 243
Cauchy integral formula and, 244Cauchy integral theorem and, 257
angular momentum operator, 18Argand diagram, 234associated Laguerre equation, polynomials see
Laguerre equation; Laguerre functions
associated Legendre equation, functions see
Legendre equation; Legendre functions
associated tensors, 53auxiliary (characteristic) equation, 75axial vector, 8
Bernoulli equation, 72
Bessel equation, 321
series solution, 322
Bessel functions, 323
approximations, 335
first kind /C74
n
x, 324
generating function, 330
hanging flexible chain, 328
Hankel functions, 328integral representation, 331orthogonality, 336recurrence formulas, 332second kind Y
n
xseeNeumann functions
spherical, 338
beta function, 95branch line (branch cut), 241branch points, 241
calculus of variations, 347–371
brachistochrone problem, 350
canonical equations of motion, 361constraints, 353
Euler–Lagrange equation, 348
Hamilton’s principle, 361calculus of variations ( contd )
Hamilton–Jacobi equation, 364
Lagrangian equations of motion, 355
Lagrangian multipliers, 353modified Hamilton’s principle, 364
Rayleigh–Ritz method, 359
cartesian coordinates, 3
Cauchy principal value, 289
Cauchy–Riemann conditions, 244
Cauchy’s integral formula, 260
Cauchy’s integral theorem, 257
Cayley–Hamilton theorem, 134change of
basis, 224coordinate system, 11
interval, 152
characteristic equation, 125
Christo/C128el symbol, 54
commutator, 107
complex numbers, 233
basic operations, 234polar form, 234roots, 237
connected, simply or multiply, 257contour integrals, 255
contraction of tensors, 50
contravariant tensor, 49convolution theorem Fourier transforms, 188
coordinate system, see specific coordinate system
coset, 439covariant di/C128erentiation, 55
covariant tensor, 49
cross product of vectors seevector product of
vectors
crossing conditions, 441
curl
cartesian, 24curvilinear, 32
cylindrical, 34spherical polar, 35
551
curvilinear coordinates, 27
damped oscillations, 80
De Moivre’s formula, 237
del, 22
formulas involving del, 27
delta function, Dirac, 183
Fourier integral, 183Green’s function and, 192point source, 193
determinants, 538–547di/C128erential equations, 62
first order, 63
exact 67integrating factors, 69separable variables 63
homogeneous, 63numerical solutions, 469second order, constant coecients, 72
complementary functions, 74Frobenius and Fuchs theorem, 86particular integrals, 77
singular points, 86
solution in power series, 85
direct product
matrices, 139tensors, 50
direction angles and direction cosines, 3divergence
cartesian, 22curvilinear, 30cylindrical, 33spherical polar, 35
dot product of vectors, 5dual vectors and dual spaces, 211
eigenvalues and eigenfunctions of Sturm–
Liouville equations, 340
hermitian matrices, 124
orthogonality, 129
real, 128
an operator, 217
entire functions, 247Euler’s linear equation, 83
Fourier series, 144
convergence and Dirichlet conditions, 150
di/C128erentiation, 157Euler–Fourier formulas, 145exponential form of Fourier series, 156Gibbs phenomenon, 150half-range Fourier series, 151integration, 157
interval, change of, 152
orthogonality, 162Parseval’s identity, 153vibrtating strings, 157Fourier transform, 164
convolution theorem, 188
delta function derivation, 183
Fourier integral, 164Fourier series ( contd )
Fourier sine and cosine transforms, 172
Green’s function method, 192head conduction, 179Heisenberg’s uncertainty principle, 173Parseval’s identity, 186solution of integral equation, 421transform of derivatives, 190wave packets and group velocity, 174
Fredholm integral equation, 413 seealso integral
equations
Frobenius’ method seeseries solution of
di/C128erential equations
Frobenius–Fuch’s theorem, 86function
analytic,entire,harmonic, 247
function spaces, 226
gamma function, 94
gauge transformation, 411Gauss’ law, 391Gauss’ theorem, 37generating function, for
associated Laguerre polynomials, 320
Bessel functions, 330
Hermite polynomials, 314Laguerre polynomials, 317Legendre polynomials, 301
Gibbs phenomenon, 150gradient
cartesian, 20curvilinear, 29cylindrical, 33spherical polar, 35
Gram–Schmidt orthogonalization, 209Green’s functions, 192,
construction of
one dimension, 192three dimensions, 405
delta function, 193
Green’s theorem, 43
in the plane, 44
group theory, 430
conjugate clsses, 440cosets, 439cyclic group, 433
rotation matrix, 234, 252special unitary group, SU(2), 232
definitions, 430dihedral groups, 446
generator, 451
homomorphism, 436irreducible representations, 442isomorphism, 435multiplication table, 434Lorentz group, 454
permutation group, 438
orthogonal group SO(3)
552INDE/C88
group theory ( contd )
symmetry group, 446
unitary group, 452
unitary unimodular group SU
n
Hamilton–Jacobi equation, 364Hamilton’s principle and Lagrange equations of
motion, 355
Hankel functions, 328
Hankel transforms, 385
harmonic functions, 247Helmholtz theorem, 44Hermite equation, 311Hermite polynomials, 312
generating function, 314
orthogonality, 314
recurrence relations, 313
hermitian matrices, 114
orthogonal eigenvectors, 129real eigenvalues, 128
hermitian operator, 220
completeness of eigenfunctions, 221eigenfunctions, orthogonal, 220eigenvalues, real, 220
Hilbert space, 230Hilbert-Schmidt method of solution, 421homogeneous seelinear equations
homomorphism, 436
indicial equation, 87
inertia, moment of, 135
infinity seesingularity, pole essential singularity
integral equations, 413
Abel’s equation, 426classical harmonic oscillator, 427di/C128erential equation–integral equation
transformation, 419
Fourier transform solution, 421Fredholm equation, 413Laplace transform solution, 420Neumann series, 416quantum harmonic oscillator, 427Schmidt–Hilbert method, 421
separable kernel, 414
Volterra equation, 414
integral transforms, 384 see also Fourier
transform, Hankel transform, Laplacetransform, Mellin transform
Fourier, 164
Hankel, 385
Laplace, 372Mellin, 385
integration, vector, 35
line integrals, 36surface integrals, 36
interpolation, 1461
inverse operator, uniqueness of, 218
irreducible group representations, 442
Jacobian, 29kernels of integral equations, 414
separable, 414
Kronecker delta, 6
mixed second-rank tensor, 53
Lagrangian, 355
Lagrangian multipliers, 354Laguerre equation, 316
associated Laguerre equation, 320
Laguerre functions, 317
associated Laguerre polynomials, 320generating function, 317orthogonality, 319Rodrigues’ representation, 318
Laplace equation, 389
solutions of, 392
Laplace transform, 372
existence of, 373integration of tranforms, 383inverse transformation, 373solution of integral equation, 420
the first shifting theorem, 378
the second shifting theorem, 379transform of derivatives, 382
Laplacian
cartesian, 24cylindrical, 34scalar, 24
spherical polar, 35
Laurent expansion, 274
Legendre equation, 296
associated Legendre equation, 307series solution of Legendre equation, 296
Legendre functions, 299
associated Legendre functions, 308generating function, 301orthogonality, 304recurrence relations, 302Rodrigues’ formula, 299
linear combination, 204linear independence, 204
linear operator, 212
Lorentz gauge condition, 411Lorentz group, 454Lorentz transformation, 455
mapping, 239
matrices, 100
anti-hermitian, 114commutatordefinition, 100diagonalization, 129
direct product, 139
eigenvalues and eigenvectors, 124Hermitian, 114nverse, 111matrix multiplication, 103moment of inertia, 135
orthogonal, 115
and unitary transformations, 121
553INDE/C88
matrices ( contd )
Pauli spin, 142
representation, 226
rotational, 117similarity transformation, 122symmetric and skew-symmetric, 109trace, 121transpose, 108unitary, 116
Maxwell equations, 411
derivation of wave equation, 411
Mellin transforms, 385metric tensor, 51mixed tensor, 49Morera’s theorem, 259
multipoles, 248
Neumann functions, 327
Newton’s root finding formula, 465normal modes of vibrations 136numerical methods, 459
roots of equations, 460
false position (linear interpolation), 461graphic methods, 460
Newton’s method, 464
integration, 466
rectangular rule, 455
Simpson’s rule, 469trapezoidal rule, 467
interpolation, 459
least-square fit, 477
solutions of di/C128erential equations, 469
Euler’s rule, 470Runge–Kutta method, 473Taylor series method, 472
system of equations, 476
operators
adjoint, 212
angular momentum operator, 18
commuting, 225del, 22di/C128erential operator /C68(d=dx), 78
linear, 212
orthonormal, 161, 207
oscillator,
damped, 80
integral equations for, 427simple harmonic, 427
Parseval’s identity, 153, 186partial di/C128erential equations, 387
linear second order, 388
elliptic, hyperbolic, parabolic, 388Green functions, 404Laplace transformationseparation of variables
Laplace’s equation, 392, 395, 398
wave equation, 402
Pauli spin matrices, 142phase of a complex number, 235
Poisson equation. 389
polar form of complex numbers, 234
poles, 248probability theory , 481
combinations, 485continuous distributions, 500
Gaussian, 502Maxwell–Boltzmann, 503
definition of probability, 481expectation and variance, 490fundamental theorems, 486probability distributions, 491
binomial, 491Gaussian, 497
Poisson, 495
sample space, 482
power series, 269
solution of di/C128erential equations, 85
projection operators, 222
pseudovectors, 8
quotient rule, 50rank (order), of tensor, 49
of group, 430
Rayleigh–Ritz (variational) method, 359
recurrence relations
Bessel functions, 332Hermite functions, 313Laguerre functions, 318Legendre functions, 302
residue theorem, 282
residues 279
calculus of residues, 280
Riccati equation, 98Riemann surface, 241Rodrigues’ formula
Hermite polynomials, 313Laguerre polynomials, 318
associated Laguerre polynomials, 320
Legendre polynomials, 299
associated Legendre polynomials, 308
root diagram, 238rotation
groups SO
2;SO
3, 450
of coordinates, 11, 117
of vectors, 11–13
Runge–Kutta solution, 473
scalar, definition of, 1
scalar potential, 20, 390, 411
scalar product of vectors, 5Schmidt orthogonalization seeGram–Schmidt
orthogonalization
Schro /C200dinger wave equation, 427
variational approach, 368
Schwarz–Cauchy inequality, 210secular (characteristic) equation, 125
554INDE/C88
series solution of di/C128erential equations,
Bessel’s equation, 322
Hermite’s equation, 311
Laguerre’s equation, 316Legendre’s equation, 296
associated Legendre’s equation, 307
similarity transformation, 122singularity, 86, 248
branch point, 240
di/C128erential equation, 86
Laurent series, 274on contour of integration, 290
special unitary group, SU
n, 452
spherical polar coordinates, 34step function, 380
Stirling’s asymptotic formula for n/C33, 99
Stokes’ theorem, 40
Sturm–Liouville equation, 340subgroup, 439summation convention (Einstein’s), 48symbolic software, 492
Taylor series of elementary functions, 272
tensor analysis, 47
associated tensor, 53basic operations with tensors, 49contravariant vector, 48covariant di/C128erentiation, 55
covariant vector, 48tensor analysis ( contd )
definition of second rank tensor, 49geodesic in Riemannian space, 53
metric tensor, 51quotient law, 50symmetry–antisymmetry, 50
trace (matrix), 121triple scalar product of vectors, 10triple vector product of vectors, 11
uncertainty principle in quantum theory, 173
unit group element, 430unit vectors
cartesian, 3cylindrical, 32
spherical polar, 34
variational principles, seecalculus of variations
vector and tensor analysis, 1–56
vector potential, 411vector product of vectors, 7
vector space, 13, 199
Volterra integral equation. seeintegral equations
wave equation, 389
derivation from Maxwell’s equations, 411solution of, separation of variables, 402
wave packets, 174
group velocity, 174
555INDE/C88