Arfken Math Methods 6th Ed 2005
PDF · 1195 pages · 6.4 MB
Open PDF file
Textbook by George B. Arfken and Hans J. Weber, published by Elsevier Academic Press in 2005, kept in the archive as a downloaded reference book. The contents cover vector analysis, curved coordinates and tensors, matrices, group theory, infinite series, complex variables, the gamma function, differential equations, Sturm-Liouville theory, special functions (Bessel, Legendre, Hermite and others), Fourier series, integral transforms, integral equations, and calculus of variations. The text seen covers only the front matter and table of contents through chapter 18, so later content is not confirmed.
AI-written summary; may contain errors.
Extracted text (machine-read; may contain errors)
MATHEMATICAL
METHODS FOR
PHYSICISTS
SIXTH EDITION
GeorgeB. Arfken
MiamiUniversity
Oxford,OH
Hans J.Weber
Universityof Virginia
Charlottesville,VA
Amsterdam Boston Heidelberg London NewYork Oxford
Paris SanDiego SanFrancisco Singapore Sydney Tokyo
This page intentionally left blank
MATHEMATICAL
METHODS FOR
PHYSICISTS
SIXTH EDITION
This page intentionally left blank
Acquisitions Editor TomSinger
Project Manager Simon Crump
MarketingManager Linda Beattie
Cover Design EricDeCicco
Composition VTEXTypesettingServices
Cover Printer Phoenix Color
Interior Printer TheMaple–VailBook Manufacturing Group
ElsevierAcademicPress
30 Corporate Drive,Suite 400, Burlington, MA01803, USA
525 B Street,Suite 1900, San Diego, California 92101-4495, USA
84 Theobald’s Road, London WC1X 8RR, UK
This book is printed on acid-freepaper. /circlecopyrt∞
Copyright ©2005, Elsevier Inc. Allrights reserved.
No part of this publication may be reproduced or transmitted in any form or by any means, electronic or me-
chanical, including photocopy, recording, or any information storage and retrieval system, without permission in
writing from the publisher.
Permissions may be sought directly from Elsevier’s Science & Technology Rights Department in Oxford, UK:
phone:(+44)1865843830,fax:(+44)1865853333,e-mail:[email protected]
your request on-line via the Elsevier homepage (http://elsevier.com), by selecting “Customer Support” and then
“Obtaining Permissions.”
Library of Congress Cataloging-in-Publication Data
Appication submitted
British Library Cataloguing in Publication Data
Acatalogue record for this book is availablefrom theBritish Library
ISBN: 0-12-059876-0 Case bound
ISBN: 0-12-088584-0 International Students Edition
For allinformation on all ElsevierAcademicPress Publications
visit our Website at www.books.elsevier.com
Printed inthe UnitedStates of America
050607080910987654321
CONTENTS
Preface xi
1 VectorAnalysis 1
1.1Definitions,ElementaryApproach ..................... 1
1.2RotationoftheCoordinateAxes ...................... 7
1.3ScalarorDotProduct ........................... 1 2
1.4VectororCross Product .......................... 1 8
1.5TripleScalarProduct,TripleVectorProduct ............... 2 5
1.6Gradient,∇................................. 3 2
1.7Divergence,∇................................ 3 8
1.8Curl,∇×.................................. 4 3
1.9SuccessiveApplicationsof ∇....................... 4 9
1.10VectorIntegration .............................. 5 4
1.11Gauss’Theorem ............................... 6 0
1.12Stokes’Theorem .............................. 6 4
1.13PotentialTheory .............................. 6 8
1.14Gauss’Law, Poisson’sEquation ...................... 7 9
1.15DiracDeltaFunction ............................ 8 3
1.16Helmholtz’sTheorem ............................ 9 5
AdditionalReadings ............................ 1 0 1
2 VectorAnalysisinCurvedCoordinatesandTensors 103
2.1OrthogonalCoordinatesin R3....................... 1 0 3
2.2DifferentialVectorOperators ....................... 1 1 0
2.3SpecialCoordinateSystems:Introduction ................ 1 1 4
2.4CircularCylinderCoordinates ....................... 1 1 5
2.5SphericalPolarCoordinates ........................ 1 2 3
v
vi Contents
2.6TensorAnalysis ............................... 1 3 3
2.7Contraction,DirectProduct ........................ 1 3 9
2.8QuotientRule ................................ 1 4 1
2.9Pseudotensors, DualTensors ....................... 1 4 2
2.10GeneralTensors ............................... 1 5 1
2.11TensorDerivativeOperators ........................ 1 6 0
AdditionalReadings ............................ 1 6 3
3 DeterminantsandMatrices 165
3.1Determinants ................................ 1 6 5
3.2Matrices ................................... 1 7 6
3.3OrthogonalMatrices ............................ 1 9 5
3.4HermitianMatrices,UnitaryMatrices .................. 2 0 8
3.5DiagonalizationofMatrices ........................ 2 1 5
3.6NormalMatrices .............................. 2 3 1
AdditionalReadings ............................ 2 3 9
4 GroupTheory 241
4.1IntroductiontoGroupTheory ....................... 2 4 1
4.2GeneratorsofContinuousGroups ..................... 2 4 6
4.3OrbitalAngularMomentum ........................ 2 6 1
4.4AngularMomentumCoupling ....................... 2 6 6
4.5HomogeneousLorentzGroup ....................... 2 7 8
4.6LorentzCovarianceofMaxwell’sEquations ............... 2 8 3
4.7DiscreteGroups ............................... 2 9 1
4.8DifferentialForms ............................. 3 0 4
AdditionalReadings ............................ 3 1 9
5 InfiniteSeries 321
5.1FundamentalConcepts ........................... 3 2 1
5.2ConvergenceTests ............................. 3 2 5
5.3AlternatingSeries .............................. 3 3 9
5.4AlgebraofSeries .............................. 3 4 2
5.5SeriesofFunctions ............................. 3 4 8
5.6Taylor’sExpansion ............................. 3 5 2
5.7PowerSeries ................................ 3 6 3
5.8EllipticIntegrals .............................. 3 7 0
5.9BernoulliNumbers, Euler–MaclaurinFormula .............. 3 7 6
5.10AsymptoticSeries .............................. 3 8 9
5.11InfiniteProducts .............................. 3 9 6
AdditionalReadings ............................ 4 0 1
6 FunctionsofaComplexVariableIAnalyticProperties,Mapping 403
6.1ComplexAlgebra .............................. 4 0 4
6.2Cauchy–RiemannConditions ....................... 4 1 3
6.3Cauchy’sIntegralTheorem ......................... 4 1 8
Contents vii
6.4Cauchy’sIntegralFormula ......................... 4 2 5
6.5LaurentExpansion ............................. 4 3 0
6.6Singularities ................................. 4 3 8
6.7Mapping ................................... 4 4 3
6.8ConformalMapping ............................ 4 5 1
AdditionalReadings ............................ 4 5 3
7 FunctionsofaComplexVariableII 455
7.1CalculusofResidues ............................ 4 5 5
7.2DispersionRelations ............................ 4 8 2
7.3MethodofSteepestDescents ........................ 4 8 9
AdditionalReadings ............................ 4 9 7
8 TheGammaFunction(FactorialFunction) 499
8.1Definitions,SimpleProperties ....................... 4 9 9
8.2DigammaandPolygammaFunctions ................... 5 1 0
8.3Stirling’sSeries ............................... 5 1 6
8.4TheBetaFunction ............................. 5 2 0
8.5IncompleteGammaFunction ....................... 5 2 7
AdditionalReadings ............................ 5 3 3
9 DifferentialEquations 535
9.1PartialDifferentialEquations ....................... 5 3 5
9.2First-Order DifferentialEquations .................... 5 4 3
9.3SeparationofVariables ........................... 5 5 4
9.4SingularPoints ............................... 5 6 2
9.5SeriesSolutions—Frobenius’Method ................... 5 6 5
9.6A SecondSolution .............................. 5 7 8
9.7NonhomogeneousEquation—Green’sFunction ............. 5 9 2
9.8HeatFlow, or Diffusion,PDE ....................... 6 1 1
AdditionalReadings ............................ 6 1 8
10 Sturm–LiouvilleTheory—OrthogonalFunctions 621
10.1Self-AdjointODEs ............................. 6 2 2
10.2HermitianOperators ............................ 6 3 4
10.3Gram–SchmidtOrthogonalization ..................... 6 4 2
10.4CompletenessofEigenfunctions ...................... 6 4 9
10.5Green’sFunction—EigenfunctionExpansion ............... 6 6 2
AdditionalReadings ............................ 6 7 4
11 BesselFunctions 675
11.1Bessel FunctionsoftheFirst Kind, Jν(x)................. 6 7 5
11.2Orthogonality ................................ 6 9 4
11.3NeumannFunctions ............................ 6 9 9
11.4HankelFunctions .............................. 7 0 7
11.5ModifiedBessel Functions, Iν(x)andKν(x)............... 7 1 3
viii Contents
11.6AsymptoticExpansions ........................... 7 1 9
11.7SphericalBesselFunctions ......................... 7 2 5
AdditionalReadings ............................ 7 3 9
12 LegendreFunctions 741
12.1GeneratingFunction ............................ 7 4 1
12.2RecurrenceRelations ............................ 7 4 9
12.3Orthogonality ................................ 7 5 6
12.4AlternateDefinitions ............................ 7 6 7
12.5AssociatedLegendreFunctions ...................... 7 7 1
12.6SphericalHarmonics ............................ 7 8 6
12.7OrbitalAngularMomentumOperators .................. 7 9 3
12.8AdditionTheoremforSphericalHarmonics ............... 7 9 7
12.9IntegralsofThreeY’s ............................ 8 0 3
12.10LegendreFunctionsoftheSecondKind .................. 8 0 6
12.11VectorSphericalHarmonics ........................ 8 1 3
AdditionalReadings ............................ 8 1 6
13 MoreSpecialFunctions 817
13.1HermiteFunctions ............................. 8 1 7
13.2LaguerreFunctions ............................. 8 3 7
13.3ChebyshevPolynomials .......................... 8 4 8
13.4HypergeometricFunctions ......................... 8 5 9
13.5ConfluentHypergeometricFunctions ................... 8 6 3
13.6MathieuFunctions ............................. 8 6 9
AdditionalReadings ............................ 8 7 9
14 FourierSeries 881
14.1GeneralProperties ............................. 8 8 1
14.2Advantages,Uses ofFourierSeries .................... 8 8 8
14.3ApplicationsofFourierSeries ....................... 8 9 2
14.4PropertiesofFourierSeries ........................ 9 0 3
14.5GibbsPhenomenon ............................. 9 1 0
14.6DiscreteFourierTransform ........................ 9 1 4
14.7FourierExpansionsofMathieuFunctions ................ 9 1 9
AdditionalReadings ............................ 9 2 9
15 IntegralTransforms 931
15.1IntegralTransforms ............................. 9 3 1
15.2DevelopmentoftheFourierIntegral .................... 9 3 6
15.3FourierTransforms—InversionTheorem ................. 9 3 8
15.4FourierTransformofDerivatives ..................... 9 4 6
15.5ConvolutionTheorem ............................ 9 5 1
15.6MomentumRepresentation ......................... 9 5 5
15.7TransferFunctions ............................. 9 6 1
15.8LaplaceTransforms ............................. 9 6 5
Contents ix
15.9LaplaceTransformofDerivatives ..................... 9 7 1
15.10OtherProperties .............................. 9 7 9
15.11Convolution(Faltungs) Theorem ..................... 9 9 0
15.12Inverse LaplaceTransform ......................... 9 9 4
AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1003
16 IntegralEquations 1005
16.1Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1005
16.2IntegralTransforms,GeneratingFunctions . . . . . . . . . . . . . . . . 1012
16.3NeumannSeries,Separable(Degenerate)Kernels . . . . . . . . . . . . 1018
16.4Hilbert–SchmidtTheory . . . . . . . . . . . . . . . . . . . . . . . . . . 1029
AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1036
17 CalculusofVariations 1037
17.1A DependentandanIndependentVariable . . . . . . . . . . . . . . . . 1038
17.2ApplicationsoftheEuler Equation . . . . . . . . . . . . . . . . . . . . 1044
17.3SeveralDependentVariables . . . . . . . . . . . . . . . . . . . . . . . . 1052
17.4SeveralIndependentVariables . . . . . . . . . . . . . . . . . . . . . . . 1056
17.5SeveralDependentandIndependentVariables . . . . . . . . . . . . . . 1058
17.6LagrangianMultipliers . . . . . . . . . . . . . . . . . . . . . . . . . . . 1060
17.7VariationwithConstraints . . . . . . . . . . . . . . . . . . . . . . . . . 1065
17.8Rayleigh–RitzVariationalTechnique . . . . . . . . . . . . . . . . . . . 1072
AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1076
18 NonlinearMethodsandChaos 1079
18.1Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1079
18.2TheLogisticMap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1080
18.3SensitivitytoInitialConditionsandParameters . . . . . . . . . . . . . 1085
18.4NonlinearDifferentialEquations . . . . . . . . . . . . . . . . . . . . . 1088
AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1107
19 Probability 1109
19.1Definitions,SimpleProperties . . . . . . . . . . . . . . . . . . . . . . . 1109
19.2RandomVariables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1116
19.3BinomialDistribution . . . . . . . . . . . . . . . . . . . . . . . . . . . 1128
19.4PoissonDistribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1130
19.5Gauss’NormalDistribution . . . . . . . . . . . . . . . . . . . . . . . . 1134
19.6Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1138
AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1150
GeneralReferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1150
Index 1153
This page intentionally left blank
PREFACE
Throughsixeditionsnow, MathematicalMethodsforPhysicists hasprovidedallthemath-
ematicalmethodsthataspiringsscientistsandengineersarelikelytoencounterasstudents
and beginning researchers. More than enough material is included for a two-semester un-
dergraduateorgraduatecourse.
Thebookisadvancedinthesensethatmathematicalrelationsarealmostalwaysproven,
in addition to being illustrated in terms of examples. These proofs are not what a mathe-
matician would regard as rigorous, but sketch the ideas and emphasize the relations that
are essential to the study of physics and related fields. This approach incorporates theo-
rems that are usually not cited under the most general assumptions, but are tailored to the
more restricted applications required by physics. For example, Stokes’ theorem is usually
appliedbyaphysicisttoasurfacewiththetacitunderstandingthatitbesimplyconnected.
Suchassumptionshavebeenmademoreexplicit.
PROBLEM -SOLVING SKILLS
The book also incorporates a deliberate focus on problem-solving skills. This more ad-
vancedlevelofunderstandingandactivelearningisroutineinphysicscoursesandrequires
practicebythereader.Accordingly,extensiveproblemsetsappearingineachchapterform
an integral part of the book. They have been carefully reviewed, revised and enlarged for
thisSixthEdition.
PATHWAYS THROUGH THE MATERIAL
Undergraduates may be best served if they start by reviewing Chapter 1 according to the
level of training of the class. Section 1.2 on the transformation properties of vectors, the
cross product, and the invariance of the scalar product under rotations may be postponed
until tensor analysis is started, for which these sections form the introduction and serve as
xi
xii Preface
examples. They may continue their studies with linear algebra in Chapter 3, then perhaps
tensors and symmetries (Chapters 2 and 4), and next real and complex analysis (Chap-
ters5–7), differentialequations(Chapters9, 10),andspecialfunctions(Chapters11–13).
In general, the core of a graduate one-semester course comprises Chapters 5–10 and
11–13,whichdealwithrealandcomplexanalysis,differentialequations,andspecialfunc-
tions. Depending on the level of the students in a course, some linear algebra in Chapter 3
(eigenvalues, for example), along with symmetries (group theory in Chapter 4), and ten-
sors(Chapter2)maybecoveredasneededoraccordingtotaste.Grouptheorymayalsobe
included with differential equations (Chapters 9 and 10). Appropriate relations have been
includedandarediscussedinChapters4and9.
A two-semester course can treat tensors, group theory, and special functions (Chap-
ters 11–13) more extensively, and add Fourier series (Chapter 14), integral transforms
(Chapter15),integralequations(Chapter16),andthecalculusofvariations(Chapter17).
CHANGES TO THE SIXTH EDITION
ImprovementstotheSixthEditionhavebeenmadeinnearlyallchaptersaddingexamples
and problems and more derivations of results. Numerous left-over typos caused by scan-
ning into LaTeX, an error-prone process at the rate of many errors per page, have been
corrected along with mistakes, such as in the Dirac γ-matrices in Chapter 3. A few chap-
ters have been relocated. The Gamma function is now in Chapter 8 following Chapters 6
and 7 on complex functions in one variable, as it is an application of these methods. Dif-
ferential equations are now in Chapters 9 and 10. A new chapter on probability has been
added,aswellasnewsubsectionsondifferentialformsandMathieufunctionsinresponse
to persistent demands by readers and students over the years. The new subsections are
more advanced and are written in the concise style of the book, thereby raising its level to
thegraduatelevel.Manyexampleshavebeenadded,forexampleinChapters1and2,that
are often used in physics or are standard lore of physics courses. A number of additions
have been made in Chapter 3, such as on linear dependence of vectors, dual vector spaces
and spectral decomposition of symmetric or Hermitian matrices. A subsection on the dif-
fusion equation emphasizes methods to adapt solutions of partial differential equations to
boundaryconditions.NewformulashavebeendevelopedforHermitepolynomialsandare
includedinChapter13thatareusefulfortreatingmolecularvibrations;theyareofinterest
tothechemicalphysicists.
ACKNOWLEDGMENTS
Wehavebenefitedfromtheadviceandhelpofmanypeople.Someoftherevisionsareinre-
sponsetocommentsbyreadersandformerstudents,suchasDr.K.BodoorandJ.Hughes.
WearegratefultothemandtoourEditorsBarbaraHollandandTomSingerwhoorganized
accuracychecks.WewouldliketothankinparticularDr.MichaelBozoianandProf.Frank
Harris for their invaluable help with the accuracy checking and Simon Crump, Production
Editor,for hisexpertmanagementoftheSixthEdition.
CHAPTER 1
VECTOR ANALYSIS
1.1 D EFINITIONS ,ELEMENTARY APPROACH
In science and engineering we frequently encounter quantities that have magnitude and
magnitude only: mass, time, and temperature. These we label scalarquantities, which re-
main the same no matter what coordinates we use. In contrast, many interesting physical
quantities have magnitude and, in addition, an associated direction. This second group
includes displacement, velocity, acceleration, force, momentum, and angular momentum.
Quantitieswithmagnitudeanddirectionarelabeled vectorquantities.Usually,inelemen-
tary treatments, a vector is defined as a quantity having magnitude and direction. To dis-
tinguishvectorsfrom scalars,weidentifyvectorquantitieswithboldfacetype,thatis, V.
Ourvectormaybeconvenientlyrepresentedbyanarrow,withlengthproportionaltothe
magnitude. The direction of the arrow gives the direction of the vector, the positive sense
ofdirectionbeingindicatedbythepoint.Inthisrepresentation,vectoraddition
C=A+B (1.1)
consists in placing the rear end of vector Bat the point of vector A. VectorCis then
represented by an arrow drawn from the rear of Ato the point of B. This procedure, the
triangle law of addition, assigns meaning to Eq. (1.1) and is illustrated in Fig. 1.1. By
completingtheparallelogram,weseethat
C=A+B=B+A, (1.2)
asshowninFig.1.2.In words, vectoradditionis commutative .
Forthesumofthreevectors
D=A+B+C,
Fig.1.3,wemayfirstadd AandB:
A+B=E.
1
2 Chapter 1 Vector Analysis
FIGURE 1.1Trianglelawofvector
addition.
FIGURE 1.2Parallelogramlawof
vectoraddition.
FIGURE 1.3Vectoradditionis
associative.
Thenthis sumisaddedto C:
D=E+C.
Similarly,wemayfirst add BandC:
B+C=F.
Then
D=A+F.
Intermsof theoriginalexpression,
(A+B)+C=A+(B+C).
Vectoradditionis associative .
A direct physical example of the parallelogram addition law is provided by a weight
suspended by two cords. If the junction point ( Oin Fig. 1.4) is in equilibrium, the vector
1.1 Definitions, Elementary Approach 3
FIGURE 1.4Equilibriumofforces: F1+F2=−F3.
sum of the two forces F1andF2must just cancelthe downwardforce of gravity, F3.H e r e
theparallelogramadditionlawissubjecttoimmediateexperimentalverification.1
Subtraction may be handled by defining the negative of a vector as a vector of the same
magnitudebutwithreverseddirection.Then
A−B=A+(−B).
InFig.1.3,
A=E−B.
Notethatthevectorsaretreatedasgeometricalobjectsthatareindependentofanycoor-
dinatesystem.Thisconceptofindependenceofapreferredcoordinatesystemisdeveloped
indetailinthenextsection.
The representation of vector Aby an arrow suggests a second possibility. Arrow A
(Fig.1.5),startingfromtheorigin,2terminatesatthepoint (Ax,Ay,Az).Thus,ifweagree
that the vector is to start at the origin, the positive end may be specified by giving the
Cartesiancoordinates (Ax,Ay,Az)ofthearrowhead.
Although Acouldhaverepresentedanyvectorquantity(momentum,electricfield,etc.),
one particularly important vector quantity, the displacement from the origin to the point
1Strictly speaking, the parallelogram addition was introduced as a definition. Experiments show that if we assume that the
forcesarevectorquantitiesandwecombine thembyparallelogramaddition,theequilibriumcondition ofzeroresultantforceis
satisfied.
2We could start from any point in our Cartesian reference frame; we choose the origin for simplicity. This freedom of shifting
the origin of the coordinate system without affecting the geometry is called translation invariance .
4 Chapter 1 Vector Analysis
FIGURE 1.5Cartesiancomponentsanddirectioncosinesof A.
(x,y,z), isdenotedbythespecialsymbol r.Wethenhaveachoiceofreferringtothedis-
placementaseitherthevector rorthecollection (x,y,z), thecoordinatesofitsendpoint:
r↔(x,y,z). (1.3)
Usingrfor the magnitude of vector r, we find that Fig. 1.5 shows that the endpoint coor-
dinatesandthemagnitudearerelatedby
x=rcosα, y=rcosβ, z=rcosγ. (1.4)
Herecosα,cosβ,andcosγarecalledthe directioncosines ,αbeingtheanglebetweenthe
given vector and the positive x-axis, and so on. One further bit of vocabulary: The quan-
titiesAx,Ay, andAzare known as the (Cartesian) components ofAor theprojections
ofA,with cos2α+cos2β+cos2γ=1.
Thus, any vector Amay be resolved into its components (or projected onto the coordi-
nateaxes)toyield Ax=Acosα,etc.,asinEq.(1.4).Wemaychoosetorefertothevector
as a single quantity Aor to its components (Ax,Ay,Az). Note that the subscript xinAx
denotes the xcomponent and not a dependence on the variable x. The choice between
usingAor its components (Ax,Ay,Az)is essentially a choice between a geometric and
an algebraic representation. Use either representation at your convenience. The geometric
“arrowinspace”mayaidinvisualization.Thealgebraicsetofcomponentsisusuallymore
suitablefor precisenumericalor algebraiccalculations.
Vectors enter physics in two distinct forms. (1) Vector Amay represent a single force
acting at a single point. The force of gravity acting at the center of gravity illustrates this
form. (2) Vector Amay be defined over some extended region; that is, Aand its compo-
nents may be functions of position: Ax=Ax(x,y,z), and so on. Examples of this sort
includethevelocityofafluidvaryingfrompointtopointoveragivenvolumeandelectric
and magnetic fields. These two cases may be distinguished by referring to the vector de-
fined over a region as a vector field . The concept of the vector defined over a region and
1.1 Definitions, Elementary Approach 5
being a function of position will become extremely important when we differentiate and
integratevectors.
Atthisstageitisconvenienttointroduceunitvectorsalongeachofthecoordinateaxes.
Letˆxbe a vector of unit magnitude pointing in the positive x-direction,ˆy, a vector of unit
magnitude in the positive y-direction, and ˆza vector of unit magnitude in the positive z-
direction. Then ˆxAxis a vector with magnitude equal to |Ax|and in the x-direction. By
vectoraddition,
A=ˆxAx+ˆyAy+ˆzAz. (1.5)
Notethatif Avanishes,allof itscomponentsmustvanishindividually;thatis, if
A=0,thenAx=Ay=Az=0.
Thismeansthattheseunitvectorsserveasa basis,orcompletesetofvectors,inthethree-
dimensionalEuclideanspaceintermsofwhichanyvectorcanbeexpanded.Thus,Eq.(1.5)
isanassertionthatthethreeunitvectors ˆx,ˆy,andˆzspanourrealthree-dimensionalspace:
Any vector may be written as a linear combination of ˆx,ˆy, andˆz.Sinceˆx,ˆy, andˆzare
linearly independent (no one is a linear combination of the other two), they form a basis
for the real three-dimensional Euclidean space. Finally, by the Pythagorean theorem, the
magnitudeofvector Ais
|A|=parenleftbig
A2
x+A2
y+A2
zparenrightbig1/2. (1.6)
Notethatthecoordinateunitvectorsarenottheonlycompleteset,orbasis.Thisresolution
of a vector into its components can be carried out in a variety of coordinate systems, as
shown in Chapter 2. Here we restrict ourselves to Cartesian coordinates, where the unit
vectorshavethecoordinates ˆx=(1,0,0),ˆy=(0,1,0)andˆz=(0,0,1)andareallconstant
inlengthanddirection,propertiescharacteristicof Cartesiancoordinates.
As a replacement of the graphical technique, addition and subtraction of vectors may
now be carried out in terms of their components. For A=ˆxAx+ˆyAy+ˆzAzandB=
ˆxBx+ˆyBy+ˆzBz,
A±B=ˆx(Ax±Bx)+ˆy(Ay±By)+ˆz(Az±Bz). (1.7)
It should be emphasized here that the unit vectors ˆx,ˆy, andˆzare used for convenience.
They are not essential; we can describe vectors and use them entirely in terms of their
components: A↔(Ax,Ay,Az).This is the approach of the two more powerful, more
sophisticated definitions of vector to be discussed in the next section. However, ˆx,ˆy, and
ˆzemphasizethe direction .
So far we havedefinedtheoperationsof additionand subtractionof vectors. In thenext
sections,threevarietiesofmultiplicationwillbedefinedonthebasisoftheirapplicability:
a scalar, or inner, product, a vector product peculiar to three-dimensional space, and a
direct,orouter,productyieldingasecond-ranktensor.Divisionbya vectorisnotdefined.
6 Chapter 1 Vector Analysis
Exercises
1.1.1 Showhowtofind AandB,gi v enA+BandA−B.
1.1.2 The vector Awhose magnitude is 1 .732 units makes equal angles with the coordinate
axes.Find Ax,Ay, andAz.
1.1.3 Calculate the components of a unit vector that lies in the xy-plane and makes equal
angleswiththepositivedirectionsofthe x-andy-axes.
1.1.4 The velocity of sailboat Arelative to sailboat B,vrel, is defined by the equation vrel=
vA−vB, wherevAis the velocity of AandvBis the velocity of B. Determine the
velocityof Arelativeto Bif
vA=30km/hreast
vB=40km/hrnorth.
ANS.vrel=50 km/hr, 53.1◦southofeast.
1.1.5 A sailboat sails for 1 hr at 4 km /hr (relative to the water) on a steady compass heading
of 40◦east of north. The sailboat is simultaneously carried along by a current. At the
endofthehourtheboatis6.12kmfromitsstartingpoint.Thelinefromitsstartingpoint
toitslocationlies 60◦eastofnorth.Findthe x(easterly)and y(northerly)components
ofthewater’svelocity.
ANS.veast=2.73 km/hr,vnorth≈0k m/hr.
1.1.6 Avectorequationcanbereducedtotheform A=B.Fromthisshowthattheonevector
equation is equivalent to threescalar equations. Assuming the validity of Newton’s
second law, F=ma,a savectorequation, this means that axdepends only on Fxand
isindependentof FyandFz.
1.1.7 The vertices A,B, andCof a triangle are given by the points (−1,0,2), (0,1,0), and
(1,−1,0), respectively. Find point Dso that the figure ABCDforms a plane parallel-
ogram.
ANS.(0,−2,2)or(2,0,−2).
1.1.8 A triangle is defined by the vertices of three vectors A,BandCthat extend from the
origin. In terms of A,B, andCshow that the vectorsum of the successive sides of the
triangle(AB+BC+CA)iszero, wheretheside ABisfromAtoB,etc.
1.1.9 Asphereofradius ais centeredatapoint r1.
(a) Writeoutthealgebraicequationfor thesphere.
(b) Writeouta vectorequationforthesphere.
ANS. (a) (x−x1)2+(y−y1)2+(z−z1)2=a2.
(b)r=r1+a, withr1=center.
(atakes onalldirectionsbuthasafixedmagnitude a.)
1.2 Rotation of the Coordinate Axes 7
1.1.10 A corner reflector is formed by three mutually perpendicular reflecting surfaces. Show
that a ray of light incident upon the corner reflector (striking all three surfaces) is re-
flectedbackalongalineparalleltothelineofincidence.
Hint. Consider the effect of a reflection on the components of a vector describing the
directionofthelightray.
1.1.11 Hubble’s law . Hubble found that distant galaxies are receding with a velocity propor-
tionaltotheirdistancefromwhereweareonEarth.Forthe ithgalaxy,
vi=H0ri,
with us at the origin. Show that this recession of the galaxies from us does notimply
that we are at the center of the universe. Specifically, take the galaxy at r1as a new
originandshowthatHubble’slawisstillobeyed.
1.1.12 Findthediagonalvectorsofaunitcubewithonecornerattheoriginanditsthreesides
lying along Cartesian coordinates axes. Show that there are four diagonals with length√
3.Representingtheseasvectors,whataretheircomponents?Showthatthediagonals
ofthecube’sfaceshavelength√
2 anddeterminetheircomponents.
1.2 R OTATION OF THE COORDINATE AXES3
In the preceding section vectors were defined or represented in two equivalent ways:
(1) geometrically by specifying magnitude and direction, as with an arrow, and (2) al-
gebraically by specifying the components relative to Cartesian coordinate axes. The sec-
ond definition is adequate for the vector analysis of this chapter. In this section two more
refined, sophisticated, and powerful definitions are presented. First, the vector field is de-
finedintermsofthebehaviorofitscomponentsunderrotationofthecoordinateaxes.This
transformation theory approach leads into the tensor analysis of Chapter 2 and groups of
transformations in Chapter 4. Second, the component definition of Section 1.1 is refined
andgeneralizedaccordingtothemathematician’sconceptsofvectorandvectorspace.This
approachleadstofunctionspaces, includingtheHilbertspace.
The definition of vector as a quantity with magnitude and direction is incomplete. On
the one hand, we encounter quantities, such as elastic constants and index of refraction
in anisotropic crystals, that have magnitude and direction butthat are not vectors. On
the other hand, our naïve approach is awkward to generalize to extend to more complex
quantities. We seek a new definition of vector field using our coordinate vector ras a
prototype.
Thereisaphysicalbasisforourdevelopmentofanewdefinition.Wedescribeourphys-
ical world by mathematics, but it and any physical predictions we may make must be
independent ofourmathematicalconventions.
In our specific case we assume that space is isotropic; that is, there is no preferred di-
rection, or all directions are equivalent. Then the physical system being analyzed or the
physical law being enunciated cannot and must not depend on our choice or orientation
of the coordinate axes. Specifically, if a quantity Sdoes not depend on the orientation of
thecoordinateaxes,itis calledascalar.
3This sectionis optional here.It willbe essential for Chapter2.
8 Chapter 1 Vector Analysis
FIGURE 1.6Rotationof Cartesiancoordinateaxesaboutthe z-axis.
Now we return to the concept of vector ras a geometric object independent of the
coordinate system. Let us look at rin two different systems, one rotated in relation to the
other.
For simplicity we consider first the two-dimensional case. If the x-,y-coordinates are
rotated counterclockwise through an angle ϕ,keeping r ,fixed(Fig. 1.6), we get the fol-
lowing relations between the components resolved in the original system (unprimed) and
thoseresolvedinthenewrotatedsystem(primed):
x′=xcosϕ+ysinϕ,
y′=−xsinϕ+ycosϕ.(1.8)
We saw in Section 1.1 that a vector could be represented by the coordinates of a point;
thatis,thecoordinateswereproportionaltothevectorcomponents.Hencethecomponents
of a vector must transform under rotation as coordinates of a point (such as r). Therefore
wheneveranypairofquantities AxandAyinthexy-coordinatesystemistransformedinto
(A′
x,A′
y)bythisrotationof thecoordinatesystemwith
A′
x=Axcosϕ+Aysinϕ,
A′
y=−Axsinϕ+Aycosϕ,(1.9)
wedefine4AxandAyasthecomponentsofavector A.Ourvectornowisdefinedinterms
ofthetransformationofitscomponentsunderrotationofthecoordinatesystem.If Axand
Aytransforminthesamewayas xandy,thecomponentsofthegeneraltwo-dimensional
coordinatevector r,theyarethecomponentsofavector A.IfAxandAydonotshowthis
4A scalarquantity does not depend on the orientation of coordinates; S′=Sexpresses the fact that it is invariant under rotation
of thecoordinates.
1.2 Rotation of the Coordinate Axes 9
form invariance (also called covariance ) when the coordinates are rotated, they do not
formavector.
Thevectorfieldcomponents AxandAysatisfyingthedefiningequations,Eqs.(1.9),as-
sociate a magnitude Aand a direction with each point in space. The magnitude is a scalar
quantity, invariant to the rotation of the coordinate system. The direction (relative to the
unprimed system) is likewise invariant to the rotation of the coordinate system (see Exer-
cise 1.2.1). The result of all this is that the components of a vector may vary according to
therotationof theprimedcoordinatesystem. This iswhatEqs. (1.9) say. Butthevariation
withtheangleisjustsuchthatthecomponentsintherotatedcoordinatesystem A′
xandA′
y
define a vector with the same magnitude and the same direction as the vector defined by
thecomponents AxandAyrelativetothe x-,y-coordinateaxes.(CompareExercise1.2.1.)
The components of Ain a particular coordinate system constitute the representation of
Ain that coordinate system. Equations (1.9), the transformation relations, are a guarantee
thattheentity Ais independentoftherotationofthecoordinatesystem.
Togoontothreeand,later,fourdimensions,wefinditconvenienttouseamorecompact
notation.Let
x→x1
y→x2(1.10)
a11=cosϕ, a 12=sinϕ,
a21=−sinϕ, a 22=cosϕ.(1.11)
ThenEqs. (1.8) become
x′
1=a11x1+a12x2,
x′
2=a21x1+a22x2.(1.12)
Thecoefficient aijmaybeinterpretedasadirectioncosine,thecosineoftheanglebetween
x′
iandxj;thatis,
a12=cos(x′
1,x2)=sinϕ,
a21=cos(x′
2,x1)=cosparenleftbig
ϕ+π
2parenrightbig
=−sinϕ.(1.13)
The advantage of the new notation5is that it permits us to use the summation symbolsummationtext
andtorewriteEqs. (1.12) as
x′
i=2summationdisplay
j=1aijxj,i=1,2. (1.14)
Note that iremains as a parameter that gives rise to one equation when it is set equal to 1
and to a second equation when it is set equal to 2. The index j, of course, is a summation
index, a dummy index, and, as with a variable of integration, jmay be replaced by any
otherconvenientsymbol.
5You may wonder at the replacement of one parameter ϕby four parameters aij.C l e a r l y ,t h e aijdo not constitute a minimum
set of parameters. For two dimensions the four aijare subject to the three constraints given in Eq. (1.18). The justification for
this redundant set of direction cosines is the convenience it provides. Hopefully, this convenience will become more apparent
in Chapters 2 and 3. For three-dimensional rotations (9 aijbut only three independent) alternate descriptions are provided by:
(1)theEuleranglesdiscussedinSection3.3,(2)quaternions,and(3)theCayley–Kleinparameters.Thesealternativeshavetheir
respective advantages anddisadvantages.
10 Chapter 1 Vector Analysis
Thegeneralizationtothree,four,or Ndimensionsisnowsimple.Thesetof Nquantities
Vjis said to be the components of an N-dimensional vector Vif and only if their values
relativetotherotatedcoordinateaxesaregivenby
V′
i=Nsummationdisplay
j=1aijVj,i=1,2,...,N. (1.15)
As before, aijis the cosine of the angle between x′
iandxj. Often the upper limit Nand
the corresponding range of iwill not be indicated. It is taken for granted that you know
howmanydimensionsyourspacehas.
From the definition of aijas the cosine of the angle between the positive x′
idirection
andthepositive xjdirectionwemaywrite(Cartesiancoordinates)6
aij=∂x′
i
∂xj. (1.16a)
Usingtheinverserotation( ϕ→−ϕ) yields
xj=2summationdisplay
i=1aijx′
ior∂xj
∂x′
i=aij. (1.16b)
Note that these are partial derivatives . By use of Eqs. (1.16a) and (1.16b), Eq. (1.15)
becomes
V′
i=Nsummationdisplay
j=1∂x′
i
∂xjVj=Nsummationdisplay
j=1∂xj
∂x′
iVj. (1.17)
Thedirectioncosines aijsatisfyan orthogonalitycondition
summationdisplay
iaijaik=δjk (1.18)
or,equivalently,
summationdisplay
iajiaki=δjk. (1.19)
Here,thesymbol δjkis theKroneckerdelta,definedby
δjk=1forj=k,
δjk=0forj/negationslash=k.(1.20)
It is easily verified that Eqs. (1.18) and (1.19) hold in the two-dimensional case by
substituting in the specific aijfrom Eqs. (1.11). The result is the well-known identity
sin2ϕ+cos2ϕ=1 for the nonvanishing case. To verify Eq. (1.18) in general form, we
mayusethepartialderivativeforms ofEqs. (1.16a) and(1.16b) toobtain
summationdisplay
i∂xj
∂x′
i∂xk
∂x′
i=summationdisplay
i∂xj
∂x′
i∂x′
i
∂xk=∂xj
∂xk. (1.21)
6Differentiate x′
iwithrespect to xj. Seediscussion following Eq.(1.21).
1.2 Rotation of the Coordinate Axes 11
The last step follows by the standard rules for partial differentiation, assuming that xjis
af u n c t i o no f x′
1,x′
2,x′
3, and so on. The final result, ∂xj/∂xk, is equal to δjk, sincexjand
xkas coordinate lines ( j/negationslash=k) are assumed to be perpendicular (two or three dimensions)
or orthogonal (for any number of dimensions). Equivalently, we may assume that xjand
xk(j/negationslash=k)aretotallyindependentvariables.If j=k,thepartialderivativeisclearlyequal
to 1.
In redefining a vector in terms of how its components transform under a rotation of the
coordinatesystem, weshouldemphasizetwopoints:
1. This definition is developed because it is useful and appropriate in describing our
physical world. Our vectorequationswill be independentof anyparticular coordinate
system. (The coordinate system need not even be Cartesian.) The vector equation can
always be expressed in some particular coordinate system, and, to obtain numerical
results, wemustultimatelyexpresstheequationinsomespecificcoordinatesystem.
2. Thisdefinitionissubjecttoageneralizationthatwillopenupthebranchofmathemat-
ics knownastensoranalysis(Chapter2).
A qualification is in order. The behavior of the vector components under rotation of the
coordinates is used in Section 1.3 to prove that a scalar product is a scalar, in Section 1.4
to prove that a vector product is a vector, and in Section 1.6 to show that the gradient of a
scalarψ,∇ψ, is a vector. The remainder of this chapter proceeds on the basis of the less
restrictivedefinitionsofthevectorgiveninSection1.1.
Summary: Vectors and Vector Space
It is customary in mathematics to label an ordered triple of real numbers ( x1,x2,x3)a
vector x. The number xnis called the nth component of vector x. The collection of all
such vectors (obeying the properties that follow) form a three-dimensional real vector
space.Weascribefivepropertiestoourvectors:If x=(x1,x2,x3)andy=(y1,y2,y3),
1. Vectorequality: x=ymeansxi=yi,i=1,2,3.
2. Vectoraddition: x+y=zmeansxi+yi=zi,i=1,2,3.
3. Scalarmultiplication: ax↔(ax1,ax2,ax3)(withareal).
4. Negativeofavector: −x=(−1)x↔(−x1,−x2,−x3).
5. Nullvector:Thereexistsanullvector 0↔(0,0,0).
Since our vector components are real (or complex) numbers, the following properties
alsohold:
1. Additionofvectorsiscommutative: x+y=y+x.
2. Additionofvectorsisassociative: (x+y)+z=x+(y+z).
3. Scalarmultiplicationis distributive:
a(x+y)=ax+ay,also(a+b)x=ax+bx.
4. Scalarmultiplicationis associative: (ab)x=a(bx).
12 Chapter 1 Vector Analysis
Further,thenullvector 0isunique,asis thenegativeofagivenvector x.
Sofarasthevectorsthemselvesareconcernedthisapproachmerelyformalizesthecom-
ponentdiscussionofSection1.1.Theimportanceliesintheextensions,whichwillbecon-
sidered in later chapters. In Chapter 4, we show that vectors form both an Abelian group
underadditionandalinearspacewiththetransformationsinthelinearspacedescribedby
matrices.Finally,andperhapsmostimportant,foradvancedphysicstheconceptofvectors
presentedheremaybegeneralizedto(1)complexquantities,7(2)functions,and(3)aninfi-
nitenumberofcomponents.Thisleadstoinfinite-dimensionalfunctionspaces,theHilbert
spaces, which are important in modern quantum theory. A brief introduction to function
expansionsandHilbertspaceappearsinSection10.4.
Exercises
1.2.1 (a) Show that the magnitude of a vector A,A=(A2
x+A2
y)1/2, is independent of the
orientationof therotatedcoordinatesystem,
parenleftbig
A2
x+A2
yparenrightbig1/2=parenleftbig
A′2
x+A′2
yparenrightbig1/2,
thatis, independentoftherotationangle ϕ.
This independence of angle is expressed by saying that Aisinvariant under
rotations.
(b) At a given point (x,y),Adefines an angle αrelative to the positive x-axis and
α′relative to the positive x′-axis. The angle from xtox′isϕ. Show that A=A′
definesthe samedirectioninspacewhenexpressedintermsofitsprimedcompo-
nentsas intermsofits unprimedcomponents;thatis,
α′=α−ϕ.
1.2.2 Prove the orthogonality conditionsummationtext
iajiaki=δjk. As a special case of this, the direc-
tioncosinesof Section1.1 satisfytherelation
cos2α+cos2β+cos2γ=1,
aresultthatalsofollowsfromEq. (1.6).
1.3 S CALAR OR DOTPRODUCT
Havingdefinedvectors,wenowproceedtocombinethem.Thelawsforcombiningvectors
mustbemathematicallyconsistent.Fromthepossibilitiesthatareconsistentweselecttwo
thatarebothmathematicallyandphysicallyinteresting.Athirdpossibilityisintroducedin
Chapter2, inwhichweformtensors.
The projection of a vector Aonto a coordinate axis, which gives its Cartesian compo-
nents in Eq. (1.4), defines a special geometrical case of the scalar product of Aand the
coordinateunitvectors:
Ax=Acosα≡A·ˆx,Ay=Acosβ≡A·ˆy,Az=Acosγ≡A·ˆz.(1.22)
7Then-dimensionalvectorspaceofreal n-tuplesisoftenlabeled Rnandthen-dimensionalvectorspaceofcomplex n-tuplesis
labeled Cn.
1 . 3 S c a l a ro rD o tP r o d u c t 1 3
Thisspecialcaseofascalarproductinconjunctionwithgeneralpropertiesthescalarprod-
uctis sufficienttoderivethegeneralcaseofthescalarproduct.
Just as the projection is linear in A, we want the scalar product of two vectors to be
linearinAandB, thatis, obeythedistributiveandassociativelaws
A·(B+C)=A·B+A·C (1.23a)
A·(yB)=(yA)·B=yA·B, (1.23b)
whereyisanumber.Nowwecanusethedecompositionof BintoitsCartesiancomponents
accordingtoEq.(1.5), B=Bxˆx+Byˆy+Bzˆz,toconstructthegeneralscalarordotproduct
ofthevectors AandBas
A·B=A·(Bxˆx+Byˆy+Bzˆz)
=BxA·ˆx+ByA·ˆy+BzA·ˆzuponapplyingEqs. (1.23a)and(1.23b)
=BxAx+ByAy+BzAzuponsubstitutingEq.(1.22).
Hence
A·B≡summationdisplay
iBiAi=summationdisplay
iAiBi=B·A. (1.24)
IfA=Bin Eq. (1.24), we recover the magnitude A=(summationtextA2
i)1/2ofAin Eq. (1.6) from
Eq.(1.24).
It is obvious from Eq. (1.24) that the scalar product treats AandBalike, or is sym-
metric in AandB, and is commutative. Thus, alternatively and equivalently, we can first
generalize Eqs. (1.22) to the projection ABofAonto the direction of a vector B/negationslash=0
asAB=Acosθ≡A·ˆB, whereˆB=B/Bis the unit vector in the direction of Bandθ
is the angle between AandB, as shown in Fig. 1.7. Similarly, we project BontoAas
BA=Bcosθ≡B·ˆA. Second, we make these projections symmetric in AandB, which
leadstothedefinition
A·B≡ABB=ABA=ABcosθ. (1.25)
FIGURE 1.7Scalarproduct A·B=ABcosθ.
14 Chapter 1 Vector Analysis
FIGURE 1.8Thedistributivelaw
A·(B+C)=ABA+ACA=A(B+C)A, Eq. (1.23a).
ThedistributivelawinEq.(1.23a)isillustratedinFig.1.8,whichshowsthatthesumof
the projections of BandContoA,BA+CAis equal to the projection of B+ContoA,
(B+C)A.
ItfollowsfromEqs.(1.22),(1.24),and(1.25)thatthecoordinateunitvectorssatisfythe
relations
ˆx·ˆx=ˆy·ˆy=ˆz·ˆz=1, (1.26a)
whereas
ˆx·ˆy=ˆx·ˆz=ˆy·ˆz=0. (1.26b)
Ifthecomponentdefinition,Eq.(1.24),islabeledanalgebraicdefinition,thenEq.(1.25)
is a geometric definition. One of the most common applications of the scalar product in
physics is in the calculation of work=force·displacement·cosθ, which is interpreted as
displacement times the projection of the force along the displacement direction, i.e., the
scalarproductofforce anddisplacement, W=F·S.
IfA·B=0 and we know that A/negationslash=0 andB/negationslash=0, then, from Eq. (1.25), cos θ=0, or
θ=90◦,270◦, and so on. The vectors AandBmust be perpendicular. Alternately, we
may sayAandBare orthogonal. The unit vectors ˆx,ˆy, andˆzare mutually orthogonal. To
developthis notionof orthogonalityonemorestep, supposethat nis a unitvectorand ris
anonzerovectorinthe xy-plane;thatis, r=ˆxx+ˆyy(Fig. 1.9). If
n·r=0
forallchoicesof r,thennmustbeperpendicular(orthogonal)tothe xy-plane.
Often it is convenient to replace ˆx,ˆy, andˆzby subscripted unit vectors em,m=1,2,3,
withˆx=e1, andso on.ThenEqs. (1.26a)and(1.26b)become
em·en=δmn. (1.26c)
Form/negationslash=nthe unit vectors emandenare orthogonal. For m=neach vector is normal-
ized to unity, that is, has unit magnitude. The set emis said to be orthonormal .Am a j o r
advantage of Eq. (1.26c) over Eqs. (1.26a) and (1.26b) is that Eq. (1.26c) may readily be
generalized to N-dimensional space: m,n=1,2,...,N. Finally, we are picking sets of
unitvectors emthatareorthonormalforconvenience–averygreatconvenience.
1 . 3 S c a l a ro rD o tP r o d u c t 1 5
FIGURE 1.9Anormalvector.
Invariance of the Scalar Product Under Rotations
We havenot yet shown that the word scalaris justifiedor that the scalar productis indeed
a scalar quantity. To do this, we investigate the behavior of A·Bunder a rotation of the
coordinatesystem. ByuseofEq. (1.15),
A′
xB′
x+A′
yB′
y+A′
zB′
z=summationdisplay
iaxiAisummationdisplay
jaxjBj+summationdisplay
iayiAisummationdisplay
jayjBj
+summationdisplay
iaziAisummationdisplay
jazjBj. (1.27)
Usingtheindices kandlto sum over x,y,andz, weobtain
summationdisplay
kA′
kB′
k=summationdisplay
lsummationdisplay
isummationdisplay
jaliAialjBj, (1.28)
and,byrearrangingthetermsontheright-handside, wehave
summationdisplay
kA′
kB′
k=summationdisplay
lsummationdisplay
isummationdisplay
j(alialj)AiBj=summationdisplay
isummationdisplay
jδijAiBj=summationdisplay
iAiBi.(1.29)
The last two steps follow by using Eq. (1.18), the orthogonality condition of the direction
cosines, and Eqs. (1.20), which define the Kronecker delta. The effect of the Kronecker
delta is to cancel all terms in a summation over either index except the term for which the
indices are equal. In Eq. (1.29) its effect is to set j=iand to eliminate the summation
overj. Of course, we could equally well set i=jand eliminate the summation over i.
16 Chapter 1 Vector Analysis
Equation(1.29)givesus
summationdisplay
kA′
kB′
k=summationdisplay
iAiBi, (1.30)
whichisjustourdefinitionofascalarquantity,onethatremainsinvariantundertherotation
ofthecoordinatesystem.
In a similar approach that exploits this concept of invariance, we take C=A+Band
dotit intoitself:
C·C=(A+B)·(A+B)
=A·A+B·B+2A·B. (1.31)
Since
C·C=C2, (1.32)
thesquareof themagnitudeofvector Candthusaninvariantquantity,weseethat
A·B=1
2parenleftbig
C2−A2−B2parenrightbig
,invariant. (1.33)
Since the right-hand side of Eq. (1.33) is invariant—that is, a scalar quantity—the left-
hand side, A·B, must also be invariant under rotation of the coordinate system. Hence
A·Bisascalar.
Equation(1.31) isreallyanotherformof thelawofcosines,whichis
C2=A2+B2+2ABcosθ. (1.34)
Comparing Eqs. (1.31) and (1.34), we have another verification of Eq. (1.25), or, if pre-
ferred, avectorderivationofthelawof cosines(Fig.1.10).
The dot product, given by Eq. (1.24), may be generalized in two ways. The space need
not be restricted to three dimensions. In n-dimensional space, Eq. (1.24) applies with the
sumrunningfrom1to n.Moreover, nmaybeinfinity,withthesumthenaconvergentinfi-
niteseries(Section5.2).Theothergeneralizationextendstheconceptofvectortoembrace
functions.The functionanalogofadot,or inner,productappearsinSection10.4.
FIGURE 1.10Thelawofcosines.
1 . 3 S c a l a ro rD o tP r o d u c t 1 7
Exercises
1.3.1 Twounitmagnitudevectors eiandejarerequiredtobeeitherparallelorperpendicular
to each other. Show that ei·ejprovides an interpretation of Eq. (1.18), the direction
cosineorthogonalityrelation.
1.3.2 Given that (1) the dot product of a unit vector with itself is unity and (2) this relation is
valid in all (rotated) coordinate systems, show that ˆx′·ˆx′=1 (with the primed system
rotated 45◦aboutthe z-axisrelativetotheunprimed)impliesthat ˆx·ˆy=0.
1.3.3 Thevector r,startingattheorigin,terminatesatandspecifiesthepointinspace (x,y,z).
Findthesurface sweptoutbythetipof rif
(a)(r−a)·a=0.Characterize ageometrically.
(b)(r−a)·r=0.Describethegeometricroleof a.
Thevector aisconstant(in magnitudeanddirection).
1.3.4 The interaction energy between two dipoles of moments µ1andµ2may be written in
thevectorform
V=−µ1·µ2
r3+3(µ1·r)(µ2·r)
r5
andinthescalarform
V=µ1µ2
r3(2cosθ1cosθ2−sinθ1sinθ2cosϕ).
Hereθ1andθ2are the angles of µ1andµ2relative to r, whileϕis the azimuth of µ2
relativetothe µ1–rplane(Fig.1.11). Showthatthesetwoforms areequivalent.
Hint:Equation(12.178)willbehelpful.
1.3.5 A pipe comes diagonally down the south wall of a building, making an angle of 45◦
withthehorizontal.Comingintoacorner,thepipeturnsandcontinuesdiagonallydown
a west-facing wall, still making an angle of 45◦with the horizontal. What is the angle
betweenthesouth-wallandwest-wallsectionsofthepipe?
ANS. 120◦.
1.3.6 Find the shortest distance of an observer at the point (2,1,3)from a rocket in free
flight with velocity (1,2,3)m/s. The rocket was launched at time t=0f r o m(1,1,1).
Lengthsareinkilometers.
1.3.7 Prove the law of cosines from the triangle with corners at the point of CandAin
Fig.1.10andtheprojectionofvector Bontovector A.
FIGURE 1.11Twodipolemoments.
18 Chapter 1 Vector Analysis
1.4 V ECTOR OR CROSS PRODUCT
A second form of vector multiplication employs the sine of the included angle instead
of the cosine. For instance, the angular momentum of a body shown at the point of the
distancevectorinFig.1.12is definedas
angularmomentum =radiusarm×linearmomentum
=distance×linearmomentum ×sinθ.
For convenience in treating problems relating to quantities such as angular momentum,
torque,andangularvelocity,wedefinethevectorproduct,orcross product,as
C=A×B,withC=ABsinθ. (1.35)
Unlike the preceding case of the scalar product, Cis now a vector, and we assign it a
directionperpendiculartotheplaneof AandBsuchthat A,B,andCformaright-handed
system.Withthischoiceof directionwehave
A×B=−B×A,anticommutation . (1.36a)
Fromthisdefinitionof cross productwehave
ˆx׈x=ˆy׈y=ˆz׈z=0, (1.36b)
whereas
ˆx׈y=ˆz,ˆy׈z=ˆx,ˆz׈x=ˆy,
ˆy׈x=−ˆz,ˆz׈y=−ˆx,ˆx׈z=−ˆy.(1.36c)
Amongtheexamplesofthecrossproductinmathematicalphysicsaretherelationbetween
linearmomentum pandangularmomentum L,withLdefinedas
L=r×p,
FIGURE 1.12Angularmomentum.
1 . 4 V e c t o ro rC r o s sP r o d u c t 1 9
FIGURE 1.13Parallelogramrepresentationofthevectorproduct.
andtherelationbetweenlinearvelocity vandangularvelocity ω,
v=ω×r.
Vectorsvandpdescribe properties of the particle or physical system. However, the posi-
tion vector ris determined by the choice of the origin of the coordinates. This means that
ωandLdependonthechoiceoftheorigin.
The familiar magnetic induction Bis usually defined by the vector product force equa-
tion8
FM=qv×B(mksunits) .
Herevis the velocity of the electric charge qandFMis the resulting force on the moving
charge.
The cross product has an important geometrical interpretation, which we shall use in
subsequent sections. In the parallelogram defined by AandB(Fig. 1.13), Bsinθis the
height if Ais taken as the length of the base. Then |A×B|=ABsinθis theareaof the
parallelogram.Asavector, A×Bistheareaoftheparallelogramdefinedby AandB,with
the area vector normal to the plane of the parallelogram. This suggests that area (with its
orientationinspace)maybetreatedasavectorquantity.
An alternate definition of the vector product can be derived from the special case of the
coordinateunitvectorsinEqs.(1.36c)inconjunctionwiththelinearityofthecrossproduct
inbothvectorarguments,inanalogywithEqs. (1.23) forthedotproduct,
A×(B+C)=A×B+A×C, (1.37a)
(A+B)×C=A×C+B×C, (1.37b)
A×(yB)=yA×B=(yA)×B, (1.37c)
8Theelectricfield Eis assumed hereto be zero.
20 Chapter 1 Vector Analysis
whereyisanumberagain.Usingthedecompositionof AandBintotheirCartesiancom-
ponentsaccordingtoEq.(1.5), wefind
A×B≡C=(Cx,Cy,Cz)=(Axˆx+Ayˆy+Azˆz)×(Bxˆx+Byˆy+Bzˆz)
=(AxBy−AyBx)ˆx׈y+(AxBz−AzBx)ˆx׈z
+(AyBz−AzBy)ˆy׈z
upon applying Eqs. (1.37a) and (1.37b) and substituting Eqs. (1.36a), (1.36b), and (1.36c)
sothattheCartesiancomponentsof A×Bbecome
Cx=AyBz−AzBy,C y=AzBx−AxBz,C z=AxBy−AyBx,(1.38)
or
Ci=AjBk−AkBj, i,j,k alldifferent , (1.39)
andwithcyclicpermutationoftheindices i,j,andkcorrespondingto x,y,andz,respec-
tively.Thevectorproduct Cmaybemnemonicallyrepresentedbyadeterminant,9
C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
AxAyAz
BxByBzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle≡ˆxvextendsinglevextendsinglevextendsinglevextendsingleAyAz
ByBzvextendsinglevextendsinglevextendsinglevextendsingle−ˆyvextendsinglevextendsinglevextendsinglevextendsingleAxAz
BxBzvextendsinglevextendsinglevextendsinglevextendsingle+ˆzvextendsinglevextendsinglevextendsinglevextendsingleAxAy
BxByvextendsinglevextendsinglevextendsinglevextendsingle,(1.40)
whichismeanttobeexpandedacrossthetoprowtoreproducethethreecomponentsof C
listedinEqs. (1.38).
Equation (1.35) might be called a geometric definition of the vector product. Then
Eqs. (1.38)wouldbeanalgebraicdefinition.
To show the equivalence of Eq. (1.35) and the component definition, Eqs. (1.38), let us
formA·CandB·C,usingEqs. (1.38). Wehave
A·C=A·(A×B)
=Ax(AyBz−AzBy)+Ay(AzBx−AxBz)+Az(AxBy−AyBx)
=0. (1.41)
Similarly,
B·C=B·(A×B)=0. (1.42)
Equations (1.41) and (1.42) show that Cis perpendicular to both AandB(cosθ=0,θ=
±90◦)and therefore perpendicular to the plane they determine. The positive direction is
determinedbyconsideringspecialcases,suchastheunitvectors ˆx׈y=ˆz(Cz=+AxBy).
Themagnitudeis obtainedfrom
(A×B)·(A×B)=A2B2−(A·B)2
=A2B2−A2B2cos2θ
=A2B2sin2θ. (1.43)
9SeeSection 3.1 for abrief summary of determinants.
1 . 4 V e c t o ro rC r o s sP r o d u c t 2 1
Hence
C=ABsinθ. (1.44)
The first step in Eq. (1.43) may be verified by expanding out in component form, using
Eqs. (1.38) for A×Band Eq. (1.24) for the dot product. From Eqs. (1.41), (1.42), and
(1.44)weseetheequivalenceofEqs.(1.35)and(1.38),thetwodefinitionsofvectorprod-
uct.
There still remains the problem of verifying that C=A×Bis indeed a vector, that
is, that it obeys Eq. (1.15), the vector transformation law. Starting in a rotated (primed
system),
C′
i=A′
jB′
k−A′
kB′
j,i,j,andkincyclicorder ,
=summationdisplay
lajlAlsummationdisplay
makmBm−summationdisplay
laklAlsummationdisplay
majmBm
=summationdisplay
l,m(ajlakm−aklajm)AlBm. (1.45)
Thecombinationofdirectioncosinesinparenthesesvanishesfor m=l.Wethereforehave
jandktaking on fixed values, dependent on the choice of i, and six combinations of
landm.I fi=3, thenj=1,k=2 (cyclic order), and we have the following direction
cosinecombinations:10
a11a22−a21a12=a33,
a13a21−a23a11=a32,
a12a23−a22a13=a31(1.46)
and their negatives. Equations (1.46) are identities satisfied by the direction cosines. They
maybeverifiedwiththeuseofdeterminantsandmatrices(seeExercise3.3.3).Substituting
backintoEq.(1.45),
C′
3=a33A1B2+a32A3B1+a31A2B3−a33A2B1−a32A1B3−a31A3B2
=a31C1+a32C2+a33C3
=summationdisplay
na3nCn. (1.47)
By permuting indices to pick up C′
1andC′
2, we see that Eq. (1.15) is satisfied and Cis
indeed a vector. It should be mentioned here that this vector nature of thecross product
isanaccidentassociatedwiththe three-dimensional natureofordinaryspace.11Itwillbe
seeninChapter2thatthecrossproductmayalsobetreatedasasecond-rankantisymmetric
tensor.
10Equations(1.46)holdforrotationsbecausetheypreservevolumes.Foramoregeneralorthogonaltransformation,ther.h.s.of
Eqs. (1.46) is multiplied by the determinant ofthe transformation matrix (see Chapter 3 for matricesand determinants).
11SpecificallyEqs.(1.46)holdonlyforthree-dimensionalspace.SeeD.HestenesandG.Sobczyk, CliffordAlgebratoGeometric
Calculus (Dordrecht: Reidel, 1984) for afar-reaching generalization ofthe cross product.
22 Chapter 1 Vector Analysis
If we define a vector as an ordered triplet of numbers (or functions), as in the latter part
ofSection1.2,thenthereisnoproblemidentifyingthecrossproductasavector.Thecross-
product operation maps the two triples AandBinto a third triple, C, which by definition
isavector.
We now have two ways of multiplying vectors; a third form appears in Chapter 2. But
what about division by a vector? It turns out that the ratio B/Ais not uniquely specified
(Exercise 3.2.21) unless AandBare also required to be parallel. Hence division of one
vectorbyanotheris notdefined.
Exercises
1.4.1 Showthatthemediansofatriangleintersectinthecenter,whichis 2 /3ofthemedian’s
lengthfromeachcorner.Constructanumericalexampleandplotit.
1.4.2 Provethelawofcosinesstartingfrom A2=(B−C)2.
1.4.3 Startingwith C=A+B,showthat C×C=0 leadsto
A×B=−B×A.
1.4.4 Showthat
(a)(A−B)·(A+B)=A2−B2,
(b)(A−B)×(A+B)=2A×B.
Thedistributivelawsneededhere,
A·(B+C)=A·B+A·C,
and
A×(B+C)=A×B+A×C,
mayeasilybeverified(ifdesired) byexpansioninCartesiancomponents.
1.4.5 Giventhethreevectors,
P=3ˆx+2ˆy−ˆz,
Q=−6ˆx−4ˆy+2ˆz,
R=ˆx−2ˆy−ˆz,
findtwo thatareperpendicularandtwothatareparallelor antiparallel.
1.4.6 IfP=ˆxPx+ˆyPyandQ=ˆxQx+ˆyQyare any two nonparallel (also nonantiparallel)
vectorsinthe xy-plane,showthat P×Qisinthez-direction.
1.4.7 Provethat (A×B)·(A×B)=(AB)2−(A·B)2.
1 . 4 V e c t o ro rC r o s sP r o d u c t 2 3
1.4.8 Usingthevectors
P=ˆxcosθ+ˆysinθ,
Q=ˆxcosϕ−ˆysinϕ,
R=ˆxcosϕ+ˆysinϕ,
provethefamiliartrigonometricidentities
sin(θ+ϕ)=sinθcosϕ+cosθsinϕ,
cos(θ+ϕ)=cosθcosϕ−sinθsinϕ.
1.4.9 (a) Findavector Athatis perpendicularto
U=2ˆx+ˆy−ˆz,
V=ˆx−ˆy+ˆz.
(b) What is Aif, in addition to this requirement, we demand that it have unit magni-
tude?
1.4.10 If fourvectors a,b,c,anddalllieinthesameplane,showthat
(a×b)×(c×d)=0.
Hint.Considerthedirectionsofthecross-productvectors.
1.4.11 The coordinates of the three vertices of a triangle are (2,1,5),(5,2,8),and(4,8,2).
Computeitsareabyvectormethods,itscenterandmedians.Lengthsareincentimeters.
Hint.SeeExercise1.4.1.
1.4.12 The vertices of parallelogram ABCDare(1,0,0),(2,−1,0),(0,−1,1), and(−1,0,1)
in order. Calculate the vector areas of triangle ABDand of triangle BCD.A r et h et w o
vectorareasequal?
ANS. Area ABD=−1
2(ˆx+ˆy+2ˆz).
1.4.13 The origin and the three vectors A,B, andC(all of which start at the origin) define a
tetrahedron. Taking the outward direction as positive, calculate the total vector area of
thefourtetrahedralsurfaces.
Note.In Section1.11thisresult isgeneralizedtoanyclosedsurface.
1.4.14 Findthesidesandanglesof thesphericaltriangle ABCdefinedbythethreevectors
A=(1,0,0),
B=parenleftbigg1√
2,0,1√
2parenrightbigg
,
C=parenleftbigg
0,1√
2,1√
2parenrightbigg
.
Eachvectorstarts fromtheorigin(Fig. 1.14).
24 Chapter 1 Vector Analysis
FIGURE 1.14Sphericaltriangle.
1.4.15 Derivethelawof sines(Fig.1.15):
sinα
|A|=sinβ
|B|=sinγ
|C|.
1.4.16 Themagneticinduction BisdefinedbytheLorentzforceequation,
F=q(v×B).
Carryingoutthreeexperiments,wefindthatif
v=ˆx,F
q=2ˆz−4ˆy,
v=ˆy,F
q=4ˆx−ˆz,
v=ˆz,F
q=ˆy−2ˆx.
Fromtheresultsofthesethreeseparateexperimentscalculatethemagneticinduction B.
1.4.17 Define a cross product of two vectors in two-dimensional space and give a geometrical
interpretationof yourconstruction.
1.4.18 Find the shortest distance between the paths of two rockets in free flight. Take the first
rocket path to be r=r1+t1v1with launch at r1=(1,1,1)and velocity v1=(1,2,3)
1.5 Triple Scalar Product, Triple Vector Product 25
FIGURE 1.15Lawofsines.
and the second rocket path as r=r2+t2v2withr2=(5,2,1)andv2=(−1,−1,1).
Lengthsareinkilometers,velocitiesinkilometersperhour.
1.5 T RIPLE SCALAR PRODUCT ,TRIPLE VECTOR PRODUCT
Triple Scalar Product
Sections 1.3 and 1.4 cover the two types of multiplication of interest here. However, there
arecombinationsofthreevectors, A·(B×C)andA×(B×C),thatoccurwithsufficient
frequencytodeservefurtherattention.Thecombination
A·(B×C)
is known as the triple scalar product .B×Cyields a vector that, dotted into A,g i v e sa
scalar.Wenotethat (A·B)×Crepresentsascalarcrossedintoavector,anoperationthat
is not defined. Hence, if we agree to exclude this undefined interpretation, the parentheses
maybeomittedandthetriplescalarproductwritten A·B×C.
UsingEqs. (1.38)for thecross productandEq. (1.24) forthedotproduct,weobtain
A·B×C=Ax(ByCz−BzCy)+Ay(BzCx−BxCz)+Az(BxCy−ByCx)
=B·C×A=C·A×B
=−A·C×B=−C·B×A=−B·A×C,andsoon . (1.48)
There is a high degree of symmetry in the component expansion. Every term contains the
factorsAi,Bj,andCk.Ifi,j,andkareincyclicorder (x,y,z),thesignispositive.Ifthe
orderisanticyclic,thesignisnegative.Further,thedotandthecrossmaybeinterchanged,
A·B×C=A×B·C. (1.49)
26 Chapter 1 Vector Analysis
FIGURE 1.16Parallelepipedrepresentationof triplescalarproduct.
A convenient representation of the component expansion of Eq. (1.48) is provided by the
determinant
A·B×C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleAxAyAz
BxByBz
CxCyCzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (1.50)
The rules for interchanging rows and columns of a determinant12provide an immediate
verification of the permutations listed in Eq. (1.48), whereas the symmetry of A,B, and
Cin the determinant form suggests the relation given in Eq. (1.49). The triple products
encounteredin Section 1.4, which showed that A×Bwas perpendicular to both AandB,
werespecialcasesofthegeneralresult (Eq. (1.48)).
Thetriplescalarproducthasadirectgeometricalinterpretation.Thethreevectors A,B,
andCmaybeinterpretedasdefiningaparallelepiped(Fig.1.16):
|B×C|=BCsinθ
=areaof parallelogrambase. (1.51)
The direction, of course, is normal to the base. Dotting Ainto this means multiplying the
baseareabytheprojectionof Aontothenormal,or basetimesheight.Therefore
A·B×C=volumeofparallelepipeddefinedby A,B,andC.
The triple scalar product finds an interesting and important application in the construc-
tion of a reciprocalcrystal lattice. Let a,b, andc(not necessarily mutuallyperpendicular)
12SeeSection 3.1 for a summary of the properties of determinants.
1.5 Triple Scalar Product, Triple Vector Product 27
represent the vectors that define a crystal lattice. The displacement from one lattice point
toanothermaythenbewritten
r=naa+nbb+ncc, (1.52)
withna,nb, andnctakingonintegralvalues.Withthesevectorswemayform
a′=b×c
a·b×c,b′=c×a
a·b×c,c′=a×b
a·b×c. (1.53a)
We see that a′is perpendicular to the plane containing bandc, and we can readily show
that
a′·a=b′·b=c′·c=1, (1.53b)
whereas
a′·b=a′·c=b′·a=b′·c=c′·a=c′·b=0. (1.53c)
It is from Eqs. (1.53b) and (1.53c) that the name reciprocal lattice is associated with the
pointsr′=n′
aa′+n′
bb′+n′
cc′.Themathematicalspaceinwhichthisreciprocallatticeex-
istsissometimescalleda Fourierspace ,onthebasisofrelationstotheFourieranalysisof
Chapters14and15.Thisreciprocallatticeisusefulinproblemsinvolvingthescatteringof
wavesfromthevariousplanesinacrystal.FurtherdetailsmaybefoundinR.B.Leighton’s
PrinciplesofModernPhysics , pp.440–448[NewYork:McGraw-Hill(1959)].
Triple Vector Product
Thesecondtripleproductofinterestis A×(B×C),whichisavector.Heretheparentheses
mustberetained,asmaybeseenfromaspecialcase (ˆx׈x)׈y=0,whileˆx×(ˆx׈y)=
ˆx׈z=−ˆy.
Example 1.5.1 ATRIPLE VECTOR PRODUCT
Forthevectors
A=ˆx+2ˆy−ˆz=(1,2,−1),B=ˆy+ˆz=(0,1,1),C=ˆx−ˆy=(0,1,1),
B×C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
011
1−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=ˆx+ˆy−ˆz,
and
A×(B×C)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
12−1
11−1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−ˆx−ˆz=−(ˆy+ˆz)−(ˆx−ˆy)
=−B−C. /squaresolid
ByrewritingtheresultinthelastlineofExample1.5.1asalinearcombinationof Band
C, we notice that, taking a geometric approach, the triple vector product is perpendicular
28 Chapter 1 Vector Analysis
FIGURE 1.17BandCar einthe xy-plane.
B×Cis perpendiculartothe xy-planeand
is shownherealongthe z-axis. Then
A×(B×C)is perpendiculartothe z-axis
andthereforeisbackinthe xy-plane.
toAand toB×C.The plane defined by BandCis perpendicular to B×C, and so the
tripleproductliesinthisplane(see Fig.1.17):
A×(B×C)=uB+vC. (1.54)
Taking the scalar product of Eq. (1.54) with Agives zero for the left-hand side, so
uA·B+vA·C=0. Hence u=wA·Candv=−wA·Bfor a suitable w. Substitut-
ingthesevaluesintoEq. (1.54) gives
A×(B×C)=wbracketleftbig
B(A·C)−C(A·B)bracketrightbig
; (1.55)
wewanttoshowthat
w=1
in Eq. (1.55), an important relation sometimes known as the BAC–CAB rule. Since
Eq. (1.55) is linear in A,B, andC,wis independent of these magnitudes. That is, we
only need to show that w=1 for unit vectors ˆA,ˆB,ˆC. Let us denote ˆB·ˆC=cosα,
ˆC·ˆA=cosβ,ˆA·ˆB=cosγ, andsquareEq.(1.55) toobtain
bracketleftbigˆA×(ˆB׈C)bracketrightbig2=ˆA2(ˆB׈C)2−bracketleftbigˆA·(ˆB׈C)bracketrightbig2
=1−cos2α−bracketleftbigˆA·(ˆB׈C)bracketrightbig2
=w2bracketleftbig
(ˆA·ˆC)2+(ˆA·ˆB)2−2(ˆA·ˆB)(ˆA·ˆC)(ˆB·ˆC)bracketrightbig
=w2parenleftbig
cos2β+cos2γ−2cosαcosβcosγparenrightbig
, (1.56)
1.5 Triple Scalar Product, Triple Vector Product 29
using(ˆA׈B)2=ˆA2ˆB2−(ˆA·ˆB)2repeatedly (see Eq. (1.43) for a proof). Consequently,
the(squared) volumespannedby ˆA,ˆB,ˆCthatoccursinEq. (1.56) canbewrittenas
bracketleftbigˆA·(ˆB׈C)bracketrightbig2=1−cos2α−w2parenleftbig
cos2β+cos2γ−2cosαcosβcosγparenrightbig
.
Herew2=1, since this volume is symmetric in α,β,γ. That is, w=±1 and is inde-
pendent ofˆA,ˆB,ˆC. Using again the special case ˆx×(ˆx׈y)=−ˆyin Eq. (1.55) finally
givesw=1.(AnalternatederivationusingtheLevi-Civitasymbol εijkofChapter2isthe
topicofExercise2.9.8.)
Itmightbenotedherethatjustasvectorsareindependentofthecoordinates,soavector
equation is independent of the particular coordinate system. The coordinate system only
determines the components. If the vector equation can be established in Cartesian coor-
dinates, it is established and valid in any of the coordinate systems to be introduced in
Chapter 2. Thus, Eq. (1.55) may be verified by a direct though not very elegant method of
expandingintoCartesiancomponents(see Exercise1.5.2).
Exercises
1.5.1 One vertex of a glass parallelepiped is at the origin (Fig. 1.18). The three adjacent
vertices are at (3,0,0),(0,0,2), and(0,3,1). All lengths are in centimeters. Calculate
the number of cubic centimeters of glass in the parallelepiped using the triple scalar
product.
1.5.2 Verifytheexpansionofthetriplevectorproduct
A×(B×C)=B(A·C)−C(A·B)
FIGURE 1.18Parallelepiped:triplescalarproduct.
30 Chapter 1 Vector Analysis
bydirectexpansioninCartesiancoordinates.
1.5.3 ShowthatthefirststepinEq. (1.43), whichis
(A×B)·(A×B)=A2B2−(A·B)2,
isconsistentwiththe BAC–CABrulefor atriplevectorproduct.
1.5.4 Youare giventhethreevectors A,B,andC,
A=ˆx+ˆy,
B=ˆy+ˆz,
C=ˆx−ˆz.
(a) Compute the triple scalar product, A·B×C. Noting that A=B+C, give a geo-
metricinterpretationofyourresultfor thetriplescalarproduct.
(b) Compute A×(B×C).
1.5.5 The orbital angular momentum Lof a particle is given by L=r×p=mr×v, where
pis the linear momentum. With linear and angular velocity related by v=ω×r,s h o w
that
L=mr2bracketleftbig
ω−ˆr(ˆr·ω)bracketrightbig
.
Hereˆris a unit vector in the r-direction. For r·ω=0 this reduces to L=Iω, with the
moment of inertia Igiven by mr2. In Section 3.5 this result is generalized to form an
inertiatensor.
1.5.6 Thekineticenergyofasingleparticleisgivenby T=1
2mv2.Forrotationalmotionthis
becomes1
2m(ω×r)2. Showthat
T=1
2mbracketleftbig
r2ω2−(r·ω)2bracketrightbig
.
Forr·ω=0 thisreducesto T=1
2Iω2, withthemomentof inertia Igivenby mr2.
1.5.7 Showthat13
a×(b×c)+b×(c×a)+c×(a×b)=0.
1.5.8 A vector Ais decomposed into a radial vector Arand a tangential vector At.I fˆris a
unitvectorintheradialdirection,showthat
(a)Ar=ˆr(A·ˆr)and
(b)At=−ˆr×(ˆr×A).
1.5.9 Prove that a necessary and sufficient condition for the three (nonvanishing) vectors A,
B,andCtobecoplanaristhevanishingofthetriplescalarproduct
A·B×C=0.
13This is Jacobi’s identity for vector products; for commutators it is important in the context of Lie algebras (see Eq. (4.16) in
Section 4.2).
1.5 Triple Scalar Product, Triple Vector Product 31
1.5.10 Threevectors A,B, andCare givenby
A=3ˆx−2ˆy+2ˆz,
B=6ˆx+4ˆy−2ˆz,
C=−3ˆx−2ˆy−4ˆz.
Computethevaluesof A·B×CandA×(B×C),C×(A×B)andB×(C×A).
1.5.11 VectorDisa linearcombinationofthreenoncoplanar(andnonorthogonal)vectors:
D=aA+bB+cC.
Showthatthecoefficientsaregivenbyaratiooftriplescalarproducts,
a=D·B×C
A·B×C,andso on.
1.5.12 Showthat
(A×B)·(C×D)=(A·C)(B·D)−(A·D)(B·C).
1.5.13 Showthat
(A×B)×(C×D)=(A·B×D)C−(A·B×C)D.
1.5.14 For aspherical trianglesuchaspicturedinFig.1.14showthat
sinA
sinBC=sinB
sinCA=sinC
sinAB.
Here sinAis the sine of the included angle at A, whileBCis the side opposite (in
radians).
1.5.15 Given
a′=b×c
a·b×c,b′=c×a
a·b×c,c′=a×b
a·b×c,
anda·b×c/negationslash=0,showthat
(a)x·y′=δxy,(x,y=a,b,c),
(b)a′·b′×c′=(a·b×c)−1,
(c)a=b′×c′
a′·b′×c′.
1.5.16 Ifx·y′=δxy,(x,y=a,b,c), prove that
a′=b×c
a·b×c.
(This istheconverseofProblem1.5.15.)
1.5.17 Showthatanyvector Vmaybeexpressedintermsofthereciprocalvectors a′,b′,c′(of
Problem1.5.15)by
V=(V·a)a′+(V·b)b′+(V·c)c′.
32 Chapter 1 Vector Analysis
1.5.18 An electric charge q1moving with velocity v1produces a magnetic induction Bgiven
by
B=µ0
4πq1v1׈r
r2(mksunits),
whereˆrpointsfrom q1tothepointatwhich Bis measured(BiotandSavartlaw).
(a) Show that the magnetic force on a second charge q2, velocity v2, is given by the
triplevectorproduct
F2=µ0
4πq1q2
r2v2×(v1׈r).
(b) Write out the corresponding magnetic force F1thatq2exerts on q1. Define your
unitradialvector.Howdo F1andF2compare?
(c) Calculate F1andF2for the case of q1andq2moving along parallel trajectories
sidebyside.
ANS.
(b)F1=−µ0
4πq1q2
r2v1×(v2׈r).
Ingeneral,thereisnosimplerelationbetween
F1andF2.Specifically,Newton’sthirdlaw, F1=−F2,
doesnothold.
(c)F1=µ0
4πq1q2
r2v2ˆr=−F2.
Mutualattraction.
1.6 G RADIENT ,∇
To provide a motivation for the vector nature of partial derivatives, we now introduce the
totalvariationof afunction F(x,y),
dF=∂F
∂xdx+∂F
∂ydy.
It consists of independent variations in the x- andy-directions. We write dFas a sum of
twoincrements,onepurelyinthe x- andtheotherinthe y-direction,
dF(x,y)≡F(x+dx,y+dy)−F(x,y)
=bracketleftbig
F(x+dx,y+dy)−F(x,y+dy)bracketrightbig
+bracketleftbig
F(x,y+dy)−F(x,y)bracketrightbig
=∂F
∂xdx+∂F
∂ydy,
byaddingandsubtracting F(x,y+dy).Themeanvaluetheorem(thatis,continuityof F)
tellsusthathere ∂F/∂x,∂F/∂yareevaluatedatsomepoint ξ,ηbetweenxandx+dx,y
1.6 Gradient, ∇ 33
andy+dy, respectively. As dx→0 anddy→0,ξ→xandη→y. This result general-
izestothreeandhigherdimensions.Forexample,for afunction ϕof threevariables,
dϕ(x,y,z)≡bracketleftbig
ϕ(x+dx,y+dy,z+dz)−ϕ(x,y+dy,z+dz)bracketrightbig
+bracketleftbig
ϕ(x,y+dy,z+dz)−ϕ(x,y,z+dz)bracketrightbig
+bracketleftbig
ϕ(x,y,z+dz)−ϕ(x,y,z)bracketrightbig
(1.57)
=∂ϕ
∂xdx+∂ϕ
∂ydy+∂ϕ
∂zdz.
Algebraically, dϕinthetotalvariationisascalarproductofthechangeinposition drand
thedirectional change of ϕ. And now we are ready to recognize the three-dimensional
partialderivativeasa vector,whichleadsustotheconceptofgradient.
Supposethat ϕ(x,y,z) isascalarpointfunction,thatis,afunctionwhosevaluedepends
onthevaluesofthecoordinates (x,y,z).Asascalar,itmusthavethesamevalueatagiven
fixedpointinspace,independentoftherotationof ourcoordinatesystem, or
ϕ′(x′
1,x′
2,x′
3)=ϕ(x1,x2,x3). (1.58)
Bydifferentiatingwithrespectto x′
iweobtain
∂ϕ′(x′
1,x′
2,x′
3)
∂x′
i=∂ϕ(x1,x2,x3)
∂x′
i=summationdisplay
j∂ϕ
∂xj∂xj
∂x′
i=summationdisplay
jaij∂ϕ
∂xj(1.59)
by the rules of partial differentiation and Eqs. (1.16a) and (1.16b). But comparison with
Eq. (1.17), the vector transformation law, now shows that we have constructed a vector
withcomponents ∂ϕ/∂xj.This vectorwelabelthegradientof ϕ.
Aconvenientsymbolismis
∇ϕ=ˆx∂ϕ
∂x+ˆy∂ϕ
∂y+ˆz∂ϕ
∂z(1.60)
or
∇=ˆx∂
∂x+ˆy∂
∂y+ˆz∂
∂z. (1.61)
∇ϕ(or delϕ) is our gradient of the scalar ϕ, whereas ∇(del) itself is a vector differential
operator (available to operate on or to differentiate a scalar ϕ). All the relationships for ∇
(del) can be derived from the hybrid nature of del in terms of both the partial derivatives
andits vectornature.
Thegradientofascalarisextremelyimportantinphysicsandengineeringinexpressing
therelationbetweenaforce fieldandapotentialfield,
forceF=−∇(potential V), (1.62)
which holds for both gravitational and electrostatic fields, among others. Note that the
minussigninEq.(1 .62)resultsinwaterflowingdownhillratherthanuphill!Ifaforcecan
be described, as in Eq. (1.62), by a single function V(r)everywhere, we call the scalar
functionVitspotential .Becausetheforceisthedirectionalderivativeofthepotential,we
canfindthepotential,ifitexists,byintegratingtheforcealongasuitablepath.Becausethe
34 Chapter 1 Vector Analysis
total variation dV=∇V·dr=−F·dris the work done against the force along the path
dr,we recognize the physical meaning of the potential (difference) as work and energy.
Moreover,inasumof pathincrementstheintermediatepointscancel,
bracketleftbig
V(r+dr1+dr2)−V(r+dr1)bracketrightbig
+bracketleftbig
V(r+dr1)−V(r)bracketrightbig
=V(r+dr2+dr1)−V(r),
sotheintegratedworkalongsomepathfromaninitialpoint ritoafinalpoint risgivenby
the potential difference V(r)−V(ri)at the endpoints of the path. Therefore, such forces
areespeciallysimpleandwellbehaved:Theyarecalled conservative .Whenthereislossof
energyduetofrictionalongthepathorsomeotherdissipation,theworkwilldependonthe
path,andsuchforces cannotbeconservative:Nopotentialexists. We discuss conservative
forcesinmoredetailinSection1.13.
Example 1.6.1 THEGRADIENT OF A POTENTIAL V(r)
Letuscalculatethegradientof V(r)=V(radicalbig
x2+y2+z2),so
∇V(r)=ˆx∂V(r)
∂x+ˆy∂V(r)
∂y+ˆz∂V(r)
∂z.
Now,V(r)dependson xthroughthedependenceof ronx.Therefore14
∂V(r)
∂x=dV(r)
dr·∂r
∂x.
Fromras afunctionof x,y,z,
∂r
∂x=∂(x2+y2+z2)1/2
∂x=x
(x2+y2+z2)1/2=x
r.
Therefore
∂V(r)
∂x=dV(r)
dr·x
r.
Permutingcoordinates (x→y,y→z,z→x)toobtainthe yandzderivatives,weget
∇V(r)=(ˆxx+ˆyy+ˆzz)1
rdV
dr
=r
rdV
dr=ˆrdV
dr.
Hereˆris a unit vector (r/r)in thepositiveradial direction. The gradient of a function of
ris a vector in the (positive or negative) radial direction. In Section 2.5, ˆris seen as one
ofthethreeorthonormalunitvectorsofsphericalpolarcoordinatesand ˆr∂/∂rastheradial
componentof ∇. /squaresolid
14This is aspecial caseof the chain rule of partial differentiation:
∂V(r,θ,ϕ)
∂x=∂V
∂r∂r
∂x+∂V
∂θ∂θ
∂x+∂V
∂ϕ∂ϕ
∂x,
where∂V/∂θ=∂V/∂ϕ=0,∂V/∂r→dV/dr.
1.6 Gradient, ∇ 35
A Geometrical Interpretation
Oneimmediateapplicationof ∇ϕistodotitintoanincrementof length
dr=ˆxdx+ˆydy+ˆzdz.
Thusweobtain
∇ϕ·dr=∂ϕ
∂xdx+∂ϕ
∂ydy+∂ϕ
∂zdz=dϕ,
thechangeinthescalarfunction ϕcorrespondingtoachangeinposition dr.Nowconsider
PandQtobetwopointsonasurface ϕ(x,y,z)=C,aconstant.Thesepointsarechosen
sothatQisadistance drfromP.Then,movingfrom PtoQ,thechangein ϕ(x,y,z)=C
isgivenby
dϕ=(∇ϕ)·dr=0 (1.63)
since we stay on the surface ϕ(x,y,z)=C. This shows that ∇ϕis perpendicular to dr.
Sincedrm a yh a v ea n yd i r e c t i o nf r o m Pas long as it stays in the surface of constant ϕ,
pointQbeingrestrictedtothesurfacebuthavingarbitrarydirection, ∇ϕisseenasnormal
tothesurface ϕ=constant(Fig. 1.19).
If we now permit drto take us from one surface ϕ=C1to an adjacent surface ϕ=C2
(Fig.1.20),
dϕ=C1−C2=/Delta1C=(∇ϕ)·dr. (1.64)
For a given dϕ,|dr|is a minimum when it is chosen parallel to ∇ϕ(cosθ=1);o r ,f o r
ag i v e n|dr|, the change in the scalar function ϕis maximized by choosing drparallel to
FIGURE 1.19Thelengthincrement drhastostayonthesurface ϕ=C.
36 Chapter 1 Vector Analysis
FIGURE 1.20Gradient.
∇ϕ.This identifies ∇ϕas a vector having the direction of the maximum space rate
of change of ϕ, an identification that will be useful in Chapter 2 when we consider non-
Cartesian coordinate systems. This identification of ∇ϕmay also be developed by using
thecalculusofvariationssubjecttoaconstraint,Exercise17.6.9.
Example 1.6.2 FORCE AS GRADIENT OF A POTENTIAL
Asaspecificexampleoftheforegoing,andasanextensionofExample1.6.1,weconsider
thesurfaces consistingofconcentricsphericalshells, Fig.1.21.Wehave
ϕ(x,y,z)=parenleftbig
x2+y2+z2parenrightbig1/2=r=C,
whereristheradius,equalto C,ourconstant. /Delta1C=/Delta1ϕ=/Delta1r,thedistancebetweentwo
shells.FromExample1.6.1
∇ϕ(r)=ˆrdϕ(r)
dr=ˆr.
Thegradientisintheradialdirectionandis normaltothesphericalsurface ϕ=C./squaresolid
Example 1.6.3 INTEGRATION BY PARTS OF GRADIENT
Letusprovetheformulaintegraltext
A(r)·∇f(r)d3r=−integraltext
f(r)∇·A(r)d3r,whereAorforboth
vanishatinfinitysothattheintegratedpartsvanish.Thisconditionissatisfiedif,forexam-
ple,Aistheelectromagneticvectorpotentialand fis abound-statewavefunction ψ(r).
1.6 Gradient, ∇ 37
FIGURE 1.21Gradientfor
ϕ(x,y,z)=(x2+y2+z2)1/2,spherical
shells:(x2
2+y2
2+z2
2)1/2=r2=C2,
(x2
1+y2
1+z2
1)1/2=r1=C1.
Writing the inner product in Cartesian coordinates, integrating each one-dimensional
integralbyparts, anddroppingtheintegratedterms,weobtain
integraldisplay
A(r)·∇f(r)d3r=integraldisplayintegraldisplaybracketleftbigg
Axf|∞
x=−∞−integraldisplay
f∂Ax
∂xdxbracketrightbigg
dydz+···
=−integraldisplayintegraldisplayintegraldisplay
f∂Ax
∂xdxdydz−integraldisplayintegraldisplayintegraldisplay
f∂Ay
∂ydydxdz−integraldisplayintegraldisplayintegraldisplay
f∂Az
∂zdzdxdy
=−integraldisplay
f(r)∇·A(r)d3r.
IfA=eikzˆedescribesanoutgoingphotoninthedirectionoftheconstantpolarizationunit
vectorˆeandf=ψ(r)is anexponentiallydecayingbound-statewavefunction,then
integraldisplay
eikzˆe·∇ψ(r)d3r=−ezintegraldisplay
ψ(r)deikz
dzd3r=−ikezintegraldisplay
ψ(r)eikzd3r,
becauseonlythe z-componentof thegradientcontributes. /squaresolid
Exercises
1.6.1 IfS(x,y,z)=(x2+y2+z2)−3/2,find
(a)∇Satthepoint (1,2,3);
(b) themagnitudeofthegradientof S,|∇S|at(1,2,3);and
(c) thedirectioncosinesof ∇Sat(1,2,3).
38 Chapter 1 Vector Analysis
1.6.2 (a) Findaunitvectorperpendiculartothesurface
x2+y2+z2=3
atthepoint (1,1,1). Lengthsare incentimeters.
(b) Derivetheequationof theplanetangenttothesurface at (1,1,1).
ANS.(a) (ˆx+ˆy+ˆz)/√
3,(b)x+y+z=3.
1.6.3 Given a vector r12=ˆx(x1−x2)+ˆy(y1−y2)+ˆz(z1−z2), show that ∇1r12(gradient
with respect to x1,y1, andz1of the magnitude r12) is a unit vector in the direction of
r12.
1.6.4 Ifavectorfunction Fdependsonbothspacecoordinates (x,y,z)andtime t,showthat
dF=(dr·∇)F+∂F
∂tdt.
1.6.5 Show that ∇(uv)=v∇u+u∇v, whereuandvare differentiable scalar functions of
x,y,andz.
(a) Show that a necessary and sufficient condition that u(x,y,z) andv(x,y,z) are
relatedbysomefunction f(u,v)=0 isthat(∇u)×(∇v)=0.
(b) Ifu=u(x,y)andv=v(x,y), show thatthe condition (∇u)×(∇v)=0 leadsto
thetwo-dimensionalJacobian
Jparenleftbiggu,v
x,yparenrightbigg
=vextendsinglevextendsinglevextendsinglevextendsingle∂u
∂x∂u
∂y
∂v
∂x∂v
∂yvextendsinglevextendsinglevextendsinglevextendsingle=0.
Thefunctions uandvare assumeddifferentiable.
1.7 D IVERGENCE ,∇
Differentiating a vector function is a simple extension of differentiating scalar quantities.
Supposer(t)describes the position of a satellite at some time t. Then, for differentiation
withrespecttotime,
dr(t)
dt=lim
/Delta1→0r(t+/Delta1t)−r(t)
/Delta1t=v,linearvelocity.
Graphically,weagainhavetheslopeofa curve,orbit, ortrajectory,asshowninFig.1.22.
If we resolve r(t)into its Cartesian components, dr/dtalways reduces directly to a
vectorsumofnotmorethanthree(forthree-dimensionalspace)scalarderivatives.Inother
coordinate systems (Chapter 2) the situation is more complicated, for the unit vectors are
no longer constant in direction. Differentiation with respect to the space coordinates is
handled in the same way as differentiation with respect to time, as seen in the following
paragraphs.
1.7 Divergence, ∇ 39
FIGURE 1.22Differentiationof avector.
In Section 1.6, ∇was defined as a vector operator. Now, paying attention to both its
vector and its differential properties, we let it operate on a vector. First, as a vector we dot
itintoasecondvectortoobtain
∇·V=∂Vx
∂x+∂Vy
∂y+∂Vz
∂z, (1.65a)
knownasthedivergenceof V.This isascalar, asdiscussedinSection1.3.
Example 1.7.1 DIVERGENCE OF COORDINATE VECTOR
Calculate ∇·r:
∇·r=parenleftbigg
ˆx∂
∂x+ˆy∂
∂y+ˆz∂
∂zparenrightbigg
·(ˆxx+ˆyy+ˆzz)
=∂x
∂x+∂y
∂y+∂z
∂z,
or∇·r=3. /squaresolid
Example 1.7.2 DIVERGENCE OF CENTRAL FORCE FIELD
GeneralizingExample1.7.1,
∇·parenleftbig
rf(r)parenrightbig
=∂
∂xbracketleftbig
xf(r)bracketrightbig
+∂
∂ybracketleftbig
yf(r)bracketrightbig
+∂
∂zbracketleftbig
zf(r)bracketrightbig
=3f(r)+x2
rdf
dr+y2
rdf
dr+z2
rdf
dr
=3f(r)+rdf
dr.
40 Chapter 1 Vector Analysis
ThemanipulationofthepartialderivativesleadingtothesecondequationinExample1.7.2
isdiscussedinExample1.6.1. Inparticular,if f(r)=rn−1,
∇·parenleftbig
rrn−1parenrightbig
=∇·ˆrrn
=3rn−1+(n−1)rn−1
=(n+2)rn−1. (1.65b)
Thisdivergencevanishesfor n=−2,exceptat r=0,animportantfactinSection1.14. /squaresolid
Example 1.7.3 INTEGRATION BY PARTS OF DIVERGENCE
Let us prove the formulaintegraltext
f(r)∇·A(r)d3r=−integraltext
A·∇fd3r,whereAorfor both
vanishatinfinity.
To show this, we proceed, as in Example 1.6.3, by integration by parts after writing
the inner product in Cartesian coordinates. Because the integrated terms are evaluated at
infinity,wheretheyvanish,weobtain
integraldisplay
f(r)∇·A(r)d3r=integraldisplay
fparenleftbigg∂Ax
∂xdxdydz+∂Ay
∂ydydxdz+∂Az
∂zdzdxdyparenrightbigg
=−integraldisplayparenleftbigg
Ax∂f
∂xdxdydz+Ay∂f
∂ydydxdz+Az∂f
∂zdzdxdyparenrightbigg
=−integraldisplay
A·∇fd3r./squaresolid
A Physical Interpretation
Todevelopafeelingforthephysicalsignificanceofthedivergence,consider ∇·(ρv)with
v(x,y,z),thevelocityofacompressiblefluid,and ρ(x,y,z) ,itsdensityatpoint (x,y,z).
Ifweconsiderasmallvolume dxdydz (Fig.1.23)at x=y=z=0,thefluidflowinginto
this volume per unit time (positive x-direction) through the face EFGHis (rate of flow
in)EFGH=ρvx|x=0=dydz. The components of the flow ρvyandρvztangential to this
face contribute nothing to the flow through this face. The rate of flow out (still positive
x-direction)throughface ABCDisρvx|x=dxdydz.Tocomparetheseflowsandtofindthe
netflowout,weexpandthislastresult,likethetotalvariationinSection1.6.15Thisyields
(rate offlowout )ABCD=ρvx|x=dxdydz
=bracketleftbigg
ρvx+∂
∂x(ρvx)dxbracketrightbigg
x=0dydz.
Herethederivativetermisafirstcorrectionterm,allowingforthepossibilityofnonuniform
densityorvelocityorboth.16Thezero-orderterm ρvx|x=0(correspondingtouniformflow)
15Herewehave the increment dxandweshow apartial derivative withrespectto xsinceρvxm ayal s od ep en do n yandz.
16Strictlyspeaking, ρvxisaveragedoverface EFGHandtheexpression ρvx+(∂/∂x)(ρv x)dxissimilarlyaveragedoverface
ABCD.Using anarbitrarily small differential volume, wefindthat the averages reduce tothe valuesemployed here.
1.7 Divergence, ∇ 41
FIGURE 1.23Differentialrectangularparallelepiped(in firstoctant).
cancelsout:
Netrateof flowout |x=∂
∂x(ρvx)dxdydz.
Equivalently,wecanarriveatthisresult by
lim
/Delta1x→0ρvx(/Delta1x,0,0)−ρvx(0,0,0)
/Delta1x≡∂[ρvx(x,y,z)]
∂xvextendsinglevextendsinglevextendsinglevextendsingle
0,0,0.
Now,the x-axisisnotentitledtoanypreferredtreatment.Theprecedingresultforthetwo
faces perpendicular to the x-axis must hold for the two faces perpendicular to the y-axis,
withxreplaced by yand the corresponding changes for yandz:y→z,z→x.T h i si s
a cyclic permutation of the coordinates. A further cyclic permutation yields the result for
the remaining two faces of our parallelepiped. Adding the net rate of flow out for all three
pairsofsurfaces of ourvolumeelement,wehave
netflowout
(per unittime)=bracketleftbigg∂
∂x(ρvx)+∂
∂y(ρvy)+∂
∂z(ρvz)bracketrightbigg
dxdydz
=∇·(ρv)dxdydz. (1.66)
Therefore the net flow of our compressible fluid out of the volume element dxdydz per
unit volume per unit time is ∇·(ρv). Hence the name divergence . A direct application is
inthecontinuityequation
∂ρ
∂t+∇·(ρv)=0, (1.67a)
which states that a net flow out of the volume results in a decreased density inside the
volume. Note that in Eq. (1.67a), ρis considered to be a possible function of time as well
as of space: ρ(x,y,z,t) . The divergence appears in a wide variety of physical problems,
42 Chapter 1 Vector Analysis
ranging from a probability current density in quantum mechanics to neutron leakage in a
nuclearreactor.
The combination ∇·(fV), in which fis a scalar function and Vis a vector function,
maybewritten
∇·(fV)=∂
∂x(fVx)+∂
∂y(fVy)+∂
∂z(fVz)
=∂f
∂xVx+f∂Vx
∂x+∂f
∂yVy+f∂Vy
∂y+∂f
∂zVz+f∂Vz
∂z
=(∇f)·V+f∇·V, (1.67b)
which is just what we would expect for the derivative of a product. Notice that ∇as a
differential operator differentiates both fandV; as a vector it is dotted into V(in each
term).
If wehavethespecialcaseofthedivergenceof avectorvanishing,
∇·B=0, (1.68)
the vector Bis said to be solenoidal , the term coming from the example in which Bis the
magnetic induction and Eq. (1.68) appears as one of Maxwell’s equations. When a vector
issolenoidal,itmaybewrittenas thecurlofanothervectorknownasthevectorpotential.
(InSection1.13weshallcalculatesuchavectorpotential.)
Exercises
1.7.1 Foraparticlemovinginacircularorbit r=ˆxrcosωt+ˆyrsinωt,
(a) evaluate r×˙r,with˙r=dr
dt=v.
(b) Showthat ¨r+ω2r=0 with¨r=dv
dt.
Theradius randtheangularvelocity ωareconstant.
ANS.(a)ˆzωr2.
1.7.2 VectorAsatisfies the vector transformation law, Eq. (1.15). Show directly that its time
derivative dA/dtalsosatisfiesEq. (1.15) andis thereforeavector.
1.7.3 Show,bydifferentiatingcomponents,that
(a)d
dt(A·B)=dA
dt·B+A·dB
dt,
(b)d
dt(A×B)=dA
dt×B+A×dB
dt,
justlikethederivativeoftheproductof twoalgebraicfunctions.
1.7.4 InChapter2itwillbeseenthattheunitvectorsinnon-Cartesiancoordinatesystemsare
usuallyfunctionsof the coordinatevariables, ei=ei(q1,q2,q3)but|ei|=1. Showthat
either∂ei/∂qj=0o r∂ei/∂qjisorthogonalto ei.
Hint.∂e2
i/∂qj=0.
1.8 Curl, ∇× 43
1.7.5 Prove ∇·(a×b)=b·(∇×a)−a·(∇×b).
Hint.Treatasatriplescalarproduct.
1.7.6 Theelectrostaticfieldofa pointcharge qis
E=q
4πε0·ˆr
r2.
Calculatethedivergenceof E.Whathappensattheorigin?
1.8 C URL ,∇×
Anotherpossibleoperationwiththevectoroperator ∇istocrossitintoavector.Weobtain
∇×V=ˆxparenleftbigg∂
∂yVz−∂
∂zVyparenrightbigg
+ˆyparenleftbigg∂
∂zVx−∂
∂xVzparenrightbigg
+ˆzparenleftbigg∂
∂xVy−∂
∂yVxparenrightbigg
=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
∂
∂x∂
∂y∂
∂z
VxVyVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (1.69)
whichiscalledthe curlofV.Inexpandingthisdeterminantwemustconsiderthederivative
nature of ∇. Specifically, V×∇is defined only as an operator, another vector differential
operator. It is certainly not equal, in general, to −∇×V.17In the case of Eq. (1.69) the
determinantmustbeexpanded fromthetopdown sothatwegetthederivativesasshown
inthemiddleportionofEq.(1.69).If ∇iscrossedintotheproductofascalarandavector,
wecanshow
∇×(fV)|x=bracketleftbigg∂
∂y(fVz)−∂
∂z(fVy)bracketrightbigg
=parenleftbigg
f∂Vz
∂y+∂f
∂yVz−f∂Vy
∂z−∂f
∂zVyparenrightbigg
=f∇×V|x+(∇f)×V|x. (1.70)
If we permute the coordinates x→y,y→z,z→xto pick up the y-component and
thenpermutethemasecondtimetopickupthe z-component,then
∇×(fV)=f∇×V+(∇f)×V, (1.71)
which is the vector product analog of Eq. (1.67b). Again, as a differential operator ∇
differentiatesboth fandV.Asavectoritiscrossedinto V(ineachterm).
17In this same spirit, if Ais a differential operator, it is not necessarily true that A×A=0. Specifically, for the quantum
mechanicalangular momentum operator L=−i(r×∇),wefin dth at L×L=iL.SeeSections 4.3 and 4.4 for more details.
44 Chapter 1 Vector Analysis
Example 1.8.1 VECTOR POTENTIAL OF A CONSTANT BFIELD
Fromelectrodynamicsweknowthat ∇·B=0,whichhasthegeneralsolution B=∇×A,
whereA(r)iscalledthevectorpotential(ofthemagneticinduction),because ∇·(∇×A)=
(∇×∇)·A≡0,asatriplescalarproductwithtwoidenticalvectors.Thislastidentitywill
not change if we add the gradient of some scalar function to the vector potential, which,
therefore,isnotunique.
In ourcase,wewanttoshowthatavectorpotentialis A=1
2(B×r).
Usingthe BAC–BACruleinconjunctionwithExample1.7.1, wefindthat
2∇×A=∇×(B×r)=(∇·r)B−(B·∇)r=3B−B=2B,
whereweindicatebytheorderingofthescalarproductofthesecondtermthatthegradient
stillactsonthecoordinatevector. /squaresolid
Example 1.8.2 CURL OF A CENTRAL FORCE FIELD
Calculate ∇×(rf(r)).
ByEq. (1.71),
∇×parenleftbig
rf(r)parenrightbig
=f(r)∇×r+bracketleftbig
∇f(r)bracketrightbig
×r. (1.72)
First,
∇×r=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
∂
∂x∂
∂y∂
∂z
xyzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (1.73)
Second,using ∇f(r)=ˆr(df/dr)(Example1.6.1), weobtain
∇×rf(r)=df
drˆr×r=0. (1.74)
Thisvectorproductvanishes,since r=ˆrrandˆr׈r=0. /squaresolid
To develop a better feeling for the physical significance of the curl, we consider the
circulationof fluidaroundadifferentialloopinthe xy-plane,Fig.1.24.
FIGURE 1.24Circulationaroundadifferentialloop.
1.8 Curl, ∇× 45
Although the circulation is technically given by a vector line integralintegraltext
V·dλ(Sec-
tion 1.10), we can set up the equivalent scalar integrals here. Let us take the circulation to
be
circulation 1234=integraldisplay
1Vx(x,y)dλ x+integraldisplay
2Vy(x,y)dλ y
+integraldisplay
3Vx(x,y)dλ x+integraldisplay
4Vy(x,y)dλ y. (1.75)
The numbers 1, 2, 3, and 4 refer to the numbered line segments in Fig. 1.24. In the first
integral,dλx=+dx; but in the third integral, dλx=−dxbecause the third line segment
istraversedinthenegative x-direction.Similarly, dλy=+dyforthesecondintegral, −dy
for the fourth. Next, the integrands are referred to the point (x0,y0)with a Taylor expan-
sion18taking into accountthe displacementof line segment3 from 1 and that of 2 from 4.
Forourdifferentiallinesegmentsthisleadsto
circulation 1234=Vx(x0,y0)dx+bracketleftbigg
Vy(x0,y0)+∂Vy
∂xdxbracketrightbigg
dy
+bracketleftbigg
Vx(x0,y0)+∂Vx
∂ydybracketrightbigg
(−dx)+Vy(x0,y0)(−dy)
=parenleftbigg∂Vy
∂x−∂Vx
∂yparenrightbigg
dxdy. (1.76)
Dividingby dxdy,weha v e
circulationperunitarea =∇×V|z. (1.77)
The circulation19about our differential area in the xy-plane is given by the z-component
of∇×V. In principle, the curl ∇×Vat(x0,y0)could be determined by inserting a
(differential) paddlewheel intothe movingfluid at point (x0,y0). The rotationof the little
paddle wheel would be a measure of the curl, and its axis would be along the direction of
∇×V,whichisperpendiculartotheplaneofcirculation.
Weshallusetheresult,Eq.(1.76),inSection1.12toderiveStokes’theorem.Whenever
thecurlof avector Vvanishes,
∇×V=0, (1.78)
Vislabeled irrotational .Themostimportantphysicalexamplesofirrotationalvectorsare
thegravitationalandelectrostaticforces. Ineachcase
V=Cˆr
r2=Cr
r3, (1.79)
whereCis a constant and ˆris the unit vector in the outward radial direction. For the
gravitationalcasewehave C=−Gm1m2,givenbyNewton’slawofuniversalgravitation.
IfC=q1q2/4πε0, we have Coulomb’s law of electrostatics (mks units). The force V
18Here,Vy(x0+dx,y0)=Vy(x0,y0)+(∂Vy
∂x)x0y0dx+···.The higher-order terms will drop out in the limit as dx→0.
Acorrection termfor the variation of Vywithyis canceledby the corresponding term in the fourth integral.
19In fluid dynamics ∇×Vis calledthe “vorticity.”
46 Chapter 1 Vector Analysis
given in Eq. (1.79) may be shown to be irrotational by direct expansion into Cartesian
components, as we did in Example 1.8.1. Another approach is developed in Chapter 2, in
whichweexpress ∇×,thecurl,intermsofsphericalpolarcoordinates.InSection1.13we
shall see that whenever a vector is irrotational, the vector may be written as the (negative)
gradient of a scalar potential. In Section 1.16 we shall prove that a vector field may be
resolved into an irrotational part and a solenoidal part (subject to conditions at infinity).
In terms of the electromagnetic field this corresponds to the resolution into an irrotational
electricfieldandasolenoidalmagneticfield.
For waves in an elastic medium, if the displacement uis irrotational, ∇×u=0, plane
waves (or spherical waves at large distances) become longitudinal. If uis solenoidal,
∇·u=0, then the waves become transverse. A seismic disturbance will produce a dis-
placement that may be resolved into a solenoidal part and an irrotational part (compare
Section 1.16). The irrotational part yields the longitudinal P(primary) earthquake waves.
Thesolenoidalpartgivesrise totheslowertransverse S(secondary)waves.
Using the gradient, divergence, and curl, and of course the BAC–CABrule, we may
construct or verify a large number of useful vector identities. For verification, complete
expansion into Cartesian components is always a possibility. Sometimes if we use insight
insteadofroutineshufflingofCartesiancomponents,theverificationprocesscanbeshort-
eneddrastically.
Rememberthat ∇is avectoroperator,ahybridcreaturesatisfyingtwosets ofrules:
1. vectorrules,and
2. partialdifferentiationrules—includingdifferentiationofa product.
Example 1.8.3 GRADIENT OF A DOTPRODUCT
Verifythat
∇(A·B)=(B·∇)A+(A·∇)B+B×(∇×A)+A×(∇×B). (1.80)
This particular example hinges on the recognition that ∇(A·B)is the type of term that
appearsinthe BAC–CABexpansionofatriplevectorproduct,Eq.(1.55). Forinstance,
A×(∇×B)=∇(A·B)−(A·∇)B,
with the ∇differentiating only B, notA. From the commutativity of factors in a scalar
productwemayinterchange AandBandwrite
B×(∇×A)=∇(A·B)−(B·∇)A,
now with ∇differentiating only A, notB. Adding these two equations, we obtain ∇dif-
ferentiating the product A·Band the identity, Eq. (1.80). This identity is used frequently
inelectromagnetictheory.Exercise1.8.13isasimpleillustration. /squaresolid
1.8 Curl, ∇× 47
Example 1.8.4 INTEGRATION BY PARTS OF CURL
Let us prove the formulaintegraltext
C(r)·(∇×A(r))d3r=integraltext
A(r)·(∇×C(r))d3r, whereAor
Cor bothvanishatinfinity.
To show this, we proceed, as in Examples 1.6.3 and 1.7.3, by integration by parts after
writing the inner product and the curl in Cartesian coordinates. Because the integrated
termsvanishatinfinityweobtain
integraldisplay
C(r)·parenleftbig
∇×A(r)parenrightbig
d3r
=integraldisplaybracketleftbigg
Czparenleftbigg∂Ay
∂x−∂Ax
∂yparenrightbigg
+Cxparenleftbigg∂Az
∂y−∂Ay
∂zparenrightbigg
+Cyparenleftbigg∂Ax
∂z−∂Az
∂xparenrightbiggbracketrightbigg
d3r
=integraldisplaybracketleftbigg
Axparenleftbigg∂Cz
∂y−∂Cy
∂zparenrightbigg
+Ayparenleftbigg∂Cx
∂z−∂Cz
∂xparenrightbigg
+Azparenleftbigg∂Cy
∂x−∂Cx
∂yparenrightbiggbracketrightbigg
d3r
=integraldisplay
A(r)·parenleftbig
∇×C(r)parenrightbig
d3r,
justrearrangingappropriatelytheterms afterintegrationbyparts. /squaresolid
Exercises
1.8.1 Show,byrotatingthecoordinates,thatthecomponentsofthecurlofavectortransform
asavector.
Hint.Thedirectioncosineidentitiesof Eq.(1.46) areavailableas needed.
1.8.2 Showthat u×vissolenoidalif uandvareeachirrotational.
1.8.3 IfAisirrotational,showthat A×ris solenoidal.
1.8.4 A rigid body is rotating with constant angular velocity ω. Show that the linear velocity
vissolenoidal.
1.8.5 Ifavectorfunction f(x,y,z)isnotirrotationalbuttheproductof fandascalarfunction
g(x,y,z) isirrotational,showthatthen
f·∇×f=0.
1.8.6 If(a)V=ˆxVx(x,y)+ˆyVy(x,y)and(b) ∇×V/negationslash=0,provethat ∇×Visperpendicular
toV.
1.8.7 Classically, orbital angular momentum is given by L=r×p, wherepis the linear
momentum. To go from classical mechanics to quantum mechanics, replace pby the
operator−i∇(Section 15.6). Show that the quantum mechanical angular momentum
48 Chapter 1 Vector Analysis
operatorhas Cartesiancomponents(inunitsof ¯h)
Lx=−iparenleftbigg
y∂
∂z−z∂
∂yparenrightbigg
,
Ly=−iparenleftbigg
z∂
∂x−x∂
∂zparenrightbigg
,
Lz=−iparenleftbigg
x∂
∂y−y∂
∂xparenrightbigg
.
1.8.8 Using the angular momentum operators previously given, show that they satisfy com-
mutationrelationsoftheform
[Lx,Ly]≡LxLy−LyLx=iLz
andhence
L×L=iL.
These commutation relations will be taken later as the defining relations of an angular
momentumoperator—Exercise3.2.15andthefollowingoneandChapter4.
1.8.9 With the commutator bracket notation [Lx,Ly]=LxLy−LyLx, the angular momen-
tumvector Lsatisfies[Lx,Ly]=iLz,et c. ,orL×L=iL.
If two other vectors aandbcommute with each other and with L, that is,[a,b]=
[a,L]=[b,L]=0,showthat
[a·L,b·L]=i(a×b)·L.
1.8.10 ForA=ˆxAx(x,y,z)andB=ˆxBx(x,y,z)evaluateeachterminthevectoridentity
∇(A·B)=(B·∇)A+(A·∇)B+B×(∇×A)+A×(∇×B)
andverifythattheidentityis satisfied.
1.8.11 Verifythevectoridentity
∇×(A×B)=(B·∇)A−(A·∇)B−B(∇·A)+A(∇·B).
1.8.12 Asanalternativetothevectoridentityof Example1.8.3showthat
∇(A·B)=(A×∇)×B+(B×∇)×A+A(∇·B)+B(∇·A).
1.8.13 Verifytheidentity
A×(∇×A)=1
2∇parenleftbig
A2parenrightbig
−(A·∇)A.
1.8.14 IfAandBareconstantvectors,showthat
∇(A·B×r)=A×B.
1.9 Successive Applications of ∇ 49
1.8.15 A distribution of electric currents creates a constant magnetic moment m=const. The
forceonminanexternalmagneticinduction Bis givenby
F=∇×(B×m).
Showthat
F=(m·∇)B.
Note.Assumingnotimedependenceofthefields,Maxwell’sequationsyield ∇×B=0.
Also,∇·B=0.
1.8.16 An electric dipole of moment pis located at the origin. The dipole creates an electric
potentialat rgivenby
ψ(r)=p·r
4πε0r3.
Findtheelectricfield, E=−∇ψatr.
1.8.17 The vector potential Aof a magnetic dipole, dipole moment m, is given by A(r)=
(µ0/4π)(m×r/r3). Showthatthemagneticinduction B=∇×Ais givenby
B=µ0
4π3ˆr(ˆr·m)−m
r3.
Note. The limiting process leading to point dipoles is discussed in Section 12.1 for
electricdipoles,inSection12.5formagneticdipoles.
1.8.18 Thevelocityof atwo-dimensionalflowofliquidis givenby
V=ˆxu(x,y)−ˆyv(x,y).
If theliquidisincompressibleandtheflowisirrotational,showthat
∂u
∂x=∂v
∂yand∂u
∂y=−∂v
∂x.
ThesearetheCauchy–Riemannconditionsof Section6.2.
1.8.19 The evaluation in this section of the four integrals for the circulation omitted Taylor
series terms such as ∂Vx/∂x,∂Vy/∂yand all second derivatives. Show that ∂Vx/∂x,
∂Vy/∂ycancel out when the four integrals are added and that the second derivative
termsdropoutinthelimitas dx→0,dy→0.
Hint.Calculatethecirculationperunitareaandthentakethelimit dx→0,dy→0.
1.9 S UCCESSIVE APPLICATIONS OF ∇
We have now defined gradient, divergence, and curl to obtain vector, scalar, and vector
quantities,respectively.Letting ∇operateoneachof thesequantities,weobtain
(a)∇·∇ϕ (b)∇×∇ϕ (c)∇∇·V
(d)∇·∇×V(e)∇×(∇×V)
50 Chapter 1 Vector Analysis
allfiveexpressionsinvolvingsecondderivativesandallfiveappearinginthesecond-order
differentialequationsof mathematicalphysics,particularlyinelectromagnetictheory.
Thefirstexpression, ∇·∇ϕ,thedivergenceofthegradient,isnamedtheLaplacianof ϕ.
We have
∇·∇ϕ=parenleftbigg
ˆx∂
∂x+ˆy∂
∂y+ˆz∂
∂zparenrightbigg
·parenleftbigg
ˆx∂ϕ
∂x+ˆy∂ϕ
∂y+ˆz∂ϕ
∂zparenrightbigg
=∂2ϕ
∂x2+∂2ϕ
∂y2+∂2ϕ
∂z2. (1.81a)
Whenϕistheelectrostaticpotential,wehave
∇·∇ϕ=0 (1.81b)
at points where the charge density vanishes, which is Laplace’s equation of electrostatics.
Oftenthecombination ∇·∇is written ∇2,or/Delta1intheEuropeanliterature.
Example 1.9.1 LAPLACIAN OF A POTENTIAL
Calculate ∇·∇V(r).
ReferringtoExamples1.6.1and1.7.2,
∇·∇V(r)=∇·ˆrdV
dr=2
rdV
dr+d2V
dr2,
replacing f(r)inExample1.7.2 by 1 /r·dV/dr.IfV(r)=rn, thisreducesto
∇·∇rn=n(n+1)rn−2.
Thisvanishesfor n=0[V(r)=constant]andfor n=−1;thatis, V(r)=1/risasolution
of Laplace’s equation, ∇2V(r)=0. This is for r/negationslash=0. Atr=0, a Dirac delta function is
involved(seeEq. (1.169)andSection9.7). /squaresolid
Expression(b) maybewritten
∇×∇ϕ=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz
∂
∂x∂
∂y∂
∂z
∂ϕ
∂x∂ϕ
∂y∂ϕ
∂zvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
Byexpandingthedeterminant,weobtain
∇×∇ϕ=ˆxparenleftbigg∂2ϕ
∂y∂z−∂2ϕ
∂z∂yparenrightbigg
+ˆyparenleftbigg∂2ϕ
∂z∂x−∂2ϕ
∂x∂zparenrightbigg
+ˆzparenleftbigg∂2ϕ
∂x∂y−∂2ϕ
∂y∂xparenrightbigg
=0, (1.82)
assuming that the order of partial differentiation may be interchanged. This is true as long
as these second partial derivatives of ϕare continuous functions. Then, from Eq. (1.82),
thecurlofagradientisidenticallyzero.Allgradients,therefore,areirrotational.Notethat
1.9 Successive Applications of ∇ 51
the zero in Eq. (1.82) comes as a mathematical identity, independent of any physics. The
zeroinEq.(1.81b) isaconsequenceof physics.
Expression(d) isa triplescalarproductthatmaybewritten
∇·∇×V=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂
∂x∂
∂y∂
∂z
∂
∂x∂
∂y∂
∂z
VxVyVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (1.83)
Again,assumingcontinuitysothattheorderofdifferentiationis immaterial,weobtain
∇·∇×V=0. (1.84)
The divergence of a curl vanishes or all curls are solenoidal. In Section 1.16 we shall see
thatvectorsmayberesolvedintosolenoidalandirrotationalpartsbyHelmholtz’stheorem.
Thetworemainingexpressionssatisfyarelation
∇×(∇×V)=∇∇·V−∇·∇V, (1.85)
valid in Cartesian coordinates (but not in curved coordinates). This follows immediately
from Eq. (1.55), the BAC–CABrule, which we rewrite so that Cappears at the extreme
rightof eachterm.The term ∇·∇Vwas notincludedin ourlist, but itmaybe definedby
Eq.(1.85).
Example 1.9.2 ELECTROMAGNETIC WAVEEQUATION
One important application of this vector relation (Eq. (1.85)) is in the derivation of the
electromagneticwaveequation.In vacuumMaxwell’sequationsbecome
∇·B=0, (1.86a)
∇·E=0, (1.86b)
∇×B=ε0µ0∂E
∂t, (1.86c)
∇×E=−∂B
∂t. (1.86d)
HereEis the electric field, Bis the magnetic induction, ε0is the electric permittivity,
andµ0is the magnetic permeability (SI units), so ε0µ0=1/c2,cbeing the velocity of
light. The relation has important consequences. Because ε0,µ0can be measured in any
frame,thevelocityof lightis thesameinanyframe.
Suppose we eliminate Bfrom Eqs. (1.86c) and (1.86d). We may do this by taking the
curlofbothsidesofEq.(1.86d)andthetimederivativeofbothsidesofEq.(1.86c).Since
thespaceandtimederivativescommute,
∂
∂t∇×B=∇×∂B
∂t,
andweobtain
∇×(∇×E)=−ε0µ0∂2E
∂t2.
52 Chapter 1 Vector Analysis
ApplicationofEqs. (1.85) and(1.86b)yields
∇·∇E=ε0µ0∂2E
∂t2, (1.87)
the electromagnetic vector wave equation. Again, if Eis expressed in Cartesian coor-
dinates, Eq. (1.87) separates into three scalar wave equations, each involving the scalar
Laplacian.
When external electric charge and current densities are kept as driving terms in
Maxwell’s equations, similar wave equations are valid for the electric potential and the
vector potential. To show this, we solve Eq. (1.86a) by writing B=∇×Aas a curl of the
vector potential. This expression is substituted into Faraday’s induction law in differential
form,Eq.(1.86d),toyield ∇×(E+∂A
∂t)=0.Thevanishingcurlimpliesthat E+∂A
∂tisa
gradient and, therefore, can be written as −∇ϕ,whereϕ(r,t)is defined as the (nonstatic)
electricpotential.Theseresultsfor the BandEfields,
B=∇×A,E=−∇ϕ−∂A
∂t, (1.88)
solvethehomogeneousMaxwell’sequations.
We nowshowthattheinhomogeneousMaxwell’sequations,
Gauss’ law: ∇·E=ρ/ε0,Oersted’slaw: ∇×B−1
c2∂E
∂t=µ0J(1.89)
indifferentialformleadtowaveequationsforthepotentials ϕandA,providedthat ∇·Ais
determinedbytheconstraint1
c2∂ϕ
∂t+∇·A=0.Thischoiceoffixingthedivergenceofthe
vectorpotential,calledthe Lorentzgauge ,servestouncouplethedifferentialequationsof
bothpotentials.This gaugeconstraintis notarestriction;ithas nophysicaleffect.
SubstitutingourelectricfieldsolutionintoGauss’ lawyields
ρ
ε0=∇·E=−∇2ϕ−∂
∂t∇·A=−∇2ϕ+1
c2∂2ϕ
∂t2, (1.90)
the wave equation for the electric potential. In the last step we have used the Lorentz
gaugetoreplacethedivergenceofthevectorpotentialbythetimederivativeoftheelectric
potentialandthusdecouple ϕfromA.
Finally, we substitute B=∇×Ainto Oersted’s law and use Eq. (1.85), which expands
∇2in terms of a longitudinal (the gradient term) and a transverse component (the curl
term).This yields
µ0J+1
c2∂E
∂t=∇×(∇×A)=∇(∇·A)−∇2A=µ0J−1
c2parenleftbigg
∇∂ϕ
∂t+∂2A
∂t2parenrightbigg
,
wherewehaveusedtheelectricfieldsolution(Eq.(1.88))inthelaststep.Nowweseethat
theLorentzgaugeconditioneliminatesthegradientterms, sothewaveequation
1
c2∂2A
∂t2−∇2A=µ0J (1.91)
1.9 Successive Applications of ∇ 53
forthevectorpotentialremains.
Finally, looking back at Oersted’s law, taking the divergence of Eq. (1.89), dropping
∇·(∇×B)=0,andsubstitutingGauss’lawfor ∇·E=ρ/ǫ0,wefindµ0∇·J=−1
ǫ0c2∂ρ
∂t,
whereǫ0µ0=1/c2,that is, the continuityequationfor the current density. This step justi-
fiestheinclusionofMaxwell’sdisplacementcurrentinthegeneralizationofOersted’slaw
tononstationarysituations. /squaresolid
Exercises
1.9.1 VerifyEq. (1.85),
∇×(∇×V)=∇∇·V−∇·∇V,
bydirectexpansioninCartesiancoordinates.
1.9.2 Showthattheidentity
∇×(∇×V)=∇∇·V−∇·∇V
followsfromthe BAC–CABruleforatriplevectorproduct.Justifyanyalterationofthe
orderof factorsinthe BACandCABterms.
1.9.3 Provethat ∇×(ϕ∇ϕ)=0.
1.9.4 You are given that the curl of Fequals the curl of G. Show that FandGmay differ by
(a) aconstantand(b) agradientof ascalarfunction.
1.9.5 TheNavier–Stokesequationofhydrodynamicscontainsanonlinearterm (v·∇)v.Show
thatthecurlof thistermmaybewrittenas −∇×[v×(∇×v)].
1.9.6 FromtheNavier–Stokesequationforthesteadyflowofanincompressibleviscousfluid
wehavetheterm
∇×bracketleftbig
v×(∇×v)bracketrightbig
,
wherevisthefluidvelocity.Showthatthistermvanishesforthespecialcase
v=ˆxv(y,z).
1.9.7 Provethat (∇u)×(∇v)issolenoidal,where uandvaredifferentiablescalarfunctions.
1.9.8 ϕis a scalar satisfying Laplace’s equation, ∇2ϕ=0. Show that ∇ϕisbothsolenoidal
andirrotational.
1.9.9 Withψascalar(wave)function,showthat
(r×∇)·(r×∇)ψ=r2∇2ψ−r2∂2ψ
∂r2−2r∂ψ
∂r.
(This canactuallybeshownmoreeasilyinsphericalpolarcoordinates,Section2.5.)
54 Chapter 1 Vector Analysis
1.9.10 Ina(nonrotating)isolatedmasssuchasastar, theconditionfor equilibriumis
∇P+ρ∇ϕ=0.
HerePis the total pressure, ρis the density, and ϕis the gravitational potential. Show
that at any given point the normals to the surfaces of constant pressure and constant
gravitationalpotentialareparallel.
1.9.11 InthePaulitheoryoftheelectron,oneencounterstheexpression
(p−eA)×(p−eA)ψ,
whereψis a scalar (wave) function. Ais the magnetic vector potential related to the
magnetic induction BbyB=∇×A.Given that p=−i∇, show that this expression
reduces to ieBψ. Show that this leads to the orbital g-factorgL=1 upon writing the
magnetic moment as µ=gLLin units of Bohr magnetons and L=−ir×∇.See also
Exercise1.13.7.
1.9.12 Showthatanysolutionof theequation
∇×(∇×A)−k2A=0
automaticallysatisfies thevectorHelmholtzequation
∇2A+k2A=0
andthesolenoidalcondition
∇·A=0.
Hint.Let∇·operateonthefirst equation.
1.9.13 Thetheoryof heatconductionleadstoanequation
∇2/Psi1=k|∇/Phi1|2,
where/Phi1isapotentialsatisfyingLaplace’sequation: ∇2/Phi1=0.Showthatasolutionof
thisequationis
/Psi1=1
2k/Phi12.
1.10 V ECTOR INTEGRATION
Thenextstepafterdifferentiatingvectorsistointegratethem.Letusstartwithlineintegrals
andthenproceedtosurfaceandvolumeintegrals.Ineachcasethemethodofattackwillbe
toreducethevectorintegraltoscalarintegralswithwhichthereaderis assumedfamiliar.
1.10 Vector Integration 55
LineIntegrals
Usinganincrementoflength dr=ˆxdx+ˆydy+ˆzdz,wemayencounterthelineintegrals
integraldisplay
Cϕdr, (1.92a)
integraldisplay
CV·dr, (1.92b)
integraldisplay
CV×dr, (1.92c)
ineachofwhichtheintegralisoversomecontour Cthatmaybeopen(withstartingpoint
and ending point separated) or closed (forming a loop). Because of its physical interpreta-
tionthatfollows,thesecondform, Eq. (1.92b)isbyfar themostimportantofthethree.
Withϕ, ascalar,thefirst integralreducesimmediatelyto
integraldisplay
Cϕdr=ˆxintegraldisplay
Cϕ(x,y,z)dx+ˆyintegraldisplay
Cϕ(x,y,z)dy+ˆzintegraldisplay
Cϕ(x,y,z)dz. (1.93)
Thisseparationhasemployedtherelation
integraldisplay
ˆxϕdx=ˆxintegraldisplay
ϕdx, (1.94)
which is permissible because the Cartesian unit vectors ˆx,ˆy, andˆzare constant in both
magnitudeanddirection.Perhapsthisrelationisobvioushere,butitwillnotbetrueinthe
non-CartesiansystemsencounteredinChapter2.
The three integrals on the right side of Eq. (1.93) are ordinary scalar integrals and, to
avoid complications, we assume that they are Riemann integrals. Note, however, that the
integral with respect to xcannot be evaluated unless yandzare known in terms of x
and similarly for the integrals with respect to yandz. This simply means that the path
of integration Cmust be specified. Unless the integrand has special properties so that
the integral depends only on the value of the end points, the value will depend on the
particular choice of contour C. For instance, if we choose the very special case ϕ=1,
Eq. (1.92a) is just the vector distance from the start of contour Cto the endpoint, in this
caseindependentofthechoiceofpathconnectingfixedendpoints.With dr=ˆxdx+ˆydy+
ˆzdz, the second and third forms also reduce to scalar integrals and, like Eq. (1.92a), are
dependent, in general, on the choice of path. The form (Eq. (1.92b)) is exactly the same
as that encountered when we calculate the work done by a force that varies along the
path,
W=integraldisplay
F·dr=integraldisplay
Fx(x,y,z)dx+integraldisplay
Fy(x,y,z)dy+integraldisplay
Fz(x,y,z)dz. (1.95a)
Inthisexpression Fis theforce exertedonaparticle.
56 Chapter 1 Vector Analysis
FIGURE 1.25Apathof integration.
Example 1.10.1 PATH-DEPENDENT WORK
The force exerted on a body is F=−ˆxy+ˆyx. The problem is to calculate the work done
goingfromtheorigintothepoint (1,1):
W=integraldisplay1,1
0,0F·dr=integraldisplay1,1
0,0(−ydx+xdy). (1.95b)
Separatingthetwointegrals,weobtain
W=−integraldisplay1
0ydx+integraldisplay1
0xdy. (1.95c)
The first integral cannot be evaluated until we specify the values of yasxranges from 0
to 1. Likewise, the second integral requires xas a function of y. Consider first the path
showninFig.1.25. Then
W=−integraldisplay1
00dx+integraldisplay1
01dy=1, (1.95d)
sincey=0alongthefirstsegmentofthepathand x=1alongthesecond.Ifweselectthe
path[x=0,0/lessorequalslanty/lessorequalslant1]and[0/lessorequalslantx/lessorequalslant1,y=1], then Eq. (1.95c) gives W=−1. For this
forcetheworkdonedependsonthechoiceofpath. /squaresolid
SurfaceIntegrals
Surfaceintegralsappearinthesameforms aslineintegrals,theelementofareaalsobeing
avector,dσ.20Oftenthisareaelementiswritten ndA,inwhich nisaunit(normal)vector
to indicate the positive direction.21There are two conventions for choosing the positive
direction. First, if the surface is a closed surface, we agree to take the outward normal
as positive. Second, if the surface is an open surface, the positive normal depends on the
direction in which the perimeter of the open surface is traversed. If the right-hand fingers
20Recallthat in Section1.4the area(of aparallelogram) is representedby across-product vector.
21Although nalwayshas unit length, its direction may wellbe afunction ofposition.
1.10 Vector Integration 57
FIGURE 1.26Right-handrulefor
thepositivenormal.
areplacedinthedirectionoftravelaroundtheperimeter,thepositivenormalisindicatedby
thethumbofthe righthand.As an illustration,acirclein the xy-plane(Fig. 1.26) mapped
out from xtoyto−xto−yand back to xwill have its positive normal parallel to the
positivez-axis (for theright-handedcoordinatesystem).
Analogous to the line integrals, Eqs. (1.92a) to (1.92c), surface integrals may appear in
theforms
integraldisplay
ϕdσ,integraldisplay
V·dσ,integraldisplay
V×dσ.
Again,thedotproductisbyfarthemostcommonlyencounteredform.Thesurfaceintegralintegraltext
V·dσmaybeinterpretedasafloworfluxthroughthegivensurface.Thisisreallywhat
we did in Section 1.7 to obtain the significance of the term divergence. This identification
reappears in Section 1.11 as Gauss’ theorem. Note that both physically and from the dot
product the tangential components of the velocity contribute nothing to the flow through
thesurface.
VolumeIntegrals
Volume integrals are somewhat simpler, for the volume element dτis a scalar quantity.22
We have
integraldisplay
VVdτ=ˆxintegraldisplay
VVxdτ+ˆyintegraldisplay
VVydτ+ˆzintegraldisplay
VVzdτ, (1.96)
againreducingthevectorintegraltoavectorsumof scalarintegrals.
22Frequently the symbols d3randd3xareused to denote avolume elementin coordinate ( xyzorx1x2x3)space.
58 Chapter 1 Vector Analysis
FIGURE 1.27Differentialrectangularparallelepiped(originatcenter).
IntegralDefinitionsof Gradient,Divergence,andCurl
One interesting and significant application of our surface and volume integrals is their use
indevelopingalternatedefinitionsof ourdifferentialrelations.We find
∇ϕ=limintegraltext
dτ→0integraltext
ϕdσintegraltext
dτ, (1.97)
∇·V=limintegraltext
dτ→0integraltext
V·dσintegraltext
dτ, (1.98)
∇×V=limintegraltext
dτ→0integraltext
dσ×Vintegraltext
dτ. (1.99)
Inthesethreeequationsintegraltext
dτisthevolumeofasmallregionofspaceand dσisthevector
area element of this volume. The identification of Eq. (1.98) as the divergence of Vwas
carried out in Section 1.7. Here we show that Eq. (1.97) is consistent with our earlier
definition of ∇ϕ(Eq. (1.60)). For simplicity we choose dτto be the differential volume
dxdydz (Fig. 1.27). This time we place the origin at the geometric center of our volume
element.Theareaintegralleadstosixintegrals,oneforeachofthesixfaces.Remembering
thatdσis outward, dσ·ˆx=−|dσ|for surface EFHG, and+|dσ|for surface ABDC,w e
have
integraldisplay
ϕdσ=−ˆxintegraldisplay
EFHGparenleftbigg
ϕ−∂ϕ
∂xdx
2parenrightbigg
dydz+ˆxintegraldisplay
ABDCparenleftbigg
ϕ+∂ϕ
∂xdx
2parenrightbigg
dydz
−ˆyintegraldisplay
AEGCparenleftbigg
ϕ−∂ϕ
∂ydy
2parenrightbigg
dxdz+ˆyintegraldisplay
BFHDparenleftbigg
ϕ+∂ϕ
∂ydy
2parenrightbigg
dxdz
−ˆzintegraldisplay
ABFEparenleftbigg
ϕ−∂ϕ
∂zdz
2parenrightbigg
dxdy+ˆzintegraldisplay
CDHGparenleftbigg
ϕ+∂ϕ
∂zdz
2parenrightbigg
dxdy.
1.10 Vector Integration 59
Using the total variations, we evaluate each integrand at the origin with a correction in-
cluded to correct for the displacement ( ±dx/2, etc.) of the center of the face from the
origin. Having chosen the total volume to be of differential size (integraltext
dτ=dxdydz) ,w e
droptheintegralsigns ontherightandobtain
integraldisplay
ϕdσ=parenleftbigg
ˆx∂ϕ
∂x+ˆy∂ϕ
∂y+ˆz∂ϕ
∂zparenrightbigg
dxdydz. (1.100)
Dividingby
integraldisplay
dτ=dxdydz,
weverifyEq. (1.97).
This verification has been oversimplified in ignoring other correction terms beyond the
first derivatives. These additional terms, which are introduced in Section 5.6 when the
Taylorexpansionisdeveloped,vanishinthelimit
integraldisplay
dτ→0(dx→0,dy→0,dz→0).
This,ofcourse,isthereasonforspecifyinginEqs.(1.97),(1.98),and(1.99)thatthislimit
be taken. Verification of Eq. (1.99) follows these same lines exactly, using a differential
volumedxdydz.
Exercises
1.10.1 Theforcefieldactingona two-dimensionallinearoscillatormaybedescribedby
F=−ˆxkx−ˆyky.
Comparetheworkdonemovingagainstthisforcefieldwhengoingfrom (1,1)to(4,4)
bythefollowingstraight-linepaths:
(a)(1,1)→(4,1)→(4,4)
(b)(1,1)→(1,4)→(4,4)
(c)(1,1)→(4,4)alongx=y.
This meansevaluating
−integraldisplay(4,4)
(1,1)F·dr
alongeachpath.
1.10.2 Findtheworkdonegoingaroundaunitcircleinthe xy-plane:
(a) counterclockwisefrom 0 to π,
(b) clockwisefrom 0 to −π, doingwork againstaforcefieldgivenby
F=−ˆxy
x2+y2+ˆyx
x2+y2.
Notethatthework donedependsonthepath.
60 Chapter 1 Vector Analysis
1.10.3 Calculatetheworkyoudoingoingfrompoint (1,1)topoint(3,3).Theforce youexert
isgivenby
F=ˆx(x−y)+ˆy(x+y).
Specifyclearlythepathyouchoose.Notethatthisforcefieldisnonconservative.
1.10.4 Evaluatecontintegraltext
r·dr.
Note.Thesymbolcontintegraltext
meansthatthepathofintegrationisaclosedloop.
1.10.5 Evaluate
1
3integraldisplay
sr·dσ
over the unit cube defined by the point (0,0,0)and the unit intercepts on the positive
x-,y-, andz-axes. Note that (a) r·dσis zero for three of the surfaces and (b) each of
thethreeremainingsurfaces contributesthesameamounttotheintegral.
1.10.6 Show,byexpansionofthesurface integral,that
limintegraltext
dτ→0integraltext
sdσ×Vintegraltext
dτ=∇×V.
Hint.Choosethevolumeintegraltext
dτtobeadifferentialvolume dxdydz.
1.11 G AUSS ’THEOREM
Herewederiveausefulrelationbetweenasurfaceintegralofavectorandthevolumeinte-
gralofthedivergenceofthatvector.Letusassumethatthevector Vanditsfirstderivatives
are continuous over the simply connected region (that does not have any holes, such as a
donut)of interest.ThenGauss’theoremstatesthat
integraldisplayintegraldisplay
/circlecopyrt
∂VV·dσ=integraldisplayintegraldisplayintegraldisplay
V∇·Vdτ. (1.101a)
In words, the surface integral of a vector over a closed surface equals the volume integral
ofthedivergenceof thatvectorintegratedoverthevolumeenclosedbythesurface.
Imagine that volume Vis subdivided into an arbitrarily large number of tiny (differen-
tial)parallelepipeds.Foreachparallelepiped
summationdisplay
six surfacesV·dσ=∇·Vdτ (1.101b)
from the analysis of Section 1.7, Eq. (1.66), with ρvreplaced by V. The summation is
overthesixfacesoftheparallelepiped.Summingoverallparallelepipeds,wefindthatthe
V·dσtermscancel(pairwise)forall interiorfaces;onlythecontributionsofthe exterior
surfaces survive (Fig. 1.28). Analogousto the definitionof a Riemannintegral as the limit
1.11 Gauss’ Theorem 61
FIGURE 1.28Exact
cancellationof dσ’s on
interiorsurfaces. No
cancellationonthe
exteriorsurface.
of a sum, we take the limit as the number of parallelepipeds approaches infinity (→∞)
andthedimensionsofeachapproachzero (→0):
summationtext
exterior surfacesV·dσ=summationtext
volumes∇·Vdτ
integraltext
SV·dσ=integraltext
V∇·Vdτ.
TheresultisEq. (1.101a), Gauss’theorem.
From a physical point of view Eq. (1.66) has established ∇·Vas the net outflow of
fluidper unitvolume.The volumeintegralthengivesthetotalnetoutflow.Butthesurface
integralintegraltext
V·dσisjustanotherwayofexpressingthissamequantity,whichistheequality,
Gauss’theorem.
Green’s Theorem
AfrequentlyusefulcorollaryofGauss’theoremisarelationknownasGreen’stheorem.If
uandvare twoscalarfunctions,wehavetheidentities
∇·(u∇v)=u∇·∇v+(∇u)·(∇v), (1.102)
∇·(v∇u)=v∇·∇u+(∇v)·(∇u). (1.103)
Subtracting Eq. (1.103) from Eq. (1.102), integrating over a volume ( u,v, and their
derivatives,assumedcontinuous),andapplyingEq. (1.101a)(Gauss’ theorem),weobtain
integraldisplayintegraldisplayintegraldisplay
V(u∇·∇v−v∇·∇u)dτ=integraldisplayintegraldisplay
/circlecopyrt
∂V(u∇v−v∇u)·dσ. (1.104)
62 Chapter 1 Vector Analysis
This is Green’s theorem. We use it for developing Green’s functions in Chapter 9. An
alternateformofGreen’stheorem,derivedfromEq. (1.102)alone,is
integraldisplayintegraldisplay
/circlecopyrt
∂Vu∇v·dσ=integraldisplayintegraldisplayintegraldisplay
Vu∇·∇vdτ+integraldisplayintegraldisplayintegraldisplay
V∇u·∇vdτ. (1.105)
Thisis theform ofGreen’stheoremusedinSection1.16.
Alternate Forms of Gauss’ Theorem
AlthoughEq.(1.101a)involvingthedivergenceisbyfarthemostimportantformofGauss’
theorem,volumeintegralsinvolvingthegradientandthecurlmayalsoappear.Suppose
V(x,y,z)=V(x,y,z) a, (1.106)
in which ais a vector with constant magnitude and constant but arbitrary direction. (You
pickthedirection,butonceyouhavechosenit, holditfixed.)Equation(1.101a)becomes
a·integraldisplayintegraldisplay
/circlecopyrt
∂VVdσ=integraldisplayintegraldisplayintegraldisplay
V∇·aVdτ=a·integraldisplayintegraldisplayintegraldisplay
V∇Vdτ (1.107)
byEq.(1.67b). Thismayberewritten
a·bracketleftbiggintegraldisplayintegraldisplay
/circlecopyrt
∂VVdσ−integraldisplayintegraldisplayintegraldisplay
V∇Vdτbracketrightbigg
=0. (1.108)
Since|a|/negationslash=0 and its direction is arbitrary, meaning that the cosine of the included angle
cannotalwaysvanish,thetermsinbracketsmustbezero.23Theresultis
integraldisplayintegraldisplay
/circlecopyrt
∂VVdσ=integraldisplayintegraldisplayintegraldisplay
V∇Vdτ . (1.109)
Inasimilarmanner,using V=a×Pinwhichais aconstantvector,wemayshow
integraldisplayintegraldisplay
/circlecopyrt
∂Vdσ×P=integraldisplayintegraldisplayintegraldisplay
V∇×Pdτ. (1.110)
TheselasttwoformsofGauss’theoremareusedinthevectorformofKirchoffdiffraction
theory. They may also be used to verify Eqs. (1.97) and (1.99). Gauss’ theorem may also
beextendedtotensors(seeSection2.11).
Exercises
1.11.1 UsingGauss’theorem,provethat
integraldisplayintegraldisplay
/circlecopyrt
Sdσ=0
ifS=∂Visaclosedsurface.
23This exploitation of the arbitrary nature of a part of a problem is a valuable and widely used technique. The arbitrary vector
isusedagaininSections1.12and1.13.OtherexamplesappearinSection1.14(integrandsequated)andinSection2.8,quotient
rule.
1.11 Gauss’ Theorem 63
1.11.2 Showthat
1
3integraldisplayintegraldisplay
/circlecopyrt
Sr·dσ=V,
whereVisthevolumeenclosedbytheclosedsurface S=∂V.
Note.This isageneralizationofExercise1.10.5.
1.11.3 IfB=∇×A,showthat
integraldisplayintegraldisplay
/circlecopyrt
SB·dσ=0
for anyclosedsurface S.
1.11.4 Over some volume Vletψbe a solution of Laplace’s equation (with the derivatives
appearing there continuous). Prove that the integral over any closed surface in Vof the
normalderivativeof ψ(∂ψ/∂n,o r∇ψ·n)willbezero.
1.11.5 In analogy to the integral definition of gradient, divergence, and curl of Section 1.10,
showthat
∇2ϕ=limintegraltext
dτ→0integraltext
∇ϕ·dσintegraltext
dτ.
1.11.6 The electric displacement vector Dsatisfies the Maxwell equation ∇·D=ρ, whereρ
is the charge density (per unit volume). At the boundary between two media there is a
surface chargedensity σ(per unitarea). Showthataboundaryconditionfor Dis
(D2−D1)·n=σ.
nis aunitvectornormaltothesurfaceandoutofmedium1.
Hint.Considerathinpillboxas showninFig.1.29.
1.11.7 From Eq. (1.67b), with Vthe electric field Eandfthe electrostatic potential ϕ,show
that,for integrationoverallspace,
integraldisplay
ρϕdτ=ε0integraldisplay
E2dτ.
Thiscorrespondstoathree-dimensionalintegrationbyparts.
Hint.E=−∇ϕ,∇·E=ρ/ε0.You may assume that ϕvanishes at large rat least as
fast asr−1.
FIGURE 1.29Pillbox.
64 Chapter 1 Vector Analysis
1.11.8 A particular steady-state electric current distribution is localized in space. Choosing a
boundingsurfacefar enoughoutsothatthecurrentdensity Jiszeroeverywhereonthe
surface, showthat
integraldisplayintegraldisplayintegraldisplay
Jdτ=0.
Hint. Take one component of Jat a time. With ∇·J=0, show that Ji=∇·(xiJ)and
applyGauss’theorem.
1.11.9 The creation of a localized system of steady electric currents (current density J) and
magneticfieldsmaybeshowntorequireanamountof work
W=1
2integraldisplayintegraldisplayintegraldisplay
H·Bdτ.
Transformthisinto
W=1
2integraldisplayintegraldisplayintegraldisplay
J·Adτ.
HereAisthemagneticvectorpotential: ∇×A=B.
Hint.InMaxwell’sequationstakethedisplacementcurrentterm ∂D/∂t=0.Ifthefields
and currents are localized, a bounding surface may be taken far enough out so that the
integralsof thefieldsandcurrentsoverthesurfaceyieldzero.
1.11.10 Provethegeneralizationof Green’stheorem:
integraldisplayintegraldisplayintegraldisplay
V(vLu−uLv)dτ=integraldisplayintegraldisplay
/circlecopyrt
∂Vp(v∇u−u∇v)·dσ.
HereListheself-adjointoperator(Section10.1),
L=∇·bracketleftbig
p(r)∇bracketrightbig
+q(r)
andp,q,u,andvarefunctionsofposition, pandqhavingcontinuousfirstderivatives
anduandvhavingcontinuoussecondderivatives.
Note.This generalizedGreen’stheoremappearsinSection9.7.
1.12 S TOKES ’THEOREM
Gauss’ theorem relates the volume integral of a derivative of a function to an integral of
the function over the closed surface bounding the volume. Here we consider an analogous
relation between the surface integral of a derivative of a function and the line integral of
thefunction,thepathofintegrationbeingtheperimeterboundingthesurface.
Let us take the surface and subdivide it into a network of arbitrarily small rectangles.
In Section 1.8 we showed that the circulation about such a differential rectangle (in the
xy-plane)is ∇×V|zdxdy.FromEq. (1.76)appliedto onedifferentialrectangle,
summationdisplay
four sidesV·dλ=∇×V·dσ. (1.111)
1.12 Stokes’ Theorem 65
FIGURE 1.30Exactcancellationon
interiorpaths.Nocancellationonthe
exteriorpath.
Wesumoverallthelittlerectangles,asinthedefinitionofaRiemannintegral.Thesurface
contributions (right-hand side of Eq. (1.111)) are added together. The line integrals (left-
hand side of Eq. (1.111)) of all interiorline segments cancel identically. Only the line
integral around the perimeter survives (Fig. 1.30). Taking the usual limit as the number of
rectanglesapproachesinfinitywhile dx→0,dy→0,wehave
summationtext
exterior line
segmentsV·dλ=summationtext
rectangles∇×V·dσ(1.112)
contintegraldisplay
V·dλ=integraldisplay
S∇×V·dσ.
This is Stokes’ theorem. The surface integral on the right is over the surface bounded
by the perimeter or contour, for the line integral on the left. The direction of the vector
representingtheareaisoutofthepaperplanetowardthereaderifthedirectionoftraversal
around the contour for the line integral is in the positive mathematical sense, as shown in
Fig.1.30.
This demonstration of Stokes’ theorem is limited by the fact that we used a Maclaurin
expansion of V(x,y,z)in establishing Eq. (1.76) in Section 1.8. Actually we need only
demand that the curl of V(x,y,z)exist and that it be integrable over the surface. A proof
of the Cauchy integral theorem analogous to the developmentof Stokes’ theorem here but
usingtheseless restrictiveconditionsappearsinSection6.3.
Stokes’theoremobviouslyappliestoanopensurface.Itispossibletoconsideraclosed
surfaceasalimitingcaseofanopensurface,withtheopening(andthereforetheperimeter)
shrinkingtozero.This isthepointof Exercise1.12.7.
66 Chapter 1 Vector Analysis
Alternate Forms of Stokes’ Theorem
As with Gauss’ theorem, other relations between surface and line integrals are possible.
Wefind
integraldisplay
Sdσ×∇ϕ=contintegraldisplay
∂Sϕdλ (1.113)
and
integraldisplay
S(dσ×∇)×P=contintegraldisplay
∂Sdλ×P. (1.114)
Equation (1.113) may readily be verified by the substitution V=aϕ, in which ais a vec-
tor of constant magnitude and of constant direction, as in Section 1.11. Substituting into
Stokes’theorem,Eq. (1.112),
integraldisplay
S(∇×aϕ)·dσ=−integraldisplay
Sa×∇ϕ·dσ
=−a·integraldisplay
S∇ϕ×dσ. (1.115)
Forthelineintegral,
contintegraldisplay
∂Saϕ·dλ=a·contintegraldisplay
∂Sϕdλ, (1.116)
andweobtain
a·parenleftbiggcontintegraldisplay
∂Sϕdλ+integraldisplay
S∇ϕ×dσparenrightbigg
=0. (1.117)
Since the choice of direction of ais arbitrary, the expression in parentheses must vanish,
thusverifyingEq.(1.113).Equation(1.114)maybederivedsimilarlybyusing V=a×P,
inwhichaisagainaconstantvector.
We can use Stokes’ theorem to derive Oersted’s and Faraday’s laws from two of
Maxwell’s equations, and vice versa, thus recognizing that the former are an integrated
formofthelatter.
Example 1.12.1 OERSTED ’SA N D FARADAY ’SLAWS
Consider the magnetic field generated by a long wire that carries a stationary current I.
StartingfromMaxwell’sdifferentiallaw ∇×H=J,Eq.(1.89)(withMaxwell’sdisplace-
ment current ∂D/∂t=0 for a stationary current case by Ohm’s law), we integrate over a
closedarea SperpendiculartoandsurroundingthewireandapplyStokes’theoremtoget
I=integraldisplay
SJ·dσ=integraldisplay
S(∇×H)·dσ=contintegraldisplay
∂SH·dr,
whichisOersted’slaw.Herethelineintegralisalong ∂S,theclosedcurvesurroundingthe
cross-sectionalarea S.
1.12 Stokes’ Theorem 67
Similarly,wecanintegrateMaxwell’sequationfor ∇×E,Eq.(1.86d),toyieldFaraday’s
induction law. Imagine moving a closed loop (∂S)of wire (of area S) across a magnetic
inductionfield B. WeintegrateMaxwell’sequationanduseStokes’theorem,yielding
integraldisplay
∂SE·dr=integraldisplay
S(∇×E)·dσ=−d
dtintegraldisplay
SB·dσ=−d/Phi1
dt,
which is Faraday’s law. The line integral on the left-hand side represents the voltage in-
duced in the wire loop, while the right-hand side is the change with time of the magnetic
flux/Phi1throughthemovingsurface Softhewire. /squaresolid
Both Stokes’ and Gauss’ theorems are of tremendous importance in a wide variety of
problems involving vector calculus. Some idea of their power and versatility may be ob-
tainedfromtheexercisesofSections1.11and1.12andthedevelopmentofpotentialtheory
inSections1.13and1.14.
Exercises
1.12.1 Given a vector t=−ˆxy+ˆyx, show, with the help of Stokes’ theorem, that the integral
aroundacontinuousclosedcurveinthe xy-plane
1
2contintegraldisplay
t·dλ=1
2contintegraldisplay
(xdy−ydx)=A,
theareaenclosedbythecurve.
1.12.2 Thecalculationofthemagneticmomentof acurrentloopleadstothelineintegral
contintegraldisplay
r×dr.
(a) Integrate around the perimeter of a current loop (in the xy-plane) and show that
thescalarmagnitudeofthislineintegralistwicetheareaoftheenclosedsurface.
(b) The perimeter of an ellipse is described by r=ˆxacosθ+ˆybsinθ. From part (a)
showthattheareaoftheellipseis πab.
1.12.3 Evaluatecontintegraltext
r×drbyusingthealternateform ofStokes’theoremgivenbyEq. (1.114):
integraldisplay
S(dσ×∇)×P=contintegraldisplay
dλ×P.
Takethelooptobeentirelyinthe xy-plane.
1.12.4 Insteadystatethemagneticfield HsatisfiestheMaxwellequation ∇×H=J,whereJ
isthecurrentdensity(persquaremeter).Attheboundarybetweentwomediathereisa
surface currentdensity K.Showthataboundaryconditionon His
n×(H2−H1)=K.
nis aunitvectornormaltothesurfaceandoutofmedium1.
Hint.Consideranarrowloopperpendiculartotheinterfaceas showninFig.1.31.
68 Chapter 1 Vector Analysis
FIGURE 1.31
Integrationpath
attheboundary
oftwomedia.
1.12.5 FromMaxwell’sequations, ∇×H=J,withJherethecurrentdensityand E=0.Show
fromthisthatcontintegraldisplay
H·dr=I,
whereIisthenetelectriccurrentenclosedbytheloopintegral.Thesearethedifferential
andintegralforms ofAmpère’slawofmagnetism.
1.12.6 Amagneticinduction Bisgeneratedbyelectriccurrentinaringofradius R.Showthat
themagnitude ofthevectorpotential A(B=∇×A)attheringcanbe
|A|=ϕ
2πR,
whereϕisthetotalmagneticfluxpassingthroughthering.
Note.Ais tangential to the ring and may be changed by adding the gradient of a scalar
function.
1.12.7 Provethatintegraldisplay
S∇×V·dσ=0,
ifSisaclosedsurface.
1.12.8 Evaluatecontintegraltext
r·dr(Exercise1.10.4)byStokes’ theorem.
1.12.9 Provethatcontintegraldisplay
u∇v·dλ=−contintegraldisplay
v∇u·dλ.
1.12.10 Provethatcontintegraldisplay
u∇v·dλ=integraldisplay
S(∇u)×(∇v)·dσ.
1.13 P OTENTIAL THEORY
Scalar Potential
If a force over a given simply connected region of space S(which means that it has no
holes)canbeexpressedasthenegativegradientof ascalarfunction ϕ,
F=−∇ϕ, (1.118)
1.13 Potential Theory 69
wecallϕascalarpotentialthatdescribestheforcebyonefunctioninsteadofthree.Ascalar
potentialisonlydetermineduptoanadditiveconstant,whichcanbeusedtoadjustitsvalue
at infinity (usually zero) or at some other point. The force Fappearing as the negative
gradient of a single-valued scalar potential is labeled a conservative force. We want to
know when a scalar potential function exists. To answer this question we establish two
otherrelationsasequivalenttoEq.(1.118). Theseare
∇×F=0 (1.119)
and
contintegraldisplay
F·dr=0, (1.120)
for every closed path in our simply connected region S. We proceed to show that each of
thesethreeequationsimpliestheothertwo.Letus startwith
F=−∇ϕ. (1.121)
Then
∇×F=−∇×∇ϕ=0 (1.122)
byEq.(1.82) orEq. (1.118)impliesEq.(1.119). Turningtothelineintegral,wehave
contintegraldisplay
F·dr=−contintegraldisplay
∇ϕ·dr=−contintegraldisplay
dϕ, (1.123)
using Eq. (1.118). Now, dϕintegrates to give ϕ. Since we have specified a closed loop,
the end points coincide and we get zero for every closed path in our region Sfor which
Eq. (1.118) holds. It is important to note the restriction here that the potential be single-
valuedandthatEq.(1.118)holdfor allpointsinS.Thisproblemmayariseinusingascalar
magnetic potential, a perfectly valid procedure as long as no net current is encircled. As
soonaswechooseapathinspacethatencirclesanetcurrent,thescalarmagneticpotential
ceasestobesingle-valuedandouranalysisnolongerapplies.
Continuing this demonstration of equivalence, let us assume that Eq. (1.120) holds. Ifcontintegraltext
F·dr=0 for all paths in S, we see that the value of the integral joining two distinct
pointsAandBisindependentof thepath(Fig.1.32). Ourpremiseisthat
contintegraldisplay
ACBDAF·dr=0. (1.124)
Therefore
integraldisplay
ACBF·dr=−integraldisplay
BDAF·dr=integraldisplay
ADBF·dr, (1.125)
reversing the sign by reversing the direction of integration. Physically, this means that
the work done in going from AtoBis independent of the path and that the work done in
goingaroundaclosedpathiszero.Thisisthereasonforlabelingsuchaforceconservative:
Energyisconserved.
70 Chapter 1 Vector Analysis
FIGURE 1.32Possiblepathsfor doingwork.
With the result shown in Eq. (1.125), we have the work done dependent only on the
endpoints AandB. Thatis,
workdonebyforce =integraldisplayB
AF·dr=ϕ(A)−ϕ(B). (1.126)
Equation (1.126) defines a scalar potential (strictly speaking, the difference in potential
between points AandB) and provides a means of calculating the potential. If point B
is taken as a variable, say, (x,y,z), then differentiation with respect to x,y, andzwill
recoverEq.(1.118).
Thechoiceofsignontheright-handsideisarbitrary.Thechoicehereismadetoachieve
agreement with Eq. (1.118) and to ensure that water will run downhill rather than uphill.
Forpoints AandBseparatedbyalength dr,Eq. (1.126)becomes
F·dr=−dϕ=−∇ϕ·dr. (1.127)
Thismayberewritten
(F+∇ϕ)·dr=0, (1.128)
andsince dris arbitrary,Eq. (1.118)mustfollow.If
contintegraldisplay
F·dr=0, (1.129)
wemayobtainEq. (1.119)byusingStokes’theorem(Eq. (1.112)):
contintegraldisplay
F·dr=integraldisplay
∇×F·dσ. (1.130)
If we take the path of integration to be the perimeter of an arbitrary differential area dσ,
theintegrandinthesurfaceintegralmustvanish.HenceEq. (1.120)impliesEq. (1.119).
Finally, if ∇×F=0, we need only reverse our statement of Stokes’ theorem
(Eq. (1.130)) to derive Eq. (1.120). Then, by Eqs. (1.126) to (1.128), the initial statement
1.13 Potential Theory 71
FIGURE 1.33Equivalentformulationsof aconservativeforce.
FIGURE 1.34Potentialenergyversusdistance(gravitational,
centrifugal,andsimpleharmonicoscillator).
F=−∇ϕis derived. The triple equivalence is demonstrated (Fig. 1.33). To summarize,
asingle-valuedscalarpotentialfunction ϕexistsifandonlyif Fisirrotationalorthework
donearoundeveryclosedloopiszero.Thegravitationalandelectrostaticforcefieldsgiven
byEq.(1.79)areirrotationalandthereforeareconservative.Gravitationalandelectrostatic
scalar potentials exist. Now, by calculating the work done (Eq. (1.126)), we proceed to
determinethreepotentials(Fig.1.34).
72 Chapter 1 Vector Analysis
Example 1.13.1 GRAVITATIONAL POTENTIAL
Findthescalarpotentialfor thegravitationalforceonaunitmass m1,
FG=−Gm1m2ˆr
r2=−kˆr
r2, (1.131)
radiallyinward. ByintegratingEq.(1.118) frominfinityintoposition r,weobtain
ϕG(r)−ϕG(∞)=−integraldisplayr
∞FG·dr=+integraldisplay∞
rFG·dr. (1.132)
By use of FG=−Fapplied, a comparison with Eq. (1.95a) shows that the potential is the
work done in bringing the unit mass in from infinity. (We can define only potential dif-
ference. Here we arbitrarily assign infinity to be a zero of potential.) The integral on the
right-hand side of Eq. (1.132) is negative, meaning that ϕG(r)is negative. Since FGis
radial,weobtaina contributionto ϕonlywhen dris radial,or
ϕG(r)=−integraldisplay∞
rkdr
r2=−k
r=−Gm1m2
r.
Thefinalnegativesignisaconsequenceof theattractiveforceof gravity. /squaresolid
Example 1.13.2 CENTRIFUGAL POTENTIAL
Calculate the scalar potential for the centrifugal force per unit mass, FC=ω2rˆr, radially
outward. Physically, you might feel this on a large horizontal spinning disk at an amuse-
ment park. Proceeding as in Example 1.13.1 but integrating from the origin outward and
takingϕC(0)=0,wehave
ϕC(r)=−integraldisplayr
0FC·dr=−ω2r2
2.
If we reverse signs, taking FSHO=−kr, we obtain ϕSHO=1
2kr2, the simple harmonic
oscillatorpotential.
The gravitational, centrifugal, and simple harmonic oscillator potentials are shown in
Fig. 1.34. Clearly, the simple harmonic oscillator yields stability and describes a restoring
force.Thecentrifugalpotentialdescribesanunstablesituation. /squaresolid
Thermodynamics — Exact Differentials
In thermodynamics, which is sometimes called a search for exact differentials, we en-
counterequationsoftheform
df=P(x,y)dx+Q(x,y)dy. (1.133a)
The usual problem is to determine whetherintegraltext
(P(x,y)dx+Q(x,y)dy) depends only on
the endpoints, that is, whether dfis indeed an exact differential. The necessary and suffi-
cientconditionis that
df=∂f
∂xdx+∂f
∂ydy (1.133b)
1.13 Potential Theory 73
orthat
P(x,y)=∂f/∂x,
Q(x,y)=∂f/∂y.(1.133c)
Equations(1.133c)dependonsatisfyingtherelation
∂P(x,y)
∂y=∂Q(x,y)
∂x. (1.133d)
This, however, is exactly analogous to Eq. (1.119), the requirement that Fbe irrotational.
Indeed,the z-componentof Eq.(1.119) yields
∂Fx
∂y=∂Fy
∂x, (1.133e)
with
Fx=∂f
∂x,F y=∂f
∂y.
VectorPotential
In some branches of physics, especially electrodynamics, it is convenient to introduce a
vectorpotential Asuchthata(force) field Bisgivenby
B=∇×A. (1.134)
Clearly, if Eq. (1.134) holds, ∇·B=0 by Eq. (1.84) and Bis solenoidal. Here we want
to develop a converse, to show that when Bis solenoidal a vector potential Aexists. We
demonstrate the existence of Aby actually calculating it. Suppose B=ˆxb1+ˆyb2+ˆzb3
andourunknown A=ˆxa1+ˆya2+ˆza3. ByEq. (1.134),
∂a3
∂y−∂a2
∂z=b1, (1.135a)
∂a1
∂z−∂a3
∂x=b2, (1.135b)
∂a2
∂x−∂a1
∂y=b3. (1.135c)
Let us assume that the coordinates have been chosen so that Ais parallel to the yz-plane;
thatis,a1=0.24Then
b2=−∂a3
∂x
b3=∂a2
∂x.(1.136)
24Clearly, this can be done at any one point. It is not at all obvious that this assumption will hold at all points; that is, Awill be
two-dimensional. The justification for the assumption is that it works; Eq.(1.141) satisfies Eq.(1.134).
74 Chapter 1 Vector Analysis
Integrating,weobtain
a2=integraldisplayx
x0b3dx+f2(y,z),
a3=−integraldisplayx
x0b2dx+f3(y,z),(1.137)
wheref2andf3are arbitrary functions of yandzbutnotfunctions of x. These two
equationscanbecheckedbydifferentiatingandrecoveringEq.(1.136).Equation(1.135a)
becomes25
∂a3
∂y−∂a2
∂z=−integraldisplayx
x0parenleftbigg∂b2
∂y+∂b3
∂zparenrightbigg
dx+∂f3
∂y−∂f2
∂z
=integraldisplayx
x0∂b1
∂xdx+∂f3
∂y−∂f2
∂z, (1.138)
using∇·B=0.Integratingwithrespectto x,weobtain
∂a3
∂y−∂a2
∂z=b1(x,y,z)−b1(x0,y,z)+∂f3
∂y−∂f2
∂z. (1.139)
Rememberingthat f3andf2are arbitraryfunctionsof yandz, wechoose
f2=0,
f3=integraldisplayy
y0b1(x0,y,z)dy,(1.140)
so that the right-hand side of Eq. (1.139) reduces to b1(x,y,z), in agreement with
Eq.(1.135a). With f2andf3givenbyEq.(1.140), wecanconstruct A:
A=ˆyintegraldisplayx
x0b3(x,y,z)dx+ˆzbracketleftbiggintegraldisplayy
y0b1(x0,y,z)dy−integraldisplayx
x0b2(x,y,z)dxbracketrightbigg
.(1.141)
However,thisisnotquitecomplete.Wemayaddanyconstantsince Bisaderivativeof A.
What is much more important, we may add any gradient of a scalar function ∇ϕwithout
affecting Bat all. Finally, the functions f2andf3are not unique. Other choices could
have been made. Instead of setting a1=0 to get Eq. (1.136) any cyclic permutation of
1,2,3,x,y,z,x 0,y0,z0wouldalsowork.
Example 1.13.3 AM AGNETIC VECTOR POTENTIAL FOR A CONSTANT MAGNETIC FIELD
To illustrate the construction of a magnetic vector potential, we take the special but still
importantcaseof aconstantmagneticinduction
B=ˆzBz, (1.142)
25Leibniz’formula in Exercise9.6.13 is useful here.
1.13 Potential Theory 75
inwhich Bzisaconstant.Equations(1.135atoc)become
∂a3
∂y−∂a2
∂z=0,
∂a1
∂z−∂a3
∂x=0, (1.143)
∂a2
∂x−∂a1
∂y=Bz.
If weassumethat a1=0,as before,thenbyEq. (1.141)
A=ˆyintegraldisplayx
Bzdx=ˆyxBz, (1.144)
setting a constant of integration equal to zero. It can readily be seen that this Asatisfies
Eq.(1.134).
To show that the choice a1=0 was not sacred or at least not required, let us try setting
a3=0.FromEq.(1.143)
∂a2
∂z=0, (1.145a)
∂a1
∂z=0, (1.145b)
∂a2
∂x−∂a1
∂y=Bz. (1.145c)
We seea1anda2areindependentof z,or
a1=a1(x,y), a 2=a2(x,y). (1.146)
Equation(1.145c)issatisfiedifwetake
a2=pintegraldisplayx
Bzdx=pxBz (1.147)
and
a1=(p−1)integraldisplayy
Bzdy=(p−1)yBz, (1.148)
withpanyconstant.Then
A=ˆx(p−1)yBz+ˆypxBz. (1.149)
Again, Eqs. (1.134), (1.142), and (1.149) are seen to be consistent. Comparison of Eqs.
(1.144) and (1.149) shows immediately that Ais not unique. The difference between
Eqs. (1.144) and (1.149) and the appearance of the parameter pin Eq. (1.149) may be
accountedfor byrewritingEq. (1.149)as
A=−1
2(ˆxy−ˆyx)Bz+parenleftbigg
p−1
2parenrightbigg
(ˆxy+ˆyx)Bz
=−1
2(ˆxy−ˆyx)Bz+parenleftbigg
p−1
2parenrightbigg
Bz∇ϕ (1.150)
76 Chapter 1 Vector Analysis
with
ϕ=xy. (1.151)
/squaresolid
The first term in Acorrespondstotheusualform
A=1
2(B×r) (1.152)
forB,a constant.
Adding a gradient of a scalar function, /Lambda1say, to the vector potential Adoes not affect
B,byEq.(1.82);thisisknownasagaugetransformation(seeExercises1.13.9and4.6.4):
A→A′=A+∇/Lambda1. (1.153)
Suppose now that the wave function ψ0solves the Schrödinger equation of quantum
mechanicswithoutmagneticinductionfield B,
braceleftbigg1
2m(−i¯h∇)2+V−Ebracerightbigg
ψ0=0, (1.154)
describingaparticlewithmass mandcharge e.WhenBisswitchedon,thewaveequation
becomes
braceleftbigg1
2m(−i¯h∇−eA)2+V−Ebracerightbigg
ψ=0. (1.155)
Its solution ψpicksupaphasefactor thatdependsonthecoordinatesingeneral,
ψ(r)=expbracketleftbiggie
¯hintegraldisplayr
A(r′)·dr′bracketrightbigg
ψ0(r). (1.156)
Fromtherelation
(−i¯h∇−eA)ψ=expbracketleftbiggie
¯hintegraldisplay
A·dr′bracketrightbiggbraceleftbigg
(−i¯h∇−eA)ψ0−i¯hψ0ie
¯hAbracerightbigg
=expbracketleftbiggie
¯hintegraldisplay
A·dr′bracketrightbigg
(−i¯h∇ψ0), (1.157)
itisobviousthat ψsolvesEq.(1.155)if ψ0solvesEq.(1.154).The gaugecovariantderiv-
ative∇−i(e/¯h)Adescribesthecouplingofachargedparticlewiththemagneticfield.Itis
often called minimal substitution and plays a central role in quantum electromagnetism,
thefirst andsimplestgaugetheoryinphysics.
To summarize this discussion of the vector potential :When a vector Bis solenoidal, a
vector potential Aexists such that B=∇×A.Ais undetermined to within an additive
gradient.Thiscorrespondstothearbitraryzeroofapotential,aconstantofintegrationfor
thescalarpotential .
In many problems the magnetic vector potential Awill be obtained from the current
distributionthatproducesthemagneticinduction B.ThismeanssolvingPoisson’s(vector)
equation(see Exercise1.14.4).
1.13 Potential Theory 77
Exercises
1.13.1 If aforce Fisgivenby
F=parenleftbig
x2+y2+z2parenrightbign(ˆxx+ˆyy+ˆzz),
find
(a)∇·F.
(b)∇×F.
(c) Ascalarpotential ϕ(x,y,z) so thatF=−∇ϕ.
(d) Forwhatvalueoftheexponent ndoesthescalarpotentialdivergeatboththeorigin
andinfinity?
ANS.(a) (2n+3)r2n,(b)0,
(c)−1
2n+2r2n+2,n/negationslash=−1,(d)n=−1,
ϕ=−lnr.
1.13.2 A sphere of radius ais uniformly charged (throughout its volume). Construct the elec-
trostaticpotential ϕ(r)for 0/lessorequalslantr<∞.
Hint. In Section 1.14 it is shown that the Coulomb force on a test charge at r=r0
depends only on the charge at distances less than r0and is independent of the charge
at distances greater than r0. Note that this applies to a spherically symmetric charge
distribution.
1.13.3 The usual problem in classical mechanics is to calculate the motion of a particle given
the potential. For a uniform density ( ρ0), nonrotating massive sphere, Gauss’ law of
Section 1.14 leads to a gravitational force on a unit mass m0at a point r0produced by
theattractionofthemassat r/lessorequalslantr0.Themassat r>r0contributesnothingtotheforce.
(a) Showthat F/m0=−(4πGρ0/3)r,0/lessorequalslantr/lessorequalslanta,whereaistheradiusofthesphere.
(b) Findthecorrespondinggravitationalpotential, 0 /lessorequalslantr/lessorequalslanta.
(c) ImagineaverticalholerunningcompletelythroughthecenteroftheEarthandout
tothefarside.NeglectingtherotationoftheEarthandassumingauniformdensity
ρ0=5.5gm/cm3,calculatethenatureofthemotionofaparticledroppedintothe
hole.Whatisits period?
Note.F∝ris actually a very poor approximation. Because of varying density,
the approximation F=constant along the outer half of a radial line and F∝r
alongtheinnerhalfis amuchcloserapproximation.
1.13.4 The origin of the Cartesian coordinates is at the Earth’s center. The moon is on the z-
axis, a fixed distance Raway (center-to-center distance). The tidal force exerted by the
moononaparticleattheEarth’ssurface (point x,y,z)i sg i v e nb y
Fx=−GMmx
R3,F y=−GMmy
R3,F z=+2GMmz
R3.
Findthepotentialthatyieldsthistidalforce.
78 Chapter 1 Vector Analysis
ANS.−GMm
R3parenleftbigg
z2−1
2x2−1
2y2parenrightbigg
.
In termsoftheLegendrepolynomialsof
Chapter12thisbecomes
−GMm
R3r2P2(cosθ).
1.13.5 A long, straight wire carrying a current Iproduces a magnetic induction Bwith com-
ponents
B=µ0I
2πparenleftbigg
−y
x2+y2,x
x2+y2,0parenrightbigg
.
Findamagneticvectorpotential A.
ANS.A=−ˆz(µ0I/4π)ln(x2+y2).(Thissolutionisnotunique.)
1.13.6 If
B=ˆr
r2=parenleftbiggx
r3,y
r3,z
r3parenrightbigg
,
finda vector Asuchthat ∇×A=B.Onepossiblesolutionis
A=ˆxyz
r(x2+y2)−ˆyxz
r(x2+y2).
1.13.7 Showthatthepairofequations
A=1
2(B×r),B=∇×A
issatisfiedbyanyconstantmagneticinduction B.
1.13.8 VectorBis formedbytheproductoftwogradients
B=(∇u)×(∇v),
whereuandvarescalarfunctions.
(a) Showthat Bis solenoidal.
(b) Showthat
A=1
2(u∇v−v∇u)
isavectorpotentialfor B, inthat
B=∇×A.
1.13.9 The magnetic induction Bis related to the magnetic vector potential AbyB=∇×A.
ByStokes’theorem
integraldisplay
B·dσ=contintegraldisplay
A·dr.
1.14 Gauss’ Law, Poisson’s Equation 79
Showthateachsideofthisequationisinvariantunderthe gaugetransformation ,A→
A+∇ϕ.
Note. Take the function ϕto be single-valued. The complete gauge transformation is
consideredinExercise4.6.4.
1.13.10 WithEtheelectricfieldand Athemagneticvectorpotential,showthat [E+∂A/∂t]is
irrotationalandthatthereforewemaywrite
E=−∇ϕ−∂A
∂t.
1.13.11 Thetotalforce ona charge qmovingwithvelocity vis
F=q(E+v×B).
Usingthescalarandvectorpotentials,showthat
F=qbracketleftbigg
−∇ϕ−dA
dt+∇(A·v)bracketrightbigg
.
Note that we now have a total time derivative of Ain place of the partial derivative of
Exercise1.13.10.
1.14 G AUSS ’LAW,POISSON ’SEQUATION
Gauss’ Law
Considerapointelectriccharge qattheoriginofourcoordinatesystem.Thisproducesan
electricfield Egivenby26
E=qˆr
4πε0r2. (1.158)
We now derive Gauss’ law, whichstates thatthe surface integralin Fig. 1.35 is q/ε0if the
closedsurface S=∂Vincludestheorigin(where qislocated)andzeroifthesurfacedoes
notincludetheorigin.Thesurface Sisanyclosedsurface; itneednotbespherical.
Using Gauss’ theorem, Eqs. (1.101a) and (1.101b) (and neglecting the q/4πε0), we
obtain
integraldisplay
Sˆr·dσ
r2=integraldisplay
V∇·parenleftbiggˆr
r2parenrightbigg
dτ=0 (1.159)
byExample1.7.2,providedthesurface Sdoesnotincludetheorigin,wheretheintegrands
arenotdefined.ThisprovesthesecondpartofGauss’law.
The first part, in which the surface Smust include the origin, may be handled by sur-
rounding the origin with a small sphere S′=∂V′of radius δ(Fig. 1.36). So that there
will be no question what is inside and what is outside, imagine the volume outside the
outer surface Sand the volume inside surface S′(r <δ)connected by a small hole. This
26Theelectricfield Eisdefinedastheforceperunitchargeonasmallstationarytestcharge qt:E=F/qt.FromCoulomb’slaw
the force on qtdue toqisF=(qqt/4πε0)(ˆr/r2).Wh enwed i v i d eb y qt, Eq.(1.158) follows.
80 Chapter 1 Vector Analysis
FIGURE 1.35Gauss’ law.
FIGURE 1.36Exclusionof theorigin.
joins surfaces SandS′, combining them into one single simply connected closed surface.
Because the radius of the imaginary hole may be made vanishingly small, there is no ad-
ditional contribution to the surface integral. The inner surface is deliberately chosen to be
1.14 Gauss’ Law, Poisson’s Equation 81
spherical so that we will be able to integrate over it. Gauss’ theorem now applies to the
volumebetween SandS′withoutanydifficulty.Wehave
integraldisplay
Sˆr·dσ
r2+integraldisplay
S′ˆr·dσ′
δ2=0. (1.160)
We may evaluate the second integral, for dσ′=−ˆrδ2d/Omega1, in which d/Omega1is an element of
solidangle.TheminussignappearsbecauseweagreedinSection1.10tohavethepositive
normalˆr′outward from the volume. In this case the outward ˆr′is in the negative radial
direction,ˆr′=−ˆr. Byintegratingoverallangles,wehave
integraldisplay
S′ˆr·dσ′
δ2=−integraldisplay
S′ˆr·ˆrδ2d/Omega1
δ2=−4π, (1.161)
independentof theradius δ. WiththeconstantsfromEq. (1.158), thisresultsin
integraldisplay
SE·dσ=q
4πε04π=q
ε0, (1.162)
completing the proof of Gauss’ law. Notice that although the surface Smay be spherical,
itneednot bespherical.Goingjustabitfurther, weconsideradistributedchargeso that
q=integraldisplay
Vρdτ. (1.163)
Equation (1.162) still applies, with qnow interpreted as the total distributed charge en-
closedbysurface S:
integraldisplay
SE·dσ=integraldisplay
Vρ
ε0dτ. (1.164)
UsingGauss’ theorem,wehave
integraldisplay
V∇·Edτ=integraldisplay
Vρ
ε0dτ. (1.165)
Sinceourvolumeiscompletelyarbitrary, theintegrandsmustbeequal,or
∇·E=ρ
ε0, (1.166)
one of Maxwell’s equations. If we reverse the argument, Gauss’ law follows immediately
fromMaxwell’sequation.
Poisson’s Equation
If wereplace Eby−∇ϕ, Eq. (1.166)becomes
∇·∇ϕ=−ρ
ε0, (1.167a)
82 Chapter 1 Vector Analysis
which is Poisson’s equation. For the condition ρ=0 this reduces to an even more famous
equation,
∇·∇ϕ=0, (1.167b)
Laplace’s equation. We encounter Laplace’s equation frequently in discussing various co-
ordinatesystems(Chapter2)andthespecialfunctionsofmathematicalphysicsthatappear
as its solutions. Poisson’s equation will be invaluable in developing the theory of Green’s
functions(Section9.7).
From direct comparison of the Coulomb electrostatic force law and Newton’s law of
universalgravitation,
FE=1
4πε0q1q2
r2ˆr,FG=−Gm1m2
r2ˆr.
All of the potential theory of this section applies equally well to gravitational potentials.
Forexample,thegravitationalPoissonequationis
∇·∇ϕ=+4πGρ, (1.168)
withρnowamass density.
Exercises
1.14.1 DevelopGauss’ lawfor thetwo-dimensionalcaseinwhich
ϕ=−qlnρ
2πε0,E=−∇ϕ=qˆρ
2πε0ρ.
Hereqisthechargeattheoriginorthelinechargeperunitlengthifthetwo-dimensional
systemisaunitthicknesssliceofathree-dimensional(circularcylindrical)system.The
variableρis measured radially outward from the line charge. ˆρis the corresponding
unitvector(seeSection2.4).
1.14.2 (a) ShowthatGauss’lawfollowsfrom Maxwell’sequation
∇·E=ρ
ε0.
Hereρistheusualchargedensity.
(b) Assumingthattheelectricfieldofapointcharge qissphericallysymmetric,show
thatGauss’ lawimpliestheCoulombinversesquareexpression
E=qˆr
4πε0r2.
1.14.3 Showthatthevalueoftheelectrostaticpotential ϕatanypoint Pisequaltotheaverage
of the potential over any spherical surface centered on P. There are no electric charges
onorwithinthesphere.
Hint.UseGreen’stheorem,Eq.(1.104),with u−1=r,thedistancefrom P,andv=ϕ.
AlsonoteEq. (1.170)inSection1.15.
1.15 Dirac Delta Function 83
1.14.4 UsingMaxwell’sequations,showthatforasystem(steadycurrent)themagneticvector
potential Asatisfies avectorPoissonequation,
∇2A=−µ0J,
providedwerequire ∇·A=0.
1.15 D IRAC DELTA FUNCTION
FromExample1.6.1 andthedevelopmentofGauss’lawinSection1.14,
integraldisplay
∇·∇parenleftbigg1
rparenrightbigg
dτ=−integraldisplay
∇·parenleftbiggˆr
r2parenrightbigg
dτ=braceleftbigg−4π
0,(1.169)
depending on whether or not the integration includes the origin r=0. This result may be
convenientlyexpressedbyintroducingtheDiracdeltafunction,
∇2parenleftbigg1
rparenrightbigg
=−4πδ(r)≡−4πδ(x)δ(y)δ(z). (1.170)
ThisDiracdeltafunctionis definedbyits assignedproperties
δ(x)=0,x/negationslash=0 (1.171a)
f(0)=integraldisplay∞
−∞f(x)δ(x)dx, (1.171b)
wheref(x)is any well-behaved function and the integration includes the origin. As a
specialcaseofEq. (1.171b),
integraldisplay∞
−∞δ(x)dx=1. (1.171c)
FromEq.(1.171b), δ(x)mustbeaninfinitelyhigh,infinitelythinspikeat x=0,asinthe
descriptionofanimpulsiveforce(Section15.9)orthechargedensityforapointcharge.27
The problem is that no such function exists , in the usual sense of function. However, the
crucial property in Eq. (1.171b) can be developed rigorously as the limit of a sequence
of functions, a distribution. For example, the delta function may be approximated by the
27The delta function is frequently invoked to describe very short-range forces, such as nuclear forces. It also appears in the
normalization of continuum wavefunctions of quantum mechanics.Compare Eq.(1.193c) for plane-waveeigenfunctions.
84 Chapter 1 Vector Analysis
FIGURE 1.37δ-Sequence
function.
FIGURE 1.38δ-Sequence
function.
sequencesoffunctions,Eqs. (1.172)to(1.175) andFigs. 1.37to1.40:
δn(x)=
0,x<−1
2n
n,−1
2n<x<1
2n
0,x>1
2n(1.172)
δn(x)=n√πexpparenleftbig
−n2x2parenrightbig
(1.173)
δn(x)=n
π·1
1+n2x2(1.174)
δn(x)=sinnx
πx=1
2πintegraldisplayn
−neixtdt. (1.175)
1.15 Dirac Delta Function 85
FIGURE 1.39δ-Sequencefunction.
FIGURE 1.40δ-Sequencefunction.
These approximations have varying degrees of usefulness. Equation (1.172) is useful in
providing a simple derivation of the integral property, Eq. (1.171b). Equation (1.173)
is convenient to differentiate. Its derivatives lead to the Hermite polynomials. Equa-
tion (1.175) is particularly useful in Fourier analysis and in its applications to quantum
mechanics. In the theory of Fourier series, Eq. (1.175) often appears (modified) as the
Dirichletkernel:
δn(x)=1
2πsin[(n+1
2)x]
sin(1
2x). (1.176)
In using these approximations in Eq. (1.171b) and later, we assume that f(x)is well be-
haved—itoffers noproblemsatlarge x.
86 Chapter 1 Vector Analysis
For most physical purposes such approximations are quite adequate. From a mathemat-
icalpointofviewthesituationis stillunsatisfactory:Thelimits
limn→∞δn(x)
donotexist .
A way out of this difficulty is provided by the theory of distributions. Recognizing that
Eq. (1.171b) is the fundamental property, we focus our attention on it rather than on δ(x)
itself.Equations(1.172)to(1.175)with n=1,2,3,...maybeinterpretedas sequences of
normalizedfunctions:
integraldisplay∞
−∞δn(x)dx=1. (1.177)
Thesequenceofintegralshasthelimit
limn→∞integraldisplay∞
−∞δn(x)f(x)dx=f(0). (1.178)
Note that Eq. (1.178) is the limit of a sequence of integrals. Again, the limit of δn(x),
n→∞, doesnotexist.(Thelimitsforallfour formsof δn(x)divergeat x=0.)
We maytreat δ(x)consistentlyintheform
integraldisplay∞
−∞δ(x)f(x)dx=limn→∞integraldisplay∞
−∞δn(x)f(x)dx. (1.179)
δ(x)is labeled a distribution (not a function) defined by the sequences δn(x)as indicated
inEq.(1.179).Wemightemphasizethattheintegralontheleft-handsideofEq.(1.179)is
nota Riemannintegral.28It isalimit.
Thisdistribution δ(x)isonlyoneofaninfinityofpossibledistributions,butitistheone
weare interestedinbecauseofEq. (1.171b).
FromthesesequencesoffunctionsweseethatDirac’sdeltafunctionmustbeevenin x,
δ(−x)=δ(x).
The integral property, Eq. (1.171b), is useful in cases where the argument of the delta
functionisafunction g(x)withsimplezerosontherealaxis, whichleadstotherules
δ(ax)=1
aδ(x), a> 0, (1.180)
δparenleftbig
g(x)parenrightbig
=summationdisplay
a,
g(a)=0,
g′(a)/negationslash=0δ(x−a)
|g′(a)|. (1.181a)
Equation(1.180)maybewritten
integraldisplay∞
−∞f(x)δ(ax)dx=1
aintegraldisplay∞
−∞fparenleftbiggy
aparenrightbigg
δ(y)dy=1
af(0),
28It can be treated as a Stieltjes integral if desired. δ(x)dxis replaced by du(x),w h e r eu(x)is the Heaviside step function
(compare Exercise1.15.13).
1.15 Dirac Delta Function 87
applying Eq. (1.171b). Equation (1.180) may be written as δ(ax)=1
|a|δ(x)fora<0.To
proveEq.(1.181a)wedecomposetheintegral
integraldisplay∞
−∞f(x)δparenleftbig
g(x)parenrightbig
dx=summationdisplay
aintegraldisplaya+ε
a−εf(x)δparenleftbig
(x−a)g′(a)parenrightbig
dx (1.181b)
intoasumofintegralsoversmallintervalscontainingthezerosof g(x).Intheseintervals,
g(x)≈g(a)+(x−a)g′(a)=(x−a)g′(a). Using Eq. (1.180) on the right-hand side of
Eq.(1.181b)weobtaintheintegralofEq. (1.181a).
Using integration by parts we can also define the derivative δ′(x)of the Dirac delta
functionbytherelation
integraldisplay∞
−∞f(x)δ′(x−x′)dx=−integraldisplay∞
−∞f′(x)δ(x−x′)dx=−f′(x′). (1.182)
We useδ(x)frequently and call it the Dirac delta function29—for historical reasons.
Remember that it is not really a function. It is essentially a shorthand notation, defined
implicitlyas thelimit of integralsina sequence, δn(x), accordingto Eq. (1.179). It should
be understood that our Dirac delta function has significance only as part of an integrand.
Inthisspirit, thelinearoperatorintegraltext
dxδ(x−x0)operateson f(x)andyields f(x0):
L(x0)f(x)≡integraldisplay∞
−∞δ(x−x0)f(x)dx=f(x0). (1.183)
It may also be classified as a linear mapping or simply as a generalized function. Shift-
ing our singularity to the point x=x′, we write the Dirac delta function as δ(x−x′).
Equation(1.171b)becomes
integraldisplay∞
−∞f(x)δ(x−x′)dx=f(x′). (1.184)
As a description of a singularity at x=x′, the Dirac delta function may be written as
δ(x−x′)orasδ(x′−x).Goingtothreedimensionsandusingsphericalpolarcoordinates,
weobtain
integraldisplay2π
0integraldisplayπ
0integraldisplay∞
0δ(r)r2drsinθdθdϕ=integraldisplayintegraldisplayintegraldisplay∞
−∞δ(x)δ(y)δ(z)dxdydz =1.(1.185)
Thiscorrespondstoasingularity(orsource)attheorigin.Again,ifoursourceisat r=r1,
Eq.(1.185) becomes
integraldisplayintegraldisplayintegraldisplay
δ(r2−r1)r2
2dr2sinθ2dθ2dϕ2=1. (1.186)
29Diracintroducedthedeltafunctiontoquantummechanics.Actually,thedeltafunctioncanbetracedbacktoKirchhoff,1882.
For further details see M. Jammer, The Conceptual Development of Quantum Mechanics . New York: McGraw–Hill (1966),
p. 301.
88 Chapter 1 Vector Analysis
Example 1.15.1 TOTAL CHARGE INSIDE A SPHERE
Consider the total electric fluxcontintegraltext
E·dσout of a sphere of radius Raround the origin
surrounding nchargesej,located at the points rjwithrj<R, that is, inside the sphere.
Theelectricfieldstrength E=−∇ϕ(r),wherethepotential
ϕ=nsummationdisplay
j=1ej
|r−rj|=integraldisplayρ(r′)
|r−r′|d3r′
isthesumoftheCoulombpotentialsgeneratedbyeachchargeandthetotalchargedensity
isρ(r)=summationtext
jejδ(r−rj).Thedeltafunctionisusedhereasanabbreviationofapointlike
density.NowweuseGauss’theoremfor
contintegraldisplay
E·dσ=−contintegraldisplay
∇ϕ·dσ=−integraldisplay
∇2ϕdτ=integraldisplayρ(r)
ε0dτ=summationtext
jej
ε0
inconjunctionwiththedifferentialform ofGauss’slaw, ∇·E=−ρ/ε0,and
summationdisplay
jejintegraldisplay
δ(r−rj)dτ=summationdisplay
jej.
/squaresolid
Example 1.15.2 PHASE SPACE
InthescatteringtheoryofrelativisticparticlesusingFeynmandiagrams,weencounterthe
followingintegraloverenergyof thescatteredparticle(wesetthevelocityoflight c=1):
integraldisplay
d4pδparenleftbig
p2−m2parenrightbig
f(p)≡integraldisplay
d3pintegraldisplay
dp0δparenleftbig
p2
0−p2−m2parenrightbig
f(p)
=integraldisplay
E>0d3pf(E,p)
2radicalbig
m2+p2+integraldisplay
E<0d3pf(E,p)
2radicalbig
m2+p2,
where we have used Eq. (1.181a) at the zeros E=±radicalbig
m2+p2of the argument of the
delta function. The physical meaning of δ(p2−m2)is that the particle of mass mand
four-momentum pµ=(p0,p)is on its mass shell, because p2=m2is equivalent to E=
±radicalbig
m2+p2. Thus, the on-mass-shell volume element in momentum space is the Lorentz
invariantd3p
2E, in contrast to the nonrelativistic d3pof momentum space. The fact that
a negative energy occurs is a peculiarity of relativistic kinematics that is related to the
antiparticle. /squaresolid
Delta Function Representation by Orthogonal
Functions
Dirac’sdeltafunction30canbeexpandedintermsofanybasisofrealorthogonalfunctions
{ϕn(x),n=0,1,2,...}. Such functions will occur in Chapter 10 as solutions of ordinary
differentialequationsof theSturm–Liouvilleform.
30This sectionis optional here. It is not needed until Chapter10.
1.15 Dirac Delta Function 89
Theysatisfytheorthogonalityrelations
integraldisplayb
aϕm(x)ϕn(x)dx=δmn, (1.187)
where the interval (a,b)may be infinite at either end or both ends. [For convenience we
assumethat ϕnhasbeendefinedtoinclude (w(x))1/2iftheorthogonalityrelationscontain
an additional positive weight function w(x).] We use the ϕnto expand the delta function
as
δ(x−t)=∞summationdisplay
n=0an(t)ϕn(x), (1.188)
where the coefficients anare functions of the variable t. Multiplying by ϕm(x)and inte-
gratingovertheorthogonalityinterval(Eq.(1.187)), wehave
am(t)=integraldisplayb
aδ(x−t)ϕm(x)dx=ϕm(t) (1.189)
or
δ(x−t)=∞summationdisplay
n=0ϕn(t)ϕn(x)=δ(t−x). (1.190)
This series is assuredly not uniformly convergent (see Chapter 5), but it may be used as
part of an integrand in which the ensuing integration will make it convergent (compare
Section5.5).
Suppose we form the integralintegraltext
F(t)δ(t−x)dx, where it is assumed that F(t)can be
expanded in a series of orthogonal functions ϕp(t), a property called completeness .W e
thenobtain
integraldisplay
F(t)δ(t−x)dt=integraldisplay∞summationdisplay
p=0apϕp(t)∞summationdisplay
n=0ϕn(x)ϕn(t)dt
=∞summationdisplay
p=0apϕp(x)=F(x), (1.191)
the cross productsintegraltext
ϕpϕndt(n/negationslash=p)vanishing by orthogonality (Eq. (1.187)). Referring
back to the definition of the Dirac delta function, Eq. (1.171b), we see that our series
representation, Eq. (1.190), satisfies the defining property of the Dirac delta function and
therefore is a representation of it. This representation of the Dirac delta function is called
closure. The assumption of completeness of a set of functions for expansion of δ(x−t)
yieldstheclosurerelation.Theconverse,thatclosureimpliescompleteness,isthetopicof
Exercise1.15.16.
90 Chapter 1 Vector Analysis
Integral Representations for the Delta Function
Integraltransforms, suchastheFourierintegral
F(ω)=integraldisplay∞
−∞f(t)exp(iωt)dt
of Chapter 15, lead to thecorrespondingintegralrepresentationsof Dirac’s deltafunction.
Forexample,take
δn(t−x)=sinn(t−x)
π(t−x)=1
2πintegraldisplayn
−nexpparenleftbig
iω(t−x)parenrightbig
dω, (1.192)
usingEq. (1.175). Wehave
f(x)=limn→∞integraldisplay∞
−∞f(t)δn(t−x)dt, (1.193a)
whereδn(t−x)isthesequenceinEq.(1.192)definingthedistribution δ(t−x).Notethat
Eq. (1.193a) assumes that f(t)is continuous at t=x. If we substitute Eq. (1.192) into
Eq.(1.193a)weobtain
f(x)=limn→∞1
2πintegraldisplay∞
−∞f(t)integraldisplayn
−nexpparenleftbig
iω(t−x)parenrightbig
dωdt. (1.193b)
Interchanging the order of integration and then taking the limit as n→∞,w eh a v et h e
Fourierintegraltheorem,Eq.(15.20).
With the understanding that it belongs under an integral sign, as in Eq. (1.193a), the
identification
δ(t−x)=1
2πintegraldisplay∞
−∞expparenleftbig
iω(t−x)parenrightbig
dω (1.193c)
providesaveryusefulintegralrepresentationof thedeltafunction.
WhentheLaplacetransform(see Sections15.1and15.9)
Lδ(s)=integraldisplay∞
0exp(−st)δ(t−t0)=exp(−st0), t 0>0 (1.194)
isinverted,weobtainthecomplexrepresentation
δ(t−t0)=1
2πiintegraldisplayγ+i∞
γ−i∞expparenleftbig
s(t−t0)parenrightbig
ds, (1.195)
whichisessentiallyequivalenttothepreviousFourierrepresentationofDirac’sdeltafunc-
tion.
1.15 Dirac Delta Function 91
Exercises
1.15.1 Let
δn(x)=
0,x<−1
2n,
n,−1
2n<x<1
2n,
0,1
2n<x.
Showthat
limn→∞integraldisplay∞
−∞f(x)δn(x)dx=f(0),
assumingthat f(x)iscontinuousat x=0.
1.15.2 Verifythatthesequence δn(x), basedonthefunction
δn(x)=braceleftbigg0,x <0,
ne−nx,x>0,
is a delta sequence (satisfying Eq. (1.178)). Note that the singularity is at +0, the posi-
tivesideof theorigin.
Hint. Replace the upper limit ( ∞)b yc/n, wherecis large but finite, and use the mean
valuetheoremof integralcalculus.
1.15.3 For
δn(x)=n
π·1
1+n2x2,
(Eq. (1.174)), showthat
integraldisplay∞
−∞δn(x)dx=1.
1.15.4 Demonstratethat δn=sinnx/πxisadeltadistributionbyshowingthat
limn→∞integraldisplay∞
−∞f(x)sinnx
πxdx=f(0).
Assumethat f(x)is continuousat x=0 andvanishesas x→±∞.
Hint.Replace xbyy/nandtake lim n→∞beforeintegrating.
1.15.5 Fejer’smethodofsummingseries isassociatedwiththefunction
δn(t)=1
2πnbracketleftbiggsin(nt/2)
sin(t/2)bracketrightbigg2
.
Showthat δn(t)isadeltadistribution,inthesensethat
limn→∞1
2πnintegraldisplay∞
−∞f(t)bracketleftbiggsin(nt/2)
sin(t/2)bracketrightbigg2
dt=f(0).
92 Chapter 1 Vector Analysis
1.15.6 Provethat
δbracketleftbig
a(x−x1)bracketrightbig
=1
aδ(x−x1).
Note.Ifδ[a(x−x1)]isconsideredeven,relativeto x1,therelationholdsfornegative a
and 1/amaybereplacedby 1 /|a|.
1.15.7 Showthat
δbracketleftbig
(x−x1)(x−x2)bracketrightbig
=bracketleftbig
δ(x−x1)+δ(x−x2)bracketrightbig
/|x1−x2|.
Hint.TryusingExercise1.15.6.
1.15.8 UsingtheGausserror curvedeltasequence( δn=n√πe−n2x2), showthat
xd
dxδ(x)=−δ(x),
treatingδ(x)anditsderivativeas inEq.(1.179).
1.15.9 Showthatintegraldisplay∞
−∞δ′(x)f(x)dx=−f′(0).
Hereweassumethat f′(x)is continuousat x=0.
1.15.10 Provethat
δparenleftbig
f(x)parenrightbig
=vextendsinglevextendsinglevextendsinglevextendsingledf(x)
dxvextendsinglevextendsinglevextendsinglevextendsingle−1
x=x0δ(x−x0),
wherex0ischosenso that f(x0)=0.
Hint.Notethat δ(f)df=δ(x)dx.
1.15.11 Show that in spherical polar coordinates (r,cosθ,ϕ)the delta function δ(r1−r2)be-
comes
1
r2
1δ(r1−r2)δ(cosθ1−cosθ2)δ(ϕ1−ϕ2).
Generalize this to the curvilinear coordinates (q1,q2,q3)of Section 2.1 with scale fac-
torsh1,h2, andh3.
1.15.12 Arigorousdevelopmentof Fouriertransforms31includesas atheoremtherelations
lima→∞2
πintegraldisplayx2
x1f(u+x)sinax
xdx
=
f(u+0)+f(u−0), x 1<0<x2
f(u+0), x 1=0<x2
f(u−0), x 1<0=x2
0,x 1<x2<0or0<x1<x2.
Verifytheseresults usingtheDiracdeltafunction.
31I. N.Sneddon, Fourier Transforms . NewYork: McGraw-Hill(1951).
1.15 Dirac Delta Function 93
FIGURE 1.411
2[1+tanhnx]andtheHeavisideunitstep
function.
1.15.13 (a) If wedefinea sequence δn(x)=n/(2cosh2nx), showthat
integraldisplay∞
−∞δn(x)dx=1,independentof n.
(b) Continuingthisanalysis,showthat32
integraldisplayx
−∞δn(x)dx=1
2[1+tanhnx]≡un(x),
limn→∞un(x)=braceleftbigg0,x<0,
1,x>0.
Thisis theHeavisideunitstepfunction(Fig.1.41).
1.15.14 Showthattheunitstepfunction u(x)mayberepresentedby
u(x)=1
2+1
2πiPintegraldisplay∞
−∞eixtdt
t,
wherePmeansCauchyprincipalvalue(Section7.1).
1.15.15 Asavariationof Eq.(1.175), take
δn(x)=1
2πintegraldisplay∞
−∞eixt−|t|/ndt.
Showthatthisreducesto (n/π)1/(1+n2x2), Eq.(1.174), andthat
integraldisplay∞
−∞δn(x)dx=1.
Note. In terms of integral transforms, the initial equation here may be interpreted as
eitheraFourierexponentialtransformof e−|t|/noraLaplacetransformof eixt.
32Manyothersymbolsareusedforthisfunction.ThisistheAMS-55(seefootnote4onp.330forthereference)notation: ufor
unit.
94 Chapter 1 Vector Analysis
1.15.16 (a) TheDiracdeltafunctionrepresentationgivenbyEq.(1.190),
δ(x−t)=∞summationdisplay
n=0ϕn(x)ϕn(t),
is often called the closure relation . For an orthonormal set of real functions,
ϕn,show that closure implies completeness, that is, Eq. (1.191) follows from
Eq.(1.190).
Hint.Onecantake
F(x)=integraldisplay
F(t)δ(x−t)dt.
(b) Following the hint of part (a) you encounter the integralintegraltext
F(t)ϕn(t)dt.H o wd o
youknowthatthisintegralis finite?
1.15.17 For the finite interval (−π,π)write the Dirac delta function δ(x−t)as a series of
sines and cosines: sin nx,cosnx,n=0,1,2,....Note that although these functions
areorthogonal,theyarenotnormalizedtounity.
1.15.18 Intheinterval (−π,π),δn(x)=n√πexp(−n2x2).
(a) Write δn(x)asaFouriercosineseries.
(b) Show that your Fourier series agrees with a Fourier expansion of δ(x)in the limit
asn→∞.
(c) Confirm the delta function nature of your Fourier series by showing that for any
f(x)thatisfiniteintheinterval [−π,π]andcontinuousat x=0,
integraldisplayπ
−πf(x)bracketleftbig
Fourierexpansionof δ∞(x)bracketrightbig
dx=f(0).
1.15.19 (a) Write δn(x)=n√πexp(−n2x2)in the interval (−∞,∞)as a Fourier integral and
comparethelimit n→∞withEq. (1.193c).
(b) Write δn(x)=nexp(−nx)as a Laplace transform and compare the limit n→∞
withEq. (1.195).
Hint.SeeEqs. (15.22) and(15.23)for (a) andEq.(15.212)for (b).
1.15.20 (a) Show that the Dirac delta function δ(x−a), expanded in a Fourier sine series in
thehalf-interval (0,L),(0<a<L) , isgivenby
δ(x−a)=2
L∞summationdisplay
n=1sinparenleftbiggnπa
Lparenrightbigg
sinparenleftbiggnπx
Lparenrightbigg
.
Notethatthisseriesactuallydescribes
−δ(x+a)+δ(x−a)intheinterval (−L,L).
(b) By integrating both sides of the preceding equation from 0 to x, show that the
cosineexpansionofthesquarewave
f(x)=braceleftbigg0,0/lessorequalslantx<a
1, a<x<L,
1.16 Helmholtz’s Theorem 95
is, for 0 /lessorequalslantx<L,
f(x)=2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
−2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
cosparenleftbiggnπx
Lparenrightbigg
.
(c) Verifythattheterm
2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
isangbracketleftbig
f(x)angbracketrightbig
≡1
LintegraldisplayL
0f(x)dx.
1.15.21 Verify the Fourier cosine expansion of the square wave, Exercise 1.15.20(b), by direct
calculationoftheFouriercoefficients.
1.15.22 Wemaydefineasequence
δn(x)=braceleftbiggn,|x|<1/2n,
0,|x|>1/2n.
(This is Eq. (1.172).) Express δn(x)as a Fourier integral (via the Fourier integral theo-
rem,inversetransform,etc.). Finally,showthatwemaywrite
δ(x)=limn→∞δn(x)=1
2πintegraldisplay∞
−∞e−ikxdk.
1.15.23 Usingthesequence
δn(x)=n√πexpparenleftbig
−n2x2parenrightbig
,
showthat
δ(x)=1
2πintegraldisplay∞
−∞e−ikxdk.
Note. Remember that δ(x)is defined in terms of its behavior as part of an integrand—
especiallyEqs. (1.178)and(1.189).
1.15.24 Derivesineandcosinerepresentationsof δ(t−x)thatarecomparabletotheexponential
representation,Eq. (1.193c).
ANS.2
πintegraltext∞
0sinωtsinωxdω,2
πintegraltext∞
0cosωtcosωxdω.
1.16 H ELMHOLTZ ’STHEOREM
InSection1.13itwasemphasizedthatthechoiceofamagneticvectorpotential Awasnot
unique.Thedivergenceof Awasstillundetermined.Inthissectiontwotheoremsaboutthe
divergenceandcurlofavectorare developed.Thefirst theoremisas follows:
A vector is uniquely specified by giving its divergence and its curl within a simply con-
nectedregion(withoutholes)anditsnormalcomponentovertheboundary .
96 Chapter 1 Vector Analysis
Note that the subregions, where the divergence and curl are defined (often in terms of
Dirac delta functions), are part of our region and are not supposed to be removed here or
inHelmholtz’stheorem,whichfollows.Letustake
∇·V1=s,
∇×V1=c,(1.196)
wheresmay be interpreted as a source (charge) density and cas a circulation (current)
density. Assuming also that the normal component V1non the boundary is given, we want
to show that V1is unique. We do this by assuming the existence of a second vector, V2,
which satisfies Eq. (1.196) and has the same normal component over the boundary, and
thenshowingthat V1−V2=0.Let
W=V1−V2.
Then
∇·W=0 (1.197)
and
∇×W=0. (1.198)
SinceWisirrotationalwemaywrite(bySection(1.13))
W=−∇ϕ. (1.199)
SubstitutingthisintoEq.(1.197), weobtain
∇·∇ϕ=0, (1.200)
Laplace’sequation.
Now we draw upon Green’s theorem in the form given in Eq. (1.105), letting uandv
eachequal ϕ. Since
Wn=V1n−V2n=0 (1.201)
ontheboundary,Green’stheoremreducesto
integraldisplay
V(∇ϕ)·(∇ϕ)dτ=integraldisplay
VW·Wdτ=0. (1.202)
Thequantity W·W=W2isnonnegativeandsowemusthave
W=V1−V2=0 (1.203)
everywhere.Thus V1is unique,provingthetheorem.
For our magnetic vector potential Athe relation B=∇×Aspecifies the curl of A.
Often for convenience we set ∇·A=0 (compare Exercise 1.14.4). Then (with boundary
conditions) Aisfixed.
This theorem may be written as a uniqueness theorem for solutions of Laplace’s equa-
tion, Exercise 1.16.1. In this form, this uniqueness theorem is of great importance in solv-
ing electrostatic and other Laplace equation boundary value problems. If we can find a
solution of Laplace’s equation that satisfies the necessary boundary conditions, then our
solution is the complete solution. Such boundary value problems are taken up in Sec-
tions12.3and12.5.
1.16 Helmholtz’s Theorem 97
Helmholtz’s Theorem
ThesecondtheoremweshallproveisHelmholtz’stheorem.
A vector Vsatisfying Eq. (1.196)with both source and circulation densities vanishing
atinfinitymaybewrittenas thesum oftwoparts, oneofwhichisirrotational,theotherof
whichis solenoidal .
Notethatourregionissimplyconnected,beingallofspace,forsimplicity.Helmholtz’s
theoremwillclearlybesatisfiedif wemaywrite Vas
V=−∇ϕ+∇×A, (1.204a)
−∇ϕbeingirrotationaland ∇×Abeingsolenoidal.We proceedtojustifyEq.(1.204a).
Vis aknownvector.Wetakethedivergenceandcurl
∇·V=s(r) (1.204b)
∇×V=c(r) (1.204c)
withs(r)andc(r)nowknownfunctionsofposition.Fromthesetwofunctionsweconstruct
ascalarpotential ϕ(r1),
ϕ(r1)=1
4πintegraldisplays(r2)
r12dτ2, (1.205a)
andavectorpotential A(r1),
A(r1)=1
4πintegraldisplayc(r2)
r12dτ2. (1.205b)
Ifs=0, thenVis solenoidal and Eq. (1.205a) implies ϕ=0. From Eq. (1.204a), V=
∇×A, withAas given in Eq. (1.141), which is consistent with Section 1.13. Further,
ifc=0, thenVis irrotational and Eq. (1.205b) implies A=0, and Eq. (1.204a) implies
V=−∇ϕ, consistentwithscalarpotentialtheoryof Section1.13.
Here the argument r1indicates (x1,y1,z1), the field point; r2, the coordinates of the
sourcepoint( x2,y2,z2), whereas
r12=bracketleftbig
(x1−x2)2+(y1−y2)2+(z1−z2)2bracketrightbig1/2. (1.206)
When a direction is associated with r12, the positive direction is taken to be away from
the source and toward the field point. Vectorially, r12=r1−r2, as shown in Fig. 1.42.
Of course, sandcmust vanish sufficiently rapidly at large distance so that the integrals
exist. The actual expansion and evaluation of integrals such as Eqs. (1.205a) and (1.205b)
istreatedinSection12.1.
From the uniqueness theorem at the beginning of this section, Vis uniquely specified
by its divergence, s, and curl, c(and boundary conditions). Returning to Eq. (1.204a), we
have
∇·V=−∇·∇ϕ, (1.207a)
thedivergenceofthecurlvanishing,and
∇×V=∇×(∇×A), (1.207b)
98 Chapter 1 Vector Analysis
FIGURE 1.42Sourceandfieldpoints.
thecurlof thegradientvanishing.If wecanshowthat
−∇·∇ϕ(r1)=s(r1) (1.207c)
and
∇×parenleftbig
∇×A(r1)parenrightbig
=c(r1), (1.207d)
thenVas given in Eq. (1.204a) will have the proper divergence and curl. Our description
willbeinternallyconsistentandEq.(1.204a) justified.33
First, weconsiderthedivergenceof V:
∇·V=−∇·∇ϕ=−1
4π∇·∇integraldisplays(r2)
r12dτ2. (1.208)
The Laplacian operator, ∇·∇,o r∇2, operates on the field coordinates (x1,y1,z1)and so
commuteswiththeintegrationwithrespectto (x2,y2,z2).We have
∇·V=−1
4πintegraldisplay
s(r2)∇2
1parenleftbigg1
r12parenrightbigg
dτ2. (1.209)
WemustmaketwominormodificationsinEq.(1.169)beforeapplyingit.First,oursource
is atr2, not at the origin. This means that a nonzero result from Gauss’ law appears if and
onlyif thesurface Sincludesthepoint r=r2. Toshowthis, werewriteEq. (1.170):
∇2parenleftbigg1
r12parenrightbigg
=−4πδ(r1−r2). (1.210)
33Alternatively, we could solve Eq. (1.207c), Poisson’s equation, and compare the solution with the constructed potential,
Eq.(1.205a). The solution of Poisson’s equation is developed in Section 9.7.
1.16 Helmholtz’s Theorem 99
Thisshift ofthesourceto r2maybeincorporatedinthedefiningequation(1.171b)as
δ(r1−r2)=0,r1/negationslash=r2, (1.211a)
integraldisplay
f(r1)δ(r1−r2)dτ1=f(r2). (1.211b)
Second, noting that differentiating r−1
12twice with respect to x2,y2,z2is the same as
differentiating twicewithrespectto x1,y1,z1,weha v e
∇2
1parenleftbigg1
r12parenrightbigg
=∇2
2parenleftbigg1
r12parenrightbigg
=−4πδ(r1−r2)
=−4πδ(r2−r1). (1.212)
RewritingEq.(1.209)andusingtheDiracdeltafunction,Eq.(1.212), wemayintegrateto
obtain
∇·V=−1
4πintegraldisplay
s(r2)∇2
2parenleftbigg1
r12parenrightbigg
dτ2
=−1
4πintegraldisplay
s(r2)(−4π)δ(r2−r1)dτ2
=s(r1). (1.213)
The final step follows from Eq. (1.211b), with the subscripts 1 and 2 exchanged. Our
result, Eq. (1.213), shows that the assumed forms of Vand of the scalar potential ϕare in
agreementwiththegivendivergence(Eq. (1.204b)).
TocompletetheproofofHelmholtz’stheorem,weneedtoshowthatourassumptionsare
consistentwithEq.(1.204c),thatis,thatthecurlof Visequalto c(r1).FromEq.(1.204a),
∇×V=∇×(∇×A)
=∇∇·A−∇2A. (1.214)
Thefirstterm, ∇∇·A,leadsto
4π∇∇·A=integraldisplay
c(r2)·∇1∇1parenleftbigg1
r12parenrightbigg
dτ2 (1.215)
byEq.(1.205b).Againreplacingthesecondderivativeswithrespectto x1,y1,z1bysecond
derivatives with respect to x2,y2,z2, we integrate each component34of Eq. (1.215) by
parts:
4π∇∇·A|x=integraldisplay
c(r2)·∇2∂
∂x2parenleftbigg1
r12parenrightbigg
dτ2
=integraldisplay
∇2·bracketleftbigg
c(r2)∂
∂x2parenleftbigg1
r12parenrightbiggbracketrightbigg
dτ2
−integraldisplaybracketleftbig
∇2·c(r2)bracketrightbig∂
∂x2parenleftbigg1
r12parenrightbigg
dτ2. (1.216)
34This avoids creating the tensor c(r2)∇2.
100 Chapter 1 Vector Analysis
The second integral vanishes because the circulation density cis solenoidal.35The first
integral may be transformed to a surface integral by Gauss’ theorem. If cis bounded in
space or vanishes faster that 1 /rfor large r, so that the integral in Eq. (1.205b) exists,
then by choosing a sufficiently large surface the first integral on the right-hand side of
Eq.(1.216) alsovanishes.
With∇∇·A=0,Eq. (1.214)nowreducesto
∇×V=−∇2A=−1
4πintegraldisplay
c(r2)∇2
1parenleftbigg1
r12parenrightbigg
dτ2. (1.217)
This is exactly like Eq. (1.209) except that the scalar s(r2)is replaced by the vector circu-
lation density c(r2). Introducing the Dirac delta function, as before, as a convenient way
ofcarryingouttheintegration,wefindthatEq.(1.217)reducestoEq.(1.196).Weseethat
our assumed forms of V, given by Eq. (1.204a), and of the vector potential A, given by
Eq.(1.205b), areinagreementwithEq. (1.196)specifyingthecurlof V.
This completes the proof of Helmholtz’s theorem, showing that a vector may be re-
solvedintoirrotationalandsolenoidalparts.Appliedtotheelectromagneticfield,wehave
resolved our field vector Vinto an irrotational electric field E, derived from a scalar po-
tentialϕ, and a solenoidal magnetic induction field B, derived from a vector potential A.
The source density s(r)may be interpreted as an electric charge density (divided by elec-
tric permittivity ε), whereas the circulation density c(r)becomes electric current density
(timesmagneticpermeability µ).
Exercises
1.16.1 Implicitinthissectionisaproofthatafunction ψ(r)isuniquelyspecifiedbyrequiring
itto(1)satisfyLaplace’sequationand(2)satisfyacompletesetofboundaryconditions.
Developthisproof explicitly.
1.16.2 (a) Assumingthat Pis a solutionof the vectorPoisson equation, ∇2
1P(r1)=−V(r1),
developanalternateproofofHelmholtz’stheorem,showingthat Vmaybewritten
as
V=−∇ϕ+∇×A,
where
A=∇×P,
and
ϕ=∇·P.
(b) SolvingthevectorPoissonequation,wefind
P(r1)=1
4πintegraldisplay
VV(r2)
r12dτ2.
Showthatthissolutionsubstitutedinto ϕandAofpart(a)leadstotheexpressions
givenfor ϕandAinSection1.16.
35Remember, c=∇×Vis known.
1.16 Additional Readings 101
AdditionalReadings
Borisenko,A.I.,andI.E.Taropov, VectorandTensorAnalysiswithApplications .EnglewoodCliffs,NJ:Prentice-
Hall(1968). Reprinted, Dover (1980).
Davis,H.F.,andA.D.Snider, Introduction to VectorAnalysis , 7th ed.Boston: Allyn & Bacon (1995).
Kellogg, O. D., Foundations of Potential Theory . New York: Dover (1953). Originally published (1929). The
classictext on potential theory.
Lewis,P.E.,andJ.P.Ward, VectorAnalysisforEngineersandScientists .Reading,MA:Addison-Wesley(1989).
Marion, J. B., Principles of Vector Analysis . New York: Academic Press (1965). A moderately advanced presen-
tation of vector analysis oriented toward tensor analysis. Rotations and other transformations are described
with the appropriate matrices.
Spiegel, M.R., VectorAnalysis . NewYork: McGraw-Hill (1989).
Tai, C.-T., Generalized VectorandDyadic Analysis . Oxford: Oxford University Press (1996).
Wrede,R.C., IntroductiontoVectorandTensorAnalysis .NewYork:Wiley(1963).Reprinted,NewYork:Dover
(1972). Fine historical introduction. Excellent discussion of differentiation of vectors and applications to me-
chanics.
This page intentionally left blank
CHAPTER 2
VECTOR ANALYSIS IN
CURVED COORDINATES
ANDTENSORS
InChapter1werestrictedourselvesalmostcompletelytorectangularorCartesiancoordi-
natesystems.ACartesiancoordinatesystemofferstheuniqueadvantagethatallthreeunit
vectors,ˆx,ˆy,andˆz,areconstantindirectionaswellasinmagnitude.Wedidintroducethe
radial distance r, but even this was treated as a function of x,y, andz. Unfortunately, not
allphysicalproblemsarewelladaptedtoasolutioninCartesiancoordinates.Forinstance,
if we have a central force problem, F=ˆrF(r), such as gravitational or electrostatic force,
Cartesiancoordinatesmaybeunusuallyinappropriate.Suchaproblemdemandstheuseof
a coordinate system in which the radial distance is taken to be one of the coordinates, that
is, sphericalpolarcoordinates.
The point is that the coordinate system should be chosen to fit the problem, to exploit
anyconstraintorsymmetrypresentinit.Thenitislikelytobemorereadilysolublethanif
wehadforceditintoaCartesianframework.
Naturally, there is a price that must be paid for the use of a non-Cartesian coordinate
system. We have not yet written expressions for gradient, divergence, or curl in any of the
non-Cartesiancoordinatesystems.SuchexpressionsaredevelopedingeneralforminSec-
tion 2.2. First, we develop a system of curvilinear coordinates, a general system that may
be specialized to any of the particular systems of interest. We shall specialize to circular
cylindricalcoordinatesinSection2.4andtosphericalpolarcoordinatesinSection2.5.
2.1 O RTHOGONAL COORDINATES IN R3
In Cartesian coordinates we deal with three mutually perpendicular families of planes:
x=constant, y=constant,and z=constant.Imaginethatwesuperimposeonthissystem
103
104 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
three other families of surfaces qi(x,y,z), i=1,2,3. The surfaces of any one family qi
neednotbeparalleltoeachotherandtheyneednotbeplanes.Ifthisisdifficulttovisualize,
the figure of a specific coordinate system, such as Fig. 2.3, may be helpful. The three new
families of surfaces need not be mutually perpendicular, but for simplicity we impose this
condition(Eq.(2.7))becauseorthogonalcoordinatesarecommoninphysicalapplications.
Thisorthogonalityhasmanyadvantages:OrthogonalcoordinatesarealmostlikeCartesian
coordinateswhereinfinitesimalareasandvolumesareproductsofcoordinatedifferentials.
Inthissectionwedevelopthegeneralformalismoforthogonalcoordinates,derivefrom
thegeometrythecoordinatedifferentials,andusethemforline,area,andvolumeelements
inmultipleintegralsandvectoroperators.Wemaydescribeanypoint (x,y,z)astheinter-
section of three planes in Cartesian coordinates or as the intersection of the three surfaces
that form our new, curvilinear coordinates. Describing the curvilinear coordinate surfaces
byq1=constant, q2=constant, q3=constant, we may identify our point by (q1,q2,q3)
aswellasby (x,y,z):
Generalcurvilinearcoordinates
q1,q2,q3Circularcylindricalcoordinates
ρ,ϕ,z
x=x(q1,q2,q3)
y=y(q1,q2,q3)
z=z(q1,q2,q3)−∞<x=ρcosϕ<∞
−∞<y=ρsinϕ<∞
−∞<z=z<∞(2.1)
specifying x,y,zintermsof q1,q2,q3andtheinverserelations
q1=q1(x,y,z) 0/lessorequalslantρ=parenleftbig
x2+y2parenrightbig1/2<∞
q2=q2(x,y,z) 0/lessorequalslantϕ=arctan(y/x)<2π
q3=q3(x,y,z)−∞<z=z<∞.(2.2)
As a specific illustration of the general, abstract q1,q2,q3, the transformation equations
for circular cylindricalcoordinates (Section 2.4) are included in Eqs. (2.1) and (2.2). With
each family of surfaces qi=constant, we can associate a unit vector ˆqinormal to the
surfaceqi=constant and in the direction of increasing qi. In general, these unit vectors
willdependonthepositioninspace.Thenavector Vmaybewritten
V=ˆq1V1+ˆq2V2+ˆq3V3, (2.3)
butthecoordinateorpositionvectoris differentingeneral,
r/negationslash=ˆq1q1+ˆq2q2+ˆq3q3,
as the special cases r=rˆrfor spherical polar coordinates and r=ρˆρ+zˆzfor cylindri-
cal coordinates demonstrate. The ˆqiare normalized to ˆq2
i=1 and form a right-handed
coordinatesystemwithvolume ˆq1·(ˆq2׈q3)>0.
Differentiationof xinEqs. (2.1) leadstothetotalvariationordifferential
dx=∂x
∂q1dq1+∂x
∂q2dq2+∂x
∂q3dq3, (2.4)
and similarly for differentiation of yandz. In vector notation dr=summationtext
i∂r
∂qidqi.From
the Pythagorean theorem in Cartesian coordinates the square of the distance between two
neighboringpointsis
ds2=dx2+dy2+dz2.
2.1 Orthogonal Coordinates in R3105
Substituting drshows that in our curvilinear coordinate space the square of the distance
elementcanbewrittenas aquadraticforminthedifferentials dqi:
ds2=dr·dr=dr2=summationdisplay
ij∂r
∂qi·∂r
∂qjdqidqj
=g11dq2
1+g12dq1dq2+g13dq1dq3
+g21dq2dq1+g22dq2
2+g23dq2dq3
+g31dq3dq1+g32dq3dq2+g33dq2
3
=summationdisplay
ijgijdqidqj, (2.5)
where nonzero mixed terms dqidqjwithi/negationslash=jsignal that these coordinates are not or-
thogonal, that is, that the tangential directions ˆqiare not mutually orthogonal. Spaces for
whichEq.(2.5) is alegitimateexpressionarecalled metricorRiemannian .
WritingEq. (2.5) moreexplicitly,weseethat
gij(q1,q2,q3)=∂x
∂qi∂x
∂qj+∂y
∂qi∂y
∂qj+∂z
∂qi∂z
∂qj=∂r
∂qi·∂r
∂qj(2.6)
arescalarproductsofthe tangentvectors∂r
∂qitothecurves rforqj=const.,j/negationslash=i.These
coefficient functions gij, which we now proceed to investigate, may be viewed as speci-
fying the nature of the coordinate system (q1,q2,q3). Collectively these coefficients are
referred to as the metricand in Section 2.10 will be shown to form a second-rank sym-
metric tensor.1In general relativity the metric components are determined by the proper-
ties of matter; that is, the gijare solutions of Einstein’s field equations with the energy–
momentum tensor as driving term; this may be articulated as “geometry is merged with
physics.”
At usual we limit ourselves to orthogonal (mutually perpendicular surfaces) coordinate
systems,whichmeans(seeExercise2.1.1)2
gij=0,i/negationslash=j, (2.7)
andˆqi·ˆqj=δij. (Nonorthogonal coordinate systems are considered in some detail in
Sections2.10and2.11intheframeworkoftensoranalysis.)Now,tosimplifythenotation,
wewrite gii=h2
i>0,so
ds2=(h1dq1)2+(h2dq2)2+(h3dq3)2=summationdisplay
i(hidqi)2. (2.8)
1The tensor nature of the set of gij’s follows from the quotient rule (Section 2.8). Then the tensor transformation law yields
Eq.(2.5).
2In relativistic cosmology the nondiagonal elements of the metric gijare usually set equal to zero as a consequence of physical
assumptions such asno rotation, as for dϕdt,dθdt .
106 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
The specific orthogonal coordinate systems are described in subsequent sections by spec-
ifying these (positive) scale factors h1,h2, andh3. Conversely, the scale factors may be
convenientlyidentifiedbytherelation
dsi=hidqi,∂r
∂qi=hiˆqi (2.9)
for any given dqi, holding all other qconstant. Here, dsiis a differential length along the
directionˆqi.Notethatthethreecurvilinearcoordinates q1,q2,q3neednotbelengths.The
scalefactors himaydependon qandtheymayhavedimensions.The producthidqimust
haveadimensionof length.Thedifferentialdistancevector drmaybewritten
dr=h1dq1ˆq1+h2dq2ˆq2+h3dq3ˆq3=summationdisplay
ihidqiˆqi.
Usingthiscurvilinearcomponentform,wefindthatalineintegralbecomes
integraldisplay
V·dr=summationdisplay
iintegraldisplay
Vihidqi.
FromEqs. (2.9) wemayimmediatelydeveloptheareaandvolumeelements
dσij=dsidsj=hihjdqidqj (2.10)
and
dτ=ds1ds2ds3=h1h2h3dq1dq2dq3. (2.11)
The expressions in Eqs. (2.10) and (2.11) agree, of course, with the results of using
the transformation equations, Eq. (2.1), and Jacobians (described shortly; see also Exer-
cise2.1.5).
FromEq. (2.10)anareaelementmaybeexpanded:
dσ=ds2ds3ˆq1+ds3ds1ˆq2+ds1ds2ˆq3
=h2h3dq2dq3ˆq1+h3h1dq3dq1ˆq2
+h1h2dq1dq2ˆq3.
Asurfaceintegralbecomes
integraldisplay
V·dσ=integraldisplay
V1h2h3dq2dq3+integraldisplay
V2h3h1dq3dq1
+integraldisplay
V3h1h2dq1dq2.
(Examplesof suchlineandsurfaceintegralsappearinSections2.4and2.5.)
2.1 Orthogonal Coordinates in R3107
In anticipation of the new forms of equations for vector calculus that appear in the
next section, let us emphasize that vector algebrais the same in orthogonal curvilinear
coordinatesasinCartesiancoordinates.Specifically,for thedotproduct,
A·B=summationdisplay
ikAiˆqi·ˆqkBk=summationdisplay
ikAiBkδik
=summationdisplay
iAiBi=A1B1+A2B2+A3B3, (2.12)
wherethesubscriptsindicatecurvilinearcomponents.Forthecross product,
A×B=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆq1ˆq2ˆq3
A1A2A3
B1B2B3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (2.13)
asinEq. (1.40).
Previously, we specialized to locally rectangular coordinates that are adapted to special
symmetries. Let us now briefly look at the more general case, where the coordinates are
not necessarily orthogonal. Surface and volume elements are part of multiple integrals,
which are common in physical applications, such as center of mass determinations and
moments of inertia. Typically, we choose coordinates according to the symmetry of the
particular problem. In Chapter 1 we used Gauss’ theorem to transform a volume integral
into a surface integral and Stokes’ theorem to transform a surface integral into a line in-
tegral. For orthogonal coordinates, the surface and volume elements are simply products
of the line elements hidqi(see Eqs. (2.10) and (2.11)). For the general case, we use the
geometric meaning of ∂r/∂qiin Eq. (2.5) as tangent vectors. We start with the Cartesian
surface element dxdy, which becomes an infinitesimal rectangle in the new coordinates
q1,q2formedbythetwoincrementalvectors
dr1=r(q1+dq1,q2)−r(q1,q2)=∂r
∂q1dq1,
dr2=r(q1,q2+dq2)−r(q1,q2)=∂r
∂q2dq2, (2.14)
whoseareais the z-componentoftheircross product,or
dxdy=dr1×dr2vextendsinglevextendsingle
z=bracketleftbigg∂x
∂q1∂y
∂q2−∂x
∂q2∂y
∂q1bracketrightbigg
dq1dq2
=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂q1∂x
∂q2
∂y
∂q1∂y
∂q2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingledq1dq2. (2.15)
Thetransformationcoefficientindeterminantformis calledthe Jacobian .
Similarly,thevolumeelement dxdydz becomesthetriplescalarproductofthethreein-
finitesimaldisplacementvectors dri=dqi∂r
∂qialongthe qidirectionsˆqi,which,according
108 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
toSection1.5, takesontheform
dxdydz=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂q1∂x
∂q2∂x
∂q3
∂y
∂q1∂y
∂q2∂y
∂q3
∂z
∂q1∂z
∂q2∂z
∂q3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingledq1dq2dq3. (2.16)
HerethedeterminantisalsocalledtheJacobian,andsooninhigherdimensions.
For orthogonal coordinates the Jacobians simplify to products of the orthogonal vec-
tors in Eq. (2.9). It follows that they are just products of the hi; for example, the volume
Jacobianbecomes
h1h2h3(ˆq1׈q2)·ˆq3=h1h2h3,
andso on.
Example 2.1.1 JACOBIANS FOR POLAR COORDINATES
LetusillustratethetransformationoftheCartesiantwo-dimensionalvolumeelement dxdy
topolarcoordinates ρ,ϕ, withx=ρcosϕ, y=ρsinϕ. (SeealsoSection2.4.) Here,
dxdy=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂ρ∂x
∂ϕ
∂y
∂ρ∂y
∂ϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsingledρdϕ=vextendsinglevextendsinglevextendsinglevextendsinglecosϕ−ρsinϕ
sinϕρcosϕvextendsinglevextendsinglevextendsinglevextendsingledρdϕ=ρdρdϕ.
Similarly, in spherical coordinates (see Section 2.5) we get, from x=rsinθcosϕ,y=
rsinθsinϕ,z=rcosθ,theJacobian
J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂r∂x
∂θ∂x
∂ϕ
∂y
∂r∂y
∂θ∂y
∂ϕ
∂z
∂r∂z
∂θ∂z
∂ϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglesinθcosϕrcosθcosϕ−rsinθsinϕ
sinθsinϕrcosθsinϕrsinθcosϕ
cosθ−rsinθ 0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle
=cosθvextendsinglevextendsinglevextendsinglevextendsinglercosθcosϕ−rsinθsinϕ
rcosθsinϕrsinθcosϕvextendsinglevextendsinglevextendsinglevextendsingle+rsinθvextendsinglevextendsinglevextendsinglevextendsinglesinθcosϕ−rsinθsinϕ
sinθsinϕrsinθcosϕvextendsinglevextendsinglevextendsinglevextendsingle
=r2parenleftbig
cos2θsinθ+sin3θparenrightbig
=r2sinθ
by expanding the determinant along the third line. Hence the volume element becomes
dxdydz=r2drsinθdθdϕ.Thevolumeintegralcanbewrittenas
integraldisplay
f(x,y,z)dxdydz =integraldisplay
fparenleftbig
x(r,θ,ϕ),y(r,θ,ϕ),z(r,θ,ϕ)parenrightbig
r2drsinθdθdϕ./squaresolid
Insummary,wehavedevelopedthegeneralformalismforvectoranalysisinorthogonal
curvilinear coordinates in R3. For most applications, locally orthogonal coordinates can
bechosenforwhichsurfaceandvolumeelementsinmultipleintegralsareproductsofline
elements.For thegeneralnonorthogonalcase, Jacobiandeterminantsapply.
2.1 Orthogonal Coordinates in R3109
Exercises
2.1.1 Show that limiting our attention to orthogonal coordinate systems implies that gij=0
fori/negationslash=j(Eq. (2.7)).
Hint.Constructatrianglewithsides ds1,ds2,andds2.Equation(2.9)mustholdregard-
less of whether gij=0. Then compare ds2from Eq. (2.5) with a calculation using the
lawof cosines.Showthat cos θ12=g12/√g11g22.
2.1.2 In the spherical polar coordinate system, q1=r,q2=θ,q3=ϕ. The transformation
equationscorrespondingtoEq. (2.1) are
x=rsinθcosϕ, y=rsinθsinϕ, z=rcosθ.
(a) Calculatethesphericalpolarcoordinatescalefactors: hr,hθ,andhϕ.
(b) Checkyourcalculatedscalefactors bytherelation dsi=hidqi.
2.1.3 Theu-,v-,z-coordinate system frequently used in electrostatics and in hydrodynamics
isdefinedby
xy=u, x2−y2=v, z=z.
Thisu-,v-,z-systemis orthogonal.
(a) In words, describe briefly the nature of each of the three families of coordinate
surfaces.
(b) Sketchthesysteminthe xy-planeshowingtheintersectionsofsurfacesofconstant
uandsurfaces of constant vwiththexy-plane.
(c) Indicatethedirectionsoftheunitvector ˆuandˆvinallfour quadrants.
(d) Finally,isthis u-,v-,z-systemright-handed (ˆu׈v=+ˆz)orleft-handed (ˆu׈v=
−ˆz)?
2.1.4 Theellipticcylindricalcoordinatesystemconsistsofthreefamiliesofsurfaces:
1)x2
a2cosh2u+y2
a2sinh2u=1;2)x2
a2cos2v−y2
a2sin2v=1;3)z=z.
Sketch the coordinate surfaces u=constant and v=constant as they intersect the first
quadrant of the xy-plane. Show the unit vectors ˆuandˆv. The range of uis 0/lessorequalslantu<∞.
Therangeof vis 0/lessorequalslantv/lessorequalslant2π.
2.1.5 A two-dimensional orthogonal system is described by the coordinates q1andq2. Show
thattheJacobian
Jparenleftbiggx,y
q1,q2parenrightbigg
≡∂(x,y)
∂(q1,q2)≡∂x
∂q1∂y
∂q2−∂x
∂q2∂y
∂q1=h1h2
isinagreementwithEq. (2.10).
Hint.It’seasiertowork withthesquareofeachsideofthisequation.
110 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
2.1.6 InMinkowskispacewedefine x1=x,x2=y,x3=z,andx0=ct.Thisisdonesothat
the metric interval becomes ds2=dx2
0–dx2
1–dx2
2–dx2
3(withc=velocity of light).
ShowthatthemetricinMinkowskispaceis
(gij)=
1 000
0−10 0
00−10
00 0 −1
.
WeuseMinkowskispaceinSections4.5and4.6fordescribingLorentztransformations.
2.2 D IFFERENTIAL VECTOR OPERATORS
Wereturntoourrestrictiontoorthogonalcoordinatesystems.
Gradient
Thestartingpointfordevelopingthegradient,divergence,andcurloperatorsincurvilinear
coordinates is the geometric interpretation of the gradient as the vector having the mag-
nitude and direction of the maximum space rate of change (compare Section 1.6). From
this interpretation the component of ∇ψ(q1,q2,q3)in the direction normal to the family
ofsurfaces q1=constantis givenby3
ˆq1·∇ψ=∇ψ|1=∂ψ
∂s1=1
h1∂ψ
∂q1, (2.17)
since this is the rate of change of ψfor varying q1, holding q2andq3fixed. The quantity
ds1is a differential length in the direction of increasing q1(compare Eqs. (2.9)). In Sec-
tion 2.1 we introduced a unit vector ˆq1to indicate this direction. By repeating Eq. (2.17)
forq2andagainfor q3andaddingvectorially,wesee thatthegradientbecomes
∇ψ(q1,q2,q3)=ˆq1∂ψ
∂s1+ˆq2∂ψ
∂s2+ˆq3∂ψ
∂s3
=ˆq11
h1∂ψ
∂q1+ˆq21
h2∂ψ
∂q2+ˆq31
h3∂ψ
∂q3
=summationdisplay
iˆqi1
hi∂ψ
∂qi. (2.18)
Exercise2.2.4offersamathematicalalternativeindependentofthisphysicalinterpretation
ofthegradient.Thetotalvariationofafunction,
dψ=∇ψ·dr=summationdisplay
i1
hi∂ψ
∂qidsi=summationdisplay
i∂ψ
∂qidqi
isconsistentwithEq. (2.18), ofcourse.
3Heretheuseof ϕtolabelafunctionisavoidedbecauseitisconventionaltousethissymboltodenoteanazimuthalcoordinate.
2.2 Differential Vector Operators 111
Divergence
Thedivergenceoperatormaybeobtainedfromtheseconddefinition(Eq.(1.98))ofChap-
ter1or equivalentlyfromGauss’theorem,Section1.11.Letususe Eq.(1.98),
∇·V(q1,q2,q3)=limintegraltext
dτ→0integraltext
V·dσintegraltext
dτ, (2.19)
with a differential volume h1h2h3dq1dq2dq3(Fig. 2.1). Note that the positive directions
havebeenchosenso that (ˆq1,ˆq2,ˆq3)formaright-handedset, ˆq1׈q2=ˆq3.
Thedifferenceofareaintegralsforthetwofaces q1=constant isgivenby
bracketleftbigg
V1h2h3+∂
∂q1(V1h2h3)dq1bracketrightbigg
dq2dq3−V1h2h3dq2dq3
=∂
∂q1(V1h2h3)dq1dq2dq3, (2.20)
exactly as in Sections 1.7 and 1.10.4Here,Vi=V·ˆqiis the projection of Vonto the
ˆqi-direction.Addinginthesimilarresults for theothertwopairs ofsurfaces, weobtain
integraldisplay
V(q1,q2,q3)·dσ
=bracketleftbigg∂
∂q1(V1h2h3)+∂
∂q2(V2h3h1)+∂
∂q3(V3h1h2)bracketrightbigg
dq1dq2dq3.
FIGURE 2.1Curvilinearvolumeelement.
4Sin cewetak eth elim it dq1,dq2,dq3→0,the second- and higher-order derivatives will drop out.
112 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Now,usingEq. (2.19), divisionbyourdifferentialvolumeyields
∇·V(q1,q2,q3)=1
h1h2h3bracketleftbigg∂
∂q1(V1h2h3)+∂
∂q2(V2h3h1)+∂
∂q3(V3h1h2)bracketrightbigg
.(2.21)
We may obtain the Laplacian by combining Eqs. (2.18) and (2.21), using V=
∇ψ(q1,q2,q3). Thisleadsto
∇·∇ψ(q1,q2,q3)
=1
h1h2h3bracketleftbigg∂
∂q1parenleftbiggh2h3
h1∂ψ
∂q1parenrightbigg
+∂
∂q2parenleftbiggh3h1
h2∂ψ
∂q2parenrightbigg
+∂
∂q3parenleftbiggh1h2
h3∂ψ
∂q3parenrightbiggbracketrightbigg
.(2.22)
Curl
Finally, to develop ∇×V, let us apply Stokes’ theorem (Section 1.12) and, as with the
divergence, take the limit as the surface area becomes vanishingly small. Working on one
component at a time, we consider a differential surface element in the curvilinear surface
q1=constant. From
integraldisplay
s∇×V·dσ=ˆq1·(∇×V)h2h3dq2dq3 (2.23)
(meanvaluetheoremofintegralcalculus),Stokes’theoremyields
ˆq1·(∇×V)h2h3dq2dq3=contintegraldisplay
V·dr, (2.24)
with the line integral lying in the surface q1=constant. Following the loop (1, 2, 3, 4) of
Fig.2.2,
contintegraldisplay
V(q1,q2,q3)·dr=V2h2dq2+bracketleftbigg
V3h3+∂
∂q2(V3h3)dq2bracketrightbigg
dq3
−bracketleftbigg
V2h2+∂
∂q3(V2h2)dq3bracketrightbigg
dq2−V3h3dq3
=bracketleftbigg∂
∂q2(h3V3)−∂
∂q3(h2V2)bracketrightbigg
dq2dq3. (2.25)
We pick up a positive sign when going in the positive direction on parts 1 and 2 and
a negative sign on parts 3 and 4 because here we are going in the negative direction.
(Higher-order terms in Maclaurin or Taylor expansions have been omitted. They will van-
ishinthelimitas thesurface becomesvanishinglysmall( dq2→0,dq3→0).)
From Eq. (2.24),
∇×V|1=1
h2h3bracketleftbigg∂
∂q2(h3V3)−∂
∂q3(h2V2)bracketrightbigg
. (2.26)
2.2 Differential Vector Operators 113
FIGURE 2.2Curvilinearsurface elementwith q1=constant.
The remaining two components of ∇×Vmay be picked up by cyclic permutation of the
indices.As inChapter1, itisoftenconvenienttowritethecurlindeterminantform:
∇×V=1
h1h2h3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆq1h1ˆq2h2ˆq3h3
∂
∂q1∂
∂q2∂
∂q3
h1V1h2V2h3V3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.27)
Rememberthat,becauseofthepresenceofthedifferentialoperators,thisdeterminantmust
be expanded from the top down. Note that this equation is notidentical with the form for
the cross product of two vectors, Eq. (2.13). ∇is not an ordinary vector; it is a vector
operator.
OurgeometricinterpretationofthegradientandtheuseofGauss’andStokes’theorems
(or integral definitions of divergence and curl) have enabled us to obtain these quantities
without having to differentiate the unit vectors ˆqi. There exist alternate ways to deter-
minegrad,div,andcurlbasedondirectdifferentiationofthe ˆqi.Oneapproachresolvesthe
ˆqiofaspecificcoordinatesystemintoitsCartesiancomponents(Exercises2.4.1and2.5.1)
anddifferentiatesthisCartesianform(Exercises2.4.3and2.5.2).Thepointhereisthatthe
derivatives of the Cartesian ˆx,ˆy, andˆzvanish sinceˆx,ˆy, andˆzare constant in direction
as well as in magnitude. A second approach [L. J. Kijewski, A m .J .P h y s . 33: 816 (1965)]
assumestheequalityof ∂2r/∂qi∂qjand∂2r/∂qj∂qianddevelopsthederivativesof ˆqiin
ageneralcurvilinearform. Exercises2.2.3and2.2.4arebasedonthismethod.
Exercises
2.2.1 Developargumentstoshowthatdotandcrossproducts(notinvolving ∇)inorthogonal
curvilinearcoordinatesin R3proceed,asinCartesiancoordinates, withnoinvolvement
ofscalefactors .
2.2.2 Withˆq1aunitvectorinthedirectionofincreasing q1, showthat
114 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
(a)∇·ˆq1=1
h1h2h3∂(h2h3)
∂q1
(b)∇׈q1=1
h1bracketleftbigg
ˆq21
h3∂h1
∂q3−ˆq31
h2∂h1
∂q2bracketrightbigg
.
Note that even though ˆq1is a unit vector, its divergence and curl do not necessarily
vanish.
2.2.3 Showthattheorthogonalunitvectors ˆqjmaybedefinedby
ˆqi=1
hi∂r
∂qi. (a)
In particular, show that ˆqi·ˆqi=1 leads to an expression for hiin agreement with
Eqs. (2.9).
Equation(a) maybetakenasastartingpointfor deriving
∂ˆqi
∂qj=ˆqj1
hi∂hj
∂qi,i/negationslash=j
and
∂ˆqi
∂qi=−summationdisplay
j/negationslash=iˆqj1
hj∂hi
∂qj.
2.2.4 Derive
∇ψ=ˆq11
h1∂ψ
∂q1+ˆq21
h2∂ψ
∂q2+ˆq31
h3∂ψ
∂q3
bydirectapplicationof Eq.(1.97),
∇ψ=limintegraltext
dτ→0integraltext
ψdσintegraltext
dτ.
Hint.Evaluation of the surface integral will lead to terms like (h1h2h3)−1(∂/∂q1)×
(ˆq1h2h3).TheresultslistedinExercise2.2.3willbehelpful.Cancellationofunwanted
termsoccurswhenthecontributionsof allthreepairs ofsurfaces areaddedtogether.
2.3 S PECIAL COORDINATE SYSTEMS :INTRODUCTION
There are at least 11 coordinate systems in which the three-dimensional Helmholtz equa-
tion can be separated into three ordinary differential equations. Some of these coordinate
systems have achieved prominence in the historical development of quantum mechanics.
Othersystems,suchasbipolarcoordinates,satisfy specialneeds.Partlybecausetheneeds
are rather infrequent but mostly because the development of computers and efficient pro-
gramming techniques reduce the need for these coordinate systems, the discussion in this
chapter is limited to (1) Cartesian coordinates, (2) spherical polar coordinates, and (3) cir-
cular cylindrical coordinates. Specifications and details of the other coordinate systems
willbefoundinthefirsttwoeditionsofthisworkandinAdditionalReadingsattheendof
thischapter(Morse andFeshbach,MargenauandMurphy).
2.4 Circular Cylinder Coordinates 115
2.4 C IRCULAR CYLINDER COORDINATES
In the circular cylindrical coordinate system the three curvilinear coordinates (q1,q2,q3)
are relabeled (ρ,ϕ,z).W ea r eu s i n g ρfor the perpendicular distance from the z-axis and
savingrfor thedistancefromtheorigin.Thelimitson ρ,ϕandzare
0/lessorequalslantρ<∞,0/lessorequalslantϕ/lessorequalslant2π,and−∞<z<∞.
Forρ=0,ϕis notwelldefined.Thecoordinatesurfaces, showninFig.2.3,are:
1. Rightcircularcylindershavingthe z-axis asacommonaxis,
ρ=parenleftbig
x2+y2parenrightbig1/2=constant.
2. Half-planesthroughthe z-axis,
ϕ=tan−1parenleftbiggy
xparenrightbigg
=constant.
3. Planesparalleltothe xy-plane,as intheCartesiansystem,
z=constant.
FIGURE 2.3Circularcylindercoordinates.
116 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
FIGURE 2.4Circularcylindrical
coordinateunitvectors.
Inverting the preceding equations for ρandϕ(or going directly to Fig. 2.3), we obtain
thetransformationrelations
x=ρcosϕ, y=ρsinϕ, z=z. (2.28)
Thez-axis remains unchanged. This is essentially a two-dimensional curvilinear system
withaCartesian z-axisaddedontoform athree-dimensionalsystem.
AccordingtoEq.(2.5) or fromthelengthelements dsi, thescalefactors are
h1=hρ=1,h 2=hϕ=ρ, h 3=hz=1. (2.29)
Theunitvectors ˆq1,ˆq2,ˆq3arerelabeled (ˆρ,ˆϕ,ˆz),asinFig.2.4.Theunitvector ˆρisnormal
to the cylindrical surface, pointing in the direction of increasing radius ρ. The unit vector
ˆϕis tangential to the cylindrical surface, perpendicular to the half plane ϕ=constant and
pointinginthedirectionofincreasingazimuthangle ϕ.Thethirdunitvector, ˆz,istheusual
Cartesianunitvector.Theyaremutuallyorthogonal,
ˆρ·ˆϕ=ˆϕ·ˆz=ˆz·ˆρ=0,
andthecoordinatevectoranda generalvector Vareexpressedas
r=ˆρρ+ˆzz,V=ˆρVρ+ˆϕVϕ+ˆzVz.
Adifferentialdisplacement drmaybewritten
dr=ˆρdsρ+ˆϕdsϕ+ˆzdz
=ˆρdρ+ˆϕρdϕ+ˆzdz. (2.30)
Example 2.4.1 AREALAW FOR PLANETARY MOTION
FirstwederiveKepler’slawincylindricalcoordinates,sayingthattheradiusvectorsweeps
outequalareasinequaltime,fromangularmomentumconservation.
2.4 Circular Cylinder Coordinates 117
Weconsiderthesunattheoriginasasourceofthe centralgravitationalforce F=f(r)ˆr.
Then the orbital angular momentum L=mr×vof a planet of mass mand velocity vis
conserved,becausethetorque
dL
dt=mdr
dt×dr
dt+r×mdv
dt=r×F=f(r)
rr×r=0.
HenceL=const. Now we can choose the z-axis to lie along the direction of the orbital
angularmomentumvector, L=Lˆz,andworkincylindricalcoordinates r=(ρ,ϕ,z)=ρˆρ
withz=0.Theplanetmovesinthe xy-planebecause randvareperpendicularto L.Thus,
weexpanditsvelocityasfollows:
v=dr
dt=˙ρˆρ+ρdˆρ
dt.
From
ˆρ=(cosϕ,sinϕ),∂ˆρ
dϕ=(−sinϕ,cosϕ)=ˆϕ,
wefindthatdˆρ
dt=dˆρ
dϕdϕ
dt=˙ϕˆϕusingthechainrule,so v=˙ρˆρ+ρdˆρ
dt=˙ρˆρ+ρ˙ϕˆϕ.When
wesubstitutetheexpansionsof ˆρandvinpolarcoordinates,weobtain
L=mρ×v=mρ(ρ˙ϕ)(ˆρ׈ϕ)=mρ2˙ϕˆz=constant.
The triangular area swept by the radius vector ρin the time dt(area law), when inte-
gratedoveronerevolution,is givenby
A=1
2integraldisplay
ρ(ρdϕ)=1
2integraldisplay
ρ2˙ϕdt=L
2mintegraldisplay
dt=Lτ
2m, (2.31)
ifwesubstitute mρ2˙ϕ=L=const.Here τistheperiod,thatis,thetimeforonerevolution
oftheplanetinitsorbit.
Kepler’s first law says that the orbit is an ellipse. Now we derive the orbit equation
ρ(ϕ)of the ellipse in polar coordinates, where in Fig. 2.5 the sun is at one focus, which is
the origin of our cylindrical coordinates. From the geometrical construction of the ellipse
we know that ρ′+ρ=2a,whereais the major half-axis; we shall show that this is
equivalenttotheconventionalformoftheellipseequation.Thedistancebetweenbothfoci
is0<2aǫ<2a,where0 <ǫ<1iscalledtheeccentricityoftheellipse.Foracircle ǫ=0
because both foci coincide with the center. There is an angle, as shown in Fig. 2.5, where
the distances ρ′=ρ=aare equal, and Pythagoras’ theorem applied to this right triangle
FIGURE 2.5Ellipseinpolarcoordinates.
118 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
givesb2+a2ǫ2=a2.As a result,√
1−ǫ2=b/ais the ratio of the minor half-axis ( b)t o
themajorhalf-axis, a.
Now consider the triangle with the sides labeled by ρ′,ρ,2aǫin Fig. 2.5 and angle
oppositeρ′equaltoπ−ϕ.Then,applyingthelawofcosines, gives
ρ′2=ρ2+4a2ǫ2+4ρaǫcosϕ.
Nowsubstituting ρ′=2a−ρ,canceling ρ2onbothsides anddividingby 4 ayields
ρ(1+ǫcosϕ)=aparenleftbig
1−ǫ2parenrightbig
≡p, (2.32)
theKeplerorbitequationinpolarcoordinates .
Alternatively, we revert to Cartesian coordinates to find, from Eq. (2.32) with x=
ρcosϕ, that
ρ2=x2+y2=(p−xǫ)2=p2+x2ǫ2−2pxǫ,
sothefamiliarellipseequationinCartesiancoordinates,
parenleftbig
1−ǫ2parenrightbigparenleftbigg
x+pǫ
1−ǫ2parenrightbigg2
+y2=p2+p2ǫ2
1−ǫ2=p2
1−ǫ2,
obtains.If wecomparethisresultwiththestandardformof theellipse,
(x−x0)2
a2+y2
b2=1,
weconfirmthat
b=p√
1−ǫ2=aradicalbig
1−ǫ2,a=p
1−ǫ2,
andthatthedistance x0betweenthecenterandfocusis aǫ,assho wninFig.2. 5. /squaresolid
The differential operations involving ∇follow from Eqs. (2.18), (2.21), (2.22), and
(2.27):
∇ψ(ρ,ϕ,z)=ˆρ∂ψ
∂ρ+ˆϕ1
ρ∂ψ
∂ϕ+ˆz∂ψ
∂z, (2.33)
∇·V=1
ρ∂
∂ρ(ρVρ)+1
ρ∂Vϕ
∂ϕ+∂Vz
∂z,(2.34)
∇2ψ=1
ρ∂
∂ρparenleftbigg
ρ∂ψ
∂ρparenrightbigg
+1
ρ2∂2ψ
∂ϕ2+∂2ψ
∂z2,(2.35)
∇×V=1
ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz
∂
∂ρ∂
∂ϕ∂
∂z
VρρVϕVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.36)
2.4 Circular Cylinder Coordinates 119
Finally, for problems such as circular wave guides and cylindrical cavity resonators the
vectorLaplacian ∇2Vresolvedincircularcylindricalcoordinatesis
∇2V|ρ=∇2Vρ−1
ρ2Vρ−2
ρ2∂Vϕ
∂ϕ,
∇2V|ϕ=∇2Vϕ−1
ρ2Vϕ+2
ρ2∂Vρ
∂ϕ, (2.37)
∇2V|z=∇2Vz,
whichfollowfromEq.(1.85).Thebasicreasonforthisparticularformofthe z-component
isthatthe z-axis isaCartesianaxis;thatis,
∇2(ˆρVρ+ˆϕVϕ+ˆzVz)=∇2(ˆρVρ+ˆϕVϕ)+ˆz∇2Vz
=ˆρf(Vρ,Vϕ)+ˆϕg(Vρ,Vϕ)+ˆz∇2Vz.
Finally,theoperator ∇2operatingonthe ˆρ,ˆϕunitvectorsstaysinthe ˆρˆϕ-plane.
Example 2.4.2 AN AVIER –STOKES TERM
TheNavier–Stokesequationsof hydrodynamicscontainanonlinearterm
∇×bracketleftbig
v×(∇×v)bracketrightbig
,
wherevisthefluidvelocity.Forfluidflowingthroughacylindricalpipeinthe z-direction,
v=ˆzv(ρ).
From Eq. (2.36),
∇×v=1
ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz
∂
∂ρ∂
∂ϕ∂
∂z
00 v(ρ)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−ˆϕ∂v
∂ρ
v×(∇×v)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρ ˆϕˆz
00 v
0−∂v
∂ρ0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=ˆρv(ρ)∂v
∂ρ.
Finally,
∇×parenleftbig
v×(∇×v)parenrightbig
=1
ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz
∂
∂ρ∂
∂ϕ∂
∂z
v∂v
∂ρ00vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0,
so,for thisparticularcase,thenonlineartermvanishes. /squaresolid
120 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Exercises
2.4.1 ResolvethecircularcylindricalunitvectorsintotheirCartesiancomponents(Fig.2.6).
ANS. ˆρ=ˆxcosϕ+ˆysinϕ,
ˆϕ=−ˆxsinϕ+ˆycosϕ,
ˆz=ˆz.
2.4.2 ResolvetheCartesianunitvectorsintotheircircularcylindricalcomponents(Fig.2.6).
ANS.ˆx=ˆρcosϕ−ˆϕsinϕ,
ˆy=ˆρsinϕ+ˆϕcosϕ,
ˆz=ˆz.
2.4.3 Fromtheresults ofExercise2.4.1showthat
∂ˆρ
∂ϕ=ˆϕ,∂ˆϕ
∂ϕ=−ˆρ
and that all other first derivatives of the circular cylindrical unit vectors with respect to
thecircularcylindricalcoordinatesvanish.
2.4.4 Compare ∇·V(Eq. (2.34)) withthegradientoperator
∇=ˆρ∂
∂ρ+ˆϕ1
ρ∂
∂ϕ+ˆz∂
∂z
(Eq. (2.33)) dotted into V. Note that the differential operators of ∇differentiate both
theunitvectorsandthecomponentsof V.
Hint. ˆϕ(1/ρ)(∂/∂ϕ)·ˆρVρbecomes ˆϕ·1
ρ∂
∂ϕ(ˆρVρ)anddoes notvanish.
2.4.5 (a) Showthat r=ˆρρ+ˆzz.
FIGURE 2.6Planepolarcoordinates.
2.4 Circular Cylinder Coordinates 121
(b) Workingentirelyincircularcylindricalcoordinates,showthat
∇·r=3 and ∇×r=0.
2.4.6 (a) Show that the parity operation (reflection through the origin) on a point (ρ,ϕ,z)
relativeto fixedx-,y-,z-axesconsistsof thetransformation
ρ→ρ, ϕ→ϕ±π, z→−z.
(b) Show that ˆρandˆϕhave odd parity (reversal of direction) and that ˆzhas even
parity.
Note.TheCartesianunitvectors ˆx,ˆy, andˆzremainconstant.
2.4.7 Arigidbodyisrotatingaboutafixedaxiswithaconstantangularvelocity ω.T ake ωto
liealongthe z-axis.Expressthepositionvector rincircularcylindricalcoordinatesand
usingcircularcylindricalcoordinates,
(a) calculate v=ω×r, (b) calculate ∇×v.
ANS.(a)v=ˆϕωρ,
(b)∇×v=2ω.
2.4.8 Find the circular cylindrical components of the velocity and acceleration of a moving
particle,
vρ=˙ρ, a ρ=¨ρ−ρ˙ϕ2,
vϕ=ρ˙ϕ, a ϕ=ρ¨ϕ+2˙ρ˙ϕ,
vz=˙z, a z=¨z.
Hint.
r(t)=ˆρ(t)ρ(t)+ˆzz(t)
=bracketleftbigˆxcosϕ(t)+ˆysinϕ(t)bracketrightbig
ρ(t)+ˆzz(t).
Note.˙ρ=dρ/dt,¨ρ=d2ρ/dt2, andso on.
2.4.9 SolveLaplace’sequation, ∇2ψ=0,incylindricalcoordinatesfor ψ=ψ(ρ).
ANS.ψ=klnρ
ρ0.
2.4.10 Inrightcircularcylindricalcoordinatesaparticularvectorfunctionisgivenby
V(ρ,ϕ)=ˆρVρ(ρ,ϕ)+ˆϕVϕ(ρ,ϕ).
Showthat ∇×Vhasonlya z-component.Notethatthisresultwillholdfor anyvector
confined to a surface q3=constant as long as the products h1V1andh2V2are each
independentof q3.
2.4.11 FortheflowofanincompressibleviscousfluidtheNavier–Stokesequationsleadto
−∇×parenleftbig
v×(∇×v)parenrightbig
=η
ρ0∇2(∇×v).
Hereηis the viscosity and ρ0is the density of the fluid. For axial flow in a cylindrical
pipewetakethevelocity vtobe
v=ˆzv(ρ).
122 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
FromExample2.4.2,
∇×parenleftbig
v×(∇×v)parenrightbig
=0
for thischoiceof v.
Showthat
∇2(∇×v)=0
leadstothedifferentialequation
1
ρd
dρparenleftbigg
ρd2v
dρ2parenrightbigg
−1
ρ2dv
dρ=0
andthatthisis satisfiedby
v=v0+a2ρ2.
2.4.12 A conducting wire along the z-axis carries a current I. The resulting magnetic vector
potentialis givenby
A=ˆzµI
2πlnparenleftbigg1
ρparenrightbigg
.
Showthatthemagneticinduction Bisgivenby
B=ˆϕµI
2πρ.
2.4.13 Aforceis describedby
F=−ˆxy
x2+y2+ˆyx
x2+y2.
(a) Express Fincircularcylindricalcoordinates.
Operatingentirelyincircularcylindricalcoordinatesfor(b) and(c),
(b) calculatethecurlof Fand
(c) calculatetheworkdoneby Fintraverstheunitcircleoncecounterclockwise.
(d) Howdoyoureconciletheresults of(b) and(c)?
2.4.14 A transverse electromagnetic wave (TEM) in a coaxial waveguide has an electric field
E=E(ρ,ϕ)ei(kz−ωt)andamagneticinductionfieldof B=B(ρ,ϕ)ei(kz−ωt).Sincethe
waveistransverse,neither EnorBhasazcomponent.Thetwofieldssatisfythe vector
Laplacianequation
∇2E(ρ,ϕ)=0
∇2B(ρ,ϕ)=0.
(a) Showthat E=ˆρE0(a/ρ)ei(kz−ωt)andB=ˆϕB0(a/ρ)ei(kz−ωt)aresolutions.Here
ais theradiusoftheinnerconductorand E0andB0areconstantamplitudes.
2.5 Spherical Polar Coordinates 123
(b) Assuming a vacuum inside the waveguide, verify that Maxwell’s equations are
satisfiedwith
B0/E0=k/ω=µ0ε0(ω/k)=1/c.
2.4.15 A calculation of the magnetohydrodynamic pinch effect involves the evaluation of
(B·∇)B.If themagneticinduction Bistakentobe B=ˆϕBϕ(ρ), showthat
(B·∇)B=−ˆρB2
ϕ/ρ.
2.4.16 The linear velocityof particles in a rigid body rotatingwith angular velocity ωis given
by
v=ˆϕρω.
Integratecontintegraltext
v·dλaroundacircleinthe xy-planeandverifythat
contintegraltext
v·dλ
area=∇×v|z.
2.4.17 Ap r o t o no fm a s s m, charge+e, and (asymptotic) momentum p=mvis incident on
a nucleus of charge +Zeat an impact parameter b. Determine the proton’s distance of
closestapproach.
2.5 S PHERICAL POLAR COORDINATES
Relabeling (q1,q2,q3)as(r,θ,ϕ), we see that the spherical polar coordinate system con-
sists ofthefollowing:
1. Concentricspheres centeredattheorigin,
r=parenleftbig
x2+y2+z2parenrightbig1/2=constant.
2. Rightcircularconescenteredonthe z-(polar) axis, verticesattheorigin,
θ=arccosz
(x2+y2+z2)1/2=constant.
3. Half-planesthroughthe z-(polar) axis,
ϕ=arctany
x=constant.
By our arbitrary choice of definitions of θ, the polar angle, and ϕ, the azimuth angle, the
z-axis is singled out for special treatment. The transformation equations corresponding to
Eq.(2.1) are
x=rsinθcosϕ, y=rsinθsinϕ, z=rcosθ, (2.38)
124 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
FIGURE 2.7Sphericalpolarcoordinatearea
elements.
measuring θfrom the positive z-axis and ϕin thexy-plane from the positive x-axis. The
ranges of values are 0 /lessorequalslantr<∞,0/lessorequalslantθ/lessorequalslantπ, and 0 /lessorequalslantϕ/lessorequalslant2π.A tr=0,θandϕare
undefined.FromdifferentiationofEq. (2.38),
h1=hr=1,
h2=hθ=r, (2.39)
h3=hϕ=rsinθ.
Thisgivesalineelement
dr=ˆrdr+ˆθrdθ+ˆϕrsinθdϕ,
so
ds2=dr·dr=dr2+r2dθ2+r2sin2θdϕ2,
the coordinates being obviously orthogonal. In this spherical coordinate system the area
element(for r=constant)is
dA=dσθϕ=r2sinθdθdϕ, (2.40)
the light, unshaded area in Fig. 2.7. Integrating over the azimuth ϕ, we find that the area
elementbecomesaringof width dθ,
dAθ=2πr2sinθdθ. (2.41)
Thisformwillappearrepeatedlyinproblemsinsphericalpolarcoordinateswithazimuthal
symmetry,suchasthescatteringofanunpolarizedbeamofparticles.Bydefinitionofsolid
radians,orsteradians,anelementofsolidangle d/Omega1isgivenby
d/Omega1=dA
r2=sinθdθdϕ. (2.42)
2.5 Spherical Polar Coordinates 125
FIGURE 2.8Sphericalpolarcoordinates.
Integratingovertheentiresphericalsurface, weobtain
integraldisplay
d/Omega1=4π.
FromEq. (2.11)thevolumeelementis
dτ=r2drsinθdθdϕ=r2drd/Omega1. (2.43)
ThesphericalpolarcoordinateunitvectorsareshowninFig.2.8.
Itmustbeemphasizedthat theunitvectors ˆr,ˆθ,and ˆϕvaryindirectionastheangles
θandϕvary.Specifically,the θandϕderivativesofthesesphericalpolarcoordinateunit
vectors do not vanish (Exercise 2.5.2). When differentiating vectors in spherical polar (or
in any non-Cartesian system), this variation of the unit vectors with position must not be
neglected.Intermsofthefixed-directionCartesianunitvectors ˆx,ˆyandˆz(cp.Eq.(2.38)),
ˆr=ˆxsinθcosϕ+ˆysinθsinϕ+ˆzcosθ,
ˆθ=ˆxcosθcosϕ+ˆycosθsinϕ−ˆzsinθ=∂ˆr
∂θ, (2.44)
ˆϕ=−ˆxsinϕ+ˆycosϕ=1
sinθ∂ˆr
∂ϕ,
whichfollowfrom
0=∂ˆr2
∂θ=2ˆr·∂ˆr
∂θ,0=∂ˆr2
∂ϕ=2ˆr·∂ˆr
∂ϕ.
126 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Note that Exercise 2.5.5 gives the inverse transformation and that a given vector can
nowbeexpressedinanumberofdifferent(butequivalent)ways.Forinstance,theposition
vectorrmaybewritten
r=ˆrr=ˆrparenleftbig
x2+y2+z2parenrightbig1/2
=ˆxx+ˆyy+ˆzz
=ˆxrsinθcosϕ+ˆyrsinθsinϕ+ˆzrcosθ. (2.45)
Selecttheform thatis mostusefulfor yourparticularproblem.
From Section 2.2, relabeling the curvilinear coordinate unit vectors ˆq1,ˆq2, andˆq3asˆr,
ˆθ, andˆϕgives
∇ψ=ˆr∂ψ
∂r+ˆθ1
r∂ψ
∂θ+ˆϕ1
rsinθ∂ψ
∂ϕ, (2.46)
∇·V=1
r2sinθbracketleftbigg
sinθ∂
∂r(r2Vr)+r∂
∂θ(sinθVθ)+r∂Vϕ
∂ϕbracketrightbigg
, (2.47)
∇·∇ψ=1
r2sinθbracketleftbigg
sinθ∂
∂rparenleftbigg
r2∂ψ
∂rparenrightbigg
+∂
∂θparenleftbigg
sinθ∂ψ
∂θparenrightbigg
+1
sinθ∂2ψ
∂ϕ2bracketrightbigg
,(2.48)
∇×V=1
r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆrrˆθrsinθˆϕ
∂
∂r∂
∂θ∂
∂ϕ
VrrVθrsinθVϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.49)
Occasionally, the vector Laplacian ∇2Vis needed in spherical polar coordinates. It is
bestobtainedbyusingthevectoridentity(Eq. (1.85))of Chapter1. Forreference
∇2V|r=parenleftbigg
−2
r2+2
r∂
∂r+∂2
∂r2+cosθ
r2sinθ∂
∂θ+1
r2∂2
∂θ2+1
r2sin2θ∂2
∂ϕ2parenrightbigg
Vr
+parenleftbigg
−2
r2∂
∂θ−2cosθ
r2sinθparenrightbigg
Vθ+parenleftbigg
−2
r2sinθ∂
∂ϕparenrightbigg
Vϕ
=∇2Vr−2
r2Vr−2
r2∂Vθ
∂θ−2cosθ
r2sinθVθ−2
r2sinθ∂Vϕ
∂ϕ, (2.50)
∇2V|θ=∇2Vθ−1
r2sin2θVθ+2
r2∂Vr
∂θ−2cosθ
r2sin2θ∂Vϕ
∂ϕ, (2.51)
∇2V|ϕ=∇2Vϕ−1
r2sin2θVϕ+2
r2sinθ∂Vr
∂ϕ+2cosθ
r2sin2θ∂Vθ
∂ϕ. (2.52)
These expressions for the components of ∇2Vare undeniably messy, but sometimes they
areneeded.
2.5 Spherical Polar Coordinates 127
Example 2.5.1 ∇,∇·,∇×FOR A CENTRAL FORCE
Using Eqs. (2.46) to (2.49), we can reproduce by inspection some of the results derived in
Chapter1bylaboriousapplicationofCartesiancoordinates.
From Eq. (2.46),
∇f(r)=ˆrdf
dr,
∇rn=ˆrnrn−1.(2.53)
FortheCoulombpotential V=Ze/(4πε0r), theelectricfieldis E=−∇V=Ze
4πε0r2ˆr.
From Eq. (2.47),
∇·ˆrf(r)=2
rf(r)+df
dr,
∇·ˆrrn=(n+2)rn−1.(2.54)
Forr>0 the charge density of the electric field of the Coulomb potential is ρ=∇·E=
Ze
4πε0∇·ˆr
r2=0 because n=−2.
From Eq. (2.48),
∇2f(r)=2
rdf
dr+d2f
dr2, (2.55)
∇2rn=n(n+1)rn−2, (2.56)
incontrasttotheordinaryradialsecondderivativeof rninvolving n−1 insteadof n+1.
Finally,from Eq.(2.49),
∇׈rf(r)=0. (2.57)
/squaresolid
Example 2.5.2 MAGNETIC VECTOR POTENTIAL
The computation of the magnetic vector potential of a single current loop in the xy-plane
usesOersted’slaw, ∇×H=J,inconjunctionwith µ0H=B=∇×A(seeExamples1.9.2
and1.12.1), andinvolvestheevaluationof
µ0J=∇×bracketleftbig
∇׈ϕAϕ(r,θ)bracketrightbig
.
Insphericalpolarcoordinatesthisreducesto
µ0J=∇×1
r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆrrˆθrsinθˆϕ
∂
∂r∂
∂θ∂
∂ϕ
00 rsinθAϕ(r,θ)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle
=∇×1
r2sinθbracketleftbigg
ˆr∂
∂θ(rsinθAϕ)−rˆθ∂
∂r(rsinθAϕ)bracketrightbigg
.
128 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Takingthecurlasecondtime,weobtain
µ0J=1
r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆr rˆθ rsinθˆϕ
∂
∂r∂
∂θ∂
∂ϕ
1
r2sinθ∂
∂θ(rsinθAϕ)−1
rsinθ∂
∂r(rsinθAϕ)0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
Byexpandingthedeterminantalongthetoprow, wehave
µ0J=−ˆϕbraceleftbigg1
r∂2
∂r2(rAϕ)+1
r2∂
∂θbracketleftbigg1
sinθ∂
∂θ(sinθAϕ)bracketrightbiggbracerightbigg
=−ˆϕbracketleftbigg
∇2Aϕ(r,θ)−1
r2sin2θAϕ(r,θ)bracketrightbigg
. (2.58)
/squaresolid
Exercises
2.5.1 Express thesphericalpolarunitvectorsinCartesianunitvectors.
ANS.ˆr=ˆxsinθcosϕ+ˆysinθsinϕ+ˆzcosθ,
ˆθ=ˆxcosθcosϕ+ˆycosθsinϕ−ˆzsinθ,
ˆϕ=−ˆxsinϕ+ˆycosϕ.
2.5.2 (a) From the results of Exercise 2.5.1, calculate the partial derivatives of ˆr,ˆθ, and ˆϕ
withrespectto r,θ,andϕ.
(b) With ∇givenby
ˆr∂
∂r+ˆθ1
r∂
∂θ+ˆϕ1
rsinθ∂
∂ϕ
(greatestspacerateofchange),usetheresultsofpart(a)tocalculate ∇·∇ψ.This
isanalternatederivationof theLaplacian.
Note.Thederivativesoftheleft-hand ∇operateontheunitvectorsoftheright-hand ∇
beforetheunitvectorsaredottedtogether.
2.5.3 Arigidbodyisrotatingaboutafixedaxiswithaconstantangularvelocity ω.T ake ωto
bealongthe z-axis. Usingsphericalpolarcoordinates,
(a) Calculate
v=ω×r.
(b) Calculate
∇×v.
ANS.(a)v=ˆϕωrsinθ,
(b)∇×v=2ω.
2.5 Spherical Polar Coordinates 129
2.5.4 Thecoordinatesystem (x,y,z)isrotatedthroughanangle /Phi1counterclockwiseaboutan
axisdefinedbytheunitvector nintosystem (x′,y′,z′).Intermsofthenewcoordinates
theradiusvectorbecomes
r′=rcos/Phi1+r×nsin/Phi1+n(n·r)(1−cos/Phi1).
(a) Derivethisexpressionfrom geometricconsiderations.
(b) Showthatitreducesasexpectedfor n=ˆz.Theanswer,inmatrixform,appearsin
Eq.(3.90).
(c) Verifythat r′2=r2.
2.5.5 ResolvetheCartesianunitvectorsintotheirsphericalpolarcomponents:
ˆx=ˆrsinθcosϕ+ˆθcosθcosϕ−ˆϕsinϕ,
ˆy=ˆrsinθsinϕ+ˆθcosθsinϕ+ˆϕcosϕ,
ˆz=ˆrcosθ−ˆθsinθ.
2.5.6 The direction of one vector is given by the angles θ1andϕ1. For a second vector the
corresponding angles are θ2andϕ2. Show that the cosine of the included angle γis
givenby
cosγ=cosθ1cosθ2+sinθ1sinθ2cos(ϕ1−ϕ2).
SeeFig. 12.15.
2.5.7 A certain vector Vhas no radial component. Its curl has no tangential components.
Whatdoesthisimplyabouttheradialdependenceofthetangentialcomponentsof V?
2.5.8 Modernphysicslaysgreatstressonthepropertyofparity—whetheraquantityremains
invariant or changes sign under an inversion of the coordinate system. In Cartesian
coordinatesthismeans x→−x,y→−y,andz→−z.
(a) Show that the inversion (reflection through the origin) of a point (r,θ,ϕ)relative
tofixedx-,y-,z-axesconsistsof thetransformation
r→r, θ→π−θ, ϕ→ϕ±π.
(b) Showthat ˆrandˆϕhaveoddparity(reversalofdirection)andthat ˆθhasevenparity.
2.5.9 WithAanyvector,
A·∇r=A.
(a) Verifythisresult inCartesiancoordinates.
(b) Verifythisresult usingsphericalpolarcoordinates.(Equation(2.46) provides ∇.)
130 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
2.5.10 Find the spherical coordinate components of the velocity and acceleration of a moving
particle:
vr=˙r,
vθ=r˙θ,
vϕ=rsinθ˙ϕ,
ar=¨r−r˙θ2−rsin2θ˙ϕ2,
aθ=r¨θ+2˙r˙θ−rsinθcosθ˙ϕ2,
aϕ=rsinθ¨ϕ+2˙rsinθ˙ϕ+2rcosθ˙θ˙ϕ.
Hint.
r(t)=ˆr(t)r(t)
=bracketleftbigˆxsinθ(t)cosϕ(t)+ˆysinθ(t)sinϕ(t)+ˆzcosθ(t)bracketrightbig
r(t).
Note.Using the Lagrangian techniques of Section 17.3, we may obtain these results
somewhat more elegantly. The dot in ˙r,˙θ,˙ϕmeans time derivative, ˙r=dr/dt,˙θ=
dθ/dt,˙ϕ=dϕ/dt.Thenotationwas originatedbyNewton.
2.5.11 Aparticle mmovesinresponsetoacentralforceaccordingtoNewton’ssecondlaw,
m¨r=ˆrf(r).
Show that r×˙r=c, a constant, and that the geometric interpretation of this leads to
Kepler’ssecondlaw.
2.5.12 Express∂/∂x,∂/∂y,∂/∂zinsphericalpolarcoordinates.
ANS.∂
∂x=sinθcosϕ∂
∂r+cosθcosϕ1
r∂
∂θ−sinϕ
rsinθ∂
∂ϕ,
∂
∂y=sinθsinϕ∂
∂r+cosθsinϕ1
r∂
∂θ+cosϕ
rsinθ∂
∂ϕ,
∂
∂z=cosθ∂
∂r−sinθ1
r∂
∂θ.
Hint.Equate ∇xyzand∇rθϕ.
2.5.13 From Exercise2.5.12showthat
−iparenleftbigg
x∂
∂y−y∂
∂xparenrightbigg
=−i∂
∂ϕ.
This is the quantum mechanical operator corresponding to the z-component of orbital
angularmomentum.
2.5.14 With the quantum mechanical orbital angular momentum operator defined as L=
−i(r×∇),showthat
(a)Lx+iLy=eiϕparenleftbigg∂
∂θ+icotθ∂
∂ϕparenrightbigg
,
2.5 Spherical Polar Coordinates 131
(b)Lx−iLy=−e−iϕparenleftbigg∂
∂θ−icotθ∂
∂ϕparenrightbigg
.
(ThesearetheraisingandloweringoperatorsofSection4.3.)
2.5.15 Verify that L×L=iLin spherical polar coordinates. L=−i(r×∇), the quantum
mechanicalorbitalangularmomentumoperator.
Hint.Use spherical polar coordinates for Lbut Cartesian components for the cross
product.
2.5.16 (a) FromEq. (2.46) showthat
L=−i(r×∇)=iparenleftbigg
ˆθ1
sinθ∂
∂ϕ−ˆϕ∂
∂θparenrightbigg
.
(b) Resolving ˆθandˆϕintoCartesiancomponents,determine Lx,Ly,andLzinterms
ofθ,ϕ, andtheirderivatives.
(c) From L2=L2
x+L2
y+L2
zshowthat
L2=−1
sinθ∂
∂θparenleftbigg
sinθ∂
∂θparenrightbigg
−1
sin2θ∂2
∂ϕ2
=−r2∇2+∂
∂rparenleftbigg
r2∂
∂rparenrightbigg
.
This latter identity is useful in relating orbital angular momentum and Legendre’s dif-
ferentialequation,Exercise9.3.8.
2.5.17 WithL=−ir×∇, verifytheoperatoridentities
(a)∇=ˆr∂
∂r−ir×L
r2,
(b)r∇2−∇parenleftbigg
1+r∂
∂rparenrightbigg
=i∇×L.
2.5.18 Showthatthefollowingthreeforms(sphericalcoordinates)of ∇2ψ(r)areequivalent:
(a)1
r2d
drbracketleftbigg
r2dψ(r)
drbracketrightbigg
;(b)1
rd2
dr2bracketleftbig
rψ(r)bracketrightbig
;(c)d2ψ(r)
dr2+2
rdψ(r)
dr.
The second form is particularly convenient in establishing a correspondence between
sphericalpolarandCartesiandescriptionsofaproblem.
2.5.19 Onemodelof thesolarcoronaassumesthatthesteady-stateequationof heatflow,
∇·(k∇T)=0,
is satisfied. Here, k, the thermal conductivity, is proportional to T5/2. Assuming that
the temperature Tis proportional to rn, show that the heat flow equation is satisfied by
T=T0(r0/r)2/7.
132 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
2.5.20 Acertainforcefieldisgivenby
F=ˆr2Pcosθ
r3+ˆθP
r3sinθ, r /greaterorequalslantP/2
(insphericalpolarcoordinates).
(a) Examine ∇×Ftosee ifapotentialexists.
(b) Calculatecontintegraltext
F·dλfor a unit circle in the plane θ=π/2. What does this indicate
abouttheforcebeingconservativeornonconservative?
(c) If you believe that Fmay be described by F=−∇ψ, findψ. Otherwise simply
statethatnoacceptablepotentialexists.
2.5.21 (a) Showthat A=−ˆϕcotθ/risasolutionof ∇×A=ˆr/r2.
(b) Show that this spherical polar coordinate solution agrees with the solution given
forExercise1.13.6:
A=ˆxyz
r(x2+y2)−ˆyxz
r(x2+y2).
Notethatthesolutiondivergesfor θ=0,πcorrespondingto x,y=0.
(c) Finally, show that A=−ˆθϕsinθ/ris a solution. Note that although this solution
does not diverge (r/negationslash=0), it is no longer single-valued for all possible azimuth
angles.
2.5.22 Amagneticvectorpotentialis givenby
A=µ0
4πm×r
r3.
Showthatthisleadstothemagneticinduction Bofapointmagneticdipolewithdipole
momentm.
ANS.form=ˆzm,
∇×A=ˆrµ0
4π2mcosθ
r3+ˆθµ0
4πmsinθ
r3.
CompareEqs. (12.133)and(12.134)
2.5.23 Atlargedistancesfromits source,electricdipoleradiationhasfields
E=aEsinθei(kr−ωt)
rˆθ,B=aBsinθei(kr−ωt)
rˆϕ.
ShowthatMaxwell’sequations
∇×E=−∂B
∂tand ∇×B=ε0µ0∂E
∂t
aresatisfied,if wetake
aE
aB=ω
k=c=(ε0µ0)−1/2.
Hint.Sincerislarge, termsoforder r−2maybedropped.
2.6 Tensor Analysis 133
2.5.24 Themagneticvectorpotentialfor auniformlychargedrotatingsphericalshellis
A=
ˆϕµ0a4σω
3·sinθ
r2,r>a
ˆϕµ0aσω
3·rcosθ, r<a.
(a=radius of spherical shell, σ=surface charge density, and ω=angular velocity.)
Findthemagneticinduction B=∇×A.
ANS.Br(r,θ)=2µ0a4σω
3·cosθ
r3,r>a,
Bθ(r,θ)=µ0a4σω
3·sinθ
r3,r>a,
B=ˆz2µ0aσω
3,r <a.
2.5.25 (a) Explainwhy ∇2inplanepolarcoordinatesfollowsfrom ∇2incircularcylindrical
coordinateswith z=constant.
(b) Explainwhytaking ∇2insphericalpolarcoordinatesandrestricting θtoπ/2does
notleadtotheplanepolarformof ∇.
Note.
∇2(ρ,ϕ)=∂2
∂ρ2+1
ρ∂
∂ρ+1
ρ2∂2
∂ϕ2.
2.6 T ENSOR ANALYSIS
Introduction, Definitions
Tensorsareimportantinmanyareasofphysics,includinggeneralrelativityandelectrody-
namics.Scalarsandvectorsarespecialcasesoftensors.InChapter1,aquantitythatdidnot
change under rotations of the coordinate system in three-dimensional space, an invariant,
was labeled a scalar. A scalaris specified by one real number and is a tensor of rank 0 .
A quantity whose components transformed under rotations like those of the distance of a
pointfromachosenorigin(Eq.(1.9),Section1.2)wascalledavector.Thetransformation
of the componentsof the vectorunder a rotationof thecoordinatespreserves the vectoras
a geometric entity (such as an arrow in space), independent of the orientation of the refer-
ence frame. In three-dimensional space, a vectoris specified by 3 =31real numbers, for
example, its Cartesian components, and is a tensor of rank 1 .Atensor of rank nhas 3n
componentsthattransforminadefiniteway.5Thistransformationphilosophyisofcentral
importance for tensor analysis and conforms with the mathematician’s concept of vector
and vector (or linear) space and the physicist’s notion that physical observables must not
dependonthechoiceofcoordinateframes.Thereisaphysicalbasisforsuchaphilosophy:
We describe the physical world by mathematics, but any physical predictions we make
5InN-dimensional spacea tensorof rank nhasNncomponents.
134 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
must be independent of our mathematical conventions, such as a coordinate system with
itsarbitraryoriginandorientationofits axes.
Thereis apossibleambiguityinthetransformationlawofa vector
A′
i=summationdisplay
jaijAj, (2.59)
inwhich aijis thecosineof theanglebetweenthe x′
i-axisandthe xj-axis.
If we start with a differential distance vector dr, then, taking dx′
ito be a functionof the
unprimedvariables,
dx′
i=summationdisplay
j∂x′
i
∂xjdxj (2.60)
bypartialdifferentiation.If weset
aij=∂x′
i
∂xj, (2.61)
Eqs.(2.59)and(2.60)areconsistent.Anysetofquantities Ajtransformingaccordingto
A′i=summationdisplay
j∂x′
i
∂xjAj(2.62a)
isdefinedasa contravariant vector,whoseindiceswewriteas superscript ;thisincludes
theCartesiancoordinatevector xi=xifrom nowon.
However,wehavealreadyencounteredaslightlydifferenttypeofvectortransformation.
Thegradientofascalar ∇ϕ, definedby
∇ϕ=ˆx∂ϕ
∂x1+ˆy∂ϕ
∂x2+ˆz∂ϕ
∂x3(2.63)
(usingx1,x2,x3forx,y,z), transformsas
∂ϕ′
∂x′i=summationdisplay
j∂ϕ
∂xj∂xj
∂x′i, (2.64)
usingϕ=ϕ(x,y,z)=ϕ(x′,y′,z′)=ϕ′,ϕdefined as a scalar quantity. Notice that this
differs from Eq. (2.62) in that we have ∂xj/∂x′iinstead of ∂x′i/∂xj. Equation (2.64)
is taken as the definition of a covariant vector, with the gradient as the prototype. The
covariantanalogof Eq.(2.62a) is
A′
i=summationdisplay
j∂xj
∂x′iAj. (2.62b)
OnlyinCartesiancoordinatesis
∂xj
∂x′i=∂x′i
∂xj=aij (2.65)
2.6 Tensor Analysis 135
so that there no difference between contravariant and covariant transformations. In other
systems, Eq. (2.65) in general does not apply, and the distinction between contravariant
and covariant is real and must be observed. This is of prime importance in the curved
Riemannianspaceofgeneralrelativity.
Intheremainderofthissectionthecomponentsofany contravariant vectoraredenoted
by asuperscript ,Ai, whereas a subscript is used for the components of a covariant
vectorAi.6
Definition of Tensors of Rank 2
Nowweproceedtodefine contravariant,mixed,andcovarianttensorsofrank2 bythe
followingequationsfortheircomponentsundercoordinatetransformations:
A′ij=summationdisplay
kl∂x′i
∂xk∂x′j
∂xlAkl,
B′ij=summationdisplay
kl∂x′i
∂xk∂xl
∂x′jBkl, (2.66)
C′
ij=summationdisplay
kl∂xk
∂x′i∂xl
∂x′jCkl.
Clearly, the rank goes as the number of partial derivatives (or direction cosines) in the de-
finition: 0 for a scalar, 1 for a vector, 2 for a second-rank tensor, and so on. Each index
(subscript or superscript) ranges over the number of dimensions of the space. The number
of indices (equal to the rank of tensor) is independent of the dimensions of the space. We
see thatAklis contravariant with respect to both indices, Cklis covariant with respect to
bothindices,and Bkltransformscontravariantlywithrespecttothefirstindex kbutcovari-
antlywithrespecttothesecondindex l.Onceagain,ifweareusingCartesiancoordinates,
all three forms of the tensors of secondrank contravariant,mixed,and covariantare—the
same.
As with the components of a vector, the transformation laws for the components of a
tensor, Eq. (2.66), yield entities (and properties) that are independent of the choice of ref-
erence frame. This is what makes tensor analysis important in physics. The independence
of reference frame (invariance) is ideal for expressing and investigating universal physical
laws.
Thesecond-ranktensor A(components Akl)maybeconvenientlyrepresentedbywriting
outits componentsinasquarearray(3 ×3 ifweareinthree-dimensionalspace):
A=
A11A12A13
A21A22A23
A31A32A33
. (2.67)
This does not mean that any square array of numbers or functions forms a tensor. The
essentialconditionisthatthecomponentstransformaccordingtoEq.(2.66).
6Thismeansthatthecoordinates (x,y,z)arewritten (x1,x2,x3)sincertransformsasacontravariantvector.Theambiguityof
x2representing both xsquared and yis theprice wepay.
136 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
In the context of matrix analysis the preceding transformation equations become (for
Cartesiancoordinates)anorthogonalsimilaritytransformation;seeSection3.3.Ageomet-
ricalinterpretationofa second-ranktensor(theinertiatensor) isdevelopedinSection3.5.
In summary, tensors are systems of components organized by one or more indices that
transform according to specific rules under a set of transformations. The number of in-
dices is called the rank of the tensor. If the transformations are coordinate rotations in
three-dimensional space, then tensor analysis amounts to what we did in the sections on
curvilinear coordinates and in Cartesian coordinates in Chapter 1. In four dimensions of
Minkowski space–time, the transformations are Lorentz transformations, and tensors of
rank1arecalledfour-vectors.
Addition and Subtraction of Tensors
The addition and subtraction of tensors is defined in terms of the individual elements, just
asfor vectors.If
A+B=C, (2.68)
then
Aij+Bij=Cij.
Of course, AandBmust be tensors of the same rank and both expressed in a space of the
samenumberofdimensions.
Summation Convention
In tensor analysis it is customary to adopt a summation convention to put Eq. (2.66) and
subsequent tensor equations in a more compact form. As long as we are distinguishing
betweencontravarianceandcovariance,letusagreethatwhenanindexappearsononeside
of an equation, once as a superscript and once as a subscript (except for the coordinates
where both are subscripts), we automatically sum over that index. Then we may write the
secondexpressioninEq. (2.66) as
B′ij=∂x′i
∂xk∂xl
∂x′jBkl, (2.69)
withthesummationoftheright-handsideover kandlimplied.ThisisEinstein’ssumma-
tion convention.7The index iis superscript because it is associated with the contravariant
x′i;likewise jissubscriptbecauseitisrelatedtothecovariantgradient.
To illustrate the use of the summation convention and some of the techniques of tensor
analysis, let us show that the now-familiar Kronecker delta, δkl, is really a mixed tensor
7In this context ∂x′i/∂xkmight better be written as ai
kand∂xl/∂x′jasbl
j.
2.6 Tensor Analysis 137
of rank 2, δkl.8The question is: Does δkltransform according to Eq. (2.66)? This is our
criterionfor callingitatensor.Wehave,usingthesummationconvention,
δkl∂x′i
∂xk∂xl
∂x′j=∂x′i
∂xk∂xk
∂x′j(2.70)
bydefinitionoftheKroneckerdelta.Now,
∂x′i
∂xk∂xk
∂x′j=∂x′i
∂x′j(2.71)
by direct partial differentiation of the right-hand side (chain rule). However, x′iandx′j
are independent coordinates, and therefore the variation of one with respect to the other
mustbezeroif theyaredifferent, unityif theycoincide;thatis,
∂x′i
∂x′j=δ′ij. (2.72)
Hence
δ′ij=∂x′i
∂xk∂xl
∂x′jδkl,
showingthatthe δklareindeedthecomponentsofamixedsecond-ranktensor.Noticethat
this result is independent of the number of dimensions of our space. The reason for the
upperindex iandlowerindex jisthesameasinEq. (2.69).
TheKroneckerdeltahasonefurtherinterestingproperty.Ithasthesamecomponentsin
all of our rotated coordinate systems and is therefore called isotropic. In Section 2.9 we
shallmeetathird-rankisotropictensorandthreefourth-rankisotropictensors.Noisotropic
first-ranktensor(vector)exists.
Symmetry–Antisymmetry
Theorderinwhichtheindicesappearinourdescriptionofatensorisimportant.Ingeneral,
Amnisindependentof Anm,buttherearesomecasesofspecialinterest.If,forall mandn,
Amn=Anm, (2.73)
wecallthetensor symmetric .If, ontheotherhand,
Amn=−Anm, (2.74)
thetensoris antisymmetric .Clearly,every(second-rank)tensorcanberesolvedintosym-
metricandantisymmetricpartsbytheidentity
Amn=1
2parenleftbig
Amn+Anmparenrightbig
+1
2parenleftbig
Amn−Anmparenrightbig
, (2.75)
the first term on the right being a symmetric tensor, the second, an antisymmetric tensor.
A similar resolution of functions into symmetric and antisymmetric parts is of extreme
importancetoquantummechanics.
8Itiscommonpracticetorefertoatensor Abyspecifyingatypicalcomponent, Aij.Aslongasthereaderrefrainsfromwriting
nonsense suchas A=Aij, no harm is done.
138 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Spinors
It was once thought that the system of scalars, vectors, tensors (second-rank), and so on
formed a complete mathematical system, one that is adequate for describing a physics
independent of the choice of reference frame. But the universe and mathematical physics
are not that simple. In the realm of elementary particles, for example, spin zero particles9
(πmesons,αparticles) may be described with scalars, spin 1 particles (deuterons) by
vectors, and spin 2 particles (gravitons) by tensors. This listing omits the most common
particles: electrons, protons, and neutrons, all with spin1
2. These particles are properly
described by spinors. A spinor is not a scalar, vector, or tensor. A brief introduction to
spinorsinthecontextof grouptheory (J=1/2)appearsinSection4.3.
Exercises
2.6.1 Show that if all the components of any tensor of any rank vanish in one particular
coordinatesystem, theyvanishinallcoordinatesystems.
Note.This point takes on special importance in the four-dimensional curved space of
generalrelativity.Ifaquantity,expressedasatensor,existsinonecoordinatesystem,it
existsinallcoordinatesystemsandisnotjustaconsequenceofa choiceofacoordinate
system(as arecentrifugalandCoriolisforcesinNewtonianmechanics).
2.6.2 The components of tensor Aare equal to the corresponding components of tensor Bin
oneparticularcoordinatesystem, denotedbythesuperscript 0;thatis,
A0
ij=B0
ij.
Showthattensor Aisequaltotensor B,Aij=Bij, inallcoordinatesystems.
2.6.3 Thelastthreecomponentsofafour-dimensionalvectorvanishineachoftworeference
frames. If the second reference frame is not merely a rotation of the first about the x0
axis, that is, if at least one of the coefficients ai0(i=1,2,3)/negationslash=0, show that the zeroth
component vanishes in all reference frames. Translated into relativistic mechanics this
means that if momentum is conserved in two Lorentz frames, then energy is conserved
inallLorentzframes.
2.6.4 From an analysis of the behavior of a general second-rank tensor under 90◦and 180◦
rotations about the coordinate axes, show that an isotropic second-rank tensor in three-
dimensionalspacemustbeamultipleof δij.
2.6.5 The four-dimensional fourth-rank Riemann–Christoffel curvature tensor of general rel-
ativity,Riklm,satisfiesthesymmetryrelations
Riklm=−Rikml=−Rkilm.
Withtheindicesrunningfrom0to3,showthatthenumberofindependentcomponents
isreducedfrom256to36andthatthecondition
Riklm=Rlmik
9The particle spin is intrinsic angular momentum (in units of ¯h). It is distinct from classical, orbital angular momentum due to
motion.
2.7 Contraction, Direct Product 139
furtherreducesthenumberofindependentcomponentsto21.Finally,ifthecomponents
satisfy an identity Riklm+Rilmk+Rimkl=0, show that the number of independent
componentsis reducedto20.
Note.Thefinalthree-termidentityfurnishesnewinformationonlyifallfourindicesare
different.Thenitreducesthenumberofindependentcomponentsbyone-third.
2.6.6 Tiklmisantisymmetricwithrespecttoallpairsofindices.Howmanyindependentcom-
ponentshasit(in three-dimensionalspace)?
2.7 C ONTRACTION ,DIRECT PRODUCT
Contraction
Whendealingwithvectors,weformedascalarproduct(Section1.3)bysummingproducts
ofcorrespondingcomponents:
A·B=AiBi(summationconvention ). (2.76)
The generalization of this expression in tensor analysis is a process known as contraction.
Twoindices,onecovariantandtheothercontravariant,aresetequaltoeachother,andthen
(as implied by the summation convention) we sum over this repeated index. For example,
letuscontractthesecond-rankmixedtensor B′ij,
B′ii=∂x′i
∂xk∂xl
∂x′iBkl=∂xl
∂xkBkl (2.77)
usingEq. (2.71), andthenbyEq. (2.72)
B′ii=δlkBkl=Bkk. (2.78)
Our contracted second-rank mixed tensor is invariant and therefore a scalar.10This is ex-
actlywhatweobtainedinSection1.3forthedotproductoftwovectorsandinSection1.7
for the divergence of a vector. In general, the operation of contraction reduces the rank of
atensorby2.Anexampleof theuseof contractionappearsinChapter4.
Direct Product
Thecomponentsofacovariantvector(first-ranktensor) aiandthoseofacontravariantvec-
tor (first-rank tensor) bjmay be multiplied component by component to give the general
termaibj. This, byEq. (2.66)is actuallyasecond-ranktensor,for
a′
ib′j=∂xk
∂x′iak∂x′j
∂xlbl=∂xk
∂x′i∂x′j
∂xlparenleftbig
akblparenrightbig
. (2.79)
Contracting,weobtain
a′
ib′i=akbk, (2.80)
10In matrix analysis this scalaris the traceof the matrix, Section 3.2.
140 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
asinEqs. (2.77) and(2.78), togivetheregularscalarproduct.
The operation of adjoining two vectors aiandbjas in the last paragraph is known as
forming the direct product . For the case of two vectors, the direct product is a tensor of
second rank. In this sense we may attach meaning to ∇E, which was not defined within
the framework of vector analysis. In general, the direct product of two tensors is a tensor
ofrankequaltothesumofthetwoinitialranks;thatis,
AijBkl=Cijkl, (2.81a)
whereCijklisatensorof fourthrank.FromEqs. (2.66),
C′ijkl=∂x′i
∂xm∂xn
∂x′j∂x′k
∂xp∂x′l
∂xqCmnpq. (2.81b)
The direct product is a technique for creating new, higher-rank tensors. Exer-
cise2.7.1isaformofthedirectproductinwhichthefirstfactoris ∇.Applicationsappear
inSection4.6.
WhenTis annth-rank Cartesian tensor, (∂/∂xi)Tjkl...,a component of ∇T,i sa
Cartesian tensor of rank n+1 (Exercise 2.7.1). However, (∂/∂xi)Tjkl...is not a tensor
inmoregeneralspaces.Innon-Cartesiansystems ∂/∂x′iwillactonthepartialderivatives
∂xp/∂x′qanddestroythesimpletensortransformationrelation(seeEq. (2.129)).
So far the distinction between a covariant transformation and a contravariant transfor-
mationhasbeenmaintainedbecauseitdoesexistinnon-Euclideanspaceandbecauseitis
ofgreatimportanceingeneralrelativity.InSections2.10and2.11weshalldevelopdiffer-
entialrelationsforgeneraltensors.Often,however,becauseofthesimplificationachieved,
werestrict ourselves to Cartesiantensors. As notedin Section2.6, thedistinctionbetween
contravarianceandcovariancedisappears.
Exercises
2.7.1 IfT···iis a tensor of rank n, show that ∂T···i/∂xjis a tensor of rank n+1( C a r t e s i a n
coordinates).
Note.Innon-Cartesiancoordinatesystemsthecoefficients aijare,ingeneral,functions
of the coordinates, and the simple derivative of a tensor of rank nis not a tensor except
in the special case of n=0. In this case the derivative does yield a covariant vector
(tensorof rank1)byEq.(2.64).
2.7.2 IfTijk···is a tensor of rank n, show thatsummationtext
j∂Tijk···/∂xjis a tensor of rank n−1
(Cartesiancoordinates).
2.7.3 Theoperator
∇2−1
c2∂2
∂t2
maybewrittenas
4summationdisplay
i=1∂2
∂x2
i,
2.8 Quotient Rule 141
usingx4=ict. This is the four-dimensional Laplacian, sometimes called the d’Alem-
bertian and denoted by /square2. Show that it is a scalaroperator, that is, is invariant under
Lorentztransformations.
2.8 Q UOTIENT RULE
IfAiandBjarevectors,asseeninSection2.7,wecaneasilyshowthat AiBjisasecond-
ranktensor.Hereweareconcernedwithavarietyofinverserelations.Considersuchequa-
tionsas
KiAi=B (2.82a)
KijAj=Bi (2.82b)
KijAjk=Bik (2.82c)
KijklAij=Bkl (2.82d)
KijAk=Bijk. (2.82e)
Inline with our restriction to Cartesian systems, we write all indices as subscripts and,
unlessspecifiedotherwise,sumrepeatedindices.
Ineachoftheseexpressions AandBareknowntensorsofrankindicatedbythenumber
of indices and Ais arbitrary. In each case Kis an unknown quantity. We wish to establish
thetransformationpropertiesof K.Thequotientruleassertsthatiftheequationofinterest
holdsinall(rotated)Cartesiancoordinatesystems, Kisatensoroftheindicatedrank.The
importance in physical theory is that the quotient rule can establish the tensor nature of
quantities. Exercise 2.8.1 is a simple illustration of this. The quotient rule (Eq. (2.82b))
shows that the inertia matrix appearing in the angular momentum equation L=Iω, Sec-
tion3.5, isa tensor.
In proving the quotient rule, we consider Eq. (2.82b) as a typical case. In our primed
coordinatesystem
K′
ijA′
j=B′
i=aikBk, (2.83)
using the vector transformation properties of B. Since the equation holds in all rotated
Cartesiancoordinatesystems,
aikBk=aik(KklAl). (2.84)
Now, transforming Aback into the primed coordinate system11(compare Eq. (2.62)), we
have
K′
ijA′
j=aikKklajlA′
j. (2.85)
Rearranging,weobtain
(K′
ij−aikajlKkl)A′
j=0. (2.86)
11Notethe order of the indices of the direction cosine ajlin thisinversetransformation. Wehave
Al=summationdisplay
j∂xl
∂x′
jA′
j=summationdisplay
jajlA′
j.
142 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Thismustholdforeachvalueoftheindex iandforeveryprimedcoordinatesystem.Since
theA′
jisarbitrary,12weconclude
K′
ij=aikajlKkl, (2.87)
whichisourdefinitionofsecond-ranktensor.
The other equations may be treated similarly, giving rise to other forms of the quotient
rule. One minor pitfall should be noted: The quotient rule does not necessarily apply if B
iszero. Thetransformationpropertiesof zeroareindeterminate.
Example 2.8.1 EQUATIONS OF MOTION AND FIELD EQUATIONS
In classical mechanics, Newton’s equations of motion m˙v=Ftell us on the basis of the
quotientrule that, if the mass is a scalar and the force a vector, then the acceleration a≡˙v
isavector.Inotherwords,thevectorcharacteroftheforceasthedrivingtermimposesits
vectorcharacterontheacceleration,providedthescalefactor misscalar.
The wave equation of electrodynamics ∂2Aµ=Jµinvolves the four-dimensional ver-
sionoftheLaplacian ∂2=∂2
c2∂t2−∇2,aLorentzscalar,andtheexternalfour-vectorcurrent
Jµas its driving term. From the quotient rule, we infer that the vector potential Aµis a
four-vector as well. If the driving current is a four-vector, the vector potential must be of
rank1bythequotientrule. /squaresolid
Thequotientruleis asubstitutefor theillegaldivisionoftensors.
Exercises
2.8.1 The double summation KijAiBjis invariant for any two vectors AiandBj. Prove that
Kijisasecond-ranktensor.
Note.In the form ds2(invariant)=gijdxidxj, this result shows that the matrix gijis
atensor.
2.8.2 Theequation KijAjk=Bikholdsforallorientationsofthecoordinatesystem.If Aand
Bare arbitrarysecond-ranktensors,showthat Kis asecond-ranktensoralso.
2.8.3 Theexponentialinaplanewaveisexp [i(k·r−ωt)].Werecognize xµ=(ct,x1,x2,x3)
asaprototypevectorinMinkowskispace.If k·r−ωtisascalarunderLorentztransfor-
mations(Section4.5), showthat kµ=(ω/c,k 1,k2,k3)isavectorinMinkowskispace.
Note.Multiplicationby ¯hyields(E/c,p)asavectorinMinkowskispace.
2.9 P SEUDOTENSORS ,DUAL TENSORS
So far our coordinate transformations have been restricted to pure passive rotations. We
nowconsidertheeffectofreflectionsor inversions.
12We might, for instance, take A′
1=1a n dA′m=0f o rm/negationslash=1. Then the equation K′
i1=aika1lKklfollows immediately. The
rest of Eq.(2.87) comes from otherspecial choicesof the arbitrary A′
j.
2.9 Pseudotensors, Dual Tensors 143
FIGURE 2.9Inversionof Cartesiancoordinates—polarvector.
If wehavetransformationcoefficients aij=−δij, thenbyEq. (2.60)
xi=−x′i, (2.88)
which is an inversion or parity transformation. Note that this transformation changes our
initial right-handed coordinate system into a left-handed coordinate system.13Our proto-
typevector rwithcomponents (x1,x2,x3)transformsto
r′=parenleftbig
x′1,x′2,x′3parenrightbig
=parenleftbig
−x1,−x2,−x3parenrightbig
.
This new vector r′has negative components, relative to the new transformed set of axes.
As shown in Fig. 2.9, reversing the directions of the coordinate axes and changing the
signs of the components gives r′=r. The vector (an arrow in space) stays exactly as it
was before the transformation was carried out. The position vector rand all other vectors
whose components behave this way (reversing sign with a reversal of the coordinate axes)
arecalled polarvectors andhaveoddparity.
Afundamentaldifferenceappearswhenweencounteravectordefinedasthecrossprod-
uct of two polar vectors. Let C=A×B, where both AandBare polar vectors. From
Eq.(1.33), thecomponentsof Caregivenby
C1=A2B3−A3B2(2.89)
andsoon.Now,whenthecoordinateaxesareinverted, Ai→−A′i,Bj→−B′
j,butfrom
itsdefinition Ck→+C′k;thatis,ourcross-productvector,vector C,doesnotbehavelike
a polar vector under inversion. To distinguish, we label it a pseudovector or axial vector
(seeFig.2.10)thathasevenparity.Theterm axialvector isfrequentlyusedbecausethese
crossproductsoftenarisefroma descriptionofrotation.
13This is an inversion of thecoordinate system or coordinate axes, objects in the physical world remaining fixed.
144 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
FIGURE 2.10InversionofCartesiancoordinates—axialvector.
Examplesare
angularvelocity, v=ω×r,
orbitalangularmomentum, L=r×p,
torque,force=F, N=r×F,
magneticinductionfield B,∂B
∂t=−∇×E.
Inv=ω×r, the axial vector is the angular velocity ω, andrandv=dr/dtare polar
vectors. Clearly, axial vectors occur frequently in physics, although this fact is usually
not pointed out. In a right-handed coordinate system an axial vector Chas a sense of
rotationassociatedwithitgivenbyaright-handrule(compareSection1.4).Intheinverted
left-handed system the sense of rotation is a left-handed rotation. This is indicated by the
curvedarrowsinFig.2.10.
The distinction between polar and axial vectors may also be illustrated by a reflection.
Apolarvectorreflectsinamirrorlikearealphysicalarrow,Fig.2.11a.InFigs.2.9and2.10
the coordinates are inverted; the physical world remains fixed. Here the coordinate axes
remain fixed; the world is reflected—as in a mirror in the xz-plane. Specifically, in this
representation we keep the axes fixed and associate a change of sign with the component
ofthevector.Foramirrorinthe xz-plane,Py→−Py.W eh a v e
P=(Px,Py,Pz)
P′=(Px,−Py,Pz)polarvector.
An axial vector such as a magnetic field Hor a magnetic moment µ(=current×area
of current loop) behaves quite differently under reflection. Consider the magnetic field
Hand magnetic moment µto be produced by an electric charge moving in a circular path
(Exercise5.8.4andExample12.5.3).Reflectionreversesthesenseofrotationofthecharge.
2.9 Pseudotensors, Dual Tensors 145
a
b
FIGURE 2.11(a) Mirror in xz-plane;(b) mirror
inxz-plane.
The two current loops and the resulting magnetic moments are shown in Fig. 2.11b. We
have
µ=(µx,µy,µz)
µ′=(−µx,µy,−µz)reflectedaxialvector.
146 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
If we agree that the universe does not care whether we use a right- or left-handed coor-
dinate system, then it does not make sense to add an axial vector to a polar vector. In the
vector equation A=B, bothAandBare either polar vectors or axial vectors.14Similar
restrictionsapplytoscalarsandpseudoscalarsand,ingeneral,tothetensorsandpseudoten-
sors consideredsubsequently.
Usually,pseudoscalars,pseudovectors,andpseudotensorswilltransformas
S′=JS, C′
i=JaijCj,A′
ij=JaikajlAkl, (2.90)
whereJis the determinant15of the array of coefficients amn, the Jacobian of the parity
transformation.In ourinversiontheJacobianis
J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−10 0
0−10
00−1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−1. (2.91)
Forareflectionofoneaxis,the x-axis,
J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−100
01 0
00 1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−1, (2.92)
andagaintheJacobian J=−1.Ontheotherhand,forallpurerotations,theJacobian Jis
always+1.RotationmatricesdiscussedfurtherinSection3.3.
In Chapter 1 the triple scalar product S=A×B·Cwas shown to be a scalar (un-
der rotations). Now by considering the parity transformation given by Eq. (2.88), we see
thatS→−S, proving that the triple scalar product is actually a pseudoscalar: This be-
havior was foreshadowed by the geometrical analogy of a volume. If all three parameters
of the volume—length, depth, and height—change from positive distances to negative
distances,theproductof thethreewillbenegative.
Levi-Civita Symbol
For future use it is convenient to introduce the three-dimensional Levi-Civita symbol εijk,
definedby
ε123=ε231=ε312=1,
ε132=ε213=ε321=−1, (2.93)
allotherεijk=0.
Note that εijkis antisymmetric with respect to all pairs of indices. Suppose now that we
have a third-rank pseudotensor δijk, which in one particular coordinate system is equal to
εijk. Then
δ′
ijk=|a|aipajqakrεpqr (2.94)
14The big exception to this is in beta decay, weak interactions. Here the universe distinguishes between right- and left-handed
systems, andweaddpolarandaxial vector interactions.
15Determinants are describedinSection3.1.
2.9 Pseudotensors, Dual Tensors 147
bydefinitionofpseudotensor.Now,
a1pa2qa3rεpqr=|a| (2.95)
by direct expansion of the determinant, showing that δ′
123=|a|2=1=ε123. Considering
theotherpossibilitiesonebyone,wefind
δ′
ijk=εijk (2.96)
for rotations and reflections. Hence εijkis a pseudotensor.16,17Furthermore, it is seen to
beanisotropicpseudotensorwiththesamecomponentsinallrotatedCartesiancoordinate
systems.
Dual Tensors
Withany antisymmetric second-ranktensor C(inthree-dimensionalspace)wemayasso-
ciateadualpseudovector Cidefinedby
Ci=1
2εijkCjk. (2.97)
Heretheantisymmetric Cmaybewritten
C=
0C12−C31
−C120C23
C31−C230
. (2.98)
Weknowthat Cimusttransformasavectorunderrotationsfromthedoublecontractionof
the fifth-rank (pseudo) tensor εijkCmnbut that it is really a pseudovector from the pseudo
natureof εijk. Specifically,thecomponentsof Caregivenby
(C1,C2,C3)=parenleftbig
C23,C31,C12parenrightbig
. (2.99)
Notice the cyclic order of the indices that comes from the cyclic order of the components
ofεijk. Eq. (2.99) means that our three-dimensional vector product may literally be taken
tobeeitherapseudovectororanantisymmetricsecond-ranktensor,dependingonhowwe
choosetowriteitout.
If wetakethree(polar)vectors A,B, andC,wemaydefinethedirectproduct
Vijk=AiBjCk. (2.100)
By an extension of the analysis of Section 2.6, Vijkis a tensor of third rank. The dual
quantity
V=1
3!εijkVijk(2.101)
16The usefulness of εpqrextends far beyond this section. For instance, the matrices Mkof Exercise 3.2.16 are derived from
(Mr)pq=−iεpqr. Much of elementary vector analysis can be written in a very compact form by using εijkand the identity of
Exercise 2.9.4SeeA.A.Evett,Permutation symbol approachtoelementaryvector analysis. Am.J.Phys. 34: 503 (1966).
17Thenumerical valueof εpqris given by the triple scalarproduct of coordinate unit vectors:
ˆxp·ˆxq׈xr.
From this point of view eachelementof εpqris a pseudoscalar, but the εpqrcollectively form athird-rank pseudotensor.
148 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
isclearlyapseudoscalar.Byexpansionit isseenthat
V=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA1B1C1
A2B2C2
A3B3C3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(2.102)
isourfamiliar triplescalarproduct.
ForuseinwritingMaxwell’sequationsincovariantform,Section4.6,wewanttoextend
this dual vector analysis to four-dimensional space and, in particular, to indicate that the
four-dimensionalvolumeelement dx0dx1dx2dx3isapseudoscalar.
We introduce the Levi-Civita symbol εijkl, the four-dimensional analog of εijk.T h i s
quantityεijklis defined as totally antisymmetric in all four indices. If (ijkl)is an even
permutation18of(0,1,2,3), thenεijklis defined as+1; if it is an odd permutation,
thenεijklis−1, and 0 if any two indices are equal. The Levi-Civita εijklm a yb ep r o v e da
pseudotensorofrank4byanalysissimilartothatusedforestablishingthetensornatureof
εijk. Introducingthedirectproductof fourvectorsasfourth-ranktensorwithcomponents
Hijkl=AiBjCkDl, (2.103)
builtfromthepolarvectors A,B,C,andD,wemaydefinethedualquantity
H=1
4!εijklHijkl, (2.104)
a pseudoscalar due to the quadruple contraction with the pseudotensor εijkl.Now we let
A,B,C, andDbe infinitesimal displacements along the four coordinate axes (Minkowski
space),
A=parenleftbig
dx0,0,0,0parenrightbig
B=parenleftbig
0,dx1,0,0parenrightbig
,andso on,(2.105)
and
H=dx0dx1dx2dx3. (2.106)
The four-dimensional volume element is now identified as a pseudoscalar. We use this
result in Section 4.6. This result could have been expected from the results of the special
theory of relativity. The Lorentz–Fitzgerald contraction of dx1dx2dx3just balances the
timedilationof dx0.
We slipped into this four-dimensional space as a simple mathematical extension of the
three-dimensional space and, indeed, we could just as easily have discussed 5-, 6-, or N-
dimensionalspace.Thisistypicalofthepowerofthecomponentanalysis.Physically,this
four-dimensionalspacemaybetakenasMinkowskispace,
parenleftbig
x0,x1,x2,x3parenrightbig
=(ct,x,y,z), (2.107)
wheretis time. This is the merger of space and time achieved in special relativity. The
transformationsthatdescribetherotationsinfour-dimensionalspacearetheLorentztrans-
formationsofspecialrelativity.WeencountertheseLorentztransformationsinSection4.6.
18A permutation is odd if it involves an odd number of interchanges of adjacent indices, such as (0123)→(0213).E v e n
permutations arise from an even number of transpositions of adjacent indices. (Actually the word adjacent is unnecessary.)
ε0123=+1.
2.9 Pseudotensors, Dual Tensors 149
Irreducible Tensors
Forsomeapplications,particularlyinthequantumtheoryofangularmomentum,ourCarte-
siantensorsarenotparticularlyconvenient.Inmathematicallanguageourgeneralsecond-
rank tensor Aijis reducible, which means that it can be decomposed into parts of lower
tensorrank.Infact, wehavealreadydonethis.FromEq.(2.78),
A=Aii (2.108)
isascalarquantity,thetraceof Aij.19
Theantisymmetricportion,
Bij=1
2(Aij−Aji), (2.109)
hasjustbeenshowntobeequivalenttoa(pseudo)vector,or
Bij=Ckcyclicpermutationof i,j,k. (2.110)
By subtracting the scalar Aand the vector Ckfrom our original tensor, we have an irre-
ducible,symmetric,zero-tracesecond-ranktensor, Sij,inwhich
Sij=1
2(Aij+Aji)−1
3Aδij, (2.111)
withfiveindependentcomponents.Then,finally,ouroriginalCartesiantensormaybewrit-
ten
Aij=1
3Aδij+Ck+Sij. (2.112)
The three quantities A,Ck, andSijform spherical tensors of rank 0, 1, and 2, respec-
tively, transforming like the spherical harmonics YM
L(Chapter 12) for L=0, 1, and 2.
Further details of such spherical tensors and their uses will be found in Chapter 4 and the
booksbyRoseandEdmondscitedthere.
A specific example of the preceding reduction is furnished by the symmetric electric
quadrupoletensor
Qij=integraldisplayparenleftbig
3xixj−r2δijparenrightbig
ρ(x1,x2,x3)d3x.
The−r2δijterm represents a subtraction of the scalar trace (the three i=jterms). The
resulting Qijhaszerotrace.
Exercises
2.9.1 Anantisymmetricsquarearrayis givenby
0C3−C2
−C30C1
C2−C10
=
0C12C13
−C120C23
−C13−C230
,
19Analternateapproach, using matrices, is given inSection3.3(see Exercise3.3.9).
150 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
where(C1,C2,C3)formapseudovector.Assumingthattherelation
Ci=1
2!εijkCjk
holds in all coordinate systems, prove that Cjkis a tensor. (This is another form of the
quotienttheorem.)
2.9.2 Showthatthevectorproductisuniquetothree-dimensionalspace;thatis,onlyinthree
dimensions can we establish a one-to-one correspondence between the components of
anantisymmetrictensor(second-rank)andthecomponentsofavector.
2.9.3 Showthatin R3
(a)δii=3,
(b)δijεijk=0,
(c)εipqεjpq=2δij,
(d)εijkεijk=6.
2.9.4 Showthatin R3
εijkεpqk=δipδjq−δiqδjp.
2.9.5 (a) Express the components of a cross-product vector C,C=A×B, in terms of εijk
andthecomponentsof AandB.
(b) Usetheantisymmetryof εijktoshowthat A·A×B=0.
ANS.(a)Ci=εijkAjBk.
2.9.6 (a) Showthattheinertiatensor(matrix)maybewritten
Iij=m(xixjδij−xixj)
fora particleofmass mat(x1,x2,x3).
(b) Showthat
Iij=−MilMlj=−mεilkxkεljmxm,
whereMil=m1/2εilkxk.Thisisthecontractionoftwosecond-ranktensorsandis
identicalwiththematrixproductof Section3.2.
2.9.7 Write ∇·∇×Aand∇×∇ϕintensor(index)notationin R3sothatitbecomesobvious
thateachexpressionvanishes.
ANS.∇·∇×A=εijk∂
∂xi∂
∂xjAk,
(∇×∇ϕ)i=εijk∂
∂xj∂
∂xkϕ.
2.9.8 ExpressingcrossproductsintermsofLevi-Civitasymbols (εijk),derivethe BAC–CAB
rule,Eq. (1.55).
Hint.TherelationofExercise2.9.4is helpful.
2.10 General Tensors 151
2.9.9 Verify that each of the following fourth-rank tensors is isotropic, that is, that it has the
sameformindependentofanyrotationofthecoordinatesystems.
(a)Aijkl=δijδkl,
(b)Bijkl=δikδjl+δilδjk,
(c)Cijkl=δikδjl−δilδjk.
2.9.10 Showthatthetwo-indexLevi-Civitasymbol εijisasecond-rankpseudotensor(intwo-
dimensionalspace).Doesthiscontradicttheuniquenessof δij(Exercise 2.6.4)?
2.9.11 Represent εijbya2×2matrix,andusingthe2 ×2rotationmatrixofSection3.3show
thatεijisinvariantunderorthogonalsimilaritytransformations.
2.9.12 GivenAk=1
2εijkBijwithBij=−Bji, antisymmetric,showthat
Bmn=εmnkAk.
2.9.13 Showthatthevectoridentity
(A×B)·(C×D)=(A·C)(B·D)−(A·D)(B·C)
(Exercise 1.5.12) follows directly from the description of a cross product with εijkand
theidentityof Exercise2.9.4.
2.9.14 Generalize the cross product of two vectors to n-dimensional space for n=4,5,....
Check the consistency of your construction and discuss concrete examples. See Exer-
cise1.4.17for thecase n=2.
2.10 G ENERAL TENSORS
The distinction between contravariant and covariant transformations was established in
Section 2.6. Then, for convenience, we restricted our attention to Cartesian coordinates
(in which the distinction disappears). Now in these two concluding sections we return to
non-Cartesiancoordinatesandresurrectthecontravariantandcovariantdependence.Asin
Section 2.6, a superscript will be used for an index denoting contravariant and a subscript
for an index denoting covariant dependence.The metric tensor of Section 2.1 will be used
torelatecontravariantandcovariantindices.
The emphasis in this section is on differentiation, culminating in the construction of
thecovariant derivative . We saw in Section 2.7 that the derivative of a vector yields a
second-rank tensor—in Cartesian coordinates. In non-Cartesian coordinate systems, it is
thecovariantderivativeofavectorratherthantheordinaryderivativethatyieldsasecond-
ranktensorbydifferentiationof avector.
Metric Tensor
Let us start with the transformation of vectors from one set of coordinates (q1,q2,q3)
to another r=(x1,x2,x3). The new coordinates are (in general nonlinear ) functions
152 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
xi(q1,q2,q3)of the old, such as spherical polar coordinates (r,θ,φ). But their differ-
entialsobeythelineartransformationlaw
dxi=∂xi
∂qjdqj, (2.113a)
or
dr=εjdqj(2.113b)
in vector notation. For convenience we take the basis vectors ε1=(∂x1
∂q1,∂x1
∂q2,∂x1
∂q3),ε2,
andε3to form a right-handed set. These vectors are not necessarily orthogonal. Also, a
limitation to three-dimensional space will be required only for the discussions of cross
products and curls. Otherwise these εimay be in N-dimensional space, including the
four-dimensional space–time of special and general relativity. The basis vectors εimay
beexpressedby
εi=∂r
∂qi, (2.114)
asinExercise2.2.3.Note,however,thatthe εiheredonotnecessarilyhaveunitmagnitude.
FromExercise2.2.3, theunitvectorsare
ei=1
hi∂r
∂qi(nosummation) ,
andtherefore
εi=hiei(nosummation) . (2.115)
Theεiarerelatedtotheunitvectors eibythescalefactors hiofSection2.2.The eihaveno
dimensions; the εihave the dimensions of hi. In spherical polar coordinates, as a specific
example,
εr=er=ˆr,εθ=reθ=rˆθ,εϕ=rsinθeϕ=rsinθˆϕ.(2.116)
In Euclidean spaces, or in Minkowski space of special relativity, the partial derivatives in
Eq.(2.113)areconstantsthatdefinethenewcoordinatesintermsoftheoldones.Weused
them to define the transformation laws of vectors in Eq. (2.59) and (2.62) and tensors in
Eq. (2.66). Generalizing, we define a contravariant vectorViundergeneralcoordinate
transformationsif itscomponentstransform accordingto
V′i=∂xi
∂qjVj, (2.117a)
or
V′=Vjεj (2.117b)
in vector notation. For covariant vectors we inspect the transformation of the gradient
operator
∂
∂xi=∂qj
∂xi∂
∂qj(2.118)
2.10 General Tensors 153
usingthechainrule. From
∂xi
∂qj∂qj
∂xk=δik (2.119)
itisclearthatEq.(2.118) isrelatedtothe inversetransformationofEq. (2.113),
dqj=∂qj
∂xidxi. (2.120)
Hencewedefinea covariant vectorViif
V′
i=∂qj
∂xiVj (2.121a)
holdsor, invectornotation,
V′=Vjεj, (2.121b)
where εjarethecontravariantvectors gjiεi=εj.
Second-ranktensorsaredefinedas inEq.(2.66),
A′ij=∂xi
∂qk∂xj
∂qlAkl, (2.122)
andtensors ofhigherranksimilarly.
AsinSection2.1, weconstructthesquareofadifferentialdisplacement
(ds)2=dr·dr=parenleftbig
εidqiparenrightbig2=εi·εjdqidqj. (2.123)
Comparing this with (ds)2of Section 2.1, Eq. (2.5), we identify εi·εjas the covariant
metrictensor
εi·εj=gij. (2.124)
Clearly,gijis symmetric. The tensor nature of gijfollows from the quotient rule, Exer-
cise2.8.1.We taketherelation
gikgkj=δij (2.125)
to define the corresponding contravariant tensor gik. Contravariant gikenters as the in-
verse20of covariant gkj. We use this contravariant gikto raise indices, converting a co-
variantindexintoacontravariantindex,asshownsubsequently.Likewisethecovariant gkj
willbeusedtolowerindices.Thechoiceof gikandgkjforthisraising–loweringoperation
isarbitrary. Anysecond-ranktensor(anditsinverse)woulddo.Specifically,wehave
gijεj=εirelatingcovariantand
contravariantbasisvectors,
gijFj=Firelatingcovariantand
contravariantvectorcomponents.(2.126)
20If thetensor gkjis written as amatrix, thetensor gikis given by the inverse matrix.
154 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Then
gijεj=εias thecorrespondingindex
gijFj=Filoweringrelations.(2.127)
It should be emphasized again that the εiandεjdonothave unit magnitude. This may
be seen in Eqs. (2.116) and in the metric tensor gijfor spherical polar coordinates and its
inversegij:
(gij)=
10 0
0r20
00r2sin2θ
parenleftbig
gijparenrightbig
=
10 0
01
r20
001
r2sin2θ
.
Christoffel Symbols
Letusform thedifferentialofascalar ψ,
dψ=∂ψ
∂qidqi. (2.128)
Since the dqiare the components of a contravariant vector, the partial derivatives
∂ψ/∂qimust form a covariant vector—by the quotient rule. The gradient of a scalar be-
comes
∇ψ=∂ψ
∂qiεi. (2.129)
Note that ∂ψ/∂qiare not the gradient components of Section 2.2—because εi/negationslash=eiof
Section2.2.
Movingontothederivativesofavector,wefindthatthesituationismuchmorecompli-
catedbecausethebasisvectors εiareingeneralnotconstant.Remember,wearenolonger
restricting ourselves to Cartesian coordinates and the nice, convenient ˆx,ˆy,ˆz! Direct dif-
ferentiationofEq. (2.117a)yields
∂V′k
∂qj=∂xk
∂qi∂Vi
∂qj+∂2xk
∂qj∂qiVi, (2.130a)
or,invectornotation,
∂V′
∂qj=∂Vi
∂qjεi+Vi∂εi
∂qj. (2.130b)
TherightsideofEq.(2.130a)differsfromthetransformationlawforasecond-rankmixed
tensor by the second term, which contains second derivatives of the coordinates xk.T h e
latterare nonzerofornonlinearcoordinatetransformations.
Now,∂εi/∂qjwillbesomelinearcombinationofthe εk,withthecoefficientdepending
on the indices iandjfrom the partial derivative and index kfrom the base vector. We
write
∂εi
∂qj=Ŵk
ijεk. (2.131a)
2.10 General Tensors 155
Multiplyingby εmandusing εm·εk=δm
kfrom Exercise2.10.2, wehave
Ŵm
ij=εm·∂εi
∂qj. (2.131b)
TheŴk
ijis a Christoffel symbol of the second kind . It is also called a coefficient of con-
nection. TheseŴk
ijarenotthird-rank tensors and the ∂Vi/∂qjof Eq. (2.130a) are not
second-ranktensors. Equations(2.131)shouldbecomparedwiththeresultsquotedinEx-
ercise2.2.3(rememberingthatingeneral εi/negationslash=ei).InCartesiancoordinates, Ŵk
ij=0forall
values of the indices i,j, andk. These Christoffel three-index symbols may be computed
by the techniques of Section 2.2. This is the topic of Exercise 2.10.8. Equation (2.138)
offers aneasiermethod.UsingEq.(2.114), weobtain
∂εi
∂qj=∂2r
∂qj∂qi=∂εj
∂qi=Ŵk
jiεk. (2.132)
HencetheseChristoffel symbolsaresymmetricinthetwolowerindices:
Ŵk
ij=Ŵk
ji. (2.133)
Christoffel Symbols as Derivatives of the Metric Tensor
ItisoftenconvenienttohaveanexplicitexpressionfortheChristoffelsymbolsintermsof
derivatives of the metric tensor. As an initial step, we define the Christoffel symbol of the
firstkind[ij,k]by
[ij,k]≡gmkŴm
ij, (2.134)
from which the symmetry [ij,k]=[ji,k]follows. Again, this [ij,k]is not a third-rank
tensor.FromEq. (2.131b),
[ij,k]=gmkεm·∂εi
∂qj
=εk·∂εi
∂qj. (2.135)
Nowwedifferentiate gij=εi·εj, Eq. (2.124):
∂gij
∂qk=∂εi
∂qk·εj+εi·∂εj
∂qk
=[ik,j]+[jk,i] (2.136)
byEq.(2.135). Then
[ij,k]=1
2braceleftbigg∂gik
∂qj+∂gjk
∂qi−∂gij
∂qkbracerightbigg
, (2.137)
and
Ŵs
ij=gks[ij,k]
=1
2gksbraceleftbigg∂gik
∂qj+∂gjk
∂qi−∂gij
∂qkbracerightbigg
. (2.138)
156 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
TheseChristoffel symbolsareappliedinthenextsection.
Covariant Derivative
WiththeChristoffelsymbols,Eq. (2.130b)mayberewritten
∂V′
∂qj=∂Vi
∂qjεi+ViŴk
ijεk. (2.139)
Now,iandkin the last term are dummy indices. Interchanging iandk(in this one term),
wehave
∂V′
∂qj=parenleftbigg∂Vi
∂qj+VkŴi
kjparenrightbigg
εi. (2.140)
Thequantityinparenthesisislabeleda covariantderivative ,Vi
;j.W eha v e
Vi
;j≡∂Vi
∂qj+VkŴi
kj. (2.141)
The;jsubscriptindicatesdifferentiationwithrespectto qj.Thedifferential dV′becomes
dV′=∂V′
∂qjdqj=[Vi
;jdqj]εi. (2.142)
A comparison with Eq. (2.113) or (2.122) shows that the quantity in square brackets is
theithcontravariantcomponentofavector.Since dqjisthejthcontravariantcomponent
of a vector (again, Eq. (2.113)), Vi
;jmust be the ijth componentof a (mixed) second-rank
tensor(quotientrule).Thecovariantderivativesofthecontravariantcomponentsofavector
formamixedsecond-ranktensor, Vi
;j.
Since the Christoffel symbols vanish in Cartesian coordinates, the covariant derivative
andtheordinarypartialderivativecoincide:
∂Vi
∂qj=Vi
;j(Cartesiancoordinates). (2.143)
Thecovariantderivativeof acovariantvector Viisgivenby(Exercise2.10.9)
Vi;j=∂Vi
∂qj−VkŴk
ij. (2.144)
LikeVi
;j,Vi;jis asecond-ranktensor.
The physical importance of the covariant derivative is that “A consistent replacement
of regular partial derivatives by covariant derivatives carries the laws of physics (in com-
ponent form) from flat space–time into the curved (Riemannian) space–time of general
relativity. Indeed, this substitution may be taken as a mathematicalstatement of Einstein’s
principleofequivalence.”21
21C. W.Misner, K.S.Thorne, andJ. A.Wheeler, Gravitation . San Francisco: W.H. Freeman (1973), p. 387.
2.10 General Tensors 157
Geodesics, Parallel Transport
The covariant derivative of vectors, tensors, and the Christoffel symbols may also be ap-
proached from geodesics. A geodesic in Euclidean space is a straight line. In general, it is
the curve of shortest length between two points and the curve along which a freely falling
particle moves. The ellipses of planets are geodesics around the sun, and the moon is in
free fall around the Earth on a geodesic. Since we can throw a particle in any direction, a
geodesic can have any direction through a given point. Hence the geodesic equation can
beobtainedfromFermat’svariationalprincipleofoptics(seeChapter17forEuler’sequa-
tion),
δintegraldisplay
ds=0, (2.145)
whereds2is themetric,Eq.(2.123), ofourspace.Usingthevariationof ds2,
2dsδds=dqidqjδgij+gijdqiδdqj+gijdqjδdqi(2.146)
inEq.(2.145) yields
1
2integraldisplaybracketleftbiggdqi
dsdqj
dsδgij+gijdqi
dsd
dsδdqj+gijdqj
dsd
dsδdqibracketrightbigg
ds=0,(2.147)
wheredsmeasuresthelengthonthegeodesic.Expressingthevariations
δgij=∂gij
∂qkδdqk≡(∂kgij)δdqk
in terms of the independent variations δdqk, shifting their derivatives in the other two
terms of Eq. (2.147) upon integrating by parts, and renaming dummy summation indices,
weobtain
1
2integraldisplaybracketleftbiggdqi
dsdqj
ds∂kgij−d
dsparenleftbigg
gikdqi
ds+gkjdqj
dsparenrightbiggbracketrightbigg
δdqkds=0. (2.148)
The integrand of Eq. (2.148), set equal to zero, is the geodesic equation. It is the Euler
equationof ourvariationalproblem.Uponexpanding
dgik
ds=(∂jgik)dqj
ds,dgkj
ds=(∂igkj)dqi
ds(2.149)
alongthegeodesicwefind
1
2dqi
dsdqj
ds(∂kgij−∂jgik−∂igkj)−gikd2qi
ds2=0. (2.150)
MultiplyingEq. (2.150)with gklandusingEq. (2.125),wefindthe geodesic equation
d2ql
ds2+dqi
dsdqj
ds1
2gkl(∂igkj+∂jgik−∂kgij)=0, (2.151)
wherethecoefficientofthevelocitiesis theChristoffelsymbol Ŵl
ijofEq. (2.138).
Geodesics are curves that are independent of the choice of coordinates. They can be
drawnthroughanypointinspaceinvariousdirections.Sincethelength dsmeasuredalong
158 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
thegeodesicisascalar,thevelocities dqi/ds(ofafreelyfallingparticlealongthegeodesic,
for example) form a contravariant vector. Hence Vkdqk/dsis a well-defined scalar on
any geodesic, which we can differentiate in order to define the covariant derivative of any
covariantvector Vk.UsingEq. (2.151)weobtainfromthescalar
d
dsparenleftbigg
Vkdqk
dsparenrightbigg
=dVk
dsdqk
ds+Vkd2qk
ds2
=∂Vk
∂qidqi
dsdqk
ds−VkŴk
ijdqi
dsdqj
ds(2.152)
=dqi
dsdqk
dsparenleftbigg∂Vk
∂qi−Ŵl
ikVlparenrightbigg
.
Whenthequotienttheoremis appliedtoEq.(2.152) ittellsusthat
Vk;i=∂Vk
∂qi−Ŵl
ikVl (2.153)
isacovarianttensorthatdefinesthecovariantderivativeof Vk,consistentwithEq.(2.144).
Similarly,higher-ordertensorsmaybederived.
ThesecondterminEq. (2.153)definesthe paralleltransportordisplacement ,
δVk=Ŵl
kiVlδqi, (2.154)
of the covariant vector Vkfrom the point with coordinates qitoqi+δqi. The parallel
transport, δUk,ofacontravariantvector Ukmaybefoundfromtheinvarianceofthescalar
productUkVkunderparalleltransport,
δ(UkVk)=δUkVk+UkδVk=0, (2.155)
inconjunctionwiththequotienttheorem.
Insummary,whenweshiftavectortoaneighboringpoint,paralleltransportpreventsit
fromstickingoutofourspace.Thiscanbeclearlyseenonthesurfaceofasphereinspher-
ical geometry, where a tangent vector is supposed to remain a tangent upon translating it
along some path on the sphere. This explains why the covariant derivative of a vector or
tensorisnaturallydefinedbytranslatingit alongageodesicinthedesireddirection.
Exercises
2.10.1 Equations (2.115) and (2.116) use the scale factor hi, citing Exercise 2.2.3. In Sec-
tion 2.2 we had restricted ourselves to orthogonal coordinate systems, yet Eq. (2.115)
holds for nonorthogonal systems. Justify the use of Eq. (2.115) for nonorthogonal sys-
tems.
2.10.2 (a) Showthat εi·εj=δi
j.
(b) Fromtheresultof part(a) showthat
Fi=F·εiandFi=F·εi.
2.10 General Tensors 159
2.10.3 For the special case of three-dimensional space ( ε1,ε2,ε3defining a right-handed co-
ordinatesystem,notnecessarilyorthogonal),showthat
εi=εj×εk
εj×εk·εi, i,j,k=1,2, 3andcyclicpermutations.
Note.These contravariant basis vectors εidefine the reciprocal lattice space of Sec-
tion1.5.
2.10.4 Provethatthecontravariantmetrictensoris givenby
gij=εi·εj.
2.10.5 If thecovariantvectors εiareorthogonal,showthat
(a)gijisdiagonal,
(b)gii=1/gii(nosummation),
(c)|εi|=1/|εi|.
2.10.6 Derive the covariant and contravariant metric tensors for circular cylindrical coordi-
nates.
2.10.7 Transformtheright-handsideofEq. (2.129),
∇ψ=∂ψ
∂qiεi,
into theeibasis, and verify that this expression agrees with the gradient developed in
Section2.2(for orthogonalcoordinates).
2.10.8 Evaluate ∂εi/∂qjfor spherical polar coordinates, and from these results calculate Ŵk
ij
for sphericalpolarcoordinates.
Note.Exercise2.5.2 offers a wayof calculatingtheneededpartialderivatives. Remem-
ber,
ε1=ˆrbut ε2=rˆθand ε3=rsinθˆϕ.
2.10.9 Showthatthecovariantderivativeof acovariantvectorisgivenby
Vi;j≡∂Vi
∂qj−VkŴk
ij.
Hint.Differentiate
εi·εj=δi
j.
2.10.10 Verifythat Vi;j=gikVk
;jbyshowingthat
∂Vi
∂qj−VsŴs
ij=gikbraceleftbigg∂Vk
∂qj+VmŴk
mjbracerightbigg
.
2.10.11 From the circular cylindrical metric tensor gij, calculate the Ŵk
ijfor circular cylindrical
coordinates.
Note.Thereareonlythreenonvanishing Ŵ.
160 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
2.10.12 Using the Ŵk
ijfrom Exercise 2.10.11, write out the covariant derivatives Vi
;jof a vector
Vincircularcylindricalcoordinates.
2.10.13 A triclinic crystal is described using an oblique coordinate system. The three covariant
basevectorsare
ε1=1.5ˆx,
ε2=0.4ˆx+1.6ˆy,
ε3=0.2ˆx+0.3ˆy+1.0ˆz.
(a) Calculatetheelementsof thecovariantmetrictensor gij.
(b) Calculate the Christoffel three-index symbols, Ŵk
ij. (This is a “by inspection” cal-
culation.)
(c) From the cross-product form of Exercise 2.10.3 calculate the contravariant base
vector ε3.
(d) Usingtheexplicitforms ε3andεi,verifythat ε3·εi=δ3i.
Note.If it were needed, the contravariant metric tensor could be determined by finding
theinverseof gijor byfindingthe εiandusing gij=εi·εj.
2.10.14 Verifythat
[ij,k]=1
2braceleftbigg∂gik
∂qj+∂gjk
∂qi−∂gij
∂qkbracerightbigg
.
Hint.SubstituteEq.(2.135)intotheright-handsideandshowthatanidentityresults.
2.10.15 Showthatfor themetrictensor gij;k=0,gij;k=0.
2.10.16 Show that parallel displacement δdqi=d2qialong a geodesic. Construct a geodesic
byparalleldisplacementof δdqi.
2.10.17 Construct the covariant derivative of a vector Viby parallel transport starting from the
limitingprocedure
lim
dqj→0Vi(qj+dqj)−Vi(qj)
dqj.
2.11 T ENSOR DERIVATIVE OPERATORS
InthissectionthecovariantdifferentiationofSection2.10isappliedtorederivethevector
differentialoperationsof Section2.2ingeneraltensorform.
Divergence
Replacingthepartialderivativebythecovariantderivative,wetakethedivergencetobe
∇·V=Vi
;i=∂Vi
∂qi+VkŴi
ik. (2.156)
2.11 Tensor Derivative Operators 161
Expressing Ŵi
ikbyEq.(2.138), wehave
Ŵi
ik=1
2gimbraceleftbigg∂gim
∂qk+∂gkm
∂qi−∂gik
∂qmbracerightbigg
. (2.157)
Whencontractedwith gimthelasttwotermsinthecurlybracketcancel,since
gim∂gkm
∂qi=gmi∂gki
∂qm=gim∂gik
∂qm. (2.158)
Then
Ŵi
ik=1
2gim∂gim
∂qk. (2.159)
Fromthetheoryofdeterminants,Section3.1,
∂g
∂qk=ggim∂gim
∂qk, (2.160)
wheregis the determinant of the metric, g=det(gij). Substituting this result into
Eq.(2.158), weobtain
Ŵi
ik=1
2g∂g
∂qk=1
g1/2∂g1/2
∂qk. (2.161)
Thisyields
∇·V=Vi
;i=1
g1/2∂
∂qkparenleftbig
g1/2Vkparenrightbig
. (2.162)
To compare this result with Eq. (2.21), note that h1h2h3=g1/2andVi(contravariant
coefficientof εi)=Vi/hi(no summation),where Viis Section2.2coefficientof ei.
Laplacian
In Section 2.2, replacement of the vector Vin∇·Vby∇ψled to the Laplacian ∇·∇ψ.
Herewehaveacontravariant Vi.Usingthemetrictensortocreateacontravariant ∇ψ,we
makethesubstitution
Vi→gik∂ψ
∂qk.
ThentheLaplacian ∇·∇ψbecomes
∇·∇ψ=1
g1/2∂
∂qiparenleftbigg
g1/2gik∂ψ
∂qkparenrightbigg
. (2.163)
Fortheorthogonal systemsofSection2.2themetrictensorisdiagonalandthecontravari-
antgii(nosummation)becomes
gii=(hi)−2.
162 Chapter 2 Vector Analysis in Curved Coordinates and Tensors
Equation(2.163)reducesto
∇·∇ψ=1
h1h2h3∂
∂qiparenleftbiggh1h2h3
h2
i∂ψ
∂qiparenrightbigg
,
inagreementwithEq.(2.22).
Curl
Thedifferenceofderivativesthatappearsinthecurl(Eq. (2.27)) willbewritten
∂Vi
∂qj−∂Vj
∂qi.
Again, remember that the components Vihere are coefficients of the contravariant
(nonunit)base vectors εi.T h eViof Section2.2 are coefficients of unit vectors ei. Adding
andsubtracting,weobtain
∂Vi
∂qj−∂Vj
∂qi=∂Vi
∂qj−VkŴk
ij−∂Vj
∂qi+VkŴk
ji
=Vi;j−Vj;i (2.164)
usingthesymmetryoftheChristoffelsymbols.Thecharacteristicdifferenceofderivatives
of the curl becomes a difference of covariant derivatives and therefore is a second-rank
tensor(covariantinbothindices).AsemphasizedinSection2.9,thespecialvectorformof
thecurlexistsonlyinthree-dimensionalspace.
From Eq. (2.138) it is clear that all the Christoffel three index symbols vanish in
Minkowskispaceandintherealspace–timeofspecialrelativitywith
gλµ=
1000
0−100
00−10
000 −1
.
Here
x0=ct, x 1=x, x 2=y,andx3=z.
Thiscompletesthedevelopmentofthedifferentialoperatorsingeneraltensorform.(The
gradient was given in Section 2.10.) In addition to the fields of elasticity and electromag-
netism, these differentials find application in mechanics (Lagrangian mechanics, Hamil-
tonianmechanics,andtheEulerequationsforrotationofrigidbody);fluidmechanics;and
perhapsmostimportantof all,thecurvedspace–timeofmoderntheoriesof gravity.
Exercises
2.11.1 VerifyEq. (2.160),
∂g
∂qk=ggim∂gim
∂qk,
for thespecificcaseof sphericalpolarcoordinates.
2.11 Additional Readings 163
2.11.2 Startingwiththedivergenceintensornotation,Eq.(2.162),developthedivergenceofa
vectorinsphericalpolarcoordinates,Eq. (2.47).
2.11.3 Thecovariantvector Aiisthegradientofascalar.Showthatthedifferenceofcovariant
derivatives Ai;j−Aj;ivanishes.
AdditionalReadings
Dirac,P.A.M., General Theory of Relativity .Princeton, NJ: Princeton University Press (1996).
Har tle,J.B., Gravity,San Francisco: Addison-Wesley (2003). This text uses aminimum of tensor analysis.
Jeffreys, H., Cartesian Tensors . Cambridge: Cambridge University Press (1952). This is an excellent discussion
of Cartesian tensors andtheirapplication to awide variety offields ofclassicalphysics.
La wd en ,D.F ., AnIntroduction to Tensor Calculus, Relativityand Cosmology , 3rd ed.NewYork: Wiley (1982).
Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nos-
trand (1956). Chapter5 covers curvilinear coordinates and 13 specificcoordinate systems.
Misner, C. W.,K.S. Thorne, andJ.A.Wheeler, Gravitation . San Francisco: W. H.Freeman (1973), p. 387.
Moller, C., The Theory of Relativity . Oxford: Oxford University Press (1955). Reprinted (1972). Most texts on
general relativity include a discussion of tensor analysis. Chapter 4 develops tensor calculus, including the
topicofdualtensors.Theextensiontonon-Cartesiansystems,asrequiredbygeneralrelativity,ispresentedin
Chapter 9.
Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 in-
cludesadescriptionofseveraldifferentcoordinatesystems.NotethatMorseandFeshbacharenotaboveusing
left-handedcoordinatesystemsevenforCartesiancoordinates.Elsewhereinthisexcellent(anddifficult)book
therearemanyexamplesoftheuseofthevariouscoordinatesystemsinsolvingphysicalproblems.Elevenad-
ditionalfascinatingbutseldom-encounteredorthogonalcoordinatesystemsarediscussedinthesecond(1970)
edition of Mathematical Methods for Physicists .
Ohanian, H. C., and R. Ruffini, Gravitation and Spacetime , 2nd ed. New York: Norton & Co. (1994). A well-
writtenintroduction to Riemannian geometry.
Sokolnikoff, I. S., Tensor Analysis—Theory and Applications , 2nd ed. New York: Wiley (1964). Particularly
useful for its extension of tensor analysis to non-Euclidean geometries.
Weinberg, S., Gravitation and Cosmology. Principles and Applications of the General Theory of Relativity .Ne w
York: Wiley (1972). This book and the one by Misner, Thorne, and Wheeler are the two leading texts on
general relativity and cosmology (withtensors in non-Cartesian space).
Young, E.C., Vectorand Tensor Analysis , 2nd ed.NewYork: MarcelDekker(1993).
This page intentionally left blank
CHAPTER 3
DETERMINANTS AND
MATRICES
3.1 D ETERMINANTS
We begin the study of matrices by solving linear equations that will lead us to determi-
nants and matrices. The concept of determinant and the notation were introduced by the
renownedGermanmathematicianandphilosopherGottfriedWilhelmvonLeibniz.
Homogeneous Linear Equations
One of the major applications of determinants is in the establishment of a condition for
the existence of a nontrivial solution for a set of linear homogeneous algebraic equations.
Supposewehavethreeunknowns x1,x2,x3(ornequationswith nunknowns):
a1x1+a2x2+a3x3=0,
b1x1+b2x2+b3x3=0, (3.1)
c1x1+c2x2+c3x3=0.
The problem is to determine under what conditions there is any solution, apart from
the trivial one x1=0,x2=0,x3=0. If we use vector notation x=(x1,x2,x3)for the
solution and three rows a=(a1,a2,a3),b=(b1,b2,b3),c=(c1,c2,c3)of coefficients,
thenthethreeequations,Eqs. (3.1), become
a·x=0,b·x=0,c·x=0. (3.2)
These three vector equations have the geometrical interpretation that xis orthogonalto
a,b, andc. If the volume spanned by a,b,cgiven by the determinant (or triple scalar
165
166 Chapter 3 Determinants and Matrices
product,seeEq. (1.50) ofSection1.5)
D3=(a×b)·c=det(a,b,c)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3
b1b2b3
c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(3.3)
isnotzero,thenthereis onlythetrivialsolution x=0.
Conversely, if the aforementioned determinant of coefficients vanishes, then one of the
row vectors is a linear combination of the other two. Let us assume that clies in the plane
spanned by aandb, that is, that the third equation is a linear combination of the first
two and not independent. Then xis orthogonal to that plane so that x∼a×b. Since
homogeneous equations can be multiplied by arbitrary numbers, only ratios of the xiare
relevant,forwhichwethenobtainratiosof 2 ×2 determinants
x1
x3=a2b3−a3b2
a1b2−a2b1
x2
x3=−a1b3−a3b1
a1b2−a2b1(3.4)
from the components of the cross product a×b, provided x3∼a1b2−a2b1/negationslash=0. This is
Cramer’srule for threehomogeneouslinearequations.
Inhomogeneous Linear Equations
Thesimplestcaseof twoequationswithtwounknowns,
a1x1+a2x2=a3,b 1x1+b2x2=b3, (3.5)
can be reduced to the previous case by imbedding it in three-dimensional space with a so-
lutionvector x=(x1,x2,−1)androwvectors a=(a1,a2,a3),b=(b1,b2,b3).Asbefore,
Eqs.(3.5)invectornotation, a·x=0 andb·x=0,implythat x∼a×b,sotheanalogof
Eqs. (3.4) holds. For this to apply, though, the third component of a×bmust not be zero,
that is,a1b2−a2b1/negationslash=0, because the third component of xis−1/negationslash=0. This yields the xi
as
(3.6a) x1=a3b2−b3a2
a1b2−a2b1=vextendsinglevextendsinglevextendsinglevextendsinglea3a2
b3b2vextendsinglevextendsinglevextendsinglevextendsingle
vextendsinglevextendsinglevextendsinglevextendsinglea1a2
b1b2vextendsinglevextendsinglevextendsinglevextendsingle,
x2=a1b3−a3b1
a1b2−a2b1=vextendsinglevextendsinglevextendsinglevextendsinglea1a3
b1b3vextendsinglevextendsinglevextendsinglevextendsingle
vextendsinglevextendsinglevextendsinglevextendsinglea1a2
b1b2vextendsinglevextendsinglevextendsinglevextendsingle. (3.6b)
The determinant in the numerator of x1(x2)is obtained from the determinant of the co-
efficientsvextendsinglevextendsinglea1a2
b2b2vextendsinglevextendsingleby replacing the first (second) column vector by the vectorparenleftbiga3
b3parenrightbig
of the
inhomogeneous side of Eq. (3.5). This is Cramer’s rule for a set of two inhomogeneous
linearequationswithtwounknowns.
3.1 Determinants 167
These solutions of linear equations in terms of determinants can be generalized to n
dimensions.Thedeterminantis asquarearray
Dn=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2···an
b1b2···bn
c1c2···cn
· · ··· ·vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(3.7)
of numbers (or functions), the coefficients of nlinear equations in our case here. The
numbernof columns (and of rows) in the array is sometimes called the orderof the
determinant. The generalization of the expansion in Eq. (1.48) of the triple scalar product
(of row vectors of three linear equations) leads to the following value of the determinant
Dninndimensions,
Dn=summationdisplay
i,j,k,...εijk···aibjck···, (3.8)
whereεijk···, analogous to the Levi-Civita symbol of Section 2.9, is +1 for even permuta-
tions1(ijk···)of(123···n),−1 for oddpermutations,andzeroif anyindexisrepeated.
Specifically,for thethird-orderdeterminant D3of Eq. (3.3), Eq.(3.8) leadsto
D3=+a1b2c3−a1b3c2−a2b1c3+a2b3c1+a3b1c2−a3b2c1.(3.9)
Thethird-orderdeterminant,then,isthisparticularlinearcombinationofproducts.Each
product contains one and only one element from each row and from each column. Each
product is added if the columns (indices) represent an even permutation of (123) and sub-
tracted if we have an odd permutation. Equation (3.3) may be considered shorthand no-
tation for Eq. (3.9). The number of terms in the sum (Eq. (3.8)) is 24 for a fourth-order
determinant, n!for annth-order determinant. Because of the appearance of the negative
signsinEq.(3.9)(andpossiblyintheindividualelementsaswell),theremaybeconsider-
able cancellation. It is quite possible that a determinant of large elements will have a very
smallvalue.
Several useful properties of the nth-order determinants follow from Eq. (3.8). Again, to
bespecific,Eq. (3.9) forthird-orderdeterminantsis usedtoillustratetheseproperties.
Laplacian Development by Minors
Equation(3.9) maybewritten
D3=a1(b2c3−b3c2)−a2(b1c3−b3c1)+a3(b1c2−b2c1)
=a1vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb2b3
c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−a2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb1b3
c1c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle+a3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb1b2
c1c2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.10)
In general, the nth-order determinant may be expanded as a linear combination of the
productsoftheelementsofanyrow(oranycolumn)andthe (n−1)th-orderdeterminants
1In a linear sequence abcd···, any single, simple transposition of adjacent elements yields an oddpermutation of the original
sequence: abcd→bacd.Twosuchtranspositionsyieldanevenpermutation.Ingeneral,anoddnumberofsuchinterchangesof
adjacentelements results in anodd permutation; an even number of suchtranspositions yields an even permutation.
168 Chapter 3 Determinants and Matrices
formedbystrikingouttherowandcolumnoftheoriginaldeterminantinwhichtheelement
appears.Thisreducedarray(2 ×2inthisspecificexample)iscalleda minor.Iftheelement
is in theith row and the jth column, the sign associated with the product is (−1)i+j.T h e
minorwiththissigniscalledthe cofactor.IfMijisusedtodesignatetheminorformedby
omitting the ith row and the jth column and Cijis the corresponding cofactor, Eq. (3.10)
becomes
D3=3summationdisplay
j=1(−1)j+1ajM1j=3summationdisplay
j=1ajC1j. (3.11)
In this case, expanding along the first row, we have i=1 and the summation over j,t h e
columns.
This Laplace expansion may be used to advantage in the evaluation of high-order de-
terminants in which a lot of the elements are zero. For example, to find the value of the
determinant
D=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0100
−10 0 0
0001
00−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.12)
weexpandacrossthetoprowtoobtain
D=(−1)1+2·(1)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−100
00 1
0−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.13)
Again,expandingacross thetoprow, weget
D=(−1)·(−1)1+1·(−1)vextendsinglevextendsinglevextendsinglevextendsingle01
−10vextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingle01
−10vextendsinglevextendsinglevextendsinglevextendsingle=1. (3.14)
(This determinant D(Eq. (3.12)) is formed from one of the Dirac matrices appearing in
Dirac’srelativisticelectrontheoryinSection3.4.)
Antisymmetry
The determinant changes sign if any two rows are interchanged or if any two columns are
interchanged. This follows from the even–odd character of the Levi-Civita εin Eq. (3.8)
orexplicitlyfrom theform ofEqs. (3.9) and(3.10).2
ThispropertywasusedinSection2.9todevelopatotallyantisymmetriclinearcombina-
tion.Itisalsofrequentlyusedinquantummechanicsintheconstructionofamany-particle
wavefunctionthat,inaccordancewiththePauliexclusionprinciple,willbeantisymmetric
under the interchange of any two identical spin1
2particles (electrons, protons, neutrons,
etc.).
2The sign reversal is reasonably obvious for the interchange of two adjacent rows (or columns), this clearly being an odd
permutation. Show that theinterchange of anytworows is still an odd permutation.
3.1 Determinants 169
•As a special case of antisymmetry, any determinant with two rows equal or two
columnsequalequalszero.
•If each element in a row or each element in a column is zero, the determinant is equal
tozero.
•If each element in a row or each element in a column is multiplied by a constant, the
determinantis multipliedbythatconstant.
•The value of a determinant is unchanged if a multiple of one row is added (column by
column)toanotherroworifamultipleofonecolumnisadded(rowbyrow)toanother
column.3
We have
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3
b1b2b3
c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1+ka2a2a3
b1+kb2b2b3
c1+kc2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.15)
UsingtheLaplacedevelopmentontheright-handside, weobtain
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1+ka2a2a3
b1+kb2b2b3
c1+kc2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3
b1b2b3
c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle+kvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea2a2a3
b2b2b3
c2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.16)
then by the property of antisymmetry the second determinant on the right-hand side of
Eq.(3.16) vanishes,verifyingEq. (3.15).
As a specialcase, adeterminantis equalto zeroif anytwo rows are proportionalor any
twocolumnsareproportional.
Some useful relations involving determinants or matrices appear in Exercises of Sec-
tions3.2and3.4.
Returning to the homogeneous Eqs. (3.1) and multiplying the determinant of the coef-
ficients by x1, then adding x2times the second column and x3times the third column, we
candirectlyestablishtheconditionfor thepresenceofa nontrivialsolutionforEqs. (3.1):
x1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3
b1b2b3
c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1x1a2a3
b1x1b2b3
c1x1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1x1+a2x2+a3x3a2a3
b1x1+b2x2+b3x3b2b3
c1x1+c2x2+c3x3c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle
=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0a2a3
0b2b3
0c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.17)
Therefore x1(andx2andx3) must be zero unless the determinant of the coefficients
vanishes.Conversely(seetextbelowEq.(3.3)),wecanshowthatifthedeterminantofthe
coefficientsvanishes,a nontrivialsolutiondoesindeedexist. Thisis usedin Section9.6to
establishthelineardependenceor independenceof asetoffunctions.
3Thisderivesfromthegeometricmeaningofthedeterminantasthevolumeoftheparallelepipedspannedbyitscolumnvectors.
Pulling it to the side without changing its height leavesthe volume unchanged.
170 Chapter 3 Determinants and Matrices
If our linear equations are inhomogeneous , that is, as in Eqs. (3.5) if the zeros on
the right-hand side of Eqs. (3.1) are replaced by a4,b4, andc4, respectively, then from
Eq.(3.17) weobtain,instead,
x1=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea4a2a3
b4b2b3
c4c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3
b1b2b3
c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.18)
whichgeneralizesEq.(3.6a)to n=3dimensions,etc.Ifthedeterminantofthecoefficients
vanishes,theinhomogeneoussetofequationshasnosolution—unlessthenumeratorsalso
vanish. In this case solutions may exist but they are not unique (see Exercise 3.1.3 for
aspecificexample).
Fornumericalwork,thisdeterminantsolution,Eq.(3.18),isexceedinglyunwieldy.The
determinant may involve large numbers with alternate signs, and in the subtraction of two
large numbers the relative error may soar to a point that makes the result worthless. Also,
although the determinant method is illustrated here with three equations and three un-
knowns, we might easily have 200 equations with 200 unknowns, which, involving up to
200! terms in each determinant, pose a challenge even to high-speed computers. There
mustbea betterway.
In fact, there are better ways. One of the best is a straightforward process often called
Gausselimination .Toillustratethis technique,considerthefollowingsetofequations.
Example 3.1.1 GAUSS ELIMINATION
Solve
3x+2y+z=11
2x+3y+z=13 (3.19)
x+y+4z=12.
Thedeterminantoftheinhomogeneouslinearequations(3.19) is18,so asolutionexists.
For convenience and for the optimum numerical accuracy, the equations are rearranged
sothatthelargestcoefficientsrunalongthemaindiagonal(upperlefttolowerright).This
hasalreadybeendoneintheprecedingset.
The Gauss technique is to use the first equation to eliminate the first unknown, x, from
the remaining equations. Then the (new) second equation is used to eliminate yfrom the
last equation. In general, we work down through the set of equations, and then, with one
unknowndetermined,weworkbackuptosolveforeachoftheotherunknownsinsucces-
sion.
Dividingeachrowbyits initialcoefficient,weseethatEqs. (3.19) become
x+2
3y+1
3z=11
3
x+3
2y+1
2z=13
2(3.20)
x+y+4z=12.
3.1 Determinants 171
Now,usingthefirst equation,weeliminate xfromthesecondandthirdequations:
x+2
3y+1
3z=11
3
5
6y+1
6z=17
6(3.21)
1
3y+11
3z=25
3
and
x+2
3y+1
3z=11
3
y+1
5z=17
5(3.22)
y+11z=25.
Repeating the technique, we use the new second equation to eliminate yfrom the third
equation:
x+2
3y+1
3z=11
3
y+1
5z=17
5(3.23)
54z=108,
or
z=2.
Finally,workingbackup, weget
y+1
5×2=17
5,
or
y=3.
Thenwith zandydetermined,
x+2
3×3+1
3×2=11
3,
and
x=1.
The technique may not seem so elegant as Eq. (3.18), but it is well adapted to computers
andis far faster thanthetimespentwithdeterminants.
This Gausstechniquemaybeusedtoconvertadeterminantintotriangularform:
D=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1b1c1
0b2c2
00c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle
forathird-orderdeterminantwhoseelementsarenottobeconfusedwiththoseinEq.(3.3).
In this form D=a1b2c3. For annth-order determinant the evaluation of the triangular
form requires only n−1 multiplications, compared with the n!required for the general
case.
172 Chapter 3 Determinants and Matrices
A variation of this progressive elimination is known as Gauss–Jordan elimination. We
startaswiththeprecedingGausselimination,buteachnewequationconsideredisusedto
eliminate a variable from allthe other equations, not just those below it. If we had used
thisGauss–Jordanelimination,Eq.(3.23) wouldbecome
x+1
5z=7
5
y+1
5z=17
5(3.24)
z=2,
using the second equation of Eqs. (3.22) to eliminate yfrom both the first and third equa-
tions.ThenthethirdequationofEqs.(3.24)isusedtoeliminate zfromthefirstandsecond,
giving
x=1
y=3 (3.25)
z=2.
WereturntothisGauss–JordantechniqueinSection3.2for invertingmatrices.
Another technique suitable for computer use is the Gauss–Seidel iteration technique.
Eachtechniquehasits advantagesanddisadvantages.TheGauss andGauss–Jordanmeth-
ods may have accuracy problems for large determinants. This is also a problem for ma-
trix inversion (Section 3.2). The Gauss–Seidel method, as an iterative method, may have
convergence problems. The IBM Scientific Subroutine Package (SSP) uses Gauss and
Gauss–Jordan techniques. The Gauss–Seidel iterative method and the Gauss and Gauss–
Jordan elimination methods are discussed in considerable detail by Ralston and Wilf and
also by Pennington.4Computer codes in FORTRAN and other programming languages
and extensive literature for the Gauss–Jordan elimination and others are also given by
Pressetal.5/squaresolid
Linear Dependence of Vectors
Twononzerotwo-dimensionalvectors
a1=parenleftbigga11
a12parenrightbigg
/negationslash=0,a2=parenleftbigga21
a22parenrightbigg
/negationslash=0
aredefinedtobe linearlydependent if twonumbers x1,x2canbefoundthatarenotboth
zero so that the linear relation x1a1+x2a2=0 holds. They are linearly independent if
x1=0=x2istheonlysolutionofthislinearrelation.WritingitinCartesiancomponents,
weobtaintwo homogeneouslinearequations
a11x1+a21x2=0,a 12x1+a22x2=0
4A. Ralston and H. Wilf, eds., Mathematical Methods for Digital Computers . New York: Wiley (1960); R. H. Pennington,
Introductory Computer Methods and Numerical Analysis . NewYork: Macmillan(1970).
5W. H. Press, B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes , 2nd ed. Cambridge, UK: Cambridge
University Press (1992), Chapter2.
3.1 Determinants 173
fromwhichweextractthefollowingcriterionforlinearindependenceoftwovectorsusing
Cramer’s rule. If a1,a2span a nonzero area , that is, their determinantvextendsinglevextendsinglea11a21a12a22vextendsinglevextendsingle/negationslash=0,
then the set of homogeneous linear equations has only the solution x1=0=x2.If
the determinant is zero ,then there is a nontrivial solution x1,x2, andour vectors are
linearly dependent . In particular, the unit vectors in the x- andy-directions are linearly
independent, the linear relation x1ˆx1+x2ˆx2=parenleftbigx1
x2parenrightbig
=parenleftbig0
0parenrightbig
having only the trivial solution
x1=0=x2.
Three or more vectors in two-dimensional space are always linearly dependent. Thus,
the maximum number of linearly independent vectors in two-dimensional space is 2. For
example,given a1,a2,a3,thelinearrelation x1a1+x2a2+x3a3=0alwayshasnontrivial
solutions.Ifoneofthevectorsiszero,lineardependenceisobviousbecausethecoefficient
ofthezerovectormaybechosentobenonzeroandthatoftheothersaszero.Soweassume
allofthemas nonzero.If a1anda2arelinearlyindependent,wewritethelinearrelation
a11x1+a21x2=−a31x3,a 12x1+a22x2=−a32x3,
asasetoftwoinhomogeneouslinearequationsandapplyCramer’srule.Sincethedetermi-
nantisnonzero,wecanfindanontrivialsolution x1,x2foranynonzero x3.Thisargument
goes through for any pair of linearly independent vectors. If all pairs are linearly depen-
dent, any of these linear relations is a linear relation among the three vectors, and we are
finished.Iftherearemorethanthreevectors,wepickanythreeofthemandapplythefore-
goingreasoningandputthecoefficientsoftheothervectors, xj=0,inthelinearrelation.
•Mutuallyorthogonalvectorsarelinearlyindependent.
Assume a linear relationsummationtext
icivi=0.Dottingvjinto this using vj·vi=0f o rj/negationslash=i,w e
obtaincjvj·vj=0,soevery cj=0 because v2
j/negationslash=0.
It is straightforward to extend these theorems to nor more vectors in n-dimensional
Euclidean space. Thus, the maximum number of linearly independent vectors in
n-dimensional space is n. The coordinate unit vectors are linearly independent be-
cause they span a nonzero parallelepiped in n-dimensional space and their determinant
isunity.
Gram–Schmidt Procedure
Inann-dimensionalvectorspacewithaninner(orscalar)product,wecanalwaysconstruct
anorthonormalbasisof nvectorswiwithwi·wj=δijstartingfrom nlinearlyindependent
vectorsvi,i=0,1,...,n−1.
We start by normalizing v0to unity, defining w0=v0√v02. Then we project v0fromv1,
formingu1=v1+a10w0, with the admixture coefficient a10chosen so that v0·u1=0.
Dottingv0intou1yieldsa10=−v0·v1radicalBig
v2
0=−v1·w0.Again,wenormalize u1definingw1=
u1radicalBig
u2
1. Here,u2
1/negationslash=0 because v0,v1arelinearlyindependent.Thisfirst stepgeneralizesto
uj=vj+aj0w0+aj1w1+···+ajj−1wj−1,
withcoefficients aji=−vj·wi. Normalizing wj=ujradicalBig
u2
jcompletesourconstruction.
174 Chapter 3 Determinants and Matrices
It will be noticed that although this Gram–Schmidt procedure is one possible way of
constructing an orthogonal or orthonormal set, the vectors wiare not unique. There is an
infinitenumberofpossibleorthonormalsets.
As an illustration of the freedom involved, consider two (nonparallel) vectors AandB
in thexy-plane. We may normalize Ato unit magnitude and then form B′=aA+Bso
thatB′is perpendicular to A. By normalizing B′we have completed the Gram–Schmidt
orthogonalizationfortwovectors.Butanytwoperpendicularunitvectors,suchas ˆxandˆy,
couldhavebeenchosenasourorthonormalset.Again,withaninfinitenumberofpossible
rotations ofˆxandˆyabout the z-axis, we have an infinite number of possible orthonormal
sets.
Example 3.1.2 VECTORS BY GRAM–SCHMIDT ORTHOGONALIZATION
Toillustratethemethod,weconsidertwovectors
v0=parenleftbigg1
1parenrightbigg
,v1=parenleftbigg1
−2parenrightbigg
,
which are neither orthogonal nor normalized. Normalizing the first vector w0=v0/√
2,
wethenconstruct u1=v1+a10w0so astobeorthogonalto v0. Thisyields
u1·v0=0=v1·v0+a10√
2v2
0=−1+a10√
2,
sotheadjustableadmixturecoefficient a10=1/√
2. Asaresult,
u1=parenleftbigg1
−2parenrightbigg
+1
2parenleftbigg1
1parenrightbigg
=3
2parenleftbigg1
−1parenrightbigg
,
sothesecondorthonormalvectorbecomes
w1=1√
2parenleftbigg1
−1parenrightbigg
.
We check that w0·w1=0. The two vectors w0,w1form an orthonormal set of vectors,
abasisof two-dimensionalEuclideanspace. /squaresolid
Exercises
3.1.1 Evaluatethefollowingdeterminants:
(a)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle101
010
100vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,(b)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle120
312
031vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,(c)1√
2vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0√
30 0√
30 2 0
020√
3
00√
30vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle.
3.1.2 Test theset oflinearhomogeneousequations
x+3y+3z=0,x−y+z=0,2x+y+3z=0
toseeif itpossessesa nontrivialsolution,andfindone.
3.1 Determinants 175
3.1.3 Giventhepairofequations
x+2y=3,2x+4y=6,
(a) Showthatthedeterminantofthecoefficientsvanishes.
(b) Showthatthenumeratordeterminants(Eq. (3.18)) alsovanish.
(c) Findatleasttwosolutions.
3.1.4 Expressthe components ofA×Bas2×2determinants.Thenshowthatthedotproduct
A·(A×B)yields a Laplacian expansionof a 3 ×3 determinant. Finally, note that two
rowsof the 3×3 determinantareidenticalandhencethat A·(A×B)=0.
3.1.5 IfCijisthecofactorofelement aij(formedbystrikingoutthe ithrowand jthcolumn
andincludingasign (−1)i+j),showthat
(a)summationtext
iaijCij=summationtext
iajiCji=|A|,where|A|isthedeterminantwiththeelements aij,
(b)summationtext
iaijCik=summationtext
iajiCki=0,j/negationslash=k.
3.1.6 A determinant with all elements of order unity may be surprisingly small. The Hilbert
determinant Hij=(i+j−1)−1,i , j=1,2,...,nis notoriousfor itssmallvalues.
(a) CalculatethevalueoftheHilbertdeterminantsoforder nforn=1,2,and3.
(b) If an appropriate subroutine is available, find the Hilbert determinants of order n
forn=4,5,and6.
ANS.nDet(Hn)
11.
28.33333×10−2
34.62963×10−4
41.65344×10−7
53.74930×10−12
65.36730×10−18
3.1.7 Solvethefollowingsetoflinearsimultaneousequations.Givetheresultstofivedecimal
places.
1.0x1+0.9x2+0.8x3+0.4x4+0.1x5=1.0
0.9x1+1.0x2+0.8x3+0.5x4+0.2x5+0.1x6=0.9
0.8x1+0.8x2+1.0x3+0.7x4+0.4x5+0.2x6=0.8
0.4x1+0.5x2+0.7x3+1.0x4+0.6x5+0.3x6=0.7
0.1x1+0.2x2+0.4x3+0.6x4+1.0x5+0.5x6=0.6
0.1x2+0.2x3+0.3x4+0.5x5+1.0x6=0.5.
Note.Theseequationsmayalsobesolvedbymatrixinversion,Section3.2.
176 Chapter 3 Determinants and Matrices
3.1.8 Solve the linear equations a·x=c,a×x+b=0f o rx=(x1,x2,x3)with constant
vectorsa/negationslash=0,bandconstant c.
ANS.x=c
a2a+(a×b)/a2.
3.1.9 Solvethelinearequations a·x=d,b·x=e,c·x=f,forx=(x1,x2,x3)withconstant
vectorsa,b,candconstants d,e,fsuchthat (a×b)·c/negationslash=0.
ANS.[(a×b)·c]x=d(b×c)+e(c×a)+f(a×b).
3.1.10 Expressinvectorformthesolution (x1,x2,x3)ofax1+bx2+cx3+d=0withconstant
vectorsa,b,c,dso that(a×b)·c/negationslash=0.
3.2 M ATRICES
Matrix analysis belongs to linear algebra because matrices are linear operators or maps
such as rotations. Suppose, for instance, we rotate the Cartesian coordinates of a two-
dimensionalspace,asinSection1.2, sothat,invectornotation,
parenleftbiggx′
1
x′
2parenrightbigg
=parenleftbiggx1cosϕ+x2sinϕ
−x2sinϕ+x2cosϕparenrightbigg
=parenleftbiggsummationtext
ja1jxjsummationtext
ja2jxjparenrightbigg
. (3.26)
We label the array of elementsparenleftbiga11a12a21a22parenrightbig
a2×2m a t r i xAconsisting of two rows and two
columns and consider the vectors x,x′as 2×1 matrices. We take the summation of
products in Eq. (3.26) as a definition of matrix multiplication involving the scalar
product of each row vector of Awith the column vector x. Thus, in matrix notation
Eq.(3.26) becomes
x′=Ax. (3.27)
Toextendthisdefinitionofmultiplicationofamatrixtimesacolumnvectortotheprod-
uctoftwo2×2matrices,letthecoordinaterotationbefollowedbyasecondrotationgiven
bymatrixBsuchthat
x′′=Bx′. (3.28)
Incomponentform,
x′′
i=summationdisplay
jbijx′
j=summationdisplay
jbijsummationdisplay
kajkxk=summationdisplay
kparenleftbiggsummationdisplay
jbijajkparenrightbigg
xk. (3.29)
Thesummationover jismatrixmultiplicationdefiningamatrix C=BAsuchthat
x′′
i=summationdisplay
kcikxk, (3.30)
orx′′=Cxin matrix notation. Again, this definition involves the scalar products of row
vectorsof Bwithcolumnvectorsof A.Thisdefinitionofmatrixmultiplicationgeneralizes
tom×nmatrices and is found useful; indeed, this usefulness is the justification for its
existence .Thegeometricalinterpretationisthatthematrixproductofthetwomatrices BA
is the rotation that carries the unprimed system directly into the double-primed coordinate
3.2 Matrices 177
system. Before passing to formal definitions, the your should note that operator Ais de-
scribedbyitseffectonthecoordinatesorbasisvectors.Thematrixelements aijconstitute
arepresentation of theoperator,arepresentationthatdependsonthechoiceof abasis.
The special case where a matrix has one column and nrows is called a column vector,
|x/angbracketright, with components xi,i=1,2,...,n.I fAis ann×nmatrix,|x/angbracketrightann-component
column vector, A|x/angbracketrightis defined as in Eqs. (3.27) and (3.26). Similarly, if a matrix has one
row and ncolumns, it is called a row vector, /angbracketleftx|with components xi,i=1,2,...,n.
Clearly,/angbracketleftx|resultsfrom|x/angbracketrightbyinterchangingrowsandcolumns,amatrixoperationcalled
transposition , and transposition for any matrix A,˜Ais called6“Atranspose” with matrix
elements (˜A)ik=Aki. Transposing a product of matrices ABreverses the order and gives
˜B˜A; similarly A|x/angbracketrighttranspose is/angbracketleftx|A. The scalar product takes the form /angbracketleftx|y/angbracketright=summationtext
ixiyi
(x∗
iinacomplexvectorspace).This Diracbra-ketnotation isusedinquantummechanics
extensivelyandinChapter10andheresubsequently.
More abstractly, we can define the dual space˜Vof linear functionals Fon a vector
spaceV, whereeachlinearfunctional Fof˜Vassignsanumber F(v)sothat
F(c1v1+c2v2)=c1F(v1)+c2F(v2)
for any vectors v1,v2from our vector space Vand numbers c1,c2. If we define the sum
oftwofunctionalsbylinearityas
(F1+F2)(v)=F1(v)+F2(v),
then˜Vis alinearspacebyconstruction.
Riesz’theorem saysthatthereisaone-to-onecorrespondencebetweenlinearfunction-
alsFin˜Vand vectors fin a vector space Vthat has an inner (or scalar) product /angbracketleftf|v/angbracketright
definedforanypairofvectors f,v.
The proof relies on the scalar product by defining a linear functional Ffor any vector f
ofVasF(v)=/angbracketleftf|v/angbracketrightforanyvofV.Thelinearityofthescalarproductin fshowsthatthese
functionals form a vector space (contained in ˜Vnecessarily). Note that a linear functional
iscompletelyspecifiedwhenitis definedfor everyvector vof agivenvectorspace.
On the other hand, starting from any nontrivial linear functional Fof˜Vwe now con-
struct a unique vector fofVso thatF(v)=f·vis given by an inner product. We start
from an orthonormal basis wiof vectors in Vusing the Gram–Schmidt procedure (see
Section 3.2). Take any vector vfromVand expand it as v=summationtext
iwi·vwi. Then the
linear functional F(v)=summationtext
iwi·vF(wi)is well defined on V. If we define the spe-
cific vector f=summationtext
iF(wi)wi, then its inner product with an arbitrary vector vis given
by/angbracketleftf|v/angbracketright=f·v=summationtext
iF(wi)wi·v=F(v),whichprovesRiesz’ theorem.
Basic Definitions
A matrix is defined as a square or rectangular array of numbers or functions that obeys
certain laws. This is a perfectly logical extension of familiar mathematical concepts. In
arithmeticwedealwithsinglenumbers.Inthetheoryofcomplexvariables(Chapter6)we
dealwithorderedpairsofnumbers, (1,2)=1+2i,inwhichtheorderingisimportant.We
6Some texts (including ours sometimes) denote Atranspose by AT.
178 Chapter 3 Determinants and Matrices
now consider numbers (or functions) ordered in a square or rectangular array. For conve-
nience in later work the numbers are distinguished by two subscripts, the first indicating
the row (horizontal) and the second indicating the column (vertical) in which the number
appears. For instance, a13is the matrix element in the first row, third column. Hence, if A
isamatrixwith mrowsand ncolumns,
A=
a11a12···a1n
a21a22···a2n
··· ··· ·
am1am2···amn
. (3.31)
Perhaps the most important fact to note is that the elements aijare not combined with
one another. A matrix is not a determinant. It is an ordered array of numbers, not a single
number.
Thematrix A,sofarjustanarrayofnumbers,hasthepropertiesweassigntoit.Literally,
this means constructing a new form of mathematics. We define that matrices A,B, andC,
withelements aij,bij,andcij, respectively,combineaccordingtothefollowingrules.
Rank
LookingbackatthehomogeneouslinearEqs.(3.1),wenotethatthematrixofcoefficients,
A, is made up of three row vectors that each represent one linear equation of the set. If
their triple scalar product is not zero, than they span a nonzero volume and are linearly
independent, and the homogeneous linear equations have only the trivial solution. In this
case the matrix is said to have rank3. Inndimensions the volume represented by the
triple scalar product becomes the determinant, det (A),for a square matrix. If det (A)/negationslash=0,
then×nmatrixAhasrankn. The case of Eqs. (3.1), where the vector clies in the plane
spanned by aandb, corresponds to rank 2 of the matrix of coefficients, because only two
of its row vectors ( a,bcorresponding to two equations) are independent. In general, the
rankrof a matrix is the maximal number of linearly independent row or column
vectorsithas,with 0≤r≤n.
Equality
MatrixA=MatrixBif and only if aij=bijfor all values of iandj. This, of course,
requiresthat AandBeachbem×narrays(mrows,ncolumns).
Addition, Subtraction
A±B=Cif and only if aij±bij=cijfor all values of iandj, the elements combining
according to the laws of ordinary algebra (or arithmetic if they are simple numbers). This
means that A+B=B+A, commutation. Also, an associative law is satisfied (A+B)+
C=A+(B+C). If all elements are zero, the matrix, called the null matrix , is denoted
byO.ForallA,
A+O=O+A=A,
3.2 Matrices 179
with
O=
000···
000···
000···
······
. (3.32)
Suchm×nmatricesforma linearspacewithrespecttoadditionandsubtraction.
Multiplication (by a Scalar)
Themultiplicationof matrix Abythescalarquantity αisdefinedas
αA=(αA), (3.33)
inwhichtheelementsof αAareαaij;thatis,eachelementofmatrix Aismultipliedbythe
scalarfactor.Thisisinstrikingcontrasttothebehaviorofdeterminantsinwhichthefactor
αmultiplies only one column or one row and not every element of the entire determinant.
Aconsequenceofthisscalarmultiplicationisthat
αA=Aα,commutation .
IfAis asquarematrix,then
det(αA)=αndet(A).
Matrix Multiplication, Inner Product
AB=Cifandonlyif7cij=summationdisplay
kaikbkj. (3.34)
Theijelementof Cis formed as a scalar productof the ith row ofAwith thejth column
ofB(which demands that Ahave the same number of columns ( n)a sBhas rows). The
dummyindex ktakesonallvalues 1 ,2,...,ninsuccession;thatis,
cij=ai1b1j+ai2b2j+ai3b3j (3.35)
forn=3. Obviously, the dummy index kmay be replaced by any other symbol that is
not already in use without altering Eq. (3.34). Perhaps the situation may be clarified by
stating that Eq. (3.34) defines the method of combining certain matrices. This method of
combination,togiveitalabel,is called matrixmultiplication .Toillustrate,considertwo
(so-calledPauli)matrices
σ1=parenleftbigg01
10parenrightbigg
andσ3=parenleftbigg10
0−1parenrightbigg
. (3.36)
7Some authors follow the summation convention here (compare Section 2.6).
180 Chapter 3 Determinants and Matrices
The11elementoftheproduct, (σ1σ3)11isgivenbythesumoftheproductsofelementsof
thefirstrowofσ1withthecorrespondingelementsof thefirst columnofσ3:
parenleftBigg
01
10parenrightBigg
10
0−1
→0·1+1·0=0.
Continuing,wehave
σ1σ3=parenleftbigg0·1+1·00·0+1·(−1)
1·1+0·01·0+0·(−1)parenrightbigg
=parenleftbigg0−1
10parenrightbigg
. (3.37)
Here
(σ1σ3)ij=σ1i1σ31j+σ1i2σ32j.
Directapplicationof thedefinitionofmatrixmultiplicationshowsthat
σ3σ1=parenleftbigg01
−10parenrightbigg
(3.38)
andbyEq. (3.37)
σ3σ1=−σ1σ3. (3.39)
Exceptinspecialcases, matrixmultiplicationisnotcommutative:8
AB/negationslash=BA. (3.40)
However,fromthedefinitionofmatrixmultiplicationwecanshow9thatanassociativelaw
holds,(AB)C=A(BC). Thereisalsoadistributivelaw, A(B+C)=AB+AC.
Theunitmatrix1haselements δij,Kroneckerdelta,andthepropertythat 1 A=A1=A
forallA,
1=
1000 ···
0100 ···
0010 ···
0001 ···
·······
. (3.41)
It should be noted that it is possible for the product of two matrices to be the null matrix
withouteitheronebeingthenullmatrix.Forexample,if
A=parenleftbigg11
00parenrightbigg
andB=parenleftbigg10
−10parenrightbigg
,
AB=O. This differs from the multiplication of real or complex numbers, which form
afield, whereas the additive and multiplicative structure of matrices is called a ringby
mathematicians. See also Exercise 3.2.6(a), from which it is evident that, if AB=0, at
8Commutationorthelackofitisconvenientlydescribedbythecommutatorbracketsymbol, [A,B]=AB−BA.Equation(3.40)
becomes[A,B]/negationslash=0.
9Notethatthebasicdefinitionsofequality,addition,andmultiplicationaregivenintermsofthematrixelements,the aij.Allour
matrix operations can be carried out in terms of the matrix elements. However, we can also treat a matrix as a single algebraic
operator, as in Eq. (3.40). Matrix elements and single operators each have their advantages, as will be seen in the following
section. Weshall use both approaches.
3.2 Matrices 181
leastoneof thematricesmust havea zerodeterminant(that is, be singularas definedafter
Eq.(3.50) inthissection).
IfAis ann×nmatrix with determinant |A|/negationslash=0, then it has a unique inverse A−1
satisfying AA−1=A−1A=1. IfBis also an n×nmatrix with inverse B−1, then the
productABhastheinverse
(AB)−1=B−1A−1(3.42)
becauseABB−1A−1=1=B−1A−1AB(see alsoExercises3.2.31and3.2.32).
Theproducttheorem ,whichsaysthatthedeterminantoftheproduct, |AB|,oftwon×n
matricesAandBisequaltotheproductofthedeterminants, |A||B|,linksmatriceswithde-
terminants. To prove this, consider the ncolumn vectors ck=(summationtext
jaijbjk,i=1,2,...,n)
of the product matrix C=ABfork=1,2,...,n. Eachck=summationtext
jkbjkkajkis a sum of n
column vectors ajk=(aijk,i=1,2,...,n). Note that we are now using a different prod-
uct summation index jkfor each column ck. Since any determinant D(b1a1+b2a2)=
b1D(a1)+b2D(a2)is linear in its column vectors, we can pull out the summation sign in
front of the determinant from each column vector in Ctogether with the common column
factorbjkksothat
|C|=summationdisplay
j′
ksbj11bj22···bjnndet(aj1aj2,...,ajn). (3.43)
Ifwerearrangethecolumnvectors ajkofthedeterminantfactorinEq.(3.43)intheproper
order, then we can pull the common factor det (a1,a2,...,an)=|A|in front of the nsum-
mationsignsinEq.(3.43).Thesecolumnpermutationsgeneratejusttherightsign εj1j2···jn
toproduceinEq.(3.43) theexpressioninEq. (3.8)for |B|so
|C|=|A|summationdisplay
j′
ksεj1j2···jnbj11bj22···bjnn=|A||B|, (3.44)
whichprovestheproducttheorem.
Direct Product
A second procedure for multiplying matrices, known as the directtensor or Kronecker
product,follows.If Aisanm×mmatrixand Bisann×nmatrix,thenthedirectproduct
is
A⊗B=C. (3.45)
Cisanmn×mnmatrixwithelements
Cαβ=AijBkl, (3.46)
with
α=m(i−1)+k, β=n(j−1)+l.
182 Chapter 3 Determinants and Matrices
Forinstance,if AandBare both 2×2 matrices,
A⊗B=parenleftbigga11Ba12B
a21Ba22Bparenrightbigg
=
a11b11a11b12a12b11a12b12
a11b21a11b22a12b21a12b22
a21b11a21b12a22b11a22b12
a21b21a21b22a22b21a22b22
. (3.47)
Thedirectproductisassociativebutnotcommutative.Asanexampleofthedirectprod-
uct, the Dirac matrices of Section 3.4 may be developed as direct products of the Pauli
matrices and the unit matrix. Other examples appear in the construction of groups (see
Chapter4) andinvectororHilbertspaceinquantumtheory.
Example 3.2.1 DIRECT PRODUCT OF VECTORS
Thedirectproductoftwotwo-dimensionalvectorsisa four-componentvector,
parenleftbiggx0
x1parenrightbigg
⊗parenleftbiggy0
y1parenrightbigg
=
x0y0
x0y1
x1y0
x1y1
;
whilethedirectproductof threesuchvectors,
parenleftbiggx0
x1parenrightbigg
⊗parenleftbiggy0
y1parenrightbigg
⊗parenleftbiggz0
z1parenrightbigg
=
x0y0z0
x0y0z1
x0y1z0
x0y1z1
x1y0z0
x1y0z1
x1y1z0
x1y1z1
,
isa(23=8)-dimensionalvector. /squaresolid
Diagonal Matrices
An important special type of matrix is the square matrix in which all the nondiagonal
elementsarezero.Specifically,if a 3 ×3m a t r i xAis diagonal,then
A=
a1100
0a220
00 a33
.
Aphysicalinterpretationofsuchdiagonalmatricesandthemethodofreducingmatricesto
thisdiagonalformareconsideredinSection3.5.Herewesimplynoteasignificantproperty
ofdiagonalmatrices—multiplicationof diagonalmatricesiscommutative,
AB=BA,ifAandBareeachdiagonal.
3.2 Matrices 183
Multiplication by a diagonal matrix [d1,d2,...,dn]that has only nonzero elements in the
diagonalis particularlysimple:
parenleftbigg10
02parenrightbiggparenleftbigg12
34parenrightbigg
=parenleftbigg12
2·32·4parenrightbigg
=parenleftbigg12
68parenrightbigg
;
whiletheoppositeordergives
parenleftbigg12
34parenrightbiggparenleftbigg10
02parenrightbigg
=parenleftbigg12·2
32·4parenrightbigg
=parenleftbigg14
38parenrightbigg
.
Thus,a diagonalmatrix does not commutewith anothermatrix unlessboth are diag-
onal,orthediagonalmatrixisproportionaltotheunitmatrix. Thisisborneoutbythe
moregeneralform
[d1,d2,...,dn]A=
d10···0
0d2···0
··· ··· ·
00···dn
a11a12···a1n
a21a22···a2n
··· ··· ·
an1an2···ann
=
d1a11d1a12···d1a1n
d2a21d2a22···d2a2n
··· ··· ·
dnan1dnan2···dnann
,
whereas
A[d1,d2,...,dn]=
a11a12···a1n
a21a22···a2n
··· ··· ·
an1an2···ann
d10···0
0d2···0
··· ··· ·
00···dn
=
d1a11d2a12···dna1n
d1a21d2a22···dna2n
··· ··· ·
d1an1d2an2···dnann
.
Herewehavedenotedby [d1,...,dn]adiagonalmatrixwithdiagonalelements d1,...,dn.
In the special case of multiplying two diagonal matrices, we simply multiply the corre-
spondingdiagonalmatrixelements,whichobviouslyiscommutative.
Trace
Inanysquarematrixthesumofthediagonalelementsis calledthe trace.
Clearlythetraceis alinearoperation:
trace(A−B)=trace(A)−trace(B).
184 Chapter 3 Determinants and Matrices
One of its interesting and useful properties is that the trace of a product of two matrices A
andBisindependentof theorder ofmultiplication:
trace(AB)=summationdisplay
i(AB)ii=summationdisplay
isummationdisplay
jaijbji
=summationdisplay
jsummationdisplay
ibjiaij=summationdisplay
j(BA)jj (3.48)
=trace(BA).
Thisholdseventhough AB/negationslash=BA.Equation(3.48)meansthatthetraceofanycommutator
[A,B]=AB−BAiszero.FromEq. (3.48)weobtain
trace(ABC)=trace(BCA)=trace(CAB),
which shows that the trace is invariant under cyclic permutation of the matrices in a prod-
uct.
For a real symmetric or a complex Hermitian matrix (see Section 3.4) the trace is the
sum, and the determinant the product, of its eigenvalues, and both are coefficients of the
characteristic polynomial. In Exercise 3.4.23 the operation of taking the trace selects one
term out of a sum of 16 terms. The trace will serve a similar function relative to matrices
asorthogonalityservesfor vectorsandfunctions.
Intermsoftensors(Section2.7)thetraceisacontractionand,likethecontractedsecond-
ranktensor,isa scalar(invariant).
Matrices are used extensively to represent the elements of groups (compare Exer-
cise3.2.7andChapter4).Thetraceofthematrixrepresentingthegroupelementisknown
in group theory as the character . The reason for the special name and special attention
is that, the trace or character remains invariant under similarity transformations (compare
Exercise3.3.9).
Matrix Inversion
Atthebeginningofthissectionmatrix Aisintroducedastherepresentationofanoperator
that (linearly) transforms the coordinate axes. A rotation would be one example of such
a linear transformation. Now we look for the inverse transformation A−1that will restore
theoriginalcoordinateaxes.This means,aseitheramatrixor anoperatorequation,10
AA−1=A−1A=1. (3.49)
With(A−1)ij≡a(−1)
ij,
a(−1)
ij≡Cji
|A|, (3.50)
10Hereandthroughoutthischapterourmatriceshavefiniterank.If Aisaninfinite-rankmatrix( n×nwithn→∞),thenlifeis
more difficult. For A−1to be theinverse wemust demand that both
AA−1=1andA−1A=1.
one relation no longer implies theother.
3.2 Matrices 185
withCjithe cofactor (see discussion preceding Eq. (3.11)) of aijand the assumption that
thedeterminantof A,|A|/negationslash=0.If itiszero, welabel Asingular.Noinverseexists.
There is a wide variety of alternative techniques. One of the best and most commonly
used is the Gauss–Jordanmatrix inversiontechnique.The theoryis based on the results of
Exercises3.2.34and3.2.35,whichshowthatthereexistmatrices MLsuchthattheproduct
MLAwillbeAbutwith
a. onerowmultipliedbyaconstant,or
b. onerowreplacedbytheoriginalrowminusamultipleof anotherrow, or
c. rows interchanged.
Other matrices MRoperating on the right (AMR)can carry out the same operations on
thecolumns ofA.
This means that the matrix rows and columns may be altered (by matrix multiplication)
as though we were dealing with determinants, so we can apply the Gauss–Jordan elimina-
tion techniques of Section 3.1 to the matrix elements. Hence there exists a matrix ML(or
MR)suchthat11
MLA=1. (3.51)
ThenML=A−1.Wedetermine MLbycarryingouttheidenticaleliminationoperationson
theunitmatrix.Then
ML1=ML. (3.52)
Toclarifythis,weconsidera specificexample.
Example 3.2.2 GAUSS–JORDAN MATRIX INVERSION
Wewanttoinvertthematrix
A=
321
231
114
. (3.53)
For convenience we write Aand 1 side by side and carry out the identical operations on
each:
321
231
114
and
100
010
001
. (3.54)
Tobesystematic,wemultiplyeachrowtoget ak1=1,
12
31
3
13
21
2
114
and
1
300
01
20
001
. (3.55)
11Remember that det (A)/negationslash=0.
186 Chapter 3 Determinants and Matrices
Subtractingthefirstrow fromthesecondandthirdrows,weobtain
12
31
3
05
61
6
01
311
3
and
1
300
−1
31
20
−1
301
. (3.56)
Then we divide the second row (of bothmatrices) by5
6and subtract2
3times it from the
firstrow and1
3timesitfrom thethirdrow. Theresults forbothmatricesare
101
5
011
5
0018
5
and
3
5−2
50
−2
53
50
−1
5−1
51
. (3.57)
We divide the third row (of bothmatrices) by18
5. Then as the last step1
5times the third
rowissubtractedfromeachof thefirst tworows(of bothmatrices).Ourfinalpairis
100
010
001
andA−1=
11
18−7
18−1
18
−7
1811
18−1
18
−1
18−1
185
18
. (3.58)
The check is to multiply the original Aby the calculated A−1to see if we really do get
theunitmatrix1. /squaresolid
AswiththeGauss–Jordansolutionofsimultaneouslinearalgebraicequations,thistech-
nique is well adapted to computers. Indeed, this Gauss–Jordan matrix inversion technique
will probably be available in the program library as a subroutine (see Sections 2.3 and 2.4
ofPressetal., loc.cit.).
Formatrices of special form, the inverse matrix can be given in closed form .F o r
example,for
A=
abc
bdb
cbe
, (3.59)
theinversematrixhasa similarbutslightlymoregeneralform,
A−1=
αβ1γ
β1δβ2
γβ2ǫ
, (3.60)
withmatrixelementsgivenby
Dα=ed−b2,Dγ=−parenleftbig
cd−b2parenrightbig
,Dβ1=(c−e)b, Dβ 2=(c−a)b,
Dδ=ae−c2,Dǫ=ad−b2,D=b2(2c−a−e)+dparenleftbig
ae−c2parenrightbig
,
whereD=det(A)isthedeterminantofthematrix A.Ife=ainA,thentheinversematrix
A−1alsosimplifiesto
β1=β2,ǫ=α, D=parenleftbig
a2−c2parenrightbig
d+2(c−a)b2.
3.2 Matrices 187
Asa check,letusworkoutthe 11-matrixelementoftheproduct AA−1=1.Wefind
aα+bβ1+cγ=1
Dbracketleftbig
aparenleftbig
ed−b2parenrightbig
+b2(c−e)−cparenleftbig
cd−b2parenrightbigbracketrightbig
=1
Dparenleftbig
−ab2+aed+2b2c−b2e−c2dparenrightbig
=D
D=1.
Similarlywecheckthatthe 12-matrixelementvanishes,
aβ1+bδ+cβ2=1
Dbracketleftbig
ab(c−e)+bparenleftbig
ae−c2parenrightbig
+cb(c−a)bracketrightbig
=0,
andso on.
Note though that we cannot always find an inverse of A−1by solving for the matrix
elements a,b,...ofA,becausenoteveryinversematrix A−1oftheforminEq.(3.60)has
acorresponding Aof thespecialform inEq. (3.59), asExample3.2.2clearlyshows.
Matricesaresquareorrectangulararraysofnumbersthatdefinelineartransformations,
suchasrotationsofacoordinatesystem.Assuch,theyarelinearoperators.Squarematri-
ces may be inverted when their determinant is nonzero. When a matrix defines a system of
linear equations, the inverse matrix solves it. Matrices with the same number of rows and
columns may be added and subtracted. They form what mathematicians call a ring with
a unit and a zero matrix. Matrices are also useful for representing group operations and
operatorsinHilbertspaces.
Exercises
3.2.1 Showthatmatrixmultiplicationisassociative, (AB)C=A(BC).
3.2.2 Showthat
(A+B)(A−B)=A2−B2
ifandonlyif AandBcommute,
[A,B]=0.
3.2.3 Showthatmatrix Aisalinearoperator byshowingthat
A(c1r1+c2r2)=c1Ar1+c2Ar2.
It can be shown that an n×nmatrix is the most general linear operator in an n-
dimensional vector space. This means that every linear operator in this n-dimensional
vectorspaceisequivalenttoamatrix.
3.2.4 (a) Complex numbers, a+ib, withaandbreal, may be represented by (or are iso-
morphicwith) 2 ×2 matrices:
a+ib↔parenleftbiggab
−baparenrightbigg
.
Showthatthismatrixrepresentationisvalidfor(i)additionand(ii)multiplication.
(b) Findthematrixcorrespondingto (a+ib)−1.
188 Chapter 3 Determinants and Matrices
3.2.5 IfAis ann×nmatrix,showthat
det(−A)=(−1)ndetA.
3.2.6 (a) The matrix equation A2=0 does not imply A=0. Show that the most general
2×2 matrixwhosesquareis zeromaybewrittenas
parenleftbiggab b2
−a2−abparenrightbigg
,
whereaandbarerealor complexnumbers.
(b) IfC=A+B, ingeneral
detC/negationslash=detA+detB.
Constructaspecificnumericalexampletoillustratethisinequality.
3.2.7 Giventhethreematrices
A=parenleftbigg−10
0−1parenrightbigg
,B=parenleftbigg01
10parenrightbigg
,C=parenleftbigg0−1
−10parenrightbigg
,
findallpossibleproductsof A,B,andC,twoatatime,includingsquares.Expressyour
answers in terms of A,B, andC, and 1, the unit matrix. These three matrices, together
with the unit matrix, form a representation of a mathematical group, the vierergruppe
(seeChapter4).
3.2.8 Given
K=
00 i
−i00
0−10
,
showthat
Kn=KKK···(nfactors)=1
(withtheproperchoiceof n,n/negationslash=0).
3.2.9 Verifythe Jacobiidentity ,
bracketleftbig
A,[B,C]bracketrightbig
=bracketleftbig
B,[A,C]bracketrightbig
−bracketleftbig
C,[A,B]bracketrightbig
.
This is useful in matrix descriptions of elementary particles (see Eq. (4.16)). As a
mnemonic aid, the you might note that the Jacobi identity has the same form as the
BAC–CABruleof Section1.5.
3.2.10 Showthatthematrices
A=
010
000
000
,B=
000
001
000
,C=
001
000
000
satisfythecommutationrelations
[A,B]=C,[A,C]=0,and[B,C]=0.
3.2 Matrices 189
3.2.11 Let
i=
0100
−1 000
0001
00−10
,j=
00 0−1
00−10
01 0 0
10 0 0
,
and
k=
00−10
00 01
10 00
0−100
.
Showthat
(a)i2=j2=k2=−1,where1is theunitmatrix.
(b)ij=−ji=k,
jk=−kj=i,
ki=−ik=j.
These three matrices ( i,j, andk) plus the unit matrix 1 form a basis for quaternions .
An alternate basis is provided by the four 2 ×2 matrices, iσ1,iσ2,−iσ3, and 1, where
theσarethePaulispinmatricesofExercise3.2.13.
3.2.12 A matrix with elements aij=0f o rj<imay be called upper right triangular. The
elementsinthelowerleft(belowandtotheleftofthemaindiagonal)vanish.Examples
are the matrices in Chapters 12 and 13, Exercise 13.1.21, relating power series and
eigenfunctionexpansions.
Showthattheproductoftwoupperrighttriangularmatricesisanupperrighttriangular
matrix.
3.2.13 ThethreePaulispinmatricesare
σ1=parenleftbigg01
10parenrightbigg
,σ 2=parenleftbigg0−i
i0parenrightbigg
,andσ3=parenleftbigg10
0−1parenrightbigg
.
Showthat
(a)(σi)2=12,
(b)σjσk=iσl,(j,k,l)=(1,2,3),(2,3,1),(3,1,2)(cyclicpermutation),
(c)σiσj+σjσi=2δij12;12is the 2×2 unitmatrix.
ThesematriceswereusedbyPauliinthenonrelativistictheoryofelectronspin.
3.2.14 UsingthePauli σiofExercise3.2.13,showthat
(σ·a)(σ·b)=a·b12+iσ·(a×b).
Here
σ≡ˆxσ1+ˆyσ2+ˆzσ3,
aandbareordinaryvectors,and 1 2is the 2×2 unitmatrix.
190 Chapter 3 Determinants and Matrices
3.2.15 Onedescriptionof spin1particlesuses thematrices
Mx=1√
2
010
101
010
,My=1√
2
0−i0
i0−i
0i0
,
and
Mz=
10 0
00 0
00−1
.
Showthat
(a)[Mx,My]=iMz, and so on12(cyclic permutation of indices). Using the Levi-
CivitasymbolofSection2.9, wemaywrite
[Mp,Mq]=iεpqrMr.
(b)M2≡M2
x+M2
y+M2
z=213, where 1 3is the 3×3 unitmatrix.
(c)[M2,Mi]=0,
[Mz,L+]=L+,
[L+,L−]=2Mz,
where
L+≡Mx+iMy,
L−≡Mx−iMy.
3.2.16 RepeatExercise3.2.15usinganalternaterepresentation,
Mx=
00 0
00−i
0i0
,My=
00i
00 0
−i00
,
and
Mz=
0−i0
i00
000
.
InChapter4thesematricesappearasthe generators oftherotationgroup.
3.2.17 Showthatthematrix–vectorequation
parenleftbigg
M·∇+131
c∂
∂tparenrightbigg
ψ=0
reproduces Maxwell’s equations in vacuum. Here ψis a column vector with compo-
nentsψj=Bj−iEj/c,j=x,y,z.Mis a vector whose elements are the angular
momentum matrices of Exercise 3.2.16. Note that ε0µ0=1/c2,13is the 3×3 unit
matrix.
12[A,B]=AB−BA.
3.2 Matrices 191
From Exercise3.2.15(b),
M2ψ=2ψ.
A comparison with the Dirac relativistic electron equation suggests that the “particle”
of electromagnetic radiation, the photon, has zero rest mass and a spin of 1 (in units
ofh).
3.2.18 RepeatExercise3.2.15,usingthematricesfor aspinof 3 /2,
Mx=1
2
0√
30 0√
30 2 0
020√
3
00√
30
,My=i
2
0−√
30 0√
30−20
020 −√
3
00√
30
,
and
Mz=1
2
30 0 0
01 0 0
00−10
00 0−3
.
3.2.19 Anoperator Pcommuteswith JxandJy,thexandycomponentsofanangularmomen-
tum operator. Show that Pcommutes with the third component of angular momentum,
thatis, that
[P,Jz]=0.
Hint.The angular momentum components must satisfy the commutation relation of
Exercise3.2.15(a).
3.2.20 TheL+andL−matrices of Exercise 3.2.15 are ladder operators (see Chapter 4): L+
operatingonasystemofspinprojection mwillraisethespinprojectionto m+1ifmis
below its maximum. L+operating on mmaxyields zero. L−reduces the spin projection
inunitsteps inasimilarfashion.Dividingby√
2,wehave
L+=
010
001
000
,L−=
000
100
010
.
Showthat
L+|−1/angbracketright=|0/angbracketright,L−|−1/angbracketright=nullcolumnvector,
L+|0/angbracketright=|1/angbracketright,L−|0/angbracketright=|−1/angbracketright,
L+|1/angbracketright=nullcolumnvector ,L−|1/angbracketright=|0/angbracketright,
where
|−1/angbracketright=
0
0
1
,|0/angbracketright=
0
1
0
,and|1/angbracketright=
1
0
0
representstatesofspinprojection −1,0,and1,respectively.
Note.Differentialoperatoranalogsof theseladderoperatorsappearinExercise12.6.7.
192 Chapter 3 Determinants and Matrices
3.2.21 VectorsAandBare relatedbythetensor T,
B=TA.
GivenAandB, show thatthere is no unique solution for the componentsof T.T h i si s
why vector division B/Ais undefined (apart from the special case of AandBparallel
andTthenascalar).
3.2.22 Wemightask foravector A−1, aninverseof agivenvector Ainthesensethat
A·A−1=A−1·A=1.
Show that this relation does not suffice to define A−1uniquely; Awould then have an
infinitenumberofinverses.
3.2.23 IfAis a diagonal matrix, with all diagonal elements different, and AandBcommute,
showthatBis diagonal.
3.2.24 IfAandBarediagonal,showthat AandBcommute.
3.2.25 Showthat trace (ABC)=trace(CBA)if anytwoofthethreematricescommute.
3.2.26 Angularmomentummatricessatisfyacommutationrelation
[Mj,Mk]=iMl, j,k,l cyclic.
Showthatthetraceof eachangularmomentummatrixvanishes.
3.2.27 (a) Theoperatortracereplacesamatrix Abyitstrace;thatis,
trace(A)=summationdisplay
iaii.
Showthattraceis a linearoperator.
(b) Theoperatordetreplacesamatrix Abyitsdeterminant;thatis,
det(A)=determinantof A.
Showthat det is notalinearoperator.
3.2.28 AandBanticommute: BA=−AB.A l s o ,A2=1,B2=1. Show that trace (A)=
trace(B)=0.
Note.ThePauliandDirac(Section3.4)matricesarespecificexamples.
3.2.29 With|x/angbracketrightanN-dimensional column vector and /angbracketlefty|anN-dimensional row vector, show
that
traceparenleftbig
|x/angbracketright/angbracketlefty|parenrightbig
=/angbracketlefty|x/angbracketright.
Note.|x/angbracketright/angbracketlefty|means direct product of column vector |x/angbracketrightwith row vector /angbracketlefty|.T h er e s u l t
isa square N×Nmatrix.
3.2.30 (a) If two nonsingular matrices anticommute, show that the trace of each one is zero.
(Nonsingular meansthatthedeterminantofthematrixnonzero.)
(b) Fortheconditionsofpart(a)tohold, AandBmustben×nmatriceswith neven.
Showthatif nisodd,acontradictionresults.
3.2 Matrices 193
3.2.31 If amatrixhas aninverse,showthattheinverseisunique.
3.2.32 IfA−1haselements
parenleftbig
A−1parenrightbig
ij=a(−1)
ij=Cji
|A|,
whereCjiis thejithcofactorof|A|, showthat
A−1A=1.
HenceA−1istheinverseof A(if|A|/negationslash=0).
3.2.33 Showthat det A−1=(detA)−1.
Hint.ApplytheproducttheoremofSection3.2.
Note.If detAiszero,then Ahas noinverse. Aissingular.
3.2.34 Findthematrices MLsuchthattheproduct MLAwillbeAbutwith:
(a) theithrowmultipliedbya constant k( aij→kaij,j=1,2,3,...);
(b) the ith row replaced by the original ith row minus a multiple of the mth row
(aij→aij−Kamj,i=1,2,3,...);
(c) theithandmthrows interchanged (aij→amj,amj→aij,j=1,2,3,...).
3.2.35 Findthematrices MRsuchthattheproduct AMRwillbeAbutwith:
(a) theithcolumnmultipliedbyaconstant k( aji→kaji,j=1,2,3,...);
(b) the ith column replaced by the original ith column minus a multiple of the mth
column(aji→aji−kajm,j=1,2,3,...);
(c) theithandmthcolumnsinterchanged (aji→ajm,ajm→aji,j=1,2,3,...).
3.2.36 Findtheinverseof
A=
321
221
114
.
3.2.37 (a) RewriteEq.(2.4)ofChapter2(andthecorrespondingequationsfor dyanddz)as
asinglematrixequation
|dxk/angbracketright=J|dqj/angbracketright.
Jisa matrixofderivatives,the Jacobian matrix.Showthat
/angbracketleftdxk|dxk/angbracketright=/angbracketleftdqi|G|dqj/angbracketright,
withthemetric(matrix) Ghavingelements gijgivenbyEq. (2.6).
(b) Showthat
det(J)dq1dq2dq3=dxdydz,
with det(J)theusualJacobian.
194 Chapter 3 Determinants and Matrices
3.2.38 Matricesarefartoousefultoremaintheexclusivepropertyofphysicists.Theymayap-
pearwherevertherearelinearrelations.Forinstance,inastudyofpopulationmovement
theinitialfractionofafixedpopulationineachof nareas(orindustriesorreligions,etc.)
is representedbyan n-componentcolumnvector P. Themovementof peoplefrom one
area to another in a given time is described by an n×n(stochastic) matrix T.H e r eTij
is the fraction of the population in the jth area that moves to the ith area. (Those not
moving are covered by i=j.) WithPdescribing the initial population distribution, the
finalpopulationdistributionis givenbythematrixequation TP=Q.
Fromitsdefinition,summationtextn
i=1Pi=1.
(a) Showthatconservationof peoplerequiresthat
nsummationdisplay
i=1Tij=1,j=1,2,...,n.
(b) Provethat
nsummationdisplay
i=1Qi=1
continuestheconservationofpeople.
3.2.39 Given a 6×6m a t r i xAwith elements aij=0.5|i−j|,i=0,1,2,...,5;i=0,1,
2,...,5,findA−1. Listits matrixelementstofivedecimalplaces.
ANS.A−1=1
3
4−2 0000
−25−2 000
0−25−20 0
00−25−20
000 −25−2
0000 −24
.
3.2.40 Exercise3.1.7maybewritteninmatrixform:
AX=C.
FindA−1andcalculate XasA−1C.
3.2.41 (a) Writea subroutine thatwillmultiply complex matrices.Assumethatthecomplex
matricesareinageneralrectangularform.
(b) TestyoursubroutinebymultiplyingpairsoftheDirac4 ×4matrices,Section3.4.
3.2.42 (a) Write a subroutine that will call the complex matrix multiplication subroutine of
Exercise 3.2.41 and will calculate the commutator bracket of two complex matri-
ces.
(b) Test your complex commutator bracket subroutine with the matrices of Exer-
cise3.2.16.
3.2.43 Interpolatingpolynomial isthenamegiventothe (n−1)-degreepolynomialdetermined
by (and passing through) npoints,(xi,yi)with all the xidistinct. This interpolating
polynomialforms abasisfor numericalquadratures.
3.3 Orthogonal Matrices 195
(a) Show that the requirement that an (n−1)-degree polynomial in xpass through
each of the npoints(xi,yi)with allxidistinct leads to nsimultaneous equations
oftheform
n−1summationdisplay
j=0ajxj
i=yi,i=1,2,...,n.
(b) Write a computer program that will read in ndata points and return the ncoeffi-
cientsaj.Useasubroutinetosolvethesimultaneousequationsifsuchasubroutine
isavailable.
(c) Rewritetheset ofsimultaneousequationsas amatrixequation
XA=Y.
(d) Repeat the computer calculation of part (b), but this time solve for vector Aby
invertingmatrix X(again,usingasubroutine).
3.2.44 Acalculationofthevaluesofelectrostaticpotentialinsideacylinderleadsto
V(0.0)=52.640V(0.6)=25.844
V(0.2)=48.292V(0.8)=12.648
V(0.4)=38.270V(1.0)=0.0.
The problem is to determine the values of the argument for which V=10, 20, 30, 40,
and50.Express V(x)asaseriessummationtext5
n=0a2nx2n.(Symmetryrequirementsintheoriginal
problem require that V(x)be an even function of x.) Determine the coefficients a2n.
WithV(x)now a known function of x, find the root of V(x)−10=0, 0≤x≤1.
Repeatfor V(x)−20,andso on.
ANS.a0=52.640,
a2=−117.676,
V(0.6851)=20.
3.3 O RTHOGONAL MATRICES
Ordinary three-dimensional space may be described with the Cartesian coordinates
(x1,x2,x3). We consider a second set of Cartesian coordinates (x′
1,x′
2,x′
3), whose ori-
gin and handedness coincides with that of the first set but whose orientation is different
(Fig. 3.1). We can say that the primed coordinate axeshave been rotatedrelative to the
initial, unprimed coordinate axes. Since this rotation is a linearoperation, we expect a
matrixequationrelatingtheprimedbasistotheunprimedbasis.
This section repeats portions of Chapters 1 and 2 in a slightly different context and
with a different emphasis. Previously, attention was focused on the vector or tensor. In
the case of the tensor, transformation properties were strongly stressed and were critical.
Here emphasis is placed on the description of the coordinate rotation itself—the matrix.
Transformation properties, the behavior of the matrix when the basis is changed, appear
at the end of this section. Sections 3.4 and 3.5 continue with transformation properties in
complexvectorspaces.
196 Chapter 3 Determinants and Matrices
FIGURE 3.1Cartesiancoordinatesystems.
Direction Cosines
A unit vector along the x′
1-axis(ˆx′
1)may be resolved into components along the x1-,x2-,
andx3-axesbytheusualprojectiontechnique:
ˆx′
1=ˆx1cos(x′
1,x1)+ˆx2cos(x′
1,x2)+ˆx3cos(x′
1,x3). (3.61)
Equation (3.61) is a specific example of the linear relations discussed at the beginning of
Section3.2.
Forconveniencethesecosines,whicharethedirectioncosines,arelabeled
cos(x′
1,x1)=ˆx′
1·ˆx1=a11,
cos(x′
1,x2)=ˆx′
1·ˆx2=a12, (3.62a)
cos(x′
1,x3)=ˆx′
1·ˆx3=a13.
Continuing,wehave
cos(x′
2,x1)=ˆx′
2·ˆx1=a21,
(3.62b)
cos(x′
2,x2)=ˆx′
2·ˆx2=a22,
andso on,where a21/negationslash=a12ingeneral.Now,Eq.(3.62) mayberewritten
ˆx′
1=ˆx1a11+ˆx2a12+ˆx3a13, (3.62c)
andalso
ˆx′
2=ˆx1a21+ˆx2a22+ˆx3a23,
(3.62d)
ˆx′
3=ˆx1a31+ˆx2a32+ˆx3a33.
3.3 Orthogonal Matrices 197
We may also go the other way by resolving ˆx1,ˆx2, andˆx3into components in the primed
system.Then
ˆx1=ˆx′
1a11+ˆx′
2a21+ˆx′
3a31,
ˆx2=ˆx′
1a12+ˆx′
2a22+ˆx′
3a32, (3.63)
ˆx3=ˆx′
1a13+ˆx′
2a23+ˆx′
3a33.
Associatingˆx1andˆx′
1with the subscript 1, ˆx2andˆx′
2with the subscript 2, ˆx3andˆx′
3
with the subscript 3, we see that in each case the first subscript of aijrefers to the primed
unit vector (ˆx′
1,ˆx′
2,ˆx′
3), whereas the second subscript refers to the unprimed unit vector
(ˆx1,ˆx2,ˆx3).
Applications to Vectors
If weconsideravectorwhosecomponentsarefunctionsofthepositioninspace,then
V(x1,x2,x3)=ˆx1V1+ˆx2V2+ˆx3V3,
(3.64)
V′(x′
1,x′
2,x′
3)=ˆx′
1V′
1+ˆx′
2V′
2+ˆx′
3V′
3,
since the point may be given both by the coordinates (x1,x2,x3)and by the coordinates
(x′
1,x′
2,x′
3).Notethat VandV′aregeometricallythesamevector(butwithdifferentcom-
ponents). The coordinate axes are being rotated; the vector stays fixed. Using Eqs. (3.62)
toeliminateˆx1,ˆx2, andˆx3,wemayseparateEq. (3.64)intothreescalarequations,
V′
1=a11V1+a12V2+a13V3,
V′
2=a21V1+a22V2+a23V3, (3.65)
V′
3=a31V1+a32V2+a33V3.
In particular, these relations will hold for the coordinates of a point (x1,x2,x3)and
(x′
1,x′
2,x′
3),gi ving
x′
1=a11x1+a12x2+a13x3,
x′
2=a21x1+a22x2+a23x3, (3.66)
x′
3=a31x1+a32x2+a33x3,
and similarly for the primed coordinates. In this notation the set of three equations (3.66)
maybewrittenas
x′
i=3summationdisplay
j=1aijxj, (3.67)
whereitakes onthevalues1,2, and3andtheresultisthree separate equations.
Now let us set aside these results and try a different approach to the same problem. We
consider two coordinate systems (x1,x2,x3)and(x′
1,x′
2,x′
3)with a common origin and
one point (x1,x2,x3)in the unprimed system, (x′
1,x′
2,x′
3)in the primed system. Note the
usual ambiguity. The same symbol xdenotes both the coordinate axis and a particular
198 Chapter 3 Determinants and Matrices
distance along that axis. Since our system is linear, x′
imust be a linear combination of
thexi.Let
x′
i=3summationdisplay
j=1aijxj. (3.68)
Theaijmay be identified as the direction cosines. This identification is carried out for the
two-dimensionalcaselater.
Ifwehavetwosetsofquantities (V1,V2,V3)intheunprimedsystemand (V′
1,V′
2,V′
3)in
theprimedsystem,relatedinthesamewayasthecoordinatesofapointinthetwodifferent
systems(Eq. (3.68)),
V′
i=3summationdisplay
j=1aijVj, (3.69)
then,asinSection1.2,thequantities (V1,V2,V3)aredefinedasthecomponentsofavector
that stays fixed while the coordinates rotate; that is, a vector is defined in terms of trans-
formation properties of its components under a rotation of the coordinate axes. In a sense
the coordinates of a point have been taken as a prototype vector. The power and useful-
ness of this definition became apparent in Chapter 2, in which it was extended to define
pseudovectorsandtensors.
From Eq. (3.67) we can derive interesting information about the aijthat describe the
orientation of coordinate system (x′
1,x′
2,x′
3)relative to the system (x1,x2,x3). The length
fromtheorigintothepointis thesameinbothsystems.Squaring,forconvenience,13
summationdisplay
ix2
i=summationdisplay
ix′2
i=summationdisplay
iparenleftbiggsummationdisplay
jaijxjparenrightbiggparenleftbiggsummationdisplay
kaikxkparenrightbigg
=summationdisplay
j,kxjxksummationdisplay
iaijaik. (3.70)
Thiscanbetrueforallpointsifandonlyif
summationdisplay
iaijaik=δjk,j,k=1,2,3. (3.71)
Note that Eq. (3.71) is equivalent to the matrix equation (3.83); see also Eqs. (3.87a)
to(3.87d).
Verification of Eq. (3.71), if needed, may be obtained by returning to Eq. (3.70) and
settingr=(x1,x2,x3)=(1,0,0),(0,1,0),(0,0,1),(1,1,0), and so on to evaluate the
ninerelationsgivenbyEq.(3.71).Thisprocessisvalid,sinceEq.(3.70)mustholdforall r
for a given set of aij. Equation (3.71), a consequence of requiring that the length remain
constant (invariant) under rotation of the coordinate system, is called the orthogonality
condition .Theaij,writtenasamatrix AsubjecttoEq.(3.71),formanorthogonalmatrix,
afirstdefinitionofanorthogonalmatrix.NotethatEq.(3.71)is notmatrixmultiplication.
Rather,itis interpretedlateras ascalarproductof twocolumnsof A.
13Note that twoindependent indices jandkareused.
3.3 Orthogonal Matrices 199
In matrixnotationEq.(3.67) becomes
|x′/angbracketright=A|x/angbracketright. (3.72)
Orthogonality Conditions — Two-Dimensional Case
Abetterunderstandingofthe aijandtheorthogonalityconditionmaybegainedbyconsid-
ering rotation in two dimensions in detail. (This can be thought of as a three-dimensional
systemwiththe x1-,x2-axesrotatedabout x3.) FromFig. 3.2,
x′
1=x1cosϕ+x2sinϕ,
x′
2=−x1sinϕ+x2cosϕ.(3.73)
ThereforebyEq.(3.72)
A=parenleftbiggcosϕsinϕ
−sinϕcosϕparenrightbigg
. (3.74)
Notice that Areduces to the unit matrix for ϕ=0. Zero angle rotation means nothing has
changed.Itis clearfrom Fig.3.2that
a11=cosϕ=cos(x′
1,x1),
(3.75)
a12=sinϕ=cosparenleftbigπ
2−ϕparenrightbig
=cos(x′
1,x2),
andsoon,thusidentifyingthematrixelements aijasthedirectioncosines.Equation(3.71),
theorthogonalitycondition,becomes
sin2ϕ+cos2ϕ=1,(3.76)sinϕcosϕ−sinϕcosϕ=0.
FIGURE 3.2Rotationof coordinates.
200 Chapter 3 Determinants and Matrices
Theextensiontothreedimensions(rotationofthecoordinatesthroughanangle ϕcoun-
terclockwiseabout x3)i ss i mp l y
A=
cosϕsinϕ0
−sinϕcosϕ0
00 1
. (3.77)
Thea33=1 expresses the fact that x′
3=x3, since the rotation has been about the x3-axis.
Thezerosguaranteethat x′
1andx′
2donotdependon x3andthatx′
3doesnotdependon x1
andx2.
Inverse Matrix, A−1
Returning to the general transformation matrix A, the inverse matrix A−1is defined such
that
|x/angbracketright=A−1|x′/angbracketright. (3.78)
That is,A−1describes the reverse of the rotation given by Aand returns the coordinate
systemtoitsoriginalposition.Symbolically,Eqs. (3.72)and(3.78)combinetogive
|x/angbracketright=A−1A|x/angbracketright, (3.79)
andsince|x/angbracketrightisarbitrary,
A−1A=1, (3.80)
theunitmatrix.Similarly,
AA−1=1, (3.81)
usingEqs. (3.72) and(3.78)andeliminating |x/angbracketrightinsteadof|x′/angbracketright.
Transpose Matrix, ˜A
We can determine the elements of our postulated inverse matrix A−1by employing the
orthogonalitycondition.Equation(3.71),theorthogonalitycondition,doesnotconformto
ourdefinitionofmatrixmultiplication,butitcanbeputintherequiredformby defininga
newmatrix˜Asuchthat
˜aji=aij. (3.82)
Equation(3.71)becomes
˜AA=1. (3.83)
This is a restatement of the orthogonality condition and may be taken as the constraint
defining an orthogonal matrix, a second definition of an orthogonal matrix. Multiplying
Eq.(3.83) by A−1fromtherightandusingEq.(3.81), wehave
˜A=A−1, (3.84)
3.3 Orthogonal Matrices 201
a third definition of an orthogonal matrix. This important result, that the inverse equals
the transpose, holds only for orthogonal matrices and indeed may be taken as a further
restatementoftheorthogonalitycondition.
MultiplyingEq.(3.84) by Afrom theleft, weobtain
A˜A=1 (3.85)
or
summationdisplay
iajiaki=δjk, (3.86)
whichisstillanotherformoftheorthogonalitycondition.
Summarizing,theorthogonalityconditionmaybestatedinseveralequivalentways:
summationdisplay
iaijaik=δjk, (3.87a)
summationdisplay
iajiaki=δjk, (3.87b)
˜AA=A˜A=1, (3.87c)
˜A=A−1. (3.87d)
Anyoneof theserelationsisanecessaryandasufficientconditionfor Atobeorthogonal.
It is now possible to see and understand why the term orthogonal is appropriate for
thesematrices.We havethegeneralform
A=
a11a12a13
a21a22a23
a31a32a33
,
a matrix of direction cosines in which aijis the cosine of the angle between x′
iandxj.
Therefore a11,a12,a13are the direction cosines of x′
1relative to x1,x2,x3. These three
elementsof Adefineaunitlengthalong x′
1, thatis, aunitvector ˆx′
1,
ˆx′
1=ˆx1a11+ˆx2a12+ˆx3a13.
The orthogonality relation (Eq. (3.86)) is simply a statement that the unit vectors ˆx′
1,ˆx′
2,
andˆx′
3aremutuallyperpendicular,ororthogonal.Ourorthogonaltransformationmatrix A
transforms one orthogonal coordinate system into a second orthogonal coordinate system
byrotationand/orreflection.
Asanexampleoftheuseofmatrices,theunitvectorsinsphericalpolarcoordinatesmay
bewrittenas
ˆr
ˆθ
ˆϕ
=C
ˆx
ˆy
ˆz
, (3.88)
202 Chapter 3 Determinants and Matrices
whereCis given in Exercise 2.5.1. This is equivalent to Eqs. (3.62) with x′
1,x′
2, andx′
3
replacedbyˆr,ˆθ,andˆϕ.Fromtheprecedinganalysis Cisorthogonal.Thereforetheinverse
relationbecomes
ˆx
ˆy
ˆz
=C−1
ˆr
ˆθ
ˆϕ
=˜C
ˆr
ˆθ
ˆϕ
, (3.89)
andExercise2.5.5issolvedbyinspection.Similarapplicationsofmatrixinversesappearin
connection with the transformation of a power series into a series of orthogonal functions
(Gram–Schmidt orthogonalization in Section 10.3) and the numerical solution of integral
equations.
Euler Angles
Our transformation matrix Acontains nine direction cosines. Clearly, only three of these
are independent, Eq. (3.71) providing six constraints. Equivalently, we may say that two
parameters ( θandϕin spherical polar coordinates) are required to fix the axis of rotation.
Then one additional parameter describes the amount of rotation about the specified axis.
(In the Lagrangian formulation of mechanics (Section 17.3) it is necessary to describe
Aby using some set of three independent parameters rather than the redundant direction
cosines.)TheusualchoiceofparametersistheEulerangles.14
The goal is to describe the orientation of a final rotated system (x′′′
1,x′′′
2,x′′′
3)relative to
some initial coordinate system (x1,x2,x3). The final system is developed in three steps,
witheachstep involvingonerotationdescribedbyoneEulerangle(Fig.3.3):
1. The coordinates are rotated about the x3-axis through an angle αcounterclockwise
intonewaxesdenotedby x′
1-,x′
2-,x′
3.( T hex3- andx′
3-axescoincide.)
FIGURE 3.3(a)Rotationabout x3throughangle α;(b)rotationabout x′
2through
angleβ; (c)rotationabout x′′
3throughangle γ.
14There are almost as many definitions of the Euler angles as there are authors. Here we follow the choice generally made by
workers in the areaof group theory and the quantum theory of angular momentum (compare Sections 4.3, 4.4).
3.3 Orthogonal Matrices 203
2. The coordinates are rotated about the x′
2-axis15through an angle βcounterclockwise
intonewaxesdenotedby x′′
1-,x′′
2-,x′′
3.( T hex′
2- andx′′
2-axescoincide.)
3. Thethirdandfinalrotationisthroughanangle γcounterclockwiseaboutthe x′′
3-axis,
yieldingthe x′′′
1,x′′′
2,x′′′
3system.(The x′′
3- andx′′′
3-axescoincide.)
Thethreematricesdescribingtheserotationsare
Rz(α)=
cosαsinα0
−sinαcosα0
00 1
, (3.90)
exactlylikeEq. (3.77),
Ry(β)=
cosβ0−sinβ
01 0
sinβ0 cosβ
(3.91)
and
Rz(γ)=
cosγsinγ0
−sinγcosγ0
00 1
. (3.92)
Thetotalrotationisdescribedbythetriplematrixproduct,
A(α,β,γ)=Rz(γ)Ry(β)Rz(α). (3.93)
Notetheorder: Rz(α)operatesfirst,then Ry(β),andfinally Rz(γ).Directmultiplication
gives
A(α,β,γ)
=
cosγcosβcosα−sinγsinαcosγcosβsinα+sinγcosα−cosγsinβ
−sinγcosβcosα−cosγsinα−sinγcosβsinα+cosγcosαsinγsinβ
sinβcosα sinβsinα cosβ
(3.94)
EquatingA(aij)withA(α,β,γ),elementbyelement,yieldsthedirectioncosinesinterms
ofthethreeEulerangles.WecouldusethisEulerangleidentificationtoverifythedirection
cosineidentities,Eq.(1.46)ofSection1.4,buttheapproachofExercise3.3.3ismuchmore
elegant.
Symmetry Properties
Our matrix description leads to the rotation group SO(3)in three-dimensional space R3,
and the Euler angle description of rotations forms a basis for developing the rotation
group in Chapter 4. Rotations may also be described by the unitary group SU(2)in two-
dimensional space C2over the complex numbers. The concept of groups such as SU(2)
and its generalizations and group theoretical techniques are often encountered in modern
15Some authors choose this second rotation to be about the x′
1-axis.
204 Chapter 3 Determinants and Matrices
particle physics, where symmetry properties play an important role. The SU(2)group is
alsoconsideredinChapter4.Thepowerandflexibilityofmatricespushedquaternionsinto
obscurityearlyinthe20thcentury.16
Itwillbenotedthatmatriceshavebeenhandledintwowaysintheforegoingdiscussion:
bytheircomponentsandassingleentities.Eachtechniquehasitsownadvantagesandboth
areuseful.
Thetransposematrixisusefulinadiscussionofsymmetryproperties.If
A=˜A,a ij=aji, (3.95)
thematrixiscalled symmetric ,whereasif
A=−˜A,a ij=−aji, (3.96)
it is called antisymmetric orskewsymmetric . The diagonal elements vanish. It is easy to
show that any (square) matrix may be written as the sum of a symmetric matrix and an
antisymmetricmatrix.Considertheidentity
A=1
2[A+˜A]+1
2[A−˜A]. (3.97)
[A+˜A]isclearly symmetric , whereas[A−˜A]isclearly antisymmetric .T h i si st h e
matrixanalogofEq.(2.75),Chapter2,fortensors.Similarly,afunctionmaybebrokenup
intoits evenandoddparts.
Sofarwehaveinterpretedtheorthogonalmatrixasrotatingthecoordinatesystem.This
changes the components of a fixed vector (not rotating with the coordinates) (Fig. 1.6,
Chapter1).However,anorthogonalmatrix Amaybeinterpretedequallywellasarotation
ofthevectorintheopposite direction(Fig.3.4).
These two possibilities, (1) rotating the vector keeping the coordinates fixed and (2)
rotating the coordinates (in the opposite sense) keeping the vector fixed, have a direct
analogy in quantum theory. Rotation (a time transformation) of the state vector gives the
Schrödingerpicture.RotationofthebasiskeepingthestatevectorfixedyieldstheHeisen-
bergpicture.
FIGURE 3.4Fixedcoordinates—
rotatedvector.
16R. J.Stephenson, Development of vector analysis from quaternions. Am.J .Ph ys. 34: 194 (1966).
3.3 Orthogonal Matrices 205
Supposeweinterpretmatrix Aas rotatinga vectorrintothepositionshownby r1; that
is, inaparticularcoordinatesystemwehavearelation
r1=Ar. (3.98)
Now let us rotate the coordinates by applying matrix B, which rotates (x,y,z)into
(x′,y′,z′),
r′
1=Br1=BAr=(Ar)′=BAparenleftbig
B−1Bparenrightbig
r
=parenleftbig
BAB−1parenrightbig
Br=parenleftbig
BAB−1parenrightbig
r′. (3.99)
Br1is justr1in the new coordinate system, with a similar interpretation holding for Br.
Henceinthisnewsystem (Br)isrotatedintoposition (Br1)bythematrix BAB−1:
Br1=(BAB−1)Br
r′
1=A′r′
In the new system the coordinates have been rotated by matrix B;Ahas the form A′,i n
which
A′=BAB−1. (3.100)
A′operatesinthe x′,y′,z′spaceasAoperatesinthe x,y,zspace.
The transformation defined by Eq. (3.100) with Bany matrix, not necessarily orthogo-
nal,isknownas a similaritytransformation .IncomponentformEq. (3.100)becomes
a′
ij=summationdisplay
k,lbikaklparenleftbig
B−1parenrightbig
lj. (3.101)
Now,ifBisorthogonal,
parenleftbig
B−1parenrightbig
lj=(˜B)lj=bjl, (3.102)
andwehave
a′
ij=summationdisplay
k,lbikbjlakl. (3.103)
Itmaybehelpfultothinkof Aagainasanoperator,possiblyasrotatingcoordinateaxes,
relatingangularmomentumandangularvelocityofarotatingsolid(Section3.5).Matrix A
istherepresentationinagivencoordinatesystem—orbasis.Buttherearedirectionsasso-
ciated with A—crystal axes, symmetry axes in the rotating solid, and so on—so that the
representation Adepends on the basis. The similarity transformation shows just how the
representationchangeswithachangeofbasis.
206 Chapter 3 Determinants and Matrices
Relation to Tensors
Comparing Eq. (3.103) with the equations of Section 2.6, we see that it is the definition
of a tensor of second rank. Hence a matrix that transforms by an orthogonal similarity
transformation is, by definition, a tensor. Clearly, then, any orthogonal matrixA, inter-
preted as rotating a vector (Eq. (3.98)), may be called a tensor. If, however, we consider
the orthogonal matrix as a collection of fixed direction cosines, giving the new orientation
ofacoordinatesystem,thereisnotensorpropertyinvolved.
Thesymmetryandantisymmetrypropertiesdefinedearlierarepreservedunder orthog-
onalsimilaritytransformations.Let Abeasymmetricmatrix, A=˜A, and
A′=BAB−1. (3.104)
Now,
˜A′=˜B−1˜A˜B=B˜AB−1, (3.105)
sinceBis orthogonal.But A=˜A. Therefore
˜A′=BAB−1=A′, (3.106)
showingthatthepropertyofsymmetryisinvariantunderanorthogonalsimilaritytransfor-
mation. In general, symmetry is notpreserved under a nonorthogonal similarity transfor-
mation.
Exercises
Note.Assumeallmatrixelementsarereal.
3.3.1 Showthattheproductoftwoorthogonalmatricesis orthogonal.
Note.This is a key step in showing that all n×northogonal matrices form a group
(Section4.1).
3.3.2 IfAis orthogonal,showthatits determinant =±1.
3.3.3 IfAisorthogonalanddet A=+1,showthat (detA)aij=Cij,whereCijisthecofactor
ofaij. This yields the identities of Eq. (1.46), used in Section 1.4 to show that a cross
productof vectors(inthree-space)is itselfavector.
Hint.NoteExercise3.2.32.
3.3.4 Anothersetof Eulerrotationsincommonuseis
(1) arotationaboutthe x3-axisthroughanangle ϕ, counterclockwise,
(2) arotationaboutthe x′
1-axisthroughanangle θ,counterclockwise,
(3) arotationaboutthe x′′
3-axisthroughanangle ψ, counterclockwise.
If
α=ϕ−π/2ϕ=α+π/2
β=θθ =β
γ=ψ+π/2ψ=γ−π/2,
showthatthefinalsystemsareidentical.
3.3 Orthogonal Matrices 207
3.3.5 Suppose the Earth is moved (rotated) so that the north pole goes to 30◦north, 20◦west
(originallatitudeandlongitudesystem)andthe10◦westmeridianpointsduesouth.
(a) WhataretheEuleranglesdescribingthisrotation?
(b) Findthecorrespondingdirectioncosines.
ANS.(b)A=
0.9551−0.2552−0.1504
0.0052 0 .5221−0.8529
0.2962 0 .8138 0 .5000
.
3.3.6 VerifythattheEuleranglerotationmatrix,Eq.(3.94),isinvariantunderthetransforma-
tion
α→α+π, β→−β, γ→γ−π.
3.3.7 ShowthattheEuleranglerotationmatrix A(α,β,γ) satisfiesthefollowingrelations:
(a)A−1(α,β,γ)=˜A(α,β,γ),
(b)A−1(α,β,γ)=A(−γ,−β,−α).
3.3.8 Showthatthetraceof theproductofasymmetricandanantisymmetricmatrixiszero.
3.3.9 Showthatthetraceof amatrixremainsinvariantundersimilaritytransformations.
3.3.10 Show that the determinant of a matrix remains invariant under similarity transforma-
tions.
Note.Exercises (3.3.9) and (3.3.10) show that the trace and the determinant are inde-
pendent of the Cartesian coordinates. They are characteristics of the matrix (operator)
itself.
3.3.11 Show that the property of antisymmetry is invariant under orthogonal similarity trans-
formations.
3.3.12 Ais 2×2 andorthogonal.Findthemostgeneralformof
A=parenleftbiggab
cdparenrightbigg
.
Comparewithtwo-dimensionalrotation.
3.3.13|x/angbracketrightand|y/angbracketrightare column vectors. Under an orthogonal transformation S,|x′/angbracketright=S|x/angbracketright,
|y′/angbracketright=S|y/angbracketright.Showthatthescalarproduct /angbracketleftx|y/angbracketrightisinvariantunderthisorthogonaltrans-
formation.
Note.Thisisequivalenttotheinvarianceofthedotproductoftwovectors,Section1.3.
3.3.14 Show that the sum of the squares of the elements of a matrix remains invariant under
orthogonalsimilaritytransformations.
3.3.15 AsageneralizationofExercise3.3.14, showthat
summationdisplay
jkSjkTjk=summationdisplay
l,mS′
lmT′
lm,
208 Chapter 3 Determinants and Matrices
where the primed and unprimed elements are related by an orthogonal similarity trans-
formation. This result is useful in deriving invariants in electromagnetic theory (com-
pareSection4.6).
Note.This product Mjk=summationtextSjkTjkis sometimes called a Hadamard product .I nt h e
framework of tensor analysis, Chapter 2, this exercise becomes a double contraction of
twosecond-ranktensors andthereforeisclearlyascalar(invariant).
3.3.16 A rotation ϕ1+ϕ2about the z-axis is carried out as two successive rotations ϕ1and
ϕ2, each about the z-axis. Use the matrix representation of the rotations to derive the
trigonometricidentities
cos(ϕ1+ϕ2)=cosϕ1cosϕ2−sinϕ1sinϕ2,
sin(ϕ1+ϕ2)=sinϕ1cosϕ2+cosϕ1sinϕ2.
3.3.17 Acolumnvector Vhascomponents V1andV2inaninitial(unprimed)system.Calculate
V′
1andV′
2for a
(a) rotationofthecoordinatesthroughanangleof θcounterclockwise ,
(b) rotationofthevectorthroughanangleof θclockwise .
Theresultsfor parts(a) and(b)shouldbeidentical.
3.3.18 Write a subroutine that will test whether a real N×Nmatrix is symmetric. Symmetry
maybedefinedas
0≤|aij−aji|≤ε,
whereεis some small tolerance (which allows for truncation error, and so on in the
computer).
3.4 H ERMITIAN MATRICES ,UNITARY MATRICES
Definitions
Thus far it has generally been assumed that our linear vector space is a real space and
that the matrix elements (the representations of the linear operators) are real. For many
calculations in classical physics, real matrix elements will suffice. However, in quantum
mechanics complex variables are unavoidable because of the form of the basic commuta-
tionrelations(ortheformofthetime-dependentSchrödingerequation).Withthisinmind,
we generalize to the case of complex matrix elements. To handle these elements, let us
define,or label,somenewproperties.
1. Complex conjugate, A∗, formed by taking the complex conjugate (i→−i)of each
element,where i=√
−1.
2. Adjoint, A†,formedbytransposing A∗,
A†=tildewiderA∗=˜A∗. (3.107)
3.4 Hermitian Matrices, Unitary Matrices 209
3. Hermitianmatrix:Thematrix Ais labeled Hermitian (orself-adjoint )if
A=A†. (3.108)
IfAis real, then A†=˜Aand real Hermitian matrices are real symmetric matrices.
In quantum mechanics (or matrix mechanics) matrices are usually constructed to be
Hermitian,or unitary.
4. Unitarymatrix:Matrix Uislabeled unitaryif
U†=U−1. (3.109)
IfUis real, then U−1=˜U, so real unitary matrices are orthogonal matrices. This
representsa generalizationof theconceptof orthogonalmatrix(compareEq. (3.84)).
5.(AB)∗=A∗B∗,(AB)†=B†A†.
If the matrix elements are complex, the physicist is almost always concerned with Her-
mitian and unitary matrices. Unitary matrices are especially important in quantum me-
chanicsbecausetheyleavethelengthofa(complex)vectorunchanged—analogoustothe
operation of an orthogonal matrix on a real vector. It is for this reason that the S matrix
of scattering theory is a unitary matrix. One important exception to this interest in unitary
matrices is the group of Lorentz matrices, Chapter 4. Using Minkowski space, we see that
thesematricesarenotunitary.
In a complex n-dimensional linear space the square of the length of a point ˜x=
xT(x1,x2,...,xn), or the square of its distance from the origin 0, is defined as x†x=summationtextx∗
ixi=summationtext|xi|2. If a coordinate transformation y=Uxleaves the distance unchanged,
thenx†x=y†y=(Ux)†Ux=x†U†Ux. Sincexis arbitrary it follows that U†U=1n;
that is,Uis a unitary n×nmatrix. If x′=Axis a linear map, then its matrix in the new
coordinatesbecomestheunitary(analogofasimilarity)transformation
A′=UAU†, (3.110)
becauseUx′=y′=UAx=UAU−1y=UAU†y.
Pauli and Dirac Matrices
Thesetofthree 2 ×2 Paulimatrices σ,
σ1=parenleftbigg01
10parenrightbigg
,σ 2=parenleftbigg0−i
i0parenrightbigg
,σ 3=parenleftbigg10
0−1parenrightbigg
, (3.111)
were introduced by W. Pauli to describe a particle of spin 1 /2 in nonrelativistic quantum
mechanics.Itcanreadilybeshownthat(compareExercises3.2.13and3.2.14)thePauli σ
satisfy
σiσj+σjσi=2δij12,anticommutation (3.112)
σiσj=iσk, i,j,k acyclicpermutationof 1,2, 3 (3.113)
(σi)2=12, (3.114)
210 Chapter 3 Determinants and Matrices
where 1 2is the 2×2 unit matrix. Thus, the vector σ/2 satisfies the same commutation
relations,
[σi,σj]≡σiσj−σjσi=2iεijkσk, (3.115)
as the orbital angular momentum L(L×L=iL, see Exercise 2.5.15 and the SO(3)and
SU(2)groupsinChapter4).
The three Pauli matrices σand the unit matrix form a complete set, so any Hermitian
2×2m a t r i xMmaybeexpandedas
M=m012+m1σ1+m2σ2+m3σ3=m0+m·σ, (3.116)
wherethe miformaconstantvector m.Using(σi)2=12andtrace(σi)=0weobtainfrom
Eq.(3.116) theexpansioncoefficients mibyformingtraces,
2m0=trace(M),2mi=trace(Mσi), i=1,2,3. (3.117)
Adding and multiplying such 2 ×2 matrices we generate the Pauli algebra.17Note that
trace(σi)=0f o ri=1,2,3.
In 1927 P. A. M. Dirac extended this formalism to fast-moving particles of spin1
2,
such as electrons (and neutrinos). To include special relativity he started from Einstein’s
energy,E2=p2c2+m2c4, instead of the nonrelativistic kinetic and potential energy,
E=p2/2m+V. ThekeytotheDiracequationistofactorize
E2−p2c2=E2−(cσ·p)2=(E−cσ·p)(E+cσ·p)=m2c4(3.118)
usingthe 2×2 matrixidentity
(σ·p)2=p212. (3.119)
The 2×2 unit matrix 1 2is not written explicitly in Eq. (3.118), and Eq. (3.119) follows
fromExercise3.2.14for a=b=p.Equivalently,wecanintroducetwomatrices γ′andγ
tofactorize E2−p2c2directly:
bracketleftbig
Eγ′⊗12−c(γ⊗σ)·pbracketrightbig2
=E2γ′2⊗12+c2γ2⊗(σ·p)2−Ec(γ′γ+γγ′)⊗σ·p
=E2−p2c2=m2c4. (3.119′)
ForEq.(3.119′) tohold,theconditions
γ′2=1=−γ2,γ′γ+γγ′=0 (3.120)
must be satisfied. Thus, the matrices γ′andγanticommute, just like the three Pauli ma-
trices; therefore they cannot be real or complex numbers. Because the conditions (3.120)
can be met by 2 ×2 matrices, we have written direct product signs (see Example 3.2.1) in
Eq.(3.119′) because γ′,γaremultipliedby 1 2,σmatrices,respectively,with
γ′=parenleftbigg10
0−1parenrightbigg
,γ=parenleftbigg01
−10parenrightbigg
. (3.121)
17For its geometrical significance, seeW. E.Baylis, J.Huschilt, andJiansu Wei, Am.J.Phys. 60: 788 (1992).
3.4 Hermitian Matrices, Unitary Matrices 211
The direct-product 4 ×4 matrices in Eq. (3 .119′) are the four conventional Dirac
γ-matrices,
γ0=γ′⊗12=parenleftbigg120
0−12parenrightbigg
=
10 0 0
01 0 0
00−10
00 0−1
,
γ1=γ⊗σ1=parenleftbigg0σ1
−σ10parenrightbigg
=
00 0 1
00 1 0
0−100
−100 0
,
γ3=γ⊗σ3=parenleftbigg0σ3
−σ30parenrightbigg
=
00 10
00 0−1
−100 0
01 00
, (3.122)
and similarly for γ2=γ⊗σ2. In vector notation γ=γ⊗σis a vector with three
components, each a 4 ×4 matrix, a generalization of the vector of Pauli matrices to a
vector of 4×4 matrices. The four matrices γiare the components of the four-vector
γµ=(γ0,γ1,γ2,γ3). If werecognizeinEq. (1 .119′)
Eγ′⊗12−c(γ⊗σ)·p=γµpµ=γ·p=(γ0,γ)·(E,cp) (3.123)
as a scalar product of two four-vectors γµandpµ(see Lorentz group in Chapter 4), then
Eq.(3.119′)withp2=p·p=E2−p2c2mayberegardedasafour-vectorgeneralization
ofEq. (3.119).
Summarizing the relativistic treatment of a spin 1/2particle, it leads to 4×4matrices,
whilethespin 1/2ofanonrelativisticparticleis describedbythe 2×2Paulimatrices σ.
By analogy with the Pauli algebra, we can form products of the basic γµmatrices
and linear combinations of them and the unit matrix 1 =14, thereby generating a 16-
dimensional (so-called Clifford18) algebra. A basis (with convenient Lorentz transforma-
tionproperties,see Chapter4) isgiven(in 2 ×2 matrixnotationofEq. (3.122))by
14,γ5=iγ0γ1γ2γ3=parenleftbigg012
120parenrightbigg
,γµ,γ5γµ,σµν=iparenleftbig
γµγν−γνγµparenrightbig
/2.(3.124)
Theγ-matricesanticommute;thatis, theirsymmetriccombinations
γµγν+γνγµ=2gµν14, (3.125)
whereg00=1=−g11=−g22=−g33, andgµν=0f o rµ/negationslash=ν, are zero or proportional
to the 4×4 unit matrix 1 4, while the six antisymmetric combinations in Eq. (3.124) give
new basis elements that transform like a tensor under Lorentz transformations (see Chap-
ter 4). Any 4×4 matrix can be expanded in terms of these 16 elements, and the expan-
sion coefficients are given by forming traces similar to the 2 ×2 case in Eq. (3.117) us-
18D.HestenesandG.Sobczyk, loc.cit.;D.Hestenes, Am.J.Phys. 39: 1013 (1971); and J. Math. Phys. 16: 556 (1975).
212 Chapter 3 Determinants and Matrices
ing trace(14)=4,trace(γ5)=0,trace(γµ)=0=trace(γ5γµ),trace(σµν)=0f o rµ,ν=
0,1,2,3 (see Exercise 3.4.23). In Chapter 4 we show that γ5is odd under parity, so γ5γµ
transformlikeanaxialvectorthathasevenparity.
The spin algebra generated by the Pauli matrices is just a matrix representation of the
four-dimensionalCliffordalgebra,whileHestenesandcoworkers(loc.cit.)havedeveloped
in theirgeometric calculus a representation-free (that is, “coordinate-free”) algebra that
containscomplexnumbers,vectors,thequaternionsubalgebra,andgeneralizedcrossprod-
uctsasdirectedareas(called bivectors).Thisalgebraic-geometricframeworkistailoredto
nonrelativisticquantummechanics,wherespinorsacquiregeometricaspectsandtheGauss
and Stokes theorems appear as components of a unified theorem. Their geometric algebra
corresponding to the 16-dimensional Clifford algebra of Dirac γ-matrices is the appropri-
atecoordinate-freeframework forrelativisticquantummechanicsandelectrodynamics.
The discussion of orthogonal matrices in Section 3.3 and unitary matrices in this sec-
tion is only a beginning. Further extensions are of vital concern in “elementary” particle
physics.WiththePauliandDiracmatrices,wecandevelop spinorwavefunctionsforelec-
trons,protons,andother(relativistic)spin1
2particles.Thecoordinatesystemrotationslead
toDj(α,β,γ), the rotation group usually represented by matrices in which the elements
are functions of the Euler angles describing the rotation. The special unitary group SU(3)
(composedof3 ×3unitarymatriceswithdeterminant +1)hasbeenusedwithconsiderable
successtodescribemesonsandbaryonsinvolvedinthestronginteractions,agaugetheory
that is now called quantum chromodynamics . These extensions are considered further in
Chapter4.
Exercises
3.4.1 Showthat
det(A∗)=(detA)∗=detparenleftbig
A†parenrightbig
.
3.4.2 Threeangularmomentummatricessatisfythebasiccommutationrelation
[Jx,Jy]=iJz
(andcyclicpermutationofindices).Iftwoofthematriceshaverealelements,showthat
theelementsofthethirdmustbepureimaginary.
3.4.3 Showthat (AB)†=B†A†.
3.4.4 Am a t r i xC=S†S. Show that the trace is positive definite unless Sis the null matrix,
inwhichcasetrace (C)=0.
3.4.5 IfAandBare Hermitian matrices, show that (AB+BA)andi(AB−BA)are also
Hermitian.
3.4.6 The matrix CisnotHermitian. Show that then C+C†andi(C−C†)are Hermitian.
Thismeansthata non-HermitianmatrixmayberesolvedintotwoHermitianparts,
C=1
2parenleftbig
C+C†parenrightbig
+1
2iiparenleftbig
C−C†parenrightbig
.
ThisdecompositionofamatrixintotwoHermitianmatrixpartsparallelsthedecompo-
sitionofacomplexnumber zintox+iy, wherex=(z+z∗)/2 andy=(z−z∗)/2i.
3.4 Hermitian Matrices, Unitary Matrices 213
3.4.7 AandBaretwononcommutingHermitianmatrices:
AB−BA=iC.
Provethat CisHermitian.
3.4.8 Show that a Hermitian matrix remains Hermitian under unitary similarity transforma-
tions.
3.4.9 Twomatrices AandBareeachHermitian.Findanecessaryandsufficientconditionfor
theirproduct ABtobeHermitian.
ANS.[A,B]=0.
3.4.10 Showthatthereciprocal(thatis,inverse)of aunitarymatrixisunitary.
3.4.11 Aparticularsimilaritytransformationyields
A′=UAU−1,
A′†=UA†U−1.
If the adjoint relationship is preserved (A†′=A′†)and detU=1, show that Umust be
unitary.
3.4.12 Twomatrices UandHarerelatedby
U=eiaH,
withareal. (The exponential function is defined by a Maclaurin expansion. This will
bedoneinSection5.6.)
(a) IfHisHermitian,showthat Uisunitary.
(b) IfUisunitary,showthat HisHermitian.( Hisindependentof a.)
Note.WithHtheHamiltonian,
ψ(x,t)=U(x,t)ψ(x, 0)=exp(−itH/¯h)ψ(x,0)
isasolutionofthetime-dependentSchrödingerequation. U(x,t)=exp(−itH/¯h)isthe
“evolutionoperator.”
3.4.13 Anoperator T(t+ε,t)describesthechangeinthewavefunctionfrom ttot+ε.F orε
realandsmallenoughsothat ε2maybeneglected,
T(t+ε,t)=1−i
¯hεH(t).
(a) IfTis unitary,showthat His Hermitian.
(b) IfHisHermitian,showthat Tis unitary.
Note.WhenH(t)isindependentoftime,thisrelationmaybeputinexponentialform—
Exercise3.4.12.
214 Chapter 3 Determinants and Matrices
3.4.14 Showthatanalternateform,
T(t+ε,t)=1−iεH(t)/2¯h
1+iεH(t)/2¯h,
agrees with the Tof part (a) of Exercise 3.4.13, neglecting ε2, and is exactly unitary
(forHHermitian).
3.4.15 Provethatthedirectproductoftwo unitarymatricesis unitary.
3.4.16 Showthat γ5anticommuteswithallfour γµ.
3.4.17 Use the four-dimensional Levi-Civita symbol ελµνρwithε0123=−1 (generalizing
Eqs. (2.93) in Section 2.9 to four dimensions) and show that (i) 2 γ5σµν=−iεµναβσαβ
using the summation convention of Section 2.6 and (ii) γλγµγν=gλµγν−gλνγµ+
gµνγλ+iελµνργργ5. Defineγµ=gµνγνusinggµν=gµνtoraiseandlowerindices.
3.4.18 Evaluatethefollowingtraces:(seeEq. (3.123)for thenotation)
(i) trace (γ·aγ·b)=4a·b,
(ii) trace (γ·aγ·bγ·c)=0,
(iii) trace (γ·aγ·bγ·cγ·d)=4(a·bc·d−a·cb·d+a·db·c),
(iv) trace (γ5γ·aγ·bγ·cγ·d)=4iεαβµνaαbβcµdν.
3.4.19 Show that (i) γµγαγµ=−2γα, (ii)γµγαγβγµ=4gαβ, and (iii) γµγαγβγνγµ=
−2γνγβγα.
3.4.20 IfM=1
2(1+γ5), showthat
M2=M.
Notethat γ5maybereplacedbyanyotherDiracmatrix(any ŴiofEq.(3.124)).If Mis
Hermitian,thenthis result, M2=M, is the definingequationfor a quantummechanical
projectionoperator.
3.4.21 Showthat
α×α=2iσ⊗12,
where α=γ0γis avector
α=(α1,α2,α3).
Notethatif αisapolarvector(Section2.4), then σisanaxialvector.
3.4.22 Provethatthe16Diracmatricesform alinearlyindependentset.
3.4.23 If we assume that a given 4 ×4m a t r i xA(with constant elements) can be written as a
linearcombinationof the16Diracmatrices
A=16summationdisplay
i=1ciŴi,
showthat
ci∼trace(AŴi).
3.5 Diagonalization of Matrices 215
3.4.24 IfC=iγ2γ0is the charge conjugation matrix, show that CγµC−1=−˜γµ, where
˜indicatestransposition.
3.4.25 Letx′
µ=/Lambda1ν
µxνbearotationbyanangle θaboutthe3-axis,
x′
0=x0,x′
1=x1cosθ+x2sinθ,
x′
2=−x1sinθ+x2cosθ, x′
3=x3.
UseR=exp(iθσ12/2)=cosθ/2+iσ12sinθ/2 (see Eq. (3.170b)) and show that
theγ’s transform just like the coordinates xµ, that is, /Lambda1ν
µγν=R−1γµR. (Note that
γµ=gµνγνand that the γµare well defined only up to a similarity transformation.)
Similarly,if x′=/Lambda1xis aboost(pureLorentztransformation)alongthe1-axis,thatis,
x′
0=x0coshζ−x1sinhζ, x′
1=−x0sinhζ+x1coshζ,
x′
2=x2,x′
3=x3,
with tanh ζ=v/candB=exp(−iζσ01/2)=coshζ/2−iσ01sinhζ/2 (see
Eq. (3.170b)),showthat /Lambda1ν
µγν=BγµB−1.
3.4.26 (a) Given r′=Ur, withUa unitary matrix and ra (column) vector with complex
elements,showthatthenorm(magnitude)of ris invariantunderthis operation.
(b) The matrix Utransforms any column vector rwith complex elements into r′,
leavingthemagnitudeinvariant: r†r=r′†r′.Showthat Uis unitary.
3.4.27 Write a subroutine that will test whether a complex n×nmatrix is self-adjoint. In
demandingequalityofmatrixelements aij=a†
ij,allowsomesmalltolerance εtocom-
pensatefor truncationerrorof thecomputer.
3.4.28 Write asubroutinethatwillformtheadjointofacomplex M×Nmatrix.
3.4.29 (a) Writeasubroutinethatwilltakeacomplex M×NmatrixAandyieldtheproduct
A†A.
Hint.Thissubroutinecancallthesubroutinesof Exercises3.2.41and3.4.28.
(b) Test your subroutine by taking Ato be one or more of the Dirac matrices,
Eq.(3.124).
3.5 D IAGONALIZATION OF MATRICES
Moment of Inertia Matrix
In many physical problems involving real symmetric or complex Hermitian matrices it is
desirable to carry out a (real) orthogonal similarity transformation or a unitary transfor-
mation (corresponding to a rotation of the coordinate system) to reduce the matrix to a
diagonal form, nondiagonal elements all equal to zero. One particularly direct example
of this is the moment of inertia matrix Iof a rigid body. From the definition of angular
momentum Lwehave
L=Iω, (3.126)
216 Chapter 3 Determinants and Matrices
ωbeingtheangularvelocity.19Theinertiamatrix Iis foundtohavediagonalcomponents
Ixx=summationdisplay
imiparenleftbig
r2
i−x2
iparenrightbig
,andsoon, (3.127)
the subscript ireferring to mass milocated at ri=(xi,yi,zi). For the nondiagonal com-
ponentswehave
Ixy=−summationdisplay
imixiyi=Iyx. (3.128)
By inspection, matrix Iis symmetric. Also, since Iappears in a physical equation of the
form (3.126), which holds for all orientations of the coordinate system, it may be consid-
eredtobea tensor(quotientrule,Section2.3).
The key now is to orient the coordinate axes (along a body-fixed frame) so that the
Ixyand the other nondiagonal elements will vanish. As a consequence of this orientation
and an indication of it, if the angular velocity is along one such realigned principal axis ,
the angular velocity and the angular momentum will be parallel. As an illustration, the
stability of rotation is used by football players when they throw the ball spinning about its
longprincipalaxis.
Eigenvectors, Eigenvalues
It is instructive to consider a geometrical picture of this problem. If the inertia matrix Iis
multiplied from each side by a unit vector of variable direction, ˆn=(α,β,γ), then in the
DiracbracketnotationofSection3.2,
/angbracketleftˆn|I|ˆn/angbracketright=I, (3.129)
whereIis the moment of inertia about the direction ˆnand a positive number (scalar).
Carryingoutthemultiplication,weobtain
I=Ixxα2+Iyyβ2+Izzγ2+2Ixyαβ+2Ixzαγ+2Iyzβγ, (3.130)
a positive definite quadratic form that must be an ellipsoid (see Fig. 3.5). From analytic
geometry it is known that the coordinate axes can always be rotated to coincide with the
axesofourellipsoid.Inmanyelementarycases,especiallywhensymmetryispresent,these
new axes, called the principal axes , can be found by inspection. We can find the axes by
locatingthelocalextremaoftheellipsoidintermsofthevariablecomponentsof n,subject
to the constraint ˆn2=1. To deal with the constraint, we introducea Lagrange multiplier λ
(Section17.6).Differentiating /angbracketleftˆn|I|ˆn/angbracketright−λ/angbracketleftˆn|ˆn/angbracketright,
∂
∂njparenleftbig
/angbracketleftˆn|I|ˆn/angbracketright−λ/angbracketleftˆn|ˆn/angbracketrightparenrightbig
=2summationdisplay
kIjknk−2λnj=0,j=1,2,3 (3.131)
yieldstheeigenvalueequations
I|ˆn/angbracketright=λ|ˆn/angbracketright. (3.132)
19Themoment of inertia matrix may also be developed from the kinetic energy of arotating body, T=1/2/angbracketleftω|I|ω/angbracketright.
3.5 Diagonalization of Matrices 217
FIGURE 3.5Momentofinertiaellipsoid.
Thesameresultcanbefoundbypurelygeometricmethods.Wenowproceedtodevelop
ageneralmethodoffindingthediagonalelementsandtheprincipalaxes.
IfR−1=˜Ris the real orthogonal matrix such that n′=Rn,o r|n′/angbracketright=R|n/angbracketrightin Dirac
notation,arethenewcoordinates,thenweobtain,using /angbracketleftn′|R=/angbracketleftn|inEq.(3.132),
/angbracketleftn|I|n/angbracketright=/angbracketleftn′|RI˜R|n′/angbracketright=I′
1n′2
1+I′
2n′2
2+I′
3n′2
3, (3.133)
where the I′
i>0 are the principal moments of inertia. The inertia matrix I′in Eq. (3.133)
isdiagonalinthenewcoordinates,
I′=RI˜R=
I′
100
0I′
20
00I′
3
. (3.134)
If werewriteEq. (3.134)using R−1=˜Rintheform
˜RI′=I˜R (3.135)
andtake˜R=(v1,v2,v3)toconsistofthreecolumnvectors,thenEq.(3.135)splitsupinto
threeeigenvalueequations,
Ivi=I′
ivi,i=1,2,3 (3.136)
witheigenvalues I′
iandeigenvectors v i. The names were introduced from the German
literature on quantum mechanics. Because these equations are linear and homogeneous
218 Chapter 3 Determinants and Matrices
(for fixed i), bySection3.1theirdeterminantshavetovanish:
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleI11−I′
iI12 I13
I12I22−I′
iI23
I13 I23I33−I′
ivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.137)
Replacing the eigenvalue I′
iby a variable λtimes the unit matrix 1, we may rewrite
Eq.(3.136) as
(I−λ1)|v/angbracketright=0. (3.136′)
Thedeterminantsettozero,
|I−λ1|=0, (3.137′)
is a cubic polynomial in λ; its three roots, of course, are the I′
i. Substituting one root at
a time back into Eq. (3.136) (or (3.136′)), we can find the corresponding eigenvectors.
Because of its applications in astronomical theories, Eq. (3.137) (or (3.137′)) is known as
thesecularequation .20Thesametreatmentappliestoanyrealsymmetricmatrix I,except
thatitseigenvaluesneednotallbepositive.Also,theorthogonalityconditioninEq.(3.87)
forRsaythat,ingeometricterms,theeigenvectors viaremutuallyorthogonalunitvectors.
Indeed they form the new coordinate system. The fact that any two eigenvectors vi,vjare
orthogonal if I′
i/negationslash=I′
jfollows from Eq. (3.136) in conjunction with the symmetry of Iby
multiplyingwith viandvj, respectively,
/angbracketleftvj|I|vi/angbracketright=I′
ivj·vi=/angbracketleftvi|I|vj/angbracketright=I′
jvi·vj. (3.138a)
SinceI′
i/negationslash=I′
jandEq. (3.138a)impliesthat (I′
j−I′
i)vi·vj=0,sovi·vj=0.
We can write the quadratic forms in Eq. (3.133) as a sum of squares in the original
coordinates|n/angbracketright,
/angbracketleftn|I|n/angbracketright=/angbracketleftn′|RI˜R|n′/angbracketright=summationdisplay
iI′
i(n·vi)2, (3.138b)
becausetherowsoftherotationmatrixin n′=Rn,or
n′
1
n′
2
n′
3
=
v1·n
v2·n
v3·n
componentwise,aremadeupoftheeigenvectors vi. Theunderlyingmatrixidentity,
I=summationdisplay
iI′
i|vi/angbracketright/angbracketleftvi|, (3.138c)
20Equation (3.126) will take on this form when ωis along one of the principal axes. Then L=λωandIω=λω.I nt h em a t h e -
matics literature λis usually calleda characteristic value ,ωacharacteristic vector .
3.5 Diagonalization of Matrices 219
maybevie wedasthe spectraldecomposition oftheinertiatensor(or anyrealsymmetric
matrix). Here, the word spectralis just another term for expansion in terms of its eigen-
values.Whenwemultiplythiseigenvalueexpansionby /angbracketleftn|ontheleftand |n/angbracketrightontheright
we reproduce the previous relation between quadratic forms. The operator Pi=|vi/angbracketright/angbracketleftvi|is
a projection operator satisfying P2
i=Pithat projects the ith component wiof any vector
|w/angbracketright=summationtext
jwj|vj/angbracketrightthat is expanded in terms of the eigenvector basis |vj/angbracketright. This is verified
by
Pi|w/angbracketright=summationdisplay
jwj|vi/angbracketright/angbracketleftvi|vj/angbracketright=wi|vi/angbracketright=vi·w|vi/angbracketright.
Finally,theidentity
summationdisplay
i|vi/angbracketright/angbracketleftvi|=1
expresses the completeness of the eigenvector basis according to which any vector |w/angbracketright=summationtext
iwi|vi/angbracketrightcan be expanded in terms of the eigenvectors. Multiplying the completeness
relationby|w/angbracketrightprovestheexpansion |w/angbracketright=summationtext
i/angbracketleftvi|w/angbracketright|vi/angbracketright.
An important extension of the spectral decomposition theorem applies to commuting
symmetric(orHermitian)matrices A,B:If[A,B]=0,thenthereisanorthogonal(unitary)
matrixthatdiagonalizesboth AandB;thatis,bothmatriceshavecommoneigenvectorsif
theeigenvaluesarenondegenerate.Thereverseofthis theoremis alsovalid.
Toprovethistheoremwediagonalize A:Avi=aivi.Multiplyingeacheigenvalueequa-
tion byBwe obtain BAvi=aiBvi=A(Bvi),which says that Bviis an eigenvector of A
with eigenvalue ai. HenceBvi=biviwith real bi. Conversely, if the vectors viare com-
mon eigenvectors of AandB,thenABvi=Abivi=aibivi=BAvi. Since the eigenvec-
torsviarecomplete,thisimplies AB=BA.
Hermitian Matrices
Forcomplexvectorspaces,Hermitianandunitarymatricesplaythesameroleassymmetric
and orthogonal matrices over real vector spaces, respectively. First, let us generalize the
important theorem about the diagonal elements and the principal axes for the eigenvalue
equation
A|r/angbracketright=λ|r/angbracketright, (3.139)
Wenowshowthatif AisaHermitianmatrix,21itseigenvaluesarerealanditseigenvectors
orthogonal.
Letλiandλjbetwoeigenvaluesand |ri/angbracketrightand|rj/angbracketright,thecorrespondingeigenvectorsof A,
aHermitianmatrix.Then
A|ri/angbracketright=λi|ri/angbracketright, (3.140)
A|rj/angbracketright=λj|rj/angbracketright. (3.141)
21IfAis real,the Hermitian requirement reduces to arequirement of symmetry.
220 Chapter 3 Determinants and Matrices
Equation(3.140)is multipliedby /angbracketleftrj|:
/angbracketleftrj|A|ri/angbracketright=λi/angbracketleftrj|ri/angbracketright. (3.142)
Equation(3.141)is multipliedby /angbracketleftri|togive
/angbracketleftri|A|rj/angbracketright=λj/angbracketleftri|rj/angbracketright. (3.143)
Takingtheadjoint22ofthisequation,wehave
/angbracketleftrj|A†|ri/angbracketright=λ∗
j/angbracketleftrj|ri/angbracketright, (3.144)
or
/angbracketleftrj|A|ri/angbracketright=λ∗
j/angbracketleftrj|ri/angbracketright (3.145)
sinceAis Hermitian.SubtractingEq. (3.145)fromEq. (3.142), weobtain
(λi−λ∗
j)/angbracketleftrj|ri/angbracketright=0. (3.146)
This is a general result for all possible combinations of iandj.F i r s t ,l e t j=i. Then
Eq.(3.146) becomes
(λi−λ∗
i)/angbracketleftri|ri/angbracketright=0. (3.147)
Since/angbracketleftri|ri/angbracketright=0 wouldbeatrivialsolutionofEq. (3.147), weconcludethat
λi=λ∗
i, (3.148)
orλiisreal, forall i.
Second,for i/negationslash=jandλi/negationslash=λj,
(λi−λj)/angbracketleftrj|ri/angbracketright=0, (3.149)
or
/angbracketleftrj|ri/angbracketright=0, (3.150)
whichmeansthattheeigenvectorsof distincteigenvaluesareorthogonal,Eq.(3.150)being
ourgeneralizationoforthogonalityinthiscomplexspace.23
Ifλi=λj(degenerate case), |ri/angbracketrightis not automatically orthogonal to |rj/angbracketright,b u ti tm a yb e
madeorthogonal.24Consider the physical problem of the momentof inertia matrix again.
Ifx1isanaxisofrotationalsymmetry,thenwewillfindthat λ2=λ3.Eigenvectors|r2/angbracketrightand
|r3/angbracketrightare each perpendicular to the symmetry axis, |r1/angbracketright, but they lie anywhere in the plane
perpendicularto |r1/angbracketright;thatis,anylinearcombinationof |r2/angbracketrightand|r3/angbracketrightisalsoaneigenvector.
Consider (a2|r2/angbracketright+a3|r3/angbracketright)witha2anda3constants.Then
Aparenleftbig
a2|r2/angbracketright+a3|r3/angbracketrightparenrightbig
=a2λ2|r2/angbracketright+a3λ3|r3/angbracketright
=λ2parenleftbig
a2|r2/angbracketright+a3|r3/angbracketrightparenrightbig
, (3.151)
22Note/angbracketleftrj|=|rj/angbracketright†for complex vectors.
23The corresponding theory for differential operators (Sturm–Liouville theory) appears in Section 10.2. The integral equation
analog (Hilbert–Schmidt theory) is given in Section 16.4.
24We are assuming here that the eigenvectors of the n-fold degenerate λispan the corresponding n-dimensional space. This
may be shown by including a parameter εin the original matrix to remove the degeneracy and then letting εapproach zero
(compareExercise3.5.30).Thisisanalogoustobreakingadegeneracyinatomicspectroscopybyapplyinganexternalmagnetic
field(Zeemaneffect).
3.5 Diagonalization of Matrices 221
as is to be expected, for x1is an axis of rotational symmetry. Therefore, if |r1/angbracketrightand|r2/angbracketright
are fixed,|r3/angbracketrightmay simply be chosen to lie in the plane perpendicular to |r1/angbracketrightand also
perpendicular to |r2/angbracketright. A general method of orthogonalizing solutions, the Gram–Schmidt
process(Section3.1), isappliedtofunctionsinSection10.3.
The set of northogonal eigenvectors |ri/angbracketrightof ourn×nHermitian matrix Aforms a
complete set, spanning the n-dimensional (complex) space,summationtext
i|ri/angbracketright/angbracketleftri|=1. This fact is
usefulinavariationalcalculationoftheeigenvalues,Section17.8.
The spectral decomposition of any Hermitian matrix Ais proved by analogy with real
symmetricmatrices
A=summationdisplay
iλi|ri/angbracketright/angbracketleftri|,
withrealeigenvalues λiandorthonormaleigenvectors |ri/angbracketright.
Eigenvalues and eigenvectors are not limited to Hermitian matrices. All matrices have
at least one eigenvalue and eigenvector. However, only Hermitian matrices have all eigen-
vectorsorthogonalandalleigenvaluesreal.
Anti-Hermitian Matrices
Occasionallyinquantumtheoryweencounteranti-Hermitianmatrices:
A†=−A.
Followingtheanalysisof thefirst portionof thissection,wecanshowthat
a. Theeigenvaluesarepureimaginary(or zero).
b. Theeigenvectorscorrespondingtodistincteigenvaluesareorthogonal.
The matrix Rformed from the normalized eigenvectors is unitary. This anti-Hermitian
propertyis preservedunderunitarytransformations.
Example 3.5.1 EIGENVALUES AND EIGENVECTORS OF A REALSYMMETRIC MATRIX
Let
A=
010
100
000
. (3.152)
Thesecularequationis
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−λ10
1−λ0
00−λvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0, (3.153)
or
−λparenleftbig
λ2−1parenrightbig
=0, (3.154)
222 Chapter 3 Determinants and Matrices
expandingbyminors.Therootsare λ=−1,0,1.Tofindtheeigenvectorcorrespondingto
λ=−1,wesubstitutethisvaluebackintotheeigenvalueequation,Eq.(3.139),
−λ10
1−λ0
00−λ
x
y
z
=
0
0
0
. (3.155)
Withλ=−1,thisyields
x+y=0,z=0. (3.156)
Within an arbitrary scale factor and an arbitrary sign (or phase factor), /angbracketleftr1|=(1,−1,0).
Note that (for real |r/angbracketrightin ordinary space) the eigenvector singles out a line in space. The
positive or negative sense is not determined. This indeterminancy could be expected if we
noted that Eq. (3.139) is homogeneous in |r/angbracketright. For convenience we will require that the
eigenvectorsbenormalizedtounity, /angbracketleftr1|r1/angbracketright=1.Withthiscondition,
/angbracketleftr1|=parenleftbigg1√
2,−1√
2,0parenrightbigg
(3.157)
isfixedexceptfor anoverallsign. For λ=0,Eq. (3.139)yields
y=0,x=0, (3.158)
/angbracketleftr2|=(0,0,1)isa suitableeigenvector.Finally,for λ=1,weget
−x+y=0,z=0, (3.159)
or
/angbracketleftr3|=parenleftbigg1√
2,1√
2,0parenrightbigg
. (3.160)
The orthogonality of r1,r2, andr3, corresponding to three distinct eigenvalues, may be
easilyverified.
Thecorrespondingspectraldecompositiongives
A=(−1)parenleftbigg1√
2,−1√
2,0parenrightbigg
1√
2
−1√
2
0
+(+1)parenleftbigg1√
2,1√
2,0parenrightbigg
1√
2
1√
2
0
+0(0,0,1)
0
0
1
=−
1
2−1
20
−1
21
20
00 0
+
1
21
20
1
21
20
000
=
010
100
000
.
/squaresolid
3.5 Diagonalization of Matrices 223
Example 3.5.2 DEGENERATE EIGENVALUES
Consider
A=
100
001
010
. (3.161)
Thesecularequationis
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle1−λ00
0−λ1
01−λvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0 (3.162)
or
(1−λ)parenleftbig
λ2−1parenrightbig
=0,λ=−1,1,1, (3.163)
adegeneratecase.If λ=−1,theeigenvalueequation(3.139)yields
2x=0,y+z=0. (3.164)
Asuitablenormalizedeigenvectoris
/angbracketleftr1|=parenleftbigg
0,1√
2,−1√
2parenrightbigg
. (3.165)
Forλ=1,weget
−y+z=0. (3.166)
Any eigenvector satisfying Eq. (3.166) is perpendicular to r1. We have an infinite number
ofchoices.Suppose,as onepossiblechoice, r2is takenas
/angbracketleftr2|=parenleftbigg
0,1√
2,1√
2parenrightbigg
, (3.167)
which clearly satisfies Eq. (3.166). Then r3must be perpendicular to r1and may be made
perpendicularto r2by25
r3=r1×r2=(1,0,0). (3.168)
Thecorrespondingspectraldecompositiongives
A=−parenleftbigg
0,1√
2,−1√
2parenrightbigg
0
1√
2
−1√
2
+parenleftbigg
0,1√
2,1√
2parenrightbigg
0
1√
2
1√
2
+(1,0,0)
1
0
0
=−
00 0
01
2−1
2
0−1
21
2
+
000
01
21
2
01
21
2
+
100
000
000
=
100
001
010
.
/squaresolid
25Theuse of thecross product is limited to three-dimensional space(see Section 1.4).
224 Chapter 3 Determinants and Matrices
Functions of Matrices
Polynomials with one or more matrix arguments are well defined and occur often. Power
series of a matrix may also be defined, provided the series converge (see Chapter 5) for
eachmatrixelement.Forexample,if Ais anyn×nmatrix,thenthepowerseries
exp(A)=∞summationdisplay
j=01
j!Aj, (3.169a)
sin(A)=∞summationdisplay
j=0(−1)j
(2j+1)!A2j+1, (3.169b)
cos(A)=∞summationdisplay
j=0(−1)j
(2j)!A2j(3.169c)
arewelldefined n×nmatrices.ForthePaulimatrices σktheEuleridentity forrealθand
k=1,2,or 3
exp(iσkθ)=12cosθ+iσksinθ, (3.170a)
follows from collecting all even and odd powers of θin separate series using σ2
k=1. For
the 4×4 Dirac matrices σjk=1 with(σjk)2=1i fj/negationslash=k=1,2 or 3 we obtain similarly
(withoutwritingtheobviousunitmatrix 14anymore)
expparenleftbig
iσjkθparenrightbig
=cosθ+iσjksinθ, (3.170b)
while
expparenleftbig
iσ0kζparenrightbig
=coshζ+iσ0ksinhζ (3.170c)
holdsfor real ζbecause(iσ0k)2=1f o rk=1,2,or 3.
ForaHermitianmatrix Athereisaunitarymatrix Uthatdiagonalizesit;thatis, UAU†=
[a1,a2,...,an].Thenthe traceformula
detparenleftbig
exp(A)parenrightbig
=expparenleftbig
trace(A)parenrightbig
(3.171)
isobtained(seeExercises3.5.2 and3.5.9) from
detparenleftbig
exp(A)parenrightbig
=detparenleftbig
Uexp(A)U†parenrightbig
=detparenleftbig
expparenleftbig
UAU†parenrightbigparenrightbig
=detexp[a1,a2,...,an]=detbracketleftbig
ea1,ea2,...,eanbracketrightbig
=productdisplay
eai=expparenleftBigsummationdisplay
aiparenrightBig
=expparenleftbig
trace(A)parenrightbig
,
usingUAiU†=(UAU†)iin the power series Eq. (3.169a) for exp (UAU†)and the product
theoremfor determinantsinSection3.2.
3.5 Diagonalization of Matrices 225
Thistraceformulaisaspecialcaseofthe spectraldecompositionlaw forany(infinitely
differentiable)function f(A)for Hermitian A:
f(A)=summationdisplay
if(λi)|ri/angbracketright/angbracketleftri|,
where|ri/angbracketrightare the common eigenvectors of AandAj. This eigenvalue expansion follows
fromAj|ri/angbracketright=λj
i|ri/angbracketright,multiplied by f(j)(0)/j!and summed over jto form the Taylor
expansion of f(λi)and yield f(A)|ri/angbracketright=f(λi)|ri/angbracketright. Finally, summing over iand using
completenessweobtain f(A)summationtext
i|ri/angbracketright/angbracketleftri|=summationtext
if(λi)|ri/angbracketright/angbracketleftri|=f(A),q.e.d.
Example 3.5.3 EXPONENTIAL OF A DIAGONAL MATRIX
If thematrix Ais diagonallike
σ3=parenleftbigg10
0−1parenrightbigg
,
then itsnth power is also diagonal with its diagonal, matrix elements raised to the nth
power:
(σ3)n=parenleftbigg10
0(−1)nparenrightbigg
.
Thensummingtheexponentialseries, elementfor element,yields
eσ3=parenleftBiggsummationtext∞
n=01
n!0
0summationtext∞
n=0(−1)n
n!parenrightBigg
=parenleftBigg
e0
01
eparenrightBigg
.
Ifwewritethegeneraldiagonalmatrixas A=[a1,a2,...,an]withdiagonalelements aj,
thenAm=[am
1,am
2,...,am
n], and summing the exponentials elementwise again we obtain
eA=[ea1,ea2,...,ean].
Usingthespectraldecompositionlawweobtaindirectly
eσ3=e+1(1,0)parenleftbigg1
0parenrightbigg
+e−1(0,1)parenleftbigg0
1parenrightbigg
=parenleftbigge0
0e−1parenrightbigg
./squaresolid
Anotherimportantrelationis the Baker–Hausdorffformula ,
exp(iG)Hexp(−iG)=H+[iG,H]+1
2bracketleftbig
iG,[iG,H]bracketrightbig
+···, (3.172)
whichfollowsfrommultiplyingthepowerseriesforexp (iG)andcollectingthetermswith
thesamepowersof iG. Herewedefine
[G,H]=GH−HG
asthecommutator ofGandH.
The preceding analysis has the advantage of exhibiting and clarifying conceptual rela-
tionships in the diagonalization of matrices. However, for matrices larger than 3 ×3, or
perhaps 4×4,theprocessrapidlybecomessocumbersomethatweturntocomputersand
226 Chapter 3 Determinants and Matrices
iterative techniques.26One such technique is the Jacobi method for determining eigenval-
ues and eigenvectors of real symmetric matrices. This Jacobi technique for determining
eigenvaluesandeigenvectorsandtheGauss–Seidelmethodofsolvingsystemsofsimulta-
neous linear equations are examples of relaxation methods. They are iterative techniques
in which the errors may decrease or relax as the iterations continue. Relaxation methods
areusedextensivelyfor thesolutionof partialdifferentialequations.
Exercises
3.5.1 (a) Startingwiththeorbitalangularmomentumofthe ith elementof mass,
Li=ri×pi=miri×(ω×ri),
derivetheinertiamatrixsuchthat L=Iω,|L/angbracketright=I|ω/angbracketright.
(b) Repeatthederivationstartingwithkineticenergy
Ti=1
2mi(ω×ri)2parenleftbigg
T=1
2/angbracketleftω|I|ω/angbracketrightparenrightbigg
.
3.5.2 Show that the eigenvalues of a matrix are unaltered if the matrix is transformed by a
similaritytransformation.
This property is not limited to symmetric or Hermitian matrices. It holds for any ma-
trix satisfying the eigenvalue equation, Eq. (3.139). If our matrix can be brought into
diagonalform byasimilaritytransformation,thentwoimmediateconsequencesare
1. The trace(sum ofeigenvalues)isinvariantunderasimilaritytransformation.
2. The determinant (product of eigenvalues) is invariant under a similarity transfor-
mation.
Note.The invariance of the trace and determinant are often demonstrated by using the
Cayley–Hamiltontheorem:Amatrixsatisfiesits owncharacteristic(secular) equation.
3.5.3 As a converse of the theorem that Hermitian matrices have real eigenvalues and that
eigenvectorscorrespondingtodistincteigenvaluesareorthogonal,showthatif
(a) theeigenvaluesofamatrixarerealand
(b) theeigenvectorssatisfy r†
irj=δij=/angbracketleftri|rj/angbracketright,
thenthematrixisHermitian.
3.5.4 Show that a real matrix that is not symmetric cannot be diagonalized by an orthogonal
similaritytransformation.
Hint.Assume that the nonsymmetric real matrix can be diagonalized and develop a
contradiction.
26In higher-dimensional systems the secular equation may be strongly ill-conditioned with respect to the determination of its
roots (the eigenvalues). Direct solution by computer may be very inaccurate. Iterative techniques for diagonalizing the original
matrix areusually preferred. SeeSections 2.7 and 2.9 ofPress et al.,loc. cit.
3.5 Diagonalization of Matrices 227
3.5.5 The matrices representing the angular momentum components Jx,Jy, andJzare all
Hermitian. Show that the eigenvalues of J2, whereJ2=J2
x+J2
y+J2
z, are real and
nonnegative.
3.5.6 Ahaseigenvalues λiandcorrespondingeigenvectors |xi/angbracketright.Showthat A−1hasthesame
eigenvectorsbutwitheigenvalues λ−1
i.
3.5.7 Asquarematrixwithzerodeterminantislabeled singular.
(a) IfAis singular,showthatthereis atleastonenonzerocolumnvector vsuchthat
A|v/angbracketright=0.
(b) If thereis anonzerovector |v/angbracketrightsuchthat
A|v/angbracketright=0,
showthatAisasingularmatrix.Thismeansthatifamatrix(oroperator)haszero
as an eigenvalue, the matrix (or operator) has no inverse and its determinant is
zero.
3.5.8 The same similarity transformation diagonalizes each of two matrices. Show that the
original matrices must commute. (This is particularly important in the matrix (Heisen-
berg)formulationofquantummechanics.)
3.5.9 Two Hermitian matrices AandBhave the same eigenvalues. Show that AandBare
relatedbyaunitarysimilaritytransformation.
3.5.10 Find the eigenvalues and an orthonormal (orthogonal and normalized) set of eigenvec-
torsfor thematricesof Exercise3.2.15.
3.5.11 Show that the inertia matrix for a single particle of mass mat(x,y,z)has a zero de-
terminant. Explain this result in terms of the invariance of the determinant of a matrix
undersimilaritytransformations(Exercise3.3.10)andapossiblerotationofthecoordi-
natesystem.
3.5.12 A certain rigid body may be represented by three point masses: m1=1a t(1,1,−2),
m2=2a t(−1,−1,0),andm3=1a t(1,1,2).
(a) Findtheinertiamatrix.
(b) Diagonalizetheinertiamatrix,obtainingtheeigenvaluesandtheprincipalaxes(as
orthonormaleigenvectors).
3.5.13 Unitmasses areplacedasshowninFig.3.6.
(a) Findthemomentofinertiamatrix.
(b) Findtheeigenvaluesandasetof orthonormaleigenvectors.
(c) Explainthedegeneracyinterms ofthesymmetryofthesystem.
ANS.I=
4−1−1
−14−1
−1−14
λ1=2
r1=(1/√
3,1/√
3,1/√
3)
λ2=λ3=5.
228 Chapter 3 Determinants and Matrices
FIGURE 3.6Mass sitesfor inertiatensor.
3.5.14 Am a s s m1=1/2 kg is located at (1,1,1)(meters), a mass m2=1/2k gi sa t
(−1,−1,−1).Thetwomassesareheldtogetherbyanideal(weightless,rigid)rod.
(a) Findtheinertiatensorof thispairofmasses.
(b) Findtheeigenvaluesandeigenvectorsofthis inertiamatrix.
(c) Explain the meaning, the physical significance of the λ=0 eigenvalue. What is
thesignificanceofthecorrespondingeigenvector?
(d) Now that you have solved this problem by rather sophisticated matrix techniques,
explainhowyoucouldobtain
(1)λ=0 andλ=? — byinspection(thatis, usingcommonsense).
(2)rλ=0=? — byinspection(thatis,usingfreshmanphysics).
3.5.15 Unitmassesareattheeightcornersofacube (±1,±1,±1).Findthemomentofinertia
matrix and show that there is a triple degeneracy. This means that so far as moments of
inertiaare concerned,thecubicstructureexhibitssphericalsymmetry.
Findtheeigenvaluesandcorrespondingorthonormaleigenvectorsof thefollowingma-
trices (as a numerical check, note that the sum of the eigenvalues equals the sum of the
diagonalelementsoftheoriginalmatrix,Exercise3.3.9).Notealsothecorrespondence
betweendet A=0 andtheexistenceof λ=0,asrequiredbyExercises3.5.2and3.5.7.
3.5.16 A=
101
010
101
.
ANS.λ=0,1,2.
3.5.17 A=
1√
20√
200
00 0
.
ANS.λ=−1,0,2.
3.5 Diagonalization of Matrices 229
3.5.18 A=
110
101
011
.
ANS.λ=−1,1,2.
3.5.19 A=
1√
80√
81√
8
0√
81
.
ANS.λ=−3,1,5.
3.5.20 A=
100
011
011
.
ANS.λ=0,1,2.
3.5.21 A=
10 0
01√
2
0√
20
.
ANS.λ=−1,1,2.
3.5.22 A=
010
101
010
.
ANS.λ=−√
2,0,√
2.
3.5.23 A=
200
011
011
.
ANS.λ=0,2,2.
3.5.24 A=
011
101
110
.
ANS.λ=−1,−1,2.
3.5.25 A=
1−1−1
−11−1
−1−11
.
ANS.λ=−1,2,2.
3.5.26 A=
111
111
111
.
ANS.λ=0,0,3.
230 Chapter 3 Determinants and Matrices
3.5.27 A=
502
010
202
.
ANS.λ=1,1,6.
3.5.28 A=
110
110
000
.
ANS.λ=0,0,2.
3.5.29 A=
50√
3
030√
30 3
.
ANS.λ=2,3,6.
3.5.30 (a) Determinetheeigenvaluesandeigenvectorsof
parenleftbigg1ε
ε1parenrightbigg
.
Note that the eigenvalues are degenerate for ε=0 but that the eigenvectors are
orthogonalfor all ε/negationslash=0 andε→0.
(b) Determinetheeigenvaluesandeigenvectorsof
parenleftbigg11
ε21parenrightbigg
.
Notethattheeigenvaluesaredegeneratefor ε=0andthatforthis(nonsymmetric)
matrixtheeigenvectors (ε=0)donotspanthespace.
(c) Find the cosine of the angle between the two eigenvectors as a function of εfor
0≤ε≤1.
3.5.31 (a) Take the coefficients of the simultaneous linear equations of Exercise 3.1.7 to be
the matrix elements aijof matrixA(symmetric). Calculate the eigenvalues and
eigenvectors.
(b) Formamatrix Rwhosecolumnsaretheeigenvectorsof A,andcalculatethetriple
matrixproduct ˜RAR.
ANS.λ=3.33163.
3.5.32 RepeatExercise3.5.31byusingthematrixofExercise3.2.39.
3.5.33 Describethegeometricpropertiesof thesurface
x2+2xy+2y2+2yz+z2=1.
Howis itorientedinthree-dimensionalspace?Is it aconicsection?If so, whichkind?
3.6 Normal Matrices 231
Table 3.1
Matrix Eigenvalues Eigenvectors
(for different eigenvalues)
Hermitian Real Orthogonal
Anti-Hermitian Pure imaginary (or zero) Orthogonal
Unitary Unit magnitude Orthogonal
Normal If Ahas eigenvalue λ, Orthogonal
A†has eigenvalue λ∗AandA†have the
same eigenvectors
3.5.34 For a Hermitian n×nmatrixAwith distinct eigenvalues λjand a function f,s h o w
thatthespectraldecompositionlawmaybeexpressedas
f(A)=nsummationdisplay
j=1f(λj)producttext
i/negationslash=j(A−λi)
producttext
i/negationslash=j(λj−λi).
Thisformulais duetoSylvester.
3.6 N ORMAL MATRICES
In Section 3.5 we concentrated primarily on Hermitian or real symmetric matrices and
on the actual process of finding the eigenvalues and eigenvectors. In this section27we
generalize to normal matrices, with Hermitian and unitary matrices as special cases. The
physicallyimportantproblemofnormalmodesofvibrationandthenumericallyimportant
problemofill-conditionedmatricesarealsoconsidered.
Anormalmatrixis amatrixthatcommuteswithitsadjoint,
bracketleftbig
A,A†bracketrightbig
=0.
ObviousandimportantexamplesareHermitianandunitarymatrices.Wewillshowthat
normalmatriceshaveorthogonaleigenvectors(see Table3.1). Weproceedintwo steps.
I.L etAhaveaneigenvector |x/angbracketrightandcorrespondingeigenvalue λ.Then
A|x/angbracketright=λ|x/angbracketright (3.173)
or
(A−λ1)|x/angbracketright=0. (3.174)
For convenience the combination A−λ1 will be labeled B. Taking the adjoint of
Eq.(3.174), weobtain
/angbracketleftx|(A−λ1)†=0=/angbracketleftx|B†. (3.175)
Because
bracketleftbig
(A−λ1)†,(A−λ1)bracketrightbig
=bracketleftbig
A,A†bracketrightbig
=0,
27Normalmatricesarethelargestclassofmatricesthatcanbediagonalizedbyunitarytransformations.Foranextensivediscus-
sion ofnormal matrices, seeP. A.Macklin,Normal matrices for physicists. Am.J.Phys. 52: 513 (1984).
232 Chapter 3 Determinants and Matrices
wehave
bracketleftbig
B,B†bracketrightbig
=0. (3.176)
Thematrix Bisalsonormal.
From Eqs. (3.174)and(3.175)weform
/angbracketleftx|B†B|x/angbracketright=0. (3.177)
Thisequals
/angbracketleftx|BB†|x/angbracketright=0 (3.178)
byEq.(3.176). NowEq. (3.178)mayberewrittenas
parenleftbig
B†|x/angbracketrightparenrightbig†parenleftbig
B†|x/angbracketrightparenrightbig
=0. (3.179)
Thus
B†|x/angbracketright=parenleftbig
A†−λ∗1parenrightbig
|x/angbracketright=0. (3.180)
We see that for normal matrices, A†has the same eigenvectors as Abut the complex con-
jugateeigenvalues.
II. Now,consideringmorethanoneeigenvector–eigenvalue,wehave
A|xi/angbracketright=λi|xi/angbracketright, (3.181)
A|xj/angbracketright=λj|xj/angbracketright. (3.182)
MultiplyingEq. (3.182)fromtheleftby /angbracketleftxi|yields
/angbracketleftxi|A|xj/angbracketright=λj/angbracketleftxi|xj/angbracketright. (3.183)
Takingthetransposeof Eq.(3.181), weobtain
/angbracketleftxi|A=parenleftbig
A†|xi/angbracketrightparenrightbig†. (3.184)
From Eq. (3.180), with A†having the same eigenvectors as Abut the complex conjugate
eigenvalues,
parenleftbig
A†|xi/angbracketrightparenrightbig†=parenleftbig
λ∗
i|xi/angbracketrightparenrightbig†=λi/angbracketleftxi|. (3.185)
SubstitutingintoEq.(3.183) wehave
λi/angbracketleftxi|xj/angbracketright=λj/angbracketleftxi|xj/angbracketright
or
(λi−λj)/angbracketleftxi|xj/angbracketright=0. (3.186)
Thisis thesameas Eq.(3.149).
Forλi/negationslash=λj,
/angbracketleftxj|xi/angbracketright=0.
Theeigenvectorscorrespondingtodifferenteigenvaluesofanormalmatrixare orthogo-
nal.Thismeansthatanormalmatrixmaybediagonalizedbyaunitarytransformation.The
required unitary matrix may be constructed from the orthonormal eigenvectors as shown
earlier,inSection3.5.
The converse of this result is also true. If Acan be diagonalized by a unitary transfor-
mation,then Ais normal.
3.6 Normal Matrices 233
Normal Modes of Vibration
WeconsiderthevibrationsofaclassicalmodeloftheCO 2molecule.Itisanillustrationof
theapplicationofmatrixtechniquestoaproblemthatdoesnotstartasamatrixproblem.It
alsoprovidesanexampleoftheeigenvaluesandeigenvectorsofanasymmetricrealmatrix.
Example 3.6.1 NORMAL MODES
Consider three masses on the x-axis joined by springs as shown in Fig. 3.7. The spring
forces are assumed to be linear (small displacements, Hooke’s law), and the mass is con-
strainedtostayonthe x-axis.
Usingadifferentcoordinateforeachmass, Newton’ssecondlawyieldsthesetofequa-
tions
¨x1=−k
M(x1−x2)
¨x2=−k
m(x2−x1)−k
m(x2−x3) (3.187)
¨x3=−k
M(x3−x2).
The system of masses is vibrating. We seek the common frequencies, ω, such that all
massesvibrateatthis samefrequency.Thesearethe normalmodes.Let
xi=xi0eiωt,i=1,2,3.
Substitutingthisset intoEq. (3.187),wemayrewriteitas
k
M−k
M0
−k
m2k
m−k
m
0−k
Mk
M
x1
x2
x3
=+ω2
x1
x2
x3
, (3.188)
with the common factor eiωtdivided out. We have a matrix–eigenvalue equation with the
matrixasymmetric.Thesecularequationis
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglek
M−ω2−k
M0
−k
m2k
m−ω2−k
m
0−k
Mk
M−ω2vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.189)
FIGURE 3.7Doubleoscillator.
234 Chapter 3 Determinants and Matrices
Thisleadsto
ω2parenleftbiggk
M−ω2parenrightbiggparenleftbigg
ω2−2k
m−k
Mparenrightbigg
=0.
Theeigenvaluesare
ω2=0,k
M,k
M+2k
m,
allreal.
Thecorrespondingeigenvectorsaredeterminedbysubstitutingtheeigenvaluesbackinto
Eq.(3.188) oneeigenvalueatatime.For ω2=0,Eq.(3.188), yields
x1−x2=0,−x1+2x2−x3=0,−x2+x3=0.
Thenweget
x1=x2=x3.
Thisdescribespuretranslationwithnorelativemotionof themasses andnovibration.
Forω2=k/M,Eq. (3.188)yields
x1=−x3,x 2=0.
Thetwooutermassesaremovinginoppositedirection.Thecentralmass isstationary.
Forω2=k/M+2k/m, theeigenvectorcomponentsare
x1=x3,x 2=−2M
mx1.
Thetwooutermassesaremovingtogether.Thecentralmassismovingoppositetothetwo
outerones. Thenetmomentumis zero.
Any displacement of the three masses along the x-axis can be described as a linear
combinationof thesethreetypesof motion:translationplustwo formsof vibration. /squaresolid
Ill-Conditioned Systems
Asystemofsimultaneouslinearequationsmaybewrittenas
A|x/angbracketright=|y/angbracketrightorA−1|y/angbracketright=|x/angbracketright, (3.190)
withAand|y/angbracketrightknown and|x/angbracketrightunknown. When a small error in |y/angbracketrightresults in a larger error
in|x/angbracketright,thenthematrix Aiscalledill-conditioned .With|δx/angbracketrightanerrorin|x/angbracketrightand|δx/angbracketrightanerror
in|y/angbracketright, therelativeerrors maybewrittenas
bracketleftbigg/angbracketleftδx|δx/angbracketright
/angbracketleftx|x/angbracketrightbracketrightbigg1/2
≤K(A)bracketleftbigg/angbracketleftδy|δy/angbracketright
/angbracketlefty|y/angbracketrightbracketrightbigg1/2
. (3.191)
HereK(A),apropertyofmatrix A,islabeledthe conditionnumber .ForAHermitianone
formoftheconditionnumberisgivenby28
K(A)=|λ|max
|λ|min. (3.192)
28G.E.Forsythe,andC.B.Moler, ComputerSolutionofLinearAlgebraicSystems .EnglewoodCliffs,NJ,PrenticeHall(1967).
3.6 Normal Matrices 235
AnapproximateformduetoTuring29is
K(A)=n[Aij]maxbracketleftbig
A−1
ijbracketrightbig
max, (3.193)
inwhich nistheorderof thematrixand [Aij]maxis themaximumelementin A.
Example 3.6.1 ANILL-CONDITIONED MATRIX
Acommonexampleofanill-conditionedmatrixistheHilbertmatrix, Hij=(i+j−1)−1.
The Hilbert matrix of order 4, H4, is encountered in a least-squares fit of data to a third-
degreepolynomial.Wehave
H4=
11
21
31
4
1
21
31
41
5
1
31
41
51
6
1
41
51
61
7
. (3.194)
Theelementsoftheinversematrix(order n)aregi v enby
parenleftbig
H−1
nparenrightbig
ij=(−1)i+j
i+j−1·(n+i−1)!(n+j−1)!
[(i−1)!(j−1)!]2(n−i)!(n−j)!.(3.195)
Forn=4,
H−1
4=
16−120 240 −140
−120 1200 −2700 1680
240−2700 6480 −4200
−140 1680 −4200 2800
. (3.196)
FromEq. (3.193)theTuringestimateof theconditionnumberfor H4becomes
KTuring=4×1×6480
=2.59×104.
This is a warning that an input error may be multiplied by 26,000 in the calculation
of the output result. It is a statement that H4is ill-conditioned. If you encounter a highly
ill-conditionedsystem, youhavetwoalternatives(besidesabandoningtheproblem).
(a) Tryadifferentmathematicalattack.
(b) Arrangetocarrymoresignificantfiguresandpushthroughbybruteforce.
As previously seen, matrix eigenvector–eigenvalue techniques are not limited to the so-
lutionofstrictlymatrixproblems.Afurtherexampleofthetransferoftechniquesfromone
area to another is seen in the application of matrix techniques to the solution of Fredholm
eigenvalue integral equations, Section 16.3. In turn, these matrix techniques are strength-
enedbyavariationalcalculationofSection17.8. /squaresolid
29CompareJ.Todd, TheConditionoftheFiniteSegmentsoftheHilbertMatrix ,AppliedMathematicsSeriesNo.313.Washing-
ton, DC: National Bureau of Standards.
236 Chapter 3 Determinants and Matrices
Exercises
3.6.1 Showthatevery2 ×2matrixhastwoeigenvectorsandcorrespondingeigenvalues.The
eigenvectorsarenotnecessarilyorthogonalandmaybedegenerate.Theeigenvaluesare
notnecessarilyreal.
3.6.2 AsanillustrationofExercise3.6.1,findtheeigenvaluesandcorrespondingeigenvectors
for
parenleftbigg24
12parenrightbigg
.
Notethattheeigenvectorsare notorthogonal.
ANS.λ1=0,r1=(2,−1);
λ2=4,r2=(2,1).
3.6.3 IfAis a 2×2 matrix,showthatitseigenvalues λsatisfy thesecularequation
λ2−λtrace(A)+detA=0.
3.6.4 Assuming a unitary matrix Uto satisfy an eigenvalue equation Ur=λr, show that the
eigenvalues of the unitary matrix have unit magnitude. This same result holds for real
orthogonalmatrices.
3.6.5 Since an orthogonal matrix describing a rotation in real three-dimensional space is a
special case of a unitary matrix, such an orthogonal matrix can be diagonalized by a
unitarytransformation.
(a) Showthatthesumofthethreeeigenvaluesis 1 +2cosϕ,whereϕisthenetangle
ofrotationaboutasinglefixedaxis.
(b) Given that one eigenvalue is 1, show that the other two eigenvalues must be eiϕ
ande−iϕ.
Ourorthogonalrotationmatrix(real elements)hascomplexeigenvalues.
3.6.6 Ais annth-order Hermitian matrix with orthonormal eigenvectors |xi/angbracketrightand real eigen-
valuesλ1≤λ2≤λ3≤···≤λn.Showthatfor aunitmagnitudevector |y/angbracketright,
λ1≤/angbracketlefty|A|y/angbracketright≤λn.
3.6.7 AparticularmatrixisbothHermitianandunitary.Showthatitseigenvaluesareall ±1.
Note.ThePauliandDiracmatricesarespecificexamples.
3.6.8 ForhisrelativisticelectrontheoryDiracrequiredasetof fouranticommutingmatrices.
AssumethatthesematricesaretobeHermitianandunitary.Iftheseare n×nmatrices,
showthat nmustbeeven.With2 ×2matricesinadequate(why?),thisdemonstratesthat
the smallest possible matrices forming a set of four anticommuting, Hermitian, unitary
matricesare 4×4.
3.6 Normal Matrices 237
3.6.9 Aisanormalmatrixwitheigenvalues λnandorthonormaleigenvectors |xn/angbracketright.Showthat
Amaybewrittenas
A=summationdisplay
nλn|xn/angbracketright/angbracketleftxn|.
Hint.Show that both this eigenvectorform of Aand the original Agive the same result
actingonanarbitraryvector |y/angbracketright.
3.6.10 Ahas eigenvalues1and −1 andcorrespondingeigenvectorsparenleftbig1
0parenrightbig
andparenleftbig0
1parenrightbig
. Construct A.
ANS.A=parenleftbigg10
0−1parenrightbigg
.
3.6.11 Anon-Hermitianmatrix Ahaseigenvalues λiandcorrespondingeigenvectors |ui/angbracketright.The
adjoint matrix A†has the same set of eigenvalues but different corresponding eigen-
vectors,|vi/angbracketright.Showthattheeigenvectorsforma biorthogonal set,inthesensethat
/angbracketleftvi|uj/angbracketright=0forλ∗
i/negationslash=λj.
3.6.12 Youare givenapairofequations:
A|fn/angbracketright=λn|gn/angbracketright
˜A|gn/angbracketright=λn|fn/angbracketrightwithAreal.
(a) Provethat |fn/angbracketrightis aneigenvectorof (˜AA)witheigenvalue λ2
n.
(b) Provethat |gn/angbracketrightisaneigenvectorof (A˜A)witheigenvalue λ2
n.
(c) Statehowyouknowthat
(1) The|fn/angbracketrightformanorthogonalset.
(2) The|gn/angbracketrightformanorthogonalset.
(3)λ2
nisreal.
3.6.13 Provethat Aof theprecedingexercisemaybewrittenas
A=summationdisplay
nλn|gn/angbracketright/angbracketleftfn|,
withthe|gn/angbracketrightand/angbracketleftfn|normalizedtounity.
Hint.Expandyourarbitraryvectorasa linearcombinationof |fn/angbracketright.
3.6.14 Given
A=1√
5parenleftbigg22
1−4parenrightbigg
,
(a) Constructthetranspose ˜Aandthesymmetricforms ˜AAandA˜A.
(b) From A˜A|gn/angbracketright=λ2
n|gn/angbracketrightfindλnand|gn/angbracketright. Normalizethe |gn/angbracketright.
(c) From˜AA|fn/angbracketright=λ2
n|gn/angbracketrightfindλn[sameas(b)] and |fn/angbracketright.Normalizethe |fn/angbracketright.
(d) Verifythat A|fn/angbracketright=λn|gn/angbracketrightand˜A|gn/angbracketright=λn|fn/angbracketright.
(e) Verifythat A=summationtext
nλn|gn/angbracketright/angbracketleftfn|.
238 Chapter 3 Determinants and Matrices
3.6.15 Giventheeigenvalues λ1=1,λ2=−1 andthecorrespondingeigenvectors
|f1/angbracketright=parenleftbigg1
0parenrightbigg
,|g1/angbracketright=1√
2parenleftbigg1
1parenrightbigg
,|f2/angbracketright=parenleftbigg0
1parenrightbigg
,and|g2/angbracketright=1√
2parenleftbigg1
−1parenrightbigg
,
(a) construct A;
(b) verifythat A|fn/angbracketright=λn|gn/angbracketright;
(c) verifythat ˜A|gn/angbracketright=λn|fn/angbracketright.
ANS.A=1√
2parenleftbigg1−1
11parenrightbigg
.
3.6.16 ThisisacontinuationofExercise3.4.12,wheretheunitarymatrix UandtheHermitian
matrixHare relatedby
U=eiaH.
(a) If trace H=0,showthat det U=+1.
(b) If det U=+1,showthattrace H=0.
Hint.Hmay be diagonalized by a similarity transformation. Then interpreting the ex-
ponentialbyaMaclaurinexpansion, Uisalsodiagonal.Thecorrespondingeigenvalues
aregivenby uj=exp(iahj).
Note.Theseproperties,andthoseofExercise3.4.12,arevitalinthedevelopmentofthe
conceptofgeneratorsingrouptheory—Section4.2.
3.6.17 Ann×nmatrixAhasneigenvalues Ai.I fB=eA, show that Bhas the same eigen-
vectorsas A, withthecorrespondingeigenvalues Bigivenby Bi=exp(Ai).
Note.eAis definedbytheMaclaurinexpansionoftheexponential:
eA=1+A+A2
2!+A3
3!+···.
3.6.18 AmatrixPisaprojectionoperator(seethediscussionfollowingEq.(3.138c))satisfying
thecondition
P2=P.
Showthatthecorrespondingeigenvalues (ρ2)λandρλsatisfy therelation
parenleftbig
ρ2parenrightbig
λ=(ρλ)2=ρλ.
Thismeansthattheeigenvaluesof Pare 0and1.
3.6.19 Inthematrixeigenvector–eigenvalueequation
A|ri/angbracketright=λi|ri/angbracketright,
Ais ann×nHermitian matrix. For simplicity assume that its nreal eigenvalues are
distinct,λ1beingthelargest.If |r/angbracketrightisanapproximationto |r1/angbracketright,
|r/angbracketright=|r1/angbracketright+nsummationdisplay
i=2δi|ri/angbracketright,
3.6 Additional Readings 239
FIGURE 3.8Tripleoscillator.
showthat
/angbracketleftr|A|r/angbracketright
/angbracketleftr|r/angbracketright≤λ1
andthattheerror in λ1isof theorder|δi|2.T ak e|δi|≪1.
Hint.Then|ri/angbracketrightformacomplete orthogonalsetspanningthe n-dimensional(complex)
space.
3.6.20 Two equal masses are connected to each other and to walls by springs as shown in
Fig.3.8. Themassesareconstrainedtostayonahorizontalline.
(a) SetuptheNewtonianaccelerationequationfor eachmass.
(b) Solvethesecularequationfortheeigenvectors.
(c) Determinetheeigenvectorsandthusthenormalmodesof motion.
3.6.21 Given a normal matrix Awith eigenvalues λj,show that A†has eigenvalues λ∗
j,its
real part (A+A†)/2 has eigenvalues ℜ(λj), and its imaginary part (A−A†)/2ihas
eigenvaluesℑ(λj).
AdditionalReadings
Aitken,A.C., DeterminantsandMatrices .NewYork:Interscience(1956).Reprinted,Greenwood(1983).Aread-
able introduction to determinants and matrices.
Barnett, S., Matrices:Methods and Applications . Oxford: Clarendon Press (1990).
Bickley, W. G., and R. S. H. G. Thompson, Matrices—Their Meaning and Manipulation . Princeton, NJ: Van
Nostrand (1964). A comprehensive account of matrices in physical problems, their analytic properties, and
numerical techniques.
Brown, W. C., Matrices and VectorSpaces . NewYork: Dekker(1991).
Gilbert,J. andL., Linear Algebra and MatrixTheory . San Diego: AcademicPress (1995).
Heading,J., MatrixTheoryforPhysicists .London:Longmans,GreenandCo.(1958).Areadableintroductionto
determinants and matrices, with applications to mechanics, electromagnetism, special relativity, and quantum
mechanics.
Vein,R.,andP. Dale, Determinants and Their Applications in Mathematical Physics .Berlin: Springer (1998).
Watkins,D. S., Fundamentals of Matrix Computations . NewYork: Wiley (1991).
This page intentionally left blank
CHAPTER 4
GROUP THEORY
Disciplined judgment, about what is neat
and symmetrical and elegant hastime and
time again provedan excellent guide to
how nature works
MURRAYGELL-MANN
4.1 I NTRODUCTION TO GROUP THEORY
In classical mechanics the symmetry of a physical system leads to conservation laws .
Conservationofangularmomentumisadirectconsequenceofrotationalsymmetry,which
meansinvariance underspatialrotations.Inthefirstthirdofthe20thcentury,Wignerand
others realized that invariance was a key concept in understanding the new quantum phe-
nomena and in developing appropriate theories. Thus, in quantum mechanics the concept
ofangularmomentumandspinhasbecomeevenmorecentral.Itsgeneralizations, isospin
in nuclear physics and the flavor symmetry in particle physics, are indispensable tools
in building and solving theories. Generalizations of the concept of gauge invariance of
classicalelectrodynamicstotheisospinsymmetryleadtotheelectroweakgaugetheory.
In each case the set of these symmetry operations forms a group. Group theory is the
mathematicaltooltotreatinvariantsandsymmetries.Itbringsunificationandformalization
of principles, such as spatial reflections, or parity, angular momentum, and geometry, that
arewidelyused byphysicists.
In geometry the fundamental role of group theory was recognized more than a cen-
turyagobymathematicians(e.g.,FelixKlein’sErlangerProgram).InEuclideangeometry
the distance between two points, the scalar product of two vectors or metric, does not
change under rotations or translations. These symmetries are characteristic of this geom-
etry. In special relativity the metric, or scalar product of four-vectors, differs from that of
241
242 Chapter 4 Group Theory
Euclidean geometry in that it is no longer positive definite and is invariant under Lorentz
transformations.
For a crystal the symmetry group contains only a finite number of rotations at discrete
values of angles or reflections. The theory of such discreteorfinitegroups, developed
originally as a branch of pure mathematics, now is a useful tool for the development of
crystallography and condensed matter physics. A brief introduction to this area appears in
Section 4.7. When the rotations depend on continuously varying angles (the Euler angles
of Section 3.3) the rotation groups have an infinite number of elements. Such continuous
(orLie1)groupsare the topic of Sections 4.2–4.6. In Section 4.8 we give an introduction
to differential forms, with applications to Maxwell’s equations and topics of Chapters 1
and2, whichallowsseeingthesetopicsfrom adifferentperspective.
Definition of a Group
A group Gmay be defined as a set of objects or operations, rotations, transformations,
called the elements of G, that may be combined, or “multiplied,” to form a well-defined
productin G, denotedbya*, thatsatisfiesthefollowingfour conditions.
1. Ifaandbare any two elements of G, then the product a∗bis also an element of G,
wherebacts before a;o r(a,b)→a∗bassociates (or maps) an element a∗bofG
with the pair (a,b)of elements of G. This property is known as “ Gis closed under
multiplicationofits ownelements.”
2. This multiplicationis associative: (a∗b)∗c=a∗(b∗c).
3. There is a unit element21i nGsuch that 1∗a=a∗1=afor every element ainG.
Theunitisunique: 1 =1′∗1=1′.
4. There is an inverse, or reciprocal, of each element aofG, labeled a−1, such that
a∗a−1=a−1∗a=1. The inverse is unique: If a−1anda′−1are both inverses of a,
thena′−1=a′−1∗(a∗a′−1)=(a′−1∗a)∗a−1=a−1.
Sincethe*formultiplicationistedioustowrite,itiscustomarytodropitandsimplyletit
beunderstood.Fromnowon, wewrite abinsteadof a∗b.
•If a subset G′ofGis closed under multiplication, it is a group and called a subgroup
ofG; thatis,G′is closedunderthe multiplicationof G. The unitof Galways forms a
subgroupof G.
•Ifgg′g−1is an element of G′for anygofGandg′ofG′, thenG′is called an in-
variant subgroup ofG. The subgroup consisting of the unit is invariant. If the group
elements are square matrices, then gg′g−1corresponds to a similarity transformation
(seeEq. (3.100)).
•Ifab=bafor alla,bofG, the group is called abelian, that is, the order in products
doesnotmatter;commutativemultiplicationisoftendenotedbya +sign.Examplesare
vector spaces whose unit is the zero vector and −ais the inverse of afor all elements
ainG.
1Afterthe NorwegianmathematicianSophus Lie.
2Following E. Wigner, the unit element of a group is often labeled E, from the German Einheit, that is, unit, or just 1, or Ifor
identity.
4.1 Introduction to Group Theory 243
Example 4.1.1 ORTHOGONAL AND UNITARY GROUPS
Orthogonal n×nmatrices form the group O(n), andSO(n)if their determinants are +1
(Sstands for “special”). If ˜Oi=O−1
ifori=1 and 2 (see Section 3.3 for orthogonal
matrices)areelementsof O(n),thentheproduct
/tildewiderO1O2=˜O2˜O1=O−1
2O−1
1=(O1O2)−1
is also an orthogonal matrix in O(n), thus proving closure under (matrix) multiplication.
Theinverseisthetranspose(orthogonal)matrix.Theunitofthegroupisthe n-dimensional
unit matrix 1 n. A real orthogonal n×nmatrix has n(n−1)/2 independent parameters.
Forn=2, there is only one parameter: one angle. For n=3, there are three independent
parameters:thethreeEuler anglesof Section3.3.
If˜Oi=O−1
i(fori=1 and 2) are elements of SO(n), then closure requires proving in
additionthattheirproducthasdeterminant +1,whichfollowsfromtheproducttheoremin
Chapter3.
Likewise, unitary n×nmatrices form the group U(n), andSU(n)if their determinants
are+1.IfU†
i=U−1
i(see Section3.4 forunitarymatrices)areelementsof U(n),then
(U1U2)†=U†
2U†
1=U−1
2U−1
1=(U1U2)−1,
sotheproductisunitaryandanelementof U(n),thusprovingclosureundermultiplication.
Eachunitarymatrixhasaninverse(itsHermitianadjoint),whichagainis unitary.
IfU†
i=U−1
iare elementsof SU(n), then closure requires us to prove thattheir product
alsohasdeterminant +1,whichfollowsfrom theproducttheoreminChapter3. /squaresolid
•Orthogonalgroupsarecalled Liegroups ;thatis,theydependoncontinuouslyvarying
parameters (the Euler angles and their generalization for higher dimensions); they are
compact because the angles vary over closed, finite intervals (containing the limit of
any converging sequence of angles). Unitary groups are also compact. Translations
formanoncompactgroupbecausethelimitoftranslationswithdistance d→∞isnot
partofthegroup.TheLorentzgroupis notcompacteither.
Homomorphism, Isomorphism
There may be a correspondence between the elements of two groups: one-to-one, two-to-
one, or many-to-one. If this correspondence preserves the group multiplication, we say
that the two groups are homomorphic . A most important homomorphic correspondence
betweentherotationgroup SO(3)andtheunitarygroup SU(2)isdevelopedinSection4.2.
If the correspondence is one-to-one, still preserving the group multiplication,3then the
groupsare isomorphic .
•If a group Gis homomorphic to a group of matrices G′, thenG′is called a represen-
tationofG.I fGandG′are isomorphic, the representation is called faithful. There
aremanyrepresentationsofgroups;theyare notunique.
3Supposetheelementsofonegrouparelabeled gi,theelementsofasecondgroup hi.Thengi↔hiisaone-to-onecorrespon-
dencefor all values of i.I fgigj=gkandhihj=hk,t h e ngkandhkmust be thecorresponding group elements.
244 Chapter 4 Group Theory
Example 4.1.2 ROTATIONS
Anotherinstructiveexampleforagroupisthesetofcounterclockwisecoordinaterotations
ofthree-dimensionalEuclideanspaceaboutits z-axis.FromChapter3weknowthatsucha
rotationisdescribedbyalineartransformationofthecoordinatesinvolvinga 3 ×3m a t r i x
made up of three rotations depending on the Euler angles. If the z-axis is fixed, the linear
transformation is through an angle ϕof thexy-coordinate system to a new orientation in
Eq.(1.8), Fig.1.6,andSection3.3:
x′
y′
z′
=Rz(ϕ)
x
y
z
≡
cosϕsinϕ0
−sinϕcosϕ0
00 1
x
y
z
(4.1)
involves only one angle of the rotation about the z-axis. As shown in Chapter 3, the linear
transformationoftwosuccessiverotationsinvolvestheproductofthematricescorrespond-
ingtothesumoftheangles.Theproductcorrespondstotworotations, Rz(ϕ1)Rz(ϕ2),and
is defined by rotating first by the angle ϕ2and then by ϕ1. According to Eq. (3.29), this
correspondstotheproductof theorthogonal 2 ×2 submatrices,
parenleftBigg
cosϕ1sinϕ1
−sinϕ1cosϕ1parenrightBiggparenleftBigg
cosϕ2sinϕ2
−sinϕ2cosϕ2parenrightBigg
=parenleftBigg
cos(ϕ1+ϕ2)sin(ϕ1+ϕ2)
−sin(ϕ1+ϕ2)cos(ϕ1+ϕ2)parenrightBigg
,(4.2)
using the addition formulas for the trigonometric functions. The unity in the lower right-
handcornerofthematrixinEq.(4.1)isalsoreproduceduponmultiplication.Theproductis
clearlyarotation,representedbytheorthogonalmatrixwithangle ϕ1+ϕ2.Theassociative
group multiplication corresponds to the associative matrix multiplication. It is commuta-
tive,orabelian,becausetheorderinwhichtheserotationsareperformeddoesnotmatter.
Theinverseoftherotationwithangle ϕisthatwithangle −ϕ.Theunitcorrespondstothe
angleϕ=0. Striking off the coordinate vectors in Eq. (4.1), we can associate the matrix
of the linear transformation with each rotation, which is a group multiplication preserving
one-to-one mapping, an isomorphism: The matrices form a faithful representation of the
rotationgroup.Theunityintheright-handcornerissuperfluousaswell,likethecoordinate
vectors, and may be deleted. This defines another isomorphism and representation by the
2×2 submatrices:
Rz(ϕ)=
cosϕsinϕ0
−sinϕcosϕ0
00 1
→R(ϕ)=parenleftBigg
cosϕsinϕ
−sinϕcosϕparenrightBigg
.(4.3)
The group’s name is SO(2), if the angle ϕvaries continuously from 0 to 2 π;SO(2) has
infinitelymanyelementsandiscompact.
Thegroupofrotations RzisobviouslyisomorphictothegroupofrotationsinEq.(4.3).
The unity with angle ϕ=0 and the rotation with ϕ=πform a finite subgroup. The finite
subgroupswithangles 2 πm/n,n anintegerand m=0,1,...,n−1a r ecyclic;thatis,the
rotationsR(2πm/n)=R(2π/n)m. /squaresolid
4.1 Introduction to Group Theory 245
In the following we shall discuss only the rotation groups SO(n)and unitary groups
SU(n)among the classical Lie groups. (More examples of finite groups will be given in
Section4.7.)
Representations — Reducible and Irreducible
The representation of group elements by matrices is a very powerful technique and has
been almost universally adopted by physicists. The use of matrices imposes no significant
restriction. It can be shown that the elements of any finite group and of the continuous
groups of Sections 4.2–4.4 may be represented by matrices. Examples are the rotations
describedinEq. (4.3).
To illustrate how matrix representations arise from a symmetry, consider the station-
ary Schrödinger equation (or some other eigenvalue equation, such as Ivi=Iivifor the
principalmomentsofinertiaofarigidbodyinclassicalmechanics,say),
Hψ=Eψ. (4.4)
Let us assume that the Hamiltonian Hstays invariant under a group Gof transformations
RinG(coordinate rotations, for example, for a central potential V(r)in the Hamiltonian
H); thatis,
HR=RHR−1=H,RH=HR. (4.5)
Now take a solution ψof Eq. (4.4) and “rotate” it: ψ→Rψ. ThenRψhas thesame
energyE becausemultiplyingEq.(4.4) by RandusingEq. (4.5) yields
RHψ=E(Rψ)=parenleftbig
RHR−1parenrightbig
Rψ=H(Rψ). (4.6)
Inotherwords,allrotatedsolutions Rψaredegenerate inenergyorformwhatphysicists
call amultiplet . For example, the spin-up and -down states of a bound electron in the
ground state of hydrogen form a doublet, and the states with projection quantum numbers
m=−l,−l+1,...,lof orbital angular momentum lform a multiplet with 2 l+1 basis
states.
Let us assume that this vector space Vψof transformed solutions has a finite dimen-
sionn.L e tψ1,ψ2,...,ψ nbe a basis. Since Rψjis a member of the multiplet, we can
expanditintermsof itsbasis,
Rψj=summationdisplay
krjkψk. (4.7)
Thus, with each RinGwe can associate a matrix (rjk). Just as in Example 4.1.2, two
successiverotationscorrespondtotheproductoftheirmatrices,sothismap R→(rjk)isa
representation of G. It is necessary for a representation to be irreducible that we can take
any element of Vψand, by rotating with allelementsRofG, transform it into allother
elements of Vψ. If not all elements of Vψare reached, then Vψsplits into a direct sum of
two or more vector subspaces, Vψ=V1⊕V2⊕···,which are mapped into themselves
by rotating their elements. For example, the 2 sstate and 2 pstates of principal quantum
numbern=2ofthehydrogenatomhavethesameenergy(thatis,aredegenerate)andform
246 Chapter 4 Group Theory
a reducible representation, because the 2 sstate cannot be rotated into the 2 pstates, and
viceversa(angularmomentumisconservedunderrotations).Inthiscasetherepresentation
is calledreducible .Then wecanfind abasis in Vψ(that is, there is a unitarymatrix U)s o
that
U(rjk)U†=
r10···
0r2···
......
(4.8)
forallRofG, andallmatrices (rjk)havesimilarblock-diagonal shape. Here r1,r2,...
arematricesoflowerdimensionthan (rjk)thatarelinedupalongthediagonalandthe 0’s
are matrices made up of zeros. We may say that the representation has been decomposed
intor1+r2+···alongwith Vψ=V1⊕V2⊕···.
The irreducible representations play a role in group theory that is roughly analogous to
the unit vectors of vector analysis. They are the simplest representations; all others can be
builtfromthem.(SeeSection4.4onClebsch–GordancoefficientsandYoungtableaux.)
Exercises
4.1.1 Showthatan n×northogonalmatrixhas n(n−1)/2 independentparameters.
Hint.Theorthogonalitycondition,Eq.(3.71), providesconstraints.
4.1.2 Showthatan n×nunitarymatrixhas n2−1 independentparameters.
Hint.Eachelementmaybecomplex,doublingthenumberofpossibleparameters.Some
oftheconstraintequationsarelikewisecomplexandcountastwoconstraints.
4.1.3 The special linear group SL(2) consists of all 2 ×2 matrices (with complex elements)
havingadeterminantof +1.Showthatsuchmatricesform agroup.
Note.T h eSL(2) group can be related to the full Lorentz group in Section 4.4, much as
theSU(2)groupisrelatedto SO(3).
4.1.4 Show that the rotations about the z-axis form a subgroup of SO(3). Is it an invariant
subgroup?
4.1.5 Show that if R,S,Tare elements of a group Gso thatRS=TandR→(rik),S→
(sik)is arepresentationaccordingtoEq.(4.7), then
(rik)(sik)=parenleftbigg
tik=summationdisplay
nrinsnkparenrightbigg
,
that is, group multiplication translates into matrix multiplication for any group repre-
sentation.
4.2 G ENERATORS OF CONTINUOUS GROUPS
Acharacteristicpropertyofcontinuousgroupsknownas Liegroups isthattheparameters
of a product element are analytic functions4of the parameters of the factors. The analytic
4Analytichere means having derivatives ofall orders.
4.2 Generators of Continuous Groups 247
natureofthefunctions(differentiability)allowsustodeveloptheconceptofgeneratorand
toreducethestudyofthewholegrouptoastudyofthegroupelementsintheneighborhood
oftheidentityelement.
Lie’s essential idea was to study elements Rin a group Gthat are infinitesimally close
to the unity of G. Let us consider the SO(2) group as a simple example. The 2 ×2r o -
tation matrices in Eq. (4.2) can be written in exponential form using the Euler identity,
Eq.(3.170a), as
R(ϕ)=parenleftBigg
cosϕsinϕ
−sinϕcosϕparenrightBigg
=12cosϕ+iσ2sinϕ=exp(iσ2ϕ). (4.9)
From the exponential form it is obvious that multiplication of these matrices is equivalent
toadditionof thearguments
R(ϕ2)R(ϕ1)=exp(iσ2ϕ2)exp(iσ2ϕ1)=expparenleftbig
iσ2(ϕ1+ϕ2)parenrightbig
=R(ϕ1+ϕ2).
Rotationscloseto1havesmallangle ϕ≈0.
This suggeststhatwelookfor anexponentialrepresentation
R=exp(iεS)=1+iεS+Oparenleftbig
ε2parenrightbig
,ε→0, (4.10)
for group elements RinGclose to the unity 1. The infinitesimal transformations are εS,
and theSare called generators of G. They form a linear space because multiplication
of the group elements Rtranslates into addition of generators S. The dimension of this
vectorspace(overthe complexnumbers) is the orderofG, thatis, thenumberof linearly
independentgeneratorsofthegroup.
IfRis a rotation, it does not change the volume element of the coordinate space that it
rotates,thatis, det (R)=1,andwemayuseEq. (3.171)toseethat
det(R)=expparenleftbig
trace(lnR)parenrightbig
=expparenleftbig
iεtrace(S)parenrightbig
=1
impliesεtrace(S)=0 and, upon dividing by the small but nonzero parameter ε, thatgen-
eratorsaretraceless ,
trace(S)=0. (4.11)
Thisis thecasenotonlyfor therotationgroups SO(n)butalsofor unitarygroups SU(n).
IfRofGin Eq. (4.10) is unitary, then S†=Sis Hermitian, which is also the case for
SO(n)andSU(n). Thisexplainswhytheextra ihasbeeninsertedinEq. (4.10).
Next we go around the unity in four steps, similar to parallel transport in differential
geometry.Weexpandthegroupelements
Ri=exp(iεiSi)=1+iεiSi−1
2ε2
iS2
i+···,
R−1
i=exp(−iεiSi)=1−iεiSi−1
2ε2
iS2
i+···,(4.12)
to second order in the small group parameter εibecause the linear terms and several
quadratictermsallcancelintheproduct(Fig. 4.1)
R−1
iR−1
jRiRj=1+εiεj[Sj,Si]+···,
=1+εiεjsummationdisplay
kck
jiSk+···, (4.13)
248 Chapter 4 Group Theory
FIGURE 4.1IllustrationofEq. (4.13).
when Eq. (4.12) is substituted into Eq. (4.13). The last line holds because the product in
Eq. (4.13) is again a group element, Rij, close to the unity in the group G. Hence its
exponent must be a linear combination of the generators Sk, and its infinitesimal group
parameter has to be proportional to the product εiεj. Comparing both lines in Eq. (4.13)
wefindthe closurerelationof thegeneratorsoftheLiegroup G,
[Si,Sj]=summationdisplay
kck
ijSk. (4.14)
The coefficients ck
ijare the structure constants of the group G. Since the commutator in
Eq.(4.14) isantisymmetricin iandj, so arethestructureconstantsinthelowerindices,
ck
ij=−ck
ji. (4.15)
If the commutator in Eq. (4.14) is taken as a multiplication law of generators, we see
thatthevectorspaceofgeneratorsbecomesanalgebra,the Liealgebra Gofthegroup G.
Analgebrahastwogroupstructures,acommutativeproductdenotedbya +symbol(this
istheadditionofinfinitesimalgeneratorsofaLiegroup)andamultiplication(thecommu-
tatorofgenerators).Oftenanalgebraisavectorspacewithamultiplication,suchasaring
ofsquarematrices.For SU(l+1)theLiealgebraiscalled Al,forSO(2l+1)itisBl,and
forSO(2l)itisDl,wherel=1,2,...isapositiveinteger,latercalledthe rankoftheLie
groupGorof itsalgebra G.
Finally,the Jacobiidentity holdsfor alldoublecommutators
bracketleftbig
[Si,Sj],Skbracketrightbig
+bracketleftbig
[Sj,Sk],Sibracketrightbig
+bracketleftbig
[Sk,Si],Sjbracketrightbig
=0, (4.16)
whichiseasilyverifiedusingthedefinitionofanycommutator [A,B]≡AB−BA.When
Eq.(4.14) issubstitutedintoEq. (4.16)wefindanotherconstraintonstructureconstants,
summationdisplay
mbraceleftbig
cm
ij[Sm,Sk]+cm
jk[Sm,Si]+cm
ki[Sm,Sj]bracerightbig
=0. (4.17)
UponinsertingEq. (4.14) again,Eq. (4.17) impliesthat
summationdisplay
mnbraceleftbig
cm
ijcn
mkSn+cm
jkcn
miSn+cm
kicn
mjSnbracerightbig
=0, (4.18)
4.2 Generators of Continuous Groups 249
wherethecommonfactor Sn(andthesumover n)maybedroppedbecausethegenerators
arelinearlyindependent.Hence
summationdisplay
mbraceleftbig
cm
ijcn
mk+cm
jkcn
mi+cm
kicn
mjbracerightbig
=0. (4.19)
The relations (4.14), (4.15), and (4.19) form the basis of Lie algebras from which finite
elementsoftheLiegroupnearits unitycanbereconstructed.
ReturningtoEq.(4.5),theinverseof RisR−1=exp(−iεS).Weexpand HRaccording
totheBaker–Hausdorff formula,Eq.(3.172),
H=HR=exp(iεS)Hexp(−iεS)=H+iε[S,H]−1
2ε2bracketleftbig
S[S,H]bracketrightbig
+··· (4.20)
We drop Hfrom Eq. (4.20), divide by the small (but nonzero), ε, and let ε→0. Then
Eq.(4.20) impliesthatthecommutator
[S,H]=0. (4.21)
IfSandHareHermitianmatrices,Eq.(4.21)impliesthat SandHcanbesimultaneously
diagonalized and have common eigenvectors (for matrices, see Section 3.5; for operators,
see Schur’s lemma in Section 4.3). If SandHare differential operators like the Hamil-
tonian and orbital angular momentum in quantum mechanics, then Eq. (4.21) implies that
SandHhave common eigenfunctions and that the degenerate eigenvalues of Hcan be
distinguished by the eigenvalues of the generators S. These eigenfunctions and eigenval-
ues,s, are solutions of separate differential equations, Sψs=sψs, so group theory (that
is, symmetries) leads to a separation of variables for a partial differential equation that is
invariantunderthetransformationsofthegroup.
Forexample,letus takethesingle-particleHamiltonian
H=−¯h2
2m1
r2∂
∂rr2∂
∂r+¯h2
2mr2L2+V(r)
that is invariant under SO(3) and, therefore, a function of the radial distance r, the radial
gradient, and the rotationally invariant operator L2ofSO(3). Upon replacing the orbital
angularmomentumoperator L2byitseigenvalue l(l+1)weobtaintheradialSchrödinger
equation(ODE),
HRl(r)=bracketleftbigg
−¯h2
2m1
r2d
drr2d
dr+¯h2l(l+1)
2mr2+V(r)bracketrightbigg
Rl(r)=ElRl(r),
whereRl(r)istheradialwavefunction.
For cylindrical symmetry, the invariance of Hunder rotations about the z-axis would
requireHtobeindependentoftherotationangle ϕ, leadingtotheODE
HRm(z,ρ)=EmRm(z,ρ),
withmtheeigenvalueof Lz=−i∂/∂ϕ,thez-componentoftheorbitalangularmomentum
operator. For more examples, see the separation of variables method for partial differen-
tial equations in Section 9.3 and special functions in Chapter 12. This is by far the most
importantapplicationofgrouptheoryinquantummechanics.
In the next subsections we shall study orthogonal and unitary groups as examples to
understandbetterthegeneralconceptsof thissection.
250 Chapter 4 Group Theory
Rotation Groups SO(2) andSO(3)
ForSO(2)asdefinedbyEq.(4.3)thereisonlyonelinearlyindependentgenerator, σ2,and
the order of SO( 2 )i s1 .W eg e t σ2from Eq. (4.9) by differentiation at the unity of SO(2),
thatis,ϕ=0,
−idR(ϕ)/dϕ|ϕ=0=−iparenleftBigg
−sinϕcosϕ
−cosϕ−sinϕparenrightBiggvextendsinglevextendsinglevextendsinglevextendsingle
ϕ=0=−iparenleftBigg
01
−10parenrightBigg
=σ2.(4.22)
For the rotations Rz(ϕ)about the z-axis described by Eq. (4.1), the generator is given
by
−idRz(ϕ)/dϕ|ϕ=0=Sz=
0−i0
i00
000
, (4.23)
wherethefactor iisinsertedtomake SzHermitian.Therotation Rz(δϕ)throughaninfin-
itesimalangle δϕmaythenbeexpandedtofirst orderinthesmall δϕas
Rz(δϕ)=13+iδϕSz. (4.24)
Afiniterotation R(ϕ)maybecompoundedofsuccessiveinfinitesimalrotations
Rz(δϕ1+δϕ2)=(1+iδϕ1Sz)(1+iδϕ2Sz). (4.25)
Letδϕ=ϕ/NforNrotations,with N→∞. Then
Rz(ϕ)=lim
N→∞bracketleftbig
1+(iϕ/N)SzbracketrightbigN=exp(iϕSz). (4.26)
This form identifies Szas the generator of the group Rz, an abelian subgroup of SO(3),
the group of rotations in three dimensions with determinant +1. Each 3×3m a t r i xRz(ϕ)
isorthogonal,henceunitary,and trace (Sz)=0,inaccordwithEq.(4.11).
Bydifferentiationof thecoordinaterotations
Rx(ψ)=
10 0
0 cosψsinψ
0−sinψcosψ
,Ry(θ)=
cosθ0−sinθ
01 0
sinθ0 cosθ
,(4.27)
wegetthegenerators
Sx=
00 0
00−i
0i0
,Sy=
00i
00 0
−i00
(4.28)
ofRx(Ry), thesubgroupofrotationsaboutthe x-(y-)axis.
4.2 Generators of Continuous Groups 251
Rotation of Functions and Orbital Angular Momentum
In the foregoing discussion the group elements are matrices that rotate the coordinates.
Any physical system being described is held fixed. Now let us hold the coordinates fixed
and rotate a function ψ(x,y,z) relative to our fixed coordinates. With Rto rotate the
coordinates,
x′=Rx, (4.29)
wedefineRonψby
Rψ(x,y,z)=ψ′(x,y,z)≡ψ(x′). (4.30)
In words, Roperates on the function ψ, creating a new function ψ′that is numerically
equal to ψ(x′), wherex′are the coordinates rotated by R.I fRrotates the coordinates
counterclockwise,theeffect of Ris torotatethepatternof thefunction ψclockwise.
Returning to Eqs. (4.30) and (4.1), consider an infinitesimal rotation again, ϕ→δϕ.
Then,using RzEq.(4.1), weobtain
Rz(δϕ)ψ(x,y,z) =ψ(x+yδϕ,y−xδϕ,z). (4.31)
Therightsidemaybeexpandedtofirstorder inthesmall δϕtogive
Rz(δϕ)ψ(x,y,z) =ψ(x,y,z)−δϕ{x∂ψ/∂y−y∂ψ/∂x}+O(δϕ)2
=(1−iδϕLz)ψ(x,y,z), (4.32)
the differential expression in curly brackets being the orbital angular momentum iLz(Ex-
ercise1.8.7). Sincearotationof first ϕandthenδϕaboutthe z-axis isgivenby
Rz(ϕ+δϕ)ψ=Rz(δϕ)Rz(ϕ)ψ=(1−iδϕLz)Rz(ϕ)ψ, (4.33)
wehave(as anoperatorequation)
dRz
dϕ=lim
δϕ→0Rz(ϕ+δϕ)−Rz(ϕ)
δϕ=−iLzRz(ϕ). (4.34)
Inthisform Eq.(4.34) integratesimmediatelyto
Rz(ϕ)=exp(−iϕLz). (4.35)
Note thatRz(ϕ)rotates functions (clockwise) relative to fixed coordinates and that Lzis
thezcomponent of the orbital angular momentum L. The constant of integration is fixed
bytheboundarycondition Rz(0)=1.
AssuggestedbyEq.(4.32), Lzisconnectedto Szby
Lz=(x,y,z)Sz
∂/∂x
∂/∂y
∂/∂z
=−iparenleftbigg
x∂
∂y−y∂
∂xparenrightbigg
, (4.36)
soLx,Ly,andLzsatisfy thesamecommutationrelations,
[Li,Lj]=iεijkLk, (4.37)
asSx,Sy, andSzandyieldthesamestructureconstants iεijkofSO(3).
252 Chapter 4 Group Theory
SU(2) —SO(3) Homomorphism
Since unitary 2 ×2 matrices transform complex two-dimensional vectors preserving their
norm, they represent the most general transformations of (a basis in the Hilbert space of)
spin1
2wavefunctionsinnonrelativisticquantummechanics.Thebasisstatesofthissystem
areconventionallychosentobe
|↑/angbracketright=parenleftbigg1
0parenrightbigg
,|↓/angbracketright=parenleftbigg0
1parenrightbigg
,
corresponding to spin1
2up and down states, respectively. We can show that the special
unitarygroupSU(2) of unitary 2 ×2 matrices with determinant +1 has all three Pauli
matricesσias generators (while the rotations of Eq. (4.3) form a one-dimensional abelian
subgroup).So SU(2)isoforder3anddependsonthreerealcontinuousparameters ξ,η,ζ,
which are often called the Cayley–Klein parameters. To construct its general element, we
start with the observation that orthogonal 2 ×2 matrices are real unitary matrices, so they
formasubgroupof SU(2). Wealsoseethat
parenleftBigg
eiα0
0e−iαparenrightBigg
isunitaryforrealangle αwithdeterminant +1.Sothesesimpleandmanifestlyunitaryma-
trices form another subgroup of SU(2) from which we can obtain all elements of SU(2),
that is, the general 2 ×2 unitary matrix of determinant +1. For a two-component spin1
2
wavefunctionofquantummechanicsthisdiagonalunitarymatrixcorrespondstomultipli-
cation of the spin-up wave function with a phase factor eiαand the spin-down component
with the inverse phase factor. Using the real angle ηinstead of ϕfor the rotation matrix
andthenmultiplyingbythediagonalunitarymatrices,weconstructa 2 ×2 unitarymatrix
thatdependsonthreeparametersandclearlyis amoregeneralelementof SU(2):
parenleftBigg
eiα0
0e−iαparenrightBiggparenleftBigg
cosηsinη
−sinηcosηparenrightBiggparenleftBigg
eiβ0
0e−iβparenrightBigg
=parenleftBigg
eiαcosηeiαsinη
−e−iαsinηe−iαcosηparenrightBiggparenleftBigg
eiβ0
0e−iβparenrightBigg
=parenleftBigg
ei(α+β)cosηei(α−β)sinη
−e−i(α−β)sinηe−i(α+β)cosηparenrightBigg
.
Defining α+β≡ξ,α−β≡ζ,wehaveinfactconstructedthegeneralelementof SU(2):
U(ξ,η,ζ)=parenleftBigg
eiξcosηeiζsinη
−e−iζsinηe−iξcosηparenrightBigg
=parenleftBigg
ab
−b∗a∗parenrightBigg
. (4.38)
To see this, we write the general SU(2) element as U=parenleftbigab
cdparenrightbig
with complex numbers
a,b,c,d so that det (U)=1. Writing unitarity, U†=U−1, and using Eq. (3.50) for the
4.2 Generators of Continuous Groups 253
inverseweobtainparenleftBigg
a∗c∗
b∗d∗parenrightBigg
=parenleftBigg
d−b
−caparenrightBigg
,
implying c=−b∗,d=a∗, as shown in Eq. (4.38). It is easy to check that the determinant
det(U)=1 andthat U†U=1=UU†hold.
Togetthegenerators,wedifferentiate(anddropirrelevantoverallfactors):
−i∂U/∂ξ|ξ=0,η=0=parenleftBigg
10
0−1parenrightBigg
=σ3, (4.39a)
−i∂U/∂η|η=0,ζ=0=parenleftBigg
0−i
i0parenrightBigg
=σ2. (4.39b)
To avoid a factor 1 /sinηforη→0 upon differentiating with respect to ζ,w eu s ei n -
stead the right-hand side of Eq. (4.38) for Ufor pure imaginary b=iβwithβ→0, so
a=radicalbig
1−β2from|a|2+|b|2=a2+β2=1. Differentiating such a U, we get the third
generator,
−i∂
∂βparenleftBiggradicalbig
1−β2iβ
iβradicalbig
1−β2parenrightBiggvextendsinglevextendsinglevextendsinglevextendsingle
β=0=−i
−β√
1−β2i
−iβ√
1−β2
vextendsinglevextendsinglevextendsinglevextendsingle
β=0=parenleftBigg
01
10parenrightBigg
=σ1.
(4.39c)
ThePaulimatricesarealltracelessandHermitian.
With the Pauli matrices as generators, the elements U1,U2,U3ofSU(2) may be gener-
atedby
U1=exp(ia1σ1/2),U2=exp(ia2σ2/2),U3=exp(ia3σ3/2).(4.40)
The three parameters aiare real. The extra factor 1 /2 is present in the exponents to make
Si=σi/2 satisfythesamecommutationrelations,
[Si,Sj]=iεijkSk, (4.41)
astheangularmomentuminEq. (4.37).
To connect and compare our results, Eq. (4.3) gives a rotation operator for rotat-
ing the Cartesian coordinates in the three-space R3. Using the angular momentum ma-
trixS3, we have as the corresponding rotation operator in two-dimensional (complex)
spaceRz(ϕ)=exp(iϕσ3/2).Forrotatingthetwo-componentvectorwavefunction(spinor)
or a spin 1 /2 particle relative to fixed coordinates, the corresponding rotation operator is
Rz(ϕ)=exp(−iϕσ3/2)accordingtoEq. (4.35).
Moregenerally,usinginEq. (4.40) theEuleridentity,Eq.(3.170a), weobtain
Uj=cosparenleftbiggaj
2parenrightbigg
+iσjsinparenleftbiggaj
2parenrightbigg
. (4.42)
Heretheparameter ajappearsasanangle,thecoefficientofanangularmomentummatrix-
likeϕinEq.(4.26).TheselectionofPaulimatricescorrespondstotheEuleranglerotations
describedinSection3.3.
254 Chapter 4 Group Theory
FIGURE 4.2Illustrationof
M′=UMU†inEq.(4.43).
As just seen, the elements of SU(2) describe rotations in a two-dimensional complex
spacethatleave |z1|2+|z2|2invariant.Thedeterminantis +1.Therearethreeindependent
real parameters. Our real orthogonal group SO(3) clearly describes rotations in ordinary
three-dimensional space with the important characteristic of leaving x2+y2+z2invari-
ant.Also,therearethreeindependentrealparameters.Therotationinterpretationsandthe
equality of numbers of parameters suggest the existence of some correspondence between
thegroups SU(2) andSO(3). Herewedevelopthiscorrespondence.
Theoperationof SU(2)onamatrixisgivenbyaunitarytransformation,Eq.(4.5),with
R=UandFig.4.2:
M′=UMU†. (4.43)
TakingMto be a 2×2 matrix, we note that any 2 ×2 matrix may be written as a linear
combination of the unit matrix and the three Pauli matrices of Section 3.4. Let Mbe the
zero-tracematrix,
M=xσ1+yσ2+zσ3=parenleftBigg
zx−iy
x+iy−zparenrightBigg
, (4.44)
theunitmatrixnotentering.Sincethetraceisinvariantunderaunitarysimilaritytransfor-
mation(Exercise3.3.9), M′musthavethesameform,
M′=x′σ1+y′σ2+z′σ3=parenleftBigg
z′x′−iy′
x′+iy′−z′parenrightBigg
. (4.45)
The determinant is also invariant under a unitary transformation (Exercise 3.3.10). There-
fore
−parenleftbig
x2+y2+z2parenrightbig
=−parenleftbig
x′2+y′2+z′2parenrightbig
, (4.46)
orx2+y2+z2is invariant under this operation of SU(2), just as with SO(3). Operations
ofSU(2) onMmust produce rotations of the coordinates x,y,zappearing therein. This
suggeststhat SU(2) andSO(3) maybeisomorphicor atleasthomomorphic.
4.2 Generators of Continuous Groups 255
Weapproachtheproblemofwhatthisoperationof SU(2)correspondstobyconsidering
specialcases. ReturningtoEq. (4.38),let a=eiξandb=0,or
U3=parenleftBigg
eiξ0
0e−iξparenrightBigg
. (4.47)
InanticipationofEq. (4.51), this Uis givenasubscript 3.
Carrying out a unitary similarity transformation, Eq. (4.43), on each of the three Pauli
σ’s ofSU(2), wehave
U3σ1U†
3=parenleftBigg
eiξ0
0e−iξparenrightBiggparenleftBigg
01
10parenrightBiggparenleftBigg
e−iξ0
0eiξparenrightBigg
=parenleftBigg
0e2iξ
e−2iξ0parenrightBigg
. (4.48)
Wereexpressthisresultintermsof thePauli σi,as inEq.(4.44), toobtain
U3xσ1U†
3=xσ1cos2ξ−xσ2sin2ξ. (4.49)
Similarly,
U3yσ2U†
3=yσ1sin2ξ+yσ2cos2ξ,
U3zσ3U†
3=zσ3. (4.50)
Fromthesedoubleangleexpressionsweseethatweshouldstartwitha
halfangle :ξ=α/2. Then, adding Eqs. (4.49) and (4.50) and comparing with Eqs. (4.44)
and(4.45), weobtain
x′=xcosα+ysinα
y′=−xsinα+ycosα (4.51)
z′=z.
The 2×2 unitary transformation using U3(α)is equivalent to the rotation operator R(α)
ofEq. (4.3).
Thecorrespondenceof
U2(β)=parenleftBigg
cosβ/2sinβ/2
−sinβ/2 cosβ/2parenrightBigg
(4.52)
andRy(β)andof
U1(ϕ)=parenleftBigg
cosϕ/2isinϕ/2
isinϕ/2 cosϕ/2parenrightBigg
(4.53)
andR1(ϕ)followsimilarly.Notethat Uk(ψ)hasthegeneralform
Uk(ψ)=12cosψ/2+iσksinψ/2, (4.54)
wherek=1,2,3.
256 Chapter 4 Group Theory
Thecorrespondence
U3(α)=parenleftBigg
eiα/20
0e−iα/2parenrightBigg
↔
cosαsinα0
−sinαcosα0
00 1
=Rz(α) (4.55)
is not a simple one-to-one correspondence. Specifically, as αinRzranges from 0 to 2 π,
theparameterin U3,α/2,goesfrom0to π. Wefind
Rz(α+2π)=Rz(α)
U3(α+2π)=parenleftBigg
−eiα/20
0−e−iα/2parenrightBigg
=−U3(α). (4.56)
Therefore bothU3(α)andU3(α+2π)=−U3(α)correspond to Rz(α). The correspon-
dence is 2 to 1, or SU(2) andSO(3) arehomomorphic . This establishment of the corre-
spondencebetweentherepresentationsof SU(2)andthoseof SO(3)meansthattheknown
representationsof SU(2) automaticallyprovideus withtherepresentationsof SO(3).
Combiningthevariousrotations,wefindthataunitarytransformationusing
U(α,β,γ)=U3(γ)U2(β)U3(α) (4.57)
correspondstothegeneralEuler rotation Rz(γ)Ry(β)Rz(α). Bydirectmultiplication,
U(α,β,γ)=parenleftBigg
eiγ/20
0e−iγ/2parenrightBiggparenleftBigg
cosβ/2sinβ/2
−sinβ/2 cosβ/2parenrightBiggparenleftBigg
eiα/20
0e−iα/2parenrightBigg
=parenleftBigg
ei(γ+α)/2cosβ/2ei(γ−α)/2sinβ/2
−e−i(γ−α)/2sinβ/2e−i(γ+α)/2cosβ/2parenrightBigg
. (4.58)
Thisis ouralternategeneralform, Eq.(4.38), with
ξ=(γ+α)/2,η=β/2,ζ=(γ−α)/2. (4.59)
Thus,from Eq.(4.58) wemayidentifytheparametersofEq. (4.38)as
a=ei(γ+α)/2cosβ/2
b=ei(γ−α)/2sinβ/2. (4.60)
SU(2)-Isospin and SU(3)-Flavor Symmetry
The application of group theory to “elementary” particles has been labeled by Wigner
the third stage of group theory and physics. The first stage was the search for the 32
crystallographic point groups and the 230 space groups giving crystal symmetries—
Section 4.7. The second stage was a search for representations such as of SO(3) and
SU(2)—Section4.2.Nowinthisstage,physicistsarebacktoa searchforgroups.
In the 1930s to 1960s the study of strongly interacting particles of nuclear and high-
energyphysicsledtothe SU(2)isospingroupandthe SU(3)flavorsymmetry.Inthe1930s,
aftertheneutronwasdiscovered,Heisenbergproposedthatthenuclearforceswerecharge
4.2 Generators of Continuous Groups 257
Table 4.1 BaryonswithSpin1
2EvenParity
Mass(MeV) YI I 3
/Xi1−1321.32 −1
2
/Xi1 −11
2
/Xi101314.9 +1
2
/Sigma1−1197.43 −1
/Sigma1/Sigma101192.55 0 1 0
/Sigma1+1189.37 +1
/Lambda1/Lambda1 1115.63 0 0 0
n 939.566 −1
2
N 11
2
p 938.272 +1
2
independent. The neutron mass differs from that of the proton by only 1.6%. If this tiny
mass difference is ignored, the neutron and proton may be considered as two charge (or
isospin) states of a doublet, called the nucleon. The isospin Ihasz-projection I3=1/2
for the proton and I3=−1/2 for the neutron. Isospin has nothing to do with spin (the
particle’s intrinsic angular momentum), but the two-component isospin state obeys the
same mathematical relations as the spin 1 /2 state. For the nucleon, I=τ/2 are the usual
Paulimatricesandthe ±1/2isospinstatesareeigenvectorsofthePaulimatrix τ3=parenleftbig10
0−1parenrightbig
.
Similarly, the three charge states of the pion ( π+,π0,π−) form a triplet. The pion is the
lightest of all strongly interacting particles and is the carrier of the nuclear force at long
distances,muchlikethephotonisthatoftheelectromagneticforce.Thestronginteraction
treats alike members of these particle families, or multiplets, and conserves isospin. The
symmetryisthe SU(2) isospingroup.
By the 1960s particles produced as resonances by accelerators had proliferated. The
eight shown in Table 4.1 attracted particular attention.5The relevant conserved quantum
numbers that are analogs and generalizations of LzandL2fromSO(3) areI3andI2for
isospinand Yforhypercharge .Particlesmaybegroupedintochargeorisospinmultiplets.
Then the hypercharge may be taken as twice the average charge of the multiplet. For the
nucleon, that is, the neutron–proton doublet, Y=2·1
2(0+1)=1. The hypercharge and
isospin values are listed in Table 4.1 for baryons like the nucleon and its (approximately
degenerate) partners. They form an octet, as shown in Fig. 4.3, after which the corre-
sponding symmetry is called the eightfold way . In 1961 Gell-Mann, and independently
Ne’eman, suggested that the strong interaction should be (approximately) invariant under
athree-dimensionalspecialunitarygroup, SU(3),thatis, has SU(3)flavorsymmetry .
The choice of SU(3) was based first on the two conserved and independent quantum
numbers, H1=I3andH2=Y(thatis,generatorswith [I3,Y]=0,notCasimirinvariants;
see the summary in Section 4.3) that call for a group of rank 2. Second, the group had
to have an eight-dimensional representation to account for the nearly degenerate baryons
and four similar octets for the mesons. In a sense, SU(3) is the simplest generalization of
SU(2)isospin.Threeofitsgeneratorsarezero-traceHermitian3 ×3matricesthatcontain
5Allmasses aregiven inenergyunits, 1MeV =106eV.
258 Chapter 4 Group Theory
FIGURE 4.3Baryonoctetweight
diagramfor SU(3).
the 2×2 isospinPaulimatrices τiintheupperleftcorner,
λi=
τi0
0
000
,i=1,2,3. (4.61a)
Thus, theSU(2)-isospin group is a subgroup of SU(3)-flavor with I3=λ3/2. Four other
generators have the off-diagonal 1’s of τ1, and−i,iofτ2in all other possible locations to
form zero-traceHermitian 3 ×3 matrices,
λ4=
001
000
100
,λ 5=
00−i
00 0
i00
,
λ6=
000
001
010
,λ 7=
00 0
00−i
0i0
.(4.61b)
The second diagonal generator has the two-dimensional unit matrix 1 2in the upper left
corner, which makes it clearly independent of the SU(2)-isospin subgroup because of its
nonzerotraceinthatsubspace,and −2 inthethirddiagonalplacetomakeittraceless,
λ8=1√
3
10 0
01 0
00−2
. (4.61c)
4.2 Generators of Continuous Groups 259
FIGURE 4.4Baryonmasssplitting.
Altogether there are 32−1=8 generators for SU(3), which has order 8. From the com-
mutatorsof thesegeneratorsthestructureconstantsof SU(3) caneasilybeobtained.
Returning to the SU(3) flavor symmetry, we imagine the Hamiltonian for our eight
baryonstobecomposedofthreeparts:
H=Hstrong+Hmedium+Helectromagnetic . (4.62)
The first part, Hstrong, has theSU(3) symmetry and leads to the eightfold degeneracy.
Introduction of the symmetry-breaking term, Hmedium, removes part of the degeneracy,
giving the four isospin multiplets (/Xi1−,/Xi10),(/Sigma1−,/Sigma10,/Sigma1+),/Lambda1, andN=(p,n)different
masses. These are still multiplets because HmediumhasSU(2)-isospin symmetry. Finally,
the presence of charge-dependent forces splits the isospin multiplets and removes the last
degeneracy.ThisimaginedsequenceisshowninFig.4.4.
The octet representation is not the simplest SU(3) representation. The simplest repre-
sentationsarethetriangularonesshowninFig.4.5,fromwhichallotherscanbegenerated
by generalized angular momentum coupling (see Section 4.4 on Young tableaux). The
fundamental representation in Fig. 4.5a contains the u(up),d(down), and s(strange)
quarks, and Fig. 4.5b contains the corresponding antiquarks. Since the meson octets can
be obtained from the quark representations as q¯q, with 32=8+1 states, this suggests
thatmesonscontainquarks(andantiquarks)astheirconstituents(seeExercise4.4.3).The
resulting quark model gives a successful description of hadronic spectroscopy. The reso-
lution of its problem with the Pauli exclusion principle eventually led to the SU(3)-color
gaugetheoryofthe stronginteraction calledquantumchromodynamics (QCD).
Tokeepgrouptheoryanditsveryrealaccomplishmentinproperperspective,weshould
emphasizethatgrouptheoryidentifiesandformalizessymmetries.It classifies(andsome-
timespredicts)particles.ButasidefromsayingthatonepartoftheHamiltonianhas SU(2)
260 Chapter 4 Group Theory
FIGURE 4.5(a) Fundamentalrepresentationof SU(3),theweightdiagramfor
theu,d,squarks;(b)weightdiagramfor theantiquarks ¯u,¯d,¯s.
symmetryandanotherparthas SU(3)symmetry,grouptheorysaysnothingaboutthepar-
ticleinteraction.Rememberthatthestatementthattheatomicpotentialissphericallysym-
metrictellsusnothingabouttheradialdependenceofthepotentialorofthewavefunction.
Incontrast,inagaugetheorytheinteractionismediatedbyvectorbosons(likethephoton
in quantum electrodynamics) and uniquely determined by the gauge covariant derivative
(seeSection1.13).
Exercises
4.2.1 (i) Show that the Pauli matrices are the generators of SU(2) without using the para-
meterization of the general unitary 2 ×2 matrix in Eq. (4.38). (ii) Derive the eight
independent generators λiofSU(3) similarly. Normalize them so that tr (λiλj)=2δij.
Thendeterminethestructureconstantsof SU(3).
Hint.Theλiare tracelessandHermitian 3 ×3 matrices.
(iii)ConstructthequadraticCasimirinvariantof SU(3).
Hint.Workbyanalogywith σ2
1+σ2
2+σ2
3ofSU(2) orL2ofSO(3).
4.2.2 Provethatthegeneralform ofa 2 ×2 unitary,unimodularmatrixis
U=parenleftBigg
ab
−b∗a∗parenrightBigg
witha∗a+b∗b=1.
4.2.3 Determinethree SU(2) subgroupsof SU(3).
4.2.4 Atranslation operatorT(a)convertsψ(x)toψ(x+a),
T(a)ψ(x)=ψ(x+a).
4.3 Orbital Angular Momentum 261
In terms of the (quantum mechanical) linear momentum operator px=−id/dx,s h o w
thatT(a)=exp(iapx), thatis,pxisthegeneratorof translations.
Hint.Expand ψ(x+a)as aTaylorseries.
4.2.5 Consider the general SU(2) element Eq. (4.38) to be built up of three Euler rotations:
(i) a rotation of a/2 about the z-axis, (ii) a rotation of b/2 about the new x-axis, and
(iii) a rotation of c/2 about the new z-axis. (All rotations are counterclockwise.) Using
thePauli σgenerators,showthattheserotationanglesare determinedby
a=ξ−ζ+π
2=α+π
2
b=2η=β
c=ξ+ζ−π
2=γ−π
2.
Note.Theangles aandbhereare notthe aandbofEq. (4.38).
4.2.6 Rotate a nonrelativistic wave function ˜ψ=(ψ↑,ψ↓)of spin 1/2 about the z-axis by
asmallangle dθ.Findthecorrespondinggenerator.
4.3 O RBITAL ANGULAR MOMENTUM
The classical concept of angular momentum, Lclass=r×p, is presented in Section 1.4
tointroducethecrossproduct.FollowingtheusualSchrödingerrepresentationofquantum
mechanics,theclassicallinearmomentum pisreplacedbytheoperator −i∇.Thequantum
mechanicalorbitalangularmomentum operator becomes6
LQM=−ir×∇. (4.63)
This is used repeatedly in Sections 1.8, 1.9, and 2.4 to illustrate vector differential oper-
ators. From Exercise 1.8.8 the angular momentum components satisfy the commutation
relations
[Li,Lj]=iεijkLk. (4.64)
Theεijkis the Levi-Civita symbol of Section 2.9. A summation over the index kis under-
stood.
Thedifferentialoperatorcorrespondingtothesquareoftheangularmomentum
L2=L·L=L2
x+L2
y+L2
z (4.65)
maybedeterminedfrom
L·L=(r×p)·(r×p), (4.66)
which is the subject of Exercises 1.9.9 and 2.5.17(b). Since L2as a scalar product is in-
variant under rotations, that is, a rotational scalar, we expect [L2,Li]=0, which can also
beverifieddirectly.
Equation(4.64)presentsthebasiccommutationrelationsofthecomponentsofthequan-
tummechanicalangularmomentum.Indeed,withintheframeworkofquantummechanics
and group theory, these commutation relations define an angular momentum operator. We
shall use them now to construct the angular momentum eigenstates and find the eigenval-
ues.FortheorbitalangularmomentumthesearethesphericalharmonicsofSection12.6.
6For simplicity, ¯his setequal to 1. This means that theangular momentum is measured in units of ¯h.
262 Chapter 4 Group Theory
Ladder Operator Approach
Letusstartwithageneralapproach,wheretheangularmomentum Jweconsidermayrep-
resentanorbitalangularmomentum L,aspin σ/2,oratotalangularmomentum L+σ/2,
etc.Weassumethat
1.Jis anHermitianoperatorwhosecomponentssatisfythecommutationrelations
[Ji,Jj]=iεijkJk,bracketleftbig
J2,Jibracketrightbig
=0. (4.67)
Otherwise Jisarbitrary. (SeeExercise4.3.l.)
2.|λM/angbracketrightissimultaneouslyanormalizedeigenfunction(oreigenvector)of Jzwitheigen-
valueMandaneigenfunction7ofJ2,
Jz|λM/angbracketright=M|λM/angbracketright,J2|λM/angbracketright=λ|λM/angbracketright,/angbracketleftλM|λM/angbracketright=1.(4.68)
We shall show that λ=J(J+1)and then find other properties of the |λM/angbracketright. The treat-
mentwillillustratethegeneralityandpowerofoperatortechniques,particularlytheuseof
ladderoperators.8
Theladderoperators aredefinedas
J+=Jx+iJy,J−=Jx−iJy. (4.69)
Intermsof theseoperators J2mayberewrittenas
J2=1
2(J+J−+J−J+)+J2
z. (4.70)
Fromthecommutationrelations,Eq.(4.67), wefind
[Jz,J+]=+J+,[Jz,J−]=−J−,[J+,J−]=2Jz. (4.71)
SinceJ+commuteswith J2(Exercise 4.3.1),
J2parenleftbig
J+|λM/angbracketrightparenrightbig
=J+parenleftbig
J2|λM/angbracketrightparenrightbig
=λparenleftbig
J+|λM/angbracketrightparenrightbig
. (4.72)
Therefore, J+|λM/angbracketrightis still an eigenfunction of J2with eigenvalue λ, and similarly for
J−|λM/angbracketright. Butfrom Eq. (4.71),
JzJ+=J+(Jz+1), (4.73)
or
Jzparenleftbig
J+|λM/angbracketrightparenrightbig
=J+(Jz+1)|λM/angbracketright=(M+1)J+|λM/angbracketright. (4.74)
7That|λM/angbracketrightcan be an eigenfunction of bothJzandJ2follows from[Jz,J2]=0 in Eq. (4.67). For SU(2),/angbracketleftλM|λM/angbracketrightis the
scalar product (of the bra and ket vector or spinors) in the bra-ket notation introduced in Section 3.1. For SO(3),|λM/angbracketrightis a
functionY(θ,ϕ)and|λM′/angbracketrightisafunction Y′(θ,ϕ)andthematrixelement /angbracketleftλM|λM′/angbracketright≡integraltext2π
ϕ=0integraltextπ
θ=0Y∗(θ,ϕ)Y′(θ,ϕ)sinθdθdϕ
is their overlap. However, in our algebraic approach only the norm in Eq. (4.68) is used and matrix elements of the angular
momentumoperatorsarereducedtothenormbymeansoftheeigenvalueequationfor Jz,Eq.(4.68),andEqs.(4.83)and(4.84).
8Ladder operators can be developed for other mathematical functions. Compare the next subsection, on other Lie groups, and
Section 13.1, for Hermite polynomials.
4.3 Orbital Angular Momentum 263
Therefore, J+|λM/angbracketrightisstillaneigenfunctionof Jzbutwitheigenvalue M+1.J+hasraised
theeigenvalueby1andsoiscalleda raisingoperator .Similarly, J−lowerstheeigenvalue
by1andis calleda loweringoperator .
Takingexpectationvaluesandusing J†
x=Jx,J†
y=Jy, weget
/angbracketleftλM|J2−J2
z|λM/angbracketright=/angbracketleftλM|J2
x+J2
y|λM/angbracketright=vextendsinglevextendsingleJx|λM/angbracketrightvextendsinglevextendsingle2+vextendsinglevextendsingleJy|λM/angbracketrightvextendsinglevextendsingle2
and see that λ−M2≥0, soMis bounded. Let Jbe thelargestM. ThenJ+|λJ/angbracketright=0,
whichimplies J−J+|λJ/angbracketright=0.Hence,combiningEqs. (4.70)and(4.71)toget
J2=J−J++Jz(Jz+1), (4.75)
wefindfromEq. (4.75) that
0=J−J+|λJ/angbracketright=parenleftbig
J2−J2
z−Jzparenrightbig
|λJ/angbracketright=parenleftbig
λ−J2−Jparenrightbig
|λJ/angbracketright.
Therefore
λ=J(J+1)≥0, (4.76)
with nonnegative J. We now relabel the states |λM/angbracketright≡|JM/angbracketright. Similarly, let J′be the
smallest M.ThenJ−|JJ′/angbracketright=0.From
J2=J+J−+Jz(Jz−1), (4.77)
wesee that
0=J+J−|JJ′/angbracketright=parenleftbig
J2+Jz−J2
zparenrightbig
|JJ′/angbracketright=parenleftbig
λ+J′−J′2parenrightbig
|JJ′/angbracketright.(4.78)
Hence
λ=J(J+1)=J′(J′−1)=(−J)(−J−1).
SoJ′=−J,andMruns inintegersteps from−Jto+J,
−J≤M≤J. (4.79)
Startingfrom|JJ/angbracketrightandapplying J−repeatedly,wereachallotherstates |JM/angbracketright.Hencethe
|JM/angbracketrightform anirreduciblerepresentationof SO(3)orSU(2);Mvariesand Jis fixed.
ThenusingEqs. (4.67), (4.75), and(4.77) weobtain
J−J+|JM/angbracketright=bracketleftbig
J(J+1)−M(M+1)bracketrightbig
|JM/angbracketright=(J−M)(J+M+1)|JM/angbracketright,
J+J−|JM/angbracketright=bracketleftbig
J(J+1)−M(M−1)bracketrightbig
|JM/angbracketright=(J+M)(J−M+1)|JM/angbracketright.(4.80)
BecauseJ+andJ−areHermitianconjugates,9
J†
+=J−,J†
−=J+, (4.81)
the eigenvalues in Eq. (4.80) must be positive or zero.10Examples of Eq. (4.81) are pro-
videdbythematricesofExercise3.2.13(spin 1 /2),3.2.15(spin1),and3.2.18(spin 3 /2).
9The Hermitian conjugation or adjoint operation is defined for matrices in Section 3.5, and for operators in general in Sec-
tion 10.1.
10For an excellent discussion of adjoint operators and Hilbert space see A. Messiah, Quantum Mechanics . New York: Wiley
1961, Chapter 7.
264 Chapter 4 Group Theory
For the orbital angular momentum ladder operators, L+, andL−, explicit forms are given
inExercises2.5.14and12.6.7. Youcannowshow(see alsoExercise12.7.2)that
/angbracketleftJM|J−parenleftbig
J+|JM/angbracketrightparenrightbig
=parenleftbig
J+|JM/angbracketrightparenrightbig†J+|JM/angbracketright. (4.82)
SinceJ+raises the eigenvalue MtoM+1, we relabel the resultant eigenfunction
|JM+1/angbracketright.The normalizationis givenbyEq. (4.80)as
J+|JM/angbracketright=radicalbig
(J−M)(J+M+1)|JM+1/angbracketright=radicalbig
J(J+1)−M(M+1)|JM+1/angbracketright,
(4.83)
taking the positive square root and not introducing any phase factor. By the same argu-
ments,
J−|JM/angbracketright=radicalbig
(J+M)(J−M+1)|JM−1/angbracketright=radicalbig
(J(J+1)−M(M−1)|JM−1/angbracketright.
(4.84)
Applying J+toEq.(4.84),weobtainthesecondlineofEq.(4.80)andverifythatEq.(4.84)
isconsistentwithEq. (4.83).
Finally,since Mrangesfrom−Jto+Jinunitsteps, 2 Jmustbeaninteger; Jiseither
an integer or half of an odd integer. As seen later, if Jis an orbital angular momentum L,
the set|LM/angbracketrightfor allMis a basis defining a representation of SO(3) andLwill then be
integral.Insphericalpolarcoordinates θ,ϕ,thefunctions |LM/angbracketrightbecomethesphericalhar-
monicsYM
L(θ,ϕ)of Section 12.6. The sets of |JM/angbracketrightstates with half-integral Jdefine rep-
resentationsof SU(2)thatarenotrepresentationsof SO(3);weget J=1/2,3/2,5/2,....
Our angular momentum is quantized, essentially as a result of the commutation relations.
All these representations are irreducible, as an application of the raising and lowering op-
eratorssuggests.
Summary of Lie Groups and Lie Algebras
The general commutation relations, Eq. (4.14) in Section 4.2, for a classical Lie group
[SO(n)andSU(n)in particular] can be simplified to look more like Eq. (4.71) for SO(3)
andSU(2) in this section. Here we merely review and, as a rule, do not provide proofs for
varioustheoremsthatweexplain.
First wechooselinearly independentand mutuallycommutinggenerators Hiwhichare
generalizationsof JzforSO(3)andSU(2).Letlbethemaximumnumberofsuch Hiwith
[Hi,Hk]=0. (4.85)
Thenliscalledthe rankoftheLiegroup GoritsLiealgebra G.Therankanddimension,
or order, of some Lie groups are given in Table 4.2. All other generators Eαcan be shown
toberaisingandloweringoperatorswithrespecttoallthe Hi,s o
[Hi,Eα]=αiEα,i=1,2,...,l. (4.86)
Thesetofso-called rootvectors (α1,α2,...,αl)form therootdiagram ofG.
When the Hicommute, they can be simultaneously diagonalized (for symmetric (or
Hermitian) matrices see Chapter 3; for operators see Chapter 10). The Hiprovide us with
asetofeigenvalues m1,m2,...,m l[projectionoradditivequantumnumbersgeneralizing
4.3 Orbital Angular Momentum 265
Table 4.2 Rank and Order of Unitary and Rotational
Groups
Liealgebra Al Bl Dl
Liegroup SU(l+1)SO(2l+1) SO(2l)
Rank ll l
Order l(l+2)l (2l+1)l (2l−1)
MofJzinSO(3) andSU(2)]. The set of so-called weight vectors (m1,m2,...,m l)for
anirreduciblerepresentation(multiplet)form a weightdiagram .
There are linvariant operators Ci, calledCasimir operators, that commute with all
generatorsandaregeneralizationsof J2,
[Ci,Hj]=0,[Ci,Eα]=0,i=1,2,...,l. (4.87)
Thefirstone, C1,isaquadraticfunctionofthegenerators;theothersaremorecomplicated.
Since the Cjcommute with all Hj, they can be simultaneously diagonalized with the Hj.
Their eigenvalues c1,c2,...,clcharacterize irreducible representations and stay constant
whiletheweightvectorvariesoveranyparticularirreduciblerepresentation.Thusthegen-
eraleigenfunctionmaybewrittenas
vextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
, (4.88)
generalizingthemultiplet |JM/angbracketrightofSO(3)andSU(2). Theireigenvalueequationsare
Hivextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=mivextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
(4.89a)
Civextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=civextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
.(4.89b)
We can now show that Eα|(c1,c2,...,cl)m1,m2,...,m l/angbracketrighthas the weight vector
(m1+α1,m2+α2,...,m l+αl)using the commutation relations, Eq. (4.86), in con-
junctionwithEqs. (4.89a)and(4.89b):
HiEαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=parenleftbig
EαHi+[Hi,Eα]parenrightbigvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=(mi+αi)Eαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
. (4.90)
Therefore
Eαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
∼vextendsinglevextendsingle(c1,...,cl)m1+α1,...,m l+αlangbracketrightbig
,
the generalization of Eqs. (4.83) and (4.84) from SO(3). These changes of eigenvalues by
theoperator Eαarecalledits selectionrules inquantummechanics.Theyaredisplayedin
therootdiagramofaLiealgebra.
Examples of root diagrams are given in Fig. 4.6 for SU(2) andSU(3). If we attach the
roots denoted by arrows in Fig. 4.6b to a weight in Figs. 4.3 or 4.5a, b, we can reach any
otherstate(representedbyadotintheweightdiagram).
HereSchur’s lemma applies: An operator Hthat commutes with all group operators,
andthereforewithallgenerators Hiofa(classical)Liegroup Ginparticular,hasaseigen-
vectors all states of a multiplet and is degenerate with the multiplet. As a consequence,
suchanoperatorcommuteswithallCasimirinvariants, [H,Ci]=0.
266 Chapter 4 Group Theory
FIGURE 4.6Rootdiagramfor(a) SU(2) and
(b)SU(3).
ThelastresultisclearbecausetheCasimirinvariantsareconstructedfromthegenerators
andraisingandloweringoperatorsofthegroup.Toprovetherest,let ψbeaneigenvector,
Hψ=Eψ. Then, for any rotation RofG,w eh a v e HRψ=ERψ, which says that Rψ
is an eigenstate with the same eigenvalue Ealong with ψ. Since[H,Ci]=0, all Casimir
invariants can be diagonalized simultaneously with Hand an eigenstate of His an eigen-
state of all the Ci. Since[Hi,Ci]=0, the rotated eigenstates Rψare eigenstates of Ci,
alongwith ψbelongingtothesamemultipletcharacterizedbytheeigenvalues ciofCi.
Finally,suchanoperator Hcannotinducetransitionsbetweendifferentmultipletsofthe
groupbecause
angbracketleftbig
(c′
1,c′
2,...,c′
l)m′
1,m′
2,...,m′
lvextendsinglevextendsingleHvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=0.
Using[H,Cj]=0 (for any j)weha v e
0=angbracketleftbig
(c′
1,c′
2,...,c′
l)m′
1,m′
2,...,m′
lvextendsinglevextendsingle[H,Cj]vextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
=(cj−c′
j)angbracketleftbig
(c′
1,c′
2,...,c′
l)m′
1,m′
2,...,m′
lvextendsinglevextendsingleHvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig
.
Ifc′
j/negationslash=cjforsome j, thenthepreviousequationfollows.
Exercises
4.3.1 Showthat(a)[J+,J2]=0,(b)[J−,J2]=0.
4.3.2 Derivetherootdiagramof SU(3) inFig.4.6bfrom thegenerators λiinEq. (4.61).
Hint.Workoutfirstthe SU(2) caseinFig. 4.6afrom thePaulimatrices.
4.4 A NGULAR MOMENTUM COUPLING
In many-body systems of classical mechanics, the total angular momentum is the sum
L=summationtext
iLiof theindividualorbitalangularmomenta.Anyisolatedparticlehas conserved
angular momentum. In quantum mechanics, conserved angular momentum arises when
particles move in a central potential, such as the Coulomb potential in atomic physics,
a shell model potential in nuclear physics, or a confinement potential of a quark model in
4.4 Angular Momentum Coupling 267
particle physics. In the relativistic Dirac equation, orbital angular momentum is no longer
conserved,but J=L+Sisconserved,thetotalangularmomentumofaparticleconsisting
ofits orbitalandintrinsicangularmomentum,calledspin S=σ/2,inunitsof¯h.
It is readily shown that the sum of angular momentum operators obeys the same com-
mutation relations in Eq. (4.37) or (4.41) as the individual angular momentum operators,
providedthosefromdifferent particlescommute.
Clebsch–Gordan Coefficients: SU(2)–SO(3)
Clearly,combiningtwocommutingangularmomenta Jitoform theirsum
J=J1+J2,[J1i,J2i]=0, (4.91)
occursofteninapplications,and Jsatisfiestheangularmomentumcommutationrelations
[Jj,Jk]=[J1j+J2j,J1k+J2k]=[J1j,J1k]+[J2j,J2k]=iεjkl(J1l+J2l)=iεjklJl.
For a single particle with spin 1 /2, for example, an electron or a quark, the total angular
momentum is a sum of orbital angular momentum and spin. For two spinless particles
their total orbital angular momentum L=L1+L2.F o rJ2andJzof Eq. (4.91) to be both
diagonal,[J2,Jz]=0 has to hold. To show this we use the obvious commutationrelations
[Jiz,J2
j]=0,and
J2=J2
1+J2
2+2J1·J2=J2
1+J2
2+J1+J2−+J1−J2++2J1zJ2z (4.91′)
inconjunctionwithEq. (4.71), for both Ji, toobtain
bracketleftbig
J2,Jzbracketrightbig
=[J1−J2++J1+J2−,J1z+J2z]
=[J1−,J1z]J2++J1−[J2+,J2z]+[J1+,J1z]J2−+J1+[J2−,J2z]
=J1−J2+−J1−J2+−J1+J2−+J1+J2−=0.
Similarly[J2,J2
i]=0 is proved. Hence the eigenvalues of J2
i,J2,Jzcan be used to label
thetotalangularmomentumstates |J1J2JM/angbracketright.
Theproductstates |J1m1/angbracketright|J2m2/angbracketrightobviouslysatisfy theeigenvalueequations
Jz|J1m1/angbracketright|J2m2/angbracketright=(J1z+J2z)|J1m1/angbracketright|J2m2/angbracketright=(m1+m2)|J1m1/angbracketright|J2m2/angbracketright
=M|J1m1/angbracketright|J2m2/angbracketright, (4.92)
J2
i|J1m1/angbracketright|J2m2/angbracketright=Ji(Ji+1)|J1m1/angbracketright|J2m2/angbracketright,
but will not have diagonal J2except for the maximally stretched states with M=
±(J1+J2)andJ=J1+J2(see Fig. 4.7a). To see this we use Eq. (4.91′) again in con-
junctionwithEqs. (4.83)and(4.84)in
J2|J1m1/angbracketrightJ2m2/angbracketright=braceleftbig
J1(J1+1)+J2(J2+1)+2m1m2bracerightbig
|J1m1/angbracketright|J2m2/angbracketright
+braceleftbig
J1(J1+1)−m1(m1+1)bracerightbig1/2braceleftbig
J2(J2+1)−m2(m2−1)bracerightbig1/2
×|J1m1+1/angbracketright|J2m2−1/angbracketright+braceleftbig
J1(J1+1)−m1(m1−1)bracerightbig1/2
×braceleftbig
J2(J2+1)−m2(m2+1)bracerightbig1/2|J1m1−1/angbracketright|J2m2+1/angbracketright. (4.93)
268 Chapter 4 Group Theory
FIGURE 4.7Couplingof twoangularmomenta:
(a) parallelstretched,(b) antiparallel,(c)general
case.
ThelasttwotermsinEq.(4.93)vanishonlywhen m1=J1andm2=J2orm1=−J1and
m2=−J2. In both cases J=J1+J2follows from the first line of Eq. (4.93). In general,
therefore,wehavetoformappropriatelinearcombinationsof productstates
|J1J2JM/angbracketright=summationdisplay
m1,m2Cparenleftbig
J1J2J|m1m2Mparenrightbig
|J1m1/angbracketright|J2m2/angbracketright, (4.94)
so thatJ2has eigenvalue J(J+1). The quantities C(J1J2J|m1m2M)in Eq. (4.94) are
calledClebsch–Gordancoefficients .FromEq.(4.92)weseethattheyvanishunless M=
m1+m2, reducing the double sum to a single sum. Applying J±to|JM/angbracketrightshows that the
eigenvalues MofJzsatisfy theusualinequalities −J≤M≤J.
Clearly,themaximal Jmax=J1+J2(seeFig.4.7a).InthiscaseEq.(4.93)reducestoa
pureproductstate
|J1J2J=J1+J2M=J1+J2/angbracketright=|J1J1/angbracketright|J2J2/angbracketright, (4.95a)
sotheClebsch–Gordancoefficient
C(J1J2J=J1+J2|J1J2J1+J2)=1. (4.95b)
Theminimal J=J1−J2(ifJ1>J2,seeFig.4.7b)and J=J2−J1forJ2>J1followif
wekeepinmindthattherearejustas manyproductstatesas |JM/angbracketrightstates;thatis,
Jmaxsummationdisplay
J=Jmin(2J+1)=(Jmax−Jmin+1)(Jmax+Jmin+1)
=(2J1+1)(2J2+1). (4.96)
This condition holds because the |J1J2JM/angbracketrightstates merely rearrange all product states into
irreduciblerepresentationsoftotalangularmomentum.Itisequivalenttothe trianglerule :
/Delta1(J1J2J)=1,if|J1−J2|≤J≤J1+J2;
/Delta1(J1J2J)=0,else.(4.97)
4.4 Angular Momentum Coupling 269
This indicates that one complete multiplet of each Jvalue from JmintoJmaxaccounts
for all the states and that all the |JM/angbracketrightstates are necessarily orthogonal. In other words,
Eq. (4.94) defines a unitary transformation from the orthogonal basis set of products of
single-particle states |J1m1;J2m2/angbracketright=|J1m1/angbracketright|J2m2/angbracketrightto the two-particle states |J1J2JM/angbracketright.
TheClebsch–Gordancoefficientsarejusttheoverlapmatrixelements
C(J1J2J|m1m2M)≡/angbracketleftJ1J2JM|J1m1;J2m2/angbracketright. (4.98)
The explicit construction in what follows shows that they are all real. The states in
Eq.(4.94) areorthonormalized,providedthattheconstraints
summationdisplay
m1,m2,m1+m2=MC(J1J2J|m1m2M)C(J 1J2J′|m1m2M′/angbracketright
=/angbracketleftJ1J2JM|J1J2J′M′/angbracketright=δJJ′δMM′(4.99a)
summationdisplay
J,MC(J1J2J|m1m2M)C(J 1J2J|m′
1m′
2M)
=/angbracketleftJ1m1|J1m′
1/angbracketright/angbracketleftJ2m2|J2m′
2/angbracketright=δm1m′
1δm2m′
2(4.99b)
hold.
Nowwearereadytoconstructmoredirectlythetotalangularmomentumstatesstarting
from|Jmax=J1+J2M=J1+J2/angbracketrightin Eq. (4.95a) and using the lowering operator J−=
J1−+J2−repeatedly.In thefirst stepweuseEq.(4.84) for
Ji−|JiJi/angbracketright=braceleftbig
Ji(Ji+1)−Ji(Ji−1)bracerightbig1/2|JiJi−1/angbracketright=(2Ji)1/2|JiJi−1/angbracketright,
which we substitute into (J1−+J2−/angbracketright|J1J1)|J2J2/angbracketright. Normalizing the resulting state with
M=J1+J2−1 properlyto1,weobtain
|J1J2J1+J2J1+J2−1/angbracketright=braceleftbig
J1/(J1+J2)bracerightbig1/2|J1J1−1/angbracketright|J2J2/angbracketright
+braceleftbig
J2/(J1+J2)bracerightbig1/2|J1J1/angbracketright|J2J2−1/angbracketright.(4.100)
Equation(4.100)yieldstheClebsch–Gordancoefficients
C(J1J2J1+J2|J1−1J2J1+J2−1)=braceleftbig
J1/(J1+J2)bracerightbig1/2,
C(J1J2J1+J2|J1J2−1J1+J2−1)=braceleftbig
J2/(J1+J2)bracerightbig1/2.(4.101)
Thenweapply J−againandnormalizethestatesobtaineduntilwereach |J1J2J1+J2M/angbracketright
withM=−(J1+J2). The Clebsch–Gordan coefficients C(J1J2J1+J2|m1m2M)may
thusbecalculatedstepbystep,andtheyareallreal.
Thenextstepistorealizethattheonlyotherstatewith M=J1+J2−1isthetopofthe
nextlowertowerof |J1+J2−1M/angbracketrightstates.Since|J1+J2−1J1+J2−1/angbracketrightisorthogonalto
|J1+J2J1+J2−1/angbracketrightinEq.(4.100),itmustbetheotherlinearcombinationwitharelative
minussign,
|J1+J2−1J1+J2−1/angbracketright=−braceleftbig
J2/(J1+J2)bracerightbig1/2|J1J1−1/angbracketright|J2J2/angbracketright
+braceleftbig
J1/(J1+J2)bracerightbig1/2|J1J1/angbracketright|J2J2−1/angbracketright,(4.102)
uptoanoverallsign.
270 Chapter 4 Group Theory
HencewehavedeterminedtheClebsch–Gordancoefficients(for J2≥J1)
C(J1J2J1+J2−1|J1−1J2J1+J2−1)=−braceleftbig
J2/(J1+J2)bracerightbig1/2,
C(J1J2J1+J2−1|J1J2−1J1+J2−1)=braceleftbig
J1/(J1+J2)bracerightbig1/2.(4.103)
Againwecontinueusing J−untilwereach M=−(J1+J2−1),andwekeepnormalizing
theresultingstates |J1+J2−1M/angbracketrightof theJ=J1+J2−1t o w e r .
In order to get to the top of the next tower, |J1+J2−2M/angbracketrightwithM=J1+J2−2, we
remember that we have already constructed two states with that M.B o t h|J1+J2J1+
J2−2/angbracketrightand|J1+J2−1J1+J2−2/angbracketrightareknownlinearcombinationsofthethreeproduct
states|J1J1/angbracketright|J2J2−2/angbracketright,|J1J1−1/angbracketright×|J2J2−1/angbracketright, and|J1J1−2/angbracketright|J2J2/angbracketright. The third linear
combination is easy to find from orthogonality to these two states, up to an overall phase,
which is chosen by the Condon–Shortley phase conventions11so that the coefficient
C(J1J2J1+J2−2|J1J2−2J1+J2−2)ofthelastproductstateispositivefor |J1J2J1+
J2−2J1+J2−2/angbracketright.Itisstraightforward,thoughabittedious,todeterminetherestofthe
Clebsch–Gordancoefficients.
Numerous recursion relations can be derived from matrix elements of various angular
momentumoperators,for whichwerefer totheliterature.12
ThesymmetrypropertiesofClebsch–Gordancoefficientsarebestdisplayedinthemore
symmetricWigner’s3 j-symbols,whicharetabulated:12
parenleftBigg
J1J2J3
m1m2m3parenrightBigg
=(−1)J1−J2−m3
(2J3+1)1/2C(J1J2J3|m1m2,−m3), (4.104a)
obeyingthesymmetryrelations
parenleftBigg
J1J2J3
m1m2m3parenrightBigg
=(−1)J1+J2+J3parenleftBigg
JkJlJn
mkmlmnparenrightBigg
(4.104b)
for(k,l,n)an odd permutation of (1,2,3). One of the most important places where
Clebsch–Gordan coefficients occur is in matrix elements of tensor operators, which are
governed by the Wigner–Eckart theorem discussed in the next section, on spherical ten-
sors. Another is coupling of operators or state vectors to total angular momentum, such
as spin-orbit coupling. Recoupling of operators and states in matrix elements leads to 6 j-
and 9j-symbols.12Clebsch–Gordan coefficients can and have been calculated for other
Liegroups,suchas SU(3).
11E.U.Condon and G.H.Shortley, Theory of AtomicSpectra . Cambridge, UK:Cambridge University Press (1935).
12There is a rich literature on this subject, e.g., A. R. Edmonds, Angular Momentum in Quantum Mechanics . Princeton, NJ:
PrincetonUniversityPress(1957);M.E.Rose, ElementaryTheoryofAngularMomentum .NewYork:Wiley(1957);A.de-Shalit
andI.Talmi, NuclearShellModel .NewYork:AcademicPress(1963);Dover(2005).Clebsch–Gordancoefficientsaretabulated
in M. Rotenberg, R. Bivins, N. Metropolis, and J. K. Wooten, Jr., The 3j- and 6j-Symbols . Cambridge, MA: Massachusetts
Institute of Technology Press (1959).
4.4 Angular Momentum Coupling 271
Spherical Tensors
In Chapter 2 the properties of Cartesian tensors are defined using the group of nonsin-
gular general linear transformations, which contains the three-dimensional rotations as a
subgroup. A tensor of a given rank that is irreducible with respect to the full group may
well become reducible for the rotation group SO(3). To explain this point, consider the
second-ranktensorwithcomponents Tjk=xjykforj,k=1,2,3.Itcontainsthesymmet-
rictensor Sjk=(xjyk+xkyj)/2andtheantisymmetrictensor Ajk=(xjyk−xkyj)/2,so
Tjk=Sjk+Ajk. This reduces TjkinSO(3). However, under rotations the scalar product
x·yis invariant and is therefore irreducible in SO(3). Thus, Sjkcan be reduced by sub-
traction of the multiple of x·ythat makes it traceless. This leads to the SO(3)-irreducible
tensor
S′
jk=1
2(xjyk+xkyj)−1
3x·yδjk.
Tensors of higher rank may be treated similarly. When we form tensors from products of
the components of the coordinate vector rthen, in polar coordinates that are tailored to
SO(3)symmetry,weendupwiththesphericalharmonicsofChapter12.
The form of the ladder operators for SO(3) in Section 4.3 leads us to introduce the
spherical components (note the different normalization and signs, though, prescribed by
theYlm) ofavector A:
A+1=−1√
2(Ax+iAy), A−1=1√
2(Ax−iAy), A 0=Az.(4.105)
Thenwehavefor thecoordinatevector rinpolarcoordinates,
r+1=−1√
2rsinθeiϕ=rradicalBig
4π
3Y11,r−1=1√
2rsinθe−iϕ=rradicalBig
4π
3Y1,−1,
r0=rradicalBig
4π
3Y10,(4.106)
whereYlm(θ,ϕ)are the spherical harmonics of Chapter 12. Again, the spherical jmcom-
ponentsoftensors Tjmof higherrank jmaybeintroducedsimilarly.
An irreducible spherical tensor operator Tjmof rankjhas 2j+1 components, just
as for spherical harmonics, and mruns from−jto+j. Under a rotation R(α), whereα
standsfortheEulerangles,the Ylmtransform as
Ylm(ˆr′)=summationdisplay
m′Ylm′(ˆr)Dl
m′m(R), (4.107a)
whereˆr′=(θ′,ϕ′)areobtainedfrom ˆr=(θ,ϕ)bytherotation Randaretheanglesofthe
samepointintherotatedframe,and
DJ
m′m(α,β,γ)=/angbracketleftJm|exp(iαJz)exp(iβJy)exp(iγJz)|Jm′/angbracketright
aretherotationmatrices.So,for theoperator Tjm, wedefine
RTjmR−1=summationdisplay
m′Tjm′Dj
m′m(α). (4.107b)
272 Chapter 4 Group Theory
Foraninfinitesimalrotation(seeEq.(4.20)inSection4.2ongenerators)theleftsideof
Eq. (4.107b) simplifies to a commutator and the right side to the matrix elements of J,t h e
infinitesimalgeneratorof therotation R:
[Jn,Tjm]=summationdisplay
m′Tjm′/angbracketleftjm′|Jn|jm/angbracketright. (4.108)
IfwesubstituteEqs.(4.83)and(4.84)forthematrixelementsof Jmweobtainthealterna-
tivetransformationlawsofatensor operator,
[J0,Tjm]=mTjm,[J±,Tjm]=Tjm±1braceleftbig
(j−m)(j±m+1)bracerightbig1/2.(4.109)
We can use the Clebsch–Gordan coefficients of the previous subsection to couple two
tensors of given rank to another rank. An example is the cross or vector product of two
vectorsaandbfromChapter1.Letuswritebothvectorsinsphericalcomponents, amand
bm. Thenweverifythatthetensor Cmofrank1 definedas
Cm≡summationdisplay
m1m2C(111|m1m2m)am1bm2=i√
2(a×b)m. (4.110)
SinceCmisasphericaltensorofrank1thatislinearinthecomponentsof aandb,itmust
beproportionaltothecrossproduct, Cm=N(a×b)m.Theconstant Ncanbedetermined
fromaspecialcase, a=ˆx,b=ˆy,essentiallywriting ˆx׈y=ˆzinsphericalcomponentsas
follows.Using
(ˆz)0=1;(ˆx)1=−1/√
2,(ˆx)−1=1/√
2;
(ˆy)1=−i/√
2,(ˆy)−1=−i/√
2,
Eq.(4.110) for m=0 becomes
C(111|1,−1,0)bracketleftbig
(ˆx)1(ˆy)−1−(ˆx)−1(ˆy)1bracketrightbig
=Nparenleftbig
(ˆz)0parenrightbig
=N
=1√
2bracketleftbigg
−1√
2parenleftbigg
−i√
2parenrightbigg
−1√
2parenleftbigg
−i√
2parenrightbiggbracketrightbigg
=i√
2,
w h e r ew eh a v eu s e d C(111|101)=1√
2from Eq. (4.103) for J1=1=J2, which implies
C(111|1,−1,0)=1√
2usingEqs. (4.104a,b):
parenleftBigg
11 1
10−1parenrightBigg
=−1√
3C(111|101)=−1
6=−parenleftBigg
111
1−10parenrightBigg
=−1√
3C(111|1,−1,0).
A bit simpler is the usual scalar product of two vectors in Chapter 1, in which aandb
arecoupledtozeroangularmomentum:
a·b≡−(ab)0√
3≡−√
3summationdisplay
mC(110|m,−m,0)amb−m. (4.111)
Again, the rank zero of our tensor product implies a·b=n(ab)0. The constant ncan
be determined from a special case, essentially writing ˆz2=1 in spherical components:
ˆz2=1=nC(110|000)=−n√
3.
4.4 Angular Momentum Coupling 273
Anotheroften-usedapplicationoftensorsisthe recoupling thatinvolves 6j-symbols for
three operators and 9 jfor four operators.12An example is the following scalar product,
forwhichit canbeshown12that
σ1·rσ2·r=1
3r2σ1·σ2+(σ1σ2)2·(rr)2, (4.112)
but which can also be rearranged by elementary means. Here the tensor operators are de-
finedas
(σ1σ2)2m=summationdisplay
m1m2C(112|m1m2m)σ1m1σ2m2, (4.113)
(rr)2m=summationdisplay
mC(112|m1m2m)rm1rm2=radicalbigg
8π
15r2Y2m(ˆr), (4.114)
andthescalarproductof tensorsofrank2 as
(σ1σ2)2·(rr)2=summationdisplay
m(−1)m(σ1σ2)2m(rr)2,−m=√
5parenleftbig
(σ1σ2)2(rr)2parenrightbig
0.(4.115)
One of the most important applications of spherical tensor operators is the Wigner–
Eckarttheorem .Itsaysthatamatrixelementofasphericaltensoroperator Tkmofrankk
betweenstatesofangularmomentum jandj′factorizesintoaClebsch–Gordancoefficient
and a so-called reduced matrix element , denoted by double bars, that no longer has any
dependenceontheprojectionquantumnumbers m,m′,n:
/angbracketleftj′m′|Tkn|jm/angbracketright=C(kjj′|nmm′)(−1)k−j+j′/angbracketleftj′/bardblTk/bardblj/angbracketright/radicalbig
(2j′+1). (4.116)
In other words, such a matrix element factors into a dynamic part, the reduced matrix
element, and a geometric part, the Clebsch–Gordan coefficient that contains the rotational
properties (expressed by the projection quantum numbers) from the SO(3) invariance. To
seethiswecouple Tknwiththeinitialstatetototalangularmomentum j′:
|j′m′/angbracketright0≡summationdisplay
nmC(kjj′|nmm′)Tkn|jm/angbracketright. (4.117)
Under rotations the state |j′m′/angbracketright0transforms just like |j′m′/angbracketright. Thus, the overlap matrix ele-
ment/angbracketleftj′m′|j′m′/angbracketright0isarotationalscalarthathasno m′dependence,sowecanaverageover
theprojections,
/angbracketleftJM|j′m′/angbracketright0=δJj′δMm′
2j′+1summationdisplay
µ/angbracketleftj′µ|j′µ/angbracketright0. (4.118)
Next we substitute our definition, Eq. (4.117), into Eq. (4.118) and invert the relation
Eq.(4.117) usingorthogonality,Eq. (4.99b),tofindthat
/angbracketleftJM|Tkn|jm/angbracketright=summationdisplay
j′m′C(kjj′|nmm′)δJj′δMm′
2J+1summationdisplay
µ/angbracketleftJµ|Jµ/angbracketright0,(4.119)
whichprovestheWigner–Eckarttheorem,Eq.(4.116).13
13Theextra factor (−1)k−j+j′/radicalbig
(2j′+1)in Eq.(4.116) is just aconvention that varies in the literature.
274 Chapter 4 Group Theory
As an application, we can write the Pauli matrix elements in terms of Clebsch–Gordan
coefficients.WeapplytheWigner–Eckarttheoremto
angbracketleftbig1
2γvextendsinglevextendsingleσαvextendsinglevextendsingle1
2βangbracketrightbig
=(σα)γβ=−1√
2Cparenleftbig
11
21
2vextendsinglevextendsingleαβγparenrightbigangbracketleftbig1
2vextenddoublevextenddoubleσvextenddoublevextenddouble1
2angbracketrightbig
. (4.120)
Since/angbracketleft1
21
2|σ0|1
21
2/angbracketright=1 withσ0=σ3andC(11
21
2|01
21
2)=−1/√
3,wefind
angbracketleftbig1
2vextenddoublevextenddoubleσvextenddoublevextenddouble1
2angbracketrightbig
=√
6, (4.121)
which,substitutedintoEq.(4.120), yields
(σα)γβ=−√
3Cparenleftbig
11
21
2vextendsinglevextendsingleαβγparenrightbig
. (4.122)
Notethatthe α=±1,0 denotethesphericalcomponentsofthePaulimatrices.
Young Tableaux for SU(n)
Young tableaux (YT) provide a powerful and elegant method for decomposing products
ofSU(n)group representations into sums of irreducible representations. The YT provide
the dimensions and symmetry types of the irreducible representations in this so-called
Clebsch–Gordan series , though not the Clebsch–Gordan coefficients by which the prod-
uct states are coupled to the quantum numbers of each irreducible representation of the
series(see Eq.(4.94)).
Products of representations correspond to multiparticle states. In this context, permuta-
tionsofparticlesareimportantwhenwedealwithseveralidenticalparticles.Permutations
ofnidentical objects form the symmetric group Sn. A close connection between irre-
ducible representations of Sn, which are the YT, and those of SU(n)is provided by this
theorem:E v e r yN-particle state of Snthat is made up of single-particle states of the fun-
damental n-dimensional SU(n)multiplet belongs to an irreducible SU(n)representation.
AproofisinChapter22ofWybourne.14
ForSU(2) the fundamental representation is a box that stands for the spin +1
2(up) and
−1
2(down)statesandhasdimension2.For SU(3)theboxcomprisesthethreequarkstates
inthetriangleofFig.4.5a;ithasdimension3.
An array of boxes shown in Fig. 4.8 with λ1boxes in the first row, λ2boxes in the
second row, ...,andλn−1boxes in the last row is called a Young tableau (YT), denoted
by[λ1,...,λn−1],andrepresentsanirreduciblerepresentationof SU(n)ifandonlyif
λ1≥λ2≥···≥λn−1. (4.123)
Boxes in the same row are symmetric representations; those in the same column are anti-
symmetric. A YT consisting of one row is totally symmetric. A YT consisting of a single
columnistotallyantisymmetric.
There are at most n−1r o w sf o r SU(n)YT because a column of nboxes is the totally
antisymmetric ( Slater determinant of single-particle states) singlet representation that
maybestruckfromtheYT.
An array of Nboxes is an N-particle state whose boxes may be labeled by positive
integerssothatthe(particlelabelsor)numbersinonerowoftheYTdonotdecreasefrom
14B. G.Wybourne, Classical Groups for Physicists .NewYork: Wiley (1974).
4.4 Angular Momentum Coupling 275
FIGURE 4.8Youngtableau(YT) for SU(n).
left to right and those in any one column increase from top to bottom. In contrast to the
possiblerepetitionsofrownumbers,thenumbersinanycolumnmustbedifferentbecause
oftheantisymmetryof thesestates.
The product of a YT with a single box, [1], is the sum of YT formed when the box is
putattheendofeachrowoftheYT,providedtheresultingYTislegitimate,thatis,obeys
Eq.(4.123).For SU(2)theproductoftwoboxes,spin 1 /2 representationsofdimension2,
generates
[1]⊗[1]=[2]⊕[1,1], (4.124)
the symmetric spin 1 representation of dimension 3 and the antisymmetric singlet of di-
mension1mentionedearlier.
Thecolumnof n−1boxesistheconjugaterepresentationofthefundamentalrepresen-
tation; its product with a single box contains the column of nboxes, which is the singlet.
ForSU(3) the conjugate representation of the single box, [1]or fundamental quark repre-
sentation, is the inverted triangle in Fig. 4.5b, [1,1], which represents the three antiquarks
¯u,¯d,¯s, obviouslyof dimension3aswell.
ThedimensionofaYT isgivenbytheratio
dimYT=N
D. (4.125)
The numerator Nis obtained by writing an nin all boxes of the YT along the diagonal,
(n+1)inallboxesimmediatelyabovethediagonal, (n−1)immediatelybelowthediago-
nal,etc.NistheproductofallthenumbersintheYT.AnexampleisshowninFig.4.9afor
the octet representation of SU(3), where N=2·3·4=24. There is a closed formula that
isequivalenttoEq.(4.125).15Thedenominator Distheproductofall hooks.16Ahookis
drawnthrougheachboxoftheYTbystartingahorizontallinefromtherighttotheboxin
questionandthencontinuingitverticallyoutoftheYT.Thenumberofboxesencountered
by the hook-line is the hook-number of the box. Dis the product of all hook-numbers of
15See, for example, M. Hamermesh, Group Theory and Its Application to Physical Problems . Reading, MA: Addison-Wesley
(1962).
16F. Close, Introduction to Quarks and Partons . NewYork: AcademicPress (1979).
276 Chapter 4 Group Theory
(a)
(b)
FIGURE 4.9Illustration
of(a)Nand(b)Din
Eq. (4.125)for theoctet
Youngtableauof SU(3).
the YT. An example is shown in Fig. 4.9b for the octet of SU(3), whose hook-number is
D=1·3·1=3.Hencethedimensionof the SU(3) octetis 24 /3=8,whenceits name.
Now we can calculate the dimensions of the YT in Eq. (4.124). For SU(2) they are
2×2=3+1=4. ForSU(3) they are 3·3=3·4/(1·2)+3·2/(2·1)=6+3=9. For
theproductofthequarktimesantiquarkYTof SU(3)weget
[1,1]⊗[1]=[2,1]⊕[1,1,1], (4.126)
that is, octet and singlet, which are precisely the meson multiplets considered in the sub-
sectionontheeightfoldway,the SU(3)flavorsymmetry,whichsuggestmesonsarebound
states of a quark and an antiquark, q¯qconfigurations. For the product of three quarks we
get
parenleftbig
[1]⊗[1]parenrightbig
⊗[1]=parenleftbig
[2]⊕[1,1]parenrightbig
⊗[1]=[3]⊕2[2,1]⊕[1,1,1],(4.127)
that is, decuplet, octet, and singlet, which are the observed multiplets for the baryons,
whichsuggeststheyareboundstatesof threequarks, q3configurations.
Aswehaveseen,YTdescribethedecompositionofaproductof SU(n)irreduciblerepre-
sentations into irreducible representations of SU(n), which is called the Clebsch–Gordan
series,whiletheClebsch–Gordancoefficientsconsideredearlierallowconstructionofthe
individualstatesinthisseries.
4.4 Angular Momentum Coupling 277
Exercises
4.4.1 Derive recursion relations for Clebsch–Gordan coefficients. Use them to calculate
C(11J|m1m2M)forJ=0,1,2.
Hint.Usetheknownmatrixelementsof J+=J1++J2+,Ji+,andJ2=(J1+J2)2,etc.
4.4.2 Show that (Ylχ)J
M=summationtextC(l1
2J|mlmsM)Ylmlχms, whereχ±1/2are the spin up and
downeigenfunctionsof σ3=σz, transforms likeasphericaltensorof rank J.
4.4.3 Whenthespinofquarksistakenintoaccount,the SU(3)flavorsymmetryisreplacedby
theSU(6) symmetry. Why? Obtain the Young tableau for the antiquark configuration
¯q. Then decompose the product q¯q. WhichSU(3) representations are contained in the
nontrivialSU(6)representationfor mesons?
Hint.DeterminethedimensionsofallYT.
4.4.4 Forl=1,Eq. (4.107a)becomes
Ym
1(θ′,ϕ′)=1summationdisplay
m′=−1D1
m′m(α,β,γ)Ym′
1(θ,ϕ).
Rewrite these spherical harmonics in Cartesian form. Show that the resulting Cartesian
coordinate equations are equivalent to the Euler rotation matrix A(α,β,γ), Eq. (3.94),
rotatingthecoordinates.
4.4.5 Assumingthat Dj(α,β,γ) is unitary,showthat
lsummationdisplay
m=−lYm∗
l(θ1,ϕ1)Ym
l(θ2,ϕ2)
is a scalar quantity (invariant under rotations). This is a spherical tensor analog of a
scalarproductof vectors.
4.4.6 (a) Showthatthe αandγdependenceof Dj(α,β,γ) maybefactoredoutsuchthat
Dj(α,β,γ)=Aj(α)dj(β)Cj(γ).
(b) Showthat Aj(α)andCj(γ)arediagonal.Findtheexplicitforms.
(c) Showthat dj(β)=Dj(0,β,0).
4.4.7 Theangularmomentum–exponentialformof theEuleranglerotationoperatorsis
R=Rz′′(γ)Ry′(β)Rz(α)
=exp(−iγJz′′)exp(−iβJy′)exp(−iαJz).
Showthatintermsoftheoriginalaxes
R=exp(iαJz)exp(−iβJy)exp(−iγJz).
Hint.T h eRoperators transform as matrices. The rotation about the y′-axis (second
Eulerrotation)maybereferredtotheoriginal y-axisby
exp(−iβJy′)=exp(−iαJz)exp(−iβJy)exp(iαJz).
278 Chapter 4 Group Theory
4.4.8 Using the Wigner–Eckart theorem, prove the decomposition theorem for a spherical
vectoroperator /angbracketleftj′m′|T1m|jm/angbracketright=/angbracketleftjm′|J·T1|jm/angbracketright
j(j+1)δjj′.
4.4.9 UsingtheWigner–Eckarttheorem,provethefactorization
/angbracketleftj′m′|JMJ·T1|jm/angbracketright=/angbracketleftjm′|JM|jm/angbracketrightδj′j/angbracketleftjm|J·T1|jm/angbracketright.
4.5 H OMOGENEOUS LORENTZ GROUP
Generalizing the approach to vectors of Section 1.2, in special relativity we demand that
ourphysicallawsbecovariant17under
a. spaceandtimetranslations,
b. rotationsinreal,three-dimensionalspace,and
c. Lorentztransformations.
The demand for covariance under translations is based on the homogeneity of space and
time. Covariance under rotations is an assertion of the isotropy of space. The requirement
of Lorentz covariance follows from special relativity. All three of these transformations
togetherformtheinhomogeneousLorentzgroupor thePoincarégroup.Whenweexclude
translations, the space rotations and the Lorentz transformations together form a group —
thehomogeneousLorentzgroup.
We first generate a subgroup, the Lorentz transformations in which the relative velocity
vis along the x=x1-axis. The generator may be determined by considering space–time
reference frames moving with a relative velocity δv, an infinitesimal.18The relations are
similar to those for rotations in real space, Sections 1.2, 2.6, and 3.3, except that here the
angleofrotationispureimaginary(compareSection4.6).
Lorentz transformations are linear not only in the space coordinates xibut in the time t
as well. They originate from Maxwell’s equations of electrodynamics, which are invariant
under Lorentz transformations, as we shall see later. Lorentz transformations leave the
quadratic form c2t2−x2
1−x2
2−x2
3=x2
0−x2
1−x2
2−x2
3invariant, where x0=ct.W e
see this if we switch on a light source at the origin of the coordinate system. At time t
lighthastraveledthedistance ct=radicalBigsummationtextx2
i,s oc2t2−x2
1−x2
2−x2
3=0.Specialrelativity
requires that in all (inertial) frames that move with velocity v≤cin any direction relative
tothexi-systemandhavethesameoriginattime t=0,c2t′2−x′2
1−x′2
2−x′2
3=0 holds
also.Four-dimensionalspace–timewiththemetric x·x=x2=x2
0−x2
1−x2
2−x2
3iscalled
Minkowskispace,withthescalarproductoftwofour-vectorsdefinedas a·b=a0b0−a·b.
Usingthemetrictensor
(gµν)=parenleftbig
gµνparenrightbig
=
1 000
0−10 0
00−10
00 0 −1
(4.128)
17To be covariant means to have the same form in different coordinate systems so that there is no preferred reference system
(compare Sections 1.2 and 2.6).
18This derivation, withaslightly different metric,appears inanarticleby J.L.Strecker, Am.J.Phys. 35: 12 (1967).
4.5 Homogeneous Lorentz Group 279
we can raise and lower the indices of a four-vector, such as the coordinates xµ=(x0,x),
so thatxµ=gµνxν=(x0,−x)andxµgµνxν=x2
0−x2, Einstein’s summationconvention
being understood. For the gradient, ∂µ=(∂/∂x0,−∇)=∂/∂xµand∂µ=(∂/∂x0,∇),s o
∂2=∂µ∂µ=(∂/∂x0)2−∇2is aLorentzscalar, justlikethemetric x2=x2
0−x2.
Forv≪c,inthenonrelativisticlimit,aLorentztransformationmustbeGalilean.Hence,
to derive the form of a Lorentz transformation along the x1-axis, we start with a Galilean
transformationfor infinitesimalrelativevelocity δv:
x′1=x1−δvt=x1−x0δβ. (4.129)
Here,β=v/c.Bysymmetrywealsowrite
x′0=x0+aδβx1, (4.129′)
withtheparameter achosenso that x2
0−x2
1isinvariant,
x′2
0−x′2
1=x2
0−x2
1. (4.130)
Remember, xµ=(x0,x)is the prototype four-dimensional vector in Minkowski space.
Thus Eq. (4.130) is simply a statement of the invariance of the square of the magnitude of
the“distance”vectorunderLorentztransformationinMinkowskispace.Hereiswherethe
specialrelativityisbroughtintoourtransformation.SquaringandsubtractingEqs.(4.129)
and (4.129′) and discarding terms of order (δβ)2, we find a=−1. Equations (4.129) and
(4.129′)maybecombinedasamatrixequation,
parenleftBigg
x′0
x′1parenrightBigg
=(12−δβσ1)parenleftBigg
x0
x1parenrightBigg
; (4.131)
σ1happens to be the Pauli matrix, σ1, and the parameter δβrepresents an infinitesimal
change.UsingthesametechniquesasinSection4.2,werepeatthetransformation Ntimes
todevelopafinitetransformationwiththevelocityparameter ρ=Nδβ. Then
parenleftBigg
x′0
x′1parenrightBigg
=parenleftbigg
12−ρσ1
NparenrightbiggNparenleftBigg
x0
x1parenrightBigg
. (4.132)
Inthelimitas N→∞,
lim
N→∞parenleftbigg
12−ρσ1
NparenrightbiggN
=exp(−ρσ1). (4.133)
AsinSection4.2,theexponentialisinterpretedbyaMaclaurinexpansion,
exp(−ρσ1)=12−ρσ1+1
2!(ρσ1)2−1
3!(ρσ1)3+···. (4.134)
Notingthat (σ1)2=12,
exp(−ρσ1)=12coshρ−σ1sinhρ. (4.135)
HenceourfiniteLorentztransformationisparenleftBigg
x′0
x′1parenrightBigg
=parenleftBigg
coshρ−sinhρ
−sinhρcoshρparenrightBiggparenleftBigg
x0
x1parenrightBigg
. (4.136)
280 Chapter 4 Group Theory
σ1has generated the representations of this pure Lorentz transformation. The quantities
coshρand sinhρmay be identified by considering the origin of the primed coordinate
system,x′1=0,orx1=vt. SubstitutingintoEq.(4.136), wehave
0=x1coshρ−x0sinhρ. (4.137)
Withx1=vtandx0=ct,
tanhρ=β=v
c.
Note that the rapidity ρ/negationslash=v/c, except in the limit as v→0. The rapidity is the additive
parameterforpureLorentztransformations(“boosts”)alongthesameaxisthatcorresponds
toanglesfor rotationsaboutthesameaxis.Using 1 −tanh2ρ=(cosh2ρ)−1,
coshρ=parenleftbig
1−β2parenrightbig−1/2≡γ,sinhρ=βγ. (4.138)
The group of Lorentz transformations is not compact, because the limit of a sequence of
rapiditiesgoingtoinfinityis nolongeranelementofthegroup.
The preceding special case of the velocity parallel to one space axis is easy, but it illus-
trates the infinitesimal velocity-exponentiation-generator technique. Now, this exact tech-
nique may be applied to derive the Lorentz transformation for the relative velocity v not
parallel to any space axis. The matrices given by Eq. (4.136) for the case of v=ˆxvxform
a subgroup. The matrices in the general case do not. The product of two Lorentz transfor-
mation matrices L(v1)andL(v2)yields a third Lorentz matrix, L(v3), if the two velocities
v1andv2are parallel. The resultant velocity, v3, is related to v1andv2by the Einstein
velocity addition law, Exercise 4.5.3. If v1andv2are not parallel, no such simple relation
exists.Specifically,considerthreereferenceframes S,S′,andS′′,withSandS′relatedby
L(v1)andS′andS′′relatedby L(v2).Ifthevelocityof S′′relativetotheoriginalsystem S
isv3,S′′isnotobtainedfrom SbyL(v3)=L(v2)L(v1).Rather, wefindthat
L(v3)=RL(v2)L(v1), (4.139)
whereRis a 3×3 space rotation matrix embedded in our four-dimensional space–time.
Withv1andv2not parallel, the final system, S′′,i srotatedrelative to S. This rotation
is the origin of the Thomas precession involved in spin-orbit coupling terms in atomic
and nuclear physics. Because of its presence, the pure Lorentz transformations L(v)by
themselvesdonotformagroup.
Kinematics and Dynamics in Minkowski Space–Time
Wehaveseenthatthepropagationoflightdeterminesthemetric
r2−c2t2=0=r′2−c2t′2,
wherexµ=(ct,r)isthecoordinatefour-vector.Foraparticlemovingwithvelocity v,the
Lorentzinvariantinfinitesimalversion
cdτ≡radicalbig
dxµdxµ=radicalbig
c2dt2−dr2=dtradicalbig
c2−v2
definestheinvariantpropertime τonitstrack.Becauseoftimedilationinmovingframes,
aproper-timeclockrideswiththeparticle(initsrestframe)andrunsattheslowestpossible
4.5 Homogeneous Lorentz Group 281
rate compared to any other inertial frame (of an observer, for example). The four-velocity
oftheparticlecannowbedefinedproperlyas
dxµ
dτ=uµ=parenleftbiggc√
c2−v2,v√
c2−v2parenrightbigg
,
sou2=1, and the four-momentum pµ=cmuµ=(E
c,p)yields Einstein’s famous energy
relation
E=mc2
radicalbig
1−v2/c2=mc2+m
2v2±···.
A consequence of u2=1 and its physical significance is that the particle is on its mass
shellp2=m2c2.
NowweformulateNewton’sequationfora singleparticle ofmassminspecialrelativity
asdpµ
dτ=Kµ, withKµdenoting the force four-vector, so its vector part of the equation
coincideswiththeusualform. For µ=1,2,3w eu s e dτ=dtradicalbig
1−v2/c2andfind
1radicalbig
1−v2/c2dp
dt=Fradicalbig
1−v2/c2=K,
determining Kin terms of the usual force F. We need to find K0. We proceed by analogy
with the derivation of energy conservation, multiplying the force equation into the four-
velocity
muνduν
dτ=m
2du2
dτ=0,
becauseu2=1=const.Theothersideof Newton’sequationyields
0=1
cu·K=K0
radicalbig
1−v2/c2−F·v/c
radicalbig
1−v2/c22,
soK0=F·v/c√
1−v2/c2isrelatedtotherateof workdonebytheforceontheparticle.
Now we turn to two-body collisions, in which energy–momentum conservation takes
the form p1+p2=p3+p4, wherepµ
iare the particle four-momenta. Because the scalar
product of any four-vector with itself is an invariant under Lorentz transformations, it is
convenient to define the Lorentz invariant energy squared s=(p1+p2)2=P2, where
Pµis the total four-momentum, and to use units where the velocity of light c=1. The
laboratory system (lab) is defined as the rest frame of the particle with four-momentum
pµ
2=(m2,0)andthecenterofmomentumframe(cms)bythetotalfour-momentum Pµ=
(E1+E2,0). Whentheincidentlabenergy EL
1isgiven,then
s=p2
1+p2
2+2p1·p2=m2
1+m2
2+2m2EL
1
isdetermined.Now,thecmsenergiesofthefourparticlesareobtainedfromscalarproducts
p1·P=E1(E1+E2)=E1√s,
282 Chapter 4 Group Theory
so
E1=p1·(p1+p2)√s=m2
1+p1·p2√s=m2
1−m2
2+s
2√s,
E2=p2·(p1+p2)√s=m2
2+p1·p2√s=m2
2−m2
1+s
2√s,
E3=p3·(p3+p4)√s=m2
3+p3·p4√s=m2
3−m2
4+s
2√s,
E4=p4·(p3+p4)√s=m2
4+p3·p4√s=m2
4−m2
3+s
2√s,
bysubstituting
2p1·p2=s−m2
1−m2
2,2p3·p4=s−m2
3−m2
4.
Thus, all cms energies Eidepend only on the incident energy but not on the scattering
angle. For elastic scattering, m3=m1,m4=m2,s oE3=E1,E4=E2. The Lorentz
invariantmomentumtransfersquared
t=(p1−p3)2=m2
1+m2
3−2p1·p3
dependslinearlyonthecosineof thescatteringangle.
Example 4.5.1 KAON DECAY AND PIONPHOTOPRODUCTION THRESHOLD
Find the kinetic energies of the muon of mass 106 MeV and massless neutrino into which
aKmesonof mass494MeVdecaysinitsrest frame.
Conservationofenergyandmomentumgives mK=Eµ+Eν=√s.Applyingtherela-
tivistickinematicsdescribedpreviouslyyields
Eµ=pµ·(pµ+pν)
mK=m2
µ+pµ·pν
mK,
Eν=pν·(pµ+pν)
mK=pµ·pν
mK.
Combiningbothresultsweobtain m2
K=m2
µ+2pµ·pν,s o
Eµ=Tµ+mµ=m2
K+m2
µ
2mK=258.4MeV,
Eν=Tν=m2
K−m2
µ
2mK=235.6MeV.
Asanotherexample,intheproductionofaneutralpionbyanincidentphotonaccordingto
γ+p→π0+p′at threshold, the neutral pion and proton are created at rest in the cms.
Therefore,
s=(pγ+p)2=m2
p+2mpEL
γ=(pπ+p′)2=(mπ+mp)2,
4.6 Lorentz Covariance of Maxwell’s Equations 283
soEL
γ=mπ+m2
π
2mp=144.7MeV. /squaresolid
Exercises
4.5.1 Two Lorentz transformations are carried out in succession: v1along the x-axis, then
v2along the y-axis. Show that the resultant transformation (given by the product of
these two successive transformations) cannotbe put in the form of a single Lorentz
transformation.
Note.Thediscrepancycorrespondstoarotation.
4.5.2 Rederive the Lorentz transformation, working entirely in the real space (x0,x1,x2,x3)
withx0=x0=ct. Show that the Lorentz transformation may be written L(v)=
exp(ρσ),with
σ=
0−λ−µ−ν
−λ000
−µ000
−ν000
andλ,µ,νthedirectioncosinesofthevelocity v.
4.5.3 Using the matrix relation, Eq. (4.136), let the rapidity ρ1relate the Lorentz reference
frames(x′0,x′1)and(x0,x1).L e tρ2relate(x′′0,x′′1)and(x′0,x′1). Finally, let ρ
relate(x′′0,x′′1)and(x0,x1).F r o mρ=ρ1+ρ2derive the Einstein velocity addition
law
v=v1+v2
1+v1v2/c2.
4.6 L ORENTZ COVARIANCE OF MAXWELL ’SEQUATIONS
If a physical law is to hold for all orientations of our (real) coordinates (that is, to be in-
variant under rotations), the terms of the equation must be covariant under rotations (Sec-
tions 1.2 and 2.6). This means that we write the physical laws in the mathematical form
scalar=scalar,vector=vector,second-ranktensor =second-ranktensor,andsoon.Sim-
ilarly,ifaphysicallawistoholdforallinertialsystems,thetermsoftheequationmustbe
covariantunderLorentztransformations.
Using Minkowski space ( ct=x0;x=x1,y=x2,z=x3), we have a four-dimensional
spacewiththemetric gµν(Eq.(4.128),Section4.5).TheLorentztransformationsarelinear
inspaceandtimeinthisfour-dimensionalrealspace.19
19Agroup theoreticderivationoftheLorentztransformation inMinkowskispaceappearsinSection4.5.SeealsoH.Goldstein,
Classical Mechanics . Cambridge, MA: Addison-Wesley (1951), Chapter 6. The metric equation x2
0−x2=0, independent of
referenceframe, leads to the Lorentz transformations.
284 Chapter 4 Group Theory
HereweconsiderMaxwell’sequations,
∇×E=−∂B
∂t, (4.140a)
∇×H=∂D
∂t+ρv, (4.140b)
∇·D=ρ, (4.140c)
∇·B=0, (4.140d)
andtherelations
D=ε0E,B=µ0H. (4.141)
The symbols have their usual meanings as given in Section 1.9. For simplicity we assume
vacuum( ε=ε0,µ=µ0).
We assume that Maxwell’s equations hold in all inertial systems; that is, Maxwell’s
equations are consistent with special relativity. (The covariance of Maxwell’s equations
under Lorentz transformations was actually shown by Lorentz and Poincaré before Ein-
steinproposedhistheoryofspecialrelativity.)OurimmediategoalistorewriteMaxwell’s
equations as tensor equations in Minkowski space. This will make the Lorentz covariance
explicit,ormanifest.
In terms of scalar, ϕ, and magnetic vector potentials, A,w em a ys o l v e20Eq. (4.140d)
andthen(4.140a)by
B=∇×A
E=−∂A
∂t−∇ϕ. (4.142)
Equation (4.142) specifies the curl of A; the divergence of Ais still undefined (compare
Section 1.16). We may, and for future convenience we do, impose a further gauge restric-
tiononthevectorpotential A:
∇·A+ε0µ0∂ϕ
∂t=0. (4.143)
This is the Lorentz gauge relation. It will serve the purpose of uncoupling the differential
equations for Aandϕthat follow. The potentials Aandϕare not yet completely fixed.
Thefreedomremainingis thetopicofExercise4.6.4.
Now we rewrite the Maxwell equations in terms of the potentials Aandϕ.F r o m
Eqs. (4.140c)for ∇·D,(4.141)and(4.142),
∇2ϕ+∇·∂A
∂t=−ρ
ε0, (4.144)
whereasEqs. (4.140b)for ∇×Hand(4.142)andEq.(1.86c) ofChapter1yield
∂2A
∂t2+∇∂ϕ
∂t+1
ε0µ0braceleftbig
∇∇·A−∇2Abracerightbig
=ρv
ε0. (4.145)
20Compare Section1.13, especiallyExercise1.13.10.
4.6 Lorentz Covariance of Maxwell’s Equations 285
UsingtheLorentzrelation,Eq. (4.143), andtherelation ε0µ0=1/c2, weobtain
bracketleftbigg
∇2−1
c2∂2
∂t2bracketrightbigg
A=−µ0ρv,
bracketleftbigg
∇2−1
c2∂2
∂t2bracketrightbigg
ϕ=−ρ
ε0. (4.146)
Now,thedifferentialoperator(see alsoExercise2.7.3)
∇2−1
c2∂2
∂t2≡−∂2≡−∂µ∂µ
is a four-dimensional Laplacian, usually called the d’Alembertian and also sometimes de-
notedby /fill50.It isascalarbyconstruction(seeExercise2.7.3).
Forconveniencewedefine
A1≡Ax
µ0c=cε0Ax,A3≡Az
µ0c=cε0Az,
A2≡Ay
µ0c=cε0Ay,A 0≡ε0ϕ=A0.(4.147)
If wefurther defineafour-vectorcurrentdensity
ρvx
c≡j1,ρvy
c≡j2,ρvz
c≡j3,ρ≡j0=j0, (4.148)
thenEq. (4.146)maybewrittenintheform
∂2Aµ=jµ. (4.149)
Thewaveequation(4.149)lookslikeafour-vectorequation,butlooksdonotconstitute
proof.Toprovethatitisafour-vectorequation,westartbyinvestigatingthetransformation
propertiesof thegeneralizedcurrent jµ.
Sinceanelectricchargeelement deis aninvariantquantity,wehave
de=ρdx1dx2dx3,invariant. (4.150)
WesawinSection2.9thatthefour-dimensionalvolumeelement dx0dx1dx2dx3wasalso
invariant,apseudoscalar.Comparingthisresult, Eq. (2.106), withEq. (4.150), wesee that
thechargedensity ρmusttransformthesamewayas dx0,thezerothcomponentofafour-
dimensionalvector dxλ.Weputρ=j0,withj0nowestablishedasthezerothcomponent
ofafour-vector.Theotherparts ofEq. (4.148)maybeexpandedas
j1=ρvx
c=ρ
cdx1
dt=j0dx1
dx0. (4.151)
Sincewehavejustshownthat j0transformsas dx0,thismeansthat j1transformsas dx1.
With similar results for j2andj3,W eh a v e jλtransforming as dxλ, proving that jλis a
four-vectorinMinkowskispace.
Equation (4.149), which follows directly from Maxwell’s equations, Eqs. (4.140), is
assumed to hold in all Cartesian systems (all Lorentz frames). Then, by the quotient rule,
Section2.8, AµisalsoavectorandEq. (4.149)is alegitimatetensorequation.
286 Chapter 4 Group Theory
Now,workingbackward,Eq. (4.142)maybewritten
ε0Ej=−∂Aj
∂x0−∂A0
∂xj,j=1,2,3,
(4.152)
1
µ0cBi=∂Ak
∂xj−∂Aj
∂xk,(i,j,k)=cyclic(1,2,3).
We defineanewtensor,
∂µAλ−∂λAµ=∂Aλ
∂xµ−∂Aµ
∂xλ≡Fµλ=−Fλµ(µ,λ=0,1,2,3),
anantisymmetricsecond-ranktensor,since Aλis avector.Writtenoutexplicitly,
Fµλ
ε0=
0ExEyEz
−Ex0−cBzcBy
−EycBz0−cBx
−Ez−cBycBx0
,Fµλ
ε0=
0−Ex−Ey−Ez
Ex0−cBzcBy
EycBz0−cBx
Ez−cBycBx0
.
(4.153)
Noticethatinourfour-dimensionalMinkowskispace EandBarenolongervectorsbutto-
getherformasecond-ranktensor.Withthistensorwemaywritethetwononhomogeneous
Maxwellequations((4.140b)and(4.140c)) combinedas atensorequation,
∂Fλµ
∂xµ=jλ. (4.154)
Theleft-handsideofEq.(4.154)isafour-dimensionaldivergenceofatensorandtherefore
a vector. This, of course, is equivalent to contracting a third-rank tensor ∂Fλµ/∂xν(com-
pareExercises2.7.1and2.7.2).ThetwohomogeneousMaxwellequations—(4.140a)for
∇×Eand(4.140d)for ∇·B— maybeexpressedinthetensorform
∂F23
∂x1+∂F31
∂x2+∂F12
∂x3=0 (4.155)
forEq. (4.140d)andthreeequationsof theform
−∂F30
∂x2−∂F02
∂x3+∂F23
∂x0=0 (4.156)
forEq. (4.140a). (Asecondequationpermutes120,athirdpermutes130.) Since
∂λFµν=∂Fµν
∂xλ≡tλµν
isatensor(of thirdrank), Eqs. (4.140a)and(4.140d)aregivenbythetensorequation
tλµν+tνλµ+tµνλ=0. (4.157)
FromEqs.(4.155)and(4.156)youwillunderstandthattheindices λ,µ,andνaresupposed
to be different. Actually Eq. (4.157) automatically reduces to 0 =0 if any two indices
coincide.Analternateform ofEq. (4.157)appearsinExercise4.6.14.
4.6 Lorentz Covariance of Maxwell’s Equations 287
Lorentz Transformation of EandB
Theconstructionofthetensorequations((4.154)and(4.157))completesourinitialgoalof
rewriting Maxwell’s equations in tensor form.21Now we exploit the tensor properties of
ourfourvectorsandthetensor Fµν.
For the Lorentz transformation corresponding to motion along the z(x3)-axis with ve-
locityv,the“directioncosines”aregivenby22
x′0=γparenleftbig
x0−βx3parenrightbig
x′3=γparenleftbig
x3−βx0parenrightbig
,(4.158)
where
β=v
c
and
γ=parenleftbig
1−β2parenrightbig−1/2. (4.159)
Using the tensor transformation properties, we may calculate the electric and magnetic
fields in the moving system in terms of the values in the original reference frame. From
Eqs. (2.66), (4.153),and(4.158)weobtain
E′
x=1radicalbig
1−β2parenleftbigg
Ex−v
c2Byparenrightbigg
,
E′
y=1radicalbig
1−β2parenleftbigg
Ey+v
c2Bxparenrightbigg
, (4.160)
E′
z=Ez
and
B′
x=1radicalbig
1−β2parenleftbigg
Bx+v
c2Eyparenrightbigg
,
B′
y=1radicalbig
1−β2parenleftbigg
By−v
c2Exparenrightbigg
, (4.161)
B′
z=Bz.
Thiscouplingof EandBistobeexpected.Consider,forinstance,thecaseofzeroelectric
fieldintheunprimedsystem
Ex=Ey=Ez=0.
21Modern theories of quantum electrodynamics and elementary particles are often written in this “manifestly covariant” form
to guarantee consistency with special relativity. Conversely, the insistence on such tensor form has been a useful guide in the
construction ofthese theories.
22Agroup theoretic derivation of theLorentz transformation appears in Section 4.5. Seealso Goldstein, loc. cit.,Chapter6.
288 Chapter 4 Group Theory
Clearly, there will be no force on a stationary charged particle. When the particle is in
motion with a small velocity valong the z-axis,23an observer on the particle sees fields
(exertingaforceonhischargedparticle)givenby
E′
x=−vBy,
E′
y=vBx,
whereBisamagneticinductionfieldintheunprimedsystem.Theseequationsmaybeput
invectorform,
E′=v×B
or (4.162)
F=qv×B,
whichisusuallytakenas theoperationaldefinitionof themagneticinduction B.
Electromagnetic Invariants
Finally, the tensor (or vector) properties allow us to construct a multitude of invariant
quantities.Amoreimportantoneisthescalarproductofthetwofour-dimensionalvectors
orfour-vectors Aλandjλ.W eha v e
Aλjλ=−cε0Axρvx
c−cε0Ayρvy
c−cε0Azρvz
c+ε0ϕρ
=ε0(ρϕ−A·J),invariant, (4.163)
withAthe usual magnetic vector potential and Jthe ordinary current density. The first
term,ρϕ, is the ordinary static electric coupling, with dimensions of energy per unit vol-
ume. Hence our newly constructed scalar invariant is an energy density. The dynamic in-
teraction of field and current is given by the product A·J. This invariant Aλjλappears in
theelectromagneticLagrangiansofExercises17.3.6and17.5.1.
OtherpossibleelectromagneticinvariantsappearinExercises4.6.9 and4.6.11.
The Lorentz group is the symmetry group of electrodynamics, of the electroweak gauge
theory, and of the strong interactions described by quantum chromodynamics: It governs
special relativity. The metric of Minkowski space–time is Lorentz invariant and expresses
the propagation of light; that is, the velocity of light is the same in all inertial frames.
Newton’sequationsofmotionarestraightforwardtoextendtospecialrelativity.Thekine-
matics of two-body collisions are important applications of vector algebra in Minkowski
space–time.
23If thevelocity is not small, arelativistic transformation offorce is needed.
4.6 Lorentz Covariance of Maxwell’s Equations 289
Exercises
4.6.1 (a) Show that every four-vector in Minkowski space may be decomposed into an or-
dinary three-space vector and a three-space scalar. Examples: (ct,r),(ρ,ρv/c),
(ε0ϕ,cε0A),(E/c,p),(ω/c,k).
Hint.Considera rotationofthethree-spacecoordinateswithtimefixed.
(b) Showthattheconverseof(a)is nottrue—everythree-vectorplusscalardoes not
formaMinkowskifour-vector.
4.6.2 (a) Showthat
∂µjµ=∂·j=∂jµ
∂xµ=0.
(b) Show how the previous tensor equation may be interpreted as a statement of con-
tinuityofchargeandcurrentinordinarythree-dimensionalspaceandtime.
(c) If this equation is known to hold in all Lorentz reference frames, why can we not
concludethat jµis avector?
4.6.3 Write the Lorentz gauge condition (Eq. (4.143)) as a tensor equation in Minkowski
space.
4.6.4 Agaugetransformationconsistsofvaryingthescalarpotential ϕ1andthevectorpoten-
tialA1accordingtotherelation
ϕ2=ϕ1+∂χ
∂t,
A2=A1−∇χ.
Thenewfunction χisrequiredtosatisfythehomogeneouswaveequation
∇2χ−1
c2∂2χ
∂t2=0.
Showthefollowing:
(a) TheLorentzgaugerelationisunchanged.
(b) The new potentials satisfy the same inhomogeneous wave equations as did the
originalpotentials.
(c) Thefields EandBareunaltered.
The invariance of our electromagnetic theory under this transformation is called gauge
invariance .
4.6.5 Achargedparticle,charge q,massm, obeystheLorentzcovariantequation
dpµ
dτ=q
ε0mcFµνpν,
wherepνis the four-momentum vector (E/c;p1,p2,p3),τis the proper time, dτ=
dtradicalbig
1−v2/c2, aLorentzscalar.Showthattheexplicitspace–timeformsare
dE
dt=qv·E;dp
dt=q(E+v×B).
290 Chapter 4 Group Theory
4.6.6 From the Lorentz transformation matrix elements (Eq. (4.158)) derive the Einstein ve-
locityadditionlaw
u′=u−v
1−(uv/c2)oru=u′+v
1+(u′v/c2),
whereu=cdx3/dx0andu′=cdx′3/dx′0.
Hint.I fL12(v)is the matrix transforming system 1 into system 2, L23(u′)the matrix
transforming system 2 into system 3, L13(u)the matrix transforming system 1 directly
into system 3, then L13(u)=L23(u′)L12(v). From this matrix relation extract the Ein-
steinvelocityadditionlaw.
4.6.7 The dual of a four-dimensional second-rank tensor Bmay be defined by ˜B, where the
elementsofthedualtensorare givenby
˜Bij=1
2!εijklBkl.
Showthat˜Btransforms as
(a) asecond-ranktensorunderrotations,
(b) apseudotensorunderinversions.
Note.Thetildeheredoes notmeantranspose.
4.6.8 Construct˜F, thedualof F, whereFis theelectromagnetictensorgivenbyEq. (4.153).
ANS.˜Fµν=ε0
0−cBx−cBy−cBz
cBx0Ez−Ey
cBy−Ez0Ex
cBzEy−Ex0
.
Thiscorrespondsto
cB→−E,
E→cB.
This transformation,sometimescalleda dualtransformation ,leavesMaxwell’sequa-
tionsinvacuum (ρ=0)invariant.
4.6.9 Because the quadruple contraction of a fourth-rank pseudotensor and two second-rank
tensorsεµλνσFµλFνσisclearlyapseudoscalar,evaluateit.
ANS.−8ε2
0cB·E.
4.6.10 (a) Ifanelectromagneticfieldispurelyelectric(orpurelymagnetic)inoneparticular
Lorentz frame, show that EandBwill be orthogonal in other Lorentz reference
systems.
(b) Conversely,if EandBareorthogonalinoneparticularLorentzframe,thereexists
aLorentzreferencesysteminwhich E(orB)vanishes.Findthatreferencesystem.
4.7 Discrete Groups 291
4.6.11 Showthat c2B2−E2is aLorentzscalar.
4.6.12 Since(dx0,dx1,dx2,dx3)is a four-vector, dxµdxµis a scalar. Evaluate this scalar
for a movingparticlein twodifferent coordinatesystems:(a) a coordinatesystemfixed
relativetoyou(labsystem),and(b)acoordinatesystemmovingwithamovingparticle
(velocity vrelative to you). With the time increment labeled dτin the particle system
anddtinthelabsystem, showthat
dτ=dtradicalBig
1−v2/c2.
τis thepropertimeoftheparticle,aLorentzinvariantquantity.
4.6.13 Expandthescalarexpression
−1
4ε0FµνFµν+1
ε0jµAµ
in terms of the fields and potentials. The resulting expression is the Lagrangian density
usedinExercise17.5.1.
4.6.14 ShowthatEq.(4.157) maybewritten
εαβγδ∂Fαβ
∂xγ=0.
4.7 D ISCRETE GROUPS
Here we consider groups with a finite number of elements. In physics, groups usually ap-
pear as a set of operations that leave a system unchanged, invariant. This is an expression
ofsymmetry.Indeed,asymmetrymaybedefinedastheinvarianceoftheHamiltonianofa
system under a group of transformations. Symmetry in this sense is important in classical
mechanics, but it becomes even more important and more profound in quantum mechan-
ics. In this section we investigate the symmetry properties of sets of objects (atoms in a
molecule or crystal). This provides additional illustrations of the group concepts of Sec-
tion4.1andleadsdirectlytodihedralgroups.Thedihedralgroupsinturnopenupthestudy
of the 32 crystallographic point groups and 230 space groups that are of such importance
in crystallography and solid-state physics. It might be noted that it was through the study
of crystal symmetries that the concepts of symmetry and group theory entered physics. In
physics, the abstract group conditions often take on direct physical meaning in terms of
transformationsof vectors,spinors,andtensors.
As a simple, but not trivial, example of a finite group, consider the set 1 ,a,b,cthat
combine according to the group multiplication table24(see Fig. 4.10). Clearly, the four
conditions of the definition of “group” are satisfied. The elements a,b,c, and 1 are ab-
stract mathematical entities, completely unrestricted except for the multiplication table of
Fig.4.10.
Now,for aspecificrepresentationofthesegroupelements,let
1→1,a→i, b→−1,c→−i, (4.164)
24Theorder of the factors is row–column: ab=cin the indicatedprevious example.
292 Chapter 4 Group Theory
FIGURE 4.10Group
multiplicationtable.
combining by ordinary multiplication. Again, the four group conditions are satisfied, and
these four elements form a group. We label this group C4. Since the multiplication of
the group elements is commutative, the group is labeled commutative ,o rabelian.O u r
group is also a cyclic group , in that the elements may be written as successive powers of
one element, in this case in,n=0,1,2,3. Note that in writing out Eq. (4.164) we have
selectedaspecificfaithfulrepresentationfor thisgroupoffour objects, C4.
We recognize that the group elements 1 ,i,−1,−imay be interpreted as successive 90◦
rotationsinthecomplexplane.Then,fromEq.(3.74),wecreatethesetoffour2 ×2matri-
ces(replacing ϕby−ϕinEq. (3.74)torotateavectorratherthanrotatethecoordinates):
R(ϕ)=parenleftBigg
cosϕ−sinϕ
sinϕcosϕparenrightBigg
,
andforϕ=0,π/2,π,and 3π/2w eh a v e
1=parenleftBigg
10
01parenrightBigg
A=parenleftBigg
0−1
10parenrightBigg
B=parenleftBigg
−10
0−1parenrightBigg
C=parenleftBigg
01
−10parenrightBigg
.(4.165)
Thissetoffourmatricesformsagroup,withthelawofcombinationbeingmatrixmultipli-
cation. Here is a second faithful representation. By matrix multiplication one verifies that
thisrepresentationisalsoabelianandcyclic.Clearly,thereisaone-to-onecorrespondence
ofthetworepresentations
1↔1↔1a↔i↔Ab↔−1↔Bc↔−i↔C. (4.166)
Inthegroup C4thetworepresentations (1,i,−1,−i)and(1,A,B,C)areisomorphic.
In contrast to this, there is no such correspondence between either of these representa-
tionsofgroup C4andanothergroupoffourobjects,thevierergruppe(Exercise3.2.7).The
Table 4.3
1V1V2V3
11V1V2V3
V1V11V3V2
V2V2V31V1
V3V3V2V11
4.7 Discrete Groups 293
vierergruppe has the multiplicationtable shown in Table 4.3. Confirming the lack of cor-
respondence between the group represented by (1,i,−1,−i)or the matrices (1,A,B,C)
of Eq. (4.165), note that although the vierergruppe is abelian, it is not cyclic. The cyclic
groupC4andthevierergruppearenotisomorphic.
Classes and Character
Consider a group element xtransformed into a group element yby a similarity transform
withrespectto gi, anelementofthegroup
gixg−1
i=y. (4.167)
The group element yisconjugate tox.Aclassis a set of mutually conjugate group ele-
ments.Ingeneral,thissetofelementsformingaclassdoesnotsatisfythegrouppostulates
and is not a group. Indeed, the unit element 1, which is always in a class by itself, is the
onlyclassthatisalsoasubgroup.Allmembersofagivenclassareequivalent,inthesense
that any one element is a similarity transform of any other element. Clearly, if a group is
abelian,everyelementis aclassbyitself. Wefindthat
1. Everyelementoftheoriginalgroupbelongstooneandonlyoneclass.
2. Thenumberofelementsinaclassis afactor oftheorderofthegroup.
We get a possible physical interpretation of the concept of class by noting that yis a
similaritytransformof x.Ifgirepresentsarotationofthecoordinatesystem,then yisthe
sameoperationas xbutrelativetothenew,relatedcoordinates.
In Section 3.3 we saw that a real matrix transforms under rotation of the coordinates
by an orthogonal similarity transformation. Depending on the choice of reference frame,
essentiallythesamematrixmaytakeonaninfinityofdifferentforms.Likewise,ourgroup
representations may be put in an infinity of different forms by using unitary transforma-
tions. But each such transformed representation is isomorphic with the original. From Ex-
ercise3.3.9thetraceofeachelement(eachmatrixofourrepresentation)isinvariantunder
unitarytransformations.Justbecauseitisinvariant,thetrace(relabeledthe character )as-
sumesaroleofsomeimportanceingrouptheory,particularlyinapplicationstosolid-state
physics. Clearly, all members of a given class (in a given representation) have the same
character. Elements of different classes may have the same character, but elements with
differentcharacterscannotbeinthesameclass.
The concept of class is important (1) because of the trace or character and (2) because
the number of nonequivalent irreducible representations of a group is equal to the
numberofclasses.
Subgroups and Cosets
Frequently a subset of the group elements (including the unit element I) will by itself
satisfythefourgrouprequirementsandthereforeisagroup.Suchasubsetiscalleda sub-
group. Every group has two trivial subgroups: the unit element alone and the group itself.
The elements 1 and bof the four-element group C4discussed earlier form a nontrivial
294 Chapter 4 Group Theory
subgroup. In Section 4.1 we consider SO(3), the (continuous) group of all rotations in or-
dinary space. The rotations about any single axis form a subgroup of SO(3). Numerous
otherexamplesofsubgroupsappearinthefollowingsections.
Considerasubgroup Hwithelements hiandagroupelement xnotinH.Thenxhiand
hixarenotinsubgroup H. Thesets generatedby
xhi,i=1,2,...andhix, i=1,2,...
are called cosets, respectively the left and right cosets of subgroup Hwith respect to x.I t
canbeshown(assumethecontraryandproveacontradiction)thatthecosetofasubgroup
has the same number of distinct elements as the subgroup. Extending this result we may
expresstheoriginalgroup Gasthesumof Handcosets:
G=H+x1H+x2H+···.
Then the order of any subgroup is a divisor of the order of the group . It is this result
that makes the concept of coset significant. In the next section the six-element group D3
(order 6) has subgroups of order 1, 2, and 3. D3cannot (and does not) have subgroups of
order4or5.
Thesimilaritytransformofasubgroup Hbyafixedgroupelement xnotinH,xHx−1,
yieldsasubgroup—Exercise4.7.8.Ifthisnewsubgroupisidenticalwith Hforallx,that
is,
xHx−1=H,
thenHis called an invariant, normal ,o rself-conjugate subgroup . Such subgroups are
involved in the analysis of multiplets of atomic and nuclear spectra and the particles dis-
cussed in Section 4.2. All subgroups of a commutative (abelian) group are automatically
invariant.
Two Objects — Twofold Symmetry Axis
Consider first the two-dimensional system of two identical atoms in the xy-plane at (1,
0) and (−1, 0), Fig. 4.11. What rotations25can be carried out (keeping both atoms in the
xy-plane) that will leave this system invariant? The first candidate is, of course, the unit
operator1.Arotationof πradiansaboutthe z-axiscompletesthelist.Sowehavearather
uninteresting group of two members (1, −1). Thez-axis is labeled a twofold symmetry
axis—correspondingtothetworotationangles,0and π, thatleavethesysteminvariant.
Our system becomes more interesting in three dimensions. Now imagine a molecule
(or part of a crystal) with atoms of element Xat±aon thex-axis, atoms of element Y
at±bon they-axis, and atoms of element Zat±con thez-axis, as show in Fig. 4.12.
Clearly,eachaxisisnowatwofoldsymmetryaxis.Using Rx(π)todesignatearotationof
πradiansaboutthe x-axis,wemay
25Herewedeliberatelyexcludereflectionsandinversions.Theymustbebroughtintodevelopthefullsetof32crystallographic
point groups.
4.7 Discrete Groups 295
FIGURE 4.11DiatomicmoleculesH 2,N2,O2,
Cl2.
FIGURE 4.12D2symmetry.
setupamatrixrepresentationoftherotationsas inSection3.3:
Rx(π)=
10 0
0−10
00−1
,Ry(π)=
−100
010
00−1
,
Rz(π)=
−100
0−10
001
, 1=
100
010
001
.(4.168)
Thesefourelements [1,Rx(π),Ry(π),Rz(π)]formanabeliangroup,withthegroupmul-
tiplicationtableshowninTable4.4.
The products shown in Table 4.4 can be obtained in either of two distinct ways:
(1) We may analyze the operations themselves—a rotation of πabout the x-axis fol-
lowed by a rotation of πabout the y-axis is equivalent to a rotation of πabout the z-axis:
Ry(π)Rx(π)=Rz(π). (2) Alternatively, once a faithful representation is established, we
296 Chapter 4 Group Theory
Table 4.4
1R x(π)Ry(π)Rz(π)
1 1R xRyRx
Rx(π)Rx1R zRy
Ry(π)RyRz1R x
Rz(π)RzRyRx1
can obtain the products by matrix multiplication. This is where the power of mathematics
isshown—whenthesystemis toocomplexfor adirectphysicalinterpretation.
Comparison with Exercises 3.2.7, 4.7.2, and 4.7.3 shows that this group is the vier-
ergruppe. The matrices of Eq. (4.168) are isomorphic with those of Exercise 3.2.7. Also,
they are reducible, being diagonal. The subgroups are (1,Rx),(1,Ry), and(1,Rz).T h e y
are invariant. It should be noted that a rotation of πabout the y-axis and a rotation of π
about the z-axis is equivalent to a rotation of πabout the x-axis:Rz(π)Ry(π)=Rx(π).
In symmetry terms, if yandzare twofold symmetry axes, xis automatically a twofold
symmetryaxis.
This symmetry group,26the vierergruppe, is often labeled D2,t h eDsignifying a dihe-
dralgroupandthesubscript2signifyingatwofoldsymmetryaxis(andnohighersymmetry
axis).
Three Objects — Threefold Symmetry Axis
Consider now three identical atoms at the vertices of an equilateral triangle, Fig. 4.13.
Rotationsofthe triangleof0,2π/3,and4π/3leavethetriangleinvariant.Inmatrixform,
wehave27
1=Rz(0)=parenleftBigg
10
01parenrightBigg
A=Rz(2π/3)=parenleftBigg
cos2π/3−sin2π/3
sin2π/3 cos2 π/3parenrightBigg
=parenleftBigg
−1/2−√
3/2
√
3/2−1/2parenrightBigg
B=Rz(4π/3)=parenleftBigg
−1/2√
3/2
−√
3/2−1/2parenrightBigg
. (4.169)
Thez-axis is a threefold symmetry axis. (1,A,B)form a cyclic group, a subgroup of the
completesix-elementgroupthatfollows.
In thexy-plane there are three additional axes of symmetry—each atom (vertex) and
thegeometriccenterdefininganaxis.Eachoftheseisatwofoldsymmetryaxis.Theserota-
tionsmaymosteasilybedescribedwithinourtwo-dimensionalframeworkbyintroducing
26Asymmetry group isagroup ofsymmetry-preserving operations, thatis,rotations, reflections,andinversions. A symmetric
group is the group of permutations of ndistinct objects—of order n!.
27Note thathere wearerotating the trianglecounterclockwise relative to fixed coordinates.
4.7 Discrete Groups 297
FIGURE 4.13Symmetryoperationsonan
equilateraltriangle.
reflections. The rotation of πabout the C-( o ry-) axis, which means the interchanging of
(structureless)atoms aandc,is justareflectionofthe x-axis:
C=RC(π)=parenleftBigg
−10
01parenrightBigg
. (4.170)
We may replace the rotation about the D-axis by a rotation of 4 π/3 (about our z-axis)
followedbyareflectionofthe x-axis(x→−x)(Fig.4.14):
D=RD(π)=CB
=parenleftBigg
−10
01parenrightBiggparenleftBigg
−1/2√
3/2
−√
3/2−1/2parenrightBigg
=parenleftBigg
1/2−√
3/2
−√
3/2−1/2parenrightBigg
. (4.171)
FIGURE 4.14Thetriangleontherightisthetriangleon
theleftrotated180◦aboutthe D-axis.D=CB.
298 Chapter 4 Group Theory
In a similar manner, the rotation of πabout the E-axis, interchanging aandb, is replaced
byarotationof 2 π/3(A)andthenareflection28ofthex-axis:
E=RE(π)=CA
=parenleftBigg
−10
01parenrightBiggparenleftBigg
−1/2−√
3/2
√
3/2−1/2parenrightBigg
=parenleftBigg
1/2√
3/2
√
3/2−1/2parenrightBigg
. (4.172)
Thecompletegroupmultiplicationtableis
1ABCDE
11ABCDE
AAB1DEC
BB1AECD
CCED1BA
DDCEA1B
EEDCBA1
Noticethateachelementofthegroupappearsonlyonceineachrowandineachcolumn,as
requiredbytherearrangementtheorem,Exercise4.7.4.Also,fromthemultiplicationtable
the group is not abelian. We have constructed a six-element group and a 2 ×2 irreducible
matrix representation of it. The only other distinct six-element group is the cyclic group
[1,R,R2,R3,R4,R5], with
R=e2πi/6orR=e−πiσ2/3=parenleftBigg
1/2−√
3/2
√
3/21/2parenrightBigg
. (4.173)
Our group[1,A,B,C,D,E]is labeled D3in crystallography, the dihedral group with a
threefold axis of symmetry. The three axes ( C,D, andE)i nt h exy-plane automatically
become twofold symmetry axes. As a consequence, (1,C),(1,D), and(1,E)all form
two-elementsubgroups.Noneofthesetwo-elementsubgroupsof D3isinvariant.
Ageneralandmostimportantresultfor finitegroupsof helementsisthat
summationdisplay
in2
i=h, (4.174)
whereniisthedimensionofthematricesofthe ithirreduciblerepresentation.Thisequal-
ity, sometimes called the dimensionality theorem , is very useful in establishing the irre-
ducible representations of a group. Here for D3we have 12+12+22=6 for our three
representations. No other irreducible representations of this symmetry group of three ob-
jects exist. (The other representations are the identity and ±1, depending upon whether a
reflectionwasinvolved.)
28Note that, as a consequence of these reflections, det (C)=det(D)=det(E)=−1. The rotations AandB, of course, have a
determinant of +1.
4.7 Discrete Groups 299
FIGURE 4.15Ruthenocene.
Dihedral Groups, Dn
Adihedralgroup Dnwithann-foldsymmetryaxisimplies naxeswithangularseparation
of2π/nradians,nisapositiveinteger,butotherwiseunrestricted.Ifweapplythesymme-
try arguments to crystal lattices , thennis limited to 1, 2, 3, 4, and 6. The requirement of
invariance of the crystal lattice under translations in the plane perpendicular to the n-fold
axis excludes n=5,7, and higher values. Try to cover a plane completely with identical
regularpentagonsandwithnooverlapping.29Forindividualmolecules,thisconstraintdoes
not exist, although the examples with n>6 are rare. n=5 is a real possibility. As an ex-
ample,thesymmetrygroupforruthenocene, (C5H5)2Ru, illustratedinFig.4.15,is D5.30
Crystallographic Point and Space Groups
The dihedral groups just considered are examples of the crystallographic point groups.
A point group is composed of combinations of rotations and reflections (including inver-
sions) that will leave some crystal lattice unchanged. Limiting the operations to rotations
and reflections (including inversions) means that one point—the origin—remains fixed,
hence the term point group . Including the cyclic groups, two cubic groups (tetrahedron
andoctahedronsymmetries),andtheimproperforms(involvingreflections),wecometoa
totalof 32crystallographicpointgroups.
29ForD6imagine a planecovered with regular hexagons and theaxis of rotation through the geometric center of one of them.
30Actuallythefull technicallabel is D5h,withhindicating invariance under a reflection of the fivefold axis.
300 Chapter 4 Group Theory
If, to the rotation and reflection operations that produced the point groups, we add the
possibility of translations and still demand that some crystal lattice remain invariant, we
come to the space groups. There are 230 distinct space groups, a number that is appalling
except,possibly,tospecialistsinthefield.Fordetails(whichcancoverhundredsofpages)
seetheAdditionalReadings.
Exercises
4.7.1 Showthatthematrices 1,A,B, andCofEq. (4.165)arereducible.Reducethem.
Note.Thismeanstransforming AandCtodiagonalform(bythesameunitarytransfor-
mation).
Hint.AandCareanti-Hermitian.Theireigenvectorswillbeorthogonal.
4.7.2 Possibleoperationsonacrystallatticeinclude Aπ(rotationby π),m(reflection),and i
(inversion).Thesethreeoperationscombineas
A2
π=m2=i2=1,
Aπ·m=i, m·i=Aπ,andi·Aπ=m.
Showthatthegroup (1,Aπ,m,i)isisomorphicwiththevierergruppe.
4.7.3 Fourpossibleoperationsinthe xy-planeare:
1. nochangebraceleftBigg
x→x
y→y
2. inversionbraceleftBigg
x→−x
y→−y
3. reflectionbraceleftBigg
x→−x
y→y
4. reflectionbraceleftBigg
x→x
y→−y.
(a) Showthatthesefour operationsformagroup.
(b) Showthatthisgroupis isomorphicwiththevierergruppe.
(c) Setupa 2 ×2 matrixrepresentation.
4.7.4 Rearrangement theorem: Given a group of n distinct elements (I,a,b,c,...,n) ,s h o w
that the set of products (aI,a2,ab,ac...an) reproduces the ndistinct elements in a
neworder.
4.7.5 Usingthe 2×2 matrixrepresentationofExercise3.2.7for thevierergruppe,
(a) Showthatthereare fourclasses, eachwithoneelement.
4.7 Discrete Groups 301
(b) Calculate the character (trace) of each class. Note that two different classes may
havethesamecharacter.
(c) Show that there are three two-element subgroups. (The unit element by itself al-
waysforms asubgroup.)
(d) For any one of the two-element subgroups show that the subgroup and a single
cosetreproducetheoriginalvierergruppe.
Notethatsubgroups,classes, andcosetsareentirelydifferent.
4.7.6 Usingthe 2×2 matrixrepresentation,Eq. (4.165),of C4,
(a) Showthatthereare fourclasses, eachwithoneelement.
(b) Calculatethecharacter(trace)ofeachclass.
(c) Showthatthereis onetwo-elementsubgroup.
(d) Showthatthesubgroupandasinglecosetreproducetheoriginalgroup.
4.7.7 Prove that the number of distinct elements in a coset of a subgroup is the same as the
numberof elementsinthesubgroup.
4.7.8 A subgroup Hhas elements hi.L e txbe a fixed element of the original group Gand
notamemberof H. The transform
xhix−1,i=1,2,...
generates a conjugate subgroup xHx−1. Show that this conjugate subgroup satisfies
eachof thefour grouppostulatesandthereforeis agroup.
4.7.9 (a) A particular group is abelian. A second group is created by replacing gibyg−1
i
foreachelementintheoriginalgroup.Showthatthetwogroupsareisomorphic.
Note.Thismeansshowingthatif aibi=ci, thena−1
ib−1
i=c−1
i.
(b) Continuing part (a), if the two groups are isomorphic, show that each must be
abelian.
4.7.10 (a) Once you have a matrix representation of any group, a one-dimensional represen-
tation can be obtained by taking the determinants of the matrices. Show that the
multiplicativerelationsarepreservedinthisdeterminantrepresentation.
(b) Usedeterminantstoobtainaone-dimensionalrepresentativeof D3.
4.7.11 Explainhowtherelation
summationdisplay
in2
i=h
appliestothevierergruppe (h=4)andtothedihedralgroup D3withh=6.
4.7.12 Showthatthesubgroup (1,A,B)ofD3is aninvariantsubgroup.
4.7.13 Thegroup D3maybediscussedasa permutation groupofthreeobjects.Matrix B,for
instance,rotatesvertex a(originallyinlocation1)tothepositionformerlyoccupiedby c
302 Chapter 4 Group Theory
(location 3). Vertex bmoves from location 2 to location 1, and so on. As a permutation
(abc)→(bca).In threedimensions
010
001
100
a
b
c
=
b
c
a
.
(a) Developanalogous 3 ×3 representationsfor theotherelementsof D3.
(b) Reduceyour 3 ×3 representationtothe 2 ×2 representationof thissection.
(This 3×3 representationmustbereducibleor Eq.(4.174)wouldbeviolated.)
Note. The actual reduction of a reducible representation may be awkward. It is often
easiertodevelopdirectlyanewrepresentationof therequireddimension.
4.7.14 (a) The permutation group of four objects P4has 4!=24 elements. Treating the four
elementsofthecyclicgroup C4aspermutations,setupa 4 ×4 matrixrepresenta-
tionofC4.C4thatbecomesasubgroupof P4.
(b) Howdoyouknowthatthis 4 ×4 matrixrepresentationof C4mustbereducible?
Note.C4is abelian and every abelian group of hobjects has only hone-dimensional
irreduciblerepresentations.
4.7.15 (a) Theobjects (abcd)arepermutedto (dacb).Writeouta4×4matrixrepresentation
ofthisonepermutation.
(b) Is thepermutation (abdc)→(dacb)oddor even?
(c) Is thispermutationapossiblememberofthe D4group?Whyor whynot?
4.7.16 Theelementsofthedihedralgroup Dnmaybewrittenintheform
SλRµ
z(2π/n), λ=0,1
µ=0,1,...,n−1,
whereRz(2π/n)representsarotationof2 π/naboutthe n-foldsymmetryaxis,whereas
Srepresentsarotationof πaboutanaxisthroughthecenteroftheregularpolygonand
oneof itsvertices.
ForS=Eshowthatthisform maydescribethematrices A,B,C, andDofD3.
Note. The elements RzandSare called the generators of this finite group. Similarly,
iis thegeneratorofthegroupgivenbyEq.(4.164).
4.7.17 Show that the cyclic group of nobjects,Cn, may be represented by rm,m=
0,1,2,...,n−1.Hereris ageneratorgivenby
r=exp(2πis/n).
Theparameter stakesonthevalues s=1,2,3,...,n,eachvalueof syieldingadiffer-
entone-dimensional(irreducible)representationof Cn.
4.7.18 Developtheirreducible2 ×2matrixrepresentationofthegroupofoperations(rotations
andreflections)thattransformasquareintoitself. Givethegroupmultiplicationtable.
Note. This is the symmetry group of a square and also the dihedral group D4. (See
Fig.4.16.)
4.7 Discrete Groups 303
FIGURE 4.16
Square.
FIGURE 4.17Hexagon.
4.7.19 Thepermutationgroupoffourobjectscontains4 !=24elements.FromExercise4.7.18,
D4, the symmetry group for a square, has far fewer than 24 elements. Explain the rela-
tionbetween D4andthepermutationgroupof fourobjects.
4.7.20 Aplaneis coveredwithregularhexagons,asshowninFig.4.17.
(a) Determinethedihedralsymmetryofanaxisperpendiculartotheplanethroughthe
common vertex of three hexagons (A). That is, if the axis has n-fold symmetry,
show (with careful explanation) what nis. Write out the 2 ×2 matrix describing
theminimum(nonzero)positiverotationofthearrayofhexagonsthatisamember
ofyourDngroup.
(b) Repeatpart(a)foranaxisperpendiculartotheplanethroughthegeometriccenter
ofonehexagon (B).
4.7.21 In a simple cubic crystal, we might have identical atoms at r=(la,ma,na) , withl,m,
andntakingonallintegralvalues.
(a) ShowthateachCartesianaxisisafourfold symmetryaxis.
(b) Thecubicgroupwillconsistofalloperations(rotations,reflections,inversion)that
leave the simple cubic crystal invariant. From a consideration of the permutation
304 Chapter 4 Group Theory
FIGURE 4.18
Multiplicationtable.
ofthepositiveandnegativecoordinateaxes,predicthowmanyelementsthiscubic
groupwillcontain.
4.7.22 (a) Fromthe D3multiplicationtableofFig.4.18constructasimilaritytransformtable
showingxyx−1, wherexandyeachrangeoverallsixelementsof D3:
(b) Divide the elements of D3into classes. Using the 2 ×2 matrix representation of
Eqs. (4.169)–(4.172)notethetrace(character)ofeachclass.
4.8 D IFFERENTIAL FORMS
InChapters1and2weadoptedtheviewthat,in ndimensions,avectorisan n-tupleofreal
numbers and that its components transform properly under changes of the coordinates. In
thissectionwestartfromthealternativeview,inwhichavectoristhoughtofasadirected
line segment, an arrow. The point of the idea is this: Although the concept of a vector as
a line segment does not generalize to curved space–time (manifolds of differential geom-
etry), except by working in the flat tangent space requiring embedding in auxiliary extra
dimensions, Elie Cartan’s differential forms are natural in curved space–time and a very
powerful tool. Calculus can be based on differential forms, as Edwards has shown by his
classic textbook (see the Additional Readings). Cartan’s calculus leads to a remarkable
unification of concepts and theorems of vector analysis that is worth pursuing. In differ-
ential geometry and advanced analysis (on manifolds) the use of differential forms is now
widespread.
Cartan’s notion of vector is based on the one-to-one correspondence between the linear
spaces of displacement vectors and directional differential operators (components of the
gradient form a basis). A crucial advantage of the latter is that they can be generalized to
curved space–time. Moreover, describing vectors in terms of directional derivatives along
curvesuniquelyspecifiesthevectoratagivenpointwithouttheneedtoinvokecoordinates.
Ultimately, since coordinates are needed to specify points, the Cartan formalism, though
anelegantmathematicaltoolfortheefficientderivationoftheoremsontensoranalysis,has
inprinciplenoadvantageoverthecomponentformalism.
1-Forms
We define dx,dy,dz in three-dimensional Euclidean space as functions assigning to a
directed line segment PQfrom the point Pto the point Qthe corresponding change in
x,y,z. The symbol dxrepresents “oriented length of the projection of a curve on the
4.8 Differential Forms 305
x-axis,” etc. Note that dx,dy,dz can be, but need not be, infinitesimally small, and they
must not be confused with the ordinary differentials that we associate with integrals
anddifferentialquotients.Afunctionof thetype
Adx+Bdy+Cdz, A,B,C realnumbers (4.175)
isdefinedasa constant1-form .
Example 4.8.1 CONSTANT 1-FORM
For a constant force F=(A,B,C) , the work done along the displacement from P=
(3,2,1)toQ=(4,5,6)is thereforegivenby
W=A(4−3)+B(5−2)+C(6−1)=A+3B+5C.
IfFis a force field, then its rectangular components A(x,y,z),B(x,y,z),C(x,y,z)
will depend on the location and the (nonconstant) 1 -formdW=F·drcorresponds to
the concept of work done against the force field F(r)alongdron a space curve. A finite
amountof work
W=integraldisplay
Cbracketleftbig
A(x,y,z)dx+B(x,y,z)dy+C(x,y,z)dzbracketrightbig
(4.176)
involves the familiar line integral along an oriented curve C, where the 1-form dWde-
scribestheamountofworkforsmalldisplacements(segmentsonthepath C).Inthislight,
the integrand f(x)dxof an integralintegraltextb
af(x)dxconsisting of the function fand of the
measure dxas the oriented length is here considered to be a 1-form. The value of the
integralisobtainedfromtheordinarylineintegral. /squaresolid
2-Forms
Consideraunitflowofmassinthe z-direction,thatis,aflowinthedirectionofincreasing
zso that a unit mass crosses a unit square of the xy-plane in unit time. The orientation
symbolizedbythesequenceofpointsinFig.4.19,
(0,0,0)→(1,0,0)→(1,1,0)→(0,1,0)→(0,0,0),
will be called counterclockwise , as usual. A unit flow in the z-direction is defined by the
functiondxdy31assigningtoorientedrectanglesinspacetheorientedareaoftheirprojec-
tions on the xy-plane. Similarly, a unit flow in the x-direction is described by dydzand a
unit flow in the y-direction by dzdx. The reverse order, dzdx, is dictated by the orienta-
tion convention, and dzdx=−dxdzby definition. This antisymmetry is consistent with
the cross product of two vectors representing oriented areas in Euclidean space. This no-
tion generalizes to polygons and curved differentiable surfaces approximated by polygons
andvolumes.
31Many authors denote this wedge product as dx∧dywithdy∧dx=−dx∧dy. Note that the product dxdy=dydxfor
ordinary differentials.
306 Chapter 4 Group Theory
FIGURE 4.19
Counterclockwise-oriented
rectangle.
Example 4.8.2 MAGNETIC FLUXACROSS AN ORIENTED SURFACE
IfB=(A,B,C) isaconstantmagneticinduction,thentheconstant 2-form
Adydz+Bdzdx+Cdxdy
describesthemagneticflux across anorientedrectangle.If Bis a magneticinductionfield
varyingacross asurface S,thentheflux
/Phi1=integraldisplay
Sbracketleftbig
Bx(r)dydz+By(r)dzdx+Bz(r)dxdybracketrightbig
(4.177)
acrosstheorientedsurface Sinvolvesthefamiliar(Riemann)integrationoverapproximat-
ingsmallorientedrectanglesfrom which Sis piecedtogether. /squaresolid
Thedefinitionofintegraltext
ωreliesondecomposing ω=summationtext
iωi,wherethedifferentialforms ωi
areeachnonzeroonlyinasmallpatchofthesurface Sthatcoversthesurface.Thenitcan
be shown thatsummationtext
iintegraltext
ωiconverges, as the patches become smaller and more numerous, to
the limitintegraltext
ω, which is independent of these decompositions. For more details and proofs,
werefer thereadertoEdwardsintheAdditionalReadings.
3-Forms
A 3-form dxdydz represents an oriented volume. For example, the determinant of three
vectors in Euclidean space changes sign if we reverse the order of two vectors. The
determinant measures the oriented volume spanned by the three vectors. In particular,integraltext
Vρ(x,y,z)dxdydz represents the total charge inside the volume Vifρis the charge
density. Higher-dimensional differential forms in higher-dimensional spaces are defined
similarlyandarecalled k-forms,with k=0,1,2,....
If a 3-form
ω=A(x1,x2,x3)dx1dx2dx3=A′(x′
1,x′
2,x′
3)dx′
1dx′
2dx′
3(4.178)
4.8 Differential Forms 307
ona 3-dimensionalmanifoldisexpressedintermsofnewcoordinates,thenthereisaone-
to-one,differentiablemap x′
i=x′
i(x1,x2,x3)betweenthesecoordinateswithJacobian
J=∂(x′
1,x′
2,x′
3)
∂(x1,x2,x3)=1,
andA=A′J=A′so thatintegraldisplay
Vω=integraldisplay
VAdx1dx2dx3=integraldisplay
V′A′dx′
1dx′
2dx′
3. (4.179)
This statement spells out the parameter independence of integrals over differential forms,
since parameterizations are essentially arbitrary. The rules governing integration of differ-
ential forms are defined on manifolds. These are continuous if we can move continuously
(actually we assume them differentiable) from point to point, oriented if the orientation of
curvesgeneralizestosurfacesandvolumesuptothedimensionofthewholemanifold.The
rulesondifferentialforms are:
•Ifω=aω1+a′ω′
1,witha,a′realnumbers,thenintegraltext
Sω=aintegraltext
Sω1+a′integraltext
Sω′
1,whereSis
acompact,oriented,continuousmanifoldwithboundary.
•If theorientationis reversed,thentheintegralintegraltext
Sωchangessign.
Exterior Derivative
Wenowintroducethe exteriorderivative dof afunction f, a0-form:
df≡∂f
∂xdx+∂f
∂ydy+∂f
∂zdz=∂f
∂xidxi, (4.180)
generating a 1-form ω1=df, the differential of f(or exterior derivative), the gradient
in standard vector analysis. Upon summing over the coordinates, we have used and will
continue to use Einstein’s summation convention. Applying the exterior derivative dto a
1-formwedefine
d(Adx+Bdy+Cdz)=dAdx+dBdy+dCdz (4.181)
with functions A,B,C. This definition in conjunction with dfas just given ties vectors
to differential operators ∂i=∂
∂xi. Similarly, we extend dtok-forms. However, applying
dtwicegiveszero, ddf=0,because
d(df)=d∂f
∂xdx+d∂f
∂ydy
=parenleftbigg∂2f
∂x2dx+∂2f
∂x∂ydyparenrightbigg
dx+parenleftbigg∂2f
∂y∂xdx+∂2f
∂y2dyparenrightbigg
dy
=parenleftbigg∂2f
∂y∂x−∂2f
∂x∂yparenrightbigg
dxdy=0. (4.182)
This follows from the fact that in mixed partial derivatives their order does not matter
provided all functions are sufficiently differentiable. Similarly we can show ddω1=0f o r
a 1-form ω1,et c.
308 Chapter 4 Group Theory
Therulesgoverningdifferentialforms,with ωkdenotinga k-form,thatwehaveusedso
far are
•dxdx=0=dydy=dzdz,dx2
i=0;
•dxdy=−dydx,dxidxj=−dxjdxi,i/negationslash=j,
•dx1dx2···dxkis totallyantisymmetricinthe dxi,i=1,2,...,k.
•df=∂f
∂xidxi;
•d(ωk+/Omega1k)=dωk+d/Omega1k, linearity;
•ddωk=0.
Now we apply the exterior derivative dto products of differential forms, starting with
functions(0-forms). Wehave
d(fg)=∂(fg)
∂xidxi=parenleftbigg
f∂g
∂xi+∂f
∂xigparenrightbigg
dxi=fd g+dfg. (4.183)
Ifω1=∂g
∂xidxiisa1-form and fisafunction,then
d(fω1)=dparenleftbigg
f∂g
∂xidxiparenrightbigg
=dparenleftbigg
f∂g
∂xiparenrightbigg
dxi
=∂parenleftbig
f∂g
∂xiparenrightbig
∂xjdxjdxi=parenleftbigg∂f
∂xj∂g
∂xi+f∂2g
∂xi∂xjparenrightbigg
dxjdxi
=dfω1+fdω1, (4.184)
asexpected.Butif ω′
1=∂f
∂xjdxjisanother1-form,then
d(ω1ω′
1)=dparenleftbigg∂g
∂xidxi∂f
∂xjdxjparenrightbigg
=dparenleftbigg∂g
∂xi∂f
∂xjparenrightbigg
dxidxj
=∂parenleftBig
∂g
∂xi∂f
∂xjparenrightBig
∂xkdxkdxidxj
=∂2g
∂xi∂xkdxkdxi∂f
∂xjdxj−∂g
∂xidxi∂2f
∂xj∂xkdxkdxj
=dω1ω′
1−ω1dω′
1. (4.185)
This proof is valid for more general 1-forms ω=fidxiwith functions fi. In general,
therefore,wedefinefor k-forms:
d(ωkω′
k)=(dωk)ω′
k+(−1)kωk(dω′
k). (4.186)
Ingeneral,theexteriorderivativeofa k-form isa (k+1)-form.
4.8 Differential Forms 309
Example 4.8.3 POTENTIAL ENERGY
Asanapplicationintwodimensions(forsimplicity),considerthepotential V(r),a0-form,
anddV, its exteriorderivative.Integrating Valonganorientedpath Cfromr1tor2gives
V(r2)−V(r1)=integraldisplay
CdV=integraldisplay
Cparenleftbigg∂V
∂xdx+∂V
∂ydyparenrightbigg
=integraldisplay
C∇V·dr, (4.187)
wherethelastintegralisthestandardformulaforthepotentialenergydifferencethatforms
part of the energy conservation theorem. The path and parameterization independence are
manifestinthisspecialcase. /squaresolid
Pullbacks
If alinearmap L2from theuv-planetothe xy-planehas theform
x=au+bv+c, y=eu+fv+g, (4.188)
oriented polygons in the uv-plane are mapped onto similar polygons in the xy-plane, pro-
videdthedeterminant af−beofthemap L2isnonzero.The2-form
dxdy=(adu+bdv)(edu+fdv )=(af−be)dudv (4.189)
can be pulled back from the xy-t ot h euv-plane. That is to say, an integral over a simply
connectedsurface Sbecomesintegraldisplay
L2(S)dxdy=(af−be)integraldisplay
Sdudv, (4.190)
and(af−be)dudv isthepullbackof dxdy,oppositetothedirectionofthemap L2from
theuv-plane to the xy-plane. Of course, the determinant af−beof the map L2is simply
theJacobian,generatedwithouteffort bythedifferentialforms inEq.(4.189).
Similarly,alinearmap L3fromtheu1u2u3-spacetothe x1x2x3-space
xi=aijuj+bi,i=1,2,3, (4.191)
automaticallygeneratesits Jacobianfromthe3-form
dx1dx2dx3=parenleftbigg3summationdisplay
j=1a1jdujparenrightbiggparenleftbigg3summationdisplay
j=1a2jdujparenrightbiggparenleftbigg3summationdisplay
j=1a3jdujparenrightbigg
=(a11a22a33−a12a21a33±···)du1du2du3
=det
a11a12a13
a21a22a23
a31a32a33
du1du2du3. (4.192)
Thus,differentialforms generatetherules governingdeterminants.
Given two linear maps in a row, it is straightforward to prove that the pullback under
a composed map is the pullback of the pullback. This theorem is the differential-forms
analogofmatrixmultiplication.
310 Chapter 4 Group Theory
Let us now consider a curve Cdefined by a parameter tin contrast to a curve defined
by an equation. For example, the circle {(cost,sint);0≤t≤2π}is a parameterization
byt, whereas the circle {(x,y);x2+y2=1}is a definition by an equation. Then the line
integral
integraldisplay
Cbracketleftbig
A(x,y)dx+B(x,y)dybracketrightbig
=integraldisplaytf
tibracketleftbigg
Adx
dt+Bdy
dtbracketrightbigg
dt (4.193)
forcontinuousfunctions A,B,dx/dt,dy/dt becomesaone-dimensionalintegraloverthe
oriented interval ti≤t≤tf. Clearly, the 1-form [Adx
dt+Bdy
dt]dton thet-line is obtained
from the 1-form Adx+Bdyon thexy-plane via the map x=x(t),y=y(t)from the
t-linetothecurve Cinthexy-plane.The1-form [Adx
dt+Bdy
dt]dtiscalledthepullbackof
the 1-form Adx+Bdyunder the map x=x(t),y=y(t). Using pullbacks we can show
thatintegralsover1-formsare independentof theparameterizationof thepath.
Inthissense,thedifferentialquotientdy
dxcanbeconsideredasthecoefficientof dxinthe
pullback of dyunder the function y=f(x),o rdy=f′(x)dx. This concept of pullback
readily generalizes to maps in three or more dimensions and to k-forms with k>1. In
particular,thechainrulecanbeseentobeapullback:If
yi=fi(x1,x2,...,xn), i=1,2,...,land
zj=gj(y1,y2,...,yl), j=1,2,...,m (4.194)
are differentiable maps from Rn→RlandRl→Rm, then the composed map Rn→Rm
is differentiable and the pullback of any k-form under the composed map is equal to the
pullback of the pullback. This theorem is useful for establishing that integrals of k-forms
areparameterindependent.
Similarly,wedefinethedifferential dfasthepullbackofthe1-form dzunderthefunc-
tionz=f(x,y):
dz=df=∂f
∂xdx+∂f
∂ydy. (4.195)
Example 4.8.4 STOKES ’THEOREM
As another application let us first sketch the standard derivation of the simplest version
of Stokes’ theorem for a rectangle S=[a≤x≤b,c≤y≤d]oriented counterclockwise,
with∂Sits boundary
integraldisplay
∂S(Adx+Bdy)=integraldisplayb
aA(x,c)dx+integraldisplayd
cB(b,y)dy+integraldisplaya
bA(x,d)dx+integraldisplayc
dB(a,y)dy
=integraldisplayd
cbracketleftbig
B(b,y)−B(a,y)bracketrightbig
dy−integraldisplayb
abracketleftbig
A(x,d)−A(x,c)bracketrightbig
dx
=integraldisplayd
cintegraldisplayb
a∂B
∂xdxdy−integraldisplayb
aintegraldisplayd
c∂A
∂ydydx
=integraldisplay
Sparenleftbigg∂B
∂x−∂A
∂yparenrightbigg
dxdy, (4.196)
4.8 Differential Forms 311
whichholdsforanysimplyconnectedsurface Sthatcanbepiecedtogetherbyrectangles.
Now we demonstrate the use of differential forms to obtain the same theorem (again in
twodimensionsforsimplicity):
d(Adx+Bdy)=dAdx+dBdy
=parenleftbigg∂A
∂xdx+∂A
∂ydyparenrightbigg
dx+parenleftbigg∂B
∂xdx+∂B
∂ydyparenrightbigg
dy=parenleftbigg∂B
∂x−∂A
∂yparenrightbigg
dxdy,
(4.197)
using the rules highlighted earlier. Integrating over a surface Sand its boundary ∂S,r e -
spectively,weobtain
integraldisplay
∂S(Adx+Bdy)=integraldisplay
Sd(Adx+Bdy)=integraldisplay
Sparenleftbigg∂B
∂x−∂A
∂yparenrightbigg
dxdy. (4.198)
Here contributions to the left-hand integral from inner boundaries cancel as usual because
they are oriented in opposite directions on adjacent rectangles. For each oriented inner
rectanglethatmakesupthesimplyconnectedsurface Swehaveused,
integraldisplay
Rddx=integraldisplay
∂Rdx=0. (4.199)
Notethattheexteriorderivativeautomaticallygeneratesthe zcomponentof thecurl.
In threedimensions,Stokes’theoremderivesfrom thedifferential-form identityinvolv-
ingthevectorpotential Aandmagneticinduction B=∇×A,
d(Axdx+Aydy+Azdz)=dAxdx+dAydy+dAzdz
=parenleftbigg∂Ax
∂xdx+∂Ax
∂ydy+∂Ax
∂zdzparenrightbigg
dx+···
=parenleftbigg∂Az
∂y−∂Ay
∂zparenrightbigg
dydz+parenleftbigg∂Ax
∂z−∂Az
∂xparenrightbigg
dzdx+parenleftbigg∂Ay
∂x−∂Ax
∂yparenrightbigg
dxdy,
(4.200)
generatingallcomponentsofthecurlinthree-dimensionalspace.Thisidentityisintegrated
over each oriented rectangle that makes up the simply connected surface S(which has no
holes, that is, where every curve contracts to a point of the surface) and then is summed
overalladjacentrectanglestoyieldthemagneticfluxacross S,
/Phi1=integraldisplay
S[Bxdydz+Bydzdx+Bzdxdy]
=integraldisplay
∂S[Axdx+Aydy+Azdz], (4.201)
or,inthestandardnotationof vectoranalysis(Stokes’theorem,Chapter1),
integraldisplay
SB·da=integraldisplay
S(∇×A)·da=integraldisplay
∂SA·dr. (4.202)
/squaresolid
312 Chapter 4 Group Theory
Example 4.8.5 GAUSS’THEOREM
Consider Gauss’ law, Section 1.14. We integrate the electric density ρ=1
ε0∇·Eover
the volume of a single parallelepiped V=[a≤x≤b,c≤y≤d,e≤z≤f]oriented by
dxdydz (right-handed), the side x=bofVis oriented by dydz(counterclockwise, as
seenfrom x>b),andsoon.Using
Ex(b,y,z)−Ex(a,y,z)=integraldisplayb
a∂Ex
∂xdx, (4.203)
we have, in the notation of differential forms, summing over all adjacent parallelepipeds
thatmakeupthevolume V,
integraldisplay
∂VExdydz=integraldisplay
V∂Ex
∂xdxdydz. (4.204)
Integratingtheelectricflux(2-form) identity
d(Exdydz+Eydzdx+Ezdxdy)=dExdydz+dEydzdx+dEzdxdy
=parenleftbigg∂Ex
∂x+∂Ey
∂y+∂Ez
∂zparenrightbigg
dxdydz (4.205)
acrossthesimplyconnectedsurface ∂VwehaveGauss’theorem,
integraldisplay
∂V(Exdydz+Eydzdx+Ezdxdy)=integraldisplay
Vparenleftbigg∂Ex
∂x+∂Ey
∂y+∂Ez
∂zparenrightbigg
dxdydz, (4.206)
or,instandardnotationofvectoranalysis,
integraldisplay
∂VE·da=integraldisplay
V∇·Ed3r=q
ε0. (4.207)
/squaresolid
Theseexamplesaredifferentcasesofasingletheoremondifferentialforms.Toexplain
why, let us begin with some terminology, a preliminary definition of a differentiable
manifold M : It is a collectionof points ( m-tuples of real numbers) that are smoothly(that
is, differentiably) connected with each other so that the neighborhood of each point looks
likeasimplyconnectedpieceofan m-dimensionalCartesianspace“closeenough”around
thepointandcontainingit.Here, m,whichstaysconstantfrompointtopoint,iscalledthe
dimension of the manifold. Examples are the m-dimensional Euclidean space Rmand the
m-dimensionalsphere
Sm=bracketleftbiggparenleftbig
x1,...,xm+1parenrightbig
;m+1summationdisplay
i=1parenleftbig
xiparenrightbig2=1bracketrightbigg
.
Any surface with sharp edges, corners, or kinks is not a manifold in our sense, that
is, is not differentiable. In differential geometry, all movements, such as translation and
paralleldisplacement,arelocal,thatis,aredefinedinfinitesimally.Ifweapplytheexterior
derivative dtoafunction f(x1,...,xm)onM, wegeneratebasic 1-forms:
df=∂f
∂xidxi, (4.208)
4.8 Differential Forms 313
wherexi(P)arecoordinatefunctions.Asbeforewehave d(df)=0 because
d(df)=dparenleftbigg∂f
∂xiparenrightbigg
dxi=∂2f
∂xj∂xidxjdxi
=summationdisplay
j<iparenleftbigg∂2f
∂xj∂xi−∂2f
∂xi∂xjparenrightbigg
dxjdxi=0 (4.209)
because the order of derivatives does not matter. Any 1-form is a linear combination ω=summationtext
iωidxiwithfunctions ωi.
Generalized Stokes’ Theorem on Differential Forms
Letωbeacontinuous (k−1)-forminx1x2···xn-spacedefinedeverywhereonacompact,
oriented, differentiable k-dimensional manifold Swith boundary ∂Sinx1x2···xn-space.
Thenintegraldisplay
∂Sω=integraldisplay
Sdω. (4.210)
Here
dω=d(Adx 1dx2···dxk−1+···)=dAdx1dx2···dxk−1+···.(4.211)
ThepotentialenergyinExample4.8.3giventhistheoremforthepotential ω=V,a0-form;
Stokes’theoreminExample4.8.4isthistheoremforthevectorpotential1-formsummationtext
iAidxi
(forEuclideanspaces dxi=dxi);andGauss’theoreminExample4.8.5isStokes’theorem
fortheelectricflux2-forminthree-dimensionalEuclideanspace.
The method of integration by parts can be generalized to differential forms using
Eq.(4.186):
integraldisplay
Sdω1ω2=integraldisplay
∂Sω1ω2−(−1)k1integraldisplay
Sω1dω2. (4.212)
Thisis provedbyintegratingtheidentity
d(ω1ω2)=dω1ω2+(−1)k1ω1dω2, (4.213)
withtheintegratedtermintegraltext
Sd(ω1ω2)=integraltext
∂Sω1ω2.
Our next goal is to cast Sections 2.10 and 2.11 in the language of differential forms. So
far wehaveworkedintwo-or three-dimensionalEuclideanspace.
Example 4.8.6 RIEMANN MANIFOLD
Let us look at the curved Riemann space–time of Sections 2.10–2.11 and reformulate
some of this tensor analysis in curved spaces in the language of differential forms. Re-
callthatdishinguishingbetweenupperandlowerindicesisimportanthere.Themetric gij
inEq.(2.123) canbewritteninterms oftangentvectors,Eq.(2.114), asfollows:
gij=∂xl
∂qi∂xl
∂qj, (4.214)
314 Chapter 4 Group Theory
where the sum over the index ldenotes the inner product of the tangent vectors. (Here
we continue to use Einstein’s summation convention over repeated indices. As before, the
metric tensor is used to raise and lower indices.) The key concept of connection involves
theChristoffelsymbols,whichweaddressfirst.Theexteriorderivativeofatangentvector
canbeexpandedintermsof thebasis oftangentvectors(compareEq.(2.131a)),
dparenleftbigg∂xl
∂qiparenrightbigg
=Ŵkij∂xl
∂qkdqj, (4.215)
thusintroducingtheChristoffelsymbolsofthesecondkind.Applying dtoEq.(4.214)we
obtain
dgij=∂gij
∂qmdqm=dparenleftbigg∂xl
∂qiparenrightbigg∂xl
∂qj+∂xl
∂qidparenleftbigg∂xl
∂qjparenrightbigg
(4.216)
=parenleftbigg
Ŵkim∂xl
∂qk∂xl
∂qj+Ŵkjm∂xl
∂qi∂xl
∂qkparenrightbigg
dqm=parenleftbig
Ŵkimgkj+Ŵkjmgikparenrightbig
dqm.
Comparingthecoefficientsof dqmyields
∂gij
∂qm=Ŵkimgkj+Ŵkjmgik. (4.217)
UsingtheChristoffel symbolof thefirst kind,
[ij,m]=gkmŴkij, (4.218)
wecanrewriteEq. (4.217)as
∂gij
∂qm=[im,j]+[jm,i], (4.219)
whichcorrespondstoEq.(2.136)andimpliesEq. (2.137). Wecheckthat
[ij,m]=1
2parenleftbigg∂gim
∂qj+∂gjm
∂qi−∂gij
∂qmparenrightbigg
(4.220)
istheuniquesolutionof Eq.(4.219)andthat
Ŵkij=gmk[ij,m]=1
2gmkparenleftbigg∂gim
∂qj+∂gjm
∂qi−∂gij
∂qmparenrightbigg
(4.221)
follows. /squaresolid
Hodge∗Operator
The differentials dxi,i=1,2,...,m, form a basis of a vector space that is (seen to be)
dualtothederivatives ∂i=∂
∂xi;theyarebasic1-forms.Forexample,thevectorspace V=
{(a1,a2,a3)}isdualtothevectorspaceofplanes(linearfunctions f)inthree-dimensional
Euclideanspace V∗={f≡a1x1+a2x2+a3x3−d=0}.Thegradient
∇f=parenleftbigg∂f
∂x1,∂f
∂x2,∂f
∂x3parenrightbigg
=(a1,a2,a3) (4.222)
4.8 Differential Forms 315
providesaone-to-one,differentiablemapfrom V∗toV.Suchdualrelationshipsaregener-
alizedbytheHodge∗operator,basedontheLevi-CivitasymbolofSection2.9.
Let theunitvectors ˆxibeanorientedorthonormalbasis ofthree-dimensionalEuclidean
space.ThentheHodge∗of scalarsis definedbythebasiselement
∗1≡1
3!εijkˆxiˆxjˆxk=ˆx1ˆx2ˆx3, (4.223)
which corresponds to (ˆx1׈x2)·ˆx3in standard vector notation. Here ˆxiˆxjˆxkis the totally
antisymmetric exterior product of the unit vectors that corresponds to (ˆxi׈xj)·ˆxkin
standardvectornotation.Forvectors,∗is definedfor thebasis ofunitvectorsas
∗ˆxi≡1
2!εijkˆxjˆxk. (4.224)
Inparticular,
∗ˆx1=ˆx2ˆx3,∗ˆx2=ˆx3ˆx1,∗ˆx3=ˆx1ˆx2. (4.225)
Fororientedareas,∗is definedonbasisareaelementsas
∗(ˆxiˆxj)≡εkijˆxk, (4.226)
so
∗(ˆx1ˆx2)=ε312ˆx3=ˆx3,∗(ˆx1ˆx3)=ε213ˆx2=−ˆx2,
∗(ˆx2ˆx3)=ε123ˆx1=ˆx1. (4.227)
Forvolumes,∗is definedas
∗(ˆx1ˆx2ˆx3)≡ε123=1. (4.228)
Example 4.8.7 CROSS PRODUCT OF VECTORS
Theexteriorproductof twovectors
a=3summationdisplay
i=1aiˆxi,b=3summationdisplay
i=1biˆxi (4.229)
isgivenby
ab=parenleftbigg3summationdisplay
i=1aiˆxiparenrightbiggparenleftbigg3summationdisplay
j=1biˆxjparenrightbigg
=summationdisplay
i<jparenleftbig
aibj−ajbiparenrightbigˆxiˆxj, (4.230)
whereasEq. (4.224)impliesthat
∗(ab)=a×b. (4.231)
/squaresolid
Next, let us analyze Sections 2.1–2.2 on curvilinear coordinates in the language of dif-
ferentialforms.
316 Chapter 4 Group Theory
Example 4.8.8 LAPLACIAN IN ORTHOGONAL COORDINATES
Considerorthogonalcoordinateswherethemetric(Eq. (2.5)) leadstolengthelements
dsi=hidqi,notsummed . (4.232)
Herethedqiareordinarydifferentials.The 1-formsassociatedwiththedirections ˆqiare
εi=hidqi,notsummed . (4.233)
Thenthegradientis definedbythe 1-form
df=∂f
∂qidqi=parenleftbigg1
hi∂f
∂qiparenrightbigg
εi. (4.234)
Weapplythehodgestar operatorto df, generatingthe 2-form
∗df=parenleftbigg1
hi∂f
∂qiparenrightbigg
∗εi=parenleftbigg1
h1∂f
∂q1parenrightbigg
ε2ε3+parenleftbigg1
h2∂f
∂q2parenrightbigg
ε3ε1+parenleftbigg1
h3∂f
∂q3parenrightbigg
ε1ε2
=parenleftbiggh2h3
h1∂f
∂q1parenrightbigg
dq2dq3+parenleftbiggh1h3
h2∂f
∂q2parenrightbigg
dq3dq1+parenleftbiggh1h2
h3∂f
∂q3parenrightbigg
dq1dq2.
(4.235)
Applyinganotherexteriorderivative d, wegettheLaplacian
d(∗df)=∂
∂q1parenleftbiggh2h3
h1∂f
∂q1parenrightbigg
dq1dq2dq3+∂
∂q2parenleftbiggh1h3
h2∂f
∂q2parenrightbigg
dq2dq1dq2dq3
+∂
∂q3parenleftbiggh1h2
h3∂f
∂q3parenrightbigg
dq3dq1dq2dq3
=1
h1h2h3bracketleftbigg∂
∂q1parenleftbiggh2h3
h1∂f
∂q1parenrightbigg
+∂
∂q2parenleftbiggh1h3
h2∂f
∂q2parenrightbigg
+∂
∂q3parenleftbiggh1h2
h3∂f
∂q3parenrightbiggbracketrightbigg
·ε1ε2ε3=∇2fdq1dq2dq3. (4.236)
DividingbythevolumeelementgivesEq.(2.22).Recallthatthevolumeelements dxdydz
andε1ε2ε3mustbeequalbecause εianddx,dy,dz areorthonormal1-formsandthemap
from thexyztotheqicoordinatesis one-to-one. /squaresolid
Example 4.8.9 MAXWELL ’SEQUATIONS
Wenowworkinfour-dimensionalMinkowskispace,thehomogeneous,flatspace–timeof
special relativity, to discuss classical electrodynamics in terms of differential forms. We
start by introducing the electromagnetic field 2-form (field tensor in standard relativistic
notation):
F=−Exdtdx−Eydtdy−Ezdtdz+Bxdydz+Bydzdx+Bzdxdy
=1
2Fµνdxµdxν, (4.237)
4.8 Differential Forms 317
which contains the electric 1-form E=Exdx+Eydy+Ezdzand the magnetic flux
2-form. Here, terms with 1-forms in opposite order have been combined. (For Eq. (4.237)
tobevalid,themagneticinductionisinunitsof c;thatis,Bi→cBi,withcthevelocityof
light; or we work in units where c=1. Also,Fis in units of 1 /ε0, the dielectric constant
of the vacuum. Moreover, the vector potential is defined as A0=ε0φ, with the nonstatic
electric potential φandA1=Ax
µ0c,...;see Section 4.6 for more details.) The field 2-form
Fencompasses Faraday’s induction law that a moving charge is acted on by magnetic
forces.
Applying the exterior derivative dgenerates Maxwell’s homogeneous equations auto-
maticallyfrom F:
dF=−parenleftbigg∂Ex
∂ydy+∂Ex
∂zdzparenrightbigg
dtdx−parenleftbigg∂Ey
∂xdx+∂Ey
∂zdzparenrightbigg
dtdy
−parenleftbigg∂Ez
∂xdx+∂Ez
∂ydyparenrightbigg
dtdz+parenleftbigg∂Bx
∂xdx+∂Bx
∂tdtparenrightbigg
dydz
+parenleftbigg∂By
∂tdt+∂By
∂ydyparenrightbigg
dzdx+parenleftbigg∂Bz
∂tdt+∂Bz
∂zdzparenrightbigg
dxdy
=parenleftbigg
−∂Ex
∂y+∂Ey
∂x+∂Bz
∂tparenrightbigg
dtdxdy+parenleftbigg
−∂Ex
∂z+∂Ez
∂x−∂By
∂tparenrightbigg
dtdxdz
+parenleftbigg
−∂Ey
∂z+∂Ez
∂y+∂Bx
∂tparenrightbigg
dtdydz=0 (4.238)
which,instandardnotationofvectoranalysis,takesthefamiliarvectorformofMaxwell’s
homogeneousequations,
∇×E+∂B
∂t=0. (4.239)
SincedF=0, that is, there is no driving term so that Fis closed, there must be a 1-form
ω=Aµdxµso thatF=dω.No w ,
dω=∂νAµdxνdxµ, (4.240)
which, in standard notation, leads to the conventional relativistic form of the electromag-
neticfieldtensor,
Fµν=∂µAν−∂νAµ. (4.241)
Maxwell’shomogeneousequations, dF=0,are thusequivalentto ∂νFµν=0.
In order to derive similarly the inhomogeneous Maxwell’s equations, we introduce the
dualelectromagneticfieldtensor
˜Fµν=1
2εµναβFαβ, (4.242)
and,intermsofdifferentialforms,
∗F=∗parenleftbig
Fµνdxµdxνparenrightbig
=Fµν∗parenleftbig
dxµdxνparenrightbig
=1
2Fµνεµναβdxαdxβ.(4.243)
318 Chapter 4 Group Theory
Applyingtheexteriorderivativeyields
d(∗F)=1
2εµναβ(∂γFµν)dxγdxαdxβ, (4.244)
theleft-handsideofMaxwell’sinhomogeneousequations,a3-form.Itsdrivingtermisthe
dualof theelectriccurrentdensity,a3-form:
∗J=Jαparenleftbig
∗dxαparenrightbig
=Jαεαµνλdxµdxνdxλ
=ρdxdydz−Jxdtdydz−Jydtdzdx−Jzdtdxdy. (4.245)
AltogetherMaxwell’sinhomogeneousequationstaketheelegantform
d(∗F)=∗J. (4.246)
/squaresolid
The differential-formframework has brought considerableunification to vector algebra
and to tensor analysis on manifolds more generally, such as uniting Stokes’ and Gauss’
theorems and providing an elegant reformulation of Maxwell’s equations and an efficient
derivationoftheLaplacianincurvedorthogonalcoordinates,amongothers.
Exercises
4.8.1 Evaluate the 1-form adx+2bdy+4cdzon the line segment PQ, withP=(3,5,7),
Q=(7,5,3).
4.8.2 Iftheforcefieldisconstantandmovingaparticlefromtheoriginto (3,0,0)requiresa
units of work, from (−1,−1,0)to(−1,1,0)takesbunits of work, and from (0,0,4)
to(0,0,5)cunitsofwork, findthe1-formof thework.
4.8.3 Evaluatetheflowdescribedbythe2-form dxdy+2dydz+3dzdxacrosstheoriented
trianglePQRwithcornersat
P=(3,1,4), Q=(−2,1,4), R=(1,4,1).
4.8.4 Arethepoints,inthisorder,
(0,1,1), (3,−1,−2), (4,2,−2), (−1,0,1)
coplanar,or dotheyformanorientedvolume(right-handedor left-handed)?
4.8.5 Write Oersted’slaw,
integraldisplay
∂SH·dr=integraldisplay
S∇×H·da∼I,
indifferentialform notation.
4.8.6 Describetheelectricfieldbythe1-form E1dx+E2dy+E3dzandthemagneticinduc-
tionbythe2-form B1dydz+B2dzdx+B3dxdy.ThenformulateFaraday’sinduction
lawintermsof theseforms.
4.8 Additional Readings 319
4.8.7 Evaluatethe1-form
xdy
x2+y2−ydx
x2+y2
ontheunitcircleabouttheoriginorientedcounterclockwise.
4.8.8 Findthepullbackof dxdzunderx=ucosv,y=u−v,z=usinv.
4.8.9 Find the pullback of the 2-form dydz+dzdx+dxdyunder the map x=sinθcosϕ,
y=sinθsinϕ,z=cosθ.
4.8.10 Parameterizethesurfaceobtainedbyrotatingthecircle (x−2)2+z2=1,y=0,about
thez-axis inacounterclockwiseorientation,as seenfromoutside.
4.8.11 A 1-form Adx+Bdyis defined as closedif∂A
∂y=∂B
∂x. It is called exactif there is a
functionfso that∂f
∂x=Aand∂f
∂y=B. Determine which of the following 1-forms are
closed,orexact,andfindthecorrespondingfunctions fforthosethatareexact:
ydx+xdy,ydx+xdy
x2+y2,bracketleftbig
ln(xy)+1bracketrightbig
dx+x
ydy,
−ydx
x2+y2+xdy
x2+y2,f(z)dzwithz=x+iy.
4.8.12 Show thatsummationtextn
i=1x2
i=a2defines a differentiable manifold of dimension D=n−1i f
a/negationslash=0 andD=0i fa=0.
4.8.13 Show that the set of orthogonal 2 ×2 matrices form a differentiable manifold, and
determineitsdimension.
4.8.14 Determine the value of the 2-form Adydz+Bdzdx+Cdxdyon a parallelogram
withsides a,b.
4.8.15 ProveLorentzinvarianceof Maxwell’sequationsinthelanguageofdifferentialforms.
AdditionalReadings
Buerger, M. J., Elementary Crystallography . New York: Wiley (1956). A comprehensive discussion of crystal
symmetries. Buerger develops all 32 point groups and all 230 space groups. Related books by this author in-
cludeContemporaryCrystallography .NewYork:McGraw-Hill(1970); CrystalStructureAnalysis .NewY ork:
Krieger (1979) (reprint, 1960); and Introduction to Crystal Geometry . New York: Krieger (1977) (reprint,
1971).
Burns,G.,andA.M.Glazer, SpaceGroupsforSolid-StateScientists .NewYork:AcademicPress(1978).Awell-
organized, readable treatment of groups and their application to the solid state.
de-Shalit, A., and I. Talmi, Nuclear Shell Model . New York: Academic Press (1963). We adopt the Condon–
Shortley phase conventions of this text.
Edmonds, A.R., Angular Momentumin Quantum Mechanics . Princeton, NJ:Princeton University Press (1957).
Edwards,H.M., AdvancedCalculus: ADifferential FormsApproach . Boston: Birkhäuser (1994).
Falicov, L. M., Group Theory and Its Physical Applications . Notes compiled by A. Luehrmann. Chicago: Uni-
versity of Chicago Press (1966). Group theory, with an emphasis on applications to crystal symmetries and
solid-state physics.
320 Chapter 4 Group Theory
Gell-Mann, M., and Y. Ne’eman, The Eightfold Way . New York: Benjamin (1965). A collection of reprints of
significant papers on SU(3) and the particles of high-energy physics. Several introductory sections by Gell-
MannandNe’emanare especiallyhelpful.
Greiner, W., and B. Müller, Quantum Mechanics Symmetries . Berlin: Springer (1989). We refer to this textbook
for more details and numerous exercises that areworked out in detail.
Hamermesh,M., GroupTheoryandItsApplicationtoPhysicalProblems .Reading,MA:Addison-Wesley(1962).
A detailed, rigorous account of both finite and continuous groups. The 32 point groups are developed. The
continuous groups are treated, with Lie algebra included. A wealth of applications to atomic and nuclear
physics.
Hassani,S., Foundations of Mathematical Physics .Boston: Allyn and Bacon (1991).
Heitler,W., TheQuantumTheoryofRadiation ,2nded.Oxford:OxfordUniversityPress(1947).Reprinted,New
York: Dover (1983).
Higman, B., Applied Group-Theoretic and Matrix Methods . Oxford: Clarendon Press (1955). A rather complete
and unusually intelligible development of matrix analysis and group theory.
Jackson, J. D., Classical Electrodynamics ,3rd ed.NewYork: Wiley (1998).
Messiah,A., Quantum Mechanics , Vol. II. Amsterdam: North-Holland (1961).
Panofsky,W.K.H.,andM.Phillips, ClassicalElectricityandMagnetism ,2nded.Reading,MA:Addison-Wesley
(1962). The Lorentz covariance of Maxwell’s equations is developed for both vacuum and material media.
Panofsky and Phillips use contravariant and covariant tensors.
Park,D.,ResourceletterSP-1onsymmetryinphysics. Am.J.Phys. 36:577–584(1968).Includesalargeselection
of basic references on group theory and its applications to physics: atoms, molecules, nuclei, solids, and
elementaryparticles.
Ram, B., Physics of the SU(3) symmetry model. Am. J. Phys. 35: 16 (1967). An excellent discussion of the
applications of SU(3) to the strongly interacting particles (baryons). For a sequel to this see R. D. Young,
Physics of the quark model. Am.J .Ph ys. 41: 472 (1973).
R o s e ,M .E . , Elementary Theory of Angular Momentum . New York: Wiley (1957). Reprinted. New York: Dover
(1995).Aspartofthedevelopmentofthequantumtheoryofangularmomentum,Roseincludesadetailedand
readable accountof the rotation group.
Wigner, E. P., Group Theory and Its Application to the Quantum Mechanics of Atomic Spectra (translated by
J.J.Griffin).NewYork:AcademicPress(1959).Thisistheclassicreferenceongrouptheoryforthephysicist.
Therotation group is treatedin considerable detail.Thereis awealthof applications to atomic physics.
CHAPTER 5
INFINITE SERIES
5.1 F UNDAMENTAL CONCEPTS
Infiniteseries,literallysummationsofaninfinitenumberofterms,occurfrequentlyinboth
pureandappliedmathematics.Theymaybeusedbythepuremathematiciantodefinefunc-
tions as a fundamental approach to the theory of functions, as well as for calculating ac-
curatevaluesoftranscendentalconstantsandtranscendentalfunctions.Inthemathematics
of science and engineering infinite series are ubiquitous, for they appear in the evaluation
of integrals (Sections 5.6 and 5.7), in the solution of differential equations (Sections 9.5
and 9.6), and as Fourier series (Chapter 14) and compete with integral representations for
thedescriptionofahostofspecialfunctions(Chapters11,12,and13).InSection16.3the
Neumann series solution for integral equations provides one more example of the occur-
renceanduseof infiniteseries.
Right at the start we face the problem of attaching meaning to the sum of an infinite
numberofterms.Theusualapproachisbypartialsums.Ifwehaveaninfinitesequenceof
termsu1,u2,u3,u4,u5,...,wedefinethe ithpartialsumas
si=isummationdisplay
n=1un. (5.1)
This is a finite summation and offers no difficulties. If the partial sums siconverge to a
(finite)limitas i→∞,
lim
i→∞si=S, (5.2)
the infinite seriessummationtext∞
n=1unis said to be convergent and to have the value S. Note that we
reasonably, plausibly, but still arbitrarily definethe infinite series as equal to Sand that a
necessary condition for this convergence to a limit is that lim n→∞un=0. This condition,
however, is not sufficient to guarantee convergence. Equation (5.2) is usually written in
formalmathematicalnotation:
321
322 Chapter 5 Infinite Series
The condition for the existence of a limit Sis that for each ε>0, there is a fixed
N=N(ε)suchthat
|S−si|<ε, fori>N.
This condition is often derived from the Cauchy criterion applied to the partial sums si.
TheCauchycriterion is:
Anecessaryandsufficientconditionthatasequence (si)convergeisthatforeach ε>0
there is a fixednumber Nsuch that
|sj−si|<ε, for alli,j >N.
This meansthat theindividual partial sumsmustcluster together aswemovefarout in
the sequence.
The Cauchy criterion may easily be extended to sequences of functions. We see it in this
form in Section 5.5 in the definition of uniform convergence and in Section 10.4 in the
development of Hilbert space. Our partial sums simay not converge to a single limit but
mayoscillate,asinthecase
∞summationdisplay
n=1un=1−1+1−1+1+···−(−1)n+···.
Clearly,si=1f o riodd butsi=0f o rieven. There is no convergence to a limit, and
series such as this one are labeled oscillatory . Whenever the sequence of partial sums
diverges (approaches ±∞), the infinite series is said to diverge. Often the term divergent
is extended to include oscillatory series as well. Because we evaluate the partial sums
by ordinary arithmetic, the convergent series, defined in terms of a limit of the partial
sums, assumes a position of supreme importance. Two examples may clarify the nature of
convergence or divergence of a series and will also serve as a basis for a further detailed
investigationinthenextsection.
Example 5.1.1 THEGEOMETRIC SERIES
Thegeometricalsequence,startingwith aandwitharatio r(=an+1/anindependentof n),
isgivenby
a+ar+ar2+ar3+···+arn−1+···.
Thenthpartialsumisgivenby1
sn=a1−rn
1−r. (5.3)
Takingthelimitas n→∞,
limn→∞sn=a
1−r,for|r|<1. (5.4)
1Multiply and divide sn=summationtextn−1
m=0armby 1−r.
5.1 Fundamental Concepts 323
Hence,bydefinition,theinfinitegeometricseries convergesfor |r|<1 andis givenby
∞summationdisplay
n=1arn−1=a
1−r. (5.5)
Ontheotherhand,if |r|≥1,thenecessarycondition un→0isnotsatisfiedandtheinfinite
seriesdiverges. /squaresolid
Example 5.1.2 THEHARMONIC SERIES
Asasecondandmoreinvolvedexample,weconsidertheharmonicseries
∞summationdisplay
n=11
n=1+1
2+1
3+1
4+···+1
n+···. (5.6)
We have the lim n→∞un=limn→∞1/n=0, but this is not sufficient to guaranteeconver-
gence.If wegrouptheterms(no changeinorder) as
1+1
2+parenleftbig1
3+1
4parenrightbig
+parenleftbig1
5+1
6+1
7+1
8parenrightbig
+parenleftbig1
9+···+1
16parenrightbig
+···, (5.7)
eachpairofparenthesesencloses ptermsof theform
1
p+1+1
p+2+···+1
p+p>p
2p=1
2. (5.8)
Formingpartialsumsbyaddingtheparentheticalgroupsonebyone,weobtain
s1=1,s 4>5
2,
s2=3
2,s5>6
2,···
s3>4
2,sn>n+1
2.(5.9)
The harmonic series considered in this way is certainly divergent.2An alternate and inde-
pendentdemonstrationofits divergenceappearsinSection5.2. /squaresolid
If theun>0are monotonically decreasing to zero, that is, un>un+1for alln,thensummationtext
nunis converging to Sif, and only if, sn−nunconverges to S.As the partial sums sn
convergeto S,thistheoremimpliesthat nun→0,forn→∞.
Toprovethis theorem,westartbyconcludingfrom 0 <un+1<unand
sn+1−(n+1)un+1=sn−nun+1=sn−nun+n(un−un+1)>sn−nun
thatsn−nunincreases as n→∞.As a consequence of sn−nun<sn≤S, sn−nun
convergestoavalue s≤S.Deletingthetailofpositiveterms ui−unfromi=ν+1t on,
2The(finite) harmonic seriesappearsin aninteresting noteon themaximum stabledisplacementofastackof coins.P.R.John-
son, The LeaningTowerof Lire. Am.J.Phys. 23: 240 (1955).
324 Chapter 5 Infinite Series
weinferfrom sn−nun>u0+(u1−un)+···+(uν−un)=sν−νunthatsn−nun≥sν
forn→∞.Hencealso s≥S,sos=Sandnun→0.
When this theorem is applied to the harmonic seriessummationtext
n1
nwithn1
n=1 it implies that
itdoesnotconverge;itdivergesto +∞.
Addition, Subtraction of Series
If we have two convergent seriessummationtext
nun→sandsummationtext
nvn→S,their sum and difference
willalsoconvergeto s±Sbecausetheirpartialsumssatisfy
vextendsinglevextendsinglesj±Sj−(si±Si)vextendsinglevextendsingle=vextendsinglevextendsinglesj−si±(Sj−Si)vextendsinglevextendsingle≤|sj−si|+|Sj−Si|<2ǫ
usingthetriangleinequality
|a|−|b|≤|a+b|≤|a|+|b|
fora=sj−si,b=Sj−Si.
A convergent seriessummationtext
nun→Smay be multiplied termwise by a real number a.The
newseries willconvergeto aSbecause
|asj−asi|=vextendsinglevextendsinglea(sj−si)vextendsinglevextendsingle=|a||sj−si|<|a|ǫ.
This multiplication by a constant can be generalized to a multiplication by terms cnof a
boundedsequenceof numbers.
Ifsummationtext
nunconverges to Sand0<cn≤Mare bounded, thensummationtext
nuncnis convergent. Ifsummationtext
nunisdivergentand cn>M>0,thensummationtext
nuncndiverges.
Toprovethis theorem wetakei,jsufficientlylargesothat |sj−si|<ǫ.Then
jsummationdisplay
i+1uncn≤Mjsummationdisplay
i+1un=M|sj−si|<Mǫ.
Thedivergentcasefollowsfrom
summationdisplay
nuncn>Msummationdisplay
nun→∞.
Usingthebinomialtheorem3(Section5.6), wemayexpandthefunction (1+x)−1:
1
1+x=1−x+x2−x3+···+(−x)n−1+···. (5.10)
If weletx→1,thisseriesbecomes
1−1+1−1+1−1+···, (5.11)
a series that we labeled oscillatory earlier in this section. Although it does not converge
in the usual sense, meaning can be attached to this series. Euler, for example, assigned a
value of 1 /2 to this oscillatory sequence on the basis of the correspondence between this
series and the well-defined function (1+x)−1. Unfortunately, such correspondence be-
tweenseries andfunctionis notunique,andthisapproachmustberefined.Othermethods
3ActuallyEq. (5.10) may be verified by multiplying both sides by 1 +x.
5.2 Convergence Tests 325
of assigning a meaning to a divergent or oscillatory series, methods of defining a sum,
have been developed. See G. H. Hardy, Divergent Series , Chelsea Publishing Co. 2nd ed.
(1992).Ingeneral,however,thisaspectofinfiniteseriesisofrelativelylittleinteresttothe
scientist or the engineer. An exception to this statement, the very important asymptotic or
semiconvergentseries, is consideredinSection 5.10.
Exercises
5.1.1 Showthat
∞summationdisplay
n=11
(2n−1)(2n+1)=1
2.
Hint.Show(bymathematicalinduction)that sm=m/(2m+1).
5.1.2 Showthat
∞summationdisplay
n=11
n(n+1)=1.
Findthepartialsum smandverifyitscorrectnessbymathematicalinduction.
Note. The method of expansion in partial fractions, Section 15.8, offers an alternative
wayofsolvingExercises5.1.1and5.1.2.
5.2 C ONVERGENCE TESTS
Although nonconvergent series may be useful in certain special cases (compare Sec-
tion 5.10), we usually insist, as a matter of convenience if not necessity, that our series be
convergent.Itthereforebecomesamatterofextremeimportancetobeabletotellwhether
a given series is convergent. We shall develop a number of possible tests, starting with the
simple and relatively insensitive tests and working up to the more complicated but quite
sensitivetests.Forthepresentletusconsidera seriesofpositiveterms an≥0,postponing
negativetermsuntilthenextsection.
Comparison Test
If term by term a series of terms 0 ≤un≤an, in which the anform a convergent series,
the seriessummationtext
nunis also convergent. If un≤anfor alln, thensummationtext
nun≤summationtext
nanandsummationtext
nun
therefore is convergent . If term byterm a series of terms vn≥bn, in whichthe bn,f o r ma
divergent series, the seriessummationtext
nvnis alsodivergent . Note that comparisons of unwithbn
orvnwithanyield no information. If vn≥bnfor alln, thensummationtext
nvn≥summationtext
nbnandsummationtext
nvn
thereforeisdivergent.
Fortheconvergentseries anwealreadyhavethegeometricseries,whereastheharmonic
series will serve as the divergent comparison series bn. As other series are identified as
either convergent or divergent, they may be used for the known series in this comparison
test.Alltestsdevelopedinthissectionareessentiallycomparisontests.Figure5.1exhibits
thesetests andtheinterrelationships.
326 Chapter 5 Infinite Series
FIGURE 5.1Comparisontests.
Example 5.2.1 AD IRICHLET SERIES
Testsummationtext∞
n=1n−p,p=0.999,forconvergence.Since n−0.999>n−1andbn=n−1formsthe
divergentharmonicseries, thecomparisontestshowsthatsummationtext
nn−0.999is divergent.Gener-
alizing,summationtext
nn−pis seen to be divergent for all p≤1 but convergent for p>1 (see Exam-
ple5.2.3). /squaresolid
Cauchy Root Test
If(an)1/n≤r<1 for all sufficiently large n, withrindependent of n, thensummationtext
nanis
convergent.If (an)1/n≥1 for allsufficientlylarge n,thensummationtext
nanisdivergent.
The first part of this test is verified easily by raising (an)1/n≤rto thenth power. We
get
an≤rn<1.
Sincernis just the nth term in a convergent geometric series,summationtext
nanis convergent by the
comparison test. Conversely, if (an)1/n≥1, thenan≥1 and the series must diverge. This
roottestisparticularlyusefulinestablishingthepropertiesofpowerseries(Section5.7).
D’Alembert (or Cauchy) Ratio Test
Ifan+1/an≤r<1 for all sufficiently large nandris independent of n, thensummationtext
nanis
convergent.If an+1/an≥1 for allsufficientlylarge n,thensummationtext
nanisdivergent.
Convergenceisprovedbydirectcomparisonwiththegeometricseries (1+r+r2+···).
In the second part, an+1≥anand divergence should be reasonably obvious. Although not
5.2 Convergence Tests 327
quitesosensitiveastheCauchyroottest,thisD’Alembertratiotestisoneoftheeasiestto
applyandiswidelyused.Analternatestatementoftheratiotestisintheformofalimit:If
limn→∞an+1
an<1,convergence ,
>1,divergence , (5.12)
=1,indeterminate .
Becauseofthisfinalindeterminatepossibility,theratiotestislikelytofailatcrucialpoints,
and more delicate, sensitive tests are necessary. The alert reader may wonder how this
indeterminacy arose. Actually it was concealed in the first statement, an+1/an≤r<1.
We might encounter an+1/an<1 for allfinitenbut be unable to choose an r<1and
independentofn suchthat an+1/an≤rforallsufficientlylarge n.Anexampleisprovided
bytheharmonicseries
an+1
an=n
n+1<1. (5.13)
Since
limn→∞an+1
an=1, (5.14)
nofixedratio r<1 existsandtheratiotestfails.
Example 5.2.2 D’A LEMBERT RATIO TEST
Testsummationtext
nn/2nfor convergence.
an+1
an=(n+1)/2n+1
n/2n=1
2·n+1
n. (5.15)
Since
an+1
an≤3
4forn≥2, (5.16)
wehaveconvergence.Alternatively,
limn→∞an+1
an=1
2(5.17)
andagain—convergence. /squaresolid
Cauchy (or Maclaurin) Integral Test
This is another sort of comparison test, in which we compare a series with an integral.
Geometrically,wecomparetheareaofaseriesofunit-widthrectangleswiththeareaunder
acurve.
328 Chapter 5 Infinite Series
FIGURE 5.2(a)Comparisonofintegralandsum-blocksleading.
(b)Comparisonofintegralandsum-blockslagging.
Letf(x)be a continuous, monotonic decreasing function in which f(n)=an. Thensummationtext
nanconverges ifintegraltext∞
1f(x)dxis finite and diverges if the integral is infinite. For the ith
partialsum,
si=isummationdisplay
n=1an=isummationdisplay
n=1f(n). (5.18)
But
si>integraldisplayi+1
1f(x)dx (5.19)
from Fig.5.2a, f(x)beingmonotonicdecreasing.Ontheotherhand,fromFig. 5.2b,
si−a1<integraldisplayi
1f(x)dx, (5.20)
in which the series is represented by the inscribed rectangles. Taking the limit as i→∞,
wehave
integraldisplay∞
1f(x)dx≤∞summationdisplay
n=1an≤integraldisplay∞
1f(x)dx+a1. (5.21)
Hence the infinite series converges or diverges as the corresponding integral converges or
diverges. This integral test is particularly useful in setting upper and lower bounds on the
remainderofaseries aftersomenumberof initialtermshavebeensummed.Thatis,
∞summationdisplay
n=1an=Nsummationdisplay
n=1an+∞summationdisplay
n=N+1an,
where
integraldisplay∞
N+1f(x)dx≤∞summationdisplay
n=N+1an≤integraldisplay∞
N+1f(x)dx+aN+1.
5.2 Convergence Tests 329
Tofreetheintegraltestfromthequiterestrictiverequirementthattheinterpolatingfunc-
tionf(x)be positive and monotonic, we show for any function f(x)with a continuous
derivativethat
Nfsummationdisplay
n=Ni+1f(n)=integraldisplayNf
Nif(x)dx+integraldisplayNf
Niparenleftbig
x−[x]parenrightbig
f′(x)dx (5.22)
holds.Here[x]denotesthelargestintegerbelow x,sox−[x]variessawtoothlikebetween
0and1. ToderiveEq. (5.22)weobservethat
integraldisplayNf
Nixf′(x)dx=Nff(Nf)−Nif(Ni)−integraldisplayNf
Nif(x)dx, (5.23)
usingintegrationbyparts.Nextweevaluatetheintegral
integraldisplayNf
Ni[x]f′(x)dx=Nf−1summationdisplay
n=Ninintegraldisplayn+1
nf′(x)dx=Nf−1summationdisplay
n=Ninbraceleftbig
f(n+1)−f(n)bracerightbig
=−Nfsummationdisplay
n=Ni+1f(n)−Nif(Ni)+Nff(Nf). (5.24)
Subtracting Eq. (5.24) from (5.23) we arrive at Eq. (5.22). Note that f(x)m a yg ou po r
down and even change sign, so Eq. (5.22) applies to alternating series (see Section 5.3) as
well. Usually f′(x)falls faster than f(x)forx→∞, so the remainder term in Eq. (5.22)
convergesbetter.ItiseasytoimproveEq.(5.22)byreplacing x−[x]byx−[x]−1
2,which
variesbetween −1
2and1
2:
summationdisplay
Ni<n≤Nff(n)=integraldisplayNf
Nif(x)dx+integraldisplayNf
Niparenleftbig
x−[x]−1
2parenrightbig
f′(x)dx
+1
2braceleftbig
f(Nf)−f(Ni)bracerightbig
. (5.25)
Thenthe f′(x)-integralbecomesevensmaller,if f′(x)doesnotchangesigntoooften.For
anapplicationofthisintegraltesttoanalternatingseriesseeExample5.3.1.
Example 5.2.3 RIEMANN ZETAFUNCTION
TheRiemannzetafunctionis definedby
ζ(p)=∞summationdisplay
n=1n−p, (5.26)
providedtheseries converges.Wemaytake f(x)=x−p, andthen
integraldisplay∞
1x−pdx=x−p+1
−p+1vextendsinglevextendsinglevextendsinglevextendsingle∞
1,p/negationslash=1
=lnx|∞
x=1,p=1. (5.27)
330 Chapter 5 Infinite Series
The integral and therefore the series are divergent for p≤1, convergent for p>1. Hence
Eq. (5.26) should carry the condition p>1. This, incidentally, is an independent proof
that the harmonic series (p=1)diverges logarithmically. The sum of the first million
termssummationtext1,000,000n−1isonly14.392 726 .... /squaresolid
ThisintegralcomparisonmayalsobeusedtosetanupperlimittotheEuler–Mascheroni
constant,4definedby
γ=limn→∞parenleftbiggnsummationdisplay
m=1m−1−lnnparenrightbigg
. (5.28)
Returningtopartialsums, Eq. (5.20)yields
sn=nsummationdisplay
m=1m−1−lnn≤integraldisplayn
1dx
x−lnn+1. (5.29)
Evaluating the integral on the right, sn<1 for all nand therefore γ≤1. Exer-
cise 5.2.12 leads to more restrictive bounds. Actually the Euler–Mascheroni constant is
0.57721566 ....
Kummer’s Test
This is the first of three tests that are somewhat more difficult to apply than the preceding
tests. Their importance lies in their power and sensitivity. Frequently, at least one of the
threewillworkwhenthesimpler,easiertestsareindecisive.Itmustberemembered,how-
ever,thatthesetests,likethosepreviouslydiscussed,areultimatelybasedoncomparisons.
It can be shown that there is no most slowly converging series and no most slowly diverg-
ingseries.Thismeansthatallconvergencetestsgivenhere,includingKummer’s,mayfail
sometime.
We consider a series of positive terms uiand a sequence of finite positive constants ai.
If
anun
un+1−an+1≥C>0 (5.30)
foralln≥N, whereNissomefixednumber,5thensummationtext∞
i=1uiconverges .If
anun
un+1−an+1≤0 (5.31)
andsummationtext∞
i=1a−1
idiverges,thensummationtext∞
i=1uidiverges.
4This is the notation of National Bureau of Standards, Handbook of Mathematical Functions , Applied Mathematics Series-55
(AMS-55). NewYork: Dover (1972).
5Withumfinite, the partial sum sNwill always be finite for Nfinite. The convergence or divergence of a series depends on the
behavior of thelast infinity of terms, not on the first Nterms.
5.2 Convergence Tests 331
The proof of this powerful test is remarkably simple. From Eq. (5.30), with Csome
positiveconstant,
CuN+1≤aNuN−aN+1uN+1
CuN+2≤aN+1uN+1−aN+2uN+2
······························
Cun≤an−1un−1−anun.(5.32)
Addinganddividingby C(andrecallingthat C/negationslash=0),weobtain
nsummationdisplay
i=N+1ui≤aNuN
C−anun
C. (5.33)
Henceforthepartialsum sn,
sn≤Nsummationdisplay
i=1ui+aNuN
C−anun
C
<Nsummationdisplay
i=1ui+aNuN
C,aconstant,independentof n. (5.34)
Thepartialsumsthereforehaveanupperbound.Withzeroasanobviouslowerbound,the
seriessummationtextuimustconverge.
Divergenceis shownasfollows.FromEq. (5.31)for un+1>0,
anun≥an−1un−1≥···≥aNuN,n>N. (5.35)
Thus,for an>0,
un≥aNuN
an(5.36)
and
∞summationdisplay
i=N+1ui≥aNuN∞summationdisplay
i=N+1a−1
i. (5.37)
Ifsummationtext∞
i=1a−1
idiverges, then by the comparison testsummationtext
iuidiverges. Equations (5.30) and
(5.31)areoftengiveninalimitform:
limn→∞parenleftbigg
anun
un+1−an+1parenrightbigg
=C. (5.38)
ThusforC>0 wehaveconvergence,whereasfor C<0 (andsummationtext
ia−1
idivergent)wehave
divergence.ItisperhapsusefultoshowthecloserelationofEq.(5.38)andEqs.(5.30)and
(5.31)andtoshowwhyindeterminacycreepsinwhenthelimit C=0.Fromthedefinition
oflimit,
vextendsinglevextendsinglevextendsinglevextendsingleanun
un+1−an+1−Cvextendsinglevextendsinglevextendsinglevextendsingle<ε (5.39)
332 Chapter 5 Infinite Series
for alln≥Nand allε>0, no matter how small εmay be. When the absolute value signs
areremoved,
C−ε<anun
un+1−an+1<C+ε. (5.40)
Now, ifC>0, Eq. (5.30) follows from εsufficiently small. On the other hand, if C<0,
Eq.(5.31)follows.However,if C=0,thecenterterm, an(un/un+1)−an+1,maybeeither
positiveornegativeandtheprooffails.TheprimaryuseofKummer’stestistoproveother
tests, suchas Raabe’s(comparealsoExercise5.2.3).
If thepositiveconstants anof Kummer’stestarechosen an=n,wehaveRaabe’stest.
Raabe’s Test
Ifun>0 andif
nparenleftbiggun
un+1−1parenrightbigg
≥P>1 (5.41)
foralln≥N,whereNisapositiveintegerindependentof n,thensummationtext
iuiconverges .Here,
P=C+1 ofKummer’stest. If
nparenleftbiggun
un+1−1parenrightbigg
≤1, (5.42)
thensummationtext
iuidiverges(assummationtext
nn−1diverges). Thelimitformof Raabe’stest is
limn→∞nparenleftbiggun
un+1−1parenrightbigg
=P. (5.43)
We have convergence for P>1, divergence for P<1, and no conclusion for P=1,
exactlyaswiththeKummertest.ThisindeterminacyispointedupbyExercise5.2.4,which
presents a convergent series and a divergent series, with both series yielding P=1i n
Eq.(5.43).
Raabe’s test is more sensitive than the d’Alembert ratio test (Exercise 5.2.3) becausesummationtext∞
n=1n−1diverges more slowly thansummationtext∞
n=11. We obtain a more sensitive test (and one
thatis stillfairly easytoapply)bychoosing an=nlnn.This isGauss’ test.
Gauss’ Test
Ifun>0 for allfinite nand
un
un+1=1+h
n+B(n)
n2, (5.44)
inwhich B(n)isaboundedfunctionof nforn→∞,thensummationtext
iuiconvergesfor h>1 and
divergesfor h≤1:Thereis noindeterminatecasehere.
5.2 Convergence Tests 333
The Gauss test is an extremely sensitive test of series convergence. It will work for all
series the physicist is likely to encounter. For h>1o rh<1 the proof follows directly
fromRaabe’stest
limn→∞nbracketleftbigg
1+h
n+B(n)
n2−1bracketrightbigg
=limn→∞bracketleftbigg
h+B(n)
nbracketrightbigg
=h. (5.45)
Ifh=1, Raabe’s test fails. However, if we return to Kummer’s test and use an=nlnn,
Eq.(5.38) leadsto
limn→∞braceleftbigg
nlnnbracketleftbigg
1+1
n+B(n)
n2bracketrightbigg
−(n+1)ln(n+1)bracerightbigg
=limn→∞bracketleftbigg
nlnn·n+1
n−(n+1)ln(n+1)bracketrightbigg
=limn→∞(n+1)bracketleftbigg
lnn−lnn−lnparenleftbigg
1+1
nparenrightbiggbracketrightbigg
. (5.46)
Borrowinga resultfromSection5.6(whichis notdependentonGauss’test), wehave
limn→∞−(n+1)lnparenleftbigg
1+1
nparenrightbigg
=limn→∞−(n+1)parenleftbigg1
n−1
2n2+1
3n3···parenrightbigg
=−1<0. (5.47)
Hence we have divergence for h=1. This is an example of a successful application of
Kummer’stestwhenRaabe’stesthadfailed.
Example 5.2.4 LEGENDRE SERIES
TherecurrencerelationfortheseriessolutionofLegendre’sequation(Exercise9.5.5)may
beputintheform
a2j+2
a2j=2j(2j+1)−l(l+1)
(2j+1)(2j+2). (5.48)
Foruj=a2jandB(j)=O(1/j2)→0 (that is,|B(j)j2|≤C,C>0, a constant) as
j→∞inGauss’ testweapplyEq. (5.45).Then, for j≫l,6
uj
uj+1→(2j+1)(2j+2)
2j(2j+1)=2j+2
2j=1+1
j. (5.49)
ByEq. (5.44)theseries isdivergent. /squaresolid
6Theldependence enters B(j)but does not affect hin Eq. (5.45).
334 Chapter 5 Infinite Series
Improvement of Convergence
This section so far has been concerned with establishing convergence as an abstract math-
ematicalproperty.Inpractice,the rateofconvergencemaybeofconsiderableimportance.
Here we present one method of improving the rate of convergence of a convergent series.
OthertechniquesaregiveninSections5.4and5.9.
The basic principle of this method, due to Kummer, is to form a linear combination of
our slowly converging series and one or more series whose sum is known. For the known
seriesthecollection
α1=∞summationdisplay
n=11
n(n+1)=1
α2=∞summationdisplay
n=11
n(n+1)(n+2)=1
4
α3=∞summationdisplay
n=11
n(n+1)(n+2)(n+3)=1
18
.........
αp=∞summationdisplay
n=11
n(n+1)···(n+p)=1
p·p!
is particularly useful.7The series are combined term by term and the coefficients in the
linearcombinationchosentocancelthemostslowlyconvergingterms.
Example 5.2.5 RIEMANN ZETAFUNCTION ,ζ(3)
Letthe series tobe summedbesummationtext∞
n=1n−3. In Section5.9 thisis identifiedas theRiemann
zetafunction, ζ(3). Weforma linearcombination
∞summationdisplay
n=1n−3+a2α2=∞summationdisplay
n=1n−3+a2
4.
α1is not included since it converges more slowly than ζ(3). Combining terms, we obtain
ontheleft-handside
∞summationdisplay
n=1braceleftbigg1
n3+a2
n(n+1)(n+2)bracerightbigg
=∞summationdisplay
n=1n2(1+a2)+3n+2
n3(n+1)(n+2).
If wechoose a2=−1,theprecedingequationsyield
ζ(3)=∞summationdisplay
n=1n−3=1
4+∞summationdisplay
n=13n+2
n3(n+1)(n+2). (5.50)
7These series sums may be verified by expanding the forms by partial fractions, writing out the initial terms, and inspecting the
pattern of cancellationof positive andnegative terms.
5.2 Convergence Tests 335
The resulting series may not be beautiful but it does converge as n−4, faster than n−3.
A more convenient form comes from Exercise 5.2.21. There, the symmetry leads to con-
vergenceas n−5. /squaresolid
The method can be extended, including a3α3to get convergence as n−5,a4α4to get
convergenceas n−6, and so on. Eventually,you have to reach a compromise between how
muchalgebrayoudoandhowmucharithmeticthecomputerdoes.Ascomputersgetfaster,
thebalanceis steadilyshiftingtolessalgebrafor youandmorearithmeticforthem.
Exercises
5.2.1 (a) Provethatif
limn→∞npun=A<∞,p>1,
theseriessummationtext∞
n=1unconverges.
(b) Provethatif
limn→∞nun=A>0,
theseries diverges.(Thetestfails for A=0.)
These two tests, known as limit tests , are often convenient for establishing the conver-
genceofaseries. Theymaybetreatedascomparisontests, comparingwith
summationdisplay
nn−q,1≤q<p .
5.2.2 If
limn→∞bn
an=K,
aconstantwith 0 <K<∞, showthatsummationtext
nbnconvergesor divergeswithsummationtextan.
Hint.Ifsummationtextanconverges,use b′
n=1
2Kbn.I fsummationtext
nandiverges,use b′′
n=2
Kbn.
5.2.3 Showthatthecompleted’AlembertratiotestfollowsdirectlyfromKummer’stestwith
ai=1.
5.2.4 ShowthatRaabe’stestisindecisivefor P=1byestablishingthat P=1fortheseries
(a)un=1
nlnnandthatthisseries diverges.
(b)un=1
n(lnn)2andthatthisseries converges.
Note. By direct additionsummationtext100,000
2[n(lnn)2]−1=2.02288. The remainder of the series
n>105yields 0.08686 by the integral comparison test. The total, then, 2 to ∞,i s
2.1097.
336 Chapter 5 Infinite Series
5.2.5 Gauss’testis oftengivenintheform ofatestof theratio
un
un+1=n2+a1n+a0
n2+b1n+b0.
Forwhatvaluesoftheparameters a1andb1is thereconvergence?divergence?
ANS.Convergentfor a1−b1>1,
divergentfor a1−b1≤1.
5.2.6 Test forconvergence
(a)∞summationdisplay
n=2(lnn)−1(d)∞summationdisplay
n=1bracketleftbig
n(n+1)bracketrightbig−1/2
(b)∞summationdisplay
n=1n!
10n(e)∞summationdisplay
n=01
2n+1.
(c)∞summationdisplay
n=11
2n(2n+1)
5.2.7 Test forconvergence
(a)∞summationdisplay
n=11
n(n+1)(d)∞summationdisplay
n=1lnparenleftbigg
1+1
nparenrightbigg
(b)∞summationdisplay
n=21
nlnn(e)∞summationdisplay
n=11
n·n1/n.
(c)∞summationdisplay
n=11
n2n
5.2.8 Forwhatvaluesof pandqwillthefollowingseries converge?summationtext∞
n=21
np(lnn)q.
ANS.ConvergentforbraceleftBigg
p>1,allq,
p=1,q >1,divergentforbraceleftBigg
p<1,allq,
p=1,q≤1.
5.2.9 Determinetherangeofconvergencefor Gauss’s hypergeometricseries
F(α,β,γ;x)=1+αβ
1!γx+α(α+1)β(β+1)
2!γ(γ+1)x2+···.
Hint. Gauss developed his test for the specific purpose of establishing the convergence
ofthisseries.
ANS.Convergentfor −1<x<1 andx=±1i fγ>α+β.
5.2.10 Apocketcalculatoryields
100summationdisplay
n=1n−3=1.202007.
5.2 Convergence Tests 337
Showthat
1.202056≤∞summationdisplay
n=1n−3≤1.202057.
Hint.Useintegralstosetupperandlowerboundsonsummationtext∞
n=101n−3.
Note. A more exact value for summation of ζ(3)=summationtext∞
n=1n−3is 1.202 056 903 ...;
ζ(3)isknowntobeanirrationalnumber,butitisnotlinkedtoknownconstantssuchas
e,π,γ,ln2.
5.2.11 Setupperandlowerboundsonsummationtext1,000,000
n=1n−1, assumingthat
(a) theEuler–Mascheroniconstantis known.
ANS. 14.392726<1,000,000summationdisplay
n=1n−1<14.392727.
(b) TheEuler–Mascheroniconstantisunknown.
5.2.12 Givensummationtext1,000
n=1n−1=7.485470...setupperandlowerboundsontheEuler–Mascheroni
constant.
ANS. 0.5767<γ<0.5778.
5.2.13 (FromOlbers’ paradox .) Assume a static universe in which the stars are uniformly
distributed. Divide all space into shells of constant thickness; the stars in any one shell
by themselves subtend a solid angle of ω0.Allowing for the blocking out of distant
stars by nearer stars , show that the total net solid angle subtended by all stars, shells
extending to infinity, is exactly4π. [Therefore the night sky should be ablaze with
light. For more details, see E. Harrison, Darkness at Night: A Riddle of the Universe .
Cambridge,MA:HarvardUniversityPress (1987).]
5.2.14 Test forconvergence
∞summationdisplay
n=1bracketleftbigg1·3·5···(2n−1)
2·4·6···(2n)bracketrightbigg2
=1
4+9
64+25
256+···.
5.2.15 TheLegendreseriessummationtext
jevenuj(x)satisfiestherecurrencerelations
uj+2(x)=(j+1)(j+2)−l(l+1)
(j+2)(j+3)x2uj(x),
in which the index jis even and lis some constant (but, in this problem, nota non-
negative odd integer). Find the range of values of xfor which this Legendre series is
convergent.Test theendpoints.
ANS.−1<x<1.
338 Chapter 5 Infinite Series
5.2.16 A series solution (Section 9.5) of the Chebyshev equation leads to successive terms
havingtheratio
uj+2(x)
uj(x)=(k+j)2−n2
(k+j+1)(k+j+2)x2,
withk=0 andk=1.Test for convergenceat x=±1.
ANS.Convergent.
5.2.17 Aseriessolutionfortheultraspherical(Gegenbauer)function Cα
n(x)leadstotherecur-
rence
aj+2=aj(k+j)(k+j+2α)−n(n+2α)
(k+j+1)(k+j+2).
Investigate the convergence of each of these series at x=±1 as a function of the para-
meterα.
ANS.Convergentfor α<1,
divergentfor α≥1.
5.2.18 Aseriesexpansionoftheincompletebetafunction(Section8.4)yields
Bx(p,q)=xpbraceleftbigg1
p+1−q
p+1x+(1−q)(2−q)
2!(p+2)x2+···
+(1−q)(2−q)···(n−q)
n!(p+n)xn+···bracerightbigg
.
Given that 0≤x≤1,p>0, andq>0, test this series for convergence. What happens
atx=1?
5.2.19 Showthatthefollowingseriesis convergent.
∞summationdisplay
s=0(2s−1)!!
(2s)!!(2s+1).
Note.(2s−1)!!=(2s−1)(2s−3)···3·1with(−1)!!=1;(2s)!!=(2s)(2s−2)···4·2
with 0!!=1. The series appears as a series expansion of sin−1(1)and equals π/2, and
sin−1x≡arcsinx/negationslash=(sinx)−1.
5.2.20 Show how to combine ζ(2)=summationtext∞
n=1n−2withα1andα2to obtain a series converging
asn−4.
Note.ζ(2)isknown: ζ(2)=π2/6 (see Section5.9).
5.2.21 The convergence improvement of Example 5.2.5 may be carried out more expediently
(in this special case) by putting α2into a more symmetric form: Replacing nbyn−1,
wehave
α′
2=∞summationdisplay
n=21
(n−1)n(n+1)=1
4.
5.3 Alternating Series 339
(a) Combine ζ(3)andα′
2toobtainconvergenceas n−5.
(b) Let α′
4beα4withn→n−2.Combine ζ(3),α′
2, andα′
4toobtainconvergenceas
n−7.
(c) Ifζ(3)istobecalculatedtosix =decimal=placeaccuracy(error5 ×10−7),how
many terms are required for ζ(3)alone? combined as in part (a)? combined as in
part(b)?
Note.Theerror maybeestimatedusingthecorrespondingintegral.
ANS. (a) ζ(3)=5
4−∞summationdisplay
n=21
n3(n2−1).
5.2.22 Catalan’sconstant (β(2)ofM.AbramowitzandI.A.Stegun,HandbookofMathemati-
calFunctionswithFormulas,Graphs,andMathematicalTables(AMS-55),Wash,D.C.
National Bureau of Standards (1972); reprinted Dover (1974), Chapter 23) is defined
by
β(2)=∞summationdisplay
k=0(−1)k(2k+1)−2=1
12−1
32+1
52···.
Calculate β(2)tosix-digitaccuracy.
Hint.Therateofconvergenceisenhancedbypairingtheterms:
(4k−1)−2−(4k+1)−2=16k
(16k2−1)2.
Ifyouhavecarriedenoughdigitsinyourseriessummation,summationtext
1≤k≤N16k/(16k2−1)2,
additionalsignificantfiguresmaybeobtainedbysettingupperandlowerboundsonthe
tail of the series,summationtext∞
k=N+1. These bounds may be set by comparison with integrals, as
intheMaclaurinintegraltest.
ANS.β(2)=0.915965594177 ....
5.3 A LTERNATING SERIES
In Section 5.2 we limited ourselves to series of positive terms. Now, in contrast, we con-
sider infinite series in which the signs alternate. The partial cancellationdue to alternating
signsmakesconvergencemorerapidandmucheasiertoidentify.WeshallprovetheLeib-
niz criterion, a general condition for the convergence of an alternating series. For series
with more irregular sign changes, like Fourier series of Chapter 14 (see Example 5.3.1),
theintegraltestof Eq. (5.25)is oftenhelpful.
Leibniz Criterion
Considertheseriessummationtext∞
n=1(−1)n+1anwithan>0.Ifan,ismonotonicallydecreasing (for
sufficientlylarge n) and lim n→∞an=0, thentheseries converges.To provethis theorem,
340 Chapter 5 Infinite Series
weexaminetheevenpartialsums
s2n=a1−a2+a3−···−a2n,
s2n+2=s2n+(a2n+1−a2n+2).(5.51)
Sincea2n+1>a2n+2,weha v e
s2n+2>s2n. (5.52)
Ontheotherhand,
s2n+2=a1−(a2−a3)−(a4−a5)−···−a2n+2. (5.53)
Hence,witheachpairofterms a2p−a2p+1>0,
s2n+2<a1. (5.54)
With the even partial sums bounded s2n<s2n+2<a1and the terms andecreasing
monotonicallyandapproachingzero,thisalternatingseriesconverges.
Onefurtherimportantresultcanbeextractedfromthepartialsumsofthesamealternat-
ingseries. Fromthedifferencebetweentheseries limit Sandthepartialsum sn,
S−sn=an+1−an+2+an+3−an+4+···
=an+1−(an+2−an+3)−(an+4−an+5)−···, (5.55)
or
S−sn<an+1. (5.56)
Equation (5.56) says that the error in cutting off an alternating series whose terms are
monotonicallydecreasing after nterms is less than an+1, the first term dropped. A knowl-
edgeoftheerror obtainedthis waymaybeofgreatpracticalimportance.
Absolute Convergence
Givenaseriesofterms uninwhichunmayvaryinsign,ifsummationtext|un|converges,thensummationtextunis
said to be absolutely convergent. Ifsummationtextunconverges butsummationtext|un|diverges, the convergence
iscalledconditional .
Thealternatingharmonicseriesisasimpleexampleofthisconditionalconvergence.We
have
∞summationdisplay
n=1(−1)n−1n−1=1−1
2+1
3−1
4+···+(−1)n−1
n+···, (5.57)
convergentbytheLeibnizcriterion;but
∞summationdisplay
n=1n−1=1+1
2+1
3+1
4+···+1
n+···
hasbeenshowntobedivergentinSections5.1 and5.2.
5.3 Alternating Series 341
NotethatmosttestsdevelopedinSection5.2assumeaseriesofpositiveterms.Therefore
thesetests inthatsectionguaranteeabsoluteconvergence.
Example 5.3.1 SERIES WITH IRREGULAR SIGNCHANGES
For 0<x<2πtheFourierseries(see Chapter14.1)
∞summationdisplay
n=1cos(nx)
n=−lnparenleftbigg
2sinx
2parenrightbigg
(5.58)
converges, having coefficients that change sign often, but not so that the Leibniz conver-
gencecriterionapplieseasily.LetusapplytheintegraltestofEq.(5.22).Usingintegration
bypartsweseeimmediatelythat
integraldisplay∞
1cos(nx)
ndn=bracketleftbiggsin(nx)
nxbracketrightbigg∞
1+1
xintegraldisplay∞
n=1sin(nx)
n2dn
converges,andtheintegralontheright-handsideevenconvergesabsolutely.Thederivative
terminEq.(5.22) hastheform
integraldisplay∞
1parenleftbig
n−[n]parenrightbigbraceleftbigg
−x
nsin(nx)−cos(nx)
n2bracerightbigg
dn,
where the second term converges absolutely and need not be considered further. Next we
observethat g(N)=integraltextN
1(n−[n])sin(nx)dnisboundedfor N→∞,justasintegraltextNsin(nx)dn
is bounded because of the periodic nature of sin (nx)and its regular sign changes. Using
integrationbypartsagain,
integraldisplay∞
1g′(n)
ndn=bracketleftbiggg(n)
nbracketrightbigg∞
n=1+integraldisplay∞
1g(n)
n2dn,
we see that the second term is absolutely convergent and that the first goes to zero at the
upper limit. Hence the series in Eq. (5.58) converges, which is hard to see from other
convergencetests.
Alternatively, we may apply the q=1 case of the Euler–Maclaurin integration formula
inEq.(5.168b),
nsummationdisplay
ν=1f(ν)=integraldisplayn
1f(x)dx+1
2braceleftbig
f(n)+f(1)bracerightbig
+1
12braceleftbig
f′(n)−f′(1)bracerightbig
−1
2integraldisplay1
0parenleftbigg
x2−x+1
6parenrightbiggn−1summationdisplay
ν=1f′′(x+ν)dx,
whichisstraightforwardbutmoretediousbecauseof thesecondderivative. /squaresolid
342 Chapter 5 Infinite Series
Exercises
5.3.1 (a) From the electrostatic two-hemisphere problem (Exercise 12.3.20) we obtain the
series
∞summationdisplay
s=0(−1)s(4s+3)(2s−1)!!
(2s+2)!!.
Testit forconvergence.
(b) Thecorrespondingseriesfor thesurfacechargedensityis
∞summationdisplay
s=0(−1)s(4s+3)(2s−1)!!
(2s)!!.
Testit forconvergence.
The!!notationisexplainedinSection8.1andExercise5.2.19.
5.3.2 Showbydirectnumericalcomputationthatthesumof thefirst10termsof
lim
x→1ln(1+x)=ln2=∞summationdisplay
n=1(−1)n−1n−1
differs from ln2 bylessthantheeleventhterm: ln2 =0.6931471806 ....
5.3.3 In Exercise 5.2.9 the hypergeometric series is shown convergent for x=±1,ifγ>
α+β. Show that there is conditional convergence for x=−1f o rγdown to γ>
α+β−1.
Hint. The asymptotic behavior of the factorial function is given by Stirling’s series,
Section8.3.
5.4 A LGEBRA OF SERIES
The establishment of absolute convergence is important because it can be proved that ab-
solutely convergent series may be reordered according to the familiar rules of algebra or
arithmetic.
•If aninfiniteseries isabsolutelyconvergent,theseries sumis independentof theorder
inwhichthetermsareadded.
•Theseriesmaybemultipliedwithanotherabsolutelyconvergentseries.Thelimitofthe
productwillbetheproductoftheindividualserieslimits.Theproductseries,adouble
series, willalso convergeabsolutely.
Nosuchguaranteescanbegivenforconditionallyconvergentseries.Againconsiderthe
alternatingharmonicseries. If wewrite
1−1
2+1
3−1
4+···=1−parenleftbig1
2−1
3parenrightbig
−parenleftbig1
4−1
5parenrightbig
−···, (5.59)
5.4 Algebra of Series 343
itisclearthatthesum
∞summationdisplay
n=1(−1)n−1n−1<1. (5.60)
However, if we rearrange the terms slightly, we may make the alternating harmonic series
convergeto3
2. WeregroupthetermsofEq. (5.59), taking
parenleftbig
1+1
3+1
5parenrightbig
−parenleftbig1
2parenrightbig
+parenleftbig1
7+1
9+1
11+1
13+1
15parenrightbig
−parenleftbig1
4parenrightbig
+parenleftbig1
17+···+1
25parenrightbig
−parenleftbig1
6parenrightbig
+parenleftbig1
27+···+1
35parenrightbig
−parenleftbig1
8parenrightbig
+···.(5.61)
Treating the terms grouped in parentheses as single terms for convenience, we obtain the
partialsums
s1=1.5333 s2=1.0333
s3=1.5218 s4=1.2718
s5=1.5143 s6=1.3476
s7=1.5103 s8=1.3853
s9=1.5078 s10=1.4078.
From this tabulation of snand the plot of snversusnin Fig. 5.3, the convergence to
3
2is fairly clear. We have rearranged the terms, taking positive terms until the partial sum
was equal to or greater than3
2and then adding in negative terms until the partial sum just
fell below3
2and so on. As the series extends to infinity, all original terms will eventually
appear,butthepartialsumsof thisrearrangedalternatingharmonicseries convergeto3
2.
By a suitable rearrangement of terms, a conditionally convergent series may be made
to converge to any desired value or even to diverge. This statement is sometimes given
FIGURE 5.3Alternatingharmonicseries—terms
rearrangedtogiveconvergenceto1.5.
344 Chapter 5 Infinite Series
asRiemann’s theorem . Obviously, conditionally convergent series must be treated with
caution.
Absolutely convergent series can be multiplied without problems. This follows as a
special case from the rearrangement of double series. However, conditionally convergent
series cannot always be multiplied to yield convergent series, as the following example
shows.
Example 5.4.1 SQUARE OF A CONDITIONALLY CONVERGENT SERIES MAYDIVERGE
Theseriessummationtext∞
n=1(−1)n−1
√nconverges,bytheLeibnizcriterion.Its square,
bracketleftbiggsummationdisplay
n(−1)n−1
√nbracketrightbigg2
=summationdisplay
n(−1)nbracketleftbigg1√
11√
n−1+1√
21√
n−2+···+1√
n−11√
1bracketrightbigg
,
hasthegeneralterminbracketsconsistingof n−1additiveterms,eachofwhichisgreater
than1√
n−1√
n−1,so the product term in brackets is greater thann−1
n−1and does not go to
zero.Hencethisproductoscillatesandthereforediverges. /squaresolid
Hence for a product of two series to converge, we have to demand as a sufficient con-
dition that at least one of them converge absolutely. To prove this product convergence
theorem thatifsummationtext
nunconvergesabsolutelyto U,summationtext
nvnconvergesto V,then
summationdisplay
ncn,c n=nsummationdisplay
m=0umvn−m
converges to UV,it is sufficient to show that the difference terms Dn≡c0+c1+···+
c2n−UnVn→0f o rn→∞,whereUn,Vnarethepartialsumsofourseries.Asaresult,
thepartialsumdifferences
Dn=u0v0+(u0v1+u1v0)+···+(u0v2n+u1v2n−1+···+u2nv0)
−(u0+u1+···+un)(v0+v1+···+vn)
=u0(vn+1+···+v2n)+u1(vn+1+···+v2n−1)+···+un+1vn+1
+vn+1(v0+···+vn−1)+···+u2nv0,
sofor allsufficientlylarge n,
|Dn|<ǫparenleftbig
|u0|+···+| un−1|parenrightbig
+Mparenleftbig
|un+1|+···+| u2n|parenrightbig
<ǫ(a+M),
because|vn+1+vn+2+···+vn+m|<ǫforsufficientlylarge nandallpositiveintegers m
assummationtextvnconverges,andthepartialsums Vn<Bofsummationtext
nvnareboundedby M,becausethe
sumconverges.Finallywecallsummationtext
n|un|=a,assummationtextunconvergesabsolutely.
Two series can be multiplied, provided one of them converges absolutely. Addition and
subtractionofseries is alsovalidtermwiseifoneseries convergesabsolutely.
5.4 Algebra of Series 345
Improvement of Convergence,
Rational Approximations
Theseries
ln(1+x)=∞summationdisplay
n=1(−1)n−1xn
n,−1<x≤1, (5.61a)
converges very slowly as xapproaches+1. Therateof convergence may be improved
substantially by multiplying both sides of Eq. (5.61a) by a polynomial and adjusting the
polynomial coefficients to cancel the more slowly converging portions of the series. Con-
siderthesimplestpossibility:Multiply ln (1+x)by 1+a1x:
(1+a1x)ln(1+x)=∞summationdisplay
n=1(−1)n−1xn
n+a1∞summationdisplay
n=1(−1)n−1xn+1
n.
Combiningthetwoseries ontheright,termbyterm,weobtain
(1+a1x)ln(1+x)=x+∞summationdisplay
n=2(−1)n−1parenleftbigg1
n−a1
n−1parenrightbigg
xn
=x+∞summationdisplay
n=2(−1)n−1n(1−a1)−1
n(n−1)xn.
Clearly, if we take a1=1, thenin the numerator disappears and our combined series
convergesas n−2.
Continuing this process, we find that (1+2x+x2)ln(1+x)vanishes as n−3and that
(1+3x+3x2+x3)ln(1+x)vanishes as n−4. In effect we are shifting from a simple
series expansion of Eq. (5.61a) to a rational fraction representation in which the function
ln(1+x)isrepresentedbytheratioofa seriesandapolynomial:
ln(1+x)=x+summationtext∞
n=2(−1)nxn/[n(n−1)]
1+x.
Suchrationalapproximationsmaybebothcompactandaccurate.
Rearrangement of Double Series
Another aspect of the rearrangement of series appears in the treatment of double series
(Fig.5.4):
∞summationdisplay
m=0∞summationdisplay
n=0an,m.
Letussubstitute
n=q≥0,m=p−q≥0(q≤p).
346 Chapter 5 Infinite Series
FIGURE 5.4Double
series—summationover n
indicatedbyverticaldashed
lines.
Thisresults intheidentity
∞summationdisplay
m=0∞summationdisplay
n=0an,m=∞summationdisplay
p=0psummationdisplay
q=0aq,p−q. (5.62)
Thesummationover pandqof Eq.(5.62) isillustratedinFig.5.5.Thesubstitution
n=s≥0,m=r−2s≥0parenleftbigg
s≤r
2parenrightbigg
leadsto
∞summationdisplay
m=0∞summationdisplay
n=0an,m=∞summationdisplay
r=0[r/2]summationdisplay
s=0as,r−2s, (5.63)
FIGURE 5.5Doubleseries
—again,thefirst summation
is representedbyvertical
dashedlines,butthese
verticallinescorrespondto
diagonalsinFig.5.4.
5.4 Algebra of Series 347
FIGURE 5.6Doubleseries. The
summationover scorrespondstoa
summationalongthealmost-horizontal
dashedlinesinFig.5.4.
with[r/2]=r/2f o rreven and (r−1)/2f o rrodd. The summation over randsof
Eq. (5.63) is shown in Fig. 5.6. Equations (5.62) and (5.63) are clearly rearrangements of
the array of coefficients anm, rearrangements that are valid as long as we have absolute
convergence.
ThecombinationofEqs. (5.62) and(5.63),
∞summationdisplay
p=0psummationdisplay
q=0aq,p−q=∞summationdisplay
r=0[r/2]summationdisplay
s=0as,r−2s, (5.64)
isusedinSection12.1inthedeterminationoftheseriesformoftheLegendrepolynomials.
Exercises
5.4.1 Giventheseries(derivedinSection5.6)
ln(1+x)=x−x2
2+x3
3−x4
4···,−1<x≤1,
showthat
lnparenleftbigg1+x
1−xparenrightbigg
=2parenleftbigg
x+x3
3+x5
5+···parenrightbigg
,−1<x<1.
The original series, ln (1+x), appears in an analysis of binding energy in crystals. It
is1
2the Madelung constant (2ln2)for a chain of atoms. The second series is useful
in normalizing the Legendre polynomials (Section 12.3) and in developing a second
solutionfor Legendre’sdifferentialequation(Section12.10).
5.4.2 Determine the values of the coefficients a1,a2, anda3that will make
(1+a1x+a2x2+a3x3)ln(1+x)convergeas n−4. Findtheresultingseries.
5.4.3 Showthat
(a)∞summationdisplay
n=2bracketleftbig
ζ(n)−1bracketrightbig
=1, (b)∞summationdisplay
n=2(−1)nbracketleftbig
ζ(n)−1bracketrightbig
=1
2,
whereζ(n)istheRiemannzetafunction.
348 Chapter 5 Infinite Series
5.4.4 Writeaprogramthatwillrearrangethetermsofthealternatingharmonicseriestomake
theseriesconvergeto1.5.GroupyourtermsasindicatedinEq.(5.61).Listthefirst100
successivepartialsumsthatjustclimbabove1.5orjustdropbelow1.5,andlistthenew
termsincludedineachsuchpartialsum.
ANS.n1 2 3 4 5
sn1.5333 1.0333 1.5218 1.2718 1.5143
5.5 S ERIES OF FUNCTIONS
Weextendourconceptofinfiniteseriestoincludethepossibilitythateachterm unmaybe
afunctionofsomevariable, un=un(x).Numerousillustrationsofsuchseriesoffunctions
appearinChapters11–14.The partialsumsbecomefunctionsof thevariable x,
sn(x)=u1(x)+u2(x)+···+un(x), (5.65)
asdoestheseriessum, definedasthelimitofthepartialsums:
∞summationdisplay
n=1un(x)=S(x)=limn→∞sn(x). (5.66)
So far we have concerned ourselves with the behavior of the partial sums as a function
ofn.Nowweconsiderhowtheforegoingquantitiesdependon x.The keyconcepthereis
thatof uniformconvergence.
Uniform Convergence
Ifforanysmall ε>0thereexistsanumber N,independentof xintheinterval[a,b](that
is,a≤x≤b)suchthat
vextendsinglevextendsingleS(x)−sn(x)vextendsinglevextendsingle<ε, foralln≥N, (5.67)
then the series is said to be uniformly convergent in the interval [a,b]. This says that for
our series to be uniformly convergent, it must be possible to find a finite Nso that the tail
of the infinite series, |summationtext∞
i=N+1ui(x)|, will be less than an arbitrarily small εfor allxin
thegiveninterval.
Thiscondition,Eq.(5.67),whichdefinesuniformconvergence,isillustratedinFig.5.7.
Thepointisthatnomatterhowsmall εistakentobe,wecanalwayschoose nlargeenough
so that the absolute magnitude of the difference between S(x)andsn(x)is less than εfor
allx,a≤x≤b.Ifthiscannotbedone,thensummationtextun(x)isnotuniformlyconvergentin [a,b].
Example 5.5.1 NONUNIFORM CONVERGENCE
∞summationdisplay
n=1un(x)=∞summationdisplay
n=1x
[(n−1)x+1][nx+1]. (5.68)
5.5 Series of Functions 349
FIGURE 5.7Uniformconvergence.
Thepartialsum sn(x)=nx(nx+1)−1,asmaybeverifiedby mathematicalinduction .
By inspection this expression for sn(x)holds for n=1,2. We assume it holds for nterms
andthenproveitholdsfor n+1t e r m s :
sn+1(x)=sn(x)+x
[nx+1][(n+1)x+1]
=nx
[nx+1]+x
[nx+1][(n+1)x+1]
=(n+1)x
(n+1)x+1,
completingtheproof.
Lettingnapproachinfinity,weobtain
S(0)=limn→∞sn(0)=0,
S(x/negationslash=0)=limn→∞sn(x/negationslash=0)=1.
We have a discontinuity in our series limit at x=0. However, sn(x)is a continuous func-
tion ofx,0≤x≤1, for all finite n. No matter how small εmay be, Eq. (5.67) will be
violatedfor allsufficientlysmall x. Ourseriesdoesnotconvergeuniformly. /squaresolid
Weierstrass M(Majorant) Test
The most commonly encountered test for uniform convergence is the Weierstrass Mtest.
If we can construct a series of numberssummationtext∞
1Mi, in which Mi≥|ui(x)|for allxin the
interval[a,b]andsummationtext∞
1Miis convergent, our series ui(x)will beuniformly convergent
in[a,b].
350 Chapter 5 Infinite Series
TheproofofthisWeierstrass Mtestisdirectandsimple.Sincesummationtext
iMiconverges,some
numberNexistssuchthatfor n+1≥N,
∞summationdisplay
i=n+1Mi<ε. (5.69)
This follows from our definition of convergence. Then, with |ui(x)|≤Mifor allxin the
intervala≤x≤b,
∞summationdisplay
i=n+1vextendsinglevextendsingleui(x)vextendsinglevextendsingle<ε. (5.70)
Hence
vextendsinglevextendsingleS(x)−sn(x)vextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingle∞summationdisplay
i=n+1ui(x)vextendsinglevextendsinglevextendsinglevextendsingle<ε, (5.71)
and by definitionsummationtext∞
i=1ui(x)is uniformly convergent in [a,b]. Since we have specified
absolute values in the statement of the Weierstrass Mtest, the seriessummationtext∞
i=1ui(x)is also
seentobe absolutely convergent.
Note that uniform convergence and absolute convergence are independent properties.
Neitherimpliestheother.Forspecificexamples,
∞summationdisplay
n=1(−1)n
n+x2,−∞<x<∞, (5.72)
and
∞summationdisplay
n=1(−1)n−1xn
n=ln(1+x),0≤x≤1, (5.73)
converge uniformly in the indicated intervals but do not converge absolutely. On the other
hand,
∞summationdisplay
n=0(1−x)xn=1,0≤x<1
=0,x=1, (5.74)
convergesabsolutelybutdoesnotconvergeuniformlyin [0,1].
Fromthedefinitionofuniformconvergencewemayshowthatanyseries
f(x)=∞summationdisplay
n=1un(x) (5.75)
cannotconvergeuniformlyinanyintervalthatincludesadiscontinuityof f(x)ifallun(x)
arecontinuous.
Since the Weierstrass Mtest establishes both uniform and absolute convergence, it will
necessarilyfail forseries thatareuniformlybutconditionallyconvergent.
5.5 Series of Functions 351
Abel’s Test
Asomewhatmoredelicatetestfor uniformconvergencehasbeengivenbyAbel.If
un(x)=anfn(x),
summationdisplay
an=A,convergent
and the functions fn(x)are monotonic [fn+1(x)≤fn(x)]and bounded, 0 ≤fn(x)≤M,
forallxin[a,b], thensummationtext
nun(x)convergesuniformly in[a,b].
Thistestisespeciallyusefulinanalyzingpowerseries(compareSection5.7).Detailsof
theproofofAbel’stestandothertestsforuniformconvergencearegivenintheAdditional
Readingslistedattheendof thischapter.
Uniformlyconvergentseries havethreeparticularlyusefulproperties.
1. If theindividualterms un(x)arecontinuous,theseries sum
f(x)=∞summationdisplay
n=1un(x) (5.76)
isalsocontinuous.
2. If the individual terms un(x)are continuous, the series may be integrated term by
term.Thesumoftheintegralsis equaltotheintegralofthesum.
integraldisplayb
af(x)dx=∞summationdisplay
n=1integraldisplayb
aun(x)dx. (5.77)
3. The derivative of the series sum f(x)equals the sum of the individual term deriva-
tives:
d
dxf(x)=∞summationdisplay
n=1d
dxun(x), (5.78)
providedthefollowingconditionsaresatisfied:
un(x)anddun(x)
dxarecontinuousin [a,b].
∞summationdisplay
n=1dun(x)
dxisuniformlyconvergentin [a,b].
Term-by-term integration of a uniformly convergent series8requires only continuity of
the individual terms. This condition is almost always satisfied in physical applications.
Term-by-term differentiation of a series is often not valid because more restrictive condi-
tions must be satisfied. Indeed, we shall encounter Fourier series in Chapter 14 in which
term-by-termdifferentiationof auniformlyconvergentseries leadstoadivergentseries.
8Term-by-term integration may also bevalid in the absenceof uniform convergence.
352 Chapter 5 Infinite Series
Exercises
5.5.1 Findtherangeof uniformconvergenceof theDirichletseries
(a)∞summationdisplay
n=1(−1)n−1
nx,(b)ζ(x)=∞summationdisplay
n=11
nx.
ANS.(a) 0 <s≤x<∞.
(b) 1<s≤x<∞.
5.5.2 Forwhatrangeof xisthegeometricseriessummationtext∞
n=0xnuniformlyconvergent?
ANS.−1<−s≤x≤s<1.
5.5.3 Forwhatrangeofpositivevaluesof xissummationtext∞
n=01/(1+xn)
(a) convergent? (b) uniformlyconvergent?
5.5.4 Iftheseriesofthecoefficientssummationtextanandsummationtextbnareabsolutelyconvergent,showthatthe
Fourierseries
summationdisplay
(ancosnx+bnsinnx)
isuniformly convergentfor −∞<x<∞.
5.6 T AYLOR ’SEXPANSION
This is an expansion of a function into an infinite series of powers of a variable xor into
a finite series plus a remainder term. The coefficients of the successive terms of the series
involvethesuccessivederivativesofthefunction.WehavealreadyusedTaylor’sexpansion
in the establishment of a physical interpretation of divergence (Section 1.7) and in other
sectionsof Chapters1and2. NowwederivetheTaylorexpansion.
We assume that our function f(x)has a continuous nth derivative9in the interval a≤
x≤b.Then,integratingthis nthderivative ntimes,
integraldisplayx
af(n)(x1)dx1=f(n−1)(x1)vextendsinglevextendsinglevextendsinglex
a=f(n−1)(x)−f(n−1)(a),
integraldisplayx
adx2integraldisplayx2
adx1f(n)(x1)=integraldisplayx
adx2bracketleftbig
f(n−1)(x2)−f(n−1)(a)bracketrightbig
(5.79)
=f(n−2)(x)−f(n−2)(a)−(x−a)f(n−1)(a).
Continuing,weobtain
integraldisplayx
adx3integraldisplayx3
adx2integraldisplayx2
adx1f(n)(x1)=f(n−3)(x)−f(n−3)(a)−(x−a)f(n−2)(a)
−(x−a)2
2!f(n−1)(a). (5.80)
9Taylor’s expansion may be derived under slightly less restrictive conditions; compare H. Jeffreys and B. S. Jeffreys, Methods
of Mathematical Physics ,3rd ed.Cambridge: Cambridge University Press (1956), Section 1.133.
5.6 Taylor’s Expansion 353
Finally,onintegratingfor the nthtime,
integraldisplayx
adxn···integraldisplayx2
adx1f(n)(x1)=f(x)−f(a)−(x−a)f′(a)−(x−a)2
2!f′′(a)
−···−(x−a)n−1
(n−1)!f(n−1)(a). (5.81)
Note that this expression is exact. No terms have been dropped, no approximations made.
Now,solvingfor f(x),weha v e
f(x)=f(a)+(x−a)f′(a)
+(x−a)2
2!f′′(a)+···+(x−a)n−1
(n−1)!f(n−1)(a)+Rn.(5.82)
Theremainder, Rn, isgivenbythe n-foldintegral
Rn=integraldisplayx
adxn···integraldisplayx2
adx1f(n)(x1). (5.83)
This remainder, Eq. (5.83), may be put into a perhaps more practical form by using the
meanvaluetheorem of integralcalculus:
integraldisplayx
ag(x)dx=(x−a)g(ξ), (5.84)
witha≤ξ≤x.Byintegrating ntimeswegettheLagrangianform10of theremainder:
Rn=(x−a)n
n!f(n)(ξ). (5.85)
With Taylor’s expansion in this form we are not concerned with any questions of infinite
series convergence. This series is finite, and the only questions concern the magnitude of
theremainder.
Whenthefunction f(x)is suchthat
limn→∞Rn=0, (5.86)
Eq.(5.82) becomesTaylor’sseries:
f(x)=f(a)+(x−a)f′(a)+(x−a)2
2!f′′(a)+···
=∞summationdisplay
n=0(x−a)n
n!f(n)(a).11(5.87)
10Analternateform derived by Cauchyis
Rn=(x−ζ)n−1(x−a)
(n−1)!f(n)(ζ),
witha≤ζ≤x.
11Notethat 0!=1 (compare Section8.1).
354 Chapter 5 Infinite Series
Our Taylor series specifies the value of a function at one point, x, in terms of the value
ofthefunctionanditsderivativesatareferencepoint a.Itisanexpansioninpowersofthe
changein the variable, /Delta1x=x−ain this case. The notation may be varied at the user’s
convenience.Withthesubstitution x→x+handa→xwehaveanalternateform,
f(x+h)=∞summationdisplay
n=0hn
n!f(n)(x).
Whenweusethe operator D=d/dx,theTaylorexpansionbecomes
f(x+h)=∞summationdisplay
n=0hnDn
n!f(x)=ehDf(x).
(The transition to the exponential form anticipates Eq. (5.90), which follows.) An equiva-
lent operator form of this Taylor expansion appears in Exercise 4.2.4. A derivation of the
TaylorexpansioninthecontextofcomplexvariabletheoryappearsinSection6.5.
Maclaurin Theorem
If weexpandabouttheorigin (a=0), Eq. (5.87)is knownasMaclaurin’sseries:
f(x)=f(0)+xf′(0)+x2
2!f′′(0)+···=∞summationdisplay
n=0xn
n!f(n)(0). (5.88)
An immediate application of the Maclaurin series (or the Taylor series) is in the expan-
sionofvarioustranscendentalfunctionsintoinfinite(power)series.
Example 5.6.1 EXPONENTIAL FUNCTION
Letf(x)=ex. Differentiating,wehave
f(n)(0)=1 (5.89)
foralln,n=1,2,3,....Then, withEq. (5.88), wehave
ex=1+x+x2
2!+x3
3!+···=∞summationdisplay
n=0xn
n!. (5.90)
This is the series expansion of the exponential function. Some authors use this series to
definetheexponentialfunction.
Althoughthisseriesisclearlyconvergentforall x,weshouldchecktheremainderterm,
Rn. ByEq. (5.85)wehave
Rn=xn
n!f(n)(ξ)=xn
n!eξ,0≤|ξ|≤x. (5.91)
5.6 Taylor’s Expansion 355
Therefore
|Rn|≤xnex
n!(5.92)
and
limn→∞Rn=0 (5.93)
for allfinitevalues of x, which indicates that this Maclaurin expansion of exconverges
absolutelyovertherange −∞<x<∞. /squaresolid
Example 5.6.2 LOGARITHM
Letf(x)=ln(1+x). Bydifferentiating,weobtain
f′(x)=(1+x)−1,
f(n)(x)=(−1)n−1(n−1)!(1+x)−n. (5.94)
TheMaclaurinexpansion(Eq. (5.88)) yields
ln(1+x)=x−x2
2+x3
3−x4
4+···+Rn
=nsummationdisplay
p=1(−1)p−1xp
p+Rn. (5.95)
Inthiscaseourremainderis givenby
Rn=xn
n!f(n)(ξ), 0≤ξ≤x
≤xn
n,0≤ξ≤x≤1. (5.96)
Now, the remainder approaches zero as nis increased indefinitely, provided 0 ≤x≤1.12
Asaninfiniteseries,
ln(1+x)=∞summationdisplay
n=1(−1)n−1xn
n(5.97)
converges for−1<x≤1. The range−1<x<1 is easily established by the d’Alembert
ratiotest(Section5.2).Convergenceat x=1followsbytheLeibnizcriterion(Section5.3).
Inparticular,at x=1w eh a v e
ln2=1−1
2+1
3−1
4+1
5−···=∞summationdisplay
n=1(−1)n−1n−1, (5.98)
theconditionallyconvergentalternatingharmonicseries. /squaresolid
12This range caneasilybe extendedto −1<x≤1 but not to x=−1.
356 Chapter 5 Infinite Series
Binomial Theorem
A second, extremely important application of the Taylor and Maclaurin expansions is the
derivationofthebinomialtheoremfornegativeand/ornonintegralpowers.
Letf(x)=(1+x)m, in which mmay be negative and is not limited to integral values.
Directapplicationof Eq.(5.88) gives
(1+x)m=1+mx+m(m−1)
2!x2+···+Rn. (5.99)
Forthisfunctiontheremainderis
Rn=xn
n!(1+ξ)m−nm(m−1)···(m−n+1) (5.100)
andξlies between 0 and x,0≤ξ≤x.N o w ,f o r n>m ,(1+ξ)m−nis a maximum for
ξ=0.Therefore
Rn≤xn
n!m(m−1)···(m−n+1). (5.101)
Notethatthe mdependentfactorsdonotyieldazerounless misanonnegativeinteger; Rn
tends to zero as n→∞ifxis restricted to the range 0 ≤x<1. The binomial expansion
thereforeisshowntobe
(1+x)m=1+mx+m(m−1)
2!x2+m(m−1)(m−2)
3!x3+···.(5.102)
Inother,equivalentnotation,
(1+x)m=∞summationdisplay
n=0m!
n!(m−n)!xn=∞summationdisplay
n=0parenleftbiggm
nparenrightbigg
xn. (5.103)
The quantityparenleftbigm
nparenrightbig
, which equals m!/[n!(m−n)!], is called a binomial coefficient .A l -
thoughwehaveonlyshownthattheremaindervanishes,
limn→∞Rn=0,
for 0≤x<1, the series in Eq. (5.102) actually may be shown to be convergent for the
extendedrange −1<x<1.Forman integer, (m−n)!=±∞ifn>m(Section8.1) and
theseries automaticallyterminatesat n=m.
Example 5.6.3 RELATIVISTIC ENERGY
Thetotalrelativisticenergyofaparticleofmass mandvelocity vis
E=mc2parenleftbigg
1−v2
c2parenrightbigg−1/2
. (5.104)
Comparethisexpressionwiththeclassicalkineticenergy, mv2/2.
5.6 Taylor’s Expansion 357
ByEq. (5.102)with x=−v2/c2andm=−1/2w eh a v e
E=mc2bracketleftbigg
1−1
2parenleftbigg
−v2
c2parenrightbigg
+(−1/2)(−3/2)
2!parenleftbigg
−v2
c2parenrightbigg2
+(−1/2)(−3/2)(−5/2)
3!parenleftbigg
−v2
c2parenrightbigg3
+···bracketrightbigg
,
or
E=mc2+1
2mv2+3
8mv2·v2
c2+5
16mv2·parenleftbiggv2
c2parenrightbigg2
+···.(5.105)
Thefirst term, mc2, isidentifiedastherest massenergy.Then
Ekinetic=1
2mv2bracketleftbigg
1+3
4v2
c2+5
8parenleftbiggv2
c2parenrightbigg2
+···bracketrightbigg
. (5.106)
For particle velocity v≪c, the velocity of light, the expression in the brackets reduces
to unity and we see that the kinetic portion of the total relativistic energy agrees with the
classicalresult. /squaresolid
Forpolynomialswecangeneralizethebinomialexpansionto
(a1+a2+···+am)n=summationdisplay n!
n1!n2!···nm!an1
1an2
2···anmm,
where the summation includes all different combinations of n1,n2,...,nmwithsummationtextm
i=1ni=n.H e r eniandnare all integral. This generalization finds considerable use
instatisticalmechanics.
MaclaurinseriesmaysometimesappearindirectlyratherthanbydirectuseofEq.(5.88).
Forinstance,themostconvenientwaytoobtaintheseriesexpansion
sin−1x=∞summationdisplay
n=0(2n−1)!!
(2n)!!·x2n+1
(2n+1)=x+x3
6+3x5
40+···, (5.106a)
istomakeuseoftherelation(from sin y=x,getdy/dx=1/√
1−x2)
sin−1x=integraldisplayx
0dt
(1−t2)1/2.
We expand (1−t2)−1/2(binomial theorem) and then integrate term by term. This term-
by-termintegrationisdiscussedinSection5.7.TheresultisEq.(5.106a).Finally,wemay
takethelimitas x→1.TheseriesconvergesbyGauss’test, Exercise5.2.5.
358 Chapter 5 Infinite Series
Taylor Expansion — More Than One Variable
If the function fhas more than one independent variable, say, f=f(x,y), the Taylor
expansionbecomes
f(x,y)=f(a,b)+(x−a)∂f
∂x+(y−b)∂f
∂y
+1
2!bracketleftbigg
(x−a)2∂2f
∂x2+2(x−a)(y−b)∂2f
∂x∂y+(y−b)2∂2f
∂y2bracketrightbigg
+1
3!bracketleftbigg
(x−a)3∂3f
∂x3+3(x−a)2(y−b)∂3f
∂x2∂y
+3(x−a)(y−b)2∂3f
∂x∂y2+(y−b)3∂3f
∂y3bracketrightbigg
+···, (5.107)
with all derivatives evaluated at the point (a,b).U s i n gαjt=xj−xj0, we may write the
Taylorexpansionfor mindependentvariablesinthesymbolicform
f(x1,...,xm)=∞summationdisplay
n=0tn
n!parenleftbiggmsummationdisplay
i=1αi∂
∂xiparenrightbiggn
f(x1,...,xm)vextendsinglevextendsinglevextendsingle
(xk=xk0,k=1,...,m).(5.108)
Aconvenientvectorform for m=3i s
ψ(r+a)=∞summationdisplay
n=01
n!(a·∇)nψ(r). (5.109)
Exercises
5.6.1 Showthat
(a) sinx=∞summationdisplay
n=0(−1)nx2n+1
(2n+1)!,
(b) cosx=∞summationdisplay
n=0(−1)nx2n
(2n)!.
InSection6.1, eixis definedbyaseries expansionsuchthat
eix=cosx+isinx.
Thisisthebasisforthepolarrepresentationofcomplexquantities.Asaspecialcasewe
find,with x=π, theintriguingrelation
eiπ=−1.
5.6 Taylor’s Expansion 359
5.6.2 Deriveaseries expansionof cot xinincreasingpowersof xbydividing cos xby sinx.
Note. The resultant series that starts with 1 /xis actually a Laurent series (Section 6.5).
Although the two series for sin xand cosxwere valid for all x, the convergence of the
series for cot xis limited by the zeros of the denominator, sin x(see Analytic Continu-
ationinSection6.5).
5.6.3 TheRaabetest forsummationtext
n(nlnn)−1leadsto
limn→∞nbracketleftbigg(n+1)ln(n+1)
nlnn−1bracketrightbigg
.
Showthatthislimitis unity(whichmeansthattheRaabetesthereisindeterminate).
5.6.4 Showbyseries expansionthat
1
2lnη0+1
η0−1=coth−1η0,|η0|>1.
Thisidentitymaybeusedtoobtainasecondsolutionfor Legendre’sequation.
5.6.5 Show that f(x)=x1/2(a) has no Maclaurin expansion but (b) has a Taylor expansion
about any point x0/negationslash=0. Find the range of convergence of the Taylor expansion about
x=x0.
5.6.6 Letxbe an approximation for a zero of f(x)and/Delta1xbe the correction. Show that by
neglectingtermsoforder (/Delta1x)2,
/Delta1x=−f(x)
f′(x).
This is Newton’s formula for finding a root. Newton’s method has the virtues of illus-
tratingseries expansionsandelementarycalculusbutis verytreacherous.
5.6.7 Expand a function /Phi1(x,y,z) by Taylor’s expansion about (0,0,0)toO(a3). Evaluate
¯/Phi1, the average value of /Phi1, averaged over a small cube of side acentered on the origin
andshowthattheLaplacianof /Phi1isameasureofdeviationof /Phi1from/Phi1(0,0,0).
5.6.8 Theratiooftwodifferentiablefunctions f(x)andg(x)takesontheindeterminateform
0/0a tx=x0. UsingTaylorexpansionsprove L’Hôpital’srule ,
limx→x0f(x)
g(x)=limx→x0f′(x)
g′(x).
5.6.9 Withn>1,showthat
(a)1
n−lnparenleftbiggn
n−1parenrightbigg
<0,(b)1
n−lnparenleftbiggn+1
nparenrightbigg
>0.
Use these inequalities to show that the limit defining the Euler–Mascheroni constant,
Eq. (5.28),is finite.
5.6.10 Expand(1−2tz+t2)−1/2inpowersof t.Assumethat tissmall.Collectthecoefficients
oft0,t1, andt2.
360 Chapter 5 Infinite Series
ANS.a0=P0(z)=1,
a1=P1(z)=z,
a2=P2(z)=1
2(3z2−1),
where an=Pn(z),t h enthLegendrepolynomial.
5.6.11 UsingthedoublefactorialnotationofSection8.1, showthat
(1+x)−m/2=∞summationdisplay
n=0(−1)n(m+2n−2)!!
2nn!(m−2)!!xn,
form=1,2,3,....
5.6.12 Usingbinomialexpansions,comparethethreeDopplershiftformulas:
(a)ν′=νparenleftbigg
1∓v
cparenrightbigg−1
movingsource ;
(b)ν′=νparenleftbigg
1±v
cparenrightbigg
movingobserver ;
(c)ν′=νparenleftbigg
1±v
cparenrightbiggparenleftbigg
1−v2
c2parenrightbigg−1/2
relativistic.
Note.Therelativisticformulaagreeswiththeclassicalformulasiftermsoforder v2/c2
canbeneglected.
5.6.13 Inthetheoryofgeneralrelativitytherearevariouswaysofrelating(defining)avelocity
ofrecessionof agalaxytoitsredshift, δ.Milne’smodel(kinematicrelativity)gives
(a)v1=cδparenleftbigg
1+1
2δparenrightbigg
,
(b)v2=cδparenleftbigg
1+1
2δparenrightbigg
(1+δ)−2,
(c) 1+δ=bracketleftbigg1+v3/c
1−v3/cbracketrightbigg1/2
.
1.Showthatfor δ≪1 (andv3/c≪1)allthreeformulasreduceto v=cδ.
2.Comparethethreevelocitiesthroughtermsof order δ2.
Note. In special relativity (with δreplaced by z), the ratio of observed wavelength λto
emittedwavelength λ0isgivenby
λ
λ0=1+z=parenleftbiggc+v
c−vparenrightbigg1/2
.
5.6.14 Therelativisticsum woftwo velocities uandvis givenby
w
c=u/c+v/c
1+uv/c2.
5.6 Taylor’s Expansion 361
If
v
c=u
c=1−α,
where 0≤α≤1,findw/cinpowersof αthroughtermsin α3.
5.6.15 The displacement xof a particle of rest mass m0, resulting from a constant force m0g
alongthe x-axis,is
x=c2
gbraceleftbiggbracketleftbigg
1+parenleftbigg
gt
cparenrightbigg2bracketrightbigg1/2
−1bracerightbigg
,
includingrelativisticeffects. Findthe displacement xas a powerseries intime t.C o m -
parewiththeclassicalresult,
x=1
2gt2.
5.6.16 By use of Dirac’s relativistic theory, the fine structure formula of atomic spectroscopy
isgivenby
E=mc2bracketleftbigg
1+γ2
(s+n−|k|)2bracketrightbigg−1/2
,
where
s=parenleftbig
|k|2−γ2parenrightbig1/2,k=±1,±2,±3,....
Expandinpowersof γ2throughorder γ4(γ2=Ze2/4πε0¯hc,withZtheatomicnum-
ber).ThisexpansionisusefulincomparingthepredictionsoftheDiracelectrontheory
withthoseofarelativisticSchrödingerelectrontheory.Experimentalresultssupportthe
Diractheory.
5.6.17 Inahead-onproton–protoncollision,theratioofthekineticenergyinthecenterofmass
systemtotheincidentkineticenergyis
R=bracketleftbigradicalBig
2mc2parenleftbig
Ek+2mc2parenrightbig
−2mc2bracketrightbig
/Ek.
Findthevalueof thisratioofkineticenergiesfor
(a)Ek≪mc2(nonrelativistic)
(b)Ek≫mc2(extreme-relativistic).
ANS.(a)1
2,(b) 0. Thelatteransweris asortof law
ofdiminishingreturnsfor high-energyparticle
accelerators(withstationarytargets).
5.6.18 Withbinomialexpansions
x
1−x=∞summationdisplay
n=1xn,x
x−1=1
1−x−1=∞summationdisplay
n=0x−n.
Addingthesetwoseries yieldssummationtext∞
n=−∞xn=0.
Hopefully,wecanagreethatthisisnonsense,butwhathasgonewrong?
362 Chapter 5 Infinite Series
5.6.19 (a) Planck’stheoryof quantizedoscillatorsleadstoanaverageenergy
/angbracketleftε/angbracketright=summationtext∞
n=1nε0exp(−nε0/kT)summationtext∞
n=0exp(−nε0/kT),
whereε0is a fixed energy. Identify the numerator and denominator as binomial
expansionsandshowthattheratiois
/angbracketleftε/angbracketright=ε0
exp(ε0/kT)−1.
(b) Showthatthe /angbracketleftε/angbracketrightofpart(a) reducesto kT,theclassicalresult,for kT≫ε0.
5.6.20 (a) ExpandbythebinomialtheoremandintegratetermbytermtoobtaintheGregory
seriesfor y=tan−1x(notethat tan y=x):
tan−1x=integraldisplayx
0dt
1+t2=integraldisplayx
0braceleftbig
1−t2+t4−t6+···bracerightbig
dt
=∞summationdisplay
n=0(−1)nx2n+1
2n+1,−1≤x≤1.
(b) Bycomparingseriesexpansions,showthat
tan−1x=i
2lnparenleftbigg1−ix
1+ixparenrightbigg
.
Hint.CompareExercise5.4.1.
5.6.21 Innumericalanalysisitisoftenconvenienttoapproximate d2ψ(x)/dx2by
d2
dx2ψ(x)≈1
h2bracketleftbig
ψ(x+h)−2ψ(x)+ψ(x−h)bracketrightbig
.
Findtheerror inthisapproximation.
ANS. Error=h2
12ψ(4)(x).
5.6.22 Y ouha v eafunction y(x)tabulatedatequallyspacedvaluesoftheargument
braceleftBigg
yn=y(xn)
xn=x+nh.
Showthatthelinearcombination
1
12h{−y2+8y1−8y−1+y−2}
yields
y′
0−h4
30y(5)
0+···.
Hence this linear combination yields y′
0if(h4/30)y(5)
0and higher powers of hand
higherderivativesof y(x)arenegligible.
5.7 Power Series 363
5.6.23 Inanumericalintegrationofapartialdifferentialequation,thethree-dimensionalLapla-
cianis replacedby
∇2ψ(x,y,z)→h−2bracketleftbig
ψ(x+h,y,z)+ψ(x−h,y,z)
+ψ(x,y+h,z)+ψ(x,y−h,z)+ψ(x,y,z+h)
+ψ(x,y,z−h)−6ψ(x,y,z)bracketrightbig
.
Determinetheerrorinthisapproximation.Here histhestepsize,thedistancebetween
adjacentpointsinthe x-,y-, orz-direction.
5.6.24 Usingdoubleprecision,calculate efromits Maclaurinseries.
Note. This simple, direct approach is the best way of calculating eto high accuracy.
Sixteen terms give eto 16 significant figures. The reciprocal factorials give very rapid
convergence.
5.7 P OWER SERIES
Thepowerseries is aspecialandextremelyusefultypeofinfiniteseriesof theform
f(x)=a0+a1x+a2x2+a3x3+···=∞summationdisplay
n=0anxn, (5.110)
wherethecoefficients aiareconstants,independentof x.13
Convergence
Equation (5.110) may readily be tested for convergence by either the Cauchy root test or
thed’Alembertratiotest(Section5.2). If
limn→∞vextendsinglevextendsinglevextendsinglevextendsinglean+1
anvextendsinglevextendsinglevextendsinglevextendsingle=R−1, (5.111)
the series converges for −R<x<R . This is the interval or radius of convergence. Since
the root and ratio tests fail when the limit is unity, the endpoints of the interval require
specialattention.
For instance, if an=n−1, thenR=1 and, from Sections 5.1, 5.2, and 5.3, the series
converges for x=−1 but diverges for x=+1. Ifan=n!, thenR=0 and the series
divergesfor all x/negationslash=0.
Uniform and Absolute Convergence
Suppose our power series (Eq. (5.110)) has been found convergent for −R<x<R ; then
itwillbeuniformlyandabsolutelyconvergentinany interiorinterval,−S≤x≤S,where
0<S<R.
This maybeproveddirectlybytheWeierstrass Mtest(Section5.5).
13Equation (5.110) may be generalized to z=x+iy, replacing x. The following two chapters will then yield uniform conver-
gence,integrability, anddifferentiability in aregion ofacomplex plane in placeof aninterval onthe x-axis.
364 Chapter 5 Infinite Series
Continuity
Since each of the terms un(x)=anxnis a continuous function of xandf(x)=summationtextanxn
converges uniformly for −S≤x≤S,f(x)must be a continuous function in the interval
ofuniformconvergence.
ThisbehavioristobecontrastedwiththestrikinglydifferentbehavioroftheFourierse-
ries(Chapter14),inwhichtheFourierseriesisusedfrequentlytorepresentdiscontinuous
functionssuchassawtoothandsquarewaves.
Differentiation and Integration
Withun(x)continuous andsummationtextanxnuniformly convergent, we find that the differentiated
series is a power series with continuous functions and the same radius of convergence as
the original series. The new factors introduced by differentiation (or integration) do not
affect either the root or the ratio test. Therefore our power series may be differentiated or
integratedasoftenasdesiredwithintheintervalofuniformconvergence(Exercise5.7.13).
In view of the rather severe restrictions placed on differentiation (Section 5.5), this is
aremarkableandvaluableresult.
Uniqueness Theorem
In the preceding section, using the Maclaurin series, we expanded exand ln(1+x)into
infinite series. In the succeeding chapters, functions are frequently represented or perhaps
definedbyinfiniteseries. We nowestablishthatthepower-seriesrepresentationisunique.
If
f(x)=∞summationdisplay
n=0anxn,−Ra<x<R a
=∞summationdisplay
n=0bnxn,−Rb<x<R b, (5.112)
withoverlappingintervalsofconvergence,includingtheorigin,then
an=bn (5.113)
for alln; that is, we assume two (different) power-series representations and then proceed
toshowthatthetwoareactuallyidentical.
FromEq. (5.112),
∞summationdisplay
n=0anxn=∞summationdisplay
n=0bnxn,−R<x<R, (5.114)
whereRis the smaller of Ra,Rb. By setting x=0 to eliminate all but the constant terms,
weobtain
a0=b0. (5.115)
5.7 Power Series 365
Now, exploiting the differentiability of our power series, we differentiate Eq. (5.114), get-
ting
∞summationdisplay
n=1nanxn−1=∞summationdisplay
n=1nbnxn−1. (5.116)
Weagainset x=0,toisolatethenewconstantterms, andfind
a1=b1. (5.117)
Byrepeatingthisprocess ntimes,weget
an=bn, (5.118)
which shows that the two series coincide. Therefore our power-series representation is
unique.
This will be a crucial point in Section 9.5, in which we use a power series to develop
solutions of differential equations. This uniqueness of power series appears frequently in
theoreticalphysics.Theestablishmentofperturbationtheoryinquantummechanicsisone
example.Thepower-seriesrepresentationoffunctionsisoftenusefulinevaluatingindeter-
minateforms,particularlywhenl’Hôpital’srulemaybeawkwardtoapply(Exercise5.7.9).
Example 5.7.1 L’HÔPITAL ’SRULE
Evaluate
lim
x→01−cosx
x2. (5.119)
Replacing cos xbyits Maclaurin-seriesexpansion,weobtain
1−cosx
x2=1−(1−1
2!x2+1
4!x4−···)
x2=1
2!−x2
4!+···.
Lettingx→0,wehave
lim
x→01−cosx
x2=1
2. (5.120)
Theuniquenessofpowerseriesmeansthatthecoefficients anmaybeidentifiedwiththe
derivativesinaMaclaurinseries. From
f(x)=∞summationdisplay
n=0anxn=∞summationdisplay
n=01
n!f(n)(0)xn
wehave
an=1
n!f(n)(0). /squaresolid
366 Chapter 5 Infinite Series
Inversion of Power Series
Supposewearegivenaseries
y−y0=a1(x−x0)+a2(x−x0)2+···=∞summationdisplay
n=1an(x−x0)n.(5.121)
This gives (y−y0)in terms of (x−x0). However, it may be desirable to have an explicit
expression for (x−x0)in terms of (y−y0). We may solve Eq. (5.121) for x−x0by
inversionofourseries. Assumethat
x−x0=∞summationdisplay
n=1bn(y−y0)n, (5.122)
with thebnto be determined in terms of the assumed known an. A brute-force approach,
whichisperfectlyadequateforthefirstfewcoefficients,issimplytosubstituteEq.(5.121)
into Eq. (5.122). By equating coefficients of (x−x0)non both sides of Eq. (5.122), since
thepowerseries isunique,weobtain
b1=1
a1,
b2=−a2
a3
1,
b3=1
a5
1parenleftbig
2a2
2−a1a3parenrightbig
, (5.123)
b4=1
a7
1parenleftbig
5a1a2a3−a2
1a4−5a3
2parenrightbig
,andsoon .
Some of the higher coefficients are listed by Dwight.14A more general and much more
elegant approach is developed by the use of complex variables in the first and second
editionsof MathematicalMethodsforPhysicists .
Exercises
5.7.1 TheclassicalLangevintheoryofparamagnetismleadstoanexpressionforthemagnetic
polarization,
P(x)=cparenleftbiggcoshx
sinhx−1
xparenrightbigg
.
ExpandP(x)asapowerseries forsmall x(lowfields, hightemperature).
14H. B. Dwight, Tables of Integrals and Other Mathematical Data , 4th ed. New York: Macmillan (1961). (Compare Formula
No.50.)
5.7 Power Series 367
5.7.2 The depolarizing factor Lfor an oblate ellipsoid in a uniform electric field parallel to
theaxisofrotationis
L=1
ε0parenleftbig
1+ζ2
0parenrightbigparenleftbig
1−ζ0cot−1ζ0parenrightbig
,
whereζ0definesanoblateellipsoidinoblatespheroidalcoordinates (ξ,ζ,ϕ).Showthat
lim
ζ0→∞L=1
3ε0(sphere),lim
ζ0→0L=1
ε0(thinsheet) .
5.7.3 Thedepolarizingfactor (Exercise5.7.2)for aprolateellipsoidis
L=1
ε0parenleftbig
η2
0−1parenrightbigparenleftbigg1
2η0lnη0+1
η0−1−1parenrightbigg
.
Showthat
limη0→∞L=1
3ε0(sphere),lim
η0→0L=0 (longneedle) .
5.7.4 Theanalysisofthediffractionpatternofacircularopeninginvolves
integraldisplay2π
0cos(ccosϕ)dϕ.
Expandtheintegrandinaseries andintegratebyusing
integraldisplay2π
0cos2nϕdϕ=(2n)!
22n(n!)2·2π,integraldisplay2π
0cos2n+1ϕdϕ=0.
Theresultis 2 πtimestheBesselfunction J0(c).
5.7.5 Neutrons are created (by a nuclear reaction) inside a hollow sphere of radius R.T h e
newly created neutrons are uniformly distributed over the spherical volume. Assuming
thatalldirectionsareequallyprobable(isotropy),whatistheaveragedistanceaneutron
will travel before striking the surface of the sphere? Assume straight-line motion and
nocollisions.
(a) Showthat
¯r=3
2Rintegraldisplay1
0integraldisplayπ
0radicalbig
1−k2sin2θk2dksinθdθ.
(b) Expandtheintegrandas aseriesandintegratetoobtain
¯r=Rbracketleftbigg
1−3∞summationdisplay
n=11
(2n−1)(2n+1)(2n+3)bracketrightbigg
.
(c) Showthatthesumof thisinfiniteseriesis 1 /12,giving¯r=3
4R.
Hint. Show that sn=1/12−[4(2n+1)(2n+3)]−1by mathematical induction. Then
letn→∞.
368 Chapter 5 Infinite Series
5.7.6 Giventhat
integraldisplay1
0dx
1+x2=tan−1xvextendsinglevextendsinglevextendsingle1
0=π
4,
expandtheintegrandintoaseries andintegratetermbytermobtaining15
π
4=1−1
3+1
5−1
7+1
9−···+(−1)n1
2n+1+···,
whichis Leibniz’sformulafor π. Comparetheconvergenceof theintegrandseries and
theintegratedseries at x=1.SeealsoExercise5.7.18.
5.7.7 Expandtheincompletefactorialfunction
γ(n+1,x)≡integraldisplayx
0e−ttndt
inaseries ofpowersof x.Whatis therangeof convergenceof theresultingseries?
ANS.integraldisplayx
0e−ttndt=xn+1bracketleftbigg1
n+1−x
n+2+x2
2!(n+3)
−···(−1)pxp
p!(n+p+1)+···bracketrightbigg
.
5.7.8 Derivetheseriesexpansionoftheincompletebetafunction
Bx(p,q)=integraldisplayx
0tp−1(1−t)q−1dt
=xpbraceleftbigg1
p+1−q
p+1x+···+(1−q)···(n−q)
n!(p+n)xn+···bracerightbigg
for 0≤x≤1,p>0,andq>0( i fx=1).
5.7.9 Evaluate
(a) lim
x→0bracketleftbig
sin(tanx)−tan(sinx)bracketrightbig
x−7,(b) lim
x→0x−njn(x), n=3,
wherejn(x)isasphericalBesselfunction(Section11.7), definedby
jn(x)=(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbiggsinx
xparenrightbigg
.
ANS.(a)−1
30,(b)1
1·3·5···(2n+1)→1
105forn=3.
15The series expansion of tan−1x(upper limit 1 replaced by x) was discovered by James Gregory in 1671, 3 years before
Leibniz. See Peter Beckmann’s entertaining book, A History of Pi , 2nd ed., Boulder, CO: Golem Press (1971) and L. Berggren,
J. andP.Borwein, Pi:ASource Book , NewYork: Springer (1997).
5.7 Power Series 369
5.7.10 Neutrontransporttheorygivesthefollowingexpressionfortheinverseneutrondiffusion
lengthof k:
a−b
ktanh−1parenleftbiggk
aparenrightbigg
=1.
By series inversion or otherwise, determine k2as a series of powers of b/a.G i v et h e
firsttwoterms oftheseries.
ANS.k2=3abparenleftbigg
1−4
5b
aparenrightbigg
.
5.7.11 Developaseries expansionof y=sinh−1x(thatis, sinh y=x)inpo wersof xby
(a) inversionoftheseriesfor sinh y,
(b) adirectMaclaurinexpansion.
5.7.12 Afunction f(z)isrepresentedbya descending powerseries
f(z)=∞summationdisplay
n=0anz−n,R≤z<∞.
Show that this series expansion is unique; that is, if f(z)=summationtext∞
n=0bnz−n,
R≤z<∞,thenan=bnfor alln.
5.7.13 A power series converges for −R<x<R . Show that the differentiated series and
the integrated series have the same interval of convergence. (Do not bother about the
endpoints x=±R.)
5.7.14 Assuming that f(x)may be expanded in a power series about the origin, f(x)=summationtext∞
n=0anxn, with some nonzero range of convergence. Use the techniques employed
inprovinguniquenessofseries toshowthatyourassumedseriesis aMaclaurinseries:
an=1
n!f(n)(0).
5.7.15 The Klein–Nishina formula for the scattering of photons by electrons contains a term
oftheform
f(ε)=(1+ε)
ε2bracketleftbigg2+2ε
1+2ε−ln(1+2ε)
εbracketrightbigg
.
Hereε=hν/mc2, theratioofthephotonenergytotheelectronrestmass energy.Find
lim
ε→0f(ε).
ANS.4
3.
5.7.16 The behavior of a neutron losing energy by colliding elastically with nuclei of mass A
isdescribedbyaparameter ξ1,
ξ1=1+(A−1)2
2AlnA−1
A+1.
370 Chapter 5 Infinite Series
Anapproximation,goodfor large A,is
ξ2=2
A+2/3.
Expandξ1andξ2in powers of A−1. Show that ξ2agrees with ξ1through(A−1)2.F i n d
thedifferenceinthecoefficientsof the (A−1)3term.
5.7.17 Showthateachof thesetwointegralsequalsCatalan’sconstant:
(a)integraldisplay1
0arctantdt
t,(b)−integraldisplay1
0lnxdx
1+x2.
Note.Seeβ(2)inSection5.9 forthevalueof Catalan’sconstant.
5.7.18 Calculate π(doubleprecision)byeachofthefollowingarctangentexpressions:
π=16tan−1(1/5)−4tan−1(1/239)
π=24tan−1(1/8)+8tan−1(1/57)+4tan−1(1/239)
π=48tan−1(1/18)+32tan−1(1/57)−20tan−1(1/239).
Obtain16significantfigures. VerifytheformulasusingExercise5.6.2.
Note.Theseformulashavebeenusedinsomeofthemoreaccuratecalculationsof π.16
5.7.19 Ananalysisof theGibbsphenomenonofSection14.5leadstotheexpression
2
πintegraldisplayπ
0sinξ
ξdξ.
(a) Expand the integrand in a series and integrate term by term. Find the numerical
valueof thisexpressiontofoursignificantfigures.
(b) EvaluatethisexpressionbytheGaussianquadratureif available.
ANS.1.178980.
5.8 E LLIPTIC INTEGRALS
Elliptic integrals are included here partly as an illustration of the use of power series and
partly for their own intrinsic interest. This interest includes the occurrence of elliptic inte-
grals in physical problems (Example 5.8.1 and Exercise 5.8.4) and applications in mathe-
maticalproblems.
Example 5.8.1 PERIOD OF A SIMPLE PENDULUM
Forsmall-amplitudeoscillations,ourpendulum(Fig.5.8)hassimpleharmonicmotionwith
aperiodT=2π(l/g)1/2.Foramaximumamplitude θMlargeenoughsothatsin θM/negationslash=θM,
Newton’ssecondlawofmotionandLagrange’sequation(Section17.7)leadtoanonlinear
differentialequation(sin θisanonlinearfunctionof θ),soweturntoadifferentapproach.
16D.Shanks and J. W.Wrench, Computation of πto 100000 decimals. Math. Comput. 16: 76 (1962).
5.8 Elliptic Integrals 371
FIGURE 5.8Simple
pendulum.
The swinging mass mhas a kinetic energy of ml2(dθ/dt)2/2 and a potential energy of
−mglcosθ(θ=π/2 taken for the arbitrary zero of potential energy). Since dθ/dt=0a t
θ=θM, conservationofenergygives
1
2ml2parenleftbiggdθ
dtparenrightbigg2
−mglcosθ=−mglcosθM. (5.124)
Solvingfor dθ/dtweobtain
dθ
dt=±parenleftbigg2g
lparenrightbigg1/2
(cosθ−cosθM)1/2, (5.125)
with the mass mcanceling out. We take tto be zero when θ=0 anddθ/dt > 0. An
integrationfrom θ=0t oθ=θMyields
integraldisplayθM
0(cosθ−cosθM)−1/2dθ=parenleftbigg2g
lparenrightbigg1/2integraldisplayt
0dt=parenleftbigg2g
lparenrightbigg1/2
t. (5.126)
This is1
4of a cycle, and therefore the time tis1
4of the period T. We note that θ≤θM,
andwithabitofclairvoyancewetry thehalf-anglesubstitution
sinparenleftbiggθ
2parenrightbigg
=sinparenleftbiggθM
2parenrightbigg
sinϕ. (5.127)
Withthis,Eq. (5.126)becomes
T=4parenleftbiggl
gparenrightbigg1/2integraldisplayπ/2
0parenleftbigg
1−sin2parenleftbiggθM
2parenrightbigg
sin2ϕparenrightbigg−1/2
dϕ. (5.128)
AlthoughnotanobviousimprovementoverEq.(5.126),theintegralnowdefinesthecom-
pleteellipticintegralofthefirstkind, K(sin2θM/2).Fromtheseriesexpansion,theperiod
ofourpendulummaybedevelopedasapowerseries—powersof sin θM/2:
T=2πparenleftbiggl
gparenrightbigg1/2braceleftbigg
1+1
4sin2θM
2+9
64sin4θM
2+···bracerightbigg
. (5.129)
/squaresolid
372 Chapter 5 Infinite Series
Definitions
GeneralizingExample5.8.1toincludetheupperlimitasavariable,the ellipticintegralof
thefirstkind isdefinedas
F(ϕ\α)=integraldisplayϕ
0parenleftbig
1−sin2αsin2θparenrightbig−1/2dθ, (5.130a)
or
F(x|m)=integraldisplayx
0bracketleftbigparenleftbig
1−t2parenrightbigparenleftbig
1−mt2parenrightbigbracketrightbig−1/2dt,0≤m<1. (5.130b)
(This is the notation of AMS-55 see footnote 4 for the reference.) For ϕ=π/2,x=1,we
havethecompleteellipticintegralof thefirstkind ,
K(m)=integraldisplayπ/2
0parenleftbig
1−msin2θparenrightbig−1/2dθ
=integraldisplay1
0bracketleftbigparenleftbig
1−t2parenrightbigparenleftbig
1−mt2parenrightbigbracketrightbig−1/2dt, (5.131)
withm=sin2α,0≤m<1.
Theellipticintegralofthesecondkind is definedby
E(ϕ\α)=integraldisplayϕ
0parenleftbig
1−sin2αsin2θparenrightbig1/2dθ (5.132a)
or
E(x|m)=integraldisplayx
0parenleftbigg1−mt2
1−t2parenrightbigg1/2
dt,0≤m≤1. (5.132b)
Again,forthecase ϕ=π/2,x=1,wehavethe completeellipticintegralofthesecond
kind:
E(m)=integraldisplayπ/2
0parenleftbig
1−msin2θparenrightbig1/2dθ
=integraldisplay1
0parenleftbigg1−mt2
1−t2parenrightbigg1/2
dt,0≤m≤1. (5.133)
Exercise5.8.1isanexampleofitsoccurrence.Figure5.9showsthebehaviorof K(m)and
E(m).ExtensivetablesareavailableinAMS-55(see Exercise5.2.22for thereference).
Series Expansion
For our range 0 ≤m<1, the denominator of K(m)may be expanded by the binomial
series
parenleftbig
1−msin2θparenrightbig−1/2=1+1
2msin2θ+3
8m2sin4θ+···
=∞summationdisplay
n=0(2n−1)!!
(2n)!!mnsin2nθ. (5.134)
5.8 Elliptic Integrals 373
FIGURE 5.9Completeellipticintegrals,
K(m)andE(m).
For any closed interval [0,mmax],mmax<1, this series is uniformly convergent and may
beintegratedterm byterm.FromExercise8.4.9,
integraldisplayπ/2
0sin2nθdθ=(2n−1)!!
(2n)!!·π
2. (5.135)
Hence
K(m)=π
2braceleftbigg
1+∞summationdisplay
n=1bracketleftbigg(2n−1)!!
(2n)!!bracketrightbigg2
mnbracerightbigg
. (5.136)
Similarly,
E(m)=π
2braceleftbigg
1−∞summationdisplay
n=1bracketleftbigg(2n−1)!!
(2n)!!bracketrightbigg2mn
2n−1bracerightbigg
(5.137)
(Exercise 5.8.2). In Section 13.5 these series are identified as hypergeometric functions,
andwehave
K(m)=π
22F1parenleftbigg1
2,1
2;1;mparenrightbigg
(5.138)
E(m)=π
22F1parenleftbigg
−1
2,1
2;1;mparenrightbigg
. (5.139)
374 Chapter 5 Infinite Series
Limiting Values
Fromtheseries Eqs. (5.136)and(5.137),or fromthedefiningintegrals,
lim
m→0K(m)=π
2, (5.140)
lim
m→0E(m)=π
2. (5.141)
Form→1 theseriesexpansionsareof littleuse.However,theintegralsyield
lim
m→1K(m)=∞, (5.142)
theintegraldiverginglogarithmically,and
lim
m→1E(m)=1. (5.143)
Theellipticintegralshavebeenusedextensivelyinthepastforevaluatingintegrals.For
instance,integralsof theform
I=integraldisplayx
0Rparenleftbig
t,radicalbig
a4t4+a3t3+a2t2+a1t1+a0parenrightbig
dt,
whereRisarationalfunctionof tandoftheradical,maybeexpressedintermsofelliptic
integrals. Jahnke and Emde, Tables of Functions with Formulae and Curves .N e wY o r k :
Dover (1943), Chapter 5, give pages of such transformations. With computers available
for direct numerical evaluation, interest in these elliptic integral techniques has declined.
However, elliptic integrals still remain of interest because of their appearance in physical
problems—seeExercises5.8.4and5.8.5.
For an extensive account of elliptic functions, integrals, and Jacobi theta functions, you
aredirectedtoWhittakerandWatson’streatise ACourseinModernAnalysis ,4thed.Cam-
bridge,UK:CambridgeUniversityPress (1962).
Exercises
5.8.1 The ellipse x2/a2+y2/b2=1 may be represented parametrically by x=asinθ,y=
bcosθ.Showthatthelengthofarc withinthefirst quadrantis
aintegraldisplayπ/2
0parenleftbig
1−msin2θparenrightbig1/2dθ=aE(m).
Here 0≤m=(a2−b2)/a2≤1.
5.8.2 Derivetheseriesexpansion
E(m)=π
2braceleftbigg
1−parenleftbigg1
2parenrightbigg2m
1−parenleftbigg1·3
2·4parenrightbigg2m2
3−···bracerightbigg
.
5.8.3 Showthat
lim
m→0(K−E)
m=π
4.
5.8 Elliptic Integrals 375
FIGURE 5.10Circularwireloop.
5.8.4 Acircularloopofwireinthe xy-plane,asshowninFig.5.10,carriesacurrent I.Gi ven
thatthevectorpotentialis
Aϕ(ρ,ϕ,z)=aµ0I
2πintegraldisplayπ
0cosαdα
(a2+ρ2+z2−2aρcosα)1/2,
showthat
Aϕ(ρ,ϕ,z)=µ0I
πkparenleftbigga
ρparenrightbigg1/2bracketleftbiggparenleftbigg
1−k2
2parenrightbigg
Kparenleftbig
k2parenrightbig
−Eparenleftbig
k2parenrightbigbracketrightbigg
,
where
k2=4aρ
(a+ρ)2+z2.
Note.ForextensionofExercise5.8.4to B, seeSmythe,p. 270.17
5.8.5 An analysis of the magnetic vector potential of a circular current loop leads to the ex-
pression
fparenleftbig
k2parenrightbig
=k−2bracketleftbigparenleftbig
2−k2parenrightbig
Kparenleftbig
k2parenrightbig
−2Eparenleftbig
k2parenrightbigbracketrightbig
,
whereK(k2)andE(k2)arethecompleteellipticintegralsofthefirstandsecondkinds.
Showthatfor k2≪1(r≫radiusof loop)
fparenleftbig
k2parenrightbig
≈πk2
16.
17W. R.Smythe, Static and DynamicElectricity ,3rd ed. NewYork: McGraw-Hill(1969).
376 Chapter 5 Infinite Series
5.8.6 Showthat
(a)dE(k2)
dk=1
k(E−K),
(b)dK(k2)
dk=E
k(1−k2)−K
k.
Hint.Forpart(b)showthat
Eparenleftbig
k2parenrightbig
=parenleftbig
1−k2parenrightbigintegraldisplayπ/2
0parenleftbig
1−ksin2θparenrightbig−3/2dθ
bycomparingseries expansions.
5.8.7 (a) Write a function subroutine that will compute E(m)from the series expansion,
Eq.(5.137).
(b) Test your function subroutine by using it to calculate E(m)over the range
m=0.0(0.1)0.9 and comparing the result with the values given by AMS-55 (see
Exercise5.2.22for thereference).
5.8.8 RepeatExercise5.8.7for K(m).
Note. These series for E(m), Eq. (5.137), and K(m), Eq. (5.136), converge only very
slowly for mnear 1. More rapidly converging series for E(m)andK(m)exist. See
Dwight’s Tables of Integrals:18No. 773.2 and 774.2. Your computer subroutine for
computing EandKprobablyuses polynomialapproximations:AMS-55,Chapter17.
5.8.9 A simple pendulum is swinging with a maximum amplitude of θM. In the limit as
θM→0, the period is 1 s. Using the elliptic integral, K(k2),k=sin(θM/2), calculate
theperiod TforθM=0(1 0◦)9 0◦.
Caution.Someellipticintegralsubroutinesrequire k=m1/2asaninputparameter,not
mitself.
Checkvalues .θM10◦50◦90◦
T(sec)1.00193 1.05033 1.18258
5.8.10 Calculate the magnetic vector potential A(ρ,ϕ,z)=ˆϕAϕ(ρ,ϕ,z)of a circular current
loop(Exercise5.8.4) for theranges ρ/a=2,3,4,andz/a=0,1,2,3,4.
Note. This elliptic integral calculationof the magneticvector potentialmay be checked
byanassociatedLegendrefunctioncalculation,Example12.5.1.
Checkvalue .F orρ/a=3 andz/a=0;Aϕ=0.029023µ0I.
5.9 B ERNOULLI NUMBERS ,
EULER –M ACLAURIN FORMULA
The Bernoulli numbers were introduced by Jacques (James, Jacob) Bernoulli. There are
several equivalent definitions, but extreme care must be taken, for some authors introduce
18H.B. Dwight, Tables of Integrals and OtherMathematical Data . NewYork: Macmillan (1947).
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 377
variations in numbering or in algebraic signs. One relatively simple approach is to define
theBernoullinumbersbytheseries19
x
ex−1=∞summationdisplay
n=0Bnxn
n!, (5.144)
which converges for |x|<2πby the ratio test substitut Eq. (5.153) (see also Exam-
ple7.1.7).Bydifferentiatingthispowerseriesrepeatedlyandthensetting x=0,weobtain
Bn=bracketleftbiggdn
dxnparenleftbiggx
ex−1parenrightbiggbracketrightbigg
x=0. (5.145)
Specifically,
B1=d
dxparenleftbiggx
ex−1parenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle
x=0=1
ex−1−xex
(ex−1)2vextendsinglevextendsinglevextendsinglevextendsingle
x=0=−1
2,(5.146)
as may be seen by series expansionof the denominators. Using B0=1 andB1=−1
2,i ti s
easytoverifythatthefunction
x
ex−1−1+x
2=∞summationdisplay
n=2Bnxn
n!=−x
e−x−1−1−x
2(5.147)
isevenin x,s oal lB2n+1=0.
ToderivearecursionrelationfortheBernoullinumbers,wemultiply
ex−1
xx
ex−1=1=braceleftbigg∞summationdisplay
m=0xm
(m+1)!bracerightbiggbraceleftbigg
1−x
2+∞summationdisplay
n=1B2nx2n
(2n)!bracerightbigg
=1+∞summationdisplay
m=1xmbraceleftbigg1
(m+1)!−1
2m!bracerightbigg
+∞summationdisplay
N=2xNsummationdisplay
1≤n≤N/2B2n
(2n)!(N−2n+1)!.(5.148)
ForN>0 thecoefficientof xNis zero,soEq. (5.148)yields
1
2(N+1)−1=summationdisplay
1≤n≤N/2B2nparenleftbiggN+1
2nparenrightbigg
=1
2(N−1), (5.149)
19The function x/(ex−1)may be considered a generating function since it generates the Bernoulli numbers. Generating
functions of the specialfunctions of mathematicalphysics appearin Chapters 11, 12, and 13.
378 Chapter 5 Infinite Series
Table 5.1 BernoulliNumbers
nB n Bn
01 1.000000000
1−1
2−0.500000000
21
60.166666667
4−1
30−0.033333333
61
420.023809524
8−1
30−0.033333333
105
660.075757576
Note. Further values are given in National Bureau of Stan-
dards,Handbook of Mathematical Functions (AMS-55).
Seefootnote4 for thereference.
whichisequivalentto
N−1
2=Nsummationdisplay
n=1B2nparenleftbigg2N+1
2nparenrightbigg
,
N−1=N−1summationdisplay
n=1B2nparenleftbigg2N
2nparenrightbigg
.(5.150)
FromEq.(5.150)theBernoullinumbersinTable5.1arereadilyobtained.Ifthevariable x
in Eq. (5.144) is replaced by 2 ixwe obtain an alternate (and equivalent) definition of B2n
(B1isset equalto−1
2byEq. (5.146))bytheexpression
xcotx=∞summationdisplay
n=0(−1)nB2n(2x)2n
(2n)!,−π<x<π . (5.151)
Using the method of residues (Section 7.1) or working from the infinite product represen-
tationof sin x(Section5.11),wefindthat
B2n=(−1)n−12(2n)!
(2π)2n∞summationdisplay
p=1p−2n,n=1,2,3,.... (5.152)
This representation of the Bernoulli numbers was discovered by Euler. It is readily seen
fromEq.(5.152)that |B2n|increaseswithoutlimitas n→∞.Numericalvalueshavebeen
calculated by Glaisher.20Illustrating the divergent behavior of the Bernoulli numbers, we
have
B20=−5.291×102
B200=−3.647×10215.
20J. W. L. Glaisher, table of the first 250 Bernoulli’s numbers (to nine figures) and their logarithms (to ten figures). Trans.
Cambridge Philos. Soc. 12: 390 (1871–1879).
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 379
SomeauthorsprefertodefinetheBernoullinumberswithamodifiedversionofEq.(5.152)
byusing
Bn=2(2n)!
(2π)2n∞summationdisplay
p=1p−2n, (5.153)
thesubscriptbeingjusthalfofoursubscriptandallsignspositive.Again,whenusingother
textsor references,youmustchecktoseeexactlyhowtheBernoullinumbersaredefined.
TheBernoullinumbersoccurfrequentlyinnumbertheory.ThevonStaudt–Clausenthe-
oremstates that
B2n=An−1
p1−1
p2−1
p3−···−1
pk, (5.154)
in which Anis an integer and p1,p2,...,pkare prime numbers so that pi−1 is a divisor
of 2n.It mayreadilybeverifiedthatthisholdsfor
B6(A3=1,p=2,3,7),
B8(A4=1,p=2,3,5), (5.155)
B10(A5=1,p=2,3,11),
andotherspecialcases.
TheBernoullinumbersappearinthesummationofintegralpowersoftheintegers,
Nsummationdisplay
j=1jp,pintegral,
and in numerous series expansions of the transcendental functions, including tan x, cotx,
ln|sinx|,(sinx)−1,l n|cosx|,l n|tanx|,(coshx)−1, tanhx,and coth x. Forexample,
tanx=x+x3
3+2
15x5+···+(−1)n−122n(22n−1)B2n
(2n)!x2n−1+···.(5.156)
TheBernoullinumbersarelikelytocomeinsuchseriesexpansionsbecauseofthedefining
equations (5.144), (5.150), and (5.151) and because of their relation to the Riemann zeta
function,
ζ(2n)=∞summationdisplay
p=1p−2n. (5.157)
Bernoulli Polynomials
If Eq.(5.144) isgeneralizedslightly,wehave
xexs
ex−1=∞summationdisplay
n=0Bn(s)xn
n!(5.158)
380 Chapter 5 Infinite Series
Table 5.2 BernoulliPolynomials
B0=1
B1=x−1
2
B2=x2−x+1
6
B3=x3−3
2x2+1
2x
B4=x4−2x3+x2−1
30
B5=x5−5
2x4+5
3x3−1
6x
B6=x6−3x5+5
2x4−1
2x2+1
42
Bn(0)=Bn, Bernoulli number
defining the Bernoulli polynomials ,Bn(s). The first seven Bernoulli polynomials are
giveninTable5.2.
Fromthegeneratingfunction,Eq.(5.158),
Bn(0)=Bn,n=0,1,2,..., (5.159)
the Bernoulli polynomial evaluated at zero equals the corresponding Bernoulli number.
Two particularly important properties of the Bernoulli polynomials follow from the defin-
ingrelation,Eq, (5.158):adifferentiationrelation
d
dsBn(s)=nBn−1(s), n=1,2,3,..., (5.160)
andasymmetryrelation(replace x→−xinEq. (5.158)andthenset s=1)
Bn(1)=(−1)nBn(0), n=1,2,3,.... (5.161)
TheserelationsareusedinthedevelopmentoftheEuler–Maclaurinintegrationformula.
Euler–Maclaurin Integration
Formula
One use of the Bernoulli functions is in the derivation of the Euler–Maclaurin integration
formula.This formulais usedinSection8.3 for thedevelopmentof anasymptoticexpres-
sionforthefactorialfunction—Stirling’sseries.
The technique is repeated integration by parts, using Eq. (5.160) to create new deriva-
tives.Westart with
integraldisplay1
0f(x)dx=integraldisplay1
0f(x)B0(x)dx. (5.162)
FromEq. (5.160)andExercise5.9.2,
B′
1(x)=B0(x)=1. (5.163)
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 381
Substituting B′
1(x)intoEq. (5.162)andintegratingbyparts, weobtain
integraldisplay1
0f(x)dx=f(1)B1(1)−f(0)B1(0)−integraldisplay1
0f′(x)B1(x)dx
=1
2bracketleftbig
f(1)+f(0)bracketrightbig
−integraldisplay1
0f′(x)B1(x)dx. (5.164)
AgainusingEq.(5.160), wehave
B1(x)=1
2B′
2(x), (5.165)
andintegratingbyparts weget
integraldisplay1
0f(x)dx=1
2bracketleftbig
f(1)+f(0)bracketrightbig
−1
2!bracketleftbig
f′(1)B2(1)−f′(0)B2(0)bracketrightbig
+1
2!integraldisplay1
0f(2)(x)B2(x)dx. (5.166)
Usingtherelations
B2n(1)=B2n(0)=B2n,n=0,1,2,...
(5.167)
B2n+1(1)=B2n+1(0)=0,n=1,2,3,...
andcontinuingthis process,wehave
integraldisplay1
0f(x)dx=1
2bracketleftbig
f(1)+f(0)bracketrightbig
−qsummationdisplay
p=11
(2p)!B2pbracketleftbig
f(2p−1)(1)−f(2p−1)(0)bracketrightbig
+1
(2q)!integraldisplay1
0f(2q)(x)B2q(x)dx. (5.168a)
Thisis theEuler–Maclaurinintegrationformula.It assumesthatthefunction f(x)has the
requiredderivatives.
TherangeofintegrationinEq.(5.168a)maybeshiftedfrom [0,1]to[1,2]byreplacing
f(x)byf(x+1).Addingsuchresultsupto [n−1,n],weobtain
integraldisplayn
0f(x)dx=1
2f(0)+f(1)+f(2)+···+f(n−1)+1
2f(n)
−qsummationdisplay
p=11
(2p)!B2pbracketleftbig
f(2p−1)(n)−f(2p−1)(0)bracketrightbig
+1
(2q)!integraldisplay1
0B2q(x)n−1summationdisplay
ν=0f(2q)(x+ν)dx. (5.168b)
The terms1
2f(0)+f(1)+···+1
2f(n)appear exactly as in trapezoidal integration, or
quadrature. The summation over pmay be interpreted as a correction to the trapezoidal
approximation. Equation (5.168b) may be seen as a generalization of Eq. (5.22); it is the
382 Chapter 5 Infinite Series
Table 5.3 RiemannZetaFunction
sζ (s)
21 .6449340668
31 .2020569032
41 .0823232337
51 .0369277551
61 .0173430620
71 .0083492774
81 .0040773562
91 .0020083928
10 1 .0009945751
formusedinExercise5.9.5forsummingpositivepowersofintegersandinSection8.3for
thederivationof Stirling’sformula.
The Euler–Maclaurin formula is often useful in summing series by converting them to
integrals.21
Riemann Zeta Function
This series,summationtext∞
p=1p−2n, was used as a comparison series for testing convergence (Sec-
tion 5.2) and in Eq. (5.152) as one definition of the Bernoulli numbers, B2n.I ta l s os e r v e s
todefinetheRiemannzetafunctionby
ζ(s)≡∞summationdisplay
n=1n−s,s>1. (5.169)
Table 5.3 lists the values of ζ(s)for integral s,s=2,3,...,10. Closed forms for even s
appear in Exercise 5.9.6. Figure 5.11 is a plot of ζ(s)−1. An integral expression for this
RiemannzetafunctionappearsinExercise8.2.21aspartofthedevelopmentofthegamma
function,andthefunctionalrelationis giveninSection14.3.
The celebrated Euler prime number product for the Riemann zeta function may be de-
rivedas
ζ(s)parenleftbig
1−2−sparenrightbig
=1+1
2s+1
3s+···−parenleftbigg1
2s+1
4s+1
6s+···parenrightbigg
; (5.170)
eliminatingallthe n−s,wherenisa multipleof 2.Then
ζ(s)parenleftbig
1−2−sparenrightbigparenleftbig
1−3−sparenrightbig
=1+1
3s+1
5s+1
7s+1
9s+···
−parenleftbigg1
3s+1
9s+1
15s+···parenrightbigg
; (5.171)
21SeeR. P.Boas and C.Stutz, Estimating sums with integrals. Am.J .Ph ys. 39: 745 (1971), for a number of examples.
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 383
FIGURE 5.11Riemannzetafunction, ζ(s)−1
versuss.
eliminating all the remaining terms in which nis a multiple of 3. Continuing, we have
ζ(s)(1−2−s)(1−3−s)(1−5−s)···(1−P−s),wherePisaprimenumber,andallterms
n−s,inwhich nis amultipleofanyintegerupthrough P, arecanceledout.As P→∞,
ζ(s)parenleftbig
1−2−sparenrightbigparenleftbig
1−3−sparenrightbig
···parenleftbig
1−P−sparenrightbig
→ζ(s)∞productdisplay
P(prime)=2parenleftbig
1−P−sparenrightbig
=1.(5.172)
Therefore
ζ(s)=∞productdisplay
P(prime)=2parenleftbig
1−P−sparenrightbig−1, (5.173)
givingζ(s)as aninfiniteproduct.22
This cancellation procedure has a clear application in numerical computation. Equa-
tion (5.170) will give ζ(s)(1−2−s)to the same accuracy as Eq. (5.169) gives ζ(s),b u t
22ThisisthestartingpointfortheextensiveapplicationsoftheRiemannzetafunctiontoanalyticnumbertheory.SeeH.M.Ed-
wards,Riemann’s Zeta Function . New York: Academic Press (1974); A. Ivi ´c,The Riemann Zeta Function . New York: Wiley
(1985); S. J. Patterson, Introduction to the Theory of the Riemann Zeta Function . Cambridge, UK: Cambridge University Press
(1988).
384 Chapter 5 Infinite Series
withonlyhalfasmanyterms.(Ineithercase,acorrectionwouldbemadefortheneglected
tailoftheseriesbytheMaclaurinintegraltesttechnique—replacingtheseriesbyaninte-
gral,Section5.2.)
Along with the Riemann zeta function, AMS-55 (Chapter 23. See Exercise 5.2.22 for
thereference.) definesthreeotherDirichletseriesrelatedto ζ(s):
η(s)=∞summationdisplay
n=1(−1)n−1n−s=parenleftbig
1−21−sparenrightbig
ζ(s),
λ(s)=∞summationdisplay
n=0(2n+1)−s=parenleftbig
1−2−sparenrightbig
ζ(s),
and
β(s)=∞summationdisplay
n=0(−1)n(2n+1)−s.
From the Bernoulli numbers (Exercise 5.9.6) or Fourier series (Example 14.3.3 and Exer-
cise14.3.13)specialvaluesare
ζ(2)=1+1
22+1
32+···=π2
6
ζ(4)=1+1
24+1
34+···=π4
90
η(2)=1−1
22+1
32+···=π2
12
η(4)=1−1
24+1
34+···=7π4
720
λ(2)=1+1
32+1
52+···=π2
8
λ(4)=1+1
34+1
54+···=π4
96
β(1)=1−1
3+1
5−···=π
4
β(3)=1−1
33+1
53−···=π3
32.
Catalan’sconstant,
β(2)=1−1
32+1
52−···=0.91596559 ...,
isthetopicof Exercise5.2.22.
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 385
Improvement of Convergence
If we are required to sum a convergent seriessummationtext∞
n=1anwhose terms are rational functions
ofn, the convergence may be improved dramatically by introducing the Riemann zeta
function.
Example 5.9.1 IMPROVEMENT OF CONVERGENCE
The problem is to evaluate the seriessummationtext∞
n=11/(1+n2). Expanding (1+n2)−1=
n−2(1+n−2)−1bydirectdivision,wehave
parenleftbig
1+n2parenrightbig−1=n−2parenleftbigg
1−n−2+n−4−n−6
1+n−2parenrightbigg
=1
n2−1
n4+1
n6−1
n8+n6.
Therefore
∞summationdisplay
n=11
1+n2=ζ(2)−ζ(4)+ζ(6)−∞summationdisplay
n=11
n8+n6.
Theζvaluesaretabulatedandtheremainderseriesconvergesas n−8.Clearly,theprocess
canbecontinuedasdesired.Youmakeachoicebetweenhowmuchalgebrayouwilldoand
how much arithmetic the computer will do. Other methods for improving computational
effectivenessaregivenattheendofSections5.2and5.4. /squaresolid
Exercises
5.9.1 Showthat
tanx=∞summationdisplay
n=1(−1)n−122n(22n−1)B2n
(2n)!x2n−1,−π
2<x<π
2.
Hint.t a nx=cotx−2cot2x.
5.9.2 ShowthatthefirstBernoullipolynomialsare
B0(s)=1
B1(s)=s−1
2
B2(s)=s2−s+1
6.
Notethat Bn(0)=Bn, theBernoullinumber.
5.9.3 Showthat B′
n(s)=nBn−1(s),n=1,2,3,....
Hint.DifferentiateEq. (5.158).
386 Chapter 5 Infinite Series
5.9.4 Showthat
Bn(1)=(−1)nBn(0).
Hint.Gobacktothegeneratingfunction,Eq.(5.158), orExercise5.9.2.
5.9.5 TheEuler–Maclaurinintegrationformulamaybeusedfortheevaluationoffiniteseries:
nsummationdisplay
m=1f(m)=integraldisplayn
0f(x)dx+1
2f(1)+1
2f(n)+B2
2!bracketleftbig
f′(n)−f′(1)bracketrightbig
+···.
Showthat
(a)nsummationdisplay
m=1m=1
2n(n+1).
(b)nsummationdisplay
m=1m2=1
6n(n+1)(2n+1).
(c)nsummationdisplay
m=1m3=1
4n2(n+1)2.
(d)nsummationdisplay
m=1m4=1
30n(n+1)(2n+1)parenleftbig
3n2+3n−1parenrightbig
.
5.9.6 From
B2n=(−1)n−12(2n)!
(2π)2nζ(2n),
showthat
(a)ζ(2)=π2
6(d)ζ(8)=π8
9450
(b)ζ(4)=π4
90(e)ζ(10)=π10
93,555.
(c)ζ(6)=π6
945
5.9.7 Planck’sblackbodyradiationlawinvolvestheintegral
integraldisplay∞
0x3dx
ex−1.
Showthatthisequals 6 ζ(4). From Exercise5.9.6,
ζ(4)=π4
90.
Hint.Makeuseofthegammafunction,Chapter8.
5.9 Bernoulli Numbers,Euler–Maclaurin Formula 387
5.9.8 Provethatintegraldisplay∞
0xnexdx
(ex−1)2=n!ζ(n).
Assuming nto be real, show that each side of the equation diverges if n=1. Hence
the preceding equation carries the condition n>1. Integrals such as this appear in the
quantumtheoryoftransporteffects—thermalandelectricalconductivity.
5.9.9 TheBloch–Gruneissenapproximationfor theresistanceinamonovalentmetalis
ρ=CT5
/Theta16integraldisplay/Theta1/T
0x5dx
(ex−1)(1−e−x),
where/Theta1istheDebyetemperaturecharacteristicofthemetal.
(a) For T→∞,showthat
ρ≈C
4·T
/Theta12.
(b) For T→0,showthat
ρ≈5!ζ(5)CT5
/Theta16.
5.9.10 Showthat
(a)integraldisplay1
0ln(1+x)
xdx=1
2ζ(2), (b) lim
a→1integraldisplaya
0ln(1−x)
xdx=ζ(2).
FromExercise5.9.6, ζ(2)=π2/6.Notethattheintegrandinpart(b)divergesfor a=1
butthatthe integrated seriesis convergent.
5.9.11 Theintegralintegraldisplay1
0bracketleftbig
ln(1−x)bracketrightbig2dx
x
appears in the fourth-order correction to the magnetic moment of the electron. Show
thatit equals 2 ζ(3).
Hint.Let1−x=e−t.
5.9.12 Showthat integraldisplay∞
0(lnz)2
1+z2dz=4parenleftbigg
1−1
33+1
53−1
73+···parenrightbigg
.
Bycontourintegration(Exercise7.1.17),this maybeshownequalto π3/8.
5.9.13 For“small”valuesof x,
ln(x!)=−γx+∞summationdisplay
n=2(−1)nζ(n)
nxn,
whereγis the Euler–Mascheroni constant and ζ(n)is the Riemann zeta function. For
whatvaluesof xdoesthisseriesconverge?
ANS.−1<x≤1.
388 Chapter 5 Infinite Series
Notethatif x=1,weobtain
γ=∞summationdisplay
n=2(−1)nζ(n)
n,
a series for the Euler–Mascheroni constant. The convergence of this series is exceed-
inglyslow.Foractualcomputationof γ,other,indirectapproachesarefarsuperior(see
Exercises5.10.11,and8.5.16).
5.9.14 Showthattheseriesexpansionof ln (x!)(Exercise5.9.13) maybewrittenas
(a) ln(x!)=1
2lnparenleftbiggπx
sinπxparenrightbigg
−γx−∞summationdisplay
n=1ζ(2n+1)
2n+1x2n+1,
(b) ln(x!)=1
2lnparenleftbiggπx
sinπxparenrightbigg
−1
2lnparenleftbigg1+x
1−xparenrightbigg
+(1−γ)x
−∞summationdisplay
n=1bracketleftbig
ζ(2n+1)−1bracketrightbigx2n+1
2n+1.
Determinetherangeofconvergenceofeachoftheseexpressions.
5.9.15 ShowthatCatalan’sconstant, β(2), maybewrittenas
β(2)=2∞summationdisplay
k=1(4k−3)−2−π2
8.
Hint.π2=6ζ(2).
5.9.16 DerivethefollowingexpansionsoftheDebyefunctionsfor n≥1:
integraldisplayx
0tndt
et−1=xnbracketleftbigg1
n−x
2(n+1)+∞summationdisplay
k=1B2kx2k
(2k+n)(2k)!bracketrightbigg
,|x|<2π;
integraldisplay∞
xtndt
et−1=∞summationdisplay
k=1e−kxbracketleftbiggxn
k+nxn−1
k2+n(n−1)xn−2
k3+···+n!
kn+1bracketrightbigg
forx>0.Thecompleteintegral (0,∞)equalsn!ζ(n+1), Exercise8.2.15.
5.9.17 (a) Showthattheequationln2 =summationtext∞
s=1(−1)s+1s−1(Exercise5.4.1)mayberewritten
as
ln2=∞summationdisplay
s=22−sζ(s)+∞summationdisplay
p=1(2p)−n−1bracketleftbigg
1−1
2pbracketrightbigg−1
.
Hint.Takethetermsinpairs.
(b) Calculate ln2 tosixsignificantfigures.
5.10 Asymptotic Series 389
5.9.18 (a) Show that the equation π/4=summationtext∞
n=1(−1)n+1(2n−1)−1(Exercise 5.7.6) may be
rewrittenas
π
4=1−2∞summationdisplay
s=14−2sζ(2s)−2∞summationdisplay
p=1(4p)−2n−2bracketleftbigg
1−1
(4p)2bracketrightbigg−1
.
(b) Calculate π/4 tosixsignificantfigures.
5.9.19 Write a function subprogram ZETA (N)that will calculate the Riemann zeta function
for integer argument. Tabulate ζ(s)fors=2,3,4,...,20. Check your values against
Table5.3andAMS-55,Chapter23.(SeeExercise5.2.22for thereference.).
Hint. If you supply the function subprogram with the known values of ζ(2),ζ(3), and
ζ(4), you avoid the more slowly converging series. Calculation time may be further
shortenedbyusingEq.(5.170).
5.9.20 Calculatethelogarithm(base 10)of |B2n|,n=10,20,...,100.
Hint.Program ζ(n)asafunctionsubprogram,Exercise5.9.19.
Checkvalues. log|B100|=78.45
log|B200|=215.56.
5.10 A SYMPTOTIC SERIES
Asymptotic series frequently occur in physics. In numerical computations they are em-
ployed for the accurate computation of a variety of functions. We consider here two types
ofintegralsthatleadtoasymptoticseries:first, anintegralof theform
I1(x)=integraldisplay∞
xe−uf(u)du,
where the variable xappears as the lower limit of an integral. Second, we consider the
form
I2(x)=integraldisplay∞
0e−ufparenleftbiggu
xparenrightbigg
du,
with the function fto be expanded as a Taylor series (binomial series). Asymptotic se-
ries often occur as solutions of differential equations. An example of this appears in Sec-
tion11.6as asolutionofBessel’s equation.
Incomplete Gamma Function
The nature of an asymptotic series is perhaps best illustrated by a specific example. Sup-
posethatwehavetheexponentialintegralfunction23
Ei(x)=integraldisplayx
−∞eu
udu, (5.174)
23This function occurs frequently in astrophysical problems involving gas with aMaxwell–Boltzmann energy distribution.
390 Chapter 5 Infinite Series
or
−Ei(−x)=integraldisplay∞
xe−u
udu=E1(x), (5.175)
to be evaluated for large values of x. Or let us take a generalization of the incomplete
factorialfunction(incompletegammafunction),24
I(x,p)=integraldisplay∞
xe−uu−pdu=Ŵ(1−p,x), (5.176)
inwhich xandparepositive.Again,weseektoevaluateit forlargevaluesof x.
Integratingbyparts,weobtain
I(x,p)=e−x
xp−pintegraldisplay∞
xe−uu−p−1du
=e−x
xp−pe−x
xp+1+p(p+1)integraldisplay∞
xe−uu−p−2du. (5.177)
Continuingtointegratebyparts,wedeveloptheseries
I(x,p)=e−xparenleftbigg1
xp−p
xp+1+p(p+1)
xp+2−···+(−1)n−1(p+n−2)!
(p−1)!xp+n−1parenrightbigg
+(−1)n(p+n−1)!
(p−1)!integraldisplay∞
xe−uu−p−ndu. (5.178)
This is a remarkable series. Checking the convergence by the d’Alembert ratio test, we
find
limn→∞|un+1|
|un|=limn→∞(p+n)!
(p+n−1)!·1
x=limn→∞p+n
x=∞ (5.179)
for all finite values of x. Therefore our series as an infinite series diverges everywhere!
BeforediscardingEq.(5.178)asworthless,letusseehowwellagivenpartialsumapprox-
imatestheincompletefactorialfunction, I(x,p):
I(x,p)−sn(x,p)=(−1)n+1(p+n)!
(p−1)!integraldisplay∞
xe−uu−p−n−1du=Rn(x,p). (5.180)
Inabsolutevalue
vextendsinglevextendsingleI(x,p)−sn(x,p)vextendsinglevextendsingle≤(p+n)!
(p−1)!integraldisplay∞
xe−uu−p−n−1du.
Whenwesubstitute u=v+x, theintegralbecomes
integraldisplay∞
xe−uu−p−n−1du=e−xintegraldisplay∞
0e−v(v+x)−p−n−1dv
=e−x
xp+n+1integraldisplay∞
0e−vparenleftbigg
1+v
xparenrightbigg−p−n−1
dv.
24SeealsoSection8.5.
5.10 Asymptotic Series 391
FIGURE 5.12Partialsumsof exE1(x)|x=5.
Forlarge xthefinalintegralapproaches1and
vextendsinglevextendsingleI(x,p)−sn(x,p)vextendsinglevextendsingle≈(p+n)!
(p−1)!·e−x
xp+n+1. (5.181)
Thismeansthatifwetake xlargeenough,ourpartialsum snisanarbitrarilygoodapprox-
imation to the function I(x,p). Our divergent series (Eq. (5.178)) therefore is perfectly
good for computations of partial sums. For this reason it is sometimes called a semicon-
vergentseries. Note that the power of xin the denominator of the remainder (p+n+1)
ishigherthanthepowerof xinthelasttermincludedin sn(x,p),(p+n).
Since the remainder Rn(x,p)alternates in sign, the successive partial sums give alter-
nately upper and lower bounds for I(x,p). The behavior of the series (with p=1) as a
functionofthenumberof termsincludedis showninFig. 5.12.Wehave
exE1(x)=exintegraldisplay∞
xe−u
udu
∼=sn(x)=1
x−1!
x2+2!
x2−3!
x4+···+(−1)nn!
xn+1,(5.182)
whichisevaluatedat x=5.Theoptimumdeterminationof exE1(x)isgivenbytheclosest
approachoftheupperandlowerbounds,thatis,between s4=s6=0.1664and s5=0.1741
forx=5.Therefore
0.1664≤exE1(x)vextendsinglevextendsingle
x=5≤0.1741. (5.183)
Actually,fromtables,
exE1(x)vextendsinglevextendsingle
x=5=0.1704, (5.184)
392 Chapter 5 Infinite Series
withinthelimitsestablishedbyourasymptoticexpansion.Notethatinclusionofadditional
terms in the series expansion beyond the optimum point literally reduces the accuracy of
the representation. As xis increased, the spread between the lowest upper bound and the
highest lower bound will diminish. By taking xlarge enough, one may compute exE1(x)
to any desired degree of accuracy. Other properties of E1(x)are derived and discussed in
Section8.5.
Cosine and Sine Integrals
Asymptoticseriesmayalsobedevelopedfromdefiniteintegrals—iftheintegrandhasthe
required behavior. As an example, the cosine and sine integrals (Section 8.5) are defined
by
Ci(x)=−integraldisplay∞
xcost
tdt, (5.185)
si(x)=−integraldisplay∞
xsint
tdt. (5.186)
Combiningthesewithregulartrigonometricfunctions,wemaydefine
f(x)=Ci(x)sinx−si(x)cosx=integraldisplay∞
0siny
y+xdy,
g(x)=−Ci(x)cosx−si(x)sinx=integraldisplay∞
0cosy
y+xdy,(5.187)
withthenewvariable y=t−x.Goingtocomplexvariables,Section6.1, wehave
g(x)+if(x)=integraldisplay∞
0eiy
y+xdy=integraldisplay∞
0ie−xu
1+iudu, (5.188)
in which u=−iy/x. The limits of integration, 0 to ∞, rather than 0 to −i∞, may be
justified by Cauchy’s theorem, Section 6.3. Rationalizing the denominator and equating
realparttorealpartandimaginaryparttoimaginarypart,weobtain
g(x)=integraldisplay∞
0ue−xu
1+u2du, f(x)=integraldisplay∞
0e−xu
1+u2du. (5.189)
Forconvergenceof theintegralswemustrequirethat ℜ(x)>0.25
25ℜ(x)=real part of (complex) x(compare Section 6.1).
5.10 Asymptotic Series 393
Now, to developthe asymptotic expansions,let v=xuand expandthe preceding factor
[1+(v/x)2]−1bythebinomialtheorem.26We have
f(x)≈1
xintegraldisplay∞
0e−vsummationdisplay
0≤n≤N(−1)nv2n
x2ndv=1
xsummationdisplay
0≤n≤N(−1)n(2n)!
x2n, (5.190)
g(x)≈1
x2integraldisplay∞
0e−vsummationdisplay
0≤n≤N(−1)nv2n+1
x2ndv=1
x2summationdisplay
0≤n≤N(−1)n(2n+1)!
x2n.
FromEqs. (5.187) and(5.190),
Ci(x)≈sinx
xsummationdisplay
0≤n≤N(−1)n(2n)!
x2n−cosx
x2summationdisplay
0≤n≤N(−1)n(2n+1)!
x2n,
si(x)≈−cosx
xsummationdisplay
0≤n≤N(−1)n(2n)!
x2n−sinx
x2summationdisplay
0≤n≤N(−1)n(2n+1)!
x2n(5.191)
arethedesiredasymptoticexpansions.
This technique of expanding the integrand of a definite integral and integrating term
by term is applied in Section 11.6 to develop an asymptotic expansion of the modified
Besselfunction KνandinSection13.5forexpansionsofthetwoconfluenthypergeometric
functions M(a,c;x)andU(a,c;x).
Definition of Asymptotic Series
The behavior of these series (Eqs. (5.178) and (5.191)), is consistent with the defining
propertiesof anasymptoticseries.27FollowingPoincaré,wetake28
xnRn(x)=xnbracketleftbig
f(x)−sn(x)bracketrightbig
, (5.192)
where
sn(x)=a0+a1
x+a2
x2+···+an
xn. (5.193)
Theasymptoticexpansionof f(x)hasthepropertiesthat
limx→∞xnRn(x)=0,for fixedn, (5.194)
and
limn→∞xnRn(x)=∞,for fixedx.29(5.195)
26This stepis validfor v≤x. Thecontributions from v≥xwill benegligible (for large x)becauseofthe negative exponential.
It is becausethe binomial expansion does not converge for v≥xthat our final series is asymptotic ratherthanconvergent.
27Itisnotnecessarythattheasymptoticseriesbeapowerseries.Therequiredpropertyisthattheremainder Rn(x)beofhigher
order than the last term kept—as in Eq.(5.194).
28Poincaré’s definition allows (or neglects) exponentially decreasing functions. The refinement of Poincaré’s definition is of
considerable importance for the advanced theory of asymptotic expansions, particularly for extensions into the complex plane.
However,forpurposesofanintroductorytreatmentandespeciallyfornumericalcomputationwith xrealandpositive,Poincaré’s
approachis perfectlysatisfactory.
394 Chapter 5 Infinite Series
See Eqs. (5.178) and (5.179) for an example of these properties. For power series, as as-
sumedintheformof sn(x),Rn(x)∼x−n−1.Withconditions(5.194)and(5.195)satisfied,
wewrite
f(x)≈∞summationdisplay
n=0anx−n. (5.196)
Note the use of ≈in place of=. The function f(x)is equal to the series only in the limit
asx→∞andafinitenumberoftermsintheseries.
Asymptotic expansions of two functions may be multiplied together, and the result will
beanasymptoticexpansionof theproductofthetwofunctions.
Theasymptoticexpansionofagivenfunction f(t)maybeintegratedtermbyterm(just
asinauniformlyconvergentseriesofcontinuousfunctions)from x≤t<∞,andtheresult
will be an asymptotic expansion ofintegraltext∞
xf(t)dt. Term-by-term differentiation, however, is
validonlyunderveryspecialconditions.
Some functions do not possess an asymptotic expansion; exis an example of such a
function. However, if a function has an asymptotic expansion, it has only one. The corre-
spondenceis notonetoone;manyfunctionsmayhavethesameasymptoticexpansion.
One of the most useful and powerful methods of generating asymptotic expansions, the
method of steepest descents, will be developed in Section 7.3. Applications include the
derivation of Stirling’s formula for the (complete) factorial function (Section 8.3) and the
asymptotic forms of the various Bessel functions (Section 11.6). Asymptotic series occur
fairlyofteninmathematicalphysics.Oneoftheearliestandstillimportantapproximations
ofquantummechanics,the WKBexpansion,isanasymptoticseries.
Exercises
5.10.1 Stirling’sformulafor thelogarithmofthefactorialfunctionis
ln(x!)=1
2ln2π+parenleftbigg
x+1
2parenrightbigg
lnx−x−Nsummationdisplay
n=1B2n
2n(2n−1)x1−2n.
TheB2nare the Bernoulli numbers (Section 5.9). Show that Stirling’s formula is an
asymptotic expansion.
5.10.2 Integratingbyparts, developasymptoticexpansionsoftheFresnelintegrals.
(a)C(x)=integraldisplayx
0cosπu2
2du,(b)s(x)=integraldisplayx
0sinπu2
2du.
Theseintegralsappearintheanalysisofaknife-edgediffractionpattern.
5.10.3 RederivetheasymptoticexpansionsofCi (x)andsi(x)byrepeatedintegrationbyparts.
Hint.Ci(x)+isi(x)=−integraltext∞
xeit
tdt.
29This excludes convergent series ofinverse powers of x. Some writers feelthat this exclusion is artificialandunnecessary.
5.10 Asymptotic Series 395
5.10.4 DerivetheasymptoticexpansionoftheGausserror function
erf(x)=2√πintegraldisplayx
0e−t2dt
≈1−e−x2
√πxparenleftbigg
1−1
2x2+1·3
22x4−1·3·5
23x6+···+(−1)n(2n−1)!!
2nx2nparenrightbigg
.
Hint:e r f(x)=1−erfc(x)=1−2√πintegraltext∞
xe−t2dt.
Normalized so that erf (∞)=1, this function plays an important role in probability
theory. It may be expressed in terms of the Fresnel integrals (Exercise 5.10.2), the in-
complete gamma functions (Section 8.5), and the confluent hypergeometric functions
(Section13.5).
5.10.5 The asymptotic expressions for the various Bessel functions, Section 11.6, contain the
series
Pν(z)∼1+∞summationdisplay
n=1(−1)nproducttext2n
s=1[4ν2−(2s−1)2]
(2n)!(8z)2n,
Qν(z)∼∞summationdisplay
n=1(−1)n+1producttext2n−1
s=1[4ν2−(2s−1)2]
(2n−1)!(8z)2n−1.
Showthatthesetwoseriesareindeedasymptoticseries.
5.10.6 Forx>1,
1
1+x=∞summationdisplay
n=0(−1)n1
xn+1.
Test thisseriestoseeif itis anasymptoticseries.
5.10.7 DerivethefollowingBernoullinumberasymptoticseriesfortheEuler–Mascheronicon-
stant:
γ=nsummationdisplay
s=1s−1−lnn−1
2n+Nsummationdisplay
k=1B2k
(2k)n2k.
Hint. Apply the Euler–Maclaurin integration formula to f(x)=x−1over the interval
[1,n]forN=1,2,....
5.10.8 Developanasymptoticseriesfor
integraldisplay∞
0e−xvparenleftbig
1+v2parenrightbig−2dv.
Takextoberealandpositive.
ANS.1
x−2!
x3+4!
x5−···+(−1)n(2n)!
x2n+1.
396 Chapter 5 Infinite Series
5.10.9 Calculatepartialsumsof exE1(x)forx=5,10,and15toexhibitthebehaviorshownin
Fig.5.11.Determinethewidthofthethroatfor x=10and15,analogoustoEq.(5.183).
ANS.Throatwidth: n=10,0.000051
n=15,0.0000002.
5.10.10 Theknife-edgediffractionpatternis describedby
I=0.5I0braceleftbigbracketleftbig
C(u0)+0.5bracketrightbig2+bracketleftbig
S(u0)+0.5bracketrightbig2bracerightbig
,
whereC(u0)andS(u0)are the Fresnel integrals of Exercise 5.10.2. Here I0is the
incident intensity and Iis the diffracted intensity; u0is proportional to the distance
away from the knife edge (measured at right angles to the incident beam). Calculate
I/I0foru0varying from−1.0t o+4.0 in steps of 0.1. Tabulate your results and, if a
plottingroutineis available,plotthem.
Checkvalue .u0=1.0,I/I0=1.259226.
5.10.11 The Euler–Maclaurin integration formula of Section 5.9 provides a way of calculating
the Euler–Mascheroni constant γto high accuracy. Using f(x)=1/xin Eq. (5.168b)
(withinterval[1,n])andthedefinitionof γ(Eq. 5.28),weobtain
γ=nsummationdisplay
s=1s−1−lnn−1
2n+Nsummationdisplay
k=1B2k
(2k)n2k.
Usingdouble-precisionarithmetic,calculate γforN=1,2,....
Note. D. E. Knuth, Euler’s constant to 1271 places. Math. Comput. 16: 275 (1962). An
evenmoreprecisecalculationappearsinExercise8.5.16.
ANS.For n=1000,N=2
γ=0.577215664901.
5.11 I NFINITE PRODUCTS
Consider a succession of positive factors f1·f2·f3·f4···fn(fi>0). Using capital pi
(producttext) toindicateproduct,as capitalsigma (summationtext)indicatesasum,wehave
f1·f2·f3···fn=nproductdisplay
i=1fi. (5.197)
Wedefine pn, apartialproduct,inanalogywith snthepartialsum,
pn=nproductdisplay
i=1fi (5.198)
andtheninvestigatethelimit,
limn→∞pn=P. (5.199)
IfPis finite (but not zero), we say the infinite product is convergent. If Pis infinite or
zero,theinfiniteproductislabeleddivergent.
5.11 Infinite Products 397
Sincetheproductwilldivergetoinfinityif
limn→∞fn>1 (5.200)
ortozerofor
limn→∞fn<1(and>0), (5.201)
itisconvenienttowriteourinfiniteproductsas
∞productdisplay
n=1(1+an).
Thecondition an→0 isthenanecessary(butnotsufficient)conditionforconvergence.
Theinfiniteproductmayberelatedtoaninfiniteseriesbytheobviousmethodoftaking
thelogarithm,
ln∞productdisplay
n=1(1+an)=∞summationdisplay
n=1ln(1+an). (5.202)
Amoreuseful relationshipis statedbythefollowingtheorem.
Convergence of Infinite Product
If 0≤an<1,theinfiniteproductsproducttext∞
n=1(1+an)andproducttext∞
n=1(1−an)convergeifsummationtext∞
n=1an
convergesanddivergeifsummationtext∞
n=1andiverges.
Consideringtheterm 1 +an,weseefromEq. (5.90) that
1+an≤ean. (5.203)
Thereforefor thepartialproduct pn, withsnthepartialsumofthe ai,
pn≤esn, (5.204)
andletting n→∞,
∞productdisplay
n=1(1+an)≤exp∞summationdisplay
n=1an, (5.205)
thusestablishinganupperboundfor theinfiniteproduct.
Todevelopalowerbound,wenotethat
pn=1+nsummationdisplay
i=1ai+nsummationdisplay
i=1nsummationdisplay
j=1aiaj+···≥sn, (5.206)
sinceai≥0.Hence
∞productdisplay
n=1(1+an)≥∞summationdisplay
n=1an. (5.207)
398 Chapter 5 Infinite Series
Iftheinfinitesumremainsfinite,theinfiniteproductwillalso.Iftheinfinitesumdiverges,
sowilltheinfiniteproduct.
The case ofproducttext(1−an)is complicated by the negative signs, but a proof that depends
ontheforegoingproofmaybedevelopedbynotingthatfor an<1
2(remember an→0f o r
convergence),
(1−an)≤(1+an)−1
and
(1−an)≥(1+2an)−1. (5.208)
Sine, Cosine, and Gamma Functions
Annth-order polynomial Pn(x)withnreal roots may be written as a product of nfactors
(seeSection6.4, Gauss’fundamentaltheoremof algebra):
Pn(x)=(x−x1)(x−x2)···(x−xn)=nproductdisplay
i=1(x−xi). (5.209)
In much the same way we may expect that a function with an infinite number of roots
may be written as an infinite product, one factor for each root. This is indeed the case for
thetrigonometricfunctions.Wehavetwoveryusefulinfiniteproductrepresentations,
sinx=x∞productdisplay
n=1parenleftbigg
1−x2
n2π2parenrightbigg
, (5.210)
cosx=∞productdisplay
n=1bracketleftbigg
1−4x2
(2n−1)2π2bracketrightbigg
. (5.211)
The most convenient and perhaps most elegant derivation of these two expressions is by
the use of complex variables.30By our theorem of convergence, Eqs. (5.210) and (5.211)
areconvergentforallfinitevaluesof x.Specifically,fortheinfiniteproductforsin x, an=
x2/n2π2,
∞summationdisplay
n=1an=x2
π2∞summationdisplay
n=1n−2=x2
π2ζ(2)=x2
6(5.212)
byExercise5.9.6. Theseries correspondingtoEq. (5.211)behavesinasimilarmanner.
Equation(5.210)leadstotwointerestingresults. First, if weset x=π/2,weobtain
1=π
2∞productdisplay
n=1bracketleftbigg
1−1
(2n)2bracketrightbigg
=π
2∞productdisplay
n=1bracketleftbigg(2n)2−1
(2n)2bracketrightbigg
. (5.213)
30SeeEqs. (7.25) and (7.26).
5.11 Infinite Products 399
Solvingfor π/2,wehave
π
2=∞productdisplay
n=1bracketleftbigg(2n)2
(2n−1)(2n+1)bracketrightbigg
=2·2
1·3·4·4
3·5·6·6
5·7···, (5.214)
whichisWallis’ famousformulafor π/2.
Thesecondresultinvolvesthegammaorfactorialfunction(Section8.1).Onedefinition
ofthegammafunctionis
Ŵ(x)=bracketleftbigg
xeγx∞productdisplay
r=1parenleftbigg
1+x
rparenrightbigg
e−x/rbracketrightbigg−1
, (5.215)
whereγis the usual Euler–Mascheroni constant (compare Section 5.2). If we take the
productof Ŵ(x)andŴ(−x), Eq.(5.215) leadsto
Ŵ(x)Ŵ(−x)=−bracketleftbigg
xeγx∞productdisplay
r=1parenleftbigg
1+x
rparenrightbigg
e−x/rxe−γx∞productdisplay
r=1parenleftbigg
1−x
rparenrightbigg
ex/rbracketrightbigg−1
=−1
x2∞productdisplay
r=1parenleftbigg
1−x2
r2parenrightbigg−1
. (5.216)
UsingEq. (5.210)with xreplacedby πx, weobtain
Ŵ(x)Ŵ(−x)=−π
xsinπx. (5.217)
AnticipatingarecurrencerelationdevelopedinSection8.1,wehave −xŴ(−x)=Ŵ(1−x).
Equation(5.217)maybewrittenas
Ŵ(x)Ŵ(1−x)=π
sinπx. (5.218)
Thiswillbeusefulintreatingthegammafunction(Chapter8).
Strictly speaking, we should check the range of xfor which Eq. (5.215) is convergent.
Clearly, individual factors will vanish for x=0,−1,−2,....The proof that the infinite
productconvergesforallother(finite)valuesof xis leftasExercise5.11.9.
Theseinfiniteproductshaveavarietyofusesinmathematics.However,becauseofrather
slowconvergence,theyarenotsuitablefor precisenumericalwork inphysics.
Exercises
5.11.1 Using
ln∞productdisplay
n=1(1±an)=∞summationdisplay
n=1ln(1±an)
andtheMaclaurinexpansionofln (1±an),showthattheinfiniteproductproducttext∞
n=1(1±an)
convergesordivergeswiththeinfiniteseriessummationtext∞
n=1an.
400 Chapter 5 Infinite Series
5.11.2 Aninfiniteproductappearsintheform
∞productdisplay
n=1parenleftbigg1+a/n
1+b/nparenrightbigg
,
whereaandbareconstants.Showthatthisinfiniteproductconvergesonlyif a=b.
5.11.3 Show that the infinite product representations of sin xand cosxare consistent with the
identity 2sin xcosx=sin2x.
5.11.4 Determinethelimittowhich
∞productdisplay
n=2parenleftbigg
1+(−1)n
nparenrightbigg
converges.
5.11.5 Showthat
∞productdisplay
n=2bracketleftbigg
1−2
n(n+1)bracketrightbigg
=1
3.
5.11.6 Provethat
∞productdisplay
n=2parenleftbigg
1−1
n2parenrightbigg
=1
2.
5.11.7 Usingtheinfiniteproductrepresentationsof sin x,showthat
xcotx=1−2∞summationdisplay
m,n=1parenleftbiggx
nπparenrightbigg2m
,
hencethattheBernoullinumber
B2n=(−1)n−12(2n)!
(2π)2nζ(2n).
5.11.8 VerifytheEuleridentity
∞productdisplay
p=1parenleftbig
1+zpparenrightbig
=∞productdisplay
q=1parenleftbig
1−z2q−1parenrightbig−1,|z|<1.
5.11.9 Show thatproducttext∞
r=1(1+x/r)e−x/rconverges for all finite x(except for the zeros of
1+x/r).
Hint.Writethe nthfactoras 1+an.
5.11.10 Calculate cos xfrom its infinite product representation, Eq. (5.211), using (a) 10,
(b) 100, and (c) 1000 factors in the product. Calculate the absolute error. Note how
slowly the partial products converge–making the infinite product quite unsuitable for
precisenumericalwork.
ANS.For1000factors, cos π=−1.00051.
5.11 Additional Readings 401
AdditionalReadings
Thetopic of infinite series is treatedinmany texts on advancedcalculus.
Bender, C. M., and S. Orszag, Advanced Mathematical Methods for Scientists and Engineers .N e wY o r k :
McGraw-Hill(1978). Particularly recommended for methods of acceleratingconvergence.
Davis, H. T., Tables of Higher Mathematical Functions . Bloomington, IN: Principia Press (1935). Volume II
contains extensive information on Bernoulli numbers and polynomials.
Dingle,R. B., Asymptotic Expansions: Their Derivation and Interpretation . NewYork: AcademicPress (1973).
Galambos, J., Representations of RealNumbers by Infinite Series .Berlin: Springer (1976).
Gradshteyn, I. S., and I. M. Ryzhik, Table of Integrals, Series and Products . Corrected and enlarged 6th edition
prepared byAlan Jeffrey. NewYork: AcademicPress (2000).
Hamming, R. W., Numerical Methods for Scientists and Engineers .Reprinted, NewYork: Dover (1987).
Hansen,E., ATableofSeriesand Products. EnglewoodCliffs, NJ:Prentice-Hall(1975). Atremendous compila-
tion of series and products.
Hardy, G. H., Divergent Series. Oxford: Clarendon Press (1956), 2nd ed., Chelsea (1992). The standard, com-
prehensive work on methods of treating divergent series. Hardy includes instructive accounts of the gradual
development ofthe concepts ofconvergence anddivergence.
Jeffrey, A., Handbook of Mathematical Formulas and Integrals . San Diego: AcademicPress (1995).
Knopp, K., Theory and Application of Infinite Series. London: Blackie and Son (2nd ed.); New York: Hafner
(1971). Reprinted: A. K. Peters Classics (1997). This is a thorough, comprehensive, and authoritative work
that covers infinite series and products. Proofs of almost all of the statements not proved in Chapter 5 will be
found in this book.
Mangulis, V., Handbook of Series for Scientists and Engineers . New York: Academic Press (1965). A most
convenientandusefulcollectionofseries.Includesalgebraicfunctions,Fourierseries,andseriesofthespecial
functions: Bessel,Legendre, andso on.
Olver, F. W. J., Asymptotics and Special Functions . New York: Academic Press (1974). A detailed, readable
development of asymptotic theory. Considerable attention is paid to error bounds for use in computation.
Rainville, E. D., Infinite Series . New York: Macmillan (1967). A readable and useful account of series constants
and functions.
Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering , 2nd ed. New York:
McGraw-Hill (1966). A long Chapter 2 (101 pages) presents infinite series in a thorough but very readable
form.Extensionstothesolutionsofdifferentialequations,tocomplexseries,andtoFourierseriesareincluded.
This page intentionally left blank
CHAPTER 6
FUNCTIONS OF A COMPLEX
VARIABLE I
ANALYTIC PROPERTIES ,MAPPING
The imaginary numbersarea wonderful flight of God’sspirit;
they arealmost anamphibian between being and not being.
GOTTFRIED WILHELM VON LEIBNIZ, 1702
Weturnnowtoastudyoffunctionsofacomplexvariable.Inthisareawedevelopsome
of the most powerful and widely useful tools in all of analysis. To indicate, at least partly,
whycomplexvariablesareimportant,wementionbrieflyseveralareasof application.
1. Formanypairs offunctions uandv,bothuandvsatisfyLaplace’sequation,
∇2ψ=∂2ψ(x,y)
∂x2+∂2ψ(x,y)
∂y2=0.
Henceeither uorvmaybeusedtodescribeatwo-dimensionalelectrostaticpotential.The
otherfunction,whichgivesafamilyofcurvesorthogonaltothoseofthefirstfunction,may
thenbeusedtodescribetheelectricfield E.Asimilarsituationholdsforthehydrodynamics
ofanidealfluidinirrotationalmotion.Thefunction umightdescribethevelocitypotential,
whereasthefunction vwouldthenbethestreamfunction.
In many cases in which the functions uandvare unknown, mapping or transforming
in the complex plane permits us to create a coordinate system tailored to the particular
problem.
2. In Chapter 9 we shall see that the second-order differential equations of interest in
physicsmaybesolvedbypowerseries.Thesamepowerseriesmaybeusedinthecomplex
plane to replace xby the complex variable z. The dependence of the solution f(z)at a
givenz0onthebehaviorof f(z)elsewheregivesusgreaterinsightintothebehaviorofour
403
404 Chapter 6 Functions of a Complex Variable I
solution and a powerful tool (analytic continuation) for extending the region in which the
solutionisvalid.
3.Thechangeofaparameter kfromrealtoimaginary, k→ik,transformstheHelmholtz
equation into the diffusion equation. The same change transforms the Helmholtz equa-
tionsolutions(BesselandsphericalBessel functions)intothediffusionequationsolutions
(modifiedBesselandmodifiedsphericalBesselfunctions).
4. Integralsinthecomplexplanehavea widevarietyof usefulapplications:
•Evaluatingdefiniteintegrals;
•Invertingpowerseries;
•Forminginfiniteproducts;
•Obtainingsolutionsofdifferentialequationsforlargevaluesofthevariable(asymptotic
solutions);
•Investigatingthestabilityof potentiallyoscillatorysystems;
•Invertingintegraltransforms.
5.Manyphysicalquantitiesthatwereoriginallyrealbecomecomplexasasimplephys-
ical theory is made more general. The real index of refraction of light becomes a complex
quantity when absorption is included. The real energy associated with an energy level be-
comescomplexwhenthefinitelifetimeofthelevelis considered.
6.1 C OMPLEX ALGEBRA
Acomplexnumberisnothingmorethananorderedpairoftworealnumbers, (a,b).Sim-
ilarly,acomplexvariableis anorderedpairoftworealvariables,1
z≡(x,y). (6.1)
The ordering is significant. In general (a,b)is not equal to (b,a)and(x,y)is not equal
to(y,x). As usual, we continue writing a real number (x,0)simply as x, and we call
i≡(0,1)theimaginaryunit.
Allourcomplexvariableanalysiscanbedevelopedintermsoforderedpairsofnumbers
(a,b),variables (x,y), andfunctions (u(x,y),v(x,y) ).
We now define addition of complex numbers in terms of their Cartesian components
as
z1+z2=(x1,y1)+(x2,y2)=(x1+x2,y1+y2), (6.2a)
that is, two-dimensional vector addition. In Chapter 1 the points in the xy-plane are
identified with the two-dimensional displacement vector r=ˆxx+ˆyy. As a result, two-
dimensional vector analogs can be developed for much of our complex analysis. Exer-
cise6.1.2is onesimpleexample;Cauchy’stheorem,Section6.3,is another.
Multiplication of complexnumbersisdefinedas
z1z2=(x1,y1)·(x2,y2)=(x1x2−y1y2,x1y2+x2y1). (6.2b)
1This is precisely how acomputer does complex arithmetic.
6.1 Complex Algebra 405
UsingEq.(6.2b)weverifythat i2=(0,1)·(0,1)=(−1,0)=−1,sowecanalsoidentify
i=√
−1, asusualandfurther rewriteEq.(6.1) as
z=(x,y)=(x,0)+(0,y)=x+(0,1)·(y,0)=x+iy. (6.2c)
Clearly, the iis not necessary here but it is convenient. It serves to keep pairs in order—
somewhatliketheunitvectorsof Chapter1.2
Permanence of Algebraic Form
All our elementary functions, ez,sinz, and so on, can be extended into the complex plane
(compare Exercise 6.1.9). For instance, they can be defined by power-series expansions,
suchas
ez=1+z
1!+z2
2!+···=∞summationdisplay
n=0zn
n!(6.3)
for the exponential. Such definitions agree with the real variable definitions along the real
x-axis and extend the corresponding real functions into the complex plane. This result is
oftencalled permanenceofthealgebraicform .
Itisconvenienttoemployagraphicalrepresentationofthecomplexvariable.Byplotting
x—therealpartof z—astheabscissaand y—theimaginarypartof z—astheordinate,
we have the complex plane, or Argand plane, shown in Fig. 6.1. If we assign specific
valuesto xandy,thenzcorrespondstoapoint (x,y)intheplane.Intermsoftheordering
mentionedbefore,itisobviousthatthepoint (x,y)doesnotcoincidewiththepoint (y,x)
exceptfor thespecialcaseof x=y.Further, fromFig.6.1wemaywrite
x=rcosθ, y=rsinθ (6.4a)
FIGURE 6.1Complex
plane—Arganddiagram.
2Thealgebra of complex numbers, (a,b), is isomorphic with that of matrices of the form
parenleftbiggab
−baparenrightbigg
(compare Exercise3.2.4).
406 Chapter 6 Functions of a Complex Variable I
and
z=r(cosθ+isinθ). (6.4b)
Using a result that is suggested (but not rigorously proved)3by Section 5.6 and Exer-
cise5.6.1,wehavetheusefulpolarrepresentation
z=r(cosθ+isinθ)=reiθ. (6.4c)
In order to prove this identity, we use i3=−i, i4=1,...in the Taylor expansion of the
exponentialandtrigonometricfunctionsandseparateevenandoddpowersin
eiθ=∞summationdisplay
n=0(iθ)n
n!=∞summationdisplay
ν=0(iθ)2ν
(2ν)!+∞summationdisplay
ν=0(iθ)2ν+1
(2ν+1)!
=∞summationdisplay
ν=0(−1)νθ2ν
(2ν)!+i∞summationdisplay
ν=0(−1)νθ2ν+1
(2ν+1)!=cosθ+isinθ.
Forthespecialvalues θ=π/2 andθ=π,weobtain
eiπ/2=cosπ
2+isinπ
2=i, eiπ=cos(π)=−1,
intriguingconnectionsbetween e,i,andπ.Moreover,theexponentialfunction eiθisperi-
odicwithperiod 2 π,just like sin θand cosθ.
Inthisrepresentation riscalledthe modulus ormagnitude ofz(r=|z|=(x2+y2)1/2)
andtheangle θ(=tan−1(y/x))islabeledtheargumentor phaseofz.(Notethatthearctan
function tan−1(y/x)hasinfinitelymanybranches.)
Thechoiceofpolarrepresentation,Eq.(6.4c),orCartesianrepresentation,Eqs.(6.1)and
(6.2c),isamatterofconvenience.Additionandsubtractionofcomplexvariablesareeasier
in the Cartesian representation, Eq. (6.2a). Multiplication, division, powers, and roots are
easiertohandleinpolarform, Eq. (6.4c).
Analytically or graphically, using the vector analogy, we may show that the modulus of
thesumoftwocomplexnumbersisnogreaterthanthesumofthemoduliandnolessthan
thedifference, Exercise6.1.3,
|z1|−|z2|≤|z1+z2|≤|z1|+|z2|. (6.5)
Becauseof thevectoranalogy,theseare calledthe triangleinequalities.
Using the polar form, Eq. (6.4c), we find that the magnitude of a product is the product
ofthemagnitudes:
|z1·z2|=|z1|·|z2|. (6.6)
Also,
arg(z1·z2)=argz1+argz2. (6.7)
3Strictly speaking, Chapter 5 was limited to real variables. The development of power-series expansions for complex functions
is takenup in Section6.5(Laurent expansion).
6.1 Complex Algebra 407
FIGURE 6.2Thefunction w(z)=u(x,y)+iv(x,y)mapspointsinthe xy-plane
intopointsinthe uv-plane.
Fromourcomplexvariable zcomplexfunctions f(z)orw(z)maybeconstructed.These
complexfunctionsmaythenberesolvedintorealandimaginaryparts,
w(z)=u(x,y)+iv(x,y), (6.8)
inwhichtheseparatefunctions u(x,y)andv(x,y)arepurereal.Forexample,if f(z)=z2,
wehave
f(z)=(x+iy)2=parenleftbig
x2−y2parenrightbig
+i2xy.
Thereal part of a function f(z)will be labeled ℜf(z), whereas the imaginary part will
belabeledℑf(z).I nE q .( 6 . 8 )
ℜw(z)=Re(w)=u(x,y),ℑw(z)=Im(w)=v(x,y).
The relationship between the independent variable zand the dependent variable wis
perhaps best pictured as a mapping operation. A given z=x+iymeans a given point in
thez-plane.Thecomplexvalueof w(z)isthenapointinthe w-plane.Pointsinthe z-plane
map into points in the w-plane and curves in the z-plane map into curves in the w-plane,
asindicatedinFig. 6.2.
Complex Conjugation
In all these steps, complex number, variable, and function, the operation of replacing iby
–iis called“takingthe complexconjugate.”The complexconjugateof zis denotedby z∗,
where4
z∗=x−iy. (6.9)
4Thecomplex conjugate is often denoted by ¯zinthe mathematicalliterature.
408 Chapter 6 Functions of a Complex Variable I
FIGURE 6.3Complexconjugatepoints.
The complex variable zand its complex conjugate z∗are mirror images of each other
reflected in the x-axis, that is, inversion of the y-axis (compare Fig. 6.3). The product zz∗
leadsto
zz∗=(x+iy)(x−iy)=x2+y2=r2. (6.10)
Hence
(zz∗)1/2=|z|,
themagnitude ofz.
Functions of a Complex Variable
All the elementary functions of real variables may be extended into the complex plane—
replacing the real variable xby the complex variable z. This is an example of the analytic
continuation mentioned in Section 6.5. The extremely important relation of Eq. (6.4c) is
anillustration.Movingintothecomplexplaneopensupnewopportunitiesfor analysis.
Example 6.1.1 DEMOIVRE ’SFORMULA
If Eq.(6.4c) (setting r=1)is raisedtothe nthpower,wehave
einθ=(cosθ+isinθ)n. (6.11)
Expandingtheexponentialnowwithargument nθ,weobtain
cosnθ+isinnθ=(cosθ+isinθ)n. (6.12)
DeMoivre’sformulaisgeneratediftheright-handsideofEq.(6.12)isexpandedbythebi-
nomialtheorem;weobtaincos nθasaseriesofpowersofcos θandsinθ,Exercise6.1.6. /squaresolid
Numerous other examples of relations among the exponential, hyperbolic, and trigono-
metricfunctionsinthecomplexplaneappearintheexercises.
Occasionally there are complications. The logarithm of a complex variable may be ex-
pandedusingthepolarrepresentation
lnz=lnreiθ=lnr+iθ. (6.13a)
6.1 Complex Algebra 409
Thisisnotcomplete.Tothephaseangle, θ,wemayaddanyintegralmultipleof2 πwithout
changing z. HenceEq.(6.13a) shouldread
lnz=lnrei(θ+2nπ)=lnr+i(θ+2nπ). (6.13b)
Theparameter nmaybeanyinteger.Thismeansthatln zisamultivalued functionhaving
an infinite number of values for a single pair of real values randθ. To avoid ambiguity,
the simplest choice is n=0 and limitation of the phase to an interval of length 2 π, such
as(−π,π).5T h el i n ei nt h e z-plane that is not crossed, the negative real axis in this case,
is labeled a cut lineorbranch cut .T h ev a l u eo fl n zwithn=0 is called the principal
valueof lnz. Further discussion of these functions, including the logarithm, appears in
Section6.7.
Exercises
6.1.1 (a) Findthereciprocalof x+iy, workingentirelyintheCartesianrepresentation.
(b) Repeat part (a), working in polar form but expressing the final result in Cartesian
form.
6.1.2 The complex quantities a=u+ivandb=x+iymay also be represented as two-
dimensionalvectors a=ˆxu+ˆyv,b=ˆxx+ˆyy.Showthat
a∗b=a·b+iˆz·a×b.
6.1.3 Provealgebraicallythatforcomplexnumbers,
|z1|−|z2|≤|z1+z2|≤|z1|+|z2|.
Interpretthisresultintermsof two-dimensionalvectors.Provethat
|z−1|<vextendsinglevextendsingleradicalbig
z2−1vextendsinglevextendsingle<|z+1|,forℜ(z)>0.
6.1.4 We may define a complex conjugation operator Ksuch that Kz=z∗. Show that Kis
notalinearoperator.
6.1.5 Show that complex numbers have square roots and that the square roots are contained
inthecomplexplane.What arethesquareroots of i?
6.1.6 Showthat
(a) cosnθ=cosnθ−parenleftbign
2parenrightbig
cosn−2θsin2θ+parenleftbign
4parenrightbig
cosn−4θsin4θ−···.
(b) sinnθ=parenleftbign
1parenrightbig
cosn−1θsinθ−parenleftbign
3parenrightbig
cosn−3θsin3θ+···.
Note.Thequantitiesparenleftbign
mparenrightbig
arebinomialcoefficients:parenleftbign
mparenrightbig
=n!/[(n−m)!m!].
6.1.7 Provethat
(a)N−1summationdisplay
n=0cosnx=sin(Nx/2)
sinx/2cos(N−1)x
2,
5Thereis no standard choice of phase; the appropriate phase depends on eachproblem.
410 Chapter 6 Functions of a Complex Variable I
(b)N−1summationdisplay
n=0sinnx=sin(Nx/2)
sinx/2sin(N−1)x
2.
Theseseriesoccurintheanalysisofthemultiple-slitdiffractionpattern.Anotherappli-
cationistheanalysisoftheGibbsphenomenon,Section14.5.
Hint. Parts (a) and (b) may be combined to form a geometric series (compare Sec-
tion5.1).
6.1.8 For−1<p<1 prove that
(a)∞summationdisplay
n=0pncosnx=1−pcosx
1−2pcosx+p2,
(b)∞summationdisplay
n=0pnsinnx=psinx
1−2pcosx+p2.
Theseseries occurinthetheoryof theFabry–Perotinterferometer.
6.1.9 Assume that the trigonometric functions and the hyperbolic functions are defined for
complexargumentbytheappropriatepowerseries
sinz=∞summationdisplay
n=1,odd(−1)(n−1)/2zn
n!=∞summationdisplay
s=0(−1)sz2s+1
(2s+1)!,
cosz=∞summationdisplay
n=0,even(−1)n/2zn
n!=∞summationdisplay
s=0(−1)sz2s
(2s)!,
sinhz=∞summationdisplay
n=1,oddzn
n!=∞summationdisplay
s=0z2s+1
(2s+1)!,
coshz=∞summationdisplay
n=0,evenzn
n!=∞summationdisplay
s=0z2s
(2s)!.
(a) Showthat
isinz=sinhiz,siniz=isinhz,
cosz=coshiz,cosiz=coshz.
(b) Verifythatfamiliarfunctionalrelationssuchas
coshz=ez+e−z
2,
sin(z1+z2)=sinz1cosz2+sinz2cosz1,
stillholdinthecomplexplane.
6.1 Complex Algebra 411
6.1.10 Usingtheidentities
cosz=eiz+e−iz
2,sinz=eiz−e−iz
2i,
establishedfromcomparisonofpowerseries, showthat
(a) sin(x+iy)=sinxcoshy+icosxsinhy,
cos(x+iy)=cosxcoshy−isinxsinhy,
(b)|sinz|2=sin2x+sinh2y,|cosz|2=cos2x+sinh2y.
Thisdemonstratesthatwemayhave |sinz|,|cosz|>1 inthecomplexplane.
6.1.11 FromtheidentitiesinExercises6.1.9 and6.1.10showthat
(a) sinh (x+iy)=sinhxcosy+icoshxsiny,
cosh(x+iy)=coshxcosy+isinhxsiny,
(b)|sinhz|2=sinh2x+sin2y,|coshz|2=cosh2x+sin2y.
6.1.12 Provethat
(a)|sinz|≥|sinx|(b)|cosz|≥|cosx|.
6.1.13 Showthattheexponentialfunction ezisperiodicwithapureimaginaryperiodof 2 πi.
6.1.14 Showthat
(a) tanhz
2=sinhx+isiny
coshx+cosy, (b) cothz
2=sinhx−isiny
coshx−cosy.
6.1.15 Findallthezerosof
(a) sinz, (b) cos z, (c) sinh z, (d) cosh z.
6.1.16 Showthat
(a) sin−1z=−ilnparenleftbig
iz±radicalbig
1−z2parenrightbig
, (d) sinh−1z=lnparenleftbig
z+radicalbig
z2+1parenrightbig
,
(b) cos−1z=−ilnparenleftbig
z±radicalbig
z2−1parenrightbig
,(e) cosh−1z=lnparenleftbig
z+radicalbig
z2−1parenrightbig
,
(c) tan−1z=i
2lnparenleftbiggi+z
i−zparenrightbigg
, (f) tanh−1z=1
2lnparenleftbigg1+z
1−zparenrightbigg
.
Hint.1. Express the trigonometric and hyperbolic functions in terms of exponentials.
2.Solvefor theexponentialandthenfor theexponent.
6.1.17 Inthequantumtheoryofthephotoionizationweencountertheidentity
parenleftbiggia−1
ia+1parenrightbiggib
=expparenleftbig
−2bcot−1aparenrightbig
,
inwhich aandbarereal.Verifythisidentity.
412 Chapter 6 Functions of a Complex Variable I
6.1.18 Aplanewaveof lightof angularfrequency ωisrepresentedby
eiω(t−nx/c).
In a certain substance the simple real index of refraction nis replaced by the complex
quantityn−ik. What is the effect of kon the wave? What does kcorrespond to phys-
ically? The generalization of a quantity from real to complex form occurs frequently
in physics. Examples range from the complex Young’s modulus of viscoelastic materi-
als to the complex (optical) potential of the “cloudy crystal ball” model of the atomic
nucleus.
6.1.19 Weseethatfor theangularmomentumcomponentsdefinedinExercise2.5.14,
Lx−iLy/negationslash=(Lx+iLy)∗.
Explainwhythisoccurs.
6.1.20 Showthatthe phaseoff(z)=u+ivisequaltotheimaginarypartofthelogarithmof
f(z). Exercise8.2.13dependsonthisresult.
6.1.21 (a) Showthat elnzalwaysequals z.
(b) Showthat ln ezdoesnotalwaysequal z.
6.1.22 The infinite product representations of Section 5.11 hold when the real variable xis
replaced by the complex variable z. From this, develop infinite product representations
for
(a) sinhz, (b) cosh z.
6.1.23 Theequationof motionofamass mrelativeto arotatingcoordinatesystem is
md2r
dt2=F−mω×(ω×r)−2mparenleftbigg
ω×dr
dtparenrightbigg
−mparenleftbiggdω
dt×rparenrightbigg
.
Consider the case F=0,r=ˆxx+ˆyy, andω=ωˆz, withωconstant. Show that the
replacementof r=ˆxx+ˆyybyz=x+iyleadsto
d2z
dt2+i2ωdz
dt−ω2z=0.
Note.This ODEmaybesolvedbythesubstitution z=fe−iωt.
6.1.24 Using the complex arithmetic available in FORTRAN, write a program that will cal-
culate the complex exponential ezfrom its series expansion (definition). Calculate ez
forz=einπ/6,n=0,1,2,...,12.Tabulatethephaseangle (θ=nπ/6),ℜz,ℑz,ℜ(ez),
ℑ(ez),|ez|, andthephaseof ez.
Checkvalue. n=5,θ=2.61799,ℜ(z)=−0.86602,
ℑz=0.50000,ℜ(ez)=0.36913,ℑ(ez)=0.20166,
|ez|=0.42062,phase (ez)=0.50000.
6.1.25 UsingthecomplexarithmeticavailableinFORTRAN,calculateandtabulate ℜ(sinhz),
ℑ(sinhz),|sinhz|, andphase (sinhz)forx=0.0(0.1)1.0 andy=0.0(0.1)1.0.
6.2 Cauchy–Riemann Conditions 413
Hint.Bewareofdividingbyzerowhencalculatinganangleas anarctangent.
Checkvalue. z=0.2+0.1i,ℜ(sinhz)=0.20033,
ℑ(sinhz)=0.10184,|sinhz|=0.22473,
phase(sinhz)=0.47030.
6.1.26 RepeatExercise6.1.25for cosh z.
6.2 C AUCHY –RIEMANN CONDITIONS
Having established complex functions of a complexvariable, we now proceed to differen-
tiatethem.Thederivativeof f(z), like thatofa realfunction,is definedby
lim
δz→0f(z+δz)−f(z)
z+δz−z=lim
δz→0δf (z)
δz=df
dz=f′(z), (6.14)
provided that the limit is independent of the particular approach to the point z. For real
variables we require that the right-hand limit ( x→x0from above) and the left-hand limit
(x→x0from below)be equalfor thederivative df(x)/dx toexistat x=x0.No w ,wi t h z
(orz0)somepointinaplane,ourrequirementthatthelimitbeindependentofthedirection
ofapproachis veryrestrictive.
Considerincrements δxandδyofthevariables xandy,respectively.Then
δz=δx+iδy. (6.15)
Also,
δf=δu+iδv, (6.16)
sothat
δf
δz=δu+iδv
δx+iδy. (6.17)
Let us take the limit indicated by Eq. (6.14) by two different approaches, as shown in
Fig.6.4.First, with δy=0,weletδx→0. Equation(6.14) yields
lim
δz→0δf
δz=lim
δx→0parenleftbiggδu
δx+iδv
δxparenrightbigg
=∂u
∂x+i∂v
∂x, (6.18)
FIGURE 6.4Alternate
approachesto z0.
414 Chapter 6 Functions of a Complex Variable I
assuming the partial derivatives exist. For a second approach, we set δx=0 and then let
δy→0.Thisleadsto
lim
δz→0δf
δz=lim
δy→0parenleftbigg
−iδu
δy+δv
δyparenrightbigg
=−i∂u
∂y+∂v
∂y. (6.19)
If we are to have a derivative df/dz, Eqs. (6.18) and (6.19) must be identical. Equating
realpartstorealpartsandimaginarypartstoimaginaryparts(likecomponentsofvectors),
weobtain
∂u
∂x=∂v
∂y,∂u
∂y=−∂v
∂x. (6.20)
Thesearethefamous Cauchy–Riemann conditions.TheywerediscoveredbyCauchyand
used extensively by Riemann in his theory of analytic functions. These Cauchy–Riemann
conditions are necessary for the existence of a derivative of f(z); that is, if df/dzexists,
theCauchy–Riemannconditionsmusthold.
Conversely, if the Cauchy–Riemann conditions are satisfied and the partial derivatives
ofu(x,y)andv(x,y)are continuous, the derivative df/dzexists. This may be shown by
writing
δf=parenleftbigg∂u
∂x+i∂v
∂xparenrightbigg
δx+parenleftbigg∂u
∂y+i∂v
∂yparenrightbigg
δy. (6.21)
The justification for this expression depends on the continuity of the partial derivatives of
uandv.Dividingby δz,weha v e
δf
δz=(∂u/∂x+i(∂v/∂x))δx+(∂u/∂y+i(∂v/∂y))δy
δx+iδy
=(∂u/∂x+i(∂v/∂x))+(∂u/∂y+i(∂v/∂y))δy/δx
1+i(δy/δx). (6.22)
Ifδf/δzistohaveauniquevalue,thedependenceon δy/δxmustbeeliminated.Apply-
ingtheCauchy–Riemannconditionstothe yderivatives,weobtain
∂u
∂y+i∂v
∂y=−∂v
∂x+i∂u
∂x. (6.23)
SubstitutingEq. (6.23)intoEq. (6.22), wemaycanceloutthe δy/δxdependenceand
δf
δz=∂u
∂x+i∂v
∂x, (6.24)
which shows that lim δf/δzis independent of the direction of approach in the complex
plane as long as the partial derivatives are continuous. Thus,df
dzexists and fis analytic
atz.
It is worthwhile noting that the Cauchy–Riemann conditions guarantee that the curves
u=c1willbe orthogonalto thecurves v=c2(compareSection2.1). This is fundamental
in application to potential problems in a variety of areas of physics. If u=c1is a line of
6.2 Cauchy–Riemann Conditions 415
electric force, then v=c2is an equipotentialline (surface), and vice versa. To see this, let
uswritetheCauchy–Riemannconditionsasa productofratiosof partialderivatives,
ux
uy·vx
vy=−1, (6.25)
withtheabbreviations
∂u
∂x≡ux,∂u
∂y≡uy,∂v
∂x≡vx,∂v
∂y≡vy.
Now recall the geometric meaning of −ux/uyas the slope of the tangent of each curve
u(x,y)=const. and similarly for v(x,y)=const. This means that the u=const. and
v=const. curvesaremutuallyorthogonalateachintersection.Alternatively,
uxdx+uydy=0=vydx−vxdy
saysthat,if (dx,dy)istangenttothe u-curve,thentheorthogonal (−dy,dx)istangentto
thev-curve at the intersection point, z=(x,y). Or equivalently, uxvx+uyvy=0 implies
that thegradient vectors (ux,uy)and(vx,vy)are perpendicular . A further implication
forpotentialtheoryis developedinExercise6.2.1.
Analytic Functions
Finally, if f(z)is differentiable at z=z0and in some small region around z0, we say that
f(z)isanalytic6atz=z0.Iff(z)isanalyticeverywhereinthe(finite)complexplane,we
callitanentirefunction.Ourtheoryofcomplexvariableshereisoneofanalyticfunctions
of a complex variable, which points up the crucial importance of the Cauchy–Riemann
conditions. The concept of analyticity carried on in advanced theories of modern physics
plays a crucial role in dispersion theory (of elementary particles). If f′(z)does not exist
atz=z0, thenz0is labeled a singular point and consideration of it is postponed until
Section6.6.
ToillustratetheCauchy–Riemannconditions,considertwoverysimpleexamples.
Example 6.2.1 z2ISANALYTIC
Letf(z)=z2.Thentherealpart u(x,y)=x2−y2andtheimaginarypart v(x,y)=2xy.
FollowingEq.(6.20),
∂u
∂x=2x=∂v
∂y,∂u
∂y=−2y=−∂v
∂x.
We see that f(z)=z2satisfies the Cauchy–Riemann conditions throughout the complex
plane. Since the partial derivatives are clearly continuous, we conclude that f(z)=z2is
analytic. /squaresolid
6Some writers use the term holomorphic orregular.
416 Chapter 6 Functions of a Complex Variable I
Example 6.2.2 z∗ISNOTANALYTIC
Letf(z)=z∗.N o wu=xandv=−y. Applying the Cauchy–Riemann conditions, we
obtain
∂u
∂x=1/negationslash=∂v
∂y=−1.
TheCauchy–Riemannconditionsarenotsatisfiedand f(z)=z∗isnotananalyticfunction
ofz. It is interesting to note that f(z)=z∗is continuous, thus providing an example of a
functionthatis everywherecontinuousbutnowheredifferentiableinthecomplexplane.
The derivative of a real function of a real variable is essentially a local characteristic, in
thatitprovidesinformationaboutthefunctiononlyinalocalneighborhood—forinstance,
as a truncated Taylor expansion. The existence of a derivative of a function of a complex
variablehasmuchmorefar-reachingimplications.Therealandimaginarypartsofouran-
alytic function must separately satisfy Laplace’s equation. This is Exercise 6.2.1. Further,
our analytic function is guaranteed derivatives of all orders, Section 6.4. In this sense the
derivative not only governs the local behavior of the complex function, but controls the
distantbehavioras well. /squaresolid
Exercises
6.2.1 The functions u(x,y)andv(x,y)are the real and imaginary parts, respectively, of an
analyticfunction w(z).
(a) Assumingthattherequiredderivativesexist,showthat
∇2u=∇2v=0.
Solutions of Laplace’s equation such as u(x,y)andv(x,y)are called harmonic
functions.
(b) Showthat
∂u
∂x∂u
∂y+∂v
∂x∂v
∂y=0,
andgiveageometricinterpretation.
Hint.ThetechniqueofSection1.6allowsyoutoconstructvectorsnormaltothecurves
u(x,y)=ciandv(x,y)=cj.
6.2.2 Showwhetheror notthefunction f(z)=ℜ(z)=xis analytic.
6.2.3 Having shown that the real part u(x,y)and the imaginary part v(x,y)of an analytic
functionw(z)each satisfy Laplace’s equation, show that u(x,y)andv(x,y)cannot
both have either a maximum or a minimum in the interior of any region in which
w(z)isanalytic.(Theycanhavesaddlepointsonly.)
6.2 Cauchy–Riemann Conditions 417
6.2.4 LetA=∂2w/∂x2,B=∂2w/∂x∂y,C=∂2w/∂y2. From the calculus of functions of
twovariables, w(x,y),weha v ea saddlepoint if
B2−AC >0.
Withf(z)=u(x,y)+iv(x,y), apply the Cauchy–Riemann conditions and show that
neitheru(x,y)norv(x,y)has a maximum or a minimum in a finite region of the
complexplane.(SeealsoSection7.3.)
6.2.5 Findtheanalyticfunction
w(z)=u(x,y)+iv(x,y)
if (a)u(x,y)=x3−3xy2,( b)v(x,y)=e−ysinx.
6.2.6 If there is some common region in which w1=u(x,y)+iv(x,y)andw2=w∗
1=
u(x,y)−iv(x,y)arebothanalytic,provethat u(x,y)andv(x,y)areconstants.
6.2.7 Thefunction f(z)=u(x,y)+iv(x,y)is analytic.Showthat f∗(z∗)is alsoanalytic.
6.2.8 Usingf(reiθ)=R(r,θ)ei/Phi1(r,θ), in which R(r,θ)and/Phi1(r,θ)are differentiable real
functions of randθ, show that the Cauchy–Riemann conditions in polar coordinates
become
(a)∂R
∂r=R
r∂/Theta1
∂θ,(b)1
r∂R
∂θ=−R∂/Theta1
∂r.
Hint.Setupthederivativefirstwith δzradialandthenwith δztangential.
6.2.9 AsanextensionofExercise6.2.8showthat /Theta1(r,θ)satisfiesLaplace’sequationinpolar
coordinates. Equation (2.35) (without the final term and set to zero) is the Laplacian in
polarcoordinates.
6.2.10 Two-dimensional irrotational fluid flow is conveniently described by a complex poten-
tialf(z)=u(x,v)+iv(x,y). We label the real part, u(x,y), the velocity potential
and the imaginary part, v(x,y), the stream function. The fluid velocity Vis given by
V=∇u.Iff(z)is analytic,
(a) Showthat df/dz=Vx−iVy;
(b) Showthat ∇·V=0 (nosourcesor sinks);
(c) Showthat ∇×V=0 (irrotational,nonturbulentflow).
6.2.11 AproofoftheSchwarzinequality(Section10.4)involvesminimizinganexpression,
f=ψaa+λψab+λ∗ψ∗
ab+λλ∗ψbb≥0.
Theψareintegralsofproductsoffunctions; ψaaandψbbarereal,ψabiscomplexand
λisacomplexparameter.
(a) Differentiatetheprecedingexpressionwithrespectto λ∗,treating λasanindepen-
dent parameter, independent of λ∗. Show that setting the derivative ∂f/∂λ∗equal
tozeroyields
λ=−ψ∗
ab
ψbb.
418 Chapter 6 Functions of a Complex Variable I
(b) Showthat ∂f/∂λ=0 leadstothesameresult.
(c) Let λ=x+iy,λ∗=x−iy. Set thexandyderivatives equal to zero and show
thatagain
λ=−ψ∗
ab
ψbb.
Thisindependenceof λandλ∗appearsagaininSection17.7.
6.2.12 The function f(z)is analytic. Show that the derivative of f(z)with respect to z∗does
notexistunless f(z)is aconstant.
Hint.Usethechainruleandtake x=(z+z∗)/2,y=(z−z∗)/2i.
Note.Thisresultemphasizesthatouranalyticfunction f(z)isnotjustacomplexfunc-
tionof tworealvariables xandy.It isafunctionofthecomplexvariable x+iy.
6.3 C AUCHY ’SINTEGRAL THEOREM
Contour Integrals
With differentiation under control, we turn to integration. The integral of a complex vari-
ableoveracontourinthecomplexplanemaybedefinedincloseanalogytothe(Riemann)
integralofarealfunctionintegratedalongthereal x-axis.
Wedividethecontourfrom z0toz′
0intonintervalsbypicking n−1intermediatepoints
z1,z2,...onthecontour(Fig.6.5). Considerthesum
Sn=nsummationdisplay
j=1f(ζj)(zj−zj−1), (6.26)
FIGURE 6.5Integrationpath.
6.3 Cauchy’s Integral Theorem 419
whereζjis apointonthecurvebetween zjandzj−1.No wl et n→∞with
|zj−zj−1|→0
for allj. If the lim n→∞Snexists and is independent of the details of choosing the points
zjandζj, then
limn→∞nsummationdisplay
j=1f(ζj)(zj−zj−1)=integraldisplayz′
0
z0f(z)dz. (6.27)
Theright-handsideofEq.(6.27)iscalledthecontourintegralof f(z)(alongthespecified
contourCfromz=z0toz=z′
0).
The preceding developmentof the contour integral is closely analogous to the Riemann
integral of a real function of a real variable. As an alternative, the contour integral may be
definedby
integraldisplayz2
z1f(z)dz=integraldisplayx2,y2
x1,y1bracketleftbig
u(x,y)+iv(x,y)bracketrightbig
[dx+idy]
=integraldisplayx2,y2
x1,y1bracketleftbig
u(x,y)dx−v(x,y)dybracketrightbig
+iintegraldisplayx2,y2
x1,y1bracketleftbig
v(x,y)dx+u(x,y)dybracketrightbig
with the path joining (x1,y1)and(x2,y2)specified. This reduces the complex integral to
thecomplexsumofrealintegrals.Itissomewhatanalogoustothereplacementofavector
integralbythevectorsumof scalarintegrals,Section1.10.
An important example is the contour integralintegraltext
Czndz, whereCis a circle of radius
r>0 around the origin z=0 in the positive mathematical sense (counterclockwise). In
polar coordinates of Eq. (6.4c) we parameterize the circle as z=reiθanddz=ireiθdθ.
Forn/negationslash=−1,naninteger,wethenobtain
1
2πiintegraldisplay
Czndz=rn+1
2πintegraldisplay2π
0expbracketleftbig
i(n+1)θbracketrightbig
dθ
=bracketleftbig
2πi(n+1)bracketrightbig−1rn+1bracketleftbig
ei(n+1)θbracketrightbigvextendsinglevextendsingle2π
0=0 (6.27a)
because 2 πis aperiodof ei(n+1)θ, whilefor n=−1
1
2πiintegraldisplay
Cdz
z=1
2πintegraldisplay2π
0dθ=1, (6.27b)
againindependentof r.
Alternatively, we can integrate around a rectangle with the corners z1,z2,z3,z4to
obtainfor n/negationslash=−1
integraldisplay
zndz=zn+1
n+1vextendsinglevextendsinglevextendsinglevextendsinglez2
z1+zn+1
n+1vextendsinglevextendsinglevextendsinglevextendsinglez3
z2+zn+1
n+1vextendsinglevextendsinglevextendsinglevextendsinglez4
z3+zn+1
n+1vextendsinglevextendsinglevextendsinglevextendsinglez1
z4=0,
because each corner point appears once as an upper and a lower limit that cancel. For
n=−1thecorrespondingrealpartsofthelogarithmscancelsimilarly,buttheirimaginary
partsinvolvetheincreasingargumentsofthepointsfrom z1toz4and,whenwecomeback
to the first corner z1, its argument has increased by 2 πdue to the multivaluedness of the
420 Chapter 6 Functions of a Complex Variable I
logarithm, so 2 πiis left over as the value of the integral. Thus, the value of the integral
involving a multivalued function must be that which is reached in a continuous fash-
ion on the path being taken . These integrals are examples of Cauchy’s integral theorem,
whichweconsiderinthenextsection.
Stokes’ Theorem Proof
Cauchy’sintegraltheoremisthefirstoftwobasictheoremsinthetheoryofthebehaviorof
functions of a complex variable. First, we offer a proof under relatively restrictive condi-
tions—conditionsthatareintolerabletothemathematiciandevelopingabeautifulabstract
theorybutthatare usuallysatisfiedinphysicalproblems.
If a function f(z)is analytic, that is, if its partial derivatives are continuous throughout
somesimply connected region R,7for every closed path C(Fig. 6.6) in R, and if it is
single-valued(assumedfor simplicityhere),thelineintegralof f(z)aroundCiszero, or
integraldisplay
Cf(z)dz=contintegraldisplay
Cf(z)dz=0. (6.27c)
Recall that in Section 1.13 such a function f(z), identified as a force, was labeled conser-
vative. The symbolcontintegraltext
is used to emphasize that the path is closed. Note that the interior
of the simply connected region bounded by a contour is that region lying to the left when
moving in the direction implied by the contour; as a rule, a simply connected region is
boundedbyasingleclosedcurve.
InthisformtheCauchyintegraltheoremmaybeprovedbydirectapplicationofStokes’
theorem(Section1.12). With f(z)=u(x,y)+iv(x,y)anddz=dx+idy,
contintegraldisplay
Cf(z)dz=contintegraldisplay
C(u+iv)(dx+idy)
=contintegraldisplay
C(udx−vdy)+icontintegraldisplay
(vdx+udy). (6.28)
ThesetwolineintegralsmaybeconvertedtosurfaceintegralsbyStokes’theorem,aproce-
dure that is justified if the partial derivatives are continuous within C. In applying Stokes’
theorem,notethatthefinaltwointegralsofEq. (6.28)arereal. Using
V=ˆxVx+ˆyVy,
Stokes’theoremsays that
contintegraldisplay
C(Vxdx+Vydy)=integraldisplayparenleftbigg∂Vy
∂x−∂Vx
∂yparenrightbigg
dxdy. (6.29)
Forthefirst integralinthelast partofEq. (6.28)let u=Vxandv=−Vy.8Then
7Any closed simple curve (one that does not intersectitself) inside asimply connectedregion or domain may becontracted toa
singlepointthatstillbelongstotheregion.Ifaregionisnotsimplyconnected,itiscalledmultiplyconnected.Asanexampleof
amultiply connected region, consider the z-plane with the interior of theunit circle excluded .
8In the proof of Stokes’ theorem, Section 1.12, VxandVyare any two functions (with continuous partial derivatives).
6.3 Cauchy’s Integral Theorem 421
FIGURE 6.6Aclosedcontour C
withinasimplyconnectedregion R.
contintegraldisplay
C(udx−vdy)=contintegraldisplay
C(Vxdx+Vydy)
=integraldisplayparenleftbigg∂Vy
∂x−∂Vx
∂yparenrightbigg
dxdy=−integraldisplayparenleftbigg∂v
∂x+∂u
∂yparenrightbigg
dxdy. (6.30)
For the second integral on the right side of Eq. (6.28) we let u=Vyandv=Vx.U s i n g
Stokes’theoremagain,weobtain
contintegraldisplay
(vdx+udy)=integraldisplayparenleftbigg∂u
∂x−∂v
∂yparenrightbigg
dxdy. (6.31)
On application of the Cauchy–Riemann conditions, which must hold, since f(z)is as-
sumedanalytic,eachintegrandvanishesand
contintegraldisplay
f(z)dz=−integraldisplayparenleftbigg∂v
∂x+∂u
∂yparenrightbigg
dxdy+iintegraldisplayparenleftbigg∂u
∂x−∂v
∂yparenrightbigg
dxdy=0.(6.32)
Cauchy–Goursat Proof
ThiscompletestheproofofCauchy’sintegraltheorem.However,theproofismarredfrom
atheoreticalpointofviewbytheneedforcontinuityofthefirstpartialderivatives.Actually,
as shown by Goursat, this conditionis not necessary. An outlineof the Goursat proof is as
follows. We subdivide the region inside the contour Cinto a network of small squares, as
indicatedinFig.6.7. Thencontintegraldisplay
Cf(z)dz=summationdisplay
jcontintegraldisplay
Cjf(z)dz, (6.33)
all integrals along interior lines canceling out. To estimate thecontintegraltext
Cjf(z)dz, we construct
thefunction
δj(z,zj)=f(z)−f(zj)
z−zj−df(z)
dzvextendsinglevextendsinglevextendsinglevextendsingle
z=zj, (6.34)
withzjan interior point of the jth subregion. Note that [f(z)−f(zj)]/(z−zj)is an
approximation to the derivative at z=zj. Equivalently, we may note that if f(z)had
422 Chapter 6 Functions of a Complex Variable I
FIGURE 6.7Cauchy–Goursatcontours.
aTaylorexpansion(whichwehavenotyetproved),then δj(z,zj)wouldbeoforder z−zj,
approaching zero as the network was made finer. But since f′(zj)exists, that is, is finite,
wemaymake
vextendsinglevextendsingleδj(z,zj)vextendsinglevextendsingle<ε, (6.35)
whereεis an arbitrarily chosen small positive quantity. Solving Eq. (6.34) for f(z)and
integratingaround Cj, weobtain
contintegraldisplay
Cjf(z)dz=contintegraldisplay
Cj(z−zj)δj(z,zj)dz, (6.36)
theintegralsoftheothertermsvanishing.9WhenEqs.(6.35)and(6.36)arecombined,one
showsthatvextendsinglevextendsinglevextendsinglevextendsinglesummationdisplay
jcontintegraldisplay
Cjf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle<Aε, (6.37)
whereAis a term of the order of the area of the enclosed region. Since εis arbitrary, we
letε→0 andconcludethatif afunction f(z)isanalyticonandwithinaclosedpath C,
contintegraldisplay
Cf(z)dz=0. (6.38)
Detailsoftheproofofthissignificantlymoregeneralandmorepowerfulformcanbefound
in Churchill in the Additional Readings. Actually we can still prove the theorem for f(z)
analyticwithintheinteriorof Candonlycontinuouson C.
The consequence of the Cauchy integral theorem is that for analytic functions the line
integralisafunctiononlyofits endpoints,independentofthepathof integration,
integraldisplayz2
z1f(z)dz=F(z2)−F(z1)=−integraldisplayz1
z2f(z)dz, (6.39)
againexactlylikethecaseof aconservativeforce, Section1.13.
9contintegraltext
dzandcontintegraltext
zdz=0 by Eq.(6.27a).
6.3 Cauchy’s Integral Theorem 423
Multiply Connected Regions
TheoriginalstatementofCauchy’sintegraltheoremdemandedasimplyconnectedregion.
This restriction may be relaxed by the creation of a barrier, a contour line. The purpose of
the following contour-line construction is to permit, within a multiply connected region,
the identification of curves that can be shrunk to a point within the region, that is, the
constructionof asubregionthatissimplyconnected.
Consider the multiply connected region of Fig. 6.8, in which f(z)is not defined for the
interior,R′. Cauchy’s integral theorem is not valid for the contour C, as shown, but we
can construct a contour C′for which the theorem holds. We draw a line from the interior
forbiddenregion, R′,totheforbiddenregionexteriorto Randthenrunanewcontour, C′,
assho wninFig.6. 9.
The new contour, C′, through ABDEFGA never crosses the contour line that literally
convertsRintoasimplyconnectedregion.Thethree-dimensionalanalogofthistechnique
wasusedinSection1.14toproveGauss’law.ByEq. (6.39),
integraldisplayA
Gf(z)dz=−integraldisplayD
Ef(z)dz, (6.40)
FIGURE 6.8Aclosedcontour Cina
multiplyconnectedregion.
FIGURE 6.9Conversionof amultiply
connectedregionintoasimplyconnected
region.
424 Chapter 6 Functions of a Complex Variable I
withf(z)having been continuous across the contour line and line segments DEandGA
arbitrarilyclosetogether.Then
contintegraldisplay
C′f(z)dz=integraldisplay
ABDf(z)dz+integraldisplay
EFGf(z)dz=0 (6.41)
by Cauchy’s integral theorem, with region Rnow simply connected. Applying Eq. (6.39)
onceagainwith ABD→C′
1andEFG→−C′
2, weobtain
contintegraldisplay
C′
1f(z)dz=contintegraldisplay
C′
2f(z)dz, (6.42)
in which C′
1andC′
2are both traversed in the same (counterclockwise, that is, positive)
direction.
Let us emphasize that the contour line here is a matter of mathematical convenience, to
permit the application of Cauchy’s integral theorem. Since f(z)is analytic in the annular
region,itis necessarilysingle-valuedandcontinuousacross anysuchcontourline.
Exercises
6.3.1 Showthatintegraltextz2
z1f(z)dz=−integraltextz1
z2f(z)dz.
6.3.2 Provethat
vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
Cf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle≤|f|max·L,
where|f|maxis the maximum value of |f(z)|along the contour CandLis the length
ofthecontour.
6.3.3 Verifythat
integraldisplay1,1
0,0z∗dz
depends on the path by evaluating the integral for the two paths shown in Fig. 6.10.
Recallthat f(z)=z∗isnotananalyticfunctionof zandthatCauchy’sintegraltheorem
thereforedoesnotapply.
6.3.4 Showthat
contintegraldisplay
Cdz
z2+z=0,
inwhichthecontour Cis acircledefinedby |z|=R>1.
Hint. Direct use of the Cauchy integral theorem is illegal. Why? The integral may be
evaluatedbytransformingtopolarcoordinatesandusingtables.Thisyields0for R>1
and 2πiforR<1.
6.4 Cauchy’s Integral Formula 425
FIGURE 6.10Contour.
6.4 C AUCHY ’SINTEGRAL FORMULA
Asintheprecedingsection,weconsiderafunction f(z)thatisanalyticonaclosedcontour
Candwithintheinteriorregionboundedby C.We seektoprovethat
1
2πicontintegraldisplay
Cf(z)
z−z0dz=f(z0), (6.43)
in which z0is any point in the interior region bounded by C. This is the second of the
two basic theorems mentioned in Section 6.3. Note that since zis on the contour Cwhile
z0is in the interior, z−z0/negationslash=0 and the integral Eq. (6.43) is well defined. Although f(z)
is assumed analytic, the integrand is f(z)/(z−z0)and is not analytic at z=z0unless
f(z0)=0. If the contour is deformed as shown in Fig. 6.11 (or Fig. 6.9, Section 6.3),
Cauchy’sintegraltheoremapplies.ByEq. (6.42),
contintegraldisplay
Cf(z)
z−z0dz−contintegraldisplay
C2f(z)
z−z0dz=0, (6.44)
whereCistheoriginaloutercontourand C2isthecirclesurroundingthepoint z0traversed
inacounterclockwise direction.Let z=z0+reiθ,usingthepolarrepresentationbecause
of the circular shape of the path around z0.H e r eris small and will eventually be made to
approachzero.Wehave(with dz=ireirθdθfromEq. (6.27a))
contintegraldisplay
C2f(z)
z−z0dz=contintegraldisplay
C2f(z0+reiθ)
reiθrieiθdθ.
Takingthelimitas r→0,weobtain
contintegraldisplay
C2f(z)
z−z0dz=if(z0)integraldisplay
C2dθ=2πif(z0), (6.45)
426 Chapter 6 Functions of a Complex Variable I
FIGURE 6.11Exclusionof a
singularpoint.
sincef(z)is analytic and therefore continuous at z=z0. This proves the Cauchy integral
formula.
Hereisaremarkableresult.Thevalueofananalyticfunction f(z)isgivenataninterior
pointz=z0once the values on the boundary Care specified. This is closely analogous to
atwo-dimensionalformofGauss’law(Section1.14)inwhichthemagnitudeofaninterior
linechargewouldbegivenintermsofthecylindricalsurfaceintegraloftheelectricfield E.
A further analogy is the determination of a function in real space by an integral of the
function and the corresponding Green’s function (and their derivatives) over the bounding
surface. Kirchhoffdiffractiontheoryisanexampleofthis.
It has been emphasized that z0is an interior point. What happens if z0is exterior to C?
In this case the entire integrand is analytic on and within C. Cauchy’s integral theorem,
Section6.3,appliesandtheintegralvanishes.Wehave
1
2πicontintegraldisplay
Cf(z)dz
z−z0=braceleftBigg
f(z0), z0interior
0,z 0exterior.
Derivatives
Cauchy’s integral formula may be used to obtain an expression for the derivative of f(z).
FromEq. (6.43), with f(z)analytic,
f(z0+δz0)−f(z0)
δz0=1
2πiδz0parenleftbiggcontintegraldisplayf(z)
z−z0−δz0dz−contintegraldisplayf(z)
z−z0dzparenrightbigg
.
Then,bydefinitionof derivative(Eq. (6.14)),
f′(z0)=lim
δz0→01
2πiδz0contintegraldisplayδz0f(z)
(z−z0−δz0)(z−z0)dz
=1
2πicontintegraldisplayf(z)
(z−z0)2dz. (6.46)
This result could have been obtained by differentiating Eq. (6.43) under the integral sign
withrespectto z0.Thisformal,orturning-the-crank,approachisvalid,butthejustification
forit iscontainedintheprecedinganalysis.
6.4 Cauchy’s Integral Formula 427
This technique for constructing derivatives may be repeated. We write f′(z0+δz0)
andf′(z0), using Eq. (6.46). Subtracting, dividing by δz0, and finally taking the limit as
δz0→0,wehave
f(2)(z0)=2
2πicontintegraldisplayf(z)dz
(z−z0)3.
Note that f(2)(z0)is independent of the direction of δz0, as it must be. Continuing, we
get10
f(n)(z0)=n!
2πicontintegraldisplayf(z)dz
(z−z0)n+1; (6.47)
that is, the requirement that f(z)be analytic guarantees not only a first derivative but
derivativesof allordersaswell!Thederivativesof f(z)areautomaticallyanalytic.Notice
thatthisstatementassumestheGoursatversionoftheCauchyintegraltheorem.Thisisalso
why Goursat’s contribution is so significant in the development of the theory of complex
variables.
Morera’s Theorem
A further application of Cauchy’s integral formula is in the proof of Morera’s theorem,
whichistheconverseofCauchy’sintegraltheorem.Thetheoremstatesthefollowing:
If afunction f(z)is continuousinasimplyconnectedregion Randcontintegraltext
Cf(z)dz=0for every closed contour CwithinR, thenf(z)is analytic
throughout R.
Let us integrate f(z)fromz1toz2. Since every closed-path integral of f(z)vanishes,
the integral is independent of path and depends only on its endpoints. We label the result
oftheintegration F(z), with
F(z2)−F(z1)=integraldisplayz2
z1f(z)dz. (6.48)
Asanidentity,
F(z2)−F(z1)
z2−z1−f(z1)=integraltextz2
z1[f(t)−f(z1)]dt
z2−z1, (6.49)
usingtasanothercomplexvariable.Nowwetakethelimitas z2→z1:
limz2→z1integraltextz2
z1[f(t)−f(z1)]dt
z2−z1=0, (6.50)
10This expression is the starting point for defining derivatives of fractional order . See A. Erdelyi (ed.), Tables of Integral
Transforms ,Vol.2.NewYork:McGraw-Hill(1954).Forrecentapplicationstomathematicalanalysis,seeT.J.Osler,Anintegral
analogue of Taylor’s series and its use in computing Fourier transforms. Math. Comput. 26: 449 (1972), and referencestherein.
428 Chapter 6 Functions of a Complex Variable I
sincef(t)is continuous.11Therefore
limz2→z1F(z2)−F(z1)
z2−z1=F′(z)vextendsinglevextendsingle
z=z1=f(z1) (6.51)
by definition of derivative (Eq. (6.14)). We have proved that F′(z)atz=z1exists and
equalsf(z1). Sincez1is any point in R, we see that F(z)is analytic. Then by Cauchy’s
integral formula (compare Eq. (6.47)), F′(z)=f(z)is also analytic, proving Morera’s
theorem.
Drawing once more on our electrostatic analog, we might use f(z)to represent the
electrostatic field E. If the net charge within every closed region in Ris zero (Gauss’
law), the charge density is everywhere zero in R. Alternatively, in terms of the analysis of
Section1.13, f(z)representsaconservativeforce(bydefinitionofconservative),andthen
wefindthatitisalwayspossibletoexpressitasthederivativeofapotentialfunction F(z).
AnimportantapplicationofCauchy’sintegralformulaisthefollowing Cauchyinequal-
ity.I ff(z)=summationtextanznis analytic and bounded, |f(z)|≤Mon a circle of radius rabout
theorigin,then
|an|rn≤M(Cauchy’sinequality ) (6.52)
gives upper bounds for the coefficients of its Taylor expansion. To prove Eq. (6.52) let us
defineM(r)=max|z|=r|f(z)|andusetheCauchyintegralfor an:
|an|=1
2πvextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay
|z|=rf(z)
zn+1dzvextendsinglevextendsinglevextendsinglevextendsingle≤M(r)2πr
2πrn+1.
An immediate consequence of the inequality (6.52) is Liouville’s theorem :I ff(z)is
analyticandboundedintheentirecomplexplaneitisaconstant.Infact,if |f(z)|≤Mfor
allz, thenCauchy’sinequality(6.52) gives |an|≤Mr−n→0a sr→∞forn>0.Hence
f(z)=a0.
Conversely,theslightestdeviationofananalyticfunctionfromaconstantvalueimplies
that there must be at least one singularity somewhere in the infinite complex plane. Apart
from the trivial constant functions, then, singularities are a fact of life, and we must learn
to live with them. But we shall do more than that. We shall next expand a function in a
Laurent series at a singularity, and we shall use singularities to develop the powerful and
usefulcalculusofresiduesinChapter7.
A famous application of Liouville’s theorem yields the fundamental theorem of alge-
bra(due to C. F. Gauss), which says that any polynomial P(z)=summationtextn
ν=0aνzνwithn>0
andan/negationslash=0 hasnroots. To prove this, suppose P(z)has no zero. Then 1 /P(z)is analytic
and bounded as |z|→∞. HenceP(z)is a constant by Liouville’s theorem, q.e.a. Thus,
P(z)has at least one root that we can divide out. Then we repeat the process for the re-
sulting polynomial of degree n−1. This leads to the conclusion that P(z)has exactly n
roots.
11Wequote the meanvalue theoremof calculus here.
6.4 Cauchy’s Integral Formula 429
Exercises
6.4.1 Showthat
contintegraldisplay
C(z−z0)ndz=braceleftBigg
2πi, n=−1,
0,n/negationslash=−1,
where the contour Cencircles the point z=z0in a positive (counterclockwise) sense.
The exponent nis an integer. See also Eq. (6.27a). The calculus of residues, Chapter 7,
isbasedonthisresult.
6.4.2 Showthat
1
2πicontintegraldisplay
zm−n−1dz, m andnintegers
(withthecontourencirclingtheoriginoncecounterclockwise)isarepresentationofthe
Kronecker δmn.
6.4.3 SolveExercise6.3.4byseparatingtheintegrandintopartialfractionsandthenapplying
Cauchy’sintegraltheoremformultiplyconnectedregions.
Note. Partial fractions are explained in Section 15.8 in connection with Laplace trans-
forms.
6.4.4 Evaluatecontintegraldisplay
Cdz
z2−1,
whereCisthecircle|z|=2.
6.4.5 Assuming that f(z)is analytic on and within a closed contour Cand that the point z0
iswithin C,showthat
contintegraldisplay
Cf′(z)
z−z0dz=contintegraldisplay
Cf(z)
(z−z0)2dz.
6.4.6 You know that f(z)is analytic on and within a closed contour C. You suspect that the
nthderivative f(n)(z0)is givenby
f(n)(z0)=n!
2πicontintegraldisplay
Cf(z)
(z−z0)n+1dz.
Usingmathematicalinduction,provethatthisexpressioniscorrect.
6.4.7 (a) A function f(z)is analytic within a closed contour C(and continuous on C). If
f(z)/negationslash=0 withinCand|f(z)|≤MonC, showthat
vextendsinglevextendsinglef(z)vextendsinglevextendsingle≤M
forallpointswithin C.
Hint.Consider w(z)=1/f (z).
(b) Iff(z)=0 within the contour C, show that the foregoing result does not hold
and that it is possible to have |f(z)|=0 at one or more points in the interior with
|f(z)|>0overtheentireboundingcontour.Citeaspecificexampleofananalytic
functionthatbehavesthisway.
430 Chapter 6 Functions of a Complex Variable I
6.4.8 Using the Cauchy integral formula for the nth derivative, convert the following Ro-
driguesformulasintothecorrespondingso-calledSchlaefliintegrals.
(a) Legendre:
Pn(x)=1
2nn!dn
dxnparenleftbig
x2−1parenrightbign.
ANS.(−1)n
2n·1
2πicontintegraldisplay(1−z2)n
(z−x)n+1dz.
(b) Hermite:
Hn(x)=(−1)nex2dn
dxne−x2.
(c) Laguerre:
Ln(x)=ex
n!dn
dxnparenleftbig
xne−xparenrightbig
.
Note. From the Schlaefli integral representations one can develop generating functions
for thesespecialfunctions.CompareSections12.4,13.1, and13.2.
6.5 L AURENT EXPANSION
Taylor Expansion
TheCauchyintegralformulaoftheprecedingsectionopensupthewayforanotherderiva-
tion of Taylor’s series (Section 5.6), but this time for functions of a complex variable.
Supposewearetryingtoexpand f(z)aboutz=z0andwehave z=z1asthenearestpoint
on the Argand diagram for which f(z)is not analytic. We construct a circle Ccentered at
z=z0with radius less than |z1−z0|(Fig. 6.12). Since z1was assumed to be the nearest
pointatwhich f(z)wasnotanalytic, f(z)is necessarilyanalyticonandwithin C.
FromEq. (6.43), theCauchyintegralformula,
f(z)=1
2πicontintegraldisplay
Cf(z′)dz′
z′−z
=1
2πicontintegraldisplay
Cf(z′)dz′
(z′−z0)−(z−z0)
=1
2πicontintegraldisplay
Cf(z′)dz′
(z′−z0)[1−(z−z0)/(z′−z0)]. (6.53)
Herez′is a point on the contour Candzis any point interior to C. It is not legal yet to
expandthedenominatoroftheintegrandinEq.(6.53)bythebinomialtheorem,forwehave
notyetprovedthebinomialtheoremfor complexvariables.Instead,wenotetheidentity
1
1−t=1+t+t2+t3+···=∞summationdisplay
n=0tn, (6.54)
6.5 Laurent Expansion 431
FIGURE 6.12CirculardomainforTaylor
expansion.
which may easily be verified by multiplying both sides by 1 −t. The infinite series, fol-
lowingthemethodsofSection5.2,is convergentfor |t|<1.
Now, for a point zinterior to C,|z−z0|<|z′−z0|, and, using Eq. (6.54), Eq. (6.53)
becomes
f(z)=1
2πicontintegraldisplay
C∞summationdisplay
n=0(z−z0)nf(z′)dz′
(z′−z0)n+1. (6.55)
Interchanging the order of integration and summation (valid because Eq. (6.54) is uni-
formlyconvergentfor |t|<1),weobtain
f(z)=1
2πi∞summationdisplay
n=0(z−z0)ncontintegraldisplay
Cf(z′)dz′
(z′−z0)n+1. (6.56)
ReferringtoEq. (6.47), weget
f(z)=∞summationdisplay
n=0(z−z0)nf(n)(z0)
n!, (6.57)
which is our desired Taylor expansion. Note that it is based only on the assumption that
f(z)isanalyticfor|z−z0|<|z1−z0|.Justasforrealvariablepowerseries(Section5.7),
thisexpansionis uniquefora given z0.
FromtheTaylorexpansionfor f(z)abinomialtheoremmaybederived(Exercise6.5.2).
Schwarz Reflection Principle
From the binomial expansion of g(z)=(z−x0)nfor integral nit is easy to see that the
complexconjugateof thefunction gis thefunctionof thecomplexconjugatefor real x0:
g∗(z)=bracketleftbig
(z−x0)nbracketrightbig∗=(z∗−x0)n=g(z∗). (6.58)
432 Chapter 6 Functions of a Complex Variable I
FIGURE 6.13Schwarzreflection.
ThisleadsustotheSchwarzreflectionprinciple:
If a function f(z)is(1)analytic over some region including the real axis and
(2)realwhen zis real,then
f∗(z)=f(z∗). (6.59)
(SeeFig.6.13.)
Expanding f(z)aboutsome(nonsingular)point x0ontherealaxis,
f(z)=∞summationdisplay
n=0(z−x0)nf(n)(x0)
n!(6.60)
by Eq. (6.56). Since f(z)is analytic at z=x0, this Taylor expansion exists. Since f(z)is
realwhen zisreal,f(n)(x0)mustberealforall n.ThenwhenweuseEq.(6.58),Eq.(6.59),
theSchwarzreflectionprinciple,followsimmediately.Exercise6.5.6isanotherformofthis
principle. This completes the proof within a circle of convergence. Analytic continuation
thenpermitsextendingthisresulttotheentireregionofanalyticity.
Analytic Continuation
Itisnaturaltothinkofthevalues f(z)ofananalyticfunction fasasingleentity,whichis
usuallydefinedinsomerestrictedregion S1ofthecomplexplane,forexample,byaTaylor
series(seeFig.6.14).Then fisanalyticinsidethe circleofconvergence C1,whoseradius
is given by the distance r1from the center of C1to thenearest singularity offatz1(in
Fig.6.14).Asingularityisanypointwhere fisnotanalytic.Ifwechooseapointinside C1
6.5 Laurent Expansion 433
FIGURE 6.14Analyticcontinuation.
thatisfartherthan r1fromthesingularity z1andmakeaTaylorexpansionof faboutit(z2
inFig.6.14),thenthecircleofconvergence, C2willusuallyextendbeyondthefirstcircle,
C1. In the overlap region of both circles, C1,C2, the function fis uniquely defined. In
the region of the circle C2that extends beyond C1,f(z)is uniquely defined by the Taylor
series about the center of C2and is analytic there, although the Taylor series about the
centerof C1isnolongerconvergentthere.AfterWeierstrassthisprocessiscalled analytic
continuation . It defines the analytic functions in terms of its original definition (in C1,
say)andallitscontinuations.
Aspecificexampleisthefunction
f(z)=1
1+z, (6.61)
which has a (simple) pole at z=−1 and is analytic elsewhere. The geometric series ex-
pansion
1
1+z=1−z+z2+···=∞summationdisplay
n=0(−z)n(6.62)
convergesfor|z|<1,thatis, insidethecircle C1inFig.6.14.
Supposeweexpand f(z)aboutz=i,s o
f(z)=1
1+z=1
1+i+(z−i)=1
(1+i)(1+(z−i)/(1+i))
=bracketleftbigg
1−z−i
1+i+(z−i)2
(1+i)2−···bracketrightbigg1
1+i(6.63)
converges for|z−i|<|1+i|=√
2. Our circle of convergence is C2in Fig. 6.14. Now
434 Chapter 6 Functions of a Complex Variable I
FIGURE 6.15|z′−z0|C1>|z−z0|;|z′−z0|C2<|z−z0|.
f(z)is defined by the expansion (6.63) in S2, which overlaps S1and extends further out
in the complex plane.12This extension is an analytic continuation, and when we have
only isolated singular points to contend with, the function can be extended indefinitely.
Equations(6.61),(6.62),and(6.63)arethreedifferentrepresentationsofthesamefunction.
Each representation has its own domain of convergence. Equation (6.62) is a Maclaurin
series.Equation(6.63)isaTaylorexpansionabout z=iandfromthefollowingparagraphs
Eq.(6.61) isseentobeaone-termLaurentseries.
Analytic continuation may take many forms, and the series expansion just considered
is not necessarily the most convenient technique. As an alternate technique we shall use a
functional relation in Section 8.1 to extend the factorial function around the isolated sin-
gular points z=−n,n=1,2,3,....As another example, the hypergeometric equation is
satisfied by the hypergeometric function defined by the series, Eq. (13.115), for |z|<1.
The integral representation given in Exercise 13.4.7 permits a continuation into the com-
plexplane.
12One of the most powerful and beautiful results of the more abstract theory of functions of a complex variable is that if two
analytic functions coincide in any region, such as the overlap of S1andS2, or coincide on any line segment, they are the same
function, in the sense that they will coincide everywhere as long as they are both well defined. In this case the agreement of the
expansions (Eqs. (6.62) and (6.63)) over the region common to S1andS2would establish the identity of the functions these
expansions represent. Then Eq. (6.63) would represent an analytic continuation or extension of f(z)into regions not covered
by Eq. (6.62). We could equally well say that f(z)=1/(1+z)is itself an analytic continuation of either of the series given by
Eqs. (6.62) and (6.63).
6.5 Laurent Expansion 435
Laurent Series
Wefrequentlyencounterfunctionsthatareanalyticandsingle-valuedinanannularregion,
say, of inner radius rand outer radius R, as shown in Fig. 6.15. Drawing an imaginary
contour line to convert our region into a simply connected region, we apply Cauchy’s
integralformula,andfortwocircles C2andC1centeredat z=z0andwithradii r2andr1,
respectively,where r<r2<r1<R,weha v e13
f(z)=1
2πicontintegraldisplay
C1f(z′)dz′
z′−z−1
2πicontintegraldisplay
C2f(z′)dz′
z′−z. (6.64)
Note that in Eq. (6.64) an explicit minus sign has been introduced so that the contour
C2(likeC1) is to be traversed in the positive (counterclockwise) sense. The treatment of
Eq. (6.64) now proceeds exactly like that of Eq. (6.53) in the development of the Taylor
series. Each denominator is written as (z′−z0)−(z−z0)and expanded by the binomial
theorem,whichnowfollowsfromtheTaylorseries(Eq. (6.57)).
Notingthatfor C1,|z′−z0|>|z−z0|whilefor C2,|z′−z0|<|z−z0|,wefind
f(z)=1
2πi∞summationdisplay
n=0(z−z0)ncontintegraldisplay
C1f(z′)dz′
(z′−z0)n+1
+1
2πi∞summationdisplay
n=1(z−z0)−ncontintegraldisplay
C2(z′−z0)n−1f(z′)dz′. (6.65)
The minus sign of Eq. (6.64) has been absorbed by the binomial expansion. Labeling the
firstseries S1andthesecond S2wehave
S1=1
2πi∞summationdisplay
n=0(z−z0)ncontintegraldisplay
C1f(z′)dz′
(z′−z0)n+1, (6.66)
which is the regular Taylor expansion, convergent for |z−z0|<|z′−z0|=r1, that is, for
allzinteriortothelargercircle, C1.Forthesecondseries inEq. (6.65)wehave
S2=1
2πi∞summationdisplay
n=1(z−z0)−ncontintegraldisplay
C2(z′−z0)n−1f(z′)dz′, (6.67)
convergentfor |z−z0|>|z′−z0|=r2,thatis,forall zexteriortothesmallercircle, C2.
Remember, C2nowgoescounterclockwise.
Thesetwoseries arecombinedintooneseries14(a Laurentseries)by
f(z)=∞summationdisplay
n=−∞an(z−z0)n, (6.68)
13We may take r2arbitrarily close to randr1arbitrarily closeto R, maximizing the areaenclosedbetween C1andC2.
14Replacenby−ninS2and add.
436 Chapter 6 Functions of a Complex Variable I
where
an=1
2πicontintegraldisplay
Cf(z′)dz′
(z′−z0)n+1. (6.69)
Since, in Eq. (6.69), convergence of a binomial expansion is no longer a problem, Cmay
beanycontourwithintheannularregion r<|z−z0|<Rencircling z0onceinacounter-
clockwisesense. If weassumethatsuchanannularregionofconvergencedoesexist,then
Eq.(6.68) istheLaurentseries, orLaurentexpansion,of f(z).
The use of the contour line (Fig. 6.15) is convenient in converting the annular region
into a simply connected region. Since our function is analytic in this annular region (and
single-valued), the contour line is not essential and, indeed, does not appear in the final
result,Eq. (6.69).
Laurent series coefficients need not come from evaluation of contour integrals (which
may be very intractable). Other techniques, such as ordinary series expansions, may pro-
videthecoefficients.
Numerous examples of Laurent series appear in Chapter 7. We limit ourselves here to
onesimpleexampletoillustratetheapplicationof Eq.(6.68).
Example 6.5.1 LAURENT EXPANSION
Letf(z)=[z(z−1)]−1. If we choose z0=0, thenr=0 andR=1,f(z)diverging at
z=1.ApartialfractionexpansionyieldstheLaurentseries
1
z(z−1)=−1
1−z−1
z=−1
z−1−z−z2−z3−···=−∞summationdisplay
n=−1zn.(6.70)
FromEqs. (6.70), (6.68), and(6.69)wethenhave
an=1
2πicontintegraldisplaydz′
(z′)n+2(z′−1)=braceleftBigg
−1forn≥−1,
0forn<−1.(6.71)
The integrals in Eq. (6.71) can also be directly evaluated by substituting the geometric-
seriesexpansionof (1−z′)−1usedalreadyinEq. (6.70) for (1−z)−1:
an=−1
2πicontintegraldisplay∞summationdisplay
m=0(z′)mdz′
(z′)n+2. (6.72)
Uponinterchangingtheorderofsummationandintegration(uniformlyconvergentseries),
wehave
an=−1
2πi∞summationdisplay
m=0contintegraldisplaydz′
(z′)n+2−m. (6.73)
6.5 Laurent Expansion 437
If weemploythepolarform, asinEq. (6.47)(or compareExercise6.4.1),
an=−1
2πi∞summationdisplay
m=0contintegraldisplayrieiθdθ
rn+2−mei(n+2−m)θ
=−1
2πi·2πi∞summationdisplay
m=0δn+2−m,1, (6.74)
whichagreeswithEq. (6.71). /squaresolid
The Laurent series differs from the Taylor series by the obvious feature of negative
powersof (z−z0).ForthisreasontheLaurentserieswillalwaysdivergeatleastat z=z0
andperhapsasfar outas somedistance r(Fig.6.15).
Exercises
6.5.1 DeveloptheTaylorexpansionof ln (1+z).
ANS.∞summationdisplay
n=1(−1)n−1zn
n.
6.5.2 Derivethebinomialexpansion
(1+z)m=1+mz+m(m−1)
1·2z2+···=∞summationdisplay
n=0parenleftbiggm
nparenrightbigg
zn
formanyrealnumber.Theexpansionis convergentfor |z|<1.Why?
6.5.3 A function f(z)is analytic on and within the unit circle. Also, |f(z)|<1f o r|z|≤1
andf(0)=0.Showthat|f(z)|<|z|for|z|≤1.
Hint. One approach is to show that f(z)/zis analytic and then to express [f(z0)/z0]n
by the Cauchy integral formula. Finally, consider absolute magnitudes and take the nth
root.This exerciseissometimescalledSchwarz’stheorem.
6.5.4 Iff(z)is a real function of the complex variable z=x+iy,t h a ti s ,i f f(x)=f∗(x),
and the Laurent expansion about the origin, f(z)=summationtextanzn, hasan=0f o rn<−N,
showthatallof thecoefficients anare real.
Hint.Showthat zNf(z)is analytic(viaMorera’s theorem,Section6.4).
6.5.5 Afunction f(z)=u(x,y)+iv(x,y)satisfiestheconditionsfortheSchwarzreflection
principle.Showthat
(a)uis anevenfunctionof y.(b)vis anoddfunctionof y.
6.5.6 A function f(z)can be expanded in a Laurent series about the origin with the coeffi-
cientsanreal.Showthatthecomplexconjugateofthisfunctionof zisthesamefunction
ofthecomplexconjugateof z;thatis,
f∗(z)=f(z∗).
Verifythis explicitlyfor
(a)f(z)=zn,naninteger, (b) f(z)=sinz.
Iff(z)=iz(a1=i), showthattheforegoingstatementdoesnothold.
438 Chapter 6 Functions of a Complex Variable I
6.5.7 The function f(z)is analytic in a domain that includes the real axis. When zis real
(z=x),f(x)is pureimaginary.
(a) Showthat
f(z∗)=−bracketleftbig
f(z)bracketrightbig∗.
(b) For the specific case f(z)=iz, develop the Cartesian forms of f(z),f(z∗), and
f∗(z). Donotquotethegeneralresultofpart (a).
6.5.8 Developthefirst threenonzeroterms oftheLaurentexpansionof
f(z)=parenleftbig
ez−1parenrightbig−1
about the origin. Notice the resemblance to the Bernoulli number–generating function,
Eq. (5.144)ofSection5.9.
6.5.9 Provethatthe Laurentexpansionofa givenfunctionaboutagivenpointis unique;that
is, if
f(z)=∞summationdisplay
n=−Nan(z−z0)n=∞summationdisplay
n=−Nbn(z−z0)n,
showthat an=bnfor alln.
Hint.UsetheCauchyintegralformula.
6.5.10 (a) Develop a Laurent expansion of f(z)=[z(z−1)]−1about the point z=1 valid
for small values of |z−1|. Specify the exact range over which your expansion
holds.Thisis ananalyticcontinuationofEq. (6.70).
(b) DeterminetheLaurentexpansionof f(z)aboutz=1b u tf o r|z−1|large.
Hint.Partialfractionthisfunctionanduse thegeometricseries.
6.5.11 (a) Given f1(z)=integraltext∞
0e−ztdt(withtreal),showthatthedomaininwhich f1(z)exists
(andisanalytic)is ℜ(z)>0.
(b) Show that f2(z)=1/zequalsf1(z)overℜ(z) >0 and is therefore an analytic
continuationof f1(z)overtheentire z-planeexceptfor z=0.
(c) Expand 1 /zabout the point z=i. You will have f3(z)=summationtext∞
n=0an(z−i)n. What
isthedomainof f3(z)?
ANS.1
z=−i∞summationdisplay
n=0in(z−i)n,|z−i|<1.
6.6 S INGULARITIES
The Laurent expansion represents a generalization of the Taylor series in the presence of
singularities. We define the point z0as anisolated singular point of the function f(z)if
f(z)is notanalyticat z=z0butisanalyticatallneighboringpoints.
6.6 Singularities 439
Poles
IntheLaurentexpansionof f(z)aboutz0,
f(z)=∞summationdisplay
m=−∞am(z−z0)m, (6.75)
ifam=0f o rm<−n<0 anda−n/negationslash=0,wesaythat z0isapoleoforder n.Forinstance,if
n=1, that is, if a−1/(z−z0)is the first nonvanishing term in the Laurent series, we have
apoleof order1,oftencalleda simplepole.
If, on the other hand, the summation continues to m=−∞, thenz0is a pole of infi-
nite order and is called an essential singularity . These essential singularities have many
pathological features. For instance, we can show that in any small neighborhood of an
essential singularity of f(z)the function f(z)comes arbitrarily close to any (and there-
fore every) preselected complex quantity w0.15Here, the entire w-plane is mapped by
finto the neighborhood of the point z0. One point of fundamental difference between a
pole of finite order nand an essential singularity is that by multiplying f(z)by(z−z0)n,
f(z)(z−z0)nis no longer singular at z0. This obviously cannot be done for an essential
singularity.
The behaviorof f(z)asz→∞is definedin terms of thebehaviorof f(1/t)ast→0.
Considerthefunction
sinz=∞summationdisplay
n=0(−1)nz2n+1
(2n+1)!. (6.76)
Asz→∞, wereplacethe zby 1/ttoobtain
sinparenleftbigg1
tparenrightbigg
=∞summationdisplay
n=0(−1)n
(2n+1)!t2n+1. (6.77)
Fromthedefinition, sin zhas anessentialsingularityatinfinity.This resultcouldbeantic-
ipatedfrom Exercise6.1.9since
sinz=siniy=isinhy,whenx=0,
which approaches infinity exponentially as y→∞. Thus, although the absolute value of
sinxfor realxis equaltoor lessthanunity,theabsolutevalueof sin zisnotbounded.
Afunctionthatisanalyticthroughoutthefinitecomplexplane exceptfor isolatedpoles
is called meromorphic , such as ratios of two polynomials or tan z, cotz. Examples are
alsoentirefunctionsthathavenosingularitiesinthefinitecomplexplane,suchas exp (z),
sinz, cosz(seeSections5.9, 5.11).
15This theorem is due to Picard. A proof is given by E. C. Titchmarsh, The Theory of Functions , 2nd ed. New York: Oxford
University Press (1939).
440 Chapter 6 Functions of a Complex Variable I
Branch Points
Thereisanothersortof singularitythatwillbeimportantinChapter7. Consider
f(z)=za,
inwhich aisnotaninteger.16Aszmovesaroundtheunitcirclefrom e0toe2πi,
f(z)→e2πai/negationslash=e0·a=1,
for nonintegral a. We have a branch point at the origin and another at infinity. If we set
z=1/t,a similar analysis of f(z)fort→0 shows that t=0; that is, z=∞is also a
branch point. The points e0iande2πiin thez-plane coincide, but these coincident points
lead to different values off(z); that is, f(z)is amultivalued function . The problem
is resolved by constructing a cut line joining both branch points so thatf(z)will be
uniquely specified for a given point in the z-plane. For za,the cut line can go out at any
angle. Note that the point at infinity must be included here; that is, the cut line may join
finitebranchpointsviathepointatinfinity.Thenextexampleisacaseinpoint.If a=p/q
isarationalnumber,then qiscalledtheorderofthebranchpoint,becauseoneneedstogo
aroundthebranchpoint qtimesbeforecomingbacktothestartingpoint.If aisirrational,
thentheorderof thebranchpointisinfinite,justasfor thelogarithm.
Note that a function with a branch point and a required cut line will not be continuous
acrossthecutline.Oftentherewillbeaphasedifferenceonoppositesidesofthiscutline.
Hencelineintegralsonoppositesidesofthisbranchpointcutlinewillnotgenerallycancel
eachother.Numerousexamplesof thiscaseappearintheexercises.
The contour line used to convert a multiply connected region into a simply connected
region(Section6.3)iscompletelydifferent.Ourfunctioniscontinuousacrossthatcontour
line,andnophasedifferenceexists.
Example 6.6.1 BRANCH POINTS OF ORDER 2
Considerthefunction
f(z)=parenleftbig
z2−1parenrightbig1/2=(z+1)1/2(z−1)1/2. (6.78)
Thefirstfactorontheright-handside, (z+1)1/2,hasabranchpointat z=−1.Thesecond
factor has a branch point at z=+1. At infinity f(z)has a simple pole. This is best seen
bysubstituting z=1/tandmakingabinomialexpansionat t=0:
parenleftbig
z2−1parenrightbig1/2=1
tparenleftbig
1−t2parenrightbig1/2=1
t∞summationdisplay
n=0parenleftbigg1/2
nparenrightbigg
(−1)nt2n=1
t−1
2t−1
8t3+···.
Thecutlinehastoconnectbothbranchpoints,soitisnotpossibletoencircleeitherbranch
pointcompletely.Tocheckonthepossibilityoftakingthelinesegmentjoining z=+1and
16z=0 is a singular point, for zahas only a finite number of derivatives, whereas an analytic function is guaranteed an infinite
numberofderivatives(Section6.4).Theproblemisthat f(z)isnotsingle-valuedasweencircletheorigin.TheCauchyintegral
formula may not be applied.
6.6 Singularities 441
FIGURE 6.16Branchcutandphasesof Table6.1.
Table 6.1 PhaseAngle
Point θϕθ+ϕ
2
10 00
20 ππ
2
30 ππ
2
4 ππ π
52 ππ3π
2
62 ππ3π
2
72 π 2π 2π
z=−1 as a cut line, let us follow the phases of these two factors as we move along the
contourshowninFig.6.16.
For convenience in following the changes of phase let z+1=reiθandz−1=ρeiϕ.
Thenthephaseof f(z)is(θ+ϕ)/2.Westartatpoint1,whereboth z+1andz−1ha v ea
phaseofzero.Movingfrompoint1topoint2, ϕ,thephaseof z−1=ρeiϕ,increasesby π.
(z−1becomesnegative.) ϕthenstaysconstantuntilthecircleiscompleted,movingfrom
6to7.θ,thephaseof z+1=reiθ,showsasimilarbehavior,increasingby2 πaswemove
from 3 to 5. The phase of the function f(z)=(z+1)1/2(z−1)1/2=r1/2ρ1/2ei(θ+ϕ)/2is
(θ+ϕ)/2.This istabulatedinthefinalcolumnofTable6.1.
Twofeaturesemerge:
1. The phase at points 5 and 6 is not the same as the phase at points 2 and 3. This
behaviorcanbeexpectedatabranchcut.
2.Thephaseatpoint7exceedsthatatpoint1by2 π,andthefunction f(z)=(z2−1)1/2
istherefore single-valued for thecontourshown,encircling bothbranchpoints.
Ifwetakethe x-axis,−1≤x≤1,asacutline, f(z)isuniquelyspecified.Alternatively,
thepositive x-axisforx>1 andthenegative x-axisforx<−1 maybetakenascutlines.
The branch points cannot be encircled, and the function remains single-valued. These two
cutlinesare, infact, onebranchcutfrom −1t o+1 viathepointatinfinity. /squaresolid
Generalizingfrom thisexample,wehavethatthephaseofafunction
f(z)=f1(z)·f2(z)·f3(z)···
isthealgebraicsumofthephaseofitsindividualfactors:
argf(z)=argf1(z)+argf2(z)+argf3(z)+···.
442 Chapter 6 Functions of a Complex Variable I
Thephaseofanindividualfactormaybetakenasthearctangentoftheratioofitsimaginary
parttoitsrealpart(choosingtheappropriatebranchofthearctanfunctiontan−1y/x,which
hasinfinitelymanybranches),
argfi(z)=tan−1parenleftbiggvi
uiparenrightbigg
.
Forthecaseof afactorof theform
fi(z)=(z−z0),
the phase corresponds to the phase angle of a two-dimensional vector from +z0toz,t h e
phase increasing by 2 πas the point+z0is encircled. Conversely, the traversal of any
closedloopnotencircling z0doesnotchangethephaseof z−z0.
Exercises
6.6.1 The function f(z)expanded in a Laurent series exhibits a pole of order matz=z0.
Showthatthecoefficientof (z−z0)−1,a−1, isgivenby
a−1=1
(m−1)!dm−1
dzm−1bracketleftbig
(z−z0)mf(z)bracketrightbig
z=z0,
with
a−1=bracketleftbig
(z−z0)f(z)bracketrightbig
z=z0,
when the pole is a simple pole (m=1). These equations for a−1are extremely useful
indeterminingtheresiduetobeusedintheresiduetheoremofSection7.1.
Hint. The technique that was so successful in proving the uniqueness of power series,
Section5.7,willworkherealso.
6.6.2 Afunction f(z)canberepresentedby
f(z)=f1(z)
f2(z),
in which f1(z)andf2(z)are analytic. The denominator, f2(z), vanishes at z=z0,
showing that f(z)has a pole at z=z0. However, f1(z0)/negationslash=0,f′
2(z0)/negationslash=0. Show that
a−1, thecoefficientof (z−z0)−1ina Laurentexpansionof f(z)atz=z0,is givenby
a−1=f1(z0)
f′
2(z0).
(This resultleadstotheHeavisideexpansiontheorem,Exercise15.12.11.)
6.6.3 InanalogywithExample6.6.1,considerindetailthephaseofeachfactorandtheresul-
tantoverallphase of f(z)=(z2+1)1/2followinga contoursimilar to thatof Fig.6.16
butencirclingthenewbranchpoints.
6.6.4 The Legendre function of the second kind, Qν(z), has branch points at z=±1. The
branchpointsarejoinedbyacutlinealongthereal (x)axis.
6.7 Mapping 443
(a) Showthat Q0(z)=1
2ln((z+1)/(z−1))issingle-valued(withtherealaxis −1≤
x≤1 takenasacutline).
(b) Forrealargument xand|x|<1 it isconvenienttotake
Q0(x)=1
2ln1+x
1−x.
Showthat
Q0(x)=1
2bracketleftbig
Q0(x+i0)+Q0(x−i0)bracketrightbig
.
Herex+i0 indicatesthat zapproachesthereal axis from above,and x−i0 indi-
catesanapproachfrom below.
6.6.5 As an example of an essential singularity, consider e1/zaszapproaches zero. For any
complexnumber z0,z0/negationslash=0,showthat
e1/z=z0
hasaninfinitenumberofsolutions.
6.7 M APPING
In the preceding sections we have defined analytic functions and developed some of their
mainfeatures.Hereweintroducesomeofthemoregeometricaspectsoffunctionsofcom-
plexvariables,aspectsthatwillbeusefulinvisualizingtheintegraloperationsinChapter7
and that are valuable in their own right in solving Laplace’s equation in two-dimensional
systems.
In ordinary analytic geometry we may take y=f(x)and then plot yversusx.O u r
problemhereismorecomplicated,for zisafunctionoftwovariables, xandy.W eusethe
notation
w=f(z)=u(x,y)+iv(x,y). (6.79)
Then for a point in the z-plane (specific values for xandy) there may correspond specific
values for u(x,y)andv(x,y)that then yield a point in the w-plane. As points in the
z-plane transform, or are mapped into points in the w-plane, lines or areas in the z-plane
will be mapped into lines or areas in the w-plane. Our immediate purpose is to see how
linesandareasmapfromthe z-planetothe w-planefor anumberofsimplefunctions.
Translation
w=z+z0. (6.80)
The function wis equal to the variable zplus a constant, z0=x0+iy0. By Eqs. (6.1) and
(6.79),
u=x+x0,v=y+y0, (6.81)
representingapuretranslationofthecoordinateaxes,asshowninFig.6.17.
444 Chapter 6 Functions of a Complex Variable I
FIGURE 6.17Translation.
Rotation
w=zz0. (6.82)
Hereit isconvenienttoreturntothepolarrepresentation,using
w=ρeiϕ,z=reiθ,andz0=r0eiθ0, (6.83)
then
ρeiϕ=rr0ei(θ+θ0), (6.84)
or
ρ=rr0,ϕ=θ+θ0. (6.85)
Two things have occurred. First, the modulus rhas been modified, either expanded or
contracted, by the factor r0. Second, the argument θhas been increased by the additive
constantθ0(Fig.6.18).Thisrepresentsarotationofthecomplexvariablethroughanangle
θ0. Forthespecialcaseof z0=i, wehaveapurerotationthrough π/2 radians.
FIGURE 6.18Rotation.
6.7 Mapping 445
Inversion
w=1
z. (6.86)
Again,usingthepolarform,wehave
ρeiϕ=1
reiθ=1
re−iθ, (6.87)
whichshowsthat
ρ=1
r,ϕ=−θ. (6.88)
The first part of Eq. (6.87) shows that inversion clearly. The interior of the unit circle
is mapped onto the exterior and vice versa (Fig. 6.19). In addition, the second part of
Eq. (6.87) shows that the polar angle is reversed in sign. Equation (6.88) therefore also
involvesareflectionof the y-axis,exactlylikethecomplexconjugateequation.
To see how curves in the z-plane transform into the w-plane, we return to the Cartesian
form:
u+iv=1
x+iy. (6.89)
Rationalizing the right-hand side by multiplying numerator and denominator by z∗and
thenequatingtherealparts andtheimaginaryparts,wehave
u=x
x2+y2,x=u
u2+v2,
v=−y
x2+y2,y=−v
u2+v2.(6.90)
FIGURE 6.19Inversion.
446 Chapter 6 Functions of a Complex Variable I
Acirclecenteredattheorigininthe z-planehastheform
x2+y2=r2(6.91)
andbyEqs. (6.90) transformsinto
u2
(u2+v2)2+v2
(u2+v2)2=r2. (6.92)
SimplifyingEq.(6.92), weobtain
u2+v2=1
r2=ρ2, (6.93)
whichdescribesacircleinthe w-planealsocenteredattheorigin.
Thehorizontalline y=c1transforms into
−v
u2+v2=c1, (6.94)
or
u2+parenleftbigg
v+1
2c1parenrightbigg2
=1
(2c1)2, (6.95)
which describes a circle in the w-plane of radius (1/2c1)and centered at u=0,v=−1
2c1(Fig.6.20).
We pick up the other three possibilities, x=±c1,y=−c1, by rotating the xy-axes. In
general, any straight line or circle in the z-plane will transform into a straight line or a
circleinthe w-plane(compareExercise6.7.1).
FIGURE 6.20Inversion,line ↔circle.
6.7 Mapping 447
Branch Points and Multivalent Functions
The three transformations just discussed have all involved one-to-one correspondence of
points in the z-plane to points in the w-plane. Now to illustrate the variety of transfor-
mations that are possible and the problems that can arise, we introduce first a two-to-one
correspondence and then a many-to-one correspondence. Finally, we take up the inverses
ofthesetwotransformations.
Considerfirst thetransformation
w=z2, (6.96)
whichleadsto
ρ=r2,ϕ=2θ. (6.97)
Clearly, our transformation is nonlinear, for the modulus is squared, but the significant
featureofEq. (6.96)is thatthephaseangleorargumentis doubled.Thismeansthatthe
•firstquadrantof z,0≤θ<π
2,→upperhalf-planeof w,0≤ϕ<π,
•upperhalf-planeof z,0≤θ<π,→wholeplaneof w,0≤ϕ<2π.
The lower half-plane of zmaps into the already covered entire plane of w, thus covering
thew-plane a second time. This is our two-to-one correspondence, that is, two distinct
pointsinthe z-plane,z0andz0eiπ=−z0, correspondingtothesinglepoint w=z2
0.
In Cartesianrepresentation,
u+iv=(x+iy)2=x2−y2+i2xy, (6.98)
leadingto
u=x2−y2,v=2xy. (6.99)
Hence the lines u=c1,v=c2in thew-plane correspond to x2−y2=c1,2xy=c2, rec-
tangular (and orthogonal) hyperbolas in the z-plane (Fig. 6.21). To every point on the
hyperbola x2−y2=c1in the right half-plane, x>0, one point on the line u=c1corre-
sponds,andviceversa.However,everypointontheline u=c1alsocorrespondstoapoint
onthehyperbola x2−y2=c1inthelefthalf-plane, x<0,asalreadyexplained.
It will be shown in Section 6.8 that if lines in the w-plane are orthogonal, the corre-
spondinglinesinthe z-planearealsoorthogonal,aslongasthetransformationisanalytic.
Sinceu=c1andv=c2are constructed perpendicular to each other, the corresponding
hyperbolasinthe z-planeareorthogonal.Wehaveconstructedaneworthogonalsystemof
hyperbolic lines (or surfaces if we add an axis perpendicular to xandy). Exercise 2.1.3
was an analysis of this system. It might be noted that if the hyperbolic lines are electric
or magnetic lines of force, then we have a quadrupole lens useful in focusing beams of
high-energyparticles.
Theinverseofthefourthtransformation(Eq. (6.96)) is
w=z1/2. (6.100)
448 Chapter 6 Functions of a Complex Variable I
FIGURE 6.21Mapping—hyperboliccoordinates.
Fromtherelation
ρeiϕ=r1/2eiθ/2(6.101)
and
2ϕ=θ, (6.102)
we now have two points in the w-plane (arguments ϕandϕ+π) corresponding to one
point in the z-plane (except for the point z=0). Or, to put it another way, θandθ+2π
correspondto ϕandϕ+π,twodistinctpointsinthe w-plane.Thisisthecomplexvariable
analog of the simple real variable equation y2=x, in which two values of y, plus and
minus,correspondtoeachvalueof x.
The important point here is that we can make the function wof Eq. (6.100) a single-
valuedfunctioninsteadofadouble-valuedfunctionifweagreetorestrict θtoarangesuch
as 0≤θ<2π. This may be done by agreeing never to cross the line θ=0i nt h ez-plane
(Fig.6.22).Suchalineofdemarcationiscalleda cutlineorbranchcut .Notethatbranch
pointsoccurinpairs.
Thecut line joins the two branch point singularities , here at 0 and ∞(for the latter,
transform z=1/tfort→0). Any line from z=0 to infinity would serve equally well.
The purpose of the cut line is to restrict the argument of z. The points zandzexp(2πi)
coincide in the z-plane but yield different points wand−w=wexp(πi)in thew-plane.
Henceintheabsenceofacutline,thefunction w=z1/2isambiguous.Alternatively,since
the function w=z1/2is double-valued, we can also glue two sheets of the complex z-
plane together along the branch cut so that arg (z)increases beyond 2 πalong the branch
cut and continues from 4 πon the second sheet to reach the same function values for z
as forze−4πi,that is, the start on the first sheet again. This construction is called the
Riemann surface ofw=z1/2. We shall encounter branch points and cut lines (branch
cuts)frequentlyinChapter7.
Thetransformation
w=ez(6.103)
leadsto
ρeiϕ=ex+iy, (6.104)
6.7 Mapping 449
FIGURE 6.22Acutline.
or
ρ=ex,ϕ=y. (6.105)
Ifyranges from 0 ≤y<2π(or−π<y≤π), thenϕcovers the same range. But this is
thewhole w-plane.Inotherwords, ahorizontalstripinthe z-planeofwidth 2 πmapsinto
theentire w-plane.Further,anypoint x+i(y+2nπ),inwhich nisanyinteger,mapsinto
the same point (by Eq. (6.104)) in the w-plane. We have a many-(infinitely many)-to-one
correspondence.
Finally,as theinverseofthefifthtransformation(Eq. (6.103)), wehave
w=lnz. (6.106)
Byexpandingit,weobtain
u+iv=lnreiθ=lnr+iθ. (6.107)
Foragivenpoint z0inthez-planetheargument θisunspecifiedwithinanintegralmultiple
of 2π.This meansthat
v=θ+2nπ, (6.108)
and, as in the exponential transformation, we have an infinitely many-to-one correspon-
dence.
Equation (6.108) has a nice physical representation. If we go around the unit circle in
thez-plane,r=1,andbyEq.(6.107), u=lnr=0;butv=θ,andθissteadilyincreasing
andcontinuestoincreaseas θcontinuespast 2 π.
The cut line joins the branch point at the origin with infinity. As θincreases past 2 π
we glue a new sheet of the complex z-plane along the cut line, etc. Going around the unit
circle in the z-plane is like the advance of a screw as it is rotated or the ascent of a person
walkingupaspiralstaircase(Fig.6.23), whichis the Riemannsurface ofw=lnz.
As in the preceding example, we can also make the correspondence unique (and
Eq. (6.106) unambiguous) by restricting θto a range such as 0 ≤θ<2πby taking the
450 Chapter 6 Functions of a Complex Variable I
FIGURE 6.23This istheRiemann
surface for ln z, amultivalued
function.
lineθ=0 (positive real axis) as a cut line. This is equivalent to taking one and only one
completeturnof thespiralstaircase.
The concept of mapping is a very broad and useful one in mathematics. Our mapping
from a complex z-plane to a complex w-plane is a simple generalization of one definition
of function: a mapping of x(from one set) into yin a second set. A more sophisticated
form of mapping appears in Section 1.15 where we use the Dirac delta function δ(x−a)
tomapafunction f(x)intoitsvalueatthepoint a.TheninChapter15integraltransforms
are used to map one function f(x)inx-space into a second (related) function F(t)in
t-space.
Exercises
6.7.1 Howdocirclescenteredontheorigininthe z-planetransform for
(a)w1(z)=z+1
z,(b)w2(z)=z−1
z,forz/negationslash=0?
Whathappenswhen |z|→1?
6.7.2 Whatpartof the z-planecorrespondstotheinterioroftheunitcircleinthe w-planeif
(a)w=z−1
z+1,(b)w=z−i
z+i?
6.7.3 Discussthetransformations
(a)w(z)=sinz,(c)w(z)=sinhz,
(b)w(z)=cosz,(d)w(z)=coshz.
Show how the lines x=c1,y=c2m a pi n t ot h e w-plane. Note that the last three trans-
formationscanbeobtainedfromthefirstonebyappropriatetranslationand/orrotation.
6.8 Conformal Mapping 451
FIGURE 6.24Besselfunctionintegrationcontour.
6.7.4 Showthatthefunction
w(z)=parenleftbig
z2−1parenrightbig1/2
issingle-valuedif wetake −1≤x≤1,y=0 asacutline.
6.7.5 Show that negative numbers have logarithms in the complex plane. In particular, find
ln(−1).
ANS. ln(−1)=iπ.
6.7.6 An integral representation of the Bessel function follows the contour in the t-plane
shown in Fig. 6.24. Map this contour into the θ-plane with t=eθ. Many additional
examplesof mappingare giveninChapters11,12,and13.
6.7.7 For noninteger m, show that the binomial expansion of Exercise 6.5.2 holds only for a
suitably defined branch of the function (1+z)m. Show how the z-plane is cut. Explain
why|z|<1 maybetakenasthecircleofconvergencefor theexpansionofthis branch,
inlightof thecutyouhavechosen.
6.7.8 The Taylor expansion of Exercises 6.5.2 and 6.7.7 is notsuitable for branches other
than the one suitably defined branch of the function (1+z)mfor noninteger m.[ N o t e
that other branches cannot have the same Taylor expansion since they must be distin-
guishable.] Using the same branch cut of the earlier exercises for all other branches,
find the corresponding Taylor expansions, detailing the phase assignments and Taylor
coefficients.
6.8 C ONFORMAL MAPPING
In Section 6.7 hyperbolas were mapped into straight lines and straight lines were mapped
into circles. Yet in all these transformations one feature stayed constant. This constancy
wasaresultof thefactthatallthetransformationsofSection6.7wereanalytic.
Aslongas w=f(z)is ananalyticfunction,wehave
df
dz=dw
dz=lim
/Delta1z→0/Delta1w
/Delta1z. (6.109)
452 Chapter 6 Functions of a Complex Variable I
FIGURE 6.25Conformalmapping—preservationofangles.
Assuming that this equation is in polar form, we may equate modulus to modulus and
argumenttoargument.Forthelatter(assumingthat df/dz/negationslash=0),
arg lim
/Delta1z→0/Delta1w
/Delta1z=lim
/Delta1z→0arg/Delta1w
/Delta1z
=lim
/Delta1z→0arg/Delta1w−lim
/Delta1z→0arg/Delta1z=argdf
dz=α, (6.110)
whereα,theargumentofthederivative,maydependon zbutisaconstantforafixed z,in-
dependentofthedirectionofapproach.Toseethesignificanceofthis,considertwocurves
Czin thez-plane and the corresponding curve Cwin thew-plane (Fig. 6.25). The incre-
ment/Delta1zis shown at an angle of θrelative to the real (x)axis, whereas the corresponding
increment /Delta1wforms anangleof ϕwiththereal (u)axis. FromEq. (6.110),
ϕ=θ+α, (6.111)
or any line in the z-plane is rotated through an angle αin thew-plane as long as wis an
analytictransformationandthederivativeisnotzero.17
Since this result holds for any line through z0, it will hold for a pair of lines. Then for
theanglebetweenthesetwolines,
ϕ2−ϕ1=(θ2+α)−(θ1+α)=θ2−θ1, (6.112)
which shows that the included angle is preserved under an analytic transformation. Such
angle-preserving transformations are called conformal . The rotation angle αwill, in gen-
eral,dependon z. In addition|f′(z)|willusuallybeafunctionof z.
Historically,theseconformaltransformationshavebeenofgreatimportancetoscientists
and engineers in solving Laplace’s equation for problems of electrostatics, hydrodynam-
ics, heat flow, and so on. Unfortunately, the conformal transformation approach, however
elegant,islimitedtoproblemsthatcanbereducedtotwodimensions.Themethodisoften
beautiful if there is a high degree of symmetry present but often impossible if the sym-
metry is broken or absent. Because of these limitations and primarily because electronic
computers offer a useful alternative (iterative solution of the partial differential equation),
thedetailsandapplicationsofconformalmappingsareomitted.
17Ifdf/dz=0, its argument orphase is undefined andthe (analytic) transformation will not necessarilypreserve angles.
6.8 Additional Readings 453
Exercises
6.8.1 Expandw(x)in a Taylor series about the point z=z0, wheref′(z0)=0. (Angles are
not preserved.) Show that if the first n−1 derivatives vanish but f(n)(z0)/negationslash=0, then
anglesinthe z-planewithverticesat z=z0appearinthe w-planemultipliedby n.
6.8.2 Developthetransformationsthatcreateeachofthefourcylindricalcoordinatesystems:
(a) Circularcylindrical: x=ρcosϕ,
y=ρsinϕ.
(b) Ellipticcylindrical: x=acoshucosv,
y=asinhusinv.
(c) Paraboliccylindrical: x=ξη,
y=1
2parenleftbig
η2−ξ2parenrightbig
.
(d) Bipolar: x=asinhη
coshη−cosξ,
y=asinξ
coshη−cosξ.
Note.Thesetransformationsarenotnecessarilyanalytic.
6.8.3 Inthetransformation
ez=a−w
a+w,
howdothecoordinatelinesinthe z-planetransform?Whatcoordinatesystemhaveyou
constructed?
AdditionalReadings
Ahlfors, L. V., Complex Analysis , 3rd ed. New York: McGraw-Hill (1979). This text is detailed, thorough, rigor-
ous, and extensive.
Churchill,R.V.,J.W.Brown,andR.F.Verkey, ComplexVariablesandApplications ,5thed.NewYork:McGraw-
Hill (1989). This is an excellent text for both the beginning and advanced student. It is readable and quite
complete. Adetailed proof of the Cauchy–Goursat theorem is given in Chapter5.
Greenleaf, F. P., Introduction to Complex Variables . Philadelphia: Saunders (1972). This very readable book has
detailed,careful explanations.
Kurala, A., Applied Functions of a Complex Variable . New York: Wiley (Interscience) (1972). An intermediate-
level text designed for scientists andengineers. Includes many physical applications.
Levinson, N., and R. M. Redheffer, Complex Variables . San Francisco: Holden-Day (1970). This text is written
for scientists andengineers whoare interested in applications.
Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 4 is
apresentation of portions of thetheory of functions of a complex variableof interest to theoretical physicists.
Remmert, R., Theory of Complex Functions . NewYork: Springer (1991).
Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering , 2nd ed. New York:
McGraw-Hill(1966). Chapter 7 covers complex variables.
Spiegel, M. R., Complex Variables . New York: McGraw-Hill (1985). An excellent summary of the theory of
complex variables for scientists.
T itch m ar sh ,E.C., The Theory of Functions , 2nd ed. NewYork: Oxford University Press (1958). A classic.
454 Chapter 6 Functions of a Complex Variable I
Watson, G. N., Complex Integration and Cauchy’s Theorem . New York: Hafner (orig. 1917, reprinted 1960).
A short work containing a rigorous development of the Cauchy integral theorem and integral formula. Appli-
cationstothecalculusofresiduesareincluded. CambridgeTractsinMathematics,andMathematicalPhysics ,
No. 15.
Otherreferencesaregivenattheendof Chapter15.
CHAPTER 7
FUNCTIONS OF A COMPLEX
VARIABLE II
In this chapter we return to the analysis that started with the Cauchy–Riemann conditions
in Chapter 6 and develop the residue theorem, with major applications to the evaluation
of definite and principal part integrals of interest to scientists and asymptotic expansion
of integrals by the method of steepest descent. We also develop further specific analytic
functions, such as pole expansions of meromorphic functions and product expansions of
entire functions. Dispersion relations are included because they represent an important
applicationofcomplexvariablemethodsforphysicists.
7.1 C ALCULUS OF RESIDUES
Residue Theorem
If the Laurent expansion of a function f(z)=summationtext∞
n=−∞an(z−z0)nis integrated term by
termbyusingaclosedcontourthatencirclesoneisolatedsingularpoint z0onceinacoun-
terclockwisesense, weobtain(Exercise6.4.1)
ancontintegraldisplay
(z−z0)ndz=an(z−z0)n+1
n+1vextendsinglevextendsinglevextendsinglevextendsinglez1
z1=0,n/negationslash=−1. (7.1)
However,if n=−1,
a−1contintegraldisplay
(z−z0)−1dz=a−1contintegraldisplayireiθdθ
reiθ=2πia−1. (7.2)
SummarizingEqs. (7.1) and(7.2), wehave
1
2πicontintegraldisplay
f(z)dz=a−1. (7.3)
455
456 Chapter 7 Functions of a Complex Variable II
FIGURE 7.1Excludingisolated
singularities.
The constant a−1,the coefficient of (z−z0)−1in the Laurent expansion, is called the
residueof f(z)atz=z0.
A set of isolated singularities can be handled by deforming our contour as shown in
Fig.7.1.Cauchy’sintegraltheorem(Section6.3) leadsto
contintegraldisplay
Cf(z)dz+contintegraldisplay
C0f(z)dz+contintegraldisplay
C1f(z)dz+contintegraldisplay
C2f(z)dz+···=0.(7.4)
Thecircularintegralaroundanygivensingularpointis givenbyEq. (7.3),
contintegraldisplay
Cif(z)dz=−2πia−1,zi, (7.5)
assuming a Laurent expansion about the singular point z=zi. The negative sign comes
from the clockwise integration, as shown in Fig. 7.1. Combining Eqs. (7.4) and (7.5), we
have
contintegraldisplay
Cf(z)dz=2πi(a−1z0+a−1z1+a−1z2+···)
=2πi×(sum ofenclosedresidues ). (7.6)
This is the residue theorem . The problem of evaluating one or more contour integrals is
replacedbythealgebraicproblemof computingresiduesattheenclosedsingularpoints.
We first use this residue theorem to develop the concept of the Cauchy principal value.
Then in the remainder of this section we apply the residue theorem to a wide variety of
definiteintegralsofmathematicalandphysicalinterest.
Using the transformation z=1/wforwapproaching 0, we can find the nature of a sin-
gularity at zgoing to∞and the residue of a function f(z)with just isolated singularities
andnobranchpoints.Insuchcasesweknowthat
summationdisplay
{residuesinthefinite z-plane}+{residueat z→∞}= 0.
7.1 Calculus of Residues 457
Cauchy Principal Value
Occasionally an isolated pole will be directly on the contour of integration, causing the
integraltodiverge.Letus illustrateaphysicalcase.
Example 7.1.1 FORCED CLASSICAL OSCILLATOR
The inhomogeneous differential equation for a classical, undamped, driven harmonic os-
cillator,
¨x(t)+ω2
0x(t)=f(t), (7.7)
may be solved by representing the driving force f(t)=integraltext
δ(t′−t)f(t′)dt′as a superpo-
sitionofimpulsesbyanalogywithanextendedchargedistributioninelectrostatics.1Ifwe
solvefirst thesimplerdifferentialequation
¨G+ω2
0G=δ(t−t′) (7.8)
forG(t,t′), which is independent of the driving term f(model dependent), then x(t)=integraltext
G(t,t′)f(t′)dt′solves the original problem. First, we verify this by substituting the in-
tegralsfor x(t)anditstimederivativesintothedifferentialequationfor x(t)usingthedif-
ferentialequationfor G.Thenwelookfor G(t,t′)=integraltext˜G(ω)eiωtdω
2πintermsofanintegral
weightedby˜G,whichissuggestedbyasimilarintegralformfor δ(t−t′)=integraltext
eiω(t−t′)dω
2π
(seeEq. (1.193c)inSection1.15).
Uponsubstituting Gand¨Gintothedifferentialequationfor G,weobtain
integraldisplaybracketleftbigparenleftbig
ω2
0−ω2parenrightbig˜G−e−iωt′bracketrightbig
eiωtdω=0. (7.9)
Because this integral is zero for all t,the expression in brackets must vanish for all ω.
Thisrelationisnolongeradifferentialequationbutanalgebraicrelationthatwecansolve
for˜G:
˜G(ω)=e−iωt′
ω2
0−ω2=e−iωt′
2ω0(ω+ω0)−e−iωt′
2ω0(ω−ω0). (7.10)
Substituting˜Gintotheintegralfor Gyields
G(t,t′)=1
4πω0integraldisplay∞
−∞bracketleftbiggeiω(t−t′)
ω+ω0−eiω(t−t′)
ω−ω0bracketrightbigg
dω. (7.11)
Here, the dependence of Gont−t′in the exponential is consistent with the same depen-
denceofδ(t−t′),itsdrivingterm.Now,theproblemisthatthisintegraldivergesbecause
theintegrandblowsupat ω=±ω0,sincetheintegrationgoesrightthroughthefirst-order
poles. To explain why this happens, we note that the δ-function driving term for Gin-
cludes all frequencies with the same amplitude. Next, we see that the equation for ˜Gat
t′=0 has its driving term equal to unity for all frequencies ω, including the resonant ω0.
1Adapted from A.Yu. Grosberg, priv. comm.
458 Chapter 7 Functions of a Complex Variable II
Weknowfromphysicsthatforcinganoscillatoratresonanceleadstoanindefinitelygrow-
ing amplitude when there is no friction. With friction, the amplitude remains finite, even
at resonance. This suggests includinga small friction term in the differential equationsfor
x(t)andG.
Withasmallfrictionterm η˙G,η>0,inthedifferentialequationfor G(t,t′)(andη˙xfor
x(t)), wecanstillsolvethealgebraicequation
parenleftbig
ω2
0−ω2+iηωparenrightbig˜G=e−iωt′(7.12)
for˜Gwithfriction. Thesolutionis
˜G=e−iωt′
ω2
0−ω2+iηω=e−iωt′
2/Omega1parenleftbigg1
ω−ω−−1
ω−ω+parenrightbigg
, (7.13)
ω±=±/Omega1+iη
2,/Omega1=ω0radicalBigg
1−parenleftbiggη
2ω0parenrightbigg2
. (7.14)
Forsmallfriction,0 <η≪ω0,/Omega1isnearlyequalto ω0andreal,whereas ω±eachpickup
asmallimaginarypart. Thismeansthattheintegrationof theintegralfor G,
G(t,t′)=1
4π/Omega1integraldisplay∞
−∞bracketleftbiggeiω(t−t′)
ω−ω−−eiω(t−t′)
ω−ω+bracketrightbigg
dω, (7.15)
nolongerencountersapoleandremainsfinite. /squaresolid
This treatment of an integral with a pole moves the pole off the contour and then con-
siders the limiting behavior as it is brought back, as in Example 7.1.1 for η→0.This
example also suggests treating ωas a complex variable in case the singularity is a first-
order pole, deforming the integration path to avoid the singularity, which is equivalent to
addingasmallimaginaryparttothepoleposition,andevaluatingtheintegralbymeansof
theresiduetheorem.
Therefore,iftheintegrationpathofanintegralintegraltextdz
z−x0forrealx0goesrightthroughthe
polex0,wemaydeformthecontourtoincludeorexcludetheresidue,asdesired,byinclud-
ingasemicirculardetourof infinitesimalradius .ThisisshowninFig.7.2.Theintegration
overthesemicirclethengives,with z−x0=δeiϕ,dz=iδeiϕdϕ(seeEq. (6.27a)),
integraldisplaydz
z−x0=iintegraldisplay2π
πdϕ=iπ,i.e.,πia−1 ifcounterclockwise ,
integraldisplaydz
z−x0=iintegraldisplay0
πdϕ=−iπ,i.e.,−πia−1ifclockwise .
This contribution, +or−, appears on the left-hand side of Eq. (7.6). If our detour were
clockwise, the residue would not be enclosed and there would be no corresponding term
ontheright-handsideof Eq. (7.6).
However, if our detour were counterclockwise, this residue would be enclosed by the
contourCandaterm 2 πia−1wouldappearontheright-handsideofEq. (7.6).
The net result for either clockwise or counterclockwise detour is that a simple pole on
the contour is counted as one-half of what it would be if it were within the contour. This
correspondstotakingtheCauchyprincipalvalue.
7.1 Calculus of Residues 459
FIGURE 7.2Bypassingsingularpoints.
FIGURE 7.3Closingthecontour
withaninfinite-radiussemicircle.
Forinstance,letus supposethat f(z)withas implepoleat z=x0isintegratedoverthe
entire real axis. The contour is closed with an infinite semicircle in the upper half-plane
(Fig.7.3). Then
contintegraldisplay
f(z)dz=integraldisplayx0−δ
−∞f(x)dx+integraldisplay
Cx0f(z)dz
+integraldisplay∞
x0+δf(x)dx+integraldisplay
Cinfinitesemicircle
=2πisummationdisplay
enclosed residues. (7.16)
If the smallsemicircle Cx0, includes x0(bygoingbelowthe x-axis, counterclockwise), x0
is enclosed, and its contribution appears twice—asπia−1inintegraltext
Cx0and as 2πia−1in the
term 2πisummationtextenclosed residues—for a net contribution of πia−1. If the upper small semi-
circle is selected, x0is excluded. The only contribution is from the clockwise integration
overCx0, which yields −πia−1. Moving this to the extreme right of Eq. (7.16), we have
+πia−1, asbefore.
The integrals along the x-axis may be combined and the semicircle radius permitted to
approachzero.Wethereforedefine
lim
δ→0braceleftbiggintegraldisplayx0−δ
−∞f(x)dx+integraldisplay∞
x0+δf(x)dxbracerightbigg
=Pintegraldisplay∞
−∞f(x)dx. (7.17)
Pindicates the Cauchy principal value and represents the preceding limiting process.
Note that the Cauchy principal value is a balancing (or canceling) process. In the vicinity
ofoursingularityat z=x0,
f(x)≈a−1
x−x0. (7.18)
460 Chapter 7 Functions of a Complex Variable II
FIGURE 7.4Cancellationatasimplepole.
Thisisodd,relativeto x0.Thesymmetricoreveninterval(relativeto x0)providescancel-
lationof theshadedareas, Fig.7.4. The contributionof thesingularityis intheintegration
aboutthesemicircle.
In general, if a function f(x)has a singularity x0somewhere inside the interval a≤
x0≤band is integrable over every portion of this interval that does not contain the point
x0,thenwedefine
integraldisplayb
af(x)dx=lim
δ1→0integraldisplayx0−δ1
af(x)dx+lim
δ2→0integraldisplayb
x0+δ2f(x)dx,
when the limit exists as δj→0independently , else the integral is said to diverge. If this
limit does not exist but the limit δ1=δ2=δ→0 exists, it is defined to be the principal
valueoftheintegral.
This samelimitingtechniqueis applicabletotheintegrationlimits ±∞.We define
integraldisplay∞
−∞f(x)dx=lim
a→−∞,b→∞integraldisplayb
af(x)dx, (7.19a)
if the integral exists with a,bapproaching their limits independently, else the integral di-
verges.Incasetheintegraldivergesbut
lima→∞integraldisplaya
−af(x)dx=Pintegraldisplay∞
−∞f(x)dx (7.19b)
exist,itisdefinedas itsprincipalvalue.
7.1 Calculus of Residues 461
Pole Expansion of Meromorphic Functions
Analyticfunctions f(z)thathaveonlyisolatedpolesassingularitiesarecalled meromor-
phic. Examples are cot z[fromd
dzlnsinzin Eq. (5.210)] and ratios of polynomials. For
simplicity we assume that these poles at finite z=anwith 0<|a1|<|a2|<···are all
simple with residues bn. Then an expansion of f(z)in terms of bn(z−an)−1depends in
a systematic way on all singularities of f(z), in contrast to the Taylor expansion about
an arbitrarily chosen analytic point z0off(z)or the Laurent expansion about one of the
singularpointsof f(z).
Let us consider a series of concentric circles Cnabout the origin so that Cnincludes
a1,a2,...,anbutnootherpoles,itsradius Rn→∞asn→∞.Toguaranteeconvergence
we assume that |f(z)|<εRnfor any small positive constant εand allzonCn. Then the
series
f(z)=f(0)+∞summationdisplay
n=1bnbraceleftbig
(z−an)−1+a−1
nbracerightbig
(7.20)
convergesto f(z).Toprovethis theorem (duetoMittag–Leffler)weusetheresiduetheo-
remtoevaluatethecontourintegralfor zinsideCn:
In=1
2πiintegraldisplay
Cnf(w)
w(w−z)dw
=nsummationdisplay
m=1bm
am(am−z)+f(z)−f(0)
z. (7.21)
OnCnwehave,for n→∞,
|In|≤2πRnmaxwonCn|f(w)|
2πRn(Rn−|z|)<εRn
Rn−|z|→ε
forRn≫|z|.U s i n gIn→0 inEq. (7.21)provesEq.(7.20).
If|f(z)|<εRp+1
n, thenweevaluatesimilarlytheintegral
In=1
2πiintegraldisplayf(w)
wp+1(w−z)dw→0asn→∞
andobtaintheanalogouspoleexpansion
f(z)=f(0)+zf′(0)+···+zpf(p)(0)
p!+∞summationdisplay
n=1bnzp+1/ap+1
n
z−an.(7.22)
NotethattheconvergenceoftheseriesinEqs.(7.20)and(7.22)isimpliedbytheboundof
|f(z)|for|z|→∞.
462 Chapter 7 Functions of a Complex Variable II
Product Expansion of Entire Functions
Afunction f(z)thatisanalyticforallfinite ziscalledan entirefunction.Thelogarithmic
derivative f′/fisameromorphicfunctionwithapoleexpansion.
Iff(z)has a simple zero at z=an, thenf(z)=(z−an)g(z)with analytic g(z)and
g(an)/negationslash=0.Hencethelogarithmicderivative
f′(z)
f(z)=(z−an)−1+g′(z)
g(z)(7.23)
has a simple pole at z=anwith residue 1, and g′/gis analytic there. If f′/fsatisfies the
conditionsthatleadtothepoleexpansioninEq. (7.20), then
f′(z)
f(z)=f′(0)
f(0)+∞summationdisplay
n=1bracketleftbigg1
an+1
z−anbracketrightbigg
(7.24)
holds.IntegratingEq.(7.24) yields
integraldisplayz
0f′(z)
f(z)dz=lnf(z)−lnf(0)
=zf′(0)
f(0)+∞summationdisplay
n=1braceleftbigg
ln(z−an)−ln(−an)+z
anbracerightbigg
,
andexponentiatingweobtaintheproductexpansion
f(z)=f(0)expparenleftbiggzf′(0)
f(0)parenrightbigg∞productdisplay
1parenleftbigg
1−z
anparenrightbigg
ez/an. (7.25)
Examplesaretheproductexpansions(seeChapter5)for
sinz=z∞productdisplay
n=−∞
n/negationslash=0parenleftbigg
1−z
nπparenrightbigg
ez/nπ=z∞productdisplay
n=1parenleftbigg
1−z2
n2π2parenrightbigg
,
cosz=∞productdisplay
n=1braceleftbigg
1−z2
(n−1/2)2π2bracerightbigg
.(7.26)
Anotherexampleistheproductexpansionofthegammafunction,whichwillbediscussed
inChapter8.
AsaconsequenceofEq.(7.23)thecontourintegralofthelogarithmicderivativemaybe
used to count the number Nfof zeros (including their multiplicities) of the function f(z)
insidethecontour C:
1
2πiintegraldisplay
Cf′(z)
f(z)dz=Nf. (7.27)
7.1 Calculus of Residues 463
Moreover, using
integraldisplayf′(z)
f(z)dz=lnf(z)=lnvextendsinglevextendsinglef(z)vextendsinglevextendsingle+iargf(z), (7.28)
weseethattherealpartinEq.(7.28)doesnotchangeas zmovesoncearoundthecontour,
whilethecorrespondingchangein arg fmustbe
/Delta1Carg(f)=2πNf. (7.29)
This leads to Rouché’s theorem :Iff(z)andg(z)are analytic inside and on a closed
contourCand|g(z)|<|f(z)|onCthenf(z)andf(z)+g(z)have the same number of
zeros inside C.
Toshowthisweuse
2πNf+g=/Delta1Carg(f+g)=/Delta1Carg(f)+/Delta1Cargparenleftbigg
1+g
fparenrightbigg
.
Since|g|<|f|onC, thepoint w=1+g(z)/f(z) is always an interiorpointof the circle
inthew-planewithcenterat1andradius1.Hencearg (1+g/f)mustreturntoitsoriginal
valuewhen zmovesaround C(itdoesnotcircletheorigin);itcannotdecreaseorincrease
byamultipleof 2 πso that/Delta1Carg(1+g/f)=0.
Rouché’s theorem may be used for an alternative proof of the fundamental theorem of
algebra:Apolynomialsummationtextn
m=0amzmwithan/negationslash=0hasnzeros.Wedefine f(z)=anzn.Then
fhas ann-fold zero at the origin and no other zeros. Let g(z)=summationtextn−1
m=0amzm. We apply
Rouché’stheoremtoacircle Cwithcenterattheoriginandradius R>1.OnC,|f(z)|=
|an|Rnand
vextendsinglevextendsingleg(z)vextendsinglevextendsingle≤|a0|+|a1|R+···+|an−1|Rn−1≤parenleftbiggn−1summationdisplay
m=0|am|parenrightbigg
Rn−1.
Hence|g(z)|<|f(z)|forzonC, provided R>(summationtextn−1
m=0|am|)/|an|. For all sufficiently
largecircles Ctherefore, f+g=summationtextn
m=0amzmhasnzerosinside CaccordingtoRouché’s
theorem.
Evaluation of Definite Integrals
Definiteintegralsappearrepeatedlyinproblemsofmathematicalphysicsaswellasinpure
mathematics. Three moderately general techniques are useful in evaluating definite inte-
grals: (1) contour integration, (2) conversion to gamma or beta functions (Chapter 8), and
(3) numerical quadrature. Other approaches include series expansion with term-by-term
integration and integral transforms. As will be seen subsequently, the method of contour
integration is perhaps the most versatile of these methods, since it is applicable to a wide
varietyof integrals.
464 Chapter 7 Functions of a Complex Variable II
Definite Integrals:/integraltext2π
0f(sinθ,cosθ)dθ
The calculus of residues is useful in evaluating a wide variety of definite integrals in both
physicalandpurelymathematicalproblems.We consider,first, integralsoftheform
I=integraldisplay2π
0f(sinθ,cosθ)dθ, (7.30)
wherefisfiniteforallvaluesof θ.Wealsorequire ftobearationalfunctionofsin θand
cosθso thatitwillbesingle-valued.Let
z=eiθ,dz=ieiθdθ.
Fromthis,
dθ=−idz
z,sinθ=z−z−1
2i,cosθ=z+z−1
2. (7.31)
Ourintegralbecomes
I=−icontintegraldisplay
fparenleftbiggz−z−1
2i,z+z−1
2parenrightbiggdz
z, (7.32)
withthepathof integrationtheunitcircle.By theresiduetheorem,Eq.(7.16),
I=(−i)2πisummationdisplay
residueswithintheunitcircle. (7.33)
Notethatweareaftertheresiduesof f/z.Illustrationsofintegralsofthistypeareprovided
byExercises7.1.7–7.1.10.
Example 7.1.2 INTEGRAL OF COS IN DENOMINATOR
Ourproblemistoevaluatethedefiniteintegral
I=integraldisplay2π
0dθ
1+εcosθ,|ε|<1.
ByEq. (7.32)thisbecomes
I=−icontintegraldisplay
unit circledz
z[1+(ε/2)(z+z−1)]
=−i2
εcontintegraldisplaydz
z2+(2/ε)z+1.
Thedenominatorhasroots
z−=−1
ε−1
εradicalbig
1−ε2andz+=−1
ε+1
εradicalbig
1−ε2,
wherez+iswithintheunitcircleand z−isoutside.ThenbyEq.(7.33)andExercise6.6.1,
I=−i2
ε·2πi1
2z+2/εvextendsinglevextendsinglevextendsinglevextendsingle
z=−1/ε+(1/ε)√
1−ε2.
7.1 Calculus of Residues 465
Weobtain
integraldisplay2π
0dθ
1+εcosθ=2π√
1−ε2,|ε|<1./squaresolid
Evaluation of Definite Integrals:/integraltext∞
−∞f( x)dx
Supposethatour definiteintegralhastheform
I=integraldisplay∞
−∞f(x)dx (7.34)
andsatisfiesthetwoconditions:
•f(z)is analytic in the upper half-plane except for a finite number of poles. (It will be
assumed that there are no poles on the real axis. If poles are present on the real axis,
theymaybeincludedorexcludedas discussedearlierinthissection.)
•f(z)vanishesasstrongly2as 1/z2for|z|→∞,0≤argz≤π.
With these conditions, we may take as a contour of integration the real axis and a semi-
circle in the upper half-plane, as shown in Fig. 7.5. We let the radius Rof the semicircle
becomeinfinitelylarge.Then
contintegraldisplay
f(z)dz=lim
R→∞integraldisplayR
−Rf(x)dx+lim
R→∞integraldisplayπ
0fparenleftbig
Reiθparenrightbig
iReiθdθ
=2πisummationdisplay
residues(upperhalf-plane) . (7.35)
Fromthesecondconditionthesecondintegral(overthesemicircle)vanishesand
integraldisplay∞
−∞f(x)dx=2πisummationdisplay
residues(upperhalf-plane) . (7.36)
FIGURE 7.5Half-circle
contour.
2Wecould use f(z)vanishes fasterthan 1 /z, and wewish to have f(z)single-valued.
466 Chapter 7 Functions of a Complex Variable II
Example 7.1.3 INTEGRAL OF MEROMORPHIC FUNCTION
Evaluate
I=integraldisplay∞
−∞dx
1+x2. (7.37)
FromEq. (7.36),
integraldisplay∞
−∞dx
1+x2=2πisummationdisplay
residues(upperhalf-plane) .
Hereandineveryothersimilarproblemwehavethequestion:Wherearethepoles?Rewrit-
ingtheintegrandas
1
z2+1=1
z+i·1
z−i, (7.38)
wesee thattherearesimplepoles(order 1)at z=iandz=−i.
Asimplepoleat z=z0indicates(andisindicatedby)aLaurentexpansionof theform
f(z)=a−1
z−z0+a0+∞summationdisplay
n=1an(z−z0)n. (7.39)
Theresidue a−1iseasilyisolatedas (Exercise6.6.1)
a−1=(z−z0)f(z)|z=z0. (7.40)
UsingEq.(7.40),wefindthattheresidueat z=iis1/2i,whereasthatat z=−iis−1/2i.
Then
integraldisplay∞
−∞dx
1+x2=2πi·1
2i=π. (7.41)
Here we have used a−1=1/2ifor the residue of the one included pole at z=i. Note that
it is possible to use the lower semicircle and that this choice will lead to the same result,
I=π. Asomewhatmoredelicateproblemis providedbythenextexample. /squaresolid
Evaluation of Definite Integrals:/integraltext∞
−∞f( x) eiaxdx
Considerthedefiniteintegral
I=integraldisplay∞
−∞f(x)eiaxdx, (7.42)
withareal and positive. (This is a Fourier transform, Chapter 15.) We assume the two
conditions:
•f(z)isanalyticintheupperhalf-planeexceptfor afinitenumberofpoles.
7.1 Calculus of Residues 467
•lim
|z|→∞f(z)=0,0≤argz≤π. (7.43)
Notethat this is a less restrictive conditionthan the second conditionimposedon f(z)for
integratingintegraltext∞
−∞f(x)dxpreviously.
We employ the contour shown in Fig. 7.5. The application of the calculus of residues
is the same as the one just considered, but here we have to work harder to show that the
integraloverthe(infinite)semicirclegoestozero.This integralbecomes
IR=integraldisplayπ
0fparenleftbig
Reiθparenrightbig
eiaRcosθ−aRsinθiReiθdθ. (7.44)
LetRbeso largethat |f(z)|=|f(Reiθ)|<ε. Then
|IR|≤εRintegraldisplayπ
0e−aRsinθdθ=2εRintegraldisplayπ/2
0e−aRsinθdθ. (7.45)
Intherange[0,π/2],
2
πθ≤sinθ.
Therefore(Fig.7.6)
|IR|≤2εRintegraldisplayπ/2
0e−2aRθ/πdθ. (7.46)
Now,integratingbyinspection,weobtain
|IR|≤2εR1−e−aR
2aR/π.
Finally,
lim
R→∞|IR|≤π
aε. (7.47)
FromEq. (7.43), ε→0a sR→∞and
lim
R→∞|IR|=0. (7.48)
FIGURE 7.6(a)y=(2/π)θ,( b )y=sinθ.
468 Chapter 7 Functions of a Complex Variable II
This useful result is sometimescalled Jordan’slemma . With it, we are prepared to tackle
Fourierintegralsof theformshowninEq.(7.42).
UsingthecontourshowninFig.7.5, wehave
integraldisplay∞
−∞f(x)eiaxdx+lim
R→∞IR=2πisummationdisplay
residues(upperhalf-plane) .
Sincetheintegralovertheuppersemicircle IRvanishesas R→∞(Jordan’s lemma),
integraldisplay∞
−∞f(x)eiaxdx=2πisummationdisplay
residues(upperhalf-plane )(a>0).(7.49)
Example 7.1.4 SIMPLE POLE ON CONTOUR OF INTEGRATION
Theproblemistoevaluate
I=integraldisplay∞
0sinx
xdx. (7.50)
Thismaybetakenas theimaginarypart3of
I2=Pintegraldisplay∞
−∞eizdz
z. (7.51)
Nowtheonlypoleisasimplepoleat z=0andtheresiduetherebyEq.(7.40)is a−1=1.
We choose the contour shown in Fig. 7.7 (1) to avoid the pole, (2) to include the real axis,
and (3) to yield a vanishingly small integrand for z=iy,y→∞. Note that in this case a
large(infinite)semicircleinthelowerhalf-planewouldbedisastrous. Wehave
contintegraldisplayeizdz
z=integraldisplay−r
−Reixdx
x+integraldisplay
C1eizdz
z+integraldisplayR
reixdx
x+integraldisplay
C2eizdz
z=0,(7.52)
FIGURE 7.7Singularityoncontour.
3One can useintegraltext
[(eiz−e−iz)/2iz]dz, but then two different contours will be needed for the two exponentials (compare Exam-
ple 7.1.5).
7.1 Calculus of Residues 469
thefinalzerocomingfromtheresiduetheorem(Eq. (7.6)). ByJordan’slemma
integraldisplay
C2eizdz
z=0, (7.53)
and
contintegraldisplayeizdz
z=integraldisplay
C1eizdz
z+Pintegraldisplay∞
−∞eixdx
x=0. (7.54)
Theintegraloverthesmallsemicircleyields (−)πitimestheresidueof1,andminus,asa
resultofgoingclockwise.Takingtheimaginarypart,4wehave
integraldisplay∞
−∞sinx
xdx=π (7.55)
orintegraldisplay∞
0sinx
xdx=π
2. (7.56)
The contour of Fig. 7.7, although convenient, is not at all unique. Another choice of
contourfor evaluatingEq. (7.50)is presentedas Exercise7.1.15. /squaresolid
Example 7.1.5 QUANTUM MECHANICAL SCATTERING
Thequantummechanicalanalysisofscatteringleadstothefunction
I(σ)=integraldisplay∞
−∞xsinxdx
x2−σ2, (7.57)
whereσis real and positive. This integral is divergent and therefore ambiguous. From the
physicalconditionsoftheproblemthereisafurtherrequirement: I(σ)istohavetheform
eiσso thatitwillrepresentanoutgoingscatteredwave.
Using
sinz=1
isinhiz=1
2ieiz−1
2ie−iz, (7.58)
wewriteEq. (7.57)inthecomplexplaneas
I(σ)=I1+I2, (7.59)
with
I1=1
2iintegraldisplay∞
−∞zeiz
z2−σ2dz,
I2=−1
2iintegraldisplay∞
−∞ze−iz
z2−σ2dz. (7.60)
4Alternatively, wemaycombine the integrals ofEq. (7.52) as
integraldisplay−r
−Reixdx
x+integraldisplayR
reixdx
x=integraldisplayR
rparenleftbig
eix−e−ixparenrightbigdx
x=2iintegraldisplayR
rsinx
xdx.
470 Chapter 7 Functions of a Complex Variable II
FIGURE 7.8Contours.
IntegralI1issimilartoExample7.1.4and,asinthatcase,wemaycompletethecontourby
aninfinitesemicircleintheupperhalf-plane,asshowninFig.7.8a.For I2theexponential
is negative and we complete the contour by an infinite semicircle in the lower half-plane,
as shown in Fig. 7.8b. As in Example 7.1.4, neither semicircle contributes anything to the
integral—Jordan’slemma.
Thereisstilltheproblemoflocatingthepolesandevaluatingtheresidues.Wefindpoles
atz=+σandz=−σon the contour of integration . The residues are (Exercises 6.6.1
and7.1.1)
z=σz=−σ
I1eiσ
2e−iσ
2
I2e−iσ
2eiσ
2
Detouring around the poles, as shown in Fig. 7.8 (it matters little whether we go above or
below),wefindthattheresiduetheoremleadsto
PI1−πiparenleftbigg1
2iparenrightbigge−iσ
2+πiparenleftbigg1
2iparenrightbiggeiσ
2=2πiparenleftbigg1
2iparenrightbiggeiσ
2, (7.61)
for we have enclosed the singularity at z=σbut excluded the one at z=−σ. In similar
fashion,butnotingthatthecontourfor I2is clockwise,
PI2−πiparenleftbigg−1
2iparenrightbiggeiσ
2+πiparenleftbigg−1
2iparenrightbigge−iσ
2=−2πiparenleftbigg−1
2iparenrightbiggeiσ
2. (7.62)
AddingEqs.(7.61) and(7.62), wehave
PI(σ)=PI1+PI2=π
2parenleftbig
eiσ+e−iσparenrightbig
=πcoshiσ=πcosσ. (7.63)
This is a perfectly good evaluation of Eq. (7.57), but unfortunately the cosine dependence
isappropriatefor astandingwaveandnotfortheoutgoingscatteredwaveas specified.
7.1 Calculus of Residues 471
To obtain the desired form, we try a different technique (compare Example 7.1.1). In-
steadofdodgingaroundthesingularpoints,letusmovethemofftherealaxis.Specifically,
letσ→σ+iγ,−σ→−σ−iγ,whereγispositivebutsmallandwilleventuallybemade
toapproachzero;thatis, for I1weincludeonepoleandfor I2theotherone,
I+(σ)=lim
γ→0I(σ+iγ). (7.64)
Withthissimplesubstitution,thefirstintegral I1becomes
I1(σ+iγ)=2πiparenleftbigg1
2iparenrightbiggei(σ+iγ)
2(7.65)
bydirectapplicationof theresiduetheorem.Also,
I2(σ+iγ)=−2πiparenleftbigg−1
2iparenrightbiggei(σ+iγ)
2. (7.66)
AddingEqs.(7.65) and(7.66)andthenletting γ→0,weobtain
I+(σ)=lim
γ→0bracketleftbig
I1(σ+iγ)+I2(σ+iγ)bracketrightbig
=lim
γ→0πei(σ+iγ)=πeiσ, (7.67)
aresultthatdoesfittheboundaryconditionsofourscatteringproblem.
It isinterestingtonotethatthesubstitution σ→σ−iγwouldhaveledto
I−(σ)=πe−iσ, (7.68)
which could represent an incoming wave. Our earlier result (Eq. (7.63)) is seen to be the
arithmeticaverageofEqs.(7.67)and(7.68).ThisaverageistheCauchyprincipalvalueof
the integral. Note that we have these possibilities (Eqs. (7.63), (7.67), and (7.68)) because
our integral is not uniquely defined until we specify the particular limiting process (or
average)tobeused. /squaresolid
Evaluation of Definite Integrals: Exponential Forms
Withexponentialorhyperbolicfunctionspresentintheintegrand,lifegetssomewhatmore
complicatedthanbefore.Insteadofageneraloverallprescription,thecontourmustbecho-
sentofitthespecificintegral.Thesecasesarealsoopportunitiestoillustratetheversatility
andpowerofcontourintegration.
Asanexample,weconsideranintegralthatwillbequiteusefulindevelopingarelation
betweenŴ(1+z)andŴ(1−z). Notice how the periodicity along the imaginary axis is
exploited.
472 Chapter 7 Functions of a Complex Variable II
FIGURE 7.9Rectangularcontour.
Example 7.1.6 FACTORIAL FUNCTION
We wish to evaluate
I=integraldisplay∞
−∞eax
1+exdx,0<a<1. (7.69)
The limits on aare sufficient (but not necessary) to prevent the integral from diverging as
x→±∞. This integral (Eq. (7.69)) may be handled by replacing the real variable xby
the complex variable zand integrating around the contour shown in Fig. 7.9. If we take
thelimitas R→∞,therealaxis,ofcourse,leadstotheintegralwewant.Thereturnpath
alongy=2πischosentoleavethedenominatoroftheintegralinvariant,atthesametime
introducingaconstantfactor ei2πainthenumerator.We have,inthecomplexplane,
contintegraldisplayeaz
1+ezdz=lim
R→∞parenleftbiggintegraldisplayR
−Reax
1+exdx−ei2πaintegraldisplayR
−Reax
1+exdxparenrightbigg
=parenleftbig
1−ei2πaparenrightbigintegraldisplay∞
−∞eax
1+exdx. (7.70)
In addition there are two vertical sections (0≤y≤2π), which vanish (exponentially) as
R→∞.
Nowwherearethepolesandwhataretheresidues?We haveapolewhen
ez=exeiy=−1. (7.71)
Equation (7.71) is satisfied at z=0+iπ. By a Laurent expansion5in powers of (z−iπ)
the pole is seen to be a simple pole with a residue of −eiπa. Then, applying the residue
theorem,
parenleftbig
1−ei2πaparenrightbigintegraldisplay∞
−∞eax
1+exdx=2πiparenleftbig
−eiπaparenrightbig
. (7.72)
Thisquicklyreducesto
integraldisplay∞
−∞eax
1+exdx=π
sinaπ,0<a<1. (7.73)
51+ez=1+ez−iπeiπ=1−ez−iπ=−(z−iπ)(1+z−iπ
2!+(z−iπ)2
3!+···).
7.1 Calculus of Residues 473
Usingthebetafunction(Section8.4),wecanshowtheintegraltobeequaltotheproduct
Ŵ(a)Ŵ(1−a). Thisresults intheinterestingandusefulfactorialfunctionrelation
Ŵ(a+1)Ŵ(1−a)=πa
sinπa. (7.74)
Although Eq. (7.73) holds for real a,0<a<1, Eq. (7.74) may be extended by analytic
continuationtoallvaluesof a, realandcomplex,excludingonlyrealintegralvalues. /squaresolid
As a final example of contour integrals of exponential functions, we consider Bernoulli
numbersagain.
Example 7.1.7 BERNOULLI NUMBERS
InSection5.9theBernoullinumbersweredefinedbytheexpansion
x
ex−1=∞summationdisplay
n=0Bn
n!xn. (7.75)
Replacing xwithz(analytic continuation), we have a Taylor series (compare Eq. (6.47))
with
Bn=n!
2πicontintegraldisplay
C0z
ez−1dz
zn+1, (7.76)
where the contour C0is around the origin counterclockwise with |z|<2πto avoid the
polesat 2 πin.
Forn=0weha v easimplepoleat z=0 witharesidueof +1.HencebyEq. (7.25),
B0=0!
2πi·2πi(1)=1. (7.77)
Forn=1thesingularityat z=0becomesasecond-orderpole.Theresiduemaybeshown
to be−1
2by series expansion of the exponential, followed by a binomial expansion. This
resultsin
B1=1!
2πi·2πiparenleftbigg
−1
2parenrightbigg
=−1
2. (7.78)
Forn≥2 this procedure becomes rather tedious, and we resort to a different means of
evaluatingEq. (7.76). Thecontouris deformed,as showninFig.7.10.
The new contour Cstill encircles the origin, as required, but now it also encircles
(in a negative direction) an infinite series of singular points along the imaginary axis at
z=±p2πi,p=1,2,3,.... The integration back and forth along the x-axis cancels out,
and forR→∞the integration over the infinite circle yields zero. Remember that n≥2.
Therefore
contintegraldisplay
C0z
ez−1dz
zn+1=−2πi∞summationdisplay
p=1residues (z=±p2πi). (7.79)
474 Chapter 7 Functions of a Complex Variable II
FIGURE 7.10Contourof
integrationforBernoullinumbers.
Atz=p2πiwe have a simple pole with a residue (p2πi)−n. Whennis odd, the residue
fromz=p2πiexactly cancels that from z=−p2πiandBn=0,n=3,5,7, and so on.
Forneventheresiduesadd,giving
Bn=n!
2πi(−2πi)2∞summationdisplay
p=11
pn(2πi)n
=−(−1)n/22n!
(2π)n∞summationdisplay
p=1p−n=−(−1)n/22n!
(2π)nζ(n) (n even),(7.80)
whereζ(n)is the Riemann zeta function introduced in Section 5.9. Equation (7.80) corre-
spondstoEq.(5.152) ofSection5.9. /squaresolid
Exercises
7.1.1 Determinethenatureofthesingularitiesofeachofthefollowingfunctionsandevaluate
theresidues (a >0).
(a)1
z2+a2.( b)1
(z2+a2)2.
(c)z2
(z2+a2)2.( d)sin1/z
z2+a2.
(e)ze+iz
z2+a2. (f)ze+iz
z2−a2.
(g)e+iz
z2−a2.( h)z−k
z+1,0<k<1.
Hint. For the point at infinity, use the transformation w=1/zfor|z|→0. For the
residue,transform f(z)dzintog(w)dw andlookatthebehaviorof g(w).
7.1 Calculus of Residues 475
7.1.2 Locatethesingularitiesandevaluatetheresidues ofeachof thefollowingfunctions.
(a)z−n(ez−1)−1,z/negationslash=0,
(b)z2ez
1+e2z.
(c) Find a closed-form expression (that is, not a sum) for the sum of the finite-plane
singularities.
(d) Usingtheresultinpart(c), whatistheresidueat |z|→∞?
Hint.SeeSection5.9forexpressionsinvolvingBernoullinumbers.NotethatEq.(5.144)
cannotbeusedtoinvestigatethesingularityat z→∞,sincethisseriesisonlyvalidfor
|z|<2π.
7.1.3 The statement that the integral halfway around a singular point is equal to one-half the
integral all the way around was limited to simple poles. Show, by a specific example,
that
integraldisplay
Semicirclef(z)dz=1
2contintegraldisplay
Circlef(z)dz
doesnotnecessarilyholdif theintegralencirclesapoleofhigherorder.
Hint.T ryf(z)=z−2.
7.1.4 A function f(z)is analytic along the real axis except for a third-order pole at z=x0.
TheLaurentexpansionabout z=x0hastheform
f(z)=a−3
(z−x0)3+a−1
z−x0+g(z),
withg(z)analyticat z=x0.ShowthattheCauchyprincipalvaluetechniqueisapplica-
ble,inthesensethat
(a) lim
δ→0braceleftbiggintegraldisplayx0−δ
−∞f(x)dx+integraldisplay∞
x0+δf(x)dxbracerightbigg
isfinite.
(b)integraldisplay
Cx0f(z)dz=±iπa−1,
whereCx0denotesa smallsemicircle aboutz=x0.
7.1.5 Theunitstepfunctionisdefinedas(compareExercise1.15.13)
u(s−a)=braceleftBigg0,s<a
1,s>a.
Showthat u(s)hastheintegralrepresentations
(a)u(s)=lim
ε→0+1
2πiintegraldisplay∞
−∞eixs
x−iεdx,
476 Chapter 7 Functions of a Complex Variable II
(b)u(s)=1
2+1
2πiPintegraldisplay∞
−∞eixs
xdx.
Note.Theparameter sis real.
7.1.6 Most of the special functions of mathematical physics may be generated (defined) by
ageneratingfunctionoftheform
g(t,x)=summationdisplay
nfn(x)tn.
Giventhefollowingintegralrepresentations,derivethecorrespondinggeneratingfunc-
tion:
(a) Bessel:
Jn(x)=1
2πicontintegraldisplay
e(x/2)(t−1/t)t−n−1dt.
(b) ModifiedBessel:
In(x)=1
2πicontintegraldisplay
e(x/2)(t+1/t)t−n−1dt.
(c) Legendre:
Pn(x)=1
2πicontintegraldisplayparenleftbig
1−2tx+t2parenrightbig−1/2t−n−1dt.
(d) Hermite:
Hn(x)=n!
2πicontintegraldisplay
e−t2+2txt−n−1dt.
(e) Laguerre:
Ln(x)=1
2πicontintegraldisplaye−xt/(1−t)
(1−t)tn+1dt.
(f) Chebyshev:
Tn(x)=1
4πicontintegraldisplay(1−t2)t−n−1
(1−2tx+t2)dt.
Eachofthecontoursencirclestheoriginandnoothersingularpoints.
7.1.7 GeneralizingExample7.1.2, showthat
integraldisplay2π
0dθ
a±bcosθ=integraldisplay2π
0dθ
a±bsinθ=2π
(a2−b2)1/2,fora>|b|.
Whathappensif |b|>|a|?
7.1.8 Showthatintegraldisplayπ
0dθ
(a+cosθ)2=πa
(a2−1)3/2,a>1.
7.1 Calculus of Residues 477
7.1.9 Showthat
integraldisplay2π
0dθ
1−2tcosθ+t2=2π
1−t2,for|t|<1.
Whathappensif |t|>1?Whathappensif |t|=1?
7.1.10 Withthecalculusofresiduesshowthat
integraldisplayπ
0cos2nθdθ=π(2n)!
22n(n!)2=π(2n−1)!!
(2n)!!,n=0,1,2,....
(ThedoublefactorialnotationisdefinedinSection8.1.)
Hint. cosθ=1
2(eiθ+e−iθ)=1
2(z+z−1),|z|=1.
7.1.11 Evaluate
integraldisplay∞
−∞cosbx−cosax
x2dx, a>b> 0.
ANS.π(a−b).
7.1.12 Provethat
integraldisplay∞
−∞sin2x
x2dx=π
2.
Hint.s i n2x=1
2(1−cos2x).
7.1.13 A quantum mechanical calculation of a transition probability leads to the function
f(t,ω)=2(1−cosωt)/ω2. Showthat
integraldisplay∞
−∞f(t,ω)dω=2πt.
7.1.14 Showthat (a >0)
(a)integraldisplay∞
−∞cosx
x2+a2dx=π
ae−a.
Howis therightsidemodifiedif cos xis replacedby cos kx?
(b)integraldisplay∞
−∞xsinx
x2+a2dx=πe−a.
Howis therightsidemodifiedif sin xisreplacedby sin kx?
These integrals may also be interpreted as Fourier cosine and sine transforms—
Chapter15.
7.1.15 Usethecontourshown(Fig.7.11)with R→∞toprovethat
integraldisplay∞
−∞sinx
xdx=π.
478 Chapter 7 Functions of a Complex Variable II
FIGURE 7.11Largesquare
contour.
7.1.16 Inthequantumtheoryofatomiccollisionsweencountertheintegral
I=integraldisplay∞
−∞sint
teiptdt,
inwhich pisreal. Showthat
I=0,|p|>1
I=π,|p|<1.
Whathappensif p=±1?
7.1.17 Evaluate
integraldisplay∞
0(lnx)2
1+x2dx
(a) byappropriateseries expansionoftheintegrandtoobtain
4∞summationdisplay
n=0(−1)n(2n+1)−3,
(b) andbycontourintegrationtoobtain
π3
8.
Hint.x→z=et. TrythecontourshowninFig.7.12,letting R→∞.
7.1.18 Showthat
integraldisplay∞
0xa
(x+1)2dx=πa
sinπa,
FIGURE 7.12Smallsquare
contour.
7.1 Calculus of Residues 479
FIGURE 7.13Contouravoiding
branchpointandpole.
where−1<a<1.Hereis stillanotherwayof derivingEq. (7.74).
Hint. Use the contour shown in Fig. 7.13, noting that z=0 is a branch point and the
positivex-axisisacutline.NotealsothecommentsonphasesfollowingExample6.6.1.
7.1.19 Showthat
integraldisplay∞
0x−a
x+1dx=π
sinaπ,
where 0<a<1. This opens up another way of deriving the factorial function relation
givenbyEq.(7.74).
Hint.Youhaveabranchpointandyouwillneedacutline.Recallthat z−a=winpolar
form is
bracketleftbig
rei(θ+2πn)bracketrightbig−a=ρeiϕ,
whichleadsto −aθ−2anπ=ϕ.Youmustrestrict ntozero(oranyothersingleinteger)
inorderthat ϕmaybeuniquelyspecified.TrythecontourshowninFig.7.14.
FIGURE 7.14Alternativecontour
avoidingbranchpoint.
480 Chapter 7 Functions of a Complex Variable II
FIGURE 7.15Anglecontour.
7.1.20 Showthat
integraldisplay∞
0dx
(x2+a2)2=π
4a3,a>0.
7.1.21 Evaluate
integraldisplay∞
−∞x2
1+x4dx.
ANS.π/√
2.
7.1.22 Showthat
integraldisplay∞
0cosparenleftbig
t2parenrightbig
dt=integraldisplay∞
0sinparenleftbig
t2parenrightbig
dt=√π
2√
2.
Hint.TrythecontourshowninFig.7.15.
Note. These are the Fresnel integrals for the special case of infinity as the upper limit.
For the general case of a varying upper limit, asymptotic expansions of the Fresnel
integralsarethetopicofExercise5.10.2.SphericalBesselexpansionsarethesubjectof
Exercise11.7.13.
7.1.23 SeveraloftheBromwichintegrals,Section15.12,involveaportionthatmaybeapprox-
imatedby
I(y)=integraldisplaya+iy
a−iyezt
z1/2dz.
Hereaandtare positiveandfinite.Showthat
limy→∞I(y)=0.
7.1 Calculus of Residues 481
FIGURE 7.16Sectorcontour.
7.1.24 Showthat
integraldisplay∞
01
1+xndx=π/n
sin(π/n).
Hint.TrythecontourshowninFig.7.16.
7.1.25 (a) Showthat
f(z)=z4−2cos2θz2+1
haszerosat eiθ,e−iθ,−eiθ, and−e−iθ.
(b) Showthat
integraldisplay∞
−∞dx
x4−2cos2θx2+1=π
2sinθ=π
21/2(1−cos2θ)1/2.
Exercise7.1.24 (n=4)isaspecialcaseofthisresult.
7.1.26 Showthat
integraldisplay∞
−∞x2dx
x4−2cos2θx2+1=π
2sinθ=π
21/2(1−cos2θ)1/2.
Exercise7.1.21isaspecialcaseofthisresult.
7.1.27 Applythetechniquesof Example7.1.5totheevaluationof theimproperintegral
I=integraldisplay∞
−∞dx
x2−σ2.
(a) Let σ→σ+iγ.
(b) Let σ→σ−iγ.
(c) TaketheCauchyprincipalvalue.
482 Chapter 7 Functions of a Complex Variable II
7.1.28 TheintegralinExercise7.1.17maybetransformedinto
integraldisplay∞
0e−yy2
1+e−2ydy=π3
16.
Evaluate this integral by the Gauss–Laguerre quadrature and compare your result with
π3/16.
ANS.Integral=1.93775(10points).
7.2 D ISPERSION RELATIONS
The concept of dispersion relations entered physics with the work of Kronig and Kramers
in optics. The name dispersion comes from optical dispersion, a result of the dependence
of the index of refraction on wavelength, or angular frequency. The index of refraction
nmay have a real part determined by the phase velocity and a (negative) imaginary part
determined by the absorption—see Eq. (7.94). Kronig and Kramers showed in 1926–
1927 that the real part of (n2−1)could be expressed as an integral of the imaginary part.
Generalizing this, we shall apply the label dispersion relations to any pair of equations
givingtherealpartofafunctionasanintegralofitsimaginarypartandtheimaginarypart
as an integral of its real part—Eqs. (7.86a) and (7.86b), which follow. The existence of
such integral relations might be suspected as an integral analog of the Cauchy–Riemann
differentialrelations,Section6.2.
The applications in modern physics are widespread. For instance, the real part of the
function might describe the forward scattering of a gamma ray in a nuclear Coulomb field
(a dispersive process). Then the imaginary part would describe the electron–positron pair
production in that same Coulomb field (the absorptive process). As will be seen later, the
dispersionrelationsmaybetakenasaconsequenceofcausalityandthereforeareindepen-
dentof thedetailsoftheparticularinteraction.
We consider a complex function f(z)that is analytic in the upper half-plane and on the
realaxis.Wealsorequirethat
lim
|z|→∞vextendsinglevextendsinglef(z)vextendsinglevextendsingle=0,0≤argz≤π, (7.81)
in order that the integral over an infinite semicircle will vanish. The point of these condi-
tionsisthatwemayexpress f(z)bytheCauchyintegralformula,Eq.(6.43),
f(z0)=1
2πicontintegraldisplayf(z)
z−z0dz. (7.82)
Theintegralovertheuppersemicircle6vanishesandwehave
f(z0)=1
2πiintegraldisplay∞
−∞f(x)
x−z0dx. (7.83)
TheintegraloverthecontourshowninFig.7.17hasbecomeanintegralalongthe x-axis.
Equation (7.83) assumes that z0is in the upper half-plane—interior to the closed con-
tour.Ifz0wereinthelowerhalf-plane,theintegralwouldyieldzerobytheCauchyintegral
6Theuse of asemicircleto closethe path of integration is convenient, not mandatory. Other paths are possible.
7.2 Dispersion Relations 483
FIGURE 7.17Semicirclecontour.
theorem, Section 6.3. Now, either letting z0approach the real axis from above (z0−x0)
or placing it on the real axis and taking an average of Eq. (7.83) and zero, we find that
Eq.(7.83) becomes
f(x0)=1
πiPintegraldisplay∞
−∞f(x)
x−x0dx, (7.84)
wherePindicatestheCauchyprincipalvalue.SplittingEq.(7.84)intorealandimaginary
parts7yields
f(x0)=u(x0)+iv(x0)
=1
πPintegraldisplay∞
−∞v(x)
x−x0dx−i
πPintegraldisplay∞
−∞u(x)
x−x0dx. (7.85)
Finally,equatingrealparttorealpartandimaginaryparttoimaginarypart,weobtain
u(x0)=1
πPintegraldisplay∞
−∞v(x)
x−x0dx(7.86a)
v(x0)=−1
πPintegraldisplay∞
−∞u(x)
x−x0dx.(7.86b)
These are the dispersion relations. The real part of our complex function is expressed as
an integral over the imaginary part. The imaginary part is expressed as an integral over
therealpart.Therealandimaginarypartsare Hilberttransforms ofeachother.Notethat
theserelationsaremeaningfulonlywhen f(x)isacomplexfunctionoftherealvariable x.
CompareExercise7.2.1.
Fromaphysicalpointofview u(x)and/orv(x)representsomephysicalmeasurements.
Thenf(z)=u(z)+iv(z)is an analytic continuation over the upper half-plane, with the
valueontherealaxisservingas aboundarycondition.
7Thesecondargument, y=0,is dropped: u(x0,0)→u(x0).
484 Chapter 7 Functions of a Complex Variable II
Symmetry Relations
On occasion f(x)will satisfy a symmetry relation and the integral from −∞to+∞
may be replaced by an integral over positive values only. This is of considerable physical
importance because the variable xmight represent a frequency and only zero and positive
frequenciesareavailableforphysicalmeasurements.Suppose8
f(−x)=f∗(x). (7.87)
Then
u(−x)+iv(−x)=u(x)−iv(x). (7.88)
The real part of f(x)is even and the imaginary part is odd.9In quantum mechanical
scattering problems these relations (Eq. (7.88)) are called crossing conditions. To exploit
thesecrossingconditions ,werewriteEq. (7.86a)as
u(x0)=1
πPintegraldisplay0
−∞v(x)
x−x0dx+1
πPintegraldisplay∞
0v(x)
x−x0dx. (7.89)
Lettingx→−xin the first integral on the right-hand side of Eq. (7.89) and substituting
v(−x)=−v(x)from Eq.(7.88), weobtain
u(x0)=1
πPintegraldisplay∞
0v(x)braceleftbigg1
x+x0+1
x−x0bracerightbigg
dx
=2
πPintegraldisplay∞
0xv(x)
x2−x2
0dx. (7.90)
Similarly,
v(x0)=−2
πPintegraldisplay∞
0x0u(x)
x2−x2
0dx. (7.91)
The original Kronig–Kramers optical dispersion relations were in this form. The asymp-
toticbehavior (x0→∞)ofEqs.(7.90)and(7.91)leadtoquantummechanical sumrules ,
Exercise7.2.4.
Optical Dispersion
Thefunction exp [i(kx−ωt)]describesanelectromagneticwavemovingalongthe x-axis
in the positive direction with velocity v=ω/k;ωis the angular frequency, kthe wave
number or propagation vector, and n=ck/ωthe index of refraction. From Maxwell’s
8This is not just a happy coincidence. It ensures that the Fourier transform of f(x)will be real. In turn, Eq. (7.87) is a conse-
quence of obtaining f(x)as the Fourier transform of areal function.
9u(x,0)=u(−x,0),v(x,0)=−v(−x,0).ComparethesesymmetryconditionswiththosethatfollowfromtheSchwarzreflec-
tion principle, Section6.5.
7.2 Dispersion Relations 485
equations, electric permittivity ε, and Ohm’s law with conductivity σ, the propagation
vectorkfor adielectricbecomes10
k2=εω2
c2parenleftbigg
1+i4πσ
ωεparenrightbigg
(7.92)
(withµ, the magnetic permeability, taken to be unity). The presence of the conductivity
(which means absorption) gives rise to an imaginary part. The propagation vector k(and
thereforetheindexof refraction n)havebecomecomplex.
Conversely, the (positive) imaginary part implies absorption. For poor conductivity
(4πσ/ωε≪1)abinomialexpansionyields
k=√εω
c+i2πσ
c√ε
and
ei(kx−ωt)=eiω(x√ε/c−t)e−2πσx/c√ε,
anattenuatedwave.
Returningtothegeneralexpressionfor k2,Eq.(7.92),wefindthattheindexofrefraction
becomes
n2=c2k2
ω2=ε+i4πσ
ω. (7.93)
We taken2to be a function of the complex variableω(withεandσdepending on ω).
However, n2does not vanish as ω→∞but instead approaches unity. So to satisfy the
condition, Eq. (7.81), one works with f(ω)=n2(ω)−1. The original Kronig–Kramers
opticaldispersionrelationswereintheform of
ℜbracketleftbig
n2(ω0)−1bracketrightbig
=2
πPintegraldisplay∞
0ωℑ[n2(ω)−1]
ω2−ω2
0dω,
(7.94)
ℑbracketleftbig
n2(ω0)−1bracketrightbig
=−2
πPintegraldisplay∞
0ω0ℜ[n2(ω)−1]
ω2−ω2
0dω.
Knowledge of the absorption coefficient at all frequencies specifies the real part of the
indexofrefraction,andviceversa.
The Parseval Relation
When the functions u(x)andv(x)are Hilbert transforms of each other (given by
Eqs. (7.86)) andeachis squareintegrable,11thetwofunctionsarerelatedby
integraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx. (7.95)
10S e eJ .D .J a c k s o n , Classical Electrodynamics , 3rd ed. New York: Wiley (1999), Sections 7.7 and 7.10. Equation (7.92) is in
Gaussianunits.
11This means thatintegraltext∞
−∞|u(x)|2dxandintegraltext∞
−∞|v(x)|2dxare finite.
486 Chapter 7 Functions of a Complex Variable II
Thisis theParsevalrelation.
ToderiveEq.(7.95), westartwith
integraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞1
πintegraldisplay∞
−∞v(s)ds
s−x1
πintegraldisplay∞
−∞v(t)dt
t−xdx,
usingEq. (7.86a)twice.Integratingfirstwithrespectto x,weha v e
integraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞v(s)dsintegraldisplay∞
−∞v(t)dt
π2integraldisplay∞
−∞dx
(s−x)(t−x).(7.96)
FromExercise7.2.8, the xintegrationyieldsadeltafunction:
1
π2integraldisplay∞
−∞dx
(s−x)(t−x)=δ(s−t).
We haveintegraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞v(t)dtintegraldisplay∞
−∞v(s)δ(s−t)ds. (7.97)
Then the sintegrationis carried out by inspection, using the defining property of the delta
function:integraldisplay∞
−∞v(s)δ(s−t)ds=v(t). (7.98)
SubstitutingEq.(7.98)intoEq.(7.97),wehaveEq.(7.95),theParsevalrelation.Again,in
terms of optics, the presence of refraction over some frequency range (n/negationslash=1)implies the
existenceof absorption,andviceversa.
Causality
Therealsignificanceofdispersionrelationsinphysicsisthattheyareadirectconsequence
of assuming that the particular physical system obeys causality. Causality is awkward to
defineprecisely,butthegeneralmeaningisthattheeffectcannotprecedethecause.Ascat-
teredwavecannotbeemittedbythescatteringcenterbeforetheincidentwavehasarrived.
For linear systems the most general relation between an input function G(the cause) and
anoutputfunction H(theeffect) maybewrittenas
H(t)=integraldisplay∞
−∞F(t−t′)G(t′)dt′. (7.99)
Causalityis imposedbyrequiringthat
F(t−t′)=0fort−t′<0.
Equation(7.99) givesthetimedependence.The frequencydependenceis obtainedbytak-
ingFouriertransforms. BytheFourierconvolutiontheorem,Section15.5,
h(ω)=f(ω)g(ω),
wheref(ω)is the Fourier transform of F(t), and so on. Conversely, F(t)is the Fourier
transform of f(ω).
7.2 Dispersion Relations 487
The connection with the dispersion relations is provided by the Titchmarsh theorem.12
This states that if f(ω)is square integrable over the real ω-axis, then any one of the fol-
lowingthreestatementsimpliestheothertwo.
1. TheFouriertransform of f(ω)is zerofor t<0:Eq. (7.99).
2. Replacing ωbyz, the function f(z)is analytic in the complex z-plane for y>0 and
approaches f(x)almosteverywhereas y→0.Further,
integraldisplay∞
−∞vextendsinglevextendsinglef(x+iy)vextendsinglevextendsingle2dx<K fory>0;
thatis, theintegralis bounded.
3. Therealandimaginarypartsof f(z)areHilberttransformsofeachother:Eqs.(7.86a)
and(7.86b).
The assumption that the relationship between the input and the output of our linear
system is causal (Eq. (7.99)) means that the first statement is satisfied. If f(ω)is square
integrable, then the Titchmarsh theorem has the third statement as a consequence and we
havedispersionrelations.
Exercises
7.2.1 The function f(z)satisfies the conditions for the dispersion relations. In addition,
f(z)=f∗(z∗), the Schwarz reflection principle, Section 6.5. Show that f(z)is identi-
callyzero.
7.2.2 Forf(z)such that we may replace the closed contour of the Cauchy integral formula
byanintegralovertherealaxiswehave
f(x0)=1
2πibraceleftbiggintegraldisplayx0−δ
−∞f(x)
x−x0dx+integraldisplay∞
x0+δf(x)
x−x0dxbracerightbigg
+1
2πiintegraldisplay
Cx0f(x)
x−x0dx.
HereCx0designates a small semicircle about x0in the lower half-plane. Show that this
reducesto
f(x0)=1
πiPintegraldisplay∞
−∞f(x)
x−x0dx,
whichisEq. (7.84).
7.2.3 (a) For f(z)=eiz,Eq.(7.81)doesnotholdattheendpoints,arg z=0,π.Show,with
thehelpofJordan’slemma,Section7.1, thatEq.(7.82) stillholds.
(b) For f(z)=eizverifythedispersionrelations,Eq.(7.89)orEqs.(7.90)and(7.91),
bydirectintegration.
7.2.4 Withf(x)=u(x)+iv(x)andf(x)=f∗(−x), showthatas x0→∞,
12RefertoE.C.Titchmarsh, IntroductiontotheTheoryofFourierIntegrals ,2nded.NewYork:OxfordUniversityPress(1937).
ForamoreinformaldiscussionoftheTitchmarshtheoremandfurtherdetailsoncausalityseeJ.Hilgevoord, DispersionRelations
and Causal Description . Amsterdam: North-Holland (1962).
488 Chapter 7 Functions of a Complex Variable II
(a)u(x0)∼−2
πx2
0integraldisplay∞
0xv(x)dx,
(b)v(x0)∼2
πx0integraldisplay∞
0u(x)dx.
Inquantummechanicsrelationsofthisform areoftencalled sumrules .
7.2.5 (a) Giventheintegralequation
1
1+x2
0=1
πPintegraldisplay∞
−∞u(x)
x−x0dx,
useHilberttransforms todetermine u(x0).
(b) Verifythattheintegralequationofpart(a) is satisfied.
(c) From f(z)|y=0=u(x)+iv(x),replacexbyzanddetermine f(z).Verifythatthe
conditionsfor theHilberttransforms aresatisfied.
(d) Arethecrossing conditionssatisfied?
ANS.(a) u(x0)=x0
1+x2
0,(c)f(z)=(z+i)−1.
7.2.6 (a) If therealpartof thecomplexindexofrefraction(squared) is constant(nooptical
dispersion),showthattheimaginarypartis zero(no absorption).
(b) Conversely, if there is absorption, show that there must be dispersion. In other
words,iftheimaginarypartof n2−1 isnotzero,showthattherealpartof n2−1
isnotconstant.
7.2.7 Givenu(x)=x/(x2+1)andv(x)=−1/(x2+1), show by direct evaluation of each
integralthat
integraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx.
ANS.integraldisplay∞
−∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx=π
2.
7.2.8 Takeu(x)=δ(x), a delta function, and assumethat the Hilbert transform equations
hold.
(a) Showthat
δ(w)=1
π2integraldisplay∞
−∞dy
y(y−w).
(b) Withchangesofvariables w=s−tandx=s−y,transformthe δrepresentation
ofpart(a) into
δ(s−t)=1
π2integraldisplay∞
−∞dx
(x−s)(s−t).
Note.TheδfunctionisdiscussedinSection1.15.
7.3 Method of Steepest Descents 489
7.2.9 Showthat
δ(x)=1
π2integraldisplay∞
−∞dt
t(t−x)
isa validrepresentationofthedeltafunctioninthesensethat
integraldisplay∞
−∞f(x)δ(x)dx=f(0).
Assumethat f(x)satisfiestheconditionfor theexistenceof aHilberttransform.
Hint.ApplyEq.(7.84) twice.
7.3 M ETHOD OF STEEPEST DESCENTS
Analytic Landscape
In analyzing problems in mathematical physics, one often finds it desirable to know the
behavior of a function for large values of the variable or some parameter s, that is, the
asymptotic behavior of the function. Specific examples are furnished by the gamma func-
tion(Chapter8)andvariousBesselfunctions(Chapter11).Alltheseanalyticfunctionsare
definedbyintegrals
I(s)=integraldisplay
CF(z,s)dz, (7.100)
whereFis analytic in zand depends on a real parameter s. We write F(z)whenever
possible.
Sofar wehaveevaluatedsuchdefiniteintegralsof analyticfunctionsalongtherealaxis
bydeformingthepath CtoC′inthecomplexplane,so |F|becomessmallforall zonC′.
This method succeeds as long as only isolated poles occur in the area between CandC′.
The poles are taken into account by applying the residue theorem of Section 7.1. The
residuesgiveameasureofthesimplepoles,where |F|→∞,whichusuallydominateand
determinethevalueof theintegral.
The behavior of the integral in Eq. (7.100) clearly depends on the absolute value |F|of
theintegrand.Moreover,thecontoursof |F|oftenbecomemorepronouncedas sbecomes
large.Letusfocusonaplotof |F(x+iy)|2=U2(x,y)+V2(x,y),ratherthantherealpart
ℜF=Uandtheimaginarypart ℑF=Vseparately.Suchaplotof |F|2overthecomplex
plane is called the analytic landscape , after Jensen, who, in 1912, proved that it has only
saddle points and troughs but no peaks . Moreover, the troughs reach down all the way
to the complex plane. In the absence of (simple) poles, saddle points are next in line to
dominate the integral in Eq. (7.100). Hence the name saddle point method . At a saddle
pointthereal(orimaginary)part UofFhasalocalmaximum,whichimpliesthat
∂U
∂x=∂U
∂y=0,
andthereforebytheuseoftheCauchy–Riemannconditionsof Section6.2,
∂V
∂x=∂V
∂y=0,
490 Chapter 7 Functions of a Complex Variable II
soVhas a minimum, or vice versa, and F′(z)=0. Jensen’s theorem prevents Uand
Vfrom having either a maximum or a minimum. See Fig. 7.18 for a typical shape (and
Exercises6.2.3and6.2.4). Ourstrategywillbetochoosethepath Csothatitrunsover
thesaddlepoint,whichgivesthedominantcontribution,andinthevalleyselsewhere.
If there are several saddle points, we treat each alike, and their contributions will add to
I(s→∞).
To prove that there are no peaks, assume there is one at z0. That is,|F(z0)|2>|F(z)|2
forallzofaneighborhood |z−z0|≤r.I f
F(z)=∞summationdisplay
n=0an(z−z0)n
is the Taylor expansion at z0, the mean value m(F)on the circle z=z0+rexp(iϕ)be-
comes
m(F)≡1
2πintegraldisplay2π
0vextendsinglevextendsingleFparenleftbig
z0+reiϕparenrightbigvextendsinglevextendsingle2dϕ
=1
2πintegraldisplay2π
0∞summationdisplay
m,n=0a∗
manrm+nei(n−m)ϕdϕ
=∞summationdisplay
n=0|an|2r2n≥|a0|2=vextendsinglevextendsingleF(z0)vextendsinglevextendsingle2, (7.101)
usingorthogonality,1
2πintegraltext2π
0expi(n−m)ϕdϕ=δnm.Sincem(F)isthemeanvalueof |F|2
onthecircleofradius r,theremustbeapoint z1onitsothat|F(z1)|2≥m(F)≥|F(z0)|2,
whichcontradictsourassumption.Hencetherecanbenosuchpeak.
Next, let us assume there is a minimum at z0so that 0<|F(z0)|2<|F(z)|2for allzof
aneighborhoodof z0.Inotherwords,thedipinthevalleydoesnotgodowntothecomplex
plane. Then|F(z)|2>0 and, since 1 /F(z)is analytic there, it has a Taylor expansion and
z0would be a peak of 1 /|F(z)|2, which is impossible. This proves Jensen’s theorem. We
nowturnour attentionbacktotheintegralinEq. (7.100).
Saddle Point Method
Since each saddle point z0necessarily lies above the complex plane, that is, |F(z0)|2>0,
we write Fin exponential form, ef(z,s), in its vicinity without loss of generality. Note
that having no zero in the complex plane is a characteristic property of the exponential
function. Moreover, any saddle point with F(z)=0 becomes a trough of |F(z)|2because
|F(z)|2≥0.A case in point is the function z2atz=0, where d(z2)/dz=2z=0.Here
z2=(x+iy)2=x2−y2+2ixy,and2xyhasasaddlepointat z=0,andsohas x2−y2,
but|z|4hasatroughthere.
Atz0thetangentialplaneishorizontal;thatis,∂F
∂z|z=z0=0,orequivalently∂f
∂z|z=z0=0.
This condition locates the saddle point. Our next goal is to determine the direction of
steepestdescent. Atz0,fhasapowerseries
f(z)=f(z0)+1
2f′′(z0)(z−z0)2+···, (7.102)
7.3 Method of Steepest Descents 491
FIGURE 7.18Asaddlepoint.
or
f(z)=f(z0)+1
2parenleftbig
f′′(z0)+εparenrightbig
(z−z0)2, (7.103)
upon collecting all higher powers in the (small) ε. Let us take f′′(z0)/negationslash=0 for simplicity.
Then
f′′(z0)(z−z0)2=−t2,treal, (7.104)
defines a line through z0(saddle point axisin Fig. 7.18). At z0,t=0. Along the axis
ℑf′′(z0)(z−z0)2is zero and v=ℑf(z)≈ℑf(z0)is constant if εin Eq. (7.103) is ne-
glected.Equation(7.104)canalsobeexpressedintermsof angles,
arg(z−z0)=π
2−1
2argf′′(z0)=constant. (7.105)
Since|F(z)|2=exp(2ℜf)varies monotonically with ℜf,|F(z)|2≈exp(−t2)falls off
exponentiallyfromitsmaximumat t=0alongthisaxis.Hencethename steepestdescent .
Thelinethrough z0definedby
f′′(z0)(z−z0)2=+t2(7.106)
isorthogonaltothisaxis( dashedinFig.7.18),whichis evidentfromits angle,
arg(z−z0)=−1
2argf′′(z0)=constant, (7.107)
whencomparedwithEq. (7.105).Here |F(z)|2growsexponentially.
492 Chapter 7 Functions of a Complex Variable II
The curvesℜf(z)=ℜf(z0)go through z0,s oℜ[(f′′(z0)+ε)(z−z0)2]=0, or
(f′′(z0)+ε)(z−z0)2=itfor realt. Expressingthisinanglesas
arg(z−z0)=π
4−1
2argparenleftbig
f′′(z0)+εparenrightbig
,t>0, (7.108a)
arg(z−z0)=−π
4−1
2argparenleftbig
f′′(z0)+εparenrightbig
,t<0, (7.108b)
and comparing with Eqs. (7.105) and (7.107) we note that these curves ( dot-dashed in
Fig. 7.18) divide the saddle point region into four sectors, two with ℜf(z)>ℜf(z0)
(hence|F(z)|>|F(z0)|), shown shaded in Fig. 7.18, and two with ℜf(z)<ℜf(z0)
(hence|F(z)|<|F(z0)|). They are at±π
4angles from the axis. Thus, the integration path
hastoavoidtheshadedareas,where |F|rises.Ifapathischosentorunuptheslopesabove
thesaddlepoint,thelargeimaginarypartof f(z)leadstorapidoscillationsof F(z)=ef(z)
andcancellingcontributionstotheintegral.
So far, our treatment has been general , except for f′′(z0)/negationslash=0, which can be relaxed.
Nowwearereadyto specializetheintegrand Ffurtherinordertotieupthepathselection
withtheasymptoticbehavioras s→∞.
We assume that sappears linearly in the exponent, that is, we replace exp f(z,s)→
exp(sf(z)). This dependence on sensures that the saddle point contribution at z0grows
withs→∞providingsteepslopes,asisthecaseinmostapplicationsinphysics.Inorder
to account for the region far away from the saddle point that is not influenced by s,w e
include another analytic function, g(z), which varies slowly near the saddle point and is
independentof s.
Altogether,then, ourintegralhasthemoreappropriateandspecificform
I(s)=integraldisplay
Cg(z)esf(z)dz. (7.109)
The path of steepest descent is the saddle point axis when we neglect the higher-order
terms,ε, in Eq. (7.103). With ε, the path of steepest descent is the curve close to the axis
within the unshaded sectors, where v=ℑf(z)is strictly constant, while ℑf(z)is only
approximately constant on the axis. We approximate I(s)by the integral along the piece
oftheaxisinsidethepatchinFig.7.18, where(comparewithEq. (7.104))
z=z0+xeiα,α=π
2−1
2argf′′(z0), a≤x≤b. (7.110)
Wefind
I(s)≈eiαintegraldisplayb
agparenleftbig
z0+xeiαparenrightbig
expbracketleftbig
sfparenleftbig
z0+xeiαparenrightbigbracketrightbig
dx, (7.111a)
and the omitted part is small and can be estimated because ℜ(f(z)−f(z0))has an upper
negative bound, −Rsay, that depends on the size of the saddle point patch in Fig. 7.18
(that is, the values of a,bin Eq. (7.110)) that we choose. In Eq. (7.111) we use the power
expansions
fparenleftbig
z0+xeiαparenrightbig
=f(z0)+1
2f′′(z0)e2iαx2+···,
(7.111b)
gparenleftbig
z0+xeiαparenrightbig
=g(z0)+g′(z0)eiαx+···,
7.3 Method of Steepest Descents 493
andrecallfromEq. (7.110)that
1
2f′′(z0)e2iα=−1
2vextendsinglevextendsinglef′′(z0)vextendsinglevextendsingle<0.
Wefindfortheleadingtermfor s→∞:
I(s)=g(z0)esf(z0)+iαintegraldisplayb
ae−1
2s|f′′(z0)|x2dx. (7.112)
Since the integrand in Eq. (7.112) is essentially zero when xdeparts appreciably from
the origin, we let b→∞anda→−∞. The small error involved is straightforward to
estimate.Notingthattheremainingintegralis justaGausserror integral,
integraldisplay∞
−∞e−1
2a2x2dx=1
aintegraldisplay∞
−∞e−1
2x2dx=√
2π
a,
wefinallyobtain
I(s)=√
2πg(z0)esf(z0)eiα
|sf′′(z0)|1/2, (7.113)
wherethephase αwas introducedinEqs. (7.110)and(7.105).
A note of warning: We assumed that the only significant contribution to the integral
came from the immediate vicinity of the saddle point(s) z=z0. This condition must be
checkedforeachnewproblem(Exercise7.3.5).
Example 7.3.1 ASYMPTOTIC FORM OF THE HANKEL FUNCTION H(1)
ν(s)
InSection11.4itisshownthattheHankelfunctions,whichsatisfyBessel’sequation,may
bedefinedby
H(1)
ν(s)=1
πiintegraldisplay∞eiπ
C1,0e(s/2)(z−1/z)dz
zν+1, (7.114)
H(2)
ν(s)=1
πiintegraldisplay0
C2,∞e−iπe(s/2)(z−1/z)dz
zν+1. (7.115)
The contour C1is the curve in the upper half-plane of Fig. 7.19. The contour C2is in the
lower half-plane. We apply the method of steepest descents to the first Hankel function,
H(1)
ν(s), whichis convenientlyintheformspecifiedbyEq.(7.109), with f(z)givenby
f(z)=1
2parenleftbigg
z−1
zparenrightbigg
. (7.116)
Bydifferentiating,weobtain
f′(z)=1
2+1
2z2. (7.117)
494 Chapter 7 Functions of a Complex Variable II
FIGURE 7.19Hankelfunctioncontours.
Settingf′(z)=0,weobtain
z=i,−i. (7.118)
Hence there are saddle points at z=+iandz=−i.A tz=i, f′′(i)=−i,or argf′′(i)=
−π/2,so the saddle point direction is given by Eq. (7.110) as α=π
2+π
4=3
4π.For the
integralfor H(1)
ν(s)wemustchoosethecontourthroughthepoint z=+isothatitstartsat
theorigin,movesouttangentiallytothepositiverealaxis,andthenmovesaroundthrough
the saddle point at z=+iin the direction given by the angle α=3π/4 and then on out to
minusinfinity,asymptoticwiththenegativerealaxis.Thepathofsteepestascent,whichwe
must avoid, has the phase −1
2argf′′(i)=π
4,according to Eq. (7.107), and is orthogonal
totheaxis, ourpathof steepestdescent.
DirectsubstitutionintoEq. (7.113)with α=3π/4 nowyields
H(1)
ν(s)=1
πi√
2πi−ν−1e(s/2)(i−1/i)e3πi/4
|(s/2)(−2/i3)|1/2
=radicalbigg
2
πse(iπ/2)(−ν−2)eisei(3π/4). (7.119)
Bycombiningterms, weobtain
H(1)
ν(s)≈radicalbigg
2
πsei(s−ν(π/2)−π/4)(7.120)
astheleadingtermoftheasymptoticexpansionoftheHankelfunction H(1)
ν(s).Additional
terms, if desired, may be picked up from the power series of fandgin Eq. (7.111b). The
otherHankelfunctioncanbetreatedsimilarlyusingthesaddlepointat z=−i. /squaresolid
Example 7.3.2 ASYMPTOTIC FORM OF THE FACTORIAL FUNCTION Ŵ(1+s)
In many physical problems, particularly in the field of statistical mechanics, it is desir-
able to have an accurate approximation of the gamma or factorial function of very large
7.3 Method of Steepest Descents 495
numbers. As developed in Section 8.1, the factorial function may be defined by the Euler
integral
Ŵ(1+s)=integraldisplay∞
0ρse−ρdρ=ss+1integraldisplay∞
0es(lnz−z)dz. (7.121)
Here we have made the substitution ρ=zsin order to convert the integral to the form
requiredbyEq.(7.109).Asbefore,weassumethat sisrealandpositive,fromwhichitfol-
lowsthattheintegrandvanishesatthelimits0and ∞.Bydifferentiatingthe z-dependence
appearingintheexponent,weobtain
df(z)
dz=d
dz(lnz−z)=1
z−1,f′′(z)=−1
z2, (7.122)
whichshowsthatthepoint z=1isasaddlepointandarg f′′(1)=arg(−1)=π.According
toEq.(7.109) welet
z−1=xeiα,α=π
2−1
2argf′′(1)=π
2−π
2=0, (7.123)
withxsmall, to describe the contour in the vicinity of the saddle point. From this we see
thatthedirectionofsteepestdescentisalongtherealaxis,aconclusionthatwecouldhave
reachedmoreorless intuitively.
DirectsubstitutionintoEq. (7.113)with α=0n o wg i v e s
Ŵ(1+s)≈√
2πss+1e−s
|s(−1−2)|1/2. (7.124)
Thusthefirsttermintheasymptoticexpansionofthefactorialfunctionis
Ŵ(1+s)≈√
2πssse−s. (7.125)
This result is the first term in Stirling’s expansion of the factorial function. The method of
steepest descent is probably the easiest way of obtaining this first term. If more terms in
theexpansionaredesired,thenthemethodof Section8.3 ispreferable. /squaresolid
In the foregoing example the calculation was carried out by assuming sto be real. This
assumption is not necessary. We may show (Exercise 7.3.6) that Eq. (7.125) also holds
whensis replaced by the complex variable w, provided only that the real part of wbe
requiredtobelargeandpositive.
Asymptotic limits of integral representations of functions are extremely important in
manyapproximationsandapplicationsinphysics :
integraldisplay
Cg(z)esf(z)dz∼√
2πg(z0)esf(z0)eiα
radicalbig
|sf′′(z0)|,f′(z0)=0.
The saddle point method is one method of choice for deriving them and belongs in the
toolkitofeveryphysicistandengineer.
496 Chapter 7 Functions of a Complex Variable II
Exercises
7.3.1 Usingthemethodofsteepestdescents,evaluatethesecondHankelfunction,givenby
H(2)
ν(s)=1
πiintegraldisplay0
−∞C2e(s/2)(z−1/z)dz
zν+1,
withcontour C2asshowninFig.7.19.
ANS.H(2)
ν(s)≈radicalbigg
2
πse−i(s−π/4−νπ/2).
7.3.2 Find the steepest path and leading asymptotic expansion for the Fresnel integralsintegraltexts
0cosx2dx,integraltexts
0sinx2dx.
Hint.Useintegraltext1
0eisz2dz.
7.3.3 (a) In applying the method of steepest descent to the Hankel function H(1)
ν(s),s h o w
that
ℜbracketleftbig
f(z)bracketrightbig
<ℜbracketleftbig
f(z0)bracketrightbig
=0
forzonthecontour C1butawayfromthepoint z=z0=i.
(b) Showthat
ℜbracketleftbig
f(z)bracketrightbig
>0for0<r<1,
π
2<θ≤π
−π≤θ<π
2
and
ℜbracketleftbig
f(z)bracketrightbig
<0forr>1,−π
2<θ<π
2
(Fig. 7.20). This is why C1may not be deformed to pass through the second saddle
point,z=−i. Comparewithandverifythedot-dashedlinesinFig.7.18forthis case.
FIGURE 7.20
7.3 Additional Readings 497
7.3.4 DeterminetheasymptoticdependenceofthemodifiedBessel functions Iν(x),gi v en
Iν(x)=1
2πiintegraldisplay
Ce(x/2)(t+1/t)dt
tν+1.
The contour starts and ends at t=−∞, encircling the origin in a positive sense. There
aretwosaddlepoints.Onlytheoneat z=+1contributessignificantlytotheasymptotic
form.
7.3.5 Determine the asymptotic dependence of the modified Bessel function of the second
kind,Kν(x),byus i ng
Kν(x)=1
2integraldisplay∞
0e(−x/2)(s+1/s)ds
s1−ν.
7.3.6 ShowthatStirling’sformula,
Ŵ(1+s)≈√
2πssse−s,
holdsforcomplexvaluesof s(withℜ(s)largeandpositive).
Hint.Thisinvolvesassigningaphaseto sandthendemandingthat ℑ[sf(z)]=constant
inthevicinityof thesaddlepoint.
7.3.7 AssumeH(1)
ν(s)tohaveanegativepower-seriesexpansionof theform
H(1)
ν(s)=radicalbigg
2
πsei(s−ν(π/2)−π/4)∞summationdisplay
n=0a−ns−n,
with the coefficient of the summation obtainedby the method of steepest descent. Sub-
stitute into Bessel’s equation and show that you reproduce the asymptotic series for
H(1)
ν(s)giveninSection11.6.
AdditionalReadings
N u s s e n z v e i g ,H .M . , Causality and Dispersion Relations , Mathematics in Science and Engineering Series,
Vol. 95. New York: Academic Press (1972). This is an advanced text covering causality and dispersion re-
lations in the first chapter and then moving on to develop the implications in a variety of areas of theoretical
physics.
Wyld, H. W., Mathematical Methods for Physics . Reading, MA: Benjamin/Cummings (1976), Perseus Books
(1999). This is arelatively advancedtext that contains anextensive discussion of the dispersion relations.
This page intentionally left blank
CHAPTER 8
THEGAMMA FUNCTION
(FACTORIAL FUNCTION )
The gamma function appears occasionally in physical problems such as the normalization
of Coulomb wave functions and the computation of probabilities in statistical mechanics.
In general, however, it has less direct physical application and interpretation than, say, the
LegendreandBesselfunctionsofChapters11and12.Rather,itsimportancestemsfromits
usefulnessindevelopingotherfunctionsthathavedirectphysicalapplication.Thegamma
function,therefore,is includedhere.
8.1 D EFINITIONS ,SIMPLE PROPERTIES
At least three different, convenient definitions of the gamma function are in common use.
Ourfirsttaskistostatethesedefinitions,todevelopsomesimple,directconsequences,and
toshowtheequivalenceof thethreeforms.
Infinite Limit (Euler)
Thefirstdefinition,namedafterEuler, is
Ŵ(z)≡limn→∞1·2·3···n
z(z+1)(z+2)···(z+n)nz,z/negationslash=0,−1,−2,−3,.... (8.1)
This definition of Ŵ(z)is useful in developing the Weierstrass infinite-product form of
Ŵ(z), Eq. (8.16), and in obtaining the derivative of ln Ŵ(z)(Section 8.2). Here and else-
499
500 Chapter 8 Gamma–Factorial Function
whereinthischapter zmaybeeitherrealorcomplex.Replacing zwithz+1,wehave
Ŵ(z+1)=limn→∞1·2·3···n
(z+1)(z+2)(z+3)···(z+n+1)nz+1
=limn→∞nz
z+n+1·1·2·3···n
z(z+1)(z+2)···(z+n)nz
=zŴ(z). (8.2)
This is the basic functional relation for the gamma function. It should be noted that it
is adifference equation. It has been shown that the gamma function is one of a general
class of functions that do not satisfy any differential equation with rational coefficients.
Specifically,thegammafunctionis oneof theveryfew functionsofmathematicalphysics
that does not satisfy either the hypergeometric differential equation (Section 13.4) or the
confluenthypergeometricequation(Section13.5).
Also,from thedefinition,
Ŵ(1)=limn→∞1·2·3···n
1·2·3···n(n+1)n=1. (8.3)
Now,applicationofEq. (8.2) gives
Ŵ(2)=1,
Ŵ(3)=2Ŵ(2)=2,... (8.4)
Ŵ(n)=1·2·3···(n−1)=(n−1)!.
Definite Integral (Euler)
Aseconddefinition,alsofrequentlycalledtheEulerintegral,is
Ŵ(z)≡integraldisplay∞
0e−ttz−1dt,ℜ(z)>0. (8.5)
The restriction on zis necessary to avoid divergence of the integral. When the gamma
function does appear in physical problems, it is often in this form or some variation, such
as
Ŵ(z)=2integraldisplay∞
0e−t2t2z−1dt,ℜ(z)>0. (8.6)
Ŵ(z)=integraldisplay1
0bracketleftbigg
lnparenleftbigg1
tparenrightbiggbracketrightbiggz−1
dt,ℜ(z)>0. (8.7)
Whenz=1
2, Eq. (8.6)is justtheGausserror integral,andwehavetheinterestingresult
Ŵparenleftbig1
2parenrightbig
=√π. (8.8)
GeneralizationsofEq.(8.6),theGaussianintegrals,areconsideredinExercise8.1.11.This
definiteintegralform of Ŵ(z), Eq. (8.5), leadstothebetafunction,Section8.4.
8.1 Definitions, Simple Properties 501
To show the equivalence of these two definitions, Eqs. (8.1) and (8.5), consider the
functionoftwovariables
F(z,n)=integraldisplayn
0parenleftbigg
1−t
nparenrightbiggn
tz−1dt,ℜ(z)>0, (8.9)
withnapositiveinteger.1Since
limn→∞parenleftbigg
1−t
nparenrightbiggn
≡e−t, (8.10)
fromthedefinitionof theexponential
limn→∞F(z,n)=F(z,∞)=integraldisplay∞
0e−ttz−1dt≡Ŵ(z) (8.11)
byEq.(8.5).
Returningto F(z,n),weevaluateitinsuccessiveintegrationsbyparts.Forconvenience
letu=t/n. Then
F(z,n)=nzintegraldisplay1
0(1−u)nuz−1du. (8.12)
Integratingbyparts, weobtain
F(z,n)
nz=(1−u)nuz
zvextendsinglevextendsinglevextendsinglevextendsingle1
0+n
zintegraldisplay1
0(1−u)n−1uzdu. (8.13)
Repeating this with the integrated part vanishing at both endpoints each time, we finally
get
F(z,n)=nzn(n−1)···1
z(z+1)···(z+n−1)integraldisplay1
0uz+n−1du
=1·2·3···n
z(z+1)(z+2)···(z+n)nz. (8.14)
Thisis identicalwiththeexpressionontherightsideof Eq.(8.1). Hence
limn→∞F(z,n)=F(z,∞)≡Ŵ(z), (8.15)
byEq.(8.1), completingtheproof.
Infinite Product (Weierstrass)
Thethirddefinition(Weierstrass’ form)is
1
Ŵ(z)≡zeγz∞productdisplay
n=1parenleftbigg
1+z
nparenrightbigg
e−z/n, (8.16)
1Theform of F(z,n)is suggested by the betafunction (compare Eq.(8.60)).
502 Chapter 8 Gamma–Factorial Function
whereγis theEuler–Mascheroniconstant,
γ=0.5772156619 .... (8.17)
This infinite-product form may be used to develop the reflection identity, Eq. (8.23), and
appliedintheexercises,suchasExercise8.1.17.Thisformcanbederivedfromtheoriginal
definition(Eq. (8.1)) byrewritingitas
Ŵ(z)=limn→∞1·2·3···n
z(z+1)···(z+n)nz=limn→∞1
znproductdisplay
m=1parenleftbigg
1+z
mparenrightbigg−1
nz.(8.18)
InvertingEq.(8.18) andusing
n−z=e(−lnn)z, (8.19)
weobtain
1
Ŵ(z)=zlimn→∞e(−lnn)znproductdisplay
m=1parenleftbigg
1+z
mparenrightbigg
. (8.20)
Multiplyinganddividingby
expbracketleftbiggparenleftbigg
1+1
2+1
3+···+1
nparenrightbigg
zbracketrightbigg
=nproductdisplay
m=1ez/m, (8.21)
weget
1
Ŵ(z)=zbraceleftbigg
limn→∞expbracketleftbiggparenleftbigg
1+1
2+1
3+···+1
n−lnnparenrightbigg
zbracketrightbiggbracerightbigg
×bracketleftbigg
limn→∞nproductdisplay
m=1parenleftbigg
1+z
mparenrightbigg
e−z/mbracketrightbigg
. (8.22)
AsshowninSection5.2,theparenthesisintheexponentapproachesalimit,namely γ,the
Euler–Mascheroniconstant.HenceEq. (8.16)follows.
It was shown in Section 5.11 that the Weierstrass infinite-product definition of Ŵ(z)led
directlytoanimportantidentity,
Ŵ(z)Ŵ(1−z)=π
sinzπ. (8.23)
Alternatively,wecanstart fromtheproductofEuler integrals,
Ŵ(z+1)Ŵ(1−z)=integraldisplay∞
0sze−sdsintegraldisplay∞
0t−ze−tdt
=integraldisplay∞
0vzdv
(v+1)2integraldisplay∞
0e−uudu=πz
sinπz,
transforming from the variables s,ttou=s+t,v=s/t, as suggested by combining the
exponentialsandthepowersintheintegrands.TheJacobianis
J=−vextendsinglevextendsinglevextendsinglevextendsingle11
1
t−s
t2vextendsinglevextendsinglevextendsinglevextendsingle=s+t
t2=(v+1)2
u,
8.1 Definitions, Simple Properties 503
where(v+1)t=u. The integralintegraltext∞
0e−uudu=1, while that over vmay be derived by
contourintegration,givingπz
sinπz.
This identity may also be derived by contour integration (Example 7.1.6 and Exer-
cises 7.1.18 and 7.1.19) and the beta function, Section 8.4. Setting z=1
2in Eq. (8.23),
weobtain
Ŵparenleftbig1
2parenrightbig
=√π (8.24a)
(takingthepositivesquareroot),inagreementwithEq. (8.8).
Similarlyonecanestablish Legendre’sduplicationformula ,
Ŵ(1+z)Ŵparenleftbig
z+1
2parenrightbig
=2−2z√πŴ(2z+1). (8.24b)
The Weierstrass definition shows immediately that Ŵ(z)has simple poles at z=
0,−1,−2,−3,...andthat[Ŵ(z)]−1hasnopolesinthefinitecomplexplane,whichmeans
thatŴ(z)hasnozeros.ThisbehaviormayalsobeseeninEq.(8.23),inwhichwenotethat
π/(sinπz)is neverequaltozero.
Actually the infinite-product definition of Ŵ(z)may be derived from the Weierstrass
factorization theorem with the specification that [Ŵ(z)]−1have simple zeros at z=
0,−1,−2,−3,....The Euler–Mascheroni constant is fixed by requiring Ŵ(1)=1. See
alsotheproductsexpansionsofentirefunctionsinSection7.1.
In probabilitytheorythegammadistribution(probabilitydensity)isgivenby
f(x)=
1
βαŴ(α)xα−1e−x/β,x>0
0,x ≤0.(8.24c)
The constant[βαŴ(α)]−1is chosen so that the total (integrated) probability will be unity.
Forx→E,kineticenergy, α→3
2,andβ→kT,Eq.(8.24c)yieldstheclassicalMaxwell–
Boltzmannstatistics.
Factorial Notation
So far this discussion has been presented in terms of the classical notation. As pointed out
by Jeffreys andothers, the −1o ft h ez−1 exponentin our seconddefinition(Eq. (8.5)) is
acontinualnuisance.Accordingly,Eq. (8.5) issometimesrewrittenas
integraldisplay∞
0e−ttzdt≡z!,ℜ(z)>−1, (8.25)
todefinea factorial function z!. Occasionally we may still encounter Gauss’ notation,producttext(z), for thefactorialfunction:
productdisplay
(z)=z!=Ŵ(z+1). (8.26)
TheŴnotation is due to Legendre. The factorial function of Eq. (8.25) is related to the
gammafunctionby
Ŵ(z)=(z−1)!orŴ(z+1)=z!. (8.27)
504 Chapter 8 Gamma–Factorial Function
FIGURE 8.1Thefactorial
function—extensiontonegative
arguments.
Ifz=n,apositiveinteger(Eq. (8.4)) showsthat
z!=n!=1·2·3···n, (8.28)
thefamiliarfactorial.However,itshouldbenotedthatsince z!isnowdefinedbyEq.(8.25)
(orequivalentlybyEq.(8.27))thefactorialfunctionisnolongerlimitedtopositiveintegral
valuesof theargument(Fig. 8.1). Thedifferencerelation(Eq. (8.2)) becomes
(z−1)!=z!
z. (8.29)
Thisshowsimmediatelythat
0!=1 (8.30)
and
n!=±∞ forn,anegative integer. (8.31)
Intermsof thefactorial, Eq.(8.23) becomes
z!(−z)!=πz
sinπz. (8.32)
Byrestrictingourselvestotherealvaluesoftheargument,wefindthat Ŵ(x+1)defines
thecurvesshowninFigs.8.1and8.2.The minimumofthecurveis
Ŵ(x+1)=x!=(0.46163...)!=0.88560.... (8.33a)
8.1 Definitions, Simple Properties 505
FIGURE 8.2Thefactorialfunctionandthefirsttwoderivativesof
ln(Ŵ(x+1)).
Double Factorial Notation
Inmanyproblemsofmathematicalphysics,particularlyinconnectionwithLegendrepoly-
nomials (Chapter 12), we encounter products of the odd positive integers and products of
the even positive integers. For convenience these are given special labels as double facto-
rials:
1·3·5···(2n+1)=(2n+1)!!
2·4·6···(2n)=(2n)!!.(8.33b)
Clearly,thesearerelatedtotheregularfactorialfunctionsby
(2n)!!=2nn!and(2n+1)!!=(2n+1)!
2nn!. (8.33c)
Wealsodefine (−1)!!=1,aspecialcasethatdoesnotfollowfromEq. (8.33c).
Integral Representation
An integral representation that is useful in developing asymptotic series for the Bessel
functionsisintegraldisplay
Ce−zzνdz=parenleftbig
e2πiν−1parenrightbig
Ŵ(ν+1), (8.34)
whereCis the contour shown in Fig. 8.3. This contour integral representation is only
useful when νis not an integer, z=0 then being a branch point . Equation (8.34) may be
506 Chapter 8 Gamma–Factorial Function
FIGURE 8.3Factorialfunctioncontour.
FIGURE 8.4Thecontourof Fig.8.3deformed.
readily verified for ν>−1 by deforming the contour as shown in Fig. 8.4. The integral
from∞into the origin yields −(ν!), placing the phase of zat 0. The integral out to ∞(in
the fourth quadrant) then yields e2πiνν!, the phase of zhaving increased to 2 π. Since the
circlearoundtheorigincontributesnothingwhen ν>−1,Eq. (8.34)follows.
It isoftenconvenienttocastthisresultintoa moresymmetricalform:
integraldisplay
Ce−z(−z)νdz=2iŴ(ν+1)sin(νπ). (8.35)
This analysis establishes Eqs. (8.34) and (8.35) for ν>−1. It is relatively simple to
extend the range to include all nonintegral ν. First, we note that the integral exists for
ν<−1 as long as we stay away from the origin. Second, integrating by parts we find
thatEq.(8.35)yieldsthefamiliardifferencerelation(Eq.(8.29)).Ifwetakethedifference
relation to define the factorial function of ν<−1, then Eqs. (8.34) and (8.35) are verified
forallν(exceptnegativeintegers).
Exercises
8.1.1 Derivetherecurrencerelations
Ŵ(z+1)=zŴ(z)
from theEulerintegral(Eq. (8.5)),
Ŵ(z)=integraldisplay∞
0e−ttz−1dt.
8.1 Definitions, Simple Properties 507
8.1.2 In a power-series solution for the Legendre functions of the second kind we encounter
theexpression
(n+1)(n+2)(n+3)···(n+2s−1)(n+2s)
2·4·6·8···(2s−2)(2s)·(2n+3)(2n+5)(2n+7)···(2n+2s+1),
inwhich sisa positiveinteger.Rewritethis expressionintermsoffactorials.
8.1.3 Showthat,as s−n→negativeinteger,
(s−n)!
(2s−2n)!→(−1)n−s(2n−2s)!
(n−s)!.
Heresandnare integers with s<n. This result can be used to avoid negative facto-
rials, such as in the series representations of the spherical Neumann functions and the
Legendrefunctionsof thesecondkind.
8.1.4 Showthat Ŵ(z)maybewritten
Ŵ(z)=2integraldisplay∞
0e−t2t2z−1dt,ℜ(z)>0,
Ŵ(z)=integraldisplay1
0bracketleftbigg
lnparenleftbigg1
tparenrightbiggbracketrightbiggz−1
dt,ℜ(z)>0.
8.1.5 In a Maxwellian distribution the fraction of particles with speed between vandv+dv
is
dN
N=4πparenleftbiggm
2πkTparenrightbigg3/2
expparenleftbigg
−mv2
2kTparenrightbigg
v2dv,
Nbeingthetotalnumberofparticles.Theaverageorexpectationvalueof vnisdefined
as/angbracketleftvn/angbracketright=N−1integraltext
vndN.Showthat
angbracketleftbig
vnangbracketrightbig
=parenleftbigg2kT
mparenrightbiggn/2Ŵparenleftbign+3
2parenrightbig
Ŵ(3/2).
8.1.6 Bytransformingtheintegralintoagammafunction,showthat
−integraldisplay1
0xklnxdx=1
(k+1)2,k>−1.
8.1.7 Showthat
integraldisplay∞
0e−x4dx=Ŵparenleftbigg5
4parenrightbigg
.
8.1.8 Showthat
lim
x→0(ax−1)!
(x−1)!=1
a.
8.1.9 Locatethepolesof Ŵ(z). Showthattheyaresimplepolesanddeterminetheresidues.
8.1.10 Showthattheequation x!=k,k/negationslash=0,hasaninfinitenumberof realroots.
8.1.11 Showthat
508 Chapter 8 Gamma–Factorial Function
(a)integraldisplay∞
0x2s+1expparenleftbig
−ax2parenrightbig
dx=s!
2as+1.
(b)integraldisplay∞
0x2sexpparenleftbig
−ax2parenrightbig
dx=(s−1
2)!
2as+1/2=(2s−1)!!
2s+1asradicalbiggπ
a.
TheseGaussianintegralsareofmajorimportanceinstatisticalmechanics.
8.1.12 (a) Developrecurrencerelationsfor (2n)!!andfor(2n+1)!!.
(b) Usetheserecurrencerelationstocalculate(or todefine) 0 !!and(−1)!!.
ANS. 0!!=1,(−1)!!=1.
8.1.13 Forsanonnegativeinteger,showthat
(−2s−1)!!=(−1)s
(2s−1)!!=(−1)s2ss!
(2s)!.
8.1.14 Express thecoefficientofthe nthtermoftheexpansionof (1+x)1/2
(a) intermsoffactorials ofintegers,
(b) intermsofthedoublefactorial( !!) functions.
ANS.an=(−1)n+1(2n−3)!
22n−2n!(n−2)!=(−1)n+1(2n−3)!!
(2n)!!,n=2,3,....
8.1.15 Express thecoefficientofthe nthtermoftheexpansionof (1+x)−1/2
(a) intermsofthefactorialsof integers,
(b) intermsofthedoublefactorial( !!) functions.
ANS.an=(−1)n(2n)!
22n(n!)2=(−1)n(2n−1)!!
(2n)!!,n=1,2,3,....
8.1.16 TheLegendrepolynomialmaybewrittenas
Pn(cosθ)=2(2n−1)!!
(2n)!!braceleftbigg
cosnθ+1
1·n
2n−1cos(n−2)θ
+1·3
1·2n(n−1)
(2n−1)(2n−3)cos(n−4)θ
+1·3·5
1·2·3n(n−1)(n−2)
(2n−1)(2n−3)(2n−5)cos(n−6)θ+···bracerightbigg
.
Letn=2s+1.Then
Pn(cosθ)=P2s+1(cosθ)=ssummationdisplay
m=0amcos(2m+1)θ.
Findamintermsoffactorials anddoublefactorials.
8.1 Definitions, Simple Properties 509
8.1.17 (a) Showthat
Ŵparenleftbig1
2−nparenrightbig
Ŵparenleftbig1
2+nparenrightbig
=(−1)nπ,
wherenis aninteger.
(b) Express Ŵ(1
2+n)andŴ(1
2−n)separatelyintermsof π1/2anda!!function.
ANS.Ŵ(1
2+n)=(2n−1)!!
2nπ1/2.
8.1.18 Fromoneof thedefinitionsofthefactorialor gammafunction,showthat
vextendsinglevextendsingle(ix)!vextendsinglevextendsingle2=πx
sinhπx.
8.1.19 Provethat
vextendsinglevextendsingleŴ(α+iβ)vextendsinglevextendsingle=vextendsinglevextendsingleŴ(α)vextendsinglevextendsingle∞productdisplay
n=0bracketleftbigg
1+β2
(α+n)2bracketrightbigg−1/2
.
Thisequationhasbeenusefulincalculationsofbetadecaytheory.
8.1.20 Showthat
vextendsinglevextendsingle(n+ib)!vextendsinglevextendsingle=parenleftbiggπb
sinhπbparenrightbigg1/2nproductdisplay
s=1parenleftbig
s2+b2parenrightbig1/2
forn,apositiveinteger.
8.1.21 Showthat
|x!|≥vextendsinglevextendsingle(x+iy)!vextendsinglevextendsingle
for allx.Thevariables xandyarereal.
8.1.22 Showthat
vextendsinglevextendsingleŴparenleftbig1
2+iyparenrightbigvextendsinglevextendsingle2=π
coshπy.
8.1.23 Theprobabilitydensityassociatedwiththenormaldistributionof statisticsisgivenby
f(x)=1
σ(2π)1/2expbracketleftbigg
−(x−µ)2
2σ2bracketrightbigg
,
with(−∞,∞)fortherangeof x.Showthat
(a) themeanvalueof x,/angbracketleftx/angbracketrightisequalto µ,
(b) thestandarddeviation (/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2)1/2is givenby σ.
8.1.24 Fromthegammadistribution
f(x)=
1
βαŴ(α)xα−1e−x/β,x>0,
0,x ≤0,
showthat
(a)/angbracketleftx/angbracketright(mean)=αβ,(b)σ2(variance)≡/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2=αβ2.
510 Chapter 8 Gamma–Factorial Function
8.1.25 The wave function of a particle scattered by a Coulomb potential is ψ(r,θ).A tt h e
originthewavefunctionbecomes
ψ(0)=e−πγ/2Ŵ(1+iγ),
whereγ=Z1Z2e2/¯hv.Showthat
vextendsinglevextendsingleψ(0)vextendsinglevextendsingle2=2πγ
e2πγ−1.
8.1.26 DerivethecontourintegralrepresentationofEq. (8.34),
2iν!sinνπ=integraldisplay
Ce−z(−z)νdz.
8.1.27 Writeafunctionsubprogram FACT(N)(fixed-pointindependentvariable)thatwillcal-
culateN!. Include provision for rejection and appropriate error message if Nis nega-
tive.
Note.For small integer N, direct multiplication is simplest. For large N, Eq. (8.55),
Stirling’sseries wouldbeappropriate.
8.1.28 (a) Write a function subprogram to calculate the double factorial ratio (2N−1)!!/
(2N)!!.Includeprovisionfor N=0andforrejectionandanerrormessageif Nis
negative.Calculateandtabulatethis ratiofor N=1(1)100.
(b) Check your function subprogram calculation of 199 !!/200!!against the value ob-
tainedfromStirling’sseries(Section8.3).
ANS.199!!
200!!=0.056348.
8.1.29 Using either the FORTRAN-supplied GAMMA or a library-supplied subroutine for
x!orŴ(x), determine the value of xfor which Ŵ(x)is a minimum (1≤x≤2)and
this minimum value of Ŵ(x). Notice that although the minimum value of Ŵ(x)may be
obtainedtoaboutsixsignificantfigures(singleprecision),thecorrespondingvalueof x
ismuchlessaccurate.Whythisrelativelylowaccuracy?
8.1.30 The factorial function expressed in integral form can be evaluated by the Gauss–
Laguerre quadrature. For a 10-point formula the resultant x!is theoretically exact for
xan integer, 0 up through 19. What happens if xis not an integer? Use the Gauss–
Laguerre quadrature to evaluate x!,x=0.0(0.1)2.0. Tabulate the absolute error as a
functionof x.
Checkvalue. x!exact−x!quadrature=0.00034 for x=1.3.
8.2 D IGAMMA AND POLYGAMMA FUNCTIONS
Digamma Functions
As may be noted from the three definitions in Section 8.1, it is inconvenient to deal with
the derivatives of the gamma or factorial function directly. Instead, it is customary to take
8.2 Digamma and Polygamma Functions 511
thenaturallogarithmofthefactorialfunction(Eq.(8.1)),converttheproducttoasum,and
thendifferentiate;thatis,
Ŵ(z+1)=zŴ(z)=limn→∞n!
(z+1)(z+2)···(z+n)nz(8.36)
and
lnŴ(z+1)=limn→∞bracketleftbig
ln(n!)+zlnn−ln(z+1)
−ln(z+2)−···−ln(z+n)bracketrightbig
, (8.37)
in which the logarithm of the limit is equal to the limit of the logarithm. Differentiating
withrespectto z, weobtain
d
dzlnŴ(z+1)≡ψ(z+1)=limn→∞parenleftbigg
lnn−1
z+1−1
z+2−···−1
z+nparenrightbigg
,(8.38)
which defines ψ(z+1), the digamma function. From the definition of the Euler–
Mascheroniconstant,2Eq. (8.38) mayberewrittenas
ψ(z+1)=−γ−∞summationdisplay
n=1parenleftbigg1
z+n−1
nparenrightbigg
=−γ+∞summationdisplay
n=1z
n(n+z). (8.39)
OneapplicationofEq.(8.39)isinthederivationoftheseriesformoftheNeumannfunction
(Section11.3).Clearly,
ψ(1)=−γ=−0.577215664901 ....3(8.40)
Another,perhapsmoreuseful, expressionfor ψ(z)isderivedinSection8.3.
Polygamma Function
The digamma function may be differentiated repeatedly, giving rise to the polygamma
function:
ψ(m)(z+1)≡dm+1
dzm+1ln(z!)
=(−1)m+1m!∞summationdisplay
n=11
(z+n)m+1,m=1,2,3,.... (8.41)
2Compare Sections 5.2 and 5.9. Weadd and substractsummationtextn
s=1s−1.
3γhas been computed to 1271 places by D. E. Knuth, Math. Comput. 16: 275 (1962), and to 3566 decimal places by
D.W. Sweeney, ibid.17: 170 (1963). It may be of interest thatthe fraction 228/395 gives γaccuratetosix places.
512 Chapter 8 Gamma–Factorial Function
Ap l o to f ψ(x+1)andψ′(x+1)is included in Fig. 8.2. Since the series in Eq. (8.41)
definestheRiemannzetafunction4(withz=0),
ζ(m)≡∞summationdisplay
n=11
nm, (8.42)
wehave
ψ(m)(1)=(−1)m+1m!ζ(m+1), m=1,2,3,.... (8.43)
The values of the polygamma functions of positive integral argument, ψ(m)(n+1),m a y
becalculatedbyusingExercise8.2.6.
In termsoftheperhapsmorecommon Ŵnotation,
dn+1
dzn+1lnŴ(z)=dn
dznψ(z)=ψ(n)(z). (8.44a)
Maclaurin Expansion, Computation
Itis nowpossibletowriteaMaclaurinexpansionfor ln Ŵ(z+1):
lnŴ(z+1)=∞summationdisplay
n=1zn
n!ψ(n−1)(1)=−γz+∞summationdisplay
n=2(−1)nzn
nζ(n) (8.44b)
convergent for |z|<1; forz=x, the range is−1<x≤1. Alternate forms of this series
appearinExercise5.9.14.Equation(8.44b)isapossiblemeansofcomputing Ŵ(z+1)for
real or complex z, but Stirling’s series (Section 8.3) is usually better, and in addition, an
excellent table of values of the gamma function for complex arguments based on the use
ofStirling’sseriesandtherecurrencerelation(Eq. (8.29)) is nowavailable.5
Series Summation
Thedigammaandpolygammafunctionsmayalsobeusedinsummingseries.Ifthegeneral
termoftheserieshastheformofarationalfraction(withthehighestpoweroftheindexin
the numerator at least two less than the highest power of the index in the denominator), it
maybetransformedbythemethodofpartialfractions(compareSection15.8).Theinfinite
series may then be expressed as a finite sum of digamma and polygamma functions. The
usefulnessofthismethoddependsontheavailabilityoftablesofdigammaandpolygamma
functions. Such tables and examples of series summation are given in AMS-55, Chapter 6
(seeAdditionalReadingsforthereference).
4SeeSection5.9. For z/negationslash=0 this series maybe usedtodefine ageneralizedzetafunction.
5TableoftheGammaFunctionforComplexArguments ,AppliedMathematicsSeriesNo.34.Washington,DC:NationalBureau
of Standards (1954).
8.2 Digamma and Polygamma Functions 513
Example 8.2.1 CATALAN ’SCONSTANT
Catalan’sconstant,Exercise5.2.22, or β(2)ofSection5.9 isgivenby
K=β(2)=∞summationdisplay
k=0(−1)k
(2k+1)2. (8.44c)
Groupingthepositiveandnegativetermsseparatelyandstartingwithunitindex(tomatch
theform of ψ(1), Eq. (8.41)), weobtain
K=1+∞summationdisplay
n=11
(4n+1)2−1
9−∞summationdisplay
n=11
(4n+3)2.
Now,quotingEq. (8.41), weget
K=8
9+1
16ψ(1)parenleftbig
1+1
4parenrightbig
−1
16ψ(1)parenleftbig
1+3
4parenrightbig
. (8.44d)
Using the values of ψ(1)from Table 6.1 of AMS-55 (see Additional Readings for the
reference),weobtain
K=0.91596559 ....
Compare this calculation of Catalan’s constant with the calculations of Chapter 5, either
directsummationoramodificationusingRiemannzetafunctionvalues. /squaresolid
Exercises
8.2.1 Verifythatthefollowingtwoformsof thedigammafunction,
ψ(x+1)=xsummationdisplay
r=11
r−γ
and
ψ(x+1)=∞summationdisplay
r=1x
r(r+x)−γ,
areequaltoeachother(for xapositiveinteger).
8.2.2 Showthat ψ(z+1)hastheseries expansion
ψ(z+1)=−γ+∞summationdisplay
n=2(−1)nζ(n)zn−1.
8.2.3 Forapower-seriesexpansionofln (z!),AMS-55(seeAdditionalReadingsforreference)
lists
ln(z!)=−ln(1+z)+z(1−γ)+∞summationdisplay
n=2(−1)n[ζ(n)−1]zn
n.
514 Chapter 8 Gamma–Factorial Function
(a) ShowthatthisagreeswithEq. (8.44b)for |z|<1.
(b) Whatistherangeofconvergenceofthisnewexpression?
8.2.4 Showthat
1
2lnparenleftbiggπz
sinπzparenrightbigg
=∞summationdisplay
n=1ζ(2n)
2nz2n,|z|<1.
Hint.TryEq. (8.32).
8.2.5 Write out a Weierstrass infinite-product definition of ln (z!). Without differentiating,
showthatthis leadsdirectlytotheMaclaurinexpansionof ln (z!), Eq.(8.44b).
8.2.6 Derivethedifferencerelationfor thepolygammafunction
ψ(m)(z+2)=ψ(m)(z+1)+(−1)mm!
(z+1)m+1,m=0,1,2,....
8.2.7 Showthatif
Ŵ(x+iy)=u+iv,
then
Ŵ(x−iy)=u−iv.
Thisis aspecialcaseof theSchwarzreflectionprinciple,Section6.5.
8.2.8 ThePochhammersymbol (a)nis definedas
(a)n=a(a+1)···(a+n−1), (a) 0=1
(for integral n).
(a) Express (a)nintermsof factorials.
(b) Find (d/da)(a) nintermsof (a)nanddigammafunctions.
ANS.d
da(a)n=(a)nbracketleftbig
ψ(a+n)−ψ(a)bracketrightbig
.
(c) Showthat
(a)n+k=(a+n)k·(a)n.
8.2.9 Verifythefollowingspecialvaluesof the ψformof thedi-andpolygammafunctions:
ψ(1)=−γ, ψ(1)(1)=ζ(2), ψ(2)(1)=−2ζ(3).
8.2.10 Derivethepolygammafunctionrecurrencerelation
ψ(m)(1+z)=ψ(m)(z)+(−1)mm!/zm+1,m=0,1,2,....
8.2.11 Verify
(a)integraldisplay∞
0e−rlnrdr=−γ.
8.2 Digamma and Polygamma Functions 515
(b)integraldisplay∞
0re−rlnrdr=1−γ.
(c)integraldisplay∞
0rne−rlnrdr=(n−1)!+nintegraldisplay∞
0rn−1e−rlnrdr , n=1,2,3,....
Hint.These may be verified by integration by parts, three parts, or differentiating the
integralformof n!withrespectto n.
8.2.12 Diracrelativisticwavefunctionsforhydrogeninvolvefactorssuchas [2(1−α2Z2)1/2]!
whereα, the fine structure constant, is1
137andZis the atomic number. Expand
[2(1−α2Z2)1/2]!ina seriesof powersof α2Z2.
8.2.13 The quantummechanicaldescriptionof a particle in a Coulombfieldrequires a knowl-
edge of the phase of the complex factorial function. Determine the phase of (1+ib)!
for small b.
8.2.14 Thetotalenergyradiatedbyablackbodyis givenby
u=8πk4T4
c3h3integraldisplay∞
0x3
ex−1dx.
Showthattheintegralinthis expressionisequalto 3 !ζ(4).
[ζ(4)=π4/90=1.0823...]Thefinalresultis theStefan–Boltzmannlaw.
8.2.15 AsageneralizationoftheresultinExercise8.2.14,showthat
integraldisplay∞
0xsdx
ex−1=s!ζ(s+1),ℜ(s)>0.
8.2.16 The neutrino energy density (Fermi distribution) in the early history of the universe is
givenby
ρν=4π
h3integraldisplay∞
0x3
exp(x/kT)+1dx.
Showthat
ρν=7π5
30h3(kT)4.
8.2.17 Provethat
integraldisplay∞
0xsdx
ex+1=s!parenleftbig
1−2−sparenrightbig
ζ(s+1),ℜ(s)>0.
Exercises8.2.15and8.2.17actuallyconstituteMellinintegraltransforms(compareSec-
tion15.1).
8.2.18 Provethat
ψ(n)(z)=(−1)n+1integraldisplay∞
0tne−zt
1−e−tdt,ℜ(z)>0.
516 Chapter 8 Gamma–Factorial Function
8.2.19 Usingdi-andpolygammafunctions,sumtheseries
(a)∞summationdisplay
n=11
n(n+1),(b)∞summationdisplay
n=21
n2−1.
Note.YoucanuseExercise8.2.6tocalculatetheneededdigammafunctions.
8.2.20 Showthat
∞summationdisplay
n=11
(n+a)(n+b)=1
(b−a)braceleftbig
ψ(1+b)−ψ(1+a)bracerightbig
,
wherea/negationslash=band neither anorbis a negative integer. It is of some interest to compare
thissummationwiththecorrespondingintegral,
integraldisplay∞
1dx
(x+a)(x+b)=1
b−abraceleftbig
ln(1+b)−ln(1+a)bracerightbig
.
Therelationbetween ψ(x)and lnxismadeexplicitinEq. (8.51)inthenextsection.
8.2.21 Verifythecontourintegralrepresentationof ζ(s),
ζ(s)=−(−s)!
2πiintegraldisplay
C(−z)s−1
ez−1dz.
Thecontour CisthesameasthatforEq.(8.35).Thepoints z=±2nπi, n=1,2,3,...,
areallexcluded.
8.2.22 Show that ζ(s)is analytic in the entire finite complex plane except at s=1, where it
hasasimplepolewitharesidueof +1.
Hint.Thecontourintegralrepresentationwillbeuseful.
8.2.23 Using the complex variable capability of FORTRAN calculate ℜ(1+ib)!,ℑ(1+ib)!,
|(1+ib)!|andphase (1+ib)!forb=0.0(0.1)1.0.Plotthephaseof (1+ib)!versusb.
Hint.Exercise8.2.3 offers aconvenientapproach.Youwillneedtocalculate ζ(n).
8.3 S TIRLING ’SSERIES
For computation of ln (z!)for very large z(statistical mechanics) and for numerical com-
putations at nonintegral values of z, a series expansion of ln (z!)in negative powers of zis
desirable.Perhapsthemostelegantwayofderivingsuchanexpansionisbythemethodof
steepest descents (Section 7.3). The following method, starting with a numerical integra-
tionformula,doesnotrequireknowledgeofcontourintegrationandis particularlydirect.
8.3 Stirling’s Series 517
Derivation from Euler–Maclaurin Integration Formula
TheEuler–Maclaurinformulafor evaluatingadefiniteintegral6is
integraldisplayn
0f(x)dx=1
2f(0)+f(1)+f(2)+···+1
2f(n)
−b2bracketleftbig
f′(n)−f′(0)bracketrightbig
−b4bracketleftbig
f′′′(n)−f′′′(0)bracketrightbig
−···,(8.45)
inwhichthe b2narerelatedtotheBernoullinumbers B2n(compareSection5.9) by
(2n)!b2n=B2n, (8.46)
B0=1,B 6=1
42,
B2=1
6,B 8=−1
30,
B4=−1
30,B 10=5
66,andsoon .(8.47)
ByapplyingEq.(8.45) tothedefiniteintegral
integraldisplay∞
0dx
(z+x)2=1
z(8.48)
(forznotonthenegativerealaxis),weobtain
1
z=1
2z2+ψ(1)(z+1)−2!b2
z3−4!b4
z5−···. (8.49)
ThisisthereasonforusingEq.(8.48).TheEuler–Maclaurinevaluationyields ψ(1)(z+1),
whichisd2lnŴ(z+1)/dz2.
UsingEq.(8.46) andsolvingfor ψ(1)(z+1),weha v e
ψ(1)(z+1)=d
dzψ(z+1)=1
z−1
2z2+B2
z3+B4
z5+···
=1
z−1
2z2+∞summationdisplay
n=1B2n
z2n+1. (8.50)
Since the Bernoulli numbers diverge strongly, this series does not converge. It is a semi-
convergent, or asymptotic, series, useful if one retains a small enough number of terms
(compareSection5.10).
Integratingonce,wegetthedigammafunction
ψ(z+1)=C1+lnz+1
2z−B2
2z2−B4
4z4−···
=C1+lnz+1
2z−∞summationdisplay
n=1B2n
2nz2n. (8.51)
IntegratingEq.(8.51)withrespectto zfromz−1tozandthenletting zapproachinfinity,
C1,theconstantofintegration,maybeshowntovanish.Thisgivesusasecondexpression
forthedigammafunction,oftenmoreusefulthanEq. (8.38) or(8.44b).
6This is obtainedby repeatedintegration byparts, Section5.9.
518 Chapter 8 Gamma–Factorial Function
Stirling’s Series
Theindefiniteintegralofthedigammafunction(Eq. (8.51)) is
lnŴ(z+1)=C2+parenleftbigg
z+1
2parenrightbigg
lnz−z+B2
2z+···+B2n
2n(2n−1)z2n−1+···,(8.52)
in which C2is another constant of integration. To fix C2we find it convenient to use the
doubling,orLegendreduplication,formuladerivedinSection8.4,
Ŵ(z+1)Ŵparenleftbig
z+1
2parenrightbig
=2−2zπ1/2Ŵ(2z+1). (8.53)
Thismaybeproveddirectlywhen zisapositiveintegerbywriting Ŵ(2z+1)asaproduct
of even terms times a product of odd terms and extracting a factor of 2 from each term
(Exercise 8.3.5). Substituting Eq. (8.52) into the logarithm of the doubling formula, we
findthatC2is
C2=1
2ln2π, (8.54)
giving
lnŴ(z+1)=1
2ln2π+parenleftbigg
z+1
2parenrightbigg
lnz−z+1
12z−1
360z3+1
1260z5−···.(8.55)
This is Stirling’s series, an asymptotic expansion. The absolute value of the error is less
thantheabsolutevalueof thefirst termomitted.
The constants of integration C1andC2may also be evaluated by comparison with the
first term of the series expansion obtained by the method of “steepest descent.” This is
carriedoutinSection7.3.
To help convey a feeling of the remarkable precision of Stirling’s series for Ŵ(s+1),
the ratio of the first term of Stirling’s approximation to Ŵ(s+1)is plotted in Fig. 8.5.
A tabulation gives the ratio of the first term in the expansion to Ŵ(s+1)and the ratio of
the first two terms in the expansion to Ŵ(s+1)(Table 8.1). The derivation of these forms
isExercise8.3.1.
Exercises
8.3.1 RewriteStirling’sseries togive Ŵ(z+1)insteadof ln Ŵ(z+1).
ANS.Ŵ(z+1)=√
2πzz+1/2e−zparenleftbigg
1+1
12z+1
288z2−139
51,840z3+···parenrightbigg
.
8.3.2 Use Stirling’s formula to estimate 52 !, the number of possible rearrangements of cards
inastandarddeckofplayingcards.
8.3.3 ByintegratingEq.(8.51)from z−1t ozandthenletting z→∞,evaluatetheconstant
C1intheasymptoticseries forthedigammafunction ψ(z).
8.3.4 Showthattheconstant C2inStirling’sformulaequals1
2ln2πbyusingthelogarithmof
thedoublingformula.
8.3 Stirling’s Series 519
FIGURE 8.5Accuracyof Stirling’sformula.
Table 8.1
s1
Ŵ(s+1)√
2πss+1/2e−s1
Ŵ(s+1)√
2πss+1/2e−sparenleftbigg
1+1
12sparenrightbigg
1 0.92213 0.99898
2 0.95950 0.99949
3 0.97270 0.99972
4 0.97942 0.99983
5 0.98349 0.99988
6 0.98621 0.99992
7 0.98817 0.99994
8 0.98964 0.99995
9 0.99078 0.99996
10 0.99170 0.99998
8.3.5 Bydirectexpansion,verifythedoublingformulafor z=n+1
2;nis aninteger.
8.3.6 WithoutusingStirling’sseriesshowthat
(a) ln(n!)<integraldisplayn+1
1lnxdx,(b)ln(n!)>integraldisplayn
1lnxdx;nis aninteger≥2.
Notice that the arithmetic mean of these two integrals gives a good approximation for
Stirling’sseries.
8.3.7 Test forconvergence
∞summationdisplay
p=0bracketleftbigg(p−1
2)!
p!bracketrightbigg2
×2p+1
2p+2=π∞summationdisplay
p=0(2p−1)!!(2p+1)!!
(2p)!!(2p+2)!!.
520 Chapter 8 Gamma–Factorial Function
This series arises in an attempt to describe the magnetic field created by and enclosed
byacurrentloop.
8.3.8 Showthat
limx→∞xb−a(x+a)!
(x+b)!=1.
8.3.9 Showthat
limn→∞(2n−1)!!
(2n)!!n1/2=π−1/2.
8.3.10 Calculatethebinomialcoefficientparenleftbig2n
nparenrightbig
tosixsignificantfiguresfor n=10,20,and30.
Checkyourvaluesby
(a) aStirlingseriesapproximationthroughtermsin n−1,
(b) adoubleprecisioncalculation.
ANS.parenleftbig20
10parenrightbig
=1.84756×105,parenleftbig40
20parenrightbig
=1.37846×1011,
parenleftbig60
30parenrightbig
=1.18264×1017.
8.3.11 Write a program (or subprogram) that will calculate log10(x!)directly from Stirling’s
series. Assume that x≥10. (Smaller values could be calculated via the factorial re-
currence relation.) Tabulate log10(x!)versusxforx=10(10)300. Check your results
againstAMS-55(seeAdditionalReadingsforthisreference)orbydirectmultiplication
(forn=10,20,and30).
Checkvalue .l o g10(100!)=157.97.
8.3.12 UsingthecomplexarithmeticcapabilityofFORTRAN,writeasubroutinethatwillcal-
culate ln(z!)for complex zbased on Stirling’s series. Include a test and an appropriate
error message if zis too close to a negative real integer. Check your subroutine against
alternatecalculationsfor zreal,zpureimaginary,and z=1+ib(Exercise8.2.23).
Checkvalues .|(i0.5)!|=0.82618
phase(i0.5)!=−0.24406.
8.4 T HEBETA FUNCTION
Using the integral definition (Eq. (8.25)), we write the product of two factorials as the
product of two integrals. To facilitate a change in variables, we take the integrals over a
finiterange:
m!n!=lim
a2→∞integraldisplaya2
0e−uumduintegraldisplaya2
0e−vvndv,ℜ(m)>−1,
ℜ(n)>−1.(8.56a)
Replacing uwithx2andvwithy2, weobtain
m!n!=lima→∞4integraldisplaya
0e−x2x2m+1dxintegraldisplaya
0e−y2y2n+1dy. (8.56b)
8.4 The Beta Function 521
FIGURE 8.6Transformationfrom
Cartesiantopolarcoordinates.
Transformingtopolarcoordinatesgivesus
m!n!=lima→∞4integraldisplaya
0e−r2r2m+2n+3drintegraldisplayπ/2
0cos2m+1θsin2n+1θdθ
=(m+n+1)!2integraldisplayπ/2
0cos2m+1θsin2n+1θdθ. (8.57)
Here the Cartesian area element dxdyhas been replaced by rdrdθ(Fig. 8.6). The last
equalityinEq. (8.57) followsfromExercise8.1.11.
Thedefiniteintegral,togetherwiththefactor2, hasbeennamedthebetafunction:
B(m+1,n+1)≡2integraldisplayπ/2
0cos2m+1θsin2n+1θdθ
=m!n!
(m+n+1)!. (8.58a)
Equivalently,intermsofthegammafunctionandnotingitssymmetry,
B(p,q)=Ŵ(p)Ŵ(q)
Ŵ(p+q),B(q,p)=B(p,q). (8.58b)
Theonlyreasonforchoosing m+1andn+1,ratherthan mandn,astheargumentsof B
istobeinagreementwiththeconventional,historicalbetafunction.
Definite Integrals, Alternate Forms
The beta function is useful in the evaluation of a wide variety of definite integrals. The
substitution t=cos2θconvertsEq.(8.58a) to7
B(m+1,n+1)=m!n!
(m+n+1)!=integraldisplay1
0tm(1−t)ndt. (8.59a)
7TheLaplacetransform convolution theorem provides analternate derivation ofEq. (8.58a), compare Exercise 15.11.2.
522 Chapter 8 Gamma–Factorial Function
Replacing tbyx2, weobtain
m!n!
2(m+n+1)!=integraldisplay1
0x2m+1parenleftbig
1−x2parenrightbigndx. (8.59b)
Thesubstitution t=u/(1+u)inEq. (8.59a)yieldsstillanotherusefulform,
m!n!
(m+n+1)!=integraldisplay∞
0um
(1+u)m+n+2du. (8.60)
The beta function as a definite integral is useful in establishing integral representations of
theBesselfunction(Exercise11.1.18)andthehypergeometricfunction(Exercise13.4.10).
Verification of πα/sinπαRelation
If wetake m=a,n=−a,−1<a<1,then
integraldisplay∞
0ua
(1+u)2du=a!(−a)!. (8.61)
By contour integration this integral may be shown to be equal to πa/sinπa(Exer-
cise7.1.18),thusprovidinganothermethodofobtainingEq. (8.32).
Derivation of Legendre Duplication Formula
The form of Eq. (8.58a) suggests that the beta function may be useful in deriving the
doubling formula used in the preceding section. From Eq. (8.59a) with m=n=zand
ℜ(z)>−1,
z!z!
(2z+1)!=integraldisplay1
0tz(1−t)zdt. (8.62)
Bysubstituting t=(1+s)/2,wehave
z!z!
(2z+1)!=2−2z−1integraldisplay1
−1parenleftbig
1−s2parenrightbigzds=2−2zintegraldisplay1
0parenleftbig
1−s2parenrightbigzds. (8.63)
The last equality holds because the integrand is even. Evaluating this integral as a beta
function(Eq. (8.59b)), weobtain
z!z!
(2z+1)!=2−2z−1z!(−1
2)!
(z+1
2)!. (8.64)
Rearranging terms and recalling that (−1
2)!=π1/2, we reduce this equation to one form
oftheLegendreduplicationformula,
z!parenleftbig
z+1
2parenrightbig
!=2−2z−1π1/2(2z+1)!. (8.65a)
Dividingby (z+1
2), weobtainanalternateformoftheduplicationformula:
z!parenleftbig
z−1
2parenrightbig
!=2−2zπ1/2(2z)!. (8.65b)
8.4 The Beta Function 523
Although the integrals used in this derivation are defined only for ℜ(z)>−1, the results
(Eqs. (8.65a)and(8.65b)holdforallregularpoints zbyanalyticcontinuation.8
Using the double factorial notation (Section 8.1), we may rewrite Eq. (8.65a) (with z=
n,aninteger)as
parenleftbig
n+1
2parenrightbig
!=π1/2(2n+1)!!/2n+1. (8.65c)
Thisis oftenconvenientfor eliminatingfactorialsoffractions.
Incomplete Beta Function
Just as there is an incomplete gamma function (Section 8.5), there is also an incomplete
betafunction,
Bx(p,q)=integraldisplayx
0tp−1(1−t)q−1dt,0≤x≤1,p>0,q>0(ifx=1).(8.66)
Clearly,Bx=1(p,q)becomes the regular (complete) beta function, Eq. (8.59a). A power-
series expansion of Bx(p,q)is the subject of Exercises 5.2.18 and 5.7.8. The relation to
hypergeometricfunctionsappearsinSection13.4.
The incomplete beta function makes an appearance in probability theory in calculating
theprobabilityofatmost ksuccessesin nindependenttrials.9
Exercises
8.4.1 Derive the doubling formula for the factorial function by integrating (sin2θ)2n+1=
(2sinθcosθ)2n+1(andusingthebetafunction).
8.4.2 Verifythefollowingbetafunctionidentities:
(a)B(a,b)=B(a+1,b)+B(a,b+1),
(b)B(a,b)=a+b
bB(a,b+1),
(c)B(a,b)=b−1
aB(a+1,b−1),
(d)B(a,b)B(a+b,c)=B(b,c)B(a,b+c).
8.4.3 (a) Showthat
integraldisplay1
−1parenleftbig
1−x2parenrightbig1/2x2ndx=
π/2,n =0
π(2n−1)!!
(2n+2)!!,n=1,2,3,....
8If 2zis anegative integer, weget the validbut unilluminating result ∞=∞.
9W. Feller, An Introduction to Probability Theory and Its Applications , 3rd ed. NewYork: Wiley (1968), Section VI.10.
524 Chapter 8 Gamma–Factorial Function
(b) Showthat
integraldisplay1
−1parenleftbig
1−x2parenrightbig−1/2x2ndx=
π, n =0
π(2n−1)!!
(2n)!!,n=1,2,3,....
8.4.4 Showthat
integraldisplay1
−1parenleftbig
1−x2parenrightbigndx=
22n+1n!n!
(2n+1)!,n>−1
2(2n)!!
(2n+1)!!,n=0,1,2,....
8.4.5 Evaluateintegraltext1
−1(1+x)a(1−x)bdxintermsofthebetafunction.
ANS. 2a+b+1B(a+1,b+1).
8.4.6 Show,bymeansofthebetafunction,that
integraldisplayz
tdx
(z−x)1−α(x−t)α=π
sinπα,0<α<1.
8.4.7 ShowthattheDirichletintegral
integraldisplayintegraldisplay
xpyqdxdy=p!q!
(p+q+2)!=B(p+1,q+1)
p+q+2,
wheretherangeofintegrationisthetriangleboundedbythepositive x-andy-axesand
thelinex+y=1.
8.4.8 Showthatintegraldisplay∞
0integraldisplay∞
0e−(x2+y2+2xycosθ)dxdy=θ
2sinθ.
Whatarethelimitson θ?
Hint.Consideroblique xy-coordinates.
ANS.−π<θ<π .
8.4.9 Evaluate(usingthebetafunction)
(a)
integraldisplayπ/2
0cos1/2θdθ=(2π)3/2
16[(1
4)!]2,
(b)
integraldisplayπ/2
0cosnθdθ=integraldisplayπ/2
0sinnθdθ=√π[(n−1)/2]!
2(n/2)!
=
(n−1)!!
n!!fornodd,
π
2·(n−1)!!
n!!forneven.
8.4 The Beta Function 525
8.4.10 Evaluateintegraltext1
0(1−x4)−1/2dxasa betafunction.
ANS.[(1
4)!]2·4
(2π)1/2=1.311028777.
8.4.11 Given
Jν(z)=2
π1/2(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplayπ/2
0sin2νθcos(zcosθ)dθ,ℜ(ν)>−1
2,
show,withtheaidofbetafunctions,thatthisreducestotheBesselseries
Jν(z)=∞summationdisplay
s=0(−1)s1
s!(s+ν)!parenleftbiggz
2parenrightbigg2s+ν
,
identifying the initial Jνas an integral representation of the Bessel function, Jν(Sec-
tion11.1).
8.4.12 GiventheassociatedLegendrefunction
Pm
m(x)=(2m−1)!!parenleftbig
1−x2parenrightbigm/2,
Section12.5,showthat
(a)integraldisplay1
−1bracketleftbig
Pm
m(x)bracketrightbig2dx=2
2m+1(2m)!,m=0,1,2,...,
(b)integraldisplay1
−1bracketleftbig
Pm
m(x)bracketrightbig2dx
1−x2=2·(2m−1)!,m=1,2,3,....
8.4.13 Showthat
(a)integraldisplay1
0x2s+1parenleftbig
1−x2parenrightbig−1/2dx=(2s)!!
(2s+1)!!,
(b)integraldisplay1
0x2pparenleftbig
1−x2parenrightbigqdx=1
2(p−1
2)!q!
(p+q+1
2)!.
8.4.14 Aparticleofmass mmovinginasymmetricpotentialthatiswelldescribedby V(x)=
A|x|nhas a total energy1
2m(dx/dt)2+V(x)=E. Solving for dx/dtand integrating
wefindthattheperiodofmotionis
τ=2√
2mintegraldisplayxmax
0dx
(E−Axn)1/2,
wherexmaxisaclassicalturningpointgivenby Axn
max=E.Showthat
τ=2
nradicalbigg
2πm
EparenleftbiggE
Aparenrightbigg1/nŴ(1/n)
Ŵ(1/n+1
2).
8.4.15 ReferringtoExercise8.4.14,
526 Chapter 8 Gamma–Factorial Function
(a) Determinethelimitas n→∞of
2
nradicalbigg
2πm
EparenleftbiggE
Aparenrightbigg1/nŴ(1/n)
Ŵ(1/n+1
2).
(b) Find limn→∞τfromthebehavioroftheintegrand (E−Axn)−1/2.
(c) Investigatethebehaviorofthephysicalsystem(potentialwell)as n→∞.Obtain
theperiodfrominspectionofthis limitingphysicalsystem.
8.4.16 Showthat
integraldisplay∞
0sinhαx
coshβxdx=1
2Bparenleftbiggα+1
2,β−α
2parenrightbigg
,−1<α<β.
Hint.Let sinh2x=u.
8.4.17 Thebetadistributionof probabilitytheoryhasaprobabilitydensity
f(x)=Ŵ(α+β)
Ŵ(α)Ŵ(β)xα−1(1−x)β−1,
withxrestrictedtotheinterval(0,1). Showthat
(a)/angbracketleftx/angbracketright(mean)=α
α+β.
(b)σ2(variance)≡/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2=αβ
(α+β)2(α+β+1).
8.4.18 From
limn→∞integraltextπ/2
0sin2nθdθ
integraltextπ/2
0sin2n+1θdθ=1
derivetheWallisformulafor π:
π
2=2·2
1·3·4·4
3·5·6·6
5·7···.
8.4.19 Tabulatethebetafunction B(p,q)forpandq=1.0(0.1)2.0 independently.
Checkvalue. B(1.3,1.7)=0.40774.
8.4.20 (a) Write a subroutine that will calculate the incomplete beta function Bx(p,q).F o r
0.5<x≤1 youwillfinditconvenienttousetherelation
Bx(p,q)=B(p,q)−B1−x(q,p).
(b) Tabulate Bx(3
2,3
2). Spot check your results by using the Gauss–Legendre quadra-
ture.
8.5 Incomplete Gamma Function 527
8.5 T HEINCOMPLETE GAMMA FUNCTIONS AND RELATED
FUNCTIONS
Generalizing the Euler definition of the gamma function (Eq. (8.5)), we define the incom-
pletegammafunctionsbythevariablelimitintegrals
γ(a,x)=integraldisplayx
0e−tta−1dt,ℜ(a)>0
and
Ŵ(a,x)=integraldisplay∞
xe−tta−1dt. (8.67)
Clearly,thetwofunctionsarerelated,for
γ(a,x)+Ŵ(a,x)=Ŵ(a). (8.68)
Thechoiceofemploying γ(a,x)orŴ(a,x)ispurelyamatterofconvenience.Ifthepara-
meterais apositiveinteger,Eq. (8.67) maybeintegratedcompletelytoyield
γ(n,x)=(n−1)!parenleftbigg
1−e−xn−1summationdisplay
s=0xs
s!parenrightbigg
Ŵ(n,x)=(n−1)!e−xn−1summationdisplay
s=0xs
s!,n=1,2,....(8.69)
Fornonintegral a,apower-seriesexpansionof γ(a,x)forsmall xandanasymptoticex-
pansionof Ŵ(a,x)(denotedas I(x,p))aredevelopedinExercise5.7.7andSection5.10:
γ(a,x)=xa∞summationdisplay
n=0(−1)nxn
n!(a+n),|x|∼1(smallx),
Ŵ(a,x)=xa−1e−x∞summationdisplay
n=0(a−1)!
(a−1−n)!·1
xn(8.70)
=xa−1e−x∞summationdisplay
n=0(−1)n(n−a)!
(−a)!·1
xn,x≫1(largex).
Theseincompletegammafunctionsmayalsobeexpressedquiteelegantlyintermsofcon-
fluenthypergeometricfunctions(compareSection13.5).
Exponential Integral
Although the incomplete gamma function Ŵ(a,x)in its general form (Eq. (8.67)) is only
infrequently encountered in physical problems, a special case is quite common and very
528 Chapter 8 Gamma–Factorial Function
FIGURE 8.7Theexponentialintegral,
E1(x)=−Ei(−x).
useful.Wedefinetheexponentialintegralby10
−Ei(−x)≡integraldisplay∞
xe−t
tdt=E1(x). (8.71)
(SeeFig.8.7.)Cautionisneededhere,fortheintegralinEq.(8.71)divergeslogarithmically
asx→0. Toobtainaseriesexpansionforsmall x,we startfrom
E1(x)=Ŵ(0,x)=lim
a→0bracketleftbig
Ŵ(a)−γ(a,x)bracketrightbig
. (8.72)
Wemaysplitthedivergenttermintheseriesexpansionfor γ(a,x),
E1(x)=lim
a→0bracketleftbiggaŴ(a)−xa
abracketrightbigg
−∞summationdisplay
n=1(−1)nxn
n·n!. (8.73)
Usingl’Hôpital’srule(Exercise5.6.8) and
d
dabraceleftbig
aŴ(a)bracerightbig
=d
daa!=d
daeln(a!)=a!ψ(a+1), (8.74)
andthenEq.(8.40),11weobtaintherapidlyconvergingseries
E1(x)=−γ−lnx−∞summationdisplay
n=1(−1)nxn
n·n!. (8.75)
An asymptotic expansion E1(x)≈e−x[1
x−1!
x2+···]forx→∞is developed in Sec-
tion5.10.
10The appearance of the two minus signs in −Ei(−x)is a historical monstrosity. AMS-55, Chapter 5, denotes this integral as
E1(x).SeeAdditional Readings for thereference.
11dxa/da=xalnx.
8.5 Incomplete Gamma Function 529
FIGURE 8.8Sineandcosineintegrals.
Further special forms related to the exponential integral are the sine integral, cosine
integral(Fig.8.8), andlogarithmicintegral,definedby12
si(x)=−integraldisplay∞
xsint
tdt
Ci(x)=−integraldisplay∞
xcost
tdt (8.76)
li(x)=integraldisplayx
0du
lnu=Ei(lnx)
fortheirprincipalbranch,withthebranchcutconventionallychosentobealongthenega-
tive real axis from the branch point at zero. By transforming from real to imaginary argu-
ment,wecanshowthat
si(x)=1
2ibracketleftbig
Ei(ix)−Ei(−ix)bracketrightbig
=1
2ibracketleftbig
E1(ix)−E1(−ix)bracketrightbig
, (8.77)
whereas
Ci(x)=1
2bracketleftbig
Ei(ix)+Ei(−ix)bracketrightbig
=−1
2bracketleftbig
E1(ix)+E1(−ix)bracketrightbig
,|argx|<π
2.(8.78)
Addingthesetworelations,weobtain
Ei(ix)=Ci(x)+isi(x), (8.79)
to show that the relation among these integrals is exactly analogous to that among eix,
cosx, and sin x. Reference to Eqs. (8.71) and (8.78) shows that Ci (x)agrees with the
definitionsof AMS-55(see AdditionalReadingsfor thereference).In termsof E1,
E1(ix)=−Ci(x)+isi(x).
Asymptotic expansions of Ci (x)and si(x)are developed in Section 5.10. Power-series
expansions about the origin for Ci (x),s i(x), and li(x)may be obtained from those for
12Another sine integral is given by Si (x)=si(x)+π/2.
530 Chapter 8 Gamma–Factorial Function
FIGURE 8.9Error function, erf x.
the exponentialintegral, E1(x), or by direct integration, Exercise 8.5.10. The exponential,
sine, and cosine integrals are tabulated in AMS-55, Chapter 5, (see Additional Readings
for the reference) and can also be accessed by symbolic software such as Mathematica,
Maple,Mathcad,andReduce.
Error Integrals
Theerror integrals
erfz=2√πintegraldisplayz
0e−t2dt,erfcz=1−erfz=2√πintegraldisplay∞
ze−t2dt (8.80a)
(normalized so that erf ∞=1) are introduced in Exercise 5.10.4 (Fig. 8.9). Asymptotic
forms are developed there. From the general form of the integrands and Eq. (8.6) we ex-
pect that erf zand erfczmay be written as incomplete gamma functions with a=1
2.T h e
relationsare
erfz=π−1/2γparenleftbig1
2,z2parenrightbig
,erfcz=π−1/2Ŵparenleftbig1
2,z2parenrightbig
. (8.80b)
Thepower-seriesexpansionof erf zfollowsdirectlyfrom Eq.(8.70).
Exercises
8.5.1 Showthat
γ(a,x)=e−x∞summationdisplay
n=0(a−1)!
(a+n)!xa+n
(a) byrepeatedlyintegratingbyparts.
(b) DemonstratethisrelationbytransformingitintoEq.(8.70).
8.5.2 Showthat
(a)dm
dxmbracketleftbig
x−aγ(a,x)bracketrightbig
=(−1)mx−a−mγ(a+m,x),
8.5 Incomplete Gamma Function 531
(b)dm
dxmbracketleftbig
exγ(a,x)bracketrightbig
=exŴ(a)
Ŵ(a−m)γ(a−m,x).
8.5.3 Showthat γ(a,x)andŴ(a,x)satisfy therecurrencerelations
(a)γ(a+1,x)=aγ(a,x)−xae−x,
(b)Ŵ(a+1,x)=aŴ(a,x)+xae−x.
8.5.4 Thepotentialproducedbya 1 Shydrogenelectron(Exercise12.8.6) isgivenby
V(r)=q
4πε0a0braceleftbigg1
2rγ(3,2r)+Ŵ(2,2r)bracerightbigg
.
(a) For r≪1,showthat
V(r)=q
4πε0a0braceleftbigg
1−2
3r2+···bracerightbigg
.
(b) For r≫1,showthat
V(r)=q
4πε0a0·1
r.
Hererisexpressedinunitsof a0,theBohrradius.
Note.Forcomputationatintermediatevaluesof r, Eqs. (8.69) areconvenient.
8.5.5 Thepotentialof a 2 Phydrogenelectronisfoundtobe(Exercise12.8.7)
V(r)=1
4πε0·q
24a0braceleftbigg1
rγ(5,r)+Ŵ(4,r)bracerightbigg
−1
4πε0·q
120a0braceleftbigg1
r3γ(7,r)+r2Ŵ(2,r)bracerightbigg
P2(cosθ).
Hereris expressed in units of a0, the Bohr radius. P2(cosθ)is a Legendre polynomial
(Section12.1).
(a) For r≪1,showthat
V(r)=1
4πε0·q
a0braceleftbigg1
4−1
120r2P2(cosθ)+···bracerightbigg
.
(b) For r≫1,showthat
V(r)=1
4πε0·q
a0rbraceleftbigg
1−6
r2P2(cosθ)+···bracerightbigg
.
8.5.6 Provethattheexponentialintegralhastheexpansion
integraldisplay∞
xe−t
tdt=−γ−lnx−∞summationdisplay
n=1(−1)nxn
n·n!,
whereγis theEuler–Mascheroniconstant.
532 Chapter 8 Gamma–Factorial Function
8.5.7 Showthat E1(z)maybewrittenas
E1(z)=e−zintegraldisplay∞
0e−zt
1+tdt.
Showalsothatwemustimposethecondition |argz|≤π/2.
8.5.8 Related to the exponential integral (Eq. (8.71)) by a simple change of variable is the
function
En(x)=integraldisplay∞
1e−xt
tndt.
Showthat En(x)satisfiestherecurrencerelation
En+1(x)=1
ne−x−x
nEn(x), n=1,2,3,....
8.5.9 WithEn(x)asdefinedinExercise8.5.8,showthat En(0)=1/(n−1),n>1.
8.5.10 Developthefollowingpower-seriesexpansions:
(a) si(x)=−π
2+∞summationdisplay
n=0(−1)nx2n+1
(2n+1)(2n+1)!,
(b) Ci(x)=γ+lnx+∞summationdisplay
n=1(−1)nx2n
2n(2n)!.
8.5.11 Ananalysisof acenter-fedlinearantennaleadstotheexpression
integraldisplayx
01−cost
tdt.
Showthatthisisequalto γ+lnx−Ci(x).
8.5.12 Usingtherelation
Ŵ(a)=γ(a,x)+Ŵ(a,x),
show that if γ(a,x)satisfies the relations of Exercise 8.5.2, then Ŵ(a,x)must satisfy
thesamerelations.
8.5.13 (a) Writeasubroutinethatwillcalculatetheincompletegammafunctions γ(n,x)and
Ŵ(n,x)fornapositiveinteger.Spotcheck Ŵ(n,x)byGauss–Laguerrequadratures.
(b) Tabulate γ(n,x)andŴ(n,x)forx=0.0(0.1)1.0 andn=1,2, 3.
8.5.14 Calculatethepotentialproducedbya1 Shydrogenelectron(Exercise8.5.4)(Fig.8.10).
TabulateV(r)/(q/ 4πε0a0)forx=0.0(0.1)4.0.Checkyourcalculationsfor r≪1 and
forr≫1 bycalculatingthelimitingforms giveninExercise8.5.4.
8.5.15 UsingEqs.(5.182) and(8.75), calculatetheexponentialintegral E1(x)for
(a)x=0.2(0.2)1.0, (b) x=6.0(2.0)10.0.
Program your own calculationbut checkeach value, using a library subroutine if avail-
able.AlsocheckyourcalculationsateachpointbyaGauss–Laguerrequadrature.
8.5 Additional Readings 533
FIGURE 8.10Distributedchargepotentialproduced
bya 1Shydrogenelectron,Exercise8.5.14.
You’ll find that the power-series converges rapidly and yields high precision for small
x.Theasymptoticseries, evenfor x=10,yieldsrelativelypooraccuracy.
Checkvalues. E1(1.0)=0.219384
E1(10.0)=4.15697×10−6.
8.5.16 Thetwoexpressionsfor E1(x),(1)Eq.(5.182),anasymptoticseriesand(2)Eq.(8.75),
a convergent power series, provide a means of calculating the Euler–Mascheroni con-
stantγto high accuracy. Using double precision, calculate γfrom Eq. (8.75), with
E1(x)evaluatedbyEq.(5.182).
Hint.As a convenient choice take xin the range 10 to 20. (Your choice of xwill set
a limit on the accuracy of your result.) To minimize errors in the alternating series of
Eq. (8.75),accumulatethepositiveandnegativetermsseparately.
ANS.For x=10 and“doubleprecision,” γ=0.57721566.
AdditionalReadings
Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and
Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover
(1974). Contains a wealth of information about gamma functions, incomplete gamma functions, exponential
integrals, error functions, and related functions—Chapters 4 to 6.
Artin,E., TheGammaFunction (translatedbyM.Butler).NewYork:Holt,RinehartandWinston(1964).Demon-
strates that if a function f(x)is smooth (log convex) and equal to (n−1)!whenx=n=integer, it is the
gamma function.
Davis, H. T., Tables of the Higher Mathematical Functions . Bloomington, IN: Principia Press (1933). Volume I
contains extensive information on the gamma function and the polygamma functions.
Gradshteyn, I. S.,and I. M.Ryzhik, Table of Integrals, Series, and Products . NewYork: AcademicPress (1980).
Lu k e,Y .L., The Special Functions and Their Approximations , Vol.1. NewYork: AcademicPress (1969).
Luke, Y. L., Mathematical Functions and Their Approximations . New York: Academic Press (1975). This is
an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical
Tables(AMS-55). Chapter1 deals with thegamma function. Chapter 4 treatsthe incomplete gamma function
anda host of relatedfunctions.
This page intentionally left blank
CHAPTER 9
DIFFERENTIAL EQUATIONS
9.1 P ARTIAL DIFFERENTIAL EQUATIONS
Introduction
Inphysics theknowledgeof theforce inanequationof motionusuallyleadstoa differen-
tial equation. Thus, almost all the elementary and numerous advanced parts of theoretical
physics are formulated in terms of differential equations. Sometimes these are ordinary
differential equations in one variable (abbreviated ODEs). More often the equations are
partialdifferentialequations( PDEs) intwoormorevariables.
Let us recall from calculus that the operation of taking an ordinary or partial derivative
isalinearoperation (L),1
d(aϕ(x)+bψ(x))
dx=adϕ
dx+bdψ
dx,
for ODEs involving derivatives in one variable xonly and no quadratic, (dψ/dx)2,o r
higherpowers. Similarly,for partialderivations,
∂(aϕ(x,y)+bψ(x,y))
∂x=a∂ϕ(x,y)
∂x+b∂ψ(x,y)
∂x.
Ingeneral
L(aϕ+bψ)=aL(ϕ)+bL(ψ).
Thus,ODEsandPDEsappearaslinearoperatorequations,
Lψ=F, (9.1)
1We are especially interested in linear operators because in quantum mechanics physical quantities are represented by linear
operators operating in acomplex, infinite-dimensional Hilbert space.
535
536 Chapter 9 Differential Equations
whereFisaknown(source)functionofone(forODEs)ormorevariables(forPDEs), Lis
alinearcombinationofderivatives,and ψistheunknownfunctionorsolution.Anylinear
combination of solutions is again a solution if F=0; this is the superposition principle
forhomogeneousPDEs.
Since the dynamics of many physical systems involve just two derivatives, for exam-
ple, acceleration in classical mechanics and the kinetic energy operator, ∼∇2, in quan-
tum mechanics, differential equations of second order occur most frequently in physics.
(Maxwell’sandDirac’sequationsarefirstorderbutinvolvetwounknownfunctions.Elim-
inating one unknown yields a second-order differential equation for the other (compare
Section1.9).)
Examples of PDEs
AmongthemostfrequentlyencounteredPDEs arethefollowing:
1. Laplace’sequation, ∇2ψ=0.
Thisverycommonandveryimportantequationoccursinstudiesof
a. electromagneticphenomena,includingelectrostatics,dielectrics,steadycurrents,
andmagnetostatics,
b. hydrodynamics(irrotationalflowofperfectfluidandsurface waves),
c. heatflow,
d. gravitation.
2. Poisson’sequation, ∇2ψ=−ρ/ε0.
IncontrasttothehomogeneousLaplaceequation,Poisson’sequationisnonhomo-
geneouswithasourceterm −ρ/ε0.
3. Thewave(Helmholtz)andtime-independentdiffusionequations, ∇2ψ±k2ψ=0.
Theseequationsappearinsuchdiversephenomenaas
a. elasticwavesinsolids,includingvibratingstrings, bars, membranes,
b. sound,oracoustics,
c. electromagneticwaves,
d. nuclearreactors.
4. Thetime-dependentdiffusionequation
∇2ψ=1
a2∂ψ
∂t
and the corresponding four-dimensional forms involving the d’Alembertian, a four-
dimensionalanalogof theLaplacianinMinkowskispace,
∂µ∂µ=∂2=1
c2∂2
∂t2−∇2.
5. Thetime-dependentwaveequation, ∂2ψ=0.
6. Thescalarpotentialequation, ∂2ψ=ρ/ε0.
Like Poisson’s equation, this equation is nonhomogeneous with a source term
ρ/ε0.
9.1 Partial Differential Equations 537
7. TheKlein–Gordonequation, ∂2ψ=−µ2ψ,andthecorrespondingvectorequations,
in which the scalar function ψis replaced by a vector function. Other, more compli-
catedforms arecommon.
8. TheSchrödingerwaveequation,
−¯h2
2m∇2ψ+Vψ=i¯h∂ψ
∂t
and
−¯h2
2m∇2ψ+Vψ=Eψ
for thetime-independentcase.
9. Theequationsfor elasticwavesandviscousfluidsandthetelegraphyequation.
10. Maxwell’s coupled partial differential equations for electric and magnetic fields and
those of Dirac for relativistic electron wave functions. For Maxwell’s equations see
theIntroductionandalsoSection1.9.
Somegeneraltechniquesfor solvingsecond-orderPDEsarediscussedinthissection.
1. Separation of variables, where the PDE is split into ODEs that are related by com-
monconstantsthatappearaseigenvaluesoflinearoperators, Lψ=lψ,usuallyinone
variable. This method is closely related to symmetries of the PDE and a group of
transformations (seeSection4.2).TheHelmholtzequation,listedexample3,hasthis
form, where the eigenvalue k2may arise by separation of the time tfrom the spatial
variables. Likewise, in example 8 the energy Eis the eigenvalue that arises in the
separation of tfromrin the Schrödinger equation. This is pursued in Chapter 10 in
greaterdetail.Section9.2servesasintroduction.ODEsmaybeattackedbyFrobenius’
power-series method in Section 9.5. It does not always work but is often the simplest
methodwhenitdoes.
2. Conversion of a PDE into an integral equation using Green’s functions applies to
inhomogeneous PDEs, such as examples 2 and 6 given above. An introduction to the
Green’sfunctiontechniqueis giveninSection9.7.
3. Other analytical methods, such as the use of integral transforms, are developed and
appliedinChapter15.
Occasionally, we encounter equations of higher order. In both the theory of the slow
motionofaviscousfluidandthetheoryof anelasticbodywefindtheequation
parenleftbig
∇2parenrightbig2ψ=0.
Fortunately, these higher-order differential equations are relatively rare and are not dis-
cussedhere.
Although not so frequently encountered and perhaps not so important as second-order
ODEs, first-order ODEs do appear in theoretical physics and are sometimes intermediate
steps for second-order ODEs. The solutions of some more important types of first-order
ODEs are developed in Section 9.2. First-order PDEs can always be reduced to ODEs.
This is a straightforward but lengthy process and involves a search for characteristics that
arebrieflyintroducedinwhatfollows;for moredetailswerefer totheliterature.
538 Chapter 9 Differential Equations
Classes of PDEs and Characteristics
Second-order PDEs form three classes: (i) Elliptic PDEs involve ∇2orc−2∂2/∂t2+∇2
(ii)parabolicPDEs, a∂/∂t+∇2;(iii)hyperbolicPDEs, c−2∂2/∂t2−∇2.Thesecanonical
operatorscomeaboutbyachangeofvariables ξ=ξ(x,y),η=η(x,y)inalinearoperator
(for twovariablesjustfor simplicity)
L=a∂2
∂x2+2b∂2
∂x∂y+c∂2
∂y2+d∂
∂x+e∂
∂y+f, (9.2)
which can be reduced to the canonical forms (i), (ii), (iii) according to whether the dis-
criminant D=ac−b2>0,=0, or<0. Ifξ(x,y)is determined from the first-order, but
nonlinear,PDE
aparenleftbigg∂ξ
∂xparenrightbigg2
+2bparenleftbigg∂ξ
∂xparenrightbiggparenleftbigg∂ξ
∂yparenrightbigg
+cparenleftbigg∂ξ
∂yparenrightbigg2
=0, (9.3)
thenthecoefficientof ∂2/∂ξ2inL(thatis,Eq.(9.3))iszero.If ηisanindependentsolution
of the same Eq. (9.3), then the coefficient of ∂2/∂η2is also zero. The remaining operator,
∂2/∂ξ∂η,i nLis characteristic of the hyperbolic case (iii) with D<0(a=0=cleads to
D=−b2<0),wherethequadraticform aλ2+2bλ+cfactorizesand,therefore,Eq.(9.3)
has two independent solutions ξ(x,y),η(x,y). In the elliptic case (i) with D>0, the
two solutions ξ,ηare complex conjugate, which, when substituted into Eq. (9.2), remove
the mixed second-order derivative instead of the other second-order terms, yielding the
canonical form (i). In the parabolic case (ii) with D=0, only∂2/∂ξ2remains in L, while
thecoefficientsof theothertwosecond-orderderivativesvanish.
If the coefficients a,b,cinLare functions of the coordinates, then this classificationis
onlylocal;thatis, itstypemaychangeasthecoordinatesvary.
Let us illustrate the physics underlying the hyperbolic case by looking at the wave
equation,Eq.(9.2) (in 1 +1 dimensionsfor simplicity)
parenleftbigg1
c2∂2
∂t2−∂2
∂x2parenrightbigg
ψ=0.
SinceEq. (9.3)nowbecomes
parenleftbigg∂ξ
∂tparenrightbigg2
−c2parenleftbigg∂ξ
∂xparenrightbigg2
=parenleftbigg∂ξ
∂t−c∂ξ
∂xparenrightbiggparenleftbigg∂ξ
∂t+c∂ξ
∂xparenrightbigg
=0
and factorizes, we determine the solution of ∂ξ/∂t−c∂ξ/∂x=0. This is an arbitrary
functionξ=F(x+ct), andξ=G(x−ct)solves∂ξ/∂t+c∂ξ/∂x=0, which is readily
verified.Bylinearsuperpositionageneralsolutionofthewaveequationis ψ=F(x+ct)+
G(x−ct). For periodic functions F,Gwe recognize the lines x+ctandx−ctas the
phases of plane waves or wave fronts, where not all second-order derivatives of ψin the
waveequationarewelldefined.Normaltothewavefrontsaretheraysofgeometricoptics.
Thus, the lines that are solutions of Eq. (9.3) and are called characteristics or sometimes
bicharacteristics (forsecond-orderPDEs)inthemathematicalliteraturecorrespondtothe
wavefronts ofthegeometricopticssolutionof thewaveequation.
9.1 Partial Differential Equations 539
For theellipticcase letusconsiderLaplace’sequation,
∂2ψ
∂x2+∂2ψ
∂y2=0,
fora potential ψoftwovariables.Herethecharacteristicsequation,
parenleftbigg∂ξ
∂xparenrightbigg2
+parenleftbigg∂ξ
∂yparenrightbigg2
=parenleftbigg∂ξ
∂x+i∂ξ
∂yparenrightbiggparenleftbigg∂ξ
∂x−i∂ξ
∂yparenrightbigg
=0,
has complex conjugate solutions: ξ=F(x+iy)for∂ξ/∂x+i(∂ξ/∂y)=0 andξ=
G(x−iy)for∂ξ/∂x−i(∂ξ/∂y)=0.AgeneralsolutionofLaplace’sequationistherefore
ψ=F(x+iy)+G(x−iy),aswellastherealandimaginarypartsof ψ,whicharecalled
harmonic functions,whilepolynomialsolutionsarecalled harmonicpolynomials .
In quantum mechanics the Wentzel–Kramers–Brillouin (WKB) form ψ=exp(−iS/¯h)
forthesolutionof theSchrödingerequation,acomplexparabolicPDE,
parenleftbigg
−¯h2
2m∇2+Vparenrightbigg
ψ=i¯h∂ψ
∂t,
leadstotheHamilton–Jacobiequationofclassicalmechanics,
1
2m(∇S)2+V=∂S
∂t, (9.4)
inthelimit¯h→0.Theclassicalaction SobeystheHamilton–Jacobiequation,whichisthe
analog of Eq. (9.3) of the Schrödinger equation. Substituting ∇ψ=−iψ∇S/¯h,∂ψ/∂t=
−iψ(∂S/∂t)/¯hintotheSchrödingerequation,droppingtheoverallnonvanishingfactor ψ,
andtakingthelimitof theresultingequationas ¯h→0,weindeedobtainEq.(9.4).
Finding solutions of PDEs by solving for the characteristics is one of several general
techniques. For more examples we refer to H. Bateman, Partial Differential Equations
of Mathematical Physics , New York: Dover (1944); K. E. Gustafson, Partial Differential
Equations and Hilbert Space Methods , 2nd ed., New York: Wiley (1987), reprinted Dover
(1998).
In order to derive and appreciate more the mathematical method behind these solutions
of hyperbolic, parabolic, and elliptic PDEs let us reconsider the PDE (9.2) with constant
coefficients and, at first, d=e=f=0 for simplicity. In accordance with the form of the
wavefrontsolutions,weseekasolution ψ=F(ξ)ofEq.(9.2)withafunction ξ=ξ(t,x)
usingthevariables t,xinsteadof x,y.Thenthepartialderivativesbecome
∂ψ
∂x=∂ξ
∂xdF
dξ,∂ψ
∂t=∂ξ
∂tdF
dξ,∂2ψ
∂x2=∂2ξ
∂x2dF
dξ+parenleftbigg∂ξ
∂xparenrightbigg2d2F
dξ2,
and
∂2ψ
∂x∂t=∂2ξ
∂x∂tdF
dξ+∂ξ
∂x∂ξ
∂td2F
dξ2,∂2ψ
∂t2=∂2ξ
∂t2dF
dξ+parenleftbigg∂ξ
∂tparenrightbigg2d2F
dξ2,
using the chain rule of differentiation. When ξdepends on xandtlinearly, these partial
derivatives of ψyield a single term only and solve our PDE (9.2) as a consequence. From
thelinear ξ=αx+βtweobtain
∂2ψ
∂x2=α2d2F
dξ2,∂2ψ
∂x∂t=αβd2F
dξ2,∂2ψ
∂t2=β2d2F
dξ2,
540 Chapter 9 Differential Equations
andourPDE(9.2) becomesequivalenttotheanalogof Eq.(9.3),
parenleftbig
α2a+2αβb+β2cparenrightbigd2F
dξ2=0. (9.5)
A solution ofd2F
dξ2=0 only leads to the trivial ψ=k1x+k2t+k3with constant kithat is
linear in the coordinates and for which all second derivatives vanish. From α2a+2αβb+
β2c=0,ontheotherhand,wegettheratios
β
α=1
cbracketleftbig
−b±parenleftbig
b2−acparenrightbig1/2bracketrightbig
≡r1,2 (9.6)
as solutions of Eq. (9.5) withd2F
dξ2/negationslash=0 in general. The lines ξ1=x+r1tandξ2=x+r2t
willsolvethePDE(9.2),with ψ(x,t)=F(ξ1)+G(ξ2)correspondingtothegeneralization
ofourprevioushyperbolicandellipticPDEexamples.
For the parabolic case, where b2=ac, there is only one ratio from Eq. (9.6), β/α=
r=−b/c, and one solution, ψ(x,t)=F(x−bt/c). In order to find the second gen-
eral solution of our PDE (9.2) we make the Ansatz (trial solution) ψ(x,t)=ψ0(x,t)·
G(x−bt/c).SubstitutingthisintoEq. (9.2) wefind
a∂2ψ0
∂x2+2b∂2ψ0
∂x∂t+c∂2ψ0
∂t2=0
forψ0since, upon replacing F→G,Gsolves Eq. (9.5) with d2G/dξ2/negationslash=0 in general.
The solution ψ0can be any solution of our PDE (9.2), including the trivial ones such as
ψ0=xandψ0=t. Thusweobtainthe generalparabolicsolution ,
ψ(x,t)=Fparenleftbigg
x−b
ctparenrightbigg
+ψ0(x,t)Gparenleftbigg
x−b
ctparenrightbigg
,
withψ0=xorψ0=t,et c.
With the same Ansatz one finds solutions of our PDE (9.2) with a source term, for
example, f/negationslash=0,butstill d=e=0 andconstant a,b,c.
Nextwedeterminethecharacteristics,thatis,curveswherethesecondorderderivatives
ofthesolution ψarenotwelldefined.Thesearethewavefrontsalongwhichthesolutions
of our hyperbolic PDE (9.2) propagate. We solve our PDE with a source term f/negationslash=0 and
Cauchy boundary conditions (see Table 9.1) that are appropriate for hyperbolic PDEs,
whereψanditsnormalderivative ∂ψ/∂narespecifiedonanopencurve
C:x=x(s), t=t(s),
with the parameter sthe length on C. Thendr=(dx,dt)is tangent and ˆnds=(dt,−dx)
is normal to the curve C, and the first-order tangential and normal derivatives are given by
thechainrule
dψ
ds=∇ψ·dr
ds=∂ψ
∂xdx
ds+∂ψ
∂tdt
ds,
dψ
dn=∇ψ·ˆn=∂ψ
∂xdt
ds−∂ψ
∂tdx
ds.
9.1 Partial Differential Equations 541
Fromthesetwolinearequations, ∂ψ/∂tand∂ψ/∂xcanbedeterminedon C, provided
vextendsinglevextendsinglevextendsinglevextendsinglevextendsingledx
dsdt
ds
dt
ds−dx
dsvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−parenleftbiggdx
dsparenrightbigg2
−parenleftbiggdt
dsparenrightbigg2
/negationslash=0.
Forthesecondderivativesweusethechainruleagain:
d
ds∂ψ
∂x=dx
ds∂2ψ
∂x2+dt
ds∂2ψ
∂x∂t, (9.7a)
d
ds∂ψ
∂t=dx
ds∂2ψ
∂x∂t+dt
ds∂2ψ
∂t2. (9.7b)
From our PDE (9.2), and Eqs. (9.7a,b), which are linear in the second-order derivatives,
theycannotbecalculatedwhenthedeterminantvanishes,thatis,
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea2bc
dx
dsdt
ds0
0dx
dsdt
dsvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=aparenleftbiggdt
dsparenrightbigg2
−2bdx
dsdt
ds+cparenleftbiggdx
dsparenrightbigg2
=0. (9.8)
FromEq.(9.8),whichdefinesthecharacteristics,wefindthatthetangentratio dx/dtobeys
cparenleftbiggdx
dtparenrightbigg2
−2bdx
dt+a=0,
so
dx
dt=1
cbracketleftbig
b±parenleftbig
b2−acparenrightbig1/2bracketrightbig
. (9.9)
For the earlier hyperbolic wave (and elliptic potential )equation examples, b=0anda,c
areconstants,sothesolutions ξi=x+trifromEq.(9.6)coincidewiththecharacteristics
ofEq.(9.9).
Nonlinear PDEs
Nonlinear ODEs and PDEs are a rapidly growing and important field. We encountered
earlierthesimplestlinearwaveequation,
∂ψ
∂t+c∂ψ
∂x=0,
asthefirst-orderPDEofthewavefrontsofthewaveequation.Thesimplestnonlinearwave
equation,
∂ψ
∂t+c(ψ)∂ψ
∂x=0, (9.10)
results if the local speed of propagation, c, is not constant but depends on the wave ψ.
When a nonlinear equation has a solution of the form ψ(x,t)=Acos(kx−ωt), where
542 Chapter 9 Differential Equations
ω(k)varies with kso thatω′′(k)/negationslash=0, then it is called dispersive . Perhaps the best-known
nonlineardispersiveequationisthe Korteweg–deVries equation,
∂ψ
∂t+ψ∂ψ
∂x+∂3ψ
∂x3=0, (9.11)
which models the lossless propagation of shallow water waves and other phenomena. It is
widely known for its solitonsolutions. A soliton is a traveling wave with the property of
persisting through an interaction with another soliton: After they pass through each other,
they emerge in the same shape and with the same velocity and acquire no more than a
phase shift. Let ψ(ξ=x−ct)be such a traveling wave. When substituted into Eq. (9.11)
thisyieldsthenonlinearODE
(ψ−c)dψ
dξ+d3ψ
dξ3=0, (9.12)
whichcanbeintegratedtoyield
d2ψ
dξ2=cψ−ψ2
2. (9.13)
There is no additive integration constant in Eq. (9.13) to ensure that d2ψ/dξ2→0 with
ψ→0f o rl a r g e ξ,s oψis localized at the characteristic ξ=0, orx=ct. Multiplying
Eq.(9.13) by dψ/dξandintegratingagainyields
parenleftbiggdψ
dξparenrightbigg2
=cψ2−ψ3
3, (9.14)
wheredψ/dξ→0f o rl a r g e ξ. Taking the root of Eq. (9.14) and integrating once more
yieldsthesolitonsolution
ψ(x−ct)=3c
cosh2parenleftbig√cx−ct
2parenrightbig. (9.15)
Some nonlinear topics, for example, the logistic equation and the onset of chaos, are re-
viewedinChapter18.Formoredetailsandliterature,seeJ.Guckenheimer,P.Holmes,and
F.John,NonlinearOscillations,DynamicalSystemsandBifurcationsofVectorFields ,re v .
ed.,NewYork:Springer-Verlag(1990).
Boundary Conditions
Usually,whenweknowaphysicalsystematsometimeandthelawgoverningthephysical
process, then we are able to predict the subsequent development. Such initial values are
themostcommonboundaryconditionsassociatedwithODEsandPDE.Findingsolutions
that match given points, curves, or surfaces corresponds to boundary value problems. So-
lutionsusuallyarerequiredtosatisfycertainimposed(forexample,asymptotic)boundary
conditions.Theseboundaryconditionsmaytakethreeforms:
1. Cauchyboundaryconditions .Thevalueofafunctionandnormalderivativespecified
ontheboundary.Inelectrostaticsthiswouldmean ϕ,thepotential,and En,thenormal
componentof theelectricfield.
9.2 First-Order Differential Equations 543
Table 9.1
Boundary Typeof partial differential equation
conditions Elliptic Hyperbolic Parabolic
Laplace,Poisson Waveequation in Diffusion equation
in(x,y) (x,t) in(x,t)
Cauchy
Opensurface Unphysical results Unique,stable Toorestrictive
(instability) solution
Closedsurface Toorestrictive Toorestrictive Toorestrictive
Dirichlet
Opensurface Insufficient Insufficient Unique, stable
solutionin one
direction
Closedsurface Unique,stable Solution not unique Too restrictive
solution
Neumann
Opensurface Insufficient Insufficient Unique, stable
solutionin one
direction
Closedsurface Unique,stable Solution not unique Too restrictive
solution
2. Dirichletboundaryconditions .Thevalueof afunctionspecifiedontheboundary.
3. Neumann boundary conditions . The normal derivative (normal gradient) of a func-
tion specified on the boundary. In the electrostatic case this would be Enand there-
foreσ, thesurface chargedensity.
Asummaryoftherelationofthesethreetypesofboundaryconditionstothethreetypes
of two-dimensional partial differential equations is given in Table 9.1. For extended dis-
cussionsofthesepartialdifferentialequationsthereadermayconsultMorseandFeshbach,
Chapter6(see AdditionalReadings).
PartsofTable9.1aresimplyamatterofmaintaininginternalconsistencyorofcommon
sense.Forinstance,for Poisson’sequationwithaclosedsurface, Dirichletconditionslead
to a unique, stable solution. Neumannconditions,independentof the Dirichletconditions,
likewise lead to a unique stable solution independent of the Dirichlet solution. Therefore
Cauchyboundaryconditions(meaningDirichletplusNeumann)couldleadtoaninconsis-
tency.
The term boundary conditions includes as a special case the concept of initial con-
ditions. For instance, specifying the initial position x0and the initial velocity v0in some
dynamicalproblemwouldcorrespondtotheCauchyboundaryconditions.Theonlydiffer-
enceinthepresentusageofboundaryconditionsintheseone-dimensionalproblemsisthat
weare goingtoapplytheconditionson bothendsoftheallowedrangeofthevariable.
9.2 F IRST -ORDER DIFFERENTIAL EQUATIONS
Physics involves some first-order differential equations. For completeness (and review) it
seems desirable to touch on them briefly. We consider here differential equations of the
544 Chapter 9 Differential Equations
generalform
dy
dx=f(x,y)=−P(x,y)
Q(x,y). (9.16)
Equation (9.16) is clearly a first-order, ordinary differential equation. It is first order be-
causeitcontainsthefirstandnohigherderivatives.Itis ordinary becausetheonlyderiva-
tive,dy/dx, is an ordinary, or total, derivative. Equation (9.16) may or may not be linear,
althoughweshalltreatthelinearcaseexplicitlylater,Eq. (9.25).
Separable Variables
FrequentlyEq. (9.16) willhavethespecialform
dy
dx=f(x,y)=−P(x)
Q(y). (9.17)
Thenit mayberewrittenas
P(x)dx+Q(y)dy=0.
Integratingfrom (x0,y0)to(x,y)yields
integraldisplayx
x0P(x)dx+integraldisplayy
y0Q(y)dy=0.
Since the lower limits, x0andy0, contribute constants, we may ignore the lower limits of
integration and simply add a constant of integration. Note that this separation of variables
techniquedoes notrequirethatthedifferentialequationbelinear.
Example 9.2.1 PARACHUTIST
We want to find the velocity of the falling parachutist as a function of time and are partic-
ularly interested in the constant limiting velocity, v0, that comes about by air drag, taken,
to be quadratic, −bv2, and opposing the force of the gravitational attraction, mg,o ft h e
Earth. We choose a coordinate system in which the positive direction is downward so that
the gravitational force is positive. For simplicity we assume that the parachute opens im-
mediately, that is, at time t=0, where v(t=0)=0, our initial condition. Newton’s law
appliedtothefallingparachutistgives
m˙v=mg−bv2,
wheremincludesthemassoftheparachute.
The terminal velocity, v0, can be found from the equation of motion as t→∞; when
thereis noacceleration, ˙v=0,so
bv2
0=mg, orv0=radicalbiggmg
b.
9.2 First-Order Differential Equations 545
Thevariables tandvseparate
dv
g−b
mv2=dt,
whichweintegratebydecomposingthedenominatorintopartialfractions.Therootsofthe
denominatorareat v=±v0.Hence
parenleftbigg
g−b
mv2parenrightbigg−1
=m
2v0bparenleftbigg1
v+v0−1
v−v0parenrightbigg
.
Integratingbothtermsyields
integraldisplayvdV
g−b
mV2=1
2radicalbiggm
gblnv0+v
v0−v=t.
Solvingfor thevelocityyields
v=e2t/T−1
e2t/T+1v0=v0sinht
T
cosht
T=v0tanht
T,
whereT=radicalBig
m
gbis the time constant governing the asymptotic approach of the velocity to
thelimitingvelocity, v0.
Putting in numerical values, g=9.8m/s2and taking b=700 kg/m,m=70 kg, gives
v0=√9.8/10∼1m/s∼3.6k m/h∼2.23 mi/h, the walking speed of a pedestrian at
landing, and T=radicalBig
m
bg=1/√
10·9.8∼0.1 s. Thus, the constant speed v0is reached
withinasecond.Finally,because it isalwaysimportantto checkthesolution ,wev erify
thatoursolutionsatisfies
˙v=cosht/T
cosht/Tv0
T−sinh2t/T
cosh2t/Tv0
T=v0
T−v2
Tv0=g−b
mv2,
that is, Newton’s equation of motion. The more realistic case, where the parachutist is in
free fall with an initial speed vi=v(0)>0 before the parachute opens, is addressed in
Exercise9.2.18. /squaresolid
Exact Differential Equations
WerewriteEq. (9.16)as
P(x,y)dx+Q(x,y)dy=0. (9.18)
Thisequationissaidtobe exactifwecanmatchtheleft-handsideofittoadifferential dϕ,
dϕ=∂ϕ
∂xdx+∂ϕ
∂ydy. (9.19)
Since Eq. (9.18) has a zero on the right, we look for an unknown function ϕ(x,y)=
constant and dϕ=0.
546 Chapter 9 Differential Equations
We have(if suchafunction ϕ(x,y)exists)
P(x,y)dx+Q(x,y)dy=∂ϕ
∂xdx+∂ϕ
∂ydy (9.20a)
and
∂ϕ
∂x=P(x,y),∂ϕ
∂y=Q(x,y). (9.20b)
The necessary and sufficient condition for our equation to be exact is that the second,
mixed partial derivatives of ϕ(x,y)(assumed continuous) are independent of the order of
differentiation:
∂2ϕ
∂y∂x=∂P(x,y)
∂y=∂Q(x,y)
∂x=∂2ϕ
∂x∂y. (9.21)
Note the resemblance to Eqs. (1.133a) of Section 1.13, “Potential Theory.” If Eq. (9.18)
correspondstoa curl(equaltozero), thenapotential, ϕ(x,y),mustexist.
Ifϕ(x,y)exists,thenfromEqs. (9.18) and(9.20a) oursolutionis
ϕ(x,y)=C.
We may construct ϕ(x,y)from its partial derivatives just as we constructed a magnetic
vectorpotentialinSection1.13fromits curl.SeeExercises9.2.7and9.2.8.
It may well turn out that Eq. (9.18) is not exact and that Eq. (9.21) is not satisfied.
However, there always exists at least one and perhaps an infinity of integrating factors
α(x,y)suchthat
α(x,y)P(x,y)dx +α(x,y)Q(x,y)dy =0
is exact. Unfortunately, an integrating factor is not always obvious or easy to find. Unlike
the case of the linear first-order differential equation to be considered next, there is no
systematicwaytodevelopanintegratingfactor for Eq.(9.18).
Adifferentialequationinwhichthevariableshavebeenseparatedisautomaticallyexact.
Anexactdifferentialequationis notnecessarilyseparable.
ThewavefrontmethodofSection9.1alsoworks forafirst-order PDE:
a(x,y)∂ψ
∂x+b(x,y)∂ψ
∂y=0. (9.22a)
We look for a solution of the form ψ=F(ξ), whereξ(x,y)=constant for varying xand
ydefinesthewavefront. Hence
dξ=∂ξ
∂xdx+∂ξ
∂ydy=0, (9.22b)
whilethePDEyields
parenleftbigg
a∂ξ
∂x+b∂ξ
∂yparenrightbiggdF
dξ=0 (9.23a)
withdF/dξ/negationslash=0 ingeneral.ComparingEqs. (9.22b)and(9.23a)yields
dx
a=dy
b, (9.23b)
9.2 First-Order Differential Equations 547
which reduces the PDE to a first-order ODE for the tangent dy/dxof the wave front
functionξ(x,y).
Whenthereis anadditionalsourceterminthePDE,
a∂ψ
∂x+b∂ψ
∂y+cψ=0, (9.23c)
thenweusetheAnsatz ψ=ψ0(x,y)F(ξ) , whichconvertsourPDEto
Fparenleftbigg
a∂ψ0
∂x+b∂ψ0
∂y+cψ0parenrightbigg
+ψ0dF
dξparenleftbigg
a∂ξ
∂x+b∂ξ
∂yparenrightbigg
=0. (9.24)
If we can guess a solution ψ0of Eq. (9.23c), then Eq. (9.24) reduces to our previous
equation,Eq.(9.23a), fromwhichtheODEofEq. (9.23b)follows.
Linear First-Order ODEs
Iff(x,y)inEq.(9.16) hastheform −p(x)y+q(x),thenEq.(9.16) becomes
dy
dx+p(x)y=q(x). (9.25)
Equation (9.25) is the most general linearfirst-order ODE. If q(x)=0, Eq. (9.25) is
homogeneous (in y). A nonzero q(x)may represent a sourceor adriving term . Equa-
tion (9.25) is linear; each term is linear in yordy/dx. There are no higher powers, that
is,y2, and no products, y(dy/dx) . Note that the linearity refers to the yanddy/dx;p(x)
andq(x)need not be linear in x. Equation (9.25), the most important of these first-order
ODEsforphysics, maybesolvedexactly.
L etusl ookf oran integratingfactor α(x)sothat
α(x)dy
dx+α(x)p(x)y=α(x)q(x) (9.26)
mayberewrittenas
d
dxbracketleftbig
α(x)ybracketrightbig
=α(x)q(x). (9.27)
The purpose of this is to make the left-hand side of Eq. (9.25) a derivative so that it can
be integrated—by inspection. It also, incidentally, makes Eq. (9.25) exact. Expanding
Eq.(9.27), weobtain
α(x)dy
dx+dα
dxy=α(x)q(x).
ComparisonwithEq. (9.26) showsthatwemustrequire
dα
dx=α(x)p(x). (9.28)
Hereisadifferentialequationfor α(x),withthevariables αandxseparable .Weseparate
variables,integrate,andobtain
α(x)=expbracketleftbiggintegraldisplayx
p(x)dxbracketrightbigg
(9.29)
548 Chapter 9 Differential Equations
asourintegratingfactor.
Withα(x)known we proceed to integrate Eq. (9.27). This, of course, was the point of
introducing αinthefirst place.Wehave
integraldisplayxd
dxbracketleftbig
α(x)y(x)bracketrightbig
dx=integraldisplayx
α(x)q(x)dx.
Nowintegratingbyinspection,wehave
α(x)y(x)=integraldisplayx
α(x)q(x)dx+C.
The constants from a constant lower limit of integration are lumped into the constant C.
Dividingby α(x), weobtain
y(x)=bracketleftbig
α(x)bracketrightbig−1braceleftbiggintegraldisplayx
α(x)q(x)dx+Cbracerightbigg
.
Finally,substitutinginEq. (9.29)for αyields
y(x)=expbracketleftbigg
−integraldisplayx
p(t)dtbracketrightbiggbraceleftbiggintegraldisplayx
expbracketleftbiggintegraldisplays
p(t)dtbracketrightbigg
q(s)ds+Cbracerightbigg
.(9.30)
Here the (dummy) variables of integration have been rewritten to make them unambigu-
ous. Equation (9.30) is the complete general solution of the linear, first-order differential
equation,Eq.(9.25). Theportion
y1(x)=Cexpbracketleftbigg
−integraldisplayx
p(t)dtbracketrightbigg
(9.31)
correspondstothecase q(x)=0andisageneralsolutionofthehomogeneousdifferential
equation.TheotherterminEq. (9.30),
y2(x)=expbracketleftbigg
−integraldisplayx
p(t)dtbracketrightbiggintegraldisplayx
expbracketleftbiggintegraldisplays
p(t)dtbracketrightbigg
q(s)ds, (9.32)
isaparticularsolutioncorrespondingto thespecificsourceterm q(x).
Note that if our linear first-order differential equation is homogeneous (q=0), then it
is separable. Otherwise, apart from special cases such as p=constant, q=constant, and
q(x)=ap(x), Eq. (9.25)is notseparable.
LetussummarizethissolutionoftheinhomogeneousODEintermsofa methodcalled
variationoftheconstant asfollows.Inthefirststep,wesolvethehomogeneousODEby
separationofvariablesasbefore, giving
y′
y=−p,lny=−integraldisplayx
p(X)dX+lnC, y(x)=Ce−integraltextxp(X)dX.
Inthesecondstep,welettheintegrationconstantbecome x-dependent,thatis, C→C(x).
This is the “variation of the constant” used to solve the inhomogeneous ODE. Differenti-
atingy(x)weobtain
y′=−pCe−integraltext
p(x)dx+C′(x)e−integraltext
p(x)dx=−py(x)+C′(x)e−integraltext
p(x)dx.
9.2 First-Order Differential Equations 549
ComparingwiththeinhomogeneousODEwefindtheODEfor C:
C′e−integraltext
p(x)dx=q,orC(x)=integraldisplayx
eintegraltextXp(Y)dYq(X)dX.
Substitutingthis Cintoy=C(x)e−integraltextxp(X)dXreproducesEq. (9.32).
Example 9.2.2 RL C IRCUIT
Foraresistance-inductancecircuitKirchhoff’slawleadsto
LdI(t)
dt+RI(t)=V(t)
forthecurrent I(t),whereListheinductanceand Ristheresistance,bothconstant. V(t)
isthetime-dependentinputvoltage.
FromEq. (9.29)ourintegratingfactor α(t)is
α(t)=expintegraldisplaytR
Ldt=eRt/L.
ThenbyEq. (9.30),
I(t)=e−Rt/Lbracketleftbiggintegraldisplayt
eRt/LV(t)
Ldt+Cbracketrightbigg
,
withtheconstant Ctobedeterminedbyaninitialcondition(aboundarycondition).
Forthespecialcase V(t)=V0, aconstant,
I(t)=e−Rt/LbracketleftbiggV0
L·L
ReRt/L+Cbracketrightbigg
=V0
R+Ce−Rt/L.
If theinitialconditionis I(0)=0,thenC=−V0/Rand
I(t)=V0
Rbracketleftbig
1−e−Rt/Lbracketrightbig
.
/squaresolid
Nowwe provethe theoremthatthesolutionoftheinhomogeneousODEisuniqueup
toanarbitrarymultipleof thesolutionof thehomogeneousODE .
Toshowthis, suppose y1,y2bothsolvetheinhomogeneousODE,Eq. (9.25);then
y′
1−y′
2+p(x)(y1−y2)=0
follows by subtracting the ODEs and says that y1−y2is a solution of the homogeneous
ODE. The solution of the homogeneous ODE can always be multiplied by an arbitrary
constant.
Wealsoprovethe theoremthatafirst-orderhomogeneousODEhasonlyonelinearly
independent solution . This is meant in the following sense. If two solutions are linearly
dependent ,bydefinitiontheysatisfy ay1(x)+by2(x)=0withnonzeroconstants a, bfor
all values of x.If the only solution of this linear relation is a=0=b, then our solutions
y1andy2ar es aidtobe linearlyindependent .
550 Chapter 9 Differential Equations
Toprovethistheorem,suppose y1,y2bothsolvethehomogeneousODE.Then
y′
1
y1=−p(x)=y′
2
y2implies W(x)≡y′
1y2−y1y′
2≡0.(9.33)
The functional determinant Wis called the Wronskian of the pair y1,y2.We now show
thatW≡0istheconditionforthemtobelinearlydependent.Assuminglineardependence,
thatis,
ay1(x)+by2(x)=0
with nonzero constants a, bfor all values of x, we differentiate this linear relation to get
anotherlinearrelation,
ay′
1(x)+by′
2(x)=0.
The conditionfor thesetwo homogeneouslinear equationsinthe unknowns a,bto havea
nontrivialsolutionisthattheirdeterminantbezero,whichis W=0.
Conversely, from W=0, there follows linear dependence, because we can find a non-
trivialsolutionoftherelation
y′
1
y1=y′
2
y2
byintegration,whichgives
lny1=lny2+lnC,ory1=Cy2.
Linear dependence and the Wronskian are generalized to three or more functions in Sec-
tion9.6.
Exercises
9.2.1 FromKirchhoff’slawthecurrent IinanRC(resistance–capacitance)circuit(Fig.9.1)
obeystheequation
RdI
dt+1
CI=0.
(a) Find I(t).
(b) For a capacitanceof 10,000 µF charged to 100 V and discharging through a resis-
tanceof1M /Omega1,findthecurrent Ifort=0 andfor t=100 seconds.
Note.Theinitialvoltageis I0RorQ/C, whereQ=integraltext∞
0I(t)dt.
9.2.2 TheLaplacetransformof Bessel’sequation (n=0)leadsto
parenleftbig
s2+1parenrightbig
f′(s)+sf(s)=0.
Solvefor f(s).
9.2 First-Order Differential Equations 551
FIGURE 9.1RCcircuit.
9.2.3 Thedecayofa populationbycatastrophictwo-bodycollisionsisdescribedby
dN
dt=−kN2.
Thisis afirst-order, nonlinear differentialequation.Derivethesolution
N(t)=N0parenleftbigg
1+t
τ0parenrightbigg−1
,
whereτ0=(kN0)−1. Thisimpliesaninfinitepopulationat t=−τ0.
9.2.4 The rate of a particular chemicalreaction A+B→Cis proportionalto the concentra-
tionsofthereactants AandB:
dC(t)
dt=αbracketleftbig
A(0)−C(t)bracketrightbigbracketleftbig
B(0)−C(t)bracketrightbig
.
(a) Find C(t)forA(0)/negationslash=B(0).
(b) Find C(t)forA(0)=B(0).
Theinitialconditionis that C(0)=0.
9.2.5 A boat, coasting through the water, experiences a resisting force proportional to vn,v
beingtheboat’sinstantaneousvelocity.Newton’ssecondlawleadsto
mdv
dt=−kvn.
Withv(t=0)=v0,x(t=0)=0, integrate to find vas a function of time and vas a
functionofdistance.
9.2.6 In the first-order differential equation dy/dx=f(x,y)the function f(x,y)is a func-
tionof theratio y/x:
dy
dx=g(y/x).
Showthatthesubstitutionof u=y/xleadstoaseparableequationin uandx.
552 Chapter 9 Differential Equations
9.2.7 Thedifferentialequation
P(x,y)dx+Q(x,y)dy=0
isexact.Constructasolution
ϕ(x,y)=integraldisplayx
x0P(x,y)dx+integraldisplayy
y0Q(x0,y)dy=constant.
9.2.8 Thedifferentialequation
P(x,y)dx+Q(x,y)dy=0
isexact.If
ϕ(x,y)=integraldisplayx
x0P(x,y)dx+integraldisplayy
y0Q(x0,y)dy,
showthat
∂ϕ
∂x=P(x,y),∂ϕ
∂y=Q(x,y).
Henceϕ(x,y)=constant isa solutionoftheoriginaldifferentialequation.
9.2.9 Prove that Eq. (9.26) is exact in the sense of Eq. (9.21), provided that α(x)satisfies
Eq. (9.28).
9.2.10 Acertaindifferentialequationhas theform
f(x)dx+g(x)h(y)dy=0,
withnoneofthefunctions f(x),g(x),h(y)identicallyzero.Showthatanecessaryand
sufficientconditionfor thisequationtobeexactisthat g(x)=constant.
9.2.11 Showthat
y(x)=expbracketleftbigg
−integraldisplayx
p(t)dtbracketrightbiggbraceleftbiggintegraldisplayx
expbracketleftbiggintegraldisplays
p(t)dtbracketrightbigg
q(s)ds+Cbracerightbigg
isa solutionof
dy
dx+p(x)y(x)=q(x)
bydifferentiatingtheexpressionfor y(x)andsubstitutingintothedifferentialequation.
9.2.12 Themotionofabodyfallinginaresistingmediummaybedescribedby
mdv
dt=mg−bv
when the retarding force is proportional to the velocity, v. Find the velocity. Evaluate
theconstantofintegrationbydemandingthat v(0)=0.
9.2 First-Order Differential Equations 553
9.2.13 Radioactivenucleidecayaccordingtothelaw
dN
dt=−λN,
Nbeing the concentration of a given nuclide and λ, the particular decay constant. In a
radioactiveseries of ndifferentnuclides,startingwith N1,
dN1
dt=−λ1N1,
dN2
dt=λ1N1−λ2N2,andsoon.
FindN2(t)for theconditions N1(0)=N0andN2(0)=0.
9.2.14 The rate of evaporation from a particular spherical drop of liquid (constant density) is
proportional to its surface area. Assuming this to be the sole mechanism of mass loss,
findtheradiusof thedropas afunctionoftime.
9.2.15 Inthelinearhomogeneousdifferentialequation
dv
dt=−av
the variables are separable. When the variables are separated, the equation is exact.
Solvethisdifferentialequationsubjectto v(0)=v0bythefollowingthreemethods:
(a) Separatingvariablesandintegrating.
(b) Treatingtheseparatedvariableequationasexact.
(c) Usingtheresultfor alinearhomogeneousdifferentialequation.
ANS.v(t)=v0e−at.
9.2.16 Bernoulli’sequation,
dy
dx+f(x)y=g(x)yn,
is nonlinear for n/negationslash=0 or 1. Show that the substitution u=y1−nreduces Bernoulli’s
equationtoalinearequation.(SeeSection18.4.)
ANS.du
dx+(1−n)f(x)u=(1−n)g(x).
9.2.17 Solve the linear, first-order equation, Eq. (9.25), by assuming y(x)=u(x)v(x), where
v(x)is a solution of the corresponding homogeneous equation [q(x)=0].T h i si st h e
methodof variationofparameters duetoLagrange.Weapplyittosecond-orderequa-
tionsinExercise9.6.25.
9.2.18 (a)SolveExample9.2.1foraninitialvelocity vi=60mi/h,whentheparachuteopens.
Findv(t).(b) For a skydiver in free fall use the friction coefficient b=0.25 kg/m and
massm=70 kg.Whatis thelimitingvelocityinthis case?
554 Chapter 9 Differential Equations
9.3 S EPARATION OF VARIABLES
TheequationsofmathematicalphysicslistedinSection9.1areallpartialdifferentialequa-
tions. Our first technique for their solution splits the partial differential equation of nvari-
ables into nordinary differential equations. Each separation introduces an arbitrary con-
stantofseparation.Ifwehave nvariables,wehavetointroduce n−1constants,determined
bytheconditionsimposedintheproblembeingsolved.
Cartesian Coordinates
InCartesiancoordinatestheHelmholtzequationbecomes
∂2ψ
∂x2+∂2ψ
∂y2+∂2ψ
∂z2+k2ψ=0, (9.34)
usingEq.(2.27)fortheLaplacian.Forthepresentlet k2beaconstant.Perhapsthesimplest
way of treating a partial differential equation such as Eq. (9.34) is to split it into a set of
ordinarydifferentialequations.This maybedoneas follows.Let
ψ(x,y,z)=X(x)Y(y)Z(z) (9.35)
andsubstitutebackintoEq.(9.34). HowdoweknowEq.(9.35) isvalid?Whenthediffer-
entialoperatorsinvariousvariablesareadditiveinthePDE,thatis,whentherearenoprod-
ucts of differential operators in different variables, the separation method usually works.
Weareproceedinginthespiritoflet’stryandseeifitworks.Ifourattemptsucceeds,then
Eq. (9.35) will be justified. If it does not succeed, we shall find out soon enough and then
we shall try another attack, such as Green’s functions, integral transforms, or brute-force
numericalanalysis.With ψassumedgivenbyEq. (9.35), Eq.(9.34) becomes
YZd2X
dx2+XZd2Y
dy2+XYd2Z
dz2+k2XYZ=0. (9.36)
Dividingby ψ=XYZandrearrangingterms,weobtain
1
Xd2X
dx2=−k2−1
Yd2Y
dy2−1
Zd2Z
dz2. (9.37)
Equation (9.37) exhibits one separation of variables. The left-hand side is a function of x
alone, whereas the right-hand side depends only on yandzand not on x.B u tx,y, andz
areallindependentcoordinates.Theequalityofbothsidesdependingondifferentvariables
means that the behavior of xas an independent variable is not determined by yandz.
Therefore,eachsidemustbeequaltoa constant,aconstantof separation.We choose2
1
Xd2X
dx2=−l2, (9.38)
−k2−1
Yd2Y
dy2−1
Zd2Z
dz2=−l2. (9.39)
2The choice of sign, completely arbitrary here, will be fixed in specific problems by the need to satisfy specific boundary
conditions.
9.3 Separation of Variables 555
Now,turningourattentiontoEq.(9.39), weobtain
1
Yd2Y
dy2=−k2+l2−1
Zd2Z
dz2, (9.40)
and a second separation has been achieved. Here we have a function of yequated to a
functionof z,asbefore.Weresolveit,asbefore,byequatingeachsidetoanotherconstant
ofseparation ,2−m2,
1
Yd2Y
dy2=−m2, (9.41)
1
Zd2Z
dz2=−k2+l2+m2=−n2, (9.42)
introducing a constant n2byk2=l2+m2+n2to produce a symmetric set of equations.
NowwehavethreeODEs((9.38),(9.41),and(9.42))toreplaceEq.(9.34).Ourassumption
(Eq. (9.35)) hassucceededandistherebyjustified.
Oursolutionshouldbelabeledaccordingtothechoiceofourconstants l,m,andn;that
is,
ψlm(x,y,z)=Xl(x)Ym(y)Zn(z). (9.43)
Subject to the conditions of the problem being solved and to the condition k2=l2+
m2+n2, we may choose l,m, andnas we like, and Eq. (9.43) will still be a solution of
Eq.(9.34),provided Xl(x)isasolutionofEq.(9.38),andsoon.Wemaydevelop themost
generalsolution ofEq. (9.34) bytakinga linearcombinationof solutions ψlm,
/Psi1=summationdisplay
l,malmψlm. (9.44)
Theconstantcoefficients almarefinallychosentopermit /Psi1tosatisfytheboundarycondi-
tionsoftheproblem,which,asarule, leadtoadiscretesetofvalues l,m.
Circular Cylindrical Coordinates
Withourunknownfunction ψdependenton ρ,ϕ,andz,theHelmholtzequationbecomes
(seeSection2.4for ∇2)
∇2ψ(ρ,ϕ,z)+k2ψ(ρ,ϕ,z)=0, (9.45)
or
1
ρ∂
∂ρparenleftbigg
ρ∂ψ
∂ρparenrightbigg
+1
ρ2∂2ψ
∂ϕ2+∂2ψ
∂z2+k2ψ=0. (9.46)
Asbefore,weassumeafactoredformfor ψ,
ψ(ρ,ϕ,z)=P(ρ)/Phi1(ϕ)Z(z). (9.47)
SubstitutingintoEq.(9.46), wehave
/Phi1Z
ρd
dρparenleftbigg
ρdP
dρparenrightbigg
+PZ
ρ2d2/Phi1
dϕ2+P/Phi1d2Z
dz2+k2P/Phi1Z=0. (9.48)
556 Chapter 9 Differential Equations
Allthepartialderivativeshavebecomeordinaryderivatives.Dividingby P/Phi1Zandmoving
thezderivativetotheright-handsideyields
1
ρPd
dρparenleftbigg
ρdP
dρparenrightbigg
+1
ρ2/Phi1d2/Phi1
dϕ2+k2=−1
Zd2Z
dz2. (9.49)
Again, a function of zon the right appears to depend on a function of ρandϕon the
left. We resolve this by setting each side of Eq. (9.49) equal to the same constant. Let us
choose3−l2. Then
d2Z
dz2=l2Z (9.50)
and
1
ρPd
dρparenleftbigg
ρdP
dρparenrightbigg
+1
ρ2/Phi1d2/Phi1
dϕ2+k2=−l2. (9.51)
Settingk2+l2=n2, multiplyingby ρ2,andrearrangingterms, weobtain
ρ
Pd
dρparenleftbigg
ρdP
dρparenrightbigg
+n2ρ2=−1
/Phi1d2/Phi1
dϕ2. (9.52)
Wemayset theright-handsideto m2and
d2/Phi1
dϕ2=−m2/Phi1. (9.53)
Finally,for the ρdependencewehave
ρd
dρparenleftbigg
ρdP
dρparenrightbigg
+parenleftbig
n2ρ2−m2parenrightbig
P=0. (9.54)
This is Bessel’s differential equation. The solutions and their properties are presented in
Chapter11.TheseparationofvariablesofLaplace’sequationinparaboliccoordinatesalso
givesrisetoBessel’sequation.ItmaybenotedthattheBesselequationisnotoriousforthe
varietyofdisguisesitmayassume.Foranextensivetabulationofpossibleformsthereader
isreferred to TablesofFunctions byJahnkeandEmde.4
The original Helmholtz equation, a three-dimensional PDE, has been replaced by three
ODEs,Eqs. (9.50), (9.53), and(9.54). AsolutionoftheHelmholtzequationis
ψ(ρ,ϕ,z)=P(ρ)/Phi1(ϕ)Z(z). (9.55)
Identifyingthespecific P,/Phi1,Zsolutionsbysubscripts,weseethatthemostgeneralsolu-
tionof theHelmholtzequationis alinearcombinationof theproductsolutions:
/Psi1(ρ,ϕ,z)=summationdisplay
m,namnPmn(ρ)/Phi1m(ϕ)Zn(z). (9.56)
3The choice of sign of the separation constant is arbitrary. However, a minus sign is chosen for the axial coordinate zin expec-
tation of a possible exponential dependence on z(from Eq. (9.50)). A positive sign is chosen for the azimuthal coordinate ϕin
expectation of aperiodic dependenceon ϕ(from Eq.(9.53)).
4E. Jahnke and F. Emde, Tables of functions , 4th rev. ed., New York: Dover (1945), p. 146; also, E. Jahnke, F. Emde, and
F. Lösch, Tables of Higher Functions , 6th ed.,NewYork: McGraw-Hill(1960).
9.3 Separation of Variables 557
Spherical Polar Coordinates
Let us try to separate the Helmholtz equation, again with k2constant, in spherical polar
coordinates.UsingEq.(2.48), weobtain
1
r2sinθbracketleftbigg
sinθ∂
∂rparenleftbigg
r2∂ψ
∂rparenrightbigg
+∂
∂θparenleftbigg
sinθ∂ψ
∂θparenrightbigg
+1
sinθ∂2ψ
∂ϕ2bracketrightbigg
=−k2ψ. (9.57)
Now,inanalogywithEq.(9.35) wetry
ψ(r,θ,ϕ)=R(r)/Theta1(θ)/Phi1(ϕ). (9.58)
BysubstitutingbackintoEq. (9.57)anddividingby R/Theta1/Phi1,weha v e
1
Rr2d
drparenleftbigg
r2dR
drparenrightbigg
+1
/Theta1r2sinθd
dθparenleftbigg
sinθd/Theta1
dθparenrightbigg
+1
/Phi1r2sin2θd2/Phi1
dϕ2=−k2.(9.59)
Note that all derivatives are now ordinary derivatives rather than partials. By multiplying
byr2sin2θ,wecanisolate (1//Phi1)(d2/Phi1/dϕ2)toobtain5
1
/Phi1d2/Phi1
dϕ2=r2sin2θbracketleftbigg
−k2−1
r2Rd
drparenleftbigg
r2dR
drparenrightbigg
−1
r2sinθ/Theta1d
dθparenleftbigg
sinθd/Theta1
dθparenrightbiggbracketrightbigg
.(9.60)
Equation (9.60) relates a function of ϕalone to a function of randθalone. Since r,θ,
andϕareindependentvariables,weequateeachsideofEq.(9.60)toaconstant.Inalmost
all physical problems ϕwill appear as an azimuth angle. This suggests a periodic solution
rather than an exponential. With this in mind, let us use −m2as the separation constant,
which,then,mustbeanintegersquared.Then
1
/Phi1d2/Phi1(ϕ)
dϕ2=−m2(9.61)
and
1
r2Rd
drparenleftbigg
r2dR
drparenrightbigg
+1
r2sinθ/Theta1d
dθparenleftbigg
sinθd/Theta1
dθparenrightbigg
−m2
r2sin2θ=−k2.(9.62)
MultiplyingEq. (9.62)by r2andrearrangingterms, weobtain
1
Rd
drparenleftbigg
r2dR
drparenrightbigg
+r2k2=−1
sinθ/Theta1d
dθparenleftbigg
sinθd/Theta1
dθparenrightbigg
+m2
sin2θ. (9.63)
Again,thevariablesareseparated.Weequateeachsidetoaconstant, Q,andfinallyobtain
1
sinθd
dθparenleftbigg
sinθd/Theta1
dθparenrightbigg
−m2
sin2θ/Theta1+Q/Theta1=0, (9.64)
1
r2d
drparenleftbigg
r2dR
drparenrightbigg
+k2R−QR
r2=0. (9.65)
5The order in which the variables are separated here is not unique. Many quantum mechanics texts show the rdependence split
off first.
558 Chapter 9 Differential Equations
Once more we have replaced a partial differential equation of three variables by three
ODEs.ThesolutionsoftheseODEsarediscussedinChapters11and12.InChapter12,for
example,Eq.(9.64)isidentifiedastheassociatedLegendreequation,inwhichtheconstant
Qbecomes l(l+1);lis a non-negative integer because θis an angular variable. If k2is
a(positive)constant,Eq. (9.65)becomesthesphericalBessel equationof Section11.7.
Again,ourmostgeneralsolutionmaybewritten
ψQm(r,θ,ϕ)=summationdisplay
Q,maQmRQ(r)/Theta1Qm(θ)/Phi1m(ϕ). (9.66)
The restriction that k2be a constant is unnecessarily severe. The separation process will
stillbepossiblefor k2as generalas
k2=f(r)+1
r2g(θ)+1
r2sin2θh(ϕ)+k′2. (9.67)
In the hydrogen atom problem, one of the most important examples of the Schrödinger
wave equation with a closed form solution is k2=f(r), withk2independent of θ,ϕ.
Equation(9.65)for thehydrogenatombecomestheassociatedLaguerreequation.
Thegreatimportanceofthisseparationofvariablesinsphericalpolarcoordinatesstems
fromthefactthatthecase k2=k2(r)coversatremendousamountofphysics:agreatdeal
ofthetheoriesofgravitation,electrostatics,andatomic,nuclear,andparticlephysics.And
withk2=k2(r), the angular dependence is isolated in Eqs. (9.61) and (9.64), which can
besolvedexactly .
Finally, as an illustration of how the constant min Eq. (9.61) is restricted, we note that
ϕin cylindrical and spherical polar coordinates is an azimuth angle. If this is a classical
problem,weshallcertainlyrequirethattheazimuthalsolution /Phi1(ϕ)besingle-valued;that
is,
/Phi1(ϕ+2π)=/Phi1(ϕ). (9.68)
Thisisequivalenttorequiringtheazimuthalsolutiontohaveaperiodof2 π.6Therefore m
mustbeaninteger.Whichintegeritisdependsonthedetailsoftheproblem.Iftheinteger
|m|>1,then/Phi1willhavetheperiod2 π/m.Wheneveracoordinatecorrespondstoanaxis
oftranslationor toanazimuthangle,theseparatedequationalways hastheform
d2/Phi1(ϕ)
dϕ2=−m2/Phi1(ϕ)
forϕ, theazimuthangle,and
d2Z(z)
dz2=±a2Z(z) (9.69)
forz, an axis of translation of the cylindrical coordinate system. The solutions, of course,
are sinazand cosazfor−a2and the corresponding hyperbolic function (or exponentials)
sinhazand coshazfor+a2.
6This also applies in most quantum mechanical problems, but the argument is much more involved. If mis not an integer,
rotation group relations and ladder operator relations (Section 4.3) are disrupted. Compare E. Merzbacher,Single valuedness of
wavefunctions. Am.J .Ph ys. 30: 237 (1962).
9.3 Separation of Variables 559
Table 9.2 SolutionsinSphericalPolarCoordinatesa
ψ=summationdisplay
l,malmψlm
1. ∇2ψ=0ψlm=braceleftBigg
rl
r−l−1bracerightBiggbraceleftBigg
Pm
l(cosθ)
Qm
l(cosθ)bracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBiggb
2.∇2ψ+k2ψ=0ψlm=braceleftBigg
jl(kr)
nl(kr)bracerightBiggbraceleftBigg
Pm
l(cosθ)
Qm
l(cosθ)bracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBiggb
3.∇2ψ−k2ψ=0ψlm=braceleftBigg
il(kr)
kl(kr)bracerightBiggbraceleftBigg
Pm
l(cosθ)
Qm
l(cosθ)bracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBiggb
aReferences for some of the functions are Pm
l(cosθ),m=0, Section 12.1; m/negationslash=0, Sec-
tion12.5; Qm
l(cosθ),Section12.10; jl(kr),nl(kr),il(kr),a ndkl(kr), Section11.7.
bcosmϕandsinmϕmay be replacedby e±imϕ.
Other occasionally encountered ODEs include the Laguerre and associated Laguerre
equationsfromthesupremelyimportanthydrogenatomprobleminquantummechanics:
xd2y
dx2+(1−x)dy
dx+αy=0, (9.70)
xd2y
dx2+(1+k−x)dy
dx+αy=0. (9.71)
Fromthequantummechanicaltheoryof thelinearoscillatorwehaveHermite’sequation,
d2y
dx2−2xdy
dx+2αy=0. (9.72)
Finally,fromtimetotimewefindtheChebyshevdifferentialequation,
parenleftbig
1−x2parenrightbigd2y
dx2−xdy
dx+n2y=0. (9.73)
For convenient reference, the forms of the solutions of Laplace’s equation, Helmholtz’s
equation, and the diffusion equation for spherical polar coordinates are collected in Ta-
ble 9.2. The solutions of Laplace’s equation in circular cylindrical coordinates are pre-
sentedinTable9.3.
Generalpropertiesfollowingfromtheformofthedifferentialequationsarediscussedin
Chapter10. TheindividualsolutionsaredevelopedandappliedinChapters11–13.
Thepracticingphysicistmayandprobablywillmeetothersecond-orderODEs,someof
which may possibly be transformed into the examples studied here. Some of these ODEs
may be solved by the techniques of Sections 9.5 and 9.6. Others may require a computer
fora numericalsolution.
We refertothesecondeditionof thistextfor otherimportantcoordinatesystems.
•ToputtheseparationmethodofsolvingPDEsinperspective,letusreviewitasaconse-
quenceofasymmetryofthePDE.TakethestationarySchrödingerequation Hψ=Eψ
asanexample,withapotential V(r)dependingonlyontheradialdistance r.Thenthis
560 Chapter 9 Differential Equations
Table 9.3 SolutionsinCircularCylindricalCoordinatesa
ψ=summationdisplay
m,αamαψmα
a.∇2ψ+α2ψ=0ψmα=braceleftBigg
Jm(αρ)
Nm(αρ)bracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBiggbraceleftBigg
e−αz
eαzbracerightBigg
b.∇2ψ−α2ψ=0ψmα=braceleftBigg
Im(αρ)
Km(αρ)bracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBiggbraceleftBigg
cosαz
sinαzbracerightBigg
c. ∇2ψ=0ψm=braceleftBigg
ρm
ρ−mbracerightBiggbraceleftBigg
cosmϕ
sinmϕbracerightBigg
aReferencesfortheradialfunctionsare Jm(αρ),Section11.1; Nm(αρ),Section11.3;
Im(αρ)andKm(αρ), Section11.5.
PDE is invariant under rotations that comprise the group SO(3). Its diagonal genera-
tor is the orbital angular momentum operator Lz=−i∂
∂ϕ, and its quadratic (Casimir)
invariant is L2. Since both commute with H(see Section 4.3), we end up with three
separateeigenvalueequations:
Hψ=Eψ, L2ψ=l(l+1)ψ, L zψ=mψ.
Uponreplacing L2
zinL2byitseigenvalue m2,theL2PDEbecomesLegendre’sODE,
andsimilarly Hψ=EψbecomestheradialODEoftheseparationmethodinspherical
polarcoordinates.
•For cylindrical coordinates the PDE is invariant under rotations about the z-axis only,
which form a subgroup of SO(3). This invariance yields the generator Lz=−i∂/∂ϕ
andseparateazimuthalODE Lzψ=mψ,asbefore.Ifthepotential Visinvariantunder
translations along the z-axis, then the generator −i∂/∂zgives the separate ODE in the
zvariable.
•Ingeneral(seeSection4.3),thereare nmutuallycommutinggenerators Hiwitheigen-
valuesmiof the (classical) Lie group Gof ranknand the corresponding Casimir in-
variantsCiwitheigenvalues ci(Chapter4), whichyieldtheseparateODEs
Hiψ=miψ, C iψ=ciψ
inadditiontothe(bynow)radialODE Hψ=Eψ.
Exercises
9.3.1 By letting the operator ∇2+k2act on the general form a1ψ1(x,y,z)+a2ψ2(x,y,z),
show that it is linear, that is, that (∇2+k2)(a1ψ1+a2ψ2)=a1(∇2+k2)ψ1+
a2(∇2+k2)ψ2.
9.3.2 ShowthattheHelmholtzequation,
∇2ψ+k2ψ=0,
9.3 Separation of Variables 561
is still separable in circular cylindrical coordinates if k2is generalized to k2+f(ρ)+
(1/ρ2)g(ϕ)+h(z).
9.3.3 SeparatevariablesintheHelmholtzequationinsphericalpolarcoordinates,splittingoff
the radial dependence first. Show that your separated equations have the same form as
Eqs. (9.61), (9.64), and(9.65).
9.3.4 Verifythat
∇2ψ(r,θ,ϕ)+bracketleftbigg
k2+f(r)+1
r2g(θ)+1
r2sin2θh(ϕ)bracketrightbigg
ψ(r,θ,ϕ)=0
is separable (in spherical polar coordinates). The functions f,g, andhare functions
onlyof thevariablesindicated; k2isaconstant.
9.3.5 An atomic (quantum mechanical) particle is confined inside a rectangular box of sides
a,b,andc.Theparticleisdescribedbyawavefunction ψthatsatisfiestheSchrödinger
waveequation
−¯h2
2m∇2ψ=Eψ.
The wave function is required to vanish at each surface of the box (but not to be identi-
callyzero).Thisconditionimposesconstraintsontheseparationconstantsandtherefore
on the energy E. What is the smallest value of Efor which such a solution can be ob-
tained?
ANS.E=π2¯h2
2mparenleftbigg1
a2+1
b2+1
c2parenrightbigg
.
9.3.6 For a homogeneous spherical solid with constant thermal diffusivity, K, and no heat
sources,theequationofheatconductionbecomes
∂T(r,t)
∂t=K∇2T(r,t).
Assumeasolutionoftheform
T=R(r)T(t)
andseparatevariables.Showthattheradialequationmaytakeonthestandardform
r2d2R
dr2+2rdR
dr+bracketleftbig
α2r2−n(n+1)bracketrightbig
R=0;n=integer.
Thesolutionsof thisequationarecalled sphericalBesselfunctions .
9.3.7 Separate variables in the thermal diffusion equation of Exercise 9.3.6 in circular cylin-
dricalcoordinates.Assumethatyoucanneglectendeffects andtake T=T(ρ,t).
9.3.8 Thequantummechanicalangularmomentumoperatorisgivenby L=−i(r×∇).Show
that
L·Lψ=l(l+1)ψ
leadstotheassociatedLegendreequation.
Hint.Exercises1.9.9and2.5.16maybehelpful.
562 Chapter 9 Differential Equations
9.3.9 The one-dimensional Schrödinger wave equation for a particle in a potential field V=
1
2kx2is
−¯h2
2md2ψ
dx2+1
2kx2ψ=Eψ(x).
(a) Using ξ=axandaconstant λ,weha v e
a=parenleftbiggmk
¯h2parenrightbigg1/4
,λ=2E
¯hparenleftbiggm
kparenrightbigg1/2
;
showthat
d2ψ(ξ)
dξ2+parenleftbig
λ−ξ2parenrightbig
ψ(ξ)=0.
(b) Substituting
ψ(ξ)=y(ξ)e−ξ2/2,
showthat y(ξ)satisfiestheHermitedifferentialequation.
9.3.10 Verifythatthefollowingaresolutionsof Laplace’sequation:
(a)ψ1=1/r,r/negationslash=0, (b) ψ2=1
2rlnr+z
r−z.
Note.T h ezderivatives of 1 /rgenerate the Legendre polynomials, Pn(cosθ),E x e r -
cise12.1.7.The zderivativesof (1/2r)ln[(r+z)/(r−z)]generatetheLegendrefunc-
tions,Qn(cosθ).
9.3.11 If/Psi1isasolutionof Laplace’sequation, ∇2/Psi1=0,showthat ∂/Psi1/∂zisalsoa solution.
9.4 S INGULAR POINTS
In this section the concept of a singular point, or singularity (as applied to a differential
equation), is introduced. The interest in this concept stems from its usefulness in (1) clas-
sifyingODEsand(2)investigatingthefeasibilityofaseriessolution.Thisfeasibilityisthe
topicofFuchs’theorem,Sections9.5and9.6.
All the ODEs listed in Section 9.3 may be solved for d2y/dx2. Using the notation
d2y/dx2=y′′,weha v e7
y′′=f(x,y,y′). (9.74)
If wewriteoursecond-orderhomogeneousdifferentialequation(in y)as
y′′+P(x)y′+Q(x)y=0, (9.75)
wearereadytodefineordinaryandsingularpoints.Ifthefunctions P(x)andQ(x)remain
finite atx=x0, pointx=x0is an ordinary point. However, if either P(x)orQ(x)(or
7This prime notation, y′=dy/dx, was introduced by Lagrange in the late 18th century as an abbreviation for Leibniz’s more
explicit but more cumbersome dy/dx.
9.4 Singular Points 563
both)divergesas x→x0,pointx0isasingularpoint.UsingEq.(9.75),wemaydistinguish
betweentwokindsofsingularpoints.
1. If either P(x)orQ(x)divergesas x→x0but(x−x0)P(x)and
(x−x0)2Q(x)remain finite as x→x0, thenx=x0is called a regular, or nonessen-
tial,singularpoint.
2. IfP(x)divergesfasterthan1 /(x−x0)sothat(x−x0)P(x)goestoinfinityas x→x0,
orQ(x)diverges faster than 1 /(x−x0)2so that(x−x0)2Q(x)goes to infinity as
x→x0, thenpoint x=x0is labeledan irregular ,oressential,singularity .
Thesedefinitionsholdforallfinitevaluesof x0.Theanalysisofpoint x→∞issimilar
tothetreatmentoffunctionsofacomplexvariable(Section6.6).Weset x=1/z,substitute
intothedifferentialequation,andthenlet z→0.Bychangingvariablesinthederivatives,
wehave
dy(x)
dx=dy(z−1)
dzdz
dx=−1
x2dy(z−1)
dz=−z2dy(z−1)
dz, (9.76)
d2y(x)
dx2=d
dzbracketleftbiggdy(x)
dxbracketrightbiggdz
dx=parenleftbig
−z2parenrightbigbracketleftbigg
−2zdy(z−1)
dz−z2d2y(z−1)
dz2bracketrightbigg
=2z3dy(z−1)
dz+z4d2y(z−1)
dz2. (9.77)
Usingtheseresults, wetransform Eq. (9.75)into
z4d2y
dz2+bracketleftbig
2z3−z2Pparenleftbig
z−1parenrightbigbracketrightbigdy
dz+Qparenleftbig
z−1parenrightbig
y=0. (9.78)
Thebehaviorat x=∞(z=0)thendependsonthebehaviorof thenewcoefficients,
2z−P(z−1)
z2andQ(z−1)
z4,
asz→0. If these two expressions remain finite, point x=∞is an ordinary point. If they
divergenomorerapidlythan 1 /zand 1/z2,respectively,point x=∞isaregularsingular
point;otherwiseitis anirregularsingularpoint(anessentialsingularity).
Example 9.4.1
Bessel’sequationis
x2y′′+xy′+parenleftbig
x2−n2parenrightbig
y=0. (9.79)
ComparingitwithEq. (9.75)wehave
P(x)=1
x,Q(x)=1−n2
x2,
which shows that point x=0 is a regular singularity. By inspection we see that there are
no other singular points in the finite range. As x→∞(z→0), from Eq. (9.78) we have
564 Chapter 9 Differential Equations
Table 9.4
Regular Irregular
singularity singularity
Equation x= x=
1. Hypergeometric 0,1,∞ –
x(x−1)y′′+[(1+a+b)x−c]y′+aby=0.
2. Legendrea−1,1,∞ –
(1−x2)y′′−2xy′+l(l+1)y=0.
3. Chebyshev −1,1,∞ –
(1−x2)y′′−xy′+n2y=0.
4. Confluent hypergeometric 0 ∞
xy′′+(c−x)y′−ay=0.
5. Bessel 0 ∞
x2y′′+xy′+(x2−n2)y=0.
6. Laguerrea0 ∞
xy′′+(1−x)y′+ay=0.
7. Simple harmonic oscillator – ∞
y′′+ω2y=0.
8. Hermite – ∞
y′′−2xy′+2αy=0.
aTheassociatedequationshavethesamesingularpoints.
thecoefficients
2z−z
z2and1−n2z2
z4.
Since the latter expression diverges as z4, pointx=∞is an irregular, or essential, singu-
larity. /squaresolid
The ordinary differential equations of Section 9.3, plus two others, the hypergeometric
andtheconfluenthypergeometric,havesingularpoints,asshowninTable9.4.
It will be seen that the first three equations in Table 9.4, hypergeometric, Legendre, and
Chebyshev,allhavethreeregularsingularpoints.Thehypergeometricequation,withregu-
larsingularitiesat0,1,and ∞istakenasthestandard,thecanonicalform.Thesolutionsof
theothertwomaythenbeexpressedintermsofitssolutions,thehypergeometricfunctions.
Thisis doneinChapter13.
In a similar manner, the confluent hypergeometric equation is taken as the canonical
form of a linear second-order differential equation with one regular and one irregular sin-
gularpoint.
Exercises
9.4.1 ShowthatLegendre’sequationhasregularsingularitiesat x=−1,1,and∞.
9.4.2 Show that Laguerre’s equation, like the Bessel equation, has a regular singularity at
x=0 andanirregularsingularityat x=∞.
9.5 Series Solutions — Frobenius’ Method 565
9.4.3 Showthatthesubstitution
x→1−x
2,a=−l, b=l+1,c=1
convertsthehypergeometricequationintoLegendre’sequation.
9.5 S ERIES SOLUTIONS —F ROBENIUS ’M ETHOD
In this section we develop a method of obtaining one solution of the linear, second-order,
homogeneousODE.Themethod,aseriesexpansion,willalwayswork,providedthepoint
ofexpansionisnoworsethanaregularsingularpoint.Inphysicsthisverygentlecondition
isalmostalwayssatisfied.
Alinear,second-order,homogeneous ODEmaybeputintheform
d2y
dx2+P(x)dy
dx+Q(x)y=0. (9.80)
The equation is homogeneous because each term contains y(x)or a derivative; linear
because each y,dy/dx,o rd2y/dx2appears as the first power—and no products. In this
section we develop (at least) one solution of Eq. (9.80). In Section 9.6 we develop the
second, independent solution and prove that no third, independent solution exists .
Thereforethe mostgeneralsolution ofEq. (9.80)maybewrittenas
y(x)=c1y1(x)+c2y2(x). (9.81)
Ourphysicalproblemmayleadtoa nonhomogeneous ,linear,second-orderODE,
d2y
dx2+P(x)dy
dx+Q(x)y=F(x). (9.82)
The function on the right, F(x), represents a source (such as electrostatic charge) or a
driving force (as in a driven oscillator). Specific solutions of this nonhomogeneous equa-
tion are touched on in Exercise 9.6.25. They are explored in some detail, using Green’s
function techniques, in Sections 9.7 and 10.5, and with a Laplace transform technique in
Section15.11.Callingthissolution yp,wemayaddtoitanysolutionofthecorresponding
homogeneousequation(Eq. (9.80)). Hencethe most generalsolution of Eq.(9.82) is
y(x)=c1y1(x)+c2y2(x)+yp(x). (9.83)
Theconstants c1andc2willeventuallybefixedbyboundaryconditions.
For the present, we assume that F(x)=0 and that our differential equation is homoge-
neous. We shall attempt to develop a solution of our linear, second-order, homogeneous
differentialequation,Eq.(9.80),bysubstitutinginapowerserieswithundeterminedcoef-
ficients. Also available as a parameter is the power of the lowest nonvanishing term of the
series. To illustrate, we apply the method to two important differential equations, first the
566 Chapter 9 Differential Equations
linear(classical)oscillatorequation
d2y
dx2+ω2y=0, (9.84)
withknownsolutions y=sinωx,cosωx.
We try
y(x)=xkparenleftbig
a0+a1x+a2x2+a3x3+···parenrightbig
=∞summationdisplay
λ=0aλxk+λ,a 0/negationslash=0, (9.85)
with the exponent kand all the coefficients aλstill undetermined. Note that kneed not be
aninteger.Bydifferentiatingtwice,weobtain
dy
dx=∞summationdisplay
λ=0aλ(k+λ)xk+λ−1,
d2y
dx2=∞summationdisplay
λ=0aλ(k+λ)(k+λ−1)xk+λ−2.
BysubstitutingintoEq.(9.84), wehave
∞summationdisplay
λ=0aλ(k+λ)(k+λ−1)xk+λ−2+ω2∞summationdisplay
λ=0aλxk+λ=0. (9.86)
From our analysis of the uniqueness of power series (Chapter 5), the coefficients of each
powerof xontheleft-handsideof Eq.(9.86) mustvanishindividually.
Thelowestpowerof xappearinginEq.(9.86)is xk−2,forλ=0 inthefirstsummation.
Therequirementthatthecoefficientvanish8yields
a0k(k−1)=0.
We had chosen a0as the coefficient of the lowest nonvanishing terms of the series
(Eq. (9.85)), hence,bydefinition, a0/negationslash=0.Thereforewehave
k(k−1)=0. (9.87)
This equation, coming from the coefficient of the lowest power of x, we call the indicial
equation. The indicial equation and its roots are of critical importance to our analysis.
Ifk=1, the coefficient a1(k+1)kofxk−1must vanish so that a1=0. Clearly, in this
examplewemustrequireeitherthat k=0o rk=1.
Beforeconsideringthesetwopossibilitiesfor k,wereturntoEq.(9.86)anddemandthat
theremainingnetcoefficients,say,thecoefficientof xk+j(j≥0),vanish.Weset λ=j+2
in the first summation and λ=jin the second. (They are independent summations and λ
isadummyindex.)Thisresults in
aj+2(k+j+2)(k+j+1)+ω2aj=0
8Seethe uniqueness of powerseries,Section 5.7.
9.5 Series Solutions — Frobenius’ Method 567
or
aj+2=−ajω2
(k+j+2)(k+j+1). (9.88)
This is a two-term recurrencerelation .9Givenaj, we may compute aj+2and thenaj+4,
aj+6, and so on up as far as desired. Note that for this example, if we start with a0,
Eq. (9.88) leads to the even coefficients a2,a4, and so on, and ignores a1,a3,a5, and
so on. Since a1is arbitrary if k=0 and necessarily zero if k=1, let us set it equal to zero
(compareExercises9.5.3and9.5.4) andthenbyEq. (9.88)
a3=a5=a7=···=0,
and all the odd-numbered coefficients vanish. The odd powers of xwill actually reappear
whenthe secondrootoftheindicialequationis used.
Returning to Eq. (9.87) our indicial equation, we first try the solution k=0. The recur-
rencerelation(Eq. (9.88))becomes
aj+2=−ajω2
(j+2)(j+1), (9.89)
whichleadsto
a2=−a0ω2
1·2=−ω2
2!a0,
a4=−a2ω2
3·4=+ω4
4!a0,
a6=−a4ω2
5·6=−ω6
6!a0,andso on.
Byinspection(andmathematicalinduction),
a2n=(−1)nω2n
(2n)!a0, (9.90)
andoursolutionis
y(x)k=0=a0bracketleftbigg
1−(ωx)2
2!+(ωx)4
4!−(ωx)6
6!+···bracketrightbigg
=a0cosωx. (9.91)
Ifwechoosetheindicialequationroot k=1 (Eq.(9.88)),therecurrencerelationbecomes
aj+2=−ajω2
(j+3)(j+2). (9.92)
9The recurrence relation may involve three terms, that is, aj+2, depending on ajandaj−2. Equation (13.2) for the Hermite
functions provides anexample ofthis behavior.
568 Chapter 9 Differential Equations
Substitutingin j=0,2,4,successively,weobtain
a2=−a0ω2
2·3=−ω2
3!a0,
a4=−a2ω2
4·5=+ω4
5!a0,
a6=−a4ω2
6·7=−ω6
7!a0,andso on.
Again,byinspectionandmathematicalinduction,
a2n=(−1)nω2n
(2n+1)!a0. (9.93)
Forthischoice, k=1,weobtain
y(x)k=1=a0xbracketleftbigg
1−(ωx)2
3!+(ωx)4
5!−(ωx)6
7!+···bracketrightbigg
=a0
ωbracketleftbigg
(ωx)−(ωx)3
3!+(ωx)5
5!−(ωx)7
7!+···bracketrightbigg
=a0
ωsinωx. (9.94)
To summarize this approach, we may write Eq . (9.86)schematically as shown in Fig .9 . 2 .
From the uniqueness of power series (Section5.7),the total coefficient of each power of x
must vanish—all by itself. The requirement that the first coefficient (1)vanish leads to
the indicial equation, Eq . (9.87).The second coefficient is handled by setting a1=0. The
vanishing of the coefficient of xk(and higher powers, taken one at a time )leads to the
recurrencerelation,Eq . (9.88).
This series substitution, known as Frobenius’ method, has given us two series solutions
of the linear oscillator equation. However, there are two points about such series solutions
thatmustbestronglyemphasized:
1. Theseriessolutionshouldalwaysbesubstitutedbackintothedifferentialequation,to
see if it works, as a precaution against algebraic and logical errors. If it works, it is
asolution.
2. Theacceptabilityofaseriessolutiondependsonitsconvergence(includingasymptotic
convergence). It is quite possible for Frobenius’ method to give a series solution that
satisfiestheoriginaldifferentialequationwhensubstitutedintheequationbutthatdoes
FIGURE 9.2Recurrencerelationfrompowerseriesexpansion.
9.5 Series Solutions — Frobenius’ Method 569
notconvergeovertheregionofinterest.Legendre’sdifferentialequationillustratesthis
situation.
Expansion About x0
Equation (9.85) is an expansionabout the origin, x0=0. It is perfectly possible to replace
Eq.(9.85) with
y(x)=∞summationdisplay
λ=0aλ(x−x0)k+λ,a 0/negationslash=0. (9.95)
Indeed,fortheLegendre,Chebyshev,andhypergeometricequationsthechoice x0=1 has
some advantages. The point x0should not be chosen at an essential singularity—or our
Frobenius method will probably fail. The resultant series ( x0an ordinary point or regular
singular point) will be valid where it converges. You can expect a divergence of some sort
when|x−x0|=|zs−x0|,wherezsistheclosestsingularityto x0(inthecomplexplane).
Symmetry of Solutions
Let us note that we obtained one solution of even symmetry, y1(x)=y1(−x), and one of
odd symmetry, y2(x)=−y2(−x). This is not just an accident but a direct consequence of
theform oftheODE.WritingageneralODEas
L(x)y(x)=0, (9.96)
in which L(x)is the differential operator, we see that for the linear oscillator equation
(Eq. (9.84)), L(x)isevenunderparity;thatis,
L(x)=L(−x). (9.97)
Wheneverthedifferentialoperatorhasaspecificparityorsymmetry,eitherevenorodd,
wemayinterchange +xand−x,andEq. (9.96)becomes
±L(x)y(−x)=0, (9.98)
+ifL(x)iseven,−ifL(x)isodd.Clearly,if y(x)isasolutionofthedifferentialequation,
y(−x)is alsoasolution.Thenanysolutionmayberesolvedintoevenandoddparts,
y(x)=1
2bracketleftbig
y(x)+y(−x)bracketrightbig
+1
2bracketleftbig
y(x)−y(−x)bracketrightbig
, (9.99)
thefirst bracketontherightgivinganevensolution,thesecondanoddsolution.
If we refer back to Section 9.4, we can see that Legendre, Chebyshev, Bessel, simple
harmonic oscillator, and Hermite equations (or differential operators) all exhibit this even
parity; that is, their P(x)in Eq. (9.80) is odd and Q(x)even. Solutions of all of them
may be presented as series of even powers of xand separate series of odd powers of x.
The Laguerre differential operator has neither even nor odd symmetry; hence its solutions
cannot be expected to exhibit even or odd parity. Our emphasis on parity stems primarily
from the importance of parity in quantum mechanics. We find that wave functions usually
are either even or odd, meaning that they have a definite parity. Most interactions (beta
decayisthebigexception)arealsoevenor odd,andtheresultis thatparityis conserved.
570 Chapter 9 Differential Equations
Limitations of Series Approach — Bessel’s Equation
This attack on the linear oscillator equation was perhaps a bit too easy. By substituting
the power series (Eq. (9.85)) into the differential equation (Eq. (9.84)), we obtained two
independentsolutionswithnotroubleatall.
Togetsomeideaof whatcanhappenwetrytosolveBessel’sequation,
x2y′′+xy′+parenleftbig
x2−n2parenrightbig
y=0, (9.100)
usingy′fordy/dxandy′′ford2y/dx2. Again,assumingasolutionoftheform
y(x)=∞summationdisplay
λ=0aλxk+λ,
wedifferentiateandsubstituteintoEq. (9.100). Theresultis
∞summationdisplay
λ=0aλ(k+λ)(k+λ−1)xk+λ+∞summationdisplay
λ=0aλ(k+λ)xk+λ
+∞summationdisplay
λ=0aλxk+λ+2−∞summationdisplay
λ=0aλn2xk+λ=0. (9.101)
By setting λ=0, we get the coefficient of xk, the lowest power of xappearing on the
left-handside,
a0bracketleftbig
k(k−1)+k−n2bracketrightbig
=0, (9.102)
andagain a0/negationslash=0 bydefinition.Equation(9.102)thereforeyieldsthe indicialequation
k2−n2=0 (9.103)
withsolutions k=±n.
It isof someinteresttoexaminethecoefficientof xk+1also.Hereweobtain
a1bracketleftbig
(k+1)k+k+1−n2bracketrightbig
=0,
or
a1(k+1−n)(k+1+n)=0. (9.104)
Fork=±n,neitherk+1−nnork+1+nvanishesandwe mustrequirea1=0.10
Proceeding to the coefficient of xk+jfork=n,w es e tλ=jin the first, second, and
fourth terms of Eq. (9.101) and λ=j−2 in the third term. By requiring the resultant
coefficientof xk+1tovanish,weobtain
ajbracketleftbig
(n+j)(n+j−1)+(n+j)−n2bracketrightbig
+aj−2=0.
Whenjis replacedby j+2,thiscanberewrittenfor j≥0a s
aj+2=−aj1
(j+2)(2n+j+2), (9.105)
10k=±n=−1
2areexceptions.
9.5 Series Solutions — Frobenius’ Method 571
which is the desired recurrence relation. Repeated application of this recurrence relation
leadsto
a2=−a01
2(2n+2)=−a0n!
221!(n+1)!,
a4=−a21
4(2n+4)=a0n!
242!(n+2)!,
a6=−a41
6(2n+6)=−a0n!
263!(n+3)!,andsoon,
andingeneral,
a2p=(−1)pa0n!
22pp!(n+p)!. (9.106)
Insertingthesecoefficientsinour assumedseries solution,wehave
y(x)=a0xnbracketleftbigg
1−n!x2
221!(n+1)!+n!x4
242!(n+2)!−···bracketrightbigg
. (9.107)
Insummationform
y(x)=a0∞summationdisplay
j=0(−1)jn!xn+2j
22jj!(n+j)!
=a02nn!∞summationdisplay
j=0(−1)j1
j!(n+j)!parenleftbiggx
2parenrightbiggn+2j
. (9.108)
In Chapter 11 the final summation is identified as the Bessel function Jn(x). Notice that
this solution, Jn(x), has either even or odd symmetry,11as might be expected from the
formofBessel’s equation.
Whenk=−nandnis not an integer, we may generate a second distinct series, to be
labeledJ−n(x).However,when −nisanegativeinteger,troubledevelops.Therecurrence
relation for the coefficients ajis still given by Eq. (9.105), but with 2 nreplaced by−2n.
Then, when j+2=2norj=2(n−1), the coefficient aj+2blows up and we have no
seriessolution.ThiscatastrophecanberemediedinEq.(9.108),asitisdoneinChapter11,
withtheresultthat
J−n(x)=(−1)nJn(x), n aninteger . (9.109)
The second solution simply reproduces the first. We have failed to construct a second in-
dependentsolutionfor Bessel’sequationbythisseries techniquewhen nis aninteger.
By substituting in an infinite series, we have obtained two solutions for the linear oscil-
lator equation and one for Bessel’s equation (two if nis not an integer). To the questions
“Can we always do this? Will this method always work?” the answer is no, we cannot
alwaysdothis.This methodofseries solutionwillnotalwayswork.
11Jn(x)is an even function if nis an even integer, an odd function if nis an odd integer. For nonintegral nthexnhas no such
simple symmetry.
572 Chapter 9 Differential Equations
Regular and Irregular Singularities
Thesuccessoftheseriessubstitutionmethoddependsontherootsoftheindicialequation
and the degree of singularity of the coefficients in the differential equation. To understand
better the effect of the equation coefficients on this naive series substitution approach,
considerfour simpleequations:
y′′−6
x2y=0, (9.110a)
y′′−6
x3y=0, (9.110b)
y′′+1
xy′−a2
x2y=0, (9.110c)
y′′+1
x2y′−a2
x2y=0. (9.110d)
Thereadermayshoweasilythatfor Eq. (9.110a)theindicialequationis
k2−k−6=0,
givingk=3,−2.Sincetheequationishomogeneousin x(counting d2/dx2asx−2),there
is no recurrence relation. However,we are left with two perfectly good solutions, x3and
x−2.
Equation (9.110b) differs from Eq. (9.110a) by only one power of x, but this sends the
indicialequationto
−6a0=0,
with no solution at all, for we have agreed that a0/negationslash=0. Our series substitution worked for
Eq. (9.110a), which had only a regular singularity, but broke down at Eq. (9.110b), which
hasanirregularsingularpointattheorigin.
ContinuingwithEq. (9.110c),wehaveaddedaterm y′/x.Theindicialequationis
k2−a2=0,
but again, there is no recurrence relation. The solutions are y=xa,x−a, both perfectly
acceptableone-termseries.
When we change the power of xin the coefficient of y′from−1t o−2, Eq. (9.110d),
there is a drastic change in the solution. The indicial equation (with only the y′term con-
tributing)becomes
k=0.
Thereisarecurrencerelation,
aj+1=+aja2−j(j−1)
j+1.
9.5 Series Solutions — Frobenius’ Method 573
Unlesstheparameter aisselectedtomaketheseriesterminate,wehave
lim
j→∞vextendsinglevextendsinglevextendsinglevextendsingleaj+1
ajvextendsinglevextendsinglevextendsinglevextendsingle=lim
j→∞j(j+1)
j+1
=lim
j→∞j2
j=∞.
Hence our series solution diverges for all x/negationslash=0. Again, our method worked for
Eq. (9.110c) with a regular singularity but failed when we had the irregular singularity
ofEq. (9.110d).
Fuchs’ Theorem
The answer to the basic question when the method of series substitution can be expected
to work is given by Fuchs’ theorem, which asserts that we can always obtain at least one
power-seriessolution,providedweareexpandingaboutapointthatisanordinarypointor
atworstaregularsingularpoint.
If we attempt an expansion about an irregular or essential singularity, our method may
fail, as it did for Eqs. (9.110b) and (9.110d). Fortunately, the more important equations
of mathematical physics, listed in Section 9.4, have no irregular singularities in the finite
plane.FurtherdiscussionofFuchs’theoremappearsinSection9.6.
From Table 9.4, Section 9.4, infinity is seen to be a singular point for all equations
considered. As a further illustration of Fuchs’ theorem, Legendre’s equation (with infinity
asaregularsingularity)hasaconvergent-seriessolutioninnegativepowersoftheargument
(Section 12.10). In contrast, Bessel’s equation (with an irregular singularity at infinity)
yieldsasymptoticseries(Sections5.10and11.6).Theseasymptoticsolutionsareextremely
useful.
Summary
If we are expanding about an ordinary point or at worst about a regular singularity, the
seriessubstitutionapproachwillyieldatleastonesolution(Fuchs’theorem).
Whether we get one or two distinct solutions depends on the roots of the indicial equa-
tion.
1. If the two roots of the indicial equation are equal, we can obtain only one solution by
thisseries substitutionmethod.
2. If the two roots differ by a nonintegral number, two independent solutions may be
obtained.
3. If thetworootsdiffer byaninteger,thelargerofthetwowillyieldasolution.
The smaller may or may not give a solution, depending on the behavior of the coeffi-
cients. In the linear oscillator equation we obtain two solutions; for Bessel’s equation, we
getonlyonesolution.
574 Chapter 9 Differential Equations
The usefulness of the series solution in terms of what the solution is (that is, numbers)
dependsontherapidityofconvergenceoftheseriesandtheavailabilityofthecoefficients.
ManyODEswillnotyieldnice,simplerecurrencerelationsforthecoefficients.Ingeneral,
the available series will probably be useful for |x|(or|x−x0|) very small. Computers
can be used to determine additional series coefficients using a symbolic language, such
as Mathematica,12Maple,13or Reduce.14Often, however, for numerical work a direct
numericalintegrationwillbepreferred.
Exercises
9.5.1 Uniqueness theorem. The function y(x)satisfies a second-order, linear, homogeneous
differential equation. At x=x0,y(x)=y0anddy/dx=y′
0. Show that y(x)is unique,
in that no other solution of this differential equation passes through the points (x0,y0)
withaslopeof y′
0.
Hint. Assume a second solution satisfying these conditions and compare the Taylor
series expansions.
9.5.2 A series solution of Eq. (9.80) is attempted, expanding about the point x=x0.I fx0is
anordinarypoint,showthattheindicialequationhasroots k=0,1.
9.5.3 In the development of a series solution of the simple harmonic oscillator (SHO) equa-
tion, the second series coefficient a1was neglected except to set it equal to zero. From
the coefficient of the next-to-the-lowest power of x,xk−1, develop a second indicial-
typeequation.
(a) (SHO equation with k=0). Show that a1, may be assigned any finite value (in-
cludingzero).
(b) (SHOequationwith k=1).Showthat a1mustbeset equaltozero.
9.5.4 Analyzetheseriessolutionsofthefollowingdifferentialequationstoseewhen a1may
besetequaltozerowithoutirrevocablylosinganythingandwhen a1mustbesetequal
tozero.
(a) Legendre,(b) Chebyshev,(c) Bessel, (d) Hermite.
ANS.(a) Legendre,(b) Chebyshev,and(d) Hermite:For k=0,a1
maybesetequaltozero;for k=1,a1mustbesetequal
tozero.
(c) Bessel: a1mustbesetequaltozero(exceptfor
k=±n=−1
2).
9.5.5 SolvetheLegendreequation
parenleftbig
1−x2parenrightbig
y′′−2xy′+n(n+1)y=0
bydirectseries substitution.
12S. Wolfram, Mathematica, ASystem for Doing Mathematics byComputer , NewYork: Addison Wesley(1991).
13A.Heck, Introduction to Maple ,NewYork: Springer (1993).
14G.Rayna, ReduceSoftware for Algebraic Computation , NewYork: Springer (1987).
9.5 Series Solutions — Frobenius’ Method 575
(a) Verifythattheindicialequationis
k(k−1)=0.
(b) Using k=0,obtainaseries ofevenpowersof x( a1=0).
yeven=a0bracketleftbigg
1−n(n+1)
2!x2+n(n−2)(n+1)(n+3)
4!x4+···bracketrightbigg
,
where
aj+2=j(j+1)−n(n+1)
(j+1)(j+2)aj.
(c) Using k=1,developaseriesof oddpowersof x( a1=1).
yodd=a1bracketleftbigg
x−(n−1)(n+2)
3!x3+(n−1)(n−3)(n+2)(n+4)
5!x5+···bracketrightbigg
,
where
aj+2=(j+1)(j+2)−n(n+1)
(j+2)(j+3)aj.
(d) Showthatbothsolutions, yevenandyodd,divergefor x=±1iftheseriescontinue
toinfinity .
(e) Finally, show that by an appropriate choice of n, one series at a time may be con-
vertedintoapolynomial,therebyavoidingthedivergencecatastrophe.Inquantum
mechanics this restriction of nto integral values corresponds to quantization of
angularmomentum .
9.5.6 Developseries solutionsfor Hermite’sdifferentialequation
(a)y′′−2xy′+2αy=0.
ANS.k(k−1)=0,indicialequation.
Fork=0,
aj+2=2ajj−α
(j+1)(j+2)(jeven),
yeven=a0bracketleftbigg
1+2(−α)x2
2!+22(−α)(2−α)x4
4!+···bracketrightbigg
.
Fork=1,
aj+2=2ajj+1−α
(j+2)(j+3)(jeven),
yodd=a1bracketleftbigg
x+2(1−α)x3
3!+22(1−α)(3−α)x5
5!+···bracketrightbigg
.
576 Chapter 9 Differential Equations
(b) Show that both series solutions are convergent for all x, the ratio of successive
coefficientsbehaving,forlargeindex,likethecorrespondingratiointheexpansion
of exp(x2).
(c) Show that by appropriate choice of αthe series solutions may be cut off and con-
verted to finite polynomials. (These polynomials, properly normalized, become
theHermitepolynomialsinSection13.1.)
9.5.7 Laguerre’sODEis
xL′′
n(x)+(1−x)L′
n(x)+nLn(x)=0.
Developaseries solutionselectingtheparameter ntomakeyourseriesapolynomial.
9.5.8 SolvetheChebyshevequation
parenleftbig
1−x2parenrightbig
T′′
n−xT′
n+n2Tn=0,
by series substitution. What restrictions are imposedon nif you demandthat the series
solutionconvergefor x=±1?
ANS.Theinfiniteseries doesconvergefor x=±1andno
restrictionon nexists(compareExercise5.2.16).
9.5.9 Solve
parenleftbig
1−x2parenrightbig
U′′
n(x)−3xU′
n(x)+n(n+2)Un(x)=0,
choosing the root of the indicial equation to obtain a series of oddpowers of x. Since
theserieswilldivergefor x=1,choose ntoconvertitintoa polynomial.
k(k−1)=0.
Fork=1,
aj+2=(j+1)(j+3)−n(n+2)
(j+2)(j+3)aj.
9.5.10 Obtainaseries solutionofthehypergeometricequation
x(x−1)y′′+bracketleftbig
(1+a+b)x−cbracketrightbig
y′+aby=0.
Test yoursolutionforconvergence.
9.5.11 Obtaintwoseriessolutionsoftheconfluenthypergeometricequation
xy′′+(c−x)y′−ay=0.
Test yoursolutionsfor convergence.
9.5.12 A quantum mechanical analysis of the Stark effect (parabolic coordinates) leads to the
differentialequation
d
dξparenleftbigg
ξdu
dξparenrightbigg
+parenleftbigg1
2Eξ+α−m2
4ξ−1
4Fξ2parenrightbigg
u=0.
Hereαis a separation constant, Eis the total energy, and Fis a constant, where Fzis
thepotentialenergyaddedtothesystembytheintroductionof anelectricfield.
9.5 Series Solutions — Frobenius’ Method 577
Using the larger root of the indicial equation, develop a power-series solution about
ξ=0.Evaluatethefirstthreecoefficientsintermsof ao.
Indicialequation k2−m2
4=0,
u(ξ)=a0ξm/2braceleftbigg
1−α
m+1ξ+bracketleftbiggα2
2(m+1)(m+2)−E
4(m+2)bracketrightbigg
ξ2+···bracerightbigg
.
Notethattheperturbation Fdoesnotappearuntil a3isincluded.
9.5.13 For the special case of no azimuthal dependence, the quantum mechanical analysis of
thehydrogenmolecularionleadstotheequation
d
dηbracketleftbiggparenleftbig
1−η2parenrightbigdu
dηbracketrightbigg
+αu+βη2u=0.
Develop a power-series solution for u(η). Evaluate the first three nonvanishing coeffi-
cientsinterms of a0.
Indicialequation k(k−1)=0,
uk=1=a0ηbraceleftbigg
1+2−α
6η2+bracketleftbigg(2−α)(12−α)
120−β
20bracketrightbigg
η4+···bracerightbigg
.
9.5.14 To a good approximation, the interaction of two nucleons may be described by a
mesonicpotential
V=Ae−ax
x,
attractive for Anegative. Develop a series solution of the resultant Schrödinger wave
equation
¯h2
2md2ψ
dx2+(E−V)ψ=0
throughthefirstthreenonvanishingcoefficients.
ψ=a0braceleftbig
x+1
2A′x2+1
6bracketleftbig1
2A′2−E′−aA′bracketrightbig
x3+···bracerightbig
,
wheretheprimeindicatesmultiplicationby 2 m/¯h2.
9.5.15 Nearthenucleusofacomplexatomthepotentialenergyof oneelectronis givenby
V=Ze2
rparenleftbig
1+b1r+b2r2parenrightbig
,
wherethecoefficients b1andb2arisefromscreeningeffects.Forthecaseofzeroangu-
larmomentumshowthatthefirstthreetermsofthesolutionoftheSchrödingerequation
have the same form as those of Exercise 9.5.14. By appropriate translation of coeffi-
cients or parameters, write out the first three terms in a series expansion of the wave
function.
578 Chapter 9 Differential Equations
9.5.16 If theparameter a2inEq.(9.110d) isequalto2, Eq.(9.110d)becomes
y′′+1
x2y′−2
x2y=0.
From the indicial equation and the recurrence relation derivea solution y=1+2x+
2x2. Verify that this is indeed a solution by substituting back into the differential equa-
tion.
9.5.17 ThemodifiedBessel function I0(x)satisfiesthedifferentialequation
x2d2
dx2I0(x)+xd
dxI0(x)−x2I0(x)=0.
FromExercise7.3.4 theleadingterminanasymptoticexpansionisfoundtobe
I0(x)∼ex
√
2πx.
Assumeaseriesof theform
I0(x)∼ex
√
2πxbraceleftbig
1+b1x−1+b2x−2+···bracerightbig
.
Determinethecoefficients b1andb2.
ANS.b1=1
8,b2=9
128.
9.5.18 Theevenpower-seriessolutionofLegendre’sequationisgivenbyExercise9.5.5.Take
a0=1 andnnot an even integer, say n=0.5. Calculate the partial sums of the series
throughx200,x400,x600,...,x2000forx=0.95(0.01)1.00. Also, write out the individ-
ualtermcorrespondingtoeachofthesepowers.
Note. This calculation does notconstitute proof of convergence at x=0.99 or diver-
genceatx=1.00,butperhapsyoucanseethedifferenceinthebehaviorofthesequence
ofpartialsumsfor thesetwovaluesof x.
9.5.19 (a) The odd power-series solution of Hermite’s equation is given by Exercise 9.5.6.
Takea0=1. Evaluate this series for α=0,x=1,2,3. Cut off your calculation
after the last term calculated has dropped below the maximum term by a factor of
106ormore.Setanupperboundtotheerrormadeinignoringtheremainingterms
intheinfiniteseries.
(b) Asacheckonthecalculationofpart(a),showthattheHermiteseries yodd(α=0)
correspondstointegraltextx
0exp(x2)dx.
(c) Calculatethisintegralfor x=1,2,3.
9.6 A S ECOND SOLUTION
In Section 9.5 a solution of a second-order homogeneous ODE was developed by substi-
tuting in a power series. By Fuchs’ theorem this is possible, provided the power series is
anexpansionaboutanordinarypointoranonessentialsingularity.15Thereisnoguarantee
15This is whythe classification ofsingularities in Section9.4is of vital importance.
9.6 A Second Solution 579
thatthisapproachwillyieldthetwoindependentsolutionsweexpectfromalinearsecond-
orderODE.Infact,weshallprovethatsuchanODEhasatmosttwolinearlyindependent
solutions.Indeed,thetechniquegaveonlyonesolutionforBessel’sequation( naninteger).
In this section we also develop two methods of obtaining a second independent solution:
an integral method and a power series containing a logarithmic term. First, however, we
considerthequestionofindependenceofaset offunctions.
Linear Independence of Solutions
Givenasetoffunctions ϕλ,thecriterionforlineardependenceistheexistenceofarelation
oftheformsummationdisplay
λkλϕλ=0, (9.111)
in which not all the coefficients kλare zero. On the other hand, if the only solution of
Eq.(9.111) is kλ=0 for allλ,thesetof functions ϕλis saidtobelinearly independent .
It may be helpful to think of linear dependence of vectors. Consider A,B, andCin
three-dimensionalspace,with A·B×C/negationslash=0.Thennonontrivialrelationoftheform
aA+bB+cC=0 (9.112)
exists.A,B, andCare linearly independent.On the other hand, any fourth vector, D,m a y
beexpressedasalinearcombinationof A,B,andC(seeSection3.1).Wecanalwayswrite
anequationof theform
D−aA−bB−cC=0, (9.113)
and the four vectors are notlinearly independent. The three noncoplanar vectors A,B,
andCspanourrealthree-dimensionalspace.
If a set of vectors or functions are mutually orthogonal, then they are automatically lin-
early independent. Orthogonality implies linear independence. This can easily be demon-
stratedbytakinginnerproducts(scalarordotproductforvectors,orthogonalityintegralof
Section10.2for functions).
Let us assume that the functions ϕλare differentiable as needed. Then, differentiating
Eq.(9.111) repeatedly,wegenerateaset ofequations
summationdisplay
λkλϕ′
λ=0, (9.114)
summationdisplay
λkλϕ′′
λ=0, (9.115)
and so on. This gives us a set of homogeneous linear equations in which kλare the un-
known quantities. By Section 3.1 there is a solution kλ/negationslash=0 only if the determinant of the
coefficientsofthe kλ’vanishes.This means
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleϕ1ϕ2···ϕn
ϕ′
1ϕ′
2···ϕ′
n
··· ··· ··· ···
ϕ(n−1)
1ϕ(n−1)
2···ϕ(n−1)
nvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (9.116)
Thisdeterminantis calledthe Wronskian .
580 Chapter 9 Differential Equations
1. If the Wronskian is not equal to zero, then Eq. (9.111) has no solution other than
kλ=0.Theset offunctions ϕλisthereforelinearlyindependent.
2. IftheWronskianvanishesatisolatedvaluesoftheargument,thisdoesnotnecessarily
provelineardependence(unlessthesetoffunctionshasonlytwofunctions).However,
if the Wronskian is zero over the entire range of the variable, the functions ϕλare
linearly dependent over this range16(compare Exercise 9.5.2 for the simple case of
twofunctions).
Example 9.6.1 LINEAR INDEPENDENCE
The solutions of the linear oscillator equation (9.84) are ϕ1=sinωx,ϕ2=cosωx.T h e
Wronskianbecomes
vextendsinglevextendsinglevextendsinglevextendsinglesinωxcosωx
ωcosωx−ωsinωxvextendsinglevextendsinglevextendsinglevextendsingle=−ω/negationslash=0.
These two solutions, ϕ1andϕ2, are therefore linearly independent. For just two functions
thismeansthatoneisnotamultipleoftheother,whichisobviouslytrueinthiscase.
Youknowthat
sinωx=±parenleftbig
1−cos2ωxparenrightbig1/2,
but this is notalinearrelation,oftheformof Eq.(9.111). /squaresolid
Examples 9.6.2 LINEAR DEPENDENCE
Foranillustrationoflineardependence,considerthesolutionsoftheone-dimensionaldif-
fusion equation. We have ϕ1=exandϕ2=e−x, and we add ϕ3=coshx, also a solution.
TheWronskianis
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleexe−xcoshx
ex−e−xsinhx
exe−xcoshxvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0.
The determinant vanishes for all xbecause the first and third rows are identical. Hence
ex,e−x, and cosh xare linearly dependent, and, indeed, we have a relation of the form of
Eq.(9.111):
ex+e−x−2coshx=0 with kλ/negationslash=0. /squaresolid
Now we are ready to prove the theorem that a second-order homogeneous ODE has
twolinearlyindependentsolutions .
Suppose y1,y2,y3are three solutions of the homogeneous ODE (9.80). Then we
form the Wronskian Wjk=yjy′
k−y′
jykof any pair yj,ykof them and recall that
16Compare H. Lass, Elements of Pure and Applied Mathematics , New York: McGraw-Hill (1957), p. 187, for proof of this
assertion. It is assumed that the functions have continuous derivatives and that at least one of the minors of the bottom row of
Eq.(9.116) (Laplaceexpansion) does not vanish in [a,b], the interval under consideration.
9.6 A Second Solution 581
W′
jk=yjy′′
k−y′′
jyk.We divide each ODE by y, getting−Qon their right-hand side,
so
y′′
j
yj+Py′
j
yj=−Q(x)=y′′
k
yk+Py′
k
yk.
Multiplyingby yjyk, wefind
(yjy′′
k−y′′
jyk)+P(yjy′
k−y′
jyk)=0,orW′
jk=−PWjk(9.117)
foranypairofsolutions.FinallyweevaluatetheWronskianofallthreesolutions,expand-
ingit alongthesecondrowandusingtheODEsfor the Wjk:
W=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingley1y2y3
y′
1y′
2y′
3
y′′
1y′′
2y′′
3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−y′
1W′
23+y′
2W′
13−y′
3W′
12
=P(y′
1W23−y′
2W13+y′
3W12)=−Pvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingley1y2y3
y′
1y′
2y′
3
y′
1y′
2y′
3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0.
The vanishing Wronskian, W=0, because of two identical rows, is just the condition for
linear dependence of the solutions yj.Thus, there are at most two linearly independent
solutions of the homogeneous ODE. Similarly one can prove that a linear homogeneous
nth-order ODE has nlinearly independent solutions yj, so the general solution y(x)=summationtextcjyj(x)isalinearcombinationofthem.
A Second Solution
Returningtoourlinear,second-order,homogeneousODEof thegeneralform
y′′+P(x)y′+Q(x)y=0, (9.118)
lety1andy2betwoindependentsolutions.ThentheWronskian,bydefinition,is
W=y1y′
2−y′
1y2. (9.119)
BydifferentiatingtheWronskian,weobtain
W′=y′
1y′
2+y1y′′
2−y′′
1y2−y′
1y′
2
=y1bracketleftbig
−P(x)y′
2−Q(x)y2bracketrightbig
−y2bracketleftbig
−P(x)y′
1−Q(x)y1bracketrightbig
=−P(x)(y 1y′
2−y′
1y2).
Theexpressioninparenthesesis just W, theWronskian, andwehave
W′=−P(x)W. (9.120)
Inthespecialcasethat P(x)=0,thatis,
y′′+Q(x)y=0, (9.121)
theWronskian
W=y1y′
2−y′
1y2=constant. (9.122)
582 Chapter 9 Differential Equations
Sinceouroriginaldifferentialequationishomogeneous,wemaymultiplythesolutions y1
andy2by whatever constants we wish and arrange to have the Wronskian equal to unity
(or−1).Thiscase, P(x)=0,appearsmorefrequentlythanmightbeexpected.Recallthat
the portion of ∇2(ψ
r)in spherical polar coordinates involving radial derivatives contains
no first radial derivative. Finally, every linear second-order differential equation can be
transformedintoanequationof theformofEq. (9.121)(compareExercise9.6.11).
For the general case, let us now assume that we have one solution of Eq. (9.118) by
a series substitution (or by guessing). We now proceed to develop a second, independent
solutionfor which W/negationslash=0.RewritingEq.(9.120) as
dW
W=−Pdx,
weintegrateoverthevariable x,fromatox,toobtain
lnW(x)
W(a)=−integraldisplayx
aP(x1)dx1,
or17
W(x)=W(a)expbracketleftbigg
−integraldisplayx
aP(x1)dx1bracketrightbigg
. (9.123)
But
W(x)=y1y′
2−y′
1y2=y2
1d
dxparenleftbiggy2
y1parenrightbigg
. (9.124)
BycombiningEqs. (9.123)and(9.124), wehave
d
dxparenleftbiggy2
y1parenrightbigg
=W(a)exp[−integraltextx
aP(x1)dx1]
y2
1. (9.125)
Finally,byintegratingEq. (9.125)from x2=btox2=xweget
y2(x)=y1(x)W(a)integraldisplayx
bexp[−integraltextx2
aP(x1)dx1]
[y1(x2)]2dx2. (9.126)
Hereaandbare arbitrary constants and a term y1(x)y2(b)/y1(b)has been dropped, for it
leadstonothingnew.Since W(a),theWronskianevaluatedat x=a,isaconstantandour
solutionsforthehomogeneousdifferentialequationalwayscontainanunknownnormaliz-
ingfactor, weset W(a)=1 andwrite
y2(x)=y1(x)integraldisplayxexp[−integraltextx2P(x1)dx1]
[y1(x2)]2dx2. (9.127)
Note that the lower limits x1=aandx2=bhave been omitted. If they are retained, they
simply make a contribution equal to a constant times the known first solution, y1(x), and
17IfP(x)remains finite in the domain of interest, W(x)/negationslash=0 unlessW(a)=0. That is, the Wronskian of our two solutions is
either identically zero or never zero. However, if P(x)does not remain finite in our interval, then W(x)can have isolated zeros
in that domain and one must be careful to choose aso thatW(a)/negationslash=0.
9.6 A Second Solution 583
hence add nothing new. If we have the important special case of P(x)=0, Eq. (9.127)
reducesto
y2(x)=y1(x)integraldisplayxdx2
[y1(x2)]2. (9.128)
This means that by using either Eq. (9.127) or Eq. (9.128) we can take one known solu-
tion and by integrating can generate a second, independent solution of Eq. (9.118). This
techniqueis used in Section 12.10 to generate a second solution of Legendre’s differential
equation.
Example 9.6.3 ASECOND SOLUTION FOR THE LINEAR OSCILLATOR EQUATION
Fromd2y/dx2+y=0 withP(x)=0 let one solution be y1=sinx. By applying
Eq.(9.128), weobtain
y2(x)=sinxintegraldisplayxdx2
sin2x2=sinx(−cotx)=−cosx,
whichisclearlyindependent(nota linearmultiple)of sin x. /squaresolid
Series Form of the Second Solution
Further insight into the nature of the second solution of our differential equation may be
obtainedbythefollowingsequenceofoperations.
1. Express P(x)andQ(x)inEq.(9.118) as
P(x)=∞summationdisplay
i=−1pixi,Q(x)=∞summationdisplay
j=−2qjxj. (9.129)
The lower limits of the summations are selected to create the strongest possible reg-
ularsingularity (at the origin). These conditions just satisfy Fuchs’ theorem and thus
helpusgainabetterunderstandingofFuchs’theorem.
2. Developthefirst fewtermsof apower-seriessolution,asinSection9.5.
3. Using this solution as y1, obtain a second series type solution, y2, with Eq. (9.127),
integratingtermbyterm.
ProceedingwithStep1,wehave
y′′+parenleftbig
p−1x−1+p0+p1x+···parenrightbig
y′+parenleftbig
q−2x−2+q−1x−1+···parenrightbig
y=0,(9.130)
inwhichpoint x=0isatworstaregularsingularpoint.If p−1=q−1=q−2=0,itreduces
toanordinarypoint.Substituting
y=∞summationdisplay
λ=0aλxk+λ
584 Chapter 9 Differential Equations
(Step2), weobtain
∞summationdisplay
λ=0(k+λ)(k+λ−1)aλxk+λ−2+∞summationdisplay
i=−1pixi∞summationdisplay
λ=0(k+λ)aλxk+λ−1
+∞summationdisplay
j=−2qjxj∞summationdisplay
λ=0aλxk+λ=0. (9.131)
Assumingthat p−1/negationslash=0,q−2/negationslash=0,our indicialequationis
k(k−1)+p−1k+q−2=0,
whichsetsthenetcoefficientof xk−2equaltozero.Thisreducesto
k2+(p−1−1)k+q−2=0. (9.132)
We denote the two roots of this indicial equation by k=αandk=α−n, wherenis zero
or a positive integer. (If nis not an integer, we expect two independent series solutions by
themethodsofSection9.5andwearedone.)Then
(k−α)(k−α+n)=0, (9.133)
or
k2+(n−2α)k+α(α−n)=0,
andequatingcoefficientsof kinEqs. (9.132)and(9.133),wehave
p−1−1=n−2α. (9.134)
Theknownseriessolutioncorrespondingtothelargerroot k=αmaybewrittenas
y1=xα∞summationdisplay
λ=0aλxλ.
Substitutingthisseries solutionintoEq.(9.127) (Step3), wearefacedwith
y2(x)=y1(x)integraldisplayxexp(−integraltextx2
asummationtext∞
i=−1pixi
1dx1)
x2α
2(summationtext∞
λ=0aλxλ
2)2dx2, (9.135)
where the solutions y1andy2have been normalized so that the Wronskian W(a)=1.
Tacklingtheexponentialfactorfirst, wehave
integraldisplayx2
a∞summationdisplay
i=−1pixi
1dx1=p−1lnx2+∞summationdisplay
k=0pk
k+1xk+1
2+f(a) (9.136)
9.6 A Second Solution 585
withf(a)anintegrationconstantthatmaydependon a.Hence,
expparenleftbigg
−integraldisplayx2
asummationdisplay
ipixi
1dx1parenrightbigg
=expbracketleftbig
−f(a)bracketrightbig
x−p−1
2expparenleftbigg
−∞summationdisplay
k=0pk
k+1xk+1
2parenrightbigg
=expbracketleftbig
−f(a)bracketrightbig
x−p−1
2bracketleftbigg
1−∞summationdisplay
k=0pk
k+1xk+1
2+1
2!parenleftbigg
−∞summationdisplay
k=0pk
k+1xk+1
2parenrightbigg2
+···bracketrightbigg
.
(9.137)
Thisfinalseriesexpansionoftheexponentialiscertainlyconvergentiftheoriginalexpan-
sionofthecoefficient P(x)was uniformlyconvergent.
ThedenominatorinEq. (9.135)maybehandledbywriting
bracketleftbigg
x2α
2parenleftbigg∞summationdisplay
λ=0aλxλ
2parenrightbigg2bracketrightbigg−1
=x−2α
2parenleftbigg∞summationdisplay
λ=0aλxλ
2parenrightbigg−2
=x−2α
2∞summationdisplay
λ=0bλxλ
2.(9.138)
Neglecting constant factors, which will be picked up anyway by the requirement that
W(a)=1,weobtain
y2(x)=y1(x)integraldisplayx
x−p−1−2α
2parenleftbigg∞summationdisplay
λ=0cλxλ
2parenrightbigg
dx2. (9.139)
ByEq. (9.134),
x−p−1−2α
2=x−n−1
2, (9.140)
andwehaveassumedherethat nisaninteger.SubstitutingthisresultintoEq.(9.139), we
obtain
y2(x)=y1(x)integraldisplayxparenleftbig
c0x−n−1
2+c1x−n
2+c2x−n+1
2+···+cnx−1
2+···parenrightbig
dx2.(9.141)
The integration indicated in Eq. (9.141) leads to a coefficient of y1(x)consisting of two
parts:
1. Apowerseries startingwith x−n.
2. Alogarithmtermfromtheintegrationof x−1(whenλ=n).Thistermalwaysappears
whennisaninteger, unlesscnfortuitouslyhappenstovanish.18
18For parity considerations, ln xis takentobe ln |x|,e v en.
586 Chapter 9 Differential Equations
Example 9.6.4 ASECOND SOLUTION OF BESSEL ’SEQUATION
FromBessel’s equation,Eq. (9.100)(dividedby x2toagreewithEq.(9.118)), wehave
P(x)=x−1Q(x)=1 forthecase n=0.
Hencep−1=1,q0=1;allother piandqjvanish.TheBesselindicialequationis
k2=0
(Eq. (9.103)with n=0).HenceweverifyEqs. (9.132)to(9.134)with nandα=0.
Our first solution is available from Eq. (9.108). Relabeling it to agree with Chapter 11
(andusing a0=1),weobtain19
y1(x)=J0(x)=1−x2
4+x4
64−Oparenleftbig
x6parenrightbig
. (9.142a)
Now, substituting all this into Eq. (9.127), we have the specific case corresponding to
Eq.(9.135):
y2(x)=J0(x)integraldisplayxexp[−integraltextx2x−1
1dx1]
[1−x2
2/4+x4
2/64−···]2dx2. (9.142b)
Fromthenumeratorof theintegrand,
expbracketleftbigg
−integraldisplayx2dx1
x1bracketrightbigg
=exp[−lnx2]=1
x2.
Thiscorrespondstothe x−p−1
2inEq.(9.137).Fromthedenominatoroftheintegrand,using
abinomialexpansion,weobtain
bracketleftbigg
1−x2
2
4+x4
2
64bracketrightbigg−2
=1+x2
2
2+5x4
2
32+···.
CorrespondingtoEq. (9.139),wehave
y2(x)=J0(x)integraldisplayx1
x2bracketleftbigg
1+x2
2
2+5x4
2
32+···bracketrightbigg
dx2
=J0(x)braceleftbigg
lnx+x2
4+5x4
128+···bracerightbigg
. (9.142c)
Letuscheckthisresult.FromEqs.(11.62)and(11.64),whichgivethestandardformof
thesecondsolution(higher-ordertermsareneeded)
N0(x)=2
π[lnx−ln2+γ]J0(x)+2
πbraceleftbiggx2
4−3x4
128+···bracerightbigg
. (9.142d)
Two points arise: (1) Since Bessel’s equation is homogeneous, we may multiply y2(x)by
any constant. To match N0(x), we multiply our y2(x)by 2/π. (2) To our second solution,
19Thecapital O(order of) as written here means terms proportional to x6and possibly higher powers of x.
9.6 A Second Solution 587
(2/π)y2(x),wemayaddanyconstantmultipleofthefirstsolution.Again,tomatch N0(x)
weadd
2
π[−ln2+γ]J0(x),
whereγistheusualEuler–Mascheroniconstant(Section5.2).20Ournew,modifiedsecond
solutionis
y2(x)=2
π[lnx−ln2+γ]J0(x)+2
πJ0(x)braceleftbiggx2
4+5x4
128+···bracerightbigg
. (9.142e)
Now the comparison with N0(x)becomes a simple multiplication of J0(x)from
Eq. (9.142a) and the curly bracket of Eq. (9.142c). The multiplication checks, through
terms of order x2andx4, which is all we carried. Our second solution from Eqs. (9.127)
and(9.135) agreeswiththestandardsecondsolution,theNeumannfunction, N0(x).
From the preceding analysis, the second solution of Eq. (9.118), y2(x), may be written
as
y2(x)=y1(x)lnx+∞summationdisplay
j=−ndjxj+α, (9.142f)
the first solution times ln xand another power series, this one starting with xα−n, which
means that we may look for a logarithmic term when the indicial equation of Sec-
tion 9.5 gives only one series solution. With the form of the second solution specified
by Eq. (9.142f), we can substitute Eq. (9.142f) into the original differential equation and
determine the coefficients djexactly as in Section 9.5. It may be worth noting that no se-
ries expansion of ln xis needed. In the substitution, ln xwill drop out; its derivatives will
survive. /squaresolid
The second solution will usually diverge at the origin because of the logarithmic factor
and the negative powers of xin the series. For this reason y2(x)is often referred to as
theirregular solution . The first series solution, y1(x), which usually converges at the
origin,iscalledthe regularsolution .Thequestionofbehaviorattheoriginisdiscussedin
more detail in Chapters 11 and 12, in which we take up Bessel functions, modified Bessel
functions,andLegendrefunctions.
Summary
The two solutions of both sections (together with the exercises) provide a complete solu-
tionof our linear, homogeneous, second-order ODE—assuming that the point of expan-
sion is no worse than a regular singularity. At least one solution can always be obtained
byseriessubstitution(Section9.5).A second,linearlyindependentsolution canbecon-
structed by the Wronskian double integral, Eq. (9.127). This is all there are: No third,
linearlyindependentsolutionexists (compareExercise9.6.10).
Thenonhomogeneous , linear, second-order ODE will have an additional solution: the
particular solution . This particular solution may be obtained by the method of variation
ofparameters,Exercise9.6.25,orbytechniquessuchasGreen’sfunction,Section9.7.
20TheNeumann function N0is definedas it is in order to achieveconvenient asymptotic properties, Sections 11.3 and11.6.
588 Chapter 9 Differential Equations
Exercises
9.6.1 Youknowthatthethreeunitvectors ˆx,ˆy,andˆzaremutuallyperpendicular(orthogonal).
Showthatˆx,ˆy,andˆzarelinearlyindependent.Specifically,showthatnorelationofthe
form of Eq.(9.111) existsfor ˆx,ˆy,andˆz.
9.6.2 Thecriterionforthelinear independence ofthreevectors A,B,andCisthattheequa-
tion
aA+bB+cC=0
(analogous to Eq. (9.111)) has no solution other than the trivial a=b=c=0. Using
components A=(A1,A2,A3), and so on, set up the determinant criterion for the exis-
tenceor nonexistenceof a nontrivialsolutionfor the coefficients a,b, andc. Show that
yourcriterionis equivalenttothetriplescalarproduct A·B×C/negationslash=0.
9.6.3 UsingtheWronskiandeterminant,showthatthesetof functions
braceleftbigg
1,xn
n!(n=1,2,...,N)bracerightbigg
islinearlyindependent.
9.6.4 If the Wronskian of two functions y1andy2is identicallyzero, show by direct integra-
tionthat
y1=cy2,
thatis, that y1andy2are dependent.Assumethefunctionshavecontinuousderivatives
andthatatleastoneofthefunctionsdoesnotvanishintheintervalunderconsideration.
9.6.5 TheWronskianoftwofunctionsisfoundtobezeroat x0−ε≤x≤x0+εforarbitrarily
smallε>0.Show that this Wronskian vanishes for all xand that the functions are
linearlydependent.
9.6.6 Thethreefunctions sin x,ex,ande−xarelinearlyindependent.Noonefunctioncanbe
written as a linear combination of the other two. Show that the Wronskian of sin x,ex,
ande−xvanishesbutonlyatisolatedpoints.
ANS.W=4sinx,
W=0f orx=±nπ, n=0,1,2,....
9.6.7 Consider two functions ϕ1=xandϕ2=|x|=xsgnx(Fig. 9.3). The function sgn xis
the sign of x. Sinceϕ′
1=1 andϕ′
2=sgnx,W(ϕ1,ϕ2)=0 for any interval, including
[−1,+1].DoesthevanishingoftheWronskianover [−1,+1]provethat ϕ1andϕ2are
linearlydependent?Clearly,theyare not.Whatiswrong?
9.6.8 Explainthat linearindependence doesnotmeantheabsenceofanydependence.Illus-
trateyourargumentwith cosh xandex.
9.6.9 Legendre’sdifferentialequation
parenleftbig
1−x2parenrightbig
y′′−2xy′+n(n+1)y=0
9.6 A Second Solution 589
FIGURE 9.3xand|x|.
has a regularsolution Pn(x)andan irregular solution Qn(x). Showthat theWronskian
ofPnandQnis givenby
Pn(x)Q′
n(x)−P′
n(x)Qn(x)=An
1−x2,
withAnindependent ofx.
9.6.10 Show, by means of the Wronskian, that a linear, second-order, homogeneous ODE of
theform
y′′(x)+P(x)y′(x)+Q(x)y(x)=0
cannot have three independent solutions . (Assume a third solution and show that the
Wronskianvanishesforall x.)
9.6.11 Transformourlinear,second-orderODE
y′′+P(x)y′+Q(x)y=0
bythesubstitution
y=zexpbracketleftbigg
−1
2integraldisplayx
P(t)dtbracketrightbigg
andshowthattheresultingdifferentialequationfor zis
z′′+q(x)z=0,
where
q(x)=Q(x)−1
2P′(x)−1
4P2(x).
Note.This substitutioncanbederivedbythetechniqueofExercise9.6.24.
9.6.12 UsetheresultofExercise9.6.11toshowthatthereplacementof ϕ(r)byrϕ(r)maybe
expected to eliminate the first derivative from the Laplacian in spherical polar coordi-
nates.SeealsoExercise2.5.18(b).
590 Chapter 9 Differential Equations
9.6.13 Bydirectdifferentiationandsubstitutionshowthat
y2(x)=y1(x)integraldisplayxexp[−integraltextsP(t)dt]
[y1(s)]2ds
satisfies(like y1(x))theODE
y′′
2(x)+P(x)y′
2(x)+Q(x)y2(x)=0.
Note.TheLeibnizformulafor thederivativeof anintegralis
d
dαintegraldisplayh(α)
g(α)f(x,α)dx=integraldisplayh(α)
g(α)∂f(x,α)
∂αdx+fbracketleftbig
h(α),αbracketrightbigdh(α)
dα−fbracketleftbig
g(α),αbracketrightbigdg(α)
dα.
9.6.14 Intheequation
y2(x)=y1(x)integraldisplayxexp[−integraltextsP(t)dt]
[y1(s)]2ds
y1(x)satisfies
y′′
1+P(x)y′
1+Q(x)y1=0.
The function y2(x)is a linearly independent second solution of the same equation.
Show that the inclusion of lower limits on the two integrals leads to nothing new, that
is,thatitgeneratesonlyanoverallconstantfactorandaconstantmultipleoftheknown
solutiony1(x).
9.6.15 Giventhatonesolutionof
R′′+1
rR′−m2
r2R=0
isR=rm,showthatEq. (9.127)predictsasecondsolution, R=r−m.
9.6.16 Usingy1(x)=summationtext∞
n=0(−1)nx2n+1/(2n+1)!as a solution of the linear oscillator equa-
tion, follow the analysis culminating in Eq. (9.142f) and show that c1=0 so that the
secondsolutiondoesnot,inthis case,containalogarithmicterm.
9.6.17 Show that when nisnotan integer in Bessel’s ODE, Eq. (9.100), the second solution
ofBessel’s equation,obtainedfrom Eq.(9.127), does notcontainalogarithmicterm.
9.6.18 (a) Onesolutionof Hermite’sdifferentialequation
y′′−2xy′+2αy=0
forα=0isy1(x)=1.Findasecondsolution, y2(x),usingEq.(9.127).Showthat
yoursecondsolutionis equivalentto yodd(Exercise 9.5.6).
(b) Find a second solution for α=1, where y1(x)=x, using Eq. (9.127). Show that
yoursecondsolutionis equivalentto yeven(Exercise 9.5.6).
9.6.19 OnesolutionofLaguerre’sdifferentialequation
xy′′+(1−x)y′+ny=0
forn=0i sy1(x)=1. Using Eq. (9.127), developa second, linearly independentsolu-
tion.Exhibitthelogarithmictermexplicitly.
9.6 A Second Solution 591
9.6.20 ForLaguerre’sequationwith n=0,
y2(x)=integraldisplayxes
sds.
(a) Write y2(x)as alogarithmplusapowerseries.
(b) Verifythattheintegralformof y2(x),previouslygiven,isasolutionofLaguerre’s
equation (n=0)by direct differentiation of the integral and substitution into the
differentialequation.
(c) Verify that the series form of y2(x), part (a), is a solution by differentiating the
seriesandsubstitutingbackintoLaguerre’sequation.
9.6.21 OnesolutionoftheChebyshevequation
parenleftbig
1−x2parenrightbig
y′′−xy′+n2y=0
forn=0i sy1=1.
(a) UsingEq. (9.127), developasecond,linearlyindependentsolution.
(b) FindasecondsolutionbydirectintegrationoftheChebyshevequation.
Hint.L e tv=y′and integrate. Compare your result with the second solution given in
Section13.3.
ANS.(a) y2=sin−1x.
(b)Thesecondsolution, Vn(x), is notdefinedfor n=0.
9.6.22 OnesolutionoftheChebyshevequation
parenleftbig
1−x2parenrightbig
y′′−xy′+n2y=0
forn=1i sy1(x)=x. Set up the Wronskian double integral solution and derive a
secondsolution, y2(x).
ANS.y2=−parenleftbig
1−x2parenrightbig1/2.
9.6.23 TheradialSchrödingerwaveequationhastheform
braceleftbigg
−¯h2
2md2
dr2+l(l+1)¯h2
2mr2+V(r)bracketrightbigg
y(r)=Ey(r).
Thepotentialenergy V(r)maybeexpandedabouttheoriginas
V(r)=b−1
r+b0+b1r+···.
(a) Showthatthereis one(regular)solutionstartingwith rl+1.
(b) FromEq. (9.128)showthattheirregularsolutiondivergesattheoriginas r−l.
9.6.24 Show that if a second solution, y2,i sa s s u m e dt oh a v et h ef o r m y2(x)=y1(x)f(x),
substitutionbackintotheoriginalequation
y′′
2+P(x)y′
2+Q(x)y2=0
592 Chapter 9 Differential Equations
leadsto
f(x)=integraldisplayxexp[−integraltextsP(t)dt]
[y1(s)]2ds,
inagreementwithEq. (9.127).
9.6.25 If our linear, second-order ODE is nonhomogeneous, that is, of the form of Eq. (9.82),
themostgeneralsolutionis
y(x)=y1(x)+y2(x)+yp(x).
(y1andy2areindependentsolutionsofthehomogeneousequation.)
Showthat
yp(x)=y2(x)integraldisplayxy1(s)F(s)ds
W{y1(s),y2(s)}−y1(x)integraldisplayxy2(s)F(s)ds
W{y1(s),y2(s)},
withW{y1(x),y2(x)}theWronskianof y1(s)andy2(s).
Hint.AsinExercise9.6.24,let yp(x)=y1(x)v(x)anddevelopafirst-orderdifferential
equationfor v′(x).
9.6.26 (a) Showthat
y′′+1−α2
4x2y=0
hastwosolutions:
y1(x)=a0x(1+α)/2,
y2(x)=a0x(1−α)/2.
(b) Forα=0 the two linearly independentsolutions of part (a) reduceto y10=a0x1/2.
UsingEq.(9.128)deriveasecondsolution,
y20(x)=a0x1/2lnx.
Verifythat y20is indeedasolution.
(c)Showthatthesecondsolutionfrompart(b)maybeobtainedasalimitingcasefrom
thetwosolutionsof part(a):
y20(x)=lim
α→0parenleftbiggy1−y2
αparenrightbigg
.
9.7 N ONHOMOGENEOUS EQUATION —G REEN ’SFUNCTION
The series substitution of Section 9.5 and the Wronskian double integral of Section 9.6
provide the most general solution of the homogeneous , linear, second-order ODE. The
specific solution, yp, linearly dependent on the source term ( F(x)of Eq. (9.82)) may be
crankedoutbythevariationofparametersmethod,Exercise9.6.25.Inthissectionweturn
toadifferentmethodofsolution—Green’sfunction.
For a brief introduction to Green’s function method, as applied to the solution of a non-
homogeneous PDE, it is helpful to use the electrostatic analog. In the presence of charges
9.7 Nonhomogeneous Equation — Green’s Function 593
the electrostatic potential ψsatisfies Poisson’s nonhomogeneous equation (compare Sec-
tion1.14),
∇2ψ=−ρ
ε0(mksunits) , (9.143)
andLaplace’shomogeneousequation,
∇2ψ=0, (9.144)
intheabsenceofelectriccharge (ρ=0).Ifthechargesarepointcharges qi,weknowthat
thesolutionis
ψ=1
4πε0summationdisplay
iqi
ri, (9.145)
asuperpositionofsingle-pointchargesolutionsobtainedfromCoulomb’slawfortheforce
betweentwopointcharges q1andq2,
F=q1q2ˆr
4πε0r2. (9.146)
Byreplacementofthediscretepointchargeswithasmeared-outdistributedcharge,charge
densityρ,Eq.(9.145) becomes
ψ(r=0)=1
4πε0integraldisplayρ(r)
rdτ (9.147)
or,for thepotentialat r=r1awayfromtheoriginandthechargeat r=r2,
ψ(r1)=1
4πε0integraldisplayρ(r2)
|r1−r2|dτ2. (9.148)
We useψas the potential corresponding to the given distribution of charge and there-
fore satisfying Poisson’s equation (9.143), whereas a function G, which we label Green’s
function, is required to satisfy Poisson’s equation with a point source at the point defined
byr2:
∇2G=−δ(r1−r2). (9.149)
Physically, then, Gis the potential at r1corresponding to a unit source at r2. By Green’s
theorem(Section1.11,Eq. (11.104))
integraldisplayparenleftbig
ψ∇2G−G∇2ψparenrightbig
dτ2=integraldisplay
(ψ∇G−G∇ψ)·dσ. (9.150)
Assuming that the integrand falls off faster than r−2we may simplify our problem by
takingthevolumeso largethatthesurfaceintegralvanishes,leaving
integraldisplay
ψ∇2Gdτ2=integraldisplay
G∇2ψdτ2, (9.151)
or,bysubstitutinginEqs. (9.143)and(9.149), wehave
−integraldisplay
ψ(r2)δ(r1−r2)dτ2=−integraldisplayG(r1,r2)ρ(r2)
ε0dτ2. (9.152)
594 Chapter 9 Differential Equations
IntegrationbyemployingthedefiningpropertyoftheDiracdeltafunction(Eq.(1.171b))
produces
ψ(r1)=1
ε0integraldisplay
G(r1,r2)ρ(r2)dτ2. (9.153)
Note that we have used Eq. (9.149) to eliminate ∇2Gbut that the function Gitself is
stillunknown.InSection1.14,Gausslaw,wefoundthat
integraldisplay
∇2parenleftbigg1
rparenrightbigg
dτ=braceleftbigg0,
−4π,(9.154)
0 if thevolumedidnot includethe originand −4πif theorigin wereincluded.This result
from Section1.14mayberewrittenasinEq. (1.170),or
∇2parenleftbigg1
4πrparenrightbigg
=−δ(r),or ∇2parenleftbigg1
4πr12parenrightbigg
=−δ(r1−r2), (9.155)
corresponding to a shift of the electrostatic charge from the origin to the position r=r2.
Herer12=|r1−r2|, and the Dirac delta function δ(r1−r2)vanishes unless r1=r2.
Therefore in a comparison of Eqs. (9.149) and (9.155) the function G(Green’s function)
isgivenby
G(r1,r2)=1
4π|r1−r2|. (9.156)
Thesolutionofour differentialequation(Poisson’sequation)is
ψ(r1)=1
4πε0integraldisplayρ(r2)
|r1−r2|dτ2, (9.157)
in complete agreement with Eq. (9.148). Actually ψ(r1), Eq. (9.157), is the particular
solution of Poisson’s equation. We may add solutions of Laplace’s equation (compare
Eq.(9.83)). Suchsolutionscoulddescribeanexternalfield.
These results will be generalized to the second-order, linear, but nonhomogeneous dif-
ferentialequation
Ly(r1)=−f(r1), (9.158)
where Lis alineardifferentialoperator.TheGreen’sfunctionistakentobeasolutionof
LG(r1,r2)=−δ(r1−r2), (9.159)
analogous to Eq. (9.149). The Green’s function depends on boundary conditions that may
nolongerbethoseofelectrostaticsinaregionofinfiniteextent.Thentheparticularsolution
y(r1)becomes
y(r1)=integraldisplay
G(r1,r2)f(r2)dτ2. (9.160)
9.7 Nonhomogeneous Equation — Green’s Function 595
(Theremayalsobeanintegraloveraboundingsurface,dependingontheconditionsspec-
ified.)
In summary, Green’s function, often written G(r1,r2)as a reminder of the name, is a
solution of Eq. (9.149)or Eq.(9.159)more generally. It enters in an integral solution
of our differential equation, as in Eqs. (9.148)and(9.153). For the simple, but impor-
tant, electrostatic case we obtain Green’s function, G(r1,r2), by Gauss’ law, comparing
Eqs.(9.149)and(9.155). Finally, from the final solution (Eq.(9.157))it is possible to
develop a physical interpretation of Green’s function. It occurs as a weighting function or
propagator function that enhances or reduces the effect of the charge element ρ(r2)dτ2
according to its distance from the field point r1. Green’s function, G(r1,r2), gives the ef-
fectofaunitpointsourceat r2inproducingapotentialat r1.Thisishowitwasintroduced
inEq.(9.149);this ishowitappearsinEq. (9.157).
Symmetry of Green’s Function
AnimportantpropertyofGreen’sfunctionis thesymmetryof itstwovariables;thatis,
G(r1,r2)=G(r2,r1). (9.161)
Although this is obvious in the electrostatic case just considered, it can be proved under
moregeneralconditions.Inplaceof Eq.(9.149), letusrequirethat G(r,r1)satisfy21
∇·bracketleftbig
p(r)∇G(r,r1)bracketrightbig
+λq(r)G(r,r1)=−δ(r−r1), (9.162)
corresponding to a mathematical point source at r=r1. Here the functions p(r)andq(r)
are well-behaved but otherwise arbitrary functions of r. The Green’s function, G(r,r2),
satisfies the same equation, but the subscript 1 is replaced by subscript 2. The Green’s
functions, G(r,r1)andG(r,r2), have the same values over a given surface Sof some
volume of finite or infinite extent, and their normal derivatives have the same values over
thesurface S,or theseGreen’sfunctionsvanishon S(Dirichletboundaryconditions,Sec-
tion9.1).22ThenG(r,r2)is asortof potentialat r, createdbyaunitpointsourceat r2.
We multiply the equation for G(r,r1)byG(r,r2)and the equation for G(r,r2)by
G(r,r1)andthensubtractthetwo:
G(r,r2)∇·bracketleftbig
p(r)∇G(r,r1)bracketrightbig
−G(r,r1)∇·bracketleftbig
p(r)∇G(r,r2)bracketrightbig
=−G(r,r2)δ(r−r1)+G(r,r1)δ(r−r2). (9.163)
ThefirstterminEq. (9.163),
G(r,r2)∇·bracketleftbig
p(r)∇G(r,r1)bracketrightbig
,
maybereplacedby
∇·bracketleftbig
G(r,r2)p(r)∇G(r,r1)bracketrightbig
−∇G(r,r2)·p(r)∇G(r,r1).
21Equation (9.162) is a three-dimensional, inhomogeneous version of the self-adjoint eigenvalue equation, Eq. (10.8).
22Any attempt to demand that the normal derivatives vanish at the surface (Neumann’s conditions, Section 9.1) leads to trouble
with Gauss’ law. It is like demanding thatintegraltext
E·dσ=0 when you know perfectly well that there is some electric charge inside
the surface.
596 Chapter 9 Differential Equations
A similar transformation is carried out on the second term. Then integrating over the vol-
umewhosesurfaceis SandusingGreen’stheorem,weobtaina surfaceintegral:
integraldisplay
Sbracketleftbig
G(r,r2)p(r)∇G(r,r1)−G(r,r1)p(r)∇G(r,r2)bracketrightbig
·dσ
=−G(r1,r2)+G(r2,r1). (9.164)
The terms on the right-hand side appear when we use the Dirac delta functions in
Eq. (9.163) and carry out the volume integration. With the boundary conditions earlier
imposedontheGreen’sfunction,thesurface integralvanishesand
G(r1,r2)=G(r2,r1), (9.165)
whichshowsthatGreen’sfunctionissymmetric.Iftheeigenfunctionsarecomplex,bound-
ary conditions corresponding to Eqs. (10.19) to (10.20) are appropriate. Equation (9.165)
becomes
G(r1,r2)=G∗(r2,r1). (9.166)
Note that this symmetry property holds for Green’s functions in every equation in the
form of Eq. (9.162). In Chapter 10 we shall call equations in this form self-adjoint .T h e
symmetry is the basis of various reciprocity theorems; the effect of a charge at r2on the
potentialat r1is thesameas theeffectofachargeat r1onthepotentialat r2.
This use of Green’s functions is a powerful technique for solving many of the more
difficultproblemsof mathematicalphysics.
Form of Green’s Functions
Letusassumethat Lisaself-adjointdifferentialoperatorofthegeneralform23
L1=∇1·bracketleftbig
p(r1)∇1bracketrightbig
+q(r1). (9.167)
Herethe subscript 1 on Lemphasizesthat Loperateson r1. Then, as a simplegeneraliza-
tionof Green’stheorem,Eq. (1.104),wehave
integraldisplay
(vL2u−uL2v)dτ2=integraldisplay
p(v∇2u−u∇2v)·dσ2, (9.168)
inwhichallquantitieshave r2astheirargument.(ToverifyEq.(9.168),takethedivergence
of the integrand of the surface integral.) We let u(r2)=y(r2)so that Eq. (9.158) applies
andv(r2)=G(r1,r2)so that Eq. (9.159) applies. (Remember, G(r1,r2)=G(r2,r1).)
SubstitutingintoGreen’stheoremweget
integraldisplaybraceleftbig
−G(r1,r2)f(r2)+y(r2)δ(r1−r2)bracerightbig
dτ2
=integraldisplay
p(r2)braceleftbig
G(r1,r2)∇2y(r2)−y(r2)∇2G(r1,r2)bracerightbig
·dσ2.(9.169)
23L1may be in 1, 2, or3 dimensions (with appropriate interpretation of ∇1).
9.7 Nonhomogeneous Equation — Green’s Function 597
WhenweintegrateovertheDiracdeltafunction
y(r1)=integraldisplay
G(r1,r2)f(r2)dτ2
+integraldisplay
p(r2)braceleftbig
G(r1,r2)∇2y(r2)−y(r2)∇2G(r1,r2)bracerightbig
·dσ2,(9.170)
our solution to Eq. (9.158) appears as a volume integral plus a surface integral. If yand
Gbothsatisfy Dirichlet boundary conditions or if both satisfy Neumann boundary con-
ditions, the surface integral vanishes and we regain Eq. (9.160). The volume integral is a
weighted integral over the source term f(r2)with our Green’s function G(r1,r2)as the
weightingfunction.
Forthespecialcaseof p(r1)=1andq(r1)=0,Lis∇2,theLaplacian.Letusintegrate
∇2
1G(r1,r2)=−δ(r1−r2) (9.171)
overasmallvolumeincludingthepointsource.Then
integraldisplay
∇1·∇1G(r1,r2)dτ1=−integraldisplay
δ(r1−r2)dτ1=−1. (9.172)
The volumeintegralon the left maybe transformed by Gauss’ theorem,as in the develop-
mentofGauss’ law—Section1.14.Wefindthatintegraldisplay
∇1G(r1,r2)·dσ1=−1. (9.173)
This shows, incidentally, that it may not be possible to impose a Neumann boundary con-
dition, that the normal derivative of the Green’s function, ∂G/∂n, vanishes over the entire
surface.
If weareinthree-dimensionalspace,Eq. (9.173)is satisfiedbytaking
∂
∂r12G(r1,r2)=−1
4π·1
|r1−r2|2,r 12=|r1−r2|. (9.174)
Theintegrationisoverthesurfaceofaspherecenteredat r2.TheintegralofEq.(9.174)is
G(r1,r2)=1
4π·1
|r1−r2|, (9.175)
inagreementwithSection1.14.
If weareintwo-dimensionalspace,Eq. (9.173)is satisfiedbytaking
∂
∂ρ12G(ρ1,ρ2)=1
2π·1
|ρ1−ρ2|, (9.176)
withrbeingreplacedby ρ,ρ=(x2+y2)1/2,andtheintegrationbeingoverthecircumfer-
enceofa circlecenteredon ρ2.Her eρ12=|ρ1−ρ2|. IntegratingEq.(9.176), weobtain
G(ρ1,ρ2)=−1
2πln|ρ1−ρ2|. (9.177)
ToG(ρ1,ρ2)(and toG(r1,r2)) we may add any multiple of the regular solution of the
homogeneous(Laplace’s)equationasneededtosatisfy boundaryconditions.
598 Chapter 9 Differential Equations
Table 9.5 Green’sFunctionsa
Laplace Helmholtz Modified Helmholtz
∇2∇2+k2∇2−k2
One-dimensional space No solutioni
2kexp(ik|x1−x2|)1
2kexp(−k|x1−x2|)
for(−∞,∞)
Two-dimensional space −1
2πln|ρ1−ρ2|i
4H(1)
0(k|ρ1−ρ2|)1
2πK0(k|ρ1−ρ2|)
Three-dimensional space1
4π·1
|r1−r2|exp(ik|r1−r2|)
4π|r1−r2|exp(−k|r1−r2|)
4π|r1−r2|
aThese are the Green’s functions satisfying the boundary condition G(r1,r2)=0asr1→∞for the Laplace and modified
Helmholtz operators. For the Helmholtz operator, G(r1,r2)corresponds to an outgoing wave. H(1)
0is the Hankel function of
Section11.4. K0ishemodifiedBessel functionofSection11.5.
ThebehavioroftheLaplaceoperatorGreen’sfunctioninthevicinityofthesourcepoint
r1=r2shown by Eqs. (9.175) and (9.177) facilitates the identification of the Green’s
functionsfor theothercases, suchastheHelmholtzandmodifiedHelmholtzequations.
1. Forr1/negationslash=r2,G(r1,r2)mustsatisfy the homogeneous differentialequation
L1G(r1,r2)=0,r1/negationslash=r2. (9.178)
2. Asr1→r2(orρ1→ρ2),
G(ρ1,ρ2)≈−1
2πln|ρ1−ρ2|,two-dimensionalspace, (9.179)
G(r1,r2)≈1
4π·1
|r1−r2|,three-dimensionalspace. (9.180)
The term±k2in the operator does not affect the behavior of Gnear the singular point
r1=r2. For convenience, the Green’s functions for the Laplace, Helmholtz, and modified
HelmholtzoperatorsarelistedinTable9.5.
Spherical Polar Coordinate Expansion24
AsanalternatedeterminationoftheGreen’sfunctionoftheLaplaceoperator,letusassume
asphericalharmonicexpansionof theform
G(r1,r2)=∞summationdisplay
l=0lsummationdisplay
m=−lgl(r1,r2)Ym
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2), (9.181)
where the summation index lis the same for the spherical harmonics, as a consequence
of the symmetry of the Green’s function. We will now determine the radial functions
24This section is optional here and may be postponed to Chapter 12.
9.7 Nonhomogeneous Equation — Green’s Function 599
gl(r1,r2). FromExercises1.15.11and12.6.6,
δ(r1−r2)=1
r2
1δ(r1−r2)δ(cosθ1−cosθ2)δ(ϕ1−ϕ2)
=1
r2
1δ(r1−r2)∞summationdisplay
l=0lsummationdisplay
m=−lYm
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2). (9.182)
Substituting Eqs. (9.181) and (9.182) into the Green’s function differential equation,
Eq. (9.171), and making use of the orthogonality of the spherical harmonics, we obtain
aradialequation:
r1d2
dr2
1bracketleftbig
r1gl(r1,r2)bracketrightbig
−l(l+1)gl(r1,r2)=−δ(r1−r2). (9.183)
This is now a one-dimensional problem. The solutions25of the corresponding homoge-
neous equation are rl
1andr−l−1
1. If we demand that glremain finite as r1→0 and vanish
asr1→∞,thetechniqueof Section10.5leadsto
gl(r1,r2)=1
2l+1
rl
1
rl+1
2,r1<r2,
rl
2
rl+1
1,r1>r2,(9.184)
or
gl(r1,r2)=1
2l+1·rl
<
rl+1>. (9.185)
HenceourGreen’sfunctionis
G(r1,r2)=∞summationdisplay
l=0lsummationdisplay
m=−l1
2l+1rl
<
rl+1>Ym
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2). (9.186)
Sincewealreadyhave G(r1,r2)inclosedform, Eq. (9.175),wemaywrite
1
4π·1
|r1−r2|=∞summationdisplay
l=0lsummationdisplay
m=−l1
2l+1rl
<
rl+1>Ym
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2). (9.187)
One immediate use for this spherical harmonic expansion of the Green’s function is
in the development of an electrostatic multipole expansion. The potential for an arbitrary
chargedistributionis
ψ(r1)=1
4πε0integraldisplayρ(r2)
|r1−r2|dτ2
25Compare Table9.2.
600 Chapter 9 Differential Equations
(whichis Eq.(9.148)). SubstitutingEq.(9.187), weget
ψ(r1)=1
ε0∞summationdisplay
l=0lsummationdisplay
m=−lbraceleftbigg1
2l+1Ym
l(θ1,ϕ1)
rl+1
1
·integraldisplay
ρ(r2)Ym∗
l(θ2,ϕ2)rl
2dϕ2sinθ2dθ2r2
2dr2bracerightbigg
,forr1>r2.
Thisisthe multipoleexpansion .Therelativeimportanceofthevarioustermsinthedouble
sumdependsontheform ofthesource, ρ(r2).
Legendre Polynomial Addition Theorem26
Fromthegeneratingexpressionfor Legendrepolynomials,Eq. (12.4a),
1
4π·1
|r1−r2|=1
4π∞summationdisplay
l=0rl
<
rl+1>Pl(cosγ), (9.188)
whereγis the angle included between vectors r1andr2, Fig. 9.4. Equating Eqs. (9.187)
and(9.188), wehavetheLegendrepolynomialadditiontheorem:
Pl(cosγ)=4π
2l+1lsummationdisplay
m=−lYm
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2). (9.189)
FIGURE 9.4Sphericalpolarcoordinates.
26This section is optional here and may be postponed to Chapter 12.
9.7 Nonhomogeneous Equation — Green’s Function 601
It is instructive to compare this derivation with the relatively cumbersome derivation of
Section12.8leadingtoEq. (12.177).
Circular Cylindrical Coordinate Expansion27
Inanalogywiththeprecedingsphericalpolarcoordinateexpansion,wewrite
δ(r1−r2)=1
ρ1δ(ρ1−ρ2)δ(ϕ1−ϕ2)δ(z1−z2)
=1
ρ1δ(ρ1−ρ2)1
4π2∞summationdisplay
m=−∞eim(ϕ1−ϕ2)integraldisplay∞
−∞eik(z1−z2)dk,(9.190)
usingExercise12.6.5andEq.(1.193c)andtheCauchyprincipalvalue.Butwhyasumma-
tion for the ϕ-dependence and an integration for the z-dependence? The requirement that
the azimuthal dependence be single-valued quantizes m, hence the summation. No such
restrictionappliesto k.
Toavoidproblemslaterwithnegativevaluesof k,werewriteEq. (9.190)as
δ(r1−r2)=1
ρ1δ(ρ1−ρ2)1
2π∞summationdisplay
m=−∞eim(ϕ1−ϕ2)1
πintegraldisplay∞
0cosk(z1−z2)dk.(9.191)
Weassumeasimilarexpansionof theGreen’sfunction,
G(r1,r2)=1
2π2∞summationdisplay
m=−∞gm(ρ1,ρ2)eim(ϕ1−ϕ2)integraldisplay∞
0cosk(z1−z2)dk, (9.192)
with the ρ-dependent coefficients gm(ρ1,ρ2)to be determined. Substituting into
Eq.(9.171), nowincircularcylindricalcoordinates,wefindthatif g(ρ1,ρ2)satisfies
d
dρ1bracketleftbigg
ρ1dgm
dρ1bracketrightbigg
−bracketleftbigg
k2ρ1+m2
ρ1bracketrightbigg
gm=−δ(ρ1−ρ2), (9.193)
thenEq. (9.171)is satisfied.
The operator in Eq. (9.193) is identified as the modified Bessel operator (in self-
adjoint form). Hence the solutions of the corresponding homogeneous equation are u1=
Im(kρ),u2=Km(kρ). As in the spherical polar coordinate case, we demand that Gbe
finiteatρ1=0 andvanishas ρ1→∞.ThenthetechniqueofSection10.5yields
gm(ρ1,ρ2)=−1
AIm(kρ<)Km(kρ>). (9.194)
This corresponds to Eq. (9.155). The constant Acomes from the Wronskian (see
Eq.(9.120)):
Im(kρ)K′
m(kρ)−I′
m(kρ)Km(kρ)=A
P(kρ). (9.195)
27This section is optional here and may be postponed to Chapter 11.
602 Chapter 9 Differential Equations
FromExercise11.5.10, A=−1 and
gm(ρ1,ρ2)=Im(kρ<)Km(kρ>). (9.196)
ThereforeourcircularcylindricalcoordinateGreen’sfunctionis
G(r1,r2)=1
4π·1
|r1−r2|
=1
2π2∞summationdisplay
m=−∞integraldisplay∞
0Im(kρ<)Km(kρ>)eim(ϕ1−ϕ2)cosk(z1−z2)dk.
(9.197)
Exercise9.7.14is aspecialcaseof thisresult.
Example 9.7.1 QUANTUM MECHANICAL SCATTERING —N EUMANN SERIES SOLUTION
The quantum theory of scattering provides a nice illustration of integral equation tech-
niques and an application of a Green’s function. Our physical picture of scattering is as
follows. A beam of particles moves along the negative z-axis toward the origin. A small
fraction of the particles is scattered by the potential V(r)and goes off as an outgoing
spherical wave. Our wave function ψ(r)must satisfy the time-independent Schrödinger
equation
−¯h2
2m∇2ψ(r)+V(r)ψ(r)=Eψ(r), (9.198a)
or
∇2ψ(r)+k2ψ(r)=−bracketleftbigg
−2m
¯h2V(r)ψ(r)bracketrightbigg
,k2=2mE
¯h2. (9.198b)
From the physical picture just presented we look for a solution having an asymptotic
form
ψ(r)∼eik0·r+fk(θ,ϕ)eikr
r. (9.199)
Hereeik0·ris the incident plane wave28withk0the propagation vector carrying the sub-
script 0 to indicate that it is in the θ=0(z-axis) direction. The magnitudes k0andkare
equal (ignoring recoil), and eikr/ris the outgoing spherical wave with an angular (and
energy) dependent amplitude factor fk(θ,ϕ).29Vectorkhas the direction of the outgoing
scattered wave. In quantum mechanics texts it is shown that the differential probability of
scattering, dσ/d/Omega1,thescatteringcross sectionperunitsolidangle,is givenby |fk(θ,ϕ|2.
Identifying[−(2m/¯h2)V(r)ψ(r)]withf(r)of Eq.(9.158), wehave
ψ(r1)=−integraldisplay2m
¯h2V(r2)ψ(r2)G(r1,r2)d3r2 (9.200)
28Forsimplicityweassumeacontinuousincidentbeam.Inamoresophisticatedandmorerealistictreatment,Eq.(9.199)would
be one component of aFourier wavepacket.
29IfV(r)represents a centralforce, fkwill be afunction of θonly, independent of azimuth.
9.7 Nonhomogeneous Equation — Green’s Function 603
byEq.(9.170).ThisdoesnothavethedesiredasymptoticformofEq.(9.199),butwemay
add to Eq. (9.200) eik0·r1, a solution of the homogeneous equation, and put ψ(r)into the
desiredform:
ψ(r1)=eik0·r1−integraldisplay2m
¯h2V(r2)ψ(r2)G(r1,r2)d3r2. (9.201)
Our Green’s function is the Green’s function of the operator L=∇2+k2(Eq. (9.198)),
satisfyingtheboundaryconditionthatitdescribeanoutgoingwave.Then,fromTable9.5,
G(r1,r2)=exp(ik|r1−r2|)/(4π|r1−r2|)and
ψ(r1)=eik0·r1−integraldisplay2m
¯h2V(r2)ψ(r2)eik|r1−r2|
4π|r1−r2|d3r2. (9.202)
ThisintegralequationanalogoftheoriginalSchrödingerwaveequationis exact.Employ-
ing the Neumann series technique of Section 16.3 (remember, the scattering probability is
verysmall), wehave
ψ0(r1)=eik0·r1, (9.203a)
whichhasthephysicalinterpretationof noscattering.
Substituting ψ0(r2)=eik0·r2intotheintegral,weobtainthefirst correctionterm,
ψ1(r1)=eik0·r1−integraldisplay2m
¯h2V(r2)eik|r1−r2|
4π|r1−r2|eik0·r2d3r2.(9.203b)
Thisisthefamous Bornapproximation .Itisexpectedtobemostaccurateforweakpoten-
tials and high incident energy. If a more accurate approximation is desired, the Neumann
seriesmaybecontinued.30/squaresolid
Example 9.7.2 QUANTUM MECHANICAL SCATTERING —G REEN ’SFUNCTION
Again, we consider the Schrödinger wave equation (Eq. (9.198b)) for the scattering prob-
lem. This time we use Fourier transform techniques and derive the desired form of the
Green’s function by contour integration. Substituting the desired asymptotic form of the
solution(with kreplacedby k0),
ψ(r)∼eik0z+fk0(θ,ϕ)eik0r
r=eik0z+/Phi1(r), (9.204)
intotheSchrödingerwaveequation,Eq. (9.198b),yields
parenleftbig
∇2+k2
0parenrightbig
/Phi1(r)=U(r)eik0z+U(r)/Phi1(r). (9.205a)
Here
¯h2
2mU(r)=V(r),
30ThisassumestheNeumannseriesisconvergent.Insomephysicalsituationsitisnotconvergentandthenothertechniquesare
needed.
604 Chapter 9 Differential Equations
thescattering(perturbing)potential.Sincetheprobabilityofscatteringismuchlessthan1,
thesecondtermontheright-handsideofEq.(9.205a)isexpectedtobenegligible(relative
tothefirsttermontheright-handside)andthuswedropit.Notethatweare approximat-
ingour differentialequationwith
parenleftbig
∇2+k2
0parenrightbig
/Phi1(r)=U(r)eik0z. (9.205b)
We now proceed to solve Eq. (9.205b), a nonhomogeneous PDE. The differential oper-
ator∇2generatesa continuoussetof eigenfunctions
∇2ψk(r)=−k2ψk(r), (9.206)
where
ψk(r)=(2π)−3/2eik·r.
Theseplane-waveeigenfunctionsform acontinuousbutorthonormalset,inthesensethat
integraldisplay
ψ∗
k1(r)ψk2(r)d3r=δ(k1−k2)
(compareEq. (15.21d)).31Weusetheseeigenfunctionstoderivea Green’sfunction.
We expandtheunknownfunction /Phi1(r1)intheseeigenfunctions,
/Phi1(r1)=integraldisplay
Ak1ψk1(r1)d3k1, (9.207)
a Fourier integral with Ak1, the unknown coefficients. Substituting Eq. (9.207) into
Eq.(9.205b)andusingEq.(9.206), weobtain
integraldisplay
Akparenleftbig
k2
0−k2parenrightbig
ψk(r)d3k=U(r)eik0z. (9.208)
Using the now-familiar technique of multiplying by ψ∗
k2(r)and integrating over the space
coordinates,wehave
integraldisplay
Ak1parenleftbig
k2
0−k2
1parenrightbig
d3k1integraldisplay
ψ∗
k2(r)ψk1(r)d3r=Ak2parenleftbig
k2
0−k2
2parenrightbig
=integraldisplay
ψ∗
k2(r)U(r)eik0zd3r.(9.209)
Solvingfor Ak2andsubstitutingintoEq. (9.207)wehave
/Phi1(r2)=integraldisplaybracketleftbiggparenleftbig
k2
0−k2
2parenrightbig−1integraldisplay
ψ∗
k2(r1)U(r1)eik0z1d3r1bracketrightbigg
ψk2(r2)d3k2.(9.210)
Hence
/Phi1(r1)=integraldisplay
ψk1(r1)parenleftbig
k2
0−k2
1parenrightbig−1d3k1integraldisplay
ψ∗
k1(r2)U(r2)eik0z2d3r2, (9.211)
31d3r=dxdydz, a(three-dimensional) volume element in r-space.
9.7 Nonhomogeneous Equation — Green’s Function 605
replacing k2byk1andr1byr2to agree with Eq. (9.207). Reversing the order of integra-
tion,wehave
/Phi1(r1)=−integraldisplay
Gk0(r1,r2)U(r2)eik0z2d3r2, (9.212)
whereGk0(r1,r2), ourGreen’sfunction,is givenby
Gk0(r1,r2)=integraldisplayψ∗
k1(r2)ψk1(r1)
k2
1−k2
0d3k1, (9.213)
analogous to Eq. (10.90) of Section 10.5 for discrete eigenfunctions. Equation (9.212)
shouldbecomparedwiththeGreen’sfunctionsolutionofPoisson’sequation(9.157).
It is perhaps worth evaluatingthis integralto emphasizeoncemore the vitalrole played
bytheboundaryconditions.UsingtheeigenfunctionsfromEq. (9.206)and
d3k=k2dksinθdθdϕ,
weobtain
Gk0(r1,r2)=1
(2π)3integraldisplay∞
0integraldisplayπ
0integraldisplay2π
0eikρcosθ
k2−k2
0dϕsinθdθk2dk. (9.214)
Herekρcosθhas replaced k·(r1−r2), with ρ=r1−r2indicating the polar axis in k-
space. Integrating over ϕby inspection, we pick up a 2 π.T h eθ-integration then leads
to
Gk0(r1,r2)=1
4π2ρiintegraldisplay∞
0eikρ−e−ikρ
k2−k2
0kdk, (9.215)
andsincetheintegrandis anevenfunctionof k,wemayset
Gk0(r1,r2)=1
8π2ρiintegraldisplay∞
−∞(eiκ−e−iκ)
κ2−σ2κdκ. (9.216)
Thelatterstepistakeninanticipationoftheevaluationof Gk(r1,r2)asacontourintegral.
Thesymbols κandσ(σ>0)represent kρandk0ρ,respectively.
If the integral in Eq. (9.216) is interpreted as a Riemann integral, the integral does not
exist. This implies that L−1does not exist, and in a literal sense it does not. L=∇2+k2
is singular since there exist nontrivial solutions ψfor which the homogeneous equation
Lψ=0. We avoid this problem by introducing a parameter γ, defining a different opera-
torL−1
γ, andtakingthelimitas γ→0.
Splittingtheintegralintotwopartssothateachpartmaybewrittenasasuitablecontour
integralgivesus
G(r1,r2)=1
8π2ρicontintegraldisplay
C1κeiκdκ
κ2−σ2+1
8π2ρicontintegraldisplay
C2κe−iκdκ
κ2−σ2. (9.217)
ContourC1is closed by a semicircle in the upper half-plane, C2by a semicircle in
the lower half-plane. These integrals were evaluated in Chapter 7 by using appropriately
choseninfinitesimalsemicirclestogoaroundthesingularpoints κ=±σ.Asanalternative
procedure, let us first displace the singular points from the real axis by replacing σby
σ+iγandthen,afterevaluation,takingthelimitas γ→0 (Fig.9.5).
606 Chapter 9 Differential Equations
FIGURE 9.5PossibleGreen’s
functioncontoursof integration.
Forγpositive, contour C1encloses the singular point κ=σ+iγand the first integral
contributes
2πi·1
2ei(σ+iγ).
Fromthesecondintegralwealsoobtain
2πi·1
2ei(σ+iγ),
theenclosedsingularitybeing κ=−(σ+iγ).ReturningtoEq.(9.217)andletting γ→0,
wehave
G(r1,r2)=1
4πρeiσ=eik0|r1−r2|
4π|r1−r2|, (9.218)
infullagreementwithExercise9.7.16.Thisresultdependsonstartingwith γpositive.Had
wechosen γnegative,ourGreen’sfunctionwouldhaveincluded e−iσ,whichcorresponds
to anincoming wave. The choice of positive γis dictated by the boundary conditions we
wishtosatisfy.
Equations (9.212) and (9.218) reproduce the scattered wave in Eq. (9.203b) and consti-
tuteanexactsolutionof theapproximateEq. (9.205b).Exercises9.7.18and9.7.20extend
theseresults. /squaresolid
9.7 Nonhomogeneous Equation — Green’s Function 607
Exercises
9.7.1 VerifyEq. (9.168),
integraldisplay
(vL2u−uL2v)dτ2=integraldisplay
p(v∇2u−u∇2v)·dσ2.
9.7.2 Showthattheterms +k2intheHelmholtzoperatorand −k2inthemodifiedHelmholtz
operatordonotaffectthebehaviorof G(r1,r2)intheimmediatevicinityofthesingular
pointr1=r2. Specifically,showthat
lim
|r1−r2|→0integraldisplay
k2G(r1,r2)dτ2=1.
9.7.3 Showthat
exp(ik|r1−r2|)
4π|r1−r2|
satisfies the two appropriate criteria and therefore is a Green’s function for the
Helmholtzequation.
9.7.4 (a) Find the Green’s function for the three-dimensional Helmholtz equation, Exer-
cise9.7.3,whenthewaveisastandingwave.
(b) Howis thisGreen’sfunctionrelatedtothesphericalBessel functions?
9.7.5 ThehomogeneousHelmholtzequation
∇2ϕ+λ2ϕ=0
haseigenvalues λ2
iandeigenfunctions ϕi.ShowthatthecorrespondingGreen’sfunction
thatsatisfies
∇2G(r1,r2)+λ2G(r1,r2)=−δ(r1−r2)
maybewrittenas
G(r1,r2)=∞summationdisplay
i=1ϕi(r1)ϕi(r2)
λ2
i−λ2.
An expansion of this form is called a bilinearexpansion. If the Green’s function is
availablein closedform, thisprovidesameansofgeneratingfunctions.
9.7.6 Anelectrostaticpotential(mksunits)is
ϕ(r)=Z
4πε0·e−ar
r.
Reconstruct the electrical charge distribution that will produce this potential. Note that
ϕ(r)vanishesexponentiallyfor large r, showingthatthenetchargeis zero.
ANS.ρ(r)=Zδ(r)−Za2
4πe−ar
r.
608 Chapter 9 Differential Equations
9.7.7 TransformtheODE
d2y(r)
dr2−k2y(r)+V0e−r
ry(r)=0
andtheboundaryconditions y(0)=y(∞)=0 intoaFredholmintegralequationofthe
form
y(r)=λintegraldisplay∞
0G(r,t)e−t
ty(t)dt.
The quantities V0=λandk2are constants. The ODE is derived from the Schrödinger
waveequationwithamesonicpotential:
G(r,t)=
1
ke−ktsinhkr,0≤r<t,
1
ke−krsinhkt, t <r < ∞.
9.7.8 Achargedconductingringofradius a(Example12.3.3)maybedescribedby
ρ(r)=q
2πa2δ(r−a)δ(cosθ).
Using the known Green’s function for this system, Eq. (9.187) find the electrostatic
potential.
Hint.Exercise12.6.3willbehelpful.
9.7.9 Changingaseparationconstantfrom k2to−k2andputtingthediscontinuityofthefirst
derivativeintothe z-dependence,showthat
1
4π|r1−r2|=1
4π∞summationdisplay
m=−∞integraldisplay∞
0eim(ϕ1−ϕ2)Jm(kρ1)Jm(kρ2)e−k|z1−z2|dk.
Hint.Therequired δ(ρ1−ρ2)maybeobtainedfromExercise15.1.2.
9.7.10 Derivetheexpansion
exp[ik|r1−r2|]
4π|r1−r2|=ik∞summationdisplay
l=0
jl(kr1)h(1)
l(kr2), r 1<r2
jl(kr2)h(1)
l(kr1), r 1>r2
×lsummationdisplay
m=−lYm
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2).
Hint.TheleftsideisaknownGreen’sfunction.Assumeasphericalharmonicexpansion
andworkontheremainingradialdependence.Thesphericalharmonicclosurerelation,
Exercise12.6.6,coverstheangulardependence.
9.7.11 ShowthatthemodifiedHelmholtzoperatorGreen’sfunction
exp(−k|r1−r2|)
4π|r1−r2|
9.7 Nonhomogeneous Equation — Green’s Function 609
hasthesphericalpolarcoordinateexpansion
exp(−k|r1−r2|)
4π|r1−r2|=k∞summationdisplay
l=0il(kr<)kl(kr>)lsummationdisplay
m=−lYm
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2).
Note. The modified spherical Bessel functions il(kr)andkl(kr)are defined in Exer-
cise11.7.15.
9.7.12 From the spherical Green’s function of Exercise 9.7.10, derive the plane-wave expan-
sion
eik·r=∞summationdisplay
l=0il(2l+1)jl(kr)Pl(cosγ),
whereγis the angle included between kandr. This is the Rayleigh equation of Exer-
cise12.4.7.
Hint.T ak er2≫r1sothat
|r1−r2|→r2−r20·r1=r2−k·r1
k.
Letr2→∞andcancelafactor of eikr2/r2.
9.7.13 Fromtheresults ofExercises9.7.10and9.7.12,showthat
eix=∞summationdisplay
l=0il(2l+1)jl(x).
9.7.14 (a) FromthecircularcylindricalcoordinateexpansionoftheLaplaceGreen’sfunction
(Eq. (9.197)), showthat
1
(ρ2+z2)1/2=2
πintegraldisplay∞
0K0(kρ)coskzdk.
Thissameresultis obtaineddirectlyinExercise15.3.11.
(b) Asaspecialcaseof part(a) showthat
integraldisplay∞
0K0(k)dk=π
2.
9.7.15 Notingthat
ψk(r)=1
(2π)3/2eik·r
isaneigenfunctionof
parenleftbig
∇2+k2parenrightbig
ψk(r)=0
(Eq. (9.206)), showthattheGreen’sfunctionof L=∇2maybeexpandedas
1
4π|r1−r2|=1
(2π)3integraldisplay
eik·(r1−r2)d3k
k2.
610 Chapter 9 Differential Equations
9.7.16 Using Fourier transforms, show that the Green’s function satisfying the nonhomoge-
neousHelmholtzequation
parenleftbig
∇2+k2
0parenrightbig
G(r1,r2)=−δ(r1−r2)
is
G(r1,r2)=1
(2π)3integraldisplayeik·(r1−r2)
k2−k2
0d3k,
inagreementwithEq. (9.213).
9.7.17 Thebasicequationof thescalarKirchhoffdiffractiontheoryis
ψ(r1)=1
4πintegraldisplay
S2bracketleftbiggeikr
r∇ψ(r2)−ψ(r2)∇parenleftbiggeikr
rparenrightbiggbracketrightbigg
·dσ2,
whereψsatisfies the homogeneous Helmholtz equation and r=|r1−r2|. Derive this
equation.Assumethat r1is interiortotheclosedsurface S2.
Hint.UseGreen’stheorem.
9.7.18 The Born approximation for the scattered wave is given by Eq. (9.203b) (and
Eq. (9.211)). Fromtheasymptoticform,Eq. (9.199),
fk(θ,ϕ)eikr
r=−2m
¯h2integraldisplay
V(r2)eik|r−r2|
4π|r−r2|eik0·r2d3r2.
Forascatteringpotential V(r2)thatis independentof anglesandfor r≫r2, showthat
fk(θ,ϕ)=−2m
¯h2integraldisplay∞
0r2V(r2)sin(|k0−k|r2)
|k0−k|dr2.
Herek0is in theθ=0 (original z-axis) direction, whereas kis in the(θ,ϕ)direction.
Themagnitudesare equal: |k0|=|k|;misthereducedmass.
Hint. You have Exercise 9.7.12 to simplify the exponential and Exercise 15.3.20 to
transform the three-dimensional Fourier exponential transform into a one-dimensional
Fouriersinetransform.
9.7.19 Calculatethescatteringamplitude fk(θ,ϕ)foramesonicpotential V(r)=V0(e−αr/αr).
Hint.ThisparticularpotentialpermitstheBornintegral,Exercise9.7.18,tobeevaluated
asaLaplacetransform.
ANS.fk(θ,ϕ)=−2mV0
¯h2α1
α2+(k0−k)2.
9.7.20 Themesonicpotential V(r)=V0(e−αr/αr)maybeusedtodescribetheCoulombscat-
tering of two charges q1andq2.W el e tα→0 andV0→0 but take the ratio V0/αto
beq1q2/4πε0.(ForGaussianunitsomitthe4 πε0.)Showthatthedifferentialscattering
cross section dσ/d/Omega1=|fk(θ,ϕ)|2isgivenby
dσ
d/Omega1=parenleftbiggq1q2
4πε0parenrightbigg21
16E2sin4(θ/2),E=p2
2m=¯h2k2
2m.
Ithappens(coincidentally)thatthisBornapproximationisinexactagreementwithboth
theexactquantummechanicalcalculationsandtheclassicalRutherfordcalculation.
9.8 Heat Flow, or Diffusion, PDE 611
9.8 H EAT FLOW ,ORDIFFUSION ,P D E
Here we return to a special PDE to developfairly general methods to adapt a special solu-
tionofaPDEtoboundaryconditionsbyintroducingparametersthatapplytoothersecond-
order PDEs with constant coefficients as well. To some extent, they are complementary to
theearlierbasicseparationmethodfor findingsolutionsinasystematicway.
We select the full time-dependent diffusion PDE for an isotropic medium. Assuming
isotropy actually is not much of a restriction because, in case we have different (constant)
ratesofdiffusionindifferentdirections,forexampleinwood,ourheatflowPDEtakesthe
form
∂ψ
∂t=a2∂2ψ
∂x2+b2∂2ψ
∂y2+c2∂2ψ
∂z2, (9.219)
if we put the coordinate axes along the principal directions of anisotropy. Now we sim-
ply rescale the coordinates using the substitutions x=aξ,y=bη,z=cζto get back the
originalisotropicform ofEq. (9.219),
∂/Phi1
∂t=∂2/Phi1
∂ξ2+∂2/Phi1
∂η2+∂2/Phi1
∂ζ2(9.220)
forthetemperaturedistributionfunction /Phi1(ξ,η,ζ,t)=ψ(x,y,z,t) .
For simplicity, we first solve the time-dependent PDE for a homogeneous one-
dimensionalmedium,alongmetalrodinthe x-direction,say,
∂ψ
∂t=a2∂2ψ
∂x2, (9.221)
where the constant ameasures the diffusivity, or heat conductivity, of the medium. We
attempt to solve this linear PDE with constant coefficients with the relevant exponential
product Ansatz ψ=eαx·eβt, which, when substituted into Eq. (9.221), solves the PDE
withtheconstraint β=a2α2fortheparameters.Weseekexponentiallydecayingsolutions
for large times, that is, solutions with negative βvalues, and therefore set α=iω,α2=
−ω2for realωandhave
ψ(x,t)=eiωxe−ω2a2t=(cosωx+isinωx)e−ω2a2t.
Formingreal linearcombinationsweobtainthesolution
ψ(x,t)=(Acosωx+Bsinωx)e−ω2a2t,
foranychoiceof A,B,ω,whichareintroducedtosatisfyboundaryconditions.Uponsum-
ming over multiples nωof the basic frequency for periodic boundary conditions or inte-
grating over the parameter ωfor general (nonperiodic boundary conditions), we find a
solution,
ψ(x,t)=integraldisplaybracketleftbig
A(ω)cosωx+B(ω)sinωxbracketrightbig
e−a2ω2tdω, (9.222)
thatisgeneralenoughtobeadaptedtoboundaryconditionsat t=0,say.Whenthebound-
ary condition gives a nonzero temperature ψ0, as for our rod, then the summation method
612 Chapter 9 Differential Equations
applies (Fourier expansion of the boundary condition). If the space is unrestricted (as for
aninfinitelyextendedrod), theFourierintegralapplies.
•Thissummationorintegrationoverparametersisoneofthestandardmethodsforgen-
eralizingspecificPDEsolutionsinordertoadaptthemtoboundaryconditions.
Example 9.8.1 ASPECIFIC BOUNDARY CONDITION
Let us solve a one-dimensional case explicitly, where the temperature at time t=0i s
ψ0(x)=1=const. in the interval between x=+1 andx=−1 and zero for x>1 and
x<1.Attheends, x=±1,thetemperatureisalways heldatzero.
Forafiniteintervalwechoosethecos (lπx/2)spatialsolutionsofEq.(9.221)forinteger
l, becausetheyvanishat x=±1.Thus, at t=0 oursolutionis aFourierseries,
ψ(x,0)=∞summationdisplay
l=1alcosπlx
2=1,−1<x<1
withcoefficients(see Section14.1.)
al=integraldisplay1
−11·cosπlx
2=2
lπsinπlx
2vextendsinglevextendsinglevextendsinglevextendsingle1
x=−1
=4
πlsinlπ
2=4(−1)m
(2m+1)π,l=2m+1;
al=0,l=2m.
Includingitstimedependence,thefullsolutionisgivenbytheseries
ψ(x,t)=4
π∞summationdisplay
m=0(−1)m
2m+1cosbracketleftbigg
(2m+1)πx
2bracketrightbigg
e−t((2m+1)πa/2)2,(9.223)
which converges absolutely for t>0 but only conditionally at t=0, as a result of the
discontinuityat x=±1.
Without the restriction to zero temperature at the endpoints of the given finite interval,
the Fourier series is replaced by a Fourier integral. The general solution is then given by
Eq. (9.222). At t=0 the given temperature distribution ψ0=1 gives the coefficients as
(seeSection15.3)
A(ω)=1
πintegraldisplay1
−1cosωxdx=1
πsinωx
ωvextendsinglevextendsinglevextendsinglevextendsingle1
x=−1=2sinω
πω,B(ω)=0.
Therefore
ψ(x,t)=2
πintegraldisplay∞
0sinω
ωcos(ωx)e−a2ω2tdω. (9.224)
/squaresolid
Inthree dimensions the corresponding exponential Ansatz ψ=eik·r/a+βtleads to a
solution with the relation β=−k2=−k2for its parameter, and the three-dimensional
9.8 Heat Flow, or Diffusion, PDE 613
formofEq. (9.221)becomes
∂2ψ
∂x2+∂2ψ
∂y2+∂2ψ
∂z2+k2ψ=0, (9.225)
which is the Helmholtz equation, which may be solved by the separation method just
like the earlier Laplace equation in Cartesian, cylindrical, or spherical coordinates under
appropriatelygeneralizedboundaryconditions.
In Cartesian coordinates, with the product Ansatz of Eq. (9.35), the separated x- andy-
ODEsfromEq.(9.221)arethesameasEqs.(9.38)and(9.41),whilethe z-ODE,Eq.(9.42),
generalizesto
1
Zd2Z
dz2=−k2+l2+m2=n2>0, (9.226)
whereweintroduceanotherseparationconstant, n2, constrainedby
k2=l2+m2−n2(9.227)
to produce a symmetric set of equations. Now, our solution of Helmholtz’s Eq. (9.225)
is labeled according to the choice of all three separation constants l,m,nsubject to the
constraint Eq. (9.227). As before the z-ODE, Eq. (9.226), yields exponentially decaying
solutions∼e−nz. The boundary conditionat z=0 fixes the expansioncoefficients alm,a s
inEq.(9.44).
Incylindricalcoordinates,wenowusetheseparationconstant l2forthez-ODE,withan
exponentiallydecayingsolutioninmind,
d2Z
dz2=l2Z>0, (9.228)
soZ∼e−lz, because the temperature goes to zero at large z.I fw es e t k2+l2=n2,
Eqs.(9.53)to(9.54)staythesame,soweendupwiththesameFourier–Besselexpansion,
Eq.(9.56), as before.
Insphericalcoordinateswithradialboundaryconditions,theseparationmethodleadsto
thesameangularODEsinEqs. (9.61) and(9.64), andtheradialODEnowbecomes
1
r2d
drparenleftbigg
r2dR
drparenrightbigg
+k2R−QR
r2=0,Q=l(l+1), (9.229)
that is, of Eq. (9.65), whose solutions are the spherical Bessel functions of Section 11.7.
TheyarelistedinTable9.2.
Therestrictionthat k2beaconstantisunnecessarilysevere.Theseparationprocesswill
stillworkwithHelmholtz’sPDEfor k2asgeneralas
k2=f(r)+1
r2g(θ)+1
r2sin2θh(ϕ)+k′2. (9.230)
Inthehydrogenatomwehave k2=f(r)intheSchrödingerwaveequation,andthisleads
toaclosed-formsolutioninvolvingLaguerrepolynomials.
614 Chapter 9 Differential Equations
Alternate Solutions
In a new approach to the heat flow PDE suggested by experiments, we now return
to the one-dimensional PDE, Eq. (9.221), seeking solutions of a new functional form
ψ(x,t)=u(x/√t), which is suggested by Example 15.1.1. Substituting u(ξ),ξ=x/√t,
intoEq. (9.221)using
∂ψ
∂x=u′
√t,∂2ψ
∂x2=u′′
t,∂ψ
∂t=−x
2√
t3u′(9.231)
withthenotation u′(ξ)≡du
dξ, thePDEis reducedtotheODE
2a2u′′(ξ)+ξu′(ξ)=0. (9.232)
WritingthisODEas
u′′
u′=−ξ
2a2,
we can integrate it once to get ln u′=−ξ2
4a2+lnC1, with an integration constant C1.E x -
ponentiatingandintegratingagainwefindthesolution
u(ξ)=C1integraldisplayξ
0e−ξ2
4a2dξ+C2, (9.233)
involving two integration constants Ci. Normalizing this solution at time t=0 to temper-
ature+1f o rx>0 and−1f o rx<0, our boundary conditions, fixes the constants Ci,
so
ψ=1
a√πintegraldisplayx√t
0e−ξ2
4a2dξ=2√πintegraldisplayx
2a√t
0e−v2dv=/Phi1parenleftbiggx
2a√tparenrightbigg
,(9.234)
where/Phi1denotes Gauss’ error function (see Exercise 5.10.4). See Example 15.1.1 for a
derivation using a Fourier transform. We need to generalize this specific solution to adapt
ittoboundaryconditions.
To this end we now generate new solutions of the PDE with constant coefficients
by differentiating a special solution , Eq. (9.234). In other words, if ψ(x,t)solves the
PDE in Eq. (9.221), so do∂ψ
∂tand∂ψ
∂x, because these derivatives and the differentiations
of the PDE commute;that is, the order in whichthey are carried out does not matter. Note
carefully that this method no longer works if any coefficient of the PDE depends on t
orxexplicitly. However, PDEs with constant coefficients dominate in physics. Examples
are Newton’s equations of motion (ODEs) in classical mechanics, the wave equations of
electrodynamics,andPoisson’sandLaplace’sequationsinelectrostaticsandgravity.Even
Einstein’s nonlinear field equations of general relativity take on this special form in local
geodesiccoordinates.
Therefore, by differentiating Eq. (9.234) with respect to x, we find the simpler, more
basicsolution
ψ1(x,t)=1
a√tπe−x2
4a2t, (9.235)
9.8 Heat Flow, or Diffusion, PDE 615
and,repeatingtheprocess, anotherbasicsolution,
ψ2(x,t)=x
2a3√
t3πe−x2
4a2t. (9.236)
Again, these solutions have to be generalized to adapt them to boundary conditions. And
there is yet another method of generating new solutions of a PDE with constant coeffi-
cients:We can translate a givensolution, for example, ψ1(x,t)→ψ1(x−α,t), and then
integrateoverthetranslationparameter α.Therefore
ψ(x,t)=1
2a√tπintegraldisplay∞
−∞C(α)e−(x−α)2
4a2tdα (9.237)
isagainasolution,whichwerewriteusingthesubstitution
ξ=x−α
2a√t,α=x−2aξ√
t, dα=−2adξ√
t. (9.238)
Thus,wefindthat
ψ(x,t)=1√πintegraldisplay∞
−∞C(x−2aξ√
t)e−ξ2dξ (9.239)
isasolutionofourPDE.Inthisformwerecognizethesignificanceoftheweightfunction
C(x)from the translation method because, at t=0,ψ ( x ,0)=C(x)=ψ0(x)is deter-
mined by the boundary condition, andintegraltext∞
−∞e−ξ2dξ=√π. Therefore, we can also write
thesolutionas
ψ(x,t)=1√πintegraldisplay∞
−∞ψ0(x−2aξ√
t)e−ξ2dξ, (9.240)
displaying the role of the boundary condition explicitly. From Eq. (9.240) we see that
the initial temperature distribution, ψ0(x), spreads out over time and is damped by the
Gaussianweightfunction.
Example 9.8.2 SPECIAL BOUNDARY CONDITION AGAIN
Let us express the solution of Example 9.8.1 in terms of the error function solution of
Eq. (9.234). The boundary condition at t=0i sψ0(x)=1f o r−1<x<1 and zero
for|x|>1. From Eq. (9.240) we find the limits on the integration variable ξby setting
x−2aξ√t=±1. This yields the integration endpoints ξ=(±1+x)/2a√t. Therefore
oursolutionbecomes
ψ(x,t)=1√πintegraldisplayx+1
2a√t
x−1
2a√te−ξ2dξ.
Usingtheerror functiondefinedinEq.(9.234) wecanalsowritethissolutionasfollows
ψ(x,t)=1
2bracketleftbigg
erfparenleftbiggx+1
2a√tparenrightbigg
−erfparenleftbiggx−1
2a√tparenrightbiggbracketrightbigg
. (9.241)
616 Chapter 9 Differential Equations
Comparing this form of our solution with that from Example 9.8.1 we see that we can
express Eq. (9.241) as the Fourier integral of Example 9.8.1, an identity that gives the
Fourierintegral,Eq.(9.224), inclosedform ofthetabulatederror function. /squaresolid
Finally,weconsidertheheatflowcaseforanextended sphericallysymmetric medium
centered at the origin, which prescribes polar coordinates r,θ,ϕ.We expect a solution of
theformψ(r,t)=u(r,t). UsingEq.(2.48) wefindthePDE
∂u
∂t=a2parenleftbigg∂2u
∂r2+2
r∂u
∂rparenrightbigg
, (9.242)
whichwetransformtotheone-dimensionalheatflowPDEbythesubstitution
u=v(r,t)
r,∂u
∂r=1
r∂v
∂r−v
r2,∂u
∂t=1
r∂v
∂t,
∂2u
∂r2=1
r∂2v
∂r2−2
r2∂v
∂r+2v
r3. (9.243)
ThisyieldsthePDE
∂v
∂t=a2∂2v
∂r2. (9.244)
Example 9.8.3 SPHERICALLY SYMMETRIC HEATFLOW
Let us apply the one-dimensional heat flow PDE with the solution Eq. (9.234) to a spheri-
cally symmetric heat flow under fairly common boundary conditions, where xis released
by the radial variable. Initially we have zero temperature everywhere. Then, at time t=0,
afiniteamountofheatenergy Qisreleasedattheorigin,spreadingevenlyinalldirections.
Whatistheresultingspatialandtemporaltemperaturedistribution?
InspectingourspecialsolutioninEq. (9.236)weseethat,for t→0,thetemperature
v(r,t)
r=C√
t3e−r2
4a2t (9.245)
goestozeroforall r/negationslash=0,sozeroinitialtemperatureisguaranteed.As t→∞,thetemper-
aturev/r→0 for allrincluding the origin, which is implicit in our boundary conditions.
Theconstant Ccanbedeterminedfrom energyconservation,whichgivestheconstraint
Q=σρintegraldisplayv
rd3r=4πσρC√
t3integraldisplay∞
0r2e−r2
4a2tdr=8radicalbig
π3σρa3C, (9.246)
whereρis the constant density of the medium and σis its specific heat. Here we have
rescaledtheintegrationvariableandintegratedbyparts toget
integraldisplay∞
0e−r2
4a2tr2dr=(2a√
t)3integraldisplay∞
0e−ξ2ξ2dξ,
integraldisplay∞
0e−ξ2ξ2dξ=−ξ
2e−ξ2vextendsinglevextendsinglevextendsinglevextendsingle∞
0+1
2integraldisplay∞
0e−ξ2dξ=√π
4.
9.8 Heat Flow, or Diffusion, PDE 617
Thetemperature,asgivenbyEq.(9.245)atanymoment,whichisatfixed t,isaGaussian
distribution that flattens out as time increases, because its width is proportional to√t.A s
a function of time the temperature is proportional to t−3/2e−T/t, withT≡r2/4a2, which
rises from zero to a maximum and then falls off to zero again for large times. To find the
maximum,weset
d
dtparenleftbig
t−3/2e−T/tparenrightbig
=t−5/2e−T/tparenleftbiggT
t−3
2parenrightbigg
=0, (9.247)
fromwhichwefind t=2T/3. /squaresolid
In the case of cylindrical symmetry (in the plane z=0 in plane polar coordinates
ρ=radicalbig
x2+y2,ϕ)welookforatemperature ψ=u(ρ,t)thatthensatisfiestheODE(using
Eq.(2.35) inthediffusionequation)
∂u
∂t=a2parenleftbigg∂2u
∂ρ2+1
ρ∂u
∂ρparenrightbigg
, (9.248)
whichistheplanaranalogofEq.(9.244).ThisODEalsohassolutionswiththefunctional
dependence ρ/√t≡r. Uponsubstituting
u=vparenleftbiggρ√tparenrightbigg
,∂u
∂t=−ρv′
2t3/2,∂u
∂ρ=v′
√t,∂2u
∂ρ2=v′
t(9.249)
intoEq. (9.248)withthenotation v′≡dv
dr, wefindtheODE
a2v′′+parenleftbigga2
r+r
2parenrightbigg
v′=0. (9.250)
This is a first-order ODE for v′, which we can integrate when we separate the variables v
andras
v′′
v′=−parenleftbigg1
r+r
2a2parenrightbigg
. (9.251)
Thisyields
v(r)=C
re−r2
4a2=C√t
ρe−ρ2
4a2t. (9.252)
Thisspecialsolutionforcylindricalsymmetrycanbesimilarlygeneralizedandadaptedto
boundary conditions, as for the spherical case. Finally, the z-dependence can be factored
in,because zseparatesfrom theplanepolarradialvariable ρ.
Insummary,PDEscanbesolvedwithinitialconditions,justlikeODEs,orwithbound-
aryconditionsprescribingthevalueofthesolutionoritsderivativeonboundarysurfaces,
curves, or points. When the solution is prescribed on the boundary, the PDE is called a
Dirichlet problem;if the normal derivative of the solution is prescribed on the boundary,
thePDE iscalleda Neumann problem.
Whentheinitialtemperatureisprescribedfortheone-dimensionalorthree-dimensional
heatequation (withsphericalorcylindricalsymmetry )itbecomesaweightfunctionofthe
solution,intermsofanintegraloverthegenericGaussiansolution.Thethree-dimensional
618 Chapter 9 Differential Equations
heat equation, with spherical or cylindrical boundary conditions, is solved by separation
of the variables, leading to eigenfunctions in each separated variable and eigenvalues as
separationconstants.Forfiniteboundaryintervalsineachspatialcoordinate,thesumover
separationconstantsleadstoaFourier-seriessolution,whileinfiniteboundaryconditions
lead to a Fourier-integral solution. The separation of variables method attempts to solve
a PDE by writing the solution as a product of functions of one variable each. General
conditions for the separation method to work are provided by the symmetry properties of
thePDE, towhichcontinuousgrouptheoryapplies.
AdditionalReadings
Bateman, H., Partial Differential Equations of Mathematical Physics . New York: Dover (1944), 1st ed. (1932).
A wealth of applications of various partial differential equations in classical physics. Excellent examples of
the use of different coordinate systems—ellipsoidal, paraboloidal, toroidal coordinates, and so on.
Cohen, H., Mathematics for Scientists and Engineers .Englewood Cliffs, NJ: Prentice-Hall (1992).
Courant, R., and D. Hilbert, Methods of Mathematical Physics , Vol. 1 (English edition). New York: Interscience
(1953), Wiley (1989). This is one of the classic works of mathematical physics. Originally published in Ger-
manin1924,therevisedEnglisheditionisanexcellentreferenceforarigoroustreatmentofGreen’sfunctions
and for a widevariety of othertopics on mathematical physics.
Davis,P.J.,andP.Rabinowitz, Numerical Integration . Waltham,MA:Blaisdell (1967). Thisbook covers agreat
deal of material in a relatively easy-to-read form. Appendix 1 ( On the Practical Evaluation of Integrals by
M.Abramowitz) is excellentas anoverall view.
Garcia,A.L., Numerical Methods for Physics .Englewood Cliffs, NJ: Prentice-Hall(1994).
Hamming, R. W., Numerical Methods for Scientists and Engineers , 2nd ed. New York: McGraw-Hill (1973),
reprinted Dover (1987). This well-written text discusses a wide variety of numerical methods from zeros of
functionstothefastFouriertransform.Alltopicsareselectedanddevelopedwithamoderncomputerinmind.
Hubbard, J., andB. H.West, Differential Equations . Berlin: Springer (1995).
Ince,E.L., OrdinaryDifferentialEquations .NewYork:Dover(1956).Theclassicworkinthetheoryofordinary
differential equations.
Lapidus, L., and J. H. Seinfeld, Numerical Solutions of Ordinary Differential Equations . New York: Academic
Press(1971).Adetailedandcomprehensivediscussionofnumericaltechniques,withemphasisontheRunge–
Kuttaandpredictor–correctormethods.Recentworkontheimprovementofcharacteristicssuchasstabilityis
clearlypresented.
Margenau, H., and G. M. Murhpy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nos-
trand (1956). Chapter5 covers curvilinear coordinates and 13 specificcoordinate systems.
Miller,R. K.,andA.N.Michel, Ordinary DifferentialEquations . NewYork: AcademicPress (1982).
Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5
includes a description of several different coordinate systems. Note that Morse and Feshbach are not above
usingleft-handedcoordinatesystemsevenforCartesiancoordinates.Elsewhereinthisexcellent(anddifficult)
bookaremanyexamplesoftheuseofthevariouscoordinatesystemsinsolvingphysicalproblems.Chapter7
is a particularly detailed, complete discussion of Green’s functions from the point of view of mathematical
physics. Note, however, that Morse and Feshbach frequently choose a source of 4 πδ(r−r′)in place of our
δ(r−r′).Considerable attention is devoted to bounded regions.
Murphy,G.M., OrdinaryDifferentialEquationsandTheirSolutions .Princeton,NJ:VanNostrand(1960).Athor-
ough, relatively readabletreatment of ordinary differential equations, both linear and nonlinear.
Press, W. H., B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes , 2nd ed. Cambridge, UK:
Cambridge University Press (1992).
Ralston, A.,andH.Wilf, eds., Mathematical Methods for Digital Computers . NewYork: Wiley (1960).
9.8 Additional Readings 619
Ritger, P. D.,andN.J.Rose, Differential Equations with Applications . NewYork: McGraw-Hill(1968).
Stakgold, I., Green’s Functions and Boundary ValueProblems , 2nd ed.NewYork: Wiley (1997).
Stoer, J.,and R. Burlirsch, Introduction to Numerical Analysis . NewYork: Springer-Verlag (1992).
Stroud, A. H., Numerical Quadrature and Solution of Ordinary Differential Equations , Applied Mathematics
Series, Vol. 10. New York: Springer-Verlag (1974). A balanced, readable, and very helpful discussion of var-
ious methods of integrating differential equations. Stroud is familiar with the work in this field and provides
numerous references.
This page intentionally left blank
CHAPTER 10
STURM –LIOUVILLE
THEORY —O RTHOGONAL
FUNCTIONS
In the preceding chapter we developed two linearly independent solutions of the second-
order linear homogeneous differential equation and proved that no third, linearly inde-
pendent solution existed. In this chapter the emphasis shifts from solving the differential
equation to developing and understanding general properties of the solutions. There is a
close analogy between the concepts in this chapter and those of linear algebra in Chap-
ter 3. Functions here play the role of vectors there, and linear operators that of matri-
ces in Chapter 3. The diagonalization of a real symmetric matrix in Chapter 3 corre-
sponds here to the solution of an ODE defined by a self-adjoint operator Lin terms
of its eigenfunctions, which are the “continuous” analog of the eigenvectors in Chap-
ter 3. Examples for the corresponding analogy between Hermitian matrices and Her-
mitian operators are Hamiltonians in quantum mechanics and their energy eigenfunc-
tions.
InSection10.1theconceptsofself-adjointoperator,eigenfunction,eigenvalue,andHer-
mitian operator are presented. The concept of adjoint operator, given first in terms of dif-
ferential equations, is then redefined in accordance with usage in quantum mechanics,
where eigenfunctions take complex values. The vital properties of reality of eigenvalues
and orthogonality of eigenfunctions are derived in Section 10.2. In Section 10.3 we dis-
cuss the Gram–Schmidt procedure for systematically constructuring sets of orthogonal
functions. Finally, the general property of the completeness of a set of eigenfunctions is
explored in Section 10.4, and Green’s functions from Chapter 9 are continued in Sec-
tion10.5.
621
622 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.1 S ELF-ADJOINT ODE S
In Chapter 9 we studied, classified, and solved linear, second-order ODEs corresponding
tolinear,second-orderdifferentialoperatorsofthegeneralform
Lu(x)=p0(x)d2
dx2u(x)+p1(x)d
dxu(x)+p2(x)u(x). (10.1)
The coefficients p0(x),p1(x), andp2(x)are real functions of x, and over the region
of interest, a≤x≤b, the first 2−iderivatives of pi(x)are continuous. Reference to
Eq.(9.118)showsthat P(x)=p1(x)/p0(x)andQ(x)=p2(x)/p0(x).Hence,p0(x)must
not vanish for a<x<b . Now, the zeros of p0(x)are singular points (Section 9.4), and
the preceding statement means that our interval [a,b]must be given so that there are no
singularpointsintheinterioroftheinterval.Theremaybeandoftenaresingularpointson
theboundaries.
For a linear operator L, the analog of a quadratic form for a matrix in Chapter 3 is the
integral
/angbracketleftu|L|u/angbracketright≡/angbracketleftu|Lu/angbracketright≡integraldisplayb
au(x)Lu(x)dx
=integraldisplayb
au{p0u′′+p1u′+p2u}dx, (10.2)
wheretheprimesontherealfunction u(x)denotederivatives,asusual,and,forsimplicity,
u(x)is taken to be real. If we shift the derivatives to the first factor, u, in Eq. (10.2) by
integratingbypartsonceor twice,weareledtotheequivalentexpression,
/angbracketleftu|L|u/angbracketright=bracketleftbig
u(x)(p1−p′
0)u(x)bracketrightbigb
x=a
+integraldisplayb
abraceleftbiggd2
dx2[p0u]−d
dx[p1u]+p2ubracerightbigg
udx. (10.3)
If we require that the integrals in Eqs. (10.2) and (10.3) be identical for all (twice differ-
entiable)functions u, thentheintegrandshavetobeequal.Thecomparisonthenyields
u(p′′
0−p′
1)u+2u(p′
0−p1)u′=0,
or
p′
0(x)=p1(x), (10.4)
and,asabonus,thetermsattheboundaries x=aandx=binEq.(10.3)thenalsovanish.
BecauseoftheanalogywiththetransposedmatrixinChapter3,itisconvenienttodefine
thelinearoperatorinEq.(10.3),
¯Lu=d2
dx2[p0u]−d
dx[p1u]+p2u
=p0d2u
dx2+(2p′
0−p1)du
dx+(p′′
0−p′
1+p2)u, (10.5)
10.1 Self-Adjoint ODEs 623
astheadjoint1operator¯L.Wehavedefinedtheadjointoperator ¯Landhaveshownthatif
Eq.(10.4)issatisfied, /angbracketleft¯Lu|u/angbracketright=/angbracketleftu|Lu/angbracketright.Followingthesameprocedurewecanshowmore
generallythat/angbracketleftv|Lu/angbracketright=/angbracketleftLv|u/angbracketright.Whenthis conditionissatisfied,
¯Lu=Lu=d
dxbracketleftbigg
p(x)du(x)
dxbracketrightbigg
+q(x)u(x), (10.6)
the operator Lis said to be self-adjoint . Here, for the self-adjoint case, p0(x)is replaced
byp(x)andp2(x)byq(x)to avoid unnecessary subscripts. The form of Eq. (10.6) al-
lows carrying out two integrations by parts in Eq. (10.3) (and Eq. (10.22) and following)
withoutintegratedterms.2Notethatagivenoperatorisnotinherentlyself-adjoint;itsself-
adjointness depends on the properties of the function space in which it acts and on the
boundaryconditions.
In a survey of the ODEs introduced in Section 9.3, Legendre’s equation and the linear
oscillatorequationareself-adjoint,butothers,suchastheLaguerreandHermiteequations,
are not. However, the theory of linear, second-order, self-adjoint differential equations is
perfectly general because we can alwaystransform the non-self-adjoint operator into the
requiredself-adjointform.ConsiderEq. (10.1) with p′
0/negationslash=p1. If wemultiply Lby3
1
p0(x)expbracketleftbiggintegraldisplayxp1(t)
p0(t)dtbracketrightbigg
,
weobtain
1
p0(x)expbracketleftbiggintegraldisplayxp1(t)
p0(t)dtbracketrightbigg
Lu(x)=d
dxbraceleftbigg
expbracketleftbiggintegraldisplayxp1(t)
p0(t)dtbracketrightbiggdu(x)
dxbracerightbigg
+p2(x)
p0(x)·expbracketleftbiggintegraldisplayxp1(t)
p0(t)dtbracketrightbigg
u,(10.7)
which is clearly self-adjoint (see Eq. (10.6)). Notice the p0(x)in the denominator. This is
whywerequire p0(x)/negationslash=0,a<x<b .Inthefollowingdevelopmentweassumethat Lhas
beenputintoself-adjointform.
1Theadjointoperator bears a somewhat forced relationship to the adjointmatrix. A better justification for the nomenclature
is found in a comparison of the self-adjoint operator (plus appropriate boundary conditions) with the self-adjoint matrix. The
significant properties aredeveloped in Section10.2. Becauseofthese properties, weareinterested in self-adjoint operators.
2The full importance of the self-adjoint form (plus boundary conditions) will become apparent in Section 10.2. In addition,
self-adjoint forms will be required for developing Green’s functions in Section 10.5.
3If wemultiply Lbyf(x)/p0(x)andthendemandthat
f′(x)=fp1
p0,
so that the newoperator will be self-adjoint, weobtain
f(x)=expbracketleftbiggintegraldisplayxp1(t)
p0(t)dtbracketrightbigg
.
624 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Eigenfunctions, Eigenvalues
Schrödinger’swaveequation
Hψ(x)=Eψ(x)
is the major example of an eigenvalue equation in physics; here the differential operator
Lis defined by the Hamiltonian Hand may no longer be real, and the eigenvalue be-
comes the total energy Eof the system. The eigenfunction ψ(x)may be complex and is
usually called a wave function . A variational formulation of this Schrödinger equation
appears in Section 17.7. Based on spherical, cylindrical, or some other symmetry prop-
erties, a three- or four-dimensional PDE or eigenvalue equation such as the Schrödinger
equation may separate into eigenvalue equations in a single variable each. Examples are
Eqs. (9.41), (9.42), (9.50), and (9.53). However, sometimes an eigenvalue equation takes
themoregeneralself-adjointform
Lu(x)+λw(x)u(x)=0, (10.8)
where the constant λis the eigenvalue4andw(x)is a known weight or density func-
tion;w(x) >0 except possibly at isolated points at which w(x)=0. (In Section 10.1,
w(x)≡1.) For a given choice of the parameter λ, a function uλ(x), which satisfies
Eq.(10.8) andtheimposedboundaryconditions ,iscalledan eigenfunction correspond-
ingtoλ.Theconstant λisthencalledan eigenvalue bymathematicians.Thereisnoguar-
antee that an eigenfunction uλ(x)will exist for an arbitrary choice of the parameter λ.
Indeed,therequirementthattherebeaneigenfunctionoftenrestrictstheacceptablevalues
ofλto a discrete set. Examples of this for the Legendre, Hermite, and Chebyshev equa-
tions appear in the exercises of Section 9.5. Here we have the mathematical approach to
theprocess ofquantizationinquantummechanics.
The inner product of two functions, /angbracketleftv|u/angbracketright=integraltextb
av∗(x)w(x)u(x)dx, depends on the
weight function and generalizes our previous definition, where w(x)≡1.The weight
function also modifies the definition of orthogonality of two eigenfunctions: They are
orthogonal if their inner product /angbracketleftuλ′|uλ/angbracketright=0.The extra weight function w(x)appears
sometimes as an asymptotic wave function ψ∞that is a common factor in all solutions
of a PDE such as the Schrödinger equation, for example, when the potential V(x)→0a s
x→∞inH=T+V. We can find ψ∞when we set V=0 in the Schrödinger equa-
tion. Another source for w(x)may be a nonzero angular momentum barrier l(l+1)/x2
in a PDE or separated ODE Eq. (9.65) that has a regular singularity and dominates at
x→0. In such a case the indicial equation, such as Eq. (9.87) or (9.103), shows that
the wave function has xlas an overall factor. Since the wave function enters twice in
matrix elements and orthogonality relations, the weight functions in Table 10.1 come
from these common factors in both radial wave functions. This is how the exp (−x)for
Laguerre polynomials arises and xkexp(−x)for associated Laguerre polynomials in Ta-
ble10.1.
4Notethat this mathematicaldefinition of the eigenvalue differs by a sign from the usage in physics.
10.1 Self-Adjoint ODEs 625
Table 10.1
Equation p(x) q(x) λ w(x)
Legendrea1−x20 l(l+1) 1
Shifted Legendreax(1−x) 0 l(l+1) 1
AssociatedLegendrea1−x2−m2/(1−x2)l(l+1) 1
Chebyshev I (1−x2)1/20 n2(1−x2)−1/2
Shifted Chebyshev I [x(1−x)]1/20 n2[x(1−x)]−1/2
Chebyshev II (1−x2)3/20 n(n+2)(1−x2)1/2
Ultraspherical (Gegenbauer) (1−x2)α+1/20 n(n+2α) (1−x2)α−1/2
Besselb,0≤x≤ax −n2/x a2x
Laguerre, 0≤x<∞ xe−x0 αe−x
AssociatedLaguerrecxk+1e−x0 α−kxke−x
Hermite, 0≤x<∞ e−x202 αe−x2
Simple harmonic oscillatord10 n21
al=0,1,...,−l≤m≤lareintegersand −1≤x≤1,0≤x≤1forshiftedLegendre.
bOrthogonalityofBesselfunctionsisratherspecial.CompareSection11.2.fordetails.Asecondtypeoforthogonality
isdevelopedinEq.(11.174).
ckis anon-negativeinteger.Formore details,seeTable10.2.
dThiswillformthebasisforChapter14,Fourierseries.
Example 10.1.1 LEGENDRE ’SEQUATION
Legendre’sequationis givenby
parenleftbig
1−x2parenrightbig
u′′−2xu′+n(n+1)u=0,−1≤x≤1. (10.9)
FromEqs. (10.1), (10.8), and(10.9),
p0(x)=1−x2=p, w(x)=1,
p1(x)=−2x=p′,λ=n(n+1),
p2(x)=0=q.
RecallthatourseriessolutionsofLegendre’sequation(Exercise9.5.5)5divergedunless n
wasrestrictedtooneoftheintegers.This representsaquantizationoftheeigenvalue λ./squaresolid
When the equations of Chapter 9 are transformed into the self-adjoint form, we find
the following values of the coefficients and parameters (Table 10.1). The coefficient p(x)
is the coefficient of the second derivative of the eigenfunction. The eigenvalue λis the
parameterthatis availablein a termof theform λw(x)u(x) ;a n yxdependenceapartfrom
theeigenfunctionbecomestheweightingfunction w(x).Ifthereisanothertermcontaining
theeigenfunction(notthederivatives),thecoefficientoftheeigenfunctioninthisadditional
termisidentifiedas q(x). If nosuchtermispresent, q(x)is zero.
5Compare also Exercise5.2.15 and 12.10.
626 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Example 10.1.2 DEUTERON
Further insight into the concepts of eigenfunction and eigenvalue may be provided by an
extremely simple model of the deuteron, a bound state of a neutron and proton. From
experiment,thebindingenergyofabout2MeV ≪Mc2,withM=Mp=Mn,thecommon
neutron and proton mass whose small mass difference we neglect. Due to the short range
of the nuclear force, the deuteron properties do not depend much on the detailed shape of
the interactionpotential. Thus, the neutron–protonnuclear interaction may be modeled by
asphericallysymmetricsquarewellpotential: V=V0<0f o r0≤r<a,V=0f o rr>a.
TheSchrödingerwaveequationis
−¯h2
M∇2ψ+Vψ=Eψ, (10.10)
wheretheenergyeigenvalue E<0foraboundstate.Forthegroundstatetheorbitalangu-
lar momentum l=0 because for l/negationslash=0 there is the additional positive angular momentum
barrier. So, with ψ=ψ(r), we may write u(r)=rψ(r), and, using Exercise 2.5.18, the
waveequationbecomes
d2u
dr2+k2
1u=0, (10.11)
with
k2
1=M
¯h2(E−V0)>0 (10.12)
fortheinteriorrange, 0 ≤r<a.F ora<r<∞,weha v e
d2u
dr2−k2
2u=0, (10.13)
with
k2
2=−ME
¯h2>0. (10.14)
Theboundaryconditionthat ψremainfiniteat r=0 implies u(0)=0 and
u1(r)=sink1r,0≤r<a. (10.15)
In the range outside the potential well, we have a linear combination of the two exponen-
tials,
u2(r)=Aexpk2r+Bexp(−k2r), a<r< ∞. (10.16)
Continuity of particle and current density demand that u1(a)=u2(a)and thatu′
1(a)=
u′
2(a). Thesejoining,ormatching,conditions give
sink1a=Aexpk2a+Bexp(−k2a),
k1cosk1a=k2Aexpk2a−k2Bexp(−k2a).(10.17)
The condition that we actually have a bound proton–neutron combination is thatintegraltext∞
0u2(r)dr=1. This constraint can be met if we impose a boundarycondition that ψ(r)
10.1 Self-Adjoint ODEs 627
FIGURE 10.1Adeuteroneigenfunction.
remain finite as r→∞. And this, in turn, means that A=0. Dividing the preceding pair
ofequations(tocancel B), weobtain
tank1a=−k1
k2=−radicalbigg
E−V0
−E, (10.18)
atranscendentalequationfortheenergy Ewithonlycertaindiscretesolutions.If Eissuch
that Eq. (10.18) can be satisfied, our solutions u1(r)andu2(r)can satisfy the boundary
conditions. If Eq. (10.18) is not satisfied, no acceptable solution exists . The values of
Efor which Eq. (10.18) is satisfied are the eigenvalues; the corresponding functions u1
andu2(orψ)aretheeigenfunctions.Forthedeuteron,problemthereisone(andonlyone)
negativevalueof EsatisfyingEq.(10.18);thatis,thedeuteronhasoneandonlyonebound
state.
Now, what happens if Edoes not satisfy Eq. (10.18), that is, if E/negationslash=E0is not an
eigenvalue? In graphical form, imagine that Eand therefore k1are varied slightly. For
E=E1<E0,k1isreducedand sin k1ahasnotturneddownenoughtomatch exp (−k2a).
The joining conditions, Eq. (10.17), require A>0 and the wave function goes to +∞ex-
ponentially. For E=E2>E0,k1is larger, sin k1apeaks sooner and has descended more
rapidlyat r=a.Thejoiningconditionsdemand A<0,andthewavefunctiongoesto −∞
exponentially. Only for E=E0, an eigenvalue, will the wave function have the required
negativeexponentialasymptoticbehavior(see Fig.10.1). /squaresolid
Boundary Conditions
Intheforegoingdefinitionofeigenfunction,itwasnotedthattheeigenfunction uλ(x)was
required to satisfy certain imposed boundary conditions. The term boundary conditions
includes as a special case the concept of initial conditions . For instance, specifying the
initialposition x0andtheinitialvelocity v0insomedynamicalproblemwouldcorrespond
to the Cauchy boundary conditions. The only difference in the present usage of boundary
628 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
conditions in these one-dimensional problems is that we are going to apply the conditions
onbothendsof theallowedrangeofthevariable.
Usuallytheformofthedifferentialequationortheboundaryconditionsonthesolutions
will guarantee that at the ends of our interval (that is, at the boundary, as suggested by
Eq.(10.3)) thefollowingproductswillvanish:
p(x)v∗(x)du(x)
dxvextendsinglevextendsinglevextendsinglevextendsingle
x=a=0 and p(x)v∗(x)du(x)
dxvextendsinglevextendsinglevextendsinglevextendsingle
x=b=0.(10.19)
Hereu(x)andv(x)are solutions of the particular ODE (Eq. (10.8)) being considered.
A reason for this particular form of Eq. (10.19) is suggested shortly. If we recall the radial
wave function uof the hydrogen atom with u(0)=0 anddu/dr∼e−kr→0a sr→∞,
then both boundary conditions are satisfied. Similarly in the deuteron Example 10.1.2,
sink1r→0a sr→0 andd(e−k2r)/dr→0a sr→∞,both boundary conditions are
obeyed. We can, however, work with a somewhat less restrictive set of boundary condi-
tions,
v∗pu′vextendsinglevextendsingle
x=a=v∗pu′vextendsinglevextendsingle
x=b, (10.20)
inwhichu(x)andv(x)aresolutionsofthedifferentialequationcorrespondingtothesame
ortodifferenteigenvalues.Equation(10.20)mightwellbesatisfiedifweweredealingwith
aperiodicphysicalsystem, suchas acrystallattice.
Equations (10.19) and (10.20) are written in terms of v∗, complex conjugate. When the
solutionsarereal, v=v∗andtheasteriskmaybeignored.However,inFourierexponential
expansions and in quantum mechanics the functions will be complex and the complex
conjugatewillbeneeded.
Example 10.1.3 INTEGRATION INTERVAL [a,b]
ForL=d2/dx2, apossibleeigenvalueequationis
d2
dx2u(x)+n2u(x)=0, (10.21)
witheigenfunctions
un=cosnx, v m=sinmx.
Equation(10.20)becomes
−nsinmxsinnxvextendsinglevextendsingleb
a=0,ormcosmxcosnxvextendsinglevextendsingleb
a=0,
interchanging unandvm. Since sin mxand cosnxare periodic with period 2 π(fornand
mintegral), Eq. (10.20) is clearly satisfied if a=x0andb=x0+2π. If a problem pre-
scribes a different interval, the eigenfunctions and eigenvalues will change along with the
boundary conditions. The functions must always be chosen so that the boundary condi-
tions (Eq. (10.20) etc.) are satisfied. For this case (Fourier series) the usual choices are
x0=0 leading to (0,2π)andx0=−πleading to (−π,π). Here and throughout the fol-
lowing several chapters the orthogonality interval is so that the boundary conditions
(Eq. (10.20)) will be satisfied . The interval[a,b]and the weighting factor w(x)for the
mostcommonlyencounteredsecond-orderdifferentialequationsarelistedinTable10.2. /squaresolid
10.1 Self-Adjoint ODEs 629
Table 10.2
Equation ab w (x)
Legendre −11 1
Shifted Legendre 0 1 1
AssociatedLegendre −11 1
Chebyshev I −11 (1−x2)−1/2
Shifted Chebyshev I 0 1 [x(1−x)]−1/2
Chebyshev II −11 (1−x2)1/2
Laguerre 0 ∞ e−x
AssociatedLaguerre 0 ∞ xke−x
Hermite −∞ ∞ e−x2
Simple harmonic oscillator 0 2 π 1
−ππ 1
1. The orthogonalityinterval [a,b]is determined by the boundary condi-
tionsofSection10.1.
2. The weighting function is established by putting the ODE in self-
adjointform.
Hermitian Operators
Wenowproveanimportantpropertyoftheself-adjoint,second-orderdifferentialoperator
(Eq. (10.8)), in conjunction with solutions u(x)andv(x)that satisfy boundary conditions
givenbyEq.(10.20). Thisis motivatedbyapplicationsinquantummechanics.
By integrating v∗(complex conjugate) times the second-order self-adjoint differential
operator L(operatingon u)overtherange a≤x≤b,weobtain
integraldisplayb
av∗Ludx=integraldisplayb
av∗(pu′)′dx+integraldisplayb
av∗qudx (10.22)
usingEq. (10.6). Integratingbyparts,wehave
integraldisplayb
av∗(pu′)′dx=v∗pu′vextendsinglevextendsingleb
a−integraldisplayb
av∗′pu′dx. (10.23)
Theintegratedpartvanishesonapplicationoftheboundaryconditions(Eq.(10.20)).Inte-
gratingtheremainingintegralbyparts asecondtime,wehave
−integraldisplayb
av∗′pu′dx=−v∗′puvextendsinglevextendsingleb
a+integraldisplayb
au(pv∗′)′dx. (10.24)
Again, the integrated part vanishes in an application of Eq. (10.20). A combination of
Eqs. (10.22)to(10.24)givesus
integraldisplayb
av∗Ludx=integraldisplayb
au(Lv)∗dx. (10.25)
This property, given by Eq. (10.25), is expressed by saying that the operator Lis Her-
mitian with respect to the functions u(x)andv(x), which satisfy the boundary conditions
specifiedbyEq.(10.20).NotethatifthisHermitianpropertyfollowsfromself-adjointness
in a Hilbert space, then it includes that boundary conditions are imposed on all functions
ofthatspace.
630 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Hermitian Operators in Quantum Mechanics
The proceeding development in this section has focused on the classical second-order dif-
ferential operators of mathematical physics. Generalizing our Hermitian operator theory
as required in quantum mechanics, we have an extension: The operators need be neither
second-orderdifferentialoperatorsnorreal. px=−i¯h(∂/∂x)willbeaHermitianoperator.
Wesimplyassume(asiscustomaryinquantummechanics)thatthewavefunctionssatisfy
appropriate boundary conditions: vanishing sufficiently strongly at infinity or having peri-
odicbehavior(asinacrystallattice,orunitintensityforscatteringproblems).Theoperator
Lis calledHermitian if
integraldisplay
ψ∗
1Lψ2dτ=integraldisplay
(Lψ1)∗ψ2dτ. (10.26)
Apart from the simple extension to complex quantities, this definition is identical with
Eq.(10.25).
TheadjointA†of anoperator Aisdefinedby
integraldisplay
ψ∗
1A†ψ2dτ≡integraldisplay
(Aψ1)∗ψ2dτ. (10.27)
This generalizes our classical, second-derivative-operator–oriented definition, Eq. (10.5).
Here the adjoint is defined in terms of the resultant integral, with the A†as part of the
integrand. Clearly, if A=A†(self-adjoint ) and satisfies the aforementioned boundary
conditions,then AisHermitian.
Theexpectationvalue ofanoperator Lisdefinedas
/angbracketleftL/angbracketright=integraldisplay
ψ∗Lψdτ. (10.28a)
In the framework of quantum mechanics /angbracketleftL/angbracketrightcorresponds to the result of a measurement
of the physical quantity represented by Lwhen the physical system is in a state described
by the wave function ψ. If we require Lto be Hermitian, it is easy to show that /angbracketleftL/angbracketrightis
real (as would be expectedfrom a measurement in a physical theory). Taking the complex
conjugateof Eq.(10.28a), weobtain
/angbracketleftL/angbracketright∗=bracketleftbiggintegraldisplay
ψ∗Lψdτbracketrightbigg∗
=integraldisplay
ψL∗ψ∗dτ.
Rearrangingthefactorsintheintegrand,wehave
/angbracketleftL/angbracketright∗=integraldisplay
(Lψ)∗ψdτ.
Then,applyingour definitionofHermitianoperator,Eq.(10.26), weget
/angbracketleftL/angbracketright∗=integraldisplay
ψ∗Lψdτ=/angbracketleftL/angbracketright, (10.28b)
or/angbracketleftL/angbracketrightisreal.It is worthnotingthat ψis notnecessarilyaneigenfunctionof L.
10.1 Self-Adjoint ODEs 631
Exercises
10.1.1 ShowthatLaguerre’sODE,Eq.(13.52),maybeputintoself-adjointformbymultiply-
ingbye−xandthatw(x)=e−xis theweightingfunction.
10.1.2 Show that the Hermite ODE, Eq. (13.10), may be put into self-adjoint form by multi-
plyingby e−x2andthatthisgives w(x)=e−x2astheappropriatedensityfunction.
10.1.3 ShowthattheChebyshev(typeI)ODE,Eq.(13.100),maybeputintoself-adjointform
by multiplying by (1−x2)−1/2and that this gives w(x)=(1−x2)−1/2as the appro-
priatedensityfunction.
10.1.4 Show the following when the linear second-order differential equation is expressed in
self-adjointform:
(a) TheWronskianisequaltoa constantdividedbytheinitialcoefficient p:
W(x)=C
p(x).
(b) Asecondsolutionis givenby
y2(x)=Cy1(x)integraldisplayxdt
p(t)[y1(t)]2.
10.1.5 Un(x), theChebyshevpolynomial(typeII), satisfiestheODE,Eq.(13.101),
parenleftbig
1−x2parenrightbig
U′′
n(x)−3xU′
n(x)+n(n+2)Un(x)=0.
(a) Locate the singular points that appear in the finite plane, and show whether they
areregularor irregular.
(b) Putthis equationinself-adjointform.
(c) Identifythecompleteeigenvalue.
(d) Identifytheweightingfunction.
10.1.6 For the very special case λ=0 andq(x)=0 the self-adjoint eigenvalue equation be-
comes
d
dxbracketleftbigg
p(x)du(x)
dxbracketrightbigg
=0,
satisfiedby
du
dx=1
p(x).
Usethistoobtaina“second”solutionof thefollowing:
(a) Legendre’sequation,
(b) Laguerre’sequation,
(c) Hermite’sequation.
632 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
ANS.(a) u2(x)=1
2ln1+x
1−x,
(b)u2(x)−u2(x0)=integraldisplayx
x0etdt
t,
(c)u2(x)=integraldisplayx
0et2dt.
Thesesecondsolutionsillustratethedivergentbehaviorusuallyfoundinasecondsolu-
tion.
Note.In allthreecases u1(x)=1.
10.1.7 Given that Lu=0 andgLuis self-adjoint, show that for the adjoint operator
¯L,¯L(gu)=0.
10.1.8 Forasecond-orderdifferentialoperator Lthatis self-adjointshowthat
integraldisplayb
a[y2Ly1−y1Ly2]dx=p(y′
1y2−y1y′
2)vextendsinglevextendsingleb
a.
10.1.9 Show that if a function ψis required to satisfy Laplace’s equation in a finite region
of space and to satisfy Dirichlet boundary conditions over the entire closed bounding
surface, then ψis unique.
Hint.Oneof theforms ofGreen’stheorem,Section1.11, willbehelpful.
10.1.10 ConsiderthesolutionsoftheLegendre,Chebyshev,Hermite,andLaguerreequationsto
be polynomials. Show that the ranges of integration that guarantee that the Hermitian
operatorboundaryconditionswillbesatisfiedare
(a) Legendre [−1,1], (b) Chebyshev [−1,1],
(c) Hermite (−∞,∞), (d) Laguerre [0,∞).
10.1.11 Within the framework of quantum mechanics (Eqs. (10.26) and following), show that
thefollowingareHermitianoperators:
(a) momentum p=−i¯h∇≡−ih
2π∇
(b) angularmomentum L=−i¯hr×∇≡−ih
2πr×∇.
Hint. In Cartesian form Lis a linear combination of noncommuting Hermitian opera-
tors.
10.1.12 (a)Ais anon-Hermitianoperator.In thesenseof Eqs. (10.26)and(10.27), showthat
A+A†andi(A−A†)
areHermitianoperators.
(b) Usingtheprecedingresult,showthateverynon-Hermitianoperatormaybewritten
asalinearcombinationoftwoHermitianoperators.
10.1 Self-Adjoint ODEs 633
10.1.13 UandVare two arbitrary operators, not necessarily Hermitian. In the sense of
Eq. (10.27),showthat
(UV)†=V†U†.
NotetheresemblancetoHermitianadjointmatrices.
Hint.Applythedefinitionofadjointoperator,Eq. (10.27).
10.1.14 ProvethattheproductoftwoHermitianoperatorsisHermitian(Eq.(10.26))ifandonly
ifthetwooperatorscommute.
10.1.15 AandBarenoncommutingquantummechanicaloperators:
AB−BA=iC.
Showthat Cis Hermitian.Assumethatappropriateboundaryconditionsaresatisfied.
10.1.16 Theoperator LisHermitian.Showthat /angbracketleftL2/angbracketright≥0.
10.1.17 Aquantummechanicalexpectationvalueis definedby
/angbracketleftA/angbracketright=integraldisplay
ψ∗(x)Aψ(x)dx,
whereAis a linear operator. Show that demanding that /angbracketleftA/angbracketrightbe real means that Amust
beHermitian—withrespectto ψ(x).
10.1.18 From the definition of adjoint, Eq. (10.27), show that A††=Ain the sense thatintegraltext
ψ∗
1A††ψ2dτ=integraltext
ψ∗
1Aψ2dτ.Theadjointoftheadjointistheoriginaloperator.
Hint. The functions ψ1andψ2of Eq. (10.27) represent a class of functions. The sub-
scripts1and2maybeinterchangedorreplacedbyothersubscripts.
10.1.19 TheSchrödingerwaveequationforthedeuteron(withaWoods–Saxonpotential)is
−¯h2
2M∇2ψ+V0
1+exp[(r−r0)/a]ψ=Eψ.
HereE=−2.224 MeV, ais a “thickness parameter,” 0 .4×10−13cm. Expressing
lengths in fermis (10−13cm) and energies in million electron volts (MeV), we may
rewritethewaveequationas
d2
dr2(rψ)+1
41.47bracketleftbigg
E−V0
1+exp((r−r0)/a)bracketrightbigg
(rψ)=0.
Eis assumed known from experiment. The goal is to find V0for a specified value of
r0(say,r0=2.1). If we let y(r)=rψ(r), theny(0)=0 and we take y′(0)=1. Find
V0such that y(20.0)=0. (This should be y(∞),b u tr=20 is far enough beyond the
rangeofnuclearforces toapproximateinfinity.)
ANS.For a=0.4 andr0=2.1f m ,V0=−34.159 MeV.
10.1.20 Determine the nuclear potential well parameter V0of Exercise 10.1.19 as a function of
r0forr=2.00(0.05)2.25 fermis. Express yourresultsas apowerlaw
|V0|rν
0=k.
Determine the exponent νand the constant k. This power-law formulation is useful for
accurateinterpolation.
634 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.1.21 InExercise10.1.19itwasassumedthat20fermiswasagoodapproximationtoinfinity.
Checkonthisbycalculating V0forrψ(r)=0a t( a )r=15,(b)r=20,(c)r=25,and
(d)r=30.Sketchyourresults. Take r0=2.10 anda=0.4 (fermis).
10.1.22 For a quantum particle moving in a potential well, V(x)=1
2mω2x2, the Schrödinger
waveequationis
−¯h2
2md2ψ(x)
dx2+1
2mω2x2ψ(x)=Eψ(x),
or
d2ψ(z)
dz2−z2ψ(z)=−2E
¯hωψ(z),
wherez=(mω/¯h)1/2x. Since this operator is even, we expect solutions of definite
parity.Fortheinitialconditionsthatfollow,integrateoutfromtheoriginanddetermine
the minimum constant 2 E/¯hωthat will lead to ψ(∞)=0 in each case. (You may take
z=6 asanapproximationofinfinity.)
(a) Foraneveneigenfunction,
ψ(0)=1,ψ′(0)=0.
(b) Foranoddeigenfunction,
ψ(0)=0,ψ′(0)=1.
Note.AnalyticalsolutionsappearinSection13.1.
10.2 H ERMITIAN OPERATORS
Hermitian,orself-adjoint,operatorswithappropriateboundaryconditionshavethreeprop-
ertiesthatare ofextremeimportanceinphysics, bothclassicalandquantum.
1. TheeigenvaluesofaHermitianoperatorare real.
2. AHermitianoperatorpossessesanorthogonalsetofeigenfunctions.
3. TheeigenfunctionsofaHermitianoperatorform acompleteset.6
Real Eigenvalues
Weproceedtoprovethefirsttwo ofthesethreeproperties. Let
Lui+λiwui=0. (10.29)
6This third property is not universal. It doeshold for our linear, second-order differential operators in Sturm–Liouville (self-
adjoint)form.CompletenessisdefinedanddiscussedinSection10.4.Aproofthattheeigenfunctionsofourlinear,second-order,
self-adjoint, differential equations form a complete setmay be developed from the calculus ofvariations of Section 17.8.
10.2 Hermitian Operators 635
Assumingtheexistenceof asecondeigenvalueandeigenfunction,
Luj+λjwuj=0. (10.30)
Then,takingthecomplexconjugate,weobtain
L∗u∗
j+λ∗
jwu∗
j=0. (10.31)
Herew(x)≥0 is a real function. But we permit λk, the eigenvalues, and uk, the eigen-
functions, to be complex. Multiplying Eq. (10.29) by u∗
jand Eq. (10.31) by uiand then
subtracting,wehave
u∗
jLui−uiL∗u∗
j=(λ∗
j−λi)wuiu∗
j. (10.32)
Weintegrateovertherange a≤x≤b:
integraldisplayb
au∗
jLuidx−integraldisplayb
auiL∗u∗
jdx=(λ∗
j−λi)integraldisplayb
auiu∗
jwdx. (10.33)
SinceLis Hermitian,theleft-handsidevanishesbyEq. (10.26)and
(λ∗
j−λi)integraldisplayb
auiu∗
jwdx=0. (10.34)
Ifi=j, the integral cannot vanish [ w(x)>0, apart from isolated points], except in the
trivialcase ui=0.Hencethecoefficient (λ∗
i−λi)mustbezero,
λ∗
i=λi, (10.35)
which says that the eigenvalue is real. Since λican represent any one of the eigenvalues,
thisprovesthefirstproperty.Thisisanexactanalogofthenatureoftheeigenvaluesofreal
symmetric(andofHermitian)matrices(compareSection3.5).
The analog of the spectral decomposition of a real symmetric matrix in Section 3.5 for
aHermitianoperator Lwithadiscreteset ofeigenvalues λitakestheform
L=summationdisplay
iλi|ui/angbracketright/angbracketleftui|,f(L)=summationdisplay
if(λi)|ui/angbracketright/angbracketleftui|
witheigenvectors |ui/angbracketrightandanyinfinitelydifferentiablefunction f.
Real eigenvalues of Hermitian operators have a fundamental significance in quantum
mechanics. In quantum mechanics the eigenvalues correspond to precisely measurable
quantities, such as energy and angular momentum. With the theory formulated in terms
of Hermitian operators, this proof of real eigenvalues guarantees that the theory will pre-
dict real numbers for these measurable physical quantities. In Section 17.8 it will be seen
thatthesetof realeigenvalueshasalowerbound(for nonrelativisticproblems).
636 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Orthogonal Eigenfunctions
If we now take i/negationslash=jand ifλi/negationslash=λjin Eq. (10.34), the integral of the product of the two
differenteigenfunctionsmustvanish:
integraldisplayb
auiu∗
jwdx=0. (10.36)
This condition, called orthogonality , is the continuum analog of the vanishing of a scalar
product of two vectors.7We say that the eigenfunctions ui(x)anduj(x)are orthogonal
withrespecttotheweightingfunction w(x)overtheinterval [a,b].Equation(10.36)con-
stitutesapartialproofofthesecondpropertyofourHermitianoperators.Again,theprecise
analogywithmatrixanalysisshouldbenoted.Indeed,wecanestablishaone-to-onecorre-
spondencebetweenthisSturm–Liouvilletheoryofdifferentialequationsandthetreatment
ofHermitianmatrices.Historically,thiscorrespondencehasbeensignificantinestablishing
themathematicalequivalenceofmatrixmechanicsdevelopedbyHeisenbergandwaveme-
chanics developed by Schrödinger. Today, the two diverse approaches are merged into the
theory of quantum mechanics, and the mathematical formulation that is more convenient
foraparticularproblemisusedforthatproblem.Actuallythemathematicalalternativesdo
not end here. Integral equations, Chapter 16, form a third equivalent and sometimes more
convenientor morepowerfulapproach.
This proof of orthogonality is not quite complete. There is a loophole, because we may
haveui/negationslash=ujbut still have λi=λj. Such a case is labeled degenerate . Illustrations of
degeneracy are given at the end of this section. If λi=λj, the integral in Eq. (10.34) need
notvanish.Thismeansthatlinearlyindependenteigenfunctionscorrespondingtothesame
eigenvalue are not automatically orthogonal and that some other method must be sought
to obtain an orthogonal set. Although the eigenfunctions in this degenerate case may not
be orthogonal, they can always be made orthogonal. One method is developed in the next
section.SeealsoEq. (4.21)for degeneracyduetosymmetry.
We shall see in succeeding chapters that it is just as desirable to have a given set of
functions orthogonal as it is to have an orthogonal coordinate system. We can work with
nonorthogonal functions, but they are likely to prove as messy as an oblique coordinate
system.
Example 10.2.1 FOURIER SERIES —O RTHOGONALITY
TocontinueExample10.1.3, theeigenvalueequation,Eq. (10.21),
d2
dx2y(x)+n2y(x)=0,
7From the definition of Riemann integral,
integraldisplayb
af(x)g(x)dx=lim
N→∞parenleftbiggNsummationdisplay
i=1f(xi)g(xi)/Delta1xparenrightbigg
,
wherex0=a,xN=b,a n dxi−xi−1=/Delta1x. If we interpret f(xi)andg(xi)as theith components of an N-component vector,
then this sum (and therefore this integral) corresponds directly to a scalar product of vectors, Eq. (1.24). The vanishing of the
scalarproduct is thecondition for orthogonality of thevectors—or functions.
10.2 Hermitian Operators 637
may describe a quantum mechanical particle in a box, or perhaps a vibrating violin
string, a classical harmonic oscillator with degenerate eigenfunctions—cos nx,sinnx—
andeigenvalues n2,naninteger.
Withnreal(here takentobeintegral),theorthogonalityintegralsbecome
(a)integraldisplayx0+2π
x0sinmxsinnxdx=Cnδnm,
(b)integraldisplayx0+2π
x0cosmxcosnxdx=Dnδnm,
(c)integraldisplayx0+2π
x0sinmxcosnxdx=0.
For an interval of 2 πthe preceding analysis guarantees the Kronecker delta in (a) and
(b) but not the zero in (c) because (c) may involve degenerate eigenfunctions. However,
inspectionshows that(c) alwaysvanishesfor allintegral mandn.
OurSturm–Liouvilletheorysaysnothingaboutthevaluesof CnandDnbecausehomo-
geneousODEshavesolutionswhosescalingis arbitrary.Actualcalculationyields
Cn=braceleftBiggπ, n/negationslash=0,
0,n=0,Dn=braceleftBiggπ, n/negationslash=0,
2π, n=0.
These orthogonality integrals form the basis of the Fourier series developed in Chap-
ter14. /squaresolid
Example 10.2.2 EXPANSION IN ORTHOGONAL EIGENFUNCTIONS —SQUARE WAVE
Thepropertyofcompleteness(seeEq.(1.190)andSection10.4)meansthatcertainclasses
of functions (for example, sectionally or piecewise continuous) may be represented by a
seriesof orthogonaleigenfunctions.Considerthesquare-waveshape
f(x)=
h
2,0<x<π,
−h
2,−π<x<0.(10.37)
Thisfunctionmaybeexpandedinanyofavarietyofeigenfunctions—Legendre,Hermite,
Chebyshev,andsoon.Thechoiceofeigenfunctionismadeonthebasisofconvenienceor
an application. To illustrate the expansion technique, let us choose the eigenfunctions of
Example10.2.1, cos nxand sinnx.
Theeigenfunctionseries isconveniently(andconventionally)writtenas
f(x)=a0
2+∞summationdisplay
m=1(amcosmx+bmsinmx).
Upon multiplying f(t)by cosntor sinntand integrating, only the nth term survives, by
theorthogonalityintegralsofExample10.2.1, thusyieldingthecoefficients
an=1
πintegraldisplayπ
−πf(t)cosntdt, b n=1
πintegraldisplayπ
−πf(t)sinntdt, n=0,1,2....
638 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Directsubstitutionof ±h/2f o rf(t)yields
an=0,
whichisexpectedherebecauseof theantisymmetry, f(−x)=−f(x), and
bn=h
nπ(1−cosnπ)=
0,n even,
2h
nπ,nodd.
Hencetheeigenfunction(Fourier)expansionofthesquarewaveis
f(x)=2h
π∞summationdisplay
n=0sin(2n+1)x
2n+1. (10.38)
Additionalexamples,usingothereigenfunctions,appearinChapters11and12. /squaresolid
Degeneracy
The concept of degeneracy was introduced earlier. If Nlinearly independent eigenfunc-
tions correspond to the same eigenvalue, the eigenvalue is said to be N-fold degenerate.
A particularly simple illustration is provided by the eigenvalues and eigenfunctions of the
classical harmonic oscillator equation, Example 10.2.1. For each eigenvalue n2, there are
two possible solutions: sin nxand cosnx(and any linear combination, nan integer). We
saytheeigenfunctionsaredegenerateor theeigenvalueis degenerate.
A more involved example is furnished by the physical system of an electron in an atom
(nonrelativistictreatment,spinneglected).FromtheSchrödingerequation,Eq.(13.84)for
hydrogen,thetotalenergyoftheelectronisoureigenvalue.Wemaylabelit EnLMbyusing
thequantumnumbers n,L,andMassubscripts.Foreachdistinctsetofquantumnumbers
(n,L,M) thereisadistinct,linearlyindependenteigenfunction ψnLM(r,θ,ϕ).Forhydro-
gen, the energy EnLMis independent of LandM, reflecting the spherical (and SO(4))
symmetry of the Coulomb potential. With 0 ≤L≤n−1 and−L≤M≤L, the eigen-
value isn2-fold degenerate (including the electron spin would raise this to 2 n2). In atoms
withmorethanoneelectron,theelectrostaticpotentialisnolongerasimple r−1potential.
Theenergydependson Laswellason n,although notonM;EnLMisstill(2L+1)-fold
degenerate. This degeneracy—due to rotational invariance of the potential—may be re-
movedbyapplyinganexternalmagneticfield,breakingsphericalsymmetryandgivingrise
totheZeemaneffect.Asarule,theeigenfunctionsformaHilbertspace,thatis,acomplete
vector space of functions with a metric defined by the inner product (see Section 10.4 for
moredetailsandexamples).
Often an underlying symmetry, such as rotational invariance, is causing the degenera-
cies. States belonging to the same energy eigenvalue then will form a multiplet or repre-
sentation of the symmetry group. The powerful group-theoretical methods are treated in
Chapter4insomedetail.
10.2 Hermitian Operators 639
Exercises
10.2.1 The functions u1(x)andu2(x)are eigenfunctions of the same Hermitian operator but
fordistincteigenvalues λ1andλ2.Provethat u1(x)andu2(x)arelinearlyindependent.
10.2.2 (a) Thevectors enareorthogonaltoeachother: en·em=0forn/negationslash=m.Showthatthey
arelinearlyindependent.
(b) Thefunctions ψn(x)areorthogonaltoeachotherovertheinterval [a,b]andwith
respecttotheweightingfunction w(x).Showthatthe ψn(x)arelinearlyindepen-
dent.
10.2.3 Giventhat
P1(x)=xandQ0(x)=1
2lnparenleftbigg1+x
1−xparenrightbigg
are solutions of Legendre’s differential equation corresponding to different eigenval-
ues:
(a) Evaluatetheirorthogonalityintegral
integraldisplay1
−1x
2lnparenleftbigg1+x
1−xparenrightbigg
dx.
(b) Explain why these two functions are not orthogonal, that is, why the proof of
orthogonalitydoesnotapply.
10.2.4 T0(x)=1andV1(x)=(1−x2)1/2aresolutionsoftheChebyshevdifferentialequation
corresponding to different eigenvalues. Explain, in terms of the boundary conditions,
whythesetwofunctionsarenotorthogonal.
10.2.5 (a) Show that the first derivatives of the Legendre polynomials satisfy a self-adjoint
differentialequationwitheigenvalue λ=n(n+1)−2.
(b) ShowthattheseLegendrepolynomialderivativessatisfyanorthogonalityrelation
integraldisplay1
−1P′
m(x)P′
n(x)parenleftbig
1−x2parenrightbig
dx=0,m/negationslash=n.
Note.InSection12.5, (1−x2)1/2P′
n(x)willbelabeledanassociatedLegendrepolyno-
mial,P1
n(x).
10.2.6 Asetof functions un(x)satisfiestheSturm–Liouvilleequation
d
dxbracketleftbigg
p(x)d
dxun(x)bracketrightbigg
+λnw(x)un(x)=0.
The functions um(x)andun(x)satisfy boundary conditions that lead to orthogonality.
Thecorrespondingeigenvalues λmandλnaredistinct.Provethatforappropriatebound-
aryconditions, u′
m(x)andu′
n(x)areorthogonalwith p(x)asaweightingfunction.
640 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.2.7 A linear operator Ahasndistinct eigenvalues and ncorresponding eigenfunctions:
Aψi=λiψi. Show that the neigenfunctions are linearly independent. Ais not neces-
sarilyHermitian.
Hint. Assume linear dependence—that ψn=summationtextn−1
i=1aiψi. Use this relation and the
operator–eigenfunction equation first in one order and then in the reverse order. Show
thata contradictionresults.
10.2.8 (a) ShowthattheLiouvillesubstitution
u(x)=v(ξ)bracketleftbig
p(x)w(x)bracketrightbig−1/4,ξ=integraldisplayx
abracketleftbiggw(t)
p(t)bracketrightbigg1/2
dt
transforms
d
dxbracketleftbigg
p(x)d
dxubracketrightbigg
+bracketleftbig
λw(x)−q(x)bracketrightbig
u(x)=0
into
d2v
dξ2+bracketleftbig
λ−Q(ξ)bracketrightbig
v(ξ)=0,
where
Q(ξ)=q(x(ξ))
w(x(ξ))+bracketleftbig
pparenleftbig
x(ξ)parenrightbig
wparenleftbig
x(ξ)parenrightbigbracketrightbig−1/4d2
dξ2(pw)1/4.
(b) Ifv1(ξ)andv2(ξ)areobtainedfrom u1(x)andu2(x),respectively,byaLiouville
substitution,showthatintegraltextb
aw(x)u1u2dxistransformedintointegraltextc
0v1(ξ)v2(ξ)dξwith
c=integraltextb
a[w
p]1/2dx.
10.2.9 Theultrasphericalpolynomials C(α)
n(x)aresolutionsofthedifferentialequation
braceleftbigg
(1−x2)d2
dx2−(2α+1)xd
dx+n(n+2α)bracerightbigg
C(α)
n(x)=0.
(a) Transformthisdifferentialequationintoself-adjointform.
(b) Show that the C(α)
n(x)are orthogonal for different n. Specify the interval of inte-
grationandtheweightingfactor.
Note.Assumethatyoursolutionsarepolynomials.
10.2.10 WithLnotself-adjoint,
Lui+λiwui=0
and
¯Lvj+λjwvj=0.
10.2 Hermitian Operators 641
(a) Showthat
integraldisplayb
avjLuidx=integraldisplayb
aui¯Lvjdx,
provided
uip0v′
jvextendsinglevextendsingleb
a=vjp0u′
ivextendsinglevextendsingleb
a
and
ui(p1−p′
0)vjvextendsinglevextendsingleb
a=0.
(b) Showthattheorthogonalityintegralfor theeigenfunctions uiandvjbecomes
integraldisplayb
auivjwdx=0(λi/negationslash=λj).
10.2.11 InExercise9.5.8theseriessolutionoftheChebyshevequationisfoundtobeconvergent
foralleigenvalues n.Therefore nisnotquantizedbytheargumentusedforLegendre’s
(Exercise 9.5.5). Calculate the sum of the indicial equation k=0 Chebyshev series for
n=v=0.8,0.9,and1.0andfor x=0.0(0.1)0.9.
Note.TheChebyshevseries recurrencerelationisgiveninExercise5.2.16.
10.2.12 (a) Evaluate the n=ν=0.9, indicial equation k=0 Chebyshev series for x=
0.98,0.99,and1.00.Theseries convergesveryslowlyat x=1.00.Youmaywish
to use double precision. Upper bounds to the error in your calculation can be set
bycomparisonwiththe ν=1.0 case,whichcorrespondsto (1−x2)1/2.
(b) These series solutions for eigenvalue ν=0.9 and for ν=1.0 are obviously not
orthogonal,despitethefactthattheysatisfyaself-adjointeigenvalueequationwith
differenteigenvalues.Fromthebehaviorofthesolutionsinthevicinityof x=1.00
trytoformulateahypothesisastowhytheproofoforthogonalitydoesnotapply.
10.2.13 The Fourier expansion of the (asymmetric) square wave is given by Eq. (10.38). With
h=2,evaluatethisseriesfor x=0(π/18)π/2,usingthefirst(a)10terms,(b)100terms
oftheseries.
Note.For10termsand x=π/18,or10◦,yourFourierrepresentationhasasharphump.
ThisistheGibbsphenomenonofSection14.5.For100termsthishumphasbeenshifted
overtoabout1◦.
10.2.14 Thesymmetric squarewave
f(x)=
1,|x|<π
2
−1,π
2<|x|<π
hasaFourierexpansion
f(x)=4
π∞summationdisplay
n=0(−1)ncos(2n+1)x
2n+1.
Evaluatethisseries for x=0(π/18)π/2u s i n gt h efi r s t
642 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
(a) 10terms, (b)100termsof theseries.
Note. As in Exercise 10.2.13, the Gibbs phenomenon appears at the discontinuity. This
means that a Fourier series is not suitable for precise numerical work in the vicinity of
adiscontinuity.
10.3 G RAM –SCHMIDT ORTHOGONALIZATION
The Gram–Schmidt orthogonalization is a method that takes a nonorthogonal set of lin-
early independent vectors (see Section 3.1) or functions8and constructs an orthogonal set
of vectors or functions over an arbitrary interval and with respect to an arbitrary weight
or density factor. In the language of linear algebra, the process is equivalent to a matrix
transformation relating an orthogonal set of basis vectors (functions) to a nonorthogonal
set.AspecificexampleofthismatrixtransformationappearsinExercise12.2.1.
Next we apply the Gram–Schmidt procedure to a set of functions. The functions in-
volved may be real or complex. Here for convenience they are assumed to be real. The
generalizationtothecomplexcaseoffers nodifficulty.
Before taking up orthogonalization, we should consider normalization of functions. So
far nonormalizationhasbeenspecified.This meansthat
integraldisplayb
aϕ2
iwdx=N2
i,
but no attention has been paid to the value of Ni. Since our basic equation (Eq. (10.8)) is
linear and homogeneous, we may multiply our solution by any constant and it will still be
asolution.Wenowdemandthateachsolution ϕi(x)bemultipliedby N−1
isothatthenew
(normalized) ϕiwillsatisfy
integraldisplayb
aϕ2
i(x)w(x)dx=1 (10.39)
and
integraldisplayb
aϕi(x)ϕj(x)w(x)dx=δij. (10.40)
Equation (10.39) says that we have normalized to unity. Including the property of orthog-
onality, we have Eq. (10.40). Functions satisfying this equation are said to be orthonor-
mal(orthogonalplusunitnormalization).Othernormalizationsarecertainlypossible,and
indeed, by historical convention, each of the special functions of mathematical physics
treatedinChapters12and13willbenormalizeddifferently.
We consider three sets of functions: an original, linearly independent given set
un(x), n=0,1,2,...; an orthogonalized set ψn(x)to be constructed; and a final set
8Such a set of functions might well arise from the solutions of a PDE in which the eigenvalue was independent of one or more
of the constants of separation. As an example, we have the hydrogen atom problem (Sections 10.2 and 13.2). The eigenvalue
(energy) is independent of both the electron orbital angular momentum and its projection on the z-axis,m. Note, however, that
the origin of the setof functions is irrelevant to the Gram–Schmidt orthogonalization procedure.
10.3 Gram–Schmidt Orthogonalization 643
of functions ϕn(x), which are the normalized ψn. The original unmay be degenerate
eigenfunctions,butthisis notnecessary.We shallhavethefollowingproperties:
un(x) ψ n(x) ϕ n(x)
Linearlyindependent Linearlyindependent Linearlyindependent
Nonorthogonal Orthogonal Orthogonal
Unnormalized Unnormalized Normalized (orthonormal)
The Gram–Schmidt procedure takes the nthψfunction(ψn)to beun(x)plus an un-
knownlinearcombinationoftheprevious ϕ.Thepresenceofthenew un(x)willguarantee
linear independence. The requirement that ψn(x)be orthogonal to each of the previous
ϕyields just enough constraints to determine each of the unknown coefficients. Then the
fully determined ψnwill be normalized to unity, yielding ϕn(x). Then the sequence of
stepsis repeatedfor ψn+1(x).
We startwith n=0,letting
ψ0(x)=u0(x), (10.41)
withno“previous” ϕtoworry about.Thenwenormalize
ϕ0(x)=ψ0(x)
[integraltext
ψ2
0wdx]1/2. (10.42)
Forn=1,let
ψ1(x)=u1(x)+a1,0ϕ0(x). (10.43)
We demand that ψ1(x)be orthogonal to ϕ0(x). (At this stage the normalization of ψ1(x)
isirrelevant.) Thisorthogonalityleadsto
integraldisplay
ψ1ϕ0wdx=integraldisplay
u1ϕ0wdx+a1,0integraldisplay
ϕ2
0wdx=0. (10.44)
Sinceϕ0is normalizedtounity(Eq. (10.42)), wehave
a1,0=−integraldisplay
u1ϕ0wdx, (10.45)
fixingthevalueof a1,0.Normalizing,wedefine
ϕ1(x)=ψ1(x)
(integraltext
ψ2
1wdx)1/2. (10.46)
Finally,wegeneralizeso that
ϕi(x)=ψi(x)
(integraltext
ψ2
i(x)w(x)dx)1/2, (10.47)
where
ψi(x)=ui+ai,0ϕ0+ai,1ϕ1+···+ai,i−1ϕi−1. (10.48)
Thecoefficients ai,jaregivenby
ai,j=−integraldisplay
uiϕjwdx. (10.49)
644 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Equation(10.49)holdsfor unitnormalization.If someothernormalizationis selected,
integraldisplayb
abracketleftbig
ϕj(x)bracketrightbig2w(x)dx=N2
j,
thenEq. (10.47)is replacedby
ϕi(x)=Niψi(x)
(integraltext
ψ2
iwdx)1/2. (10.47a)
andai,jbecomes
ai,j=−integraltext
uiϕjwdx
N2
j. (10.49a)
Equations (10.48) and (10.49) may be rewritten in terms of projection operators, Pj.I f
we consider the ϕn(x)to form a linear vector space, then the integral in Eq. (10.49) may
beinterpretedastheprojectionof uiintotheϕj“coordinate,”orthe jthcomponentof ui.
With
Pjui(x)=braceleftbiggintegraldisplay
ui(t)ϕj(t)w(t)dtbracerightbigg
ϕj(x),
Eq.(10.48) becomes
ψi(x)=braceleftbigg
1−i−1summationdisplay
j=1Pjbracerightbigg
ui(x). (10.48a)
Subtractingoff thecomponents, j=1t oi−1,leaves ψi(x)orthogonaltoallthe ϕj(x).
It will be noticed that although this Gram–Schmidt procedure is one possible way of
constructing an orthogonal or orthonormal set, the functions ϕi(x)are not unique. There
is an infinite number of possible orthonormal sets for a given interval and a given density
function.
As an illustration of the freedom involved, consider two (nonparallel) vectors AandB
in thexy-plane. We may normalize Ato unit magnitude and then form B′=aA+Bso
thatB′is perpendicular to A. By normalizing B′we have completed the Gram–Schmidt
orthogonalizationfortwovectors.Butanytwoperpendicularunitvectors,suchas ˆxandˆy,
couldhavebeenchosenasourorthonormalset.Again,withaninfinitenumberofpossible
rotations ofˆxandˆyabout the z-axis, we have an infinite number of possible orthonormal
sets.
Example 10.3.1 LEGENDRE POLYNOMIALS BY GRAM–SCHMIDT ORTHOGONALIZATION
Let us form an orthonormal set from the set of functions un(x)=xn,n=0,1,2....T h e
intervalis−1≤x≤1 andthedensityfunctionis w(x)=1.
In accordancewiththeGram–Schmidtorthogonalizationprocessdescribed,
u0=1,hence ϕ0=1√
2. (10.50)
10.3 Gram–Schmidt Orthogonalization 645
Then
ψ1(x)=x+a1,01√
2(10.51)
and
a1,0=−integraldisplay1
−1x√
2dx=0 (10.52)
bysymmetry.We normalize ψ1toobtain
ϕ1(x)=radicalbigg
3
2x. (10.53)
ThenwecontinuetheGram–Schmidtprocedurewith
ψ2(x)=x2+a2,01√
2+a2,1radicalbigg
3
2x, (10.54)
where
a2,0=−integraldisplay1
−1x2
√
2dx=−√
2
3, (10.55)
a2,1=−integraldisplay1
−1radicalbigg
3
2x3dx=0, (10.56)
againbysymmetry.Therefore
ψ2(x)=x2−1
3, (10.57)
and,onnormalizingtounity,wehave
ϕ2(x)=radicalbigg
5
2·1
2parenleftbig
3x2−1parenrightbig
. (10.58)
Thenextfunction, ϕ3(x), becomes
ϕ3(x)=radicalbigg
7
2·1
2parenleftbig
5x3−3xparenrightbig
. (10.59)
ReferencetoChapter12willshowthat
ϕn(x)=radicalbigg
2n+1
2Pn(x), (10.60)
wherePn(x)is thenth-order Legendre polynomial. Our Gram–Schmidt process provides
a possible but very cumbersome method of generating the Legendre polynomials. It il-
lustrates how a power-series expansion in un(x)=xn, which is not orthogonal, can be
convertedintoanorthogonalseries. /squaresolid
646 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
The equations for Gram–Schmidt orthogonalization tend to be ill-conditioned because
ofthesubtractions,Eqs.(10.48)and(10.49).Atechniqueforavoidingthisdifficultyusing
thepolynomialrecurrencerelationis discussedbyHamming.9
InExample10.3.1wehavespecifiedanorthogonalityinterval [−1,1],aunitweighting
function, and a set of functions xnto be taken one at a time in increasing order. Given
all these specifications, the Gram–Schmidt procedure is unique (to within a normaliza-
tion factor and an overall sign, as discussed subsequently). Our resulting orthogonal set,
the Legendre polynomials, P0up through Pn, form a complete set for the description of
polynomials of order ≤nover[−1,1]. This concept of completeness is taken up in detail
in Section 10.4. Expansions of functions in series of Legendre polynomials are found in
Section12.3.
Orthogonal Polynomials
Example 10.3.1 has been chosen strictly to illustrate the Gram–Schmidt procedure. Al-
though it has the advantage of introducing the Legendre polynomials, the initial functions
un=xnare not degenerate eigenfunctions and are not solutions of Legendre’s equation.
They are simply a set of functions that we have here rearranged to create an orthonor-
mal set for the given interval and given weighting function. The fact that we obtained the
Legendrepolynomialsisnotquiteblackmagicbutadirectconsequenceofthechoiceofin-
tervalandweightingfunction.Theuseof un(x)=xnbutwithotherchoicesofintervaland
Table 10.3 OrthogonalPolynomialsGeneratedbyGram–SchmidtOrthogonalization
ofun(x)=xn,n=0,1,2,...
Weighting
Polynomials Interval function w(x) Standard normalization
Legendre −1≤x≤11integraldisplay1
−1[Pn(x)]2dx=2
2n+1
Shifted Legendre 0 ≤x≤11integraldisplay1
0[P∗
n(x)]2dx=1
2n+1
Chebyshev I −1≤x≤1(1−x2)−1/2integraldisplay1
−1[Tn(x)]2
(1−x2)1/2dx=braceleftbiggπ/2,n/negationslash=0
π, n=0
Shifted Chebyshev I 0 ≤x≤1[x(1−x)]−1/2integraldisplay1
0[T∗n(x)]2
[x(1−x)]1/2dx=braceleftbiggπ/2,n>0
π, n=0
Chebyshev II −1≤x≤1(1−x2)1/2integraldisplay1
−1[Un(x)]2(1−x2)1/2dx=π
2
Laguerre 0 ≤x<∞ e−xintegraldisplay∞
0[Ln(x)]2e−xdx=1
AssociatedLaguerre 0 ≤x<∞ xke−xintegraldisplay∞
0[Lk
n(x)]2xke−xdx=(n+k)!
n!
Hermite −∞<x<∞ e−x2integraldisplay∞
−∞[Hn(x)]2e−x2dx=2nπ1/2n!
9R. W. Hamming, Numerical Methods for Scientists and Engineers , 2nd ed.,NewYork: McGraw-Hill (1973). See Section 27.2
andreferences given there.
10.3 Gram–Schmidt Orthogonalization 647
weighting function leads to other sets of orthogonal polynomials, as shown in Table 10.3.
We consider these polynomials in detail in Chapters 12 and 13 as solutions of particular
differentialequations.
Anexaminationofthisorthogonalizationprocesswillrevealtwoarbitraryfeatures.First,
asemphasizedbefore,itisnotnecessarytonormalizethefunctionstounity.Intheexample
justgivenwecouldhaverequired
integraldisplay1
−1ϕn(x)ϕm(x)dx=2
2n+1δnm, (10.61)
andtheresultingsetwouldhavebeentheactualLegendrepolynomials.Second,thesignof
ϕnis always indeterminate. In the example we chose the sign by requiring the coefficient
of the highest power of xin the polynomial to be positive. For the Laguerre polynomials,
ontheotherhand,wewouldrequirethecoefficientof thehighestpowertobe (−1)n/n!
Exercises
10.3.1 Rework Example 10.3.1 by replacing ϕn(x)by the conventional Legendre polynomial,
Pn(x):
integraldisplay1
−1bracketleftbig
Pn(x)bracketrightbig2dx=2
2n+1.
UsingEqs. (10.47a), and(10.49a), construct P0,P1(x), andP2(x).
ANS.P0=1,P1=x,P2=3
2x2−1
2.
10.3.2 Following the Gram–Schmidt procedure, construct a set of polynomials P∗
n(x)orthog-
onal (unit weighting factor) over the range [0,1]from the set[1,x]. Normalize so that
P∗
n(1)=1.
ANS.P∗
n(x)=1,
P∗
1(x)=2x−1,
P∗
2(x)=6x2−6x+1,
P∗
3(x)=20x3−30x2+12x−1.
Thesearethefirstfour shiftedLegendrepolynomials.
Note.The“*”isthestandardnotationfor“shifted”: [0,1]insteadof[−1,1].Itdoesnot
meancomplexconjugate.
10.3.3 ApplytheGram–Schmidtproceduretoform thefirst threeLaguerrepolynomials
un(x)=xn,n=0,1,2,..., 0≤x<∞,w(x)=e−x.
Theconventionalnormalizationis
integraldisplay∞
0Lm(x)Ln(x)e−xdx=δmn.
ANS.L0=1,L1=(1−x),L2=2−4x+x2
2.
648 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.3.4 You are given
(a) asetof functions un(x)=xn,n=0,1,2,...,
(b) aninterval (0,∞),
(c) aweightingfunction w(x)=xe−x.UsetheGram–Schmidtproceduretoconstruct
the firstthree orthonormal functions from the set un(x)for this interval and this
weightingfunction.
ANS.ϕ0(x)=1,ϕ1(x)=(x−2)/√
2,ϕ2(x)=parenleftbig
x2−6x+6parenrightbig
/2√
3.
10.3.5 Using the Gram–Schmidt orthogonalization procedure, construct the lowest three Her-
mitepolynomials:
un(x)=xn,n=0,1,2,...,−∞<x<∞,w(x)=e−x2.
Forthissetof polynomialstheusualnormalizationis
integraldisplay∞
−∞Hm(x)Hn(x)w(x)dx=δmn2mm!π1/2.
ANS.H0=1,H1=2x,H2=4x2−2.
10.3.6 UsetheGram–SchmidtorthogonalizationschemetoconstructthefirstthreeChebyshev
polynomials(typeI):
un(x)=xn,n=0,1,2,...,−1≤x≤1,w(x)=parenleftbig
1−x2parenrightbig−1/2.
Takethenormalization
integraldisplay1
−1Tm(x)Tn(x)w(x)dx=δmn
π, m=n=0,
π
2,m=n≥1.
Hint.TheneededintegralsaregiveninExercise8.4.3.
ANS.T0=1,T1=x,T2=2x2−1(T3=4x3−3x).
10.3.7 UsetheGram–SchmidtorthogonalizationschemetoconstructthefirstthreeChebyshev
polynomials(typeII):
un(x)=xn,n=0,1,2,...,−1≤x≤1,w(x)=parenleftbig
1−x2parenrightbig+1/2.
Takethenormalizationtobe
integraldisplay1
−1Um(x)Un(x)w(x)dx=δmnπ
2.
Hint.
integraldisplay1
−1parenleftbig
1−x2parenrightbig1/2x2ndx=π
2×1·3·5···(2n−1)
4·6·8···(2n+2),n=1,2,3,...
=π
2,n=0.
ANS.U0=1,U1=2x,U2=4x2−1.
10.4 Completeness of Eigenfunctions 649
10.3.8 As a modification of Exercise 10.3.5, apply the Gram–Schmidt orthogonalization pro-
cedure to the set un(x)=xn,n=0,1,2,...,0≤x<∞.T a k ew(x)to be exp[−x2].
Find the first two nonvanishing polynomials. Normalize so that the coefficient of the
highestpowerof xisunity.InExercise10.3.5theinterval (−∞,∞)ledtotheHermite
polynomials.ThesearecertainlynottheHermitepolynomials.
ANS.ϕ0=1,ϕ1=x−π−1/2.
10.3.9 Form an orthogonal set over the interval 0 ≤x<∞,u s i n gun(x)=e−nx,n=
1,2,3,....Take the weighting factor, w(x), to be unity. These functions are solutions
ofu′′
n−n2un=0,whichisclearlyalreadyinSturm–Liouville(self-adjoint)form.Why
doesn’ttheSturm–Liouvilletheoryguaranteetheorthogonalityof thesefunctions?
10.4 C OMPLETENESS OF EIGENFUNCTIONS
The third important property of an Hermitian operator is that its eigenfunctions form a
complete set. This completeness means that any well-behaved (at least piecewise continu-
ous)function F(x)canbeapproximatedbyaseries
F(x)=∞summationdisplay
n=0anϕn(x) (10.62)
to any desired degree of accuracy.10More precisely, the set ϕn(x)is calledcomplete11if
thelimitofthemeansquareerror vanishes:
limm→∞integraldisplayb
abracketleftbigg
F(x)−msummationdisplay
n=0anϕn(x)bracketrightbigg2
w(x)dx=0. (10.63)
Technically, the integral here is a Lebesgue integral. We have not required that the error
vanishidenticallyin [a,b]butonlythattheintegralof theerror squaredgotozero.
This convergence in the mean, Eq. (10.63), should be compared with uniform conver-
gence (Section 5.5, Eq. (5.67)). Clearly, uniform convergence implies convergence in the
mean, but the converse does not hold; convergence in the mean is less restrictive. Specifi-
cally,Eq.(10.63)isnotupsetbypiecewisecontinuousfunctionswithonlyafinitenumber
of finite discontinuities. A relevant example is the Gibbs phenomenon of discontinuous
Fourier series discussed in Section 14.5, which occurs for other eigenfunction series as
well.
Equation (10.63) is perfectly adequate for our purposes and is far more convenient than
Eq.(5.67).Indeed,sincewefrequentlyuseeigenfunctionstodescribediscontinuousfunc-
tions,convergenceinthemeanisallwecanexpect.
In Eq.(10.62) theexpansioncoefficients ammaybedeterminedby
am=integraldisplayb
aF(x)ϕ∗
m(x)w(x)dx. (10.64)
10If wehave afinite set,as with vectors, the summation is over the number of linearly independent members of the set.
11Many authors use the term closedhere.
650 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
This follows from multiplying Eq. (10.62) by ϕ∗
m(x)w(x)and integrating. From the or-
thogonalityoftheeigenfunctions ϕn(x),onlythe mthtermsurvives.Hereweseethevalue
of orthogonality. Equation (10.64) may be compared with the dot or inner product of vec-
tors, Section 1.3, and aminterpreted as the mth projection of the function F(x).O f t e nt h e
coefficient amiscalleda generalizedFouriercoefficient .
For a known function F(x), Eq. (10.64) gives amas adefiniteintegral that can always
beevaluated,bycomputerif notanalytically.
In the language of linear algebra, we have a linear space, a function vector space.
The linearly independent, orthonormal functions ϕn(x)form the basis for this (infinite-
dimensional) space. Equation (10.62) is a statement that the functions ϕn(x)span this
linear space. With an inner product defined by Eq. (10.64), our linear space is a Hilbert
space.
Setting the weight function w(x)=1 for simplicity, completeness in operator form for
adiscreteset ofeigenfunctions |ϕi/angbracketrightbecomes
summationdisplay
i|ϕi/angbracketright/angbracketleftϕi|=1.
Multiplyingthecompletenessrelationby |F/angbracketrightweobtaintheeigenfunctionexpansion
|F/angbracketright=summationdisplay
i|ϕi/angbracketright/angbracketleftϕi|F/angbracketright
with the generalized Fourier coefficient ai=/angbracketleftϕi|F/angbracketright.Equivalently in coordinate represen-
tation,
summationdisplay
iϕ∗
i(y)ϕi(x)=δ(x−y)
implies
F(x)=integraldisplay
F(y)δ(x−y)dy=summationdisplay
iϕi(x)integraldisplay
ϕ∗
i(y)F(y)dy.
Without proof, we state that the spectrum of a linear operator Athat maps a Hilbert
spaceHintoitselfmaybedividedintoadiscrete(orpoint)spectrumwitheigenvectorsof
finite length, a continuous spectrum so that the eigenvalue equation Av=λvwithvinH
does not have a unique bounded inverse (A−λ)−1in a dense domain of Hand a residual
spectrumwhere (A−λ)−1isunboundedina domainnotdensein H.
The question of completeness of a set of functions is often determined by comparison
with a Laurent series, Section 6.5. In Section 14.1 this is done for Fourier series, thus
establishingthecompletenessofFourierseries.Forallorthogonalpolynomialsmentioned
inSection10.3it ispossibletofindapolynomialexpansionofeachpowerof z,
zn=nsummationdisplay
i=0aiPi(z), (10.65)
10.4 Completeness of Eigenfunctions 651
wherePi(z)istheithpolynomial.Exercises12.4.6,13.1.6,13.2.5,and13.3.22arespecific
examples of Eq. (10.65). Using Eq. (10.65), we may reexpress the Laurent expansion of
f(z)in terms of the polynomials, showing that the polynomial expansion exists (when it
exists, it is unique, Exercise 10.4.1). The limitation of this Laurent series development is
thatitrequiresthefunctiontobeanalytic.Equations(10.62)and(10.63)aremoregeneral.
F(x)maybeonlypiecewisecontinuous.Numerousexamplesoftherepresentationofsuch
piecewise continuous functions appear in Chapter 14 (Fourier series). A proof that our
Sturm–LiouvilleeigenfunctionsformcompletesetsappearsinCourantandHilbert.12
For examples of particular eigenfunction expansions, see the following: Fourier series,
Section10.2 andChapter14; Bessel andFourier–Besselexpansions,Section11.2; Legen-
dre series, Section 12.3; Laplace series, Section 12.6; Hermite series, Section 13.1; La-
guerreseries, Section13.2;andChebyshevseries, Section13.3.
Itmayalsohappenthattheeigenfunctionexpansion,Eq.(10.62),istheexpansionofan
unknown F(x)in a series of known eigenfunctions ϕn(x)with unknown coefficients an.
An example would be the quantum chemist’s attempt to describe an (unknown) mole-
cular wave function as a linear combination of known atomic wave functions. The un-
known coefficients anwould be determined by a variational technique—Rayleigh–Ritz,
Section17.8.
Bessel’s inequality
If the set of functions ϕn(x)does not form a complete set, possibly because we simply
have not included the required infinite number of members of an infinite set, we are led
to Bessel’s inequality. First, consider the finite case from vector analysis. Let Abe ann
componentvector,
A=e1a1+e2a2+···+enan, (10.66)
in which eiis a unit vector and aiis the corresponding component (projection) of A; that
is,
ai=A·ei. (10.67)
Then
parenleftbigg
A−summationdisplay
ieiaiparenrightbigg2
≥0. (10.68)
If we sum over all ncomponents, the summation clearly, equals Aby Eq. (10.66) and
the equality holds. If, however, the summation does not include all ncomponents, the
inequality results. By expanding Eq. (10.68) and choosing the unit vectors so as to satisfy
anorthogonalityrelation,
ei·ej=δij, (10.69)
12R. Courant and D. Hilbert, Methods of Mathematical Physics (English translation), Vol. 1, New York: Interscience (1953),
reprinted, Wiley (1989), Chapter 6,Section 3.
652 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
wehave
A2≥summationdisplay
ia2
i. (10.70)
Thisis Bessel’sinequality.
Forrealfunctionsweconsidertheintegral
integraldisplayb
abracketleftbigg
f(x)−summationdisplay
iaiϕi(x)bracketrightbigg2
w(x)dx≥0. (10.71)
This is the continuum analog of Eq. (10.68), letting n→∞and replacing the summation
byanintegration.Again,withtheweightingfactor w(x)>0,theintegrandisnonnegative.
The integral vanishes by Eq. (10.62) if we have a complete set. Otherwise it is positive.
Expandingthesquaredterm,weobtain
integraldisplayb
abracketleftbig
f(x)bracketrightbig2w(x)dx−2summationdisplay
iaiintegraldisplayb
af(x)ϕi(x)w(x)dx+summationdisplay
ia2
i≥0.(10.72)
ApplyingEq. (10.64),wehave
integraldisplayb
abracketleftbig
f(x)bracketrightbig2w(x)dx≥summationdisplay
ia2
i. (10.73)
Hence the sum of the squares of the expansion coefficients aiis less than or equal to the
weightedintegralof [f(x)]2,theequalityholdingifandonlyiftheexpansionisexact,that
is, ifthesetof functions ϕn(x)is acompleteset.
In later chapters, when we consider eigenfunctions that form complete sets (such as
Legendre polynomials), Eq. (10.73) with the equal sign holding will be called a Parseval
relation.
Bessel’s inequality has a variety of uses, including proof of convergence of the Fourier
series.
Schwarz Inequality
The frequently used Schwarz inequality is similar to the Bessel inequality. Consider the
quadraticequationwithunknown x:
nsummationdisplay
i=1(aix+bi)2=nsummationdisplay
i=1a2
iparenleftbigg
x+bi
aiparenrightbigg2
=0 (10.74)
withreal ai,bi.Ifbi/ai=constant, c,thatis,independentoftheindex i,thenthesolution
isx=−c.Ifbi/aiisnotaconstantin i,alltermscannotvanishsimultaneouslyforreal x.
Sothesolutionmustbecomplex.Expanding,wefindthat
x2nsummationdisplay
ia2
i+2xnsummationdisplay
iaibi+nsummationdisplay
ib2
i=0, (10.75)
10.4 Completeness of Eigenfunctions 653
andsince xis complex(or =−bi/ai), thequadraticformula13forxleadsto
parenleftbiggnsummationdisplay
i=1aibiparenrightbigg2
≤parenleftbiggnsummationdisplay
i=1a2
iparenrightbiggparenleftbiggnsummationdisplay
i=1b2
iparenrightbigg
, (10.76)
theequalityholdingwhen bi/aiequalsaconstant,independentof i.
Oncemore, interms ofvectors,wehave
(a·b)2=a2b2cos2θ≤a2b2, (10.77)
whereθistheangleincludedbetween aandb.
TheanalogousSchwarzinequalityforcomplexfunctionshastheform
vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplayb
af∗(x)g(x)dxvextendsinglevextendsinglevextendsinglevextendsingle2
≤integraldisplayb
af∗(x)f(x)dxintegraldisplayb
ag∗(x)g(x)dx, (10.78)
theequalityholdingifandonlyif g(x)=αf (x),αbeingaconstant.Toprovethisfunction
formoftheSchwarzinequality,14consideracomplexfunction ψ(x)=f(x)+λg(x)with
λa complex constant, where f(x)andg(x)are any two square integrable functions (for
which the integrals on the right-hand side exist). Multiplying by the complex conjugate
andintegrating,weobtain
integraldisplayb
aψ∗ψdx≡integraldisplayb
af∗fdx+λintegraldisplayb
af∗gdx+λ∗integraldisplayb
ag∗fdx
+λλ∗integraldisplayb
ag∗gdx≥0. (10.79)
The≥0appearssince ψ∗ψisnonnegative,theequal (=)signholdingonlyif ψ(x)isiden-
ticallyzero.Notingthat λandλ∗arelinearlyindependent,wedifferentiatewithrespectto
oneof themandsetthederivativeequaltozerotominimizeintegraltextb
aψ∗ψdx:
∂
∂λ∗integraldisplayb
aψ∗ψdx=integraldisplayb
ag∗fdx+λintegraldisplayb
ag∗gdx=0.
Thisyields
λ=−integraltextb
ag∗fdx
integraltextb
ag∗gdx. (10.80a)
Takingthecomplexconjugate,weobtain
λ∗=−integraltextb
af∗gdx
integraltextb
ag∗gdx. (10.80b)
Substituting these values of λandλ∗back into Eq. (10.79), we obtain Eq. (10.78), the
Schwarzinequality.
13With negative (or zero) discriminant.
14An alternatederivation is provided by theinequalityintegraltextintegraltext
[f(x)g(y)−f(y)g(x)]∗[f(x)g(y)−f(y)g(x)]dxdy≥0.
654 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
In quantum mechanics f(x)andg(x)might each represent a state or configuration of
a physical system, that is, a linear combination of wave functions. Then the Schwarz in-
equality gives an upper limit for the absolute value of the inner productintegraltextb
af∗(x)g(x)dx .
In some texts the Schwarz inequality is a key step in the derivation of the Heisenberg
uncertaintyprinciple.
ThefunctionnotationofEqs.(10.78)and(10.79)isrelativelycumbersome.Inadvanced
mathematical physics and especially in quantum mechanics it is common to use the Dirac
bra-ketnotation.Usingthisnotation,wesimplyunderstandtherangeofintegration, (a,b),
andthepresenceoftheweightingfunction w(x)≥0.InthisnotationtheSchwarzinequal-
itytakestheelegantform
vextendsinglevextendsingle/angbracketleftf|g/angbracketrightvextendsinglevextendsingle2≤/angbracketleftf|f/angbracketright/angbracketleftg|g/angbracketright. (78a)
Ifg(x)is anormalizedeigenfunction, ϕi(x), Eq. (10.78)yields(here w(x)=1)
a∗
iai≤integraldisplayb
af∗(x)f(x)dx, (10.81)
aresultthatalsofollowsfrom Eq. (10.73).
For useful representations of Dirac’s delta function in terms of orthogonal sets of func-
tionsandtherelationbetweenclosureandcompletenesswerefertotherelevantsubsection
of Section 1.15, including Exercise 1.15.16, and for coordinate versus momentum repre-
sentationsinquantummechanicstoSection15.6.
Summary — Vector Spaces, Completeness
Here we summarize some properties of vector spaces, first with the vectors taken to be
the familiar real vectors of Chapter 1 and then with the vectors taken to be ordinary func-
tions.Theconceptof completeness hasbeendevelopedforfinitevectorspaces(Chapter1,
Eq. (1.5)) and carries over into infinite vector spaces. For example, in three-dimensional
Euclidean space every vector can be written in terms of a linear combination of the three
coordinate unit vectors (representing a basis) involvingthe vector’s Cartesian components
as the expansion coefficients. Or a periodic function of an infinite vector space can be ex-
panded in terms of the set of periodic functions sin nx,cosnx,n=0,1,2,...,that form a
basis of this space. Since any periodic function with reasonable properties (spelled out in
Chapter14)canbeexpandedintermsofthesesineandcosinefunctions,theyarecomplete
andform abasisofsuchalinearfunctionspace.
1v.We shall describe our vector space with a set of nlinearly independent vectors ei,
i=1,2,...,n.I fn=3, thene1=ˆx,e2=ˆy, ande3=ˆz.T h eneispanthe linear vector
space.
1f.We shall describe our vector (function) space with a set of nlinearly independent
functions, ϕi(x),i=0,1,...,n−1. The index istarts with 0 to agree with the labeling
of the classical polynomials. Here ϕi(x)is assumed to be a polynomial of degree i.T h e
nϕi(x)spanthelinearvector(function)space.
2v.Thevectorsinourvectorspacesatisfythefollowingrelations(Section1.2;thevector
componentsarenumbers):
10.4 Completeness of Eigenfunctions 655
a. Vectoradditionis commutative u+v=v+u
b. Vectoradditionis associative [u+v]+w=u+[v+w]
c. Thereisa nullvector 0+v=v
d. Multiplicationbyascalar
Distributive a[u+v]=au+av
Distributive (a+b)u=au+bu
Associative a[bu]=(ab)u
e. Multiplication
Byunitscalar 1 u=u
Byzero 0 u=0
f. Negativevector (−1)u=−u.
2f.The functions in our linear function space satisfy the properties listed for vectors
(substitute“function”for “vector”):
f(x)+g(x)=g(x)+f(x)
bracketleftbig
f(x)+g(x)bracketrightbig
+h(x)=f(x)+bracketleftbig
g(x)+h(x)bracketrightbig
0+f(x)=f(x)
abracketleftbig
f(x)+g(x)bracketrightbig
=af(x)+ag(x)
(a+b)f(x)=af(x)+bf (x)
abracketleftbig
bf (x)bracketrightbig
=(ab)f(x)
1·f(x)=f(x)
0·f(x)=0
(−1)·f(x)=−f(x).
3v.Inn-dimensionalvectorspaceanarbitraryvector cisdescribedbyits ncomponents
(c1,c2,...,cn),or
c=nsummationdisplay
i=1ciei.
Whennei(1) are linearly independent and (2) span the n-dimensional vector space, then
theeiforma basisandconstitutea complete set.
3f.Inn-dimensionalfunctionspaceapolynomialof degree m≤n−1 isdescribedby
f(x)=n−1summationdisplay
i=0ciϕi(x).
When the nϕi(x)(1) are linearly independent and (2) span the n-dimensional function
space,thenthe ϕi(x)formabasisandconstitutea complete set(fordescribingpolynomi-
alsofde gree m≤n−1).
4v.Aninnerproduct(scalar, dotproduct)of avectorspaceisdefinedby
c·d=nsummationdisplay
i=1cidi.
656 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Ifcanddhavecomplexcomponentsinanorthogonalcoordinatesystem,theinnerproduct
isdefinedassummationtextn
i=1c∗
idi.Theinnerproducthasthepropertiesof
a. Distributivelawofaddition c·(d+e)=c·d+c·e
b. Scalarmultiplication c·ad=ac·d
c. Complexconjugation c·d=(d·c)∗.
4f.Aninnerproductof alinearspaceof functionsis definedby
/angbracketleftf|g/angbracketright=integraldisplayb
af∗(x)g(x)w(x)dx.
The choice of the weighting function w(x)and the interval (a,b)follows from the dif-
ferential equation satisfied by ϕi(x)and the boundary conditions—Section 10.1. In ma-
trix terminology, Section 3.2, |g/angbracketrightis a column vector and /angbracketleftf|is a row vector, the adjoint
of|f/angbracketright,where both may have infinitely many components. For example, if we expand
g(x)=summationtext
igiϕi(x),then|g/angbracketrighthas theith component giin a column vector and |f/angbracketrighthasf∗
i
asitsithcomponentinarow vector.
Theinnerproducthasthepropertieslistedforvectors:
a./angbracketleftf|g+h/angbracketright=/angbracketleftf|g/angbracketright+/angbracketleftf|h/angbracketright
b./angbracketleftf|ag/angbracketright=a/angbracketleftf|g/angbracketright
c./angbracketleftf|g/angbracketright=/angbracketleftg|f/angbracketright∗.
5v. Orthogonality:
ej·ej=0,i/negationslash=j.
If theneiare not already orthogonal, the Gram–Schmidt process may be used to create
anorthogonalset.
5f. Orthogonality:
/angbracketleftϕi|ϕj/angbracketright=integraldisplayb
aϕ∗
i(x)ϕj(x)w(x)dx=0,i/negationslash=j.
Ifthenϕi(x)arenotalreadyorthogonal,theGram–Schmidtprocess(Section10.3)maybe
usedtocreateanorthogonalset.
6v.Definitionof norm:
|c|=(c·c)1/2=parenleftbiggnsummationdisplay
i=1c2
iparenrightbigg1/2
.
The basis vectors eiare taken to have unit norm (length) ei·ei=1. The components of c
aregivenby
ci=ei·c,i=1,2,...,n.
6f.Definitionof norm:
/bardblf/bardbl=/angbracketleftf|f/angbracketright1/2=bracketleftbiggintegraldisplayb
avextendsinglevextendsinglef(x)vextendsinglevextendsingle2w(x)dxbracketrightbigg1/2
=bracketleftbiggn−1summationdisplay
i=0|ci|2bracketrightbigg1/2
,
10.4 Completeness of Eigenfunctions 657
Parseval’s identity. /bardblf/bardbl>0 unless f(x)is identically zero. The basis functions ϕi(x)
maybetakentohaveunitnorm(unitnormalization),
/bardblϕi/bardbl=1.
Theexpansioncoefficientsof ourpolynomial f(x)aregivenby
ci=/angbracketleftϕi|f/angbracketright,i=0,1,...,n−1.
7v. Bessel’sinequality:
c·c≥summationdisplay
ic2
i.
If the equals sign holds for all c, it indicates that the eispan the vector space; that is, they
arecomplete.
7f. Bessel’sinequality:
/angbracketleftf|f/angbracketright=integraldisplayb
avextendsinglevextendsinglef(x)vextendsinglevextendsingle2w(x)dx≥summationdisplay
i|ci|2.
If the equals sign holds for all allowable f, it indicates that the ϕi(x)span the function
space;thatis, theyarecomplete.
8v. Schwarz’inequality:
|c·d|≤|c|·|d|.
The equals sign holds when cis a multiple of d. If the angle included between canddis
θ,then|cosθ|≤1.
8f. Schwarz’inequality:
vextendsinglevextendsingle/angbracketleftf|g/angbracketrightvextendsinglevextendsingle≤/angbracketleftf|f/angbracketright1/2/angbracketleftg|g/angbracketright1/2=/bardblf/bardbl·/bardblg/bardbl.
The equals sign holds when f(x)andg(x)are linearly dependent, that is, when f(x)is a
multipleof g(x).
Now,letn→∞,forminganinfinite-dimensionallinearvectorspace, l2.
9v.Inaninfinite-dimensionalspaceourvector cis
c=∞summationdisplay
i=1ciei.
Werequirethat
∞summationdisplay
i=1c2
i<∞.
Thecomponentsof caregivenby
ci=ei·c,i=1,2,...,∞,
exactlyas inafinite-dimensionalvectorspace.
Then let n→∞, forming an infinite-dimensional vector (function) space L2. ThenL
standsforLebesgue,thesuperscript2forthequadraticnorm,thatis,the2in |f(x)|2.Our
functionsneednolongerbepolynomials,butwedorequirethat f(x)beatleastpiecewise
658 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
continuous (Dirichlet conditions for Fourier series) and that /angbracketleftf|f/angbracketright=integraltextb
a|f(x)|2w(x)dx
exist.Thislatterconditionis oftenstatedas arequirementthat f(x)besquareintegrable.
9f.Cauchy sequence (generalized Fourier expansion): Expand f(x)=summationtext∞
i=0fiϕi(x)
andlet
fn(x)=nsummationdisplay
i=0fiϕi(x).
If
vextenddoublevextenddoublef(x)−fn(x)vextenddoublevextenddouble→0asn→∞
or
limn→∞integraldisplayvextendsinglevextendsinglevextendsinglevextendsinglef(x)−nsummationdisplay
i=0fiϕi(x)vextendsinglevextendsinglevextendsinglevextendsingle2
w(x)dx=0,
then we have convergence in the mean. This is analogous to the partial sum–Cauchy se-
quencecriterionfor theconvergenceof aninfiniteseries, Section5.1.
IfeveryCauchysequenceofallowablevectors(squareintegrable,piecewisecontinuous
functions) converges to a limit vector in our linear space, the space is said to be complete.
Then
f(x)=∞summationdisplay
i=0ciϕi(x) (almosteverywhere )
inthesenseofconvergenceinthemean.Asnotedbefore,thisisaweakerrequirementthan
pointwiseconvergence(fixedvalueof x)oruniformconvergence.
Expansion Coefficients
Forafunction fitsexpansioncoefficientsaredefinedas
ci=/angbracketleftϕi|f/angbracketright,i=0,1,...,∞,
exactlyas inafinite-dimensionalvectorspace.Hence
f(x)=summationdisplay
i/angbracketleftϕi|f/angbracketrightϕi(x).
A linear space (finite- or infinite-dimensional) that (1) has an inner product defined
(/angbracketleftf|g/angbracketright)and(2)is completeis a Hilbertspace .
Infinite-dimensionalHilbertspaceprovidesanaturalmathematicalframe-workformod-
ern quantummechanics.Away from quantummechanics,Hilbert space retains its abstract
mathematicalpowerandbeautyandhasmanyuses.
10.4 Completeness of Eigenfunctions 659
Exercises
10.4.1 Afunction f(x)is expandedinaseriesof orthonormaleigenfunctions
f(x)=∞summationdisplay
n=0anϕn(x).
Show that the series expansion is unique for a given set of ϕn(x). The functions ϕn(x)
arebeingtakenhereasthe basisvectorsinaninfinite-dimensionalHilbertspace.
10.4.2 Afunction f(x)is representedbyafinitesetof basisfunctions ϕi(x),
f(x)=Nsummationdisplay
i=1ciϕi(x).
Showthatthecomponents ciare unique,thatnodifferent set c′
iexists.
Note. Your basis functions are automatically linearly independent. They are not neces-
sarilyorthogonal.
10.4.3 A function f(x)is approximated by a power seriessummationtextn−1
i=0cixiover the interval [0,1].
Showthatminimizingthemeansquareerrorleadstoaset oflinearequations
Ac=b,
where
Aij=integraldisplay1
0xi+jdx=1
i+j+1,i,j=0,1,2,...,n−1
and
bi=integraldisplay1
0xif(x)dx, i =0,1,2,...,n−1.
Note.TheAijaretheelementsoftheHilbertmatrixoforder n.Thedeterminantofthis
Hilbertmatrixisarapidlydecreasingfunctionof n.Forn=5,detA=3.7×10−12and
thesetofequations Ac=bisbecomingill-conditionedandunstable.
10.4.4 Inplaceof theexpansionofa function F(x)givenby
F(x)=∞summationdisplay
n=0anϕn(x),
with
an=integraldisplayb
aF(x)ϕn(x)w(x)dx,
takethefiniteseries approximation
F(x)≈msummationdisplay
n=0cnϕn(x).
660 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Showthatthemeansquareerror
integraldisplayb
abracketleftbigg
F(x)−msummationdisplay
n=0cnϕn(x)bracketrightbigg2
w(x)dx
isminimizedbytaking cn=an.
Note.Thevaluesofthecoefficientsareindependentofthenumberoftermsinthefinite
series. This independence is a consequence of orthogonality and would not hold for a
least-squaresfitusingpowersof x.
10.4.5 FromExample10.2.2,
f(x)=
h
2,0<x<π
−h
2,−π<x<0
=2h
π∞summationdisplay
n=0sin(2n+1)x
2n+1.
(a) Showthat
integraldisplayπ
−πbracketleftbig
f(x)bracketrightbig2dx=π
2h2=4h2
π∞summationdisplay
n=0(2n+1)−2.
For a finite upper limit this would be Bessel’s inequality. For the upper limit ∞,
thisisParseval’sidentity.
(b) Verifythat
π
2h2=4h2
π∞summationdisplay
n=0(2n+1)−2
byevaluatingtheseries.
Hint.Theseriescanbeexpressedas theRiemannzetafunction.
10.4.6 DifferentiateEq. (10.79),
/angbracketleftψ|ψ/angbracketright=/angbracketleftf|f/angbracketright+λ/angbracketleftf|g/angbracketright+λ∗/angbracketleftg|f/angbracketright+λλ∗/angbracketleftg|g/angbracketright,
withrespectto λ∗andshowthatyougettheSchwarzinequality,Eq. (10.78).
10.4.7 DerivetheSchwarzinequalityfromtheidentity
bracketleftbiggintegraldisplayb
af(x)g(x)dxbracketrightbigg2
=integraldisplayb
abracketleftbig
f(x)bracketrightbig2dxintegraldisplayb
abracketleftbig
g(x)bracketrightbig2dx
−1
2integraldisplayb
aintegraldisplayb
abracketleftbig
f(x)g(y)−f(y)g(x)bracketrightbig2dxdy.
10.4.8 Ifthefunctions f(x)andg(x)oftheSchwarzinequality,Eq.(10.78),maybeexpanded
in a series of eigenfunctions ϕi(x), show that Eq. (10.78) reduces to Eq. (10.76) (with
npossiblyinfinite).
10.4 Completeness of Eigenfunctions 661
Notethedescriptionof f(x)asavectorinafunctionspaceinwhich ϕi(x)corresponds
totheunitvector e1.
10.4.9 Theoperator HisHermitianandpositivedefinite;thatis, for all f:
integraldisplayb
af∗Hf dx > 0.
ProvethegeneralizedSchwarzinequality:
vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplayb
af∗Hgdxvextendsinglevextendsinglevextendsinglevextendsingle2
≤integraldisplayb
af∗Hf dxintegraldisplayb
ag∗Hgdx.
10.4.10 A normalized wave function ψ(x)=summationtext∞
n=0anϕn(x). The expansion coefficients anare
knownasprobabilityamplitudes.Wemaydefineadensitymatrix ρwithelements ρij=
aia∗
j. Showthat
parenleftbig
ρ2parenrightbig
ij=ρij,
or
ρ2=ρ.
Thisresult, bydefinition,makes ρaprojectionoperator.
Hint:Use
integraldisplay
ψ∗ψdx=1.
10.4.11 Showthat
(a) theoperator
vextendsinglevextendsingleϕi(x)angbracketrightbigangbracketleftbig
ϕi(t)vextendsinglevextendsingle
operatingon
f(t)=summationdisplay
jcjvextendsinglevextendsingleϕj(t)angbracketrightbig
yields
civextendsinglevextendsingleϕi(x)angbracketrightbig
.
(b)summationdisplay
ivextendsinglevextendsingleϕi(x)angbracketrightbigangbracketleftbig
ϕi(x)vextendsinglevextendsingle=1.
This operator is a projection operator projecting f(x)onto theith coordinate,
selectivelypickingoutthe ithcomponent ci|ϕi(x)/angbracketrightoff(x).
Hint.Theoperatoroperatesviathewell-definedinnerproduct.
662 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.5 G REEN ’SFUNCTION —E IGENFUNCTION EXPANSION
Aseriessomewhatsimilartothatrepresenting δ(x−t)resultswhenweexpandtheGreen’s
functionintheeigenfunctionsofthecorrespondinghomogeneousequation.Intheinhomo-
geneousHelmholtzequationwehave
∇2ψ(r)+k2ψ(r)=−ρ(r). (10.82)
ThehomogeneousHelmholtzequationis satisfiedbyits orthonormaleigenfunctions ϕn,
∇2ϕn(r)+k2
nϕn(r)=0. (10.83)
As outlined in Section 9.7, the Green’s function G(r1,r2)satisfies the point source equa-
tion
∇2G(r1,r2)+k2G(r1,r2)=−δ(r1−r2) (10.84)
and the boundary conditions imposed on the solutions of the homogeneous equation. Be-
causeGis real, we expand the Green’s function in a series of real eigenfunctions of the
homogeneousequation(10.83);thatis,
G(r1,r2)=∞summationdisplay
n=0an(r2)ϕn(r1), (10.85)
andbysubstitutingintoEq.(10.84) weobtain
−∞summationdisplay
n=0an(r2)k2
nϕn(r1)+k2∞summationdisplay
n=0an(r2)ϕn(r1)=−∞summationdisplay
n=0ϕn(r1)ϕn(r2).(10.86)
Hereδ(r1−r2)has been replaced by its eigenfunction expansion, Eq. (1.190). When we
employtheorthogonalityof ϕn(r1)toisolate an,thisyields
∞summationdisplay
m=0am(r2)parenleftbig
k2−k2
mparenrightbigintegraldisplay
ϕn(r1)ϕm(r1)d3r1=−∞summationdisplay
m=0ϕm(r2)integraldisplay
ϕn(r1)ϕm(r1)d3r1,
or
an(r2)parenleftbig
k2−k2
nparenrightbig
=−ϕn(r2).
Thensubstitutingthis intoEq. (10.85), theGreen’sfunctionbecomes
G(r1,r2)=∞summationdisplay
n=0ϕn(r1)ϕn(r2)
k2n−k2, (10.87)
a bilinear expansion, symmetric with respect to r1andr2, as expected. Finally, ψ(r1),t h e
desiredsolutionoftheinhomogeneousequation,is givenby
ψ(r1)=integraldisplay
G(r1,r2)ρ(r2)dτ2. (10.88)
10.5 Green’s Function — Eigenfunction Expansion 663
If wegeneralizeourinhomogeneousdifferentialequationto
Lψ+λψ=−ρ, (10.89)
where Lis aHermitianoperator,wefindthat
G(r1,r2)=∞summationdisplay
n=0ϕn(r1)ϕn(r2)
λn−λ, (10.90)
whereλnis thenth eigenvalue and ϕnis the corresponding orthonormal eigenfunction of
thehomogeneousdifferentialequation
Lψ+λψ=0. (10.91)
The eigenfunction expansion of the Green’s function in Eq. (10.90) makes the symmetry
propertyG(r1,r2)=G(r2,r1)explicitandisoftenusefulwhencomparingwithsolutions
obtainedbyothermeans.
Green’s Functions — One-Dimensional
The development of the Green’s function for two- and three-dimensional systems was the
topicdiscussedintheprecedingmaterialandinSection9.7.Here,forsimplicity,werestrict
ourselvestoone-dimensionalcasesandfollowa somewhatdifferent approach.
DefiningProperties
Inourone-dimensionalanalysisweconsiderfirst theinhomogeneousequation
Ly(x)+f(x)=0, (10.92)
inwhich Listheself-adjoint differentialoperator
L=d
dxparenleftbigg
p(x)d
dxparenrightbigg
+q(x). (10.93)
AsinSection10.1, y(x)isrequiredtosatisfycertainboundaryconditionsattheendpoints
aandbof ourinterval [a,b].
We now proceed to define a rather strange and arbitrary function Gover the interval
[a,b]. At this stage the most that can be said in defense of Gis that the defining prop-
erties are legitimate, or mathematically acceptable. Later, Gwill appear as a reasonable
tool for obtaining solutions of the inhomogeneous ODE, Eq. (10.92); this role dictates its
properties.
1. The interval a≤x≤bis divided by a parameter t. We label G(x)=G1(x)fora≤
x<tandG(x)=G2(x)fort<x≤b.
2. Thefunctions G1(x)andG2(x)eachsatisfy thehomogeneous15equation;thatis,
LG1(x)=0,a≤x<t,
LG2(x)=0,t<x≤b.(10.94)
15Homogeneous with respectto the unknown function. Thefunction f(x)in Eq.(10.92) is setequal to zero.
664 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
3. Atx=a,G1(x)satisfies the boundary conditions we impose on y(x),as o l u t i o no f
the inhomogeneous ODE, Eq. (10.92). At x=b,G2(x)satisfies the boundary condi-
tions imposed on y(x)at this endpoint of the interval. For convenience, the boundary
conditionsaretakentobehomogeneous;thatis, at x=a,
y(a)=0,ory′(a)=0,orαy(a)+βy′(a)=0
andsimilarlyat x=b.
4. We demandthat G(x)becontinuous ,16
limx→t−G1(x)=limx→t+G2(x). (10.95)
5. We requirethat G′(x)bediscontinuous ,specificallythat15
d
dxG2(x)vextendsinglevextendsingle
t−d
dxG1(x)vextendsinglevextendsingle
t=−1
p(t), (10.96)
wherep(t)comes from the self-adjoint operator, Eq. (10.93). Note that with the first
derivativediscontinuous,thesecondderivativedoesnotexist.
These requirements, in effect, make G a function of two variables, G(x,t).A l s o ,w e
notethat G(x,t)dependsonboththeformofthedifferentialoperator Landtheboundary
conditions that y(x)must satisfy. Note that we have described the properties of Green’s
functions for second-order differential equations. Note that for Green’s functions for first-
orderdifferentialequations,thediscontinuitiesarisein Gitself.
Now, assuming that we can find a function G(x,t)that has these properties, we label it
aGreen’sfunctionandproceedtoshowthatasolutionof Eq.(10.92) is
y(x)=integraldisplayb
aG(x,t)f(t)dt. (10.97)
To do this we first construct the Green’s function G(x,t).L e tu(x)be a solution of the
homogeneous equation that satisfies the boundary conditions at x=a, and letv(x)be a
solutionthatsatisfiestheboundaryconditionsat x=b.Thenwemaytake17
G(x,t)=braceleftBiggc1u(x), a≤x<t,
c2v(x), t <x ≤b.(10.98)
Continuityat x=t(Eq. (10.95))requires
c2v(t)−c1u(t)=0. (10.99)
Finally,thediscontinuityinthefirstderivative(Eq. (10.96)) becomes
c2v′(t)−c1u′(t)=−1
p(t). (10.100)
16Strictly speaking, this is the limit as x→t.
17The“constants” c1andc2areindependent of x, but they may (and do) depend on the other variable, t.
10.5 Green’s Function — Eigenfunction Expansion 665
There will be a unique solution for our unknown coefficients c1andc2if the Wronskian
determinantvextendsinglevextendsinglevextendsinglevextendsinglevextendsingleu(t) v(t)
u′(t) v′(t)vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=u(t)v′(t)−v(t)u′(t)
does not vanish. We have seen in Section 9.6 that the nonvanishing of this determinant is
a necessary condition for linear independence. Let us assume u(x)andv(x)to be inde-
pendent.(If u(x)andv(x)arelinearlydependent,thesituationbecomesmorecomplicated
andisnotconsideredhere. SeeCourantandHilbertinAdditionalReadingsof Chapter9.)
For independent u(x)andv(x)we have the Wronskian (again from Section 9.6 or Exer-
cise10.1.4)
u(t)v′(t)−v(t)u′(t)=A
p(t), (10.101)
inwhichAisaconstant.Equation(10.101)issometimescalled Abel’sformula .Numerous
examples have appeared in connection with Bessel and Legendre functions. Now, from
Eq.(10.100), weidentify
c1=−v(t)
A,c 2=−u(t)
A. (10.102)
Equation(10.99)isclearlysatisfied.SubstitutionintoEq.(10.98)yieldsourGreen’sfunc-
tion
G(x,t)=
−1
Au(x)v(t), a ≤x<t,
−1
Au(t)v(x), t <x ≤b.(10.103)
Notethat G(x,t)=G(t,x). Thisisthesymmetrypropertythatwas provedearlierinSec-
tion9.7.Itsphysicalinterpretationisgivenbythereciprocityprinciple(viaourpropagator
function)—a cause at tyields the same effect at xas a cause at xproduces at t.I nt e r m s
of our electrostaticanalogythis is obvious,the propagatorfunctiondependingonly on the
magnitudeofthedistancebetweenthetwopoints:
|r1−r2|=|r2−r1|.
Green’s Function Integral — Differential Equation
We have constructed G(x,t), but there still remains the task of showing that the integral
(Eq.(10.97))withournewGreen’sfunctionisindeedasolutionoftheoriginaldifferential
equation(10.92). This we do by direct substitution. With G(x,t)given by Eq. (10.103),18
Eq.(10.97) becomes
y(x)=−1
Aintegraldisplayx
av(x)u(t)f(t)dt −1
Aintegraldisplayb
xu(x)v(t)f(t)dt. (10.104)
18In thefirst integral, a≤t≤x. HenceG(x,t)=G2(x,t)=−(1/A)u(t)v(x) . Similarly, the second integral requires G=G1.
666 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Differentiating,weobtain
y′(x)=−1
Aintegraldisplayx
av′(x)u(t)f(t)dt −1
Aintegraldisplayb
xu′(x)v(t)f(t)dt, (10.105)
thederivativesofthelimitscanceling.Aseconddifferentiationyields
y′′(x)=−1
Aintegraldisplayx
av′′(x)u(t)f(t)dt −1
Aintegraldisplayb
xu′′(x)v(t)f(t)dt
−1
Abracketleftbig
u(x)v′(x)−v(x)u′(x)bracketrightbig
f(x). (10.106)
ByEqs. (10.100)and(10.102)thismayberewrittenas
y′′(x)=−v′′(x)
Aintegraldisplayx
au(t)f(t)dt−u′′(x)
Aintegraldisplayb
xv(t)f(t)dt−f(x)
p(x).(10.107)
Now,bysubstitutingintoEq. (10.93),wehave
Ly(x)=−Lv(x)
Aintegraldisplayx
au(t)f(t)dt−Lu(x)
Aintegraldisplayb
xv(t)f(t)dt−f(x). (10.108)
Sinceu(x)andv(x)were chosen to satisfy the homogeneous equation, the L-factors are
zeroandtheintegraltermsvanish,andwesee thatEq. (10.92)is satisfied.
Wemustalsocheckthat y(x)satisfiestherequiredboundaryconditions.Atpoint x=a,
y(a)=−u(a)
Aintegraldisplayb
av(t)f(t)dt=cu(a), (10.109)
y′(a)=−u′(a)
Aintegraldisplayb
av(t)f(t)dt=cu′(a), (10.110)
sincethedefiniteintegralisa constant.Wechose u(x)tosatisfy
αu(a)+βu′(a)=0. (10.111)
Multiplying by the constant c, we verify that y(x)also satisfies Eq. (10.111). This illus-
trates the utility of the homogeneous boundary conditions: The normalization does not
matter. In quantum mechanical problems the boundary condition on the wave function is
oftenexpressedintermsoftheratio
ψ′(x)
ψ(x)=d
dxlnψ(x), comparedtod
dxlnu(x)vextendsinglevextendsingle
x=a=−α
β,
Eq.(10.111). Theadvantageis thatthewavefunctionneednotbenormalizedyet.
Summarizing,wehaveEq. (10.97),
y(x)=integraldisplayb
aG(x,t)f(t)dt,
whichsatisfiesthedifferentialequation(Eq. (10.92)),
Ly(x)+f(x)=0,
10.5 Green’s Function — Eigenfunction Expansion 667
andtheboundaryconditions,theseboundaryconditionshavingbeenbuiltintotheGreen’s
function, G(x,t).
Basically, what we have done is to use the solutions of the homogeneous equation
Eq.(10.94)toconstructasolutionoftheinhomogeneousequation.Again,Poisson’sequa-
tionisanillustration.Thesolution(Eq.(9.148))representsaweighted [ρ(r2)]combination
of solutions of the corresponding homogeneous Laplace’s equation. (We followed these
samesteps earlyinthis section.)
It should be noted that our y(x), Eq. (10.97), is actually the particular solution of the
differential equation, Eq. (10.92). Our boundary conditions exclude the addition of so-
lutions of the homogeneous equation. In an actual physical problem we may well have
both types of solutions. In electrostatics, for instance (compare Section 9.7), the Green’s
functionsolutionofPoisson’sequationgivesthepotentialcreatedbythegivenchargedis-
tribution.Inaddition,theremaybeexternalfieldssuperimposed.Thesewouldbedescribed
bysolutionsofthehomogeneousequation,Laplace’sequation.
Eigenfunction, Eigenvalue Equation
Theprecedinganalysisplacednospecialrestrictionsonour f(x).Letusnowassumethat
f(x)=λρ(x)y(x) .19Thenwehave
y(x)=λintegraldisplayb
aG(x,t)ρ(t)y(t)dt (10.112)
asasolutionof
Ly(x)+λρ(x)y(x)=0 (10.113)
anditsboundaryconditions.Equation(10.112)isahomogeneousFredholmintegralequa-
tion of the second kind, and Eq. (10.113) is the homogeneous eigenvalue equation (with
theweightingfunction w(x)replacedby ρ(x)).
ThereisachangeintheinterpretationofourGreen’sfunction.Itstartedasapropagator
function, a weighting function giving the importance of the charge ρ(r2)in producing
the potential ϕ(r1). The charge ρwas the inhomogeneous term in the inhomogeneous
differential equation (10.92). Now the differential equation and the integral equation are
bothhomogeneous .G(x,t)has become a link relating the two equations, differential and
integral.
To complete the discussion of this differential equation–integral equation equivalence,
let us now show that Eq. (10.113) implies Eq. (10.112), that is, that a solution of our
differential equation (10.113) with its boundary conditions satisfies the integral equa-
tion (10.112). We multiply Eq. (10.113) by G(x,t), the appropriate Green’s function, and
integratefrom x=atox=btoobtain
integraldisplayb
aG(x,t) Ly(x)dx+λintegraldisplayb
aG(x,t)ρ(x)y(x)dx =0. (10.114)
19Thefunction ρ(x)is some weighting function, not acharge density.
668 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
Thefirstintegralissplitintwo (x<t,x>t) ,accordingtotheconstructionofourGreen’s
function,giving
−integraldisplayt
aG1(x,t)Ly(x)dx−integraldisplayb
tG2(x,t)Ly(x)dx=λintegraldisplayb
aG(x,t)ρ(x)y(x)dx. (10.115)
Note that tis the upper limit for the G1integrals and the lower limit for the G2integrals.
We are going to reduce the left-hand side of Eq. (10.115) to y(t). Then, with G(x,t)=
G(t,x), wehaveEq. (10.112)(with xandtinterchanged).
ApplyingGreen’stheoremtotheleft-handsideor,equivalently,integratingbyparts,we
obtain
−integraldisplayt
aG1(t,x)bracketleftbiggd
dxparenleftbigg
p(x)d
dxy(x)parenrightbigg
+q(x)y(x)bracketrightbigg
dx
=−bracketleftbig
G1(x,t)p(x)y′(x)bracketrightbigvextendsinglevextendsinglex=t
x=a+integraldisplayt
aparenleftbigg∂
∂xG1(x,t)parenrightbigg
p(x)y′(x)dx
−integraldisplayt
aG1(x,t)q(x)y(x)dx, (10.116)
withanequivalentexpressionforthesecondintegral.Asecondintegrationbyparts yields
−integraldisplayt
aG1(x,t)Ly(x)dx=−integraldisplayt
ay(x)LG1(x,t)dx
−bracketleftbig
G1(x,t)p(x)y′(x)bracketrightbigvextendsinglevextendsinglex=t
x=a
+bracketleftbig
G′
1(x,t)p(x)y(x)bracketrightbigvextendsinglevextendsinglex=t
x=a. (10.117)
The integral on the right vanishes because LG1=0. By combining the integrated terms
withthosefromintegrating G2,weha v e
−p(t)bracketleftbigg
G1(t,t)y′(t)−y(t)∂
∂xG1(x,t)vextendsinglevextendsingle
x=t−G2(t,t)y′(t)+y(t)∂
∂xG2(x,t)vextendsinglevextendsingle
x=tbracketrightbigg
+p(a)bracketleftbigg
y′(a)G1(a,t)−y(a)∂
∂xG1(x,t)vextendsinglevextendsingle
x=abracketrightbigg
−p(b)bracketleftbigg
G2(b,t)y′(b)−y(b)∂
∂xG2(x,t)vextendsinglevextendsingle
x=bbracketrightbigg
. (10.118)
Each of the last two expressions vanishes, for G(x,t)andy(x)satisfy the same boundary
conditions.Thefirstexpression,withthehelpofEqs.(10.95)and(10.96),reducesto y(t).
Substituting into Eq. (10.115), we have Eq. (10.112), thus completing the demonstration
of the equivalence of the integral equation and the differential equation plus boundary
conditions.
Example 10.5.1 LINEAR OSCILLATOR
Asasimpleexample,considerthelinearoscillatorequation(for avibratingstring):
y′′(x)+λy(x)=0. (10.119)
10.5 Green’s Function — Eigenfunction Expansion 669
We impose the conditions y(0)=y(1)=0, which correspond to a string clamped at both
ends.Now,toconstructourGreen’sfunction,weneedsolutionsofthehomogeneousequa-
tionLy(x)=0,whichis y′′(x)=0.Tosatisfytheboundaryconditions,wemusthaveone
solutionvanishat x=0,theotherat x=1.Suchsolutions(unnormalized)are
u(x)=x, v(x)=1−x. (10.120)
Wefindthat
uv′−vu′=−1 (10.121)
or,byEq. (10.101)with p(x)=1,A=−1.OurGreen’sfunctionbecomes
G(x,t)=braceleftBiggx(1−t),0≤x<t,
t(1−x), t<x≤1.(10.122)
HencebyEq.(10.112)ourclampedvibratingstringsatisfies
y(x)=λintegraldisplay1
0G(x,t)y(t)dt. (10.123)
YoumayshowthattheknownsolutionsofEq. (10.119),
y=sinnπx, λ=n2π2,
doindeedsatisfy Eq. (10.123).Notethatoureigenvalue λisnotthewavelength. /squaresolid
Green’s Function and the Dirac Delta Function
One more approach to the Green’s function may shed additional light on our formulation
and particularly on its relation to physical problems. Let us refer once more to Poisson’s
equation,thistimefor apointcharge:
∇2ϕ(r)=−ρpoint
ε0. (10.124)
The Green’s function solution of this equation was developed in Section 9.7. This time let
ustakeaone-dimensionalanalog
Ly(x)+f(x)point=0. (10.125)
Heref(x)pointrefers to a unit point “charge,” or a point force. We may represent it by a
numberofforms, butperhapsthemostconvenientis
f(x)point=
1
2ε,t−ε<x<t+ε,
0,elsewhere ,(10.126)
whichisessentiallythesameasEq. (1.172). Then,integratingEq. (10.125),wehave
integraldisplayt+ε
t−εLy(x)dx=−integraldisplayt+ε
t−εf(x)pointdx=−1 (10.127)
670 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
fromthedefinitionof f(x). Letus examine Ly(x)moreclosely.We have
integraldisplayt+ε
t−εd
dxbracketleftbig
p(x)y′(x)bracketrightbig
dx+integraldisplayt+ε
t−εq(x)y(x)dx
=vextendsinglevextendsinglep(x)y′(x)vextendsinglevextendsinglet+ε
t−ε+integraldisplayt+ε
t−εq(x)y(x)dx=−1. (10.128)
In the limit ε→0 we may satisfy this relation by permitting y′(x)to have a discontinu-
ity of−1/p(x)atx=t,y(x)itself remaining continuous.20These, however, are just the
properties used to define our Green’s function, G(x,t). In addition, we note that in the
limitε→0,
f(x)point=δ(x−t), (10.129)
inwhichδ(x−t)isourDiracdeltafunction,definedinthismannerinSection1.15.Hence
Eq.(10.125)has become
LG(x,t)=−δ(x−t). (10.130)
Thisisaone-dimensionalversionofEq.(9.159),whichweexploitforthedevelopmentof
Green’s functions in two and three dimensions—Section 9.7. It will be recalled that we
usedthisrelationinSection9.7todetermineourGreen’sfunctions.
Equation (10.130) could have been expected since it is actually a consequence of our
differential equation, Eq. (10.92), and Green’s function integral solution, Eq. (10.97). If
we let Lx(subscript to emphasize that it operates on the x-dependence) operate on both
sidesofEq. (10.97), then
Lxy(x)=Lxintegraldisplayb
aG(x,t)f(t)dt.
By Eq. (10.92) the left-hand side is just −f(x). On the right Lx, is independent of the
variableofintegration t, so wemaywrite
−f(x)=integraldisplayb
abraceleftbig
LxG(x,t)bracerightbig
f(t)dt.
BydefinitionofDirac’sdeltafunction,Eqs.(1.171b)and(1.183),wehaveEq.(10.130).
Exercises
10.5.1 Showthat
G(x,t)=braceleftBiggx,0≤x<t,
t, t<x≤1,
istheGreen’sfunctionfor theoperator L=d2/dx2andtheboundaryconditions
y(0)=0,y′(1)=0.
20The functions p(x)andq(x)appearing in the operator Lare continuous functions. With y(x)remaining continuous,integraltext
q(x)y(x)dx is certainly continuous. Hencethis integral over an interval 2 ε(Eq. (10.128)) vanishes as εvanishes.
10.5 Green’s Function — Eigenfunction Expansion 671
10.5.2 FindtheGreen’sfunctionfor
(a)Ly(x)=d2y(x)
dx2+y(x),braceleftBiggy(0)=0,
y′(1)=0.
(b)Ly(x)=d2y(x)
dx2−y(x), y(x) finitefor−∞<x<∞.
10.5.3 FindtheGreen’sfunctionfor theoperators
(a)Ly(x)=d
dxparenleftbigg
xdy(x)
dxparenrightbigg
.
ANS.G(x,t)=braceleftBigg−lnt,0≤x<t,
−lnx, t<x≤1.
(b)Ly(x)=d
dxparenleftbigg
xdy(x)
dxparenrightbigg
−n2
xy(x), withy(0)finiteand y(1)=0.
ANS.G(x,t)=
1
2nbracketleftbiggparenleftbiggx
tparenrightbiggn
−(xt)nbracketrightbigg
,0≤x<t,
1
2nbracketleftbiggparenleftbiggt
xparenrightbiggn
−(xt)nbracketrightbigg
,t<x≤1.
ThecombinationofoperatorandintervalspecifiedinExercise10.5.3(a)ispathological,
in that one of the endpoints of the interval (zero) is a singular point of the operator. As
a consequence, the integrated part (the surface integral of Green’s theorem) does not
vanish.Thenextfour exercisesexplorethissituation.
10.5.4 (a) Showthattheparticularsolutionof
d
dxbracketleftbigg
xd
dxy(x)bracketrightbigg
=−1
isyP(x)=−x.
(b) Showthat
yP(x)=−x/negationslash=integraldisplay1
0G(x,t)(−1)dt,
whereG(x,t)is theGreen’sfunctionof Exercise10.5.3(a).
10.5.5 Show that Green’s theorem, Eq. (1.104) in one dimension with a Sturm–Liouville-type
operator(d/dt)p(t)(d/dt) replacing ∇·∇, mayberewrittenas
integraldisplayb
abracketleftbigg
u(t)d
dtparenleftbigg
p(t)dv(t)
dtparenrightbigg
−v(t)d
dtparenleftbigg
p(t)du(t)
dtparenrightbiggbracketrightbigg
dt
=bracketleftbigg
u(t)p(t)dv(t)
dt−v(t)p(t)du(t)
dtbracketrightbiggvextendsinglevextendsinglevextendsinglevextendsingleb
a.
672 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.5.6 Usingtheone-dimensionalformof Green’stheoremofExercise10.5.5,let
v(t)=y(t)andd
dtparenleftbigg
p(t)dy(t)
dtparenrightbigg
=−f(t),
u(t)=G(x,t) andd
dtparenleftbigg
p(t)∂G(x,t)
∂tparenrightbigg
=−δ(x−t).
ShowthatGreen’stheoremyields
y(x)=integraldisplayb
aG(x,t)f(t)dt +bracketleftbigg
G(x,t)p(t)dy(t)
dt−y(t)p(t)∂
∂tG(x,t)bracketrightbiggvextendsinglevextendsinglevextendsinglevextendsinglet=b
t=a.
10.5.7 Forp(t)=t,y(t)=−t,
G(x,t)=braceleftBigg−lnt,0≤x<t
−lnx, t<x≤1,
verifythattheintegratedpartdoesnotvanish.
10.5.8 ConstructtheGreen’sfunctionfor
x2d2y
dx2+xdy
dx+parenleftbig
k2x2−1parenrightbig
y=0,
subjecttotheboundaryconditions
y(0)=0,y(1)=0.
10.5.9 Giventhat
L=parenleftbig
1−x2parenrightbigd2
dx2−2xd
dx
and
G(±1,t)remainsfinite ,
showthatnoGreen’sfunctioncanbeconstructedbythetechniquesofthissection.( u(x)
andv(x)are linearlydependent.)
10.5.10 Constructtheone-dimensionalGreen’sfunctionfor theHelmholtzequation
parenleftbiggd2
dx2+k2parenrightbigg
ψ(x)=g(x).
The boundary conditions are those for a wave advancing in the positive x-direction—
assumingatimedependence e−iwt.
ANS.G(x1,x2)=i
2kexpparenleftbig
ik|x1−x2|parenrightbig
.
10.5.11 Constructtheone-dimensionalGreen’sfunctionfor themodifiedHelmholtzequation
parenleftbiggd2
dx2−k2parenrightbigg
ψ(x)=f(x).
10.5 Green’s Function — Eigenfunction Expansion 673
The boundary conditions are that the Green’s function must vanish for x→∞and
x→−∞.
ANS.G(x1,x2)=1
2kexpparenleftbig
−k|x1−x2|parenrightbig
.
10.5.12 FromtheeigenfunctionexpansionoftheGreen’sfunctionshowthat
(a)2
π2∞summationdisplay
n=1sinnπxsinnπt
n2=braceleftBiggx(1−t),0≤x<t,
t(1−x), t<x≤1.
(b)2
π2∞summationdisplay
n=0sin(n+1
2)πxsin(n+1
2)πt
(n+1
2)2=braceleftBiggx,0≤x<t,
t, t<x≤1.
Note.InSection10.4theGreen’sfunctionof L+λisexpandedineigenfunctions.The
λthereis anadjustableparameter,notaneigenvalue.
10.5.13 IntheFredholmequation,
f(x)=λ2integraldisplayb
aG(x,t)ϕ(t)dt,
G(x,t)is aGreen’sfunctiongivenby
G(x,t)=∞summationdisplay
n=1ϕn(x)ϕn(t)
λ2n−λ2.
Showthatthesolutionis
ϕ(x)=∞summationdisplay
n=1λ2
n−λ2
λ2ϕn(x)integraldisplayb
af(t)ϕn(t)dt.
10.5.14 ShowthattheGreen’sfunctionintegraltransform operator
integraldisplayb
aG(x,t)[]dt
isequalto−L−1, inthesensethat
(a)Lxintegraldisplayb
aG(x,t)y(t)dt =−y(x),
(b)integraldisplayb
aG(x,t) Lty(t)dt=−y(x).
Note.T ak e Ly(x)+f(x)=0,Eq. (10.92).
10.5.15 SubstituteEq.(10.87),theeigenfunctionexpansionofGreen’sfunction,intoEq.(10.88)
and then show that Eq. (10.88) is indeed a solution of the inhomogeneous Helmholtz
equation(10.82).
674 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions
10.5.16 (a) Startingwithaone-dimensionalinhomogeneousdifferentialequation(Eq.(10.89)),
assume that ψ(x)andρ(x)may be represented by eigenfunction expansions.
WithoutanyuseoftheDiracdeltafunctionorits representations,showthat
ψ(x)=∞summationdisplay
n=0integraltextb
aρ(t)ϕn(t)dt
λn−λϕn(x).
Note that (1) if ρ=0, no solution exists unless λ=λnand (2) if λ=λn,n o
solutionexistsunless ρisorthogonalto ϕn.Thissamebehaviorwillreappearwith
integralequationsinSection16.4.
(b) Interchanging summation and integration, show that you have constructed the
Green’sfunctioncorrespondingtoEq. (10.90).
10.5.17 The eigenfunctions of the Schrödinger equation are often complex. In this case the
orthogonalityintegral,Eq. (10.40), isreplacedby
integraldisplayb
aϕ∗
i(x)ϕj(x)w(x)dx=δij.
InsteadofEq. (1.189),wehave
δ(r1−r2)=∞summationdisplay
n=0ϕn(r1)ϕ∗
n(r2).
ShowthattheGreen’sfunction,Eq.(10.87), becomes
G(r1,r2)=∞summationdisplay
n=0ϕn(r1)ϕ∗
n(r2)
k2n−k2=G∗(r2,r1).
AdditionalReadings
Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics . Reading, MA: Addison-
Wesley (1969).
Dennery, P.,andA.Krzywicki, Mathematics for Physicists .Reprinted. NewYork: Dover (1996).
Hirsch,M., DifferentialEquations,DynamicalSystems,andLinearAlgebra .SanDiego:AcademicPress(1974).
Miller,K.S., Linear Differential Equations in the RealDomain . New York: Norton (1963).
Titchmarsh, E. C., Eigenfunction Expansions Associated with Second-Order Differential Equations , 2nd ed.,
Vol. 1.London: Oxford University Press (1962), Vol.II (1958).
CHAPTER 11
BESSEL FUNCTIONS
11.1 B ESSEL FUNCTIONS OF THE FIRST KIND,Jν(x)
Bessel functions appear in a wide variety of physical problems. In Section 9.3, separa-
tion of the Helmholtz, or wave, equation in circular cylindrical coordinates led to Bessel’s
equation. In Section 11.7 we will see that the Helmholtz equation in spherical polar co-
ordinates also leads to a form of Bessel’s equation. Bessel functions may also appear in
integral form—integral representations. This may result from integral transforms (Chap-
ter 15) or from the mathematical elegance of starting the study of Bessel functions with
Hankelfunctions,Section11.4.
Besselfunctionsandcloselyrelatedfunctionsformarichareaofmathematicalanalysis
with many representations, many interesting and useful properties, and many interrela-
tions. Some of the major interrelations are developed in Section 11.1 and in succeeding
sections.NotethatBesselfunctionsarenotrestrictedtoChapter11.Theasymptoticforms
are developed in Section 7.3 as well as in Section 11.6. The confluent hypergeometric
representationsappearinSection13.5.
Generating Function for Integral Order
AlthoughBesselfunctionsareofinterestprimarilyassolutionsofdifferentialequations,it
isinstructiveandconvenienttodevelopthemfromacompletelydifferentapproach,thatof
thegeneratingfunction.1Thisapproachalsohastheadvantageoffocusingonthefunctions
themselvesratherthanonthedifferentialequationstheysatisfy.Letusintroduceafunction
oftwovariables,
g(x,t)=e(x/2)(t−1/t). (11.1)
1Generating functions have already been used in Chapter 5. In Section 5.6 the generating function (1+x)nwas used to derive
the binomial coefficients. InSection5.9the generating function x(ex−1)−1wasused to derive the Bernoulli numbers.
675
676 Chapter 11 Bessel Functions
ExpandingthisfunctioninaLaurentseries (Section6.5), weobtain
e(x/2)(t−1/t)=∞summationdisplay
n=−∞Jn(x)tn. (11.2)
Itis instructivetocompareEq.(11.2) withtheequivalentEqs. (11.23)and(11.25).
Thecoefficientof tn,Jn(x),isdefinedtobeaBesselfunctionofthefirstkind,ofintegral
ordern. Expanding the exponentials, we have a product of Maclaurin series in xt/2 and
−x/2t,respectively,
ext/2·e−x/2t=∞summationdisplay
r=0parenleftbiggx
2parenrightbiggrtr
r!∞summationdisplay
s=0(−1)sparenleftbiggx
2parenrightbiggst−s
s!. (11.3)
Here,thesummationindex rischangedto n,withn=r−sandsummationlimits n=−s
to∞, and the order of the summations is interchanged, which is justified by absolute
convergence.Therangeofthesummationover nbecomes−∞to∞,whilethesummation
oversextendsfrom max (−n,0)to∞.For a given swegettn(n≥0)fromr=n+s:
parenleftbiggx
2parenrightbiggn+stn+s
(n+s)!(−1)sparenleftbiggx
2parenrightbiggst−s
s!. (11.4)
Thecoefficientof tnisthen2
Jn(x)=∞summationdisplay
s=0(−1)s
s!(n+s)!parenleftbiggx
2parenrightbiggn+2s
=xn
2nn!−xn+2
2n+2(n+1)!+···. (11.5)
ThisseriesformexhibitsisbehavioroftheBesselfunction Jn(x)forsmall xandpermits
numericalevaluationof Jn(x). The results for J0,J1, andJ2are shown in Fig. 11.1. From
Section 5.3 the error in using only a finite number of terms of this alternating series in
numerical evaluation is less than the first term omitted. For instance, if we want Jn(x)
FIGURE 11.1Besselfunctions, J0(x),J1(x), andJ2(x).
2Fromthestepsleadingtothisseriesandfromitsconvergencecharacteristicsitshouldbeclearthatthisseriesmaybeusedwith
xreplacedby zandwith zany point in the finite complex plane.
11.1 Bessel Functions of the First Kind, Jν(x) 677
to±1% accuracy, the first term alone of Eq. (11.5) will suffice, provided the ratio of the
second term to the first is less than 1% (in magnitude) or x<0.2(n+1)1/2. The Bessel
functionsoscillatebutare notperiodic—exceptinthelimitas x→∞(Section11.6).The
amplitudeof Jn(x)isnotconstantbutdecreasesasymptoticallyas x−1/2.(SeeEq.(11.137)
forthis envelope.)
Forn<0,Eq.(11.5) gives
J−n(x)=∞summationdisplay
s=0(−1)s
s!(s−n)!parenleftbiggx
2parenrightbigg2s−n
. (11.6)
Sincenisaninteger(here), (s−n)!→∞fors=0,...,(n−1).Hencetheseriesmaybe
consideredtostartwith s=n.Replacing sbys+n,weobtain
J−n(x)=∞summationdisplay
s=0(−1)s+n
s!(s+n)!parenleftbiggx
2parenrightbiggn+2s
, (11.7)
showingimmediatelythat Jn(x)andJ−n(x)arenotindependentbutarerelatedby
J−n(x)=(−1)nJn(x) (integraln). (11.8)
These series expressions (Eqs. (11.5) and (11.6)) may be used with nreplaced by νto
defineJν(x)andJ−ν(x)for nonintegral ν(compareExercise11.1.7).
Recurrence Relations
The recurrence relations for Jn(x)and its derivatives may all be obtained by operating
on the series, Eq. (11.5), although this requires a bit of clairvoyance (or a lot of trial and
error). Verification of the known recurrence relations is straightforward, Exercise 11.1.7.
Here it is convenient to obtain them from the generating function, g(x,t). Differentiating
bothsidesof Eq. (11.1)withrespectto t, wefindthat
∂
∂tg(x,t)=1
2xparenleftbigg
1+1
t2parenrightbigg
e(x/2)(t−1/t)
=∞summationdisplay
n=−∞nJn(x)tn−1, (11.9)
andsubstitutingEq.(11.2)fortheexponentialandequatingthecoefficientsoflikepowers
oft,3weobtain
Jn−1(x)+Jn+1(x)=2n
xJn(x). (11.10)
This is a three-term recurrence relation. Given J0andJ1, for example, J2(and any other
integralorder Jn)maybecomputed.
3This depends onthe factthat thepower-series representation is unique (Sections 5.7 and6.5).
678 Chapter 11 Bessel Functions
DifferentiatingEq. (11.1)withrespectto x,weha v e
∂
∂xg(x,t)=1
2parenleftbigg
t−1
tparenrightbigg
e(x/2)(t−1/t)=∞summationdisplay
n=−∞J′
n(x)tn. (11.11)
Again,substitutinginEq.(11.2)andequatingthecoefficientsoflikepowersof t,weobtain
theresult
Jn−1(x)−Jn+1(x)=2J′
n(x). (11.12)
Asaspecialcaseof thisgeneralrecurrencerelation,
J′
0(x)=−J1(x). (11.13)
AddingEqs.(11.10) and(11.12) anddividingby2,wehave
Jn−1(x)=n
xJn(x)+J′
n(x). (11.14)
Multiplyingby xnandrearrangingtermsproduces
d
dxbracketleftbig
xnJn(x)bracketrightbig
=xnJn−1(x). (11.15)
SubtractingEq. (11.12)fromEq. (11.10)anddividingby2yields
Jn+1(x)=n
xJn(x)−J′
n(x). (11.16)
Multiplyingby x−nandrearrangingterms, weobtain
d
dxbracketleftbig
x−nJn(x)bracketrightbig
=−x−nJn+1(x). (11.17)
Bessel’s Differential Equation
Suppose we consider a set of functions Zν(x)that satisfies the basic recurrence relations
(Eqs. (11.10) and (11.12)), but with νnot necessarily an integer and Zνnot necessarily
givenbytheseries (Eq. (11.5)). Equation(11.14)mayberewritten (n→ν)as
xZ′
ν(x)=xZν−1(x)−νZν(x). (11.18)
Ondifferentiatingwithrespectto x,weha v e
xZ′′
ν(x)+(ν+1)Z′
ν−xZ′
ν−1−Zν−1=0. (11.19)
Multiplyingby xandthensubtractingEq. (11.18)multipliedby νgivesus
x2Z′′
ν+xZ′
ν−ν2Zν+(ν−1)xZν−1−x2Z′
ν−1=0. (11.20)
NowwerewriteEq. (11.16)andreplace nbyν−1:
xZ′
ν−1=(ν−1)Zν−1−xZν. (11.21)
UsingEq. (11.21)toeliminate Zν−1andZ′
ν−1fromEq. (11.20), wefinallyget
x2Z′′
ν+xZ′
ν+parenleftbig
x2−ν2parenrightbig
Zν=0, (11.22)
11.1 Bessel Functions of the First Kind, Jν(x) 679
which is Bessel’s ODE. Hence any functions Zν(x)that satisfy the recurrence relations
(Eqs. (11.10) and (11.12), (11.14) and (11.16), or (11.15) and (11.17)) satisfy Bessel’s
equation; that is, the unknown Zνare Bessel functions. In particular, we have shown that
thefunctions Jn(x), definedbyourgeneratingfunction,satisfyBessel’sODE.If theargu-
mentiskρratherthan x,Eq.(11.22) becomes
ρ2d2
dρ2Zν(kρ)+ρd
dρZν(kρ)+parenleftbig
k2ρ2−ν2parenrightbig
Zν(kρ)=0. (11.22a)
Integral Representation
A particularly useful and powerful way of treating Bessel functions employs integral rep-
resentations.Ifwereturntothegeneratingfunction(Eq.(11.2)),andsubstitute t=eiθ,we
get
eixsinθ=J0(x)+2bracketleftbig
J2(x)cos2θ+J4(x)cos4θ+···bracketrightbig
+2ibracketleftbig
J1(x)sinθ+J3(x)sin3θ+···bracketrightbig
, (11.23)
inwhichwehaveusedtherelations
J1(x)eiθ+J−1(x)e−iθ=J1(x)parenleftbig
eiθ−e−iθparenrightbig
=2iJ1(x)sinθ, (11.24)
J2(x)e2iθ+J−2(x)e−2iθ=2J2(x)cos2θ,
andso on.
In summationnotation,
cos(xsinθ)=J0(x)+2∞summationdisplay
n=1J2n(x)cos(2nθ),
(11.25)
sin(xsinθ)=2∞summationdisplay
n=1J2n−1(x)sinbracketleftbig
(2n−1)θbracketrightbig
,
equatingrealandimaginarypartsofEq. (11.23).
Byemployingtheorthogonalitypropertiesofcosineandsine,4
integraldisplayπ
0cosnθcosmθdθ=π
2δnm, (11.26a)
integraldisplayπ
0sinnθsinmθdθ=π
2δnm, (11.26b)
4They are eigenfunctions of a self-adjoint equation (linear oscillator equation) and satisfy appropriate boundary conditions
(compare Sections 10.2 and 14.1).
680 Chapter 11 Bessel Functions
inwhich nandmarepositiveintegers(zerois excluded),5weobtain
1
πintegraldisplayπ
0cos(xsinθ)cosnθdθ=braceleftbiggJn(x), n even,
0,n odd,(11.27)
1
πintegraldisplayπ
0sin(xsinθ)sinnθdθ=braceleftbigg0,n even,
Jn(x), n odd.(11.28)
If thesetwo equationsareaddedtogether,
Jn(x)=1
πintegraldisplayπ
0bracketleftbig
cos(xsinθ)cosnθ+sin(xsinθ)sinnθbracketrightbig
dθ
=1
πintegraldisplayπ
0cos(nθ−xsinθ)dθ, n=0,1,2,3,.... (11.29)
Asaspecialcase(integrateEq. (11.25)over (0,π)toget)
J0(x)=1
πintegraldisplayπ
0cos(xsinθ)dθ. (11.30)
Noting that cos (xsinθ)repeats itself in all four quadrants, we may write Eq. (11.30)
as
J0(x)=1
2πintegraldisplay2π
0cos(xsinθ)dθ. (11.30a)
Ontheotherhand, sin (xsinθ)reversesits signinthethirdandfourthquadrants,so
1
2πintegraldisplay2π
0sin(xsinθ)dθ=0. (11.30b)
Adding Eq. (11.30a) and itimes Eq. (11.30b), we obtain the complex exponential repre-
sentation
J0(x)=1
2πintegraldisplay2π
0eixsinθdθ=1
2πintegraldisplay2π
0eixcosθdθ. (11.30c)
This integral representation (Eq. (11.29)) may be obtained somewhat more directly by
employing contour integration (compare Exercise 11.1.16).6Many other integral repre-
sentationsexist(compareExercise11.1.18).
Example 11.1.1 FRAUNHOFER DIFFRACTION ,CIRCULAR APERTURE
Inthetheoryof diffractionthroughacircularapertureweencountertheintegral
/Phi1∼integraldisplaya
0rdrintegraldisplay2π
0eibrcosθdθ (11.31)
5Equations (11.26a) and (11.26b) hold for either morn=0. If both mandn=0, the constant in (11.26a) becomes π;t h e
constant in Eq.(11.26b) becomes 0.
6Forn=0 a simple integration over θfrom 0 to 2 πwillconvert Eq.(11.23) into Eq. (11.30c).
11.1 Bessel Functions of the First Kind, Jν(x) 681
FIGURE 11.2Fraunhoferdiffraction,circularaperture.
for/Phi1,theamplitudeofthediffractedwave.7Hereθisanazimuthangleintheplaneofthe
circular aperture of radius a, andαis the angle defined by a point on a screen below the
circular aperture relative to the normal through the center point. The parameter bis given
by
b=2π
λsinα, (11.32)
withλthe wavelength of the incident wave. The other symbols are defined by Fig. 11.2.
FromEq. (11.30c)weget8
/Phi1∼2πintegraldisplaya
0J0(br)rdr. (11.33)
Equation(11.15)enablesus tointegrateEq. (11.33)immediatelytoobtain
/Phi1∼2πab
b2J1(ab)∼λa
sinαJ1parenleftbigg2πa
λsinαparenrightbigg
. (11.34)
Noteherethat J1(0)=0.Theintensityofthelightinthediffractionpatternisproportional
to/Phi12and
/Phi12∼braceleftbiggJ1[(2πa/λ)sinα]
sinαbracerightbigg2
. (11.35)
7Theexponent ibrcosθgivesthephaseofthewaveonthedistantscreenatangle αrelativetothephaseofthewaveincidenton
the aperture at the point (r,θ). The imaginary exponential form of this integrand means that the integral is technically a Fourier
transform, Chapter 15. In general, theFraunhofer diffraction pattern is given by theFourier transform of the aperture.
8Wecould alsoreferto Exercise11.1.16(b).
682 Chapter 11 Bessel Functions
Table 11.1 Zeros oftheBesselFunctionsandTheirFirstDerivatives
Numberof zero J0(x) J 1(x) J 2(x) J 3(x) J 4(x) J 5(x)
12 .4048 3 .8317 5 .1356 6 .3802 7 .5883 8 .7715
25 .5201 7 .0156 8 .4172 9 .7610 11 .0647 12 .3386
38 .6537 10 .1735 11 .6198 13 .0152 14 .3725 15 .7002
41 1.7915 13 .3237 14 .7960 16 .2235 17 .6160 18 .9801
51 4.9309 16 .4706 17 .9598 19 .4094 20 .8269 22 .2178
J′
0(x)aJ′
1(x) J′
2(x) J′
3(x)
13 .8317 1 .8412 3 .0542 4 .2012
27 .0156 5 .3314 6 .7061 8 .0152
31 0.1735 8 .5363 9 .9695 11 .3459
aJ′
0(x)=−J1(x).
FromTable11.1,whichliststhezerosoftheBesselfunctionsandtheirfirstderivatives,9
Eq.(11.35) willhaveazeroat
2πa
λsinα=3.8317..., (11.36)
or
sinα=3.8317λ
2πa. (11.37)
Forgreenlight, λ=5.5×10−5cm.Hence,if a=0.5c m ,
α≈sinα=6.7×10−5(radian)≈14secondsofarc , (11.38)
whichshowsthatthebendingorspreadingofthelightrayisextremelysmall.Ifthisanaly-
sis had been known in the seventeenth century, the arguments against the wave theory of
light would have collapsed. In mid-twentieth century this same diffraction pattern appears
in the scattering of nuclear particles by atomic nuclei—a striking demonstration of the
wavepropertiesofthenuclearparticles. /squaresolid
A further example of the use of Bessel functions and their roots is provided by the
electromagnetic resonant cavity (Example 11.1.2) and the example and exercises of Sec-
tion11.2.
Example 11.1.2 CYLINDRICAL RESONANT CAVITY
The propagation of electromagnetic waves in hollow metallic cylinders is important in
many practical devices. If the cylinder has end surfaces, it is called a cavity. Resonant
cavitiesplayacrucialroleinmanyparticleaccelerators.
9Additional roots of the Bessel functions and their first derivatives may be found in C. L. Beattie, Table of first 700 zeros of
Bessel functions. Bell Syst. Tech. J. 37: 689 (1958), and Bell Monogr. 3055. Roots may be accessed in Mathematica and other
symbolic software and are on the Web.
11.1 Bessel Functions of the First Kind, Jν(x) 683
FIGURE 11.3Cylindricalresonant
cavity.
Wetakethe z-axisalongthecenterofthecavitywithendsurfacesat z=0andz=land
use cylindricalcoordinates suggested by the geometry. Its walls are perfect conductors, so
thetangentialelectricfieldvanishesonthem(as inFig.11.3):
Ez=0=Eϕforρ=a, E ρ=0=Eϕforz=0,l.
Inside the cavity we have a vacuum, so ε0µ0=1/c2.In the interior of a resonant cav-
ity, electromagnetic waves oscillate with harmonic time dependence e−iωt,which follows
from separating the time from the spatial variables in Maxwell’s equations (Section 1.9),
so
∇×∇×E=−1
c2∂2E
∂t2=α2E,α=ω
c.
With∇·E=0 (vacuum, no charges) and Eq. (1.85), we obtain for the space part of the
electricfield
∇2E+α2E=0,
which is called the vector Helmholtz PDE .T h ez-component ( Ez, space part only) satis-
fiesthescalarHelmholtzequation,
∇2Ez+α2Ez=0. (11.39)
The transverse electric field components E⊥=(Eρ,Eϕ)obey the same PDE but different
boundaryconditions,givenearlier.Once Ezisknown,Maxwell’sequationsdetermine Eϕ
fully.SeeJackson, Electrodynamics inAdditionalReadingsfor details.
684 Chapter 11 Bessel Functions
We separate the zvariable from ρandϕ,because there are no mixed derivatives∂2Ez
∂z∂ρ,
etc.Theproductsolution, Ez=v(ρ,ϕ)w(z), issubstitutedintotheHelmholtzPDEfor Ez
usingEq. (2.35) for ∇2incylindricalcoordinates,andthenwedivideby vw,yielding
1
w(z)d2w
dz2+1
vparenleftbigg∂2v
∂ρ2+1
ρ∂v
∂ρ+1
ρ2∂2v
∂ϕ2+α2parenrightbigg
v(ρ,ϕ)=0.
Thisimplies
−1
w(z)d2w
dz2=1
v(ρ,ϕ)parenleftbigg∂2v
∂ρ2+1
ρ∂v
∂ρ+1
ρ2∂2v
∂ϕ2+α2vparenrightbigg
=k2.
Here,k2isaseparationconstant,becausetheleft-andright-handsidesdependondifferent
variables.For w(z)wefindtheharmonicoscillatorODEwithstandingwavesolution(not
transients)thatweseek,
w(z)=Asinkz+Bcoskz,
withA,Bconstants.For v(ρ,ϕ)weobtain
∂2v
∂ρ2+1
ρ∂v
∂ρ+1
ρ2∂2v
∂ϕ2+γ2v=0,γ2=α2−k2.
In this PDE we can separate the ρandϕvariables, because there is no mixed term∂2v
∂ρ∂ϕ.
Theproductform v=u(ρ)/Phi1(ϕ) yields
ρ2
u(ρ)parenleftbiggd2u
dρ2+1
ρdu
dρ+γ2parenrightbigg
=−1
/Phi1(ϕ)d2/Phi1
dϕ2=m2,
wherethe separationconstant m2mustbeaninteger ,becausetheangularsolution /Phi1=
eimϕoftheODE
d2/Phi1
dϕ2+m2/Phi1=0
mustbeperiodicintheazimuthalangle.
This leavesuswiththeradialODE
d2u
dρ2+1
ρdu
dρ+parenleftbigg
γ2−m2
ρ2parenrightbigg
u=0.
Dimensionalargumentssuggestrescaling ρ→r=γρanddividingby γ2,whichyields
d2u
dr2+1
rdu
dr+parenleftbigg
1−m2
r2parenrightbigg
u=0.
ThisisBessel’sODEfor ν=m.Weusetheregularsolution Jm(γρ)becausethe(irregular)
second independent solution is singular at the origin, which is unacceptable here. The
completesolutionis
Ez=Jm(γρ)eimϕ(Asinkz+Bcoskz), (11.40a)
wheretheconstant γisdeterminedfromthe boundarycondition Ez=0onthecavitysur-
faceρ=a,thatis,that γabearootoftheBesselfunction Jm(seeTable11.1).Thisgives
adiscreteset ofvalues γ=γmn,wherendesignatesthe nthrootof Jm(see Table11.1).
11.1 Bessel Functions of the First Kind, Jν(x) 685
ForthetransversemagneticorTMmodeofoscillationwith Hz=0Maxwell’sequations
imply. (See again Resonant Cavities in J. D. Jackson’s Electrodynamics in Additional
Readings.)
E⊥∼∇⊥∂Ez
∂z,∇⊥=parenleftbigg∂
∂ρ,1
ρ∂
∂ϕparenrightbigg
.
Theformofthisresultsuggests Ez∼coskz,thatis,setting A=0 sothatE⊥∼sinkz=0
atz=0,lcanbesatisfiedby
k=pπ
l,p=0,1,2,.... (11.41)
Thus,the tangential electricfields EρandEϕvanishat z=0andl.Inotherwords, A=0
correspondsto dEz/dz=0a tz=0 andz=lfor theTM mode.Altogetherthen,wehave
γ2=ω2
c2−k2=ω2
c2−p2π2
l2, (11.42)
with
γ=γmn=αmn
a, (11.43)
whereαmnis thenthzeroof Jm. Thegeneralsolution
Ez=summationdisplay
m,n,pJm(γmnρ)e±imϕBmnpcospπz
l, (11.40b)
withconstants Bmnp,nowfollowsfromthesuperpositionprinciple.
The result of the two boundary conditions and the separation constant m2is that the
angularfrequencyofouroscillationdependsonthreediscreteparameters:
ωmnp=cradicalBigg
α2mn
a2+p2π2
l2,
m=0,1,2,...,
n=1,2,3,...,
p=0,1,2....(11.44)
ThesearetheallowableresonantfrequenciesforourTMmode.TheTEmodeofoscillation
isthetopicof Exercise11.1.26. /squaresolid
Alternate Approaches
Bessel functions are introduced here by means of a generating function, Eq. (11.2). Other
approachesarepossible.Listingthevariouspossibilities,wehave:
1. Generatingfunction(magic),Eq. (11.2).
2. SeriessolutionofBessel’s differentialequation,Section9.5.
3. Contour integrals: Some writers prefer to start with contour integral definitions of the
Hankel functions, Section 7.3 and 11.4, and develop the Bessel function Jν(x)from
theHankelfunctions.
686 Chapter 11 Bessel Functions
4. Direct solution of physical problems: Example 11.1.1. Fraunhofer diffraction with a
circular aperture, illustrates this. Incidentally, Eq. (11.31) can be treated by series ex-
pansion,ifdesired.Feynman10developsBesselfunctionsfromaconsiderationofcav-
ityresonators.
In case the generating function seems too arbitrary, it can be derived from a contour inte-
gral, Exercise 11.1.16, or from the Bessel function recurrence relations, Exercise 11.1.6.
Notethatthecontourintegralisnotlimitedtointeger ν,thusprovidingastartingpointfor
developingBesselfunctions.
Bessel Functions of Nonintegral Order
These different approaches are not exactly equivalent. The generating function approach
is very convenient for deriving two recurrence relations, Bessel’s differential equation,
integral representations, addition theorems (Exercise 11.1.2), and upper and lower bounds
(Exercise 11.1.1). However, you will probably have noticed that the generating function
defined only Bessel functions of integral order, J0,J1,J2, and so on. This is a limitation
of the generating function approach that can be avoided by using the contour integral in
Exercise11.1.16instead,thusleadingtoforegoingapproach(3).ButtheBesselfunctionof
thefirstkind, Jν(x),mayeasilybedefinedfornonintegral νbyusingtheseries(Eq.(11.5))
asanewdefinition.
Therecurrencerelationsmaybeverifiedbysubstitutingintheseriesformof Jν(x)(Ex-
ercise11.1.7).FromtheserelationsBessel’sequationfollows.Infact,if νisnotaninteger,
there is actually an important simplification. It is found that JνandJ−νare independent,
fornorelationoftheformofEq.(11.8)exists.Ontheotherhand,for ν=n,aninteger,we
need another solution. The development of this second solution and an investigation of its
propertiesform thesubjectofSection11.3.
Exercises
11.1.1 Fromtheproductofthegeneratingfunctions g(x,t)·g(x,−t)showthat
1=bracketleftbig
J0(x)bracketrightbig2+2bracketleftbig
J1(x)bracketrightbig2+2bracketleftbig
J2(x)bracketrightbig2+···
andthereforethat |J0(x)|≤1 and|Jn(x)|≤1/√
2,n=1,2,3,....
Hint.Useuniquenessofpowerseries, Section5.7.
11.1.2 Usingageneratingfunction g(x,t)=g(u+v,t)=g(u,t)·g(v,t), showthat
(a)Jn(u+v)=∞summationdisplay
s=−∞Js(u)·Jn−s(v),
(b)J0(u+v)=J0(u)J0(v)+2∞summationdisplay
s=1Js(u)J−s(v).
10R. P. Feynman, R. B. Leighton, and M. Sands, The Feynman Lectures on Physics , Vol. II. Reading, MA: Addison-Wesley
(1964), Chapter 23.
11.1 Bessel Functions of the First Kind, Jν(x) 687
Theseareadditiontheoremsfor theBesselfunctions.
11.1.3 Usingonlythegeneratingfunction
e(x/2)(t−1/t)=∞summationdisplay
n=−∞Jn(x)tn
andnottheexplicitseriesformof Jn(x),showthat Jn(x)hasoddorevenparityaccord-
ingtowhether nisoddoreven,thatis,11
Jn(x)=(−1)nJn(−x).
11.1.4 DerivetheJacobi–Angerexpansion
eizcosθ=∞summationdisplay
m=−∞imJm(z)eimθ.
Thisis anexpansionof aplanewaveinaseries ofcylindricalwaves.
11.1.5 Showthat
(a) cosx=J0(x)+2∞summationdisplay
n=1(−1)nJ2n(x),
(b) sinx=2∞summationdisplay
n=0(−1)nJ2n+1(x).
11.1.6 To help remove the generating function from the realm of magic, show that it can be
derivedfrom therecurrencerelation,Eq.(11.10).
Hint.
(a) Assumeageneratingfunctionof theform
g(x,t)=∞summationdisplay
m=−∞Jm(x)tm.
(b) MultiplyEq.(11.10) by tnandsumover n.
(c) Rewritetheprecedingresultas
parenleftbigg
t+1
tparenrightbigg
g(x,t)=2t
x∂g(x,t)
∂t.
(d) Integrate and adjust the “constant” of integration (a function of x) so that the
coefficientof thezerothpower, t0,isJ0(x), as givenbyEq.(11.5).
11.1.7 Show,bydirectdifferentiation,that
Jν(x)=∞summationdisplay
s=0(−1)s
s!(s+ν)!parenleftbiggx
2parenrightbiggν+2s
11This is easily seenfrom theseries form (Eq. (11.5)).
688 Chapter 11 Bessel Functions
satisfiesthetworecurrencerelations
Jν−1(x)+Jν+1(x)=2ν
xJν(x),
Jν−1(x)−Jν+1(x)=2J′
ν(x),
andBessel’s differentialequation
x2J′′
ν(x)+xJ′
ν(x)+parenleftbig
x2−ν2parenrightbig
Jν(x)=0.
11.1.8 Provethat
sinx
x=integraldisplayπ/2
0J0(xcosθ)cosθdθ,1−cosx
x=integraldisplayπ/2
0J1(xcosθ)dθ.
Hint.Thedefiniteintegral
integraldisplayπ/2
0cos2s+1θdθ=2·4·6···(2s)
1·3·5···(2s+1)
maybeuseful.
11.1.9 Showthat
J0(x)=2
πintegraldisplay1
0cosxt√
1−t2dt.
This integral is a Fourier cosine transform (compare Section 15.3). The corresponding
Fouriersinetransform,
J0(x)=2
πintegraldisplay∞
1sinxt√
t2−1dt,
is established in Section 11.4 (Exercise 11.4.6) using a Hankel function integral repre-
sentation.
11.1.10 Derive
Jn(x)=(−1)nxnparenleftbigg1
xd
dxparenrightbiggn
J0(x).
Hint.Trymathematicalinduction.
11.1.11 Show that between any two consecutive zeros of Jn(x)there is one and only one zero
ofJn+1(x).
Hint.Equations(11.15)and(11.17)maybeuseful.
11.1.12 An analysis of antenna radiation patterns for a system with a circular aperture involves
theequation
g(u)=integraldisplay1
0f(r)J0(ur)rdr.
Iff(r)=1−r2, showthat
g(u)=2
u2J2(u).
11.1 Bessel Functions of the First Kind, Jν(x) 689
11.1.13 The differential cross section in a nuclear scattering experiment is given by dσ/d/Omega1=
|f(θ)|2. Anapproximatetreatmentleadsto
f(θ)=−ik
2πintegraldisplay2π
0integraldisplayR
0exp[ikρsinθsinϕ]ρdρdϕ.
Hereθis an angle through which the scattered particle is scattered. Ris the nuclear
radius.Showthat
dσ
d/Omega1=parenleftbig
πR2parenrightbig1
πbracketleftbiggJ1(kRsinθ)
sinθbracketrightbigg2
.
11.1.14 Asetof functions Cn(x)satisfiestherecurrencerelations
Cn−1(x)−Cn+1(x)=2n
xCn(x),
Cn−1(x)+Cn+1(x)=2C′
n(x).
(a) Whatlinearsecond-orderODEdoesthe Cn(x)satisfy?
(b) By a change of variable transform your ODE into Bessel’s equation. This sug-
gests that Cn(x)may be expressed in terms of Bessel functions of transformed
argument.
11.1.15 A particle (mass m) is contained in a right circular cylinder (pillbox) of radius Rand
heightH.TheparticleisdescribedbyawavefunctionsatisfyingtheSchrödingerwave
equation
−¯h2
2m∇2ψ(ρ,ϕ,z)=Eψ(ρ,ϕ,z)
andtheconditionthatthewavefunctiongotozerooverthesurfaceofthepillbox.Find
thelowest(zeropoint)permittedenergy.
ANS.E=¯h2
2mbracketleftbiggparenleftbiggzpq
Rparenrightbigg2
+parenleftbiggnπ
Hparenrightbigg2bracketrightbigg
,
Emin=¯h2
2mbracketleftbiggparenleftbigg2.405
Rparenrightbigg2
+parenleftbiggπ
Hparenrightbigg2bracketrightbigg
,
wherezpqis theqthzeroof Jpandtheindex pis fixedbytheazimuthaldependence.
11.1.16 (a) Showbydirectdifferentiationandsubstitutionthat
Jν(x)=1
2πiintegraldisplay
Ce(x/2)(t−1/t)t−ν−1dt
orthattheequivalentequation,
Jν(x)=1
2πiparenleftbiggx
2parenrightbiggνintegraldisplay
es−x2/4ss−ν−1ds,
satisfies Bessel’s equation. Cis the contour shown in Fig. 11.4. The negative real
axisisthecutline.
690 Chapter 11 Bessel Functions
FIGURE 11.4Besselfunctioncontour.
Hint.Showthatthetotalintegrand(aftersubstitutinginBessel’sdifferentialequa-
tion)maybewrittenas atotalderivative:
d
dtbraceleftbigg
expbracketleftbiggx
2parenleftbigg
t−1
tparenrightbiggbracketrightbigg
t−νbracketleftbigg
ν+x
2parenleftbigg
t+1
tparenrightbiggbracketrightbiggbracerightbigg
.
(b) Showthatthefirst integral(with naninteger)maybetransformedinto
Jn(x)=1
2πintegraldisplay2π
0ei(xsinθ−nθ)dθ=i−n
2πintegraldisplay2π
0ei(xcosθ+nθ)dθ.
11.1.17 The contour Cin Exercise 11.1.16 is deformed to the path −∞to−1, unit circle e−iπ
toeiπ,andfinally−1t o−∞. Showthat
Jν(x)=1
πintegraldisplayπ
0cos(νθ−xsinθ)dθ−sinνπ
πintegraldisplay∞
0e−νθ−xsinhθdθ.
Thisis Bessel’s integral.
Hint.Thenegativevaluesofthevariableof integration umaybehandledbyusing
u=te±ix.
11.1.18 (a) Showthat
Jν(x)=2
π1/2(ν−1
2)!parenleftbiggx
2parenrightbiggνintegraldisplayπ/2
0cos(xsinθ)cos2νθdθ,
whereν>−1
2.
Hint. Here is a chance to use series expansion and term-by-term integration. The
formulasofSection8.4willproveuseful.
11.1 Bessel Functions of the First Kind, Jν(x) 691
(b) Transform theintegralinpart(a)into
Jν(x)=1
π1/2(ν−1
2)!parenleftbiggx
2parenrightbiggνintegraldisplayπ
0cos(xcosθ)sin2νθdθ
=1
π1/2(ν−1
2)!parenleftbiggx
2parenrightbiggνintegraldisplayπ
0e±ixcosθsin2νθdθ
=1
π1/2(ν−1
2)!parenleftbiggx
2parenrightbiggνintegraldisplay1
−1e±ipx(1−p2)ν−1/2dp.
Thesearealternateintegralrepresentationsof Jν(x).
11.1.19 (a) From
Jν(x)=1
2πiparenleftbiggx
2parenrightbiggνintegraldisplay
t−ν−1et−x2/4tdt
derivetherecurrencerelation
J′
ν(x)=ν
xJν(x)−Jν+1(x).
(b) From
Jν(x)=1
2πiintegraldisplay
t−ν−1e(x/2)(t−1/t)dt
derivetherecurrencerelation
J′
ν(x)=1
2bracketleftbig
Jν−1(x)−Jν+1(x)bracketrightbig
.
11.1.20 Showthattherecurrencerelation
J′
n(x)=1
2bracketleftbig
Jn−1(x)−Jn+1(x)bracketrightbig
followsdirectlyfromdifferentiationof
Jn(x)=1
πintegraldisplayπ
0cos(nθ−xsinθ)dθ.
11.1.21 Evaluate
integraldisplay∞
0e−axJ0(bx)dx, a,b> 0.
Actuallytheresults holdfor a≥0,−∞<b<∞.This isaLaplacetransformof J0.
Hint.Eitheranintegralrepresentationof J0oraseries expansionwillbehelpful.
11.1.22 Usingtrigonometricforms, verifythat
J0(br)=1
2πintegraldisplay2π
0eibrsinθdθ.
11.1.23 (a) Plot the intensity ( /Phi12of Eq. (11.35)) as a function of (sinα/λ)along a diameter
ofthecirculardiffractionpattern.Locatethefirst twominima.
692 Chapter 11 Bessel Functions
(b) Whatfractionofthetotallightintensityfalls withinthecentralmaximum?
Hint.[J1(x)]2/xmay be written as a derivative and the area integral of the intensity
integratedbyinspection.
11.1.24 Thefractionoflightincidentonacircularaperture(normalincidence)thatistransmitted
isgivenby
T=2integraldisplay2ka
0J2(x)dx
x−1
2kaintegraldisplay2ka
0J2(x)dx.
Hereaistheradiusoftheapertureand kisthewavenumber, 2 π/λ.Showthat
(a)T=1−1
ka∞summationdisplay
n=0J2n+1(2ka), (b)T=1−1
2kaintegraldisplay2ka
0J0(x)dx.
11.1.25 Theamplitude U(ρ,ϕ,t) ofavibratingcircularmembraneofradius asatisfiesthewave
equation
∇2U−1
v2∂2U
∂t2=0.
Herevis the phase velocity of the wave fixed by the elastic constants and whatever
dampingis imposed.
(a) Showthatasolutionis
U(ρ,ϕ,t)=Jm(kρ)parenleftbig
a1eimϕ+a2e−imϕparenrightbigparenleftbig
b1eiωt+b2e−iωtparenrightbig
.
(b) From the Dirichlet boundary condition, Jm(ka)=0, find the allowable values of
thewavelength λ(k=2π/λ).
Note. There are other Bessel functions besides Jm, but they all diverge at ρ=0.
This is shown explicitly in Section 11.3. The divergent behavior is actually implicit
inEq. (11.6).
11.1.26 Example 11.1.2 describes the TM modes of electromagnetic cavity oscillation. The
transverse electric (TE) modes differ, in that we work from the zcomponent of the
magneticinduction B:
∇2Bz+α2Bz=0
withboundaryconditions
Bz(0)=Bz(l)=0 and∂Bz
∂ρvextendsinglevextendsinglevextendsinglevextendsingle
ρ=0=0.
ShowthattheTEresonantfrequenciesaregivenby
ωmnp=cradicalBigg
β2mn
a2+p2π2
l2,p=1,2,3,....
11.1 Bessel Functions of the First Kind, Jν(x) 693
11.1.27 Plot the three lowest TM and the three lowest TE angular resonant frequencies, ωmnp,
asafunctionoftheradius/length (a/l)ratiofor 0≤a/l≤1.5.
Hint.Tryplotting ω2(inunitsof c2/a2)v e r s u s(a/l)2. Whythischoice?
11.1.28 A thin conducting disk of radius acarries a charge q. Show that the potential is de-
scribedby
ϕ(r,z)=q
4πε0aintegraldisplay∞
0e−k|z|J0(kr)sinka
kdk,
whereJ0is the usual Bessel function and randzare the familiar cylindrical coordi-
nates.
Note. This is a difficult problem. One approach is through Fourier transforms such as
Exercise15.3.11.ForadiscussionofthephysicalproblemseeJackson( ClassicalElec-
trodynamics inAdditionalReadings).
11.1.29 Showthat
integraldisplaya
0xmJn(x)dx, m ≥n≥0,
(a) is integrable in terms of Bessel functions and powers of x(such asapJq(a))f o r
m+nodd;
(b) maybereducedtointegratedtermsplusintegraltexta
0J0(x)dxform+neven.
11.1.30 Showthat
integraldisplayα0n
0parenleftbigg
1−y
α0nparenrightbigg
J0(y)ydy=1
α0nintegraldisplayα0n
0J0(y)dy.
Hereα0nis thenth root of J0(y). This relation is useful (see Exercise 11.2.11): The
expression on the right is easier and quicker to evaluate—and much more accurate.
Taking the difference of two terms in the expression on the left leads to a large relative
error.
11.1.31 The circular aperature diffraction amplitude /Phi1of Eq. (17.35) is proportional to f(z)=
J1(z)/z. The corresponding single slit diffraction amplitude is proportional to g(z)=
sinz/z.
(a) Calculateandplot f(z)andg(z)forz=0.0(0.2)12.0.
(b) Locatethetwolowestvaluesof z(z>0)forwhich f(z)takesonanextremevalue.
Calculatethecorrespondingvaluesof f(z).
(c) Locatethetwolowestvaluesof z(z>0)forwhich g(z)takesonanextremevalue.
Calculatethecorrespondingvaluesof g(z).
11.1.32 Calculate the electrostatic potential of a charged disk ϕ(r,z)from the integral
form of Exercise 11.1.28. Calculate the potential for r/a=0.0(0.5)2.0 andz/a=
0.25(0.25)1.25. Why is z/a=0 omitted? Exercise 12.3.17 is a spherical harmonic
versionof thissameproblem.
694 Chapter 11 Bessel Functions
11.2 O RTHOGONALITY
IfBessel’sequation,Eq.(11.22a),isdividedby ρ,weseethatitbecomesself-adjoint,and
therefore, by the Sturm–Liouville theory, Section 10.2, the solutions are expected to be
orthogonal—if we can arrange to have appropriate boundary conditions satisfied. To take
care of the boundary conditions for a finite interval [0,a], we introduce parameters aand
ανmintotheargumentof JνtogetJν(ανmρ/a).Hereaistheupperlimitofthecylindrical
radialcoordinate ρ. FromEq. (11.22a),
ρd2
dρ2Jνparenleftbigg
ανmρ
aparenrightbigg
+d
dρJνparenleftbigg
ανmρ
aparenrightbigg
+parenleftbiggα2
νmρ
a2−ν2
ρparenrightbigg
Jνparenleftbigg
ανmρ
aparenrightbigg
=0.(11.45)
Changingtheparameter ανmtoανn,wefindthat Jν(ανnρ/a)satisfies
ρd2
dρ2Jνparenleftbigg
ανnρ
aparenrightbigg
+d
dρJνparenleftbigg
ανnρ
aparenrightbigg
+parenleftbiggα2
νnρ
a2−ν2
ρparenrightbigg
Jνparenleftbigg
ανnρ
aparenrightbigg
=0.(11.45a)
Proceeding as in Section 10.2, we multiply Eq. (11.45) by Jν(ανnρ/a)and Eq. (11.45a)
byJν(ανmρ/a)andsubtract,obtaining
Jνparenleftbigg
ανnρ
aparenrightbiggd
dρbracketleftbigg
ρd
dρJνparenleftbigg
ανmρ
aparenrightbiggbracketrightbigg
−Jνparenleftbigg
ανmρ
aparenrightbiggd
dρbracketleftbigg
ρd
dρJνparenleftbigg
ανnρ
aparenrightbiggbracketrightbigg
=α2
νn−α2
νm
a2ρJνparenleftbigg
ανmρ
aparenrightbigg
Jνparenleftbigg
ανnρ
aparenrightbigg
. (11.46)
Integratingfrom ρ=0t oρ=a, weobtain
integraldisplaya
0Jνparenleftbigg
ανnρ
aparenrightbiggd
dρbracketleftbigg
ρd
dρJνparenleftbigg
ανmρ
aparenrightbiggbracketrightbigg
dρ−integraldisplaya
0Jνparenleftbigg
ανmρ
aparenrightbiggd
dρbracketleftbigg
ρd
dρJνparenleftbigg
ανnρ
aparenrightbiggbracketrightbigg
dρ
=α2
νn−α2
νm
a2integraldisplaya
0Jνparenleftbigg
ανmρ
aparenrightbigg
Jνparenleftbigg
ανnρ
aparenrightbigg
ρdρ. (11.47)
Uponintegratingbyparts, wesee thattheleft-handsideof Eq.(11.47) becomes
vextendsinglevextendsinglevextendsinglevextendsingleρJνparenleftbigg
ανnρ
aparenrightbiggd
dρJνparenleftbigg
ανmρ
aparenrightbiggvextendsinglevextendsinglevextendsinglevextendsinglea
0−vextendsinglevextendsinglevextendsinglevextendsingleρJνparenleftbigg
ανmρ
aparenrightbiggd
dρJνparenleftbigg
ανnρ
aparenrightbiggvextendsinglevextendsinglevextendsinglevextendsinglea
0.(11.48)
Forν≥0 the factor ρguarantees a zero at the lower limit, ρ=0. Actually the lower
limit on the index νmay be extended down to ν>−1, Exercise 11.2.4.12Atρ=a, each
expression vanishes if we choose the parameters ανnandανmto be zeros, or roots of Jν;
thatis,Jν(ανm)=0.Thesubscriptsnowbecomemeaningful: ανmis themthzeroof Jν.
Withthischoiceofparameters,theleft-handsidevanishes(theSturm–Liouvillebound-
aryconditionsaresatisfied)andfor m/negationslash=n,
integraldisplaya
0Jνparenleftbigg
ανmρ
aparenrightbigg
Jνparenleftbigg
ανnρ
aparenrightbigg
ρdρ=0. (11.49)
Thisgivesusorthogonalityovertheinterval [0,a].
12Thecase ν=−1 reverts to ν=+1, Eq.(11.8).
11.2 Orthogonality 695
Normalization
The normalization integral may be developed by returning to Eq. (11.48), setting ανn=
ανm+ε, and taking the limit ε→0 (compare Exercise 11.2.2). With the aid of the recur-
rencerelation,Eq. (11.16),theresultmaybewrittenas
integraldisplaya
0bracketleftbigg
Jνparenleftbigg
ανmρ
aparenrightbiggbracketrightbigg2
ρd ρ=a2
2bracketleftbig
Jν+1(ανm)bracketrightbig2. (11.50)
Bessel Series
If we assume that the set of Bessel functions Jν(ανmρ/a))(νfixed,m=1,2,3,...)i s
complete, then any well-behaved but otherwise arbitrary function f(ρ)may be expanded
inaBesselseries (Bessel–FourierorFourier–Bessel)
f(ρ)=∞summationdisplay
m=1cνmJνparenleftbigg
ανmρ
aparenrightbigg
,0≤ρ≤a, ν>−1.(11.51)
Thecoefficients cνmaredeterminedbyusingEq. (11.50),
cνm=2
a2[Jν+1(ανm)]2integraldisplaya
0f(ρ)Jνparenleftbigg
ανmρ
aparenrightbigg
ρdρ. (11.52)
A similar series expansion involving Jν(βνmρ/a)with(d/dρ)J ν(βνmρ/a)|ρ=a=0i s
includedinExercises11.2.3and11.2.6(b).
Example 11.2.1 ELECTROSTATIC POTENTIAL IN A HOLLOW CYLINDER
From Table 9.3 of Section 9.3 (with αreplaced by k), our solution of Laplace’s equation
incircularcylindricalcoordinatesisalinearcombinationof
ψkm(ρ,ϕ,z)=Jm(kρ)[amsinmϕ+bmcosmϕ]bracketleftbig
c1ekz+c2e−kzbracketrightbig
.(11.53)
Theparticularlinearcombinationisdeterminedbytheboundaryconditionstobesatisfied.
Ourcylinderherehasaradius aandaheight l.Thetopendsectionhasapotentialdistrib-
utionψ(ρ,ϕ). Elsewhere on the surface the potential is zero.13The problem is to find the
electrostaticpotential
ψ(ρ,ϕ,z)=summationdisplay
k,mψkm(ρ,ϕ,z) (11.54)
everywhereintheinterior.
For convenience, the circular cylindrical coordinates are placed as shown in Fig. 11.3.
Sinceψ(ρ,ϕ,0)=0, we take c1=−c2=1
2.T h ezdependence becomes sinh kz, vanish-
ing atz=0. The requirement that ψ=0 on the cylindrical sides is met by requiring the
separationconstant ktobe
k=kmn=αmn
a, (11.55)
13Ifψ=0a tz=0,l,b u tψ/negationslash=0f o rρ=a, the modified Bessel functions, Section 11.5, are involved.
696 Chapter 11 Bessel Functions
where the first subscript, m, gives the index of the Bessel function, whereas the second
subscriptidentifiestheparticularzeroof Jm.
Theelectrostaticpotentialbecomes
ψ(ρ,ϕ,z)=∞summationdisplay
m=0∞summationdisplay
n=1Jmparenleftbigg
αmnρ
aparenrightbigg
·[amnsinmϕ+bmncosmϕ]·sinhparenleftbigg
αmnz
aparenrightbigg
.(11.56)
Equation(11.56)is adoubleseries:aBessel seriesin ρandaFourierseriesin ϕ.
Atz=l,ψ=ψ(ρ,ϕ),aknownfunctionof ρandϕ. Therefore
ψ(ρ,ϕ)=∞summationdisplay
m=0∞summationdisplay
n=1Jmparenleftbigg
αmnρ
aparenrightbigg
·[amnsinmϕ+bmncosmϕ]·sinhparenleftbigg
αmnl
aparenrightbigg
. (11.57)
The constants amnandbmnare evaluated by using Eqs. (11.49) and (11.50) and the corre-
spondingequationsfor sin ϕand cosϕ(Example10.2.1andEqs.(14.2), (14.3),(14.15)to
(14.17)). Wefind14
amn
bmnbracerightbigg
=2bracketleftbigg
πa2sinhparenleftbigg
αmnl
aparenrightbigg
J2
m+1(αmn)bracketrightbigg−1
·integraldisplay2π
0integraldisplaya
0ψ(ρ,ϕ)J mparenleftbigg
αmnρ
aparenrightbiggbraceleftbiggsinmϕ
cosmϕbracerightbigg
ρdρdϕ. (11.58)
These are definite integrals, that is, numbers. Substitutingback into Eq. (11.56), the series
isspecifiedandthepotential ψ(ρ,ϕ,z) is determined. /squaresolid
Continuum Form
The Bessel series, Eq. (11.51), and Exercise 11.2.6 apply to expansions over the finite
interval[0,a].I fa→∞, then the series forms may be expected to go over into integrals.
Thediscreteroots ανmbecomeacontinuousvariable α.Asimilarsituationisencountered
intheFourierseries,Section15.2.ThedevelopmentoftheBesselintegralfromtheBessel
seriesis leftasExercise11.2.8.
ForoperationswithacontinuumofBesselfunctions, Jν(αρ),akeyrelationistheBessel
functionclosureequation,
integraldisplay∞
0Jν(αρ)Jν(α′ρ)ρdρ=1
αδ(α−α′), ν >−1
2. (11.59)
ThismaybeprovedbytheuseofHankeltransforms,Section15.1.Analternateapproach,
startingfromarelationsimilartoEq.(10.82),isgivenbyMorseandFeshbach,Section6.3.
Asecondkindoforthogonality(varyingtheindex)isdevelopedforsphericalBesselfunc-
tionsinSection11.7.
14Ifm=0,the factor 2 is omitted (compare Eq. (14.16)).
11.2 Orthogonality 697
Exercises
11.2.1 Showthat
parenleftbig
a2−b2parenrightbigintegraldisplayP
0Jν(ax)Jν(bx)xdx=Pbracketleftbig
bJν(aP)J′
ν(bP)−aJ′
ν(aP)Jν(bP)bracketrightbig
,
with
J′
ν(aP)=d
d(ax)Jν(ax)vextendsinglevextendsingle
x=P,
integraldisplayP
0bracketleftbig
Jν(ax)bracketrightbig2xdx=P2
2braceleftbiggbracketleftbig
J′
ν(aP)bracketrightbig2+parenleftbigg
1−ν2
a2P2parenrightbiggbracketleftbig
Jν(aP)bracketrightbig2bracerightbigg
,ν>−1.
Thesetwointegralsareusuallycalledthe firstandsecondLommelintegrals .
Hint. We have the development of the orthogonality of the Bessel functions as an anal-
ogy.
11.2.2 Showthat
integraldisplaya
0bracketleftbigg
Jνparenleftbigg
ανmρ
aparenrightbiggbracketrightbigg2
ρdρ=a2
2bracketleftbig
Jν+1(ανm)bracketrightbig2,ν>−1.
Hereανmis themthzeroof Jν.
Hint.Withανn=ανm+ε,expandJν[(ανm+ε)ρ/a]aboutανmρ/abyaTaylorexpan-
sion.
11.2.3 (a) Ifβνmis themth zero of (d/dρ)J ν(βνmρ/a), show that the Bessel functions are
orthogonalovertheinterval [0,a]withanorthogonalityintegral
integraldisplaya
0Jνparenleftbigg
βνmρ
aparenrightbigg
Jνparenleftbigg
βνnρ
aparenrightbigg
ρdρ=0,m/negationslash=n, ν >−1.
(b) Derivethecorrespondingnormalizationintegral (m=n).
ANS.a2
2parenleftbigg
1−ν2
β2νmparenrightbiggbracketleftbig
Jν(βνm)bracketrightbig2,ν>−1.
11.2.4 Verify that the orthogonality equation, Eq. (11.49), and the normalization equation,
Eq. (11.50),holdfor ν>−1.
Hint.Usingpower-seriesexpansions,examinethebehaviorof Eq.(11.48) as ρ→0.
11.2.5 From Eq. (11.49) develop a proof that Jν(z),ν >−1, has no complex roots (with
nonzeroimaginarypart).
Hint.
(a) Usetheseries form of Jν(z)toexcludepureimaginaryroots.
(b) Assume ανmtobecomplexandtake ανntobeα∗
νm.
11.2.6 (a) Intheseriesexpansion
f(ρ)=∞summationdisplay
m=1cνmJνparenleftbigg
ανmρ
aparenrightbigg
,0≤ρ≤a, ν>−1,
698 Chapter 11 Bessel Functions
withJν(ανm)=0,showthatthecoefficientsare givenby
cνm=2
a2[Jν+1(ανm)]2integraldisplaya
0f(ρ)Jνparenleftbigg
ανmρ
aparenrightbigg
ρdρ.
(b) Intheseriesexpansion
f(ρ)=∞summationdisplay
m=1dνmJνparenleftbigg
βνmρ
aparenrightbigg
,0≤ρ≤a, ν>−1,
with(d/dρ)J ν(βνmρ/a)|ρ=a=0,showthatthecoefficientsaregivenby
dνm=2
a2(1−ν2/β2νm)[Jν(βνm)]2integraldisplaya
0f(ρ)Jνparenleftbigg
βνmρ
aparenrightbigg
ρdρ.
11.2.7 Arightcircularcylinderhasanelectrostaticpotentialof ψ(ρ,ϕ)onbothends.Thepo-
tentialonthecurvedcylindricalsurface iszero.Findthepotentialatallinteriorpoints.
Hint. Choose your coordinate system and adjust your zdependenceto exploit the sym-
metryofyourpotential.
11.2.8 Forthecontinuumcase,showthatEqs. (11.51) and(11.52) arereplacedby
f(ρ)=integraldisplay∞
0a(α)Jν(αρ)dα,
a(α)=αintegraldisplay∞
0f(ρ)Jν(αρ)ρdρ.
Hint.ThecorrespondingcaseforsinesandcosinesisworkedoutinSection15.2.These
are Hankel transforms. A derivation for the special case ν=0 is the topic of Exer-
cise15.1.1.
11.2.9 Afunction f(x)is expressedasaBessel series:
f(x)=∞summationdisplay
n=1anJm(αmnx),
withαmnthenthrootof Jm. ProvetheParsevalrelation,
integraldisplay1
0bracketleftbig
f(x)bracketrightbig2xdx=1
2∞summationdisplay
n=1a2
nbracketleftbig
Jm+1(αmn)bracketrightbig2.
11.2.10 Provethat
∞summationdisplay
n=1(αmn)−2=1
4(m+1).
Hint.Expand xminaBesselseries andapplytheParsevalrelation.
11.3 Neumann Functions 699
11.2.11 Arightcircularcylinderof length lhasapotential
ψparenleftbigg
z=±l
2parenrightbigg
=100parenleftbigg
1−ρ
aparenrightbigg
,
whereais the radius. The potential over the curved surface (side) is zero. Using
the Bessel series from Exercise 11.2.7, calculate the electrostatic potential for ρ/a=
0.0(0.2)1.0 andz/l=0.0(0.1)0.5.Takea/l=0.5.
Hint.FromExercise11.1.30youhave
integraldisplayα0n
0parenleftbigg
1−y
α0nparenrightbigg
J0(y)ydy.
Showthatthisequals
1
α0nintegraldisplayα0n
0J0(y)dy.
Numerical evaluation of this latter form rather than the former is both faster and more
accurate.
Note.F o rρ/a=0.0 andz/l=0.5 the convergence is slow, 20 terms giving only 98.4
ratherthan100.
Checkvalue .F orρ/a=0.4andz/l=0.3,
ψ=24.558.
11.3 N EUMANN FUNCTIONS ,BESSEL FUNCTIONS
OF THE SECOND KIND
FromthetheoryofODEsitisknownthatBessel’sequationhastwoindependentsolutions.
Indeed, for nonintegral order νwe have already found two solutions and labeled them
Jν(x)andJ−ν(x), using the infinite series (Eq. (11.5)). The trouble is that when νis
integral, Eq. (11.8) holds and we have but one independent solution. A second solution
may be developed by the methods of Section 9.6. This yields a perfectly good second
solutionofBessel’s equationbutis notthestandardform.
Definition and Series Form
Asanalternateapproach,wetaketheparticularlinearcombinationof Jν(x)andJ−ν(x)
Nν(x)=cosνπJν(x)−J−ν(x)
sinνπ. (11.60)
This is the Neumann function (Fig. 11.5).15For nonintegral ν,Nν(x)clearly satisfies
Bessel’s equation, for it is a linear combination of known solutions Jν(x)andJ−ν(x).
15In AMS-55 (see footnote 4 in Chapter 5 or Additional Readings of Chapter 8 p. for this ref.) and in most mathematics tables,
this is labeled Yν(x).
700 Chapter 11 Bessel Functions
FIGURE 11.5Neumannfunctions N0(x),N1(x), andN2(x).
Substitutingthepower-seriesEq. (11.6)for n→ν(giveninExercise11.1.7) yields
Nν(x)=−(ν−1)!
πparenleftbigg2
xparenrightbiggν
+···,16(11.61)
forν>0.However,forintegral ν,ν=n,Eq.(11.8)appliesandEq.(11.60)16becomesin-
determinate. The definition of Nν(x)was chosen deliberately for this indeterminate prop-
erty. Again substituting the power series and evaluating Nν(x)forν→0 by l’Hôpital’s
ruleforindeterminateforms, weobtainthelimitingvalue
N0(x)=2
π(lnx+γ−ln2)+Oparenleftbig
x2parenrightbig
(11.62)
forn=0 andx→0,using
ν!(−ν)!=πν
sinπν(11.63)
fromEq.(8.32).ThefirstandthirdtermsinEq.(11.62)comefromusing (d/dν)(x/ 2)ν=
(x/2)νln(x/2),whileγcomesfrom (d/dν)ν!forν→0 usingEqs.(8.38)and(8.40).For
n>0 weobtainsimilarly
Nn(x)=−1
π(n−1)!parenleftbigg2
xparenrightbiggn
+···+2
πparenleftbiggx
2parenrightbiggn1
n!lnparenleftbiggx
2parenrightbigg
+···. (11.64)
Equations(11.62)and(11.64)exhibitthelogarithmicdependencethatwastobeexpected.
This, ofcourse,verifiestheindependenceof JnandNn.
16Note thatthis limiting form applies to both integral and nonintegral values of the index ν.
11.3 Neumann Functions 701
Other Forms
As with all the other Bessel functions, Nν(x)has integral representations. For N0(x)we
have
N0(x)=−2
πintegraldisplay∞
0cos(xcosht)dt=−2
πintegraldisplay∞
1cos(xt)
(t2−1)1/2dt, x> 0.
These forms can be derived as the imaginary part of the Hankel representations of Exer-
cise11.4.7.Thelatterformis aFouriercosinetransform.
Toverifythat Nν(x),ourNeumannfunction(Fig.11.5)orBesselfunctionofthesecond
kind, actually does satisfy Bessel’s equation for integral n, we may proceed as follows.
L’Hôpital’sruleappliedtoEq. (11.60)yields
Nn(x)=(d/dν)[cosνπJν(x)−J−ν(x)]
(d/dν)sinνπvextendsinglevextendsinglevextendsinglevextendsingle
ν=n
=−πsinnπJn(x)+[cosnπ∂Jν/∂ν−∂J−ν/∂ν]|ν=n
πcosnπ
=1
πbracketleftbigg∂Jν(x)
∂ν−(−1)n∂J−ν(x)
∂νbracketrightbiggvextendsinglevextendsinglevextendsinglevextendsingle
ν=n. (11.65)
DifferentiatingBessel’s equationfor J±ν(x)withrespectto ν,weha v e
x2d2
dx2parenleftbigg∂J±ν
∂νparenrightbigg
+xd
dxparenleftbigg∂J±ν
∂νparenrightbigg
+parenleftbig
x2−ν2parenrightbig∂J±ν
∂ν=2νJ±ν. (11.66)
Multiplying the equation for J−νby(−1)ν, subtracting from the equation for Jν(as sug-
gestedbyEq. (11.65)), andtakingthelimit ν→n,weobtain
x2d2
dx2Nn+xd
dxNn+parenleftbig
x2−n2parenrightbig
Nn=2n
πbracketleftbig
Jn−(−1)nJ−nbracketrightbig
. (11.67)
Forν=n, an integer, the right-hand side vanishes by Eq. (11.8) and Nn(x)is seen to be a
solutionofBessel’sequation.Themostgeneralsolutionforany νcanthereforebewritten
as
y(x)=AJν(x)+BNν(x). (11.68)
It is seen from Eqs. (11.62) and (11.64) that Nndiverges, at least logarithmically. Any
boundary condition that requires the solution to be finite at the origin (as in our vibrat-
ing circular membrane (Section 11.1)) automatically excludes Nn(x). Conversely, in the
absenceofsucharequirement, Nn(x)mustbeconsidered.
To a certain extent the definition of the Neumann function Nn(x)is arbitrary. Equa-
tions (11.62) and (11.64) contain terms of the form anJn(x). Clearly, any finite value of
the constant anwould still give us a second solution of Bessel’s equation. Why should an
havetheparticularvalueimplicitin Eqs. (11.62)and(11.64)?Theanswerinvolvestheas-
ymptotic dependence developed in Section 11.6. If Jncorresponds to a cosine wave, then
Nncorresponds to a sine wave. This simple and convenient asymptotic phase relationship
isaconsequenceof theparticularadmixtureof JninNn.
702 Chapter 11 Bessel Functions
Recurrence Relations
SubstitutingEq.(11.60)for Nν(x)(nonintegral ν)intotherecurrencerelations(Eqs.(11.10)
and(11.12)for Jn(x),weseeimmediatelythat Nν(x)satisfiesthesesamerecurrencerela-
tions. This actually constitutes another proof that Nνis a solution. Note that the converse
is not necessarily true. All solutions need not satisfy the same recurrence relations. An
exampleofthis sortof troubleappearsinSection11.5.
Wronskian Formulas
From Section 9.6 and Exercise 10.1.4 we have the Wronskian formula17for solutions of
theBessel equation,
uν(x)v′
ν(x)−u′
ν(x)vν(x)=Aν
x, (11.69)
in which Aνis a parameter that depends on the particular Bessel functions uν(x)and
vν(x)being considered. Aνis a constant in the sense that it is independent of x. Consider
thespecialcase
uν(x)=Jν(x), v ν(x)=J−ν(x), (11.70)
JνJ′
−ν−J′
νJ−ν=Aν
x. (11.71)
SinceAνis a constant, it may be identified at any convenient point, such as x=0. Using
thefirst termsintheseriesexpansions(Eqs. (11.5) and(11.6)), weobtain
Jν→xν
2νν!,J−ν→2νx−ν
(−ν)!
J′
ν→νxν−1
2νν!,J′
−ν→−ν2νx−ν−1
(−ν)!. (11.72)
SubstitutionintoEq.(11.69) yields
Jν(x)J′
−ν(x)−J′
ν(x)J−ν(x)=−2ν
xν!(−ν)!=−2sinνπ
πx, (11.73)
usingEq.(8.32).Notethat Aνvanishesforintegral ν,asitmust,sincethenonvanishingof
the Wronskian is a test of the independence of the two solutions. By Eq. (11.73), Jnand
J−nare clearlylinearlydependent.
Using our recurrence relations, we may readily develop a large number of alternate
forms, amongwhichare
JνJ−ν+1+J−νJν−1=2sinνπ
πx, (11.74)
17This result depends on P(x)of Section 9.5 being equal to p′(x)/p(x), the corresponding coefficient of the self-adjoint form
of Section 10.1.
11.3 Neumann Functions 703
JνJ−ν−1+J−νJν+1=−2sinνπ
πx, (11.75)
JνN′
ν−J′
νNν=2
πx, (11.76)
JνNν+1−Jν+1Nν=−2
πx. (11.77)
Manymorewillbefoundinthereferencesgivenatchapter’send.
You will recall that in Chapter 9 Wronskians were of great value in two respects: (1) in
establishingthelinearindependenceorlineardependenceofsolutionsofdifferentialequa-
tions and (2) in developing an integral form of a second solution. Here the specific forms
oftheWronskiansandWronskian-derivedcombinationsofBesselfunctionsareusefulpri-
marilytoillustratethegeneralbehaviorofthevariousBesselfunctions.Wronskiansareof
great use in checking tables of Bessel functions. In Section 10.5 Wronskians appeared in
connectionwithGreen’sfunctions.
Example 11.3.1 COAXIAL WAVEGUIDES
Weareinterestedinanelectromagneticwaveconfinedbetweentheconcentric,conducting
cylindricalsurfaces ρ=aandρ=b.MostofthemathematicsisworkedoutinSection9.3
andExample11.1.2.Togofromthestandingwaveoftheseexamplestothetravelingwave
here,welet A=iB,A=amn,B=bmninEq. (11.40a)andobtain
Ez=summationdisplay
m,nbmnJm(γρ)e±imϕei(kz−ωt). (11.78)
Additional properties of the components of the electromagnetic wave in the simple cylin-
drical wave guide are explored in Exercises 11.3.8 and 11.3.9. For the coaxial wave guide
one generalization is needed. The origin, ρ=0, is now excluded (0<a≤ρ≤b). Hence
theNeumannfunction Nm(γρ)maynotbeexcluded. Ez(ρ,ϕ,z,t) becomes
Ez=summationdisplay
m,nbracketleftbig
bmnJm(γρ)+cmnNm(γρ)bracketrightbig
e±imϕei(kz−ωt). (11.79)
Withthecondition
Hz=0, (11.80)
wehavethebasicequationsfora TM(transversemagnetic)wave.
The(tangential)electricfieldmustvanishattheconductingsurfaces(Dirichletboundary
condition),or
bmnJm(γa)+cmnNm(γa)=0, (11.81)
bmnJm(γb)+cmnNm(γb)=0. (11.82)
704 Chapter 11 Bessel Functions
These transcendental equations may be solved for γ(γmn)and the ratio cmn/bmn.F r o m
Example11.1.2,
k2=ω2µ0ε0−γ2=ω2
c2−γ2. (11.83)
Sincek2must be positive for a real wave, the minimum frequency that will be propagated
(inthisTM mode)is
ω=γc, (11.84)
withγfixed by the boundary conditions, Eqs. (11.81) and (11.82). This is the cutoff fre-
quencyof thewaveguide.
ThereisalsoaTE(transverseelectric)mode,with Ez=0 andHzgivenbyEq.(11.79).
ThenwehaveNeumannboundaryconditionsinplaceofEqs.(11.81)and(11.82).Finally,
for the coaxial guide (not for the plain cylindrical guide, a=0), a TEM (transverse elec-
tromagnetic)mode, Ez=Hz=0, is possible. This corresponds to a plane wave, as in free
space.
The simpler cases (no Neumann functions, simpler boundary conditions) of a circular
waveguideareincludedasExercises11.3.8and11.3.9.
ToconcludethisdiscussionofNeumannfunctions,weintroducetheNeumannfunction
Nν(x)forthefollowingreasons:
1. Itisasecond,independentsolutionofBessel’sequation,whichcompletesthegeneral
solution.
2. It is required for specific physical problems such as electromagnetic waves in coaxial
cablesandquantummechanicalscatteringtheory.
3. It leadstoaGreen’sfunctionfor theBesselequation(Sections9.7 and10.5).
4. It leadsdirectlytothetwoHankelfunctions(Section11.4)./squaresolid
Exercises
11.3.1 ProvethattheNeumannfunctions Nn(withnaninteger)satisfytherecurrencerelations
Nn−1(x)+Nn+1(x)=2n
xNn(x),
Nn−1(x)−Nn+1(x)=2N′
n(x).
Hint.Theserelationsmaybeprovedbydifferentiatingtherecurrencerelationsfor Jνor
byusingthelimitformof Nνbutnotdividingeverythingbyzero.
11.3.2 Showthat
N−n(x)=(−1)nNn(x).
11.3.3 Showthat
N′
0(x)=−N1(x).
11.3 Neumann Functions 705
11.3.4 IfYandZare anytwosolutionsof Bessel’sequation,showthat
Yν(x)Z′
ν(x)−Y′
ν(x)Zν(x)=Aν
x,
in which Aνmay depend on νbut is independent of x. This is a special case of Exer-
cise10.1.4.
11.3.5 VerifytheWronskianformulas
Jν(x)J−ν+1(x)+J−ν(x)Jν−1(x)=2sinνπ
πx,
Jν(x)N′
ν(x)−J′
ν(x)Nν(x)=2
πx.
11.3.6 Asanalternativetoletting xapproachzerointheevaluationoftheWronskianconstant,
we may invoke uniqueness of power series (Section 5.7). The coefficient of x−1in the
seriesexpansionof uν(x)v′
ν(x)−u′
ν(x)vν(x)isthenAν.Showbyseriesexpansionthat
thecoefficientsof x0andx1ofJν(x)J′
−ν(x)−J′
ν(x)J−ν(x)are eachzero.
11.3.7 (a) BydifferentiatingandsubstitutingintoBessel’s ODE,showthat
integraldisplay∞
0cos(xcosht)dt
isasolution.
Hint.Youcanrearrangethefinalintegralas
integraldisplay∞
0d
dtbraceleftbig
xsin(xcosht)sinhtbracerightbig
dt.
(b) Showthat
N0(x)=−2
πintegraldisplay∞
0cos(xcosht)dt
islinearlyindependentof J0(x).
11.3.8 A cylindrical wave guide has radius r0. Find the nonvanishing components of the elec-
tricandmagneticfieldsfor
(a) TM 01, transversemagneticwave (Hz=Hρ=Eϕ=0),
(b) TE 01, transverseelectricwave (Ez=Eρ=Hϕ=0).
The subscripts 01 indicate that the longitudinal component ( EzorHz)i n v o l v e s J0and
theboundaryconditionissatisfiedbythe firstzeroofJ0orJ′
0.
Hint.Allcomponentsofthewavehavethesamefactor: exp i(kz−ωt).
11.3.9 Foragivenmodeofoscillationthe minimum frequencythatwillbepassedbyacircular
cylindricalwaveguide(radius r0)i s
νmin=c
λc,
706 Chapter 11 Bessel Functions
inwhich λcis fixedbytheboundarycondition
Jnparenleftbigg2πr0
λcparenrightbigg
=0forTMnmmode,
J′
nparenleftbigg2πr0
λcparenrightbigg
=0forTEnmmode.
The subscript ndenotes the order of the Bessel function and mindicates the zero
used. Find this cutoff wavelength λcfor the three TM and three TE modes with the
longest cutoff wavelengths. Explain your results in terms of the graph of J0,J1, andJ2
(Fig.11.1).
11.3.10 Write a program that will compute successive roots of the Neumann function Nn(x),
that isαns, whereNn(αns)=0. Tabulate the first five roots of N0,N1, andN2. Check
your values for the roots against those listed in AMS-55 (see Additional Readings of
Chapter8forthefullref.).
Checkvalue. α12=5.42968.
11.3.11 For the case m=0,a=1, andb=2, the coaxial wave guide boundary conditions lead
to
f(x)=J0(2x)
N0(2x)−J0(x)
N0(x)
(Fig.11.6).
(a) Calculate f(x)forx=0.0(0.1)10.0 and plot f(x)versusxto find the approxi-
matelocationoftheroots.
FIGURE 11.6f(x)of Exercise11.3.11.
11.4 Hankel Functions 707
(b) Callaroot-findingsubroutinetodeterminethefirstthreerootstohigherprecision.
ANS.3.1230,6.2734,9.4182.
Note. The higher roots can be expected to appear at intervals whose length approaches
n. Why? AMS-55 (see Additional Readings of Chapter 8 for the reference), gives an
approximateformulafortheroots.Thefunction g(x)=J0(x)N0(2x)−J0(2x)N0(x)is
muchbetterbehavedthan f(x)previouslydiscussed.
11.4 H ANKEL FUNCTIONS
ManyauthorsprefertointroducetheHankelfunctionsbymeansofintegralrepresentations
andthentousethemtodefinetheNeumannfunction Nν(z).Anoutlineofthisapproachis
givenattheendofthis section.
Definitions
Because we have already obtained the Neumann function by more elementary (and less
powerful)techniques,wemayuseittodefinetheHankelfunctions H(1)
ν(x)andH(2)
ν(x):
H(1)
ν(x)=Jν(x)+iNν(x) (11.85)
and
H(2)
ν(x)=Jν(x)−iNν(x). (11.86)
Thisis exactlyanalogoustotaking
e±iθ=cosθ±isinθ. (11.87)
Forrealarguments, H(1)
νandH(2)
νare complexconjugates.The extentof theanalogywill
be seen even better when the asymptotic forms are considered (Section 11.6). Indeed, it is
theirasymptoticbehaviorthatmakestheHankelfunctionsuseful.
Seriesexpansionof H(1)
ν(x)andH(2)
ν(x)maybeobtainedbycombiningEqs.(11.5)and
(11.63).Oftenonlythefirst term isof interest;itis givenby
H(1)
0(x)≈i2
πlnx+1+i2
π(γ−ln2)+···, (11.88)
H(1)
ν(x)≈−i(ν−1)!
πparenleftbigg2
xparenrightbiggν
+···,ν>0, (11.89)
H(2)
0(x)≈−i2
πlnx+1−i2
π(γ−ln2)+···, (11.90)
H(2)
ν(x)≈i(ν−1)!
πparenleftbigg2
xparenrightbiggν
+···,ν>0. (11.91)
708 Chapter 11 Bessel Functions
Since the Hankel functions are linear combinations (with constant coefficients) of Jν
andNν,theysatisfy thesamerecurrencerelations(Eqs. (11.10)and(11.12))
Hν−1(x)+Hν+1(x)=2ν
xHν(x), (11.92)
Hν−1(x)−Hν+1(x)=2H′
ν(x), (11.93)
forbothH(1)
ν(x)andH(2)
ν(x).
AvarietyofWronskianformulascanbedeveloped:
H(2)
νH(1)
ν+1−H(1)
νH(2)
ν+1=4
iπx, (11.94)
Jν−1H(1)
ν−JνH(1)
ν−1=2
iπx, (11.95)
JνH(2)
ν−1−Jν−1H(2)
ν=2
iπx. (11.96)
Example 11.4.1 CYLINDRICAL TRAVELING WAVES
AsanillustrationoftheuseofHankelfunctions,consideratwo-dimensionalwaveproblem
similartothevibratingcircularmembraneofExercise11.1.25.Nowimaginethatthewaves
are generated at r=0 and move outward to infinity. We replace our standing waves by
traveling ones. The differential equation remains the same, but the boundary conditions
change.Wenowdemandthatfor large rthewavebehavelike
U∼ei(kr−ωt)(11.97)
to describe an outgoing wave. As before, kis the wave number. This assumes, for sim-
plicity, that there is no azimuthaldependence,that is, no angularmomentum,or m=0.In
Sections7.3and11.6, H(1)
0(kr)isshowntohavetheasymptoticbehavior(for r→∞)
H(1)
0(kr)∼eikr. (11.98)
Thisboundaryconditionatinfinitythendeterminesourwavesolutionas
U(r,t)=H(1)
0(kr)e−iωt. (11.99)
This solution diverges as r→0, which is the behavior to be expected with a source at the
origin.
Thechoiceofatwo-dimensionalwaveproblemtoillustratetheHankelfunction H(1)
0(z)
is not accidental. Bessel functions may appear in a variety of ways, such as in the sepa-
ration of conical coordinates. However, they enter most commonly in the radial equations
from the separation of variables in the Helmholtz equation in cylindrical and in spheri-
cal polar coordinates. We have taken a degenerate form of cylindrical coordinates for this
illustration. Had we used spherical polar coordinates (spherical waves), we should have
encounteredindex ν=n+1
2,nan integer.These specialvalues yieldthe sphericalBessel
functionstobediscussedinSection11.7. /squaresolid
11.4 Hankel Functions 709
Contour Integral Representation of
the Hankel Functions
Theintegralrepresentation(Schlaefliintegral)
Jν(x)=1
2πicontintegraldisplay
Ce(x/2)(t−1/t)dt
tν+1(11.100)
may easily be established as a Cauchy integral for ν=n, an integer (by recognizing that
the numerator is the generating function (Eq. (11.1)) and integrating around the origin).
Ifνis not an integer, the integrand is not single-valued and a cut line is needed in our
complexplane.Choosingthenegativerealaxisasthecutlineandusingthecontourshown
in Fig. 11.7, we can extend Eq. (11.100) to nonintegral ν. Substituting Eq. (11.100) into
Bessel’s ODE, we can represent the combined integrand by an exact differential that van-
ishesast→∞e±iπ(compareExercise11.1.16).
We now deform the contour so that it approaches the origin along the positive real axis,
as shown in Fig. 11.8. For x>0,this particular approach guarantees that the exact differ-
entialmentionedwillvanishas t→0 becauseofthe e−x/2t→0 factor.Henceeachofthe
separate portions ( ∞e−iπto 0) and (0 to ∞eiπ) is a solution of Bessel’s equation. We
define
H(1)
ν(x)=1
πiintegraldisplay∞eiπ
0e(x/2)(t−1/t)dt
tν+1, (11.101)
H(2)
ν(x)=1
πiintegraldisplay0
∞e−iπe(x/2)(t−1/t)dt
tν+1. (11.102)
Theseexpressionsareparticularlyconvenientbecausetheymaybehandledbythemethod
of steepest descents (Section 7.3). H(1)
ν(x)has a saddle point at t=+i, whereas H(2)
ν(x)
hasasaddlepointat t=−i.
FIGURE 11.7Besselfunctioncontour.
710 Chapter 11 Bessel Functions
FIGURE 11.8Hankelfunctioncontours.
TheproblemofrelatingEqs.(11.101)and(11.102)toourearlierdefinitionoftheHankel
function (Eqs. (11.85) and (11.86)) remains. Since Eqs. (11.100) to (11.102) combined
yield
Jν(x)=1
2bracketleftbig
H(1)
ν(x)+H(2)
ν(x)bracketrightbig
(11.103)
byinspection,weneedonlyshowthat
Nν(x)=1
2ibracketleftbig
H(1)
ν(x)−H(2)
ν(x)bracketrightbig
. (11.104)
Thismaybeaccomplishedbythefollowingsteps:
1. Withthesubstitutions t=eiπ/sforH(1)
νandt=e−iπ/sforH(2)
ν, weobtain
H(1)
ν(x)=e−iνπH(1)
−ν(x), (11.105)
H(2)
ν(x)=eiνπH(2)
−ν(x). (11.106)
2. FromEqs. (11.103) (ν→−ν), (11.105), and(11.106),
J−ν(x)=1
2bracketleftbig
eiνπH(1)
ν(x)+e−iνπH(2)
ν(x)bracketrightbig
. (11.107)
3. Finally substitute Jν(Eq. (11.103)) and J−ν(Eq. (11.107)) into the defining equation
forNν, Eq. (11.60). This leads to Eq. (11.104) and establishes the contour integrals
Eqs. (11.101)and(11.102)astheHankelfunctions.
Integral representations have appeared before: Eq. (8.35) for Ŵ(z)and various representa-
tionsofJν(z)inSection11.1.WiththeseintegralrepresentationsoftheHankelfunctions,
it is perhaps appropriate to ask why we are interested in integral representations. There
are at least four reasons. The first is simply aesthetic appeal. Second, the integral repre-
sentationshelptodistinguishbetweentwolinearlyindependentsolutions.InFig.11.6,the
contoursC1andC2crossdifferent saddlepoints(Section7.3).FortheLegendrefunctions
thecontourfor Pn(z)(Fig.12.11)andthatfor Qn(z)encircledifferent singularpoints.
11.4 Hankel Functions 711
Third, the integral representations facilitate manipulations, analysis, and the develop-
ment of relations among the various special functions. Fourth, and probably most impor-
tant of all, the integral representations are extremely useful in developing asymptotic ex-
pansions.Oneapproach,themethodofsteepestdescents,appearsinSection7.3.Asecond
approach,thedirectexpansionofanintegralrepresentationisgiveninSection11.6forthe
modified Bessel function Kν(z). This same technique may be used to obtain asymptotic
expansionsof theconfluenthypergeometricfunctions MandU—Exercise13.5.13.
In conclusion,theHankelfunctionsareintroducedhereforthefollowingreasons:
•Asanalogsof e±ixtheyare usefulfordescribingtravelingwaves.
•Theyofferanalternate(contourintegral)andaratherelegantdefinitionofBesselfunc-
tions.
•H(1)
νis usedtodefinethemodifiedBesselfunction KνofSection11.5.
Exercises
11.4.1 VerifytheWronskianformulas
(a)Jν(x)H(1)′
ν(x)−J′
ν(x)H(1)
ν(x)=2i
πx,
(b)Jν(x)H(2)′
ν(x)−J′
ν(x)H(2)
ν(x)=−2i
πx,
(c)Nν(x)H(1)′
ν(x)−N′
ν(x)H(1)
ν(x)=−2
πx,
(d)Nν(x)H(2)′
ν(x)−N′
ν(x)H(2)
ν(x)=−2
πx,
(e)H(1)
ν(x)H(2)′
ν(x)−H(1)′
ν(x)H(2)
ν(x)=−4i
πx,
(f)H(2)
ν(x)H(1)
ν+1(x)−H(1)
ν(x)H(2)
ν+1(x)=4
iπx,
(g)Jν−1(x)H(1)
ν(x)−Jν(x)H(1)
ν−1(x)=2
iπx.
11.4.2 Showthattheintegralforms
(a)1
iπintegraldisplay∞eiπ
0C1e(x/2)(t−1/t)dt
tν+1=H(1)
ν(x),
(b)1
iπintegraldisplay0
∞e−iπC2e(x/2)(t−1/t)dt
tν+1=H(2)
ν(x)
satisfyBessel’s ODE.Thecontours C1andC2areshowninFig.11.8.
11.4.3 Usingtheintegralsandcontoursgiveninproblem11.4.2,showthat
1
2ibracketleftbig
H(1)
ν(x)−H(2)
ν(x)bracketrightbig
=Nν(x).
11.4.4 ShowthattheintegralsinExercise11.4.2maybetransformedtoyield
(a)H(1)
ν(x)=1
πiintegraldisplay
C3exsinhγ−νγdγ,(b)H(2)
ν(x)=1
πiintegraldisplay
C4exsinhγ−νγdγ
712 Chapter 11 Bessel Functions
FIGURE 11.9Hankelfunctioncontours.
(seeFig.11.9).
11.4.5 (a) Transform H(1)
0(x), Eq. (11.101),into
H(1)
0(x)=1
iπintegraldisplay
Ceixcoshsds,
where the contour Cruns from−∞−iπ/2 through the origin of the s-plane to
∞+iπ/2.
(b) Justifyrewriting H(1)
0(x)as
H(1)
0(x)=2
iπintegraldisplay∞+iπ/2
0eixcoshsds.
(c) Verify that this integral representation actually satisfies Bessel’s differential equa-
tion.(The iπ/2intheupperlimitisnotessential.Itservesasaconvergencefactor.
Wecanreplaceitby iaπ/2 andtakethelimit.)
11.4.6 From
H(1)
0(x)=2
iπintegraldisplay∞
0eixcoshsds
showthat
(a)J0(x)=2
πintegraldisplay∞
0sin(xcoshs)ds, (b)J0(x)=2
πintegraldisplay∞
1sin(xt)√
t2−1dt.
Thislastresult isaFouriersinetransform.
11.4.7 From (seeExercises11.4.4and11.4.5)
H(1)
0(x)=2
iπintegraldisplay∞
0eixcoshsds
showthat
(a)N0(x)=−2
πintegraldisplay∞
0cos(xcoshs)ds.
11.5 Modified Bessel Functions, Iν(x)andKν(x) 713
(b)N0(x)=−2
πintegraldisplay∞
1cos(xt)radicalbig
t2−1)dt.
ThesearetheintegralrepresentationsinSection11.3(OtherForms).
Thislastresult isaFouriercosinetransform.
11.5 M ODIFIED BESSEL FUNCTIONS ,Iν(x) ANDKν(x)
TheHelmholtzequation,
∇2ψ+k2ψ=0,
separated in circular cylindrical coordinates, leads to Eq. (11.22a), the Bessel equation.
Equation (11.22a) is satisfied by the Bessel and Neumann functions Jν(kρ)andNν(kρ)
and any linear combination, such as the Hankel functions H(1)
ν(kρ)andH(2)
ν(kρ).N o w ,
the Helmholtz equation describes the space part of wave phenomena. If instead we have a
diffusionproblem,thentheHelmholtzequationis replacedby
∇2ψ−k2ψ=0. (11.108)
TheanalogtoEq.(11.22a)is
ρ2d2
dρ2Yν(kρ)+ρd
dρYν(kρ)−parenleftbig
k2ρ2+ν2parenrightbig
Yν(kρ)=0.(11.109)
The Helmholtz equation may be transformed into the diffusion equation by the trans-
formation k→ik. Similarly, k→ikchanges Eq. (11.22a) into Eq. (11.109) and shows
that
Yν(kρ)=Zν(ikρ).
The solutions of Eq. (11.109) are Bessel functions of imaginary argument. To obtain a
solution that is regular at the origin, we take Zνas the regular Bessel function Jν.I ti s
customary(andconvenient)tochoosethenormalizationsothat
Yν(x)=Iν(x)≡i−νJν(ix). (11.110)
(Here the variable kρis being replaced by xfor simplicity.) The extra i−νnormalization
cancelsthe iνfrom eachtermandleaves Iν(x)real.Oftenthisis writtenas
Iν(x)=e−νπi/2Jνparenleftbig
xeiπ/2parenrightbig
. (11.111)
I0andI1areshowninFig.11.10.
714 Chapter 11 Bessel Functions
FIGURE 11.10ModifiedBessel
functions.
Series Form
In terms of infinite series this is equivalent to removing the (−1)ssign in Eq. (11.5) and
writing
Iν(x)=∞summationdisplay
s=01
s!(s+ν)!parenleftbiggx
2parenrightbigg2s+ν
,I−ν(x)=∞summationdisplay
s=01
s!(s−ν)!parenleftbiggx
2parenrightbigg2s−ν
.(11.112)
For integral νthis yields
In(x)=I−n(x). (11.113)
Recurrence Relations
The recurrence relations satisfied by Iν(x)may be developed from the series expansions,
but it is perhaps easier to work from the existing recurrence relations for Jν(x). Let us
replacexby−ixandrewriteEq.(11.110)as
Jν(x)=iνIν(−ix). (11.114)
ThenEq. (11.10)becomes
iν−1Iν−1(−ix)+iν+1Iν+1(−ix)=2ν
xiνIν(−ix).
Replacing xbyix, wehavearecurrencerelationfor Iν(x),
Iν−1(x)−Iν+1(x)=2ν
xIν(x). (11.115)
11.5 Modified Bessel Functions, Iν(x)andKν(x) 715
Equation(11.12)transforms to
Iν−1(x)+Iν+1(x)=2I′
ν(x). (11.116)
ThesearetherecurrencerelationsusedinExercise11.1.14.Itisworthemphasizingthatal-
thoughtworecurrencerelations,Eqs.(11.115)and(11.116)orExercise11.5.7,specifythe
second-orderODE,theconverseisnottrue.TheODEdoesnotuniquelyfixtherecurrence
relations.Equations(11.115)and(11.116)andExercise11.5.7provideanexample.
From Eq. (11.113) it is seen that we have but one independent solution when νis an
integer,exactlyasintheBesselfunctions Jν.Thechoiceofasecond,independentsolution
of Eq. (11.108) is essentially a matter of convenience. The second solution given here
is selected on the basis of its asymptotic behavior—as shown in the next section. The
confusion of choice and notation for this solution is perhaps greater than anywhere else
in this field.18Many authors19choose to define a second solution in terms of the Hankel
functionH(1)
ν(x)by
Kν(x)≡π
2iν+1H(1)
ν(ix)=π
2iν+1bracketleftbig
Jν(ix)+iNν(ix)bracketrightbig
. (11.117)
Thefactor iν+1makesKν(x)realwhen xisreal.UsingEqs.(11.60)and(11.110),wemay
transform Eq. (11.117)to20
Kν(x)=π
2I−ν(x)−Iν(x)
sinνπ, (11.118)
analogoustoEq.(11.60)for Nν(x).ThechoiceofEq.(11.117)asadefinitionissomewhat
unfortunate in that the function Kν(x)does not satisfy the same recurrence relations as
Iν(x)(compare Exercises 11.5.7 and 11.5.8). To avoid this annoyance, other authors21
haveincludedanadditionalfactorofcos νπ.Thispermits Kνtosatisfythesamerecurrence
relationsas Iν, butithasthedisadvantageof making Kν=0f o rν=1
2,3
5,5
2,....
The series expansion of Kν(x)follows directly from the series form of H(1)
ν(ix).T h e
lowest-orderterms are(cf. Eqs. (11.61)and(11.62))
K0(x)=−lnx−γ+ln2+···,
Kν(x)=2ν−1(ν−1)!x−ν+···. (11.119)
BecausethemodifiedBessel function Iνis relatedtotheBessel function Jν, muchassinh
is related to sine, Iνand the second solution Kνare sometimes referred to as hyperbolic
Besselfunctions. K0andK1are showninFig.11.10.
I0(x)andK0(x)havetheintegralrepresentations
I0(x)=1
πintegraldisplayπ
0cosh(xcosθ)dθ, (11.120)
K0(x)=integraldisplay∞
0cos(xsinht)dt=integraldisplay∞
0cos(xt)dt
(t2+1)1/2,x>0. (11.121)
18Adiscussion and comparison of notations will befound in Math. Tables Aids Comput. 1: 207–308 (1944).
19Watson, Morse and Feshbach, Jeffreys and Jeffreys (without the π/2).
20For integral index nwetakethe limit as ν→n.
21Whittaker and Watson, seeAdditional Readings of Chapter 13.
716 Chapter 11 Bessel Functions
Equation(11.120)maybederivedfromEq.(11.30)for J0(x)ormaybetakenasaspecial
caseofExercise11.5.4, ν=0.Theintegralrepresentationof K0,Eq.(11.121),isaFourier
transform and may best be derived with Fourier transforms, Chapter 15, or with Green’s
functionsSection9.7.Avarietyofotherformsofintegralrepresentations(including ν/negationslash=0)
appearintheexercises.Theseintegralrepresentationsareusefulindevelopingasymptotic
forms(Section11.6) andinconnectionwithFouriertransforms, Chapter15.
To put the modified Bessel functions Iν(x)andKν(x)in proper perspective, we intro-
ducethemherebecause:
•Thesefunctionsare solutionsofthefrequentlyencounteredmodifiedBessel equation.
•Theyare neededfor specificphysicalproblems,suchas diffusionproblems.
•Kν(x)providesa Green’sfunction,Section9.7.
•Kν(x)leadstoaconvenientdeterminationof asymptoticbehavior(Section11.6).
Exercises
11.5.1 Showthat
e(x/2)(t+1/t)=∞summationdisplay
n=−∞In(x)tn,
thusgeneratingmodifiedBesselfunctions, In(x).
11.5.2 Verifythefollowingidentities
(a) 1=I0(x)+2∞summationdisplay
n=1(−1)nI2n(x),
(b)ex=I0(x)+2∞summationdisplay
n=1In(x),
(c)e−x=I0(x)+2∞summationdisplay
n=1(−1)nIn(x),
(d) cosh x=I0(x)+2∞summationdisplay
n=1I2n(x),
(e) sinh x=2∞summationdisplay
n=1I2n−1(x).
11.5.3 (a) FromthegeneratingfunctionofExercise11.5.1showthat
In(x)=1
2πicontintegraldisplay
expbracketleftbig
(x/2)(t+1/t)bracketrightbigdt
tn+1.
11.5 Modified Bessel Functions, Iν(x)andKν(x) 717
(b) For n=ν, not an integer, show that the preceding integral representation may be
generalizedto
Iν(x)=1
2πiintegraldisplay
Cexpbracketleftbig
(x/2)(t+1/t)bracketrightbigdt
tν+1.
Thecontour Cisthesameasthatfor Jν(x), Fig.11.7.
11.5.4 Forν>−1
2showthat Iν(z)mayberepresentedby
Iν(z)=1
π1/2(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplayπ
0e±zcosθsin2νθdθ
=1
π1/2(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplay1
−1e±zpparenleftbig
1−p2parenrightbigν−1/2dp
=2
π1/2(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplayπ/2
0cosh(zcosθ)sin2νθdθ.
11.5.5 A cylindrical cavity has a radius aand height l, Fig. 11.3. The ends, z=0 andl,a r ea t
zeropotential.Thecylindricalwalls, ρ=a, haveapotential V=V(ϕ,z).
(a) Showthattheelectrostaticpotential /Phi1(ρ,ϕ,z) has thefunctionalform
/Phi1(ρ,ϕ,z)=∞summationdisplay
m=0∞summationdisplay
n=1Im(knρ)sinknz·(amnsinmϕ+bmncosmϕ),
wherekn=nπ/l.
(b) Showthatthecoefficients amnandbmnaregivenby22
amn
bmnbracerightbigg
=2
πlIm(kna)integraldisplay2π
0integraldisplayl
0V(ϕ,z)sinknz·braceleftbiggsinmϕ
cosmϕbracerightbigg
dzdϕ.
Hint. Expand V(ϕ,z)as a double series and use the orthogonality of the trigonometric
functions.
11.5.6 Verifythat Kν(x)is givenby
Kν(x)=π
2I−ν(x)−Iν(x)
sinνπ
andfrom thisshowthat
Kν(x)=K−ν(x).
11.5.7 Showthat Kν(x)satisfiestherecurrencerelations
Kν−1(x)−Kν+1(x)=−2ν
xKν(x),
Kν−1(x)+Kν+1(x)=−2K′
ν(x).
22Whenm=0, the2in thecoefficient is replacedby 1.
718 Chapter 11 Bessel Functions
11.5.8 IfKν=eνπiKν, showthat Kνsatisfiesthesamerecurrencerelationsas Iν.
11.5.9 Forν>−1
2showthat Kν(z)mayberepresentedby
Kν(z)=π1/2
(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplay∞
0e−zcoshtsinh2νtdt,−π
2<argz<π
2
=π1/2
(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplay∞
1e−zp(p2−1)ν−1/2dp.
11.5.10 Showthat Iν(x)andKν(x)satisfy theWronskianrelation
Iν(x)K′
ν(x)−I′
ν(x)Kν(x)=−1
x.
Thisresult isquotedinSection9.7inthedevelopmentof aGreen’sfunction.
11.5.11 Ifr=(x2+y2)1/2, prove that
1
r=2
πintegraldisplay∞
0cos(xt)K0(yt)dt.
Thisis aFouriercosinetransformof K0.
11.5.12 (a) Verifythat
I0(x)=1
πintegraldisplayπ
0cosh(xcosθ)dθ
satisfiesthemodifiedBessel equation, ν=0.
(b) Showthatthisintegralcontainsnoadmixtureof K0(x),theirregularsecondsolu-
tion.
(c) Verifythenormalizationfactor 1 /π.
11.5.13 Verifythattheintegralrepresentations
In(z)=1
πintegraldisplayπ
0ezcostcos(nt)dt,
Kν(z)=integraldisplay∞
0e−zcoshtcosh(νt)dt,ℜ(z)>0,
satisfy the modified Bessel equation by direct substitution into that equation. How can
you show that the first form does not contain an admixture of Knand that the second
formdoesnotcontainanadmixtureof Iν?Howcanyoucheckthenormalization?
11.5.14 Derivetheintegralrepresentation
In(x)=1
πintegraldisplayπ
0excosθcos(nθ)dθ.
Hint. Start with the corresponding integral representation of Jn(x). Equation (11.120)
isa specialcaseofthisrepresentation.
11.6 Asymptotic Expansions 719
11.5.15 Showthat
K0(z)=integraldisplay∞
0e−zcoshtdt
satisfies the modified Bessel equation. How can you establish that this form is linearly
independentof I0(z)?
11.5.16 Showthat
eax=I0(a)T0(x)+2∞summationdisplay
n=1In(a)Tn(x),−1≤x≤1.
Tn(x)isthenth-orderChebyshevpolynomial,Section13.3.
Hint.AssumeaChebyshevseriesexpansion.Usingtheorthogonalityandnormalization
oftheTn(x), solvefor thecoefficientsof theChebyshevseries.
11.5.17 (a) Write a double precision subroutine to calculate In(x)to 12-decimal-place accu-
racy forn=0,1,2,3,...and 0≤x≤1. Check your results against the 10-place
valuesgiveninAMS-55,Table9.11,seeAdditionalReadingsofChapter8forthe
reference.
(b) Referring to Exercise 11.5.16, calculate the coefficients in the Chebyshev expan-
sionsof cosh xandof sinh x.
11.5.18 Thecylindricalcavityof Exercise11.5.5hasa potentialalongthecylinderwalls:
V(z)=braceleftBigg
100z
l, 0≤z
l≤1
2,
100parenleftbig
1−z
lparenrightbig
,1
2≤z
l≤1.
Withtheradius–heightratio a/l=0.5,calculatethepotentialfor z/l=0.1(0.1)0.5and
ρ/a=0.0(0.2)1.0.
Checkvalue. Forz/l=0.3 andρ/a=0.8,V=26.396.
11.6 A SYMPTOTIC EXPANSIONS
Frequently in physical problems there is a need to know how a given Bessel or modified
Bessel function behaves for large values of the argument, that is, the asymptotic behavior.
This is one occasion when computers are not very helpful. One possible approach is to
developapower-seriessolutionofthedifferentialequation,asinSection9.5,butnowusing
negative powers. This is Stokes’ method, Exercise 11.6.5. The limitation is that starting
from some positive value of the argument (for convergence of the series), we do not know
whatmixtureofsolutionsormultipleofagivensolutionwehave.Theproblemistorelate
theasymptoticseries(usefulforlargevaluesofthevariable)tothepower-seriesorrelated
definition (useful for small values of the variable). This relationship can be established by
introducingasuitable integralrepresentation andthenusingeitherthemethodofsteepest
descent,Section7.3, orthedirectexpansionasdevelopedinthissection.
720 Chapter 11 Bessel Functions
Expansion of an Integral Representation
Asadirectapproach,considertheintegralrepresentation(Exercise11.5.9)
Kν(z)=π1/2
(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplay∞
1e−zxparenleftbig
x2−1parenrightbigν−1/2dx, ν>−1
2.(11.122)
For the present let us take zto be real, although Eq. (11.122) may be established for
−π/2<argz<π/2(ℜ(z)>0). Wehavethreetasks:
1. To show that Kνas given in Eq. (11.122) actually satisfies the modified Bessel equa-
tion(11.109).
2. Toshowthattheregularsolution Iνisabsent.
3. ToshowthatEq.(11.122) hasthepropernormalization.
1.ThefactthatEq.(11.122)isasolutionofthemodifiedBesselequationmaybeverified
bydirectsubstitution.Weobtain
zν+1integraldisplay∞
1d
dxbracketleftbig
e−zxparenleftbig
x2−1parenrightbigν+1/2bracketrightbig
dx=0,
which transforms the combined integrand into the derivative of a function that vanishes at
bothendpoints.Hencetheintegralissomelinearcombinationof IνandKν.
2. The rejection of the possibility that this solution contains Iνconstitutes Exer-
cise11.6.1.
3. The normalization may be verified by showing that, in the limit z→0,Kν(z)is in
agreementwithEq.(11.119). Bysubstituting x=1+t/z,
π1/2
(ν−1
2)!parenleftbiggz
2parenrightbiggνintegraldisplay∞
1e−zxparenleftbig
x2−1parenrightbigν−1/2dx
=π1/2
(ν−1
2)!parenleftbiggz
2parenrightbiggν
e−zintegraldisplay∞
0e−tparenleftbiggt2
z2+2t
zparenrightbiggν−1/2dt
z(11.123a)
=π1/2
(ν−1
2)!e−z
2νzνintegraldisplay∞
0e−tt2ν−1parenleftbigg
1+2z
tparenrightbiggν−1/2
dt, (11.123b)
takingout t2/z2asafactor.Thissubstitutionhaschangedthelimitsofintegrationtoamore
convenient range and has isolated the negative exponential dependence e−z. The integral
inEq.(11.123b)maybeevaluatedfor z=0toyield (2ν−1)!.Then,usingtheduplication
formula(Section8.4), wehave
lim
z→0Kν(z)=(ν−1)!2ν−1
zν,ν>0, (11.124)
inagreementwithEq.(11.119), whichthuschecksthenormalization.23
23Forν→0 the integral diverges logarithmically, in agreement with the logarithmic divergence of K0(z)forz→0 (Sec-
tion 11.5).
11.6 Asymptotic Expansions 721
Now,todevelopanasymptoticseries for Kν(z), wemayrewriteEq. (11.123a)as
Kν(z)=radicalbiggπ
2ze−z
(ν−1
2)!integraldisplay∞
0e−ttν−1/2parenleftbigg
1+t
2zparenrightbiggν−1/2
dt (11.125)
(takingout 2 t/zasafactor).
We expand (1+t/2z)ν−1/2bythebinomialtheoremtoobtain
Kν(z)=radicalbiggπ
2ze−z
(ν−1
2)!∞summationdisplay
r=0(ν−1
2)!
r!(ν−r−1
2)!(2z)−rintegraldisplay∞
0e−ttν+r−1/2dt.(11.126)
Term-by-term integration (valid for asymptotic series) yields the desired asymptotic ex-
pansionof Kν(z):
Kν(z)∼radicalbiggπ
2ze−zbracketleftbigg
1+(4ν2−12)
1!8z+(4ν2−12)(4ν2−32)
2!(8z)2+···bracketrightbigg
.(11.127)
Althoughthe integralof Eq. (11.122), integratingalong the real axis, was convergentonly
for−π/2<argz<π/2, Eq. (11.127) may be extended to −3π/2<argz<3π/2. Con-
sidered as an infinite series, Eq. (11.127) is actually divergent.24However, this series is
asymptotic, in the sense that for large enough z,Kν(z)may be approximated to any fixed
degree of accuracy with a small number of terms. (Compare Section 5.10 for a definition
anddiscussionofasymptoticseries.)
It isconvenienttorewriteEq.(11.127)as
Kν(z)=radicalbiggπ
2ze−zbracketleftbig
Pν(iz)+iQν(iz)bracketrightbig
, (11.128)
where
Pν(z)∼1−(µ−1)(µ−9)
2!(8z)2+(µ−1)(µ−9)(µ−25)(µ−49)
4!(8z)4−···,(11.129a)
Qν(z)∼µ−1
1!(8z)−(µ−1)(µ−9)(µ−25)
3!(8z)3+···, (11.129b)
and
µ=4ν2.
Itshouldbenotedthatalthough Pν(z)ofEq.(11.129a)and Qν(z)ofEq.(11.129b)have
alternating signs, the series for Pν(iz)andQν(iz)of Eq. (11.128) have all signs positive.
Finally,for zlarge,Pνdominates.
Thenwiththeasymptoticformof Kν(z),Eq.(11.128),wecanobtainexpansionsforall
otherBessel andhyperbolicBessel functionsbydefiningrelations:
24Our binomial expansion is valid only for t<2zand we have integrated tout to infinity. The exponential decrease of the
integrandpreventsadisaster,buttheresultantseriesisstillonlyasymptotic,notconvergent.ByTable9.3, z=∞isanessential
singularity of theBessel(and modified Bessel)equations. Fuchs’theoremdoes not guaranteeaconvergent series andwedo not
get aconvergent series.
722 Chapter 11 Bessel Functions
1. From
π
2iν+1H(1)
ν(iz)=Kν(z) (11.130)
wehave
H(1)
ν(z)=radicalbigg
2
πzexpbraceleftbigg
ibracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbiggbracerightbigg
·bracketleftbig
Pν(z)+iQν(z)bracketrightbig
,−π<argz<2π.(11.131)
2. The second Hankel function is just the complex conjugate of the first (for real argu-
ment),
H(2)
ν(z)=radicalbigg
2
πzexpbraceleftbigg
−ibracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbiggbracerightbigg
·bracketleftbig
Pν(z)−iQν(z)bracketrightbig
,−2π<argz<π. (11.132)
An alternate derivation of the asymptotic behavior of the Hankel functions appears in
Section7.3as anapplicationofthemethodofsteepestdescents.
3. Since Jν(z)istherealpartof H(1)
ν(z)for realz,
Jν(z)=radicalbigg
2
πzbraceleftbigg
Pν(z)cosbracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbigg
−Qν(z)sinbracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbiggbracerightbigg
,−π<argz<π,(11.133)
holds for real z,that is, arg z=0,π. Once Eq. (11.133) is established for real z,the
relationisvalidforcomplex zinthegivenrangeof argument.
4. TheNeumannfunctionistheimaginarypartof H(1)
ν(z)for realz,or
Nν(z)=radicalbigg
2
πzbraceleftbigg
Pν(z)sinbracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbigg
+Qν(z)cosbracketleftbigg
z−parenleftbigg
ν+1
2parenrightbiggπ
2bracketrightbiggbracerightbigg
,−π<argz<π.(11.134)
Initially, this relation is established for real z,but it may be extended to the complex
domainas shown.
5. Finally,theregularhyperbolicormodifiedBesselfunction Iν(z)isgivenby
Iν(z)=i−νJν(iz) (11.135)
or
Iν(z)=ez
√
2πzbracketleftbig
Pν(iz)−iQν(iz)bracketrightbig
,−π
2<argz<π
2. (11.136)
11.6 Asymptotic Expansions 723
FIGURE 11.11Asymptoticapproximationof J0(x).
This completes our determination of the asymptotic expansions. However, it is perhaps
worth noting the primary characteristics. Apart from the ubiquitous z−1/2,JνandNνbe-
have as cosine and sine, respectively. The zeros are almostevenly spaced at intervals of
π;thespacingbecomesexactly πinthelimitas z→∞. TheHankelfunctionshavebeen
defined to behave like the imaginary exponentials, and the modified Bessel functions Iν
andKνgo into the positive and negative exponentials. This asymptotic behavior may be
sufficienttoeliminateimmediatelyoneofthesefunctionsasasolutionforaphysicalprob-
lem. We should also note that the asymptotic series Pν(z)andQν(z), Eqs. (11.129a) and
(11.129b),terminatefor ν=±1/2,±3/2,...andbecomepolynomials(innegativepowers
ofz).Forthesespecialvaluesof νtheasymptoticapproximationsbecomeexactsolutions.
It is of some interest to consider the accuracy of the asymptotic forms, taking just the
firstterm, forexample(Fig. 11.11),
Jn(x)≈radicalbigg
2
πxcosbracketleftbigg
x−parenleftbigg
n+1
2parenrightbiggparenleftbiggπ
2parenrightbiggbracketrightbigg
. (11.137)
Clearly, the condition for the validity of Eq. (11.137) is that the sine term be negligible;
thatis,
8x≫4n2−1. (11.138)
Fornorν>1 theasymptoticregionmaybefar out.
AspointedoutinSection11.3,theasymptoticformsmaybeusedtoevaluatethevarious
Wronskianformulas(compareExercise11.6.3).
Exercises
11.6.1 Incheckingthenormalizationoftheintegralrepresentationof Kν(z)(Eq.(11.122)),we
assumed that Iν(z)was not present. How do we know that the integral representation
(Eq. (11.122))doesnotyield Kν(z)+εIν(z)withε/negationslash=0?
724 Chapter 11 Bessel Functions
FIGURE 11.12ModifiedBessel functioncontours.
11.6.2 (a) Showthat
y(z)=zνintegraldisplay
e−ztparenleftbig
t2−1parenrightbigν−1/2dt
satisfiesthemodifiedBessel equation,providedthecontouris chosenso that
e−ztparenleftbig
t2−1parenrightbigν+1/2
hasthesamevalueattheinitialandfinalpointsof thecontour.
(b) VerifythatthecontoursshowninFig.11.12are suitablefor thisproblem.
11.6.3 UsetheasymptoticexpansionstoverifythefollowingWronskianformulas:
(a)Jν(x)J−ν−1(x)+J−ν(x)Jν+1(x)=−2sinνπ/πx,
(b)Jν(x)Nν+1(x)−Jν+1(x)Nν(x)=−2/πx,
(c)Jν(x)H(2)
ν−1(x)−Jν−1(x)H(2)
ν(x)=2/iπx,
(d)Iν(x)K′
ν(x)−I′
ν(x)Kν(x)=−1/x,
(e)Iν(x)Kν+1(x)+Iν+1(x)Kν(x)=1/x.
11.6.4 From the asymptotic form of Kν(z), Eq. (11.127), derive the asymptotic form of
H(1)
ν(z), Eq.(11.131). Noteparticularlythephase, (ν+1
2)π/2.
11.6.5 Stokes’method.
(a) ReplacetheBesselfunctioninBessel’sequationby x−1/2y(x)andshowthat y(x)
satisfies
y′′(x)+parenleftbigg
1−ν2−1
4
x2parenrightbigg
y(x)=0.
(b) Develop a power-series solution with negative powers of xstarting with the as-
sumedform
y(x)=eix∞summationdisplay
n=0anx−n.
Determine the recurrence relation giving an+1in terms of an. Check your result
againsttheasymptoticseries, Eq.(11.131).
(c) Fromtheresults ofSection7.4determinetheinitialcoefficient, a0.
11.7 Spherical Bessel Functions 725
11.6.6 Calculate the first 15 partial sums of P0(x)andQ0(x), Eqs. (11.129a) and (11.129b).
Letxvary from 4 to 10 in unit steps. Determine the number of terms to be retained
for maximum accuracy and the accuracy achieved as a function of x. Specifically, how
smallmay xbewithoutraisingtheerrorabove 3 ×10−6?
ANS.xmin=6.
11.6.7 (a) Using the asymptotic series (partial sums) P0(x)andQ0(x)determined in Exer-
cise 11.6.6, write a function subprogram FCT(X) that will calculate J0(x),xreal,
forx≥xmin.
(b) Test your function by comparing it with the J0(x)(tables or computer library
subroutine)for x=xmin(10)xmin+10.
Note.Amoreaccurateandperhapssimplerasymptoticformfor J0(x)isgiveninAMS-
55,Eq. (9.4.3), seeAdditionalReadingsofChapter8for thereference.
11.7 S PHERICAL BESSEL FUNCTIONS
WhentheHelmholtzequationisseparatedinsphericalcoordinates,theradialequationhas
theform
r2d2R
dr2+2rdR
dr+bracketleftbig
k2r2−n(n+1)bracketrightbig
R=0. (11.139)
This is Eq. (9.65) of Section 9.3. The parameter kenters from the original Helmholtz
equation, while n(n+1)is a separation constant. From the behavior of the polar angle
function (Legendre’s equation, Sections 9.5 and 12.5), the separation constant must have
this form, with na nonnegative integer. Equation (11.139) has the virtue of being self-
adjoint,butclearlyitis notBessel’s equation.However,if wesubstitute
R(kr)=Z(kr)
(kr)1/2,
Equation(11.139)becomes
r2d2Z
dr2+rdZ
dr+bracketleftbigg
k2r2−parenleftbigg
n+1
2parenrightbigg2bracketrightbigg
Z=0, (11.140)
whichisBessel’s equation. Zis a Bessel function of order n+1
2(nan integer). Because
oftheimportanceof sphericalcoordinates,thiscombination,thatis,
Zn+1/2(kr)
(kr)1/2,
occursquiteoften.
726 Chapter 11 Bessel Functions
Definitions
ItisconvenienttolabelthesefunctionssphericalBesselfunctionswiththefollowingdefin-
ingequations:
jn(x)=radicalbiggπ
2xJn+1/2(x),
nn(x)=radicalbiggπ
2xNn+1/2(x)=(−1)n+1radicalbiggπ
2xJ−n−1/2(x),25
h(1)
n(x)=radicalbiggπ
2xH(1)
n+1/2(x)=jn(x)+inn(x),
h(2)
n(x)=radicalbiggπ
2xH(2)
n+1/2(x)=jn(x)−inn(x).(11.141)
ThesesphericalBesselfunctions(Figs.11.13and11.14)canbeexpressedinseriesform
byusingtheseries (Eq. (11.5))for Jn, replacing nwithn+1
2:
Jn+1/2(x)=∞summationdisplay
s=0(−1)s
s!(s+n+1
2)!parenleftbiggx
2parenrightbigg2s+n+1/2
. (11.142)
UsingtheLegendreduplicationformula,
z!(z+1
2)!=2−2z−1π1/2(2z+1)!, (11.143)
wehave
jn(x)=radicalbiggπ
2x∞summationdisplay
s=0(−1)s22s+2n+1(s+n)!
π1/2(2s+2n+1)!s!parenleftbiggx
2parenrightbigg2s+n+1/2
=2nxn∞summationdisplay
s=0(−1)s(s+n)!
s!(2s+2n+1)!x2s. (11.144)
Now,Nn+1/2(x)=(−1)n+1J−n−1/2(x)andfromEq. (11.5)wefindthat
J−n−1/2(x)=∞summationdisplay
s=0(−1)s
s!(s−n−1
2)!parenleftbiggx
2parenrightbigg2s−n−1/2
. (11.145)
This yields
nn(x)=(−1)n+12nπ1/2
xn+1∞summationdisplay
s=0(−1)s
s!(s−n−1
2)!parenleftbiggx
2parenrightbigg2s
. (11.146)
25This is possible because cos (n+1
2)π=0,seeEq. (11.60).
11.7 Spherical Bessel Functions 727
FIGURE 11.13SphericalBesselfunctions.
FIGURE 11.14SphericalNeumannfunctions.
728 Chapter 11 Bessel Functions
TheLegendreduplicationformulacanbeusedagaintogive
nn(x)=(−1)n+1
2nxn+1∞summationdisplay
s=0(−1)s(s−n)!
s!(2s−2n)!x2s. (11.147)
Theseseriesforms,Eqs.(11.144)and(11.147),areusefulinthreeways:(1)limitingvalues
asx→0, (2) closed-form representations for n=0, and, as an extension of this, (3) an
indicationthatthesphericalBesselfunctionsarecloselyrelatedtosineandcosine.
Forthespecialcase n=0 wefindfrom Eq.(11.144) that
j0(x)=∞summationdisplay
s=0(−1)s
(2s+1)!x2s=sinx
x, (11.148)
whereasfor n0,Eq. (11.147)yields
n0(x)=−cosx
x. (11.149)
FromthedefinitionofthesphericalHankelfunctions(Eq. (11.141)),
h(1)
0(x)=1
x(sinx−icosx)=−i
xeix,
h(2)
0(x)=1
x(sinx+icosx)=i
xe−ix. (11.150)
Equations (11.148) and (11.149) suggest expressing all spherical Bessel functions as
combinationsofsineandcosine.Theappropriatecombinationscanbedevelopedfromthe
power-seriessolutions,Eqs.(11.144)and(11.147),butthisapproachisawkward.Actually
thetrigonometricformsarealreadyavailableastheasymptoticexpansionofSection11.6.
FromEqs. (11.131)and(11.129a),
h(1)
n(x)=radicalbiggπ
2zH(1)
n+1/2(z)
=(−i)n+1eiz
zbraceleftbig
Pn+1/2(z)+iQn+1/2(z)bracerightbig
. (11.151)
Now,Pn+1/2andQn+1/2arepolynomials .ThismeansthatEq.(11.151)ismathematically
exact,notsimplyanasymptoticapproximation.Weobtain
h(1)
n(z)=(−i)n+1eiz
znsummationdisplay
s=0is
s!(8z)s(2n+2s)!!
(2n−2s)!!
=(−i)n+1eiz
znsummationdisplay
s=0is
s!(2z)s(n+s)!
(n−s)!. (11.152)
Often a factor (−i)n=(e−iπ/2)nwill be combined with the eizto giveei(z−nπ/2).F o r
zreal,jn(z)is the real part of this, nn(z)the imaginary part, and h(2)
n(z)the complex
conjugate.Specifically,
h(1)
1(x)=eixparenleftbigg
−1
x−i
x2parenrightbigg
, (11.153a)
11.7 Spherical Bessel Functions 729
h(1)
2(x)=eixparenleftbiggi
x−3
x2−3i
x3parenrightbigg
, (11.153b)
j1(x)=sinx
x2−cosx
x,
(11.154)
j2(x)=parenleftbigg3
x3−1
xparenrightbigg
sinx−3
x2cosx,
n1(x)=−cosx
x2−sinx
x,
(11.155)
n2(x)=−parenleftbigg3
x3−1
xparenrightbigg
cosx−3
x2sinx,
andso on.
Limiting Values
Forx≪1,26Eqs. (11.144)and(11.147)yield
jn(x)≈2nn!
(2n+1)!xn=xn
(2n+1)!!, (11.156)
nn(x)≈(−1)n+1
2n·(−n)!
(−2n)!x−n−1
=−(2n)!
2nn!x−n−1=−(2n−1)!!x−n−1. (11.157)
The transformation of factorials in the expressions for nn(x)employs Exercise 8.1.3. The
limitingvaluesof thesphericalHankelfunctionsgoas ±inn(x).
Theasymptoticvaluesof jn,nn,h(2)
n,andh(1)
nmaybeobtainedfromtheBesselasymp-
toticforms, Section11.6.We find
jn(x)∼1
xsinparenleftbigg
x−nπ
2parenrightbigg
, (11.158)
nn(x)∼−1
xcosparenleftbigg
x−nπ
2parenrightbigg
, (11.159)
h(1)
n(x)∼(−i)n+1eix
x=−iei(x−nπ/2)
x, (11.160a)
h(2)
n(x)∼in+1e−ix
x=ie−i(x−nπ/2)
x. (11.160b)
26The condition that the second term in the series be negligible compared to the first is actually x≪2[(2n+2)(2n+3)/
(n+1)]1/2forjn(x).
730 Chapter 11 Bessel Functions
The condition for these spherical Bessel forms is that x≫n(n+1)/2. From these as-
ymptotic values we see that jn(x)andnn(x)are appropriate for a description of standing
sphericalwaves ;h(1)
n(x)andh(2)
n(x)correspondto travelingsphericalwaves .Ifthetime
dependence for the traveling waves is taken to be e−iωt, thenh(1)
n(x)yields an outgoing
travelingsphericalwave, h(2)
n(x)anincomingwave.Radiationtheoryinelectromagnetism
andscatteringtheoryinquantummechanicsprovidemanyapplications.
Recurrence Relations
Therecurrencerelationstowhichwenowturnprovideaconvenientwayofdevelopingthe
higher-order spherical Bessel functions. These recurrence relations may be derived from
theseries,but,aswiththemodifiedBesselfunctions,itiseasiertosubstituteintotheknown
recurrencerelations(Eqs. (11.10)and(11.12)). Thisgives
fn−1(x)+fn+1(x)=2n+1
xfn(x), (11.161)
nfn−1(x)−(n+1)fn+1(x)=(2n+1)f′
n(x). (11.162)
Rearrangingtheserelations(orsubstitutingintoEqs. (11.15)and(11.17)), weobtain
d
dxbracketleftbig
xn+1fn(x)bracketrightbig
=xn+1fn−1(x), (11.163)
d
dxbracketleftbig
x−nfn(x)bracketrightbig
=−x−nfn+1(x). (11.164)
Herefnmayrepresent jn,nn,h(1)
n,orh(2)
n.
The specific forms, Eqs. (11.154) and (11.155), may also be readily obtained from
Eq.(11.164).
BymathematicalinductionwemayestablishtheRayleighformulas
jn(x)=(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbiggsinx
xparenrightbigg
, (11.165)
nn(x)=−(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbiggcosx
xparenrightbigg
, (11.166)
h(1)
n(x)=−i(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbiggeix
xparenrightbigg
,
(11.167)
h(2)
n(x)=i(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbigge−ix
xparenrightbigg
.
11.7 Spherical Bessel Functions 731
Orthogonality
WemaytaketheorthogonalityintegralfortheordinaryBesselfunctions(Eqs.(11.49)and
(11.50)),
integraldisplaya
0Jνparenleftbigg
ανpρ
aparenrightbigg
Jνparenleftbigg
ανqρ
aparenrightbigg
ρdρ=a2
2bracketleftbig
Jν+1(ανp)bracketrightbig2δpq, (11.168)
andsubstituteintheexpressionfor jntoobtain
integraldisplaya
0jnparenleftbigg
αnpρ
aparenrightbigg
jnparenleftbigg
αnqρ
aparenrightbigg
ρ2dρ=a3
2bracketleftbig
jn+1(αnp)bracketrightbig2δpq. (11.169)
Hereαnpandαnqarerootsof jn.
This representsorthogonalitywithrespecttotherootsof theBesselfunctions.Anillus-
trationofthissortoforthogonalityisprovidedinExample11.7.1,theproblemofaparticle
in a sphere. Equation (11.169) guarantees orthogonality of the wave functions jn(r)for
fixedn.(Ifnvaries,theaccompanyingsphericalharmonicwillprovideorthogonality.)
Example 11.7.1 PARTICLE IN A SPHERE
An illustration of the use of the spherical Bessel functions is provided by the problem of
a quantum mechanical particle in a sphere of radius a. Quantum theory requires that the
wavefunction ψ,describingourparticle,satisfy
−¯h2
2m∇2ψ=Eψ, (11.170)
and the boundary conditions (1) ψ(r≤a)remains finite, (2) ψ(a)=0. This corresponds
to a square-well potential V=0,r≤a, andV=∞,r>a.H e r e¯his Planck’s constant
divided by 2 π,mis the mass of our particle, and Eis, its energy. Let us determine the
minimum value of the energy for which our wave equation has an acceptable solution.
Equation (11.170) is Helmholtz’s equation with a radial part (compare Section 9.3 for
separationofvariables):
d2R
dr2+2
rdR
dr+bracketleftbigg
k2−n(n+1)
r2bracketrightbigg
R=0, (11.171)
withk2=2mE/¯h2. HencebyEq.(11.139), with n=0,
R=Aj0(kr)+Bn0(kr).
Wechoosetheorbitalangularmomentumindex n=0,for anyangulardependencewould
raise the energy. The spherical Neumann function is rejected because of its divergent be-
haviorattheorigin.Tosatisfythesecondboundarycondition(for allangles),werequire
ka=√
2mE
¯ha=α, (11.172)
732 Chapter 11 Bessel Functions
whereαis a root of j0, that is, j0(α)=0. This has the effect of limiting the allowable
energies to a certain discrete set, or, in other words, application of boundary condition (2)
quantizestheenergy E. Thesmallest αisthefirst zeroof j0,
α=π,
and
Emin=π2¯h2
2ma2=h2
8ma2, (11.173)
which means that for any finite sphere the particle energy will have a positive minimum
or zero-pointenergy.This is an illustrationof theHeisenberguncertaintyprinciplefor /Delta1p
with/Delta1r≤a.
In solid-state physics, astrophysics, and other areas of physics, we may wish to know
how many different solutions (energy states) correspond to energies less than or equal to
some fixed energy E0. For a cubic volume (Exercise 9.3.5) the problem is fairly simple.
The considerably more difficult spherical case is worked out by R. H. Lambert, Am. J.
Phys.36:417,1169(1968).
Therelevantorthogonalityrelationforthe jn(kr)canbederivedfromtheintegralgiven
inExercise11.7.23. /squaresolid
Anotherform, orthogonalitywithrespecttotheindices,maybewrittenas
integraldisplay∞
−∞jm(x)jn(x)dx=0,m/negationslash=n, m,n≥0. (11.174)
Theproofisleft asExercise11.7.10.If m=n(compareExercise11.7.11),wehave
integraldisplay∞
−∞bracketleftbig
jn(x)bracketrightbig2dx=π
2n+1. (11.175)
Most physical applications of orthogonal Bessel and spherical Bessel functions involve
orthogonalitywithvaryingroots and an interval [0,a]and Eqs. (11.168) and (11.169) and
Exercise11.7.23for continuous-energyeigenvalues.
The spherical Bessel functions will enter again in connection with spherical waves, but
furtherconsiderationispostponeduntilthecorrespondingangularfunctions,theLegendre
functions,havebeenintroduced.
Exercises
11.7.1 Showthatif
nn(x)=radicalbiggπ
2xNn+1/2(x),
itautomaticallyequals
(−1)n+1radicalbiggπ
2xJ−n−1/2(x).
11.7 Spherical Bessel Functions 733
11.7.2 Derivethetrigonometric-polynomialforms of jn(z)andnn(z).27
jn(z)=1
zsinparenleftbigg
z−nπ
2parenrightbigg[n/2]summationdisplay
s=0(−1)s(n+2s)!
(2s)!(2z)2s(n−2s)!
+1
zcosparenleftbigg
z−nπ
2parenrightbigg[(n−1)/2]summationdisplay
s=0(−1)s(n+2s+1)!
(2s+1)!(2z)2s(n−2s−1)!,
nn(z)=(−1)n+1
zcosparenleftbigg
z+nπ
2parenrightbigg[n/2]summationdisplay
s=0(−1)s(n+2s)!
(2s)!(2z)2s(n−2s)!
+(−1)n+1
zsinparenleftbigg
z+nπ
2parenrightbigg[(n−1)/2]summationdisplay
s=0(−1)s(n+2s+1)!
(2s+1)!(2z)2s+1(n−2s−1)!.
11.7.3 Usetheintegralrepresentationof Jν(x),
Jν(x)=1
π1/2(ν−1
2)!parenleftbiggx
2parenrightbiggνintegraldisplay1
−1e±ixpparenleftbig
1−p2parenrightbigν−1/2dp,
to show that the spherical Bessel functions jn(x)are expressible in terms of trigono-
metricfunctions;thatis, for example,
j0(x)=sinx
x,j 1(x)=sinx
x2−cosx
x.
11.7.4 (a) Derivetherecurrencerelations
fn−1(x)+fn+1(x)=2n+1
xfn(x),
nfn−1(x)−(n+1)fn+1(x)=(2n+1)f′
n(x)
satisfiedbythesphericalBessel functions jn(x),nn(x),h(1)
n(x), andh(2)
n(x).
(b) Show,fromthesetworecurrencerelations,thatthesphericalBesselfunction fn(x)
satisfiesthedifferentialequation
x2f′′
n(x)+2xf′
n(x)+bracketleftbig
x2−n(n+1)bracketrightbig
fn(x)=0.
11.7.5 Provebymathematicalinductionthat
jn(x)=(−1)nxnparenleftbigg1
xd
dxparenrightbiggnparenleftbiggsinx
xparenrightbigg
fornanarbitrarynonnegativeinteger.
11.7.6 From the discussion of orthogonality of the spherical Bessel functions, show that a
Wronskianrelationfor jn(x)andnn(x)is
jn(x)n′
n(x)−j′
n(x)nn(x)=1
x2.
27Theupper limit on the summation [n/2]means the largest integerthat does not exceed n/2.
734 Chapter 11 Bessel Functions
11.7.7 Verify
h(1)
n(x)h(2)′
n(x)−h(1)′
n(x)h(2)
n(x)=−2i
x2.
11.7.8 VerifyPoisson’sintegralrepresentationofthesphericalBesselfunction,
jn(z)=zn
2n+1n!integraldisplayπ
0cos(zcosθ)sin2n+1θdθ.
11.7.9 Showthatintegraldisplay∞
0Jµ(x)Jν(x)dx
x=2
πsin[(µ−ν)π/2]
µ2−ν2,µ+ν>−1.
11.7.10 DeriveEq. (11.174):
integraldisplay∞
−∞jm(x)jn(x)dx=0,m/negationslash=n
m,n≥0.
11.7.11 DeriveEq. (11.175):
integraldisplay∞
−∞bracketleftbig
jn(x)bracketrightbig2dx=π
2n+1.
11.7.12 Set up the orthogonality integral for jL(kr)in a sphere of radius Rwith the boundary
condition
jL(kR)=0.
The result is used in classifying electromagnetic radiation according to its angular mo-
mentum.
11.7.13 The Fresnel integrals (Fig. 11.15 and Exercise 5.10.2) occurring in diffraction theory
aregivenby
x(t)=radicalbiggπ
2Cparenleftbiggradicalbiggπ
2tparenrightbigg
=integraldisplayt
0cosparenleftbig
v2parenrightbig
dv, y(t)=radicalbiggπ
2Sparenleftbiggradicalbiggπ
2tparenrightbigg
=integraldisplayt
0sinparenleftbig
v2parenrightbig
dv.
Showthattheseintegralsmaybeexpandedinseries ofsphericalBesselfunctions
x(s)=1
2integraldisplays
0j−1(u)u1/2du=s1/2∞summationdisplay
n=0j2n(s),
y(s)=1
2integraldisplays
0j0(u)u1/2du=s1/2∞summationdisplay
n=0j2n+1(s).
Hint. To establish the equality of the integral and the sum, you may wish to work with
theirderivatives.ThesphericalBesselanalogsof Eqs. (11.12)and(11.14)arehelpful.
11.7.14 Ahollowsphereofradius a(Helmholtzresonator)containsstandingsoundwaves.Find
theminimumfrequencyofoscillationintermsoftheradius aandthevelocityofsound
v.Thesoundwavessatisfy thewaveequation
∇2ψ=1
v2∂2ψ
∂t2
11.7 Spherical Bessel Functions 735
FIGURE 11.15Fresnelintegrals.
andtheboundarycondition
∂ψ
∂r=0,r=a.
This is a Neumann boundary condition. Example 11.7.1 has the same PDE but with a
Dirichletboundarycondition.
ANS.νmin=0.3313v/a,λmax=3.018a.
11.7.15 DefiningthesphericalmodifiedBessel functions(Fig.11.16)by
in(x)=radicalbiggπ
2xIn+1/2(x), k n(x)=radicalbigg
2
πxKn+1/2(x),
showthat
i0(x)=sinhx
x,k 0(x)=e−x
x.
Notethatthenumericalfactors inthedefinitionsof inandknare notidentical.
11.7.16 (a) Showthattheparityof in(x)is(−1)n.
(b) Showthat kn(x)hasnodefiniteparity.
736 Chapter 11 Bessel Functions
FIGURE 11.16SphericalmodifiedBessel
functions.
11.7.17 ShowthatthesphericalmodifiedBessel functionssatisfy thefollowingrelations:
(a)in(x)=i−njn(ix),
kn(x)=−inh(1)
n(ix),
(b)in+1(x)=xnd
dxparenleftbig
x−ninparenrightbig
,
kn+1(x)=−xnd
dxparenleftbig
x−nknparenrightbig
,
(c)in(x)=xnparenleftbigg1
xd
dxparenrightbiggnsinhx
x,
kn(x)=(−1)nxnparenleftbigg1
xd
dxparenrightbiggne−x
x.
11.7.18 Showthattherecurrencerelationsfor in(x)andkn(x)are
(a)in−1(x)−in+1(x)=2n+1
xin(x),
nin−1(x)+(n+1)in+1(x)=(2n+1)i′
n(x),
11.7 Spherical Bessel Functions 737
(b)kn−1(x)−kn+1(x)=−2n+1
xkn(x),
nkn−1(x)+(n+1)kn+1(x)=−(2n+1)k′
n(x).
11.7.19 Derivethelimitingvaluesfor thesphericalmodifiedBesselfunctions
(a)in(x)≈xn
(2n+1)!!,k n(x)≈(2n−1)!!
xn+1,x≪1.
(b)in(x)∼ex
2x,k n(x)∼e−x
x,x≫1
2n(n+1).
11.7.20 ShowthattheWronskianofthesphericalmodifiedBesselfunctionsis givenby
in(x)k′
n(x)−i′
n(x)kn(x)=−1
x2.
11.7.21 Aquantumparticleofmass Mistrappedina“square”wellofradius a.TheSchrödinger
equationpotentialis
V(r)=braceleftBigg
−V0,0≤r<a
0,r>a.
Theparticle’senergy Eis negative(aneigenvalue).
(a) Show that the radial part of the wave function is given by jl(k1r)for 0≤r<a
andkl(k2r)forr>a. (We require that ψ(0)be finite and ψ(∞)→0.) Here
k2
1=2M(E+V0)/¯h2,k2
2=−2ME/¯h2, andlis the angular momentum ( nin
Eq.(11.139)).
(b) Theboundaryconditionat r=aisthatthewavefunction ψ(r)anditsfirstderiv-
ativebecontinuous.Showthatthismeans
(d/dr)j l(k1r)
jl(k1r)vextendsinglevextendsinglevextendsinglevextendsingle
r=a=(d/dr)k l(k2r)
kl(k2r)vextendsinglevextendsinglevextendsinglevextendsingle
r=a.
Thisequationdeterminestheenergyeigenvalues.
Note.This isageneralizationofExample10.1.2.
11.7.22 Thequantummechanicalradialwavefunctionfor ascatteredwaveisgivenby
ψk=sin(kr+δ0)
kr,
wherekis the wave number, k=√2mE/¯h, andδ0is the scattering phase shift. Show
thatthenormalizationintegralis
integraldisplay∞
0ψk(r)ψk′(r)r2dr=π
2kδ(k−k′).
Hint.YoucanuseasinerepresentationoftheDiracdeltafunction.SeeExercise15.3.8.
738 Chapter 11 Bessel Functions
11.7.23 DerivethesphericalBessel functionclosurerelation
2a2
πintegraldisplay∞
0jn(ar)jn(br)r2dr=δ(a−b).
Note. An interesting derivation involving Fourier transforms, the Rayleigh plane-wave
expansion, and spherical harmonics has been given by P. Ugincius, Am. J. Phys. 40:
1690(1972).
11.7.24 (a) Writea subroutinethatwillgeneratethesphericalBesselfunctions, jn(x), thatis,
willgeneratethenumericalvalueof jn(x)givenxandn.
Note.Onepossibilityistousetheexplicitknownformsof j0andj1andtodevelop
thehigherindex jn, byrepeatedapplicationof therecurrencerelation.
(b) Check your subroutine by an independent calculation, such as Eq. (11.154). If
possible, compare the machine time needed for this check with the time required
foryoursubroutine.
11.7.25 The wave function of a particle in a sphere (Example 11.7.1) with angular momen-
tumlisψ(r,θ,ϕ)=Ajl((√
2ME)r/¯h)Ym
l(θ,ϕ).T h eYm
l(θ,ϕ)is a spherical har-
monic, described in Section 12.6. From the boundary condition ψ(a,θ,ϕ)=0o r
jl((√
2ME)a/¯h)=0 calculate the 10 lowest-energy states. Disregard the mdegen-
eracy (2l+1 values of mfor each choice of l). Check your results against AMS-55,
Table10.6, seeAdditionalReadingsfor Chapter8forthereference.
Hint.YoucanuseyoursphericalBessel subroutineandaroot-findingsubroutine.
Checkvalues. jl(αls)=0,
α01=3.1416
α11=4.4934
α21=5.7635
α02=6.2832.
11.7.26 LetExample11.7.1bemodifiedso thatthepotentialisafinite V0outside(r >a).
(a) For E<V0showthat
ψout(r,θ,ϕ)∼klparenleftbiggr
¯hradicalbig
2M(V0−E)parenrightbigg
.
(b) Thenewboundaryconditionstobesatisfiedat r=aare
ψin(a,θ,ϕ)=ψout(a,θ,ϕ),
∂
∂rψin(a,θ,ϕ)=∂
∂rψout(a,θ,ϕ)
or
1
ψin∂ψin
∂rvextendsinglevextendsinglevextendsinglevextendsingle
r=a=1
ψout∂ψout
∂rvextendsinglevextendsinglevextendsinglevextendsingle
r=a.
Forl=0 showthattheboundaryconditionat r=aleadsto
f(E)=kbraceleftbigg
cotka−1
kabracerightbigg
+k′braceleftbigg
1+1
k′abracerightbigg
=0,
wherek=√
2ME/¯handk′=√2M(V0−E)/¯h.
11.7 Additional Readings 739
(c) With a=4πε0¯h2/Me2(Bohr radius) and V0=4Me4/2¯h2, compute the possible
boundstates (0<E<V 0).
Hint. Call a root-finding subroutine after you know the approximate location of
theroots of
f(E)=0(0≤E≤V0).
(d) Show that when a=4πε0¯h2/Me2the minimum value of V0for which a bound
stateexistsis V0=2.4674Me4/2¯h2.
11.7.27 In some nuclear stripping reactions the differential cross section is proportional to
jl(x)2,wherelistheangularmomentum.Thelocationofthemaximumonthecurveof
experimentaldatapermitsadeterminationof l,ifthelocationofthe(first)maximumof
jl(x)is known. Compute the location of the first maximum of j1(x),j2(x), andj3(x).
Note.Forbetteraccuracylookforthefirstzeroof j′
l(x).Whyisthismoreaccuratethan
directlocationof themaximum?
AdditionalReadings
Jackson, J. D., Classical Electrodynamics , 3rd ed.,NewYork: J. Wiley (1999).
McBride,E.B., ObtainingGeneratingFunctions .NewYork:Springer-Verlag(1971).Anintroductiontomethods
of obtaining generating functions.
Watson, G. N., A Treatise on the Theory of Bessel Functions , 2nd ed. Cambridge, UK: Cambridge University
Press (1952). This is the definitive text on Bessel functions and their properties. Although difficult reading, it
is invaluable asthe ultimate reference.
Watson,G.N., ATreatiseontheTheoryofBesselFunctions ,1sted.Cambridge,UK:CambridgeUniversityPress
(1922). Seealso the references listed atthe endofChapter 13.
This page intentionally left blank
CHAPTER 12
LEGENDRE FUNCTIONS
12.1 G ENERATING FUNCTION
Legendre polynomials appear in many different mathematical and physical situations.
(1) They may originate as solutions of the Legendre ODE which we have already en-
countered in the separation of variables (Section 9.3) for Laplace’s equation, Helmholtz’s
equation,andsimilarODEsinsphericalpolarcoordinates.(2)Theyenterasaconsequence
of a Rodrigues’ formula (Section 12.4). (3) They arise as a consequence of demanding a
complete, orthogonal set of functions over the interval [−1,1](Gram–Schmidt orthogo-
nalization, Section 10.3). (4) In quantum mechanics they (really the spherical harmonics,
Sections 12.6 and 12.7) represent angular momentum eigenfunctions. (5) They are gen-
erated by a generating function. We introduce Legendre polynomials here by way of a
generatingfunction.
Physical Basis — Electrostatics
AswithBesselfunctions,itisconvenienttointroducetheLegendrepolynomialsbymeans
of a generating function, which here appears in a physical context. Consider an electric
chargeqplacedonthe z-axisatz=a.AsshowninFig.12.1,theelectrostaticpotentialof
chargeqis
ϕ=1
4πε0·q
r1(SIunits). (12.1)
We want to express the electrostatic potential in terms of the spherical polar coordinates r
andθ(the coordinate ϕis absent because of symmetry about the z-axis). Using the law of
cosinesinFig. 12.1,weobtain
ϕ=q
4πε0parenleftbig
r2+a2−2arcosθparenrightbig−1/2. (12.2)
741
742 Chapter 12 Legendre Functions
FIGURE 12.1Electrostaticpotential.
Chargeqdisplacedfrom origin.
Legendre Polynomials
Consider the case of r>aor, more precisely, r2>|a2−2arcosθ|. The radical in
Eq. (12.2) may be expanded in a binomial series and then rearranged in powers of (a/r).
The Legendre polynomial Pn(cosθ)(see Fig. 12.2) is defined as the coefficient of the nth
powerin
ϕ=q
4πε0r∞summationdisplay
n=0Pn(cosθ)parenleftbigga
rparenrightbiggn
. (12.3)
FIGURE 12.2Legendre
polynomials P2(x),P3(x),
P4(x), andP5(x).
12.1 Generating Function 743
Dropping the factor q/4πε0rand using xandtinstead of cos θanda/r, respectively, we
have
g(t,x)=parenleftbig
1−2xt+t2parenrightbig−1/2=∞summationdisplay
n=0Pn(x)tn,|t|<1. (12.4)
Equation (12.4) is our generating function formula. In the next section it is shown that
|Pn(cosθ)|≤1, whichmeansthat theseries expansion(Eq. (12.4)) is convergentfor |t|<
1.1Indeed,theseries isconvergentfor |t|=1 exceptfor|x|=1.
In physicalapplicationsEq. (12.4)oftenappearsinthevectorform(see Section9.7)
1
|r1−r2|=1
r>∞summationdisplay
n=0parenleftbiggr<
r>parenrightbiggn
Pn(cosθ), (12.4a)
where
r>=|r1|
r<=|r2|bracerightbigg
for|r1|>|r2|, (12.4b)
and
r>=|r2|
r<=|r1|bracerightbigg
for|r2|>|r1|. (12.4c)
Usingthebinomialtheorem(Section5.6)andExercise8.1.15,weexpandthegenerating
functionas(compareEq.(12.33))
parenleftbig
1−2xt+t2parenrightbig−1/2=∞summationdisplay
n=0(2n)!
22n(n!)2parenleftbig
2xt−t2parenrightbign
=1+∞summationdisplay
n=1(2n−1)!!
(2n)!!parenleftbig
2xt−t2parenrightbign. (12.5)
ForthefirstfewLegendrepolynomials,say, P0,P1,andP2,weneedthecoefficientsof t0,
t1, andt2. These powers of tappear only in the terms n=0,1, and 2, and hence we may
limitourattentiontothefirst threetermsoftheinfiniteseries:
0!
20(0!)2parenleftbig
2xt−t2parenrightbig0+2!
22(1!)2parenleftbig
2xt−t2parenrightbig1+4!
24(2!)2parenleftbig
2xt−t2parenrightbig2
=1t0+xt1+parenleftbigg3
2x2−1
2parenrightbigg
t2+Oparenleftbig
t3parenrightbig
.
Then,fromEq. (12.4)(anduniquenessof powerseries),
P0(x)=1,P 1(x)=x, P 2(x)=3
2x2−1
2.
Werepeatthislimiteddevelopmentinavectorframeworklaterinthissection.
1Note that the series in Eq. (12.3) is convergent for r>a, even though the binomial expansion involved is valid only for
r>( a2+2ar)1/2and cosθ=−1,orr>a(1+√
2).
744 Chapter 12 Legendre Functions
Inemployingageneraltreatment,wefindthatthebinomialexpansionofthe (2xt−t2)n
factoryieldsthedoubleseries
parenleftbig
1−2xt+t2parenrightbig−1/2=∞summationdisplay
n=0(2n)!
22n(n!)2tnnsummationdisplay
k=0(−1)kn!
k!(n−k)!(2x)n−ktk
=∞summationdisplay
n=0nsummationdisplay
k=0(−1)k(2n)!
22nn!k!(n−k)!·(2x)n−ktn+k.(12.6)
FromEq. (5.64) ofSection5.4(rearrangingtheorder ofsummation),Eq. (12.6) becomes
parenleftbig
1−2xt+t2parenrightbig−1/2=∞summationdisplay
n=0[n/2]summationdisplay
k=0(−1)k(2n−2k)!
22n−2kk!(n−k)!(n−2k)!·(2x)n−2ktn,(12.7)
with thetnindependent of the index k.2Now, equating our two power series (Eqs. (12.4)
and(12.7)) termbyterm,wehave3
Pn(x)=[n/2]summationdisplay
k=0(−1)k(2n−2k)!
2nk!(n−k)!(n−2k)!xn−2k. (12.8)
Hence,for neven,Pnhasonlyevenpowersof xandevenparity(seeEq.(12.37)),andodd
powersandoddparityfor odd n.
Linear Electric Multipoles
Returning to the electric charge on the z-axis, we demonstrate the usefulness and power
of the generating function by adding a charge −qatz=−a, as shown in Fig. 12.3. The
FIGURE 12.3Electricdipole.
2[n/2]=n/2f o rneven,(n−1)/2f o rnodd.
3Equation (12.8) starts with xn. By changing the index, we can transform it into a series that starts with x0forneven and x1
fornodd. Theseascending series aregiven as hypergeometric functions in Eqs. (13.138) and (13.139), Section 13.4.
12.1 Generating Function 745
potentialbecomes
ϕ=q
4πε0parenleftbigg1
r1−1
r2parenrightbigg
, (12.9)
andbyusingthelawofcosines,wehave
ϕ=q
4πε0rbraceleftbiggbracketleftbigg
1−2parenleftbigga
rparenrightbigg
cosθ+parenleftbigga
rparenrightbigg2bracketrightbigg−1/2
−bracketleftbigg
1+2parenleftbigga
rparenrightbigg
cosθ+parenleftbigga
rparenrightbigg2bracketrightbigg−1/2bracerightbigg
,(r>a).
Clearly, the second radical is like the first, except that ahas been replaced by −a. Then,
usingEq. (12.4), weobtain
ϕ=q
4πε0rbracketleftbigg∞summationdisplay
n=0Pn(cosθ)parenleftbigga
rparenrightbiggn
−∞summationdisplay
n=0Pn(cosθ)(−1)nparenleftbigga
rparenrightbiggnbracketrightbigg
=2q
4πε0rbracketleftbigg
P1(cosθ)parenleftbigga
rparenrightbigg
+P3(cosθ)parenleftbigga
rparenrightbigg3
+···bracketrightbigg
. (12.10)
Thefirstterm(anddominanttermfor r≫a)i s
ϕ=2aq
4πε0·P1(cosθ)
r2, (12.11)
which is the electric dipole potential, and 2 aqis the dipole moment (Fig. 12.3). This
analysis may be extended by placing additional charges on the z-axis so that the P1term,
as well as the P0(monopole) term, is canceled. For instance, charges of qatz=aand
z=−a,−2qatz=0giverisetoapotentialwhoseseriesexpansionstartswith P2(cosθ).
This is a linear electric quadrupole. Two linear quadrupoles may be placed so that the
quadrupoletermis canceledbutthe P3,theoctupoleterm,survives.
Vector Expansion
Weconsidertheelectrostaticpotentialproducedbyadistributedcharge ρ(r2):
ϕ(r1)=1
4πε0integraldisplayρ(r2)
|r1−r2|d3r2. (12.12a)
This expression has already appeared in Sections 1.16 and 9.7. Taking the denominator
of the integrand, using first the law of cosines and then a binomial expansion, yields (see
Fig.1.42)
1
|r1−r2|=parenleftbig
r2
1−2r1·r2+r2
2parenrightbig−1/2(12.12b)
=1
r1bracketleftbigg
1+parenleftbigg
−2r1·r2
r2
1+r2
2
r2
1parenrightbiggbracketrightbigg−1/2
,forr1>r2
=1
r1bracketleftbigg
1+r1·r2
r2
1−1
2r2
2
r2
1+3
2(r1·r2)2
r4
1+Oparenleftbiggr2
r1parenrightbigg3bracketrightbigg
.
746 Chapter 12 Legendre Functions
(Forr1=1,r2=t, andr1·r2=xt, Eq. (12.12b) reduces to the generating function,
Eq.(12.4).)
Thefirst terminthesquarebracket,1, yieldsa potential
ϕ0(r1)=1
4πε01
r1integraldisplay
ρ(r2)d3r2. (12.12c)
Theintegralisjustthetotalcharge.Thispartofthetotalpotentialisanelectric monopole .
Thesecondtermyields
ϕ1(r1)=1
4πε0r1·
r3
1integraldisplay
r2ρ(r2)d3r2, (12.12d)
where the integral is the dipole moment whose charge density ρ(r2)is weighted by a mo-
ment arm r2. We have an electric dipole potential. For atomic or nuclear states of definite
parity,ρ(r2)is anevenfunctionandthedipoleintegralis identicallyzero.
The last two terms, both of order (r2/r1)2, may be handled by using Cartesian coordi-
nates:
(r1·r2)2=3summationdisplay
i=1x1ix2i3summationdisplay
j=1x1jx2j.
Rearrangingvariablestotakethe x1componentsoutsidetheintegralyields
ϕ2(r1)=1
4πε01
2r5
13summationdisplay
i,j=1x1ix1jintegraldisplaybracketleftbig
3x2ix2j−δijr2
2bracketrightbig
ρ(r2)d3r2. (12.12e)
This is the electric quadrupole term. We note that the square bracket in the integrand
formsasymmetric,zero-tracetensor.
AgeneralelectrostaticmultipoleexpansioncanalsobedevelopedbyusingEq.(12.12a)
forthepotential ϕ(r1)andreplacing1 /(4π|r1−r2|)byGreen’sfunction,Eq.(9.187).This
yields the potential ϕ(r1)as a (double) series of the spherical harmonics Ym
l(θ1,ϕ1)and
Ym
l(θ2,ϕ2).
Before leavingmultipolefields,perhapsweshouldemphasizethreepoints.
•First, an electric (or magnetic) multipole is isolated and well defined only if all lower-
order multipoles vanish. For instance, the potential of one charge qatz=awas ex-
panded in a series of Legendre polynomials. Although we refer to the P1(cosθ)term
in this expansion as a dipole term, it should be remembered that this term exists only
becauseofour choiceofcoordinates.We alsohaveamonopole, P0(cosθ).
•Second, in physical systems we do not encounter pure multipoles. As an example,
the potential of the finite dipole ( qatz=a,−qatz=−a) contained a P3(cosθ)
term.Thesehigher-ordertermsmaybeeliminatedbyshrinkingthemultipoletoapoint
multipole, in this case keeping the product qaconstant(a→0,q→∞)to maintain
thesamedipolemoment.
12.1 Generating Function 747
•Third,themultipoletheoryisnotrestrictedtoelectricalphenomena.Planetaryconfigu-
rationsaredescribedintermsofmassmultipoles,Sections12.3and12.6.Gravitational
radiation depends on the time behavior of mass quadrupoles. (The gravitational radia-
tion field is a tensorfield. The radiation quanta, gravitons, carry two units of angular
momentum.)
It might also be noted that a multipole expansion is actually a decomposition into the
irreduciblerepresentationsoftherotationgroup(Section4.2).
Extension to Ultraspherical Polynomials
The generating function used here, g(t,x), is actually a special case of a more general
generatingfunction,
1
(1−2xt+t2)α=∞summationdisplay
n=0C(α)
n(x)tn. (12.13)
The coefficients C(α)
n(x)are the ultraspherical polynomials (proportional to the Gegen-
bauer polynomials). For α=1/2 this equation reduces to Eq. (12.4); that is, C(1/2)
n(x)=
Pn(x). The cases a=0 andα=1 are considered in Chapter 13 in connection with the
Chebyshevpolynomials.
Exercises
12.1.1 Develop the electrostatic potential for the array of charges shown. This is a linear elec-
tricquadrupole(Fig. 12.4).
12.1.2 Calculate the electrostatic potential of the array of charges shown in Fig. 12.5. Here
is an example of two equal but oppositely directed dipoles. The dipole contributions
cancel.Theoctupoletermsdonotcancel.
12.1.3 Showthattheelectrostaticpotentialproducedbya charge qatz=aforr<ais
ϕ(r)=q
4πε0a∞summationdisplay
n=0parenleftbiggr
aparenrightbiggn
Pn(cosθ).
FIGURE 12.4Linearelectricquadrupole.
748 Chapter 12 Legendre Functions
FIGURE 12.5Linearelectricoctupole.
FIGURE 12.6
12.1.4 UsingE=−∇ϕ, determine the components of the electric field corresponding to the
(pure)electricdipolepotential
ϕ(r)=2aqP1(cosθ)
4πε0r2.
Hereitisassumedthat r≫a.
ANS.Er=+4aqcosθ
4πε0r3,Eθ=+2aqsinθ
4πε0r3,Eϕ=0.
12.1.5 Apointelectricdipoleofstrength p(1)isplacedat z=a;asecondpointelectricdipole
of equal but opposite strength is at the origin. Keeping the product p(1)aconstant, let
a→0.Showthatthis resultsinapointelectricquadrupole.
Hint.Exercise12.2.5(whenproved)willbehelpful.
12.1.6 A point charge qis in the interior of a hollow conducting sphere of radius r0.T h e
chargeqis displaced a distance afrom the center of the sphere. If the conducting
sphere is grounded, show that the potential in the interior produced by qand the dis-
tributed induced charge is the same as that produced by qand its image charge q′.T h e
image charge is at a distance a′=r2
0/afrom the center, collinear with qand the origin
(Fig.12.6).
Hint. Calculate the electrostatic potential for a<r0<a′. Show that the potential van-
ishesforr=r0ifwetake q′=−qr0/a.
12.1.7 Provethat
Pn(cosθ)=(−1)nrn+1
n!∂n
∂znparenleftbigg1
rparenrightbigg
.
Hint.ComparetheLegendrepolynomialexpansionofthegeneratingfunction( a→/Delta1z,
Fig.12.1)withaTaylorseriesexpansionof1 /r,wherezdependenceof rchangesfrom
ztoz−/Delta1z(Fig.12.7).
12.1.8 Bydifferentiationanddirectsubstitutionoftheseriesform,Eq.(12.8),showthat Pn(x)
satisfies the Legendre ODE. Note that there is no restriction upon x. We may have any
x,−∞<x<∞, andindeedany zintheentirefinitecomplexplane.
12.2 Recurrence Relations 749
FIGURE 12.7
12.1.9 TheChebyshevpolynomials(typeII) are generatedby(Eq. (13.93), Section13.3)
1
1−2xt+t2=∞summationdisplay
n=0Un(x)tn.
Using the techniques of Section 5.4 for transforming series, develop a series represen-
tationofUn(x).
ANS.Un(x)=[n/2]summationdisplay
k=0(−1)k(n−k)!
k!(n−2k)!(2x)n−2k.
12.2 R ECURRENCE RELATIONS AND SPECIAL PROPERTIES
Recurrence Relations
The Legendre polynomial generating function provides a convenient way of deriving the
recurrence relations4and some special properties. If our generating function (Eq. (12.4))
isdifferentiatedwithrespectto t, weobtain
∂g(t,x)
∂t=x−t
(1−2xt+t2)3/2=∞summationdisplay
n=0nPn(x)tn−1. (12.14)
BysubstitutingEq. (12.4)intothisandrearrangingterms,wehave
parenleftbig
1−2xt+t2parenrightbig∞summationdisplay
n=0nPn(x)tn−1+(t−x)∞summationdisplay
n=0Pn(x)tn=0.(12.15)
The left-hand side is a power series in t. Since this power series vanishes for all values of
t, the coefficient of each power of tis equal to zero; that is, our power series is unique
(Section 5.7). These coefficients are found by separating the individual summations and
4Wecanalso apply theexplicit series form Eq.(12.8) directly.
750 Chapter 12 Legendre Functions
usingdistinctivesummationindices:
∞summationdisplay
m=0mPm(x)tm−1−∞summationdisplay
n=02nxPn(x)tn+∞summationdisplay
s=0sPs(x)ts+1
+∞summationdisplay
s=0Ps(x)ts+1−∞summationdisplay
n=0xPn(x)tn=0. (12.16)
Now,letting m=n+1,s=n−1,wefind
(2n+1)xPn(x)=(n+1)Pn+1(x)+nPn−1(x), n=1,2,3,.... (12.17)
This is another three-term recurrence relation, similar to (but not identicalwith) the recur-
rence relation for Bessel functions. With this recurrence relation we may easily construct
the higher Legendre polynomials. If we take n=1 and insert the easily found values of
P0(x)andP1(x)(Exercise12.1.7orEq. (12.8)), weobtain
3xP1(x)=2P2(x)+P0(x), (12.18)
or
P2(x)=1
2parenleftbig
3x2−1parenrightbig
. (12.19)
This process may be continued indefinitely, the first few Legendre polynomials are listed
inTable12.1.
As cumbersome as it may appear at first, this technique is actually more efficient for
adigitalcomputerthanisdirectevaluationoftheseries(Eq.(12.8)).Forgreaterstability(to
avoid undue accumulation and magnification of round-off error), Eq. (12.17) is rewritten
as
Pn+1(x)=2xPn(x)−Pn−1(x)−1
n+1bracketleftbig
xPn(x)−Pn−1(x)bracketrightbig
. (12.17a)
Onestartswith P0(x)=1,P1(x)=x,andcomputesthe numerical valuesofallthe Pn(x)
for a given value of xup to the desired PN(x). The values of Pn(x),0≤n<N,a r e
availableas afringebenefit.
Table 12.1 LegendrePolynomials
P0(x)=1
P1(x)=x
P2(x)=1
2(3x2−1)
P3(x)=1
2(5x3−3x)
P4(x)=1
8(35x4−30x2+3)
P5(x)=1
8(63x5−70x3+15x)
P6(x)=1
16(231x6−315x4+105x2−5)
P7(x)=1
16(429x7−693x5+315x3−35x)
P8(x)=1
128(6435x8−12012x6+6930x4−1260x2+35)
12.2 Recurrence Relations 751
Differential Equations
More information about the behavior of the Legendre polynomials can be obtained if we
nowdifferentiateEq.(12.4) withrespectto x.Thisgives
∂g(t,x)
∂x=t
(1−2xt+t2)3/2=∞summationdisplay
n=0P′
n(x)tn, (12.20)
or
parenleftbig
1−2xt+t2parenrightbig∞summationdisplay
n=0P′
n(x)tn−t∞summationdisplay
n=0Pn(x)tn=0. (12.21)
Asbefore,thecoefficientofeachpowerof tisset equaltozeroandweobtain
P′
n+1(x)+P′
n−1(x)=2xP′
n(x)+Pn(x). (12.22)
AmoreusefulrelationmaybefoundbydifferentiatingEq.(12.17)withrespectto xand
multiplying by 2. To this we add (2n+1)times Eq. (12.22), canceling the P′
nterm. The
resultis
P′
n+1(x)−P′
n−1(x)=(2n+1)Pn(x). (12.23)
From Eqs. (12.22) and (12.23) numerous additional equations may be developed,5in-
cluding
P′
n+1(x)=(n+1)Pn(x)+xP′
n(x), (12.24)
P′
n−1(x)=−nPn(x)+xP′
n(x), (12.25)
parenleftbig
1−x2parenrightbig
P′
n(x)=nPn−1(x)−nxPn(x), (12.26)
parenleftbig
1−x2parenrightbig
P′
n(x)=(n+1)xPn(x)−(n+1)Pn+1(x). (12.27)
By differentiating Eq. (12.26) and using Eq. (12.25) to eliminate P′
n−1(x), we find that
Pn(x)satisfiesthelinearsecond-orderODE
parenleftbig
1−x2parenrightbig
P′′
n(x)−2xP′
n(x)+n(n+1)Pn(x)=0. (12.28)
The previous equations, Eqs. (12.22) to (12.27), are all first-order ODEs, but with poly-
nomials of two different indices. The price for having all indices alike is a second-order
5Usingthe equation number in parentheses to denote the left-hand side ofthe equation, wemay writethe derivatives as
2·d
dx(12.17)+(2n+1)·(12.22)⇒(12.23),
1
2braceleftbig
(12.22)+(12.23)bracerightbig
⇒(12.24),
1
2braceleftbig
(12.22)−(12.23)bracerightbig
⇒(12.25),
(12.24)n→n−1+x·(12.25)⇒(12.26),
d
dx(12.26)+n·(12.25)⇒(12.28).
752 Chapter 12 Legendre Functions
differentialequation.Equation(12.28)is Legendre’s ODE.Wenowseethatthepolynomi-
alsPn(x)generatedbythepowerseriesfor (1−2xt+t2)−1/2satisfyLegendre’sequation,
which,of course,iswhytheyarecalledLegendrepolynomials.
In Eq. (12.28) differentiation is with respect to x( x=cosθ). Frequently, we encounter
Legendre’sequationexpressedintermsofdifferentiationwithrespectto θ:
1
sinθd
dθparenleftbigg
sinθdPn(cosθ)
dθparenrightbigg
+n(n+1)Pn(cosθ)=0. (12.29)
Special Values
Our generating function provides still more information about the Legendre polynomials.
If weset x=1,Eq.(12.4) becomes
1
(1−2t+t2)1/2=1
1−t=∞summationdisplay
n=0tn, (12.30)
usingabinomialexpansionorthegeometricseries,Example5.1.1.ButEq.(12.4)for x=1
defines
1
(1−2t+t2)1/2=∞summationdisplay
n=0Pn(1)tn.
Comparingthetwoseries expansions(uniquenessof powerseries, Section5.7), wehave
Pn(1)=1. (12.31)
If weletx=−1 inEq. (12.4) anduse
1
(1+2t+t2)1/2=1
1+t,
thisshowsthat
Pn(−1)=(−1)n. (12.32)
For obtaining these results, we find that the generating function is more convenient than
theexplicitseries form,Eq. (12.8).
If wetake x=0 inEq. (12.4), usingthebinomialexpansion
parenleftbig
1+t2parenrightbig−1/2=1−1
2t2+3
8t4+···+(−1)n1·3···(2n−1)
2nn!t2n+···,(12.33)
wehave6
P2n(0)=(−1)n1·3···(2n−1)
2nn!=(−1)n(2n−1)!!
(2n)!!=(−1)n(2n)!
22n(n!)2(12.34)
P2n+1(0)=0,n=0,1,2.... (12.35)
TheseresultsalsofollowfromEq. (12.8)byinspection.
6Thedouble factorial notation is defined in Section8.1:
(2n)!!=2·4·6···(2n), ( 2n−1)!!=1·3·5···(2n−1), (−1)!!=1.
12.2 Recurrence Relations 753
Parity
SomeoftheseresultsarespecialcasesoftheparitypropertyoftheLegendrepolynomials.
We refer once more to Eqs. (12.4) and (12.8). If we replace xby−xandtby−t,t h e
generatingfunctionis unchanged.Hence
g(t,x)=g(−t,−x)=bracketleftbig
1−2(−t)(−x)+(−t)2bracketrightbig−1/2
=∞summationdisplay
n=0Pn(−x)(−t)n=∞summationdisplay
n=0Pn(x)tn. (12.36)
Comparingthesetwo series, wehave
Pn(−x)=(−1)nPn(x); (12.37)
thatis,thepolynomialfunctionsareoddoreven(withrespectto x=0,θ=π/2)according
to whether the index nis odd or even. This is the parity,7or reflection, property that plays
such an important role in quantum mechanics. For central forces the index nis a measure
oftheorbitalangularmomentum,thuslinkingparityandorbitalangularmomentum.
This parity property is confirmed by the series solution and for the special values tabu-
lated in Table 12.1. It might also be noted that Eq. (12.37) may be predicted by inspection
of Eq. (12.17), the recurrence relation. Specifically, if Pn−1(x)andxPn(x)are even, then
Pn+1(x)mustbeeven.
Upper and Lower Bounds for Pn(cosθ)
Finally,inadditiontotheseresults,ourgeneratingfunctionenablesustosetanupperlimit
on|Pn(cosθ)|.W eha v e
parenleftbig
1−2tcosθ+t2parenrightbig−1/2=parenleftbig
1−teiθparenrightbig−1/2parenleftbig
1−te−iθparenrightbig−1/2
=parenleftbig
1+1
2teiθ+3
8t2e2iθ+···parenrightbig
·parenleftbig
1+1
2te−iθ+3
8t2e−2iθ+···parenrightbig
,(12.38)
with all coefficients positive. Our Legendre polynomial, Pn(cosθ), still the coefficient of
tn, maynowbewrittenasasumof termsoftheform
1
2amparenleftbig
eimθ+e−imθparenrightbig
=amcosmθ (12.39a)
withallthe ampositiveandmandnbothevenor oddso that
Pn(cosθ)=nsummationdisplay
m=0o r1amcosmθ. (12.39b)
7In spherical polar coordinates the inversion of the point (r,θ,ϕ)through the origin is accomplished by the transformation
[r→r,θ→π−θ,andϕ→ϕ±π].Then,cos θ→cos(π−θ)=−cosθ,correspondingto x→−x(compareExercise2.5.8).
754 Chapter 12 Legendre Functions
This series, Eq. (12.39b), is clearly a maximum when θ=0 and cos mθ=1. But for x=
cosθ=1,Eq.(12.31) showsthat Pn(1)=1.Therefore
vextendsinglevextendsinglePn(cosθ)vextendsinglevextendsingle≤Pn(1)=1. (12.39c)
A fringe benefit of Eq. (12.39b) is that it shows that our Legendre polynomial is a linear
combination of cos mθ. This means that the Legendre polynomials form a complete set
for any functions that may be expanded by a Fourier cosine series (Section 14.1) over the
interval[0,π].
•In this section various useful properties of the Legendre polynomials are derived from
thegeneratingfunction,Eq.(12.4).
•The explicit series representation, Eq. (12.8), offers an alternate and sometimes supe-
rior approach.
Exercises
12.2.1 Giventheseries
α0+α2cos2θ+α4cos4θ+α6cos6θ=a0P0+a2P2+a4P4+a6P6,
express the coefficients αias a column vector αand the coefficients aias a column
vectoraanddeterminethematrices AandBsuchthat
Aα=aandBa=α.
Checkyourcomputationbyshowingthat AB=1(unitmatrix).Repeatfortheoddcase
α1cosθ+α3cos3θ+α5cos5θ+α7cos7θ=a1P1+a3P3+a5P5+a7P7.
Note.Pn(cosθ)and cosnθare tabulated in terms of each other in AMS-55 (see Addi-
tionalReadingsof Chapter8forthecompletereference).
12.2.2 By differentiating the generating function g(t,x)with respect to t, multiplying by 2 t,
andthenadding g(t,x), showthat
1−t2
(1−2tx+t2)3/2=∞summationdisplay
n=0(2n+1)Pn(x)tn.
This result is useful in calculating the charge induced on a grounded metal sphere by a
pointcharge q.
12.2.3 (a) DeriveEq. (12.27),
parenleftbig
1−x2parenrightbig
P′
n(x)=(n+1)xPn(x)−(n+1)Pn+1(x).
(b) Write out the relation of Eq. (12.27) to preceding equations in symbolic form
analogoustothesymbolicforms forEqs. (12.23) to(12.26).
12.2 Recurrence Relations 755
12.2.4 A point electric octupole may be constructed by placing a point electric quadrupole
(pole strength p(2)in thez-direction) at z=aand an equal but opposite point elec-
tric quadrupole at z=0 and then letting a→0, subject to p(2)a=constant. Find the
electrostatic potential corresponding to a point electric octupole. Show from the con-
structionofthepointelectricoctupolethatthecorrespondingpotentialmaybeobtained
bydifferentiatingthepointquadrupolepotential.
12.2.5 Operatingin sphericalpolarcoordinates ,showthat
∂
∂zbracketleftbiggPn(cosθ)
rn+1bracketrightbigg
=−(n+1)Pn+1(cosθ)
rn+2.
This is the key step in the mathematical argument that the derivative of one multipole
leadstothenexthighermultipole.
Hint.CompareExercise2.5.12.
12.2.6 From
PL(cosθ)=1
L!∂L
∂tLparenleftbig
1−2tcosθ+t2parenrightbig−1/2vextendsinglevextendsingle
t=0
showthat
PL(1)=1,P L(−1)=(−1)L.
12.2.7 Provethat
P′
n(1)=d
dxPn(x)vextendsinglevextendsingle
x=1=1
2n(n+1).
12.2.8 Show that Pn(cosθ)=(−1)nPn(−cosθ)by use of the recurrence relation relating
Pn,Pn+1, andPn−1andyourknowledgeof P0andP1.
12.2.9 From Eq. (12.38) write out the coefficient of t2in terms of cos nθ,n≤2. This coeffi-
cientisP2(cosθ).
12.2.10 Write a program that will generate the coefficients asin the polynomial form of the
Legendrepolynomial
Pn(x)=nsummationdisplay
s=0asxs.
12.2.11 (a) Calculate P10(x)overtherange [0,1]andplotyourresults.
(b) Calculateprecise(atleasttofivedecimalplaces)valuesofthefivepositiverootsof
P10(x). Compare your values with the values listed in AMS-55, Table 25.4. (For
thecompletereference,seeAdditionalReadingsofChapter8.)
12.2.12 (a) Calculatethe largestrootofPn(x)forn=2(1)50.
(b) Develop an approximation for the largest root from the hypergeometric represen-
tation of Pn(x)(Section 13.4) and compare your values from part (a) with your
hypergeometric approximation. Compare also with the values listed in AMS-55,
Table25.4.(For thecompletereference,seeAdditionalReadingsof Chapter8.)
756 Chapter 12 Legendre Functions
12.2.13 (a) From Exercise 12.2.1 and AMS-55, Table 22.9, develop the 6 ×6m a t r i xBthat
will transform a series of even-order Legendre polynomials through P10(x)into a
powerseriessummationtext5
n=0α2nx2n.
(b) Calculate AasB−1.Checktheelementsof AagainstthevalueslistedinAMS-55,
Table22.9.(For thecompletereference,seeAdditionalReadingsof Chapter8.)
(c) By using matrix multiplication, transform some even power seriessummationtext5
n=0α2nx2n
intoaLegendreseries.
12.2.14 Write a subroutinethatwill transform a finitepower seriessummationtextN
n=0anxnintoa Legendre
seriessummationtextN
n=0bnPn(x).Usetherecurrencerelation,Eq.(12.17),andfollowthetechnique
outlinedinSection13.3for aChebyshevseries.
12.3 O RTHOGONALITY
Legendre’sODE(12.28)maybewrittenintheform
d
dxbracketleftbigparenleftbig
1−x2parenrightbig
P′
n(x)bracketrightbig
+n(n+1)Pn(x)=0, (12.40)
showing clearly that it is self-adjoint. Subject to satisfying certain boundary condi-
tions, then, it is known that the solutions Pn(x)will be orthogonal. Upon comparing
Eq. (12.40) with Eqs. (10.6) and (10.8) we see that the weight function w(x)=1,L=
(d/dx)(1−x2)(d/dx),p(x)=1−x2, and the eigenvalue λ=n(n+1). The integration
limitson xare±1,where p(±1)=0.Thenfor m/negationslash=n,Eq. (10.34)becomes
integraldisplay1
−1Pn(x)Pm(x)dx=0,8(12.41)
integraldisplayπ
0Pn(cosθ)Pm(cosθ)sinθdθ=0, (12.42)
showing that Pn(x)andPm(x)are orthogonal for the interval [−1,1]. This orthogonality
mayalsobedemonstratedbyusingRodrigues’definitionof Pn(x)(compareSection12.4,
Exercise12.4.2).
Weshallneedtoevaluatetheintegral(Eq.(12.41))when n=m.Certainlyitisnolonger
zero.Fromourgeneratingfunction,
parenleftbig
1−2tx+t2parenrightbig−1=bracketleftbigg∞summationdisplay
n=0Pn(x)tnbracketrightbigg2
. (12.43)
Integratingfrom x=−1t ox=+1,wehave
integraldisplay1
−1dx
1−2tx+t2=∞summationdisplay
n=0t2nintegraldisplay1
−1bracketleftbig
Pn(x)bracketrightbig2dx; (12.44)
8In Section 10.4 suchintegrals are interpreted as inner products in a linearvector(function) space.Alternate notations are
integraldisplay1
−1bracketleftbig
Pn(x)bracketrightbig∗Pm(x)dx≡/angbracketleftPn|Pm/angbracketright≡(Pn,Pm).
The/angbracketleft/angbracketrightform, popularized by Dirac, is common in the physics literature. The () form is more common in the mathematics
literature.
12.3 Orthogonality 757
the cross terms in the series vanish by means of Eq. (12.42). Using y=1−2tx+t2,
dy=−2tdx,weobtain
integraldisplay1
−1dx
1−2tx+t2=1
2tintegraldisplay(1+t)2
(1−t)2dy
y=1
tlnparenleftbigg1+t
1−tparenrightbigg
. (12.45)
Expandingthisina powerseries (Exercise5.4.1)givesus
1
tlnparenleftbigg1+t
1−tparenrightbigg
=2∞summationdisplay
n=0t2n
2n+1. (12.46)
Comparingpower-seriescoefficientsofEqs. (12.44) and(12.46), wemusthave
integraldisplay1
−1bracketleftbig
Pn(x)bracketrightbig2dx=2
2n+1. (12.47)
CombiningEq. (12.42)withEq.(12.47) wehavetheorthonormalitycondition
integraldisplay1
−1Pm(x)Pn(x)dx=2δmn
2n+1. (12.48)
We shall return to this result in Section 12.6 when we construct the orthonormal spherical
harmonics.
Expansion of Functions, Legendre Series
Inadditiontoorthogonality,theSturm–LiouvilletheoryimpliesthattheLegendrepolyno-
mialsform acompleteset. Letusassume, then,thattheseries
∞summationdisplay
n=0anPn(x)=f(x) (12.49)
converges in the mean (Section 10.4) in the interval [−1,1]. This demands that f(x)and
f′(x)be at least sectionally continuous in this interval. The coefficients anare found by
multiplying the series by Pm(x)and integrating term by term. Using the orthogonality
propertyexpressedinEqs. (12.42)and(12.48),weobtain
2
2m+1am=integraldisplay1
−1f(x)Pm(x)dx. (12.50)
Wereplacethevariableof integration xbytandtheindex mbyn.Then,substitutinginto
Eq.(12.49), wehave
f(x)=∞summationdisplay
n=02n+1
2parenleftbiggintegraldisplay1
−1f(t)Pn(t)dtparenrightbigg
Pn(x). (12.51)
This expansion in a series of Legendre polynomials is usually referred to as a Legendre
series.9Its properties are quite similar to the more familiar Fourier series (Chapter 14). In
9Notethat Eq. (12.50) gives amasadefiniteintegral, that is, anumber for a given f(x).
758 Chapter 12 Legendre Functions
particular, we can use the orthogonality property (Eq. (12.48)) to show that the series is
unique.
On a more abstract (and more powerful) level, Eq. (12.51) gives the representation of
f(x)inthevectorspaceofLegendrepolynomials(a Hilbertspace,Section10.4).
From the viewpoint of integral transforms (Chapter 15), Eq. (12.50) may be considered
afiniteLegendretransformof f(x).Equation(12.51)isthentheinversetransform.Itmay
also be interpreted in terms of the projection operators of quantum theory. We may take
Pmin
[Pmf](x)≡Pm(x)2m+1
2integraldisplay1
−1Pm(t)bracketleftbig
f(t)bracketrightbig
dt
asan(integral)operator,readytooperateon f(t).(Thef(t)wouldgointhesquarebracket
asafactor intheintegrand.)Then,fromEq. (12.50),
[Pmf](x)=amPm(x).10
Theoperator Pmprojectsoutthe mthcomponentof thefunction f.
Equation (12.3), which leads directly to the generating function definition of Legendre
polynomials, is a Legendre expansion of 1 /r1. This Legendre expansion of 1 /r1or 1/r12
appears in several exercises of Section 12.8. Going beyond a Coulomb field, the 1 /r12is
oftenreplacedbyapotential V(|r1−r2|),andthesolutionoftheproblemisagaineffected
byaLegendreexpansion.
The Legendre series, Eq. (12.49), has been treated as a knownfunctionf(x)that we
arbitrarily chose to expandin a series of Legendrepolynomials.Sometimesthe origin and
nature of the Legendre series are different. In the next examples we consider unknown
functions we know can be represented by a Legendre series because of the differential
equation the unknown functions satisfy. As before, the problem is to determine the un-
known coefficients in the series expansion. Here, however, the coefficients are not found
by Eq. (12.50). Rather, they are determined by demanding that the Legendre series match
aknownsolutionataboundary.These areboundaryvalueproblems.
Example 12.3.1 EARTH ’SGRAVITATIONAL FIELD
AnexampleofaLegendreseriesisprovidedbythedescriptionoftheEarth’sgravitational
potential U(for exteriorpoints),neglectingazimuthaleffects. With
R=equatorialradius =6378.1±0.1km
GM
R=62.494±0.001km2/s2,
wewrite
U(r,θ)=GM
RbracketleftbiggR
r−∞summationdisplay
n=2anparenleftbiggR
rparenrightbiggn+1
Pn(cosθ)bracketrightbigg
, (12.52)
10Thedependent variables arearbitrary. Here xcamefrom the xinPm.
12.3 Orthogonality 759
aLegendreseries. Artificialsatellitemotionshaveshownthat
a2=(1,082,635±11)×10−9,
a3=(−2,531±7)×10−9,
a4=(−1,600±12)×10−9.
This is the famous pear-shaped deformation of the Earth. Other coefficients have been
computed through n=20. Note that P1is omitted because the origin from which ris
measuredis theEarth’s centerof mass( P1wouldrepresentadisplacement).
More recent satellite data permit a determination of the longitudinal dependence of the
Earth’s gravitational field. Such dependence may be described by a Laplace series (Sec-
tion12.6). /squaresolid
Example 12.3.2 SPHERE IN A UNIFORM FIELD
Another illustration of the use of Legendre polynomials is provided by the problem of
a neutral conducting sphere (radius r0) placed in a (previously) uniform electric field
(Fig. 12.8). The problem is to find the new, perturbed, electrostatic potential. If we call
theelectrostaticpotential11V, itsatisfies
∇2V=0, (12.53)
Laplace’sequation.Weselectsphericalpolarcoordinatesbecauseofthesphericalshapeof
the conductor. (This will simplify the application of the boundary condition at the surface
oftheconductor.)SeparatingvariablesandglancingatTable9.2,wecanwritetheunknown
potential V(r,θ)intheregionoutsidethesphereasalinearcombinationofsolutions:
V(r,θ)=∞summationdisplay
n=0anrnPn(cosθ)+∞summationdisplay
n=0bnPn(cosθ)
rn+1. (12.54)
FIGURE 12.8Conductingspherein
auniformfield.
11It should beemphasizedthatthis isnot apresentationofaLegendre-series expansion ofaknown V(cosθ).Hereweareback
toboundary value problems of PDEs.
760 Chapter 12 Legendre Functions
Noϕ-dependence appears because of the axial symmetry of our problem. (The center of
theconductingsphereistakenastheoriginandthe z-axisisorientedparalleltotheoriginal
uniformfield.)
It might be noted here that nis an integer, because only for integral nis theθdepen-
dence well behaved at cos θ=±1. For nonintegral nthe solutions of Legendre’s equation
divergeattheendsoftheinterval [−1,1],thepoles θ=0,πofthesphere(compareExam-
ple5.2.4andExercises5.2.15and9.5.5).Itisforthissamereasonthatthesecondsolution
ofLegendre’sequation, Qn,is alsoexcluded.
Now we turn to our (Dirichlet) boundary conditions to determine the unknown anand
bnof our series solution, Eq. (12.54). If the original, unperturbed electrostatic field is E0,
werequire, asoneboundarycondition,
V(r→∞)=−E0z=−E0rcosθ=−E0rP1(cosθ). (12.55)
SinceourLegendreseriesisunique,wemayequatecoefficientsof Pn(cosθ)inEq.(12.54)
(r→∞)andEq.(12.55) toobtain
an=0,n>1 and n=0,a 1=−E0. (12.56)
Ifan/negationslash=0f o rn>1, these terms would dominate at large rand the boundary condition
(Eq. (12.55))couldnotbesatisfied.
As a second boundary condition, we may choose the conducting sphere and the plane
θ=π/2 tobeatzeropotential,whichmeansthatEq. (12.54)nowbecomes
V(r=r0)=b0
r0+parenleftbiggb1
r2
0−E0r0parenrightbigg
P1(cosθ)+∞summationdisplay
n=2bnPn(cosθ)
rn+1
0=0.(12.57)
Inorderthatthismayholdforallvaluesof θ,eachcoefficientof Pn(cosθ)mustvanish.12
Hence
b0=0,13bn=0,n≥2, (12.58)
whereas
b1=E0r3
0. (12.59)
Theelectrostaticpotential(outsidethesphere) isthen
V=−E0rP1(cosθ)+E0r3
0
r2P1(cosθ)=−E0rP1(cosθ)parenleftbigg
1−r3
0
r3parenrightbigg
.(12.60)
InSection1.16itwasshownthatasolutionofLaplace’sequationthatsatisfiedthebound-
aryconditionsovertheentireboundarywasunique.Theelectrostaticpotential V,asgi ven
byEq.(12.60),isasolutionofLaplace’sequation.Itsatisfiesourboundaryconditionsand
thereforeisthesolutionofLaplace’sequationfor thisproblem.
12Again,thisisequivalenttosayingthataseriesexpansioninLegendrepolynomials(oranycompleteorthogonalset)isunique.
13The coefficient of P0isb0/r0.W es e tb0=0 because there is no net charge on the sphere. If there is a net charge q,t h e n
b0/negationslash=0.
12.3 Orthogonality 761
It may further be shown (Exercise 12.3.13) that there is an induced surface charge den-
sity
σ=−ε0∂V
∂rvextendsinglevextendsinglevextendsinglevextendsingle
r=r0=3ε0E0cosθ (12.61)
onthesurface ofthesphereandaninducedelectricdipolemoment(Exercise12.3.13)
P=4πr3
0ε0E0. (12.62)
/squaresolid
Example 12.3.3 ELECTROSTATIC POTENTIAL OF A RING OF CHARGE
As a further example, consider the electrostatic potential produced by a conducting ring
carrying a total electric charge q(Fig. 12.9). From electrostatics (and Section 1.14) the
potential ψsatisfiesLaplace’sequation.Separatingvariablesinsphericalpolarcoordinates
(compareTable9.2), weobtain
ψ(r,θ)=∞summationdisplay
n=0cnan
rn+1Pn(cosθ), r>a. (12.63a)
Hereais the radius of the ring that is assumed to be in the θ=π/2 plane. There is no
ϕ(azimuthal) dependence because of the cylindrical symmetry of the system. The terms
with positive exponent in the radial dependence have been rejected because the potential
musthaveanasymptoticbehavior,
ψ∼q
4πε0·1
r,r≫a. (12.63b)
The problem is to determine the coefficients cnin Eq. (12.63a). This may be done by
evaluating ψ(r,θ)atθ=0,r=z, and comparing with an independent calculation of the
FIGURE 12.9Charged,
conductingring.
762 Chapter 12 Legendre Functions
potential from Coulomb’s law. In effect, we are using a boundary condition along the z-
axis.FromCoulomb’slaw(withallchargeequidistant),
ψ(r,θ)=q
4πε0·1
(z2+a2)1/2,braceleftbiggθ=0
r=z,
=q
4πε0z∞summationdisplay
s=0(−1)s(2s)!
22s(s!)2parenleftbigga
zparenrightbigg2s
,z>a. (12.63c)
ThelaststepusestheresultofExercise8.1.15.Now,Eq.(12.63a)evaluatedat θ=0,r=z
(withPn(1)=1),yields
ψ(r,θ)=∞summationdisplay
n=0cnan
zn+1,r=z. (12.63d)
ComparingEqs. (12.63c)and(12.63d), weget cn=0f o rnodd.Setting n=2s,weha v e
c2s=q
4πε0(−1)s(2s)!
22s(s!)2, (12.63e)
andourelectrostaticpotential ψ(r,θ)is givenby
ψ(r,θ)=q
4πε0r∞summationdisplay
s=0(−1)s(2s)!
22s(s!)2parenleftbigga
rparenrightbigg2s
P2s(cosθ), r>a. (12.63f)
Themagneticanalogofthis problemappearsinExample12.5.3. /squaresolid
Exercises
12.3.1 YouhaveconstructedasetoforthogonalfunctionsbytheGram–Schmidtprocess(Sec-
tion 10.3), taking un(x)=xn,n=0,1,2,...,in increasing order with w(x)=1 and
an interval−1≤x≤1. Prove that the nth such function constructed is proportional to
Pn(x).
Hint.Usemathematicalinduction.
12.3.2 Expand the Dirac delta function in a series of Legendre polynomials using the interval
−1≤x≤1.
12.3.3 VerifytheDiracdeltafunctionexpansions
δ(1−x)=∞summationdisplay
n=02n+1
2Pn(x)
δ(1+x)=∞summationdisplay
n=0(−1)n2n+1
2Pn(x).
These expressions appear in a resolution of the Rayleigh plane-wave expansion (Exer-
cise12.4.7)intoincomingandoutgoingsphericalwaves.
Note. Assume that the entireDirac delta function is covered when integrating over
[−1,1].
12.3 Orthogonality 763
12.3.4 Neutrons (mass 1) are being scattered by a nucleus of mass A( A>1). In the center-
of-masssystemthescatteringisisotropic.Then,inthelaboratorysystemtheaverageof
thecosineof theangleof deflectionof theneutronis
/angbracketleftcosψ/angbracketright=1
2integraldisplayπ
0Acosθ+1
(A2+2Acosθ+1)1/2sinθdθ.
Show,byexpansionofthedenominator,that /angbracketleftcosψ/angbracketright=2/3A.
12.3.5 A particular function f(x)defined over the interval [−1,1]is expanded in a Legendre
seriesoverthissameinterval.Showthattheexpansionisunique.
12.3.6 Afunction f(x)is expandedinaLegendreseries f(x)=summationtext∞
n=0anPn(x). Showthat
integraldisplay1
−1bracketleftbig
f(x)bracketrightbig2dx=∞summationdisplay
n=02a2
n
2n+1.
ThisistheLegendreformoftheFourierseriesParsevalidentity,Exercise14.4.2.Italso
illustratesBessel’sinequality,Eq. (10.72),becominganequalityfor acompleteset.
12.3.7 Derivetherecurrencerelation
parenleftbig
1−x2parenrightbig
P′
n(x)=nPn−1(x)−nxPn(x)
fromtheLegendrepolynomialgeneratingfunction.
12.3.8 Evaluateintegraltext1
0Pn(x)dx.
ANS.n=2s;1 fors=0,0f ors>0,
n=2s+1;P2s(0)/(2s+2)=(−1)s(2s−1)!!/1(2s+2)!!
Hint. Use a recurrence relation to replace Pn(x)by derivatives and then integrate by
inspection.Alternatively,youcanintegratethegeneratingfunction.
12.3.9 (a) For
f(x)=braceleftbigg+1,0<x<1
−1,−1<x<0,
showthat
integraldisplay1
−1bracketleftbig
f(x)bracketrightbig2dx=2∞summationdisplay
n=0(4n+3)bracketleftbigg(2n−1)!!
(2n+2)!!bracketrightbigg2
.
(b) Bytestingtheseries, provethattheseries is convergent.
12.3.10 Provethat
integraldisplay1
−1xparenleftbig
1−x2parenrightbig
P′
nP′
mdx=0,unlessm=n±1,
=2n(n2−1)
4n2−1δm,n−1,ifm<n.
=2n(n+2)(n+1)
(2n+1)(2n+3)δm,n+1,ifm>n.
764 Chapter 12 Legendre Functions
12.3.11 Theamplitudeof ascatteredwaveis givenby
f(θ)=1
k∞summationdisplay
l=0(2l+1)exp[iδl]sinδlPl(cosθ).
Hereθis the angle of scattering, lis the angular momentum eigenvalue, ¯hkis the
incident momentum, and δlis the phase shift produced by the central potential that is
doingthescattering.Thetotalcross sectionis σtot=integraltext
|f(θ)|2d/Omega1.Showthat
σtot=4π
k2∞summationdisplay
l=0(2l+1)sin2δl.
12.3.12 The coincidence counting rate, W(θ), in a gamma–gamma angular correlation experi-
menthastheform
W(θ)=∞summationdisplay
n=0a2nP2n(cosθ).
Show that data in the range π/2≤θ≤πcan, in principle, define the function W(θ)
(and permit a determination of the coefficients a2n). This means that although data in
therange 0≤θ<π /2 maybeusefulasacheck,theyarenotessential.
12.3.13 A conducting sphere of radius r0is placed in an initially uniform electric field, E0.
Showthefollowing:
(a) Theinducedsurface chargedensityis
σ=3ε0E0cosθ.
(b) Theinducedelectricdipolemomentis
P=4πr3
0ε0E0.
The induced electric dipole moment can be calculated either from the surface
charge [part (a)] or by noting that the final electric field Eis the result of su-
perimposinga dipolefieldontheoriginaluniformfield.
12.3.14 A charge qis displaced a distance aalong the z-axis from the center of a spherical
cavityofradius R.
(a) Showthattheelectricfieldaveragedoverthevolume a≤r≤Riszero.
(b) Showthattheelectricfieldaveragedoverthevolume 0 ≤r≤ais
E=ˆzEz=−ˆzq
4πε0a2(SI units)=−ˆznqa
3ε0,
wherenisthenumberofsuchdisplacedchargesperunitvolume.Thisisabasiccalcu-
lationinthepolarizationofa dielectric.
Hint.E=−∇ϕ.
12.3 Orthogonality 765
FIGURE 12.10Charged,
conductingdisk.
12.3.15 Determine the electrostatic potential (Legendre expansion) of a circular ring of electric
chargefor r<a.
12.3.16 Calculate the electric field produced by the charged conductingring of Example 12.3.3
for
(a)r>a,(b)r<a.
12.3.17 As an extension of Example 12.3.3, find the potential ψ(r,θ)produced by a charged
conducting disk, Fig. 12.10, for r>a, the radius of the disk. The charge density σ(on
eachsideof thedisk) is
σ(ρ)=q
4πa(a2−ρ2)1/2,ρ2=x2+y2.
Hint. The definite integral you get can be evaluated as a beta function, Section 8.4. For
moredetailsseeSection5.03ofSmytheinAdditionalReadings.
ANS.ψ(r,θ)=q
4πε0r∞summationdisplay
l=0(−1)l1
2l+1parenleftbigga
rparenrightbigg2l
P2l(cosθ).
12.3.18 From the result of Exercise 12.3.17 calculate the potential of the disk. Since you are
violatingthecondition r>a, justifyyourcalculation.
Hint.Youmayrunintotheseries giveninExercise5.2.9.
12.3.19 The hemisphere defined by r=a,0≤θ<π /2, has an electrostatic potential +V0.
The hemisphere r=a,π/2<θ≤πhas an electrostatic potential −V0. Show that the
potentialatinteriorpointsis
V=V0∞summationdisplay
n=04n+3
2n+2parenleftbiggr
aparenrightbigg2n+1
P2n(0)P2n+1(cosθ)
=V0∞summationdisplay
n=0(−1)n(4n+3)(2n−1)!!
(2n+2)!!parenleftbiggr
aparenrightbigg2n+1
P2n+1(cosθ).
Hint.YouneedExercise12.3.8.
12.3.20 Aconductingsphereofradius aisdividedintotwoelectricallyseparatehemispheresby
a thin insulating barrier at its equator. The top hemisphere is maintained at a potential
V0,thebottomhemisphereat −V0.
766 Chapter 12 Legendre Functions
(a) Showthattheelectrostaticpotential exteriortothetwohemispheresis
V(r,θ)=V0∞summationdisplay
s=0(−1)s(4s+3)(2s−1)!!
(2s+2)!!parenleftbigga
rparenrightbigg2s+2
P2s+1(cosθ).
(b) Calculatetheelectricchargedensity σontheoutsidesurface.Notethatyourseries
diverges at cos θ=±1, as you expect from the infinite capacitance of this system
(zerothicknessfor theinsulatingbarrier).
ANS.σ=ε0En=−ε0∂V
∂rvextendsinglevextendsinglevextendsinglevextendsingle
r=a
=ε0V0∞summationdisplay
s=0(−1)s(4s+3)(2s−1)!!
(2s)!!P2s+1(cosθ).
12.3.21 In the notation of Section 10.4, ϕs(x)=√(2s+1)/2Ps(x), a Legendre polynomial is
renormalized to unity. Explain how |ϕs/angbracketright/angbracketleftϕs|acts as a projection operator. In particular,
showthatif|f/angbracketright=summationtext
na′
n|ϕn/angbracketright,then
|ϕs/angbracketright/angbracketleftϕs|f/angbracketright=a′
s|ϕs/angbracketright.
12.3.22 Expandx8asaLegendreseries.DeterminetheLegendrecoefficientsfromEq.(12.50),
am=2m+1
2integraldisplay1
−1x8Pm(x)dx.
CheckyourvaluesagainstAMS-55,Table22.9.(Forthecompletereference,seeAddi-
tional Readingsin Chapter 8). This illustrates the expansionof a simple function f(x).
Actually if f(x)is expressed as a power series, the technique of Exercise 12.2.14 is
bothfaster andmoreaccurate.
Hint.Gaussianquadraturecanbeusedtoevaluatetheintegral.
12.3.23 Calculate and tabulate the electrostatic potential created by a ring of charge, Exam-
ple12.3.3,for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry termsthrough P22(cosθ).
Note. The convergence of your series will be slow for r/a=1.5. Truncating the series
atP22limitsyoutoaboutfour-significant-figureaccuracy.
Checkvalue. Forr/a=2.5 andθ=60◦,ψ=0.40272(q/4πε0r).
12.3.24 Calculate and tabulate the electrostatic potential created by a charged disk, Ex-
ercise 12.3.17, for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry terms through
P22(cosθ).
Checkvalue. Forr/a=2.0 andθ=15◦,ψ=0.46638(q/4πε0r).
12.3.25 Calculate the first five (nonvanishing) coefficients in the Legendre series expansion of
f(x)=1−|x|using Eq. (12.51)—numerical integration. Actually these coefficients
can be obtained in closed form. Compare your coefficients with those obtained from
Exercise13.3.28.
ANS.a0=0.5000,a2=−0.6250,a4=0.1875,a6=−0.1016,a8=0.0664.
12.4 Alternate Definitions 767
12.3.26 Calculate and tabulate the exterior electrostatic potential created by the two charged
hemispheres of Exercise 12.3.20, for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry
termsthrough P23(cosθ).
Checkvalue. Forr/a=2.0 andθ=45◦,V=0.27066V0.
12.3.27 (a) Given f(x)=2.0,|x|<0.5;f(x)=0,0.5<|x|<1.0, expand f(x)in a Legen-
dreseriesandcalculatethecoefficients anthrougha80(analytically).
(b) Evaluatesummationtext80
n=0anPn(x)forx=0.400(0.005)0.600.Plotyourresults.
Note.ThisillustratestheGibbsphenomenonofSection14.5andthedangeroftryingto
calculatewithaseriesexpansioninthevicinityof adiscontinuity.
12.4 A LTERNATE DEFINITIONS OF LEGENDRE POLYNOMIALS
Rodrigues’ Formula
The series form of the Legendre polynomials (Eq. (12.8)) of Section 12.1 may be trans-
formedasfollows.FromEq. (12.8),
Pn(x)=[n/2]summationdisplay
r=0(−1)r(2n−2r)!
2nr!(n−2r)!(n−r)!xn−2r. (12.64)
Fornaninteger,
Pn(x)=[n/2]summationdisplay
r=0(−1)r1
2nr!(n−r)!parenleftbiggd
dxparenrightbiggn
x2n−2r
=1
2nn!parenleftbiggd
dxparenrightbiggnnsummationdisplay
r=0(−1)rn!
r!(n−r)!x2n−2r. (12.64a)
Note the extension of the upper limit. The reader is asked to show in Exercise 12.4.1 that
the additional terms [n/2]+1t onin the summation contribute nothing. However, the
effectoftheseextratermsistopermitthereplacementofthenewsummationby (x2−1)n
(binomialtheoremonceagain)toobtain
Pn(x)=1
2nn!parenleftbiggd
dxparenrightbiggnparenleftbig
x2−1parenrightbign. (12.65)
This is Rodrigues’ formula. It is useful in proving many of the properties of the Legendre
polynomials, such as orthogonality. A related application is seen in Exercise 12.4.3. The
Rodrigues definition is extended in Section 12.5 to define the associated Legendre func-
tions.InSection12.7itis usedtoidentifytheorbitalangularmomentumeigenfunctions.
768 Chapter 12 Legendre Functions
Schlaefli Integral
Rodrigues’ formula provides a means of developing an integral representation of Pn(z).
UsingCauchy’sintegralformula(Section6.4)
f(z)=1
2πicontintegraldisplayf(t)
t−zdt (12.66)
with
f(z)=parenleftbig
z2−1parenrightbign, (12.67)
wehave
parenleftbig
z2−1parenrightbign=1
2πicontintegraldisplay(t2−1)n
t−zdt. (12.68)
Differentiating ntimeswithrespectto zandmultiplyingby 1 /2nn!gives
Pn(z)=1
2nn!dn
dznparenleftbig
z2−1parenrightbign=2−n
2πicontintegraldisplay(t2−1)n
(t−z)n+1dt, (12.69)
withthecontourenclosingthepoint t=z.
This is the Schlaefli integral. Margenau and Murphy14use this to derive the recurrence
relationsweobtainedfromthegeneratingfunction.
The Schlaefli integral may readily be shown to satisfy Legendre’s equation by differen-
tiationanddirectsubstitution(Fig.12.11). Weobtain
parenleftbig
1−z2parenrightbigd2Pn
dz2−2zdPn
dz+n(n+1)Pn=n+1
2n2πicontintegraldisplayd
dtbracketleftbigg(t2−1)n+1
(t−z)n+2bracketrightbigg
dt.(12.70)
Forintegral nourfunction (t2−1)n+1/(t−z)n+2issingle-valued,andtheintegralaround
the closed path vanishes. The Schlaefli integral may also be used to define Pν(z)for non-
integralνintegrating around the points t=z,t=1, but not crossing the cut line −1t o
−∞. We could equally well encircle the points t=zandt=−1, but this would lead to
FIGURE 12.11Schlaefliintegralcontour.
14H. Margenau and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed., Princeton, NJ: Van Nostrand (1956),
Section3.5.
12.4 Alternate Definitions 769
nothing new. A contour about t=+1 andt=−1 will lead to a second solution, Qν(z),
Section12.10.
Exercises
12.4.1 Showthat eachterminthesummation
nsummationdisplay
r=[n/2]+1parenleftbiggd
dxparenrightbiggn(−1)rn!
r!(n−r)!x2n−2r
vanishes( randnintegral).
12.4.2 UsingRodrigues’ formula,showthatthe Pn(x)areorthogonalandthat
integraldisplay1
−1bracketleftbig
Pn(x)bracketrightbig2dx=2
2n+1.
Hint.UseRodrigues’formulaandintegratebyparts.
12.4.3 Showthatintegraltext1
−1xmPn(x)dx=0 whenm<n.
Hint.UseRodrigues’formulaor expand xminLegendrepolynomials.
12.4.4 Showthat
integraldisplay1
−1xnPn(x)dx=2n+1n!n!
(2n+1)!.
Note.YouareexpectedtouseRodrigues’formulaandintegratebyparts,butalsoseeif
youcangettheresult fromEq. (12.8)byinspection.
12.4.5 Showthat
integraldisplay1
−1x2rP2n(x)dx=22n+1(2r)!(r+n!)
(2r+2n+1)!(r−n)!,r≥n.
12.4.6 As a generalization of Exercises 12.4.4 and 12.4.5, show that the Legendre expansions
ofxsare
(a)x2r=rsummationdisplay
n=022n(4n+1)(2r)!(r+n)!
(2r+2n+1)!(r−n)!P2n(x), s=2r,
(b)x2r+1=rsummationdisplay
n=022n+1(4n+3)(2r+1)!(r+n+1)!
(2r+2n+3)!(r−n)!P2n+1(x), s=2r+1.
12.4.7 AplanewavemaybeexpandedinaseriesofsphericalwavesbytheRayleighequation,
eikrcosγ=∞summationdisplay
n=0anjn(kr)Pn(cosγ).
Showthat an=in(2n+1).
770 Chapter 12 Legendre Functions
Hint.
1. Use theorthogonalityof the Pntosolvefor anjn(kr).
2. Differentiate ntimes with respect to (kr)and setr=0 to eliminate the r-
dependence.
3. EvaluatetheremainingintegralbyExercise12.4.4.
Note.Thisproblemmayalsobetreatedbynotingthatbothsidesoftheequationsatisfy
the Heemholtz equation. The equality can be established by showing that the solutions
have the same behavior at the origin and also behave alike at large distances. A “by
inspection”typeofsolutionis developedinSection9.7usingGreen’sfunctions.
12.4.8 VerifytheRayleighequationofExercise12.4.7bystartingwiththefollowingsteps:
(a) Differentiatewithrespectto (kr)toestablish
summationdisplay
nanj′
n(kr)Pn(cosγ)=isummationdisplay
nanjn(kr)cosγPn(cosγ).
(b) Use a recurrence relation to replace cos γPn(cosγ)by a linear combination of
Pn−1andPn+1.
(c) Usearecurrencerelationtoreplace j′
nbyalinearcombinationof jn−1andjn+1.
12.4.9 From Exercise12.4.7showthat
jn(kr)=1
2inintegraldisplay1
−1eikrµPn(µ)dµ.
This means that (apart from a constant factor) the spherical Bessel function jn(kr)is
theFouriertransformof theLegendrepolynomial Pn(µ).
12.4.10 TheLegendrepolynomialsandthesphericalBesselfunctionsarerelatedby
jn(z)=1
2(−i)nintegraldisplayπ
0eizcosθPn(cosθ)sinθdθ, n=0,1,2,....
Verifythis relationbytransformingtheright-handsideinto
zn
2n+1n!integraldisplayπ
0cos(zcosθ)sin2n+1θdθ
andusingExercise11.7.8.
12.4.11 Bydirectevaluationof theSchlaefliintegralshowthat Pn(1)=1.
12.4.12 Explain why the contour of the Schlaefli integral, Eq. (12.69), is chosen to enclose the
pointst=zandt=1 whenn→ν, notaninteger.
12.4.13 Innumericalwork(forexample,theGauss–Legendrequadrature)itisusefultoestablish
thatPn(x)hasnrealzerosintheinteriorof [−1,1].Showthatthisis so.
Hint. Rolle’s theorem shows that the first derivative of (x2−1)2nhas one zero in the
interior of[−1,1]. Extend this argument to the second, third, and ultimately the nth
derivative.
12.5 Associated Legendre Functions 771
12.5 A SSOCIATED LEGENDRE FUNCTIONS
When Helmholtz’s equation is separated in spherical polar coordinates (Section 9.3), one
oftheseparatedODEsis theassociatedLegendreequation
1
sinθd
dθparenleftbigg
sinθdPm
n(cosθ)
dθparenrightbigg
+bracketleftbigg
n(n+1)−m2
sin2θbracketrightbigg
Pm
n(cosθ)=0.(12.71)
Withx=cosθ,this becomes
parenleftbig
1−x2parenrightbigd2
dx2Pm
n(x)−2xd
dxPm
n(x)+bracketleftbigg
n(n+1)−m2
1−x2bracketrightbigg
Pm
n(x)=0.(12.72)
If the azimuthal separation constant m2=0, we have Legendre’s equation, Eq. (12.28).
Theregularsolutions Pm
n(x)(withmnotnecessarilyzero,butaninteger)are
v≡Pm
n(x)=parenleftbig
1−x2parenrightbigm/2dm
dxmPn(x) (12.73a)
withm≥0 aninteger.
One way of developing the solution of the associated Legendre equation is to start with
theregularLegendreequationandconvertitintotheassociatedLegendreequationbyusing
multiple differentiation. These multiple differentiations are suggested by Eq. (12.73a), the
generation of associated Legendre polynomials, and spherical harmonics of Section 12.6
moregenerally,inSection4.3usingraisingorloweringoperatorsofEq.(4.69)repeatedly.
Fortheirderivativeform seeExercise12.6.8. WetakeLegendre’sequation
parenleftbig
1−x2parenrightbig
P′′
n−2xP′
n+n(n+1)Pn=0, (12.74)
andwiththehelpofLeibniz’formula15differentiate mtimes.Theresultis
parenleftbig
1−x2parenrightbig
u′′−2x(m+1)u′+(n−m)(n+m+1)u=0, (12.75)
where
u≡dm
dxmPn(x). (12.76)
Equation(12.74)isnotself-adjoint.Toputitintoself-adjointformandconverttheweight-
ingfunctionto1, wereplace u(x)by
v(x)=parenleftbig
1−x2parenrightbigm/2u(x)=parenleftbig
1−x2parenrightbigm/2dmPn(x)
dxm. (12.73b)
15Leibniz’ formula for the nth derivative of aproduct is
dn
dxnbracketleftbig
A(x)B(x)bracketrightbig
=nsummationdisplay
s=0parenleftbiggn
sparenrightbiggdn−s
dxn−sA(x)ds
dxsB(x),parenleftbiggn
sparenrightbigg
=n!
(n−s)!s!,
abinomial coefficient.
772 Chapter 12 Legendre Functions
Solvingfor uanddifferentiating,weobtain
u′=parenleftbigg
v′+mxv
1−x2parenrightbiggparenleftbig
1−x2parenrightbig−m/2, (12.77)
u′′=bracketleftbigg
v′′+2mxv′
1−x2+mv
1−x2+m(m+2)x2v
(1−x2)2bracketrightbigg
·parenleftbig
1−x2parenrightbig−m/2.(12.78)
Substituting into Eq. (12.74), we find that the new function vsatisfies the self-adjoint
ODE
parenleftbig
1−x2parenrightbig
v′′−2xv′+bracketleftbigg
n(n+1)−m2
1−x2bracketrightbigg
v=0, (12.79)
whichistheassociatedLegendreequation;itreducestoLegendre’sequationwhen misset
equal to zero. Expressed in spherical polar coordinates, the associated Legendre equation
is
1
sinθd
dθparenleftbigg
sinθdv
dθparenrightbigg
+bracketleftbigg
n(n+1)−m2
sin2θbracketrightbigg
v=0. (12.80)
Associated Legendre Polynomials
Theregularsolutions,relabeled Pm
n(x),ar e
v≡Pm
n(x)=parenleftbig
1−x2parenrightbigm/2dm
dxmPn(x). (12.73c)
These are the associated Legendre functions.16Since the highest power of xinPn(x)is
xn,w em u s th a v e m≤n(or them-fold differentiation will drive our function to zero).
In quantum mechanics the requirement that m≤nhas the physical interpretation that the
expectation value of the square of the zcomponent of the angular momentum is less than
orequaltotheexpectationvalueofthesquareoftheangularmomentumvector L,
angbracketleftbig
L2
zangbracketrightbig
≤angbracketleftbig
L2angbracketrightbig
≡integraldisplay
ψ∗
lmL2ψlmd3r.
FromtheformofEq.(12.73c)wemightexpect mtobenonnegative.However,if Pn(x)
isexpressedbyRodrigues’formula,thislimitationon misrelaxedandwemayhave −n≤
m≤n,negativeaswellaspositivevaluesof mbeingpermitted.Theselimitsareconsistent
withthoseobtainedbymeansofraisingandloweringoperatorsinChapter4.Inparticular,
|m|>nis ruled out. This also follows from Eq. (12.73c). Using Leibniz’ differentiation
formulaonceagain,wecanshow(Exercise12.5.1)that Pm
n(x)andP−m
n(x)arerelatedby
P−m
n(x)=(−1)m(n−m)!
(n+m)!Pm
n(x). (12.81)
16Occasionally (as in AMS-55; for the complete reference, see the Additional Readings of Chapter 8), one finds the associated
Legendre functions defined with an additional factor of (−1)m.T h i s(−1)mseems an unnecessary complication at this point.
It will be included in the definition of the spherical harmonics Ymn(θ,ϕ)in Section 12.6. Our definition agrees with Jackson’s
Electrodynamics (seeAdditionalReadingsofChapter11forthisreference).Notealsothattheupperindex misnotanexponent.
12.5 Associated Legendre Functions 773
Fromour definitionoftheassociatedLegendrefunctions Pm
n(x),
P0
n(x)=Pn(x). (12.82)
AgeneratingfunctionfortheassociatedLegendrefunctionsisobtained,viaEq.(12.71),
fromthatoftheordinaryLegendrepolynomials:
(2m)!(1−x2)m/2
2mm!(1−2tx+t2)m+1/2=∞summationdisplay
s=0Pm
s+m(x)ts. (12.83)
If we drop the factor (1−x2)m/2=sinmθfrom this formula and define the polynomi-
alsPm
s+m(x)=Pm
s+m(x)(1−x2)−m/2,then we obtain a practical form of the generating
function,
gm(x,t)≡(2m)!
2mm!(1−2tx+t2)m+1/2=∞summationdisplay
s=0Pm
s+m(x)ts. (12.84)
WecanderivearecursionrelationforassociatedLegendrepolynomialsthatisanalogous
toEqs. (12.14)and(12.17)bydifferentiationasfollows:
parenleftbig
1−2tx+t2parenrightbig∂gm
∂t=(2m+1)(x−t)gm(x,t).
Substitutingthedefiningexpansionsfor associatedLegendrepolynomialsweget
parenleftbig
1−2tx+t2parenrightbigsummationdisplay
ssPm
s+m(x)ts−1=(2m+1)summationdisplay
sbracketleftbig
xPm
s+mts−Pm
s+mts+1bracketrightbig
.
Comparing coefficients of powers of tin these power series, we obtain the recurrence
relation
(s+1)Pm
s+m+1−(2m+1+2s)xPm
s+m+(s+2m)Pm
s+m−1=0.(12.85)
Form=0 ands=nthisrelationis Eq.(12.17).
Before we can use this relation we need to initialize it, that is, relate the associated
Legendre polynomials to ordinary Legendre polynomials. We can use Pm
m=(2m−1)!!
fromEq.(12.73c).Also,since |m|≤n,wemayset Pn+1
n=0andusethistoobtainstarting
valuesfor variousrecursiveprocesses.We observethat
parenleftbig
1−2xt+t2parenrightbig
g1(x,t)=parenleftbig
1−2xt+t2parenrightbig−1/2=summationdisplay
sPs(x)ts,(12.86)
souponinsertingEq. (12.84)wegettherecursion
P1
s+1−2xP1
s+P1
s−1=Ps(x). (12.87)
Moregenerally,wealsohavetheidentity
parenleftbig
1−2xt+t2parenrightbig
gm+1(x,t)=(2m+1)gm(x,t), (12.88)
fromwhichweextracttherecursion
Pm+1
s+m+1−2xPm+1
s+m+Pm+1
s+m−1=(2m+1)Pm
s+m(x), (12.89)
whichrelatestheassociatedLegendrepolynomialswithsuperindex m+1tothosewith m.
Form=0 werecovertheinitialrecursionEq. (12.87).
774 Chapter 12 Legendre Functions
Table 12.2 AssociatedLegendreFunctions
P1
1(x)=(1−x2)1/2=sinθ
P1
2(x)=3x(1−x2)1/2=3cosθsinθ
P2
2(x)=3(1−x2)=3sin2θ
P1
3(x)=3
2(5x2−1)(1−x2)1/2=3
2(5cos2θ−1)sinθ
P2
3(x)=15x(1−x2)=15cosθsin2θ
P3
3(x)=15(1−x2)3/2=15sin3θ
P1
4(x)=5
2(7x3−3x)(1−x2)1/2=5
2(7cos3θ−3cosθ)sinθ
P2
4(x)=15
2(7x2−1)(1−x2)=15
2(7cos2θ−1)sin2θ
P3
4(x)=105x(1−x2)3/2=105cosθsin3θ
P4
4(x)=105(1−x2)2=105sin4θ
Example 12.5.1 LOWEST ASSOCIATED LEGENDRE POLYNOMIALS
Now we are ready to derive the entries of Table 12.2. For m=1 ands=0, Eq. (12.87)
yields P1
1=1, because P1
0=0=P1
−1do not occur in the definition, Eq. (12.84), of the
associated Legendre polynomials. Multiplying by (1−x2)1/2=sinθwe get the first line
ofTable12.2. For s=1 wefind,from Eq.(12.87),
P1
2(x)=P1+2xP1
1=x+2x=3x,
from which the second line of Table 12.2, 3cos θsinθ, follows upon multiplying by sin θ.
Fors=2 weget
P1
3(x)=P2+2xP1
2−P1
1=1
2parenleftbig
3x2−1parenrightbig
+6x2−1=15
2x2−3
2,
inagreementwithline4ofTable12.2.Togetline3weuseEq.(12.88).For m=1,s=0,
this gives P2
2(x)=3P1
1(x)=3, and multiplying by 1 −x2=sin2θreproduces line 3 of
Table12.2.Forlines5,8,9,Eq.(12.84)maybeused,whichweleaveasanexercise.More
generally, we use Eq. (12.89) instead of Eq. (12.87) to get a starting value of Pm
m. Then
Eq. (12.85) reduces to a two-term formula for Pm
m,g i v i n g(2m−1)!!. Note that, if m=0,
thisis(−1)!!=1. /squaresolid
Example 12.5.2 SPECIAL VALUES
Forx=1w eu s e
parenleftbig
1−2t+t2parenrightbig−m−1/2=(1−t)−2m−1=∞summationdisplay
s=0parenleftbigg−2m−1
sparenrightbigg
ts
inEq.(12.84) andfind
Pm
s+m(1)=(2m)!
2mm!parenleftbigg−2m−1
sparenrightbigg
, (12.90)
12.5 Associated Legendre Functions 775
whereparenleftbigg−m
sparenrightbigg
=1
fors=0 andparenleftbigg−m
sparenrightbigg
=(−m)(−m−1)···(1−s−m)
s!
fors≥1. Form=1,s=0w eh a v e P1
1(1)=parenleftbig−3
0parenrightbig
=1; fors=1,P1
2(1)=−parenleftbig−3
1parenrightbig
=3;
fors=2,P1
3(1)=parenleftbig−3
2parenrightbig
=(−3)(−4)
2=6=3
2(5−1), which all agree with Table 12.2. For
x=0 wecanalsousethebinomialexpansion,whichweleaveasanexercise. /squaresolid
Recurrence Relations
As expected and already seen, the associated Legendre functions satisfy recurrence rela-
tions. Because of the existence of two indices instead of just one, we have a wide variety
ofrecurrencerelations:
Pm+1
n−2mx
(1−x2)1/2Pm
n+bracketleftbig
n(n+1)−m(m−1)bracketrightbig
Pm−1
n=0, (12.91)
(2n+1)xPm
n=(n+m)Pm
n−1+(n−m+1)Pm
n+1, (12.92)
(2n+1)parenleftbig
1−x2parenrightbig1/2Pm
n=Pm+1
n+1−Pm+1
n−1
=(n+m)(n+m−1)Pm−1
n−1
−(n−m+1)(n−m+2)Pm−1
n+1,(12.93)
parenleftbig
1−x2parenrightbig1/2Pm′
n=1
2Pm+1
n−1
2(n+m)(n−m+1)Pm−1
n.(12.94)
These relations, and many other similar ones, may be verified by use of the generat-
ing function (Eq. (12.4)), by substitution of the series solution of the associated Legen-
dre equation (12.79) or reduction to the Legendre polynomial recurrence relations, us-
ing Eq. (12.73c). As an example of the last method, consider Eq. (12.93). It is similar to
Eq.(12.23):
(2n+1)Pn(x)=P′
n+1(x)−P′
n−1(x). (12.95)
LetusdifferentiatethisLegendrepolynomialrecurrencerelation mtimestoobtain
(2n+1)dm
dxmPn(x)=dm
dxmP′
n+1(x)−dm
dxmP′
n−1(x)
=dm+1
dxm+1Pn+1(x)−dm+1
dxm+1Pn−1(x). (12.96)
Now multiplying by (1−x2)(m+1)/2and using the definition of Pn(x), we obtain the first
partofEq. (12.93).
776 Chapter 12 Legendre Functions
Parity
The parity relation satisfied by the associated Legendre functions may be determined by
examination of the defining equation (12.73c). As x→−x, we already know that Pn(x)
contributesa (−1)n.T hem-fold differentiationyieldsafactor of (−1)m.Hencewehave
Pm
n(−x)=(−1)n+mPm
n(x). (12.97)
AglanceatTable12.2verifiesthis for 1 ≤m≤n≤4.
Also,from thedefinitioninEq. (12.73c),
Pm
n(±1)=0,form/negationslash=0. (12.98)
Orthogonality
Theorthogonalityofthe Pm
n(x)followsfromtheODE,justasforthe Pn(x)(Section12.3),
ifmis the same for both functions. However, it is instructive to demonstrate the orthogo-
nalitybyanothermethod,amethodthatwillalsoprovidethenormalizationconstant.
UsingthedefinitioninEq.(12.73c)andRodrigues’formula(Eq.(12.65))for Pn(x),we
find
integraldisplay1
−1Pm
p(x)Pm
q(x)dx=(−1)m
2p+qp!q!integraldisplay1
−1Xmparenleftbiggdp+m
dxp+mXpparenrightbiggdq+m
dxq+mXqdx.(12.99)
Thefunction Xisgivenby X≡(x2−1).Ifp/negationslash=q,letusassumethat p<q.Noticethatthe
superscript mis the same for both functions. This is an essential condition. The technique
is to integrate repeatedly by parts; all the integrated parts will vanish as long as there is a
factorX=x2−1.Letusintegrate q+mtimestoobtain
integraldisplay1
−1Pm
p(x)Pm
q(x)dx=(−1)m(−1)q+m
2p+qp!q!integraldisplay1
−1Xqdq+m
dxq+mparenleftbigg
Xmdp+m
dxp+mXpparenrightbigg
dx.(12.100)
Theintegrandontheright-handsideis nowexpandedbyLeibniz’formulatogive
Xqdq+m
dxq+mparenleftbigg
Xmdp+m
dxp+mXpparenrightbigg
=Xqq+msummationdisplay
i=0(q+m)!
i!(q+m−i)!parenleftbiggdq+m−i
dxq+m−iXmparenrightbiggdp+m+i
dxp+m+iXp.(12.101)
Sincetheterm Xmcontainsnopowerof xgreaterthan x2m,wemustha v e
q+m−i≤2m (12.102)
orthederivativewillvanish.Similarly,
p+m+i≤2p. (12.103)
Addingbothinequalitiesyields
q≤p, (12.104)
12.5 Associated Legendre Functions 777
which contradicts our assumption that p<q. Hence, there is no solution for iand the
integralvanishes.Thesameresult obviouslywillfollowif p>q.
For the remaining case, p=q, we have the single term corresponding to i=q−m.
PuttingEq.(12.101)intoEq. (12.100),wehave
integraldisplay1
−1bracketleftbig
Pm
q(x)bracketrightbig2dx=(−1)q+2m(q+m)!
22qq!q!(2m)!(q−m)!integraldisplay1
−1Xqparenleftbiggd2m
dx2mXmparenrightbiggparenleftbiggd2q
dx2qXqparenrightbigg
dx.
(12.105)
Since
Xm=parenleftbig
x2−1parenrightbigm=x2m−mx2m−2+···, (12.106)
d2m
dx2mXm=(2m)!, (12.107)
Eq.(12.105)reducesto
integraldisplay1
−1bracketleftbig
Pm
q(x)bracketrightbig2dx=(−1)q+2m(2q)!(q+m)!
22qq!q!(q−m)!integraldisplay1
−1Xqdx. (12.108)
Theintegralontherightis just
(−1)qintegraldisplayπ
0sin2q+1θdθ=(−1)q22q+1q!q!
(2q+1)!(12.109)
(compare Exercise 8.4.9). Combining Eqs. (12.108) and (12.109), we have the orthogo-
nalityintegral ,
integraldisplay1
−1Pm
p(x)Pm
q(x)dx=2
2q+1·(q+m)!
(q−m)!δpq, (12.110)
or,insphericalpolarcoordinates,
integraldisplayπ
0Pm
p(cosθ)Pm
q(cosθ)sinθdθ=2
2q+1·(q+m)!
(q−m)!δpq. (12.111)
The orthogonality of the Legendre polynomials is a special case of this result, obtained
by setting mequal to zero; that is, for m=0,Eq. (12.110) reduces to Eqs. (12.47) and
(12.48). In both Eqs. (12.110) and (12.111), our Sturm–Liouville theory of Chapter 10
could provide the Kronecker delta. A special calculation, such as the analysis here, is re-
quiredfor thenormalizationconstant.
The orthogonality of the associated Legendre functions over the same interval and with
the same weighting factor as the Legendre polynomials does not contradict the unique-
ness of the Gram–Schmidt construction of the Legendre polynomials, Example 10.3.1.
Table12.2suggests(andSection12.4verifies)thatintegraltext1
−1Pm
p(x)Pm
q(x)dxmaybewrittenas
integraldisplay1
−1Pm
p(x)Pm
q(x)parenleftbig
1−x2parenrightbigmdx,
wherewedefinedearlier
Pm
p(x)parenleftbig
1−x2parenrightbigm/2=Pm
p(x).
778 Chapter 12 Legendre Functions
Thefunctions Pm
p(x)maybeconstructedbytheGram–Schmidtprocedurewiththeweight-
ingfunction w(x)=(1−x2)m.
It is possible to develop an orthogonality relation for associated Legendre functions of
thesamelowerindexbutdifferentupperindex.Wefind
integraldisplay1
−1Pm
n(x)Pk
n(x)parenleftbig
1−x2parenrightbig−1dx=(n+m)!
m(n−m!)δm,k. (12.112)
Notethatanewweightingfactor, (1−x2)−1,hasbeenintroduced.Thisrelationisamath-
ematicalcuriosity.InphysicalproblemswithsphericalsymmetrysolutionsofEqs.(12.80)
and (9.64) appear in conjunction with those of Eq. (9.61), and orthogonality of the az-
imuthaldependencemakesthetwoupperindicesequalandalwaysleadstoEq.(12.111).
Example 12.5.3 MAGNETIC INDUCTION FIELD OF A CURRENT LOOP
Like the other ODEs of mathematical physics, the associated Legendre equation is likely
to pop up quite unexpectedly. As an illustration, consider the magnetic induction field B
and magnetic vector potential Acreated by a single circular current loop in the equatorial
plane(Fig.12.12).
We know from electromagnetic theory that the contribution of current element Idλto
themagneticvectorpotentialis
dA=µ0
4πIdλ
r. (12.113)
(ThisfollowsfromExercise1.14.4andSection9.7)Equation(12.113),plusthesymmetry
ofoursystem,showsthat Ahasonlyaˆϕcomponentandthatthecomponentisindependent
FIGURE 12.12Circularcurrent
loop.
12.5 Associated Legendre Functions 779
ofϕ,17
A=ˆϕAϕ(r,θ). (12.114)
ByMaxwell’sequations,
∇×H=J,∂D
∂t=0(SI units). (12.115)
Since
µ0H=B=∇×A, (12.116)
wehave
∇×(∇×A)=µ0J, (12.117)
whereJis the current density. In our problem Jis zero everywhere except in the current
loop.Therefore,awayfrom theloop,
∇×∇׈ϕAϕ(r,θ)=0, (12.118)
usingEq. (12.114).
From the expression for the curl in spherical polar coordinates (Section 2.5), we obtain
(Example2.5.2)
∇×bracketleftbig
∇׈ϕAϕ(r,θ)bracketrightbig
=ˆϕbracketleftbigg
−∂2Aϕ
∂r2−2
r∂Aϕ
∂r−1
r2∂2Aϕ
∂θ2−1
r2∂
∂θ(cotθAϕ)bracketrightbigg
=0. (12.119)
LettingAϕ(r,θ)=R(r)/Theta1(θ) andseparatingvariables,wehave
r2d2R
dr2+2dR
dr−n(n+1)R=0, (12.120)
d2/Theta1
dθ2+cotθd/Theta1
dθ+n(n+1)/Theta1−/Theta1
sin2θ=0. (12.121)
Thesecondequationis theassociatedLegendreequation(12.80) with m=1,andwemay
immediatelywrite
/Theta1(θ)=P1
n(cosθ). (12.122)
Theseparationconstant n(n+1),nanonnegativeinteger,waschosentokeepthissolution
wellbehaved.
By trial, letting R(r)=rα, we find that α=n,o r−n−1. The first possibility is dis-
carded,for oursolutionmustvanishas r→∞. Hence
Aϕn=bn
rn+1P1
n(cosθ)=cnparenleftbigga
rparenrightbiggn+1
P1
n(cosθ) (12.123)
17Pair off corresponding current elements Idλ(ϕ1)andIdλ(ϕ2),wh er eϕ−ϕ1=ϕ2−ϕ.
780 Chapter 12 Legendre Functions
and
Aϕ(r,θ)=∞summationdisplay
n=1cnparenleftbigga
rparenrightbiggn+1
P1
n(cosθ) (r >a). (12.124)
Hereaistheradiusofthecurrentloop.
SinceAϕmust be invarianttoreflectionin the equatorialplane,by the symmetryof our
problem,
Aϕ(r,cosθ)=Aϕ(r,−cosθ), (12.125)
theparitypropertyof Pm
n(cosθ)(Eq. (12.97)) showsthat cn=0f o rneven.
To complete the evaluation of the constants, we may use Eq. (12.124) to calculate Bz
along the z-axis(Bz=Br(r,θ=0))and compare with the expression obtained from the
Biot–Savartlaw.ThisisthesametechniqueasusedinExample12.3.3.Wehave(compare
Eq.(2.47))
Br=∇×Avextendsinglevextendsingle
r=1
rsinθbracketleftbigg∂
∂θ(sinθAϕ)bracketrightbigg
=cotθ
rAϕ+1
r∂Aϕ
∂θ. (12.126)
Using
∂P1
n(cosθ)
∂θ=−sinθdP1
n(cosθ)
d(cosθ)=−1
2P2
n+n(n+1)
2P0
n (12.127)
(Eq. (12.94))andthenEq. (12.91)with m=1,
P2
n(cosθ)−2cosθ
sinθP1
n(cosθ)+n(n+1)Pn(cosθ)=0, (12.128)
weobtain
Br(r,θ)=∞summationdisplay
n=1cnn(n+1)an+1
rn+2Pn(cosθ), r>a, (12.129)
(for allθ). Inparticular,for θ=0,
Br(r,0)=∞summationdisplay
n=1cnn(n+1)an+1
rn+2. (12.130)
Wemayalso obtain
Bθ(r,θ)=−1
r∂(rAϕ)
∂r=∞summationdisplay
n=1cnnan+1
rn+2P1
n(cosθ), r>a, (12.131)
TheBiot–Savartlawstatesthat
dB=µ0
4πIdλ׈r
r2(SI units). (12.132)
We now integrate over the perimeter of our loop (radius a). The geometry is shown in
Fig.12.13.Theresultingmagneticinductionfieldis ˆzBz, alongthe z-axis, with
Bz=µ0I
2a2parenleftbig
a2+z2parenrightbig−3/2=µ0I
2a2
z3parenleftbigg
1+a2
z2parenrightbigg−3/2
. (12.133)
12.5 Associated Legendre Functions 781
FIGURE 12.13Biot–Savartlawappliedtoacircularloop.
Expandingbythebinomialtheorem,weobtain
Bz=µ0I
2a2
z3bracketleftbigg
1−3
2parenleftbigga
zparenrightbigg2
+15
8parenleftbigga
zparenrightbigg4
−···bracketrightbigg
=µ0I
2a2
z3∞summationdisplay
s=0(−1)s(2s+1)!!
(2s)!!parenleftbigga
zparenrightbigg2s
,z>a. (12.134)
EquatingEqs. (12.130)and(12.134)termbyterm(with r=z),18wefind
c1=µ0I
4,c 3=−µ0I
16,c 2=c4=···=0.
cn=(−1)(n−1)/2µ0I
2n(n+1)·(n/2)!
[(n−1)/2]!(1
2)!,nodd.(12.135)
Equivalently,wemaywrite
c2n+1=(−1)nµ0I
22n+2·(2n)!
n!(n+1)!=(−1)nµ0I
2·(2n−1)!!
(2n+2)!!(12.136)
18Thedescending powerseries is alsounique.
782 Chapter 12 Legendre Functions
and
Aϕ(r,θ)=parenleftbigga
rparenrightbigg2∞summationdisplay
n=0c2n+1parenleftbigga
rparenrightbigg2n
P1
2n+1(cosθ), (12.137)
Br(r,θ)=a2
r3∞summationdisplay
n=0c2n+1(2n+1)(2n+2)parenleftbigga
rparenrightbigg2n
P2n+1(cosθ),(12.138)
Bθ(r,θ)=a2
r3∞summationdisplay
n=0c2n+1(2n+1)parenleftbigga
rparenrightbigg2n
P1
2n+1(cosθ). (12.139)
These fields may be described in closed form by the use of elliptic integrals. Exer-
cise 5.8.4 is an illustration of this approach. A third possibility is direct integration of
Eq. (12.113) by expanding the denominator of the integral for Aϕin Exercise 5.8.4 as
a Legendre polynomial generating function. The current is specified by Dirac delta func-
tions.Thesemethodshavetheadvantageof yieldingtheconstants cndirectly.
Acomparisonofmagneticcurrentloopdipolefieldsandfiniteelectricdipolefieldsmay
beofinterest.Forthemagneticcurrentloopdipole,theprecedinganalysisgives
Br(r,θ)=µ0I
2a2
r3bracketleftbigg
P1−3
2parenleftbigga
rparenrightbigg2
P3+···bracketrightbigg
, (12.140)
Bθ(r,θ)=µ0I
4a2
r3bracketleftbigg
P1
1−3
4parenleftbigga
rparenrightbigg2
P1
3+···bracketrightbigg
. (12.141)
Fromthefiniteelectricdipolepotentialof Section12.1wehave
Er(r,θ)=qa
πε0r3bracketleftbigg
P1+2parenleftbigga
rparenrightbigg2
P3+···bracketrightbigg
, (12.142)
Eθ(r,θ)=qa
2πε0r3bracketleftbigg
P1
1+parenleftbigga
rparenrightbigg2
P1
3+···bracketrightbigg
. (12.143)
The two fields agree in form as far as the leading term is concerned (r−3P1), and this is
thebasis forcallingthembothdipolefields.
As with electric multipoles, it is sometimes convenient to discuss pointmagnetic mul-
tipoles (see Fig. 12.14). For the dipole case, Eqs. (12.140) and (12.141), the point dipole
is formed by taking the limit a→0,I→∞, withIa2held constant. With na unit vector
normal to the current loop (positive sense by right-hand rule, Section 1.10), the magnetic
momentmisgivenby m=nIπa2. /squaresolid
FIGURE 12.14Electricdipole.
12.5 Associated Legendre Functions 783
Exercises
12.5.1 Provethat
P−m
n(x)=(−1)m(n−m)!
(n+m)!Pm
n(x),
wherePm
n(x)is definedby
Pm
n(x)=1
2nn!parenleftbig
1−x2parenrightbigm/2dn+m
dxn+mparenleftbig
x2−1parenrightbign.
Hint.Oneapproachis toapplyLeibniz’formulato (x+1)n(x−1)n.
12.5.2 Showthat
P1
2n(0)=0,
P1
2n+1(0)=(−1)n(2n+1)!
(2nn!)2=(−1)n(2n+1)!!
(2n)!!,
byeachofthesethreemethods:
(a) useofrecurrencerelations,
(b) expansionof thegeneratingfunction,
(c) Rodrigues’formula.
12.5.3 Evaluate Pm
n(0).
ANS.Pm
n(0)=
(−1)(n−m)/2 (n+m)!
2n((n−m)/2)!((n+m)/2!),n+meven,
0,n +modd.
Also,Pm
n(0)=(−1)(n−m)/2(n+m−1)!!
(n−m)!!,n+meven.
12.5.4 Showthat
Pn
n(cosθ)=(2n−1)!!sinnθ, n=0,1,2,....
12.5.5 DerivetheassociatedLegendrerecurrencerelation
Pm+1
n(x)−2mx
(1−x2)1/2Pm
n(x)+bracketleftbig
n(n+1)−m(m−1)bracketrightbig
Pm−1
n(x)=0.
12.5.6 Developarecurrencerelationthatwillyield P1
n(x)as
P1
n(x)=f1(x,n)P n(x)+f2(x,n)P n−1(x).
Followeither(a) or(b).
(a) Derive a recurrence relation of the preceding form. Give f1(x,n)andf2(x,n)
explicitly.
(b) Findtheappropriaterecurrencerelationinprint.
(1) Givethesource.
784 Chapter 12 Legendre Functions
(2) Verifytherecurrencerelation.
ANS.P1
n(x)=−nx
(1−x2)1/2Pn+n
(1−x2)1/2Pn−1.
12.5.7 Showthat
sinθd
dcosθPn(cosθ)=P1
n(cosθ).
12.5.8 Showthat
(a)integraldisplayπ
0parenleftbiggdPm
n
dθdPm
n′
dθ+m2Pm
nPm
n′
sin2θparenrightbigg
sinθdθ=2n(n+1)
2n+1(n+m)!
(n−m)!δnn′,
(b)integraldisplayπ
0parenleftbiggP1
n
sinθdP1
n′
dθ+P1
n′
sinθdP1
n
dθparenrightbigg
sinθdθ=0.
Theseintegralsoccurinthetheoryofscatteringofelectromagneticwavesbyspheres.
12.5.9 AsarepeatofExercise12.3.10,show,usingassociatedLegendrefunctions,that
integraldisplay1
−1xparenleftbig
1−x2parenrightbig
P′
n(x)P′
m(x)dx=n+1
2n+1·2
2n−1·n!
(n−2)!δm,n−1
+n
2n+1·2
2n+3·(n+2)!
n!δm,n+1.
12.5.10 Evaluate
integraldisplayπ
0sin2θP1
n(cosθ)dθ.
12.5.11 TheassociatedLegendrepolynomial Pm
n(x)satisfiestheself-adjointODE
parenleftbig
1−x2parenrightbigd2Pm
n(x)
dx2−2xdPm
n(x)
dx+bracketleftbigg
n(n+1)−m2
1−x2bracketrightbigg
Pm
n(x)=0.
Fromthedifferentialequationsfor Pm
n(x)andPk
n(x)showthat
integraldisplay1
−1Pm
n(x)Pk
n(x)dx
1−x2=0,
fork/negationslash=m.
12.5.12 Determinethevectorpotentialofamagneticquadrupolebydifferentiatingthemagnetic
dipolepotential.
ANS.AMQ=µ0
2parenleftbig
Ia2parenrightbig
(dz)ˆϕP1
2(cosθ)
r3+higher-orderterms .
BMQ=µ0parenleftbig
Ia2parenrightbig
(dz)bracketleftbigg
ˆr3P2(cosθ)
r4+ˆθP1
2(cosθ)
r4bracketrightbigg
.
12.5 Associated Legendre Functions 785
This corresponds to placing a current loop of radius aatz→dzand an oppositely
directed current loop at z→−dzand letting a→0 subject to (dz)a(dipole strength)
equalconstant.
Anotherapproachtothisproblemwouldbetointegrate dA(Eq.(12.113),toexpandthe
denominator in a series of Legendre polynomials, and to use the Legendre polynomial
additiontheorem(Section12.8).
12.5.13 Asingleloopofwireof radius acarriesaconstantcurrent I.
(a) Findthemagneticinduction Bforr<a,θ=π/2.
(b) Calculate the integral of the magnetic flux (B·dσ)over the area of the current
loop,thatis
integraldisplaya
0integraldisplay2π
0Bzparenleftbigg
r,θ=π
2parenrightbigg
dϕrdr.
ANS.∞.
The Earth is within such a ring current, in which Iapproximates millions of amperes
arisingfromthedriftof chargedparticlesintheVanAllenbelt.
12.5.14 (a) Showthatinthepointdipolelimitthemagneticinductionfieldofthecurrentloop
becomes
Br(r,θ)=µ0
2πm
r3P1(cosθ),
Bθ(r,θ)=µ0
2πm
r3P1
1(cosθ)
withm=Iπa2.
(b) Comparetheseresultswiththemagneticinductionofthepointmagneticdipoleof
Exercise1.8.17. Take m=ˆzm.
12.5.15 Auniformlychargedsphericalshellisrotatingwithconstantangularvelocity.
(a) Calculatethemagneticinduction Balongtheaxisofrotationoutsidethesphere.
(b) Using the vector potential series of Section 12.5, find Aand then Bfor all space
outsidethesphere.
12.5.16 In the liquid drop model of the nucleus, the spherical nucleus is subjected to small
deformations. Consider a sphere of radius r0that is deformed so that its new surface is
givenby
r=r0bracketleftbig
1+α2P2(cosθ)bracketrightbig
.
Findtheareaofthedeformedspherethroughtermsoforder α2
2.
Hint.
dA=bracketleftbigg
r2+parenleftbiggdr
dθparenrightbigg2bracketrightbigg1/2
rsinθdθdϕ.
ANS.A=4πr2
0bracketleftbig
1+4
5α2
2+Oparenleftbig
α3
2parenrightbigbracketrightbig
.
786 Chapter 12 Legendre Functions
Note. The area element dAfollows from noting that the line element dsfor fixed ϕis
givenby
ds=parenleftbig
r2dθ2+dr2parenrightbig1/2=parenleftbigg
r2+parenleftbiggdr
dθparenrightbigg2parenrightbigg1/2
dθ.
12.5.17 A nuclear particle is in a spherical square well potential V(r,θ,ϕ)=0f o r0≤r<a
and∞forr>a.Theparticleisdescribedbyawavefunction ψ(r,θ,ϕ) whichsatisfies
thewaveequation
−¯h2
2M∇2ψ+V0ψ=Eψ, r<a,
andtheboundarycondition
ψ(r=a)=0.
Show that for the energy Eto be a minimum there must be no angular dependence in
thewavefunction;thatis, ψ=ψ(r).
Hint.Theproblemcentersontheboundaryconditionontheradialfunction.
12.5.18 (a) Write a subroutine to calculate the numerical value of the associated Legendre
functionP1
N(x)for givenvaluesof Nandx.
Hint. With the known forms of P1
1andP1
2you can use the recurrence relation
Eq.(12.92) togenerate P1
N,N>2.
(b) Checkyoursubroutinebyhavingitcalculate P1
N(x)forx=0.0(0.5)1.0andN=
1(1)10. Check these numerical values against the known values of P1
N(0)and
P1
N(1)andagainstthetabulatedvaluesof P1
N(0.5).
12.5.19 Calculatethemagneticvectorpotentialofacurrentloop,Example12.5.1.Tabulateyour
resultsfor r/a=1.5(0.5)5.0andθ=0◦(15◦)90◦.Includetermsintheseriesexpansion,
Eq.(12.137),untiltheabsolutevaluesofthetermsdropbelowtheleadingtermbyafac-
torof 105ormore.
Note. This associated Legendre expansion can be checked by comparison with the el-
lipticintegralsolution,Exercise5.8.4.
Checkvalue. Forr/a=4.0andθ=20◦,
Aϕ/µ0I=4.9398×10−3.
12.6 S PHERICAL HARMONICS
In the separation of variables of (1) Laplace’s equation, (2) Helmholtz’s or the space-
dependence of the electromagnetic wave equation, and (3) the Schrödinger wave equation
forcentralforcefields,
∇2ψ+k2f(r)ψ=0, (12.144)
12.6 Spherical Harmonics 787
theangulardependence,comingentirelyfromtheLaplacianoperator,is19
/Phi1(ϕ)
sinθd
dθparenleftbigg
sinθd/Theta1
dθparenrightbigg
+/Theta1(θ)
sin2θd2/Phi1(ϕ)
dϕ2+n(n+1)/Theta1(θ)/Phi1(ϕ)=0.(12.145)
Azimuthal Dependence — Orthogonality
Theseparatedazimuthalequationis
1
/Phi1(ϕ)d2/Phi1(ϕ)
dϕ2=−m2, (12.146)
withsolutions
/Phi1(ϕ)=e−imϕ,eimϕ, (12.147)
withminteger,whichsatisfy theorthogonalcondition
integraldisplay2π
0e−im1ϕeim2ϕdϕ=2πδm1m2. (12.148)
Notice that it is the product /Phi1∗
m1(ϕ)/Phi1m2(ϕ)that is taken and that∗is used to indicate the
complex conjugate function. This choice is not required, but it is convenient for quantum
mechanicalcalculations.We couldhaveused
/Phi1=sinmϕ,cosmϕ (12.149)
and the conditions of orthogonality that form the basis for Fourier series (Chapter 14).
For applications such as describing the Earth’s gravitational or magnetic field, sin mϕand
cosmϕwouldbethepreferred choice(seeExample12.6.1).
Inelectrostaticsandmostotherphysicalproblemswerequire mtobeanintegerinorder
that/Phi1(ϕ)be a single-valued function of the azimuth angle. In quantum mechanics the
questionismuchmoreinvolved:ComparethefootnoteinSection9.3.
BymeansofEq. (12.148),
/Phi1m=1√
2πeimϕ(12.150)
is orthonormal (orthogonal and normalized) with respect to integration over the azimuth
angleϕ.
19For a separation constant of the form n(n+1)withnan integer, a Legendre-equation-series solution becomes a polynomial.
Otherwise both series solutions diverge, Exercise 9.5.5.
788 Chapter 12 Legendre Functions
Polar Angle Dependence
Splitting off the azimuthal dependence, the polar angle dependence (θ)leads to the asso-
ciatedLegendreequation(12.80), whichis satisfied bytheassociatedLegendrefunctions;
that is,/Theta1(θ)=Pm
n(cosθ). To include negative values of m, we use Rodrigues’ formula,
Eq.(12.65), inthedefinitionof Pm
n(cosθ).This leadsto
Pm
n(cosθ)=1
2nn!parenleftbig
1−x2parenrightbigm/2dm+n
dxm+nparenleftbig
x2−1parenrightbign,−n≤m≤n.(12.151)
Pm
n(cosθ)andP−m
n(cosθ)arerelatedasindicatedinExercise12.5.1.Anadvantageofthis
approach over simply defining Pm
n(cosθ)for 0≤m≤nand requiring that P−m
n=Pm
nis
thattherecurrencerelationsvalidfor 0 ≤m≤nremainvalidfor −n≤m<0.
Normalizing the associated Legendre function by Eq. (12.110), we obtain the orthonor-
malfunctionsradicalBigg
2n+1
2(n−m)!
(n+m)!Pm
n(cosθ),−n≤m≤n, (12.152)
whichareorthonormalwithrespecttothepolarangle θ.
Spherical Harmonics
The function /Phi1m(ϕ)(Eq. (12.150)) is orthonormal with respect to the azimuthal an-
gleϕ. We take the product of /Phi1m(ϕ)and the orthonormal function in polar angle from
Eq.(12.152)anddefine
Ym
n(θ,ϕ)≡(−1)mradicalBigg
2n+1
4π(n−m)!
(n+m)!Pm
n(cosθ)eimϕ(12.153)
to obtain functions of two angles (and two indices) that are orthonormal over the spheri-
cal surface. These Ym
n(θ,ϕ)are spherical harmonics, of which the first few are plotted in
Fig.12.15.Thecompleteorthogonalityintegralbecomes
integraldisplay2π
ϕ=0integraldisplayπ
θ=0Ym1∗
n1(θ,ϕ)Ym2n2(θ,ϕ)sinθdθdϕ=δn1n2δm1m2. (12.154)
The extra (−1)mincluded in the defining equation of Ym
n(θ,ϕ)deserves some comment.
It is clearly legitimate, since Eq. (12.144) is linear and homogeneous. It is not necessary,
but in moving on to certain quantum mechanical calculations, particularly in the quan-
tum theory of angular momentum (Section 12.7), it is most convenient. The factor (−1)m
is a phase factor, often called the Condon–Shortley phase, after the authors of a classic
text on atomic spectroscopy. The effect of this (−1)m(Eq. (12.153)) and the (−1)mof
Eq. (12.73c) for P−m
n(cosθ)is to introduce an alternation of sign among the positive m
sphericalharmonics.Thisis showninTable12.3.
The functions Ym
n(θ,ϕ)acquired the name spherical harmonics first because they are
defined over the surface of a sphere with θthe polar angle and ϕthe azimuth. The har-
monicwas included because solutions of Laplace’s equation were called harmonic func-
tionsand Ym
n(cos,ϕ)is theangularpartofsuchasolution.
12.6 Spherical Harmonics 789
FIGURE 12.15[ℜYm
l(θ,ϕ)]2for 0≤l≤3,m=0,...,3.
790 Chapter 12 Legendre Functions
Table 12.3 Spherical Harmon-
ics(Condon–ShortleyPhase)
Y0
0(θ,ϕ)=1√
4π
Y1
1(θ,ϕ)=−radicalbigg
3
8πsinθeiϕ
Y0
1(θ,ϕ)=radicalbigg
3
4πcosθ
Y−1
1(θ,ϕ)=+radicalbigg
3
8πsinθe−iϕ
Y2
2(θ,ϕ)=radicalbigg
5
96π3sin2θe2iϕ
Y1
2(θ,ϕ)=−radicalbigg
5
24π3sinθcosθeiϕ
Y0
2(θ,ϕ)=radicalbigg
5
4πparenleftbigg3
2cos2θ−1
2parenrightbigg
Y−1
2(θ,ϕ)=+radicalbigg
5
24π3sinθcosθe−iϕ
Y−2
2(θ,ϕ)=radicalbigg
5
96π3sin2θe−2iϕ
In the framework of quantum mechanics Eq. (12.145) becomes an orbital angular mo-
mentum equation and the solution YM
L(θ,ϕ)(nreplaced by L,mreplaced by M)i sa n
angular momentum eigenfunction, Lbeing the angular momentum quantum number and
Mthez-axis projection of L. These relationships are developed in more detail in Sec-
tions4.3and12.7.
Laplace Series, Expansion Theorem
Part of the importance of spherical harmonics lies in the completeness property, a con-
sequence of the Sturm–Liouville form of Laplace’s equation. This property, in this case,
means that any function f(θ,ϕ)(with sufficient continuity properties) evaluated over the
surfaceofthespherecanbeexpandedinauniformlyconvergentdoubleseriesofspherical
harmonics20(Laplace’sseries):
f(θ,ϕ)=summationdisplay
m,namnYm
n(θ,ϕ). (12.155)
Iff(θ,ϕ)is known, the coefficients can be immediately found by the use of the orthogo-
nalityintegral.
20For a proof of this fundamental theorem see E. W. Hobson, The Theory of Spherical and Ellipsoidal Harmonics ,N e wY o r k :
Chelsea (1955), Chapter VII. If f(θ,ϕ)is discontinuous wemay still have convergence in the mean,Section10.4.
12.6 Spherical Harmonics 791
Table 12.4 GravityFieldCoefficients,Eq. (12.156)
CoefficientaEarth Moon Mars
C20 1.083×10−3(0.200±0.002)×10−3(1.96±0.01)×10−3
C22 0.16×10−5(2.4±0.5)×10−5(−5±1)×10−5
S22 −0.09×10−5(0.5±0.6)×10−5(3±1)×10−5
aC20represents an equatorial bulge, whereas C22andS22represent an azimuthal dependence of the
gravitationalfield.
Example 12.6.1 LAPLACE SERIES —G RAVITY FIELDS
The gravity fields of the Earth, the Moon, and Mars have been described by a Laplace
serieswithrealeigenfunctions:
U(r,θ,ϕ)=GM
RbracketleftbiggR
r−∞summationdisplay
n=2nsummationdisplay
m=0parenleftbiggR
rparenrightbiggn+1braceleftbig
CnmYe
mn(θ,ϕ)+SnmYo
mn(θ,ϕ)bracerightbigbracketrightbigg
.(12.156)
HereMisthemassofthebodyand Ristheequatorialradius.Therealfunctions Ye
mnand
Yo
mnare definedby
Ye
mn(θ,ϕ)=Pm
n(cosθ)cosmϕ, Yo
mn(θ,ϕ)=Pm
n(cosθ)sinmϕ.
For applications such as this, the real trigonometric forms are preferred to the imaginary
exponential form of YM
L(θ,ϕ). Satellite measurements have led to the numerical values
showninTable12.4. /squaresolid
Exercises
12.6.1 Show that the parity of YM
L(θ,ϕ)is(−1)L. Note the disappearance of any Mdepen-
dence.
Hint.FortheparityoperationinsphericalpolarcoordinatesseeExercise2.5.8andfoot-
note7inSection12.2.
12.6.2 Provethat
YM
L(0,ϕ)=parenleftbigg2L+1
4πparenrightbigg1/2
δM,0.
12.6.3 InthetheoryofCoulombexcitationofnucleiweencounter YM
L(π/2,0). Showthat
YM
Lparenleftbiggπ
2,0parenrightbigg
=parenleftbigg2L+1
4πparenrightbigg1/2[(L−M)!(L+M)!]1/2
(L−M)!!(L+M)!!(−1)(L+M)/2forL+Meven,
=0forL+Modd.
Here
(2n)!!=2n(2n−2)···6·4·2,
(2n+1)!!=(2n+1)(2n−1)···5·3·1.
792 Chapter 12 Legendre Functions
12.6.4 (a) Express the elements of the quadrupole moment tensor xixjas a linear combina-
tionof thesphericalharmonics Ym
2(andY0
0).
Note.Thetensor xixjisreducible.The Y0
0indicatesthepresenceofascalarcom-
ponent.
(b) Thequadrupolemomenttensorisusuallydefinedas
Qij=integraldisplayparenleftbig
3xixj−r2δijparenrightbig
ρ(r)dτ,
withρ(r)the charge density. Express the components of (3xixj−r2δij)in terms
ofr2YM
2.
(c) Whatisthesignificanceof the −r2δijterm?
Hint.CompareSections2.9and4.4.
12.6.5 TheorthogonalazimuthalfunctionsyieldausefulrepresentationoftheDiracdeltafunc-
tion.Showthat
δ(ϕ1−ϕ2)=1
2π∞summationdisplay
m=−∞expbracketleftbig
im(ϕ1−ϕ2)bracketrightbig
.
12.6.6 Derivethesphericalharmonicclosurerelation
∞summationdisplay
l=0+lsummationdisplay
m=−lYm
l(θ1,ϕ1)Ym∗
l(θ2,ϕ2)=1
sinθ1δ(θ1−θ2)δ(ϕ1−ϕ2)
=δ(cosθ1−cosθ2)δ(ϕ1−ϕ2).
12.6.7 Thequantummechanicalangularmomentumoperators Lx±iLyaregivenby
Lx+iLy=eiϕparenleftbigg∂
∂θ+icotθ∂
∂ϕparenrightbigg
,
Lx−iLy=−e−iϕparenleftbigg∂
∂θ−icotθ∂
∂ϕparenrightbigg
.
Showthat
(a)(Lx+iLy)YM
L(θ,ϕ)=radicalbig
(L−M)(L+M+1)YM+1
L(θ,ϕ),
(b)(Lx−iLy)YM
L(θ,ϕ)=radicalbig
(L+M)(L−M+1)YM−1
L(θ,ϕ).
12.6.8 WithL±givenby
L±=Lx±iLy=±e±iϕbracketleftbigg∂
∂θ±icotθ∂
∂ϕbracketrightbigg
,
showthat
(a)Ym
l=radicalBigg
(l+m)!
(2l)!(l−m)!(L−)l−mYl
l,
12.7 Orbital Angular Momentum Operators 793
(b)Ym
l=radicalBigg
(l−m)!
(2l)!(l+m)!(L+)l+mY−l
l.
12.6.9 Insomecircumstancesitisdesirabletoreplacetheimaginaryexponentialofourspher-
ical harmonic by sine or cosine. Morse and Feshbach (see the General References at
book’send)define
Ye
mn=Pm
n(cosθ)cosmϕ,
Yo
mn=Pm
n(cosθ)sinmϕ,
where
integraldisplay2π
0integraldisplayπ
0bracketleftbig
Yeoro
mn(θ,ϕ)bracketrightbig2sinθdθdϕ=4π
2(2n+1)(n+m)!
(n−m)!,n=1,2,...
=4πforn=0(Yo
00isundefined ).
These spherical harmonics are often named according to the patterns of their positive
and negative regions on the surface of a sphere—zonal harmonics for m=0, sectoral
harmonics for m=n, and tesseral harmonics for 0 <m<n.F o rYe
mn,n=4,m=
0,2,4,indicateonadiagramofahemisphere(onediagramforeachsphericalharmonic)
theregionsinwhichthesphericalharmonicis positive.
12.6.10 Afunction f(r,θ,ϕ) maybeexpressedasa Laplaceseries
f(r,θ,ϕ)=summationdisplay
l,malmrlYm
l(θ,ϕ).
With/angbracketleft/angbracketrightsphereusedtomeantheaverageoverasphere(centeredontheorigin),showthat
angbracketleftbig
f(r,θ,ϕ)angbracketrightbig
sphere=f(0,0,0).
12.7 O RBITAL ANGULAR MOMENTUM OPERATORS
Now we return to the specific orbital angular momentum operators Lx,Ly, andLzof
quantummechanicsintroducedinSection4.3. Equation(4.68)becomes
LzψLM(θ,ϕ)=MψLM(θ,ϕ),
andwewanttoshowthat
ψLM(θ,ϕ)=YM
L(θ,ϕ)
aretheeigenfunctions |LM/angbracketrightofL2andLzofSection4.3insphericalpolarcoordinates,the
sphericalharmonics.Theexplicitformof Lz=−i∂/∂ϕfromExercise2.5.13indicatesthat
ψLMhas aϕdependence of exp (iMϕ)—withMan integer to keep ψLMsingle-valued.
AndifMis aninteger,then Lisanintegeralso.
To determine the θdependence of ψLM(θ,ϕ), we proceed in two main steps: (1) the
determination of ψLL(θ,ϕ)and (2) the development of ψLM(θ,ϕ)in terms of ψLLwith
thephasefixedby ψL0.L et
ψLM(θ,ϕ)=/Theta1LM(θ)eiMϕ. (12.157)
794 Chapter 12 Legendre Functions
FromL+ψLL=0,Lbeing the largest M, using the form of L+given in Exercises 2.5.14
and12.6.7, wehave
ei(L+1)ϕbracketleftbiggd
dθ−Lcotθbracketrightbigg
/Theta1LL(θ)=0, (12.158)
andthus
ψLL(θ,ϕ)=cLsinLθeiLϕ. (12.159)
Normalizing,weobtain
c∗
LcLintegraldisplay2π
0integraldisplayπ
0sin2L+1θdθdϕ=1. (12.160)
Theθintegralmaybeevaluatedasabetafunction(Exercise8.4.9) and
|cL|=radicalBigg
(2L+1)!!
4π(2L)!!=√(2L)!
2LL!radicalbigg
2L+1
4π. (12.161)
Thiscompletesourfirst step.
To obtain the ψLM,M/negationslash=±L, we return to the ladder operators. From Eqs. (4.83) and
(4.84)andasshowninExercise12.7.2( J+replacedby L+andJ−replacedby L−),
ψLM(θ,ϕ)=radicalBigg
(L+M)!
(2L)!(L−M)!(L−)L−MψLL(θ,ϕ),
(12.162)
ψLM(θ,ϕ)=radicalBigg
(L−M)!
(2L)!(L+M)!(L+)L+MψL,−L(θ,ϕ).
Again, note that the relative phases are set by the ladder operators. L+andL−operating
on/Theta1LM(θ)eiMϕmaybewrittenas
L+/Theta1LM(θ)eiMϕ=ei(M+1)ϕbracketleftbiggd
dθ−Mcotθbracketrightbigg
/Theta1LM(θ)
=ei(M+1)ϕsin1+Mθd
d(cosθ)sin−M/Theta1LM(θ),
(12.163)
L−/Theta1LM(θ)eiMϕ=−ei(M−1)ϕbracketleftbiggd
dθ+Mcotθbracketrightbigg
/Theta1LM(θ)
=ei(M−1)ϕsin1−Mθd
d(cosθ)sinMθ/Theta1LM(θ).
Repeatingtheseoperations ntimesyields
(L+)n/Theta1LM(θ)eiMϕ=(−1)nei(M+n)ϕsinn+Mθdnsin−Mθ/Theta1LM(θ)
d(cosθ)n,
(12.164)
(L−)n/Theta1LM(θ)eiMϕ=ei(M−n)ϕsinn−MθdnsinMθ/Theta1LM(θ)
d(cosθ)n.
12.7 Orbital Angular Momentum Operators 795
FromEq. (12.163),
ψLM(θ,ϕ)=cLradicalBigg
(L+M)!
(2L)!(L−M)!eiMϕsin−MθdL−M
d(cosθ)L−Msin2Lθ,(12.165)
andforM=−L:
ψL,−L(θ,ϕ)=cL
(2L)!e−iLϕsinLθd2L
d(cosθ)2Lsin2Lθ
=(−1)LcLsinLθe−iLϕ. (12.166)
Notethecharacteristic (−1)LphaseofψL,−Lrelativeto ψL,L.This(−1)Lentersfrom
sin2Lθ=parenleftbig
1−x2parenrightbigL=(−1)Lparenleftbig
x2−1parenrightbigL. (12.167)
CombiningEqs. (12.163), (12.163),and(12.166),weobtain
ψLM(θ,ϕ)=(−1)LcLradicalBigg
(L−M)!
(2L)!(L+M)!(−1)L+MeiMϕsinMθdL+Msin2Lθ
d(cosθ)L+M.(12.168)
Equations(12.165)and(12.168)agreeif
ψL0(θ,ϕ)=cL1√(2L)!dL
(dcosθ)Lsin2Lθ. (12.169)
UsingRodrigues’formula,Eq. (12.65), wehave
ψL0(θ,ϕ)=(−1)LcL2LL!√(2L)!PL(cosθ)
=(−1)LcL
|cL|radicalbigg
2L+1
4πPL(cosθ). (12.170)
The last equality follows from Eq. (12.161). We now demand that ψL0(0,0)be real and
positive.Therefore
cL=(−1)L|cL|=(−1)L√(2L)!
2LL!radicalbigg
2L+1
4π. (12.171)
With(−1)LcL/|cL|=1,ψL0(θ,ϕ)in Eq. (12.170) may be identified with the spherical
harmonic Y0
L(θ,ϕ)ofSection12.6.
796 Chapter 12 Legendre Functions
Whenwesubstitutethevalueof (−1)LcLintoEq.(12.168),
ψLM(θ,ϕ)=√(2L)!
2LL!radicalbigg
2L+1
4πradicalBigg
(L−M)!
(2L)!(L+M)!(−1)L+M
·eiMϕsinMθdL+M
d(cosθ)L+Msin2Lθ
=radicalbigg
2L+1
4πradicalBigg
(L−M)!
(L+M)!eiMϕ(−1)M
·braceleftbigg1
2LL!parenleftbig
1−x2parenrightbigM/2dL+M
dxL+Mparenleftbig
x2−1parenrightbigLbracerightbigg
,x=cosθ, M≥0.
(12.172)
The expression in the curly bracket is identified as the associated Legendre function
(Eq. (12.151),andwehave
ψLM(θ,ϕ)=YM
L(θ,ϕ)
=(−1)MradicalBigg
2L+1
4π·(L−M)!
(L+M)!·PM
L(cosθ)eiMϕ,M≥0,
(12.173)
in complete agreement with Section 12.6. Then by Eq. (12.73c), YM
Lfor negative super-
scriptisgivenby
Y−M
L(θ,ϕ)=(−1)Mbracketleftbig
YM
L(θ,ϕ)bracketrightbig∗. (12.174)
•Our angular momentum eigenfunctions ψLM(θ,ϕ)are identified with the spherical
harmonics. The phase factor (−1)Mis associated with the positive values of Mand is
seentobeaconsequenceoftheladderoperators.
•OurdevelopmentofsphericalharmonicsheremaybeconsideredaportionofLiealge-
bra—relatedtogrouptheory,Section4.3.
Exercises
12.7.1 Usingtheknownforms of L+andL−(Exercises2.5.14 and12.6.7), showthat
integraldisplaybracketleftbig
YM
Lbracketrightbig∗L−parenleftbig
L+YM
Lparenrightbig
d/Omega1=integraldisplayparenleftbig
L+YM
Lparenrightbig∗parenleftbig
L+YM
Lparenrightbig
d/Omega1.
12.7.2 Derivetherelations
(a)ψLM(θ,ϕ)=radicalBigg
(L+M)!
(2L)!(L−M)!(L−)L−MψLL(θ,ϕ),
(b)ψLM(θ,ϕ)=radicalBigg
(L−M)!
(2L)!(L+M)!(L+)L+MψL,−L(θ,ϕ).
12.8 Addition Theorem for Spherical Harmonics 797
Hint.Equations(4.83) and(4.84) maybehelpful.
12.7.3 Derivethemultipleoperatorequations
(L+)n/Theta1LM(θ)eiMϕ=(−1)nei(M+n)ϕsinn+Mθdnsin−Mθ/Theta1LM(θ)
d(cosθ)n,
(L−)n/Theta1LM(θ)eiMϕ=ei(M−n)ϕsinn−MθdnsinMθ/Theta1LM(θ)
d(cosθ)n.
Hint.Trymathematicalinduction.
12.7.4 Show,using (L−)n, that
Y−M
L(θ,ϕ)=(−1)MY∗M
L(θ,ϕ).
12.7.5 Verifybyexplicitcalculationthat
(a)L+Y0
1(θ,ϕ)=−radicalbigg
3
4πsinθeiϕ=√
2Y1
1(θ,ϕ),
(b)L−Y0
1(θ,ϕ)=+radicalbigg
3
4πsinθe−iϕ=√
2Y−1
1(θ,ϕ).
The signs (Condon–Shortley phase) are a consequence of the ladder operators L+and
L−.
12.8 T HEADDITION THEOREM FOR SPHERICAL HARMONICS
Trigonometric Identity
In the following discussion, (θ1,ϕ1)and(θ2,ϕ2)denote two different directions in our
spherical coordinate system (x1,y1,z1), separated by an angle γ(Fig. 12.16). The polar
anglesθ1,θ2aremeasuredfromthe z1-axis.Theseanglessatisfythetrigonometricidentity
cosγ=cosθ1cosθ2+sinθ1sinθ2cos(ϕ1−ϕ2), (12.175)
whichisperhapsmosteasilyprovedbyvectormethods(compareChapter1).
Theadditiontheorem,then,asserts that
Pn(cosγ)=4π
2n+1nsummationdisplay
m=−n(−1)mYm
n(θ1,ϕ1)Y−m
n(θ2,ϕ2), (12.176)
orequivalently,
Pn(cosγ)=4π
2n+1nsummationdisplay
m=−nYm
n(θ1,ϕ1)bracketleftbig
Ym
n(θ2,ϕ2)bracketrightbig∗.21(12.177)
21Theasterisk for complex conjugation may go on either spherical harmonic.
798 Chapter 12 Legendre Functions
Intermsof theassociatedLegendrefunctions,theadditiontheoremis
Pn(cosγ)=Pn(cosθ1)Pn(cosθ2)
+2nsummationdisplay
m=1(n−m)!
(n+m)!Pm
n(cosθ1)Pm
n(cosθ2)cosm(ϕ1−ϕ2).
(12.178)
Equation(12.175)isaspecialcaseofEq. (12.178), n=1.
Derivation of Addition Theorem
WenowderiveEq.(12.177).Let (γ,ξ)betheanglesthatspecifythedirection (θ1,ϕ1)ina
coordinatesystem (x2,y2,z2)whoseaxisis alignedwith (θ2,ϕ2).(Actually,thechoiceof
the 0 azimuthangle ξinFig.12.16isirrelevant.)First,weexpand Ym
n(θ1,ϕ1)inspherical
harmonicsinthe (γ,ξ)angularvariables:
Ym
n(θ1,ϕ1)=nsummationdisplay
σ=−nam
nσYσ
n(γ,ξ). (12.179)
Wewritenosummationover ninEq.(12.179)becausetheangularmomentum nofYm
nis
conserved(seeSection4.3);asasphericalharmonic, Ym
n(θ1,ϕ1)isaneigenfunctionof L2
witheigenvalue n(n+1).
FIGURE 12.16Twodirectionsseparatedbyan
angleγ.
12.8 Addition Theorem for Spherical Harmonics 799
Weneedforourproofonlythecoefficient am
n0,whichwegetbymultiplyingEq.(12.179)
by[Y0
n(γ,ξ)]∗andintegratingoverthesphere:
am
n0=integraldisplay
Ym
n(θ1,ϕ1)bracketleftbig
Y0
n(γ,ξ)bracketrightbig∗d/Omega1γ,ξ. (12.180)
Similarly,weexpand Pn(cosγ)intermsofsphericalharmonics Ym
n(θ1,ϕ1):
Pn(cosγ)=parenleftbigg4π
2n+1parenrightbigg1/2
Y0
n(γ,ξ)=nsummationdisplay
m=−nbnmYm
n(θ1,ϕ1), (12.181)
where the bnmwill, of course, depend on θ2,ϕ2, that is, on the orientation of the z2-axis.
Multiplyingby [Ym
n(θ1,ϕ1)]∗andintegratingwithrespectto θ1andϕ1overthesphere,we
have
bnm=integraldisplay
Pn(cosγ)Ym∗
n(θ1,ϕ1)d/Omega1θ1,ϕ1. (12.182)
Intermsof sphericalharmonicsEq. (12.182)becomes
parenleftbigg4π
2n+1parenrightbigg1/2integraldisplay
Y0
n(γ,ψ)bracketleftbig
Ym
n(θ1,ϕ1)bracketrightbig∗d/Omega1=bnm. (12.183)
Note that the subscripts have been dropped from the solid angle element d/Omega1. Since the
range of integration is over all solid angles, the choice of polar axis is irrelevant. Then
comparingEqs. (12.180)and(12.183),weseethat
b∗
nm=am
n0parenleftbigg4π
2n+1parenrightbigg1/2
. (12.184)
Now we evaluate Ym
n(θ2,ϕ2)using the expansion of Eq. (12.179) and noting that the
valuesof (γ,ξ)correspondingto (θ1,ϕ1)=(θ2,ϕ2)are(0,0).Theresultis
Ym
n(θ2,ϕ2)=am
n0Y0
n(0,0)=am
n0parenleftbigg2n+1
4πparenrightbigg1/2
, (12.185)
alltermswithnonzero σvanishing.SubstitutingthisbackintoEq. (12.184),weobtain
bnm=4π
2n+1bracketleftbig
Ym
n(θ2,ϕ2)bracketrightbig∗. (12.186)
Finally, substituting this expression for bnminto the summation, Eq. (12.181) yields
Eq.(12.177), thusprovingouradditiontheorem.
Those familiar with group theory will find a much more elegant proof of Eq. (12.177)
byusingtherotationgroup.22Thisis Exercise4.4.5.
One application of the addition theorem is in the construction of a Green’s function for
the three-dimensional Laplace equation in spherical polar coordinates. If the source is on
22C o m p ar eM.E.R o s e, ElementaryTheory of Angular Momentum , NewYork: Wiley (1957).
800 Chapter 12 Legendre Functions
thepolaraxisatthepoint (r=a,θ=0,ϕ=0),then,byEq. (12.4a),
1
R=1
|r−ˆza|=∞summationdisplay
n=0Pn(cosγ)an
rn+1,r>a
=∞summationdisplay
n=0Pn(cosγ)rn
an+1,r<a. (12.187)
Rotatingourcoordinatesystemtoputthesourceat (a,θ2,ϕ2)andthepointofobservation
at(r,θ1,ϕ1), weobtain
G(r,θ1,ϕ1,a,θ2,ϕ2)=1
R
=∞summationdisplay
n=0nsummationdisplay
m=−n4π
2n+1bracketleftbig
Ym
n(θ1,ϕ1)bracketrightbig∗Ym
n(θ2,ϕ2)an
rn+1,r>a,
=∞summationdisplay
n=0nsummationdisplay
m=−n4π
2n+1bracketleftbig
Ym
n(θ1,ϕ1)bracketrightbig∗Ym
n(θ2,ϕ2)rn
an+1,r<a.(12.188)
In Section 9.7 this argument is reversed to provide another derivation of the Legendre
polynomialadditiontheorem.
Exercises
12.8.1 In proving the addition theorem, we assumed that Yk
n(θ1,ϕ1)could be expanded in a
series of Ym
n(θ2,ϕ2), in which mvaried from−nto+nbutnwas held fixed. What
arguments can you develop to justify summing only over the upper index, m, andnot
overthelowerindex, n?
Hints. One possibility is to examine the homogeneity of the Ym
n, that is, Ym
nmay be
expressed entirely in terms of the form cosn−pθsinpθ,o rxn−p−sypzs/rn. Another
possibility is to examine the behavior of the Legendre equation under rotation of the
coordinatesystem.
12.8.2 Anatomicelectronwithangularmomentum Landmagneticquantumnumber Mhasa
wavefunction
ψ(r,θ,ϕ)=f(r)YM
L(θ,ϕ).
Show that the sum of the electron densities in a given complete shell is spherically
symmetric;thatis,summationtextL
M=−Lψ∗(r,θ,ϕ)ψ(r,θ,ϕ) isindependentof θandϕ.
12.8.3 Thepotentialof anelectronatpoint reinthefieldof Zprotonsatpoints rpis
/Phi1=−e2
4πε0Zsummationdisplay
p=11
|re−rp|.
12.8 Addition Theorem for Spherical Harmonics 801
Showthatthismaybewrittenas
/Phi1=−e2
4πε0reZsummationdisplay
p=1summationdisplay
L,Mparenleftbiggrp
reparenrightbiggL4π
2L+1bracketleftbig
YM
L(θp,ϕp)bracketrightbig∗YM
L(θe,ϕe),
wherere>rp.Howshould /Phi1bewrittenfor re<rp?
12.8.4 Two protons are uniformly distributed within the same spherical volume. If the coor-
dinates of one element of charge are (r1,θ1,ϕ1)and the coordinates of the other are
(r2,θ2,ϕ2)andr12is the distance between them, the element of energy of repulsion
willbegivenby
dψ=ρ2dτ1dτ2
r12=ρ2r2
1dr1sinθ1dθ1dϕ1r2
2dr2sinθ2dθ2dϕ2
r12.
Here
ρ=charge
volume=3e
4πR3,chargedensity ,
r2
12=r2
1+r2
2−2r1r2cosγ.
Calculatethetotalelectrostaticenergy(ofrepulsion)ofthetwoprotons.Thiscalculation
isusedinaccountingfor themassdifferencein“mirror”nuclei,suchasO15andN15.
ANS.6
5e2
R.
This isdoublethat required to create a uniformly charged sphere because we have two
separatecloudchargesinteracting,notonechargeinteractingwithitself(withpermuta-
tionof pairs notconsidered).
12.8.5 Eachofthetwo1 Selectronsinheliummaybedescribedbyahydrogenicwavefunction
ψ(r)=parenleftbiggZ3
πa3
0parenrightbigg1/2
e−Zr/a0
in the absence of the other electron. Here Z, the atomic number, is 2. The symbol a0is
the Bohr radius, ¯h2/me2. Find the mutual potential energy of the two electrons, given
by
integraldisplay
ψ∗(r1)ψ∗(r2)e2
r12ψ(r1)ψ(r2)d3r1d3r2.
ANS.5e2Z
8a0.
Note.d3r1=r2dr1sinθ1dθ1dϕ1≡dτ1,r12=|r1−r2|.
12.8.6 Theprobabilityoffindinga1 Shydrogenelectroninavolumeelement r2drsinθdθdϕ
is
1
πa3
0exp[−2r/a0]r2drsinθdθdϕ.
802 Chapter 12 Legendre Functions
Findthecorrespondingelectrostaticpotential.Calculatethepotentialfrom
V(r1)=q
4πε0integraldisplayρ(r2)
r12d3r2,
withr1notonthez-axis.Expand r12.ApplytheLegendrepolynomialadditiontheorem
andshowthattheangulardependenceof V(r1)dropsout.
ANS.V(r1)=q
4πε0braceleftbigg1
2r1γparenleftbigg
3,2r1
a0parenrightbigg
+1
a0Ŵparenleftbigg
2,2r1
a0parenrightbiggbracerightbigg
.
12.8.7 Ahydrogenelectronina 2 Porbithas achargedistribution
ρ=q
64πa5
0r2e−r/a0sin2θ,
wherea0is the Bohr radius, ¯h2/me2. Find the electrostatic potential corresponding to
thischargedistribution.
12.8.8 Theelectriccurrentdensityproducedbya 2 Pelectroninahydrogenatomis
J=ˆϕq¯h
32ma5
0e−r/a0rsinθ.
Using
A(r1)=µ0
4πintegraldisplayJ(r2)
|r1−r2|d3r2,
findthemagneticvectorpotentialproducedbythishydrogenelectron.
Hint. Resolve into Cartesian components. Use the addition theorem to eliminate γ,t h e
angleincludedbetween r1andr2.
12.8.9 (a) As a Laplace series and as an example of Eq. (1.190) (now with complex func-
tions),showthat
δ(/Omega11−/Omega12)=∞summationdisplay
n=0nsummationdisplay
m=−nYm∗
n(θ2,ϕ2)Ym
n(θ1,ϕ1).
(b) Showalsothatthis sameDiracdeltafunctionmaybewrittenas
δ(/Omega11−/Omega12)=∞summationdisplay
n=02n+1
4πPn(cosγ).
Now, if you can justify equating the summations over nterm by term , you have an
alternatederivationofthesphericalharmonicadditiontheorem.
12.9 Integrals of Three Y’s 803
12.9 I NTEGRALS OF PRODUCTS OF THREE SPHERICAL
HARMONICS
Frequentlyinquantummechanicsweencounterintegralsofthegeneralform
angbracketleftbig
YM1
L1vextendsinglevextendsingleYM2
L2vextendsinglevextendsingleYM3
L3angbracketrightbig
=integraldisplay2π
0integraldisplayπ
0bracketleftbig
YM1
L1bracketrightbig∗YM2
L2YM3
L3sinθdθdϕ
=radicalBigg
(2L2+1)(2L3+1)
4π(2L1+1)C(L2L3L1|000)C(L2L3L1|M2M3M1),
(12.189)
inwhichallsphericalharmonicsdependon θ,ϕ.Thefirstfactorintheintegrandmaycome
fromthewavefunctionofafinalstateandthethirdfactorfromaninitialstate,whereasthe
middlefactormayrepresentanoperatorthatisbeingevaluatedorwhose“matrixelement”
isbeingdetermined.
By using group theoretical methods, as in the quantum theory of angular momentum,
we may give a general expression for the forms listed. The analysis involves the vector–
additionorClebsch–GordancoefficientsfromSection4.4,whicharetabulated.Threegen-
eralrestrictionsappear.
1. The integral vanishes unless the triangle condition of the L’s (angular momentum) is
zero,|L1−L3|≤L2≤L1+L3.
2. Theintegralvanishesunless M2+M3=M1.Herewehavethetheoreticalfoundation
of thevectormodelofatomicspectroscopy.
3. Finally,theintegralvanishesunlesstheproduct [YM1
L1]∗YM2
L2YM3
L3iseven,thatis,unless
L1+L2+L3isaneveninteger.This isaparityconservationlaw.
ThekeytothedeterminationoftheintegralinEq.(12.189)istheexpansionoftheprod-
uct of two spherical harmonics depending on the same angles (in contrast to the addition
theorem),whicharecoupledbyClebsch–Gordancoefficientstoangularmomentum L,M,
which, from its rotational transformation properties, must be proportional to YM
L(θ,ϕ);
thatis,summationdisplay
M1,M2C(L2L3L1|M2M3M1)YM2
L2(θ,ϕ)YM3
L3(θ,ϕ)∼YM1
L1(θ,ϕ).
Fordetailswerefer toEdmonds.23
LetusoutlinesomeofthestepsofthisgeneralandpowerfulapproachusingSection4.4.
TheWigner–EckarttheoremappliedtothematrixelementinEq. (12.189)yields
angbracketleftbig
YM1
L1vextendsinglevextendsingleYM2
L2vextendsinglevextendsingleYM3
L3angbracketrightbig
=(−1)L2−L3+L1C(L2L3L1|M2M3M1)
·/angbracketleftYL1/bardblYL2/bardblYL3/angbracketright√(2L1+1), (12.190)
23E. U. Condon and G. H. Shortley, The Theory of Atomic Spectra , Cambridge, UK: Cambridge University Press (1951); M.
E. Rose,Elementary Theory of Angular Momentum , New York: Wiley (1957); A. Edmonds, Angular Momentum in Quantum
Mechanics , Princeton, NJ: Princeton University Press (1957); E. P. Wigner, Group Theory and Its Applications to Quantum
Mechanics (translated by J. J. Griffin), NewYork: AcademicPress (1959).
804 Chapter 12 Legendre Functions
where the double bars denote the reduced matrix element, which no longer depends on
theMi. Selection rules (1) and (2) mentioned earlier follow directly from the Clebsch–
Gordan coefficient in Eq. (12.190). Next we use Eq. (12.190) for M1=M2=M3=0i n
conjunctionwithEq. (12.153)for m=0,whichyields
angbracketleftbig
Y0
L1vextendsinglevextendsingleY0
L2vextendsinglevextendsingleY0
L3angbracketrightbig
=(−1)L2−L3+L1
√2L1+1C(L2L3L1|000)·/angbracketleftYL1/bardblYL2/bardblYL3/angbracketright
=radicalbigg
(2L1+1)(2L2+1)(2L3+1)
4π
·1
2·integraldisplay1
−1PL1(x)PL2(x)PL3(x)dx, (12.191)
wherex=cosθ.Byelementarymethodsitcanbeshownthat
integraldisplay1
−1PL1(x)PL2(x)PL3(x)dx=2
2L1+1C(L2L3L1|000)2. (12.192)
SubstitutingEq. (12.192)into(12.191)weobtain
/angbracketleftYL1/bardblYL2/bardblYL3/angbracketright=(−1)L2−L3+L1C(L2L3L1|000)radicalbigg
(2L2+1)(2L3+1)
4π.(12.193)
The aforementioned parity selection rule (3) above follows from Eq. (12.193) in conjunc-
tionwiththephaserelation
C(L2L3L1|−M2,−M3,−M1)=(−1)L2+L3−L1C(L2L3L1|M2M3M1).(12.194)
Note that the vector-addition coefficients are developed in terms of the Condon–Shortley
phaseconvention,23inwhichthe (−1)mof Eq. (12.153)isassociatedwiththepositive m.
It is possibleto evaluatemanyof the commonlyencounteredintegrals of this form with
the techniques already developed. The integration over azimuth may be carried out by
inspection:
integraldisplay2π
0e−iM1ϕeiM2ϕeiM3ϕdϕ=2πδM2+M3−M1,0. (12.195)
Physicallythiscorrespondstotheconservationofthe zcomponentofangularmomentum.
Application of Recurrence Relations
A glance at Table 12.3 will show that the θ-dependence of YM2
L2, that is,PM2
L2(θ), can be
expressedintermsof cos θand sinθ.However,afactorof cos θor sinθmaybecombined
withtheYM3
L3factorbyusingtheassociatedLegendrepolynomialrecurrencerelations.For
12.9 Integrals of Three Y’s 805
instance,fromEqs. (12.92) and(12.93) weget
cosθYM
L=+bracketleftbigg(L−M+1)(L+M+1)
(2L+1)(2L+3)bracketrightbigg1/2
YM
L+1
+bracketleftbigg(L−M)(L+M)
(2L−1)(2L+1)bracketrightbigg1/2
YM
L−1 (12.196)
eiϕsinθYM
L=−bracketleftbigg(L+M+1)(L+M+2)
(2L+1)(2L+3)bracketrightbigg1/2
YM+1
L+1
+bracketleftbigg(L−M)(L−M−1)
(2L−1)(2L+1)bracketrightbigg1/2
YM+1
L−1(12.197)
e−iϕsinθYM
L=+bracketleftbigg(L−M+1)(L−M+2)
(2L+1)(2L+3)bracketrightbigg1/2
YM−1
L+1
−bracketleftbigg(L+M)(L+M−1)
(2L−1)(2L+1)bracketrightbigg1/2
YM−1
L−1. (12.198)
Usingtheseequations,weobtain
integraldisplay
YM1∗
L1cosθYM
Ld/Omega1=bracketleftbigg(L−M+1)(L+M+1)
(2L+1)(2L+3)bracketrightbigg1/2
δM1,MδL1,L+1
+bracketleftbigg(L−M)(L+M)
(2L−1)(2L+1)bracketrightbigg1/2
δM1,MδL1,L−1.(12.199)
The occurrence of the Kronecker delta (L1,L±1)is an aspect of the conservation of
angular momentum. Physically, this integral arises in a consideration of ordinary atomic
electromagneticradiation(electricdipole).Itleadstothefamiliarselectionrulethattransi-
tionstoanatomiclevelwithorbitalangularmomentumquantumnumber L1canoriginate
only from atomic levels with quantum numbers L1−1o rL1+1. The application to ex-
pressionssuchas
quadrupolemoment ∼integraldisplay
YM∗
L(θ,ϕ)P 2(cosθ)YM
L(θ,ϕ)d/Omega1
ismoreinvolvedbutperfectlystraightforward.
Exercises
12.9.1 Verify
(a)integraldisplay
YM
L(θ,ϕ)Y0
0(θ,ϕ)YM∗
L(θ,ϕ)d/Omega1=1√
4π,
(b)integraldisplay
YM
LY0
1YM∗
L+1d/Omega1=radicalbigg
3
4πradicalBigg
(L+M+1)(L−M+1)
(2L+1)(2L+3),
806 Chapter 12 Legendre Functions
(c)integraldisplay
YM
LY1
1YM+1∗
L+1d/Omega1=radicalbigg
3
8πradicalBigg
(L+M+1)(L+M+2)
(2L+1)(2L+3),
(d)integraldisplay
YM
LY1
1YM+1∗
L−1d/Omega1=−radicalbigg
3
8πradicalBigg
(L−M)(L−M−1)
(2L−1)(2L+1).
These integrals were used in an investigation of the angular correlation of internal con-
versionelectrons.
12.9.2 Showthat
(a)integraldisplay1
−1xPL(x)PN(x)dx=
2(L+1)
(2L+1)(2L+3),N=L+1,
2L
(2L−1)(2L+1),N=L−1,
(b)integraldisplay1
−1x2PL(x)PN(x)dx=
2(L+1)(L+2)
(2L+1)(2L+3)(2L+5),N=L+2,
2(2L2+2L−1)
(2L−1)(2L+1)(2L+3),N=L,
2L(L−1)
(2L−3)(2L−1)(2L+1),N=L−2.
12.9.3 SincexPn(x)is a polynomial (degree n+1), it may be represented by the Legendre
series
xPn(x)=∞summationdisplay
s=0asPs(x).
(a) Showthat as=0f o rs<n−1 ands>n+1.
(b) Calculate an−1,an, andan+1and show that you have reproduced the recurrence
relation,Eq. 12.17.
Note. This argument may be put in a general form to demonstrate the existence of a
three-termrecurrencerelationfor anyof ourcompletesets oforthogonalpolynomials:
xϕn=an+1ϕn+1+anϕn+an−1ϕn−1.
12.9.4 Show that Eq. (12.199) is a special case of Eq. (12.190) and derive the reduced matrix
element/angbracketleftYL1/bardblY1/bardblYL/angbracketright.
ANS./angbracketleftYL1/bardblY1/bardblYL/angbracketright=(−1)L1+1−LC(1LL1|000)√3(2L+1)
4π.
12.10 L EGENDRE FUNCTIONS OF THE SECOND KIND
In all the analysis so far in this chapter we have been dealing with one solution of Legen-
dre’sequation,thesolution Pn(cosθ),whichisregular(finite)atthetwosingularpointsof
12.10 Legendre Functions of the Second Kind 807
the differential equation, cos θ=±1. From the general theory of differential equations it
isknownthatasecondsolutionexists.Wedevelopthissecondsolution, Qn,withnonneg-
ative integer n(because Qnin applications will occur in conjunction with Pn) ,b yas e r i e s
solutionofLegendre’sequation.Later aclosedform willbeobtained.
Series Solutions of Legendre’s Equation
To solve
d
dxbracketleftbigg
(1−x2)dy
dxbracketrightbigg
+n(n+1)y=0 (12.200)
weproceedasinChapter9,letting24
y=∞summationdisplay
λ=0aλxk+λ, (12.201)
with
y′=∞summationdisplay
λ=0(k+λ)aλxk+λ−1, (12.202)
y′′=∞summationdisplay
λ=0(k+λ)(k+λ−1)aλxk+λ−2. (12.203)
Substitutionintotheoriginaldifferentialequationgives
∞summationdisplay
λ=0(k+λ)(k+λ−1)aλxk+λ−2
+∞summationdisplay
λ=0bracketleftbig
n(n+1)−2(k+λ)−(k+λ)(k+λ−1)bracketrightbig
aλxk+λ=0.(12.204)
Theindicialequation is
k(k−1)=0, (12.205)
withsolutions k=0,1.Wetryfirst k=0witha0=1,a1=0.Thenourseriesisdescribed
bytherecurrencerelation
(λ+2)(λ+1)aλ+2+bracketleftbig
n(n+1)−2λ−λ(λ−1)bracketrightbig
aλ=0, (12.206)
whichbecomes
aλ+2=−(n+λ+1)(n−λ)
(λ+1)(λ+2)aλ. (12.207)
24Notethat xmaybe replacedby the complex variable z.
808 Chapter 12 Legendre Functions
Labelingthisseries, from Eq. (12.201), y(x)=pn(x),weha v e
pn(x)=1−n(n+1)
2!x2+(n−2)n(n+1)(n+3)
4!x4+···. (12.208)
The second solution of the indicial equation, k=1, witha0=0,a1=1, leads to the
recurrencerelation
aλ+2=−(n+λ+2)(n−λ−1)
(λ+2)(λ+3)aλ. (12.209)
Labelingthisseries, from Eq. (12.201), y(x)=qn(x), weobtain
qn(x)=x−(n−1)(n+2)
3!x3+(n−3)(n−1)(n+2)(n+4)
5!x5−···.(12.210)
Ourgeneralsolutionof Eq.(12.200), then,is
yn(x)=Anpn(x)+Bnqn(x), (12.211)
providedwehaveconvergence .FromGauss’test,Section5.2(seeExample5.2.4),wedo
nothaveconvergenceat x=±1.Togetoutofthisdifficulty,wesettheseparationconstant
nequaltoaninteger(Exercise9.5.5) andconverttheinfiniteseries intoapolynomial.
Forna positive even integer (or zero), series pnterminates, and with a proper choice
of a normalizing factor (selected to obtain agreement with the definition of Pn(x)in Sec-
tion12.1)
Pn(x)=(−1)n/2n!
2n[(n/2)!]2pn(x)=(−1)s(2s)!
22s(s!)2p2s(x)
=(−1)s(2s−1)!!
(2s)!!p2s(x), forn=2s. (12.212)
Ifnis a positive odd integer, series qnterminates after a finite number of terms, and we
write
Pn(x)=(−1)n−1)/2 n!
2n−1{[n−1)/2]!}2qn(x)
=(−1)s(2s+1)!
22s(s!)2q2s+1(x)=(−1)s(2s+1)!!
(2s)!!q2s+1(x), forn=2s+1.
(12.213)
Note that these expressions hold for all real values of x,−∞<x<∞, and for complex
values in the finite complex plane. The constants that multiply pnandqnare chosen to
makePnagreewithLegendrepolynomialsgivenbythegeneratingfunction.
Equations (12.208) and (12.210) may still be used with n=ν, not an integer, but now
the series no longer terminates, and the range of convergence becomes −1<x<1. The
endpoints, x=±1,are notincluded.
It is sometimes convenient to reverse the order of the terms in the series. This may be
donebyputting
s=n
2−λ inthefirstform of Pn(x), n even,
s=n−1
2−λinthesecondformof Pn(x), n odd,
12.10 Legendre Functions of the Second Kind 809
sothatEqs. (12.212)and(12.213)become
Pn(x)=[n/2]summationdisplay
s=0(−1)s(2n−2s)!
2ns!(n−s)!(n−2s)!xn−2s, (12.214)
where the upper limit s=n/2 (forneven) or (n−1)/2( f o rnodd). This reproduces
Eq. (12.8) of Section 12.1, which is obtained directly from the generating function. This
agreement with Eq. (12.8) is the reason for the particular choice of normalization in
Eqs. (12.212)and(12.213).
Qn(x)Functions of the Second Kind
Itwillbenoticedthatwehaveusedonly pnfornevenand qnfornodd(becausetheyter-
minatedforthischoiceof n).WemaynowdefineasecondsolutionofLegendre’sequation
(Fig.12.17)by
Qn(x)=(−1)n/2[n/2]!22n
n!qn(x)
=(−1)s(2s)!!
(2s−1)!!q2s(x), forneven,n=2s,(12.215)
FIGURE 12.17SecondLegendrefunction, Qn(x),
0≤x<1.
810 Chapter 12 Legendre Functions
FIGURE 12.18Second
Legendrefunction, Qn(x),
x>1.
Qn(x)=(−1)(n+1)/2{[(n−1)/2]!}22n−1
n!pn(x)
=(−1)s+1(2s)!!
(2s+1)!!p2s+1(x), fornodd,n=2s+1.(12.216)
This choice of normalizing factors forces Qnto satisfy the same recurrence relations as
Pn. This may be verified by substituting Eqs. (12.215) and (12.216) into Eqs. (12.17) and
(12.26).Inspectionofthe(series)recurrencerelations(Eqs.(12.207)and(12.209)),thatis,
by theCauchy ratio test, shows that Qn(x)willconvergefor −1<x<1.If|x|≥1, these
series forms of our second solution diverge. A solution in a series of negative powers of
xcan be developed for the region |x|>1 (Fig. 12.18), but we proceed to a closed-form
solution that can be used over the entire complex plane (apart from the singular points
x=±1 andwithcareoncutlines).
Closed-Form Solutions
Frequently,aclosedformofthesecondsolution, Qn(z),isdesirable.Thismaybeobtained
bythemethoddiscussedinSection9.6.We write
Qn(z)=Pn(z)braceleftbigg
An+Bnintegraldisplayzdx
(1−x2)[Pn(x)]2bracerightbigg
, (12.217)
inwhichtheconstant Anreplacestheevaluationoftheintegralatthearbitrarylowerlimit.
Bothconstants, AnandBn,maybedeterminedfor specialcases.
12.10 Legendre Functions of the Second Kind 811
Forn=0,Eq.(12.217)yields
Q0(z)=P0(z)braceleftbigg
A0+B0integraldisplayzdx
(1−x2)[P0(x)]2bracerightbigg
=A0+B01
2ln1+z
1−z
=A0+B0parenleftbigg
z+z3
3+z5
5+···+z2s+1
2s+1+···parenrightbigg
, (12.218)
thelastexpressionfollowingfromaMaclaurinexpansionofthelogarithm.Comparingthis
withtheseriessolution(Eq. (12.210)),weobtain
Q0(z)=q0(z)=z+z3
3+z5
5+···+z2s+1
2s+1+···, (12.219)
wehaveA0=0,B0=1.Similarresultsfollowfor n=1.We obtain
Q1(z)=zbracketleftbigg
A1+B1integraldisplayzdx
(1−x2)x2bracketrightbigg
=A1z+B1zparenleftbigg1
2ln1+z
1−z−1
zparenrightbigg
. (12.220)
Expandingin a powerseries and comparingwith Q1(z)=−p1(z),w eh a v e A1=0,B1=
1.Thereforewemaywrite
Q0(z)=1
2ln1+z
1−z,Q 1(z)=1
2zln1+z
1−z−1,|z|<1.(12.221)
Perhaps the best way of determining the higher-order Qn(z)is to use the recurrence
relation(Eq.(12.17)),whichmaybeverifiedforboth x2<1andx2>1bysubstitutingin
theseries forms. Thisrecurrencerelationtechniqueyields
Q2(z)=1
2P2(z)ln1+z
1−z−3
2P1(z). (12.222)
Repeatedapplicationoftherecurrenceformulaleadsto
Qn(z)=1
2Pn(z)ln1+z
1−z−2n−1
1·nPn−1(z)−2n−5
3(n−1)Pn−3(z)−···.(12.223)
From the form ln [(1+z)/(1−z)]it will be seen that for real zthese expressions hold in
the range−1<x<1. If we wish to have closed forms valid outside this range, we need
onlyreplace
ln1+x
1−xby lnz+1
z−1.
When using the latter form, valid for large z, we take the line interval −1≤x≤1 as a cut
line.Valuesof Qn(x)onthecutlineare customarilyassignedbytherelation
Qn(x)=1
2bracketleftbig
Qn(x+i0)+Qn(x−i0)bracketrightbig
, (12.224)
812 Chapter 12 Legendre Functions
the arithmetic average of approaches from the positive imaginary side and from the nega-
tive imaginary side. Note that for z→x>1,z−1→(1−x)e±iπ. The result is that for
allz, exceptontherealaxis, −1≤x≤1,wehave
Q0(z)=1
2lnz+1
z−1, (12.225)
Q1(z)=1
2zlnz+1
z−1−1, (12.226)
andso on.
Forconvenientreferencesomespecialvaluesof Qn(z)aregiven.
1.Qn(1)=∞,from thelogarithmicterm(Eq. (12.223)).
2.Qn(∞)=0. This is best obtained from a representation of Qn(x)as a series of nega-
tivepowersof x,Exercise12.10.4.
3.Qn(−z)=(−1)n+1Qn(z). This follows from the series form. It may also be derived
byusingQ0(z),Q1(z)andtherecurrencerelation(Eq. (12.17)).
4.Qn(0)=0,forneven,by(3).
5.Qn(0)=(−1)(n+1)/2{[(n−1)/2]!}2
n!2n−1
=(−1)s+1(2s)!!
(2s+1)!!,fornodd,n=2s+1.
Thislastresultcomesfrom theseries form(Eq. (12.216))with pn(0)=1.
Exercises
12.10.1 Derivetheparityrelationfor Qn(x).
12.10.2 FromEqs. (12.212)and(12.213)showthat
(a)P2n(x)=(−1)n
22n−1nsummationdisplay
s=0(−1)s(2n+2s−1)!
(2s)!(n+s−1)!(n−s)!x2s.
(b)P2n+1(x)=(−1)n
22nnsummationdisplay
s=0(−1)s(2n+2s+1)!
(2s+1)!(n+s)!(n−s)!x2s+1.
Checkthenormalizationbyshowingthatonetermofeachseriesagreeswiththecorre-
spondingtermofEq. (12.8).
12.10.3 Showthat
(a)Q2n(x)=(−1)n22nnsummationdisplay
s=0(−1)s(n+s)!(n−s)!
(2s+1)!(2n−2s)!x2s+1
+22n∞summationdisplay
s=n+1(n+s)!(2s−2n)!
(2s+1)!(s−n)!x2s+1,|x|<1.
12.11 Vector Spherical Harmonics 813
(b)Q2n+1(x)=(−1)n+122nnsummationdisplay
s=0(−1)s(n+s)!(n−s)!
(2s)!(2n−2s+1)!x2s
+22n+1∞summationdisplay
s=n+1(n+s)!(2s−2n−2)!
(2s)!(s−n−1)!x2s,|x|<1.
12.10.4 (a) Startingwiththeassumedform
Qn(x)=∞summationdisplay
λ=0b−λxk−λ,
showthat
Qn(x)=b0x−n−1∞summationdisplay
s=0(n+s)!(n+2s)!(2n+1)!
s!(n!)2(2n+2s+1)!x−2s.
(b) Thestandardchoiceof b0is
b0=2n(n!)2
(2n+1)!.
Show that this choice of bobrings this negative power-series form of Qn(x)into
agreementwiththeclosed-formsolutions.
12.10.5 Verify that the Legendre functions of the second kind, Qn(x), satisfy the same recur-
rencerelationsas Pn(x), bothfor|x|<1 andfor|x|>1:
(2n+1)xQn(x)=(n+1)Qn+1(x)+nQn−1(x),
(2n+1)Qn(x)=Q′
n+1(x)−Q′
n−1(x).
12.10.6 (a) Usingtherecurrencerelations,prove(independentoftheWronskianrelation)that
nbracketleftbig
Pn(x)Qn−1(x)−Pn−1(x)Qn(x)bracketrightbig
=P1(x)Q0(x)−P0(x)Q1(x).
(b) Bydirectsubstitutionshowthattheright-handsideofthis equationequals1.
12.10.7 (a) Write a subroutine that will generate Qn(x)andQ0throughQn−1based on the
recurrence relation for these Legendre functions of the second kind. Take xto be
within(−1,1)—excludingtheendpoints.
Hint.T ak eQ0(x)andQ1(x)tobeknown.
(b) Test your subroutine for accuracy by computing Q10(x)and comparing with the
valuestabulatedinAMS-55(foracompletereference,seeAdditionalReadingsin
Chapter8).
12.11 V ECTOR SPHERICAL HARMONICS
Most of our attention in this chapter has been directed toward solving the equations of
scalarfields,suchastheelectrostaticpotential.Thiswasdoneprimarilybecausethescalar
fields are easier to handle than vector fields. However, with scalar field problems under
firmcontrol,moreandmoreattentionis beingpaidtovectorfieldproblems.
814 Chapter 12 Legendre Functions
Maxwell’s equations for the vacuum, where the external current and charge densities
vanish,leadtothewave(orvectorHelmholtz)equationforthevectorpotential A.Inapar-
tialwaveexpansionof Ainsphericalpolarcoordinateswewanttouseangulareigenfunc-
tions that are vectors. To this end we write the coordinate unit vectors ˆx,ˆy,ˆzin spherical
notation(see Section4.4),
ˆe+1=−ˆx+iˆy√
2,ˆe0=ˆz,ˆe−1=ˆx−iˆy√
2, (12.227)
sothatˆemformasphericaltensorofrank1.Ifwecouplethesphericalharmonicswiththe
ˆemto total angular momentum Jusing the relevant Clebsch–Gordan coefficients, we are
ledtothevectorsphericalharmonics:
YJLMJ(θ,ϕ)=summationdisplay
m,MC(L1J|MmMJ)YM
L(θ,ϕ)ˆem. (12.228)
Itis obviousthattheyobeytheorthogonalityrelations
integraldisplay
Y∗
JLMJ(θ,ϕ)·YJ′L′M′
J(θ,ϕ)d/Omega1=δJJ′δLL′δMJM′
M. (12.229)
GivenJ,theselectionrulesofangularmomentumcouplingtellusthat Lcanonlytakeon
the values J+1,J, andJ−1. If we look up the Clebsch–Gordan coefficients and invert
Eq.(12.228)weget
ˆrYM
L(θ,ϕ)=−bracketleftbiggL+1
2L+1bracketrightbigg1/2
YLL+1M+bracketleftbiggL
2L+1bracketrightbigg1/2
YLL−1M, (12.230)
displayingthevectorcharacterofthe Yandtheorbitalangularmomentumcontents, L+1
andL−1,ofˆrYM
L.
Undertheparityoperations(coordinateinversion)thevectorsphericalharmonicstrans-
form as
YLL+1M(θ′,ϕ′)=(−1)L+1YLL+1M(θ,ϕ),
YLL−1M(θ′,ϕ′)=(−1)L+1YLL−1M(θ,ϕ), (12.231)
YLLM(θ′,ϕ′)=(−1)LYLLM(θ,ϕ),
where
θ′=π−θϕ′=π+ϕ. (12.232)
The vector spherical harmonics are useful in a further development of the gradient
(Eq. (2.46)), divergence (Eq. (2.47)) and curl (Eq. (2.49)) operators in spherical polar co-
12.11 Vector Spherical Harmonics 815
ordinates:
∇bracketleftbig
F(r)YM
L(θ,ϕ)bracketrightbig
=−bracketleftbiggL+1
2L+1bracketrightbigg1/2bracketleftbiggd
dr−L
rbracketrightbigg
FYLL+1M
+bracketleftbiggL
2L+1bracketrightbigg1/2bracketleftbiggd
dr+L+1
rbracketrightbigg
FYLL−1M,(12.233)
∇·bracketleftbig
F(r)YLL+1M(θ,ϕ)bracketrightbig
=−parenleftbiggL+1
2L+1parenrightbigg1/2bracketleftbiggdF
dr+L+2
rFbracketrightbigg
YM
L(θ,ϕ), (12.234)
∇·bracketleftbig
F(r)YLL−1M(θ,ϕ)bracketrightbig
=parenleftbiggL
2L+1parenrightbigg1/2bracketleftbiggdF
dr−L−1
rFbracketrightbigg
YM
L(θ,ϕ), (12.235)
∇·bracketleftbig
F(r)YLLM(θ,ϕ)bracketrightbig
=0, (12.236)
∇×bracketleftbig
F(r)YLL+1Mbracketrightbig
=ibracketleftbiggL
2L+1bracketrightbigg1/2bracketleftbiggdF
dr+L+2
rFbracketrightbigg
YLLM,(12.237)
∇×bracketleftbig
F(r)YLLMbracketrightbig
=iparenleftbiggL
2L+1parenrightbigg1/2bracketleftbiggdF
dr−L
rFbracketrightbigg
YLL+1M
+iparenleftbiggL+1
2L+1parenrightbigg1/2bracketleftbiggdF
dr+L+1
rFbracketrightbigg
YLL−1M,(12.238)
∇×bracketleftbig
F(r)YLL−1Mbracketrightbig
=ibracketleftbiggL+1
2L+1bracketrightbigg1/2bracketleftbiggdF
dr−L−1
rFbracketrightbigg
YLLM.(12.239)
If we substitute Eq. (12.230) into the radial component ˆr∂/∂rof the gradient operator,
for example, we obtain both dF/drterms in Eq. (12.233). For a complete derivation of
Eqs. (12.233) to (12.239) we refer to the literature.25These relations play an important
roleinthepartialwaveexpansionofclassicalandquantumelectrodynamics.
Thedefinitionsofthevectorsphericalharmonicsgivenherearedictatedbyconvenience,
primarily in quantum mechanical calculations, in which the angular momentum is a sig-
nificant parameter. Further examples of the usefulness and power of the vector spherical
harmonics will be found in Blatt and Weisskopf,25in Morse and Feshbach (see General
Referencesbook’send),andinJackson’s ClassicalElectrodynamics ,3rded.,NewYork:J.
Wiley & Sons (1998), which use vector spherical harmonics in a description of multipole
radiationandrelatedelectromagneticproblems.
•Vector spherical harmonics are developed from coupling Lunits of orbital angular
momentum and 1 unit of spin angular momentum. An extension, coupling Lunits
of orbital angular momentum and 2 units of spin angular momentum to form tensor
sphericalharmonics,ispresentedbyMathews.26
25E. H. Hill, Theory of vector spherical harmonics, A m .J .P h y s . 22: 211 (1954); also J. M. Blatt and V. Weisskopf, Theoret-
ical Nuclear Physics , New York: Wiley (1952). Note that Hill assigns phases in accordance with the Condon–Shortley phase
convention (Section 4.4). In Hill’s notation XLM=YLLM,VLM=YLL+1M,WLM=YLL−1M.
26J. Mathews, Gravitational multipole radiation, in In Memoriam (H.P. Robertson, ed.), Philadelphia: Society for Industrial and
AppliedMathematics(1963).
816 Chapter 12 Legendre Functions
•The major application of tensor spherical harmonics is in the investigation of gravita-
tionalradiation.
Exercises
12.11.1 Constructthe l=0,m=0 andl=1,m=0 vectorsphericalharmonics.
ANS.Y010=−ˆr(4π)−1/2
Y000=0
Y120=−ˆr(2π)−1/2cosθ−ˆθ(8π)−1/2sinθ
Y110=ˆϕi(3/8π)1/2sinθ
Y100=ˆr(4π)−1/2cosθ−ˆθ(4π)−1/2sinθ.
12.11.2 Verify that the parity of YLL+1Mis(−1)L+1, that of YLLMis(−1)L, and that of
YLL−1Mis(−1)L+1.What happenedtothe M-dependenceoftheparity?
Hint.ˆrandˆϕhaveoddparity; ˆθhasevenparity(compareExercise2.5.8).
12.11.3 Verifytheorthonormalityof thevectorsphericalharmonics YJLMJ.
12.11.4 InJackson’s ClassicalElectrodynamics ,3rded.,(seeAdditionalReadingsofChapter11
for thereference)defines YLLMbytheequation
YLLM(θ,ϕ)=1√L(L+1)LYM
L(θ,ϕ),
inwhichtheangularmomentumoperator Lis givenby
L=−i(r×∇).
ShowthatthisdefinitionagreeswithEq.(12.228).
12.11.5 Showthat
Lsummationdisplay
M=−LY∗
LLM(θ,ϕ)·YLLM(θ,ϕ)=2L+1
4π.
Hint. One way is to use Exercise 12.11.4 with Lexpanded in Cartesian coordinates
usingtheraisingandloweringoperatorsofSection4.3.
12.11.6 Showthatintegraldisplay
YLLM·(ˆr×YLLM)d/Omega1=0.
The integrand represents an interference term in electromagnetic radiation that con-
tributestoangulardistributionsbutnottototalintensity.
AdditionalReadings
Hobson, E. W., The Theory of Spherical and Ellipsoidal Harmonics . New York: Chelsea (1955). This is a very
complete reference and theclassictext on Legendre polynomials and allrelatedfunctions.
Smythe, W. R., Static and DynamicElectricity ,3rd ed.NewYork: McGraw-Hill(1989).
Seealsothe referenceslisted in Sections 4.4 and12.9 andatthe endofChapter 13.
CHAPTER 13
MORESPECIAL FUNCTIONS
Inthischapterweshallstudyfoursetsoforthogonalpolynomials,Hermite,Laguerre,and
Chebyshev1of first and second kinds. Although these four sets are of less importance in
mathematical physics than are the Bessel and Legendre functions of Chapters 11 and 12,
they are used and therefore deserve attention. For example, Hermite polynomials occur in
solutionsofthesimpleharmonicoscillatorofquantummechanicsandLaguerrepolynomi-
als in wave functions of the hydrogen atom. Because the general mathematical techniques
duplicate those of the preceding two chapters, the development of these functions is only
outlined. Detailed proofs, along the lines of Chapters 11 and 12, are left to the reader. We
express these polynomials and other functions in terms of hypergeometric and confluent
hypergeometric functions. To conclude the chapter, we give an introduction to Mathieu
functions,whichariseassolutionsofODEsandPDEswithellipticalboundaryconditions.
13.1 H ERMITE FUNCTIONS
Generating Functions — Hermite Polynomials
TheHermitepolynomials(Fig. 13.1), Hn(x), maybedefinedbythegeneratingfunction2
g(x,t)=e−t2+2tx=∞summationdisplay
n=0Hn(x)tn
n!. (13.1)
1This is the spelling choice of AMS-55 (for the complete reference see footnote 4 in Chapter 5). However, a variety of names,
suchas Tschebyscheff, is encountered.
2Aderivation of this Hermite-generating function is outlined in Exercise 13.1.1.
817
818 Chapter 13 More Special Functions
FIGURE 13.1Hermite
polynomials.
Recurrence Relations
Notetheabsenceofasuperscript,whichdistinguishesHermitepolynomialsfromtheunre-
latedHankelfunctions.FromthegeneratingfunctionwefindthattheHermitepolynomials
satisfytherecurrencerelations
Hn+1(x)=2xHn(x)−2nHn−1(x) (13.2)
and
H′
n(x)=2nHn−1(x). (13.3)
Equation(13.2)is obtainedbydifferentiatingthegeneratingfunctionwithrespectto t:
∂g
∂t=(−2t+2x)e−t2+2tx=∞summationdisplay
n=0Hn+1(x)tn
n!
=−2∞summationdisplay
n=0Hn(x)tn+1
n!+2x∞summationdisplay
n=0Hn(x)tn
n!,
whichcanberewrittenas
∞summationdisplay
n=0tn
n!bracketleftbig
Hn+1(x)−2xHn(x)+2nHn−1(x)bracketrightbig
=0.
Becauseeachcoefficientofthispowerseriesvanishes,Eq.(13.2)isestablished.Similarly,
differentiationwithrespectto xleadsto
∂g
∂x=2te−t2+2tx=∞summationdisplay
n=0H′
n(x)tn
n!=2∞summationdisplay
n=0Hn(x)tn+1
n!,
whichyieldsEq. (13.3) uponshiftingthesummationindex ninthelastsum n+1→n.
13.1 Hermite Functions 819
Table 13.1 HermitePolynomials
H0(x)=1
H1(x)=2x
H2(x)=4x2−2
H3(x)=8x3−12x
H4(x)=16x4−48x2+12
H5(x)=32x5−160x3+120x
H6(x)=64x6−480x4+720x2−120
TheMaclaurinexpansionofthegeneratingfunction
e−t2+2tx=∞summationdisplay
n=0(2tx−t2)n
n!=1+parenleftbig
2tx−t2parenrightbig
+··· (13.4)
givesH0(x)=1 andH1(x)=2x, and then the recursion Eq. (13.2) permits the construc-
tion of any Hn(x)desired (integral n). For convenient reference the first several Hermite
polynomialsarelistedinTable13.1.
Special values of the Hermite polynomials follow from the generating function for
x=0:
e−t2=∞summationdisplay
n=0(−t2)n
n!=∞summationdisplay
n=0Hn(0)tn
n!,
thatis,
H2n(0)=(−1)n(2n)!
n!,H 2n+1(0)=0,n=0,1,.... (13.5)
Wealsoobtainfromthegeneratingfunctiontheimportantparityrelation
Hn(x)=(−1)nHn(−x) (13.6)
bynotingthatEq. (13.1)yields
g(−x,−t)=∞summationdisplay
n=0Hn(−x)(−t)n
n!=g(x,t)=∞summationdisplay
n=0Hn(x)tn
n!.
Alternate Representations
TheRodriguesrepresentationof Hn(x)is
Hn(x)=(−1)nex2dn
dxne−x2. (13.7)
Letusshowthisusingmathematicalinductionasfollows.
820 Chapter 13 More Special Functions
Example 13.1.1 RODRIGUES REPRESENTATION
Werewritethegeneratingfunctionas g(x,t)=ex2e−(t−x)2andnotethat
∂
∂te−(t−x)2=−∂
∂xe−(t−x)2.
Thisyields
∂g
∂tvextendsinglevextendsinglevextendsinglevextendsingle
t=0=(2x−2t)gvextendsinglevextendsinglevextendsingle
t=0=2x=H1(x)=−ex2d
dxe−x2,
whichistheinitial n=1 case.Assumingthecase nofEq.(13.7)asvalid,wenowusethe
operatoridentityd
dxex2=2xex2+ex2d
dxin
(−1)n+1ex2dn+1
dxn+1e−x2=(−1)n+1bracketleftbiggd
dxex2−2xex2bracketrightbiggdn
dxne−x2
=−d
dxHn(x)+2xHn(x)=Hn+1(x)
toestablishthe n+1 case,withthelastequalityfollowingfromEqs.(13.2) and(13.3).
More directly, differentiation of the generating function ntimes with respect to tand
thensetting tequaltozeroyields
Hn(x)=∂n
∂tnparenleftbig
e−t2+2txparenrightbigvextendsinglevextendsinglevextendsingle
t=0=(−1)nex2∂n
∂xne−(t−x)2vextendsinglevextendsinglevextendsingle
t=0=(−1)nex2dn
dxne−x2.
/squaresolid
Asecondrepresentationmaybeobtainedbyusingthecalculusofresidues(Section7.1).
IfwemultiplyEq.(13.1)by t−m−1andintegratearoundtheorigininthecomplex t-plane,
onlythetermwith Hm(x)willsurvive:
Hm(x)=m!
2πicontintegraldisplay
t−m−1e−t2+2txdt. (13.8)
Also, from the Maclaurin expansion, Eq. (13.4), we can derive our Hermite polynomial
Hn(x)inseriesform:Usingthebinomialexpansionof (2x−t)νandtheindex N=s+ν,
e−t2+2tx=∞summationdisplay
ν=0tν
ν!(2x−t)ν=∞summationdisplay
ν=0tν
ν!νsummationdisplay
s=0parenleftbiggν
sparenrightbigg
(2x)ν−s(−t)s
=∞summationdisplay
N=0tN
N![N/2]summationdisplay
s=0(2x)N−2s(−1)sN!
(N−s)!parenleftbiggN−s
sparenrightbigg
,
where[N/2]is the largest integer less than or equal to N/2. Writing the binomial coeffi-
cientintermsof factorialsandusingEq. (13.1)weobtain
HN(x)=[N/2]summationdisplay
s=0(2x)N−2s(−1)sN!
s!(N−2s)!.
13.1 Hermite Functions 821
Moreexplicitly,replacing N→n,weha v e
Hn(x)=(2x)n−2n!
(n−2)!2!(2x)n−2+4n!
(n−4)!4!(2x)n−41·3···
=[n/2]summationdisplay
s=0(−2)s(2x)n−2sparenleftbiggn
2sparenrightbigg
1·3·5···(2s−1)
=[n/2]summationdisplay
s=0(−1)s(2x)n−2sn!
(n−2s)!s!. (13.9)
Thisseries terminatesfor integral nandyieldsourHermitepolynomial.
Orthogonality
If we substitute the recursion Eq. (13.3) into Eq. (13.2) we can eliminate the index n−1,
obtaining
Hn+1(x)=2xHn(x)−H′
n(x),
which was used already in Example 13.1.1. If we differentiate this recursion relation and
substituteEq. (13.3)for theindex n+1 wefind
H′
n+1(x)=2(n+1)Hn(x)=2Hn(x)+2xH′
n(x)−H′′
n(x),
which can be rearranged to the second-order ODE for Hermite polynomials. Thus, the
recurrencerelations(Eqs. (13.2) and(13.3)) leadtothesecond-orderODE
H′′
n(x)−2xH′
n(x)+2nHn(x)=0, (13.10)
whichisclearly notself-adjoint.
ToputtheODEinself-adjointform, followingSection10.1, wemultiplyby exp (−x2),
Exercise10.1.2. Thisleadstotheorthogonalityintegral
integraldisplay∞
−∞Hm(x)Hn(x)e−x2dx=0,m/negationslash=n, (13.11)
withtheweightingfunctionexp (−x2),aconsequenceofputtingtheODEintoself-adjoint
form. The interval (−∞,∞)is chosen to obtain the Hermitian operator boundary condi-
tions, Section 10.1. It is sometimes convenient to absorb the weighting function into the
Hermitepolynomials.We maydefine
ϕn(x)=e−x2/2Hn(x), (13.12)
withϕn(x)nolongerapolynomial.
SubstitutionintoEq. (13.10)yieldsthedifferentialequationfor ϕn(x),
ϕ′′
n(x)+parenleftbig
2n+1−x2parenrightbig
ϕn(x)=0. (13.13)
This is the differential equation for a quantum mechanical, simple harmonic oscillator,
which is perhaps the most important physics application of the Hermite polynomials.
822 Chapter 13 More Special Functions
Equation (13.13) is self-adjoint, and the solutions ϕn(x)are orthogonal for the interval
(−∞<x<∞)witha unitweightingfunction.
Theproblemofnormalizingthesefunctionsremains.ProceedingasinSection12.3,we
multiplyEq. (13.1) byitself andthenby e−x2.This yields
e−x2e−s2+2sxe−t2+2tx=∞summationdisplay
m,n=0e−x2Hm(x)Hn(x)smtn
m!n!.
When we integrate this relation over xfrom−∞to+∞, the cross terms of the double
sumdropoutbecauseof theorthogonalityproperty:3
∞summationdisplay
n=0(st)n
n!n!integraldisplay∞
−∞e−x2bracketleftbig
Hn(x)bracketrightbig2dx=integraldisplay∞
−∞e−x2−s2+2sx−t2+2txdx
=integraldisplay∞
−∞e−(x−s−t)2e2stdx
=π1/2e2st=π1/2∞summationdisplay
n=02n(st)n
n!,(13.14)
usingEqs. (8.6) and(8.8). Byequatingcoefficientsoflikepowersof st, weobtain
integraldisplay∞
−∞e−x2bracketleftbig
Hn(x)bracketrightbig2dx=2nπ1/2n!. (13.15)
Quantum Mechanical Simple Harmonic Oscillator
The following development of Hermite polynomials via simple harmonic oscillator wave
functions φn(x)is analogous to the use of the raising and lowering operators for angu-
lar momentum operators presented in Section 4.3. This means that we derive the eigen-
valuesn+1/2 and eigenfunctions (the Hn(x)) without assuming the development that
led to Eq. (13.13). The key aspect of the eigenvalue Eq. (13.13), (d2
dx2−x2)ϕn(x)=
−(2n+1)ϕn(x), isthattheHamiltonian
−2H≡d2
dx2−x2=parenleftbiggd
dx−xparenrightbiggparenleftbiggd
dx+xparenrightbigg
+bracketleftbigg
x,d
dxbracketrightbigg
(13.16)
almostfactorizes.Usingnaively a2−b2=(a−b)(a+b),thebasiccommutator [px,x]=
¯h/iof quantum mechanics (with momentum px=(¯h/i)d/dx ) enters as a correction in
Eq. (13.16). (Because pxis Hermitian, d/dxis anti-Hermitian, (d/dx)†=−d/dx.) This
commutator can be evaluated as follows. Imagine the differential operator d/dxacts on a
wavefunction ϕ(x)totheright,asinEq. (13.13), so
d
dx(xϕ)=xd
dxϕ+ϕ, (13.17)
3The cross terms (m/negationslash=n)may be left in, if desired. Then, when the coefficients of sαtβare equated, the orthogonality will be
apparent.
13.1 Hermite Functions 823
bytheproductrule.Droppingthewavefunction ϕfromEq.(13.17),werewriteEq.(13.17)
as
d
dxx−xd
dx≡bracketleftbiggd
dx,xbracketrightbigg
=1, (13.18)
a constant, and then verify Eq. (13.16) directly by expanding the product of operators.
The product form of Eq. (13.16), up to the constant commutator, suggests introducing the
non-Hermitianoperators
ˆa†≡1√
2parenleftbigg
x−d
dxparenrightbigg
,ˆa≡1√
2parenleftbigg
x+d
dxparenrightbigg
, (13.19)
with(ˆa)†=ˆa†,whichareadjointsofeachother.Theyobeythecommutationrelations
bracketleftbig
ˆa,ˆa†bracketrightbig
=bracketleftbiggd
dx,xbracketrightbigg
=1,[ˆa,ˆa]=0=bracketleftbig
ˆa†,ˆa†bracketrightbig
, (13.20)
which are characteristic of these operators and straightforward to derive from Eq. (13.18)
andbracketleftbiggd
dx,d
dxbracketrightbigg
=0=[x,x]andbracketleftbigg
x,d
dxbracketrightbigg
=−bracketleftbiggd
dx,xbracketrightbigg
.
ReturningtoEq. (13.16)andusingEq. (13.19)werewritetheHamiltonianas
H=ˆa†ˆa+1
2=ˆa†ˆa+1
2parenleftbig
ˆa†ˆa+ˆaˆa†parenrightbig
=1
2parenleftbig
ˆa†ˆa+ˆaˆa†parenrightbig
(13.21)
and introduce the Hermitian number operator N=ˆa†ˆaso that H=N+1/2.Let|n/angbracketrightbe
aneigenfunctionof H,
H|n/angbracketright=λn|n/angbracketright,
whose eigenvalue λnis unknown at this point. Now we prove the key property that Nhas
nonnegativeintegereigenvalues
N|n/angbracketright=parenleftbigg
λn−1
2parenrightbigg
|n/angbracketright=n|n/angbracketright,n=0,1,2..., (13.22)
thatis,λn=n+1/2.Sinceˆa|n/angbracketrightiscomplexconjugateto /angbracketleftn|ˆa†,thenormalizationintegral
/angbracketleftn|ˆa†ˆa|n/angbracketright≥0 andis finite.From
parenleftbig
ˆa|n/angbracketrightparenrightbig†ˆa|n/angbracketright=/angbracketleftn|ˆa†ˆa|n/angbracketright=parenleftbigg
λn−1
2parenrightbigg
≥0 (13.23)
wesee that Nhasnonnegativeeigenvalues.
We now show that if ˆa|n/angbracketrightis nonzero it is an eigenfunction with eigenvalue λn−1=
λn−1. After normalizing ˆa|n/angbracketright, this state is designated |n−1/angbracketright.T h i si sp r o v e db yt h e
commutationrelations
bracketleftbig
N,ˆa†bracketrightbig
=ˆa†,[N,ˆa]=−ˆa, (13.24)
824 Chapter 13 More Special Functions
whichfollowfromEq.(13.20).Thesecommutationrelationscharacterize Nasthenumber
operator.Toseethis,wedeterminetheeigenvalueof Nforthestatesˆa†|n/angbracketrightandˆa|n/angbracketright.Using
ˆaˆa†=N+[ˆa,ˆa†]=N+1,wefindthat
Nparenleftbig
ˆa†|n/angbracketrightparenrightbig
=ˆa†parenleftbig
ˆaˆa†parenrightbig
|n/angbracketright=ˆa†parenleftbigbracketleftbig
ˆa,ˆa†bracketrightbig
+Nparenrightbig
|n/angbracketright
=ˆa†(N+1)|n/angbracketright=parenleftbigg
λn+1
2parenrightbigg
ˆa†|n/angbracketright=(n+1)ˆa†|n/angbracketright,(13.25)
Nparenleftbig
ˆa|n/angbracketrightparenrightbig
=parenleftbig
ˆaˆa†−1parenrightbig
ˆa|n/angbracketright=ˆa(N−1)|n/angbracketright=(n−1)ˆa|n/angbracketright.
Inotherwords, Nactingonˆa†|n/angbracketrightshowsthatˆa†hasraisedtheeigenvalue ncorresponding
to|n/angbracketrightbyoneunit,whenceitsname raising,orcreation,operator.Applyingˆa†repeatedly,
we can reach all higher states. There is no upper limit to the sequence of eigenvalues.
Similarly,ˆalowers the eigenvalue nby one unit; hence it is a lowering (orannihilation )
operator because.Therefore,
ˆa†|n/angbracketright∼|n+1/angbracketright,ˆa|n/angbracketright∼|n−1/angbracketright. (13.26)
Applyingˆarepeatedly,wecanreachthelowest,orground,state |0/angbracketrightwitheigenvalue λ0.W e
cannotsteplowerbecause λ0≥1/2.Thereforeˆa|0/angbracketright≡0,suggestingweconstruct ψ0=|0/angbracketright
from the(factored) first-order ODE
√
2ˆaψ0=parenleftbiggd
dx+xparenrightbigg
ψ0=0. (13.27)
Integrating
ψ′
0
ψ0=−x, (13.28)
weobtain
lnψ0=−1
2x2+lnc0, (13.29)
wherec0is anintegrationconstant.Thesolution,
ψ0(x)=c0e−x2/2, (13.30)
is a Gaussian that can be normalized, with c0=π−1/4using the error integral, Eqs. (8.6)
and(8.8). Substituting ψ0intoEq.(13.13) wefind
H|0/angbracketright=parenleftbigg
ˆa†ˆa+1
2parenrightbigg
|0/angbracketright=1
2|0/angbracketright, (13.31)
so its energy eigenvalue is λ0=1/2 and its number eigenvalue is n=0, confirming the
notation|0/angbracketright.Applyingˆa†repeatedlyto ψ0=|0/angbracketright,allothereigenvaluesareconfirmedtobe
λn=n+1/2,provingEq.(13.13).ThenormalizationsneededforEq.(13.26)followfrom
Eqs. (13.25)and(13.23)and
/angbracketleftn|ˆaˆa†|n/angbracketright=/angbracketleftn|ˆa†ˆa+1|n/angbracketright=n+1, (13.32)
13.1 Hermite Functions 825
showing
√
n+1|n+1/angbracketright=ˆa†|n/angbracketright,√n|n−1/angbracketright=ˆa|n/angbracketright. (13.33)
Thus, the excited-state wave functions, ψ1,ψ2, and so on, are generated by the raising
operator
|1/angbracketright=ˆa†|0/angbracketright=1√
2parenleftbigg
x−d
dxparenrightbigg
ψ0(x)=x√
2
π1/4e−x2/2, (13.34)
yielding(andleadingtoupcomingEq. (13.38))
ψn(x)=NnHn(x)e−x2/2,N n≡π−1/4parenleftbig
2nn!parenrightbig−1/2, (13.35)
whereHnaretheHermitepolynomials(Fig.13.2).
Asshown,theHermitepolynomialsareusedinanalyzingthequantummechanicalsim-
ple harmonic oscillator. For a potential energy V=1
2Kz2=1
2mω2z2(forceF=−∇V=
−Kzˆz),theSchrödingerwaveequationis
−¯h2
2m∇2/Psi1(z)+1
2Kz2/Psi1(z)=E/Psi1(z). (13.36)
Ouroscillatingparticlehasmass mandtotalenergy E.By useoftheabbreviations
x=αzwith α4=mK
¯h2=m2ω2
¯h2,
λ=2E
¯hparenleftbiggm
Kparenrightbigg1/2
=2E
¯hω,(13.37)
in which ωis the angular frequency of the corresponding classical oscillator, Eq. (13.36)
becomes(with /Psi1(z)=/Psi1(x/α)=ψ(x))
d2ψ(x)
dx2+parenleftbig
λ−x2parenrightbig
ψ(x)=0. (13.38)
Thisis Eq.(13.13) with λ=2n+1.Hence(Fig.13.2),
ψn(x)=2−n/2π−1/4(n!)−1/2e−x2/2Hn(x), (normalized ). (13.39)
Alternatively, the requirement that nbe an integer is dictated by the boundary conditions
ofthequantummechanicalsystem,
lim
z→±∞/Psi1(z)=0.
Specifically,if n→ν,notaninteger,apower-seriessolutionofEq.(13.13)(Exercise9.5.6)
shows that Hν(x)will behave as xνex2for large x. The functions ψν(x)and/Psi1ν(z)will
thereforeblowupatinfinity,anditwillbeimpossibletonormalizethewavefunction /Psi1(z).
Withthisrequirement,theenergy Ebecomes
E=parenleftbigg
n+1
2parenrightbigg
¯hω. (13.40)
826 Chapter 13 More Special Functions
FIGURE 13.2Quantummechanical
oscillatorwavefunctions:The
heavybaronthe x-axisindicatesthe
allowedrangeoftheclassical
oscillatorwiththesametotalenergy.
Asnrangesoverintegralvalues (n≥0),weseethattheenergyisquantizedandthatthere
isaminimumorzeropointenergy
Emin=1
2¯hω. (13.41)
This zero point energy is an aspect of the uncertainty principle, a genuine quantum phe-
nomenon.
13.1 Hermite Functions 827
In quantum mechanical problems, particularly in molecular spectroscopy, a number of
integralsoftheform
integraldisplay∞
−∞xre−x2Hn(x)Hm(x)dx
areneeded.Examplesfor r=1andr=2(withn=m)areincludedintheexercisesatthe
endofthissection.AlargenumberofotherexamplesarecontainedinWilson,Decius,and
Cross.4
In the dynamics and spectroscopy of molecules in the Born–Oppenheimer approxima-
tion, the motion of a molecule is separated into electronic, vibrational and rotational mo-
tion. Each vibrating atom contributes to a matrix element two Hermite polynomials, its
initial state and another one to its final state. Thus, integrals of products of Hermite poly-
nomialsareneeded.
Example 13.1.2 THREEFOLD HERMITE FORMULA
Wewanttocalculatethefollowingintegralinvolving m=3 Hermitepolynomials:
I3≡integraldisplay∞
−∞e−x2HN1(x)HN2(x)HN3(x)dx, (13.42)
whereNi≥0 are integers. The formula (due to E. C. Titchmarsh, J. Lond. Math. Soc. 23:
15(1948),seeGradshteynandRyzhik,p.838,intheAdditionalReadings)generalizesthe
m=2 case needed for the orthogonality and normalization of Hermite polynomials. To
derive it, we start with the product of three generating functions of Hermite polynomials,
multiplyby e−x2, andintegrateover xinordertogenerate I3:
Z3≡integraldisplay∞
−∞e−x23productdisplay
j=1e2xtj−t2
jdx=integraldisplay∞
−∞e−(summationtext3
j=1tj−x)2+2(t1t2+t1t3+t2t3)dx
=√πe2(t1t2+t1t3+t2t3). (13.43)
The last equality follows from substituting y=x−summationtext
jtjand using the error integralintegraltext∞
−∞e−y2dy=√π, Eqs. (8.6) and (8.8). Expanding the generating functions in terms of
Hermitepolynomialsweobtain
Z3=∞summationdisplay
N1,N2,N3=0tN1
1tN2
2tN3
3
N1!N2!N3!integraldisplay∞
−∞e−x2HN1(x)HN2(x)HN3(x)dx
=√π∞summationdisplay
N=02N
N!(t1t2+t1t3+t2t3)N
=√π∞summationdisplay
N=02N
N!summationdisplay
0≤ni≤N,summationtext
ini=NN!
n1!n2!n3!(t1t2)n1(t1t3)n2(t2t3)n3,
4E.B.Wilson,Jr.,J.C.Decius,andP.C.Cross, MolecularVibrations ,NewYork:McGraw-Hill(1955),reprintedDover(1980).
828 Chapter 13 More Special Functions
usingthepolynomialexpansion
parenleftbiggmsummationdisplay
j=1ajparenrightbiggN
=summationdisplay
0≤ni≤mN!
n1!···nm!an1
1···anmm.
Thepowersoftheforegoing tjtkbecome
(t1t2)n1(t1t3)n2(t2t3)n3=tN1
1tN2
2tN3
3;
N1=n1+n2,N 2=n1+n3,N 3=n2+n3.
Thatis, from
2N=2(n1+n2+n3)=N1+N2+N3
therefollows
2N=2n1+2N3=2n2+2N2=2n3+2N1,
soweobtain
n1=N−N3,n 2=N−N2,n 3=N−N1.
Theniare all fixed (making this case special and easy) because the Niare fixed, and
2N=3summationtext
i=1Ni, withN≥0 an integer by parity. Hence, upon comparing the foregoing like
t1t2t3powers,
I3=√π2NN1!N2!N3!
(N−N1)!(N−N2)!(N−N3)!, (13.44)
which is the desired formula. If we orderN1≥N2≥N3≥0, thenn1≥n2≥n3≥0
follows,beingequivalentto N−N3≥N−N2≥N−N1≥0,whichoccurinthedenom-
inatorsofthefactorialsof I3. /squaresolid
Example 13.1.3 DIRECT EXPANSION OF PRODUCTS OF HERMITE POLYNOMIALS
Inanalternativeapproach,wenowstartagainfromthegeneratingfunctionidentity
∞summationdisplay
N1,N2=0HN1(x)HN2(x)tN1
1
N1!tN2
2
N2!=e2x(t1+t2)−t2
1−t2
2=e2x(t1+t2)−(t1+t2)2·e2t1t2
=∞summationdisplay
N=0HN(x)(t1+t2)N
N!∞summationdisplay
ν=0(2t1t2)ν
ν!.
13.1 Hermite Functions 829
Usingthebinomialexpansionandthencomparinglikepowersof t1t2weextractanidentity
duetoE. Feldheim( J. Lond.Math.Soc. 13:22(1938)):
HN1(x)HN2(x)=min(N1,N2)summationdisplay
ν=0HN1+N2−2νN1!N2!2ν
ν!(N1+N2−2ν)!parenleftbiggN1+N2−2ν
N1−νparenrightbigg
=summationdisplay
0≤ν≤min(N1,N2)HN1+N2−2ν2νν!parenleftbiggN1
νparenrightbiggparenleftbiggN2
νparenrightbigg
. (13.45)
Forν=0 thecoefficientof HN1+N2isobviouslyunity.Specialcases, suchas
H2
1=H2+2,H 1H2=H3+4H1,H2
2=H4+8H2+8,H 1H3=H4+6H2,
canbederivedfrom Table13.1andagreewiththegeneraltwofoldproductformula.
This compact formula can be generalized to products of mHermite polynomials, and
thisinturnyieldsa newclosedformresultfor theintegral Im.
Let us beginwitha newresultfor I4containinga productof four Hermitepolynomials.
InsertingtheFeldheimidentityfor HN1HN2andHN3HN4andusingorthogonality
integraldisplay∞
−∞e−x2HN1HN2dx=√π2N1N1!δN1N2
fortheremainingproductof twoHermitepolynomialsyields
I4=integraldisplay∞
−∞e−x2HN1HN2HN3HN4dx
=summationdisplay
0≤µ≤min(N1,N2);0≤ν≤min(N3,N4)2µµ!
·parenleftbiggN1
µparenrightbiggparenleftbiggN2
µparenrightbigg
2νν!parenleftbiggN3
νparenrightbiggparenleftbiggN4
νparenrightbiggintegraldisplay∞
−∞e−x2HN1+N2−2µHN3+N4−2νdx
=N4summationdisplay
ν=0√π2M(N3+N4−2ν)!N1!N2!N3!N4!
(M−N3−N4−ν)!(M−N1+ν)!(M−N2+ν)!(N3−ν)!(N4−ν)!ν!.
(13.46)
Hereweusethenotation M=(N1+N2+N3+N4)/2andwritethebinomialcoefficients
explicitly,so
1
2(N1+N2−N3−N4)=M−N3−N4,
1
2(N1−N2+N3+N4)=M−N2,
1
2(N2−N1+N3+N4)=M−N1.
From orthogonality we have µ=(N1+N2−N3−N4)/2+ν. The upper limit
ofνis min(N3,N4,M−N1,M−N2)=min(N4,M−N1)and the lower limit is
max(0,N3+N4−M)=0,if weorder N1≥N2≥N3≥N4.
830 Chapter 13 More Special Functions
Nowwereturntotheproductexpansionof mHermitepolynomialsandthecorrespond-
ingnewresultfrom itfor Im. Weprovea generalizedFeldheimidentity,
HN1(x)···HNm(x)=summationdisplay
ν1,...,νm−1HM(x)aν1,...,νm−1, (13.47)
where
M=m−1summationdisplay
i=1(Ni−2νi)+Nm,
by mathematical induction. Multiplying this by HNm+1and using the Feldheim identity,
we end up with the same formula for m+1 Hermite polynomials, including the recursion
relation
aν1,...,νm=aν1,...,νm−12νmνm!parenleftbiggNm+1
νmparenrightbiggparenleftbiggsummationtextm−1
i=1(Ni−2νi)+Nm+1
νmparenrightbigg
.
Its solutionis
aν1,...,νm−1=m−1productdisplay
i=1parenleftbiggNi+1
νiparenrightbiggparenleftbiggsummationtexti−1
j=1(Nj−2νj)+Ni
νiparenrightbigg
2νiνi!.(13.48)
Thelimitsof thesummationindicesare
0≤ν1≤min(N1,N2),0≤ν2≤min(N3,N1+N2−2ν1),...,
0≤νm−1≤minparenleftbigg
Nm,m−2summationdisplay
i=1(Ni−2νi)+Nm−1parenrightbigg
.(13.49)
We now apply this generalized Feldheim identity, with indices ordered as N1≥N2≥
···≥Nm,t oIm, grouping HN2···HNmtogether and using orthogonality for the re-
maining product of two Hermite polynomials HN1Hsummationtextm−1
i=2(Ni−2νi)+Nm. This yields N1=
summationtextm−1
i=2(Ni−2νi)+Nm,fixingνm−1, and
Im=√π2N1N1!summationdisplay
ν2,...,νm−1m−1productdisplay
i=2parenleftbiggNi+1
νiparenrightbiggparenleftbiggsummationtexti−1
j=2(Nj−2νj)+Ni
νiparenrightbigg
νi!2νi,(13.50)
wherethelimitsonthesummationindicesare
0≤ν2≤min(N3,N2),..., 0≤νm−1≤minparenleftbigg
Nm,m−2summationdisplay
i=2(Ni−2νi)+Nm−1parenrightbigg
.(13.51)
/squaresolid
13.1 Hermite Functions 831
Example 13.1.4 APPLICATIONS OF THE PRODUCT FORMULAS
Tochecktheexpression Imform=3,wenotethatthesumsummationtexti−1
j=2withi−1=m−2=1
in the second binomial coefficient in Imis empty (that is, zero), so only Ni=Nm−1=N2
remains. Also, with νm−2=ν1the sum over the νiis that over ν2, which is fixed by the
constraint on the summation index ν2:N1=N2−2ν2+N3. Henceν2=(N2+N3−
N1)/2=N−N1, withN=(N1+N2+N3)/2. That is, only the product remains in Im.
Thegeneralformulafor Imthereforegives
I3=√π2N1N1!parenleftbiggN3
ν2parenrightbiggparenleftbiggN2
ν2parenrightbigg
ν2!2ν2=√π2NN1!N2!N3!
(N−N1)!(N−N2)!(N−N3)!,
which agrees with our earlier result of Example 13.1.2. The last expression is based on
the following observations. The power of 2 has the exponent N1+ν2=N. The factorials
from the binomial coefficients are N3−ν2=(N1+N3−N2)/2=N−N2,N2−ν2=
(N1+N2−N3)/2=N−N3.
Next let us consider m=4, where we do not order the Hermite indices Nias yet. The
reason is that the general Imexpression was derived with a different grouping of the Her-
mite polynomials than the separate calculation of I4with which we compare. That is why
we’ll have to permute the indices to get the earlier result for I4. That is a general conclu-
sion: Different groupings of the Hermite polynomials just give different permutations of
theHermiteindicesinthegeneralresult.
We havetwosummationsover ν2andνm−1=ν3,whichisfixedbytheconstraint N1=
N2−2ν2+N3−2ν3+N4. Hence
ν3=1
2(N2+N3+N4−N1)−ν2=M−N1−ν2
withM=1
2(N1+N2+N3+N4).Theexponentof 2 is N1+ν2+ν3=M.Thereforefor
m=4t h eImformulagives
I4=√π2N1N1!summationdisplay
ν2≥0parenleftbiggN3
ν2parenrightbiggparenleftbiggN4
ν3parenrightbiggparenleftbiggN2
ν2parenrightbiggparenleftbiggN2−2ν2+N3
ν3parenrightbigg
ν2!2ν2ν3!2ν3
=summationdisplay
ν2≥0√π2MN1!N2!N3!N4!(N2−2ν2+N3)!
ν2!ν3!(N2−ν2)!(N3−ν2)!(N4−ν3)!(N2+N3−2ν2−ν3)!
=summationdisplay
ν2≥0√π2MN1!N2!N3!N4!(N2+N3−2ν2)!
(N2−ν2)!(N3−ν2)!(N4−ν3)!ν2!ν3!(N2+N3−2ν2−ν3)!
=summationdisplay
ν2≥0√π2MN1!N2!N3!N4!(N2+N3−2ν2)!
ν2!(M−N1−ν2)!(N3−ν2)!(N2−ν2)!(M−N2−N3+ν2)!(M−N4−ν2)!.
Inthelastexpressionwehavesubstituted ν3andused
N4−ν3=(N1−N2−N3+N4)+ν2=M−N2−N3+ν2,
N2+N3−2ν2−ν3=N1+N2+N3−N4
2−ν2=M−N4−ν2.
832 Chapter 13 More Special Functions
The upper limit is ν2≤min(N2,N3,M−N1,M−N4), and the lower limit is ν2≥
max(0,N2+N3−M). If we make the permutation N2↔N4,ν2→ν, then our previ-
ousI4resultisobtainedwithupperlimit ν≤min(N4,M−N1)=1
2(N2+N3+N4−N1)
and lower limit ν≥max(0,N3+N4−M)=0 because N3+N4−N1−N2≤0f o r
N1≥N2≥N3≥N4≥0. /squaresolid
The Hermite polynomial product formula also applies to products of simple harmonic
oscillatorwavefunctions,integraltext∞
−∞e−mx2/2HN1(x)···HNm(x)dx,withadifferentexponential
weight function. To evaluate such integrals we use the generalized Feldheim identity for
HN2···HNmin conjunction with the integral (see Gradshteyn and Ryzhik, p. 845, in the
AdditionalReadings),
integraldisplay∞
−∞e−a2x2Hm(x)Hn(x)dx=1
2parenleftbigg2
aparenrightbiggm+n+1parenleftbig
1−a2parenrightbig(m+n)/2Ŵparenleftbiggm+n+1
2parenrightbigg
·2F1parenleftbigg
−m,−n;1−m−n
2;a2
2(a2−1)parenrightbigg
,
instead of the standard orthogonality integral for the remaining product of two Hermite
polynomials.Herethehypergeometricfunctionis thefinitesum
2F1parenleftbigg
−m,−n;1−m−n
2;a2
2(a2−1)parenrightbigg
=min(m,n)−1summationdisplay
ν=0(−m)ν(−n)ν
ν!(1−m−n
2)νparenleftbigga2
2(a2−1)parenrightbiggν
with(−m)ν=(−m)(1−m)···(ν−1−m)and(−m)0≡1.Thisyieldsaresultsimilarto
Imbutsomewhatmorecomplicated.
The oscillator potential has also been employed extensively in calculations of nuclear
structure(nuclearshellmodel)quarkmodelsof hadronsandthenuclearforce.
There is a second independent solution of Eq. (13.13). This Hermite function of the
second kind is an infinite series (Sections 9.5 and 9.6) and is of no physical interest, at
leastnotyet.
Exercises
13.1.1 Assume the Hermite polynomials are known to be solutions of the differential equa-
tion (13.13). From this the recurrence relation, Eq. (13.3), and the values of Hn(0)are
alsoknown.
(a) Assumetheexistenceofageneratingfunction
g(x,t)=∞summationdisplay
n=0Hn(x)tn
n!.
(b) Differentiate g(x,t)with respect to xand using the recurrence relation develop a
first-orderPDEfor g(x,t).
(c) Integratewithrespectto x,holding tfixed.
(d) Evaluate g(0,t)usingEq. (13.5). Finally,showthat
g(x,t)=expparenleftbig
−t2+2txparenrightbig
.
13.1 Hermite Functions 833
13.1.2 In developing the properties of the Hermite polynomials, start at a number of different
points,suchas:
1. Hermite’sODE,Eq. (13.13),
2. Rodrigues’formula, Eq.(13.7),
3. Integralrepresentation,Eq. (13.8),
4. Generatingfunction,Eq. (13.1),
5. Gram–Schmidt construction of a complete set of orthogonal polynomials over
(−∞,∞)withaweightingfactor of exp (−x2),Section10.3.
Outlinehowyoucangofrom anyoneofthesestartingpointstoalltheotherpoints.
13.1.3 Provethat
parenleftbigg
2x−d
dxparenrightbiggn
1=Hn(x).
Hint.Checkoutthefirstcoupleof examplesandthenusemathematicalinduction.
13.1.4 Provethat
vextendsinglevextendsingleHn(x)vextendsinglevextendsingle≤vextendsinglevextendsingleHn(ix)vextendsinglevextendsingle.
13.1.5 Rewritetheseries formof Hn(x), Eq. (13.9), asan ascending powerseries.
ANS.H2n(x)=(−1)nnsummationdisplay
s=0(−1)2s(2x)2s(2n)!
(2s)!(n−s)!,
H2n+1(x)=(−1)nnsummationdisplay
s=0(−1)s(2x)2s+1(2n+1)!
(2s+1)!(n−s)!.
13.1.6 (a) Expand x2rinaseriesof even-orderHermitepolynomials.
(b) Expand x2r+1inaseries ofodd-orderHermitepolynomials.
ANS.(a) x2r=(2r)!
22rrsummationdisplay
n=0H2n(x)
(2n)!(r−n)!
(b)x2r+1=(2r+1)!
22r+1rsummationdisplay
n=0H2n+1(x)
(2n+1)!(r−n)!,r=0,1,2,....
Hint.UseaRodriguesrepresentationandintegratebyparts.
13.1.7 Showthat
(a)integraldisplay∞
−∞Hn(x)expbracketleftbigg
−x2
2bracketrightbigg
dx=braceleftbigg2πn!/(n/2)!,neven
0,n odd.
(b)integraldisplay∞
−∞xHn(x)expbracketleftbigg
−x2
2bracketrightbigg
dx=
0,n even
2π(n+1)!
((n+1)/2)!,nodd.
834 Chapter 13 More Special Functions
13.1.8 Showthat
integraldisplay∞
−∞xme−x2Hn(x)dx=0formaninteger ,0≤m≤n−1.
13.1.9 Thetransitionprobabilitybetweentwooscillatorstates mandndependson
integraldisplay∞
−∞xe−x2Hn(x)Hm(x)dx.
Show that this integral equals π1/22n−1n!δm,n−1+π1/22n(n+1)!δm,n+1. This result
shows that such transitions can occur only between states of adjacent energy levels,
m=n±1.
Hint. Multiply the generating function (Eq. (13.1)) by itself using two different sets of
variables (x,s)and(x,t). Alternatively, the factor xmay be eliminated by the recur-
rencerelation,Eq. (13.2).
13.1.10 Showthat
integraldisplay∞
−∞x2e−x2Hn(x)Hn(x)dx=π1/22nn!parenleftbigg
n+1
2parenrightbigg
.
Thisintegraloccursinthecalculationofthemean-squaredisplacementofourquantum
oscillator.
Hint.Usetherecurrencerelation,Eq.(13.2), andtheorthogonalityintegral.
13.1.11 Evaluate
integraldisplay∞
−∞x2e−x2Hn(x)Hm(x)dx
intermsof nandmandappropriateKroneckerdeltafunctions.
ANS. 2n−1π1/2(2n+1)n!δnm+2nπ1/2(n+2)!δn+2,m+2n−2π1/2n!δn−2,m.
13.1.12 Showthat
integraldisplay∞
−∞xre−x2Hn(x)Hn+p(x)dx=braceleftbigg0,p >r
2nπ1/2(n+r)!,p=r,
withn,p, andrnonnegativeintegers.
Hint.Usetherecurrencerelation,Eq.(13.2), ptimes.
13.1.13 (a) Using the Cauchy integral formula, develop an integral representation of Hn(x)
basedonEq. (13.1)withthecontourenclosingthepoint z=−x.
ANS.Hn(x)=n!
2πiex2contintegraldisplaye−z2
(z+x)n+1dz.
(b) Showbydirectsubstitutionthatthisresultsatisfies theHermiteequation.
13.1 Hermite Functions 835
13.1.14 With
ψn(x)=e−x2/2Hn(x)
(2nn!π1/2)1/2,
verifythat
ˆanψn(x)=1√
2parenleftbigg
x+d
dxparenrightbigg
ψn(x)=n1/2ψn−1(x),
ˆa†
nψn(x)=1√
2parenleftbigg
x−d
dxparenrightbigg
ψn(x)=(n+1)1/2ψn+1(x).
Note. The usual quantum mechanical operator approach establishes these raising and
loweringpropertiesbeforetheform of ψn(x)is known.
13.1.15 (a) Verifytheoperatoridentity
x−d
dx=−expbracketleftbiggx2
2bracketrightbiggd
dxexpbracketleftbigg
−x2
2bracketrightbigg
.
(b) Thenormalizedsimpleharmonicoscillatorwavefunctionis
ψn(x)=parenleftbig
π1/22nn!parenrightbig−1/2expbracketleftbigg
−x2
2bracketrightbigg
Hn(x).
Showthatthismaybewrittenas
ψn(x)=parenleftbig
π1/22nn!parenrightbig−1/2parenleftbigg
x−d
dxparenrightbiggn
expbracketleftbigg
−x2
2bracketrightbigg
.
Note. This corresponds to an n-fold application of the raising operator of Exer-
cise13.1.14.
13.1.16 (a) ShowthatthesimpleoscillatorHamiltonian(from Eq.(13.38)) maybewrittenas
H=−1
2d2
dx2+1
2x2=1
2parenleftbig
ˆaˆa†+ˆa†ˆaparenrightbig
.
Hint.Express Einunitsof¯hω.
(b) Usingthecreation–annihilationoperatorformulationofpart(a), showthat
Hψ(x)=parenleftbigg
n+1
2parenrightbigg
ψ(x).
This means the energy eigenvalues are E=(n+1
2)(¯hω), in agreement with
Eq.(13.40).
13.1.17 Write a program that will generate the coefficients as, in the polynomial form of the
Hermitepolynomial Hn(x)=summationtextn
s=0asxs.
13.1.18 Afunction f(x)is expandedinaHermiteseries:
f(x)=∞summationdisplay
n=0anHn(x).
836 Chapter 13 More Special Functions
From the orthogonality and normalization of the Hermite polynomials the coefficient
anis givenby
an=1
2nπ1/2n!integraldisplay∞
−∞f(x)Hn(x)e−x2dx.
Forf(x)=x8determinetheHermitecoefficients anbytheGauss–Hermitequadrature.
Check your coefficients against AMS-55, Table 22.12 (for the reference see footnote 4
inChapter5ortheGeneralReferencesatbook’send).
13.1.19 (a) In analogy with Exercise 12.2.13, set up the matrix of even Hermite polynomial
coefficientsthatwilltransformanevenHermiteseries intoanevenpowerseries:
B=
1−212···
04−48···
001 6···
.........···
.
ExtendBtohandleanevenpolynomialseriesthrough H8(x).
(b) Invert your matrix to obtain matrix A, which will transform an even power series
(through x8) into a series of even Hermite polynomials. Check the elements of A
against those listed in AMS-55 (Table 22.12, in the General References at book’s
end).
(c) Finally, using matrix multiplication, determine the Hermite series equivalent to
f(x)=x8.
13.1.20 Write asubroutinethatwilltransform afinitepowerseries,summationtextN
n=0anxn, intoaHermite
series,summationtextN
n=0bnHn(x). Usetherecurrencerelation,Eq.(13.2).
Note.BothExercises13.1.19and13.1.20arefasterandmoreaccuratethantheGaussian
quadrature,Exercise13.1.18,if f(x)isavailableasapowerseries.
13.1.21 Write asubroutineforevaluatingHermitepolynomialmatrixelementsof theform
Mpqr=integraldisplay∞
−∞Hp(x)Hq(x)xre−x2dx,
using the 10-point Gauss–Hermite quadrature (for p+q+r≤19). Include a parity
check and set equal to zero the integrals with odd-parity integrand. Also, check to see
ifris in the range |p−q|≤r. Otherwise Mpqr=0. Check your results against the
specificcaseslistedinExercises13.1.9,13.1.10, 13.1.11,and13.1.12.
13.1.22 Calculateandtabulatethenormalizedlinearoscillatorwavefunctions
ψn(x)=2−n/2π−1/4(n!)−1/2Hn(x)expparenleftbigg
−x2
2parenrightbigg
forx=0.0(0.1)5.0
andn=0(1)5.If aplottingroutineisavailable,plotyourresults.
13.1.23 Evaluateintegraltext∞
−∞e−2x2HN1(x)···HN4(x)dxinclosedform.
Hint.integraltext∞
−∞e−2x2HN1(x)HN2(x)HN3(x)dx=1
π2(N1+N2+N3−1)/2·Ŵ(s−N1)Ŵ(s−N2)
·Ŵ(s−N3),s=(N1+N2+N3+1)/2o rintegraltext∞
−∞e−2x2HN1(x)HN2(x)dx=
(−1)(N1+N2−1)/22(N1+N2−1)/2·Ŵ((N1+N2+1)/2)may be helpful. Prove these for-
mulas(seeGradshteynandRyzhik,no.7.375onp.844,intheAdditionalReadings).
13.2 Laguerre Functions 837
13.2 L AGUERRE FUNCTIONS
Differential Equation — Laguerre Polynomials
If we start with the appropriate generating function, it is possible to develop the Laguerre
polynomialsinanalogywiththeHermitepolynomials.Alternatively,aseriessolutionmay
be developed by the methods of Section 9.5. Instead, to illustrate a different technique, let
usstartwithLaguerre’sODEandobtainasolutionintheformofacontourintegral,aswe
didwiththeintegralrepresentationforthemodifiedBesselfunction Kν(x)(Section11.6).
Fromthis integralrepresentationageneratingfunctionwillbederived.
Laguerre’s ODE (which derives from the radial ODE of Schrödinger’s PDE for the hy-
drogenatom)is
xy′′(x)+(1−x)y′(x)+ny(x)=0. (13.52)
We shall attempt to represent y, or rather yn, sinceywill depend on the parameter n,
anonnegativeinteger,bythecontourintegral
yn(x)=1
2πicontintegraldisplaye−xz/(1−z)
(1−z)zn+1dz (13.53a)
anddemonstratethatitsatisfiesLaguerre’sODE.Thecontourincludestheoriginbutdoes
notenclosethepoint z=1.By differentiatingtheexponentialinEq. (13.53a)weobtain
y′
n(x)=−1
2πicontintegraldisplaye−xz/(1−z)
(1−z)2zndz, (13.53b)
y′′
n(x)=1
2πicontintegraldisplaye−xz/(1−z)
(1−z)3zn−1dz. (13.53c)
Substitutingintotheleft-handsideofEq. (13.52), weobtain
1
2πicontintegraldisplaybracketleftbiggx
(1−z)3zn−1−1−x
(1−z)2zn+n
(1−z)zn+1bracketrightbigg
e−xz/(1−z)dz,
whichisequalto
−1
2πicontintegraldisplayd
dzbracketleftbigge−xz/(1−z)
(1−z)znbracketrightbigg
dz. (13.54)
Ifweintegrateourperfectdifferentialaroundaclosedcontour(Fig.13.3),theintegralwill
vanish,thusverifyingthat yn(x)(Eq. (13.53a)) isasolutionofLaguerre’sequation.
It hasbecomecustomarytodefine Ln(x), theLaguerrepolynomial(Fig.13.4), by5
Ln(x)=1
2πicontintegraldisplaye−xz/(1−z)
(1−z)zn+1dz. (13.55)
5Other definitions of Ln(x)are in use. The definitions here of the Laguerre polynomial Ln(x)and the associated Laguerre
polynomial Lkn(x)agree with AMS-55, Chapter 22. (For the full ref. see footnote 4 in Chapter 5 or the General References at
book’s end.)
838 Chapter 13 More Special Functions
FIGURE 13.3Laguerre
polynomialcontour.
FIGURE 13.4Laguerre
polynomials.
Thisis exactlywhatwewouldobtainfromtheseries
g(x,z)=e−xz/(1−z)
1−z=∞summationdisplay
n=0Ln(x)zn,|z|<1, (13.56)
ifwemultiplied g(x,z)byz−n−1andintegratedaroundtheorigin.Asinthedevelopment
of the calculus of residues (Section 7.1), only the z−1term in the series survives. On this
basisweidentify g(x,z)as thegeneratingfunctionfor theLaguerrepolynomials.
13.2 Laguerre Functions 839
Withthetransformation
xz
1−z=s−x,orz=s−x
s, (13.57)
Ln(x)=ex
2πicontintegraldisplaysne−s
(s−x)n+1ds, (13.58)
the new contour enclosing the point s=xin thes-plane. By Cauchy’s integral formula
(for derivatives),
Ln(x)=ex
n!dn
dxnparenleftbig
xne−xparenrightbig
(integraln), (13.59)
givingRodrigues’formulaforLaguerrepolynomials.Fromtheserepresentationsof Ln(x)
wefindtheseriesform (for integral n),
Ln(x)=(−1)n
n!bracketleftbigg
xn−n2
1!xn−1+n2(n−1)2
2!xn−2−···+(−1)nn!bracketrightbigg
=nsummationdisplay
m=0(−1)mn!xm
(n−m)!m!m!=nsummationdisplay
s=0(−1)n−sn!xn−s
(n−s)!(n−s)!s!(13.60)
and the specific polynomials listed in Table 13.2 (Exercise 13.2.1). Clearly, the defini-
tion of Laguerre polynomials in Eqs. (13.55), (13.56), (13.59), and (13.60) are equivalent.
Practical applications will decide which approach is used as one’s starting point. Equa-
tion (13.59) is most convenient for generating Table 13.2, Eq. (13.56) for deriving recur-
sionrelationsfrom whichtheODE(13.52)is recovered.
By differentiating the generating function in Eq. (13.56) with respect to xandz,w e
obtain recurrence relations for the Laguerre polynomials as follows. Using the product
rulefordifferentiationweverifytheidentities
(1−z)2∂g
∂z=(1−x−z)g(x,z), (z −1)∂g
∂x=zg(x,z). (13.61)
Table 13.2 LaguerrePolynomials
L0(x)=1
L1(x)=−x+1
2!L2(x)=x2−4x+2
3!L3(x)=−x3+9x2−18x+6
4!L4(x)=x4−16x3+72x2−96x+24
5!L5(x)=−x5+25x4−200x3+600x2−600x+120
6!L6(x)=x6−36x5+450x4−2400x3+5400x2−4320x+720
840 Chapter 13 More Special Functions
Writingtheleft-handandright-handsidesofthefirstidentityintermsofLaguerrepolyno-
mialsusingEq. (13.56)weobtain
summationdisplay
nbracketleftbig
(n+1)Ln+1(x)−2nLn(x)+(n−1)Ln−1(x)bracketrightbig
zn
=summationdisplay
nbracketleftbig
(1−x)Ln(x)−Ln−1(x)bracketrightbig
zn.
Equatingcoefficientsof znyields
(n+1)Ln+1(x)=(2n+1−x)Ln(x)−nLn−1(x). (13.62)
To get the second recursion relation we use both identities of Eqs. (13.61) to verify the
thirdidentity,
x∂g
∂x=z∂g
∂z−z∂(zg)
∂z, (13.63)
which, when written similarly in terms of Laguerre polynomials, is seen to be equivalent
to
xL′
n(x)=nLn(x)−nLn−1(x). (13.64)
Equation(13.61),modifiedtoread
Ln+1(x)=2Ln(x)−Ln−1(x)−1
n+1bracketleftbig
(1+x)Ln(x)−Ln−1(x)bracketrightbig
,(13.65)
for reasons of economy and numerical stability, is used for computation of numerical val-
ues ofLn(x). The computer starts with known numerical values of L0(x)andL1(x),T a -
ble 13.2, and works up step by step. This is the same technique discussed for computing
Legendrepolynomials,Section12.2.
Also,from Eq.(13.56) wefind
g(0,z)=1
1−z=∞summationdisplay
n=0zn=∞summationdisplay
n=0Ln(0)zn,
whichyieldsthespecialvaluesof Laguerrepolynomials
Ln(0)=1. (13.66)
As is seen from the form of the generating function, from the form of Laguerre’s ODE, or
from Table13.2, theLaguerre polynomialshaveneitheroddnor evensymmetryunder the
paritytransformation x→−x.
The Laguerre ODE is not self-adjoint, and the Laguerre polynomials Ln(x)do not by
themselves form an orthogonal set. However, following the method of Section 10.1, if we
multiplyEq. (13.52)by e−x(Exercise10.1.1)weobtain
integraldisplay∞
0e−xLm(x)Ln(x)dx=δmn. (13.67)
13.2 Laguerre Functions 841
Thisorthogonalityisa consequenceof theSturm–Liouvilletheory,Section10.1. Thenor-
malization follows from the generating function. It is sometimes convenient to define or-
thogonalizedLaguerrefunctions(withunitweightingfunction)by
ϕn(x)=e−x/2Ln(x). (13.68)
Ourneworthonormalfunction, ϕn(x), satisfiestheODE
xϕ′′
n(x)+ϕ′
n(x)+parenleftbigg
n+1
2−x
4parenrightbigg
ϕn(x)=0, (13.69)
which is seen to have the (self-adjoint) Sturm–Liouville form. Note that the interval
(0≤x<∞)was used because Sturm–Liouville boundary conditions are satisfied at its
endpoints.
Associated Laguerre Polynomials
Inmanyapplications,particularlyinquantummechanics,weneedtheassociatedLaguerre
polynomialsdefinedby6
Lk
n(x)=(−1)kdk
dxkLn+k(x). (13.70)
From the series form of Ln(x)we verify that the lowest associated Laguerre polynomials
aregivenby
Lk
0(x)=1,
Lk
1(x)=−x+k+1,
Lk
2(x)=x2
2−(k+2)x+(k+2)(k+1)
2. (13.71)
Ingeneral,
Lk
n(x)=nsummationdisplay
m=0(−1)m(n+k)!
(n−m)!(k+m)!m!xm,k>−1. (13.72)
A generating function may be developed by differentiating the Laguerre generating func-
tionktimestoyield
(−1)kdk
dxke−xz/(1−z)
1−z=(−1)k∞summationdisplay
n=0dk
dxkLn+k(x)zn+k=∞summationdisplay
n=0Lk
n(x)zn+k
=parenleftbiggz
1−zparenrightbiggkexz/(1−z)
1−z.
6Some authors use Lk
n+k(x)=(dk/dxk)[Ln+k(x)]. Henceour Lkn(x)=(−1)kLk
n+k(x).
842 Chapter 13 More Special Functions
Fromthelasttwomembersofthisequation,cancelingthecommonfactor zk, weobtain
e−xz/(1−z)
(1−z)k+1=∞summationdisplay
n=0Lk
n(x)zn,|z|<1. (13.73)
Fromthis, for x=0,thebinomialexpansion
1
(1−z)k+1=∞summationdisplay
n=0parenleftbigg−k−1
nparenrightbigg
(−z)n=∞summationdisplay
n=0Lk
n(0)zn
yields
Lk
n(0)=(n+k)!
n!k!. (13.74)
Recurrence relations can be derived from the generating function or by differentiating the
Laguerrepolynomialrecurrencerelations.Amongthenumerouspossibilitiesare
(n+1)Lk
n+1(x)=(2n+k+1−x)Lk
n(x)−(n+k)Lk
n−1(x), (13.75)
xdLk
n(x)
dx=nLk
n(x)−(n+k)Lk
n−1(x). (13.76)
Thus,weobtainfrom differentiatingLaguerre’sODEonce
xdL′′
n
dx+L′′
n−L′
n+(1−x)dL′
n
dx+ndLn
dx=0,
andeventuallyfrom differentiatingLaguerre’sODE ktimes
xdk
dxkL′′
n+kdk−1
dxk−1L′′
n−kdk−1
dxk−1L′
n+(1−x)dk
dxkL′
n+ndk
dxkLn=0.
Adjustingtheindex n→n+k,wehavetheassociatedLaguerreODE
xd2Lk
n(x)
dx2+(k+1−x)dLk
n(x)
dx+nLk
n(x)=0. (13.77)
When associated Laguerre polynomials appear in a physical problem it is usually because
that physical problem involves Eq. (13.77). The most important application is the bound
statesofthehydrogenatom,whicharederivedinupcomingExample13.2.1.
ARodriguesrepresentationoftheassociatedLaguerrepolynomial
Lk
n(x)=exx−k
n!dn
dxnparenleftbig
e−xxn+kparenrightbig
(13.78)
maybeobtainedfromsubstitutingEq.(13.59)intoEq.(13.70).Notethatalltheseformulas
for associated Legendre polynomials Lk
n(x)reduce to the corresponding expressions for
Ln(x)whenk=0.
13.2 Laguerre Functions 843
The associated Laguerre equation (13.77) is not self-adjoint, but it can be put in self-
adjoint form by multiplying by e−xxk, which becomes the weighting function (Sec-
tion10.1). Weobtain
integraldisplay∞
0e−xxkLk
n(x)Lk
m(x)dx=(n+k)!
n!δmn. (13.79)
Equation (13.79) shows the same orthogonality interval (0,∞)as that for the Laguerre
polynomials, but with a new weighting function we have a new set of orthogonal polyno-
mials,theassociatedLaguerrepolynomials.
Byletting ψk
n(x)=e−x/2xk/2Lk
n(x),ψk
n(x)satisfiestheself-adjointODE
xd2ψk
n(x)
dx2+dψk
n(x)
dx+parenleftbigg
−x
4+2n+k+1
2−k2
4xparenrightbigg
ψk
n(x)=0. (13.80)
Theψk
n(x)aresometimescalled Laguerrefunctions .Equation(13.67)isthespecialcase
k=0 ofEq. (13.79).
Afurther usefulformisgivenbydefining7
/Phi1k
n(x)=e−x/2x(k+1)/2Lk
n(x). (13.81)
SubstitutionintotheassociatedLaguerreequationyields
d2/Phi1k
n(x)
dx2+parenleftbigg
−1
4+2n+k+1
2x−k2−1
4x2parenrightbigg
/Phi1k
n(x)=0. (13.82)
Thecorrespondingnormalizationintegralintegraltext∞
0|/Phi1k
n(x)|2dxis
integraldisplay∞
0e−xxk+1bracketleftbig
Lk
n(x)bracketrightbig2dx=(n+k)!
n!(2n+k+1). (13.83)
Noticethatthe /Phi1k
n(x)donotformanorthogonalset(exceptwith x−1asaweightingfunc-
tion) because of the x−1in the term (2n+k+1)/2x. (The Laguerre functions Lµ
ν(x)in
which the indices νandµarenotintegers may be defined using the confluent hypergeo-
metricfunctionsofSection13.5.)
Example 13.2.1 THEHYDROGEN ATOM
The most important application of the Laguerre polynomials is in the solution of the
Schrödingerequationfor thehydrogenatom.This equationis
−¯h2
2m∇2ψ−Ze2
4πǫ0rψ=Eψ, (13.84)
in which Z=1 for hydrogen, 2 for ionized helium, and so on. Separating variables, we
find that the angular dependence of ψis the spherical harmonics YM
L(θ,ϕ). The radial
part,R(r), satisfiestheequation
−¯h2
2m1
r2d
drparenleftbigg
r2dR
drparenrightbigg
−Ze2
4πǫ0rR+¯h2
2mL(L+1)
r2R=ER. (13.85)
7This corresponds to modifying the function ψin Eq.(13.80) to eliminate the first derivative (compare Exercise 9.6.11).
844 Chapter 13 More Special Functions
Forboundstates, R→0asr→∞,andRisfiniteattheorigin, r=0.Wedonotconsider
continuumstateswithpositiveenergy.Onlywhenthelatterareincludeddohydrogenwave
functionsforma completeset.
By use of the abbreviations (resulting from rescaling rto the dimensionless radial vari-
ableρ)
ρ=αrwithα2=−8mE
¯h2,E<0,λ=mZe2
2πǫ0α¯h2,(13.86)
Eq.(13.85) becomes
1
ρ2d
dρparenleftbigg
ρ2dχ(ρ)
dρparenrightbigg
+parenleftbiggλ
ρ−1
4−L(L+1)
ρ2parenrightbigg
χ(ρ)=0, (13.87)
whereχ(ρ)=R(ρ/α). A comparison with Eq. (13.82) for /Phi1k
n(x)shows that Eq. (13.87)
issatisfiedby
ρχ(ρ)=e−ρ/2ρL+1L2L+1
λ−L−1(ρ), (13.88)
inwhich kis replacedby 2 L+1 andnbyλ−L−1,uponusing
1
ρ2d
dρρ2dχ
dρ=1
ρd2
dρ2(ρχ).
Wemustrestricttheparameter λbyrequiringittobeaninteger n,n=1,2,3,....8This
isnecessarybecausetheLaguerrefunctionofnonintegral nwoulddiverge9asρneρ,which
isunacceptableforour physicalproblem,inwhich
limr→∞R(r)=0.
This restriction on λ, imposed by our boundary condition, has the effect of quantizing the
energy,
En=−Z2m
2n2¯h2parenleftbigge2
4πǫ0parenrightbigg2
. (13.89)
The negativesign reflects the fact that we are dealing here with bound states ( E<0), cor-
responding to an electron that is unable to escape to infinity, where the Coulomb potential
goestozero.Usingthis resultfor En,weha v e
α=me2
2πǫ0¯h2·Z
n=2Z
na0,ρ=2Z
na0r, (13.90)
with
a0=4πǫ0¯h2
me2,theBohrradius.
8This is the conventional notation for λ.It is not the same nas the index nin/Phi1kn(x).
9This maybe shown, as in Exercise9.5.5.
13.2 Laguerre Functions 845
Thus,thefinalnormalizedhydrogenwavefunctioniswrittenas
ψnLM(r,θ,ϕ)=bracketleftbiggparenleftbigg2Z
na0parenrightbigg3(n−L−1)!
2n(n+L)!bracketrightbigg1/2
e−αr/2(αr)LL2L+1
n−L−1(αr)YM
L(θ,ϕ).
(13.91)
Regular solutions exist for n≥L+1, so the lowest state with L=1 (called a 2P state)
occursonlywith n=2. /squaresolid
Exercises
13.2.1 Show with the aid of the Leibniz formula that the series expansion of Ln(x)
(Eq. (13.60))followsfrom theRodriguesrepresentation(Eq. (13.59)).
13.2.2 (a) Usingtheexplicitseriesform (Eq. (13.60)) showthat
L′
n(0)=−n,
L′′
n(0)=1
2n(n−1).
(b) Repeatwithoutusingtheexplicitseriesform of Ln(x).
13.2.3 FromthegeneratingfunctionderivetheRodriguesrepresentation
Lk
n(x)=exx−k
n!dn
dxnparenleftbig
e−xxn+kparenrightbig
.
13.2.4 Derivethenormalizationrelation(Eq.(13.79))fortheassociatedLaguerrepolynomials.
13.2.5 Expandxrin aseries of associatedLaguerrepolynomials Lk
n(x),kfixedand nranging
from 0to r(or to∞ifris notaninteger).
Hint.TheRodriguesform of Lk
n(x)willbeuseful.
ANS.xr=(r+k)!r!rsummationdisplay
n=0(−1)nLk
n(x)
(n+k)!(r−n)!,0≤x<∞.
13.2.6 Expande−axinaseriesofassociatedLaguerrepolynomials Lk
n(x),kfixedand nrang-
ingfrom 0to∞.
(a) Evaluatedirectlythecoefficientsinyourassumedexpansion.
(b) Developthedesiredexpansionfromthegeneratingfunction.
ANS.e−ax=1
(1+a)1+k∞summationdisplay
n=0parenleftbigga
1+aparenrightbiggn
Lk
n(x), 0≤x<∞.
13.2.7 Showthatintegraldisplay∞
0e−xxk+1Lk
n(x)Lk
n(x)dx=(n+k)!
n!(2n+k+1).
Hint.Notethat
xLk
n=(2n+k+1)Lk
n−(n+k)Lk
n−1−(n+1)Lk
n+1.
846 Chapter 13 More Special Functions
13.2.8 AssumethataparticularprobleminquantummechanicshasledtotheODE
d2y
dx2−bracketleftbiggk2−1
4x2−2n+k+1
2x+1
4bracketrightbigg
y=0
for nonnegativeintegers n,k.Writey(x)as
y(x)=A(x)B(x)C(x),
withtherequirementthat
(a)A(x)be anegative exponential giving the required asymptotic behavior of y(x)
and
(b)B(x)beapositivepowerof xgivingthebehaviorof y(x)for 0≤x≪1.
Determine A(x)andB(x).Findtherelationbetween C(x)andtheassociatedLaguerre
polynomial.
ANS.A(x)=e−x/2,B(x)=x(k+1)/2,C(x)=Lk
n(x).
13.2.9 FromEq.(13.91) thenormalizedradialpartof thehydrogenicwavefunctionis
RnL(r)=bracketleftbigg
α3(n−L−1)!
2n(n+L)!bracketrightbigg1/2
e−αr(αr)LL2L+1
n−L−1(αr),
inwhich α=2Z/na0=2Zme2/4πε0¯h2.Evaluate
(a)/angbracketleftr/angbracketright=integraldisplay∞
0rRnL(αr)RnL(αr)r2dr,
(b)angbracketleftbig
r−1angbracketrightbig
=integraldisplay∞
0r−1RnL(αr)RnL(αr)r2dr.
The quantity/angbracketleftr/angbracketrightis the average displacement of the electron from the nucleus, whereas
/angbracketleftr−1/angbracketrightistheaverageofthereciprocaldisplacement.
ANS./angbracketleftr/angbracketright=a0
2bracketleftbig
3n2−L(L+1)bracketrightbig
,angbracketleftbig
r−1angbracketrightbig
=1
n2a0.
13.2.10 Derivetherecurrencerelationforthehydrogenwavefunctionexpectationvalues:
s+2
n2angbracketleftbig
rs+1angbracketrightbig
−(2s+3)a0angbracketleftbig
rsangbracketrightbig
+s+1
4bracketleftbig
(2L+1)2−(s+1)2bracketrightbig
a2
0angbracketleftbig
rs−1angbracketrightbig
=0,
withs≥−2L−1,/angbracketleftrs/angbracketright≡/overbarrs.
Hint.TransformEq.(13.87)intoaformanalogoustoEq.(13.80).Multiplyby ρs+2u′−
cρs+1u.Her eu=ρ/Phi1. Adjustctocanceltermsthatdonotyieldexpectationvalues.
13.2.11 The hydrogen wave functions, Eq. (13.91), are mutually orthogonal, as they should be,
sincetheyareeigenfunctionsof theself-adjointSchrödingerequation
integraldisplay
ψ∗
n1L1M1ψn2L2M2r2drd/Omega1=δn1n2δL1L2δM1M2.
13.2 Laguerre Functions 847
Yettheradialintegralhas the(misleading)form
integraldisplay∞
0e−αr/2(αr)LL2L+1
n1−L−1(αr)e−αr/2(αr)LL2L+1
n2−L−1(αr)r2dr,
whichappearsto match Eq. (13.83) and not the associated Laguerre orthogonality re-
lation,Eq. (13.79). Howdoyouresolvethisparadox?
ANS.The parameter αis dependent on n. The first three α,p r e v i -
ously shown, are 2 Z/n1a0. The last three are 2 Z/n2a0.F o r
n1=n2Eq.(13.83)applies.For n1/negationslash=n2neitherEq.(13.79)
nor Eq.(13.83) isapplicable.
13.2.12 A quantum mechanical analysis of the Stark effect (parabolic coordinate) leads to the
ODE
d
dξparenleftbigg
ξdu
dξparenrightbigg
+parenleftbigg1
2Eξ+L−m2
4ξ−1
4Fξ2parenrightbigg
u=0.
HereFis a measure of the perturbation energy introduced by an external electric field.
Find the unperturbed wave functions (F=0)in terms of associated Laguerre polyno-
mials.
ANS.u(ξ)=e−εξ/2ξm/2Lm
p(εξ), withε=√
−2E>0,
p=L/ε−(m+1)/2, anonnegativeinteger.
13.2.13 Thewaveequationforthethree-dimensionalharmonicoscillatoris
−¯h2
2M∇2ψ+1
2Mω2r2ψ=Eψ.
Hereωistheangularfrequencyofthecorrespondingclassicaloscillator.Showthatthe
radial part of ψ(in spherical polar coordinates ) may be written in terms of associated
Laguerrefunctionsof argument (βr2), whereβ=Mω/¯h.
Hint. As in Exercise 13.2.8, split off radial factors of rlande−βr2/2. The associated
Laguerrefunctionwillhavetheform Ll+1/2
1/2(n−l−1)(βr2).
13.2.14 Write acomputerprogramthatwillgeneratethecoefficients asinthepolynomialform
oftheLaguerrepolynomial Ln(x)=summationtextn
s=0asxs.
13.2.15 Write a computer program that will transform a finite power seriessummationtextN
n=0anxninto a
LaguerreseriessummationtextN
n=0bnLn(x). Usetherecurrencerelation,Eq. (13.62).
13.2.16 Tabulate L10(x)forx=0.0(0.1)30.0. This will include the 10 roots of L10. Beyond
x=30.0,L10(x)is monotonic increasing. If graphic software is available, plot your
results.
Checkvalue. Eighthroot=16.279.
13.2.17 Determinethe10rootsof L10(x)usingroot-findingsoftware.Youmayuseyourknowl-
edgeoftheapproximatelocationoftherootsordevelopasearchroutinetolookforthe
roots.The10rootsof L10(x)aretheevaluationpointsforthe10-pointGauss–Laguerre
quadrature. Check your values by comparing with AMS-55, Table 25.9. (For the refer-
enceseefootnote4 inChapter5or theGeneralReferencesatbook’send.)
848 Chapter 13 More Special Functions
13.2.18 Calculate the coefficients of a Laguerre series expansion (Ln(x),k=0)of the ex-
ponential e−x. Evaluate the coefficients by the Gauss–Laguerre quadrature (compare
Eq. (10.64)). CheckyourresultsagainstthevaluesgiveninExercise13.2.6.
Note. Direct application of the Gauss–Laguerre quadrature with f(x)Ln(x)e−xgives
poor accuracy because of the extra e−x. Try a change of variable, y=2x, so that the
functionappearingintheintegrandwillbesimply Ln(y/2).
13.2.19 (a) WriteasubroutinetocalculatetheLaguerrematrixelements
Mmnp=integraldisplay∞
0Lm(x)Ln(x)xpe−xdx.
Include a check of the condition |m−n|≤p≤m+n. (Ifpis outside this range,
Mmnp=0.Why?)
Note. A 10-point Gauss–Laguerre quadrature will give accurate results for
m+n+p≤19.
(b) Call your subroutine to calculate a variety of Laguerre matrix elements. Check
Mmn1againstExercise13.2.7.
13.2.20 Writeasubroutinetocalculatethenumericalvalueof Lk
n(x)forspecifiedvaluesof n,k,
andx.Requirethat nandkbenonnegativeintegersand x≥0.
Hint.Startingwithknownvaluesof Lk
0andLk
1(x), we mayusethe recurrencerelation,
Eq. (13.75),togenerate Lk
n(x),n=2,3,4,....
13.2.21 Showthatintegraltext∞
−∞xne−x2Hn(xy)dx=√πn!Pn(y), wherePnisaLegendrepolynomial.
13.2.22 Write a program to calculate the normalized hydrogen radial wave function ψnL(r).
This isψnLMof Eq. (13.91), omitting the spherical harmonic YM
L(θ,ϕ).T a k eZ=1
anda0=1 (which means that ris being expressed in units of Bohr radii). Accept n
andLas input data. Tabulate ψnL(r)forr=0.0(0.2)RwithRtaken large enough to
exhibit the significant features of ψ. This means roughly R=5f o rn=1,R=10 for
n=2,andR=30 forn=3.
13.3 C HEBYSHEV POLYNOMIALS
In this section two types of Chebyshev polynomials are developed as special cases of ul-
traspherical polynomials. Their properties follow from the ultraspherical polynomial gen-
erating function. The primary importance of the Chebyshev polynomials is in numerical
analysis.
Generating Functions
InSection12.1thegeneratingfunctionfor theultraspherical,orGegenbauer,polynomials
1
(1−2xt+t2)α=∞summationdisplay
n=0C(α)
n(x)tn,|x|<1,|t|<1 (13.92)
wasmentioned,with α=1
2givingrisetotheLegendrepolynomials.Inthissectionwefirst
takeα=1 and then α=0 to generate two sets of polynomials known as the Chebyshev
polynomials.
13.3 Chebyshev Polynomials 849
Type II
Withα=1 andC(1)
n(x)=Un(x), Eq.(13.92) gives
1
1−2xt+t2=∞summationdisplay
n=0Un(x)tn,|x|<1,|t|<1. (13.93)
Thesefunctions Un(x)generatedby (1−2xt+t2)−1arelabeledChebyshevpolynomials
type II. Although these polynomials have few applications in mathematical physics, one
unusualapplicationisinthedevelopmentoffour-dimensionalsphericalharmonicsusedin
angularmomentumtheory.
Type I
Withα=0 there is a difficulty. Indeed, our generating function reduces to the constant 1.
WemayavoidthisproblembyfirstdifferentiatingEq.(13.92)withrespectto t.Thisyields
−α(−2x+2t)
(1−2xt+t2)α+1=∞summationdisplay
n=1nC(α)
n(x)tn−1, (13.94)
or
x−t
(1−2xt+t2)α+1=∞summationdisplay
n=1n
2bracketleftbiggC(α)
n(x)
αbracketrightbigg
tn−1. (13.95)
Wedefine C(0)
n(x)by
C(0)
n(x)=lim
α→0C(α)
n(x)
α. (13.96)
The purpose of differentiating with respect to twas to get αin the denominator and
to create an indeterminate form. Now multiplying Eq. (13.95) by 2 tand adding 1 =
(1−2xt+t2)/(1−2xt+t2),weobtain
1−t2
1−2xt+t2=1+2∞summationdisplay
n=1n
2C(0)
n(x)tn. (13.97)
Wedefine Tn(x)by
Tn(x)=braceleftBigg1,n =0
n
2C(0)
n(x), n> 0.(13.98)
Noticethespecialtreatmentfor n=0.Thisissimilartothetreatmentofthe n=0t e r mi n
theFourierseries.Also,notethat C(0)
nisthelimitindicatedinEq.(13.96)andnotaliteral
substitutionof α=0 intothegeneratingfunctionseries. Withthesenewlabels,
1−t2
1−2xt+t2=T0(x)+2∞summationdisplay
n=1Tn(x)tn,|x|≤1,|t|<1. (13.99)
850 Chapter 13 More Special Functions
WecallTn(x)thetypeIChebyshevpolynomials.Notethatthenotationandspellingofthe
name for these functions differs from reference to reference. Here we follow the usage of
AMS-55(for thefullreferenceseefootnote4inChapter5).
Differentiating the generating function (Eqs. (13.99)) with respect to tand multiplying
bythedenominator, 1 −2xt+t2, weobtain
−t−(t−x)bracketleftbigg
T0(x)+2∞summationdisplay
n=1Tn(x)tnbracketrightbigg
=parenleftbig
1−2xt+t2parenrightbig∞summationdisplay
n=1nTn(x)tn−1
=∞summationdisplay
n=1bracketleftbig
nTntn−1−2xnTntn+nTntn+1bracketrightbig
,
fromwhichtherecurrencerelation
Tn+1(x)−2xTn(x)+Tn−1(x)=0 (13.100)
follows by shifting the summation index so as to get the same power, tn, in each term and
thencomparingcoefficientsof tn. SimilarlytreatingEq.(13.93) wefind
−2(t−x)
1−2xt+t2=parenleftbig
1−2xt+t2parenrightbig∞summationdisplay
n=1nUn(x)tn−1
fromwhichtherecursionrelation
Un+1(x)−2xUn(x)+Un−1(x)=0 (13.101)
followsuponcomparingcoefficientsoflikepowersof t(seeTable13.3).
Then, using the generating functions for the first few values of nand these recurrence
relationsforthehigher-orderpolynomials,wegetTables13.4and13.5(seealsoFigs.13.5
and13.6).
As with the Hermite polynomials, Section 13.1, the recurrence relations, Eqs. (13.100)
and(13.101),togetherwiththeknownvaluesof T0(x),T1(x),U0(x),andU1(x),providea
convenient—thatis,foracomputer—meansofgettingthenumericalvalueofany Tn(x0)
orUn(x0),withx0agivennumber.
Table 13.3 RecursionrelationaPn+1(x)=
(Anx+Bn)Pn(x)−CnPn−1(x)
Pn(x) A n BnCn
Legendre Pn(x)2n+1
n+101
n+1
Chebyshev I Tn(x) 20 1
Shifted Chebyshev I T∗n(x) 4−21
Chebyshev II Un(x) 20 1
Shifted Chebyshev II U∗n(x) 4−21
AssociatedLaguerre L(k)
n(x)−1
n+12n+k+1
n+1n+k
n+1
Hermite Hn(x) 20 2 n
aPndenotesany oftheorthogonalpolynomials.
13.3 Chebyshev Polynomials 851
Table 13.4 Chebyshev
polynomials,typeI
T0=1
T1=x
T2=2x2−1
T3=4x3−3x
T4=8x4−8x2+1
T5=16x5−20x3+5x
T6=32x6−48x4+18x2−1Table 13.5 Chebyshev
polynomials,typeII
U0=1
U1=2x
U2=4x2−1
U3=8x3−4x
U4=16x4−12x2+1
U5=32x5−32x3+6x
U6=64x6−80x4+24x2−1
FIGURE 13.5Chebyshevpolynomials Tn(x).
FIGURE 13.6Chebyshevpolynomials Un(x).
852 Chapter 13 More Special Functions
Again, from the generating functions, we can obtain the special values of various poly-
nomials:
Tn(1)=1,T n(−1)=(−1)n,
(13.102)
T2n(0)=(−1)n,T 2n+1(0)=0;
Un(1)=n+1,U n(−1)=(−1)n(n+1),
(13.103)
U2n(0)=(−1)n,U 2n+1(0)=0.
Forexample,comparingthepowerseries
1−t2
(1−t)2=1+t
1−t=∞summationdisplay
n=0parenleftbig
tn+tn+1parenrightbig
withEq.(13.99)for x=1gi vesTn(1),andforx=−1asimilarexpansionof (1−t)/(1+
t)givesTn(−1),whilereplacing t→−t2inthefirstpowerseriesyields Tn(0).Thepower
seriesfor (1±t)−2and(1+t2)−1generatethecorresponding Un(±1),Un(0).
The parity relations for TnandUnfollow from their generating functions, with the sub-
stitutions t→−t,x→−x,whichleavetheminvariant;theseare
Tn(x)=(−1)nTn(−x), U n(x)=(−1)nUn(−x). (13.104)
Rodrigues’representationsof Tn(x)andUn(x)are
Tn(x)=(−1)nπ1/2(1−x2)1/2
2n(n−1
2)!dn
dxnbracketleftbigparenleftbig
1−x2parenrightbign−1/2bracketrightbig
(13.105)
and
Un(x)=(−1)n(n+1)π1/2
2n+1(n+1
2)!(1−x2)1/2dn
dxnbracketleftbigparenleftbig
1−x2parenrightbign+1/2bracketrightbig
.(13.106)
Recurrence Relations — Derivatives
Differentiationofthegeneratingfunctionsfor Tn(x)andUn(x)withrespecttothevariable
xleads to a variety of recurrence relations involving derivatives. For example, from Eq.
(13.99)wethusobtain
parenleftbig
1−2xt+t2parenrightbig
2∞summationdisplay
n=1T′
n(x)tn=2tbracketleftbigg
T0(x)+2∞summationdisplay
n=1Tn(x)tnbracketrightbigg
,
fromwhichweextracttherecursion
2Tn−1(x)=T′
n(x)−2xT′
n−1(x)+T′
n−2(x), (13.107)
which is the derivative of Eq. (13.100) for n→n−1. Among the more useful recursions
wethusfindare
parenleftbig
1−x2parenrightbig
T′
n(x)=−nxTn(x)+nTn−1(x) (13.108)
13.3 Chebyshev Polynomials 853
and
parenleftbig
1−x2parenrightbig
U′
n(x)=−nxUn(x)+(n+1)Un−1(x). (13.109)
ManipulatingavarietyoftheserecursionsasinSection12.2forLegendrepolynomialsone
can eliminate the index n−1a l s oi nf a v o ro f T′′
nand establish that Tn(x), the Chebyshev
polynomialtypeI, satisfiestheODE
parenleftbig
1−x2parenrightbig
T′′
n(x)−xT′
n(x)+n2Tn(x)=0. (13.110)
TheChebyshevpolynomialoftypeII, Un(x), satisfies
parenleftbig
1−x2parenrightbig
U′′
n(x)−3xU′
n(x)+n(n+2)Un(x)=0. (13.111)
Chebyshev polynomials may be defined starting from these ODEs, but our emphasis has
beenongeneratingfunctions.
Theultrasphericalequation
parenleftbig
1−x2parenrightbigd2
dx2C(α)
n(x)−(2α+1)xd
dxC(α)
n(x)+n(n+2α)C(α)
n(x)=0 (13.112)
is a generalization of these differential equations, reducing to Eq. (13.110) for α=0 and
Eq.(13.111)for α=1 (andtoLegendre’sequationfor α=1
2).
Trigonometric Form
AtthispointinthedevelopmentofthepropertiesoftheChebyshevsolutionsitisbeneficial
to change variables, replacing xby cosθ. Withx=cosθandd/dx=(−1/sinθ)(d/dθ) ,
weverifythat
parenleftbig
1−x2parenrightbigd2Tn
dx2=d2Tn
dθ2−cotθdTn
dθ,xT′
n=−cotθdTn
dθ.
Addingtheseterms, Eq.(13.110) becomes
d2Tn
dθ2+n2Tn=0, (13.113)
the simple harmonic oscillator equation with solutions cos nθand sinnθ. The special val-
ues(boundaryconditionsat x=0,1)identify
Tn=cosnθ=cosn(arccosx). (13.114a)
Asecondlinearlyindependentsolutionof Eqs. (13.110)and(13.113)is labeled
Vn=sinnθ=sinn(arccosx). (13.114b)
ThecorrespondingsolutionsofthetypeII Chebyshevequation,Eq.(13.111), become
Un=sin(n+1)θ
sinθ, (13.115a)
Wn=cos(n+1)θ
sinθ. (13.115b)
854 Chapter 13 More Special Functions
Thetwosetsof solutions,typeI andtypeII, arerelatedby
Vn(x)=parenleftbig
1−x2parenrightbig1/2Un−1(x), (13.116a)
Wn(x)=parenleftbig
1−x2parenrightbig−1/2Tn+1(x). (13.116b)
As already seen from generating functions, Tn(x)andUn(x)are polynomials. Clearly,
Vn(x)andWn(x)arenotpolynomials.From
Tn(x)+iVn(x)=cosnθ+isinnθ
=(cosθ+isinθ)n=bracketleftbig
x+iparenleftbig
1−x2parenrightbig1/2bracketrightbign,|x|≤1 (13.117)
weobtainexpansions
Tn(x)=xn−parenleftbiggn
2parenrightbigg
xn−2parenleftbig
1−x2parenrightbig
+parenleftbiggn
4parenrightbigg
xn−4parenleftbig
1−x2parenrightbig2−··· (13.118a)
and
Vn(x)=radicalbig
1−x2bracketleftbiggparenleftbiggn
1parenrightbigg
xn−1−parenleftbiggn
3parenrightbigg
xn−3parenleftbig
1−x2parenrightbig
+···bracketrightbigg
. (13.118b)
Fromthegeneratingfunctions,orfrom theODEs, power-seriesrepresentationsare
Tn(x)=n
2[n/2]summationdisplay
m=0(−1)m(n−m−1)!
m!(n−2m)!(2x)n−2m, (13.119a)
forn≥1,with[n/2]thelargestintegerbelow n/2 and
Un(x)=[n/2]summationdisplay
m=0(−1)m(n−m)!
m!(n−2m)!(2x)n−2m. (13.119b)
Orthogonality
IfEq.(13.110)isputintoself-adjointform(Section10.1),weobtain w(x)=(1−x2)−1/2
asaweightingfactor.ForEq.(13.111)thecorrespondingweightingfactoris (1−x2)+1/2.
Theresultingorthogonalityintegrals,
integraldisplay1
−1Tm(x)Tn(x)parenleftbig
1−x2parenrightbig−1/2dx=
0,m/negationslash=n,π
2,m=n/negationslash=0,
π, m=n=0,(13.120)
integraldisplay1
−1Vm(x)Vn(x)parenleftbig
1−x2parenrightbig−1/2dx=
0,m/negationslash=n,
π
2,m=n/negationslash=0,
0,m=n=0,(13.121)
integraldisplay1
−1Um(x)Un(x)parenleftbig
1−x2parenrightbig1/2dx=π
2δm,n, (13.122)
13.3 Chebyshev Polynomials 855
and
integraldisplay1
−1Wm(x)Wn(x)parenleftbig
1−x2parenrightbig1/2dx=π
2δm,n, (13.123)
are a direct consequence of the Sturm–Liouville theory, Chapter 10. The normalization
values may best be obtained by using x=cosθand converting these four integrals into
Fouriernormalizationintegrals(for thehalf-periodinterval [0,π]).
Exercises
13.3.1 AnotherChebyshevgeneratingfunctionis
1−xt
1−2xt+t2=∞summationdisplay
n=0Xn(x)tn,|t|<1.
HowisXn(x)relatedto Tn(x)andUn(x)?
13.3.2 Given
parenleftbig
1−x2parenrightbig
U′′
n(x)−3xU′
n(x)+n(n+2)Un(x)=0,
showthat Vn(x)(Eq. (13.116a))satisfies
parenleftbig
1−x2parenrightbig
V′′
n(x)−xV′
n(x)+n2Vn(x)=0,
whichisChebyshev’sequation.
13.3.3 ShowthattheWronskianof Tn(x)andVn(x)is givenby
Tn(x)V′
n(x)−T′
n(x)Vn(x)=−n
(1−x2)1/2.
This verifies that TnandVn(n/negationslash=0)are independent solutions of Eq. (13.110). Con-
versely,for n=0,wedonothavelinearindependence.Whathappensat n=0?Where
isthe“second”solution?
13.3.4 Showthat Wn(x)=(1−x2)−1/2Tn+1(x)isasolutionof
parenleftbig
1−x2parenrightbig
W′′
n(x)−3xW′
n(x)+n(n+2)Wn(x)=0.
13.3.5 EvaluatetheWronskianof Un(x)andWn(x)=(1−x2)−1/2Tn+1(x).
13.3.6 Vn(x)=(1−x2)1/2Un−1(x)is not defined for n=0. Show that a second and inde-
pendent solution of the Chebyshev differential equation for Tn(x) (n=0)isV0(x)=
arccosx(or arcsin x).
13.3.7 Show that Vn(x)satisfies the same three-term recurrence relation as Tn(x)
(Eq. (13.100)).
13.3.8 Verifytheseries solutionsfor Tn(x)andUn(x)(Eqs. (13.109a)and(13.119b)).
856 Chapter 13 More Special Functions
13.3.9 Transform theseriesform of Tn(x), Eq.(13.119a),intoan ascending powerseries.
ANS.T2n(x)=(−1)nnnsummationdisplay
m=0(−1)m(n+m−1)!
(n−m)!(2m)!(2x)2m,n≥1,
T2n+1(x)=2n+1
2nsummationdisplay
m=0(−1)m+n(n+m)!
(n−m)!(2m+1)!(2x)2m+1.
13.3.10 Rewritetheseries formof Un(x), Eq. (13.119b),as anascendingpowerseries.
ANS.U2n(x)=(−1)nnsummationdisplay
m=0(−1)m(n+m)!
(n−m)!(2m)!(2x)2m,
U2n+1(x)=(−1)nnsummationdisplay
m=0(−1)m(n+m+1)!
(n−m)!(2m+1)!(2x)2m+1.
13.3.11 DerivetheRodriguesrepresentationof Tn(x),
Tn(x)=(−1)nπ1/2(1−x2)1/2
2n(n−1
2)!dn
dxnbracketleftbigparenleftbig
1−x2parenrightbign−1/2bracketrightbig
.
Hint.Onepossibilityis tousethehypergeometricfunctionrelation
2F1(a,b;c;z)=(1−z)−a2F1parenleftbigg
a,c−b;c;−z
1−zparenrightbigg
,
withz=(1−x)/2.Analternateapproachistodevelopafirst-orderdifferentialequation
fory=(1−x2)n−1/2.RepeateddifferentiationofthisequationleadstotheChebyshev
equation.
13.3.12 (a) Fromthedifferentialequationfor Tn(in self-adjointform) showthat
integraldisplay1
−1dTm(x)
dxdTn(x)
dxparenleftbig
1−x2parenrightbig1/2dx=0,m/negationslash=n.
(b) Confirmtheprecedingresultbyshowingthat
dTn(x)
dx=nUn−1(x).
13.3.13 Theexpansionof apowerof xinaChebyshevseriesleadstotheintegral
Imn=integraldisplay1
−1xmTn(x)dx√
1−x2.
(a) Showthatthisintegralvanishesfor m<n.
(b) Showthatthisintegralvanishesfor m+nodd.
13.3.14 Evaluatetheintegral
Imn=integraldisplay1
−1xmTn(x)dx√
1−x2
form≥nandm+nevenbyeachof twomethods:
13.3 Chebyshev Polynomials 857
(a) Operatewith xasthevariablereplacing TnbyitsRodriguesrepresentation.
(b) Using x=cosθtransform theintegraltoa formwith θasthevariable.
ANS.Imn=πm!
(m−n)!(m−n−1)!!
(m+n)!!,m≥n, m+neven.
13.3.15 Establishthefollowingbounds, −1≤x≤1:
(a)|Un(x)|≤n+1, (b)vextendsinglevextendsinglevextendsinglevextendsingled
dxTn(x)vextendsinglevextendsinglevextendsinglevextendsingle≤n2.
13.3.16 (a) Establishthefollowingbound, −1≤x≤1:|Vn(x)|≤1.
(b) Showthat Wn(x)is unboundedin −1≤x≤1.
13.3.17 Verifytheorthogonality-normalizationintegralsfor
(a)Tm(x),Tn(x),(b)Vm(x),Vn(x),
(c)Um(x),Un(x),(d)Wm(x),Wn(x).
Hint.AllthesecanbeconvertedtoFourierorthogonality-normalizationintegrals.
13.3.18 Showwhether
(a)Tm(x)andVn(x)are or are not orthogonal over the interval [−1,1]with respect
totheweightingfactor (1−x2)−1/2.
(b)Um(x)andWn(x)are or are not orthogonal over the interval [−1,1]with respect
totheweightingfactor (1−x2)1/2.
13.3.19 Derive
(a)Tn+1(x)+Tn−1(x)=2xTn(x),
(b)Tm+n(x)+Tm−n(x)=2Tm(x)Tn(x),
fromthe“corresponding”cosineidentities.
13.3.20 A number of equations relate the two types of Chebyshev polynomials. As examples
showthat
Tn(x)=Un(x)−xUn−1(x)
and
parenleftbig
1−x2parenrightbig
Un(x)=xTn+1(x)−Tn+2(x).
13.3.21 Showthat
dVn(x)
dx=−nTn(x)√
1−x2
(a) usingthetrigonometricformsof VnandTn,
(b) usingtheRodriguesrepresentation.
858 Chapter 13 More Special Functions
13.3.22 Startingwith x=cosθandTn(cosθ)=cosnθ,expand
xk=parenleftbiggeiθ+e−iθ
2parenrightbiggk
andshowthat
xk=1
2k−1bracketleftbigg
Tk(x)+parenleftbiggk
1parenrightbigg
Tk−2(x)+parenleftbiggk
2parenrightbigg
Tk−4+···bracketrightbigg
,
theseriesinbracketsterminatingwithparenleftbigk
mparenrightbig
T1(x)fork=2m+1or1
2parenleftbigk
mparenrightbig
T0fork=2m.
13.3.23 (a) Calculate and tabulate the Chebyshev functions V1(x),V2(x), andV3(x)forx=
−1.0(0.1)1.0.
(b) A second solution of the Chebyshev differential equation, Eq. (13.100), for
n=0i sy(x)=sin−1x. Tabulate and plot this function over the same range:
−1.0(0.1)1.0.
13.3.24 Write acomputerprogramthatwillgeneratethecoefficients asinthepolynomialform
oftheChebyshevpolynomial Tn(x)=summationtextn
s=0asxs.
13.3.25 Tabulate T10(x)for 0.00(0.01)1.00. This will include the five positive roots of T10.I f
graphicssoftware isavailable,plotyourresults.
13.3.26 Determine the five positive roots of T10(x)by calling a root-finding subroutine. Use
your knowledge of the approximate location of these roots from Exercise 13.3.25 or
writeasearchroutinetolookfortheroots.Thesefivepositiveroots(andtheirnegatives)
aretheevaluationpointsofthe10-pointGauss–Chebyshevquadraturemethod.
Checkvalues. xk=cosbracketleftbig
(2k−1)π/20bracketrightbig
,k=1,2,3,4,5.
13.3.27 DevelopthefollowingChebyshevexpansions(for [−1,1]):
(a)parenleftbig
1−x2parenrightbig1/2=2
πbracketleftbigg
1−2∞summationdisplay
s=1parenleftbig
4s2−1parenrightbig−1T2s(x)bracketrightbigg
.
(b)+1,0<x≤1
−1,−1≤x<0bracerightbigg
=4
π∞summationdisplay
s=0(−1)s(2s+1)−1T2s+1(x).
13.3.28 (a) Fortheinterval [−1,1]showthat
|x|=1
2+∞summationdisplay
s=1(−1)s+1(2s−3)!!
(2s+2)!!(4s+1)P2s(x)
=2
π+4
π∞summationdisplay
s=1(−1)s+11
4s2−1T2s(x).
(b) Showthattheratioofthecoefficientof T2s(x)tothatof P2s(x)approaches (πs)−1
ass→∞. This illustrates the relatively rapid convergence of the Chebyshev se-
ries.
13.4 Hypergeometric Functions 859
Hint. With the Legendre recurrence relations, rewrite xPn(x)as a linear combination
ofderivatives.Thetrigonometricsubstitution x=cosθ,Tn(x)=cosnθismosthelpful
for theChebyshevpart.
13.3.29 Showthat
π2
8=1+2∞summationdisplay
s=1parenleftbig
4s2−1parenrightbig−2.
Hint. Apply Parseval’s identity (or the completeness relation) to the results of Exer-
cise13.3.28.
13.3.30 Showthat
(a) cos−1x=π
2−4
π∞summationdisplay
n=01
(2n+1)2T2n+1(x).
(b) sin−1x=4
π∞summationdisplay
n=01
(2n+1)2T2n+1(x).
13.4 H YPERGEOMETRIC FUNCTIONS
InChapter9thehypergeometricequation10
x(1−x)y′′(x)+bracketleftbig
c−(a+b+1)xbracketrightbig
y′(x)−aby(x)=0 (13.124)
wasintroducedasacanonicalformofalinearsecond-orderODEwithregularsingularities
atx=0,1,and∞. Onesolutionis
y(x)=2F1(a,b;c;x)
=1+a·b
cx
1!+a(a+1)b(b+1)
c(c+1)x2
2!+···,c/negationslash=0,−1,−2,−3,...,
(13.125)
which is known as the hypergeometric function orhypergeometric series . The range of
convergencefor c>a+bis|x|<1 andx=1,andisx=−1f o rc>a+b−1.Interms
oftheoften-usedPochhammersymbol,
(a)n=a(a+1)(a+2)···(a+n−1)=(a+n−1)!
(a−1)!,
(a)0=1, (13.126)
thehypergeometricfunctionbecomes
2F1(a,b;c;x)=∞summationdisplay
n=0(a)n(b)n
(c)nxn
n!. (13.127)
10This is sometimes calledGauss’ODE.Thesolutions thenbecome Gauss functions.
860 Chapter 13 More Special Functions
In this form the subscripts 2 and 1 become clear. The leading subscript 2 indicates that
two Pochhammer symbols appear in the numerator and the final subscript 1 indicates one
Pochhammer symbol in the denominator.11(The confluent hypergeometric function 1F1
with one Pochhammer symbol in the numerator and one in the denominator appears in
Section13.5.)
FromtheformofEq.(13.125)weseethattheparameter cmaynotbezerooranegative
integer.Ontheotherhand,if aorbequals0oranegativeinteger,theseriesterminatesand
the hypergeometric function becomes a polynomial. Many more or less elementary func-
tionscanberepresentedbythehypergeometricfunction.12Comparingthepowerserieswe
verifythat
ln(1+x)=x2F1(1,1;2;−x). (13.128)
Forthecompleteellipticintegrals KandE,
Kparenleftbig
k2parenrightbig
=integraldisplayπ/2
0parenleftbig
1−k2sin2θparenrightbig−1/2dθ=π
22F1parenleftbigg1
2,1
2;1;k2parenrightbigg
,(13.129)
Eparenleftbig
k2parenrightbig
=integraldisplayπ/2
0parenleftbig
1−k2sinθparenrightbig1/2dθ=π
22F1parenleftbigg1
2,−1
2;1;k2parenrightbigg
.(13.130)
The explicit series forms and other properties of the elliptic integrals are developed in
Section5.8.
The hypergeometric equation as a second-order linear ODE has a second independent
solution.Theusualform is
y(x)=x1−c2F1(a+1−c,b+1−c;2−c;x), c/negationslash=2,3,4,.... (13.131)
Ifcis an integer either the two solutions coincide or (barring a rescue by integral aor
integralb) one of the solutions will blow up (see Exercise 13.4.1). In such a case the
secondsolutionis expectedtoincludealogarithmicterm.
Alternateforms ofthehypergeometricODEinclude
parenleftbig
1−z2parenrightbigd2
dz2yparenleftbigg1−z
2parenrightbigg
−bracketleftbig
(a+b+1)z−(a+b+1−2c)bracketrightbigd
dzyparenleftbigg1−z
2parenrightbigg
−abyparenleftbigg1−z
2parenrightbigg
=0, (13.132)
parenleftbig
1−z2parenrightbigd2
dz2y(z2)−bracketleftbigg
(2a+2b+1)z+1−2c
zbracketrightbiggd
dzyparenleftbig
z2parenrightbig
−4abyparenleftbig
z2parenrightbig
=0.(13.133)
11ThePochhammer symbol is often useful in other expressions involving factorials, for instance,
(1−z)−a=∞summationdisplay
n=0(a)nzn/n!,|z|<1.
12With threeparameters, a,b,a n dc,wecan represent almost anything.
13.4 Hypergeometric Functions 861
Contiguous Function Relations
The parameters a,b, andcenter in the same way as the parameter nof Bessel, Legendre,
and other special functions. As we found with these functions, we expect recurrence rela-
tions involving unit changes in the parameters a,b, andc. The usual nomenclature for the
hypergeometric functions, in which one parameter changes by +or−1, is a “contiguous
function.” Generalizing this term to include simultaneous unit changes in more than one
parameter, we find 26 functions contiguous to 2F1(a,b;c;x). Taking them two at a time,
wecandeveloptheformidabletotalof325equationsamongthecontiguousfunctions.One
typicalexampleis
(a−b)braceleftbig
c(a+b−1)+1−a2−b2+bracketleftbig
(a−b)2−1bracketrightbig
(1−x)bracerightbig
2F1(a,b;c;x)
=(c−a)(a−b+1)b2F1(a−1,b+1;c;x)
+(c−b)(a−b−1)a2F1(a+1,b−1;c;x). (13.134)
AnothercontiguousfunctionrelationappearsinExercise13.4.10.
Hypergeometric Representations
Sincetheultrasphericalequation(13.112)inSection13.3isaspecialcaseofEq.(13.124),
we see that ultraspherical functions (and Legendre and Chebyshev functions) may be ex-
pressedashypergeometricfunctions.Fortheultrasphericalfunctionweobtain
Cβ
n(x)=(n+2β)!
2βn!β!2F1parenleftbigg
−n,n+2β+1;1+β;1−x
2parenrightbigg
(13.135)
upon comparing its ODE with Eq. (13.124) and the power-series solutions. For Legendre
andassociatedLegendrefunctionswefindsimilarly
Pn(x)=2F1parenleftbigg
−n,n+1;1;1−x
2parenrightbigg
, (13.136)
Pm
n(x)=(n+m)!
(n−m)!(1−x2)m/2
2mm!2F1parenleftbigg
m−n,m+n+1;m+1;1−x
2parenrightbigg
.(13.137)
Alternateformsare
P2n(x)=(−1)n(2n)!
22nn!n!2F1parenleftbigg
−n,n+1
2;1
2;x2parenrightbigg
=(−1)n(2n−1)!!
(2n)!!2F1parenleftbigg
−n,n+1
2;1
2;x2parenrightbigg
, (13.138)
P2n+1(x)=(−1)n(2n+1)!
22nn!n!2F1parenleftbigg
−n,n+3
2;3
2;x2parenrightbigg
x
=(−1)n(2n+1)!!
(2n)!!2F1parenleftbigg
−n,n+3
2;3
2;x2parenrightbigg
x. (13.139)
862 Chapter 13 More Special Functions
Intermsof hypergeometricfunctions,theChebyshevfunctionsbecome
Tn(x)=2F1parenleftbigg
−n,n;1
2;1−x
2parenrightbigg
, (13.140)
Un(x)=(n+1)2F1parenleftbigg
−n,n+2;3
2;1−x
2parenrightbigg
, (13.141)
Vn(x)=nradicalbig
1−x22F1parenleftbigg
−n+1,n+1;3
2;1−x
2parenrightbigg
. (13.142)
The leading factors are determined by direct comparison of complete power series, com-
parisonofcoefficientsofparticularpowersofthevariable,orevaluationat x=0 or1,and
soon.
Thehypergeometricseriesmaybeusedtodefinefunctionswithnonintegralindices.The
physicalapplicationsare minimal.
Exercises
13.4.1 (a) For c,aninteger,and aandbnonintegral,showthat
2F1(a,b;c;x)andx1−c2F1(a+1−c,b+1−c;2−c;x)
yieldonlyonesolutiontothehypergeometricequation.
(b) Whathappensif aisaninteger,say, a=−1,andc=−2?
13.4.2 Find the Legendre, Chebyshev I, and Chebyshev II recurrence relations corresponding
tothecontiguoushypergeometricfunctionEq.(13.134).
13.4.3 Transformthefollowingpolynomialsintohypergeometricfunctionsofargument x2.(a)
T2n(x);( b )x−1T2n+1(x);( c )U2n(x);( d)x−1U2n+1(x).
ANS.(a) T2n(x)=(−1)n2F1(−n,n;1
2;x2).
(b)x−1T2n+1(x)=(−1)n(2n+1)2F1(−n,n+1;3
2;x2).
(c)U2n(x)=(−1)n2F1(−n,n+1;1
2;x2).
(d)x−1U2n+1(x)=(−1)n(2n+2)2F1(−n,n+2;3
2;x2).
13.4.4 Derive or verify the leading factor in the hypergeometric representations of the Cheby-
shevfunctions.
13.4.5 VerifythattheLegendrefunctionofthesecondkind, Qν(z), is givenby
Qν(z)=π1/2ν!
(ν+1
2)!(2z)ν+12F1parenleftbiggν
2+1
2,ν
2+1;ν
2+3
2;z−2parenrightbigg
,
|z|>1,|argz|<π, ν/negationslash=−1,−2,−3,....
13.4.6 Analogoustotheincompletegammafunction,wemaydefineanincompletebetafunc-
tionby
Bx(a,b)=integraldisplayx
0ta−1(1−t)b−1dt.
13.5 Confluent Hypergeometric Functions 863
Showthat
Bx(a,b)=a−1xa2F1(a,1−b;a+1;x).
13.4.7 Verifytheintegralrepresentation
2F1(a,b;c;z)=Ŵ(c)
Ŵ(b)Ŵ(c−b)integraldisplay1
0tb−1(1−t)c−b−1(1−tz)−adt.
Whatrestrictionsmustbeplacedontheparameters bandcandonthevariable z?
Note. The restriction on |z|can be dropped—analytic continuation. For nonintegral a
therealaxisinthe z-planefrom1to ∞isacutline.
Hint.The integralis suspiciouslylikea betafunctionandcanbe expandedintoa series
ofbetafunctions.
ANS.ℜ(c)>ℜ(b)>0,and|z|<1.
13.4.8 Provethat
2F1(a,b;c;1)=Ŵ(c)Ŵ(c−a−b)
Ŵ(c−a)Ŵ(c−b),c/negationslash=0,−1,−2,... c>a +b.
Hint.Hereis achancetousetheintegralrepresentation,Exercise13.4.7.
13.4.9 Provethat
2F1(a,b;c;x)=(1−x)−a2F1parenleftbigg
a,c−b;c;−x
1−xparenrightbigg
.
Hint.Tryanintegralrepresentation.
Note.ThisrelationisusefulindevelopingaRodriguesrepresentationof Tn(x)(compare
Exercise13.3.11).
13.4.10 Verifythat
2F1(−n,b;c;1)=(c−b)n
(c)n.
Hint. Here is a chance to use the contiguous function relation [2a−c+(b−a)x]·
2F1(a,b;c;x)=a(1−x)2F1(a+1,b;c;x)−(c−a)2F1(a−1,b;c;x)and mathe-
maticalinduction.Alternatively,usetheintegralrepresentationandthebetafunction.
13.5 C ONFLUENT HYPERGEOMETRIC FUNCTIONS
Theconfluenthypergeometricequation13
xy′′(x)+(c−x)y′(x)−ay(x)=0 (13.143)
13This is often called Kummer’s equation . The solutions, then, are Kummerfunctions .
864 Chapter 13 More Special Functions
hasaregularsingularityat x=0andanirregularoneat x=∞.Itisobtainedfromthehy-
pergeometricequationofSection13.4bymerging(byhand: x(1−x)→xinEq.(13.124))
two of the latter’s three singularities. One solution of the confluent hypergeometric equa-
tionis
y(x)=1F1(a;c;x)=M(a,c,x)
=1+a
cx
1!+a(a+1)
c(c+1)x2
2!+···,c/negationslash=0,−1,−2,.... (13.144)
This solution is convergent for all finite x(or complex z). In terms of the Pochhammer
symbols,wehave
M(a,c,x)=∞summationdisplay
n=0(a)n
(c)nxn
n!. (13.145)
Clearly,M(a,c,x) becomes a polynomial if the parameter ais 0 or a negative integer.
Numerous more or less elementary functions may be represented by the confluent hyper-
geometric function. Examples are the error function and the incomplete gamma function
(from Eq. (8.69)):
erf(x)=2
π1/2integraldisplayx
0e−t2dt=2
π1/2xMparenleftbigg1
2,3
2,−x2parenrightbigg
, (13.146)
γ(a,x)=integraldisplayx
0e−tta−1dt=a−1xaM(a,a+1,−x),ℜ(a)>0.(13.147)
Clearly, this coincides with the first solution for c=a. The error function and the incom-
pletegammafunctionarediscussedfurther inSection8.5.
AsecondsolutionofEq. (13.143)isgivenby
y(x)=x1−cM(a+1−c,2−c,x), c/negationslash=2,3,4,.... (13.148)
The standard form of the second solution of Eq. (13.143) is a linear combination of
Eqs. (13.144)and(13.148):
U(a,c,x)=π
sinπcbracketleftbiggM(a,c,x)
(a−c)!(c−1)!−x1−cM(a+1−c,2−c,x)
(a−1)!(1−c)!bracketrightbigg
.(13.149)
Note the resemblance to our definition of the Neumann function, Eq. (11.60). As with our
Neumannfunction,Eq.(11.60),thisdefinitionof U(a,c,x) becomesindeterminateinthis
caseforcaninteger.
An alternate form of the confluent hypergeometric equation that will be useful later is
obtainedbychangingtheindependentvariablefrom xtox2:
d2
dx2yparenleftbig
x2parenrightbig
+bracketleftbigg2c−1
x−2xbracketrightbiggd
dxyparenleftbig
x2parenrightbig
−4ayparenleftbig
x2parenrightbig
=0. (13.150)
As with the hypergeometric functions, contiguous functions exist in which the para-
metersaandcare changed by ±1. Including the cases of simultaneous changes in the
13.5 Confluent Hypergeometric Functions 865
twoparameters,14wehaveeightpossibilities.Takingtheoriginalfunctionandpairsofthe
contiguousfunctions,wecandevelopatotalof 28equations.15
Integral Representations
Itisfrequentlyconvenienttohavetheconfluenthypergeometricfunctionsinintegralform.
Wefind(Exercise13.5.10)
M(a,c,x)=Ŵ(c)
Ŵ(a)Ŵ(c−a)integraldisplay1
0extta−1(1−t)c−a−1dt,ℜ(c)>ℜ(a)>0,
(13.151)
U(a,c,x)=1
Ŵ(a)integraldisplay∞
0e−xtta−1(1+t)c−a−1dt,ℜ(x)>0,ℜ(a)>0.
(13.152)
Three important techniques for deriving or verifying integral representations are as fol-
lows:
1. TransformationofgeneratingfunctionexpansionsandRodriguesrepresentations:The
Bessel andLegendrefunctionsprovideexamplesof thisapproach.
2. Directintegrationtoyieldaseries:ThisdirecttechniqueisusefulforaBesselfunction
representation(Exercise11.1.18)andahypergeometricintegral(Exercise13.4.7).
3. (a) Verification that the integral representation satisfies the ODE. (b) Exclusion of
the other solution. (c) Verification of normalization. This is the method used in Sec-
tion11.5toestablishanintegralrepresentationofthemodifiedBesselfunction Kν(z).
It willwork heretoestablishEqs. (13.151)and(13.152).
Bessel and Modified Bessel Functions
Kummer’sfirstformula,
M(a,c,x)=exM(c−a,c,−x), (13.153)
isusefulinrepresentingtheBesselandmodifiedBesselfunctions.Theformulamaybever-
ifiedbyseriesexpansionorbyuseofanintegralrepresentation(compareExercise13.5.10).
As expected from the form of the confluent hypergeometric equation and the character
of its singularities, the confluent hypergeometric functions are useful in representing a
numberofthespecialfunctionsofmathematicalphysics.FortheBesselfunctions,
Jν(x)=e−ix
ν!parenleftbiggx
2parenrightbiggν
Mparenleftbigg
ν+1
2,2ν+1,2ixparenrightbigg
, (13.154)
whereasfor themodifiedBessel functionsof thefirst kind,
Iν(x)=e−x
ν!parenleftbiggx
2parenrightbiggν
Mparenleftbigg
ν+1
2,2ν+1,2xparenrightbigg
. (13.155)
14Slaterrefers tothese as associated functions .
15Therecurrence relations for Bessel,Hermite,and Laguerrefunctions arespecial casesofthese equations.
866 Chapter 13 More Special Functions
Hermite Functions
TheHermitefunctionsaregivenby
H2n(x)=(−1)n(2n)!
n!Mparenleftbigg
−n,1
2,x2parenrightbigg
, (13.156)
H2n+1(x)=(−1)n2(2n+1)!
n!xMparenleftbigg
−n,3
2,x2parenrightbigg
, (13.157)
usingEq. (13.150).
ComparingtheLaguerreODEwiththeconfluenthypergeometricequation(13.143),we
have
Ln(x)=M(−n,1,x). (13.158)
TheconstantisfixedasunitybynotingEq.(13.66)for x=0.FortheassociatedLaguerre
functions,
Lm
n(x)=(−1)mdm
dxmLn+m(x)=(n+m)!
n!m!M(−n,m+1,x). (13.159)
Alternate verification is obtained by comparing Eq. (13.159) with the power-series so-
lution(Eq.(13.72)ofSection13.2).Notethatinthehypergeometricform,asdistinctfrom
a Rodrigues representation, the indices nandmneed not be integers, and, if they are not
integers,Lm
n(x)willnotbeapolynomial.
Miscellaneous Cases
There are certain advantages in expressing our special functions in terms of hypergeomet-
ricandconfluenthypergeometricfunctions.Ifthegeneralbehaviorofthelatterfunctionsis
known,thebehaviorofthespecialfunctionswehaveinvestigatedfollowsasaseriesofspe-
cialcases.Thismaybeusefulindeterminingasymptoticbehaviororevaluatingnormaliza-
tion integrals. The asymptotic behavior of M(a,c,x) andU(a,c,x) may be conveniently
obtained from integral representations of these functions, Eqs. (13.151) and (13.152). The
further advantage is that the relations between the special functions are clarified. For in-
stance,anexaminationofEqs.(13.156),(13.157),and(13.159)suggeststhattheLaguerre
andHermitefunctionsarerelated.
The confluent hypergeometric equation (13.143) is clearly not self-adjoint. For this and
otherreasons itisconvenienttodefine
Mkµ(x)=e−x/2xµ+1/2Mparenleftbigg
µ−k+1
2,2µ+1,xparenrightbigg
. (13.160)
Thisnewfunction, Mkµ(x), is aWhittakerfunctionthatsatisfiestheself-adjointequation
M′′
kµ(x)+parenleftbigg
−1
4+k
x+1
4−µ2
x2parenrightbigg
Mkµ(x)=0. (13.161)
Thecorrespondingsecondsolutionis
Wkµ(x)=e−x/2xµ+1/2Uparenleftbigg
µ−k+1
2,2µ+1,xparenrightbigg
. (13.162)
13.5 Confluent Hypergeometric Functions 867
Exercises
13.5.1 Verifytheconfluenthypergeometricrepresentationoftheerrorfunction
erf(x)=2x
π1/2Mparenleftbigg1
2,3
2,−x2parenrightbigg
.
13.5.2 Show that the Fresnel integrals C(x)andS(x)of Exercise 5.10.2 may be expressed in
termsoftheconfluenthypergeometricfunctionas
C(x)+iS(x)=xMparenleftbigg1
2,3
2,iπx2
2parenrightbigg
.
13.5.3 Bydirectdifferentiationandsubstitutionverifythat
y=ax−aintegraldisplayx
0e−tta−1dt=ax−aγ(a,x)
satisfies
xy′′+(a+1+x)y′+ay=0.
13.5.4 ShowthatthemodifiedBessel functionofthesecondkind, Kν(x), isgivenby
Kν(x)=π1/2e−x(2x)νUparenleftbigg
ν+1
2,2ν+1,2xparenrightbigg
.
13.5.5 Show that the cosine and sine integrals of Section 8.5 may be expressed in terms of
confluenthypergeometricfunctionsas
Ci(x)+isi(x)=−eixU(1,1,−ix).
ThisrelationisusefulinnumericalcomputationofCi (x)andsi(x)forlargevaluesof x.
13.5.6 Verify the confluent hypergeometric form of the Hermite polynomial H2n+1(x)
(Eq. (13.157))byshowingthat
(a)H2n+1(x)/xsatisfies the confluent hypergeometric equation with a=−n,c=3
2
andargument x2,
(b) lim
x→0H2n+1(x)
x=(−1)n2(2n+1)!
n!.
13.5.7 Showthatthecontiguousconfluenthypergeometricfunctionequation
(c−a)M(a−1,c,x)+(2a−c+x)M(a,c,x)−aM(a+1,c,x)=0
leadstotheassociatedLaguerrefunctionrecurrencerelation(Eq. (13.75)).
13.5.8 VerifytheKummertransformations:
(a)M(a,c,x)=exM(c−a,c,−x)
(b)U(a,c,x)=x1−cU(a−c+1,2−c,x).
868 Chapter 13 More Special Functions
13.5.9 Provethat
(a)dn
dxnM(a,c,x)=(a)n
(b)nM(a+n,b+n,x),
(b)dn
dxnU(a,c,x)=(−1)n(a)nU(a+n,c+n,x).
13.5.10 Verifythefollowingintegralrepresentations:
(a)M(a,c,x)=Ŵ(c)
Ŵ(a)Ŵ(c−a)integraldisplay1
0extta−1(1−t)c−a−1dt,ℜ(c)>ℜ(a)>0.
(b)U(a,c,x)=1
Ŵ(a)integraldisplay∞
0e−xtta−1(1+t)c−a−1dt,ℜ(x)>0,ℜ(a)>0.
Underwhatconditionscanyouaccept ℜ(x)=0 inpart(b)?
13.5.11 Fromtheintegralrepresentationof M(a,c,x) , Exercise13.5.10(a),showthat
M(a,c,x)=exM(c−a,c,−x).
Hint. Replace the variable of integration tby 1−sto release a factor exfrom the
integral.
13.5.12 Fromtheintegralrepresentationof U(a,c,x) ,Exercise13.5.10(b),showthattheexpo-
nentialintegralisgivenby
E1(x)=e−xU(1,1,x).
Hint.Replacethevariableofintegration tinE1(x)byx(1+s).
13.5.13 From the integral representations of M(a,c,x) andU(a,c,x) in Exercise 13.5.10 de-
velopasymptoticexpansionsof
(a)M(a,c,x) ,(b)U(a,c,x) .
Hint.Youcanusethetechniquethatwasemployedwith Kν(z), Section11.6.
ANS.(a)Ŵ(c)
Ŵ(a)ex
xc−abraceleftbigg
1+(1−a)(c−a)
1!x+
(1−a)(2−a)(c−a)(c−a+1)
2!x2+···bracerightbigg
(b)1
xabraceleftbigg
1+a(1+a−c)
1!(−x)+a(a+1)(1+a−c)(2+a−c)
2!(−x)2+···bracerightbigg
.
13.5.14 ShowthattheWronskianofthetwoconfluenthypergeometricfunctions M(a,c,x) and
U(a,c,x) is givenby
MU′−M′U=−(c−1)!
(a−1)!ex
xc.
Whathappensif ais0or anegativeinteger?
13.6 Mathieu Functions 869
13.5.15 The Coulomb wave equation (radial part of the Schrödinger equation with Coulomb
potential)is
d2y
dρ2+bracketleftbigg
1−2η
ρ−L(L+1)
ρ2bracketrightbigg
y=0.
Showthataregularsolution y=FL(η,ρ)isgivenby
FL(η,ρ)=CL(η)ρL+1e−iρM(L+1−iη,2L+2,2iρ).
13.5.16 (a) Showthattheradialpartofthehydrogenwavefunction,Eq.(13.81),maybewrit-
tenas
e−αr/2(αr)LL2L+1
n−L−1(αr)
=(n+L)!
(n−L−1)!(2L+1)!e−αr/2(αr)LM(L+1−n,2L+2,αr).
(b) Itwasassumedpreviouslythatthetotal(kinetic +potential)energy Eoftheelec-
tron was negative. Rewrite the (unnormalized) radial wave function for the free
electron, E>0.
ANS.eiαr/2(αr)LM(L+1−in,2L+2,−iαr), outgoing wave.
This representation provides a powerful alternative tech-
nique for the calculation of photoionization and recombina-
tioncoefficients.
13.5.17 Evaluate
(a)integraldisplay∞
0bracketleftbig
Mkµ(x)bracketrightbig2dx,(b)integraldisplay∞
0bracketleftbig
Mkµ(x)bracketrightbig2dx
x,
(c)integraldisplay∞
0bracketleftbig
Mkµ(x)bracketrightbig2dx
x1−a,
where 2µ=0,1,2,...,k−µ−1
2=0,1,2,...,a>−2µ−1.
ANS.(a) (2µ)!2k.(b)(2µ)!.( c)(2µ)!(2k)a.
13.6 M ATHIEU FUNCTIONS
When PDEs such as Laplace’s, Poisson’s, and the wave equation are solved with cylin-
drical or spherical boundary conditions by separating variables in polar coordinates, we
find radial solutions, which are the Bessel functions of Chapter 11, and angular solutions,
which are sin mϕ,cosmϕin cylindrical cases and spherical harmonics in spherical cases.
Examples are electromagnetic waves in resonant cavities, vibrating circular drumheads,
andcoaxialwaveguides.
When in such cylindrical problems the circular boundary condition becomes elliptical
we are led to the angular and radial Mathieu functions, which, therefore, might be called
elliptic cylinder functions. In fact, in 1868 Mathieu developed the leading terms of series
solutionsof thevibratingellipticaldrumhead,andWhittaker andothersin theearly 1900s
derivedhigher-orderterms aswell.
870 Chapter 13 More Special Functions
Here our goal is to give an introduction to the rich and complex properties of Mathieu
functions.
Separation of Variables in Elliptical Coordinates
Ellipticalcylinder coordinates ξ,η,z, whichare appropriate for ellipticalboundary condi-
tions,areexpressedinrectangularcoordinatesas
x=ccoshξcosη, y=csinhξsinη, z=z, (13.163)
0≤ξ<∞,0≤η≤2π,
where the parameter 2 c>0 is the distance between the foci of the confocal ellipses de-
scribed by these coordinates (Fig. 13.7). We want to show that in the limit c→0 the foci
of the ellipses coalesce to the center of circles. We work at constant z-coordinate mostly,
z=0,say.Indeedforfixedradialvariable ξ=const.wecaneliminatetheangularvariable
ηtoobtainfromEq. (13.163)
x2
c2cosh2ξ+y2
c2sinh2ξ=1, (13.164)
describing confocal ellipses centered at the origin of the x,y-plane with major and minor
half-axes
a=ccoshξ, b=csinhξ, (13.165)
respectively.Since
b
a=tanhξ=radicalBigg
1−1
cosh2ξ≡radicalbig
1−e2, (13.166)
the eccentricity e=1/coshξof the ellipse with 0 ≤e≤1, and the distance between the
foci 2ae=2c, providing a geometrical interpretation of the radial coordinate ξand the
FIGURE 13.7Ellipticalcoordinates ξ,η.
13.6 Mathieu Functions 871
parameter c.A sξ→∞,e→0 and the ellipses become circles, which is indicated in
Fig. 13.7. As ξ→0, the ellipse becomes more elongated until, at ξ=0, it has shrunk to
thelinesegmentbetweenthefoci.
Whenη=const.weeliminate ξtofindconfocalhyperbolas
x2
c2cos2η−y2
c2sin2η=1, (13.167)
whicharealsoplottedinFig.13.7.Differentiatingtheellipse,weobtain
xdx
cosh2ξ+ydy
sinh2ξ=0, (13.168)
which means that the tangent vector (dx,dy)of the ellipse is perpendicular to the vector
(x
cosh2ξ,y
sinh2ξ). Forthehyperbolatheorthogonalityconditionis
xdx
cos2η−ydy
sin2η=0, (13.169)
so the scalar product of the ellipse and hyperbola tangent vectors at each of their intersec-
tionpoints (x,y)of Eq.(13.163)obey
x2
cosh2ξcos2η−y2
sinh2ξsin2η=c2−c2=0. (13.170)
This means that these confocal ellipses and hyperbolas form an orthogonal coordinate
system,inthesenseofSection2.1.Toextractthescalefactors hξ,hηfromthedifferentials
oftheellipticalcoordinates
dx=csinhξcosηdξ−ccoshξsinηdη,
(13.171)
dy=ccoshξsinηdξ+csinhξcosηdη,
wesumtheirsquares, finding
dx2+dy2=c2parenleftbig
sinh2ξcos2η+cosh2ξsin2ηparenrightbigparenleftbig
dξ2+dη2parenrightbig
=c2parenleftbig
cosh2ξ−cos2ηparenrightbigparenleftbig
dξ2+dη2parenrightbig
≡h2
ξdξ2+h2
ηdη2(13.172)
andyielding
hξ=hη=cparenleftbig
cosh2ξ−cos2ηparenrightbig1/2. (13.173)
Note that there is no cross term involving dξdη, showing again that we are dealing with
orthogonalcoordinates.
Nowweare readytoderiveMathieu’sdifferentialequations.
872 Chapter 13 More Special Functions
Example 13.6.1 ELLIPTICAL DRUM
We consider vibrations of an elliptical drumhead with vertical displacement z=z(x,y,t)
governedbythewaveequation
∂2z
∂x2+∂2z
∂y2=1
v2∂2z
∂t2, (13.174)
wherethevelocitysquared v2=T/ρwithtension Tandmassdensity ρisaconstant.We
firstseparatetheharmonictimedependence,writing
z(x,y,t)=u(x,y)w(t), (13.175)
wherew(t)=cos(ωt+δ), withωthe frequency and δa constant phase. Substituting this
functionzintoEq.(13.174)yields
1
uparenleftbigg∂2u
∂x2+∂2u
∂y2parenrightbigg
=1
v2w∂2w
∂t2=−ω2
v2=−k2=const., (13.176)
that is, the two-dimensional Helmholtz equation for the displacement u. We now use Eq.
(2.22) to convert the Laplacian ∇2to the elliptical coordinates, where we drop the z-
coordinate.Thisgives
∂2u
∂x2+∂2u
∂y2+k2u=1
h2
ξparenleftbigg∂2u
∂ξ2+∂2u
∂η2parenrightbigg
+k2u=0, (13.177)
thatis, theHelmholtzequationinelliptical ξ,ηcoordinates,
∂2u
∂ξ2+∂2u
∂η2+c2k2parenleftbig
cosh2ξ−cos2ηparenrightbig
u=0. (13.178)
Lastly,weseparate ξandη,writingu(ξ,η)=R(ξ)/Phi1(η) , whichyields
1
Rd2R
dξ2+c2k2cosh2ξ=c2k2cos2η−1
/Phi1d2/Phi1
dη2=λ+1
2c2k2, (13.179)
whereλ+c2k2/2 is the separation constant. Writing cosh2 ξ,cos2ηinstead of cosh2ξ,
cos2η(which motivates the special form of the separation constant in Eq. (13.179)) we
findthelinear,second-orderODE
d2R
dξ2−(λ−2qcosh2ξ)R(ξ)=0,q=1
4c2k2, (13.180)
whichisalsocalledthe radialMathieuequation ,and
d2/Phi1
dη2+(λ−2qcos2η)/Phi1(η)=0, (13.181)
theangular,o rmodified,Mathieu equation . Note that the eigenvalue λ(q)is a function
of the continuous parameter qin the Mathieu ODEs. It is this parameter dependence that
complicates the analysis of Mathieu functions and makes them among the most difficult
specialfunctionsusedinphysics. /squaresolid
13.6 Mathieu Functions 873
Clearly, all finite points are regular points of both ODEs, while infinity is an essential
singularity for both ODEs, which are of the Sturm–Liouville type (Chapter 10) with coef-
ficientfunctions p≡1 and
q(ξ)=−λ+2qcosh2ξ, q(η)=λ−2qcos2η. (13.182)
(These functions qmust not be confused with the parameter q.) As a consequence, their
solutionsformorthogonalsetsoffunctions.Thesubstitution η→iξtransformstheangular
totheradialMathieuODE,so theirsolutionsarecloselyrelated.
Using the Lindemann–Stieltjes substitution z=cos2η, dz/dη=−sin2η, the angular
Mathieu ODE is transformed into an ODE with coefficients that are algebraic in the vari-
ablez(usingd
dη=dz
dηd
dz=−sin2ηd
dzandd2
dη2=−2cos2ηd
dz+sin22ηd2
dz2):
4z(1−z)d2/Phi1
dz2+2(1−2z)d/Phi1
dz+bracketleftbig
λ+2q(1−2z)bracketrightbig
/Phi1=0. (13.183)
This ODE has regular singularities at z=0 andz=1, whereas the point at infinity is
an essential singularity (Chapter 9). By comparison, the hypergeometric ODE has three
regular singularities. But not all ODEs with two regular singularities and one essential
singularitycanbetransformedintoanODEof theMathieutype.
Example 13.6.2 THEQUANTUM PENDULUM
A planependulumof length landmass mwithgravitationalpotential V(θ)=−mglcosθ
iscalleda quantumpendulum if itswavefunction /Psi1obeystheSchrödingerequation
−¯h2
2ml2d2/Psi1
dθ2+bracketleftbig
V(θ)−Ebracketrightbig
/Psi1=0, (13.184)
where the variable θis the angular displacement from the vertical direction. (For fur-
ther details and illustrations we refer to Gutiérrez-Vega et al. in the Additional Readings.)
A boundary condition applies to /Psi1so as to be single-valued; that is, /Psi1(θ+2π)=/Psi1(θ).
Substituting
θ=2η, λ=8Eml2
¯h2,q=−4m2gl3
¯h2(13.185)
into the Schrödinger equation yields the angular Mathieu ODE for /Psi1(2(η+π))=
/Psi1(2η). /squaresolid
For many other applications involving Mathieu functions we refer to Ruby in the Addi-
tionalReadings.
Our main focus will be on the solutions of the angular Mathieu ODE, which has the
importantpropertythatitscoefficientfunctionisperiodicwithperiod π.
874 Chapter 13 More Special Functions
General Properties of Mathieu Functions
InphysicsapplicationstheangularMathieufunctionsarerequiredtobesingle-valued,that
is, periodic with period 2 π. Let us start with some nomenclature. Since Mathieu’s ODEs
are invariant under parity ( η→−η), Mathieu functions have definite parity. Those of odd
paritythathaveperiod2 πand,forsmall q,startwithsin (2n+1)ηarecalledse 2n+1(η,q),
withnan integer, n=0,1,2,...(se is short for sine-elliptic). Mathieu functions of odd
parity and period πthat start with sin2 nηfor small qare called se 2n(η,q), withn=
1,2,....Mathieu functions of even parity, period πthat start with cos2 nηfor small q
are called ce 2n(η,q)(ce is short for cosine-elliptic), while those with period 2 πthat start
withcos(2n+1)η,n=0,1,...,forsmall qarecalledce 2n+1(η,q).Inthelimitwherethe
parameter q→0 (and theMathieuODEbecomestheclassicalharmonicoscillatorODE),
Mathieufunctionsreducetothesetrigonometricfunctions.
The periodicity condition /Phi1(η+2π)=/Phi1(η)is sufficient to determine a set of eigen-
valuesλin terms of q. An elementaryanalogof this result is the fact that a solution of the
classicalharmonicoscillatorODE u′′(η)+λu(η)=0hasperiod2 πif,andonlyif, λ=n2
is the square of an integer. Such problems will be pursued in Section 14.7 as applications
ofFourierseries.
Example 13.6.3 RADIAL MATHIEU FUNCTIONS
Upon replacing the angular elliptic variable η→iξ, the angular Mathieu ODE,
Eq. (13.181), becomes the radial ODE, Eq. (13.180). This motivates the definitions of
radialMathieufunctionsas
Ce2n+p(ξ,q)=ce2n+p(iξ,q), p =0,1;n=0,1,...,
Se2n+p(ξ,q)=−ise2n+p(iξ,q), p =0,1;n=1,2,....
Because these functions are differentiable, they correspond to the regular solutions of the
radialMathieuODE.Ofcourse,theyarenolongerperiodicbutareoscillatory(Fig.13.8).
In physical problems involving elliptical coordinates, the radial Mathieu ODE,
Eq. (13.180), plays a role corresponding to Bessel’s ODE in cylindrical geometry. Be-
cause there are four families of independent Bessel functions—the regular solutions Jn
and irregular Neumann functions Nn, along with the modified Bessel functions Inand
Kn—we expect four kinds of radial Mathieu functions. Because of parity, the solutions
splitintoevenandoddMathieufunctionsandsothereareeightkinds. For q>0,
Je2n(ξ,q)=Ce2n(ξ,q), Je2n+1(ξ,q)=Ce2n+1(ξ,q),
Jo2n(ξ,q)=Se2n(ξ,q), Jo2n+1(ξ,q)=Se2n+1(ξ,q), regularorfirst kind ;
Nen(ξ,q),Non(ξ,q), irregularorsecondkind ;
forq<0,thesolutionsoftheradialMathieuODEaredenotedby
Ien(ξ,q),Ion(ξ,q), regularorfirst kind ,
Ken(ξ,q),Kon(ξ,q), irregularorsecondkind
13.6 Mathieu Functions 875
FIGURE 13.8RadialMathieufunctions: q=1 (solidline), q=2 (dashed
line),q=3 (dottedline).(FromGutiérrez-Vega etal.,Am. J.Phys. 71:
233(2003).)
and are known as the evanescent radial Mathieu functions . Mathieu functions corre-
sponding to the Hankel functions can be similarly defined. In Fig. 13.8 some of them are
plotted.
In applications such as a vibrating drumhead with elliptical boundary conditions (see
Example13.6.1),thesolutioncanbeexpandedinevenandoddMathieufunctions:
zen≡Jen(ξ,q)cen(η,q)cos(ωnt), m≥0,
zon≡Jon(ξ,q)sen(η,q)cos(ωnt), m≥1.
TheyobeyDirichletboundaryconditions, zen(ξ0,η,t)=0=zon(ξ0,η,t),whichholdpro-
vided the radial functions satisfy Je n(ξ0,q)=0=Jon(ξ0,q)at the elliptical boundary,
whereξ=ξ0.
Whenthefocaldistance c→0,theangularMathieufunctionsbecometheconventional
trigonometricfunctions,whiletheradialMathieufunctionsbecomeBesselfunctions.
In the case of oscillations of a confocal annular elliptic lake, the modes have to include
theMathieufunctionsofthesecondkindandare thusgivenby
zen≡bracketleftbig
AJen(ξ,q)+BNen(ξ,q)bracketrightbig
cen(η,q)cos(ωnt), m≥0,
zon≡bracketleftbig
AJon(ξ,q)+BNon(ξ,q)bracketrightbig
sen(η,q)cos(ωnt), m≥1,
withA,Bconstants. These standing wave solutions must obey Neumann boundary con-
ditions at the inner ( ξ=ξ0) and outer ( ξ=ξ1) elliptical boundaries; that is, the normal
876 Chapter 13 More Special Functions
derivatives (a prime denotes d/dξ)o fz enand zonvanish at each point of the boundaries.
For even modes, we have ze′
n(ξ0,η,t)=0=ze′
n(ξ1,η,t). The implied radial constraints
are similar to Eqs. (11.81) and (11.82) of Example 11.3.1. Numerical examples and plots,
alsofortravelingwaves,aregiveninGutiérrez-Vega etal.intheAdditionalReadings. /squaresolid
ForzerosofMathieufunctions,theirasymptoticexpansions,andamorecompletelisting
offormulaswerefertoAbramowitzandStegun(AMS-55)intheAdditionalReadings, Am.
J.Phys.71,JahnkeandEmdeandGradshteynandRyzhikintheAdditionalReadings.
To illustrate and support the nomenclature, we want to show16that there is an angular
Mathieufunctionthatis
•even inηandofperiod πifandonlyif /Phi1′
1(π/2)=0;
•oddandof period πifandonlyif /Phi12(π/2)=0;
•evenandofperiod 2 πif andonlyif /Phi11(π/2)=0;
•oddandof period 2 πifandonlyif /Phi1′
2(π/2)=0,
where/Phi11(η),/Phi12(η)are two linearly independent solutions of the angular Mathieu ODE
sothat
/Phi11(0)=1,/Phi1′
1(0)=0;/Phi12(0)=0,/Phi1′
2(0)=1. (13.186)
Since the Mathieu ODE is a linear second-order ODE, we know (Chapter 9) that these
initial conditions are realistic. The first case just given corresponds to ce 2n(η,q), with
/Phi1′
1(π/2)=−2nsin2nη|η=π/2+···=0f o rn=1,2,....The second is the se 2n(η,q),
with/Phi12(π/2)=sin2nη|π/2+···=0.Thethirdcaseisthece 2n+1(η,q),with/Phi11(π/2)=
cos(2n+1)π/2+···=0.Thefourthcaseisthese 2n+1(η,q).
The key to the proof is Floquet’s approach to linear second-order ODEs with periodic
coefficient functions, such as Mathieu’s angular ODE or the simple pendulum (Exercise
13.6.1). If /Phi11(η),/Phi12(η)are two linearly independent solutions of the ODE, any other
solution/Phi1canbeexpressedas
/Phi1(η)=c1/Phi11(η)+c2/Phi12(η), (13.187)
withconstants c1,c2.Now ,/Phi1k(η+2π)arealsosolutionsbecausesuchanODEisinvariant
underthetranslation η→η+2π, andinparticular
/Phi11(η+2π)=a1/Phi11(η)+a2/Phi12(η),
/Phi12(η+2π)=b1/Phi11(η)+b2/Phi12(η), (13.188)
withconstants ai,bj. SubstitutingEq.(13.188)intoEq. (13.187)weget
/Phi1(η+2π)=(c1a1+c2b1)/Phi11(η)+(c2b2+c1a2)/Phi12(η), (13.189)
wheretheconstants cicanbechosenassolutionsoftheeigenvalueequations
a1c1+b1c2=λc1,
a2c1+b2c2=λc2. (13.190)
16SeeHochstadt in the Additional Readings.
13.6 Mathieu Functions 877
ThenFloquet’stheorem statesthat /Phi1(η+2π)=λ/Phi1(η), whereλisarootof
vextendsinglevextendsinglevextendsinglevextendsinglea1−λb 1
a2b2−λvextendsinglevextendsinglevextendsinglevextendsingle=0. (13.191)
Au s e f u l corollary is obtained if we define µandybyλ=exp(2πµ)andy(η)=
exp(−µη)/Phi1(η),s o
y(η+2π)=e−µηe−2πµ/Phi1(η+2π)=e−µη/Phi1(η)=y(η). (13.192)
Thus,/Phi1(η)=eµηy(η), withyaperiodicfunctionof ηwithperiod 2 π.
LetusapplyFloquet’sargumenttothe /Phi1k(η+π),whicharealsosolutionsofMathieu’s
ODEbecausethelatteris invariantunderthetranslation
η→η+π. UsingthespecialvaluesinEq. (13.186)weknowthat
/Phi11(η+π)=/Phi11(π)/Phi11(η)+/Phi1′
1(π)/Phi12(η),
/Phi12(η+π)=/Phi12(π)/Phi11(η)+/Phi1′
2(π)/Phi12(η), (13.193)
because these linear combinations of /Phi1k(η)are solutions of Mathieu’s ODE with the cor-
rectvalues /Phi1i(η+π),/Phi1′
i(η+π)forη=0.Therefore,
/Phi1i(η+π)=λi/Phi1i(η), (13.194)
wherethe λiaretheroots of
vextendsinglevextendsinglevextendsinglevextendsingle/Phi11(π)−λ/Phi1 2(π)
/Phi1′
1(π) /Phi1′
2(π)−λvextendsinglevextendsinglevextendsinglevextendsingle=0. (13.195)
Theconstantterminthecharacteristicpolynomialis givenbytheWronskian
Wparenleftbig
/Phi11(η),/Phi12(η)parenrightbig
=C, (13.196)
aconstantbecausethecoefficientof d/Phi1/dηintheangularMathieuODEvanishes,imply-
ingdW/dη=0.In fact, usingEq. (13.186),
Wparenleftbig
/Phi11(0),/Phi12(0)parenrightbig
=/Phi11(0)/Phi1′
2(0)−/Phi1′
1(0)/Phi12(0)=1
=Wparenleftbig
/Phi11(π),/Phi12(π)parenrightbig
, (13.197)
sotheeigenvalueEq. (13.195)for λbecomes
parenleftbig
/Phi11(π)−λparenrightbigparenleftbig
/Phi1′
2(π)−λparenrightbig
−/Phi12(π)/Phi1′
1(π)=0
=λ2−bracketleftbig
/Phi11(π)+/Phi1′
2(π)bracketrightbig
λ+1, (13.198)
withλ1·λ2=1 andλ1+λ2=/Phi11(π)+/Phi1′
2(π).
If|λ1|=|λ2|=1,thenλ1=exp(iφ)andλ2=exp(−iφ),soλ1+λ2=2cosφ.Forφ/negationslash=
0,π,2π,...this case corresponds to |/Phi11(π)+/Phi1′
2(π)|<2, where both solutions remain
bounded as η→∞in steps of πusing Eq. (13.194). These cases do not yield periodic
Mathieu functions, and this is also the case when |/Phi11(π)+/Phi1′
2(π)|>2. Ifφ=0, that
is,λ1=1=λ2is a double root, then the /Phi1ihave period πand|/Phi11(π)+/Phi1′
2(π)|=2. If
φ=π, that is,λ1=−1=λ2is again a double root, then |/Phi11(π)+/Phi1′
2(π)|=−2 and the
/Phi1ihaveperiod 2 πwith/Phi1i(η+π)=−/Phi1i(η).
878 Chapter 13 More Special Functions
Because the angular Mathieu ODE is invariant under a parity transformation η→−η,
itisconvenienttoconsidersolutions
/Phi1e(η)=1
2bracketleftbig
/Phi1(η)+/Phi1(−η)bracketrightbig
,/Phi1 o(η)=1
2bracketleftbig
/Phi1(η)−/Phi1(−η)bracketrightbig
(13.199)
of definite parity, which obey the same initial conditions as /Phi1i. We now relabel /Phi1e→
/Phi11,/Phi1o→/Phi12, taking/Phi11to be even and /Phi12to be odd under parity. These solutions of
definite parity of Mathieu’s ODE are called Mathieu functions and are labeled according
toournomenclaturediscussedearlier.
If/Phi11(η)has period π, then/Phi1′
1(η+π)=/Phi1′
1(η)also has period πbut is odd under
parity.Substituting η=−π/2 weobtain
/Phi1′
1parenleftbiggπ
2parenrightbigg
=/Phi1′
1parenleftbigg
−π
2parenrightbigg
=−/Phi1′
1parenleftbiggπ
2parenrightbigg
,so/Phi1′
1parenleftbiggπ
2parenrightbigg
=0. (13.200)
Conversely,if /Phi1′
1(π/2)=0,then/Phi11(η)hasperiod π. Toseethis,weuse
/Phi11(η+π)=c1/Phi11(η)+c2/Phi12(η). (13.201)
This expansion is valid because /Phi11(η+π)is a solution of the angular Mathieu ODE. We
now determine the coefficients ci, settingη=−π/2, and recall that /Phi11and/Phi1′
2are even
underparity,whereas /Phi12and/Phi1′
1areodd.This yields
/Phi11parenleftbiggπ
2parenrightbigg
=c1/Phi11parenleftbiggπ
2parenrightbigg
−c2/Phi12parenleftbiggπ
2parenrightbigg
,
(13.202)
/Phi1′
1parenleftbiggπ
2parenrightbigg
=−c1/Phi1′
1parenleftbiggπ
2parenrightbigg
+c2/Phi1′
2parenleftbiggπ
2parenrightbigg
.
Since/Phi1′
1(π/2)=0,/Phi1′
2(π/2)/negationslash=0, or the Wronskian would vanish and /Phi12∼/Phi11would
follow. Hence c2=1 follows from the second equation and c1=1 from the first. Thus,
/Phi11(η+π)=/Phi11(η). Theotherbulletedcaseslistedearliercanbeprovedsimilarly.
Because the Mathieu ODEs are of the Sturm–Liouville type, Mathieu functions repre-
sent orthogonal systems of functions. So, for m,nnonnegative integers, the orthogonality
relationsandnormalizationsare
integraldisplayπ
−πcemcendη=integraldisplayπ
−πsemsendη=0,ifm/negationslash=n;
integraldisplayπ
−πcemsendη=0; (13.203)
integraldisplayπ
−π[ce2n]2dη=integraldisplayπ
−π[se2n]2dη=π,ifn≥1;integraldisplayπ
0bracketleftbig
ce0(η,q)bracketrightbig2dη=π.
If a function f(η)is periodic with period π, then it can be expanded in a series of orthog-
onalMathieufunctionsas
f(η)=1
2a0ce0(η,q)+∞summationdisplay
n=1bracketleftbig
ance2n(η,q)+bnse2n(η,q)bracketrightbig
(13.204)
13.6 Additional Readings 879
with
an=1
πintegraldisplayπ
−πf(η)ce2n(η,q)dη, n ≥0;
(13.205)
bn=1
πintegraldisplayπ
−πf(η)se2n(η,q)dη, n ≥1.
Similarexpansionsexistfor functionsof period 2 πintermsof ce 2n+1andse2n+1.
SeriesexpansionsofMathieufunctionswillbederivedinSection14.7.
Exercises
13.6.1 For the simple pendulum ODE of Section 5.8, apply Floquet’s method and derive the
propertiesofitssolutionssimilartothosemarkedbybulletsbeforeEq. (13.186).
13.6.2 Derive a Mathieu function analog for the Rayleigh expansion of a plane wave for
cos(kcosηcosθ)and sin(kcosηcosθ).
AdditionalReadings
Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions , Applied Mathematics Series-
55 (AMS-55). Washington, DC: National Bureau of Standards (1964). Paperback edition, New York: Dover
(1974). Chapter 22 is a detailed summary of the properties and representations of orthogonal polynomials.
Other chapters summarize properties of Bessel, Legendre, hypergeometric, and confluent hypergeometric
functions and much more.
Buchholz, H., The Confluent Hypergeometric Function . New York: Springer-Verlag (1953); translated (1969).
BuchholzstronglyemphasizestheWhittakerratherthantheKummerforms.Applicationstoavarietyofother
transcendental functions.
Erdelyi, A., W. Magnus, F. Oberhettinger, and F. G. Tricomi, Higher Transcendental Functions ,3v o l s .N e w
York: McGraw-Hill (1953). Reprinted Krieger (1981). A detailed, almost exhaustive listing of the properties
of thespecial functions of mathematicalphysics.
Fox,L.andI.B.Parker, ChebyshevPolynomialsinNumericalAnalysis .Oxford:OxfordUniversityPress(1968).
Adetailed,thorough,butveryreadableaccountofChebyshevpolynomialsandtheirapplicationsinnumerical
analysis.
Gradshteyn, I. S.,and I. M.Ryzhik, Table of Integrals, Series and Products , NewYork: AcademicPress (1980).
Gutiérrez-Vega, J. C., R. M. Rodríguez-Dagnino, M. A. Meneses-Nava and S. Chávez-Cerda, A m .J .P h y s . 71:
233 (2003).
Hochstadt, H., Special Functions of Mathematical Physics . New York: Holt, Rinehart and Winston (1961),
reprinted Dover (1986).
Jahnke,E.,andF.Emde, Table of Functions . Leipzig:Teubner (1933); NewYork: Dover (1943).
Lebedev,N.N., SpecialFunctionsandtheirApplications (translatedbyR.A.Silverman).EnglewoodCliffs,NJ:
Prentice-Hall (1965). Paperback, NewYork: Dover (1972).
Luke,Y .L., TheSpecialFunctionsandTheirApproximations .NewYork:AcademicPress(1969).Twovolumes:
Volume 1 is a thorough theoretical treatment of gamma functions, hypergeometric functions, confluent hy-
pergeometric functions, and related functions. Volume 2 develops approximations and other techniques for
numerical work.
Luke, Y. L., Mathematical Functions and Their Approximations . New York: Academic Press (1975). This is
an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs and Mathematical
Tables(AMS-55).
880 Chapter 13 More Special Functions
Mathieu,E., J.de Math. Pures etAppl. 13: 137–203 (1868).
McLachlan,N.W., Theory and Applications of Mathieu Functions . Oxford, UK:Clarendon Press (1947).
Magnus,W.,F.Oberhettinger,andR.P.Soni, FormulasandTheoremsfortheSpecialFunctionsofMathematical
Physics. NewYork: Springer (1966). An excellent summary of just what the title says, including the topics of
Chapters 10 to 13.
Rainville, E. D., Special Functions . New York: Macmillan (1960), reprinted Chelsea (1971). This book is a
coherent,comprehensiveaccountofalmostallthespecialfunctionsofmathematicalphysicsthatthereaderis
likelytoencounter.
Rowland, D.R., Am.J .Ph ys. 72: 758–766 (2004).
Ruby, L., Am.J .Ph ys. 64: 39–44 (1996).
Sansone, G., Orthogonal Functions (translated by A. H. Diamond). New York: Interscience (1959). Reprinted
Dover (1991).
Slater, L. J., Confluent Hypergeometric Functions . Cambridge, UK: Cambridge University Press (1960). This is
a clear and detailed development of the properties of the confluent hypergeometric functions and of relations
of the confluent hypergeometric equation to other ODEsofmathematical physics.
Sneddon, I. N., Special Functions of Mathematical Physicsand Chemistry , 3rd ed.NewYork: Longman (1980).
Whittaker,E.T.,andG.N.Watson, ACourseofModernAnalysis .Cambridge,UK:CambridgeUniversityPress,
reprinted (1997). Theclassic text on special functions andreal andcomplex analysis.
CHAPTER 14
FOURIER SERIES
14.1 G ENERAL PROPERTIES
Periodicphenomenainvolvingwaves,rotatingmachines(harmonicmotion),orotherrepet-
itive driving forces are described by periodic functions. Fourier series are a basic tool for
solving ordinary differential equations (ODEs) and partial differential equations (PDEs)
with periodic boundary conditions. Fourier integrals for nonperiodic phenomena are de-
velopedinChapter15.Thecommonnameforthefieldis Fourieranalysis .
A Fourier series is defined as an expansion of a function or representation of a function
inaseriesof sinesandcosines,suchas
f(x)=a0
2+∞summationdisplay
n=1ancosnx+∞summationdisplay
n=1bnsinnx. (14.1)
The coefficients a0,an, andbnare related to the periodic function f(x)by definite inte-
grals:
an=1
πintegraldisplay2π
0f(x)cosnxdx, (14.2)
bn=1
πintegraldisplay2π
0f(x)sinnxdx, n=0,1,2,..., (14.3)
which are subject to the requirement that the integrals exist. Notice that a0is singled out
for special treatment by the inclusion of the factor1
2. This is done so that Eq. (14.2) will
applytoall an,n=0a sw e l la s n>0.
The conditions imposed on f(x)to make Eq. (14.1) valid are that f(x)have only a
finitenumberoffinitediscontinuitiesandonlyafinitenumberofextremevalues,maxima,
and minima in the interval [0,2π].1Functions satisfying these conditions may be called
1Theseconditions are sufficient but notnecessary .
881
882 Chapter 14 Fourier Series
piecewise regular . The conditions themselves are known as the Dirichlet conditions. Al-
thoughtherearesomefunctionsthatdonotobeytheseDirichletconditions,theymaywell
be labeled pathological for purposes of Fourier expansions. In the vast majority of physi-
cal problems involving a Fourier series these conditions will be satisfied. In most physical
problemsweshallbeinterestedinfunctionsthataresquareintegrable(intheHilbertspace
L2of Section 10.4). In this space the sines and cosines form a complete orthogonal set.
AndthisinturnmeansthatEq. (14.1) isvalid,inthesenseofconvergenceinthemean.
Expressing cos nxand sinnxinexponentialform, wemayrewriteEq.(14.1) as
f(x)=∞summationdisplay
n=−∞cneinx, (14.4)
inwhich
cn=1
2(an−ibn), c−n=1
2(an+ibn), n> 0, (14.5a)
and
c0=1
2a0. (14.5b)
Complex Variables — Abel’s Theorem
Considerafunction f(z)representedbyaconvergentpowerseries
f(z)=∞summationdisplay
n=0Cnzn=∞summationdisplay
n=0Cnrneinθ. (14.6)
This is our Fourier exponential series, Eq. (14.4). Separating real and imaginary parts we
get
u(r,θ)=∞summationdisplay
n=0Cnrncosnθ, v(r,θ) =∞summationdisplay
n=1Cnrnsinnθ, (14.7a)
the Fourier cosine and sine series. Abel’s theorem asserts that if u(1,θ)andv(1,θ)are
convergentfor agiven θ,then
u(1,θ)+iv(1,θ)=lim
r→1fparenleftbig
reiθparenrightbig
. (14.7b)
Anapplicationof thisappearsasExercise14.1.9andinExample14.1.1.
Example 14.1.1 SUMMATION OF A FOURIER SERIES
Usually in this chapter we shall be concerned with finding the coefficients of the Fourier
expansion of a known function. Occasionally, we may wish to reverse this process and
determinethefunctionrepresentedbya givenFourierseries.
14.1 General Properties 883
Consider the seriessummationtext∞
n=1(1/n)cosnx,x∈(0,2π). Since this series is only condition-
allyconvergent(anddivergesat x=0),wetake
∞summationdisplay
n=1cosnx
n=lim
r→1∞summationdisplay
n=1rncosnx
n, (14.8)
absolutely convergent for |r|<1. Our procedure is to try forming power series by trans-
formingthetrigonometricfunctionsintoexponentialform:
∞summationdisplay
n=1rncosnx
n=1
2∞summationdisplay
n=1rneinx
n+1
2∞summationdisplay
n=1rne−inx
n. (14.9)
Now, these power series may be identified as Maclaurin expansions of −ln(1−z),z=
reix,re−ix(Eq. (5.95)), and
∞summationdisplay
n=1rncosnx
n=−1
2bracketleftbig
lnparenleftbig
1−reixparenrightbig
+lnparenleftbig
1−re−ixparenrightbigbracketrightbig
=−lnbracketleftbigparenleftbig
1+r2parenrightbig
−2rcosxbracketrightbig1/2. (14.10)
Lettingr=1 andusingAbel’stheorem,weseethat
∞summationdisplay
n=1cosnx
n=−ln(2−2cosx)1/2
=−lnparenleftbigg
2sinx
2parenrightbigg
,x∈(0,2π).2(14.11)
Bothsidesof thisexpressiondivergeas x→0 and 2π. /squaresolid
Completeness
The problem of establishing completeness may be approached in a number of different
ways. One way is to transform the trigonometric Fourier series into exponential form and
tocompareitwithaLaurentseries.Ifweexpand f(z)inaLaurentseries3(assuming f(z)
isanalytic),
f(z)=∞summationdisplay
n=−∞dnzn. (14.12)
Ontheunitcircle z=eiθand
f(z)=f(eiθ)=∞summationdisplay
n=−∞dneinθ. (14.13)
2Thelimits maybe shifted to [−π,π](andx/negationslash=0)using|x|on the right-hand side.
3Section6.5.
884 Chapter 14 Fourier Series
FIGURE 14.1Fourier
representationof sawtooth
wave.
TheLaurentexpansionontheunitcircle(Eq.(14.13))hasthesameformasthecomplex
Fourier series (Eq. (14.12)), which shows the equivalence between the two expansions.
SincetheLaurentseriesasapowerserieshasthepropertyofcompleteness,weseethatthe
Fourier functions einxform a complete set. There is a significant limitation here. Laurent
seriesandcomplexpowerseriescannothandlediscontinuitiessuchasasquarewaveorthe
sawtoothwaveofFig.14.1, exceptonthecircleof convergence.
Thetheoryofvectorspacesprovidesasecondapproachtothecompletenessofthesines
andcosines.HerecompletenessisestablishedbytheWeierstrasstheoremfortwovariables.
TheFourierexpansionandthecompletenesspropertymaybeexpected,forthefunctions
sinnx,cosnx,einxare alleigenfunctionsofaself-adjointlinearODE,
y′′+n2y=0. (14.14)
Weobtainorthogonaleigenfunctionsfordifferentvaluesoftheeigenvalue nfortheinterval
[0,2π]that satisfy the boundary conditions in the Sturm–Liouville theory (Chapter 10).
Differenteigenfunctionsfor thesameeigenvalue nareorthogonal.We have
integraldisplay2π
0sinmxsinnxdx=braceleftbiggπδmn,m/negationslash=0,
0,m=0,(14.15)
integraldisplay2π
0cosmxcosnxdx=braceleftbiggπδmn,m/negationslash=0,
2π, m=n=0,(14.16)
integraldisplay2π
0sinmxcosnxdx=0 forallintegral mandn. (14.17)
Note that any interval x0≤x≤x0+2πwill be equally satisfactory. Frequently, we shall
usex0=−πtoobtaintheinterval −π≤x≤π.Forthecomplexeigenfunctions e±inxor-
thogonalityisusually definedintermsofthecomplexconjugateofoneofthetwofactors,
integraldisplay2π
0parenleftbig
eimxparenrightbig∗einxdx=2πδmn. (14.18)
Thisagreeswiththetreatmentof thesphericalharmonics(Section12.6).
14.1 General Properties 885
Sturm–Liouville Theory
The Sturm–Liouville theory guarantees the validity of Eq. (14.1) (for functions satis-
fying the Dirichlet conditions) and, by use of the orthogonality relations, Eqs. (14.15),
(14.16), and (14.17), allows us to compute the expansion coefficients an,bn,a ss h o w ni n
Eqs. (14.2), and (14.3). Substituting Eqs. (14.2) and (14.3) into Eq. (14.1), we write our
Fourierexpansionas
f(x)=1
2πintegraldisplay2π
0f(t)dt
+1
π∞summationdisplay
n=1parenleftbigg
cosnxintegraldisplay2π
0f(t)cosntdt+sinnxintegraldisplay2π
0f(t)sinntdtparenrightbigg
=1
2πintegraldisplay2π
0f(t)dt+1
π∞summationdisplay
n=1integraldisplay2π
0f(t)cosn(t−x)dt, (14.19)
the first (constant) term being the average value of f(x)over the interval [0,2π]. Equa-
tion (14.19) offers one approach to the development of the Fourier integral and Fourier
transforms, Section15.1.
Another way of describing what we are doing here is to say that f(x)is part of an
infinite-dimensional Hilbert space, with the orthogonal cos nxand sinnxas the basis.
(They can always be renormalized to unity if desired.) The statement that cos nxand
sinnx (n=0,1,2,...)span this Hilbert space is equivalent to saying that they form a
completeset.Finally,theexpansioncoefficients anandbncorrespondtotheprojectionsof
f(x), with the integral inner products (Eqs. (14.2) and (14.3)) playing the role of the dot
productofSection1.3. Thesepointsare outlinedinSection10.4.
Example 14.1.2 SAWTOOTH WAVE
An idea of the convergence of a Fourier series and the error in using only a finite number
oftermsintheseries maybeobtainedbyconsideringtheexpansionof
f(x)=braceleftbiggx, 0≤x<π,
x−2π, π <x≤2π.(14.20)
This is a sawtooth wave, and for convenience we shall shift our interval from [0,2π]to
[−π,π]. In this interval we have f(x)=x. Using Eqs. (14.2) and (14.3), we show the
expansiontobe
f(x)=x=2bracketleftbigg
sinx−sin2x
2+sin3x
3−···+(−1)n+1sinnx
n+···bracketrightbigg
.(14.21)
Figure14.1shows f(x)for0≤x<πforthesumof4,6,and10termsoftheseries.Three
featuresdeservecomment.
1. Thereisasteadyincreaseintheaccuracyoftherepresentationasthenumberofterms
includedisincreased.
886 Chapter 14 Fourier Series
2. Allthecurvespass throughthemidpoint, f(x)=0,atx=π.
3. Inthevicinityof x=πthereisanovershootthatpersistsandshowsnosignofdimin-
ishing.
As a matter of incidental interest, setting x=π/2 in Eq. (14.21) provides an alternate
derivationofLeibniz’formula,Exercise5.7.6. /squaresolid
Behavior of Discontinuities
The behavior of the sawtooth wave f(x)atx=πis an example of a general rule that at
a finite discontinuity the series converges to the arithmetic mean. For a discontinuity at
x=x0theseries yields
f(x0)=1
2bracketleftbig
f(x0+0)+f(x0−0)bracketrightbig
, (14.22)
the arithmetic mean of the right and left approaches to x=x0. A general proof using
partial sums, as in Section 14.5, is given by Jeffreys and Jeffreys and by Carslaw (see the
Additional Readings). The proof may be simplified by the use of Dirac delta functions—
Exercise14.5.1.
The overshoot of the sawtooth wave just before x=πin Fig. 14.1 is an example of the
Gibbsphenomenon,discussedinSection14.5.
Exercises
14.1.1 Afunction f(x)(quadraticallyintegrable)istoberepresentedbya finiteFourierseries.
A convenient measure of the accuracy of the series is given by the integrated square of
thedeviation,
/Delta1p=integraldisplay2π
0bracketleftbigg
f(x)−a0
2−psummationdisplay
n=1(ancosnx+bnsinnx)bracketrightbigg2
dx.
Showthattherequirementthat /Delta1pbeminimized,thatis,
∂/Delta1p
∂an=0,∂/Delta1p
∂bn=0,
for alln,leadstochoosing anandbnas giveninEqs. (14.2)and(14.3).
Note. Your coefficients anandbnare independent of p. This independence is a con-
sequence of orthogonality and would not hold for powers of x, fitting a curve with
polynomials.
14.1.2 Intheanalysisofacomplexwaveform(oceantides,earthquakes,musicaltones,etc.)it
mightbemoreconvenienttohavetheFourierseries writtenas
f(x)=a0
2+∞summationdisplay
n=1αncos(nx−θn).
14.1 General Properties 887
ShowthatthisisequivalenttoEq. (14.1)with
an=αncosθn,α2
n=a2
n+b2
n,
bn=αnsinθn,tanθn=bn/an.
Note. The coefficients α2
nas a function of ndefine what is called the power spectrum .
Theimportanceof α2
nlies intheirinvarianceunderashiftinthephase θn.
14.1.3 Afunction f(x)is expandedinanexponentialFourierseries
f(x)=∞summationdisplay
n=−∞cneinx.
Iff(x)isreal,f(x)=f∗(x), whatrestrictionis imposedonthecoefficients cn?
14.1.4 Assumingthatintegraltextπ
−π[f(x)]2dxis finite,showthat
limm→∞am=0,limm→∞bm=0.
Hint.I n t e g r a t e[f(x)−sn(x)]2, wheresn(x)is thenth partial sum, and use Bessel’s
inequality, Section 10.4. For our finite interval the assumption that f(x)is square inte-
grable (integraltextπ
−π|f(x)|2dxis finite) implies thatintegraltextπ
−π|f(x)|dxis also finite. The converse
doesnothold.
14.1.5 Applythesummationtechniqueof thissectiontoshowthat
∞summationdisplay
n=1sinnx
n=braceleftBigg1
2(π−x), 0<x≤π
−1
2(π+x),−π≤x<0
(Fig.14.2).
FIGURE 14.2Reversesawtoothwave.
888 Chapter 14 Fourier Series
14.1.6 Sumthetrigonometricseries
∞summationdisplay
n=1(−1)n+1sinnx
n
andshowthatitequals x/2.
14.1.7 Sumthetrigonometricseries
∞summationdisplay
n=0sin(2n+1)x
2n+1
andshowthatitequals braceleftbiggπ/4,0<x<π
−π/4,−π<x<0.
14.1.8 Calculate the sum of the finite Fourier sine series for the sawtooth wave, f(x)=
x,(−π,π), Eq. (14.21). Use 4-, 6-, 8-, and 10-term series and x/π=0.00(0.02)1.00.
If aplottingroutineis available,plotyourresultsandcomparewithFig. 14.1.
14.1.9 Letf(z)=ln(1+z)=summationtext∞
n=1(−1)n+1zn/n. (This series converges to ln (1+z)for
|z|≤1,exceptatthepoint z=−1.)
(a) Fromtherealparts showthat
lnparenleftbigg
2cosθ
2parenrightbigg
=∞summationdisplay
n=1(−1)n+1cosnθ
n,−π<θ<π .
(b) Usingachangeof variable,transform part(a)into
−lnparenleftbigg
2sinθ
2parenrightbigg
=∞summationdisplay
n=1cosnθ
n,0<θ<2π.
14.2 A DVANTAGES ,USES OF FOURIER SERIES
Discontinuous Functions
OneoftheadvantagesofaFourierrepresentationoversomeotherrepresentation,suchasa
Taylorseries,isthatitcanrepresentadiscontinuousfunction.Anexampleisthesawtooth
wave in the preceding section. Other examples are considered in Section 14.3 and in the
exercises.
Periodic Functions
Related to this advantage is the usefulness of a Fourier series in representing a periodic
function.If f(x)hasaperiodof 2 π,perhapsitisonlynaturalthatweexpanditinaseries
of functions with period 2 π,2π/2,2π/3,.... This guarantees that if our periodic f(x)is
representedoveroneinterval [0,2π]or[−π,π], therepresentationholdsforallfinite x.
14.2 Advantages, Uses of Fourier Series 889
At this point we may conveniently consider the properties of symmetry. Using the in-
terval[−π,π],sinxis odd and cos xis an even function of x. Hence, by Eqs. (14.2) and
(14.3),4iff(x)is odd,all an=0 andiff(x)iseven,all bn=0.In otherwords,
f(x)=a0
2+∞summationdisplay
n=1ancosnx, f(x) even, (14.23)
f(x)=∞summationdisplay
n=1bnsinnx, f(x) odd, (14.24)
Frequentlythesepropertiesarehelpfulinexpandingagivenfunction.
We have noted that the Fourier series is periodic. This is important in considering
whetherEq.(14.1) holdsoutsidetheinitialinterval.Supposewearegivenonlythat
f(x)=x,0≤x<π (14.25)
and are asked to represent f(x)by a series expansion. Let us take three of the infinite
numberofpossibleexpansions.
1. If weassumeaTaylorexpansion,wehave
f(x)=x, (14.26)
aone-termseries. This(one-term)series isdefinedfor allfinite x.
2. Using the Fourier cosine series (Eq. (14.23)), thereby assuming the function is repre-
sented faithfully in the interval [0,π)and extended to neighboring intervals using the
knownsymmetryproperties,wepredictthat
f(x)=−x,−π<x≤0,
f(x)=2π−x, π <x< 2π.(14.27)
3. Finally,fromtheFouriersineseries (Eq. (14.24)), wehave
f(x)=x,−π<x≤0,
f(x)=x−2π, π <x< 2π.(14.28)
Thesethreepossibilities—Taylorseries,Fouriercosineseries,andFouriersineseries—
are each perfectly valid in the original interval, [0,π]. Outside, however, their behavior is
strikinglydifferent(compareFig.14.3).Whichofthethree,then,iscorrect?Thisquestion
has no answer, unless we are given more information about f(x). It may be any of the
three or none of them. Our Fourier expansionsare valid overthe basic interval. Unless the
functionf(x)isknowntobeperiodicwithaperiodequaltoourbasicintervalorto (1/n)th
ofourbasicinterval,thereisnoassurancewhateverthattherepresentation(Eq.(14.1))will
haveanymeaningoutsidethebasicinterval.
Inadditiontotheadvantagesofrepresentingdiscontinuousandperiodicfunctions,there
is a third very real advantage in using a Fourier series. Suppose that we are solving the
equationofmotionofanoscillatingparticlesubjecttoaperiodicdrivingforce.TheFourier
4With the range of integration −π≤x≤π.
890 Chapter 14 Fourier Series
FIGURE 14.3ComparisonofFouriercosineseries, Fouriersineseries,
andTaylorseries.
expansion of the driving force then gives us the fundamental term and a series of harmon-
ics. The (linear) ODE may be solved for each of these harmonics individually, a process
that may be much easier than dealing with the original driving force. Then, as long as the
ODEislinear,allthesolutionsmaybeaddedtogethertoobtainthefinalsolution.5This is
morethanjustaclevermathematicaltrick.
•It corresponds to finding the response of the system to the fundamental frequency and
toeachof theharmonicfrequencies.
Onequestionthatissometimesraisedis:“Weretheharmonicsthereallalong,orwerethey
createdbyourFourieranalysis?”Oneanswercomparesthefunctionalresolutionintohar-
monics with the resolution of a vector into rectangular components. The components may
have been present, in the sense that they may be isolated and observed, but the resolution
is certainly not unique. Hence many authors prefer to say that the harmonics were created
by our choice of expansion. Other expansions in other sets of orthogonal functions would
give different results. For further discussion we refer to a series of notes and letters in the
AmericanJournalofPhysics .6
Change of Interval
So far attention has been restricted to an interval of length 2 π. This restriction may easily
berelaxed.If f(x)isperiodicwithaperiod 2 L,wemaywrite
5Oneof the nastier features of nonlinear differential equations is that this principle of superposition is not valid.
6B.L.Robinson,Concerningfrequenciesresultingfromdistortion. Am.J.Phys. 21:391(1953);F.W.VanName,Jr.,Concerning
frequencies resulting from distortion. ibid.22: 94 (1954).
14.2 Advantages, Uses of Fourier Series 891
f(x)=a0
2+∞summationdisplay
n=1bracketleftbigg
ancosnπx
L+bnsinnπx
Lbracketrightbigg
, (14.29)
with
an=1
LintegraldisplayL
−Lf(t)cosnπt
Ldt, n=0,1,2,3,..., (14.30)
bn=1
LintegraldisplayL
−Lf(t)sinnπt
Ldt, n=1,2,3,..., (14.31)
replacing xin Eq. (14.1) with πx/Landtin Eqs. (14.2) and (14.3) with πt/L.( F o r
conveniencetheintervalinEqs.(14.2)and(14.3)isshiftedto −π≤t≤π.)Thechoiceof
thesymmetricinterval (−L,L)isnotessential.For f(x)periodicwithaperiodof2 L,any
interval(x0,x0+2L)will do. The choice is a matter of convenience or literally personal
preference.
Exercises
14.2.1 The boundary conditions (such as ψ(0)=ψ(l)=0) may suggest solutions of the form
sin(nπx/l)andeliminatethecorrespondingcosines.
(a) Verify that the boundary conditions used in the Sturm–Liouville theory are satis-
fiedfor theinterval (0,l).Notethatthisis onlyhalftheusualFourierinterval.
(b) Show that the set of functions ϕn(x)=sin(nπx/l), n=1,2,3,...,satisfies an
orthogonalityrelation
integraldisplayl
0ϕm(x)ϕn(x)dx=l
2δmn,n>0.
14.2.2 (a) Expand f(x)=xin the interval (0,2L). Sketch the series you have found (right-
handsideofAns.) over (−2L,2L).
ANS.x=L−2L
π∞summationdisplay
n=11
nsinparenleftbiggnπx
Lparenrightbigg
.
(b) Expand f(x)=xas a sine series in the halfinterval(0,L). Sketch the series you
havefound(right-handsideofAns.) over (−2L,2L).
ANS.x=4L
π∞summationdisplay
n=01
2n+1sinparenleftbigg(2n+1)πx
Lparenrightbigg
.
14.2.3 In some problems it is convenient to approximate sin πxover the interval [0,1]by a
parabola ax(1−x), whereais a constant. To get a feeling for the accuracy of this
approximation,expand 4 x(1−x)inaFouriersineseries (−1≤x≤1):
f(x)=braceleftbigg4x(1−x), 0≤x≤1
4x(1+x),−1≤x≤0bracerightbigg
=∞summationdisplay
n=1bnsinnπx.
892 Chapter 14 Fourier Series
FIGURE 14.4Parabolicsinewave.
ANS.bn=32
π3·1
n3,nodd
bn=0, neven.
(Fig.14.4.)
14.3 A PPLICATIONS OF FOURIER SERIES
Example 14.3.1 SQUARE WAVE—H IGHFREQUENCIES
OneapplicationofFourierseries,theanalysisofa“square”wave(Fig.14.5)intermsofits
Fouriercomponents,occursinelectroniccircuitsdesignedtohandlesharplyrisingpulses.
Supposethatour waveisdefinedby
f(x)=0,−π<x<0,
f(x)=h,0<x<π. (14.32)
FromEqs. (14.2) and(14.3)wefind
a0=1
πintegraldisplayπ
0hdt=h, (14.33)
an=1
πintegraldisplayπ
0hcosntdt=0,n=1,2,3,..., (14.34)
bn=1
πintegraldisplayπ
0hsinntdt=h
nπ(1−cosnπ); (14.35)
bn=2h
nπ,nodd, (14.36)
bn=0,neven. (14.37)
Theresultingseries is
f(x)=h
2+2h
πparenleftbiggsinx
1+sin3x
3+sin5x
5+···parenrightbigg
. (14.38)
14.3 Applications of Fourier Series 893
FIGURE 14.5Squarewave.
Except for the first term, which represents an average of f(x)over the interval [−π,π],
allthecosinetermshavevanished.Since f(x)−h/2isodd,wehaveaFouriersineseries.
Althoughonlytheoddtermsinthesineseriesoccur,theyfallonlyas n−1.Thisconditional
convergence is like that of the alternating harmonic series. Physically this means that our
squarewavecontainsalotof high-frequencycomponents .Iftheelectronicapparatuswill
not pass these components, our square-wave input will emerge more or less rounded off,
perhapsas anamorphousblob. /squaresolid
Example 14.3.2 FULL-WAVERECTIFIER
As a second example, let us ask how well the output of a full-wave rectifier approaches
puredirectcurrent(Fig.14.6).Ourrectifiermaybethoughtofashavingpassedthepositive
peaksof anincomingsinewaveandinvertingthenegativepeaks.This yields
f(t)=sinωt,0<ωt<π,
(14.39)
f(t)=−sinωt,−π<ω t< 0.
Sincef(t)defined here is even, no terms of the form sin nωtwill appear. Again, from
Eqs. (14.2)and(14.3), wehave
a0=−1
πintegraldisplay0
−πsinωtd(ωt)+1
πintegraldisplayπ
0sinωtd(ωt)
=2
πintegraldisplayπ
0sinωtd(ωt)=4
π, (14.40)
an=2
πintegraldisplayπ
0sinωtcosnωtd(ωt)
=−2
π2
n2−1,neven,
=0,nodd. (14.41)
894 Chapter 14 Fourier Series
FIGURE 14.6Full-waverectifier.
Notethat[0,π]isnotanorthogonalityintervalforbothsinesandcosinestogetherandwe
donotgetzerofor even n.Theresultingseriesis
f(t)=2
π−4
π∞summationdisplay
n=2,4,6,...cosnωt
n2−1. (14.42)
The original frequency, ω, has been eliminated. The lowest-frequency oscillation is 2 ω.
The high-frequency components fall off as n−2, showing that the full-wave rectifier does
a fairly good job of approximating direct current. Whether this good approximation is
adequatedepends on the particular application.If the remainingac componentsare objec-
tionable,theymaybefurthersuppressedbyappropriatefiltercircuits.Thesetwoexamples
bringouttwofeaturescharacteristicof Fourierexpansions.7
•Iff(x)has discontinuities (as in the square wave in Example 14.3.1), we can expect
thenthcoefficienttobedecreasingas O(1/n). Convergenceis conditionalonly.
•Iff(x)is continuous (although possibly with discontinuous derivatives, as in the full-
wave rectifier of Example 14.3.2), we can expect the nth coefficient to be decreasing
as 1/n2, thatis, absoluteconvergence./squaresolid
Example 14.3.3 INFINITE SERIES ,RIEMANN ZETAFUNCTION
Asafinalexample,weconsidertheproblemof expanding x2.Let
f(x)=x2,−π<x<π . (14.43)
Sincef(x)iseven,all bn=0.Forthe anwehave
a0=1
πintegraldisplayπ
−πx2dx=2π2
3, (14.44)
7G.Raisbeek, Order of magnitude of Fourier coefficients. Am.Math. Mon. 62: 149–155 (1955).
14.3 Applications of Fourier Series 895
an=2
πintegraldisplayπ
0x2cosnxdx
=2
π·(−1)n2π
n2
=(−1)n4
n2. (14.45)
Fromthis weobtain
x2=π2
3+4∞summationdisplay
n=1(−1)ncosnx
n2. (14.46)
Asitstands,Eq. (14.46)is ofnoparticularimportance.Butif weset x=π,
cosnπ=(−1)n(14.47)
andEq. (14.46)becomes8
π2=π2
3+4∞summationdisplay
n=11
n2, (14.48)
or
π2
6=∞summationdisplay
n=11
n2≡ζ(2), (14.49)
thus yielding the Riemann zeta function, ζ(2), in closed form (in agreement with the
BernoullinumberresultofSection5.9).Fromourexpansionof x2andexpansionsofother
powersof x,numerousotherinfiniteseriescanbeevaluated.Afewareincludedinthislist
ofexercises:
Fourier series Reference
1.∞summationdisplay
n=11
nsinnx=braceleftBigg
−1
2(π+x),−π≤x<0
1
2(π−x), 0≤x<πExercise14 .1.5
Exercise14 .3.3
2.∞summationdisplay
n=1(−1)n+11
nsinnx=1
2x,−π<x<πExercise14 .1.6
Exercise14 .3.2
3.∞summationdisplay
n=01
2n+1sin(2n+1)x=braceleftbigg−π/4,−π<x<0
+π/4,0<x<πExercise14 .1.7
Eq.(14.38)
4.∞summationdisplay
n=1cosnx
n=−lnbracketleftbigg
2sinparenleftbigg|x|
2parenrightbiggbracketrightbigg
,−π<x<πEq.(14.11)
Exercise14.1.9(b)
5.∞summationdisplay
n=1(−1)n1
ncosnx=−lnbracketleftbigg
2cosparenleftbiggx
2parenrightbiggbracketrightbigg
,−π<x<π Exercise 14.1.9(a)
6.∞summationdisplay
n=01
2n+1cos(2n+1)x=1
2lnbracketleftbigg
cot|x|
2bracketrightbigg
,−π<x<π
8Notethat the point x=πis not a point of discontinuity.
896 Chapter 14 Fourier Series
Thesquare-waveFourierseries fromEq. (14.38)anditem(3) inthetable,
g(x)=∞summationdisplay
n=0sin(2n+1)x
2n+1=(−1)mπ
4,mπ<x<(m+1)π, (14.50)
can be used to derive Riemann’s functional equation for the zeta function . Its defining
Dirichletseriescanbewritteninvariousforms:
ζ(s)=∞summationdisplay
n=1n−s=1+∞summationdisplay
n=1(2n)−s+∞summationdisplay
n=1(2n+1)−s
=2−sζ(s)+∞summationdisplay
n=0(2n+1)−s
implyingthatthefunction λ(s)definedinSection5.9(alongwith η(s)) satisfies
λ(s)≡∞summationdisplay
n=0(2n+1)−s=parenleftbig
1−2−sparenrightbig
ζ(s). (14.51)
Heresisacomplexvariable.BothDirichletseriesconvergefor σ=ℜs>1.Alternatively,
usingEq. (14.51), wehave
η(s)≡∞summationdisplay
n=1(−1)n−1n−s=∞summationdisplay
n=0(2n+1)−s−∞summationdisplay
n=1(2n)−s=parenleftbig
1−21−sparenrightbig
ζ(s), (14.52)
which converges already for ℜs>0 using the Leibniz convergence criterion (see Sec-
tion5.3).
AnotherapproachtoDirichletseriesstartsfromEuler’sintegralforthegammafunction,
integraldisplay∞
0ys−1e−nydy=n−sintegraldisplay∞
0e−yys−1dy=n−sŴ(s), (14.53)
whichmaybesummedusingthegeometricseries
∞summationdisplay
n=1e−ny=e−y
1−e−y=1
ey−1
toyieldtheintegralrepresentationfor thezetafunction:
integraldisplay∞
0ys−1
ey−1dy=ζ(s)Ŵ(s). (14.54)
If wecombinethealternativeforms ofEq. (14.53),
integraldisplay∞
0ys−1e−inydy=n−sŴ(s)e−iπs/2,
integraldisplay∞
0ys−1einydy=n−sŴ(s)eiπs/2,
14.3 Applications of Fourier Series 897
weobtain
integraldisplay∞
0ys−1sin(ny)dy=n−sŴ(s)sinπs
2. (14.55)
DividingbothsidesofEq.(14.55)by nandsummingoverallodd nyields,for σ=ℜ(s)>
0,
integraldisplay∞
0g(y)ys−1dy=parenleftbig
1−2−s−1parenrightbig
ζ(s+1)Ŵ(s)sinπs
2, (14.56)
usingEqs. (14.50) and (14.51). Here, the interchangeof summationandintegrationis jus-
tifiedbyuniformconvergence.Thisrelationisattheheartofthefunctionalequation.Ifwe
divide the integration range into intervals mπ<y<(m +1)πand substitute Eq. (14.50)
intoEq. (14.56)wefind
integraldisplay∞
0g(y)ys−1dy=π
4∞summationdisplay
m=0(−1)mintegraldisplay(m+1)π
mπys−1dy
=πs+1
4sbraceleftbigg∞summationdisplay
m=1(−1)mbracketleftbig
(m+1)s−msbracketrightbig
+1bracerightbigg
=πs+1
2sparenleftbig
1−2s+1parenrightbig
ζ(−s), (14.57)
using Eq. (14.52). The series in Eq. (14.57) converges for ℜs<1 to an analytic func-
tion. Comparing Eqs. (14.56) and (14.57) for the common area of convergence to analytic
functions, 0 <σ=ℜs<1,wegetthe functionalequation
πs+1
2sparenleftbig
1−2s+1parenrightbig
ζ(−s)=parenleftbig
1−2−s−1parenrightbig
ζ(s+1)Ŵ(s)sinπs
2,
whichcanberewrittenas
ζ(1−s)=2(2π)−sζ(s)Ŵ(s)cosπs
2. (14.58)
This functional equation provides an analytic continuation of ζ(s)into the negative half-
plane ofs.F o rs→1 the pole of ζ(s)and the zero of cos (πs/2)cancel in Eq. (14.58), so
ζ(0)=−1/2results.Sincecos (πs/2)=0fors=2m+1=oddinteger,Eq.(14.58)gives
ζ(−2m)=0,thetrivialzerosofthezetafunctionfor m=1,2,....Allotherzerosmustlie
in the“criticalstrip” 0 <σ=ℜs<1. Theyare closely relatedto thedistributionof prime
numbers because the prime number product for ζ(s)(see Section 5.9) can be converted
into a Dirichlet series over prime powers for ζ′/ζ=dlnζ(s)/ds.F r o mh e r eo nw es k e t c h
ideasonly,withoutproofs.UsingtheinverseMellintransform(seeSection16.2)yieldsthe
relation
summationdisplay
pm<x,p=prime
m=1,2,...lnp=−1
2πiintegraldisplayσ+i∞
σ−i∞ζ′(s)
ζ(s)sxsds (14.59)
forσ>1, which is a cornerstone of analytic number theory. Since zeros of ζ(s)become
simple poles of ζ′/ζ, the asymptotic distribution of prime numbers is directly related by
898 Chapter 14 Fourier Series
Eq. (14.59) to the zeros of the Riemann zeta function. Riemann conjectured that all zeros
lieontheline σ=1/2,thatis,havetheform1 /2+itwithreal t.Ifso,onecouldshiftthe
lineofintegrationtotheleftto σ=1/2+ε,thesimplepoleof ζ(s)ats=1 givingriseto
the residue x, while the integral along the line σ=1/2+εis of order O(x1/2+ε). Hence,
theremarkablysmallremainderintheasymptoticestimate
summationdisplay
p<xlnp∼x+Oparenleftbig
x1/2+εparenrightbig
,x→∞
would result for arbitrarily small ε. This is equivalent to the estimate for the number of
primesbelow x,
π(x)=summationdisplay
p<x1=integraldisplayx
2(lnt)−1dt+Oparenleftbig
x1/2+εparenrightbig
,x→∞.
In fact, numerical studies have shown that the first 300 ×109zeros are simple and lie
all on the critical line σ=1/2. For more details the reader is referred to the classic text
by E. C. Titchmarsh and D. R. Heath-Brown, The Theory of the Riemann Zeta Function ,
Oxford, UK: Clarendon Press (1986); H. M. Edwards, Riemann’s Zeta Function ,N e w
York: Academic Press (1974) and Dover (2003); J. Van de Lune, H. J. J. Te Riele, and
D. T. Winter, On the zeros of the Riemann zeta function in the critical strip. IV. Math.
Comput.47:667(1986).PopularaccountscanbefoundinM.duSautoy, TheMusicofthe
Primes:SearchingtoSolvetheGreatestMysteryinMathematics ,NewYork:HarperCollins
(2003); J. Derbyshire, Prime Obsession: Bernhard Riemann and the Greatest Unsolved
Problem in Mathematics . Washington, DC: Joseph Henry Press (2003); K. Sabbagh, The
Riemann Hypothesis: The Greatest Unsolved Problem in Mathematics , New York: Farrar,
StrausandGiroux(2003).
More recently the statistics of the zeros ρof the Riemann zeta function on the critical
line played a prominent role in the development of theories of chaos (see Chapter 18 for
an introduction). Assuming that there is a quantum mechanical system whose energies
are the imaginary parts of the ρ, then primes determine the primitive periodic orbits of
the associated classically chaotic system. For this case Gutzwiller’s trace formula, which
relates quantum energy levels and classical periodic orbits, plays a central role and can be
better understood using properties of the zeta function and primes. For more details see
Sections12.6and12.7byJ.Keating,in TheNatureofChaos (T.Mullin,ed.),Oxford,UK:
ClarendonPress (1993),andreferencestherein. /squaresolid
Exercises
14.3.1 DeveloptheFourierseriesrepresentationof
f(t)=braceleftbigg0,−π≤ωt≤0,
sinωt,0≤ωt≤π.
Thisistheoutputofasimplehalf-waverectifier.Itisalsoanapproximationofthesolar
thermaleffectthatproduces“tides”intheatmosphere.
ANS.f(t)=1
π+1
2sinωt−2
π∞summationdisplay
n=2,4,6,...cosnωt
n2−1.
14.3 Applications of Fourier Series 899
FIGURE 14.7Triangularwave.
14.3.2 Asawtoothwaveisgivenby
f(x)=x,−π<x<π .
Showthat
f(x)=2∞summationdisplay
n=1(−1)n+1
nsinnx.
14.3.3 Adifferentsawtoothwaveis describedby
f(x)=braceleftBigg
−1
2(π+x),−π≤x<0
+1
2(π−x),0<x≤π.
Showthat f(x)=summationtext∞
n=1(sinnx/n).
14.3.4 Atriangularwave(Fig. 14.7)is representedby
f(x)=braceleftBiggx,0<x<π
−x,−π<x<0.
Represent f(x)bya Fourierseries.
ANS.f(x)=π
2−4
πsummationdisplay
n=1,3,5,...cosnx
n2.
14.3.5 Expand
f(x)=braceleftBigg1,x2<x2
0
0,x2>x2
0
intheinterval[−π,π].
Note.This variable-widthsquarewaveis ofsomeimportanceinelectronicmusic.
900 Chapter 14 Fourier Series
FIGURE 14.8Cross section
ofsplittube.
14.3.6 Ametalcylindricaltubeofradius aissplitlengthwiseintotwonontouchinghalves.The
top half is maintained at a potential +V, the bottom half at a potential −V(Fig. 14.8).
Separate the variables in Laplace’s equation and solve for the electrostatic potential for
r≤a. Observe the resemblance between your solution for r=aand the Fourier series
for asquarewave.
14.3.7 A metal cylinder is placed in a (previously) uniform electric field, E0, with the axis of
thecylinderperpendiculartothatof theoriginalfield.
(a) Findtheperturbedelectrostaticpotential.
(b) Findtheinducedsurfacechargeonthecylinderasa functionofangularposition.
14.3.8 Transform the Fourier expansion of a square wave, Eq. (14.38), into a power series.
Show that the coefficients of x1form adivergent series. Repeat for the coefficients
ofx3.
A power series cannot handle a discontinuity. These infinite coefficients are the result
ofattemptingtobeatthisbasiclimitationonpowerseries.
14.3.9 (a) ShowthattheFourierexpansionof cos axis
cosax=2asinaπ
πbraceleftbigg1
2a2−cosx
a2−12+cos2x
a2−22−···bracerightbigg
,
an=(−1)n2asinaπ
π(a2−n2).
(b) Fromtheprecedingresultshowthat
aπcotaπ=1−2∞summationdisplay
p=1ζ(2p)a2p.
ThisprovidesanalternatederivationoftherelationbetweentheRiemannzetafunction
andtheBernoullinumbers,Eq. (5.152).
14.3 Applications of Fourier Series 901
14.3.10 DerivetheFourierseriesexpansionoftheDiracdeltafunction δ(x)intheinterval−π<
x<π.
(a) Whatsignificancecanbeattachedtotheconstantterm?
(b) Inwhatregionisthisrepresentationvalid?
(c) Withtheidentity
Nsummationdisplay
n=1cosnx=sin(Nx/2)
sin(x/2)cosbracketleftbiggparenleftbigg
N+1
2parenrightbiggx
2bracketrightbigg
,
showthatyourFourierrepresentationof δ(x)is consistentwithEq. (1.190).
14.3.11 Expandδ(x−t)in a Fourier series. Compare your result with the bilinear form of
Eq. (1.190).
ANS.δ( x−t)=1
2π+1
π∞summationdisplay
n=1(cosnxcosnt+sinnxsinnt)
=1
2π+1
π∞summationdisplay
n=1cosn(x−t).
14.3.12 Verifythat
δ(ϕ1−ϕ2)=1
2π∞summationdisplay
m=−∞eim(ϕ1−ϕ2)
isaDiracdeltafunctionbyshowingthatitsatisfiesthedefinitionofaDiracdeltafunc-
tion:
integraldisplayπ
−πf(ϕ1)1
2π∞summationdisplay
m=−∞eim(ϕ1−ϕ2)dϕ1=f(ϕ2).
Hint.Represent f(ϕ1)byanexponentialFourierseries.
Note. The continuum analog of this expression is developed in Section 15.2. The most
important application of this expression is in the determination of Green’s functions,
Section9.7.
14.3.13 (a) Using
f(x)=x2,−π<x<π ,
showthat
∞summationdisplay
n=1(−1)n+1
n2=π2
12=η(2).
(b) Using the Fourier series for a triangular wave developed in Exercise 14.3.4, show
that
∞summationdisplay
n=11
(2n−1)2=π2
8=λ(2).
902 Chapter 14 Fourier Series
(c) Using
f(x)=x4,−π<x<π ,
showthat
∞summationdisplay
n=11
n4=π4
90=ζ(4),∞summationdisplay
n=1(−1)n+1
n4=7π4
720=η(4).
(d) Using
f(x)=braceleftbiggx(π−x),0<x<π,
x(π+x), π <x< 0,
derive
f(x)=8
π∞summationdisplay
n=1,3,5,...sinnx
n3
andshowthat
∞summationdisplay
n=1,3,5,...(−1)(n−1)/21
n3=1−1
33+1
53−1
73+···=π3
32=β(3).
(e) UsingtheFourierseriesfor asquarewave,showthat
∞summationdisplay
n=1,3,5,...(−1)(n−1)/21
n=1−1
3+1
5−1
7+···=π
4=β(1).
ThisisLeibniz’formulafor π,obtainedbyadifferenttechniqueinExercise5.7.6.
Note.T h eη(2),η(4),λ(2),β(1), andβ(3)functions are defined by the indicated
series. GeneraldefinitionsappearinSection5.9.
14.3.14 (a) FindtheFourierseries representationof
f(x)=braceleftbigg0,−π<x≤0
x,0≤x<π.
(b) FromtheFourierexpansionshowthat
π2
8=1+1
32+1
52+···.
14.3.15 Asymmetrictriangularpulseofadjustableheightandwidthisdescribedby
f(x)=braceleftbigga(1−x/b), 0≤|x|≤b
0,b ≤|x|≤π.
(a) ShowthattheFouriercoefficientsare
a0=ab
π,a n=2ab
π(nb)2(1−cosnb).
Sum the finite Fourier series through n=10 and through n=100 forx/π=
0(1/9)1.Takea=1 andb=π/2.
14.4 Properties of Fourier Series 903
(b) CallaFourieranalysissubroutine(ifavailable)tocalculatetheFouriercoefficients
off(x),a 0througha10.
14.3.16 (a) Using a Fourier analysis subroutine, calculate the Fourier cosine coefficients a0
througha10of
f(x)=bracketleftbigg
1−parenleftbiggx
πparenrightbigg2bracketrightbigg1/2
,x∈[−π,π].
(b) Spot-check by calculating some of the preceding coefficients by direct numerical
quadrature.
Checkvalues. a0=0.785,a2=0.284.
14.3.17 Using a Fourier analysis subroutine, calculate the Fourier coefficients through a10and
b10for
(a) afull-waverectifier,Example14.3.2,
(b) ahalf-waverectifier,Exercise14.3.1.Checkyourresultsagainsttheanalyticforms
given(Eq. (14.41)andExercise14.3.1).
14.4 P ROPERTIES OF FOURIER SERIES
Convergence
Itmightbenoted,first,thatourFourierseriesshouldnotbeexpectedtobeuniformlycon-
vergentifitrepresentsadiscontinuousfunction.Auniformlyconvergentseriesofcontinu-
ous functions (sinnx,cosnx)always yields a continuous function (compare Section 5.5).
If, however,
(a)f(x)iscontinuous,−π≤x≤π,
(b)f(−π)=f(+π), and
(c)f′(x)issectionallycontinuous,
theFourierseries for f(x)willconvergeuniformly.These restrictions donotdemandthat
f(x)beperiodic,buttheywillbesatisfiedbycontinuous,differentiable,periodicfunctions
(period of 2 π). For a proof of uniform convergence we refer to the literature.9With or
without a discontinuity in f(x), the Fourier series will yield convergence in the mean,
Section10.4.
9See, for instance, R. V. Churchill, Fourier Series and Boundary Value Problems , 5th ed., New York: McGraw-Hill (1993),
Section 38.
904 Chapter 14 Fourier Series
Integration
Term-by-termintegrationoftheseries
f(x)=a0
2+∞summationdisplay
n=1ancosnx+∞summationdisplay
n=1bnsinnx (14.60)
yields
integraldisplayx
x0f(x)dx=a0x
2vextendsinglevextendsinglevextendsinglevextendsinglex
x0+∞summationdisplay
n=1an
nsinnxvextendsinglevextendsinglevextendsinglex
x0−∞summationdisplay
n=1bn
ncosnxvextendsinglevextendsinglevextendsinglex
x0. (14.61)
Clearly, the effect of integration is to place an additional power of nin the denominator
of each coefficient. This results in more rapid convergence than before. Consequently, a
convergent Fourier series may always be integrated term by term, the resulting series con-
verginguniformlytotheintegraloftheoriginalfunction.Indeed,term-by-termintegration
may be valid even if the original series (Eq. (14.60)) is not itself convergent. The func-
tionf(x)need only be integrable. A discussion will be found in Jeffreys and Jeffreys,
Section14.06(see theAdditionalReadings).
Strictlyspeaking,Eq.(14.61)maynotbeaFourierseries;thatis,if a0/negationslash=0,therewillbe
ater m1
2a0x. However,
integraldisplayx
x0f(x)dx−1
2a0x (14.62)
willstillbeaFourierseries.
Differentiation
The situation regarding differentiation is quite different from that of integration. Here the
wordiscaution.Considertheseries for
f(x)=x,−π<x<π . (14.63)
Wereadilyfind(compareExercise14.3.2)thattheFourierseries is
x=2∞summationdisplay
n=1(−1)n+1sinnx
n,−π<x<π . (14.64)
Differentiatingtermbyterm, weobtain
1=2∞summationdisplay
n=1(−1)n+1cosnx, (14.65)
whichisnotconvergent. Warning :Checkyourderivativeforconvergence.
For a triangular wave (Exercise 14.3.4), in which the convergence is more rapid (and
uniform),
f(x)=π
2−4
π∞summationdisplay
n=1,oddcosnx
n2. (14.66)
14.4 Properties of Fourier Series 905
Differentiatingtermbytermweget
f′(x)=4
π∞summationdisplay
n=1,oddsinnx
n, (14.67)
whichistheFourierexpansionofasquarewave,
f′(x)=braceleftbigg1,0<x<π,
−1,−π<x<0.(14.68)
InspectionofFig.14.7verifiesthatthisis indeedthederivativeofourtriangularwave.
•As the inverse of integration, the operation of differentiation has placed an additional
factornin the numerator of each term. This reduces the rate of convergence and may,
asinthefirst casementioned,renderthedifferentiatedseriesdivergent.
•Ingeneral,term-by-termdifferentiationispermissibleunderthesameconditionslisted
for uniformconvergence.
Exercises
14.4.1 ShowthatintegrationoftheFourierexpansionof f(x)=x,−π<x<π ,leadsto
π2
12=∞summationdisplay
n=1(−1)n+1
n2=1−1
4+1
9−1
16+···.
14.4.2 Parseval’sidentity.
(a) AssumingthattheFourierexpansionof f(x)isuniformlyconvergent,showthat
1
πintegraldisplayπ
−πbracketleftbig
f(x)bracketrightbig2dx=a2
0
2+∞summationdisplay
n=1parenleftbig
a2
n+b2
nparenrightbig
.
ThisisParseval’sidentity.Itisactuallyaspecialcaseofthecompletenessrelation,
Eq.(10.73).
(b) Given
x2=π2
3+4∞summationdisplay
n=1(−1)ncosnx
n2,−π≤x≤π,
applyParseval’sidentitytoobtain ζ(4)inclosedform.
(c) Theconditionofuniformconvergenceisnotnecessary.Showthisbyapplyingthe
Parsevalidentitytothesquarewave
f(x)=braceleftbigg−1,−π<x<0
1,0<x<π
=4
π∞summationdisplay
n=1sin(2n−1)x
2n−1.
906 Chapter 14 Fourier Series
FIGURE 14.9Rectangularpulse.
14.4.3 Show that integrating the Fourier expansion of the Dirac delta function (Exer-
cise 14.3.10) leads to the Fourier representation of the square wave, Eq. (14.38), with
h=1.
Note. Integrating the constant term (1/2π)leads to a term x/2π. What are you going
todowiththis?
14.4.4 IntegratetheFourierexpansionoftheunitstepfunction
f(x)=braceleftbigg0,−π<x<0
x,0<x<π.
Showthatyourintegratedseries agreeswithExercise14.3.14.
14.4.5 Intheinterval (−π,π),
δn(x)=braceleftBiggn,for|x|<1
2n,
0,for|x|>1
2n
(Fig.14.9).
(a) Expand δn(x)as aFouriercosineseries.
(b) Show that your Fourier series agrees with a Fourier expansion of δ(x)in the limit
asn→∞.
14.4.6 Confirm the delta function nature of your Fourier series of Exercise 14.4.4 by showing
thatfor any f(x)thatis finiteintheinterval [−π,π]andcontinuousat x=0,
integraldisplayπ
−πf(x)bracketleftbig
Fourierexpansionof δ∞(x)bracketrightbig
dx=f(0).
14.4.7 (a) Show that the Dirac delta function δ(x−a), expanded in a Fourier sine series in
thehalf-interval (0,L)(0<a<L) , is givenby
δ(x−a)=2
L∞summationdisplay
n=1sinparenleftbiggnπa
Lparenrightbigg
sinparenleftbiggnπx
Lparenrightbigg
.
Notethatthisseriesactuallydescribes
−δ(x+a)+δ(x−a)intheinterval (−L,L).
14.4 Properties of Fourier Series 907
(b) By integrating both sides of the preceding equation from 0 to x, show that the
cosineexpansionofthesquarewave
f(x)=braceleftbigg0,0≤x<a
1, a<x<L,
is
f(x)=2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
−2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
cosparenleftbiggnπx
Lparenrightbigg
,
for 0≤x<L.
(c) Verifythattheterm
2
π∞summationdisplay
n=11
nsinparenleftbiggnπa
Lparenrightbigg
is/angbracketleftf(x)/angbracketright.
14.4.8 Verify the Fourier cosine expansion of the square wave, Exercise 14.4.7(b), by direct
calculationoftheFouriercoefficients.
14.4.9 (a) A string is clamped at both ends x=0 andx=L. Assuming small-amplitude
vibrations,wefindthattheamplitude y(x,t)satisfiesthewaveequation
∂2y
∂x2=1
v2∂2y
∂t2.
Herevisthewavevelocity.Thestringissetinvibrationbyasharpblowat x=a.
Hencewehave
y(x,0)=0,∂y(x,t)
∂t=Lv0δ(x−a)att=0.
The constant Lis included to compensate for the dimensions (inverse length) of
δ(x−a). Withδ(x−a)given by Exercise 14.4.7(a), solve the wave equation
subjecttotheseinitialconditions.
ANS.y(x,t)=2v0L
πv∞summationdisplay
n=11
nsinnπa
Lsinnπx
Lsinnπvt
L.
(b) Showthatthetransversevelocityofthestring ∂y(x,t)/∂t is givenby
∂y(x,t)
∂t=2v0∞summationdisplay
n=1sinnπa
Lsinnπx
Lcosnπvt
L.
14.4.10 A string, clamped at x=0 and atx=1, is vibrating freely. Its motion is described by
thewaveequation
∂2u(x,t)
∂t2=v2∂2u(x,t)
∂x2.
908 Chapter 14 Fourier Series
AssumeaFourierexpansionoftheform
u(x,t)=∞summationdisplay
n=1bn(t)sinnπx
l
anddeterminethecoefficients bn(t). Theinitialconditionsare
u(x,0)=f(x) and∂
∂tu(x,0)=g(x).
Note.ThisisonlyhalftheconventionalFourierorthogonalityintegralinterval.However,
aslongasonlythesinesareincludedhere,theSturm–Liouvilleboundaryconditionsare
stillsatisfiedandthefunctionsare orthogonal.
ANS.bn(t)=Ancosnπvt
l+Bnsinnπvt
l,
An=2
lintegraldisplayl
0f(x)sinnπx
ldx,Bn=2
nπvintegraldisplayl
0g(x)sinnπx
ldx.
14.4.11 (a) Let us continue the vibrating string problem, Exercise 14.4.10. The presence of a
resistingmediumwilldampthevibrationsaccordingtotheequation
∂2u(x,t)
∂t2=v2∂2u(x,t)
∂x2−k∂u(x,t)
∂t.
Assumea Fourierexpansion
u(x,t)=∞summationdisplay
n=1bn(t)sinnπx
l
and again determine the coefficients bn(t). Take the initial and boundary condi-
tionstobethesameasinExercise14.4.10.Assumethedampingtobesmall.
(b) Repeat,butassumethedampingtobelarge.
ANS.(a)bn(t)=e−kt/2{Ancosωnt+Bnsinωnt},
An=2
lintegraldisplayl
0f(x)sinnπx
ldx,
Bn=2
ωnlintegraldisplayl
0g(x)sinnπx
ldx+k
2ωnAn,
ω2
n=parenleftbiggnπv
lparenrightbigg
−parenleftbiggk
2parenrightbigg2
>0.
14.4 Properties of Fourier Series 909
(b)bn(t)=e−kt/2{Ancoshσnt+Bnsinhσnt},
An=2
lintegraldisplayl
0f(x)sinnπx
ldx,
Bn=2
σnlintegraldisplayl
0g(x)sinnπx
ldx+k
2σnAn,
σ2
n=parenleftbiggk
2parenrightbigg2
−parenleftbiggnπv
lparenrightbigg2
>0.
14.4.12 Find the charge distribution over the interior surfaces of the semicircles of Exer-
cise14.3.6.
Note. You obtain a divergent series and this Fourier approach fails. Using conformal
mapping techniques, we may show the charge density to be proportional to csc θ. Does
cscθhaveaFourierexpansion?
14.4.13 Given
ϕ1(x)=∞summationdisplay
n=1sinnx
n=braceleftBigg−1
2(π+x),−π≤x<0,
1
2(π−x), 0<x≤π,
showbyintegratingthat
ϕ2(x)≡∞summationdisplay
n=1cosnx
n2=
1
4(π+x)2−π2
12,−π≤x≤0
1
4(π−x)2−π2
12,0≤x≤π.
14.4.14 Given
ψ2s(x)=∞summationdisplay
n=1sinnx
n2s,ψ 2s+1(x)=∞summationdisplay
n=1cosnx
n2s+1,
developthefollowingrecurrencerelations:
(a)ψ2s(x)=integraldisplayx
0ψ2s−1(x)dx
(b)ψ2s+1(x)=ζ(2s+1)−integraldisplayx
0ψ2s(x)dx.
Note. The functions ψs(x)and theϕs(x)of the preceding two exercises are known as
Clausenfunctions .Intheorytheymaybeusedtoimprovetherateofconvergenceofa
Fourierseries.AswiththeseriesofChapter5,thereisalwaysthequestionofhowmuch
analyticalworkwedoandhowmucharithmeticworkwedemandthatthecomputerdo.
As computers become steadily more powerful, the balance progressively shifts so that
wearedoingless anddemandingthattheydomore.
14.4.15 Showthat
f(x)=∞summationdisplay
n=1cosnx
n+1
910 Chapter 14 Fourier Series
maybewrittenas
f(x)=ψ1(x)−ϕ2(x)+∞summationdisplay
n=1cosnx
n2(n+1).
Note.ψ1(x)andϕ2(x)aredefinedintheprecedingexercises.
14.5 G IBBS PHENOMENON
TheGibbsphenomenonisanovershoot,apeculiarityoftheFourierseriesandothereigen-
functionseriesata simplediscontinuity.AnexampleisseeninFig.14.1.
Summation of Series
In Section 14.1 the sum of the first several terms of the Fourier series for a sawtooth wave
was plotted (Fig. 14.10). Now we develop analytic methods of summing the first rterms
ofourFourierseries.
FromEq. (14.19),
ancosnx+bnsinnx=1
πintegraldisplayπ
−πf(t)cosn(t−x)dt. (14.69)
Thenthe rthpartialsumbecomes10
sr(x)=a0
2+rsummationdisplay
n=1(ancosnx+bnsinnx)
=ℜ1
πintegraldisplayπ
−πf(t)bracketleftbigg1
2+rsummationdisplay
n=1e−i(t−x)nbracketrightbigg
dt. (14.70)
Summingthefiniteseries ofexponentials(geometricprogression),11weobtain
sr(x)=1
2πintegraldisplayπ
−πf(t)sin[(r+1
2)(t−x)]
sin1
2(t−x)dt. (14.71)
Thisis convergentatallpoints,including t=x. Thefactor
sin[(r+1
2)(t−x)]
2πsin1
2(t−x)
istheDirichletkernelmentionedinSection1.15asaDiracdeltadistribution.
10It is of some interest to note that this series also occurs in theanalysis of the diffraction grating ( rslits).
11Compare Exercise 6.1.7 with initial value n=1.
14.5 Gibbs Phenomenon 911
Square Wave
For convenience of numerical calculation we consider the behavior of the Fourier series
thatrepresentstheperiodicsquarewave
f(x)=braceleftBiggh
2,0<x<π,
−h
2,−π<x<0.(14.72)
This is essentially the square wave used in Section 14.3, and we immediately see that the
solutionis
f(x)=2h
πparenleftbiggsinx
1+sin3x
3+sin5x
5+···parenrightbigg
. (14.73)
ApplyingEq.(14.71)tooursquarewave(Eq.(14.72)),wehavethesumofthefirst rterms
(plus1
2a0, whichiszerohere):
sr(x)=h
4πintegraldisplayπ
0sin[(r+1
2)(t−x)]
sin1
2(t−x)dt−h
4πintegraldisplay0
−πsin[(r+1
2)(t−x)]
sin1
2(t−x)dt
=h
4πintegraldisplayπ
0sin[(r+1
2)(t−x)]
sin1
2(t−x)dt−h
4πintegraldisplayπ
0sin[(r+1
2)(t+x)]
sin1
2(t+x)dt.(14.74)
Thislastresultfollowsfrom thetransformation
/vectort−t
inthesecondintegral.Replacing t−xinthefirsttermwith sandt+xinthesecondterm
withs,weobtain
sr(x)=h
4πintegraldisplayπ−x
−xsin(r+1
2)s
sin1
2sds−h
4πintegraldisplayπ+x
xsin(r+1
2)s
sin1
2sds. (14.75)
The intervals of integration are shown in Fig. 14.10(top). Because the integrands have
the same mathematical form, the integrals from xtoπ−xcancel, leaving the integral
rangesshowninthebottomportionof Fig.14.10:
sr(x)=h
4πintegraldisplayx
−xsin(r+1
2)s
sin1
2sds−h
4πintegraldisplayπ+x
π−xsin(r+1
2)s
sin1
2sds. (14.76)
Considerthepartialsuminthevicinityofthediscontinuityat x=0.Asx→0,thesec-
ond integral becomes negligible, and we associate the first integral with the discontinuity
atx=0.Usingr+1
2=pandps=ξweobtain
sr(x)=h
2πintegraldisplaypx
0sinξ
sin(ξ/2p)dξ
p. (14.77)
912 Chapter 14 Fourier Series
FIGURE 14.10Intervalsof integration—Eq.(14.75).
Calculation of Overshoot
Our partial sum sr(x)starts at zero when x=0 (in agreement with Eq. (14.22)) and in-
creases until ξ=ps=π, at which point the numerator, sin ξ, goes negative. For large r,
andthereforeforlarge p,ourdenominatorremainspositive.Wegetthemaximumvalueof
the partial sum by taking the upper limit px=π. Right here we see that x, the locationof
theovershootmaximum,isinverselyproportionaltothenumberofterms taken:
x=π
p≈π
r.
Themaximumvalueofthepartialsumis then
sr(x)max=h
2·1
πintegraldisplayπ
0sinξdξ
sin(ξ/2p)p≈h
2·2
πintegraldisplayπ
0sinξ
ξdξ. (14.78)
Intermsof thesineintegral, si (x)of Section8.5,
integraldisplayπ
0sinξ
ξdξ=π
2+si(π). (14.79)
Theintegralisclearlygreaterthan π/2,sinceitcanbewrittenas
parenleftbiggintegraldisplay∞
0−integraldisplay3π
π−integraldisplay5π
3π−···parenrightbiggsinξ
ξdξ=integraldisplayπ
0sinξ
ξdξ. (14.80)
WesawinExample7.1.4thattheintegralfrom0to ∞isπ/2.Fromthisintegralweare
subtracting a series of negative terms. A Gaussian quadrature or a power-series expansion
andterm-by-termintegrationyields
2
πintegraldisplayπ
0sinξ
ξdξ=1.1789797..., (14.81)
whichmeansthattheFourierseriestendstoovershootthepositivecornerbysome18per-
centandtoundershootthenegativecornerbythesameamount,assuggestedinFig.14.11.
14.5 Gibbs Phenomenon 913
FIGURE 14.11Squarewave—Gibbsphenomenon.
The inclusion of more terms (increasing r) does nothing to remove this overshoot but
merely moves it closer to the point of discontinuity. The overshoot is the Gibbs phenom-
enon, and because of it the Fourier series representation may be highly unreliable for pre-
cisenumericalwork, especiallyinthevicinityof adiscontinuity.
The Gibbs phenomenon is not limited to the Fourier series. It occurs with other eigen-
functionexpansions.Exercise12.3.27isanexampleoftheGibbsphenomenonforaLegen-
dreseries.Formoredetails,seeW.J.Thompson,FourierseriesandtheGibbsphenomenon,
Am.J. Phys. 60:425(1992).
Exercises
14.5.1 With the partial sum summation techniques of this section, show that at a discontinuity
inf(x)the Fourier series for f(x)takes on the arithmetic mean of the right- and left-
handlimits:
f(x0)=1
2bracketleftbig
f(x0+0)+f(x0−0)bracketrightbig
.
Inevaluatinglim r→∞sr(x0)youmayfinditconvenienttoidentifypartoftheintegrand
asaDiracdeltafunction.
14.5.2 Determinethepartialsum, sn, oftheseries inEq.(14.73) byusing
(a)sinmx
m=integraldisplayx
0cosmy dy,(b)nsummationdisplay
p=1cos(2p−1)y=sin2ny
2siny.
DoyouagreewiththeresultgiveninEq. (14.79)?
14.5.3 Evaluate the finite step function series, Eq. (14.73), h=2, using 100, 200, 300, 400,
and 500 terms for x=0.0000(0.0005)0.0200. Sketch your results (five curves) or, if a
plottingroutineis available,plotyourresults.
914 Chapter 14 Fourier Series
14.5.4 (a) CalculatethevalueoftheGibbsphenomenonintegral
I=2
πintegraldisplayπ
0sint
tdt
bynumericalquadratureaccurateto12significantfigures.
(b) Check your result by (1) expanding the integrand as a series, (2) integrating term
by term, and (3) evaluating the integrated series. This calls for double precision
calculation.
ANS.I=1.178979744472 .
14.6 D ISCRETE FOURIER TRANSFORM
For many physicists the Fourier transform is automatically the continuous Fourier trans-
form of Chapter 15. The use of digital computers, however, necessarily replaces a contin-
uumofvaluesbyadiscreteset;anintegrationisreplacedbyasummation.Thecontinuous
FouriertransformbecomesthediscreteFouriertransformandanappropriatetopicforthis
chapter.
Orthogonality over Discrete Points
The orthogonality of the trigonometric functions and the imaginary exponentials is ex-
pressedinEqs.(14.15)to(14.18).Thisistheusualorthogonalityforfunctions: integration
ofaproductoffunctionsovertheorthogonalityinterval.Thesines,cosines,andimaginary
exponentials have the remarkable property that they are also orthogonal over a series of
discrete,equallyspacedpointsovertheperiod(theorthogonalityinterval).
Consideraset of 2 Ntimevalues
tk=0,T
2N,2T
2N,...,(2N−1)T
2N(14.82)
forthetimeinterval (0,T). Then
tk=kT
2N,k=0,1,2,...,2N−1. (14.83)
We shall prove that the exponential functions exp (2πiptk/T)and exp(2πiqtk/T)satisfy
anorthogonalityrelationoverthediscretepoints tk:
2N−1summationdisplay
k=0bracketleftbigg
expparenleftbigg2πiptk
Tparenrightbiggbracketrightbigg∗
expparenleftbigg2πiqtk
Tparenrightbigg
=2Nδp,q±2nN. (14.84)
Heren,p, andqareallintegers.
Replacing q−pbys,wefindthattheleft-handsideofEq. (14.84)becomes
2N−1summationdisplay
k=0expparenleftbigg2πistk
Tparenrightbigg
=2N−1summationdisplay
k=0expparenleftbigg2πisk
2Nparenrightbigg
.
14.6 Discrete Fourier Transform 915
Thisright-handsideisobtainedbyusingEq.(14.83)toreplace T.Thisisafinitegeometric
serieswithaninitialterm1andaratio
r=expparenleftbiggπis
Nparenrightbigg
.
FromEq. (5.3),
2N−1summationdisplay
k=0expparenleftbigg2πistk
Tparenrightbigg
=
1−r2N
1−r=0,r/negationslash=1
2N, r =1,(14.85)
establishing Eq. (14.84), our basic orthogonality relation. The upper value, zero, is a con-
sequenceof
r2N=exp(2πis)=1
forsan integer. The lower value, 2 N,f o rr=1 corresponds to p=q. The orthogonality
ofthecorrespondingtrigonometricfunctionsisleftas Exercise14.6.1.
Discrete Fourier Transform
To simplify the notation and to make more direct contact with physics, we introduce the
(reciprocal) ω-space,or angularfrequency,with
ωp=2πp
T,p=0,1,2,...,2N−1. (14.86)
We make prange over the same integers as k. The exponential exp (±2πiptk/T)of
Eq. (14.84) becomes exp (±iωptk). The choice of whether to use the +or the−sign is
amatterofconvenienceorconvention.Inquantummechanicsthenegativesignisselected
whenexpressingthetimedependence.
Consider a function of time defined (measured) at the discrete time values tk. Then we
construct
F(ωp)=1
2N2N−1summationdisplay
k=0f(tk)eiωptk. (14.87)
Employingtheorthogonalityrelation,weobtain
1
2N2N−1summationdisplay
p=0parenleftbig
eiωptmparenrightbig∗eiωptk=δmk, (14.88)
andthenreplacingthesubscript mbyk, wefindthattheamplitudes f(tk)become
f(tk)=2N−1summationdisplay
p=0F(ωp)e−iωptk. (14.89)
The time function f(tk),k=0,1,2,...,2N−1, and the frequency function F(ωp),p=
0,1,2,...,2N−1, arediscreteFouriertransforms ofeachother.12CompareEqs. (14.87)
12Thetwo transform equations may be symmetrized with aresulting (2N)−1/2ineachequation ifdesired.
916 Chapter 14 Fourier Series
and (14.89) with the corresponding continuous Fourier transforms, Eqs. (15.22) and
(15.23)ofChapter15.
Limitations
Taken as a pair of mathematical relations, the discrete Fourier transforms are exact. We
can say that the 2 N2N-component vectors exp (−iωptk),k=0,1,2,...,2N−1, form
a complete set13spanning the tk-space. Then f(tk)in Eq. (14.89) is simply a particular
linear combination of these vectors. Alternatively, we may take the 2 Nmeasured compo-
nentsf(tk)as defining a 2 N-component vector in tk-space. Then, Eq. (14.87) yields the
2N-component vector F(ωp)in thereciprocal ωp-space. Equations (14.87) and (14.89)
becomematrixequations,with exp (iωptk)/(2N)1/2theelementsof aunitarymatrix.
The limitations of the discrete Fourier transform arise when we apply Eqs. (14.87) and
(14.89) to physical systems and attempt physical interpretation and the limit F(ωp)→
F(ω).Example14.6.1illustratestheproblemsthatcanoccur.Themostimportantprecau-
tion to be taken to avoid trouble is to take Nsufficiently large so that there is no angular
frequency component of a higher angular frequency than ωN=2πN/T. For details on
errors and limitations in the use of the discrete Fourier transform we refer to Hamming in
theAdditionalReadings.
Example 14.6.1 DISCRETE FOURIER TRANSFORM —A LIASING
Considerthesimplecaseof T=2π,N=2,andf(tk)=costk.From
tk=kT
4=kπ
2,k=0,1,2,3, (14.90)
f(tk)=cos(tk)isrepresentedbythefour-componentvector
f(tk)=(1,0,−1,0). (14.91)
Thefrequencies, ωp, aregivenbyEq. (14.86):
ωp=2πp
T=p. (14.92)
Clearly, cos tkimpliesa p=1 componentandnootherfrequencycomponents.
Thetransformationmatrix
exp(iωptk)
2N=exp(ipkπ/2)
2N
becomes
1
4
1 111
1i−1−i
1−11−1
1−i−1i
. (14.93)
13By Eq. (14.85) thesevectors areorthogonal and aretherefore linearly independent.
14.6 Discrete Fourier Transform 917
Note that the 2 N×2Nmatrix has only 2 Nindependent components. It is the repetition
ofvaluesthat makesthefast Fouriertransform techniquepossible .
Operatingoncolumnvector f(tk), wefindthatthismatrixyieldsacolumnvector
F(ωp)=parenleftbig
0,1
2,0,1
2parenrightbig
. (14.94)
Apparently, there is a p=3 frequency component present. We reconstruct f(tk)by
Eq.(14.89), obtaining
f(tk)=1
2e−itk+1
2e−3itk. (14.95)
Takingrealparts, wecanrewritetheequationas
ℜf(tk)=1
2costk+1
2cos3tk. (14.96)
Obviously, this result, Eq. (14.96), is not identical with our original f(tk)costk.B u t
costk=1
2costk+1
2cos3tkattk=0,π/2,π; and 3π/2. The cos tkand cos3 tkmimic
each other because of the limited number of data points (and the particular choice of data
points).Thiserrorofonefrequencymimickinganotherisknownas aliasing.Theproblem
canbeminimizedbytakingmoredatapoints. /squaresolid
Fast Fourier Transform
The fast Fourier transform is a particular way of factoring and rearranging the terms in
the sums of the discrete Fourier transform. Brought to the attention of the scientific com-
munity by Cooley and Tukey,14its importance lies in the drastic reduction in the number
of numerical operations required. Because of the tremendous increase in speed achieved
(and reduction in cost), the fast Fourier transform has been hailed as one of the few really
significantadvancesinnumericalanalysisinthepastfew decades.
ForNtime values (measurements), a direct calculation of a discrete Fourier transform
wouldmeanabout N2multiplications.For Napowerof2,thefastFouriertransformtech-
nique of Cooley and Tukey cuts the number of multiplications required to (N/2)log2N.
IfN=1024(=210), the fast Fourier transform achieves a computational reduction by a
factor of over 200. This is why the fast Fourier transform is called fast and why it has rev-
olutionized the digital processing of waveforms. Details on the internal operation will be
foundinthepaperbyCooleyandTukeyandinthepaperbyBergland.15
14J. W.Cooley and J.W. Tukey, Math. Comput. 19: 297 (1965).
15G. D. Bergland, A guided tour of the fast Fourier transform, IEEE Spectrum , July, pp. 41–52 (1969); see also, W. H. Press,
B.P.Flannery,S.A.Teukolsky,andW.T.Vetterling, NumericalRecipes ,2nded.,Cambridge,UK:CambridgeUniversityPress
(1996), Section 12.3.
918 Chapter 14 Fourier Series
Exercises
14.6.1 Derivethetrigonometricforms ofdiscreteorthogonalitycorrespondingtoEq. (14.84):
2N−1summationdisplay
k=0cosparenleftbigg2πptk
Tparenrightbigg
sinparenleftbigg2πqtk
Tparenrightbigg
=0
2N−1summationdisplay
k=0cosparenleftbigg2πptk
Tparenrightbigg
cosparenleftbigg2πqtk
Tparenrightbigg
=
0,p/negationslash=q
N, p=q/negationslash=0,N
2N, p=q=0,N
2N−1summationdisplay
k=0sinparenleftbigg2πptk
Tparenrightbigg
sinparenleftbigg2πqtk
Tparenrightbigg
=
0,p/negationslash=q
N, p=q/negationslash=0,N
0,p=q=0,N.
Hint.Trigonometricidentitiessuchas
sinAcosB=1
2bracketleftbig
sin(A+B)+sin(A−B)bracketrightbig
areuseful.
14.6.2 Equation (14.84) exhibits orthogonality summing over time points. Show that we have
thesameorthogonalitysummingoverfrequencypoints
1
2N2N−1summationdisplay
p=0parenleftbig
eiωptmparenrightbig∗eiωptk=δmk.
14.6.3 Showindetailhowtogofrom
F(ωp)=1
2N2N−1summationdisplay
k=0f(tk)eiωptk
to
f(tk)=2N−1summationdisplay
p=0F(ωp)e−iωptk.
14.6.4 Thefunctions f(tk)andF(ωp)arediscreteFouriertransformsofeachother.Derivethe
followingsymmetryrelations:
(a) Iff(tk)is real,F(ωp)isHermitiansymmetric;thatis,
F(ωp)=F∗parenleftbigg4πN
T−ωpparenrightbigg
.
(b) Iff(tk)is pureimaginary,
F(ωp)=−F∗parenleftbigg4πN
T−ωpparenrightbigg
.
Note. The symmetry of part (a) is an illustration of aliasing. The frequency 4 πN/T−
ωpmasqueradesasthefrequency ωp.
14.7 Fourier Expansions of Mathieu Functions 919
14.6.5 GivenN=2,T=2π,andf(tk)=sintk,
(a) find F(ωp),p=0,1,2,3,and
(b) reconstruct f(tk)fromF(ωp)andexhibitthealiasingof ω1=1 andω3=3.
ANS.(a) F(ωp)=(0,i/2,0,−i/2)
(b)f(tk)=1
2sintk−1
2sin3tk.
14.6.6 ShowthattheChebyshevpolynomials Tm(x)satisfy adiscreteorthogonalityrelation
1
2Tm(−1)Tn(−1)+N−1summationdisplay
s=1Tm(xs)Tn(xs)+1
2Tm(1)Tn(1)=
0,m/negationslash=n
N/2,m=n/negationslash=0
N, m=n=0.
Here,xs=cosθs, wherethe (N+1)θsare equallyspacedalongthe θ-axis:
θs=sπ
N,s=0,1,2,...,N.
14.7 F OURIER EXPANSIONS OF MATHIEU FUNCTIONS
As a realistic application of Fourier series we now derive first integral equations satisfied
byMathieufunctions,from whichsubsequentlytheirFourierseriesare obtained.
Integral Equations and Fourier Series for Mathieu
Functions
Our first goal is to establish Whittaker’s integral equations that Mathieu functions satisfy,
fromwhichwethenobtaintheirFourierseriesrepresentations.
We startfromanintegralrepresentation
V(r)=integraldisplayπ
−πf(z+ixcosθ+iysinθ,θ)dθ (14.97)
of a solution Vof Laplace’s equation with a twice differentiable function f(v,θ).Apply-
ing∇2toVweverifythatitobeysLaplace’sPDE.SeparatingvariablesinLaplace’sPDE
suggestschoosingtheproductform f(v,θ)=ekvφ(θ).Substitutingtheellipticalvariables
ofEq. (13.163)werewrite Vas
R(ξ)/Phi1(η)ekz=integraldisplayπ
−πφ(θ)ek(z+iccoshξcosηcosθ+icsinhξsinηsinθ)dθ (14.98)
with normalization R(0)=1.Sinceξandηare independent variables we may set ξ=0,
whichleadstoWhittaker’s integralrepresentation
/Phi1(η)=integraldisplayπ
−πφ(θ)exp(ickcosθcosη)dθ, (14.99)
920 Chapter 14 Fourier Series
whereck=2√qfrom Eq. (13.180). Clearly, /Phi1is evenin the variable ηandperiodicwith
periodπ.In order to prove that φ∼/Phi1we check how φ(θ)is constrained when /Phi1(η)is
takentoobeytheangularMathieuODE
d2/Phi1
dη2+(λ−2qcos2η)/Phi1(η)
=integraldisplayπ
−πφ(θ)exp(ickcosθcosη)
·bracketleftbig
λ−2qcos2η+(ickcosθsinη)2−ickcosθcosηbracketrightbig
dθ.(14.100)
Hereweintegratethelasttermontheright-handsidebyparts, obtaining
d2/Phi1
dη2+(λ−2qcos2η)/Phi1(η)
=φ(θ)(−ickcosηsinθ)exp(ickcosθcosη)vextendsinglevextendsinglevextendsingleπ
θ=−π
+integraldisplayπ
−πφ(θ)exp(ickcosθcosη)[λ−2qcos2η−ickcosθcosη]dθ
+integraldisplayπ
−πbracketleftbig
−φ′(θ)(−ickcosηsinθ)+φ(θ)ickcosηcosθbracketrightbig
exp(ickcosθcosη)dθ
=integraldisplayπ
−πexpickcosθcosηbracketleftbig
φ(θ)(λ−2qcos2η)+φ′(θ)ickcosηsinθbracketrightbig
dθ,(14.101)
where the integrated term vanishes if φ(−π)=φ(π),which we assume to be the case.
Integratingoncemorebyparts yields
d2/Phi1
dη2+(λ−2qcos2η)/Phi1(η)=−φ′(θ)expickcosθcosηvextendsinglevextendsinglevextendsingleπ
θ=−π
+integraldisplayπ
−πexp(ickcosθcosη)bracketleftbig
φ(θ)(λ−2qcos2η)+φ′′(θ)bracketrightbig
dθ,(14.102)
where the integrated term vanishes if φ′is periodic with period π, which we assume is
thecase.Therefore,if φ(θ)obeystheangularMathieuODE,sodoestheintegral /Phi1(η),in
Eq. (14.99). As a consequence, φ(θ)∼/Phi1(θ), where the constantmay be a functionof the
parameter q.
Thus,wehavethemainresultthatasolution /Phi1(η)ofMathieu’sODEthatiseveninthe
variableηsatisfies theintegralequation
/Phi1(η)=/Lambda1n(q)integraldisplayπ
−πe2i√qcosθcosη/Phi1(θ)dθ. (14.103)
When these Mathieu functions are expanded in a Fourier cosine series and normalized so
thattheleadingtermis cos nη,theyare denotedbyce n(η,q).
14.7 Fourier Expansions of Mathieu Functions 921
Similarly, solutions of Mathieu’s ODE that are odd in ηwith leading term sin nηin a
Fourierseriesaredenotedbyse n(η,q),andtheycansimilarlybeshowntoobeytheintegral
equation
sen(η,q)=sn(q)integraldisplayπ
−πsinparenleftbig
2i√qsinηsinθparenrightbig
sen(θ,q)dθ. (14.104)
We now come to the Fourier expansion for the angular Mathieu functions and start
with
se1(η,q)=sinη+∞summationdisplay
ν=1βν(q)sin(2ν+1)η, βν(q)=∞summationdisplay
µ=νβ(ν)
µqµ(14.105)
as a paradigm for the systematic construction of Mathieu functions of odd parity. Notice
the key point that the coefficient βνof sin(2ν+1)ηin the Fourier series depends on the
parameter qand is expanded in a power series. Moreover, se 1is normalized so that the
coefficient of the leading term, sin η,is unity, that is, independent of q.This feature will
become important when se 1is substituted into the angular Mathieu ODE to determine the
eigenvalue λ(q).
Thefactthatthe βνpowerseriesin qstartswithexponent νcanbeprovedbyasimpler
butsimilarseries for se 1(η,q):
se1(η,q)=∞summationdisplay
ν=0γν(q)sin2ν+1η, γ ν(q)=∞summationdisplay
µ=0γ(ν)
µqµ,(14.106)
whichisusefulfor thisdemonstrationalone.However,sinceweneedtoexpand
sin2ν+1η=νsummationdisplay
m=0Bνmsin(2m+1)η (14.107)
withFouriercoefficients
Bνm=1
πintegraldisplayπ
−πsin2ν+1ηsin(2m+1)ηdη=(−1)m
22νparenleftbigg2ν+1
ν−mparenrightbigg
(14.108)
that we can look up in a table of integrals (see Gradshteyn and Ryzhik in the Additional
Readings of Chapter 13), this proof gives us an opportunity to introduce the Bνmthat are
nonzeroonlyif m≤νandareimportantingredientsoftherecursionrelationsforthelead-
ing terms of se 1(and all other Mathieu functions of odd parity). Substituting Eq. (14.107)
intoEq. (14.106)weobtain
se1(η,q)=∞summationdisplay
ν=0νsummationdisplay
m=0Bνmγν(q)sin(2m+1)η. (14.109)
Comparingthisexpressionfor se 1withEq.(14.105)wefind
βν(q)=∞summationdisplay
m=νBmνγm(q). (14.110)
Here,thesumstartswith m=νbecauseBmν=0f o rm<ν.
922 Chapter 14 Fourier Series
NextwesubstituteEq.(14.106)intotheintegralEq.(14.104)for n=1,whereweinsert
thepowerseries for sin (2i√qsinηsinθ).This yields
se1(η,q)
2πs1(q)=1
2πintegraldisplayπ
−πsin(2i√qsinηsinθ)se1(θ,q)dθ
=1
2πs1(q)∞summationdisplay
m=0γm(q)sin2m+1η
=i√q∞summationdisplay
m,ν=0qmγν(q)sin2m+1η22m+1
(2m+1)!1
2πintegraldisplayπ
−πsin2ν+2m+2θdθ,
(14.111)
fromwhichweobtaintherecursionrelations
γm(q)=22m+1
(2m+1)!qmi√qs1(q)∞summationdisplay
ν=0γν(q)integraldisplayπ
−πsin2ν+2m+2θdθ, (14.112)
uponcomparingcoefficientsofsin2m+1η.Thisshowsthatthepowerseriesfor γm(q)starts
withqm.Using Eq. (14.110) proves that the power series for βm(q)also starts with qm,
and this confirms Eq. (14.105). The integral in Eq. (14.112) can be evaluated analytically
and expressed via the beta function (Chapter 8) in terms of ratios of factorials, but we do
notneedthisformulahere.
Our next goal is to establish a recursion relation for the leading term β(ν)
νof se1,i n
Eq.(14.105).WesubstituteEq.(14.105)intotheintegralEq.(14.104)for n=1,wherewe
insertthepowerseries for sin (2i√qsinηsinθ)again,alongwiththeexpansion
1
2πs1(q)=i∞summationdisplay
m=0αmqm+1/2. (14.113)
Here, the extra factor, i√q, cancels the corresponding factor from the sine in the integral
equation.Thisyields
∞summationdisplay
ν=0β(λ)
µqµ+νsin2ν+1η22ν+1
(2ν+1)!1
2πintegraldisplayπ
−πsin2ν+2λ+2θdθ
=∞summationdisplay
m,µ=0αmqm+µβ(ν)
µsin(2ν+1)η. (14.114)
Here, we replace sin2ν+1ηby sin(2m+1)ηusing Eq. (14.107). Upon comparing the co-
efficientsof qNsin(2ν+1)ηforN=µ+νweobtaintherecursionrelation
Nsummationdisplay
ν=nN−νsummationdisplay
λ=0β(λ)
N−ν22ν
(2ν+1)!BνnBνλ=N−nsummationdisplay
m=0αmβ(n)
N−m. (14.115)
14.7 Fourier Expansions of Mathieu Functions 923
Now we substitute Eq. (14.108) to obtain the main recursion relation for the leading
coefficients β(ν)
νof se1:
Nsummationdisplay
ν=nN−νsummationdisplay
λ=0β(λ)
N−ν
22ν(2ν+1)!parenleftbigg2ν+1
ν−nparenrightbiggparenleftbigg2ν+1
ν−λparenrightbigg
=N−nsummationdisplay
m=0αmβ(n)
N−m.(14.116)
Example 14.7.1 LEADING COEFFICIENTS OF se1
We evaluate Eq. (14.116) starting with N=0,n=0.For this case we find β(0)
0=α0β(0)
0,
orα0=1 because the coefficient of sin ηin se1,β(0)
0=1,by normalization. For N=1,
n=0 Eq. (14.116)yields
α0β(0)
1+α1β(0)
0=β(0)
1+1
4·3!parenleftbigg3
1parenrightbiggbracketleftbigg
β(0)
0parenleftbigg3
1parenrightbigg
+β(1)
0parenleftbigg3
0parenrightbiggbracketrightbigg
,(14.117)
whereβ(1)
0=0 andβ(0)
1drops out, a general feature. Of course, β(0)
1=0 because sin ηin
se1hascoefficientunity.Thisyields α1=3/8.
Thecase N=1,n=1 yields
−1
4·3!parenleftbigg3
0parenrightbigg
β(0)
0parenleftbigg3
1parenrightbigg
=α0β(1)
1, (14.118)
orβ(1)
1=−1/8.Theleadingtermisobtainedfromthegeneralcase n=N,
(−1)N
22N(2N+1)!parenleftbigg2N+1
0parenrightbigg
β(0)
0parenleftbigg2N+1
Nparenrightbigg
=α0β(N)
N, (14.119)
as
β(N)
N=(−1)N
22N(2N+1)!parenleftbigg2N+1
Nparenrightbigg
, (14.120)
which was first derived by Mathieu. For N=1 this formula reproduces our earlier result,
β(1)
1=−1/8. /squaresolid
In order to determine the first nonleading term β(N)
N+1of se1,Eq. (14.105), and the
eigenvalue λ1(q)we substitute se 1into the angular Mathieu ODE, Eq. (13.181), using
thetrigonometricidentities
2cos2ηsin(2ν+1)η=sin(2ν+3)η+sin(2ν−1)η
and
d2sin(2ν+1)η
dη2=−(2ν+1)2sin(2ν+1)η.
924 Chapter 14 Fourier Series
Thisyields
0=d2se1
dη2+(λ1−2qcos2η)se1=q(sinη−sin3η)+λ1sinη−sinη
+∞summationdisplay
ν=1bracketleftbig
λ1−(2ν+1)2bracketrightbigbracketleftbigg(−1)νqν
22νν!(ν+1)!+β(ν)
ν+1qν+1+···bracketrightbigg
sin(2ν+1)η
−q∞summationdisplay
ν=1bracketleftbigg(−q)ν
22νν!(ν+1)!+β(ν)
ν+1qν+1+···bracketrightbiggparenleftbig
sin(2ν+3)η+sin(2ν−1)ηparenrightbig
=parenleftbigg
λ1−1+q−qbracketleftbigg
−q
222!+β(1)
2q2+···bracketrightbiggparenrightbigg
sinη
+sin3ηbracketleftbigg
−q−qparenleftbiggq2
242!3!+β(2)
3q3parenrightbigg
+parenleftbig
λ1−32parenrightbigparenleftbigg
−q
222!+β(1)
2q2parenrightbiggbracketrightbigg
+sin(2ν+1)ηbracketleftbig
λ1−(2ν+1)2bracketrightbigparenleftbigg(−q)ν
22νν!(ν+1)!+β(ν)
ν+1qν+1parenrightbigg
−qsin(2ν+1)ηparenleftbigg(−q)ν+1
22(ν+1)(ν+1)!(ν+2)!+β(ν+1)
ν+2qν+2parenrightbigg
−qsin(2ν+1)ηparenleftbigg(−q)ν−1
22(ν−1)(ν−1)!ν!+β(ν−1)
νqνparenrightbigg
+···. (14.121)
In this series the coefficient of each power of qwithin different sine terms must vanish;
thatof sin ηbeingzeroyieldstheeigenvalue
λ1(q)=1−q−1
8q2+β(1)
2q3+···, (14.122)
withβ(1)
2=1/26coming from the vanishing coefficient of q2in sin3η.Setting the coeffi-
cientof(−q)νin sin(2ν+1)ηequaltozeroyieldstheidentity
bracketleftbig
1−(2ν+1)2bracketrightbig1
22νν!(ν+1)!+1
22(ν−1)(ν−1)!ν!=0, (14.123)
which verifies the correct determination of the leading terms β(ν)
νin Eq. (14.120). The
vanishingcoefficientof qν+1in sin(2ν+1)ηyields
(−1)ν+1
22νν!(ν+1)!+bracketleftbig
1−(2ν+1)2bracketrightbig
β(ν)
ν+1−β(ν−1)
ν=0, (14.124)
whichimpliesthemain recursionrelationfornonleadingcoefficients ,
4ν(ν+1)β(ν)
ν+1=−β(ν−1)
ν+(−1)ν+1
22νν!(ν+1)!, (14.125)
forthefirst nonleadingterms.We verifythat
β(ν)
ν+1=(−1)ν+1ν
22ν+2(ν+1)!2(14.126)
14.7 Fourier Expansions of Mathieu Functions 925
satisfies this recursion relation. Higher nonleading terms may be obtained by setting to
zerothecoefficientof qν+2, etc.AltogetherwehavederivedtheFourierseries for
se1(η,q)=sinη+∞summationdisplay
ν=1bracketleftbigg(−q)ν
22νν!(ν+1)!+(−q)ν+1ν
22ν+2(ν+1)!2+···bracketrightbigg
·sin(2ν+1)η.
(14.127)
AsimilartreatmentyieldstheFourierseriesforse 2n+1(η,q)andse2n(η,q).Aninvariance
ofMathieu’sODEleadstothesymmetryrelation
ce2n+1(η,q)=(−1)nse2n+1(η+π/2,−q), (14.128)
which allows us to determine the ce 2n+1of period 2 πfrom se 2n+1.Similarly,
ce2n(η+π/2,−q)=se2n(η,q)relatestheseMathieufunctionsofperiod πtoeachother.
Finally,webrieflyoutlineaderivationoftheFourierseries for
ce0(η,q)=1+∞summationdisplay
n=1βn(q)cos2nη, β n(q)=∞summationdisplay
m=nβ(n)
mqm,(14.129)
as a paradigm for the Mathieu functions of period π.Note that this normalization agrees
with Whittaker and Watson and with Hochstadtin the AdditionalReadings of Chapter 13,
whereas in AMS-55 (for the full reference see footnote 4 in Chapter 5) ce 0differs by a
factorof 1 /√
2.ThesymmetryrelationfromtheMathieuODE,
ce0parenleftbiggπ
2−η,−qparenrightbigg
=ce0(η,q), (14.130)
implies
βn(−q)=(−1)nβn(q); (14.131)
thatis,β2ncontainsonlyevenpowersof qandβ2n+1onlyoddpowers.
The fact that the power series for βn(q)in Eq. (14.129) starts with qncan be proved by
thesimilarexpansion
ce0(η,q)=∞summationdisplay
n=0γn(q)cos2nη, γ n(q)=∞summationdisplay
µ=0γ(n)
µqµ, (14.132)
asforse 1inEqs.(14.105)to(14.112).SubstitutingEq.(14.132)intotheintegralequation
ce0(η,q)=c0(q)integraldisplayπ
−πe2i√qcosθcosηce(θ,q)dθ, (14.133)
inserting the power series for the exponential function (odd powers cos2m+1θdrop out)
andequatingthecoefficientsof cos2mηyields
γm(q)=c1(q)(−q)m22m
(2m)!∞summationdisplay
µ=0γµ(q)integraldisplayπ
−πcos2m+2µθdθ. (14.134)
926 Chapter 14 Fourier Series
Thisrecursionrelationshowsthatthepowerseries for γm(q)startswith qm.Weexpand
cos2nη=nsummationdisplay
m=0Anmcos2mη (14.135)
withFouriercoefficients
Anm=2
πintegraldisplayπ/2
−π/2cos2nηcos2mηdη=1
22n−1parenleftbigg2n
n−mparenrightbigg
, (14.136)
which are nonzero only when m≤n.Using this result to replace the cosine powers in
Eq.(14.132)by cos2 mηweobtain
βn(q)=∞summationdisplay
m=nAmnγm(q), (14.137)
confirmingEq. (14.129).
Proceeding as for se 1in Eqs. (14.113) to (14.120) we substitute Eq. (14.129) into the
integralEq.(14.133)andobtain
∞summationdisplay
m,µ,ν,λ=0(−1)m22m
(2m)!qµ+mmsummationdisplay
ν=0Amνcos(2νη)Amλ
=∞summationdisplay
m,µ,ν=0αmβ(ν)
µqm+µcos(2νη), (14.138)
with
1
2πc1(q)=∞summationdisplay
m=0αmqm. (14.139)
Uponcomparingthecoefficientof qNcos(2νη)withN=m+µ,weextractthe recursion
relationforleadingcoefficients β(n)
nof ce0
Nsummationdisplay
m=ν(−1)m22m
(2m)!Amνmsummationdisplay
λ=0β(λ)
N−mAmλ=Nsummationdisplay
m=0αmβ(ν)
N−m, (14.140)
withAnminEq. (14.136).
Example 14.7.2 LEADING COEFFICIENTS FOR ce0
Thecase N=0,ν=0 ofEq. (14.140)yields
A2
00β(0)
0=α0β(0)
0, (14.141)
withA00=1andβ(0)
0=1fromnormalizingtheleadingtermofce 0tounitysothat α0=1
results.
Thecase N=1,ν=0 yields
A00β(0)
1A00−2A10bracketleftbig
β(0)
0A10+β(1)
0A11bracketrightbig
=α0β(0)
1+α1β(0)
0,(14.142)
14.7 Fourier Expansions of Mathieu Functions 927
withβ(1)
0=0 byEq. (14.129).This simplifiesto
β(0)
1−1
2β(0)
0=α1β(0)
0+β(0)
1, (14.143)
whereβ(0)
1drops out. We know already that β(0)
1=0 from the leading term unity of ce 0.
Therefore, α1=−1/2.
Forthecase N=1,ν=1 weobtain
−2A11β(0)
0A10=α0β(1)
1, (14.144)
withA10=1/2=A11,from which β(1)
1=−1/2 follows. For the case N=2,ν=2w e
find
24
4!23β(0)
0A20=α0β(2)
2, (14.145)
withA20=3/8,from which β(2)
2=2−4follows.Thegeneralcase N,ν=Nyields
(−1)N22N
(2N)!ANNβ(0)
0AN0=α0β(N)
N, (14.146)
withANN=2−2N+1,AN0=1
22N−1parenleftbig2N
Nparenrightbig
, fromwhichtheleadingterm
β(N)
N=(−1)N
22N−1(2N)!parenleftbigg2N
Nparenrightbigg
=(−1)N
22N−1N!2(14.147)
follows. /squaresolid
The nonleading terms β(N)
N+1of ce0are best determined from the angular Mathieu ODE
by substitution of Eq. (14.129), in analogy with se 1,Eqs. (14.121) to (14.127). Using the
identities
2cos(2nη)cos2η=cos(2n+2)η+cos(2n−2)η,
d2
dη2cos(2nη)=−(2n)2cos(2nη), (14.148)
weobtain
d2ce0
dη2+parenleftbig
λ0(q)−2qcos2ηparenrightbig
ce0=0
=λ0(q)−qparenleftbigg
−q
2+7
27q3parenrightbigg
+···
+∞summationdisplay
n=1parenleftbig
λ0−4n2parenrightbig
cos(2nη)bracketleftbigg(−q)n
22n−1n!2+β(n)
n+2qn+2bracketrightbigg
−q∞summationdisplay
n=1bracketleftbigg(−q)n
22n−1n!2+β(n)
n+2qn+2bracketrightbigg
×bracketleftbig
cos(2n+2)η+cos(2n−2)ηbracketrightbig
. (14.149)
928 Chapter 14 Fourier Series
Settingthecoefficientof cos (2nη)forn=0 tozeroyieldstheeigenvalue
λ0=−1
2q2+7
27q4+···. (14.150)
Thecoefficientof cos (2nη)qnyieldsanidentity,
(−1)n+14n2
22n−1n!2+(−1)n
22n−3(n−1)!2=0, (14.151)
which shows that the leading term in Eq. (14.147) was correctly determined. The coeffi-
cientofqn+2cos(2nη)yieldstherecursionrelation
−4n2β(n)
n+2−β(n−1)
n+1+(−1)n+1
22nn!2+(−1)n
22n+1(n+1)!2=0. (14.152)
Itis straightforwardtocheckthat
β(n)
n+2=(−1)n+1n(3n+4)
22n+3(n+1)!2(14.153)
satisfiesthisrecursionrelation.Altogetherwehavederivedtheformula
ce0(η,q)=1+cos2ηbracketleftbigg
−1
2q2+7
27q3+···bracketrightbigg
+cos4ηbracketleftbiggq2
25+···bracketrightbigg
+cos6ηbracketleftbigg
−q3
2732+···bracketrightbigg
=1+∞summationdisplay
n=1cos(2nη)bracketleftbigg(−q)n
22n−1n!2+(−1)n+1n(3n+4)qn+2
22n+3(n+1)!2+···bracketrightbigg
.
(14.154)
Similarlyonecanderive
ce1(η,q)=cosη+∞summationdisplay
n=1cos(2n+1)ηbracketleftbigg(−q)n
22nn!(n+1)!−(−q)n+1n
22n+2(n+1)12+···bracketrightbigg
,
(14.155)
whoseeigenvalueis givenbythepowerseries
λ1(q)=1+q−1
8q2−1
26q3+···. (14.156)
14.7 Additional Readings 929
FIGURE 14.12AngularMathieufunctions.(FromGutiérrez-Vega etal.,
Am. J.Phys. 71:233(2003).)
Exercises
14.7.1 Determinethenonleadingcoefficients β(n)
n+2forse1.Deriveasuitablerecursionrelation.
14.7.2 Determinethenonleadingcoefficients β(n)
n+4force0.Derivethecorrespondingrecursion
relation.
14.7.3 Derivetheformulafor ce 1, Eq.(14.155), andits eigenvalue,Eq. (14.156).
AdditionalReadings
Carslaw, H. S., Introduction to the Theory of Fourier’s Series and Integrals , 2nd ed. London: Macmillan (1921);
3rd ed., paperback, New York: Dover (1952). This is a detailed and classic work; includes a considerable
discussion of Gibbs phenomenon in Chapter IX.
Hamming, R. W., Numerical Methods for Scientists and Engineers , 2nd ed. New York: McGraw-Hill (1973),
reprinted Dover (1987). Chapter 33 provides an excellentdescription ofthe fast Fourier transform.
Jeffreys,H.,andB.S.Jeffreys, MethodsofMathematicalPhysics ,3rded.Cambridge,UK:CambridgeUniversity
Press (1972).
Kufner, A., and J. Kadlec, Fourier Series . London: Iliffe (1971). This book is a clear account of Fourier series in
the context ofHilbert space.
930 Chapter 14 Fourier Series
Lanczos, C., Applied Analysis , Englewood Cliffs, NJ: Prentice-Hall (1956), reprinted Dover (1988). The book
givesawell-writtenpresentationoftheLanczosconvergencetechnique(whichsuppressestheGibbsphenom-
enon oscillations). This and several other topics are presented from the point of view of a mathematician who
wantsuseful numerical results and not just abstractexistence theorems.
Oberhettinger, F., Fourier Expansions, ACollection of Formulas . NewYork, AcademicPress (1973).
Zygmund, A., Trigonometric Series . Cambridge, UK: Cambridge University Press (1988). The volume contains
an extremely complete exposition, including relatively recent results in therealmof pure mathematics.
CHAPTER 15
INTEGRAL TRANSFORMS
15.1 I NTEGRAL TRANSFORMS
Frequently in mathematical physics we encounter pairs of functions related by an expres-
sionoftheform
g(α)=integraldisplayb
af(t)K(α,t)dt. (15.1)
The function g(α)is called the (integral) transform of f(t)by the kernel K(α,t).
The operation may also be described as mapping a function f(t)int-space into another
function, g(α),i nα-space. This interpretation takes on physical significance in the time-
frequency relation of Fourier transforms, as in Example 15.3.1, and in the real space–
momentumspacerelationsinquantumphysicsof Section15.6.
Fourier Transform
One of the most useful of the infinite number of possible transforms is the Fourier trans-
form,givenby
g(ω)=1√
2πintegraldisplay∞
−∞f(t)eiωtdt. (15.2)
Two modifications of this form, developed in Section 15.3, are the Fourier cosine and
Fouriersinetransforms:
gc(ω)=radicalbigg
2
πintegraldisplay∞
0f(t)cosωtdt, (15.3)
gs(ω)=radicalbigg
2
πintegraldisplay∞
0f(t)sinωtdt. (15.4)
931
932 Chapter 15 Integral Transforms
TheFouriertransformisbasedonthekernel eiωtanditsrealandimaginarypartstakensep-
arately, cos ωtand sinωt. Because these kernels are the functions used to describe waves,
Fouriertransformsappearfrequentlyinstudiesofwavesandtheextractionofinformation
from waves, particularly when phase information is involved. The output of a stellar in-
terferometer, for instance, involves a Fourier transform of the brightness across a stellar
disk.TheelectrondistributioninanatommaybeobtainedfromaFouriertransformofthe
amplitude of scattered X-rays. In quantum mechanics the physical origin of the Fourier
relationsofSection15.6isthewavenatureofmatterandourdescriptionofmatterinterms
ofwaves.
Example 15.1.1 FOURIER TRANSFORM OF GAUSSIAN
TheFouriertransform ofaGaussianfunction e−a2t2,
g(ω)=1√
2πintegraldisplay∞
−∞e−a2t2eiωtdt,
canbedoneanalyticallybycompletingthesquareintheexponent,
−a2t2+iωt=−a2parenleftbigg
t−iω
2a2parenrightbigg2
−ω2
4a2,
whichwecheckbyevaluatingthesquare.Substitutingthisidentityweobtain
g(ω)=1√
2πe−ω2/4a2integraldisplay∞
−∞e−a2t2dt,
upon shifting the integration variable t→t+iω
2a2.This is justified by an application of
Cauchy’s theorem to the rectangle with vertices −T, T, T+iω
2a2,−T+iω
2a2forT→
∞,noting that the integrand has no singularities in this region and that the integrals over
the sides from ±Tto±T+iω
2a2become negligible for T→∞.Finally we rescale the
integrationvariableas ξ=atintheintegral(seeEqs. (8.6) and(8.8)):
integraldisplay∞
−∞e−a2t2dt=1
aintegraldisplay∞
−∞e−ξ2dξ=√π
a.
Substitutingtheseresultswefind
g(ω)=1
a√
2expparenleftbigg
−ω2
4a2parenrightbigg
,
againaGaussian,butin ω-space.Thebigger ais,thatis,thenarrowertheoriginalGaussian
e−a2t2is, thewiderisits Fouriertransform ∼e−ω2/4a2. /squaresolid
15.1 Integral Transforms 933
Laplace, Mellin, and Hankel Transforms
Threeotherusefulkernelsare
e−αt,tJn(αt), tα−1.
Thesegiverise tothefollowingtransforms
g(α)=integraldisplay∞
0f(t)e−αtdt,Laplacetransform (15.5)
g(α)=integraldisplay∞
0f(t)tJn(αt)dt, Hankeltransform (Fourier–Bessel) (15.6)
g(α)=integraldisplay∞
0f(t)tα−1dt,Mellintransform . (15.7)
Clearly,thepossibletypesareunlimited.Thesetransformshavebeenusefulinmathemati-
calanalysisandinphysicalapplications.WehaveactuallyusedtheMellintransformwith-
out calling it by name; that is, g(α)=(α−1)!is the Mellin transform of f(t)=e−t. See
E.C. Titchmarsh, IntroductiontotheTheoryofFourierIntegrals ,2nded.,NewYork:Ox-
fordUniversityPress(1937),formoreMellintransforms.Ofcourse,wecouldjustaswell
sayg(α)=n!/αn+1istheLaplacetransformof f(t)=tn.Ofthethree,theLaplacetrans-
formisbyfarthemostused.ItisdiscussedatlengthinSections15.8to15.12.TheHankel
transform, a Fourier transform for a Bessel function expansion, represents a limiting case
of a Fourier–Bessel series. It occurs in potential problems in cylindrical coordinates and
hasbeenappliedextensivelyinacoustics.
Linearity
Alltheseintegraltransforms arelinear;thatis,
integraldisplayb
abracketleftbig
c1f1(t)+c2f2(t)bracketrightbig
K(α,t)dt
=c1integraldisplayb
af1(t)K(α,t)dt+c2integraldisplayb
af2(t)K(α,t)dt, (15.8)
integraldisplayb
acf(t)K(α,t)dt =cintegraldisplayb
af(t)K(α,t)dt, (15.9)
wherec1andc2are constants and f1(t)andf2(t)are functions for which the transform
operationisdefined.
Representingourlinearintegraltransformbytheoperator L,weobtain
g(α)=Lf(t). (15.10)
934 Chapter 15 Integral Transforms
FIGURE 15.1Schematicintegraltransforms.
Weexpectaninverseoperator L−1existssuchthat1
f(t)=L−1g(α). (15.11)
ForourthreeFouriertransforms L−1isgiveninSection15.3.Ingeneral,thedetermination
of the inverse transform is the main problem in using integral transforms. The inverse
Laplace transform is discussed in Section 15.12. For details of the inverse Hankel and
inverseMellintransformswerefer totheAdditionalReadingsattheendof thechapter.
Integral transforms have many special physical applications and interpretations that
are noted in the remainder of this chapter. The most common application is outlined in
Fig. 15.1. Perhaps an original problem can be solved only with difficulty, if at all, in the
original coordinates (space). It often happens that the transform of the problem can be
solved relatively easily. Then the inverse transform returns the solution from the trans-
formcoordinatestotheoriginalsystem.Example15.4.1andExercise15.4.1illustratethis
technique.
Exercises
15.1.1 TheFouriertransforms for afunctionoftwovariablesare
F(u,v)=1
2πintegraldisplay∞
−∞integraldisplay
f(x,y)ei(ux+vy)dxdy,
f(x,y)=1
2πintegraldisplay∞
−∞integraldisplay
F(u,v)e−i(ux+vy)dudv.
Usingf(x,y)=f([x2+y2]1/2),showthatthezero-orderHankeltransforms
F(ρ)=integraldisplay∞
0rf(r)J0(ρr)dr,
f(r)=integraldisplay∞
0ρF(ρ)J 0(ρr)dρ,
areaspecialcaseof theFouriertransforms.
1Expectationisnot proof, andhereproofofexistenceiscomplicatedbecauseweareactuallyin an infinite-dimensional Hilbert
space.Weshall prove existence inthe specialcasesof interest byactual construction.
15.1 Integral Transforms 935
This technique may be generalized to derive the Hankel transforms of order ν=
0,1
2,1,1
2,...(compare I. N. Sneddon, Fourier Transforms , New York: McGraw-Hill
(1951)).Amoregeneralapproach,validfor ν>−1
2,ispresentedinSneddon’s TheUse
of Integral Transforms (New York: McGraw-Hill (1972)). It might also be noted that
the Hankel transforms of nonintegral order ν=±1
2reduce to Fourier sine and cosine
transforms.
15.1.2 AssumingthevalidityoftheHankeltransform–inversetransformpairof equations
g(α)=integraldisplay∞
0f(t)Jn(αt)tdt,
f(t)=integraldisplay∞
0g(α)Jn(αt)αdα,
showthattheDiracdeltafunctionhasaBessel integralrepresentation
δ(t−t′)=tintegraldisplay∞
0Jn(αt)Jn(αt′)αdα.
This expression is useful in developing Green’s functions in cylindrical coordinates,
wheretheeigenfunctionsareBessel functions.
15.1.3 FromtheFouriertransforms, Eqs. (15.22)and(15.23),showthatthetransformation
t→lnx
iω→α−γ
leadsto
G(α)=integraldisplay∞
0F(x)xα−1dx
and
F(x)=1
2πiintegraldisplayγ+i∞
γ−i∞G(α)x−αdα.
These are the Mellin transforms. A similar change of variables is employed in Sec-
tion15.12toderivetheinverseLaplacetransform.
15.1.4 VerifythefollowingMellintransforms:
(a)integraldisplay∞
0xα−1sin(kx)dx=k−α(α−1)!sinπα
2,−1<α<1.
(b)integraldisplay∞
0xα−1cos(kx)dx=k−α(α−1)!cosπα
2,0<α<1.
Hint.Youcanforcetheintegralsintoatractableformbyinsertingaconvergencefactor
e−bxand(after integrating)letting b→0.Also, cos kx+isinkx=expikx.
936 Chapter 15 Integral Transforms
15.2 D EVELOPMENT OF THE FOURIER INTEGRAL
In Chapter 14 it was shown that Fourier series are useful in representing certain func-
tions (1) over a limited range [0,2π],[−L,L], and so on, or (2) for the infinite interval
(−∞,∞),if the function is periodic . We now turn our attention to the problem of rep-
resenting a nonperiodic function over the infinite range. Physically this means resolving a
singlepulseor wavepacketintosinusoidalwaves.
We have seen (Section 14.2) that for the interval [−L,L]the coefficients anandbn
couldbewrittenas
an=1
LintegraldisplayL
−Lf(t)cosnπt
Ldt, (15.12)
bn=1
LintegraldisplayL
−Lf(t)sinnπt
Ldt. (15.13)
TheresultingFourierseriesis
f(x)=1
2LintegraldisplayL
−Lf(t)dt+1
L∞summationdisplay
n=1cosnπx
LintegraldisplayL
−Lf(t)cosnπt
Ldt
+1
L∞summationdisplay
n=1sinnπx
LintegraldisplayL
−Lf(t)sinnπt
Ldt, (15.14)
or
f(x)=1
2LintegraldisplayL
−Lf(t)dt+1
L∞summationdisplay
n=1integraldisplayL
−Lf(t)cosnπ
L(t−x)dt. (15.15)
Wenowlettheparameter Lapproachinfinity,transformingthefiniteinterval [−L,L]into
theinfiniteinterval (−∞,∞).W es e t
nπ
L=ω,π
L=/Delta1ω, withL→∞.
Thenwehave
f(x)→1
π∞summationdisplay
n=1/Delta1ωintegraldisplay∞
−∞f(t)cosω(t−x)dt, (15.16)
or
f(x)=1
πintegraldisplay∞
0dωintegraldisplay∞
−∞f(t)cosω(t−x)dt, (15.17)
replacing the infinite sum by the integral over ω. The first term (corresponding to a0) has
vanished,assumingthatintegraltext∞
−∞f(t)dtexists.
It must be emphasized that this result (Eq. (15.17)) is purely formal. It is not intended
as a rigorous derivation, but it can be made rigorous (compare I. N. Sneddon, Fourier
Transforms , Section 3.2). We take Eq. (15.17) as the Fourier integral. It is subject to the
conditions that f(x)is (1) piecewise continuous, (2) piecewise differentiable, and (3) ab-
solutelyintegrable—thatis,integraltext∞
−∞|f(x)|dxisfinite.
15.2 Development of the Fourier Integral 937
Fourier Integral — Exponential Form
OurFourierintegral(Eq. (15.17)) maybeputintoexponentialformbynotingthat
f(x)=1
2πintegraldisplay∞
−∞dωintegraldisplay∞
−∞f(t)cosω(t−x)dt, (15.18)
whereas
1
2πintegraldisplay∞
−∞dωintegraldisplay∞
−∞f(t)sinω(t−x)dt=0; (15.19)
cosω(t−x)is an even function of ωand sinω(t−x)is an odd function of ω. Adding
Eqs. (15.18)and(15.19)(withafactor i), weobtainthe Fourierintegraltheorem
f(x)=1
2πintegraldisplay∞
−∞e−iωxdωintegraldisplay∞
−∞f(t)eiωtdt. (15.20)
The variable ωintroduced here is an arbitrary mathematical variable. In many physical
problems, however, it corresponds to the angular frequency ω. We may then interpret
Eq. (15.18) or (15.20) as a representation of f(x)in terms of a distribution of infinitely
long sinusoidal wave trains of angular frequency ω, in which this frequency is a continu-
ousvariable.
Dirac Delta Function Derivation
If theorder ofintegrationof Eq.(15.20) isreversed,wemayrewriteitas
f(x)=integraldisplay∞
−∞f(t)braceleftbigg1
2πintegraldisplay∞
−∞eiω(t−x)dωbracerightbigg
dt. (15.20a)
Apparently the quantity in curly brackets behaves as a delta function δ(t−x). We might
take Eq. (15.20a) as presenting us with a representation of the Dirac delta function. Alter-
natively,wetakeitas acluetoanewderivationoftheFourierintegraltheorem.
FromEq. (1.171b)(shiftingthesingularityfrom t=0t ot=x),
f(x)=limn→∞integraldisplay∞
−∞f(t)δn(t−x)dt, (15.21a)
whereδn(t−x)is a sequence defining the distribution δ(t−x). Note that Eq. (15.21a)
assumesthat f(t)is continuousat t=x.W etak e δn(t−x)tobe
δn(t−x)=sinn(t−x)
π(t−x)=1
2πintegraldisplayn
−neiω(t−x)dω, (15.21b)
usingEq. (1.174). SubstitutingintoEq.(15.21a), wehave
f(x)=limn→∞1
2πintegraldisplay∞
−∞f(t)integraldisplayn
−neiω(t−x)dωdt. (15.21c)
938 Chapter 15 Integral Transforms
Interchanging the order of integration and then taking the limit as n→∞,w eh a v e
Eq.(15.20), theFourierintegraltheorem.
With the understanding that it belongs under an integral sign, as in Eq. (15.21a), the
identification
δ(t−x)=1
2πintegraldisplay∞
−∞eiω(t−x)dω (15.21d)
providesaveryusefulrepresentationofthedeltafunction.
15.3 F OURIER TRANSFORMS —I NVERSION THEOREM
Letusdefineg(ω), theFouriertransformofthefunction f(t),by
g(ω)≡1√
2πintegraldisplay∞
−∞f(t)eiωtdt. (15.22)
Exponential Transform
Then,fromEq. (15.20), wehavetheinverserelation,
f(t)=1√
2πintegraldisplay∞
−∞g(ω)e−iωtdω. (15.23)
Note that Eqs. (15.22) and (15.23) are almost but not quite symmetrical, differing in the
signofi.
Here two points deserve comment. First, the 1 /√
2πsymmetry is a matter of choice,
not of necessity. Many authors will attach the entire 1 /2πfactor of Eq. (15.20) to one
of the two equations: Eq. (15.22) or Eq. (15.23). Second, although the Fourier integral,
Eq.(15.20),hasreceivedmuchattentioninthemathematicsliterature,weshallbeprimar-
ilyinterestedintheFouriertransformanditsinverse.Theyaretheequationswithphysical
significance.
WhenwemovetheFouriertransformpairtothree-dimensionalspace,itbecomes
g(k)=1
(2π)3/2integraldisplay
f(r)eik·rd3r, (15.23a)
f(r)=1
(2π)3/2integraldisplay
g(k)e−ik·rd3k. (15.23b)
The integrals are over all space. Verification, if desired, follows immediately by substitut-
ingtheleft-handsideofoneequationintotheintegrandoftheotherequationandusingthe
three-dimensional delta function.2Equation (15.23b) may be interpreted as an expansion
of a function f(r)in a continuum of plane wave eigenfunctions; g(k)then becomes the
amplitudeofthewave, exp (−ik·r).
2δ(r1−r2)=δ(x1−x2)δ(y1−y2)δ(z1−z2)with Fourier integral δ(x1−x2)=1
2πintegraltext∞
−∞exp[ik1(x1−x2)]dk1,etc.
15.3 Fourier Transforms — Inversion Theorem 939
Cosine Transform
Iff(x)is odd or even, these transforms may be expressed in a somewhat different
form. Consider first an even function fcwithfc(x)=fc(−x). Writing the exponential
ofEq. (15.22)intrigonometricform,wehave
gc(ω)=1√
2πintegraldisplay∞
−∞fc(t)(cosωt+isinωt)dt
=radicalbigg
2
πintegraldisplay∞
0fc(t)cosωtdt, (15.24)
the sinωtdependence vanishing on integration over the symmetric interval (−∞,∞).
Similarly,since cos ωtis even,Eqs. (15.23) transformsto
fc(x)=radicalbigg
2
πintegraldisplay∞
0gc(ω)cosωxdω. (15.25)
Equations(15.24)and(15.25)areknownas Fouriercosinetransforms.
Sine Transform
The corresponding pair of Fourier sine transforms is obtained by assuming that fs(x)=
−fs(−x),odd,andapplyingthesamesymmetryarguments.Theequationsare
gs(ω)=radicalbigg
2
πintegraldisplay∞
0fs(t)sinωtdt,3(15.26)
fs(x)=radicalbigg
2
πintegraldisplay∞
0gs(ω)sinωxdω. (15.27)
From the last equation we may develop the physical interpretation that f(x)is being
describedbyacontinuumofsinewaves.Theamplitudeof sin ωxisgivenby√2/πgs(ω),
inwhich gs(ω)istheFouriersinetransformof f(x).ItwillbeseenthatEq.(15.27)isthe
integral analog of the summation (Eq. (14.24)). Similar interpretations hold for the cosine
andexponentialcases.
If we take Eqs. (15.22), (15.24), and (15.26) as the direct integral transforms, de-
scribed by Lin Eq. (15.10) (Section 15.1), the corresponding inverse transforms, L−1
ofEq. (15.11),are givenbyEqs.(15.23), (15.25), and(15.27).
Note that the Fourier cosine transforms and the Fourier sine transforms each involve
onlypositivevalues(andzero)ofthearguments.Weusetheparityof f(x)toestablishthe
transforms; but once the transforms are established, the behavior of the functions fandg
for negative argument is irrelevant. In effect, the transform equations themselves impose
adefinite parity :even for the Fourier cosine transform and odd for the Fourier sine
transform.
3Notethat afactor −ihas beenabsorbed into this g(ω).
940 Chapter 15 Integral Transforms
FIGURE 15.2Finitewavetrain.
Example 15.3.1 FINITE WAVETRAIN
An important application of the Fourier transform is the resolution of a finite pulse into
sinusoidal waves. Imagine that an infinite wave train sin ω0tis clipped by Kerr cell or
saturabledyecellshutters sothatwehave
f(t)=
sinω0t,|t|<Nπ
ω0,
0,|t|>Nπ
ω0.(15.28)
This corresponds to Ncycles of our original wave train (Fig. 15.2). Since f(t)is odd, we
mayusetheFouriersinetransform(Eq. (15.26))toobtain
gs(ω)=radicalbigg
2
πintegraldisplayNπ/ω0
0sinω0tsinωtdt. (15.29)
Integrating,wefindouramplitudefunction:
gs(ω)=radicalbigg
2
πbracketleftbiggsin[(ω0−ω)(Nπ/ω 0)]
2(ω0−ω)−sin[(ω0+ω)(Nπ/ω 0)]
2(ω0+ω)bracketrightbigg
.(15.30)
It is of considerable interest to see how gs(ω)depends on frequency. For large ω0and
ω≈ω0, only the first term will be of any importance because of the denominators. It is
plottedinFig.15.3.Thisis theamplitudecurvefor thesingle-slitdiffractionpattern.
Therearezeros at
ω0−ω
ω0=/Delta1ω
ω0=±1
N,±2
N,andso on . (15.31)
Forlarge N,gs(ω)mayalsobeinterpretedasaDiracdeltadistribution,asinSection1.15.
Sincethecontributionsoutsidethecentralmaximumaresmallinthiscase,wemaytake
/Delta1ω=ω0
N(15.32)
as a good measure of the spread in frequency of our wave pulse. Clearly, if Nis large
(alongpulse),thefrequencyspreadwillbesmall.Ontheotherhand,ifourpulseisclipped
15.3 Fourier Transforms — Inversion Theorem 941
FIGURE 15.3Fouriertransform
of finitewavetrain.
short,Nsmall,thefrequencydistributionwillbewiderandthesecondarymaximaaremore
important. /squaresolid
Uncertainty Principle
Hereisaclassicalanalogofthefamousuncertaintyprincipleofquantummechanics.Ifwe
aredealingwithelectromagneticwaves,
hω
2π=E,energy(ofourphoton )
h/Delta1ω
2π=/Delta1E, (15.33)
hbeing Planck’s constant. Here /Delta1Erepresents an uncertainty in the energy of our pulse.
Thereisalsoanuncertaintyinthetime,forourwaveof Ncyclesrequires2 Nπ/ω0seconds
topass.Taking
/Delta1t=2Nπ
ω0, (15.34)
wehavetheproductof thesetwouncertainties:
/Delta1E·/Delta1t=h/Delta1ω
2π·2πN
ω0=hω0
2πN·2πN
ω0=h. (15.35)
TheHeisenberguncertaintyprincipleactuallystates
/Delta1E·/Delta1t≥h
4π, (15.36)
andthisis clearlysatisfiedinourexample.
942 Chapter 15 Integral Transforms
Exercises
15.3.1 (a) Show that g(−ω)=g∗(ω)is a necessary and sufficient condition for f(x)to be
real.
(b) Showthat g(−ω)=−g∗(ω)isanecessaryandsufficientconditionfor f(x)tobe
pureimaginary.
Note. Theconditionof part(a) is used inthedevelopmentof thedispersionrelationsof
Section7.2.
15.3.2 LetF(ω)betheFourier(exponential)transformof f(x)andG(ω)betheFouriertrans-
form ofg(x)=f(x+a). Showthat
G(ω)=e−iaωF(ω).
15.3.3 Thefunction
f(x)=braceleftbigg1,|x|<1
0,|x|>1
isa symmetricalfinitestepfunction.
(a) Findthe gc(ω), Fouriercosinetransformof f(x).
(b) Takingtheinversecosinetransform, showthat
f(x)=2
πintegraldisplay∞
0sinωcosωx
ωdω.
(c) Frompart (b)showthat
integraldisplay∞
0sinωcosωx
ωdω=
0,|x|>1,
π
4,|x|=1,
π
2,|x|<1.
15.3.4 (a) ShowthattheFouriersineandcosinetransforms of e−atare
gs(ω)=radicalbigg
2
πω
ω2+a2,g c(ω)=radicalbigg
2
πa
ω2+a2.
Hint.Eachof thetransforms canberelatedtotheotherbyintegrationbyparts.
(b) Showthat
integraldisplay∞
0ωsinωx
ω2+a2dω=π
2e−ax,x>0,
integraldisplay∞
0cosωx
ω2+a2dω=π
2ae−ax,x>0.
Theseresultsare alsoobtainedbycontourintegration(Exercise7.1.14).
15.3.5 FindtheFouriertransformofthetriangularpulse(Fig.15.4).
f(x)=braceleftBigghparenleftbig
1−a|x|parenrightbig
,|x|<1
a,
0, |x|>1
a.
Note.This functionprovidesanotherdeltasequencewith h=aanda→∞.
15.3 Fourier Transforms — Inversion Theorem 943
FIGURE 15.4Triangularpulse.
15.3.6 Defineasequence
δn(x)=braceleftBiggn,|x|<1
2n,
0,|x|>1
2n.
(This is Eq. (1.172).) Express δn(x)as a Fourier integral (via the Fourier integral theo-
rem,inversetransform,etc.). Finally,showthatwemaywrite
δ(x)=limn→∞δn(x)=1
2πintegraldisplay∞
−∞e−ikxdk.
15.3.7 Usingthesequence
δn(x)=n√πexpparenleftbig
−n2x2parenrightbig
,
showthat
δ(x)=1
2πintegraldisplay∞
−∞e−ikxdk.
Note. Remember that δ(x)is defined in terms of its behavior as part of an integrand
(Section1.15), especiallyEqs. (1.178) and(1.179).
15.3.8 Derivesineandcosinerepresentationsof δ(t−x)thatarecomparabletotheexponential
representation,Eq. (15.21d).
ANS.2
πintegraldisplay∞
0sinωtsinωxdω,2
πintegraldisplay∞
0cosωtcosωxdω.
15.3.9 Inaresonantcavityanelectromagneticoscillationoffrequency ω0diesoutas
A(t)=A0e−ω0t/2Qe−iω0t,t>0.
(TakeA(t)=0fort<0.)Theparameter Qisameasureoftheratioofstoredenergyto
energylosspercycle.Calculatethefrequencydistributionoftheoscillation, a∗(ω)a(ω),
wherea(ω)is theFouriertransformof A(t).
Note.Thelarger Qis, thesharperyourresonancelinewillbe.
ANS.a∗(ω)a(ω)=A2
0
2π1
(ω−ω0)2+(ω0/2Q)2.
944 Chapter 15 Integral Transforms
15.3.10 Provethat
¯h
2πiintegraldisplay∞
−∞e−iωtdω
E0−iŴ/2−¯hω=braceleftBigg
expparenleftbig
−Ŵt
2¯hparenrightbig
expparenleftbig
−iE0t
¯hparenrightbig
,t>0,
0,t <0.
This Fourier integral appears in a variety of problems in quantum mechanics: WKB
barrierpenetration,scattering,time-dependentperturbationtheory,andso on.
Hint.Trycontourintegration.
15.3.11 VerifythatthefollowingareFourierintegraltransforms ofoneanother:
(a)radicalbigg
2
π·1√
a2−x2,|x|<a, andJ0(ay),
0, |x|>a,
(b)0, |x|<a,
−radicalbigg
2
π1√
x2+a2,|x|>a, andN0(a|y|),
(c)radicalbiggπ
2·1√
x2+a2andK0parenleftbig
a|y|parenrightbig
.
(d) Canyousuggestwhy I0(ay)is notincludedinthislist?
Hint.J0,N0, andK0may be transformed most easily by using an exponential repre-
sentation, reversing the order of integration, and employing the Dirac delta function
exponential representation (Section 15.2). These cases can be treated equally well as
Fouriercosinetransforms.
Note.T h eK0relation appears as a consequence of a Green’s function equation in Ex-
ercise9.7.14.
15.3.12 A calculation of the magnetic field of a circular current loop in circular cylindrical
coordinatesleadstotheintegral
integraldisplay∞
0coskzkK1(ka)dk.
Showthatthisintegralisequalto
πa
2(z2+a2)3/2.
Hint.TrydifferentiatingExercise15.3.11(c).
15.3.13 AsanextensionofExercise15.3.11,showthat
(a)integraldisplay∞
0J0(y)dy=1, (b)integraldisplay∞
0N0(y)dy=0, (c)integraldisplay∞
0K0(y)dy=π
2.
15.3.14 The Fourier integral, Eq. (15.18), has been held meaningless for f(t)=cosαt. Show
thattheFourierintegralcanbeextendedtocover f(t)=cosαtbyuseoftheDiracdelta
function.
15.3 Fourier Transforms — Inversion Theorem 945
15.3.15 Showthat
integraldisplay∞
0sinkaJ0(kρ)dk=braceleftbiggparenleftbig
a2−ρ2parenrightbig−1/2,ρ<a,
0,ρ >a.
Hereaandρarepositive.Theequationcomesfromthedeterminationofthedistribution
of charge on an isolated conducting disk, radius a. Note that the function on the right
hasaninfinitediscontinuityat ρ=a.
Note.ALaplacetransformapproachappearsinExercise15.10.8.
15.3.16 Thefunction f(r)hasaFourierexponentialtransform,
g(k)=1
(2π)3/2integraldisplay
f(r)eik·rd3r=1
(2π)3/2k2.
Determine f(r).
Hint.Usesphericalpolarcoordinatesin k-space.
ANS.f(r)=1
4πr.
15.3.17 (a) CalculatetheFourierexponentialtransform of f(x)=e−a|x|.
(b) Calculate the inverse transform by employing the calculus of residues (Sec-
tion7.1).
15.3.18 ShowthatthefollowingareFouriertransforms ofeachother
inJn(t)and
radicalbigg
2
πTn(x)parenleftbig
1−x2parenrightbig−1/2,|x|<1,
0, |x|>1.
Tn(x)isthenth-orderChebyshevpolynomial.
Hint. WithTn(cosθ)=cosnθ, the transform of Tn(x)(1−x2)−1/2leads to an integral
representationof Jn(t).
15.3.19 ShowthattheFourierexponentialtransformof
f(µ)=braceleftbiggPn(µ),|µ|≤1,
0,|µ|>1
is(2in/2π)jn(kr).H e r ePn(µ)is a Legendre polynomial and jn(kr)is a spherical
Besselfunction.
15.3.20 Show that the three-dimensional Fourier exponential transform of a radially symmetric
functionmayberewrittenasaFouriersinetransform:
1
(2π)3/2integraldisplay∞
−∞f(r)eik·rd3x=1
kradicalbigg
2
πintegraldisplay∞
0bracketleftbig
rf(r)bracketrightbig
sinkrdr.
946 Chapter 15 Integral Transforms
15.3.21 (a) Show that f(x)=x−1/2is aself-reciprocal under both Fourier cosine and sine
transforms;thatis,
radicalbigg
2
πintegraldisplay∞
0x−1/2cosxtdx=t−1/2,
radicalbigg
2
πintegraldisplay∞
0x−1/2sinxtds=t−1/2.
(b) Use the preceding results to evaluate the Fresnel integralsintegraltext∞
0cos(y2)dyandintegraltext∞
0sin(y2)dy.
15.4 F OURIER TRANSFORM OF DERIVATIVES
In Section 15.1, Fig. 15.1 outlines the overall technique of using Fourier transforms and
inversetransforms tosolveaproblem.Herewetakeaninitialstepinsolvingadifferential
equation—obtainingtheFouriertransformofaderivative.
Usingtheexponentialform, wedeterminethattheFouriertransformof f(x)is
g(ω)=1√
2πintegraldisplay∞
−∞f(x)eiωxdx (15.37)
andfordf(x)/dx
g1(ω)=1√
2πintegraldisplay∞
−∞df(x)
dxeiωxdx. (15.38)
IntegratingEq. (15.38)byparts, weobtain
g1(ω)=eiωx
√
2πf(x)vextendsinglevextendsinglevextendsingle∞
−∞−iω√
2πintegraldisplay∞
−∞f(x)eiωxdx. (15.39)
Iff(x)vanishes4asx→±∞,weha v e
g1(ω)=−iωg(ω); (15.40)
thatis,thetransformofthederivativeis (−iω)timesthetransformoftheoriginalfunction.
Thismayreadilybegeneralizedtothe nthderivativetoyield
gn(ω)=(−iω)ng(ω), (15.41)
provided all the integrated parts vanish as x→±∞. This is the power of the Fourier
transform,thereasonitissousefulinsolving(partial)differentialequations.Theoperation
ofdifferentiationhasbeenreplacedbya multiplicationin ω-space.
4Apart from casessuch asExercise 15.3.6, f(x)must vanish as x→±∞in order for the Fourier transform of f(x)to exist.
15.4 Fourier Transform of Derivatives 947
Example 15.4.1 WAVEEQUATION
ThistechniquemaybeusedtoadvantageinhandlingPDEs.Toillustratethetechnique,let
usderiveafamiliarexpressionofelementaryphysics.Aninfinitelylongstringisvibrating
freely.Theamplitude yofthe(small)vibrationssatisfiesthewaveequation
∂2y
∂x2=1
v2∂2y
∂t2. (15.42)
Weshallassumeaninitialcondition
y(x,0)=f(x), (15.43)
wherefis localized,thatis, approacheszeroatlarge x.
Applying our Fourier transform in x, which means multiplying by eiαxand integrating
overx,weobtain
integraldisplay∞
−∞∂2y(x,t)
∂x2eiαxdx=1
v2integraldisplay∞
−∞∂2y(x,t)
∂t2eiαxdx (15.44)
or
(−iα)2Y(α,t)=1
v2∂2Y(α,t)
∂t2. (15.45)
Herewehaveused
Y(α,t)=1√
2πintegraldisplay∞
−∞y(x,t)eiαxdx (15.46)
and Eq. (15.41) for the second derivative. Note that the integrated part of Eq. (15.39) van-
ishes: The wave has not yet gone to ±∞because it is propagating forward in time, and
there is no source at infinity because f(±∞)=0.Since no derivatives with respect to
αappear, Eq. (15.45) is actually an ODE—in fact, the linear oscillator equation. This
transformation,fromaPDEtoanODE,isasignificantachievement.WesolveEq.(15.45)
subject to the appropriate initial conditions. At t=0, applying Eq. (15.43), Eq. (15.46)
reducesto
Y(α,0)=1√
2πintegraldisplay∞
−∞f(x)eiαxdx=F(α). (15.47)
Thegeneralsolutionof Eq.(15.45) inexponentialformis
Y(α,t)=F(α)e±ivαt. (15.48)
Usingtheinversionformula(Eq. (15.23)), wehave
y(x,t)=1√
2πintegraldisplay∞
−∞Y(α,t)e−iαxdα, (15.49)
and,byEq.(15.48),
y(x,t)=1√
2πintegraldisplay∞
−∞F(α)e−iα(x∓vt)dα. (15.50)
948 Chapter 15 Integral Transforms
Sincef(x)istheFourierinversetransform of F(α),
y(x,t)=f(x∓vt), (15.51)
correspondingtowavesadvancinginthe +x-and−x-directions,respectively.
The particular linear combinations of waves is given by the boundary condition of
Eq.(15.43) andsomeotherboundarycondition,suchas arestrictionon ∂y/∂t. /squaresolid
TheaccomplishmentoftheFouriertransform heredeservesspecialemphasis.
•Our Fourier transform converted a PDE into an ODE, where the “degree of transcen-
dence”oftheproblemwasreduced.
In Section 15.9 Laplace transforms are used to convert ODEs (with constant coefficients)
into algebraic equations. Again, the degree of transcendence is reduced. The problem is
simplified—as outlinedinFig.15.1.
Example 15.4.2 HEATFLOW PDE
To illustrate another transformation of a PDE into an ODE, let us Fourier transform the
heatflowpartialdifferentialequation
∂ψ
∂t=a2∂2ψ
∂x2,
where the solution ψ(x,t)is the temperature in space as a function of time. By taking the
Fourier transform of both sides of this equation (note that here only ωis the transform
variableconjugateto xbecausetis thetimeintheheatflowPDE),where
/Psi1(ω,t)=1√
2πintegraldisplay∞
−∞ψ(x,t)eiωxdx,
thisyieldsanODEfor theFouriertransform /Psi1ofψinthetimevariable t,
∂/Psi1(ω,t)
∂t=−a2ω2/Psi1(ω,t).
Integratingweobtain
ln/Psi1=−a2ω2t+lnC,or/Psi1=Ce−a2ω2t,
where the integration constant Cmay still depend on ωand, in general, is determined
by initial conditions. In fact, C=/Psi1(ω,0)is the initial spatial distribution of /Psi1,so it is
given by the transform (in x) of the initial distribution of ψ,namely,ψ(x,0).Putting this
solutionbackintoourinverseFouriertransform, thisyields
ψ(x,t)=1√
2πintegraldisplay∞
−∞C(ω)e−iωxe−a2ω2tdω.
For simplicity, we here take Cω-independent (assuming a delta-function initial temper-
ature distribution) and integrate by completing the square in ω,as in Example 15.1.1,
15.4 Fourier Transform of Derivatives 949
making appropriate changes of variables and parameters ( a2→a2t, ω→x,t→−ω).
ThisyieldstheparticularsolutionoftheheatflowPDE,
ψ(x,t)=C
a√
2texpparenleftbigg
−x2
4a2tparenrightbigg
,
whichappearsasacleverguessinChapter8.Ineffect,wehaveshownthat ψistheinverse
Fouriertransform of Cexp(−a2ω2t). /squaresolid
Example 15.4.3 INVERSION OF PDE
DeriveaFourierintegralfortheGreen’sfunction G0ofPoisson’sPDE,whichisasolution
of
∇2G0(r,r′)=−δ(r−r′).
OnceG0is known,thegeneralsolutionofPoisson’sPDE,
∇2/Phi1=−4πρ(r)
ofelectrostatics,isgivenas
/Phi1(r)=integraldisplay
G0(r,r′)4πρ(r′)d3r′.
Applying ∇2to/Phi1andusingthePDEtheGreen’sfunctionsatisfies,wecheckthat
∇2/Phi1(r)=integraldisplay
∇2G0(r,r′)4πρ(r′)d3r′=−integraldisplay
δ(r−r′)4πρ(r′)d3r′=−4πρ(r).
NowweusetheFouriertransformof G0,whichisg0,andofthatofthe δfunction,writing
∇2integraldisplay
g0(p)eip·(r−r′)d3p
(2π)3=−integraldisplay
eip·(r−r′)d3p
(2π)3.
Because the integrands of equal Fourier integrals must be the same (almost) everywhere,
whichfollowsfromtheinverseFouriertransform,andwith
∇eip·(r−r′)=ipeip·(r−r′),
this yields−p2g0(p)=−1.Therefore, application of the Laplacian to a Fourier integral
f(r)correspondstomultiplyingitsFouriertransform g(p)by−p2.Substitutingthissolu-
tionintotheinverseFouriertransform for G0gives
G0(r,r′)=integraldisplay
eip·(r−r′)d3p
(2π)3p2=1
4π|r−r′|.
We can verify the last part of this result by applying ∇2toG0again and recalling from
Chapter1that ∇21
|r−r′|=−4πδ(r−r′).
950 Chapter 15 Integral Transforms
The inverse Fourier transform can be evaluated using polar coordinates, exploiting the
sphericalsymmetryof p2.Forsimplicity,wewrite R=r−r′andcallθtheanglebetween
Randp,
integraldisplay
eip·Rd3p
p2=integraldisplay∞
0dpintegraldisplay1
−1eipRcosθdcosθintegraldisplay2π
0dϕ
=2π
iRintegraldisplay∞
0dp
peipRcosθvextendsinglevextendsinglevextendsingle1
cosθ=−1=4π
Rintegraldisplay∞
0sinpR
pdp
=4π
Rintegraldisplay∞
0sinpR
pRd(pR)=2π2
R,
whereθandϕare the angles of pandintegraltext∞
0sinx
xdx=π
2, from Example 7.1.4. Dividing by
(2π)3,we obtain G0(R)=1/(4πR),as claimed. An evaluation of this Fourier transform
bycontourintegrationis giveninExample9.7.2. /squaresolid
Exercises
15.4.1 Theone-dimensionalFermiageequationfor thediffusionofneutronsslowingdownin
somemedium(suchas graphite)is
∂2q(x,τ)
∂x2=∂q(x,τ)
∂τ.
Hereqis the number of neutrons that slow down, falling below some given energy per
secondperunitvolume.TheFermiage, τ, isameasureoftheenergyloss.
Ifq(x,0)=Sδ(x), corresponding to a plane source of neutrons at x=0, emitting S
neutronsper unitareapersecond,derivethesolution
q=Se−x2/4τ
√
4πτ.
Hint.Replace q(x,τ)with
p(k,τ)=1√
2πintegraldisplay∞
−∞q(x,τ)eikxdx.
Thisis analogoustothediffusionof heatinaninfinitemedium.
15.4.2 Equation(15.41)yields
g2(ω)=−ω2g(ω)
fortheFouriertransformofthesecondderivativeof f(x).Thecondition f(x)→0f o r
x→±∞may be relaxed slightly. Find the least restrictive condition for the preceding
equationfor g2(ω)tohold.
ANS.bracketleftbiggdf(x)
dx−iωf(x)bracketrightbigg
eiωxvextendsinglevextendsinglevextendsingle∞
−∞=0.
15.5 Convolution Theorem 951
15.4.3 Theone-dimensionalneutrondiffusionequationwitha(plane)sourceis
−Dd2ϕ(x)
dx2+K2Dϕ(x)=Qδ(x),
whereϕ(x)istheneutronflux, Qδ(x)isthe(plane)sourceat x=0,andDandK2are
constants.ApplyaFouriertransform.Solvetheequationintransformspace.Transform
yoursolutionbackinto x-space.
ANS.ϕ(x)=Q
2KDe−|Kx|.
15.4.4 For a point source at the origin, the three-dimensional neutron diffusion equation be-
comes
−D∇2ϕ(r)+K2Dϕ(r)=Qδ(r).
Apply a three-dimensional Fourier transform. Solve the transformed equation. Trans-
formthesolutionbackinto r-space.
15.4.5 (a) Given that F(k)is the three-dimensional Fourier transform of f(r)andF1(k)is
thethree-dimensionalFouriertransformof ∇f(r), showthat
F1(k)=(−ik)F(k).
Thisis athree-dimensionalgeneralizationofEq. (15.40).
(b) Showthatthethree-dimensionalFouriertransformof ∇·∇f(r)is
F2(k)=(−ik)2F(k).
Note. Vectorkis a vector in the transform space. In Section 15.6 we shall have
¯hk=p,linearmomentum.
15.5 C ONVOLUTION THEOREM
We shall employ convolutions to solve differential equations, to normalize momentum
wavefunctions(Section15.6),andtoinvestigatetransfer functions(Section15.7).
Let us consider two functions f(x)andg(x)with Fourier transforms F(t)andG(t),
respectively.Wedefinetheoperation
f∗g≡1√
2πintegraldisplay∞
−∞g(y)f(x−y)dy (15.52)
as theconvolution of the two functions fandgover the interval (−∞,∞). This form of
an integral appears in probability theory in the determination of the probability density of
two random, independent variables. Our solution of Poisson’s equation, Eq. (9.148), may
be interpreted as a convolution of a charge distribution, ρ(r2), and a weighting function,
(4πε0|r1−r2|)−1. In other works this is sometimes referred to as the Faltung,t ou s et h e
952 Chapter 15 Integral Transforms
FIGURE 15.5
German term for “folding.”5We now transform the integral in Eq. (15.52) by introducing
theFouriertransforms:
integraldisplay∞
−∞g(y)f(x−y)dy=1√
2πintegraldisplay∞
−∞g(y)integraldisplay∞
−∞F(t)e−it(x−y)dtdy
=1√
2πintegraldisplay∞
−∞F(t)bracketleftbiggintegraldisplay∞
−∞g(y)eitydybracketrightbigg
e−itxdt
=integraldisplay∞
−∞F(t)G(t)e−itxdt, (15.53)
interchanging the order of integration and transforming g(y). This result may be inter-
preted as follows: The Fourier inverse transform of a productof Fourier transforms is the
convolutionof theoriginalfunctions, f∗g.
Forthespecialcase x=0w eh a v e
integraldisplay∞
−∞F(t)G(t)dt=integraldisplay∞
−∞f(−y)g(y)dy. (15.54)
Theminussignin −ysuggeststhatmodificationsbetried.Wenowdothiswith g∗instead
ofgusingadifferent technique.
Parseval’s Relation
Results analogous to Eqs. (15.53) and (15.54) may be derived for the Fourier sine and co-
sinetransforms(Exercises15.5.1and15.5.3).Equation(15.54)andthecorrespondingsine
and cosine convolutions are often labeled Parseval’s relations by analogy with Parseval’s
theoremfor Fourierseries (Chapter14,Exercise14.4.2).
5Forf(y)=e−y,f(y)andf(x−y)are plotted in Fig. 15.5. Clearly, f(y)andf(x−y)are mirror images of each other in
relationto thevertical line y=x/2,that is,wecould generate f(x−y)by folding over f(y)on the line y=x/2.
15.5 Convolution Theorem 953
TheParsevalrelation6,7
integraldisplay∞
−∞F(ω)G∗(ω)dω=integraldisplay∞
−∞f(t)g∗(t)dt (15.55)
may be derived elegantly using the Dirac delta function representation, Eq. (15.21d). We
have
integraldisplay∞
−∞f(t)g∗(t)dt=integraldisplay∞
−∞1√
2πintegraldisplay∞
−∞F(ω)e−iωtdω·1√
2πintegraldisplay∞
−∞G∗(x)eixtdxdt,
(15.56)
withattentiontothecomplexconjugationinthe G∗(x)tog∗(t)transform.Integratingover
tfirst, andusingEq.(15.21d), weobtain
integraldisplay∞
−∞f(t)g∗(t)dt=integraldisplay∞
−∞F(ω)integraldisplay∞
−∞G∗(x)δ(x−ω)dxdω
=integraldisplay∞
−∞F(ω)G∗(ω)dω, (15.57)
our desired Parseval relation. If f(t)=g(t), then the integrals in the Parseval relation are
normalization integrals (Section 10.4). Equation (15.57) guarantees that if a function f(t)
isnormalizedtounity,itstransform F(ω)islikewisenormalizedtounity.Thisisextremely
importantinquantummechanicsasdevelopedinthenextsection.
ItmaybeshownthattheFouriertransformisaunitaryoperation(intheHilbertspace L2,
squareintegrablefunctions).TheParsevalrelationisareflectionofthisunitaryproperty—
analogoustoExercise3.4.26for matrices.
In Fraunhofer diffraction optics the diffraction pattern (amplitude) appears as the trans-
form of the function describing the aperture (compare Exercise 15.5.5). With intensity
proportional to the square of the amplitude the Parseval relation implies that the energy
passing through the aperture seems to be somewhere in the diffraction pattern—a state-
ment of the conservation of energy. Parseval’s relations may be developed independently
of the inverse Fourier transform and then used rigorously to derive the inverse transform.
DetailsaregivenbyMorse andFeshbach,8Section4.8(see alsoExercise15.5.4).
Exercises
15.5.1 WorkouttheconvolutionequationcorrespondingtoEq.(15.53) for
(a) Fouriersinetransforms
1
2integraldisplay∞
0g(y)bracketleftbig
f(y+x)+f(y−x)bracketrightbig
dy=integraldisplay∞
0Fs(s)Gs(s)cossxds,
wherefandgare oddfunctions.
6Notethat all arguments arepositive, in contrast to Eq.(15.54).
7Some authors prefer to restrict Parseval’s name to series and refer to Eq. (15.55) as Rayleigh’s theorem .
8P.M.Morse andH.Feshbach, Methods of Theoretical Physics ,NewYork: McGraw-Hill(1953).
954 Chapter 15 Integral Transforms
(b) Fouriercosinetransforms
1
2integraldisplay∞
0g(y)bracketleftbig
f(y+x)+f(x−y)bracketrightbig
dy=integraldisplay∞
0Fc(s)Gc(s)cossxds,
wherefandgare evenfunctions.
15.5.2 F(ρ)andG(ρ)are the Hankel transforms of f(r)andg(r), respectively (Exer-
cise15.1.1). DerivetheHankeltransformParsevalrelation:
integraldisplay∞
0F∗(ρ)G(ρ)ρdρ=integraldisplay∞
0f∗(r)g(r)rdr.
15.5.3 Show that for both Fourier sine and Fourier cosine transforms Parseval’s relation has
theform
integraldisplay∞
0F(t)G(t)dt=integraldisplay∞
0f(y)g(y)dy.
15.5.4 Starting from Parseval’s relation (Eq. (15.54)), let g(y)=1, 0≤y≤α, and zero else-
where.FromthisderivetheFourierinversetransform (Eq. (15.23)).
Hint.Differentiatewithrespectto α.
15.5.5 (a) Arectangularpulseis describedby
f(x)=braceleftbigg1,|x|<a,
0,|x|>a.
ShowthattheFourierexponentialtransform is
F(t)=radicalbigg
2
πsinat
t.
This is the single-slit diffraction problem of physical optics. The slit is described
byf(x).Thediffractionpattern amplitude isgivenbytheFouriertransform F(t).
(b) UsetheParsevalrelationtoevaluate
integraldisplay∞
−∞sin2t
t2dt.
This integral may also be evaluated by using the calculus of residues, Exer-
cise7.1.12.
ANS.(b)π.
15.5.6 Solve Poisson’s equation, ∇2ψ(r)=−ρ(r)/ε0, by the following sequence of opera-
tions:
(a) Take the Fourier transform of both sides of this equation. Solve for the Fourier
transform of ψ(r).
(b) CarryouttheFourierinversetransformbyusingathree-dimensionalanalogofthe
convolutiontheorem,Eq. (15.53).
15.6 Momentum Representation 955
15.5.7 (a) Given f(x)=1−|x/2|,−2≤x≤2, and zero elsewhere, show that the Fourier
transform of f(x)is
F(t)=radicalbigg
2
πparenleftbiggsint
tparenrightbigg2
.
(b) UsingtheParsevalrelation,evaluate
integraldisplay∞
−∞parenleftbiggsint
tparenrightbigg4
dt.
ANS.(b)2π
3.
15.5.8 WithF(t)andG(t)theFouriertransforms of f(x)andg(x), respectively,showthat
integraldisplay∞
−∞vextendsinglevextendsinglef(x)−g(x)vextendsinglevextendsingle2dx=integraldisplay∞
−∞vextendsinglevextendsingleF(t)−G(t)vextendsinglevextendsingle2dt.
Ifg(x)is an approximation to f(x), the preceding relation indicates that the mean
squaredeviationin t-spaceis equaltothemeansquaredeviationin x-space.
15.5.9 UsetheParsevalrelationtoevaluate
(a)integraldisplay∞
−∞dω
(ω2+a2)2,(b)integraldisplay∞
−∞ω2dω
(ω2+a2)2.
Hint.CompareExercise15.3.4.
ANS.(a)π
2a3,( b)π
2a.
15.6 M OMENTUM REPRESENTATION
In advanced dynamics and in quantum mechanics, linear momentum and spatial position
occur on an equal footing. In this section we shall start with the usual space distribution
and derive the corresponding momentum distribution. For the one-dimensional case our
wavefunction ψ(x)hasthefollowingproperties:
1.ψ∗(x)ψ(x)dx is the probability density of finding a quantum particle between xand
x+dx,and
2.integraldisplay∞
−∞ψ∗(x)ψ(x)dx=1 (15.58)
correspondstoprobabilityunity.
3. In addition,wehave
/angbracketleftx/angbracketright=integraldisplay∞
−∞ψ∗(x)xψ(x)dx (15.59)
fortheaveragepositionoftheparticlealongthe x-axis.Thisisoftencalledan expec-
tationvalue .
Wewantafunction g(p)thatwillgivethesameinformationaboutthemomentum:
956 Chapter 15 Integral Transforms
1.g∗(p)g(p)dp is the probability density that our quantum particle has a momentum
betweenpandp+dp.
2.integraldisplay∞
−∞g∗(p)g(p)dp=1. (15.60)
3. /angbracketleftp/angbracketright=integraldisplay∞
−∞g∗(p)pg(p)dp. (15.61)
As subsequently shown, such a function is given by the Fourier transform of our space
functionψ(x). Specifically,9
g(p)=1√2π¯hintegraldisplay∞
−∞ψ(x)e−ipx/¯hdx, (15.62)
g∗(p)=1√2π¯hintegraldisplay∞
−∞ψ∗(x)eipx/¯hdx. (15.63)
Thecorrespondingthree-dimensionalmomentumfunctionis
g(p)=1
(2π¯h)3/2integraldisplay∞integraldisplay
−∞integraldisplay
ψ(r)e−ir·p/¯hd3r.
ToverifyEqs.(15.62) and(15.63), letuscheckonproperties2and3.
Property 2, the normalization, is automatically satisfied as a Parseval relation,
Eq. (15.55). If the space function ψ(x)is normalized to unity, the momentum function
g(p)is alsonormalizedtounity.
Tocheckonproperty3,wemustshowthat
/angbracketleftp/angbracketright=integraldisplay∞
−∞g∗(p)pg(p)dp=integraldisplay∞
−∞ψ∗(x)¯h
id
dxψ(x)dx, (15.64)
where(¯h/i)(d/dx) is the momentum operator in the space representation. We replace
the momentum functions by Fourier-transformed space functions, and the first integral
becomes
1
2π¯hintegraldisplay∞integraldisplay
−∞integraldisplay
pe−ip(x−x′)/¯hψ∗(x′)ψ(x)dpdx′dx. (15.65)
Nowweusetheplane-waveidentity
pe−ip(x−x′)/¯h=d
dxbracketleftbigg
−¯h
ie−ip(x−x′)/¯hbracketrightbigg
, (15.66)
9The¯hmay beavoided by using the wavenumber k,p=k¯h(andp=k¯h), so
ϕ(k)=1
(2π)1/2integraldisplay
ψ(x)e−ikxdx.
An example of this notation appears in Section16.1.
15.6 Momentum Representation 957
withpa constant, not an operator. Substituting into Eq. (15.65) and integrating by parts,
holdingx′andpconstant,weobtain
/angbracketleftp/angbracketright=integraldisplay∞integraldisplay
−∞bracketleftbigg1
2π¯hintegraldisplay∞
−∞e−ip(x−x′)/¯hdpbracketrightbigg
·ψ∗(x′)¯h
id
dxψ(x)dx′dx. (15.67)
Here we assume ψ(x)vanishes as x→±∞, eliminating the integrated part. Using the
Dirac delta function, Eq. (15.21c), Eq. (15.67) reduces to Eq. (15.64) to verify our mo-
mentumrepresentation.
Alternatively,if theintegrationover pis donefirst inEq.(15.65), leadingto
integraldisplay∞
−∞pe−ip(x−x′)/¯hdp=2πi¯h2δ′(x−x′),
andusingExercise1.15.9,wecandotheintegrationover x,whichcauses ψ(x)tobecome
−dψ(x′)/dx′.Theremainingintegralover x′istheright-handsideofEq. (15.64).
Example 15.6.1 HYDROGEN ATOM
Thehydrogenatomgroundstate10maybedescribedbythespatialwavefunction
ψ(r)=parenleftbigg1
πa3
0parenrightbigg1/2
e−r/a0, (15.68)
a0being the Bohr radius, 4 πε0¯h2/me2. We now have a three-dimensional wave function.
Thetransform correspondingtoEq.(15.62) is
g(p)=1
(2π¯h)3/2integraldisplay
ψ(r)e−ip·r/¯hd3r. (15.69)
SubstitutingEq. (15.68)intoEq.(15.69) andusing
integraldisplay
e−ar+ib·rd3r=8πa
(a2+b2)2, (15.70)
weobtainthehydrogenicmomentumwavefunction,
g(p)=23/2
πa3/2
0¯h5/2
(a2
0p2+¯h2)2. (15.71)
Such momentum functions have been found useful in problems like Compton scattering
fromatomicelectrons,thewavelengthdistributionofthescatteredradiation,dependingon
themomentumdistributionof thetargetelectrons.
The relation between the ordinary space representation and the momentum representa-
tion may be clarified by considering the basic commutation relations of quantum mechan-
ics.WegofromaclassicalHamiltoniantotheSchrödingerwaveequationbyrequiringthat
momentum pandposition xnotcommute.Instead,werequirethat
[p,x]≡px−xp=−i¯h. (15.72)
10See E. V. Ivash, A momentum representation treatment of the hydrogen atom problem. A m .J .P h y s . 40: 1095 (1972) for
amomentum representation treatment of the hydrogen atom l=0 states.
958 Chapter 15 Integral Transforms
Forthemultidimensionalcase,Eq.(15.72) isreplacedby
[pi,xj]=−i¯hδij. (15.73)
TheSchrödinger(space)representationis obtainedbyusing
x→x:pi→−i¯h∂
∂xi,
replacingthemomentumbyapartialspacederivative.Wesee that
[p,x]ψ(x)=−i¯hψ(x). (15.74)
However,Eq. (15.72)canequallywellbesatisfiedbyusing
p→p:xj→i¯h∂
∂pj.
Thisis themomentumrepresentation.Then
[p,x]g(p)=−i¯hg(p). (15.75)
Hencetherepresentation (x)is notunique; (p)is analternatepossibility.
Ingeneral,theSchrödingerrepresentation (x)leadingtotheSchrödingerwaveequation
is more convenient because the potential energy Vis generally given as a function of
positionV(x,y,z) .Themomentumrepresentation (p)usuallyleadstoanintegralequation
(compare Chapter 16 for the pros and cons of the integral equations). For an exception,
considertheharmonicoscillator. /squaresolid
Example 15.6.2 HARMONIC OSCILLATOR
TheclassicalHamiltonian(kineticenergy +potentialenergy =totalenergy)is
H(p,x)=p2
2m+1
2kx2=E, (15.76)
wherekistheHooke’slawconstant.
In theSchrödingerrepresentationweobtain
−¯h2
2md2ψ(x)
dx2+1
2kx2ψ(x)=Eψ(x). (15.77)
Fortotalenergy Eequalto√(k/m)¯h/2 thereisanunnormalizedsolution(Section13.1),
ψ(x)=e−(√
mk/2¯h)x2. (15.78)
Themomentumrepresentationleadsto
p2
2mg(p)−¯h2k
2d2g(p)
dp2=Eg(p). (15.79)
Again,for
E=radicalbigg
k
m¯h
2(15.80)
15.6 Momentum Representation 959
themomentumwaveequation(15.79) issatisfiedbytheunnormalized
g(p)=e−p2/(2¯h√
mk). (15.81)
Either representation, space or momentum (and an infinite number of other possibilities),
may be used, depending on which is more convenient for the particular problem under
attack.
The demonstration that g(p)is the momentum wave function corresponding to
Eq. (15.78)—that it is the Fourier inverse transform of Eq. (15.78)—is left as Exer-
cise15.6.3. /squaresolid
Exercises
15.6.1 The function eik·rdescribes a plane wave of momentum p=¯hknormalized to unit
density. (Time dependenceof e−iωtis assumed.) Show that these plane-wavefunctions
satisfyanorthogonalityrelation
integraldisplayparenleftbig
eik·rparenrightbig∗eik′·rdxdydz=(2π)3δ(k−k′).
15.6.2 Aninfiniteplanewaveinquantummechanicsmayberepresentedbythefunction
ψ(x)=eip′x/¯h.
Findthecorrespondingmomentumdistributionfunction.Notethatithasaninfinityand
thatψ(x)is notnormalized.
15.6.3 Alinearquantumoscillatorinits groundstatehasawavefunction
ψ(x)=a−1/2π−1/4e−x2/2a2.
Showthatthecorrespondingmomentumfunctionis
g(p)=a1/2π−1/4¯h−1/2e−a2p2/2¯h2.
15.6.4 Thenthexcitedstateofthelinearquantumoscillatorisdescribedby
ψn(x)=a−1/22−n/2π−1/4(n!)−1/2e−x2/2a2Hn(x/a),
whereHn(x/a)is thenth Hermite polynomial, Section 13.1. As an extension of Exer-
cise15.6.3,findthemomentumfunctioncorrespondingto ψn(x).
Hint.ψn(x)may be represented by (ˆa†)nψ0(x), whereˆa†is the raising operator, Exer-
cise13.1.14to13.1.16.
15.6.5 Afreeparticleinquantummechanicsisdescribedbyaplanewave
ψk(x,t)=ei[kx−(¯hk2/2m)t].
Combiningwaves of adjacent momentumwith an amplitudeweightingfactor ϕ(k),w e
formawavepacket
/Psi1(x,t)=integraldisplay∞
−∞ϕ(k)ei[kx−(¯hk2/2m)t]dk.
960 Chapter 15 Integral Transforms
(a) Solvefor ϕ(k)giventhat
/Psi1(x,0)=e−x2/2a2.
(b) Using the known value of ϕ(k), integrate to get the explicit form of /Psi1(x,t).N o t e
thatthiswavepacketdiffuses orspreads outwithtime.
ANS./Psi1(x,t)=e−{x2/2[a2+(i¯h/m)t]}
[1+(i¯ht/ma2)]1/2.
Note. An interesting discussion of this problem from the evolution operator point of
view is given by S. M. Blinder, Evolution of a Gaussian wave packet, Am. J. Phys. 36:
525(1968).
15.6.6 Find the time-dependent momentum wave function g(k,t)corresponding to /Psi1(x,t)of
Exercise 15.6.5. Show that the momentum wave packet g∗(k,t)g(k,t) isindependent
oftime.
15.6.7 The deuteron, Example 10.1.2, may be described reasonably well with a Hulthén wave
function
ψ(r)=A[e−αr−e−βr]/r,
withA,α,andβconstants.Find g(p), thecorrespondingmomentumfunction.
Note. The Fourier transform may be rewritten as Fourier sine and cosine transforms or
asaLaplacetransform,Section15.8.
15.6.8 The nuclear form factor F(k)and the charge distribution ρ(r)are three-dimensional
Fouriertransforms ofeachother:
F(k)=1
(2π)3/2integraldisplay
ρ(r)eik·rd3r.
If themeasuredform factor is
F(k)=(2π)−3/2parenleftbigg
1+k2
a2parenrightbigg−1
,
findthecorrespondingchargedistribution.
ANS.ρ(r)=a2
4πe−ar
r.
15.6.9 Checkthenormalizationofthehydrogenmomentumwavefunction
g(p)=23/2
πa3/2
0¯h5/2
(a2
0p2+¯h2)2
bydirectevaluationoftheintegral
integraldisplay
g∗(p)g(p)d3p.
15.6.10 Withψ(r)a wave function in ordinary space and ϕ(p)the corresponding momentum
function,showthat
15.7 Transfer Functions 961
(a)1
(2π¯h)3/2integraldisplay
rψ(r)e−ir·p/¯hd3r=i¯h∇pϕ(p),
(b)1
(2π¯h)3/2integraldisplay
r2ψ(r)e−r·p/¯hd3r=(i¯h∇p)2ϕ(p).
Note.∇pis thegradientinmomentumspace:
ˆx∂
∂px+ˆy∂
∂py+ˆz∂
∂pz.
These results may be extended to any positive integer power of rand therefore to any
(analytic)functionthatmaybeexpandedas aMaclaurinseriesin r.
15.6.11 The ordinary space wave function ψ(r,t)satisfies the time-dependent Schrödinger
equation
i¯h∂ψ(r,t)
∂t=−¯h2
2m∇2ψ+V(r)ψ.
Show that the corresponding time-dependent momentum wave function satisfies the
analogousequation,
i¯h∂ϕ(p,t)
∂t=p2
2mϕ+V(i¯h∇p)ϕ.
Note. Assume that V(r)may be expressed by a Maclaurin series and use Exer-
cise 15.6.10. V(i¯h∇p)is the same function of the variable i¯h∇pthatV(r)is of the
variabler.
15.6.12 Theone-dimensionaltime-independentSchrödingerwaveequationis
−¯h2
2md2ψ(x)
dx2+V(x)ψ(x)=Eψ(x).
For the special case of V(x)an analytic function of x, show that the corresponding
momentumwaveequationis
Vparenleftbigg
i¯hd
dpparenrightbigg
g(p)+p2
2mg(p)=Eg(p).
Derive this momentum wave equation from the Fourier transform, Eq. (15.62), and its
inverse.Donotusethesubstitution x→i¯h(d/dp)directly.
15.7 T RANSFER FUNCTIONS
A time-dependent electrical pulse may be regarded as built-up as a superposition of plane
wavesof manyfrequencies.Forangularfrequency ωwehaveacontribution
F(ω)eiωt.
Thenthecompletepulsemaybewrittenas
f(t)=1
2πintegraldisplay∞
−∞F(ω)eiωtdω. (15.82)
962 Chapter 15 Integral Transforms
FIGURE 15.6Servomechanismorastereoamplifier.
Becausetheangularfrequency ωis relatedtothelinearfrequency νby
ν=ω
2π,
itiscustomarytoassociatetheentire 1 /2πfactorwiththisintegral.
But ifωis a frequency, what about the negative frequencies? The negative ωmay be
lookedonasamathematicaldevicetoavoiddealingwithtwofunctions(cos ωtandsinωt)
separately(compareSection14.1).
Because Eq. (15.82) has the form of a Fourier transform, we may solve for F(ω)by
writingtheinversetransform,
F(ω)=integraldisplay∞
−∞f(t)e−iωtdt. (15.83)
Equation(15.83)representsa resolutionofthepulse f(t)intoitsangularfrequencycom-
ponents.Equation(15.82) isa synthesis ofthepulse fromits components.
Consider some device, such as a servomechanism or a stereo amplifier (Fig. 15.6), with
an inputf(t)and an output g(t). For an input of a single frequency ω,fω(t)=eiωt,t h e
amplifierwill alter the amplitudeand may also changethe phase. The changeswill proba-
blydependonthefrequency.Hence
gω(t)=ϕ(ω)fω(t). (15.84)
This amplitudes-and phase-modifyingfunction ϕ(ω)is calleda transferfunction. It usu-
allywillbecomplex:
ϕ(ω)=u(ω)+iv(ω), (15.85)
wherethefunctions u(ω)andv(ω)arereal.
In Eq. (15.84) we assume that the transfer function ϕ(ω)is independentof input ampli-
tude and of the presence or absence of any other frequency components. That is, we are
assuming a linear mapping of f(t)ontog(t). Then the total output may be obtained by
integratingovertheentireinput,as modifiedbytheamplifier
g(t)=1
2πintegraldisplay∞
−∞ϕ(ω)F(ω)eiωtdω. (15.86)
The transfer function is characteristic of the amplifier. Once the transfer function is
known (measured or calculated), the output g(t)can be calculated for any input f(t).
Letusconsider ϕ(ω)as theFourier(inverse)transformof somefunction /Phi1(t):
ϕ(ω)=integraldisplay∞
−∞/Phi1(t)e−iωtdt. (15.87)
15.7 Transfer Functions 963
ThenEq.(15.86)istheFouriertransformoftwoinversetransforms.FromSection15.5we
obtaintheconvolution
g(t)=integraldisplay∞
−∞f(τ)/Phi1(t−τ)dτ. (15.88)
Interpreting Eq. (15.88), we have an input—a “cause”— f(τ), modified by /Phi1(t−τ),
producing an output—an “effect”— g(t). Adopting the concept of causality —that the
causeprecedestheeffect—wemustrequire τ<t. Wedothisbyrequiring
/Phi1(t−τ)=0,τ>t. (15.89)
ThenEq. (15.88)becomes
g(t)=integraldisplayt
−∞f(τ)/Phi1(t−τ)dτ. (15.90)
TheadoptionofEq.(15.89)hasprofoundconsequenceshereandequivalentlyindisper-
siontheory,Section7.2.
Significance of /Phi1(t)
Toseethesignificanceof /Phi1,l etf(τ)bea suddenimpulsestartingat τ=0,
f(τ)=δ(τ),
whereδ(τ)isaDiracdeltadistributiononthepositivesideoftheorigin.ThenEq.(15.90)
becomes
g(t)=integraldisplayt
−∞δ(τ)/Phi1(t−τ)dτ,
(15.91)
g(t)=braceleftbigg/Phi1(t), t > 0,
0,t <0.
This identifies /Phi1(t)as the output function correspondingto a unit impulse at t=0. Equa-
tion (15.91) also serves to establish that /Phi1(t)is real. Our original transfer function gives
thesteady-stateoutputcorrespondingtoaunit-amplitudesingle-frequencyinput. /Phi1(t)and
ϕ(ω)areFouriertransforms ofeachother.
FromEq. (15.87)wenowhave
ϕ(ω)=integraldisplay∞
0/Phi1(t)e−iωtdt, (15.92)
with the lower limit set equal to zero by causality (Eq. (15.89)). With /Phi1(t)real from
Eq.(15.91) weseparaterealandimaginaryparts andwrite
u(ω)=integraldisplay∞
0/Phi1(t)cosωtdt,
(15.93)
v(ω)=−integraldisplay∞
0/Phi1(t)sinωtdt, ω> 0.
964 Chapter 15 Integral Transforms
From this we see that the real part of ϕ(ω),u(ω) , is even, whereas the imaginary part of
ϕ(ω),v(ω) ,is odd:
u(−ω)=u(ω), v(−ω)=−v(ω).
ComparethisresultwithExercise15.3.1.
InterpretingEq.(15.93)as Fouriercosineandsinetransforms, wehave
/Phi1(t)=2
πintegraldisplay∞
0u(ω)cosωtdω
=−2
πintegraldisplay∞
0v(ω)sinωtdω, t > 0. (15.94)
CombiningEqs. (15.93) and(15.94), weobtain
v(ω)=−integraldisplay∞
0sinωtbraceleftbigg2
πintegraldisplay∞
0u(ω′)cosω′tdω′bracerightbigg
dt, (15.95)
showingthatifourtransferfunctionhasarealpart,itwillalsohaveanimaginarypart(and
viceversa).Ofcourse,thisassumesthattheFouriertransformsexist,thusexcludingcases
suchas/Phi1(t)=1.
The imposition of causality has led to a mutual interdependence of the real and imagi-
nary parts of the transfer function. The reader should compare this with the results of the
dispersiontheoryofSection7.2, alsoinvolvingcausality.
It may be helpful to show that the parity properties of u(ω)andv(ω)require/Phi1(t)to
vanishfor negative t. InvertingEq. (15.87),wehave
/Phi1(t)=1
2πintegraldisplay∞
−∞bracketleftbig
u(ω)+iv(ω)bracketrightbigbracketleftbig
cosωt+isinωtbracketrightbig
dω. (15.96)
Withu(ω)evenand v(ω)odd,Eq. (15.96)becomes
/Phi1(t)=1
πintegraldisplay∞
0u(ω)cosωtdω−1
πintegraldisplay∞
0v(ω)sinωtdω. (15.97)
FromEq. (15.94),
integraldisplay∞
0u(ω)cosωtdω=−integraldisplay∞
0v(ω)sinωtdω, t > 0. (15.98)
If wereverse thesignof t,sinωtreverses signand,fromEq. (15.97),
/Phi1(t)=0,t<0
(demonstratingtheinternalconsistencyofouranalysis).
Exercise
15.7.1 Derivetheconvolution
g(t)=integraldisplay∞
−∞f(τ)/Phi1(t−τ)dτ.
15.8 Laplace Transforms 965
15.8 L APLACE TRANSFORMS
Definition
TheLaplacetransform f(s)orLof afunction F(t)is definedby11
f(s)=Lbraceleftbig
F(t)bracerightbig
=lima→∞integraldisplaya
0e−stF(t)dt=integraldisplay∞
0e−stF(t)dt. (15.99)
Afewcommentsontheexistenceoftheintegralareinorder.Theinfiniteintegralof F(t),
integraldisplay∞
0F(t)dt,
neednotexist .Forinstance, F(t)maydivergeexponentiallyforlarge t.However,ifthere
issomeconstant s0suchthat
vextendsinglevextendsinglee−s0tF(t)vextendsinglevextendsingle≤M, (15.100)
a positive constant for sufficiently large t,t>t0, the Laplace transform (Eq. (15.99)) will
exist fors>s0;F(t)is said to be of exponential order . As a counterexample, F(t)=et2
doesnotsatisfytheconditiongivenbyEq.(15.100)andis notofexponentialorder. L{et2}
doesnotexist.
The Laplace transform may also fail to exist because of a sufficiently strong singularity
inthefunction F(t)ast→0;thatis,
integraldisplay∞
0e−sttndt
divergesattheoriginfor n≤−1.TheLaplacetransform L{tn}doesnotexistfor n≤−1.
Since,fortwo functions F(t)andG(t), for whichtheintegralsexist
Lbraceleftbig
aF(t)+bG(t)bracerightbig
=aLbraceleftbig
F(t)bracerightbig
+bLbraceleftbig
G(t)bracerightbig
, (15.101)
theoperationdenotedby Lislinear.
Elementary Functions
To introduce the Laplace transform, let us apply the operation to some of the elementary
functions.In allcasesweassumethat F(t)=0f o rt<0.If
F(t)=1,t>0,
11Thisissometimescalleda one-sidedLaplacetransform ;theintegralfrom −∞to+∞isreferredtoasa two-sidedLaplace
transform . Some authors introduce an additional factor of s. This extra sappears to have little advantage and continually gets
intheway(compareJeffreysandJeffreys,Section14.13—seetheAdditionalReadings—foradditionalcomments).Generally,
wetakesto bereal and positive. It is possible to have scomplex, provided ℜ(s)>0.
966 Chapter 15 Integral Transforms
then
L{1}=integraldisplay∞
0e−stdt=1
s,fors>0. (15.102)
Again,let
F(t)=ekt,t>0.
TheLaplacetransformbecomes
Lbraceleftbig
ektbracerightbig
=integraldisplay∞
0e−stektdt=1
s−k,fors>k. (15.103)
Usingthisrelation,weobtaintheLaplacetransformofcertainotherfunctions.Since
coshkt=1
2parenleftbig
ekt+e−ktparenrightbig
,sinhkt=1
2parenleftbig
ekt−e−ktparenrightbig
, (15.104)
wehave
L{coshkt}=1
2parenleftbigg1
s−k+1
s+kparenrightbigg
=s
s2−k2,
(15.105)
L{sinhkt}=1
2parenleftbigg1
s−k−1
s+kparenrightbigg
=k
s2−k2,
bothvalidfor s>k.Wehavetherelations
coskt=coshikt,sinkt=−isinhikt. (15.106)
UsingEqs. (15.105)with kreplacedby ik,wefindthattheLaplacetransforms are
L{coskt}=s
s2+k2,
(15.107)
L{sinkt}=k
s2+k2,
both valid for s>0. Another derivation of this last transform is given in the next sec-
tion. Note that lim s→0L{sinkt}=1/k. The Laplace transform assigns a value of 1 /ktointegraltext∞
0sinktdt.
Finally,for F(t)=tn,weha v e
Lbraceleftbig
tnbracerightbig
=integraldisplay∞
0e−sttndt,
whichisjustthefactorialfunction.Hence
Lbraceleftbig
tnbracerightbig
=n!
sn+1,s>0,n>−1. (15.108)
Note that in all these transforms we have the variable sin the denominator—negative
powers of s. In particular, lim s→∞f(s)=0. The significance of this point is that if f(s)
involvespositivepowersof s(lims→∞f(s)→∞), thennoinversetransformexists.
15.8 Laplace Transforms 967
Inverse Transform
Thereislittleimportancetotheseoperationsunlesswecancarryouttheinversetransform,
asinFouriertransforms. Thatis, with
Lbraceleftbig
F(t)bracerightbig
=f(s),
then
L−1braceleftbig
f(s)bracerightbig
=F(t). (15.109)
This inverse transform is notunique. Two functions F1(t)andF2(t)may have the same
transform, f(s). However,inthis case
F1(t)−F2(t)=N(t),
whereN(t)is anullfunction(Fig.15.7), indicatingthat
integraldisplayt0
0N(t)dt=0,
for all positive t0. This result is known as Lerch’s theorem . Therefore to the physicist
and engineer N(t)may almost always be taken as zero and the inverse operation becomes
unique.
The inverse transform can be determined in various ways. (1) A table of transforms can
bebuiltupandusedtocarryouttheinversetransformation,exactlyasatableoflogarithms
canbeusedtolookupantilogarithms.Theprecedingtransforms constitutethebeginnings
of such a table. For a more complete set of Laplace transforms see upcoming Table 15.2
or AMS-55, Chapter 29 (see footnote 4 in Chapter 5 for the reference). Employing partial
fractionexpansionsandvariousoperationaltheorems,whichareconsideredinsucceeding
sections,facilitatesuse ofthetables.
•There is some justification for suspecting that these tables are probably of more value
insolvingtextbookexercisesthaninsolvingreal-worldproblems.
•(2) A general technique for L−1will be developed in Section 15.12 by using the cal-
culusofresidues.
FIGURE 15.7Apossiblenull
function.
968 Chapter 15 Integral Transforms
•(3)Forthedifficultiesandthepossibilitiesofanumericalapproach—numericalinver-
sion—werefer totheAdditionalReadings.
Partial Fraction Expansion
Utilizationofatableoftransforms(orinversetransforms)isfacilitatedbyexpanding f(s)
inpartialfractions .
Frequently f(s), our transform, occurs in the form g(s)/h(s) , whereg(s)andh(s)are
polynomials with no common factors, g(s)being of lower degree than h(s). If the factors
ofh(s)arealllinearanddistinct,thenbythemethodofpartialfractionswemaywrite
f(s)=c1
s−a1+c2
s−a2+···+cn
s−an, (15.110)
wherethe ciareindependentof s.Theaiaretherootsof h(s).Ifanyoneoftheroots,say,
a1, ismultiple(occurring mtimes), then f(s)hastheform
f(s)=c1,m
(s−a1)m+c1,m−1
(s−a1)m−1+···+c1,1
s−a1+nsummationdisplay
i=2ci
s−ai. (15.111)
Finally, if one of the factors is quadratic, (s2+ps+q), then the numerator, instead of
beingasimpleconstant,willhavetheform
as+b
s2+ps+q.
There are various ways of determining the constants introduced. For instance, in
Eq.(15.110)wemaymultiplythroughby (s−ai)andobtain
ci=lims→ai(s−ai)f(s). (15.112)
Inelementarycasesadirectsolutionisoftentheeasiest.
Example 15.8.1 PARTIAL FRACTION EXPANSION
Let
f(s)=k2
s(s2+k2)=c
s+as+b
s2+k2. (15.113)
Puttingtherightsideoftheequationoveracommondenominatorandequatinglikepowers
ofsinthenumerator,weobtain
k2
s(s2+k2)=c(s2+k2)+s(as+b)
s(s2+k2), (15.114)
c+a=0,s2;b=0,s1;ck2=k2,s0.
Solvingthese (s/negationslash=0),weha v e
c=1,b=0,a=−1,
15.8 Laplace Transforms 969
giving
f(s)=1
s−s
s2+k2, (15.115)
and
L−1braceleftbig
f(s)bracerightbig
=1−coskt (15.116)
byEqs. (15.102)and(15.106). /squaresolid
Example 15.8.2 ASTEPFUNCTION
AsoneapplicationofLaplacetransforms, considertheevaluationof
F(t)=integraldisplay∞
0sintx
xdx. (15.117)
SupposewetaketheLaplacetransformofthis definite(andimproper)integral:
Lbraceleftbiggintegraldisplay∞
0sintx
xdxbracerightbigg
=integraldisplay∞
0e−stintegraldisplay∞
0sintx
xdxdt. (15.118)
Now,interchangingtheorderof integration(whichis justified),12weget
integraldisplay∞
01
xbracketleftbiggintegraldisplay∞
0e−stsintxdtbracketrightbigg
dx=integraldisplay∞
0dx
s2+x2, (15.119)
sincethefactorinsquarebracketsisjusttheLaplacetransformofsin tx.Fromtheintegral
tables,
integraldisplay∞
0dx
s2+x2=1
stan−1parenleftbiggx
sparenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle∞
0=π
2s=f(s). (15.120)
ByEq. (15.102)wecarry outtheinversetransformationtoobtain
F(t)=π
2,t>0, (15.121)
in agreement with an evaluation by the calculus of residues (Section 7.1). It has been as-
sumed that t>0i nF(t).F o rF(−t)we need note only that sin (−tx)=−sintx,g i v i n g
F(−t)=−F(t).Finally,if t=0,F(0)is clearlyzero. Therefore
integraldisplay∞
0sintx
xdx=π
2bracketleftbig
2u(t)−1bracketrightbig
=
π
2,t>0
0,t=0
−π
2,t <0.(15.122)
Note thatintegraltext∞
0(sintx/x)dx, taken as a function of t, describes a step function (Fig. 15.8),
astepofheight πatt=0.Thisis consistentwithEq. (1.174). /squaresolid
The technique in the preceding example was to (1) introduce a second integration—
the Laplace transform, (2) reverse the order of integration and integrate, and (3) take the
12See—in the Additional Readings—Jeffreys and Jeffreys (1966), Chapter 1 (uniform convergence of integrals).
970 Chapter 15 Integral Transforms
FIGURE 15.8F(t)=integraltext∞
0sintx
xdx,
astepfunction.
inverseLaplacetransform.Therearemanyopportunitieswherethistechniqueofreversing
the order of integration can be applied and proved useful. Exercise 15.8.6 is a variation of
this.
Exercises
15.8.1 Provethat
lims→∞sf(s)=lim
t→+0F(t).
Hint.Assumethat F(t)canbeexpressedas F(t)=summationtext∞
n=0antn.
15.8.2 Showthat
1
πlim
s→0L{cosxt}=δ(x).
15.8.3 Verifythat
Lbraceleftbiggcosat−cosbt
b2−a2bracerightbigg
=s
(s2+a2)(s2+b2),a2/negationslash=b2.
15.8.4 Usingpartialfractionexpansions,showthat
(a)L−1braceleftbigg1
(s+a)(s+b)bracerightbigg
=e−at−e−bt
b−a,a/negationslash=b.
(b)L−1braceleftbiggs
(s+a)(s+b)bracerightbigg
=ae−at−be−bt
a−b,a/negationslash=b.
15.8.5 Usingpartialfractionexpansions,showthatfor a2/negationslash=b2,
(a)L−1braceleftbigg1
(s2+a2)(s2+b2)bracerightbigg
=−1
a2−b2braceleftbiggsinat
a−sinbt
bbracerightbigg
,
(b)L−1braceleftbiggs2
(s2+a2)(s2+b2)bracerightbigg
=1
a2−b2{asinat−bsinbt}.
15.9 Laplace Transform of Derivatives 971
15.8.6 The electrostatic potential of a charged conducting disk is known to have the general
form(circularcylindricalcoordinates)
/Phi1(ρ,z)=integraldisplay∞
0e−k|z|J0(kρ)f(k)dk,
withf(k)unknown. At large distances (z→∞)the potential must approach the
Coulombpotential Q/4πε0z. Showthat
lim
k→0f(k)=q
4πε0.
Hint. You may set ρ=0 and assume a Maclaurin expansion of f(k)or, using e−kz,
constructadeltasequence.
15.8.7 Showthat
(a)integraldisplay∞
0coss
sνds=π
2(ν−1)!cos(νπ/2),0<ν<1,
(b)integraldisplay∞
0sins
sνds=π
2(ν−1)!sin(νπ/2),0<ν<2,
Whyisνrestrictedto(0,1)for(a),to (0,2)for(b)?Theseintegralsmaybeinterpreted
asFouriertransformsof s−νandasMellintransformsof sin sand coss.
Hint. Replace s−νby a Laplace transform integral: L{tν−1}/(ν−1)!. Then integrate
withrespectto s.Theresultingintegralcanbetreatedasabetafunction(Section8.4).
15.8.8 Afunction F(t)canbeexpandedinapowerseries (Maclaurin);thatis,
F(t)=∞summationdisplay
n=0antn.
Then
Lbraceleftbig
F(t)bracerightbig
=integraldisplay∞
0e−st∞summationdisplay
n=0antndt=∞summationdisplay
n=0anintegraldisplay∞
0e−sttndt.
Show that f(s), the Laplace transform of F(t), contains no powers of sgreater than
s−1. Checkyourresultbycalculating L{δ(t)},andcommentonthisfiasco.
15.8.9 ShowthattheLaplacetransformof M(a,c,x) is
Lbraceleftbig
M(a,c,x)bracerightbig
=1
s2F1parenleftbigg
a,1;c,1
sparenrightbigg
.
15.9 L APLACE TRANSFORM OF DERIVATIVES
Perhaps the main application of Laplace transforms is in converting differential equations
intosimplerformsthatmaybesolvedmoreeasily.Itwillbeseen,forinstance,thatcoupled
differentialequationswithconstantcoefficientstransformtosimultaneouslinearalgebraic
equations.
972 Chapter 15 Integral Transforms
Letus transform thefirstderivativeof F(t):
Lbraceleftbig
F′(t)bracerightbig
=integraldisplay∞
0e−stdF(t)
dtdt.
Integratingbyparts, weobtain
Lbraceleftbig
F′(t)bracerightbig
=e−stF(t)vextendsinglevextendsinglevextendsingle∞
0+sintegraldisplay∞
0e−stF(t)dt
=sLbraceleftbig
F(t)bracerightbig
−F(0). (15.123)
Strictlyspeaking, F(0)=F(+0)13anddF/dtisrequiredtobeatleastpiecewisecontin-
uous for 0≤t<∞. Naturally, both F(t)and its derivative must be such that the integrals
do not diverge. Incidentally, Eq. (15.123) provides another proof of Exercise 15.8.8. An
extensiongives
Lbraceleftbig
F(2)(t)bracerightbig
=s2Lbraceleftbig
F(t)bracerightbig
−sF(+0)−F′(+0), (15.124)
Lbraceleftbig
F(n)(t)bracerightbig
=snLbraceleftbig
F(t)bracerightbig
−sn−1F(+0)−···−F(n−1)(+0).(15.125)
The Laplace transform, like the Fourier transform, replaces differentiation with multi-
plication.InthefollowingexamplesODEsbecomealgebraicequations.Hereisthepower
and the utility of the Laplace transform. But see Example 15.10.3 for what may happen if
thecoefficientsare notconstant.
Note how the initial conditions, F(+0),F′(+0), and so on, are incorporated into the
transform.Equation(15.124)maybeusedtoderive L{sinkt}.We usetheidentity
−k2sinkt=d2
dt2sinkt. (15.126)
ThenapplyingtheLaplacetransformoperation,wehave
−k2L{sinkt}=Lbraceleftbiggd2
dt2sinktbracerightbigg
=s2L{sinkt}−ssin(0)−d
dtsinktvextendsinglevextendsinglevextendsingle
t=0. (15.127)
Since sin(0)=0 andd/dtsinkt|t=0=k,
L{sinkt}=k
s2+k2, (15.128)
verifyingEq.(15.107).
13Zero is approached from thepositive side.
15.9 Laplace Transform of Derivatives 973
Example 15.9.1 SIMPLE HARMONIC OSCILLATOR
Asaphysicalexample,consideramass moscillatingundertheinfluenceofanidealspring,
springconstant k. Asusual,frictionisneglected.ThenNewton’ssecondlawbecomes
md2X(t)
dt2+kX(t)=0; (15.129)
also,wetakeasinitialconditions
X(0)=X0,X′(0)=0.
ApplyingtheLaplacetransform, weobtain
mLbraceleftbiggd2X
dt2bracerightbigg
+kLbraceleftbig
X(t)bracerightbig
=0, (15.130)
andbyuseof Eq.(15.124)this becomes
ms2x(s)−msX0+kx(s)=0, (15.131)
x(s)=X0s
s2+ω2
0,withω2
0≡k
m. (15.132)
FromEq. (15.107)thisis seentobethetransform of cos ω0t, whichgives
X(t)=X0cosω0t, (15.133)
asexpected. /squaresolid
Example 15.9.2 EARTH ’SNUTATION
A somewhat more involved example is the nutation of the earth’s poles (force-free pre-
cession). If we treat the Earth as a rigid (oblate) spheroid, the Euler equations of motion
reduceto
dX
dt=−aY,dY
dt=+aX, (15.134)
wherea≡[(Iz−Ix)/Iz]ωz,X=ωx,Y=ωywith angular velocity vector ω=
(ωx,ωy,ωz)(Fig. 15.9), Iz=moment of inertia about the z-axis and Iy=Ixmoment
of inertia about the x-( o ry-)axis. The z-axis coincides with the axis of symmetry of the
Earth.ItdiffersfromtheaxisfortheEarth’sdailyrotation, ω,bysome15meters,measured
atthepoles. Transformationofthesecoupleddifferentialequationsyields
sx(s)−X(0)=−ay(s), sy(s) −Y(0)=ax(s). (15.135)
Combiningtoeliminate y(s),w eh a v e
s2x(s)−sX(0)+aY(0)=−a2x(s),
or
x(s)=X(0)s
s2+a2−Y(0)a
s2+a2. (15.136)
974 Chapter 15 Integral Transforms
FIGURE 15.9
Hence
X(t)=X(0)cosat−Y(0)sinat. (15.137)
Similarly,
Y(t)=X(0)sinat+Y(0)cosat. (15.138)
This is seen to be a rotation of the vector (X,Y)counterclockwise (for a>0) about the
z-axiswithangle θ=atandangularvelocity a.
Adirectinterpretationmaybefoundbychoosingthetimeaxisso that Y(0)=0.Then
X(t)=X(0)cosat, Y(t)=X(0)sinat, (15.139)
whicharetheparametricequationsforrotationof (X,Y)inacircularorbitofradius X(0),
withangularvelocity ainthecounterclockwisesense.
In the case of the Earth’s angular velocity, vector X(0)is about 15 meters, whereas
a, as defined here, corresponds to a period (2π/a)of some 300 days. Actually because
of departures from the idealized rigid body assumed in setting up Euler’s equations, the
periodis about427days.14If inEq. (15.134)weset
X(t)=Lx,Y(t)=Ly,
whereLxandLyare thex- andy-components of the angular momentum L,a=−gLBz,
gLis the gyromagnetic ratio, and Bzis the magnetic field (along the z-axis), then
Eq. (15.134) describes the Larmor precession of charged bodies in a uniform magnetic
fieldBz. /squaresolid
14D. Menzel, ed., Fundamental Formulas of Physics , Englewood Cliffs, NJ: Prentice-Hall (1955), reprinted, 2nd ed., Dover
(1960), p. 695.
15.9 Laplace Transform of Derivatives 975
Dirac Delta Function
Forusewithdifferentialequationsonefurthertransformishelpful—theDiracdeltafunc-
tion:15
Lbraceleftbig
δ(t−t0)bracerightbig
=integraldisplay∞
0e−stδ(t−t0)dt=e−st0,fort0≥0, (15.140)
andfort0=0
Lbraceleftbig
δ(t)bracerightbig
=1, (15.141)
whereitis assumedthatweareusingarepresentationof thedeltafunctionsuchthat
integraldisplay∞
0δ(t)dt=1,δ(t)=0,fort>0. (15.142)
Asanalternatemethod, δ(t)maybeconsideredthelimitas ε→0o fF(t), where
F(t)=
0,t <0,
ε−1,0<t<ε,
0,t >ε.(15.143)
Bydirectcalculation
Lbraceleftbig
F(t)bracerightbig
=1−e−εs
εs. (15.144)
Takingthelimitoftheintegral(insteadoftheintegralofthelimit),wehave
lim
ε→0Lbraceleftbig
F(t)bracerightbig
=1,
orEq. (15.141),
Lbraceleftbig
δ(t)bracerightbig
=1.
This delta function is frequently called the impulsefunction because it is so useful in
describingimpulsiveforces,thatis, forceslastingonlyashort time.
Example 15.9.3 IMPULSIVE FORCE
Newton’ssecondlawfor impulsiveforceactingonaparticleof mass mbecomes
md2X
dt2=Pδ(t), (15.145)
wherePisa constant.Transforming,weobtain
ms2x(s)−msX(0)−mX′(0)=P. (15.146)
15Strictlyspeaking,theDiracdeltafunctionisundefined.However,theintegraloveritiswelldefined.Thisapproachisdeveloped
inSection1.16 using deltasequences.
976 Chapter 15 Integral Transforms
Foraparticlestartingfromrest, X′(0)=0.16We shallalsotake X(0)=0.Then
x(s)=P
ms2, (15.147)
and
X(t)=P
mt, (15.148)
dX(t)
dt=P
m,aconstant . (15.149)
Theeffectoftheimpulse Pδ(t)istotransfer(instantaneously) Punitsoflinearmomentum
totheparticle.
Asimilaranalysisappliestotheballisticgalvanometer.Thetorqueonthegalvanometer
is given initially by kι, in which ιis a pulse of current and kis a proportionality constant.
Sinceιis ofshort duration,weset
kι=kq δ(t), (15.150)
whereqis thetotalchargecarriedbythecurrent ι.Then,with Ithemomentofinertia,
Id2θ
dt2=kq δ(t), (15.151)
and, transforming as before, we find that the effect of the current pulse is a transfer of kq
unitsofangularmomentumtothegalvanometer. /squaresolid
Exercises
15.9.1 Use the expression for the transform of a second derivative to obtain the transform of
coskt.
15.9.2 Amassmisattachedtooneendofanunstretchedspring,springconstant k(Fig.15.10).
Attimet=0thefreeendofthespringexperiencesaconstantacceleration a,awayfrom
themass.UsingLaplacetransforms,
FIGURE 15.10Spring.
16This should be X′(+0). Toinclude the effect of the impulse, consider that the impulse will occurat t=εandletε→0.
15.9 Laplace Transform of Derivatives 977
(a) Findtheposition xofmasafunctionof time.
(b) Determinethelimitingformof x(t)for small t.
ANS.(a)x=1
2at2−a
ω2(1−cosωt), ω2=k
m,
(b)x=aω2
4!t4,ωt≪1.
15.9.3 Radioactivenucleidecayaccordingtothelaw
dN
dt=−λN,
Nbeingtheconcentrationofagivennuclideand λbeingtheparticulardecayconstant.
This equation may be interpreted as stating that the rate of decay is proportional to the
numberof theseradioactivenucleipresent.Theyalldecayindependently.
Inaradioactiveseriesof ndifferentnuclides,startingwith N1,
dN1
dt=−λ1N1,
dN2
dt=λ1N1−λ2N2,andso on .
dNn
dt=λn−1Nn−1,stable.
FindN1(t),N2(t),N3(t),n=3,withN1(0)=N0,N2(0)=N3(0)=0.
ANS.N1(t)=N0e−λ1t,N2(t)=N0λ1
λ2−λ1parenleftbig
e−λ1t−e−λ2tparenrightbig
,
N3(t)=N0parenleftbigg
1−λ2
λ2−λ1e−λ1t+λ1
λ2−λ1e−λ2tparenrightbigg
.
Findanapproximateexpressionfor N2andN3, validfor small twhenλ1≈λ2.
ANS.N2≈N0λ1t,N3≈N0
2λ1λ2t2.
Findapproximateexpressionsfor N2andN3, validfor large t,when
(a)λ1≫λ2,
(b)λ1≪λ2.
ANS.(a) N2≈N0e−λ2t,
N3≈N0parenleftbig
1−e−λ2tparenrightbig
,λ1t≫1.
(b)N2≈N0λ1
λ2e−λ1t,
N3≈N0parenleftbig
1−e−λ1tparenrightbig
,λ2t≫1.
15.9.4 Theformationof anisotopeinanuclearreactorisgivenby
dN2
dt=nvσ1N10−λ2N2(t)−nvσ2N2(t).
978 Chapter 15 Integral Transforms
Heretheproduct nvistheneutronflux,neutronspercubiccentimeter,timescentimeters
per second mean velocity; σ1andσ2(cm2) are measures of the probability of neutron
absorption by the original isotope, concentration N10, which is assumed constant and
the newly formed isotope, concentration N2, respectively. The radioactive decay con-
stantfortheisotopeis λ2.
(a) Findtheconcentration N2ofthenewisotopeasa functionoftime.
(b) If the original element is Eu153,σ1=400 barns=400×10−24cm2,σ2=
1000 barns=1000×10−24cm2, andλ2=1.4×10−9s−1.I fN10=1020and
(nv)=109cm−2s−1, findN2, the concentration of Eu154after one year of con-
tinuousirradiation.Is theassumptionthat N1isconstantjustified?
15.9.5 InanuclearreactorXe135isformedasbothadirectfissionproductandadecayproduct
of I135, half-life, 6.7 hours. The half-life of Xe135is 9.2 hours. Because Xe135strongly
absorbs thermalneutronsthereby“poisoning”the nuclearreactor, its concentrationis a
matterof greatinterest.Therelevantequationsare
dNI
dt=γIϕσfNU−λINI,
dNX
dt=λINI+γXϕσfNU−λXNX−ϕσXNX.
HereNI=concentrationof I135(Xe135,U235). Assume
NU=constant,
γI=yieldofI135per fission=0.060,
γX=yieldofXe135directfromfission =0.003,
λI=I135parenleftbig
Xe135parenrightbig
decayconstant =ln2
t1/2=0.693
t1/2,
σf=thermalneutronfissioncross sectionfor U235,
σX=thermalneutronabsorptioncross sectionfor Xe135
=3.5×106barns=3.5×10−18cm2.
(σItheabsorptioncross sectionof I135,is negligible. )
ϕ=neutronflux=neutrons/cm3×meanvelocity(cm /s).
(a) Find NX(t)intermsofneutronflux ϕandtheproduct σfNU.
(b) Find NX(t→∞).
(c) After NXhasreachedequilibrium,thereactorisshutdown, ϕ=0.FindNX(t)fol-
lowing shutdown. Notice the increase in NX, which may for a few hours interfere
withstartingthereactorupagain.
15.10 Other Properties 979
15.10 O THER PROPERTIES
Substitution
If we replace the parameter sbys−ain the definition of the Laplace transform
(Eq. (15.99)), wehave
f(s−a)=integraldisplay∞
0e−(s−a)tF(t)dt=integraldisplay∞
0e−steatF(t)dt
=Lbraceleftbig
eatF(t)bracerightbig
. (15.152)
Hence the replacement of swiths−acorresponds to multiplying F(t)byeat, and con-
versely. This result can be used to good advantage in extending our table of transforms.
FromEq. (15.107)wefindimmediatelythat
Lbraceleftbig
eatsinktbracerightbig
=k
(s−a)2+k2; (15.153)
also,
Lbraceleftbig
eatcosktbracerightbig
=s−a
(s−a)2+k2,s>a.
Example 15.10.1 DAMPED OSCILLATOR
These expressions are useful when we consider an oscillating mass with damping propor-
tionaltothevelocity.Equation(15.129), withsuchdampingadded,becomes
mX′′(t)+bX′(t)+kX(t)=0, (15.154)
in which bis a proportionality constant. Let us assume that the particle starts from rest at
X(0)=X0,X′(0)=0.Thetransformedequationis
mbracketleftbig
s2x(s)−sX0bracketrightbig
+bbracketleftbig
sx(s)−X0bracketrightbig
+kx(s)=0, (15.155)
and
x(s)=X0ms+b
ms2+bs+k. (15.156)
Thismaybehandledbycompletingthesquareofthedenominator:
s2+b
ms+k
m=parenleftbigg
s+b
2mparenrightbigg2
+parenleftbiggk
m−b2
4m2parenrightbigg
. (15.157)
If thedampingissmall, b2<4km, thelasttermispositiveandwillbedenotedby ω2
1:
x(s)=X0s+b/m
(s+b/2m)2+ω2
1
=X0s+b/2m
(s+b/2m)2+ω2
1+X0(b/2mω1)ω1
(s+b/2m)2+ω2
1. (15.158)
980 Chapter 15 Integral Transforms
ByEq. (15.153),
X(t)=X0e−(b/2m)tparenleftbigg
cosω1t+b
2mω1sinω1tparenrightbigg
=X0ω0
ω1e−(b/2m)tcos(ω1t−ϕ), (15.159)
where
tanϕ=b
2mω1,ω2
0=k
m.
Ofcourse,as b→0,thissolutiongoesovertotheundampedsolution(Section15.9). /squaresolid
RLC Analog
Itisworthnotingthesimilaritybetweenthisdampedsimpleharmonicoscillationofamass
on a spring and an RLCcircuit (resistance, inductance, and capacitance) (Fig. 15.11). At
any instant the sum of the potential differences around the loop must be zero (Kirchhoff’s
law,conservationof energy).Thisgives
LdI
dt+RI+1
Cintegraldisplayt
Idt=0. (15.160)
Differentiatingthecurrent Iwithrespecttotime(to eliminatetheintegral),wehave
Ld2I
dt2+RdI
dt+1
CI=0. (15.161)
If we replace I(t)withX(t),Lwithm,Rwithb, andC−1withk, then Eq. (15.161) is
identical with the mechanical problem. It is but one example of the unification of diverse
branchesofphysicsbymathematics.AmorecompletediscussionwillbefoundinOlson’s
book.17
FIGURE 15.11RLCcircuit.
17H.F.Olson, Dynamical Analogies , NewYork: VanNostrand (1943).
15.10 Other Properties 981
FIGURE 15.12Translation.
Translation
Thistimelet f(s)bemultipliedby e−bs,b>0:
e−bsf(s)=e−bsintegraldisplay∞
0e−stF(t)dt
=integraldisplay∞
0e−s(t+b)F(t)dt. (15.162)
Nowlett+b=τ. Equation(15.162)becomes
e−bsf(s)=integraldisplay∞
be−sτF(τ−b)dτ
=integraldisplay∞
0e−sτF(τ−b)u(τ−b)dτ, (15.163)
whereu(τ−b)istheunitstepfunction.Thisrelationisoftencalledthe Heavisideshifting
theorem (Fig. 15.12).
SinceF(t)is assumed to be equal to zero for t<0,F(τ−b)=0f o r0≤τ<b.
Thereforewecanextendthelowerlimittozerowithoutchangingthevalueoftheintegral.
Then,notingthat τis onlyavariableof integration,weobtain
e−bsf(s)=Lbraceleftbig
F(t−b)bracerightbig
. (15.164)
Example 15.10.2 ELECTROMAGNETIC WAVES
The electromagnetic wave equation with E=EyorEz, a transverse wave propagating
alongthe x-axis,is
∂2E(x,t)
∂x2−1
v2∂2E(x,t)
∂t2=0. (15.165)
Transformingthisequationwithrespectto t, weget
∂2
∂x2Lbraceleftbig
E(x,t)bracerightbig
−s2
v2Lbraceleftbig
E(x,t)bracerightbig
+s
v2E(x,0)+1
v2∂E(x,t)
∂tvextendsinglevextendsinglevextendsinglevextendsingle
t=0=0.(15.166)
982 Chapter 15 Integral Transforms
If wehavetheinitialcondition E(x,0)=0 and
∂E(x,t)
∂tvextendsinglevextendsinglevextendsinglevextendsingle
t=0=0,
then
∂2
∂x2Lbraceleftbig
E(x,t)bracerightbig
=s2
v2Lbraceleftbig
E(x,t)bracerightbig
. (15.167)
Thesolution(of this ODE)is
Lbraceleftbig
E(x,t)bracerightbig
=c1e−(s/v)x+c2e+(s/v)x. (15.168)
The “constants” c1andc2are obtained by additional boundary conditions. They are
constant with respect to xbut may depend on s. If our wave remains finite as x→
∞,L{E(x,t)}will also remain finite. Hence c2=0. IfE(0,t)is denoted by F(t), then
c1=f(s)and
Lbraceleftbig
E(x,t)bracerightbig
=e−(s/v)xf(s). (15.169)
Fromthetranslationproperty(Eq. (15.164))wefindimmediatelythat
E(x,t)=braceleftBiggFparenleftbig
t−x
vparenrightbig
,t≥x
v,
0,t <x
v.(15.170)
Differentiation and substitution into Eq. (15.165) verifies Eq. (15.170). Our solution rep-
resents a wave (or pulse) moving in the positive x-direction with velocity v. Note that for
x>v tthe region remains undisturbed; the pulse has not had time to get there. If we had
wanted a signal propagated along the negative x-axis,c1would have been set equal to 0
andwewouldhaveobtained
E(x,t)=braceleftBigg
Fparenleftbig
t+x
vparenrightbig
,t≥−x
v,
0,t <−x
v,(15.171)
awavealongthenegative x-axis. /squaresolid
Derivative of a Transform
WhenF(t), which is at least piecewise continuous, and sare chosen so that e−stF(t)
convergesexponentiallyforlarge s, theintegral
integraldisplay∞
0e−stF(t)dt
is uniformly convergent and may be differentiated (under the integral sign) with respect
tos. Then
f′(s)=integraldisplay∞
0(−t)e−stF(t)dt=Lbraceleftbig
−tF(t)bracerightbig
. (15.172)
Continuingthisprocess, weobtain
f(n)(s)=Lbraceleftbig
(−t)nF(t)bracerightbig
. (15.173)
15.10 Other Properties 983
Alltheintegralssoobtainedwillbeuniformlyconvergentbecauseofthedecreasingexpo-
nentialbehaviorof e−stF(t).
This sametechniquemaybeappliedtogeneratemoretransforms. Forexample,
Lbraceleftbig
ektbracerightbig
=integraldisplay∞
0e−stektdt=1
s−k,s>k. (15.174)
Differentiatingwithrespectto s(orwithrespectto k), weobtain
Lbraceleftbig
tektbracerightbig
=1
(s−k)2,s>k. (15.175)
Example 15.10.3 BESSEL ’SEQUATION
An interesting application of a differentiated Laplace transform appears in the solution of
Bessel’sequationwith n=0.FromChapter11wehave
x2y′′(x)+xy′(x)+x2y(x)=0. (15.176)
Dividing by xand substituting t=xandF(t)=y(x)to agree with the present notation,
wesee thattheBesselequationbecomes
tF′′(t)+F′(t)+tF(t)=0. (15.177)
We need a regular solution, in particular, F(0)=1. From Eq. (15.177) with t=0,
F′(+0)=0. Also, we assume that our unknown F(t)has a transform. Transforming and
usingEqs. (15.123), (15.124),and(15.172),wehave
−d
dsbracketleftbig
s2f(s)−sbracketrightbig
+sf(s)−1−d
dsf(s)=0. (15.178)
RearrangingEq. (15.178),weobtain
parenleftbig
s2+1parenrightbig
f′(s)+sf(s)=0, (15.179)
or
df
f=−sds
s2+1, (15.180)
afirst-order ODE.Byintegration,
lnf(s)=−1
2lnparenleftbig
s2+1parenrightbig
+lnC, (15.181)
whichmayberewrittenas
f(s)=C√
s2+1. (15.182)
To make use of Eq. (15.108), we expand f(s)in a series of negative powers of s, conver-
gentfors>1:
f(s)=C
sparenleftbigg
1+1
s2parenrightbigg−1/2
=C
sbracketleftbigg
1−1
2s2+1·3
22·2!s4−···+(−1)n(2n)!
(2nn!)2s2n+···bracketrightbigg
.(15.183)
984 Chapter 15 Integral Transforms
Inverting,term byterm,weobtain
F(t)=C∞summationdisplay
n=0(−1)nt2n
(2nn!)2. (15.184)
WhenCis set equal to 1, as required by the initial condition F(0)=1,F(t)is justJ0(t),
ourfamiliarBesselfunctionoforderzero.Hence
Lbraceleftbig
J0(t)bracerightbig
=1√
s2+1. (15.185)
Notethatweassumed s>1.Theprooffor s>0 is leftasaproblem.
It is worth noting that this application was successful and relatively easy because we
tookn=0 in Bessel’s equation. This made it possible to divide out a factor of x(ort).
If this had not been done, the terms of the form t2F(t)would have introduced a second
derivative of f(s). The resulting equation would have been no easier to solve than the
originalone.
WhenwegobeyondlinearODEswithconstantcoefficients,theLaplacetransformmay
stillbeapplied,butthereis noguaranteethatitwillbehelpful.
The application to Bessel’s equation, n/negationslash=0, will be found in the references. Alterna-
tively,wecanshowthat
Lbraceleftbig
Jn(at)bracerightbig
=a−n(√
s2+a2−s)n
√
s2+a2(15.186)
byexpressing Jn(t)as aninfiniteseriesandtransformingtermbyterm. /squaresolid
Integration of Transforms
Again, with F(t)at least piecewise continuous and xlarge enough so that e−xtF(t)de-
creasesexponentially(as x→∞), theintegral
f(x)=integraldisplay∞
0e−xtF(t)dt (15.187)
is uniformly convergent with respect to x. This justifies reversing the order of integration
inthefollowingequation:
integraldisplayb
sf(x)dx=integraldisplayb
sdxintegraldisplay∞
0dte−xtF(t)
=integraldisplay∞
0F(t)
tparenleftbig
e−st−e−btparenrightbig
dt, (15.188)
on integrating with respect to x. The lower limit sis chosen large enough so that f(s)is
withintheregionofuniformconvergence.Nowletting b→∞,weha v e
integraldisplay∞
sf(x)dx=integraldisplay∞
0F(t)
te−stdt=LbraceleftbiggF(t)
tbracerightbigg
, (15.189)
providedthat F(t)/tisfiniteat t=0 ordivergeslessstronglythan t−1(sothat L{F(t)/t}
willexist).
15.10 Other Properties 985
Limits of Integration — Unit Step Function
TheactuallimitsofintegrationfortheLaplacetransformmaybespecifiedwiththe(Heav-
iside)unitstepfunction
u(t−k)=braceleftbigg0,t<k
1,t>k.
Forinstance,
Lbraceleftbig
u(t−k)bracerightbig
=integraldisplay∞
ke−stdt=1
se−ks.
A rectangular pulse of width kand unit height is described by F(t)=u(t)−u(t−k).
TakingtheLaplacetransform,weobtain
Lbraceleftbig
u(t)−u(t−k)bracerightbig
=integraldisplayk
0e−stdt=1
sparenleftbig
1−e−ksparenrightbig
.
The unit step function is also used in Eq. (15.163) and could be invoked in Exer-
cise15.10.13.
Exercises
15.10.1 Solve Eq. (15.154), which describes a damped simple harmonic oscillator for X(0)=
X0,X′(0)=0,and
(a)b2=4 km(criticallydamped),
(b)b2>4 km(overdamped).
ANS.(a)X( t)=X0e−(b/2m)tparenleftbigg
1+b
2mtparenrightbigg
.
15.10.2 SolveEq.(15.154),whichdescribesadampedsimpleharmonicoscillatorfor X(0)=0,
X′(0)=v0, and
(a)b2<4 km(underdamped),
(b)b2=4 km(criticallydamped),
(c)b2>4 km(overdamped).
ANS.(a) X(t)=v0
ω1e−(b/2m)tsinω1t,
(b)X(t)=v0te−(b/2m)t.
15.10.3 Themotionofabodyfallinginaresistingmediummaybedescribedby
md2X(t)
dt2=mg−bdX(t)
dt
986 Chapter 15 Integral Transforms
FIGURE 15.13Ringingcircuit.
whentheretardingforceisproportionaltothevelocity.Find X(t)anddX(t)/dt forthe
initialconditions
X(0)=dX
dtvextendsinglevextendsinglevextendsinglevextendsingle
t=0=0.
15.10.4 Ringing circuit . In certain electronic circuits, resistance, inductance, and capacitance
are placed in the plate circuit in parallel (Fig. 15.13). A constant voltage is maintained
across the parallel elements, keeping the capacitor charged. At time t=0 the circuit
is disconnected from the voltage source. Find the voltages across the parallel elements
R,L,andCasafunctionoftime.Assume Rtobelarge.
Hint.ByKirchhoff’slaws
IR+IC+IL=0 and ER=EC=EL,
where
ER=IRR, E L=LdIL
dt
EC=q0
C+1
Cintegraldisplayt
0ICdt,
q0=initialchargeofcapacitor.
WiththeDCimpedanceof L=0,letIL(0)=I0,EL(0)=0.Thismeans q0=0.
15.10.5 WithJ0(t)expressed as a contour integral, apply the Laplace transform operation, re-
versetheorderofintegration,andthusshowthat
Lbraceleftbig
J0(t)bracerightbig
=parenleftbig
s2+1parenrightbig−1/2,fors>0.
15.10.6 Develop the Laplace transform of Jn(t)fromL{J0(t)}by using the Bessel function
recurrencerelations.
Hint.Hereis achancetousemathematicalinduction.
15.10.7 A calculation of the magnetic field of a circular current loop in circular cylindrical
coordinatesleadstotheintegral
integraldisplay∞
0e−kzkJ1(ka)dk,ℜ(z)≥0.
Showthatthisintegralisequalto a/(z2+a2)3/2.
15.10 Other Properties 987
15.10.8 The electrostatic potential of a point charge qat the origin in circular cylindrical coor-
dinatesis
q
4πε0integraldisplay∞
0e−kzJ0(kρ)dk=q
4πε0·1
(ρ2+z2)1/2,ℜ(z)≥0.
FromthisrelationshowthattheFouriercosineandsinetransformsof J0(kρ)are
(a)radicalbiggπ
2Fcbraceleftbig
J0(kρ)bracerightbig
=integraldisplay∞
0J0(kρ)coskζdk=braceleftbiggparenleftbig
ρ2−ζ2parenrightbig−1/2,ρ>ζ,
0,ρ <ζ.
(b)radicalbiggπ
2Fsbraceleftbig
J0(kρ)bracerightbig
=integraldisplay∞
0J0(kρ)sinkζdk=braceleftbigg0,ρ >ζ,parenleftbig
ρ2−ζ2parenrightbig−1/2,ρ<ζ.
Hint.Replace zbyz+iζandtakethelimitas z→0.
15.10.9 Showthat
Lbraceleftbig
I0(at)bracerightbig
=parenleftbig
s2−a2parenrightbig−1/2,s>a.
15.10.10 VerifythefollowingLaplacetransforms:
(a)Lbraceleftbig
j0(at)bracerightbig
=Lbraceleftbiggsinat
atbracerightbigg
=1
acot−1parenleftbiggs
aparenrightbigg
,
(b)Lbraceleftbig
n0(at)bracerightbig
doesnotexist,
(c)Lbraceleftbig
i0(at)bracerightbig
=Lbraceleftbiggsinhat
atbracerightbigg
=1
2alns+a
s−a=1
acoth−1parenleftbiggs
aparenrightbigg
,
(d)Lbraceleftbig
k0(at)bracerightbig
doesnotexist.
15.10.11 DevelopaLaplacetransformsolutionofLaguerre’sequation
tF′′(t)+(1−t)F′(t)+nF(t)=0.
Note that you need a derivative of a transform and a transform of derivatives. Go as far
asyoucanwith n;then(andonlythen)set n=0.
15.10.12 ShowthattheLaplacetransformoftheLaguerrepolynomial Ln(at)is givenby
Lbraceleftbig
Ln(at)bracerightbig
=(s−a)n
sn+1,s>0.
15.10.13 Showthat
Lbraceleftbig
E1(t)bracerightbig
=1
sln(s+1), s > 0,
where
E1(t)=integraldisplay∞
te−τ
τdτ=integraldisplay∞
1e−xt
xdx.
E1(t)is theexponential-integralfunction.
988 Chapter 15 Integral Transforms
15.10.14 (a) FromEq. (15.189)showthat
integraldisplay∞
0f(x)dx=integraldisplay∞
0F(t)
tdt,
providedtheintegralsexist.
(b) Fromtheprecedingresultshowthat
integraldisplay∞
0sint
tdt=π
2,
inagreementwithEqs. (15.122)and(7.56).
15.10.15 (a) Showthat
Lbraceleftbiggsinkt
tbracerightbigg
=cot−1parenleftbiggs
kparenrightbigg
.
(b) Usingthisresult (with k=1), prove that
Lbraceleftbig
si(t)bracerightbig
=−1
stan−1s,
where
si(t)=−integraldisplay∞
tsinx
xdx,thesineintegral .
15.10.16 IfF(t)is periodic (Fig. 15.14) with a period aso thatF(t+a)=F(t)for allt≥0,
showthat
Lbraceleftbig
F(t)bracerightbig
=integraltexta
0e−stF(t)dt
1−e−as,
withtheintegrationnowoveronlythe first period ofF(t).
15.10.17 FindtheLaplacetransformof thesquarewave(period a) definedby
F(t)=braceleftBigg
1,0<t<a
2
0,a
2<t<a.
ANS.f(s)=1
s·1−e−as/2
1−e−as.
FIGURE 15.14Periodicfunction.
15.10 Other Properties 989
15.10.18 Showthat
(a)L{coshatcosat}=s3
s4+4a4,(c)L{sinhatcosat}=as2−2a3
s4+4a4,
(b)L{coshatsinat}=as2+2a3
s4+4a4,(d)L{sinhatsinat}=2a2s
s4+4a4.
15.10.19 Showthat
(a)L−1braceleftbigparenleftbig
s2+a2parenrightbig−2bracerightbig
=1
2a3sinat−1
2a2tcosat,
(b)L−1braceleftbig
sparenleftbig
s2+a2parenrightbig−2bracerightbig
=1
2atsinat,
(c)L−1braceleftbig
s2parenleftbig
s2+a2parenrightbig−2bracerightbig
=1
2asinat+1
2tcosat,
(d)L−1braceleftbig
s3parenleftbig
s2+a2parenrightbig−2bracerightbig
=cosat−a
2tsinat.
15.10.20 Showthat
Lbraceleftbigparenleftbig
t2−k2parenrightbig−1/2u(t−k)bracerightbig
=K0(ks).
Hint. Try transforming an integral representation of K0(ks)into the Laplace transform
integral.
15.10.21 TheLaplacetransform
integraldisplay∞
0e−xsxJ0(x)dx=s
(s2+1)3/2
mayberewrittenas
1
s2integraldisplay∞
0e−yyJ0parenleftbiggy
sparenrightbigg
dy=s
(s2+1)3/2,
whichisinGauss–Laguerrequadratureform.Evaluatethisintegralfor s=1.0,0.9,0.8,
...,decreasing sin steps of 0.1 until the relative error rises to 10 percent. (The effect
ofdecreasing sistomaketheintegrandoscillatemorerapidlyperunitlengthof y,thus
decreasingtheaccuracyof thenumericalquadrature.)
15.10.22 (a) Evaluate
integraldisplay∞
0e−kzkJ1(ka)dk
bytheGauss–Laguerrequadrature.Take a=1 andz=0.1(0.1)1.0.
(b) From the analytic form, Exercise 15.10.7, calculate the absolute error and the rel-
ativeerror.
990 Chapter 15 Integral Transforms
15.11 C ONVOLUTION (FALTUNGS )THEOREM
One of the most important properties of the Laplace transform is that given by the convo-
lution,or Faltungs,theorem.18Wetaketwotransforms,
f1(s)=Lbraceleftbig
F1(t)bracerightbig
andf2(s)=Lbraceleftbig
F2(t)bracerightbig
, (15.190)
and multiplythem together. To avoid complicationswhen changingvariables, we hold the
upperlimitsfinite:
f1(s)f2(s)=lima→∞integraldisplaya
0e−sxF1(x)dxintegraldisplaya−x
0e−syF2(y)dy. (15.191)
The upper limits are chosen so that the area of integration, shown in Fig. 15.15a, is the
shaded triangle, not the square. If we integrate over a square in the xy-plane, we have
a parallelogram in the tz-plane, which simply adds complications. This modification is
permissiblebecausethetwointegrandsareassumedtodecreaseexponentially.Inthelimit
a→∞, the integral over the unshaded triangle will give zero contribution. Substituting
x=t−z,y=z,theregionofintegrationismappedintothetriangleshowninFig.15.15b.
To verify the mapping, map the vertices: t=x+y,z=y. Using Jacobians to transform
theelementof area,wehave
dxdy=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂t∂y
∂t
∂x
∂z∂y
∂zvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingledtdz=vextendsinglevextendsinglevextendsinglevextendsingle10
−11vextendsinglevextendsinglevextendsinglevextendsingledtdz (15.192)
ordxdy=dtdz.Withthis substitutionEq.(15.191)becomes
f1(s)f2(s)=lima→∞integraldisplaya
0e−stintegraldisplayt
0F1(t−z)F2(z)dzdt
=Lbraceleftbiggintegraldisplayt
0F1(t−z)F2(z)dzbracerightbigg
. (15.193)
a
b
FIGURE 15.15
Changeofvariables,
(a)xy-plane(b) zt-plane.
18An alternatederivation employs the Bromwich integral (Section 15.12). This is Exercise 15.12.3.
15.11 Convolution (Faltungs) Theorem 991
Forconveniencethisintegralisrepresentedbythesymbol
integraldisplayt
0F1(t−z)F2(z)dz≡F1∗F2 (15.194)
and referred to as the convolution , closely analogous to the Fourier convolution (Sec-
tion15.5). If wesubstitute w=t−z, wefind
F1∗F2=F2∗F1, (15.195)
showingthattherelationis symmetric.
Carryingouttheinversetransform, wealsofind
L−1braceleftbig
f1(s)f2(s)bracerightbig
=integraldisplayt
0F1(t−z)F2(z)dz. (15.196)
This can be useful in the development of new transforms or as an alternative to a partial
fractionexpansion.Oneimmediateapplicationisinthesolutionofintegralequations(Sec-
tion16.2).Sincetheupperlimit, t,isvariable,thisLaplaceconvolutionisusefulintreating
Volterraintegralequations.TheFourierconvolutionwithfixed(infinite)limitswouldapply
toFredholmintegralequations.
Example 15.11.1 DRIVEN OSCILLATOR WITH DAMPING
As one illustration of the use of the convolution theorem, let us return to the mass mon
a spring, with damping and a driving force F(t). The equation of motion ((15.129) or
(15.154))nowbecomes
mX′′(t)+bX′(t)+kX(t)=F(t). (15.197)
Initial conditions X(0)=0,X′(0)=0 are used to simplify this illustration, and the trans-
formedequationis
ms2x(s)+bsx(s)+kx(s)=f(s), (15.198)
or
x(s)=f(s)
m1
(s+b/2m)2+ω2
1, (15.199)
whereω2
1≡k/m−b2/4m2, asbefore.
Bytheconvolutiontheorem(Eq. (15.193)or(15.196)),
X(t)=1
mω1integraldisplayt
0F(t−z)e−(b/2m)zsinω1zdz. (15.200)
If theforce isimpulsive, F(t)=Pδ(t),19
X(t)=P
mω1e−(b/2m)tsinω1t. (15.201)
19Notethat δ(t)liesinsidethe interval[0,t].
992 Chapter 15 Integral Transforms
Prepresents the momentum transferred by the impulse, and the constant P/mtakes the
placeofaninitialvelocity X′(0).
IfF(t)=F0sinωt,Eq.(15.200)maybeused,butapartialfractionexpansionisperhaps
moreconvenient.With
f(s)=F0ω
s2+ω2
Eq.(15.199)becomes
x(s)=F0ω
m·1
s2+ω2·1
(s+b/2m)2+ω2
1
=F0ω
mbracketleftbigga′s+b′
s2+ω2+c′s+d′
(s+b/2m)2+ω2
1bracketrightbigg
. (15.202)
Thecoefficients a′,b′,c′, andd′areindependentof s. Directcalculationshows
−1
a′=b
mω2+m
bparenleftbig
ω2
0−ω2parenrightbig2,
−1
b′=−m
bparenleftbig
ω2
0−ω2parenrightbigbracketleftbiggb
mω2+m
bparenleftbig
ω2
0−ω2parenrightbig2bracketrightbigg
.
Sincec′andd′will lead to exponentially decreasing terms (transients), they will be dis-
cardedhere.Carryingouttheinverseoperation,wefindforthesteady-statesolution
X(t)=F0
[b2ω2+m2(ω2
0−ω2)2]1/2sin(ωt−ϕ), (15.203)
where
tanϕ=bω
m(ω2
0−ω2).
Differentiatingthedenominator,wefindthattheamplitudehasamaximumwhen
ω2=ω2
0−b2
2m2=ω2
1−b2
4m2. (15.204)
This is the resonance condition.20At resonance the amplitude becomes F0/bω1, showing
that the mass mgoes into infinite oscillationat resonance if dampingis neglected (b=0).
Itis worthnotingthatwehavehadthreedifferentcharacteristicfrequencies:
ω2
2=ω2
0−b2
2m2,
resonancefor forcedoscillations,withdamping;
ω2
1=ω2
0−b2
4m2,
20Theamplitude (squared) has thetypical resonance denominator, the Lorentz line shape,Exercise 15.3.9.
15.11 Convolution (Faltungs) Theorem 993
freeoscillationfrequency,withdamping;and
ω2
0=k
m,
freeoscillationfrequency,nodamping.Theycoincideonlyif thedampingis zero. /squaresolid
Returning to Eqs. (15.197) and (15.199), Eq. (15.197) is our ODE for the response of
adynamicalsystemtoanarbitrarydrivingforce.Thefinalresponseclearlydependsonboth
the driving force and the characteristics of our system. This dual dependence is separated
in the transform space. In Eq. (15.199) the transform of the response (output) appears as
theproductoftwofactors,onedescribingthedrivingforce(input)andtheotherdescribing
the dynamical system. This latter part, which modifies the input and yields the output, is
oftencalleda transferfunction .Specifically,[(s+b/2m)2+ω2
1]−1isthetransferfunction
correspondingtothisdampedoscillator.Theconceptofatransferfunctionisofgreatusein
thefieldofservomechanisms.Oftenthecharacteristicsofaparticularservomechanismare
described by giving its transfer function. The convolution theorem then yields the output
signalfor aparticularinputsignal.
Exercises
15.11.1 Fromtheconvolutiontheoremshowthat
1
sf(s)=Lbraceleftbiggintegraldisplayt
0F(x)dxbracerightbigg
,
wheref(s)=L{F(t)}.
15.11.2 IfF(t)=taandG(t)=tb,a>−1,b>−1:
(a) Showthattheconvolution
F∗G=ta+b+1integraldisplay1
0ya(1−y)bdy.
(b) Byusingtheconvolutiontheorem,showthat
integraldisplay1
0ya(1−y)bdy=a!b!
(a+b+1)!.
Thisis theEulerformulafor thebetafunction(Eq. (8.59a)).
15.11.3 Usingtheconvolutionintegral,calculate
L−1braceleftbiggs
(s2+a2)(s2+b2)bracerightbigg
,a2/negationslash=b2.
15.11.4 Anundampedoscillatorisdrivenbyaforce F0sinωt.Findthedisplacementasafunc-
tion of time. Notice that it is a linear combination of two simple harmonic motions,
one with the frequency of the driving force and one with the frequency ω0of the free
oscillator.(Assume X(0)=X′(0)=0.)
ANS.X(t)=F0/m
ω2−ω2
0parenleftbiggω
ω0sinω0t−sinωtparenrightbigg
.
994 Chapter 15 Integral Transforms
OtherexercisesinvolvingtheLaplaceconvolutionappearinSection16.2.
15.12 I NVERSE LAPLACE TRANSFORM
Bromwich Integral
We now develop an expression for the inverse Laplace transform L−1appearing in the
equation
F(t)=L−1braceleftbig
f(s)bracerightbig
. (15.205)
One approach lies in the Fourier transform, for which we know the inverse relation. There
is a difficulty, however. Our Fourier transformable function had to satisfy the Dirichlet
conditions.Inparticular,werequiredthat
limω→∞G(ω)=0 (15.206)
so that the infinite integral would be well defined.21Now we wish to treat functions F(t)
that may diverge exponentially. To surmount this difficulty, we extract an exponential fac-
tor,eγt, fromour(possibly) divergentLaplacefunctionandwrite
F(t)=eγtG(t). (15.207)
IfF(t)divergesas eαt,werequire γtobegreaterthan αsothatG(t)willbeconvergent .
Now,with G(t)=0fort<0andotherwisesuitablyrestrictedsothatitmayberepresented
byaFourierintegral(Eq. (15.20)),
G(t)=1
2πintegraldisplay∞
−∞eiutduintegraldisplay∞
0G(v)e−iuvdv. (15.208)
UsingEq. (15.207),wemayrewrite(15.208)as
F(t)=eγt
2πintegraldisplay∞
−∞eiutduintegraldisplay∞
0F(v)e−γve−iuvdv. (15.209)
Now,withthechangeofvariable,
s=γ+iu, (15.210)
the integral over visthrownintotheformof aLaplacetransform,
integraldisplay∞
0f(v)e−svdv=f(s); (15.211)
21If deltafunctions areincluded, G(ω)may be acosine. Although this does not satisfy Eq.(15.206), G(ω)is still bounded.
15.12 Inverse Laplace Transform 995
FIGURE 15.16Singularities
ofestf(s).
sis now a complex variable, and ℜ(s)≥γto guarantee convergence. Notice that the
Laplace transform has mapped a function specified on the positive real axis onto the com-
plexplane,ℜ(s)≥γ.22
Withγas aconstant, ds=idu.SubstitutingEq. (15.211)intoEq.(15.209), weobtain
F(t)=1
2πiintegraldisplayγ+i∞
γ−i∞estf(s)ds. (15.212)
Here is our inverse transform . We have rotated the line of integration through 90◦(by
usingds=idu). The path has become an infinite vertical line in the complex plane, the
constantγhavingbeenchosensothatallthesingularitiesof f(s)areontheleft-handside
(Fig.15.16).
Equation (15.212), our inverse transformation, is usually known as the Bromwich in-
tegral, although sometimes it is referred to as the Fourier–Mellin theorem orFourier–
Mellin integral . This integral may now be evaluated by the regular methods of contour
integration(Chapter7).If t>0,thecontourmaybeclosedbyaninfinitesemicircleinthe
lefthalf-plane.Thenbytheresiduetheorem(Section7.1)
F(t)=/Sigma1(residuesincludedfor ℜ(s)<γ). (15.213)
Possibly this means of evaluation with ℜ(s)ranging through negative values seems para-
doxical in view of our previous requirement that ℜ(s)≥γ. The paradox disappears when
we recall that the requirement ℜ(s)≥γwas imposed to guarantee convergence of the
Laplace transform integral that defined f(s). Oncef(s)is obtained, we may then pro-
ceed to exploit its properties as an analytical function in the complex plane wherever we
choose.23In effect we are employing analytic continuation to get L{F(t)}in the left half-
plane, exactly as the recurrence relation for the factorial function was used to extend the
Eulerintegraldefinition(Eq. (8.5)) tothelefthalf-plane.
PerhapsapairofexamplesmayclarifytheevaluationofEq. (15.212).
22For a derivation of the inverse Laplace transform using only real variables, see C. L. Bohn and R. W. Flynn, Real variable
inversion of Laplacetransforms: Anapplication in plasma physics. Am.J.Phys. 46: 1250 (1978).
23In numerical work f(s)may well be available only for discrete real, positive values of s. Then numerical procedures are
indicated. SeeKrylov and Skoblya in the Additional Reading.
996 Chapter 15 Integral Transforms
Example 15.12.1 INVERSION VIA CALCULUS OF RESIDUES
Iff(s)=a/(s2−a2), then
estf(s)=aest
s2−a2=aest
(s+a)(s−a). (15.214)
TheresiduesmaybefoundbyusingExercise6.6.1orvariousothermeans.Thefirststepis
to identify the singularities, the poles. Here we have one simple pole at s=aand another
simple pole at s=−a. By Exercise 6.6.1, the residue at s=ais(1
2)eatand the residue at
s=−ais(−1
2)e−at. Then
Residues=parenleftbig1
2parenrightbigparenleftbig
eat−e−atparenrightbig
=sinhat=F(t), (15.215)
inagreementwithEq.(15.105). /squaresolid
Example 15.12.2
If
f(s)=1−e−as
s,
thenes(t−a)grows exponentially for t<aon the semicircle in the left-hand s-plane, so
contour integration and the residue theorem are not applicable. However, we can evaluate
theintegralexplicitlyas follows.Welet γ→0 andsubstitute s=iy,so
F(t)=1
2πiintegraldisplayγ+i∞
γ−i∞estf(s)=1
2πintegraldisplay∞
−∞bracketleftbig
eiyt−eiy(t−a)bracketrightbigdy
y. (15.216)
UsingtheEuleridentity,onlythesinessurvivethatareoddin yandweobtain
F(t)=1
πintegraldisplay∞
−∞bracketleftbiggsinty
y−sin(t−a)y
ybracketrightbigg
. (15.217)
Ifk>0,thenintegraltext∞
0sinky
ydygivesπ/2, and it gives −π/2i fk<0.As a consequence,
F(t)=0i ft>a>0 and if t<0.If 0<t<a, thenF(t)=1.This can be written
compactlyintermsoftheHeavisideunitstepfunction u(t)asfollows:
F(t)=u(t)−u(t−a)=
0,t<0,
1,0<t<a,
0,t>a,(15.218)
astepfunctionofunitheightandlength a(Fig.15.17). /squaresolid
Twogeneralcommentsmaybeinorder.First,thesetwoexampleshardlybegintoshow
the usefulness and power of the Bromwich integral. It is always available for inverting a
complicatedtransform whenthetablesproveinadequate.
Second, this derivation is not presented as a rigorous one. Rather, it is given more as
a plausibility argument, although it can be made rigorous. The determination of the in-
verse transform is somewhat similar to the solution of a differential equation. It makes
15.12 Inverse Laplace Transform 997
FIGURE 15.17
Finite-lengthstepfunction
u(t)−u(t−a).
little difference how you get the solution. Guess at it if you want. The solution can al-
ways be checked by substitution back into the original differential equation. Similarly,
F(t)can(and,tocheckoncarelesserrors,should)becheckedbydeterminingwhether,by
Eq.(15.99),
Lbraceleftbig
F(t)bracerightbig
=f(s).
Two alternate derivations of the Bromwich integral are the subjects of Exercises 15.12.1
and15.12.2.
As a final illustration of the use of the Laplace inverse transform, we have some results
fromtheworkofBrillouinandSommerfeld(1914)inelectromagnetictheory.
Example 15.12.3 VELOCITY OF ELECTROMAGNETIC WAVES IN A DISPERSIVE MEDIUM
Thegroupvelocity uof travelingwavesisrelatedtothephasevelocity vbytheequation
u=v−λdv
dλ. (15.219)
Hereλis the wavelength. In the vicinity of an absorption line (resonance), dv/dλmay be
sufficiently negative so that u>c(Fig. 15.18). The question immediately arises whether
a signal can be transmitted faster than c, the velocity of light in vacuum. This question,
which assumes that such a group velocity is meaningful, is of fundamental importance to
thetheoryofspecialrelativity.
We needasolutiontothewaveequation
∂2ψ
∂x2=1
v2∂2ψ
∂t2, (15.220)
correspondingtoaharmonicvibrationstartingattheoriginattimezero.Sinceourmedium
isdispersive, visafunctionoftheangularfrequency.Imagine,forinstance,aplanewave,
angular frequency ω, incident on a shutter at the origin. At t=0 the shutter is (instanta-
neously)opened,andthewaveis permittedtoadvancealongthepositive x-axis.
998 Chapter 15 Integral Transforms
FIGURE 15.18Opticaldispersion.
Let us then build up a solution starting at x=0. It is convenient to use the Cauchy
integralformula,Eq.(6.43),
ψ(0,t)=1
2πicontintegraldisplaye−izt
z−z0dz=e−iz0t
(for a contour encircling z=z0in the positive sense). Using s=−izandz0=ω,w e
obtain
ψ(0,t)=1
2πiintegraldisplayγ+i∞
γ−i∞est
s+iωds=braceleftbigg0,t <0,
e−iωt,t>0.(15.221)
To be complete, the loop integral is along the vertical line ℜ(s)=γandan infinite semi-
circle, as shown in Fig. 15.19. The location of the infinite semicircle is chosen so that the
integral over it vanishes. This means a semicircle in the left half-plane for t>0 and the
residue is enclosed. For t<0 we pick the right half-plane and no singularity is enclosed.
Thefactthatthisisjust theBromwichintegralmaybeverifiedbynotingthat
F(t)=braceleftbigg0,t <0,
e−iωt,t>0(15.222)
FIGURE 15.19Possibleclosedcontours.
15.12 Inverse Laplace Transform 999
andapplyingtheLaplacetransform.Thetransformedfunction f(s)becomes
f(s)=1
s+iω. (15.223)
Our Cauchy–Bromwich integral provides us with the time dependence of a signal leav-
ingtheoriginat t=0.Toincludethespacedependence,wenotethat
es(t−x/v)
satisfiesthewaveequation.Withthisasaclue,wereplace tbyt−x/vandwriteasolution:
ψ(x,t)=1
2πiintegraldisplayγ+i∞
γ−i∞es(t−x/v)
s+iωds. (15.224)
It was seen in the derivation of the Bromwich integral that our variable sreplaces the ω
oftheFouriertransformation.Hencethewavevelocity vmaybecomeafunctionof s,that
is,v(s).Itsparticularformneednotconcernushere.Weneedonlytheproperty v≤cand
lim
|s|→∞v(s)=constant,c . (15.225)
ThisissuggestedbytheasymptoticbehaviorofthecurveontherightsideofFig.15.18.24
EvaluatingEq.(15.225)bythecalculusofresidues,wemayclosethepathofintegration
byasemicircleintherighthalf-plane,provided
t−x
c<0.
Hence
ψ(x,t)=0,t−x
c<0, (15.226)
which means that the velocity of our signal cannot exceed the velocity of light in the vac-
uum,c.Thissimplebutverysignificantresultwas extendedbySommerfeldandBrillouin
toshowjusthowthewaveadvancedinthedispersivemedium. /squaresolid
Summary — Inversion of Laplace Transform
•Direct use of tables, Table 15.2, and references; use of partial fractions (Section 15.8)
andtheoperationaltheoremsofTable15.1.
•Bromwichintegral,Eq. (15.212),andthecalculusofresidues.
•Numericalinversion,seetheAdditionalReadings.
24Equation (15.225) follows rigorously from the theory of anomalous dispersion. See also the Kronig–Kramers optical disper-
sion relations of Section 7.2.
1000 Chapter 15 Integral Transforms
Table 15.1 LaplaceTransformOperations
Operations Equation
1. Laplacetransform f(s)=L{F(t)}=integraldisplay∞
0e−stF(t)dt (15.99)
2. Transform of derivative sf(s)−F(+0)=L{F′(t)} (15.123)
s2f(s)−sF(+0)−F′(+0)=L{F′′(t)}(15.124)
3. Transform of integral1
sf(s)=Lbraceleftbiggintegraldisplayt
0F(x)dxbracerightbigg
(Exercise 15.11.1)
4. Substitution f(s−a)=L{eatF(t)} (15.152)
5. Translation e−bsf(s)=L{F(t−b)} (15.164)
6. Derivativeof transform f(n)(s)=L{(−t)nF(t)} (15.173)
7. Integral of transformintegraldisplay∞
sf(x)dx=LbraceleftbiggF(t)
tbracerightbigg
(15.189)
8. Convolution f1(s)f2(s)=Lbraceleftbiggintegraldisplayt
0F1(t−z)F2(z)dzbracerightbigg
(15.193)
9. Inverse transform, Bromwich integral1
2πiintegraldisplayγ+i∞
γ−i∞estf(s)ds=F(t) (15.212)
Exercises
15.12.1 DerivetheBromwichintegralfromCauchy’sintegralformula.
Hint.Applytheinversetransform L−1to
f(s)=1
2πilimα→∞integraldisplayγ+iα
γ−iαf(z)
s−zdz,
wheref(z)is analyticforℜ(z)≥γ.
15.12.2 Startingwith
1
2πiintegraldisplayγ+i∞
γ−i∞estf(s)ds,
showthatbyintroducing
f(s)=integraldisplay∞
0e−szF(z)dz,
we can convert one integral into the Fourier representation of a Dirac delta function.
FromthisderivetheinverseLaplacetransform.
15.12.3 Derive the Laplace transformation convolution theorem by use of the Bromwich inte-
gral.
15.12.4 Find
L−1braceleftbiggs
s2−k2bracerightbigg
(a) byapartialfractionexpansion.
(b) Repeat,usingtheBromwichintegral.
15.12 Inverse Laplace Transform 1001
Table 15.2 LaplaceTransforms
f(s) F(t) Limitation Equation
1. 1 δ(t) Singularity at+0 (15.141)
2.1
s1 s>0 (15.102)
3.n!
sn+1tns>0 (15.108)
n>−1
4.1
s−kekts>k (15.103)
5.1
(s−k)2tekts>k (15.175)
6.s
s2−k2coshkt s>k (15.105)
7.k
s2−k2sinhkt s>k (15.105)
8.s
s2+k2coskt s> 0 (15.107)
9.k
s2+k2sinkt s> 0 (15.107)
10.s−a
(s−a)2+k2eatcoskt s>a (15.153)
11.k
(s−a)2+k2eatsinkt s>a (15.153)
12.s2−k2
(s2+k2)2tcoskt s> 0 (Exercise 15.10.19)
13.2ks
(s2+k2)2tsinkt s> 0 (Exercise 15.10.19)
14.(s2+a2)−1/2J0(at) s > 0 (15.185)
15.(s2−a2)−1/2I0(at) s >a (Exercise 15.10.9)
16.1
acot−1parenleftbiggs
aparenrightbigg
j0(at) s > 0 (Exercise 15.10.10)
17.1
2alns+a
s−a
1
acoth−1parenleftbiggs
aparenrightbigg
i0(at) s >a (Exercise 15.10.10)
18.(s−a)n
sn+1Ln(at) s > 0 (Exercise 15.10.12)
19.1
sln(s+1)E 1(x)=−Ei(−x) s> 0 (Exercise 15.10.13)
20.lns
s−lnt−γs >0 (Exercise 15.12.9)
Amore extensivetableof LaplacetransformsappearsinChapter 29ofAMS-55(see footnote4 inChapter5 forthereference).
15.12.5 Find
L−1braceleftbiggk2
s(s2+k2)bracerightbigg
1002 Chapter 15 Integral Transforms
(a) byusingapartialfractionexpansion.
(b) Repeatusingtheconvolutiontheorem.
(c) RepeatusingtheBromwichintegral.
ANS.F(t)=1−coskt.
15.12.6 Use the Bromwich integral to find the function whose transform is f(s)=s−1/2.N o t e
thatf(s)hasabranchpointat s=0.Thenegative x-axismaybetakenasa cutline.
ANS.F(t)=(πt)−1/2.
15.12.7 Showthat
L−1braceleftbigparenleftbig
s2+1parenrightbig−1/2bracerightbig
=J0(t)
byevaluationof theBromwichintegral.
Hint. Convert your Bromwich integral into an integral representation of J0(t).F i g -
ure15.20showsapossiblecontour.
15.12.8 EvaluatetheinverseLaplacetransform
L−1braceleftbigparenleftbig
s2−a2parenrightbig−1/2bracerightbig
byeachofthefollowingmethods:
(a) Expansioninaseries andterm-by-terminversion.
(b) DirectevaluationoftheBromwichintegral.
(c) ChangeofvariableintheBromwichintegral: s=(a/2)(z+z−1).
FIGURE 15.20Apossible
contourfor theinversionof
J0(t).
15.12 Additional Readings 1003
15.12.9 Showthat
L−1braceleftbigglns
sbracerightbigg
=−lnt−γ,
whereγ=0.5772...,theEuler–Mascheroniconstant.
15.12.10 EvaluatetheBromwichintegralfor
f(s)=s
(s2+a2)2.
15.12.11 Heavisideexpansiontheorem .If thetransform f(s)maybewrittenas aratio
f(s)=g(s)
h(s),
whereg(s)andh(s)areanalyticfunctions, h(s)havingsimple,isolatedzerosat s=si,
showthat
F(t)=L−1braceleftbiggg(s)
h(s)bracerightbigg
=summationdisplay
ig(si)
h′(si)esit.
Hint.SeeExercise6.6.2.
15.12.12 Using the Bromwich integral, invert f(s)=s−2e−ks. Express F(t)=L−1{f(s)}in
termsofthe(shifted) unitstepfunction u(t−k).
ANS.F(t)=(t−k)u(t−k).
15.12.13 Youhavea Laplacetransform:
f(s)=1
(s+a)(s+b),a/negationslash=b.
Invertthistransform byeachof threemethods:
(a) Partialfractionsanduse oftables.
(b) Convolutiontheorem.
(c) Bromwichintegral.
ANS.F(t)=e−bt−e−at
a−b,a/negationslash=b.
AdditionalReadings
Champeney, D. C., Fourier Transforms and Their Physical Applications. New York: Academic Press (1973).
Fourier transforms are developed in a careful, easy-to-follow manner. Approximately 60% of the book is
devoted to applications of interest in physics and engineering.
Erdelyi, A.,W. Magnus, F. Oberhettinger, and F. G. Tricomi, Tables of Integral Transforms ,2v o l s .N e wY o r k :
McGraw–Hill (1954). This text contains extensive tables of Fourier sine, cosine, and exponential transforms,
Laplace and inverse Laplace transforms, Mellin and inverse Mellin transforms, Hankel transforms, and other,
more specializedintegral transforms.
1004 Chapter 15 Integral Transforms
Hanna, J. R., Fourier Series and Integrals of Boundary Value Problems . Somerset, NJ: Wiley (1990). This book
is a broad treatment of the Fourier solution of boundary value problems. The concepts of convergence and
completeness aregiven careful attention.
Jeffreys,H.,andB.S.Jeffreys, MethodsofMathematicalPhysics ,3rded.Cambridge,UK:CambridgeUniversity
Press (1972).
Krylov, V. I., and N. S. Skoblya, Handbook of Numerical Inversion of Laplace Transform. Jerusalem: Israel
Program for ScientificTranslations (1969).
Lepage, W. R., Complex Variables and the Laplace Transform for Engineers . New York: McGraw-Hill (1961);
New York: Dover (1980). A complex variable analysis that is carefully developed and then applied to Fourier
and Laplacetransforms. It is written to be readby students, but intended for the serious student.
McCollum, P. A., and B. F. Brown, Laplace Transform Tables and Theorems . New York: Holt, Rinehart and
Winston (1965).
Miles, J. W., Integral Transforms in Applied Mathematics . Cambridge, UK: Cambridge University Press (1971).
Thisisabriefbutinterestingandusefultreatmentfortheadvancedundergraduate.Itemphasizesapplications
ratherthan abstractmathematical theory.
Papoulis, A., The Fourier Integral and Its Applications . New York: McGraw-Hill (1962). This is a rigorous
development ofFourier and Laplacetransforms and has extensive applications in scienceand engineering.
Roberts, G.E.,andH.Kaufman, Table of Laplace Transforms . Philadelphia: Saunders (1966).
Sneddon, I. N., Fourier Transforms . New York: McGraw-Hill (1951), reprinted, Dover (1995). A detailed com-
prehensive treatment, this book is loaded with applications to a wide variety of fields of modern and classical
physics.
Sneddon, I. H., The Useof Integral Transforms . NewYork: McGraw-Hill(1972). Written for students in science
and engineering in terms they can understand, this book covers all the integral transforms mentioned in this
chapteras wellasinseveral others. Manyapplications areincluded.
Van der Pol, B., and H. Bremmer, Operational Calculus Based on the Two-sided Laplace Integral , 3rd ed. Cam-
bridge, UK: Cambridge University Press (1987). Here is a development based on the integral range −∞to
+∞, rather than the useful 0 to ∞. Chapter V contains a detailed study of the Dirac delta function (impulse
function).
W o l f ,K .B . , Integral Transforms in Science and Engineering . New York: Plenum Press (1979). This book is a
very comprehensive treatment of integral transforms and their applications.
CHAPTER 16
INTEGRAL EQUATIONS
16.1 I NTRODUCTION
Withtheexceptionoftheintegraltransformsofthelastchapter,wehavebeenconsidering
relations between the unknown function ϕ(x)and one or more of its derivatives. We now
proceed to investigate equations containing the unknown function within an integral. As
withdifferentialequations,weshallconfineourattentiontolinearrelations,linearintegral
equations.Integralequationsareclassifiedintwoways:
•Ifthelimitsofintegrationarefixed ,wecalltheequationa Fredholm equation;if one
limitis variable ,itisaVolterra equation.
•Iftheunknownfunction appearsonlyundertheintegral sign,welabelit firstkind .
If itappearsboth insideandoutside theintegral,itis labeled secondkind .
Definitions
Symbolically,wehavea Fredholmequationofthefirst kind ,
f(x)=integraldisplayb
aK(x,t)ϕ(t)dt; (16.1)
theFredholmequationofthesecondkind ,withλbeingtheeigenvalue,
ϕ(x)=f(x)+λintegraldisplayb
aK(x,t)ϕ(t)dt; (16.2)
theVolterraequationofthefirst kind ,
f(x)=integraldisplayx
aK(x,t)ϕ(t)dt; (16.3)
1005
1006 Chapter 16 Integral Equations
andtheVolterraequationof thesecondkind ,
ϕ(x)=f(x)+integraldisplayx
aK(x,t)ϕ(t)dt. (16.4)
In all four cases ϕ(t)is the unknown function. K(x,t), which we call the kernel, and
f(x)areassumedtobeknown.When f(x)=0,theequationis saidtobe homogeneous .
Why do we bother about integral equations? After all, the differential equations have
done a rather good job of describing our physical world so far. There are several reasons
forintroducingintegralequationshere.
We have placed considerable emphasis on the solution of differential equations subject
to particular boundary conditions . For instance, the boundary condition at r=0 deter-
mines whether the Neumann function Nn(r)is present when Bessel’s equation is solved.
The boundary condition for r→∞determines whether the In(r)is present in our solu-
tion of the modified Bessel equation. The integral equation relates the unknown function
notonly toits valuesat neighboringpoints(derivatives) but alsoto its values throughouta
region, including the boundary. In a very real sense the boundary conditions are built into
theintegralequationratherthanimposedatthefinalstageofthesolution.Itcanbeseenin
Section10.5,wherekernelsareconstructed,thattheformofthekerneldependsontheval-
uesontheboundary.Theintegralequation,then,iscompactandmayturnouttobeamore
convenientorpowerfulformthanthedifferentialequation.Mathematicalproblemssuchas
existence, uniqueness, and completeness may often be handled more easily and elegantly
in integral form. Finally, whether or not we like it, there are some problems, such as some
diffusion and transport phenomena, that cannot be represented by differential equations.
If we wish to solve such problems, we are forced to handle integral equations. Finally, an
integralequationmayalsoappearasamatterofdeliberatechoicebasedonconvenienceor
theneedforthemathematicalpowerof anintegralequationformulation.
Example 16.1.1 MOMENTUM REPRESENTATION IN QUANTUM MECHANICS
TheSchrödingerequation(inordinaryspacerepresentation)is
−¯h2
2m∇2ψ(r)+V(r)ψ(r)=Eψ(r), (16.5)
or
parenleftbig
∇2+a2parenrightbig
ψ(r)=v(r)ψ(r), (16.6)
where
a2=2m
¯h2E, v( r)=2m
¯h2V(r). (16.7)
If wegeneralizeEq. (16.6)to
parenleftbig
∇2+a2parenrightbig
ψ(r)=integraldisplay
v(r,r′)ψ(r′)d3r′, (16.8)
then,forthespecialcaseof
v(r,r′)=v(r′)δ(r−r′), (16.9)
16.1 Introduction 1007
alocalinteraction,Eq.(16.8)reducestoEq.(16.6).ConsidertheFouriertransformpair ψ
and/Psi1(comparefootnote9inSection15.6):
/Psi1(k)=1
(2π)3/2integraldisplay
ψ(r)e−ik·rd3r, ψ( r)=1
(2π)3/2integraldisplay
/Psi1(k)eik·rd3k,(16.10)
withtheabbreviation pfor momentumso that
p
¯h=k(wavenumber). (16.11)
MultiplyingEq. (16.8)bytheplane-wave e−ik·r, weobtain
integraldisplay
e−ik·rparenleftbig
∇2+a2parenrightbig
ψ(r)d3r=integraldisplay
d3re−ik·rintegraldisplay
v(r,r′)ψ(r′)d3r′. (16.12)
Note that the ∇2on the left operates only on the ψ(r). Integrating the left-hand side by
partsandsubstitutingEq.(16.10) for ψ(r′)ontheright,weget
integraldisplayparenleftbig
−k2+a2parenrightbig
ψ(r)e−ik·rd3r=(2π)3/2parenleftbig
−k2+a2parenrightbig
/Psi1(k)
=1
(2π)3/2integraldisplayintegraldisplayintegraldisplay
v(r,r′)/Psi1(k′)e−i(k·r−k′·r′)d3r′d3rd3k′.(16.13)
If weuse
f(k,k′)=1
(2π)3/2integraldisplayintegraldisplay
v(r,r′)e−i(k·r−k′·r′)d3r′d3r, (16.14)
Eq.(16.13) becomes
parenleftbig
−k2+a2parenrightbig
/Psi1(k)=integraldisplay
f(k,k′)/Psi1(k′)d3k′, (16.15)
a Fredholm equation of the second kind in which the parameter a2corresponds to the
eigenvalue.
Forourspecialbutimportantcaseoflocalinteraction,applicationofEq.(16.9)leadsto
f(k,k′)=f(k−k′). (16.16)
Thisisourmomentumrepresentation,equivalenttoanordinarystaticinteractionpoten-
tialincoordinatespace.Ourmomentumwavefunction /Psi1(k)satisfiestheintegralequation
Eq.(16.15).Itmustbeemphasizedthatallthroughherewehaveassumedthattherequired
Fourier integrals exist. For a harmonic oscillator potential, V(r)=r2, the required inte-
grals would not exist. Equation (16.10) would lead to divergent oscillations and we would
havenoEq.(16.15). /squaresolid
1008 Chapter 16 Integral Equations
Transformation of a Differential Equation into an
Integral Equation
Often we find that we have a choice. The physical problem may be represented by a dif-
ferential or an integral equation. Let us assume that we have the differential equation and
wishtotransformitintoanintegralequation.Startingwitha linearsecond-orderODE
y′′+A(x)y′+B(x)y=g(x) (16.17)
withinitialconditions
y(a)=y0,y′(a)=y′
0,
weintegratetoobtain
y′(x)=−integraldisplayx
aA(t)y′(t)dt−integraldisplayx
aB(t)y(t)dt+integraldisplayx
ag(t)dt+y′
0. (16.18)
Integratingthefirst integralontherightbypartsyields
y′(x)=−Ay(x)−integraldisplayx
a(B−A′)y(t)dt+integraldisplayx
ag(t)dt+A(a)y0+y′
0.(16.19)
Notice how the initial conditions are being absorbed into our new version. Integrating a
secondtime,weobtain
y(x)=−integraldisplayx
aAy dx−integraldisplayx
aduintegraldisplayu
abracketleftbig
B(t)−A′(t)bracketrightbig
y(t)dt
+integraldisplayx
aduintegraldisplayu
ag(t)dt+bracketleftbig
A(a)y0+y′
0bracketrightbig
(x−a)+y0.(16.20)
Totransformthisequationintoaneaterform, weusetherelation
integraldisplayx
aduintegraldisplayu
af(t)dt=integraldisplayx
a(x−t)f(t)dt. (16.21)
This may be verified by differentiating both sides. Since the derivatives are equal, the
original expressions can differ only by a constant. Letting x→a, the constant vanishes
andEq. (16.21)is established.ApplyingittoEq.(16.20), weobtain
y(x)=−integraldisplayx
abraceleftbig
A(t)+(x−t)bracketleftbig
B(t)−A′(t)bracketrightbigbracerightbig
y(t)dt
+integraldisplayx
a(x−t)g(t)dt+bracketleftbig
A(a)y0+y′
0bracketrightbig
(x−a)+y0.(16.22)
If wenowintroducetheabbreviations
K(x,t)=(t−x)bracketleftbig
B(t)−A′(t)bracketrightbig
−A(t),
(16.23)
f(x)=integraldisplayx
a(x−t)g(t)dt+bracketleftbig
A(a)y0+y′
0bracketrightbig
(x−a)+y0,
16.1 Introduction 1009
Eq.(16.22) becomes
y(x)=f(x)+integraldisplayx
aK(x,t)y(t)dt, (16.24)
which is a Volterra equation of the second kind. This reformulation as a Volterra integral
equationoffers certainadvantagesininvestigatingquestionsofexistenceanduniqueness.
Example 16.1.2 LINEAR OSCILLATOR EQUATION
Asanillustration,considerthelinearoscillatorequation
y′′+ω2y=0 (16.25)
with
y(0)=0,y′(0)=1.
Thisyields(comparewithEq. (16.17))
A(x)=0,B(x)=ω2,g(x)=0.
Substituting into Eq. (16.22) (or Eqs. (16.23) and (16.24)), we find that the integral equa-
tionbecomes
y(x)=x+ω2integraldisplayx
0(t−x)y(t)dt. (16.26)
•This integral equation, Eq. (16.26), is equivalent to the original differential equation
plustheinitialconditions.
Acheckshowsthateachformis indeedsatisfiedby y(x)=(1/ω)sinωx. /squaresolid
Let us reconsider the linear oscillator equation (16.25) but now with the boundary con-
ditions
y(0)=0,y(b)=0.
Sincey′(0)is notgiven,wemustmodifytheprocedure.Thefirst integrationgives
y′=−ω2integraldisplayx
0ydx+y′(0). (16.27)
Integratinga secondtimeandagainusingEq. (16.21),wehave
y=−ω2integraldisplayx
0(x−t)y(t)dt+y′(0)x. (16.28)
Toeliminatetheunknown y′(0), wenowimposethecondition y(b)=0.Thisgives
ω2integraldisplayb
0(b−t)y(t)dt=by′(0). (16.29)
1010 Chapter 16 Integral Equations
FIGURE 16.1
SubstitutingthisbackintoEq. (16.28), weobtain
y(x)=−ω2integraldisplayx
0(x−t)y(t)dt+ω2x
bintegraldisplayb
0(b−t)y(t)dt. (16.30)
Nowletusbreaktheinterval [0,b]intotwointervals, [0,x]and[x,b]. Since
x
b(b−t)−(x−t)=t
b(b−x), (16.31)
wefind
y(x)=ω2integraldisplayx
0t
b(b−x)y(t)dt+ω2integraldisplayb
xx
b(b−t)y(t)dt. (16.32)
Finally,if wedefineakernel(Fig. 16.1)
K(x,t)=
t
b(b−x), t<x,
x
b(b−t), t>x,(16.33)
wehave
y(x)=ω2integraldisplayb
0K(x,t)y(t)dt, (16.34)
ahomogeneousFredholmequationof thesecondkind.
Ournewkernel, K(x,t), has someinterestingproperties.
1. It issymmetric, K(x,t)=K(t,x).
2. It iscontinuous,inthesensethat
t
b(b−x)vextendsinglevextendsinglevextendsingle
t=x=x
b(b−t)vextendsinglevextendsinglevextendsingle
t=x.
3. Itsderivativewithrespectto tisdiscontinuous .Astincreasesthroughthepoint t=x,
thereis adiscontinuityof −1i n∂K(x,t)/∂t .
AccordingtothesepropertiesinSection9.7weidentify K(x,t)asaGreen’sfunction.
1. Inthetransformationofalinear,second-orderODEintoanintegralequation,theinitial
orboundaryconditionsplayadecisiverole.Ifwehave initialconditions(onlyoneend
of our interval), the differential equation transforms into a Volterra integral equation.
For the case of the linear oscillator equation with boundary conditions (both ends
16.1 Introduction 1011
of our interval), the differential equation leads to a Fredholm integral equation with
akernelthatwillbeaGreen’sfunction.
2. Note that the reverse transformation (integral equation to differential equation) is not
alwayspossible.Thereexistintegralequationsforwhichnocorrespondingdifferential
equationisknown.
Exercises
16.1.1 Starting with the ODE, integrate twice and derive the Volterra integral equation corre-
spondingto
(a)y′′(x)−y(x)=0;y(0)=0,y′(0)=1.
ANS.y=integraldisplayx
0(x−t)y(t)dt+x.
(b)y′′(x)−y(x)=0;y(0)=1,y′(0)=−1.
ANS.y=integraldisplayx
0(x−t)y(t)dt−x+1.
Checkyourresults withEq. (16.23).
16.1.2 DeriveaFredholmintegralequationcorrespondingto
y′′(x)−y(x)=0,y(1)=1,y(−1)=1,
(a) byintegratingtwice,
(b) byformingtheGreen’sfunction.
ANS.y(x)=1−integraldisplay1
−1K(x,t)y(t)dt ,
K(x,t)=braceleftBigg1
2(1−x)(t+1), x >t,
1
2(1−t)(x+1), x <t.
16.1.3 (a) Starting with the given answers of Exercise 16.1.1, differentiate and recover the
originalODEs andtheboundaryconditions .
(b) Repeatfor Exercise16.1.2.
16.1.4 Thegeneralsecond-orderlinearODEwithconstantcoefficientsis
y′′(x)+a1y′(x)+a2y(x)=0.
Giventheboundaryconditions
y(0)=y(1)=0,
integratetwiceanddeveloptheintegralequation
y(x)=integraldisplay1
0K(x,t)y(t)dt,
1012 Chapter 16 Integral Equations
with
K(x,t)=braceleftBigg
a2t(1−x)+a1(x−1), t <x,
a2x(1−t)+a1x, x<t.
Note that K(x,t)is symmetric and continuous if a1=0. How is this related to self-
adjointnessoftheODE?
16.1.5 Verify thatintegraltextx
aintegraltextx
af(t)dtdx=integraltextx
a(x−t)f(t)dt for allf(t)(for which the integrals
exist).
16.1.6 Givenϕ(x)=x−integraltextx
0(t−x)ϕ(t)dt , solve this integral equation by converting it to an
ODE(plusboundaryconditions)andsolvingtheODE(byinspection).
16.1.7 ShowthatthehomogeneousVolterraequationof thesecondkind
ψ(x)=λintegraldisplayx
0K(x,t)ψ(t)dt
hasnosolution(apartfrom thetrivial ψ=0).
Hint.Develop a Maclaurin expansion of ψ(x). Assume ψ(x)andK(x,t)are differen-
tiablewithrespectto xasneeded.
16.2 I NTEGRAL TRANSFORMS ,GENERATING FUNCTIONS
Analogous to differentiation, linear ODEs are solved in Chapter 9. Analogous to integra-
tion, there is no general method available for solving integral equations. However, certain
special cases may be treated with our integral transforms (Chapter 15). For convenience
theseare listedhere. If
ψ(x)=1√
2πintegraldisplay∞
−∞eixtϕ(t)dt,
then
ϕ(x)=1√
2πintegraldisplay∞
−∞e−ixtψ(t)dt (Fourier). (16.35)
If
ψ(x)=integraldisplay∞
0e−xtϕ(t)dt,
then
ϕ(x)=1
2πiintegraldisplayγ+i∞
γ−i∞extψ(t)dt (Laplace). (16.36)
If
ψ(x)=integraldisplay∞
0tx−1ϕ(t)dt,
16.2 Integral Transforms, Generating Functions 1013
then
ϕ(x)=1
2πiintegraldisplayγ+i∞
γ−i∞x−tψ(t)dt (Mellin). (16.37)
If
ψ(x)=integraldisplay∞
0tϕ(t)Jν(xt)dt,
then
ϕ(x)=integraldisplay∞
0tψ(t)Jν(xt)dt (Hankel). (16.38)
Actually the usefulness of the integral transform technique extends a bit beyond these
fourratherspecializedforms.
Example 16.2.1 FOURIER TRANSFORM SOLUTION
Let us consider a Fredholm equation of the first kind with a kernel of the general type
k(x−t),
f(x)=integraldisplay∞
−∞k(x−t)ϕ(t)dt, (16.39)
in which ϕ(t)is our unknown function. Assuming that the needed transforms exist ,w e
applytheFourierconvolutiontheorem(Section15.5)toobtain
f(x)=integraldisplay∞
−∞K(ω)/Phi1(ω)e−iωxdω. (16.40)
Thefunctions K(ω),/Phi1(ω) ,andF(ω)aretheFouriertransformsof k(x),ϕ(x) ,andf(x),
respectively. Taking the Fourier transform of both sides of Eq. (16.40), by Eq. (16.35) we
have
K(ω)/Phi1(ω)=1
2πintegraldisplay∞
−∞f(x)eiωxdx=F(ω)√
2π. (16.41)
Then
/Phi1(ω)=1√
2π·F(ω)
K(ω), (16.42)
and,usingtheinverseFouriertransform, wehave
ϕ(x)=1
2πintegraldisplay∞
−∞F(ω)
K(ω)e−iωxdω. (16.43)
For a rigorous justification of this result one can follow Morse and Feshbach (see the
Additional Readings) (1953) across complex planes. An extension of this transformation
solutionappearsasExercise16.2.1. /squaresolid
1014 Chapter 16 Integral Equations
Example 16.2.2 GENERALIZED ABELEQUATION ,CONVOLUTION THEOREM
ThegeneralizedAbelequationis
f(x)=integraldisplayx
0ϕ(t)
(x−t)αdt,0<α<1,withbraceleftbiggf(x)known,
ϕ(t)unknown .(16.44)
TakingtheLaplacetransformofbothsidesof thisequation,weobtain
L{f(x)}=Lbraceleftbiggintegraldisplayx
0ϕ(t)
(x−t)αdtbracerightbigg
=Lbraceleftbig
x−αbracerightbig
Lbraceleftbig
ϕ(x)bracerightbig
, (16.45)
thelast stepfollowingbytheLaplaceconvolutiontheorem(Section15.11).Then
Lbraceleftbig
ϕ(x)bracerightbig
=s1−αL{f(x)}
(−α)!. (16.46)
Dividingby s,1weobtain
1
sLbraceleftbig
ϕ(x)bracerightbig
=s−αL{f(x)}
(−α)!=L{xα−1}L{f(x)}
(α−1)!(−α)!. (16.47)
Combiningthefactorials(Eq.(8.32))andapplyingtheLaplaceconvolutiontheoremagain,
wediscoverthat
1
sLbraceleftbig
ϕ(x)bracerightbig
=sinπα
πLbraceleftbiggintegraldisplayx
0f(t)
(x−t)1−αdtbracerightbigg
. (16.48)
InvertingwiththeaidofExercise15.11.1,weget
integraldisplayx
0ϕ(t)dt=sinπα
πintegraldisplayx
0f(t)
(x−t)1−αdt, (16.49)
andfinally,bydifferentiating,
ϕ(x)=sinπα
πd
dxintegraldisplayx
0f(t)
(x−t)1−αdt. (16.50)
/squaresolid
Generating Functions
Occasionally, the reader may encounter integral equations that involve generating func-
tions.Supposewehavetheadmittedlyspecialcase
f(x)=integraldisplay1
−1ϕ(t)
(1−2xt+x2)1/2dt,−1≤x≤1. (16.51)
Wenoticetwoimportantfeatures:
1.(1−2xt+x2)−1/2generatestheLegendrepolynomials.
2.[−1,1]istheorthogonalityintervalfor theLegendrepolynomials.
1s1−αdoes not have an inverse for 0 <α<1.
16.2 Integral Transforms, Generating Functions 1015
Ifwenowexpandthedenominator(property1)andassumethatourunknown ϕ(t)may
bewrittenasaseries ofthesesameLegendrepolynomials,
f(x)=integraldisplay1
−1∞summationdisplay
n=0anPn(t)∞summationdisplay
r=0Pr(t)xrdt. (16.52)
UtilizingtheorthogonalityoftheLegendrepolynomials(property2), weobtain
f(x)=∞summationdisplay
r=02ar
2r+1xr. (16.53)
Wemayidentifythe anbydifferentiating ntimesandthensetting x=0:
f(n)(0)=n!2
2n+1an. (16.54)
Hence
ϕ(t)=∞summationdisplay
n=02n+1
2f(n)(0)
n!Pn(t). (16.55)
Similar results may be obtained with the other generating functions (compare Exer-
cise7.1.6).
•This technique of expanding in a series of special functions is always available. It is
worth a try whenever the expansion is possible (and convenient) and the interval is
appropriate.
Exercises
16.2.1 ThekernelofaFredholmequationofthesecondkind,
ϕ(x)=f(x)+λintegraldisplay∞
−∞K(x,t)ϕ(t)dt,
isof theform k(x−t).2Assumingthattherequiredtransforms exist,showthat
ϕ(x)=1√
2πintegraldisplay∞
−∞F(t)e−ixtdt
1−√
2πλK(t).
F(t)andK(t)aretheFouriertransformsof f(x)andk(x), respectively.
16.2.2 ThekernelofaVolterraequationof thefirst kind,
f(x)=integraldisplayx
0K(x,t)ϕ(t)dt,
2Thiskernelandarange 0 ≤x<∞arethecharacteristicsofintegralequationsoftheWiener–Hopftype.Detailswillbefound
in Chapter 8 of Morse and Feshbach (1953); seethe Additional Readings.
1016 Chapter 16 Integral Equations
hastheform k(x−t).Assumingthattherequiredtransforms exist,showthat
ϕ(x)=1
2πiintegraldisplayγ+i∞
γ−i∞F(s)
K(s)exsds.
F(s)andK(s)aretheLaplacetransforms of f(x)andk(x), respectively.
16.2.3 ThekernelofaVolterraequationof thesecondkind,
ϕ(x)=f(x)+λintegraldisplayx
0K(x,t)ϕ(t)dt,
hastheform k(x−t).Assumingthattherequiredtransforms exist,showthat
ϕ(x)=1
2πiintegraldisplayγ+i∞
γ−i∞F(s)
1−λK(s)exsds.
16.2.4 UsingtheLaplacetransform solution(Exercise16.2.3), solve
(a)ϕ(x)=x+integraldisplayx
0(t−x)ϕ(t)dt .
ANS.ϕ(x)=sinx.
(b)ϕ(x)=x−integraldisplayx
0(t−x)ϕ(t)dt .
ANS.ϕ(x)=sinhx.
Checkyourresults bysubstitutingbackintotheoriginalintegralequations.
16.2.5 Reformulate the equations of Example 16.2.1 (Eqs. (16.39) to (16.43)), using Fourier
cosinetransforms.
16.2.6 GiventheFredholmintegralequation,
e−x2=integraldisplay∞
−∞e−(x−t)2ϕ(t)dt,
applytheFourierconvolutiontechniqueof Example16.2.1tosolvefor ϕ(t).
16.2.7 SolveAbel’sequation,
f(x)=integraldisplayx
0ϕ(t)
(x−t)αdt,0<α<1,
bythefollowingmethod:
(a) Multiply both sides by (z−x)α−1and integrate with respect to xover the range
0≤x≤z.
(b) Reverse the order of integration and evaluate the integral on the right-hand side
(withrespectto x)bythebetafunction.
16.2 Integral Transforms, Generating Functions 1017
Note.integraldisplayz
tdx
(z−x)1−α(x−t)α=B(1−α,α)=(−α)!(α−1)!=π
sinπα.
16.2.8 GiventhegeneralizedAbelequationwith f(x)=1,
1=integraldisplayx
0ϕ(t)
(x−t)αdt,0<α<1,
solvefor ϕ(t)andverifythat ϕ(t)is asolutionof thegivenequation.
ANS.ϕ(t)=sinπα
πtα−1.
16.2.9 AFredholmequationof thefirst kindhas akernel e−(x−t)2:
f(x)=integraldisplay∞
−∞e−(x−t)2ϕ(t)dt.
Showthatthesolutionis
ϕ(x)=1√π∞summationdisplay
π=0f(n)(0)
2nn!Hn(x),
inwhich Hn(x)is annth-orderHermitepolynomial.
16.2.10 Solvetheintegralequation
f(x)=integraldisplay1
−1ϕ(t)
(1−2xt+x2)1/2dt,−1≤x≤1,
for theunknownfunction ϕ(t)if
(a)f(x)=x2s,(b)f( x)=x2s+1.
ANS.(a) ϕ(t)=4s+1
2P2s(t),(b)ϕ(t)=4s+3
2P2s+1(t).
16.2.11 AKirchhoffdiffractiontheoryanalysisofalaserleadstotheintegralequation
v(r2)=γintegraldisplayintegraldisplay
K(r1,r2)v(r1)dA.
The unknown, v(r1), gives the geometric distribution of the radiation field over one
mirror surface; the range of integration is over the surface of that mirror. For square
confocalsphericalmirrors theintegralequationbecomes
v(x2,y2)=−iγeikb
λbintegraldisplaya
−aintegraldisplaya
−ae−(ik/b)(x 1x2+y1y2)v(x1,y1)dx1dy1,
in which bis the centerline distance between the laser mirrors. This can be put in a
somewhatsimplerformbythesubstitutions
kx2
i
b=ξ2
i,ky2
i
b=η2
i,andka2
b=2πa2
λb=α2.
(a) Showthatthevariablesseparateandwegettwo integralequations.
1018 Chapter 16 Integral Equations
(b) Show that the new limits, ±α, may be approximated by ±∞for a mirror dimen-
siona≫λ.
(c) Solvetheresultingintegralequations.
16.3 N EUMANN SERIES ,SEPARABLE (DEGENERATE )
KERNELS
Many and probably most integral equations cannot be solved by the specialized integral
transform techniques of the preceding section. Here we develop three rather general tech-
niques for solving integral equations. The first, due largely to Neumann, Liouville, and
Volterra, develops the unknown function ϕ(x)as a power series in λ, whereλis a given
constant.Themethodis applicablewhenevertheseries converges.
Thesecondmethodissomewhatrestrictedbecauseitrequiresthatthetwovariablesap-
pearingin the kernel K(x,t)be separable. However,there are two major rewards: (1) The
relation between an integral equation and a set of simultaneous linear algebraic equations
is shown explicitly, and (2) the method leads to eigenvalues and eigenfunctions—in close
analogytoSection3.5.
Third, a technique for numerical solution of Fredholm equations of both the first and
secondkindisoutlined.Theproblemposedbyill-conditionedmatricesis emphasized.
Neumann Series
We solve a linear integral equation of the second kind by successive approximations; our
integralequationis theFredholmequation,
ϕ(x)=f(x)+λintegraldisplayb
aK(x,t)ϕ(t)dt, (16.56)
in which f(x)/negationslash=0. If the upper limit of the integral is a variable (Volterra equation), the
followingdevelopmentwillstillhold,butwithminormodifications.Letustry(thereisno
guarantee thatitwillwork)toapproximateourunknownfunctionby
ϕ(x)≈ϕ0(x)=f(x). (16.57)
This choice is not mandatory. If you can make a better guess, go ahead and guess. The
choice here is equivalent to saying that the integral or the constant λis small. To improve
thisfirstcrudeapproximation,wefeed ϕ0(x)backintotheintegral,Eq. (16.56),andget
ϕ1(x)=f(x)+λintegraldisplayb
aK(x,t)f(t)dt. (16.58)
Repeatingthisprocessofsubstitutingthenew ϕn(x)backintoEq.(16.56),wedevelopthe
sequence
ϕ2(x)=f(x)+λintegraldisplayb
aK(x,t1)f(t1)dt1
+λ2integraldisplayb
aintegraldisplayb
aK(x,t1)K(t1,t2)f(t2)dt2dt1 (16.59)
16.3 Neumann Series, Separable (Degenerate) Kernels 1019
and
ϕn(x)=nsummationdisplay
i=0λiui(x), (16.60)
where
u0(x)=f(x),
u1(x)=integraldisplayb
aK(x,t1)f(t1)dt1,
(16.61)
u2(x)=integraldisplayb
aintegraldisplayb
aK(x,t1)K(t1,t2)f(t2)dt2dt1,
un(x)=integraldisplayintegraldisplay
···integraldisplay
K(x,t1)K(t1,t2)···K(tn−1,tn)·f(tn)dtn···dt1.
Weexpectthatoursolution ϕ(x)willbe
ϕ(x)=limn→∞ϕn(x)=limn→∞nsummationdisplay
i=0λiui(x), (16.62)
providedthatourinfiniteseriesconverges .Wemayconvenientlychecktheconvergence
bytheCauchyratiotest,Section5.2, notingthat
vextendsinglevextendsingleλnun(x)vextendsinglevextendsingle≤vextendsinglevextendsingleλnvextendsinglevextendsingle·|f|max·|K|n
max·|b−a|n, (16.63)
using|f|maxto represent the maximum value of|f(x)|in the interval [a,b]and|K|max
to represent the maximum value of |K(x,t)|in its domain in the x,t-plane. We have
convergenceif
|λ|·|K|max·|b−a|<1. (16.64)
Notethat λ|un(max)|isbeingusedasa comparison series.Ifitconverges,ouractualseries
must converge. If this condition is not satisfied, we may or may not have convergence.
A more sensitive test is required. Of course, even if the Neumann series diverges, there
stillmaybeasolutionobtainablebyanothermethod.
To see what has been done with this iterative manipulation, we may find it helpful to
rewrite the Neumann series solution, Eq. (16.59), in operator form. We start by rewriting
Eq.(16.56) as
ϕ=λKϕ+f,
whereKrepresents the integraloperatorintegraltextb
aK(x,t)[]dt.Solvingfor ϕ, weobtain
ϕ=(1−λK)−1f.
Binomial expansion leads to Eq. (16.59). The convergence of the Neumann series is a
demonstrationthattheinverseoperator (1−λK)−1exists.
1020 Chapter 16 Integral Equations
Example 16.3.1 NEUMANN SERIES SOLUTION
ToillustratetheNeumannmethod,weconsidertheintegralequation
ϕ(x)=x+1
2integraldisplay1
−1(t−x)ϕ(t)dt. (16.65)
TostarttheNeumannseries, wetake
ϕ0(x)=x. (16.66)
Then
ϕ1(x)=x+1
2integraldisplay1
−1(t−x)tdt=x+1
2parenleftbigg1
3t3−1
2t2xparenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle1
−1=x+1
3.
Substituting ϕ1(x)backintoEq. (16.65), weget
ϕ2(x)=x+1
2integraldisplay1
−1(t−x)tdt+1
2integraldisplay1
−1(t−x)1
3dt=x+1
3−x
3.
Continuingthisprocess ofsubstitutingbackintoEq. (16.65),weobtain
ϕ3(x)=x+1
3−x
3−1
32,
andbyinduction
ϕ2n(x)=x+nsummationdisplay
s=1(−1)s−13−s−xnsummationdisplay
s=1(−1)s−13−s. (16.67)
Lettingn→∞, weget
ϕ(x)=3
4x+1
4. (16.68)
This solution can (and should) be checked by substituting back into the original equation,
Eq.(16.65). /squaresolid
It is interesting to note that our series converged easily even though Eq. (16.64) is not
satisfied in this particular case. Actually Eq. (16.64) is a rather crude upper bound on λ.
It can be shown that a necessary and sufficient condition for the convergence of our series
solution is that |λ|<|λe|, whereλeis the eigenvalue of smallest magnitude of the cor-
responding homogeneous equation [f(x)=0)]. For this particular example λe=√
3/2.
Clearly,λ=1
2<λe=√
3/2.
Oneapproachtothecalculationoftime-dependentperturbationsinquantummechanics
startswiththeintegralequationfor theevolutionoperator
U(t,t0)=1−i
¯hintegraldisplayt
t0V(t1)U(t1,t0)dt1. (16.69a)
Iterationleadsto
U(t,t0)=1−i
¯hintegraldisplayt
t0V(t1)dt1+parenleftbiggi
¯hparenrightbigg2integraldisplayt
t0integraldisplayt1
t0V(t1)V(t2)dt2dt1+···.(16.69b)
16.3 Neumann Series, Separable (Degenerate) Kernels 1021
Theevolutionoperatorisobtainedasaseriesofmultipleintegralsoftheperturbingpoten-
tialV(t), closely analogous to the Neumann series, Eq. (16.60). For V=V0, independent
oft, the evolution operator becomes (see Exercise 3.4.13, replace t→/Delta1t, and construct
Ufrom productsof T(t+/Delta1t,t)asinEq. (4.26))
U(t1,t0)=expbracketleftbigg
−i
¯h(t−t0)V0bracketrightbigg
.
AsecondandsimilarrelationshipbetweentheNeumannseriesandquantummechanics
appears when the Schrödinger wave equation for scattering is reformulated as an integral
equation. The first term in a Neumann series solution is the incident (unperturbed) wave.
Thesecondtermis thefirst-order Bornapproximation,Eq. (9.203b)of Section9.7.
The Neumann method may also be applied to Volterra integral equations of the second
kind, Eq. (16.4) or Eq. (16.56) with the fixed upper limit, b, replaced by a variable, x.I n
the Volterra case the Neumann series converges for all λas long as the kernel is square
integrable.
Separable Kernel
Thetechniqueofreplacingourintegralequationbysimultaneousalgebraicequationsmay
alsobeusedwheneverourkernel K(x,t)is separable,inthesensethat
K(x,t)=nsummationdisplay
j=1Mj(x)Nj(t), (16.70)
wheren, the upper limit of the sum, is finite. Such kernels are sometimes called degener-
ate. Our class of separable kernels includes all polynomials and many of the elementary
transcendentalfunctions;thatis,
cos(t−x)=costcosx+sintsinx. (16.70a)
If Eq. (16.70) is satisfied, substitution into the Fredholm equation of the second kind, Eq.
(16.2), yields
ϕ(x)=f(x)+λnsummationdisplay
j=1Mj(x)integraldisplayb
aNj(t)ϕ(t)dt, (16.71)
interchangingintegrationandsummation.Now,theintegralwithrespectto tisaconstant,
integraldisplayb
aNj(t)ϕ(t)dt=cj. (16.72)
HenceEq.(16.71) becomes
ϕ(x)=f(x)+λnsummationdisplay
j=1cjMj(x). (16.73)
This gives us ϕ(x), our solution, once the constants cihave been determined. Equa-
tion (16.73) further tells us the form of ϕ(x):f(x), plus a linear combination of the
x-dependentfactors oftheseparablekernel.
1022 Chapter 16 Integral Equations
We may find ciby multiplying Eq. (16.73) by Ni(x)and integrating to eliminate the
x-dependence.UseofEq. (16.72)yields
ci=bi+λnsummationdisplay
j=1aijcj, (16.74)
where
bi=integraldisplayb
aNi(x)f(x)dx, a ij=integraldisplayb
aNi(x)Mj(x)dx. (16.75)
It isperhapshelpfultowriteEq. (16.74)inmatrixform, with A=(aij):
b=c−λAc=(1−λA)c, (16.76a)
or3
c=(1−λA)−1b. (16.76b)
Equation(16.76a)isequivalenttoasetof simultaneouslinearalgebraicequations
(1−λa11)c1−λa12c2−λa13c3−···=b1,
−λa21c1+(1−λa22)c2−λa23c3−···=b2, (16.77)
−λa31c1−λa32c2+(1−λa33)c3−···=b3,andso on .
If our integral equation is homogeneous, [f(x)=0], thenb=0. To get a solution, we set
thedeterminantofthecoefficientsof ciequaltozero,
|1−λA|=0, (16.78)
exactly as in Section 3.5. The roots of Eq. (16.78) yield our eigenvalues. Substituting into
(1−λA)c=0,wefindthe ci,andthenEq. (16.73)givesoursolution.
Example 16.3.2
Toillustratethistechniquefordeterminingeigenvaluesandeigenfunctionsofthehomoge-
neousFredholmequation,weconsiderthecase
ϕ(x)=λintegraldisplay1
−1(t+x)ϕ(t)dt. (16.79)
Here(comparewithEqs. (16.71)and(16.77))
M1=1,M 2(x)=x,
N1(t)=t, N 2=1.
Equation(16.75)yields
a11=a22=0,a 12=2
3,a 21=2;b1=0=b2.
3Noticethe similarity to the operator form ofthe Neumann series.
16.3 Neumann Series, Separable (Degenerate) Kernels 1023
Equation(16.78),our secularequation,becomes
vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle1−2λ
3
−2λ1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (16.80)
Expanding,weobtain
1−4λ2
3=0,λ=±√
3
2. (16.81)
Substitutingtheeigenvalues λ=±√
3/2 intoEq.(16.76), wehave
c1∓c2√
3=0. (16.82)
Finally,witha choiceof c1=1,Eq.(16.73) gives
ϕ1(x)=√
3
2(1+√
3x), λ=√
3
2, (16.83)
ϕ2(x)=−√
3
2(1−√
3x), λ=−√
3
2. (16.84)
Sinceourequationishomogeneous,thenormalizationof ϕ(x)is arbitrary. /squaresolid
If the kernel is not separable in the sense of Eq. (16.70), there is still the possibility that
itmaybeapproximatedbyakernelthatisseparable.Thenwecangettheexactsolutionof
anapproximateequation,anequationthatapproximatestheoriginalequation.Thesolution
oftheseparableapproximatekernelproblemcanthenbecheckedbysubstitutingbackinto
theoriginal,unseparablekernelproblem.
Numerical Solution
There is extensive literature on the numerical solution of integral equations, and much of
it concerns special techniques for certain situations. One method of fair generality is the
replacement of the single integral equation by a set of simultaneous algebraic equations.
And again matrix techniques are invoked. This simultaneous algebraic equation–matrix
approach is applied here to two different cases. For the homogeneous Fredholm equation
ofthesecondkindthismethodworks well.FortheFredholmequationofthefirstkindthe
methodis adisaster.First wedealwiththedisaster.
We considertheFredholmintegralequationofthefirst kind,
f(x)=integraldisplayb
aK(x,t)ϕ(t)dt, (16.84a)
withf(x)andK(x,t)known and ϕ(t)unknown. The integral can be evaluated (in prin-
ciple) by quadrature techniques. For maximum accuracy the Gaussian method is recom-
mended(ifthekerneliscontinuousandhascontinuousderivatives).Thenumericalquadra-
turereplacestheintegralbyasummation,
f(xi)=nsummationdisplay
k=1AkK(xi,tk)ϕ(tk), (16.84b)
1024 Chapter 16 Integral Equations
withAkthe quadrature coefficients. We abbreviate f(xi)asfi,ϕ(tk)asϕk, and
AkK(xi,tk)asBik. In effect we are changing from a function description to a vector–
matrix description, with the ncomponents of the vector (fi)defined as the values of the
functionatthe ndiscretepoints [f(xi)]. Equation(16.84b)becomes
fi=nsummationdisplay
k=1Bikϕk,
amatrixequation.Inverting (Bik), weobtain
ϕ(xk)=ϕk=nsummationdisplay
k=1B−1
kifi, (16.84c)
andEq.(16.84a)issolved—inprinciple.Inpractice,thequadraturecoefficient–kernelma-
trix is often “ill-conditioned” (with respect to inversion). This means that in the inversion
process small (numerical) errors are multiplied by large factors. In the inversion process
allsignificantfiguresmaybelost andEq. (16.84c)becomesnumericalnonsense.
This disaster should not be entirely unexpected. Integration is essentially a smoothing
operation. f(x)is relatively insensitive to local variation of ϕ(t). Conversely, ϕ(t)may
be exceedingly sensitive to small changes in f(x). Small errors in f(x)or inB−1are
magnified and accuracy disappears. This same behavior shows up in attempts to invert
Laplacetransformsnumerically.
When the quadrature–matrix technique is applied to the integral equation eigenvalue
problem,thesymmetrickernel,homogeneousFredholmequationofthesecondkind,4
λϕ(x)=integraldisplayb
aK(x,t)ϕ(t)dt, (16.84d)
the technique is far more successful. Replacing the integral by a set of simultaneous alge-
braicequations(numericalquadrature),wehave
λϕi=nsummationdisplay
k=1AkKikϕk, (16.84e)
withϕi=ϕ(xi), as before. The points xi,i=1,2,...,n, are taken to be the same (nu-
merically) as tk,k=1,2,...,n,s oKikwill be symmetric. The system is symmetrized by
multiplyingby A1/2
isothat
λparenleftbig
A1/2
iϕiparenrightbig
=nsummationdisplay
k=1parenleftbig
A1/2
iKikA1/2
kparenrightbigparenleftbig
A1/2
kϕkparenrightbig
. (16.84f)
Replacing A1/2
iϕibyψiandA1/2
iKikA1/2
kbySik, weobtain
λψ=Sψ, (16.84g)
withSsymmetric (since the kernel K(x,t)was assumed symmetric). Of course, ψhas
components ψi=ψ(xi).Equation(16.84g)isourmatrixeigenvalueequation,Eq.(3.136).
4The eigenvalue λhas been written on the left side, multiplying the eigenfunction, as is customary in matrix analysis (Section
3.5). In this form λwilltakeon a maximum value .
16.3 Neumann Series, Separable (Degenerate) Kernels 1025
Theeigenvaluesarereadilyobtainedbycallingacannedeigenroutine.5Forkernelssuchas
those of Exercise 16.3.15 and using a 10-point Gauss–Legendre quadrature, the eigenrou-
tine determines the largest eigenvalue to within about 0.5 percent for the cases where the
kernel has discontinuities in its derivatives. If the derivatives are continuous, the accuracy
ismuchbetter.
Linz6hasdescribedaninterestingvariationalrefinementinthedeterminationof λmaxto
highaccuracy.ThekeytohismethodisExercise17.8.7.Thecomponentsoftheeigenfunc-
tionvectorareobtainedfromEq.(16.84d)with ϕ(tk)nowknownand ϕi=ϕ(xi)generated
asrequired.(The xiarenolongertiedtothe tk.)
Exercises
16.3.1 UsingtheNeumannseries, solve
(a)ϕ(x)=1−2integraldisplayx
0tϕ(t)dt,
(b)ϕ(x)=x+integraldisplayx
0(t−x)ϕ(t)dt ,
(c)ϕ(x)=x−integraldisplayx
0(t−x)ϕ(t)dt .
ANS.(a) ϕ(x)=e−x2.
16.3.2 Solvetheequation
ϕ(x)=x+1
2integraldisplay1
−1(t+x)ϕ(t)dt
bytheseparablekernelmethod.ComparewiththeNeumannmethodsolutionofSection
16.3.
ANS.ϕ(x)=1
2(3x−1).
16.3.3 Findtheeigenvaluesandeigenfunctionsof
ϕ(x)=λintegraldisplay1
−1(t−x)ϕ(t)dt.
16.3.4 Findtheeigenvaluesandeigenfunctionsof
ϕ(x)=λintegraldisplay2π
0cos(x−t)ϕ(t)dt.
ANS.λ1=λ2=1
π,ϕ(x)=Acosx+Bsinx.
5SeeW.H.Press,B.P.Flannery,S.A.Teukolsky,andW.T.Vetterling, NumericalRecipes ,2nded.,Cambridge,UK:Cambridge
UniversityPress(1992),Chapter11,fordetails,references,andcomputercodes.Thesymbolicsoftware Mathematica andMaple
also include matrix functions for computing eigenvalues and eigenvectors.
6P. Linz, On the numerical computation of eigenvalues and eigenvectors of symmetric integral equations. Math. Comput. 24:
905 (1970).
1026 Chapter 16 Integral Equations
16.3.5 Findtheeigenvaluesandeigenfunctionsof
y(x)=λintegraldisplay1
−1(x−t)2y(t)dt.
Hint.This problem may be treated by the separable kernel method or by a Legendre
expansion.
16.3.6 IftheseparablekerneltechniqueofthissectionisappliedtoaFredholmequationofthe
firstkind(Eq. (16.1)), showthatEq.(16.76) isreplacedby
c=A−1b.
Ingeneralthesolutionfor theunknown ϕ(t)isnotunique.
16.3.7 Solve
ψ(x)=x+integraldisplay1
0(1+xt)ψ(t)dt
byeachofthefollowingmethods:
(a) theNeumannseriestechnique,
(b) theseparablekerneltechnique,
(c) educatedguessing.
16.3.8 Usetheseparablekerneltechniquetoshowthat
ψ(x)=λintegraldisplayπ
0cosxsintψ(t)dt
hasnosolution(apartfromthetrivial ψ=0).Explainthisresultintermsofseparability
andsymmetry.
16.3.9 Solve
ϕ(x)=1+λ2integraldisplayx
0(x−t)ϕ(t)dt
byeachofthefollowingmethods:
(a) reductiontoanODE(findtheboundaryconditions),
(b) theNeumannseries,
(c) theuse ofLaplacetransforms.
ANS.ϕ(x)=coshλx.
16.3.10 (a) In Eq. (16.69a) take V=V0, independent of t. Without using Eq. (16.69b), show
thatEq. (16.69a)leadsdirectlyto
U(t−t0)=expbracketleftbigg
−i
¯h(t−t0)V0bracketrightbigg
.
(b) Repeatfor Eq.(16.69b) withoutusingEq. (16.69a).
16.3 Neumann Series, Separable (Degenerate) Kernels 1027
16.3.11 Givenϕ(x)=λintegraltext1
0(1+xt)ϕ(t)dt , solve for the eigenvalues and the eigenfunctions by
theseparablekerneltechnique.
16.3.12 Knowingtheformof thesolutionscanbeagreatadvantage,for theintegralequation
ϕ(x)=λintegraldisplay1
0(1+xt)ϕ(t)dt,
assumeϕ(x)to have the form 1 +bx. Substitute into the integral equation. Integrate
andsolvefor bandλ.
16.3.13 Theintegralequation
ϕ(x)=λintegraldisplay1
0J0(αxt)ϕ(t)dt, J 0(α)=0,
isapproximatedby
ϕ(x)=λintegraldisplay1
0bracketleftbig
1−x2t2bracketrightbig
ϕ(t)dt.
Find the minimum eigenvalue λand the corresponding eigenfunction ϕ(t)of the ap-
proximateequation.
ANS.λmin=1.112486,ϕ(x)=1−0.303337x2.
16.3.14 Youare giventheintegralequation
ϕ(x)=λintegraldisplay1
0sinπxtϕ(t)dt.
Approximatethekernelby
K(x,t)=4xt(1−xt)≈sinπxt.
Find the positive eigenvalue and the corresponding eigenfunction for the approximate
integralequation.
Note.ForK(x,t)=sinπxt,λ=1.6334.
ANS.λ=1.5678,ϕ(x)=x−0.6955x2
(λ+=√
31−4,λ−=−√
31−4).
16.3.15 Theequation
f(x)=integraldisplayb
aK(x,t)ϕ(t)dt
hasadegeneratekernel K(x,t)=summationtextn
i=1Mi(x)Ni(t).
(a) Showthatthisintegralequationhasnosolutionunless f(x)canbewrittenas
f(x)=nsummationdisplay
i=1fiMi(x),
withtheficonstants.
1028 Chapter 16 Integral Equations
(b) Showthattoanysolution ϕ(x)wemayadd ψ(x),provided ψ(x)isorthogonalto
allNi(x):
integraldisplayb
aNi(x)ψ(x)dx=0 for all i.
16.3.16 Usingnumericalquadrature,convert
ϕ(x)=λintegraldisplay1
0J0(αxt)ϕ(t)dt, J 0(α)=0,
toasetof simultaneouslinearequations.
(a) Findtheminimumeigenvalue λ.
(b) Determine ϕ(x)at discrete values of xand plotϕ(x)versusx. Compare with the
approximateeigenfunctionofExercise16.3.13.
ANS.(a) λmin=1.14502.
16.3.17 Usingnumericalquadrature,convert
ϕ(x)=λintegraldisplay1
0sinπxtϕ(t)dt
toasetof simultaneouslinearequations.
(a) Findtheminimumeigenvalue λ.
(b) Determine ϕ(x)at discrete values of xand plotϕ(x)versusx. Compare with the
approximateeigenfunctionofExercise16.3.14.
ANS.(a) λmin=1.6334.
16.3.18 GivenahomogeneousFredholmequationof thesecondkind
λϕ(x)=integraldisplay1
0K(x,t)ϕ(t)dt.
(a) Calculate the largest eigenvalue λ0. Use the 10-point Gauss–Legendre quadrature
technique.ForcomparisontheeigenvalueslistedbyLinzaregivenas λexact.
(b) Tabulate ϕ(xk), wherethe xkarethe10evaluationpointsin [0,1].
(c) Tabulatetheratio
1
λ0ϕ(x)integraldisplay1
0K(x,t)ϕ(t)dt forx=xk.
Thisis thetestofwhetheror notyoureallyhaveasolution.
(a)K(x,t)=ext.
ANS.λexact=1.35303.
16.4 Hilbert–Schmidt Theory 1029
(b)K(x,t)=braceleftBigg1
2x(2−t), x<t,
1
2t(2−x), x>t.
ANS.λexact=0.24296.
(c)K(x,t)=|x−t|.
ANS.λexact=0.34741.
(d)K(x,t)=braceleftBiggx, x<t,
t, x>t.
ANS.λexact=0.40528.
Note.(1) The evaluation points xiof Gauss–Legendre quadrature for [−1,1]may be
linearlytransformedinto [0,1],
xi[0,1]=1
2parenleftbig
xi[−1,1]+1parenrightbig
.
Thentheweightingfactors Aiare reducedinproportiontothelengthoftheinterval:
Ai[0,1]=1
2Ai[−1,1].
16.3.19 UsingthematrixvariationaltechniqueofExercise17.8.7,refineyourcalculationofthe
eigenvalueofExercise16.3.18(c) [K(x,t)=|x−t|].T rya40×40 matrix.
Note.Your matrix should be symmetric so that the (unknown) eigenvectors will be or-
thogonal.
ANS.(40-pointGauss–Legendrequadrature)0.34727.
16.4 H ILBERT –SCHMIDT THEORY
Symmetrization of Kernels
Thisisthedevelopmentofthepropertiesoflinearintegralequations(Fredholmtype)with
symmetrickernels:
K(x,t)=K(t,x). (16.85)
Beforeplungingintothetheory,wenotethatsomeimportantnonsymmetrickernelscanbe
symmetrized.If wehavetheequation
ϕ(x)=f(x)+λintegraldisplayb
aK(x,t)ρ(t)ϕ(t)dt, (16.86)
thetotalkernelisactually K(x,t)ρ(t) ,clearlynotsymmetricif K(x,t)aloneissymmetric.
However,ifwemultiplyEq.(16.86) by√ρ(x)andsubstitute
radicalbig
ρ(x)ϕ(x)=ψ(x), (16.87)
1030 Chapter 16 Integral Equations
weobtain
ψ(x)=radicalbig
ρ(x)f(x)+λintegraldisplayb
abracketleftbig
K(x,t)radicalbig
ρ(x)ρ(t)bracketrightbig
ψ(t)dt, (16.88)
with a symmetric total kernel K(x,t)√ρ(x)ρ(t). We shall meet ρ(x)later as a positive
weightingfactor inthis integralequationSturm–Liouvilletheory.
Orthogonal Eigenfunctions
WenowfocusonthehomogeneousFredholmequationofthesecondkind:
ϕ(x)=λintegraldisplayb
aK(x,t)ϕ(t)dt. (16.89)
We assume that the kernel K(x,t)is symmetric and real. Perhaps one of the first ques-
tions we might ask about the equation is: “Does it make sense?” or more precisely, “Does
an eigenvalue λsatisfying this equation exist?” With the aid of the Schwarz and Bessel
inequalities, Chapter 10 and Courant and Hilbert (Chapter III, Section 4—see the Addi-
tional Readings) show that if K(x,t)is continuous, there is at least one such eigenvalue
andpossiblyaninfinitenumberofthem.
We show that the eigenvalues, λ, are real and that the corresponding eigenfunctions,
ϕi(x), are orthogonal. Let λi,λjbe twodifferent eigenvalues and ϕi(x),ϕj(x)be the
correspondingeigenfunctions.Equation(16.89)thenbecomes
ϕi(x)=λiintegraldisplayb
aK(x,t)ϕ i(t)dt, (16.90a)
ϕj(x)=λjintegraldisplayb
aK(x,t)ϕ j(t)dt. (16.90b)
If we multiply Eq. (16.90a) by λjϕj(x)and Eq. (16.90b) by λiϕi(x)and then integrate
withrespectto x, thetwoequationsbecome7
λjintegraldisplayb
aϕi(x)ϕj(x)dx=λiλjintegraldisplayb
aintegraldisplayb
aK(x,t)ϕ i(t)ϕj(x)dtdx, (16.91a)
λiintegraldisplayb
aϕi(x)ϕj(x)dx=λiλjintegraldisplayb
aintegraldisplayb
aK(x,t)ϕ j(t)ϕi(x)dtdx, (16.91b)
Sincewehavedemandedthat K(x,t)bysymmetric,Eq. (16.91b)mayberewrittenas
λiintegraldisplayb
aϕi(x)ϕj(x)dx=λiλjintegraldisplayb
aintegraldisplayb
aK(x,t)ϕ i(t)ϕj(x)dtdx. (16.92)
SubtractingEq. (16.92)fromEq. (16.91a), weobtain
(λj−λi)integraldisplayb
aϕi(x)ϕj(x)dx=0. (16.93)
7Weassume that the necessaryintegrals exist. For anexample of asimple pathological case,seeExercise 16.4.3.
16.4 Hilbert–Schmidt Theory 1031
ThishasthesameformasEq. (10.34)intheSturm–Liouvilletheory.Since λi/negationslash=λj,
integraldisplayb
aϕi(x)ϕj(x)dx=0,i/negationslash=j, (16.94)
proving orthogonality. Note that with a real symmetric kernel, no complex conjugates are
involvedinEq.(16.94). Fortheself-adjointorHermitiankernel,see Exercise16.4.1.
Iftheeigenvalue λiisdegenerate,8theeigenfunctionsforthatparticulareigenvaluemay
be orthogonalized by the Gram–Schmidt method (Section 10.3). Our orthogonal eigen-
functions may, of course, be normalized, and we assume that this has been done. The
resultis
integraldisplayb
aϕi(x)ϕj(x)dx=δij. (16.95)
To demonstrate that the λiare real, we need to admit complex conjugates. Taking the
complexconjugateof Eq. (16.90a),wehave
ϕ∗
i(x)=λ∗
iintegraldisplayb
aK(x,t)ϕ∗
i(t)dt, (16.96)
providedthekernel K(x,t)isreal.Now,usingEq.(16.96)insteadofEq.(16.90b),wesee
thattheanalysisleadsto
(λ∗
i−λi)integraldisplayb
aϕ∗
i(x)ϕi(x)dx=0. (16.97)
Thistimetheintegralcannotvanish(unlesswehavethetrivialsolution, ϕi(x)=0)and
λ∗
i=λi, (16.98)
orλi, oureigenvalue,isreal.
This is the thirdtime we have passed this way, first with Hermitian matrices, then with
Sturm–Liouville (self-adjoint) ODEs, and now with Hilbert–Schmidt integral equations.
The correspondence between the Hermitian matrices and the self-adjoint ODEs shows up
in physics as the two outstanding formulations of quantum mechanics—the Heisenberg
matrix approach and the Schrödinger differential operator approach. In Section 17.8 and
Exercise 17.7.6 we shall explore further the correspondence between the Hilbert–Schmidt
symmetrickernelintegralequationsandtheSturm–Liouvilleself-adjointdifferentialequa-
tions.
The eigenfunctions of our integral equations form a complete set,9in the sense that any
functiong(x)thatcanbegeneratedbytheintegral
g(x)=integraldisplay
K(x,t)h(t)dt, (16.99)
8If more than one distinct eigenfunction corresponds to the same eigenvalue (satisfying Eq. (16.89)), that eigenvalue is said to
be degenerate(see Chapters 3 and 4).
9For aproof of this statement, seeCourant and Hilbert (1953), Chapter III, Section 5, in the Additional Readings.
1032 Chapter 16 Integral Equations
in which h(t)is any piecewise continuous function, can be represented by a series of
eigenfunctions,
g(x)=∞summationdisplay
n=1anϕn(x). (16.100)
Theseriesconvergesuniformlyandabsolutely.
Letus extendthis tothekernel K(x,t)byassertingthat
K(x,t)=∞summationdisplay
n=1anϕn(t), (16.101)
andan=an(x).Substitutingintotheoriginalintegralequation(Eq.(16.89))andusingthe
orthogonalityintegral,weobtain
ϕi(x)=λiai(x). (16.102)
Therefore for our homogeneous Fredholm equation of the second kind, the kernel may be
expressedintermsof theeigenfunctionsandeigenvaluesby
K(x,t)=∞summationdisplay
n=1ϕn(x)ϕn(t)
λn(zero notaneigenvalue) . (16.103)
Herewehaveabilinearexpansion,alinearexpansionin ϕn(x)andlinearin ϕn(t).Similar
bilinear expansions appear in Section 9.7. It is possible that the expansion given by Eq.
(16.101) may not exist. As an illustration of the sort of pathological behavior that may
occur,youareinvitedtoapplythisanalysisto
ϕ(x)=λintegraldisplay∞
0e−xtϕ(t)dt
(compareExercise16.4.3).
ItshouldbeemphasizedthatthisHilbert–Schmidttheoryisconcernedwiththeestablish-
ment of properties of the eigenvalues (real) and eigenfunctions (orthogonality, complete-
ness), properties that may be of great interest and value. The Hilbert–Schmidttheory does
not solve the homogeneous integral equation for us any more than the Sturm–Liouville
theory of Chapter 10 solved the ODEs. The solutions of the integral equation come from
Sections16.2and16.3(includingnumericalanalysis).
Nonhomogeneous Integral Equation
Weneedasolutionofthenonhomogeneousequation
ϕ(x)=f(x)+λintegraldisplayb
aK(x,t)ϕ(t)dt. (16.104)
Let us assume that the solutions of the corresponding homogeneous integral equation are
known:
ϕn(x)=λnintegraldisplayb
aK(x,t)ϕ n(t)dt, (16.105)
16.4 Hilbert–Schmidt Theory 1033
the solution ϕn(x)corresponding to the eigenvalue λn. We expand both ϕ(x)andf(x)in
termsofthisset ofeigenfunctions:
ϕ(x)=∞summationdisplay
n=1anϕn(x)(anunknown) , (16.106)
f(x)=∞summationdisplay
n=1bnϕn(x)(bnknown). (16.107)
SubstitutingintoEq.(16.104), weobtain
∞summationdisplay
n=1anϕn(x)=∞summationdisplay
n=1bnϕn(x)+λintegraldisplayb
aK(x,t)∞summationdisplay
n=1anϕn(t)dt. (16.108)
By interchangingthe order of integrationand summation,we may evaluatethe integral by
Eq.(16.105), andweget
∞summationdisplay
n=1anϕn(x)=∞summationdisplay
n=1bnϕn(x)+λ∞summationdisplay
n=1anϕn(x)
λn. (16.109)
If wemultiplyby ϕi(x)andintegratefrom x=atox=b,the orthogonalityof our eigen-
functionsleadsto
ai=bi+λai
λi. (16.110)
Thiscanberewrittenas
ai=bi+λ
λi−λbi, (16.111)
whichbringsustooursolution
ϕ(x)=f(x)+λ∞summationdisplay
i=1integraltextb
af(t)ϕi(t)dt
λi−λϕi(x). (16.112)
Here it is assumed that the eigenfunctions ϕi(x)are normalized to unity. Note that if
f(x)=0there is no solution unless λ=λi. This means that our homogeneous equation
hasnosolution(exceptthetrivial ϕ(x)=0)unless λisaneigenvalue, λi.
In the event that λfor the nonhomogeneous equation (16.104) is equal to one of the
eigenvalues λpof the homogeneous equation, our solution (Eq. (16.112)) blows up. To
repairthedamagewereturntoEq. (16.110)andgivethevalue
ap=bp+λpap
λp=bp+ap (16.113)
specialattention.Clearly, apdropsoutandisnolongerdeterminedby bp,whereas bp=0.
This implies thatintegraltext
f(x)ϕp(x)dx=0; that is, f(x)is orthogonal to the eigenfunction
ϕp(x).I fthisis notthecase,wehavenosolution.
1034 Chapter 16 Integral Equations
Equation (16.111) still holds for i/negationslash=p, so we multiply by ϕi(x)and sum over i(i/negationslash=p)
toobtain
ϕ(x)=f(x)+apϕp+λp∞summationdisplay
i=1
i/negationslash=pintegraltextb
af(t)ϕi(t)dt
λi−λpϕi(x). (16.114)
Inthissolutionthe apremainsas anundeterminedconstant.10
Exercises
16.4.1 IntheFredholmequation
ϕ(x)=λintegraldisplayb
aK(x,t)ϕ(t)dt
thekernel K(x,t)is self-adjointor Hermitian:
K(x,t)=K∗(t,x).
Showthat
(a) theeigenfunctionsareorthogonal,inthesense that
integraldisplayb
aϕ∗
m(x)ϕn(x)dx=0,m/negationslash=n( λm/negationslash=λn),
(b) theeigenvaluesarereal.
16.4.2 Solvetheintegralequation
ϕ(x)=x+1
2integraldisplay1
−1(t+x)ϕ(t)dt
(compareExercise16.3.2)bytheHilbert–Schmidtmethod.
Note. The application of the Hilbert–Schmidt technique here is somewhat like using
a shotgun to kill a mosquito, especially when the equation can be solved quickly by
expandinginLegendrepolynomials.
16.4.3 SolvetheFredholmintegralequation
ϕ(x)=λintegraldisplay∞
0e−xtϕ(t)dt.
Note.A series expansion of the kernel e−xtwould permit a separable kernel-type solu-
tion(Section16.3),exceptthattheseriesisinfinite.Thissuggestsaninfinitenumberof
eigenvaluesandeigenfunctions.If youstopwith
ϕ(x)=x−1/2,λ=π−1/2,
10This is like the inhomogeneous linear ODE. We may add to its solution any constant times a solution of the corresponding
homogeneous ODE.
16.4 Hilbert–Schmidt Theory 1035
you will have missed most of the solutions. Show that the normalization integrals of
the eigenfunctions do notexist. A basic reason for this anomalous behavior is that the
rangeofintegrationis infinite,makingthisa“singular”integralequation.
16.4.4 Given
y(x)=x+λintegraldisplay1
0xty(t)dt.
(a) Determine y(x)asaNeumannseries.
(b) Find the range of λfor which your Neumann series solution is convergent. Com-
parewiththevalueobtainedfrom
|λ|·|K|max<1.
(c) Find the eigenvalue and the eigenfunction of the corresponding homogeneous in-
tegralequation.
(d) Bytheseparablekernelmethodshowthatthesolutionis
y(x)=3x
3−λ.
(e) Find y(x)bytheHilbert–Schmidtmethod.
16.4.5 InExercise16.3.4,
K(x,t)=cos(x−t).
The(unnormalized)eigenfunctionsare cos xand sinx.
(a) Show that there is a function h(t)such that K(x,s), considered as a function of s
alone,maybewrittenas
K(x,s)=integraldisplay2π
0K(s,t)h(t)dt.
(b) Showthat K(x,t)maybeexpandedas
K(x,t)=2summationdisplay
n=1ϕn(x)ϕn(t)
λn.
16.4.6 Theintegralequation ϕ(x)=λintegraltext1
0(1+xt)ϕ(t)dt haseigenvalues λ1=0.7889and λ2=
15.211 andeigenfunctions ϕ1=1+0.5352xandϕ2=1−1.8685x.
(a) Showthattheseeigenfunctionsareorthogonalovertheinterval [0,1].
(b) Normalizetheeigenfunctionstounity.
(c) Showthat
K(x,t)=ϕ1(x)ϕ1(t)
λ1+ϕ2(x)ϕ2(t)
λ2.
1036 Chapter 16 Integral Equations
ANS.(b) ϕ1(x)=0.7831+0.4191x
ϕ2(x)=1.8403−3.4386x.
16.4.7 An alternate form of the solution to the nonhomogeneous integral equation, Eq.
(16.104),is
ϕ(x)=∞summationdisplay
i=1biλi
λi−λϕi(x).
(a) DerivethisformwithoutusingEq. (16.112).
(b) ShowthatthisformandEq.(16.112)are equivalent.
16.4.8 (a) Showthattheeigenfunctionsof Exercise16.3.5areorthogonal.
(b) Showthattheeigenfunctionsof Exercise16.3.11are orthogonal.
AdditionalReadings
Bocher, M., An Introduction to the Study of Integral Equations , Cambridge Tracts in Mathematics and Mathe-
maticalPhysics, No.10. NewYork: Hafner (1960). This is ahelpful introduction to integral equations.
Cochran, J. A., The Analysis of Linear Integral Equations . New York: McGraw-Hill (1972). This is a compre-
hensivetreatmentoflinearintegralequationswhichisintendedforappliedmathematiciansandmathematical
physicists. It assumes amoderate to high levelof mathematicalcompetence onthe part of thereader.
Courant, R., and D. Hilbert, Methods of Mathematical Physics , Vol.1 (English edition). New York: Interscience
(1953). This is one of the classic works of mathematical physics. Originally published in German in 1924,
the revised English edition is an excellent reference for a rigorous treatment of integral equations, Green’s
functions, and awidevariety of other topics on mathematicalphysics.
Golberg, M. A., ed., Solution Methods of Integral Equations . New York: Plenum Press (1979). This is a set of
papers from a conference on integral equations. The initial chapter is excellent for up-to-date orientation and
awealthofreferences.
Kanval, R. P., Linear Integral Equations . New York: Academic Press (1971), reprinted, Birkhäuser (1996). This
book is a detailedbut readabletreatment of avariety oftechniques for solving linearintegral equations.
Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 7
is a particularly detailed, complete discussion of Green’s functions from the point of view of mathematical
physics. Note, however, that Morse and Feshbach frequently choose a source of 4 πδ(r−r′)in place of our
δ(r−r′).Considerable attention is devoted to bounded regions.
Muskhelishvili, N.I., Singular Integral Equations , 2nd ed.,NewYork: Dover (1992).
Stakgold, I., Green’s Functions and Boundary ValueProblems . NewYork: Wiley (1979).
CHAPTER 17
CALCULUS OF VARIATIONS
Uses of the Calculus of Variations
We now address problems where we search for a function or curve, rather than a value of
somevariable,thatmakes agivenquantitystationary,usuallyan energyoractionintegral.
Becauseafunctionisvaried,theseproblemsarecalled variational .Variationalprinciples,
such as D’Alembert’s and Hamilton’s, have been developed in classical mechanics, and
Lagrangiantechniquesoccurinquantummechanicsandfieldtheory,forexample,Fermat’s
principle of the shortest optical path in electrodynamics. Before plunging into this rather
differentbranchofmathematicalphysics,letussummarizesomeofitsusesinbothphysics
andmathematics.
1. In existingphysicaltheories:
a. Unificationofdiverseareasofphysicsusingenergyasakeyconcept.
b. Convenienceinanalysis—Lagrangeequations,Section17.3.
c. Eleganttreatmentofconstraints,Section17.7.
2. Starting point for new, complex areas of physics and engineering . In general rela-
tivity the geodesic is taken as the minimum path of a light pulse or the free-fall path
of a particle in curved Riemannian space (see geodesics in Section 2.10). Variational
principles appear in quantum field theory. Variational principles have been applied
extensivelyincontroltheory.
3. Mathematical unification . Variational analysis provides a proof of the completeness
of the Sturm–Liouville eigenfunctions, Chapter 10, and establishes a lower bound for
the eigenvalues. Similar results follow for the eigenvalues and eigenfunctions of the
Hilbert–Schmidtintegralequation,Section16.4.
4. Calculationtechniques ,Section17.8.Calculationoftheeigenfunctionsandeigenval-
uesoftheSturm–Liouvilleequation.Integralequationeigenfunctionsandeigenvalues
maybecalculatedusingnumericalquadratureandmatrixtechniques,Section16.3.
1037
1038 Chapter 17 Calculus of Variations
17.1 A D EPENDENT AND AN INDEPENDENT VARIABLE
Concept of Variation
The calculus of variations involves problems in which the quantity to be minimized (or
maximized)appearsasastationaryintegral,afunctional,becauseafunction y(x,α)needs
to be determined from a class described by an infinitesimal parameter α. As the simplest
case,let
J=integraldisplayx2
x1f(y,yx,x)dx. (17.1)
HereJis the quantity that takes on a stationary value. Under the integral sign, fis a
knownfunctionoftheindicatedvariables xandα,asarey(x,α),y x(x,α)≡∂y(x,α)/∂x ,
but the dependence of yonx(andα)is not yet known; that is, y(x)isunknown .T h i s
meansthatalthoughtheintegralisfrom x1tox2,theexactpathofintegrationisnotknown
(Fig.17.1).Wearetochoosethepathofintegrationthroughpoints (x1,y1)and(x2,y2)to
minimize J. Strictly speaking, we determine stationary values of J: minima, maxima, or
saddle points. In most cases of physical interest the stationary value will be a minimum.
This problem is considerably more difficult than the corresponding problem of a function
y(x)in differential calculus. Indeed, there may be no solution. In differential calculus the
minimum is determined by comparing y(x0)withy(x), wherexranges over neighboring
points. Here we assume the existence of an optimum path, that is, an acceptable path for
whichJis stationary, and then compare Jfor our (unknown) optimum path with that
obtainedfromneighboringpaths.InFig.17.1twopossiblepathsareshown.(Therearean
infinite number of possibilities.) The difference between these two for a given xis called
the variation of y,δy, and is conveniently described by introducing a new function, η(x),
to define the arbitrary deformation of the path and a scale factor, α, to give the magnitude
ofthevariation.Thefunction η(x)is arbitraryexceptfortworestrictions. First,
η(x1)=η(x2)=0, (17.2)
FIGURE 17.1Avariedpath.
17.1 A Dependent and an Independent Variable 1039
which means that all varied paths must pass through the fixed endpoints. Second, as will
beseenshortly, η(x)mustbedifferentiable;thatis, wemaynotuse
η(x)=1,x=x0,
(17.3)
=0,x/negationslash=x0,
but we can choose η(x)to have a form similar to the functions used to represent the Dirac
deltafunction(Chapter1)sothat η(x)differsfromzeroonlyoveraninfinitesimalregion.1
Then,withthepathdescribedby αandη(x),
y(x,α)=y(x,0)+αη(x) (17.4)
and
δy=y(x,α)−y(x,0)=αη(x). (17.5)
Let us choose y(x,α=0)as the unknown path that will minimize J. Theny(x,α)
for nonzero αdescribes a neighboring path. In Eq. (17.1), Jis now a function2of our
parameter α:
J(α)=integraldisplayx2
x1fbracketleftbig
y(x,α),y x(x,α),xbracketrightbig
dx, (17.6)
andourconditionfor anextremevalueisthat
bracketleftbigg∂J(α)
∂αbracketrightbigg
α=0=0, (17.7)
analogoustothevanishingof thederivative dy/dxindifferentialcalculus.
Now, the α-dependence of the integral is contained in y(x,α)andyx(x,α)=
(∂/∂x)y(x,α) . Therefore3
∂J(α)
∂α=integraldisplayx2
x1bracketleftbigg∂f
∂y∂y
∂α+∂f
∂yx∂yx
∂αbracketrightbigg
dx. (17.8)
FromEq. (17.4),
∂y(x,α)
∂α=η(x), (17.9)
∂yx(x,α)
∂α=dη(x)
dx, (17.10)
soEq. (17.8)becomes
∂J(α)
∂α=integraldisplayx2
x1parenleftbigg∂f
∂yη(x)+∂f
∂yxdη(x)
dxparenrightbigg
dx. (17.11)
1Compare H. Jeffreys and B. S. Jeffreys, Methods of Mathematical Physics , 3rd ed., Cambridge, UK: Cambridge University
Press (1966), Chapter 10, for amore complete discussion of this point.
2Technically, Jis afunctional of y,yx, but a function of αdepending on the functions y(x,α)andyx(x,α):J[y(x,α),
yx(x,α)].
3Notethat yandyxarebeingtreatedas independent variables.
1040 Chapter 17 Calculus of Variations
Integrating the second term by parts to get η(x)as a common and arbitrary nonvanishing
factor,weobtain
integraldisplayx2
x1dη(x)
dx∂f
∂yxdx=η(x)∂f
∂yxvextendsinglevextendsinglevextendsinglevextendsinglex2
x1−integraldisplayx2
x1η(x)d
dx∂f
∂yxdx. (17.12)
TheintegratedpartvanishesbyEq.(17.2), andEq. (17.11)becomes
integraldisplayx2
x1bracketleftbigg∂f
∂y−d
dx∂f
∂yxbracketrightbigg
η(x)dx=0. (17.13)
Inthisform αhasbeensetequaltozero,correspondingtothesolutionpath,and,ineffect,
isnolongerpartoftheproblem.
Occasionally we will see Eq. (17.13) multiplied by δα, which gives, upon using
η(x)δα=δy,
integraldisplayx2
x1parenleftbigg∂f
∂y−d
dx∂f
∂yxparenrightbigg
δydx=δαbracketleftbigg∂J
∂αbracketrightbigg
α=0=δJ=0. (17.14)
Sinceη(x)isarbitrary,wemaychooseittohavethesamesignasthebracketedexpression
inEq.(17.13)wheneverthelatterdiffersfromzero.Hencetheintegrandisalwaysnonneg-
ative. Equation (17.13), our condition for the existence of a stationary value, can then be
satisfied only if the bracketed term itself is zero almost everywhere. The condition for our
stationaryvalueis thusaPDE,4
∂f
∂y−d
dx∂f
∂yx=0, (17.15)
known as the Euler equation, which can be expressed in various other forms. Sometimes
solutionsaremissedwhentheyarenottwicedifferentiable,asrequiredbyEq.(17.15).An
exampleisGoldschmidt’sdiscontinuoussolutionofSection17.2.ItisclearthatEq.(17.15)
must be satisfied for Jto take on a stationary value, that is, for Eq. (17.14) to be satisfied.
Equation(17.15)isnecessary,butitisbynomeanssufficient.5CourantandRobbins(1996;
see the Additional Readings) illustrate this very nicely by considering the distance over a
sphere between points on the sphere, AandB, Fig. 17.2. Path (1), a great circle, is found
from Eq. (17.15). But path (2), the remainder of the great circle through points AandB,
also satisfies the Euler equation. Path (2) is a maximum, but only if we demand that it be
a great circle and then only if we make less than one circuit; that is, path (2) +ncomplete
revolutions is also a solution. If the path is not required to be a great circle, any deviation
from (2) will increase the length. This is hardly the property of a local maximum, and
that is why it is important to check the properties of solutions of Eq. (17.15) to see if they
satisfythephysicalconditionsofthegivenproblem.
4It is important to watchthemeaning of ∂/∂xandd/dxclosely. For example, if f=f[y(x),yx,x],
df
dx=∂f
∂x+∂f
∂ydy
dx+∂f
∂yxd2y
dx2.
The first term on the right gives the explicitx-dependence. The second and third terms give the implicitx-dependence via y
andyx.
5Foradiscussionofsufficiencyconditionsandthedevelopmentofthecalculusofvariationsasapartofmathematics,seeG.M.
Ewing,Calculus of Variations with Applications , New York: Norton (1969). Sufficiency conditions are also covered by Sagan
(in theAdditional Readings atthe end of this chapter).
17.1 A Dependent and an Independent Variable 1041
FIGURE 17.2Stationary
pathsoverasphere.
Example 17.1.1 OPTICAL PATHNEAREVENT HORIZON OF A BLACK HOLE
Determine the optical path in an atmosphere where the velocity of light increases in pro-
portion to the height, v(y)=y/b,withb>0 some parameter describing the light speed.
Sov=0a ty=0,which simulates the conditions at the surface of a black hole, called its
event horizon , where the gravitational force is so strong that the velocity of light goes to
zero,thuseventrappinglight.
Becauselighttakestheshortesttime,thevariationalproblemtakestheform
/Delta1t=integraldisplayt2
t1dt=integraldisplayds
v=bintegraldisplayradicalbig
dx2+dy2
ydt=minimum .
Herev=ds/dt=y/bis the velocity of light in this environment, the ycoordinate being
the height. A look at the variational functional suggests choosing yas the independent
variable because xdoes not appear in the integrand. We can bring dyoutside the radical
and change the role of xandyinJof Eq. (17.1) and the resulting Euler equation. With
x=x(y), x′=dx/dy,weobtain
bintegraldisplay√
x′2+1
ydy=minimum ,
andtheEulerequationbecomes
∂f
∂x−d
dy∂f
∂x′=0.
Since∂f/∂x=0,thiscanbeintegrated,giving
x′
y√
x′2+1=C1=const.,orx′2=C2
1y2parenleftbig
x′2+1parenrightbig
.
Separating dxanddyinthisfirst-order ODEwefindtheintegral
integraldisplayx
dx=integraldisplayyC1ydyradicalBig
1−C2
1y2,
1042 Chapter 17 Calculus of Variations
FIGURE 17.3Circularopticalpathinmedium.
whichyields
x+C2=−1
C1radicalBig
1−C2
1y2,or(x+C2)2+y2=1
C2
1.
This is a circular light path with center on the x-axis along the event horizon. (See
Fig. 17.3.) This example may be adapted to a mirage (Fata Morgana) in a desert with
hot air near the ground and cooler air aloft (the index of refraction changes with height
in cool versus hot air), thus changing the velocity law from v=y/b→v0−y/b.In this
case, the circular light path is no longer convex with center on the x-axis, but becomes
concave. /squaresolid
Alternate Forms of Euler Equations
Oneotherform(Exercise17.1.1), whichis oftenuseful, is
∂f
∂x−d
dxparenleftbigg
f−yx∂f
∂yxparenrightbigg
=0. (17.16)
In problems in which f=f(y,yx), that is, in which xdoes not appear explicitly,
Eq.(17.16) reducesto
d
dxparenleftbigg
f−yx∂f
∂yxparenrightbigg
=0, (17.17)
or
f−yx∂f
∂yx=constant. (17.18)
Example 17.1.2 Missing Dependent Variables
Consider the variational problemintegraltext
f(˙r)dt=minimum. Here ris absent from the inte-
grand.ThereforetheEulerequationsbecome
d
dt∂f
∂˙x=0,d
dt∂f
∂˙y=0,d
dt∂f
∂˙z=0,
17.1 A Dependent and an Independent Variable 1043
withr=(x,y,z), sof˙r=c=const.Solvingthesethreeequationsforthethreeunknowns
˙x,˙y,˙zyields˙r=c1=const. Integrating this constant velocity gives r=c1t+c2.The
solutionsarestraightlines,despitethegeneralnatureof thefunction f.
A physical example illustrating this case is the propagation of light in a crystal, where
thevelocityoflightdependsonthe(crystal)directionsbutnotonthelocationinthecrystal,
becauseacrystalisananisotropic homogeneousmedium .The variationalproblem
integraldisplayds
v=integraldisplay√
˙r2
v(˙r)dt=minimum
hastheformofourexample.Notethat tneednotbethetime,butitparameterizesthelight
path. /squaresolid
Exercises
17.1.1 Fordy/dx≡yx/negationslash=0,showtheequivalenceofthetwoforms ofEuler’s equation:
∂f
∂x−d
dx∂f
∂yx=0
and
∂f
∂y−d
dxparenleftbigg
f−yx∂f
∂yxparenrightbigg
=0.
17.1.2 DeriveEuler’sequationbyexpandingtheintegrandof
J(α)=integraldisplayx2
x1fbracketleftbig
y(x,α),y x(x,α),xbracketrightbig
dx
inpowersof α,usingaTaylor(Maclaurin)expansionwith yandyxasthetwovariables
(Section5.6).
Note. The stationary condition is ∂J(α)/∂α=0, evaluated at α=0. The terms
quadraticin αmaybeusefulinestablishingthenatureofthestationarysolution(maxi-
mum,minimum,or saddlepoint).
17.1.3 FindtheEulerequationcorrespondingtoEq.(17.15) if f=f(yxx,yx,y,x).
ANS.d2
dx2parenleftbigg∂f
∂yxxparenrightbigg
−d
dxparenleftbigg∂f
∂yxparenrightbigg
+∂f
∂y=0,
η(x1)=η(x2)=0,ηx(x1)=ηx(x2)=0.
17.1.4 Theintegrand f(y,yx,x)ofEq. (17.1)hastheform
f(y,yx,x)=f1(x,y)+f2(x,y)y x.
(a) ShowthattheEuler equationleadsto
∂f1
∂y−∂f2
∂x=0.
(b) Whatdoesthisimplyforthedependenceoftheintegral Juponthechoiceofpath?
1044 Chapter 17 Calculus of Variations
17.1.5 Showthattheconditionthat
J=integraldisplay
f(x,y)dx
hasastationaryvalue
(a) leadsto f(x,y)independentof yand
(b) yieldsnoinformationaboutany x-dependence.
Wegetno(continuous,differentiable)solution.Tobeameaningfulvariationalproblem,
dependenceon yorhigherderivativesis essential.
Note. The situation will change when constraints are introduced (compare Exer-
cise17.7.7).
17.2 A PPLICATIONS OF THE EULER EQUATION
Example 17.2.1 STRAIGHT LINE
PerhapsthesimplestapplicationoftheEulerequationisinthedeterminationoftheshortest
distancebetweentwopointsintheEuclidean xy-plane.Sincetheelementofdistanceis
ds=bracketleftbig
(dx)2+(dy)2bracketrightbig1/2=bracketleftbig
1+y2
xbracketrightbig1/2dx, (17.19)
thedistance Jmaybewrittenas
J=integraldisplayx2,y2
x1,y1ds=integraldisplayx2
x1bracketleftbig
1+y2
xbracketrightbig1/2dx. (17.20)
ComparisonwithEq. (17.1) showsthat
f(y,yx,x)=parenleftbig
1+y2
xparenrightbig1/2. (17.21)
SubstitutingintoEq.(17.16), weobtain
−d
dxbracketleftbigg1
(1+y2x)1/2bracketrightbigg
=0, (17.22)
or
1
(1+y2x)1/2=C,aconstant . (17.23)
Thisis satisfiedby
yx=a,asecondconstant , (17.24)
and
y=ax+b, (17.25)
which is the familiar equation for a straight line. The constants aandbare chosen so
that the line passes through the two points (x1,y1)and(x2,y2). Hence the Euler equation
17.2 Applications of the Euler Equation 1045
predictsthattheshortest6distancebetweentwofixedpointsinEuclideanspaceisastraight
line. /squaresolid
Thegeneralizationofthisincurvedfour-dimensionalspace–timeleadstotheimportant
conceptof thegeodesicingeneralrelativity(seeSection2.10).
Example 17.2.2 SOAPFILM
As a second illustration (Fig. 17.4), consider two parallel coaxial wire circles to be con-
nected by a surface of minimum area that is generated by revolving a curve y(x)about
thex-axis.Thecurveisrequiredtopassthroughfixedendpoints (x1,y1)and(x2,y2).The
variationalproblemistochoosethecurve y(x)sothattheareaoftheresultingsurfacewill
beaminimum.
Fortheelementof areashowninFig.17.4,
dA=2πyds=2πyparenleftbig
1+y2
xparenrightbig1/2dx. (17.26)
Thevariationalequationis then
J=integraldisplayx2
x12πyparenleftbig
1+y2
xparenrightbig1/2dx. (17.27)
Neglectingthe 2 π,weobtain
f(y,yx,x)=yparenleftbig
1+y2
xparenrightbig1/2. (17.28)
FIGURE 17.4Surfaceofrotation—soapfilm
problem.
6Technically,wehave astationary value. From the α2terms it can beidentified asa minimum (Exercise 17.2.2).
1046 Chapter 17 Calculus of Variations
Since∂f/∂x=0,wemayapplyEq. (17.18)directlyandget
yparenleftbig
1+y2
xparenrightbig1/2−yy2
x1
(1+y2x)1/2=c1, (17.29)
or
y
(1+y2x)1/2=c1. (17.30)
Squaring,weget
y2
1+y2x=c2
1withc2
1≤y2
min, (17.31)
and
(yx)−1=dx
dy=c1radicalBig
y2−c2
1. (17.32)
Thismaybeintegratedtogive
x=c1cosh−1y
c1+c2. (17.33)
Solvingfor y,weha v e
y=c1coshparenleftbiggx−c2
c1parenrightbigg
, (17.34)
and again c1andc2are determined by requiring the hyperbolic cosine to pass through the
points(x1,y1)and(x2,y2).Our“minimum”-areasurfaceisaspecialcaseofacatenaryof
revolution,ora catenoid. /squaresolid
Soap Film — Minimum Area
This calculus of variations contains many pitfalls for the unwary. (Remember, the Euler
equationisa necessary conditionassuminga differentiablesolution .Thesufficiencycon-
ditions are quite involved. See the Additional Readings for details.) Respect for some of
these hazards may be developed by considering a specific physical problem, for example,
aminimum-areaproblemwith (x1,y1)=(−x0,1),(x2,y2)=(+x0,1).Theminimumsur-
faceisasoapfilmstretchedbetweenthetworingsofunitradiusat x=±x0.Theproblem
istopredictthecurve y(x)assumedbythesoapfilm.
By referring to Eq. (17.34), we find that c2=0 by the symmetry of the problem about
x=0.Then
y=c1coshparenleftbiggx
c1parenrightbigg
,c 1coshparenleftbiggx0
c1parenrightbigg
=1. (17.34a)
If wetake x0=1
2weobtainatranscendentalequationfor c1,vi z.
1=c1coshparenleftbigg1
2c1parenrightbigg
. (17.35)
17.2 Applications of the Euler Equation 1047
We find that this equation has two solutions: c1=0.2350, leading to a “deep” curve, and
c1=0.8483, leading to a “flat” curve. Which curve is assumed by the soap film? Before
answering this question, consider the physical situation with the rings moved apart so that
x0=1.ThenEq.(17.34a)becomes
1=c1coshparenleftbigg1
c1parenrightbigg
, (17.36)
whichhas norealsolutions .Thephysicalsignificanceisthatastheunit-radiusringswere
moved out from the origin, a point was reached at which the soap film could no longer
maintain the same horizontal force over each vertical section. Stable equilibrium was no
longer possible. The soap film broke (irreversible process) and formed a circular film over
each ring (with a total area of 2 π=6.2832...). This is the Goldschmidt discontinuous
solution.
Thenextquestionis:Howlargemay x0beandstillgivearealsolutionforEq.(17.34a)?7
Lettingc−1
1=p, Eq.(17.34a)becomes
p=coshpx0. (17.37)
To findx0maxwe could solve for x0(as in Eq. (17.33)) and then differentiate with respect
top. Finally, with an eye on Fig. 17.5, dx0/dpwould be set equal to zero. Alternatively,
directdifferentiationof Eq.(17.37) withrespectto pyields
1=bracketleftbigg
x0+pdx0
dpbracketrightbigg
sinhpx0.
FIGURE 17.5Solutionsof
Eq. (17.34a)forunit-radiusrings
atx=±x0.
7Fromanumericalpointofviewitiseasiertoinverttheproblem.Pickavalueof c1andsolvefor x0.Equation(17.34a)becomes
x0=c1cosh−1(1/c1).This has numerical solutions in the range 0 <c1≤1.
1048 Chapter 17 Calculus of Variations
Therequirementthat dx0/dpvanishleadsto
1=x0sinhpx0. (17.38)
Equations(17.37)and(17.38)maybecombinedtoform
px0=cothpx0, (17.39)
withtheroot
px0=1.1997. (17.40)
SubstitutingintoEq.(17.37) or(17.38), weobtain
p=1.810,c 1=0.5524 (17.41)
and
x0max=0.6627. (17.42)
Returning to the question of the solution of Eq. (17.35) that describes the soap film, let
uscalculatetheareacorrespondingtoeachsolution.Wehave
A=4πintegraldisplayx0
0yparenleftbig
1+y2
xparenrightbig1/2dx=4π
c1integraldisplayx0
0y2dx (byEq.(17.30))
=4πc1integraldisplayx0
0parenleftbigg
coshx
c1parenrightbigg2
dx=πc2
1bracketleftbigg
sinhparenleftbigg2x0
c1parenrightbigg
+2x0
c1bracketrightbigg
. (17.43)
Forx0=1
2, Eq.(17.35) leadsto
c1=0.2350→A=6.8456,
c1=0.8483→A=5.9917,
showing that the former can at most be only a local minimum. A more detailed investiga-
tion(compareBliss, CalculusofVariations ,ChapterIV)showsthatthissurfaceisnoteven
alocalminimum.For x0=1
2, thesoapfilmwillbedescribedbytheflatcurve
y=0.8483coshparenleftbiggx
0.8483parenrightbigg
. (17.44)
This flat or shallow catenoid (catenary of revolution) will be an absolute minimum for
0≤x0<0.528.However,for 0 .528<x<0.6627 its areais greaterthanthatof theGold-
schmidtdiscontinuoussolution(6.2832)anditis onlyarelativeminimum(Fig.17.6).
For an excellent discussion of both the mathematical problems and experiments with
soap films, we refer to Courant and Robbins (1996) in the Additional Readings at the end
ofthechapter.
17.2 Applications of the Euler Equation 1049
FIGURE 17.6Catenoidarea(unit-radiusringsat
x=±x0).
Exercises
17.2.1 A soap film is stretched across the space between two rings of unit radius centered at
±x0on thex-axis and perpendicular to the x-axis. Using the solution developed in
Section 17.2, set up the transcendental equations for the condition that x0is such that
the area of the curved surface of rotation equals the area of the two rings (Goldschmidt
discontinuoussolution).Solvefor x0(Fig.17.7).
17.2.2 In Example 17.2.1, expand J[y(x,α)]−J[y(x,0)]in powers of α. The term linear in
αleads to the Euler equation and to the straight-line solution, Eq. (17.25). Investigate
FIGURE 17.7Surfaceof rotation.
1050 Chapter 17 Calculus of Variations
theα2term and show that the stationary value of J, the straight-line distance, is a
minimum .
17.2.3 (a) Showthattheintegral
J=integraldisplayx2
x1f(y,yx,x)dx, withf=y(x),
hasnoextremevalues.
(b) Iff(y,yx,x)=y2(x), find a discontinuous solution similar to the Goldschmidt
solutionfor thesoapfilmproblem.
17.2.4 Fermat’sprincipleof opticsstates thatalightraywillfollowthepath y(x)for which
integraldisplayx2,y2
x1,y1n(y,x)ds
is a minimum when nis the index of refraction. For y2=y1=1,−x1=x2=1, find
theraypathif
(a)n=ey,(b)n=a(y−y0), y >y 0.
17.2.5 A frictionless particle moves from point Aon the surface of the Earth to point Bby
sliding through a tunnel. Find the differential equation to be satisfied if the transit time
istobeaminimum.
Note.AssumetheEarthtobenonrotatingsphereofuniformdensity.
ANS.(Eq. (17.15)): rϕϕ(r3−ra2)+r2
ϕ(2a2−r2)+a2r2=0,
r(ϕ=0)=r0,rϕ(ϕ=0)=0,r ( ϕ=ϕA)=a, r(ϕ=ϕB)=a.
Eq. (17.18): r2
ϕ=a2r2
r2
0·r2−r2
0
a2−r2. The solution of these equations is a hypocycloid, gener-
ated by a circle of radius1
2(a−r0)rolling inside the circle of radius a. You might like
toshowthatthetransittimeis
t=π(a2−r2
0)1/2
(ag)1/2.
For details see P. W. Cooper, A m .J .P h y s . 34: 68 (1966); G. Veneziano et al., ibid. , pp.
701–704.
17.2.6 A ray of light follows a straight-line path in a first homogeneous medium, is refracted
at an interface, and then follows a new straight-line path in the second medium. Use
Fermat’sprincipleof opticstoderiveSnell’slawofrefraction:
n1sinθ1=n2sinθ2.
Hint. Keep the points (x1,y1)and(x2,y2)fixed and vary x0to satisfy Fermat
(Fig. 17.8). This is notan Euler equation problem. (The light path is not differentiable
atx0.)
17.2 Applications of the Euler Equation 1051
FIGURE 17.8Snell’slaw.
17.2.7 A second soap film configuration for the unit-radius rings at x=±x0consists of a
circular disk, radius a,i nt h ex=0 plane and two catenoids of revolution, one joining
thediskandeachring.Onecatenoidmaybedescribedby
y=c1coshparenleftbiggx
c1+c3parenrightbigg
.
(a) Imposeboundaryconditionsat x=0 andx=x0.
(b) Althoughnotnecessary,itisconvenienttorequirethatthecatenoidsformanangle
of 120◦where they join the central disk. Express this third boundary condition in
mathematicalterms.
(c) Showthatthetotalareaofcatenoidspluscentraldiskis
A=c2
1bracketleftbigg
sinhparenleftbigg2x0
c1+2c3parenrightbigg
+2x0
c1bracketrightbigg
.
Note. Although this soap film configuration is physically realizable and stable, the area
is larger than that of the simple catenoid for all ring separations for which both films
exist.
ANS.(a)
1=c1coshparenleftbiggx0
c1+c3parenrightbigg
a=c1coshc3,(b)dy
dx=tan30◦=sinhc3.
17.2.8 For the soap film described in Exercise 17.2.7, find (numerically) the maximum value
ofx0.
Note.Thiscallsforapocketcalculatorwithhyperbolicfunctionsoratableofhyperbolic
cotangents.
ANS.x0max=0.4078.
1052 Chapter 17 Calculus of Variations
17.2.9 Find the root of px0=cothpx0(Eq. (17.39)) and determine the corresponding values
ofpandx0(Eqs.(17.41)and(17.42)).Calculateyourvaluestofivesignificantfigures.
17.2.10 Forthetwo-ringsoapfilmproblemofthissectioncalculateandtabulate x0,p,p−1,and
A,thesoapfilmareafor px0=0.00(0.02)1.30.
17.2.11 Findthevalueof x0(tofivesignificantfigures)thatleadstoasoapfilmarea,Eq.(17.43),
equalto 2 π,theGoldschmidtdiscontinuoussolution.
ANS.x0=0.52770.
17.2.12 Find the curve of quickest descent from (0,0)to(x0,y0)for a particle sliding under
gravity and without friction. Show that the ratio of times taken by the particle along a
straight line joining the two points compared to along the curve of quickest descent is
(1+4/π2)1/2.
Hint.T a k eyto increase downwards. Apply Eq. (17.18) to obtain y2
x=(1−c2y)/c2y,
wherecis an integration constant. Then make the substitution y=(sin2ϕ/2)/c2to
parametrizethecycloidandtake (x0,y0)=(π/2c2,1/c2).
17.3 S EVERAL DEPENDENT VARIABLES
Our original variational problem, Eq. (17.1), may be generalized in several respects. In
this section we consider the integrand fto be a function of several dependent vari-
ablesy1(x),y2(x),y3(x),...,all of which depend on x, the independent variable. In Sec-
tion 17.4 fagain will contain only one unknown function y,b u tywill be a function of
several independent variables (over which we integrate). In Section 17.5 these two gener-
alizations are combined. In Section 17.7 the stationary value is restricted by one or more
constraints.
Formorethanonedependentvariable,Eq.(17.1) becomes
J=integraldisplayx2
x1fbracketleftbig
y1(x),y2(x),...,y n(x),y1x(x),y2x(x),...,y nx(x),xbracketrightbig
dx. (17.45)
As in Section 17.1, we determine the extreme value of Jby comparing neighboring
paths.Let
yi(x,α)=yi(x,0)+αηi(x), i=1,2,...,n, (17.46)
with the ηiindependent of one another but subject to the restrictions discussed in Sec-
tion 17.1. By differentiating Eq. (17.45) with respect to αand setting α=0, since
Eq.(17.7) stillapplies,weobtain
integraldisplayx2
x1summationdisplay
iparenleftbigg∂f
∂yiηi+∂f
∂yixηixparenrightbigg
dx=0, (17.47)
thesubscript xdenotingpartialdifferentiationwithrespectto x;thatis,yix=∂yi/∂x,and
so on. Again, each of the terms (∂f/∂y ix)ηixis integrated by parts. The integrated part
vanishesandEq. (17.47)becomes
integraldisplayx2
x1summationdisplay
iparenleftbigg∂f
∂yi−d
dx∂f
∂yixparenrightbigg
ηidx=0. (17.48)
17.3 Several Dependent Variables 1053
Since the ηiare arbitrary and independent of one another,8each of the terms in the sum
mustvanish independently .W eha v e
∂f
∂yi−d
dx∂f
∂(∂yi/∂x)=0,i=1,2,...,n, (17.49)
awholesetof Eulerequations,eachof whichmustbesatisfiedfor anextremevalue.
Hamilton’s Principle
The most important application of Eq. (17.45) occurs when the integrand fis taken to
be a Lagrangian L. The Langrangian (for nonrelativistic systems; see Exercise 17.3.5 for
a relativistic particle) is defined as the difference of kinetic and potential energies of a
system:
L≡T−V. (17.50)
Using time as an independent variable instead of xandxi(t)as the dependent variables,
weget
x→t, y i→xi(t), y ix→˙xi(t);
xi(t)is the location and ˙xi=dxi/dtis the velocity of particle ias a function of time.
The equation δJ=0 is then a mathematicalstatement of Hamilton’sprinciple of classical
mechanics,
δintegraldisplayt2
t1L(x1,x2,...,xn,˙x1,˙x2,...,˙xn;t)dt=0. (17.51)
In words, Hamilton’s principle asserts that the motion of the system from time t1tot2
is such that the time integral of the Lagrangian L, or action, has a stationary value. The
resultingEulerequationsareusuallycalledtheLagrangianequationsof motion,
d
dt∂L
∂˙xi−∂L
∂xi=0. (17.52)
TheseLagrangianequationscanbederivedfromNewton’sequationsofmotion,andNew-
ton’s equations can be derived from Lagrange’s. The two sets of equations are equally
“fundamental.”
The Lagrangian formulation has advantages over the conventional Newtonian laws.
Whereas Newton’s equations are vector equations, we see that Lagrange’s equations in-
volveonlyscalarquantities.Thecoordinates x1,x2,...neednotbeanystandardsetofco-
ordinatesorlengths.Theycanbeselectedtomatchtheconditionsofthephysicalproblem.
TheLagrangeequationsareinvariantwithrespecttothechoiceofcoordinatesystem.New-
ton’s equations (in component form) are not manifestly invariant. Exercise 2.5.10 shows
whathappensto F=maresolvedinsphericalpolarcoordinates.
8For example, we could set η2=η3=η4=···=0, eliminating all but one term of the sum, and then treat η1exactly as in
Section 17.1.
1054 Chapter 17 Calculus of Variations
Exploitingtheconceptofenergy,wemayeasilyextendtheLagrangianformulationfrom
mechanicstodiversefields,suchaselectricalnetworksandacousticalsystems.Extensions
to electromagnetism appear in the exercises. The result is a unity of otherwise-separate
areas of physics. In the development of new areas, the quantization of Lagrangian particle
mechanics provided a model for the quantization of electromagnetic fields and led to the
gaugetheoryofquantumelectrodynamics.
One of the most valuable advantages of the Hamilton principle—Lagrange equation
formulation—istheeaseinseeingarelationbetweenasymmetryandaconservationlaw.
As an example, let xi=ϕ, an azimuthal angle. If our Lagrangian is independent of ϕ
(that is, if ϕis an ignorable coordinate), there are two consequences: (1) the conservation
or invariance of a component of angular momentum and (2) from Eq. (17.52) ∂L/∂˙ϕ=
constant.Similarly,invarianceundertranslationleadstoconservationoflinearmomentum.
Noether’stheoremisageneralizationofthisinvariance(symmetry)—theconservationlaw
relation.
Example 17.3.1 MOVING PARTICLE —C ARTESIAN COORDINATES
ConsiderEq.(17.50), whichdescribesoneparticlewithkineticenergy
T=1
2m˙x2(17.53)
and potential energy V(x), in which, as usual, the force is given by the negative gradient
ofthepotential,
F(x)=−dV(x)
dx. (17.54)
FromEq. (17.52),
d
dt(m˙x)−∂(T−V)
∂x=m¨x−F(x)=0, (17.55)
whichisNewton’ssecondlawofmotion. /squaresolid
Example 17.3.2 MOVING PARTICLE —C IRCULAR CYLINDRICAL COORDINATES
Now let us describe a moving particle in cylindrical coordinates of the xy-plane, that is,
z=0.Thekineticenergyis
T=1
2mparenleftbig
˙x2+˙y2parenrightbig
=1
2mparenleftbig
˙ρ2+ρ2˙ϕ2parenrightbig
, (17.56)
andwetake V=0 for simplicity.
The transformation of ˙x2+˙y2into circular cylindrical coordinates could be carried out
by taking x(ρ,ϕ)andy(ρ,ϕ), Eq. (2.28), and differentiating with respect to time and
squaring. It is much easier to interpret ˙x2+˙y2asv2and just write down the components
ofvasˆρ(dsρ/dt)=ˆρ˙ρ,andsoon.(The dsρisanincrementof length,ρchangingby dρ,
ϕremainingconstant.SeeSections2.1and2.4.)
TheLagrangianequationsyield
d
dt(m˙ρ)−mρ˙ϕ2=0,d
dtparenleftbig
mρ2˙ϕparenrightbig
=0. (17.57)
17.3 Several Dependent Variables 1055
Thesecondequationisastatementofconservationofangularmomentum.Thefirstmaybe
interpretedasradialacceleration9equatedtocentrifugalforce.Inthissensethecentrifugal
force is a real force. It is of some interest that this interpretation of centrifugal force as a
realforceissupportedbythegeneraltheoryofrelativity. /squaresolid
Exercises
17.3.1 (a) Developtheequationsofmotioncorrespondingto L=1
2m(˙x2+˙y2).
(b) Inwhatsense doyoursolutionsminimizetheintegralintegraltextt2
t1Ldt?
Comparetheresultfor yoursolutionwith x=const.,y=const.
17.3.2 From the Lagrangian equations of motion, Eq. (17.52), show that a system in stable
equilibriumhasa minimumpotentialenergy.
17.3.3 Write out the Lagrangian equations of motion of a particle in spherical coordinates for
potential Vequaltoaconstant.Identifythetermscorrespondingto(a)centrifugalforce
and(b) Coriolisforce.
17.3.4 The spherical pendulum consists of a mass on a wire of length l, free to move in polar
angleθandazimuthangle ϕ(Fig.17.9).
(a) SetuptheLagrangianfor thisphysicalsystem.
(b) DeveloptheLagrangianequationsofmotion.
17.3.5 ShowthattheLagrangian
L=m0c2parenleftbigg
1−radicalBigg
1−v2
c2parenrightbigg
−V(r)
FIGURE 17.9Spherical
pendulum.
9Hereis asecond method ofattackingExercise 2.4.8.
1056 Chapter 17 Calculus of Variations
leadstoarelativisticformof Newton’ssecondlawof motion,
d
dtparenleftbiggm0viradicalbig
1−v2/c2parenrightbigg
=Fi,
inwhichtheforcecomponentsare Fi=−∂V/∂xi.
17.3.6 The Lagrangian for a particle with charge qin an electromagnetic field described by
scalarpotential ϕandvectorpotential Ais
L=1
2mv2−qϕ+qA·v.
Findtheequationofmotionofthechargedparticle.
Hint.(d/dt)A j=∂Aj/∂t+summationtext
i(∂Aj/∂xi)˙xi.Thedependenceoftheforcefields Eand
Buponthepotentials ϕandAisdevelopedinSection1.13(compareExercise1.13.10).
ANS.m¨xi=q[E+v×B]i.
17.3.7 ConsiderasysteminwhichtheLagrangianisgivenby
L(qi,˙qi)=T(qi,˙qi)−V(qi),
whereqiand˙qirepresent sets of variables. The potential energy Vis independent of
velocityandneither TnorVhasanyexplicittimedependence.
(a) Showthat
d
dtparenleftbiggsummationdisplay
j˙qj∂L
∂˙qj−Lparenrightbigg
=0.
(b) Theconstantquantity
summationdisplay
j˙qj∂L
∂˙qj−L
defines the Hamiltonian H. Show that under the preceding assumed conditions,
H=T+V, thetotalenergy.
Note.Thekineticenergy Tis aquadraticfunctionofthe ˙qi.
17.4 S EVERAL INDEPENDENT VARIABLES
Sometimes the integrand fof Eq. (17.1) will contain one unknown function, u, that is a
function of several independent variables, u=u(x,y,z) , for the three-dimensional case,
forexample.Equation(17.1) becomes
J=integraldisplayintegraldisplayintegraldisplay
f[u,ux,uy,uz,x,y,z]dxdydz, (17.58)
ux=∂u/∂x,andsoon.Thevariationalproblemistofindthefunction u(x,y,z) forwhich
Jisstationary,
δJ=δα∂J
∂αvextendsinglevextendsinglevextendsinglevextendsingle
α=0=0. (17.59)
17.4 Several Independent Variables 1057
GeneralizingSection17.1,welet
u(x,y,z,α)=u(x,y,z, 0)+αη(x,y,z), (17.60)
whereu(x,y,z,α=0)represents the (unknown) function for which Eq. (17.59) is satis-
fied, whereas again η(x,y,z) is the arbitrary deviation that describes the varied function
u(x,y,z,α) . This deviation η(x,y,z) is required to be differentiable and to vanish at the
endpoints.ThenfromEq. (17.60),
ux(x,y,z,α)=ux(x,y,z,0)+αηx, (17.61)
andsimilarlyfor uyanduz.
Differentiating the integral Eq. (17.58) with respect to the parameter αand then setting
α=0,weobtain
∂J
∂αvextendsinglevextendsinglevextendsinglevextendsingle
α=0=integraldisplayintegraldisplayintegraldisplayparenleftbigg∂f
∂uη+∂f
∂uxηx+∂f
∂uyηy+∂f
∂uzηzparenrightbigg
dxdydz=0.(17.62)
Again, we integrate each of the terms (∂f/∂u i)ηiby parts. The integrated part vanishes
attheendpoints(becausethedeviation ηis requiredtogotozeroattheendpoints)and
integraldisplayintegraldisplayintegraldisplayparenleftbigg∂f
∂u−∂
∂x∂f
∂ux−∂
∂y∂f
∂uy−∂
∂z∂f
∂uzparenrightbigg
η(x,y,z)dxdydz =0.10(17.63)
Sincethevariation η(x,y,z) isarbitrary,theterminlargeparenthesesissetequaltozero.
ThisyieldstheEuler equationfor (three)independentvariables,
∂f
∂y−∂
∂x∂f
∂ux−∂
∂y∂f
∂uy−∂
∂z∂f
∂uz=0. (17.64)
Example 17.4.1 LAPLACE ’SEQUATION
Anexampleofthissortofvariationalproblemisprovidedbyelectrostatics.Theenergyof
anelectrostaticfieldis
energydensity =1
2εE2, (17.65)
inwhichEis theusualelectrostaticforcefield.Interms ofthestaticpotential ϕ,
energydensity =1
2ε(∇ϕ)2. (17.66)
Now let us impose the requirement that the electrostatic energy (associated with the field)
inagivenvolumebeaminimum.(Boundaryconditionson Eandϕmuststillbesatisfied.)
Wehavethevolumeintegral11
J=integraldisplayintegraldisplayintegraldisplay
(∇ϕ)2dxdydz=integraldisplayintegraldisplayintegraldisplayparenleftbig
ϕ2
x+ϕ2
y+ϕ2
zparenrightbig
dxdydz. (17.67)
10Recallthat ∂/∂xisapartialderivative,where yandzareheldconstant.But ∂/∂xalsoactson implicitx-dependenceaswell
as onexplicitx-dependence. In this sense,for example,
∂
∂xparenleftbigg∂f
∂uxparenrightbigg
=∂2f
∂x∂ux+∂2f
∂u∂uxux+∂2f
∂u2xuxx+∂2f
∂uy∂uxuxy+∂2f
∂uz∂uxuxz.
11Thesubscript xindicatesthe x-partial derivative, not an x-component.
1058 Chapter 17 Calculus of Variations
With
f(ϕ,ϕx,ϕy,ϕz,x,y,z)=ϕ2
x+ϕ2
y+ϕ2
z, (17.68)
thefunction ϕreplacingthe uofEq. (17.64), Euler’sequation(Eq. (17.64)) yields
−2(ϕxx+ϕyy+ϕzz)=0, (17.69)
or
∇2ϕ(x,y,z)=0, (17.70)
whichisLaplace’sequationof electrostatics.
Closer investigation shows that this stationary value is indeed a minimum. Thus the
demandthatthefieldenergybeminimizedleadstoLaplace’sPDE. /squaresolid
Exercises
17.4.1 TheLagrangianfor avibratingstring(small-amplitudevibrations)is
L=integraldisplayparenleftbig1
2ρu2
t−1
2τu2
xparenrightbig
dx,
whereρis the (constant) linear mass density and τis the (constant) tension. The x-
integrationisoverthelengthofthestring.ShowthatapplicationofHamilton’sprinciple
totheLagrangiandensity(theintegrand),nowwithtwoindependentvariables,leadsto
theclassicalwaveequation
∂2u
∂x2=ρ
τ∂2u
∂t2.
17.4.2 Show that the stationary value of the total energy of the electrostatic field of Exam-
ple17.4.1isa minimum .
Hint.UseEq.(17.61) andinvestigatethe α2terms.
17.5 S EVERAL DEPENDENT AND INDEPENDENT VARIABLES
In some cases our integrand fcontains more than one dependent variable and more than
oneindependentvariable.Consider
f=fbracketleftbig
p(x,y,z),p x,py,pz,q(x,y,z),q x,qy,qz,r(x,y,z),r x,ry,rz,x,y,zbracketrightbig
.(17.71)
Weproceedas beforewith
p(x,y,z,α)=p(x,y,z, 0)+αξ(x,y,z),
q(x,y,z,α)=q(x,y,z, 0)+αη(x,y,z), (17.72)
r(x,y,z,α)=r(x,y,z, 0)+αζ(x,y,z), andso on .
17.5 Several Dependent and Independent Variables 1059
Keeping in mind that ξ,η, andζare independent of one another, as were the ηiin Sec-
tion17.3, thesamedifferentiationandthenintegrationbyparts leadsto
∂f
∂p−∂
∂x∂f
∂px−∂
∂y∂f
∂py−∂
∂z∂f
∂pz=0, (17.73)
with similar equations for functions qandr. Replacing p,q,r,... withyiandx,y,z,...
withxy, wecanputEq. (17.73)inamorecompactform:
∂f
∂yi−summationdisplay
j∂
∂xjparenleftbigg∂f
∂yijparenrightbigg
=0,i=1,2,..., (17.73a)
inwhich
yij≡∂yi
∂xj.
Anapplicationof Eq.(17.73) appearsinSection17.7.
Relation to Physics
The calculus of variations as developed so far provides an elegant description of a wide
variety of physical phenomena. The physics includes classical mechanics in Section 17.3;
relativistic mechanics, Exercise 17.3.5; electrostatics, Example 17.4.1; and electromag-
netictheoryinExercise17.5.1.Theconvenienceshouldnotbeminimized,butatthesame
timeweshouldbeawarethatinthesecasesthecalculusofvariationshasonlyprovidedan
alternate description of what was already known. The situation does change with incom-
pletetheories.
•If the basic physics is not yet known, a postulated variational principle can be a useful
startingpoint.
Exercise
17.5.1 TheLagrangian(perunitvolume)ofanelectromagneticfieldwithachargedensity ρis
givenby
L=1
2parenleftbigg
ε0E2−1
µ0B2parenrightbigg
−ρϕ+ρv·A.
Show that Lagrange’s equations lead to two of Maxwell’s equations. (The remaining
twoareaconsequenceofthedefinitionof EandBintermsof Aandϕ.)ThisLagrange
densitycomesfromascalarexpressioninSection4.6.
Hint.T a k eA1,A2,A3, andϕasdependent variables, x,y,z, andtasindependent
variables. EandBare givenintermsof AandϕbyEq. (4.142)andEq. (1.88).
1060 Chapter 17 Calculus of Variations
17.6 L AGRANGIAN MULTIPLIERS
In this section the concept of a constraint is introduced. To simplify the treatment, the
constraint appears as a simple function rather than as an integral. In this section we are
not concerned with the calculus of variations, but in Section 17.7 the constraints, with our
newlydevelopedLagrangianmultipliers,are incorporatedintothecalculusofvariations.
Consider a function of three independent variables, f(x,y,z) . For the function fto be
amaximum(or extreme),12
df=0. (17.74)
Thenecessaryandsufficientconditionfor thisis
∂f
∂x=∂f
∂y=∂f
∂z=0, (17.75)
inwhich
df=∂f
∂xdx+∂f
∂ydy+∂f
∂zdz. (17.76)
Often in physical problems the variables x,y,zare subject to constraints so that they
are no longer all independent. It is possible, at least in principle, to use each constraint to
eliminateonevariableandtoproceedwithanewandsmallersetofindependentvariables.
The use of Lagrangian multipliers is an alternate technique that may be applied when
thiseliminationofvariablesisinconvenientorundesirable.Letour equationofconstraint
be
ϕ(x,y,z)=0, (17.77)
from which z(x,y)may be extracted if x,yare taken as the independent coordinates.
Returning to Eq. (17.74), Eq. (17.75) no longer follows because there are now only two
independent variables, so dzis no longer arbitrary. From the total differential dϕ=0,we
thenobtain
−∂ϕ
∂zdz=∂ϕ
∂xdx+∂ϕ
∂ydy (17.78)
andtherefore
df=∂f
∂xdx+∂f
∂ydy+λparenleftbigg∂ϕ
∂xdx+∂ϕ
∂xdxparenrightbigg
,λ=−fz
ϕz,
assuming ϕz=∂ϕ
∂z/negationslash=0.Thus, we may add Eq. (17.76) and a multiple of Eq. (17.78) to
obtain
df+λdϕ=parenleftbigg∂f
∂x+λ∂ϕ
∂xparenrightbigg
dx+parenleftbigg∂f
∂y+λ∂ϕ
∂yparenrightbigg
dy+parenleftbigg∂f
∂z+λ∂ϕ
∂zparenrightbigg
dz=0.(17.79)
Inotherwords, ourLagrangianmultiplier λischosensothat
∂f
∂z+λ∂ϕ
∂z=0, (17.80)
12Including asaddle point.
17.6 Lagrangian Multipliers 1061
assumingthat ∂ϕ/∂z/negationslash=0.Equation(17.79)nowbecomes
parenleftbigg∂f
∂x+λ∂ϕ
∂xparenrightbigg
dx+parenleftbigg∂f
∂y+λ∂ϕ
∂yparenrightbigg
dy=0. (17.81)
However,now dxanddyarearbitraryandthequantitiesinparenthesesmustvanish:
∂f
∂x+λ∂ϕ
∂x=0,∂f
∂y+λ∂ϕ
∂y=0. (17.82)
When Eqs. (17.80) and (17.82) are satisfied, df=0 andfis an extremum. Notice that
there are now four unknowns: x,y,z, andλ. The fourth equation is, of course, the con-
straintEq.(17.77).Wewantonly x,y,andz,soλneednotbedetermined.Forthisreason
λis sometimes called Lagrange’s undetermined multiplier . This method will fail if all
the coefficients of λvanish at the extremum, ∂ϕ/∂x,∂ϕ/∂y,∂ϕ/∂z=0. It is then impos-
s ibletos olv ef or λ.
NotethatfromtheformofEqs.(17.80)and(17.82),wecouldidentify fasthefunction
takinganextremevaluesubjectto ϕ,theconstraint,orwecouldidentify fastheconstraint
andϕas thefunction.
If wehavea set ofconstraints ϕk, thenEqs. (17.80)and(17.82)become
∂f
∂xi+summationdisplay
kλk∂ϕk
∂xi=0,i=1,2,...,n,
withaseparateLagrangemultiplier λkfor eachϕk.
Example 17.6.1 PARTICLE IN A BOX
As an example of the use of Lagrangian multipliers, consider the quantum mechanical
problemofaparticle(mass m)inabox.Theboxisarectangularparallelepipedwithsides
a,b,andc.Theground-stateenergyof theparticleis givenby
E=h2
8mparenleftbigg1
a2+1
b2+1
c2parenrightbigg
. (17.83)
We seektheshapeoftheboxthatwillminimizetheenergy E,subjecttoconstraintthat
thevolumeis constant,
V(a,b,c)=abc=k. (17.84)
Withf(a,b,c)=E(a,b,c) andϕ(a,b,c)=abc−k=0,weobtain
∂E
∂a+λ∂ϕ
∂a=−h2
4ma3+λbc=0. (17.85)
Also,
−h2
4mb3+λac=0,−h2
4mc3+λab=0.
1062 Chapter 17 Calculus of Variations
Multiplying the first of these expressions by a, the second by b, and the third by c,w e
have
λabc=h2
4ma2=h2
4mb2=h2
4mc2. (17.86)
Thereforeoursolutionis
a=b=c,acube. (17.87)
Noticethat λhas notbeendeterminedbutfollowsfromEq. (17.86). /squaresolid
Example 17.6.2 CYLINDRICAL NUCLEAR REACTOR
A further example is provided by the nuclear reactor theory. Suppose a (thermal) nuclear
reactor is to have the shape of a right circular cylinder of radius Rand height H. Neutron
diffusiontheorysuppliesaconstraint:
ϕ(R,H)=parenleftbigg2.4048
Rparenrightbigg2
+parenleftbiggπ
Hparenrightbigg2
=constant.13(17.88)
Wewishtominimizethevolumeofthereactorvessel,
f(R,H)=πR2H. (17.89)
ApplicationofEq. (17.82)leadsto
∂f
∂R+λ∂ϕ
∂R=2πRH−2λ(2.4048)2
R3=0,
∂f
∂H+λ∂ϕ
∂H=πR2−2λπ2
H3=0. (17.90)
Bymultiplyingthefirstof theseequationsby R/2 andthesecondby H, weobtain
πR2H=λ(2.4048)2
R2=λ2π2
H2, (17.91)
orheight
H=√
2πR
2.4048=1.847R, (17.92)
fortheminimum-volumeright-circularcylindricalreactor.
Strictly speaking, we have found only an extremum. Its identification as a minimum
followsfromaconsiderationoftheoriginalequations. /squaresolid
132.4048...is the lowest root of Besselfunction J0(R)(compare Section 11.1).
17.6 Lagrangian Multipliers 1063
Exercises
ThefollowingproblemsaretobesolvedbyusingLagrangianmultipliers.
17.6.1 The ground-state energy of a quantum particle of mass min a pillbox (right-circular
cylinder)is givenby
E=¯h2
2mparenleftbigg(2.4048)2
R2+π2
H2parenrightbigg
,
inwhich Ristheradiusand Histheheightofthepillbox.Findtheratioof RtoHthat
willminimizetheenergyfor afixedvolume.
17.6.2 Find the ratio of R(radius) to H(height) that will minimize the total surface area of a
right-circularcylinderof fixedvolume.
17.6.3 TheU.S.PostOfficelimitsfirstclassmailtoCanadatoatotalof36inches,lengthplus
girth. Using a Lagrange multiplier, find the maximum volume and the dimensions of a
(rectangularparallelepiped)packagesubjecttothisconstraint.
17.6.4 Athermalnuclearreactoris subjecttotheconstraint
ϕ(a,b,c)=parenleftbiggπ
aparenrightbigg2
+parenleftbiggπ
bparenrightbigg2
+parenleftbiggπ
cparenrightbigg2
=B2,a constant .
Findtheratiosofthesidesoftherectangularparallelepipedreactorofminimumvolume.
ANS.a=b=c,cube.
17.6.5 For a lens of focal length f, the object distance pand the image distance qare related
by 1/p+1/q=1/f. Find the minimum object–image distance (p+q)for fixed f.
Assumerealobjectandimage( pandqbothpositive).
17.6.6 You have an ellipse (x/a)2+(y/b)2=1. Find the inscribed rectangle of maximum-
area. Show that the ratio of the area of the maximum-area rectangle to the area of the
ellipseis 2 /π=0.6366.
17.6.7 A rectangular parallelepiped is inscribed in an ellipsoid of semiaxes a,b, andc. Maxi-
mize the volume of the inscribed rectangular parallelepiped. Show that the ratio of the
maximumvolumetothevolumeof theellipsoidis 2 /π√
3≈0.367.
17.6.8 Adeformed sphere has a radius given by r=r0{α0+α2P2(cosθ)}, whereα0≈1 and
|α2|≪|α0|. FromExercise12.5.16theareaandvolumeare
A=4πr2
0α2
0braceleftbigg
1+4
5parenleftbiggα2
α0parenrightbigg2bracerightbigg
,V=4πr3
0
3a3
0braceleftbigg
1+3
5parenleftbiggα2
α0parenrightbigg2bracerightbigg
.
Terms oforder α3
2havebeenneglected.
(a) Withtheconstraintthattheenclosedvolumebeheldconstant,thatis, V=4πr3
0/3,
showthattheboundingsurface ofminimumareaisasphere( α0=1,α2=0).
(b) With the constraint that the area of the bounding surface be held constant, that
is,A=4πr2
0, show that the enclosed volume is a maximum when the surface is
asphere.
1064 Chapter 17 Calculus of Variations
17.6.9 Findthemaximumvalueofthedirectionalderivativeof ϕ(x,y,z) ,
dϕ
ds=∂ϕ
∂xcosα+∂ϕ
∂ycosβ+∂ϕ
∂zcosγ,
subjecttotheconstraint
cos2α+cos2β+cos2γ=1.
ANS.parenleftbiggdϕ
dsparenrightbigg
=|∇ϕ|.
Note concerning the following exercises: In a quantum mechanical system there are gi
distinct quantum states between energies EiandEi+dEi. The problem is to describe
howniparticlesaredistributedamongthesestatessubjecttotwoconstraints:
(a) fixednumberofparticles,
summationdisplay
ini=n.
(b) fixedtotalenergy,
summationdisplay
iniEi=E.
17.6.10 For identical particles obeying the Pauli exclusion principle, the probability of a given
arrangementis
WFD=productdisplay
igi!
ni!(gi−ni)!.
Show that maximizing WFD, subject to a fixed number of particles and fixed total en-
ergy,leadsto
ni=gi
eλ1+λ2Ei+1.
Withλ1=−E0/kTandλ2=1/kT, thisyieldsFermi–Diracstatistics.
Hint.Tryworkingwithln WandusingStirling’sformula,Section8.3.Thejustification
fordifferentiation withrespectto niisthatwearedealingherewithalargenumberof
particles, /Delta1ni/ni≪1.
17.6.11 For identical particles but no restriction on the number in a given state, the probability
ofagivenarrangementis
WBE=productdisplay
i(ni+gi−1)!
ni!(gi−1)!.
Show that maximizing WBE, subject to a fixed number of particles and fixed total en-
ergy,leadsto
ni=gi
eλ1+λ2Ei−1.
Withλ1=−E0/kTandλ2=1/kT, thisyieldsBose–Einsteinstatistics.
Note.Assumethat gi≫1.
17.7 Variation with Constraints 1065
17.6.12 Photonssatisfy WBEandtheconstraintthattotalenergyisconstant.Theyclearlydo not
satisfy the fixed-number constraint. Show that eliminating the fixed-number constraint
leadstotheforegoingresultbutwith λ1=0.
17.7 V ARIATION WITH CONSTRAINTS
Asintheprecedingsections,weseekthepaththatwillmaketheintegral
J=integraldisplay
fparenleftbigg
yi,∂yi
∂xj,xjparenrightbigg
dxj (17.93)
stationary. This is the general case in which xjrepresents a set of independent variables
andyiasetof dependentvariables.Again,
δJ=0. (17.94)
Now,however,we introduceone or moreconstraints. This meansthat the yiare no longer
independent of each other. Not all the ηimay be varied arbitrarily, and Eqs. (17.62)
and(17.73a)wouldnotapply.Theconstraintmayhavetheform
ϕk(yi,xj)=0, (17.95)
as in Section 17.6. In this case we may multiply by a function of xj,s a y ,λk(xj), and
integrateoverthesamerangeasinEq. (17.93)toobtain
integraldisplay
λk(xj)ϕk(yi,xj)dxj=0. (17.96)
Thenclearly
δintegraldisplay
λk(xj)ϕk(yi,xj)dxj=0. (17.97)
Alternatively,theconstraintmayappearintheformofanintegral
integraldisplay
ϕk(yi,∂yi/∂xj,xj)dxj=constant. (17.98)
We may introduce any constant Lagrangian multiplier, and again Eq. (17.97) follows—
nowwith λaconstant.
In either case, by adding Eqs. (17.94) and (17.97), possibly with more than one con-
straint,weobtain
δintegraldisplaybracketleftbigg
fparenleftbigg
yi,∂yi
∂xj,xjparenrightbigg
+summationdisplay
kλkϕk(yi,xj)bracketrightbigg
dxj=0. (17.99)
The Lagrangian multiplier λkmay depend on xjwhenϕ(yi,xj)is given in the form of
Eq.(17.95).
Treatingtheentireintegrandas anewfunction,
gparenleftbigg
yi,∂yi
∂xj,xjparenrightbigg
,
1066 Chapter 17 Calculus of Variations
weobtain
gparenleftbigg
yi,∂yi
∂xj,xjparenrightbigg
=f+summationdisplay
kλkϕk. (17.100)
If we have Nyi(i=1,2,...,N)andmconstraints (k=1,2,...m), thenN−mof
theηimay be taken as arbitrary. For the remaining mηi,t h eλmay, in principle, be cho-
sen so that the remaining Euler–Lagrange equations are satisfied, completely analogous
to Eq. (17.80). The result is that our composite function gmust satisfy the usual Euler–
Lagrangeequations,
∂g
∂yi−summationdisplay
j∂
∂xj∂g
(∂yi/∂xj)=0, (17.101)
withonesuchequationforeachdependentvariable yi(compareEqs.(17.64)and(17.73)).
These Euler equations and the equations of constraint are then solved simultaneously to
findthefunctionyieldingastationaryvalue.
Lagrangian Equations
In the absence of constraints, Lagrange’s equations of motion (Eq. (17.52)) were found to
be14
d
dt∂L
∂˙qi−∂L
∂qi=0,
witht(time)theoneindependentvariableand qi(t)(particlepositions)asetofdependent
variables.Usuallythegeneralizedcoordinates qiarechosentoeliminatetheforcesofcon-
straint, but this is not necessary and not always desirable. In the presence of (holonomic)
constraints, ϕk=0,Hamilton’sprincipleis
δintegraldisplaybracketleftbigg
L(qi,˙qi,t)+summationdisplay
kλk(t)ϕk(qi,t)bracketrightbigg
dt=0, (17.102)
andtheconstrainedLagrangianequationsofmotionare
d
dt∂L
∂˙qi−∂L
∂qi=summationdisplay
kaikλk. (17.103)
Usuallyϕk=ϕk(qi,t), independent of the generalized velocities ˙qi. In this case the coef-
ficientaikis givenby
aik=∂ϕk
∂qi. (17.104)
Thenaikλk(no summation) represents the force of the kth constraint in the qi-direction,
appearinginEq. (17.103)inexactlythesamewayas −∂V/∂qi.
14The symbol qis customary in classical mechanics. It serves to emphasize that the variable is not necessarily a Cartesian
variable (and not necessarily alength).
17.7 Variation with Constraints 1067
FIGURE 17.10
Simplependulum.
Example 17.7.1 SIMPLE PENDULUM
To illustrate, consider the simple pendulum, a mass mconstrained by a wire of length lto
swinginanarc(Fig. 17.10).Intheabsenceof theoneconstraint
ϕ1=r−l=0 (17.105)
there are two generalized coordinates randθ(motion in vertical plane). The Lagrangian
is
L=T−V=1
2mparenleftbig
˙r2+r2˙θ2parenrightbig
+mgrcosθ, (17.106)
taking the potential Vto be zero when the pendulum is horizontal, θ=π/2. By
Eq.(17.103)theequationsofmotionare
d
dt∂L
∂˙r−∂L
∂r=λ1,d
dt∂L
∂˙θ−∂L
∂θ=0(ar1=1,aθ1=0),(17.107)
or
d
dt(m˙r)−mr˙θ2−mgcosθ=λ1,
d
dtparenleftbig
mr2˙θparenrightbig
+mgrsinθ=0. (17.108)
Substitutingintheequationof constraint (r=l,˙r=0),weha v e
ml˙θ2+mgcosθ=−λ1,ml2¨θ+mglsinθ=0. (17.109)
The second equation may be solved for θ(t)to yield simple harmonic motion if the am-
plitude is small (sinθ∼θ), whereas the first equation expresses the tension in the wire in
termsofθand˙θ.
Notethatsincetheequationofconstraint,Eq.(17.105),isintheformofEq.(17.95),the
Lagrangemultiplier λmaybe(andhereis) afunctionof t(orofθ). /squaresolid
1068 Chapter 17 Calculus of Variations
FIGURE 17.11Aparticle
slidingonacylindrical
surface.
Example 17.7.2 SLIDING OFF A LOG
Closely related to this is the problem of a particle sliding on a cylindrical surface. The
object is to find the critical angle θcat which the particle flies off from the surface. This
critical angle is the angle at which the radial force of constraint goes to zero (Fig. 17.11).
We have
L=T−V=1
2mparenleftbig
˙r2+r2˙θ2parenrightbig
−mgrcosθ (17.110)
andtheoneequationof constraint,
ϕ1=r−l=0. (17.111)
Proceedingas inExample17.7.1with ar1=1,
m¨r−mr˙θ2+mgcosθ=λ1(θ),
mr2¨θ+2mr˙r˙θ−mgrsinθ=0, (17.112)
inwhichtheconstrainingforce λ1(θ)isafunctionoftheangle θ.15Sincer=l,¨r=˙r=0,
Eq.(17.112)reducesto
−ml˙θ2+mgcosθ=λ1(θ), (17.113a)
ml2¨θ−mglsinθ=0. (17.113b)
DifferentiatingEq.(17.113a)withrespecttotimeandrememberingthat
df(θ)
dt=df(θ)
dθ˙θ, (17.114)
weobtain
−2ml¨θ−mgsinθ=dλ1(θ)
dθ. (17.115)
15Note that λ1is theradialforce exerted by the cylinder on the particle. Consideration of the physical problem shows that
λ1must depend on the angle θ. We permitted λ=λ(t). Now we are replacing the time dependence by an (unknown) angular
dependence using θ=θ(t).
17.7 Variation with Constraints 1069
UsingEq. (17.113b)toeliminatethe ¨θtermandthenintegrating,wehave
λ1(θ)=3mgcosθ+C. (17.116)
Since
λ1(0)=mg, (17.117)
C=−2mg. (17.118)
The particle mwill stay on the surface as long as the force of constraint is nonnegative,
thatis, as longasthesurface hastopushoutwardontheparticle:
λ1(θ)=3mgcosθ−2mg≥0. (17.119)
The critical angle lies where λ1(θc)=0, the force of constraint going to zero. From
Eq.(17.119),
cosθc=2
3,orθc=48◦11′(17.120)
fromthevertical.Atthisangle(neglectingallfriction)ourparticletakesoff.
Itmustbeadmittedthatthisresultcanbeobtainedmoreeasilybyconsideringavarying
centripetalforcefurnishedbytheradialcomponentofthegravitationalforce.Theexample
was chosen to illustrate the use of Lagrange’s undetermined multiplier without confusing
thereaderwithacomplicatedphysicalsystem. /squaresolid
Example 17.7.3 THESCHRÖDINGER WAVEEQUATION
Asafinalillustrationofaconstrainedminimum,letusfindtheEulerequationsforaquan-
tummechanicalproblem
δintegraldisplayintegraldisplayintegraldisplay
ψ∗(x,y,z)Hψ(x,y,z)dxdydz =0, (17.121)
withthenormalizationconstraintintegraldisplayintegraldisplayintegraldisplay
ψ∗ψdxdydz=1. (17.122)
Equation (17.121) is a statement that the energy of the system is stationary, Hbeing the
quantummechanicalHamiltonianfor aparticleofmass m, adifferentialoperator,
H=−¯h2
2m∇2+V(x,y,z). (17.123)
Equation (17.122) is a bound-state constraint, ψis the usual wave function, a dependent
variable,and ψ∗, itscomplexconjugate,istreatedas a second16dependentvariable.
The integrand in Eq. (17.121) involves secondderivatives, which can be converted to
firstderivativesbyintegratingbyparts:
integraldisplay
ψ∗∂2ψ
∂x2dx=ψ∗∂ψ
∂xvextendsinglevextendsinglevextendsinglevextendsingle−integraldisplay∂ψ∗
∂x∂ψ
∂xdx. (17.124)
We assume either periodic boundary conditions (as in the Sturm–Liouville theory,
Chapter10) or that the volume of integration is so large that ψandψ∗vanish rapidly
16Compare Section 6.1.
1070 Chapter 17 Calculus of Variations
enough17at the boundary. Then the integrated part vanishes and Eq. (17.121) may be
rewrittenas
δintegraldisplayintegraldisplayintegraldisplaybracketleftbigg
¯h2
2m∇ψ∗·∇ψ+Vψ∗ψbracketrightbigg
dxdydz=0. (17.125)
Thefunction gof Eq.(17.100)is
g=¯h2
2m∇ψ∗·∇ψ+Vψ∗ψ−λψ∗ψ
=¯h2
2m(ψ∗
xψx+ψ∗
yψy+ψ∗
zψz)+Vψ∗ψ−λψ∗ψ, (17.126)
againusingthesubscript xtodenote ∂/∂x.F o ryi=ψ∗, Eq. (17.101)becomes
∂g
∂ψ∗−∂
∂x∂g
∂ψ∗x−∂
∂y∂g
∂ψ∗y−∂
∂z∂g
∂ψ∗z=0.
Thisyields
Vψ−λψ−¯h2
2m(ψxx+ψyy+ψzz)=0,
or
−¯h2
2m∇2ψ+Vψ=λψ. (17.127)
ReferencetoEq.(17.123)enablesustoidentify λphysicallyastheenergyofthequantum
mechanical system. With this interpretation, Eq. (17.127) is the celebrated Schrödinger
waveequation. /squaresolid
This variational approach is more than just a matter of academic curiosity. It provides a
verypowerfulmethodofobtainingapproximatesolutionsofthewaveequation(Rayleigh–
Ritzvariationalmethod,Section17.8).
Exercises
17.7.1 A particle, mass m, is on a frictionless horizontal surface. It is constrained to move so
thatθ=ωt(rotatingradialarm,nofriction).Withtheinitialconditions
t=0,r=r0,˙r=0,
(a) findtheradialpositionsas afunctionoftime.
ANS.r(t)=r0coshωt.
(b) findtheforceexertedontheparticlebytheconstraint.
ANS.F(c)=2m˙rω=2mr0ω2sinhωt.
17For example, lim r→∞rψ(r)=0.
17.7 Variation with Constraints 1071
17.7.2 A point mass mis moving over a flat, horizontal, frictionless plane. The mass is con-
strained by a string to move radially inward at a constant rate. Using plane polar coor-
dinates(ρ,ϕ),ρ=ρ0−kt,
(a) SetuptheLagrangian.
(b) ObtaintheconstrainedLagrangeequations.
(c) Solve the ϕ-dependent Lagrange equation to obtain ω(t), the angular velocity.
What is the physical significance of the constant of integration that you get from
your“free”integration?
(d) Usingthe ω(t)frompart(b),solvethe ρ-dependent(constrained)Lagrangeequa-
tion to obtain λ(t). In other words, explain what is happening to the forceof con-
straintas ρ→0.
17.7.3 A flexible cable is suspended from two fixed points. The length of the cable is fixed.
Findthecurvethatwillminimizethetotalgravitationalpotentialenergyof thecable.
ANS. Hyperboliccosine.
17.7.4 Afixedvolumeofwaterisrotatinginacylinderwithconstantangularvelocity ω.Find
the curve of the water surface that will minimize the total potential energy of the water
inthecombinedgravitational-centrifugalforcefield.
ANS.Parabola.
17.7.5 (a) Showthatforafixed-lengthperimeterthefigurewithmaximumareaisacircle.
(b) Showthatforafixedareathecurvewithminimumperimeterisacircle.
Hint.Theradiusofcurvature Ris givenby
R=(r2+r2
θ)3/2
rrθθ−2r2
θ−r2.
Note. The problems of this section, variation subject to constraints, are often called
isoperimetric . The term arose from problems of maximizing area subject to a fixed
perimeter—asinExercise17.7.5(a).
17.7.6 Showthatrequiring J,givenby
J=integraldisplayb
abracketleftbig
p(x)y2
x−q(x)y2bracketrightbig
dx,
tohaveastationaryvaluesubjecttothenormalizingcondition
integraldisplayb
ay2w(x)dx=1
leadstotheSturm–LiouvilleequationofChapter10:
d
dxparenleftbigg
pdy
dxparenrightbigg
+qy+λwy=0.
Note.Theboundarycondition
pyxyvextendsinglevextendsinglevextendsingleb
a=0
isusedinSection10.1inestablishingtheHermitianpropertyoftheoperator.
1072 Chapter 17 Calculus of Variations
17.7.7 Showthatrequiring J,givenby
J=integraldisplayb
aintegraldisplayb
aK(x,t)ϕ(x)ϕ(t)dxdt,
tohaveastationaryvaluesubjecttothenormalizingcondition
integraldisplayb
aϕ2(x)dx=1
leadstotheHilbert–Schmidtintegralequation,Eq.(16.89).
Note.Thekernel K(x,t)issymmetric.
17.8 R AYLEIGH –RITZVARIATIONAL TECHNIQUE
Exercise 17.7.6 opens up a relation between the calculus of variations and eigenfunction–
eigenvalueproblems.Wemayrewritetheexpressionof Exercise17.7.6as
Fbracketleftbig
y(x)bracketrightbig
=integraltextb
a(py2
x−qy2)dx
integraltextb
ay2wdx, (17.128)
in which the constraint appears in the denominator as a normalizing condition. After the
unconstrained minimum of Fhas been found, ycan be normalized without changing the
stationary value of Fbecause stationary values of Jcorrespond to stationary values of F.
ThenfromExercise17.7.6,when y(x)issuchthat JandFtakeonastationaryvalue,the
optimumfunction y(x)satisfiestheSturm–Liouvilleequation
d
dxparenleftbigg
pdy
dxparenrightbigg
+qy+λwy=0, (17.129)
withλtheeigenvalue( notaLagrangianmultiplier).Integratingthefirstterminthenumer-
atorofEq. (17.128)bypartsandusingthe boundarycondition ,
pyxyvextendsinglevextendsinglevextendsingleb
a=0, (17.130)
weobtain
Fbracketleftbig
y(x)bracketrightbig
=−integraldisplayb
aybraceleftbiggd
dxparenleftbigg
pdy
dxparenrightbigg
+qybracerightbigg
dxslashBigintegraldisplayb
ay2wdx. (17.131)
ThensubstitutinginEq. (17.129), thestationaryvaluesof F[y(x)]aregivenby
Fbracketleftbig
y(x)bracketrightbig
=λn, (17.132)
withλnthe eigenvalue corresponding to the eigenfunction yn. Equation (17.132) with F
given by either Eq. (17.128) or Eq. (17.131) forms the basis of the Rayleigh–Ritz method
forthecomputationofeigenfunctionsandeigenvalues.
17.8 Rayleigh–Ritz Variational Technique 1073
Ground State Eigenfunction
Suppose that we seek to compute the ground-state eigenfunction y0and eigenvalue18λ0
of some complicated atomic or nuclear system. The classical example, for which no exact
solutionexists,istheheliumatomproblem.Theeigenfunction y0isunknown ,butweshall
assumewecanmakeaprettygoodguessatanapproximatefunction y,somathematically
wemaywrite19
y=y0+∞summationdisplay
i=1ciyi. (17.133)
Theciare small quantities. (How small depends on how good our guess was.) The yiare
orthonormalized eigenfunctions (also unknown), and therefore our trial function yis not
normalized.
Substitutingtheapproximatefunction yintoEq.(17.131)andnotingthat
integraldisplayb
ayibraceleftbiggd
dxparenleftbigg
pdyj
dxparenrightbigg
+qyibracerightbigg
dx=−λiδij, (17.134)
Fbracketleftbig
y(x)bracketrightbig
=λ0+summationtext∞
i=1c2
iλi
1+summationtext∞
i=1c2
i. (17.135)
Herewehavetakentheeigenfunctionstobeorthonormal—sincetheyaresolutionsofthe
Sturm–Liouvilleequation,Eq.(17.129).Wealsoassumethat y0isnondegenerate.Now,if
wereplacesummationtext
ic2
iλi→summationtext
ic2
iλ0+summationtext
ic2
i(λi−λ0)weobtain
Fbracketleftbig
y(x)bracketrightbig
=λ0+summationtext∞
i=1c2
i(λi−λ0)
1+summationtext∞
i=1c2
i. (17.136)
Equation(17.136)containstwoimportantresults.
•Whereastheerrorintheeigenfunction ywasO(ci),theerrorin λisonly O(c2
i).Ev en
a poor approximation of the eigenfunctions may yield an accurate calculation of the
eigenvalue.
•Ifλ0is thelowesteigenvalue(groundstate),thensince λi−λ0>0,
Fbracketleftbig
y(x)bracketrightbig
=λ≥λ0, (17.137)
or our approximation is always on the high side and becoming lower, converging on
λ0as our approximate eigenfunction yimproves (ci→0). Note that Eq. (17.137)
is a direct consequence of Eq. (17.135). More directly, F[y(x)]in Eq. (17.135) is
a positively weighted average of the λiand, therefore, must be no smaller than the
smallestλi, to wit,λ0. In practical problems in quantum mechanics, yoften depends
on parameters that may be varied to minimize Fand thereby improve the estimate
of the ground-state energy λ0. This is the “variational method” discussed in quantum
mechanicstexts.
18Thismeansthat λ0isthelowesteigenvalue.ItisclearfromEq.(17.128) thatif p(x)≥0a n dq(x)≤0 (compareTable10.1),
thenF[y(x)]has alowerbound and this lowerbound is nonnegative. Recallfrom Section 10.1 that w(x)≥0.
19Weareguessing atthe form of the function. Thenormalization is irrelevant.
1074 Chapter 17 Calculus of Variations
Example 17.8.1 VIBRATING STRING
Avibratingstring,clampedat x=0 and1,satisfiestheeigenvalueequation
d2y
dx2+λy=0 (17.138)
andtheboundarycondition y(0)=y(1)=0.Forthissimpleexamplewerecognizeimme-
diately that y0(x)=sinπx(unnormalized) and λ0=π2. But let us try out the Rayleigh–
Ritztechnique.
Withoneeyeontheboundaryconditions,wetry
y(x)=x(1−x). (17.139)
Thenwith p=1 andw=1,Eq. (17.128)yields
F[y(x)]=integraltext1
0(1−2x)2dx
integraltext1
0x2(1−x)2dx=1/3
1/30=10. (17.140)
This result, λ=10, is a fairly good approximation (1.3% error)20ofλ0=π2=9.8696.
You may have noted that y(x), Eq. (17.139), is not normalized to unity. The denominator
inF[y(x)]compensatesforthelackofunitnormalization. Fmayalsobecalculatedfrom
Eq.(17.131)sinceEq. (17.130)issatisfiedby yfrom Eq.(17.139).
In the usual scientific calculation the eigenfunction would be improved by introducing
moreterms andadjustableparameters,suchas
y=x(1−x)+a2x2(1−x)2. (17.141)
Itisconvenienttohavetheadditionaltermsorthogonal,butitisnotnecessary.Theparame-
tera2isadjustedto minimize F[y(x)].Inthiscase,choosing a2=1.1353 drives F[y(x)]
downto9.8697,veryclosetothecorrecteigenvaluevalue. /squaresolid
Exercises
17.8.1 From Eq. (17.128) develop in detail the argument when λ≥0o rλ<0. Explain the
circumstancesunderwhich λ=0,andillustratewithseveralexamples.
17.8.2 Anunknownfunctionsatisfies thedifferentialequation
y′′+parenleftbiggπ
2parenrightbigg2
y=0
andtheboundaryconditions
y(0)=1,y(1)=0.
20The closeness of the fit may be checked by a Fourier sine expansion (compare Exercise 14.2.3 over the half-interval [0,1]or,
equivalently, over the interval [−1,1], withy(x)taken to be odd). Because of the even symmetry relative to x=1/2, only odd
nterms appear:
y(x)=x(1−x)=parenleftbigg8
π3parenrightbiggbracketleftbigg
sinπx+sin3πx
33+sin5πx
53+···bracketrightbigg
.
17.8 Rayleigh–Ritz Variational Technique 1075
(a) Calculatetheapproximation
λ=F[ytrial]
for
ytrial=1−x2.
(b) Comparewiththeexacteigenvalue.
ANS.(a) λ=2.5, (b) λ/λexact=1.013.
17.8.3 InExercise17.8.2useatrialfunction
y=1−xn.
(a) Findthevalueof nthatwillminimize F[ytrial].
(b) Showthattheoptimumvalueof ndrivestheratio λ/λexactdownto1.003.
ANS.(a)n=1.7247.
17.8.4 Aquantummechanicalparticleina sphere(Example11.7.1)satisfies
∇2ψ+k2ψ=0,
withk2=2mE/¯h2.Theboundaryconditionisthat ψ(r=a)=0,whereaistheradius
ofthesphere.Forthegroundstate[where ψ=ψ(r)]tryanapproximatewavefunction
ψa(r)=1−parenleftbiggr
aparenrightbigg2
andcalculateanapproximateeigenvalue k2
a.
Hint. To determine p(r)andw(r), put your equation in self-adjoint form (in spherical
polarcoordinates).
ANS.k2
a=10.5
a2,k2
exact=π2
a2.
17.8.5 Thewaveequationforthequantummechanicaloscillatormaybewrittenas
d2ψ(x)
dx2+parenleftbig
λ−x2parenrightbig
ψ(x)=0,
withλ=1 forthegroundstate(Eq. (13.18)). Take
ψtrial=braceleftBigg
1−x2
a2,x2≤a2
0,x2>a2
for the ground-state wave function (with a2an adjustable parameter) and calculate the
correspondingground-stateenergy.Howmucherror doyouhave?
Note.YourparabolaisreallynotaverygoodapproximationtoaGaussianexponential.
Whatimprovementscanyousuggest?
1076 Chapter 17 Calculus of Variations
17.8.6 TheSchrödingerequationfor acentralpotentialmaybewrittenas
Lu(r)+¯h2l(l+1)
2Mr2u(r)=Eu(r).
Thel(l+1)term, the angular momentum barrier, comes from splitting off the angu-
lar dependence (Section 9.3). Treating this term as a perturbation, use your variational
techniquetoshowthat E>E0,whereE0istheenergyeigenvalueof Lu0=E0u0cor-
responding to l=0. This means that the minimum energy state will have l=0, zero
angularmomentum.
Hint.Youcanexpand u(r)asu0(r)+summationtext∞
i=1ciui, where Lui=Eiui,Ei>E0.
17.8.7 Inthematrixeigenvector,eigenvalueequation
Ari=λiri,
whereλisann×nHermitianmatrix.Forsimplicity,assumethatits nrealeigenvalues
(Section3.5) aredistinct, λ1beingthelargest. If ris anapproximationto r1,
r=r1+nsummationdisplay
i=2δiri,
showthat
r†Ar
r†r≤λ1
andthattheerror in λ1isof theorder|δi|2.T ak e|δi|≪1.
Hint.T h enriform a complete orthogonal set spanning the n-dimensional (complex)
space.
17.8.8 The variational solution of Example 17.8.1 may be refined by taking y=x(1−x)+
a2x2(1−x)2.Usingthenumericalquadrature,calculate λapprox=F[y(x)],Eq.(17.128),
forafixedvalueof a2.V arya2tominimize λ.Calculatethevalueof a2thatminimizes
λand calculate λitself, both to five significant figures. Compare your eigenvalue λ
withπ2.
AdditionalReadings
Bliss, G. A., Calculus of Variations . The Mathematical Association of America. LaSalle, IL: Open Court Pub-
lishing Co. (1925). As one of the older texts, this is still a valuable reference for details of problems such as
minimum-area problems.
Courant,R.,andH.Robbins, WhatIsMathematics? 2nded.NewYork:OxfordUniversityPress(1996).Chapter
VII contains a fine discussion of the calculus of variations, including soap film solutions to minimum-area
problems.
Lanczos, C., The Variational Principles of Mechanics , 4th ed. Toronto: University of Toronto Press (1970),
reprinted, Dover (1986). This book is a very complete treatment of variational principles and their applica-
tions tothe development of classicalmechanics.
Sagan, H., Boundary and Eigenvalue Problems in Mathematical Physics . New York: Wiley (1961), reprinted,
Dover(1989).ThisdelightfultextcouldalsobelistedasareferenceforSturm–Liouvilletheory,Legendreand
Besselfunctions,andFourierSeries.Chapter1isanintroductiontothecalculusofvariations,withapplications
to mechanics. Chapter7 picks up the calculus ofvariations again andapplies it to eigenvalue problems.
17.8 Additional Readings 1077
Sagan, H., Introduction to the Calculus of Variations . New York: McGraw-Hill (1969), reprinted, Dover (1983).
Thisisanexcellentintroductiontothemoderntheoryofthecalculusofvariations,whichismoresophisticated
and complete than his 1961 text. Sagan covers sufficiency conditions and relates the calculus of variations to
problems ofspace technology.
Weinstock, R., Calculus of Variations . New York: McGraw-Hill (1952); New York: Dover (1974). A detailed,
systematic development of the calculus of variations and applications to Sturm–Liouville theory and physical
problems in elasticity, electrostatics,andquantum mechanics.
Yourgrau,W.,andS.Mandelstam, VariationalPrinciplesinDynamicsandQuantumTheory ,3rded.Philadelphia:
Saunders (1968); New York: Dover (1979). This is a comprehensive, authoritative treatment of variational
principles. The discussions of the historical development and the many metaphysical pitfalls are of particular
interest.
This page intentionally left blank
CHAPTER 18
NONLINEAR METHODS
ANDCHAOS
Our mind would lose itself in the complexity of the world if that complexity were not
harmonious;liketheshort–sighted,itwouldonlyseethedetails, andwouldbeobliged
toforgeteachofthesedetailsbeforeexaminingthenext,becauseitwouldbeincapable
oftakinginthewhole.Theonlyfactsworthyofourattentionarethosewhichintroduce
orderinto this complexity and somake it accessible to us.
HENRIPOINCARÉ
18.1 I NTRODUCTION
The origin of nonlinear dynamics goes back to the work of the renowned French mathe-
matician Henri Poincaré on celestial mechanics at the turn of the twentieth century. Clas-
sical mechanics is, in general, nonlinear in its dependence on the coordinates of the par-
ticles and the velocities, one example being vibrations with a nonlinear restoring force.
The Navier–Stokes equations are nonlinear, which makes hydrodynamics difficult to han-
dle. For almost four centuries however, following the lead of Galileo, Newton, and others,
physicists have focused on predictable, effectively linear responses of classical systems,
whichusuallyhavelinearandnonlinearproperties.
Poincaréwasthefirsttounderstandthepossibilityofcompletelyirregular,or“chaotic,”
behavior of solutions of nonlinear differential equations that are characterized by an ex-
treme sensitivity to initial conditions: Given slightly different initial conditions, from er-
rors in measurements for example, solutions can grow exponentially apart with time, so
the system soon becomes effectively unpredictable, or “chaotic.” This property of chaos,
often called the “butterfly” effect, will be discussed in Section 18.3. Since the rediscovery
ofthiseffectbyLorenzinmeteorologyintheearly1960s,thefieldofnonlineardynamics
hasgrowntremendously.Thus,nonlineardynamicsandchaostheorynowhaveenteredthe
mainstreamofphysics.
1079
1080 Chapter 18 Nonlinear Methods and Chaos
Numerousexamplesofnonlinearsystemshavebeenfoundtodisplayirregularbehavior.
Surprisingly,order,inthesenseofquantitativesimilaritiesasuniversalproperties,orother
regularities may arise spontaneously in chaos; a first example. Feigenbaum’s universal
numbers αandδwill appear in Section 18.2. Dynamical chaos is not a rare phenomenon
but is ubiquitous in nature. It includes irregular shapes of clouds, coast lines, and other
landscapes, which are examples of fractals, to be discussed in Section 18.3, and turbulent
flowoffluids,waterdrippingfromafaucet,andtheweather,ofcourse.Thedamped,driven
pendulumis amongthesimplestsystemsdisplayingchaoticmotion.
Necessaryconditionsforchaoticmotionindynamicalsystemsdescribedby first-order
differentialequationsare
•atleastthreedynamicalvariables,and
•oneormorenonlineartermscouplingtwoor severalof them.
Asinclassicalmechanics,thespaceofthetime-dependentdynamicalvariablesofasystem
of coupled differential equations is called its phase space . In such deterministic systems,
trajectories in phase space are not allowed to cross. If they did, the system would have a
choiceateachintersectionandwouldnotbedeterministic.Intwodimensionssuchnonlin-
earsystemsallowonlyforfixedpoints.Anexampleisadampedpendulum,whosesecond
derivative,¨θ=f(˙θ,θ), can be written as two first-order derivatives, ω=˙θ,˙ω=f(ω,θ),
involving just two dynamic variables, ω(t)andθ(t). In the undamped case, there will
onlybeperiodicmotionandequilibriumpoints.Withthreeormoredynamicvariables(for
example, damped,drivenpendulum, written as first-order coupled ODEs again), more
complicated nonintersecting trajectories are possible. These include chaotic motion and
arecalled deterministicchaos .
Acentralthemeinchaosistheevolutionof complex formsfromtherepetitionof simple
butnonlinear operations;thisisbeingrecognizedasa fundamentalorganizingprinciple
ofnature .Whilenonlineardifferentialequationsareanaturalplaceinphysicsforchaosto
occur,themathematicallysimpleriterationofnonlinearfunctionsprovidesaquickerentry
to chaos theory, which we will pursue first in Section 18.2. In this context, chaos already
arisesincertainnonlinearfunctionsofa singlevariable.
18.2 T HELOGISTIC MAP
Thenonlinearone-dimensionaliteration,ordifferenceequation,
xn+1=µxn(1−xn), x n∈[0,1];1<µ<4, (18.1)
is called the logistic map . It is patterned after the nonlinear differential equation dx/dt=
µx(1−x),usedbyP.F.Verhulstin1845tomodelthedevelopmentofabreedingpopula-
tion whose generations do not overlap. The density of the population at time nisxn.T h e
lineartermsimulatesthebirthrateandthenonlineartermthedeathrateofthespeciesina
constantenvironmentcontrolledbytheparameter µ.
Thequadraticfunction fµ(x)=µx(1−x)ischosenbecauseithasonemaximuminthe
interval[0,1]andiszeroattheendpoints, fµ(0)=0=fµ(1).Themaximumat xm=1/2
18.2 The Logistic Map 1081
FIGURE 18.1Cycle(x0,x1,...)forthelogisticmapfor µ=2,
startingvalue x0=0.1 andattractor x∗=1/2.
isdeterminedfrom f′(x)=0,thatis,
f′
µ(xm)=µ(1−2xm)=0,x m=1
2, (18.2)
wherefµ(1/2)=µ/4.
•Varying the single parameter µcontrols a rich and complex behavior, including one-
dimensionalchaos,asweshallsee.Moreparametersoradditionalvariablesarehardly
necessaryatthispointtoincreasethecomplexity.Inaratherqualitativesensethesim-
plelogisticmap of Eq. (18.1) is representative of many dynamical systems in biology,
chemistry,andphysics.
Figure (18.1) shows a plot of fµ(x)=µx(1−x)along with the diagonal and a series
of points (x0,x1,...)called acycle. To construct a cycle for a fixed value of µ(=2i n
Fig. 18.1), we choose some x0∈[0,1][x0=0.1 in Eq. (18.1)]. The vertical line through
x0intersectsthecurve fµ(x)atx1=fµ(x0)(=0.18inFig.18.1).Proceedinghorizontally
fromx1leads us to x1on the diagonal. Going vertically from the abscissa x1givesx2=
fµ(x1)on the curve ( x2=0.2952 in Fig. 18.1), etc. That is, straight vertical lines show
the intersections with the curve fµand horizontal lines convert fµ(xi)=xi+1to the next
abscissa.
For any initial value x0with 0<x0<1,thexiconverge toward the fixed point x∗,o r
attractor [=(0.5,0.5)inFig.18.1]:
fµ(x∗)=µx∗(1−x∗)=x∗,i.e.,x∗=1−1
µ. (18.3)
The interval (0,1)defines a basin of attraction for the fixed point x∗. The attractor x∗is
stableprovided the slope |f′
µ(x∗)|=|2−µ|<1, or 1<µ<3. This can be seen from a
Taylorexpansionofaniterationneartheattractor:
xn+1=fµ(xn)=fµ(x∗)+f′
µ(x∗)(xn−x∗)+···,i.e.,xn+1−x∗
xn−x∗=f′
µ(x∗),
1082 Chapter 18 Nonlinear Methods and Chaos
FIGURE 18.2Partof thebifurcationplotfor thelogistic
map:fixedpoints x∗versusµ.
upon dropping all higher-order terms. Thus, if |f′
µ(x∗)|<1,the next iterate, xn+1, lies
closer to x∗than does xn, implying convergence to and stability of the fixed point. How-
ever, if|f′
µ(x∗)|>1,xn+1moves farther from x∗than does xnimplying divergence and
instability. Given the continuity of f′
µinµ,the fixed point and its properties persist when
theparameter(here µ)is slightlyvaried.
Forµ>1 andx0<0o rx0>1, it is easy to verify graphically or analytically that
thexi→−∞. The origin, x=0, is arepellent fixed point since f′
µ(0)=µ>1 and the
iteratesmoveawayfrom it.Since f′
µ(1)=−µ,thepoint x=1 is arepellorfor µ>1.
When
f′
µ(x∗)=µ(1−2x∗)=2−µ=−1
isreachedfor µ=3,twofixedpointsoccur ,shownasthetwobranchesinFig.18.2,as µ
increasesbeyondthevalue3. Theycanbelocatedbysolving
x∗
2=fµparenleftbig
fµ(x∗
2)parenrightbig
=µ2x∗
2(1−x∗
2)bracketleftbig
1−µx∗
2(1−x∗
2)bracketrightbig
forx∗
2. Here it is convenient to abbreviate f(1)(x)=fµ(x),f(2)(x)=fµ(fµ(x))for the
second iterate, etc. Now we drop the common x∗
2and then reduce the remaining third-
order polynomial to second-order by recalling that a fixed point of fµis also a fixed point
off(2)becausefµ(fµ(x∗))=fµ(x∗)=x∗.S ox∗
2=x∗is one solution. Factoring out the
quadraticpolynomialweobtain
0=µ2bracketleftbig
1−(µ+1)x∗
2+2(x∗
2)2−µ(x∗
2)3bracketrightbig
−1
=(µ−1−µx∗
2)bracketleftbig
µ+1−µ(µ+1)x∗
2+µ2(x∗
2)2bracketrightbig
.
Therootsofthequadraticpolynomialare
x∗
2=1
2µparenleftbig
µ+1±radicalbig
(µ+1)(µ−3)parenrightbig
,
whicharethetwobranchesinFig.18.2for µ>3startingat x∗
2=2/3.Thisshowsthatboth
fixed points bifurcate at the same value of µ. Eachx∗
2is a point of period 2 and invariant
18.2 The Logistic Map 1083
under two iterations of the map fµ. The iterates oscillate between both branches of fixed
pointsx∗
2. A point xnis defined as a periodic point of period nforfµiff(n)(x0)=
x0,b u tf(i)(x0)/negationslash=x0for 0<i<n. Thus, for 3 <µ<3.45 (see Fig. 18.2) the stable
attractorbifurcates ,orsplits,intotwofixedpoints x∗
2.Thebifurcationfor µ=3,wherethe
doublingoccurs,iscalleda pitchfork bifurcationbecauseofitscharacteristic(roundedY-)
shape. A bifurcation is a sudden change in the evolution of the system, such as a splitting
ofonecurveintotwocurves.
Asµincreases beyond 3, the derivative df(2)/dxdecreases from unity to −1. Forµ=
1+√
6∼3.44949,whichcanbederivedfrom
df(2)
dxvextendsinglevextendsinglevextendsinglevextendsingle
x=x∗=−1,f(2)(x∗)=x∗,
each branch of fixed points bifurcates again, so x∗
4=f(4)(x∗
4), that is, has period 4. For
µ=1+√
6 theseare x∗
4=0.43996 and x∗
4=0.849938.
Withincreasingperioddoublingsitbecomesimpossibletoobtainanalyticsolutions.The
iterations are better done numerically on a programmable pocket calculator or a personal
computer, whose rapid improvements (computer-driven graphics, in particular) and wide
distributionsincethe1970shasacceleratedthedevelopmentofchaostheory.Thesequence
of bifurcations continues with ever longer periods until we reach µ∞=3.5699456...,
where an infinite number of bifurcations occur. Near bifurcation points, fluctuations,
roundingerrorsininitialconditions,etc.,playanincreasingrolebecausethesystemhasto
choose between two possible branches and becomes much more sensitive to small pertur-
bations.Inthepresentcasethe xnneverrepeat.Thebandsoffixedpoints x∗beginforming
a continuum (shown dark in Fig. 18.2); this is where chaos starts. This increasing period
doublingis the route to chaos for the logistic map that is characterizedby a universalcon-
stantδ,calleda Feigenbaumnumber .Ifthefirstbifurcationoccursat µ1=3,thesecond
atµ2=3.45,...,thentheratioof spacingsbetweenthe µnconvergesto δ:
limn→∞µn−µn−1
µn+1−µn=δ=4.66920161 .... (18.4)
FromthebifurcationplotinFig.18.2weobtain
µ2−µ1
µ3−µ2=3.45−3.00
3.54−3.45=5.0
as a first approximation for the dimensionless δ. The corresponding critical-period-2 n
pointsx∗
nleadtoanotheruniversalanddimensionlessquantity:
limn→∞x∗
n−x∗
n−1
x∗
n+1−x∗n=α=2.5029.... (18.5)
Againreadingoff Fig.18.2’sapproximatevaluesfor x∗
nweobtain
0.44−0.67
0.37−0.44=3.3
asafirst approximationfor α.
1084 Chapter 18 Nonlinear Methods and Chaos
The Feigenbaum number δis universal for the route to chaos via period doublings for
all maps with a quadratic maximum similar to the logistic map. It is an example of or-
der in chaos. Experience shows that its validity is even wider, including two-dimensional
(dissipative)systemsandtwicecontinuouslydifferentiablefunctionswithsubharmonicbi-
furcations.1When the maps behave like |x−xm|1+εnear their maximum xmfor some ε
between 0 and 1, the Feigenbaum number will depend on the exponent ε; thusδ(ε)varies
betweenδ(1)giveninEq. (18.4)for quadraticmapsto δ(0)=2f o rε=0.2
Exercises
18.2.1 Show that x∗=1 is a nontrivial fixed point of the map xn+1=xnexp[r(1−xn)]with
aslope 1−r, sothattheequilibriumisstableif 0 <r<2.
18.2.2 Drawabifurcationdiagramfor theexponentialmapofExercise18.2.1for r>1.9.
18.2.3 Determine fixed points of the cubic map xn+1=ax3
n+(1−a)xnfor 0<a<4 and
0<xn<1.
18.2.4 Writethetime-delayedlogisticalmap xn+1=µxn(1−xn−1)asatwo-dimensionalmap
xn+1=µxn(1−yn),yn+1=xn, anddeterminesomeof itsfixedpoints.
18.2.5 Show that the second bifurcation for the logistical map that leads to cycles of period 4
islocatedat µ=1+√
6.
18.2.6 Construct a nonlinear iteration function with Feigenbaum δin the interval 2 <δ<
4.6692....
18.2.7 Determine the Feigenbaum δfor (a) the exponential map of Exercise 18.2.1, (b) some
cubicmapofExercise18.2.3,(c) thetime-delayedlogisticmapof Exercise18.2.4.
18.2.8 RepeatExercise18.2.7for Feigenbaum’s αinsteadof δ.
18.2.9 Findnumericallythefirstfourpoints µforperioddoublingofthelogisticmap,andthen
obtain the first two approximations to the Feigenbaum δ. Compare with Fig. 18.2 and
Eq. (18.4).
18.2.10 Find numerically the values µwhere the cycle of period 1, 3, 4, 5, 6 begins and then
whereit becomesunstable.
Checkvalues. Forperiod3,µ=3.8284,
4,µ=3.9601,
5,µ=3.7382,
6,µ=3.6265.
18.2.11 RepeatExercise18.2.9for Feigenbaum’s α.
1More details and computer codes for the logistic map are given by G. L. Baker and J. P. Gollub, Chaotic Dynamics: An
Introduction , Cambridge, UK: Cambridge University Press (1990).
2For other maps and a discussion of the fascinating history how chaos became again a hot research topic, see D. Holton and
R. M. May in The Nature of Chaos (T. Mullin, ed.), Oxford, UK: Clarendon Press (1993), Section 5, p. 95; and Gleick’s Chaos
(1987)—seethe Additional Readings.
18.3 Sensitivity to Initial Conditions and Parameters 1085
18.3 S ENSITIVITY TO INITIAL CONDITIONS AND PARAMETERS
Lyapunov Exponents
InSection18.2wedescribedhow,asweapproachtheperiod-doublingaccumulationpara-
meter value µ∞=3.5699...from below, the period n+1 of cycles ( x0,x1,...,xn) with
xn+1=x0getslonger.It isalsoeasytocheckthatthedistances
dn=vextendsinglevextendsinglef(n)(x0+ε)−f(n)(x0)vextendsinglevextendsingle (18.6)
grow as well for small ε>0. From experience with chaotic behavior we find that this
distanceincreasesexponentiallywith n→∞;thatis,dn/ε=eλn,or
λ=1
nlnparenleftbigg|f(n)(x0+ε)−f(n)(x0)|
εparenrightbigg
, (18.7)
whereλis aLyapunov exponent for the cycle. For ε→0 we may rewrite Eq. (18.7) in
termsofderivativesas
λ=1
nlnvextendsinglevextendsinglevextendsinglevextendsingledf(n)(x0)
dxvextendsinglevextendsinglevextendsinglevextendsingle=1
nnsummationdisplay
i=0lnvextendsinglevextendsinglef′(xi)vextendsinglevextendsingle, (18.8)
usingthechainruleof differentiationfor df(n)(x)/dx, where
df(2)(x0)
dx=dfµ
dxvextendsinglevextendsinglevextendsinglevextendsingle
x=fµ(x0)dfµ
dxvextendsinglevextendsinglevextendsinglevextendsingle
x=x0=f′
µ(x1)f′
µ(x0) (18.9)
andf′
µ=dfµ/dx, etc. Our Lyapunov exponent has been calculated at the point x0, and
Eq.(18.8) isexactfor one-dimensionalmaps.
Asameasureofthesensitivityofthesystemtochangesininitialconditions,onepointis
notenoughtodetermine λinhigher-dimensionaldynamicalsystemsingeneral,wherethe
motionoftenisbounded,sothe dncannotgoto∞.Insuchcases,werepeattheprocedure
forseveralpointsonthetrajectoryandaverageoverthem.Thisway,weobtainthe average
Lyapunov exponent for the sample. This average value is often called and taken as the
Lyapunovexponent.
TheLyapunovexponent λisaquantitativemeasureofchaos:Aone-dimensionaliterated
function similar to the logistic map has chaoticcycles (x0,x1,...) for the parameter µif
the average Lyapunov exponent is positive for that value of µ. Any such initial point
x0is called a strangeorchaotic attractor (the shaded region in Fig. 18.2). For cycles
of finite period, λis negative. This is the case for µ<3, forµ<µ∞, and even in the
periodicwindowat µ∼3.627insidethechaoticregionofFig. 18.2.Atbifurcationpoints,
λ=0. Forµ>µ∞the Lyapunov exponent is positive, except in the periodic windows,
whereλ<0,andλgrowswith µ.Inotherwords,thesystembecomesmorechaoticasthe
controlparameter µincreases.
In the chaos region of the logistic map there is a scaling law for the average Lyapunov
exponent(wedonotderiveit),
λ(µ)=λ0(µ−µ∞)ln2/lnδ, (18.10)
1086 Chapter 18 Nonlinear Methods and Chaos
where ln2 /lnδ∼0.445,δis the universal Feigenbaum number of Section 18.2, and λ0
is a constant. This relation (18.10) is reminiscent of a physical observable at a (second-
order) phase transition. The exponent in Eq. (18.10) is a universal number; the Lyapunov
exponent plays the role of an order parameter , whileµ−µ∞is the analog of T−Tc,
whereTcis thecriticaltemperatureatwhichthephasetransitionoccurs.
Fractals
In dissipative chaotic systems (but rarely in conservative Hamiltonian systems) often new
geometric objects with intricate shapes appear that are called fractalsbecause of their
noninteger dimension. Fractals are irregular geometric objects whose dimension is typi-
cally not integral and that exist at many scales, so their smaller parts resemble their larger
parts. Intuitively a fractal is a set which is (approximately) self-similar under magnifica-
tion.Asetofattractingpointswithnonintegerdimensionis calleda strangeattractor.
Weneedaquantitativemeasureofdimensionalityinordertodescribefractals.Unfortu-
nately, there are several definitions with usually different numerical values, none of which
has yet become a standard. For strictly self-similar, sets, one measure suffices. More com-
plicated(forinstance,onlyapproximatelyself-similar)setsrequiremoremeasuresfortheir
complete description. The simplest is the box-counting dimension , due to Kolmogorov
and Hausdorff. For a one-dimensional set, we cover the curve by line segments of length
R. In two dimensions the boxes are squares of area R2, in three dimensions cubes of vol-
umeR3,etc.Thenwecountthenumber N(R)ofboxesneededtocovertheset.Letting R
go to zero we expect Nto scale as N(R)∼R−d. Taking the logarithm the box-counting
dimension isdefinedas
d≡lim
R→0lnN(R)
lnR. (18.11)
For example, in a two-dimensional space a single point is covered by one square, so
lnN(R)=0 andd=0. A finite set of isolated points also has dimension d=0. For a
differentiable curve of length L,N(R)∼L/RasR→0, sod=1 from Eq. (18.11), as
expected.
Let us now construct a more irregular set, the Kochcurve. We start with a line segment
of unit length in Fig. 18.3 and remove the middle third. Then we replace it with two seg-
ments of length 1 /3, which form a triangle in Fig. 18.3. We iterate this procedure with
each segment ad infinitum. The resulting Koch curve is infinitely long and is nowhere dif-
ferentiable because of the infinitely many discontinuous changes of slope. At the nth step
each line segment has length Rn=3−nand there are N(Rn)=4nsegments. Hence its
dimension is d=ln4/ln3=1.26...,which is more than a curve but less than a surface.
BecausetheKochcurveresults fromiterationofthefirst step,itis strictlyself-similar.
For the logistic map the box-counting dimension at a period–doubling accumulation
pointµ∞is 0.5388...,which is a universal number for iterations of functions in one
variable with a quadratic maximum. To see roughly how this comes about, consider the
pairsoflinesegmentsoriginatingfromsuccessivebifurcationpointsforagivenparameter
µinthechaosregime(seeFig.18.2).Imagineremovingtheinteriorspacefromthechaotic
bands. When we go to the next bifurcation, the relevant scale parameter is α=2.5029...
from Eq. (18.5). Suppose we need 2nline segments of length Rto cover 2nbands. In the
18.3 Sensitivity to Initial Conditions and Parameters 1087
FIGURE 18.3ConstructionoftheKochcurveby
iterations.
next stage then we need 2n+1segments of length R/αto cover the bands. This yields a
dimension d=−ln(2n/2n+1)/lnα=0.4498.... This crude estimatecan be improvedby
taking into account that the width between neighboring pairs of line segments differs by
1/α(see Fig. 18.2). The improved estimate, 0.543, is closer to 0.5388 .... This example
suggests that when the fractal set does not have a simple self-similar structure, then the
box-countingdimensiondependsonthebox-constructionmethod.
Finally, we turn to the beautiful fractals that are surprisingly easy to generate and
whose color pictures had considerable impact. For complex c=a+ib, the correspond-
ingquadraticcomplexmapinvolvingthecomplexvariable z=x+iy,
zn+1=z2
n+c, (18.12)
looks deceptively simple, but the equivalent two-dimensional map in terms of the real
variables
xn+1=x2
n−y2
n+a, y n+1=2xnyn+b (18.13)
revealsalreadymoreofitscomplexity.ThismapformsthebasisforsomeofMandelbrot’s
beautifulmulticolorfractalpictures(wereferthereadertoMandelbrot(1988)andPeitgen
and Richter (1986) in the Additional Readings), and it has been found to generate rather
intricate shapes for various c/negationslash=0. For example, the Julia set of a map zn+1=F(zn)is
defined as the set of all its repelling fixed or periodic points. Thus it forms the boundary
betweeninitialconditionsofatwo-dimensionaliteratedmapleadingtoiteratesthatdiverge
and those that stay within some finite region of the complex plane. For the case c=0 and
F(z)=z2, the Julia set can be shown to be just a circle about the origin of the complex
plane. Yet, just by adding a constant c/negationslash=0, the Julia set becomes fractal. For instance, for
c=−1 one finds a fractal necklace with infinitely many loops (see Devaney (1989) in the
AdditionalReadings).
While the Julia set is drawn in the complex plane, the Mandelbrot set is constructed
in the two-dimensional parameter space c=(a,b)=a+bi. It is constructed as follows.
1088 Chapter 18 Nonlinear Methods and Chaos
Starting from the initial value z0=0=(0,0)one searches Eq. (18.12) for parameter val-
uescsothattheiterated {zn}donotdivergeto∞.Eachcoloroutsidethefractalboundary
of the Mandelbrot set represents a given number of iterations m, say, needed for the zn
to go beyond a specified absolute (real) value R,|zm|>R>|zm−1|. For real parame-
ter value c=a, the resulting map, xn+1=x2
n+a, is equivalent to the logistic map with
period-doubling bifurcations (see Section 18.2) as aincreases on the real axis inside the
Mandelbrotset.
Exercises
18.3.1 Use a programmable pocket calculator (or a personal computer with BASIC or FOR-
TRAN or symbolic software such as Mathematica or Maple) to obtain the iterates xi
of an initial 0 <x0<1 andf′
µ(xi)for the logistic map. Then calculate the Lyapunov
exponent for cycles of period 2 ,3,...of the logistic map for 2 <µ<3.7. Show that
forµ<µ∞theLyapunovexponent λis0atbifurcationpointsandnegativeelsewhere,
whilefor µ>µ∞itis positiveexceptinperiodicwindows.
Hint.SeeFig.9.3 ofHilborn(1994)intheAdditionalReadings.
18.3.2 Considerthemap xn+1=F(xn)with
F(x)=braceleftBigg
a+bx, x < 1,
c+dx, x> 1,
forb>0andd<0.ShowthatitsLyapunovexponentispositivewhen b>1,d<−1.
Plota fewiterationsinthe (xn+1,xn)plane.
18.4 N ONLINEAR DIFFERENTIAL EQUATIONS
In Section 18.1 we mentioned nonlinear differential equations (abbreviated as NDEs) as
the natural place in physics for chaos to occur, but continued with the simpler iteration of
nonlinearfunctionsofonevariable(maps).Herewebrieflyaddressthemuchbroaderarea
of NDEs and the far greater complexity in the behavior of their solutions. However, maps
and systems of solutions of NDEs are closely related. The latter can often be analyzed in
terms of discrete maps. One prescription is the so-called Poincaré section of a system of
NDE solutions. Placing a plane transverse into a trajectory (of a solution of a NDE), it in-
tersectstheplaneinaseriesofpointsatincreasingdiscretetimes,forexample,inFig.18.4
(x(t1),y(t1))=(x1,y1),(x2,y2),...,which are recorded and graphically or numerically
analyzed for fixed points, period-doubling bifurcations, etc. This method is useful when
solutions of NDEs are obtained numerically in computer simulations so that one can gen-
erate Poincaré sections at various locations and with different orientations, with further
analysisleadingtotwo-dimensionaliteratedmaps
xn+1=F1(xn,yn), y n+1=F2(xn,yn) (18.14)
stored by the computer. Extracting the functions Fjanalytically or graphically is not al-
wayseasy,though.
Let us start with a few classical examples of NDEs. In Chapter 9 we have already dis-
cussedthesolitonsolutionofthenonlinearKorteweg–deVries PDE,Eq. (9.11).
18.4 Nonlinear Differential Equations 1089
FIGURE 18.4Schematicof aPoincarésection.
Exercise
18.4.1 Forthedampedharmonicoscillator
¨x+2a˙x+x=0,
consider the Poincaré section {x>0,y=˙x=0}.T a k e0<a≪1 and show that the
mapis givenby xn+1=bxnwithb<1.Findanestimatefor b.
Bernoulli and Riccati Equations
Bernoulliequationsarealsononlinear,havingtheform
y′(x)=p(x)y(x)+q(x)bracketleftbig
y(x)bracketrightbign, (18.15)
wherepandqare real functions and n/negationslash=0, 1 to exclude first-order linear ODEs. If we
substitute
u(x)=bracketleftbig
y(x)bracketrightbig1−n, (18.16)
thenEq. (18.15)becomesafirst-order linearODE,
u′=(1−n)y−ny′=(1−n)bracketleftbig
p(x)u(x)+q(x)bracketrightbig
, (18.17)
whichwecansolveasdescribedinSection9.2.
Riccatiequationsare quadraticin y(x):
y′=p(x)y2+q(x)y+r(x), (18.18)
1090 Chapter 18 Nonlinear Methods and Chaos
wherep/negationslash=0toexcludelinearODEsand r/negationslash=0toexcludeBernoulliequations.Thereisno
general method for solving Riccati equations. However, when a special solution y0(x)of
Eq. (18.18) is known by a guess or inspection, then one can write the general solution in
theformy=y0+u, withusatisfyingtheBernoulliequation
u′=pu2+(2py0+q)u, (18.19)
becausesubstitutionof y=y0+uintoEq. (18.18)removes r(x)fromEq. (18.18).
JustasforRiccatiequationstherearenogeneralmethodsforobtainingexactsolutionsof
other nonlinear ODEs. It is more important to develop methods for finding the qualitative
behavior of solutions. In Chapter 9 we mentioned that power-series solutions of ODEs
exist except (possibly) at regular or essential singularities, which are directly given by
localanalysisofthecoefficientfunctionsoftheODE.Suchlocalanalysisprovidesuswith
theasymptoticbehaviorofsolutionsaswell.
Fixed and Movable Singularities, Special Solutions
Solutions of NDEs also have such singular points, independent of the initial or boundary
conditionsandcalled fixedsingularities .Inadditiontheymayhave spontaneous ,ormov-
able, singularities that vary with the initial or boundary conditions. They complicate the
(asymptotic)analysisofNDEs.ThispointisillustratedbyacomparisonofthelinearODE
y′+y
x−1=0, (18.20)
which has the obvious regular singularity at x=1, with the NDE y′=y2. Both have the
same solution with initial condition y(0)=1, namely, y(x)=1/(1−x).F o ry(0)=2,
though, the pole in the (obvious, but check) solution y(x)=2/(1−2x)of the NDE has
moved to x=1/2.
For a second-order ODE we have a complete description of (the asymptotic behavior
of) its solutions when (that of) two linearly independent solutions are known. For NDEs
there may still be special solutions whose asymptotic behavior is not obtainable from
two independent solutions. This is another characteristic property of NDEs, which we
illustrateagainbyanexample.
ThegeneralsolutionoftheNDE y′′=yy′/xis givenby
y(x)=2c1tan(c1lnx+c2)−1, (18.21)
whereciare integration constants. An obvious (check it) special solution is y=c3=
constant, which cannot be obtained from Eq. (18.20) for any choice of the parameters
c1,c2. Note that using the substitution x=et,Y(t)=y(et)so thatxdy/dx=dY/dt,w e
obtain the ODE Y′′=Y′(Y+1). This ODE can be integrated once to give Y′=1
2Y2+
Y+cwithc=2(c2
1+1/4)an integration constant, and again according to Section 9.2 to
leadtothesolutionofEq. (18.21).
18.4 Nonlinear Differential Equations 1091
Autonomous Differential Equations
Differential equations that do not explicitly contain the independent variable, taken to be
the timethere, are called autonomous . Verhulst’s NDE ˙y=dy/dt=µy(1−y), which
weencounteredbrieflyin Section18.2 as motivationfor thelogistic map,is a specialcase
of this wide and important class of ODEs.3For one dependent variable y(t)they can be
writtenas
˙y=f(y), (18.22a)
andfor severaldependentvariablesas asystem
˙yi=fi(y1,y2,...,yn), i=1,2,...,n, (18.22b)
with sufficiently differentiable functions f,fi. A solution of Eq. (18.22b) is a curve or
trajectory y(t)forn=1 and in general a trajectory (y1(t),y2(t),...,y n(t))in ann-
dimensional(so-called) phasespace .AsdiscussedalreadyinSection18.1,twotrajectories
cannot cross because of the uniqueness of the solutions of ODEs. Clearly, solutions of the
algebraicsystem
fi(y1,y2,...,yn)=0 (18.23)
arespecialpointsinphasespace,wherethepositionvector( y1,y2,...,yn)doesnotmove
onthetrajectory;theyarecalled critical(orfixed)points.Itturnsoutthatalocalanalysis
of solutions near critical points leads to an understanding of the global behavior of the
solutions.Firstletuslookatasimpleexample.
ForVerhulst’sODE, f(y)=µy(1−y)=0g i v e sy=0andy=1 asthecriticalpoints.
For the logistic map, y=0 andy=1 are repellent fixed points because df/dy(0)=µ
aty=0 anddf/dy(1)=−µaty=1f o rµ>1. A local analysis near y=0 suggests
neglecting the y2term and solving ˙y=µyinstead. Integratingintegraltext
dy/y=µt+lncgives
the solution y(t)=ceµt, which diverges as t→∞,s oy=0 is a repellent critical point.
(Notethatfor µ<0ofthelogisticmapthecriticalpoint y=0wouldbeattracting,leading
to a converging y∼eµtsolution.) Similarly at y=1,integraltext
dy/(1−y)=µt−lncleads to
y(t)=1−ce−µt→1f o rt→∞. Hencey=1 is an attracting critical point. Because the
ODEis separable,itsgeneralsolutionis givenby
integraldisplaydy
y(1−y)=integraldisplay
dybracketleftbigg1
y+1
1−ybracketrightbigg
=lny
1−y=µt+lnc.
Hencey(t)=ceµt/(1+ceµt)fort→∞convergesto 1, thus confirming the local analy-
sis. This examplemotivatesus to look nextat the properties of fixedpoints in more detail.
Foranarbitraryfunction f,itiseasytoseethat
•inone dimension , fixed points yiwithf(yi)=0d i v i d et h e y-axis into dynamically
separate intervals because, given an initial value in one of the intervals, the trajectory
y(t)willstaythere,for itcannotgobeyondeitherfixedpointwhere ˙y=0.
3Solutions of nonautonomous equations canbe much more complicated.
1092 Chapter 18 Nonlinear Methods and Chaos
FIGURE 18.5Fixedpoints:(a) repellor,(b) sink.
Iff′(y0)>0 atthefixedpoint y0wheref(y0)=0,thenat y0+εforε>0 sufficiently
small,˙y=f′(y0)ε+O(ε2)>0inaneighborhoodtotherightof y0,sothetrajectory y(t)
keepsmovingtotheright,awayfromthefixedpoint y0.Totheleftof y0,˙y=−f′(y0)ε+
O(ε2)<0,sothetrajectorymovesawayfrom thefixedpointhereaswell.Hence,
•a fixed point [with f(y0)=0] aty0withf′(y0)>0, as shown in Fig. 18.5a, repels
trajectories;thatis,alltrajectoriesmoveawayfromthecriticalpoint: ···←·→··· ;it
isarepellor.Similarly,weseethat
•a fixed point at y0withf′(y0)<0, as shown in Fig. 18.5b, attracts trajectories; that
is, all trajectories converge toward the critical point y0:···→·←··· ;i ti sasink or
node.
Letusnowconsidertheremainingcasewhenalso f′(y0)=0.
Let us assume f′′(y0)>0. Then at y0+εto the right of fixed point y0,˙y=
f′′(y0)ε2/2+O(ε3)>0, so the trajectory moves away from the fixed point there, while
to the left it moves closer to y0. In other words, we have a saddle point .F o rf′′(y0)<0,
the sign of˙yis reversed, so we deal again with a saddle point with the motion to the right
ofy0toward the fixed point and at left away from it. Let us summarize the local behavior
oftrajectoriesnearsuchafixedpoint y0:W eha v ea
•a saddle point at y0whenf(y0)=0, andf′(y0)=0, as shown in Fig. 18.6a,b corre-
sponding to the cases where (a) f′′(y0)>0 and trajectories on one side of the critical
point, converge toward it and diverge from it on the other side: ···→·→··· ; and
(b)f′′(y0)<0. Here the direction is simply reversed compared to (a). Figure 18.6(c)
showsthecaseswhere f′′(y0)=0.
So far we have ignored the additional dependence of f(y)on one or more parameters,
such asµfor the logistic map. When a critical point maintains its properties qualitatively
as we adjust a parameter slightly, we call it structurally stable . This is reasonable be-
cause structurally unstable objects are unlikely to occur in reality because noise and other
neglected degrees of freedom act as perturbations on the system that effectively prevent
such unstable points from being observed. Let us now look at fixed points from this point
of view. Upon varying such a control parameter slightly we deform the function f,o rw e
18.4 Nonlinear Differential Equations 1093
FIGURE 18.6Saddlepoints.
may just shift fup or down or sideways in Fig. 18.5 a bit. This will move a little the
locationy0of the fixed point with f(y0)=0, but maintain the sign of f′(y0). Thus, both
sinks and repellors are stable , while a saddle point in general is not. For example, shift-
ingfinFig. 18.6adownabitcreatestwo fixedpoints,oneasinkandtheotherarepellor,
and removes the saddle point. Since two conditions must be satisfied at a saddle point,
they are less common and important, being unstable with respect to variations of parame-
ters. However, they mark the border between different types of dynamics and are useful
and meaningful for the global analysis of the dynamics. We are now ready to consider the
richer,butmorecomplicated,higher-dimensionalcases.
Local and Global Behavior in Higher Dimensions
In two or more dimensions we start the local analysis at a fixed point (y0
1,y0
2,...)with
˙yi=fi(y0
1,y0
2,...)=0 using the same Taylor expansion of the fiin Eq. (18.22b) as
for the one-dimensional case. Retaining only the first-order derivatives, this approach lin-
earizes the coupled NDEs of Eq. (18.22b) and reduces their solution to linear algebra as
follows. We abbreviate the constant derivatives at the fixed point as a matrix Fwith ele-
ments
fij≡∂fi
∂yjvextendsinglevextendsinglevextendsinglevextendsingle
(y0
1,y0
2,···). (18.24)
1094 Chapter 18 Nonlinear Methods and Chaos
IncontrasttothestandardlinearalgebrainChapter3,however, Fisneithersymmetricnor
Hermitianingeneral.Asaresult,itseigenvaluesmaynotbereal.Ifweshiftthefixedpoint
to the origin and call the shifted coordinates xi=yi−y0
i, then the coupled NDEs of Eq.
(18.22b)become
˙xi=summationdisplay
jfijxj, (18.25)
that is, coupled linear ODEs with constant coefficients. We solve Eq. (18.25) with the
standardexponentialAnsatz,
xi(t)=summationdisplay
jcijeλjt, (18.26)
with constant exponents λjand a constant matrix Cof coefficients cij,s ocj=(cij,i=
1,2,...)formsthe jthcolumnvectorof C.SubstitutingEq.(18.26)intoEq.(18.25)yields
alinearcombinationof exponentialfunctions,
summationdisplay
jcijλjeλjt=summationdisplay
j,kfikckjeλjt, (18.27)
whichareindependentif λi/negationslash=λj.Thisisthegeneralcaseonwhichwefocus,whiledegen-
eracieswheretwoormore λareequalrequirespecialtreatmentsimilartosaddlepointsin
one dimension. Comparing coefficients of exponential functions with the same exponent
yieldsthelineareigenvalueequations
summationdisplay
kfikckj=λjcij,orFcj=λjcj. (18.28)
A nontrivial solution comprising the eigenvalue λjand eigenvector cjof the homoge-
neous linear equations (18.28) requires λjto be a root of the secular equation (compare
withSection3.5):
det(F−λ·1)=0. (18.29)
Equation (18.28) means that Cdiagonalizes F, so we can write Eq. (18.28) also
as
C−1FC=[λ1,λ2,...]. (18.30)
In the new but in general nonorthogonal coordinates ξj, defined as Cξ=x,w eh a v e
a fixed point for each direction ξj,a s˙ξj=λjξj, where the λjplay the role of f′(y0)
in the one-dimensional case. The λarecharacteristic exponents and complex num-
bers in general. This is seen by substituting x=Cξinto Eq. (18.25) in conjunction with
Eqs. (18.28) and (18.30). Thus, this solution represents the independent combination of
one-dimensional fixed points, one for each component of ξand each independent of the
other components. In two dimensions for λ1<0 andλ2<0, then, we have a sink in all
directions.Whenboth λaregreaterthan 0,wehaverepellorinalldirections.
18.4 Nonlinear Differential Equations 1095
Example 18.4.1 STABLE SINK
ThecoupledODEs
˙x=−x,˙y=−x−3y
haveanequilibriumpointattheorigin.Thesolutionshavetheform
x(t)=c11eλ1t,y(t)=c21eλ1t+c22eλ2t,
so the eigenvalue λ1=−1 results from λ1c11=−c11, and the solution is x=c11e−t.T h e
determinantofEq. (18.29),
vextendsinglevextendsinglevextendsinglevextendsingle−1−λ0
−1−3−λvextendsinglevextendsinglevextendsinglevextendsingle=(1+λ)(3+λ)=0,
yieldstheeigenvalues λ1=−1,λ2=−3.Becausebotharenegativewehaveastablesink
attheorigin.TheODEfor ygivesthelinearrelations
λ1c21=−c11−3c21=−c21,λ 2c22=−3c22,
from which we infer 2 c21=−c11,orc21=−c11/2. Because the general solution will
containtwoconstants,itis givenby
x(t)=c11e−t,y(t)=−c11
2e−t+c22e−3t.
Asthetime t→∞,wehavey∼−x/2andx→0andy→0,whilefor t→−∞,y∼x3
andx,y→±∞. The motion toward the sink is indicated by arrows in Fig. 18.7. To find
theorbit, weeliminatetheindependentvariable, t, andfindthecubics:
y=−x
2+c22
c3
11x3./squaresolid
When both λare greater than 0, we have repellor. In this case the motion is away from
the fixed point. However, when the λhave different signs, we have a saddle point, that is,
acombinationofasinkinonedimensionandarepellorintheother.Thistypeofbehavior
generalizestohigherdimensions.
Example 18.4.2 SADDLE POINT
ThecoupledODEs
˙x=−2x−y,˙y=−x+2y
haveafixedpointattheorigin.Thesolutionshavetheform
x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t.
Theeigenvalues λ=±√
5 aredeterminedfrom
vextendsinglevextendsinglevextendsinglevextendsingle−2−λ−1
−12−λvextendsinglevextendsinglevextendsinglevextendsingle=λ2−5=0.
1096 Chapter 18 Nonlinear Methods and Chaos
FIGURE 18.7Stablesink.
SubstitutingthegeneralsolutionsintotheODEsyieldsthelinearequations
λ1c11=−2c11−c21=√
5c11,λ 1c21=−c11+2c21=√
5c21,
λ2c12=−2c12−c22=−√
5c12,λ 2c22=−c12+2c22=−√
5c22,
or
(√
5+2)c11=−c21,(√
5−2)c12=c22,
(√
5−2)c21=−c11,(√
5+2)c22=c12,
soc21=−(2+√
5)c11,c22=(√
5−2)c12. The family of solutions depends on two
parameters, c11,c12. For large time t→∞, the positive exponent prevails and y∼
−(√
5+2)x,while for t→−∞we have y=(√
5−2)x. These straight lines are the
asymptotesoftheorbits.Because −(√
5+2)(√
5−2)=−1 theyareorthogonal.Wefind
the orbits by eliminating the independent variable, t, as follows. Substituting the c2jwe
write
y=−2x−√
5parenleftbig
c11e√
5t−c12e−√
5tparenrightbig
,soy+2x√
5=−c11e√
5t+c12e−√
5t.
Nowweaddandsubtractthesolution x(t)toget
1√
5(y+2x)+x=2c12e−√
5t,1√
5(y+2x)−x=−2c11e√
5t,
18.4 Nonlinear Differential Equations 1097
FIGURE 18.8Saddlepoint.
whichwemultiplytoobtain
1
5(y+2x)2−x2=−4c12c11=const.
The resulting quadratic form, y2+4xy−x2=const., is a hyperbola because of the
negative sign. The hyperbola is rotated in the sense that its asymptotes are not aligned
with the x,y-axes (Fig. 18.8). Its orientation is given by the direction of the as-
ymptotes that we found earlier. Alternatively we could find the direction of minimal
distance from the origin, proceeding as follows. We set the xandyderivatives of
f+/Lambda1g≡x2+y2+/Lambda1(y2+4xy−x2)equaltozero,where /Lambda1istheLagrangemultiplier
for the hyperbolic constraint. The four branches of hyperbolas correspond to the differ-
ent signs of the parameters c11andc12. Figure 18.8 is plotted for the cases c11=±1,
c12=±2. /squaresolid
However, a new kind of behavior arises for a pair of complex conjugate eigenvalues
λ1,2=ρ±iκ. If we write the complex solutions ξ1,2=exp(ρt±iκt)in real variables
ξ+=(ξ1+ξ2)/2,ξ−=(ξ1−ξ2)/2iuponusingtheEuleridentityexp (ix)=cosx+isinx
(seeSection6.1),
ξ+=exp(ρt)cos(κt), ξ −=exp(ρt)sin(κt) (18.31)
describe a trajectory that spirals inward to the fixed point at the origin for ρ<0, aspiral
node,andspiralsawayfromthefixedpointfor ρ>0,aspiralrepellor .
1098 Chapter 18 Nonlinear Methods and Chaos
Example 18.4.3 SPIRAL FIXED POINT
ThecoupledODEs
˙x=−x+3y,˙y=−3x+2y
haveafixedpointattheoriginandsolutionsoftheform
x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t.
Theexponents λ1,2aresolutionsof
vextendsinglevextendsinglevextendsinglevextendsingle−1−λ3
−32−λvextendsinglevextendsinglevextendsinglevextendsingle=(1+λ)(λ−2)+9=0,
orλ2−λ+7=0.Theeigenvaluesarecomplexconjugate, λ=1/2±i√
27/2,sowedeal
withaspiralfixedpointattheorigin(arepellorbecause1 /2>0).Substitutingthegeneral
solutionsintotheODEsyieldsthelinearequations
λ1c11=−c11+3c21,λ 1c21=−3c11+2c21,
λ2c12=−c12+3c22,λ 2c22=−3c12+2c22,
or
(λ1+1)c11=3c21,(λ1−2)c21=−3c11,
(λ2+1)c12=3c22,(λ2−2)c22=−3c12,
which,usingthevaluesof λ1,2, implythefamilyofcurves
x(t)=et/2parenleftbig
c11ei√
27t/2+c12e−i√
27t/2parenrightbig
,
y(t)=x
2+√
27
6et/2iparenleftbig
c11ei√
27t/2−c12e−i√
27t/2parenrightbig
,
whichdependsontwoparameters, c11,c12.Tosimplifywecanseparaterealandimaginary
partsofx(t)andy(t)usingtheEuleridentity eix=cosx+isinx.Itisequivalent,butmore
convenient, to choose c11=c12=c/2 and rescale t→2t,so with the Euler identity we
have
x(t)=cetcos(√
27t), y(t)=x
2−√
27
6cetsin(√
27t).
Herewecaneliminate tandfindtheorbit
x2+4
3parenleftbigg
y−x
2parenrightbigg2
=parenleftbig
cetparenrightbig2.
For fixed tthis is the positive definite quadratic form x2−xy+y2=const., that is, an
ellipse. But there is no ellipse in the solutions because tis not fixed. Nonetheless, it is
useful to find its orientation. We proceed as follows. With /Lambda1the Lagrange multiplier for
the elliptical constraint we seek the directions of maximal and minimal distance from the
origin,forming
f(x,y)+/Lambda1g(x,y)≡x2+y2+/Lambda1parenleftbig
x2−xy+y2parenrightbig
18.4 Nonlinear Differential Equations 1099
FIGURE 18.9Spiralpoint.
andsetting
∂(f+/Lambda1g)
∂x=2x+2/Lambda1x−/Lambda1y=0,∂(f+/Lambda1g)
∂y=2y+2/Lambda1y−/Lambda1x=0.
From
2(/Lambda1+1)x=/Lambda1y, 2(/Lambda1+1)y=/Lambda1x
weobtainthedirections
x
y=/Lambda1
2(/Lambda1+1)=2(/Lambda1+1)
/Lambda1,
or/Lambda12+8
3/Lambda1+4
3=0. This yields the values /Lambda1=−2/3,−2 and the directions y=±x.
In other words, our ellipse is centered at the origin and rotated by 45◦. As we vary the
independent variable, t, the size of the ellipse changes, so we get the rotated spiral shown
inFig.18.9for c=1. /squaresolid
In the special case when ρ=0 in Eq. (18.31), the circular trajectory is called a cycle.
Whentrajectoriesnearitareattractedastimegoeson,itiscalleda limitcycle ,representing
periodicmotionforautonomoussystems.
1100 Chapter 18 Nonlinear Methods and Chaos
FIGURE 18.10Center.
Example 18.4.4 CENTER OR CYCLE
TheundampedlinearharmonicoscillatorODE ¨x+ω2x=0canbewrittenastwocoupled
ODEs:
˙x=−ωy,˙y=ωx.
Integrating the resulting ODE ˙xx+˙yy=0 yields the circular orbits x2+y2=const.,
whichdefineacenterattheoriginandareshowninFig.18.10.Thesolutionscanbepara-
meterizedas x=Rcost,y=Rsint,whereRistheradiusparameter.Theycorrespondto
thecomplexconjugateeigenvalues λ1,2=±iω.Wecancheckthemifwewritethegeneral
solutionas
x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t.
Thentheeigenvaluesfollowfrom
vextendsinglevextendsinglevextendsinglevextendsingle−λ−ω
ω−λvextendsinglevextendsinglevextendsinglevextendsingle=λ2+ω2=0=0.
/squaresolid
Anotherclassicattractorisquasiperiodicmotion,suchasthetrajectory
x(t)=A1sin(ω1t+b1)+A2sin(ω2t+b2), (18.32)
18.4 Nonlinear Differential Equations 1101
where the ratio ω1/ω2is an irrational number. Such combined oscillations occur as solu-
tionsofadampedanharmonicoscillator(VanderPolnonautonomoussystem)
¨x+2γ˙x+ω2
2x+βx3=fcos(ω1t). (18.33)
In three dimensions, when there is a positive characteristic exponent and an attracting
complex conjugate pair in the other directions, we have a spiral saddle point as a new
feature.Conversely,anegativecharacteristicexponentinconjunctionwitharepellingpair
also gives rise to a spiral saddle point, where trajectories spiral out in two dimensions but
areattractedinathirddirection.
In general, when some form of damping (or dissipation of energy) is present, the tran-
sients decay and the system settles either in equilibrium, that is, a single point, or in pe-
riodic or quasiperiodic motion. Chaotic motion in dissipative systems is now recognized
as a fourth state, and its attractors are often called strange. In dissipative systems, initial
conditions are not important because trajectories end up on some attractor. They are cru-
cialinHamiltoniansystems.InnonintegrableHamiltoniansystems,chaosmayalsooccur,
and then it is called conservative chaos. We refer to Chapter 8 of Hilborn (1994) in the
AdditionalReadingsfor thismorecomplicatedtopic.
For the driven damped pendulum when trajectories near a center (closed orbit) are at-
tracted to it as time goes on, this closed orbit is defined as a limit cycle , representing
periodicmotionforautonomoussystems.Adampedpendulumusuallyspiralsintotheori-
gin (the position at rest); that is, the origin is a spiral fixed point in its phase space. When
we turn on a driving force, then the system formally becomes nonautonomous,because of
itsexplicittimedependence,butalsomoreinteresting.Inthiscase,wecancalltheexplicit
timeinasinusoidaldrivingforceanewvariable, ϕ,whereω0isafixedrate,intheequation
ofmotion
˙ω+γω+sinθ=fsinϕ, ω=˙θ, ϕ=ω0t.
Then we increase the dimension of our phase space by 1 (adding one variable, ϕ) because
˙ϕ=ω0=const., but we keep the coupled ODEs autonomous. This driven damped pen-
dulum has trajectories that cross a closed orbit in phase space and spiral back to it; it is
calledlimit cycle . This happens for a range of strength fof the driving force, the control
parameter of the system. As we increase f, the phase space trajectories go through sev-
eral neighboring limit cycles and eventually become aperiodic and chaotic. Such closed
limit cycles are called Hopf bifurcations of the pendulum on its road to chaos, after the
mathematician E. Hopf, who generalized Poincaré’s results on such bifurcations to higher
dimensionsofphasespace.
Such spiral sinks or saddle points cannot occur in one dimension, but we might ask if
theyarestablewhentheyoccurinhigherdimensions.AnanswerisgivenbythePoincaré–
Bendixson theorem, which says that either trajectories (in the finite region to be specified
in a moment) are attracted to a fixed point as time goes on, or they approach a limit cycle
provided the relevant two-dimensional subsystem stays inside a finite region, that is, does
not diverge there as t→∞. For a proof we refer to Hirsch and Smale (1974) and Jackson
(1989)intheAdditionalReadings.
In general, when some form of damping is present, the transients decay and the system
settles either in equilibrium, that is, a single point, or in periodic or quasiperiodic motion.
Chaoticmotionisnowrecognizedas afourthstate,anditsattractorsareoften strange.
1102 Chapter 18 Nonlinear Methods and Chaos
Exercise
18.4.2 Showthatthe(Rössler) coupledODEs
˙x1=−x2−x3,˙x2=x1+a1x2,˙x3=a2+(x1−a3)x3
(a) havetwofixedpointsfor a2=2,a3=4,and 0<a1<2,
(b) haveaspiralrepellorattheorigin,and
(c) haveaspiralchaoticattractorfor a1=0.398.
Dissipation in Dynamical Systems
Dissipativeforcesofteninvolvevelocities,thatis,first-ordertimederivatives,suchasfric-
tion (for example, for the damped oscillator). Let us look for a measure of dissipation,
that is, how a small area A=c1,2/Delta1ξ1/Delta1ξ2at a fixed point shrinks or expands, first in two
dimensions for simplicity. Here c1,2≡sin(ˆξ1,ˆξ2), involving the sine of the characteristic
directions, is a time-independentangular factor that takes into accountthe nonorthogonal-
ityofthecharacteristicdirections ˆξ1andˆξ2ofEq.(18.28).Ifwetakethetimederivativeof
Aand use˙ξj=λjξjof the characteristic coordinates, implying ˙/Delta1ξj=λj/Delta1ξj, we obtain,
tolowestorder inthe /Delta1ξj,
˙A=c1,2[/Delta1ξ1λ2/Delta1ξ2+/Delta1ξ2λ1/Delta1ξ1]=c1,2/Delta1ξ1/Delta1ξ2(λ1+λ2). (18.34)
Inthelimit /Delta1ξj→0,wefindfromEq. (18.34)thattherate
˙A
A=λ1+λ2=trace(F)=∇·f|y0, (18.35)
withf=(f1,f2)thevectoroftimeevolutionfunctionsofEq.(18.22b).Notethatthetime-
independent sine of the angle between ξ1andξ2drops out of the rate. The generalization
tohigherdimensionsis obvious.Moreover,in ndimensions,
trace(F)=summationdisplay
iλi. (18.36)
This trace formula follows from the invariance of the secular polynomial in Eq. (18.29)
under a linear transformation, Cξ=xin particular, and it is a result of its determinental
formusingtheproducttheoremfor determinants(see Section3.2), viz.
det(F−λ·1)=bracketleftbig
det(C)bracketrightbig−1det(F−λ·1)bracketleftbig
det(C)bracketrightbig
=detparenleftbig
C−1(F−λ·1)Cparenrightbig
=detparenleftbig
C−1FC−λ·1parenrightbig
=nproductdisplay
i=1(λi−λ). (18.37)
HeretheproductformcomesaboutbysubstitutingEq. (18.30). Now, trace (F)is thecoef-
ficient of (−λ)n−1upon expanding det (F−λ·1)in powers of λ, while it issummationtext
iλifrom
theproductformproducttext
i(λi−λ),whichprovesEq.(18.36).Clearly,accordingtoEqs.(18.35)
and(18.36),
18.4 Nonlinear Differential Equations 1103
•itisthesign(andmorepreciselythetrace)ofthecharacteristicexponentsofthederiv-
ative matrix at the fixed point that determines whether there is expansion or shrinkage
ofareas andvolumesinhigherdimensionsnearacriticalpoint.
In summary then, Eq. (18.35) states that dissipation requires ∇·f(y)/negationslash=0, where˙yj=fj,
anddoesnotoccurinHamiltoniansystemswhere ∇·f=0.
Moreover,intwoormoredimensions,therearethefollowingglobalpossibilities:
•Thetrajectorymaydescribeaclosedorbit(cycle).
•The trajectory may approach a closed orbit (spiraling inward or outward toward the
orbit)ast→∞.Inthis casewehavealimitcycle.
The local behavior of a trajectory near a critical point is also more varied in general
than in one dimension: At a stable critical point all trajectories may approach the critical
point along straight lines or spiral inward (toward the spiral node )o rm a yf o l l o wam o r e
complicated path. If all time-reversed trajectories move toward the critical point in spirals
ast→−∞, then the critical point is a divergent spiral point, or spiral repellor . When
some trajectories approach the critical point while others move away from it, then it is
calledasaddlepoint .Whenalltrajectoriesformclosedorbitsaboutthecriticalpoint,itis
calledacenter.
Bifurcations in Dynamical Systems
A bifurcation is a sudden change in dynamics for specific parameter values, such as the
birthofanode–repellorpairoffixedpointsortheirdisappearanceuponadjustingacontrol
parameter; that is, the motions before and after the bifurcation are topologically different.
At a bifurcation point, not only are solutions unstable when one or more parameters are
changed slightly, but the character of the bifurcation in phase space or in the parameter
manifold may change. Thus we are dealing with fairly sudden events of nonlinear dynam-
ics. Rather sudden changes from regular to random behavior of trajectories are character-
istic of bifurcations, as is sensitive dependence on initial conditions: Nearby initial condi-
tions can lead to very different long-term behavior. If a bifurcation does not change quali-
tatively with parameter adjustments, it is called structurally stable . Note that structurally
unstable bifurcations are unlikely to occur in reality because noise and other neglected
degrees of freedom act as perturbations on the system that effectively eliminate unstable
bifurcations from our view. Bifurcations (such as doublings in maps) are important as one
among many routes to chaos. Others are sudden changes in trajectories associated with
several critical points called global bifurcations . Often they involve changes in basins of
attractionand/orotherglobalstructures.Thetheoryofglobalbifurcationsisfairlycompli-
catedandis stillinitsinfancyatpresent.
Bifurcations that are linked to sudden changes in the qualitative behavior of dynamical
systems at a single fixed point are called local bifurcations . More specifically, a change
in stability occurs in parameter space where the real part of a characteristic exponent of
thefixedpointaltersitssign,thatis,movesfromattractingtorepellingtrajectories,orvice
versa. The center–manifold theorem says that at a local bifurcation only those degrees
1104 Chapter 18 Nonlinear Methods and Chaos
of freedom matter that are involved with characteristic exponents going to zero: ℜλi=0.
Locating the set of these points is the first step in a bifurcation analysis. Another step
consists in cataloguing the types of bifurcations in dynamical systems, to which we turn
next.
The conventional normal forms of dynamical equations represent a start in classifying
bifurcations. For systems with one parameter (that is, a one-dimensional center manifold)
wewritethegeneralcaseofNDEas follows:
˙x=∞summationdisplay
j=0a(0)
jxj+c∞summationdisplay
j=0a(1)
jxj+c2∞summationdisplay
j=0a(2)
jxj+···, (18.38)
wherethesuperscriptonthe a(m)denotesthepoweroftheparameter ctheyareassociated
with. One-dimensional iterated nonlinear maps such as the logistic map of Section 18.2
(which occur in Poincaré sections) of nonlinear dynamical systems can be classified simi-
larly,viz.
xn+1=∞summationdisplay
j=0a(0)
jxj
n+c∞summationdisplay
j=0a(1)
jxj
n+c2∞summationdisplay
j=0a(2)
jxj
n+···. (18.39)
Thus,oneofthesimplestNDEswithabifurcationis
˙x=x2−c, (18.40)
which corresponds to all a(m)
j=0 except for a(1)
0=−1 anda(0)
2=1. Forc>0, there are
two fixed points (recall, ˙x=0)x±=±√cwith characteristic exponents 2 x±,s ox−is a
node and x+is a repellor. For c<0 there are no fixed points. Therefore, as c→0t h e
fixed point pair disappears suddenly; that is, the parameter value c=0 is a repellor-node
bifurcation that is structurally unstable. This complex map (with c→−c) generates the
fractalJuliaandMandelbrotsets discussedinSection18.3.
Apitchforkbifurcationoccursfortheundamped(nondissipativeandspecialcaseofthe
Duffing)oscillatorwithacubicanharmonicity
¨x+ax+bx3=0,b>0. (18.41)
It has a continuous frequency spectrum and is, among others, a model for a ball bouncing
between two walls. When the control parameter a>0, there is only one fixed point, at
x=0, a node, while for a<0 there are two more nodes, at x±=±√−a/b.T h u s ,w e
havea pitchforkbifurcationof anodeattheorgin intoasaddlepointattheoriginandtwo
nodes, at x±/negationslash=0. In terms of a potential formulation, V(x)=ax2/2+bx4/4 is a single
wellfora>0 butadoublewell(withamaximumat x=0)fora<0.
Whenapairofcomplexconjugatecharacteristicexponents ρ±iκcrossesfromaspiral
node (ρ<0) to a repelling spiral ( ρ>0) and periodic motion (limit cycle) emerges, then
we call the qualitative change a Hopf bifurcation . They occur in the quasiperiodic route
tochaosthatwillbediscussedinthenextsection,onchaos.
In a global analysis we piece together the motions near various critical points, such as
nodes and bifurcations, to bundles of trajectories that flow more or less together in two
dimensions.(Thisgeometricviewisthecurrentmodeofanalyzingsolutionsofdynamical
systems.) But this flow is no longer collective in the case of three dimensions, where they
diverge from each other in general, because chaotic motion is possible that typically fills
theplaneofaPoincarésectionwithpoints.
18.4 Nonlinear Differential Equations 1105
Chaos in Dynamical Systems
Our previous summaries of intricate and complicated features of dynamical systems due
to nonlinearities in one and two dimensions do not include chaos, although some of them,
such as bifurcations, sometimes are precursors to chaos. In three- or more-dimensional
NDEs, chaoticmotionmayoccur, often whena constantof the motion(an energyintegral
forNDEsdefinedbyaHamiltonian,forexample)restrictsthetrajectoriestoafinitevolume
inphasespaceandwhentherearenocriticalpoints.Anothercharacteristicsignalforchaos
iswhenforeachtrajectorytherearenearbyones,someofwhichmoveawayfromit,while
others approach it with increasing time. The notion of exponential divergence of nearby
trajectories is made quantitative by the Lyapunov exponent λ(see Section 18.3 for more
details)ofiteratedmapsofPoincarésectionsassociatedwiththedynamicalsystem.Iftwo
nearby trajectories are at a distance d0at timet=0 but diverge with a distance d(t)at a
latertime t,thend(t)≈d0eλtholds.Thus,byanalyzingtheseriesofpoints,thatis,iterated
maps generated on Poincaré sections, one can study routes to chaos of three-dimensional
dynamical systems. This is the key method for studying chaos. As one varies the location
and orientation of the Poincaré plane, a fixed point on it often is recognized to originate
from a limit cycle in the three-dimensional phase space whose structural stability can be
checked there. For example, attracting limit cycles show up as nodes in Poincaré sections,
repelling limit cycles as repellors of Poincaré maps, and saddle cycles as saddle points of
associatedPoincarémaps.
Threeormoredimensionsofphasespacearerequiredforchaostooccurbecauseofthe
interplayofthenecessaryconditionswejustdiscussed,viz.
•boundedtrajectories(areoftenthecaseforHamiltoniansystems),
•exponential divergence of nearby trajectories (is guaranteed by positive Lyapunov ex-
ponentsof correspondingPoincarémaps),
•nointersectionoftrajectories.
The last condition is obeyed by deterministic systems in particular, as we discussed in
Section 18.1. A surprising feature of chaos, mentioned in Section 18.1, is how prevalent
it is and how universal the routes to chaos often are, despite the overwhelming variety of
NDEs.
An example for spatially complex patterns in classical mechanics is the planar pendu-
lum,whoseone-dimensionalequationofmotion
Idθ
dt=L,dL
dt=−lmgsinθ (18.42)
isnonlinearinthedynamicvariable θ(t).HereIisthemomentofinertia, listhedistance
tothecenterofmass, misthemass,and gisthegravitationalaccelerationconstant.When
allparametersinEq.(18.42)areconstantintimeandspace,thenthesolutionsaregivenin
termsofellipticintegrals(seeSection5.8)andnochaosexists.However,apendulumunder
aperiodicexternalforcecanexhibitchaoticdynamics,for example,for theLagrangian
L=m
2˙r2−mg(l−z), (x−x0)2+y2+z2=l2, (18.43)
x0=εlcosωt. (18.44)
1106 Chapter 18 Nonlinear Methods and Chaos
(SeeMoon(1992)intheAdditionalReadings.)
Goodcandidatesfor chaosaremultiplewellpotentialproblems,
d2r
dt2+∇V(r)=Fparenleftbigg
r,dr
dt,tparenrightbigg
, (18.45)
whereFrepresentsdissipativeand/ordrivingforces.Anotherclassicexampleisrigid-body
rotation,whosenonlinearthree-dimensionalEulerequationsarefamiliar,viz.
d
dtI1ω1=(I2−I3)ω2ω3+M1,
d
dtI2ω2=(I3−I1)ω1ω3+M2, (18.46)
d
dtI3ω3=(I1−I2)ω1ω2+M3.
HeretheIjaretheprincipalmomentsofinertiaand ωistheangularvelocitywithcompo-
nentsωjaboutthebody-fixedprincipalaxes.Evenfreerigid-bodyrotationcanbechaotic,
foritsnonlinearcouplingsandthree-dimensionalformsatisfyallrequirementsforchaosto
occur (see Section 18.1). A rigid-body exampleof chaos in our solar system is the chaotic
tumbling of Hyperion, one of Saturn’s moons that is highly nonspherical. It is a world
where the Saturn rise and set is so irregular as to be unpredictable. Another is Halley’s
comet, whose orbit is perturbed by Jupiter and Saturn. In general, when three or more ce-
lestial bodies interact gravitationally, stochastic dynamics are possible. Note, though, that
computer simulations over large time intervals are required to ascertain chaotic dynamics
in the solar system. For more details on chaos in such conservative Hamiltonian systems
werefer toChapter8ofHilborn(1994)intheAdditionalReadings.
Exercise
18.4.3 ConstructaPoincarémapfor theDuffingoscillatorinEq. (18.41).
Routes to Chaos in Dynamical Systems
Letusnowlookatsomeroutestochaos.Theperiod-doublingroutetochaosisexemplified
by the logistic map in Section 18.2, and the universal Feigenbaum numbers α,δare its
quantitativefeatures,alongwithLyapunovexponents.Itiscommonindynamicalsystems.
Itmaybeginwithlimitcycle(periodic)motionthatshowsupasafixedpointinaPoincaré
section. The limit cycle may have originated in a bifurcation from a node or some other
fixedpoint.Asacontrolparameterchanges,thefixedpointofthePoincarémapsplitsinto
two points; that is, the limit cycle has a characteristic exponent going through zero from
attracting to repelling, say. The periodic motion now has a period twice as long as before,
etc. We refer to Chapter 11 of Barger and Olsson (1995) in the Additional Readings for
period-doublingplotsofPoincarésectionsfortheDuffingequation(18.41)withaperiodic
externalforce.Anotherexampleforperioddoublingisaforcedoscillatorwithfriction(see
HellemaninCvitanovic(1989)intheAdditionalReadings).
18.4 Additional Readings 1107
Thequasiperiodicroutetochaosisalsoquitecommonindynamicalsystems,forexam-
ple,startingfrom atime-independentnode,a fixedpoint.If weadjustacontrolparameter,
the system undergoes a Hopf bifurcation to the periodic motion corresponding to a limit
cycle in phase space. With further change of the control parameter, a second frequency
appears. If the frequency ratio is an irrational number, the trajectories are quasiperiodic,
eventuallycoveringthesurfaceofatorusinphasespace;thatis,quasiperiodicorbitsnever
close or repeat. Further changes of the control parameter may lead to a third frequency or
directly to chaotic motion. Bands of chaotic motion can alternate with quasiperiodic mo-
tion in parameter space. An example for such a dynamic system is a periodically driven
pendulum.
A third route to chaos goes via intermittency, where the dynamical system switches
between two qualitatively different motions at fixed control parameters. For example, at
thebeginning,periodicmotionalternateswithanoccasionalburstofchaoticmotion.With
achangeofthecontrolparameter,thechaoticburststypicallylengthenuntil,eventually,no
periodicmotionremains.Thechaoticpartsareirregularanddonotresembleeachother,but
oneneedstocheckforapositiveLyapunovexponenttodemonstratechaos.Intermittencies
of various types are common features of turbulent states in fluid dynamics. The Lorenz
coupledNDEsalsoshowintermittency.
Exercise
18.4.4 Plot the intermittency region of the logistic map at µ=3.8319. What is the period of
thecycles?Whathappensat µ=1+2√
2?
ANS.Thereisatangentbifurcationtoperiod3cycles.
AdditionalReadings
Amann, H., Ordinary Differential Equations: An Introduction To Nonlinear Analysis . New York: de Gruyter
(1990).
Baker, G. L., and J. P. Gollub, Chaotic Dynamics: An Introduction , 2nd ed. Cambridge, UK: Cambridge Univer-
sity Press (1996).
B ar g er ,V .D.,an dM.G.Ol s s o n , Classical Mechanics , 2nd ed. NewYork: McGraw-Hill (1995).
Bender, C. M., and S. A. Orszag, Advanced Mathematical Methods For Scientists and Engineers .N e wY o r k :
McGraw-Hill(1978), Chapter 4inparticular.
Bergé,P., Y.Pomeau, andC.Vidal, Order within Chaos . NewYork: Wiley (1987).
Cvitanovic, P.,ed., Universalityin Chaos , 2nd ed.Bristol, UK:AdamHilger(1989).
Devaney,R.L., AnIntroductiontoChaoticDynamicalSystems .MenloPark,CA:Benjamin/Cummings;2nded.,
Perseus (1989).
Earnshaw, J.C., and D.Haughey, Lyapunov exponents for pedestrians. Am.J .Ph ys. 61: 401 (1993).
Gleick,J., Chaos. New York: Penguin Books (1987).
Hilborn, R.C., Chaos and Nonlinear Dynamics .NewYork: Oxford University Press (1994).
Hirsch,M.W.,andS.Smale, DifferentialEquations, DynamicalSystems,and Linear Algebra .NewYork: Acad-
emic Press (1974).
Infeld,E.,andG.Rowlands, NonlinearWaves,SolitonsandChaos .Cambridge,UK:CambridgeUniversityPress
(1990).
1108 Chapter 18 Nonlinear Methods and Chaos
Jackson, E.A., Perspectivesof Nonlinear Dynamics .Cambridge, UK:Cambridge University Press (1989).
Jordan,D.W.,andP.Smith, Nonlinear OrdinaryDifferentialEquations , 2nded.Oxford, UK:Oxford University
Press (1987).
Lyapunov, A.M., The General Problemof the Stability of Motion . Bristol, PA: Taylor & Francis (1992).
Mandelbrot, B. B., The Fractal Geometryof Nature . San Francisco: W. H.Freeman, reprinted (1988).
Moon, F.C., Chaotic and Fractal Dynamics .NewYork: Wiley (1992).
Peitgen,H.-O.,andP. H.Richter, The Beautyof Fractals . NewYork: Springer (1986).
Sachdev,P. L., Nonlinear Differential Equations and their Applications . NewYork: MarcelDekker(1991).
Tufillaro,N.B.,T.Abbott,andJ.Reilly, AnExperimentalApproachtoNonlinearDynamicsandChaos .Redwood
City, CA: Addison-Wesley (1992).
CHAPTER 19
PROBABILITY
Probabilitiesariseinmanyproblemsdealingwithrandomeventsorlargenumbersofparti-
clesdefiningrandomvariables.Aneventiscalled randomifitispracticallyimpossibleto
predict from the initial state. This includes those cases where we have merely incomplete
information about initial states and/or the dynamics, as in statistical mechanics, where we
may know the energy of the system that corresponds to very many possible microscopic
configurations,preventingusfrompredictingindividualoutcomes.Oftentheaverageprop-
ertiesofmanysimilareventsarepredictable,asinquantumtheory.Thisiswhyprobability
theorycanbeandhasbeendeveloped.
Randomvariablesareinvolvedwhendatadependonchance,suchasweatherreportsand
stockprices.Thetheoryofprobabilitydescribesmathematicalmodelsofchanceprocesses
in terms of probability distributions of random variables that describe how some “random
events” are more likely than others. In this sense probability is a measure of our igno-
rance, giving quantitative meaning to qualitative statements such as “It will probably rain
tomorrow” and “I’m unlikely to draw the heart queen.” Probabilities are of fundamental
importance in quantum mechanics and statistical mechanics and are applied in meteorol-
ogy,economics,games,andmanyotherareasof dailylife.
Toamathematician,probabilitiesarebasedonaxioms,butwewilldiscussherepractical
ways of calculating probabilities for random events. Because experiments in the sciences
are always subject to errors, theories of errors and their propagation involve probabilities.
Instatisticswedealwiththeapplicationsofprobabilitytheorytoexperimentaldata.
19.1 D EFINITIONS ,SIMPLE PROPERTIES
All possible mutually exclusive1outcomes of an experiment that is subject to chance
represent the events (or points) of the sample space S. For example, each time we toss a
coinwe givethe trial a number i=1,2,...andobserve the outcomes xi. Here the sample
1This means that given thatone particular event did occur, theothers could not have occurred.
1109
1110 Chapter 19 Probability
consists of two events: heads and tails, and the xirepresent a discrete random variable
that takes on one of two values, heads or tails. When two coins are tossed, the sample
contains the events two heads, one head and one tail, two tails; the number of heads is a
good value to assign to the random variable, so the possible values are 2, 1, and 0. There
are four equally probable outcomes, of which one has value 2, two have value 1, and one
hasvalue0.Sotheprobabilitiesofthethreevaluesoftherandomvariableare 1 /4f o rt w o
heads(value2), 1 /4 for noheads(value0),and 1 /2 for value1.In otherwords, wedefine
thetheoreticalprobability Pof aneventdenotedbythepoint xiofthesampleas
P(xi)≡numberofoutcomesof event xi
totalnumberofallevents. (19.1)
An experimentaldefinition applies when the total number of events is not well defined (or
isdifficulttoobtain)orequallylikelyoutcomesdonotalways occur.Then
P(xi)≡numberof timesevent xioccurs
totalnumberoftrials(19.2)
is more appropriate. A large, thoroughly mixed pile of black and white sand grains of the
samesizeandinequalproportionsisarelevantexample,becauseitisimpracticaltocount
themall.Butwecancountthegrainsinasmallsamplevolumethatwepick.Thiswaywe
cancheckthatwhiteandblackgrainsturnupwithroughlyequalprobability1 /2,provided
we put back each sample and mix the pile again. It is found that the larger the sample
volume, the smaller the spread about 1 /2 will be. The more trials we run, the closer the
average occurrence of all trial counts will be to 1 /2.We could even pick single grains and
check if the probability 1 /4 of picking two black grains in a row equals that of two white
grains,etc.Therearelotsofstatisticsquestionswecanpursue.Thus,pilesofcoloredsand
provideforinstructiveexperiments.
Thefollowingaxiomsareself-evident.
•Probabilities satisfy 0 ≤P≤1.Probability 1 means certainty; probability 0 means
impossibility.
•The entire sample has probability 1 .For example, drawing an arbitrary card has prob-
ability 1.
•The probabilities for mutually exclusive events add. The probability for getting one
headintwocointossesis1 /4+1/4=1/2becauseitis1 /4forheadfirstandthentail,
plus 1/4 for tailfirst andthenhead.
Example 19.1.1 PROBABILITY FOR AORB
Whatistheprobabilityfordrawing2acluborajackfromashuffleddeckofcards?Because
there are 52 cards in a deck, each being equally likely, 13 cards for each suit and 4 jacks,
there are 13 clubs including the club jack, and 3 other jacks; that is, there are 16 possible
cardsoutof 52,givingtheprobability (13+3)/52=16/52=4/13. /squaresolid
2Theseareexamples of non-mutually exclusive events.
19.1 Definitions, Simple Properties 1111
Ifwerepresentthesamplespacebyaset Sofpoints,theneventsaresubsets A,B,...of
S,denotedas A⊂S,etc.Twosets A,Bareequalif Aiscontainedin B,A⊂B,andBis
containedin A,B⊂A.TheunionA∪Bconsistsofallpoints(events)thatarein AorB
or both (see Fig. 19.1). The intersection A∩Bconsists of all points, that are in both A
andB.IfAandBhavenocommonpoints,theirintersectionisthe emptyset ,A∩B=∅,
whichhas noelements(events). The set of pointsin Athat are notin theintersectionof A
andBisdenotedby A−A∩B,definingasubtractionofsets .Ifwetaketheclubsuitin
Example 19.1.1 as set Aand the four jacks as set B,then their union comprises all clubs
andjacks,andtheirintersectionistheclubjackonly.
Each subset Ahas its probability P(A)≥0.In terms of these set theory concepts and
notations,theprobabilitylawswejustdiscussedbecome
0≤P(A)≤1.
The entire sample space has P(S)=1. The probability of the union A∪Bof mutually
exclusiveeventsis thesum
P(A∪B)=P(A)+P(B), A∩B=∅.
Theadditionrule for probabilitiesofarbitrarysets isgivenbythefollowingtheorem.
ADDITION RULE :
P(A∪B)=P(A)+P(B)−P(A∩B). (19.3)
To prove this, we decompose the union into two mutually exclusive sets A∪B=A∪
(B−B∩A),subtracting the intersection of AandBfromBbefore joining them. Their
probabilitiesare P(A),P(B)−P(B∩A),whichweadd.Wecouldalsohavedecomposed
A∪B=(A−A∩B)∪B,from which our theorem follows similarly by adding these
probabilities, P(A∪B)=[P(A)−P(A∩B)]+P(B). Note that A∩B=B∩A.(See
Fig.19.1.)
Sometimestherulesanddefinitionsofprobabilitiesthatwehavediscussedsofararenot
sufficient,however.
FIGURE 19.1Theshadedareagivesthe
intersection A∩B, correspondingtothe
AandBevents,thedashedlineencloses
A∪B, correspondingtothe AorB
events.
1112 Chapter 19 Probability
Example 19.1.2 CONDITIONAL PROBABILITY
Asimpleexampleconsistsofaboxof10identicalredand20identicalbluepens,arranged
in random order, from which we remove pens successively, that is, without putting them
back.Supposewedrawaredpenfirst,event A.Thatwillhappenwithprobability P(A)=
10/30=1/3 if the pens are thoroughly mixed up. The conditional probability P(B|A)
of drawing a blue pen in the next round, event B,however, will depend on the fact that
we drew a red pen in the first round. It is given by 20 /29.There are 10·20 possible
sample points (red/blue pen events) in two rounds, and the sample has 30 ·29 events, so
thecombinedprobabilityis
P(A,B)=10
3020
29=10·20
30·29=20
87. /squaresolid
In general, the combined probability P(A,B) thatAandBhappen (in this order) is
given by the product of the probability that Ahappens, P(A),and the probability that B
happensif Adoes,P(B|A):
P(A,B)=P(A)P(B|A). (19.4)
Inotherwords, the conditionalprobability P(B|A)is givenbytheratio
P(B|A)=P(A,B)
P(A). (19.5)
If the conditionalprobability P(B|A)=P(B)is independentof A,then the events Aand
Barecalled independent ,andthecombinedprobability
P(A∩B)=P(A)P(B) (19.6)
issimplythe productofbothprobabilities .
Example 19.1.3 SCHOLASTIC APTITUDE TESTS
CollegesanduniversitiesrelyontheverbalandmathematicsSATscores,amongothers,as
predictors of a student’s success in passing courses and graduating. A research university
is known to admit mostly students with a combined verbal and mathematics score above
1400 points. The graduation rate is 95%; that is, 5% drop out or transfer elsewhere. Of
thosewhograduate,97%haveanSATscoreofmorethan1400points,while80%ofthose
who drop out have an SAT score below 1400 .Suppose a student has an SAT score below
1400.Whatis his/herprobabilityof graduating?
LetAbethecaseshavinganSATtestscorebelow1400 ,Brepresentthoseabove1400 ,
mutuallyexclusiveeventswith P(A)+P(B)=1,andCbethosestudentswhograduate.
That is, we want to know the conditional probabilities P(C|A)andP(C|B).To apply
Eq. (19.5) we need P(A)andP(B).There are 3% of students with scores below 1400
amongthosewhograduate(95%)and80%ofthose5%whodonotgraduate,so
P(A)=0.03·0.95+4
50.05=0.0685,P(B)=0.97·0.95+0.05
5=0.9315,
19.1 Definitions, Simple Properties 1113
andalso
P(C∩A)=0.03·0.95=0.0285 and P(C∩B)=0.97·0.95=0.9215.
Here the combined probabilities P(C,A)=P(C∩A),P(C,B)=P(C∩B)asCandA
(andCandB) arepartsofthesamesamplespace.Therefore,
P(C|A)=P(C∩A)
P(A)=0.0285
0.0685∼41.6%,
P(C|B)=P(C∩B)
P(B)=0.9215
0.9315∼98.9%;
that is, a little less than 42% is the probability for a student with a score below 1400 to
graduateatthisparticularuniversity. /squaresolid
As a corollary to the definition of a conditional probability, Eq. (19.5), we compare
P(A|B)=P(A∩B)/P(B) andP(B|A)=P(A∩B)/P(A), whichleadstothefollowing
theorem.
BAYES THEOREM :
P(A|B)=P(A)
P(B)P(B|A). (19.7)
This canbegeneralizedtothefollowing.
THEOREM:If the random events Aiwith probabilities P(Ai)>0are mutually exclusive
andtheirunionrepresentstheentiresample S,thenanarbitraryrandomevent B⊂Shas
theprobability
P(B)=nsummationdisplay
i=1P(Ai)P(B|Ai). (19.8)
FIGURE 19.2Theshadedarea Bis
composedof mutuallyexclusivesubsetsof
Bbelongingalsoto A1,A2,A3,wherethe
Aiaremutuallyexclusive.
1114 Chapter 19 Probability
This decomposition law resembles the expansion of a vector into a basis of unit vectors
defining the components of the vector. This relation follows from the obvious decomposi-
tionB=uniontext
i(B∩Ai),Fig. 19.2, which implies P(B)=summationtext
iP(B∩Ai)for the probabil-
ities because the components B∩Aiare mutually exclusive. For each i, we know from
Eq.(19.5) that P(B∩Ai)=P(Ai)P(B|Ai),whichprovesthetheorem.
Counting of Permutations and Combinations
Countingparticlesinsamplescanhelpusfindprobabilities,asinstatisticalmechanics.
If we have ndifferent molecules, let us ask in how many ways we can arrange them in
arow,thatis,permutethem.Thisnumberisdefinedasthenumberoftheir permutations .
Thus, by definition, the order matters in permutations . There are nchoices of picking
the first molecule, n−1 for the second, etc. Altogether there are n!permutations of n
different moleculesor objects.
Generalizingthis,supposethereare npeoplebutonly k<nchairstoseatthem.Inhow
manywayscanweseat kpeopleinthechairs?Countingasbefore, weget
n(n−1)···(n−k+1)=n!
(n−k)!
forthenumberof permutationsof ndifferentobjects, katatime.
Wenowconsiderthenumberof combinations ofobjectswhentheir orderisirrelevant
by definition. For example, three letters a,b,ccan be combined, two letters at a time, in
3=3!
2!ways:ab,ac,bc.Ifletterscanberepeated,thenweaddthepairs aa,bb,ccandhave
sixcombinations.Thus,a combination ofdifferentparticlesdiffersfromapermutationin
thattheir orderdoesnotmatter .Combinationsoccurwithrepetition(themathematician’s
wayoftreatingindistinguishableobjects)andwithout,wherenotwosetscontainthesame
particles.
Thenumberofdifferentcombinationsof nparticles, katatimeandwithoutrepetitions,
isgivenbythebinomialcoefficient
n(n−1)···(n−k+1)
k!=parenleftBign
kparenrightBig
.
If repetitionisallowed,thenthenumberis
parenleftbiggn+k−1
kparenrightbigg
.
In the number n!/(n−k)!of permutations of nparticles, kat a time, we have to divide
out the number k!of permutations of the groups of kparticles because their order does
not matter in a combination. This proves the first claim. The second one is shown by
mathematicalinduction.
In statistical mechanics, we ask in how many ways we can put nparticles in kboxes so
that there will be ni(distinguishable) particles in the ith box, without regard to order in
each box, withsummationtextk
i=1ni=n.Counting as before, there are nchoices for selecting the first
particle,n−1forpickingthesecond,etc.,butthe n1!permutationswithinthefirstboxare
19.1 Definitions, Simple Properties 1115
discounted,and n2!permutationswithinthesecondboxaredisregarded,etc.Thereforethe
numberofcombinationsis
n!
n1!n2!···nk!,n 1+n2+···+nk=n.
In statisticalmechanics,particlesthatobey
•Maxwell–Boltzmann (MB) statistics are distinguishable, without restriction on their
numberineachstate;
•Bose–Einstein (BE) statistics are indistinguishable, with no restriction on the number
ofparticlesineachquantumstate;
•Fermi–Dirac(FD)statisticsareindistinguishable,withatmostoneparticleper state.
Forexample,puttingthreeparticlesinfourboxes,thereare43equallylikelyarrangements
fortheMBcase,becauseeachparticlecanbeputintoanyboxinfourways,givingatotalof
43choices.ForBEstatistics,thenumberofcombinationswithrepetitionsisparenleftbig3+4−1
3parenrightbig
=parenleftbig6
3parenrightbig
for the Bose–Einstein case. For FD statistics, it isparenleftbig3+1
3parenrightbig
=parenleftbig4
3parenrightbig
. More generally, for MB
statistics the number of distinct arrangements of nparticles among kstates (boxes) is kn,
forBE statisticsitisparenleftbign+k−1
nparenrightbig
, andfor FDstatisticsitisparenleftbigk
nparenrightbig
.
Exercises
19.1.1 A card is drawn from a shuffled deck. (a) What is the probability that it is black, (b) a
rednine,(c) oraqueenof spades?
19.1.2 Find the probability of drawing two kings from a shuffled deck of cards (a) if the first
cardisputbackbeforethesecondisdrawn,and(b)ifthefirstcardisnotputbackafter
beingdrawn.
19.1.3 When two fair dice are thrown, what is the probability of (a) observing a number less
than 4 or(b) anumbergreaterthanor equalto 4 butlessthan 6?
19.1.4 Rollingthreefair dice,whatis theprobabilityofobtainingsixpoints?
19.1.5 Determinetheprobability P(A∩B∩C)interms of P(A),P(B),P(C), etc.
19.1.6 Determine directly or by mathematical induction the probability of a distribution of N
(Maxwell–Boltzmann) particles in kboxes with N1in box 1, N2in box 2,...,N kin
thekthboxforanynumbers Nj≥1withN1+N2+···+Nk=N,k<N.Repeatthis
for Fermi–DiracandBose–Einsteinparticles.
19.1.7 Showthat P(A∪B∪C)=P(A)+P(B)+P(C)−P(A∩B)−P(A∩C)
−P(B∩C)+P(A∩B∩C).
19.1.8 Determinetheprobabilitythatapositiveinteger n≤100isdivisiblebyaprimenumber
p≤100.Verifyyourresultfor p=3,5,7.
19.1.9 PuttwoparticlesobeyingMaxwell–Boltzmann(Fermi–Dirac,orBose–Einstein)statis-
ticsinthreeboxes.Howmanywaysarethereineachcase?
1116 Chapter 19 Probability
19.2 R ANDOM VARIABLES
Each time we toss a die, we give the trial a number i=1,2,...and observe the point
xi=1,or 2, 3, 4, 5, 6 with probability 1 /6.Ifidenotes the trial number, then xiis a
discreterandomvariablethattakesthediscretevaluesfrom1to6withadefiniteprobability
P(xi)=1/6.
Example 19.2.1 DISCRETE RANDOM VARIABLE
If we toss two dice and record the sum of the points shown in each trial, then this sum is
also a discrete random variable, which takes on the value 2 when both dice show 1 with
probability (1/6)2; the value 3 when one die has 1 and the other 2 ,hence with proba-
bility(1/6)2+(1/6)2=1/18; the value 4 when both dice have 2 or one has 1 and the
other3,sowithprobability (1/6)2+(1/6)2+(1/6)2=1/12;thevalue5withprobability
4(1/6)2=1/9; the value 6 with probability 5 /36; the value 7 with the maximum proba-
bility, 6(1/6)2=1/6; up to the value 12 when both dice show 6 points with probability
(1/6)2.Thisprobabilitydistributionissymmetricabout7 .Thissymmetryisobviousfrom
Fig. 19.3 and becomes visible algebraically when we write the rising and falling linear
partsas
P(x)=x−1
36=6−(7−x)
36,x=2,3,...,7,
P(x)=13−x
36=6+(7−x)
36,x=7,8,...,12./squaresolid
In summary,then,
•The different values xithat a random variable Xassumes denote and distinguish the
eventsinthesamplespaceofanexperiment;eacheventoccursbychancewithaprob-
FIGURE 19.3Probabilitydistribution P(x)of thesumof pointswhentwo
diceare tossed.
19.2 Random Variables 1117
abilityP(X=xi)=pi≥0 that is a function of the random variable X.A random
variableX(ei)=xiis definedonthesamplespace,thatis, fortheevents ei∈S.
•Wedefinetheprobabilitydensity f(x)ofacontinuousrandomvariable Xas
P(x≤X≤x+dx)=f(x)dx; (19.9)
that is,f(x)dxis the probability that Xlies in the interval x≤X≤x+dx.For
f(x)to be a probability density, it has to satisfy f(x)≥0 andintegraltext
f(x)dx=1.The
generalization to probability distributions depending on several random variables is
straightforward.Quantumphysicsaboundsinexamples.
Example 19.2.2 CONTINUOUS RANDOM VARIABLE :HYDROGEN ATOM
Quantum mechanics gives the probability |ψ|2d3rof finding a 1 selectron in a hydrogen
atominvolume3d3r,whereψ=Ne−r/ais thewavefunctionthatis normalizedto
1=integraldisplay
|ψ|2dV=4πN2integraldisplay∞
0e−2r/ar2dr=πa3N2,dV=r2drdcosθdϕ
being the volume element and athe Bohr radius. The radial integral is found by repeated
integrationbypartsor byrescalingittothegammafunction
integraldisplay∞
0e−2r/ar2dr=parenleftbigga
2parenrightbigg3integraldisplay∞
0e−xx2dx=a3
8Ŵ(3)=a3
4.
Here all points in space constitute the sample and represent three random variables, but
theprobabilitydensity |ψ|2in thiscase dependsonlyon theradial variablebecauseof the
sphericalsymmetryof the 1 sstate.
A measure for the size of the Hatom is givenby theaverageradial distanceof the elec-
tron from the proton at the center, which in quantum mechanics is called the expectation
value:
/angbracketleft1s|r|1s/angbracketright=integraldisplay
r|ψ|2dV=4πN2integraldisplay∞
0re−2r/ar2dr=3
2a.
Weshalldefinethisconceptforarbitraryprobabilitydistributionsshortly. /squaresolid
•A random variable that takes only discrete values x1,x2,...,xnwith probabilities
p1,p2,...,pn,respectively, is called a discrete random variable, sosummationtext
ipi=1. If an
“experiment”ortrialis performed,someoutcomemustoccur,withunitprobability.
•If the values comprise a continuous range of values a≤x≤b,then we deal with
a continuous random variable, whose probability distribution may or may not be a
continuousfunctionaswell.
3Notethat|ψ|24πr2drgives the probability for the electronto befound between randr+dr, at anyangle.
1118 Chapter 19 Probability
When we measure a quantity xntimes, obtaining the values xj,we define the average
value
¯x=1
nnsummationdisplay
j=1xj (19.10)
of the trials, also called the meanorexpectation value , where this formula assumes that
everyobservedvalue xiisequallylikelyandoccurswithprobability1 /n.Thisconnection
isthekeylinkofexperimentaldatawithprobabilitytheory.Thisobservationandpractical
experiencesuggestdefiningthe meanvaluefor adiscreterandomvariable Xas
/angbracketleftX/angbracketright≡summationdisplay
ixipi (19.11)
andthatfor a continuousrandomvariable characterizedbyprobabilitydensity f(x)as
/angbracketleftX/angbracketright=integraldisplay
xf(x)dx. (19.12)
Thesearelinearaverages.Othernotationsintheliteratureare ¯XandE(X).
The use of the arithmetic mean ¯xofnmeasurements as the average value is suggested
bysimplicityandplainexperience,assumingequalprobabilityforeach xiagain.Butwhy
dowenotconsiderthegeometricmean
xg=(x1·x2·····xn)1/n(19.13)
ortheharmonicmean xhdeterminedbytherelation
1
xh=1
nparenleftbigg1
x1+1
x1+···+1
xnparenrightbigg
(19.14)
or that value˜xthat minimizes the sum of absolute deviations |xi−˜x|? Here the xiare
taken to increase monotonically. When we plot O(x)=summationtext2n+1
i=1|xi−x|, as in Fig. 19.4a,
for an odd number of points, we realize that it has a minimum at its central value i=n,
FIGURE 19.4(a)summationtext3
i=1|xi−x|for anodd
numberofpoints;(b)summationtext4
i=1|xi−x|foraneven
numberof points.
19.2 Random Variables 1119
while for an even number of points E(x)=summationtext2n
i=1|xi−x|is flat in its central region, as
shown in Fig. 19.4b. These properties make these functions unacceptable for determining
averagevalues.Instead, whenweminimizethesumofquadraticdeviations,
nsummationdisplay
i=1(x−xi)2=minimum , (19.15)
settingthederivativeequaltozeroyields 2summationtext
i(x−xi)=0,or
x=1
nsummationdisplay
ixi≡¯x,
thatis,thearithmeticmean.Ithasanotherimportantproperty:Ifwedenoteby vi=xi−¯x
the deviations, thensummationtext
ivi=0,that is, the sum of positive deviations equals the sum of
negative deviations. This principle of minimizing the quadratic sum of deviations, called
themethodof leastsquares ,isduetoC. F.Gauss, amongothers.
How close a fit of the mean value to a set of data points is depends on the spread of the
individual measurements from this mean. Again, we reject the average sum of deviationssummationtextn
i=1|xi−¯x|/nas a measure of the spread because it selects the central measurement
as the best value for no good reason. A more appropriate definition of the spread is the
averageof thedeviationsfrom themean,squared,or standarddeviation
σ=radicaltpradicalvertexradicalvertexradicalbt1
nnsummationdisplay
i=1(xi−¯x)2,
wherethesquarerootismotivatedbydimensionalanalysis.
Example 19.2.3 STANDARD DEVIATION OF MEASUREMENTS
From the measurements x1=7,x2=9,x3=10,x4=11,x5=13 we extract ¯x=10
for the mean value and σ=√(9+1+1+9)/4=2.2361 for the standard deviation, or
spread,usingtheexperimentalformula(19.2)becausetheprobabilitiesarenotknown. /squaresolid
There is yet another interpretation of the standard variation, in terms of the sum of
squaresofmeasurementdifferences
summationdisplay
i<k(xi−xk)2=1
2nsummationdisplay
i=1nsummationdisplay
k=1parenleftbig
x2
i+x2
k−2xixkparenrightbig
=1
2parenleftbig
2n2angbracketleftbig
x2angbracketrightbig
−2n2/angbracketleftx/angbracketright2parenrightbig
=n2σ2, (19.16)
becausebymultiplyingoutthesquareinthedefinitionof σ2weobtain
σ2=1
nsummationdisplay
iparenleftbig
xi−/angbracketleftx/angbracketrightparenrightbig2=1
nsummationdisplay
ix2
i−2/angbracketleftx/angbracketright
nsummationdisplay
ixi+/angbracketleftx/angbracketright2
=1
nsummationdisplay
ix2
i−/angbracketleftx/angbracketright2=angbracketleftbig
x2angbracketrightbig
−/angbracketleftx/angbracketright2. (19.17)
Thisformulaisoftenusedandwidelyappliedfor σ2.
1120 Chapter 19 Probability
Now we are ready to generalize the spread in a set of nmeasurements with equal prob-
ability 1/nto thevariance of an arbitrary probability distribution. For a discrete random
variableXwithprobabilities piatX=xiwedefinethe variance
σ2=summationdisplay
jparenleftbig
xj−/angbracketleftX/angbracketrightparenrightbig2pj, (19.18)
andsimilarlyfor acontinuousprobabilitydistribution
σ2=integraldisplay∞
−∞parenleftbig
x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx. (19.19)
Thesedefinitionsimplythefollowing.
THEOREM:Ifarandomvariable Y=aX+bislinearlyrelatedto X,thenwecanimme-
diately derive the mean value /angbracketleftY/angbracketright=a/angbracketleftX/angbracketright+band variance σ2(Y)=a2σ2(X)from these
definitions.
Weprovethistheoremonlyforacontinuousdistributionandleavethecaseofadiscrete
random variable as an exercise for the reader. For the infinitesimal probability we know
thatf(x)dx=g(y)dywithy=ax+b,becausethelineartransformationhastopreserve
probability,so
/angbracketleftY/angbracketright=integraldisplay∞
−∞yg(y)dy=integraldisplay∞
−∞(ax+b)f(x)dx=a/angbracketleftX/angbracketright+b,
sinceintegraltext
f(x)dx=1.Forthevariancewesimilarlyobtain
σ2(Y)=integraldisplay∞
−∞parenleftbig
y−/angbracketleftY/angbracketrightparenrightbig2g(y)dy=integraldisplay∞
−∞parenleftbig
ax+b−a/angbracketleftX/angbracketright−bparenrightbig2f(x)dx
=a2σ2(X)
aftersubstitutingourresultfor themeanvalue /angbracketleftY/angbracketright.
Finallyweprovethegeneral Chebychevinequality
Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥kσparenrightbig
≤1
k2, (19.20)
which demonstrates why the standard deviation serves as a measure of the spread of an
arbitrary probability distribution from its mean value /angbracketleftX/angbracketrightand shows why experimental or
other data are often characterized according to their spread in numbers of standard devia-
tions.Wefirst showthesimplerinequality
P(Y≥K)≤/angbracketleftY/angbracketright
K
for a continuous random variable Ywhose values y≥0.(The proof for a discrete random
variablefollowsalongsimilarlines.) Thisinequalityfollowsfrom
/angbracketleftY/angbracketright=integraldisplay∞
0yf(y)dy=integraldisplayK
0yf(y)dy+integraldisplay∞
Kyf(y)dy
≥integraldisplay∞
Kyf(y)dy≥Kintegraldisplay∞
Kf(y)dy=KP(Y≥K).
19.2 Random Variables 1121
Nextweapplythesamemethodtothepositivevarianceintegral
σ2=integraldisplayparenleftbig
x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx≥integraldisplay
|x−/angbracketleftX/angbracketright|≥kσparenleftbig
x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx
≥k2σ2integraldisplay
|x−/angbracketleftX/angbracketright|≥kσf(x)dx=k2σ2Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥kσparenrightbig
,
decreasing the right-hand side first by omitting the part of the positive integral with
|x−/angbracketleftX/angbracketright|≤kσand then again by replacing (x−/angbracketleftX/angbracketright)2in the remaining integral by its
lowest limit, k2σ2.This proves the Chebychev inequality. For k=3 we have the conven-
tionalthree-standard-deviationestimate
Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥3σparenrightbig
≤1
9. (19.21)
It is straightforward to generalize the mean value to higher moments of probability dis-
tributionsrelativetothemeanvalue /angbracketleftX/angbracketright:
angbracketleftbigparenleftbig
X−/angbracketleftX/angbracketrightparenrightbigkangbracketrightbig
=summationdisplay
jparenleftbig
xj−/angbracketleftX/angbracketrightparenrightbigkpj,discretedistribution , (19.22)
angbracketleftbigparenleftbig
X−/angbracketleftX/angbracketrightparenrightbigkangbracketrightbig
=integraldisplay∞
−∞parenleftbig
x−/angbracketleftX/angbracketrightparenrightbigkf(x)dx, continuousdistribution .
Themoment-generatingfunction
angbracketleftbig
etXangbracketrightbig
=integraldisplay
etxf(x)dx=1+t/angbracketleftX/angbracketright+t2
2!angbracketleftbig
X2angbracketrightbig
+··· (19.23)
is a weighted sum of the moments of the continuous random variable Xupon substituting
the Taylor expansion of the exponential functions. So /angbracketleftX/angbracketright=d/angbracketleftetX/angbracketright
dtvextendsinglevextendsingle
t=0.Notice that the
moments here are not relative to the expectation value; they are called central moments .
Thenthcentralmoment, /angbracketleftXn/angbracketright=dn/angbracketleftetX/angbracketright
dtnvextendsinglevextendsingle
t=0,isgivenbythe nthderivativeofthemoment-
generating function at t=0.By a change of the parameter t→itthe moment-generating
function is related to the characteristic function /angbracketlefteitX/angbracketright,which is the often-used Fourier
transformoftheprobabilitydensity f(x).
Moreover, mean values, moments, and variance can be defined similarly for probability
distributions that depend on several random variables. For simplicity, let us restrict our
attentiontotwocontinuousrandomvariables X, Yandlistthecorrespondingquantities:
/angbracketleftX/angbracketright=integraldisplay∞
−∞integraldisplay∞
−∞xf(x,y)dx dy,
/angbracketleftY/angbracketright=integraldisplay∞
−∞integraldisplay∞
−∞yf(x,y)dx dy, (19.24)
σ2(X)=integraldisplay∞
−∞integraldisplay∞
−∞parenleftbig
x−/angbracketleftX/angbracketrightparenrightbig2f(x,y)dxdy,
σ2(Y)=integraldisplay∞
−∞integraldisplay∞
−∞parenleftbig
y−/angbracketleftY/angbracketrightparenrightbig2f(x,y)dxdy. (19.25)
1122 Chapter 19 Probability
Two random variables are said to be independent if the probability density f(x,y)
factorizes into a product f(x)g(y) of probability distributions of one random variable
each.
Thecovariance,definedas
cov(X,Y)=angbracketleftbigparenleftbig
X−/angbracketleftX/angbracketrightparenrightbigparenleftbig
Y−/angbracketleftY/angbracketrightparenrightbigangbracketrightbig
, (19.26)
is a measureof how muchthe randomvariables X, Yare correlated (or related): It is zero
forindependentrandomvariablesbecause
cov(X,Y)=integraldisplayparenleftbig
x−/angbracketleftX/angbracketrightparenrightbigparenleftbig
y−/angbracketleftY/angbracketrightparenrightbig
f(x,y)dxdy
=integraldisplayparenleftbig
x−/angbracketleftX/angbracketrightparenrightbig
f(x)dxintegraldisplayparenleftbig
y−/angbracketleftY/angbracketrightparenrightbig
g(y)dy=parenleftbig
/angbracketleftX/angbracketright−/angbracketleftX/angbracketrightparenrightbigparenleftbig
/angbracketleftY/angbracketright−/angbracketleftY/angbracketrightparenrightbig
=0.
Thenormalizedcovariancecov(X,Y)
σ(X)σ(Y),whichhasvaluesbetween −1and+1,isoftencalled
correlation .
In ordertodemonstratethatthecorrelationisboundedby
−1≤cov(X,Y)
σ(X)σ(Y)≤1
weanalyzethepositivemeanvalue
Q=angbracketleftbigbracketleftbig
uparenleftbig
X−/angbracketleftX/angbracketrightparenrightbig
+vparenleftbig
Y−/angbracketleftY/angbracketrightparenrightbigbracketrightbig2angbracketrightbig
=u2angbracketleftbigbracketleftbig
X−/angbracketleftX/angbracketrightbracketrightbig2angbracketrightbig
+2uvangbracketleftbigbracketleftbig
X−/angbracketleftX/angbracketrightbracketrightbigbracketleftbig
Y−/angbracketleftY/angbracketrightbracketrightbigangbracketrightbig
+v2angbracketleftbigbracketleftbig
Y−/angbracketleftY/angbracketrightbracketrightbig2angbracketrightbig
=u2σ(X)2+2uvcov(X,Y)+v2σ(Y)2≥0, (19.27)
whereu, vare numbers, not functions. For this quadratic form to be nonnegative, its dis-
criminantmustobeycov (X,Y)2−σ(X)2σ(Y)2≤0,whichprovesthedesiredinequality.
Theusefulnessofthecorrelationasaquantitativemeasureisemphasizedbythefollow-
ing.
THEOREM:P(Y=aX+b)=1isvalidif,andonlyif,thecorrelationisequalto ±1.
Thistheoremstatesthata ±100%correlationbetween X, Yimpliesnotonlysomefunc-
tionalrelationbetweenbothrandomvariablesbutalsoa linearrelation betweenthem.We
denoteby/angbracketleftB|A/angbracketrighttheexpectationvalueoftheconditionalprobabilitydistribution P(B|A).
To prove this strong correlation property, we apply Bayes’ decomposition law
(Eq. (19.8)) to the mean value and variance of the random variable Y, assuming first that
P(Y=aX+b)=1,soP(Y/negationslash=aX+b)=0.Thisyields
/angbracketleftY/angbracketright=P(Y=aX+b)/angbracketleftY|Y=aX+b/angbracketright
+P(Y/negationslash=aX+b)/angbracketleftY|Y/negationslash=aX+b/angbracketright
=/angbracketleftaX+b/angbracketright=a/angbracketleftX/angbracketright+b,
σ(Y)2=P(Y=aX+b)angbracketleftbigbracketleftbig
Y−/angbracketleftY/angbracketrightbracketrightbig2vextendsinglevextendsingleY=aX+bangbracketrightbig
+P(Y/negationslash=aX+b)angbracketleftbigbracketleftbig
Y−/angbracketleftY/angbracketrightbracketrightbig2vextendsinglevextendsingleY/negationslash=aX+bangbracketrightbig
=angbracketleftbigbracketleftbig
aX+b−/angbracketleftY/angbracketrightbracketrightbig2angbracketrightbig
=angbracketleftbig
a2bracketleftbig
X−/angbracketleftX/angbracketrightbracketrightbig2angbracketrightbig
=a2σ(X)2,
19.2 Random Variables 1123
substituting/angbracketleftY/angbracketright=a/angbracketleftX/angbracketright+b.Similarly we obtain cov (X,Y)=a2σ(X)2.These results
showthatthecorrelationis ±1.
Conversely, we start from cov (X,Y)2=σ(X)2σ(Y)2.Hence the quadratic form in
Eq.(19.27) mustbezerofor (practically)all xfor some (u0,v0)/negationslash=(0,0):
angbracketleftbigbracketleftbig
u0parenleftbig
X−/angbracketleftX/angbracketrightparenrightbig
+v0parenleftbig
Y−/angbracketleftY/angbracketrightparenrightbigbracketrightbig2angbracketrightbig
=0.
Because the argument of this mean value is positive definite, this relationship is satisfied
only ifP(u0(X−/angbracketleftX/angbracketright)+v0(Y−/angbracketleftY/angbracketright)=0)=1,which means that YandXare linearly
related.
Whenweintegrateoutonerandomvariable,weareleftwiththeprobabilitydistribution
oftheotherrandomvariable,
F(x)=integraldisplay
f(x,y)dy, orG(y)=integraldisplay
f(x,y)dx, (19.28)
andanalogouslyfordiscreteprobabilitydistributions.Whenoneormorerandomvariables
are integrated out, the remaining probability distribution is called marginal, motivated
by the geometric aspects of projection. It is straightforward to show that these marginal
distributionssatisfyalltherequirementsofproperlynormalizedprobabilitydistributions.
If we are interested in the distribution of the randomvariable Xfor a definitevalue y=
y0oftheotherrandomvariable,thenwedealwitha conditionalprobabilitydistribution
P(X=x|Y=y0).Thecorrespondingcontinuousprobabilitydensityis f(x,y0).
Example 19.2.4 REPEATED DRAWS OF CARDS
When we draw cards repeatedly, we shuffle the deck often because we want to make sure
these events stay independent. So we draw the first card at random from a bridge deck
containing 52 cardsandthenput itbackata randomplace.Nowwerepeattheprocess for
asecondcard.Thenthedeckis reshuffled,etc.Wenowdefinetherandomvariables
•X=number ofso-calledhonors,thatis,10s, jacks,queens,kings,oraces;
•Y=number of 2sor 3s.
In a single draw the probability of a 10 to an ace is a=5·4/52=5/13,andb=
2·4/52=2/13fortwoorthreetobedrawnand c=(13−5−2)/13=6/13foranything
else,with a+b+c=1.
In two drawings XandYcan bex=0=y,when no 10 to ace show up or 2 or 3 .This
case has probability c2.In general X,Yhave the values 0, 1, and 2 ,so 0≤x+y≤2
because we will have drawn two cards. The probability function of ( X=x,Y=y)i s
givenbytheproductoftheprobabilitiesof thethreepossibilities ax,by,c2−x−ytimesthe
number of distributions (or permutations) of two cards over the three cases with probabil-
itiesa,b,c, which is 2!/[x!y!(2−x−y)!].This number is the coefficient of the power
axbyc2−x−yinthegeneralizedbinomialexpansionofallpossibilitiesintwodrawingswith
1124 Chapter 19 Probability
probability 1:
1=(a+b+c)2=summationdisplay
0≤x+y≤22!
x!y!(2−x−y)!axbyc2−x−y
=a2+b2+c2+2(ab+ac+bc). (19.29)
Hencetheprobabilitydistributionofourdiscreterandomvariablesisgivenby
f(X=x,Y=y)=2!
x!y!(2−x−y)!parenleftbigg5
13parenrightbiggxparenleftbigg2
13parenrightbiggyparenleftbigg6
13parenrightbigg2−x−y
,
x,y=0,1,2;0≤x+y≤2, (19.30)
ormoreexplicitlyas
f(0,0)=parenleftbigg6
13parenrightbigg2
,f(1,0)=2·5
13·6
13=60
132,
f(2,0)=parenleftbigg5
13parenrightbigg2
,f(0,1)=22
13·6
13=24
132,
f(0,2)=parenleftbigg2
13parenrightbigg2
,f(1,1)=25
13·2
13=20
132.
The probability distribution is properly normalized according to Eq. (19.29). Its expecta-
tionvaluesaregivenby
/angbracketleftX/angbracketright=summationdisplay
0≤x+y≤2xf(x,y)=f(1,0)+f(1,1)+2f(2,0)
=60
132+20
132+2parenleftbigg5
13parenrightbigg2
=130
132=10
13=2a,
and
/angbracketleftY/angbracketright=summationdisplay
0≤x+y≤2yf(x,y)=f(0,1)+f(1,1)+2f(0,2)
=24
132+20
132+2parenleftbigg2
13parenrightbigg2
=52
132=4
13=2b,
asexpectedbecausewearedrawingacardtwotimes.Thevariancesare
σ2(X)=summationdisplay
0≤x+y≤2parenleftbigg
x−10
13parenrightbigg2
f(x,y)
=parenleftbigg10
13parenrightbigg2bracketleftbig
f(0,0)+f(0,1)+f(0,2)bracketrightbig
+parenleftbigg3
13parenrightbigg2bracketleftbig
f(1,0)+f(1,1)bracketrightbig
+parenleftbigg16
13parenrightbigg2
f(2,0)
=102·64+32·80+162·52
134=42·5·169
134=80
132,
19.2 Random Variables 1125
σ2(Y)=summationdisplay
0≤x+y≤2parenleftbigg
y−4
13parenrightbigg2
f(x,y)
=parenleftbigg4
13parenrightbigg2bracketleftbig
f(0,0)+f(1,0)+f(2,0)bracketrightbig
+parenleftbigg9
13parenrightbigg2bracketleftbig
f(0,1)+f(1,1)bracketrightbig
+parenleftbigg22
13parenrightbigg2
f(0,2)
=42·112+92·44+222·22
134=11·4·169
134=44
132.
It is reasonable that σ2(Y)<σ2(X),becauseYtakes only two values, 2 and 3, while X
variesoverthefivehonors. Thecovarianceis givenby
cov(X,Y)=summationdisplay
0≤x+y≤2parenleftbigg
x−10
13parenrightbiggparenleftbigg
y−4
13parenrightbigg
f(x,y)
=10·4
132·62
132−10·9
132·24
132−10·22
132·4
132−3·4
132·60
132
+3·9
132·20
132−16·4
132·52
132=−20·169
134=−20
132.
Thereforethecorrelationoftherandomvariables X,Yisgivenby
cov(X,Y)
σ(X)σ(Y)=−20
8√
5·11=−1
2radicalbigg
5
11=−0.3371,
which means that there is a small (negative) correlation between these random variables,
becauseifacardis anhonoritcannotb ea2ora3,andvicev er s a.
Finally,letus determinethemarginaldistribution:
F(X=x)=2summationdisplay
y=0f(x,y), (19.31)
orexplicitly
F(0)=f(0,0)+f(0,1)+f(0,2)=parenleftbigg6
13parenrightbigg2
+24
132+parenleftbigg2
13parenrightbigg2
=parenleftbigg8
13parenrightbigg2
,
F(1)=f(1,0)+f(1,1)=60
132+20
132=80
132,
F(2)=f(2,0)=parenleftbigg5
13parenrightbigg2
,
whichisproperlynormalizedbecause
F(0)+F(1)+F(2)=64+80+25
132=169
132=1.
1126 Chapter 19 Probability
Its meanvalueis givenby
/angbracketleftX/angbracketrightF=2summationdisplay
x=0xF(x)=F(1)+2F(2)=80+2·25
132=130
132=10
13=/angbracketleftX/angbracketright,
andits variance
σ2
F=2summationdisplay
x=0parenleftbigg
x−10
13parenrightbigg2
F(x)=parenleftbigg10
13parenrightbigg2
·parenleftbigg8
13parenrightbigg2
+parenleftbigg3
13parenrightbigg280
132+parenleftbigg16
13parenrightbigg2
·parenleftbigg5
13parenrightbigg2
=80·169
134=80
132=σ2(X).
Fromthedefinitionsitfollowsthattheseresults holdgenerally. /squaresolid
Finally we address the transformation of two random variables X,YintoU(X,Y),
V(X,Y). We treatthecontinuouscase,leavingthediscretecase,as anexercise.If
u=u(x,y), v =v(x,y);x=x(u,v), y =y(u,v) (19.32)
describe the transformation and its inverse, then the probability stays invariant and the
integral of the density transforms according to the rules of Jacobians of Chapter 2, so the
transformedprobabilitydensitybecomes
g(u,v)=fparenleftbig
x(u,v),y(u,v)parenrightbig
|J|, (19.33)
withtheJacobian
J=∂(x,y)
∂(u,v)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x
∂u∂x
∂v
∂y
∂u∂y
∂vvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (19.34)
Example 19.2.5 SUM,PRODUCT ,AND RATIO OF RANDOM VARIABLES
Let us consider three examples. (1) The sum Z=X+Y,where the transformation may
betakentobe
x=x, z=x+y, J=vextendsinglevextendsinglevextendsinglevextendsingle11
01vextendsinglevextendsinglevextendsinglevextendsingle,
using
∂x
∂x=1,∂(z−y)
∂z=1,∂y
∂x=0,∂(z−x)
∂z=1,
sotheprobabilityis givenby
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞f(x,z−x)dxdz. (19.35)
If therandomvariables X,Yare independentwithdensities f1,f2,then
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞f1(x)f2(z−x)dxdz. (19.36)
19.2 Random Variables 1127
(2) Theproduct Z=XY,takingX,Zas thenewvariables,leadstotheJacobian
J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle11
y
01
xvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=1
x,
using
∂x
∂x=1,∂(z
y)
∂z=1
y,∂y
∂x=0,∂(z
x)
∂z=1
x,
sotheprobabilityis givenby
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞fparenleftbigg
x,z
xparenrightbiggdx
|x|dz. (19.37)
If therandomvariables X,Yare independentwithdensities f1,f2,then
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞f1(x)f2parenleftbiggz
xparenrightbiggdx
|x|dz. (19.38)
(3) Theratio Z=X
Y, takingY,Zasthenewvariables,hastheJacobian
J=vextendsinglevextendsinglevextendsinglevextendsinglezy
10vextendsinglevextendsinglevextendsinglevextendsingle=−y,
using
∂(yz)
∂y=z,∂(yz)
∂z=y,∂y
∂y=1,∂y
∂z=0,
sotheprobabilityis givenby
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞f(yz,y)|y|dydz. (19.39)
If therandomvariables X,Yare independentwithdensities f1,f2,then
F(Z)=integraldisplayZ
−∞integraldisplay∞
−∞f1(yz)f2(y)|y|dydz. (19.40)
/squaresolid
Exercises
19.2.1 Show that adding a constant cto a random variable Xchanges the expectation value
/angbracketleftX/angbracketrightby that same constant but not the variance. Show also that multiplying a random
variablebyaconstantmultipliesboththemeanandvariancebythatconstant.Showthat
therandomvariable X−/angbracketleftX/angbracketrighthasmeanvaluezero.
19.2.2 If/angbracketleftX/angbracketright,/angbracketleftY/angbracketrightare the average values of two independent random variables X,Y,what is
theexpectationvalueoftheproduct X·Y?
1128 Chapter 19 Probability
19.2.3 A velocity vj=xj/tjis measured by recording the distances xjat the corresponding
timestj.Showthat¯x/¯tisagoodapproximationfortheaveragevelocity v,providedall
the errors|xj−¯x|≪|¯x|and|tj−¯t|≪|¯t|are small.
19.2.4 Define the random variable Yin Example 19.2.4 as the number of 4s, 5s, 6s, 7s, 8s, or
9s.Thendeterminethecorrelationof the XandYrandomvariables.
19.2.5 IfXandYare two independent random variables with different probability densities
andthefunction f(x,y)hasderivativesofanyorder,express /angbracketleftf(X,Y)/angbracketrightintermsof/angbracketleftX/angbracketright
and/angbracketleftY/angbracketright.Developsimilarlythecovarianceandcorrelation.
19.2.6 Letf(x,y)be the joint probability density of two random variables X,Y.Find the
varianceσ2(aX+bY),wherea,bare constants. What happens when X,Yare inde-
pendent?
19.2.7 The probability that a particle of an ideal gas travels a small distance dxbetween col-
lisions is∼e−x/fdx,wherefis the constant mean free path. Verify that fis the aver-
age distance between collisions, and determine the probability of a free path of length
l≥3f.
19.2.8 Determinethe probabilitydensityfor a particleinsimple harmonicmotionin theinter-
val−A≤x≤A.
Hint.The probability that the particle is between xandx+dxis proportional to the
timeittakes totravelacrosstheinterval.
19.3 B INOMIAL DISTRIBUTION
Example 19.3.1 REPEATED TOSSES OF DICE
Whatistheprobabilityofthree 6sinfourtosses,alltrialsbeingindependent?Gettingone
6 in a single toss of a fair die has probability a=1/6,and anything else has probability
b=5/6 witha+b=1.Let the random variable X=xbe the number of 6s. In four
tosses, 0≤x≤4.The probability distribution f(X)is given by the product of the two
possibilities, axandb4−x, times the number of combinations of four tosses over the two
cases with probabilities a,b.This number is the coefficient of the power axb4−xin the
binomialexpansionofallpossibilitiesinfour tosseswithprobability 1:
1=(a+b)4=4summationdisplay
x=04!
x!(4−x)!axb4−x
=a4+b4+4a3b+4ab3+6a2b2. (19.41)
Hencetheprobabilitydistributionofourdiscreterandomvariableisgivenby
f(X=x)=4!
x!(4−x)!axb4−x,0≤x≤4,
ormoreexplicitly
f(0)=b4,f(1)=4ab3,f(2)=6a2b2,f(3)=4a3b, f( 4)=a4.
19.3 Binomial Distribution 1129
The probability distribution is properly normalized according to Eq. (19.41). The proba-
bilityofthree 6sinfourtosses is
4a3b=45/6
63=5
4·34,
fairlysmall. /squaresolid
This case dealt with repeated independent trials, each with two possible outcomes of
constant probability pfor a hit and q=1−pfor a miss, and it is typical of many ap-
plications, such as defective products, hits or misses of a target, and decays of radioactive
atoms. The generalization to X=xsuccesses in ntrials is given by the binomial proba-
bilitydistribution
f(X=x)=n!
x!(n−x)!pxqn−x=parenleftbiggn
xparenrightbigg
pxqn−x, (19.42)
usingthebinomialcoefficients(seeChapter5).Thisdistributionisnormalizedtotheprob-
ability 1 ofallpossibilitiesin ntrials,as canbeseenfromthebinomialexpansion
1=(p+q)n=pn+npn−1q+···+npqn−1+qn. (19.43)
Figure19.5showstypicalhistograms.Therandomvariable Xtakesthevalues0 ,1,2,...,n
indiscretestepsandcanalsobeviewedasacompositionsummationtext
iXiofnindependentrandom
variables Xi,one for each trial, that have the value 0 for a miss and 1 for a hit. This
observationallowsus toemploythemoment-generatingfunctions
angbracketleftbig
etXiangbracketrightbig
=P(Xi=0)+etP(Xi=1)=q+pet(19.44)
and
angbracketleftbig
etXangbracketrightbig
=productdisplay
iangbracketleftbig
etXiangbracketrightbig
=parenleftbig
pet+qparenrightbign, (19.45)
FIGURE 19.5Binomialprobabilitydistributions
forn=20 andp=0.1,0.3,0.5.
1130 Chapter 19 Probability
from which the mean values and higher moments can be read off upon differentiating and
settingt=0.Using
∂/angbracketleftetX/angbracketright
∂t=npetparenleftbig
pet+qparenrightbign−1,
∂/angbracketleftetX/angbracketright
∂tvextendsinglevextendsinglevextendsinglevextendsingle
t=0=/angbracketleftX/angbracketright=summationdisplay
ixif(xi)=np,
∂2/angbracketleftetX/angbracketright
∂t2=npetparenleftbig
pet+qparenrightbign−1+n(n−1)p2e2tparenleftbig
pet+qparenrightbign−2,
angbracketleftbig
X2angbracketrightbig
=∂2/angbracketleftetX/angbracketright
∂t2vextendsinglevextendsinglevextendsinglevextendsingle
t=0=summationdisplay
ix2
if(xi)=np+n(n−1)p2,
weobtain,withEq. (19.17),
σ2(X)=angbracketleftbig
X2angbracketrightbig
−/angbracketleftX/angbracketright2=np+n(n−1)p2−n2p2
=np(1−p)=npq. (19.46)
Figure 19.5 illustrates these results with peaks at x=np=2,6,10,which widen with
increasing p.
Exercises
19.3.1 Show that the variable X=xnumber of heads in ncoin tosses is a random variable,
anddetermineitsprobabilitydistribution.Describethesamplespace.Whatareitsmean
value, the variance, and the standard deviation? Plot the probability function f(x)=
n!/(x!(n−x)!2n)forn=10,20,30 usinggraphicalsoftware.
19.3.2 Plotthebinomialprobabilityfunctionfortheprobabilities p=1/6,q=5/6andn=6
throwsofa die.
19.3.3 A hardware company knows that the probability of mass-producing nails includes a
small probability p=0.03 of defective nails (without a sharp tip usually). What is the
probabilityoffindingmorethantwodefectivenailsinitscommercialboxof100nails?
19.3.4 Four cards are drawn from a shuffled bridge deck. What is the probability that they are
all red? that they are all hearts? that they are honors? Compare the probabilities when
thecardsareputbackatrandomplaces,or not.
19.3.5 Show that for the binomial distribution of Eq. (19.42) the most probable value of xis
np.
19.4 P OISSON DISTRIBUTION
The Poisson distribution typically occurs in situations involving an event repeated at a
constant rate of probability, thereby depleting the population. The decay of a radioactive
19.4 Poisson Distribution 1131
sample is a case in point because, once a particle decays, it does not decay again. If the
observation time dtis small enough so that the emission of two or more particles is neg-
ligible, then the probability that one particle (He4inαdecay or an electron in βdecay) is
emitted is µdtwith constant µandµdt≪1.We can set up a recursion relation for the
probability Pn(t)of observing ncounts during a time interval t.Forn>0 the probability
Pn(t+dt)iscomposedoftwomutuallyexclusiveeventsthat(i) nparticlesareemittedin
thetimet,noneindt, and(ii)n−1 particlesareemittedintime t,oneindt.Therefore
Pn(t+dt)=Pn(t)P0(dt)+Pn−1(t)P1(dt).
Herewesubstitutetheprobabilityofobservingoneparticle, P1(dt)=µdt,andnoparticle,
P0(dt)=1−P1(dt),intimedt.Thisyields
Pn(t+dt)=Pn(t)(1−µdt)+Pn−1(t)µdt.
So,afterrearranginganddividingby dt,weget
dPn(t)
dt=Pn(t+dt)−Pn(t)
dt=µPn−1(t)−µPn(t). (19.47)
Forn=0thisdifferentialrecursionrelationsimplifies,becausethereisnoparticleintimes
tanddtgiving
dP0(t)
dt=−µP0(t). (19.48)
The ODE says that particles have a constant decay probability and decay removes them
fromthedistribution.ThisODEintegratesto P0(t)=e−µtiftheprobabilitythatnoparticle
is emitted during a zero time interval P0(0)=1 is used. Here P0(0)=1 means no decay
takesplaceat t≤0.
NowwegobacktoEq.(19.47) for n=1,
˙P1=µparenleftbig
e−µt−P1parenrightbig
,P 1(0)=0, (19.49)
and solve the homogeneous equation, which is the same for P1as Eq. (19.48). This yields
P1(t)=µ1e−µt.Then we solve the inhomogeneous ODE (Eq. (19.49)) by varying the
constantµ1tofind˙µ1=µ,soP1(t)=µte−µt.Thegeneralsolutionis
Pn(t)=(µt)n
n!e−µt, (19.50)
as may be confirmed by substitution into Eq. (19.47) and verifying the initial conditions,
Pn(0)=0,n>0.This isanexampleof thePoissondistribution.
ThePoissondistributionisdefinedwiththeprobabilities
p(n)=µn
n!e−µ,X=n=0,1,2,... (19.51)
and is exhibited in Fig. 19.6. The random variable Xis discrete. The probabilities are
properlynormalizedbecause e−µsummationtext∞
n=0µn
n!=1.Themeanvalueandvariance,
/angbracketleftX/angbracketright=e−µ∞summationdisplay
n=1nµn
n!=µe−µ∞summationdisplay
n=0µn
n!=µ,
σ2=angbracketleftbig
X2angbracketrightbig
−/angbracketleftX/angbracketright2=µ(µ+1)−µ2=µ, (19.52)
1132 Chapter 19 Probability
FIGURE 19.6Poissondistribution
comparedwithbinomialdistribution.
followfrom thecharacteristicfunction
angbracketleftbig
eitXangbracketrightbig
=∞summationdisplay
n=0eitn−µµn
n!=e−µ∞summationdisplay
n=0(µeit)n
n!=eµ(eit−1)
bydifferentiationandsetting t=0,usingEq.(19.17).
APoissondistributionbecomesagoodapproximationofthebinomialdistributionfora
largenumber noftrials andsmallprobability p∼µ/n,µaconstant.
THEOREM:In the limit n→∞andp→0so that the mean value np→µstays finite,
thebinomialdistributionbecomesaPoissondistribution.
To prove this theorem, we apply Stirling’s formula (Chapter 8) n!∼√
2πn(n/e)nfor
largento the factorials in Eq. (19.42), keeping xfinite while n→∞.This yields for
n→∞:
n!
(n−x)!∼parenleftbiggn
eparenrightbiggnparenleftbigge
n−xparenrightbiggn−x
∼parenleftbiggn
eparenrightbiggxparenleftbiggn
n−xparenrightbiggn−x
∼parenleftbiggn
eparenrightbiggxparenleftbigg
1+x
n−xparenrightbiggn−x
∼parenleftbiggn
eparenrightbiggx
ex∼nx,
andforn→∞,p→0,withnp→µ:
(1−p)n−x∼parenleftbigg
1−pn
nparenrightbiggn
∼parenleftbigg
1−µ
nparenrightbiggn
∼e−µ.
19.4 Poisson Distribution 1133
Table 19.1
i→0 12345678 9 1 0
ni→57 203 383 525 532 408 273 139 45 27 16
Finally,pxnx→µx,so altogether
n!
x!(n−x)!px(1−p)n−x→µx
x!e−µ,n→∞, (19.53)
which is a Poisson distribution for the random variable X=xwith 0≤x<∞.This limit
theoremisaparticularexampleofthe lawsoflargenumbers .
Exercises
19.4.1 Radioactive decays are governed by the Poisson distribution. In a Rutherford–Geiger
experiment the number niof emitted αparticles is counted in n=2608 time intervals
of7.5secondseach.InTable19.1 niisthenumberoftimeintervalsinwhich iparticles
wereemitted.Determinetheaveragenumber λofemittedparticles,andcomparethe ni
ofTable19.1with npicomputedfromthePoissondistributionwithmeanvalue λ.
19.4.2 DerivethestandarddeviationofaPoissondistributionof meanvalue µ.
19.4.3 Thenumberof αdecayparticlesofaradiumsampleiscountedperminutefor40hours.
The total number is 5000 .How many 1-minute intervals are there expected to be with
(a) 2,(b) 5 αparticles?
19.4.4 For a radioactive sample, 10 decays are counted on average in 100 seconds. Use the
Poissondistributiontoestimatetheprobabilityofcounting 3 decaysin 10 seconds.
19.4.5238Uhasahalf-lifeof4 .51×109years.Itsdecayseriesendswiththestableleadisotope
206Pb. The ratio of the number of206Pb to238U atoms in a rock sample is measured as
0.0058. Estimate the age of the rock assuming that all the lead in the rock is from the
initialdecayofthe238U,whichdeterminestherateoftheentiredecayprocess,because
thesubsequentstepstakeplacefar morerapidly.
Hint.The decay constant λin the decay law N(t)=Ne−λtis related to the half-life T
byT=ln2/λ.
ANS. 3.8×107years.
19.4.6 Theprobabilityofhittingatargetinoneshotisknowntobe 20% .Iffiveshotsarefired
independently,whatistheprobabilityofstrikingthetargetatleastonce?
19.4.7 Apieceofuraniumisknowntocontaintheisotopes235
92Uand238
92Uaswellasfrom0 .80
gof206
82Pbpergramofuranium.Estimatetheageofthepiece(andthusEarth)inyears.
Hint.Assume the lead comes only from the238
92U. Use the decay constant from Exer-
cise19.4.5.
1134 Chapter 19 Probability
19.5 G AUSS ’NORMAL DISTRIBUTION
Thebell-shapedGaussdistributionisdefinedbytheprobabilitydensity
f(x)=1
σ√
2πexpparenleftbigg
−[x−µ]2
2σ2parenrightbigg
,−∞<x<∞, (19.54)
with mean value µand variance σ2.It is by far the most importantcontinuous probability
distributionandis displayedinFig.19.7.
It isproperlynormalizedbecause,substituting y=x−µ
σ√
2,weobtain
1
σ√
2πintegraldisplay∞
−∞e−(x−µ)2
2σ2dx=1√πintegraldisplay∞
−∞e−y2dy=2√πintegraldisplay∞
0e−y2dy=1.
Similarly,substituting y=x−µ,wesee that
/angbracketleftX/angbracketright−µ=integraldisplay∞
−∞x−µ
σ√
2πe−(x−µ)2
2σ2dx=integraldisplay∞
−∞y
σ√
2πe−y2
2σ2dy=0,
the integrand being odd in y,so the integral over y>0 cancels that over y<0.Similarly
wecheckthatthestandarddeviationis σ.
Fromthenormaldistribution(bythesubstitution y=x−/angbracketleftX/angbracketright
σ)
PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle>kσparenrightbig
=Pparenleftbigg|X−/angbracketleftX/angbracketright|
σ>kparenrightbigg
=Pparenleftbig
|Y|>kparenrightbig
=radicalbigg
2
πintegraldisplay∞
ke−y2/2dy=radicalbigg
4
πintegraldisplay∞
k/√
2e−z2dz=erfck√
2,
FIGURE 19.7NormalGaussdistributionformeanvaluezeroandvarious
standarddeviations h=1/σ√
2.
19.5 Gauss’ Normal Distribution 1135
we can evaluate the integral for k=1,2,3 and thus extract the following numerical rela-
tionsforanormallydistributedrandomvariable:
PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥σparenrightbig
∼0.3173,PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥2σparenrightbig
∼0.0455,
PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥3σparenrightbig
∼0.0027, (19.55)
of which the last one is interesting to compare with Chebychev’s inequality (see
Eq. (19.21).) giving ≤1/9 for anarbitrary probability distribution instead of ∼0.0027
forthe 3σ-ruleofthe normaldistribution.
ADDITION THEOREM :If the random variables X,Yhave the same normal distributions,
that is, the same mean value and variance, then Z=X+Yhas normal distribution with
twicethemeanvalueandtwicethevarianceof XandY.
Toprovethistheorem,wetaketheGaussdensityas
f(x)=1√
2πe−x2/2,with1√
2πintegraldisplay∞
−∞e−x2/2dx=1,
withoutloss ofgenerality.Thentheprobabilitydensityof (X,Y)is theproduct
f(x,y)=1√
2πe−x2/21√
2πe−y2/2=1
2πe−(x2+y2)/2.
Also,Eq. (17.36)givesthedensityfor Z=X+Yas
g(z)=integraldisplay∞
−∞1√
2πe−x2/21√
2πe−(x−z)2/2dx.
Completingthesquareintheexponent,
2x2−2xz+z2=parenleftbigg
x√
2−z√
2parenrightbigg2
+z2
2,
weobtain
g(z)=1
2πe−z2/4integraldisplay∞
−∞expparenleftbigg
−1
2parenleftbigg
x√
2−z√
2parenrightbigg2parenrightbigg
dx.
Usingthesubstitution u=x−z
2,wefindthattheintegraltransformsinto
integraldisplay∞
−∞expparenleftbigg
−1
2parenleftbigg
x√
2−z√
2parenrightbigg2parenrightbigg
dx=integraldisplay∞
−∞e−u2du=√π,
sothedensityfor Z=X+Yis
g(z)=1
2√πe−z2/4, (19.56)
whichmeansithas meanvaluezeroandvariance 2 ,twicethatof XandY.
In a special limit the discrete Poisson probability distribution is closely related to the
continuous Gauss distribution. This limit theorem is another example of the laws of large
numbers ,whichareoftendominatedbythebell-shapednormaldistribution.
1136 Chapter 19 Probability
THEOREM:For large nand mean value µ, the Poisson distribution approaches a Gauss
distribution.
To prove this theorem for n→∞,we approximate the factorial in the Poisson’s proba-
bilityp(n)ofEq. (19.51)byStirling’sasymptoticformula(seeChapter8),
n!∼√
2nπparenleftbiggn
eparenrightbiggn
,n→∞,
and choose the deviation v=n−µfrom the mean value as the new variable. We let the
mean value µ→∞and treat v/µas small but v2/µas finite. Substituting n=µ+vand
expandingthelogarithminaMacLaurinseries, keepingtwoterms, weobtain
lnp(n)=−µ+nlnµ−nlnn+n−ln√
2nπ
=(µ+v)lnµ−(µ+v)ln(µ+v)+v−lnradicalbig
2π(µ+v)
=(µ+v)lnparenleftbigg
1−v
µ+vparenrightbigg
+v−lnradicalbig
2πµ
=(µ+v)parenleftbigg
−v
µ+v−v2
2(µ+v)2parenrightbigg
+v−lnradicalbig
2πµ
∼−v2
2µ−lnradicalbig
2πµ,
replacing µ+v→µbecause|v|≪µ.Exponentiating this result we find that for large n
andµ
p(n)→1√2πµe−v2/2µ, (19.57)
whichis a Gauss distributionof the continuousvariable vwithmean value 0 and standard
deviation σ=√µ.
In a special limit the discrete binomial probability distribution is also closely related to
the continuous Gauss distribution. This limit theorem is another example of the laws of
largenumbers .
THEOREM:In the limit n→∞, so that the mean value np→∞,the binomial distribu-
tion becomes Gauss’ normal distribution. Recall from Section 19.4that, when np→µ<
∞,thebinomialdistributionbecomesaPoissondistribution.
Instead of the large number xof successes in ntrials, we use the deviation v=x−pn
fromthe(large)meanvalue pnasournewcontinuousrandomvariable,underthecondition
that|v|≪pnbutv2/nis finite as n→∞.Thus, we replace xbyv+pnandn−xby
qn−vinthefactorialsofEq.(19.42), f(x)→W(v)asn→∞,andthenapplyStirling’s
formula.Thisyields
W(v)=pxqn−xnn+1/2e−n+x+(n−x)
√
2π(v+pn)x+1/2(qn−v)n−x+1/2.
19.5 Gauss’ Normal Distribution 1137
Herewefactor outthedominantpowersof nandcancelpowersof pandqtofind
W(v)=1√2πpqnparenleftbigg
1+v
pnparenrightbigg−(v+pn+1/2)parenleftbigg
1−v
qnparenrightbigg−(qn−v+1/2)
.
Intermsof thelogarithmwehave
lnW(v)=ln1√2πpqn−(v+pn+1/2)lnparenleftbigg
1+v
pnparenrightbigg
−(qn−v+1/2)lnparenleftbigg
1−v
qnparenrightbigg
=ln1√2πpqn−(v+pn+1/2)parenleftbiggv
pn−v2
2p2n2+···parenrightbigg
−(qn−v+1/2)parenleftbigg
−v
qn−v2
2q2n2+···parenrightbigg
=ln1√2πpqn−bracketleftbiggv
nparenleftbigg1
2p−1
2qparenrightbigg
+v2
nparenleftbigg1
2p+1
2qparenrightbigg
+···bracketrightbigg
,
where
v
n→0,v2
n
isfiniteand
parenleftbigg1
2p+1
2qparenrightbigg
=p+q
2pq=1
2pq.
Neglectinghigherordersin v/n,suchasv2/p2n2andv2/q2n2,wefindthelarge nlimit
W(v)=1√2πpqne−v2/2pqn, (19.58)
whichisaGaussiandistributioninthedeviations x−pn,withmeanvalue 0 andstandard
deviation σ=√npq.The large mean value pn(and the discarded terms) restricts the
validityofthetheoremtothecentralpartoftheGaussianbellshape,excludingthetails.
Exercises
19.5.1 What is the probability for a normally distributed random variable to differ by more
than 4σfrom its mean value? Compare your result with the corresponding one from
Chebychev’sinequality.Explainthedifferenceinyourownwords.
19.5.2 LetX1,X2,...,X nbeindependentnormalrandomvariableswiththesamemean ¯xand
varianceσ2.Showthatsummationtext
iXi/n−¯x√nσisnormalwithmeanzeroandvariance 1 .
1138 Chapter 19 Probability
19.5.3 An instructor grades a final exam of a large undergraduate class, obtaining the mean
value of points Mand the variance σ2. Assuming a normal distribution for the number
Mof points, he defines a grade F when M<m−3σ/2,D whenm−3σ/2<M<
m−σ/2,C whenm−σ/2<M<m+σ/2,B whenm+σ/2<M<m+3σ/2,
A whenM>m+3σ/2.What is the percentage of As, Fs; Bs, Ds; Cs? Redesign the
cutoffs so that there are equal percentages of As and Fs (5%), 25% Bs and Ds, and
40% Cs.
19.5.4 If the random variable Xis normal with mean value 29 and standard deviation 3 ,what
arethedistributionsof 2 X−1 and 3X+2?
19.5.5 Foranormaldistributionofmeanvalue mandvariance σ2,findthedistance rsuchthat
halftheareaunderthebellshapeis between m−randm+r.
19.6 S TATISTICS
In statistics, probability theory is applied to the evaluation of data from random experi-
ments or to samples to test some hypothesis because the data have random fluctuations
due to lack of complete control over the experimental conditions. Typically one attempts
to estimate the mean value and variance of the distributions, from which the samples de-
rive, and to generalize properties valid for a sample to the rest of the events at a prescised
confidence level. Any assumption about an unknown probability distribution is called a
statistical hypothesis . The concepts of tests and confidence intervals are among the most
importantdevelopmentsof statistics.
Error Propagation
When we measure a quantity xrepeatedly, obtaining the values xjat random, or select a
samplefor testing,wedeterminethemeanvalue(see Eq.(19.10)) andthevariance,
¯x=1
nnsummationdisplay
j=1xj,σ2=1
nnsummationdisplay
j=1(xj−¯x)2,
as a measure for the error, or spread from the mean value ¯x.We can write xj=¯x+ej,
wheretheerror ejisthedeviationfromthemeanvalue,andweknowthatsummationtext
jej=0.(See
thediscussionafterEq. (19.15).)
Now suppose we want to determine a known function f(x)from these measurements;
that is, we have a set fj=f(xj)from the measurements of x.Substituting xj=¯x+ej
andformingthemeanvaluefrom
¯f=1
nsummationdisplay
jf(xj)=1
nsummationdisplay
jf(¯x+ej)
=f(¯x)+1
nf′(¯x)summationdisplay
jej+1
2nf′′(¯x)summationdisplay
je2
j+···
=f(¯x)+1
2σ2f′′(¯x)+···, (19.59)
19.6 Statistics 1139
we obtain the average value ¯fasf(¯x)in lowest order, as expected. But in second order
thereisacorrectiongivenbyhalfthevariancewithascalefactor f′′(¯x).Itisinterestingto
compare this correction of the mean value with the average spread of individual fjfrom
the mean value ¯f,the variance of f.To lowest order, this is given by the average of the
sumofsquaresof thedeviations,inwhichweapproximate fj≈¯f+f′(¯x)ej,yielding
σ2(f)≡1
nsummationdisplay
j(fj−¯f)2=parenleftbig
f′(¯x)parenrightbig21
nsummationdisplay
je2
j=parenleftbig
f′(¯x)parenrightbig2σ2.(19.60)
Insummarywemayformulatesomewhatsymbolically
f(¯x±σ)=f(¯x)±f′(¯x)σ
asthesimplestform oferror propagationbyafunctionofonemeasuredvariable.
Forafunction f(xj,yk)oftwomeasuredquantities xj=¯x+uj,yk=¯y+vk,weobtain
similarly
¯f=1
rsrsummationdisplay
j=1ssummationdisplay
k=1fjk=1
rsrsummationdisplay
j=1ssummationdisplay
k=1f(¯x+uj,¯y+vk)
=f(¯x,¯y)+1
rfxsummationdisplay
juj+1
sfxsummationdisplay
kvk+···,
wheresummationtext
juj=0=summationtext
kvk,so again¯f=f(¯x,¯y)inlowestorder.Here
fx=∂f
∂x(¯x,¯y), f y=∂f
∂y(¯x,¯y) (19.61)
denote partial derivatives. The sum of squares of the deviations from the mean value is
givenby
rsummationdisplay
j=1ssummationdisplay
k+1(fjk−¯f)2=summationdisplay
j,k(ujfx+vkfy)2=sf2
xsummationdisplay
ju2
j+rf2
ysummationdisplay
kv2
k,
becausesummationtext
j,kujvk=summationtext
jujsummationtext
kvk=0.Thereforethevarianceis
σ2(f)=1
rssummationdisplay
j,k(fjk−¯f)2=f2
xσ2
x+f2
yσ2
y, (19.62)
withfx,fyfrom Eq.(19.61);and
σ2
x=1
rsummationdisplay
ju2
j,σ2
y=1
ssummationdisplay
kv2
k
are the variances of the xandydata points. Symbolically the error propagation for a
functionoftwomeasuredvariablesmaybesummarizedas
f(¯x±σx,¯y±σy)=f(¯x,¯y)±radicalBig
f2xσ2x+f2yσ2y.
As an application and generalization of the last result, we now calculate the error of
the mean value ¯x=1
nsummationtextn
j=1xjof a sample of nindividual measurements xj,each with
1140 Chapter 19 Probability
spreadσ.Inthiscasethepartialderivativesaregivenby fx=1
n=fy=···andσx=σ=
σy=···.Thus,ourlasterrorpropagationruletellsusthaterrorsofasumofvariablesadd
quadratically,so theuncertaintyofthearithmeticmeanisgivenby
¯σ=1
nradicalbig
nσ2=σ√n, (19.63)
decreasingwiththenumberof measurements n.
As the number nof measurements increases, we expect the arithmetic mean ¯xto con-
vergetosometruevalue x.Let¯xdifferfrom xbyαandvj=xj−xbethetruedeviations;
then
summationdisplay
j(xj−¯x)2=summationdisplay
je2
j=summationdisplay
jv2
j+nα2.
Taking into account the error of the arithmetic mean, we determine the spread of the indi-
vidualpointsabouttheunknowntruemeanvaluetobe
σ2=1
nsummationdisplay
je2
j=1
nsummationdisplay
jv2
j+α2.
AccordingtoourearlierdiscussionleadingtoEq.(19.63), α2=1
nσ2.Asaresult
σ2=1
nsummationdisplay
jv2
j+σ2
n,
fromwhichthe standarddeviationofasampleinstatistics follows:
σ=radicalBiggsummationtext
jv2
j
n−1=radicalBiggsummationtext
j(xj−x)2
n−1, (19.64)
withn−1 being the number of control measurements of the sample. This modified mean
errorincludestheexpectederror inthearithmeticmean.
Because the spread is not well defined when there is no comparison measurement, that
is, whenn=1,the variance is sometimes defined by Eq. (19.64), in which we replace the
numbernofmeasurementsbythenumber n−1 ofcontrolmeasurementsinstatistics.
Fitting Curves to Data
Supposewehaveasampleofmeasurements yj(forexample,aparticlemovingfreely,that
is, no force) taken at known times tj(which are taken to be practically free of errors; that
is, the time tis an ordinary independent variable) that we expect to be linearly related as
y=at,ourhypothesis.Wewanttofitthislinetothedata.
Firstweminimizethesumofdeviationssummationtext
j(atj−yj)2todeterminetheslopeparame-
tera,also called the regression coefficient , using the method of least squares. Differenti-
atingwithrespectto aweobtain
2summationdisplay
j(atj−yj)tj=0,
19.6 Statistics 1141
FIGURE 19.8Straightlinefitto
datapoints (tj,yj)withtj
known,yjmeasured.
fromwhich
a=summationtext
jtjyjsummationtext
jt2
j(19.65)
follows.Notethatthenumeratorisbuiltlikeasamplecovariance,thescalarproductofthe
variables t,yof the sample. As shown in Fig. 19.8, the measured values yjdo not lie on
thelineasarule.Theyhavethespread(orrootmeansquaredeviationfromthefittedline)
σ=radicalBiggsummationtext
j(yj−atj)2
n−1.
Alternatively,letthe yjvaluesbeknown(withouterror) while tjare measurements.As
suggestedbyFig.19.9,inthiscaseweneedtointerchangetheroleof tandyandtofitthe
linet=bytothedatapoints.Weminimizesummationtext
j(byj−tj)2,setthederivativewithrespect
FIGURE 19.9Straightlinefitto
datapoints (tj,yj)withyj
known,tjmeasured.
1142 Chapter 19 Probability
FIGURE 19.10(a) Straightlinefittodatapoints (tj,yj).(b) Geometryof
deviations uj,vj,dj.
tobequaltozero,andfindsimilarlytheslopeparameter
b=summationtext
jtjyjsummationtext
jy2
j. (19.66)
In case both tjandyjhave errors (we take tandyto have the same units), we have to
minimizethesumofsquaresofthedeviationsofbothvariablesandfittoaparameterization
tsinα−ycosα=0,wheretandyoccuronanequalfooting.AsdisplayedinFig.(19.10a)
thismeansgeometricallythatthelinehastobedrawnsothatthesumofthesquaresofthe
distances djofthepoints (tj,yj)fromthelinebecomesaminimum.(SeeFig.19.10band
Chapter 1.) Here dj=tjsinα−yjcosα,sosummationtext
jd2
j=minimum must be solved for the
angleα.Settingthederivativewithrespecttotheangleequaltozero,
summationdisplay
j(tjsinα−yjcosα)(tjcosα+yjsinα)=0,
yields
sinαcosαsummationdisplay
jparenleftbig
t2
j−y2
jparenrightbig
−parenleftbig
cos2α−sin2αparenrightbigsummationdisplay
jtjyj=0.
Thereforetheangleofthestraight-linefitis givenby
tan2α=2summationtext
jtjyjsummationtext
j(t2
j−y2
j). (19.67)
This least-squares fitting applies when the measurement errors are unknown. It allows as-
signing at least some kind of error bar to the measured points. Recall that we did not use
errors for the points. Our parameter a(orα) is most likely to reproduce the data under
these circumstances. More precisely, the least-squares method is a maximum-likelihood
estimate of the fitted parameters when it is reasonable to assume that the errors are in-
dependent and normally distributed with the same deviation for all points . This fairly
19.6 Statistics 1143
strong assumption can be relaxed in “weighted” least-squares fits called chi square fits .4
(SeealsoExample19.6.1.)
Theχ2Distribution
This distribution is typically applied to fits of a curve y(t,a,...) with parameters a,...
to datatjusing the method of least squares involving the weighted sum of squares of
deviations;thatis,
χ2=Nsummationdisplay
j=1parenleftbiggyj−y(tj,a,...)
/Delta1yjparenrightbigg2
isminimized,where Nisthenumberofpointsand risthenumberofadjustedparameters
a,....This quadratic merit function gives more weight to points with small measurement
uncertainties /Delta1yj.
We represent each point by a normally distributed random variable Xwith zero mean
value and variance σ2=1,the latter in view of the weights in the χ2function. In a first
step, we determine the probability density for the random variable Y=X2of a single
point that takes only positive values. Assuming a zero mean value is no loss of generality
because,if/angbracketleftX/angbracketright=m/negationslash=0,wewouldconsidertheshiftedvariable Y=X−m,whosemean
valueis zero.Weshowthatif XhasaGaussnormaldensity
f(x)=1
σ√
2πe−x2/2σ2,−∞<x<∞,
thentheprobabilityof therandomvariable Yis zeroif y≤0,and
P(Y<y)=P(X2<y)=P(−√y<X<√y)ify>0.
From the continuous normal distribution P(y)=integraltexty
−∞f(x)dx, we obtain the probability
densityg(y)bydifferentiation:
g(y)=d
dybracketleftbig
P(√y)−P(−√y)bracketrightbig
=1
2√yparenleftbig
f(√y)+f(−√y)parenrightbig
=1
σ√2πye−y/2σ2,y>0. (19.68)
This density,∼e−y/2σ2/√y,corresponds to the integrand of the Euler integral of the
gammafunction.Suchaprobabilitydistribution
g(y)=yp−1
Ŵ(p)(2σ2)pe−y/2σ2
is called a gamma distribution with parameters pandσ.Its characteristic function for
ourcase, p=1/2,is proportionaltotheFouriertransform
angbracketleftbig
eitYangbracketrightbig
=1
σ√
2πintegraldisplay∞
0e−y(1/2σ2−it)dy√y=1
σ√
2π(1
2σ2−it)1/2integraldisplay∞
0e−xdx√x
=parenleftbig
1−2itσ2parenrightbig−1/2.
4For more details,seeChapter 14 of Press etal. in theAdditional Readings of Chapter9.
1144 Chapter 19 Probability
Sincethe χ2samplefunctioncontainsasumofsquares,weneedthefollowingtheorem.
ADDITION THEOREM :for the gamma distributions :If the independent random variables
Y1andY2have a gamma distribution with p=1/2,and the same σthenY1+Y2has a
gammadistributionwith p=1.
SinceY1andY2are independent, the product of their densities (Eq. (19.36)) generates
thecharacteristicfunction
angbracketleftbig
eit(Y1+Y2)angbracketrightbig
=angbracketleftbig
eitY1eitY2angbracketrightbig
=angbracketleftbig
eitY1angbracketrightbigangbracketleftbig
eitY2angbracketrightbig
=parenleftbig
1−2itσ2parenrightbig−1. (19.69)
Now we come to the second step. We assess the quality of the fit by the random variable
Y=summationtextn
j=1X2
j,wheren=N−ris the number of degrees of freedom for Ndata points
andrfitted parameters. The independent random variables Xjare taken to be normally
distributed with the (sample) variance σ2.(In our case r=1 andσ=1.)T h eχ2analysis
does not really test the assumptions of normality and independence, but if these are not
approximately valid, there will be many outlying points in the fit. The addition theorem
givestheprobabilitydensity(Fig.19.11)for Y,
gn(y)=yn
2−1
2n/2σnŴ(n
2)e−y/2σ2,y>0,
andgn(y)=0ify<0,whichisthe χ2distributioncorrespondingto ndegreesoffreedom.
Its characteristicfunctionis
angbracketleftbig
eitYangbracketrightbig
=parenleftbig
1−2itσ2parenrightbig−n/2.
Differentiatingandsetting t=0 weobtainitsmeanvalueandvariance
/angbracketleftY/angbracketright=nσ2,σ2(Y)=2nσ4. (19.70)
FIGURE 19.11χ2probabilitydensity
gn(y).
19.6 Statistics 1145
Table 19.2 χ2Distribution
nv=0.8 v=0.7 v=0.5 v=0.3 v=0.2 v=0.1
1 0.064 0.148 0.455 1.074 1.642 2 .706
2 0.446 0.713 1.386 2.408 3.219 4 .605
3 1.005 1.424 2.366 3.665 4.642 6 .251
4 1.649 2.195 3.357 4.878 5.989 7 .779
5 2.343 3.000 4.351 6.064 7.289 9 .236
6 3.070 3.828 5.348 7.231 8.558 10 .645
Entries are χvfor the probabilities v=P(χ2≥χ2v)=1
2n/2Ŵ(n/2)integraltext∞
χ2ve−y/2y(n/2)−1dyforσ=1.
Tablesgivevaluesfor the χ2probabilityfor ndegreesoffreedom,
Pparenleftbig
χ2≥y0parenrightbig
=1
2n/2σnŴ(n
2)integraldisplay∞
y0yn/2−1e−y/2σ2dy
forσ=1andy0>0.TouseTable19.2for σ/negationslash=1,rescale y0=v0σ2sothatP(χ2≥v0σ2)
corresponds to P(χ2≥v0)of Table 19.2. The following example will illustrate the whole
process.
Example 19.6.1
Let us apply the χ2function to the fit in Fig. 19.8. The measured points (tj,yj±/Delta1yj)
witherrors /Delta1yjare
(1,0.8±0.1), (2,1.5±0.05), (3,3±0.2).
Forcomparison,themaximum-likelihoodfit, Eq.(19.65), gives
a=1·0.8+2·1.5+3·3
1+4+9=12.8
14=0.914.
Minimizinginstead,
χ2=summationdisplay
jparenleftbiggyj−atj
/Delta1yjparenrightbigg2
,
gives
0=∂χ2
∂a=−2summationdisplay
jtj(yj−atj)
(/Delta1yj)2,
or
a=summationtext
jtjyj
(/Delta1yj)2
summationtext
jt2
j
(/Delta1yj)2.
1146 Chapter 19 Probability
Inourcase
a=1·0.8
0.12+2·1.5
0.052+3·3
0.22
12
0.12+22
0.052+32
0.22=1505
1925=0.782
is dominated by the middle point with the smallest error, /Delta1y2=0.05.The error propaga-
tionformula(Eq. (19.62)) givesus thevariance σ2
aof theestimateof a,
σ2
a=summationdisplay
j(/Delta1yj)2parenleftbigg∂a
∂yjparenrightbigg2
=summationdisplay
jt2
j
(/Delta1yj)2
parenleftbigsummationtext
kt2
k
(/Delta1yk)2parenrightbig2=1
summationtext
jt2
j
(/Delta1yj)2
using
∂a
∂yj=tj
(/Delta1yj)2
summationtext
kt2
k
(/Delta1yk)2.
Forourcase, σa=1/√
1925=0.023;thatis, ourslopeparameteris a=0.782±0.023.
To estimate the quality of this fit of a,we compute the χ2probability that the two
independent(control)pointsmissthefitbytwostandarddeviations;thatis,onaverageeach
point misses by one standard deviation. We apply the χ2distribution to the fit involving
N=3datapointsand r=1parameter,thatis,for n=3−1=2degreesoffreedom.From
Eq.(19.70)the χ2distributionhasameanvalue2andavariance4.Aruleofthumbisthat
χ2≈nfor a reasonably good fit. Then P(χ2≥2)∼0.496 is read off Table 19.2, where
weinterpolatebetween P(χ2≥1.3862)=0.50 andP(χ2≥2.4082)=0.30 asfollows:
Pparenleftbig
χ2≥2parenrightbig
=Pparenleftbig
χ2≥1.3862parenrightbig
−2−1.3862
2.4082−1.3862bracketleftbig
Pparenleftbig
χ2≥1.3862parenrightbig
−Pparenleftbig
χ2≥2.4082parenrightbigbracketrightbig
=0.5−0.02·0.2=0.496.
Thus the χ2probability that, on average, each point misses by one standard deviation is
nearly 50% andfairlylarge. /squaresolid
Our next goal is to compute a confidence interval for the slope parameter of our fit.
Aconfidenceintervalforanaprioriunknownparameterofsomedistribution(forexample,
adetermined by our fit) is an interval that contains anot with certainty but with a high
probability p,theconfidencelevel,whichwecanchoose.Suchanintervaliscomputedfor
agivensample.SuchananalysisinvolvestheStudent tdistribution.
The Student tDistribution
Becausewealwayscomputethearithmeticmeanofmeasuredpoints,wenowconsiderthe
samplefunction
¯X=1
nnsummationdisplay
j=1Xj,
19.6 Statistics 1147
where the random variables Xjare assumed independent with a normal distribution of
the same mean value mand variance σ2. The addition theorem for the Gauss distrib-
ution tells us that X1+···+Xnhas the mean value nmand variance nσ2.Therefore
(X1+···+Xn)/nisnormalwithmeanvalue mandvariance nσ2/n2=σ2/n.Theprob-
abilitydensityof thevariable ¯X−mis theGaussdistribution
¯f(¯x−m)=√n
σ√
2πexpparenleftbigg
−n(¯x−m)2
2σ2parenrightbigg
. (19.71)
The key problem solved by the Student tdistribution is to provide estimates for the mean
valuem, whenσis not known , in terms of a sample function whose distribution is inde-
pendentofσ.Tothis end,wedefinearescaledsamplefunction(traditionallycalled) t:
t=¯X−m
S√
n−1,S2=1
nnsummationdisplay
j=1(Xj−¯X)2. (19.72)
It can be shown that tandSare independent random variables. Following the arguments
leading to the χ2distribution, the density of the denominator variable Sis given by the
gammadistribution
d(s)=n(n−1)/2sn−2e−ns2/2σ2
2n−3
2Ŵ(n−1
2)σn−1. (19.73)
The probability for the ratio Z=X/Yof two independent random variables X,Ywith
normaldensityfor ¯fanddasgivenbyEqs. (19.71)and(19.73)is (Eq. (19.40))
R(z)=integraldisplayz
−∞integraldisplay∞
−∞f(yz)d(y)|y|dydz, (19.74)
sothevariable V=(¯X−m)/Shasthedensity
r(v)=integraldisplay∞
0√n
σ√
2πexpparenleftbigg
−nv2s2
2σ2parenrightbiggn(n−1)/2sn−2e−ns2/2σ2
2n−3
2Ŵ(n−1
2)σn−1sds
=nn/2
σn√π2(n−2)/2Ŵ(n−1
2)integraldisplay∞
0e−ns2(v2+1)/2σ2sn−1ds.
Herewesubstitute z=s2andobtain
r(v)=nn/2
σn√π2n/2Ŵ(n−1
2)integraldisplay∞
0e−nz(v2+1)/2σ2z(n−2)/2dz.
Nowwesubstitute Ŵ(1/2)=√π,definetheparameter
a=n(v2+1)
2σ2,
andtransform theintegralinto Ŵ(n/2)/an/2tofind
r(v)=Ŵ(n/2)
√πŴ(n−1
2)(v2+1)n/2,−∞<v<∞.
1148 Chapter 19 Probability
FIGURE 19.12Studenttprobability
densitygn(y)forn=3.
Table 19.3 StudenttDistribution
pn =1 n=2 n=3 n=4 n=5
0.8 1 .38 1 .06 0 .98 0.94 0.92
0.9 3 .08 1 .89 1 .64 1.53 1.48
0.95 6 .31 2 .92 2 .35 2.13 2.02
0.975 12 .74 .30 3 .18 2.78 2.57
0.99 31 .86 .96 4 .54 3.75 3.36
0.999 318 .32 2.31 0.2 7.17 5.89
Entries arethe values CinP(C)=KnintegraltextC
−∞(1+t2
n)−(n+1)/2dt=p,nis the number of degrees of freedom.
Finally we rescale this expression to the variable tin Eq. (19.72) with the density
(Fig.19.12)
g(t)=Ŵ(n/2)
√π(n−1)Ŵ(n−1
2)(1+t2
n−1)n/2,−∞<t<∞,(19.75)
fortheStudent tdistribution,whichmanifestlydoesnotdependon morσ.Theprobability
fort1<t<t2isgivenbytheintegral
P(t1,t2)=Ŵ(n/2)√π(n−1)Ŵ(n−1
2)integraldisplayt2
t1dt
(1+t2
n−1)n/2, (19.76)
andP(z)≡P(−∞,z)is tabulated. (See Table 19.3 for example.) Also, P(∞,−∞)=1
andP(−z)=1−P(z),becausetheintegrandinEq(19.76) isevenin t,so
integraldisplay−z
−∞dt
(1+t2
n−1)n/2=integraldisplay∞
zdt
(1+t2
n−1)n/2
and
integraldisplay∞
zdt
(1+t2
n−1)n/2=integraldisplay∞
−∞dt
(1+t2
n−1)n/2−integraldisplayz
−∞dt
(1+t2
n−1)n/2.
19.6 Statistics 1149
Multiplying this by the factor preceding the integral in Eq. (19.76) yields P(−z)=1−
P(z).In the following example we show how to apply the Student tdistribution to our fit
ofExample19.6.1.
Example 19.6.2 CONFIDENCE INTERVAL
Here we want to determine a confidence interval for the slope ain the linear y=atfit of
Fig.19.8.Weassume
•firstthatthesamplepoints (tj,yj)are randomandindependent,and
•secondthat,foreachfixedvalue t,therandomvariable Yisnormalwithmean µ(t)=
atandvariance σ2independentof t.
These values yjare measurements of the random variable Y,but we will regard them
as single measurements of the independent random variables Yjwith the same normal
distributionas Y(whosevariancewedonotknow).
We chooseaconfidencelevel, p=95%,say.ThentheStudentprobabilityis
P(−C,C)=P(C)−P(−C)=p=−1+2P(C),
hence
P(C)=1
2(1+p),
usingP(−C)=1−P(C),and
P(C)=1
2(1+p)=0.975=KnintegraldisplayC
−∞parenleftbigg
1+t2
nparenrightbigg−(n+1)/2
dt,
whereKn−1is the factor preceding the integral in Eq. (19.76). Now we determine a solu-
tionC=4.3 from Table 19.3 of Student’s tdistribution, with n=N−r=3−1=2t h e
numberofdegreesoffreedom,notingthat (1+p)/2 correspondsto pinTable19.3.
Then we compute A=Cσa/√
Nfor sample size N=3.The confidence interval is
givenby
a−A≤a≤a+A,atp=95%confidencelevel .
From the χ2analysis of Example 17.6.1 we use the slope a=0.782 and variance σ2
a=
0.0232,soA=4.30.023√
3=0.057,and the confidence interval is determined by a−A=
0.782−0.057=0.725,a+A=0.839,or
0.725<a<0.839 at95%confidencelevel .
Comparedto σa,theuncertaintyof ahasincreasedduetothehighconfidencelevel.Alook
at Table 19.3 shows that a decrease in confidence level, p,reduces the uncertainty inter-
val, and increasing the number of degrees of freedom, n,would also lower the range of
uncertainty. /squaresolid
1150 Chapter 19 Probability
Exercises
19.6.1 Let/Delta1Abetheerror ofameasurementof A,etc.Useerror propagationtoshowthat
parenleftbiggσ(C)
Cparenrightbigg2
=parenleftbiggσ(A)
Aparenrightbigg2
+parenleftbiggσ(B)
Bparenrightbigg2
holdsfortheproduct C=ABandtheratio C=A/B.
19.6.2 Find the mean value and standard deviation of the sample of measurements x1=
6.0,x2=6.5,x3=5.9,x5=6.2.If the point x6=6.1 is added to the sample, how
doesthechangeaffect themeanvalueandstandarddeviation?
19.6.3 (a)Carryouta χ2analysisofthefitofcase binFig.19.9assumingthesameerrorsfor
theti,/Delta1 ti=/Delta1yi,a sf o rt h e yiused in the χ2analysis of the fit in Fig. 19.8. (b) Deter-
minetheconfidenceintervalat95% confidencelevel.
19.6.4 Ifx1,x2,...,xnareasampleofmeasurementswithmeanvaluegivenbythearithmetic
mean¯xand the corresponding random variables Xjthat take the values xjwith the
same probability are independent and have mean value µand variance σ2,then show
that/angbracketleft¯x/angbracketright=µandσ2(¯x)=σ2/n.If¯σ2=1
nsummationtext
j(xj−¯x)2is the sample variance, show
that/angbracketleft¯σ2/angbracketright=n−1
nσ2.
AdditionalReadings
Kreyszig, E., Introductory Mathematical Statistics: Principles and Methods. NewYork: Wiley (1970).
Suhir, E., Applied Probability for Engineers and Scientists. NewYork: McGraw-Hill(1997).
Papoulis, A., Probability, Random Variables, and Stochastic Processes ,3rd ed.NewYork: McGraw-Hill(1991).
Ross, S. M., FirstCourse in Probability , 5th ed.,Vol. A.NewYork: Prentice-Hall (1997).
Ross, S. M., Introduction to Probability Models ,7th ed.NewYork: AcademicPress (2000).
Ross,S.M., IntroductiontoProbabilityandStatisticsforEngineersandScientists ,2nded.NewYork:Academic
Press (1999).
Chung, K.L., ACourse in Probability Theory Revised ,3rd ed. NewYork: AcademicPress (2000).
Devore,J.L., ProbabilityandStatisticsforEngineeringandtheSciences ,5thed.NewYork:DuxburyPr.(1999).
Montgomery,D.C.,andG.C.Runger, AppliedStatisticsandProbabilityforEngineers ,2nded.NewYork:Wiley
(1998).
Degroot, M.H., Probability and Statistics , 2nd ed.NewYork: Addison-Wesley (1986).
Bevington,P.R.,andD.K.Robinson, DataReductionandErrorAnalysisforthePhysicalSciences ,3rd.ed.New
York: McGraw-Hill(2003).
GeneralReferences
Additional, more specializedreferences arelistedatthe endof eachchapter.
1. E. T. Whittaker and G. N. Watson, A Course of modern Analysis , 4th ed. Cambridge, UK: Cambridge Uni-
versity Press (1962), paperback. Although this is the oldest of the references (original edition 1902), it still
is the classicreference.It leansstrongly toward pure mathematics,as of1902, with full mathematicalrigor.
2. P.M.MorseandH.Feshbach, MethodsofTheoreticalPhysics ,2vols.NewYork:McGraw-Hill(1953).This
work presents the mathematics of much of theoretical physics in detail but at a rather advanced level. It is
recommended as the outstanding source of information for supplementary reading and advancedstudy.
19.6 General References 1151
3. H. S. Jeffreys and B. S. Jeffreys, Methods of Mathematical Physics , 3rd ed. Cambridge, UK: Cambridge
University Press (1972). This is a scholarly treatment of a wide range of mathematical analysis, in which
considerableattentionispaidtomathematicalrigor.Applicationsaretoclassicalphysicsandtogeophysics.
4. R. Courant and D. Hilbert, Methods of Mathematical Physics , Vol. 1 (1st English ed.). New York: Wiley
(Interscience)(1953). Asareferencebookformathematicalphysics,itisparticularly valuableforexistence
theoremsanddiscussionsofareassuchaseigenvalueproblems,integralequations,andcalculusofvariations.
5. F. W. Byron Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics , Reading, MA: Addison-
Wesley, reprinted, Dover (1992). This is an advanced text that presupposes a moderate knowledge of math-
ematicalphysics.
6. C. M. Bender and S. A. Orszag, Advanced Mathematical Methods for Scientists and Engineers .N e wY o r k :
McGraw-Hill(1978).
7.Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables, Applied Math-
ematics Series-55 (AMS-55). Washington, DC: National Bureau of Standards, U.S. Department of Com-
merce;reprinted,Dover(1974).Asatremendouscompilationofjustwhatthetitlesays,thisisanextremely
useful reference.
This page intentionally left blank
INDEX
Numbers
1-forms, 304–5
2-forms, 305–6
3-forms, 306–7
A
Abeliangroup, 242
Abel’sequation, 1016
Abel’stest, 351, 665, 882
Abel’stheorem, 882
absolute convergence, 340–42, 350, 363
addition
of matrices, 178–79
of series, 324–25
of tensors, 136
addition rule, 1111
addition theorem forspherical harmonics,
797–802
Bessel functions, 636
derivation of addition theorem, 798–800
Legendre polynomials, 798
trigonometric identity, 797–98
adjoint operator property, 208
algebraicform, 405
aliasing, 916–17
alternating series, 339–42
absolute convergence, 340–42
exercises, 342
Leibniz criterion, 339–40
overview, 339
analyticcontinuation, 432–34
analyticfunctions, 415–18
z∗,416
z2,415
analyticlandscape, 489–90
angularMathieuequation, 872
angular momentum, 18, 215, 267, seealsoorbital
angular momentum
coupling, 266–78
Clebsch–Gordan coefficients SU(2)and
SO(3), 267–70
exercises, 277–78
overview, 266–67
spherical tensors, 271–74
young tableaux for SU(n), 274–77angular momentum operators, 261
Clebsch–Gordan coefficients, 803
orbital, 793–96
spherical harmonics, 793
vector spherical harmonics, 813–16
annihilation operator, 824
anomalous dispersion, 998
anticommutation relation, 18
anti-Hermitian matrices,221–23
eigenvalues
degenerate,223
andeigenvectors ofreal symmetric matrices,
221–22
overview, 221
antisymmetry
antisymmetric matrices, 204
and determinants, 168–72
Gausselimination, 170–72
overview, 168–70
of tensors, 137, 147
arealawfor planetary motion, 116–19
Argand diagram, 405
associatedLegendre equation functions, see
Legendre equation, functions polynomials
associative, 2
asymptotic expansions, 719–25
asymptotic forms
offactorial function Ŵ(1+s),494–95
of Hankelfunction, 493–94
asymptotic series, 389–96
Bessel functions, 722
confluent hypergeometric functions, 393
cosine and sine integrals, 392–93
definition of, 393–94
exercises, 394–96
incomplete gamma function, 389–92
overview, 389
overview: integral representation expansion,
719–23
steepestdescent, 489–95
Stokes’ method, 719, 724
asymptotic values, Besselfunctions, 722, 729
attractor, 1081
autonomous differential equations, 1091–93
average value,1117–18
axes,seerotations; symmetry
1153
1154 Index
axialvector, 143–44
axis,coordinate, 4–5, 7–11, 195–99
azimuthal dependence—orthogonality, 787
B
Baker–Hausdorff formula, 225
basin of attraction,1081
Bernoulli and Riccatiequations, 1089–90
Bernoulli numbers, 376–89, 473–74
Euler–Maclaurin integration formula, 380–82
exercises,385–89
improvement of convergence, 385
overview, 376–79
polynomials, 379–80
Riemann Zetafunction, 382–84
Bessel functions, 675–739, 865
asymptotic expansions, 719–25
exercises, 723–25
expansion ofan integral representation,
720–23
asymptotic values,722
closure equation, 696
of first kind, 675–93
alternateapproaches, 685–86
Bessel functions of nonintegral order, 686
Bessel’s differential equation, 678–79
Bessel’s differential equation: self-adjoint
form, 694
confluent hypergeometric representation, 865
cylindrical resonant cavity, 682–85
cylindrical waveguide, 705
exercises, 686–93
Fourier transform, 933
Fraunhofer diffraction, circularaperture,
680–82
generatingfunction for integralorder,675–77
integral representation, 679–80
Laplacetransform solution, 983–84
orthogonality, 694
recurrence relations, 677–78
second kinds, 699–707
series solution, 570–72, 676–79
singularities, 564
spherical, 725–39
Wronskian, 702–5
Fraunhofer diffraction, 680–82
Hankelfunctions, 707–13
contour integral representation of, 709–11
cylindrical traveling waves,708
definitions, 707–8
exercises, 711–13
Helmholtz equation, 683–84, 725
Laplace’sequation, 695
modified, 713–19
asymptotic expansion, 711, 719exercises,716–19
Fourier transform, 716
generating function, 709
integral representation, 720–23
Laplacetransform, 933
recurrencerelations, 714–16
seriesform, 714
Neumann functions, Bessel functions of second
kind, 699–707
coaxialwaveguides, 703–4
definition and series form, 699–700
exercises,704–7
other forms, 701
recurrencerelations, 702
Wronskian formulas, 702–3
of nonintegral order, 686
orthogonality, 694–99
Besselseries, 695
continuum form, 696
electrostaticpotential in a hollow cylinder,
695–96
exercises,697–99
normalization, 695
recurrence relations, 677
spherical, 725–39
asymptotic values,729
definitions, 726–29
exercises,732–39
limiting values, 729–30
orthogonality, 731
particlein a sphere, 731–32
recurrencerelations, 730
spherical waves,730
in wave guides, 703–4
zeros,682
Bessel’s differential equation, 678–79, 684
self-adjoint form, 694
Bessel’s equation, 983–84
Bessel series,695
Bessel’s inequality, 651–52
betafunction, 520–26
definite integrals, alternate forms, 521–22
derivation of Legendre duplication formula,
522–23
incomplete, 523
Laplaceconvolution, 993
verification of πα/sinπαrelation, 522
bifurcates, 1083
bifurcations in dynamical systems, 1103–4
Hopf, 1101, 1103–4, 1107
pitchfork, 1083, 1086, 1101, 1107
binomial coefficient,356
binomial distribution, 1128–30
binomial expansion, 1129
binomial probability distribution, 1129
Index 1155
binomial theorem, 356–57
Biot and Savart law, 780–82
blackhole, optical pathnearevent horizon of,
1041–42
Bohr radius, 844
Born approximation, 603
Bose–Einstein statistics, 1064, 1115
boundary conditions, 542–43
Cauchy, 542
Dirichlet, 543
hollow cylinder, 695
magnetic field of current loop, 778–82
Neumann, 543
ring of charge,761
sphere in uniform electricfield, 759
Sturm–Liouville theory, 627–29
waveguide, coaxial cable,703
bound state,627
box counting dimension, 1086
branchcut (cut line), 409
branch points, 440–42
and multivalent functions, 447–50
of order 2, 440–42
Bromwich integral, 994–95
Butterfly effect,1079
C
calculusof residues, 455–82, seealsodefinite
integrals
Cauchy principal value, 457–60
exercises, 474–82
Jordan’s lemma, 466–68
overview, 455
pole expansion of meromorphic functions, 461
product expansion of entire functions, 462–63
residue theorem, 455–56
calculusof variations, 1037–77
applications of theEuler equation, 1044–52
exercises, 1049–52
soap film, 1045–46
soap film—minimum area,1046–49
straight line, 1044–45
dependent and an independent variable,
1038–44
alternateforms of Eulerequations, 1042
concept of variation, 1038–41
exercises, 1043–44
missing dependent variables, 1042–43
optical pathnearevent horizon ofablack
hole, 1041–42
Lagrangian multipliers, 1060–65
constraints, 1060–72
cylindrical nuclearreactor, 1062
exercises, 1063–65
particle in abox, 1061–62Rayleigh–Ritz variational technique, 1072–76
exercises,1074–76
ground stateeigenfunction, 1073
Sturm–Liouville equation, 1072
vibrating string, 1074
several dependent and independent variables,
1058–59
exercises,1059
relation to physics, 1059
several dependent variables, 1052–58
exercises,1055–56
Hamilton’s principle, 1053–54
Laplace’sequation, 1057–58
moving particle—Cartesian coordinates,
1054
moving particle—circular cylindrical
coordinates, 1054–55
several independent variables,exercises, 1058
surface of revolution, 1046
uses of, 1037
variation with constraints, 1065–72
exercises,1070–72
Lagrangian equations, 1066–67
Schrödinger wave equation, 1069–70
simple pendulum, 1067–68
sliding off a log, 1068–69
Cartesian components, 4
Cartesian coordinates, 554–55
unit vectors, 5
Casimir operators, 265
Catalan’s constant, 384, 513
catenoid,catenary ofrevolution, 1046
Cauchy (Maclaurin) integral test, 327–30, 418–30
contour integrals, 418–20
derivatives, 426–27
exercises, 424–25, 429–30
Goursat proof, 421–23
Morera’s theorem, 427–28
multiply connected regions, 423–24
overview, 418, 425–26
Stokes’ theorem proof, 420–21
Cauchy boundary conditions, 542
Cauchy criterion, 322
Cauchy inequality, 428
Cauchy principal value,457–60, 471
Cauchy–Riemann conditions, 413–18
analyticfunctions, 415–18
z∗,416
z2,415
exercises, 416–18
overview, 413–15
causality, 486–87
cavities,cylindrical, 682–85
Cayley–Klein parameters,252
centeror cycle,1100–1101
1156 Index
centralforce, 117
centralforce field, 39–40, 44–46
centrifugal potentials, 72
chainrule, 34
chaos in dynamical systems, 1105–6
chaoticattractor, 1085
character,184, 293
characteristics,538–41
Chebyshev differential equation, 559
Chebyshev polynomials, 848–59
generating functions, 848
Gram–Schmidt construction, 646
hypergeometric representations, 862
orthogonality, 854–55
recurrencerelation, 850
recurrencerelations—derivatives, 852–53
shifted, 646, 850
trigonometric form, 853–54
type I, 849–52
type II, 849
chi-squared ( χ2) distribution, 1143–45
Christoffel symbols, 154–56, 314
circularcylinder coordinates, 115–23
arealaw for planetary motion, 116–19
exercises,120–23
Navier–Stokes term, 119
overview, 115–16
circularcylindrical coordinates, 555–56
expansion, 601–2
circularmembrane, Bessel functions, 693, 708
classesandcharacter,293
Clausen functions, 909
Clebsch–Gordan coefficients, 267–70, 803
Clifford algebra,211–12
closed-form solutions, 810–12
closure, of Bessel function, 89, 248, 696
closure, of spherical harmonics, 790, 792
coaxialwaveguides, 703–4
commutative, 1
commutator, 180, 225, 231, 249, 253, 262, 264
comparison tests, 325–26
completeness of eigenfunctions
of Fourier series: of Sturm–Liouville
eigenfunctions, 649–51
of Hilbert–Schmidt: of integral equations,
1031–33
complex variables,403–54, 455–97, seealso
calculusof residues; Cauchy–Riemann
conditions; functions; mapping; saddle
points (steepest descent method);
singularities
algebra using, 404–13
calculus of residues, 455–82
complex conjugation, 407–8
exercises, 409–13overview, 404–5
permanenceof algebraicform, 405–7
Cauchy’s integral formula, 425–30
derivatives, 426–27
exercises,429–30
Morera’s theorem, 427–28
overview, 425–26
Cauchy’s integral theorem, 418–25
Cauchy–Goursat proof, 421–23
contour integrals, 418–20
exercises,424–25
multiply connected regions, 423–24
overview, 418
Stokes’ theorem proof, 420–21
dispersion relations, 482–89
causality,486–87
exercises,487–89
opticaldispersion, 484–85
overview, 482–83
Parseval relation, 485–86
symmetry relations, 484
Laurentexpansion, 430–38
analyticcontinuation, 432–34
exercises,437–38
Schwarzreflection principle, 431–32
Taylorexpansion, 430–31
overview, 403–4, 455
conditional convergence, 340
conditional probability, 1112
condition number, of ill-conditioned systems, 234
Condon–Shortley phase conventions, 270
confidence interval, 1146, 1149
confluent hypergeometric functions, 863–69
asymptotic expansions, 866
Bessel and modified Bessel functions, 865
Hermitefunctions, 866
integral representations, 865
Laguerrefunctions, 837–48
miscellaneous cases,866
Whittaker functions, 866
Wronskian, 868
conformal mapping, 451–54
conjugation, complex, 407–8
connected,simply or multiply, 60, 95, 420, 423,
426, 435
conservation theorem, 309
conservative force, 34, 69
constant 1-forms, 305
constantBfield, vector potentials of, 44
contiguous function relations, 861
continuation, analytic,432–34
continuity equation, 40–42
continuity of power series, 364
continuous random variable, 1117–19
continuum form, 696
Index 1157
contour integral representation, Hankelfunctions,
709–11
contour integrals, 418–20
contour of integration, simple pole on, 468–69
contraction, 139
contravariant tensor, 135, 156–58, 160, 162
contravariant vector, 134
convergence, rate of, 334, 345
convergence of infinite series,321, 903
absolute, 340–42, 350
improvement of, and Bernoulli numbers, 385
of infinite product, 397–98
of powerseries, 363
rate, 334, 345
and rational approximations, 345
tests, 325–39, seealsoCauchy(Maclaurin)
integral test
comparison, 325–26
exercises, 335–39
Gauss’, 332–33, 357
improvement of, 334–39
Kummer’s, 330–32
overview, 325
partial sum approximation, 390
Raabe’s, 332
uniform and nonuniform, 348–49
convolution (Faltungs) theorem, 951–55, 990–94
driven oscillator with damping, 991–93
Fourier transform, 931–32, 936–45
Laplacetransform, 965–1003
Parseval’s relation, 952–53
coordinates, seealsocircularcylinder coordinates;
curved coordinates and vectors; orthogonal
coordinates; spherical polarcoordinates
axes, rotation of, seerotations
curvilinear, 104, 105, 110, 111, 112
divergence of coordinate vector, 39
Laplacianin orthogonal, 316
rotation of, 199
correlation, 1122
cosets and subgroups, 293–94
cosines
asymptotic expansion, 392–93
confluent hypergeometric representation, 867
cosine transform, 939
direction, 196–97
direction cosines (orthogonal matrices), 4,
196–97, 201
functions of infinite products, 398–99
infinite product, 398, 462
integral, 392
integral of in denominator, 464–65
integrals in asymptotic series,392–93
law of, 16–17, 118, 745
law of, theorem, 16theorem, 16
coupling, angular momentum, seeangular
momentum
covariance,1122
covariance ofMaxwell’sequations, Lorentz, see
Lorentz covarianceof Maxwell’sequations
covariant derivative, 151, 156
covariant vectors, 134, 152–53
tensor, 135, 156–58, 160, 162
Cramer’s rule, 166
creation operator, 824
criterion, Leibniz, 339–40
criticalpoint, 1091–93
criticalstrip, 897
criticaltemperature, 1086
crossing conditions, 484
cross product, 18–22, seealsotriple vector
products
exercises, 22–25
overview, 18–22
of vectors, 315
crystallographic point and spacegroups, 299–300
curl,∇×, 43–49
centralforce field, 44–46
asdifferential vector operator, 112–13
exercises, 47–49
gradient of dot product, 46
integral definitions of gradient, divergence and,
58–59
integration by parts of, 47
overview, 43
astensor derivative operator, 162–63
vector potential ofconstant Bfield,44
curl,∇×
centralforce field
in circularcylindrical coordinates, 118
in curvilinearcoordinates, 112–13
in spherical polarcoordinates, 126
irrotational, 45
curved coordinates and vectors, 103–33, seealso
circularcylinder coordinates; orthogonal
coordinates; spherical polar coordinates
differential operators, 110–14
curl,112–13
divergence, 111–12
exercises,113–14
gradient, 110
overview, 110
overview, 103
specialcoordinate systems, 114–33
curves, fitting to data, 1140–43
curvilinear coordinates, 104, 105, 110, 111, 112
cut line (branch cut),409
cylinder coordinates, circular, seecircularcylinder
coordinates
1158 Index
cylindrical coordinates, 104
cylindrical symmetry, 617
cylindrical traveling waves,708
D
d’Alembertian, 141
d’Alembert ratio test,326–27
damped oscillator, 979–80
damped simple harmonic oscillation, 980
decay,kaon, 282–83
definite integral (Euler), 500–501
definite integrals
evaluation of, 463
exponential forms, 471–82
Bernoulli numbers, 473–74
factorial function, 472–73integraltext∞
−∞f(x)dx, 465–66integraltext∞
−∞f(x)eiaxdx, 466–71
quantum mechanicalscattering, 469–71
simplepoleoncontourofintegration,468–69integraltext2π
0f(sinθ,cosθ)dθ, 464–65
degeneracy, of Schrödinger’s waveequation, 638
degenerateeigenfunctions, 638
degenerateeigenvalues, 223, 638
Del(∇), 42, 43
forcentral force, 127
successiveapplications of, 49–54
electromagnetic waveequation, 51–53
exercises, 53–54
Laplacianofpotential, 50–51
overview, 49–50
deltafunction, Dirac,83–85, 669–70, 975
Besselrepresentation, 935
in circularcylindrical coordinates, 601
derivation, 937–38
eigenfunction expansion, 89, 650
exercises,91–95
Fourier integral, 90
Fourier representation, 90
Green’sfunction and, 592–610
impulse force, 975
integral representations for, 90
Laplacetransform, 975
overview, 83–87
phasespace,88
point source, 88, 593
quantum theory, 955–61
representation by orthogonal functions, 88–89
sequences,83, 86
sine,cosine representations, 943
in spherical polarcoordinates, 82, 599
theory of distributions, 86
totalcharge inside sphere, 88
DeMoivre’s formula, 408–13denominator, integral of cos in, 464–65
dependent and independent variables,1038–44
alternateforms of Euler equations, 1042
conceptof variation, 1038–41
missing dependent variables, 1042–43
optical pathnearevent horizon of ablackhole,
1041–42
derivative operators, tensor
curl, 162–63
divergence, 160–61
exercises, 162–63
Laplacian,161–62
overview, 160
derivatives, 426–27, seealsoexterior derivative
covariant, 156
gauge covariant, 76
descending power series solutions, 369, 781
descent,steepest, seesaddle points (steepest
descentmethod)
determinants, 165–239
antisymmetry, 168–72
Gausselimination, 170–72
Gauss–Jordan elimination, inversion, 185–86
overview, 168–70
exercises, 174–76
Gram–Schmidt procedure, 173–74
overview, 173–74
vectors by orthogonalization, 174
homogeneous linear equations, 165–66
inhomogeneous linearequations, 166–67
Laplaciandevelopment by minors, 167–68
lineardependenceof vectors, 172–73
overview, 165
product theorem, 181
representation ofavector product, 20
secularequation, 218
solution of aset of homogenous equations,
165
solution of a set of nonhomogenous
equations, 166
deuteron, 626–27
diagonal matrices,182–83, 215–31, seealso
anti-Hermitian matrices
eigenvectors andeigenvalues, 216–19
exercises, 226–31
functions of, 224–26
Hermitian, 219–21
moment of inertia, 215–16
differential equations, 535–619, 751–52
first-order differential equations, 543–53
exactdifferential equations, 545–47
exercises,550–53
linearfirst-order ODEs,547–50
nonlinear, 1088–1102
parachutist, 544–49
Index 1159
RL circuit, 549–50
separable variables, 544–45
Fuchs’ theorem, 573
heat flow, or diffusion, PDE,611–18
alternatesolutions, 614–15
special boundary condition again,615–16
specific boundary condition, 612–13
spherically symmetric heatflow, 616–17
homogeneous, 536, 548–50, 565
linearindependence of solutions
second solution, 581–83
series form of the second solution, 583–85
nonhomogeneous equation—Green’s function,
592–610
circularcylindrical coordinate expansion, 27,
601–2
exercises, 607–10
form of Green’s functions, 596–98
Legendre polynomial addition theorem,
600–601
quantum mechanicalscattering—Green’s
function, 603–6
quantum mechanicalscattering—Neumann
series solution, 602–3
spherical polar coordinate expansion,
598–600
symmetry of Green’sfunction, 595–96
partial differential equations, 535–43
boundary conditions, 542–43
classes ofPDEsand characteristics,538–41
examples of, 536–38
introduction, 535–36
nonlinear PDEs,541–42
particular solution, 548, 565
second solution, 573–78
exercises, 588–92
lineardependence,580–81
linearindependence, 580
linear independence of solutions, 579–80
second solution for the linear oscillator
equation, 583
second solution of Bessel’s equation, 586–87
second solution logarithmic term, 585, 700
separation of variables,554–62
Cartesian coordinates, 554–55
circularcylindrical coordinates, 555–56
exercises, 560–62
spherical polar coordinates, 557–60
series solutions—Frobenius’ method, 565–78
exercises, 574–78
expansion about x0,569
Fuchs’ theorem, 573
limitations ofseries approach—Bessel’s
equation, 570–71
regular and irregular singularities, 572–73symmetry of solutions, 569
singular points, 562–65
differential forms, 304–19, seealsopullbacks
1-forms, 304–5
2-forms, 305–6
3-forms, 306–7
exercises, 318–20
exterior derivative, 307–9
Hodge operator *, 314–19
cross product of vectors, 315
Laplacianin orthogonal coordinates, 316
Maxwell’sequations, 316–20
overview, 314–15
Legendre’s equation, 752
overview, 304
Stokes’ theorem on, 313–14
differentials, exact, seethermodynamics
differential vector operators, 110–14
adjoint, 621, 623, 634–35
curl, 112–13
del,33
divergence, 111–12
exercises, 113–14
gradient, 110
overview, 110
differentiation, 904–5
differentiation ofpowerseries, 364
diffraction, 680–82, 1017
diffusion equation, seedifferential equations, heat
flow PDE
digamma and polygamma functions, 510–16
digamma functions, 510–11
Maclaurinexpansion, computation, 512
polygamma function, 511–12
series summation, 512
dihedral groups, Dn, 299
dimension
box-counting, 1086
Hausdorff, 1086
Kolmogorov, 1086
dimensionality theorem, 298
dipoles, interactionenergy, magnetic dipole,
radiation fields, seeelectricdipole
Diracbra-ket notation, 177
Diracdeltafunction, seedeltafunction, Dirac
Diracmatrices, 209–12
direction cosines (orthogonal matrices), 4,
196–97, 201
direct product, 139–41
exercises, 140–41
and matrix multiplication, 181–82
overview, 139–40
of tensors, 140
direct tensor, 181
direct tensor product, 139
1160 Index
Dirichletconditions, 543, 882
kernel,910
Dirichletproblem, 617
Dirichletseries, 326
discontinuities, behavior of, 886
discontinuous functions, 888
discreteFourier transform, 914–19
aliasing,916–17
fastFourier transform, 917
limitations, 916
orthogonality over discretepoints, 914–15
discrete groups, 291–304
classesandcharacter,293
crystallographic point and spacegroups,
299–300
dihedral groups, Dn,299
exercises,300–304
subgroups and cosets,293–94
threefold symmetry axis, 296–99
twofold symmetry axis, 294–96
discreterandom variable,1116–17
dispersion relations, 482–89
causality,486–87
crossing relations, 484
exercises,487–89
Hilbert transform, 485
opticaldispersion, 484–85
overview, 482–83
Parseval relation, 485–86
sum rules, 484
symmetry relations, 484
displacement, 158
dissipation, 1101–3
distributive, 13
divergence,∇, 38–43
of central force field,39–40
circularcylindrical, 115–23
coordinates, Cartesian, 4–7
of coordinate vector, 39
curvilinear coordinates, 111
asdifferential vectoroperator, 111–12
exercises,42–43
integral definitions of gradient, curl, and,58–59
integration by parts of, 40
overview, 38–39
physical interpretation, 40–42
solenoidal, 42
spherical polar, 123–33
astensor derivative operator, 160–61
divergent series, 325–26
Doppler shift, 360
dot products, 12–17
exercises,17
gradient of, 46invariance ofscalarproduct under rotations,
15–17
overview, 12–15
double series,rearrangement of, 345–48
driven oscillator with damping, 991–93
dual tensors, 147–48
duplication formula for factorialfunctions, see
Legendre duplication formula
dynamical systems, dissipation in,1102–3
E
E, Lorentztransformation of the electricfield,
287–88
Earth’s gravitational field, 758–59
Earth’s nutation, 973–74
eigenfunctions, 624, 626–27
Bessel’s inequality, 651–52
completeness of, 649–61
of Fourier series: of Sturm–Liouville
eigenfunctions, 649–51
of Hilbert–Schmidt: of integral equations,
1031–33
eigenvalue equation, 667–68
expansion, Green’s function, 662–74
of Diracdelta function, 88–90
of Hermitian differential operators, 635,
649–51
of square wave,637
expansion coefficients, 658
orthogonal, 636, 637, 1030–32
Schwarzinequality, 652–54
summary—vector spaces,completeness,
654–58
variational calculation,1072–76
eigenvalues, 217–18, 223, 624–27, 634–35
of Hermitial differential operator, 634
of Hermitial matrices, 219–23
of Hilbert–Schmidt integral equation, 1030–36
of normal matrices, 231–32
of real symmetric matrices,215–19
variational principle for, 1072
eigenvectors, 216–19, 221–22
eightfold way(weight diagram), 257
Einstein’s energy relation, 281
Einstein’s summation convention, 136
velocity addition law,283, 290
electricalcharge inside spheres, 88
electricdipole potential, Legendre polynomial
expansion, 745
electromagnetic invariants, 288
electromagnetic waveequation, 51–53
electromagnetic waves,981–82
electrostaticmultipole expansion, 599
electrostaticpotential, 593
in hollow cylinder, 695–96
Index 1161
of ring of charge,761–62
electrostatics,physical basis, 741–42
elimination, seedeterminants
ellipticaldrum, 872–73
ellipticintegrals, 370–76
definitions of, 372
exercises, 374–76
of first kind, 372
hypergeometric representations, 373, 860
limiting values, 374
overview, 370
period of simple pendulum, 370–71
of second kind, 372
series expansion, 372–73
ellipticPDEs,538
empty set, 1111
energy
potentials, 309
relativistic, 356–57
equality of matrices,178
equations, seealsolinearequations; Maxwell’s
equations; Poisson’s equation
of motion and field,142
error integrals, 530
asymptotic expansion, 393
confluent hypergeometric representation, 864
error propagation, 1138–40
essential(irregular) singular point, 563
Eulerangles, 202–3
Eulerequation, 1040
alternate forms of, 1042
applications of, 1044–52
soap film, 1045–46
soap film—minimum area,1046–49
straight line, 1044–45
Euleridentity, 224
product formula, 382–84
Eulerintegrals, 502
Euler–Maclaurin integration formula, 380–82, 517
Euler–Mascheroni constant, 330
Eulerproduct for Riemann Zetafunction, 382–84
event horizon, 1041
exactdifferential equations, 545–47
exactdifferentials, seethermodynamics
expansion, seealsoLaurent expansion; Taylor’s
expansion
pole, of meromorphic functions, 461
product, of entire function, 462–63
of series, 372–73
expansion coefficients, 658
expansion of functions, Legendre series, 757–58
expectation value,630, 955–56, 1117–18
exponential forms, 471–82
Bernoulli numbers, 473–74
factorial function, 472–73exponential function, of Maclaurintheorem,
354–55
exponential integral, 527–30
exponential of diagonal matrix, 225–26
exponential transform, 938
exterior derivative, 307–9
extreme or stationary value,36, 881, 1038–40
F
Factorial function Ŵ(1+s)
asymptotic form of, 494–95
complex argument, 499–506
contour integrals, 505
digamma function, 510
double factorialnotation, 505
Gammafunctional relation, 503
infinite product, 501
integral (Euler) representation, 500
Legendre duplication formula, 503
Maclaurinexpansion, 512
polygamma functions, 511
steepestdescent asymptotic formula, 495
Stirling’s (formula) series, 516–18
factorial notation, 503–5
faithful group, 243
Faraday’s law,66–68
fast Fourier transform, 917
Feigenbaum number, 1083–84
Fermage equation, 950
Fermat’s principle, 157, 1041, 1050
Fermi–Dirac statistics, 1064, 1115
field equations, 142
finite wave train, 940–41
first-order differential equations, 543–53
exactdifferential equations, 545–47
linearfirst-order ODEs,547–50
separablevariables, 544–45
fixed andmovable singularities, specialsolutions,
1090
Floquet’s theorem, 877
force as gradient of potentials, 36
forced classicaloscillators, 457–60
force field, central, seecentral force field
Fourier–Bessel series, 695
Fourier expansions of Mathieufunctions, 919–29
integral equations and Fourier series for
Mathieu functions, 919–23
leadingcoefficients for ce 0,926–29
leadingcoefficients ofse 1,923–26
Fourier integral theorem, 937
development of, 936–38
exponential form, 937
Fourier–Mellin integral, 995
Fourier representation, of Diracdeltafunction, 90
Fourier series,881–930
1162 Index
advantages,uses of, 888–92
change of interval, 890–91
completeness, seegeneral properties of
Fourier series
convergence, 893, 903–10
differentiation, seegeneral properties of
Fourier series
discontinuous functions, 888
exercises, 891–92
integration, seegeneral properties ofFourier
series
periodic functions, 888–90
applications of, 892–903
exercises, 898–903
full-wave rectifier, 893–94
infinite series,Riemann Zetafunction,
894–98
square wave—high frequencies, 892–93
discreteFourier transform, 914–19
discrete fourier transform, 915–16
discreteFouriertransform—aliasing,916–17
exercises, 918–19
fast Fourier transform, 917
limitations, 916
orthogonality over discrete points, 914–15
Fourier expansions of Mathieu functions,
919–29
exercises, 929
integral equations and Fourier seriesfor
Mathieu functions, 919–23
leading coefficientsfor ce 0, 926–29
leading coefficientsof se 1,923–26
generalproperties, 881–88
behavior of discontinuities, 886
completeness, 883–84
complex variables—Abel’s theorem, 882
exercises, 886–88
sawtooth wave,885–86
square wave,892
Sturm–Liouville theory, 885
summation of aFourier series,882–83
Gibbs phenomenon, 910–14
calculation of overshoot, 912–13
exercises, 913–14
square wave,911–12
summation of series,910
orthogonality, 636–37
properties of, 903–10
convergence, 903
differentiation, 904–5
exercises, 905–10
integration, 904
summation of, 882–83
Fourier transform, of Gaussian,932
Fourier transform of derivatives, 946–51heatflow PDE,948–49
inversion of PDE,949–50
wave equation, 947–48
Fourier transforms, 486–87, 931–32
aliasing, 916
convolution (Faltungs) theorem, 951
deltafunction derivation, 937
transfer functions, 961
Fourier transforms—inversion theorem, 938–46
cosine transform, 939
exponential transform, 938
fastFourier transform, 917
finite wave train, 940–41
Fourier integral, 931, 936–39
momentum spacerepresentation, 955
sine transform, 939–40
uncertainty principle, 941
Fourier transform solution, 1013
fractals,1086–88
fractional order, 427
Fraunhofer diffraction, Bessel function, 680–82
Fredholm equation, 1005–6
Frobenius’ method, series solutions, 565–78
Fuchs’ theorem, 573
full-wave rectifier, 893–94
functional equation, Riemann Zetafunction,
896
Gamma function, 500
functions, 817–80, seealsoanalytic functions
Chebyshev polynomials, 848–59
exercises,855–59
generating functions, 848
orthogonality, 854–55
recurrencerelations—derivatives, 852–53
trigonometric form, 853–54
type I, 849–52
type II, 849
of complex variable, 408–13
confluent hypergeometric functions, 863–69
Besseland modified Besselfunctions, 865
exercises,867–69
Hermitefunctions, 866
integral representations, 865
miscellaneous cases,866
entire,415, 451, 462–63
exponential, of Maclaurin theorem, 354–55
factorial,472–73
Hermitefunctions, 817–36
alternaterepresentations, 819
applications of theproduct formulas, 831–32
directexpansion ofproducts of Hermite
polynomials, 828–30
exercises,832–36
generating functions—Hermitepolynomials,
817–18
Index 1163
orthogonality, 821–22
quantum mechanicalsimple harmonic
oscillator, 822–27
recurrence relations, 818–19
Rodrigues’ representation, 820–21
threefold Hermite formula, 827–28
hypergeometric functions, 859–63
contiguous function relations, 861
exercises, 862–63
hypergeometric representations, 861–62
Laguerre functions, 837–48
associated Laguerrepolynomials, 841–43
differential equation—Laguerre
polynomials, 837–41
exercises, 845–48
hydrogen atom, 843–45
Mathieu functions, 869–79
elliptical drum, 872–73
exercises, 879
general properties of Mathieufunctions, 874
quantum pendulum, 873
radial Mathieu functions, 874–79
separation ofvariables in elliptical
coordinates, 870–71
of matrices, 224–26
meromorphic, 451, 461–62, 478
multivalent, and branch points, 447–50
rotation of, 251
series of, 348–52
Abel’s test,351
exercises, 352
overview, 348
uniform and nonuniform convergence,
348–49
Weierstrass Mtest,349–50
G
gamma distribution, 1143
gamma function, seealsofactorialfunction,
499–533
beta function, 520–26
definite integrals, alternateforms, 521–22
derivation of Legendre duplication formula,
522–23
exercises, 523–26
incomplete betafunction, 523
verification of πα/sinπαrelation, 522
definitions, simple properties, 499–510
definite integral (Euler), 500–501
double factorial notation, 505
exercises, 506–10
factorial notation, 503–5
infinite limit (Euler), 499–500
infinite product (Weierstrass), 501–3
integral representation, 505–6digamma and polygamma functions, 510–16
Catalan’sconstant, 513
digamma functions, 510–11
exercises,513–16
Maclaurinexpansion, computation, 512
polygamma function, 511–12
seriessummation, 512
incomplete gamma functions andrelated
functions, 527–33
error integrals, 530
exercises,530–33
exponential integral, 527–30
of infinite product, 398–99
Stirling’s series,516–20
derivation from Euler–Maclaurin integration
formula, 517
exercises,518–20
Stirling’s series, 518
gauge covariant derivative, 76
theory, 259
transformation, 76
gauge theory, 76, 241
Gauss elimination, 170–72
Gauss’ differential equation, 312–13, 318,
614–15, 617–18
hypergeometric differential equation, 576,
859–62, 873
Gauss error integral, 500
asymptotic expansion, 530
Gauss’ fundamental theorem of algebra,428, 463
Gauss–Jordan matrix inversion, 185–87
Gauss’ law, 52, 79–83, 594
Gauss’ normal distribution, 1134–38
Gauss’ notation, 503
Gauss–Seidel iteration technique, 172
Gauss–Seidel method, 226
Gauss’ test, 332–33, 357
Legendre series,333
Gauss’ theorem, 60–64
alternateforms of, 62–64
exercises,62–64
overview, 62
overview, 60–61
pullbacks, 312–13
Gegenbauer polynomials, seeultraspherical
polynomials
general parabolic solution, 540
general properties, 881–88
behavior of discontinuities, 886
completeness, 883–84
complex variables—Abel’s theorem, 882
of Mathieu functions, 874
sawtooth wave,885–86
Sturm–Liouville theory, 885
summation of aFourier series, 882–83
1164 Index
general tensors, 151–60, seealsoChristoffel
symbols
covariant derivative, 156
exercises,158–60
geodesics and paralleltransport, 157–60
metrictensor, 151–54
overview, 151
generating function, 741–49, 848, 1014–15
associatedLaguerre polynomials, 624–25,
841–45
associatedLegendre functions, 771–82, 788
associatedLegendre polynomials, 773
Bernoulli numbers, 376–89, 473–74, 517
Bernoulli polynomials, 379–80
Besselfunctions, modified, 711, 713–19,
723–24, 865
Chebyshev polynomials, 848–59
extension to ultraspherical polynomials, 747
Hermitepolynomials, 817–30
for integral order, 675–77
Laguerre polynomials, 647, 837–41, 843–45
Legendre polynomials, 742–44
linearelectricmultipoles, 744–45
physical basis—electrostatics,741–42
ultraspherical polynomials, 747
vector expansion, 745–47
generators of continuous groups, 246–61, seealso
rotations;SU(2)
exercises,260–61
overview, 246–50
geodesic equation, 157
geodesics, 157–60
geometrical interpretation of gradient, 35–38
integration by parts of, 36–37
of potential, force as, 36
geometric series, 322–23
Gibbs phenomenon, 886, 910–14
calculationof overshoot, 912–13
square wave,911–12
summation of series,910
global behavior, 1093–1101
Goursat proof of Cauchy’s integral theorem,
421–23
gradient
in Cartesian coordinates, 37, 51, 113, 134
in circularcylindrical coordinates, 118
in curvilinearcoordinates, 110
in spherical polarcoordinates, 126
gradient, curvilinear coordinates, 110
gradient,∇,32–38
asdifferential vectoroperator, 110
of dot product, 46
exercises,37–38
geometrical interpretation, 35–38
integration by parts of, 36–37of potential, force as, 36
integral definitions of divergence, curl, and,
58–59
overview, 32–34
of potential, 34
Gram–Schmidt orthogonalization, 642–49
Gram–Schmidt procedure, 173–74
overview, 173–74
vectors by orthogonalization, 174
gravitational potentials, 72
great circle,1040
Green’s function, 662–74
construction of, one dimension, 598, 599,
663–65, 670
construction of, twodimension, 597–98
construction of, three dimension, 597–98
and Diracdeltafunction, 669–70
eigenfunction, eigenvalue equation, 667–68
eigenfunction expansion, 662–82
electrostaticanalog, 592, 665
form of, 596–98
Helmholtz, 598
Helmholtz equation, 662
integral—differential equation, 665–67
Laplaceoperator, 598
circularcylindrical expansion, 601–6
spherical polar expansion, 598–600
linear oscillator, 668–69
modified Helmholtz, 598
nonhomogeneous equation, 592–610
one-dimensional, 663–65
Poisson’s equation, 669–70
symmetry of, 595–96
Green’s theorem, 61–62, 593
Gregory series, 362, 368
ground stateeigenfunction, 1073
group theory, 241–320, seealsoangular
momentum; differential forms; generators
of continuous groups; homogeneous
Lorentz group
character,184
definition of, 242–43
discrete,291–93
classesandcharacter,293
crystallographic point and space,299–300
dihedral, Dn,299
exercises,300–304
subgroups and cosets,293–94
threefold symmetry axis, 296–99
twofold symmetry axis, 294–96
faithfulness, 243
homomorphic, 243
homomorphism and isomorphism, 243–45
homomorphism SU(2)–SO(3),252–6
isomorphic, 243
Index 1165
Lorentz covariance ofMaxwell’s equations,
283–91
electromagnetic invariants, 288
exercises, 289–91
overview, 283–86
transformation of EandB,287–88
overview, 241–42
permutation groups, 301–3
quantum chromodynamics (QCD), 258
reducible and irreducible representations,
245–46
special unitary group SU(2), 251
vierergruppe, 292–93, 296
Gutzwiller’strace formula, 898
H
Hadamard product, 208
Hamilton–Jacobi equation, 539
Hamilton’s principle, 1053–54
Hankelfunctions and Lagrangeequations of
motion, 493–94, 707–13
asymptotic forms, 493–94, 723
contour integral representation ofthe Hankel
functions, 709–11
cylindrical traveling waves,708
definition, by Neumann function, 707–8
definitions, 707–8
series expansion, 707, 715
spherical, 728–29
Wronskian formulas, 708
Hankeltransforms, 933
harmonic functions, 539
harmonic oscillator, 822–27, 958–59
harmonics, seealsospherical harmonics, vector
spherical harmonics
harmonic series,323–24
Hausdorff, 225, 249, 1086
heatflow PDE,948–49
Heavisideexpansion theorem, 442, 1003
Heavisideshifting theorem, 981
Heavisideunit step function, 93, 996
Heisenberg uncertainty principle, 732, 941
Helmholtz diffusion equation, 536, 537
Helmholtz equation, 556, 557, 613
Bessel function, 683–84, 725
Green’s function, 662
spherical coordinates, 725
Helmholtz operators, 598
Helmholtz’s theorem, 95–101
exercises, 100–101
overview, 95–96
Hermitefunctions, 817–36, 866
alternate representations, 819
applications of theproduct formulas, 831–32
confluent hypergeometric representation, 866direct expansion of products of Hermite
polynomials, 828–30
generating functions—Hermite polynomials,
817–18
Gram–Schmidt construction, 642
orthogonality, 821–22
quantum mechanicalsimple harmonic
oscillator, 822–27
recurrence relations, 818–19
Rodrigues representation, 820–21
threefold Hermite formula, 827–28
Hermite polynomials
direct expansion of products of, 828–30
generating functions, 817–18
orthogonality integral, 821
recurrence relations, 818–19
Rodrigues representation, 820–21
Hermitian matrices, 184, 209
and matrix diagonalization, 219–21
unitary and, 208–15
exercises,212–15
overview, 208–9
Pauli and Dirac,209–12
Hermitian matrices, anti-, 221–23
eigenvalues
degenerate,223
andeigenvectors ofreal symmetric matrices,
221–22
overview, 221
Hermitian operators, 629–30, 634–42
completeness of eigenfunctions, 649–58
degeneracy, 638
expansion in orthogonal
eigenfunctions—square wave,637
Fourier series—orthogonality, 636–37
integration interval, 628–29
orthogonal eigenfunctions, 636
properties of, 634–38
in quantum mechanics, 630
realeigenvalues, 634–35
Hilbert matrix, determinant, 235
Hilbert–Schmidt theory, 1029–36
nonhomogeneous integral equation, 1032–34
orthogonal eigenfunctions, 1030–32
symmetrization of kernels, 1029–30
Hilbert space,7, 535, 629, 638, 658, 885
Hilbert transforms, 483
Hodge * operator, 314–19
cross product of vectors, 315
Laplacianin orthogonal coordinates, 316
Maxwell’sequations, 316–20
overview, 314–15
holomorphic functions (analytic or regular
functions), 415
homogeneous equations, 165–66, 565
1166 Index
homogeneous Lorentz group, 278–83
exercises,283
kinematics and dynamics in Minkowski
space–time, 280–83
overview, 278–80
homomorphic group, 243
homomorphism, 243–45
overview, 243
rotations, 244–45
SU(2)andSO(3), 252–56
hooks, 275–76
Hubble’s law, 7
hydrogen atom, 843–45, 957–58
Schrödinger’s wave equation, 843
hyperbolic PDEs,538
hypercharge, 257
hypergeometric equation
alternateforms, 860, 864
second independent solution, 860
singularities, 564, 859, 864, 865, 873
hypergeometric functions, 859–63
contiguous function relations, 861
hypergeometric representations, 861–62
I
ill-conditioned matrices,234–35
imaginary part, 407
impulsive force, 975–76
incomplete gamma function, 389–92
confluent hypergeometric representation, 500,
864
recurrencerelations, 399, 512
independence, linear,173, 579–81, 643, 665, 703
of solutions of ordinary differential equations,
579–81
of vectors, 173, 579
indicial equation, 566
inertia matrix, moment of, 215–16
infinite limit (Euler), 499–500
infinite product (Weierstrass), 499–500, 501–3
infinite products, 378, 383, 396–99, 499, 501–3
convergence, 397–98
cosine,398–99
entirefunctions, 462
gamma function, 398–99
sine,398–99
infinite series, 321–401, seealsoalternating
series; powerseries; Taylor’s expansion
algebra of, 342–48
alternating series,342–43
convergence, 342–45
convergence: absolute, 342–44
convergence: Cauchy integral, 327–29
convergence: Cauchy root, 326
convergence: comparison, 325–26convergence: conditional, Leibniz criterion,
344
convergence: D’Alembert ratio, 326–27
convergence: Gauss’, 332–33
convergence: improvement of, 345
convergence: Kummer’s, 330–32
convergence: Maclaurinintegral, 327–30
convergence: Raabe’s, 332
convergence: tests of, 325–35
convergence: uniform, 348–51, 363–64
divergence of squares, 344
double series, 345–47
exercises,347–48
overview, 342–44
rearrangement of double, 345–48
asymptotic series,389–96
cosine and sine integrals, 392–93
definition of, 393–94
exercises,394–96
incomplete gamma function, 389–92
overview, 389
Bernoulli numbers, 376–89
Euler–Maclaurin integration formula, 380–82
exercises,385–89
improvement of convergence, 385
overview, 376–79
polynomials, 379–80
Riemann Zetafunction, 382–84
ellipticintegrals, 370–76
definitions of, 372
exercises,374–76
limiting values, 374
overview, 370
period of simple pendulum, 370–71
seriesexpansion, 372–73
of functions, 348–52
Abel’stest, 351
exercises,352
overview, 348
uniform and nonuniform convergence,
348–49
Weierstrass Mtest, 349–50
fundamental concepts, 321–25
addition and subtraction of, 324–25
exercises,325
geometric, 322–23
harmonic, 323–24
overview, 321–22
power series,363–66, 578–79
products of, 396–401
convergence of, 397–98
exercises,399–401
overview, 396–97
sine,cosine, and gamma functions, 398–99
Riemann’s theorem, 894–98
Index 1167
infinity,seesingularity, pole, essentialsingularity
inhomogeneous linearequations, 166–67
inhomogeneous ordinary differential equation
(ODE),Green’s function solutions, 663–64
inner product and matrix multiplication, 179–81
integral definitions of gradient, divergence, and
curl, 58–59
integral—differential equation, 665–67
integral equations, 1005–36
and Fourier series for Mathieu functions,
919–23
Fredholm equations, 1005, 1007, 1010, 1013,
1018, 1021–24, 1030–32
Hilbert–Schmidt theory, 1029–36
exercises, 1034–36
nonhomogeneous integral equation, 1032–34
orthogonal eigenfunctions, 1030–32
symmetrization of kernels, 1029–30
integral transforms, generating functions,
1012–18
exercises, 1015–18
Fourier transform solution, 1013
generalizedAbelequation, convolution
theorem, 1014
generating functions, 1014–15
introduction, 1005–12
definitions, 1005–6
exercises, 1011
linear oscillator equation, 1009–11
momentum representation in quantum
mechanics,1006–7
transformation of adifferential equation into
an integral equation, 1008–9
Neumann series,separable (degenerate) kernels,
1018–29
exercises, 1025–29
Neumann series,1018–19
Neumann seriessolution, 1020–21
numerical solution, 1023–25
separable kernel,1021–22
Volterra equations, 991, 1005–6, 1009–11, 1021
integral form, Neumann functions, 701
integral representations, 505–6, 679–80
for Dirac deltafunction, 90
expansion of, 720–23
integrals, see alsoCauchy(Maclaurin) integral
test;definite integrals; ellipticintegrals
contour integration, 463, 471, 503, 522, 603
differentiation of, 590
evaluation of, 810
Lebesgue, 649, 657
line, 55–56, 65–67, 440
of products of three spherical harmonics, 803–6
Riemann, 55, 60–61, 605, 636
Stieltjes, 86, 873surface,56–57
volume, 57–58
integral test,Cauchy, seeCauchy(Maclaurin)
integral test
integral transforms, 931–1004
convolution (Faltungs) theorem, 990–94
driven oscillator withdamping, 991–93
exercises,993
convolution theorem, 951–55
exercises,953–55
Parseval’s relation, 952–53
development of theFourier integral, 936–38
Diracdeltafunction derivation, 937–38
Fourier integral—exponential form, 937
Fourier transform, 931–32
Fourier transform of derivatives, 946–51
heatflow PDE,948–49
inversion of PDE,949–50
wave equation, 947–48
Fourier transform of Gaussian, 932
Fourier transforms—inversion theorem,
938–46
cosinetransform, 939
exercises,942–46
exponential transform, 938
finite wave train, 940–41
sine transform, 939–40
uncertainty principle, 941
Fourier transforms of derivatives, 950–51
generating functions, 1012–18
Fourier transform solution, 1013
generalizedAbelequation, convolution
theorem, 1014
generating functions, 1014–15
integral transforms, 931–35
exercises,934–35
Fourier transform, 931–32
Fourier transform of Gaussian,932
Laplace,Mellin, andHankeltransforms, 933
linearity, 933–34
inverse Laplacetransform, 994–1003
Bromwich integral, 994–95
exercises,1000–1003
inversion via calculusof residues, 996
summary—inversion of Laplacetransform,
999–1000
velocityof electromagneticwaves in a
dispersive medium, 997–99
Laplace,Mellin,andHankeltransforms, 933
Laplacetransform of derivatives, 971–78
Diracdeltafunction, 975
Earth’s nutation, 973–74
exercises,976–78
impulsive force, 975–76
simple harmonic oscillator, 973
1168 Index
Laplacetransforms, 965–71
definition, 965
elementary functions, 965–66
exercises, 970–71
inverse transform, 967–68
partial fraction expansion, 968–69
step function, 969–70
linearity, 933–34
momentum representation, 955–61
exercises, 959–61
harmonic oscillator, 958–59
hydrogen atom, 957–58
other properties, 979–89
Bessel’s equation, 983–84
damped oscillator, 979–80
derivative of a transform, 982–83
electromagnetic waves,981–82
exercises, 985–89
integration of transforms, 984
limits of integration—unit step function, 985
RLCanalog, 980–81
substitution, 979
translation, 981
transfer functions, 961–64
exercises, 964
significance of /Phi1(t),963–64
integration, 904, seealsopath-dependent work
Euler–Maclaurin formula, 380–82
by parts of curl, 47
byparts of divergence, 40
by parts of gradient, 36
of powerseries, 364
simple pole on contour of, 468–69
of transforms, 984
of vectors, 54–60
exercises, 59–60
overview, 54–56
integration interval [a,b], 628–29
interpolating polynomials, 194
interpretation, seegeometrical interpretation of
gradient; physical interpretation of
divergence
interitem, 1111
invarianceofscalarproductunderrotations,15–17
invariants, electromagnetic,288
inverse Laplacetransform, 994–1003
Bromwich integral, 994–95
inversion via calculusof residues, 996
summary—inversion of Laplacetransform,
999–1000
velocityof electromagneticwaves in a
dispersive medium, 997–99
inverse matrix, 200
inverse operator, 934, 1019
inverse transform, 967–68inversion, 445–46
matrix, 184–87
Gauss–Jordan, 185–87
overview, 184–85
of PDE,949–50
of powerseries, 366
via calculusof residues, 996
irreducible representations, 245–46
irreducible tensors, 149–51
irregular (essential) singular point, 563
irregular sign changes, series with,341–42
irregular singularities, 572–73
irrotational, 45
isomorphic group, 243
isomorphism, 243–45
overview, 243
rotations, 244–45
isospin,SU(2),256–60
isospinI, 257
J
Jacobian, 107–8
parity transformation, 146
Jacobians for polar coordinates, 108–10
Jacobi–Anger expansion, 687
Jacobi identity, 248
Jacobi technique, 226
Jordan’s lemma, 468
Julia set,1087
K
kaon decay, 282–83
Kepler’s lawsof planetary motion, 116–17
kinematics and dynamics in Minkowski
space–time,280–83
Kirchhoff diffraction theory, 426
Klein–Gordon equation, 537
Korteweg–deVries equation, 542
Kroneckerdelta,10, 136–37
Kroneckerproduct, 181
Kronig–Kramers optical dispersion relations, 484,
485
Kummer’s equation, seeconfluent hypergeometric
equation
Kummer’s test, 330–32
L
ladderoperators, approach to orbital angular
momentum, 262–64
Lagrange’s equations, 1066
Lagrangian, 1053–54
Lagrangian equations, 1066–67
Lagrangian multipliers, 1060–65
cylindrical nuclearreactor, 1062
particlein a box, 1061–62
Index 1169
Laguerrefunctions, 837–48
associated Laguerre polynomials, 841–43
differential equation—Laguerrepolynomials,
837–41
hydrogen atom, 843–45
Laguerrepolynomials, 624
associated, 624–25
confluent hypergeometric representation, 866
generating function, 841–42
integral representation, 843
orthogonality, 843
recurrence relations, 842
Rodrigues’ representation, 767, 842
confluent hypergeometric representation, 866
differential equation, 837–41
generating function, 837–39
Gram–Schmidt construction, 647
orthogonality, 624, 647
recurrence relations, 840, 842
Rodrigues’ formula, 839
Schrödinger’s wave equation, 843
self-adjoint form, 624, 840
singularities, 564
Laplace,Mellin, andHankeltransforms, 933
Laplacefunction, 598
Laplace’sequation, 536, 1057–58
Bessel functions, 695
Legendre polynomials, 760, 761
solutions, 50–51, 96, 443, 452, 539, 559–60
Laplaceseries
expansion theorem, 790–91
gravity fields, 791
Laplacetransforms, 965–71
convolution theorem, 521, 1000, 1003, 1014
definition, 965
of derivatives, 971–78
Diracdelta function, 975
Earth’s nutation, 973–74
impulsive force, 975–76
simple harmonic oscillator, 973
elementary functions, 965–66
inverse transform, 967–68
partial fraction expansion, 968–69
step function, 969–70
table of transforms, 967–68, 979
translation, 1000
Laplacian
in Cartesian coordinates, 51–52, 554–55
in circularcylindrical coordinates, 119
development by minors, 167–68
in orthogonal coordinates, 316
of potentials, 50–51
spherical polarcoordinates, 126
as tensor derivative operator, 161–62
Laurentexpansion, 430–38, 466, 472analyticcontinuation, 432–34
exercises, 437–38
Schwarzreflection principle, 431–32
Taylorexpansion, 430–31
lawof cosines, 16–17, 118, 745
leadingcoefficients for ce 0,926–29
leadingcoefficients of se 1, 923–26
leastsquares, method of, 1119
Legendre duplication formula, derivation of,
522–23
Legendre equation, Maxwell’sequation, 779
self-adjoint form, 623, 625
Legendre functions, 741–816
addition theorem, 798
addition theorem for spherical harmonics,
797–802
derivation of addition theorem, 798–800
exercises,800–802
trigonometric identity, 797–98
alternatedefinitions of Legendre polynomials,
767–70
exercises,769–70
Rodrigues’ formula, 767
Schlaefliintegral, 768–69
associated,772
associatedLegendre functions, 771–86
associatedLegendre polynomials, 772–74
equation, 558, 771–72, 778–82, 788
Fourier transform, 770
Gram–Schmidt construction, 789
hypergeometric representation, 861
lowestassociatedLegendre polynomials, 774
magneticinduction field of a current loop,
778–82
orthogonality, 776–78
parity, 776
poles, 760, 782
recurrencerelations, 775
Rodrigues’ formula, 772–73
Schlaefliintegral, 768
second kind, 806–12
self-adjoint form, 771–72
specialvalues,774–75
generating function, 741–49
exercises,747–49
extension to ultraspherical polynomials, 747
Legendre polynomials, 742–44
linearelectricmultipoles, 744–45
physical basis—electrostatics,741–42
vector expansion, 745–47
integrals ofproducts of three spherical
harmonics, 803–6
application ofrecurrence relations, 804–5
exercises,805–6
orbital angularmomentum operators, 793–97
1170 Index
orbital angular momentum operators, 796–97
orthogonality, 756–67
Earth’s gravitational field, 758–59
electrostaticpotential ofa ring ofcharge,
761–62
exercises, 762–66
expansion of functions, Legendre series,
757–58
polarization of dielectric,764
ring of electriccharge,761–62
sphere in a uniform field, 759–61
recurrencerelations andspecialproperties,
749–56
differential equations, 751–52
exercises, 754–56
parity, 753
recurrence relations, 749–50
special values,752
sphere in uniform electricfield, 759–61
upper and lower bounds for Pn(cosθ),753–4
of second kind, 806–13
closed-form solutions, 810–12
exercises, 812–13
Qn(x)functions of the second kind, 809–10
series solutions of Legendre’s equation,
807–9
spherical harmonics, 786–93
azimuthal dependence—orthogonality, 787
exercises, 791–93
Laplaceseries, expansion theorem, 790–91
Laplaceseries—gravity fields, 791
polar angle dependence,788
spherical harmonics, 788–90
vector spherical harmonics, 813–16
Legendre polynomials, 644–46, 741, 742–44
associated
generating function, 773
recurrence relations, 775
generating function, 743
by Gram–Schmidt orthogonalization, 644–46
Laplace’sequation, 760, 761
orthogonality integral, 777
recurrencerelations, 749–50
Rodrigues’ formula, 761
Schlaefliintegral, 768
Legendre’s duplication formula, 503
Legendre’s equation, 625
differential form, 752
Legendre’s equation
self-adjoint form, 625, 771
Legendre’s equation
seriessolutions of, 807–9
Legendre series,333, 807–9
recurrencerelations, 807–8
Leibnizcriterion, 339–40formula for differentiating an integral, 590, 776
formula for differentiating aproduct, 771
Lerch’s theorem, 967
Levi-Civita symbol, 146–47
L’Hôpital’s rule,365
Lie groups and algebras, 243, 248, 264–66
limits of integration—unit step function, 985
limits to values of elliptic integrals, 374
linearelectricmultipoles, 744–45
linearequations
homogeneous, 165–66
inhomogeneous, 165–66
linear independence ofsolutions, 581–83
linearity, 933–34
linearly dependent solutions, 549
linearly independent solutions, 549
linear operator, 87, 176, 208, 535, 622, 650
differential operator, 42–43, 249, 261, 285, 304,
307, 554, 569, 629, 634, 664
integral operator, 768, 1019
linearoscillator
Green’sfunction, 668–69
linear oscillator equation, 1009–11
linear transformation law,152
line integrals, 55
Liouville’s theorem, 428
liquid drop model, 785
logistic map,1080–84
Lommel integrals, 697
Lorentz covarianceof Maxwell’sequations,
283–91
electromagneticinvariants, 288
exercises, 289–91
overview, 283–86
transformation of EandB,287–88
Lorentz–Fitzgerald contraction, 148
Lorentz gauge,52
Lorentz group, seehomogeneous Lorentzgroup
lowering operator, 263
lowest associatedLegendre polynomials, 774
Lyapunov exponents, 1085–86
M
Maclaurinexpansion, series computation, 512
Maclaurinintegral test, 327–30
Riemann Zetafunction, 329–30
Maclaurintheorem, 354–55
exponential function, 354–55
logarithm, 355
overview, 354
Madelung constant, 347
magnetic, 19, 46, 51, 66–67, 69, 74–76, 96, 100,
127–28, 144–45, 283–84, 288, 306, 311,
317, 447, 537, 638, 685, 703–5, 746,
778–82, 974
Index 1171
magnetic field constant ( Bfield), 44
magnetic fluxacross an oriented surface,306
magnetic induction field of current loop, 778–82
magnetic moments, 145
magnetic vector potentials, 44, 74–76, 127–28
Mandelbrot set, 1087–88
mapping
complex variables, 443–51
branch points and multivalent functions,
447–50
exercises, 450–51
inversion, 445–46
overview, 443
rotation, 444
translation, 443–44
conformal, 451–54
exercises, 453–54
Mathieuequation
angular, 872
modified, 872
radial, 872
Mathieufunctions, 869–79, 921
elliptical drum, 872–73
Fourier expansions of, 919–29
integral equations and Fourier seriesfor
Mathieu functions, 919–23
leading coefficientsfor ce 0, 926–29
leading coefficientsof se 1,923–26
general properties ofMathieu functions, 874
quantum pendulum, 873
radial Mathieu functions, 874–79
separation of variables inellipticalcoordinates,
870–71
matrices,165–239, seealsodeterminants;
diagonal matrices; orthogonal matrices
addition and subtraction, 178–79
adjoint, 208
angular momentum matrices, 253, 273
anticommuting sets,236
antihermitian, 221–23, 231
antisymmetric, 204
definition, 176
diagonalization, 215–26
direct product, 181–82
equality, 178
Euler angle rotation, 202–3
exercises, 187–95
Gauss–Jordan matrix inversion technique, 185
Hermitian and unitary, 208–15
exercises, 212–15
overview, 208–9
Pauli and Dirac,209–12
inversion of, 184–87
Gauss–Jordan, 185–87
overview, 184–85ladder operators, 263
matrix multiplication, 176, 179–81
moment of intertia, 215–16, 220
multiplication, 179–80
directproduct, 181–82
inner product, 179–81
byscalar,179
normal, 231–39
exercises,236–39
ill-conditioned systems, 234–35
normal modes of vibration, 233–34
overview, 231–32
null matrix, 178
orthogonal matrix, 201, 206, 209
overview, 176–78
product theorem, 181
quaternions, 204, 212
rank, 178
relation to tensor, 206
representation, 177, 184, 205, 208, 212
self-adjoint, 209
similarity transformation, 205
skewsymmetric, 204
symmetric, 204
traces,139, 183–84
transposition, 177
unitary, 209
vector transformation law,198
Maxwell’sequations, 51, 284, 316–20
derivation of waveequations, 52
dual transformation, 290
Gauss’ law,52–53
Legendre equation, 779
Lorentzcovariance of, 283–91
electromagneticinvariants, 288
exercises,289–91
overview, 283–86
transformation of EandB, 287–88
Oersted’s law,52–53
mean valuetheorem, 353
Mellin transforms, 897, 933
meromorphic functions, 439, 461
integral of, 466
pole expansion of, 461
metric, curvilinearcoordinates, 105
metric tensor, 151–54
Christoffel symbols asderivatives of, 155–56
minimal substitution, 76
Minkowski space,136, 278–79
Minkowski space–time, seekinematics and
dynamics in Minkowski space–time
minor, 168
minors, Laplaciandevelopment by, 167–68
missing dependent variables, 1042–43
Mittag-Lefflertheorem, 461
1172 Index
mixed tensor, 136, 139, 154
modes of vibration, normal, 233–34
modified Bessel functions, 713–19
asymptotic expansion, 711, 719
Fourier transform, 716
generating function, 709
integral representation, 720–23
Laplacetransform, 933
recurrencerelations, 714–16
seriesform, 714
modified Helmholtz operator, 598
modified Mathieuequation, 872
modulus, 406
moment of inertia matrix, 215–16
momentum, seeangular momentum; orbital
angular momentum
momentum representation, 955–61
harmonic oscillator, 958–59
hydrogen atom, 957–58
Schrödinger wave equation, 957–58
momentum representation in quantum mechanics,
1006–7
monopole, 745–46
Morera’s theorem, 427–28
motion
arealaw for planetary, 116–19
equations of, 142
moving particle
Cartesian coordinates, 1054
circularcylindrical coordinates, 1054–55
multiplet, 245
multiplication
ofmatrices
direct product, 181–82
of vectors, 182
inner product, 179–81
by scalar,179
multiply connected regions, 423–24
multipole expansion, electrostatic,599
multivalent functions, 447–50
multivalued function, 409
mutually exclusive, 1109–10
N
Navier–Stokes equations, 119
Neumann boundary conditions, 543
Neumann functions, 511
asymptotic form, 602–3
Besselfunctions of second kind, 699–707
coaxial waveguides, 703–4
definition and series form, 699–700
other forms, 701
recurrence relations, 702
Wronskian formulas, 702–3
Fourier transform, 602Hankelfunction definition, 707–8
integral form, 701
recurrence relations, 702
spherical, 507, 727
Wronskian formulas, 702–3
Neumann functions, integral form, 701
Neumann problem, 617
Neumann series, 1018–19
Neumannseries, separable(degenerate) kernels,
1018–29
numerical solution, 1023–25
separablekernel, 1021–22
Neumann series solution, 1020–21
neutron diffusion theory, 369, 951
Newton’s second laws,130, 233, 370, 973, 1054
node, spiral, 1097, 1103–4
non-Cartesian tensors, 140
nonessential (regular) singular point, 563
nonhomogeneous equation—Green’s function,
592–610
circularcylindricalcoordinateexpansion, 601–2
form of Green’sfunctions, 596–98
Legendre polynomial addition theorem,
600–601
spherical polar coordinate expansion, 598–600
symmetry of Green’s function, 595–96
nonhomogeneous integral equation, 1032–34
nonlinear differential equations (NDEs),1088–89
autonomous differential equations, 1091–93
Bernoulli and Riccatiequations, 1089–90
bifurcations in dynamical systems, 1103–4
centeror cycle,1100–1101
chaos in dynamical systems, 1105–6
dissipation in dynamical systems, 1102–3
fixedandmovable singularities, special
solutions, 1090
localandglobal behavior in higher dimensions,
1093–94
routes to chaos in dynamical systems, 1106–7
saddle point, 1095–97
spiral fixed point, 1098–1100
stablesink, 1095
nonlinear methods and chaos, 1079–1108
introduction, 1079–80
logistic map, 1080–84
exercises,1084
nonlinear differential equations (NDEs),
1088–89
autonomous differential equations, 1091–93
Bernoulli and Riccatiequations, 1089–90
bifurcations in dynamical systems, 1103–4
centerorcycle,1100–1101
chaosin dynamical systems, 1105–6
dissipation in dynamical systems, 1102–3
exercises,1089, 1102, 1106, 1107
Index 1173
fixed andmovable singularities, special
solutions, 1090
local andglobal behavior inhigher
dimensions, 1093–94
routestochaosindynamicalsystems,1106–7
saddle point, 1095–97
spiral fixed point, 1098–1100
stable sink, 1095
sensitivity to initial conditions and parameters,
1085–88
exercises, 1088
fractals, 1086–88
Lyapunov exponents, 1085–86
nonuniform convergence, 348–49
normalization, 695
normal matrices,231–39
exercises, 236–39
ill-conditioned, 234–35
normal modes of vibration, 233–34
overview, 231–32
null matrix, 178
number operator, 823
numbers, Bernoulli, seeBernoulli numbers
numerical solution, 1023–25
nutation, 973–74
O
Oersted’s law,52, 66–68
Olbers’ paradox, 337
operators, differential vector, seedifferential
vector operators
optical dispersion, 484–85
optical pathnearevent horizon of blackhole,
1041–42
orbital angular momentum, 251, 261–66
exercises, 266
ladder operator approach, 262–64
Liegroups and algebras, 264–66
Liegroups and operators, order of, 264–65
operators, 793–97
overview, 261
rotation of, 251
order 2 branch points, 440–42
order parameter, 1086
ordinary differential equations (ODEs), linear
first-order, 547–50
orientedsurface, magnetic fluxacross, 306
orthogonal coordinates
Laplacianin, 316
inR3, 103–10
exercises, 109–10
Jacobians for polar coordinates, 108
overview, 103–8
orthogonal eigenfunctions, 636, 637, 1030–32orthogonal functions, representation ofDiracdelta
function by, 88–89
orthogonal groups, 243
orthogonal group SO(3), 254, 256
orthogonality, 694–99, 731, 776–78, 821–22,
854–55
Bessel series,695
continuum form, 696
curvilinear coordinates, 104
Earth’s gravitational field, 758–59
electrostaticpotential in hollow cylinder,
695–96
electrostaticpotential of ring ofcharge, 761–62
expansionoffunctions,Legendreseries,757–58
Fourier series, 636–37
Fourier series: Hilbert–Schmidt integral
equations, 919–29
normalization, 695
over discrete points, 914–15
sphere in a uniform field, 759–61
Sturm–Liouville differential equations, 1031
of vectors, 14
orthogonality condition, 10, 198
orthogonality integral
Hermitepolynomials, 821
Legendre polynomials, 777
spherical harmonics, 788
orthogonality relations, spherical harmonics, 814
orthogonalization, Gram–Schmidt, 173, 642–7
orthogonal matrices, 195–208, 209
applications to vectors, 197–98
direction cosines, 196–97
Eulerangles, 202–3
exercises, 206–8
inverse, 200
overview, 195–96
relation to tensors, 206
symmetry properties, 203–5
transpose matrix, ˜A,200–202
two-dimensional conditions for, 199–200
orthonormal functions, 642–44
polynomials, 646–47
vectors, 174
oscillator
damping, 979–80, 991–93, 1101
driven, 457, 565, 991–93
forced classical,457–60
harmonic, 822–27
integral equation for, 991–93
Laplacetransform solution, 991–93
linear, 668–69
momentum spacewavefunction, 569
self-adjoint equation, 679
series solution of differential equation, 568–69
1174 Index
singularities in harmonic oscillator differential
equation, 564
oscillatory series,322
overshoot, calculation of, 912–13
P
parabolic PDEs,538
parallelogram addition law,2–3
paralleltransport, 157–60
parity, 753, 776, 939
Besselfunctions, 687, 735
Chebyshev functions, 569
differential operator, 569
Fourier cosine, sine transforms, 939
Hermitefunctions, 569
Legendre functions, 569
Legendre functions, associated,776
Legendre functions, second kind, 812
spherical harmonics, 776, 791
vector spherical harmonics, 814
parity transformation, 143
Parseval relation, 485–86, 952–53
partial differential equations (PDEs), 535–43
bicharacteristicsof, 538
boundary conditions, 542–43
characteristicsof, 538–41
classesof, 538–41
elliptic,538
examples of, 536–38
harmonic functions, 539
hyperbolic, 538
introduction, 535–36
inversion of, 949–50
nonlinear, 541–42
parabolic,538
partial fraction expansion, 968–69
partial sum approximation, 390
particle
in a box, 1061–62
in asphere, 731–32
particlemotion
Cartesian coordinates, 1054
circularcylindrical coordinates, 1054–55
quantum mechanical,827
in rectangular box, 1061–62
in right circularcylinder, 1062
in sphere, 731–32
path-dependent work, 56–60
integral definitions ofgradient, divergence, and
curl, 58–59
overview, 56
surfaceintegrals, 56–57
volume integrals, 57–58
Pauli matrices, 209–12
pendulums, period of simple, 370–71periodic functions, 888–90
boundary conditions, 636, 883–5
permutations and combinations, counting of,
1114–15
phase of a complex function, 409
phase of a complex number, 409
phase space,88, 1080, 1091
physical interpretation ofdivergence, 40–42
pi,π,219, 379, 396–400, 586
Leibniz formula, 590, 776, 886
Wallisformula, 399
pion photoproduction threshold, 282–83
pi-sine relation, verification of, 522
planetary motion, arealaw for, 116–19
Pochhammer symbol, 859–60, 864
Poincaré section, 1088–89, 1104–6
point and spacegroups, crystallographic, 299–300
point source equation, 662
Poisson distribution, 1130–33
Poisson’s equation, 536
and Gauss’ law,81–83
Green’sfunction, 669–70
polar angledependence, 788
polar coordinates, Jacobians for, 108–10, seealso
spherical polar coordinates
polar vectors, 143–44
pole expansion of meromorphic functions, 461
poles, 439
simple, on contour of integration, 468–69
polygamma functions, 511–12
Catalan’sconstant, 513
polynomials, Bernoulli, 379–80
potential energy, 309
potentials, 68–79, seealsothermodynamics
gradient of, 34
force as,36
Laplacianof, 50–51
overview, 68
scalar,68–72
centrifugal, 72
gravitational, 72
overview, 68–71
vector, 73–79
of constant Bfield,44
exercises,77–79
magnetic,74–76
potential theory
conservative force, 69
electrostaticpotential, 593
scalarpotential, 70
vector potential, 73, 311
power series, 363–70
continuity, 364
convergence, 363
uniform and absolute, 363
Index 1175
differentiation and integration, 364
exercises, 366–70
inversion of, 366
overview, 363
uniqueness theorem, 364–65
L’Hôpital’s rule, 365
prime numbers, 379, 382–83, 897
prime number theorem, asymptotic, 897–8
primitive, 898
principal axis, 216
principal value, 409
probability, 1109–51
binomial distribution, 1128–30
exercises, 1129
repeatedtosses of dice,1128–30
definitions, simple properties, 1109–15
conditional probability, 1112
counting of permutations and combinations,
1114–15
exercises, 1115
probability for AorB,1110–11
scholastic aptitude tests,1112–14
Gauss’ normal distribution, 1134–38
exercises, 1137–38
Poisson distribution, 1130–33
exercises, 1133
random variables, 1116–28
continuous random variable: hydrogen atom,
1117–19
discrete random variable, 1116–17
exercises, 1127–28
repeated draws of cards,1123–26
standarddeviationofmeasurements,1119–23
sum, product, and ratio of random variables,
1126–27
statistics, 1138
χ2distribution, 1143–45
confidence interval, 1149
error propagation, 1138–40
exercises, 1150
fitting curves to data,1140–43
studenttdistribution, 1146–49
probability for AorB,1110–11
product convergence theorem, 344
products, seealsocross product; direct product;
dot products; scalars
expansion of entirefunctions, 462–63
of infinite series,396–401
convergence of infinite, 397–98
exercises, 399–401
overview, 396–97
sine, cosine, and gamma functions, 398–99
product theorem, 181
projection operators, 644
projections, of vectors, 4, 12pseudoscalars, 146
pseudotensors, 142–51
dual tensors, 147–48
exercises, 149–51
irreducible tensors, 149–51
Levi-Civita symbol, 146–47
overview, 142–46
pseudovectors, 146
pullbacks, 309–13
Gauss’ theorem, differential form, 312–13
overview, 309–10
Stokes’ theorem, 310–11
differential form, 311
Q
QCD(quantum chromodynamics), 259
Qn(x)functions of thesecond kind, 809–10
quadrupole, 149, 745–47, 805
quantization, 624–25, 1054
quantum chromodynamics (QCD), 259
quantum mechanicalscattering, 469–71
quantum mechanicalsimple harmonic oscillator,
822–27
quantum mechanics, sum rules, 484
angular momentum, 189–92, 251–6, 261–4,
266–70
configuration spacerepresentation, 957–58
expectation values, 263
hydrogen atom, 245–46
momentum representation, 1006–7
Schrödinger representation, 1006–7
quantum pendulum, 873
quasiperiodic, 1100–1101, 1107
quaternions, 189, 204, 212
quotient rule, 141–42
equations of motion and field equations, 142
exercises, 142
overview, 141–42
R
R3,orthogonal coordinates in, seeorthogonal
coordinates
Raabe’stest, 332
radial Mathieuequation, 872
radial Mathieufunctions, 874–79
radioactive decay,553, 977–78, 1129, 1130–31,
1133
raising operator, 263
random variables, 1116–28
continuous random variable: hydrogen atom,
1117–19
discreterandom variable,1116–17
standard deviation of measurements, 1119–23
sum, product, and ratio of random variables,
1126–27
1176 Index
rank, 264
of matrices, 178
of tensor, 133
rapidity, 280
rational approximations, 345
ratio test,Cauchy, d’Alembert, 326
Rayleigh formulas, 730
Rayleigh–Ritz variational technique, 1072–76
ground stateeigenfunction, 1073
vibrating string, 1074
realpart, 407
rearrangement of double series, 345–48
reciprocal,916
reciprocal lattice,27
reciprocity principle, 665
rectifier,full-wave, 893–94
recurrence relations, 567, 714–16, 749–50, 775,
818–19
application of, 804–5
associatedLegendre polynomials, 775
Bernoulli numbers, 377
Besselfunctions, 677
Besselfunctions, spherical, 730
Chebyshev polynomials, 850
confluent hypergeometric functions, 866
derivatives, 852–53
exponential integral, 529
factorialfunction, gamma, 504, 506, 529, 995
Hankelfunctions, 708
Hermitepolynomials, 818–19
hypergeometric functions, 861
Laguerrefunctions, associated,842
Legendre polynomials, 749–50
Legendre series,807–8
modified Bessel functions, 714–16
Neumann functions, 702
polygamma functions, 512
and specialproperties, 749–56
differential equations, 751–52
parity, 753
recurrence relations, 749–50
special values,752
upper and lower bounds for Pn(cosθ),
753–54
spherical Besselfunctions, 730
reducible representations, 245–46
reflection principle, Schwarz,431–32
regression coefficient,1140–41
regular (nonessential) singular point, 563
regular functions (holomorphic or analytic
functions), 415
regular singularities, 572–73
relations, seedispersion relations
relativistic energy, 356–57
repeateddraws ofcards, 1123–26repellant, 1082
repellor, 1092
representation
fundamental, 259–60, 274–75
irreducible, 246, 263, 265, 268, 274, 276, 293,
298
reducible, 246
residue theorem, 455–56, 472; see alsocalculus of
residues
resonant cavity, 682–85
Riccatiequation, 1089–90
Riemann–Christoffel curvature tensor, 138
Riemann integral, 60–61
Riemann manifold, 313–14
Riemann’s theorem, 344
Riemann surface,448–49
Riemann Zetafunction, 329–30, 334–39
and Bernoulli numbers, 382–84
Fourier series evaluation, 329
tableof values,382
infinite series, 894–98
Riesz’ theorem, 177
RLC analog, 980–81
RL circuit,549–50
Rodrigues’ formula, 767
Laguerrepolynomials, 839
associated,767, 842
Rodrigues representation, 820–21
Hermitepolynomials, 820–21
root diagram, 264
root test,Cauchy, 326
rotations, 176, 444
of coordinate axes, 7–12
exercises,12
vectors andvector space,11–12
of coordinates, 199
offunctionsandorbitalangularmomentum,251
groupsSO(2)andSO(3), 250
invariance of scalar product under, 15–17
isomorphic and homomorphic, 244–45
Rouché’s theorem, 463
routes to chaos in dynamical systems, 1106–7
S
saddle points (steepest descent method), 489–97,
1092, 1095–97, 1103
analyticlandscape, 489–90
asymptotic forms
offactorial function Ŵ, 494–95
ofHankelfunction H(1)
ν(s),493–94
exercises, 496–97
factorialfunction Ŵ(z),494–95
Hankelfunction H(1)
ν,493–94
overview, 489
sample space,1109–10
Index 1177
sawtooth wave,885–86
scalarpotential, 33–34, 70, 536
scalarquantities, 1
scalars,7, 133
multiplication of matricesby, 179
potentials, 68–72
centrifugal, 72
gravitational, 72
overview, 68–71
products, 12–17
exercises, 17
invariance of under rotations, 15–17
overview, 12–15
triple, 25–27
scattering, quantum mechanical,469–71
Schlaefliintegral, 709, 768–69
Legendre polynomials, 768
Schmidt orthogonalonization, seeGram–Schmidt
scholastic aptitude tests, 1112–14
Schrödinger’s wave equation, 624
degeneracy of, 638
hydrogen atom, 843
Schrödinger wave equation, 76, 537, 1069–70
momentum space representation, 957–58
variational derivations, 1069–70
Schur’s lemma,265
Schwarzinequality, 652–54
generalized, 661
Schwarzreflection principle, 431–32
second-rank tensors, 135–36
section,seePoincaré section
secularequation, 218
selectionrules, 265
self-adjoint eigenvalue equations, 595
self-adjoint matrices, 209
self-adjoint ODEs,622–34
boundary conditions, 627–28
deuteron, 626–27
eigenfunctions, eigenvalues, 624–25
Hermitian operators, 629
Hermitian operators in quantum mechanics,630
integration interval [a,b], 628–29
Legendre’s equation, 625
self-adjoint operator, 623, 630
Sturm–Liouville theory, differential equations,
622–30
semiconvergent series,391
sensitivity to initial conditions and parameters,
1085–88
fractals, 1086–88
Lyapunov exponents, 1085–86
separablekernel, 1021–22
separablevariables, 544–45
separationof variables inellipticalcoordinates,
870–71series,seeinfinite series
series approach, 570–71
Bessel’s equation, limitations of, 570–71
Chebyshev, 338, 569
Hermite,567
hypergeometric, 569, 859, 862
incomplete beta,523
Laguerre,569
Legendre,333, 507, 569, 757–62
shifted polynomials, 836
Chebyshev, 850
Legendre,629
ultraspherical, 338
series form, 714
of second solution, 583–85
series solutions—Frobenius’ method, 565–78
expansion about x0, 569
Fuchs’ theorem, 573
limitations of seriesapproach—Bessel’s
equation, 570–71
regular and irregular singularities, 572–73
symmetry of solutions, 569
sign changes,series with alternating, irregular,
341–42
similarity transformation, 205
simple harmonic oscillator, 973
simple pendulum, 1067–68
simple pole on contour of integration, 468–69
sine
confluent hypergeometric representation, 393
functions of infinite products, 398–99
integrals in asymptotic series, 392–93
sine transform, 939–40
singularities, 438–43
branch points, 440–42
of order 2, 440–42
exercises, 442–43
fixed, 1090
Laurentseries, 439
movable, 1090
overview, 438
poles, 439
singular points, 562–65
essential(irregular), 563
irregular (essential), 563
nonessential (regular), 563
regular (nonessential), 563
sink, 1092–96, 1101
sinx
infinite product representation, 378, 392–93,
398, 439, 468–69, 583, 679–80, 728–29,
730, 885, 889, 892, 911, 950, 1021,
1097–98
power series,358, 406, 410
skewsymmetric matrices, 204
1178 Index
Slaterdeterminant, 274
small oscillations, 370–71
SO(2)rotation groups, 250
SO(3)
Clebsch–Gordan coefficients, 267–70
homomorphism, 252–56
rotation groups, 250
soap film, 1045–46
soap film—minimum area,1046–49
solenoidal, 42
soliton solutions, 542
space,vector, seevectors
spaceand point groups, crystallographic, 299–300
space–time,Minkowski, seekinematics and
dynamics in Minkowski space–time
specialunitary groups
SU(2), Pauli spin matrices,189, 203–4,
250–66, 267–70, 274–76
SU(3), Gell-Mannmatrices, 212, 256–60,
265–66, 274–76
SU(n), Young tableaux,274–76
specialunitary group SU(2), 252
specialvalues, 774–75
spectraldecomposition, 219, 225, 635
sphere in a uniform field, 759–61
spheres, total charge inside, 88
spherical Bessel functions, 725–39
asymptotic values,729
definitions, 726–29
limiting values, 729–30
recurrencerelations, 730
spherical components, 271
spherical coordinates, Helmholtz equation, 725
spherical harmonics, 264, 786–93
addition theorem for, 797–802
derivation of addition theorem, 798–800
trigonometric identity, 797–98
angular momentum operators, 793
azimuthal dependence—orthogonality, 787
Condon–Shortley phase conventions, 270
integrals of, 804
ladder operators, 796–97
Laplaceseries,expansion theorem, 790–91
orthogonality integral, 788
orthogonality relations, 814
polarangle dependence,788
spherical harmonics, 788–90
vector spherical harmonics, 813–16
spherical polar coordinates, 123–33, 557–60
exercises,128–33
expansion, 598–600
magneticvector potential, 127–28
∇,∇·,∇×for centralforce, 127
overview, 123–26
unit vectors, 123spherical symmetry, 616
spherical tensor operator, 271
spherical tensors, 271–74
spherical waves,Bessel functions, 730
spinors, 138–39
exercises, 138–39
overview, 138
spinor wave functions, 212
spiral fixed point, 1098–1100
spiral node, 1097, 1103
spiral repellor, 1097, 1103
square integrable, 485, 487, 653, 658, 882
squares ofseries, divergent, 344
square wave,911–12
square wave—high frequencies, 892–93
stable sink, 1095
standard deviation, 1119, 1140
standard deviation of measurements, 1119–23
stark effect,576, 847
statisticalhypothesis, 1138
statistics, 1138
χ2distribution, 1143–45
confidence interval, 1149
error propagation, 1138–40
fitting curves to data,1140–43
studenttdistribution, 1146–49
steepestdescent, method of, 489–96
factorialfunction, 494–95
Hankelfunctions, 493–94
modified Bessel functions, 720
step function, 969–70
Stirling’s expansion, 495
Stirling’s series,516–20
derivation from Euler–Maclaurin integration
formula, 517
Stokes’ theorem, 64–68
alternateforms of, 66–68
Oersted’s and Faraday’s laws, 66–68
overview, 66
on differential forms, 313–14
Riemann manifold, 313–14
exercises, 67–68
overview, 64–65
proof, 420–21
pullbacks, 310–11
straight line, 1044–45
strange attractor, 1085, 1086
string, Lagrangian of avibrational, 1058, 1071
structure constants, 248, 251, 259
studenttdistribution, 1146–49
Sturm–Liouville theory, 885
Sturm–Liouville theory—orthogonal functions,
621–74
completeness of egenfunctions, 649–61
Bessel’s inequality, 651–52
Index 1179
expansion coefficients,658
Schwarzinequality, 652–54
summary—vectorspaces,completeness,
654–58
Sturm–Liouville theory—orthogonal functions
completeness of Eigenfunction, 659–61
Gram–Schmidt orthogonalization, 642–49
exercises, 647–49
Legendre polynomials by Gram–Schmidt
orthogonalization, 644–46
Green’s function—eigenfunction expansion,
662–74
eigenfunction, eigenvalue equation, 667–68
exercises, 670–74
Green’sfunctionandtheDiracdeltafunction,
669–70
Green’s function integral—differential
equation, 665–67
Green’s functions—one-dimensional,
663–65
linearocillator, 668–69
Hermitian operators, 634–42
degeneracy, 638
exercises, 639–42
expansion in orthogonal
eigenfunctions—square wave,637
Fourier series—orthogonality, 636–37
orthogonal eigenfunctions, 636
real eigenvalues, 634–35
self-adjoint ODEs,622–34
boundary conditions, 627–28
deuteron, 626–27
eigenfunctions, eigenvalues, 624–25
exercises, 631–34
Hermitian operators, 629
Hermitian operators in quantum mechanics,
630
integration interval [a,b], 628–29
Legendre’s equation, 625
SU(2)
Clebsch–Gordan coefficients, 267–70
isospin and SU(3)flavor symmetry, 256–60
andSO(3)homomorphism, 252–56
SU(3)flavor symmetry, 256–60
subgroups and cosets,293–94
substitution, 979
subtraction
of matrices, 178–79
of series, 324–25
of sets, 1111
of tensors, 136
sum, product, and ratio of random variables,
1126–27
summation convention, 136–37, 139
summation of series, 910sum rules, 484
SU(n), young tableaux for, 274–78
superposition principle for homogenous ODEs,
PDEs,536
surface integrals, 56–57
symmetric matrices, 204
symmetric tensor, 137
symmetrization of kernels, 1029–30
symmetry, 889
axes
threefold, 296–99
twofold, 294–96
cylindrical, 617
properties of orthogonal matrices,203–5
relations, 484
of solutions, 569
spherical, 616
SU(3)flavor, 256–60
of tensors, 137
T
tableauxfor SU(n), Young, 274–78
Taylor’s expansion, 352–63, 430–31
binomial theorem, 356–57
relativistic energy, 356–57
exercises, 358–63
Maclaurintheorem, 354–55
exponential function, 354–55
logarithm, 355
overview, 354
multiple variables,358
overview, 352–54
tensor analysis
contravariant tensor, 135–36, 153
contravariant vector, 134–35, 139, 152–54, 158
covariant tensor, 135–36, 158
covariant vector, 134–35, 139–40, 152–54, 156,
158
definition, 133–36
displacement, 158
isotropic tensor, 137
non-Cartesian tensors, 140
paralleltransport, 158
scalarquantity, 8, 15, 57, 134, 149, 179
spherical components, 271
spherical tensor operator, 271
symmetry–asymmetry, 137
tensor density, seepseudotensors
tensor derivative operators
curl, 162–63
divergence, 160–61
exercises, 162–63
Laplacian,161–62
overview, 160
1180 Index
tensors,seealsoderivative operators, tensor;
direct product; general tensors;
pseudotensors; quotient rule; spinors
generaltensors, 151–60, seealsoChristoffel
symbols
covariant derivative, 156
exercises, 158–60
geodesics and paralleltransport, 157–60
metric tensor, 151–54
overview, 151
relation to orthogonal matrices, 206
spherical,271–74
vector analysis in, 133–63
addition and subtraction of, 136
contraction, 139
overview, 133–35
second-rank, 135–36
summation convention, 136–37
symmetry–antisymmetry, 137
thermodynamics, 72–79
exactdifferentials, 72–76
overview, 72–73
vector potential, 73–79
exercises, 77–79
magnetic, 74–76
Thomas precession, 280
threefold Hermite formula, 827–28
threefold symmetry axis,296–99
time-dependent diffusion equation, 536
time-independent diffusion equation, 536
Titchmarsh theorem, 487
trace,139
traceformula, 224–25
Gutzwiller’s,898
tracesof matrices,183–84
trajectory, 38, 1088, 1091–92, 1097, 1099, 1100,
1103, 1105
transfer function, 962
transfer functions, 961–64
significanceof /Phi1(t), 963–64
transform, derivative of, 982–83; seealsocosines;
exponential; Fourier; Fourier–Bessel;
Hankel;Laplace;Mellin; sine
transformation law, 134
transformation of differential equation into
integral equation, 1008–9
transformation of EandB, Lorentz,287–88
translation, 443–44, 981
transport, parallel,157–60
transpose matrix, ˜A,200–202
transposition, 177
triangle inequalities, 406
triangle rule, 268–69
trigonometric form, 853–54
trigonometric identity, 797–98triple scalar products, 25–27, 165–66
triple vector products, 27–29, 46
BAC–CABrule, 28, 46, 51
exercises, 27–32
overview, 27
Tschebyscheff, seeChebyshev
two-dimensional conditions for orthogonal
matrices,199–200
twofold symmetry axis,294–96
U
ultraspherical polynomials, extension to, 747
equation, 853, 861
self-adjoint form, 854
uncertainty principle in quantum theory, 941
uniform convergence, 348–49, 363
union of sets,1111
uniqueness theorem, 364–65
descending power series,781
inverse operator, 181
Laurentexpansion, 433
of power series, 364–65
uniqueness theorem, L’Hôpital’s rule, 365
unitary groups, 243
unitary matrices, seealsoHermitian matrices
algebra,210–12
Heaviside,93, 981, 985, 996
ring, 180, 187
unit element of group, 293
unit step function, 191
vectorspace,208
unit vectors, 123
Cartesian coordinates, 5
circularcylindrical, 116
orthogonality relation, 651
spherical polar, 201–2
spherical polarcoordinates, 126
upper and lower bounds for Pn(cosθ),753–4
V
values, limiting, of elliptic integrals, 374
variables, seealsocomplex variables
dependent, 1052–58
Hamilton’s Principle, 1053–54
Laplace’sequation, 1057–58
moving particle—Cartesian coordinates,
1054
moving particle—circular cylindrical
coordinates, 1054–55
dependent and independent, 1038–44, 1058–59
alternateforms of Euler equations, 1042
conceptof variation, 1038–41
missing dependent variables,1042–43
opticalpath nearevent horizon ofa black
hole, 1041–42
Index 1181
multiple, of Taylor’s expansion, 358
separation of, 554–62
variance,1120
variation, concept of, 1038–41
variation of the constant, 548
variations, seealsocalculus ofvariations
variation with constraints, 1065–72
Lagrangian equations, 1066–67
Schrödinger wave equation, 1069–70
simple pendulum, 1067–68
sliding off a log, 1068–69
vectoranalysis
parallelogram addition law,2–3
reciprocal lattice,27
rotation of coordinates, 205
transformation law, 134
vector definition, 8–9
representation of, 9
vector expansion, 745–47
vectorfield, 7
vectorintegrals
line integrals, 55
surface integrals, 56–57
volume integrals, 57
vector potential, 73, 311
vector product, 5, 11, 18–22, 25–29, 43–44, 46,
147, 272
vectorquantities, 1
vectors, 1–101, seealsocurved coordinates and
vectors; divergence, ∇;g r ad i en t ,∇;
integration; potentials; rotations; Stokes’
theorem; tensors
applications of orthogonal matrices to, 197–98
components, 8, 11, 153, 654
contravariant, 134–35, 139, 152–54, 158
covariant, 134–35, 139, 152–54, 156, 158
cross product of, 315
curl,∇×,43–49
of centralforce field, 44–46
exercises, 47–49
gradient of dot product, 46
integration by parts of, 47
overview, 43
potential ofconstant Bfield,44
definitions and elementary approach, 1–7
exercises, 6–7
differential vector operators, 110–14
curl, 112–13
divergence, 111–12
exercises, 113–14
gradient, 110
overview, 110
Dirac deltafunction, 83–95
exercises, 91–95
integral representations for, 90–95overview, 83–87
phasespace,88
representation by orthogonal functions,
88–89
totalcharge inside sphere, 88
direction, 5
direct product of, 182
elementary approach to, 1–7
elementary approach to, exercises, 6–7
Gauss’ law,79–83
exercises,82–83
Poisson’s equation, 81–83
Gauss’ theorem, 60–64
alternateforms of, 62
exercises,62–64
Green’stheorem, 61–62
overview, 60–61
by Gram–Schmidt orthogonalization, 174–76
Helmholtz’s theorem, 95–101
exercises,100–101
overview, 95–96
irrotational, 45–46
linear dependenceof, 172–73
normal, 15, 56
or cross product, 18–22
exercises,22–25
orthogonal, 108, 173
overview, 1
potentials, magnetic, 127–28
scalaror dot product, 12–17
exercises,17
invariance of under rotations, 15–16
serve as abasis, 5
space,11–12
exercises,12
successive applications of ∇, 49–54
electromagneticwaveequation, 51–53
exercises,53–54
Laplacianof potential, 50–51
overview, 49–50
triangle law of addition, 1–2
triple product of, 27–29
exercises,27–32
triple scalarproduct, 25–27
vectorspace,linearspace,7, 11–12, 173, 177,
208, 219, 247–48, 314, 638, 644, 650,
654–55, 657–58, 884
vector spaces, completeness, 654–58
vector spherical harmonics, 813–16
vector transformation law, 21
velocity ofelectromagnetic wavesin a dispersive
medium, 997–99
vibrating string, 1074
vibration, normal modes of, 233–34
vierergruppe, 292–93, 296
1182 Index
Volterra equation, 1005–6
volume integrals, 57–58
von Staudt–Clausen theorem, 379
W
Wallis’ formula, 399
wavediffusion equation (Helmholtz diffusion
equation), 536, 537
wave equation, 947–48
anomalous dispersion, 999
derivation from Maxwell’sequation, 52
Fourier transform solution, 947–48
Laplacetransform solution, 948
waveequation, electromagnetic, 51–53
wave functions, 624
spinor, 212
waveguides, coaxial,Bessel functions, 703–4
Weierstrass infinite-product form of Ŵ(z),
499–500
Weierstrass Mtest, 349–50
weight diagram, 257, 258, 260, 265
weight vectors, 265
Whittaker functions, 866, 919
Wigner–Eckart theorem, 273
WKBexpansion, 394work, potential, 34, 70–72
Wronskian, 550
Wronskian determinant, 579–80
Wronskian formulas, 702–3
absenceof third solution, 587, 589
Bessel functions, 702–4
Bessel functions, spherical, 601, 723
Chebyshev functions, 582–83, 591
confluent hypergeometric functions, 868
Green’sfunction, construction of, 592, 703
linear independence of functions, 579–80, 665,
703
second solution ofdifferential equations,
583–87
solutions of self-adjoint differential equation,
702
Y
Young tableaux for SU(n),274–77
Z
zero-point energy, 732, 826
zeros, Besselfunction, 682
zetafunction, seeRiemannZetafunction