Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books / PDF Originals

Arfken,weber,Mathematical_Methods_for_Physicists_6_Ed

PDF · 1196 pages · 7.8 MB
Open PDF file

A published textbook by George B. Arfken and Hans J. Weber, not Phil's own work, kept in a folder of downloaded math methods books. The contents list covers vector analysis, tensors, matrices, group theory, infinite series, complex variables, the gamma function, differential equations, Sturm-Liouville theory, Bessel, Legendre and other special functions, Fourier series, integral transforms, integral equations and calculus of variations. Only the front matter and table of contents were read.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
=MATHEMATICAL Methods forPhysicists ARFKEN &WEBER MATHEMATICAL METHODS FOR PHYSICISTS SIXTH EDITION GeorgeB. Arfken MiamiUniversity Oxford,OH Hans J.Weber Universityof Virginia Charlottesville,VA Amsterdam Boston Heidelberg London NewYork Oxford Paris SanDiego SanFrancisco Singapore Sydney Tokyo This page intentionally left blank MATHEMATICAL METHODS FOR PHYSICISTS SIXTH EDITION This page intentionally left blank Acquisitions Editor TomSinger Project Manager Simon Crump MarketingManager Linda Beattie Cover Design EricDeCicco Composition VTEXTypesettingServices Cover Printer Phoenix Color Interior Printer TheMaple–VailBook Manufacturing Group ElsevierAcademicPress 30 Corporate Drive,Suite 400, Burlington, MA01803, USA 525 B Street,Suite 1900, San Diego, California 92101-4495, USA 84 Theobald’s Road, London WC1X 8RR, UK This book is printed on acid-freepaper. /circlecopyrt∞ Copyright ©2005, Elsevier Inc. Allrights reserved. No part of this publication may be reproduced or transmitted in any form or by any means, electronic or me- chanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Permissions may be sought directly from Elsevier’s Science & Technology Rights Department in Oxford, UK: phone:(+44)1865843830,fax:(+44)1865853333,e-mail:[email protected] your request on-line via the Elsevier homepage (http://elsevier.com), by selecting “Customer Support” and then “Obtaining Permissions.” Library of Congress Cataloging-in-Publication Data Appication submitted British Library Cataloguing in Publication Data Acatalogue record for this book is availablefrom theBritish Library ISBN: 0-12-059876-0 Case bound ISBN: 0-12-088584-0 International Students Edition For allinformation on all ElsevierAcademicPress Publications visit our Website at www.books.elsevier.com Printed inthe UnitedStates of America 050607080910987654321 CONTENTS Preface xi 1 VectorAnalysis 1 1.1Definitions,ElementaryApproach ..................... 1 1.2RotationoftheCoordinateAxes ...................... 7 1.3ScalarorDotProduct ........................... 1 2 1.4VectororCross Product .......................... 1 8 1.5TripleScalarProduct,TripleVectorProduct ............... 2 5 1.6Gradient,∇................................. 3 2 1.7Divergence,∇................................ 3 8 1.8Curl,∇×.................................. 4 3 1.9SuccessiveApplicationsof ∇....................... 4 9 1.10VectorIntegration .............................. 5 4 1.11Gauss’Theorem ............................... 6 0 1.12Stokes’Theorem .............................. 6 4 1.13PotentialTheory .............................. 6 8 1.14Gauss’Law, Poisson’sEquation ...................... 7 9 1.15DiracDeltaFunction ............................ 8 3 1.16Helmholtz’sTheorem ............................ 9 5 AdditionalReadings ............................ 1 0 1 2 VectorAnalysisinCurvedCoordinatesandTensors 103 2.1OrthogonalCoordinatesin R3....................... 1 0 3 2.2DifferentialVectorOperators ....................... 1 1 0 2.3SpecialCoordinateSystems:Introduction ................ 1 1 4 2.4CircularCylinderCoordinates ....................... 1 1 5 2.5SphericalPolarCoordinates ........................ 1 2 3 v vi Contents 2.6TensorAnalysis ............................... 1 3 3 2.7Contraction,DirectProduct ........................ 1 3 9 2.8QuotientRule ................................ 1 4 1 2.9Pseudotensors, DualTensors ....................... 1 4 2 2.10GeneralTensors ............................... 1 5 1 2.11TensorDerivativeOperators ........................ 1 6 0 AdditionalReadings ............................ 1 6 3 3 DeterminantsandMatrices 165 3.1Determinants ................................ 1 6 5 3.2Matrices ................................... 1 7 6 3.3OrthogonalMatrices ............................ 1 9 5 3.4HermitianMatrices,UnitaryMatrices .................. 2 0 8 3.5DiagonalizationofMatrices ........................ 2 1 5 3.6NormalMatrices .............................. 2 3 1 AdditionalReadings ............................ 2 3 9 4 GroupTheory 241 4.1IntroductiontoGroupTheory ....................... 2 4 1 4.2GeneratorsofContinuousGroups ..................... 2 4 6 4.3OrbitalAngularMomentum ........................ 2 6 1 4.4AngularMomentumCoupling ....................... 2 6 6 4.5HomogeneousLorentzGroup ....................... 2 7 8 4.6LorentzCovarianceofMaxwell’sEquations ............... 2 8 3 4.7DiscreteGroups ............................... 2 9 1 4.8DifferentialForms ............................. 3 0 4 AdditionalReadings ............................ 3 1 9 5 InfiniteSeries 321 5.1FundamentalConcepts ........................... 3 2 1 5.2ConvergenceTests ............................. 3 2 5 5.3AlternatingSeries .............................. 3 3 9 5.4AlgebraofSeries .............................. 3 4 2 5.5SeriesofFunctions ............................. 3 4 8 5.6Taylor’sExpansion ............................. 3 5 2 5.7PowerSeries ................................ 3 6 3 5.8EllipticIntegrals .............................. 3 7 0 5.9BernoulliNumbers, Euler–MaclaurinFormula .............. 3 7 6 5.10AsymptoticSeries .............................. 3 8 9 5.11InfiniteProducts .............................. 3 9 6 AdditionalReadings ............................ 4 0 1 6 FunctionsofaComplexVariableIAnalyticProperties,Mapping 403 6.1ComplexAlgebra .............................. 4 0 4 6.2Cauchy–RiemannConditions ....................... 4 1 3 6.3Cauchy’sIntegralTheorem ......................... 4 1 8 Contents vii 6.4Cauchy’sIntegralFormula ......................... 4 2 5 6.5LaurentExpansion ............................. 4 3 0 6.6Singularities ................................. 4 3 8 6.7Mapping ................................... 4 4 3 6.8ConformalMapping ............................ 4 5 1 AdditionalReadings ............................ 4 5 3 7 FunctionsofaComplexVariableII 455 7.1CalculusofResidues ............................ 4 5 5 7.2DispersionRelations ............................ 4 8 2 7.3MethodofSteepestDescents ........................ 4 8 9 AdditionalReadings ............................ 4 9 7 8 TheGammaFunction(FactorialFunction) 499 8.1Definitions,SimpleProperties ....................... 4 9 9 8.2DigammaandPolygammaFunctions ................... 5 1 0 8.3Stirling’sSeries ............................... 5 1 6 8.4TheBetaFunction ............................. 5 2 0 8.5IncompleteGammaFunction ....................... 5 2 7 AdditionalReadings ............................ 5 3 3 9 DifferentialEquations 535 9.1PartialDifferentialEquations ....................... 5 3 5 9.2First-Order DifferentialEquations .................... 5 4 3 9.3SeparationofVariables ........................... 5 5 4 9.4SingularPoints ............................... 5 6 2 9.5SeriesSolutions—Frobenius’Method ................... 5 6 5 9.6A SecondSolution .............................. 5 7 8 9.7NonhomogeneousEquation—Green’sFunction ............. 5 9 2 9.8HeatFlow, or Diffusion,PDE ....................... 6 1 1 AdditionalReadings ............................ 6 1 8 10 Sturm–LiouvilleTheory—OrthogonalFunctions 621 10.1Self-AdjointODEs ............................. 6 2 2 10.2HermitianOperators ............................ 6 3 4 10.3Gram–SchmidtOrthogonalization ..................... 6 4 2 10.4CompletenessofEigenfunctions ...................... 6 4 9 10.5Green’sFunction—EigenfunctionExpansion ............... 6 6 2 AdditionalReadings ............................ 6 7 4 11 BesselFunctions 675 11.1Bessel FunctionsoftheFirst Kind, Jν(x)................. 6 7 5 11.2Orthogonality ................................ 6 9 4 11.3NeumannFunctions ............................ 6 9 9 11.4HankelFunctions .............................. 7 0 7 11.5ModifiedBessel Functions, Iν(x)andKν(x)............... 7 1 3 viii Contents 11.6AsymptoticExpansions ........................... 7 1 9 11.7SphericalBesselFunctions ......................... 7 2 5 AdditionalReadings ............................ 7 3 9 12 LegendreFunctions 741 12.1GeneratingFunction ............................ 7 4 1 12.2RecurrenceRelations ............................ 7 4 9 12.3Orthogonality ................................ 7 5 6 12.4AlternateDefinitions ............................ 7 6 7 12.5AssociatedLegendreFunctions ...................... 7 7 1 12.6SphericalHarmonics ............................ 7 8 6 12.7OrbitalAngularMomentumOperators .................. 7 9 3 12.8AdditionTheoremforSphericalHarmonics ............... 7 9 7 12.9IntegralsofThreeY’s ............................ 8 0 3 12.10LegendreFunctionsoftheSecondKind .................. 8 0 6 12.11VectorSphericalHarmonics ........................ 8 1 3 AdditionalReadings ............................ 8 1 6 13 MoreSpecialFunctions 817 13.1HermiteFunctions ............................. 8 1 7 13.2LaguerreFunctions ............................. 8 3 7 13.3ChebyshevPolynomials .......................... 8 4 8 13.4HypergeometricFunctions ......................... 8 5 9 13.5ConfluentHypergeometricFunctions ................... 8 6 3 13.6MathieuFunctions ............................. 8 6 9 AdditionalReadings ............................ 8 7 9 14 FourierSeries 881 14.1GeneralProperties ............................. 8 8 1 14.2Advantages,Uses ofFourierSeries .................... 8 8 8 14.3ApplicationsofFourierSeries ....................... 8 9 2 14.4PropertiesofFourierSeries ........................ 9 0 3 14.5GibbsPhenomenon ............................. 9 1 0 14.6DiscreteFourierTransform ........................ 9 1 4 14.7FourierExpansionsofMathieuFunctions ................ 9 1 9 AdditionalReadings ............................ 9 2 9 15 IntegralTransforms 931 15.1IntegralTransforms ............................. 9 3 1 15.2DevelopmentoftheFourierIntegral .................... 9 3 6 15.3FourierTransforms—InversionTheorem ................. 9 3 8 15.4FourierTransformofDerivatives ..................... 9 4 6 15.5ConvolutionTheorem ............................ 9 5 1 15.6MomentumRepresentation ......................... 9 5 5 15.7TransferFunctions ............................. 9 6 1 15.8LaplaceTransforms ............................. 9 6 5 Contents ix 15.9LaplaceTransformofDerivatives ..................... 9 7 1 15.10OtherProperties .............................. 9 7 9 15.11Convolution(Faltungs) Theorem ..................... 9 9 0 15.12Inverse LaplaceTransform ......................... 9 9 4 AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1003 16 IntegralEquations 1005 16.1Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1005 16.2IntegralTransforms,GeneratingFunctions . . . . . . . . . . . . . . . . 1012 16.3NeumannSeries,Separable(Degenerate)Kernels . . . . . . . . . . . . 1018 16.4Hilbert–SchmidtTheory . . . . . . . . . . . . . . . . . . . . . . . . . . 1029 AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1036 17 CalculusofVariations 1037 17.1A DependentandanIndependentVariable . . . . . . . . . . . . . . . . 1038 17.2ApplicationsoftheEuler Equation . . . . . . . . . . . . . . . . . . . . 1044 17.3SeveralDependentVariables . . . . . . . . . . . . . . . . . . . . . . . . 1052 17.4SeveralIndependentVariables . . . . . . . . . . . . . . . . . . . . . . . 1056 17.5SeveralDependentandIndependentVariables . . . . . . . . . . . . . . 1058 17.6LagrangianMultipliers . . . . . . . . . . . . . . . . . . . . . . . . . . . 1060 17.7VariationwithConstraints . . . . . . . . . . . . . . . . . . . . . . . . . 1065 17.8Rayleigh–RitzVariationalTechnique . . . . . . . . . . . . . . . . . . . 1072 AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1076 18 NonlinearMethodsandChaos 1079 18.1Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1079 18.2TheLogisticMap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1080 18.3SensitivitytoInitialConditionsandParameters . . . . . . . . . . . . . 1085 18.4NonlinearDifferentialEquations . . . . . . . . . . . . . . . . . . . . . 1088 AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1107 19 Probability 1109 19.1Definitions,SimpleProperties . . . . . . . . . . . . . . . . . . . . . . . 1109 19.2RandomVariables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1116 19.3BinomialDistribution . . . . . . . . . . . . . . . . . . . . . . . . . . . 1128 19.4PoissonDistribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1130 19.5Gauss’NormalDistribution . . . . . . . . . . . . . . . . . . . . . . . . 1134 19.6Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1138 AdditionalReadings . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1150 GeneralReferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1150 Index 1153 This page intentionally left blank PREFACE Throughsixeditionsnow, MathematicalMethodsforPhysicists hasprovidedallthemath- ematicalmethodsthataspiringsscientistsandengineersarelikelytoencounterasstudents and beginning researchers. More than enough material is included for a two-semester un- dergraduateorgraduatecourse. Thebookisadvancedinthesensethatmathematicalrelationsarealmostalwaysproven, in addition to being illustrated in terms of examples. These proofs are not what a mathe- matician would regard as rigorous, but sketch the ideas and emphasize the relations that are essential to the study of physics and related fields. This approach incorporates theo- rems that are usually not cited under the most general assumptions, but are tailored to the more restricted applications required by physics. For example, Stokes’ theorem is usually appliedbyaphysicisttoasurfacewiththetacitunderstandingthatitbesimplyconnected. Suchassumptionshavebeenmademoreexplicit. PROBLEM -SOLVING SKILLS The book also incorporates a deliberate focus on problem-solving skills. This more ad- vancedlevelofunderstandingandactivelearningisroutineinphysicscoursesandrequires practicebythereader.Accordingly,extensiveproblemsetsappearingineachchapterform an integral part of the book. They have been carefully reviewed, revised and enlarged for thisSixthEdition. PATHWAYS THROUGH THE MATERIAL Undergraduates may be best served if they start by reviewing Chapter 1 according to the level of training of the class. Section 1.2 on the transformation properties of vectors, the cross product, and the invariance of the scalar product under rotations may be postponed until tensor analysis is started, for which these sections form the introduction and serve as xi xii Preface examples. They may continue their studies with linear algebra in Chapter 3, then perhaps tensors and symmetries (Chapters 2 and 4), and next real and complex analysis (Chap- ters5–7), differentialequations(Chapters9, 10),andspecialfunctions(Chapters11–13). In general, the core of a graduate one-semester course comprises Chapters 5–10 and 11–13,whichdealwithrealandcomplexanalysis,differentialequations,andspecialfunc- tions. Depending on the level of the students in a course, some linear algebra in Chapter 3 (eigenvalues, for example), along with symmetries (group theory in Chapter 4), and ten- sors(Chapter2)maybecoveredasneededoraccordingtotaste.Grouptheorymayalsobe included with differential equations (Chapters 9 and 10). Appropriate relations have been includedandarediscussedinChapters4and9. A two-semester course can treat tensors, group theory, and special functions (Chap- ters 11–13) more extensively, and add Fourier series (Chapter 14), integral transforms (Chapter15),integralequations(Chapter16),andthecalculusofvariations(Chapter17). CHANGES TO THE SIXTH EDITION ImprovementstotheSixthEditionhavebeenmadeinnearlyallchaptersaddingexamples and problems and more derivations of results. Numerous left-over typos caused by scan- ning into LaTeX, an error-prone process at the rate of many errors per page, have been corrected along with mistakes, such as in the Dirac γ-matrices in Chapter 3. A few chap- ters have been relocated. The Gamma function is now in Chapter 8 following Chapters 6 and 7 on complex functions in one variable, as it is an application of these methods. Dif- ferential equations are now in Chapters 9 and 10. A new chapter on probability has been added,aswellasnewsubsectionsondifferentialformsandMathieufunctionsinresponse to persistent demands by readers and students over the years. The new subsections are more advanced and are written in the concise style of the book, thereby raising its level to thegraduatelevel.Manyexampleshavebeenadded,forexampleinChapters1and2,that are often used in physics or are standard lore of physics courses. A number of additions have been made in Chapter 3, such as on linear dependence of vectors, dual vector spaces and spectral decomposition of symmetric or Hermitian matrices. A subsection on the dif- fusion equation emphasizes methods to adapt solutions of partial differential equations to boundaryconditions.NewformulashavebeendevelopedforHermitepolynomialsandare includedinChapter13thatareusefulfortreatingmolecularvibrations;theyareofinterest tothechemicalphysicists. ACKNOWLEDGMENTS Wehavebenefitedfromtheadviceandhelpofmanypeople.Someoftherevisionsareinre- sponsetocommentsbyreadersandformerstudents,suchasDr.K.BodoorandJ.Hughes. WearegratefultothemandtoourEditorsBarbaraHollandandTomSingerwhoorganized accuracychecks.WewouldliketothankinparticularDr.MichaelBozoianandProf.Frank Harris for their invaluable help with the accuracy checking and Simon Crump, Production Editor,for hisexpertmanagementoftheSixthEdition. CHAPTER 1 VECTOR ANALYSIS 1.1 D EFINITIONS ,ELEMENTARY APPROACH In science and engineering we frequently encounter quantities that have magnitude and magnitude only: mass, time, and temperature. These we label scalarquantities, which re- main the same no matter what coordinates we use. In contrast, many interesting physical quantities have magnitude and, in addition, an associated direction. This second group includes displacement, velocity, acceleration, force, momentum, and angular momentum. Quantitieswithmagnitudeanddirectionarelabeled vectorquantities.Usually,inelemen- tary treatments, a vector is defined as a quantity having magnitude and direction. To dis- tinguishvectorsfrom scalars,weidentifyvectorquantitieswithboldfacetype,thatis, V. Ourvectormaybeconvenientlyrepresentedbyanarrow,withlengthproportionaltothe magnitude. The direction of the arrow gives the direction of the vector, the positive sense ofdirectionbeingindicatedbythepoint.Inthisrepresentation,vectoraddition C=A+B (1.1) consists in placing the rear end of vector Bat the point of vector A. VectorCis then represented by an arrow drawn from the rear of Ato the point of B. This procedure, the triangle law of addition, assigns meaning to Eq. (1.1) and is illustrated in Fig. 1.1. By completingtheparallelogram,weseethat C=A+B=B+A, (1.2) asshowninFig.1.2.In words, vectoradditionis commutative . Forthesumofthreevectors D=A+B+C, Fig.1.3,wemayfirstadd AandB: A+B=E. 1 2 Chapter 1 Vector Analysis FIGURE 1.1Trianglelawofvector addition. FIGURE 1.2Parallelogramlawof vectoraddition. FIGURE 1.3Vectoradditionis associative. Thenthis sumisaddedto C: D=E+C. Similarly,wemayfirst add BandC: B+C=F. Then D=A+F. Intermsof theoriginalexpression, (A+B)+C=A+(B+C). Vectoradditionis associative . A direct physical example of the parallelogram addition law is provided by a weight suspended by two cords. If the junction point ( Oin Fig. 1.4) is in equilibrium, the vector 1.1 Definitions, Elementary Approach 3 FIGURE 1.4Equilibriumofforces: F1+F2=−F3. sum of the two forces F1andF2must just cancelthe downwardforce of gravity, F3.H e r e theparallelogramadditionlawissubjecttoimmediateexperimentalverification.1 Subtraction may be handled by defining the negative of a vector as a vector of the same magnitudebutwithreverseddirection.Then A−B=A+(−B). InFig.1.3, A=E−B. Notethatthevectorsaretreatedasgeometricalobjectsthatareindependentofanycoor- dinatesystem.Thisconceptofindependenceofapreferredcoordinatesystemisdeveloped indetailinthenextsection. The representation of vector Aby an arrow suggests a second possibility. Arrow A (Fig.1.5),startingfromtheorigin,2terminatesatthepoint (Ax,Ay,Az).Thus,ifweagree that the vector is to start at the origin, the positive end may be specified by giving the Cartesiancoordinates (Ax,Ay,Az)ofthearrowhead. Although Acouldhaverepresentedanyvectorquantity(momentum,electricfield,etc.), one particularly important vector quantity, the displacement from the origin to the point 1Strictly speaking, the parallelogram addition was introduced as a definition. Experiments show that if we assume that the forcesarevectorquantitiesandwecombine thembyparallelogramaddition,theequilibriumcondition ofzeroresultantforceis satisfied. 2We could start from any point in our Cartesian reference frame; we choose the origin for simplicity. This freedom of shifting the origin of the coordinate system without affecting the geometry is called translation invariance . 4 Chapter 1 Vector Analysis FIGURE 1.5Cartesiancomponentsanddirectioncosinesof A. (x,y,z), isdenotedbythespecialsymbol r.Wethenhaveachoiceofreferringtothedis- placementaseitherthevector rorthecollection (x,y,z), thecoordinatesofitsendpoint: r↔(x,y,z). (1.3) Usingrfor the magnitude of vector r, we find that Fig. 1.5 shows that the endpoint coor- dinatesandthemagnitudearerelatedby x=rcosα, y=rcosβ, z=rcosγ. (1.4) Herecosα,cosβ,andcosγarecalledthe directioncosines ,αbeingtheanglebetweenthe given vector and the positive x-axis, and so on. One further bit of vocabulary: The quan- titiesAx,Ay, andAzare known as the (Cartesian) components ofAor theprojections ofA,with cos2α+cos2β+cos2γ=1. Thus, any vector Amay be resolved into its components (or projected onto the coordi- nateaxes)toyield Ax=Acosα,etc.,asinEq.(1.4).Wemaychoosetorefertothevector as a single quantity Aor to its components (Ax,Ay,Az). Note that the subscript xinAx denotes the xcomponent and not a dependence on the variable x. The choice between usingAor its components (Ax,Ay,Az)is essentially a choice between a geometric and an algebraic representation. Use either representation at your convenience. The geometric “arrowinspace”mayaidinvisualization.Thealgebraicsetofcomponentsisusuallymore suitablefor precisenumericalor algebraiccalculations. Vectors enter physics in two distinct forms. (1) Vector Amay represent a single force acting at a single point. The force of gravity acting at the center of gravity illustrates this form. (2) Vector Amay be defined over some extended region; that is, Aand its compo- nents may be functions of position: Ax=Ax(x,y,z), and so on. Examples of this sort includethevelocityofafluidvaryingfrompointtopointoveragivenvolumeandelectric and magnetic fields. These two cases may be distinguished by referring to the vector de- fined over a region as a vector field . The concept of the vector defined over a region and 1.1 Definitions, Elementary Approach 5 being a function of position will become extremely important when we differentiate and integratevectors. Atthisstageitisconvenienttointroduceunitvectorsalongeachofthecoordinateaxes. Letˆxbe a vector of unit magnitude pointing in the positive x-direction,ˆy, a vector of unit magnitude in the positive y-direction, and ˆza vector of unit magnitude in the positive z- direction. Then ˆxAxis a vector with magnitude equal to |Ax|and in the x-direction. By vectoraddition, A=ˆxAx+ˆyAy+ˆzAz. (1.5) Notethatif Avanishes,allof itscomponentsmustvanishindividually;thatis, if A=0,thenAx=Ay=Az=0. Thismeansthattheseunitvectorsserveasa basis,orcompletesetofvectors,inthethree- dimensionalEuclideanspaceintermsofwhichanyvectorcanbeexpanded.Thus,Eq.(1.5) isanassertionthatthethreeunitvectors ˆx,ˆy,andˆzspanourrealthree-dimensionalspace: Any vector may be written as a linear combination of ˆx,ˆy, andˆz.Sinceˆx,ˆy, andˆzare linearly independent (no one is a linear combination of the other two), they form a basis for the real three-dimensional Euclidean space. Finally, by the Pythagorean theorem, the magnitudeofvector Ais |A|=parenleftbig A2 x+A2y+A2zparenrightbig1/2. (1.6) Notethatthecoordinateunitvectorsarenottheonlycompleteset,orbasis.Thisresolution of a vector into its components can be carried out in a variety of coordinate systems, as shown in Chapter 2. Here we restrict ourselves to Cartesian coordinates, where the unit vectorshavethecoordinates ˆx=(1,0,0),ˆy=(0,1,0)andˆz=(0,0,1)andareallconstant inlengthanddirection,propertiescharacteristicof Cartesiancoordinates. As a replacement of the graphical technique, addition and subtraction of vectors may now be carried out in terms of their components. For A=ˆxAx+ˆyAy+ˆzAzandB= ˆxBx+ˆyBy+ˆzBz, A±B=ˆx(Ax±Bx)+ˆy(Ay±By)+ˆz(Az±Bz). (1.7) It should be emphasized here that the unit vectors ˆx,ˆy, andˆzare used for convenience. They are not essential; we can describe vectors and use them entirely in terms of their components: A↔(Ax,Ay,Az).This is the approach of the two more powerful, more sophisticated definitions of vector to be discussed in the next section. However, ˆx,ˆy, and ˆzemphasizethe direction . So far we havedefinedtheoperationsof additionand subtractionof vectors. In thenext sections,threevarietiesofmultiplicationwillbedefinedonthebasisoftheirapplicability: a scalar, or inner, product, a vector product peculiar to three-dimensional space, and a direct,orouter,productyieldingasecond-ranktensor.Divisionbya vectorisnotdefined. 6 Chapter 1 Vector Analysis Exercises 1.1.1 Showhowtofind AandB,gi v enA+BandA−B. 1.1.2 The vector Awhose magnitude is 1 .732 units makes equal angles with the coordinate axes.Find Ax,Ay, andAz. 1.1.3 Calculate the components of a unit vector that lies in the xy-plane and makes equal angleswiththepositivedirectionsofthe x-andy-axes. 1.1.4 The velocity of sailboat Arelative to sailboat B,vrel, is defined by the equation vrel= vA−vB, wherevAis the velocity of AandvBis the velocity of B. Determine the velocityof Arelativeto Bif vA=30km/hreast vB=40km/hrnorth. ANS.vrel=50 km/hr, 53.1◦southofeast. 1.1.5 A sailboat sails for 1 hr at 4 km /hr (relative to the water) on a steady compass heading of 40◦east of north. The sailboat is simultaneously carried along by a current. At the endofthehourtheboatis6.12kmfromitsstartingpoint.Thelinefromitsstartingpoint toitslocationlies 60◦eastofnorth.Findthe x(easterly)and y(northerly)components ofthewater’svelocity. ANS.veast=2.73 km/hr,vnorth≈0k m/hr. 1.1.6 Avectorequationcanbereducedtotheform A=B.Fromthisshowthattheonevector equation is equivalent to threescalar equations. Assuming the validity of Newton’s second law, F=ma,a savectorequation, this means that axdepends only on Fxand isindependentof FyandFz. 1.1.7 The vertices A,B, andCof a triangle are given by the points (−1,0,2), (0,1,0), and (1,−1,0), respectively. Find point Dso that the figure ABCDforms a plane parallel- ogram. ANS.(0,−2,2)or(2,0,−2). 1.1.8 A triangle is defined by the vertices of three vectors A,BandCthat extend from the origin. In terms of A,B, andCshow that the vectorsum of the successive sides of the triangle(AB+BC+CA)iszero, wheretheside ABisfromAtoB,etc. 1.1.9 Asphereofradius ais centeredatapoint r1. (a) Writeoutthealgebraicequationfor thesphere. (b) Writeouta vectorequationforthesphere. ANS. (a) (x−x1)2+(y−y1)2+(z−z1)2=a2. (b)r=r1+a, withr1=center. (atakes onalldirectionsbuthasafixedmagnitude a.) 1.2 Rotation of the Coordinate Axes 7 1.1.10 A corner reflector is formed by three mutually perpendicular reflecting surfaces. Show that a ray of light incident upon the corner reflector (striking all three surfaces) is re- flectedbackalongalineparalleltothelineofincidence. Hint. Consider the effect of a reflection on the components of a vector describing the directionofthelightray. 1.1.11 Hubble’s law . Hubble found that distant galaxies are receding with a velocity propor- tionaltotheirdistancefromwhereweareonEarth.Forthe ithgalaxy, vi=H0ri, with us at the origin. Show that this recession of the galaxies from us does notimply that we are at the center of the universe. Specifically, take the galaxy at r1as a new originandshowthatHubble’slawisstillobeyed. 1.1.12 Findthediagonalvectorsofaunitcubewithonecornerattheoriginanditsthreesides lying along Cartesian coordinates axes. Show that there are four diagonals with length√ 3.Representingtheseasvectors,whataretheircomponents?Showthatthediagonals ofthecube’sfaceshavelength√ 2 anddeterminetheircomponents. 1.2 R OTATION OF THE COORDINATE AXES3 In the preceding section vectors were defined or represented in two equivalent ways: (1) geometrically by specifying magnitude and direction, as with an arrow, and (2) al- gebraically by specifying the components relative to Cartesian coordinate axes. The sec- ond definition is adequate for the vector analysis of this chapter. In this section two more refined, sophisticated, and powerful definitions are presented. First, the vector field is de- finedintermsofthebehaviorofitscomponentsunderrotationofthecoordinateaxes.This transformation theory approach leads into the tensor analysis of Chapter 2 and groups of transformations in Chapter 4. Second, the component definition of Section 1.1 is refined andgeneralizedaccordingtothemathematician’sconceptsofvectorandvectorspace.This approachleadstofunctionspaces, includingtheHilbertspace. The definition of vector as a quantity with magnitude and direction is incomplete. On the one hand, we encounter quantities, such as elastic constants and index of refraction in anisotropic crystals, that have magnitude and direction butthat are not vectors. On the other hand, our naïve approach is awkward to generalize to extend to more complex quantities. We seek a new definition of vector field using our coordinate vector ras a prototype. Thereisaphysicalbasisforourdevelopmentofanewdefinition.Wedescribeourphys- ical world by mathematics, but it and any physical predictions we may make must be independent ofourmathematicalconventions. In our specific case we assume that space is isotropic; that is, there is no preferred di- rection, or all directions are equivalent. Then the physical system being analyzed or the physical law being enunciated cannot and must not depend on our choice or orientation of the coordinate axes. Specifically, if a quantity Sdoes not depend on the orientation of thecoordinateaxes,itis calledascalar. 3This sectionis optional here.It willbe essential for Chapter2. 8 Chapter 1 Vector Analysis FIGURE 1.6Rotationof Cartesiancoordinateaxesaboutthe z-axis. Now we return to the concept of vector ras a geometric object independent of the coordinate system. Let us look at rin two different systems, one rotated in relation to the other. For simplicity we consider first the two-dimensional case. If the x-,y-coordinates are rotated counterclockwise through an angle ϕ,keeping r ,fixed(Fig. 1.6), we get the fol- lowing relations between the components resolved in the original system (unprimed) and thoseresolvedinthenewrotatedsystem(primed): x′=xcosϕ+ysinϕ, y′=−xsinϕ+ycosϕ.(1.8) We saw in Section 1.1 that a vector could be represented by the coordinates of a point; thatis,thecoordinateswereproportionaltothevectorcomponents.Hencethecomponents of a vector must transform under rotation as coordinates of a point (such as r). Therefore wheneveranypairofquantities AxandAyinthexy-coordinatesystemistransformedinto (A′ x,A′y)bythisrotationof thecoordinatesystemwith A′ x=Axcosϕ+Aysinϕ, A′y=−Axsinϕ+Aycosϕ,(1.9) wedefine4AxandAyasthecomponentsofavector A.Ourvectornowisdefinedinterms ofthetransformationofitscomponentsunderrotationofthecoordinatesystem.If Axand Aytransforminthesamewayas xandy,thecomponentsofthegeneraltwo-dimensional coordinatevector r,theyarethecomponentsofavector A.IfAxandAydonotshowthis 4A scalarquantity does not depend on the orientation of coordinates; S′=Sexpresses the fact that it is invariant under rotation of thecoordinates. 1.2 Rotation of the Coordinate Axes 9 form invariance (also called covariance ) when the coordinates are rotated, they do not formavector. Thevectorfieldcomponents AxandAysatisfyingthedefiningequations,Eqs.(1.9),as- sociate a magnitude Aand a direction with each point in space. The magnitude is a scalar quantity, invariant to the rotation of the coordinate system. The direction (relative to the unprimed system) is likewise invariant to the rotation of the coordinate system (see Exer- cise 1.2.1). The result of all this is that the components of a vector may vary according to therotationof theprimedcoordinatesystem. This iswhatEqs. (1.9) say. Butthevariation withtheangleisjustsuchthatthecomponentsintherotatedcoordinatesystem A′ xandA′y define a vector with the same magnitude and the same direction as the vector defined by thecomponents AxandAyrelativetothe x-,y-coordinateaxes.(CompareExercise1.2.1.) The components of Ain a particular coordinate system constitute the representation of Ain that coordinate system. Equations (1.9), the transformation relations, are a guarantee thattheentity Ais independentoftherotationofthecoordinatesystem. Togoontothreeand,later,fourdimensions,wefinditconvenienttouseamorecompact notation.Let x→x1 y→x2(1.10) a11=cosϕ, a 12=sinϕ, a21=−sinϕ, a 22=cosϕ.(1.11) ThenEqs. (1.8) become x′ 1=a11x1+a12x2, x′ 2=a21x1+a22x2.(1.12) Thecoefficient aijmaybeinterpretedasadirectioncosine,thecosineoftheanglebetween x′ iandxj;thatis, a12=cos(x′ 1,x2)=sinϕ, a21=cos(x′ 2,x1)=cosparenleftbig ϕ+π 2parenrightbig =−sinϕ.(1.13) The advantage of the new notation5is that it permits us to use the summation symbolsummationtext andtorewriteEqs. (1.12) as x′ i=2summationdisplay j=1aijxj,i=1,2. (1.14) Note that iremains as a parameter that gives rise to one equation when it is set equal to 1 and to a second equation when it is set equal to 2. The index j, of course, is a summation index, a dummy index, and, as with a variable of integration, jmay be replaced by any otherconvenientsymbol. 5You may wonder at the replacement of one parameter ϕby four parameters aij.C l e a r l y ,t h e aijdo not constitute a minimum set of parameters. For two dimensions the four aijare subject to the three constraints given in Eq. (1.18). The justification for this redundant set of direction cosines is the convenience it provides. Hopefully, this convenience will become more apparent in Chapters 2 and 3. For three-dimensional rotations (9 aijbut only three independent) alternate descriptions are provided by: (1)theEuleranglesdiscussedinSection3.3,(2)quaternions,and(3)theCayley–Kleinparameters.Thesealternativeshavetheir respective advantages anddisadvantages. 10 Chapter 1 Vector Analysis Thegeneralizationtothree,four,or Ndimensionsisnowsimple.Thesetof Nquantities Vjis said to be the components of an N-dimensional vector Vif and only if their values relativetotherotatedcoordinateaxesaregivenby V′ i=Nsummationdisplay j=1aijVj,i=1,2,...,N. (1.15) As before, aijis the cosine of the angle between x′ iandxj. Often the upper limit Nand the corresponding range of iwill not be indicated. It is taken for granted that you know howmanydimensionsyourspacehas. From the definition of aijas the cosine of the angle between the positive x′ idirection andthepositive xjdirectionwemaywrite(Cartesiancoordinates)6 aij=∂x′ i ∂xj. (1.16a) Usingtheinverserotation( ϕ→−ϕ) yields xj=2summationdisplay i=1aijx′ ior∂xj ∂x′ i=aij. (1.16b) Note that these are partial derivatives . By use of Eqs. (1.16a) and (1.16b), Eq. (1.15) becomes V′ i=Nsummationdisplay j=1∂x′ i ∂xjVj=Nsummationdisplay j=1∂xj ∂x′ iVj. (1.17) Thedirectioncosines aijsatisfyan orthogonalitycondition summationdisplay iaijaik=δjk (1.18) or,equivalently, summationdisplay iajiaki=δjk. (1.19) Here,thesymbol δjkis theKroneckerdelta,definedby δjk=1forj=k, δjk=0forj/negationslash=k.(1.20) It is easily verified that Eqs. (1.18) and (1.19) hold in the two-dimensional case by substituting in the specific aijfrom Eqs. (1.11). The result is the well-known identity sin2ϕ+cos2ϕ=1 for the nonvanishing case. To verify Eq. (1.18) in general form, we mayusethepartialderivativeforms ofEqs. (1.16a) and(1.16b) toobtain summationdisplay i∂xj ∂x′ i∂xk ∂x′ i=summationdisplay i∂xj ∂x′ i∂x′ i ∂xk=∂xj ∂xk. (1.21) 6Differentiate x′ iwithrespect to xj. Seediscussion following Eq.(1.21). 1.2 Rotation of the Coordinate Axes 11 The last step follows by the standard rules for partial differentiation, assuming that xjis af u n c t i o no f x′ 1,x′ 2,x′ 3, and so on. The final result, ∂xj/∂xk, is equal to δjk, sincexjand xkas coordinate lines ( j/negationslash=k) are assumed to be perpendicular (two or three dimensions) or orthogonal (for any number of dimensions). Equivalently, we may assume that xjand xk(j/negationslash=k)aretotallyindependentvariables.If j=k,thepartialderivativeisclearlyequal to 1. In redefining a vector in terms of how its components transform under a rotation of the coordinatesystem, weshouldemphasizetwopoints: 1. This definition is developed because it is useful and appropriate in describing our physical world. Our vectorequationswill be independentof anyparticular coordinate system. (The coordinate system need not even be Cartesian.) The vector equation can always be expressed in some particular coordinate system, and, to obtain numerical results, wemustultimatelyexpresstheequationinsomespecificcoordinatesystem. 2. Thisdefinitionissubjecttoageneralizationthatwillopenupthebranchofmathemat- ics knownastensoranalysis(Chapter2). A qualification is in order. The behavior of the vector components under rotation of the coordinates is used in Section 1.3 to prove that a scalar product is a scalar, in Section 1.4 to prove that a vector product is a vector, and in Section 1.6 to show that the gradient of a scalarψ,∇ψ, is a vector. The remainder of this chapter proceeds on the basis of the less restrictivedefinitionsofthevectorgiveninSection1.1. Summary: Vectors and Vector Space It is customary in mathematics to label an ordered triple of real numbers ( x1,x2,x3)a vector x. The number xnis called the nth component of vector x. The collection of all such vectors (obeying the properties that follow) form a three-dimensional real vector space.Weascribefivepropertiestoourvectors:If x=(x1,x2,x3)andy=(y1,y2,y3), 1. Vectorequality: x=ymeansxi=yi,i=1,2,3. 2. Vectoraddition: x+y=zmeansxi+yi=zi,i=1,2,3. 3. Scalarmultiplication: ax↔(ax1,ax2,ax3)(withareal). 4. Negativeofavector: −x=(−1)x↔(−x1,−x2,−x3). 5. Nullvector:Thereexistsanullvector 0↔(0,0,0). Since our vector components are real (or complex) numbers, the following properties alsohold: 1. Additionofvectorsiscommutative: x+y=y+x. 2. Additionofvectorsisassociative: (x+y)+z=x+(y+z). 3. Scalarmultiplicationis distributive: a(x+y)=ax+ay,also(a+b)x=ax+bx. 4. Scalarmultiplicationis associative: (ab)x=a(bx). 12 Chapter 1 Vector Analysis Further,thenullvector 0isunique,asis thenegativeofagivenvector x. Sofarasthevectorsthemselvesareconcernedthisapproachmerelyformalizesthecom- ponentdiscussionofSection1.1.Theimportanceliesintheextensions,whichwillbecon- sidered in later chapters. In Chapter 4, we show that vectors form both an Abelian group underadditionandalinearspacewiththetransformationsinthelinearspacedescribedby matrices.Finally,andperhapsmostimportant,foradvancedphysicstheconceptofvectors presentedheremaybegeneralizedto(1)complexquantities,7(2)functions,and(3)aninfi- nitenumberofcomponents.Thisleadstoinfinite-dimensionalfunctionspaces,theHilbert spaces, which are important in modern quantum theory. A brief introduction to function expansionsandHilbertspaceappearsinSection10.4. Exercises 1.2.1 (a) Show that the magnitude of a vector A,A=(A2 x+A2y)1/2, is independent of the orientationof therotatedcoordinatesystem, parenleftbig A2 x+A2yparenrightbig1/2=parenleftbig A′2 x+A′2 yparenrightbig1/2, thatis, independentoftherotationangle ϕ. This independence of angle is expressed by saying that Aisinvariant under rotations. (b) At a given point (x,y),Adefines an angle αrelative to the positive x-axis and α′relative to the positive x′-axis. The angle from xtox′isϕ. Show that A=A′ definesthe samedirectioninspacewhenexpressedintermsofitsprimedcompo- nentsas intermsofits unprimedcomponents;thatis, α′=α−ϕ. 1.2.2 Prove the orthogonality conditionsummationtext iajiaki=δjk. As a special case of this, the direc- tioncosinesof Section1.1 satisfytherelation cos2α+cos2β+cos2γ=1, aresultthatalsofollowsfromEq. (1.6). 1.3 S CALAR OR DOTPRODUCT Havingdefinedvectors,wenowproceedtocombinethem.Thelawsforcombiningvectors mustbemathematicallyconsistent.Fromthepossibilitiesthatareconsistentweselecttwo thatarebothmathematicallyandphysicallyinteresting.Athirdpossibilityisintroducedin Chapter2, inwhichweformtensors. The projection of a vector Aonto a coordinate axis, which gives its Cartesian compo- nents in Eq. (1.4), defines a special geometrical case of the scalar product of Aand the coordinateunitvectors: Ax=Acosα≡A·ˆx,Ay=Acosβ≡A·ˆy,Az=Acosγ≡A·ˆz.(1.22) 7Then-dimensionalvectorspaceofreal n-tuplesisoftenlabeled Rnandthen-dimensionalvectorspaceofcomplex n-tuplesis labeled Cn. 1 . 3 S c a l a ro rD o tP r o d u c t 1 3 Thisspecialcaseofascalarproductinconjunctionwithgeneralpropertiesthescalarprod- uctis sufficienttoderivethegeneralcaseofthescalarproduct. Just as the projection is linear in A, we want the scalar product of two vectors to be linearinAandB, thatis, obeythedistributiveandassociativelaws A·(B+C)=A·B+A·C (1.23a) A·(yB)=(yA)·B=yA·B, (1.23b) whereyisanumber.Nowwecanusethedecompositionof BintoitsCartesiancomponents accordingtoEq.(1.5), B=Bxˆx+Byˆy+Bzˆz,toconstructthegeneralscalarordotproduct ofthevectors AandBas A·B=A·(Bxˆx+Byˆy+Bzˆz) =BxA·ˆx+ByA·ˆy+BzA·ˆzuponapplyingEqs. (1.23a)and(1.23b) =BxAx+ByAy+BzAzuponsubstitutingEq.(1.22). Hence A·B≡summationdisplay iBiAi=summationdisplay iAiBi=B·A. (1.24) IfA=Bin Eq. (1.24), we recover the magnitude A=(summationtextA2 i)1/2ofAin Eq. (1.6) from Eq.(1.24). It is obvious from Eq. (1.24) that the scalar product treats AandBalike, or is sym- metric in AandB, and is commutative. Thus, alternatively and equivalently, we can first generalize Eqs. (1.22) to the projection ABofAonto the direction of a vector B/negationslash=0 asAB=Acosθ≡A·ˆB, whereˆB=B/Bis the unit vector in the direction of Bandθ is the angle between AandB, as shown in Fig. 1.7. Similarly, we project BontoAas BA=Bcosθ≡B·ˆA. Second, we make these projections symmetric in AandB, which leadstothedefinition A·B≡ABB=ABA=ABcosθ. (1.25) FIGURE 1.7Scalarproduct A·B=ABcosθ. 14 Chapter 1 Vector Analysis FIGURE 1.8Thedistributivelaw A·(B+C)=ABA+ACA=A(B+C)A, Eq. (1.23a). ThedistributivelawinEq.(1.23a)isillustratedinFig.1.8,whichshowsthatthesumof the projections of BandContoA,BA+CAis equal to the projection of B+ContoA, (B+C)A. ItfollowsfromEqs.(1.22),(1.24),and(1.25)thatthecoordinateunitvectorssatisfythe relations ˆx·ˆx=ˆy·ˆy=ˆz·ˆz=1, (1.26a) whereas ˆx·ˆy=ˆx·ˆz=ˆy·ˆz=0. (1.26b) Ifthecomponentdefinition,Eq.(1.24),islabeledanalgebraicdefinition,thenEq.(1.25) is a geometric definition. One of the most common applications of the scalar product in physics is in the calculation of work=force·displacement·cosθ, which is interpreted as displacement times the projection of the force along the displacement direction, i.e., the scalarproductofforce anddisplacement, W=F·S. IfA·B=0 and we know that A/negationslash=0 andB/negationslash=0, then, from Eq. (1.25), cos θ=0, or θ=90◦,270◦, and so on. The vectors AandBmust be perpendicular. Alternately, we may sayAandBare orthogonal. The unit vectors ˆx,ˆy, andˆzare mutually orthogonal. To developthis notionof orthogonalityonemorestep, supposethat nis a unitvectorand ris anonzerovectorinthe xy-plane;thatis, r=ˆxx+ˆyy(Fig. 1.9). If n·r=0 forallchoicesof r,thennmustbeperpendicular(orthogonal)tothe xy-plane. Often it is convenient to replace ˆx,ˆy, andˆzby subscripted unit vectors em,m=1,2,3, withˆx=e1, andso on.ThenEqs. (1.26a)and(1.26b)become em·en=δmn. (1.26c) Form/negationslash=nthe unit vectors emandenare orthogonal. For m=neach vector is normal- ized to unity, that is, has unit magnitude. The set emis said to be orthonormal .Am a j o r advantage of Eq. (1.26c) over Eqs. (1.26a) and (1.26b) is that Eq. (1.26c) may readily be generalized to N-dimensional space: m,n=1,2,...,N. Finally, we are picking sets of unitvectors emthatareorthonormalforconvenience–averygreatconvenience. 1 . 3 S c a l a ro rD o tP r o d u c t 1 5 FIGURE 1.9Anormalvector. Invariance of the Scalar Product Under Rotations We havenot yet shown that the word scalaris justifiedor that the scalar productis indeed a scalar quantity. To do this, we investigate the behavior of A·Bunder a rotation of the coordinatesystem. ByuseofEq. (1.15), A′ xB′ x+A′yB′ y+A′zB′ z=summationdisplay iaxiAisummationdisplay jaxjBj+summationdisplay iayiAisummationdisplay jayjBj +summationdisplay iaziAisummationdisplay jazjBj. (1.27) Usingtheindices kandlto sum over x,y,andz, weobtain summationdisplay kA′ kB′ k=summationdisplay lsummationdisplay isummationdisplay jaliAialjBj, (1.28) and,byrearrangingthetermsontheright-handside, wehave summationdisplay kA′kB′ k=summationdisplay lsummationdisplay isummationdisplay j(alialj)AiBj=summationdisplay isummationdisplay jδijAiBj=summationdisplay iAiBi.(1.29) The last two steps follow by using Eq. (1.18), the orthogonality condition of the direction cosines, and Eqs. (1.20), which define the Kronecker delta. The effect of the Kronecker delta is to cancel all terms in a summation over either index except the term for which the indices are equal. In Eq. (1.29) its effect is to set j=iand to eliminate the summation overj. Of course, we could equally well set i=jand eliminate the summation over i. 16 Chapter 1 Vector Analysis Equation(1.29)givesus summationdisplay kA′ kB′ k=summationdisplay iAiBi, (1.30) whichisjustourdefinitionofascalarquantity,onethatremainsinvariantundertherotation ofthecoordinatesystem. In a similar approach that exploits this concept of invariance, we take C=A+Band dotit intoitself: C·C=(A+B)·(A+B) =A·A+B·B+2A·B. (1.31) Since C·C=C2, (1.32) thesquareof themagnitudeofvector Candthusaninvariantquantity,weseethat A·B=1 2parenleftbig C2−A2−B2parenrightbig ,invariant. (1.33) Since the right-hand side of Eq. (1.33) is invariant—that is, a scalar quantity—the left- hand side, A·B, must also be invariant under rotation of the coordinate system. Hence A·Bisascalar. Equation(1.31) isreallyanotherformof thelawofcosines,whichis C2=A2+B2+2ABcosθ. (1.34) Comparing Eqs. (1.31) and (1.34), we have another verification of Eq. (1.25), or, if pre- ferred, avectorderivationofthelawof cosines(Fig.1.10). The dot product, given by Eq. (1.24), may be generalized in two ways. The space need not be restricted to three dimensions. In n-dimensional space, Eq. (1.24) applies with the sumrunningfrom1to n.Moreover, nmaybeinfinity,withthesumthenaconvergentinfi- niteseries(Section5.2).Theothergeneralizationextendstheconceptofvectortoembrace functions.The functionanalogofadot,or inner,productappearsinSection10.4. FIGURE 1.10Thelawofcosines. 1 . 3 S c a l a ro rD o tP r o d u c t 1 7 Exercises 1.3.1 Twounitmagnitudevectors eiandejarerequiredtobeeitherparallelorperpendicular to each other. Show that ei·ejprovides an interpretation of Eq. (1.18), the direction cosineorthogonalityrelation. 1.3.2 Given that (1) the dot product of a unit vector with itself is unity and (2) this relation is valid in all (rotated) coordinate systems, show that ˆx′·ˆx′=1 (with the primed system rotated 45◦aboutthe z-axisrelativetotheunprimed)impliesthat ˆx·ˆy=0. 1.3.3 Thevector r,startingattheorigin,terminatesatandspecifiesthepointinspace (x,y,z). Findthesurface sweptoutbythetipof rif (a)(r−a)·a=0.Characterize ageometrically. (b)(r−a)·r=0.Describethegeometricroleof a. Thevector aisconstant(in magnitudeanddirection). 1.3.4 The interaction energy between two dipoles of moments µ1andµ2may be written in thevectorform V=−µ1·µ2 r3+3(µ1·r)(µ2·r) r5 andinthescalarform V=µ1µ2 r3(2cosθ1cosθ2−sinθ1sinθ2cosϕ). Hereθ1andθ2are the angles of µ1andµ2relative to r, whileϕis the azimuth of µ2 relativetothe µ1–rplane(Fig.1.11). Showthatthesetwoforms areequivalent. Hint:Equation(12.178)willbehelpful. 1.3.5 A pipe comes diagonally down the south wall of a building, making an angle of 45◦ withthehorizontal.Comingintoacorner,thepipeturnsandcontinuesdiagonallydown a west-facing wall, still making an angle of 45◦with the horizontal. What is the angle betweenthesouth-wallandwest-wallsectionsofthepipe? ANS. 120◦. 1.3.6 Find the shortest distance of an observer at the point (2,1,3)from a rocket in free flight with velocity (1,2,3)m/s. The rocket was launched at time t=0f r o m(1,1,1). Lengthsareinkilometers. 1.3.7 Prove the law of cosines from the triangle with corners at the point of CandAin Fig.1.10andtheprojectionofvector Bontovector A. FIGURE 1.11Twodipolemoments. 18 Chapter 1 Vector Analysis 1.4 V ECTOR OR CROSS PRODUCT A second form of vector multiplication employs the sine of the included angle instead of the cosine. For instance, the angular momentum of a body shown at the point of the distancevectorinFig.1.12is definedas angularmomentum =radiusarm×linearmomentum =distance×linearmomentum ×sinθ. For convenience in treating problems relating to quantities such as angular momentum, torque,andangularvelocity,wedefinethevectorproduct,orcross product,as C=A×B,withC=ABsinθ. (1.35) Unlike the preceding case of the scalar product, Cis now a vector, and we assign it a directionperpendiculartotheplaneof AandBsuchthat A,B,andCformaright-handed system.Withthischoiceof directionwehave A×B=−B×A,anticommutation . (1.36a) Fromthisdefinitionof cross productwehave ˆx׈x=ˆy׈y=ˆz׈z=0, (1.36b) whereas ˆx׈y=ˆz,ˆy׈z=ˆx,ˆz׈x=ˆy, ˆy׈x=−ˆz,ˆz׈y=−ˆx,ˆx׈z=−ˆy.(1.36c) Amongtheexamplesofthecrossproductinmathematicalphysicsaretherelationbetween linearmomentum pandangularmomentum L,withLdefinedas L=r×p, FIGURE 1.12Angularmomentum. 1 . 4 V e c t o ro rC r o s sP r o d u c t 1 9 FIGURE 1.13Parallelogramrepresentationofthevectorproduct. andtherelationbetweenlinearvelocity vandangularvelocity ω, v=ω×r. Vectorsvandpdescribe properties of the particle or physical system. However, the posi- tion vector ris determined by the choice of the origin of the coordinates. This means that ωandLdependonthechoiceoftheorigin. The familiar magnetic induction Bis usually defined by the vector product force equa- tion8 FM=qv×B(mksunits) . Herevis the velocity of the electric charge qandFMis the resulting force on the moving charge. The cross product has an important geometrical interpretation, which we shall use in subsequent sections. In the parallelogram defined by AandB(Fig. 1.13), Bsinθis the height if Ais taken as the length of the base. Then |A×B|=ABsinθis theareaof the parallelogram.Asavector, A×Bistheareaoftheparallelogramdefinedby AandB,with the area vector normal to the plane of the parallelogram. This suggests that area (with its orientationinspace)maybetreatedasavectorquantity. An alternate definition of the vector product can be derived from the special case of the coordinateunitvectorsinEqs.(1.36c)inconjunctionwiththelinearityofthecrossproduct inbothvectorarguments,inanalogywithEqs. (1.23) forthedotproduct, A×(B+C)=A×B+A×C, (1.37a) (A+B)×C=A×C+B×C, (1.37b) A×(yB)=yA×B=(yA)×B, (1.37c) 8Theelectricfield Eis assumed hereto be zero. 20 Chapter 1 Vector Analysis whereyisanumberagain.Usingthedecompositionof AandBintotheirCartesiancom- ponentsaccordingtoEq.(1.5), wefind A×B≡C=(Cx,Cy,Cz)=(Axˆx+Ayˆy+Azˆz)×(Bxˆx+Byˆy+Bzˆz) =(AxBy−AyBx)ˆx׈y+(AxBz−AzBx)ˆx׈z +(AyBz−AzBy)ˆy׈z upon applying Eqs. (1.37a) and (1.37b) and substituting Eqs. (1.36a), (1.36b), and (1.36c) sothattheCartesiancomponentsof A×Bbecome Cx=AyBz−AzBy,C y=AzBx−AxBz,C z=AxBy−AyBx,(1.38) or Ci=AjBk−AkBj, i,j,k alldifferent , (1.39) andwithcyclicpermutationoftheindices i,j,andkcorrespondingto x,y,andz,respec- tively.Thevectorproduct Cmaybemnemonicallyrepresentedbyadeterminant,9 C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz AxAyAz BxByBzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle≡ˆxvextendsinglevextendsinglevextendsinglevextendsingleAyAz ByBzvextendsinglevextendsinglevextendsinglevextendsingle−ˆyvextendsinglevextendsinglevextendsinglevextendsingleAxAz BxBzvextendsinglevextendsinglevextendsinglevextendsingle+ˆzvextendsinglevextendsinglevextendsinglevextendsingleAxAy BxByvextendsinglevextendsinglevextendsinglevextendsingle,(1.40) whichismeanttobeexpandedacrossthetoprowtoreproducethethreecomponentsof C listedinEqs. (1.38). Equation (1.35) might be called a geometric definition of the vector product. Then Eqs. (1.38)wouldbeanalgebraicdefinition. To show the equivalence of Eq. (1.35) and the component definition, Eqs. (1.38), let us formA·CandB·C,usingEqs. (1.38). Wehave A·C=A·(A×B) =Ax(AyBz−AzBy)+Ay(AzBx−AxBz)+Az(AxBy−AyBx) =0. (1.41) Similarly, B·C=B·(A×B)=0. (1.42) Equations (1.41) and (1.42) show that Cis perpendicular to both AandB(cosθ=0,θ= ±90◦)and therefore perpendicular to the plane they determine. The positive direction is determinedbyconsideringspecialcases,suchastheunitvectors ˆx׈y=ˆz(Cz=+AxBy). Themagnitudeis obtainedfrom (A×B)·(A×B)=A2B2−(A·B)2 =A2B2−A2B2cos2θ =A2B2sin2θ. (1.43) 9SeeSection 3.1 for abrief summary of determinants. 1 . 4 V e c t o ro rC r o s sP r o d u c t 2 1 Hence C=ABsinθ. (1.44) The first step in Eq. (1.43) may be verified by expanding out in component form, using Eqs. (1.38) for A×Band Eq. (1.24) for the dot product. From Eqs. (1.41), (1.42), and (1.44)weseetheequivalenceofEqs.(1.35)and(1.38),thetwodefinitionsofvectorprod- uct. There still remains the problem of verifying that C=A×Bis indeed a vector, that is, that it obeys Eq. (1.15), the vector transformation law. Starting in a rotated (primed system), C′ i=A′ jB′ k−A′kB′ j,i,j,andkincyclicorder , =summationdisplay lajlAlsummationdisplay makmBm−summationdisplay laklAlsummationdisplay majmBm =summationdisplay l,m(ajlakm−aklajm)AlBm. (1.45) Thecombinationofdirectioncosinesinparenthesesvanishesfor m=l.Wethereforehave jandktaking on fixed values, dependent on the choice of i, and six combinations of landm.I fi=3, thenj=1,k=2 (cyclic order), and we have the following direction cosinecombinations:10 a11a22−a21a12=a33, a13a21−a23a11=a32, a12a23−a22a13=a31(1.46) and their negatives. Equations (1.46) are identities satisfied by the direction cosines. They maybeverifiedwiththeuseofdeterminantsandmatrices(seeExercise3.3.3).Substituting backintoEq.(1.45), C′ 3=a33A1B2+a32A3B1+a31A2B3−a33A2B1−a32A1B3−a31A3B2 =a31C1+a32C2+a33C3 =summationdisplay na3nCn. (1.47) By permuting indices to pick up C′ 1andC′ 2, we see that Eq. (1.15) is satisfied and Cis indeed a vector. It should be mentioned here that this vector nature of thecross product isanaccidentassociatedwiththe three-dimensional natureofordinaryspace.11Itwillbe seeninChapter2thatthecrossproductmayalsobetreatedasasecond-rankantisymmetric tensor. 10Equations(1.46)holdforrotationsbecausetheypreservevolumes.Foramoregeneralorthogonaltransformation,ther.h.s.of Eqs. (1.46) is multiplied by the determinant ofthe transformation matrix (see Chapter 3 for matricesand determinants). 11SpecificallyEqs.(1.46)holdonlyforthree-dimensionalspace.SeeD.HestenesandG.Sobczyk, CliffordAlgebratoGeometric Calculus (Dordrecht: Reidel, 1984) for afar-reaching generalization ofthe cross product. 22 Chapter 1 Vector Analysis If we define a vector as an ordered triplet of numbers (or functions), as in the latter part ofSection1.2,thenthereisnoproblemidentifyingthecrossproductasavector.Thecross- product operation maps the two triples AandBinto a third triple, C, which by definition isavector. We now have two ways of multiplying vectors; a third form appears in Chapter 2. But what about division by a vector? It turns out that the ratio B/Ais not uniquely specified (Exercise 3.2.21) unless AandBare also required to be parallel. Hence division of one vectorbyanotheris notdefined. Exercises 1.4.1 Showthatthemediansofatriangleintersectinthecenter,whichis 2 /3ofthemedian’s lengthfromeachcorner.Constructanumericalexampleandplotit. 1.4.2 Provethelawofcosinesstartingfrom A2=(B−C)2. 1.4.3 Startingwith C=A+B,showthat C×C=0 leadsto A×B=−B×A. 1.4.4 Showthat (a)(A−B)·(A+B)=A2−B2, (b)(A−B)×(A+B)=2A×B. Thedistributivelawsneededhere, A·(B+C)=A·B+A·C, and A×(B+C)=A×B+A×C, mayeasilybeverified(ifdesired) byexpansioninCartesiancomponents. 1.4.5 Giventhethreevectors, P=3ˆx+2ˆy−ˆz, Q=−6ˆx−4ˆy+2ˆz, R=ˆx−2ˆy−ˆz, findtwo thatareperpendicularandtwothatareparallelor antiparallel. 1.4.6 IfP=ˆxPx+ˆyPyandQ=ˆxQx+ˆyQyare any two nonparallel (also nonantiparallel) vectorsinthe xy-plane,showthat P×Qisinthez-direction. 1.4.7 Provethat (A×B)·(A×B)=(AB)2−(A·B)2. 1 . 4 V e c t o ro rC r o s sP r o d u c t 2 3 1.4.8 Usingthevectors P=ˆxcosθ+ˆysinθ, Q=ˆxcosϕ−ˆysinϕ, R=ˆxcosϕ+ˆysinϕ, provethefamiliartrigonometricidentities sin(θ+ϕ)=sinθcosϕ+cosθsinϕ, cos(θ+ϕ)=cosθcosϕ−sinθsinϕ. 1.4.9 (a) Findavector Athatis perpendicularto U=2ˆx+ˆy−ˆz, V=ˆx−ˆy+ˆz. (b) What is Aif, in addition to this requirement, we demand that it have unit magni- tude? 1.4.10 If fourvectors a,b,c,anddalllieinthesameplane,showthat (a×b)×(c×d)=0. Hint.Considerthedirectionsofthecross-productvectors. 1.4.11 The coordinates of the three vertices of a triangle are (2,1,5),(5,2,8),and(4,8,2). Computeitsareabyvectormethods,itscenterandmedians.Lengthsareincentimeters. Hint.SeeExercise1.4.1. 1.4.12 The vertices of parallelogram ABCDare(1,0,0),(2,−1,0),(0,−1,1), and(−1,0,1) in order. Calculate the vector areas of triangle ABDand of triangle BCD.A r et h et w o vectorareasequal? ANS. Area ABD=−1 2(ˆx+ˆy+2ˆz). 1.4.13 The origin and the three vectors A,B, andC(all of which start at the origin) define a tetrahedron. Taking the outward direction as positive, calculate the total vector area of thefourtetrahedralsurfaces. Note.In Section1.11thisresult isgeneralizedtoanyclosedsurface. 1.4.14 Findthesidesandanglesof thesphericaltriangle ABCdefinedbythethreevectors A=(1,0,0), B=parenleftbigg1√ 2,0,1√ 2parenrightbigg , C=parenleftbigg 0,1√ 2,1√ 2parenrightbigg . Eachvectorstarts fromtheorigin(Fig. 1.14). 24 Chapter 1 Vector Analysis FIGURE 1.14Sphericaltriangle. 1.4.15 Derivethelawof sines(Fig.1.15): sinα |A|=sinβ |B|=sinγ |C|. 1.4.16 Themagneticinduction BisdefinedbytheLorentzforceequation, F=q(v×B). Carryingoutthreeexperiments,wefindthatif v=ˆx,F q=2ˆz−4ˆy, v=ˆy,F q=4ˆx−ˆz, v=ˆz,F q=ˆy−2ˆx. Fromtheresultsofthesethreeseparateexperimentscalculatethemagneticinduction B. 1.4.17 Define a cross product of two vectors in two-dimensional space and give a geometrical interpretationof yourconstruction. 1.4.18 Find the shortest distance between the paths of two rockets in free flight. Take the first rocket path to be r=r1+t1v1with launch at r1=(1,1,1)and velocity v1=(1,2,3) 1.5 Triple Scalar Product, Triple Vector Product 25 FIGURE 1.15Lawofsines. and the second rocket path as r=r2+t2v2withr2=(5,2,1)andv2=(−1,−1,1). Lengthsareinkilometers,velocitiesinkilometersperhour. 1.5 T RIPLE SCALAR PRODUCT ,TRIPLE VECTOR PRODUCT Triple Scalar Product Sections 1.3 and 1.4 cover the two types of multiplication of interest here. However, there arecombinationsofthreevectors, A·(B×C)andA×(B×C),thatoccurwithsufficient frequencytodeservefurtherattention.Thecombination A·(B×C) is known as the triple scalar product .B×Cyields a vector that, dotted into A,g i v e sa scalar.Wenotethat (A·B)×Crepresentsascalarcrossedintoavector,anoperationthat is not defined. Hence, if we agree to exclude this undefined interpretation, the parentheses maybeomittedandthetriplescalarproductwritten A·B×C. UsingEqs. (1.38)for thecross productandEq. (1.24) forthedotproduct,weobtain A·B×C=Ax(ByCz−BzCy)+Ay(BzCx−BxCz)+Az(BxCy−ByCx) =B·C×A=C·A×B =−A·C×B=−C·B×A=−B·A×C,andsoon . (1.48) There is a high degree of symmetry in the component expansion. Every term contains the factorsAi,Bj,andCk.Ifi,j,andkareincyclicorder (x,y,z),thesignispositive.Ifthe orderisanticyclic,thesignisnegative.Further,thedotandthecrossmaybeinterchanged, A·B×C=A×B·C. (1.49) 26 Chapter 1 Vector Analysis FIGURE 1.16Parallelepipedrepresentationof triplescalarproduct. A convenient representation of the component expansion of Eq. (1.48) is provided by the determinant A·B×C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleAxAyAz BxByBz CxCyCzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (1.50) The rules for interchanging rows and columns of a determinant12provide an immediate verification of the permutations listed in Eq. (1.48), whereas the symmetry of A,B, and Cin the determinant form suggests the relation given in Eq. (1.49). The triple products encounteredin Section 1.4, which showed that A×Bwas perpendicular to both AandB, werespecialcasesofthegeneralresult (Eq. (1.48)). Thetriplescalarproducthasadirectgeometricalinterpretation.Thethreevectors A,B, andCmaybeinterpretedasdefiningaparallelepiped(Fig.1.16): |B×C|=BCsinθ =areaof parallelogrambase. (1.51) The direction, of course, is normal to the base. Dotting Ainto this means multiplying the baseareabytheprojectionof Aontothenormal,or basetimesheight.Therefore A·B×C=volumeofparallelepipeddefinedby A,B,andC. The triple scalar product finds an interesting and important application in the construc- tion of a reciprocalcrystal lattice. Let a,b, andc(not necessarily mutuallyperpendicular) 12SeeSection 3.1 for a summary of the properties of determinants. 1.5 Triple Scalar Product, Triple Vector Product 27 represent the vectors that define a crystal lattice. The displacement from one lattice point toanothermaythenbewritten r=naa+nbb+ncc, (1.52) withna,nb, andnctakingonintegralvalues.Withthesevectorswemayform a′=b×c a·b×c,b′=c×a a·b×c,c′=a×b a·b×c. (1.53a) We see that a′is perpendicular to the plane containing bandc, and we can readily show that a′·a=b′·b=c′·c=1, (1.53b) whereas a′·b=a′·c=b′·a=b′·c=c′·a=c′·b=0. (1.53c) It is from Eqs. (1.53b) and (1.53c) that the name reciprocal lattice is associated with the pointsr′=n′ aa′+n′ bb′+n′ cc′.Themathematicalspaceinwhichthisreciprocallatticeex- istsissometimescalleda Fourierspace ,onthebasisofrelationstotheFourieranalysisof Chapters14and15.Thisreciprocallatticeisusefulinproblemsinvolvingthescatteringof wavesfromthevariousplanesinacrystal.FurtherdetailsmaybefoundinR.B.Leighton’s PrinciplesofModernPhysics , pp.440–448[NewYork:McGraw-Hill(1959)]. Triple Vector Product Thesecondtripleproductofinterestis A×(B×C),whichisavector.Heretheparentheses mustberetained,asmaybeseenfromaspecialcase (ˆx׈x)׈y=0,whileˆx×(ˆx׈y)= ˆx׈z=−ˆy. Example 1.5.1 ATRIPLE VECTOR PRODUCT Forthevectors A=ˆx+2ˆy−ˆz=(1,2,−1),B=ˆy+ˆz=(0,1,1),C=ˆx−ˆy=(0,1,1), B×C=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz 011 1−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=ˆx+ˆy−ˆz, and A×(B×C)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz 12−1 11−1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−ˆx−ˆz=−(ˆy+ˆz)−(ˆx−ˆy) =−B−C. /squaresolid ByrewritingtheresultinthelastlineofExample1.5.1asalinearcombinationof Band C, we notice that, taking a geometric approach, the triple vector product is perpendicular 28 Chapter 1 Vector Analysis FIGURE 1.17BandCar einthe xy-plane. B×Cis perpendiculartothe xy-planeand is shownherealongthe z-axis. Then A×(B×C)is perpendiculartothe z-axis andthereforeisbackinthe xy-plane. toAand toB×C.The plane defined by BandCis perpendicular to B×C, and so the tripleproductliesinthisplane(see Fig.1.17): A×(B×C)=uB+vC. (1.54) Taking the scalar product of Eq. (1.54) with Agives zero for the left-hand side, so uA·B+vA·C=0. Hence u=wA·Candv=−wA·Bfor a suitable w. Substitut- ingthesevaluesintoEq. (1.54) gives A×(B×C)=wbracketleftbig B(A·C)−C(A·B)bracketrightbig ; (1.55) wewanttoshowthat w=1 in Eq. (1.55), an important relation sometimes known as the BAC–CAB rule. Since Eq. (1.55) is linear in A,B, andC,wis independent of these magnitudes. That is, we only need to show that w=1 for unit vectors ˆA,ˆB,ˆC. Let us denote ˆB·ˆC=cosα, ˆC·ˆA=cosβ,ˆA·ˆB=cosγ, andsquareEq.(1.55) toobtain bracketleftbigˆA×(ˆB׈C)bracketrightbig2=ˆA2(ˆB׈C)2−bracketleftbigˆA·(ˆB׈C)bracketrightbig2 =1−cos2α−bracketleftbigˆA·(ˆB׈C)bracketrightbig2 =w2bracketleftbig (ˆA·ˆC)2+(ˆA·ˆB)2−2(ˆA·ˆB)(ˆA·ˆC)(ˆB·ˆC)bracketrightbig =w2parenleftbig cos2β+cos2γ−2cosαcosβcosγparenrightbig , (1.56) 1.5 Triple Scalar Product, Triple Vector Product 29 using(ˆA׈B)2=ˆA2ˆB2−(ˆA·ˆB)2repeatedly (see Eq. (1.43) for a proof). Consequently, the(squared) volumespannedby ˆA,ˆB,ˆCthatoccursinEq. (1.56) canbewrittenas bracketleftbigˆA·(ˆB׈C)bracketrightbig2=1−cos2α−w2parenleftbig cos2β+cos2γ−2cosαcosβcosγparenrightbig . Herew2=1, since this volume is symmetric in α,β,γ. That is, w=±1 and is inde- pendent ofˆA,ˆB,ˆC. Using again the special case ˆx×(ˆx׈y)=−ˆyin Eq. (1.55) finally givesw=1.(AnalternatederivationusingtheLevi-Civitasymbol εijkofChapter2isthe topicofExercise2.9.8.) Itmightbenotedherethatjustasvectorsareindependentofthecoordinates,soavector equation is independent of the particular coordinate system. The coordinate system only determines the components. If the vector equation can be established in Cartesian coor- dinates, it is established and valid in any of the coordinate systems to be introduced in Chapter 2. Thus, Eq. (1.55) may be verified by a direct though not very elegant method of expandingintoCartesiancomponents(see Exercise1.5.2). Exercises 1.5.1 One vertex of a glass parallelepiped is at the origin (Fig. 1.18). The three adjacent vertices are at (3,0,0),(0,0,2), and(0,3,1). All lengths are in centimeters. Calculate the number of cubic centimeters of glass in the parallelepiped using the triple scalar product. 1.5.2 Verifytheexpansionofthetriplevectorproduct A×(B×C)=B(A·C)−C(A·B) FIGURE 1.18Parallelepiped:triplescalarproduct. 30 Chapter 1 Vector Analysis bydirectexpansioninCartesiancoordinates. 1.5.3 ShowthatthefirststepinEq. (1.43), whichis (A×B)·(A×B)=A2B2−(A·B)2, isconsistentwiththe BAC–CABrulefor atriplevectorproduct. 1.5.4 Youare giventhethreevectors A,B,andC, A=ˆx+ˆy, B=ˆy+ˆz, C=ˆx−ˆz. (a) Compute the triple scalar product, A·B×C. Noting that A=B+C, give a geo- metricinterpretationofyourresultfor thetriplescalarproduct. (b) Compute A×(B×C). 1.5.5 The orbital angular momentum Lof a particle is given by L=r×p=mr×v, where pis the linear momentum. With linear and angular velocity related by v=ω×r,s h o w that L=mr2bracketleftbig ω−ˆr(ˆr·ω)bracketrightbig . Hereˆris a unit vector in the r-direction. For r·ω=0 this reduces to L=Iω, with the moment of inertia Igiven by mr2. In Section 3.5 this result is generalized to form an inertiatensor. 1.5.6 Thekineticenergyofasingleparticleisgivenby T=1 2mv2.Forrotationalmotionthis becomes1 2m(ω×r)2. Showthat T=1 2mbracketleftbig r2ω2−(r·ω)2bracketrightbig . Forr·ω=0 thisreducesto T=1 2Iω2, withthemomentof inertia Igivenby mr2. 1.5.7 Showthat13 a×(b×c)+b×(c×a)+c×(a×b)=0. 1.5.8 A vector Ais decomposed into a radial vector Arand a tangential vector At.I fˆris a unitvectorintheradialdirection,showthat (a)Ar=ˆr(A·ˆr)and (b)At=−ˆr×(ˆr×A). 1.5.9 Prove that a necessary and sufficient condition for the three (nonvanishing) vectors A, B,andCtobecoplanaristhevanishingofthetriplescalarproduct A·B×C=0. 13This is Jacobi’s identity for vector products; for commutators it is important in the context of Lie algebras (see Eq. (4.16) in Section 4.2). 1.5 Triple Scalar Product, Triple Vector Product 31 1.5.10 Threevectors A,B, andCare givenby A=3ˆx−2ˆy+2ˆz, B=6ˆx+4ˆy−2ˆz, C=−3ˆx−2ˆy−4ˆz. Computethevaluesof A·B×CandA×(B×C),C×(A×B)andB×(C×A). 1.5.11 VectorDisa linearcombinationofthreenoncoplanar(andnonorthogonal)vectors: D=aA+bB+cC. Showthatthecoefficientsaregivenbyaratiooftriplescalarproducts, a=D·B×C A·B×C,andso on. 1.5.12 Showthat (A×B)·(C×D)=(A·C)(B·D)−(A·D)(B·C). 1.5.13 Showthat (A×B)×(C×D)=(A·B×D)C−(A·B×C)D. 1.5.14 For aspherical trianglesuchaspicturedinFig.1.14showthat sinA sinBC=sinB sinCA=sinC sinAB. Here sinAis the sine of the included angle at A, whileBCis the side opposite (in radians). 1.5.15 Given a′=b×c a·b×c,b′=c×a a·b×c,c′=a×b a·b×c, anda·b×c/negationslash=0,showthat (a)x·y′=δxy,(x,y=a,b,c), (b)a′·b′×c′=(a·b×c)−1, (c)a=b′×c′ a′·b′×c′. 1.5.16 Ifx·y′=δxy,(x,y=a,b,c), prove that a′=b×c a·b×c. (This istheconverseofProblem1.5.15.) 1.5.17 Showthatanyvector Vmaybeexpressedintermsofthereciprocalvectors a′,b′,c′(of Problem1.5.15)by V=(V·a)a′+(V·b)b′+(V·c)c′. 32 Chapter 1 Vector Analysis 1.5.18 An electric charge q1moving with velocity v1produces a magnetic induction Bgiven by B=µ0 4πq1v1׈r r2(mksunits), whereˆrpointsfrom q1tothepointatwhich Bis measured(BiotandSavartlaw). (a) Show that the magnetic force on a second charge q2, velocity v2, is given by the triplevectorproduct F2=µ0 4πq1q2 r2v2×(v1׈r). (b) Write out the corresponding magnetic force F1thatq2exerts on q1. Define your unitradialvector.Howdo F1andF2compare? (c) Calculate F1andF2for the case of q1andq2moving along parallel trajectories sidebyside. ANS. (b)F1=−µ0 4πq1q2 r2v1×(v2׈r). Ingeneral,thereisnosimplerelationbetween F1andF2.Specifically,Newton’sthirdlaw, F1=−F2, doesnothold. (c)F1=µ0 4πq1q2 r2v2ˆr=−F2. Mutualattraction. 1.6 G RADIENT ,∇ To provide a motivation for the vector nature of partial derivatives, we now introduce the totalvariationof afunction F(x,y), dF=∂F ∂xdx+∂F ∂ydy. It consists of independent variations in the x- andy-directions. We write dFas a sum of twoincrements,onepurelyinthe x- andtheotherinthe y-direction, dF(x,y)≡F(x+dx,y+dy)−F(x,y) =bracketleftbig F(x+dx,y+dy)−F(x,y+dy)bracketrightbig +bracketleftbig F(x,y+dy)−F(x,y)bracketrightbig =∂F ∂xdx+∂F ∂ydy, byaddingandsubtracting F(x,y+dy).Themeanvaluetheorem(thatis,continuityof F) tellsusthathere ∂F/∂x,∂F/∂yareevaluatedatsomepoint ξ,ηbetweenxandx+dx,y 1.6 Gradient, ∇ 33 andy+dy, respectively. As dx→0 anddy→0,ξ→xandη→y. This result general- izestothreeandhigherdimensions.Forexample,for afunction ϕof threevariables, dϕ(x,y,z)≡bracketleftbig ϕ(x+dx,y+dy,z+dz)−ϕ(x,y+dy,z+dz)bracketrightbig +bracketleftbig ϕ(x,y+dy,z+dz)−ϕ(x,y,z+dz)bracketrightbig +bracketleftbig ϕ(x,y,z+dz)−ϕ(x,y,z)bracketrightbig (1.57) =∂ϕ ∂xdx+∂ϕ ∂ydy+∂ϕ ∂zdz. Algebraically, dϕinthetotalvariationisascalarproductofthechangeinposition drand thedirectional change of ϕ. And now we are ready to recognize the three-dimensional partialderivativeasa vector,whichleadsustotheconceptofgradient. Supposethat ϕ(x,y,z) isascalarpointfunction,thatis,afunctionwhosevaluedepends onthevaluesofthecoordinates (x,y,z).Asascalar,itmusthavethesamevalueatagiven fixedpointinspace,independentoftherotationof ourcoordinatesystem, or ϕ′(x′ 1,x′ 2,x′ 3)=ϕ(x1,x2,x3). (1.58) Bydifferentiatingwithrespectto x′ iweobtain ∂ϕ′(x′ 1,x′ 2,x′ 3) ∂x′ i=∂ϕ(x1,x2,x3) ∂x′ i=summationdisplay j∂ϕ ∂xj∂xj ∂x′ i=summationdisplay jaij∂ϕ ∂xj(1.59) by the rules of partial differentiation and Eqs. (1.16a) and (1.16b). But comparison with Eq. (1.17), the vector transformation law, now shows that we have constructed a vector withcomponents ∂ϕ/∂xj.This vectorwelabelthegradientof ϕ. Aconvenientsymbolismis ∇ϕ=ˆx∂ϕ ∂x+ˆy∂ϕ ∂y+ˆz∂ϕ ∂z(1.60) or ∇=ˆx∂ ∂x+ˆy∂ ∂y+ˆz∂ ∂z. (1.61) ∇ϕ(or delϕ) is our gradient of the scalar ϕ, whereas ∇(del) itself is a vector differential operator (available to operate on or to differentiate a scalar ϕ). All the relationships for ∇ (del) can be derived from the hybrid nature of del in terms of both the partial derivatives andits vectornature. Thegradientofascalarisextremelyimportantinphysicsandengineeringinexpressing therelationbetweenaforce fieldandapotentialfield, forceF=−∇(potential V), (1.62) which holds for both gravitational and electrostatic fields, among others. Note that the minussigninEq.(1 .62)resultsinwaterflowingdownhillratherthanuphill!Ifaforcecan be described, as in Eq. (1.62), by a single function V(r)everywhere, we call the scalar functionVitspotential .Becausetheforceisthedirectionalderivativeofthepotential,we canfindthepotential,ifitexists,byintegratingtheforcealongasuitablepath.Becausethe 34 Chapter 1 Vector Analysis total variation dV=∇V·dr=−F·dris the work done against the force along the path dr,we recognize the physical meaning of the potential (difference) as work and energy. Moreover,inasumof pathincrementstheintermediatepointscancel, bracketleftbig V(r+dr1+dr2)−V(r+dr1)bracketrightbig +bracketleftbig V(r+dr1)−V(r)bracketrightbig =V(r+dr2+dr1)−V(r), sotheintegratedworkalongsomepathfromaninitialpoint ritoafinalpoint risgivenby the potential difference V(r)−V(ri)at the endpoints of the path. Therefore, such forces areespeciallysimpleandwellbehaved:Theyarecalled conservative .Whenthereislossof energyduetofrictionalongthepathorsomeotherdissipation,theworkwilldependonthe path,andsuchforces cannotbeconservative:Nopotentialexists. We discuss conservative forcesinmoredetailinSection1.13. Example 1.6.1 THEGRADIENT OF A POTENTIAL V(r) Letuscalculatethegradientof V(r)=V(radicalbig x2+y2+z2),so ∇V(r)=ˆx∂V(r) ∂x+ˆy∂V(r) ∂y+ˆz∂V(r) ∂z. Now,V(r)dependson xthroughthedependenceof ronx.Therefore14 ∂V(r) ∂x=dV(r) dr·∂r ∂x. Fromras afunctionof x,y,z, ∂r ∂x=∂(x2+y2+z2)1/2 ∂x=x (x2+y2+z2)1/2=x r. Therefore ∂V(r) ∂x=dV(r) dr·x r. Permutingcoordinates (x→y,y→z,z→x)toobtainthe yandzderivatives,weget ∇V(r)=(ˆxx+ˆyy+ˆzz)1 rdV dr =r rdV dr=ˆrdV dr. Hereˆris a unit vector (r/r)in thepositiveradial direction. The gradient of a function of ris a vector in the (positive or negative) radial direction. In Section 2.5, ˆris seen as one ofthethreeorthonormalunitvectorsofsphericalpolarcoordinatesand ˆr∂/∂rastheradial componentof ∇. /squaresolid 14This is aspecial caseof the chain rule of partial differentiation: ∂V(r,θ,ϕ) ∂x=∂V ∂r∂r ∂x+∂V ∂θ∂θ ∂x+∂V ∂ϕ∂ϕ ∂x, where∂V/∂θ=∂V/∂ϕ=0,∂V/∂r→dV/dr. 1.6 Gradient, ∇ 35 A Geometrical Interpretation Oneimmediateapplicationof ∇ϕistodotitintoanincrementof length dr=ˆxdx+ˆydy+ˆzdz. Thusweobtain ∇ϕ·dr=∂ϕ ∂xdx+∂ϕ ∂ydy+∂ϕ ∂zdz=dϕ, thechangeinthescalarfunction ϕcorrespondingtoachangeinposition dr.Nowconsider PandQtobetwopointsonasurface ϕ(x,y,z)=C,aconstant.Thesepointsarechosen sothatQisadistance drfromP.Then,movingfrom PtoQ,thechangein ϕ(x,y,z)=C isgivenby dϕ=(∇ϕ)·dr=0 (1.63) since we stay on the surface ϕ(x,y,z)=C. This shows that ∇ϕis perpendicular to dr. Sincedrm a yh a v ea n yd i r e c t i o nf r o m Pas long as it stays in the surface of constant ϕ, pointQbeingrestrictedtothesurfacebuthavingarbitrarydirection, ∇ϕisseenasnormal tothesurface ϕ=constant(Fig. 1.19). If we now permit drto take us from one surface ϕ=C1to an adjacent surface ϕ=C2 (Fig.1.20), dϕ=C1−C2=/Delta1C=(∇ϕ)·dr. (1.64) For a given dϕ,|dr|is a minimum when it is chosen parallel to ∇ϕ(cosθ=1);o r ,f o r ag i v e n|dr|, the change in the scalar function ϕis maximized by choosing drparallel to FIGURE 1.19Thelengthincrement drhastostayonthesurface ϕ=C. 36 Chapter 1 Vector Analysis FIGURE 1.20Gradient. ∇ϕ.This identifies ∇ϕas a vector having the direction of the maximum space rate of change of ϕ, an identification that will be useful in Chapter 2 when we consider non- Cartesian coordinate systems. This identification of ∇ϕmay also be developed by using thecalculusofvariationssubjecttoaconstraint,Exercise17.6.9. Example 1.6.2 FORCE AS GRADIENT OF A POTENTIAL Asaspecificexampleoftheforegoing,andasanextensionofExample1.6.1,weconsider thesurfaces consistingofconcentricsphericalshells, Fig.1.21.Wehave ϕ(x,y,z)=parenleftbig x2+y2+z2parenrightbig1/2=r=C, whereristheradius,equalto C,ourconstant. /Delta1C=/Delta1ϕ=/Delta1r,thedistancebetweentwo shells.FromExample1.6.1 ∇ϕ(r)=ˆrdϕ(r) dr=ˆr. Thegradientisintheradialdirectionandis normaltothesphericalsurface ϕ=C./squaresolid Example 1.6.3 INTEGRATION BY PARTS OF GRADIENT Letusprovetheformulaintegraltext A(r)·∇f(r)d3r=−integraltext f(r)∇·A(r)d3r,whereAorforboth vanishatinfinitysothattheintegratedpartsvanish.Thisconditionissatisfiedif,forexam- ple,Aistheelectromagneticvectorpotentialand fis abound-statewavefunction ψ(r). 1.6 Gradient, ∇ 37 FIGURE 1.21Gradientfor ϕ(x,y,z)=(x2+y2+z2)1/2,spherical shells:(x2 2+y2 2+z2 2)1/2=r2=C2, (x2 1+y2 1+z2 1)1/2=r1=C1. Writing the inner product in Cartesian coordinates, integrating each one-dimensional integralbyparts, anddroppingtheintegratedterms,weobtain integraldisplay A(r)·∇f(r)d3r=integraldisplayintegraldisplaybracketleftbigg Axf|∞ x=−∞−integraldisplay f∂Ax ∂xdxbracketrightbigg dydz+··· =−integraldisplayintegraldisplayintegraldisplay f∂Ax ∂xdxdydz−integraldisplayintegraldisplayintegraldisplay f∂Ay ∂ydydxdz−integraldisplayintegraldisplayintegraldisplay f∂Az ∂zdzdxdy =−integraldisplay f(r)∇·A(r)d3r. IfA=eikzˆedescribesanoutgoingphotoninthedirectionoftheconstantpolarizationunit vectorˆeandf=ψ(r)is anexponentiallydecayingbound-statewavefunction,then integraldisplay eikzˆe·∇ψ(r)d3r=−ezintegraldisplay ψ(r)deikz dzd3r=−ikezintegraldisplay ψ(r)eikzd3r, becauseonlythe z-componentof thegradientcontributes. /squaresolid Exercises 1.6.1 IfS(x,y,z)=(x2+y2+z2)−3/2,find (a)∇Satthepoint (1,2,3); (b) themagnitudeofthegradientof S,|∇S|at(1,2,3);and (c) thedirectioncosinesof ∇Sat(1,2,3). 38 Chapter 1 Vector Analysis 1.6.2 (a) Findaunitvectorperpendiculartothesurface x2+y2+z2=3 atthepoint (1,1,1). Lengthsare incentimeters. (b) Derivetheequationof theplanetangenttothesurface at (1,1,1). ANS.(a) (ˆx+ˆy+ˆz)/√ 3,(b)x+y+z=3. 1.6.3 Given a vector r12=ˆx(x1−x2)+ˆy(y1−y2)+ˆz(z1−z2), show that ∇1r12(gradient with respect to x1,y1, andz1of the magnitude r12) is a unit vector in the direction of r12. 1.6.4 Ifavectorfunction Fdependsonbothspacecoordinates (x,y,z)andtime t,showthat dF=(dr·∇)F+∂F ∂tdt. 1.6.5 Show that ∇(uv)=v∇u+u∇v, whereuandvare differentiable scalar functions of x,y,andz. (a) Show that a necessary and sufficient condition that u(x,y,z) andv(x,y,z) are relatedbysomefunction f(u,v)=0 isthat(∇u)×(∇v)=0. (b) Ifu=u(x,y)andv=v(x,y), show thatthe condition (∇u)×(∇v)=0 leadsto thetwo-dimensionalJacobian Jparenleftbiggu,v x,yparenrightbigg =vextendsinglevextendsinglevextendsinglevextendsingle∂u ∂x∂u ∂y ∂v ∂x∂v ∂yvextendsinglevextendsinglevextendsinglevextendsingle=0. Thefunctions uandvare assumeddifferentiable. 1.7 D IVERGENCE ,∇ Differentiating a vector function is a simple extension of differentiating scalar quantities. Supposer(t)describes the position of a satellite at some time t. Then, for differentiation withrespecttotime, dr(t) dt=lim /Delta1→0r(t+/Delta1t)−r(t) /Delta1t=v,linearvelocity. Graphically,weagainhavetheslopeofa curve,orbit, ortrajectory,asshowninFig.1.22. If we resolve r(t)into its Cartesian components, dr/dtalways reduces directly to a vectorsumofnotmorethanthree(forthree-dimensionalspace)scalarderivatives.Inother coordinate systems (Chapter 2) the situation is more complicated, for the unit vectors are no longer constant in direction. Differentiation with respect to the space coordinates is handled in the same way as differentiation with respect to time, as seen in the following paragraphs. 1.7 Divergence, ∇ 39 FIGURE 1.22Differentiationof avector. In Section 1.6, ∇was defined as a vector operator. Now, paying attention to both its vector and its differential properties, we let it operate on a vector. First, as a vector we dot itintoasecondvectortoobtain ∇·V=∂Vx ∂x+∂Vy ∂y+∂Vz ∂z, (1.65a) knownasthedivergenceof V.This isascalar, asdiscussedinSection1.3. Example 1.7.1 DIVERGENCE OF COORDINATE VECTOR Calculate ∇·r: ∇·r=parenleftbigg ˆx∂ ∂x+ˆy∂ ∂y+ˆz∂ ∂zparenrightbigg ·(ˆxx+ˆyy+ˆzz) =∂x ∂x+∂y ∂y+∂z ∂z, or∇·r=3. /squaresolid Example 1.7.2 DIVERGENCE OF CENTRAL FORCE FIELD GeneralizingExample1.7.1, ∇·parenleftbig rf(r)parenrightbig =∂ ∂xbracketleftbig xf(r)bracketrightbig +∂ ∂ybracketleftbig yf(r)bracketrightbig +∂ ∂zbracketleftbig zf(r)bracketrightbig =3f(r)+x2 rdf dr+y2 rdf dr+z2 rdf dr =3f(r)+rdf dr. 40 Chapter 1 Vector Analysis ThemanipulationofthepartialderivativesleadingtothesecondequationinExample1.7.2 isdiscussedinExample1.6.1. Inparticular,if f(r)=rn−1, ∇·parenleftbig rrn−1parenrightbig =∇·ˆrrn =3rn−1+(n−1)rn−1 =(n+2)rn−1. (1.65b) Thisdivergencevanishesfor n=−2,exceptat r=0,animportantfactinSection1.14. /squaresolid Example 1.7.3 INTEGRATION BY PARTS OF DIVERGENCE Let us prove the formulaintegraltext f(r)∇·A(r)d3r=−integraltext A·∇fd3r,whereAorfor both vanishatinfinity. To show this, we proceed, as in Example 1.6.3, by integration by parts after writing the inner product in Cartesian coordinates. Because the integrated terms are evaluated at infinity,wheretheyvanish,weobtain integraldisplay f(r)∇·A(r)d3r=integraldisplay fparenleftbigg∂Ax ∂xdxdydz+∂Ay ∂ydydxdz+∂Az ∂zdzdxdyparenrightbigg =−integraldisplayparenleftbigg Ax∂f ∂xdxdydz+Ay∂f ∂ydydxdz+Az∂f ∂zdzdxdyparenrightbigg =−integraldisplay A·∇fd3r./squaresolid A Physical Interpretation Todevelopafeelingforthephysicalsignificanceofthedivergence,consider ∇·(ρv)with v(x,y,z),thevelocityofacompressiblefluid,and ρ(x,y,z) ,itsdensityatpoint (x,y,z). Ifweconsiderasmallvolume dxdydz (Fig.1.23)at x=y=z=0,thefluidflowinginto this volume per unit time (positive x-direction) through the face EFGHis (rate of flow in)EFGH=ρvx|x=0=dydz. The components of the flow ρvyandρvztangential to this face contribute nothing to the flow through this face. The rate of flow out (still positive x-direction)throughface ABCDisρvx|x=dxdydz.Tocomparetheseflowsandtofindthe netflowout,weexpandthislastresult,likethetotalvariationinSection1.6.15Thisyields (rate offlowout )ABCD=ρvx|x=dxdydz =bracketleftbigg ρvx+∂ ∂x(ρvx)dxbracketrightbigg x=0dydz. Herethederivativetermisafirstcorrectionterm,allowingforthepossibilityofnonuniform densityorvelocityorboth.16Thezero-orderterm ρvx|x=0(correspondingtouniformflow) 15Herewehave the increment dxandweshow apartial derivative withrespectto xsinceρvxm ayal s od ep en do n yandz. 16Strictlyspeaking, ρvxisaveragedoverface EFGHandtheexpression ρvx+(∂/∂x)(ρv x)dxissimilarlyaveragedoverface ABCD.Using anarbitrarily small differential volume, wefindthat the averages reduce tothe valuesemployed here. 1.7 Divergence, ∇ 41 FIGURE 1.23Differentialrectangularparallelepiped(in firstoctant). cancelsout: Netrateof flowout |x=∂ ∂x(ρvx)dxdydz. Equivalently,wecanarriveatthisresult by lim /Delta1x→0ρvx(/Delta1x,0,0)−ρvx(0,0,0) /Delta1x≡∂[ρvx(x,y,z)] ∂xvextendsinglevextendsinglevextendsinglevextendsingle 0,0,0. Now,the x-axisisnotentitledtoanypreferredtreatment.Theprecedingresultforthetwo faces perpendicular to the x-axis must hold for the two faces perpendicular to the y-axis, withxreplaced by yand the corresponding changes for yandz:y→z,z→x.T h i si s a cyclic permutation of the coordinates. A further cyclic permutation yields the result for the remaining two faces of our parallelepiped. Adding the net rate of flow out for all three pairsofsurfaces of ourvolumeelement,wehave netflowout (per unittime)=bracketleftbigg∂ ∂x(ρvx)+∂ ∂y(ρvy)+∂ ∂z(ρvz)bracketrightbigg dxdydz =∇·(ρv)dxdydz. (1.66) Therefore the net flow of our compressible fluid out of the volume element dxdydz per unit volume per unit time is ∇·(ρv). Hence the name divergence . A direct application is inthecontinuityequation ∂ρ ∂t+∇·(ρv)=0, (1.67a) which states that a net flow out of the volume results in a decreased density inside the volume. Note that in Eq. (1.67a), ρis considered to be a possible function of time as well as of space: ρ(x,y,z,t) . The divergence appears in a wide variety of physical problems, 42 Chapter 1 Vector Analysis ranging from a probability current density in quantum mechanics to neutron leakage in a nuclearreactor. The combination ∇·(fV), in which fis a scalar function and Vis a vector function, maybewritten ∇·(fV)=∂ ∂x(fVx)+∂ ∂y(fVy)+∂ ∂z(fVz) =∂f ∂xVx+f∂Vx ∂x+∂f ∂yVy+f∂Vy ∂y+∂f ∂zVz+f∂Vz ∂z =(∇f)·V+f∇·V, (1.67b) which is just what we would expect for the derivative of a product. Notice that ∇as a differential operator differentiates both fandV; as a vector it is dotted into V(in each term). If wehavethespecialcaseofthedivergenceof avectorvanishing, ∇·B=0, (1.68) the vector Bis said to be solenoidal , the term coming from the example in which Bis the magnetic induction and Eq. (1.68) appears as one of Maxwell’s equations. When a vector issolenoidal,itmaybewrittenas thecurlofanothervectorknownasthevectorpotential. (InSection1.13weshallcalculatesuchavectorpotential.) Exercises 1.7.1 Foraparticlemovinginacircularorbit r=ˆxrcosωt+ˆyrsinωt, (a) evaluate r×˙r,with˙r=dr dt=v. (b) Showthat ¨r+ω2r=0 with¨r=dv dt. Theradius randtheangularvelocity ωareconstant. ANS.(a)ˆzωr2. 1.7.2 VectorAsatisfies the vector transformation law, Eq. (1.15). Show directly that its time derivative dA/dtalsosatisfiesEq. (1.15) andis thereforeavector. 1.7.3 Show,bydifferentiatingcomponents,that (a)d dt(A·B)=dA dt·B+A·dB dt, (b)d dt(A×B)=dA dt×B+A×dB dt, justlikethederivativeoftheproductof twoalgebraicfunctions. 1.7.4 InChapter2itwillbeseenthattheunitvectorsinnon-Cartesiancoordinatesystemsare usuallyfunctionsof the coordinatevariables, ei=ei(q1,q2,q3)but|ei|=1. Showthat either∂ei/∂qj=0o r∂ei/∂qjisorthogonalto ei. Hint.∂e2 i/∂qj=0. 1.8 Curl, ∇× 43 1.7.5 Prove ∇·(a×b)=b·(∇×a)−a·(∇×b). Hint.Treatasatriplescalarproduct. 1.7.6 Theelectrostaticfieldofa pointcharge qis E=q 4πε0·ˆr r2. Calculatethedivergenceof E.Whathappensattheorigin? 1.8 C URL ,∇× Anotherpossibleoperationwiththevectoroperator ∇istocrossitintoavector.Weobtain ∇×V=ˆxparenleftbigg∂ ∂yVz−∂ ∂zVyparenrightbigg +ˆyparenleftbigg∂ ∂zVx−∂ ∂xVzparenrightbigg +ˆzparenleftbigg∂ ∂xVy−∂ ∂yVxparenrightbigg =vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz ∂ ∂x∂ ∂y∂ ∂z VxVyVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (1.69) whichiscalledthe curlofV.Inexpandingthisdeterminantwemustconsiderthederivative nature of ∇. Specifically, V×∇is defined only as an operator, another vector differential operator. It is certainly not equal, in general, to −∇×V.17In the case of Eq. (1.69) the determinantmustbeexpanded fromthetopdown sothatwegetthederivativesasshown inthemiddleportionofEq.(1.69).If ∇iscrossedintotheproductofascalarandavector, wecanshow ∇×(fV)|x=bracketleftbigg∂ ∂y(fVz)−∂ ∂z(fVy)bracketrightbigg =parenleftbigg f∂Vz ∂y+∂f ∂yVz−f∂Vy ∂z−∂f ∂zVyparenrightbigg =f∇×V|x+(∇f)×V|x. (1.70) If we permute the coordinates x→y,y→z,z→xto pick up the y-component and thenpermutethemasecondtimetopickupthe z-component,then ∇×(fV)=f∇×V+(∇f)×V, (1.71) which is the vector product analog of Eq. (1.67b). Again, as a differential operator ∇ differentiatesboth fandV.Asavectoritiscrossedinto V(ineachterm). 17In this same spirit, if Ais a differential operator, it is not necessarily true that A×A=0. Specifically, for the quantum mechanicalangular momentum operator L=−i(r×∇),wefin dth at L×L=iL.SeeSections 4.3 and 4.4 for more details. 44 Chapter 1 Vector Analysis Example 1.8.1 VECTOR POTENTIAL OF A CONSTANT BFIELD Fromelectrodynamicsweknowthat ∇·B=0,whichhasthegeneralsolution B=∇×A, whereA(r)iscalledthevectorpotential(ofthemagneticinduction),because ∇·(∇×A)= (∇×∇)·A≡0,asatriplescalarproductwithtwoidenticalvectors.Thislastidentitywill not change if we add the gradient of some scalar function to the vector potential, which, therefore,isnotunique. In ourcase,wewanttoshowthatavectorpotentialis A=1 2(B×r). Usingthe BAC–BACruleinconjunctionwithExample1.7.1, wefindthat 2∇×A=∇×(B×r)=(∇·r)B−(B·∇)r=3B−B=2B, whereweindicatebytheorderingofthescalarproductofthesecondtermthatthegradient stillactsonthecoordinatevector. /squaresolid Example 1.8.2 CURL OF A CENTRAL FORCE FIELD Calculate ∇×(rf(r)). ByEq. (1.71), ∇×parenleftbig rf(r)parenrightbig =f(r)∇×r+bracketleftbig ∇f(r)bracketrightbig ×r. (1.72) First, ∇×r=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz ∂ ∂x∂ ∂y∂ ∂z xyzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (1.73) Second,using ∇f(r)=ˆr(df/dr)(Example1.6.1), weobtain ∇×rf(r)=df drˆr×r=0. (1.74) Thisvectorproductvanishes,since r=ˆrrandˆr׈r=0. /squaresolid To develop a better feeling for the physical significance of the curl, we consider the circulationof fluidaroundadifferentialloopinthe xy-plane,Fig.1.24. FIGURE 1.24Circulationaroundadifferentialloop. 1.8 Curl, ∇× 45 Although the circulation is technically given by a vector line integralintegraltext V·dλ(Sec- tion 1.10), we can set up the equivalent scalar integrals here. Let us take the circulation to be circulation 1234=integraldisplay 1Vx(x,y)dλ x+integraldisplay 2Vy(x,y)dλ y +integraldisplay 3Vx(x,y)dλ x+integraldisplay 4Vy(x,y)dλ y. (1.75) The numbers 1, 2, 3, and 4 refer to the numbered line segments in Fig. 1.24. In the first integral,dλx=+dx; but in the third integral, dλx=−dxbecause the third line segment istraversedinthenegative x-direction.Similarly, dλy=+dyforthesecondintegral, −dy for the fourth. Next, the integrands are referred to the point (x0,y0)with a Taylor expan- sion18taking into accountthe displacementof line segment3 from 1 and that of 2 from 4. Forourdifferentiallinesegmentsthisleadsto circulation 1234=Vx(x0,y0)dx+bracketleftbigg Vy(x0,y0)+∂Vy ∂xdxbracketrightbigg dy +bracketleftbigg Vx(x0,y0)+∂Vx ∂ydybracketrightbigg (−dx)+Vy(x0,y0)(−dy) =parenleftbigg∂Vy ∂x−∂Vx ∂yparenrightbigg dxdy. (1.76) Dividingby dxdy,weha v e circulationperunitarea =∇×V|z. (1.77) The circulation19about our differential area in the xy-plane is given by the z-component of∇×V. In principle, the curl ∇×Vat(x0,y0)could be determined by inserting a (differential) paddlewheel intothe movingfluid at point (x0,y0). The rotationof the little paddle wheel would be a measure of the curl, and its axis would be along the direction of ∇×V,whichisperpendiculartotheplaneofcirculation. Weshallusetheresult,Eq.(1.76),inSection1.12toderiveStokes’theorem.Whenever thecurlof avector Vvanishes, ∇×V=0, (1.78) Vislabeled irrotational .Themostimportantphysicalexamplesofirrotationalvectorsare thegravitationalandelectrostaticforces. Ineachcase V=Cˆr r2=Cr r3, (1.79) whereCis a constant and ˆris the unit vector in the outward radial direction. For the gravitationalcasewehave C=−Gm1m2,givenbyNewton’slawofuniversalgravitation. IfC=q1q2/4πε0, we have Coulomb’s law of electrostatics (mks units). The force V 18Here,Vy(x0+dx,y0)=Vy(x0,y0)+(∂Vy ∂x)x0y0dx+···.The higher-order terms will drop out in the limit as dx→0. Acorrection termfor the variation of Vywithyis canceledby the corresponding term in the fourth integral. 19In fluid dynamics ∇×Vis calledthe “vorticity.” 46 Chapter 1 Vector Analysis given in Eq. (1.79) may be shown to be irrotational by direct expansion into Cartesian components, as we did in Example 1.8.1. Another approach is developed in Chapter 2, in whichweexpress ∇×,thecurl,intermsofsphericalpolarcoordinates.InSection1.13we shall see that whenever a vector is irrotational, the vector may be written as the (negative) gradient of a scalar potential. In Section 1.16 we shall prove that a vector field may be resolved into an irrotational part and a solenoidal part (subject to conditions at infinity). In terms of the electromagnetic field this corresponds to the resolution into an irrotational electricfieldandasolenoidalmagneticfield. For waves in an elastic medium, if the displacement uis irrotational, ∇×u=0, plane waves (or spherical waves at large distances) become longitudinal. If uis solenoidal, ∇·u=0, then the waves become transverse. A seismic disturbance will produce a dis- placement that may be resolved into a solenoidal part and an irrotational part (compare Section 1.16). The irrotational part yields the longitudinal P(primary) earthquake waves. Thesolenoidalpartgivesrise totheslowertransverse S(secondary)waves. Using the gradient, divergence, and curl, and of course the BAC–CABrule, we may construct or verify a large number of useful vector identities. For verification, complete expansion into Cartesian components is always a possibility. Sometimes if we use insight insteadofroutineshufflingofCartesiancomponents,theverificationprocesscanbeshort- eneddrastically. Rememberthat ∇is avectoroperator,ahybridcreaturesatisfyingtwosets ofrules: 1. vectorrules,and 2. partialdifferentiationrules—includingdifferentiationofa product. Example 1.8.3 GRADIENT OF A DOTPRODUCT Verifythat ∇(A·B)=(B·∇)A+(A·∇)B+B×(∇×A)+A×(∇×B). (1.80) This particular example hinges on the recognition that ∇(A·B)is the type of term that appearsinthe BAC–CABexpansionofatriplevectorproduct,Eq.(1.55). Forinstance, A×(∇×B)=∇(A·B)−(A·∇)B, with the ∇differentiating only B, notA. From the commutativity of factors in a scalar productwemayinterchange AandBandwrite B×(∇×A)=∇(A·B)−(B·∇)A, now with ∇differentiating only A, notB. Adding these two equations, we obtain ∇dif- ferentiating the product A·Band the identity, Eq. (1.80). This identity is used frequently inelectromagnetictheory.Exercise1.8.13isasimpleillustration. /squaresolid 1.8 Curl, ∇× 47 Example 1.8.4 INTEGRATION BY PARTS OF CURL Let us prove the formulaintegraltext C(r)·(∇×A(r))d3r=integraltext A(r)·(∇×C(r))d3r, whereAor Cor bothvanishatinfinity. To show this, we proceed, as in Examples 1.6.3 and 1.7.3, by integration by parts after writing the inner product and the curl in Cartesian coordinates. Because the integrated termsvanishatinfinityweobtain integraldisplay C(r)·parenleftbig ∇×A(r)parenrightbig d3r =integraldisplaybracketleftbigg Czparenleftbigg∂Ay ∂x−∂Ax ∂yparenrightbigg +Cxparenleftbigg∂Az ∂y−∂Ay ∂zparenrightbigg +Cyparenleftbigg∂Ax ∂z−∂Az ∂xparenrightbiggbracketrightbigg d3r =integraldisplaybracketleftbigg Axparenleftbigg∂Cz ∂y−∂Cy ∂zparenrightbigg +Ayparenleftbigg∂Cx ∂z−∂Cz ∂xparenrightbigg +Azparenleftbigg∂Cy ∂x−∂Cx ∂yparenrightbiggbracketrightbigg d3r =integraldisplay A(r)·parenleftbig ∇×C(r)parenrightbig d3r, justrearrangingappropriatelytheterms afterintegrationbyparts. /squaresolid Exercises 1.8.1 Show,byrotatingthecoordinates,thatthecomponentsofthecurlofavectortransform asavector. Hint.Thedirectioncosineidentitiesof Eq.(1.46) areavailableas needed. 1.8.2 Showthat u×vissolenoidalif uandvareeachirrotational. 1.8.3 IfAisirrotational,showthat A×ris solenoidal. 1.8.4 A rigid body is rotating with constant angular velocity ω. Show that the linear velocity vissolenoidal. 1.8.5 Ifavectorfunction f(x,y,z)isnotirrotationalbuttheproductof fandascalarfunction g(x,y,z) isirrotational,showthatthen f·∇×f=0. 1.8.6 If(a)V=ˆxVx(x,y)+ˆyVy(x,y)and(b) ∇×V/negationslash=0,provethat ∇×Visperpendicular toV. 1.8.7 Classically, orbital angular momentum is given by L=r×p, wherepis the linear momentum. To go from classical mechanics to quantum mechanics, replace pby the operator−i∇(Section 15.6). Show that the quantum mechanical angular momentum 48 Chapter 1 Vector Analysis operatorhas Cartesiancomponents(inunitsof ¯h) Lx=−iparenleftbigg y∂ ∂z−z∂ ∂yparenrightbigg , Ly=−iparenleftbigg z∂ ∂x−x∂ ∂zparenrightbigg , Lz=−iparenleftbigg x∂ ∂y−y∂ ∂xparenrightbigg . 1.8.8 Using the angular momentum operators previously given, show that they satisfy com- mutationrelationsoftheform [Lx,Ly]≡LxLy−LyLx=iLz andhence L×L=iL. These commutation relations will be taken later as the defining relations of an angular momentumoperator—Exercise3.2.15andthefollowingoneandChapter4. 1.8.9 With the commutator bracket notation [Lx,Ly]=LxLy−LyLx, the angular momen- tumvector Lsatisfies[Lx,Ly]=iLz,et c. ,orL×L=iL. If two other vectors aandbcommute with each other and with L, that is,[a,b]= [a,L]=[b,L]=0,showthat [a·L,b·L]=i(a×b)·L. 1.8.10 ForA=ˆxAx(x,y,z)andB=ˆxBx(x,y,z)evaluateeachterminthevectoridentity ∇(A·B)=(B·∇)A+(A·∇)B+B×(∇×A)+A×(∇×B) andverifythattheidentityis satisfied. 1.8.11 Verifythevectoridentity ∇×(A×B)=(B·∇)A−(A·∇)B−B(∇·A)+A(∇·B). 1.8.12 Asanalternativetothevectoridentityof Example1.8.3showthat ∇(A·B)=(A×∇)×B+(B×∇)×A+A(∇·B)+B(∇·A). 1.8.13 Verifytheidentity A×(∇×A)=1 2∇parenleftbig A2parenrightbig −(A·∇)A. 1.8.14 IfAandBareconstantvectors,showthat ∇(A·B×r)=A×B. 1.9 Successive Applications of ∇ 49 1.8.15 A distribution of electric currents creates a constant magnetic moment m=const. The forceonminanexternalmagneticinduction Bis givenby F=∇×(B×m). Showthat F=(m·∇)B. Note.Assumingnotimedependenceofthefields,Maxwell’sequationsyield ∇×B=0. Also,∇·B=0. 1.8.16 An electric dipole of moment pis located at the origin. The dipole creates an electric potentialat rgivenby ψ(r)=p·r 4πε0r3. Findtheelectricfield, E=−∇ψatr. 1.8.17 The vector potential Aof a magnetic dipole, dipole moment m, is given by A(r)= (µ0/4π)(m×r/r3). Showthatthemagneticinduction B=∇×Ais givenby B=µ0 4π3ˆr(ˆr·m)−m r3. Note. The limiting process leading to point dipoles is discussed in Section 12.1 for electricdipoles,inSection12.5formagneticdipoles. 1.8.18 Thevelocityof atwo-dimensionalflowofliquidis givenby V=ˆxu(x,y)−ˆyv(x,y). If theliquidisincompressibleandtheflowisirrotational,showthat ∂u ∂x=∂v ∂yand∂u ∂y=−∂v ∂x. ThesearetheCauchy–Riemannconditionsof Section6.2. 1.8.19 The evaluation in this section of the four integrals for the circulation omitted Taylor series terms such as ∂Vx/∂x,∂Vy/∂yand all second derivatives. Show that ∂Vx/∂x, ∂Vy/∂ycancel out when the four integrals are added and that the second derivative termsdropoutinthelimitas dx→0,dy→0. Hint.Calculatethecirculationperunitareaandthentakethelimit dx→0,dy→0. 1.9 S UCCESSIVE APPLICATIONS OF ∇ We have now defined gradient, divergence, and curl to obtain vector, scalar, and vector quantities,respectively.Letting ∇operateoneachof thesequantities,weobtain (a)∇·∇ϕ (b)∇×∇ϕ (c)∇∇·V (d)∇·∇×V(e)∇×(∇×V) 50 Chapter 1 Vector Analysis allfiveexpressionsinvolvingsecondderivativesandallfiveappearinginthesecond-order differentialequationsof mathematicalphysics,particularlyinelectromagnetictheory. Thefirstexpression, ∇·∇ϕ,thedivergenceofthegradient,isnamedtheLaplacianof ϕ. We have ∇·∇ϕ=parenleftbigg ˆx∂ ∂x+ˆy∂ ∂y+ˆz∂ ∂zparenrightbigg ·parenleftbigg ˆx∂ϕ ∂x+ˆy∂ϕ ∂y+ˆz∂ϕ ∂zparenrightbigg =∂2ϕ ∂x2+∂2ϕ ∂y2+∂2ϕ ∂z2. (1.81a) Whenϕistheelectrostaticpotential,wehave ∇·∇ϕ=0 (1.81b) at points where the charge density vanishes, which is Laplace’s equation of electrostatics. Oftenthecombination ∇·∇is written ∇2,or/Delta1intheEuropeanliterature. Example 1.9.1 LAPLACIAN OF A POTENTIAL Calculate ∇·∇V(r). ReferringtoExamples1.6.1and1.7.2, ∇·∇V(r)=∇·ˆrdV dr=2 rdV dr+d2V dr2, replacing f(r)inExample1.7.2 by 1 /r·dV/dr.IfV(r)=rn, thisreducesto ∇·∇rn=n(n+1)rn−2. Thisvanishesfor n=0[V(r)=constant]andfor n=−1;thatis, V(r)=1/risasolution of Laplace’s equation, ∇2V(r)=0. This is for r/negationslash=0. Atr=0, a Dirac delta function is involved(seeEq. (1.169)andSection9.7). /squaresolid Expression(b) maybewritten ∇×∇ϕ=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆxˆyˆz ∂ ∂x∂ ∂y∂ ∂z ∂ϕ ∂x∂ϕ ∂y∂ϕ ∂zvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. Byexpandingthedeterminant,weobtain ∇×∇ϕ=ˆxparenleftbigg∂2ϕ ∂y∂z−∂2ϕ ∂z∂yparenrightbigg +ˆyparenleftbigg∂2ϕ ∂z∂x−∂2ϕ ∂x∂zparenrightbigg +ˆzparenleftbigg∂2ϕ ∂x∂y−∂2ϕ ∂y∂xparenrightbigg =0, (1.82) assuming that the order of partial differentiation may be interchanged. This is true as long as these second partial derivatives of ϕare continuous functions. Then, from Eq. (1.82), thecurlofagradientisidenticallyzero.Allgradients,therefore,areirrotational.Notethat 1.9 Successive Applications of ∇ 51 the zero in Eq. (1.82) comes as a mathematical identity, independent of any physics. The zeroinEq.(1.81b) isaconsequenceof physics. Expression(d) isa triplescalarproductthatmaybewritten ∇·∇×V=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂ ∂x∂ ∂y∂ ∂z ∂ ∂x∂ ∂y∂ ∂z VxVyVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (1.83) Again,assumingcontinuitysothattheorderofdifferentiationis immaterial,weobtain ∇·∇×V=0. (1.84) The divergence of a curl vanishes or all curls are solenoidal. In Section 1.16 we shall see thatvectorsmayberesolvedintosolenoidalandirrotationalpartsbyHelmholtz’stheorem. Thetworemainingexpressionssatisfyarelation ∇×(∇×V)=∇∇·V−∇·∇V, (1.85) valid in Cartesian coordinates (but not in curved coordinates). This follows immediately from Eq. (1.55), the BAC–CABrule, which we rewrite so that Cappears at the extreme rightof eachterm.The term ∇·∇Vwas notincludedin ourlist, but itmaybe definedby Eq.(1.85). Example 1.9.2 ELECTROMAGNETIC WAVEEQUATION One important application of this vector relation (Eq. (1.85)) is in the derivation of the electromagneticwaveequation.In vacuumMaxwell’sequationsbecome ∇·B=0, (1.86a) ∇·E=0, (1.86b) ∇×B=ε0µ0∂E ∂t, (1.86c) ∇×E=−∂B ∂t. (1.86d) HereEis the electric field, Bis the magnetic induction, ε0is the electric permittivity, andµ0is the magnetic permeability (SI units), so ε0µ0=1/c2,cbeing the velocity of light. The relation has important consequences. Because ε0,µ0can be measured in any frame,thevelocityof lightis thesameinanyframe. Suppose we eliminate Bfrom Eqs. (1.86c) and (1.86d). We may do this by taking the curlofbothsidesofEq.(1.86d)andthetimederivativeofbothsidesofEq.(1.86c).Since thespaceandtimederivativescommute, ∂ ∂t∇×B=∇×∂B ∂t, andweobtain ∇×(∇×E)=−ε0µ0∂2E ∂t2. 52 Chapter 1 Vector Analysis ApplicationofEqs. (1.85) and(1.86b)yields ∇·∇E=ε0µ0∂2E ∂t2, (1.87) the electromagnetic vector wave equation. Again, if Eis expressed in Cartesian coor- dinates, Eq. (1.87) separates into three scalar wave equations, each involving the scalar Laplacian. When external electric charge and current densities are kept as driving terms in Maxwell’s equations, similar wave equations are valid for the electric potential and the vector potential. To show this, we solve Eq. (1.86a) by writing B=∇×Aas a curl of the vector potential. This expression is substituted into Faraday’s induction law in differential form,Eq.(1.86d),toyield ∇×(E+∂A ∂t)=0.Thevanishingcurlimpliesthat E+∂A ∂tisa gradient and, therefore, can be written as −∇ϕ,whereϕ(r,t)is defined as the (nonstatic) electricpotential.Theseresultsfor the BandEfields, B=∇×A,E=−∇ϕ−∂A ∂t, (1.88) solvethehomogeneousMaxwell’sequations. We nowshowthattheinhomogeneousMaxwell’sequations, Gauss’ law: ∇·E=ρ/ε0,Oersted’slaw: ∇×B−1 c2∂E ∂t=µ0J(1.89) indifferentialformleadtowaveequationsforthepotentials ϕandA,providedthat ∇·Ais determinedbytheconstraint1 c2∂ϕ ∂t+∇·A=0.Thischoiceoffixingthedivergenceofthe vectorpotential,calledthe Lorentzgauge ,servestouncouplethedifferentialequationsof bothpotentials.This gaugeconstraintis notarestriction;ithas nophysicaleffect. SubstitutingourelectricfieldsolutionintoGauss’ lawyields ρ ε0=∇·E=−∇2ϕ−∂ ∂t∇·A=−∇2ϕ+1 c2∂2ϕ ∂t2, (1.90) the wave equation for the electric potential. In the last step we have used the Lorentz gaugetoreplacethedivergenceofthevectorpotentialbythetimederivativeoftheelectric potentialandthusdecouple ϕfromA. Finally, we substitute B=∇×Ainto Oersted’s law and use Eq. (1.85), which expands ∇2in terms of a longitudinal (the gradient term) and a transverse component (the curl term).This yields µ0J+1 c2∂E ∂t=∇×(∇×A)=∇(∇·A)−∇2A=µ0J−1 c2parenleftbigg ∇∂ϕ ∂t+∂2A ∂t2parenrightbigg , wherewehaveusedtheelectricfieldsolution(Eq.(1.88))inthelaststep.Nowweseethat theLorentzgaugeconditioneliminatesthegradientterms, sothewaveequation 1 c2∂2A ∂t2−∇2A=µ0J (1.91) 1.9 Successive Applications of ∇ 53 forthevectorpotentialremains. Finally, looking back at Oersted’s law, taking the divergence of Eq. (1.89), dropping ∇·(∇×B)=0,andsubstitutingGauss’lawfor ∇·E=ρ/ǫ0,wefindµ0∇·J=−1 ǫ0c2∂ρ ∂t, whereǫ0µ0=1/c2,that is, the continuityequationfor the current density. This step justi- fiestheinclusionofMaxwell’sdisplacementcurrentinthegeneralizationofOersted’slaw tononstationarysituations. /squaresolid Exercises 1.9.1 VerifyEq. (1.85), ∇×(∇×V)=∇∇·V−∇·∇V, bydirectexpansioninCartesiancoordinates. 1.9.2 Showthattheidentity ∇×(∇×V)=∇∇·V−∇·∇V followsfromthe BAC–CABruleforatriplevectorproduct.Justifyanyalterationofthe orderof factorsinthe BACandCABterms. 1.9.3 Provethat ∇×(ϕ∇ϕ)=0. 1.9.4 You are given that the curl of Fequals the curl of G. Show that FandGmay differ by (a) aconstantand(b) agradientof ascalarfunction. 1.9.5 TheNavier–Stokesequationofhydrodynamicscontainsanonlinearterm (v·∇)v.Show thatthecurlof thistermmaybewrittenas −∇×[v×(∇×v)]. 1.9.6 FromtheNavier–Stokesequationforthesteadyflowofanincompressibleviscousfluid wehavetheterm ∇×bracketleftbig v×(∇×v)bracketrightbig , wherevisthefluidvelocity.Showthatthistermvanishesforthespecialcase v=ˆxv(y,z). 1.9.7 Provethat (∇u)×(∇v)issolenoidal,where uandvaredifferentiablescalarfunctions. 1.9.8 ϕis a scalar satisfying Laplace’s equation, ∇2ϕ=0. Show that ∇ϕisbothsolenoidal andirrotational. 1.9.9 Withψascalar(wave)function,showthat (r×∇)·(r×∇)ψ=r2∇2ψ−r2∂2ψ ∂r2−2r∂ψ ∂r. (This canactuallybeshownmoreeasilyinsphericalpolarcoordinates,Section2.5.) 54 Chapter 1 Vector Analysis 1.9.10 Ina(nonrotating)isolatedmasssuchasastar, theconditionfor equilibriumis ∇P+ρ∇ϕ=0. HerePis the total pressure, ρis the density, and ϕis the gravitational potential. Show that at any given point the normals to the surfaces of constant pressure and constant gravitationalpotentialareparallel. 1.9.11 InthePaulitheoryoftheelectron,oneencounterstheexpression (p−eA)×(p−eA)ψ, whereψis a scalar (wave) function. Ais the magnetic vector potential related to the magnetic induction BbyB=∇×A.Given that p=−i∇, show that this expression reduces to ieBψ. Show that this leads to the orbital g-factorgL=1 upon writing the magnetic moment as µ=gLLin units of Bohr magnetons and L=−ir×∇.See also Exercise1.13.7. 1.9.12 Showthatanysolutionof theequation ∇×(∇×A)−k2A=0 automaticallysatisfies thevectorHelmholtzequation ∇2A+k2A=0 andthesolenoidalcondition ∇·A=0. Hint.Let∇·operateonthefirst equation. 1.9.13 Thetheoryof heatconductionleadstoanequation ∇2/Psi1=k|∇/Phi1|2, where/Phi1isapotentialsatisfyingLaplace’sequation: ∇2/Phi1=0.Showthatasolutionof thisequationis /Psi1=1 2k/Phi12. 1.10 V ECTOR INTEGRATION Thenextstepafterdifferentiatingvectorsistointegratethem.Letusstartwithlineintegrals andthenproceedtosurfaceandvolumeintegrals.Ineachcasethemethodofattackwillbe toreducethevectorintegraltoscalarintegralswithwhichthereaderis assumedfamiliar. 1.10 Vector Integration 55 LineIntegrals Usinganincrementoflength dr=ˆxdx+ˆydy+ˆzdz,wemayencounterthelineintegrals integraldisplay Cϕdr, (1.92a) integraldisplay CV·dr, (1.92b) integraldisplay CV×dr, (1.92c) ineachofwhichtheintegralisoversomecontour Cthatmaybeopen(withstartingpoint and ending point separated) or closed (forming a loop). Because of its physical interpreta- tionthatfollows,thesecondform, Eq. (1.92b)isbyfar themostimportantofthethree. Withϕ, ascalar,thefirst integralreducesimmediatelyto integraldisplay Cϕdr=ˆxintegraldisplay Cϕ(x,y,z)dx+ˆyintegraldisplay Cϕ(x,y,z)dy+ˆzintegraldisplay Cϕ(x,y,z)dz. (1.93) Thisseparationhasemployedtherelation integraldisplay ˆxϕdx=ˆxintegraldisplay ϕdx, (1.94) which is permissible because the Cartesian unit vectors ˆx,ˆy, andˆzare constant in both magnitudeanddirection.Perhapsthisrelationisobvioushere,butitwillnotbetrueinthe non-CartesiansystemsencounteredinChapter2. The three integrals on the right side of Eq. (1.93) are ordinary scalar integrals and, to avoid complications, we assume that they are Riemann integrals. Note, however, that the integral with respect to xcannot be evaluated unless yandzare known in terms of x and similarly for the integrals with respect to yandz. This simply means that the path of integration Cmust be specified. Unless the integrand has special properties so that the integral depends only on the value of the end points, the value will depend on the particular choice of contour C. For instance, if we choose the very special case ϕ=1, Eq. (1.92a) is just the vector distance from the start of contour Cto the endpoint, in this caseindependentofthechoiceofpathconnectingfixedendpoints.With dr=ˆxdx+ˆydy+ ˆzdz, the second and third forms also reduce to scalar integrals and, like Eq. (1.92a), are dependent, in general, on the choice of path. The form (Eq. (1.92b)) is exactly the same as that encountered when we calculate the work done by a force that varies along the path, W=integraldisplay F·dr=integraldisplay Fx(x,y,z)dx+integraldisplay Fy(x,y,z)dy+integraldisplay Fz(x,y,z)dz. (1.95a) Inthisexpression Fis theforce exertedonaparticle. 56 Chapter 1 Vector Analysis FIGURE 1.25Apathof integration. Example 1.10.1 PATH-DEPENDENT WORK The force exerted on a body is F=−ˆxy+ˆyx. The problem is to calculate the work done goingfromtheorigintothepoint (1,1): W=integraldisplay1,1 0,0F·dr=integraldisplay1,1 0,0(−ydx+xdy). (1.95b) Separatingthetwointegrals,weobtain W=−integraldisplay1 0ydx+integraldisplay1 0xdy. (1.95c) The first integral cannot be evaluated until we specify the values of yasxranges from 0 to 1. Likewise, the second integral requires xas a function of y. Consider first the path showninFig.1.25. Then W=−integraldisplay1 00dx+integraldisplay1 01dy=1, (1.95d) sincey=0alongthefirstsegmentofthepathand x=1alongthesecond.Ifweselectthe path[x=0,0/lessorequalslanty/lessorequalslant1]and[0/lessorequalslantx/lessorequalslant1,y=1], then Eq. (1.95c) gives W=−1. For this forcetheworkdonedependsonthechoiceofpath. /squaresolid SurfaceIntegrals Surfaceintegralsappearinthesameforms aslineintegrals,theelementofareaalsobeing avector,dσ.20Oftenthisareaelementiswritten ndA,inwhich nisaunit(normal)vector to indicate the positive direction.21There are two conventions for choosing the positive direction. First, if the surface is a closed surface, we agree to take the outward normal as positive. Second, if the surface is an open surface, the positive normal depends on the direction in which the perimeter of the open surface is traversed. If the right-hand fingers 20Recallthat in Section1.4the area(of aparallelogram) is representedby across-product vector. 21Although nalwayshas unit length, its direction may wellbe afunction ofposition. 1.10 Vector Integration 57 FIGURE 1.26Right-handrulefor thepositivenormal. areplacedinthedirectionoftravelaroundtheperimeter,thepositivenormalisindicatedby thethumbofthe righthand.As an illustration,acirclein the xy-plane(Fig. 1.26) mapped out from xtoyto−xto−yand back to xwill have its positive normal parallel to the positivez-axis (for theright-handedcoordinatesystem). Analogous to the line integrals, Eqs. (1.92a) to (1.92c), surface integrals may appear in theforms integraldisplay ϕdσ,integraldisplay V·dσ,integraldisplay V×dσ. Again,thedotproductisbyfarthemostcommonlyencounteredform.Thesurfaceintegralintegraltext V·dσmaybeinterpretedasafloworfluxthroughthegivensurface.Thisisreallywhat we did in Section 1.7 to obtain the significance of the term divergence. This identification reappears in Section 1.11 as Gauss’ theorem. Note that both physically and from the dot product the tangential components of the velocity contribute nothing to the flow through thesurface. VolumeIntegrals Volume integrals are somewhat simpler, for the volume element dτis a scalar quantity.22 We have integraldisplay VVdτ=ˆxintegraldisplay VVxdτ+ˆyintegraldisplay VVydτ+ˆzintegraldisplay VVzdτ, (1.96) againreducingthevectorintegraltoavectorsumof scalarintegrals. 22Frequently the symbols d3randd3xareused to denote avolume elementin coordinate ( xyzorx1x2x3)space. 58 Chapter 1 Vector Analysis FIGURE 1.27Differentialrectangularparallelepiped(originatcenter). IntegralDefinitionsof Gradient,Divergence,andCurl One interesting and significant application of our surface and volume integrals is their use indevelopingalternatedefinitionsof ourdifferentialrelations.We find ∇ϕ=limintegraltext dτ→0integraltext ϕdσintegraltext dτ, (1.97) ∇·V=limintegraltext dτ→0integraltext V·dσintegraltext dτ, (1.98) ∇×V=limintegraltext dτ→0integraltext dσ×Vintegraltext dτ. (1.99) Inthesethreeequationsintegraltext dτisthevolumeofasmallregionofspaceand dσisthevector area element of this volume. The identification of Eq. (1.98) as the divergence of Vwas carried out in Section 1.7. Here we show that Eq. (1.97) is consistent with our earlier definition of ∇ϕ(Eq. (1.60)). For simplicity we choose dτto be the differential volume dxdydz (Fig. 1.27). This time we place the origin at the geometric center of our volume element.Theareaintegralleadstosixintegrals,oneforeachofthesixfaces.Remembering thatdσis outward, dσ·ˆx=−|dσ|for surface EFHG, and+|dσ|for surface ABDC,w e have integraldisplay ϕdσ=−ˆxintegraldisplay EFHGparenleftbigg ϕ−∂ϕ ∂xdx 2parenrightbigg dydz+ˆxintegraldisplay ABDCparenleftbigg ϕ+∂ϕ ∂xdx 2parenrightbigg dydz −ˆyintegraldisplay AEGCparenleftbigg ϕ−∂ϕ ∂ydy 2parenrightbigg dxdz+ˆyintegraldisplay BFHDparenleftbigg ϕ+∂ϕ ∂ydy 2parenrightbigg dxdz −ˆzintegraldisplay ABFEparenleftbigg ϕ−∂ϕ ∂zdz 2parenrightbigg dxdy+ˆzintegraldisplay CDHGparenleftbigg ϕ+∂ϕ ∂zdz 2parenrightbigg dxdy. 1.10 Vector Integration 59 Using the total variations, we evaluate each integrand at the origin with a correction in- cluded to correct for the displacement ( ±dx/2, etc.) of the center of the face from the origin. Having chosen the total volume to be of differential size (integraltext dτ=dxdydz) ,w e droptheintegralsigns ontherightandobtain integraldisplay ϕdσ=parenleftbigg ˆx∂ϕ ∂x+ˆy∂ϕ ∂y+ˆz∂ϕ ∂zparenrightbigg dxdydz. (1.100) Dividingby integraldisplay dτ=dxdydz, weverifyEq. (1.97). This verification has been oversimplified in ignoring other correction terms beyond the first derivatives. These additional terms, which are introduced in Section 5.6 when the Taylorexpansionisdeveloped,vanishinthelimit integraldisplay dτ→0(dx→0,dy→0,dz→0). This,ofcourse,isthereasonforspecifyinginEqs.(1.97),(1.98),and(1.99)thatthislimit be taken. Verification of Eq. (1.99) follows these same lines exactly, using a differential volumedxdydz. Exercises 1.10.1 Theforcefieldactingona two-dimensionallinearoscillatormaybedescribedby F=−ˆxkx−ˆyky. Comparetheworkdonemovingagainstthisforcefieldwhengoingfrom (1,1)to(4,4) bythefollowingstraight-linepaths: (a)(1,1)→(4,1)→(4,4) (b)(1,1)→(1,4)→(4,4) (c)(1,1)→(4,4)alongx=y. This meansevaluating −integraldisplay(4,4) (1,1)F·dr alongeachpath. 1.10.2 Findtheworkdonegoingaroundaunitcircleinthe xy-plane: (a) counterclockwisefrom 0 to π, (b) clockwisefrom 0 to −π, doingwork againstaforcefieldgivenby F=−ˆxy x2+y2+ˆyx x2+y2. Notethatthework donedependsonthepath. 60 Chapter 1 Vector Analysis 1.10.3 Calculatetheworkyoudoingoingfrompoint (1,1)topoint(3,3).Theforce youexert isgivenby F=ˆx(x−y)+ˆy(x+y). Specifyclearlythepathyouchoose.Notethatthisforcefieldisnonconservative. 1.10.4 Evaluatecontintegraltext r·dr. Note.Thesymbolcontintegraltext meansthatthepathofintegrationisaclosedloop. 1.10.5 Evaluate 1 3integraldisplay sr·dσ over the unit cube defined by the point (0,0,0)and the unit intercepts on the positive x-,y-, andz-axes. Note that (a) r·dσis zero for three of the surfaces and (b) each of thethreeremainingsurfaces contributesthesameamounttotheintegral. 1.10.6 Show,byexpansionofthesurface integral,that limintegraltext dτ→0integraltext sdσ×Vintegraltext dτ=∇×V. Hint.Choosethevolumeintegraltext dτtobeadifferentialvolume dxdydz. 1.11 G AUSS ’THEOREM Herewederiveausefulrelationbetweenasurfaceintegralofavectorandthevolumeinte- gralofthedivergenceofthatvector.Letusassumethatthevector Vanditsfirstderivatives are continuous over the simply connected region (that does not have any holes, such as a donut)of interest.ThenGauss’theoremstatesthat integraldisplayintegraldisplay /circlecopyrt ∂VV·dσ=integraldisplayintegraldisplayintegraldisplay V∇·Vdτ. (1.101a) In words, the surface integral of a vector over a closed surface equals the volume integral ofthedivergenceof thatvectorintegratedoverthevolumeenclosedbythesurface. Imagine that volume Vis subdivided into an arbitrarily large number of tiny (differen- tial)parallelepipeds.Foreachparallelepiped summationdisplay six surfacesV·dσ=∇·Vdτ (1.101b) from the analysis of Section 1.7, Eq. (1.66), with ρvreplaced by V. The summation is overthesixfacesoftheparallelepiped.Summingoverallparallelepipeds,wefindthatthe V·dσtermscancel(pairwise)forall interiorfaces;onlythecontributionsofthe exterior surfaces survive (Fig. 1.28). Analogousto the definitionof a Riemannintegral as the limit 1.11 Gauss’ Theorem 61 FIGURE 1.28Exact cancellationof dσ’s on interiorsurfaces. No cancellationonthe exteriorsurface. of a sum, we take the limit as the number of parallelepipeds approaches infinity (→∞) andthedimensionsofeachapproachzero (→0): summationtext exterior surfacesV·dσ=summationtext volumes∇·Vdτ integraltext SV·dσ=integraltext V∇·Vdτ. TheresultisEq. (1.101a), Gauss’theorem. From a physical point of view Eq. (1.66) has established ∇·Vas the net outflow of fluidper unitvolume.The volumeintegralthengivesthetotalnetoutflow.Butthesurface integralintegraltext V·dσisjustanotherwayofexpressingthissamequantity,whichistheequality, Gauss’theorem. Green’s Theorem AfrequentlyusefulcorollaryofGauss’theoremisarelationknownasGreen’stheorem.If uandvare twoscalarfunctions,wehavetheidentities ∇·(u∇v)=u∇·∇v+(∇u)·(∇v), (1.102) ∇·(v∇u)=v∇·∇u+(∇v)·(∇u). (1.103) Subtracting Eq. (1.103) from Eq. (1.102), integrating over a volume ( u,v, and their derivatives,assumedcontinuous),andapplyingEq. (1.101a)(Gauss’ theorem),weobtain integraldisplayintegraldisplayintegraldisplay V(u∇·∇v−v∇·∇u)dτ=integraldisplayintegraldisplay /circlecopyrt ∂V(u∇v−v∇u)·dσ. (1.104) 62 Chapter 1 Vector Analysis This is Green’s theorem. We use it for developing Green’s functions in Chapter 9. An alternateformofGreen’stheorem,derivedfromEq. (1.102)alone,is integraldisplayintegraldisplay /circlecopyrt ∂Vu∇v·dσ=integraldisplayintegraldisplayintegraldisplay Vu∇·∇vdτ+integraldisplayintegraldisplayintegraldisplay V∇u·∇vdτ. (1.105) Thisis theform ofGreen’stheoremusedinSection1.16. Alternate Forms of Gauss’ Theorem AlthoughEq.(1.101a)involvingthedivergenceisbyfarthemostimportantformofGauss’ theorem,volumeintegralsinvolvingthegradientandthecurlmayalsoappear.Suppose V(x,y,z)=V(x,y,z) a, (1.106) in which ais a vector with constant magnitude and constant but arbitrary direction. (You pickthedirection,butonceyouhavechosenit, holditfixed.)Equation(1.101a)becomes a·integraldisplayintegraldisplay /circlecopyrt ∂VVdσ=integraldisplayintegraldisplayintegraldisplay V∇·aVdτ=a·integraldisplayintegraldisplayintegraldisplay V∇Vdτ (1.107) byEq.(1.67b). Thismayberewritten a·bracketleftbiggintegraldisplayintegraldisplay /circlecopyrt ∂VVdσ−integraldisplayintegraldisplayintegraldisplay V∇Vdτbracketrightbigg =0. (1.108) Since|a|/negationslash=0 and its direction is arbitrary, meaning that the cosine of the included angle cannotalwaysvanish,thetermsinbracketsmustbezero.23Theresultis integraldisplayintegraldisplay /circlecopyrt ∂VVdσ=integraldisplayintegraldisplayintegraldisplay V∇Vdτ . (1.109) Inasimilarmanner,using V=a×Pinwhichais aconstantvector,wemayshow integraldisplayintegraldisplay /circlecopyrt ∂Vdσ×P=integraldisplayintegraldisplayintegraldisplay V∇×Pdτ. (1.110) TheselasttwoformsofGauss’theoremareusedinthevectorformofKirchoffdiffraction theory. They may also be used to verify Eqs. (1.97) and (1.99). Gauss’ theorem may also beextendedtotensors(seeSection2.11). Exercises 1.11.1 UsingGauss’theorem,provethat integraldisplayintegraldisplay /circlecopyrt Sdσ=0 ifS=∂Visaclosedsurface. 23This exploitation of the arbitrary nature of a part of a problem is a valuable and widely used technique. The arbitrary vector isusedagaininSections1.12and1.13.OtherexamplesappearinSection1.14(integrandsequated)andinSection2.8,quotient rule. 1.11 Gauss’ Theorem 63 1.11.2 Showthat 1 3integraldisplayintegraldisplay /circlecopyrt Sr·dσ=V, whereVisthevolumeenclosedbytheclosedsurface S=∂V. Note.This isageneralizationofExercise1.10.5. 1.11.3 IfB=∇×A,showthat integraldisplayintegraldisplay /circlecopyrt SB·dσ=0 for anyclosedsurface S. 1.11.4 Over some volume Vletψbe a solution of Laplace’s equation (with the derivatives appearing there continuous). Prove that the integral over any closed surface in Vof the normalderivativeof ψ(∂ψ/∂n,o r∇ψ·n)willbezero. 1.11.5 In analogy to the integral definition of gradient, divergence, and curl of Section 1.10, showthat ∇2ϕ=limintegraltext dτ→0integraltext ∇ϕ·dσintegraltext dτ. 1.11.6 The electric displacement vector Dsatisfies the Maxwell equation ∇·D=ρ, whereρ is the charge density (per unit volume). At the boundary between two media there is a surface chargedensity σ(per unitarea). Showthataboundaryconditionfor Dis (D2−D1)·n=σ. nis aunitvectornormaltothesurfaceandoutofmedium1. Hint.Considerathinpillboxas showninFig.1.29. 1.11.7 From Eq. (1.67b), with Vthe electric field Eandfthe electrostatic potential ϕ,show that,for integrationoverallspace, integraldisplay ρϕdτ=ε0integraldisplay E2dτ. Thiscorrespondstoathree-dimensionalintegrationbyparts. Hint.E=−∇ϕ,∇·E=ρ/ε0.You may assume that ϕvanishes at large rat least as fast asr−1. FIGURE 1.29Pillbox. 64 Chapter 1 Vector Analysis 1.11.8 A particular steady-state electric current distribution is localized in space. Choosing a boundingsurfacefar enoughoutsothatthecurrentdensity Jiszeroeverywhereonthe surface, showthat integraldisplayintegraldisplayintegraldisplay Jdτ=0. Hint. Take one component of Jat a time. With ∇·J=0, show that Ji=∇·(xiJ)and applyGauss’theorem. 1.11.9 The creation of a localized system of steady electric currents (current density J) and magneticfieldsmaybeshowntorequireanamountof work W=1 2integraldisplayintegraldisplayintegraldisplay H·Bdτ. Transformthisinto W=1 2integraldisplayintegraldisplayintegraldisplay J·Adτ. HereAisthemagneticvectorpotential: ∇×A=B. Hint.InMaxwell’sequationstakethedisplacementcurrentterm ∂D/∂t=0.Ifthefields and currents are localized, a bounding surface may be taken far enough out so that the integralsof thefieldsandcurrentsoverthesurfaceyieldzero. 1.11.10 Provethegeneralizationof Green’stheorem: integraldisplayintegraldisplayintegraldisplay V(vLu−uLv)dτ=integraldisplayintegraldisplay /circlecopyrt ∂Vp(v∇u−u∇v)·dσ. HereListheself-adjointoperator(Section10.1), L=∇·bracketleftbig p(r)∇bracketrightbig +q(r) andp,q,u,andvarefunctionsofposition, pandqhavingcontinuousfirstderivatives anduandvhavingcontinuoussecondderivatives. Note.This generalizedGreen’stheoremappearsinSection9.7. 1.12 S TOKES ’THEOREM Gauss’ theorem relates the volume integral of a derivative of a function to an integral of the function over the closed surface bounding the volume. Here we consider an analogous relation between the surface integral of a derivative of a function and the line integral of thefunction,thepathofintegrationbeingtheperimeterboundingthesurface. Let us take the surface and subdivide it into a network of arbitrarily small rectangles. In Section 1.8 we showed that the circulation about such a differential rectangle (in the xy-plane)is ∇×V|zdxdy.FromEq. (1.76)appliedto onedifferentialrectangle, summationdisplay four sidesV·dλ=∇×V·dσ. (1.111) 1.12 Stokes’ Theorem 65 FIGURE 1.30Exactcancellationon interiorpaths.Nocancellationonthe exteriorpath. Wesumoverallthelittlerectangles,asinthedefinitionofaRiemannintegral.Thesurface contributions (right-hand side of Eq. (1.111)) are added together. The line integrals (left- hand side of Eq. (1.111)) of all interiorline segments cancel identically. Only the line integral around the perimeter survives (Fig. 1.30). Taking the usual limit as the number of rectanglesapproachesinfinitywhile dx→0,dy→0,wehave summationtext exterior line segmentsV·dλ=summationtext rectangles∇×V·dσ(1.112) contintegraldisplay V·dλ=integraldisplay S∇×V·dσ. This is Stokes’ theorem. The surface integral on the right is over the surface bounded by the perimeter or contour, for the line integral on the left. The direction of the vector representingtheareaisoutofthepaperplanetowardthereaderifthedirectionoftraversal around the contour for the line integral is in the positive mathematical sense, as shown in Fig.1.30. This demonstration of Stokes’ theorem is limited by the fact that we used a Maclaurin expansion of V(x,y,z)in establishing Eq. (1.76) in Section 1.8. Actually we need only demand that the curl of V(x,y,z)exist and that it be integrable over the surface. A proof of the Cauchy integral theorem analogous to the developmentof Stokes’ theorem here but usingtheseless restrictiveconditionsappearsinSection6.3. Stokes’theoremobviouslyappliestoanopensurface.Itispossibletoconsideraclosed surfaceasalimitingcaseofanopensurface,withtheopening(andthereforetheperimeter) shrinkingtozero.This isthepointof Exercise1.12.7. 66 Chapter 1 Vector Analysis Alternate Forms of Stokes’ Theorem As with Gauss’ theorem, other relations between surface and line integrals are possible. Wefind integraldisplay Sdσ×∇ϕ=contintegraldisplay ∂Sϕdλ (1.113) and integraldisplay S(dσ×∇)×P=contintegraldisplay ∂Sdλ×P. (1.114) Equation (1.113) may readily be verified by the substitution V=aϕ, in which ais a vec- tor of constant magnitude and of constant direction, as in Section 1.11. Substituting into Stokes’theorem,Eq. (1.112), integraldisplay S(∇×aϕ)·dσ=−integraldisplay Sa×∇ϕ·dσ =−a·integraldisplay S∇ϕ×dσ. (1.115) Forthelineintegral, contintegraldisplay ∂Saϕ·dλ=a·contintegraldisplay ∂Sϕdλ, (1.116) andweobtain a·parenleftbiggcontintegraldisplay ∂Sϕdλ+integraldisplay S∇ϕ×dσparenrightbigg =0. (1.117) Since the choice of direction of ais arbitrary, the expression in parentheses must vanish, thusverifyingEq.(1.113).Equation(1.114)maybederivedsimilarlybyusing V=a×P, inwhichaisagainaconstantvector. We can use Stokes’ theorem to derive Oersted’s and Faraday’s laws from two of Maxwell’s equations, and vice versa, thus recognizing that the former are an integrated formofthelatter. Example 1.12.1 OERSTED ’SA N D FARADAY ’SLAWS Consider the magnetic field generated by a long wire that carries a stationary current I. StartingfromMaxwell’sdifferentiallaw ∇×H=J,Eq.(1.89)(withMaxwell’sdisplace- ment current ∂D/∂t=0 for a stationary current case by Ohm’s law), we integrate over a closedarea SperpendiculartoandsurroundingthewireandapplyStokes’theoremtoget I=integraldisplay SJ·dσ=integraldisplay S(∇×H)·dσ=contintegraldisplay ∂SH·dr, whichisOersted’slaw.Herethelineintegralisalong ∂S,theclosedcurvesurroundingthe cross-sectionalarea S. 1.12 Stokes’ Theorem 67 Similarly,wecanintegrateMaxwell’sequationfor ∇×E,Eq.(1.86d),toyieldFaraday’s induction law. Imagine moving a closed loop (∂S)of wire (of area S) across a magnetic inductionfield B. WeintegrateMaxwell’sequationanduseStokes’theorem,yielding integraldisplay ∂SE·dr=integraldisplay S(∇×E)·dσ=−d dtintegraldisplay SB·dσ=−d/Phi1 dt, which is Faraday’s law. The line integral on the left-hand side represents the voltage in- duced in the wire loop, while the right-hand side is the change with time of the magnetic flux/Phi1throughthemovingsurface Softhewire. /squaresolid Both Stokes’ and Gauss’ theorems are of tremendous importance in a wide variety of problems involving vector calculus. Some idea of their power and versatility may be ob- tainedfromtheexercisesofSections1.11and1.12andthedevelopmentofpotentialtheory inSections1.13and1.14. Exercises 1.12.1 Given a vector t=−ˆxy+ˆyx, show, with the help of Stokes’ theorem, that the integral aroundacontinuousclosedcurveinthe xy-plane 1 2contintegraldisplay t·dλ=1 2contintegraldisplay (xdy−ydx)=A, theareaenclosedbythecurve. 1.12.2 Thecalculationofthemagneticmomentof acurrentloopleadstothelineintegral contintegraldisplay r×dr. (a) Integrate around the perimeter of a current loop (in the xy-plane) and show that thescalarmagnitudeofthislineintegralistwicetheareaoftheenclosedsurface. (b) The perimeter of an ellipse is described by r=ˆxacosθ+ˆybsinθ. From part (a) showthattheareaoftheellipseis πab. 1.12.3 Evaluatecontintegraltext r×drbyusingthealternateform ofStokes’theoremgivenbyEq. (1.114): integraldisplay S(dσ×∇)×P=contintegraldisplay dλ×P. Takethelooptobeentirelyinthe xy-plane. 1.12.4 Insteadystatethemagneticfield HsatisfiestheMaxwellequation ∇×H=J,whereJ isthecurrentdensity(persquaremeter).Attheboundarybetweentwomediathereisa surface currentdensity K.Showthataboundaryconditionon His n×(H2−H1)=K. nis aunitvectornormaltothesurfaceandoutofmedium1. Hint.Consideranarrowloopperpendiculartotheinterfaceas showninFig.1.31. 68 Chapter 1 Vector Analysis FIGURE 1.31 Integrationpath attheboundary oftwomedia. 1.12.5 FromMaxwell’sequations, ∇×H=J,withJherethecurrentdensityand E=0.Show fromthisthatcontintegraldisplay H·dr=I, whereIisthenetelectriccurrentenclosedbytheloopintegral.Thesearethedifferential andintegralforms ofAmpère’slawofmagnetism. 1.12.6 Amagneticinduction Bisgeneratedbyelectriccurrentinaringofradius R.Showthat themagnitude ofthevectorpotential A(B=∇×A)attheringcanbe |A|=ϕ 2πR, whereϕisthetotalmagneticfluxpassingthroughthering. Note.Ais tangential to the ring and may be changed by adding the gradient of a scalar function. 1.12.7 Provethatintegraldisplay S∇×V·dσ=0, ifSisaclosedsurface. 1.12.8 Evaluatecontintegraltext r·dr(Exercise1.10.4)byStokes’ theorem. 1.12.9 Provethatcontintegraldisplay u∇v·dλ=−contintegraldisplay v∇u·dλ. 1.12.10 Provethatcontintegraldisplay u∇v·dλ=integraldisplay S(∇u)×(∇v)·dσ. 1.13 P OTENTIAL THEORY Scalar Potential If a force over a given simply connected region of space S(which means that it has no holes)canbeexpressedasthenegativegradientof ascalarfunction ϕ, F=−∇ϕ, (1.118) 1.13 Potential Theory 69 wecallϕascalarpotentialthatdescribestheforcebyonefunctioninsteadofthree.Ascalar potentialisonlydetermineduptoanadditiveconstant,whichcanbeusedtoadjustitsvalue at infinity (usually zero) or at some other point. The force Fappearing as the negative gradient of a single-valued scalar potential is labeled a conservative force. We want to know when a scalar potential function exists. To answer this question we establish two otherrelationsasequivalenttoEq.(1.118). Theseare ∇×F=0 (1.119) and contintegraldisplay F·dr=0, (1.120) for every closed path in our simply connected region S. We proceed to show that each of thesethreeequationsimpliestheothertwo.Letus startwith F=−∇ϕ. (1.121) Then ∇×F=−∇×∇ϕ=0 (1.122) byEq.(1.82) orEq. (1.118)impliesEq.(1.119). Turningtothelineintegral,wehave contintegraldisplay F·dr=−contintegraldisplay ∇ϕ·dr=−contintegraldisplay dϕ, (1.123) using Eq. (1.118). Now, dϕintegrates to give ϕ. Since we have specified a closed loop, the end points coincide and we get zero for every closed path in our region Sfor which Eq. (1.118) holds. It is important to note the restriction here that the potential be single- valuedandthatEq.(1.118)holdfor allpointsinS.Thisproblemmayariseinusingascalar magnetic potential, a perfectly valid procedure as long as no net current is encircled. As soonaswechooseapathinspacethatencirclesanetcurrent,thescalarmagneticpotential ceasestobesingle-valuedandouranalysisnolongerapplies. Continuing this demonstration of equivalence, let us assume that Eq. (1.120) holds. Ifcontintegraltext F·dr=0 for all paths in S, we see that the value of the integral joining two distinct pointsAandBisindependentof thepath(Fig.1.32). Ourpremiseisthat contintegraldisplay ACBDAF·dr=0. (1.124) Therefore integraldisplay ACBF·dr=−integraldisplay BDAF·dr=integraldisplay ADBF·dr, (1.125) reversing the sign by reversing the direction of integration. Physically, this means that the work done in going from AtoBis independent of the path and that the work done in goingaroundaclosedpathiszero.Thisisthereasonforlabelingsuchaforceconservative: Energyisconserved. 70 Chapter 1 Vector Analysis FIGURE 1.32Possiblepathsfor doingwork. With the result shown in Eq. (1.125), we have the work done dependent only on the endpoints AandB. Thatis, workdonebyforce =integraldisplayB AF·dr=ϕ(A)−ϕ(B). (1.126) Equation (1.126) defines a scalar potential (strictly speaking, the difference in potential between points AandB) and provides a means of calculating the potential. If point B is taken as a variable, say, (x,y,z), then differentiation with respect to x,y, andzwill recoverEq.(1.118). Thechoiceofsignontheright-handsideisarbitrary.Thechoicehereismadetoachieve agreement with Eq. (1.118) and to ensure that water will run downhill rather than uphill. Forpoints AandBseparatedbyalength dr,Eq. (1.126)becomes F·dr=−dϕ=−∇ϕ·dr. (1.127) Thismayberewritten (F+∇ϕ)·dr=0, (1.128) andsince dris arbitrary,Eq. (1.118)mustfollow.If contintegraldisplay F·dr=0, (1.129) wemayobtainEq. (1.119)byusingStokes’theorem(Eq. (1.112)): contintegraldisplay F·dr=integraldisplay ∇×F·dσ. (1.130) If we take the path of integration to be the perimeter of an arbitrary differential area dσ, theintegrandinthesurfaceintegralmustvanish.HenceEq. (1.120)impliesEq. (1.119). Finally, if ∇×F=0, we need only reverse our statement of Stokes’ theorem (Eq. (1.130)) to derive Eq. (1.120). Then, by Eqs. (1.126) to (1.128), the initial statement 1.13 Potential Theory 71 FIGURE 1.33Equivalentformulationsof aconservativeforce. FIGURE 1.34Potentialenergyversusdistance(gravitational, centrifugal,andsimpleharmonicoscillator). F=−∇ϕis derived. The triple equivalence is demonstrated (Fig. 1.33). To summarize, asingle-valuedscalarpotentialfunction ϕexistsifandonlyif Fisirrotationalorthework donearoundeveryclosedloopiszero.Thegravitationalandelectrostaticforcefieldsgiven byEq.(1.79)areirrotationalandthereforeareconservative.Gravitationalandelectrostatic scalar potentials exist. Now, by calculating the work done (Eq. (1.126)), we proceed to determinethreepotentials(Fig.1.34). 72 Chapter 1 Vector Analysis Example 1.13.1 GRAVITATIONAL POTENTIAL Findthescalarpotentialfor thegravitationalforceonaunitmass m1, FG=−Gm1m2ˆr r2=−kˆr r2, (1.131) radiallyinward. ByintegratingEq.(1.118) frominfinityintoposition r,weobtain ϕG(r)−ϕG(∞)=−integraldisplayr ∞FG·dr=+integraldisplay∞ rFG·dr. (1.132) By use of FG=−Fapplied, a comparison with Eq. (1.95a) shows that the potential is the work done in bringing the unit mass in from infinity. (We can define only potential dif- ference. Here we arbitrarily assign infinity to be a zero of potential.) The integral on the right-hand side of Eq. (1.132) is negative, meaning that ϕG(r)is negative. Since FGis radial,weobtaina contributionto ϕonlywhen dris radial,or ϕG(r)=−integraldisplay∞ rkdr r2=−k r=−Gm1m2 r. Thefinalnegativesignisaconsequenceof theattractiveforceof gravity. /squaresolid Example 1.13.2 CENTRIFUGAL POTENTIAL Calculate the scalar potential for the centrifugal force per unit mass, FC=ω2rˆr, radially outward. Physically, you might feel this on a large horizontal spinning disk at an amuse- ment park. Proceeding as in Example 1.13.1 but integrating from the origin outward and takingϕC(0)=0,wehave ϕC(r)=−integraldisplayr 0FC·dr=−ω2r2 2. If we reverse signs, taking FSHO=−kr, we obtain ϕSHO=1 2kr2, the simple harmonic oscillatorpotential. The gravitational, centrifugal, and simple harmonic oscillator potentials are shown in Fig. 1.34. Clearly, the simple harmonic oscillator yields stability and describes a restoring force.Thecentrifugalpotentialdescribesanunstablesituation. /squaresolid Thermodynamics — Exact Differentials In thermodynamics, which is sometimes called a search for exact differentials, we en- counterequationsoftheform df=P(x,y)dx+Q(x,y)dy. (1.133a) The usual problem is to determine whetherintegraltext (P(x,y)dx+Q(x,y)dy) depends only on the endpoints, that is, whether dfis indeed an exact differential. The necessary and suffi- cientconditionis that df=∂f ∂xdx+∂f ∂ydy (1.133b) 1.13 Potential Theory 73 orthat P(x,y)=∂f/∂x, Q(x,y)=∂f/∂y.(1.133c) Equations(1.133c)dependonsatisfyingtherelation ∂P(x,y) ∂y=∂Q(x,y) ∂x. (1.133d) This, however, is exactly analogous to Eq. (1.119), the requirement that Fbe irrotational. Indeed,the z-componentof Eq.(1.119) yields ∂Fx ∂y=∂Fy ∂x, (1.133e) with Fx=∂f ∂x,F y=∂f ∂y. VectorPotential In some branches of physics, especially electrodynamics, it is convenient to introduce a vectorpotential Asuchthata(force) field Bisgivenby B=∇×A. (1.134) Clearly, if Eq. (1.134) holds, ∇·B=0 by Eq. (1.84) and Bis solenoidal. Here we want to develop a converse, to show that when Bis solenoidal a vector potential Aexists. We demonstrate the existence of Aby actually calculating it. Suppose B=ˆxb1+ˆyb2+ˆzb3 andourunknown A=ˆxa1+ˆya2+ˆza3. ByEq. (1.134), ∂a3 ∂y−∂a2 ∂z=b1, (1.135a) ∂a1 ∂z−∂a3 ∂x=b2, (1.135b) ∂a2 ∂x−∂a1 ∂y=b3. (1.135c) Let us assume that the coordinates have been chosen so that Ais parallel to the yz-plane; thatis,a1=0.24Then b2=−∂a3 ∂x b3=∂a2 ∂x.(1.136) 24Clearly, this can be done at any one point. It is not at all obvious that this assumption will hold at all points; that is, Awill be two-dimensional. The justification for the assumption is that it works; Eq.(1.141) satisfies Eq.(1.134). 74 Chapter 1 Vector Analysis Integrating,weobtain a2=integraldisplayx x0b3dx+f2(y,z), a3=−integraldisplayx x0b2dx+f3(y,z),(1.137) wheref2andf3are arbitrary functions of yandzbutnotfunctions of x. These two equationscanbecheckedbydifferentiatingandrecoveringEq.(1.136).Equation(1.135a) becomes25 ∂a3 ∂y−∂a2 ∂z=−integraldisplayx x0parenleftbigg∂b2 ∂y+∂b3 ∂zparenrightbigg dx+∂f3 ∂y−∂f2 ∂z =integraldisplayx x0∂b1 ∂xdx+∂f3 ∂y−∂f2 ∂z, (1.138) using∇·B=0.Integratingwithrespectto x,weobtain ∂a3 ∂y−∂a2 ∂z=b1(x,y,z)−b1(x0,y,z)+∂f3 ∂y−∂f2 ∂z. (1.139) Rememberingthat f3andf2are arbitraryfunctionsof yandz, wechoose f2=0, f3=integraldisplayy y0b1(x0,y,z)dy,(1.140) so that the right-hand side of Eq. (1.139) reduces to b1(x,y,z), in agreement with Eq.(1.135a). With f2andf3givenbyEq.(1.140), wecanconstruct A: A=ˆyintegraldisplayx x0b3(x,y,z)dx+ˆzbracketleftbiggintegraldisplayy y0b1(x0,y,z)dy−integraldisplayx x0b2(x,y,z)dxbracketrightbigg .(1.141) However,thisisnotquitecomplete.Wemayaddanyconstantsince Bisaderivativeof A. What is much more important, we may add any gradient of a scalar function ∇ϕwithout affecting Bat all. Finally, the functions f2andf3are not unique. Other choices could have been made. Instead of setting a1=0 to get Eq. (1.136) any cyclic permutation of 1,2,3,x,y,z,x 0,y0,z0wouldalsowork. Example 1.13.3 AM AGNETIC VECTOR POTENTIAL FOR A CONSTANT MAGNETIC FIELD To illustrate the construction of a magnetic vector potential, we take the special but still importantcaseof aconstantmagneticinduction B=ˆzBz, (1.142) 25Leibniz’formula in Exercise9.6.13 is useful here. 1.13 Potential Theory 75 inwhich Bzisaconstant.Equations(1.135atoc)become ∂a3 ∂y−∂a2 ∂z=0, ∂a1 ∂z−∂a3 ∂x=0, (1.143) ∂a2 ∂x−∂a1 ∂y=Bz. If weassumethat a1=0,as before,thenbyEq. (1.141) A=ˆyintegraldisplayx Bzdx=ˆyxBz, (1.144) setting a constant of integration equal to zero. It can readily be seen that this Asatisfies Eq.(1.134). To show that the choice a1=0 was not sacred or at least not required, let us try setting a3=0.FromEq.(1.143) ∂a2 ∂z=0, (1.145a) ∂a1 ∂z=0, (1.145b) ∂a2 ∂x−∂a1 ∂y=Bz. (1.145c) We seea1anda2areindependentof z,or a1=a1(x,y), a 2=a2(x,y). (1.146) Equation(1.145c)issatisfiedifwetake a2=pintegraldisplayx Bzdx=pxBz (1.147) and a1=(p−1)integraldisplayy Bzdy=(p−1)yBz, (1.148) withpanyconstant.Then A=ˆx(p−1)yBz+ˆypxBz. (1.149) Again, Eqs. (1.134), (1.142), and (1.149) are seen to be consistent. Comparison of Eqs. (1.144) and (1.149) shows immediately that Ais not unique. The difference between Eqs. (1.144) and (1.149) and the appearance of the parameter pin Eq. (1.149) may be accountedfor byrewritingEq. (1.149)as A=−1 2(ˆxy−ˆyx)Bz+parenleftbigg p−1 2parenrightbigg (ˆxy+ˆyx)Bz =−1 2(ˆxy−ˆyx)Bz+parenleftbigg p−1 2parenrightbigg Bz∇ϕ (1.150) 76 Chapter 1 Vector Analysis with ϕ=xy. (1.151) /squaresolid The first term in Acorrespondstotheusualform A=1 2(B×r) (1.152) forB,a constant. Adding a gradient of a scalar function, /Lambda1say, to the vector potential Adoes not affect B,byEq.(1.82);thisisknownasagaugetransformation(seeExercises1.13.9and4.6.4): A→A′=A+∇/Lambda1. (1.153) Suppose now that the wave function ψ0solves the Schrödinger equation of quantum mechanicswithoutmagneticinductionfield B, braceleftbigg1 2m(−i¯h∇)2+V−Ebracerightbigg ψ0=0, (1.154) describingaparticlewithmass mandcharge e.WhenBisswitchedon,thewaveequation becomes braceleftbigg1 2m(−i¯h∇−eA)2+V−Ebracerightbigg ψ=0. (1.155) Its solution ψpicksupaphasefactor thatdependsonthecoordinatesingeneral, ψ(r)=expbracketleftbiggie ¯hintegraldisplayr A(r′)·dr′bracketrightbigg ψ0(r). (1.156) Fromtherelation (−i¯h∇−eA)ψ=expbracketleftbiggie ¯hintegraldisplay A·dr′bracketrightbiggbraceleftbigg (−i¯h∇−eA)ψ0−i¯hψ0ie ¯hAbracerightbigg =expbracketleftbiggie ¯hintegraldisplay A·dr′bracketrightbigg (−i¯h∇ψ0), (1.157) itisobviousthat ψsolvesEq.(1.155)if ψ0solvesEq.(1.154).The gaugecovariantderiv- ative∇−i(e/¯h)Adescribesthecouplingofachargedparticlewiththemagneticfield.Itis often called minimal substitution and plays a central role in quantum electromagnetism, thefirst andsimplestgaugetheoryinphysics. To summarize this discussion of the vector potential :When a vector Bis solenoidal, a vector potential Aexists such that B=∇×A.Ais undetermined to within an additive gradient.Thiscorrespondstothearbitraryzeroofapotential,aconstantofintegrationfor thescalarpotential . In many problems the magnetic vector potential Awill be obtained from the current distributionthatproducesthemagneticinduction B.ThismeanssolvingPoisson’s(vector) equation(see Exercise1.14.4). 1.13 Potential Theory 77 Exercises 1.13.1 If aforce Fisgivenby F=parenleftbig x2+y2+z2parenrightbign(ˆxx+ˆyy+ˆzz), find (a)∇·F. (b)∇×F. (c) Ascalarpotential ϕ(x,y,z) so thatF=−∇ϕ. (d) Forwhatvalueoftheexponent ndoesthescalarpotentialdivergeatboththeorigin andinfinity? ANS.(a) (2n+3)r2n,(b)0, (c)−1 2n+2r2n+2,n/negationslash=−1,(d)n=−1, ϕ=−lnr. 1.13.2 A sphere of radius ais uniformly charged (throughout its volume). Construct the elec- trostaticpotential ϕ(r)for 0/lessorequalslantr<∞. Hint. In Section 1.14 it is shown that the Coulomb force on a test charge at r=r0 depends only on the charge at distances less than r0and is independent of the charge at distances greater than r0. Note that this applies to a spherically symmetric charge distribution. 1.13.3 The usual problem in classical mechanics is to calculate the motion of a particle given the potential. For a uniform density ( ρ0), nonrotating massive sphere, Gauss’ law of Section 1.14 leads to a gravitational force on a unit mass m0at a point r0produced by theattractionofthemassat r/lessorequalslantr0.Themassat r>r0contributesnothingtotheforce. (a) Showthat F/m0=−(4πGρ0/3)r,0/lessorequalslantr/lessorequalslanta,whereaistheradiusofthesphere. (b) Findthecorrespondinggravitationalpotential, 0 /lessorequalslantr/lessorequalslanta. (c) ImagineaverticalholerunningcompletelythroughthecenteroftheEarthandout tothefarside.NeglectingtherotationoftheEarthandassumingauniformdensity ρ0=5.5gm/cm3,calculatethenatureofthemotionofaparticledroppedintothe hole.Whatisits period? Note.F∝ris actually a very poor approximation. Because of varying density, the approximation F=constant along the outer half of a radial line and F∝r alongtheinnerhalfis amuchcloserapproximation. 1.13.4 The origin of the Cartesian coordinates is at the Earth’s center. The moon is on the z- axis, a fixed distance Raway (center-to-center distance). The tidal force exerted by the moononaparticleattheEarth’ssurface (point x,y,z)i sg i v e nb y Fx=−GMmx R3,F y=−GMmy R3,F z=+2GMmz R3. Findthepotentialthatyieldsthistidalforce. 78 Chapter 1 Vector Analysis ANS.−GMm R3parenleftbigg z2−1 2x2−1 2y2parenrightbigg . In termsoftheLegendrepolynomialsof Chapter12thisbecomes −GMm R3r2P2(cosθ). 1.13.5 A long, straight wire carrying a current Iproduces a magnetic induction Bwith com- ponents B=µ0I 2πparenleftbigg −y x2+y2,x x2+y2,0parenrightbigg . Findamagneticvectorpotential A. ANS.A=−ˆz(µ0I/4π)ln(x2+y2).(Thissolutionisnotunique.) 1.13.6 If B=ˆr r2=parenleftbiggx r3,y r3,z r3parenrightbigg , finda vector Asuchthat ∇×A=B.Onepossiblesolutionis A=ˆxyz r(x2+y2)−ˆyxz r(x2+y2). 1.13.7 Showthatthepairofequations A=1 2(B×r),B=∇×A issatisfiedbyanyconstantmagneticinduction B. 1.13.8 VectorBis formedbytheproductoftwogradients B=(∇u)×(∇v), whereuandvarescalarfunctions. (a) Showthat Bis solenoidal. (b) Showthat A=1 2(u∇v−v∇u) isavectorpotentialfor B, inthat B=∇×A. 1.13.9 The magnetic induction Bis related to the magnetic vector potential AbyB=∇×A. ByStokes’theorem integraldisplay B·dσ=contintegraldisplay A·dr. 1.14 Gauss’ Law, Poisson’s Equation 79 Showthateachsideofthisequationisinvariantunderthe gaugetransformation ,A→ A+∇ϕ. Note. Take the function ϕto be single-valued. The complete gauge transformation is consideredinExercise4.6.4. 1.13.10 WithEtheelectricfieldand Athemagneticvectorpotential,showthat [E+∂A/∂t]is irrotationalandthatthereforewemaywrite E=−∇ϕ−∂A ∂t. 1.13.11 Thetotalforce ona charge qmovingwithvelocity vis F=q(E+v×B). Usingthescalarandvectorpotentials,showthat F=qbracketleftbigg −∇ϕ−dA dt+∇(A·v)bracketrightbigg . Note that we now have a total time derivative of Ain place of the partial derivative of Exercise1.13.10. 1.14 G AUSS ’LAW,POISSON ’SEQUATION Gauss’ Law Considerapointelectriccharge qattheoriginofourcoordinatesystem.Thisproducesan electricfield Egivenby26 E=qˆr 4πε0r2. (1.158) We now derive Gauss’ law, whichstates thatthe surface integralin Fig. 1.35 is q/ε0if the closedsurface S=∂Vincludestheorigin(where qislocated)andzeroifthesurfacedoes notincludetheorigin.Thesurface Sisanyclosedsurface; itneednotbespherical. Using Gauss’ theorem, Eqs. (1.101a) and (1.101b) (and neglecting the q/4πε0), we obtain integraldisplay Sˆr·dσ r2=integraldisplay V∇·parenleftbiggˆr r2parenrightbigg dτ=0 (1.159) byExample1.7.2,providedthesurface Sdoesnotincludetheorigin,wheretheintegrands arenotdefined.ThisprovesthesecondpartofGauss’law. The first part, in which the surface Smust include the origin, may be handled by sur- rounding the origin with a small sphere S′=∂V′of radius δ(Fig. 1.36). So that there will be no question what is inside and what is outside, imagine the volume outside the outer surface Sand the volume inside surface S′(r <δ)connected by a small hole. This 26Theelectricfield Eisdefinedastheforceperunitchargeonasmallstationarytestcharge qt:E=F/qt.FromCoulomb’slaw the force on qtdue toqisF=(qqt/4πε0)(ˆr/r2).Wh enwed i v i d eb y qt, Eq.(1.158) follows. 80 Chapter 1 Vector Analysis FIGURE 1.35Gauss’ law. FIGURE 1.36Exclusionof theorigin. joins surfaces SandS′, combining them into one single simply connected closed surface. Because the radius of the imaginary hole may be made vanishingly small, there is no ad- ditional contribution to the surface integral. The inner surface is deliberately chosen to be 1.14 Gauss’ Law, Poisson’s Equation 81 spherical so that we will be able to integrate over it. Gauss’ theorem now applies to the volumebetween SandS′withoutanydifficulty.Wehave integraldisplay Sˆr·dσ r2+integraldisplay S′ˆr·dσ′ δ2=0. (1.160) We may evaluate the second integral, for dσ′=−ˆrδ2d/Omega1, in which d/Omega1is an element of solidangle.TheminussignappearsbecauseweagreedinSection1.10tohavethepositive normalˆr′outward from the volume. In this case the outward ˆr′is in the negative radial direction,ˆr′=−ˆr. Byintegratingoverallangles,wehave integraldisplay S′ˆr·dσ′ δ2=−integraldisplay S′ˆr·ˆrδ2d/Omega1 δ2=−4π, (1.161) independentof theradius δ. WiththeconstantsfromEq. (1.158), thisresultsin integraldisplay SE·dσ=q 4πε04π=q ε0, (1.162) completing the proof of Gauss’ law. Notice that although the surface Smay be spherical, itneednot bespherical.Goingjustabitfurther, weconsideradistributedchargeso that q=integraldisplay Vρdτ. (1.163) Equation (1.162) still applies, with qnow interpreted as the total distributed charge en- closedbysurface S: integraldisplay SE·dσ=integraldisplay Vρ ε0dτ. (1.164) UsingGauss’ theorem,wehave integraldisplay V∇·Edτ=integraldisplay Vρ ε0dτ. (1.165) Sinceourvolumeiscompletelyarbitrary, theintegrandsmustbeequal,or ∇·E=ρ ε0, (1.166) one of Maxwell’s equations. If we reverse the argument, Gauss’ law follows immediately fromMaxwell’sequation. Poisson’s Equation If wereplace Eby−∇ϕ, Eq. (1.166)becomes ∇·∇ϕ=−ρ ε0, (1.167a) 82 Chapter 1 Vector Analysis which is Poisson’s equation. For the condition ρ=0 this reduces to an even more famous equation, ∇·∇ϕ=0, (1.167b) Laplace’s equation. We encounter Laplace’s equation frequently in discussing various co- ordinatesystems(Chapter2)andthespecialfunctionsofmathematicalphysicsthatappear as its solutions. Poisson’s equation will be invaluable in developing the theory of Green’s functions(Section9.7). From direct comparison of the Coulomb electrostatic force law and Newton’s law of universalgravitation, FE=1 4πε0q1q2 r2ˆr,FG=−Gm1m2 r2ˆr. All of the potential theory of this section applies equally well to gravitational potentials. Forexample,thegravitationalPoissonequationis ∇·∇ϕ=+4πGρ, (1.168) withρnowamass density. Exercises 1.14.1 DevelopGauss’ lawfor thetwo-dimensionalcaseinwhich ϕ=−qlnρ 2πε0,E=−∇ϕ=qˆρ 2πε0ρ. Hereqisthechargeattheoriginorthelinechargeperunitlengthifthetwo-dimensional systemisaunitthicknesssliceofathree-dimensional(circularcylindrical)system.The variableρis measured radially outward from the line charge. ˆρis the corresponding unitvector(seeSection2.4). 1.14.2 (a) ShowthatGauss’lawfollowsfrom Maxwell’sequation ∇·E=ρ ε0. Hereρistheusualchargedensity. (b) Assumingthattheelectricfieldofapointcharge qissphericallysymmetric,show thatGauss’ lawimpliestheCoulombinversesquareexpression E=qˆr 4πε0r2. 1.14.3 Showthatthevalueoftheelectrostaticpotential ϕatanypoint Pisequaltotheaverage of the potential over any spherical surface centered on P. There are no electric charges onorwithinthesphere. Hint.UseGreen’stheorem,Eq.(1.104),with u−1=r,thedistancefrom P,andv=ϕ. AlsonoteEq. (1.170)inSection1.15. 1.15 Dirac Delta Function 83 1.14.4 UsingMaxwell’sequations,showthatforasystem(steadycurrent)themagneticvector potential Asatisfies avectorPoissonequation, ∇2A=−µ0J, providedwerequire ∇·A=0. 1.15 D IRAC DELTA FUNCTION FromExample1.6.1 andthedevelopmentofGauss’lawinSection1.14, integraldisplay ∇·∇parenleftbigg1 rparenrightbigg dτ=−integraldisplay ∇·parenleftbiggˆr r2parenrightbigg dτ=braceleftbigg−4π 0,(1.169) depending on whether or not the integration includes the origin r=0. This result may be convenientlyexpressedbyintroducingtheDiracdeltafunction, ∇2parenleftbigg1 rparenrightbigg =−4πδ(r)≡−4πδ(x)δ(y)δ(z). (1.170) ThisDiracdeltafunctionis definedbyits assignedproperties δ(x)=0,x/negationslash=0 (1.171a) f(0)=integraldisplay∞ −∞f(x)δ(x)dx, (1.171b) wheref(x)is any well-behaved function and the integration includes the origin. As a specialcaseofEq. (1.171b), integraldisplay∞ −∞δ(x)dx=1. (1.171c) FromEq.(1.171b), δ(x)mustbeaninfinitelyhigh,infinitelythinspikeat x=0,asinthe descriptionofanimpulsiveforce(Section15.9)orthechargedensityforapointcharge.27 The problem is that no such function exists , in the usual sense of function. However, the crucial property in Eq. (1.171b) can be developed rigorously as the limit of a sequence of functions, a distribution. For example, the delta function may be approximated by the 27The delta function is frequently invoked to describe very short-range forces, such as nuclear forces. It also appears in the normalization of continuum wavefunctions of quantum mechanics.Compare Eq.(1.193c) for plane-waveeigenfunctions. 84 Chapter 1 Vector Analysis FIGURE 1.37δ-Sequence function. FIGURE 1.38δ-Sequence function. sequencesoffunctions,Eqs. (1.172)to(1.175) andFigs. 1.37to1.40: δn(x)=  0,x<−1 2n n,−1 2n<x<1 2n 0,x>1 2n(1.172) δn(x)=n√πexpparenleftbig −n2x2parenrightbig (1.173) δn(x)=n π·1 1+n2x2(1.174) δn(x)=sinnx πx=1 2πintegraldisplayn −neixtdt. (1.175) 1.15 Dirac Delta Function 85 FIGURE 1.39δ-Sequencefunction. FIGURE 1.40δ-Sequencefunction. These approximations have varying degrees of usefulness. Equation (1.172) is useful in providing a simple derivation of the integral property, Eq. (1.171b). Equation (1.173) is convenient to differentiate. Its derivatives lead to the Hermite polynomials. Equa- tion (1.175) is particularly useful in Fourier analysis and in its applications to quantum mechanics. In the theory of Fourier series, Eq. (1.175) often appears (modified) as the Dirichletkernel: δn(x)=1 2πsin[(n+1 2)x] sin(1 2x). (1.176) In using these approximations in Eq. (1.171b) and later, we assume that f(x)is well be- haved—itoffers noproblemsatlarge x. 86 Chapter 1 Vector Analysis For most physical purposes such approximations are quite adequate. From a mathemat- icalpointofviewthesituationis stillunsatisfactory:Thelimits limn→∞δn(x) donotexist . A way out of this difficulty is provided by the theory of distributions. Recognizing that Eq. (1.171b) is the fundamental property, we focus our attention on it rather than on δ(x) itself.Equations(1.172)to(1.175)with n=1,2,3,...maybeinterpretedas sequences of normalizedfunctions: integraldisplay∞ −∞δn(x)dx=1. (1.177) Thesequenceofintegralshasthelimit limn→∞integraldisplay∞ −∞δn(x)f(x)dx=f(0). (1.178) Note that Eq. (1.178) is the limit of a sequence of integrals. Again, the limit of δn(x), n→∞, doesnotexist.(Thelimitsforallfour formsof δn(x)divergeat x=0.) We maytreat δ(x)consistentlyintheform integraldisplay∞ −∞δ(x)f(x)dx=limn→∞integraldisplay∞ −∞δn(x)f(x)dx. (1.179) δ(x)is labeled a distribution (not a function) defined by the sequences δn(x)as indicated inEq.(1.179).Wemightemphasizethattheintegralontheleft-handsideofEq.(1.179)is nota Riemannintegral.28It isalimit. Thisdistribution δ(x)isonlyoneofaninfinityofpossibledistributions,butitistheone weare interestedinbecauseofEq. (1.171b). FromthesesequencesoffunctionsweseethatDirac’sdeltafunctionmustbeevenin x, δ(−x)=δ(x). The integral property, Eq. (1.171b), is useful in cases where the argument of the delta functionisafunction g(x)withsimplezerosontherealaxis, whichleadstotherules δ(ax)=1 aδ(x), a> 0, (1.180) δparenleftbig g(x)parenrightbig =summationdisplay a, g(a)=0, g′(a)/negationslash=0δ(x−a) |g′(a)|. (1.181a) Equation(1.180)maybewritten integraldisplay∞ −∞f(x)δ(ax)dx=1 aintegraldisplay∞ −∞fparenleftbiggy aparenrightbigg δ(y)dy=1 af(0), 28It can be treated as a Stieltjes integral if desired. δ(x)dxis replaced by du(x),w h e r eu(x)is the Heaviside step function (compare Exercise1.15.13). 1.15 Dirac Delta Function 87 applying Eq. (1.171b). Equation (1.180) may be written as δ(ax)=1 |a|δ(x)fora<0.To proveEq.(1.181a)wedecomposetheintegral integraldisplay∞ −∞f(x)δparenleftbig g(x)parenrightbig dx=summationdisplay aintegraldisplaya+ε a−εf(x)δparenleftbig (x−a)g′(a)parenrightbig dx (1.181b) intoasumofintegralsoversmallintervalscontainingthezerosof g(x).Intheseintervals, g(x)≈g(a)+(x−a)g′(a)=(x−a)g′(a). Using Eq. (1.180) on the right-hand side of Eq.(1.181b)weobtaintheintegralofEq. (1.181a). Using integration by parts we can also define the derivative δ′(x)of the Dirac delta functionbytherelation integraldisplay∞ −∞f(x)δ′(x−x′)dx=−integraldisplay∞ −∞f′(x)δ(x−x′)dx=−f′(x′). (1.182) We useδ(x)frequently and call it the Dirac delta function29—for historical reasons. Remember that it is not really a function. It is essentially a shorthand notation, defined implicitlyas thelimit of integralsina sequence, δn(x), accordingto Eq. (1.179). It should be understood that our Dirac delta function has significance only as part of an integrand. Inthisspirit, thelinearoperatorintegraltext dxδ(x−x0)operateson f(x)andyields f(x0): L(x0)f(x)≡integraldisplay∞ −∞δ(x−x0)f(x)dx=f(x0). (1.183) It may also be classified as a linear mapping or simply as a generalized function. Shift- ing our singularity to the point x=x′, we write the Dirac delta function as δ(x−x′). Equation(1.171b)becomes integraldisplay∞ −∞f(x)δ(x−x′)dx=f(x′). (1.184) As a description of a singularity at x=x′, the Dirac delta function may be written as δ(x−x′)orasδ(x′−x).Goingtothreedimensionsandusingsphericalpolarcoordinates, weobtain integraldisplay2π 0integraldisplayπ 0integraldisplay∞ 0δ(r)r2drsinθdθdϕ=integraldisplayintegraldisplayintegraldisplay∞ −∞δ(x)δ(y)δ(z)dxdydz =1.(1.185) Thiscorrespondstoasingularity(orsource)attheorigin.Again,ifoursourceisat r=r1, Eq.(1.185) becomes integraldisplayintegraldisplayintegraldisplay δ(r2−r1)r2 2dr2sinθ2dθ2dϕ2=1. (1.186) 29Diracintroducedthedeltafunctiontoquantummechanics.Actually,thedeltafunctioncanbetracedbacktoKirchhoff,1882. For further details see M. Jammer, The Conceptual Development of Quantum Mechanics . New York: McGraw–Hill (1966), p. 301. 88 Chapter 1 Vector Analysis Example 1.15.1 TOTAL CHARGE INSIDE A SPHERE Consider the total electric fluxcontintegraltext E·dσout of a sphere of radius Raround the origin surrounding nchargesej,located at the points rjwithrj<R, that is, inside the sphere. Theelectricfieldstrength E=−∇ϕ(r),wherethepotential ϕ=nsummationdisplay j=1ej |r−rj|=integraldisplayρ(r′) |r−r′|d3r′ isthesumoftheCoulombpotentialsgeneratedbyeachchargeandthetotalchargedensity isρ(r)=summationtext jejδ(r−rj).Thedeltafunctionisusedhereasanabbreviationofapointlike density.NowweuseGauss’theoremfor contintegraldisplay E·dσ=−contintegraldisplay ∇ϕ·dσ=−integraldisplay ∇2ϕdτ=integraldisplayρ(r) ε0dτ=summationtext jej ε0 inconjunctionwiththedifferentialform ofGauss’slaw, ∇·E=−ρ/ε0,and summationdisplay jejintegraldisplay δ(r−rj)dτ=summationdisplay jej. /squaresolid Example 1.15.2 PHASE SPACE InthescatteringtheoryofrelativisticparticlesusingFeynmandiagrams,weencounterthe followingintegraloverenergyof thescatteredparticle(wesetthevelocityoflight c=1): integraldisplay d4pδparenleftbig p2−m2parenrightbig f(p)≡integraldisplay d3pintegraldisplay dp0δparenleftbig p2 0−p2−m2parenrightbig f(p) =integraldisplay E>0d3pf(E,p) 2radicalbig m2+p2+integraldisplay E<0d3pf(E,p) 2radicalbig m2+p2, where we have used Eq. (1.181a) at the zeros E=±radicalbig m2+p2of the argument of the delta function. The physical meaning of δ(p2−m2)is that the particle of mass mand four-momentum pµ=(p0,p)is on its mass shell, because p2=m2is equivalent to E= ±radicalbig m2+p2. Thus, the on-mass-shell volume element in momentum space is the Lorentz invariantd3p 2E, in contrast to the nonrelativistic d3pof momentum space. The fact that a negative energy occurs is a peculiarity of relativistic kinematics that is related to the antiparticle. /squaresolid Delta Function Representation by Orthogonal Functions Dirac’sdeltafunction30canbeexpandedintermsofanybasisofrealorthogonalfunctions {ϕn(x),n=0,1,2,...}. Such functions will occur in Chapter 10 as solutions of ordinary differentialequationsof theSturm–Liouvilleform. 30This sectionis optional here. It is not needed until Chapter10. 1.15 Dirac Delta Function 89 Theysatisfytheorthogonalityrelations integraldisplayb aϕm(x)ϕn(x)dx=δmn, (1.187) where the interval (a,b)may be infinite at either end or both ends. [For convenience we assumethat ϕnhasbeendefinedtoinclude (w(x))1/2iftheorthogonalityrelationscontain an additional positive weight function w(x).] We use the ϕnto expand the delta function as δ(x−t)=∞summationdisplay n=0an(t)ϕn(x), (1.188) where the coefficients anare functions of the variable t. Multiplying by ϕm(x)and inte- gratingovertheorthogonalityinterval(Eq.(1.187)), wehave am(t)=integraldisplayb aδ(x−t)ϕm(x)dx=ϕm(t) (1.189) or δ(x−t)=∞summationdisplay n=0ϕn(t)ϕn(x)=δ(t−x). (1.190) This series is assuredly not uniformly convergent (see Chapter 5), but it may be used as part of an integrand in which the ensuing integration will make it convergent (compare Section5.5). Suppose we form the integralintegraltext F(t)δ(t−x)dx, where it is assumed that F(t)can be expanded in a series of orthogonal functions ϕp(t), a property called completeness .W e thenobtain integraldisplay F(t)δ(t−x)dt=integraldisplay∞summationdisplay p=0apϕp(t)∞summationdisplay n=0ϕn(x)ϕn(t)dt =∞summationdisplay p=0apϕp(x)=F(x), (1.191) the cross productsintegraltext ϕpϕndt(n/negationslash=p)vanishing by orthogonality (Eq. (1.187)). Referring back to the definition of the Dirac delta function, Eq. (1.171b), we see that our series representation, Eq. (1.190), satisfies the defining property of the Dirac delta function and therefore is a representation of it. This representation of the Dirac delta function is called closure. The assumption of completeness of a set of functions for expansion of δ(x−t) yieldstheclosurerelation.Theconverse,thatclosureimpliescompleteness,isthetopicof Exercise1.15.16. 90 Chapter 1 Vector Analysis Integral Representations for the Delta Function Integraltransforms, suchastheFourierintegral F(ω)=integraldisplay∞ −∞f(t)exp(iωt)dt of Chapter 15, lead to thecorrespondingintegralrepresentationsof Dirac’s deltafunction. Forexample,take δn(t−x)=sinn(t−x) π(t−x)=1 2πintegraldisplayn −nexpparenleftbig iω(t−x)parenrightbig dω, (1.192) usingEq. (1.175). Wehave f(x)=limn→∞integraldisplay∞ −∞f(t)δn(t−x)dt, (1.193a) whereδn(t−x)isthesequenceinEq.(1.192)definingthedistribution δ(t−x).Notethat Eq. (1.193a) assumes that f(t)is continuous at t=x. If we substitute Eq. (1.192) into Eq.(1.193a)weobtain f(x)=limn→∞1 2πintegraldisplay∞ −∞f(t)integraldisplayn −nexpparenleftbig iω(t−x)parenrightbig dωdt. (1.193b) Interchanging the order of integration and then taking the limit as n→∞,w eh a v et h e Fourierintegraltheorem,Eq.(15.20). With the understanding that it belongs under an integral sign, as in Eq. (1.193a), the identification δ(t−x)=1 2πintegraldisplay∞ −∞expparenleftbig iω(t−x)parenrightbig dω (1.193c) providesaveryusefulintegralrepresentationof thedeltafunction. WhentheLaplacetransform(see Sections15.1and15.9) Lδ(s)=integraldisplay∞ 0exp(−st)δ(t−t0)=exp(−st0), t 0>0 (1.194) isinverted,weobtainthecomplexrepresentation δ(t−t0)=1 2πiintegraldisplayγ+i∞ γ−i∞expparenleftbig s(t−t0)parenrightbig ds, (1.195) whichisessentiallyequivalenttothepreviousFourierrepresentationofDirac’sdeltafunc- tion. 1.15 Dirac Delta Function 91 Exercises 1.15.1 Let δn(x)=  0,x<−1 2n, n,−1 2n<x<1 2n, 0,1 2n<x. Showthat limn→∞integraldisplay∞ −∞f(x)δn(x)dx=f(0), assumingthat f(x)iscontinuousat x=0. 1.15.2 Verifythatthesequence δn(x), basedonthefunction δn(x)=braceleftbigg0,x <0, ne−nx,x>0, is a delta sequence (satisfying Eq. (1.178)). Note that the singularity is at +0, the posi- tivesideof theorigin. Hint. Replace the upper limit ( ∞)b yc/n, wherecis large but finite, and use the mean valuetheoremof integralcalculus. 1.15.3 For δn(x)=n π·1 1+n2x2, (Eq. (1.174)), showthat integraldisplay∞ −∞δn(x)dx=1. 1.15.4 Demonstratethat δn=sinnx/πxisadeltadistributionbyshowingthat limn→∞integraldisplay∞ −∞f(x)sinnx πxdx=f(0). Assumethat f(x)is continuousat x=0 andvanishesas x→±∞. Hint.Replace xbyy/nandtake lim n→∞beforeintegrating. 1.15.5 Fejer’smethodofsummingseries isassociatedwiththefunction δn(t)=1 2πnbracketleftbiggsin(nt/2) sin(t/2)bracketrightbigg2 . Showthat δn(t)isadeltadistribution,inthesensethat limn→∞1 2πnintegraldisplay∞ −∞f(t)bracketleftbiggsin(nt/2) sin(t/2)bracketrightbigg2 dt=f(0). 92 Chapter 1 Vector Analysis 1.15.6 Provethat δbracketleftbig a(x−x1)bracketrightbig =1 aδ(x−x1). Note.Ifδ[a(x−x1)]isconsideredeven,relativeto x1,therelationholdsfornegative a and 1/amaybereplacedby 1 /|a|. 1.15.7 Showthat δbracketleftbig (x−x1)(x−x2)bracketrightbig =bracketleftbig δ(x−x1)+δ(x−x2)bracketrightbig /|x1−x2|. Hint.TryusingExercise1.15.6. 1.15.8 UsingtheGausserror curvedeltasequence( δn=n√πe−n2x2), showthat xd dxδ(x)=−δ(x), treatingδ(x)anditsderivativeas inEq.(1.179). 1.15.9 Showthatintegraldisplay∞ −∞δ′(x)f(x)dx=−f′(0). Hereweassumethat f′(x)is continuousat x=0. 1.15.10 Provethat δparenleftbig f(x)parenrightbig =vextendsinglevextendsinglevextendsinglevextendsingledf(x) dxvextendsinglevextendsinglevextendsinglevextendsingle−1 x=x0δ(x−x0), wherex0ischosenso that f(x0)=0. Hint.Notethat δ(f)df=δ(x)dx. 1.15.11 Show that in spherical polar coordinates (r,cosθ,ϕ)the delta function δ(r1−r2)be- comes 1 r2 1δ(r1−r2)δ(cosθ1−cosθ2)δ(ϕ1−ϕ2). Generalize this to the curvilinear coordinates (q1,q2,q3)of Section 2.1 with scale fac- torsh1,h2, andh3. 1.15.12 Arigorousdevelopmentof Fouriertransforms31includesas atheoremtherelations lima→∞2 πintegraldisplayx2 x1f(u+x)sinax xdx =  f(u+0)+f(u−0), x 1<0<x2 f(u+0), x 1=0<x2 f(u−0), x 1<0=x2 0,x 1<x2<0or0<x1<x2. Verifytheseresults usingtheDiracdeltafunction. 31I. N.Sneddon, Fourier Transforms . NewYork: McGraw-Hill(1951). 1.15 Dirac Delta Function 93 FIGURE 1.411 2[1+tanhnx]andtheHeavisideunitstep function. 1.15.13 (a) If wedefinea sequence δn(x)=n/(2cosh2nx), showthat integraldisplay∞ −∞δn(x)dx=1,independentof n. (b) Continuingthisanalysis,showthat32 integraldisplayx −∞δn(x)dx=1 2[1+tanhnx]≡un(x), limn→∞un(x)=braceleftbigg0,x<0, 1,x>0. Thisis theHeavisideunitstepfunction(Fig.1.41). 1.15.14 Showthattheunitstepfunction u(x)mayberepresentedby u(x)=1 2+1 2πiPintegraldisplay∞ −∞eixtdt t, wherePmeansCauchyprincipalvalue(Section7.1). 1.15.15 Asavariationof Eq.(1.175), take δn(x)=1 2πintegraldisplay∞ −∞eixt−|t|/ndt. Showthatthisreducesto (n/π)1/(1+n2x2), Eq.(1.174), andthat integraldisplay∞ −∞δn(x)dx=1. Note. In terms of integral transforms, the initial equation here may be interpreted as eitheraFourierexponentialtransformof e−|t|/noraLaplacetransformof eixt. 32Manyothersymbolsareusedforthisfunction.ThisistheAMS-55(seefootnote4onp.330forthereference)notation: ufor unit. 94 Chapter 1 Vector Analysis 1.15.16 (a) TheDiracdeltafunctionrepresentationgivenbyEq.(1.190), δ(x−t)=∞summationdisplay n=0ϕn(x)ϕn(t), is often called the closure relation . For an orthonormal set of real functions, ϕn,show that closure implies completeness, that is, Eq. (1.191) follows from Eq.(1.190). Hint.Onecantake F(x)=integraldisplay F(t)δ(x−t)dt. (b) Following the hint of part (a) you encounter the integralintegraltext F(t)ϕn(t)dt.H o wd o youknowthatthisintegralis finite? 1.15.17 For the finite interval (−π,π)write the Dirac delta function δ(x−t)as a series of sines and cosines: sin nx,cosnx,n=0,1,2,....Note that although these functions areorthogonal,theyarenotnormalizedtounity. 1.15.18 Intheinterval (−π,π),δn(x)=n√πexp(−n2x2). (a) Write δn(x)asaFouriercosineseries. (b) Show that your Fourier series agrees with a Fourier expansion of δ(x)in the limit asn→∞. (c) Confirm the delta function nature of your Fourier series by showing that for any f(x)thatisfiniteintheinterval [−π,π]andcontinuousat x=0, integraldisplayπ −πf(x)bracketleftbig Fourierexpansionof δ∞(x)bracketrightbig dx=f(0). 1.15.19 (a) Write δn(x)=n√πexp(−n2x2)in the interval (−∞,∞)as a Fourier integral and comparethelimit n→∞withEq. (1.193c). (b) Write δn(x)=nexp(−nx)as a Laplace transform and compare the limit n→∞ withEq. (1.195). Hint.SeeEqs. (15.22) and(15.23)for (a) andEq.(15.212)for (b). 1.15.20 (a) Show that the Dirac delta function δ(x−a), expanded in a Fourier sine series in thehalf-interval (0,L),(0<a<L) , isgivenby δ(x−a)=2 L∞summationdisplay n=1sinparenleftbiggnπa Lparenrightbigg sinparenleftbiggnπx Lparenrightbigg . Notethatthisseriesactuallydescribes −δ(x+a)+δ(x−a)intheinterval (−L,L). (b) By integrating both sides of the preceding equation from 0 to x, show that the cosineexpansionofthesquarewave f(x)=braceleftbigg0,0/lessorequalslantx<a 1, a<x<L, 1.16 Helmholtz’s Theorem 95 is, for 0 /lessorequalslantx<L, f(x)=2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg −2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg cosparenleftbiggnπx Lparenrightbigg . (c) Verifythattheterm 2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg isangbracketleftbig f(x)angbracketrightbig ≡1 LintegraldisplayL 0f(x)dx. 1.15.21 Verify the Fourier cosine expansion of the square wave, Exercise 1.15.20(b), by direct calculationoftheFouriercoefficients. 1.15.22 Wemaydefineasequence δn(x)=braceleftbiggn,|x|<1/2n, 0,|x|>1/2n. (This is Eq. (1.172).) Express δn(x)as a Fourier integral (via the Fourier integral theo- rem,inversetransform,etc.). Finally,showthatwemaywrite δ(x)=limn→∞δn(x)=1 2πintegraldisplay∞ −∞e−ikxdk. 1.15.23 Usingthesequence δn(x)=n√πexpparenleftbig −n2x2parenrightbig , showthat δ(x)=1 2πintegraldisplay∞ −∞e−ikxdk. Note. Remember that δ(x)is defined in terms of its behavior as part of an integrand— especiallyEqs. (1.178)and(1.189). 1.15.24 Derivesineandcosinerepresentationsof δ(t−x)thatarecomparabletotheexponential representation,Eq. (1.193c). ANS.2 πintegraltext∞ 0sinωtsinωxdω,2 πintegraltext∞ 0cosωtcosωxdω. 1.16 H ELMHOLTZ ’STHEOREM InSection1.13itwasemphasizedthatthechoiceofamagneticvectorpotential Awasnot unique.Thedivergenceof Awasstillundetermined.Inthissectiontwotheoremsaboutthe divergenceandcurlofavectorare developed.Thefirst theoremisas follows: A vector is uniquely specified by giving its divergence and its curl within a simply con- nectedregion(withoutholes)anditsnormalcomponentovertheboundary . 96 Chapter 1 Vector Analysis Note that the subregions, where the divergence and curl are defined (often in terms of Dirac delta functions), are part of our region and are not supposed to be removed here or inHelmholtz’stheorem,whichfollows.Letustake ∇·V1=s, ∇×V1=c,(1.196) wheresmay be interpreted as a source (charge) density and cas a circulation (current) density. Assuming also that the normal component V1non the boundary is given, we want to show that V1is unique. We do this by assuming the existence of a second vector, V2, which satisfies Eq. (1.196) and has the same normal component over the boundary, and thenshowingthat V1−V2=0.Let W=V1−V2. Then ∇·W=0 (1.197) and ∇×W=0. (1.198) SinceWisirrotationalwemaywrite(bySection(1.13)) W=−∇ϕ. (1.199) SubstitutingthisintoEq.(1.197), weobtain ∇·∇ϕ=0, (1.200) Laplace’sequation. Now we draw upon Green’s theorem in the form given in Eq. (1.105), letting uandv eachequal ϕ. Since Wn=V1n−V2n=0 (1.201) ontheboundary,Green’stheoremreducesto integraldisplay V(∇ϕ)·(∇ϕ)dτ=integraldisplay VW·Wdτ=0. (1.202) Thequantity W·W=W2isnonnegativeandsowemusthave W=V1−V2=0 (1.203) everywhere.Thus V1is unique,provingthetheorem. For our magnetic vector potential Athe relation B=∇×Aspecifies the curl of A. Often for convenience we set ∇·A=0 (compare Exercise 1.14.4). Then (with boundary conditions) Aisfixed. This theorem may be written as a uniqueness theorem for solutions of Laplace’s equa- tion, Exercise 1.16.1. In this form, this uniqueness theorem is of great importance in solv- ing electrostatic and other Laplace equation boundary value problems. If we can find a solution of Laplace’s equation that satisfies the necessary boundary conditions, then our solution is the complete solution. Such boundary value problems are taken up in Sec- tions12.3and12.5. 1.16 Helmholtz’s Theorem 97 Helmholtz’s Theorem ThesecondtheoremweshallproveisHelmholtz’stheorem. A vector Vsatisfying Eq. (1.196)with both source and circulation densities vanishing atinfinitymaybewrittenas thesum oftwoparts, oneofwhichisirrotational,theotherof whichis solenoidal . Notethatourregionissimplyconnected,beingallofspace,forsimplicity.Helmholtz’s theoremwillclearlybesatisfiedif wemaywrite Vas V=−∇ϕ+∇×A, (1.204a) −∇ϕbeingirrotationaland ∇×Abeingsolenoidal.We proceedtojustifyEq.(1.204a). Vis aknownvector.Wetakethedivergenceandcurl ∇·V=s(r) (1.204b) ∇×V=c(r) (1.204c) withs(r)andc(r)nowknownfunctionsofposition.Fromthesetwofunctionsweconstruct ascalarpotential ϕ(r1), ϕ(r1)=1 4πintegraldisplays(r2) r12dτ2, (1.205a) andavectorpotential A(r1), A(r1)=1 4πintegraldisplayc(r2) r12dτ2. (1.205b) Ifs=0, thenVis solenoidal and Eq. (1.205a) implies ϕ=0. From Eq. (1.204a), V= ∇×A, withAas given in Eq. (1.141), which is consistent with Section 1.13. Further, ifc=0, thenVis irrotational and Eq. (1.205b) implies A=0, and Eq. (1.204a) implies V=−∇ϕ, consistentwithscalarpotentialtheoryof Section1.13. Here the argument r1indicates (x1,y1,z1), the field point; r2, the coordinates of the sourcepoint( x2,y2,z2), whereas r12=bracketleftbig (x1−x2)2+(y1−y2)2+(z1−z2)2bracketrightbig1/2. (1.206) When a direction is associated with r12, the positive direction is taken to be away from the source and toward the field point. Vectorially, r12=r1−r2, as shown in Fig. 1.42. Of course, sandcmust vanish sufficiently rapidly at large distance so that the integrals exist. The actual expansion and evaluation of integrals such as Eqs. (1.205a) and (1.205b) istreatedinSection12.1. From the uniqueness theorem at the beginning of this section, Vis uniquely specified by its divergence, s, and curl, c(and boundary conditions). Returning to Eq. (1.204a), we have ∇·V=−∇·∇ϕ, (1.207a) thedivergenceofthecurlvanishing,and ∇×V=∇×(∇×A), (1.207b) 98 Chapter 1 Vector Analysis FIGURE 1.42Sourceandfieldpoints. thecurlof thegradientvanishing.If wecanshowthat −∇·∇ϕ(r1)=s(r1) (1.207c) and ∇×parenleftbig ∇×A(r1)parenrightbig =c(r1), (1.207d) thenVas given in Eq. (1.204a) will have the proper divergence and curl. Our description willbeinternallyconsistentandEq.(1.204a) justified.33 First, weconsiderthedivergenceof V: ∇·V=−∇·∇ϕ=−1 4π∇·∇integraldisplays(r2) r12dτ2. (1.208) The Laplacian operator, ∇·∇,o r∇2, operates on the field coordinates (x1,y1,z1)and so commuteswiththeintegrationwithrespectto (x2,y2,z2).We have ∇·V=−1 4πintegraldisplay s(r2)∇2 1parenleftbigg1 r12parenrightbigg dτ2. (1.209) WemustmaketwominormodificationsinEq.(1.169)beforeapplyingit.First,oursource is atr2, not at the origin. This means that a nonzero result from Gauss’ law appears if and onlyif thesurface Sincludesthepoint r=r2. Toshowthis, werewriteEq. (1.170): ∇2parenleftbigg1 r12parenrightbigg =−4πδ(r1−r2). (1.210) 33Alternatively, we could solve Eq. (1.207c), Poisson’s equation, and compare the solution with the constructed potential, Eq.(1.205a). The solution of Poisson’s equation is developed in Section 9.7. 1.16 Helmholtz’s Theorem 99 Thisshift ofthesourceto r2maybeincorporatedinthedefiningequation(1.171b)as δ(r1−r2)=0,r1/negationslash=r2, (1.211a) integraldisplay f(r1)δ(r1−r2)dτ1=f(r2). (1.211b) Second, noting that differentiating r−1 12twice with respect to x2,y2,z2is the same as differentiating twicewithrespectto x1,y1,z1,weha v e ∇2 1parenleftbigg1 r12parenrightbigg =∇2 2parenleftbigg1 r12parenrightbigg =−4πδ(r1−r2) =−4πδ(r2−r1). (1.212) RewritingEq.(1.209)andusingtheDiracdeltafunction,Eq.(1.212), wemayintegrateto obtain ∇·V=−1 4πintegraldisplay s(r2)∇2 2parenleftbigg1 r12parenrightbigg dτ2 =−1 4πintegraldisplay s(r2)(−4π)δ(r2−r1)dτ2 =s(r1). (1.213) The final step follows from Eq. (1.211b), with the subscripts 1 and 2 exchanged. Our result, Eq. (1.213), shows that the assumed forms of Vand of the scalar potential ϕare in agreementwiththegivendivergence(Eq. (1.204b)). TocompletetheproofofHelmholtz’stheorem,weneedtoshowthatourassumptionsare consistentwithEq.(1.204c),thatis,thatthecurlof Visequalto c(r1).FromEq.(1.204a), ∇×V=∇×(∇×A) =∇∇·A−∇2A. (1.214) Thefirstterm, ∇∇·A,leadsto 4π∇∇·A=integraldisplay c(r2)·∇1∇1parenleftbigg1 r12parenrightbigg dτ2 (1.215) byEq.(1.205b).Againreplacingthesecondderivativeswithrespectto x1,y1,z1bysecond derivatives with respect to x2,y2,z2, we integrate each component34of Eq. (1.215) by parts: 4π∇∇·A|x=integraldisplay c(r2)·∇2∂ ∂x2parenleftbigg1 r12parenrightbigg dτ2 =integraldisplay ∇2·bracketleftbigg c(r2)∂ ∂x2parenleftbigg1 r12parenrightbiggbracketrightbigg dτ2 −integraldisplaybracketleftbig ∇2·c(r2)bracketrightbig∂ ∂x2parenleftbigg1 r12parenrightbigg dτ2. (1.216) 34This avoids creating the tensor c(r2)∇2. 100 Chapter 1 Vector Analysis The second integral vanishes because the circulation density cis solenoidal.35The first integral may be transformed to a surface integral by Gauss’ theorem. If cis bounded in space or vanishes faster that 1 /rfor large r, so that the integral in Eq. (1.205b) exists, then by choosing a sufficiently large surface the first integral on the right-hand side of Eq.(1.216) alsovanishes. With∇∇·A=0,Eq. (1.214)nowreducesto ∇×V=−∇2A=−1 4πintegraldisplay c(r2)∇2 1parenleftbigg1 r12parenrightbigg dτ2. (1.217) This is exactly like Eq. (1.209) except that the scalar s(r2)is replaced by the vector circu- lation density c(r2). Introducing the Dirac delta function, as before, as a convenient way ofcarryingouttheintegration,wefindthatEq.(1.217)reducestoEq.(1.196).Weseethat our assumed forms of V, given by Eq. (1.204a), and of the vector potential A, given by Eq.(1.205b), areinagreementwithEq. (1.196)specifyingthecurlof V. This completes the proof of Helmholtz’s theorem, showing that a vector may be re- solvedintoirrotationalandsolenoidalparts.Appliedtotheelectromagneticfield,wehave resolved our field vector Vinto an irrotational electric field E, derived from a scalar po- tentialϕ, and a solenoidal magnetic induction field B, derived from a vector potential A. The source density s(r)may be interpreted as an electric charge density (divided by elec- tric permittivity ε), whereas the circulation density c(r)becomes electric current density (timesmagneticpermeability µ). Exercises 1.16.1 Implicitinthissectionisaproofthatafunction ψ(r)isuniquelyspecifiedbyrequiring itto(1)satisfyLaplace’sequationand(2)satisfyacompletesetofboundaryconditions. Developthisproof explicitly. 1.16.2 (a) Assumingthat Pis a solutionof the vectorPoisson equation, ∇2 1P(r1)=−V(r1), developanalternateproofofHelmholtz’stheorem,showingthat Vmaybewritten as V=−∇ϕ+∇×A, where A=∇×P, and ϕ=∇·P. (b) SolvingthevectorPoissonequation,wefind P(r1)=1 4πintegraldisplay VV(r2) r12dτ2. Showthatthissolutionsubstitutedinto ϕandAofpart(a)leadstotheexpressions givenfor ϕandAinSection1.16. 35Remember, c=∇×Vis known. 1.16 Additional Readings 101 AdditionalReadings Borisenko,A.I.,andI.E.Taropov, VectorandTensorAnalysiswithApplications .EnglewoodCliffs,NJ:Prentice- Hall(1968). Reprinted, Dover (1980). Davis,H.F.,andA.D.Snider, Introduction to VectorAnalysis , 7th ed.Boston: Allyn & Bacon (1995). Kellogg, O. D., Foundations of Potential Theory . New York: Dover (1953). Originally published (1929). The classictext on potential theory. Lewis,P.E.,andJ.P.Ward, VectorAnalysisforEngineersandScientists .Reading,MA:Addison-Wesley(1989). Marion, J. B., Principles of Vector Analysis . New York: Academic Press (1965). A moderately advanced presen- tation of vector analysis oriented toward tensor analysis. Rotations and other transformations are described with the appropriate matrices. Spiegel, M.R., VectorAnalysis . NewYork: McGraw-Hill (1989). Tai, C.-T., Generalized VectorandDyadic Analysis . Oxford: Oxford University Press (1996). Wrede,R.C., IntroductiontoVectorandTensorAnalysis .NewYork:Wiley(1963).Reprinted,NewYork:Dover (1972). Fine historical introduction. Excellent discussion of differentiation of vectors and applications to me- chanics. This page intentionally left blank CHAPTER 2 VECTOR ANALYSIS IN CURVED COORDINATES ANDTENSORS InChapter1werestrictedourselvesalmostcompletelytorectangularorCartesiancoordi- natesystems.ACartesiancoordinatesystemofferstheuniqueadvantagethatallthreeunit vectors,ˆx,ˆy,andˆz,areconstantindirectionaswellasinmagnitude.Wedidintroducethe radial distance r, but even this was treated as a function of x,y, andz. Unfortunately, not allphysicalproblemsarewelladaptedtoasolutioninCartesiancoordinates.Forinstance, if we have a central force problem, F=ˆrF(r), such as gravitational or electrostatic force, Cartesiancoordinatesmaybeunusuallyinappropriate.Suchaproblemdemandstheuseof a coordinate system in which the radial distance is taken to be one of the coordinates, that is, sphericalpolarcoordinates. The point is that the coordinate system should be chosen to fit the problem, to exploit anyconstraintorsymmetrypresentinit.Thenitislikelytobemorereadilysolublethanif wehadforceditintoaCartesianframework. Naturally, there is a price that must be paid for the use of a non-Cartesian coordinate system. We have not yet written expressions for gradient, divergence, or curl in any of the non-Cartesiancoordinatesystems.SuchexpressionsaredevelopedingeneralforminSec- tion 2.2. First, we develop a system of curvilinear coordinates, a general system that may be specialized to any of the particular systems of interest. We shall specialize to circular cylindricalcoordinatesinSection2.4andtosphericalpolarcoordinatesinSection2.5. 2.1 O RTHOGONAL COORDINATES IN R3 In Cartesian coordinates we deal with three mutually perpendicular families of planes: x=constant, y=constant,and z=constant.Imaginethatwesuperimposeonthissystem 103 104 Chapter 2 Vector Analysis in Curved Coordinates and Tensors three other families of surfaces qi(x,y,z), i=1,2,3. The surfaces of any one family qi neednotbeparalleltoeachotherandtheyneednotbeplanes.Ifthisisdifficulttovisualize, the figure of a specific coordinate system, such as Fig. 2.3, may be helpful. The three new families of surfaces need not be mutually perpendicular, but for simplicity we impose this condition(Eq.(2.7))becauseorthogonalcoordinatesarecommoninphysicalapplications. Thisorthogonalityhasmanyadvantages:OrthogonalcoordinatesarealmostlikeCartesian coordinateswhereinfinitesimalareasandvolumesareproductsofcoordinatedifferentials. Inthissectionwedevelopthegeneralformalismoforthogonalcoordinates,derivefrom thegeometrythecoordinatedifferentials,andusethemforline,area,andvolumeelements inmultipleintegralsandvectoroperators.Wemaydescribeanypoint (x,y,z)astheinter- section of three planes in Cartesian coordinates or as the intersection of the three surfaces that form our new, curvilinear coordinates. Describing the curvilinear coordinate surfaces byq1=constant, q2=constant, q3=constant, we may identify our point by (q1,q2,q3) aswellasby (x,y,z): Generalcurvilinearcoordinates q1,q2,q3Circularcylindricalcoordinates ρ,ϕ,z x=x(q1,q2,q3) y=y(q1,q2,q3) z=z(q1,q2,q3)−∞<x=ρcosϕ<∞ −∞<y=ρsinϕ<∞ −∞<z=z<∞(2.1) specifying x,y,zintermsof q1,q2,q3andtheinverserelations q1=q1(x,y,z) 0/lessorequalslantρ=parenleftbig x2+y2parenrightbig1/2<∞ q2=q2(x,y,z) 0/lessorequalslantϕ=arctan(y/x)<2π q3=q3(x,y,z)−∞<z=z<∞.(2.2) As a specific illustration of the general, abstract q1,q2,q3, the transformation equations for circular cylindricalcoordinates (Section 2.4) are included in Eqs. (2.1) and (2.2). With each family of surfaces qi=constant, we can associate a unit vector ˆqinormal to the surfaceqi=constant and in the direction of increasing qi. In general, these unit vectors willdependonthepositioninspace.Thenavector Vmaybewritten V=ˆq1V1+ˆq2V2+ˆq3V3, (2.3) butthecoordinateorpositionvectoris differentingeneral, r/negationslash=ˆq1q1+ˆq2q2+ˆq3q3, as the special cases r=rˆrfor spherical polar coordinates and r=ρˆρ+zˆzfor cylindri- cal coordinates demonstrate. The ˆqiare normalized to ˆq2 i=1 and form a right-handed coordinatesystemwithvolume ˆq1·(ˆq2׈q3)>0. Differentiationof xinEqs. (2.1) leadstothetotalvariationordifferential dx=∂x ∂q1dq1+∂x ∂q2dq2+∂x ∂q3dq3, (2.4) and similarly for differentiation of yandz. In vector notation dr=summationtext i∂r ∂qidqi.From the Pythagorean theorem in Cartesian coordinates the square of the distance between two neighboringpointsis ds2=dx2+dy2+dz2. 2.1 Orthogonal Coordinates in R3105 Substituting drshows that in our curvilinear coordinate space the square of the distance elementcanbewrittenas aquadraticforminthedifferentials dqi: ds2=dr·dr=dr2=summationdisplay ij∂r ∂qi·∂r ∂qjdqidqj =g11dq2 1+g12dq1dq2+g13dq1dq3 +g21dq2dq1+g22dq2 2+g23dq2dq3 +g31dq3dq1+g32dq3dq2+g33dq2 3 =summationdisplay ijgijdqidqj, (2.5) where nonzero mixed terms dqidqjwithi/negationslash=jsignal that these coordinates are not or- thogonal, that is, that the tangential directions ˆqiare not mutually orthogonal. Spaces for whichEq.(2.5) is alegitimateexpressionarecalled metricorRiemannian . WritingEq. (2.5) moreexplicitly,weseethat gij(q1,q2,q3)=∂x ∂qi∂x ∂qj+∂y ∂qi∂y ∂qj+∂z ∂qi∂z ∂qj=∂r ∂qi·∂r ∂qj(2.6) arescalarproductsofthe tangentvectors∂r ∂qitothecurves rforqj=const.,j/negationslash=i.These coefficient functions gij, which we now proceed to investigate, may be viewed as speci- fying the nature of the coordinate system (q1,q2,q3). Collectively these coefficients are referred to as the metricand in Section 2.10 will be shown to form a second-rank sym- metric tensor.1In general relativity the metric components are determined by the proper- ties of matter; that is, the gijare solutions of Einstein’s field equations with the energy– momentum tensor as driving term; this may be articulated as “geometry is merged with physics.” At usual we limit ourselves to orthogonal (mutually perpendicular surfaces) coordinate systems,whichmeans(seeExercise2.1.1)2 gij=0,i/negationslash=j, (2.7) andˆqi·ˆqj=δij. (Nonorthogonal coordinate systems are considered in some detail in Sections2.10and2.11intheframeworkoftensoranalysis.)Now,tosimplifythenotation, wewrite gii=h2 i>0,so ds2=(h1dq1)2+(h2dq2)2+(h3dq3)2=summationdisplay i(hidqi)2. (2.8) 1The tensor nature of the set of gij’s follows from the quotient rule (Section 2.8). Then the tensor transformation law yields Eq.(2.5). 2In relativistic cosmology the nondiagonal elements of the metric gijare usually set equal to zero as a consequence of physical assumptions such asno rotation, as for dϕdt,dθdt . 106 Chapter 2 Vector Analysis in Curved Coordinates and Tensors The specific orthogonal coordinate systems are described in subsequent sections by spec- ifying these (positive) scale factors h1,h2, andh3. Conversely, the scale factors may be convenientlyidentifiedbytherelation dsi=hidqi,∂r ∂qi=hiˆqi (2.9) for any given dqi, holding all other qconstant. Here, dsiis a differential length along the directionˆqi.Notethatthethreecurvilinearcoordinates q1,q2,q3neednotbelengths.The scalefactors himaydependon qandtheymayhavedimensions.The producthidqimust haveadimensionof length.Thedifferentialdistancevector drmaybewritten dr=h1dq1ˆq1+h2dq2ˆq2+h3dq3ˆq3=summationdisplay ihidqiˆqi. Usingthiscurvilinearcomponentform,wefindthatalineintegralbecomes integraldisplay V·dr=summationdisplay iintegraldisplay Vihidqi. FromEqs. (2.9) wemayimmediatelydeveloptheareaandvolumeelements dσij=dsidsj=hihjdqidqj (2.10) and dτ=ds1ds2ds3=h1h2h3dq1dq2dq3. (2.11) The expressions in Eqs. (2.10) and (2.11) agree, of course, with the results of using the transformation equations, Eq. (2.1), and Jacobians (described shortly; see also Exer- cise2.1.5). FromEq. (2.10)anareaelementmaybeexpanded: dσ=ds2ds3ˆq1+ds3ds1ˆq2+ds1ds2ˆq3 =h2h3dq2dq3ˆq1+h3h1dq3dq1ˆq2 +h1h2dq1dq2ˆq3. Asurfaceintegralbecomes integraldisplay V·dσ=integraldisplay V1h2h3dq2dq3+integraldisplay V2h3h1dq3dq1 +integraldisplay V3h1h2dq1dq2. (Examplesof suchlineandsurfaceintegralsappearinSections2.4and2.5.) 2.1 Orthogonal Coordinates in R3107 In anticipation of the new forms of equations for vector calculus that appear in the next section, let us emphasize that vector algebrais the same in orthogonal curvilinear coordinatesasinCartesiancoordinates.Specifically,for thedotproduct, A·B=summationdisplay ikAiˆqi·ˆqkBk=summationdisplay ikAiBkδik =summationdisplay iAiBi=A1B1+A2B2+A3B3, (2.12) wherethesubscriptsindicatecurvilinearcomponents.Forthecross product, A×B=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆq1ˆq2ˆq3 A1A2A3 B1B2B3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (2.13) asinEq. (1.40). Previously, we specialized to locally rectangular coordinates that are adapted to special symmetries. Let us now briefly look at the more general case, where the coordinates are not necessarily orthogonal. Surface and volume elements are part of multiple integrals, which are common in physical applications, such as center of mass determinations and moments of inertia. Typically, we choose coordinates according to the symmetry of the particular problem. In Chapter 1 we used Gauss’ theorem to transform a volume integral into a surface integral and Stokes’ theorem to transform a surface integral into a line in- tegral. For orthogonal coordinates, the surface and volume elements are simply products of the line elements hidqi(see Eqs. (2.10) and (2.11)). For the general case, we use the geometric meaning of ∂r/∂qiin Eq. (2.5) as tangent vectors. We start with the Cartesian surface element dxdy, which becomes an infinitesimal rectangle in the new coordinates q1,q2formedbythetwoincrementalvectors dr1=r(q1+dq1,q2)−r(q1,q2)=∂r ∂q1dq1, dr2=r(q1,q2+dq2)−r(q1,q2)=∂r ∂q2dq2, (2.14) whoseareais the z-componentoftheircross product,or dxdy=dr1×dr2vextendsinglevextendsingle z=bracketleftbigg∂x ∂q1∂y ∂q2−∂x ∂q2∂y ∂q1bracketrightbigg dq1dq2 =vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂q1∂x ∂q2 ∂y ∂q1∂y ∂q2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingledq1dq2. (2.15) Thetransformationcoefficientindeterminantformis calledthe Jacobian . Similarly,thevolumeelement dxdydz becomesthetriplescalarproductofthethreein- finitesimaldisplacementvectors dri=dqi∂r ∂qialongthe qidirectionsˆqi,which,according 108 Chapter 2 Vector Analysis in Curved Coordinates and Tensors toSection1.5, takesontheform dxdydz=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂q1∂x ∂q2∂x ∂q3 ∂y ∂q1∂y ∂q2∂y ∂q3 ∂z ∂q1∂z ∂q2∂z ∂q3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingledq1dq2dq3. (2.16) HerethedeterminantisalsocalledtheJacobian,andsooninhigherdimensions. For orthogonal coordinates the Jacobians simplify to products of the orthogonal vec- tors in Eq. (2.9). It follows that they are just products of the hi; for example, the volume Jacobianbecomes h1h2h3(ˆq1׈q2)·ˆq3=h1h2h3, andso on. Example 2.1.1 JACOBIANS FOR POLAR COORDINATES LetusillustratethetransformationoftheCartesiantwo-dimensionalvolumeelement dxdy topolarcoordinates ρ,ϕ, withx=ρcosϕ, y=ρsinϕ. (SeealsoSection2.4.) Here, dxdy=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂ρ∂x ∂ϕ ∂y ∂ρ∂y ∂ϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsingledρdϕ=vextendsinglevextendsinglevextendsinglevextendsinglecosϕ−ρsinϕ sinϕρcosϕvextendsinglevextendsinglevextendsinglevextendsingledρdϕ=ρdρdϕ. Similarly, in spherical coordinates (see Section 2.5) we get, from x=rsinθcosϕ,y= rsinθsinϕ,z=rcosθ,theJacobian J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂r∂x ∂θ∂x ∂ϕ ∂y ∂r∂y ∂θ∂y ∂ϕ ∂z ∂r∂z ∂θ∂z ∂ϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglesinθcosϕrcosθcosϕ−rsinθsinϕ sinθsinϕrcosθsinϕrsinθcosϕ cosθ−rsinθ 0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle =cosθvextendsinglevextendsinglevextendsinglevextendsinglercosθcosϕ−rsinθsinϕ rcosθsinϕrsinθcosϕvextendsinglevextendsinglevextendsinglevextendsingle+rsinθvextendsinglevextendsinglevextendsinglevextendsinglesinθcosϕ−rsinθsinϕ sinθsinϕrsinθcosϕvextendsinglevextendsinglevextendsinglevextendsingle =r2parenleftbig cos2θsinθ+sin3θparenrightbig =r2sinθ by expanding the determinant along the third line. Hence the volume element becomes dxdydz=r2drsinθdθdϕ.Thevolumeintegralcanbewrittenas integraldisplay f(x,y,z)dxdydz =integraldisplay fparenleftbig x(r,θ,ϕ),y(r,θ,ϕ),z(r,θ,ϕ)parenrightbig r2drsinθdθdϕ./squaresolid Insummary,wehavedevelopedthegeneralformalismforvectoranalysisinorthogonal curvilinear coordinates in R3. For most applications, locally orthogonal coordinates can bechosenforwhichsurfaceandvolumeelementsinmultipleintegralsareproductsofline elements.For thegeneralnonorthogonalcase, Jacobiandeterminantsapply. 2.1 Orthogonal Coordinates in R3109 Exercises 2.1.1 Show that limiting our attention to orthogonal coordinate systems implies that gij=0 fori/negationslash=j(Eq. (2.7)). Hint.Constructatrianglewithsides ds1,ds2,andds2.Equation(2.9)mustholdregard- less of whether gij=0. Then compare ds2from Eq. (2.5) with a calculation using the lawof cosines.Showthat cos θ12=g12/√g11g22. 2.1.2 In the spherical polar coordinate system, q1=r,q2=θ,q3=ϕ. The transformation equationscorrespondingtoEq. (2.1) are x=rsinθcosϕ, y=rsinθsinϕ, z=rcosθ. (a) Calculatethesphericalpolarcoordinatescalefactors: hr,hθ,andhϕ. (b) Checkyourcalculatedscalefactors bytherelation dsi=hidqi. 2.1.3 Theu-,v-,z-coordinate system frequently used in electrostatics and in hydrodynamics isdefinedby xy=u, x2−y2=v, z=z. Thisu-,v-,z-systemis orthogonal. (a) In words, describe briefly the nature of each of the three families of coordinate surfaces. (b) Sketchthesysteminthe xy-planeshowingtheintersectionsofsurfacesofconstant uandsurfaces of constant vwiththexy-plane. (c) Indicatethedirectionsoftheunitvector ˆuandˆvinallfour quadrants. (d) Finally,isthis u-,v-,z-systemright-handed (ˆu׈v=+ˆz)orleft-handed (ˆu׈v= −ˆz)? 2.1.4 Theellipticcylindricalcoordinatesystemconsistsofthreefamiliesofsurfaces: 1)x2 a2cosh2u+y2 a2sinh2u=1;2)x2 a2cos2v−y2 a2sin2v=1;3)z=z. Sketch the coordinate surfaces u=constant and v=constant as they intersect the first quadrant of the xy-plane. Show the unit vectors ˆuandˆv. The range of uis 0/lessorequalslantu<∞. Therangeof vis 0/lessorequalslantv/lessorequalslant2π. 2.1.5 A two-dimensional orthogonal system is described by the coordinates q1andq2. Show thattheJacobian Jparenleftbiggx,y q1,q2parenrightbigg ≡∂(x,y) ∂(q1,q2)≡∂x ∂q1∂y ∂q2−∂x ∂q2∂y ∂q1=h1h2 isinagreementwithEq. (2.10). Hint.It’seasiertowork withthesquareofeachsideofthisequation. 110 Chapter 2 Vector Analysis in Curved Coordinates and Tensors 2.1.6 InMinkowskispacewedefine x1=x,x2=y,x3=z,andx0=ct.Thisisdonesothat the metric interval becomes ds2=dx2 0–dx2 1–dx2 2–dx2 3(withc=velocity of light). ShowthatthemetricinMinkowskispaceis (gij)= 1 000 0−10 0 00−10 00 0 −1 . WeuseMinkowskispaceinSections4.5and4.6fordescribingLorentztransformations. 2.2 D IFFERENTIAL VECTOR OPERATORS Wereturntoourrestrictiontoorthogonalcoordinatesystems. Gradient Thestartingpointfordevelopingthegradient,divergence,andcurloperatorsincurvilinear coordinates is the geometric interpretation of the gradient as the vector having the mag- nitude and direction of the maximum space rate of change (compare Section 1.6). From this interpretation the component of ∇ψ(q1,q2,q3)in the direction normal to the family ofsurfaces q1=constantis givenby3 ˆq1·∇ψ=∇ψ|1=∂ψ ∂s1=1 h1∂ψ ∂q1, (2.17) since this is the rate of change of ψfor varying q1, holding q2andq3fixed. The quantity ds1is a differential length in the direction of increasing q1(compare Eqs. (2.9)). In Sec- tion 2.1 we introduced a unit vector ˆq1to indicate this direction. By repeating Eq. (2.17) forq2andagainfor q3andaddingvectorially,wesee thatthegradientbecomes ∇ψ(q1,q2,q3)=ˆq1∂ψ ∂s1+ˆq2∂ψ ∂s2+ˆq3∂ψ ∂s3 =ˆq11 h1∂ψ ∂q1+ˆq21 h2∂ψ ∂q2+ˆq31 h3∂ψ ∂q3 =summationdisplay iˆqi1 hi∂ψ ∂qi. (2.18) Exercise2.2.4offersamathematicalalternativeindependentofthisphysicalinterpretation ofthegradient.Thetotalvariationofafunction, dψ=∇ψ·dr=summationdisplay i1 hi∂ψ ∂qidsi=summationdisplay i∂ψ ∂qidqi isconsistentwithEq. (2.18), ofcourse. 3Heretheuseof ϕtolabelafunctionisavoidedbecauseitisconventionaltousethissymboltodenoteanazimuthalcoordinate. 2.2 Differential Vector Operators 111 Divergence Thedivergenceoperatormaybeobtainedfromtheseconddefinition(Eq.(1.98))ofChap- ter1or equivalentlyfromGauss’theorem,Section1.11.Letususe Eq.(1.98), ∇·V(q1,q2,q3)=limintegraltext dτ→0integraltext V·dσintegraltext dτ, (2.19) with a differential volume h1h2h3dq1dq2dq3(Fig. 2.1). Note that the positive directions havebeenchosenso that (ˆq1,ˆq2,ˆq3)formaright-handedset, ˆq1׈q2=ˆq3. Thedifferenceofareaintegralsforthetwofaces q1=constant isgivenby bracketleftbigg V1h2h3+∂ ∂q1(V1h2h3)dq1bracketrightbigg dq2dq3−V1h2h3dq2dq3 =∂ ∂q1(V1h2h3)dq1dq2dq3, (2.20) exactly as in Sections 1.7 and 1.10.4Here,Vi=V·ˆqiis the projection of Vonto the ˆqi-direction.Addinginthesimilarresults for theothertwopairs ofsurfaces, weobtain integraldisplay V(q1,q2,q3)·dσ =bracketleftbigg∂ ∂q1(V1h2h3)+∂ ∂q2(V2h3h1)+∂ ∂q3(V3h1h2)bracketrightbigg dq1dq2dq3. FIGURE 2.1Curvilinearvolumeelement. 4Sin cewetak eth elim it dq1,dq2,dq3→0,the second- and higher-order derivatives will drop out. 112 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Now,usingEq. (2.19), divisionbyourdifferentialvolumeyields ∇·V(q1,q2,q3)=1 h1h2h3bracketleftbigg∂ ∂q1(V1h2h3)+∂ ∂q2(V2h3h1)+∂ ∂q3(V3h1h2)bracketrightbigg .(2.21) We may obtain the Laplacian by combining Eqs. (2.18) and (2.21), using V= ∇ψ(q1,q2,q3). Thisleadsto ∇·∇ψ(q1,q2,q3) =1 h1h2h3bracketleftbigg∂ ∂q1parenleftbiggh2h3 h1∂ψ ∂q1parenrightbigg +∂ ∂q2parenleftbiggh3h1 h2∂ψ ∂q2parenrightbigg +∂ ∂q3parenleftbiggh1h2 h3∂ψ ∂q3parenrightbiggbracketrightbigg .(2.22) Curl Finally, to develop ∇×V, let us apply Stokes’ theorem (Section 1.12) and, as with the divergence, take the limit as the surface area becomes vanishingly small. Working on one component at a time, we consider a differential surface element in the curvilinear surface q1=constant. From integraldisplay s∇×V·dσ=ˆq1·(∇×V)h2h3dq2dq3 (2.23) (meanvaluetheoremofintegralcalculus),Stokes’theoremyields ˆq1·(∇×V)h2h3dq2dq3=contintegraldisplay V·dr, (2.24) with the line integral lying in the surface q1=constant. Following the loop (1, 2, 3, 4) of Fig.2.2, contintegraldisplay V(q1,q2,q3)·dr=V2h2dq2+bracketleftbigg V3h3+∂ ∂q2(V3h3)dq2bracketrightbigg dq3 −bracketleftbigg V2h2+∂ ∂q3(V2h2)dq3bracketrightbigg dq2−V3h3dq3 =bracketleftbigg∂ ∂q2(h3V3)−∂ ∂q3(h2V2)bracketrightbigg dq2dq3. (2.25) We pick up a positive sign when going in the positive direction on parts 1 and 2 and a negative sign on parts 3 and 4 because here we are going in the negative direction. (Higher-order terms in Maclaurin or Taylor expansions have been omitted. They will van- ishinthelimitas thesurface becomesvanishinglysmall( dq2→0,dq3→0).) From Eq. (2.24), ∇×V|1=1 h2h3bracketleftbigg∂ ∂q2(h3V3)−∂ ∂q3(h2V2)bracketrightbigg . (2.26) 2.2 Differential Vector Operators 113 FIGURE 2.2Curvilinearsurface elementwith q1=constant. The remaining two components of ∇×Vmay be picked up by cyclic permutation of the indices.As inChapter1, itisoftenconvenienttowritethecurlindeterminantform: ∇×V=1 h1h2h3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆq1h1ˆq2h2ˆq3h3 ∂ ∂q1∂ ∂q2∂ ∂q3 h1V1h2V2h3V3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.27) Rememberthat,becauseofthepresenceofthedifferentialoperators,thisdeterminantmust be expanded from the top down. Note that this equation is notidentical with the form for the cross product of two vectors, Eq. (2.13). ∇is not an ordinary vector; it is a vector operator. OurgeometricinterpretationofthegradientandtheuseofGauss’andStokes’theorems (or integral definitions of divergence and curl) have enabled us to obtain these quantities without having to differentiate the unit vectors ˆqi. There exist alternate ways to deter- minegrad,div,andcurlbasedondirectdifferentiationofthe ˆqi.Oneapproachresolvesthe ˆqiofaspecificcoordinatesystemintoitsCartesiancomponents(Exercises2.4.1and2.5.1) anddifferentiatesthisCartesianform(Exercises2.4.3and2.5.2).Thepointhereisthatthe derivatives of the Cartesian ˆx,ˆy, andˆzvanish sinceˆx,ˆy, andˆzare constant in direction as well as in magnitude. A second approach [L. J. Kijewski, A m .J .P h y s . 33: 816 (1965)] assumestheequalityof ∂2r/∂qi∂qjand∂2r/∂qj∂qianddevelopsthederivativesof ˆqiin ageneralcurvilinearform. Exercises2.2.3and2.2.4arebasedonthismethod. Exercises 2.2.1 Developargumentstoshowthatdotandcrossproducts(notinvolving ∇)inorthogonal curvilinearcoordinatesin R3proceed,asinCartesiancoordinates, withnoinvolvement ofscalefactors . 2.2.2 Withˆq1aunitvectorinthedirectionofincreasing q1, showthat 114 Chapter 2 Vector Analysis in Curved Coordinates and Tensors (a)∇·ˆq1=1 h1h2h3∂(h2h3) ∂q1 (b)∇׈q1=1 h1bracketleftbigg ˆq21 h3∂h1 ∂q3−ˆq31 h2∂h1 ∂q2bracketrightbigg . Note that even though ˆq1is a unit vector, its divergence and curl do not necessarily vanish. 2.2.3 Showthattheorthogonalunitvectors ˆqjmaybedefinedby ˆqi=1 hi∂r ∂qi. (a) In particular, show that ˆqi·ˆqi=1 leads to an expression for hiin agreement with Eqs. (2.9). Equation(a) maybetakenasastartingpointfor deriving ∂ˆqi ∂qj=ˆqj1 hi∂hj ∂qi,i/negationslash=j and ∂ˆqi ∂qi=−summationdisplay j/negationslash=iˆqj1 hj∂hi ∂qj. 2.2.4 Derive ∇ψ=ˆq11 h1∂ψ ∂q1+ˆq21 h2∂ψ ∂q2+ˆq31 h3∂ψ ∂q3 bydirectapplicationof Eq.(1.97), ∇ψ=limintegraltext dτ→0integraltext ψdσintegraltext dτ. Hint.Evaluation of the surface integral will lead to terms like (h1h2h3)−1(∂/∂q1)× (ˆq1h2h3).TheresultslistedinExercise2.2.3willbehelpful.Cancellationofunwanted termsoccurswhenthecontributionsof allthreepairs ofsurfaces areaddedtogether. 2.3 S PECIAL COORDINATE SYSTEMS :INTRODUCTION There are at least 11 coordinate systems in which the three-dimensional Helmholtz equa- tion can be separated into three ordinary differential equations. Some of these coordinate systems have achieved prominence in the historical development of quantum mechanics. Othersystems,suchasbipolarcoordinates,satisfy specialneeds.Partlybecausetheneeds are rather infrequent but mostly because the development of computers and efficient pro- gramming techniques reduce the need for these coordinate systems, the discussion in this chapter is limited to (1) Cartesian coordinates, (2) spherical polar coordinates, and (3) cir- cular cylindrical coordinates. Specifications and details of the other coordinate systems willbefoundinthefirsttwoeditionsofthisworkandinAdditionalReadingsattheendof thischapter(Morse andFeshbach,MargenauandMurphy). 2.4 Circular Cylinder Coordinates 115 2.4 C IRCULAR CYLINDER COORDINATES In the circular cylindrical coordinate system the three curvilinear coordinates (q1,q2,q3) are relabeled (ρ,ϕ,z).W ea r eu s i n g ρfor the perpendicular distance from the z-axis and savingrfor thedistancefromtheorigin.Thelimitson ρ,ϕandzare 0/lessorequalslantρ<∞,0/lessorequalslantϕ/lessorequalslant2π,and−∞<z<∞. Forρ=0,ϕis notwelldefined.Thecoordinatesurfaces, showninFig.2.3,are: 1. Rightcircularcylindershavingthe z-axis asacommonaxis, ρ=parenleftbig x2+y2parenrightbig1/2=constant. 2. Half-planesthroughthe z-axis, ϕ=tan−1parenleftbiggy xparenrightbigg =constant. 3. Planesparalleltothe xy-plane,as intheCartesiansystem, z=constant. FIGURE 2.3Circularcylindercoordinates. 116 Chapter 2 Vector Analysis in Curved Coordinates and Tensors FIGURE 2.4Circularcylindrical coordinateunitvectors. Inverting the preceding equations for ρandϕ(or going directly to Fig. 2.3), we obtain thetransformationrelations x=ρcosϕ, y=ρsinϕ, z=z. (2.28) Thez-axis remains unchanged. This is essentially a two-dimensional curvilinear system withaCartesian z-axisaddedontoform athree-dimensionalsystem. AccordingtoEq.(2.5) or fromthelengthelements dsi, thescalefactors are h1=hρ=1,h 2=hϕ=ρ, h 3=hz=1. (2.29) Theunitvectors ˆq1,ˆq2,ˆq3arerelabeled (ˆρ,ˆϕ,ˆz),asinFig.2.4.Theunitvector ˆρisnormal to the cylindrical surface, pointing in the direction of increasing radius ρ. The unit vector ˆϕis tangential to the cylindrical surface, perpendicular to the half plane ϕ=constant and pointinginthedirectionofincreasingazimuthangle ϕ.Thethirdunitvector, ˆz,istheusual Cartesianunitvector.Theyaremutuallyorthogonal, ˆρ·ˆϕ=ˆϕ·ˆz=ˆz·ˆρ=0, andthecoordinatevectoranda generalvector Vareexpressedas r=ˆρρ+ˆzz,V=ˆρVρ+ˆϕVϕ+ˆzVz. Adifferentialdisplacement drmaybewritten dr=ˆρdsρ+ˆϕdsϕ+ˆzdz =ˆρdρ+ˆϕρdϕ+ˆzdz. (2.30) Example 2.4.1 AREALAW FOR PLANETARY MOTION FirstwederiveKepler’slawincylindricalcoordinates,sayingthattheradiusvectorsweeps outequalareasinequaltime,fromangularmomentumconservation. 2.4 Circular Cylinder Coordinates 117 Weconsiderthesunattheoriginasasourceofthe centralgravitationalforce F=f(r)ˆr. Then the orbital angular momentum L=mr×vof a planet of mass mand velocity vis conserved,becausethetorque dL dt=mdr dt×dr dt+r×mdv dt=r×F=f(r) rr×r=0. HenceL=const. Now we can choose the z-axis to lie along the direction of the orbital angularmomentumvector, L=Lˆz,andworkincylindricalcoordinates r=(ρ,ϕ,z)=ρˆρ withz=0.Theplanetmovesinthe xy-planebecause randvareperpendicularto L.Thus, weexpanditsvelocityasfollows: v=dr dt=˙ρˆρ+ρdˆρ dt. From ˆρ=(cosϕ,sinϕ),∂ˆρ dϕ=(−sinϕ,cosϕ)=ˆϕ, wefindthatdˆρ dt=dˆρ dϕdϕ dt=˙ϕˆϕusingthechainrule,so v=˙ρˆρ+ρdˆρ dt=˙ρˆρ+ρ˙ϕˆϕ.When wesubstitutetheexpansionsof ˆρandvinpolarcoordinates,weobtain L=mρ×v=mρ(ρ˙ϕ)(ˆρ׈ϕ)=mρ2˙ϕˆz=constant. The triangular area swept by the radius vector ρin the time dt(area law), when inte- gratedoveronerevolution,is givenby A=1 2integraldisplay ρ(ρdϕ)=1 2integraldisplay ρ2˙ϕdt=L 2mintegraldisplay dt=Lτ 2m, (2.31) ifwesubstitute mρ2˙ϕ=L=const.Here τistheperiod,thatis,thetimeforonerevolution oftheplanetinitsorbit. Kepler’s first law says that the orbit is an ellipse. Now we derive the orbit equation ρ(ϕ)of the ellipse in polar coordinates, where in Fig. 2.5 the sun is at one focus, which is the origin of our cylindrical coordinates. From the geometrical construction of the ellipse we know that ρ′+ρ=2a,whereais the major half-axis; we shall show that this is equivalenttotheconventionalformoftheellipseequation.Thedistancebetweenbothfoci is0<2aǫ<2a,where0 <ǫ<1iscalledtheeccentricityoftheellipse.Foracircle ǫ=0 because both foci coincide with the center. There is an angle, as shown in Fig. 2.5, where the distances ρ′=ρ=aare equal, and Pythagoras’ theorem applied to this right triangle FIGURE 2.5Ellipseinpolarcoordinates. 118 Chapter 2 Vector Analysis in Curved Coordinates and Tensors givesb2+a2ǫ2=a2.As a result,√ 1−ǫ2=b/ais the ratio of the minor half-axis ( b)t o themajorhalf-axis, a. Now consider the triangle with the sides labeled by ρ′,ρ,2aǫin Fig. 2.5 and angle oppositeρ′equaltoπ−ϕ.Then,applyingthelawofcosines, gives ρ′2=ρ2+4a2ǫ2+4ρaǫcosϕ. Nowsubstituting ρ′=2a−ρ,canceling ρ2onbothsides anddividingby 4 ayields ρ(1+ǫcosϕ)=aparenleftbig 1−ǫ2parenrightbig ≡p, (2.32) theKeplerorbitequationinpolarcoordinates . Alternatively, we revert to Cartesian coordinates to find, from Eq. (2.32) with x= ρcosϕ, that ρ2=x2+y2=(p−xǫ)2=p2+x2ǫ2−2pxǫ, sothefamiliarellipseequationinCartesiancoordinates, parenleftbig 1−ǫ2parenrightbigparenleftbigg x+pǫ 1−ǫ2parenrightbigg2 +y2=p2+p2ǫ2 1−ǫ2=p2 1−ǫ2, obtains.If wecomparethisresultwiththestandardformof theellipse, (x−x0)2 a2+y2 b2=1, weconfirmthat b=p√ 1−ǫ2=aradicalbig 1−ǫ2,a=p 1−ǫ2, andthatthedistance x0betweenthecenterandfocusis aǫ,assho wninFig.2. 5. /squaresolid The differential operations involving ∇follow from Eqs. (2.18), (2.21), (2.22), and (2.27): ∇ψ(ρ,ϕ,z)=ˆρ∂ψ ∂ρ+ˆϕ1 ρ∂ψ ∂ϕ+ˆz∂ψ ∂z, (2.33) ∇·V=1 ρ∂ ∂ρ(ρVρ)+1 ρ∂Vϕ ∂ϕ+∂Vz ∂z,(2.34) ∇2ψ=1 ρ∂ ∂ρparenleftbigg ρ∂ψ ∂ρparenrightbigg +1 ρ2∂2ψ ∂ϕ2+∂2ψ ∂z2,(2.35) ∇×V=1 ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz ∂ ∂ρ∂ ∂ϕ∂ ∂z VρρVϕVzvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.36) 2.4 Circular Cylinder Coordinates 119 Finally, for problems such as circular wave guides and cylindrical cavity resonators the vectorLaplacian ∇2Vresolvedincircularcylindricalcoordinatesis ∇2V|ρ=∇2Vρ−1 ρ2Vρ−2 ρ2∂Vϕ ∂ϕ, ∇2V|ϕ=∇2Vϕ−1 ρ2Vϕ+2 ρ2∂Vρ ∂ϕ, (2.37) ∇2V|z=∇2Vz, whichfollowfromEq.(1.85).Thebasicreasonforthisparticularformofthe z-component isthatthe z-axis isaCartesianaxis;thatis, ∇2(ˆρVρ+ˆϕVϕ+ˆzVz)=∇2(ˆρVρ+ˆϕVϕ)+ˆz∇2Vz =ˆρf(Vρ,Vϕ)+ˆϕg(Vρ,Vϕ)+ˆz∇2Vz. Finally,theoperator ∇2operatingonthe ˆρ,ˆϕunitvectorsstaysinthe ˆρˆϕ-plane. Example 2.4.2 AN AVIER –STOKES TERM TheNavier–Stokesequationsof hydrodynamicscontainanonlinearterm ∇×bracketleftbig v×(∇×v)bracketrightbig , wherevisthefluidvelocity.Forfluidflowingthroughacylindricalpipeinthe z-direction, v=ˆzv(ρ). From Eq. (2.36), ∇×v=1 ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz ∂ ∂ρ∂ ∂ϕ∂ ∂z 00 v(ρ)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−ˆϕ∂v ∂ρ v×(∇×v)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρ ˆϕˆz 00 v 0−∂v ∂ρ0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=ˆρv(ρ)∂v ∂ρ. Finally, ∇×parenleftbig v×(∇×v)parenrightbig =1 ρvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆρρˆϕˆz ∂ ∂ρ∂ ∂ϕ∂ ∂z v∂v ∂ρ00vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0, so,for thisparticularcase,thenonlineartermvanishes. /squaresolid 120 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Exercises 2.4.1 ResolvethecircularcylindricalunitvectorsintotheirCartesiancomponents(Fig.2.6). ANS. ˆρ=ˆxcosϕ+ˆysinϕ, ˆϕ=−ˆxsinϕ+ˆycosϕ, ˆz=ˆz. 2.4.2 ResolvetheCartesianunitvectorsintotheircircularcylindricalcomponents(Fig.2.6). ANS.ˆx=ˆρcosϕ−ˆϕsinϕ, ˆy=ˆρsinϕ+ˆϕcosϕ, ˆz=ˆz. 2.4.3 Fromtheresults ofExercise2.4.1showthat ∂ˆρ ∂ϕ=ˆϕ,∂ˆϕ ∂ϕ=−ˆρ and that all other first derivatives of the circular cylindrical unit vectors with respect to thecircularcylindricalcoordinatesvanish. 2.4.4 Compare ∇·V(Eq. (2.34)) withthegradientoperator ∇=ˆρ∂ ∂ρ+ˆϕ1 ρ∂ ∂ϕ+ˆz∂ ∂z (Eq. (2.33)) dotted into V. Note that the differential operators of ∇differentiate both theunitvectorsandthecomponentsof V. Hint. ˆϕ(1/ρ)(∂/∂ϕ)·ˆρVρbecomes ˆϕ·1 ρ∂ ∂ϕ(ˆρVρ)anddoes notvanish. 2.4.5 (a) Showthat r=ˆρρ+ˆzz. FIGURE 2.6Planepolarcoordinates. 2.4 Circular Cylinder Coordinates 121 (b) Workingentirelyincircularcylindricalcoordinates,showthat ∇·r=3 and ∇×r=0. 2.4.6 (a) Show that the parity operation (reflection through the origin) on a point (ρ,ϕ,z) relativeto fixedx-,y-,z-axesconsistsof thetransformation ρ→ρ, ϕ→ϕ±π, z→−z. (b) Show that ˆρandˆϕhave odd parity (reversal of direction) and that ˆzhas even parity. Note.TheCartesianunitvectors ˆx,ˆy, andˆzremainconstant. 2.4.7 Arigidbodyisrotatingaboutafixedaxiswithaconstantangularvelocity ω.T ake ωto liealongthe z-axis.Expressthepositionvector rincircularcylindricalcoordinatesand usingcircularcylindricalcoordinates, (a) calculate v=ω×r, (b) calculate ∇×v. ANS.(a)v=ˆϕωρ, (b)∇×v=2ω. 2.4.8 Find the circular cylindrical components of the velocity and acceleration of a movingparticle, vρ=˙ρ, a ρ=¨ρ−ρ˙ϕ2, vϕ=ρ˙ϕ, a ϕ=ρ¨ϕ+2˙ρ˙ϕ, vz=˙z, a z=¨z. Hint. r(t)=ˆρ(t)ρ(t)+ˆzz(t) =bracketleftbigˆxcosϕ(t)+ˆysinϕ(t)bracketrightbig ρ(t)+ˆzz(t). Note.˙ρ=dρ/dt,¨ρ=d2ρ/dt2, andso on. 2.4.9 SolveLaplace’sequation, ∇2ψ=0,incylindricalcoordinatesfor ψ=ψ(ρ). ANS.ψ=klnρ ρ0. 2.4.10 Inrightcircularcylindricalcoordinatesaparticularvectorfunctionisgivenby V(ρ,ϕ)=ˆρVρ(ρ,ϕ)+ˆϕVϕ(ρ,ϕ). Showthat ∇×Vhasonlya z-component.Notethatthisresultwillholdfor anyvector confined to a surface q3=constant as long as the products h1V1andh2V2are each independentof q3. 2.4.11 FortheflowofanincompressibleviscousfluidtheNavier–Stokesequationsleadto −∇×parenleftbig v×(∇×v)parenrightbig =η ρ0∇2(∇×v). Hereηis the viscosity and ρ0is the density of the fluid. For axial flow in a cylindrical pipewetakethevelocity vtobe v=ˆzv(ρ). 122 Chapter 2 Vector Analysis in Curved Coordinates and Tensors FromExample2.4.2, ∇×parenleftbig v×(∇×v)parenrightbig =0 for thischoiceof v. Showthat ∇2(∇×v)=0 leadstothedifferentialequation 1 ρd dρparenleftbigg ρd2v dρ2parenrightbigg −1 ρ2dv dρ=0 andthatthisis satisfiedby v=v0+a2ρ2. 2.4.12 A conducting wire along the z-axis carries a current I. The resulting magnetic vector potentialis givenby A=ˆzµI 2πlnparenleftbigg1 ρparenrightbigg . Showthatthemagneticinduction Bisgivenby B=ˆϕµI 2πρ. 2.4.13 Aforceis describedby F=−ˆxy x2+y2+ˆyx x2+y2. (a) Express Fincircularcylindricalcoordinates. Operatingentirelyincircularcylindricalcoordinatesfor(b) and(c), (b) calculatethecurlof Fand (c) calculatetheworkdoneby Fintraverstheunitcircleoncecounterclockwise. (d) Howdoyoureconciletheresults of(b) and(c)? 2.4.14 A transverse electromagnetic wave (TEM) in a coaxial waveguide has an electric field E=E(ρ,ϕ)ei(kz−ωt)andamagneticinductionfieldof B=B(ρ,ϕ)ei(kz−ωt).Sincethe waveistransverse,neither EnorBhasazcomponent.Thetwofieldssatisfythe vector Laplacianequation ∇2E(ρ,ϕ)=0 ∇2B(ρ,ϕ)=0. (a) Showthat E=ˆρE0(a/ρ)ei(kz−ωt)andB=ˆϕB0(a/ρ)ei(kz−ωt)aresolutions.Here ais theradiusoftheinnerconductorand E0andB0areconstantamplitudes. 2.5 Spherical Polar Coordinates 123 (b) Assuming a vacuum inside the waveguide, verify that Maxwell’s equations are satisfiedwith B0/E0=k/ω=µ0ε0(ω/k)=1/c. 2.4.15 A calculation of the magnetohydrodynamic pinch effect involves the evaluation of (B·∇)B.If themagneticinduction Bistakentobe B=ˆϕBϕ(ρ), showthat (B·∇)B=−ˆρB2 ϕ/ρ. 2.4.16 The linear velocityof particles in a rigid body rotatingwith angular velocity ωis given by v=ˆϕρω. Integratecontintegraltext v·dλaroundacircleinthe xy-planeandverifythat contintegraltext v·dλ area=∇×v|z. 2.4.17 Ap r o t o no fm a s s m, charge+e, and (asymptotic) momentum p=mvis incident on a nucleus of charge +Zeat an impact parameter b. Determine the proton’s distance of closestapproach. 2.5 S PHERICAL POLAR COORDINATES Relabeling (q1,q2,q3)as(r,θ,ϕ), we see that the spherical polar coordinate system con- sists ofthefollowing: 1. Concentricspheres centeredattheorigin, r=parenleftbig x2+y2+z2parenrightbig1/2=constant. 2. Rightcircularconescenteredonthe z-(polar) axis, verticesattheorigin, θ=arccosz (x2+y2+z2)1/2=constant. 3. Half-planesthroughthe z-(polar) axis, ϕ=arctany x=constant. By our arbitrary choice of definitions of θ, the polar angle, and ϕ, the azimuth angle, the z-axis is singled out for special treatment. The transformation equations corresponding to Eq.(2.1) are x=rsinθcosϕ, y=rsinθsinϕ, z=rcosθ, (2.38) 124 Chapter 2 Vector Analysis in Curved Coordinates and Tensors FIGURE 2.7Sphericalpolarcoordinatearea elements. measuring θfrom the positive z-axis and ϕin thexy-plane from the positive x-axis. The ranges of values are 0 /lessorequalslantr<∞,0/lessorequalslantθ/lessorequalslantπ, and 0 /lessorequalslantϕ/lessorequalslant2π.A tr=0,θandϕare undefined.FromdifferentiationofEq. (2.38), h1=hr=1, h2=hθ=r, (2.39) h3=hϕ=rsinθ. Thisgivesalineelement dr=ˆrdr+ˆθrdθ+ˆϕrsinθdϕ, so ds2=dr·dr=dr2+r2dθ2+r2sin2θdϕ2, the coordinates being obviously orthogonal. In this spherical coordinate system the area element(for r=constant)is dA=dσθϕ=r2sinθdθdϕ, (2.40) the light, unshaded area in Fig. 2.7. Integrating over the azimuth ϕ, we find that the area elementbecomesaringof width dθ, dAθ=2πr2sinθdθ. (2.41) Thisformwillappearrepeatedlyinproblemsinsphericalpolarcoordinateswithazimuthal symmetry,suchasthescatteringofanunpolarizedbeamofparticles.Bydefinitionofsolid radians,orsteradians,anelementofsolidangle d/Omega1isgivenby d/Omega1=dA r2=sinθdθdϕ. (2.42) 2.5 Spherical Polar Coordinates 125 FIGURE 2.8Sphericalpolarcoordinates. Integratingovertheentiresphericalsurface, weobtain integraldisplay d/Omega1=4π. FromEq. (2.11)thevolumeelementis dτ=r2drsinθdθdϕ=r2drd/Omega1. (2.43) ThesphericalpolarcoordinateunitvectorsareshowninFig.2.8. Itmustbeemphasizedthat theunitvectors ˆr,ˆθ,and ˆϕvaryindirectionastheangles θandϕvary.Specifically,the θandϕderivativesofthesesphericalpolarcoordinateunit vectors do not vanish (Exercise 2.5.2). When differentiating vectors in spherical polar (or in any non-Cartesian system), this variation of the unit vectors with position must not be neglected.Intermsofthefixed-directionCartesianunitvectors ˆx,ˆyandˆz(cp.Eq.(2.38)), ˆr=ˆxsinθcosϕ+ˆysinθsinϕ+ˆzcosθ, ˆθ=ˆxcosθcosϕ+ˆycosθsinϕ−ˆzsinθ=∂ˆr ∂θ, (2.44) ˆϕ=−ˆxsinϕ+ˆycosϕ=1 sinθ∂ˆr ∂ϕ, whichfollowfrom 0=∂ˆr2 ∂θ=2ˆr·∂ˆr ∂θ,0=∂ˆr2 ∂ϕ=2ˆr·∂ˆr ∂ϕ. 126 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Note that Exercise 2.5.5 gives the inverse transformation and that a given vector can nowbeexpressedinanumberofdifferent(butequivalent)ways.Forinstance,theposition vectorrmaybewritten r=ˆrr=ˆrparenleftbig x2+y2+z2parenrightbig1/2 =ˆxx+ˆyy+ˆzz =ˆxrsinθcosϕ+ˆyrsinθsinϕ+ˆzrcosθ. (2.45) Selecttheform thatis mostusefulfor yourparticularproblem. From Section 2.2, relabeling the curvilinear coordinate unit vectors ˆq1,ˆq2, andˆq3asˆr, ˆθ, andˆϕgives ∇ψ=ˆr∂ψ ∂r+ˆθ1 r∂ψ ∂θ+ˆϕ1 rsinθ∂ψ ∂ϕ, (2.46) ∇·V=1 r2sinθbracketleftbigg sinθ∂ ∂r(r2Vr)+r∂ ∂θ(sinθVθ)+r∂Vϕ ∂ϕbracketrightbigg , (2.47) ∇·∇ψ=1 r2sinθbracketleftbigg sinθ∂ ∂rparenleftbigg r2∂ψ ∂rparenrightbigg +∂ ∂θparenleftbigg sinθ∂ψ ∂θparenrightbigg +1 sinθ∂2ψ ∂ϕ2bracketrightbigg ,(2.48) ∇×V=1 r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆrrˆθrsinθˆϕ ∂ ∂r∂ ∂θ∂ ∂ϕ VrrVθrsinθVϕvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (2.49) Occasionally, the vector Laplacian ∇2Vis needed in spherical polar coordinates. It is bestobtainedbyusingthevectoridentity(Eq. (1.85))of Chapter1. Forreference ∇2V|r=parenleftbigg −2 r2+2 r∂ ∂r+∂2 ∂r2+cosθ r2sinθ∂ ∂θ+1 r2∂2 ∂θ2+1 r2sin2θ∂2 ∂ϕ2parenrightbigg Vr +parenleftbigg −2 r2∂ ∂θ−2cosθ r2sinθparenrightbigg Vθ+parenleftbigg −2 r2sinθ∂ ∂ϕparenrightbigg Vϕ =∇2Vr−2 r2Vr−2 r2∂Vθ ∂θ−2cosθ r2sinθVθ−2 r2sinθ∂Vϕ ∂ϕ, (2.50) ∇2V|θ=∇2Vθ−1 r2sin2θVθ+2 r2∂Vr ∂θ−2cosθ r2sin2θ∂Vϕ ∂ϕ, (2.51) ∇2V|ϕ=∇2Vϕ−1 r2sin2θVϕ+2 r2sinθ∂Vr ∂ϕ+2cosθ r2sin2θ∂Vθ ∂ϕ. (2.52) These expressions for the components of ∇2Vare undeniably messy, but sometimes they areneeded. 2.5 Spherical Polar Coordinates 127 Example 2.5.1 ∇,∇·,∇×FOR A CENTRAL FORCE Using Eqs. (2.46) to (2.49), we can reproduce by inspection some of the results derived in Chapter1bylaboriousapplicationofCartesiancoordinates. From Eq. (2.46), ∇f(r)=ˆrdf dr, ∇rn=ˆrnrn−1.(2.53) FortheCoulombpotential V=Ze/(4πε0r), theelectricfieldis E=−∇V=Ze 4πε0r2ˆr. From Eq. (2.47), ∇·ˆrf(r)=2 rf(r)+df dr, ∇·ˆrrn=(n+2)rn−1.(2.54) Forr>0 the charge density of the electric field of the Coulomb potential is ρ=∇·E= Ze 4πε0∇·ˆr r2=0 because n=−2. From Eq. (2.48), ∇2f(r)=2 rdf dr+d2f dr2, (2.55) ∇2rn=n(n+1)rn−2, (2.56) incontrasttotheordinaryradialsecondderivativeof rninvolving n−1 insteadof n+1. Finally,from Eq.(2.49), ∇׈rf(r)=0. (2.57) /squaresolid Example 2.5.2 MAGNETIC VECTOR POTENTIAL The computation of the magnetic vector potential of a single current loop in the xy-plane usesOersted’slaw, ∇×H=J,inconjunctionwith µ0H=B=∇×A(seeExamples1.9.2 and1.12.1), andinvolvestheevaluationof µ0J=∇×bracketleftbig ∇׈ϕAϕ(r,θ)bracketrightbig . Insphericalpolarcoordinatesthisreducesto µ0J=∇×1 r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆrrˆθrsinθˆϕ ∂ ∂r∂ ∂θ∂ ∂ϕ 00 rsinθAϕ(r,θ)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle =∇×1 r2sinθbracketleftbigg ˆr∂ ∂θ(rsinθAϕ)−rˆθ∂ ∂r(rsinθAϕ)bracketrightbigg . 128 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Takingthecurlasecondtime,weobtain µ0J=1 r2sinθvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleˆr rˆθ rsinθˆϕ ∂ ∂r∂ ∂θ∂ ∂ϕ 1 r2sinθ∂ ∂θ(rsinθAϕ)−1 rsinθ∂ ∂r(rsinθAϕ)0vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. Byexpandingthedeterminantalongthetoprow, wehave µ0J=−ˆϕbraceleftbigg1 r∂2 ∂r2(rAϕ)+1 r2∂ ∂θbracketleftbigg1 sinθ∂ ∂θ(sinθAϕ)bracketrightbiggbracerightbigg =−ˆϕbracketleftbigg ∇2Aϕ(r,θ)−1 r2sin2θAϕ(r,θ)bracketrightbigg . (2.58) /squaresolid Exercises 2.5.1 Express thesphericalpolarunitvectorsinCartesianunitvectors. ANS.ˆr=ˆxsinθcosϕ+ˆysinθsinϕ+ˆzcosθ, ˆθ=ˆxcosθcosϕ+ˆycosθsinϕ−ˆzsinθ, ˆϕ=−ˆxsinϕ+ˆycosϕ. 2.5.2 (a) From the results of Exercise 2.5.1, calculate the partial derivatives of ˆr,ˆθ, and ˆϕ withrespectto r,θ,andϕ. (b) With ∇givenby ˆr∂ ∂r+ˆθ1 r∂ ∂θ+ˆϕ1 rsinθ∂ ∂ϕ (greatestspacerateofchange),usetheresultsofpart(a)tocalculate ∇·∇ψ.This isanalternatederivationof theLaplacian. Note.Thederivativesoftheleft-hand ∇operateontheunitvectorsoftheright-hand ∇ beforetheunitvectorsaredottedtogether. 2.5.3 Arigidbodyisrotatingaboutafixedaxiswithaconstantangularvelocity ω.T ake ωto bealongthe z-axis. Usingsphericalpolarcoordinates, (a) Calculate v=ω×r. (b) Calculate ∇×v. ANS.(a)v=ˆϕωrsinθ, (b)∇×v=2ω. 2.5 Spherical Polar Coordinates 129 2.5.4 Thecoordinatesystem (x,y,z)isrotatedthroughanangle /Phi1counterclockwiseaboutan axisdefinedbytheunitvector nintosystem (x′,y′,z′).Intermsofthenewcoordinates theradiusvectorbecomes r′=rcos/Phi1+r×nsin/Phi1+n(n·r)(1−cos/Phi1). (a) Derivethisexpressionfrom geometricconsiderations. (b) Showthatitreducesasexpectedfor n=ˆz.Theanswer,inmatrixform,appearsin Eq.(3.90). (c) Verifythat r′2=r2. 2.5.5 ResolvetheCartesianunitvectorsintotheirsphericalpolarcomponents: ˆx=ˆrsinθcosϕ+ˆθcosθcosϕ−ˆϕsinϕ, ˆy=ˆrsinθsinϕ+ˆθcosθsinϕ+ˆϕcosϕ, ˆz=ˆrcosθ−ˆθsinθ. 2.5.6 The direction of one vector is given by the angles θ1andϕ1. For a second vector the corresponding angles are θ2andϕ2. Show that the cosine of the included angle γis givenby cosγ=cosθ1cosθ2+sinθ1sinθ2cos(ϕ1−ϕ2). SeeFig. 12.15. 2.5.7 A certain vector Vhas no radial component. Its curl has no tangential components. Whatdoesthisimplyabouttheradialdependenceofthetangentialcomponentsof V? 2.5.8 Modernphysicslaysgreatstressonthepropertyofparity—whetheraquantityremains invariant or changes sign under an inversion of the coordinate system. In Cartesian coordinatesthismeans x→−x,y→−y,andz→−z. (a) Show that the inversion (reflection through the origin) of a point (r,θ,ϕ)relative tofixedx-,y-,z-axesconsistsof thetransformation r→r, θ→π−θ, ϕ→ϕ±π. (b) Showthat ˆrandˆϕhaveoddparity(reversalofdirection)andthat ˆθhasevenparity. 2.5.9 WithAanyvector, A·∇r=A. (a) Verifythisresult inCartesiancoordinates. (b) Verifythisresult usingsphericalpolarcoordinates.(Equation(2.46) provides ∇.) 130 Chapter 2 Vector Analysis in Curved Coordinates and Tensors 2.5.10 Find the spherical coordinate components of the velocity and acceleration of a moving particle: vr=˙r, vθ=r˙θ, vϕ=rsinθ˙ϕ, ar=¨r−r˙θ2−rsin2θ˙ϕ2, aθ=r¨θ+2˙r˙θ−rsinθcosθ˙ϕ2, aϕ=rsinθ¨ϕ+2˙rsinθ˙ϕ+2rcosθ˙θ˙ϕ. Hint. r(t)=ˆr(t)r(t) =bracketleftbigˆxsinθ(t)cosϕ(t)+ˆysinθ(t)sinϕ(t)+ˆzcosθ(t)bracketrightbig r(t). Note.Using the Lagrangian techniques of Section 17.3, we may obtain these results somewhat more elegantly. The dot in ˙r,˙θ,˙ϕmeans time derivative, ˙r=dr/dt,˙θ= dθ/dt,˙ϕ=dϕ/dt.Thenotationwas originatedbyNewton. 2.5.11 Aparticle mmovesinresponsetoacentralforceaccordingtoNewton’ssecondlaw, m¨r=ˆrf(r). Show that r×˙r=c, a constant, and that the geometric interpretation of this leads to Kepler’ssecondlaw. 2.5.12 Express∂/∂x,∂/∂y,∂/∂zinsphericalpolarcoordinates. ANS.∂ ∂x=sinθcosϕ∂ ∂r+cosθcosϕ1 r∂ ∂θ−sinϕ rsinθ∂ ∂ϕ, ∂ ∂y=sinθsinϕ∂ ∂r+cosθsinϕ1 r∂ ∂θ+cosϕ rsinθ∂ ∂ϕ, ∂ ∂z=cosθ∂ ∂r−sinθ1 r∂ ∂θ. Hint.Equate ∇xyzand∇rθϕ. 2.5.13 From Exercise2.5.12showthat −iparenleftbigg x∂ ∂y−y∂ ∂xparenrightbigg =−i∂ ∂ϕ. This is the quantum mechanical operator corresponding to the z-component of orbital angularmomentum. 2.5.14 With the quantum mechanical orbital angular momentum operator defined as L= −i(r×∇),showthat (a)Lx+iLy=eiϕparenleftbigg∂ ∂θ+icotθ∂ ∂ϕparenrightbigg , 2.5 Spherical Polar Coordinates 131 (b)Lx−iLy=−e−iϕparenleftbigg∂ ∂θ−icotθ∂ ∂ϕparenrightbigg . (ThesearetheraisingandloweringoperatorsofSection4.3.) 2.5.15 Verify that L×L=iLin spherical polar coordinates. L=−i(r×∇), the quantum mechanicalorbitalangularmomentumoperator. Hint.Use spherical polar coordinates for Lbut Cartesian components for the cross product. 2.5.16 (a) FromEq. (2.46) showthat L=−i(r×∇)=iparenleftbigg ˆθ1 sinθ∂ ∂ϕ−ˆϕ∂ ∂θparenrightbigg . (b) Resolving ˆθandˆϕintoCartesiancomponents,determine Lx,Ly,andLzinterms ofθ,ϕ, andtheirderivatives. (c) From L2=L2 x+L2y+L2zshowthat L2=−1 sinθ∂ ∂θparenleftbigg sinθ∂ ∂θparenrightbigg −1 sin2θ∂2 ∂ϕ2 =−r2∇2+∂ ∂rparenleftbigg r2∂ ∂rparenrightbigg . This latter identity is useful in relating orbital angular momentum and Legendre’s dif- ferentialequation,Exercise9.3.8. 2.5.17 WithL=−ir×∇, verifytheoperatoridentities (a)∇=ˆr∂ ∂r−ir×L r2, (b)r∇2−∇parenleftbigg 1+r∂ ∂rparenrightbigg =i∇×L. 2.5.18 Showthatthefollowingthreeforms(sphericalcoordinates)of ∇2ψ(r)areequivalent: (a)1 r2d drbracketleftbigg r2dψ(r) drbracketrightbigg ;(b)1 rd2 dr2bracketleftbig rψ(r)bracketrightbig ;(c)d2ψ(r) dr2+2 rdψ(r) dr. The second form is particularly convenient in establishing a correspondence between sphericalpolarandCartesiandescriptionsofaproblem. 2.5.19 Onemodelof thesolarcoronaassumesthatthesteady-stateequationof heatflow, ∇·(k∇T)=0, is satisfied. Here, k, the thermal conductivity, is proportional to T5/2. Assuming that the temperature Tis proportional to rn, show that the heat flow equation is satisfied by T=T0(r0/r)2/7. 132 Chapter 2 Vector Analysis in Curved Coordinates and Tensors 2.5.20 Acertainforcefieldisgivenby F=ˆr2Pcosθ r3+ˆθP r3sinθ, r /greaterorequalslantP/2 (insphericalpolarcoordinates). (a) Examine ∇×Ftosee ifapotentialexists. (b) Calculatecontintegraltext F·dλfor a unit circle in the plane θ=π/2. What does this indicate abouttheforcebeingconservativeornonconservative? (c) If you believe that Fmay be described by F=−∇ψ, findψ. Otherwise simply statethatnoacceptablepotentialexists. 2.5.21 (a) Showthat A=−ˆϕcotθ/risasolutionof ∇×A=ˆr/r2. (b) Show that this spherical polar coordinate solution agrees with the solution given forExercise1.13.6: A=ˆxyz r(x2+y2)−ˆyxz r(x2+y2). Notethatthesolutiondivergesfor θ=0,πcorrespondingto x,y=0. (c) Finally, show that A=−ˆθϕsinθ/ris a solution. Note that although this solution does not diverge (r/negationslash=0), it is no longer single-valued for all possible azimuth angles. 2.5.22 Amagneticvectorpotentialis givenby A=µ0 4πm×r r3. Showthatthisleadstothemagneticinduction Bofapointmagneticdipolewithdipole momentm. ANS.form=ˆzm, ∇×A=ˆrµ0 4π2mcosθ r3+ˆθµ0 4πmsinθ r3. CompareEqs. (12.133)and(12.134) 2.5.23 Atlargedistancesfromits source,electricdipoleradiationhasfields E=aEsinθei(kr−ωt) rˆθ,B=aBsinθei(kr−ωt) rˆϕ. ShowthatMaxwell’sequations ∇×E=−∂B ∂tand ∇×B=ε0µ0∂E ∂t aresatisfied,if wetake aE aB=ω k=c=(ε0µ0)−1/2. Hint.Sincerislarge, termsoforder r−2maybedropped. 2.6 Tensor Analysis 133 2.5.24 Themagneticvectorpotentialfor auniformlychargedrotatingsphericalshellis A=  ˆϕµ0a4σω 3·sinθ r2,r>a ˆϕµ0aσω 3·rcosθ, r<a. (a=radius of spherical shell, σ=surface charge density, and ω=angular velocity.) Findthemagneticinduction B=∇×A. ANS.Br(r,θ)=2µ0a4σω 3·cosθ r3,r>a, Bθ(r,θ)=µ0a4σω 3·sinθ r3,r>a, B=ˆz2µ0aσω 3,r <a. 2.5.25 (a) Explainwhy ∇2inplanepolarcoordinatesfollowsfrom ∇2incircularcylindrical coordinateswith z=constant. (b) Explainwhytaking ∇2insphericalpolarcoordinatesandrestricting θtoπ/2does notleadtotheplanepolarformof ∇. Note. ∇2(ρ,ϕ)=∂2 ∂ρ2+1 ρ∂ ∂ρ+1 ρ2∂2 ∂ϕ2. 2.6 T ENSOR ANALYSIS Introduction, Definitions Tensorsareimportantinmanyareasofphysics,includinggeneralrelativityandelectrody- namics.Scalarsandvectorsarespecialcasesoftensors.InChapter1,aquantitythatdidnot change under rotations of the coordinate system in three-dimensional space, an invariant, was labeled a scalar. A scalaris specified by one real number and is a tensor of rank 0 . A quantity whose components transformed under rotations like those of the distance of a pointfromachosenorigin(Eq.(1.9),Section1.2)wascalledavector.Thetransformation of the componentsof the vectorunder a rotationof thecoordinatespreserves the vectoras a geometric entity (such as an arrow in space), independent of the orientation of the refer- ence frame. In three-dimensional space, a vectoris specified by 3 =31real numbers, for example, its Cartesian components, and is a tensor of rank 1 .Atensor of rank nhas 3n componentsthattransforminadefiniteway.5Thistransformationphilosophyisofcentral importance for tensor analysis and conforms with the mathematician’s concept of vector and vector (or linear) space and the physicist’s notion that physical observables must not dependonthechoiceofcoordinateframes.Thereisaphysicalbasisforsuchaphilosophy: We describe the physical world by mathematics, but any physical predictions we make 5InN-dimensional spacea tensorof rank nhasNncomponents. 134 Chapter 2 Vector Analysis in Curved Coordinates and Tensors must be independent of our mathematical conventions, such as a coordinate system with itsarbitraryoriginandorientationofits axes. Thereis apossibleambiguityinthetransformationlawofa vector A′ i=summationdisplay jaijAj, (2.59) inwhich aijis thecosineof theanglebetweenthe x′ i-axisandthe xj-axis. If we start with a differential distance vector dr, then, taking dx′ ito be a functionof the unprimedvariables, dx′ i=summationdisplay j∂x′ i ∂xjdxj (2.60) bypartialdifferentiation.If weset aij=∂x′ i ∂xj, (2.61) Eqs.(2.59)and(2.60)areconsistent.Anysetofquantities Ajtransformingaccordingto A′i=summationdisplay j∂x′ i ∂xjAj(2.62a) isdefinedasa contravariant vector,whoseindiceswewriteas superscript ;thisincludes theCartesiancoordinatevector xi=xifrom nowon. However,wehavealreadyencounteredaslightlydifferenttypeofvectortransformation. Thegradientofascalar ∇ϕ, definedby ∇ϕ=ˆx∂ϕ ∂x1+ˆy∂ϕ ∂x2+ˆz∂ϕ ∂x3(2.63) (usingx1,x2,x3forx,y,z), transformsas ∂ϕ′ ∂x′i=summationdisplay j∂ϕ ∂xj∂xj ∂x′i, (2.64) usingϕ=ϕ(x,y,z)=ϕ(x′,y′,z′)=ϕ′,ϕdefined as a scalar quantity. Notice that this differs from Eq. (2.62) in that we have ∂xj/∂x′iinstead of ∂x′i/∂xj. Equation (2.64) is taken as the definition of a covariant vector, with the gradient as the prototype. The covariantanalogof Eq.(2.62a) is A′ i=summationdisplay j∂xj ∂x′iAj. (2.62b) OnlyinCartesiancoordinatesis ∂xj ∂x′i=∂x′i ∂xj=aij (2.65) 2.6 Tensor Analysis 135 so that there no difference between contravariant and covariant transformations. In other systems, Eq. (2.65) in general does not apply, and the distinction between contravariant and covariant is real and must be observed. This is of prime importance in the curved Riemannianspaceofgeneralrelativity. Intheremainderofthissectionthecomponentsofany contravariant vectoraredenoted by asuperscript ,Ai, whereas a subscript is used for the components of a covariant vectorAi.6 Definition of Tensors of Rank 2 Nowweproceedtodefine contravariant,mixed,andcovarianttensorsofrank2 bythe followingequationsfortheircomponentsundercoordinatetransformations: A′ij=summationdisplay kl∂x′i ∂xk∂x′j ∂xlAkl, B′ij=summationdisplay kl∂x′i ∂xk∂xl ∂x′jBkl, (2.66) C′ ij=summationdisplay kl∂xk ∂x′i∂xl ∂x′jCkl. Clearly, the rank goes as the number of partial derivatives (or direction cosines) in the de- finition: 0 for a scalar, 1 for a vector, 2 for a second-rank tensor, and so on. Each index (subscript or superscript) ranges over the number of dimensions of the space. The number of indices (equal to the rank of tensor) is independent of the dimensions of the space. We see thatAklis contravariant with respect to both indices, Cklis covariant with respect to bothindices,and Bkltransformscontravariantlywithrespecttothefirstindex kbutcovari- antlywithrespecttothesecondindex l.Onceagain,ifweareusingCartesiancoordinates, all three forms of the tensors of secondrank contravariant,mixed,and covariantare—the same. As with the components of a vector, the transformation laws for the components of a tensor, Eq. (2.66), yield entities (and properties) that are independent of the choice of ref- erence frame. This is what makes tensor analysis important in physics. The independence of reference frame (invariance) is ideal for expressing and investigating universal physical laws. Thesecond-ranktensor A(components Akl)maybeconvenientlyrepresentedbywriting outits componentsinasquarearray(3 ×3 ifweareinthree-dimensionalspace): A= A11A12A13 A21A22A23 A31A32A33. (2.67) This does not mean that any square array of numbers or functions forms a tensor. The essentialconditionisthatthecomponentstransformaccordingtoEq.(2.66). 6Thismeansthatthecoordinates (x,y,z)arewritten (x1,x2,x3)sincertransformsasacontravariantvector.Theambiguityof x2representing both xsquared and yis theprice wepay. 136 Chapter 2 Vector Analysis in Curved Coordinates and Tensors In the context of matrix analysis the preceding transformation equations become (for Cartesiancoordinates)anorthogonalsimilaritytransformation;seeSection3.3.Ageomet- ricalinterpretationofa second-ranktensor(theinertiatensor) isdevelopedinSection3.5. In summary, tensors are systems of components organized by one or more indices that transform according to specific rules under a set of transformations. The number of in- dices is called the rank of the tensor. If the transformations are coordinate rotations in three-dimensional space, then tensor analysis amounts to what we did in the sections on curvilinear coordinates and in Cartesian coordinates in Chapter 1. In four dimensions of Minkowski space–time, the transformations are Lorentz transformations, and tensors of rank1arecalledfour-vectors. Addition and Subtraction of Tensors The addition and subtraction of tensors is defined in terms of the individual elements, just asfor vectors.If A+B=C, (2.68) then Aij+Bij=Cij. Of course, AandBmust be tensors of the same rank and both expressed in a space of the samenumberofdimensions. Summation Convention In tensor analysis it is customary to adopt a summation convention to put Eq. (2.66) and subsequent tensor equations in a more compact form. As long as we are distinguishing betweencontravarianceandcovariance,letusagreethatwhenanindexappearsononeside of an equation, once as a superscript and once as a subscript (except for the coordinates where both are subscripts), we automatically sum over that index. Then we may write the secondexpressioninEq. (2.66) as B′ij=∂x′i ∂xk∂xl ∂x′jBkl, (2.69) withthesummationoftheright-handsideover kandlimplied.ThisisEinstein’ssumma- tion convention.7The index iis superscript because it is associated with the contravariant x′i;likewise jissubscriptbecauseitisrelatedtothecovariantgradient. To illustrate the use of the summation convention and some of the techniques of tensor analysis, let us show that the now-familiar Kronecker delta, δkl, is really a mixed tensor 7In this context ∂x′i/∂xkmight better be written as ai kand∂xl/∂x′jasbl j. 2.6 Tensor Analysis 137 of rank 2, δkl.8The question is: Does δkltransform according to Eq. (2.66)? This is our criterionfor callingitatensor.Wehave,usingthesummationconvention, δkl∂x′i ∂xk∂xl ∂x′j=∂x′i ∂xk∂xk ∂x′j(2.70) bydefinitionoftheKroneckerdelta.Now, ∂x′i ∂xk∂xk ∂x′j=∂x′i ∂x′j(2.71) by direct partial differentiation of the right-hand side (chain rule). However, x′iandx′j are independent coordinates, and therefore the variation of one with respect to the other mustbezeroif theyaredifferent, unityif theycoincide;thatis, ∂x′i ∂x′j=δ′ij. (2.72) Hence δ′ij=∂x′i ∂xk∂xl ∂x′jδkl, showingthatthe δklareindeedthecomponentsofamixedsecond-ranktensor.Noticethat this result is independent of the number of dimensions of our space. The reason for the upperindex iandlowerindex jisthesameasinEq. (2.69). TheKroneckerdeltahasonefurtherinterestingproperty.Ithasthesamecomponentsin all of our rotated coordinate systems and is therefore called isotropic. In Section 2.9 we shallmeetathird-rankisotropictensorandthreefourth-rankisotropictensors.Noisotropic first-ranktensor(vector)exists. Symmetry–Antisymmetry Theorderinwhichtheindicesappearinourdescriptionofatensorisimportant.Ingeneral, Amnisindependentof Anm,buttherearesomecasesofspecialinterest.If,forall mandn, Amn=Anm, (2.73) wecallthetensor symmetric .If, ontheotherhand, Amn=−Anm, (2.74) thetensoris antisymmetric .Clearly,every(second-rank)tensorcanberesolvedintosym- metricandantisymmetricpartsbytheidentity Amn=1 2parenleftbig Amn+Anmparenrightbig +1 2parenleftbig Amn−Anmparenrightbig , (2.75) the first term on the right being a symmetric tensor, the second, an antisymmetric tensor. A similar resolution of functions into symmetric and antisymmetric parts is of extreme importancetoquantummechanics. 8Itiscommonpracticetorefertoatensor Abyspecifyingatypicalcomponent, Aij.Aslongasthereaderrefrainsfromwriting nonsense suchas A=Aij, no harm is done. 138 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Spinors It was once thought that the system of scalars, vectors, tensors (second-rank), and so on formed a complete mathematical system, one that is adequate for describing a physics independent of the choice of reference frame. But the universe and mathematical physics are not that simple. In the realm of elementary particles, for example, spin zero particles9 (πmesons,αparticles) may be described with scalars, spin 1 particles (deuterons) by vectors, and spin 2 particles (gravitons) by tensors. This listing omits the most common particles: electrons, protons, and neutrons, all with spin1 2. These particles are properly described by spinors. A spinor is not a scalar, vector, or tensor. A brief introduction to spinorsinthecontextof grouptheory (J=1/2)appearsinSection4.3. Exercises 2.6.1 Show that if all the components of any tensor of any rank vanish in one particular coordinatesystem, theyvanishinallcoordinatesystems. Note.This point takes on special importance in the four-dimensional curved space of generalrelativity.Ifaquantity,expressedasatensor,existsinonecoordinatesystem,it existsinallcoordinatesystemsandisnotjustaconsequenceofa choiceofacoordinate system(as arecentrifugalandCoriolisforcesinNewtonianmechanics). 2.6.2 The components of tensor Aare equal to the corresponding components of tensor Bin oneparticularcoordinatesystem, denotedbythesuperscript 0;thatis, A0 ij=B0 ij. Showthattensor Aisequaltotensor B,Aij=Bij, inallcoordinatesystems. 2.6.3 Thelastthreecomponentsofafour-dimensionalvectorvanishineachoftworeference frames. If the second reference frame is not merely a rotation of the first about the x0 axis, that is, if at least one of the coefficients ai0(i=1,2,3)/negationslash=0, show that the zeroth component vanishes in all reference frames. Translated into relativistic mechanics this means that if momentum is conserved in two Lorentz frames, then energy is conserved inallLorentzframes. 2.6.4 From an analysis of the behavior of a general second-rank tensor under 90◦and 180◦ rotations about the coordinate axes, show that an isotropic second-rank tensor in three- dimensionalspacemustbeamultipleof δij. 2.6.5 The four-dimensional fourth-rank Riemann–Christoffel curvature tensor of general rel- ativity,Riklm,satisfiesthesymmetryrelations Riklm=−Rikml=−Rkilm. Withtheindicesrunningfrom0to3,showthatthenumberofindependentcomponents isreducedfrom256to36andthatthecondition Riklm=Rlmik 9The particle spin is intrinsic angular momentum (in units of ¯h). It is distinct from classical, orbital angular momentum due to motion. 2.7 Contraction, Direct Product 139 furtherreducesthenumberofindependentcomponentsto21.Finally,ifthecomponents satisfy an identity Riklm+Rilmk+Rimkl=0, show that the number of independent componentsis reducedto20. Note.Thefinalthree-termidentityfurnishesnewinformationonlyifallfourindicesare different.Thenitreducesthenumberofindependentcomponentsbyone-third. 2.6.6 Tiklmisantisymmetricwithrespecttoallpairsofindices.Howmanyindependentcom- ponentshasit(in three-dimensionalspace)? 2.7 C ONTRACTION ,DIRECT PRODUCT Contraction Whendealingwithvectors,weformedascalarproduct(Section1.3)bysummingproducts ofcorrespondingcomponents: A·B=AiBi(summationconvention ). (2.76) The generalization of this expression in tensor analysis is a process known as contraction. Twoindices,onecovariantandtheothercontravariant,aresetequaltoeachother,andthen (as implied by the summation convention) we sum over this repeated index. For example, letuscontractthesecond-rankmixedtensor B′ij, B′ii=∂x′i ∂xk∂xl ∂x′iBkl=∂xl ∂xkBkl (2.77) usingEq. (2.71), andthenbyEq. (2.72) B′ii=δlkBkl=Bkk. (2.78) Our contracted second-rank mixed tensor is invariant and therefore a scalar.10This is ex- actlywhatweobtainedinSection1.3forthedotproductoftwovectorsandinSection1.7 for the divergence of a vector. In general, the operation of contraction reduces the rank of atensorby2.Anexampleof theuseof contractionappearsinChapter4. Direct Product Thecomponentsofacovariantvector(first-ranktensor) aiandthoseofacontravariantvec- tor (first-rank tensor) bjmay be multiplied component by component to give the general termaibj. This, byEq. (2.66)is actuallyasecond-ranktensor,for a′ ib′j=∂xk ∂x′iak∂x′j ∂xlbl=∂xk ∂x′i∂x′j ∂xlparenleftbig akblparenrightbig . (2.79) Contracting,weobtain a′ ib′i=akbk, (2.80) 10In matrix analysis this scalaris the traceof the matrix, Section 3.2. 140 Chapter 2 Vector Analysis in Curved Coordinates and Tensors asinEqs. (2.77) and(2.78), togivetheregularscalarproduct. The operation of adjoining two vectors aiandbjas in the last paragraph is known as forming the direct product . For the case of two vectors, the direct product is a tensor of second rank. In this sense we may attach meaning to ∇E, which was not defined within the framework of vector analysis. In general, the direct product of two tensors is a tensor ofrankequaltothesumofthetwoinitialranks;thatis, AijBkl=Cijkl, (2.81a) whereCijklisatensorof fourthrank.FromEqs. (2.66), C′ijkl=∂x′i ∂xm∂xn ∂x′j∂x′k ∂xp∂x′l ∂xqCmnpq. (2.81b) The direct product is a technique for creating new, higher-rank tensors. Exer- cise2.7.1isaformofthedirectproductinwhichthefirstfactoris ∇.Applicationsappear inSection4.6. WhenTis annth-rank Cartesian tensor, (∂/∂xi)Tjkl...,a component of ∇T,i sa Cartesian tensor of rank n+1 (Exercise 2.7.1). However, (∂/∂xi)Tjkl...is not a tensor inmoregeneralspaces.Innon-Cartesiansystems ∂/∂x′iwillactonthepartialderivatives ∂xp/∂x′qanddestroythesimpletensortransformationrelation(seeEq. (2.129)). So far the distinction between a covariant transformation and a contravariant transfor- mationhasbeenmaintainedbecauseitdoesexistinnon-Euclideanspaceandbecauseitis ofgreatimportanceingeneralrelativity.InSections2.10and2.11weshalldevelopdiffer- entialrelationsforgeneraltensors.Often,however,becauseofthesimplificationachieved, werestrict ourselves to Cartesiantensors. As notedin Section2.6, thedistinctionbetween contravarianceandcovariancedisappears. Exercises 2.7.1 IfT···iis a tensor of rank n, show that ∂T···i/∂xjis a tensor of rank n+1( C a r t e s i a n coordinates). Note.Innon-Cartesiancoordinatesystemsthecoefficients aijare,ingeneral,functions of the coordinates, and the simple derivative of a tensor of rank nis not a tensor except in the special case of n=0. In this case the derivative does yield a covariant vector (tensorof rank1)byEq.(2.64). 2.7.2 IfTijk···is a tensor of rank n, show thatsummationtext j∂Tijk···/∂xjis a tensor of rank n−1 (Cartesiancoordinates). 2.7.3 Theoperator ∇2−1 c2∂2 ∂t2 maybewrittenas 4summationdisplay i=1∂2 ∂x2 i, 2.8 Quotient Rule 141 usingx4=ict. This is the four-dimensional Laplacian, sometimes called the d’Alem- bertian and denoted by /square2. Show that it is a scalaroperator, that is, is invariant under Lorentztransformations. 2.8 Q UOTIENT RULE IfAiandBjarevectors,asseeninSection2.7,wecaneasilyshowthat AiBjisasecond- ranktensor.Hereweareconcernedwithavarietyofinverserelations.Considersuchequa- tionsas KiAi=B (2.82a) KijAj=Bi (2.82b) KijAjk=Bik (2.82c) KijklAij=Bkl (2.82d) KijAk=Bijk. (2.82e) Inline with our restriction to Cartesian systems, we write all indices as subscripts and, unlessspecifiedotherwise,sumrepeatedindices. Ineachoftheseexpressions AandBareknowntensorsofrankindicatedbythenumber of indices and Ais arbitrary. In each case Kis an unknown quantity. We wish to establish thetransformationpropertiesof K.Thequotientruleassertsthatiftheequationofinterest holdsinall(rotated)Cartesiancoordinatesystems, Kisatensoroftheindicatedrank.The importance in physical theory is that the quotient rule can establish the tensor nature of quantities. Exercise 2.8.1 is a simple illustration of this. The quotient rule (Eq. (2.82b)) shows that the inertia matrix appearing in the angular momentum equation L=Iω, Sec- tion3.5, isa tensor. In proving the quotient rule, we consider Eq. (2.82b) as a typical case. In our primed coordinatesystem K′ ijA′ j=B′ i=aikBk, (2.83) using the vector transformation properties of B. Since the equation holds in all rotated Cartesiancoordinatesystems, aikBk=aik(KklAl). (2.84) Now, transforming Aback into the primed coordinate system11(compare Eq. (2.62)), we have K′ ijA′j=aikKklajlA′j. (2.85) Rearranging,weobtain (K′ ij−aikajlKkl)A′ j=0. (2.86) 11Notethe order of the indices of the direction cosine ajlin thisinversetransformation. Wehave Al=summationdisplay j∂xl ∂x′ jA′ j=summationdisplay jajlA′j. 142 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Thismustholdforeachvalueoftheindex iandforeveryprimedcoordinatesystem.Since theA′ jisarbitrary,12weconclude K′ ij=aikajlKkl, (2.87) whichisourdefinitionofsecond-ranktensor. The other equations may be treated similarly, giving rise to other forms of the quotient rule. One minor pitfall should be noted: The quotient rule does not necessarily apply if B iszero. Thetransformationpropertiesof zeroareindeterminate. Example 2.8.1 EQUATIONS OF MOTION AND FIELD EQUATIONS In classical mechanics, Newton’s equations of motion m˙v=Ftell us on the basis of the quotientrule that, if the mass is a scalar and the force a vector, then the acceleration a≡˙v isavector.Inotherwords,thevectorcharacteroftheforceasthedrivingtermimposesits vectorcharacterontheacceleration,providedthescalefactor misscalar. The wave equation of electrodynamics ∂2Aµ=Jµinvolves the four-dimensional ver- sionoftheLaplacian ∂2=∂2 c2∂t2−∇2,aLorentzscalar,andtheexternalfour-vectorcurrent Jµas its driving term. From the quotient rule, we infer that the vector potential Aµis a four-vector as well. If the driving current is a four-vector, the vector potential must be of rank1bythequotientrule. /squaresolid Thequotientruleis asubstitutefor theillegaldivisionoftensors. Exercises 2.8.1 The double summation KijAiBjis invariant for any two vectors AiandBj. Prove that Kijisasecond-ranktensor. Note.In the form ds2(invariant)=gijdxidxj, this result shows that the matrix gijis atensor. 2.8.2 Theequation KijAjk=Bikholdsforallorientationsofthecoordinatesystem.If Aand Bare arbitrarysecond-ranktensors,showthat Kis asecond-ranktensoralso. 2.8.3 Theexponentialinaplanewaveisexp [i(k·r−ωt)].Werecognize xµ=(ct,x1,x2,x3) asaprototypevectorinMinkowskispace.If k·r−ωtisascalarunderLorentztransfor- mations(Section4.5), showthat kµ=(ω/c,k 1,k2,k3)isavectorinMinkowskispace. Note.Multiplicationby ¯hyields(E/c,p)asavectorinMinkowskispace. 2.9 P SEUDOTENSORS ,DUAL TENSORS So far our coordinate transformations have been restricted to pure passive rotations. We nowconsidertheeffectofreflectionsor inversions. 12We might, for instance, take A′ 1=1a n dA′m=0f o rm/negationslash=1. Then the equation K′ i1=aika1lKklfollows immediately. The rest of Eq.(2.87) comes from otherspecial choicesof the arbitrary A′ j. 2.9 Pseudotensors, Dual Tensors 143 FIGURE 2.9Inversionof Cartesiancoordinates—polarvector. If wehavetransformationcoefficients aij=−δij, thenbyEq. (2.60) xi=−x′i, (2.88) which is an inversion or parity transformation. Note that this transformation changes our initial right-handed coordinate system into a left-handed coordinate system.13Our proto- typevector rwithcomponents (x1,x2,x3)transformsto r′=parenleftbig x′1,x′2,x′3parenrightbig =parenleftbig −x1,−x2,−x3parenrightbig . This new vector r′has negative components, relative to the new transformed set of axes. As shown in Fig. 2.9, reversing the directions of the coordinate axes and changing the signs of the components gives r′=r. The vector (an arrow in space) stays exactly as it was before the transformation was carried out. The position vector rand all other vectors whose components behave this way (reversing sign with a reversal of the coordinate axes) arecalled polarvectors andhaveoddparity. Afundamentaldifferenceappearswhenweencounteravectordefinedasthecrossprod- uct of two polar vectors. Let C=A×B, where both AandBare polar vectors. From Eq.(1.33), thecomponentsof Caregivenby C1=A2B3−A3B2(2.89) andsoon.Now,whenthecoordinateaxesareinverted, Ai→−A′i,Bj→−B′ j,butfrom itsdefinition Ck→+C′k;thatis,ourcross-productvector,vector C,doesnotbehavelike a polar vector under inversion. To distinguish, we label it a pseudovector or axial vector (seeFig.2.10)thathasevenparity.Theterm axialvector isfrequentlyusedbecausethese crossproductsoftenarisefroma descriptionofrotation. 13This is an inversion of thecoordinate system or coordinate axes, objects in the physical world remaining fixed. 144 Chapter 2 Vector Analysis in Curved Coordinates and Tensors FIGURE 2.10InversionofCartesiancoordinates—axialvector. Examplesare angularvelocity, v=ω×r, orbitalangularmomentum, L=r×p, torque,force=F, N=r×F, magneticinductionfield B,∂B ∂t=−∇×E. Inv=ω×r, the axial vector is the angular velocity ω, andrandv=dr/dtare polar vectors. Clearly, axial vectors occur frequently in physics, although this fact is usually not pointed out. In a right-handed coordinate system an axial vector Chas a sense of rotationassociatedwithitgivenbyaright-handrule(compareSection1.4).Intheinverted left-handed system the sense of rotation is a left-handed rotation. This is indicated by the curvedarrowsinFig.2.10. The distinction between polar and axial vectors may also be illustrated by a reflection. Apolarvectorreflectsinamirrorlikearealphysicalarrow,Fig.2.11a.InFigs.2.9and2.10 the coordinates are inverted; the physical world remains fixed. Here the coordinate axes remain fixed; the world is reflected—as in a mirror in the xz-plane. Specifically, in this representation we keep the axes fixed and associate a change of sign with the component ofthevector.Foramirrorinthe xz-plane,Py→−Py.W eh a v e P=(Px,Py,Pz) P′=(Px,−Py,Pz)polarvector. An axial vector such as a magnetic field Hor a magnetic moment µ(=current×area of current loop) behaves quite differently under reflection. Consider the magnetic field Hand magnetic moment µto be produced by an electric charge moving in a circular path (Exercise5.8.4andExample12.5.3).Reflectionreversesthesenseofrotationofthecharge. 2.9 Pseudotensors, Dual Tensors 145 a b FIGURE 2.11(a) Mirror in xz-plane;(b) mirror inxz-plane. The two current loops and the resulting magnetic moments are shown in Fig. 2.11b. We have µ=(µx,µy,µz) µ′=(−µx,µy,−µz)reflectedaxialvector. 146 Chapter 2 Vector Analysis in Curved Coordinates and Tensors If we agree that the universe does not care whether we use a right- or left-handed coor- dinate system, then it does not make sense to add an axial vector to a polar vector. In the vector equation A=B, bothAandBare either polar vectors or axial vectors.14Similar restrictionsapplytoscalarsandpseudoscalarsand,ingeneral,tothetensorsandpseudoten- sors consideredsubsequently. Usually,pseudoscalars,pseudovectors,andpseudotensorswilltransformas S′=JS, C′ i=JaijCj,A′ ij=JaikajlAkl, (2.90) whereJis the determinant15of the array of coefficients amn, the Jacobian of the parity transformation.In ourinversiontheJacobianis J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−10 0 0−10 00−1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−1. (2.91) Forareflectionofoneaxis,the x-axis, J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−100 01 0 00 1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−1, (2.92) andagaintheJacobian J=−1.Ontheotherhand,forallpurerotations,theJacobian Jis always+1.RotationmatricesdiscussedfurtherinSection3.3. In Chapter 1 the triple scalar product S=A×B·Cwas shown to be a scalar (un- der rotations). Now by considering the parity transformation given by Eq. (2.88), we see thatS→−S, proving that the triple scalar product is actually a pseudoscalar: This be- havior was foreshadowed by the geometrical analogy of a volume. If all three parameters of the volume—length, depth, and height—change from positive distances to negative distances,theproductof thethreewillbenegative. Levi-Civita Symbol For future use it is convenient to introduce the three-dimensional Levi-Civita symbol εijk, definedby ε123=ε231=ε312=1, ε132=ε213=ε321=−1, (2.93) allotherεijk=0. Note that εijkis antisymmetric with respect to all pairs of indices. Suppose now that we have a third-rank pseudotensor δijk, which in one particular coordinate system is equal to εijk. Then δ′ ijk=|a|aipajqakrεpqr (2.94) 14The big exception to this is in beta decay, weak interactions. Here the universe distinguishes between right- and left-handed systems, andweaddpolarandaxial vector interactions. 15Determinants are describedinSection3.1. 2.9 Pseudotensors, Dual Tensors 147 bydefinitionofpseudotensor.Now, a1pa2qa3rεpqr=|a| (2.95) by direct expansion of the determinant, showing that δ′ 123=|a|2=1=ε123. Considering theotherpossibilitiesonebyone,wefind δ′ ijk=εijk (2.96) for rotations and reflections. Hence εijkis a pseudotensor.16,17Furthermore, it is seen to beanisotropicpseudotensorwiththesamecomponentsinallrotatedCartesiancoordinate systems. Dual Tensors Withany antisymmetric second-ranktensor C(inthree-dimensionalspace)wemayasso- ciateadualpseudovector Cidefinedby Ci=1 2εijkCjk. (2.97) Heretheantisymmetric Cmaybewritten C= 0C12−C31 −C120C23 C31−C230. (2.98) Weknowthat Cimusttransformasavectorunderrotationsfromthedoublecontractionof the fifth-rank (pseudo) tensor εijkCmnbut that it is really a pseudovector from the pseudo natureof εijk. Specifically,thecomponentsof Caregivenby (C1,C2,C3)=parenleftbig C23,C31,C12parenrightbig . (2.99) Notice the cyclic order of the indices that comes from the cyclic order of the components ofεijk. Eq. (2.99) means that our three-dimensional vector product may literally be taken tobeeitherapseudovectororanantisymmetricsecond-ranktensor,dependingonhowwe choosetowriteitout. If wetakethree(polar)vectors A,B, andC,wemaydefinethedirectproduct Vijk=AiBjCk. (2.100) By an extension of the analysis of Section 2.6, Vijkis a tensor of third rank. The dual quantity V=1 3!εijkVijk(2.101) 16The usefulness of εpqrextends far beyond this section. For instance, the matrices Mkof Exercise 3.2.16 are derived from (Mr)pq=−iεpqr. Much of elementary vector analysis can be written in a very compact form by using εijkand the identity of Exercise 2.9.4SeeA.A.Evett,Permutation symbol approachtoelementaryvector analysis. Am.J.Phys. 34: 503 (1966). 17Thenumerical valueof εpqris given by the triple scalarproduct of coordinate unit vectors: ˆxp·ˆxq׈xr. From this point of view eachelementof εpqris a pseudoscalar, but the εpqrcollectively form athird-rank pseudotensor. 148 Chapter 2 Vector Analysis in Curved Coordinates and Tensors isclearlyapseudoscalar.Byexpansionit isseenthat V=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleA1B1C1 A2B2C2 A3B3C3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(2.102) isourfamiliar triplescalarproduct. ForuseinwritingMaxwell’sequationsincovariantform,Section4.6,wewanttoextend this dual vector analysis to four-dimensional space and, in particular, to indicate that the four-dimensionalvolumeelement dx0dx1dx2dx3isapseudoscalar. We introduce the Levi-Civita symbol εijkl, the four-dimensional analog of εijk.T h i s quantityεijklis defined as totally antisymmetric in all four indices. If (ijkl)is an even permutation18of(0,1,2,3), thenεijklis defined as+1; if it is an odd permutation, thenεijklis−1, and 0 if any two indices are equal. The Levi-Civita εijklm a yb ep r o v e da pseudotensorofrank4byanalysissimilartothatusedforestablishingthetensornatureof εijk. Introducingthedirectproductof fourvectorsasfourth-ranktensorwithcomponents Hijkl=AiBjCkDl, (2.103) builtfromthepolarvectors A,B,C,andD,wemaydefinethedualquantity H=1 4!εijklHijkl, (2.104) a pseudoscalar due to the quadruple contraction with the pseudotensor εijkl.Now we let A,B,C, andDbe infinitesimal displacements along the four coordinate axes (Minkowski space), A=parenleftbig dx0,0,0,0parenrightbig B=parenleftbig 0,dx1,0,0parenrightbig ,andso on,(2.105) and H=dx0dx1dx2dx3. (2.106) The four-dimensional volume element is now identified as a pseudoscalar. We use this result in Section 4.6. This result could have been expected from the results of the special theory of relativity. The Lorentz–Fitzgerald contraction of dx1dx2dx3just balances the timedilationof dx0. We slipped into this four-dimensional space as a simple mathematical extension of the three-dimensional space and, indeed, we could just as easily have discussed 5-, 6-, or N- dimensionalspace.Thisistypicalofthepowerofthecomponentanalysis.Physically,this four-dimensionalspacemaybetakenasMinkowskispace, parenleftbig x0,x1,x2,x3parenrightbig =(ct,x,y,z), (2.107) wheretis time. This is the merger of space and time achieved in special relativity. The transformationsthatdescribetherotationsinfour-dimensionalspacearetheLorentztrans- formationsofspecialrelativity.WeencountertheseLorentztransformationsinSection4.6. 18A permutation is odd if it involves an odd number of interchanges of adjacent indices, such as (0123)→(0213).E v e n permutations arise from an even number of transpositions of adjacent indices. (Actually the word adjacent is unnecessary.) ε0123=+1. 2.9 Pseudotensors, Dual Tensors 149 Irreducible Tensors Forsomeapplications,particularlyinthequantumtheoryofangularmomentum,ourCarte- siantensorsarenotparticularlyconvenient.Inmathematicallanguageourgeneralsecond- rank tensor Aijis reducible, which means that it can be decomposed into parts of lower tensorrank.Infact, wehavealreadydonethis.FromEq.(2.78), A=Aii (2.108) isascalarquantity,thetraceof Aij.19 Theantisymmetricportion, Bij=1 2(Aij−Aji), (2.109) hasjustbeenshowntobeequivalenttoa(pseudo)vector,or Bij=Ckcyclicpermutationof i,j,k. (2.110) By subtracting the scalar Aand the vector Ckfrom our original tensor, we have an irre- ducible,symmetric,zero-tracesecond-ranktensor, Sij,inwhich Sij=1 2(Aij+Aji)−1 3Aδij, (2.111) withfiveindependentcomponents.Then,finally,ouroriginalCartesiantensormaybewrit- ten Aij=1 3Aδij+Ck+Sij. (2.112) The three quantities A,Ck, andSijform spherical tensors of rank 0, 1, and 2, respec- tively, transforming like the spherical harmonics YM L(Chapter 12) for L=0, 1, and 2. Further details of such spherical tensors and their uses will be found in Chapter 4 and the booksbyRoseandEdmondscitedthere. A specific example of the preceding reduction is furnished by the symmetric electric quadrupoletensor Qij=integraldisplayparenleftbig 3xixj−r2δijparenrightbig ρ(x1,x2,x3)d3x. The−r2δijterm represents a subtraction of the scalar trace (the three i=jterms). The resulting Qijhaszerotrace. Exercises 2.9.1 Anantisymmetricsquarearrayis givenby  0C3−C2 −C30C1 C2−C10=0C12C13 −C120C23 −C13−C230, 19Analternateapproach, using matrices, is given inSection3.3(see Exercise3.3.9). 150 Chapter 2 Vector Analysis in Curved Coordinates and Tensors where(C1,C2,C3)formapseudovector.Assumingthattherelation Ci=1 2!εijkCjk holds in all coordinate systems, prove that Cjkis a tensor. (This is another form of the quotienttheorem.) 2.9.2 Showthatthevectorproductisuniquetothree-dimensionalspace;thatis,onlyinthree dimensions can we establish a one-to-one correspondence between the components of anantisymmetrictensor(second-rank)andthecomponentsofavector. 2.9.3 Showthatin R3 (a)δii=3, (b)δijεijk=0, (c)εipqεjpq=2δij, (d)εijkεijk=6. 2.9.4 Showthatin R3 εijkεpqk=δipδjq−δiqδjp. 2.9.5 (a) Express the components of a cross-product vector C,C=A×B, in terms of εijk andthecomponentsof AandB. (b) Usetheantisymmetryof εijktoshowthat A·A×B=0. ANS.(a)Ci=εijkAjBk. 2.9.6 (a) Showthattheinertiatensor(matrix)maybewritten Iij=m(xixjδij−xixj) fora particleofmass mat(x1,x2,x3). (b) Showthat Iij=−MilMlj=−mεilkxkεljmxm, whereMil=m1/2εilkxk.Thisisthecontractionoftwosecond-ranktensorsandis identicalwiththematrixproductof Section3.2. 2.9.7 Write ∇·∇×Aand∇×∇ϕintensor(index)notationin R3sothatitbecomesobvious thateachexpressionvanishes. ANS.∇·∇×A=εijk∂ ∂xi∂ ∂xjAk, (∇×∇ϕ)i=εijk∂ ∂xj∂ ∂xkϕ. 2.9.8 ExpressingcrossproductsintermsofLevi-Civitasymbols (εijk),derivethe BAC–CAB rule,Eq. (1.55). Hint.TherelationofExercise2.9.4is helpful. 2.10 General Tensors 151 2.9.9 Verify that each of the following fourth-rank tensors is isotropic, that is, that it has the sameformindependentofanyrotationofthecoordinatesystems. (a)Aijkl=δijδkl, (b)Bijkl=δikδjl+δilδjk, (c)Cijkl=δikδjl−δilδjk. 2.9.10 Showthatthetwo-indexLevi-Civitasymbol εijisasecond-rankpseudotensor(intwo- dimensionalspace).Doesthiscontradicttheuniquenessof δij(Exercise 2.6.4)? 2.9.11 Represent εijbya2×2matrix,andusingthe2 ×2rotationmatrixofSection3.3show thatεijisinvariantunderorthogonalsimilaritytransformations. 2.9.12 GivenAk=1 2εijkBijwithBij=−Bji, antisymmetric,showthat Bmn=εmnkAk. 2.9.13 Showthatthevectoridentity (A×B)·(C×D)=(A·C)(B·D)−(A·D)(B·C) (Exercise 1.5.12) follows directly from the description of a cross product with εijkand theidentityof Exercise2.9.4. 2.9.14 Generalize the cross product of two vectors to n-dimensional space for n=4,5,.... Check the consistency of your construction and discuss concrete examples. See Exer- cise1.4.17for thecase n=2. 2.10 G ENERAL TENSORS The distinction between contravariant and covariant transformations was established in Section 2.6. Then, for convenience, we restricted our attention to Cartesian coordinates (in which the distinction disappears). Now in these two concluding sections we return to non-Cartesiancoordinatesandresurrectthecontravariantandcovariantdependence.Asin Section 2.6, a superscript will be used for an index denoting contravariant and a subscript for an index denoting covariant dependence.The metric tensor of Section 2.1 will be used torelatecontravariantandcovariantindices. The emphasis in this section is on differentiation, culminating in the construction of thecovariant derivative . We saw in Section 2.7 that the derivative of a vector yields a second-rank tensor—in Cartesian coordinates. In non-Cartesian coordinate systems, it is thecovariantderivativeofavectorratherthantheordinaryderivativethatyieldsasecond- ranktensorbydifferentiationof avector. Metric Tensor Let us start with the transformation of vectors from one set of coordinates (q1,q2,q3) to another r=(x1,x2,x3). The new coordinates are (in general nonlinear ) functions 152 Chapter 2 Vector Analysis in Curved Coordinates and Tensors xi(q1,q2,q3)of the old, such as spherical polar coordinates (r,θ,φ). But their differ- entialsobeythelineartransformationlaw dxi=∂xi ∂qjdqj, (2.113a) or dr=εjdqj(2.113b) in vector notation. For convenience we take the basis vectors ε1=(∂x1 ∂q1,∂x1 ∂q2,∂x1 ∂q3),ε2, andε3to form a right-handed set. These vectors are not necessarily orthogonal. Also, a limitation to three-dimensional space will be required only for the discussions of cross products and curls. Otherwise these εimay be in N-dimensional space, including the four-dimensional space–time of special and general relativity. The basis vectors εimay beexpressedby εi=∂r ∂qi, (2.114) asinExercise2.2.3.Note,however,thatthe εiheredonotnecessarilyhaveunitmagnitude. FromExercise2.2.3, theunitvectorsare ei=1 hi∂r ∂qi(nosummation) , andtherefore εi=hiei(nosummation) . (2.115) Theεiarerelatedtotheunitvectors eibythescalefactors hiofSection2.2.The eihaveno dimensions; the εihave the dimensions of hi. In spherical polar coordinates, as a specific example, εr=er=ˆr,εθ=reθ=rˆθ,εϕ=rsinθeϕ=rsinθˆϕ.(2.116) In Euclidean spaces, or in Minkowski space of special relativity, the partial derivatives in Eq.(2.113)areconstantsthatdefinethenewcoordinatesintermsoftheoldones.Weused them to define the transformation laws of vectors in Eq. (2.59) and (2.62) and tensors in Eq. (2.66). Generalizing, we define a contravariant vectorViundergeneralcoordinate transformationsif itscomponentstransform accordingto V′i=∂xi ∂qjVj, (2.117a) or V′=Vjεj (2.117b) in vector notation. For covariant vectors we inspect the transformation of the gradient operator ∂ ∂xi=∂qj ∂xi∂ ∂qj(2.118) 2.10 General Tensors 153 usingthechainrule. From ∂xi ∂qj∂qj ∂xk=δik (2.119) itisclearthatEq.(2.118) isrelatedtothe inversetransformationofEq. (2.113), dqj=∂qj ∂xidxi. (2.120) Hencewedefinea covariant vectorViif V′ i=∂qj ∂xiVj (2.121a) holdsor, invectornotation, V′=Vjεj, (2.121b) where εjarethecontravariantvectors gjiεi=εj. Second-ranktensorsaredefinedas inEq.(2.66), A′ij=∂xi ∂qk∂xj ∂qlAkl, (2.122) andtensors ofhigherranksimilarly. AsinSection2.1, weconstructthesquareofadifferentialdisplacement (ds)2=dr·dr=parenleftbig εidqiparenrightbig2=εi·εjdqidqj. (2.123) Comparing this with (ds)2of Section 2.1, Eq. (2.5), we identify εi·εjas the covariant metrictensor εi·εj=gij. (2.124) Clearly,gijis symmetric. The tensor nature of gijfollows from the quotient rule, Exer- cise2.8.1.We taketherelation gikgkj=δij (2.125) to define the corresponding contravariant tensor gik. Contravariant gikenters as the in- verse20of covariant gkj. We use this contravariant gikto raise indices, converting a co- variantindexintoacontravariantindex,asshownsubsequently.Likewisethecovariant gkj willbeusedtolowerindices.Thechoiceof gikandgkjforthisraising–loweringoperation isarbitrary. Anysecond-ranktensor(anditsinverse)woulddo.Specifically,wehave gijεj=εirelatingcovariantand contravariantbasisvectors, gijFj=Firelatingcovariantand contravariantvectorcomponents.(2.126) 20If thetensor gkjis written as amatrix, thetensor gikis given by the inverse matrix. 154 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Then gijεj=εias thecorrespondingindex gijFj=Filoweringrelations.(2.127) It should be emphasized again that the εiandεjdonothave unit magnitude. This may be seen in Eqs. (2.116) and in the metric tensor gijfor spherical polar coordinates and its inversegij: (gij)= 10 0 0r20 00r2sin2θ parenleftbig gijparenrightbig = 10 0 01 r20 001 r2sin2θ . Christoffel Symbols Letusform thedifferentialofascalar ψ, dψ=∂ψ ∂qidqi. (2.128) Since the dqiare the components of a contravariant vector, the partial derivatives ∂ψ/∂qimust form a covariant vector—by the quotient rule. The gradient of a scalar be- comes ∇ψ=∂ψ ∂qiεi. (2.129) Note that ∂ψ/∂qiare not the gradient components of Section 2.2—because εi/negationslash=eiof Section2.2. Movingontothederivativesofavector,wefindthatthesituationismuchmorecompli- catedbecausethebasisvectors εiareingeneralnotconstant.Remember,wearenolonger restricting ourselves to Cartesian coordinates and the nice, convenient ˆx,ˆy,ˆz! Direct dif- ferentiationofEq. (2.117a)yields ∂V′k ∂qj=∂xk ∂qi∂Vi ∂qj+∂2xk ∂qj∂qiVi, (2.130a) or,invectornotation, ∂V′ ∂qj=∂Vi ∂qjεi+Vi∂εi ∂qj. (2.130b) TherightsideofEq.(2.130a)differsfromthetransformationlawforasecond-rankmixed tensor by the second term, which contains second derivatives of the coordinates xk.T h e latterare nonzerofornonlinearcoordinatetransformations. Now,∂εi/∂qjwillbesomelinearcombinationofthe εk,withthecoefficientdepending on the indices iandjfrom the partial derivative and index kfrom the base vector. We write ∂εi ∂qj=Ŵk ijεk. (2.131a) 2.10 General Tensors 155 Multiplyingby εmandusing εm·εk=δm kfrom Exercise2.10.2, wehave Ŵm ij=εm·∂εi ∂qj. (2.131b) TheŴk ijis a Christoffel symbol of the second kind . It is also called a coefficient of con- nection. TheseŴk ijarenotthird-rank tensors and the ∂Vi/∂qjof Eq. (2.130a) are not second-ranktensors. Equations(2.131)shouldbecomparedwiththeresultsquotedinEx- ercise2.2.3(rememberingthatingeneral εi/negationslash=ei).InCartesiancoordinates, Ŵk ij=0forall values of the indices i,j, andk. These Christoffel three-index symbols may be computed by the techniques of Section 2.2. This is the topic of Exercise 2.10.8. Equation (2.138) offers aneasiermethod.UsingEq.(2.114), weobtain ∂εi ∂qj=∂2r ∂qj∂qi=∂εj ∂qi=Ŵk jiεk. (2.132) HencetheseChristoffel symbolsaresymmetricinthetwolowerindices: Ŵk ij=Ŵk ji. (2.133) Christoffel Symbols as Derivatives of the Metric Tensor ItisoftenconvenienttohaveanexplicitexpressionfortheChristoffelsymbolsintermsof derivatives of the metric tensor. As an initial step, we define the Christoffel symbol of the firstkind[ij,k]by [ij,k]≡gmkŴm ij, (2.134) from which the symmetry [ij,k]=[ji,k]follows. Again, this [ij,k]is not a third-rank tensor.FromEq. (2.131b), [ij,k]=gmkεm·∂εi ∂qj =εk·∂εi ∂qj. (2.135) Nowwedifferentiate gij=εi·εj, Eq. (2.124): ∂gij ∂qk=∂εi ∂qk·εj+εi·∂εj ∂qk =[ik,j]+[jk,i] (2.136) byEq.(2.135). Then [ij,k]=1 2braceleftbigg∂gik ∂qj+∂gjk ∂qi−∂gij ∂qkbracerightbigg , (2.137) and Ŵs ij=gks[ij,k] =1 2gksbraceleftbigg∂gik ∂qj+∂gjk ∂qi−∂gij ∂qkbracerightbigg . (2.138) 156 Chapter 2 Vector Analysis in Curved Coordinates and Tensors TheseChristoffel symbolsareappliedinthenextsection. Covariant Derivative WiththeChristoffelsymbols,Eq. (2.130b)mayberewritten ∂V′ ∂qj=∂Vi ∂qjεi+ViŴk ijεk. (2.139) Now,iandkin the last term are dummy indices. Interchanging iandk(in this one term), wehave ∂V′ ∂qj=parenleftbigg∂Vi ∂qj+VkŴi kjparenrightbigg εi. (2.140) Thequantityinparenthesisislabeleda covariantderivative ,Vi ;j.W eha v e Vi ;j≡∂Vi ∂qj+VkŴi kj. (2.141) The;jsubscriptindicatesdifferentiationwithrespectto qj.Thedifferential dV′becomes dV′=∂V′ ∂qjdqj=[Vi ;jdqj]εi. (2.142) A comparison with Eq. (2.113) or (2.122) shows that the quantity in square brackets is theithcontravariantcomponentofavector.Since dqjisthejthcontravariantcomponent of a vector (again, Eq. (2.113)), Vi ;jmust be the ijth componentof a (mixed) second-rank tensor(quotientrule).Thecovariantderivativesofthecontravariantcomponentsofavector formamixedsecond-ranktensor, Vi ;j. Since the Christoffel symbols vanish in Cartesian coordinates, the covariant derivative andtheordinarypartialderivativecoincide: ∂Vi ∂qj=Vi ;j(Cartesiancoordinates). (2.143) Thecovariantderivativeof acovariantvector Viisgivenby(Exercise2.10.9) Vi;j=∂Vi ∂qj−VkŴk ij. (2.144) LikeVi ;j,Vi;jis asecond-ranktensor. The physical importance of the covariant derivative is that “A consistent replacement of regular partial derivatives by covariant derivatives carries the laws of physics (in com- ponent form) from flat space–time into the curved (Riemannian) space–time of general relativity. Indeed, this substitution may be taken as a mathematicalstatement of Einstein’s principleofequivalence.”21 21C. W.Misner, K.S.Thorne, andJ. A.Wheeler, Gravitation . San Francisco: W.H. Freeman (1973), p. 387. 2.10 General Tensors 157 Geodesics, Parallel Transport The covariant derivative of vectors, tensors, and the Christoffel symbols may also be ap- proached from geodesics. A geodesic in Euclidean space is a straight line. In general, it is the curve of shortest length between two points and the curve along which a freely falling particle moves. The ellipses of planets are geodesics around the sun, and the moon is in free fall around the Earth on a geodesic. Since we can throw a particle in any direction, a geodesic can have any direction through a given point. Hence the geodesic equation can beobtainedfromFermat’svariationalprincipleofoptics(seeChapter17forEuler’sequa- tion), δintegraldisplay ds=0, (2.145) whereds2is themetric,Eq.(2.123), ofourspace.Usingthevariationof ds2, 2dsδds=dqidqjδgij+gijdqiδdqj+gijdqjδdqi(2.146) inEq.(2.145) yields 1 2integraldisplaybracketleftbiggdqi dsdqj dsδgij+gijdqi dsd dsδdqj+gijdqj dsd dsδdqibracketrightbigg ds=0,(2.147) wheredsmeasuresthelengthonthegeodesic.Expressingthevariations δgij=∂gij ∂qkδdqk≡(∂kgij)δdqk in terms of the independent variations δdqk, shifting their derivatives in the other two terms of Eq. (2.147) upon integrating by parts, and renaming dummy summation indices, weobtain 1 2integraldisplaybracketleftbiggdqi dsdqj ds∂kgij−d dsparenleftbigg gikdqi ds+gkjdqj dsparenrightbiggbracketrightbigg δdqkds=0. (2.148) The integrand of Eq. (2.148), set equal to zero, is the geodesic equation. It is the Euler equationof ourvariationalproblem.Uponexpanding dgik ds=(∂jgik)dqj ds,dgkj ds=(∂igkj)dqi ds(2.149) alongthegeodesicwefind 1 2dqi dsdqj ds(∂kgij−∂jgik−∂igkj)−gikd2qi ds2=0. (2.150) MultiplyingEq. (2.150)with gklandusingEq. (2.125),wefindthe geodesic equation d2ql ds2+dqi dsdqj ds1 2gkl(∂igkj+∂jgik−∂kgij)=0, (2.151) wherethecoefficientofthevelocitiesis theChristoffelsymbol Ŵl ijofEq. (2.138). Geodesics are curves that are independent of the choice of coordinates. They can be drawnthroughanypointinspaceinvariousdirections.Sincethelength dsmeasuredalong 158 Chapter 2 Vector Analysis in Curved Coordinates and Tensors thegeodesicisascalar,thevelocities dqi/ds(ofafreelyfallingparticlealongthegeodesic, for example) form a contravariant vector. Hence Vkdqk/dsis a well-defined scalar on any geodesic, which we can differentiate in order to define the covariant derivative of any covariantvector Vk.UsingEq. (2.151)weobtainfromthescalar d dsparenleftbigg Vkdqk dsparenrightbigg =dVk dsdqk ds+Vkd2qk ds2 =∂Vk ∂qidqi dsdqk ds−VkŴk ijdqi dsdqj ds(2.152) =dqi dsdqk dsparenleftbigg∂Vk ∂qi−Ŵl ikVlparenrightbigg . Whenthequotienttheoremis appliedtoEq.(2.152) ittellsusthat Vk;i=∂Vk ∂qi−Ŵl ikVl (2.153) isacovarianttensorthatdefinesthecovariantderivativeof Vk,consistentwithEq.(2.144). Similarly,higher-ordertensorsmaybederived. ThesecondterminEq. (2.153)definesthe paralleltransportordisplacement , δVk=Ŵl kiVlδqi, (2.154) of the covariant vector Vkfrom the point with coordinates qitoqi+δqi. The parallel transport, δUk,ofacontravariantvector Ukmaybefoundfromtheinvarianceofthescalar productUkVkunderparalleltransport, δ(UkVk)=δUkVk+UkδVk=0, (2.155) inconjunctionwiththequotienttheorem. Insummary,whenweshiftavectortoaneighboringpoint,paralleltransportpreventsit fromstickingoutofourspace.Thiscanbeclearlyseenonthesurfaceofasphereinspher- ical geometry, where a tangent vector is supposed to remain a tangent upon translating it along some path on the sphere. This explains why the covariant derivative of a vector or tensorisnaturallydefinedbytranslatingit alongageodesicinthedesireddirection. Exercises 2.10.1 Equations (2.115) and (2.116) use the scale factor hi, citing Exercise 2.2.3. In Sec- tion 2.2 we had restricted ourselves to orthogonal coordinate systems, yet Eq. (2.115) holds for nonorthogonal systems. Justify the use of Eq. (2.115) for nonorthogonal sys- tems. 2.10.2 (a) Showthat εi·εj=δi j. (b) Fromtheresultof part(a) showthat Fi=F·εiandFi=F·εi. 2.10 General Tensors 159 2.10.3 For the special case of three-dimensional space ( ε1,ε2,ε3defining a right-handed co- ordinatesystem,notnecessarilyorthogonal),showthat εi=εj×εk εj×εk·εi, i,j,k=1,2, 3andcyclicpermutations. Note.These contravariant basis vectors εidefine the reciprocal lattice space of Sec- tion1.5. 2.10.4 Provethatthecontravariantmetrictensoris givenby gij=εi·εj. 2.10.5 If thecovariantvectors εiareorthogonal,showthat (a)gijisdiagonal, (b)gii=1/gii(nosummation), (c)|εi|=1/|εi|. 2.10.6 Derive the covariant and contravariant metric tensors for circular cylindrical coordi- nates. 2.10.7 Transformtheright-handsideofEq. (2.129), ∇ψ=∂ψ ∂qiεi, into theeibasis, and verify that this expression agrees with the gradient developed in Section2.2(for orthogonalcoordinates). 2.10.8 Evaluate ∂εi/∂qjfor spherical polar coordinates, and from these results calculate Ŵk ij for sphericalpolarcoordinates. Note.Exercise2.5.2 offers a wayof calculatingtheneededpartialderivatives. Remem- ber, ε1=ˆrbut ε2=rˆθand ε3=rsinθˆϕ. 2.10.9 Showthatthecovariantderivativeof acovariantvectorisgivenby Vi;j≡∂Vi ∂qj−VkŴk ij. Hint.Differentiate εi·εj=δi j. 2.10.10 Verifythat Vi;j=gikVk ;jbyshowingthat ∂Vi ∂qj−VsŴs ij=gikbraceleftbigg∂Vk ∂qj+VmŴk mjbracerightbigg . 2.10.11 From the circular cylindrical metric tensor gij, calculate the Ŵk ijfor circular cylindrical coordinates. Note.Thereareonlythreenonvanishing Ŵ. 160 Chapter 2 Vector Analysis in Curved Coordinates and Tensors 2.10.12 Using the Ŵk ijfrom Exercise 2.10.11, write out the covariant derivatives Vi ;jof a vector Vincircularcylindricalcoordinates. 2.10.13 A triclinic crystal is described using an oblique coordinate system. The three covariant basevectorsare ε1=1.5ˆx, ε2=0.4ˆx+1.6ˆy, ε3=0.2ˆx+0.3ˆy+1.0ˆz. (a) Calculatetheelementsof thecovariantmetrictensor gij. (b) Calculate the Christoffel three-index symbols, Ŵk ij. (This is a “by inspection” cal- culation.) (c) From the cross-product form of Exercise 2.10.3 calculate the contravariant base vector ε3. (d) Usingtheexplicitforms ε3andεi,verifythat ε3·εi=δ3i. Note.If it were needed, the contravariant metric tensor could be determined by finding theinverseof gijor byfindingthe εiandusing gij=εi·εj. 2.10.14 Verifythat [ij,k]=1 2braceleftbigg∂gik ∂qj+∂gjk ∂qi−∂gij ∂qkbracerightbigg . Hint.SubstituteEq.(2.135)intotheright-handsideandshowthatanidentityresults. 2.10.15 Showthatfor themetrictensor gij;k=0,gij;k=0. 2.10.16 Show that parallel displacement δdqi=d2qialong a geodesic. Construct a geodesic byparalleldisplacementof δdqi. 2.10.17 Construct the covariant derivative of a vector Viby parallel transport starting from the limitingprocedure lim dqj→0Vi(qj+dqj)−Vi(qj) dqj. 2.11 T ENSOR DERIVATIVE OPERATORS InthissectionthecovariantdifferentiationofSection2.10isappliedtorederivethevector differentialoperationsof Section2.2ingeneraltensorform. Divergence Replacingthepartialderivativebythecovariantderivative,wetakethedivergencetobe ∇·V=Vi ;i=∂Vi ∂qi+VkŴi ik. (2.156) 2.11 Tensor Derivative Operators 161 Expressing Ŵi ikbyEq.(2.138), wehave Ŵi ik=1 2gimbraceleftbigg∂gim ∂qk+∂gkm ∂qi−∂gik ∂qmbracerightbigg . (2.157) Whencontractedwith gimthelasttwotermsinthecurlybracketcancel,since gim∂gkm ∂qi=gmi∂gki ∂qm=gim∂gik ∂qm. (2.158) Then Ŵi ik=1 2gim∂gim ∂qk. (2.159) Fromthetheoryofdeterminants,Section3.1, ∂g ∂qk=ggim∂gim ∂qk, (2.160) wheregis the determinant of the metric, g=det(gij). Substituting this result into Eq.(2.158), weobtain Ŵi ik=1 2g∂g ∂qk=1 g1/2∂g1/2 ∂qk. (2.161) Thisyields ∇·V=Vi ;i=1 g1/2∂ ∂qkparenleftbig g1/2Vkparenrightbig . (2.162) To compare this result with Eq. (2.21), note that h1h2h3=g1/2andVi(contravariant coefficientof εi)=Vi/hi(no summation),where Viis Section2.2coefficientof ei. Laplacian In Section 2.2, replacement of the vector Vin∇·Vby∇ψled to the Laplacian ∇·∇ψ. Herewehaveacontravariant Vi.Usingthemetrictensortocreateacontravariant ∇ψ,we makethesubstitution Vi→gik∂ψ ∂qk. ThentheLaplacian ∇·∇ψbecomes ∇·∇ψ=1 g1/2∂ ∂qiparenleftbigg g1/2gik∂ψ ∂qkparenrightbigg . (2.163) Fortheorthogonal systemsofSection2.2themetrictensorisdiagonalandthecontravari- antgii(nosummation)becomes gii=(hi)−2. 162 Chapter 2 Vector Analysis in Curved Coordinates and Tensors Equation(2.163)reducesto ∇·∇ψ=1 h1h2h3∂ ∂qiparenleftbiggh1h2h3 h2 i∂ψ ∂qiparenrightbigg , inagreementwithEq.(2.22). Curl Thedifferenceofderivativesthatappearsinthecurl(Eq. (2.27)) willbewritten ∂Vi ∂qj−∂Vj ∂qi. Again, remember that the components Vihere are coefficients of the contravariant (nonunit)base vectors εi.T h eViof Section2.2 are coefficients of unit vectors ei. Adding andsubtracting,weobtain ∂Vi ∂qj−∂Vj ∂qi=∂Vi ∂qj−VkŴk ij−∂Vj ∂qi+VkŴk ji =Vi;j−Vj;i (2.164) usingthesymmetryoftheChristoffelsymbols.Thecharacteristicdifferenceofderivatives of the curl becomes a difference of covariant derivatives and therefore is a second-rank tensor(covariantinbothindices).AsemphasizedinSection2.9,thespecialvectorformof thecurlexistsonlyinthree-dimensionalspace. From Eq. (2.138) it is clear that all the Christoffel three index symbols vanish in Minkowskispaceandintherealspace–timeofspecialrelativitywith gλµ= 1000 0−100 00−10 000 −1 . Here x0=ct, x 1=x, x 2=y,andx3=z. Thiscompletesthedevelopmentofthedifferentialoperatorsingeneraltensorform.(The gradient was given in Section 2.10.) In addition to the fields of elasticity and electromag- netism, these differentials find application in mechanics (Lagrangian mechanics, Hamil- tonianmechanics,andtheEulerequationsforrotationofrigidbody);fluidmechanics;and perhapsmostimportantof all,thecurvedspace–timeofmoderntheoriesof gravity. Exercises 2.11.1 VerifyEq. (2.160), ∂g ∂qk=ggim∂gim ∂qk, for thespecificcaseof sphericalpolarcoordinates. 2.11 Additional Readings 163 2.11.2 Startingwiththedivergenceintensornotation,Eq.(2.162),developthedivergenceofa vectorinsphericalpolarcoordinates,Eq. (2.47). 2.11.3 Thecovariantvector Aiisthegradientofascalar.Showthatthedifferenceofcovariant derivatives Ai;j−Aj;ivanishes. AdditionalReadings Dirac,P.A.M., General Theory of Relativity .Princeton, NJ: Princeton University Press (1996). Har tle,J.B., Gravity,San Francisco: Addison-Wesley (2003). This text uses aminimum of tensor analysis. Jeffreys, H., Cartesian Tensors . Cambridge: Cambridge University Press (1952). This is an excellent discussion of Cartesian tensors andtheirapplication to awide variety offields ofclassicalphysics. La wd en ,D.F ., AnIntroduction to Tensor Calculus, Relativityand Cosmology , 3rd ed.NewYork: Wiley (1982). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nos- trand (1956). Chapter5 covers curvilinear coordinates and 13 specificcoordinate systems. Misner, C. W.,K.S. Thorne, andJ.A.Wheeler, Gravitation . San Francisco: W. H.Freeman (1973), p. 387. Moller, C., The Theory of Relativity . Oxford: Oxford University Press (1955). Reprinted (1972). Most texts on general relativity include a discussion of tensor analysis. Chapter 4 develops tensor calculus, including the topicofdualtensors.Theextensiontonon-Cartesiansystems,asrequiredbygeneralrelativity,ispresentedin Chapter 9. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 in- cludesadescriptionofseveraldifferentcoordinatesystems.NotethatMorseandFeshbacharenotaboveusing left-handedcoordinatesystemsevenforCartesiancoordinates.Elsewhereinthisexcellent(anddifficult)book therearemanyexamplesoftheuseofthevariouscoordinatesystemsinsolvingphysicalproblems.Elevenad- ditionalfascinatingbutseldom-encounteredorthogonalcoordinatesystemsarediscussedinthesecond(1970) edition of Mathematical Methods for Physicists . Ohanian, H. C., and R. Ruffini, Gravitation and Spacetime , 2nd ed. New York: Norton & Co. (1994). A well- writtenintroduction to Riemannian geometry. Sokolnikoff, I. S., Tensor Analysis—Theory and Applications , 2nd ed. New York: Wiley (1964). Particularly useful for its extension of tensor analysis to non-Euclidean geometries. Weinberg, S., Gravitation and Cosmology. Principles and Applications of the General Theory of Relativity .Ne w York: Wiley (1972). This book and the one by Misner, Thorne, and Wheeler are the two leading texts on general relativity and cosmology (withtensors in non-Cartesian space). Young, E.C., Vectorand Tensor Analysis , 2nd ed.NewYork: MarcelDekker(1993). This page intentionally left blank CHAPTER 3 DETERMINANTS AND MATRICES 3.1 D ETERMINANTS We begin the study of matrices by solving linear equations that will lead us to determi- nants and matrices. The concept of determinant and the notation were introduced by the renownedGermanmathematicianandphilosopherGottfriedWilhelmvonLeibniz. Homogeneous Linear Equations One of the major applications of determinants is in the establishment of a condition for the existence of a nontrivial solution for a set of linear homogeneous algebraic equations. Supposewehavethreeunknowns x1,x2,x3(ornequationswith nunknowns): a1x1+a2x2+a3x3=0, b1x1+b2x2+b3x3=0, (3.1) c1x1+c2x2+c3x3=0. The problem is to determine under what conditions there is any solution, apart from the trivial one x1=0,x2=0,x3=0. If we use vector notation x=(x1,x2,x3)for the solution and three rows a=(a1,a2,a3),b=(b1,b2,b3),c=(c1,c2,c3)of coefficients, thenthethreeequations,Eqs. (3.1), become a·x=0,b·x=0,c·x=0. (3.2) These three vector equations have the geometrical interpretation that xis orthogonalto a,b, andc. If the volume spanned by a,b,cgiven by the determinant (or triple scalar 165 166 Chapter 3 Determinants and Matrices product,seeEq. (1.50) ofSection1.5) D3=(a×b)·c=det(a,b,c)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3 b1b2b3 c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(3.3) isnotzero,thenthereis onlythetrivialsolution x=0. Conversely, if the aforementioned determinant of coefficients vanishes, then one of the row vectors is a linear combination of the other two. Let us assume that clies in the plane spanned by aandb, that is, that the third equation is a linear combination of the first two and not independent. Then xis orthogonal to that plane so that x∼a×b. Since homogeneous equations can be multiplied by arbitrary numbers, only ratios of the xiare relevant,forwhichwethenobtainratiosof 2 ×2 determinants x1 x3=a2b3−a3b2 a1b2−a2b1 x2 x3=−a1b3−a3b1 a1b2−a2b1(3.4) from the components of the cross product a×b, provided x3∼a1b2−a2b1/negationslash=0. This is Cramer’srule for threehomogeneouslinearequations. Inhomogeneous Linear Equations Thesimplestcaseof twoequationswithtwounknowns, a1x1+a2x2=a3,b 1x1+b2x2=b3, (3.5) can be reduced to the previous case by imbedding it in three-dimensional space with a so- lutionvector x=(x1,x2,−1)androwvectors a=(a1,a2,a3),b=(b1,b2,b3).Asbefore, Eqs.(3.5)invectornotation, a·x=0 andb·x=0,implythat x∼a×b,sotheanalogof Eqs. (3.4) holds. For this to apply, though, the third component of a×bmust not be zero, that is,a1b2−a2b1/negationslash=0, because the third component of xis−1/negationslash=0. This yields the xi as (3.6a) x1=a3b2−b3a2 a1b2−a2b1=vextendsinglevextendsinglevextendsinglevextendsinglea3a2 b3b2vextendsinglevextendsinglevextendsinglevextendsingle vextendsinglevextendsinglevextendsinglevextendsinglea1a2 b1b2vextendsinglevextendsinglevextendsinglevextendsingle, x2=a1b3−a3b1 a1b2−a2b1=vextendsinglevextendsinglevextendsinglevextendsinglea1a3 b1b3vextendsinglevextendsinglevextendsinglevextendsingle vextendsinglevextendsinglevextendsinglevextendsinglea1a2 b1b2vextendsinglevextendsinglevextendsinglevextendsingle. (3.6b) The determinant in the numerator of x1(x2)is obtained from the determinant of the co- efficientsvextendsinglevextendsinglea1a2 b2b2vextendsinglevextendsingleby replacing the first (second) column vector by the vectorparenleftbiga3 b3parenrightbig of the inhomogeneous side of Eq. (3.5). This is Cramer’s rule for a set of two inhomogeneous linearequationswithtwounknowns. 3.1 Determinants 167 These solutions of linear equations in terms of determinants can be generalized to n dimensions.Thedeterminantis asquarearray Dn=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2···an b1b2···bn c1c2···cn · · ··· ·vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle(3.7) of numbers (or functions), the coefficients of nlinear equations in our case here. The numbernof columns (and of rows) in the array is sometimes called the orderof the determinant. The generalization of the expansion in Eq. (1.48) of the triple scalar product (of row vectors of three linear equations) leads to the following value of the determinant Dninndimensions, Dn=summationdisplay i,j,k,...εijk···aibjck···, (3.8) whereεijk···, analogous to the Levi-Civita symbol of Section 2.9, is +1 for even permuta- tions1(ijk···)of(123···n),−1 for oddpermutations,andzeroif anyindexisrepeated. Specifically,for thethird-orderdeterminant D3of Eq. (3.3), Eq.(3.8) leadsto D3=+a1b2c3−a1b3c2−a2b1c3+a2b3c1+a3b1c2−a3b2c1.(3.9) Thethird-orderdeterminant,then,isthisparticularlinearcombinationofproducts.Each product contains one and only one element from each row and from each column. Each product is added if the columns (indices) represent an even permutation of (123) and sub- tracted if we have an odd permutation. Equation (3.3) may be considered shorthand no- tation for Eq. (3.9). The number of terms in the sum (Eq. (3.8)) is 24 for a fourth-order determinant, n!for annth-order determinant. Because of the appearance of the negative signsinEq.(3.9)(andpossiblyintheindividualelementsaswell),theremaybeconsider- able cancellation. It is quite possible that a determinant of large elements will have a very smallvalue. Several useful properties of the nth-order determinants follow from Eq. (3.8). Again, to bespecific,Eq. (3.9) forthird-orderdeterminantsis usedtoillustratetheseproperties. Laplacian Development by Minors Equation(3.9) maybewritten D3=a1(b2c3−b3c2)−a2(b1c3−b3c1)+a3(b1c2−b2c1) =a1vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb2b3 c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−a2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb1b3 c1c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle+a3vextendsinglevextendsinglevextendsinglevextendsinglevextendsingleb1b2 c1c2vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.10) In general, the nth-order determinant may be expanded as a linear combination of the productsoftheelementsofanyrow(oranycolumn)andthe (n−1)th-orderdeterminants 1In a linear sequence abcd···, any single, simple transposition of adjacent elements yields an oddpermutation of the original sequence: abcd→bacd.Twosuchtranspositionsyieldanevenpermutation.Ingeneral,anoddnumberofsuchinterchangesof adjacentelements results in anodd permutation; an even number of suchtranspositions yields an even permutation. 168 Chapter 3 Determinants and Matrices formedbystrikingouttherowandcolumnoftheoriginaldeterminantinwhichtheelement appears.Thisreducedarray(2 ×2inthisspecificexample)iscalleda minor.Iftheelement is in theith row and the jth column, the sign associated with the product is (−1)i+j.T h e minorwiththissigniscalledthe cofactor.IfMijisusedtodesignatetheminorformedby omitting the ith row and the jth column and Cijis the corresponding cofactor, Eq. (3.10) becomes D3=3summationdisplay j=1(−1)j+1ajM1j=3summationdisplay j=1ajC1j. (3.11) In this case, expanding along the first row, we have i=1 and the summation over j,t h e columns. This Laplace expansion may be used to advantage in the evaluation of high-order de- terminants in which a lot of the elements are zero. For example, to find the value of the determinant D=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0100 −10 0 0 0001 00−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.12) weexpandacrossthetoprowtoobtain D=(−1)1+2·(1)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−100 00 1 0−10vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.13) Again,expandingacross thetoprow, weget D=(−1)·(−1)1+1·(−1)vextendsinglevextendsinglevextendsinglevextendsingle01 −10vextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingle01 −10vextendsinglevextendsinglevextendsinglevextendsingle=1. (3.14) (This determinant D(Eq. (3.12)) is formed from one of the Dirac matrices appearing in Dirac’srelativisticelectrontheoryinSection3.4.) Antisymmetry The determinant changes sign if any two rows are interchanged or if any two columns are interchanged. This follows from the even–odd character of the Levi-Civita εin Eq. (3.8) orexplicitlyfrom theform ofEqs. (3.9) and(3.10).2 ThispropertywasusedinSection2.9todevelopatotallyantisymmetriclinearcombina- tion.Itisalsofrequentlyusedinquantummechanicsintheconstructionofamany-particle wavefunctionthat,inaccordancewiththePauliexclusionprinciple,willbeantisymmetric under the interchange of any two identical spin1 2particles (electrons, protons, neutrons, etc.). 2The sign reversal is reasonably obvious for the interchange of two adjacent rows (or columns), this clearly being an odd permutation. Show that theinterchange of anytworows is still an odd permutation. 3.1 Determinants 169 •As a special case of antisymmetry, any determinant with two rows equal or two columnsequalequalszero. •If each element in a row or each element in a column is zero, the determinant is equal tozero. •If each element in a row or each element in a column is multiplied by a constant, the determinantis multipliedbythatconstant. •The value of a determinant is unchanged if a multiple of one row is added (column by column)toanotherroworifamultipleofonecolumnisadded(rowbyrow)toanother column.3 We have vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3 b1b2b3 c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1+ka2a2a3 b1+kb2b2b3 c1+kc2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (3.15) UsingtheLaplacedevelopmentontheright-handside, weobtain vextendsingle vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1+ka2a2a3 b1+kb2b2b3 c1+kc2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3 b1b2b3 c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle+kvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea2a2a3 b2b2b3 c2c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.16) then by the property of antisymmetry the second determinant on the right-hand side of Eq.(3.16) vanishes,verifyingEq. (3.15). As a specialcase, adeterminantis equalto zeroif anytwo rows are proportionalor any twocolumnsareproportional. Some useful relations involving determinants or matrices appear in Exercises of Sec- tions3.2and3.4. Returning to the homogeneous Eqs. (3.1) and multiplying the determinant of the coef- ficients by x1, then adding x2times the second column and x3times the third column, we candirectlyestablishtheconditionfor thepresenceofa nontrivialsolutionforEqs. (3.1): x1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3 b1b2b3 c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1x1a2a3 b1x1b2b3 c1x1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1x1+a2x2+a3x3a2a3 b1x1+b2x2+b3x3b2b3 c1x1+c2x2+c3x3c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle =vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0a2a3 0b2b3 0c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.17) Therefore x1(andx2andx3) must be zero unless the determinant of the coefficients vanishes.Conversely(seetextbelowEq.(3.3)),wecanshowthatifthedeterminantofthe coefficientsvanishes,a nontrivialsolutiondoesindeedexist. Thisis usedin Section9.6to establishthelineardependenceor independenceof asetoffunctions. 3Thisderivesfromthegeometricmeaningofthedeterminantasthevolumeoftheparallelepipedspannedbyitscolumnvectors. Pulling it to the side without changing its height leavesthe volume unchanged. 170 Chapter 3 Determinants and Matrices If our linear equations are inhomogeneous , that is, as in Eqs. (3.5) if the zeros on the right-hand side of Eqs. (3.1) are replaced by a4,b4, andc4, respectively, then from Eq.(3.17) weobtain,instead, x1=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea4a2a3 b4b2b3 c4c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1a2a3 b1b2b3 c1c2c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle, (3.18) whichgeneralizesEq.(3.6a)to n=3dimensions,etc.Ifthedeterminantofthecoefficients vanishes,theinhomogeneoussetofequationshasnosolution—unlessthenumeratorsalso vanish. In this case solutions may exist but they are not unique (see Exercise 3.1.3 for aspecificexample). Fornumericalwork,thisdeterminantsolution,Eq.(3.18),isexceedinglyunwieldy.The determinant may involve large numbers with alternate signs, and in the subtraction of two large numbers the relative error may soar to a point that makes the result worthless. Also, although the determinant method is illustrated here with three equations and three un- knowns, we might easily have 200 equations with 200 unknowns, which, involving up to 200! terms in each determinant, pose a challenge even to high-speed computers. There mustbea betterway. In fact, there are better ways. One of the best is a straightforward process often called Gausselimination .Toillustratethis technique,considerthefollowingsetofequations. Example 3.1.1 GAUSS ELIMINATION Solve 3x+2y+z=11 2x+3y+z=13 (3.19) x+y+4z=12. Thedeterminantoftheinhomogeneouslinearequations(3.19) is18,so asolutionexists. For convenience and for the optimum numerical accuracy, the equations are rearranged sothatthelargestcoefficientsrunalongthemaindiagonal(upperlefttolowerright).This hasalreadybeendoneintheprecedingset. The Gauss technique is to use the first equation to eliminate the first unknown, x, from the remaining equations. Then the (new) second equation is used to eliminate yfrom the last equation. In general, we work down through the set of equations, and then, with one unknowndetermined,weworkbackuptosolveforeachoftheotherunknownsinsucces- sion. Dividingeachrowbyits initialcoefficient,weseethatEqs. (3.19) become x+2 3y+1 3z=11 3 x+3 2y+1 2z=13 2(3.20) x+y+4z=12. 3.1 Determinants 171 Now,usingthefirst equation,weeliminate xfromthesecondandthirdequations: x+2 3y+1 3z=11 3 5 6y+1 6z=17 6(3.21) 1 3y+11 3z=25 3 and x+2 3y+1 3z=11 3 y+1 5z=17 5(3.22) y+11z=25. Repeating the technique, we use the new second equation to eliminate yfrom the third equation: x+2 3y+1 3z=11 3 y+1 5z=17 5(3.23) 54z=108, or z=2. Finally,workingbackup, weget y+1 5×2=17 5, or y=3. Thenwith zandydetermined, x+2 3×3+1 3×2=11 3, and x=1. The technique may not seem so elegant as Eq. (3.18), but it is well adapted to computers andis far faster thanthetimespentwithdeterminants. This Gausstechniquemaybeusedtoconvertadeterminantintotriangularform: D=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea1b1c1 0b2c2 00c3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle forathird-orderdeterminantwhoseelementsarenottobeconfusedwiththoseinEq.(3.3). In this form D=a1b2c3. For annth-order determinant the evaluation of the triangular form requires only n−1 multiplications, compared with the n!required for the general case. 172 Chapter 3 Determinants and Matrices A variation of this progressive elimination is known as Gauss–Jordan elimination. We startaswiththeprecedingGausselimination,buteachnewequationconsideredisusedto eliminate a variable from allthe other equations, not just those below it. If we had used thisGauss–Jordanelimination,Eq.(3.23) wouldbecome x+1 5z=7 5 y+1 5z=17 5(3.24) z=2, using the second equation of Eqs. (3.22) to eliminate yfrom both the first and third equa- tions.ThenthethirdequationofEqs.(3.24)isusedtoeliminate zfromthefirstandsecond, giving x=1 y=3 (3.25) z=2. WereturntothisGauss–JordantechniqueinSection3.2for invertingmatrices. Another technique suitable for computer use is the Gauss–Seidel iteration technique. Eachtechniquehasits advantagesanddisadvantages.TheGauss andGauss–Jordanmeth- ods may have accuracy problems for large determinants. This is also a problem for ma- trix inversion (Section 3.2). The Gauss–Seidel method, as an iterative method, may have convergence problems. The IBM Scientific Subroutine Package (SSP) uses Gauss and Gauss–Jordan techniques. The Gauss–Seidel iterative method and the Gauss and Gauss– Jordan elimination methods are discussed in considerable detail by Ralston and Wilf and also by Pennington.4Computer codes in FORTRAN and other programming languages and extensive literature for the Gauss–Jordan elimination and others are also given by Pressetal.5/squaresolid Linear Dependence of Vectors Twononzerotwo-dimensionalvectors a1=parenleftbigga11 a12parenrightbigg /negationslash=0,a2=parenleftbigga21 a22parenrightbigg /negationslash=0 aredefinedtobe linearlydependent if twonumbers x1,x2canbefoundthatarenotboth zero so that the linear relation x1a1+x2a2=0 holds. They are linearly independent if x1=0=x2istheonlysolutionofthislinearrelation.WritingitinCartesiancomponents, weobtaintwo homogeneouslinearequations a11x1+a21x2=0,a 12x1+a22x2=0 4A. Ralston and H. Wilf, eds., Mathematical Methods for Digital Computers . New York: Wiley (1960); R. H. Pennington, Introductory Computer Methods and Numerical Analysis . NewYork: Macmillan(1970). 5W. H. Press, B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes , 2nd ed. Cambridge, UK: Cambridge University Press (1992), Chapter2. 3.1 Determinants 173 fromwhichweextractthefollowingcriterionforlinearindependenceoftwovectorsusing Cramer’s rule. If a1,a2span a nonzero area , that is, their determinantvextendsinglevextendsinglea11a21a12a22vextendsinglevextendsingle/negationslash=0, then the set of homogeneous linear equations has only the solution x1=0=x2.If the determinant is zero ,then there is a nontrivial solution x1,x2, andour vectors are linearly dependent . In particular, the unit vectors in the x- andy-directions are linearly independent, the linear relation x1ˆx1+x2ˆx2=parenleftbigx1 x2parenrightbig =parenleftbig0 0parenrightbig having only the trivial solution x1=0=x2. Three or more vectors in two-dimensional space are always linearly dependent. Thus, the maximum number of linearly independent vectors in two-dimensional space is 2. For example,given a1,a2,a3,thelinearrelation x1a1+x2a2+x3a3=0alwayshasnontrivial solutions.Ifoneofthevectorsiszero,lineardependenceisobviousbecausethecoefficient ofthezerovectormaybechosentobenonzeroandthatoftheothersaszero.Soweassume allofthemas nonzero.If a1anda2arelinearlyindependent,wewritethelinearrelation a11x1+a21x2=−a31x3,a 12x1+a22x2=−a32x3, asasetoftwoinhomogeneouslinearequationsandapplyCramer’srule.Sincethedetermi- nantisnonzero,wecanfindanontrivialsolution x1,x2foranynonzero x3.Thisargument goes through for any pair of linearly independent vectors. If all pairs are linearly depen- dent, any of these linear relations is a linear relation among the three vectors, and we are finished.Iftherearemorethanthreevectors,wepickanythreeofthemandapplythefore- goingreasoningandputthecoefficientsoftheothervectors, xj=0,inthelinearrelation. •Mutuallyorthogonalvectorsarelinearlyindependent. Assume a linear relationsummationtext icivi=0.Dottingvjinto this using vj·vi=0f o rj/negationslash=i,w e obtaincjvj·vj=0,soevery cj=0 because v2 j/negationslash=0. It is straightforward to extend these theorems to nor more vectors in n-dimensional Euclidean space. Thus, the maximum number of linearly independent vectors in n-dimensional space is n. The coordinate unit vectors are linearly independent be- cause they span a nonzero parallelepiped in n-dimensional space and their determinant isunity. Gram–Schmidt Procedure Inann-dimensionalvectorspacewithaninner(orscalar)product,wecanalwaysconstruct anorthonormalbasisof nvectorswiwithwi·wj=δijstartingfrom nlinearlyindependent vectorsvi,i=0,1,...,n−1. We start by normalizing v0to unity, defining w0=v0 √v02. Then we project v0fromv1, formingu1=v1+a10w0, with the admixture coefficient a10chosen so that v0·u1=0. Dottingv0intou1yieldsa10=−v0·v1radicalBig v2 0=−v1·w0.Again,wenormalize u1definingw1= u1radicalBig u2 1. Here,u2 1/negationslash=0 because v0,v1arelinearlyindependent.Thisfirst stepgeneralizesto uj=vj+aj0w0+aj1w1+···+ajj−1wj−1, withcoefficients aji=−vj·wi. Normalizing wj=ujradicalBig u2 jcompletesourconstruction. 174 Chapter 3 Determinants and Matrices It will be noticed that although this Gram–Schmidt procedure is one possible way of constructing an orthogonal or orthonormal set, the vectors wiare not unique. There is an infinitenumberofpossibleorthonormalsets. As an illustration of the freedom involved, consider two (nonparallel) vectors AandB in thexy-plane. We may normalize Ato unit magnitude and then form B′=aA+Bso thatB′is perpendicular to A. By normalizing B′we have completed the Gram–Schmidt orthogonalizationfortwovectors.Butanytwoperpendicularunitvectors,suchas ˆxandˆy, couldhavebeenchosenasourorthonormalset.Again,withaninfinitenumberofpossible rotations ofˆxandˆyabout the z-axis, we have an infinite number of possible orthonormal sets. Example 3.1.2 VECTORS BY GRAM–SCHMIDT ORTHOGONALIZATION Toillustratethemethod,weconsidertwovectors v0=parenleftbigg1 1parenrightbigg ,v1=parenleftbigg1 −2parenrightbigg , which are neither orthogonal nor normalized. Normalizing the first vector w0=v0/√ 2, wethenconstruct u1=v1+a10w0so astobeorthogonalto v0. Thisyields u1·v0=0=v1·v0+a10√ 2v2 0=−1+a10√ 2, sotheadjustableadmixturecoefficient a10=1/√ 2. Asaresult, u1=parenleftbigg1 −2parenrightbigg +1 2parenleftbigg1 1parenrightbigg =3 2parenleftbigg1 −1parenrightbigg , sothesecondorthonormalvectorbecomes w1=1√ 2parenleftbigg1 −1parenrightbigg . We check that w0·w1=0. The two vectors w0,w1form an orthonormal set of vectors, abasisof two-dimensionalEuclideanspace. /squaresolid Exercises 3.1.1 Evaluatethefollowingdeterminants: (a)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle101 010 100vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,(b)vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle120 312 031vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle,(c)1√ 2vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle0√ 30 0√ 30 2 0 020√ 3 00√ 30vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. 3.1.2 Test theset oflinearhomogeneousequations x+3y+3z=0,x−y+z=0,2x+y+3z=0 toseeif itpossessesa nontrivialsolution,andfindone. 3.1 Determinants 175 3.1.3 Giventhepairofequations x+2y=3,2x+4y=6, (a) Showthatthedeterminantofthecoefficientsvanishes. (b) Showthatthenumeratordeterminants(Eq. (3.18)) alsovanish. (c) Findatleasttwosolutions. 3.1.4 Expressthe components ofA×Bas2×2determinants.Thenshowthatthedotproduct A·(A×B)yields a Laplacian expansionof a 3 ×3 determinant. Finally, note that two rowsof the 3×3 determinantareidenticalandhencethat A·(A×B)=0. 3.1.5 IfCijisthecofactorofelement aij(formedbystrikingoutthe ithrowand jthcolumn andincludingasign (−1)i+j),showthat (a)summationtext iaijCij=summationtext iajiCji=|A|,where|A|isthedeterminantwiththeelements aij, (b)summationtext iaijCik=summationtext iajiCki=0,j/negationslash=k. 3.1.6 A determinant with all elements of order unity may be surprisingly small. The Hilbert determinant Hij=(i+j−1)−1,i , j=1,2,...,nis notoriousfor itssmallvalues. (a) CalculatethevalueoftheHilbertdeterminantsoforder nforn=1,2,and3. (b) If an appropriate subroutine is available, find the Hilbert determinants of order n forn=4,5,and6. ANS.nDet(Hn) 11. 28.33333×10−2 34.62963×10−4 41.65344×10−7 53.74930×10−12 65.36730×10−18 3.1.7 Solvethefollowingsetoflinearsimultaneousequations.Givetheresultstofivedecimal places. 1.0x1+0.9x2+0.8x3+0.4x4+0.1x5=1.0 0.9x1+1.0x2+0.8x3+0.5x4+0.2x5+0.1x6=0.9 0.8x1+0.8x2+1.0x3+0.7x4+0.4x5+0.2x6=0.8 0.4x1+0.5x2+0.7x3+1.0x4+0.6x5+0.3x6=0.7 0.1x1+0.2x2+0.4x3+0.6x4+1.0x5+0.5x6=0.6 0.1x2+0.2x3+0.3x4+0.5x5+1.0x6=0.5. Note.Theseequationsmayalsobesolvedbymatrixinversion,Section3.2. 176 Chapter 3 Determinants and Matrices 3.1.8 Solve the linear equations a·x=c,a×x+b=0f o rx=(x1,x2,x3)with constant vectorsa/negationslash=0,bandconstant c. ANS.x=c a2a+(a×b)/a2. 3.1.9 Solvethelinearequations a·x=d,b·x=e,c·x=f,forx=(x1,x2,x3)withconstant vectorsa,b,candconstants d,e,fsuchthat (a×b)·c/negationslash=0. ANS.[(a×b)·c]x=d(b×c)+e(c×a)+f(a×b). 3.1.10 Expressinvectorformthesolution (x1,x2,x3)ofax1+bx2+cx3+d=0withconstant vectorsa,b,c,dso that(a×b)·c/negationslash=0. 3.2 M ATRICES Matrix analysis belongs to linear algebra because matrices are linear operators or maps such as rotations. Suppose, for instance, we rotate the Cartesian coordinates of a two- dimensionalspace,asinSection1.2, sothat,invectornotation, parenleftbiggx′ 1 x′ 2parenrightbigg =parenleftbiggx1cosϕ+x2sinϕ −x2sinϕ+x2cosϕparenrightbigg =parenleftbiggsummationtext ja1jxjsummationtext ja2jxjparenrightbigg . (3.26) We label the array of elementsparenleftbiga11a12a21a22parenrightbig a2×2m a t r i xAconsisting of two rows and two columns and consider the vectors x,x′as 2×1 matrices. We take the summation of products in Eq. (3.26) as a definition of matrix multiplication involving the scalar product of each row vector of Awith the column vector x. Thus, in matrix notation Eq.(3.26) becomes x′=Ax. (3.27) Toextendthisdefinitionofmultiplicationofamatrixtimesacolumnvectortotheprod- uctoftwo2×2matrices,letthecoordinaterotationbefollowedbyasecondrotationgiven bymatrixBsuchthat x′′=Bx′. (3.28) Incomponentform, x′′ i=summationdisplay jbijx′ j=summationdisplay jbijsummationdisplay kajkxk=summationdisplay kparenleftbiggsummationdisplay jbijajkparenrightbigg xk. (3.29) Thesummationover jismatrixmultiplicationdefiningamatrix C=BAsuchthat x′′ i=summationdisplay kcikxk, (3.30) orx′′=Cxin matrix notation. Again, this definition involves the scalar products of row vectorsof Bwithcolumnvectorsof A.Thisdefinitionofmatrixmultiplicationgeneralizes tom×nmatrices and is found useful; indeed, this usefulness is the justification for its existence .Thegeometricalinterpretationisthatthematrixproductofthetwomatrices BA is the rotation that carries the unprimed system directly into the double-primed coordinate 3.2 Matrices 177 system. Before passing to formal definitions, the your should note that operator Ais de- scribedbyitseffectonthecoordinatesorbasisvectors.Thematrixelements aijconstitute arepresentation of theoperator,arepresentationthatdependsonthechoiceof abasis. The special case where a matrix has one column and nrows is called a column vector, |x/angbracketright, with components xi,i=1,2,...,n.I fAis ann×nmatrix,|x/angbracketrightann-component column vector, A|x/angbracketrightis defined as in Eqs. (3.27) and (3.26). Similarly, if a matrix has one row and ncolumns, it is called a row vector, /angbracketleftx|with components xi,i=1,2,...,n. Clearly,/angbracketleftx|resultsfrom|x/angbracketrightbyinterchangingrowsandcolumns,amatrixoperationcalled transposition , and transposition for any matrix A,˜Ais called6“Atranspose” with matrix elements (˜A)ik=Aki. Transposing a product of matrices ABreverses the order and gives ˜B˜A; similarly A|x/angbracketrighttranspose is/angbracketleftx|A. The scalar product takes the form /angbracketleftx|y/angbracketright=summationtext ixiyi (x∗ iinacomplexvectorspace).This Diracbra-ketnotation isusedinquantummechanics extensivelyandinChapter10andheresubsequently. More abstractly, we can define the dual space˜Vof linear functionals Fon a vector spaceV, whereeachlinearfunctional Fof˜Vassignsanumber F(v)sothat F(c1v1+c2v2)=c1F(v1)+c2F(v2) for any vectors v1,v2from our vector space Vand numbers c1,c2. If we define the sum oftwofunctionalsbylinearityas (F1+F2)(v)=F1(v)+F2(v), then˜Vis alinearspacebyconstruction. Riesz’theorem saysthatthereisaone-to-onecorrespondencebetweenlinearfunction- alsFin˜Vand vectors fin a vector space Vthat has an inner (or scalar) product /angbracketleftf|v/angbracketright definedforanypairofvectors f,v. The proof relies on the scalar product by defining a linear functional Ffor any vector f ofVasF(v)=/angbracketleftf|v/angbracketrightforanyvofV.Thelinearityofthescalarproductin fshowsthatthese functionals form a vector space (contained in ˜Vnecessarily). Note that a linear functional iscompletelyspecifiedwhenitis definedfor everyvector vof agivenvectorspace. On the other hand, starting from any nontrivial linear functional Fof˜Vwe now con- struct a unique vector fofVso thatF(v)=f·vis given by an inner product. We start from an orthonormal basis wiof vectors in Vusing the Gram–Schmidt procedure (see Section 3.2). Take any vector vfromVand expand it as v=summationtext iwi·vwi. Then the linear functional F(v)=summationtext iwi·vF(wi)is well defined on V. If we define the spe- cific vector f=summationtext iF(wi)wi, then its inner product with an arbitrary vector vis given by/angbracketleftf|v/angbracketright=f·v=summationtext iF(wi)wi·v=F(v),whichprovesRiesz’ theorem. Basic Definitions A matrix is defined as a square or rectangular array of numbers or functions that obeys certain laws. This is a perfectly logical extension of familiar mathematical concepts. In arithmeticwedealwithsinglenumbers.Inthetheoryofcomplexvariables(Chapter6)we dealwithorderedpairsofnumbers, (1,2)=1+2i,inwhichtheorderingisimportant.We 6Some texts (including ours sometimes) denote Atranspose by AT. 178 Chapter 3 Determinants and Matrices now consider numbers (or functions) ordered in a square or rectangular array. For conve- nience in later work the numbers are distinguished by two subscripts, the first indicating the row (horizontal) and the second indicating the column (vertical) in which the number appears. For instance, a13is the matrix element in the first row, third column. Hence, if A isamatrixwith mrowsand ncolumns, A= a11a12···a1n a21a22···a2n ··· ··· · am1am2···amn . (3.31) Perhaps the most important fact to note is that the elements aijare not combined with one another. A matrix is not a determinant. It is an ordered array of numbers, not a single number. Thematrix A,sofarjustanarrayofnumbers,hasthepropertiesweassigntoit.Literally, this means constructing a new form of mathematics. We define that matrices A,B, andC, withelements aij,bij,andcij, respectively,combineaccordingtothefollowingrules. Rank LookingbackatthehomogeneouslinearEqs.(3.1),wenotethatthematrixofcoefficients, A, is made up of three row vectors that each represent one linear equation of the set. If their triple scalar product is not zero, than they span a nonzero volume and are linearly independent, and the homogeneous linear equations have only the trivial solution. In this case the matrix is said to have rank3. Inndimensions the volume represented by the triple scalar product becomes the determinant, det (A),for a square matrix. If det (A)/negationslash=0, then×nmatrixAhasrankn. The case of Eqs. (3.1), where the vector clies in the plane spanned by aandb, corresponds to rank 2 of the matrix of coefficients, because only two of its row vectors ( a,bcorresponding to two equations) are independent. In general, the rankrof a matrix is the maximal number of linearly independent row or column vectorsithas,with 0≤r≤n. Equality MatrixA=MatrixBif and only if aij=bijfor all values of iandj. This, of course, requiresthat AandBeachbem×narrays(mrows,ncolumns). Addition, Subtraction A±B=Cif and only if aij±bij=cijfor all values of iandj, the elements combining according to the laws of ordinary algebra (or arithmetic if they are simple numbers). This means that A+B=B+A, commutation. Also, an associative law is satisfied (A+B)+ C=A+(B+C). If all elements are zero, the matrix, called the null matrix , is denoted byO.ForallA, A+O=O+A=A, 3.2 Matrices 179 with O= 000··· 000··· 000··· ······ . (3.32) Suchm×nmatricesforma linearspacewithrespecttoadditionandsubtraction. Multiplication (by a Scalar) Themultiplicationof matrix Abythescalarquantity αisdefinedas αA=(αA), (3.33) inwhichtheelementsof αAareαaij;thatis,eachelementofmatrix Aismultipliedbythe scalarfactor.Thisisinstrikingcontrasttothebehaviorofdeterminantsinwhichthefactor αmultiplies only one column or one row and not every element of the entire determinant. Aconsequenceofthisscalarmultiplicationisthat αA=Aα,commutation . IfAis asquarematrix,then det(αA)=αndet(A). Matrix Multiplication, Inner Product AB=Cifandonlyif7cij=summationdisplay kaikbkj. (3.34) Theijelementof Cis formed as a scalar productof the ith row ofAwith thejth column ofB(which demands that Ahave the same number of columns ( n)a sBhas rows). The dummyindex ktakesonallvalues 1 ,2,...,ninsuccession;thatis, cij=ai1b1j+ai2b2j+ai3b3j (3.35) forn=3. Obviously, the dummy index kmay be replaced by any other symbol that is not already in use without altering Eq. (3.34). Perhaps the situation may be clarified by stating that Eq. (3.34) defines the method of combining certain matrices. This method of combination,togiveitalabel,is called matrixmultiplication .Toillustrate,considertwo (so-calledPauli)matrices σ1=parenleftbigg01 10parenrightbigg andσ3=parenleftbigg100−1parenrightbigg . (3.36) 7Some authors follow the summation convention here (compare Section 2.6). 180 Chapter 3 Determinants and Matrices The11elementoftheproduct, (σ1σ3)11isgivenbythesumoftheproductsofelementsof thefirstrowofσ1withthecorrespondingelementsof thefirst columnofσ3: parenleftBigg 01 10parenrightBigg 10 0−1 →0·1+1·0=0. Continuing,wehave σ1σ3=parenleftbigg0·1+1·00·0+1·(−1) 1·1+0·01·0+0·(−1)parenrightbigg =parenleftbigg0−1 10parenrightbigg . (3.37) Here (σ1σ3)ij=σ1i1σ31j+σ1i2σ32j. Directapplicationof thedefinitionofmatrixmultiplicationshowsthat σ3σ1=parenleftbigg01 −10parenrightbigg (3.38) andbyEq. (3.37) σ3σ1=−σ1σ3. (3.39) Exceptinspecialcases, matrixmultiplicationisnotcommutative:8 AB/negationslash=BA. (3.40) However,fromthedefinitionofmatrixmultiplicationwecanshow9thatanassociativelaw holds,(AB)C=A(BC). Thereisalsoadistributivelaw, A(B+C)=AB+AC. Theunitmatrix1haselements δij,Kroneckerdelta,andthepropertythat 1 A=A1=A forallA, 1= 1000 ··· 0100 ··· 0010 ··· 0001 ··· ······· . (3.41) It should be noted that it is possible for the product of two matrices to be the null matrix withouteitheronebeingthenullmatrix.Forexample,if A=parenleftbigg11 00parenrightbigg andB=parenleftbigg10 −10parenrightbigg , AB=O. This differs from the multiplication of real or complex numbers, which form afield, whereas the additive and multiplicative structure of matrices is called a ringby mathematicians. See also Exercise 3.2.6(a), from which it is evident that, if AB=0, at 8Commutationorthelackofitisconvenientlydescribedbythecommutatorbracketsymbol, [A,B]=AB−BA.Equation(3.40) becomes[A,B]/negationslash=0. 9Notethatthebasicdefinitionsofequality,addition,andmultiplicationaregivenintermsofthematrixelements,the aij.Allour matrix operations can be carried out in terms of the matrix elements. However, we can also treat a matrix as a single algebraic operator, as in Eq. (3.40). Matrix elements and single operators each have their advantages, as will be seen in the following section. Weshall use both approaches. 3.2 Matrices 181 leastoneof thematricesmust havea zerodeterminant(that is, be singularas definedafter Eq.(3.50) inthissection). IfAis ann×nmatrix with determinant |A|/negationslash=0, then it has a unique inverse A−1 satisfying AA−1=A−1A=1. IfBis also an n×nmatrix with inverse B−1, then the productABhastheinverse (AB)−1=B−1A−1(3.42) becauseABB−1A−1=1=B−1A−1AB(see alsoExercises3.2.31and3.2.32). Theproducttheorem ,whichsaysthatthedeterminantoftheproduct, |AB|,oftwon×n matricesAandBisequaltotheproductofthedeterminants, |A||B|,linksmatriceswithde- terminants. To prove this, consider the ncolumn vectors ck=(summationtext jaijbjk,i=1,2,...,n) of the product matrix C=ABfork=1,2,...,n. Eachck=summationtext jkbjkkajkis a sum of n column vectors ajk=(aijk,i=1,2,...,n). Note that we are now using a different prod- uct summation index jkfor each column ck. Since any determinant D(b1a1+b2a2)= b1D(a1)+b2D(a2)is linear in its column vectors, we can pull out the summation sign in front of the determinant from each column vector in Ctogether with the common column factorbjkksothat |C|=summationdisplay j′ ksbj11bj22···bjnndet(aj1aj2,...,ajn). (3.43) Ifwerearrangethecolumnvectors ajkofthedeterminantfactorinEq.(3.43)intheproper order, then we can pull the common factor det (a1,a2,...,an)=|A|in front of the nsum- mationsignsinEq.(3.43).Thesecolumnpermutationsgeneratejusttherightsign εj1j2···jn toproduceinEq.(3.43) theexpressioninEq. (3.8)for |B|so |C|=|A|summationdisplay j′ ksεj1j2···jnbj11bj22···bjnn=|A||B|, (3.44) whichprovestheproducttheorem. Direct Product A second procedure for multiplying matrices, known as the directtensor or Kronecker product,follows.If Aisanm×mmatrixand Bisann×nmatrix,thenthedirectproduct is A⊗B=C. (3.45) Cisanmn×mnmatrixwithelements Cαβ=AijBkl, (3.46) with α=m(i−1)+k, β=n(j−1)+l. 182 Chapter 3 Determinants and Matrices Forinstance,if AandBare both 2×2 matrices, A⊗B=parenleftbigga11Ba12B a21Ba22Bparenrightbigg = a11b11a11b12a12b11a12b12 a11b21a11b22a12b21a12b22 a21b11a21b12a22b11a22b12 a21b21a21b22a22b21a22b22 . (3.47) Thedirectproductisassociativebutnotcommutative.Asanexampleofthedirectprod- uct, the Dirac matrices of Section 3.4 may be developed as direct products of the Pauli matrices and the unit matrix. Other examples appear in the construction of groups (see Chapter4) andinvectororHilbertspaceinquantumtheory. Example 3.2.1 DIRECT PRODUCT OF VECTORS Thedirectproductoftwotwo-dimensionalvectorsisa four-componentvector, parenleftbiggx0 x1parenrightbigg ⊗parenleftbiggy0 y1parenrightbigg = x0y0 x0y1 x1y0 x1y1 ; whilethedirectproductof threesuchvectors, parenleftbiggx0 x1parenrightbigg ⊗parenleftbiggy0 y1parenrightbigg ⊗parenleftbiggz0 z1parenrightbigg = x0y0z0 x0y0z1 x0y1z0 x0y1z1 x1y0z0 x1y0z1 x1y1z0 x1y1z1 , isa(23=8)-dimensionalvector. /squaresolid Diagonal Matrices An important special type of matrix is the square matrix in which all the nondiagonal elementsarezero.Specifically,if a 3 ×3m a t r i xAis diagonal,then A= a1100 0a220 00 a33. Aphysicalinterpretationofsuchdiagonalmatricesandthemethodofreducingmatricesto thisdiagonalformareconsideredinSection3.5.Herewesimplynoteasignificantproperty ofdiagonalmatrices—multiplicationof diagonalmatricesiscommutative, AB=BA,ifAandBareeachdiagonal. 3.2 Matrices 183 Multiplication by a diagonal matrix [d1,d2,...,dn]that has only nonzero elements in the diagonalis particularlysimple: parenleftbigg10 02parenrightbiggparenleftbigg1234parenrightbigg =parenleftbigg12 2·32·4parenrightbigg =parenleftbigg1268parenrightbigg ; whiletheoppositeordergives parenleftbigg12 34parenrightbiggparenleftbigg1002parenrightbigg =parenleftbigg12·2 32·4parenrightbigg =parenleftbigg1438parenrightbigg . Thus,a diagonalmatrix does not commutewith anothermatrix unlessboth are diag- onal,orthediagonalmatrixisproportionaltotheunitmatrix. Thisisborneoutbythe moregeneralform [d1,d2,...,dn]A= d10···0 0d2···0 ··· ··· · 00···dn  a11a12···a1n a21a22···a2n ··· ··· · an1an2···ann  = d1a11d1a12···d1a1n d2a21d2a22···d2a2n ··· ··· · dnan1dnan2···dnann , whereas A[d1,d2,...,dn]= a11a12···a1n a21a22···a2n ··· ··· · an1an2···ann  d10···0 0d2···0 ··· ··· · 00···dn  = d1a11d2a12···dna1n d1a21d2a22···dna2n ··· ··· · d1an1d2an2···dnann . Herewehavedenotedby [d1,...,dn]adiagonalmatrixwithdiagonalelements d1,...,dn. In the special case of multiplying two diagonal matrices, we simply multiply the corre- spondingdiagonalmatrixelements,whichobviouslyiscommutative. Trace Inanysquarematrixthesumofthediagonalelementsis calledthe trace. Clearlythetraceis alinearoperation: trace(A−B)=trace(A)−trace(B). 184 Chapter 3 Determinants and Matrices One of its interesting and useful properties is that the trace of a product of two matrices A andBisindependentof theorder ofmultiplication: trace(AB)=summationdisplay i(AB)ii=summationdisplay isummationdisplay jaijbji =summationdisplay jsummationdisplay ibjiaij=summationdisplay j(BA)jj (3.48) =trace(BA). Thisholdseventhough AB/negationslash=BA.Equation(3.48)meansthatthetraceofanycommutator [A,B]=AB−BAiszero.FromEq. (3.48)weobtain trace(ABC)=trace(BCA)=trace(CAB), which shows that the trace is invariant under cyclic permutation of the matrices in a prod- uct. For a real symmetric or a complex Hermitian matrix (see Section 3.4) the trace is the sum, and the determinant the product, of its eigenvalues, and both are coefficients of the characteristic polynomial. In Exercise 3.4.23 the operation of taking the trace selects one term out of a sum of 16 terms. The trace will serve a similar function relative to matrices asorthogonalityservesfor vectorsandfunctions. Intermsoftensors(Section2.7)thetraceisacontractionand,likethecontractedsecond- ranktensor,isa scalar(invariant). Matrices are used extensively to represent the elements of groups (compare Exer- cise3.2.7andChapter4).Thetraceofthematrixrepresentingthegroupelementisknown in group theory as the character . The reason for the special name and special attention is that, the trace or character remains invariant under similarity transformations (compare Exercise3.3.9). Matrix Inversion Atthebeginningofthissectionmatrix Aisintroducedastherepresentationofanoperator that (linearly) transforms the coordinate axes. A rotation would be one example of such a linear transformation. Now we look for the inverse transformation A−1that will restore theoriginalcoordinateaxes.This means,aseitheramatrixor anoperatorequation,10 AA−1=A−1A=1. (3.49) With(A−1)ij≡a(−1) ij, a(−1) ij≡Cji |A|, (3.50) 10Hereandthroughoutthischapterourmatriceshavefiniterank.If Aisaninfinite-rankmatrix( n×nwithn→∞),thenlifeis more difficult. For A−1to be theinverse wemust demand that both AA−1=1andA−1A=1. one relation no longer implies theother. 3.2 Matrices 185 withCjithe cofactor (see discussion preceding Eq. (3.11)) of aijand the assumption that thedeterminantof A,|A|/negationslash=0.If itiszero, welabel Asingular.Noinverseexists. There is a wide variety of alternative techniques. One of the best and most commonly used is the Gauss–Jordanmatrix inversiontechnique.The theoryis based on the results of Exercises3.2.34and3.2.35,whichshowthatthereexistmatrices MLsuchthattheproduct MLAwillbeAbutwith a. onerowmultipliedbyaconstant,or b. onerowreplacedbytheoriginalrowminusamultipleof anotherrow, or c. rows interchanged. Other matrices MRoperating on the right (AMR)can carry out the same operations on thecolumns ofA. This means that the matrix rows and columns may be altered (by matrix multiplication) as though we were dealing with determinants, so we can apply the Gauss–Jordan elimina- tion techniques of Section 3.1 to the matrix elements. Hence there exists a matrix ML(or MR)suchthat11 MLA=1. (3.51) ThenML=A−1.Wedetermine MLbycarryingouttheidenticaleliminationoperationson theunitmatrix.Then ML1=ML. (3.52) Toclarifythis,weconsidera specificexample. Example 3.2.2 GAUSS–JORDAN MATRIX INVERSION Wewanttoinvertthematrix A= 321 231 114 . (3.53) For convenience we write Aand 1 side by side and carry out the identical operations on each:321 231 114 and100 010 001 . (3.54) Tobesystematic,wemultiplyeachrowtoget ak1=1,  12 31 3 13 21 2 114 and 1 300 01 20 001 . (3.55) 11Remember that det (A)/negationslash=0. 186 Chapter 3 Determinants and Matrices Subtractingthefirstrow fromthesecondandthirdrows,weobtain  12 31 3 05 61 6 01 311 3 and 1 300 −1 31 20 −1 301 . (3.56) Then we divide the second row (of bothmatrices) by5 6and subtract2 3times it from the firstrow and1 3timesitfrom thethirdrow. Theresults forbothmatricesare  101 5 011 5 0018 5 and 3 5−2 50 −2 53 50 −1 5−1 51 . (3.57) We divide the third row (of bothmatrices) by18 5. Then as the last step1 5times the third rowissubtractedfromeachof thefirst tworows(of bothmatrices).Ourfinalpairis  100 010 001 andA−1= 11 18−7 18−1 18 −7 1811 18−1 18 −1 18−1 185 18 . (3.58) The check is to multiply the original Aby the calculated A−1to see if we really do get theunitmatrix1. /squaresolid AswiththeGauss–Jordansolutionofsimultaneouslinearalgebraicequations,thistech- nique is well adapted to computers. Indeed, this Gauss–Jordan matrix inversion technique will probably be available in the program library as a subroutine (see Sections 2.3 and 2.4 ofPressetal., loc.cit.). Formatrices of special form, the inverse matrix can be given in closed form .F o r example,for A= abc bdb cbe, (3.59) theinversematrixhasa similarbutslightlymoregeneralform, A−1=αβ1γ β1δβ2 γβ2ǫ, (3.60) withmatrixelementsgivenby Dα=ed−b2,Dγ=−parenleftbig cd−b2parenrightbig ,Dβ1=(c−e)b, Dβ 2=(c−a)b, Dδ=ae−c2,Dǫ=ad−b2,D=b2(2c−a−e)+dparenleftbig ae−c2parenrightbig , whereD=det(A)isthedeterminantofthematrix A.Ife=ainA,thentheinversematrix A−1alsosimplifiesto β1=β2,ǫ=α, D=parenleftbig a2−c2parenrightbig d+2(c−a)b2. 3.2 Matrices 187 Asa check,letusworkoutthe 11-matrixelementoftheproduct AA−1=1.Wefind aα+bβ1+cγ=1 Dbracketleftbig aparenleftbig ed−b2parenrightbig +b2(c−e)−cparenleftbig cd−b2parenrightbigbracketrightbig =1 Dparenleftbig −ab2+aed+2b2c−b2e−c2dparenrightbig =D D=1. Similarlywecheckthatthe 12-matrixelementvanishes, aβ1+bδ+cβ2=1 Dbracketleftbig ab(c−e)+bparenleftbig ae−c2parenrightbig +cb(c−a)bracketrightbig =0, andso on. Note though that we cannot always find an inverse of A−1by solving for the matrix elements a,b,...ofA,becausenoteveryinversematrix A−1oftheforminEq.(3.60)has acorresponding Aof thespecialform inEq. (3.59), asExample3.2.2clearlyshows. Matricesaresquareorrectangulararraysofnumbersthatdefinelineartransformations, suchasrotationsofacoordinatesystem.Assuch,theyarelinearoperators.Squarematri- ces may be inverted when their determinant is nonzero. When a matrix defines a system of linear equations, the inverse matrix solves it. Matrices with the same number of rows and columns may be added and subtracted. They form what mathematicians call a ring with a unit and a zero matrix. Matrices are also useful for representing group operations and operatorsinHilbertspaces. Exercises 3.2.1 Showthatmatrixmultiplicationisassociative, (AB)C=A(BC). 3.2.2 Showthat (A+B)(A−B)=A2−B2 ifandonlyif AandBcommute, [A,B]=0. 3.2.3 Showthatmatrix Aisalinearoperator byshowingthat A(c1r1+c2r2)=c1Ar1+c2Ar2. It can be shown that an n×nmatrix is the most general linear operator in an n- dimensional vector space. This means that every linear operator in this n-dimensional vectorspaceisequivalenttoamatrix. 3.2.4 (a) Complex numbers, a+ib, withaandbreal, may be represented by (or are iso- morphicwith) 2 ×2 matrices: a+ib↔parenleftbiggab −baparenrightbigg . Showthatthismatrixrepresentationisvalidfor(i)additionand(ii)multiplication. (b) Findthematrixcorrespondingto (a+ib)−1. 188 Chapter 3 Determinants and Matrices 3.2.5 IfAis ann×nmatrix,showthat det(−A)=(−1)ndetA. 3.2.6 (a) The matrix equation A2=0 does not imply A=0. Show that the most general 2×2 matrixwhosesquareis zeromaybewrittenas parenleftbiggab b2 −a2−abparenrightbigg , whereaandbarerealor complexnumbers. (b) IfC=A+B, ingeneral detC/negationslash=detA+detB. Constructaspecificnumericalexampletoillustratethisinequality. 3.2.7 Giventhethreematrices A=parenleftbigg−10 0−1parenrightbigg ,B=parenleftbigg01 10parenrightbigg ,C=parenleftbigg0−1 −10parenrightbigg , findallpossibleproductsof A,B,andC,twoatatime,includingsquares.Expressyour answers in terms of A,B, andC, and 1, the unit matrix. These three matrices, together with the unit matrix, form a representation of a mathematical group, the vierergruppe (seeChapter4). 3.2.8 Given K= 00 i −i00 0−10, showthat Kn=KKK···(nfactors)=1 (withtheproperchoiceof n,n/negationslash=0). 3.2.9 Verifythe Jacobiidentity , bracketleftbig A,[B,C]bracketrightbig =bracketleftbig B,[A,C]bracketrightbig −bracketleftbig C,[A,B]bracketrightbig . This is useful in matrix descriptions of elementary particles (see Eq. (4.16)). As a mnemonic aid, the you might note that the Jacobi identity has the same form as the BAC–CABruleof Section1.5. 3.2.10 Showthatthematrices A= 010 000 000 ,B=000 001 000 ,C=001 000 000  satisfythecommutationrelations [A,B]=C,[A,C]=0,and[B,C]=0. 3.2 Matrices 189 3.2.11 Let i= 0100 −1 000 0001 00−10 ,j= 00 0−1 00−10 01 0 0 10 0 0 , and k= 00−10 00 01 10 00 0−100 . Showthat (a)i2=j2=k2=−1,where1is theunitmatrix. (b)ij=−ji=k, jk=−kj=i, ki=−ik=j. These three matrices ( i,j, andk) plus the unit matrix 1 form a basis for quaternions . An alternate basis is provided by the four 2 ×2 matrices, iσ1,iσ2,−iσ3, and 1, where theσarethePaulispinmatricesofExercise3.2.13. 3.2.12 A matrix with elements aij=0f o rj<imay be called upper right triangular. The elementsinthelowerleft(belowandtotheleftofthemaindiagonal)vanish.Examples are the matrices in Chapters 12 and 13, Exercise 13.1.21, relating power series and eigenfunctionexpansions. Showthattheproductoftwoupperrighttriangularmatricesisanupperrighttriangular matrix. 3.2.13 ThethreePaulispinmatricesare σ1=parenleftbigg01 10parenrightbigg ,σ 2=parenleftbigg0−i i0parenrightbigg ,andσ3=parenleftbigg100−1parenrightbigg . Showthat (a)(σi)2=12, (b)σjσk=iσl,(j,k,l)=(1,2,3),(2,3,1),(3,1,2)(cyclicpermutation), (c)σiσj+σjσi=2δij12;12is the 2×2 unitmatrix. ThesematriceswereusedbyPauliinthenonrelativistictheoryofelectronspin. 3.2.14 UsingthePauli σiofExercise3.2.13,showthat (σ·a)(σ·b)=a·b12+iσ·(a×b). Here σ≡ˆxσ1+ˆyσ2+ˆzσ3, aandbareordinaryvectors,and 1 2is the 2×2 unitmatrix. 190 Chapter 3 Determinants and Matrices 3.2.15 Onedescriptionof spin1particlesuses thematrices Mx=1√ 2 010 101 010 ,My=1√ 2 0−i0 i0−i 0i0, and Mz= 10 0 00 0 00−1 . Showthat (a)[Mx,My]=iMz, and so on12(cyclic permutation of indices). Using the Levi- CivitasymbolofSection2.9, wemaywrite [Mp,Mq]=iεpqrMr. (b)M2≡M2 x+M2y+M2z=213, where 1 3is the 3×3 unitmatrix. (c)[M2,Mi]=0, [Mz,L+]=L+, [L+,L−]=2Mz, where L+≡Mx+iMy, L−≡Mx−iMy. 3.2.16 RepeatExercise3.2.15usinganalternaterepresentation, Mx= 00 0 00−i 0i0 ,My=00i 00 0 −i00, and Mz= 0−i0 i00 000. InChapter4thesematricesappearasthe generators oftherotationgroup. 3.2.17 Showthatthematrix–vectorequation parenleftbigg M·∇+131 c∂ ∂tparenrightbigg ψ=0 reproduces Maxwell’s equations in vacuum. Here ψis a column vector with compo- nentsψj=Bj−iEj/c,j=x,y,z.Mis a vector whose elements are the angular momentum matrices of Exercise 3.2.16. Note that ε0µ0=1/c2,13is the 3×3 unit matrix. 12[A,B]=AB−BA. 3.2 Matrices 191 From Exercise3.2.15(b), M2ψ=2ψ. A comparison with the Dirac relativistic electron equation suggests that the “particle” of electromagnetic radiation, the photon, has zero rest mass and a spin of 1 (in units ofh). 3.2.18 RepeatExercise3.2.15,usingthematricesfor aspinof 3 /2, Mx=1 2 0√ 30 0√ 30 2 0 020√ 3 00√ 30 ,My=i 2 0−√ 30 0√ 30−20 020 −√ 3 00√ 30 , and Mz=1 2 30 0 0 01 0 0 00−10 00 0−3 . 3.2.19 Anoperator Pcommuteswith JxandJy,thexandycomponentsofanangularmomen- tum operator. Show that Pcommutes with the third component of angular momentum, thatis, that [P,Jz]=0. Hint.The angular momentum components must satisfy the commutation relation of Exercise3.2.15(a). 3.2.20 TheL+andL−matrices of Exercise 3.2.15 are ladder operators (see Chapter 4): L+ operatingonasystemofspinprojection mwillraisethespinprojectionto m+1ifmis below its maximum. L+operating on mmaxyields zero. L−reduces the spin projection inunitsteps inasimilarfashion.Dividingby√ 2,wehave L+= 010 001 000 ,L−=000 100 010 . Showthat L+|−1/angbracketright=|0/angbracketright,L−|−1/angbracketright=nullcolumnvector, L+|0/angbracketright=|1/angbracketright,L−|0/angbracketright=|−1/angbracketright, L+|1/angbracketright=nullcolumnvector ,L−|1/angbracketright=|0/angbracketright, where |−1/angbracketright= 0 0 1 ,|0/angbracketright=0 1 0 ,and|1/angbracketright=1 0 0  representstatesofspinprojection −1,0,and1,respectively. Note.Differentialoperatoranalogsof theseladderoperatorsappearinExercise12.6.7. 192 Chapter 3 Determinants and Matrices 3.2.21 VectorsAandBare relatedbythetensor T, B=TA. GivenAandB, show thatthere is no unique solution for the componentsof T.T h i si s why vector division B/Ais undefined (apart from the special case of AandBparallel andTthenascalar). 3.2.22 Wemightask foravector A−1, aninverseof agivenvector Ainthesensethat A·A−1=A−1·A=1. Show that this relation does not suffice to define A−1uniquely; Awould then have an infinitenumberofinverses. 3.2.23 IfAis a diagonal matrix, with all diagonal elements different, and AandBcommute, showthatBis diagonal. 3.2.24 IfAandBarediagonal,showthat AandBcommute. 3.2.25 Showthat trace (ABC)=trace(CBA)if anytwoofthethreematricescommute. 3.2.26 Angularmomentummatricessatisfyacommutationrelation [Mj,Mk]=iMl, j,k,l cyclic. Showthatthetraceof eachangularmomentummatrixvanishes. 3.2.27 (a) Theoperatortracereplacesamatrix Abyitstrace;thatis, trace(A)=summationdisplay iaii. Showthattraceis a linearoperator. (b) Theoperatordetreplacesamatrix Abyitsdeterminant;thatis, det(A)=determinantof A. Showthat det is notalinearoperator. 3.2.28 AandBanticommute: BA=−AB.A l s o ,A2=1,B2=1. Show that trace (A)= trace(B)=0. Note.ThePauliandDirac(Section3.4)matricesarespecificexamples. 3.2.29 With|x/angbracketrightanN-dimensional column vector and /angbracketlefty|anN-dimensional row vector, show that traceparenleftbig |x/angbracketright/angbracketlefty|parenrightbig =/angbracketlefty|x/angbracketright. Note.|x/angbracketright/angbracketlefty|means direct product of column vector |x/angbracketrightwith row vector /angbracketlefty|.T h er e s u l t isa square N×Nmatrix. 3.2.30 (a) If two nonsingular matrices anticommute, show that the trace of each one is zero. (Nonsingular meansthatthedeterminantofthematrixnonzero.) (b) Fortheconditionsofpart(a)tohold, AandBmustben×nmatriceswith neven. Showthatif nisodd,acontradictionresults. 3.2 Matrices 193 3.2.31 If amatrixhas aninverse,showthattheinverseisunique. 3.2.32 IfA−1haselements parenleftbig A−1parenrightbig ij=a(−1) ij=Cji |A|, whereCjiis thejithcofactorof|A|, showthat A−1A=1. HenceA−1istheinverseof A(if|A|/negationslash=0). 3.2.33 Showthat det A−1=(detA)−1. Hint.ApplytheproducttheoremofSection3.2. Note.If detAiszero,then Ahas noinverse. Aissingular. 3.2.34 Findthematrices MLsuchthattheproduct MLAwillbeAbutwith: (a) theithrowmultipliedbya constant k( aij→kaij,j=1,2,3,...); (b) the ith row replaced by the original ith row minus a multiple of the mth row (aij→aij−Kamj,i=1,2,3,...); (c) theithandmthrows interchanged (aij→amj,amj→aij,j=1,2,3,...). 3.2.35 Findthematrices MRsuchthattheproduct AMRwillbeAbutwith: (a) theithcolumnmultipliedbyaconstant k( aji→kaji,j=1,2,3,...); (b) the ith column replaced by the original ith column minus a multiple of the mth column(aji→aji−kajm,j=1,2,3,...); (c) theithandmthcolumnsinterchanged (aji→ajm,ajm→aji,j=1,2,3,...). 3.2.36 Findtheinverseof A= 321 221 114 . 3.2.37 (a) RewriteEq.(2.4)ofChapter2(andthecorrespondingequationsfor dyanddz)as asinglematrixequation |dxk/angbracketright=J|dqj/angbracketright. Jisa matrixofderivatives,the Jacobian matrix.Showthat /angbracketleftdxk|dxk/angbracketright=/angbracketleftdqi|G|dqj/angbracketright, withthemetric(matrix) Ghavingelements gijgivenbyEq. (2.6). (b) Showthat det(J)dq1dq2dq3=dxdydz, with det(J)theusualJacobian. 194 Chapter 3 Determinants and Matrices 3.2.38 Matricesarefartoousefultoremaintheexclusivepropertyofphysicists.Theymayap- pearwherevertherearelinearrelations.Forinstance,inastudyofpopulationmovement theinitialfractionofafixedpopulationineachof nareas(orindustriesorreligions,etc.) is representedbyan n-componentcolumnvector P. Themovementof peoplefrom one area to another in a given time is described by an n×n(stochastic) matrix T.H e r eTij is the fraction of the population in the jth area that moves to the ith area. (Those not moving are covered by i=j.) WithPdescribing the initial population distribution, the finalpopulationdistributionis givenbythematrixequation TP=Q. Fromitsdefinition,summationtextn i=1Pi=1. (a) Showthatconservationof peoplerequiresthat nsummationdisplay i=1Tij=1,j=1,2,...,n. (b) Provethat nsummationdisplay i=1Qi=1 continuestheconservationofpeople. 3.2.39 Given a 6×6m a t r i xAwith elements aij=0.5|i−j|,i=0,1,2,...,5;i=0,1, 2,...,5,findA−1. Listits matrixelementstofivedecimalplaces. ANS.A−1=1 3 4−2 0000 −25−2 000 0−25−20 0 00−25−20 000 −25−2 0000 −24 . 3.2.40 Exercise3.1.7maybewritteninmatrixform: AX=C. FindA−1andcalculate XasA−1C. 3.2.41 (a) Writea subroutine thatwillmultiply complex matrices.Assumethatthecomplex matricesareinageneralrectangularform. (b) TestyoursubroutinebymultiplyingpairsoftheDirac4 ×4matrices,Section3.4. 3.2.42 (a) Write a subroutine that will call the complex matrix multiplication subroutine of Exercise 3.2.41 and will calculate the commutator bracket of two complex matri- ces. (b) Test your complex commutator bracket subroutine with the matrices of Exer- cise3.2.16. 3.2.43 Interpolatingpolynomial isthenamegiventothe (n−1)-degreepolynomialdetermined by (and passing through) npoints,(xi,yi)with all the xidistinct. This interpolating polynomialforms abasisfor numericalquadratures. 3.3 Orthogonal Matrices 195 (a) Show that the requirement that an (n−1)-degree polynomial in xpass through each of the npoints(xi,yi)with allxidistinct leads to nsimultaneous equations oftheform n−1summationdisplay j=0ajxj i=yi,i=1,2,...,n. (b) Write a computer program that will read in ndata points and return the ncoeffi- cientsaj.Useasubroutinetosolvethesimultaneousequationsifsuchasubroutine isavailable. (c) Rewritetheset ofsimultaneousequationsas amatrixequation XA=Y. (d) Repeat the computer calculation of part (b), but this time solve for vector Aby invertingmatrix X(again,usingasubroutine). 3.2.44 Acalculationofthevaluesofelectrostaticpotentialinsideacylinderleadsto V(0.0)=52.640V(0.6)=25.844 V(0.2)=48.292V(0.8)=12.648 V(0.4)=38.270V(1.0)=0.0. The problem is to determine the values of the argument for which V=10, 20, 30, 40, and50.Express V(x)asaseriessummationtext5 n=0a2nx2n.(Symmetryrequirementsintheoriginal problem require that V(x)be an even function of x.) Determine the coefficients a2n. WithV(x)now a known function of x, find the root of V(x)−10=0, 0≤x≤1. Repeatfor V(x)−20,andso on. ANS.a0=52.640, a2=−117.676, V(0.6851)=20. 3.3 O RTHOGONAL MATRICES Ordinary three-dimensional space may be described with the Cartesian coordinates (x1,x2,x3). We consider a second set of Cartesian coordinates (x′ 1,x′ 2,x′ 3), whose ori- gin and handedness coincides with that of the first set but whose orientation is different (Fig. 3.1). We can say that the primed coordinate axeshave been rotatedrelative to the initial, unprimed coordinate axes. Since this rotation is a linearoperation, we expect a matrixequationrelatingtheprimedbasistotheunprimedbasis. This section repeats portions of Chapters 1 and 2 in a slightly different context and with a different emphasis. Previously, attention was focused on the vector or tensor. In the case of the tensor, transformation properties were strongly stressed and were critical. Here emphasis is placed on the description of the coordinate rotation itself—the matrix. Transformation properties, the behavior of the matrix when the basis is changed, appear at the end of this section. Sections 3.4 and 3.5 continue with transformation properties in complexvectorspaces. 196 Chapter 3 Determinants and Matrices FIGURE 3.1Cartesiancoordinatesystems. Direction Cosines A unit vector along the x′ 1-axis(ˆx′ 1)may be resolved into components along the x1-,x2-, andx3-axesbytheusualprojectiontechnique: ˆx′1=ˆx1cos(x′ 1,x1)+ˆx2cos(x′ 1,x2)+ˆx3cos(x′ 1,x3). (3.61) Equation (3.61) is a specific example of the linear relations discussed at the beginning of Section3.2. Forconveniencethesecosines,whicharethedirectioncosines,arelabeled cos(x′ 1,x1)=ˆx′ 1·ˆx1=a11, cos(x′ 1,x2)=ˆx′1·ˆx2=a12, (3.62a) cos(x′ 1,x3)=ˆx′1·ˆx3=a13. Continuing,wehave cos(x′ 2,x1)=ˆx′2·ˆx1=a21, (3.62b) cos(x′ 2,x2)=ˆx′2·ˆx2=a22, andso on,where a21/negationslash=a12ingeneral.Now,Eq.(3.62) mayberewritten ˆx′1=ˆx1a11+ˆx2a12+ˆx3a13, (3.62c) andalso ˆx′2=ˆx1a21+ˆx2a22+ˆx3a23, (3.62d) ˆx′3=ˆx1a31+ˆx2a32+ˆx3a33. 3.3 Orthogonal Matrices 197 We may also go the other way by resolving ˆx1,ˆx2, andˆx3into components in the primed system.Then ˆx1=ˆx′ 1a11+ˆx′2a21+ˆx′3a31, ˆx2=ˆx′1a12+ˆx′2a22+ˆx′3a32, (3.63) ˆx3=ˆx′1a13+ˆx′2a23+ˆx′3a33. Associatingˆx1andˆx′1with the subscript 1, ˆx2andˆx′2with the subscript 2, ˆx3andˆx′3 with the subscript 3, we see that in each case the first subscript of aijrefers to the primed unit vector (ˆx′1,ˆx′2,ˆx′3), whereas the second subscript refers to the unprimed unit vector (ˆx1,ˆx2,ˆx3). Applications to Vectors If weconsideravectorwhosecomponentsarefunctionsofthepositioninspace,then V(x1,x2,x3)=ˆx1V1+ˆx2V2+ˆx3V3, (3.64) V′(x′ 1,x′ 2,x′ 3)=ˆx′ 1V′ 1+ˆx′2V′ 2+ˆx′3V′ 3, since the point may be given both by the coordinates (x1,x2,x3)and by the coordinates (x′ 1,x′ 2,x′ 3).Notethat VandV′aregeometricallythesamevector(butwithdifferentcom- ponents). The coordinate axes are being rotated; the vector stays fixed. Using Eqs. (3.62) toeliminateˆx1,ˆx2, andˆx3,wemayseparateEq. (3.64)intothreescalarequations, V′ 1=a11V1+a12V2+a13V3, V′ 2=a21V1+a22V2+a23V3, (3.65) V′ 3=a31V1+a32V2+a33V3. In particular, these relations will hold for the coordinates of a point (x1,x2,x3)and (x′ 1,x′ 2,x′ 3),gi ving x′ 1=a11x1+a12x2+a13x3, x′ 2=a21x1+a22x2+a23x3, (3.66) x′ 3=a31x1+a32x2+a33x3, and similarly for the primed coordinates. In this notation the set of three equations (3.66) maybewrittenas x′ i=3summationdisplay j=1aijxj, (3.67) whereitakes onthevalues1,2, and3andtheresultisthree separate equations. Now let us set aside these results and try a different approach to the same problem. We consider two coordinate systems (x1,x2,x3)and(x′ 1,x′ 2,x′ 3)with a common origin and one point (x1,x2,x3)in the unprimed system, (x′ 1,x′ 2,x′ 3)in the primed system. Note the usual ambiguity. The same symbol xdenotes both the coordinate axis and a particular 198 Chapter 3 Determinants and Matrices distance along that axis. Since our system is linear, x′ imust be a linear combination of thexi.Let x′ i=3summationdisplay j=1aijxj. (3.68) Theaijmay be identified as the direction cosines. This identification is carried out for the two-dimensionalcaselater. Ifwehavetwosetsofquantities (V1,V2,V3)intheunprimedsystemand (V′ 1,V′ 2,V′ 3)in theprimedsystem,relatedinthesamewayasthecoordinatesofapointinthetwodifferent systems(Eq. (3.68)), V′ i=3summationdisplay j=1aijVj, (3.69) then,asinSection1.2,thequantities (V1,V2,V3)aredefinedasthecomponentsofavector that stays fixed while the coordinates rotate; that is, a vector is defined in terms of trans- formation properties of its components under a rotation of the coordinate axes. In a sense the coordinates of a point have been taken as a prototype vector. The power and useful- ness of this definition became apparent in Chapter 2, in which it was extended to define pseudovectorsandtensors. From Eq. (3.67) we can derive interesting information about the aijthat describe the orientation of coordinate system (x′ 1,x′ 2,x′ 3)relative to the system (x1,x2,x3). The length fromtheorigintothepointis thesameinbothsystems.Squaring,forconvenience,13 summationdisplay ix2 i=summationdisplay ix′2 i=summationdisplay iparenleftbiggsummationdisplay jaijxjparenrightbiggparenleftbiggsummationdisplay kaikxkparenrightbigg =summationdisplay j,kxjxksummationdisplay iaijaik. (3.70) Thiscanbetrueforallpointsifandonlyif summationdisplay iaijaik=δjk,j,k=1,2,3. (3.71) Note that Eq. (3.71) is equivalent to the matrix equation (3.83); see also Eqs. (3.87a) to(3.87d). Verification of Eq. (3.71), if needed, may be obtained by returning to Eq. (3.70) and settingr=(x1,x2,x3)=(1,0,0),(0,1,0),(0,0,1),(1,1,0), and so on to evaluate the ninerelationsgivenbyEq.(3.71).Thisprocessisvalid,sinceEq.(3.70)mustholdforall r for a given set of aij. Equation (3.71), a consequence of requiring that the length remain constant (invariant) under rotation of the coordinate system, is called the orthogonality condition .Theaij,writtenasamatrix AsubjecttoEq.(3.71),formanorthogonalmatrix, afirstdefinitionofanorthogonalmatrix.NotethatEq.(3.71)is notmatrixmultiplication. Rather,itis interpretedlateras ascalarproductof twocolumnsof A. 13Note that twoindependent indices jandkareused. 3.3 Orthogonal Matrices 199 In matrixnotationEq.(3.67) becomes |x′/angbracketright=A|x/angbracketright. (3.72) Orthogonality Conditions — Two-Dimensional Case Abetterunderstandingofthe aijandtheorthogonalityconditionmaybegainedbyconsid- ering rotation in two dimensions in detail. (This can be thought of as a three-dimensional systemwiththe x1-,x2-axesrotatedabout x3.) FromFig. 3.2, x′ 1=x1cosϕ+x2sinϕ, x′ 2=−x1sinϕ+x2cosϕ.(3.73) ThereforebyEq.(3.72) A=parenleftbiggcosϕsinϕ −sinϕcosϕparenrightbigg . (3.74) Notice that Areduces to the unit matrix for ϕ=0. Zero angle rotation means nothing has changed.Itis clearfrom Fig.3.2that a11=cosϕ=cos(x′ 1,x1), (3.75) a12=sinϕ=cosparenleftbigπ 2−ϕparenrightbig =cos(x′ 1,x2), andsoon,thusidentifyingthematrixelements aijasthedirectioncosines.Equation(3.71), theorthogonalitycondition,becomes sin2ϕ+cos2ϕ=1,(3.76)sinϕcosϕ−sinϕcosϕ=0. FIGURE 3.2Rotationof coordinates. 200 Chapter 3 Determinants and Matrices Theextensiontothreedimensions(rotationofthecoordinatesthroughanangle ϕcoun- terclockwiseabout x3)i ss i mp l y A= cosϕsinϕ0 −sinϕcosϕ0 00 1. (3.77) Thea33=1 expresses the fact that x′ 3=x3, since the rotation has been about the x3-axis. Thezerosguaranteethat x′ 1andx′ 2donotdependon x3andthatx′ 3doesnotdependon x1 andx2. Inverse Matrix, A−1 Returning to the general transformation matrix A, the inverse matrix A−1is defined such that |x/angbracketright=A−1|x′/angbracketright. (3.78) That is,A−1describes the reverse of the rotation given by Aand returns the coordinate systemtoitsoriginalposition.Symbolically,Eqs. (3.72)and(3.78)combinetogive |x/angbracketright=A−1A|x/angbracketright, (3.79) andsince|x/angbracketrightisarbitrary, A−1A=1, (3.80) theunitmatrix.Similarly, AA−1=1, (3.81) usingEqs. (3.72) and(3.78)andeliminating |x/angbracketrightinsteadof|x′/angbracketright. Transpose Matrix, ˜A We can determine the elements of our postulated inverse matrix A−1by employing the orthogonalitycondition.Equation(3.71),theorthogonalitycondition,doesnotconformto ourdefinitionofmatrixmultiplication,butitcanbeputintherequiredformby defininga newmatrix˜Asuchthat ˜aji=aij. (3.82) Equation(3.71)becomes ˜AA=1. (3.83) This is a restatement of the orthogonality condition and may be taken as the constraint defining an orthogonal matrix, a second definition of an orthogonal matrix. Multiplying Eq.(3.83) by A−1fromtherightandusingEq.(3.81), wehave ˜A=A−1, (3.84) 3.3 Orthogonal Matrices 201 a third definition of an orthogonal matrix. This important result, that the inverse equals the transpose, holds only for orthogonal matrices and indeed may be taken as a further restatementoftheorthogonalitycondition. MultiplyingEq.(3.84) by Afrom theleft, weobtain A˜A=1 (3.85) or summationdisplay iajiaki=δjk, (3.86) whichisstillanotherformoftheorthogonalitycondition. Summarizing,theorthogonalityconditionmaybestatedinseveralequivalentways: summationdisplay iaijaik=δjk, (3.87a) summationdisplay iajiaki=δjk, (3.87b) ˜AA=A˜A=1, (3.87c) ˜A=A−1. (3.87d) Anyoneof theserelationsisanecessaryandasufficientconditionfor Atobeorthogonal. It is now possible to see and understand why the term orthogonal is appropriate for thesematrices.We havethegeneralform A= a11a12a13 a21a22a23 a31a32a33, a matrix of direction cosines in which aijis the cosine of the angle between x′ iandxj. Therefore a11,a12,a13are the direction cosines of x′ 1relative to x1,x2,x3. These three elementsof Adefineaunitlengthalong x′ 1, thatis, aunitvector ˆx′ 1, ˆx′1=ˆx1a11+ˆx2a12+ˆx3a13. The orthogonality relation (Eq. (3.86)) is simply a statement that the unit vectors ˆx′1,ˆx′2, andˆx′3aremutuallyperpendicular,ororthogonal.Ourorthogonaltransformationmatrix A transforms one orthogonal coordinate system into a second orthogonal coordinate system byrotationand/orreflection. Asanexampleoftheuseofmatrices,theunitvectorsinsphericalpolarcoordinatesmay bewrittenas  ˆr ˆθ ˆϕ=Cˆx ˆy ˆz, (3.88) 202 Chapter 3 Determinants and Matrices whereCis given in Exercise 2.5.1. This is equivalent to Eqs. (3.62) with x′ 1,x′2, andx′3 replacedbyˆr,ˆθ,andˆϕ.Fromtheprecedinganalysis Cisorthogonal.Thereforetheinverse relationbecomes  ˆx ˆy ˆz=C−1ˆr ˆθ ˆϕ=˜Cˆr ˆθ ˆϕ, (3.89) andExercise2.5.5issolvedbyinspection.Similarapplicationsofmatrixinversesappearin connection with the transformation of a power series into a series of orthogonal functions (Gram–Schmidt orthogonalization in Section 10.3) and the numerical solution of integral equations. Euler Angles Our transformation matrix Acontains nine direction cosines. Clearly, only three of these are independent, Eq. (3.71) providing six constraints. Equivalently, we may say that two parameters ( θandϕin spherical polar coordinates) are required to fix the axis of rotation. Then one additional parameter describes the amount of rotation about the specified axis. (In the Lagrangian formulation of mechanics (Section 17.3) it is necessary to describe Aby using some set of three independent parameters rather than the redundant direction cosines.)TheusualchoiceofparametersistheEulerangles.14 The goal is to describe the orientation of a final rotated system (x′′′ 1,x′′′ 2,x′′′ 3)relative to some initial coordinate system (x1,x2,x3). The final system is developed in three steps, witheachstep involvingonerotationdescribedbyoneEulerangle(Fig.3.3): 1. The coordinates are rotated about the x3-axis through an angle αcounterclockwise intonewaxesdenotedby x′ 1-,x′ 2-,x′ 3.( T hex3- andx′ 3-axescoincide.) FIGURE 3.3(a)Rotationabout x3throughangle α;(b)rotationabout x′ 2through angleβ; (c)rotationabout x′′ 3throughangle γ. 14There are almost as many definitions of the Euler angles as there are authors. Here we follow the choice generally made by workers in the areaof group theory and the quantum theory of angular momentum (compare Sections 4.3, 4.4). 3.3 Orthogonal Matrices 203 2. The coordinates are rotated about the x′ 2-axis15through an angle βcounterclockwise intonewaxesdenotedby x′′ 1-,x′′ 2-,x′′ 3.( T hex′ 2- andx′′ 2-axescoincide.) 3. Thethirdandfinalrotationisthroughanangle γcounterclockwiseaboutthe x′′ 3-axis, yieldingthe x′′′ 1,x′′′ 2,x′′′ 3system.(The x′′ 3- andx′′′ 3-axescoincide.) Thethreematricesdescribingtheserotationsare Rz(α)= cosαsinα0 −sinαcosα0 00 1, (3.90) exactlylikeEq. (3.77), Ry(β)= cosβ0−sinβ 01 0 sinβ0 cosβ (3.91) and Rz(γ)= cosγsinγ0 −sinγcosγ0 00 1. (3.92) Thetotalrotationisdescribedbythetriplematrixproduct, A(α,β,γ)=Rz(γ)Ry(β)Rz(α). (3.93) Notetheorder: Rz(α)operatesfirst,then Ry(β),andfinally Rz(γ).Directmultiplication gives A(α,β,γ) =cosγcosβcosα−sinγsinαcosγcosβsinα+sinγcosα−cosγsinβ −sinγcosβcosα−cosγsinα−sinγcosβsinα+cosγcosαsinγsinβ sinβcosα sinβsinα cosβ (3.94) EquatingA(aij)withA(α,β,γ),elementbyelement,yieldsthedirectioncosinesinterms ofthethreeEulerangles.WecouldusethisEulerangleidentificationtoverifythedirection cosineidentities,Eq.(1.46)ofSection1.4,buttheapproachofExercise3.3.3ismuchmore elegant. Symmetry Properties Our matrix description leads to the rotation group SO(3)in three-dimensional space R3, and the Euler angle description of rotations forms a basis for developing the rotation group in Chapter 4. Rotations may also be described by the unitary group SU(2)in two- dimensional space C2over the complex numbers. The concept of groups such as SU(2) and its generalizations and group theoretical techniques are often encountered in modern 15Some authors choose this second rotation to be about the x′ 1-axis. 204 Chapter 3 Determinants and Matrices particle physics, where symmetry properties play an important role. The SU(2)group is alsoconsideredinChapter4.Thepowerandflexibilityofmatricespushedquaternionsinto obscurityearlyinthe20thcentury.16 Itwillbenotedthatmatriceshavebeenhandledintwowaysintheforegoingdiscussion: bytheircomponentsandassingleentities.Eachtechniquehasitsownadvantagesandboth areuseful. Thetransposematrixisusefulinadiscussionofsymmetryproperties.If A=˜A,a ij=aji, (3.95) thematrixiscalled symmetric ,whereasif A=−˜A,a ij=−aji, (3.96) it is called antisymmetric orskewsymmetric . The diagonal elements vanish. It is easy to show that any (square) matrix may be written as the sum of a symmetric matrix and an antisymmetricmatrix.Considertheidentity A=1 2[A+˜A]+1 2[A−˜A]. (3.97) [A+˜A]isclearly symmetric , whereas[A−˜A]isclearly antisymmetric .T h i si st h e matrixanalogofEq.(2.75),Chapter2,fortensors.Similarly,afunctionmaybebrokenup intoits evenandoddparts. Sofarwehaveinterpretedtheorthogonalmatrixasrotatingthecoordinatesystem.This changes the components of a fixed vector (not rotating with the coordinates) (Fig. 1.6, Chapter1).However,anorthogonalmatrix Amaybeinterpretedequallywellasarotation ofthevectorintheopposite direction(Fig.3.4). These two possibilities, (1) rotating the vector keeping the coordinates fixed and (2) rotating the coordinates (in the opposite sense) keeping the vector fixed, have a direct analogy in quantum theory. Rotation (a time transformation) of the state vector gives the Schrödingerpicture.RotationofthebasiskeepingthestatevectorfixedyieldstheHeisen- bergpicture. FIGURE 3.4Fixedcoordinates— rotatedvector. 16R. J.Stephenson, Development of vector analysis from quaternions. Am.J .Ph ys. 34: 194 (1966). 3.3 Orthogonal Matrices 205 Supposeweinterpretmatrix Aas rotatinga vectorrintothepositionshownby r1; that is, inaparticularcoordinatesystemwehavearelation r1=Ar. (3.98) Now let us rotate the coordinates by applying matrix B, which rotates (x,y,z)into (x′,y′,z′), r′ 1=Br1=BAr=(Ar)′=BAparenleftbig B−1Bparenrightbig r =parenleftbig BAB−1parenrightbig Br=parenleftbig BAB−1parenrightbig r′. (3.99) Br1is justr1in the new coordinate system, with a similar interpretation holding for Br. Henceinthisnewsystem (Br)isrotatedintoposition (Br1)bythematrix BAB−1: Br1=(BAB−1)Br r′ 1=A′r′ In the new system the coordinates have been rotated by matrix B;Ahas the form A′,i n which A′=BAB−1. (3.100) A′operatesinthe x′,y′,z′spaceasAoperatesinthe x,y,zspace. The transformation defined by Eq. (3.100) with Bany matrix, not necessarily orthogo- nal,isknownas a similaritytransformation .IncomponentformEq. (3.100)becomes a′ ij=summationdisplay k,lbikaklparenleftbig B−1parenrightbig lj. (3.101) Now,ifBisorthogonal, parenleftbig B−1parenrightbig lj=(˜B)lj=bjl, (3.102) andwehave a′ ij=summationdisplay k,lbikbjlakl. (3.103) Itmaybehelpfultothinkof Aagainasanoperator,possiblyasrotatingcoordinateaxes, relatingangularmomentumandangularvelocityofarotatingsolid(Section3.5).Matrix A istherepresentationinagivencoordinatesystem—orbasis.Buttherearedirectionsasso- ciated with A—crystal axes, symmetry axes in the rotating solid, and so on—so that the representation Adepends on the basis. The similarity transformation shows just how the representationchangeswithachangeofbasis. 206 Chapter 3 Determinants and Matrices Relation to Tensors Comparing Eq. (3.103) with the equations of Section 2.6, we see that it is the definition of a tensor of second rank. Hence a matrix that transforms by an orthogonal similarity transformation is, by definition, a tensor. Clearly, then, any orthogonal matrixA, inter- preted as rotating a vector (Eq. (3.98)), may be called a tensor. If, however, we consider the orthogonal matrix as a collection of fixed direction cosines, giving the new orientation ofacoordinatesystem,thereisnotensorpropertyinvolved. Thesymmetryandantisymmetrypropertiesdefinedearlierarepreservedunder orthog- onalsimilaritytransformations.Let Abeasymmetricmatrix, A=˜A, and A′=BAB−1. (3.104) Now, ˜A′=˜B−1˜A˜B=B˜AB−1, (3.105) sinceBis orthogonal.But A=˜A. Therefore ˜A′=BAB−1=A′, (3.106) showingthatthepropertyofsymmetryisinvariantunderanorthogonalsimilaritytransfor- mation. In general, symmetry is notpreserved under a nonorthogonal similarity transfor- mation. Exercises Note.Assumeallmatrixelementsarereal. 3.3.1 Showthattheproductoftwoorthogonalmatricesis orthogonal. Note.This is a key step in showing that all n×northogonal matrices form a group (Section4.1). 3.3.2 IfAis orthogonal,showthatits determinant =±1. 3.3.3 IfAisorthogonalanddet A=+1,showthat (detA)aij=Cij,whereCijisthecofactor ofaij. This yields the identities of Eq. (1.46), used in Section 1.4 to show that a cross productof vectors(inthree-space)is itselfavector. Hint.NoteExercise3.2.32. 3.3.4 Anothersetof Eulerrotationsincommonuseis (1) arotationaboutthe x3-axisthroughanangle ϕ, counterclockwise, (2) arotationaboutthe x′ 1-axisthroughanangle θ,counterclockwise, (3) arotationaboutthe x′′ 3-axisthroughanangle ψ, counterclockwise. If α=ϕ−π/2ϕ=α+π/2 β=θθ =β γ=ψ+π/2ψ=γ−π/2, showthatthefinalsystemsareidentical. 3.3 Orthogonal Matrices 207 3.3.5 Suppose the Earth is moved (rotated) so that the north pole goes to 30◦north, 20◦west (originallatitudeandlongitudesystem)andthe10◦westmeridianpointsduesouth. (a) WhataretheEuleranglesdescribingthisrotation? (b) Findthecorrespondingdirectioncosines. ANS.(b)A= 0.9551−0.2552−0.1504 0.0052 0 .5221−0.8529 0.2962 0 .8138 0 .5000. 3.3.6 VerifythattheEuleranglerotationmatrix,Eq.(3.94),isinvariantunderthetransforma- tion α→α+π, β→−β, γ→γ−π. 3.3.7 ShowthattheEuleranglerotationmatrix A(α,β,γ) satisfiesthefollowingrelations: (a)A−1(α,β,γ)=˜A(α,β,γ), (b)A−1(α,β,γ)=A(−γ,−β,−α). 3.3.8 Showthatthetraceof theproductofasymmetricandanantisymmetricmatrixiszero. 3.3.9 Showthatthetraceof amatrixremainsinvariantundersimilaritytransformations. 3.3.10 Show that the determinant of a matrix remains invariant under similarity transforma- tions. Note.Exercises (3.3.9) and (3.3.10) show that the trace and the determinant are inde- pendent of the Cartesian coordinates. They are characteristics of the matrix (operator) itself. 3.3.11 Show that the property of antisymmetry is invariant under orthogonal similarity trans- formations. 3.3.12 Ais 2×2 andorthogonal.Findthemostgeneralformof A=parenleftbiggab cdparenrightbigg . Comparewithtwo-dimensionalrotation. 3.3.13|x/angbracketrightand|y/angbracketrightare column vectors. Under an orthogonal transformation S,|x′/angbracketright=S|x/angbracketright, |y′/angbracketright=S|y/angbracketright.Showthatthescalarproduct /angbracketleftx|y/angbracketrightisinvariantunderthisorthogonaltrans- formation. Note.Thisisequivalenttotheinvarianceofthedotproductoftwovectors,Section1.3. 3.3.14 Show that the sum of the squares of the elements of a matrix remains invariant under orthogonalsimilaritytransformations. 3.3.15 AsageneralizationofExercise3.3.14, showthat summationdisplay jkSjkTjk=summationdisplay l,mS′ lmT′ lm, 208 Chapter 3 Determinants and Matrices where the primed and unprimed elements are related by an orthogonal similarity trans- formation. This result is useful in deriving invariants in electromagnetic theory (com- pareSection4.6). Note.This product Mjk=summationtextSjkTjkis sometimes called a Hadamard product .I nt h e framework of tensor analysis, Chapter 2, this exercise becomes a double contraction of twosecond-ranktensors andthereforeisclearlyascalar(invariant). 3.3.16 A rotation ϕ1+ϕ2about the z-axis is carried out as two successive rotations ϕ1and ϕ2, each about the z-axis. Use the matrix representation of the rotations to derive the trigonometricidentities cos(ϕ1+ϕ2)=cosϕ1cosϕ2−sinϕ1sinϕ2, sin(ϕ1+ϕ2)=sinϕ1cosϕ2+cosϕ1sinϕ2. 3.3.17 Acolumnvector Vhascomponents V1andV2inaninitial(unprimed)system.Calculate V′ 1andV′ 2for a (a) rotationofthecoordinatesthroughanangleof θcounterclockwise , (b) rotationofthevectorthroughanangleof θclockwise . Theresultsfor parts(a) and(b)shouldbeidentical. 3.3.18 Write a subroutine that will test whether a real N×Nmatrix is symmetric. Symmetry maybedefinedas 0≤|aij−aji|≤ε, whereεis some small tolerance (which allows for truncation error, and so on in the computer). 3.4 H ERMITIAN MATRICES ,UNITARY MATRICES Definitions Thus far it has generally been assumed that our linear vector space is a real space and that the matrix elements (the representations of the linear operators) are real. For many calculations in classical physics, real matrix elements will suffice. However, in quantum mechanics complex variables are unavoidable because of the form of the basic commuta- tionrelations(ortheformofthetime-dependentSchrödingerequation).Withthisinmind, we generalize to the case of complex matrix elements. To handle these elements, let us define,or label,somenewproperties. 1. Complex conjugate, A∗, formed by taking the complex conjugate (i→−i)of each element,where i=√ −1. 2. Adjoint, A†,formedbytransposing A∗, A†=tildewiderA∗=˜A∗. (3.107) 3.4 Hermitian Matrices, Unitary Matrices 209 3. Hermitianmatrix:Thematrix Ais labeled Hermitian (orself-adjoint )if A=A†. (3.108) IfAis real, then A†=˜Aand real Hermitian matrices are real symmetric matrices. In quantum mechanics (or matrix mechanics) matrices are usually constructed to be Hermitian,or unitary. 4. Unitarymatrix:Matrix Uislabeled unitaryif U†=U−1. (3.109) IfUis real, then U−1=˜U, so real unitary matrices are orthogonal matrices. This representsa generalizationof theconceptof orthogonalmatrix(compareEq. (3.84)). 5.(AB)∗=A∗B∗,(AB)†=B†A†. If the matrix elements are complex, the physicist is almost always concerned with Her- mitian and unitary matrices. Unitary matrices are especially important in quantum me- chanicsbecausetheyleavethelengthofa(complex)vectorunchanged—analogoustothe operation of an orthogonal matrix on a real vector. It is for this reason that the S matrix of scattering theory is a unitary matrix. One important exception to this interest in unitary matrices is the group of Lorentz matrices, Chapter 4. Using Minkowski space, we see that thesematricesarenotunitary. In a complex n-dimensional linear space the square of the length of a point ˜x= xT(x1,x2,...,xn), or the square of its distance from the origin 0, is defined as x†x=summationtextx∗ ixi=summationtext|xi|2. If a coordinate transformation y=Uxleaves the distance unchanged, thenx†x=y†y=(Ux)†Ux=x†U†Ux. Sincexis arbitrary it follows that U†U=1n; that is,Uis a unitary n×nmatrix. If x′=Axis a linear map, then its matrix in the new coordinatesbecomestheunitary(analogofasimilarity)transformation A′=UAU†, (3.110) becauseUx′=y′=UAx=UAU−1y=UAU†y. Pauli and Dirac Matrices Thesetofthree 2 ×2 Paulimatrices σ, σ1=parenleftbigg01 10parenrightbigg ,σ 2=parenleftbigg0−i i0parenrightbigg ,σ 3=parenleftbigg100−1parenrightbigg , (3.111) were introduced by W. Pauli to describe a particle of spin 1 /2 in nonrelativistic quantum mechanics.Itcanreadilybeshownthat(compareExercises3.2.13and3.2.14)thePauli σ satisfy σiσj+σjσi=2δij12,anticommutation (3.112) σiσj=iσk, i,j,k acyclicpermutationof 1,2, 3 (3.113) (σi)2=12, (3.114) 210 Chapter 3 Determinants and Matrices where 1 2is the 2×2 unit matrix. Thus, the vector σ/2 satisfies the same commutation relations, [σi,σj]≡σiσj−σjσi=2iεijkσk, (3.115) as the orbital angular momentum L(L×L=iL, see Exercise 2.5.15 and the SO(3)and SU(2)groupsinChapter4). The three Pauli matrices σand the unit matrix form a complete set, so any Hermitian 2×2m a t r i xMmaybeexpandedas M=m012+m1σ1+m2σ2+m3σ3=m0+m·σ, (3.116) wherethe miformaconstantvector m.Using(σi)2=12andtrace(σi)=0weobtainfrom Eq.(3.116) theexpansioncoefficients mibyformingtraces, 2m0=trace(M),2mi=trace(Mσi), i=1,2,3. (3.117) Adding and multiplying such 2 ×2 matrices we generate the Pauli algebra.17Note that trace(σi)=0f o ri=1,2,3. In 1927 P. A. M. Dirac extended this formalism to fast-moving particles of spin1 2, such as electrons (and neutrinos). To include special relativity he started from Einstein’s energy,E2=p2c2+m2c4, instead of the nonrelativistic kinetic and potential energy, E=p2/2m+V. ThekeytotheDiracequationistofactorize E2−p2c2=E2−(cσ·p)2=(E−cσ·p)(E+cσ·p)=m2c4(3.118) usingthe 2×2 matrixidentity (σ·p)2=p212. (3.119) The 2×2 unit matrix 1 2is not written explicitly in Eq. (3.118), and Eq. (3.119) follows fromExercise3.2.14for a=b=p.Equivalently,wecanintroducetwomatrices γ′andγ tofactorize E2−p2c2directly: bracketleftbig Eγ′⊗12−c(γ⊗σ)·pbracketrightbig2 =E2γ′2⊗12+c2γ2⊗(σ·p)2−Ec(γ′γ+γγ′)⊗σ·p =E2−p2c2=m2c4. (3.119′) ForEq.(3.119′) tohold,theconditions γ′2=1=−γ2,γ′γ+γγ′=0 (3.120) must be satisfied. Thus, the matrices γ′andγanticommute, just like the three Pauli ma- trices; therefore they cannot be real or complex numbers. Because the conditions (3.120) can be met by 2 ×2 matrices, we have written direct product signs (see Example 3.2.1) in Eq.(3.119′) because γ′,γaremultipliedby 1 2,σmatrices,respectively,with γ′=parenleftbigg10 0−1parenrightbigg ,γ=parenleftbigg01 −10parenrightbigg . (3.121) 17For its geometrical significance, seeW. E.Baylis, J.Huschilt, andJiansu Wei, Am.J.Phys. 60: 788 (1992). 3.4 Hermitian Matrices, Unitary Matrices 211 The direct-product 4 ×4 matrices in Eq. (3 .119′) are the four conventional Dirac γ-matrices, γ0=γ′⊗12=parenleftbigg120 0−12parenrightbigg = 10 0 0 01 0 0 00−10 00 0−1 , γ1=γ⊗σ1=parenleftbigg0σ1 −σ10parenrightbigg = 00 0 1 00 1 0 0−100 −100 0 , γ3=γ⊗σ3=parenleftbigg0σ3 −σ30parenrightbigg = 00 10 00 0−1 −100 0 01 00 , (3.122) and similarly for γ2=γ⊗σ2. In vector notation γ=γ⊗σis a vector with three components, each a 4 ×4 matrix, a generalization of the vector of Pauli matrices to a vector of 4×4 matrices. The four matrices γiare the components of the four-vector γµ=(γ0,γ1,γ2,γ3). If werecognizeinEq. (1 .119′) Eγ′⊗12−c(γ⊗σ)·p=γµpµ=γ·p=(γ0,γ)·(E,cp) (3.123) as a scalar product of two four-vectors γµandpµ(see Lorentz group in Chapter 4), then Eq.(3.119′)withp2=p·p=E2−p2c2mayberegardedasafour-vectorgeneralization ofEq. (3.119). Summarizing the relativistic treatment of a spin 1/2particle, it leads to 4×4matrices, whilethespin 1/2ofanonrelativisticparticleis describedbythe 2×2Paulimatrices σ. By analogy with the Pauli algebra, we can form products of the basic γµmatrices and linear combinations of them and the unit matrix 1 =14, thereby generating a 16- dimensional (so-called Clifford18) algebra. A basis (with convenient Lorentz transforma- tionproperties,see Chapter4) isgiven(in 2 ×2 matrixnotationofEq. (3.122))by 14,γ5=iγ0γ1γ2γ3=parenleftbigg012 120parenrightbigg ,γµ,γ5γµ,σµν=iparenleftbig γµγν−γνγµparenrightbig /2.(3.124) Theγ-matricesanticommute;thatis, theirsymmetriccombinations γµγν+γνγµ=2gµν14, (3.125) whereg00=1=−g11=−g22=−g33, andgµν=0f o rµ/negationslash=ν, are zero or proportional to the 4×4 unit matrix 1 4, while the six antisymmetric combinations in Eq. (3.124) give new basis elements that transform like a tensor under Lorentz transformations (see Chap- ter 4). Any 4×4 matrix can be expanded in terms of these 16 elements, and the expan- sion coefficients are given by forming traces similar to the 2 ×2 case in Eq. (3.117) us- 18D.HestenesandG.Sobczyk, loc.cit.;D.Hestenes, Am.J.Phys. 39: 1013 (1971); and J. Math. Phys. 16: 556 (1975). 212 Chapter 3 Determinants and Matrices ing trace(14)=4,trace(γ5)=0,trace(γµ)=0=trace(γ5γµ),trace(σµν)=0f o rµ,ν= 0,1,2,3 (see Exercise 3.4.23). In Chapter 4 we show that γ5is odd under parity, so γ5γµ transformlikeanaxialvectorthathasevenparity. The spin algebra generated by the Pauli matrices is just a matrix representation of the four-dimensionalCliffordalgebra,whileHestenesandcoworkers(loc.cit.)havedeveloped in theirgeometric calculus a representation-free (that is, “coordinate-free”) algebra that containscomplexnumbers,vectors,thequaternionsubalgebra,andgeneralizedcrossprod- uctsasdirectedareas(called bivectors).Thisalgebraic-geometricframeworkistailoredto nonrelativisticquantummechanics,wherespinorsacquiregeometricaspectsandtheGauss and Stokes theorems appear as components of a unified theorem. Their geometric algebra corresponding to the 16-dimensional Clifford algebra of Dirac γ-matrices is the appropri- atecoordinate-freeframework forrelativisticquantummechanicsandelectrodynamics. The discussion of orthogonal matrices in Section 3.3 and unitary matrices in this sec- tion is only a beginning. Further extensions are of vital concern in “elementary” particle physics.WiththePauliandDiracmatrices,wecandevelop spinorwavefunctionsforelec- trons,protons,andother(relativistic)spin1 2particles.Thecoordinatesystemrotationslead toDj(α,β,γ), the rotation group usually represented by matrices in which the elements are functions of the Euler angles describing the rotation. The special unitary group SU(3) (composedof3 ×3unitarymatriceswithdeterminant +1)hasbeenusedwithconsiderable successtodescribemesonsandbaryonsinvolvedinthestronginteractions,agaugetheory that is now called quantum chromodynamics . These extensions are considered further in Chapter4. Exercises 3.4.1 Showthat det(A∗)=(detA)∗=detparenleftbig A†parenrightbig . 3.4.2 Threeangularmomentummatricessatisfythebasiccommutationrelation [Jx,Jy]=iJz (andcyclicpermutationofindices).Iftwoofthematriceshaverealelements,showthat theelementsofthethirdmustbepureimaginary. 3.4.3 Showthat (AB)†=B†A†. 3.4.4 Am a t r i xC=S†S. Show that the trace is positive definite unless Sis the null matrix, inwhichcasetrace (C)=0. 3.4.5 IfAandBare Hermitian matrices, show that (AB+BA)andi(AB−BA)are also Hermitian. 3.4.6 The matrix CisnotHermitian. Show that then C+C†andi(C−C†)are Hermitian. Thismeansthata non-HermitianmatrixmayberesolvedintotwoHermitianparts, C=1 2parenleftbig C+C†parenrightbig +1 2iiparenleftbig C−C†parenrightbig . ThisdecompositionofamatrixintotwoHermitianmatrixpartsparallelsthedecompo- sitionofacomplexnumber zintox+iy, wherex=(z+z∗)/2 andy=(z−z∗)/2i. 3.4 Hermitian Matrices, Unitary Matrices 213 3.4.7 AandBaretwononcommutingHermitianmatrices: AB−BA=iC. Provethat CisHermitian. 3.4.8 Show that a Hermitian matrix remains Hermitian under unitary similarity transforma- tions. 3.4.9 Twomatrices AandBareeachHermitian.Findanecessaryandsufficientconditionfor theirproduct ABtobeHermitian. ANS.[A,B]=0. 3.4.10 Showthatthereciprocal(thatis,inverse)of aunitarymatrixisunitary. 3.4.11 Aparticularsimilaritytransformationyields A′=UAU−1, A′†=UA†U−1. If the adjoint relationship is preserved (A†′=A′†)and detU=1, show that Umust be unitary. 3.4.12 Twomatrices UandHarerelatedby U=eiaH, withareal. (The exponential function is defined by a Maclaurin expansion. This will bedoneinSection5.6.) (a) IfHisHermitian,showthat Uisunitary. (b) IfUisunitary,showthat HisHermitian.( Hisindependentof a.) Note.WithHtheHamiltonian, ψ(x,t)=U(x,t)ψ(x, 0)=exp(−itH/¯h)ψ(x,0) isasolutionofthetime-dependentSchrödingerequation. U(x,t)=exp(−itH/¯h)isthe “evolutionoperator.” 3.4.13 Anoperator T(t+ε,t)describesthechangeinthewavefunctionfrom ttot+ε.F orε realandsmallenoughsothat ε2maybeneglected, T(t+ε,t)=1−i ¯hεH(t). (a) IfTis unitary,showthat His Hermitian. (b) IfHisHermitian,showthat Tis unitary. Note.WhenH(t)isindependentoftime,thisrelationmaybeputinexponentialform— Exercise3.4.12. 214 Chapter 3 Determinants and Matrices 3.4.14 Showthatanalternateform, T(t+ε,t)=1−iεH(t)/2¯h 1+iεH(t)/2¯h, agrees with the Tof part (a) of Exercise 3.4.13, neglecting ε2, and is exactly unitary (forHHermitian). 3.4.15 Provethatthedirectproductoftwo unitarymatricesis unitary. 3.4.16 Showthat γ5anticommuteswithallfour γµ. 3.4.17 Use the four-dimensional Levi-Civita symbol ελµνρwithε0123=−1 (generalizing Eqs. (2.93) in Section 2.9 to four dimensions) and show that (i) 2 γ5σµν=−iεµναβσαβ using the summation convention of Section 2.6 and (ii) γλγµγν=gλµγν−gλνγµ+ gµνγλ+iελµνργργ5. Defineγµ=gµνγνusinggµν=gµνtoraiseandlowerindices. 3.4.18 Evaluatethefollowingtraces:(seeEq. (3.123)for thenotation) (i) trace (γ·aγ·b)=4a·b, (ii) trace (γ·aγ·bγ·c)=0, (iii) trace (γ·aγ·bγ·cγ·d)=4(a·bc·d−a·cb·d+a·db·c), (iv) trace (γ5γ·aγ·bγ·cγ·d)=4iεαβµνaαbβcµdν. 3.4.19 Show that (i) γµγαγµ=−2γα, (ii)γµγαγβγµ=4gαβ, and (iii) γµγαγβγνγµ= −2γνγβγα. 3.4.20 IfM=1 2(1+γ5), showthat M2=M. Notethat γ5maybereplacedbyanyotherDiracmatrix(any ŴiofEq.(3.124)).If Mis Hermitian,thenthis result, M2=M, is the definingequationfor a quantummechanical projectionoperator. 3.4.21 Showthat α×α=2iσ⊗12, where α=γ0γis avector α=(α1,α2,α3). Notethatif αisapolarvector(Section2.4), then σisanaxialvector. 3.4.22 Provethatthe16Diracmatricesform alinearlyindependentset. 3.4.23 If we assume that a given 4 ×4m a t r i xA(with constant elements) can be written as a linearcombinationof the16Diracmatrices A=16summationdisplay i=1ciŴi, showthat ci∼trace(AŴi). 3.5 Diagonalization of Matrices 215 3.4.24 IfC=iγ2γ0is the charge conjugation matrix, show that CγµC−1=−˜γµ, where ˜indicatestransposition. 3.4.25 Letx′ µ=/Lambda1ν µxνbearotationbyanangle θaboutthe3-axis, x′ 0=x0,x′ 1=x1cosθ+x2sinθ, x′ 2=−x1sinθ+x2cosθ, x′ 3=x3. UseR=exp(iθσ12/2)=cosθ/2+iσ12sinθ/2 (see Eq. (3.170b)) and show that theγ’s transform just like the coordinates xµ, that is, /Lambda1νµγν=R−1γµR. (Note that γµ=gµνγνand that the γµare well defined only up to a similarity transformation.) Similarly,if x′=/Lambda1xis aboost(pureLorentztransformation)alongthe1-axis,thatis, x′ 0=x0coshζ−x1sinhζ, x′ 1=−x0sinhζ+x1coshζ, x′ 2=x2,x′ 3=x3, with tanh ζ=v/candB=exp(−iζσ01/2)=coshζ/2−iσ01sinhζ/2 (see Eq. (3.170b)),showthat /Lambda1ν µγν=BγµB−1. 3.4.26 (a) Given r′=Ur, withUa unitary matrix and ra (column) vector with complex elements,showthatthenorm(magnitude)of ris invariantunderthis operation. (b) The matrix Utransforms any column vector rwith complex elements into r′, leavingthemagnitudeinvariant: r†r=r′†r′.Showthat Uis unitary. 3.4.27 Write a subroutine that will test whether a complex n×nmatrix is self-adjoint. In demandingequalityofmatrixelements aij=a† ij,allowsomesmalltolerance εtocom- pensatefor truncationerrorof thecomputer. 3.4.28 Write asubroutinethatwillformtheadjointofacomplex M×Nmatrix. 3.4.29 (a) Writeasubroutinethatwilltakeacomplex M×NmatrixAandyieldtheproduct A†A. Hint.Thissubroutinecancallthesubroutinesof Exercises3.2.41and3.4.28. (b) Test your subroutine by taking Ato be one or more of the Dirac matrices, Eq.(3.124). 3.5 D IAGONALIZATION OF MATRICES Moment of Inertia Matrix In many physical problems involving real symmetric or complex Hermitian matrices it is desirable to carry out a (real) orthogonal similarity transformation or a unitary transfor- mation (corresponding to a rotation of the coordinate system) to reduce the matrix to a diagonal form, nondiagonal elements all equal to zero. One particularly direct example of this is the moment of inertia matrix Iof a rigid body. From the definition of angular momentum Lwehave L=Iω, (3.126) 216 Chapter 3 Determinants and Matrices ωbeingtheangularvelocity.19Theinertiamatrix Iis foundtohavediagonalcomponents Ixx=summationdisplay imiparenleftbig r2 i−x2 iparenrightbig ,andsoon, (3.127) the subscript ireferring to mass milocated at ri=(xi,yi,zi). For the nondiagonal com- ponentswehave Ixy=−summationdisplay imixiyi=Iyx. (3.128) By inspection, matrix Iis symmetric. Also, since Iappears in a physical equation of the form (3.126), which holds for all orientations of the coordinate system, it may be consid- eredtobea tensor(quotientrule,Section2.3). The key now is to orient the coordinate axes (along a body-fixed frame) so that the Ixyand the other nondiagonal elements will vanish. As a consequence of this orientation and an indication of it, if the angular velocity is along one such realigned principal axis , the angular velocity and the angular momentum will be parallel. As an illustration, the stability of rotation is used by football players when they throw the ball spinning about its longprincipalaxis. Eigenvectors, Eigenvalues It is instructive to consider a geometrical picture of this problem. If the inertia matrix Iis multiplied from each side by a unit vector of variable direction, ˆn=(α,β,γ), then in the DiracbracketnotationofSection3.2, /angbracketleftˆn|I|ˆn/angbracketright=I, (3.129) whereIis the moment of inertia about the direction ˆnand a positive number (scalar). Carryingoutthemultiplication,weobtain I=Ixxα2+Iyyβ2+Izzγ2+2Ixyαβ+2Ixzαγ+2Iyzβγ, (3.130) a positive definite quadratic form that must be an ellipsoid (see Fig. 3.5). From analytic geometry it is known that the coordinate axes can always be rotated to coincide with the axesofourellipsoid.Inmanyelementarycases,especiallywhensymmetryispresent,these new axes, called the principal axes , can be found by inspection. We can find the axes by locatingthelocalextremaoftheellipsoidintermsofthevariablecomponentsof n,subject to the constraint ˆn2=1. To deal with the constraint, we introducea Lagrange multiplier λ (Section17.6).Differentiating /angbracketleftˆn|I|ˆn/angbracketright−λ/angbracketleftˆn|ˆn/angbracketright, ∂ ∂njparenleftbig /angbracketleftˆn|I|ˆn/angbracketright−λ/angbracketleftˆn|ˆn/angbracketrightparenrightbig =2summationdisplay kIjknk−2λnj=0,j=1,2,3 (3.131) yieldstheeigenvalueequations I|ˆn/angbracketright=λ|ˆn/angbracketright. (3.132) 19Themoment of inertia matrix may also be developed from the kinetic energy of arotating body, T=1/2/angbracketleftω|I|ω/angbracketright. 3.5 Diagonalization of Matrices 217 FIGURE 3.5Momentofinertiaellipsoid. Thesameresultcanbefoundbypurelygeometricmethods.Wenowproceedtodevelop ageneralmethodoffindingthediagonalelementsandtheprincipalaxes. IfR−1=˜Ris the real orthogonal matrix such that n′=Rn,o r|n′/angbracketright=R|n/angbracketrightin Dirac notation,arethenewcoordinates,thenweobtain,using /angbracketleftn′|R=/angbracketleftn|inEq.(3.132), /angbracketleftn|I|n/angbracketright=/angbracketleftn′|RI˜R|n′/angbracketright=I′ 1n′2 1+I′ 2n′2 2+I′ 3n′2 3, (3.133) where the I′ i>0 are the principal moments of inertia. The inertia matrix I′in Eq. (3.133) isdiagonalinthenewcoordinates, I′=RI˜R= I′ 100 0I′ 20 00I′ 3. (3.134) If werewriteEq. (3.134)using R−1=˜Rintheform ˜RI′=I˜R (3.135) andtake˜R=(v1,v2,v3)toconsistofthreecolumnvectors,thenEq.(3.135)splitsupinto threeeigenvalueequations, Ivi=I′ ivi,i=1,2,3 (3.136) witheigenvalues I′ iandeigenvectors v i. The names were introduced from the German literature on quantum mechanics. Because these equations are linear and homogeneous 218 Chapter 3 Determinants and Matrices (for fixed i), bySection3.1theirdeterminantshavetovanish: vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleI11−I′ iI12 I13 I12I22−I′ iI23 I13 I23I33−I′ ivextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.137) Replacing the eigenvalue I′ iby a variable λtimes the unit matrix 1, we may rewrite Eq.(3.136) as (I−λ1)|v/angbracketright=0. (3.136′) Thedeterminantsettozero, |I−λ1|=0, (3.137′) is a cubic polynomial in λ; its three roots, of course, are the I′ i. Substituting one root at a time back into Eq. (3.136) (or (3.136′)), we can find the corresponding eigenvectors. Because of its applications in astronomical theories, Eq. (3.137) (or (3.137′)) is known as thesecularequation .20Thesametreatmentappliestoanyrealsymmetricmatrix I,except thatitseigenvaluesneednotallbepositive.Also,theorthogonalityconditioninEq.(3.87) forRsaythat,ingeometricterms,theeigenvectors viaremutuallyorthogonalunitvectors. Indeed they form the new coordinate system. The fact that any two eigenvectors vi,vjare orthogonal if I′ i/negationslash=I′ jfollows from Eq. (3.136) in conjunction with the symmetry of Iby multiplyingwith viandvj, respectively, /angbracketleftvj|I|vi/angbracketright=I′ ivj·vi=/angbracketleftvi|I|vj/angbracketright=I′ jvi·vj. (3.138a) SinceI′ i/negationslash=I′ jandEq. (3.138a)impliesthat (I′ j−I′ i)vi·vj=0,sovi·vj=0. We can write the quadratic forms in Eq. (3.133) as a sum of squares in the original coordinates|n/angbracketright, /angbracketleftn|I|n/angbracketright=/angbracketleftn′|RI˜R|n′/angbracketright=summationdisplay iI′ i(n·vi)2, (3.138b) becausetherowsoftherotationmatrixin n′=Rn,or  n′ 1 n′2 n′3 = v1·n v2·n v3·n componentwise,aremadeupoftheeigenvectors vi. Theunderlyingmatrixidentity, I=summationdisplay iI′ i|vi/angbracketright/angbracketleftvi|, (3.138c) 20Equation (3.126) will take on this form when ωis along one of the principal axes. Then L=λωandIω=λω.I nt h em a t h e - matics literature λis usually calleda characteristic value ,ωacharacteristic vector . 3.5 Diagonalization of Matrices 219 maybevie wedasthe spectraldecomposition oftheinertiatensor(or anyrealsymmetric matrix). Here, the word spectralis just another term for expansion in terms of its eigen- values.Whenwemultiplythiseigenvalueexpansionby /angbracketleftn|ontheleftand |n/angbracketrightontheright we reproduce the previous relation between quadratic forms. The operator Pi=|vi/angbracketright/angbracketleftvi|is a projection operator satisfying P2 i=Pithat projects the ith component wiof any vector |w/angbracketright=summationtext jwj|vj/angbracketrightthat is expanded in terms of the eigenvector basis |vj/angbracketright. This is verified by Pi|w/angbracketright=summationdisplay jwj|vi/angbracketright/angbracketleftvi|vj/angbracketright=wi|vi/angbracketright=vi·w|vi/angbracketright. Finally,theidentity summationdisplay i|vi/angbracketright/angbracketleftvi|=1 expresses the completeness of the eigenvector basis according to which any vector |w/angbracketright=summationtext iwi|vi/angbracketrightcan be expanded in terms of the eigenvectors. Multiplying the completeness relationby|w/angbracketrightprovestheexpansion |w/angbracketright=summationtext i/angbracketleftvi|w/angbracketright|vi/angbracketright. An important extension of the spectral decomposition theorem applies to commuting symmetric(orHermitian)matrices A,B:If[A,B]=0,thenthereisanorthogonal(unitary) matrixthatdiagonalizesboth AandB;thatis,bothmatriceshavecommoneigenvectorsif theeigenvaluesarenondegenerate.Thereverseofthis theoremis alsovalid. Toprovethistheoremwediagonalize A:Avi=aivi.Multiplyingeacheigenvalueequa- tion byBwe obtain BAvi=aiBvi=A(Bvi),which says that Bviis an eigenvector of A with eigenvalue ai. HenceBvi=biviwith real bi. Conversely, if the vectors viare com- mon eigenvectors of AandB,thenABvi=Abivi=aibivi=BAvi. Since the eigenvec- torsviarecomplete,thisimplies AB=BA. Hermitian Matrices Forcomplexvectorspaces,Hermitianandunitarymatricesplaythesameroleassymmetric and orthogonal matrices over real vector spaces, respectively. First, let us generalize the important theorem about the diagonal elements and the principal axes for the eigenvalue equation A|r/angbracketright=λ|r/angbracketright, (3.139) Wenowshowthatif AisaHermitianmatrix,21itseigenvaluesarerealanditseigenvectors orthogonal. Letλiandλjbetwoeigenvaluesand |ri/angbracketrightand|rj/angbracketright,thecorrespondingeigenvectorsof A, aHermitianmatrix.Then A|ri/angbracketright=λi|ri/angbracketright, (3.140) A|rj/angbracketright=λj|rj/angbracketright. (3.141) 21IfAis real,the Hermitian requirement reduces to arequirement of symmetry. 220 Chapter 3 Determinants and Matrices Equation(3.140)is multipliedby /angbracketleftrj|: /angbracketleftrj|A|ri/angbracketright=λi/angbracketleftrj|ri/angbracketright. (3.142) Equation(3.141)is multipliedby /angbracketleftri|togive /angbracketleftri|A|rj/angbracketright=λj/angbracketleftri|rj/angbracketright. (3.143) Takingtheadjoint22ofthisequation,wehave /angbracketleftrj|A†|ri/angbracketright=λ∗ j/angbracketleftrj|ri/angbracketright, (3.144) or /angbracketleftrj|A|ri/angbracketright=λ∗j/angbracketleftrj|ri/angbracketright (3.145) sinceAis Hermitian.SubtractingEq. (3.145)fromEq. (3.142), weobtain (λi−λ∗j)/angbracketleftrj|ri/angbracketright=0. (3.146) This is a general result for all possible combinations of iandj.F i r s t ,l e t j=i. Then Eq.(3.146) becomes (λi−λ∗ i)/angbracketleftri|ri/angbracketright=0. (3.147) Since/angbracketleftri|ri/angbracketright=0 wouldbeatrivialsolutionofEq. (3.147), weconcludethat λi=λ∗ i, (3.148) orλiisreal, forall i. Second,for i/negationslash=jandλi/negationslash=λj, (λi−λj)/angbracketleftrj|ri/angbracketright=0, (3.149) or /angbracketleftrj|ri/angbracketright=0, (3.150) whichmeansthattheeigenvectorsof distincteigenvaluesareorthogonal,Eq.(3.150)being ourgeneralizationoforthogonalityinthiscomplexspace.23 Ifλi=λj(degenerate case), |ri/angbracketrightis not automatically orthogonal to |rj/angbracketright,b u ti tm a yb e madeorthogonal.24Consider the physical problem of the momentof inertia matrix again. Ifx1isanaxisofrotationalsymmetry,thenwewillfindthat λ2=λ3.Eigenvectors|r2/angbracketrightand |r3/angbracketrightare each perpendicular to the symmetry axis, |r1/angbracketright, but they lie anywhere in the plane perpendicularto |r1/angbracketright;thatis,anylinearcombinationof |r2/angbracketrightand|r3/angbracketrightisalsoaneigenvector. Consider (a2|r2/angbracketright+a3|r3/angbracketright)witha2anda3constants.Then Aparenleftbig a2|r2/angbracketright+a3|r3/angbracketrightparenrightbig =a2λ2|r2/angbracketright+a3λ3|r3/angbracketright =λ2parenleftbig a2|r2/angbracketright+a3|r3/angbracketrightparenrightbig , (3.151) 22Note/angbracketleftrj|=|rj/angbracketright†for complex vectors. 23The corresponding theory for differential operators (Sturm–Liouville theory) appears in Section 10.2. The integral equation analog (Hilbert–Schmidt theory) is given in Section 16.4. 24We are assuming here that the eigenvectors of the n-fold degenerate λispan the corresponding n-dimensional space. This may be shown by including a parameter εin the original matrix to remove the degeneracy and then letting εapproach zero (compareExercise3.5.30).Thisisanalogoustobreakingadegeneracyinatomicspectroscopybyapplyinganexternalmagnetic field(Zeemaneffect). 3.5 Diagonalization of Matrices 221 as is to be expected, for x1is an axis of rotational symmetry. Therefore, if |r1/angbracketrightand|r2/angbracketright are fixed,|r3/angbracketrightmay simply be chosen to lie in the plane perpendicular to |r1/angbracketrightand also perpendicular to |r2/angbracketright. A general method of orthogonalizing solutions, the Gram–Schmidt process(Section3.1), isappliedtofunctionsinSection10.3. The set of northogonal eigenvectors |ri/angbracketrightof ourn×nHermitian matrix Aforms a complete set, spanning the n-dimensional (complex) space,summationtext i|ri/angbracketright/angbracketleftri|=1. This fact is usefulinavariationalcalculationoftheeigenvalues,Section17.8. The spectral decomposition of any Hermitian matrix Ais proved by analogy with real symmetricmatrices A=summationdisplay iλi|ri/angbracketright/angbracketleftri|, withrealeigenvalues λiandorthonormaleigenvectors |ri/angbracketright. Eigenvalues and eigenvectors are not limited to Hermitian matrices. All matrices have at least one eigenvalue and eigenvector. However, only Hermitian matrices have all eigen- vectorsorthogonalandalleigenvaluesreal. Anti-Hermitian Matrices Occasionallyinquantumtheoryweencounteranti-Hermitianmatrices: A†=−A. Followingtheanalysisof thefirst portionof thissection,wecanshowthat a. Theeigenvaluesarepureimaginary(or zero). b. Theeigenvectorscorrespondingtodistincteigenvaluesareorthogonal. The matrix Rformed from the normalized eigenvectors is unitary. This anti-Hermitian propertyis preservedunderunitarytransformations. Example 3.5.1 EIGENVALUES AND EIGENVECTORS OF A REALSYMMETRIC MATRIX Let A= 010 100 000 . (3.152) Thesecularequationis vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle−λ10 1−λ0 00−λvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0, (3.153) or −λparenleftbig λ2−1parenrightbig =0, (3.154) 222 Chapter 3 Determinants and Matrices expandingbyminors.Therootsare λ=−1,0,1.Tofindtheeigenvectorcorrespondingto λ=−1,wesubstitutethisvaluebackintotheeigenvalueequation,Eq.(3.139),  −λ10 1−λ0 00−λx y z =0 0 0 . (3.155) Withλ=−1,thisyields x+y=0,z=0. (3.156) Within an arbitrary scale factor and an arbitrary sign (or phase factor), /angbracketleftr1|=(1,−1,0). Note that (for real |r/angbracketrightin ordinary space) the eigenvector singles out a line in space. The positive or negative sense is not determined. This indeterminancy could be expected if we noted that Eq. (3.139) is homogeneous in |r/angbracketright. For convenience we will require that the eigenvectorsbenormalizedtounity, /angbracketleftr1|r1/angbracketright=1.Withthiscondition, /angbracketleftr1|=parenleftbigg1 √ 2,−1√ 2,0parenrightbigg (3.157) isfixedexceptfor anoverallsign. For λ=0,Eq. (3.139)yields y=0,x=0, (3.158) /angbracketleftr2|=(0,0,1)isa suitableeigenvector.Finally,for λ=1,weget −x+y=0,z=0, (3.159) or /angbracketleftr3|=parenleftbigg1√ 2,1√ 2,0parenrightbigg . (3.160) The orthogonality of r1,r2, andr3, corresponding to three distinct eigenvalues, may be easilyverified. Thecorrespondingspectraldecompositiongives A=(−1)parenleftbigg1√ 2,−1√ 2,0parenrightbigg 1√ 2 −1√ 2 0 +(+1)parenleftbigg1√ 2,1√ 2,0parenrightbigg 1√ 2 1√ 2 0 +0(0,0,1) 0 0 1  =− 1 2−1 20 −1 21 20 00 0 + 1 21 20 1 21 20 000 = 010 100 000 . /squaresolid 3.5 Diagonalization of Matrices 223 Example 3.5.2 DEGENERATE EIGENVALUES Consider A= 100 001 010 . (3.161) Thesecularequationis vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle1−λ00 0−λ1 01−λvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0 (3.162) or (1−λ)parenleftbig λ2−1parenrightbig =0,λ=−1,1,1, (3.163) adegeneratecase.If λ=−1,theeigenvalueequation(3.139)yields 2x=0,y+z=0. (3.164) Asuitablenormalizedeigenvectoris /angbracketleftr1|=parenleftbigg 0,1 √ 2,−1√ 2parenrightbigg . (3.165) Forλ=1,weget −y+z=0. (3.166) Any eigenvector satisfying Eq. (3.166) is perpendicular to r1. We have an infinite number ofchoices.Suppose,as onepossiblechoice, r2is takenas /angbracketleftr2|=parenleftbigg 0,1√ 2,1√ 2parenrightbigg , (3.167) which clearly satisfies Eq. (3.166). Then r3must be perpendicular to r1and may be made perpendicularto r2by25 r3=r1×r2=(1,0,0). (3.168) Thecorrespondingspectraldecompositiongives A=−parenleftbigg 0,1√ 2,−1√ 2parenrightbigg 0 1√ 2 −1√ 2 +parenleftbigg 0,1√ 2,1√ 2parenrightbigg 0 1√ 2 1√ 2 +(1,0,0) 1 0 0  =− 00 0 01 2−1 2 0−1 21 2 + 000 01 21 2 01 21 2 + 100 000 000 =100 001 010 . /squaresolid 25Theuse of thecross product is limited to three-dimensional space(see Section 1.4). 224 Chapter 3 Determinants and Matrices Functions of Matrices Polynomials with one or more matrix arguments are well defined and occur often. Power series of a matrix may also be defined, provided the series converge (see Chapter 5) for eachmatrixelement.Forexample,if Ais anyn×nmatrix,thenthepowerseries exp(A)=∞summationdisplay j=01 j!Aj, (3.169a) sin(A)=∞summationdisplay j=0(−1)j (2j+1)!A2j+1, (3.169b) cos(A)=∞summationdisplay j=0(−1)j (2j)!A2j(3.169c) arewelldefined n×nmatrices.ForthePaulimatrices σktheEuleridentity forrealθand k=1,2,or 3 exp(iσkθ)=12cosθ+iσksinθ, (3.170a) follows from collecting all even and odd powers of θin separate series using σ2 k=1. For the 4×4 Dirac matrices σjk=1 with(σjk)2=1i fj/negationslash=k=1,2 or 3 we obtain similarly (withoutwritingtheobviousunitmatrix 14anymore) expparenleftbig iσjkθparenrightbig =cosθ+iσjksinθ, (3.170b) while expparenleftbig iσ0kζparenrightbig =coshζ+iσ0ksinhζ (3.170c) holdsfor real ζbecause(iσ0k)2=1f o rk=1,2,or 3. ForaHermitianmatrix Athereisaunitarymatrix Uthatdiagonalizesit;thatis, UAU†= [a1,a2,...,an].Thenthe traceformula detparenleftbig exp(A)parenrightbig =expparenleftbig trace(A)parenrightbig (3.171) isobtained(seeExercises3.5.2 and3.5.9) from detparenleftbig exp(A)parenrightbig =detparenleftbig Uexp(A)U†parenrightbig =detparenleftbig expparenleftbig UAU†parenrightbigparenrightbig =detexp[a1,a2,...,an]=detbracketleftbig ea1,ea2,...,eanbracketrightbig =productdisplay eai=expparenleftBigsummationdisplay aiparenrightBig =expparenleftbig trace(A)parenrightbig , usingUAiU†=(UAU†)iin the power series Eq. (3.169a) for exp (UAU†)and the product theoremfor determinantsinSection3.2. 3.5 Diagonalization of Matrices 225 Thistraceformulaisaspecialcaseofthe spectraldecompositionlaw forany(infinitely differentiable)function f(A)for Hermitian A: f(A)=summationdisplay if(λi)|ri/angbracketright/angbracketleftri|, where|ri/angbracketrightare the common eigenvectors of AandAj. This eigenvalue expansion follows fromAj|ri/angbracketright=λj i|ri/angbracketright,multiplied by f(j)(0)/j!and summed over jto form the Taylor expansion of f(λi)and yield f(A)|ri/angbracketright=f(λi)|ri/angbracketright. Finally, summing over iand using completenessweobtain f(A)summationtext i|ri/angbracketright/angbracketleftri|=summationtext if(λi)|ri/angbracketright/angbracketleftri|=f(A),q.e.d. Example 3.5.3 EXPONENTIAL OF A DIAGONAL MATRIX If thematrix Ais diagonallike σ3=parenleftbigg10 0−1parenrightbigg , then itsnth power is also diagonal with its diagonal, matrix elements raised to the nth power: (σ3)n=parenleftbigg100(−1)nparenrightbigg . Thensummingtheexponentialseries, elementfor element,yields eσ3=parenleftBiggsummationtext∞ n=01 n!0 0summationtext∞ n=0(−1)n n!parenrightBigg =parenleftBigg e0 01 eparenrightBigg . Ifwewritethegeneraldiagonalmatrixas A=[a1,a2,...,an]withdiagonalelements aj, thenAm=[am 1,am 2,...,am n], and summing the exponentials elementwise again we obtain eA=[ea1,ea2,...,ean]. Usingthespectraldecompositionlawweobtaindirectly eσ3=e+1(1,0)parenleftbigg1 0parenrightbigg +e−1(0,1)parenleftbigg01parenrightbigg =parenleftbigge0 0e−1parenrightbigg ./squaresolid Anotherimportantrelationis the Baker–Hausdorffformula , exp(iG)Hexp(−iG)=H+[iG,H]+1 2bracketleftbig iG,[iG,H]bracketrightbig +···, (3.172) whichfollowsfrommultiplyingthepowerseriesforexp (iG)andcollectingthetermswith thesamepowersof iG. Herewedefine [G,H]=GH−HG asthecommutator ofGandH. The preceding analysis has the advantage of exhibiting and clarifying conceptual rela- tionships in the diagonalization of matrices. However, for matrices larger than 3 ×3, or perhaps 4×4,theprocessrapidlybecomessocumbersomethatweturntocomputersand 226 Chapter 3 Determinants and Matrices iterative techniques.26One such technique is the Jacobi method for determining eigenval- ues and eigenvectors of real symmetric matrices. This Jacobi technique for determining eigenvaluesandeigenvectorsandtheGauss–Seidelmethodofsolvingsystemsofsimulta- neous linear equations are examples of relaxation methods. They are iterative techniques in which the errors may decrease or relax as the iterations continue. Relaxation methods areusedextensivelyfor thesolutionof partialdifferentialequations. Exercises 3.5.1 (a) Startingwiththeorbitalangularmomentumofthe ith elementof mass, Li=ri×pi=miri×(ω×ri), derivetheinertiamatrixsuchthat L=Iω,|L/angbracketright=I|ω/angbracketright. (b) Repeatthederivationstartingwithkineticenergy Ti=1 2mi(ω×ri)2parenleftbigg T=1 2/angbracketleftω|I|ω/angbracketrightparenrightbigg . 3.5.2 Show that the eigenvalues of a matrix are unaltered if the matrix is transformed by a similaritytransformation. This property is not limited to symmetric or Hermitian matrices. It holds for any ma- trix satisfying the eigenvalue equation, Eq. (3.139). If our matrix can be brought into diagonalform byasimilaritytransformation,thentwoimmediateconsequencesare 1. The trace(sum ofeigenvalues)isinvariantunderasimilaritytransformation. 2. The determinant (product of eigenvalues) is invariant under a similarity transfor- mation. Note.The invariance of the trace and determinant are often demonstrated by using the Cayley–Hamiltontheorem:Amatrixsatisfiesits owncharacteristic(secular) equation. 3.5.3 As a converse of the theorem that Hermitian matrices have real eigenvalues and that eigenvectorscorrespondingtodistincteigenvaluesareorthogonal,showthatif (a) theeigenvaluesofamatrixarerealand (b) theeigenvectorssatisfy r† irj=δij=/angbracketleftri|rj/angbracketright, thenthematrixisHermitian. 3.5.4 Show that a real matrix that is not symmetric cannot be diagonalized by an orthogonal similaritytransformation. Hint.Assume that the nonsymmetric real matrix can be diagonalized and develop a contradiction. 26In higher-dimensional systems the secular equation may be strongly ill-conditioned with respect to the determination of its roots (the eigenvalues). Direct solution by computer may be very inaccurate. Iterative techniques for diagonalizing the original matrix areusually preferred. SeeSections 2.7 and 2.9 ofPress et al.,loc. cit. 3.5 Diagonalization of Matrices 227 3.5.5 The matrices representing the angular momentum components Jx,Jy, andJzare all Hermitian. Show that the eigenvalues of J2, whereJ2=J2 x+J2 y+J2 z, are real and nonnegative. 3.5.6 Ahaseigenvalues λiandcorrespondingeigenvectors |xi/angbracketright.Showthat A−1hasthesame eigenvectorsbutwitheigenvalues λ−1 i. 3.5.7 Asquarematrixwithzerodeterminantislabeled singular. (a) IfAis singular,showthatthereis atleastonenonzerocolumnvector vsuchthat A|v/angbracketright=0. (b) If thereis anonzerovector |v/angbracketrightsuchthat A|v/angbracketright=0, showthatAisasingularmatrix.Thismeansthatifamatrix(oroperator)haszero as an eigenvalue, the matrix (or operator) has no inverse and its determinant is zero. 3.5.8 The same similarity transformation diagonalizes each of two matrices. Show that the original matrices must commute. (This is particularly important in the matrix (Heisen- berg)formulationofquantummechanics.) 3.5.9 Two Hermitian matrices AandBhave the same eigenvalues. Show that AandBare relatedbyaunitarysimilaritytransformation. 3.5.10 Find the eigenvalues and an orthonormal (orthogonal and normalized) set of eigenvec- torsfor thematricesof Exercise3.2.15. 3.5.11 Show that the inertia matrix for a single particle of mass mat(x,y,z)has a zero de- terminant. Explain this result in terms of the invariance of the determinant of a matrix undersimilaritytransformations(Exercise3.3.10)andapossiblerotationofthecoordi- natesystem. 3.5.12 A certain rigid body may be represented by three point masses: m1=1a t(1,1,−2), m2=2a t(−1,−1,0),andm3=1a t(1,1,2). (a) Findtheinertiamatrix. (b) Diagonalizetheinertiamatrix,obtainingtheeigenvaluesandtheprincipalaxes(as orthonormaleigenvectors). 3.5.13 Unitmasses areplacedasshowninFig.3.6. (a) Findthemomentofinertiamatrix. (b) Findtheeigenvaluesandasetof orthonormaleigenvectors. (c) Explainthedegeneracyinterms ofthesymmetryofthesystem. ANS.I= 4−1−1 −14−1 −1−14λ1=2 r1=(1/√ 3,1/√ 3,1/√ 3) λ2=λ3=5. 228 Chapter 3 Determinants and Matrices FIGURE 3.6Mass sitesfor inertiatensor. 3.5.14 Am a s s m1=1/2 kg is located at (1,1,1)(meters), a mass m2=1/2k gi sa t (−1,−1,−1).Thetwomassesareheldtogetherbyanideal(weightless,rigid)rod. (a) Findtheinertiatensorof thispairofmasses. (b) Findtheeigenvaluesandeigenvectorsofthis inertiamatrix. (c) Explain the meaning, the physical significance of the λ=0 eigenvalue. What is thesignificanceofthecorrespondingeigenvector? (d) Now that you have solved this problem by rather sophisticated matrix techniques, explainhowyoucouldobtain (1)λ=0 andλ=? — byinspection(thatis, usingcommonsense). (2)rλ=0=? — byinspection(thatis,usingfreshmanphysics). 3.5.15 Unitmassesareattheeightcornersofacube (±1,±1,±1).Findthemomentofinertia matrix and show that there is a triple degeneracy. This means that so far as moments of inertiaare concerned,thecubicstructureexhibitssphericalsymmetry. Findtheeigenvaluesandcorrespondingorthonormaleigenvectorsof thefollowingma- trices (as a numerical check, note that the sum of the eigenvalues equals the sum of the diagonalelementsoftheoriginalmatrix,Exercise3.3.9).Notealsothecorrespondence betweendet A=0 andtheexistenceof λ=0,asrequiredbyExercises3.5.2and3.5.7. 3.5.16 A= 101 010 101 . ANS.λ=0,1,2. 3.5.17 A=1√ 20√ 200 00 0 . ANS.λ=−1,0,2. 3.5 Diagonalization of Matrices 229 3.5.18 A= 110 101 011 . ANS.λ=−1,1,2. 3.5.19 A=1√ 80√ 81√ 8 0√ 81 . ANS.λ=−3,1,5. 3.5.20 A= 100 011 011 . ANS.λ=0,1,2. 3.5.21 A=10 0 01√ 2 0√ 20 . ANS.λ=−1,1,2. 3.5.22 A= 010 101 010 . ANS.λ=−√ 2,0,√ 2. 3.5.23 A= 200 011 011 . ANS.λ=0,2,2. 3.5.24 A=011 101 110 . ANS.λ=−1,−1,2. 3.5.25 A=1−1−1 −11−1 −1−11. ANS.λ=−1,2,2. 3.5.26 A= 111 111 111 . ANS.λ=0,0,3. 230 Chapter 3 Determinants and Matrices 3.5.27 A= 502 010 202 . ANS.λ=1,1,6. 3.5.28 A=110 110 000 . ANS.λ=0,0,2. 3.5.29 A= 50√ 3 030√ 30 3 . ANS.λ=2,3,6. 3.5.30 (a) Determinetheeigenvaluesandeigenvectorsof parenleftbigg1ε ε1parenrightbigg . Note that the eigenvalues are degenerate for ε=0 but that the eigenvectors are orthogonalfor all ε/negationslash=0 andε→0. (b) Determinetheeigenvaluesandeigenvectorsof parenleftbigg11 ε21parenrightbigg . Notethattheeigenvaluesaredegeneratefor ε=0andthatforthis(nonsymmetric) matrixtheeigenvectors (ε=0)donotspanthespace. (c) Find the cosine of the angle between the two eigenvectors as a function of εfor 0≤ε≤1. 3.5.31 (a) Take the coefficients of the simultaneous linear equations of Exercise 3.1.7 to be the matrix elements aijof matrixA(symmetric). Calculate the eigenvalues and eigenvectors. (b) Formamatrix Rwhosecolumnsaretheeigenvectorsof A,andcalculatethetriple matrixproduct ˜RAR. ANS.λ=3.33163. 3.5.32 RepeatExercise3.5.31byusingthematrixofExercise3.2.39. 3.5.33 Describethegeometricpropertiesof thesurface x2+2xy+2y2+2yz+z2=1. Howis itorientedinthree-dimensionalspace?Is it aconicsection?If so, whichkind? 3.6 Normal Matrices 231 Table 3.1 Matrix Eigenvalues Eigenvectors (for different eigenvalues) Hermitian Real Orthogonal Anti-Hermitian Pure imaginary (or zero) Orthogonal Unitary Unit magnitude Orthogonal Normal If Ahas eigenvalue λ, Orthogonal A†has eigenvalue λ∗AandA†have the same eigenvectors 3.5.34 For a Hermitian n×nmatrixAwith distinct eigenvalues λjand a function f,s h o w thatthespectraldecompositionlawmaybeexpressedas f(A)=nsummationdisplay j=1f(λj)producttext i/negationslash=j(A−λi) producttext i/negationslash=j(λj−λi). Thisformulais duetoSylvester. 3.6 N ORMAL MATRICES In Section 3.5 we concentrated primarily on Hermitian or real symmetric matrices and on the actual process of finding the eigenvalues and eigenvectors. In this section27we generalize to normal matrices, with Hermitian and unitary matrices as special cases. The physicallyimportantproblemofnormalmodesofvibrationandthenumericallyimportant problemofill-conditionedmatricesarealsoconsidered. Anormalmatrixis amatrixthatcommuteswithitsadjoint, bracketleftbig A,A†bracketrightbig =0. ObviousandimportantexamplesareHermitianandunitarymatrices.Wewillshowthat normalmatriceshaveorthogonaleigenvectors(see Table3.1). Weproceedintwo steps. I.L etAhaveaneigenvector |x/angbracketrightandcorrespondingeigenvalue λ.Then A|x/angbracketright=λ|x/angbracketright (3.173) or (A−λ1)|x/angbracketright=0. (3.174) For convenience the combination A−λ1 will be labeled B. Taking the adjoint of Eq.(3.174), weobtain /angbracketleftx|(A−λ1)†=0=/angbracketleftx|B†. (3.175) Because bracketleftbig (A−λ1)†,(A−λ1)bracketrightbig =bracketleftbig A,A†bracketrightbig =0, 27Normalmatricesarethelargestclassofmatricesthatcanbediagonalizedbyunitarytransformations.Foranextensivediscus- sion ofnormal matrices, seeP. A.Macklin,Normal matrices for physicists. Am.J.Phys. 52: 513 (1984). 232 Chapter 3 Determinants and Matrices wehave bracketleftbig B,B†bracketrightbig =0. (3.176) Thematrix Bisalsonormal. From Eqs. (3.174)and(3.175)weform /angbracketleftx|B†B|x/angbracketright=0. (3.177) Thisequals /angbracketleftx|BB†|x/angbracketright=0 (3.178) byEq.(3.176). NowEq. (3.178)mayberewrittenas parenleftbig B†|x/angbracketrightparenrightbig†parenleftbig B†|x/angbracketrightparenrightbig =0. (3.179) Thus B†|x/angbracketright=parenleftbig A†−λ∗1parenrightbig |x/angbracketright=0. (3.180) We see that for normal matrices, A†has the same eigenvectors as Abut the complex con- jugateeigenvalues. II. Now,consideringmorethanoneeigenvector–eigenvalue,wehave A|xi/angbracketright=λi|xi/angbracketright, (3.181) A|xj/angbracketright=λj|xj/angbracketright. (3.182) MultiplyingEq. (3.182)fromtheleftby /angbracketleftxi|yields /angbracketleftxi|A|xj/angbracketright=λj/angbracketleftxi|xj/angbracketright. (3.183) Takingthetransposeof Eq.(3.181), weobtain /angbracketleftxi|A=parenleftbig A†|xi/angbracketrightparenrightbig†. (3.184) From Eq. (3.180), with A†having the same eigenvectors as Abut the complex conjugate eigenvalues, parenleftbig A†|xi/angbracketrightparenrightbig†=parenleftbig λ∗ i|xi/angbracketrightparenrightbig†=λi/angbracketleftxi|. (3.185) SubstitutingintoEq.(3.183) wehave λi/angbracketleftxi|xj/angbracketright=λj/angbracketleftxi|xj/angbracketright or (λi−λj)/angbracketleftxi|xj/angbracketright=0. (3.186) Thisis thesameas Eq.(3.149). Forλi/negationslash=λj, /angbracketleftxj|xi/angbracketright=0. Theeigenvectorscorrespondingtodifferenteigenvaluesofanormalmatrixare orthogo- nal.Thismeansthatanormalmatrixmaybediagonalizedbyaunitarytransformation.The required unitary matrix may be constructed from the orthonormal eigenvectors as shown earlier,inSection3.5. The converse of this result is also true. If Acan be diagonalized by a unitary transfor- mation,then Ais normal. 3.6 Normal Matrices 233 Normal Modes of Vibration WeconsiderthevibrationsofaclassicalmodeloftheCO 2molecule.Itisanillustrationof theapplicationofmatrixtechniquestoaproblemthatdoesnotstartasamatrixproblem.It alsoprovidesanexampleoftheeigenvaluesandeigenvectorsofanasymmetricrealmatrix. Example 3.6.1 NORMAL MODES Consider three masses on the x-axis joined by springs as shown in Fig. 3.7. The spring forces are assumed to be linear (small displacements, Hooke’s law), and the mass is con- strainedtostayonthe x-axis. Usingadifferentcoordinateforeachmass, Newton’ssecondlawyieldsthesetofequa- tions ¨x1=−k M(x1−x2) ¨x2=−k m(x2−x1)−k m(x2−x3) (3.187) ¨x3=−k M(x3−x2). The system of masses is vibrating. We seek the common frequencies, ω, such that all massesvibrateatthis samefrequency.Thesearethe normalmodes.Let xi=xi0eiωt,i=1,2,3. Substitutingthisset intoEq. (3.187),wemayrewriteitas  k M−k M0 −k m2k m−k m 0−k Mk M  x1 x2 x3=+ω2x1 x2 x3, (3.188) with the common factor eiωtdivided out. We have a matrix–eigenvalue equation with the matrixasymmetric.Thesecularequationis vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglek M−ω2−k M0 −k m2k m−ω2−k m 0−k Mk M−ω2vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (3.189) FIGURE 3.7Doubleoscillator. 234 Chapter 3 Determinants and Matrices Thisleadsto ω2parenleftbiggk M−ω2parenrightbiggparenleftbigg ω2−2k m−k Mparenrightbigg =0. Theeigenvaluesare ω2=0,k M,k M+2k m, allreal. Thecorrespondingeigenvectorsaredeterminedbysubstitutingtheeigenvaluesbackinto Eq.(3.188) oneeigenvalueatatime.For ω2=0,Eq.(3.188), yields x1−x2=0,−x1+2x2−x3=0,−x2+x3=0. Thenweget x1=x2=x3. Thisdescribespuretranslationwithnorelativemotionof themasses andnovibration. Forω2=k/M,Eq. (3.188)yields x1=−x3,x 2=0. Thetwooutermassesaremovinginoppositedirection.Thecentralmass isstationary. Forω2=k/M+2k/m, theeigenvectorcomponentsare x1=x3,x 2=−2M mx1. Thetwooutermassesaremovingtogether.Thecentralmassismovingoppositetothetwo outerones. Thenetmomentumis zero. Any displacement of the three masses along the x-axis can be described as a linear combinationof thesethreetypesof motion:translationplustwo formsof vibration. /squaresolid Ill-Conditioned Systems Asystemofsimultaneouslinearequationsmaybewrittenas A|x/angbracketright=|y/angbracketrightorA−1|y/angbracketright=|x/angbracketright, (3.190) withAand|y/angbracketrightknown and|x/angbracketrightunknown. When a small error in |y/angbracketrightresults in a larger error in|x/angbracketright,thenthematrix Aiscalledill-conditioned .With|δx/angbracketrightanerrorin|x/angbracketrightand|δx/angbracketrightanerror in|y/angbracketright, therelativeerrors maybewrittenas bracketleftbigg/angbracketleftδx|δx/angbracketright /angbracketleftx|x/angbracketrightbracketrightbigg1/2 ≤K(A)bracketleftbigg/angbracketleftδy|δy/angbracketright /angbracketlefty|y/angbracketrightbracketrightbigg1/2 . (3.191) HereK(A),apropertyofmatrix A,islabeledthe conditionnumber .ForAHermitianone formoftheconditionnumberisgivenby28 K(A)=|λ|max |λ|min. (3.192) 28G.E.Forsythe,andC.B.Moler, ComputerSolutionofLinearAlgebraicSystems .EnglewoodCliffs,NJ,PrenticeHall(1967). 3.6 Normal Matrices 235 AnapproximateformduetoTuring29is K(A)=n[Aij]maxbracketleftbig A−1 ijbracketrightbig max, (3.193) inwhich nistheorderof thematrixand [Aij]maxis themaximumelementin A. Example 3.6.1 ANILL-CONDITIONED MATRIX Acommonexampleofanill-conditionedmatrixistheHilbertmatrix, Hij=(i+j−1)−1. The Hilbert matrix of order 4, H4, is encountered in a least-squares fit of data to a third- degreepolynomial.Wehave H4= 11 21 31 4 1 21 31 41 5 1 31 41 51 6 1 41 51 61 7 . (3.194) Theelementsoftheinversematrix(order n)aregi v enby parenleftbig H−1 nparenrightbig ij=(−1)i+j i+j−1·(n+i−1)!(n+j−1)! [(i−1)!(j−1)!]2(n−i)!(n−j)!.(3.195) Forn=4, H−1 4= 16−120 240 −140 −120 1200 −2700 1680 240−2700 6480 −4200 −140 1680 −4200 2800 . (3.196) FromEq. (3.193)theTuringestimateof theconditionnumberfor H4becomes KTuring=4×1×6480 =2.59×104. This is a warning that an input error may be multiplied by 26,000 in the calculation of the output result. It is a statement that H4is ill-conditioned. If you encounter a highly ill-conditionedsystem, youhavetwoalternatives(besidesabandoningtheproblem). (a) Tryadifferentmathematicalattack. (b) Arrangetocarrymoresignificantfiguresandpushthroughbybruteforce. As previously seen, matrix eigenvector–eigenvalue techniques are not limited to the so- lutionofstrictlymatrixproblems.Afurtherexampleofthetransferoftechniquesfromone area to another is seen in the application of matrix techniques to the solution of Fredholm eigenvalue integral equations, Section 16.3. In turn, these matrix techniques are strength- enedbyavariationalcalculationofSection17.8. /squaresolid 29CompareJ.Todd, TheConditionoftheFiniteSegmentsoftheHilbertMatrix ,AppliedMathematicsSeriesNo.313.Washing- ton, DC: National Bureau of Standards. 236 Chapter 3 Determinants and Matrices Exercises 3.6.1 Showthatevery2 ×2matrixhastwoeigenvectorsandcorrespondingeigenvalues.The eigenvectorsarenotnecessarilyorthogonalandmaybedegenerate.Theeigenvaluesare notnecessarilyreal. 3.6.2 AsanillustrationofExercise3.6.1,findtheeigenvaluesandcorrespondingeigenvectorsfor parenleftbigg24 12parenrightbigg . Notethattheeigenvectorsare notorthogonal. ANS.λ1=0,r1=(2,−1); λ2=4,r2=(2,1). 3.6.3 IfAis a 2×2 matrix,showthatitseigenvalues λsatisfy thesecularequation λ2−λtrace(A)+detA=0. 3.6.4 Assuming a unitary matrix Uto satisfy an eigenvalue equation Ur=λr, show that the eigenvalues of the unitary matrix have unit magnitude. This same result holds for real orthogonalmatrices. 3.6.5 Since an orthogonal matrix describing a rotation in real three-dimensional space is a special case of a unitary matrix, such an orthogonal matrix can be diagonalized by a unitarytransformation. (a) Showthatthesumofthethreeeigenvaluesis 1 +2cosϕ,whereϕisthenetangle ofrotationaboutasinglefixedaxis. (b) Given that one eigenvalue is 1, show that the other two eigenvalues must be eiϕ ande−iϕ. Ourorthogonalrotationmatrix(real elements)hascomplexeigenvalues. 3.6.6 Ais annth-order Hermitian matrix with orthonormal eigenvectors |xi/angbracketrightand real eigen- valuesλ1≤λ2≤λ3≤···≤λn.Showthatfor aunitmagnitudevector |y/angbracketright, λ1≤/angbracketlefty|A|y/angbracketright≤λn. 3.6.7 AparticularmatrixisbothHermitianandunitary.Showthatitseigenvaluesareall ±1. Note.ThePauliandDiracmatricesarespecificexamples. 3.6.8 ForhisrelativisticelectrontheoryDiracrequiredasetof fouranticommutingmatrices. AssumethatthesematricesaretobeHermitianandunitary.Iftheseare n×nmatrices, showthat nmustbeeven.With2 ×2matricesinadequate(why?),thisdemonstratesthat the smallest possible matrices forming a set of four anticommuting, Hermitian, unitary matricesare 4×4. 3.6 Normal Matrices 237 3.6.9 Aisanormalmatrixwitheigenvalues λnandorthonormaleigenvectors |xn/angbracketright.Showthat Amaybewrittenas A=summationdisplay nλn|xn/angbracketright/angbracketleftxn|. Hint.Show that both this eigenvectorform of Aand the original Agive the same result actingonanarbitraryvector |y/angbracketright. 3.6.10 Ahas eigenvalues1and −1 andcorrespondingeigenvectorsparenleftbig1 0parenrightbig andparenleftbig01parenrightbig . Construct A. ANS.A=parenleftbigg10 0−1parenrightbigg . 3.6.11 Anon-Hermitianmatrix Ahaseigenvalues λiandcorrespondingeigenvectors |ui/angbracketright.The adjoint matrix A†has the same set of eigenvalues but different corresponding eigen- vectors,|vi/angbracketright.Showthattheeigenvectorsforma biorthogonal set,inthesensethat /angbracketleftvi|uj/angbracketright=0forλ∗ i/negationslash=λj. 3.6.12 Youare givenapairofequations: A|fn/angbracketright=λn|gn/angbracketright ˜A|gn/angbracketright=λn|fn/angbracketrightwithAreal. (a) Provethat |fn/angbracketrightis aneigenvectorof (˜AA)witheigenvalue λ2 n. (b) Provethat |gn/angbracketrightisaneigenvectorof (A˜A)witheigenvalue λ2n. (c) Statehowyouknowthat (1) The|fn/angbracketrightformanorthogonalset. (2) The|gn/angbracketrightformanorthogonalset. (3)λ2 nisreal. 3.6.13 Provethat Aof theprecedingexercisemaybewrittenas A=summationdisplay nλn|gn/angbracketright/angbracketleftfn|, withthe|gn/angbracketrightand/angbracketleftfn|normalizedtounity. Hint.Expandyourarbitraryvectorasa linearcombinationof |fn/angbracketright. 3.6.14 Given A=1 √ 5parenleftbigg22 1−4parenrightbigg , (a) Constructthetranspose ˜Aandthesymmetricforms ˜AAandA˜A. (b) From A˜A|gn/angbracketright=λ2 n|gn/angbracketrightfindλnand|gn/angbracketright. Normalizethe |gn/angbracketright. (c) From˜AA|fn/angbracketright=λ2n|gn/angbracketrightfindλn[sameas(b)] and |fn/angbracketright.Normalizethe |fn/angbracketright. (d) Verifythat A|fn/angbracketright=λn|gn/angbracketrightand˜A|gn/angbracketright=λn|fn/angbracketright. (e) Verifythat A=summationtext nλn|gn/angbracketright/angbracketleftfn|. 238 Chapter 3 Determinants and Matrices 3.6.15 Giventheeigenvalues λ1=1,λ2=−1 andthecorrespondingeigenvectors |f1/angbracketright=parenleftbigg1 0parenrightbigg ,|g1/angbracketright=1√ 2parenleftbigg1 1parenrightbigg ,|f2/angbracketright=parenleftbigg01parenrightbigg ,and|g2/angbracketright=1 √ 2parenleftbigg1 −1parenrightbigg , (a) construct A; (b) verifythat A|fn/angbracketright=λn|gn/angbracketright; (c) verifythat ˜A|gn/angbracketright=λn|fn/angbracketright. ANS.A=1√ 2parenleftbigg1−1 11parenrightbigg . 3.6.16 ThisisacontinuationofExercise3.4.12,wheretheunitarymatrix UandtheHermitian matrixHare relatedby U=eiaH. (a) If trace H=0,showthat det U=+1. (b) If det U=+1,showthattrace H=0. Hint.Hmay be diagonalized by a similarity transformation. Then interpreting the ex- ponentialbyaMaclaurinexpansion, Uisalsodiagonal.Thecorrespondingeigenvalues aregivenby uj=exp(iahj). Note.Theseproperties,andthoseofExercise3.4.12,arevitalinthedevelopmentofthe conceptofgeneratorsingrouptheory—Section4.2. 3.6.17 Ann×nmatrixAhasneigenvalues Ai.I fB=eA, show that Bhas the same eigen- vectorsas A, withthecorrespondingeigenvalues Bigivenby Bi=exp(Ai). Note.eAis definedbytheMaclaurinexpansionoftheexponential: eA=1+A+A2 2!+A3 3!+···. 3.6.18 AmatrixPisaprojectionoperator(seethediscussionfollowingEq.(3.138c))satisfying thecondition P2=P. Showthatthecorrespondingeigenvalues (ρ2)λandρλsatisfy therelation parenleftbig ρ2parenrightbig λ=(ρλ)2=ρλ. Thismeansthattheeigenvaluesof Pare 0and1. 3.6.19 Inthematrixeigenvector–eigenvalueequation A|ri/angbracketright=λi|ri/angbracketright, Ais ann×nHermitian matrix. For simplicity assume that its nreal eigenvalues are distinct,λ1beingthelargest.If |r/angbracketrightisanapproximationto |r1/angbracketright, |r/angbracketright=|r1/angbracketright+nsummationdisplay i=2δi|ri/angbracketright, 3.6 Additional Readings 239 FIGURE 3.8Tripleoscillator. showthat /angbracketleftr|A|r/angbracketright /angbracketleftr|r/angbracketright≤λ1 andthattheerror in λ1isof theorder|δi|2.T ak e|δi|≪1. Hint.Then|ri/angbracketrightformacomplete orthogonalsetspanningthe n-dimensional(complex) space. 3.6.20 Two equal masses are connected to each other and to walls by springs as shown in Fig.3.8. Themassesareconstrainedtostayonahorizontalline. (a) SetuptheNewtonianaccelerationequationfor eachmass. (b) Solvethesecularequationfortheeigenvectors. (c) Determinetheeigenvectorsandthusthenormalmodesof motion. 3.6.21 Given a normal matrix Awith eigenvalues λj,show that A†has eigenvalues λ∗ j,its real part (A+A†)/2 has eigenvalues ℜ(λj), and its imaginary part (A−A†)/2ihas eigenvaluesℑ(λj). AdditionalReadings Aitken,A.C., DeterminantsandMatrices .NewYork:Interscience(1956).Reprinted,Greenwood(1983).Aread- able introduction to determinants and matrices. Barnett, S., Matrices:Methods and Applications . Oxford: Clarendon Press (1990). Bickley, W. G., and R. S. H. G. Thompson, Matrices—Their Meaning and Manipulation . Princeton, NJ: Van Nostrand (1964). A comprehensive account of matrices in physical problems, their analytic properties, and numerical techniques. Brown, W. C., Matrices and VectorSpaces . NewYork: Dekker(1991). Gilbert,J. andL., Linear Algebra and MatrixTheory . San Diego: AcademicPress (1995). Heading,J., MatrixTheoryforPhysicists .London:Longmans,GreenandCo.(1958).Areadableintroductionto determinants and matrices, with applications to mechanics, electromagnetism, special relativity, and quantum mechanics. Vein,R.,andP. Dale, Determinants and Their Applications in Mathematical Physics .Berlin: Springer (1998). Watkins,D. S., Fundamentals of Matrix Computations . NewYork: Wiley (1991). This page intentionally left blank CHAPTER 4 GROUP THEORY Disciplined judgment, about what is neat and symmetrical and elegant hastime and time again provedan excellent guide to how nature works MURRAYGELL-MANN 4.1 I NTRODUCTION TO GROUP THEORY In classical mechanics the symmetry of a physical system leads to conservation laws . Conservationofangularmomentumisadirectconsequenceofrotationalsymmetry,which meansinvariance underspatialrotations.Inthefirstthirdofthe20thcentury,Wignerand others realized that invariance was a key concept in understanding the new quantum phe- nomena and in developing appropriate theories. Thus, in quantum mechanics the concept ofangularmomentumandspinhasbecomeevenmorecentral.Itsgeneralizations, isospin in nuclear physics and the flavor symmetry in particle physics, are indispensable tools in building and solving theories. Generalizations of the concept of gauge invariance of classicalelectrodynamicstotheisospinsymmetryleadtotheelectroweakgaugetheory. In each case the set of these symmetry operations forms a group. Group theory is the mathematicaltooltotreatinvariantsandsymmetries.Itbringsunificationandformalization of principles, such as spatial reflections, or parity, angular momentum, and geometry, that arewidelyused byphysicists. In geometry the fundamental role of group theory was recognized more than a cen- turyagobymathematicians(e.g.,FelixKlein’sErlangerProgram).InEuclideangeometry the distance between two points, the scalar product of two vectors or metric, does not change under rotations or translations. These symmetries are characteristic of this geom- etry. In special relativity the metric, or scalar product of four-vectors, differs from that of 241 242 Chapter 4 Group Theory Euclidean geometry in that it is no longer positive definite and is invariant under Lorentz transformations. For a crystal the symmetry group contains only a finite number of rotations at discrete values of angles or reflections. The theory of such discreteorfinitegroups, developed originally as a branch of pure mathematics, now is a useful tool for the development of crystallography and condensed matter physics. A brief introduction to this area appears in Section 4.7. When the rotations depend on continuously varying angles (the Euler angles of Section 3.3) the rotation groups have an infinite number of elements. Such continuous (orLie1)groupsare the topic of Sections 4.2–4.6. In Section 4.8 we give an introduction to differential forms, with applications to Maxwell’s equations and topics of Chapters 1 and2, whichallowsseeingthesetopicsfrom adifferentperspective. Definition of a Group A group Gmay be defined as a set of objects or operations, rotations, transformations, called the elements of G, that may be combined, or “multiplied,” to form a well-defined productin G, denotedbya*, thatsatisfiesthefollowingfour conditions. 1. Ifaandbare any two elements of G, then the product a∗bis also an element of G, wherebacts before a;o r(a,b)→a∗bassociates (or maps) an element a∗bofG with the pair (a,b)of elements of G. This property is known as “ Gis closed under multiplicationofits ownelements.” 2. This multiplicationis associative: (a∗b)∗c=a∗(b∗c). 3. There is a unit element21i nGsuch that 1∗a=a∗1=afor every element ainG. Theunitisunique: 1 =1′∗1=1′. 4. There is an inverse, or reciprocal, of each element aofG, labeled a−1, such that a∗a−1=a−1∗a=1. The inverse is unique: If a−1anda′−1are both inverses of a, thena′−1=a′−1∗(a∗a′−1)=(a′−1∗a)∗a−1=a−1. Sincethe*formultiplicationistedioustowrite,itiscustomarytodropitandsimplyletit beunderstood.Fromnowon, wewrite abinsteadof a∗b. •If a subset G′ofGis closed under multiplication, it is a group and called a subgroup ofG; thatis,G′is closedunderthe multiplicationof G. The unitof Galways forms a subgroupof G. •Ifgg′g−1is an element of G′for anygofGandg′ofG′, thenG′is called an in- variant subgroup ofG. The subgroup consisting of the unit is invariant. If the group elements are square matrices, then gg′g−1corresponds to a similarity transformation (seeEq. (3.100)). •Ifab=bafor alla,bofG, the group is called abelian, that is, the order in products doesnotmatter;commutativemultiplicationisoftendenotedbya +sign.Examplesare vector spaces whose unit is the zero vector and −ais the inverse of afor all elements ainG. 1Afterthe NorwegianmathematicianSophus Lie. 2Following E. Wigner, the unit element of a group is often labeled E, from the German Einheit, that is, unit, or just 1, or Ifor identity. 4.1 Introduction to Group Theory 243 Example 4.1.1 ORTHOGONAL AND UNITARY GROUPS Orthogonal n×nmatrices form the group O(n), andSO(n)if their determinants are +1 (Sstands for “special”). If ˜Oi=O−1 ifori=1 and 2 (see Section 3.3 for orthogonal matrices)areelementsof O(n),thentheproduct /tildewiderO1O2=˜O2˜O1=O−1 2O−1 1=(O1O2)−1 is also an orthogonal matrix in O(n), thus proving closure under (matrix) multiplication. Theinverseisthetranspose(orthogonal)matrix.Theunitofthegroupisthe n-dimensional unit matrix 1 n. A real orthogonal n×nmatrix has n(n−1)/2 independent parameters. Forn=2, there is only one parameter: one angle. For n=3, there are three independent parameters:thethreeEuler anglesof Section3.3. If˜Oi=O−1 i(fori=1 and 2) are elements of SO(n), then closure requires proving in additionthattheirproducthasdeterminant +1,whichfollowsfromtheproducttheoremin Chapter3. Likewise, unitary n×nmatrices form the group U(n), andSU(n)if their determinants are+1.IfU† i=U−1 i(see Section3.4 forunitarymatrices)areelementsof U(n),then (U1U2)†=U† 2U†1=U−1 2U−1 1=(U1U2)−1, sotheproductisunitaryandanelementof U(n),thusprovingclosureundermultiplication. Eachunitarymatrixhasaninverse(itsHermitianadjoint),whichagainis unitary. IfU† i=U−1 iare elementsof SU(n), then closure requires us to prove thattheir product alsohasdeterminant +1,whichfollowsfrom theproducttheoreminChapter3. /squaresolid •Orthogonalgroupsarecalled Liegroups ;thatis,theydependoncontinuouslyvarying parameters (the Euler angles and their generalization for higher dimensions); they are compact because the angles vary over closed, finite intervals (containing the limit of any converging sequence of angles). Unitary groups are also compact. Translations formanoncompactgroupbecausethelimitoftranslationswithdistance d→∞isnot partofthegroup.TheLorentzgroupis notcompacteither. Homomorphism, Isomorphism There may be a correspondence between the elements of two groups: one-to-one, two-to- one, or many-to-one. If this correspondence preserves the group multiplication, we say that the two groups are homomorphic . A most important homomorphic correspondence betweentherotationgroup SO(3)andtheunitarygroup SU(2)isdevelopedinSection4.2. If the correspondence is one-to-one, still preserving the group multiplication,3then the groupsare isomorphic . •If a group Gis homomorphic to a group of matrices G′, thenG′is called a represen- tationofG.I fGandG′are isomorphic, the representation is called faithful. There aremanyrepresentationsofgroups;theyare notunique. 3Supposetheelementsofonegrouparelabeled gi,theelementsofasecondgroup hi.Thengi↔hiisaone-to-onecorrespon- dencefor all values of i.I fgigj=gkandhihj=hk,t h e ngkandhkmust be thecorresponding group elements. 244 Chapter 4 Group Theory Example 4.1.2 ROTATIONS Anotherinstructiveexampleforagroupisthesetofcounterclockwisecoordinaterotations ofthree-dimensionalEuclideanspaceaboutits z-axis.FromChapter3weknowthatsucha rotationisdescribedbyalineartransformationofthecoordinatesinvolvinga 3 ×3m a t r i x made up of three rotations depending on the Euler angles. If the z-axis is fixed, the linear transformation is through an angle ϕof thexy-coordinate system to a new orientation in Eq.(1.8), Fig.1.6,andSection3.3:  x′ y′ z′=Rz(ϕ)x y z ≡cosϕsinϕ0 −sinϕcosϕ0 00 1x y z  (4.1) involves only one angle of the rotation about the z-axis. As shown in Chapter 3, the linear transformationoftwosuccessiverotationsinvolvestheproductofthematricescorrespond- ingtothesumoftheangles.Theproductcorrespondstotworotations, Rz(ϕ1)Rz(ϕ2),and is defined by rotating first by the angle ϕ2and then by ϕ1. According to Eq. (3.29), this correspondstotheproductof theorthogonal 2 ×2 submatrices, parenleftBigg cosϕ1sinϕ1 −sinϕ1cosϕ1parenrightBiggparenleftBigg cosϕ2sinϕ2 −sinϕ2cosϕ2parenrightBigg =parenleftBigg cos(ϕ1+ϕ2)sin(ϕ1+ϕ2) −sin(ϕ1+ϕ2)cos(ϕ1+ϕ2)parenrightBigg ,(4.2) using the addition formulas for the trigonometric functions. The unity in the lower right- handcornerofthematrixinEq.(4.1)isalsoreproduceduponmultiplication.Theproductis clearlyarotation,representedbytheorthogonalmatrixwithangle ϕ1+ϕ2.Theassociative group multiplication corresponds to the associative matrix multiplication. It is commuta- tive,orabelian,becausetheorderinwhichtheserotationsareperformeddoesnotmatter. Theinverseoftherotationwithangle ϕisthatwithangle −ϕ.Theunitcorrespondstothe angleϕ=0. Striking off the coordinate vectors in Eq. (4.1), we can associate the matrix of the linear transformation with each rotation, which is a group multiplication preserving one-to-one mapping, an isomorphism: The matrices form a faithful representation of the rotationgroup.Theunityintheright-handcornerissuperfluousaswell,likethecoordinate vectors, and may be deleted. This defines another isomorphism and representation by the 2×2 submatrices: Rz(ϕ)= cosϕsinϕ0 −sinϕcosϕ0 00 1 →R(ϕ)=parenleftBigg cosϕsinϕ −sinϕcosϕparenrightBigg .(4.3) The group’s name is SO(2), if the angle ϕvaries continuously from 0 to 2 π;SO(2) has infinitelymanyelementsandiscompact. Thegroupofrotations RzisobviouslyisomorphictothegroupofrotationsinEq.(4.3). The unity with angle ϕ=0 and the rotation with ϕ=πform a finite subgroup. The finite subgroupswithangles 2 πm/n,n anintegerand m=0,1,...,n−1a r ecyclic;thatis,the rotationsR(2πm/n)=R(2π/n)m. /squaresolid 4.1 Introduction to Group Theory 245 In the following we shall discuss only the rotation groups SO(n)and unitary groups SU(n)among the classical Lie groups. (More examples of finite groups will be given in Section4.7.) Representations — Reducible and Irreducible The representation of group elements by matrices is a very powerful technique and has been almost universally adopted by physicists. The use of matrices imposes no significant restriction. It can be shown that the elements of any finite group and of the continuous groups of Sections 4.2–4.4 may be represented by matrices. Examples are the rotations describedinEq. (4.3). To illustrate how matrix representations arise from a symmetry, consider the station- ary Schrödinger equation (or some other eigenvalue equation, such as Ivi=Iivifor the principalmomentsofinertiaofarigidbodyinclassicalmechanics,say), Hψ=Eψ. (4.4) Let us assume that the Hamiltonian Hstays invariant under a group Gof transformations RinG(coordinate rotations, for example, for a central potential V(r)in the Hamiltonian H); thatis, HR=RHR−1=H,RH=HR. (4.5) Now take a solution ψof Eq. (4.4) and “rotate” it: ψ→Rψ. ThenRψhas thesame energyE becausemultiplyingEq.(4.4) by RandusingEq. (4.5) yields RHψ=E(Rψ)=parenleftbig RHR−1parenrightbig Rψ=H(Rψ). (4.6) Inotherwords,allrotatedsolutions Rψaredegenerate inenergyorformwhatphysicists call amultiplet . For example, the spin-up and -down states of a bound electron in the ground state of hydrogen form a doublet, and the states with projection quantum numbers m=−l,−l+1,...,lof orbital angular momentum lform a multiplet with 2 l+1 basis states. Let us assume that this vector space Vψof transformed solutions has a finite dimen- sionn.L e tψ1,ψ2,...,ψ nbe a basis. Since Rψjis a member of the multiplet, we can expanditintermsof itsbasis, Rψj=summationdisplay krjkψk. (4.7) Thus, with each RinGwe can associate a matrix (rjk). Just as in Example 4.1.2, two successiverotationscorrespondtotheproductoftheirmatrices,sothismap R→(rjk)isa representation of G. It is necessary for a representation to be irreducible that we can take any element of Vψand, by rotating with allelementsRofG, transform it into allother elements of Vψ. If not all elements of Vψare reached, then Vψsplits into a direct sum of two or more vector subspaces, Vψ=V1⊕V2⊕···,which are mapped into themselves by rotating their elements. For example, the 2 sstate and 2 pstates of principal quantum numbern=2ofthehydrogenatomhavethesameenergy(thatis,aredegenerate)andform 246 Chapter 4 Group Theory a reducible representation, because the 2 sstate cannot be rotated into the 2 pstates, and viceversa(angularmomentumisconservedunderrotations).Inthiscasetherepresentation is calledreducible .Then wecanfind abasis in Vψ(that is, there is a unitarymatrix U)s o that U(rjk)U†= r10··· 0r2··· ...... (4.8) forallRofG, andallmatrices (rjk)havesimilarblock-diagonal shape. Here r1,r2,... arematricesoflowerdimensionthan (rjk)thatarelinedupalongthediagonalandthe 0’s are matrices made up of zeros. We may say that the representation has been decomposed intor1+r2+···alongwith Vψ=V1⊕V2⊕···. The irreducible representations play a role in group theory that is roughly analogous to the unit vectors of vector analysis. They are the simplest representations; all others can be builtfromthem.(SeeSection4.4onClebsch–GordancoefficientsandYoungtableaux.) Exercises 4.1.1 Showthatan n×northogonalmatrixhas n(n−1)/2 independentparameters. Hint.Theorthogonalitycondition,Eq.(3.71), providesconstraints. 4.1.2 Showthatan n×nunitarymatrixhas n2−1 independentparameters. Hint.Eachelementmaybecomplex,doublingthenumberofpossibleparameters.Some oftheconstraintequationsarelikewisecomplexandcountastwoconstraints. 4.1.3 The special linear group SL(2) consists of all 2 ×2 matrices (with complex elements) havingadeterminantof +1.Showthatsuchmatricesform agroup. Note.T h eSL(2) group can be related to the full Lorentz group in Section 4.4, much as theSU(2)groupisrelatedto SO(3). 4.1.4 Show that the rotations about the z-axis form a subgroup of SO(3). Is it an invariant subgroup? 4.1.5 Show that if R,S,Tare elements of a group Gso thatRS=TandR→(rik),S→ (sik)is arepresentationaccordingtoEq.(4.7), then (rik)(sik)=parenleftbigg tik=summationdisplay nrinsnkparenrightbigg , that is, group multiplication translates into matrix multiplication for any group repre- sentation. 4.2 G ENERATORS OF CONTINUOUS GROUPS Acharacteristicpropertyofcontinuousgroupsknownas Liegroups isthattheparameters of a product element are analytic functions4of the parameters of the factors. The analytic4Analytichere means having derivatives ofall orders. 4.2 Generators of Continuous Groups 247 natureofthefunctions(differentiability)allowsustodeveloptheconceptofgeneratorand toreducethestudyofthewholegrouptoastudyofthegroupelementsintheneighborhood oftheidentityelement. Lie’s essential idea was to study elements Rin a group Gthat are infinitesimally close to the unity of G. Let us consider the SO(2) group as a simple example. The 2 ×2r o - tation matrices in Eq. (4.2) can be written in exponential form using the Euler identity, Eq.(3.170a), as R(ϕ)=parenleftBigg cosϕsinϕ −sinϕcosϕparenrightBigg =12cosϕ+iσ2sinϕ=exp(iσ2ϕ). (4.9) From the exponential form it is obvious that multiplication of these matrices is equivalent toadditionof thearguments R(ϕ2)R(ϕ1)=exp(iσ2ϕ2)exp(iσ2ϕ1)=expparenleftbig iσ2(ϕ1+ϕ2)parenrightbig =R(ϕ1+ϕ2). Rotationscloseto1havesmallangle ϕ≈0. This suggeststhatwelookfor anexponentialrepresentation R=exp(iεS)=1+iεS+Oparenleftbig ε2parenrightbig ,ε→0, (4.10) for group elements RinGclose to the unity 1. The infinitesimal transformations are εS, and theSare called generators of G. They form a linear space because multiplication of the group elements Rtranslates into addition of generators S. The dimension of this vectorspace(overthe complexnumbers) is the orderofG, thatis, thenumberof linearly independentgeneratorsofthegroup. IfRis a rotation, it does not change the volume element of the coordinate space that it rotates,thatis, det (R)=1,andwemayuseEq. (3.171)toseethat det(R)=expparenleftbig trace(lnR)parenrightbig =expparenleftbig iεtrace(S)parenrightbig =1 impliesεtrace(S)=0 and, upon dividing by the small but nonzero parameter ε, thatgen- eratorsaretraceless , trace(S)=0. (4.11) Thisis thecasenotonlyfor therotationgroups SO(n)butalsofor unitarygroups SU(n). IfRofGin Eq. (4.10) is unitary, then S†=Sis Hermitian, which is also the case for SO(n)andSU(n). Thisexplainswhytheextra ihasbeeninsertedinEq. (4.10). Next we go around the unity in four steps, similar to parallel transport in differential geometry.Weexpandthegroupelements Ri=exp(iεiSi)=1+iεiSi−1 2ε2 iS2 i+···, R−1 i=exp(−iεiSi)=1−iεiSi−1 2ε2 iS2 i+···,(4.12) to second order in the small group parameter εibecause the linear terms and several quadratictermsallcancelintheproduct(Fig. 4.1) R−1 iR−1 jRiRj=1+εiεj[Sj,Si]+···, =1+εiεjsummationdisplay kck jiSk+···, (4.13) 248 Chapter 4 Group Theory FIGURE 4.1IllustrationofEq. (4.13). when Eq. (4.12) is substituted into Eq. (4.13). The last line holds because the product in Eq. (4.13) is again a group element, Rij, close to the unity in the group G. Hence its exponent must be a linear combination of the generators Sk, and its infinitesimal group parameter has to be proportional to the product εiεj. Comparing both lines in Eq. (4.13) wefindthe closurerelationof thegeneratorsoftheLiegroup G, [Si,Sj]=summationdisplay kck ijSk. (4.14) The coefficients ck ijare the structure constants of the group G. Since the commutator in Eq.(4.14) isantisymmetricin iandj, so arethestructureconstantsinthelowerindices, ck ij=−ck ji. (4.15) If the commutator in Eq. (4.14) is taken as a multiplication law of generators, we see thatthevectorspaceofgeneratorsbecomesanalgebra,the Liealgebra Gofthegroup G. Analgebrahastwogroupstructures,acommutativeproductdenotedbya +symbol(this istheadditionofinfinitesimalgeneratorsofaLiegroup)andamultiplication(thecommu- tatorofgenerators).Oftenanalgebraisavectorspacewithamultiplication,suchasaring ofsquarematrices.For SU(l+1)theLiealgebraiscalled Al,forSO(2l+1)itisBl,and forSO(2l)itisDl,wherel=1,2,...isapositiveinteger,latercalledthe rankoftheLie groupGorof itsalgebra G. Finally,the Jacobiidentity holdsfor alldoublecommutators bracketleftbig [Si,Sj],Skbracketrightbig +bracketleftbig [Sj,Sk],Sibracketrightbig +bracketleftbig [Sk,Si],Sjbracketrightbig =0, (4.16) whichiseasilyverifiedusingthedefinitionofanycommutator [A,B]≡AB−BA.When Eq.(4.14) issubstitutedintoEq. (4.16)wefindanotherconstraintonstructureconstants, summationdisplay mbraceleftbig cm ij[Sm,Sk]+cm jk[Sm,Si]+cm ki[Sm,Sj]bracerightbig =0. (4.17) UponinsertingEq. (4.14) again,Eq. (4.17) impliesthat summationdisplay mnbraceleftbig cm ijcn mkSn+cm jkcn miSn+cm kicn mjSnbracerightbig =0, (4.18) 4.2 Generators of Continuous Groups 249 wherethecommonfactor Sn(andthesumover n)maybedroppedbecausethegenerators arelinearlyindependent.Hence summationdisplay mbraceleftbig cm ijcn mk+cm jkcn mi+cm kicn mjbracerightbig =0. (4.19) The relations (4.14), (4.15), and (4.19) form the basis of Lie algebras from which finite elementsoftheLiegroupnearits unitycanbereconstructed. ReturningtoEq.(4.5),theinverseof RisR−1=exp(−iεS).Weexpand HRaccording totheBaker–Hausdorff formula,Eq.(3.172), H=HR=exp(iεS)Hexp(−iεS)=H+iε[S,H]−1 2ε2bracketleftbig S[S,H]bracketrightbig +··· (4.20) We drop Hfrom Eq. (4.20), divide by the small (but nonzero), ε, and let ε→0. Then Eq.(4.20) impliesthatthecommutator [S,H]=0. (4.21) IfSandHareHermitianmatrices,Eq.(4.21)impliesthat SandHcanbesimultaneously diagonalized and have common eigenvectors (for matrices, see Section 3.5; for operators, see Schur’s lemma in Section 4.3). If SandHare differential operators like the Hamil- tonian and orbital angular momentum in quantum mechanics, then Eq. (4.21) implies that SandHhave common eigenfunctions and that the degenerate eigenvalues of Hcan be distinguished by the eigenvalues of the generators S. These eigenfunctions and eigenval- ues,s, are solutions of separate differential equations, Sψs=sψs, so group theory (that is, symmetries) leads to a separation of variables for a partial differential equation that is invariantunderthetransformationsofthegroup. Forexample,letus takethesingle-particleHamiltonian H=−¯h2 2m1 r2∂ ∂rr2∂ ∂r+¯h2 2mr2L2+V(r) that is invariant under SO(3) and, therefore, a function of the radial distance r, the radial gradient, and the rotationally invariant operator L2ofSO(3). Upon replacing the orbital angularmomentumoperator L2byitseigenvalue l(l+1)weobtaintheradialSchrödinger equation(ODE), HRl(r)=bracketleftbigg −¯h2 2m1 r2d drr2d dr+¯h2l(l+1) 2mr2+V(r)bracketrightbigg Rl(r)=ElRl(r), whereRl(r)istheradialwavefunction. For cylindrical symmetry, the invariance of Hunder rotations about the z-axis would requireHtobeindependentoftherotationangle ϕ, leadingtotheODE HRm(z,ρ)=EmRm(z,ρ), withmtheeigenvalueof Lz=−i∂/∂ϕ,thez-componentoftheorbitalangularmomentum operator. For more examples, see the separation of variables method for partial differen- tial equations in Section 9.3 and special functions in Chapter 12. This is by far the most importantapplicationofgrouptheoryinquantummechanics. In the next subsections we shall study orthogonal and unitary groups as examples to understandbetterthegeneralconceptsof thissection. 250 Chapter 4 Group Theory Rotation Groups SO(2) andSO(3) ForSO(2)asdefinedbyEq.(4.3)thereisonlyonelinearlyindependentgenerator, σ2,and the order of SO( 2 )i s1 .W eg e t σ2from Eq. (4.9) by differentiation at the unity of SO(2), thatis,ϕ=0, −idR(ϕ)/dϕ|ϕ=0=−iparenleftBigg −sinϕcosϕ −cosϕ−sinϕparenrightBiggvextendsinglevextendsinglevextendsinglevextendsingle ϕ=0=−iparenleftBigg 01 −10parenrightBigg =σ2.(4.22) For the rotations Rz(ϕ)about the z-axis described by Eq. (4.1), the generator is given by −idRz(ϕ)/dϕ|ϕ=0=Sz= 0−i0 i00 000 , (4.23) wherethefactor iisinsertedtomake SzHermitian.Therotation Rz(δϕ)throughaninfin- itesimalangle δϕmaythenbeexpandedtofirst orderinthesmall δϕas Rz(δϕ)=13+iδϕSz. (4.24) Afiniterotation R(ϕ)maybecompoundedofsuccessiveinfinitesimalrotations Rz(δϕ1+δϕ2)=(1+iδϕ1Sz)(1+iδϕ2Sz). (4.25) Letδϕ=ϕ/NforNrotations,with N→∞. Then Rz(ϕ)=lim N→∞bracketleftbig 1+(iϕ/N)SzbracketrightbigN=exp(iϕSz). (4.26) This form identifies Szas the generator of the group Rz, an abelian subgroup of SO(3), the group of rotations in three dimensions with determinant +1. Each 3×3m a t r i xRz(ϕ) isorthogonal,henceunitary,and trace (Sz)=0,inaccordwithEq.(4.11). Bydifferentiationof thecoordinaterotations Rx(ψ)= 10 0 0 cosψsinψ 0−sinψcosψ ,Ry(θ)= cosθ0−sinθ 01 0 sinθ0 cosθ ,(4.27) wegetthegenerators Sx= 00 0 00−i 0i0 ,Sy= 00i 00 0 −i00  (4.28) ofRx(Ry), thesubgroupofrotationsaboutthe x-(y-)axis. 4.2 Generators of Continuous Groups 251 Rotation of Functions and Orbital Angular Momentum In the foregoing discussion the group elements are matrices that rotate the coordinates. Any physical system being described is held fixed. Now let us hold the coordinates fixed and rotate a function ψ(x,y,z) relative to our fixed coordinates. With Rto rotate the coordinates, x′=Rx, (4.29) wedefineRonψby Rψ(x,y,z)=ψ′(x,y,z)≡ψ(x′). (4.30) In words, Roperates on the function ψ, creating a new function ψ′that is numerically equal to ψ(x′), wherex′are the coordinates rotated by R.I fRrotates the coordinates counterclockwise,theeffect of Ris torotatethepatternof thefunction ψclockwise. Returning to Eqs. (4.30) and (4.1), consider an infinitesimal rotation again, ϕ→δϕ. Then,using RzEq.(4.1), weobtain Rz(δϕ)ψ(x,y,z) =ψ(x+yδϕ,y−xδϕ,z). (4.31) Therightsidemaybeexpandedtofirstorder inthesmall δϕtogive Rz(δϕ)ψ(x,y,z) =ψ(x,y,z)−δϕ{x∂ψ/∂y−y∂ψ/∂x}+O(δϕ)2 =(1−iδϕLz)ψ(x,y,z), (4.32) the differential expression in curly brackets being the orbital angular momentum iLz(Ex- ercise1.8.7). Sincearotationof first ϕandthenδϕaboutthe z-axis isgivenby Rz(ϕ+δϕ)ψ=Rz(δϕ)Rz(ϕ)ψ=(1−iδϕLz)Rz(ϕ)ψ, (4.33) wehave(as anoperatorequation) dRz dϕ=lim δϕ→0Rz(ϕ+δϕ)−Rz(ϕ) δϕ=−iLzRz(ϕ). (4.34) Inthisform Eq.(4.34) integratesimmediatelyto Rz(ϕ)=exp(−iϕLz). (4.35) Note thatRz(ϕ)rotates functions (clockwise) relative to fixed coordinates and that Lzis thezcomponent of the orbital angular momentum L. The constant of integration is fixed bytheboundarycondition Rz(0)=1. AssuggestedbyEq.(4.32), Lzisconnectedto Szby Lz=(x,y,z)Sz ∂/∂x ∂/∂y ∂/∂z =−iparenleftbigg x∂ ∂y−y∂ ∂xparenrightbigg , (4.36) soLx,Ly,andLzsatisfy thesamecommutationrelations, [Li,Lj]=iεijkLk, (4.37) asSx,Sy, andSzandyieldthesamestructureconstants iεijkofSO(3). 252 Chapter 4 Group Theory SU(2) —SO(3) Homomorphism Since unitary 2 ×2 matrices transform complex two-dimensional vectors preserving their norm, they represent the most general transformations of (a basis in the Hilbert space of) spin1 2wavefunctionsinnonrelativisticquantummechanics.Thebasisstatesofthissystem areconventionallychosentobe |↑/angbracketright=parenleftbigg1 0parenrightbigg ,|↓/angbracketright=parenleftbigg01parenrightbigg , corresponding to spin1 2up and down states, respectively. We can show that the special unitarygroupSU(2) of unitary 2 ×2 matrices with determinant +1 has all three Pauli matricesσias generators (while the rotations of Eq. (4.3) form a one-dimensional abelian subgroup).So SU(2)isoforder3anddependsonthreerealcontinuousparameters ξ,η,ζ, which are often called the Cayley–Klein parameters. To construct its general element, we start with the observation that orthogonal 2 ×2 matrices are real unitary matrices, so they formasubgroupof SU(2). Wealsoseethat parenleftBigg eiα0 0e−iαparenrightBigg isunitaryforrealangle αwithdeterminant +1.Sothesesimpleandmanifestlyunitaryma- trices form another subgroup of SU(2) from which we can obtain all elements of SU(2), that is, the general 2 ×2 unitary matrix of determinant +1. For a two-component spin1 2 wavefunctionofquantummechanicsthisdiagonalunitarymatrixcorrespondstomultipli- cation of the spin-up wave function with a phase factor eiαand the spin-down component with the inverse phase factor. Using the real angle ηinstead of ϕfor the rotation matrix andthenmultiplyingbythediagonalunitarymatrices,weconstructa 2 ×2 unitarymatrix thatdependsonthreeparametersandclearlyis amoregeneralelementof SU(2): parenleftBigg eiα0 0e−iαparenrightBiggparenleftBigg cosηsinη −sinηcosηparenrightBiggparenleftBigg eiβ0 0e−iβparenrightBigg =parenleftBigg eiαcosηeiαsinη −e−iαsinηe−iαcosηparenrightBiggparenleftBigg eiβ0 0e−iβparenrightBigg =parenleftBigg ei(α+β)cosηei(α−β)sinη −e−i(α−β)sinηe−i(α+β)cosηparenrightBigg . Defining α+β≡ξ,α−β≡ζ,wehaveinfactconstructedthegeneralelementof SU(2): U(ξ,η,ζ)=parenleftBigg eiξcosηeiζsinη −e−iζsinηe−iξcosηparenrightBigg =parenleftBigg ab −b∗a∗parenrightBigg . (4.38) To see this, we write the general SU(2) element as U=parenleftbigab cdparenrightbig with complex numbers a,b,c,d so that det (U)=1. Writing unitarity, U†=U−1, and using Eq. (3.50) for the 4.2 Generators of Continuous Groups 253 inverseweobtainparenleftBigg a∗c∗ b∗d∗parenrightBigg =parenleftBigg d−b −caparenrightBigg , implying c=−b∗,d=a∗, as shown in Eq. (4.38). It is easy to check that the determinant det(U)=1 andthat U†U=1=UU†hold. Togetthegenerators,wedifferentiate(anddropirrelevantoverallfactors): −i∂U/∂ξ|ξ=0,η=0=parenleftBigg 10 0−1parenrightBigg =σ3, (4.39a) −i∂U/∂η|η=0,ζ=0=parenleftBigg 0−i i0parenrightBigg =σ2. (4.39b) To avoid a factor 1 /sinηforη→0 upon differentiating with respect to ζ,w eu s ei n - stead the right-hand side of Eq. (4.38) for Ufor pure imaginary b=iβwithβ→0, so a=radicalbig 1−β2from|a|2+|b|2=a2+β2=1. Differentiating such a U, we get the third generator, −i∂ ∂βparenleftBiggradicalbig 1−β2iβ iβradicalbig 1−β2parenrightBiggvextendsinglevextendsinglevextendsinglevextendsingle β=0=−i −β√ 1−β2i −iβ√ 1−β2 vextendsinglevextendsinglevextendsinglevextendsingle β=0=parenleftBigg 01 10parenrightBigg =σ1. (4.39c) ThePaulimatricesarealltracelessandHermitian. With the Pauli matrices as generators, the elements U1,U2,U3ofSU(2) may be gener- atedby U1=exp(ia1σ1/2),U2=exp(ia2σ2/2),U3=exp(ia3σ3/2).(4.40) The three parameters aiare real. The extra factor 1 /2 is present in the exponents to make Si=σi/2 satisfythesamecommutationrelations, [Si,Sj]=iεijkSk, (4.41) astheangularmomentuminEq. (4.37). To connect and compare our results, Eq. (4.3) gives a rotation operator for rotat- ing the Cartesian coordinates in the three-space R3. Using the angular momentum ma- trixS3, we have as the corresponding rotation operator in two-dimensional (complex) spaceRz(ϕ)=exp(iϕσ3/2).Forrotatingthetwo-componentvectorwavefunction(spinor) or a spin 1 /2 particle relative to fixed coordinates, the corresponding rotation operator is Rz(ϕ)=exp(−iϕσ3/2)accordingtoEq. (4.35). Moregenerally,usinginEq. (4.40) theEuleridentity,Eq.(3.170a), weobtain Uj=cosparenleftbiggaj 2parenrightbigg +iσjsinparenleftbiggaj 2parenrightbigg . (4.42) Heretheparameter ajappearsasanangle,thecoefficientofanangularmomentummatrix- likeϕinEq.(4.26).TheselectionofPaulimatricescorrespondstotheEuleranglerotations describedinSection3.3. 254 Chapter 4 Group Theory FIGURE 4.2Illustrationof M′=UMU†inEq.(4.43). As just seen, the elements of SU(2) describe rotations in a two-dimensional complex spacethatleave |z1|2+|z2|2invariant.Thedeterminantis +1.Therearethreeindependent real parameters. Our real orthogonal group SO(3) clearly describes rotations in ordinary three-dimensional space with the important characteristic of leaving x2+y2+z2invari- ant.Also,therearethreeindependentrealparameters.Therotationinterpretationsandthe equality of numbers of parameters suggest the existence of some correspondence between thegroups SU(2) andSO(3). Herewedevelopthiscorrespondence. Theoperationof SU(2)onamatrixisgivenbyaunitarytransformation,Eq.(4.5),with R=UandFig.4.2: M′=UMU†. (4.43) TakingMto be a 2×2 matrix, we note that any 2 ×2 matrix may be written as a linear combination of the unit matrix and the three Pauli matrices of Section 3.4. Let Mbe the zero-tracematrix, M=xσ1+yσ2+zσ3=parenleftBigg zx−iy x+iy−zparenrightBigg , (4.44) theunitmatrixnotentering.Sincethetraceisinvariantunderaunitarysimilaritytransfor- mation(Exercise3.3.9), M′musthavethesameform, M′=x′σ1+y′σ2+z′σ3=parenleftBigg z′x′−iy′ x′+iy′−z′parenrightBigg . (4.45) The determinant is also invariant under a unitary transformation (Exercise 3.3.10). There- fore −parenleftbig x2+y2+z2parenrightbig =−parenleftbig x′2+y′2+z′2parenrightbig , (4.46) orx2+y2+z2is invariant under this operation of SU(2), just as with SO(3). Operations ofSU(2) onMmust produce rotations of the coordinates x,y,zappearing therein. This suggeststhat SU(2) andSO(3) maybeisomorphicor atleasthomomorphic. 4.2 Generators of Continuous Groups 255 Weapproachtheproblemofwhatthisoperationof SU(2)correspondstobyconsidering specialcases. ReturningtoEq. (4.38),let a=eiξandb=0,or U3=parenleftBigg eiξ0 0e−iξparenrightBigg . (4.47) InanticipationofEq. (4.51), this Uis givenasubscript 3. Carrying out a unitary similarity transformation, Eq. (4.43), on each of the three Pauli σ’s ofSU(2), wehave U3σ1U† 3=parenleftBigg eiξ0 0e−iξparenrightBiggparenleftBigg 01 10parenrightBiggparenleftBigg e−iξ0 0eiξparenrightBigg =parenleftBigg 0e2iξ e−2iξ0parenrightBigg . (4.48) Wereexpressthisresultintermsof thePauli σi,as inEq.(4.44), toobtain U3xσ1U† 3=xσ1cos2ξ−xσ2sin2ξ. (4.49) Similarly, U3yσ2U†3=yσ1sin2ξ+yσ2cos2ξ, U3zσ3U†3=zσ3. (4.50) Fromthesedoubleangleexpressionsweseethatweshouldstartwitha halfangle :ξ=α/2. Then, adding Eqs. (4.49) and (4.50) and comparing with Eqs. (4.44) and(4.45), weobtain x′=xcosα+ysinα y′=−xsinα+ycosα (4.51) z′=z. The 2×2 unitary transformation using U3(α)is equivalent to the rotation operator R(α) ofEq. (4.3). Thecorrespondenceof U2(β)=parenleftBigg cosβ/2sinβ/2 −sinβ/2 cosβ/2parenrightBigg (4.52) andRy(β)andof U1(ϕ)=parenleftBigg cosϕ/2isinϕ/2 isinϕ/2 cosϕ/2parenrightBigg (4.53) andR1(ϕ)followsimilarly.Notethat Uk(ψ)hasthegeneralform Uk(ψ)=12cosψ/2+iσksinψ/2, (4.54) wherek=1,2,3. 256 Chapter 4 Group Theory Thecorrespondence U3(α)=parenleftBigg eiα/20 0e−iα/2parenrightBigg ↔ cosαsinα0 −sinαcosα0 00 1 =Rz(α) (4.55) is not a simple one-to-one correspondence. Specifically, as αinRzranges from 0 to 2 π, theparameterin U3,α/2,goesfrom0to π. Wefind Rz(α+2π)=Rz(α) U3(α+2π)=parenleftBigg −eiα/20 0−e−iα/2parenrightBigg =−U3(α). (4.56) Therefore bothU3(α)andU3(α+2π)=−U3(α)correspond to Rz(α). The correspon- dence is 2 to 1, or SU(2) andSO(3) arehomomorphic . This establishment of the corre- spondencebetweentherepresentationsof SU(2)andthoseof SO(3)meansthattheknown representationsof SU(2) automaticallyprovideus withtherepresentationsof SO(3). Combiningthevariousrotations,wefindthataunitarytransformationusing U(α,β,γ)=U3(γ)U2(β)U3(α) (4.57) correspondstothegeneralEuler rotation Rz(γ)Ry(β)Rz(α). Bydirectmultiplication, U(α,β,γ)=parenleftBigg eiγ/20 0e−iγ/2parenrightBiggparenleftBigg cosβ/2sinβ/2 −sinβ/2 cosβ/2parenrightBiggparenleftBigg eiα/20 0e−iα/2parenrightBigg =parenleftBigg ei(γ+α)/2cosβ/2ei(γ−α)/2sinβ/2 −e−i(γ−α)/2sinβ/2e−i(γ+α)/2cosβ/2parenrightBigg . (4.58) Thisis ouralternategeneralform, Eq.(4.38), with ξ=(γ+α)/2,η=β/2,ζ=(γ−α)/2. (4.59) Thus,from Eq.(4.58) wemayidentifytheparametersofEq. (4.38)as a=ei(γ+α)/2cosβ/2 b=ei(γ−α)/2sinβ/2. (4.60) SU(2)-Isospin and SU(3)-Flavor Symmetry The application of group theory to “elementary” particles has been labeled by Wigner the third stage of group theory and physics. The first stage was the search for the 32 crystallographic point groups and the 230 space groups giving crystal symmetries— Section 4.7. The second stage was a search for representations such as of SO(3) and SU(2)—Section4.2.Nowinthisstage,physicistsarebacktoa searchforgroups. In the 1930s to 1960s the study of strongly interacting particles of nuclear and high- energyphysicsledtothe SU(2)isospingroupandthe SU(3)flavorsymmetry.Inthe1930s, aftertheneutronwasdiscovered,Heisenbergproposedthatthenuclearforceswerecharge 4.2 Generators of Continuous Groups 257 Table 4.1 BaryonswithSpin1 2EvenParity Mass(MeV) YI I 3 /Xi1−1321.32 −1 2 /Xi1 −11 2 /Xi101314.9 +1 2 /Sigma1−1197.43 −1 /Sigma1/Sigma101192.55 0 1 0 /Sigma1+1189.37 +1 /Lambda1/Lambda1 1115.63 0 0 0 n 939.566 −1 2 N 11 2 p 938.272 +1 2 independent. The neutron mass differs from that of the proton by only 1.6%. If this tiny mass difference is ignored, the neutron and proton may be considered as two charge (or isospin) states of a doublet, called the nucleon. The isospin Ihasz-projection I3=1/2 for the proton and I3=−1/2 for the neutron. Isospin has nothing to do with spin (the particle’s intrinsic angular momentum), but the two-component isospin state obeys the same mathematical relations as the spin 1 /2 state. For the nucleon, I=τ/2 are the usual Paulimatricesandthe ±1/2isospinstatesareeigenvectorsofthePaulimatrix τ3=parenleftbig10 0−1parenrightbig . Similarly, the three charge states of the pion ( π+,π0,π−) form a triplet. The pion is the lightest of all strongly interacting particles and is the carrier of the nuclear force at long distances,muchlikethephotonisthatoftheelectromagneticforce.Thestronginteraction treats alike members of these particle families, or multiplets, and conserves isospin. The symmetryisthe SU(2) isospingroup. By the 1960s particles produced as resonances by accelerators had proliferated. The eight shown in Table 4.1 attracted particular attention.5The relevant conserved quantum numbers that are analogs and generalizations of LzandL2fromSO(3) areI3andI2for isospinand Yforhypercharge .Particlesmaybegroupedintochargeorisospinmultiplets. Then the hypercharge may be taken as twice the average charge of the multiplet. For the nucleon, that is, the neutron–proton doublet, Y=2·1 2(0+1)=1. The hypercharge and isospin values are listed in Table 4.1 for baryons like the nucleon and its (approximately degenerate) partners. They form an octet, as shown in Fig. 4.3, after which the corre- sponding symmetry is called the eightfold way . In 1961 Gell-Mann, and independently Ne’eman, suggested that the strong interaction should be (approximately) invariant under athree-dimensionalspecialunitarygroup, SU(3),thatis, has SU(3)flavorsymmetry . The choice of SU(3) was based first on the two conserved and independent quantum numbers, H1=I3andH2=Y(thatis,generatorswith [I3,Y]=0,notCasimirinvariants; see the summary in Section 4.3) that call for a group of rank 2. Second, the group had to have an eight-dimensional representation to account for the nearly degenerate baryons and four similar octets for the mesons. In a sense, SU(3) is the simplest generalization of SU(2)isospin.Threeofitsgeneratorsarezero-traceHermitian3 ×3matricesthatcontain 5Allmasses aregiven inenergyunits, 1MeV =106eV. 258 Chapter 4 Group Theory FIGURE 4.3Baryonoctetweight diagramfor SU(3). the 2×2 isospinPaulimatrices τiintheupperleftcorner, λi= τi0 0 000 ,i=1,2,3. (4.61a) Thus, theSU(2)-isospin group is a subgroup of SU(3)-flavor with I3=λ3/2. Four other generators have the off-diagonal 1’s of τ1, and−i,iofτ2in all other possible locations to form zero-traceHermitian 3 ×3 matrices, λ4= 001 000 100 ,λ 5= 00−i 00 0 i00 , λ6= 000 001 010 ,λ 7= 00 0 00−i 0i0 .(4.61b) The second diagonal generator has the two-dimensional unit matrix 1 2in the upper left corner, which makes it clearly independent of the SU(2)-isospin subgroup because of its nonzerotraceinthatsubspace,and −2 inthethirddiagonalplacetomakeittraceless, λ8=1√ 3 10 0 01 0 00−2 . (4.61c) 4.2 Generators of Continuous Groups 259 FIGURE 4.4Baryonmasssplitting. Altogether there are 32−1=8 generators for SU(3), which has order 8. From the com- mutatorsof thesegeneratorsthestructureconstantsof SU(3) caneasilybeobtained. Returning to the SU(3) flavor symmetry, we imagine the Hamiltonian for our eight baryonstobecomposedofthreeparts: H=Hstrong+Hmedium+Helectromagnetic . (4.62) The first part, Hstrong, has theSU(3) symmetry and leads to the eightfold degeneracy. Introduction of the symmetry-breaking term, Hmedium, removes part of the degeneracy, giving the four isospin multiplets (/Xi1−,/Xi10),(/Sigma1−,/Sigma10,/Sigma1+),/Lambda1, andN=(p,n)different masses. These are still multiplets because HmediumhasSU(2)-isospin symmetry. Finally, the presence of charge-dependent forces splits the isospin multiplets and removes the last degeneracy.ThisimaginedsequenceisshowninFig.4.4. The octet representation is not the simplest SU(3) representation. The simplest repre- sentationsarethetriangularonesshowninFig.4.5,fromwhichallotherscanbegenerated by generalized angular momentum coupling (see Section 4.4 on Young tableaux). The fundamental representation in Fig. 4.5a contains the u(up),d(down), and s(strange) quarks, and Fig. 4.5b contains the corresponding antiquarks. Since the meson octets can be obtained from the quark representations as q¯q, with 32=8+1 states, this suggests thatmesonscontainquarks(andantiquarks)astheirconstituents(seeExercise4.4.3).The resulting quark model gives a successful description of hadronic spectroscopy. The reso- lution of its problem with the Pauli exclusion principle eventually led to the SU(3)-color gaugetheoryofthe stronginteraction calledquantumchromodynamics (QCD). Tokeepgrouptheoryanditsveryrealaccomplishmentinproperperspective,weshould emphasizethatgrouptheoryidentifiesandformalizessymmetries.It classifies(andsome- timespredicts)particles.ButasidefromsayingthatonepartoftheHamiltonianhas SU(2) 260 Chapter 4 Group Theory FIGURE 4.5(a) Fundamentalrepresentationof SU(3),theweightdiagramfor theu,d,squarks;(b)weightdiagramfor theantiquarks ¯u,¯d,¯s. symmetryandanotherparthas SU(3)symmetry,grouptheorysaysnothingaboutthepar- ticleinteraction.Rememberthatthestatementthattheatomicpotentialissphericallysym- metrictellsusnothingabouttheradialdependenceofthepotentialorofthewavefunction. Incontrast,inagaugetheorytheinteractionismediatedbyvectorbosons(likethephoton in quantum electrodynamics) and uniquely determined by the gauge covariant derivative (seeSection1.13). Exercises 4.2.1 (i) Show that the Pauli matrices are the generators of SU(2) without using the para- meterization of the general unitary 2 ×2 matrix in Eq. (4.38). (ii) Derive the eight independent generators λiofSU(3) similarly. Normalize them so that tr (λiλj)=2δij. Thendeterminethestructureconstantsof SU(3). Hint.Theλiare tracelessandHermitian 3 ×3 matrices. (iii)ConstructthequadraticCasimirinvariantof SU(3). Hint.Workbyanalogywith σ2 1+σ2 2+σ2 3ofSU(2) orL2ofSO(3). 4.2.2 Provethatthegeneralform ofa 2 ×2 unitary,unimodularmatrixis U=parenleftBigg ab −b∗a∗parenrightBigg witha∗a+b∗b=1. 4.2.3 Determinethree SU(2) subgroupsof SU(3). 4.2.4 Atranslation operatorT(a)convertsψ(x)toψ(x+a), T(a)ψ(x)=ψ(x+a). 4.3 Orbital Angular Momentum 261 In terms of the (quantum mechanical) linear momentum operator px=−id/dx,s h o w thatT(a)=exp(iapx), thatis,pxisthegeneratorof translations. Hint.Expand ψ(x+a)as aTaylorseries. 4.2.5 Consider the general SU(2) element Eq. (4.38) to be built up of three Euler rotations: (i) a rotation of a/2 about the z-axis, (ii) a rotation of b/2 about the new x-axis, and (iii) a rotation of c/2 about the new z-axis. (All rotations are counterclockwise.) Using thePauli σgenerators,showthattheserotationanglesare determinedby a=ξ−ζ+π 2=α+π 2 b=2η=β c=ξ+ζ−π 2=γ−π 2. Note.Theangles aandbhereare notthe aandbofEq. (4.38). 4.2.6 Rotate a nonrelativistic wave function ˜ψ=(ψ↑,ψ↓)of spin 1/2 about the z-axis by asmallangle dθ.Findthecorrespondinggenerator. 4.3 O RBITAL ANGULAR MOMENTUM The classical concept of angular momentum, Lclass=r×p, is presented in Section 1.4 tointroducethecrossproduct.FollowingtheusualSchrödingerrepresentationofquantum mechanics,theclassicallinearmomentum pisreplacedbytheoperator −i∇.Thequantum mechanicalorbitalangularmomentum operator becomes6 LQM=−ir×∇. (4.63) This is used repeatedly in Sections 1.8, 1.9, and 2.4 to illustrate vector differential oper- ators. From Exercise 1.8.8 the angular momentum components satisfy the commutation relations [Li,Lj]=iεijkLk. (4.64) Theεijkis the Levi-Civita symbol of Section 2.9. A summation over the index kis under- stood. Thedifferentialoperatorcorrespondingtothesquareoftheangularmomentum L2=L·L=L2 x+L2y+L2z (4.65) maybedeterminedfrom L·L=(r×p)·(r×p), (4.66) which is the subject of Exercises 1.9.9 and 2.5.17(b). Since L2as a scalar product is in- variant under rotations, that is, a rotational scalar, we expect [L2,Li]=0, which can also beverifieddirectly. Equation(4.64)presentsthebasiccommutationrelationsofthecomponentsofthequan- tummechanicalangularmomentum.Indeed,withintheframeworkofquantummechanics and group theory, these commutation relations define an angular momentum operator. We shall use them now to construct the angular momentum eigenstates and find the eigenval- ues.FortheorbitalangularmomentumthesearethesphericalharmonicsofSection12.6. 6For simplicity, ¯his setequal to 1. This means that theangular momentum is measured in units of ¯h. 262 Chapter 4 Group Theory Ladder Operator Approach Letusstartwithageneralapproach,wheretheangularmomentum Jweconsidermayrep- resentanorbitalangularmomentum L,aspin σ/2,oratotalangularmomentum L+σ/2, etc.Weassumethat 1.Jis anHermitianoperatorwhosecomponentssatisfythecommutationrelations [Ji,Jj]=iεijkJk,bracketleftbig J2,Jibracketrightbig =0. (4.67) Otherwise Jisarbitrary. (SeeExercise4.3.l.) 2.|λM/angbracketrightissimultaneouslyanormalizedeigenfunction(oreigenvector)of Jzwitheigen- valueMandaneigenfunction7ofJ2, Jz|λM/angbracketright=M|λM/angbracketright,J2|λM/angbracketright=λ|λM/angbracketright,/angbracketleftλM|λM/angbracketright=1.(4.68) We shall show that λ=J(J+1)and then find other properties of the |λM/angbracketright. The treat- mentwillillustratethegeneralityandpowerofoperatortechniques,particularlytheuseof ladderoperators.8 Theladderoperators aredefinedas J+=Jx+iJy,J−=Jx−iJy. (4.69) Intermsof theseoperators J2mayberewrittenas J2=1 2(J+J−+J−J+)+J2 z. (4.70) Fromthecommutationrelations,Eq.(4.67), wefind [Jz,J+]=+J+,[Jz,J−]=−J−,[J+,J−]=2Jz. (4.71) SinceJ+commuteswith J2(Exercise 4.3.1), J2parenleftbig J+|λM/angbracketrightparenrightbig =J+parenleftbig J2|λM/angbracketrightparenrightbig =λparenleftbig J+|λM/angbracketrightparenrightbig . (4.72) Therefore, J+|λM/angbracketrightis still an eigenfunction of J2with eigenvalue λ, and similarly for J−|λM/angbracketright. Butfrom Eq. (4.71), JzJ+=J+(Jz+1), (4.73) or Jzparenleftbig J+|λM/angbracketrightparenrightbig =J+(Jz+1)|λM/angbracketright=(M+1)J+|λM/angbracketright. (4.74) 7That|λM/angbracketrightcan be an eigenfunction of bothJzandJ2follows from[Jz,J2]=0 in Eq. (4.67). For SU(2),/angbracketleftλM|λM/angbracketrightis the scalar product (of the bra and ket vector or spinors) in the bra-ket notation introduced in Section 3.1. For SO(3),|λM/angbracketrightis a functionY(θ,ϕ)and|λM′/angbracketrightisafunction Y′(θ,ϕ)andthematrixelement /angbracketleftλM|λM′/angbracketright≡integraltext2π ϕ=0integraltextπ θ=0Y∗(θ,ϕ)Y′(θ,ϕ)sinθdθdϕ is their overlap. However, in our algebraic approach only the norm in Eq. (4.68) is used and matrix elements of the angular momentumoperatorsarereducedtothenormbymeansoftheeigenvalueequationfor Jz,Eq.(4.68),andEqs.(4.83)and(4.84). 8Ladder operators can be developed for other mathematical functions. Compare the next subsection, on other Lie groups, and Section 13.1, for Hermite polynomials. 4.3 Orbital Angular Momentum 263 Therefore, J+|λM/angbracketrightisstillaneigenfunctionof Jzbutwitheigenvalue M+1.J+hasraised theeigenvalueby1andsoiscalleda raisingoperator .Similarly, J−lowerstheeigenvalue by1andis calleda loweringoperator . Takingexpectationvaluesandusing J† x=Jx,J† y=Jy, weget /angbracketleftλM|J2−J2 z|λM/angbracketright=/angbracketleftλM|J2 x+J2 y|λM/angbracketright=vextendsinglevextendsingleJx|λM/angbracketrightvextendsinglevextendsingle2+vextendsinglevextendsingleJy|λM/angbracketrightvextendsinglevextendsingle2 and see that λ−M2≥0, soMis bounded. Let Jbe thelargestM. ThenJ+|λJ/angbracketright=0, whichimplies J−J+|λJ/angbracketright=0.Hence,combiningEqs. (4.70)and(4.71)toget J2=J−J++Jz(Jz+1), (4.75) wefindfromEq. (4.75) that 0=J−J+|λJ/angbracketright=parenleftbig J2−J2 z−Jzparenrightbig |λJ/angbracketright=parenleftbig λ−J2−Jparenrightbig |λJ/angbracketright. Therefore λ=J(J+1)≥0, (4.76) with nonnegative J. We now relabel the states |λM/angbracketright≡|JM/angbracketright. Similarly, let J′be the smallest M.ThenJ−|JJ′/angbracketright=0.From J2=J+J−+Jz(Jz−1), (4.77) wesee that 0=J+J−|JJ′/angbracketright=parenleftbig J2+Jz−J2 zparenrightbig |JJ′/angbracketright=parenleftbig λ+J′−J′2parenrightbig |JJ′/angbracketright.(4.78) Hence λ=J(J+1)=J′(J′−1)=(−J)(−J−1). SoJ′=−J,andMruns inintegersteps from−Jto+J, −J≤M≤J. (4.79) Startingfrom|JJ/angbracketrightandapplying J−repeatedly,wereachallotherstates |JM/angbracketright.Hencethe |JM/angbracketrightform anirreduciblerepresentationof SO(3)orSU(2);Mvariesand Jis fixed. ThenusingEqs. (4.67), (4.75), and(4.77) weobtain J−J+|JM/angbracketright=bracketleftbig J(J+1)−M(M+1)bracketrightbig |JM/angbracketright=(J−M)(J+M+1)|JM/angbracketright, J+J−|JM/angbracketright=bracketleftbig J(J+1)−M(M−1)bracketrightbig |JM/angbracketright=(J+M)(J−M+1)|JM/angbracketright.(4.80) BecauseJ+andJ−areHermitianconjugates,9 J† +=J−,J† −=J+, (4.81) the eigenvalues in Eq. (4.80) must be positive or zero.10Examples of Eq. (4.81) are pro- videdbythematricesofExercise3.2.13(spin 1 /2),3.2.15(spin1),and3.2.18(spin 3 /2). 9The Hermitian conjugation or adjoint operation is defined for matrices in Section 3.5, and for operators in general in Sec- tion 10.1. 10For an excellent discussion of adjoint operators and Hilbert space see A. Messiah, Quantum Mechanics . New York: Wiley 1961, Chapter 7. 264 Chapter 4 Group Theory For the orbital angular momentum ladder operators, L+, andL−, explicit forms are given inExercises2.5.14and12.6.7. Youcannowshow(see alsoExercise12.7.2)that /angbracketleftJM|J−parenleftbig J+|JM/angbracketrightparenrightbig =parenleftbig J+|JM/angbracketrightparenrightbig†J+|JM/angbracketright. (4.82) SinceJ+raises the eigenvalue MtoM+1, we relabel the resultant eigenfunction |JM+1/angbracketright.The normalizationis givenbyEq. (4.80)as J+|JM/angbracketright=radicalbig (J−M)(J+M+1)|JM+1/angbracketright=radicalbig J(J+1)−M(M+1)|JM+1/angbracketright, (4.83) taking the positive square root and not introducing any phase factor. By the same argu- ments, J−|JM/angbracketright=radicalbig (J+M)(J−M+1)|JM−1/angbracketright=radicalbig (J(J+1)−M(M−1)|JM−1/angbracketright. (4.84) Applying J+toEq.(4.84),weobtainthesecondlineofEq.(4.80)andverifythatEq.(4.84) isconsistentwithEq. (4.83). Finally,since Mrangesfrom−Jto+Jinunitsteps, 2 Jmustbeaninteger; Jiseither an integer or half of an odd integer. As seen later, if Jis an orbital angular momentum L, the set|LM/angbracketrightfor allMis a basis defining a representation of SO(3) andLwill then be integral.Insphericalpolarcoordinates θ,ϕ,thefunctions |LM/angbracketrightbecomethesphericalhar- monicsYM L(θ,ϕ)of Section 12.6. The sets of |JM/angbracketrightstates with half-integral Jdefine rep- resentationsof SU(2)thatarenotrepresentationsof SO(3);weget J=1/2,3/2,5/2,.... Our angular momentum is quantized, essentially as a result of the commutation relations. All these representations are irreducible, as an application of the raising and lowering op- eratorssuggests. Summary of Lie Groups and Lie Algebras The general commutation relations, Eq. (4.14) in Section 4.2, for a classical Lie group [SO(n)andSU(n)in particular] can be simplified to look more like Eq. (4.71) for SO(3) andSU(2) in this section. Here we merely review and, as a rule, do not provide proofs for varioustheoremsthatweexplain. First wechooselinearly independentand mutuallycommutinggenerators Hiwhichare generalizationsof JzforSO(3)andSU(2).Letlbethemaximumnumberofsuch Hiwith [Hi,Hk]=0. (4.85) Thenliscalledthe rankoftheLiegroup GoritsLiealgebra G.Therankanddimension, or order, of some Lie groups are given in Table 4.2. All other generators Eαcan be shown toberaisingandloweringoperatorswithrespecttoallthe Hi,s o [Hi,Eα]=αiEα,i=1,2,...,l. (4.86) Thesetofso-called rootvectors (α1,α2,...,αl)form therootdiagram ofG. When the Hicommute, they can be simultaneously diagonalized (for symmetric (or Hermitian) matrices see Chapter 3; for operators see Chapter 10). The Hiprovide us with asetofeigenvalues m1,m2,...,m l[projectionoradditivequantumnumbersgeneralizing 4.3 Orbital Angular Momentum 265 Table 4.2 Rank and Order of Unitary and Rotational Groups Liealgebra Al Bl Dl Liegroup SU(l+1)SO(2l+1) SO(2l) Rank ll l Order l(l+2)l (2l+1)l (2l−1) MofJzinSO(3) andSU(2)]. The set of so-called weight vectors (m1,m2,...,m l)for anirreduciblerepresentation(multiplet)form a weightdiagram . There are linvariant operators Ci, calledCasimir operators, that commute with all generatorsandaregeneralizationsof J2, [Ci,Hj]=0,[Ci,Eα]=0,i=1,2,...,l. (4.87) Thefirstone, C1,isaquadraticfunctionofthegenerators;theothersaremorecomplicated. Since the Cjcommute with all Hj, they can be simultaneously diagonalized with the Hj. Their eigenvalues c1,c2,...,clcharacterize irreducible representations and stay constant whiletheweightvectorvariesoveranyparticularirreduciblerepresentation.Thusthegen- eraleigenfunctionmaybewrittenas vextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig , (4.88) generalizingthemultiplet |JM/angbracketrightofSO(3)andSU(2). Theireigenvalueequationsare Hivextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =mivextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig (4.89a) Civextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =civextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig .(4.89b) We can now show that Eα|(c1,c2,...,cl)m1,m2,...,m l/angbracketrighthas the weight vector (m1+α1,m2+α2,...,m l+αl)using the commutation relations, Eq. (4.86), in con- junctionwithEqs. (4.89a)and(4.89b): HiEαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =parenleftbig EαHi+[Hi,Eα]parenrightbigvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =(mi+αi)Eαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig . (4.90) Therefore Eαvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig ∼vextendsinglevextendsingle(c1,...,cl)m1+α1,...,m l+αlangbracketrightbig , the generalization of Eqs. (4.83) and (4.84) from SO(3). These changes of eigenvalues by theoperator Eαarecalledits selectionrules inquantummechanics.Theyaredisplayedin therootdiagramofaLiealgebra. Examples of root diagrams are given in Fig. 4.6 for SU(2) andSU(3). If we attach the roots denoted by arrows in Fig. 4.6b to a weight in Figs. 4.3 or 4.5a, b, we can reach any otherstate(representedbyadotintheweightdiagram). HereSchur’s lemma applies: An operator Hthat commutes with all group operators, andthereforewithallgenerators Hiofa(classical)Liegroup Ginparticular,hasaseigen- vectors all states of a multiplet and is degenerate with the multiplet. As a consequence, suchanoperatorcommuteswithallCasimirinvariants, [H,Ci]=0. 266 Chapter 4 Group Theory FIGURE 4.6Rootdiagramfor(a) SU(2) and (b)SU(3). ThelastresultisclearbecausetheCasimirinvariantsareconstructedfromthegenerators andraisingandloweringoperatorsofthegroup.Toprovetherest,let ψbeaneigenvector, Hψ=Eψ. Then, for any rotation RofG,w eh a v e HRψ=ERψ, which says that Rψ is an eigenstate with the same eigenvalue Ealong with ψ. Since[H,Ci]=0, all Casimir invariants can be diagonalized simultaneously with Hand an eigenstate of His an eigen- state of all the Ci. Since[Hi,Ci]=0, the rotated eigenstates Rψare eigenstates of Ci, alongwith ψbelongingtothesamemultipletcharacterizedbytheeigenvalues ciofCi. Finally,suchanoperator Hcannotinducetransitionsbetweendifferentmultipletsofthe groupbecause angbracketleftbig (c′ 1,c′ 2,...,c′ l)m′ 1,m′2,...,m′ lvextendsinglevextendsingleHvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =0. Using[H,Cj]=0 (for any j)weha v e 0=angbracketleftbig (c′ 1,c′ 2,...,c′ l)m′ 1,m′2,...,m′ lvextendsinglevextendsingle[H,Cj]vextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig =(cj−c′ j)angbracketleftbig (c′ 1,c′ 2,...,c′ l)m′ 1,m′2,...,m′ lvextendsinglevextendsingleHvextendsinglevextendsingle(c1,c2,...,cl)m1,m2,...,m langbracketrightbig . Ifc′ j/negationslash=cjforsome j, thenthepreviousequationfollows. Exercises 4.3.1 Showthat(a)[J+,J2]=0,(b)[J−,J2]=0. 4.3.2 Derivetherootdiagramof SU(3) inFig.4.6bfrom thegenerators λiinEq. (4.61). Hint.Workoutfirstthe SU(2) caseinFig. 4.6afrom thePaulimatrices. 4.4 A NGULAR MOMENTUM COUPLING In many-body systems of classical mechanics, the total angular momentum is the sum L=summationtext iLiof theindividualorbitalangularmomenta.Anyisolatedparticlehas conserved angular momentum. In quantum mechanics, conserved angular momentum arises when particles move in a central potential, such as the Coulomb potential in atomic physics, a shell model potential in nuclear physics, or a confinement potential of a quark model in 4.4 Angular Momentum Coupling 267 particle physics. In the relativistic Dirac equation, orbital angular momentum is no longer conserved,but J=L+Sisconserved,thetotalangularmomentumofaparticleconsisting ofits orbitalandintrinsicangularmomentum,calledspin S=σ/2,inunitsof¯h. It is readily shown that the sum of angular momentum operators obeys the same com- mutation relations in Eq. (4.37) or (4.41) as the individual angular momentum operators, providedthosefromdifferent particlescommute. Clebsch–Gordan Coefficients: SU(2)–SO(3) Clearly,combiningtwocommutingangularmomenta Jitoform theirsum J=J1+J2,[J1i,J2i]=0, (4.91) occursofteninapplications,and Jsatisfiestheangularmomentumcommutationrelations [Jj,Jk]=[J1j+J2j,J1k+J2k]=[J1j,J1k]+[J2j,J2k]=iεjkl(J1l+J2l)=iεjklJl. For a single particle with spin 1 /2, for example, an electron or a quark, the total angular momentum is a sum of orbital angular momentum and spin. For two spinless particles their total orbital angular momentum L=L1+L2.F o rJ2andJzof Eq. (4.91) to be both diagonal,[J2,Jz]=0 has to hold. To show this we use the obvious commutationrelations [Jiz,J2 j]=0,and J2=J2 1+J22+2J1·J2=J21+J22+J1+J2−+J1−J2++2J1zJ2z (4.91′) inconjunctionwithEq. (4.71), for both Ji, toobtain bracketleftbig J2,Jzbracketrightbig =[J1−J2++J1+J2−,J1z+J2z] =[J1−,J1z]J2++J1−[J2+,J2z]+[J1+,J1z]J2−+J1+[J2−,J2z] =J1−J2+−J1−J2+−J1+J2−+J1+J2−=0. Similarly[J2,J2 i]=0 is proved. Hence the eigenvalues of J2i,J2,Jzcan be used to label thetotalangularmomentumstates |J1J2JM/angbracketright. Theproductstates |J1m1/angbracketright|J2m2/angbracketrightobviouslysatisfy theeigenvalueequations Jz|J1m1/angbracketright|J2m2/angbracketright=(J1z+J2z)|J1m1/angbracketright|J2m2/angbracketright=(m1+m2)|J1m1/angbracketright|J2m2/angbracketright =M|J1m1/angbracketright|J2m2/angbracketright, (4.92) J2i|J1m1/angbracketright|J2m2/angbracketright=Ji(Ji+1)|J1m1/angbracketright|J2m2/angbracketright, but will not have diagonal J2except for the maximally stretched states with M= ±(J1+J2)andJ=J1+J2(see Fig. 4.7a). To see this we use Eq. (4.91′) again in con- junctionwithEqs. (4.83)and(4.84)in J2|J1m1/angbracketrightJ2m2/angbracketright=braceleftbig J1(J1+1)+J2(J2+1)+2m1m2bracerightbig |J1m1/angbracketright|J2m2/angbracketright +braceleftbig J1(J1+1)−m1(m1+1)bracerightbig1/2braceleftbig J2(J2+1)−m2(m2−1)bracerightbig1/2 ×|J1m1+1/angbracketright|J2m2−1/angbracketright+braceleftbig J1(J1+1)−m1(m1−1)bracerightbig1/2 ×braceleftbig J2(J2+1)−m2(m2+1)bracerightbig1/2|J1m1−1/angbracketright|J2m2+1/angbracketright. (4.93) 268 Chapter 4 Group Theory FIGURE 4.7Couplingof twoangularmomenta: (a) parallelstretched,(b) antiparallel,(c)general case. ThelasttwotermsinEq.(4.93)vanishonlywhen m1=J1andm2=J2orm1=−J1and m2=−J2. In both cases J=J1+J2follows from the first line of Eq. (4.93). In general, therefore,wehavetoformappropriatelinearcombinationsof productstates |J1J2JM/angbracketright=summationdisplay m1,m2Cparenleftbig J1J2J|m1m2Mparenrightbig |J1m1/angbracketright|J2m2/angbracketright, (4.94) so thatJ2has eigenvalue J(J+1). The quantities C(J1J2J|m1m2M)in Eq. (4.94) are calledClebsch–Gordancoefficients .FromEq.(4.92)weseethattheyvanishunless M= m1+m2, reducing the double sum to a single sum. Applying J±to|JM/angbracketrightshows that the eigenvalues MofJzsatisfy theusualinequalities −J≤M≤J. Clearly,themaximal Jmax=J1+J2(seeFig.4.7a).InthiscaseEq.(4.93)reducestoa pureproductstate |J1J2J=J1+J2M=J1+J2/angbracketright=|J1J1/angbracketright|J2J2/angbracketright, (4.95a) sotheClebsch–Gordancoefficient C(J1J2J=J1+J2|J1J2J1+J2)=1. (4.95b) Theminimal J=J1−J2(ifJ1>J2,seeFig.4.7b)and J=J2−J1forJ2>J1followif wekeepinmindthattherearejustas manyproductstatesas |JM/angbracketrightstates;thatis, Jmaxsummationdisplay J=Jmin(2J+1)=(Jmax−Jmin+1)(Jmax+Jmin+1) =(2J1+1)(2J2+1). (4.96) This condition holds because the |J1J2JM/angbracketrightstates merely rearrange all product states into irreduciblerepresentationsoftotalangularmomentum.Itisequivalenttothe trianglerule : /Delta1(J1J2J)=1,if|J1−J2|≤J≤J1+J2; /Delta1(J1J2J)=0,else.(4.97) 4.4 Angular Momentum Coupling 269 This indicates that one complete multiplet of each Jvalue from JmintoJmaxaccounts for all the states and that all the |JM/angbracketrightstates are necessarily orthogonal. In other words, Eq. (4.94) defines a unitary transformation from the orthogonal basis set of products of single-particle states |J1m1;J2m2/angbracketright=|J1m1/angbracketright|J2m2/angbracketrightto the two-particle states |J1J2JM/angbracketright. TheClebsch–Gordancoefficientsarejusttheoverlapmatrixelements C(J1J2J|m1m2M)≡/angbracketleftJ1J2JM|J1m1;J2m2/angbracketright. (4.98) The explicit construction in what follows shows that they are all real. The states in Eq.(4.94) areorthonormalized,providedthattheconstraints summationdisplay m1,m2,m1+m2=MC(J1J2J|m1m2M)C(J 1J2J′|m1m2M′/angbracketright =/angbracketleftJ1J2JM|J1J2J′M′/angbracketright=δJJ′δMM′(4.99a) summationdisplay J,MC(J1J2J|m1m2M)C(J 1J2J|m′ 1m′2M) =/angbracketleftJ1m1|J1m′1/angbracketright/angbracketleftJ2m2|J2m′2/angbracketright=δm1m′ 1δm2m′2(4.99b) hold. Nowwearereadytoconstructmoredirectlythetotalangularmomentumstatesstarting from|Jmax=J1+J2M=J1+J2/angbracketrightin Eq. (4.95a) and using the lowering operator J−= J1−+J2−repeatedly.In thefirst stepweuseEq.(4.84) for Ji−|JiJi/angbracketright=braceleftbig Ji(Ji+1)−Ji(Ji−1)bracerightbig1/2|JiJi−1/angbracketright=(2Ji)1/2|JiJi−1/angbracketright, which we substitute into (J1−+J2−/angbracketright|J1J1)|J2J2/angbracketright. Normalizing the resulting state with M=J1+J2−1 properlyto1,weobtain |J1J2J1+J2J1+J2−1/angbracketright=braceleftbig J1/(J1+J2)bracerightbig1/2|J1J1−1/angbracketright|J2J2/angbracketright +braceleftbig J2/(J1+J2)bracerightbig1/2|J1J1/angbracketright|J2J2−1/angbracketright.(4.100) Equation(4.100)yieldstheClebsch–Gordancoefficients C(J1J2J1+J2|J1−1J2J1+J2−1)=braceleftbig J1/(J1+J2)bracerightbig1/2, C(J1J2J1+J2|J1J2−1J1+J2−1)=braceleftbig J2/(J1+J2)bracerightbig1/2.(4.101) Thenweapply J−againandnormalizethestatesobtaineduntilwereach |J1J2J1+J2M/angbracketright withM=−(J1+J2). The Clebsch–Gordan coefficients C(J1J2J1+J2|m1m2M)may thusbecalculatedstepbystep,andtheyareallreal. Thenextstepistorealizethattheonlyotherstatewith M=J1+J2−1isthetopofthe nextlowertowerof |J1+J2−1M/angbracketrightstates.Since|J1+J2−1J1+J2−1/angbracketrightisorthogonalto |J1+J2J1+J2−1/angbracketrightinEq.(4.100),itmustbetheotherlinearcombinationwitharelative minussign, |J1+J2−1J1+J2−1/angbracketright=−braceleftbig J2/(J1+J2)bracerightbig1/2|J1J1−1/angbracketright|J2J2/angbracketright +braceleftbig J1/(J1+J2)bracerightbig1/2|J1J1/angbracketright|J2J2−1/angbracketright,(4.102) uptoanoverallsign. 270 Chapter 4 Group Theory HencewehavedeterminedtheClebsch–Gordancoefficients(for J2≥J1) C(J1J2J1+J2−1|J1−1J2J1+J2−1)=−braceleftbig J2/(J1+J2)bracerightbig1/2, C(J1J2J1+J2−1|J1J2−1J1+J2−1)=braceleftbig J1/(J1+J2)bracerightbig1/2.(4.103) Againwecontinueusing J−untilwereach M=−(J1+J2−1),andwekeepnormalizing theresultingstates |J1+J2−1M/angbracketrightof theJ=J1+J2−1t o w e r . In order to get to the top of the next tower, |J1+J2−2M/angbracketrightwithM=J1+J2−2, we remember that we have already constructed two states with that M.B o t h|J1+J2J1+ J2−2/angbracketrightand|J1+J2−1J1+J2−2/angbracketrightareknownlinearcombinationsofthethreeproduct states|J1J1/angbracketright|J2J2−2/angbracketright,|J1J1−1/angbracketright×|J2J2−1/angbracketright, and|J1J1−2/angbracketright|J2J2/angbracketright. The third linear combination is easy to find from orthogonality to these two states, up to an overall phase, which is chosen by the Condon–Shortley phase conventions11so that the coefficient C(J1J2J1+J2−2|J1J2−2J1+J2−2)ofthelastproductstateispositivefor |J1J2J1+ J2−2J1+J2−2/angbracketright.Itisstraightforward,thoughabittedious,todeterminetherestofthe Clebsch–Gordancoefficients. Numerous recursion relations can be derived from matrix elements of various angular momentumoperators,for whichwerefer totheliterature.12 ThesymmetrypropertiesofClebsch–Gordancoefficientsarebestdisplayedinthemore symmetricWigner’s3 j-symbols,whicharetabulated:12 parenleftBigg J1J2J3 m1m2m3parenrightBigg =(−1)J1−J2−m3 (2J3+1)1/2C(J1J2J3|m1m2,−m3), (4.104a) obeyingthesymmetryrelations parenleftBigg J1J2J3 m1m2m3parenrightBigg =(−1)J1+J2+J3parenleftBigg JkJlJn mkmlmnparenrightBigg (4.104b) for(k,l,n)an odd permutation of (1,2,3). One of the most important places where Clebsch–Gordan coefficients occur is in matrix elements of tensor operators, which are governed by the Wigner–Eckart theorem discussed in the next section, on spherical ten- sors. Another is coupling of operators or state vectors to total angular momentum, such as spin-orbit coupling. Recoupling of operators and states in matrix elements leads to 6 j- and 9j-symbols.12Clebsch–Gordan coefficients can and have been calculated for other Liegroups,suchas SU(3). 11E.U.Condon and G.H.Shortley, Theory of AtomicSpectra . Cambridge, UK:Cambridge University Press (1935). 12There is a rich literature on this subject, e.g., A. R. Edmonds, Angular Momentum in Quantum Mechanics . Princeton, NJ: PrincetonUniversityPress(1957);M.E.Rose, ElementaryTheoryofAngularMomentum .NewYork:Wiley(1957);A.de-Shalit andI.Talmi, NuclearShellModel .NewYork:AcademicPress(1963);Dover(2005).Clebsch–Gordancoefficientsaretabulated in M. Rotenberg, R. Bivins, N. Metropolis, and J. K. Wooten, Jr., The 3j- and 6j-Symbols . Cambridge, MA: Massachusetts Institute of Technology Press (1959). 4.4 Angular Momentum Coupling 271 Spherical Tensors In Chapter 2 the properties of Cartesian tensors are defined using the group of nonsin- gular general linear transformations, which contains the three-dimensional rotations as a subgroup. A tensor of a given rank that is irreducible with respect to the full group may well become reducible for the rotation group SO(3). To explain this point, consider the second-ranktensorwithcomponents Tjk=xjykforj,k=1,2,3.Itcontainsthesymmet- rictensor Sjk=(xjyk+xkyj)/2andtheantisymmetrictensor Ajk=(xjyk−xkyj)/2,so Tjk=Sjk+Ajk. This reduces TjkinSO(3). However, under rotations the scalar product x·yis invariant and is therefore irreducible in SO(3). Thus, Sjkcan be reduced by sub- traction of the multiple of x·ythat makes it traceless. This leads to the SO(3)-irreducible tensor S′ jk=1 2(xjyk+xkyj)−1 3x·yδjk. Tensors of higher rank may be treated similarly. When we form tensors from products of the components of the coordinate vector rthen, in polar coordinates that are tailored to SO(3)symmetry,weendupwiththesphericalharmonicsofChapter12. The form of the ladder operators for SO(3) in Section 4.3 leads us to introduce the spherical components (note the different normalization and signs, though, prescribed by theYlm) ofavector A: A+1=−1√ 2(Ax+iAy), A−1=1√ 2(Ax−iAy), A 0=Az.(4.105) Thenwehavefor thecoordinatevector rinpolarcoordinates, r+1=−1√ 2rsinθeiϕ=rradicalBig 4π 3Y11,r−1=1√ 2rsinθe−iϕ=rradicalBig 4π 3Y1,−1, r0=rradicalBig 4π 3Y10,(4.106) whereYlm(θ,ϕ)are the spherical harmonics of Chapter 12. Again, the spherical jmcom- ponentsoftensors Tjmof higherrank jmaybeintroducedsimilarly. An irreducible spherical tensor operator Tjmof rankjhas 2j+1 components, just as for spherical harmonics, and mruns from−jto+j. Under a rotation R(α), whereα standsfortheEulerangles,the Ylmtransform as Ylm(ˆr′)=summationdisplay m′Ylm′(ˆr)Dl m′m(R), (4.107a) whereˆr′=(θ′,ϕ′)areobtainedfrom ˆr=(θ,ϕ)bytherotation Randaretheanglesofthe samepointintherotatedframe,and DJ m′m(α,β,γ)=/angbracketleftJm|exp(iαJz)exp(iβJy)exp(iγJz)|Jm′/angbracketright aretherotationmatrices.So,for theoperator Tjm, wedefine RTjmR−1=summationdisplay m′Tjm′Dj m′m(α). (4.107b) 272 Chapter 4 Group Theory Foraninfinitesimalrotation(seeEq.(4.20)inSection4.2ongenerators)theleftsideof Eq. (4.107b) simplifies to a commutator and the right side to the matrix elements of J,t h e infinitesimalgeneratorof therotation R: [Jn,Tjm]=summationdisplay m′Tjm′/angbracketleftjm′|Jn|jm/angbracketright. (4.108) IfwesubstituteEqs.(4.83)and(4.84)forthematrixelementsof Jmweobtainthealterna- tivetransformationlawsofatensor operator, [J0,Tjm]=mTjm,[J±,Tjm]=Tjm±1braceleftbig (j−m)(j±m+1)bracerightbig1/2.(4.109) We can use the Clebsch–Gordan coefficients of the previous subsection to couple two tensors of given rank to another rank. An example is the cross or vector product of two vectorsaandbfromChapter1.Letuswritebothvectorsinsphericalcomponents, amand bm. Thenweverifythatthetensor Cmofrank1 definedas Cm≡summationdisplay m1m2C(111|m1m2m)am1bm2=i√ 2(a×b)m. (4.110) SinceCmisasphericaltensorofrank1thatislinearinthecomponentsof aandb,itmust beproportionaltothecrossproduct, Cm=N(a×b)m.Theconstant Ncanbedetermined fromaspecialcase, a=ˆx,b=ˆy,essentiallywriting ˆx׈y=ˆzinsphericalcomponentsas follows.Using (ˆz)0=1;(ˆx)1=−1/√ 2,(ˆx)−1=1/√ 2; (ˆy)1=−i/√ 2,(ˆy)−1=−i/√ 2, Eq.(4.110) for m=0 becomes C(111|1,−1,0)bracketleftbig (ˆx)1(ˆy)−1−(ˆx)−1(ˆy)1bracketrightbig =Nparenleftbig (ˆz)0parenrightbig =N =1√ 2bracketleftbigg −1√ 2parenleftbigg −i√ 2parenrightbigg −1√ 2parenleftbigg −i√ 2parenrightbiggbracketrightbigg =i√ 2, w h e r ew eh a v eu s e d C(111|101)=1√ 2from Eq. (4.103) for J1=1=J2, which implies C(111|1,−1,0)=1√ 2usingEqs. (4.104a,b): parenleftBigg 11 1 10−1parenrightBigg =−1√ 3C(111|101)=−1 6=−parenleftBigg 111 1−10parenrightBigg =−1√ 3C(111|1,−1,0). A bit simpler is the usual scalar product of two vectors in Chapter 1, in which aandb arecoupledtozeroangularmomentum: a·b≡−(ab)0√ 3≡−√ 3summationdisplay mC(110|m,−m,0)amb−m. (4.111) Again, the rank zero of our tensor product implies a·b=n(ab)0. The constant ncan be determined from a special case, essentially writing ˆz2=1 in spherical components: ˆz2=1=nC(110|000)=−n√ 3. 4.4 Angular Momentum Coupling 273 Anotheroften-usedapplicationoftensorsisthe recoupling thatinvolves 6j-symbols for three operators and 9 jfor four operators.12An example is the following scalar product, forwhichit canbeshown12that σ1·rσ2·r=1 3r2σ1·σ2+(σ1σ2)2·(rr)2, (4.112) but which can also be rearranged by elementary means. Here the tensor operators are de- finedas (σ1σ2)2m=summationdisplay m1m2C(112|m1m2m)σ1m1σ2m2, (4.113) (rr)2m=summationdisplay mC(112|m1m2m)rm1rm2=radicalbigg 8π 15r2Y2m(ˆr), (4.114) andthescalarproductof tensorsofrank2 as (σ1σ2)2·(rr)2=summationdisplay m(−1)m(σ1σ2)2m(rr)2,−m=√ 5parenleftbig (σ1σ2)2(rr)2parenrightbig 0.(4.115) One of the most important applications of spherical tensor operators is the Wigner– Eckarttheorem .Itsaysthatamatrixelementofasphericaltensoroperator Tkmofrankk betweenstatesofangularmomentum jandj′factorizesintoaClebsch–Gordancoefficient and a so-called reduced matrix element , denoted by double bars, that no longer has any dependenceontheprojectionquantumnumbers m,m′,n: /angbracketleftj′m′|Tkn|jm/angbracketright=C(kjj′|nmm′)(−1)k−j+j′/angbracketleftj′/bardblTk/bardblj/angbracketright/radicalbig (2j′+1). (4.116) In other words, such a matrix element factors into a dynamic part, the reduced matrix element, and a geometric part, the Clebsch–Gordan coefficient that contains the rotational properties (expressed by the projection quantum numbers) from the SO(3) invariance. To seethiswecouple Tknwiththeinitialstatetototalangularmomentum j′: |j′m′/angbracketright0≡summationdisplay nmC(kjj′|nmm′)Tkn|jm/angbracketright. (4.117) Under rotations the state |j′m′/angbracketright0transforms just like |j′m′/angbracketright. Thus, the overlap matrix ele- ment/angbracketleftj′m′|j′m′/angbracketright0isarotationalscalarthathasno m′dependence,sowecanaverageover theprojections, /angbracketleftJM|j′m′/angbracketright0=δJj′δMm′ 2j′+1summationdisplay µ/angbracketleftj′µ|j′µ/angbracketright0. (4.118) Next we substitute our definition, Eq. (4.117), into Eq. (4.118) and invert the relation Eq.(4.117) usingorthogonality,Eq. (4.99b),tofindthat /angbracketleftJM|Tkn|jm/angbracketright=summationdisplay j′m′C(kjj′|nmm′)δJj′δMm′ 2J+1summationdisplay µ/angbracketleftJµ|Jµ/angbracketright0,(4.119) whichprovestheWigner–Eckarttheorem,Eq.(4.116).13 13Theextra factor (−1)k−j+j′/radicalbig (2j′+1)in Eq.(4.116) is just aconvention that varies in the literature. 274 Chapter 4 Group Theory As an application, we can write the Pauli matrix elements in terms of Clebsch–Gordan coefficients.WeapplytheWigner–Eckarttheoremto angbracketleftbig1 2γvextendsinglevextendsingleσαvextendsinglevextendsingle1 2βangbracketrightbig =(σα)γβ=−1√ 2Cparenleftbig 11 21 2vextendsinglevextendsingleαβγparenrightbigangbracketleftbig1 2vextenddoublevextenddoubleσvextenddoublevextenddouble1 2angbracketrightbig . (4.120) Since/angbracketleft1 21 2|σ0|1 21 2/angbracketright=1 withσ0=σ3andC(11 21 2|01 21 2)=−1/√ 3,wefind angbracketleftbig1 2vextenddoublevextenddoubleσvextenddoublevextenddouble1 2angbracketrightbig =√ 6, (4.121) which,substitutedintoEq.(4.120), yields (σα)γβ=−√ 3Cparenleftbig 11 21 2vextendsinglevextendsingleαβγparenrightbig . (4.122) Notethatthe α=±1,0 denotethesphericalcomponentsofthePaulimatrices. Young Tableaux for SU(n) Young tableaux (YT) provide a powerful and elegant method for decomposing products ofSU(n)group representations into sums of irreducible representations. The YT provide the dimensions and symmetry types of the irreducible representations in this so-called Clebsch–Gordan series , though not the Clebsch–Gordan coefficients by which the prod- uct states are coupled to the quantum numbers of each irreducible representation of the series(see Eq.(4.94)). Products of representations correspond to multiparticle states. In this context, permuta- tionsofparticlesareimportantwhenwedealwithseveralidenticalparticles.Permutations ofnidentical objects form the symmetric group Sn. A close connection between irre- ducible representations of Sn, which are the YT, and those of SU(n)is provided by this theorem:E v e r yN-particle state of Snthat is made up of single-particle states of the fun- damental n-dimensional SU(n)multiplet belongs to an irreducible SU(n)representation. AproofisinChapter22ofWybourne.14 ForSU(2) the fundamental representation is a box that stands for the spin +1 2(up) and −1 2(down)statesandhasdimension2.For SU(3)theboxcomprisesthethreequarkstates inthetriangleofFig.4.5a;ithasdimension3. An array of boxes shown in Fig. 4.8 with λ1boxes in the first row, λ2boxes in the second row, ...,andλn−1boxes in the last row is called a Young tableau (YT), denoted by[λ1,...,λn−1],andrepresentsanirreduciblerepresentationof SU(n)ifandonlyif λ1≥λ2≥···≥λn−1. (4.123) Boxes in the same row are symmetric representations; those in the same column are anti- symmetric. A YT consisting of one row is totally symmetric. A YT consisting of a single columnistotallyantisymmetric. There are at most n−1r o w sf o r SU(n)YT because a column of nboxes is the totally antisymmetric ( Slater determinant of single-particle states) singlet representation that maybestruckfromtheYT. An array of Nboxes is an N-particle state whose boxes may be labeled by positive integerssothatthe(particlelabelsor)numbersinonerowoftheYTdonotdecreasefrom 14B. G.Wybourne, Classical Groups for Physicists .NewYork: Wiley (1974). 4.4 Angular Momentum Coupling 275 FIGURE 4.8Youngtableau(YT) for SU(n). left to right and those in any one column increase from top to bottom. In contrast to the possiblerepetitionsofrownumbers,thenumbersinanycolumnmustbedifferentbecause oftheantisymmetryof thesestates. The product of a YT with a single box, [1], is the sum of YT formed when the box is putattheendofeachrowoftheYT,providedtheresultingYTislegitimate,thatis,obeys Eq.(4.123).For SU(2)theproductoftwoboxes,spin 1 /2 representationsofdimension2, generates [1]⊗[1]=[2]⊕[1,1], (4.124) the symmetric spin 1 representation of dimension 3 and the antisymmetric singlet of di- mension1mentionedearlier. Thecolumnof n−1boxesistheconjugaterepresentationofthefundamentalrepresen- tation; its product with a single box contains the column of nboxes, which is the singlet. ForSU(3) the conjugate representation of the single box, [1]or fundamental quark repre- sentation, is the inverted triangle in Fig. 4.5b, [1,1], which represents the three antiquarks ¯u,¯d,¯s, obviouslyof dimension3aswell. ThedimensionofaYT isgivenbytheratio dimYT=N D. (4.125) The numerator Nis obtained by writing an nin all boxes of the YT along the diagonal, (n+1)inallboxesimmediatelyabovethediagonal, (n−1)immediatelybelowthediago- nal,etc.NistheproductofallthenumbersintheYT.AnexampleisshowninFig.4.9afor the octet representation of SU(3), where N=2·3·4=24. There is a closed formula that isequivalenttoEq.(4.125).15Thedenominator Distheproductofall hooks.16Ahookis drawnthrougheachboxoftheYTbystartingahorizontallinefromtherighttotheboxin questionandthencontinuingitverticallyoutoftheYT.Thenumberofboxesencountered by the hook-line is the hook-number of the box. Dis the product of all hook-numbers of 15See, for example, M. Hamermesh, Group Theory and Its Application to Physical Problems . Reading, MA: Addison-Wesley (1962). 16F. Close, Introduction to Quarks and Partons . NewYork: AcademicPress (1979). 276 Chapter 4 Group Theory (a) (b) FIGURE 4.9Illustration of(a)Nand(b)Din Eq. (4.125)for theoctet Youngtableauof SU(3). the YT. An example is shown in Fig. 4.9b for the octet of SU(3), whose hook-number is D=1·3·1=3.Hencethedimensionof the SU(3) octetis 24 /3=8,whenceits name. Now we can calculate the dimensions of the YT in Eq. (4.124). For SU(2) they are 2×2=3+1=4. ForSU(3) they are 3·3=3·4/(1·2)+3·2/(2·1)=6+3=9. For theproductofthequarktimesantiquarkYTof SU(3)weget [1,1]⊗[1]=[2,1]⊕[1,1,1], (4.126) that is, octet and singlet, which are precisely the meson multiplets considered in the sub- sectionontheeightfoldway,the SU(3)flavorsymmetry,whichsuggestmesonsarebound states of a quark and an antiquark, q¯qconfigurations. For the product of three quarks we get parenleftbig [1]⊗[1]parenrightbig ⊗[1]=parenleftbig [2]⊕[1,1]parenrightbig ⊗[1]=[3]⊕2[2,1]⊕[1,1,1],(4.127) that is, decuplet, octet, and singlet, which are the observed multiplets for the baryons, whichsuggeststheyareboundstatesof threequarks, q3configurations. Aswehaveseen,YTdescribethedecompositionofaproductof SU(n)irreduciblerepre- sentations into irreducible representations of SU(n), which is called the Clebsch–Gordan series,whiletheClebsch–Gordancoefficientsconsideredearlierallowconstructionofthe individualstatesinthisseries. 4.4 Angular Momentum Coupling 277 Exercises 4.4.1 Derive recursion relations for Clebsch–Gordan coefficients. Use them to calculate C(11J|m1m2M)forJ=0,1,2. Hint.Usetheknownmatrixelementsof J+=J1++J2+,Ji+,andJ2=(J1+J2)2,etc. 4.4.2 Show that (Ylχ)J M=summationtextC(l1 2J|mlmsM)Ylmlχms, whereχ±1/2are the spin up and downeigenfunctionsof σ3=σz, transforms likeasphericaltensorof rank J. 4.4.3 Whenthespinofquarksistakenintoaccount,the SU(3)flavorsymmetryisreplacedby theSU(6) symmetry. Why? Obtain the Young tableau for the antiquark configuration ¯q. Then decompose the product q¯q. WhichSU(3) representations are contained in the nontrivialSU(6)representationfor mesons? Hint.DeterminethedimensionsofallYT. 4.4.4 Forl=1,Eq. (4.107a)becomes Ym 1(θ′,ϕ′)=1summationdisplay m′=−1D1 m′m(α,β,γ)Ym′ 1(θ,ϕ). Rewrite these spherical harmonics in Cartesian form. Show that the resulting Cartesian coordinate equations are equivalent to the Euler rotation matrix A(α,β,γ), Eq. (3.94), rotatingthecoordinates. 4.4.5 Assumingthat Dj(α,β,γ) is unitary,showthat lsummationdisplay m=−lYm∗ l(θ1,ϕ1)Ym l(θ2,ϕ2) is a scalar quantity (invariant under rotations). This is a spherical tensor analog of a scalarproductof vectors. 4.4.6 (a) Showthatthe αandγdependenceof Dj(α,β,γ) maybefactoredoutsuchthat Dj(α,β,γ)=Aj(α)dj(β)Cj(γ). (b) Showthat Aj(α)andCj(γ)arediagonal.Findtheexplicitforms. (c) Showthat dj(β)=Dj(0,β,0). 4.4.7 Theangularmomentum–exponentialformof theEuleranglerotationoperatorsis R=Rz′′(γ)Ry′(β)Rz(α) =exp(−iγJz′′)exp(−iβJy′)exp(−iαJz). Showthatintermsoftheoriginalaxes R=exp(iαJz)exp(−iβJy)exp(−iγJz). Hint.T h eRoperators transform as matrices. The rotation about the y′-axis (second Eulerrotation)maybereferredtotheoriginal y-axisby exp(−iβJy′)=exp(−iαJz)exp(−iβJy)exp(iαJz). 278 Chapter 4 Group Theory 4.4.8 Using the Wigner–Eckart theorem, prove the decomposition theorem for a spherical vectoroperator /angbracketleftj′m′|T1m|jm/angbracketright=/angbracketleftjm′|J·T1|jm/angbracketright j(j+1)δjj′. 4.4.9 UsingtheWigner–Eckarttheorem,provethefactorization /angbracketleftj′m′|JMJ·T1|jm/angbracketright=/angbracketleftjm′|JM|jm/angbracketrightδj′j/angbracketleftjm|J·T1|jm/angbracketright. 4.5 H OMOGENEOUS LORENTZ GROUP Generalizing the approach to vectors of Section 1.2, in special relativity we demand that ourphysicallawsbecovariant17under a. spaceandtimetranslations, b. rotationsinreal,three-dimensionalspace,and c. Lorentztransformations. The demand for covariance under translations is based on the homogeneity of space and time. Covariance under rotations is an assertion of the isotropy of space. The requirement of Lorentz covariance follows from special relativity. All three of these transformations togetherformtheinhomogeneousLorentzgroupor thePoincarégroup.Whenweexclude translations, the space rotations and the Lorentz transformations together form a group — thehomogeneousLorentzgroup. We first generate a subgroup, the Lorentz transformations in which the relative velocity vis along the x=x1-axis. The generator may be determined by considering space–time reference frames moving with a relative velocity δv, an infinitesimal.18The relations are similar to those for rotations in real space, Sections 1.2, 2.6, and 3.3, except that here the angleofrotationispureimaginary(compareSection4.6). Lorentz transformations are linear not only in the space coordinates xibut in the time t as well. They originate from Maxwell’s equations of electrodynamics, which are invariant under Lorentz transformations, as we shall see later. Lorentz transformations leave the quadratic form c2t2−x2 1−x2 2−x2 3=x2 0−x2 1−x2 2−x2 3invariant, where x0=ct.W e see this if we switch on a light source at the origin of the coordinate system. At time t lighthastraveledthedistance ct=radicalBigsummationtextx2 i,s oc2t2−x2 1−x2 2−x2 3=0.Specialrelativity requires that in all (inertial) frames that move with velocity v≤cin any direction relative tothexi-systemandhavethesameoriginattime t=0,c2t′2−x′2 1−x′2 2−x′2 3=0 holds also.Four-dimensionalspace–timewiththemetric x·x=x2=x2 0−x2 1−x2 2−x2 3iscalled Minkowskispace,withthescalarproductoftwofour-vectorsdefinedas a·b=a0b0−a·b. Usingthemetrictensor (gµν)=parenleftbig gµνparenrightbig = 1 000 0−10 0 00−10 00 0 −1 (4.128) 17To be covariant means to have the same form in different coordinate systems so that there is no preferred reference system (compare Sections 1.2 and 2.6). 18This derivation, withaslightly different metric,appears inanarticleby J.L.Strecker, Am.J.Phys. 35: 12 (1967). 4.5 Homogeneous Lorentz Group 279 we can raise and lower the indices of a four-vector, such as the coordinates xµ=(x0,x), so thatxµ=gµνxν=(x0,−x)andxµgµνxν=x2 0−x2, Einstein’s summationconvention being understood. For the gradient, ∂µ=(∂/∂x0,−∇)=∂/∂xµand∂µ=(∂/∂x0,∇),s o ∂2=∂µ∂µ=(∂/∂x0)2−∇2is aLorentzscalar, justlikethemetric x2=x2 0−x2. Forv≪c,inthenonrelativisticlimit,aLorentztransformationmustbeGalilean.Hence, to derive the form of a Lorentz transformation along the x1-axis, we start with a Galilean transformationfor infinitesimalrelativevelocity δv: x′1=x1−δvt=x1−x0δβ. (4.129) Here,β=v/c.Bysymmetrywealsowrite x′0=x0+aδβx1, (4.129′) withtheparameter achosenso that x2 0−x2 1isinvariant, x′2 0−x′2 1=x2 0−x2 1. (4.130) Remember, xµ=(x0,x)is the prototype four-dimensional vector in Minkowski space. Thus Eq. (4.130) is simply a statement of the invariance of the square of the magnitude of the“distance”vectorunderLorentztransformationinMinkowskispace.Hereiswherethe specialrelativityisbroughtintoourtransformation.SquaringandsubtractingEqs.(4.129) and (4.129′) and discarding terms of order (δβ)2, we find a=−1. Equations (4.129) and (4.129′)maybecombinedasamatrixequation, parenleftBigg x′0 x′1parenrightBigg =(12−δβσ1)parenleftBigg x0 x1parenrightBigg ; (4.131) σ1happens to be the Pauli matrix, σ1, and the parameter δβrepresents an infinitesimal change.UsingthesametechniquesasinSection4.2,werepeatthetransformation Ntimes todevelopafinitetransformationwiththevelocityparameter ρ=Nδβ. Then parenleftBigg x′0 x′1parenrightBigg =parenleftbigg 12−ρσ1 NparenrightbiggNparenleftBigg x0 x1parenrightBigg . (4.132) Inthelimitas N→∞, lim N→∞parenleftbigg 12−ρσ1 NparenrightbiggN =exp(−ρσ1). (4.133) AsinSection4.2,theexponentialisinterpretedbyaMaclaurinexpansion, exp(−ρσ1)=12−ρσ1+1 2!(ρσ1)2−1 3!(ρσ1)3+···. (4.134) Notingthat (σ1)2=12, exp(−ρσ1)=12coshρ−σ1sinhρ. (4.135) HenceourfiniteLorentztransformationisparenleftBigg x′0 x′1parenrightBigg =parenleftBigg coshρ−sinhρ −sinhρcoshρparenrightBiggparenleftBigg x0 x1parenrightBigg . (4.136) 280 Chapter 4 Group Theory σ1has generated the representations of this pure Lorentz transformation. The quantities coshρand sinhρmay be identified by considering the origin of the primed coordinate system,x′1=0,orx1=vt. SubstitutingintoEq.(4.136), wehave 0=x1coshρ−x0sinhρ. (4.137) Withx1=vtandx0=ct, tanhρ=β=v c. Note that the rapidity ρ/negationslash=v/c, except in the limit as v→0. The rapidity is the additive parameterforpureLorentztransformations(“boosts”)alongthesameaxisthatcorresponds toanglesfor rotationsaboutthesameaxis.Using 1 −tanh2ρ=(cosh2ρ)−1, coshρ=parenleftbig 1−β2parenrightbig−1/2≡γ,sinhρ=βγ. (4.138) The group of Lorentz transformations is not compact, because the limit of a sequence of rapiditiesgoingtoinfinityis nolongeranelementofthegroup. The preceding special case of the velocity parallel to one space axis is easy, but it illus- trates the infinitesimal velocity-exponentiation-generator technique. Now, this exact tech- nique may be applied to derive the Lorentz transformation for the relative velocity v not parallel to any space axis. The matrices given by Eq. (4.136) for the case of v=ˆxvxform a subgroup. The matrices in the general case do not. The product of two Lorentz transfor- mation matrices L(v1)andL(v2)yields a third Lorentz matrix, L(v3), if the two velocities v1andv2are parallel. The resultant velocity, v3, is related to v1andv2by the Einstein velocity addition law, Exercise 4.5.3. If v1andv2are not parallel, no such simple relation exists.Specifically,considerthreereferenceframes S,S′,andS′′,withSandS′relatedby L(v1)andS′andS′′relatedby L(v2).Ifthevelocityof S′′relativetotheoriginalsystem S isv3,S′′isnotobtainedfrom SbyL(v3)=L(v2)L(v1).Rather, wefindthat L(v3)=RL(v2)L(v1), (4.139) whereRis a 3×3 space rotation matrix embedded in our four-dimensional space–time. Withv1andv2not parallel, the final system, S′′,i srotatedrelative to S. This rotation is the origin of the Thomas precession involved in spin-orbit coupling terms in atomic and nuclear physics. Because of its presence, the pure Lorentz transformations L(v)by themselvesdonotformagroup. Kinematics and Dynamics in Minkowski Space–Time Wehaveseenthatthepropagationoflightdeterminesthemetric r2−c2t2=0=r′2−c2t′2, wherexµ=(ct,r)isthecoordinatefour-vector.Foraparticlemovingwithvelocity v,the Lorentzinvariantinfinitesimalversion cdτ≡radicalbig dxµdxµ=radicalbig c2dt2−dr2=dtradicalbig c2−v2 definestheinvariantpropertime τonitstrack.Becauseoftimedilationinmovingframes, aproper-timeclockrideswiththeparticle(initsrestframe)andrunsattheslowestpossible 4.5 Homogeneous Lorentz Group 281 rate compared to any other inertial frame (of an observer, for example). The four-velocity oftheparticlecannowbedefinedproperlyas dxµ dτ=uµ=parenleftbiggc√ c2−v2,v√ c2−v2parenrightbigg , sou2=1, and the four-momentum pµ=cmuµ=(E c,p)yields Einstein’s famous energy relation E=mc2 radicalbig 1−v2/c2=mc2+m 2v2±···. A consequence of u2=1 and its physical significance is that the particle is on its mass shellp2=m2c2. NowweformulateNewton’sequationfora singleparticle ofmassminspecialrelativity asdpµ dτ=Kµ, withKµdenoting the force four-vector, so its vector part of the equation coincideswiththeusualform. For µ=1,2,3w eu s e dτ=dtradicalbig 1−v2/c2andfind 1radicalbig 1−v2/c2dp dt=Fradicalbig 1−v2/c2=K, determining Kin terms of the usual force F. We need to find K0. We proceed by analogy with the derivation of energy conservation, multiplying the force equation into the four- velocity muνduν dτ=m 2du2 dτ=0, becauseu2=1=const.Theothersideof Newton’sequationyields 0=1 cu·K=K0 radicalbig 1−v2/c2−F·v/c radicalbig 1−v2/c22, soK0=F·v/c√ 1−v2/c2isrelatedtotherateof workdonebytheforceontheparticle. Now we turn to two-body collisions, in which energy–momentum conservation takes the form p1+p2=p3+p4, wherepµ iare the particle four-momenta. Because the scalar product of any four-vector with itself is an invariant under Lorentz transformations, it is convenient to define the Lorentz invariant energy squared s=(p1+p2)2=P2, where Pµis the total four-momentum, and to use units where the velocity of light c=1. The laboratory system (lab) is defined as the rest frame of the particle with four-momentum pµ 2=(m2,0)andthecenterofmomentumframe(cms)bythetotalfour-momentum Pµ= (E1+E2,0). Whentheincidentlabenergy EL 1isgiven,then s=p2 1+p2 2+2p1·p2=m2 1+m22+2m2EL 1 isdetermined.Now,thecmsenergiesofthefourparticlesareobtainedfromscalarproducts p1·P=E1(E1+E2)=E1√ s, 282 Chapter 4 Group Theory so E1=p1·(p1+p2)√s=m2 1+p1·p2√s=m2 1−m22+s 2√s, E2=p2·(p1+p2)√s=m2 2+p1·p2√s=m2 2−m21+s 2√s, E3=p3·(p3+p4)√s=m2 3+p3·p4√s=m2 3−m24+s 2√s, E4=p4·(p3+p4)√s=m2 4+p3·p4√s=m2 4−m23+s 2√s, bysubstituting 2p1·p2=s−m2 1−m22,2p3·p4=s−m23−m24. Thus, all cms energies Eidepend only on the incident energy but not on the scattering angle. For elastic scattering, m3=m1,m4=m2,s oE3=E1,E4=E2. The Lorentz invariantmomentumtransfersquared t=(p1−p3)2=m21+m23−2p1·p3 dependslinearlyonthecosineof thescatteringangle. Example 4.5.1 KAON DECAY AND PIONPHOTOPRODUCTION THRESHOLD Find the kinetic energies of the muon of mass 106 MeV and massless neutrino into which aKmesonof mass494MeVdecaysinitsrest frame. Conservationofenergyandmomentumgives mK=Eµ+Eν=√ s.Applyingtherela- tivistickinematicsdescribedpreviouslyyields Eµ=pµ·(pµ+pν) mK=m2 µ+pµ·pν mK, Eν=pν·(pµ+pν) mK=pµ·pν mK. Combiningbothresultsweobtain m2 K=m2 µ+2pµ·pν,s o Eµ=Tµ+mµ=m2 K+m2 µ 2mK=258.4MeV, Eν=Tν=m2 K−m2 µ 2mK=235.6MeV. Asanotherexample,intheproductionofaneutralpionbyanincidentphotonaccordingto γ+p→π0+p′at threshold, the neutral pion and proton are created at rest in the cms. Therefore, s=(pγ+p)2=m2 p+2mpEL γ=(pπ+p′)2=(mπ+mp)2, 4.6 Lorentz Covariance of Maxwell’s Equations 283 soEL γ=mπ+m2 π 2mp=144.7MeV. /squaresolid Exercises 4.5.1 Two Lorentz transformations are carried out in succession: v1along the x-axis, then v2along the y-axis. Show that the resultant transformation (given by the product of these two successive transformations) cannotbe put in the form of a single Lorentz transformation. Note.Thediscrepancycorrespondstoarotation. 4.5.2 Rederive the Lorentz transformation, working entirely in the real space (x0,x1,x2,x3) withx0=x0=ct. Show that the Lorentz transformation may be written L(v)= exp(ρσ),with σ= 0−λ−µ−ν −λ000 −µ000 −ν000  andλ,µ,νthedirectioncosinesofthevelocity v. 4.5.3 Using the matrix relation, Eq. (4.136), let the rapidity ρ1relate the Lorentz reference frames(x′0,x′1)and(x0,x1).L e tρ2relate(x′′0,x′′1)and(x′0,x′1). Finally, let ρ relate(x′′0,x′′1)and(x0,x1).F r o mρ=ρ1+ρ2derive the Einstein velocity addition law v=v1+v2 1+v1v2/c2. 4.6 L ORENTZ COVARIANCE OF MAXWELL ’SEQUATIONS If a physical law is to hold for all orientations of our (real) coordinates (that is, to be in- variant under rotations), the terms of the equation must be covariant under rotations (Sec- tions 1.2 and 2.6). This means that we write the physical laws in the mathematical form scalar=scalar,vector=vector,second-ranktensor =second-ranktensor,andsoon.Sim- ilarly,ifaphysicallawistoholdforallinertialsystems,thetermsoftheequationmustbe covariantunderLorentztransformations. Using Minkowski space ( ct=x0;x=x1,y=x2,z=x3), we have a four-dimensional spacewiththemetric gµν(Eq.(4.128),Section4.5).TheLorentztransformationsarelinear inspaceandtimeinthisfour-dimensionalrealspace.19 19Agroup theoreticderivationoftheLorentztransformation inMinkowskispaceappearsinSection4.5.SeealsoH.Goldstein, Classical Mechanics . Cambridge, MA: Addison-Wesley (1951), Chapter 6. The metric equation x2 0−x2=0, independent of referenceframe, leads to the Lorentz transformations. 284 Chapter 4 Group Theory HereweconsiderMaxwell’sequations, ∇×E=−∂B ∂t, (4.140a) ∇×H=∂D ∂t+ρv, (4.140b) ∇·D=ρ, (4.140c) ∇·B=0, (4.140d) andtherelations D=ε0E,B=µ0H. (4.141) The symbols have their usual meanings as given in Section 1.9. For simplicity we assume vacuum( ε=ε0,µ=µ0). We assume that Maxwell’s equations hold in all inertial systems; that is, Maxwell’s equations are consistent with special relativity. (The covariance of Maxwell’s equations under Lorentz transformations was actually shown by Lorentz and Poincaré before Ein- steinproposedhistheoryofspecialrelativity.)OurimmediategoalistorewriteMaxwell’s equations as tensor equations in Minkowski space. This will make the Lorentz covariance explicit,ormanifest. In terms of scalar, ϕ, and magnetic vector potentials, A,w em a ys o l v e20Eq. (4.140d) andthen(4.140a)by B=∇×A E=−∂A ∂t−∇ϕ. (4.142) Equation (4.142) specifies the curl of A; the divergence of Ais still undefined (compare Section 1.16). We may, and for future convenience we do, impose a further gauge restric- tiononthevectorpotential A: ∇·A+ε0µ0∂ϕ ∂t=0. (4.143) This is the Lorentz gauge relation. It will serve the purpose of uncoupling the differential equations for Aandϕthat follow. The potentials Aandϕare not yet completely fixed. Thefreedomremainingis thetopicofExercise4.6.4. Now we rewrite the Maxwell equations in terms of the potentials Aandϕ.F r o m Eqs. (4.140c)for ∇·D,(4.141)and(4.142), ∇2ϕ+∇·∂A ∂t=−ρ ε0, (4.144) whereasEqs. (4.140b)for ∇×Hand(4.142)andEq.(1.86c) ofChapter1yield ∂2A ∂t2+∇∂ϕ ∂t+1 ε0µ0braceleftbig ∇∇·A−∇2Abracerightbig =ρv ε0. (4.145) 20Compare Section1.13, especiallyExercise1.13.10. 4.6 Lorentz Covariance of Maxwell’s Equations 285 UsingtheLorentzrelation,Eq. (4.143), andtherelation ε0µ0=1/c2, weobtain bracketleftbigg ∇2−1 c2∂2 ∂t2bracketrightbigg A=−µ0ρv, bracketleftbigg ∇2−1 c2∂2 ∂t2bracketrightbigg ϕ=−ρ ε0. (4.146) Now,thedifferentialoperator(see alsoExercise2.7.3) ∇2−1 c2∂2 ∂t2≡−∂2≡−∂µ∂µ is a four-dimensional Laplacian, usually called the d’Alembertian and also sometimes de- notedby /fill50.It isascalarbyconstruction(seeExercise2.7.3). Forconveniencewedefine A1≡Ax µ0c=cε0Ax,A3≡Az µ0c=cε0Az, A2≡Ay µ0c=cε0Ay,A 0≡ε0ϕ=A0.(4.147) If wefurther defineafour-vectorcurrentdensity ρvx c≡j1,ρvy c≡j2,ρvz c≡j3,ρ≡j0=j0, (4.148) thenEq. (4.146)maybewrittenintheform ∂2Aµ=jµ. (4.149) Thewaveequation(4.149)lookslikeafour-vectorequation,butlooksdonotconstitute proof.Toprovethatitisafour-vectorequation,westartbyinvestigatingthetransformation propertiesof thegeneralizedcurrent jµ. Sinceanelectricchargeelement deis aninvariantquantity,wehave de=ρdx1dx2dx3,invariant. (4.150) WesawinSection2.9thatthefour-dimensionalvolumeelement dx0dx1dx2dx3wasalso invariant,apseudoscalar.Comparingthisresult, Eq. (2.106), withEq. (4.150), wesee that thechargedensity ρmusttransformthesamewayas dx0,thezerothcomponentofafour- dimensionalvector dxλ.Weputρ=j0,withj0nowestablishedasthezerothcomponent ofafour-vector.Theotherparts ofEq. (4.148)maybeexpandedas j1=ρvx c=ρ cdx1 dt=j0dx1 dx0. (4.151) Sincewehavejustshownthat j0transformsas dx0,thismeansthat j1transformsas dx1. With similar results for j2andj3,W eh a v e jλtransforming as dxλ, proving that jλis a four-vectorinMinkowskispace. Equation (4.149), which follows directly from Maxwell’s equations, Eqs. (4.140), is assumed to hold in all Cartesian systems (all Lorentz frames). Then, by the quotient rule, Section2.8, AµisalsoavectorandEq. (4.149)is alegitimatetensorequation. 286 Chapter 4 Group Theory Now,workingbackward,Eq. (4.142)maybewritten ε0Ej=−∂Aj ∂x0−∂A0 ∂xj,j=1,2,3, (4.152) 1 µ0cBi=∂Ak ∂xj−∂Aj ∂xk,(i,j,k)=cyclic(1,2,3). We defineanewtensor, ∂µAλ−∂λAµ=∂Aλ ∂xµ−∂Aµ ∂xλ≡Fµλ=−Fλµ(µ,λ=0,1,2,3), anantisymmetricsecond-ranktensor,since Aλis avector.Writtenoutexplicitly, Fµλ ε0= 0ExEyEz −Ex0−cBzcBy −EycBz0−cBx −Ez−cBycBx0 ,Fµλ ε0= 0−Ex−Ey−Ez Ex0−cBzcBy EycBz0−cBx Ez−cBycBx0 . (4.153) Noticethatinourfour-dimensionalMinkowskispace EandBarenolongervectorsbutto- getherformasecond-ranktensor.Withthistensorwemaywritethetwononhomogeneous Maxwellequations((4.140b)and(4.140c)) combinedas atensorequation, ∂Fλµ ∂xµ=jλ. (4.154) Theleft-handsideofEq.(4.154)isafour-dimensionaldivergenceofatensorandtherefore a vector. This, of course, is equivalent to contracting a third-rank tensor ∂Fλµ/∂xν(com- pareExercises2.7.1and2.7.2).ThetwohomogeneousMaxwellequations—(4.140a)for ∇×Eand(4.140d)for ∇·B— maybeexpressedinthetensorform ∂F23 ∂x1+∂F31 ∂x2+∂F12 ∂x3=0 (4.155) forEq. (4.140d)andthreeequationsof theform −∂F30 ∂x2−∂F02 ∂x3+∂F23 ∂x0=0 (4.156) forEq. (4.140a). (Asecondequationpermutes120,athirdpermutes130.) Since ∂λFµν=∂Fµν ∂xλ≡tλµν isatensor(of thirdrank), Eqs. (4.140a)and(4.140d)aregivenbythetensorequation tλµν+tνλµ+tµνλ=0. (4.157) FromEqs.(4.155)and(4.156)youwillunderstandthattheindices λ,µ,andνaresupposed to be different. Actually Eq. (4.157) automatically reduces to 0 =0 if any two indices coincide.Analternateform ofEq. (4.157)appearsinExercise4.6.14. 4.6 Lorentz Covariance of Maxwell’s Equations 287 Lorentz Transformation of EandB Theconstructionofthetensorequations((4.154)and(4.157))completesourinitialgoalof rewriting Maxwell’s equations in tensor form.21Now we exploit the tensor properties of ourfourvectorsandthetensor Fµν. For the Lorentz transformation corresponding to motion along the z(x3)-axis with ve- locityv,the“directioncosines”aregivenby22 x′0=γparenleftbig x0−βx3parenrightbig x′3=γparenleftbig x3−βx0parenrightbig ,(4.158) where β=v c and γ=parenleftbig 1−β2parenrightbig−1/2. (4.159) Using the tensor transformation properties, we may calculate the electric and magnetic fields in the moving system in terms of the values in the original reference frame. From Eqs. (2.66), (4.153),and(4.158)weobtain E′ x=1radicalbig 1−β2parenleftbigg Ex−v c2Byparenrightbigg , E′ y=1radicalbig 1−β2parenleftbigg Ey+v c2Bxparenrightbigg , (4.160) E′ z=Ez and B′ x=1radicalbig 1−β2parenleftbigg Bx+v c2Eyparenrightbigg , B′ y=1radicalbig 1−β2parenleftbigg By−v c2Exparenrightbigg , (4.161) B′ z=Bz. Thiscouplingof EandBistobeexpected.Consider,forinstance,thecaseofzeroelectric fieldintheunprimedsystem Ex=Ey=Ez=0. 21Modern theories of quantum electrodynamics and elementary particles are often written in this “manifestly covariant” form to guarantee consistency with special relativity. Conversely, the insistence on such tensor form has been a useful guide in the construction ofthese theories. 22Agroup theoretic derivation of theLorentz transformation appears in Section 4.5. Seealso Goldstein, loc. cit.,Chapter6. 288 Chapter 4 Group Theory Clearly, there will be no force on a stationary charged particle. When the particle is in motion with a small velocity valong the z-axis,23an observer on the particle sees fields (exertingaforceonhischargedparticle)givenby E′ x=−vBy, E′ y=vBx, whereBisamagneticinductionfieldintheunprimedsystem.Theseequationsmaybeput invectorform, E′=v×B or (4.162) F=qv×B, whichisusuallytakenas theoperationaldefinitionof themagneticinduction B. Electromagnetic Invariants Finally, the tensor (or vector) properties allow us to construct a multitude of invariant quantities.Amoreimportantoneisthescalarproductofthetwofour-dimensionalvectors orfour-vectors Aλandjλ.W eha v e Aλjλ=−cε0Axρvx c−cε0Ayρvy c−cε0Azρvz c+ε0ϕρ =ε0(ρϕ−A·J),invariant, (4.163) withAthe usual magnetic vector potential and Jthe ordinary current density. The first term,ρϕ, is the ordinary static electric coupling, with dimensions of energy per unit vol- ume. Hence our newly constructed scalar invariant is an energy density. The dynamic in- teraction of field and current is given by the product A·J. This invariant Aλjλappears in theelectromagneticLagrangiansofExercises17.3.6and17.5.1. OtherpossibleelectromagneticinvariantsappearinExercises4.6.9 and4.6.11. The Lorentz group is the symmetry group of electrodynamics, of the electroweak gauge theory, and of the strong interactions described by quantum chromodynamics: It governs special relativity. The metric of Minkowski space–time is Lorentz invariant and expresses the propagation of light; that is, the velocity of light is the same in all inertial frames. Newton’sequationsofmotionarestraightforwardtoextendtospecialrelativity.Thekine- matics of two-body collisions are important applications of vector algebra in Minkowski space–time. 23If thevelocity is not small, arelativistic transformation offorce is needed. 4.6 Lorentz Covariance of Maxwell’s Equations 289 Exercises 4.6.1 (a) Show that every four-vector in Minkowski space may be decomposed into an or- dinary three-space vector and a three-space scalar. Examples: (ct,r),(ρ,ρv/c), (ε0ϕ,cε0A),(E/c,p),(ω/c,k). Hint.Considera rotationofthethree-spacecoordinateswithtimefixed. (b) Showthattheconverseof(a)is nottrue—everythree-vectorplusscalardoes not formaMinkowskifour-vector. 4.6.2 (a) Showthat ∂µjµ=∂·j=∂jµ ∂xµ=0. (b) Show how the previous tensor equation may be interpreted as a statement of con- tinuityofchargeandcurrentinordinarythree-dimensionalspaceandtime. (c) If this equation is known to hold in all Lorentz reference frames, why can we not concludethat jµis avector? 4.6.3 Write the Lorentz gauge condition (Eq. (4.143)) as a tensor equation in Minkowski space. 4.6.4 Agaugetransformationconsistsofvaryingthescalarpotential ϕ1andthevectorpoten- tialA1accordingtotherelation ϕ2=ϕ1+∂χ ∂t, A2=A1−∇χ. Thenewfunction χisrequiredtosatisfythehomogeneouswaveequation ∇2χ−1 c2∂2χ ∂t2=0. Showthefollowing: (a) TheLorentzgaugerelationisunchanged. (b) The new potentials satisfy the same inhomogeneous wave equations as did the originalpotentials. (c) Thefields EandBareunaltered. The invariance of our electromagnetic theory under this transformation is called gauge invariance . 4.6.5 Achargedparticle,charge q,massm, obeystheLorentzcovariantequation dpµ dτ=q ε0mcFµνpν, wherepνis the four-momentum vector (E/c;p1,p2,p3),τis the proper time, dτ= dtradicalbig 1−v2/c2, aLorentzscalar.Showthattheexplicitspace–timeformsare dE dt=qv·E;dp dt=q(E+v×B). 290 Chapter 4 Group Theory 4.6.6 From the Lorentz transformation matrix elements (Eq. (4.158)) derive the Einstein ve- locityadditionlaw u′=u−v 1−(uv/c2)oru=u′+v 1+(u′v/c2), whereu=cdx3/dx0andu′=cdx′3/dx′0. Hint.I fL12(v)is the matrix transforming system 1 into system 2, L23(u′)the matrix transforming system 2 into system 3, L13(u)the matrix transforming system 1 directly into system 3, then L13(u)=L23(u′)L12(v). From this matrix relation extract the Ein- steinvelocityadditionlaw. 4.6.7 The dual of a four-dimensional second-rank tensor Bmay be defined by ˜B, where the elementsofthedualtensorare givenby ˜Bij=1 2!εijklBkl. Showthat˜Btransforms as (a) asecond-ranktensorunderrotations, (b) apseudotensorunderinversions. Note.Thetildeheredoes notmeantranspose. 4.6.8 Construct˜F, thedualof F, whereFis theelectromagnetictensorgivenbyEq. (4.153). ANS.˜Fµν=ε0 0−cBx−cBy−cBz cBx0Ez−Ey cBy−Ez0Ex cBzEy−Ex0 . Thiscorrespondsto cB→−E, E→cB. This transformation,sometimescalleda dualtransformation ,leavesMaxwell’sequa- tionsinvacuum (ρ=0)invariant. 4.6.9 Because the quadruple contraction of a fourth-rank pseudotensor and two second-rank tensorsεµλνσFµλFνσisclearlyapseudoscalar,evaluateit. ANS.−8ε2 0cB·E. 4.6.10 (a) Ifanelectromagneticfieldispurelyelectric(orpurelymagnetic)inoneparticular Lorentz frame, show that EandBwill be orthogonal in other Lorentz reference systems. (b) Conversely,if EandBareorthogonalinoneparticularLorentzframe,thereexists aLorentzreferencesysteminwhich E(orB)vanishes.Findthatreferencesystem. 4.7 Discrete Groups 291 4.6.11 Showthat c2B2−E2is aLorentzscalar. 4.6.12 Since(dx0,dx1,dx2,dx3)is a four-vector, dxµdxµis a scalar. Evaluate this scalar for a movingparticlein twodifferent coordinatesystems:(a) a coordinatesystemfixed relativetoyou(labsystem),and(b)acoordinatesystemmovingwithamovingparticle (velocity vrelative to you). With the time increment labeled dτin the particle system anddtinthelabsystem, showthat dτ=dtradicalBig 1−v2/c2. τis thepropertimeoftheparticle,aLorentzinvariantquantity. 4.6.13 Expandthescalarexpression −1 4ε0FµνFµν+1 ε0jµAµ in terms of the fields and potentials. The resulting expression is the Lagrangian density usedinExercise17.5.1. 4.6.14 ShowthatEq.(4.157) maybewritten εαβγδ∂Fαβ ∂xγ=0. 4.7 D ISCRETE GROUPS Here we consider groups with a finite number of elements. In physics, groups usually ap- pear as a set of operations that leave a system unchanged, invariant. This is an expression ofsymmetry.Indeed,asymmetrymaybedefinedastheinvarianceoftheHamiltonianofa system under a group of transformations. Symmetry in this sense is important in classical mechanics, but it becomes even more important and more profound in quantum mechan- ics. In this section we investigate the symmetry properties of sets of objects (atoms in a molecule or crystal). This provides additional illustrations of the group concepts of Sec- tion4.1andleadsdirectlytodihedralgroups.Thedihedralgroupsinturnopenupthestudy of the 32 crystallographic point groups and 230 space groups that are of such importance in crystallography and solid-state physics. It might be noted that it was through the study of crystal symmetries that the concepts of symmetry and group theory entered physics. In physics, the abstract group conditions often take on direct physical meaning in terms of transformationsof vectors,spinors,andtensors. As a simple, but not trivial, example of a finite group, consider the set 1 ,a,b,cthat combine according to the group multiplication table24(see Fig. 4.10). Clearly, the four conditions of the definition of “group” are satisfied. The elements a,b,c, and 1 are ab- stract mathematical entities, completely unrestricted except for the multiplication table of Fig.4.10. Now,for aspecificrepresentationofthesegroupelements,let 1→1,a→i, b→−1,c→−i, (4.164) 24Theorder of the factors is row–column: ab=cin the indicatedprevious example. 292 Chapter 4 Group Theory FIGURE 4.10Group multiplicationtable. combining by ordinary multiplication. Again, the four group conditions are satisfied, and these four elements form a group. We label this group C4. Since the multiplication of the group elements is commutative, the group is labeled commutative ,o rabelian.O u r group is also a cyclic group , in that the elements may be written as successive powers of one element, in this case in,n=0,1,2,3. Note that in writing out Eq. (4.164) we have selectedaspecificfaithfulrepresentationfor thisgroupoffour objects, C4. We recognize that the group elements 1 ,i,−1,−imay be interpreted as successive 90◦ rotationsinthecomplexplane.Then,fromEq.(3.74),wecreatethesetoffour2 ×2matri- ces(replacing ϕby−ϕinEq. (3.74)torotateavectorratherthanrotatethecoordinates): R(ϕ)=parenleftBigg cosϕ−sinϕ sinϕcosϕparenrightBigg , andforϕ=0,π/2,π,and 3π/2w eh a v e 1=parenleftBigg 10 01parenrightBigg A=parenleftBigg 0−1 10parenrightBigg B=parenleftBigg −10 0−1parenrightBigg C=parenleftBigg 01 −10parenrightBigg .(4.165) Thissetoffourmatricesformsagroup,withthelawofcombinationbeingmatrixmultipli- cation. Here is a second faithful representation. By matrix multiplication one verifies that thisrepresentationisalsoabelianandcyclic.Clearly,thereisaone-to-onecorrespondence ofthetworepresentations 1↔1↔1a↔i↔Ab↔−1↔Bc↔−i↔C. (4.166) Inthegroup C4thetworepresentations (1,i,−1,−i)and(1,A,B,C)areisomorphic. In contrast to this, there is no such correspondence between either of these representa- tionsofgroup C4andanothergroupoffourobjects,thevierergruppe(Exercise3.2.7).The Table 4.3 1V1V2V3 11V1V2V3 V1V11V3V2 V2V2V31V1 V3V3V2V11 4.7 Discrete Groups 293 vierergruppe has the multiplicationtable shown in Table 4.3. Confirming the lack of cor- respondence between the group represented by (1,i,−1,−i)or the matrices (1,A,B,C) of Eq. (4.165), note that although the vierergruppe is abelian, it is not cyclic. The cyclic groupC4andthevierergruppearenotisomorphic. Classes and Character Consider a group element xtransformed into a group element yby a similarity transform withrespectto gi, anelementofthegroup gixg−1 i=y. (4.167) The group element yisconjugate tox.Aclassis a set of mutually conjugate group ele- ments.Ingeneral,thissetofelementsformingaclassdoesnotsatisfythegrouppostulates and is not a group. Indeed, the unit element 1, which is always in a class by itself, is the onlyclassthatisalsoasubgroup.Allmembersofagivenclassareequivalent,inthesense that any one element is a similarity transform of any other element. Clearly, if a group is abelian,everyelementis aclassbyitself. Wefindthat 1. Everyelementoftheoriginalgroupbelongstooneandonlyoneclass. 2. Thenumberofelementsinaclassis afactor oftheorderofthegroup. We get a possible physical interpretation of the concept of class by noting that yis a similaritytransformof x.Ifgirepresentsarotationofthecoordinatesystem,then yisthe sameoperationas xbutrelativetothenew,relatedcoordinates. In Section 3.3 we saw that a real matrix transforms under rotation of the coordinates by an orthogonal similarity transformation. Depending on the choice of reference frame, essentiallythesamematrixmaytakeonaninfinityofdifferentforms.Likewise,ourgroup representations may be put in an infinity of different forms by using unitary transforma- tions. But each such transformed representation is isomorphic with the original. From Ex- ercise3.3.9thetraceofeachelement(eachmatrixofourrepresentation)isinvariantunder unitarytransformations.Justbecauseitisinvariant,thetrace(relabeledthe character )as- sumesaroleofsomeimportanceingrouptheory,particularlyinapplicationstosolid-state physics. Clearly, all members of a given class (in a given representation) have the same character. Elements of different classes may have the same character, but elements with differentcharacterscannotbeinthesameclass. The concept of class is important (1) because of the trace or character and (2) because the number of nonequivalent irreducible representations of a group is equal to the numberofclasses. Subgroups and Cosets Frequently a subset of the group elements (including the unit element I) will by itself satisfythefourgrouprequirementsandthereforeisagroup.Suchasubsetiscalleda sub- group. Every group has two trivial subgroups: the unit element alone and the group itself. The elements 1 and bof the four-element group C4discussed earlier form a nontrivial 294 Chapter 4 Group Theory subgroup. In Section 4.1 we consider SO(3), the (continuous) group of all rotations in or- dinary space. The rotations about any single axis form a subgroup of SO(3). Numerous otherexamplesofsubgroupsappearinthefollowingsections. Considerasubgroup Hwithelements hiandagroupelement xnotinH.Thenxhiand hixarenotinsubgroup H. Thesets generatedby xhi,i=1,2,...andhix, i=1,2,... are called cosets, respectively the left and right cosets of subgroup Hwith respect to x.I t canbeshown(assumethecontraryandproveacontradiction)thatthecosetofasubgroup has the same number of distinct elements as the subgroup. Extending this result we may expresstheoriginalgroup Gasthesumof Handcosets: G=H+x1H+x2H+···. Then the order of any subgroup is a divisor of the order of the group . It is this result that makes the concept of coset significant. In the next section the six-element group D3 (order 6) has subgroups of order 1, 2, and 3. D3cannot (and does not) have subgroups of order4or5. Thesimilaritytransformofasubgroup Hbyafixedgroupelement xnotinH,xHx−1, yieldsasubgroup—Exercise4.7.8.Ifthisnewsubgroupisidenticalwith Hforallx,that is, xHx−1=H, thenHis called an invariant, normal ,o rself-conjugate subgroup . Such subgroups are involved in the analysis of multiplets of atomic and nuclear spectra and the particles dis- cussed in Section 4.2. All subgroups of a commutative (abelian) group are automatically invariant. Two Objects — Twofold Symmetry Axis Consider first the two-dimensional system of two identical atoms in the xy-plane at (1, 0) and (−1, 0), Fig. 4.11. What rotations25can be carried out (keeping both atoms in the xy-plane) that will leave this system invariant? The first candidate is, of course, the unit operator1.Arotationof πradiansaboutthe z-axiscompletesthelist.Sowehavearather uninteresting group of two members (1, −1). Thez-axis is labeled a twofold symmetry axis—correspondingtothetworotationangles,0and π, thatleavethesysteminvariant. Our system becomes more interesting in three dimensions. Now imagine a molecule (or part of a crystal) with atoms of element Xat±aon thex-axis, atoms of element Y at±bon they-axis, and atoms of element Zat±con thez-axis, as show in Fig. 4.12. Clearly,eachaxisisnowatwofoldsymmetryaxis.Using Rx(π)todesignatearotationof πradiansaboutthe x-axis,wemay 25Herewedeliberatelyexcludereflectionsandinversions.Theymustbebroughtintodevelopthefullsetof32crystallographic point groups. 4.7 Discrete Groups 295 FIGURE 4.11DiatomicmoleculesH 2,N2,O2, Cl2. FIGURE 4.12D2symmetry. setupamatrixrepresentationoftherotationsas inSection3.3: Rx(π)= 10 0 0−10 00−1 ,Ry(π)= −100 010 00−1 , Rz(π)= −100 0−10 001 , 1= 100 010 001 .(4.168) Thesefourelements [1,Rx(π),Ry(π),Rz(π)]formanabeliangroup,withthegroupmul- tiplicationtableshowninTable4.4. The products shown in Table 4.4 can be obtained in either of two distinct ways: (1) We may analyze the operations themselves—a rotation of πabout the x-axis fol- lowed by a rotation of πabout the y-axis is equivalent to a rotation of πabout the z-axis: Ry(π)Rx(π)=Rz(π). (2) Alternatively, once a faithful representation is established, we 296 Chapter 4 Group Theory Table 4.4 1R x(π)Ry(π)Rz(π) 1 1R xRyRx Rx(π)Rx1R zRy Ry(π)RyRz1R x Rz(π)RzRyRx1 can obtain the products by matrix multiplication. This is where the power of mathematics isshown—whenthesystemis toocomplexfor adirectphysicalinterpretation. Comparison with Exercises 3.2.7, 4.7.2, and 4.7.3 shows that this group is the vier- ergruppe. The matrices of Eq. (4.168) are isomorphic with those of Exercise 3.2.7. Also, they are reducible, being diagonal. The subgroups are (1,Rx),(1,Ry), and(1,Rz).T h e y are invariant. It should be noted that a rotation of πabout the y-axis and a rotation of π about the z-axis is equivalent to a rotation of πabout the x-axis:Rz(π)Ry(π)=Rx(π). In symmetry terms, if yandzare twofold symmetry axes, xis automatically a twofold symmetryaxis. This symmetry group,26the vierergruppe, is often labeled D2,t h eDsignifying a dihe- dralgroupandthesubscript2signifyingatwofoldsymmetryaxis(andnohighersymmetry axis). Three Objects — Threefold Symmetry Axis Consider now three identical atoms at the vertices of an equilateral triangle, Fig. 4.13. Rotationsofthe triangleof0,2π/3,and4π/3leavethetriangleinvariant.Inmatrixform, wehave27 1=Rz(0)=parenleftBigg 10 01parenrightBigg A=Rz(2π/3)=parenleftBigg cos2π/3−sin2π/3 sin2π/3 cos2 π/3parenrightBigg =parenleftBigg −1/2−√ 3/2 √ 3/2−1/2parenrightBigg B=Rz(4π/3)=parenleftBigg −1/2√ 3/2 −√ 3/2−1/2parenrightBigg . (4.169) Thez-axis is a threefold symmetry axis. (1,A,B)form a cyclic group, a subgroup of the completesix-elementgroupthatfollows. In thexy-plane there are three additional axes of symmetry—each atom (vertex) and thegeometriccenterdefininganaxis.Eachoftheseisatwofoldsymmetryaxis.Theserota- tionsmaymosteasilybedescribedwithinourtwo-dimensionalframeworkbyintroducing 26Asymmetry group isagroup ofsymmetry-preserving operations, thatis,rotations, reflections,andinversions. A symmetric group is the group of permutations of ndistinct objects—of order n!. 27Note thathere wearerotating the trianglecounterclockwise relative to fixed coordinates. 4.7 Discrete Groups 297 FIGURE 4.13Symmetryoperationsonan equilateraltriangle. reflections. The rotation of πabout the C-( o ry-) axis, which means the interchanging of (structureless)atoms aandc,is justareflectionofthe x-axis: C=RC(π)=parenleftBigg −10 01parenrightBigg . (4.170) We may replace the rotation about the D-axis by a rotation of 4 π/3 (about our z-axis) followedbyareflectionofthe x-axis(x→−x)(Fig.4.14): D=RD(π)=CB =parenleftBigg −10 01parenrightBiggparenleftBigg −1/2√ 3/2 −√ 3/2−1/2parenrightBigg =parenleftBigg 1/2−√ 3/2 −√ 3/2−1/2parenrightBigg . (4.171) FIGURE 4.14Thetriangleontherightisthetriangleon theleftrotated180◦aboutthe D-axis.D=CB. 298 Chapter 4 Group Theory In a similar manner, the rotation of πabout the E-axis, interchanging aandb, is replaced byarotationof 2 π/3(A)andthenareflection28ofthex-axis: E=RE(π)=CA =parenleftBigg −10 01parenrightBiggparenleftBigg −1/2−√ 3/2 √ 3/2−1/2parenrightBigg =parenleftBigg 1/2√ 3/2 √ 3/2−1/2parenrightBigg . (4.172) Thecompletegroupmultiplicationtableis 1ABCDE 11ABCDE AAB1DEC BB1AECD CCED1BA DDCEA1B EEDCBA1 Noticethateachelementofthegroupappearsonlyonceineachrowandineachcolumn,as requiredbytherearrangementtheorem,Exercise4.7.4.Also,fromthemultiplicationtable the group is not abelian. We have constructed a six-element group and a 2 ×2 irreducible matrix representation of it. The only other distinct six-element group is the cyclic group [1,R,R2,R3,R4,R5], with R=e2πi/6orR=e−πiσ2/3=parenleftBigg 1/2−√ 3/2 √ 3/21/2parenrightBigg . (4.173) Our group[1,A,B,C,D,E]is labeled D3in crystallography, the dihedral group with a threefold axis of symmetry. The three axes ( C,D, andE)i nt h exy-plane automatically become twofold symmetry axes. As a consequence, (1,C),(1,D), and(1,E)all form two-elementsubgroups.Noneofthesetwo-elementsubgroupsof D3isinvariant. Ageneralandmostimportantresultfor finitegroupsof helementsisthat summationdisplay in2 i=h, (4.174) whereniisthedimensionofthematricesofthe ithirreduciblerepresentation.Thisequal- ity, sometimes called the dimensionality theorem , is very useful in establishing the irre- ducible representations of a group. Here for D3we have 12+12+22=6 for our three representations. No other irreducible representations of this symmetry group of three ob- jects exist. (The other representations are the identity and ±1, depending upon whether a reflectionwasinvolved.) 28Note that, as a consequence of these reflections, det (C)=det(D)=det(E)=−1. The rotations AandB, of course, have a determinant of +1. 4.7 Discrete Groups 299 FIGURE 4.15Ruthenocene. Dihedral Groups, Dn Adihedralgroup Dnwithann-foldsymmetryaxisimplies naxeswithangularseparation of2π/nradians,nisapositiveinteger,butotherwiseunrestricted.Ifweapplythesymme- try arguments to crystal lattices , thennis limited to 1, 2, 3, 4, and 6. The requirement of invariance of the crystal lattice under translations in the plane perpendicular to the n-fold axis excludes n=5,7, and higher values. Try to cover a plane completely with identical regularpentagonsandwithnooverlapping.29Forindividualmolecules,thisconstraintdoes not exist, although the examples with n>6 are rare. n=5 is a real possibility. As an ex- ample,thesymmetrygroupforruthenocene, (C5H5)2Ru, illustratedinFig.4.15,is D5.30 Crystallographic Point and Space Groups The dihedral groups just considered are examples of the crystallographic point groups. A point group is composed of combinations of rotations and reflections (including inver- sions) that will leave some crystal lattice unchanged. Limiting the operations to rotations and reflections (including inversions) means that one point—the origin—remains fixed, hence the term point group . Including the cyclic groups, two cubic groups (tetrahedron andoctahedronsymmetries),andtheimproperforms(involvingreflections),wecometoa totalof 32crystallographicpointgroups. 29ForD6imagine a planecovered with regular hexagons and theaxis of rotation through the geometric center of one of them. 30Actuallythefull technicallabel is D5h,withhindicating invariance under a reflection of the fivefold axis. 300 Chapter 4 Group Theory If, to the rotation and reflection operations that produced the point groups, we add the possibility of translations and still demand that some crystal lattice remain invariant, we come to the space groups. There are 230 distinct space groups, a number that is appalling except,possibly,tospecialistsinthefield.Fordetails(whichcancoverhundredsofpages) seetheAdditionalReadings. Exercises 4.7.1 Showthatthematrices 1,A,B, andCofEq. (4.165)arereducible.Reducethem. Note.Thismeanstransforming AandCtodiagonalform(bythesameunitarytransfor- mation). Hint.AandCareanti-Hermitian.Theireigenvectorswillbeorthogonal. 4.7.2 Possibleoperationsonacrystallatticeinclude Aπ(rotationby π),m(reflection),and i (inversion).Thesethreeoperationscombineas A2 π=m2=i2=1, Aπ·m=i, m·i=Aπ,andi·Aπ=m. Showthatthegroup (1,Aπ,m,i)isisomorphicwiththevierergruppe. 4.7.3 Fourpossibleoperationsinthe xy-planeare: 1. nochangebraceleftBigg x→x y→y 2. inversionbraceleftBigg x→−x y→−y 3. reflectionbraceleftBigg x→−x y→y 4. reflectionbraceleftBigg x→x y→−y. (a) Showthatthesefour operationsformagroup. (b) Showthatthisgroupis isomorphicwiththevierergruppe. (c) Setupa 2 ×2 matrixrepresentation. 4.7.4 Rearrangement theorem: Given a group of n distinct elements (I,a,b,c,...,n) ,s h o w that the set of products (aI,a2,ab,ac...an) reproduces the ndistinct elements in a neworder. 4.7.5 Usingthe 2×2 matrixrepresentationofExercise3.2.7for thevierergruppe, (a) Showthatthereare fourclasses, eachwithoneelement. 4.7 Discrete Groups 301 (b) Calculate the character (trace) of each class. Note that two different classes may havethesamecharacter. (c) Show that there are three two-element subgroups. (The unit element by itself al- waysforms asubgroup.) (d) For any one of the two-element subgroups show that the subgroup and a single cosetreproducetheoriginalvierergruppe. Notethatsubgroups,classes, andcosetsareentirelydifferent. 4.7.6 Usingthe 2×2 matrixrepresentation,Eq. (4.165),of C4, (a) Showthatthereare fourclasses, eachwithoneelement. (b) Calculatethecharacter(trace)ofeachclass. (c) Showthatthereis onetwo-elementsubgroup. (d) Showthatthesubgroupandasinglecosetreproducetheoriginalgroup. 4.7.7 Prove that the number of distinct elements in a coset of a subgroup is the same as the numberof elementsinthesubgroup. 4.7.8 A subgroup Hhas elements hi.L e txbe a fixed element of the original group Gand notamemberof H. The transform xhix−1,i=1,2,... generates a conjugate subgroup xHx−1. Show that this conjugate subgroup satisfies eachof thefour grouppostulatesandthereforeis agroup. 4.7.9 (a) A particular group is abelian. A second group is created by replacing gibyg−1 i foreachelementintheoriginalgroup.Showthatthetwogroupsareisomorphic. Note.Thismeansshowingthatif aibi=ci, thena−1 ib−1 i=c−1 i. (b) Continuing part (a), if the two groups are isomorphic, show that each must be abelian. 4.7.10 (a) Once you have a matrix representation of any group, a one-dimensional represen- tation can be obtained by taking the determinants of the matrices. Show that the multiplicativerelationsarepreservedinthisdeterminantrepresentation. (b) Usedeterminantstoobtainaone-dimensionalrepresentativeof D3. 4.7.11 Explainhowtherelation summationdisplay in2 i=h appliestothevierergruppe (h=4)andtothedihedralgroup D3withh=6. 4.7.12 Showthatthesubgroup (1,A,B)ofD3is aninvariantsubgroup. 4.7.13 Thegroup D3maybediscussedasa permutation groupofthreeobjects.Matrix B,for instance,rotatesvertex a(originallyinlocation1)tothepositionformerlyoccupiedby c 302 Chapter 4 Group Theory (location 3). Vertex bmoves from location 2 to location 1, and so on. As a permutation (abc)→(bca).In threedimensions  010 001 100  a b c = b c a . (a) Developanalogous 3 ×3 representationsfor theotherelementsof D3. (b) Reduceyour 3 ×3 representationtothe 2 ×2 representationof thissection. (This 3×3 representationmustbereducibleor Eq.(4.174)wouldbeviolated.) Note. The actual reduction of a reducible representation may be awkward. It is often easiertodevelopdirectlyanewrepresentationof therequireddimension. 4.7.14 (a) The permutation group of four objects P4has 4!=24 elements. Treating the four elementsofthecyclicgroup C4aspermutations,setupa 4 ×4 matrixrepresenta- tionofC4.C4thatbecomesasubgroupof P4. (b) Howdoyouknowthatthis 4 ×4 matrixrepresentationof C4mustbereducible? Note.C4is abelian and every abelian group of hobjects has only hone-dimensional irreduciblerepresentations. 4.7.15 (a) Theobjects (abcd)arepermutedto (dacb).Writeouta4×4matrixrepresentation ofthisonepermutation. (b) Is thepermutation (abdc)→(dacb)oddor even? (c) Is thispermutationapossiblememberofthe D4group?Whyor whynot? 4.7.16 Theelementsofthedihedralgroup Dnmaybewrittenintheform SλRµ z(2π/n), λ=0,1 µ=0,1,...,n−1, whereRz(2π/n)representsarotationof2 π/naboutthe n-foldsymmetryaxis,whereas Srepresentsarotationof πaboutanaxisthroughthecenteroftheregularpolygonand oneof itsvertices. ForS=Eshowthatthisform maydescribethematrices A,B,C, andDofD3. Note. The elements RzandSare called the generators of this finite group. Similarly, iis thegeneratorofthegroupgivenbyEq.(4.164). 4.7.17 Show that the cyclic group of nobjects,Cn, may be represented by rm,m= 0,1,2,...,n−1.Hereris ageneratorgivenby r=exp(2πis/n). Theparameter stakesonthevalues s=1,2,3,...,n,eachvalueof syieldingadiffer- entone-dimensional(irreducible)representationof Cn. 4.7.18 Developtheirreducible2 ×2matrixrepresentationofthegroupofoperations(rotations andreflections)thattransformasquareintoitself. Givethegroupmultiplicationtable. Note. This is the symmetry group of a square and also the dihedral group D4. (See Fig.4.16.) 4.7 Discrete Groups 303 FIGURE 4.16 Square. FIGURE 4.17Hexagon. 4.7.19 Thepermutationgroupoffourobjectscontains4 !=24elements.FromExercise4.7.18, D4, the symmetry group for a square, has far fewer than 24 elements. Explain the rela- tionbetween D4andthepermutationgroupof fourobjects. 4.7.20 Aplaneis coveredwithregularhexagons,asshowninFig.4.17. (a) Determinethedihedralsymmetryofanaxisperpendiculartotheplanethroughthe common vertex of three hexagons (A). That is, if the axis has n-fold symmetry, show (with careful explanation) what nis. Write out the 2 ×2 matrix describing theminimum(nonzero)positiverotationofthearrayofhexagonsthatisamember ofyourDngroup. (b) Repeatpart(a)foranaxisperpendiculartotheplanethroughthegeometriccenter ofonehexagon (B). 4.7.21 In a simple cubic crystal, we might have identical atoms at r=(la,ma,na) , withl,m, andntakingonallintegralvalues. (a) ShowthateachCartesianaxisisafourfold symmetryaxis. (b) Thecubicgroupwillconsistofalloperations(rotations,reflections,inversion)that leave the simple cubic crystal invariant. From a consideration of the permutation 304 Chapter 4 Group Theory FIGURE 4.18 Multiplicationtable. ofthepositiveandnegativecoordinateaxes,predicthowmanyelementsthiscubic groupwillcontain. 4.7.22 (a) Fromthe D3multiplicationtableofFig.4.18constructasimilaritytransformtable showingxyx−1, wherexandyeachrangeoverallsixelementsof D3: (b) Divide the elements of D3into classes. Using the 2 ×2 matrix representation of Eqs. (4.169)–(4.172)notethetrace(character)ofeachclass. 4.8 D IFFERENTIAL FORMS InChapters1and2weadoptedtheviewthat,in ndimensions,avectorisan n-tupleofreal numbers and that its components transform properly under changes of the coordinates. In thissectionwestartfromthealternativeview,inwhichavectoristhoughtofasadirected line segment, an arrow. The point of the idea is this: Although the concept of a vector as a line segment does not generalize to curved space–time (manifolds of differential geom- etry), except by working in the flat tangent space requiring embedding in auxiliary extra dimensions, Elie Cartan’s differential forms are natural in curved space–time and a very powerful tool. Calculus can be based on differential forms, as Edwards has shown by his classic textbook (see the Additional Readings). Cartan’s calculus leads to a remarkable unification of concepts and theorems of vector analysis that is worth pursuing. In differ- ential geometry and advanced analysis (on manifolds) the use of differential forms is now widespread. Cartan’s notion of vector is based on the one-to-one correspondence between the linear spaces of displacement vectors and directional differential operators (components of the gradient form a basis). A crucial advantage of the latter is that they can be generalized to curved space–time. Moreover, describing vectors in terms of directional derivatives along curvesuniquelyspecifiesthevectoratagivenpointwithouttheneedtoinvokecoordinates. Ultimately, since coordinates are needed to specify points, the Cartan formalism, though anelegantmathematicaltoolfortheefficientderivationoftheoremsontensoranalysis,has inprinciplenoadvantageoverthecomponentformalism. 1-Forms We define dx,dy,dz in three-dimensional Euclidean space as functions assigning to a directed line segment PQfrom the point Pto the point Qthe corresponding change in x,y,z. The symbol dxrepresents “oriented length of the projection of a curve on the 4.8 Differential Forms 305 x-axis,” etc. Note that dx,dy,dz can be, but need not be, infinitesimally small, and they must not be confused with the ordinary differentials that we associate with integrals anddifferentialquotients.Afunctionof thetype Adx+Bdy+Cdz, A,B,C realnumbers (4.175) isdefinedasa constant1-form . Example 4.8.1 CONSTANT 1-FORM For a constant force F=(A,B,C) , the work done along the displacement from P= (3,2,1)toQ=(4,5,6)is thereforegivenby W=A(4−3)+B(5−2)+C(6−1)=A+3B+5C. IfFis a force field, then its rectangular components A(x,y,z),B(x,y,z),C(x,y,z) will depend on the location and the (nonconstant) 1 -formdW=F·drcorresponds to the concept of work done against the force field F(r)alongdron a space curve. A finite amountof work W=integraldisplay Cbracketleftbig A(x,y,z)dx+B(x,y,z)dy+C(x,y,z)dzbracketrightbig (4.176) involves the familiar line integral along an oriented curve C, where the 1-form dWde- scribestheamountofworkforsmalldisplacements(segmentsonthepath C).Inthislight, the integrand f(x)dxof an integralintegraltextb af(x)dxconsisting of the function fand of the measure dxas the oriented length is here considered to be a 1-form. The value of the integralisobtainedfromtheordinarylineintegral. /squaresolid 2-Forms Consideraunitflowofmassinthe z-direction,thatis,aflowinthedirectionofincreasing zso that a unit mass crosses a unit square of the xy-plane in unit time. The orientation symbolizedbythesequenceofpointsinFig.4.19, (0,0,0)→(1,0,0)→(1,1,0)→(0,1,0)→(0,0,0), will be called counterclockwise , as usual. A unit flow in the z-direction is defined by the functiondxdy31assigningtoorientedrectanglesinspacetheorientedareaoftheirprojec- tions on the xy-plane. Similarly, a unit flow in the x-direction is described by dydzand a unit flow in the y-direction by dzdx. The reverse order, dzdx, is dictated by the orienta- tion convention, and dzdx=−dxdzby definition. This antisymmetry is consistent with the cross product of two vectors representing oriented areas in Euclidean space. This no- tion generalizes to polygons and curved differentiable surfaces approximated by polygons andvolumes. 31Many authors denote this wedge product as dx∧dywithdy∧dx=−dx∧dy. Note that the product dxdy=dydxfor ordinary differentials. 306 Chapter 4 Group Theory FIGURE 4.19 Counterclockwise-oriented rectangle. Example 4.8.2 MAGNETIC FLUXACROSS AN ORIENTED SURFACE IfB=(A,B,C) isaconstantmagneticinduction,thentheconstant 2-form Adydz+Bdzdx+Cdxdy describesthemagneticflux across anorientedrectangle.If Bis a magneticinductionfield varyingacross asurface S,thentheflux /Phi1=integraldisplay Sbracketleftbig Bx(r)dydz+By(r)dzdx+Bz(r)dxdybracketrightbig (4.177) acrosstheorientedsurface Sinvolvesthefamiliar(Riemann)integrationoverapproximat- ingsmallorientedrectanglesfrom which Sis piecedtogether. /squaresolid Thedefinitionofintegraltext ωreliesondecomposing ω=summationtext iωi,wherethedifferentialforms ωi areeachnonzeroonlyinasmallpatchofthesurface Sthatcoversthesurface.Thenitcan be shown thatsummationtext iintegraltext ωiconverges, as the patches become smaller and more numerous, to the limitintegraltext ω, which is independent of these decompositions. For more details and proofs, werefer thereadertoEdwardsintheAdditionalReadings. 3-Forms A 3-form dxdydz represents an oriented volume. For example, the determinant of three vectors in Euclidean space changes sign if we reverse the order of two vectors. The determinant measures the oriented volume spanned by the three vectors. In particular,integraltext Vρ(x,y,z)dxdydz represents the total charge inside the volume Vifρis the charge density. Higher-dimensional differential forms in higher-dimensional spaces are defined similarlyandarecalled k-forms,with k=0,1,2,.... If a 3-form ω=A(x1,x2,x3)dx1dx2dx3=A′(x′ 1,x′ 2,x′ 3)dx′ 1dx′ 2dx′ 3(4.178) 4.8 Differential Forms 307 ona 3-dimensionalmanifoldisexpressedintermsofnewcoordinates,thenthereisaone- to-one,differentiablemap x′ i=x′ i(x1,x2,x3)betweenthesecoordinateswithJacobian J=∂(x′ 1,x′ 2,x′ 3) ∂(x1,x2,x3)=1, andA=A′J=A′so thatintegraldisplay Vω=integraldisplay VAdx1dx2dx3=integraldisplay V′A′dx′ 1dx′ 2dx′ 3. (4.179) This statement spells out the parameter independence of integrals over differential forms, since parameterizations are essentially arbitrary. The rules governing integration of differ- ential forms are defined on manifolds. These are continuous if we can move continuously (actually we assume them differentiable) from point to point, oriented if the orientation of curvesgeneralizestosurfacesandvolumesuptothedimensionofthewholemanifold.The rulesondifferentialforms are: •Ifω=aω1+a′ω′ 1,witha,a′realnumbers,thenintegraltext Sω=aintegraltext Sω1+a′integraltext Sω′ 1,whereSis acompact,oriented,continuousmanifoldwithboundary. •If theorientationis reversed,thentheintegralintegraltext Sωchangessign. Exterior Derivative Wenowintroducethe exteriorderivative dof afunction f, a0-form: df≡∂f ∂xdx+∂f ∂ydy+∂f ∂zdz=∂f ∂xidxi, (4.180) generating a 1-form ω1=df, the differential of f(or exterior derivative), the gradient in standard vector analysis. Upon summing over the coordinates, we have used and will continue to use Einstein’s summation convention. Applying the exterior derivative dto a 1-formwedefine d(Adx+Bdy+Cdz)=dAdx+dBdy+dCdz (4.181) with functions A,B,C. This definition in conjunction with dfas just given ties vectors to differential operators ∂i=∂ ∂xi. Similarly, we extend dtok-forms. However, applying dtwicegiveszero, ddf=0,because d(df)=d∂f ∂xdx+d∂f ∂ydy =parenleftbigg∂2f ∂x2dx+∂2f ∂x∂ydyparenrightbigg dx+parenleftbigg∂2f ∂y∂xdx+∂2f ∂y2dyparenrightbigg dy =parenleftbigg∂2f ∂y∂x−∂2f ∂x∂yparenrightbigg dxdy=0. (4.182) This follows from the fact that in mixed partial derivatives their order does not matter provided all functions are sufficiently differentiable. Similarly we can show ddω1=0f o r a 1-form ω1,et c. 308 Chapter 4 Group Theory Therulesgoverningdifferentialforms,with ωkdenotinga k-form,thatwehaveusedso far are •dxdx=0=dydy=dzdz,dx2 i=0; •dxdy=−dydx,dxidxj=−dxjdxi,i/negationslash=j, •dx1dx2···dxkis totallyantisymmetricinthe dxi,i=1,2,...,k. •df=∂f ∂xidxi; •d(ωk+/Omega1k)=dωk+d/Omega1k, linearity; •ddωk=0. Now we apply the exterior derivative dto products of differential forms, starting with functions(0-forms). Wehave d(fg)=∂(fg) ∂xidxi=parenleftbigg f∂g ∂xi+∂f ∂xigparenrightbigg dxi=fd g+dfg. (4.183) Ifω1=∂g ∂xidxiisa1-form and fisafunction,then d(fω1)=dparenleftbigg f∂g ∂xidxiparenrightbigg =dparenleftbigg f∂g ∂xiparenrightbigg dxi =∂parenleftbig f∂g ∂xiparenrightbig ∂xjdxjdxi=parenleftbigg∂f ∂xj∂g ∂xi+f∂2g ∂xi∂xjparenrightbigg dxjdxi =dfω1+fdω1, (4.184) asexpected.Butif ω′ 1=∂f ∂xjdxjisanother1-form,then d(ω1ω′ 1)=dparenleftbigg∂g ∂xidxi∂f ∂xjdxjparenrightbigg =dparenleftbigg∂g ∂xi∂f ∂xjparenrightbigg dxidxj =∂parenleftBig ∂g ∂xi∂f ∂xjparenrightBig ∂xkdxkdxidxj =∂2g ∂xi∂xkdxkdxi∂f ∂xjdxj−∂g ∂xidxi∂2f ∂xj∂xkdxkdxj =dω1ω′ 1−ω1dω′ 1. (4.185) This proof is valid for more general 1-forms ω=fidxiwith functions fi. In general, therefore,wedefinefor k-forms: d(ωkω′ k)=(dωk)ω′ k+(−1)kωk(dω′ k). (4.186) Ingeneral,theexteriorderivativeofa k-form isa (k+1)-form. 4.8 Differential Forms 309 Example 4.8.3 POTENTIAL ENERGY Asanapplicationintwodimensions(forsimplicity),considerthepotential V(r),a0-form, anddV, its exteriorderivative.Integrating Valonganorientedpath Cfromr1tor2gives V(r2)−V(r1)=integraldisplay CdV=integraldisplay Cparenleftbigg∂V ∂xdx+∂V ∂ydyparenrightbigg =integraldisplay C∇V·dr, (4.187) wherethelastintegralisthestandardformulaforthepotentialenergydifferencethatforms part of the energy conservation theorem. The path and parameterization independence are manifestinthisspecialcase. /squaresolid Pullbacks If alinearmap L2from theuv-planetothe xy-planehas theform x=au+bv+c, y=eu+fv+g, (4.188) oriented polygons in the uv-plane are mapped onto similar polygons in the xy-plane, pro- videdthedeterminant af−beofthemap L2isnonzero.The2-form dxdy=(adu+bdv)(edu+fdv )=(af−be)dudv (4.189) can be pulled back from the xy-t ot h euv-plane. That is to say, an integral over a simply connectedsurface Sbecomesintegraldisplay L2(S)dxdy=(af−be)integraldisplay Sdudv, (4.190) and(af−be)dudv isthepullbackof dxdy,oppositetothedirectionofthemap L2from theuv-plane to the xy-plane. Of course, the determinant af−beof the map L2is simply theJacobian,generatedwithouteffort bythedifferentialforms inEq.(4.189). Similarly,alinearmap L3fromtheu1u2u3-spacetothe x1x2x3-space xi=aijuj+bi,i=1,2,3, (4.191) automaticallygeneratesits Jacobianfromthe3-form dx1dx2dx3=parenleftbigg3summationdisplay j=1a1jdujparenrightbiggparenleftbigg3summationdisplay j=1a2jdujparenrightbiggparenleftbigg3summationdisplay j=1a3jdujparenrightbigg =(a11a22a33−a12a21a33±···)du1du2du3 =det a11a12a13 a21a22a23 a31a32a33du1du2du3. (4.192) Thus,differentialforms generatetherules governingdeterminants. Given two linear maps in a row, it is straightforward to prove that the pullback under a composed map is the pullback of the pullback. This theorem is the differential-forms analogofmatrixmultiplication. 310 Chapter 4 Group Theory Let us now consider a curve Cdefined by a parameter tin contrast to a curve defined by an equation. For example, the circle {(cost,sint);0≤t≤2π}is a parameterization byt, whereas the circle {(x,y);x2+y2=1}is a definition by an equation. Then the line integral integraldisplay Cbracketleftbig A(x,y)dx+B(x,y)dybracketrightbig =integraldisplaytf tibracketleftbigg Adx dt+Bdy dtbracketrightbigg dt (4.193) forcontinuousfunctions A,B,dx/dt,dy/dt becomesaone-dimensionalintegraloverthe oriented interval ti≤t≤tf. Clearly, the 1-form [Adx dt+Bdy dt]dton thet-line is obtained from the 1-form Adx+Bdyon thexy-plane via the map x=x(t),y=y(t)from the t-linetothecurve Cinthexy-plane.The1-form [Adx dt+Bdy dt]dtiscalledthepullbackof the 1-form Adx+Bdyunder the map x=x(t),y=y(t). Using pullbacks we can show thatintegralsover1-formsare independentof theparameterizationof thepath. Inthissense,thedifferentialquotientdy dxcanbeconsideredasthecoefficientof dxinthe pullback of dyunder the function y=f(x),o rdy=f′(x)dx. This concept of pullback readily generalizes to maps in three or more dimensions and to k-forms with k>1. In particular,thechainrulecanbeseentobeapullback:If yi=fi(x1,x2,...,xn), i=1,2,...,land zj=gj(y1,y2,...,yl), j=1,2,...,m (4.194) are differentiable maps from Rn→RlandRl→Rm, then the composed map Rn→Rm is differentiable and the pullback of any k-form under the composed map is equal to the pullback of the pullback. This theorem is useful for establishing that integrals of k-forms areparameterindependent. Similarly,wedefinethedifferential dfasthepullbackofthe1-form dzunderthefunc- tionz=f(x,y): dz=df=∂f ∂xdx+∂f ∂ydy. (4.195) Example 4.8.4 STOKES ’THEOREM As another application let us first sketch the standard derivation of the simplest version of Stokes’ theorem for a rectangle S=[a≤x≤b,c≤y≤d]oriented counterclockwise, with∂Sits boundary integraldisplay ∂S(Adx+Bdy)=integraldisplayb aA(x,c)dx+integraldisplayd cB(b,y)dy+integraldisplaya bA(x,d)dx+integraldisplayc dB(a,y)dy =integraldisplayd cbracketleftbig B(b,y)−B(a,y)bracketrightbig dy−integraldisplayb abracketleftbig A(x,d)−A(x,c)bracketrightbig dx =integraldisplayd cintegraldisplayb a∂B ∂xdxdy−integraldisplayb aintegraldisplayd c∂A ∂ydydx =integraldisplay Sparenleftbigg∂B ∂x−∂A ∂yparenrightbigg dxdy, (4.196) 4.8 Differential Forms 311 whichholdsforanysimplyconnectedsurface Sthatcanbepiecedtogetherbyrectangles. Now we demonstrate the use of differential forms to obtain the same theorem (again in twodimensionsforsimplicity): d(Adx+Bdy)=dAdx+dBdy =parenleftbigg∂A ∂xdx+∂A ∂ydyparenrightbigg dx+parenleftbigg∂B ∂xdx+∂B ∂ydyparenrightbigg dy=parenleftbigg∂B ∂x−∂A ∂yparenrightbigg dxdy, (4.197) using the rules highlighted earlier. Integrating over a surface Sand its boundary ∂S,r e - spectively,weobtain integraldisplay ∂S(Adx+Bdy)=integraldisplay Sd(Adx+Bdy)=integraldisplay Sparenleftbigg∂B ∂x−∂A ∂yparenrightbigg dxdy. (4.198) Here contributions to the left-hand integral from inner boundaries cancel as usual because they are oriented in opposite directions on adjacent rectangles. For each oriented inner rectanglethatmakesupthesimplyconnectedsurface Swehaveused, integraldisplay Rddx=integraldisplay ∂Rdx=0. (4.199) Notethattheexteriorderivativeautomaticallygeneratesthe zcomponentof thecurl. In threedimensions,Stokes’theoremderivesfrom thedifferential-form identityinvolv- ingthevectorpotential Aandmagneticinduction B=∇×A, d(Axdx+Aydy+Azdz)=dAxdx+dAydy+dAzdz =parenleftbigg∂Ax ∂xdx+∂Ax ∂ydy+∂Ax ∂zdzparenrightbigg dx+··· =parenleftbigg∂Az ∂y−∂Ay ∂zparenrightbigg dydz+parenleftbigg∂Ax ∂z−∂Az ∂xparenrightbigg dzdx+parenleftbigg∂Ay ∂x−∂Ax ∂yparenrightbigg dxdy, (4.200) generatingallcomponentsofthecurlinthree-dimensionalspace.Thisidentityisintegrated over each oriented rectangle that makes up the simply connected surface S(which has no holes, that is, where every curve contracts to a point of the surface) and then is summed overalladjacentrectanglestoyieldthemagneticfluxacross S, /Phi1=integraldisplay S[Bxdydz+Bydzdx+Bzdxdy] =integraldisplay ∂S[Axdx+Aydy+Azdz], (4.201) or,inthestandardnotationof vectoranalysis(Stokes’theorem,Chapter1), integraldisplay SB·da=integraldisplay S(∇×A)·da=integraldisplay ∂SA·dr. (4.202) /squaresolid 312 Chapter 4 Group Theory Example 4.8.5 GAUSS’THEOREM Consider Gauss’ law, Section 1.14. We integrate the electric density ρ=1 ε0∇·Eover the volume of a single parallelepiped V=[a≤x≤b,c≤y≤d,e≤z≤f]oriented by dxdydz (right-handed), the side x=bofVis oriented by dydz(counterclockwise, as seenfrom x>b),andsoon.Using Ex(b,y,z)−Ex(a,y,z)=integraldisplayb a∂Ex ∂xdx, (4.203) we have, in the notation of differential forms, summing over all adjacent parallelepipeds thatmakeupthevolume V, integraldisplay ∂VExdydz=integraldisplay V∂Ex ∂xdxdydz. (4.204) Integratingtheelectricflux(2-form) identity d(Exdydz+Eydzdx+Ezdxdy)=dExdydz+dEydzdx+dEzdxdy =parenleftbigg∂Ex ∂x+∂Ey ∂y+∂Ez ∂zparenrightbigg dxdydz (4.205) acrossthesimplyconnectedsurface ∂VwehaveGauss’theorem, integraldisplay ∂V(Exdydz+Eydzdx+Ezdxdy)=integraldisplay Vparenleftbigg∂Ex ∂x+∂Ey ∂y+∂Ez ∂zparenrightbigg dxdydz, (4.206) or,instandardnotationofvectoranalysis, integraldisplay ∂VE·da=integraldisplay V∇·Ed3r=q ε0. (4.207) /squaresolid Theseexamplesaredifferentcasesofasingletheoremondifferentialforms.Toexplain why, let us begin with some terminology, a preliminary definition of a differentiable manifold M : It is a collectionof points ( m-tuples of real numbers) that are smoothly(that is, differentiably) connected with each other so that the neighborhood of each point looks likeasimplyconnectedpieceofan m-dimensionalCartesianspace“closeenough”around thepointandcontainingit.Here, m,whichstaysconstantfrompointtopoint,iscalledthe dimension of the manifold. Examples are the m-dimensional Euclidean space Rmand the m-dimensionalsphere Sm=bracketleftbiggparenleftbig x1,...,xm+1parenrightbig ;m+1summationdisplay i=1parenleftbig xiparenrightbig2=1bracketrightbigg . Any surface with sharp edges, corners, or kinks is not a manifold in our sense, that is, is not differentiable. In differential geometry, all movements, such as translation and paralleldisplacement,arelocal,thatis,aredefinedinfinitesimally.Ifweapplytheexterior derivative dtoafunction f(x1,...,xm)onM, wegeneratebasic 1-forms: df=∂f ∂xidxi, (4.208) 4.8 Differential Forms 313 wherexi(P)arecoordinatefunctions.Asbeforewehave d(df)=0 because d(df)=dparenleftbigg∂f ∂xiparenrightbigg dxi=∂2f ∂xj∂xidxjdxi =summationdisplay j<iparenleftbigg∂2f ∂xj∂xi−∂2f ∂xi∂xjparenrightbigg dxjdxi=0 (4.209) because the order of derivatives does not matter. Any 1-form is a linear combination ω=summationtext iωidxiwithfunctions ωi. Generalized Stokes’ Theorem on Differential Forms Letωbeacontinuous (k−1)-forminx1x2···xn-spacedefinedeverywhereonacompact, oriented, differentiable k-dimensional manifold Swith boundary ∂Sinx1x2···xn-space. Thenintegraldisplay ∂Sω=integraldisplay Sdω. (4.210) Here dω=d(Adx 1dx2···dxk−1+···)=dAdx1dx2···dxk−1+···.(4.211) ThepotentialenergyinExample4.8.3giventhistheoremforthepotential ω=V,a0-form; Stokes’theoreminExample4.8.4isthistheoremforthevectorpotential1-formsummationtext iAidxi (forEuclideanspaces dxi=dxi);andGauss’theoreminExample4.8.5isStokes’theorem fortheelectricflux2-forminthree-dimensionalEuclideanspace. The method of integration by parts can be generalized to differential forms using Eq.(4.186): integraldisplay Sdω1ω2=integraldisplay ∂Sω1ω2−(−1)k1integraldisplay Sω1dω2. (4.212) Thisis provedbyintegratingtheidentity d(ω1ω2)=dω1ω2+(−1)k1ω1dω2, (4.213) withtheintegratedtermintegraltext Sd(ω1ω2)=integraltext ∂Sω1ω2. Our next goal is to cast Sections 2.10 and 2.11 in the language of differential forms. So far wehaveworkedintwo-or three-dimensionalEuclideanspace. Example 4.8.6 RIEMANN MANIFOLD Let us look at the curved Riemann space–time of Sections 2.10–2.11 and reformulate some of this tensor analysis in curved spaces in the language of differential forms. Re- callthatdishinguishingbetweenupperandlowerindicesisimportanthere.Themetric gij inEq.(2.123) canbewritteninterms oftangentvectors,Eq.(2.114), asfollows: gij=∂xl ∂qi∂xl ∂qj, (4.214) 314 Chapter 4 Group Theory where the sum over the index ldenotes the inner product of the tangent vectors. (Here we continue to use Einstein’s summation convention over repeated indices. As before, the metric tensor is used to raise and lower indices.) The key concept of connection involves theChristoffelsymbols,whichweaddressfirst.Theexteriorderivativeofatangentvector canbeexpandedintermsof thebasis oftangentvectors(compareEq.(2.131a)), dparenleftbigg∂xl ∂qiparenrightbigg =Ŵkij∂xl ∂qkdqj, (4.215) thusintroducingtheChristoffelsymbolsofthesecondkind.Applying dtoEq.(4.214)we obtain dgij=∂gij ∂qmdqm=dparenleftbigg∂xl ∂qiparenrightbigg∂xl ∂qj+∂xl ∂qidparenleftbigg∂xl ∂qjparenrightbigg (4.216) =parenleftbigg Ŵkim∂xl ∂qk∂xl ∂qj+Ŵkjm∂xl ∂qi∂xl ∂qkparenrightbigg dqm=parenleftbig Ŵkimgkj+Ŵkjmgikparenrightbig dqm. Comparingthecoefficientsof dqmyields ∂gij ∂qm=Ŵkimgkj+Ŵkjmgik. (4.217) UsingtheChristoffel symbolof thefirst kind, [ij,m]=gkmŴkij, (4.218) wecanrewriteEq. (4.217)as ∂gij ∂qm=[im,j]+[jm,i], (4.219) whichcorrespondstoEq.(2.136)andimpliesEq. (2.137). Wecheckthat [ij,m]=1 2parenleftbigg∂gim ∂qj+∂gjm ∂qi−∂gij ∂qmparenrightbigg (4.220) istheuniquesolutionof Eq.(4.219)andthat Ŵkij=gmk[ij,m]=1 2gmkparenleftbigg∂gim ∂qj+∂gjm ∂qi−∂gij ∂qmparenrightbigg (4.221) follows. /squaresolid Hodge∗Operator The differentials dxi,i=1,2,...,m, form a basis of a vector space that is (seen to be) dualtothederivatives ∂i=∂ ∂xi;theyarebasic1-forms.Forexample,thevectorspace V= {(a1,a2,a3)}isdualtothevectorspaceofplanes(linearfunctions f)inthree-dimensional Euclideanspace V∗={f≡a1x1+a2x2+a3x3−d=0}.Thegradient ∇f=parenleftbigg∂f ∂x1,∂f ∂x2,∂f ∂x3parenrightbigg =(a1,a2,a3) (4.222) 4.8 Differential Forms 315 providesaone-to-one,differentiablemapfrom V∗toV.Suchdualrelationshipsaregener- alizedbytheHodge∗operator,basedontheLevi-CivitasymbolofSection2.9. Let theunitvectors ˆxibeanorientedorthonormalbasis ofthree-dimensionalEuclidean space.ThentheHodge∗of scalarsis definedbythebasiselement ∗1≡1 3!εijkˆxiˆxjˆxk=ˆx1ˆx2ˆx3, (4.223) which corresponds to (ˆx1׈x2)·ˆx3in standard vector notation. Here ˆxiˆxjˆxkis the totally antisymmetric exterior product of the unit vectors that corresponds to (ˆxi׈xj)·ˆxkin standardvectornotation.Forvectors,∗is definedfor thebasis ofunitvectorsas ∗ˆxi≡1 2!εijkˆxjˆxk. (4.224) Inparticular, ∗ˆx1=ˆx2ˆx3,∗ˆx2=ˆx3ˆx1,∗ˆx3=ˆx1ˆx2. (4.225) Fororientedareas,∗is definedonbasisareaelementsas ∗(ˆxiˆxj)≡εkijˆxk, (4.226) so ∗(ˆx1ˆx2)=ε312ˆx3=ˆx3,∗(ˆx1ˆx3)=ε213ˆx2=−ˆx2, ∗(ˆx2ˆx3)=ε123ˆx1=ˆx1. (4.227) Forvolumes,∗is definedas ∗(ˆx1ˆx2ˆx3)≡ε123=1. (4.228) Example 4.8.7 CROSS PRODUCT OF VECTORS Theexteriorproductof twovectors a=3summationdisplay i=1aiˆxi,b=3summationdisplay i=1biˆxi (4.229) isgivenby ab=parenleftbigg3summationdisplay i=1aiˆxiparenrightbiggparenleftbigg3summationdisplay j=1biˆxjparenrightbigg =summationdisplay i<jparenleftbig aibj−ajbiparenrightbigˆxiˆxj, (4.230) whereasEq. (4.224)impliesthat ∗(ab)=a×b. (4.231) /squaresolid Next, let us analyze Sections 2.1–2.2 on curvilinear coordinates in the language of dif- ferentialforms. 316 Chapter 4 Group Theory Example 4.8.8 LAPLACIAN IN ORTHOGONAL COORDINATES Considerorthogonalcoordinateswherethemetric(Eq. (2.5)) leadstolengthelements dsi=hidqi,notsummed . (4.232) Herethedqiareordinarydifferentials.The 1-formsassociatedwiththedirections ˆqiare εi=hidqi,notsummed . (4.233) Thenthegradientis definedbythe 1-form df=∂f ∂qidqi=parenleftbigg1 hi∂f ∂qiparenrightbigg εi. (4.234) Weapplythehodgestar operatorto df, generatingthe 2-form ∗df=parenleftbigg1 hi∂f ∂qiparenrightbigg ∗εi=parenleftbigg1 h1∂f ∂q1parenrightbigg ε2ε3+parenleftbigg1 h2∂f ∂q2parenrightbigg ε3ε1+parenleftbigg1 h3∂f ∂q3parenrightbigg ε1ε2 =parenleftbiggh2h3 h1∂f ∂q1parenrightbigg dq2dq3+parenleftbiggh1h3 h2∂f ∂q2parenrightbigg dq3dq1+parenleftbiggh1h2 h3∂f ∂q3parenrightbigg dq1dq2. (4.235) Applyinganotherexteriorderivative d, wegettheLaplacian d(∗df)=∂ ∂q1parenleftbiggh2h3 h1∂f ∂q1parenrightbigg dq1dq2dq3+∂ ∂q2parenleftbiggh1h3 h2∂f ∂q2parenrightbigg dq2dq1dq2dq3 +∂ ∂q3parenleftbiggh1h2 h3∂f ∂q3parenrightbigg dq3dq1dq2dq3 =1 h1h2h3bracketleftbigg∂ ∂q1parenleftbiggh2h3 h1∂f ∂q1parenrightbigg +∂ ∂q2parenleftbiggh1h3 h2∂f ∂q2parenrightbigg +∂ ∂q3parenleftbiggh1h2 h3∂f ∂q3parenrightbiggbracketrightbigg ·ε1ε2ε3=∇2fdq1dq2dq3. (4.236) DividingbythevolumeelementgivesEq.(2.22).Recallthatthevolumeelements dxdydz andε1ε2ε3mustbeequalbecause εianddx,dy,dz areorthonormal1-formsandthemap from thexyztotheqicoordinatesis one-to-one. /squaresolid Example 4.8.9 MAXWELL ’SEQUATIONS Wenowworkinfour-dimensionalMinkowskispace,thehomogeneous,flatspace–timeof special relativity, to discuss classical electrodynamics in terms of differential forms. We start by introducing the electromagnetic field 2-form (field tensor in standard relativistic notation): F=−Exdtdx−Eydtdy−Ezdtdz+Bxdydz+Bydzdx+Bzdxdy =1 2Fµνdxµdxν, (4.237) 4.8 Differential Forms 317 which contains the electric 1-form E=Exdx+Eydy+Ezdzand the magnetic flux 2-form. Here, terms with 1-forms in opposite order have been combined. (For Eq. (4.237) tobevalid,themagneticinductionisinunitsof c;thatis,Bi→cBi,withcthevelocityof light; or we work in units where c=1. Also,Fis in units of 1 /ε0, the dielectric constant of the vacuum. Moreover, the vector potential is defined as A0=ε0φ, with the nonstatic electric potential φandA1=Ax µ0c,...;see Section 4.6 for more details.) The field 2-form Fencompasses Faraday’s induction law that a moving charge is acted on by magnetic forces. Applying the exterior derivative dgenerates Maxwell’s homogeneous equations auto- maticallyfrom F: dF=−parenleftbigg∂Ex ∂ydy+∂Ex ∂zdzparenrightbigg dtdx−parenleftbigg∂Ey ∂xdx+∂Ey ∂zdzparenrightbigg dtdy −parenleftbigg∂Ez ∂xdx+∂Ez ∂ydyparenrightbigg dtdz+parenleftbigg∂Bx ∂xdx+∂Bx ∂tdtparenrightbigg dydz +parenleftbigg∂By ∂tdt+∂By ∂ydyparenrightbigg dzdx+parenleftbigg∂Bz ∂tdt+∂Bz ∂zdzparenrightbigg dxdy =parenleftbigg −∂Ex ∂y+∂Ey ∂x+∂Bz ∂tparenrightbigg dtdxdy+parenleftbigg −∂Ex ∂z+∂Ez ∂x−∂By ∂tparenrightbigg dtdxdz +parenleftbigg −∂Ey ∂z+∂Ez ∂y+∂Bx ∂tparenrightbigg dtdydz=0 (4.238) which,instandardnotationofvectoranalysis,takesthefamiliarvectorformofMaxwell’s homogeneousequations, ∇×E+∂B ∂t=0. (4.239) SincedF=0, that is, there is no driving term so that Fis closed, there must be a 1-form ω=Aµdxµso thatF=dω.No w , dω=∂νAµdxνdxµ, (4.240) which, in standard notation, leads to the conventional relativistic form of the electromag- neticfieldtensor, Fµν=∂µAν−∂νAµ. (4.241) Maxwell’shomogeneousequations, dF=0,are thusequivalentto ∂νFµν=0. In order to derive similarly the inhomogeneous Maxwell’s equations, we introduce the dualelectromagneticfieldtensor ˜Fµν=1 2εµναβFαβ, (4.242) and,intermsofdifferentialforms, ∗F=∗parenleftbig Fµνdxµdxνparenrightbig =Fµν∗parenleftbig dxµdxνparenrightbig =1 2Fµνεµναβdxαdxβ.(4.243) 318 Chapter 4 Group Theory Applyingtheexteriorderivativeyields d(∗F)=1 2εµναβ(∂γFµν)dxγdxαdxβ, (4.244) theleft-handsideofMaxwell’sinhomogeneousequations,a3-form.Itsdrivingtermisthe dualof theelectriccurrentdensity,a3-form: ∗J=Jαparenleftbig ∗dxαparenrightbig =Jαεαµνλdxµdxνdxλ =ρdxdydz−Jxdtdydz−Jydtdzdx−Jzdtdxdy. (4.245) AltogetherMaxwell’sinhomogeneousequationstaketheelegantform d(∗F)=∗J. (4.246) /squaresolid The differential-formframework has brought considerableunification to vector algebra and to tensor analysis on manifolds more generally, such as uniting Stokes’ and Gauss’ theorems and providing an elegant reformulation of Maxwell’s equations and an efficient derivationoftheLaplacianincurvedorthogonalcoordinates,amongothers. Exercises 4.8.1 Evaluate the 1-form adx+2bdy+4cdzon the line segment PQ, withP=(3,5,7), Q=(7,5,3). 4.8.2 Iftheforcefieldisconstantandmovingaparticlefromtheoriginto (3,0,0)requiresa units of work, from (−1,−1,0)to(−1,1,0)takesbunits of work, and from (0,0,4) to(0,0,5)cunitsofwork, findthe1-formof thework. 4.8.3 Evaluatetheflowdescribedbythe2-form dxdy+2dydz+3dzdxacrosstheoriented trianglePQRwithcornersat P=(3,1,4), Q=(−2,1,4), R=(1,4,1). 4.8.4 Arethepoints,inthisorder, (0,1,1), (3,−1,−2), (4,2,−2), (−1,0,1) coplanar,or dotheyformanorientedvolume(right-handedor left-handed)? 4.8.5 Write Oersted’slaw, integraldisplay ∂SH·dr=integraldisplay S∇×H·da∼I, indifferentialform notation. 4.8.6 Describetheelectricfieldbythe1-form E1dx+E2dy+E3dzandthemagneticinduc- tionbythe2-form B1dydz+B2dzdx+B3dxdy.ThenformulateFaraday’sinduction lawintermsof theseforms. 4.8 Additional Readings 319 4.8.7 Evaluatethe1-form xdy x2+y2−ydx x2+y2 ontheunitcircleabouttheoriginorientedcounterclockwise. 4.8.8 Findthepullbackof dxdzunderx=ucosv,y=u−v,z=usinv. 4.8.9 Find the pullback of the 2-form dydz+dzdx+dxdyunder the map x=sinθcosϕ, y=sinθsinϕ,z=cosθ. 4.8.10 Parameterizethesurfaceobtainedbyrotatingthecircle (x−2)2+z2=1,y=0,about thez-axis inacounterclockwiseorientation,as seenfromoutside. 4.8.11 A 1-form Adx+Bdyis defined as closedif∂A ∂y=∂B ∂x. It is called exactif there is a functionfso that∂f ∂x=Aand∂f ∂y=B. Determine which of the following 1-forms are closed,orexact,andfindthecorrespondingfunctions fforthosethatareexact: ydx+xdy,ydx+xdy x2+y2,bracketleftbig ln(xy)+1bracketrightbig dx+x ydy, −ydx x2+y2+xdy x2+y2,f(z)dzwithz=x+iy. 4.8.12 Show thatsummationtextn i=1x2 i=a2defines a differentiable manifold of dimension D=n−1i f a/negationslash=0 andD=0i fa=0. 4.8.13 Show that the set of orthogonal 2 ×2 matrices form a differentiable manifold, and determineitsdimension. 4.8.14 Determine the value of the 2-form Adydz+Bdzdx+Cdxdyon a parallelogram withsides a,b. 4.8.15 ProveLorentzinvarianceof Maxwell’sequationsinthelanguageofdifferentialforms. AdditionalReadings Buerger, M. J., Elementary Crystallography . New York: Wiley (1956). A comprehensive discussion of crystal symmetries. Buerger develops all 32 point groups and all 230 space groups. Related books by this author in- cludeContemporaryCrystallography .NewYork:McGraw-Hill(1970); CrystalStructureAnalysis .NewY ork: Krieger (1979) (reprint, 1960); and Introduction to Crystal Geometry . New York: Krieger (1977) (reprint, 1971). Burns,G.,andA.M.Glazer, SpaceGroupsforSolid-StateScientists .NewYork:AcademicPress(1978).Awell- organized, readable treatment of groups and their application to the solid state. de-Shalit, A., and I. Talmi, Nuclear Shell Model . New York: Academic Press (1963). We adopt the Condon– Shortley phase conventions of this text. Edmonds, A.R., Angular Momentumin Quantum Mechanics . Princeton, NJ:Princeton University Press (1957). Edwards,H.M., AdvancedCalculus: ADifferential FormsApproach . Boston: Birkhäuser (1994). Falicov, L. M., Group Theory and Its Physical Applications . Notes compiled by A. Luehrmann. Chicago: Uni- versity of Chicago Press (1966). Group theory, with an emphasis on applications to crystal symmetries and solid-state physics. 320 Chapter 4 Group Theory Gell-Mann, M., and Y. Ne’eman, The Eightfold Way . New York: Benjamin (1965). A collection of reprints of significant papers on SU(3) and the particles of high-energy physics. Several introductory sections by Gell- MannandNe’emanare especiallyhelpful. Greiner, W., and B. Müller, Quantum Mechanics Symmetries . Berlin: Springer (1989). We refer to this textbook for more details and numerous exercises that areworked out in detail. Hamermesh,M., GroupTheoryandItsApplicationtoPhysicalProblems .Reading,MA:Addison-Wesley(1962). A detailed, rigorous account of both finite and continuous groups. The 32 point groups are developed. The continuous groups are treated, with Lie algebra included. A wealth of applications to atomic and nuclear physics. Hassani,S., Foundations of Mathematical Physics .Boston: Allyn and Bacon (1991). Heitler,W., TheQuantumTheoryofRadiation ,2nded.Oxford:OxfordUniversityPress(1947).Reprinted,New York: Dover (1983). Higman, B., Applied Group-Theoretic and Matrix Methods . Oxford: Clarendon Press (1955). A rather complete and unusually intelligible development of matrix analysis and group theory. Jackson, J. D., Classical Electrodynamics ,3rd ed.NewYork: Wiley (1998). Messiah,A., Quantum Mechanics , Vol. II. Amsterdam: North-Holland (1961). Panofsky,W.K.H.,andM.Phillips, ClassicalElectricityandMagnetism ,2nded.Reading,MA:Addison-Wesley (1962). The Lorentz covariance of Maxwell’s equations is developed for both vacuum and material media. Panofsky and Phillips use contravariant and covariant tensors. Park,D.,ResourceletterSP-1onsymmetryinphysics. Am.J.Phys. 36:577–584(1968).Includesalargeselection of basic references on group theory and its applications to physics: atoms, molecules, nuclei, solids, and elementaryparticles. Ram, B., Physics of the SU(3) symmetry model. Am. J. Phys. 35: 16 (1967). An excellent discussion of the applications of SU(3) to the strongly interacting particles (baryons). For a sequel to this see R. D. Young, Physics of the quark model. Am.J .Ph ys. 41: 472 (1973). R o s e ,M .E . , Elementary Theory of Angular Momentum . New York: Wiley (1957). Reprinted. New York: Dover (1995).Aspartofthedevelopmentofthequantumtheoryofangularmomentum,Roseincludesadetailedand readable accountof the rotation group. Wigner, E. P., Group Theory and Its Application to the Quantum Mechanics of Atomic Spectra (translated by J.J.Griffin).NewYork:AcademicPress(1959).Thisistheclassicreferenceongrouptheoryforthephysicist. Therotation group is treatedin considerable detail.Thereis awealthof applications to atomic physics. CHAPTER 5 INFINITE SERIES 5.1 F UNDAMENTAL CONCEPTS Infiniteseries,literallysummationsofaninfinitenumberofterms,occurfrequentlyinboth pureandappliedmathematics.Theymaybeusedbythepuremathematiciantodefinefunc- tions as a fundamental approach to the theory of functions, as well as for calculating ac- curatevaluesoftranscendentalconstantsandtranscendentalfunctions.Inthemathematics of science and engineering infinite series are ubiquitous, for they appear in the evaluation of integrals (Sections 5.6 and 5.7), in the solution of differential equations (Sections 9.5 and 9.6), and as Fourier series (Chapter 14) and compete with integral representations for thedescriptionofahostofspecialfunctions(Chapters11,12,and13).InSection16.3the Neumann series solution for integral equations provides one more example of the occur- renceanduseof infiniteseries. Right at the start we face the problem of attaching meaning to the sum of an infinite numberofterms.Theusualapproachisbypartialsums.Ifwehaveaninfinitesequenceof termsu1,u2,u3,u4,u5,...,wedefinethe ithpartialsumas si=isummationdisplay n=1un. (5.1) This is a finite summation and offers no difficulties. If the partial sums siconverge to a (finite)limitas i→∞, lim i→∞si=S, (5.2) the infinite seriessummationtext∞ n=1unis said to be convergent and to have the value S. Note that we reasonably, plausibly, but still arbitrarily definethe infinite series as equal to Sand that a necessary condition for this convergence to a limit is that lim n→∞un=0. This condition, however, is not sufficient to guarantee convergence. Equation (5.2) is usually written in formalmathematicalnotation: 321 322 Chapter 5 Infinite Series The condition for the existence of a limit Sis that for each ε>0, there is a fixed N=N(ε)suchthat |S−si|<ε, fori>N. This condition is often derived from the Cauchy criterion applied to the partial sums si. TheCauchycriterion is: Anecessaryandsufficientconditionthatasequence (si)convergeisthatforeach ε>0 there is a fixednumber Nsuch that |sj−si|<ε, for alli,j >N. This meansthat theindividual partial sumsmustcluster together aswemovefarout in the sequence. The Cauchy criterion may easily be extended to sequences of functions. We see it in this form in Section 5.5 in the definition of uniform convergence and in Section 10.4 in the development of Hilbert space. Our partial sums simay not converge to a single limit but mayoscillate,asinthecase ∞summationdisplay n=1un=1−1+1−1+1+···−(−1)n+···. Clearly,si=1f o riodd butsi=0f o rieven. There is no convergence to a limit, and series such as this one are labeled oscillatory . Whenever the sequence of partial sums diverges (approaches ±∞), the infinite series is said to diverge. Often the term divergent is extended to include oscillatory series as well. Because we evaluate the partial sums by ordinary arithmetic, the convergent series, defined in terms of a limit of the partial sums, assumes a position of supreme importance. Two examples may clarify the nature of convergence or divergence of a series and will also serve as a basis for a further detailed investigationinthenextsection. Example 5.1.1 THEGEOMETRIC SERIES Thegeometricalsequence,startingwith aandwitharatio r(=an+1/anindependentof n), isgivenby a+ar+ar2+ar3+···+arn−1+···. Thenthpartialsumisgivenby1 sn=a1−rn 1−r. (5.3) Takingthelimitas n→∞, limn→∞sn=a 1−r,for|r|<1. (5.4) 1Multiply and divide sn=summationtextn−1 m=0armby 1−r. 5.1 Fundamental Concepts 323 Hence,bydefinition,theinfinitegeometricseries convergesfor |r|<1 andis givenby ∞summationdisplay n=1arn−1=a 1−r. (5.5) Ontheotherhand,if |r|≥1,thenecessarycondition un→0isnotsatisfiedandtheinfinite seriesdiverges. /squaresolid Example 5.1.2 THEHARMONIC SERIES Asasecondandmoreinvolvedexample,weconsidertheharmonicseries ∞summationdisplay n=11 n=1+1 2+1 3+1 4+···+1 n+···. (5.6) We have the lim n→∞un=limn→∞1/n=0, but this is not sufficient to guaranteeconver- gence.If wegrouptheterms(no changeinorder) as 1+1 2+parenleftbig1 3+1 4parenrightbig +parenleftbig1 5+1 6+1 7+1 8parenrightbig +parenleftbig1 9+···+1 16parenrightbig +···, (5.7) eachpairofparenthesesencloses ptermsof theform 1 p+1+1 p+2+···+1 p+p>p 2p=1 2. (5.8) Formingpartialsumsbyaddingtheparentheticalgroupsonebyone,weobtain s1=1,s 4>5 2, s2=3 2,s5>6 2,··· s3>4 2,sn>n+1 2.(5.9) The harmonic series considered in this way is certainly divergent.2An alternate and inde- pendentdemonstrationofits divergenceappearsinSection5.2. /squaresolid If theun>0are monotonically decreasing to zero, that is, un>un+1for alln,thensummationtext nunis converging to Sif, and only if, sn−nunconverges to S.As the partial sums sn convergeto S,thistheoremimpliesthat nun→0,forn→∞. Toprovethis theorem,westartbyconcludingfrom 0 <un+1<unand sn+1−(n+1)un+1=sn−nun+1=sn−nun+n(un−un+1)>sn−nun thatsn−nunincreases as n→∞.As a consequence of sn−nun<sn≤S, sn−nun convergestoavalue s≤S.Deletingthetailofpositiveterms ui−unfromi=ν+1t on, 2The(finite) harmonic seriesappearsin aninteresting noteon themaximum stabledisplacementofastackof coins.P.R.John- son, The LeaningTowerof Lire. Am.J.Phys. 23: 240 (1955). 324 Chapter 5 Infinite Series weinferfrom sn−nun>u0+(u1−un)+···+(uν−un)=sν−νunthatsn−nun≥sν forn→∞.Hencealso s≥S,sos=Sandnun→0. When this theorem is applied to the harmonic seriessummationtext n1 nwithn1 n=1 it implies that itdoesnotconverge;itdivergesto +∞. Addition, Subtraction of Series If we have two convergent seriessummationtext nun→sandsummationtext nvn→S,their sum and difference willalsoconvergeto s±Sbecausetheirpartialsumssatisfy vextendsinglevextendsinglesj±Sj−(si±Si)vextendsinglevextendsingle=vextendsinglevextendsinglesj−si±(Sj−Si)vextendsinglevextendsingle≤|sj−si|+|Sj−Si|<2ǫ usingthetriangleinequality |a|−|b|≤|a+b|≤|a|+|b| fora=sj−si,b=Sj−Si. A convergent seriessummationtext nun→Smay be multiplied termwise by a real number a.The newseries willconvergeto aSbecause |asj−asi|=vextendsinglevextendsinglea(sj−si)vextendsinglevextendsingle=|a||sj−si|<|a|ǫ. This multiplication by a constant can be generalized to a multiplication by terms cnof a boundedsequenceof numbers. Ifsummationtext nunconverges to Sand0<cn≤Mare bounded, thensummationtext nuncnis convergent. Ifsummationtext nunisdivergentand cn>M>0,thensummationtext nuncndiverges. Toprovethis theorem wetakei,jsufficientlylargesothat |sj−si|<ǫ.Then jsummationdisplay i+1uncn≤Mjsummationdisplay i+1un=M|sj−si|<Mǫ. Thedivergentcasefollowsfrom summationdisplay nuncn>Msummationdisplay nun→∞. Usingthebinomialtheorem3(Section5.6), wemayexpandthefunction (1+x)−1: 1 1+x=1−x+x2−x3+···+(−x)n−1+···. (5.10) If weletx→1,thisseriesbecomes 1−1+1−1+1−1+···, (5.11) a series that we labeled oscillatory earlier in this section. Although it does not converge in the usual sense, meaning can be attached to this series. Euler, for example, assigned a value of 1 /2 to this oscillatory sequence on the basis of the correspondence between this series and the well-defined function (1+x)−1. Unfortunately, such correspondence be- tweenseries andfunctionis notunique,andthisapproachmustberefined.Othermethods 3ActuallyEq. (5.10) may be verified by multiplying both sides by 1 +x. 5.2 Convergence Tests 325 of assigning a meaning to a divergent or oscillatory series, methods of defining a sum, have been developed. See G. H. Hardy, Divergent Series , Chelsea Publishing Co. 2nd ed. (1992).Ingeneral,however,thisaspectofinfiniteseriesisofrelativelylittleinteresttothe scientist or the engineer. An exception to this statement, the very important asymptotic or semiconvergentseries, is consideredinSection 5.10. Exercises 5.1.1 Showthat ∞summationdisplay n=11 (2n−1)(2n+1)=1 2. Hint.Show(bymathematicalinduction)that sm=m/(2m+1). 5.1.2 Showthat ∞summationdisplay n=11 n(n+1)=1. Findthepartialsum smandverifyitscorrectnessbymathematicalinduction. Note. The method of expansion in partial fractions, Section 15.8, offers an alternative wayofsolvingExercises5.1.1and5.1.2. 5.2 C ONVERGENCE TESTS Although nonconvergent series may be useful in certain special cases (compare Sec- tion 5.10), we usually insist, as a matter of convenience if not necessity, that our series be convergent.Itthereforebecomesamatterofextremeimportancetobeabletotellwhether a given series is convergent. We shall develop a number of possible tests, starting with the simple and relatively insensitive tests and working up to the more complicated but quite sensitivetests.Forthepresentletusconsidera seriesofpositiveterms an≥0,postponing negativetermsuntilthenextsection. Comparison Test If term by term a series of terms 0 ≤un≤an, in which the anform a convergent series, the seriessummationtext nunis also convergent. If un≤anfor alln, thensummationtext nun≤summationtext nanandsummationtext nun therefore is convergent . If term byterm a series of terms vn≥bn, in whichthe bn,f o r ma divergent series, the seriessummationtext nvnis alsodivergent . Note that comparisons of unwithbn orvnwithanyield no information. If vn≥bnfor alln, thensummationtext nvn≥summationtext nbnandsummationtext nvn thereforeisdivergent. Fortheconvergentseries anwealreadyhavethegeometricseries,whereastheharmonic series will serve as the divergent comparison series bn. As other series are identified as either convergent or divergent, they may be used for the known series in this comparison test.Alltestsdevelopedinthissectionareessentiallycomparisontests.Figure5.1exhibits thesetests andtheinterrelationships. 326 Chapter 5 Infinite Series FIGURE 5.1Comparisontests. Example 5.2.1 AD IRICHLET SERIES Testsummationtext∞ n=1n−p,p=0.999,forconvergence.Since n−0.999>n−1andbn=n−1formsthe divergentharmonicseries, thecomparisontestshowsthatsummationtext nn−0.999is divergent.Gener- alizing,summationtext nn−pis seen to be divergent for all p≤1 but convergent for p>1 (see Exam- ple5.2.3). /squaresolid Cauchy Root Test If(an)1/n≤r<1 for all sufficiently large n, withrindependent of n, thensummationtext nanis convergent.If (an)1/n≥1 for allsufficientlylarge n,thensummationtext nanisdivergent. The first part of this test is verified easily by raising (an)1/n≤rto thenth power. We get an≤rn<1. Sincernis just the nth term in a convergent geometric series,summationtext nanis convergent by the comparison test. Conversely, if (an)1/n≥1, thenan≥1 and the series must diverge. This roottestisparticularlyusefulinestablishingthepropertiesofpowerseries(Section5.7). D’Alembert (or Cauchy) Ratio Test Ifan+1/an≤r<1 for all sufficiently large nandris independent of n, thensummationtext nanis convergent.If an+1/an≥1 for allsufficientlylarge n,thensummationtext nanisdivergent. Convergenceisprovedbydirectcomparisonwiththegeometricseries (1+r+r2+···). In the second part, an+1≥anand divergence should be reasonably obvious. Although not 5.2 Convergence Tests 327 quitesosensitiveastheCauchyroottest,thisD’Alembertratiotestisoneoftheeasiestto applyandiswidelyused.Analternatestatementoftheratiotestisintheformofalimit:If limn→∞an+1 an<1,convergence , >1,divergence , (5.12) =1,indeterminate . Becauseofthisfinalindeterminatepossibility,theratiotestislikelytofailatcrucialpoints, and more delicate, sensitive tests are necessary. The alert reader may wonder how this indeterminacy arose. Actually it was concealed in the first statement, an+1/an≤r<1. We might encounter an+1/an<1 for allfinitenbut be unable to choose an r<1and independentofn suchthat an+1/an≤rforallsufficientlylarge n.Anexampleisprovided bytheharmonicseries an+1 an=n n+1<1. (5.13) Since limn→∞an+1 an=1, (5.14) nofixedratio r<1 existsandtheratiotestfails. Example 5.2.2 D’A LEMBERT RATIO TEST Testsummationtext nn/2nfor convergence. an+1 an=(n+1)/2n+1 n/2n=1 2·n+1 n. (5.15) Since an+1 an≤3 4forn≥2, (5.16) wehaveconvergence.Alternatively, limn→∞an+1 an=1 2(5.17) andagain—convergence. /squaresolid Cauchy (or Maclaurin) Integral Test This is another sort of comparison test, in which we compare a series with an integral. Geometrically,wecomparetheareaofaseriesofunit-widthrectangleswiththeareaunder acurve. 328 Chapter 5 Infinite Series FIGURE 5.2(a)Comparisonofintegralandsum-blocksleading. (b)Comparisonofintegralandsum-blockslagging. Letf(x)be a continuous, monotonic decreasing function in which f(n)=an. Thensummationtext nanconverges ifintegraltext∞ 1f(x)dxis finite and diverges if the integral is infinite. For the ith partialsum, si=isummationdisplay n=1an=isummationdisplay n=1f(n). (5.18) But si>integraldisplayi+1 1f(x)dx (5.19) from Fig.5.2a, f(x)beingmonotonicdecreasing.Ontheotherhand,fromFig. 5.2b, si−a1<integraldisplayi 1f(x)dx, (5.20) in which the series is represented by the inscribed rectangles. Taking the limit as i→∞, wehave integraldisplay∞ 1f(x)dx≤∞summationdisplay n=1an≤integraldisplay∞ 1f(x)dx+a1. (5.21) Hence the infinite series converges or diverges as the corresponding integral converges or diverges. This integral test is particularly useful in setting upper and lower bounds on the remainderofaseries aftersomenumberof initialtermshavebeensummed.Thatis, ∞summationdisplay n=1an=Nsummationdisplay n=1an+∞summationdisplay n=N+1an, where integraldisplay∞ N+1f(x)dx≤∞summationdisplay n=N+1an≤integraldisplay∞ N+1f(x)dx+aN+1. 5.2 Convergence Tests 329 Tofreetheintegraltestfromthequiterestrictiverequirementthattheinterpolatingfunc- tionf(x)be positive and monotonic, we show for any function f(x)with a continuous derivativethat Nfsummationdisplay n=Ni+1f(n)=integraldisplayNf Nif(x)dx+integraldisplayNf Niparenleftbig x−[x]parenrightbig f′(x)dx (5.22) holds.Here[x]denotesthelargestintegerbelow x,sox−[x]variessawtoothlikebetween 0and1. ToderiveEq. (5.22)weobservethat integraldisplayNf Nixf′(x)dx=Nff(Nf)−Nif(Ni)−integraldisplayNf Nif(x)dx, (5.23) usingintegrationbyparts.Nextweevaluatetheintegral integraldisplayNf Ni[x]f′(x)dx=Nf−1summationdisplay n=Ninintegraldisplayn+1 nf′(x)dx=Nf−1summationdisplay n=Ninbraceleftbig f(n+1)−f(n)bracerightbig =−Nfsummationdisplay n=Ni+1f(n)−Nif(Ni)+Nff(Nf). (5.24) Subtracting Eq. (5.24) from (5.23) we arrive at Eq. (5.22). Note that f(x)m a yg ou po r down and even change sign, so Eq. (5.22) applies to alternating series (see Section 5.3) as well. Usually f′(x)falls faster than f(x)forx→∞, so the remainder term in Eq. (5.22) convergesbetter.ItiseasytoimproveEq.(5.22)byreplacing x−[x]byx−[x]−1 2,which variesbetween −1 2and1 2: summationdisplay Ni<n≤Nff(n)=integraldisplayNf Nif(x)dx+integraldisplayNf Niparenleftbig x−[x]−1 2parenrightbig f′(x)dx +1 2braceleftbig f(Nf)−f(Ni)bracerightbig . (5.25) Thenthe f′(x)-integralbecomesevensmaller,if f′(x)doesnotchangesigntoooften.For anapplicationofthisintegraltesttoanalternatingseriesseeExample5.3.1. Example 5.2.3 RIEMANN ZETAFUNCTION TheRiemannzetafunctionis definedby ζ(p)=∞summationdisplay n=1n−p, (5.26) providedtheseries converges.Wemaytake f(x)=x−p, andthen integraldisplay∞ 1x−pdx=x−p+1 −p+1vextendsinglevextendsinglevextendsinglevextendsingle∞ 1,p/negationslash=1 =lnx|∞ x=1,p=1. (5.27) 330 Chapter 5 Infinite Series The integral and therefore the series are divergent for p≤1, convergent for p>1. Hence Eq. (5.26) should carry the condition p>1. This, incidentally, is an independent proof that the harmonic series (p=1)diverges logarithmically. The sum of the first million termssummationtext1,000,000n−1isonly14.392 726 .... /squaresolid ThisintegralcomparisonmayalsobeusedtosetanupperlimittotheEuler–Mascheroni constant,4definedby γ=limn→∞parenleftbiggnsummationdisplay m=1m−1−lnnparenrightbigg . (5.28) Returningtopartialsums, Eq. (5.20)yields sn=nsummationdisplay m=1m−1−lnn≤integraldisplayn 1dx x−lnn+1. (5.29) Evaluating the integral on the right, sn<1 for all nand therefore γ≤1. Exer- cise 5.2.12 leads to more restrictive bounds. Actually the Euler–Mascheroni constant is 0.57721566 .... Kummer’s Test This is the first of three tests that are somewhat more difficult to apply than the preceding tests. Their importance lies in their power and sensitivity. Frequently, at least one of the threewillworkwhenthesimpler,easiertestsareindecisive.Itmustberemembered,how- ever,thatthesetests,likethosepreviouslydiscussed,areultimatelybasedoncomparisons. It can be shown that there is no most slowly converging series and no most slowly diverg- ingseries.Thismeansthatallconvergencetestsgivenhere,includingKummer’s,mayfail sometime. We consider a series of positive terms uiand a sequence of finite positive constants ai. If anun un+1−an+1≥C>0 (5.30) foralln≥N, whereNissomefixednumber,5thensummationtext∞ i=1uiconverges .If anun un+1−an+1≤0 (5.31) andsummationtext∞ i=1a−1 idiverges,thensummationtext∞i=1uidiverges. 4This is the notation of National Bureau of Standards, Handbook of Mathematical Functions , Applied Mathematics Series-55 (AMS-55). NewYork: Dover (1972). 5Withumfinite, the partial sum sNwill always be finite for Nfinite. The convergence or divergence of a series depends on the behavior of thelast infinity of terms, not on the first Nterms. 5.2 Convergence Tests 331 The proof of this powerful test is remarkably simple. From Eq. (5.30), with Csome positiveconstant, CuN+1≤aNuN−aN+1uN+1 CuN+2≤aN+1uN+1−aN+2uN+2 ······························ Cun≤an−1un−1−anun.(5.32) Addinganddividingby C(andrecallingthat C/negationslash=0),weobtain nsummationdisplay i=N+1ui≤aNuN C−anun C. (5.33) Henceforthepartialsum sn, sn≤Nsummationdisplay i=1ui+aNuN C−anun C <Nsummationdisplay i=1ui+aNuN C,aconstant,independentof n. (5.34) Thepartialsumsthereforehaveanupperbound.Withzeroasanobviouslowerbound,the seriessummationtextuimustconverge. Divergenceis shownasfollows.FromEq. (5.31)for un+1>0, anun≥an−1un−1≥···≥aNuN,n>N. (5.35) Thus,for an>0, un≥aNuN an(5.36) and ∞summationdisplay i=N+1ui≥aNuN∞summationdisplay i=N+1a−1 i. (5.37) Ifsummationtext∞ i=1a−1 idiverges, then by the comparison testsummationtext iuidiverges. Equations (5.30) and (5.31)areoftengiveninalimitform: limn→∞parenleftbigg anun un+1−an+1parenrightbigg =C. (5.38) ThusforC>0 wehaveconvergence,whereasfor C<0 (andsummationtext ia−1 idivergent)wehave divergence.ItisperhapsusefultoshowthecloserelationofEq.(5.38)andEqs.(5.30)and (5.31)andtoshowwhyindeterminacycreepsinwhenthelimit C=0.Fromthedefinition oflimit, vextendsinglevextendsinglevextendsinglevextendsingleanun un+1−an+1−Cvextendsinglevextendsinglevextendsinglevextendsingle<ε (5.39) 332 Chapter 5 Infinite Series for alln≥Nand allε>0, no matter how small εmay be. When the absolute value signs areremoved, C−ε<anun un+1−an+1<C+ε. (5.40) Now, ifC>0, Eq. (5.30) follows from εsufficiently small. On the other hand, if C<0, Eq.(5.31)follows.However,if C=0,thecenterterm, an(un/un+1)−an+1,maybeeither positiveornegativeandtheprooffails.TheprimaryuseofKummer’stestistoproveother tests, suchas Raabe’s(comparealsoExercise5.2.3). If thepositiveconstants anof Kummer’stestarechosen an=n,wehaveRaabe’stest. Raabe’s Test Ifun>0 andif nparenleftbiggun un+1−1parenrightbigg ≥P>1 (5.41) foralln≥N,whereNisapositiveintegerindependentof n,thensummationtext iuiconverges .Here, P=C+1 ofKummer’stest. If nparenleftbiggun un+1−1parenrightbigg ≤1, (5.42) thensummationtext iuidiverges(assummationtext nn−1diverges). Thelimitformof Raabe’stest is limn→∞nparenleftbiggun un+1−1parenrightbigg =P. (5.43) We have convergence for P>1, divergence for P<1, and no conclusion for P=1, exactlyaswiththeKummertest.ThisindeterminacyispointedupbyExercise5.2.4,which presents a convergent series and a divergent series, with both series yielding P=1i n Eq.(5.43). Raabe’s test is more sensitive than the d’Alembert ratio test (Exercise 5.2.3) becausesummationtext∞ n=1n−1diverges more slowly thansummationtext∞n=11. We obtain a more sensitive test (and one thatis stillfairly easytoapply)bychoosing an=nlnn.This isGauss’ test. Gauss’ Test Ifun>0 for allfinite nand un un+1=1+h n+B(n) n2, (5.44) inwhich B(n)isaboundedfunctionof nforn→∞,thensummationtext iuiconvergesfor h>1 and divergesfor h≤1:Thereis noindeterminatecasehere. 5.2 Convergence Tests 333 The Gauss test is an extremely sensitive test of series convergence. It will work for all series the physicist is likely to encounter. For h>1o rh<1 the proof follows directly fromRaabe’stest limn→∞nbracketleftbigg 1+h n+B(n) n2−1bracketrightbigg =limn→∞bracketleftbigg h+B(n) nbracketrightbigg =h. (5.45) Ifh=1, Raabe’s test fails. However, if we return to Kummer’s test and use an=nlnn, Eq.(5.38) leadsto limn→∞braceleftbigg nlnnbracketleftbigg 1+1 n+B(n) n2bracketrightbigg −(n+1)ln(n+1)bracerightbigg =limn→∞bracketleftbigg nlnn·n+1 n−(n+1)ln(n+1)bracketrightbigg =limn→∞(n+1)bracketleftbigg lnn−lnn−lnparenleftbigg 1+1 nparenrightbiggbracketrightbigg . (5.46) Borrowinga resultfromSection5.6(whichis notdependentonGauss’test), wehave limn→∞−(n+1)lnparenleftbigg 1+1 nparenrightbigg =limn→∞−(n+1)parenleftbigg1 n−1 2n2+1 3n3···parenrightbigg =−1<0. (5.47) Hence we have divergence for h=1. This is an example of a successful application of Kummer’stestwhenRaabe’stesthadfailed. Example 5.2.4 LEGENDRE SERIES TherecurrencerelationfortheseriessolutionofLegendre’sequation(Exercise9.5.5)may beputintheform a2j+2 a2j=2j(2j+1)−l(l+1) (2j+1)(2j+2). (5.48) Foruj=a2jandB(j)=O(1/j2)→0 (that is,|B(j)j2|≤C,C>0, a constant) as j→∞inGauss’ testweapplyEq. (5.45).Then, for j≫l,6 uj uj+1→(2j+1)(2j+2) 2j(2j+1)=2j+2 2j=1+1 j. (5.49) ByEq. (5.44)theseries isdivergent. /squaresolid 6Theldependence enters B(j)but does not affect hin Eq. (5.45). 334 Chapter 5 Infinite Series Improvement of Convergence This section so far has been concerned with establishing convergence as an abstract math- ematicalproperty.Inpractice,the rateofconvergencemaybeofconsiderableimportance. Here we present one method of improving the rate of convergence of a convergent series. OthertechniquesaregiveninSections5.4and5.9. The basic principle of this method, due to Kummer, is to form a linear combination of our slowly converging series and one or more series whose sum is known. For the known seriesthecollection α1=∞summationdisplay n=11 n(n+1)=1 α2=∞summationdisplay n=11 n(n+1)(n+2)=1 4 α3=∞summationdisplay n=11 n(n+1)(n+2)(n+3)=1 18 ......... αp=∞summationdisplay n=11 n(n+1)···(n+p)=1 p·p! is particularly useful.7The series are combined term by term and the coefficients in the linearcombinationchosentocancelthemostslowlyconvergingterms. Example 5.2.5 RIEMANN ZETAFUNCTION ,ζ(3) Letthe series tobe summedbesummationtext∞ n=1n−3. In Section5.9 thisis identifiedas theRiemann zetafunction, ζ(3). Weforma linearcombination ∞summationdisplay n=1n−3+a2α2=∞summationdisplay n=1n−3+a2 4. α1is not included since it converges more slowly than ζ(3). Combining terms, we obtain ontheleft-handside ∞summationdisplay n=1braceleftbigg1 n3+a2 n(n+1)(n+2)bracerightbigg =∞summationdisplay n=1n2(1+a2)+3n+2 n3(n+1)(n+2). If wechoose a2=−1,theprecedingequationsyield ζ(3)=∞summationdisplay n=1n−3=1 4+∞summationdisplay n=13n+2 n3(n+1)(n+2). (5.50) 7These series sums may be verified by expanding the forms by partial fractions, writing out the initial terms, and inspecting the pattern of cancellationof positive andnegative terms. 5.2 Convergence Tests 335 The resulting series may not be beautiful but it does converge as n−4, faster than n−3. A more convenient form comes from Exercise 5.2.21. There, the symmetry leads to con- vergenceas n−5. /squaresolid The method can be extended, including a3α3to get convergence as n−5,a4α4to get convergenceas n−6, and so on. Eventually,you have to reach a compromise between how muchalgebrayoudoandhowmucharithmeticthecomputerdoes.Ascomputersgetfaster, thebalanceis steadilyshiftingtolessalgebrafor youandmorearithmeticforthem. Exercises 5.2.1 (a) Provethatif limn→∞npun=A<∞,p>1, theseriessummationtext∞ n=1unconverges. (b) Provethatif limn→∞nun=A>0, theseries diverges.(Thetestfails for A=0.) These two tests, known as limit tests , are often convenient for establishing the conver- genceofaseries. Theymaybetreatedascomparisontests, comparingwith summationdisplay nn−q,1≤q<p . 5.2.2 If limn→∞bn an=K, aconstantwith 0 <K<∞, showthatsummationtext nbnconvergesor divergeswithsummationtextan. Hint.Ifsummationtextanconverges,use b′ n=1 2Kbn.I fsummationtext nandiverges,use b′′ n=2 Kbn. 5.2.3 Showthatthecompleted’AlembertratiotestfollowsdirectlyfromKummer’stestwith ai=1. 5.2.4 ShowthatRaabe’stestisindecisivefor P=1byestablishingthat P=1fortheseries (a)un=1 nlnnandthatthisseries diverges. (b)un=1 n(lnn)2andthatthisseries converges. Note. By direct additionsummationtext100,000 2[n(lnn)2]−1=2.02288. The remainder of the series n>105yields 0.08686 by the integral comparison test. The total, then, 2 to ∞,i s 2.1097. 336 Chapter 5 Infinite Series 5.2.5 Gauss’testis oftengivenintheform ofatestof theratio un un+1=n2+a1n+a0 n2+b1n+b0. Forwhatvaluesoftheparameters a1andb1is thereconvergence?divergence? ANS.Convergentfor a1−b1>1, divergentfor a1−b1≤1. 5.2.6 Test forconvergence (a)∞summationdisplay n=2(lnn)−1(d)∞summationdisplay n=1bracketleftbig n(n+1)bracketrightbig−1/2 (b)∞summationdisplay n=1n! 10n(e)∞summationdisplay n=01 2n+1. (c)∞summationdisplay n=11 2n(2n+1) 5.2.7 Test forconvergence (a)∞summationdisplay n=11 n(n+1)(d)∞summationdisplay n=1lnparenleftbigg 1+1 nparenrightbigg (b)∞summationdisplay n=21 nlnn(e)∞summationdisplay n=11 n·n1/n. (c)∞summationdisplay n=11 n2n 5.2.8 Forwhatvaluesof pandqwillthefollowingseries converge?summationtext∞ n=21 np(lnn)q. ANS.ConvergentforbraceleftBigg p>1,allq, p=1,q >1,divergentforbraceleftBigg p<1,allq, p=1,q≤1. 5.2.9 Determinetherangeofconvergencefor Gauss’s hypergeometricseries F(α,β,γ;x)=1+αβ 1!γx+α(α+1)β(β+1) 2!γ(γ+1)x2+···. Hint. Gauss developed his test for the specific purpose of establishing the convergence ofthisseries. ANS.Convergentfor −1<x<1 andx=±1i fγ>α+β. 5.2.10 Apocketcalculatoryields 100summationdisplay n=1n−3=1.202007. 5.2 Convergence Tests 337 Showthat 1.202056≤∞summationdisplay n=1n−3≤1.202057. Hint.Useintegralstosetupperandlowerboundsonsummationtext∞ n=101n−3. Note. A more exact value for summation of ζ(3)=summationtext∞ n=1n−3is 1.202 056 903 ...; ζ(3)isknowntobeanirrationalnumber,butitisnotlinkedtoknownconstantssuchas e,π,γ,ln2. 5.2.11 Setupperandlowerboundsonsummationtext1,000,000 n=1n−1, assumingthat (a) theEuler–Mascheroniconstantis known. ANS. 14.392726<1,000,000summationdisplay n=1n−1<14.392727. (b) TheEuler–Mascheroniconstantisunknown. 5.2.12 Givensummationtext1,000 n=1n−1=7.485470...setupperandlowerboundsontheEuler–Mascheroni constant. ANS. 0.5767<γ<0.5778. 5.2.13 (FromOlbers’ paradox .) Assume a static universe in which the stars are uniformly distributed. Divide all space into shells of constant thickness; the stars in any one shell by themselves subtend a solid angle of ω0.Allowing for the blocking out of distant stars by nearer stars , show that the total net solid angle subtended by all stars, shells extending to infinity, is exactly4π. [Therefore the night sky should be ablaze with light. For more details, see E. Harrison, Darkness at Night: A Riddle of the Universe . Cambridge,MA:HarvardUniversityPress (1987).] 5.2.14 Test forconvergence ∞summationdisplay n=1bracketleftbigg1·3·5···(2n−1) 2·4·6···(2n)bracketrightbigg2 =1 4+9 64+25 256+···. 5.2.15 TheLegendreseriessummationtext jevenuj(x)satisfiestherecurrencerelations uj+2(x)=(j+1)(j+2)−l(l+1) (j+2)(j+3)x2uj(x), in which the index jis even and lis some constant (but, in this problem, nota non- negative odd integer). Find the range of values of xfor which this Legendre series is convergent.Test theendpoints. ANS.−1<x<1. 338 Chapter 5 Infinite Series 5.2.16 A series solution (Section 9.5) of the Chebyshev equation leads to successive terms havingtheratio uj+2(x) uj(x)=(k+j)2−n2 (k+j+1)(k+j+2)x2, withk=0 andk=1.Test for convergenceat x=±1. ANS.Convergent. 5.2.17 Aseriessolutionfortheultraspherical(Gegenbauer)function Cα n(x)leadstotherecur- rence aj+2=aj(k+j)(k+j+2α)−n(n+2α) (k+j+1)(k+j+2). Investigate the convergence of each of these series at x=±1 as a function of the para- meterα. ANS.Convergentfor α<1, divergentfor α≥1. 5.2.18 Aseriesexpansionoftheincompletebetafunction(Section8.4)yields Bx(p,q)=xpbraceleftbigg1 p+1−q p+1x+(1−q)(2−q) 2!(p+2)x2+··· +(1−q)(2−q)···(n−q) n!(p+n)xn+···bracerightbigg . Given that 0≤x≤1,p>0, andq>0, test this series for convergence. What happens atx=1? 5.2.19 Showthatthefollowingseriesis convergent. ∞summationdisplay s=0(2s−1)!! (2s)!!(2s+1). Note.(2s−1)!!=(2s−1)(2s−3)···3·1with(−1)!!=1;(2s)!!=(2s)(2s−2)···4·2 with 0!!=1. The series appears as a series expansion of sin−1(1)and equals π/2, and sin−1x≡arcsinx/negationslash=(sinx)−1. 5.2.20 Show how to combine ζ(2)=summationtext∞ n=1n−2withα1andα2to obtain a series converging asn−4. Note.ζ(2)isknown: ζ(2)=π2/6 (see Section5.9). 5.2.21 The convergence improvement of Example 5.2.5 may be carried out more expediently (in this special case) by putting α2into a more symmetric form: Replacing nbyn−1, wehave α′ 2=∞summationdisplay n=21 (n−1)n(n+1)=1 4. 5.3 Alternating Series 339 (a) Combine ζ(3)andα′ 2toobtainconvergenceas n−5. (b) Let α′ 4beα4withn→n−2.Combine ζ(3),α′ 2, andα′ 4toobtainconvergenceas n−7. (c) Ifζ(3)istobecalculatedtosix =decimal=placeaccuracy(error5 ×10−7),how many terms are required for ζ(3)alone? combined as in part (a)? combined as in part(b)? Note.Theerror maybeestimatedusingthecorrespondingintegral. ANS. (a) ζ(3)=5 4−∞summationdisplay n=21 n3(n2−1). 5.2.22 Catalan’sconstant (β(2)ofM.AbramowitzandI.A.Stegun,HandbookofMathemati- calFunctionswithFormulas,Graphs,andMathematicalTables(AMS-55),Wash,D.C. National Bureau of Standards (1972); reprinted Dover (1974), Chapter 23) is defined by β(2)=∞summationdisplay k=0(−1)k(2k+1)−2=1 12−1 32+1 52···. Calculate β(2)tosix-digitaccuracy. Hint.Therateofconvergenceisenhancedbypairingtheterms: (4k−1)−2−(4k+1)−2=16k (16k2−1)2. Ifyouhavecarriedenoughdigitsinyourseriessummation,summationtext 1≤k≤N16k/(16k2−1)2, additionalsignificantfiguresmaybeobtainedbysettingupperandlowerboundsonthe tail of the series,summationtext∞ k=N+1. These bounds may be set by comparison with integrals, as intheMaclaurinintegraltest. ANS.β(2)=0.915965594177 .... 5.3 A LTERNATING SERIES In Section 5.2 we limited ourselves to series of positive terms. Now, in contrast, we con- sider infinite series in which the signs alternate. The partial cancellationdue to alternating signsmakesconvergencemorerapidandmucheasiertoidentify.WeshallprovetheLeib- niz criterion, a general condition for the convergence of an alternating series. For series with more irregular sign changes, like Fourier series of Chapter 14 (see Example 5.3.1), theintegraltestof Eq. (5.25)is oftenhelpful. Leibniz Criterion Considertheseriessummationtext∞ n=1(−1)n+1anwithan>0.Ifan,ismonotonicallydecreasing (for sufficientlylarge n) and lim n→∞an=0, thentheseries converges.To provethis theorem, 340 Chapter 5 Infinite Series weexaminetheevenpartialsums s2n=a1−a2+a3−···−a2n, s2n+2=s2n+(a2n+1−a2n+2).(5.51) Sincea2n+1>a2n+2,weha v e s2n+2>s2n. (5.52) Ontheotherhand, s2n+2=a1−(a2−a3)−(a4−a5)−···−a2n+2. (5.53) Hence,witheachpairofterms a2p−a2p+1>0, s2n+2<a1. (5.54) With the even partial sums bounded s2n<s2n+2<a1and the terms andecreasing monotonicallyandapproachingzero,thisalternatingseriesconverges. Onefurtherimportantresultcanbeextractedfromthepartialsumsofthesamealternat- ingseries. Fromthedifferencebetweentheseries limit Sandthepartialsum sn, S−sn=an+1−an+2+an+3−an+4+··· =an+1−(an+2−an+3)−(an+4−an+5)−···, (5.55) or S−sn<an+1. (5.56) Equation (5.56) says that the error in cutting off an alternating series whose terms are monotonicallydecreasing after nterms is less than an+1, the first term dropped. A knowl- edgeoftheerror obtainedthis waymaybeofgreatpracticalimportance. Absolute Convergence Givenaseriesofterms uninwhichunmayvaryinsign,ifsummationtext|un|converges,thensummationtextunis said to be absolutely convergent. Ifsummationtextunconverges butsummationtext|un|diverges, the convergence iscalledconditional . Thealternatingharmonicseriesisasimpleexampleofthisconditionalconvergence.We have ∞summationdisplay n=1(−1)n−1n−1=1−1 2+1 3−1 4+···+(−1)n−1 n+···, (5.57) convergentbytheLeibnizcriterion;but ∞summationdisplay n=1n−1=1+1 2+1 3+1 4+···+1 n+··· hasbeenshowntobedivergentinSections5.1 and5.2. 5.3 Alternating Series 341 NotethatmosttestsdevelopedinSection5.2assumeaseriesofpositiveterms.Therefore thesetests inthatsectionguaranteeabsoluteconvergence. Example 5.3.1 SERIES WITH IRREGULAR SIGNCHANGES For 0<x<2πtheFourierseries(see Chapter14.1) ∞summationdisplay n=1cos(nx) n=−lnparenleftbigg 2sinx 2parenrightbigg (5.58) converges, having coefficients that change sign often, but not so that the Leibniz conver- gencecriterionapplieseasily.LetusapplytheintegraltestofEq.(5.22).Usingintegration bypartsweseeimmediatelythat integraldisplay∞ 1cos(nx) ndn=bracketleftbiggsin(nx) nxbracketrightbigg∞ 1+1 xintegraldisplay∞ n=1sin(nx) n2dn converges,andtheintegralontheright-handsideevenconvergesabsolutely.Thederivative terminEq.(5.22) hastheform integraldisplay∞ 1parenleftbig n−[n]parenrightbigbraceleftbigg −x nsin(nx)−cos(nx) n2bracerightbigg dn, where the second term converges absolutely and need not be considered further. Next we observethat g(N)=integraltextN 1(n−[n])sin(nx)dnisboundedfor N→∞,justasintegraltextNsin(nx)dn is bounded because of the periodic nature of sin (nx)and its regular sign changes. Using integrationbypartsagain, integraldisplay∞ 1g′(n) ndn=bracketleftbiggg(n) nbracketrightbigg∞ n=1+integraldisplay∞ 1g(n) n2dn, we see that the second term is absolutely convergent and that the first goes to zero at the upper limit. Hence the series in Eq. (5.58) converges, which is hard to see from other convergencetests. Alternatively, we may apply the q=1 case of the Euler–Maclaurin integration formula inEq.(5.168b), nsummationdisplay ν=1f(ν)=integraldisplayn 1f(x)dx+1 2braceleftbig f(n)+f(1)bracerightbig +1 12braceleftbig f′(n)−f′(1)bracerightbig −1 2integraldisplay1 0parenleftbigg x2−x+1 6parenrightbiggn−1summationdisplay ν=1f′′(x+ν)dx, whichisstraightforwardbutmoretediousbecauseof thesecondderivative. /squaresolid 342 Chapter 5 Infinite Series Exercises 5.3.1 (a) From the electrostatic two-hemisphere problem (Exercise 12.3.20) we obtain the series ∞summationdisplay s=0(−1)s(4s+3)(2s−1)!! (2s+2)!!. Testit forconvergence. (b) Thecorrespondingseriesfor thesurfacechargedensityis ∞summationdisplay s=0(−1)s(4s+3)(2s−1)!! (2s)!!. Testit forconvergence. The!!notationisexplainedinSection8.1andExercise5.2.19. 5.3.2 Showbydirectnumericalcomputationthatthesumof thefirst10termsof lim x→1ln(1+x)=ln2=∞summationdisplay n=1(−1)n−1n−1 differs from ln2 bylessthantheeleventhterm: ln2 =0.6931471806 .... 5.3.3 In Exercise 5.2.9 the hypergeometric series is shown convergent for x=±1,ifγ> α+β. Show that there is conditional convergence for x=−1f o rγdown to γ> α+β−1. Hint. The asymptotic behavior of the factorial function is given by Stirling’s series, Section8.3. 5.4 A LGEBRA OF SERIES The establishment of absolute convergence is important because it can be proved that ab- solutely convergent series may be reordered according to the familiar rules of algebra or arithmetic. •If aninfiniteseries isabsolutelyconvergent,theseries sumis independentof theorder inwhichthetermsareadded. •Theseriesmaybemultipliedwithanotherabsolutelyconvergentseries.Thelimitofthe productwillbetheproductoftheindividualserieslimits.Theproductseries,adouble series, willalso convergeabsolutely. Nosuchguaranteescanbegivenforconditionallyconvergentseries.Againconsiderthe alternatingharmonicseries. If wewrite 1−1 2+1 3−1 4+···=1−parenleftbig1 2−1 3parenrightbig −parenleftbig1 4−1 5parenrightbig −···, (5.59) 5.4 Algebra of Series 343 itisclearthatthesum ∞summationdisplay n=1(−1)n−1n−1<1. (5.60) However, if we rearrange the terms slightly, we may make the alternating harmonic series convergeto3 2. WeregroupthetermsofEq. (5.59), taking parenleftbig 1+1 3+1 5parenrightbig −parenleftbig1 2parenrightbig +parenleftbig1 7+1 9+1 11+1 13+1 15parenrightbig −parenleftbig1 4parenrightbig +parenleftbig1 17+···+1 25parenrightbig −parenleftbig1 6parenrightbig +parenleftbig1 27+···+1 35parenrightbig −parenleftbig1 8parenrightbig +···.(5.61) Treating the terms grouped in parentheses as single terms for convenience, we obtain the partialsums s1=1.5333 s2=1.0333 s3=1.5218 s4=1.2718 s5=1.5143 s6=1.3476 s7=1.5103 s8=1.3853 s9=1.5078 s10=1.4078. From this tabulation of snand the plot of snversusnin Fig. 5.3, the convergence to 3 2is fairly clear. We have rearranged the terms, taking positive terms until the partial sum was equal to or greater than3 2and then adding in negative terms until the partial sum just fell below3 2and so on. As the series extends to infinity, all original terms will eventually appear,butthepartialsumsof thisrearrangedalternatingharmonicseries convergeto3 2. By a suitable rearrangement of terms, a conditionally convergent series may be made to converge to any desired value or even to diverge. This statement is sometimes given FIGURE 5.3Alternatingharmonicseries—terms rearrangedtogiveconvergenceto1.5. 344 Chapter 5 Infinite Series asRiemann’s theorem . Obviously, conditionally convergent series must be treated with caution. Absolutely convergent series can be multiplied without problems. This follows as a special case from the rearrangement of double series. However, conditionally convergent series cannot always be multiplied to yield convergent series, as the following example shows. Example 5.4.1 SQUARE OF A CONDITIONALLY CONVERGENT SERIES MAYDIVERGE Theseriessummationtext∞ n=1(−1)n−1 √nconverges,bytheLeibnizcriterion.Its square, bracketleftbiggsummationdisplay n(−1)n−1 √nbracketrightbigg2 =summationdisplay n(−1)nbracketleftbigg1√ 11√ n−1+1√ 21√ n−2+···+1√ n−11√ 1bracketrightbigg , hasthegeneralterminbracketsconsistingof n−1additiveterms,eachofwhichisgreater than1√ n−1√ n−1,so the product term in brackets is greater thann−1 n−1and does not go to zero.Hencethisproductoscillatesandthereforediverges. /squaresolid Hence for a product of two series to converge, we have to demand as a sufficient con- dition that at least one of them converge absolutely. To prove this product convergence theorem thatifsummationtext nunconvergesabsolutelyto U,summationtext nvnconvergesto V,then summationdisplay ncn,c n=nsummationdisplay m=0umvn−m converges to UV,it is sufficient to show that the difference terms Dn≡c0+c1+···+ c2n−UnVn→0f o rn→∞,whereUn,Vnarethepartialsumsofourseries.Asaresult, thepartialsumdifferences Dn=u0v0+(u0v1+u1v0)+···+(u0v2n+u1v2n−1+···+u2nv0) −(u0+u1+···+un)(v0+v1+···+vn) =u0(vn+1+···+v2n)+u1(vn+1+···+v2n−1)+···+un+1vn+1 +vn+1(v0+···+vn−1)+···+u2nv0, sofor allsufficientlylarge n, |Dn|<ǫparenleftbig |u0|+···+| un−1|parenrightbig +Mparenleftbig |un+1|+···+| u2n|parenrightbig <ǫ(a+M), because|vn+1+vn+2+···+vn+m|<ǫforsufficientlylarge nandallpositiveintegers m assummationtextvnconverges,andthepartialsums Vn<Bofsummationtext nvnareboundedby M,becausethe sumconverges.Finallywecallsummationtext n|un|=a,assummationtextunconvergesabsolutely. Two series can be multiplied, provided one of them converges absolutely. Addition and subtractionofseries is alsovalidtermwiseifoneseries convergesabsolutely. 5.4 Algebra of Series 345 Improvement of Convergence, Rational Approximations Theseries ln(1+x)=∞summationdisplay n=1(−1)n−1xn n,−1<x≤1, (5.61a) converges very slowly as xapproaches+1. Therateof convergence may be improved substantially by multiplying both sides of Eq. (5.61a) by a polynomial and adjusting the polynomial coefficients to cancel the more slowly converging portions of the series. Con- siderthesimplestpossibility:Multiply ln (1+x)by 1+a1x: (1+a1x)ln(1+x)=∞summationdisplay n=1(−1)n−1xn n+a1∞summationdisplay n=1(−1)n−1xn+1 n. Combiningthetwoseries ontheright,termbyterm,weobtain (1+a1x)ln(1+x)=x+∞summationdisplay n=2(−1)n−1parenleftbigg1 n−a1 n−1parenrightbigg xn =x+∞summationdisplay n=2(−1)n−1n(1−a1)−1 n(n−1)xn. Clearly, if we take a1=1, thenin the numerator disappears and our combined series convergesas n−2. Continuing this process, we find that (1+2x+x2)ln(1+x)vanishes as n−3and that (1+3x+3x2+x3)ln(1+x)vanishes as n−4. In effect we are shifting from a simple series expansion of Eq. (5.61a) to a rational fraction representation in which the function ln(1+x)isrepresentedbytheratioofa seriesandapolynomial: ln(1+x)=x+summationtext∞ n=2(−1)nxn/[n(n−1)] 1+x. Suchrationalapproximationsmaybebothcompactandaccurate. Rearrangement of Double Series Another aspect of the rearrangement of series appears in the treatment of double series (Fig.5.4): ∞summationdisplay m=0∞summationdisplay n=0an,m. Letussubstitute n=q≥0,m=p−q≥0(q≤p). 346 Chapter 5 Infinite Series FIGURE 5.4Double series—summationover n indicatedbyverticaldashed lines. Thisresults intheidentity ∞summationdisplay m=0∞summationdisplay n=0an,m=∞summationdisplay p=0psummationdisplay q=0aq,p−q. (5.62) Thesummationover pandqof Eq.(5.62) isillustratedinFig.5.5.Thesubstitution n=s≥0,m=r−2s≥0parenleftbigg s≤r 2parenrightbigg leadsto ∞summationdisplay m=0∞summationdisplay n=0an,m=∞summationdisplay r=0[r/2]summationdisplay s=0as,r−2s, (5.63) FIGURE 5.5Doubleseries —again,thefirst summation is representedbyvertical dashedlines,butthese verticallinescorrespondto diagonalsinFig.5.4. 5.4 Algebra of Series 347 FIGURE 5.6Doubleseries. The summationover scorrespondstoa summationalongthealmost-horizontal dashedlinesinFig.5.4. with[r/2]=r/2f o rreven and (r−1)/2f o rrodd. The summation over randsof Eq. (5.63) is shown in Fig. 5.6. Equations (5.62) and (5.63) are clearly rearrangements of the array of coefficients anm, rearrangements that are valid as long as we have absolute convergence. ThecombinationofEqs. (5.62) and(5.63), ∞summationdisplay p=0psummationdisplay q=0aq,p−q=∞summationdisplay r=0[r/2]summationdisplay s=0as,r−2s, (5.64) isusedinSection12.1inthedeterminationoftheseriesformoftheLegendrepolynomials. Exercises 5.4.1 Giventheseries(derivedinSection5.6) ln(1+x)=x−x2 2+x3 3−x4 4···,−1<x≤1, showthat lnparenleftbigg1+x 1−xparenrightbigg =2parenleftbigg x+x3 3+x5 5+···parenrightbigg ,−1<x<1. The original series, ln (1+x), appears in an analysis of binding energy in crystals. It is1 2the Madelung constant (2ln2)for a chain of atoms. The second series is useful in normalizing the Legendre polynomials (Section 12.3) and in developing a second solutionfor Legendre’sdifferentialequation(Section12.10). 5.4.2 Determine the values of the coefficients a1,a2, anda3that will make (1+a1x+a2x2+a3x3)ln(1+x)convergeas n−4. Findtheresultingseries. 5.4.3 Showthat (a)∞summationdisplay n=2bracketleftbig ζ(n)−1bracketrightbig =1, (b)∞summationdisplay n=2(−1)nbracketleftbig ζ(n)−1bracketrightbig =1 2, whereζ(n)istheRiemannzetafunction. 348 Chapter 5 Infinite Series 5.4.4 Writeaprogramthatwillrearrangethetermsofthealternatingharmonicseriestomake theseriesconvergeto1.5.GroupyourtermsasindicatedinEq.(5.61).Listthefirst100 successivepartialsumsthatjustclimbabove1.5orjustdropbelow1.5,andlistthenew termsincludedineachsuchpartialsum. ANS.n1 2 3 4 5 sn1.5333 1.0333 1.5218 1.2718 1.5143 5.5 S ERIES OF FUNCTIONS Weextendourconceptofinfiniteseriestoincludethepossibilitythateachterm unmaybe afunctionofsomevariable, un=un(x).Numerousillustrationsofsuchseriesoffunctions appearinChapters11–14.The partialsumsbecomefunctionsof thevariable x, sn(x)=u1(x)+u2(x)+···+un(x), (5.65) asdoestheseriessum, definedasthelimitofthepartialsums: ∞summationdisplay n=1un(x)=S(x)=limn→∞sn(x). (5.66) So far we have concerned ourselves with the behavior of the partial sums as a function ofn.Nowweconsiderhowtheforegoingquantitiesdependon x.The keyconcepthereis thatof uniformconvergence. Uniform Convergence Ifforanysmall ε>0thereexistsanumber N,independentof xintheinterval[a,b](that is,a≤x≤b)suchthat vextendsinglevextendsingleS(x)−sn(x)vextendsinglevextendsingle<ε, foralln≥N, (5.67) then the series is said to be uniformly convergent in the interval [a,b]. This says that for our series to be uniformly convergent, it must be possible to find a finite Nso that the tail of the infinite series, |summationtext∞ i=N+1ui(x)|, will be less than an arbitrarily small εfor allxin thegiveninterval. Thiscondition,Eq.(5.67),whichdefinesuniformconvergence,isillustratedinFig.5.7. Thepointisthatnomatterhowsmall εistakentobe,wecanalwayschoose nlargeenough so that the absolute magnitude of the difference between S(x)andsn(x)is less than εfor allx,a≤x≤b.Ifthiscannotbedone,thensummationtextun(x)isnotuniformlyconvergentin [a,b]. Example 5.5.1 NONUNIFORM CONVERGENCE ∞summationdisplay n=1un(x)=∞summationdisplay n=1x [(n−1)x+1][nx+1]. (5.68) 5.5 Series of Functions 349 FIGURE 5.7Uniformconvergence. Thepartialsum sn(x)=nx(nx+1)−1,asmaybeverifiedby mathematicalinduction . By inspection this expression for sn(x)holds for n=1,2. We assume it holds for nterms andthenproveitholdsfor n+1t e r m s : sn+1(x)=sn(x)+x [nx+1][(n+1)x+1] =nx [nx+1]+x [nx+1][(n+1)x+1] =(n+1)x (n+1)x+1, completingtheproof. Lettingnapproachinfinity,weobtain S(0)=limn→∞sn(0)=0, S(x/negationslash=0)=limn→∞sn(x/negationslash=0)=1. We have a discontinuity in our series limit at x=0. However, sn(x)is a continuous func- tion ofx,0≤x≤1, for all finite n. No matter how small εmay be, Eq. (5.67) will be violatedfor allsufficientlysmall x. Ourseriesdoesnotconvergeuniformly. /squaresolid Weierstrass M(Majorant) Test The most commonly encountered test for uniform convergence is the Weierstrass Mtest. If we can construct a series of numberssummationtext∞ 1Mi, in which Mi≥|ui(x)|for allxin the interval[a,b]andsummationtext∞ 1Miis convergent, our series ui(x)will beuniformly convergent in[a,b]. 350 Chapter 5 Infinite Series TheproofofthisWeierstrass Mtestisdirectandsimple.Sincesummationtext iMiconverges,some numberNexistssuchthatfor n+1≥N, ∞summationdisplay i=n+1Mi<ε. (5.69) This follows from our definition of convergence. Then, with |ui(x)|≤Mifor allxin the intervala≤x≤b, ∞summationdisplay i=n+1vextendsinglevextendsingleui(x)vextendsinglevextendsingle<ε. (5.70) Hence vextendsingle vextendsingleS(x)−sn(x)vextendsinglevextendsingle=vextendsinglevextendsinglevextendsinglevextendsingle∞summationdisplay i=n+1ui(x)vextendsinglevextendsinglevextendsinglevextendsingle<ε, (5.71) and by definitionsummationtext∞ i=1ui(x)is uniformly convergent in [a,b]. Since we have specified absolute values in the statement of the Weierstrass Mtest, the seriessummationtext∞ i=1ui(x)is also seentobe absolutely convergent. Note that uniform convergence and absolute convergence are independent properties. Neitherimpliestheother.Forspecificexamples, ∞summationdisplay n=1(−1)n n+x2,−∞<x<∞, (5.72) and ∞summationdisplay n=1(−1)n−1xn n=ln(1+x),0≤x≤1, (5.73) converge uniformly in the indicated intervals but do not converge absolutely. On the other hand, ∞summationdisplay n=0(1−x)xn=1,0≤x<1 =0,x=1, (5.74) convergesabsolutelybutdoesnotconvergeuniformlyin [0,1]. Fromthedefinitionofuniformconvergencewemayshowthatanyseries f(x)=∞summationdisplay n=1un(x) (5.75) cannotconvergeuniformlyinanyintervalthatincludesadiscontinuityof f(x)ifallun(x) arecontinuous. Since the Weierstrass Mtest establishes both uniform and absolute convergence, it will necessarilyfail forseries thatareuniformlybutconditionallyconvergent. 5.5 Series of Functions 351 Abel’s Test Asomewhatmoredelicatetestfor uniformconvergencehasbeengivenbyAbel.If un(x)=anfn(x), summationdisplay an=A,convergent and the functions fn(x)are monotonic [fn+1(x)≤fn(x)]and bounded, 0 ≤fn(x)≤M, forallxin[a,b], thensummationtext nun(x)convergesuniformly in[a,b]. Thistestisespeciallyusefulinanalyzingpowerseries(compareSection5.7).Detailsof theproofofAbel’stestandothertestsforuniformconvergencearegivenintheAdditional Readingslistedattheendof thischapter. Uniformlyconvergentseries havethreeparticularlyusefulproperties. 1. If theindividualterms un(x)arecontinuous,theseries sum f(x)=∞summationdisplay n=1un(x) (5.76) isalsocontinuous. 2. If the individual terms un(x)are continuous, the series may be integrated term by term.Thesumoftheintegralsis equaltotheintegralofthesum. integraldisplayb af(x)dx=∞summationdisplay n=1integraldisplayb aun(x)dx. (5.77) 3. The derivative of the series sum f(x)equals the sum of the individual term deriva- tives: d dxf(x)=∞summationdisplay n=1d dxun(x), (5.78) providedthefollowingconditionsaresatisfied: un(x)anddun(x) dxarecontinuousin [a,b]. ∞summationdisplay n=1dun(x) dxisuniformlyconvergentin [a,b]. Term-by-term integration of a uniformly convergent series8requires only continuity of the individual terms. This condition is almost always satisfied in physical applications. Term-by-term differentiation of a series is often not valid because more restrictive condi- tions must be satisfied. Indeed, we shall encounter Fourier series in Chapter 14 in which term-by-termdifferentiationof auniformlyconvergentseries leadstoadivergentseries. 8Term-by-term integration may also bevalid in the absenceof uniform convergence. 352 Chapter 5 Infinite Series Exercises 5.5.1 Findtherangeof uniformconvergenceof theDirichletseries (a)∞summationdisplay n=1(−1)n−1 nx,(b)ζ(x)=∞summationdisplay n=11 nx. ANS.(a) 0 <s≤x<∞. (b) 1<s≤x<∞. 5.5.2 Forwhatrangeof xisthegeometricseriessummationtext∞ n=0xnuniformlyconvergent? ANS.−1<−s≤x≤s<1. 5.5.3 Forwhatrangeofpositivevaluesof xissummationtext∞n=01/(1+xn) (a) convergent? (b) uniformlyconvergent? 5.5.4 Iftheseriesofthecoefficientssummationtextanandsummationtextbnareabsolutelyconvergent,showthatthe Fourierseries summationdisplay (ancosnx+bnsinnx) isuniformly convergentfor −∞<x<∞. 5.6 T AYLOR ’SEXPANSION This is an expansion of a function into an infinite series of powers of a variable xor into a finite series plus a remainder term. The coefficients of the successive terms of the series involvethesuccessivederivativesofthefunction.WehavealreadyusedTaylor’sexpansion in the establishment of a physical interpretation of divergence (Section 1.7) and in other sectionsof Chapters1and2. NowwederivetheTaylorexpansion. We assume that our function f(x)has a continuous nth derivative9in the interval a≤ x≤b.Then,integratingthis nthderivative ntimes, integraldisplayx af(n)(x1)dx1=f(n−1)(x1)vextendsinglevextendsinglevextendsinglex a=f(n−1)(x)−f(n−1)(a), integraldisplayx adx2integraldisplayx2 adx1f(n)(x1)=integraldisplayx adx2bracketleftbig f(n−1)(x2)−f(n−1)(a)bracketrightbig (5.79) =f(n−2)(x)−f(n−2)(a)−(x−a)f(n−1)(a). Continuing,weobtain integraldisplayx adx3integraldisplayx3 adx2integraldisplayx2 adx1f(n)(x1)=f(n−3)(x)−f(n−3)(a)−(x−a)f(n−2)(a) −(x−a)2 2!f(n−1)(a). (5.80) 9Taylor’s expansion may be derived under slightly less restrictive conditions; compare H. Jeffreys and B. S. Jeffreys, Methods of Mathematical Physics ,3rd ed.Cambridge: Cambridge University Press (1956), Section 1.133. 5.6 Taylor’s Expansion 353 Finally,onintegratingfor the nthtime, integraldisplayx adxn···integraldisplayx2 adx1f(n)(x1)=f(x)−f(a)−(x−a)f′(a)−(x−a)2 2!f′′(a) −···−(x−a)n−1 (n−1)!f(n−1)(a). (5.81) Note that this expression is exact. No terms have been dropped, no approximations made. Now,solvingfor f(x),weha v e f(x)=f(a)+(x−a)f′(a) +(x−a)2 2!f′′(a)+···+(x−a)n−1 (n−1)!f(n−1)(a)+Rn.(5.82) Theremainder, Rn, isgivenbythe n-foldintegral Rn=integraldisplayx adxn···integraldisplayx2 adx1f(n)(x1). (5.83) This remainder, Eq. (5.83), may be put into a perhaps more practical form by using the meanvaluetheorem of integralcalculus: integraldisplayx ag(x)dx=(x−a)g(ξ), (5.84) witha≤ξ≤x.Byintegrating ntimeswegettheLagrangianform10of theremainder: Rn=(x−a)n n!f(n)(ξ). (5.85) With Taylor’s expansion in this form we are not concerned with any questions of infinite series convergence. This series is finite, and the only questions concern the magnitude of theremainder. Whenthefunction f(x)is suchthat limn→∞Rn=0, (5.86) Eq.(5.82) becomesTaylor’sseries: f(x)=f(a)+(x−a)f′(a)+(x−a)2 2!f′′(a)+··· =∞summationdisplay n=0(x−a)n n!f(n)(a).11(5.87) 10Analternateform derived by Cauchyis Rn=(x−ζ)n−1(x−a) (n−1)!f(n)(ζ), witha≤ζ≤x. 11Notethat 0!=1 (compare Section8.1). 354 Chapter 5 Infinite Series Our Taylor series specifies the value of a function at one point, x, in terms of the value ofthefunctionanditsderivativesatareferencepoint a.Itisanexpansioninpowersofthe changein the variable, /Delta1x=x−ain this case. The notation may be varied at the user’s convenience.Withthesubstitution x→x+handa→xwehaveanalternateform, f(x+h)=∞summationdisplay n=0hn n!f(n)(x). Whenweusethe operator D=d/dx,theTaylorexpansionbecomes f(x+h)=∞summationdisplay n=0hnDn n!f(x)=ehDf(x). (The transition to the exponential form anticipates Eq. (5.90), which follows.) An equiva- lent operator form of this Taylor expansion appears in Exercise 4.2.4. A derivation of the TaylorexpansioninthecontextofcomplexvariabletheoryappearsinSection6.5. Maclaurin Theorem If weexpandabouttheorigin (a=0), Eq. (5.87)is knownasMaclaurin’sseries: f(x)=f(0)+xf′(0)+x2 2!f′′(0)+···=∞summationdisplay n=0xn n!f(n)(0). (5.88) An immediate application of the Maclaurin series (or the Taylor series) is in the expan- sionofvarioustranscendentalfunctionsintoinfinite(power)series. Example 5.6.1 EXPONENTIAL FUNCTION Letf(x)=ex. Differentiating,wehave f(n)(0)=1 (5.89) foralln,n=1,2,3,....Then, withEq. (5.88), wehave ex=1+x+x2 2!+x3 3!+···=∞summationdisplay n=0xn n!. (5.90) This is the series expansion of the exponential function. Some authors use this series to definetheexponentialfunction. Althoughthisseriesisclearlyconvergentforall x,weshouldchecktheremainderterm, Rn. ByEq. (5.85)wehave Rn=xn n!f(n)(ξ)=xn n!eξ,0≤|ξ|≤x. (5.91) 5.6 Taylor’s Expansion 355 Therefore |Rn|≤xnex n!(5.92) and limn→∞Rn=0 (5.93) for allfinitevalues of x, which indicates that this Maclaurin expansion of exconverges absolutelyovertherange −∞<x<∞. /squaresolid Example 5.6.2 LOGARITHM Letf(x)=ln(1+x). Bydifferentiating,weobtain f′(x)=(1+x)−1, f(n)(x)=(−1)n−1(n−1)!(1+x)−n. (5.94) TheMaclaurinexpansion(Eq. (5.88)) yields ln(1+x)=x−x2 2+x3 3−x4 4+···+Rn =nsummationdisplay p=1(−1)p−1xp p+Rn. (5.95) Inthiscaseourremainderis givenby Rn=xn n!f(n)(ξ), 0≤ξ≤x ≤xn n,0≤ξ≤x≤1. (5.96) Now, the remainder approaches zero as nis increased indefinitely, provided 0 ≤x≤1.12 Asaninfiniteseries, ln(1+x)=∞summationdisplay n=1(−1)n−1xn n(5.97) converges for−1<x≤1. The range−1<x<1 is easily established by the d’Alembert ratiotest(Section5.2).Convergenceat x=1followsbytheLeibnizcriterion(Section5.3). Inparticular,at x=1w eh a v e ln2=1−1 2+1 3−1 4+1 5−···=∞summationdisplay n=1(−1)n−1n−1, (5.98) theconditionallyconvergentalternatingharmonicseries. /squaresolid 12This range caneasilybe extendedto −1<x≤1 but not to x=−1. 356 Chapter 5 Infinite Series Binomial Theorem A second, extremely important application of the Taylor and Maclaurin expansions is the derivationofthebinomialtheoremfornegativeand/ornonintegralpowers. Letf(x)=(1+x)m, in which mmay be negative and is not limited to integral values. Directapplicationof Eq.(5.88) gives (1+x)m=1+mx+m(m−1) 2!x2+···+Rn. (5.99) Forthisfunctiontheremainderis Rn=xn n!(1+ξ)m−nm(m−1)···(m−n+1) (5.100) andξlies between 0 and x,0≤ξ≤x.N o w ,f o r n>m ,(1+ξ)m−nis a maximum for ξ=0.Therefore Rn≤xn n!m(m−1)···(m−n+1). (5.101) Notethatthe mdependentfactorsdonotyieldazerounless misanonnegativeinteger; Rn tends to zero as n→∞ifxis restricted to the range 0 ≤x<1. The binomial expansion thereforeisshowntobe (1+x)m=1+mx+m(m−1) 2!x2+m(m−1)(m−2) 3!x3+···.(5.102) Inother,equivalentnotation, (1+x)m=∞summationdisplay n=0m! n!(m−n)!xn=∞summationdisplay n=0parenleftbiggm nparenrightbigg xn. (5.103) The quantityparenleftbigm nparenrightbig , which equals m!/[n!(m−n)!], is called a binomial coefficient .A l - thoughwehaveonlyshownthattheremaindervanishes, limn→∞Rn=0, for 0≤x<1, the series in Eq. (5.102) actually may be shown to be convergent for the extendedrange −1<x<1.Forman integer, (m−n)!=±∞ifn>m(Section8.1) and theseries automaticallyterminatesat n=m. Example 5.6.3 RELATIVISTIC ENERGY Thetotalrelativisticenergyofaparticleofmass mandvelocity vis E=mc2parenleftbigg 1−v2 c2parenrightbigg−1/2 . (5.104) Comparethisexpressionwiththeclassicalkineticenergy, mv2/2. 5.6 Taylor’s Expansion 357 ByEq. (5.102)with x=−v2/c2andm=−1/2w eh a v e E=mc2bracketleftbigg 1−1 2parenleftbigg −v2 c2parenrightbigg +(−1/2)(−3/2) 2!parenleftbigg −v2 c2parenrightbigg2 +(−1/2)(−3/2)(−5/2) 3!parenleftbigg −v2 c2parenrightbigg3 +···bracketrightbigg , or E=mc2+1 2mv2+3 8mv2·v2 c2+5 16mv2·parenleftbiggv2 c2parenrightbigg2 +···.(5.105) Thefirst term, mc2, isidentifiedastherest massenergy.Then Ekinetic=1 2mv2bracketleftbigg 1+3 4v2 c2+5 8parenleftbiggv2 c2parenrightbigg2 +···bracketrightbigg . (5.106) For particle velocity v≪c, the velocity of light, the expression in the brackets reduces to unity and we see that the kinetic portion of the total relativistic energy agrees with the classicalresult. /squaresolid Forpolynomialswecangeneralizethebinomialexpansionto (a1+a2+···+am)n=summationdisplay n! n1!n2!···nm!an1 1an2 2···anmm, where the summation includes all different combinations of n1,n2,...,nmwithsummationtextm i=1ni=n.H e r eniandnare all integral. This generalization finds considerable use instatisticalmechanics. MaclaurinseriesmaysometimesappearindirectlyratherthanbydirectuseofEq.(5.88). Forinstance,themostconvenientwaytoobtaintheseriesexpansion sin−1x=∞summationdisplay n=0(2n−1)!! (2n)!!·x2n+1 (2n+1)=x+x3 6+3x5 40+···, (5.106a) istomakeuseoftherelation(from sin y=x,getdy/dx=1/√ 1−x2) sin−1x=integraldisplayx 0dt (1−t2)1/2. We expand (1−t2)−1/2(binomial theorem) and then integrate term by term. This term- by-termintegrationisdiscussedinSection5.7.TheresultisEq.(5.106a).Finally,wemay takethelimitas x→1.TheseriesconvergesbyGauss’test, Exercise5.2.5. 358 Chapter 5 Infinite Series Taylor Expansion — More Than One Variable If the function fhas more than one independent variable, say, f=f(x,y), the Taylor expansionbecomes f(x,y)=f(a,b)+(x−a)∂f ∂x+(y−b)∂f ∂y +1 2!bracketleftbigg (x−a)2∂2f ∂x2+2(x−a)(y−b)∂2f ∂x∂y+(y−b)2∂2f ∂y2bracketrightbigg +1 3!bracketleftbigg (x−a)3∂3f ∂x3+3(x−a)2(y−b)∂3f ∂x2∂y +3(x−a)(y−b)2∂3f ∂x∂y2+(y−b)3∂3f ∂y3bracketrightbigg +···, (5.107) with all derivatives evaluated at the point (a,b).U s i n gαjt=xj−xj0, we may write the Taylorexpansionfor mindependentvariablesinthesymbolicform f(x1,...,xm)=∞summationdisplay n=0tn n!parenleftbiggmsummationdisplay i=1αi∂ ∂xiparenrightbiggn f(x1,...,xm)vextendsinglevextendsinglevextendsingle (xk=xk0,k=1,...,m).(5.108) Aconvenientvectorform for m=3i s ψ(r+a)=∞summationdisplay n=01 n!(a·∇)nψ(r). (5.109) Exercises 5.6.1 Showthat (a) sinx=∞summationdisplay n=0(−1)nx2n+1 (2n+1)!, (b) cosx=∞summationdisplay n=0(−1)nx2n (2n)!. InSection6.1, eixis definedbyaseries expansionsuchthat eix=cosx+isinx. Thisisthebasisforthepolarrepresentationofcomplexquantities.Asaspecialcasewe find,with x=π, theintriguingrelation eiπ=−1. 5.6 Taylor’s Expansion 359 5.6.2 Deriveaseries expansionof cot xinincreasingpowersof xbydividing cos xby sinx. Note. The resultant series that starts with 1 /xis actually a Laurent series (Section 6.5). Although the two series for sin xand cosxwere valid for all x, the convergence of the series for cot xis limited by the zeros of the denominator, sin x(see Analytic Continu- ationinSection6.5). 5.6.3 TheRaabetest forsummationtext n(nlnn)−1leadsto limn→∞nbracketleftbigg(n+1)ln(n+1) nlnn−1bracketrightbigg . Showthatthislimitis unity(whichmeansthattheRaabetesthereisindeterminate). 5.6.4 Showbyseries expansionthat 1 2lnη0+1 η0−1=coth−1η0,|η0|>1. Thisidentitymaybeusedtoobtainasecondsolutionfor Legendre’sequation. 5.6.5 Show that f(x)=x1/2(a) has no Maclaurin expansion but (b) has a Taylor expansion about any point x0/negationslash=0. Find the range of convergence of the Taylor expansion about x=x0. 5.6.6 Letxbe an approximation for a zero of f(x)and/Delta1xbe the correction. Show that by neglectingtermsoforder (/Delta1x)2, /Delta1x=−f(x) f′(x). This is Newton’s formula for finding a root. Newton’s method has the virtues of illus- tratingseries expansionsandelementarycalculusbutis verytreacherous. 5.6.7 Expand a function /Phi1(x,y,z) by Taylor’s expansion about (0,0,0)toO(a3). Evaluate ¯/Phi1, the average value of /Phi1, averaged over a small cube of side acentered on the origin andshowthattheLaplacianof /Phi1isameasureofdeviationof /Phi1from/Phi1(0,0,0). 5.6.8 Theratiooftwodifferentiablefunctions f(x)andg(x)takesontheindeterminateform 0/0a tx=x0. UsingTaylorexpansionsprove L’Hôpital’srule , limx→x0f(x) g(x)=limx→x0f′(x) g′(x). 5.6.9 Withn>1,showthat (a)1 n−lnparenleftbiggn n−1parenrightbigg <0,(b)1 n−lnparenleftbiggn+1 nparenrightbigg >0. Use these inequalities to show that the limit defining the Euler–Mascheroni constant, Eq. (5.28),is finite. 5.6.10 Expand(1−2tz+t2)−1/2inpowersof t.Assumethat tissmall.Collectthecoefficients oft0,t1, andt2. 360 Chapter 5 Infinite Series ANS.a0=P0(z)=1, a1=P1(z)=z, a2=P2(z)=1 2(3z2−1), where an=Pn(z),t h enthLegendrepolynomial. 5.6.11 UsingthedoublefactorialnotationofSection8.1, showthat (1+x)−m/2=∞summationdisplay n=0(−1)n(m+2n−2)!! 2nn!(m−2)!!xn, form=1,2,3,.... 5.6.12 Usingbinomialexpansions,comparethethreeDopplershiftformulas: (a)ν′=νparenleftbigg 1∓v cparenrightbigg−1 movingsource ; (b)ν′=νparenleftbigg 1±v cparenrightbigg movingobserver ; (c)ν′=νparenleftbigg 1±v cparenrightbiggparenleftbigg 1−v2 c2parenrightbigg−1/2 relativistic. Note.Therelativisticformulaagreeswiththeclassicalformulasiftermsoforder v2/c2 canbeneglected. 5.6.13 Inthetheoryofgeneralrelativitytherearevariouswaysofrelating(defining)avelocity ofrecessionof agalaxytoitsredshift, δ.Milne’smodel(kinematicrelativity)gives (a)v1=cδparenleftbigg 1+1 2δparenrightbigg , (b)v2=cδparenleftbigg 1+1 2δparenrightbigg (1+δ)−2, (c) 1+δ=bracketleftbigg1+v3/c 1−v3/cbracketrightbigg1/2 . 1.Showthatfor δ≪1 (andv3/c≪1)allthreeformulasreduceto v=cδ. 2.Comparethethreevelocitiesthroughtermsof order δ2. Note. In special relativity (with δreplaced by z), the ratio of observed wavelength λto emittedwavelength λ0isgivenby λ λ0=1+z=parenleftbiggc+v c−vparenrightbigg1/2 . 5.6.14 Therelativisticsum woftwo velocities uandvis givenby w c=u/c+v/c 1+uv/c2. 5.6 Taylor’s Expansion 361 If v c=u c=1−α, where 0≤α≤1,findw/cinpowersof αthroughtermsin α3. 5.6.15 The displacement xof a particle of rest mass m0, resulting from a constant force m0g alongthe x-axis,is x=c2 gbraceleftbiggbracketleftbigg 1+parenleftbigg gt cparenrightbigg2bracketrightbigg1/2 −1bracerightbigg , includingrelativisticeffects. Findthe displacement xas a powerseries intime t.C o m - parewiththeclassicalresult, x=1 2gt2. 5.6.16 By use of Dirac’s relativistic theory, the fine structure formula of atomic spectroscopy isgivenby E=mc2bracketleftbigg 1+γ2 (s+n−|k|)2bracketrightbigg−1/2 , where s=parenleftbig |k|2−γ2parenrightbig1/2,k=±1,±2,±3,.... Expandinpowersof γ2throughorder γ4(γ2=Ze2/4πε0¯hc,withZtheatomicnum- ber).ThisexpansionisusefulincomparingthepredictionsoftheDiracelectrontheory withthoseofarelativisticSchrödingerelectrontheory.Experimentalresultssupportthe Diractheory. 5.6.17 Inahead-onproton–protoncollision,theratioofthekineticenergyinthecenterofmass systemtotheincidentkineticenergyis R=bracketleftbigradicalBig 2mc2parenleftbig Ek+2mc2parenrightbig −2mc2bracketrightbig /Ek. Findthevalueof thisratioofkineticenergiesfor (a)Ek≪mc2(nonrelativistic) (b)Ek≫mc2(extreme-relativistic). ANS.(a)1 2,(b) 0. Thelatteransweris asortof law ofdiminishingreturnsfor high-energyparticle accelerators(withstationarytargets). 5.6.18 Withbinomialexpansions x 1−x=∞summationdisplay n=1xn,x x−1=1 1−x−1=∞summationdisplay n=0x−n. Addingthesetwoseries yieldssummationtext∞ n=−∞xn=0. Hopefully,wecanagreethatthisisnonsense,butwhathasgonewrong? 362 Chapter 5 Infinite Series 5.6.19 (a) Planck’stheoryof quantizedoscillatorsleadstoanaverageenergy /angbracketleftε/angbracketright=summationtext∞ n=1nε0exp(−nε0/kT)summationtext∞ n=0exp(−nε0/kT), whereε0is a fixed energy. Identify the numerator and denominator as binomial expansionsandshowthattheratiois /angbracketleftε/angbracketright=ε0 exp(ε0/kT)−1. (b) Showthatthe /angbracketleftε/angbracketrightofpart(a) reducesto kT,theclassicalresult,for kT≫ε0. 5.6.20 (a) ExpandbythebinomialtheoremandintegratetermbytermtoobtaintheGregory seriesfor y=tan−1x(notethat tan y=x): tan−1x=integraldisplayx 0dt 1+t2=integraldisplayx 0braceleftbig 1−t2+t4−t6+···bracerightbig dt =∞summationdisplay n=0(−1)nx2n+1 2n+1,−1≤x≤1. (b) Bycomparingseriesexpansions,showthat tan−1x=i 2lnparenleftbigg1−ix 1+ixparenrightbigg . Hint.CompareExercise5.4.1. 5.6.21 Innumericalanalysisitisoftenconvenienttoapproximate d2ψ(x)/dx2by d2 dx2ψ(x)≈1 h2bracketleftbig ψ(x+h)−2ψ(x)+ψ(x−h)bracketrightbig . Findtheerror inthisapproximation. ANS. Error=h2 12ψ(4)(x). 5.6.22 Y ouha v eafunction y(x)tabulatedatequallyspacedvaluesoftheargument braceleftBigg yn=y(xn) xn=x+nh. Showthatthelinearcombination 1 12h{−y2+8y1−8y−1+y−2} yields y′ 0−h4 30y(5) 0+···. Hence this linear combination yields y′ 0if(h4/30)y(5) 0and higher powers of hand higherderivativesof y(x)arenegligible. 5.7 Power Series 363 5.6.23 Inanumericalintegrationofapartialdifferentialequation,thethree-dimensionalLapla- cianis replacedby ∇2ψ(x,y,z)→h−2bracketleftbig ψ(x+h,y,z)+ψ(x−h,y,z) +ψ(x,y+h,z)+ψ(x,y−h,z)+ψ(x,y,z+h) +ψ(x,y,z−h)−6ψ(x,y,z)bracketrightbig . Determinetheerrorinthisapproximation.Here histhestepsize,thedistancebetween adjacentpointsinthe x-,y-, orz-direction. 5.6.24 Usingdoubleprecision,calculate efromits Maclaurinseries. Note. This simple, direct approach is the best way of calculating eto high accuracy. Sixteen terms give eto 16 significant figures. The reciprocal factorials give very rapid convergence. 5.7 P OWER SERIES Thepowerseries is aspecialandextremelyusefultypeofinfiniteseriesof theform f(x)=a0+a1x+a2x2+a3x3+···=∞summationdisplay n=0anxn, (5.110) wherethecoefficients aiareconstants,independentof x.13 Convergence Equation (5.110) may readily be tested for convergence by either the Cauchy root test or thed’Alembertratiotest(Section5.2). If limn→∞vextendsinglevextendsinglevextendsinglevextendsinglean+1 anvextendsinglevextendsinglevextendsinglevextendsingle=R−1, (5.111) the series converges for −R<x<R . This is the interval or radius of convergence. Since the root and ratio tests fail when the limit is unity, the endpoints of the interval require specialattention. For instance, if an=n−1, thenR=1 and, from Sections 5.1, 5.2, and 5.3, the series converges for x=−1 but diverges for x=+1. Ifan=n!, thenR=0 and the series divergesfor all x/negationslash=0. Uniform and Absolute Convergence Suppose our power series (Eq. (5.110)) has been found convergent for −R<x<R ; then itwillbeuniformlyandabsolutelyconvergentinany interiorinterval,−S≤x≤S,where 0<S<R. This maybeproveddirectlybytheWeierstrass Mtest(Section5.5). 13Equation (5.110) may be generalized to z=x+iy, replacing x. The following two chapters will then yield uniform conver- gence,integrability, anddifferentiability in aregion ofacomplex plane in placeof aninterval onthe x-axis. 364 Chapter 5 Infinite Series Continuity Since each of the terms un(x)=anxnis a continuous function of xandf(x)=summationtextanxn converges uniformly for −S≤x≤S,f(x)must be a continuous function in the interval ofuniformconvergence. ThisbehavioristobecontrastedwiththestrikinglydifferentbehavioroftheFourierse- ries(Chapter14),inwhichtheFourierseriesisusedfrequentlytorepresentdiscontinuous functionssuchassawtoothandsquarewaves. Differentiation and Integration Withun(x)continuous andsummationtextanxnuniformly convergent, we find that the differentiated series is a power series with continuous functions and the same radius of convergence as the original series. The new factors introduced by differentiation (or integration) do not affect either the root or the ratio test. Therefore our power series may be differentiated or integratedasoftenasdesiredwithintheintervalofuniformconvergence(Exercise5.7.13). In view of the rather severe restrictions placed on differentiation (Section 5.5), this is aremarkableandvaluableresult. Uniqueness Theorem In the preceding section, using the Maclaurin series, we expanded exand ln(1+x)into infinite series. In the succeeding chapters, functions are frequently represented or perhaps definedbyinfiniteseries. We nowestablishthatthepower-seriesrepresentationisunique. If f(x)=∞summationdisplay n=0anxn,−Ra<x<R a =∞summationdisplay n=0bnxn,−Rb<x<R b, (5.112) withoverlappingintervalsofconvergence,includingtheorigin,then an=bn (5.113) for alln; that is, we assume two (different) power-series representations and then proceed toshowthatthetwoareactuallyidentical. FromEq. (5.112), ∞summationdisplay n=0anxn=∞summationdisplay n=0bnxn,−R<x<R, (5.114) whereRis the smaller of Ra,Rb. By setting x=0 to eliminate all but the constant terms, weobtain a0=b0. (5.115) 5.7 Power Series 365 Now, exploiting the differentiability of our power series, we differentiate Eq. (5.114), get- ting ∞summationdisplay n=1nanxn−1=∞summationdisplay n=1nbnxn−1. (5.116) Weagainset x=0,toisolatethenewconstantterms, andfind a1=b1. (5.117) Byrepeatingthisprocess ntimes,weget an=bn, (5.118) which shows that the two series coincide. Therefore our power-series representation is unique. This will be a crucial point in Section 9.5, in which we use a power series to develop solutions of differential equations. This uniqueness of power series appears frequently in theoreticalphysics.Theestablishmentofperturbationtheoryinquantummechanicsisone example.Thepower-seriesrepresentationoffunctionsisoftenusefulinevaluatingindeter- minateforms,particularlywhenl’Hôpital’srulemaybeawkwardtoapply(Exercise5.7.9). Example 5.7.1 L’HÔPITAL ’SRULE Evaluate lim x→01−cosxx2. (5.119) Replacing cos xbyits Maclaurin-seriesexpansion,weobtain 1−cosx x2=1−(1−1 2!x2+1 4!x4−···) x2=1 2!−x2 4!+···. Lettingx→0,wehave lim x→01−cosx x2=1 2. (5.120) Theuniquenessofpowerseriesmeansthatthecoefficients anmaybeidentifiedwiththe derivativesinaMaclaurinseries. From f(x)=∞summationdisplay n=0anxn=∞summationdisplay n=01 n!f(n)(0)xn wehave an=1 n!f(n)(0). /squaresolid 366 Chapter 5 Infinite Series Inversion of Power Series Supposewearegivenaseries y−y0=a1(x−x0)+a2(x−x0)2+···=∞summationdisplay n=1an(x−x0)n.(5.121) This gives (y−y0)in terms of (x−x0). However, it may be desirable to have an explicit expression for (x−x0)in terms of (y−y0). We may solve Eq. (5.121) for x−x0by inversionofourseries. Assumethat x−x0=∞summationdisplay n=1bn(y−y0)n, (5.122) with thebnto be determined in terms of the assumed known an. A brute-force approach, whichisperfectlyadequateforthefirstfewcoefficients,issimplytosubstituteEq.(5.121) into Eq. (5.122). By equating coefficients of (x−x0)non both sides of Eq. (5.122), since thepowerseries isunique,weobtain b1=1 a1, b2=−a2 a3 1, b3=1 a5 1parenleftbig 2a2 2−a1a3parenrightbig , (5.123) b4=1 a7 1parenleftbig 5a1a2a3−a2 1a4−5a3 2parenrightbig ,andsoon . Some of the higher coefficients are listed by Dwight.14A more general and much more elegant approach is developed by the use of complex variables in the first and second editionsof MathematicalMethodsforPhysicists . Exercises 5.7.1 TheclassicalLangevintheoryofparamagnetismleadstoanexpressionforthemagnetic polarization, P(x)=cparenleftbiggcoshx sinhx−1 xparenrightbigg . ExpandP(x)asapowerseries forsmall x(lowfields, hightemperature). 14H. B. Dwight, Tables of Integrals and Other Mathematical Data , 4th ed. New York: Macmillan (1961). (Compare Formula No.50.) 5.7 Power Series 367 5.7.2 The depolarizing factor Lfor an oblate ellipsoid in a uniform electric field parallel to theaxisofrotationis L=1 ε0parenleftbig 1+ζ2 0parenrightbigparenleftbig 1−ζ0cot−1ζ0parenrightbig , whereζ0definesanoblateellipsoidinoblatespheroidalcoordinates (ξ,ζ,ϕ).Showthat lim ζ0→∞L=1 3ε0(sphere),lim ζ0→0L=1 ε0(thinsheet) . 5.7.3 Thedepolarizingfactor (Exercise5.7.2)for aprolateellipsoidis L=1 ε0parenleftbig η2 0−1parenrightbigparenleftbigg1 2η0lnη0+1 η0−1−1parenrightbigg . Showthat limη0→∞L=1 3ε0(sphere),lim η0→0L=0 (longneedle) . 5.7.4 Theanalysisofthediffractionpatternofacircularopeninginvolves integraldisplay2π 0cos(ccosϕ)dϕ. Expandtheintegrandinaseries andintegratebyusing integraldisplay2π 0cos2nϕdϕ=(2n)! 22n(n!)2·2π,integraldisplay2π 0cos2n+1ϕdϕ=0. Theresultis 2 πtimestheBesselfunction J0(c). 5.7.5 Neutrons are created (by a nuclear reaction) inside a hollow sphere of radius R.T h e newly created neutrons are uniformly distributed over the spherical volume. Assuming thatalldirectionsareequallyprobable(isotropy),whatistheaveragedistanceaneutron will travel before striking the surface of the sphere? Assume straight-line motion and nocollisions. (a) Showthat ¯r=3 2Rintegraldisplay1 0integraldisplayπ 0radicalbig 1−k2sin2θk2dksinθdθ. (b) Expandtheintegrandas aseriesandintegratetoobtain ¯r=Rbracketleftbigg 1−3∞summationdisplay n=11 (2n−1)(2n+1)(2n+3)bracketrightbigg . (c) Showthatthesumof thisinfiniteseriesis 1 /12,giving¯r=3 4R. Hint. Show that sn=1/12−[4(2n+1)(2n+3)]−1by mathematical induction. Then letn→∞. 368 Chapter 5 Infinite Series 5.7.6 Giventhat integraldisplay1 0dx 1+x2=tan−1xvextendsinglevextendsinglevextendsingle1 0=π 4, expandtheintegrandintoaseries andintegratetermbytermobtaining15 π 4=1−1 3+1 5−1 7+1 9−···+(−1)n1 2n+1+···, whichis Leibniz’sformulafor π. Comparetheconvergenceof theintegrandseries and theintegratedseries at x=1.SeealsoExercise5.7.18. 5.7.7 Expandtheincompletefactorialfunction γ(n+1,x)≡integraldisplayx 0e−ttndt inaseries ofpowersof x.Whatis therangeof convergenceof theresultingseries? ANS.integraldisplayx 0e−ttndt=xn+1bracketleftbigg1 n+1−x n+2+x2 2!(n+3) −···(−1)pxp p!(n+p+1)+···bracketrightbigg . 5.7.8 Derivetheseriesexpansionoftheincompletebetafunction Bx(p,q)=integraldisplayx 0tp−1(1−t)q−1dt =xpbraceleftbigg1 p+1−q p+1x+···+(1−q)···(n−q) n!(p+n)xn+···bracerightbigg for 0≤x≤1,p>0,andq>0( i fx=1). 5.7.9 Evaluate (a) lim x→0bracketleftbig sin(tanx)−tan(sinx)bracketrightbig x−7,(b) lim x→0x−njn(x), n=3, wherejn(x)isasphericalBesselfunction(Section11.7), definedby jn(x)=(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbiggsinx xparenrightbigg . ANS.(a)−1 30,(b)1 1·3·5···(2n+1)→1 105forn=3. 15The series expansion of tan−1x(upper limit 1 replaced by x) was discovered by James Gregory in 1671, 3 years before Leibniz. See Peter Beckmann’s entertaining book, A History of Pi , 2nd ed., Boulder, CO: Golem Press (1971) and L. Berggren, J. andP.Borwein, Pi:ASource Book , NewYork: Springer (1997). 5.7 Power Series 369 5.7.10 Neutrontransporttheorygivesthefollowingexpressionfortheinverseneutrondiffusion lengthof k: a−b ktanh−1parenleftbiggk aparenrightbigg =1. By series inversion or otherwise, determine k2as a series of powers of b/a.G i v et h e firsttwoterms oftheseries. ANS.k2=3abparenleftbigg 1−4 5b aparenrightbigg . 5.7.11 Developaseries expansionof y=sinh−1x(thatis, sinh y=x)inpo wersof xby (a) inversionoftheseriesfor sinh y, (b) adirectMaclaurinexpansion. 5.7.12 Afunction f(z)isrepresentedbya descending powerseries f(z)=∞summationdisplay n=0anz−n,R≤z<∞. Show that this series expansion is unique; that is, if f(z)=summationtext∞ n=0bnz−n, R≤z<∞,thenan=bnfor alln. 5.7.13 A power series converges for −R<x<R . Show that the differentiated series and the integrated series have the same interval of convergence. (Do not bother about the endpoints x=±R.) 5.7.14 Assuming that f(x)may be expanded in a power series about the origin, f(x)=summationtext∞ n=0anxn, with some nonzero range of convergence. Use the techniques employed inprovinguniquenessofseries toshowthatyourassumedseriesis aMaclaurinseries: an=1 n!f(n)(0). 5.7.15 The Klein–Nishina formula for the scattering of photons by electrons contains a term oftheform f(ε)=(1+ε) ε2bracketleftbigg2+2ε 1+2ε−ln(1+2ε) εbracketrightbigg . Hereε=hν/mc2, theratioofthephotonenergytotheelectronrestmass energy.Find lim ε→0f(ε). ANS.4 3. 5.7.16 The behavior of a neutron losing energy by colliding elastically with nuclei of mass A isdescribedbyaparameter ξ1, ξ1=1+(A−1)2 2AlnA−1 A+1. 370 Chapter 5 Infinite Series Anapproximation,goodfor large A,is ξ2=2 A+2/3. Expandξ1andξ2in powers of A−1. Show that ξ2agrees with ξ1through(A−1)2.F i n d thedifferenceinthecoefficientsof the (A−1)3term. 5.7.17 Showthateachof thesetwointegralsequalsCatalan’sconstant: (a)integraldisplay1 0arctantdt t,(b)−integraldisplay1 0lnxdx 1+x2. Note.Seeβ(2)inSection5.9 forthevalueof Catalan’sconstant. 5.7.18 Calculate π(doubleprecision)byeachofthefollowingarctangentexpressions: π=16tan−1(1/5)−4tan−1(1/239) π=24tan−1(1/8)+8tan−1(1/57)+4tan−1(1/239) π=48tan−1(1/18)+32tan−1(1/57)−20tan−1(1/239). Obtain16significantfigures. VerifytheformulasusingExercise5.6.2. Note.Theseformulashavebeenusedinsomeofthemoreaccuratecalculationsof π.16 5.7.19 Ananalysisof theGibbsphenomenonofSection14.5leadstotheexpression 2 πintegraldisplayπ 0sinξ ξdξ. (a) Expand the integrand in a series and integrate term by term. Find the numerical valueof thisexpressiontofoursignificantfigures. (b) EvaluatethisexpressionbytheGaussianquadratureif available. ANS.1.178980. 5.8 E LLIPTIC INTEGRALS Elliptic integrals are included here partly as an illustration of the use of power series and partly for their own intrinsic interest. This interest includes the occurrence of elliptic inte- grals in physical problems (Example 5.8.1 and Exercise 5.8.4) and applications in mathe- maticalproblems. Example 5.8.1 PERIOD OF A SIMPLE PENDULUM Forsmall-amplitudeoscillations,ourpendulum(Fig.5.8)hassimpleharmonicmotionwith aperiodT=2π(l/g)1/2.Foramaximumamplitude θMlargeenoughsothatsin θM/negationslash=θM, Newton’ssecondlawofmotionandLagrange’sequation(Section17.7)leadtoanonlinear differentialequation(sin θisanonlinearfunctionof θ),soweturntoadifferentapproach. 16D.Shanks and J. W.Wrench, Computation of πto 100000 decimals. Math. Comput. 16: 76 (1962). 5.8 Elliptic Integrals 371 FIGURE 5.8Simple pendulum. The swinging mass mhas a kinetic energy of ml2(dθ/dt)2/2 and a potential energy of −mglcosθ(θ=π/2 taken for the arbitrary zero of potential energy). Since dθ/dt=0a t θ=θM, conservationofenergygives 1 2ml2parenleftbiggdθ dtparenrightbigg2 −mglcosθ=−mglcosθM. (5.124) Solvingfor dθ/dtweobtain dθ dt=±parenleftbigg2g lparenrightbigg1/2 (cosθ−cosθM)1/2, (5.125) with the mass mcanceling out. We take tto be zero when θ=0 anddθ/dt > 0. An integrationfrom θ=0t oθ=θMyields integraldisplayθM 0(cosθ−cosθM)−1/2dθ=parenleftbigg2g lparenrightbigg1/2integraldisplayt 0dt=parenleftbigg2g lparenrightbigg1/2 t. (5.126) This is1 4of a cycle, and therefore the time tis1 4of the period T. We note that θ≤θM, andwithabitofclairvoyancewetry thehalf-anglesubstitution sinparenleftbiggθ 2parenrightbigg =sinparenleftbiggθM 2parenrightbigg sinϕ. (5.127) Withthis,Eq. (5.126)becomes T=4parenleftbiggl gparenrightbigg1/2integraldisplayπ/2 0parenleftbigg 1−sin2parenleftbiggθM 2parenrightbigg sin2ϕparenrightbigg−1/2 dϕ. (5.128) AlthoughnotanobviousimprovementoverEq.(5.126),theintegralnowdefinesthecom- pleteellipticintegralofthefirstkind, K(sin2θM/2).Fromtheseriesexpansion,theperiod ofourpendulummaybedevelopedasapowerseries—powersof sin θM/2: T=2πparenleftbiggl gparenrightbigg1/2braceleftbigg 1+1 4sin2θM 2+9 64sin4θM 2+···bracerightbigg . (5.129) /squaresolid 372 Chapter 5 Infinite Series Definitions GeneralizingExample5.8.1toincludetheupperlimitasavariable,the ellipticintegralof thefirstkind isdefinedas F(ϕ\α)=integraldisplayϕ 0parenleftbig 1−sin2αsin2θparenrightbig−1/2dθ, (5.130a) or F(x|m)=integraldisplayx 0bracketleftbigparenleftbig 1−t2parenrightbigparenleftbig 1−mt2parenrightbigbracketrightbig−1/2dt,0≤m<1. (5.130b) (This is the notation of AMS-55 see footnote 4 for the reference.) For ϕ=π/2,x=1,we havethecompleteellipticintegralof thefirstkind , K(m)=integraldisplayπ/2 0parenleftbig 1−msin2θparenrightbig−1/2dθ =integraldisplay1 0bracketleftbigparenleftbig 1−t2parenrightbigparenleftbig 1−mt2parenrightbigbracketrightbig−1/2dt, (5.131) withm=sin2α,0≤m<1. Theellipticintegralofthesecondkind is definedby E(ϕ\α)=integraldisplayϕ 0parenleftbig 1−sin2αsin2θparenrightbig1/2dθ (5.132a) or E(x|m)=integraldisplayx 0parenleftbigg1−mt2 1−t2parenrightbigg1/2 dt,0≤m≤1. (5.132b) Again,forthecase ϕ=π/2,x=1,wehavethe completeellipticintegralofthesecond kind: E(m)=integraldisplayπ/2 0parenleftbig 1−msin2θparenrightbig1/2dθ =integraldisplay1 0parenleftbigg1−mt2 1−t2parenrightbigg1/2 dt,0≤m≤1. (5.133) Exercise5.8.1isanexampleofitsoccurrence.Figure5.9showsthebehaviorof K(m)and E(m).ExtensivetablesareavailableinAMS-55(see Exercise5.2.22for thereference). Series Expansion For our range 0 ≤m<1, the denominator of K(m)may be expanded by the binomial series parenleftbig 1−msin2θparenrightbig−1/2=1+1 2msin2θ+3 8m2sin4θ+··· =∞summationdisplay n=0(2n−1)!! (2n)!!mnsin2nθ. (5.134) 5.8 Elliptic Integrals 373 FIGURE 5.9Completeellipticintegrals, K(m)andE(m). For any closed interval [0,mmax],mmax<1, this series is uniformly convergent and may beintegratedterm byterm.FromExercise8.4.9, integraldisplayπ/2 0sin2nθdθ=(2n−1)!! (2n)!!·π 2. (5.135) Hence K(m)=π 2braceleftbigg 1+∞summationdisplay n=1bracketleftbigg(2n−1)!! (2n)!!bracketrightbigg2 mnbracerightbigg . (5.136) Similarly, E(m)=π 2braceleftbigg 1−∞summationdisplay n=1bracketleftbigg(2n−1)!! (2n)!!bracketrightbigg2mn 2n−1bracerightbigg (5.137) (Exercise 5.8.2). In Section 13.5 these series are identified as hypergeometric functions, andwehave K(m)=π 22F1parenleftbigg1 2,1 2;1;mparenrightbigg (5.138) E(m)=π 22F1parenleftbigg −1 2,1 2;1;mparenrightbigg . (5.139) 374 Chapter 5 Infinite Series Limiting Values Fromtheseries Eqs. (5.136)and(5.137),or fromthedefiningintegrals, lim m→0K(m)=π 2, (5.140) lim m→0E(m)=π 2. (5.141) Form→1 theseriesexpansionsareof littleuse.However,theintegralsyield lim m→1K(m)=∞, (5.142) theintegraldiverginglogarithmically,and lim m→1E(m)=1. (5.143) Theellipticintegralshavebeenusedextensivelyinthepastforevaluatingintegrals.For instance,integralsof theform I=integraldisplayx 0Rparenleftbig t,radicalbig a4t4+a3t3+a2t2+a1t1+a0parenrightbig dt, whereRisarationalfunctionof tandoftheradical,maybeexpressedintermsofelliptic integrals. Jahnke and Emde, Tables of Functions with Formulae and Curves .N e wY o r k : Dover (1943), Chapter 5, give pages of such transformations. With computers available for direct numerical evaluation, interest in these elliptic integral techniques has declined. However, elliptic integrals still remain of interest because of their appearance in physical problems—seeExercises5.8.4and5.8.5. For an extensive account of elliptic functions, integrals, and Jacobi theta functions, you aredirectedtoWhittakerandWatson’streatise ACourseinModernAnalysis ,4thed.Cam- bridge,UK:CambridgeUniversityPress (1962). Exercises 5.8.1 The ellipse x2/a2+y2/b2=1 may be represented parametrically by x=asinθ,y= bcosθ.Showthatthelengthofarc withinthefirst quadrantis aintegraldisplayπ/2 0parenleftbig 1−msin2θparenrightbig1/2dθ=aE(m). Here 0≤m=(a2−b2)/a2≤1. 5.8.2 Derivetheseriesexpansion E(m)=π 2braceleftbigg 1−parenleftbigg1 2parenrightbigg2m 1−parenleftbigg1·3 2·4parenrightbigg2m2 3−···bracerightbigg . 5.8.3 Showthat lim m→0(K−E) m=π 4. 5.8 Elliptic Integrals 375 FIGURE 5.10Circularwireloop. 5.8.4 Acircularloopofwireinthe xy-plane,asshowninFig.5.10,carriesacurrent I.Gi ven thatthevectorpotentialis Aϕ(ρ,ϕ,z)=aµ0I 2πintegraldisplayπ 0cosαdα (a2+ρ2+z2−2aρcosα)1/2, showthat Aϕ(ρ,ϕ,z)=µ0I πkparenleftbigga ρparenrightbigg1/2bracketleftbiggparenleftbigg 1−k2 2parenrightbigg Kparenleftbig k2parenrightbig −Eparenleftbig k2parenrightbigbracketrightbigg , where k2=4aρ (a+ρ)2+z2. Note.ForextensionofExercise5.8.4to B, seeSmythe,p. 270.17 5.8.5 An analysis of the magnetic vector potential of a circular current loop leads to the ex- pression fparenleftbig k2parenrightbig =k−2bracketleftbigparenleftbig 2−k2parenrightbig Kparenleftbig k2parenrightbig −2Eparenleftbig k2parenrightbigbracketrightbig , whereK(k2)andE(k2)arethecompleteellipticintegralsofthefirstandsecondkinds. Showthatfor k2≪1(r≫radiusof loop) fparenleftbig k2parenrightbig ≈πk2 16. 17W. R.Smythe, Static and DynamicElectricity ,3rd ed. NewYork: McGraw-Hill(1969). 376 Chapter 5 Infinite Series 5.8.6 Showthat (a)dE(k2) dk=1 k(E−K), (b)dK(k2) dk=E k(1−k2)−K k. Hint.Forpart(b)showthat Eparenleftbig k2parenrightbig =parenleftbig 1−k2parenrightbigintegraldisplayπ/2 0parenleftbig 1−ksin2θparenrightbig−3/2dθ bycomparingseries expansions. 5.8.7 (a) Write a function subroutine that will compute E(m)from the series expansion, Eq.(5.137). (b) Test your function subroutine by using it to calculate E(m)over the range m=0.0(0.1)0.9 and comparing the result with the values given by AMS-55 (see Exercise5.2.22for thereference). 5.8.8 RepeatExercise5.8.7for K(m). Note. These series for E(m), Eq. (5.137), and K(m), Eq. (5.136), converge only very slowly for mnear 1. More rapidly converging series for E(m)andK(m)exist. See Dwight’s Tables of Integrals:18No. 773.2 and 774.2. Your computer subroutine for computing EandKprobablyuses polynomialapproximations:AMS-55,Chapter17. 5.8.9 A simple pendulum is swinging with a maximum amplitude of θM. In the limit as θM→0, the period is 1 s. Using the elliptic integral, K(k2),k=sin(θM/2), calculate theperiod TforθM=0(1 0◦)9 0◦. Caution.Someellipticintegralsubroutinesrequire k=m1/2asaninputparameter,not mitself. Checkvalues .θM10◦50◦90◦ T(sec)1.00193 1.05033 1.18258 5.8.10 Calculate the magnetic vector potential A(ρ,ϕ,z)=ˆϕAϕ(ρ,ϕ,z)of a circular current loop(Exercise5.8.4) for theranges ρ/a=2,3,4,andz/a=0,1,2,3,4. Note. This elliptic integral calculationof the magneticvector potentialmay be checked byanassociatedLegendrefunctioncalculation,Example12.5.1. Checkvalue .F orρ/a=3 andz/a=0;Aϕ=0.029023µ0I. 5.9 B ERNOULLI NUMBERS , EULER –M ACLAURIN FORMULA The Bernoulli numbers were introduced by Jacques (James, Jacob) Bernoulli. There are several equivalent definitions, but extreme care must be taken, for some authors introduce 18H.B. Dwight, Tables of Integrals and OtherMathematical Data . NewYork: Macmillan (1947). 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 377 variations in numbering or in algebraic signs. One relatively simple approach is to define theBernoullinumbersbytheseries19 x ex−1=∞summationdisplay n=0Bnxn n!, (5.144) which converges for |x|<2πby the ratio test substitut Eq. (5.153) (see also Exam- ple7.1.7).Bydifferentiatingthispowerseriesrepeatedlyandthensetting x=0,weobtain Bn=bracketleftbiggdn dxnparenleftbiggx ex−1parenrightbiggbracketrightbigg x=0. (5.145) Specifically, B1=d dxparenleftbiggx ex−1parenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle x=0=1 ex−1−xex (ex−1)2vextendsinglevextendsinglevextendsinglevextendsingle x=0=−1 2,(5.146) as may be seen by series expansionof the denominators. Using B0=1 andB1=−1 2,i ti s easytoverifythatthefunction x ex−1−1+x 2=∞summationdisplay n=2Bnxn n!=−x e−x−1−1−x 2(5.147) isevenin x,s oal lB2n+1=0. ToderivearecursionrelationfortheBernoullinumbers,wemultiply ex−1 xx ex−1=1=braceleftbigg∞summationdisplay m=0xm (m+1)!bracerightbiggbraceleftbigg 1−x 2+∞summationdisplay n=1B2nx2n (2n)!bracerightbigg =1+∞summationdisplay m=1xmbraceleftbigg1 (m+1)!−1 2m!bracerightbigg +∞summationdisplay N=2xNsummationdisplay 1≤n≤N/2B2n (2n)!(N−2n+1)!.(5.148) ForN>0 thecoefficientof xNis zero,soEq. (5.148)yields 1 2(N+1)−1=summationdisplay 1≤n≤N/2B2nparenleftbiggN+1 2nparenrightbigg =1 2(N−1), (5.149) 19The function x/(ex−1)may be considered a generating function since it generates the Bernoulli numbers. Generating functions of the specialfunctions of mathematicalphysics appearin Chapters 11, 12, and 13. 378 Chapter 5 Infinite Series Table 5.1 BernoulliNumbers nB n Bn 01 1.000000000 1−1 2−0.500000000 21 60.166666667 4−1 30−0.033333333 61 420.023809524 8−1 30−0.033333333 105 660.075757576 Note. Further values are given in National Bureau of Stan- dards,Handbook of Mathematical Functions (AMS-55). Seefootnote4 for thereference. whichisequivalentto N−1 2=Nsummationdisplay n=1B2nparenleftbigg2N+1 2nparenrightbigg , N−1=N−1summationdisplay n=1B2nparenleftbigg2N 2nparenrightbigg .(5.150) FromEq.(5.150)theBernoullinumbersinTable5.1arereadilyobtained.Ifthevariable x in Eq. (5.144) is replaced by 2 ixwe obtain an alternate (and equivalent) definition of B2n (B1isset equalto−1 2byEq. (5.146))bytheexpression xcotx=∞summationdisplay n=0(−1)nB2n(2x)2n (2n)!,−π<x<π . (5.151) Using the method of residues (Section 7.1) or working from the infinite product represen- tationof sin x(Section5.11),wefindthat B2n=(−1)n−12(2n)! (2π)2n∞summationdisplay p=1p−2n,n=1,2,3,.... (5.152) This representation of the Bernoulli numbers was discovered by Euler. It is readily seen fromEq.(5.152)that |B2n|increaseswithoutlimitas n→∞.Numericalvalueshavebeen calculated by Glaisher.20Illustrating the divergent behavior of the Bernoulli numbers, we have B20=−5.291×102 B200=−3.647×10215. 20J. W. L. Glaisher, table of the first 250 Bernoulli’s numbers (to nine figures) and their logarithms (to ten figures). Trans. Cambridge Philos. Soc. 12: 390 (1871–1879). 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 379 SomeauthorsprefertodefinetheBernoullinumberswithamodifiedversionofEq.(5.152) byusing Bn=2(2n)! (2π)2n∞summationdisplay p=1p−2n, (5.153) thesubscriptbeingjusthalfofoursubscriptandallsignspositive.Again,whenusingother textsor references,youmustchecktoseeexactlyhowtheBernoullinumbersaredefined. TheBernoullinumbersoccurfrequentlyinnumbertheory.ThevonStaudt–Clausenthe- oremstates that B2n=An−1 p1−1 p2−1 p3−···−1 pk, (5.154) in which Anis an integer and p1,p2,...,pkare prime numbers so that pi−1 is a divisor of 2n.It mayreadilybeverifiedthatthisholdsfor B6(A3=1,p=2,3,7), B8(A4=1,p=2,3,5), (5.155) B10(A5=1,p=2,3,11), andotherspecialcases. TheBernoullinumbersappearinthesummationofintegralpowersoftheintegers, Nsummationdisplay j=1jp,pintegral, and in numerous series expansions of the transcendental functions, including tan x, cotx, ln|sinx|,(sinx)−1,l n|cosx|,l n|tanx|,(coshx)−1, tanhx,and coth x. Forexample, tanx=x+x3 3+2 15x5+···+(−1)n−122n(22n−1)B2n (2n)!x2n−1+···.(5.156) TheBernoullinumbersarelikelytocomeinsuchseriesexpansionsbecauseofthedefining equations (5.144), (5.150), and (5.151) and because of their relation to the Riemann zeta function, ζ(2n)=∞summationdisplay p=1p−2n. (5.157) Bernoulli Polynomials If Eq.(5.144) isgeneralizedslightly,wehave xexs ex−1=∞summationdisplay n=0Bn(s)xn n!(5.158) 380 Chapter 5 Infinite Series Table 5.2 BernoulliPolynomials B0=1 B1=x−1 2 B2=x2−x+1 6 B3=x3−3 2x2+1 2x B4=x4−2x3+x2−1 30 B5=x5−5 2x4+5 3x3−1 6x B6=x6−3x5+5 2x4−1 2x2+1 42 Bn(0)=Bn, Bernoulli number defining the Bernoulli polynomials ,Bn(s). The first seven Bernoulli polynomials are giveninTable5.2. Fromthegeneratingfunction,Eq.(5.158), Bn(0)=Bn,n=0,1,2,..., (5.159) the Bernoulli polynomial evaluated at zero equals the corresponding Bernoulli number. Two particularly important properties of the Bernoulli polynomials follow from the defin- ingrelation,Eq, (5.158):adifferentiationrelation d dsBn(s)=nBn−1(s), n=1,2,3,..., (5.160) andasymmetryrelation(replace x→−xinEq. (5.158)andthenset s=1) Bn(1)=(−1)nBn(0), n=1,2,3,.... (5.161) TheserelationsareusedinthedevelopmentoftheEuler–Maclaurinintegrationformula. Euler–Maclaurin Integration Formula One use of the Bernoulli functions is in the derivation of the Euler–Maclaurin integration formula.This formulais usedinSection8.3 for thedevelopmentof anasymptoticexpres- sionforthefactorialfunction—Stirling’sseries. The technique is repeated integration by parts, using Eq. (5.160) to create new deriva- tives.Westart with integraldisplay1 0f(x)dx=integraldisplay1 0f(x)B0(x)dx. (5.162) FromEq. (5.160)andExercise5.9.2, B′ 1(x)=B0(x)=1. (5.163) 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 381 Substituting B′ 1(x)intoEq. (5.162)andintegratingbyparts, weobtain integraldisplay1 0f(x)dx=f(1)B1(1)−f(0)B1(0)−integraldisplay1 0f′(x)B1(x)dx =1 2bracketleftbig f(1)+f(0)bracketrightbig −integraldisplay1 0f′(x)B1(x)dx. (5.164) AgainusingEq.(5.160), wehave B1(x)=1 2B′ 2(x), (5.165) andintegratingbyparts weget integraldisplay1 0f(x)dx=1 2bracketleftbig f(1)+f(0)bracketrightbig −1 2!bracketleftbig f′(1)B2(1)−f′(0)B2(0)bracketrightbig +1 2!integraldisplay1 0f(2)(x)B2(x)dx. (5.166) Usingtherelations B2n(1)=B2n(0)=B2n,n=0,1,2,... (5.167) B2n+1(1)=B2n+1(0)=0,n=1,2,3,... andcontinuingthis process,wehave integraldisplay1 0f(x)dx=1 2bracketleftbig f(1)+f(0)bracketrightbig −qsummationdisplay p=11 (2p)!B2pbracketleftbig f(2p−1)(1)−f(2p−1)(0)bracketrightbig +1 (2q)!integraldisplay1 0f(2q)(x)B2q(x)dx. (5.168a) Thisis theEuler–Maclaurinintegrationformula.It assumesthatthefunction f(x)has the requiredderivatives. TherangeofintegrationinEq.(5.168a)maybeshiftedfrom [0,1]to[1,2]byreplacing f(x)byf(x+1).Addingsuchresultsupto [n−1,n],weobtain integraldisplayn 0f(x)dx=1 2f(0)+f(1)+f(2)+···+f(n−1)+1 2f(n) −qsummationdisplay p=11 (2p)!B2pbracketleftbig f(2p−1)(n)−f(2p−1)(0)bracketrightbig +1 (2q)!integraldisplay1 0B2q(x)n−1summationdisplay ν=0f(2q)(x+ν)dx. (5.168b) The terms1 2f(0)+f(1)+···+1 2f(n)appear exactly as in trapezoidal integration, or quadrature. The summation over pmay be interpreted as a correction to the trapezoidal approximation. Equation (5.168b) may be seen as a generalization of Eq. (5.22); it is the 382 Chapter 5 Infinite Series Table 5.3 RiemannZetaFunction sζ (s) 21 .6449340668 31 .2020569032 41 .0823232337 51 .0369277551 61 .0173430620 71 .0083492774 81 .0040773562 91 .0020083928 10 1 .0009945751 formusedinExercise5.9.5forsummingpositivepowersofintegersandinSection8.3for thederivationof Stirling’sformula. The Euler–Maclaurin formula is often useful in summing series by converting them to integrals.21 Riemann Zeta Function This series,summationtext∞ p=1p−2n, was used as a comparison series for testing convergence (Sec- tion 5.2) and in Eq. (5.152) as one definition of the Bernoulli numbers, B2n.I ta l s os e r v e s todefinetheRiemannzetafunctionby ζ(s)≡∞summationdisplay n=1n−s,s>1. (5.169) Table 5.3 lists the values of ζ(s)for integral s,s=2,3,...,10. Closed forms for even s appear in Exercise 5.9.6. Figure 5.11 is a plot of ζ(s)−1. An integral expression for this RiemannzetafunctionappearsinExercise8.2.21aspartofthedevelopmentofthegamma function,andthefunctionalrelationis giveninSection14.3. The celebrated Euler prime number product for the Riemann zeta function may be de- rivedas ζ(s)parenleftbig 1−2−sparenrightbig =1+1 2s+1 3s+···−parenleftbigg1 2s+1 4s+1 6s+···parenrightbigg ; (5.170) eliminatingallthe n−s,wherenisa multipleof 2.Then ζ(s)parenleftbig 1−2−sparenrightbigparenleftbig 1−3−sparenrightbig =1+1 3s+1 5s+1 7s+1 9s+··· −parenleftbigg1 3s+1 9s+1 15s+···parenrightbigg ; (5.171) 21SeeR. P.Boas and C.Stutz, Estimating sums with integrals. Am.J .Ph ys. 39: 745 (1971), for a number of examples. 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 383 FIGURE 5.11Riemannzetafunction, ζ(s)−1 versuss. eliminating all the remaining terms in which nis a multiple of 3. Continuing, we have ζ(s)(1−2−s)(1−3−s)(1−5−s)···(1−P−s),wherePisaprimenumber,andallterms n−s,inwhich nis amultipleofanyintegerupthrough P, arecanceledout.As P→∞, ζ(s)parenleftbig 1−2−sparenrightbigparenleftbig 1−3−sparenrightbig ···parenleftbig 1−P−sparenrightbig →ζ(s)∞productdisplay P(prime)=2parenleftbig 1−P−sparenrightbig =1.(5.172) Therefore ζ(s)=∞productdisplay P(prime)=2parenleftbig 1−P−sparenrightbig−1, (5.173) givingζ(s)as aninfiniteproduct.22 This cancellation procedure has a clear application in numerical computation. Equa- tion (5.170) will give ζ(s)(1−2−s)to the same accuracy as Eq. (5.169) gives ζ(s),b u t 22ThisisthestartingpointfortheextensiveapplicationsoftheRiemannzetafunctiontoanalyticnumbertheory.SeeH.M.Ed- wards,Riemann’s Zeta Function . New York: Academic Press (1974); A. Ivi ´c,The Riemann Zeta Function . New York: Wiley (1985); S. J. Patterson, Introduction to the Theory of the Riemann Zeta Function . Cambridge, UK: Cambridge University Press (1988). 384 Chapter 5 Infinite Series withonlyhalfasmanyterms.(Ineithercase,acorrectionwouldbemadefortheneglected tailoftheseriesbytheMaclaurinintegraltesttechnique—replacingtheseriesbyaninte- gral,Section5.2.) Along with the Riemann zeta function, AMS-55 (Chapter 23. See Exercise 5.2.22 for thereference.) definesthreeotherDirichletseriesrelatedto ζ(s): η(s)=∞summationdisplay n=1(−1)n−1n−s=parenleftbig 1−21−sparenrightbig ζ(s), λ(s)=∞summationdisplay n=0(2n+1)−s=parenleftbig 1−2−sparenrightbig ζ(s), and β(s)=∞summationdisplay n=0(−1)n(2n+1)−s. From the Bernoulli numbers (Exercise 5.9.6) or Fourier series (Example 14.3.3 and Exer- cise14.3.13)specialvaluesare ζ(2)=1+1 22+1 32+···=π2 6 ζ(4)=1+1 24+1 34+···=π4 90 η(2)=1−1 22+1 32+···=π2 12 η(4)=1−1 24+1 34+···=7π4 720 λ(2)=1+1 32+1 52+···=π2 8 λ(4)=1+1 34+1 54+···=π4 96 β(1)=1−1 3+1 5−···=π 4 β(3)=1−1 33+1 53−···=π3 32. Catalan’sconstant, β(2)=1−1 32+1 52−···=0.91596559 ..., isthetopicof Exercise5.2.22. 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 385 Improvement of Convergence If we are required to sum a convergent seriessummationtext∞ n=1anwhose terms are rational functions ofn, the convergence may be improved dramatically by introducing the Riemann zeta function. Example 5.9.1 IMPROVEMENT OF CONVERGENCE The problem is to evaluate the seriessummationtext∞ n=11/(1+n2). Expanding (1+n2)−1= n−2(1+n−2)−1bydirectdivision,wehave parenleftbig 1+n2parenrightbig−1=n−2parenleftbigg 1−n−2+n−4−n−6 1+n−2parenrightbigg =1 n2−1 n4+1 n6−1 n8+n6. Therefore ∞summationdisplay n=11 1+n2=ζ(2)−ζ(4)+ζ(6)−∞summationdisplay n=11 n8+n6. Theζvaluesaretabulatedandtheremainderseriesconvergesas n−8.Clearly,theprocess canbecontinuedasdesired.Youmakeachoicebetweenhowmuchalgebrayouwilldoand how much arithmetic the computer will do. Other methods for improving computational effectivenessaregivenattheendofSections5.2and5.4. /squaresolid Exercises 5.9.1 Showthat tanx=∞summationdisplay n=1(−1)n−122n(22n−1)B2n (2n)!x2n−1,−π 2<x<π 2. Hint.t a nx=cotx−2cot2x. 5.9.2 ShowthatthefirstBernoullipolynomialsare B0(s)=1 B1(s)=s−1 2 B2(s)=s2−s+1 6. Notethat Bn(0)=Bn, theBernoullinumber. 5.9.3 Showthat B′ n(s)=nBn−1(s),n=1,2,3,.... Hint.DifferentiateEq. (5.158). 386 Chapter 5 Infinite Series 5.9.4 Showthat Bn(1)=(−1)nBn(0). Hint.Gobacktothegeneratingfunction,Eq.(5.158), orExercise5.9.2. 5.9.5 TheEuler–Maclaurinintegrationformulamaybeusedfortheevaluationoffiniteseries: nsummationdisplay m=1f(m)=integraldisplayn 0f(x)dx+1 2f(1)+1 2f(n)+B2 2!bracketleftbig f′(n)−f′(1)bracketrightbig +···. Showthat (a)nsummationdisplay m=1m=1 2n(n+1). (b)nsummationdisplay m=1m2=1 6n(n+1)(2n+1). (c)nsummationdisplay m=1m3=1 4n2(n+1)2. (d)nsummationdisplay m=1m4=1 30n(n+1)(2n+1)parenleftbig 3n2+3n−1parenrightbig . 5.9.6 From B2n=(−1)n−12(2n)! (2π)2nζ(2n), showthat (a)ζ(2)=π2 6(d)ζ(8)=π8 9450 (b)ζ(4)=π4 90(e)ζ(10)=π10 93,555. (c)ζ(6)=π6 945 5.9.7 Planck’sblackbodyradiationlawinvolvestheintegral integraldisplay∞ 0x3dx ex−1. Showthatthisequals 6 ζ(4). From Exercise5.9.6, ζ(4)=π4 90. Hint.Makeuseofthegammafunction,Chapter8. 5.9 Bernoulli Numbers,Euler–Maclaurin Formula 387 5.9.8 Provethatintegraldisplay∞ 0xnexdx (ex−1)2=n!ζ(n). Assuming nto be real, show that each side of the equation diverges if n=1. Hence the preceding equation carries the condition n>1. Integrals such as this appear in the quantumtheoryoftransporteffects—thermalandelectricalconductivity. 5.9.9 TheBloch–Gruneissenapproximationfor theresistanceinamonovalentmetalis ρ=CT5 /Theta16integraldisplay/Theta1/T 0x5dx (ex−1)(1−e−x), where/Theta1istheDebyetemperaturecharacteristicofthemetal. (a) For T→∞,showthat ρ≈C 4·T /Theta12. (b) For T→0,showthat ρ≈5!ζ(5)CT5 /Theta16. 5.9.10 Showthat (a)integraldisplay1 0ln(1+x) xdx=1 2ζ(2), (b) lim a→1integraldisplaya 0ln(1−x) xdx=ζ(2). FromExercise5.9.6, ζ(2)=π2/6.Notethattheintegrandinpart(b)divergesfor a=1 butthatthe integrated seriesis convergent. 5.9.11 Theintegralintegraldisplay1 0bracketleftbig ln(1−x)bracketrightbig2dx x appears in the fourth-order correction to the magnetic moment of the electron. Show thatit equals 2 ζ(3). Hint.Let1−x=e−t. 5.9.12 Showthat integraldisplay∞ 0(lnz)2 1+z2dz=4parenleftbigg 1−1 33+1 53−1 73+···parenrightbigg . Bycontourintegration(Exercise7.1.17),this maybeshownequalto π3/8. 5.9.13 For“small”valuesof x, ln(x!)=−γx+∞summationdisplay n=2(−1)nζ(n) nxn, whereγis the Euler–Mascheroni constant and ζ(n)is the Riemann zeta function. For whatvaluesof xdoesthisseriesconverge? ANS.−1<x≤1. 388 Chapter 5 Infinite Series Notethatif x=1,weobtain γ=∞summationdisplay n=2(−1)nζ(n) n, a series for the Euler–Mascheroni constant. The convergence of this series is exceed- inglyslow.Foractualcomputationof γ,other,indirectapproachesarefarsuperior(see Exercises5.10.11,and8.5.16). 5.9.14 Showthattheseriesexpansionof ln (x!)(Exercise5.9.13) maybewrittenas (a) ln(x!)=1 2lnparenleftbiggπx sinπxparenrightbigg −γx−∞summationdisplay n=1ζ(2n+1) 2n+1x2n+1, (b) ln(x!)=1 2lnparenleftbiggπx sinπxparenrightbigg −1 2lnparenleftbigg1+x 1−xparenrightbigg +(1−γ)x −∞summationdisplay n=1bracketleftbig ζ(2n+1)−1bracketrightbigx2n+1 2n+1. Determinetherangeofconvergenceofeachoftheseexpressions. 5.9.15 ShowthatCatalan’sconstant, β(2), maybewrittenas β(2)=2∞summationdisplay k=1(4k−3)−2−π2 8. Hint.π2=6ζ(2). 5.9.16 DerivethefollowingexpansionsoftheDebyefunctionsfor n≥1: integraldisplayx 0tndt et−1=xnbracketleftbigg1 n−x 2(n+1)+∞summationdisplay k=1B2kx2k (2k+n)(2k)!bracketrightbigg ,|x|<2π; integraldisplay∞ xtndt et−1=∞summationdisplay k=1e−kxbracketleftbiggxn k+nxn−1 k2+n(n−1)xn−2 k3+···+n! kn+1bracketrightbigg forx>0.Thecompleteintegral (0,∞)equalsn!ζ(n+1), Exercise8.2.15. 5.9.17 (a) Showthattheequationln2 =summationtext∞ s=1(−1)s+1s−1(Exercise5.4.1)mayberewritten as ln2=∞summationdisplay s=22−sζ(s)+∞summationdisplay p=1(2p)−n−1bracketleftbigg 1−1 2pbracketrightbigg−1 . Hint.Takethetermsinpairs. (b) Calculate ln2 tosixsignificantfigures. 5.10 Asymptotic Series 389 5.9.18 (a) Show that the equation π/4=summationtext∞ n=1(−1)n+1(2n−1)−1(Exercise 5.7.6) may be rewrittenas π 4=1−2∞summationdisplay s=14−2sζ(2s)−2∞summationdisplay p=1(4p)−2n−2bracketleftbigg 1−1 (4p)2bracketrightbigg−1 . (b) Calculate π/4 tosixsignificantfigures. 5.9.19 Write a function subprogram ZETA (N)that will calculate the Riemann zeta function for integer argument. Tabulate ζ(s)fors=2,3,4,...,20. Check your values against Table5.3andAMS-55,Chapter23.(SeeExercise5.2.22for thereference.). Hint. If you supply the function subprogram with the known values of ζ(2),ζ(3), and ζ(4), you avoid the more slowly converging series. Calculation time may be further shortenedbyusingEq.(5.170). 5.9.20 Calculatethelogarithm(base 10)of |B2n|,n=10,20,...,100. Hint.Program ζ(n)asafunctionsubprogram,Exercise5.9.19. Checkvalues. log|B100|=78.45 log|B200|=215.56. 5.10 A SYMPTOTIC SERIES Asymptotic series frequently occur in physics. In numerical computations they are em- ployed for the accurate computation of a variety of functions. We consider here two types ofintegralsthatleadtoasymptoticseries:first, anintegralof theform I1(x)=integraldisplay∞ xe−uf(u)du, where the variable xappears as the lower limit of an integral. Second, we consider the form I2(x)=integraldisplay∞ 0e−ufparenleftbiggu xparenrightbigg du, with the function fto be expanded as a Taylor series (binomial series). Asymptotic se- ries often occur as solutions of differential equations. An example of this appears in Sec- tion11.6as asolutionofBessel’s equation. Incomplete Gamma Function The nature of an asymptotic series is perhaps best illustrated by a specific example. Sup- posethatwehavetheexponentialintegralfunction23 Ei(x)=integraldisplayx −∞eu udu, (5.174) 23This function occurs frequently in astrophysical problems involving gas with aMaxwell–Boltzmann energy distribution. 390 Chapter 5 Infinite Series or −Ei(−x)=integraldisplay∞ xe−u udu=E1(x), (5.175) to be evaluated for large values of x. Or let us take a generalization of the incomplete factorialfunction(incompletegammafunction),24 I(x,p)=integraldisplay∞ xe−uu−pdu=Ŵ(1−p,x), (5.176) inwhich xandparepositive.Again,weseektoevaluateit forlargevaluesof x. Integratingbyparts,weobtain I(x,p)=e−x xp−pintegraldisplay∞ xe−uu−p−1du =e−x xp−pe−x xp+1+p(p+1)integraldisplay∞ xe−uu−p−2du. (5.177) Continuingtointegratebyparts,wedeveloptheseries I(x,p)=e−xparenleftbigg1 xp−p xp+1+p(p+1) xp+2−···+(−1)n−1(p+n−2)! (p−1)!xp+n−1parenrightbigg +(−1)n(p+n−1)! (p−1)!integraldisplay∞ xe−uu−p−ndu. (5.178) This is a remarkable series. Checking the convergence by the d’Alembert ratio test, we find limn→∞|un+1| |un|=limn→∞(p+n)! (p+n−1)!·1 x=limn→∞p+n x=∞ (5.179) for all finite values of x. Therefore our series as an infinite series diverges everywhere! BeforediscardingEq.(5.178)asworthless,letusseehowwellagivenpartialsumapprox- imatestheincompletefactorialfunction, I(x,p): I(x,p)−sn(x,p)=(−1)n+1(p+n)! (p−1)!integraldisplay∞ xe−uu−p−n−1du=Rn(x,p). (5.180) Inabsolutevalue vextendsinglevextendsingleI(x,p)−sn(x,p)vextendsinglevextendsingle≤(p+n)! (p−1)!integraldisplay∞ xe−uu−p−n−1du. Whenwesubstitute u=v+x, theintegralbecomes integraldisplay∞ xe−uu−p−n−1du=e−xintegraldisplay∞ 0e−v(v+x)−p−n−1dv =e−x xp+n+1integraldisplay∞ 0e−vparenleftbigg 1+v xparenrightbigg−p−n−1 dv. 24SeealsoSection8.5. 5.10 Asymptotic Series 391 FIGURE 5.12Partialsumsof exE1(x)|x=5. Forlarge xthefinalintegralapproaches1and vextendsinglevextendsingleI(x,p)−sn(x,p)vextendsinglevextendsingle≈(p+n)! (p−1)!·e−x xp+n+1. (5.181) Thismeansthatifwetake xlargeenough,ourpartialsum snisanarbitrarilygoodapprox- imation to the function I(x,p). Our divergent series (Eq. (5.178)) therefore is perfectly good for computations of partial sums. For this reason it is sometimes called a semicon- vergentseries. Note that the power of xin the denominator of the remainder (p+n+1) ishigherthanthepowerof xinthelasttermincludedin sn(x,p),(p+n). Since the remainder Rn(x,p)alternates in sign, the successive partial sums give alter- nately upper and lower bounds for I(x,p). The behavior of the series (with p=1) as a functionofthenumberof termsincludedis showninFig. 5.12.Wehave exE1(x)=exintegraldisplay∞ xe−u udu ∼=sn(x)=1 x−1! x2+2! x2−3! x4+···+(−1)nn! xn+1,(5.182) whichisevaluatedat x=5.Theoptimumdeterminationof exE1(x)isgivenbytheclosest approachoftheupperandlowerbounds,thatis,between s4=s6=0.1664and s5=0.1741 forx=5.Therefore 0.1664≤exE1(x)vextendsinglevextendsingle x=5≤0.1741. (5.183) Actually,fromtables, exE1(x)vextendsinglevextendsingle x=5=0.1704, (5.184) 392 Chapter 5 Infinite Series withinthelimitsestablishedbyourasymptoticexpansion.Notethatinclusionofadditional terms in the series expansion beyond the optimum point literally reduces the accuracy of the representation. As xis increased, the spread between the lowest upper bound and the highest lower bound will diminish. By taking xlarge enough, one may compute exE1(x) to any desired degree of accuracy. Other properties of E1(x)are derived and discussed in Section8.5. Cosine and Sine Integrals Asymptoticseriesmayalsobedevelopedfromdefiniteintegrals—iftheintegrandhasthe required behavior. As an example, the cosine and sine integrals (Section 8.5) are defined by Ci(x)=−integraldisplay∞ xcost tdt, (5.185) si(x)=−integraldisplay∞ xsint tdt. (5.186) Combiningthesewithregulartrigonometricfunctions,wemaydefine f(x)=Ci(x)sinx−si(x)cosx=integraldisplay∞ 0siny y+xdy, g(x)=−Ci(x)cosx−si(x)sinx=integraldisplay∞ 0cosy y+xdy,(5.187) withthenewvariable y=t−x.Goingtocomplexvariables,Section6.1, wehave g(x)+if(x)=integraldisplay∞ 0eiy y+xdy=integraldisplay∞ 0ie−xu 1+iudu, (5.188) in which u=−iy/x. The limits of integration, 0 to ∞, rather than 0 to −i∞, may be justified by Cauchy’s theorem, Section 6.3. Rationalizing the denominator and equating realparttorealpartandimaginaryparttoimaginarypart,weobtain g(x)=integraldisplay∞ 0ue−xu 1+u2du, f(x)=integraldisplay∞ 0e−xu 1+u2du. (5.189) Forconvergenceof theintegralswemustrequirethat ℜ(x)>0.25 25ℜ(x)=real part of (complex) x(compare Section 6.1). 5.10 Asymptotic Series 393 Now, to developthe asymptotic expansions,let v=xuand expandthe preceding factor [1+(v/x)2]−1bythebinomialtheorem.26We have f(x)≈1 xintegraldisplay∞ 0e−vsummationdisplay 0≤n≤N(−1)nv2n x2ndv=1 xsummationdisplay 0≤n≤N(−1)n(2n)! x2n, (5.190) g(x)≈1 x2integraldisplay∞ 0e−vsummationdisplay 0≤n≤N(−1)nv2n+1 x2ndv=1 x2summationdisplay 0≤n≤N(−1)n(2n+1)! x2n. FromEqs. (5.187) and(5.190), Ci(x)≈sinx xsummationdisplay 0≤n≤N(−1)n(2n)! x2n−cosx x2summationdisplay 0≤n≤N(−1)n(2n+1)! x2n, si(x)≈−cosx xsummationdisplay 0≤n≤N(−1)n(2n)! x2n−sinx x2summationdisplay 0≤n≤N(−1)n(2n+1)! x2n(5.191) arethedesiredasymptoticexpansions. This technique of expanding the integrand of a definite integral and integrating term by term is applied in Section 11.6 to develop an asymptotic expansion of the modified Besselfunction KνandinSection13.5forexpansionsofthetwoconfluenthypergeometric functions M(a,c;x)andU(a,c;x). Definition of Asymptotic Series The behavior of these series (Eqs. (5.178) and (5.191)), is consistent with the defining propertiesof anasymptoticseries.27FollowingPoincaré,wetake28 xnRn(x)=xnbracketleftbig f(x)−sn(x)bracketrightbig , (5.192) where sn(x)=a0+a1 x+a2 x2+···+an xn. (5.193) Theasymptoticexpansionof f(x)hasthepropertiesthat limx→∞xnRn(x)=0,for fixedn, (5.194) and limn→∞xnRn(x)=∞,for fixedx.29(5.195) 26This stepis validfor v≤x. Thecontributions from v≥xwill benegligible (for large x)becauseofthe negative exponential. It is becausethe binomial expansion does not converge for v≥xthat our final series is asymptotic ratherthanconvergent. 27Itisnotnecessarythattheasymptoticseriesbeapowerseries.Therequiredpropertyisthattheremainder Rn(x)beofhigher order than the last term kept—as in Eq.(5.194). 28Poincaré’s definition allows (or neglects) exponentially decreasing functions. The refinement of Poincaré’s definition is of considerable importance for the advanced theory of asymptotic expansions, particularly for extensions into the complex plane. However,forpurposesofanintroductorytreatmentandespeciallyfornumericalcomputationwith xrealandpositive,Poincaré’s approachis perfectlysatisfactory. 394 Chapter 5 Infinite Series See Eqs. (5.178) and (5.179) for an example of these properties. For power series, as as- sumedintheformof sn(x),Rn(x)∼x−n−1.Withconditions(5.194)and(5.195)satisfied, wewrite f(x)≈∞summationdisplay n=0anx−n. (5.196) Note the use of ≈in place of=. The function f(x)is equal to the series only in the limit asx→∞andafinitenumberoftermsintheseries. Asymptotic expansions of two functions may be multiplied together, and the result will beanasymptoticexpansionof theproductofthetwofunctions. Theasymptoticexpansionofagivenfunction f(t)maybeintegratedtermbyterm(just asinauniformlyconvergentseriesofcontinuousfunctions)from x≤t<∞,andtheresult will be an asymptotic expansion ofintegraltext∞ xf(t)dt. Term-by-term differentiation, however, is validonlyunderveryspecialconditions. Some functions do not possess an asymptotic expansion; exis an example of such a function. However, if a function has an asymptotic expansion, it has only one. The corre- spondenceis notonetoone;manyfunctionsmayhavethesameasymptoticexpansion. One of the most useful and powerful methods of generating asymptotic expansions, the method of steepest descents, will be developed in Section 7.3. Applications include the derivation of Stirling’s formula for the (complete) factorial function (Section 8.3) and the asymptotic forms of the various Bessel functions (Section 11.6). Asymptotic series occur fairlyofteninmathematicalphysics.Oneoftheearliestandstillimportantapproximations ofquantummechanics,the WKBexpansion,isanasymptoticseries. Exercises 5.10.1 Stirling’sformulafor thelogarithmofthefactorialfunctionis ln(x!)=1 2ln2π+parenleftbigg x+1 2parenrightbigg lnx−x−Nsummationdisplay n=1B2n 2n(2n−1)x1−2n. TheB2nare the Bernoulli numbers (Section 5.9). Show that Stirling’s formula is an asymptotic expansion. 5.10.2 Integratingbyparts, developasymptoticexpansionsoftheFresnelintegrals. (a)C(x)=integraldisplayx 0cosπu2 2du,(b)s(x)=integraldisplayx 0sinπu2 2du. Theseintegralsappearintheanalysisofaknife-edgediffractionpattern. 5.10.3 RederivetheasymptoticexpansionsofCi (x)andsi(x)byrepeatedintegrationbyparts. Hint.Ci(x)+isi(x)=−integraltext∞ xeit tdt. 29This excludes convergent series ofinverse powers of x. Some writers feelthat this exclusion is artificialandunnecessary. 5.10 Asymptotic Series 395 5.10.4 DerivetheasymptoticexpansionoftheGausserror function erf(x)=2√πintegraldisplayx 0e−t2dt ≈1−e−x2 √πxparenleftbigg 1−1 2x2+1·3 22x4−1·3·5 23x6+···+(−1)n(2n−1)!! 2nx2nparenrightbigg . Hint:e r f(x)=1−erfc(x)=1−2√πintegraltext∞ xe−t2dt. Normalized so that erf (∞)=1, this function plays an important role in probability theory. It may be expressed in terms of the Fresnel integrals (Exercise 5.10.2), the in- complete gamma functions (Section 8.5), and the confluent hypergeometric functions (Section13.5). 5.10.5 The asymptotic expressions for the various Bessel functions, Section 11.6, contain the series Pν(z)∼1+∞summationdisplay n=1(−1)nproducttext2n s=1[4ν2−(2s−1)2] (2n)!(8z)2n, Qν(z)∼∞summationdisplay n=1(−1)n+1producttext2n−1 s=1[4ν2−(2s−1)2] (2n−1)!(8z)2n−1. Showthatthesetwoseriesareindeedasymptoticseries. 5.10.6 Forx>1, 1 1+x=∞summationdisplay n=0(−1)n1 xn+1. Test thisseriestoseeif itis anasymptoticseries. 5.10.7 DerivethefollowingBernoullinumberasymptoticseriesfortheEuler–Mascheronicon- stant: γ=nsummationdisplay s=1s−1−lnn−1 2n+Nsummationdisplay k=1B2k (2k)n2k. Hint. Apply the Euler–Maclaurin integration formula to f(x)=x−1over the interval [1,n]forN=1,2,.... 5.10.8 Developanasymptoticseriesfor integraldisplay∞ 0e−xvparenleftbig 1+v2parenrightbig−2dv. Takextoberealandpositive. ANS.1 x−2! x3+4! x5−···+(−1)n(2n)! x2n+1. 396 Chapter 5 Infinite Series 5.10.9 Calculatepartialsumsof exE1(x)forx=5,10,and15toexhibitthebehaviorshownin Fig.5.11.Determinethewidthofthethroatfor x=10and15,analogoustoEq.(5.183). ANS.Throatwidth: n=10,0.000051 n=15,0.0000002. 5.10.10 Theknife-edgediffractionpatternis describedby I=0.5I0braceleftbigbracketleftbig C(u0)+0.5bracketrightbig2+bracketleftbig S(u0)+0.5bracketrightbig2bracerightbig , whereC(u0)andS(u0)are the Fresnel integrals of Exercise 5.10.2. Here I0is the incident intensity and Iis the diffracted intensity; u0is proportional to the distance away from the knife edge (measured at right angles to the incident beam). Calculate I/I0foru0varying from−1.0t o+4.0 in steps of 0.1. Tabulate your results and, if a plottingroutineis available,plotthem. Checkvalue .u0=1.0,I/I0=1.259226. 5.10.11 The Euler–Maclaurin integration formula of Section 5.9 provides a way of calculating the Euler–Mascheroni constant γto high accuracy. Using f(x)=1/xin Eq. (5.168b) (withinterval[1,n])andthedefinitionof γ(Eq. 5.28),weobtain γ=nsummationdisplay s=1s−1−lnn−1 2n+Nsummationdisplay k=1B2k (2k)n2k. Usingdouble-precisionarithmetic,calculate γforN=1,2,.... Note. D. E. Knuth, Euler’s constant to 1271 places. Math. Comput. 16: 275 (1962). An evenmoreprecisecalculationappearsinExercise8.5.16. ANS.For n=1000,N=2 γ=0.577215664901. 5.11 I NFINITE PRODUCTS Consider a succession of positive factors f1·f2·f3·f4···fn(fi>0). Using capital pi (producttext) toindicateproduct,as capitalsigma (summationtext)indicatesasum,wehave f1·f2·f3···fn=nproductdisplay i=1fi. (5.197) Wedefine pn, apartialproduct,inanalogywith snthepartialsum, pn=nproductdisplay i=1fi (5.198) andtheninvestigatethelimit, limn→∞pn=P. (5.199) IfPis finite (but not zero), we say the infinite product is convergent. If Pis infinite or zero,theinfiniteproductislabeleddivergent. 5.11 Infinite Products 397 Sincetheproductwilldivergetoinfinityif limn→∞fn>1 (5.200) ortozerofor limn→∞fn<1(and>0), (5.201) itisconvenienttowriteourinfiniteproductsas ∞productdisplay n=1(1+an). Thecondition an→0 isthenanecessary(butnotsufficient)conditionforconvergence. Theinfiniteproductmayberelatedtoaninfiniteseriesbytheobviousmethodoftaking thelogarithm, ln∞productdisplay n=1(1+an)=∞summationdisplay n=1ln(1+an). (5.202) Amoreuseful relationshipis statedbythefollowingtheorem. Convergence of Infinite Product If 0≤an<1,theinfiniteproductsproducttext∞ n=1(1+an)andproducttext∞n=1(1−an)convergeifsummationtext∞n=1an convergesanddivergeifsummationtext∞ n=1andiverges. Consideringtheterm 1 +an,weseefromEq. (5.90) that 1+an≤ean. (5.203) Thereforefor thepartialproduct pn, withsnthepartialsumofthe ai, pn≤esn, (5.204) andletting n→∞, ∞productdisplay n=1(1+an)≤exp∞summationdisplay n=1an, (5.205) thusestablishinganupperboundfor theinfiniteproduct. Todevelopalowerbound,wenotethat pn=1+nsummationdisplay i=1ai+nsummationdisplay i=1nsummationdisplay j=1aiaj+···≥sn, (5.206) sinceai≥0.Hence ∞productdisplay n=1(1+an)≥∞summationdisplay n=1an. (5.207) 398 Chapter 5 Infinite Series Iftheinfinitesumremainsfinite,theinfiniteproductwillalso.Iftheinfinitesumdiverges, sowilltheinfiniteproduct. The case ofproducttext(1−an)is complicated by the negative signs, but a proof that depends ontheforegoingproofmaybedevelopedbynotingthatfor an<1 2(remember an→0f o r convergence), (1−an)≤(1+an)−1 and (1−an)≥(1+2an)−1. (5.208) Sine, Cosine, and Gamma Functions Annth-order polynomial Pn(x)withnreal roots may be written as a product of nfactors (seeSection6.4, Gauss’fundamentaltheoremof algebra): Pn(x)=(x−x1)(x−x2)···(x−xn)=nproductdisplay i=1(x−xi). (5.209) In much the same way we may expect that a function with an infinite number of roots may be written as an infinite product, one factor for each root. This is indeed the case for thetrigonometricfunctions.Wehavetwoveryusefulinfiniteproductrepresentations, sinx=x∞productdisplay n=1parenleftbigg 1−x2 n2π2parenrightbigg , (5.210) cosx=∞productdisplay n=1bracketleftbigg 1−4x2 (2n−1)2π2bracketrightbigg . (5.211) The most convenient and perhaps most elegant derivation of these two expressions is by the use of complex variables.30By our theorem of convergence, Eqs. (5.210) and (5.211) areconvergentforallfinitevaluesof x.Specifically,fortheinfiniteproductforsin x, an= x2/n2π2, ∞summationdisplay n=1an=x2 π2∞summationdisplay n=1n−2=x2 π2ζ(2)=x2 6(5.212) byExercise5.9.6. Theseries correspondingtoEq. (5.211)behavesinasimilarmanner. Equation(5.210)leadstotwointerestingresults. First, if weset x=π/2,weobtain 1=π 2∞productdisplay n=1bracketleftbigg 1−1 (2n)2bracketrightbigg =π 2∞productdisplay n=1bracketleftbigg(2n)2−1 (2n)2bracketrightbigg . (5.213) 30SeeEqs. (7.25) and (7.26). 5.11 Infinite Products 399 Solvingfor π/2,wehave π 2=∞productdisplay n=1bracketleftbigg(2n)2 (2n−1)(2n+1)bracketrightbigg =2·2 1·3·4·4 3·5·6·6 5·7···, (5.214) whichisWallis’ famousformulafor π/2. Thesecondresultinvolvesthegammaorfactorialfunction(Section8.1).Onedefinition ofthegammafunctionis Ŵ(x)=bracketleftbigg xeγx∞productdisplay r=1parenleftbigg 1+x rparenrightbigg e−x/rbracketrightbigg−1 , (5.215) whereγis the usual Euler–Mascheroni constant (compare Section 5.2). If we take the productof Ŵ(x)andŴ(−x), Eq.(5.215) leadsto Ŵ(x)Ŵ(−x)=−bracketleftbigg xeγx∞productdisplay r=1parenleftbigg 1+x rparenrightbigg e−x/rxe−γx∞productdisplay r=1parenleftbigg 1−x rparenrightbigg ex/rbracketrightbigg−1 =−1 x2∞productdisplay r=1parenleftbigg 1−x2 r2parenrightbigg−1 . (5.216) UsingEq. (5.210)with xreplacedby πx, weobtain Ŵ(x)Ŵ(−x)=−π xsinπx. (5.217) AnticipatingarecurrencerelationdevelopedinSection8.1,wehave −xŴ(−x)=Ŵ(1−x). Equation(5.217)maybewrittenas Ŵ(x)Ŵ(1−x)=π sinπx. (5.218) Thiswillbeusefulintreatingthegammafunction(Chapter8). Strictly speaking, we should check the range of xfor which Eq. (5.215) is convergent. Clearly, individual factors will vanish for x=0,−1,−2,....The proof that the infinite productconvergesforallother(finite)valuesof xis leftasExercise5.11.9. Theseinfiniteproductshaveavarietyofusesinmathematics.However,becauseofrather slowconvergence,theyarenotsuitablefor precisenumericalwork inphysics. Exercises 5.11.1 Using ln∞productdisplay n=1(1±an)=∞summationdisplay n=1ln(1±an) andtheMaclaurinexpansionofln (1±an),showthattheinfiniteproductproducttext∞ n=1(1±an) convergesordivergeswiththeinfiniteseriessummationtext∞n=1an. 400 Chapter 5 Infinite Series 5.11.2 Aninfiniteproductappearsintheform ∞productdisplay n=1parenleftbigg1+a/n 1+b/nparenrightbigg , whereaandbareconstants.Showthatthisinfiniteproductconvergesonlyif a=b. 5.11.3 Show that the infinite product representations of sin xand cosxare consistent with the identity 2sin xcosx=sin2x. 5.11.4 Determinethelimittowhich ∞productdisplay n=2parenleftbigg 1+(−1)n nparenrightbigg converges. 5.11.5 Showthat ∞productdisplay n=2bracketleftbigg 1−2 n(n+1)bracketrightbigg =1 3. 5.11.6 Provethat ∞productdisplay n=2parenleftbigg 1−1 n2parenrightbigg =1 2. 5.11.7 Usingtheinfiniteproductrepresentationsof sin x,showthat xcotx=1−2∞summationdisplay m,n=1parenleftbiggx nπparenrightbigg2m , hencethattheBernoullinumber B2n=(−1)n−12(2n)! (2π)2nζ(2n). 5.11.8 VerifytheEuleridentity ∞productdisplay p=1parenleftbig 1+zpparenrightbig =∞productdisplay q=1parenleftbig 1−z2q−1parenrightbig−1,|z|<1. 5.11.9 Show thatproducttext∞ r=1(1+x/r)e−x/rconverges for all finite x(except for the zeros of 1+x/r). Hint.Writethe nthfactoras 1+an. 5.11.10 Calculate cos xfrom its infinite product representation, Eq. (5.211), using (a) 10, (b) 100, and (c) 1000 factors in the product. Calculate the absolute error. Note how slowly the partial products converge–making the infinite product quite unsuitable for precisenumericalwork. ANS.For1000factors, cos π=−1.00051. 5.11 Additional Readings 401 AdditionalReadings Thetopic of infinite series is treatedinmany texts on advancedcalculus. Bender, C. M., and S. Orszag, Advanced Mathematical Methods for Scientists and Engineers .N e wY o r k : McGraw-Hill(1978). Particularly recommended for methods of acceleratingconvergence. Davis, H. T., Tables of Higher Mathematical Functions . Bloomington, IN: Principia Press (1935). Volume II contains extensive information on Bernoulli numbers and polynomials. Dingle,R. B., Asymptotic Expansions: Their Derivation and Interpretation . NewYork: AcademicPress (1973). Galambos, J., Representations of RealNumbers by Infinite Series .Berlin: Springer (1976). Gradshteyn, I. S., and I. M. Ryzhik, Table of Integrals, Series and Products . Corrected and enlarged 6th edition prepared byAlan Jeffrey. NewYork: AcademicPress (2000). Hamming, R. W., Numerical Methods for Scientists and Engineers .Reprinted, NewYork: Dover (1987). Hansen,E., ATableofSeriesand Products. EnglewoodCliffs, NJ:Prentice-Hall(1975). Atremendous compila- tion of series and products. Hardy, G. H., Divergent Series. Oxford: Clarendon Press (1956), 2nd ed., Chelsea (1992). The standard, com- prehensive work on methods of treating divergent series. Hardy includes instructive accounts of the gradual development ofthe concepts ofconvergence anddivergence. Jeffrey, A., Handbook of Mathematical Formulas and Integrals . San Diego: AcademicPress (1995). Knopp, K., Theory and Application of Infinite Series. London: Blackie and Son (2nd ed.); New York: Hafner (1971). Reprinted: A. K. Peters Classics (1997). This is a thorough, comprehensive, and authoritative work that covers infinite series and products. Proofs of almost all of the statements not proved in Chapter 5 will be found in this book. Mangulis, V., Handbook of Series for Scientists and Engineers . New York: Academic Press (1965). A most convenientandusefulcollectionofseries.Includesalgebraicfunctions,Fourierseries,andseriesofthespecial functions: Bessel,Legendre, andso on. Olver, F. W. J., Asymptotics and Special Functions . New York: Academic Press (1974). A detailed, readable development of asymptotic theory. Considerable attention is paid to error bounds for use in computation. Rainville, E. D., Infinite Series . New York: Macmillan (1967). A readable and useful account of series constants and functions. Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering , 2nd ed. New York: McGraw-Hill (1966). A long Chapter 2 (101 pages) presents infinite series in a thorough but very readable form.Extensionstothesolutionsofdifferentialequations,tocomplexseries,andtoFourierseriesareincluded. This page intentionally left blank CHAPTER 6 FUNCTIONS OF A COMPLEX VARIABLE I ANALYTIC PROPERTIES ,MAPPING The imaginary numbersarea wonderful flight of God’sspirit; they arealmost anamphibian between being and not being. GOTTFRIED WILHELM VON LEIBNIZ, 1702 Weturnnowtoastudyoffunctionsofacomplexvariable.Inthisareawedevelopsome of the most powerful and widely useful tools in all of analysis. To indicate, at least partly, whycomplexvariablesareimportant,wementionbrieflyseveralareasof application. 1. Formanypairs offunctions uandv,bothuandvsatisfyLaplace’sequation, ∇2ψ=∂2ψ(x,y) ∂x2+∂2ψ(x,y) ∂y2=0. Henceeither uorvmaybeusedtodescribeatwo-dimensionalelectrostaticpotential.The otherfunction,whichgivesafamilyofcurvesorthogonaltothoseofthefirstfunction,may thenbeusedtodescribetheelectricfield E.Asimilarsituationholdsforthehydrodynamics ofanidealfluidinirrotationalmotion.Thefunction umightdescribethevelocitypotential, whereasthefunction vwouldthenbethestreamfunction. In many cases in which the functions uandvare unknown, mapping or transforming in the complex plane permits us to create a coordinate system tailored to the particular problem. 2. In Chapter 9 we shall see that the second-order differential equations of interest in physicsmaybesolvedbypowerseries.Thesamepowerseriesmaybeusedinthecomplex plane to replace xby the complex variable z. The dependence of the solution f(z)at a givenz0onthebehaviorof f(z)elsewheregivesusgreaterinsightintothebehaviorofour 403 404 Chapter 6 Functions of a Complex Variable I solution and a powerful tool (analytic continuation) for extending the region in which the solutionisvalid. 3.Thechangeofaparameter kfromrealtoimaginary, k→ik,transformstheHelmholtz equation into the diffusion equation. The same change transforms the Helmholtz equa- tionsolutions(BesselandsphericalBessel functions)intothediffusionequationsolutions (modifiedBesselandmodifiedsphericalBesselfunctions). 4. Integralsinthecomplexplanehavea widevarietyof usefulapplications: •Evaluatingdefiniteintegrals; •Invertingpowerseries; •Forminginfiniteproducts; •Obtainingsolutionsofdifferentialequationsforlargevaluesofthevariable(asymptotic solutions); •Investigatingthestabilityof potentiallyoscillatorysystems; •Invertingintegraltransforms. 5.Manyphysicalquantitiesthatwereoriginallyrealbecomecomplexasasimplephys- ical theory is made more general. The real index of refraction of light becomes a complex quantity when absorption is included. The real energy associated with an energy level be- comescomplexwhenthefinitelifetimeofthelevelis considered. 6.1 C OMPLEX ALGEBRA Acomplexnumberisnothingmorethananorderedpairoftworealnumbers, (a,b).Sim- ilarly,acomplexvariableis anorderedpairoftworealvariables,1 z≡(x,y). (6.1) The ordering is significant. In general (a,b)is not equal to (b,a)and(x,y)is not equal to(y,x). As usual, we continue writing a real number (x,0)simply as x, and we call i≡(0,1)theimaginaryunit. Allourcomplexvariableanalysiscanbedevelopedintermsoforderedpairsofnumbers (a,b),variables (x,y), andfunctions (u(x,y),v(x,y) ). We now define addition of complex numbers in terms of their Cartesian components as z1+z2=(x1,y1)+(x2,y2)=(x1+x2,y1+y2), (6.2a) that is, two-dimensional vector addition. In Chapter 1 the points in the xy-plane are identified with the two-dimensional displacement vector r=ˆxx+ˆyy. As a result, two- dimensional vector analogs can be developed for much of our complex analysis. Exer- cise6.1.2is onesimpleexample;Cauchy’stheorem,Section6.3,is another. Multiplication of complexnumbersisdefinedas z1z2=(x1,y1)·(x2,y2)=(x1x2−y1y2,x1y2+x2y1). (6.2b) 1This is precisely how acomputer does complex arithmetic. 6.1 Complex Algebra 405 UsingEq.(6.2b)weverifythat i2=(0,1)·(0,1)=(−1,0)=−1,sowecanalsoidentify i=√ −1, asusualandfurther rewriteEq.(6.1) as z=(x,y)=(x,0)+(0,y)=x+(0,1)·(y,0)=x+iy. (6.2c) Clearly, the iis not necessary here but it is convenient. It serves to keep pairs in order— somewhatliketheunitvectorsof Chapter1.2 Permanence of Algebraic Form All our elementary functions, ez,sinz, and so on, can be extended into the complex plane (compare Exercise 6.1.9). For instance, they can be defined by power-series expansions, suchas ez=1+z 1!+z2 2!+···=∞summationdisplay n=0zn n!(6.3) for the exponential. Such definitions agree with the real variable definitions along the real x-axis and extend the corresponding real functions into the complex plane. This result is oftencalled permanenceofthealgebraicform . Itisconvenienttoemployagraphicalrepresentationofthecomplexvariable.Byplotting x—therealpartof z—astheabscissaand y—theimaginarypartof z—astheordinate, we have the complex plane, or Argand plane, shown in Fig. 6.1. If we assign specific valuesto xandy,thenzcorrespondstoapoint (x,y)intheplane.Intermsoftheordering mentionedbefore,itisobviousthatthepoint (x,y)doesnotcoincidewiththepoint (y,x) exceptfor thespecialcaseof x=y.Further, fromFig.6.1wemaywrite x=rcosθ, y=rsinθ (6.4a) FIGURE 6.1Complex plane—Arganddiagram. 2Thealgebra of complex numbers, (a,b), is isomorphic with that of matrices of the form parenleftbiggab −baparenrightbigg (compare Exercise3.2.4). 406 Chapter 6 Functions of a Complex Variable I and z=r(cosθ+isinθ). (6.4b) Using a result that is suggested (but not rigorously proved)3by Section 5.6 and Exer- cise5.6.1,wehavetheusefulpolarrepresentation z=r(cosθ+isinθ)=reiθ. (6.4c) In order to prove this identity, we use i3=−i, i4=1,...in the Taylor expansion of the exponentialandtrigonometricfunctionsandseparateevenandoddpowersin eiθ=∞summationdisplay n=0(iθ)n n!=∞summationdisplay ν=0(iθ)2ν (2ν)!+∞summationdisplay ν=0(iθ)2ν+1 (2ν+1)! =∞summationdisplay ν=0(−1)νθ2ν (2ν)!+i∞summationdisplay ν=0(−1)νθ2ν+1 (2ν+1)!=cosθ+isinθ. Forthespecialvalues θ=π/2 andθ=π,weobtain eiπ/2=cosπ 2+isinπ 2=i, eiπ=cos(π)=−1, intriguingconnectionsbetween e,i,andπ.Moreover,theexponentialfunction eiθisperi- odicwithperiod 2 π,just like sin θand cosθ. Inthisrepresentation riscalledthe modulus ormagnitude ofz(r=|z|=(x2+y2)1/2) andtheangle θ(=tan−1(y/x))islabeledtheargumentor phaseofz.(Notethatthearctan function tan−1(y/x)hasinfinitelymanybranches.) Thechoiceofpolarrepresentation,Eq.(6.4c),orCartesianrepresentation,Eqs.(6.1)and (6.2c),isamatterofconvenience.Additionandsubtractionofcomplexvariablesareeasier in the Cartesian representation, Eq. (6.2a). Multiplication, division, powers, and roots are easiertohandleinpolarform, Eq. (6.4c). Analytically or graphically, using the vector analogy, we may show that the modulus of thesumoftwocomplexnumbersisnogreaterthanthesumofthemoduliandnolessthan thedifference, Exercise6.1.3, |z1|−|z2|≤|z1+z2|≤|z1|+|z2|. (6.5) Becauseof thevectoranalogy,theseare calledthe triangleinequalities. Using the polar form, Eq. (6.4c), we find that the magnitude of a product is the product ofthemagnitudes: |z1·z2|=|z1|·|z2|. (6.6) Also, arg(z1·z2)=argz1+argz2. (6.7) 3Strictly speaking, Chapter 5 was limited to real variables. The development of power-series expansions for complex functions is takenup in Section6.5(Laurent expansion). 6.1 Complex Algebra 407 FIGURE 6.2Thefunction w(z)=u(x,y)+iv(x,y)mapspointsinthe xy-plane intopointsinthe uv-plane. Fromourcomplexvariable zcomplexfunctions f(z)orw(z)maybeconstructed.These complexfunctionsmaythenberesolvedintorealandimaginaryparts, w(z)=u(x,y)+iv(x,y), (6.8) inwhichtheseparatefunctions u(x,y)andv(x,y)arepurereal.Forexample,if f(z)=z2, wehave f(z)=(x+iy)2=parenleftbig x2−y2parenrightbig +i2xy. Thereal part of a function f(z)will be labeled ℜf(z), whereas the imaginary part will belabeledℑf(z).I nE q .( 6 . 8 ) ℜw(z)=Re(w)=u(x,y),ℑw(z)=Im(w)=v(x,y). The relationship between the independent variable zand the dependent variable wis perhaps best pictured as a mapping operation. A given z=x+iymeans a given point in thez-plane.Thecomplexvalueof w(z)isthenapointinthe w-plane.Pointsinthe z-plane map into points in the w-plane and curves in the z-plane map into curves in the w-plane, asindicatedinFig. 6.2. Complex Conjugation In all these steps, complex number, variable, and function, the operation of replacing iby –iis called“takingthe complexconjugate.”The complexconjugateof zis denotedby z∗, where4 z∗=x−iy. (6.9) 4Thecomplex conjugate is often denoted by ¯zinthe mathematicalliterature. 408 Chapter 6 Functions of a Complex Variable I FIGURE 6.3Complexconjugatepoints. The complex variable zand its complex conjugate z∗are mirror images of each other reflected in the x-axis, that is, inversion of the y-axis (compare Fig. 6.3). The product zz∗ leadsto zz∗=(x+iy)(x−iy)=x2+y2=r2. (6.10) Hence (zz∗)1/2=|z|, themagnitude ofz. Functions of a Complex Variable All the elementary functions of real variables may be extended into the complex plane— replacing the real variable xby the complex variable z. This is an example of the analytic continuation mentioned in Section 6.5. The extremely important relation of Eq. (6.4c) is anillustration.Movingintothecomplexplaneopensupnewopportunitiesfor analysis. Example 6.1.1 DEMOIVRE ’SFORMULA If Eq.(6.4c) (setting r=1)is raisedtothe nthpower,wehave einθ=(cosθ+isinθ)n. (6.11) Expandingtheexponentialnowwithargument nθ,weobtain cosnθ+isinnθ=(cosθ+isinθ)n. (6.12) DeMoivre’sformulaisgeneratediftheright-handsideofEq.(6.12)isexpandedbythebi- nomialtheorem;weobtaincos nθasaseriesofpowersofcos θandsinθ,Exercise6.1.6. /squaresolid Numerous other examples of relations among the exponential, hyperbolic, and trigono- metricfunctionsinthecomplexplaneappearintheexercises. Occasionally there are complications. The logarithm of a complex variable may be ex- pandedusingthepolarrepresentation lnz=lnreiθ=lnr+iθ. (6.13a) 6.1 Complex Algebra 409 Thisisnotcomplete.Tothephaseangle, θ,wemayaddanyintegralmultipleof2 πwithout changing z. HenceEq.(6.13a) shouldread lnz=lnrei(θ+2nπ)=lnr+i(θ+2nπ). (6.13b) Theparameter nmaybeanyinteger.Thismeansthatln zisamultivalued functionhaving an infinite number of values for a single pair of real values randθ. To avoid ambiguity, the simplest choice is n=0 and limitation of the phase to an interval of length 2 π, such as(−π,π).5T h el i n ei nt h e z-plane that is not crossed, the negative real axis in this case, is labeled a cut lineorbranch cut .T h ev a l u eo fl n zwithn=0 is called the principal valueof lnz. Further discussion of these functions, including the logarithm, appears in Section6.7. Exercises 6.1.1 (a) Findthereciprocalof x+iy, workingentirelyintheCartesianrepresentation. (b) Repeat part (a), working in polar form but expressing the final result in Cartesian form. 6.1.2 The complex quantities a=u+ivandb=x+iymay also be represented as two- dimensionalvectors a=ˆxu+ˆyv,b=ˆxx+ˆyy.Showthat a∗b=a·b+iˆz·a×b. 6.1.3 Provealgebraicallythatforcomplexnumbers, |z1|−|z2|≤|z1+z2|≤|z1|+|z2|. Interpretthisresultintermsof two-dimensionalvectors.Provethat |z−1|<vextendsinglevextendsingleradicalbig z2−1vextendsinglevextendsingle<|z+1|,forℜ(z)>0. 6.1.4 We may define a complex conjugation operator Ksuch that Kz=z∗. Show that Kis notalinearoperator. 6.1.5 Show that complex numbers have square roots and that the square roots are contained inthecomplexplane.What arethesquareroots of i? 6.1.6 Showthat (a) cosnθ=cosnθ−parenleftbign 2parenrightbig cosn−2θsin2θ+parenleftbign 4parenrightbig cosn−4θsin4θ−···. (b) sinnθ=parenleftbign 1parenrightbig cosn−1θsinθ−parenleftbign 3parenrightbig cosn−3θsin3θ+···. Note.Thequantitiesparenleftbign mparenrightbig arebinomialcoefficients:parenleftbign mparenrightbig =n!/[(n−m)!m!]. 6.1.7 Provethat (a)N−1summationdisplay n=0cosnx=sin(Nx/2) sinx/2cos(N−1)x 2, 5Thereis no standard choice of phase; the appropriate phase depends on eachproblem. 410 Chapter 6 Functions of a Complex Variable I (b)N−1summationdisplay n=0sinnx=sin(Nx/2) sinx/2sin(N−1)x 2. Theseseriesoccurintheanalysisofthemultiple-slitdiffractionpattern.Anotherappli- cationistheanalysisoftheGibbsphenomenon,Section14.5. Hint. Parts (a) and (b) may be combined to form a geometric series (compare Sec- tion5.1). 6.1.8 For−1<p<1 prove that (a)∞summationdisplay n=0pncosnx=1−pcosx 1−2pcosx+p2, (b)∞summationdisplay n=0pnsinnx=psinx 1−2pcosx+p2. Theseseries occurinthetheoryof theFabry–Perotinterferometer. 6.1.9 Assume that the trigonometric functions and the hyperbolic functions are defined for complexargumentbytheappropriatepowerseries sinz=∞summationdisplay n=1,odd(−1)(n−1)/2zn n!=∞summationdisplay s=0(−1)sz2s+1 (2s+1)!, cosz=∞summationdisplay n=0,even(−1)n/2zn n!=∞summationdisplay s=0(−1)sz2s (2s)!, sinhz=∞summationdisplay n=1,oddzn n!=∞summationdisplay s=0z2s+1 (2s+1)!, coshz=∞summationdisplay n=0,evenzn n!=∞summationdisplay s=0z2s (2s)!. (a) Showthat isinz=sinhiz,siniz=isinhz, cosz=coshiz,cosiz=coshz. (b) Verifythatfamiliarfunctionalrelationssuchas coshz=ez+e−z 2, sin(z1+z2)=sinz1cosz2+sinz2cosz1, stillholdinthecomplexplane. 6.1 Complex Algebra 411 6.1.10 Usingtheidentities cosz=eiz+e−iz 2,sinz=eiz−e−iz 2i, establishedfromcomparisonofpowerseries, showthat (a) sin(x+iy)=sinxcoshy+icosxsinhy, cos(x+iy)=cosxcoshy−isinxsinhy, (b)|sinz|2=sin2x+sinh2y,|cosz|2=cos2x+sinh2y. Thisdemonstratesthatwemayhave |sinz|,|cosz|>1 inthecomplexplane. 6.1.11 FromtheidentitiesinExercises6.1.9 and6.1.10showthat (a) sinh (x+iy)=sinhxcosy+icoshxsiny, cosh(x+iy)=coshxcosy+isinhxsiny, (b)|sinhz|2=sinh2x+sin2y,|coshz|2=cosh2x+sin2y. 6.1.12 Provethat (a)|sinz|≥|sinx|(b)|cosz|≥|cosx|. 6.1.13 Showthattheexponentialfunction ezisperiodicwithapureimaginaryperiodof 2 πi. 6.1.14 Showthat (a) tanhz 2=sinhx+isiny coshx+cosy, (b) cothz 2=sinhx−isiny coshx−cosy. 6.1.15 Findallthezerosof (a) sinz, (b) cos z, (c) sinh z, (d) cosh z. 6.1.16 Showthat (a) sin−1z=−ilnparenleftbig iz±radicalbig 1−z2parenrightbig , (d) sinh−1z=lnparenleftbig z+radicalbig z2+1parenrightbig , (b) cos−1z=−ilnparenleftbig z±radicalbig z2−1parenrightbig ,(e) cosh−1z=lnparenleftbig z+radicalbig z2−1parenrightbig , (c) tan−1z=i 2lnparenleftbiggi+z i−zparenrightbigg , (f) tanh−1z=1 2lnparenleftbigg1+z 1−zparenrightbigg . Hint.1. Express the trigonometric and hyperbolic functions in terms of exponentials. 2.Solvefor theexponentialandthenfor theexponent. 6.1.17 Inthequantumtheoryofthephotoionizationweencountertheidentity parenleftbiggia−1 ia+1parenrightbiggib =expparenleftbig −2bcot−1aparenrightbig , inwhich aandbarereal.Verifythisidentity. 412 Chapter 6 Functions of a Complex Variable I 6.1.18 Aplanewaveof lightof angularfrequency ωisrepresentedby eiω(t−nx/c). In a certain substance the simple real index of refraction nis replaced by the complex quantityn−ik. What is the effect of kon the wave? What does kcorrespond to phys- ically? The generalization of a quantity from real to complex form occurs frequently in physics. Examples range from the complex Young’s modulus of viscoelastic materi- als to the complex (optical) potential of the “cloudy crystal ball” model of the atomic nucleus. 6.1.19 Weseethatfor theangularmomentumcomponentsdefinedinExercise2.5.14, Lx−iLy/negationslash=(Lx+iLy)∗. Explainwhythisoccurs. 6.1.20 Showthatthe phaseoff(z)=u+ivisequaltotheimaginarypartofthelogarithmof f(z). Exercise8.2.13dependsonthisresult. 6.1.21 (a) Showthat elnzalwaysequals z. (b) Showthat ln ezdoesnotalwaysequal z. 6.1.22 The infinite product representations of Section 5.11 hold when the real variable xis replaced by the complex variable z. From this, develop infinite product representations for (a) sinhz, (b) cosh z. 6.1.23 Theequationof motionofamass mrelativeto arotatingcoordinatesystem is md2r dt2=F−mω×(ω×r)−2mparenleftbigg ω×dr dtparenrightbigg −mparenleftbiggdω dt×rparenrightbigg . Consider the case F=0,r=ˆxx+ˆyy, andω=ωˆz, withωconstant. Show that the replacementof r=ˆxx+ˆyybyz=x+iyleadsto d2z dt2+i2ωdz dt−ω2z=0. Note.This ODEmaybesolvedbythesubstitution z=fe−iωt. 6.1.24 Using the complex arithmetic available in FORTRAN, write a program that will cal- culate the complex exponential ezfrom its series expansion (definition). Calculate ez forz=einπ/6,n=0,1,2,...,12.Tabulatethephaseangle (θ=nπ/6),ℜz,ℑz,ℜ(ez), ℑ(ez),|ez|, andthephaseof ez. Checkvalue. n=5,θ=2.61799,ℜ(z)=−0.86602, ℑz=0.50000,ℜ(ez)=0.36913,ℑ(ez)=0.20166, |ez|=0.42062,phase (ez)=0.50000. 6.1.25 UsingthecomplexarithmeticavailableinFORTRAN,calculateandtabulate ℜ(sinhz), ℑ(sinhz),|sinhz|, andphase (sinhz)forx=0.0(0.1)1.0 andy=0.0(0.1)1.0. 6.2 Cauchy–Riemann Conditions 413 Hint.Bewareofdividingbyzerowhencalculatinganangleas anarctangent. Checkvalue. z=0.2+0.1i,ℜ(sinhz)=0.20033, ℑ(sinhz)=0.10184,|sinhz|=0.22473, phase(sinhz)=0.47030. 6.1.26 RepeatExercise6.1.25for cosh z. 6.2 C AUCHY –RIEMANN CONDITIONS Having established complex functions of a complexvariable, we now proceed to differen- tiatethem.Thederivativeof f(z), like thatofa realfunction,is definedby lim δz→0f(z+δz)−f(z) z+δz−z=lim δz→0δf (z) δz=df dz=f′(z), (6.14) provided that the limit is independent of the particular approach to the point z. For real variables we require that the right-hand limit ( x→x0from above) and the left-hand limit (x→x0from below)be equalfor thederivative df(x)/dx toexistat x=x0.No w ,wi t h z (orz0)somepointinaplane,ourrequirementthatthelimitbeindependentofthedirection ofapproachis veryrestrictive. Considerincrements δxandδyofthevariables xandy,respectively.Then δz=δx+iδy. (6.15) Also, δf=δu+iδv, (6.16) sothat δf δz=δu+iδv δx+iδy. (6.17) Let us take the limit indicated by Eq. (6.14) by two different approaches, as shown in Fig.6.4.First, with δy=0,weletδx→0. Equation(6.14) yields lim δz→0δf δz=lim δx→0parenleftbiggδu δx+iδv δxparenrightbigg =∂u ∂x+i∂v ∂x, (6.18) FIGURE 6.4Alternate approachesto z0. 414 Chapter 6 Functions of a Complex Variable I assuming the partial derivatives exist. For a second approach, we set δx=0 and then let δy→0.Thisleadsto lim δz→0δf δz=lim δy→0parenleftbigg −iδu δy+δv δyparenrightbigg =−i∂u ∂y+∂v ∂y. (6.19) If we are to have a derivative df/dz, Eqs. (6.18) and (6.19) must be identical. Equating realpartstorealpartsandimaginarypartstoimaginaryparts(likecomponentsofvectors), weobtain ∂u ∂x=∂v ∂y,∂u ∂y=−∂v ∂x. (6.20) Thesearethefamous Cauchy–Riemann conditions.TheywerediscoveredbyCauchyand used extensively by Riemann in his theory of analytic functions. These Cauchy–Riemann conditions are necessary for the existence of a derivative of f(z); that is, if df/dzexists, theCauchy–Riemannconditionsmusthold. Conversely, if the Cauchy–Riemann conditions are satisfied and the partial derivatives ofu(x,y)andv(x,y)are continuous, the derivative df/dzexists. This may be shown by writing δf=parenleftbigg∂u ∂x+i∂v ∂xparenrightbigg δx+parenleftbigg∂u ∂y+i∂v ∂yparenrightbigg δy. (6.21) The justification for this expression depends on the continuity of the partial derivatives of uandv.Dividingby δz,weha v e δf δz=(∂u/∂x+i(∂v/∂x))δx+(∂u/∂y+i(∂v/∂y))δy δx+iδy =(∂u/∂x+i(∂v/∂x))+(∂u/∂y+i(∂v/∂y))δy/δx 1+i(δy/δx). (6.22) Ifδf/δzistohaveauniquevalue,thedependenceon δy/δxmustbeeliminated.Apply- ingtheCauchy–Riemannconditionstothe yderivatives,weobtain ∂u ∂y+i∂v ∂y=−∂v ∂x+i∂u ∂x. (6.23) SubstitutingEq. (6.23)intoEq. (6.22), wemaycanceloutthe δy/δxdependenceand δf δz=∂u ∂x+i∂v ∂x, (6.24) which shows that lim δf/δzis independent of the direction of approach in the complex plane as long as the partial derivatives are continuous. Thus,df dzexists and fis analytic atz. It is worthwhile noting that the Cauchy–Riemann conditions guarantee that the curves u=c1willbe orthogonalto thecurves v=c2(compareSection2.1). This is fundamental in application to potential problems in a variety of areas of physics. If u=c1is a line of 6.2 Cauchy–Riemann Conditions 415 electric force, then v=c2is an equipotentialline (surface), and vice versa. To see this, let uswritetheCauchy–Riemannconditionsasa productofratiosof partialderivatives, ux uy·vx vy=−1, (6.25) withtheabbreviations ∂u ∂x≡ux,∂u ∂y≡uy,∂v ∂x≡vx,∂v ∂y≡vy. Now recall the geometric meaning of −ux/uyas the slope of the tangent of each curve u(x,y)=const. and similarly for v(x,y)=const. This means that the u=const. and v=const. curvesaremutuallyorthogonalateachintersection.Alternatively, uxdx+uydy=0=vydx−vxdy saysthat,if (dx,dy)istangenttothe u-curve,thentheorthogonal (−dy,dx)istangentto thev-curve at the intersection point, z=(x,y). Or equivalently, uxvx+uyvy=0 implies that thegradient vectors (ux,uy)and(vx,vy)are perpendicular . A further implication forpotentialtheoryis developedinExercise6.2.1. Analytic Functions Finally, if f(z)is differentiable at z=z0and in some small region around z0, we say that f(z)isanalytic6atz=z0.Iff(z)isanalyticeverywhereinthe(finite)complexplane,we callitanentirefunction.Ourtheoryofcomplexvariableshereisoneofanalyticfunctions of a complex variable, which points up the crucial importance of the Cauchy–Riemann conditions. The concept of analyticity carried on in advanced theories of modern physics plays a crucial role in dispersion theory (of elementary particles). If f′(z)does not exist atz=z0, thenz0is labeled a singular point and consideration of it is postponed until Section6.6. ToillustratetheCauchy–Riemannconditions,considertwoverysimpleexamples. Example 6.2.1 z2ISANALYTIC Letf(z)=z2.Thentherealpart u(x,y)=x2−y2andtheimaginarypart v(x,y)=2xy. FollowingEq.(6.20), ∂u ∂x=2x=∂v ∂y,∂u ∂y=−2y=−∂v ∂x. We see that f(z)=z2satisfies the Cauchy–Riemann conditions throughout the complex plane. Since the partial derivatives are clearly continuous, we conclude that f(z)=z2is analytic. /squaresolid 6Some writers use the term holomorphic orregular. 416 Chapter 6 Functions of a Complex Variable I Example 6.2.2 z∗ISNOTANALYTIC Letf(z)=z∗.N o wu=xandv=−y. Applying the Cauchy–Riemann conditions, we obtain ∂u ∂x=1/negationslash=∂v ∂y=−1. TheCauchy–Riemannconditionsarenotsatisfiedand f(z)=z∗isnotananalyticfunction ofz. It is interesting to note that f(z)=z∗is continuous, thus providing an example of a functionthatis everywherecontinuousbutnowheredifferentiableinthecomplexplane. The derivative of a real function of a real variable is essentially a local characteristic, in thatitprovidesinformationaboutthefunctiononlyinalocalneighborhood—forinstance, as a truncated Taylor expansion. The existence of a derivative of a function of a complex variablehasmuchmorefar-reachingimplications.Therealandimaginarypartsofouran- alytic function must separately satisfy Laplace’s equation. This is Exercise 6.2.1. Further, our analytic function is guaranteed derivatives of all orders, Section 6.4. In this sense the derivative not only governs the local behavior of the complex function, but controls the distantbehavioras well. /squaresolid Exercises 6.2.1 The functions u(x,y)andv(x,y)are the real and imaginary parts, respectively, of an analyticfunction w(z). (a) Assumingthattherequiredderivativesexist,showthat ∇2u=∇2v=0. Solutions of Laplace’s equation such as u(x,y)andv(x,y)are called harmonic functions. (b) Showthat ∂u ∂x∂u ∂y+∂v ∂x∂v ∂y=0, andgiveageometricinterpretation. Hint.ThetechniqueofSection1.6allowsyoutoconstructvectorsnormaltothecurves u(x,y)=ciandv(x,y)=cj. 6.2.2 Showwhetheror notthefunction f(z)=ℜ(z)=xis analytic. 6.2.3 Having shown that the real part u(x,y)and the imaginary part v(x,y)of an analytic functionw(z)each satisfy Laplace’s equation, show that u(x,y)andv(x,y)cannot both have either a maximum or a minimum in the interior of any region in which w(z)isanalytic.(Theycanhavesaddlepointsonly.) 6.2 Cauchy–Riemann Conditions 417 6.2.4 LetA=∂2w/∂x2,B=∂2w/∂x∂y,C=∂2w/∂y2. From the calculus of functions of twovariables, w(x,y),weha v ea saddlepoint if B2−AC >0. Withf(z)=u(x,y)+iv(x,y), apply the Cauchy–Riemann conditions and show that neitheru(x,y)norv(x,y)has a maximum or a minimum in a finite region of the complexplane.(SeealsoSection7.3.) 6.2.5 Findtheanalyticfunction w(z)=u(x,y)+iv(x,y) if (a)u(x,y)=x3−3xy2,( b)v(x,y)=e−ysinx. 6.2.6 If there is some common region in which w1=u(x,y)+iv(x,y)andw2=w∗ 1= u(x,y)−iv(x,y)arebothanalytic,provethat u(x,y)andv(x,y)areconstants. 6.2.7 Thefunction f(z)=u(x,y)+iv(x,y)is analytic.Showthat f∗(z∗)is alsoanalytic. 6.2.8 Usingf(reiθ)=R(r,θ)ei/Phi1(r,θ), in which R(r,θ)and/Phi1(r,θ)are differentiable real functions of randθ, show that the Cauchy–Riemann conditions in polar coordinates become (a)∂R ∂r=R r∂/Theta1 ∂θ,(b)1 r∂R ∂θ=−R∂/Theta1 ∂r. Hint.Setupthederivativefirstwith δzradialandthenwith δztangential. 6.2.9 AsanextensionofExercise6.2.8showthat /Theta1(r,θ)satisfiesLaplace’sequationinpolar coordinates. Equation (2.35) (without the final term and set to zero) is the Laplacian in polarcoordinates. 6.2.10 Two-dimensional irrotational fluid flow is conveniently described by a complex poten-tialf(z)=u(x,v)+iv(x,y). We label the real part, u(x,y), the velocity potential and the imaginary part, v(x,y), the stream function. The fluid velocity Vis given by V=∇u.Iff(z)is analytic, (a) Showthat df/dz=Vx−iVy; (b) Showthat ∇·V=0 (nosourcesor sinks); (c) Showthat ∇×V=0 (irrotational,nonturbulentflow). 6.2.11 AproofoftheSchwarzinequality(Section10.4)involvesminimizinganexpression, f=ψaa+λψab+λ∗ψ∗ ab+λλ∗ψbb≥0. Theψareintegralsofproductsoffunctions; ψaaandψbbarereal,ψabiscomplexand λisacomplexparameter. (a) Differentiatetheprecedingexpressionwithrespectto λ∗,treating λasanindepen- dent parameter, independent of λ∗. Show that setting the derivative ∂f/∂λ∗equal tozeroyields λ=−ψ∗ ab ψbb. 418 Chapter 6 Functions of a Complex Variable I (b) Showthat ∂f/∂λ=0 leadstothesameresult. (c) Let λ=x+iy,λ∗=x−iy. Set thexandyderivatives equal to zero and show thatagain λ=−ψ∗ ab ψbb. Thisindependenceof λandλ∗appearsagaininSection17.7. 6.2.12 The function f(z)is analytic. Show that the derivative of f(z)with respect to z∗does notexistunless f(z)is aconstant. Hint.Usethechainruleandtake x=(z+z∗)/2,y=(z−z∗)/2i. Note.Thisresultemphasizesthatouranalyticfunction f(z)isnotjustacomplexfunc- tionof tworealvariables xandy.It isafunctionofthecomplexvariable x+iy. 6.3 C AUCHY ’SINTEGRAL THEOREM Contour Integrals With differentiation under control, we turn to integration. The integral of a complex vari- ableoveracontourinthecomplexplanemaybedefinedincloseanalogytothe(Riemann) integralofarealfunctionintegratedalongthereal x-axis. Wedividethecontourfrom z0toz′ 0intonintervalsbypicking n−1intermediatepoints z1,z2,...onthecontour(Fig.6.5). Considerthesum Sn=nsummationdisplay j=1f(ζj)(zj−zj−1), (6.26) FIGURE 6.5Integrationpath. 6.3 Cauchy’s Integral Theorem 419 whereζjis apointonthecurvebetween zjandzj−1.No wl et n→∞with |zj−zj−1|→0 for allj. If the lim n→∞Snexists and is independent of the details of choosing the points zjandζj, then limn→∞nsummationdisplay j=1f(ζj)(zj−zj−1)=integraldisplayz′ 0 z0f(z)dz. (6.27) Theright-handsideofEq.(6.27)iscalledthecontourintegralof f(z)(alongthespecified contourCfromz=z0toz=z′ 0). The preceding developmentof the contour integral is closely analogous to the Riemann integral of a real function of a real variable. As an alternative, the contour integral may be definedby integraldisplayz2 z1f(z)dz=integraldisplayx2,y2 x1,y1bracketleftbig u(x,y)+iv(x,y)bracketrightbig [dx+idy] =integraldisplayx2,y2 x1,y1bracketleftbig u(x,y)dx−v(x,y)dybracketrightbig +iintegraldisplayx2,y2 x1,y1bracketleftbig v(x,y)dx+u(x,y)dybracketrightbig with the path joining (x1,y1)and(x2,y2)specified. This reduces the complex integral to thecomplexsumofrealintegrals.Itissomewhatanalogoustothereplacementofavector integralbythevectorsumof scalarintegrals,Section1.10. An important example is the contour integralintegraltext Czndz, whereCis a circle of radius r>0 around the origin z=0 in the positive mathematical sense (counterclockwise). In polar coordinates of Eq. (6.4c) we parameterize the circle as z=reiθanddz=ireiθdθ. Forn/negationslash=−1,naninteger,wethenobtain 1 2πiintegraldisplay Czndz=rn+1 2πintegraldisplay2π 0expbracketleftbig i(n+1)θbracketrightbig dθ =bracketleftbig 2πi(n+1)bracketrightbig−1rn+1bracketleftbig ei(n+1)θbracketrightbigvextendsinglevextendsingle2π 0=0 (6.27a) because 2 πis aperiodof ei(n+1)θ, whilefor n=−1 1 2πiintegraldisplay Cdz z=1 2πintegraldisplay2π 0dθ=1, (6.27b) againindependentof r. Alternatively, we can integrate around a rectangle with the corners z1,z2,z3,z4to obtainfor n/negationslash=−1 integraldisplay zndz=zn+1 n+1vextendsinglevextendsinglevextendsinglevextendsinglez2 z1+zn+1 n+1vextendsinglevextendsinglevextendsinglevextendsinglez3 z2+zn+1 n+1vextendsinglevextendsinglevextendsinglevextendsinglez4 z3+zn+1 n+1vextendsinglevextendsinglevextendsinglevextendsinglez1 z4=0, because each corner point appears once as an upper and a lower limit that cancel. For n=−1thecorrespondingrealpartsofthelogarithmscancelsimilarly,buttheirimaginary partsinvolvetheincreasingargumentsofthepointsfrom z1toz4and,whenwecomeback to the first corner z1, its argument has increased by 2 πdue to the multivaluedness of the 420 Chapter 6 Functions of a Complex Variable I logarithm, so 2 πiis left over as the value of the integral. Thus, the value of the integral involving a multivalued function must be that which is reached in a continuous fash- ion on the path being taken . These integrals are examples of Cauchy’s integral theorem, whichweconsiderinthenextsection. Stokes’ Theorem Proof Cauchy’sintegraltheoremisthefirstoftwobasictheoremsinthetheoryofthebehaviorof functions of a complex variable. First, we offer a proof under relatively restrictive condi- tions—conditionsthatareintolerabletothemathematiciandevelopingabeautifulabstract theorybutthatare usuallysatisfiedinphysicalproblems. If a function f(z)is analytic, that is, if its partial derivatives are continuous throughout somesimply connected region R,7for every closed path C(Fig. 6.6) in R, and if it is single-valued(assumedfor simplicityhere),thelineintegralof f(z)aroundCiszero, or integraldisplay Cf(z)dz=contintegraldisplay Cf(z)dz=0. (6.27c) Recall that in Section 1.13 such a function f(z), identified as a force, was labeled conser- vative. The symbolcontintegraltext is used to emphasize that the path is closed. Note that the interior of the simply connected region bounded by a contour is that region lying to the left when moving in the direction implied by the contour; as a rule, a simply connected region is boundedbyasingleclosedcurve. InthisformtheCauchyintegraltheoremmaybeprovedbydirectapplicationofStokes’ theorem(Section1.12). With f(z)=u(x,y)+iv(x,y)anddz=dx+idy, contintegraldisplay Cf(z)dz=contintegraldisplay C(u+iv)(dx+idy) =contintegraldisplay C(udx−vdy)+icontintegraldisplay (vdx+udy). (6.28) ThesetwolineintegralsmaybeconvertedtosurfaceintegralsbyStokes’theorem,aproce- dure that is justified if the partial derivatives are continuous within C. In applying Stokes’ theorem,notethatthefinaltwointegralsofEq. (6.28)arereal. Using V=ˆxVx+ˆyVy, Stokes’theoremsays that contintegraldisplay C(Vxdx+Vydy)=integraldisplayparenleftbigg∂Vy ∂x−∂Vx ∂yparenrightbigg dxdy. (6.29) Forthefirst integralinthelast partofEq. (6.28)let u=Vxandv=−Vy.8Then 7Any closed simple curve (one that does not intersectitself) inside asimply connectedregion or domain may becontracted toa singlepointthatstillbelongstotheregion.Ifaregionisnotsimplyconnected,itiscalledmultiplyconnected.Asanexampleof amultiply connected region, consider the z-plane with the interior of theunit circle excluded . 8In the proof of Stokes’ theorem, Section 1.12, VxandVyare any two functions (with continuous partial derivatives). 6.3 Cauchy’s Integral Theorem 421 FIGURE 6.6Aclosedcontour C withinasimplyconnectedregion R. contintegraldisplay C(udx−vdy)=contintegraldisplay C(Vxdx+Vydy) =integraldisplayparenleftbigg∂Vy ∂x−∂Vx ∂yparenrightbigg dxdy=−integraldisplayparenleftbigg∂v ∂x+∂u ∂yparenrightbigg dxdy. (6.30) For the second integral on the right side of Eq. (6.28) we let u=Vyandv=Vx.U s i n g Stokes’theoremagain,weobtain contintegraldisplay (vdx+udy)=integraldisplayparenleftbigg∂u ∂x−∂v ∂yparenrightbigg dxdy. (6.31) On application of the Cauchy–Riemann conditions, which must hold, since f(z)is as- sumedanalytic,eachintegrandvanishesand contintegraldisplay f(z)dz=−integraldisplayparenleftbigg∂v ∂x+∂u ∂yparenrightbigg dxdy+iintegraldisplayparenleftbigg∂u ∂x−∂v ∂yparenrightbigg dxdy=0.(6.32) Cauchy–Goursat Proof ThiscompletestheproofofCauchy’sintegraltheorem.However,theproofismarredfrom atheoreticalpointofviewbytheneedforcontinuityofthefirstpartialderivatives.Actually, as shown by Goursat, this conditionis not necessary. An outlineof the Goursat proof is as follows. We subdivide the region inside the contour Cinto a network of small squares, as indicatedinFig.6.7. Thencontintegraldisplay Cf(z)dz=summationdisplay jcontintegraldisplay Cjf(z)dz, (6.33) all integrals along interior lines canceling out. To estimate thecontintegraltext Cjf(z)dz, we construct thefunction δj(z,zj)=f(z)−f(zj) z−zj−df(z) dzvextendsinglevextendsinglevextendsinglevextendsingle z=zj, (6.34) withzjan interior point of the jth subregion. Note that [f(z)−f(zj)]/(z−zj)is an approximation to the derivative at z=zj. Equivalently, we may note that if f(z)had 422 Chapter 6 Functions of a Complex Variable I FIGURE 6.7Cauchy–Goursatcontours. aTaylorexpansion(whichwehavenotyetproved),then δj(z,zj)wouldbeoforder z−zj, approaching zero as the network was made finer. But since f′(zj)exists, that is, is finite, wemaymake vextendsinglevextendsingleδj(z,zj)vextendsinglevextendsingle<ε, (6.35) whereεis an arbitrarily chosen small positive quantity. Solving Eq. (6.34) for f(z)and integratingaround Cj, weobtain contintegraldisplay Cjf(z)dz=contintegraldisplay Cj(z−zj)δj(z,zj)dz, (6.36) theintegralsoftheothertermsvanishing.9WhenEqs.(6.35)and(6.36)arecombined,one showsthatvextendsinglevextendsinglevextendsinglevextendsinglesummationdisplay jcontintegraldisplay Cjf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle<Aε, (6.37) whereAis a term of the order of the area of the enclosed region. Since εis arbitrary, we letε→0 andconcludethatif afunction f(z)isanalyticonandwithinaclosedpath C, contintegraldisplay Cf(z)dz=0. (6.38) Detailsoftheproofofthissignificantlymoregeneralandmorepowerfulformcanbefound in Churchill in the Additional Readings. Actually we can still prove the theorem for f(z) analyticwithintheinteriorof Candonlycontinuouson C. The consequence of the Cauchy integral theorem is that for analytic functions the line integralisafunctiononlyofits endpoints,independentofthepathof integration, integraldisplayz2 z1f(z)dz=F(z2)−F(z1)=−integraldisplayz1 z2f(z)dz, (6.39) againexactlylikethecaseof aconservativeforce, Section1.13. 9contintegraltext dzandcontintegraltext zdz=0 by Eq.(6.27a). 6.3 Cauchy’s Integral Theorem 423 Multiply Connected Regions TheoriginalstatementofCauchy’sintegraltheoremdemandedasimplyconnectedregion. This restriction may be relaxed by the creation of a barrier, a contour line. The purpose of the following contour-line construction is to permit, within a multiply connected region, the identification of curves that can be shrunk to a point within the region, that is, the constructionof asubregionthatissimplyconnected. Consider the multiply connected region of Fig. 6.8, in which f(z)is not defined for the interior,R′. Cauchy’s integral theorem is not valid for the contour C, as shown, but we can construct a contour C′for which the theorem holds. We draw a line from the interior forbiddenregion, R′,totheforbiddenregionexteriorto Randthenrunanewcontour, C′, assho wninFig.6. 9. The new contour, C′, through ABDEFGA never crosses the contour line that literally convertsRintoasimplyconnectedregion.Thethree-dimensionalanalogofthistechnique wasusedinSection1.14toproveGauss’law.ByEq. (6.39), integraldisplayA Gf(z)dz=−integraldisplayD Ef(z)dz, (6.40) FIGURE 6.8Aclosedcontour Cina multiplyconnectedregion. FIGURE 6.9Conversionof amultiply connectedregionintoasimplyconnected region. 424 Chapter 6 Functions of a Complex Variable I withf(z)having been continuous across the contour line and line segments DEandGA arbitrarilyclosetogether.Then contintegraldisplay C′f(z)dz=integraldisplay ABDf(z)dz+integraldisplay EFGf(z)dz=0 (6.41) by Cauchy’s integral theorem, with region Rnow simply connected. Applying Eq. (6.39) onceagainwith ABD→C′ 1andEFG→−C′ 2, weobtain contintegraldisplay C′ 1f(z)dz=contintegraldisplay C′ 2f(z)dz, (6.42) in which C′ 1andC′ 2are both traversed in the same (counterclockwise, that is, positive) direction. Let us emphasize that the contour line here is a matter of mathematical convenience, to permit the application of Cauchy’s integral theorem. Since f(z)is analytic in the annular region,itis necessarilysingle-valuedandcontinuousacross anysuchcontourline. Exercises 6.3.1 Showthatintegraltextz2 z1f(z)dz=−integraltextz1 z2f(z)dz. 6.3.2 Provethat vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay Cf(z)dzvextendsinglevextendsinglevextendsinglevextendsingle≤|f|max·L, where|f|maxis the maximum value of |f(z)|along the contour CandLis the length ofthecontour. 6.3.3 Verifythat integraldisplay1,1 0,0z∗dz depends on the path by evaluating the integral for the two paths shown in Fig. 6.10. Recallthat f(z)=z∗isnotananalyticfunctionof zandthatCauchy’sintegraltheorem thereforedoesnotapply. 6.3.4 Showthat contintegraldisplay Cdz z2+z=0, inwhichthecontour Cis acircledefinedby |z|=R>1. Hint. Direct use of the Cauchy integral theorem is illegal. Why? The integral may be evaluatedbytransformingtopolarcoordinatesandusingtables.Thisyields0for R>1 and 2πiforR<1. 6.4 Cauchy’s Integral Formula 425 FIGURE 6.10Contour. 6.4 C AUCHY ’SINTEGRAL FORMULA Asintheprecedingsection,weconsiderafunction f(z)thatisanalyticonaclosedcontour Candwithintheinteriorregionboundedby C.We seektoprovethat 1 2πicontintegraldisplay Cf(z) z−z0dz=f(z0), (6.43) in which z0is any point in the interior region bounded by C. This is the second of the two basic theorems mentioned in Section 6.3. Note that since zis on the contour Cwhile z0is in the interior, z−z0/negationslash=0 and the integral Eq. (6.43) is well defined. Although f(z) is assumed analytic, the integrand is f(z)/(z−z0)and is not analytic at z=z0unless f(z0)=0. If the contour is deformed as shown in Fig. 6.11 (or Fig. 6.9, Section 6.3), Cauchy’sintegraltheoremapplies.ByEq. (6.42), contintegraldisplay Cf(z) z−z0dz−contintegraldisplay C2f(z) z−z0dz=0, (6.44) whereCistheoriginaloutercontourand C2isthecirclesurroundingthepoint z0traversed inacounterclockwise direction.Let z=z0+reiθ,usingthepolarrepresentationbecause of the circular shape of the path around z0.H e r eris small and will eventually be made to approachzero.Wehave(with dz=ireirθdθfromEq. (6.27a)) contintegraldisplay C2f(z) z−z0dz=contintegraldisplay C2f(z0+reiθ) reiθrieiθdθ. Takingthelimitas r→0,weobtain contintegraldisplay C2f(z) z−z0dz=if(z0)integraldisplay C2dθ=2πif(z0), (6.45) 426 Chapter 6 Functions of a Complex Variable I FIGURE 6.11Exclusionof a singularpoint. sincef(z)is analytic and therefore continuous at z=z0. This proves the Cauchy integral formula. Hereisaremarkableresult.Thevalueofananalyticfunction f(z)isgivenataninterior pointz=z0once the values on the boundary Care specified. This is closely analogous to atwo-dimensionalformofGauss’law(Section1.14)inwhichthemagnitudeofaninterior linechargewouldbegivenintermsofthecylindricalsurfaceintegraloftheelectricfield E. A further analogy is the determination of a function in real space by an integral of the function and the corresponding Green’s function (and their derivatives) over the bounding surface. Kirchhoffdiffractiontheoryisanexampleofthis. It has been emphasized that z0is an interior point. What happens if z0is exterior to C? In this case the entire integrand is analytic on and within C. Cauchy’s integral theorem, Section6.3,appliesandtheintegralvanishes.Wehave 1 2πicontintegraldisplay Cf(z)dz z−z0=braceleftBigg f(z0), z0interior 0,z 0exterior. Derivatives Cauchy’s integral formula may be used to obtain an expression for the derivative of f(z). FromEq. (6.43), with f(z)analytic, f(z0+δz0)−f(z0) δz0=1 2πiδz0parenleftbiggcontintegraldisplayf(z) z−z0−δz0dz−contintegraldisplayf(z) z−z0dzparenrightbigg . Then,bydefinitionof derivative(Eq. (6.14)), f′(z0)=lim δz0→01 2πiδz0contintegraldisplayδz0f(z) (z−z0−δz0)(z−z0)dz =1 2πicontintegraldisplayf(z) (z−z0)2dz. (6.46) This result could have been obtained by differentiating Eq. (6.43) under the integral sign withrespectto z0.Thisformal,orturning-the-crank,approachisvalid,butthejustification forit iscontainedintheprecedinganalysis. 6.4 Cauchy’s Integral Formula 427 This technique for constructing derivatives may be repeated. We write f′(z0+δz0) andf′(z0), using Eq. (6.46). Subtracting, dividing by δz0, and finally taking the limit as δz0→0,wehave f(2)(z0)=2 2πicontintegraldisplayf(z)dz (z−z0)3. Note that f(2)(z0)is independent of the direction of δz0, as it must be. Continuing, we get10 f(n)(z0)=n! 2πicontintegraldisplayf(z)dz (z−z0)n+1; (6.47) that is, the requirement that f(z)be analytic guarantees not only a first derivative but derivativesof allordersaswell!Thederivativesof f(z)areautomaticallyanalytic.Notice thatthisstatementassumestheGoursatversionoftheCauchyintegraltheorem.Thisisalso why Goursat’s contribution is so significant in the development of the theory of complex variables. Morera’s Theorem A further application of Cauchy’s integral formula is in the proof of Morera’s theorem, whichistheconverseofCauchy’sintegraltheorem.Thetheoremstatesthefollowing: If afunction f(z)is continuousinasimplyconnectedregion Randcontintegraltext Cf(z)dz=0for every closed contour CwithinR, thenf(z)is analytic throughout R. Let us integrate f(z)fromz1toz2. Since every closed-path integral of f(z)vanishes, the integral is independent of path and depends only on its endpoints. We label the result oftheintegration F(z), with F(z2)−F(z1)=integraldisplayz2 z1f(z)dz. (6.48) Asanidentity, F(z2)−F(z1) z2−z1−f(z1)=integraltextz2 z1[f(t)−f(z1)]dt z2−z1, (6.49) usingtasanothercomplexvariable.Nowwetakethelimitas z2→z1: limz2→z1integraltextz2 z1[f(t)−f(z1)]dt z2−z1=0, (6.50) 10This expression is the starting point for defining derivatives of fractional order . See A. Erdelyi (ed.), Tables of Integral Transforms ,Vol.2.NewYork:McGraw-Hill(1954).Forrecentapplicationstomathematicalanalysis,seeT.J.Osler,Anintegral analogue of Taylor’s series and its use in computing Fourier transforms. Math. Comput. 26: 449 (1972), and referencestherein. 428 Chapter 6 Functions of a Complex Variable I sincef(t)is continuous.11Therefore limz2→z1F(z2)−F(z1) z2−z1=F′(z)vextendsinglevextendsingle z=z1=f(z1) (6.51) by definition of derivative (Eq. (6.14)). We have proved that F′(z)atz=z1exists and equalsf(z1). Sincez1is any point in R, we see that F(z)is analytic. Then by Cauchy’s integral formula (compare Eq. (6.47)), F′(z)=f(z)is also analytic, proving Morera’s theorem. Drawing once more on our electrostatic analog, we might use f(z)to represent the electrostatic field E. If the net charge within every closed region in Ris zero (Gauss’ law), the charge density is everywhere zero in R. Alternatively, in terms of the analysis of Section1.13, f(z)representsaconservativeforce(bydefinitionofconservative),andthen wefindthatitisalwayspossibletoexpressitasthederivativeofapotentialfunction F(z). AnimportantapplicationofCauchy’sintegralformulaisthefollowing Cauchyinequal- ity.I ff(z)=summationtextanznis analytic and bounded, |f(z)|≤Mon a circle of radius rabout theorigin,then |an|rn≤M(Cauchy’sinequality ) (6.52) gives upper bounds for the coefficients of its Taylor expansion. To prove Eq. (6.52) let us defineM(r)=max|z|=r|f(z)|andusetheCauchyintegralfor an: |an|=1 2πvextendsinglevextendsinglevextendsinglevextendsingleintegraldisplay |z|=rf(z) zn+1dzvextendsinglevextendsinglevextendsinglevextendsingle≤M(r)2πr 2πrn+1. An immediate consequence of the inequality (6.52) is Liouville’s theorem :I ff(z)is analyticandboundedintheentirecomplexplaneitisaconstant.Infact,if |f(z)|≤Mfor allz, thenCauchy’sinequality(6.52) gives |an|≤Mr−n→0a sr→∞forn>0.Hence f(z)=a0. Conversely,theslightestdeviationofananalyticfunctionfromaconstantvalueimplies that there must be at least one singularity somewhere in the infinite complex plane. Apart from the trivial constant functions, then, singularities are a fact of life, and we must learn to live with them. But we shall do more than that. We shall next expand a function in a Laurent series at a singularity, and we shall use singularities to develop the powerful and usefulcalculusofresiduesinChapter7. A famous application of Liouville’s theorem yields the fundamental theorem of alge- bra(due to C. F. Gauss), which says that any polynomial P(z)=summationtextn ν=0aνzνwithn>0 andan/negationslash=0 hasnroots. To prove this, suppose P(z)has no zero. Then 1 /P(z)is analytic and bounded as |z|→∞. HenceP(z)is a constant by Liouville’s theorem, q.e.a. Thus, P(z)has at least one root that we can divide out. Then we repeat the process for the re- sulting polynomial of degree n−1. This leads to the conclusion that P(z)has exactly n roots. 11Wequote the meanvalue theoremof calculus here. 6.4 Cauchy’s Integral Formula 429 Exercises 6.4.1 Showthat contintegraldisplay C(z−z0)ndz=braceleftBigg 2πi, n=−1, 0,n/negationslash=−1, where the contour Cencircles the point z=z0in a positive (counterclockwise) sense. The exponent nis an integer. See also Eq. (6.27a). The calculus of residues, Chapter 7, isbasedonthisresult. 6.4.2 Showthat 1 2πicontintegraldisplay zm−n−1dz, m andnintegers (withthecontourencirclingtheoriginoncecounterclockwise)isarepresentationofthe Kronecker δmn. 6.4.3 SolveExercise6.3.4byseparatingtheintegrandintopartialfractionsandthenapplying Cauchy’sintegraltheoremformultiplyconnectedregions. Note. Partial fractions are explained in Section 15.8 in connection with Laplace trans- forms. 6.4.4 Evaluatecontintegraldisplay Cdz z2−1, whereCisthecircle|z|=2. 6.4.5 Assuming that f(z)is analytic on and within a closed contour Cand that the point z0 iswithin C,showthat contintegraldisplay Cf′(z) z−z0dz=contintegraldisplay Cf(z) (z−z0)2dz. 6.4.6 You know that f(z)is analytic on and within a closed contour C. You suspect that the nthderivative f(n)(z0)is givenby f(n)(z0)=n! 2πicontintegraldisplay Cf(z) (z−z0)n+1dz. Usingmathematicalinduction,provethatthisexpressioniscorrect. 6.4.7 (a) A function f(z)is analytic within a closed contour C(and continuous on C). If f(z)/negationslash=0 withinCand|f(z)|≤MonC, showthat vextendsinglevextendsinglef(z)vextendsinglevextendsingle≤M forallpointswithin C. Hint.Consider w(z)=1/f (z). (b) Iff(z)=0 within the contour C, show that the foregoing result does not hold and that it is possible to have |f(z)|=0 at one or more points in the interior with |f(z)|>0overtheentireboundingcontour.Citeaspecificexampleofananalytic functionthatbehavesthisway. 430 Chapter 6 Functions of a Complex Variable I 6.4.8 Using the Cauchy integral formula for the nth derivative, convert the following Ro- driguesformulasintothecorrespondingso-calledSchlaefliintegrals. (a) Legendre: Pn(x)=1 2nn!dn dxnparenleftbig x2−1parenrightbign. ANS.(−1)n 2n·1 2πicontintegraldisplay(1−z2)n (z−x)n+1dz. (b) Hermite: Hn(x)=(−1)nex2dn dxne−x2. (c) Laguerre: Ln(x)=ex n!dn dxnparenleftbig xne−xparenrightbig . Note. From the Schlaefli integral representations one can develop generating functions for thesespecialfunctions.CompareSections12.4,13.1, and13.2. 6.5 L AURENT EXPANSION Taylor Expansion TheCauchyintegralformulaoftheprecedingsectionopensupthewayforanotherderiva- tion of Taylor’s series (Section 5.6), but this time for functions of a complex variable. Supposewearetryingtoexpand f(z)aboutz=z0andwehave z=z1asthenearestpoint on the Argand diagram for which f(z)is not analytic. We construct a circle Ccentered at z=z0with radius less than |z1−z0|(Fig. 6.12). Since z1was assumed to be the nearest pointatwhich f(z)wasnotanalytic, f(z)is necessarilyanalyticonandwithin C. FromEq. (6.43), theCauchyintegralformula, f(z)=1 2πicontintegraldisplay Cf(z′)dz′ z′−z =1 2πicontintegraldisplay Cf(z′)dz′ (z′−z0)−(z−z0) =1 2πicontintegraldisplay Cf(z′)dz′ (z′−z0)[1−(z−z0)/(z′−z0)]. (6.53) Herez′is a point on the contour Candzis any point interior to C. It is not legal yet to expandthedenominatoroftheintegrandinEq.(6.53)bythebinomialtheorem,forwehave notyetprovedthebinomialtheoremfor complexvariables.Instead,wenotetheidentity 1 1−t=1+t+t2+t3+···=∞summationdisplay n=0tn, (6.54) 6.5 Laurent Expansion 431 FIGURE 6.12CirculardomainforTaylor expansion. which may easily be verified by multiplying both sides by 1 −t. The infinite series, fol- lowingthemethodsofSection5.2,is convergentfor |t|<1. Now, for a point zinterior to C,|z−z0|<|z′−z0|, and, using Eq. (6.54), Eq. (6.53) becomes f(z)=1 2πicontintegraldisplay C∞summationdisplay n=0(z−z0)nf(z′)dz′ (z′−z0)n+1. (6.55) Interchanging the order of integration and summation (valid because Eq. (6.54) is uni- formlyconvergentfor |t|<1),weobtain f(z)=1 2πi∞summationdisplay n=0(z−z0)ncontintegraldisplay Cf(z′)dz′ (z′−z0)n+1. (6.56) ReferringtoEq. (6.47), weget f(z)=∞summationdisplay n=0(z−z0)nf(n)(z0) n!, (6.57) which is our desired Taylor expansion. Note that it is based only on the assumption that f(z)isanalyticfor|z−z0|<|z1−z0|.Justasforrealvariablepowerseries(Section5.7), thisexpansionis uniquefora given z0. FromtheTaylorexpansionfor f(z)abinomialtheoremmaybederived(Exercise6.5.2). Schwarz Reflection Principle From the binomial expansion of g(z)=(z−x0)nfor integral nit is easy to see that the complexconjugateof thefunction gis thefunctionof thecomplexconjugatefor real x0: g∗(z)=bracketleftbig (z−x0)nbracketrightbig∗=(z∗−x0)n=g(z∗). (6.58) 432 Chapter 6 Functions of a Complex Variable I FIGURE 6.13Schwarzreflection. ThisleadsustotheSchwarzreflectionprinciple: If a function f(z)is(1)analytic over some region including the real axis and (2)realwhen zis real,then f∗(z)=f(z∗). (6.59) (SeeFig.6.13.) Expanding f(z)aboutsome(nonsingular)point x0ontherealaxis, f(z)=∞summationdisplay n=0(z−x0)nf(n)(x0) n!(6.60) by Eq. (6.56). Since f(z)is analytic at z=x0, this Taylor expansion exists. Since f(z)is realwhen zisreal,f(n)(x0)mustberealforall n.ThenwhenweuseEq.(6.58),Eq.(6.59), theSchwarzreflectionprinciple,followsimmediately.Exercise6.5.6isanotherformofthis principle. This completes the proof within a circle of convergence. Analytic continuation thenpermitsextendingthisresulttotheentireregionofanalyticity. Analytic Continuation Itisnaturaltothinkofthevalues f(z)ofananalyticfunction fasasingleentity,whichis usuallydefinedinsomerestrictedregion S1ofthecomplexplane,forexample,byaTaylor series(seeFig.6.14).Then fisanalyticinsidethe circleofconvergence C1,whoseradius is given by the distance r1from the center of C1to thenearest singularity offatz1(in Fig.6.14).Asingularityisanypointwhere fisnotanalytic.Ifwechooseapointinside C1 6.5 Laurent Expansion 433 FIGURE 6.14Analyticcontinuation. thatisfartherthan r1fromthesingularity z1andmakeaTaylorexpansionof faboutit(z2 inFig.6.14),thenthecircleofconvergence, C2willusuallyextendbeyondthefirstcircle, C1. In the overlap region of both circles, C1,C2, the function fis uniquely defined. In the region of the circle C2that extends beyond C1,f(z)is uniquely defined by the Taylor series about the center of C2and is analytic there, although the Taylor series about the centerof C1isnolongerconvergentthere.AfterWeierstrassthisprocessiscalled analytic continuation . It defines the analytic functions in terms of its original definition (in C1, say)andallitscontinuations. Aspecificexampleisthefunction f(z)=1 1+z, (6.61) which has a (simple) pole at z=−1 and is analytic elsewhere. The geometric series ex- pansion 1 1+z=1−z+z2+···=∞summationdisplay n=0(−z)n(6.62) convergesfor|z|<1,thatis, insidethecircle C1inFig.6.14. Supposeweexpand f(z)aboutz=i,s o f(z)=1 1+z=1 1+i+(z−i)=1 (1+i)(1+(z−i)/(1+i)) =bracketleftbigg 1−z−i 1+i+(z−i)2 (1+i)2−···bracketrightbigg1 1+i(6.63) converges for|z−i|<|1+i|=√ 2. Our circle of convergence is C2in Fig. 6.14. Now 434 Chapter 6 Functions of a Complex Variable I FIGURE 6.15|z′−z0|C1>|z−z0|;|z′−z0|C2<|z−z0|. f(z)is defined by the expansion (6.63) in S2, which overlaps S1and extends further out in the complex plane.12This extension is an analytic continuation, and when we have only isolated singular points to contend with, the function can be extended indefinitely. Equations(6.61),(6.62),and(6.63)arethreedifferentrepresentationsofthesamefunction. Each representation has its own domain of convergence. Equation (6.62) is a Maclaurin series.Equation(6.63)isaTaylorexpansionabout z=iandfromthefollowingparagraphs Eq.(6.61) isseentobeaone-termLaurentseries. Analytic continuation may take many forms, and the series expansion just considered is not necessarily the most convenient technique. As an alternate technique we shall use a functional relation in Section 8.1 to extend the factorial function around the isolated sin- gular points z=−n,n=1,2,3,....As another example, the hypergeometric equation is satisfied by the hypergeometric function defined by the series, Eq. (13.115), for |z|<1. The integral representation given in Exercise 13.4.7 permits a continuation into the com- plexplane. 12One of the most powerful and beautiful results of the more abstract theory of functions of a complex variable is that if two analytic functions coincide in any region, such as the overlap of S1andS2, or coincide on any line segment, they are the same function, in the sense that they will coincide everywhere as long as they are both well defined. In this case the agreement of the expansions (Eqs. (6.62) and (6.63)) over the region common to S1andS2would establish the identity of the functions these expansions represent. Then Eq. (6.63) would represent an analytic continuation or extension of f(z)into regions not covered by Eq. (6.62). We could equally well say that f(z)=1/(1+z)is itself an analytic continuation of either of the series given by Eqs. (6.62) and (6.63). 6.5 Laurent Expansion 435 Laurent Series Wefrequentlyencounterfunctionsthatareanalyticandsingle-valuedinanannularregion, say, of inner radius rand outer radius R, as shown in Fig. 6.15. Drawing an imaginary contour line to convert our region into a simply connected region, we apply Cauchy’s integralformula,andfortwocircles C2andC1centeredat z=z0andwithradii r2andr1, respectively,where r<r2<r1<R,weha v e13 f(z)=1 2πicontintegraldisplay C1f(z′)dz′ z′−z−1 2πicontintegraldisplay C2f(z′)dz′ z′−z. (6.64) Note that in Eq. (6.64) an explicit minus sign has been introduced so that the contour C2(likeC1) is to be traversed in the positive (counterclockwise) sense. The treatment of Eq. (6.64) now proceeds exactly like that of Eq. (6.53) in the development of the Taylor series. Each denominator is written as (z′−z0)−(z−z0)and expanded by the binomial theorem,whichnowfollowsfromtheTaylorseries(Eq. (6.57)). Notingthatfor C1,|z′−z0|>|z−z0|whilefor C2,|z′−z0|<|z−z0|,wefind f(z)=1 2πi∞summationdisplay n=0(z−z0)ncontintegraldisplay C1f(z′)dz′ (z′−z0)n+1 +1 2πi∞summationdisplay n=1(z−z0)−ncontintegraldisplay C2(z′−z0)n−1f(z′)dz′. (6.65) The minus sign of Eq. (6.64) has been absorbed by the binomial expansion. Labeling the firstseries S1andthesecond S2wehave S1=1 2πi∞summationdisplay n=0(z−z0)ncontintegraldisplay C1f(z′)dz′ (z′−z0)n+1, (6.66) which is the regular Taylor expansion, convergent for |z−z0|<|z′−z0|=r1, that is, for allzinteriortothelargercircle, C1.Forthesecondseries inEq. (6.65)wehave S2=1 2πi∞summationdisplay n=1(z−z0)−ncontintegraldisplay C2(z′−z0)n−1f(z′)dz′, (6.67) convergentfor |z−z0|>|z′−z0|=r2,thatis,forall zexteriortothesmallercircle, C2. Remember, C2nowgoescounterclockwise. Thesetwoseries arecombinedintooneseries14(a Laurentseries)by f(z)=∞summationdisplay n=−∞an(z−z0)n, (6.68) 13We may take r2arbitrarily close to randr1arbitrarily closeto R, maximizing the areaenclosedbetween C1andC2. 14Replacenby−ninS2and add. 436 Chapter 6 Functions of a Complex Variable I where an=1 2πicontintegraldisplay Cf(z′)dz′ (z′−z0)n+1. (6.69) Since, in Eq. (6.69), convergence of a binomial expansion is no longer a problem, Cmay beanycontourwithintheannularregion r<|z−z0|<Rencircling z0onceinacounter- clockwisesense. If weassumethatsuchanannularregionofconvergencedoesexist,then Eq.(6.68) istheLaurentseries, orLaurentexpansion,of f(z). The use of the contour line (Fig. 6.15) is convenient in converting the annular region into a simply connected region. Since our function is analytic in this annular region (and single-valued), the contour line is not essential and, indeed, does not appear in the final result,Eq. (6.69). Laurent series coefficients need not come from evaluation of contour integrals (which may be very intractable). Other techniques, such as ordinary series expansions, may pro- videthecoefficients. Numerous examples of Laurent series appear in Chapter 7. We limit ourselves here to onesimpleexampletoillustratetheapplicationof Eq.(6.68). Example 6.5.1 LAURENT EXPANSION Letf(z)=[z(z−1)]−1. If we choose z0=0, thenr=0 andR=1,f(z)diverging at z=1.ApartialfractionexpansionyieldstheLaurentseries 1 z(z−1)=−1 1−z−1 z=−1 z−1−z−z2−z3−···=−∞summationdisplay n=−1zn.(6.70) FromEqs. (6.70), (6.68), and(6.69)wethenhave an=1 2πicontintegraldisplaydz′ (z′)n+2(z′−1)=braceleftBigg −1forn≥−1, 0forn<−1.(6.71) The integrals in Eq. (6.71) can also be directly evaluated by substituting the geometric- seriesexpansionof (1−z′)−1usedalreadyinEq. (6.70) for (1−z)−1: an=−1 2πicontintegraldisplay∞summationdisplay m=0(z′)mdz′ (z′)n+2. (6.72) Uponinterchangingtheorderofsummationandintegration(uniformlyconvergentseries), wehave an=−1 2πi∞summationdisplay m=0contintegraldisplaydz′ (z′)n+2−m. (6.73) 6.5 Laurent Expansion 437 If weemploythepolarform, asinEq. (6.47)(or compareExercise6.4.1), an=−1 2πi∞summationdisplay m=0contintegraldisplayrieiθdθ rn+2−mei(n+2−m)θ =−1 2πi·2πi∞summationdisplay m=0δn+2−m,1, (6.74) whichagreeswithEq. (6.71). /squaresolid The Laurent series differs from the Taylor series by the obvious feature of negative powersof (z−z0).ForthisreasontheLaurentserieswillalwaysdivergeatleastat z=z0 andperhapsasfar outas somedistance r(Fig.6.15). Exercises 6.5.1 DeveloptheTaylorexpansionof ln (1+z). ANS.∞summationdisplay n=1(−1)n−1zn n. 6.5.2 Derivethebinomialexpansion (1+z)m=1+mz+m(m−1) 1·2z2+···=∞summationdisplay n=0parenleftbiggm nparenrightbigg zn formanyrealnumber.Theexpansionis convergentfor |z|<1.Why? 6.5.3 A function f(z)is analytic on and within the unit circle. Also, |f(z)|<1f o r|z|≤1 andf(0)=0.Showthat|f(z)|<|z|for|z|≤1. Hint. One approach is to show that f(z)/zis analytic and then to express [f(z0)/z0]n by the Cauchy integral formula. Finally, consider absolute magnitudes and take the nth root.This exerciseissometimescalledSchwarz’stheorem. 6.5.4 Iff(z)is a real function of the complex variable z=x+iy,t h a ti s ,i f f(x)=f∗(x), and the Laurent expansion about the origin, f(z)=summationtextanzn, hasan=0f o rn<−N, showthatallof thecoefficients anare real. Hint.Showthat zNf(z)is analytic(viaMorera’s theorem,Section6.4). 6.5.5 Afunction f(z)=u(x,y)+iv(x,y)satisfiestheconditionsfortheSchwarzreflection principle.Showthat (a)uis anevenfunctionof y.(b)vis anoddfunctionof y. 6.5.6 A function f(z)can be expanded in a Laurent series about the origin with the coeffi- cientsanreal.Showthatthecomplexconjugateofthisfunctionof zisthesamefunction ofthecomplexconjugateof z;thatis, f∗(z)=f(z∗). Verifythis explicitlyfor (a)f(z)=zn,naninteger, (b) f(z)=sinz. Iff(z)=iz(a1=i), showthattheforegoingstatementdoesnothold. 438 Chapter 6 Functions of a Complex Variable I 6.5.7 The function f(z)is analytic in a domain that includes the real axis. When zis real (z=x),f(x)is pureimaginary. (a) Showthat f(z∗)=−bracketleftbig f(z)bracketrightbig∗. (b) For the specific case f(z)=iz, develop the Cartesian forms of f(z),f(z∗), and f∗(z). Donotquotethegeneralresultofpart (a). 6.5.8 Developthefirst threenonzeroterms oftheLaurentexpansionof f(z)=parenleftbig ez−1parenrightbig−1 about the origin. Notice the resemblance to the Bernoulli number–generating function, Eq. (5.144)ofSection5.9. 6.5.9 Provethatthe Laurentexpansionofa givenfunctionaboutagivenpointis unique;that is, if f(z)=∞summationdisplay n=−Nan(z−z0)n=∞summationdisplay n=−Nbn(z−z0)n, showthat an=bnfor alln. Hint.UsetheCauchyintegralformula. 6.5.10 (a) Develop a Laurent expansion of f(z)=[z(z−1)]−1about the point z=1 valid for small values of |z−1|. Specify the exact range over which your expansion holds.Thisis ananalyticcontinuationofEq. (6.70). (b) DeterminetheLaurentexpansionof f(z)aboutz=1b u tf o r|z−1|large. Hint.Partialfractionthisfunctionanduse thegeometricseries. 6.5.11 (a) Given f1(z)=integraltext∞ 0e−ztdt(withtreal),showthatthedomaininwhich f1(z)exists (andisanalytic)is ℜ(z)>0. (b) Show that f2(z)=1/zequalsf1(z)overℜ(z) >0 and is therefore an analytic continuationof f1(z)overtheentire z-planeexceptfor z=0. (c) Expand 1 /zabout the point z=i. You will have f3(z)=summationtext∞ n=0an(z−i)n. What isthedomainof f3(z)? ANS.1 z=−i∞summationdisplay n=0in(z−i)n,|z−i|<1. 6.6 S INGULARITIES The Laurent expansion represents a generalization of the Taylor series in the presence of singularities. We define the point z0as anisolated singular point of the function f(z)if f(z)is notanalyticat z=z0butisanalyticatallneighboringpoints. 6.6 Singularities 439 Poles IntheLaurentexpansionof f(z)aboutz0, f(z)=∞summationdisplay m=−∞am(z−z0)m, (6.75) ifam=0f o rm<−n<0 anda−n/negationslash=0,wesaythat z0isapoleoforder n.Forinstance,if n=1, that is, if a−1/(z−z0)is the first nonvanishing term in the Laurent series, we have apoleof order1,oftencalleda simplepole. If, on the other hand, the summation continues to m=−∞, thenz0is a pole of infi- nite order and is called an essential singularity . These essential singularities have many pathological features. For instance, we can show that in any small neighborhood of an essential singularity of f(z)the function f(z)comes arbitrarily close to any (and there- fore every) preselected complex quantity w0.15Here, the entire w-plane is mapped by finto the neighborhood of the point z0. One point of fundamental difference between a pole of finite order nand an essential singularity is that by multiplying f(z)by(z−z0)n, f(z)(z−z0)nis no longer singular at z0. This obviously cannot be done for an essential singularity. The behaviorof f(z)asz→∞is definedin terms of thebehaviorof f(1/t)ast→0. Considerthefunction sinz=∞summationdisplay n=0(−1)nz2n+1 (2n+1)!. (6.76) Asz→∞, wereplacethe zby 1/ttoobtain sinparenleftbigg1 tparenrightbigg =∞summationdisplay n=0(−1)n (2n+1)!t2n+1. (6.77) Fromthedefinition, sin zhas anessentialsingularityatinfinity.This resultcouldbeantic- ipatedfrom Exercise6.1.9since sinz=siniy=isinhy,whenx=0, which approaches infinity exponentially as y→∞. Thus, although the absolute value of sinxfor realxis equaltoor lessthanunity,theabsolutevalueof sin zisnotbounded. Afunctionthatisanalyticthroughoutthefinitecomplexplane exceptfor isolatedpoles is called meromorphic , such as ratios of two polynomials or tan z, cotz. Examples are alsoentirefunctionsthathavenosingularitiesinthefinitecomplexplane,suchas exp (z), sinz, cosz(seeSections5.9, 5.11). 15This theorem is due to Picard. A proof is given by E. C. Titchmarsh, The Theory of Functions , 2nd ed. New York: Oxford University Press (1939). 440 Chapter 6 Functions of a Complex Variable I Branch Points Thereisanothersortof singularitythatwillbeimportantinChapter7. Consider f(z)=za, inwhich aisnotaninteger.16Aszmovesaroundtheunitcirclefrom e0toe2πi, f(z)→e2πai/negationslash=e0·a=1, for nonintegral a. We have a branch point at the origin and another at infinity. If we set z=1/t,a similar analysis of f(z)fort→0 shows that t=0; that is, z=∞is also a branch point. The points e0iande2πiin thez-plane coincide, but these coincident points lead to different values off(z); that is, f(z)is amultivalued function . The problem is resolved by constructing a cut line joining both branch points so thatf(z)will be uniquely specified for a given point in the z-plane. For za,the cut line can go out at any angle. Note that the point at infinity must be included here; that is, the cut line may join finitebranchpointsviathepointatinfinity.Thenextexampleisacaseinpoint.If a=p/q isarationalnumber,then qiscalledtheorderofthebranchpoint,becauseoneneedstogo aroundthebranchpoint qtimesbeforecomingbacktothestartingpoint.If aisirrational, thentheorderof thebranchpointisinfinite,justasfor thelogarithm. Note that a function with a branch point and a required cut line will not be continuous acrossthecutline.Oftentherewillbeaphasedifferenceonoppositesidesofthiscutline. Hencelineintegralsonoppositesidesofthisbranchpointcutlinewillnotgenerallycancel eachother.Numerousexamplesof thiscaseappearintheexercises. The contour line used to convert a multiply connected region into a simply connected region(Section6.3)iscompletelydifferent.Ourfunctioniscontinuousacrossthatcontour line,andnophasedifferenceexists. Example 6.6.1 BRANCH POINTS OF ORDER 2 Considerthefunction f(z)=parenleftbig z2−1parenrightbig1/2=(z+1)1/2(z−1)1/2. (6.78) Thefirstfactorontheright-handside, (z+1)1/2,hasabranchpointat z=−1.Thesecond factor has a branch point at z=+1. At infinity f(z)has a simple pole. This is best seen bysubstituting z=1/tandmakingabinomialexpansionat t=0: parenleftbig z2−1parenrightbig1/2=1 tparenleftbig 1−t2parenrightbig1/2=1 t∞summationdisplay n=0parenleftbigg1/2 nparenrightbigg (−1)nt2n=1 t−1 2t−1 8t3+···. Thecutlinehastoconnectbothbranchpoints,soitisnotpossibletoencircleeitherbranch pointcompletely.Tocheckonthepossibilityoftakingthelinesegmentjoining z=+1and 16z=0 is a singular point, for zahas only a finite number of derivatives, whereas an analytic function is guaranteed an infinite numberofderivatives(Section6.4).Theproblemisthat f(z)isnotsingle-valuedasweencircletheorigin.TheCauchyintegral formula may not be applied. 6.6 Singularities 441 FIGURE 6.16Branchcutandphasesof Table6.1. Table 6.1 PhaseAngle Point θϕθ+ϕ 2 10 00 20 ππ 2 30 ππ 2 4 ππ π 52 ππ3π 2 62 ππ3π 2 72 π 2π 2π z=−1 as a cut line, let us follow the phases of these two factors as we move along the contourshowninFig.6.16. For convenience in following the changes of phase let z+1=reiθandz−1=ρeiϕ. Thenthephaseof f(z)is(θ+ϕ)/2.Westartatpoint1,whereboth z+1andz−1ha v ea phaseofzero.Movingfrompoint1topoint2, ϕ,thephaseof z−1=ρeiϕ,increasesby π. (z−1becomesnegative.) ϕthenstaysconstantuntilthecircleiscompleted,movingfrom 6to7.θ,thephaseof z+1=reiθ,showsasimilarbehavior,increasingby2 πaswemove from 3 to 5. The phase of the function f(z)=(z+1)1/2(z−1)1/2=r1/2ρ1/2ei(θ+ϕ)/2is (θ+ϕ)/2.This istabulatedinthefinalcolumnofTable6.1. Twofeaturesemerge: 1. The phase at points 5 and 6 is not the same as the phase at points 2 and 3. This behaviorcanbeexpectedatabranchcut. 2.Thephaseatpoint7exceedsthatatpoint1by2 π,andthefunction f(z)=(z2−1)1/2 istherefore single-valued for thecontourshown,encircling bothbranchpoints. Ifwetakethe x-axis,−1≤x≤1,asacutline, f(z)isuniquelyspecified.Alternatively, thepositive x-axisforx>1 andthenegative x-axisforx<−1 maybetakenascutlines. The branch points cannot be encircled, and the function remains single-valued. These two cutlinesare, infact, onebranchcutfrom −1t o+1 viathepointatinfinity. /squaresolid Generalizingfrom thisexample,wehavethatthephaseofafunction f(z)=f1(z)·f2(z)·f3(z)··· isthealgebraicsumofthephaseofitsindividualfactors: argf(z)=argf1(z)+argf2(z)+argf3(z)+···. 442 Chapter 6 Functions of a Complex Variable I Thephaseofanindividualfactormaybetakenasthearctangentoftheratioofitsimaginary parttoitsrealpart(choosingtheappropriatebranchofthearctanfunctiontan−1y/x,which hasinfinitelymanybranches), argfi(z)=tan−1parenleftbiggvi uiparenrightbigg . Forthecaseof afactorof theform fi(z)=(z−z0), the phase corresponds to the phase angle of a two-dimensional vector from +z0toz,t h e phase increasing by 2 πas the point+z0is encircled. Conversely, the traversal of any closedloopnotencircling z0doesnotchangethephaseof z−z0. Exercises 6.6.1 The function f(z)expanded in a Laurent series exhibits a pole of order matz=z0. Showthatthecoefficientof (z−z0)−1,a−1, isgivenby a−1=1 (m−1)!dm−1 dzm−1bracketleftbig (z−z0)mf(z)bracketrightbig z=z0, with a−1=bracketleftbig (z−z0)f(z)bracketrightbig z=z0, when the pole is a simple pole (m=1). These equations for a−1are extremely useful indeterminingtheresiduetobeusedintheresiduetheoremofSection7.1. Hint. The technique that was so successful in proving the uniqueness of power series, Section5.7,willworkherealso. 6.6.2 Afunction f(z)canberepresentedby f(z)=f1(z) f2(z), in which f1(z)andf2(z)are analytic. The denominator, f2(z), vanishes at z=z0, showing that f(z)has a pole at z=z0. However, f1(z0)/negationslash=0,f′ 2(z0)/negationslash=0. Show that a−1, thecoefficientof (z−z0)−1ina Laurentexpansionof f(z)atz=z0,is givenby a−1=f1(z0) f′ 2(z0). (This resultleadstotheHeavisideexpansiontheorem,Exercise15.12.11.) 6.6.3 InanalogywithExample6.6.1,considerindetailthephaseofeachfactorandtheresul- tantoverallphase of f(z)=(z2+1)1/2followinga contoursimilar to thatof Fig.6.16 butencirclingthenewbranchpoints. 6.6.4 The Legendre function of the second kind, Qν(z), has branch points at z=±1. The branchpointsarejoinedbyacutlinealongthereal (x)axis. 6.7 Mapping 443 (a) Showthat Q0(z)=1 2ln((z+1)/(z−1))issingle-valued(withtherealaxis −1≤ x≤1 takenasacutline). (b) Forrealargument xand|x|<1 it isconvenienttotake Q0(x)=1 2ln1+x 1−x. Showthat Q0(x)=1 2bracketleftbig Q0(x+i0)+Q0(x−i0)bracketrightbig . Herex+i0 indicatesthat zapproachesthereal axis from above,and x−i0 indi- catesanapproachfrom below. 6.6.5 As an example of an essential singularity, consider e1/zaszapproaches zero. For any complexnumber z0,z0/negationslash=0,showthat e1/z=z0 hasaninfinitenumberofsolutions. 6.7 M APPING In the preceding sections we have defined analytic functions and developed some of their mainfeatures.Hereweintroducesomeofthemoregeometricaspectsoffunctionsofcom- plexvariables,aspectsthatwillbeusefulinvisualizingtheintegraloperationsinChapter7 and that are valuable in their own right in solving Laplace’s equation in two-dimensional systems. In ordinary analytic geometry we may take y=f(x)and then plot yversusx.O u r problemhereismorecomplicated,for zisafunctionoftwovariables, xandy.W eusethe notation w=f(z)=u(x,y)+iv(x,y). (6.79) Then for a point in the z-plane (specific values for xandy) there may correspond specific values for u(x,y)andv(x,y)that then yield a point in the w-plane. As points in the z-plane transform, or are mapped into points in the w-plane, lines or areas in the z-plane will be mapped into lines or areas in the w-plane. Our immediate purpose is to see how linesandareasmapfromthe z-planetothe w-planefor anumberofsimplefunctions. Translation w=z+z0. (6.80) The function wis equal to the variable zplus a constant, z0=x0+iy0. By Eqs. (6.1) and (6.79), u=x+x0,v=y+y0, (6.81) representingapuretranslationofthecoordinateaxes,asshowninFig.6.17. 444 Chapter 6 Functions of a Complex Variable I FIGURE 6.17Translation. Rotation w=zz0. (6.82) Hereit isconvenienttoreturntothepolarrepresentation,using w=ρeiϕ,z=reiθ,andz0=r0eiθ0, (6.83) then ρeiϕ=rr0ei(θ+θ0), (6.84) or ρ=rr0,ϕ=θ+θ0. (6.85) Two things have occurred. First, the modulus rhas been modified, either expanded or contracted, by the factor r0. Second, the argument θhas been increased by the additive constantθ0(Fig.6.18).Thisrepresentsarotationofthecomplexvariablethroughanangle θ0. Forthespecialcaseof z0=i, wehaveapurerotationthrough π/2 radians. FIGURE 6.18Rotation. 6.7 Mapping 445 Inversion w=1 z. (6.86) Again,usingthepolarform,wehave ρeiϕ=1 reiθ=1 re−iθ, (6.87) whichshowsthat ρ=1 r,ϕ=−θ. (6.88) The first part of Eq. (6.87) shows that inversion clearly. The interior of the unit circle is mapped onto the exterior and vice versa (Fig. 6.19). In addition, the second part of Eq. (6.87) shows that the polar angle is reversed in sign. Equation (6.88) therefore also involvesareflectionof the y-axis,exactlylikethecomplexconjugateequation. To see how curves in the z-plane transform into the w-plane, we return to the Cartesian form: u+iv=1 x+iy. (6.89) Rationalizing the right-hand side by multiplying numerator and denominator by z∗and thenequatingtherealparts andtheimaginaryparts,wehave u=x x2+y2,x=u u2+v2, v=−y x2+y2,y=−v u2+v2.(6.90) FIGURE 6.19Inversion. 446 Chapter 6 Functions of a Complex Variable I Acirclecenteredattheorigininthe z-planehastheform x2+y2=r2(6.91) andbyEqs. (6.90) transformsinto u2 (u2+v2)2+v2 (u2+v2)2=r2. (6.92) SimplifyingEq.(6.92), weobtain u2+v2=1 r2=ρ2, (6.93) whichdescribesacircleinthe w-planealsocenteredattheorigin. Thehorizontalline y=c1transforms into −v u2+v2=c1, (6.94) or u2+parenleftbigg v+1 2c1parenrightbigg2 =1 (2c1)2, (6.95) which describes a circle in the w-plane of radius (1/2c1)and centered at u=0,v=−1 2c1(Fig.6.20). We pick up the other three possibilities, x=±c1,y=−c1, by rotating the xy-axes. In general, any straight line or circle in the z-plane will transform into a straight line or a circleinthe w-plane(compareExercise6.7.1). FIGURE 6.20Inversion,line ↔circle. 6.7 Mapping 447 Branch Points and Multivalent Functions The three transformations just discussed have all involved one-to-one correspondence of points in the z-plane to points in the w-plane. Now to illustrate the variety of transfor- mations that are possible and the problems that can arise, we introduce first a two-to-one correspondence and then a many-to-one correspondence. Finally, we take up the inverses ofthesetwotransformations. Considerfirst thetransformation w=z2, (6.96) whichleadsto ρ=r2,ϕ=2θ. (6.97) Clearly, our transformation is nonlinear, for the modulus is squared, but the significant featureofEq. (6.96)is thatthephaseangleorargumentis doubled.Thismeansthatthe •firstquadrantof z,0≤θ<π 2,→upperhalf-planeof w,0≤ϕ<π, •upperhalf-planeof z,0≤θ<π,→wholeplaneof w,0≤ϕ<2π. The lower half-plane of zmaps into the already covered entire plane of w, thus covering thew-plane a second time. This is our two-to-one correspondence, that is, two distinct pointsinthe z-plane,z0andz0eiπ=−z0, correspondingtothesinglepoint w=z2 0. In Cartesianrepresentation, u+iv=(x+iy)2=x2−y2+i2xy, (6.98) leadingto u=x2−y2,v=2xy. (6.99) Hence the lines u=c1,v=c2in thew-plane correspond to x2−y2=c1,2xy=c2, rec- tangular (and orthogonal) hyperbolas in the z-plane (Fig. 6.21). To every point on the hyperbola x2−y2=c1in the right half-plane, x>0, one point on the line u=c1corre- sponds,andviceversa.However,everypointontheline u=c1alsocorrespondstoapoint onthehyperbola x2−y2=c1inthelefthalf-plane, x<0,asalreadyexplained. It will be shown in Section 6.8 that if lines in the w-plane are orthogonal, the corre- spondinglinesinthe z-planearealsoorthogonal,aslongasthetransformationisanalytic. Sinceu=c1andv=c2are constructed perpendicular to each other, the corresponding hyperbolasinthe z-planeareorthogonal.Wehaveconstructedaneworthogonalsystemof hyperbolic lines (or surfaces if we add an axis perpendicular to xandy). Exercise 2.1.3 was an analysis of this system. It might be noted that if the hyperbolic lines are electric or magnetic lines of force, then we have a quadrupole lens useful in focusing beams of high-energyparticles. Theinverseofthefourthtransformation(Eq. (6.96)) is w=z1/2. (6.100) 448 Chapter 6 Functions of a Complex Variable I FIGURE 6.21Mapping—hyperboliccoordinates. Fromtherelation ρeiϕ=r1/2eiθ/2(6.101) and 2ϕ=θ, (6.102) we now have two points in the w-plane (arguments ϕandϕ+π) corresponding to one point in the z-plane (except for the point z=0). Or, to put it another way, θandθ+2π correspondto ϕandϕ+π,twodistinctpointsinthe w-plane.Thisisthecomplexvariable analog of the simple real variable equation y2=x, in which two values of y, plus and minus,correspondtoeachvalueof x. The important point here is that we can make the function wof Eq. (6.100) a single- valuedfunctioninsteadofadouble-valuedfunctionifweagreetorestrict θtoarangesuch as 0≤θ<2π. This may be done by agreeing never to cross the line θ=0i nt h ez-plane (Fig.6.22).Suchalineofdemarcationiscalleda cutlineorbranchcut .Notethatbranch pointsoccurinpairs. Thecut line joins the two branch point singularities , here at 0 and ∞(for the latter, transform z=1/tfort→0). Any line from z=0 to infinity would serve equally well. The purpose of the cut line is to restrict the argument of z. The points zandzexp(2πi) coincide in the z-plane but yield different points wand−w=wexp(πi)in thew-plane. Henceintheabsenceofacutline,thefunction w=z1/2isambiguous.Alternatively,since the function w=z1/2is double-valued, we can also glue two sheets of the complex z- plane together along the branch cut so that arg (z)increases beyond 2 πalong the branch cut and continues from 4 πon the second sheet to reach the same function values for z as forze−4πi,that is, the start on the first sheet again. This construction is called the Riemann surface ofw=z1/2. We shall encounter branch points and cut lines (branch cuts)frequentlyinChapter7. Thetransformation w=ez(6.103) leadsto ρeiϕ=ex+iy, (6.104) 6.7 Mapping 449 FIGURE 6.22Acutline. or ρ=ex,ϕ=y. (6.105) Ifyranges from 0 ≤y<2π(or−π<y≤π), thenϕcovers the same range. But this is thewhole w-plane.Inotherwords, ahorizontalstripinthe z-planeofwidth 2 πmapsinto theentire w-plane.Further,anypoint x+i(y+2nπ),inwhich nisanyinteger,mapsinto the same point (by Eq. (6.104)) in the w-plane. We have a many-(infinitely many)-to-one correspondence. Finally,as theinverseofthefifthtransformation(Eq. (6.103)), wehave w=lnz. (6.106) Byexpandingit,weobtain u+iv=lnreiθ=lnr+iθ. (6.107) Foragivenpoint z0inthez-planetheargument θisunspecifiedwithinanintegralmultiple of 2π.This meansthat v=θ+2nπ, (6.108) and, as in the exponential transformation, we have an infinitely many-to-one correspon- dence. Equation (6.108) has a nice physical representation. If we go around the unit circle in thez-plane,r=1,andbyEq.(6.107), u=lnr=0;butv=θ,andθissteadilyincreasing andcontinuestoincreaseas θcontinuespast 2 π. The cut line joins the branch point at the origin with infinity. As θincreases past 2 π we glue a new sheet of the complex z-plane along the cut line, etc. Going around the unit circle in the z-plane is like the advance of a screw as it is rotated or the ascent of a person walkingupaspiralstaircase(Fig.6.23), whichis the Riemannsurface ofw=lnz. As in the preceding example, we can also make the correspondence unique (and Eq. (6.106) unambiguous) by restricting θto a range such as 0 ≤θ<2πby taking the 450 Chapter 6 Functions of a Complex Variable I FIGURE 6.23This istheRiemann surface for ln z, amultivalued function. lineθ=0 (positive real axis) as a cut line. This is equivalent to taking one and only one completeturnof thespiralstaircase. The concept of mapping is a very broad and useful one in mathematics. Our mapping from a complex z-plane to a complex w-plane is a simple generalization of one definition of function: a mapping of x(from one set) into yin a second set. A more sophisticated form of mapping appears in Section 1.15 where we use the Dirac delta function δ(x−a) tomapafunction f(x)intoitsvalueatthepoint a.TheninChapter15integraltransforms are used to map one function f(x)inx-space into a second (related) function F(t)in t-space. Exercises 6.7.1 Howdocirclescenteredontheorigininthe z-planetransform for (a)w1(z)=z+1 z,(b)w2(z)=z−1 z,forz/negationslash=0? Whathappenswhen |z|→1? 6.7.2 Whatpartof the z-planecorrespondstotheinterioroftheunitcircleinthe w-planeif (a)w=z−1 z+1,(b)w=z−i z+i? 6.7.3 Discussthetransformations (a)w(z)=sinz,(c)w(z)=sinhz, (b)w(z)=cosz,(d)w(z)=coshz. Show how the lines x=c1,y=c2m a pi n t ot h e w-plane. Note that the last three trans- formationscanbeobtainedfromthefirstonebyappropriatetranslationand/orrotation. 6.8 Conformal Mapping 451 FIGURE 6.24Besselfunctionintegrationcontour. 6.7.4 Showthatthefunction w(z)=parenleftbig z2−1parenrightbig1/2 issingle-valuedif wetake −1≤x≤1,y=0 asacutline. 6.7.5 Show that negative numbers have logarithms in the complex plane. In particular, find ln(−1). ANS. ln(−1)=iπ. 6.7.6 An integral representation of the Bessel function follows the contour in the t-plane shown in Fig. 6.24. Map this contour into the θ-plane with t=eθ. Many additional examplesof mappingare giveninChapters11,12,and13. 6.7.7 For noninteger m, show that the binomial expansion of Exercise 6.5.2 holds only for a suitably defined branch of the function (1+z)m. Show how the z-plane is cut. Explain why|z|<1 maybetakenasthecircleofconvergencefor theexpansionofthis branch, inlightof thecutyouhavechosen. 6.7.8 The Taylor expansion of Exercises 6.5.2 and 6.7.7 is notsuitable for branches other than the one suitably defined branch of the function (1+z)mfor noninteger m.[ N o t e that other branches cannot have the same Taylor expansion since they must be distin- guishable.] Using the same branch cut of the earlier exercises for all other branches, find the corresponding Taylor expansions, detailing the phase assignments and Taylor coefficients. 6.8 C ONFORMAL MAPPING In Section 6.7 hyperbolas were mapped into straight lines and straight lines were mapped into circles. Yet in all these transformations one feature stayed constant. This constancy wasaresultof thefactthatallthetransformationsofSection6.7wereanalytic. Aslongas w=f(z)is ananalyticfunction,wehave df dz=dw dz=lim /Delta1z→0/Delta1w /Delta1z. (6.109) 452 Chapter 6 Functions of a Complex Variable I FIGURE 6.25Conformalmapping—preservationofangles. Assuming that this equation is in polar form, we may equate modulus to modulus and argumenttoargument.Forthelatter(assumingthat df/dz/negationslash=0), arg lim /Delta1z→0/Delta1w /Delta1z=lim /Delta1z→0arg/Delta1w /Delta1z =lim /Delta1z→0arg/Delta1w−lim /Delta1z→0arg/Delta1z=argdf dz=α, (6.110) whereα,theargumentofthederivative,maydependon zbutisaconstantforafixed z,in- dependentofthedirectionofapproach.Toseethesignificanceofthis,considertwocurves Czin thez-plane and the corresponding curve Cwin thew-plane (Fig. 6.25). The incre- ment/Delta1zis shown at an angle of θrelative to the real (x)axis, whereas the corresponding increment /Delta1wforms anangleof ϕwiththereal (u)axis. FromEq. (6.110), ϕ=θ+α, (6.111) or any line in the z-plane is rotated through an angle αin thew-plane as long as wis an analytictransformationandthederivativeisnotzero.17 Since this result holds for any line through z0, it will hold for a pair of lines. Then for theanglebetweenthesetwolines, ϕ2−ϕ1=(θ2+α)−(θ1+α)=θ2−θ1, (6.112) which shows that the included angle is preserved under an analytic transformation. Such angle-preserving transformations are called conformal . The rotation angle αwill, in gen- eral,dependon z. In addition|f′(z)|willusuallybeafunctionof z. Historically,theseconformaltransformationshavebeenofgreatimportancetoscientists and engineers in solving Laplace’s equation for problems of electrostatics, hydrodynam- ics, heat flow, and so on. Unfortunately, the conformal transformation approach, however elegant,islimitedtoproblemsthatcanbereducedtotwodimensions.Themethodisoften beautiful if there is a high degree of symmetry present but often impossible if the sym- metry is broken or absent. Because of these limitations and primarily because electronic computers offer a useful alternative (iterative solution of the partial differential equation), thedetailsandapplicationsofconformalmappingsareomitted. 17Ifdf/dz=0, its argument orphase is undefined andthe (analytic) transformation will not necessarilypreserve angles. 6.8 Additional Readings 453 Exercises 6.8.1 Expandw(x)in a Taylor series about the point z=z0, wheref′(z0)=0. (Angles are not preserved.) Show that if the first n−1 derivatives vanish but f(n)(z0)/negationslash=0, then anglesinthe z-planewithverticesat z=z0appearinthe w-planemultipliedby n. 6.8.2 Developthetransformationsthatcreateeachofthefourcylindricalcoordinatesystems: (a) Circularcylindrical: x=ρcosϕ, y=ρsinϕ. (b) Ellipticcylindrical: x=acoshucosv, y=asinhusinv. (c) Paraboliccylindrical: x=ξη, y=1 2parenleftbig η2−ξ2parenrightbig . (d) Bipolar: x=asinhη coshη−cosξ, y=asinξ coshη−cosξ. Note.Thesetransformationsarenotnecessarilyanalytic. 6.8.3 Inthetransformation ez=a−w a+w, howdothecoordinatelinesinthe z-planetransform?Whatcoordinatesystemhaveyou constructed? AdditionalReadings Ahlfors, L. V., Complex Analysis , 3rd ed. New York: McGraw-Hill (1979). This text is detailed, thorough, rigor- ous, and extensive. Churchill,R.V.,J.W.Brown,andR.F.Verkey, ComplexVariablesandApplications ,5thed.NewYork:McGraw- Hill (1989). This is an excellent text for both the beginning and advanced student. It is readable and quite complete. Adetailed proof of the Cauchy–Goursat theorem is given in Chapter5. Greenleaf, F. P., Introduction to Complex Variables . Philadelphia: Saunders (1972). This very readable book has detailed,careful explanations. Kurala, A., Applied Functions of a Complex Variable . New York: Wiley (Interscience) (1972). An intermediate- level text designed for scientists andengineers. Includes many physical applications. Levinson, N., and R. M. Redheffer, Complex Variables . San Francisco: Holden-Day (1970). This text is written for scientists andengineers whoare interested in applications. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 4 is apresentation of portions of thetheory of functions of a complex variableof interest to theoretical physicists. Remmert, R., Theory of Complex Functions . NewYork: Springer (1991). Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering , 2nd ed. New York: McGraw-Hill(1966). Chapter 7 covers complex variables. Spiegel, M. R., Complex Variables . New York: McGraw-Hill (1985). An excellent summary of the theory of complex variables for scientists. T itch m ar sh ,E.C., The Theory of Functions , 2nd ed. NewYork: Oxford University Press (1958). A classic. 454 Chapter 6 Functions of a Complex Variable I Watson, G. N., Complex Integration and Cauchy’s Theorem . New York: Hafner (orig. 1917, reprinted 1960). A short work containing a rigorous development of the Cauchy integral theorem and integral formula. Appli- cationstothecalculusofresiduesareincluded. CambridgeTractsinMathematics,andMathematicalPhysics , No. 15. Otherreferencesaregivenattheendof Chapter15. CHAPTER 7 FUNCTIONS OF A COMPLEX VARIABLE II In this chapter we return to the analysis that started with the Cauchy–Riemann conditions in Chapter 6 and develop the residue theorem, with major applications to the evaluation of definite and principal part integrals of interest to scientists and asymptotic expansion of integrals by the method of steepest descent. We also develop further specific analytic functions, such as pole expansions of meromorphic functions and product expansions of entire functions. Dispersion relations are included because they represent an important applicationofcomplexvariablemethodsforphysicists. 7.1 C ALCULUS OF RESIDUES Residue Theorem If the Laurent expansion of a function f(z)=summationtext∞ n=−∞an(z−z0)nis integrated term by termbyusingaclosedcontourthatencirclesoneisolatedsingularpoint z0onceinacoun- terclockwisesense, weobtain(Exercise6.4.1) ancontintegraldisplay (z−z0)ndz=an(z−z0)n+1 n+1vextendsinglevextendsinglevextendsinglevextendsinglez1 z1=0,n/negationslash=−1. (7.1) However,if n=−1, a−1contintegraldisplay (z−z0)−1dz=a−1contintegraldisplayireiθdθ reiθ=2πia−1. (7.2) SummarizingEqs. (7.1) and(7.2), wehave 1 2πicontintegraldisplay f(z)dz=a−1. (7.3) 455 456 Chapter 7 Functions of a Complex Variable II FIGURE 7.1Excludingisolated singularities. The constant a−1,the coefficient of (z−z0)−1in the Laurent expansion, is called the residueof f(z)atz=z0. A set of isolated singularities can be handled by deforming our contour as shown in Fig.7.1.Cauchy’sintegraltheorem(Section6.3) leadsto contintegraldisplay Cf(z)dz+contintegraldisplay C0f(z)dz+contintegraldisplay C1f(z)dz+contintegraldisplay C2f(z)dz+···=0.(7.4) Thecircularintegralaroundanygivensingularpointis givenbyEq. (7.3), contintegraldisplay Cif(z)dz=−2πia−1,zi, (7.5) assuming a Laurent expansion about the singular point z=zi. The negative sign comes from the clockwise integration, as shown in Fig. 7.1. Combining Eqs. (7.4) and (7.5), we have contintegraldisplay Cf(z)dz=2πi(a−1z0+a−1z1+a−1z2+···) =2πi×(sum ofenclosedresidues ). (7.6) This is the residue theorem . The problem of evaluating one or more contour integrals is replacedbythealgebraicproblemof computingresiduesattheenclosedsingularpoints. We first use this residue theorem to develop the concept of the Cauchy principal value. Then in the remainder of this section we apply the residue theorem to a wide variety of definiteintegralsofmathematicalandphysicalinterest. Using the transformation z=1/wforwapproaching 0, we can find the nature of a sin- gularity at zgoing to∞and the residue of a function f(z)with just isolated singularities andnobranchpoints.Insuchcasesweknowthat summationdisplay {residuesinthefinite z-plane}+{residueat z→∞}= 0. 7.1 Calculus of Residues 457 Cauchy Principal Value Occasionally an isolated pole will be directly on the contour of integration, causing the integraltodiverge.Letus illustrateaphysicalcase. Example 7.1.1 FORCED CLASSICAL OSCILLATOR The inhomogeneous differential equation for a classical, undamped, driven harmonic os- cillator, ¨x(t)+ω2 0x(t)=f(t), (7.7) may be solved by representing the driving force f(t)=integraltext δ(t′−t)f(t′)dt′as a superpo- sitionofimpulsesbyanalogywithanextendedchargedistributioninelectrostatics.1Ifwe solvefirst thesimplerdifferentialequation ¨G+ω2 0G=δ(t−t′) (7.8) forG(t,t′), which is independent of the driving term f(model dependent), then x(t)=integraltext G(t,t′)f(t′)dt′solves the original problem. First, we verify this by substituting the in- tegralsfor x(t)anditstimederivativesintothedifferentialequationfor x(t)usingthedif- ferentialequationfor G.Thenwelookfor G(t,t′)=integraltext˜G(ω)eiωtdω 2πintermsofanintegral weightedby˜G,whichissuggestedbyasimilarintegralformfor δ(t−t′)=integraltext eiω(t−t′)dω 2π (seeEq. (1.193c)inSection1.15). Uponsubstituting Gand¨Gintothedifferentialequationfor G,weobtain integraldisplaybracketleftbigparenleftbig ω2 0−ω2parenrightbig˜G−e−iωt′bracketrightbig eiωtdω=0. (7.9) Because this integral is zero for all t,the expression in brackets must vanish for all ω. Thisrelationisnolongeradifferentialequationbutanalgebraicrelationthatwecansolve for˜G: ˜G(ω)=e−iωt′ ω2 0−ω2=e−iωt′ 2ω0(ω+ω0)−e−iωt′ 2ω0(ω−ω0). (7.10) Substituting˜Gintotheintegralfor Gyields G(t,t′)=1 4πω0integraldisplay∞ −∞bracketleftbiggeiω(t−t′) ω+ω0−eiω(t−t′) ω−ω0bracketrightbigg dω. (7.11) Here, the dependence of Gont−t′in the exponential is consistent with the same depen- denceofδ(t−t′),itsdrivingterm.Now,theproblemisthatthisintegraldivergesbecause theintegrandblowsupat ω=±ω0,sincetheintegrationgoesrightthroughthefirst-order poles. To explain why this happens, we note that the δ-function driving term for Gin- cludes all frequencies with the same amplitude. Next, we see that the equation for ˜Gat t′=0 has its driving term equal to unity for all frequencies ω, including the resonant ω0. 1Adapted from A.Yu. Grosberg, priv. comm. 458 Chapter 7 Functions of a Complex Variable II Weknowfromphysicsthatforcinganoscillatoratresonanceleadstoanindefinitelygrow- ing amplitude when there is no friction. With friction, the amplitude remains finite, even at resonance. This suggests includinga small friction term in the differential equationsfor x(t)andG. Withasmallfrictionterm η˙G,η>0,inthedifferentialequationfor G(t,t′)(andη˙xfor x(t)), wecanstillsolvethealgebraicequation parenleftbig ω2 0−ω2+iηωparenrightbig˜G=e−iωt′(7.12) for˜Gwithfriction. Thesolutionis ˜G=e−iωt′ ω2 0−ω2+iηω=e−iωt′ 2/Omega1parenleftbigg1 ω−ω−−1 ω−ω+parenrightbigg , (7.13) ω±=±/Omega1+iη 2,/Omega1=ω0radicalBigg 1−parenleftbiggη 2ω0parenrightbigg2 . (7.14) Forsmallfriction,0 <η≪ω0,/Omega1isnearlyequalto ω0andreal,whereas ω±eachpickup asmallimaginarypart. Thismeansthattheintegrationof theintegralfor G, G(t,t′)=1 4π/Omega1integraldisplay∞ −∞bracketleftbiggeiω(t−t′) ω−ω−−eiω(t−t′) ω−ω+bracketrightbigg dω, (7.15) nolongerencountersapoleandremainsfinite. /squaresolid This treatment of an integral with a pole moves the pole off the contour and then con- siders the limiting behavior as it is brought back, as in Example 7.1.1 for η→0.This example also suggests treating ωas a complex variable in case the singularity is a first- order pole, deforming the integration path to avoid the singularity, which is equivalent to addingasmallimaginaryparttothepoleposition,andevaluatingtheintegralbymeansof theresiduetheorem. Therefore,iftheintegrationpathofanintegralintegraltextdz z−x0forrealx0goesrightthroughthe polex0,wemaydeformthecontourtoincludeorexcludetheresidue,asdesired,byinclud- ingasemicirculardetourof infinitesimalradius .ThisisshowninFig.7.2.Theintegration overthesemicirclethengives,with z−x0=δeiϕ,dz=iδeiϕdϕ(seeEq. (6.27a)), integraldisplaydz z−x0=iintegraldisplay2π πdϕ=iπ,i.e.,πia−1 ifcounterclockwise , integraldisplaydz z−x0=iintegraldisplay0 πdϕ=−iπ,i.e.,−πia−1ifclockwise . This contribution, +or−, appears on the left-hand side of Eq. (7.6). If our detour were clockwise, the residue would not be enclosed and there would be no corresponding term ontheright-handsideof Eq. (7.6). However, if our detour were counterclockwise, this residue would be enclosed by the contourCandaterm 2 πia−1wouldappearontheright-handsideofEq. (7.6). The net result for either clockwise or counterclockwise detour is that a simple pole on the contour is counted as one-half of what it would be if it were within the contour. This correspondstotakingtheCauchyprincipalvalue. 7.1 Calculus of Residues 459 FIGURE 7.2Bypassingsingularpoints. FIGURE 7.3Closingthecontour withaninfinite-radiussemicircle. Forinstance,letus supposethat f(z)withas implepoleat z=x0isintegratedoverthe entire real axis. The contour is closed with an infinite semicircle in the upper half-plane (Fig.7.3). Then contintegraldisplay f(z)dz=integraldisplayx0−δ −∞f(x)dx+integraldisplay Cx0f(z)dz +integraldisplay∞ x0+δf(x)dx+integraldisplay Cinfinitesemicircle =2πisummationdisplay enclosed residues. (7.16) If the smallsemicircle Cx0, includes x0(bygoingbelowthe x-axis, counterclockwise), x0 is enclosed, and its contribution appears twice—asπia−1inintegraltext Cx0and as 2πia−1in the term 2πisummationtextenclosed residues—for a net contribution of πia−1. If the upper small semi- circle is selected, x0is excluded. The only contribution is from the clockwise integration overCx0, which yields −πia−1. Moving this to the extreme right of Eq. (7.16), we have +πia−1, asbefore. The integrals along the x-axis may be combined and the semicircle radius permitted to approachzero.Wethereforedefine lim δ→0braceleftbiggintegraldisplayx0−δ −∞f(x)dx+integraldisplay∞ x0+δf(x)dxbracerightbigg =Pintegraldisplay∞ −∞f(x)dx. (7.17) Pindicates the Cauchy principal value and represents the preceding limiting process. Note that the Cauchy principal value is a balancing (or canceling) process. In the vicinity ofoursingularityat z=x0, f(x)≈a−1 x−x0. (7.18) 460 Chapter 7 Functions of a Complex Variable II FIGURE 7.4Cancellationatasimplepole. Thisisodd,relativeto x0.Thesymmetricoreveninterval(relativeto x0)providescancel- lationof theshadedareas, Fig.7.4. The contributionof thesingularityis intheintegration aboutthesemicircle. In general, if a function f(x)has a singularity x0somewhere inside the interval a≤ x0≤band is integrable over every portion of this interval that does not contain the point x0,thenwedefine integraldisplayb af(x)dx=lim δ1→0integraldisplayx0−δ1 af(x)dx+lim δ2→0integraldisplayb x0+δ2f(x)dx, when the limit exists as δj→0independently , else the integral is said to diverge. If this limit does not exist but the limit δ1=δ2=δ→0 exists, it is defined to be the principal valueoftheintegral. This samelimitingtechniqueis applicabletotheintegrationlimits ±∞.We define integraldisplay∞ −∞f(x)dx=lim a→−∞,b→∞integraldisplayb af(x)dx, (7.19a) if the integral exists with a,bapproaching their limits independently, else the integral di- verges.Incasetheintegraldivergesbut lima→∞integraldisplaya −af(x)dx=Pintegraldisplay∞ −∞f(x)dx (7.19b) exist,itisdefinedas itsprincipalvalue. 7.1 Calculus of Residues 461 Pole Expansion of Meromorphic Functions Analyticfunctions f(z)thathaveonlyisolatedpolesassingularitiesarecalled meromor- phic. Examples are cot z[fromd dzlnsinzin Eq. (5.210)] and ratios of polynomials. For simplicity we assume that these poles at finite z=anwith 0<|a1|<|a2|<···are all simple with residues bn. Then an expansion of f(z)in terms of bn(z−an)−1depends in a systematic way on all singularities of f(z), in contrast to the Taylor expansion about an arbitrarily chosen analytic point z0off(z)or the Laurent expansion about one of the singularpointsof f(z). Let us consider a series of concentric circles Cnabout the origin so that Cnincludes a1,a2,...,anbutnootherpoles,itsradius Rn→∞asn→∞.Toguaranteeconvergence we assume that |f(z)|<εRnfor any small positive constant εand allzonCn. Then the series f(z)=f(0)+∞summationdisplay n=1bnbraceleftbig (z−an)−1+a−1 nbracerightbig (7.20) convergesto f(z).Toprovethis theorem (duetoMittag–Leffler)weusetheresiduetheo- remtoevaluatethecontourintegralfor zinsideCn: In=1 2πiintegraldisplay Cnf(w) w(w−z)dw =nsummationdisplay m=1bm am(am−z)+f(z)−f(0) z. (7.21) OnCnwehave,for n→∞, |In|≤2πRnmaxwonCn|f(w)| 2πRn(Rn−|z|)<εRn Rn−|z|→ε forRn≫|z|.U s i n gIn→0 inEq. (7.21)provesEq.(7.20). If|f(z)|<εRp+1 n, thenweevaluatesimilarlytheintegral In=1 2πiintegraldisplayf(w) wp+1(w−z)dw→0asn→∞ andobtaintheanalogouspoleexpansion f(z)=f(0)+zf′(0)+···+zpf(p)(0) p!+∞summationdisplay n=1bnzp+1/ap+1 n z−an.(7.22) NotethattheconvergenceoftheseriesinEqs.(7.20)and(7.22)isimpliedbytheboundof |f(z)|for|z|→∞. 462 Chapter 7 Functions of a Complex Variable II Product Expansion of Entire Functions Afunction f(z)thatisanalyticforallfinite ziscalledan entirefunction.Thelogarithmic derivative f′/fisameromorphicfunctionwithapoleexpansion. Iff(z)has a simple zero at z=an, thenf(z)=(z−an)g(z)with analytic g(z)and g(an)/negationslash=0.Hencethelogarithmicderivative f′(z) f(z)=(z−an)−1+g′(z) g(z)(7.23) has a simple pole at z=anwith residue 1, and g′/gis analytic there. If f′/fsatisfies the conditionsthatleadtothepoleexpansioninEq. (7.20), then f′(z) f(z)=f′(0) f(0)+∞summationdisplay n=1bracketleftbigg1 an+1 z−anbracketrightbigg (7.24) holds.IntegratingEq.(7.24) yields integraldisplayz 0f′(z) f(z)dz=lnf(z)−lnf(0) =zf′(0) f(0)+∞summationdisplay n=1braceleftbigg ln(z−an)−ln(−an)+z anbracerightbigg , andexponentiatingweobtaintheproductexpansion f(z)=f(0)expparenleftbiggzf′(0) f(0)parenrightbigg∞productdisplay 1parenleftbigg 1−z anparenrightbigg ez/an. (7.25) Examplesaretheproductexpansions(seeChapter5)for sinz=z∞productdisplay n=−∞ n/negationslash=0parenleftbigg 1−z nπparenrightbigg ez/nπ=z∞productdisplay n=1parenleftbigg 1−z2 n2π2parenrightbigg , cosz=∞productdisplay n=1braceleftbigg 1−z2 (n−1/2)2π2bracerightbigg .(7.26) Anotherexampleistheproductexpansionofthegammafunction,whichwillbediscussed inChapter8. AsaconsequenceofEq.(7.23)thecontourintegralofthelogarithmicderivativemaybe used to count the number Nfof zeros (including their multiplicities) of the function f(z) insidethecontour C: 1 2πiintegraldisplay Cf′(z) f(z)dz=Nf. (7.27) 7.1 Calculus of Residues 463 Moreover, using integraldisplayf′(z) f(z)dz=lnf(z)=lnvextendsinglevextendsinglef(z)vextendsinglevextendsingle+iargf(z), (7.28) weseethattherealpartinEq.(7.28)doesnotchangeas zmovesoncearoundthecontour, whilethecorrespondingchangein arg fmustbe /Delta1Carg(f)=2πNf. (7.29) This leads to Rouché’s theorem :Iff(z)andg(z)are analytic inside and on a closed contourCand|g(z)|<|f(z)|onCthenf(z)andf(z)+g(z)have the same number of zeros inside C. Toshowthisweuse 2πNf+g=/Delta1Carg(f+g)=/Delta1Carg(f)+/Delta1Cargparenleftbigg 1+g fparenrightbigg . Since|g|<|f|onC, thepoint w=1+g(z)/f(z) is always an interiorpointof the circle inthew-planewithcenterat1andradius1.Hencearg (1+g/f)mustreturntoitsoriginal valuewhen zmovesaround C(itdoesnotcircletheorigin);itcannotdecreaseorincrease byamultipleof 2 πso that/Delta1Carg(1+g/f)=0. Rouché’s theorem may be used for an alternative proof of the fundamental theorem of algebra:Apolynomialsummationtextn m=0amzmwithan/negationslash=0hasnzeros.Wedefine f(z)=anzn.Then fhas ann-fold zero at the origin and no other zeros. Let g(z)=summationtextn−1 m=0amzm. We apply Rouché’stheoremtoacircle Cwithcenterattheoriginandradius R>1.OnC,|f(z)|= |an|Rnand vextendsinglevextendsingleg(z)vextendsinglevextendsingle≤|a0|+|a1|R+···+|an−1|Rn−1≤parenleftbiggn−1summationdisplay m=0|am|parenrightbigg Rn−1. Hence|g(z)|<|f(z)|forzonC, provided R>(summationtextn−1 m=0|am|)/|an|. For all sufficiently largecircles Ctherefore, f+g=summationtextn m=0amzmhasnzerosinside CaccordingtoRouché’s theorem. Evaluation of Definite Integrals Definiteintegralsappearrepeatedlyinproblemsofmathematicalphysicsaswellasinpure mathematics. Three moderately general techniques are useful in evaluating definite inte- grals: (1) contour integration, (2) conversion to gamma or beta functions (Chapter 8), and (3) numerical quadrature. Other approaches include series expansion with term-by-term integration and integral transforms. As will be seen subsequently, the method of contour integration is perhaps the most versatile of these methods, since it is applicable to a wide varietyof integrals. 464 Chapter 7 Functions of a Complex Variable II Definite Integrals:/integraltext2π 0f(sinθ,cosθ)dθ The calculus of residues is useful in evaluating a wide variety of definite integrals in both physicalandpurelymathematicalproblems.We consider,first, integralsoftheform I=integraldisplay2π 0f(sinθ,cosθ)dθ, (7.30) wherefisfiniteforallvaluesof θ.Wealsorequire ftobearationalfunctionofsin θand cosθso thatitwillbesingle-valued.Let z=eiθ,dz=ieiθdθ. Fromthis, dθ=−idz z,sinθ=z−z−1 2i,cosθ=z+z−1 2. (7.31) Ourintegralbecomes I=−icontintegraldisplay fparenleftbiggz−z−1 2i,z+z−1 2parenrightbiggdz z, (7.32) withthepathof integrationtheunitcircle.By theresiduetheorem,Eq.(7.16), I=(−i)2πisummationdisplay residueswithintheunitcircle. (7.33) Notethatweareaftertheresiduesof f/z.Illustrationsofintegralsofthistypeareprovided byExercises7.1.7–7.1.10. Example 7.1.2 INTEGRAL OF COS IN DENOMINATOR Ourproblemistoevaluatethedefiniteintegral I=integraldisplay2π 0dθ 1+εcosθ,|ε|<1. ByEq. (7.32)thisbecomes I=−icontintegraldisplay unit circledz z[1+(ε/2)(z+z−1)] =−i2 εcontintegraldisplaydz z2+(2/ε)z+1. Thedenominatorhasroots z−=−1 ε−1 εradicalbig 1−ε2andz+=−1 ε+1 εradicalbig 1−ε2, wherez+iswithintheunitcircleand z−isoutside.ThenbyEq.(7.33)andExercise6.6.1, I=−i2 ε·2πi1 2z+2/εvextendsinglevextendsinglevextendsinglevextendsingle z=−1/ε+(1/ε)√ 1−ε2. 7.1 Calculus of Residues 465 Weobtain integraldisplay2π 0dθ 1+εcosθ=2π√ 1−ε2,|ε|<1./squaresolid Evaluation of Definite Integrals:/integraltext∞ −∞f( x)dx Supposethatour definiteintegralhastheform I=integraldisplay∞ −∞f(x)dx (7.34) andsatisfiesthetwoconditions: •f(z)is analytic in the upper half-plane except for a finite number of poles. (It will be assumed that there are no poles on the real axis. If poles are present on the real axis, theymaybeincludedorexcludedas discussedearlierinthissection.) •f(z)vanishesasstrongly2as 1/z2for|z|→∞,0≤argz≤π. With these conditions, we may take as a contour of integration the real axis and a semi- circle in the upper half-plane, as shown in Fig. 7.5. We let the radius Rof the semicircle becomeinfinitelylarge.Then contintegraldisplay f(z)dz=lim R→∞integraldisplayR −Rf(x)dx+lim R→∞integraldisplayπ 0fparenleftbig Reiθparenrightbig iReiθdθ =2πisummationdisplay residues(upperhalf-plane) . (7.35) Fromthesecondconditionthesecondintegral(overthesemicircle)vanishesand integraldisplay∞ −∞f(x)dx=2πisummationdisplay residues(upperhalf-plane) . (7.36) FIGURE 7.5Half-circle contour. 2Wecould use f(z)vanishes fasterthan 1 /z, and wewish to have f(z)single-valued. 466 Chapter 7 Functions of a Complex Variable II Example 7.1.3 INTEGRAL OF MEROMORPHIC FUNCTION Evaluate I=integraldisplay∞ −∞dx 1+x2. (7.37) FromEq. (7.36), integraldisplay∞ −∞dx 1+x2=2πisummationdisplay residues(upperhalf-plane) . Hereandineveryothersimilarproblemwehavethequestion:Wherearethepoles?Rewrit- ingtheintegrandas 1 z2+1=1 z+i·1 z−i, (7.38) wesee thattherearesimplepoles(order 1)at z=iandz=−i. Asimplepoleat z=z0indicates(andisindicatedby)aLaurentexpansionof theform f(z)=a−1 z−z0+a0+∞summationdisplay n=1an(z−z0)n. (7.39) Theresidue a−1iseasilyisolatedas (Exercise6.6.1) a−1=(z−z0)f(z)|z=z0. (7.40) UsingEq.(7.40),wefindthattheresidueat z=iis1/2i,whereasthatat z=−iis−1/2i. Then integraldisplay∞ −∞dx 1+x2=2πi·1 2i=π. (7.41) Here we have used a−1=1/2ifor the residue of the one included pole at z=i. Note that it is possible to use the lower semicircle and that this choice will lead to the same result, I=π. Asomewhatmoredelicateproblemis providedbythenextexample. /squaresolid Evaluation of Definite Integrals:/integraltext∞ −∞f( x) eiaxdx Considerthedefiniteintegral I=integraldisplay∞ −∞f(x)eiaxdx, (7.42) withareal and positive. (This is a Fourier transform, Chapter 15.) We assume the two conditions: •f(z)isanalyticintheupperhalf-planeexceptfor afinitenumberofpoles. 7.1 Calculus of Residues 467 •lim |z|→∞f(z)=0,0≤argz≤π. (7.43) Notethat this is a less restrictive conditionthan the second conditionimposedon f(z)for integratingintegraltext∞ −∞f(x)dxpreviously. We employ the contour shown in Fig. 7.5. The application of the calculus of residues is the same as the one just considered, but here we have to work harder to show that the integraloverthe(infinite)semicirclegoestozero.This integralbecomes IR=integraldisplayπ 0fparenleftbig Reiθparenrightbig eiaRcosθ−aRsinθiReiθdθ. (7.44) LetRbeso largethat |f(z)|=|f(Reiθ)|<ε. Then |IR|≤εRintegraldisplayπ 0e−aRsinθdθ=2εRintegraldisplayπ/2 0e−aRsinθdθ. (7.45) Intherange[0,π/2], 2 πθ≤sinθ. Therefore(Fig.7.6) |IR|≤2εRintegraldisplayπ/2 0e−2aRθ/πdθ. (7.46) Now,integratingbyinspection,weobtain |IR|≤2εR1−e−aR 2aR/π. Finally, lim R→∞|IR|≤π aε. (7.47) FromEq. (7.43), ε→0a sR→∞and lim R→∞|IR|=0. (7.48) FIGURE 7.6(a)y=(2/π)θ,( b )y=sinθ. 468 Chapter 7 Functions of a Complex Variable II This useful result is sometimescalled Jordan’slemma . With it, we are prepared to tackle Fourierintegralsof theformshowninEq.(7.42). UsingthecontourshowninFig.7.5, wehave integraldisplay∞ −∞f(x)eiaxdx+lim R→∞IR=2πisummationdisplay residues(upperhalf-plane) . Sincetheintegralovertheuppersemicircle IRvanishesas R→∞(Jordan’s lemma), integraldisplay∞ −∞f(x)eiaxdx=2πisummationdisplay residues(upperhalf-plane )(a>0).(7.49) Example 7.1.4 SIMPLE POLE ON CONTOUR OF INTEGRATION Theproblemistoevaluate I=integraldisplay∞ 0sinx xdx. (7.50) Thismaybetakenas theimaginarypart3of I2=Pintegraldisplay∞ −∞eizdz z. (7.51) Nowtheonlypoleisasimplepoleat z=0andtheresiduetherebyEq.(7.40)is a−1=1. We choose the contour shown in Fig. 7.7 (1) to avoid the pole, (2) to include the real axis, and (3) to yield a vanishingly small integrand for z=iy,y→∞. Note that in this case a large(infinite)semicircleinthelowerhalf-planewouldbedisastrous. Wehave contintegraldisplayeizdz z=integraldisplay−r −Reixdx x+integraldisplay C1eizdz z+integraldisplayR reixdx x+integraldisplay C2eizdz z=0,(7.52) FIGURE 7.7Singularityoncontour. 3One can useintegraltext [(eiz−e−iz)/2iz]dz, but then two different contours will be needed for the two exponentials (compare Exam- ple 7.1.5). 7.1 Calculus of Residues 469 thefinalzerocomingfromtheresiduetheorem(Eq. (7.6)). ByJordan’slemma integraldisplay C2eizdz z=0, (7.53) and contintegraldisplayeizdz z=integraldisplay C1eizdz z+Pintegraldisplay∞ −∞eixdx x=0. (7.54) Theintegraloverthesmallsemicircleyields (−)πitimestheresidueof1,andminus,asa resultofgoingclockwise.Takingtheimaginarypart,4wehave integraldisplay∞ −∞sinx xdx=π (7.55) orintegraldisplay∞ 0sinx xdx=π 2. (7.56) The contour of Fig. 7.7, although convenient, is not at all unique. Another choice of contourfor evaluatingEq. (7.50)is presentedas Exercise7.1.15. /squaresolid Example 7.1.5 QUANTUM MECHANICAL SCATTERING Thequantummechanicalanalysisofscatteringleadstothefunction I(σ)=integraldisplay∞ −∞xsinxdx x2−σ2, (7.57) whereσis real and positive. This integral is divergent and therefore ambiguous. From the physicalconditionsoftheproblemthereisafurtherrequirement: I(σ)istohavetheform eiσso thatitwillrepresentanoutgoingscatteredwave. Using sinz=1 isinhiz=1 2ieiz−1 2ie−iz, (7.58) wewriteEq. (7.57)inthecomplexplaneas I(σ)=I1+I2, (7.59) with I1=1 2iintegraldisplay∞ −∞zeiz z2−σ2dz, I2=−1 2iintegraldisplay∞ −∞ze−iz z2−σ2dz. (7.60) 4Alternatively, wemaycombine the integrals ofEq. (7.52) as integraldisplay−r −Reixdx x+integraldisplayR reixdx x=integraldisplayR rparenleftbig eix−e−ixparenrightbigdx x=2iintegraldisplayR rsinx xdx. 470 Chapter 7 Functions of a Complex Variable II FIGURE 7.8Contours. IntegralI1issimilartoExample7.1.4and,asinthatcase,wemaycompletethecontourby aninfinitesemicircleintheupperhalf-plane,asshowninFig.7.8a.For I2theexponential is negative and we complete the contour by an infinite semicircle in the lower half-plane, as shown in Fig. 7.8b. As in Example 7.1.4, neither semicircle contributes anything to the integral—Jordan’slemma. Thereisstilltheproblemoflocatingthepolesandevaluatingtheresidues.Wefindpoles atz=+σandz=−σon the contour of integration . The residues are (Exercises 6.6.1 and7.1.1) z=σz=−σ I1eiσ 2e−iσ 2 I2e−iσ 2eiσ 2 Detouring around the poles, as shown in Fig. 7.8 (it matters little whether we go above or below),wefindthattheresiduetheoremleadsto PI1−πiparenleftbigg1 2iparenrightbigge−iσ 2+πiparenleftbigg1 2iparenrightbiggeiσ 2=2πiparenleftbigg1 2iparenrightbiggeiσ 2, (7.61) for we have enclosed the singularity at z=σbut excluded the one at z=−σ. In similar fashion,butnotingthatthecontourfor I2is clockwise, PI2−πiparenleftbigg−1 2iparenrightbiggeiσ 2+πiparenleftbigg−1 2iparenrightbigge−iσ 2=−2πiparenleftbigg−1 2iparenrightbiggeiσ 2. (7.62) AddingEqs.(7.61) and(7.62), wehave PI(σ)=PI1+PI2=π 2parenleftbig eiσ+e−iσparenrightbig =πcoshiσ=πcosσ. (7.63) This is a perfectly good evaluation of Eq. (7.57), but unfortunately the cosine dependence isappropriatefor astandingwaveandnotfortheoutgoingscatteredwaveas specified. 7.1 Calculus of Residues 471 To obtain the desired form, we try a different technique (compare Example 7.1.1). In- steadofdodgingaroundthesingularpoints,letusmovethemofftherealaxis.Specifically, letσ→σ+iγ,−σ→−σ−iγ,whereγispositivebutsmallandwilleventuallybemade toapproachzero;thatis, for I1weincludeonepoleandfor I2theotherone, I+(σ)=lim γ→0I(σ+iγ). (7.64) Withthissimplesubstitution,thefirstintegral I1becomes I1(σ+iγ)=2πiparenleftbigg1 2iparenrightbiggei(σ+iγ) 2(7.65) bydirectapplicationof theresiduetheorem.Also, I2(σ+iγ)=−2πiparenleftbigg−1 2iparenrightbiggei(σ+iγ) 2. (7.66) AddingEqs.(7.65) and(7.66)andthenletting γ→0,weobtain I+(σ)=lim γ→0bracketleftbig I1(σ+iγ)+I2(σ+iγ)bracketrightbig =lim γ→0πei(σ+iγ)=πeiσ, (7.67) aresultthatdoesfittheboundaryconditionsofourscatteringproblem. It isinterestingtonotethatthesubstitution σ→σ−iγwouldhaveledto I−(σ)=πe−iσ, (7.68) which could represent an incoming wave. Our earlier result (Eq. (7.63)) is seen to be the arithmeticaverageofEqs.(7.67)and(7.68).ThisaverageistheCauchyprincipalvalueof the integral. Note that we have these possibilities (Eqs. (7.63), (7.67), and (7.68)) because our integral is not uniquely defined until we specify the particular limiting process (or average)tobeused. /squaresolid Evaluation of Definite Integrals: Exponential Forms Withexponentialorhyperbolicfunctionspresentintheintegrand,lifegetssomewhatmore complicatedthanbefore.Insteadofageneraloverallprescription,thecontourmustbecho- sentofitthespecificintegral.Thesecasesarealsoopportunitiestoillustratetheversatility andpowerofcontourintegration. Asanexample,weconsideranintegralthatwillbequiteusefulindevelopingarelation betweenŴ(1+z)andŴ(1−z). Notice how the periodicity along the imaginary axis is exploited. 472 Chapter 7 Functions of a Complex Variable II FIGURE 7.9Rectangularcontour. Example 7.1.6 FACTORIAL FUNCTION We wish to evaluate I=integraldisplay∞ −∞eax 1+exdx,0<a<1. (7.69) The limits on aare sufficient (but not necessary) to prevent the integral from diverging as x→±∞. This integral (Eq. (7.69)) may be handled by replacing the real variable xby the complex variable zand integrating around the contour shown in Fig. 7.9. If we take thelimitas R→∞,therealaxis,ofcourse,leadstotheintegralwewant.Thereturnpath alongy=2πischosentoleavethedenominatoroftheintegralinvariant,atthesametime introducingaconstantfactor ei2πainthenumerator.We have,inthecomplexplane, contintegraldisplayeaz 1+ezdz=lim R→∞parenleftbiggintegraldisplayR −Reax 1+exdx−ei2πaintegraldisplayR −Reax 1+exdxparenrightbigg =parenleftbig 1−ei2πaparenrightbigintegraldisplay∞ −∞eax 1+exdx. (7.70) In addition there are two vertical sections (0≤y≤2π), which vanish (exponentially) as R→∞. Nowwherearethepolesandwhataretheresidues?We haveapolewhen ez=exeiy=−1. (7.71) Equation (7.71) is satisfied at z=0+iπ. By a Laurent expansion5in powers of (z−iπ) the pole is seen to be a simple pole with a residue of −eiπa. Then, applying the residue theorem, parenleftbig 1−ei2πaparenrightbigintegraldisplay∞ −∞eax 1+exdx=2πiparenleftbig −eiπaparenrightbig . (7.72) Thisquicklyreducesto integraldisplay∞ −∞eax 1+exdx=π sinaπ,0<a<1. (7.73) 51+ez=1+ez−iπeiπ=1−ez−iπ=−(z−iπ)(1+z−iπ 2!+(z−iπ)2 3!+···). 7.1 Calculus of Residues 473 Usingthebetafunction(Section8.4),wecanshowtheintegraltobeequaltotheproduct Ŵ(a)Ŵ(1−a). Thisresults intheinterestingandusefulfactorialfunctionrelation Ŵ(a+1)Ŵ(1−a)=πa sinπa. (7.74) Although Eq. (7.73) holds for real a,0<a<1, Eq. (7.74) may be extended by analytic continuationtoallvaluesof a, realandcomplex,excludingonlyrealintegralvalues. /squaresolid As a final example of contour integrals of exponential functions, we consider Bernoulli numbersagain. Example 7.1.7 BERNOULLI NUMBERS InSection5.9theBernoullinumbersweredefinedbytheexpansion x ex−1=∞summationdisplay n=0Bn n!xn. (7.75) Replacing xwithz(analytic continuation), we have a Taylor series (compare Eq. (6.47)) with Bn=n! 2πicontintegraldisplay C0z ez−1dz zn+1, (7.76) where the contour C0is around the origin counterclockwise with |z|<2πto avoid the polesat 2 πin. Forn=0weha v easimplepoleat z=0 witharesidueof +1.HencebyEq. (7.25), B0=0! 2πi·2πi(1)=1. (7.77) Forn=1thesingularityat z=0becomesasecond-orderpole.Theresiduemaybeshown to be−1 2by series expansion of the exponential, followed by a binomial expansion. This resultsin B1=1! 2πi·2πiparenleftbigg −1 2parenrightbigg =−1 2. (7.78) Forn≥2 this procedure becomes rather tedious, and we resort to a different means of evaluatingEq. (7.76). Thecontouris deformed,as showninFig.7.10. The new contour Cstill encircles the origin, as required, but now it also encircles (in a negative direction) an infinite series of singular points along the imaginary axis at z=±p2πi,p=1,2,3,.... The integration back and forth along the x-axis cancels out, and forR→∞the integration over the infinite circle yields zero. Remember that n≥2. Therefore contintegraldisplay C0z ez−1dz zn+1=−2πi∞summationdisplay p=1residues (z=±p2πi). (7.79) 474 Chapter 7 Functions of a Complex Variable II FIGURE 7.10Contourof integrationforBernoullinumbers. Atz=p2πiwe have a simple pole with a residue (p2πi)−n. Whennis odd, the residue fromz=p2πiexactly cancels that from z=−p2πiandBn=0,n=3,5,7, and so on. Forneventheresiduesadd,giving Bn=n! 2πi(−2πi)2∞summationdisplay p=11 pn(2πi)n =−(−1)n/22n! (2π)n∞summationdisplay p=1p−n=−(−1)n/22n! (2π)nζ(n) (n even),(7.80) whereζ(n)is the Riemann zeta function introduced in Section 5.9. Equation (7.80) corre- spondstoEq.(5.152) ofSection5.9. /squaresolid Exercises 7.1.1 Determinethenatureofthesingularitiesofeachofthefollowingfunctionsandevaluate theresidues (a >0). (a)1 z2+a2.( b)1 (z2+a2)2. (c)z2 (z2+a2)2.( d)sin1/z z2+a2. (e)ze+iz z2+a2. (f)ze+iz z2−a2. (g)e+iz z2−a2.( h)z−k z+1,0<k<1. Hint. For the point at infinity, use the transformation w=1/zfor|z|→0. For the residue,transform f(z)dzintog(w)dw andlookatthebehaviorof g(w). 7.1 Calculus of Residues 475 7.1.2 Locatethesingularitiesandevaluatetheresidues ofeachof thefollowingfunctions. (a)z−n(ez−1)−1,z/negationslash=0, (b)z2ez 1+e2z. (c) Find a closed-form expression (that is, not a sum) for the sum of the finite-plane singularities. (d) Usingtheresultinpart(c), whatistheresidueat |z|→∞? Hint.SeeSection5.9forexpressionsinvolvingBernoullinumbers.NotethatEq.(5.144) cannotbeusedtoinvestigatethesingularityat z→∞,sincethisseriesisonlyvalidfor |z|<2π. 7.1.3 The statement that the integral halfway around a singular point is equal to one-half the integral all the way around was limited to simple poles. Show, by a specific example, that integraldisplay Semicirclef(z)dz=1 2contintegraldisplay Circlef(z)dz doesnotnecessarilyholdif theintegralencirclesapoleofhigherorder. Hint.T ryf(z)=z−2. 7.1.4 A function f(z)is analytic along the real axis except for a third-order pole at z=x0. TheLaurentexpansionabout z=x0hastheform f(z)=a−3 (z−x0)3+a−1 z−x0+g(z), withg(z)analyticat z=x0.ShowthattheCauchyprincipalvaluetechniqueisapplica- ble,inthesensethat (a) lim δ→0braceleftbiggintegraldisplayx0−δ −∞f(x)dx+integraldisplay∞ x0+δf(x)dxbracerightbigg isfinite. (b)integraldisplay Cx0f(z)dz=±iπa−1, whereCx0denotesa smallsemicircle aboutz=x0. 7.1.5 Theunitstepfunctionisdefinedas(compareExercise1.15.13) u(s−a)=braceleftBigg0,s<a 1,s>a. Showthat u(s)hastheintegralrepresentations (a)u(s)=lim ε→0+1 2πiintegraldisplay∞ −∞eixs x−iεdx, 476 Chapter 7 Functions of a Complex Variable II (b)u(s)=1 2+1 2πiPintegraldisplay∞ −∞eixs xdx. Note.Theparameter sis real. 7.1.6 Most of the special functions of mathematical physics may be generated (defined) by ageneratingfunctionoftheform g(t,x)=summationdisplay nfn(x)tn. Giventhefollowingintegralrepresentations,derivethecorrespondinggeneratingfunc- tion: (a) Bessel: Jn(x)=12πicontintegraldisplay e(x/2)(t−1/t)t−n−1dt. (b) ModifiedBessel: In(x)=1 2πicontintegraldisplay e(x/2)(t+1/t)t−n−1dt. (c) Legendre: Pn(x)=1 2πicontintegraldisplayparenleftbig 1−2tx+t2parenrightbig−1/2t−n−1dt. (d) Hermite: Hn(x)=n! 2πicontintegraldisplay e−t2+2txt−n−1dt. (e) Laguerre: Ln(x)=1 2πicontintegraldisplaye−xt/(1−t) (1−t)tn+1dt. (f) Chebyshev: Tn(x)=1 4πicontintegraldisplay(1−t2)t−n−1 (1−2tx+t2)dt. Eachofthecontoursencirclestheoriginandnoothersingularpoints. 7.1.7 GeneralizingExample7.1.2, showthat integraldisplay2π 0dθ a±bcosθ=integraldisplay2π 0dθ a±bsinθ=2π (a2−b2)1/2,fora>|b|. Whathappensif |b|>|a|? 7.1.8 Showthatintegraldisplayπ 0dθ (a+cosθ)2=πa (a2−1)3/2,a>1. 7.1 Calculus of Residues 477 7.1.9 Showthat integraldisplay2π 0dθ 1−2tcosθ+t2=2π 1−t2,for|t|<1. Whathappensif |t|>1?Whathappensif |t|=1? 7.1.10 Withthecalculusofresiduesshowthat integraldisplayπ 0cos2nθdθ=π(2n)! 22n(n!)2=π(2n−1)!! (2n)!!,n=0,1,2,.... (ThedoublefactorialnotationisdefinedinSection8.1.) Hint. cosθ=1 2(eiθ+e−iθ)=1 2(z+z−1),|z|=1. 7.1.11 Evaluate integraldisplay∞ −∞cosbx−cosax x2dx, a>b> 0. ANS.π(a−b). 7.1.12 Provethat integraldisplay∞ −∞sin2x x2dx=π 2. Hint.s i n2x=1 2(1−cos2x). 7.1.13 A quantum mechanical calculation of a transition probability leads to the function f(t,ω)=2(1−cosωt)/ω2. Showthat integraldisplay∞ −∞f(t,ω)dω=2πt. 7.1.14 Showthat (a >0) (a)integraldisplay∞ −∞cosx x2+a2dx=π ae−a. Howis therightsidemodifiedif cos xis replacedby cos kx? (b)integraldisplay∞ −∞xsinx x2+a2dx=πe−a. Howis therightsidemodifiedif sin xisreplacedby sin kx? These integrals may also be interpreted as Fourier cosine and sine transforms— Chapter15. 7.1.15 Usethecontourshown(Fig.7.11)with R→∞toprovethat integraldisplay∞ −∞sinx xdx=π. 478 Chapter 7 Functions of a Complex Variable II FIGURE 7.11Largesquare contour. 7.1.16 Inthequantumtheoryofatomiccollisionsweencountertheintegral I=integraldisplay∞ −∞sint teiptdt, inwhich pisreal. Showthat I=0,|p|>1 I=π,|p|<1. Whathappensif p=±1? 7.1.17 Evaluate integraldisplay∞ 0(lnx)2 1+x2dx (a) byappropriateseries expansionoftheintegrandtoobtain 4∞summationdisplay n=0(−1)n(2n+1)−3, (b) andbycontourintegrationtoobtain π3 8. Hint.x→z=et. TrythecontourshowninFig.7.12,letting R→∞. 7.1.18 Showthat integraldisplay∞ 0xa (x+1)2dx=πa sinπa, FIGURE 7.12Smallsquare contour. 7.1 Calculus of Residues 479 FIGURE 7.13Contouravoiding branchpointandpole. where−1<a<1.Hereis stillanotherwayof derivingEq. (7.74). Hint. Use the contour shown in Fig. 7.13, noting that z=0 is a branch point and the positivex-axisisacutline.NotealsothecommentsonphasesfollowingExample6.6.1. 7.1.19 Showthat integraldisplay∞ 0x−a x+1dx=π sinaπ, where 0<a<1. This opens up another way of deriving the factorial function relation givenbyEq.(7.74). Hint.Youhaveabranchpointandyouwillneedacutline.Recallthat z−a=winpolar form is bracketleftbig rei(θ+2πn)bracketrightbig−a=ρeiϕ, whichleadsto −aθ−2anπ=ϕ.Youmustrestrict ntozero(oranyothersingleinteger) inorderthat ϕmaybeuniquelyspecified.TrythecontourshowninFig.7.14. FIGURE 7.14Alternativecontour avoidingbranchpoint. 480 Chapter 7 Functions of a Complex Variable II FIGURE 7.15Anglecontour. 7.1.20 Showthat integraldisplay∞ 0dx (x2+a2)2=π 4a3,a>0. 7.1.21 Evaluate integraldisplay∞ −∞x2 1+x4dx. ANS.π/√ 2. 7.1.22 Showthat integraldisplay∞ 0cosparenleftbig t2parenrightbig dt=integraldisplay∞ 0sinparenleftbig t2parenrightbig dt=√π 2√ 2. Hint.TrythecontourshowninFig.7.15. Note. These are the Fresnel integrals for the special case of infinity as the upper limit. For the general case of a varying upper limit, asymptotic expansions of the Fresnel integralsarethetopicofExercise5.10.2.SphericalBesselexpansionsarethesubjectof Exercise11.7.13. 7.1.23 SeveraloftheBromwichintegrals,Section15.12,involveaportionthatmaybeapprox- imatedby I(y)=integraldisplaya+iy a−iyezt z1/2dz. Hereaandtare positiveandfinite.Showthat limy→∞I(y)=0. 7.1 Calculus of Residues 481 FIGURE 7.16Sectorcontour. 7.1.24 Showthat integraldisplay∞ 01 1+xndx=π/n sin(π/n). Hint.TrythecontourshowninFig.7.16. 7.1.25 (a) Showthat f(z)=z4−2cos2θz2+1 haszerosat eiθ,e−iθ,−eiθ, and−e−iθ. (b) Showthat integraldisplay∞ −∞dx x4−2cos2θx2+1=π 2sinθ=π 21/2(1−cos2θ)1/2. Exercise7.1.24 (n=4)isaspecialcaseofthisresult. 7.1.26 Showthat integraldisplay∞ −∞x2dx x4−2cos2θx2+1=π 2sinθ=π 21/2(1−cos2θ)1/2. Exercise7.1.21isaspecialcaseofthisresult. 7.1.27 Applythetechniquesof Example7.1.5totheevaluationof theimproperintegral I=integraldisplay∞ −∞dx x2−σ2. (a) Let σ→σ+iγ. (b) Let σ→σ−iγ. (c) TaketheCauchyprincipalvalue. 482 Chapter 7 Functions of a Complex Variable II 7.1.28 TheintegralinExercise7.1.17maybetransformedinto integraldisplay∞ 0e−yy2 1+e−2ydy=π3 16. Evaluate this integral by the Gauss–Laguerre quadrature and compare your result with π3/16. ANS.Integral=1.93775(10points). 7.2 D ISPERSION RELATIONS The concept of dispersion relations entered physics with the work of Kronig and Kramers in optics. The name dispersion comes from optical dispersion, a result of the dependence of the index of refraction on wavelength, or angular frequency. The index of refraction nmay have a real part determined by the phase velocity and a (negative) imaginary part determined by the absorption—see Eq. (7.94). Kronig and Kramers showed in 1926– 1927 that the real part of (n2−1)could be expressed as an integral of the imaginary part. Generalizing this, we shall apply the label dispersion relations to any pair of equations givingtherealpartofafunctionasanintegralofitsimaginarypartandtheimaginarypart as an integral of its real part—Eqs. (7.86a) and (7.86b), which follow. The existence of such integral relations might be suspected as an integral analog of the Cauchy–Riemann differentialrelations,Section6.2. The applications in modern physics are widespread. For instance, the real part of the function might describe the forward scattering of a gamma ray in a nuclear Coulomb field (a dispersive process). Then the imaginary part would describe the electron–positron pair production in that same Coulomb field (the absorptive process). As will be seen later, the dispersionrelationsmaybetakenasaconsequenceofcausalityandthereforeareindepen- dentof thedetailsoftheparticularinteraction. We consider a complex function f(z)that is analytic in the upper half-plane and on the realaxis.Wealsorequirethat lim |z|→∞vextendsinglevextendsinglef(z)vextendsinglevextendsingle=0,0≤argz≤π, (7.81) in order that the integral over an infinite semicircle will vanish. The point of these condi- tionsisthatwemayexpress f(z)bytheCauchyintegralformula,Eq.(6.43), f(z0)=1 2πicontintegraldisplayf(z) z−z0dz. (7.82) Theintegralovertheuppersemicircle6vanishesandwehave f(z0)=1 2πiintegraldisplay∞ −∞f(x) x−z0dx. (7.83) TheintegraloverthecontourshowninFig.7.17hasbecomeanintegralalongthe x-axis. Equation (7.83) assumes that z0is in the upper half-plane—interior to the closed con- tour.Ifz0wereinthelowerhalf-plane,theintegralwouldyieldzerobytheCauchyintegral 6Theuse of asemicircleto closethe path of integration is convenient, not mandatory. Other paths are possible. 7.2 Dispersion Relations 483 FIGURE 7.17Semicirclecontour. theorem, Section 6.3. Now, either letting z0approach the real axis from above (z0−x0) or placing it on the real axis and taking an average of Eq. (7.83) and zero, we find that Eq.(7.83) becomes f(x0)=1 πiPintegraldisplay∞ −∞f(x) x−x0dx, (7.84) wherePindicatestheCauchyprincipalvalue.SplittingEq.(7.84)intorealandimaginary parts7yields f(x0)=u(x0)+iv(x0) =1 πPintegraldisplay∞ −∞v(x) x−x0dx−i πPintegraldisplay∞ −∞u(x) x−x0dx. (7.85) Finally,equatingrealparttorealpartandimaginaryparttoimaginarypart,weobtain u(x0)=1 πPintegraldisplay∞ −∞v(x) x−x0dx(7.86a) v(x0)=−1 πPintegraldisplay∞ −∞u(x) x−x0dx.(7.86b) These are the dispersion relations. The real part of our complex function is expressed as an integral over the imaginary part. The imaginary part is expressed as an integral over therealpart.Therealandimaginarypartsare Hilberttransforms ofeachother.Notethat theserelationsaremeaningfulonlywhen f(x)isacomplexfunctionoftherealvariable x. CompareExercise7.2.1. Fromaphysicalpointofview u(x)and/orv(x)representsomephysicalmeasurements. Thenf(z)=u(z)+iv(z)is an analytic continuation over the upper half-plane, with the valueontherealaxisservingas aboundarycondition. 7Thesecondargument, y=0,is dropped: u(x0,0)→u(x0). 484 Chapter 7 Functions of a Complex Variable II Symmetry Relations On occasion f(x)will satisfy a symmetry relation and the integral from −∞to+∞ may be replaced by an integral over positive values only. This is of considerable physical importance because the variable xmight represent a frequency and only zero and positive frequenciesareavailableforphysicalmeasurements.Suppose8 f(−x)=f∗(x). (7.87) Then u(−x)+iv(−x)=u(x)−iv(x). (7.88) The real part of f(x)is even and the imaginary part is odd.9In quantum mechanical scattering problems these relations (Eq. (7.88)) are called crossing conditions. To exploit thesecrossingconditions ,werewriteEq. (7.86a)as u(x0)=1 πPintegraldisplay0 −∞v(x) x−x0dx+1 πPintegraldisplay∞ 0v(x) x−x0dx. (7.89) Lettingx→−xin the first integral on the right-hand side of Eq. (7.89) and substituting v(−x)=−v(x)from Eq.(7.88), weobtain u(x0)=1 πPintegraldisplay∞ 0v(x)braceleftbigg1 x+x0+1 x−x0bracerightbigg dx =2 πPintegraldisplay∞ 0xv(x) x2−x2 0dx. (7.90) Similarly, v(x0)=−2 πPintegraldisplay∞ 0x0u(x) x2−x2 0dx. (7.91) The original Kronig–Kramers optical dispersion relations were in this form. The asymp- toticbehavior (x0→∞)ofEqs.(7.90)and(7.91)leadtoquantummechanical sumrules , Exercise7.2.4. Optical Dispersion Thefunction exp [i(kx−ωt)]describesanelectromagneticwavemovingalongthe x-axis in the positive direction with velocity v=ω/k;ωis the angular frequency, kthe wave number or propagation vector, and n=ck/ωthe index of refraction. From Maxwell’s 8This is not just a happy coincidence. It ensures that the Fourier transform of f(x)will be real. In turn, Eq. (7.87) is a conse- quence of obtaining f(x)as the Fourier transform of areal function. 9u(x,0)=u(−x,0),v(x,0)=−v(−x,0).ComparethesesymmetryconditionswiththosethatfollowfromtheSchwarzreflec- tion principle, Section6.5. 7.2 Dispersion Relations 485 equations, electric permittivity ε, and Ohm’s law with conductivity σ, the propagation vectorkfor adielectricbecomes10 k2=εω2 c2parenleftbigg 1+i4πσ ωεparenrightbigg (7.92) (withµ, the magnetic permeability, taken to be unity). The presence of the conductivity (which means absorption) gives rise to an imaginary part. The propagation vector k(and thereforetheindexof refraction n)havebecomecomplex. Conversely, the (positive) imaginary part implies absorption. For poor conductivity (4πσ/ωε≪1)abinomialexpansionyields k=√εω c+i2πσ c√ε and ei(kx−ωt)=eiω(x√ε/c−t)e−2πσx/c√ε, anattenuatedwave. Returningtothegeneralexpressionfor k2,Eq.(7.92),wefindthattheindexofrefraction becomes n2=c2k2 ω2=ε+i4πσ ω. (7.93) We taken2to be a function of the complex variableω(withεandσdepending on ω). However, n2does not vanish as ω→∞but instead approaches unity. So to satisfy the condition, Eq. (7.81), one works with f(ω)=n2(ω)−1. The original Kronig–Kramers opticaldispersionrelationswereintheform of ℜbracketleftbig n2(ω0)−1bracketrightbig =2 πPintegraldisplay∞ 0ωℑ[n2(ω)−1] ω2−ω2 0dω, (7.94) ℑbracketleftbig n2(ω0)−1bracketrightbig =−2 πPintegraldisplay∞ 0ω0ℜ[n2(ω)−1] ω2−ω2 0dω. Knowledge of the absorption coefficient at all frequencies specifies the real part of the indexofrefraction,andviceversa. The Parseval Relation When the functions u(x)andv(x)are Hilbert transforms of each other (given by Eqs. (7.86)) andeachis squareintegrable,11thetwofunctionsarerelatedby integraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx. (7.95) 10S e eJ .D .J a c k s o n , Classical Electrodynamics , 3rd ed. New York: Wiley (1999), Sections 7.7 and 7.10. Equation (7.92) is in Gaussianunits. 11This means thatintegraltext∞ −∞|u(x)|2dxandintegraltext∞ −∞|v(x)|2dxare finite. 486 Chapter 7 Functions of a Complex Variable II Thisis theParsevalrelation. ToderiveEq.(7.95), westartwith integraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞1 πintegraldisplay∞ −∞v(s)ds s−x1 πintegraldisplay∞ −∞v(t)dt t−xdx, usingEq. (7.86a)twice.Integratingfirstwithrespectto x,weha v e integraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞v(s)dsintegraldisplay∞ −∞v(t)dt π2integraldisplay∞ −∞dx (s−x)(t−x).(7.96) FromExercise7.2.8, the xintegrationyieldsadeltafunction: 1 π2integraldisplay∞ −∞dx (s−x)(t−x)=δ(s−t). We haveintegraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞v(t)dtintegraldisplay∞ −∞v(s)δ(s−t)ds. (7.97) Then the sintegrationis carried out by inspection, using the defining property of the delta function:integraldisplay∞ −∞v(s)δ(s−t)ds=v(t). (7.98) SubstitutingEq.(7.98)intoEq.(7.97),wehaveEq.(7.95),theParsevalrelation.Again,in terms of optics, the presence of refraction over some frequency range (n/negationslash=1)implies the existenceof absorption,andviceversa. Causality Therealsignificanceofdispersionrelationsinphysicsisthattheyareadirectconsequence of assuming that the particular physical system obeys causality. Causality is awkward to defineprecisely,butthegeneralmeaningisthattheeffectcannotprecedethecause.Ascat- teredwavecannotbeemittedbythescatteringcenterbeforetheincidentwavehasarrived. For linear systems the most general relation between an input function G(the cause) and anoutputfunction H(theeffect) maybewrittenas H(t)=integraldisplay∞ −∞F(t−t′)G(t′)dt′. (7.99) Causalityis imposedbyrequiringthat F(t−t′)=0fort−t′<0. Equation(7.99) givesthetimedependence.The frequencydependenceis obtainedbytak- ingFouriertransforms. BytheFourierconvolutiontheorem,Section15.5, h(ω)=f(ω)g(ω), wheref(ω)is the Fourier transform of F(t), and so on. Conversely, F(t)is the Fourier transform of f(ω). 7.2 Dispersion Relations 487 The connection with the dispersion relations is provided by the Titchmarsh theorem.12 This states that if f(ω)is square integrable over the real ω-axis, then any one of the fol- lowingthreestatementsimpliestheothertwo. 1. TheFouriertransform of f(ω)is zerofor t<0:Eq. (7.99). 2. Replacing ωbyz, the function f(z)is analytic in the complex z-plane for y>0 and approaches f(x)almosteverywhereas y→0.Further, integraldisplay∞ −∞vextendsinglevextendsinglef(x+iy)vextendsinglevextendsingle2dx<K fory>0; thatis, theintegralis bounded. 3. Therealandimaginarypartsof f(z)areHilberttransformsofeachother:Eqs.(7.86a) and(7.86b). The assumption that the relationship between the input and the output of our linear system is causal (Eq. (7.99)) means that the first statement is satisfied. If f(ω)is square integrable, then the Titchmarsh theorem has the third statement as a consequence and we havedispersionrelations. Exercises 7.2.1 The function f(z)satisfies the conditions for the dispersion relations. In addition, f(z)=f∗(z∗), the Schwarz reflection principle, Section 6.5. Show that f(z)is identi- callyzero. 7.2.2 Forf(z)such that we may replace the closed contour of the Cauchy integral formula byanintegralovertherealaxiswehave f(x0)=1 2πibraceleftbiggintegraldisplayx0−δ −∞f(x) x−x0dx+integraldisplay∞ x0+δf(x) x−x0dxbracerightbigg +1 2πiintegraldisplay Cx0f(x) x−x0dx. HereCx0designates a small semicircle about x0in the lower half-plane. Show that this reducesto f(x0)=1 πiPintegraldisplay∞ −∞f(x) x−x0dx, whichisEq. (7.84). 7.2.3 (a) For f(z)=eiz,Eq.(7.81)doesnotholdattheendpoints,arg z=0,π.Show,with thehelpofJordan’slemma,Section7.1, thatEq.(7.82) stillholds. (b) For f(z)=eizverifythedispersionrelations,Eq.(7.89)orEqs.(7.90)and(7.91), bydirectintegration. 7.2.4 Withf(x)=u(x)+iv(x)andf(x)=f∗(−x), showthatas x0→∞, 12RefertoE.C.Titchmarsh, IntroductiontotheTheoryofFourierIntegrals ,2nded.NewYork:OxfordUniversityPress(1937). ForamoreinformaldiscussionoftheTitchmarshtheoremandfurtherdetailsoncausalityseeJ.Hilgevoord, DispersionRelations and Causal Description . Amsterdam: North-Holland (1962). 488 Chapter 7 Functions of a Complex Variable II (a)u(x0)∼−2 πx2 0integraldisplay∞ 0xv(x)dx, (b)v(x0)∼2 πx0integraldisplay∞ 0u(x)dx. Inquantummechanicsrelationsofthisform areoftencalled sumrules . 7.2.5 (a) Giventheintegralequation 1 1+x2 0=1 πPintegraldisplay∞ −∞u(x) x−x0dx, useHilberttransforms todetermine u(x0). (b) Verifythattheintegralequationofpart(a) is satisfied. (c) From f(z)|y=0=u(x)+iv(x),replacexbyzanddetermine f(z).Verifythatthe conditionsfor theHilberttransforms aresatisfied. (d) Arethecrossing conditionssatisfied? ANS.(a) u(x0)=x0 1+x2 0,(c)f(z)=(z+i)−1. 7.2.6 (a) If therealpartof thecomplexindexofrefraction(squared) is constant(nooptical dispersion),showthattheimaginarypartis zero(no absorption). (b) Conversely, if there is absorption, show that there must be dispersion. In other words,iftheimaginarypartof n2−1 isnotzero,showthattherealpartof n2−1 isnotconstant. 7.2.7 Givenu(x)=x/(x2+1)andv(x)=−1/(x2+1), show by direct evaluation of each integralthat integraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx. ANS.integraldisplay∞ −∞vextendsinglevextendsingleu(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞vextendsinglevextendsinglev(x)vextendsinglevextendsingle2dx=π 2. 7.2.8 Takeu(x)=δ(x), a delta function, and assumethat the Hilbert transform equations hold. (a) Showthat δ(w)=1 π2integraldisplay∞ −∞dy y(y−w). (b) Withchangesofvariables w=s−tandx=s−y,transformthe δrepresentation ofpart(a) into δ(s−t)=1 π2integraldisplay∞ −∞dx (x−s)(s−t). Note.TheδfunctionisdiscussedinSection1.15. 7.3 Method of Steepest Descents 489 7.2.9 Showthat δ(x)=1 π2integraldisplay∞ −∞dt t(t−x) isa validrepresentationofthedeltafunctioninthesensethat integraldisplay∞ −∞f(x)δ(x)dx=f(0). Assumethat f(x)satisfiestheconditionfor theexistenceof aHilberttransform. Hint.ApplyEq.(7.84) twice. 7.3 M ETHOD OF STEEPEST DESCENTS Analytic Landscape In analyzing problems in mathematical physics, one often finds it desirable to know the behavior of a function for large values of the variable or some parameter s, that is, the asymptotic behavior of the function. Specific examples are furnished by the gamma func- tion(Chapter8)andvariousBesselfunctions(Chapter11).Alltheseanalyticfunctionsare definedbyintegrals I(s)=integraldisplay CF(z,s)dz, (7.100) whereFis analytic in zand depends on a real parameter s. We write F(z)whenever possible. Sofar wehaveevaluatedsuchdefiniteintegralsof analyticfunctionsalongtherealaxis bydeformingthepath CtoC′inthecomplexplane,so |F|becomessmallforall zonC′. This method succeeds as long as only isolated poles occur in the area between CandC′. The poles are taken into account by applying the residue theorem of Section 7.1. The residuesgiveameasureofthesimplepoles,where |F|→∞,whichusuallydominateand determinethevalueof theintegral. The behavior of the integral in Eq. (7.100) clearly depends on the absolute value |F|of theintegrand.Moreover,thecontoursof |F|oftenbecomemorepronouncedas sbecomes large.Letusfocusonaplotof |F(x+iy)|2=U2(x,y)+V2(x,y),ratherthantherealpart ℜF=Uandtheimaginarypart ℑF=Vseparately.Suchaplotof |F|2overthecomplex plane is called the analytic landscape , after Jensen, who, in 1912, proved that it has only saddle points and troughs but no peaks . Moreover, the troughs reach down all the way to the complex plane. In the absence of (simple) poles, saddle points are next in line to dominate the integral in Eq. (7.100). Hence the name saddle point method . At a saddle pointthereal(orimaginary)part UofFhasalocalmaximum,whichimpliesthat ∂U ∂x=∂U ∂y=0, andthereforebytheuseoftheCauchy–Riemannconditionsof Section6.2, ∂V ∂x=∂V ∂y=0, 490 Chapter 7 Functions of a Complex Variable II soVhas a minimum, or vice versa, and F′(z)=0. Jensen’s theorem prevents Uand Vfrom having either a maximum or a minimum. See Fig. 7.18 for a typical shape (and Exercises6.2.3and6.2.4). Ourstrategywillbetochoosethepath Csothatitrunsover thesaddlepoint,whichgivesthedominantcontribution,andinthevalleyselsewhere. If there are several saddle points, we treat each alike, and their contributions will add to I(s→∞). To prove that there are no peaks, assume there is one at z0. That is,|F(z0)|2>|F(z)|2 forallzofaneighborhood |z−z0|≤r.I f F(z)=∞summationdisplay n=0an(z−z0)n is the Taylor expansion at z0, the mean value m(F)on the circle z=z0+rexp(iϕ)be- comes m(F)≡1 2πintegraldisplay2π 0vextendsinglevextendsingleFparenleftbig z0+reiϕparenrightbigvextendsinglevextendsingle2dϕ =1 2πintegraldisplay2π 0∞summationdisplay m,n=0a∗ manrm+nei(n−m)ϕdϕ =∞summationdisplay n=0|an|2r2n≥|a0|2=vextendsinglevextendsingleF(z0)vextendsinglevextendsingle2, (7.101) usingorthogonality,1 2πintegraltext2π 0expi(n−m)ϕdϕ=δnm.Sincem(F)isthemeanvalueof |F|2 onthecircleofradius r,theremustbeapoint z1onitsothat|F(z1)|2≥m(F)≥|F(z0)|2, whichcontradictsourassumption.Hencetherecanbenosuchpeak. Next, let us assume there is a minimum at z0so that 0<|F(z0)|2<|F(z)|2for allzof aneighborhoodof z0.Inotherwords,thedipinthevalleydoesnotgodowntothecomplex plane. Then|F(z)|2>0 and, since 1 /F(z)is analytic there, it has a Taylor expansion and z0would be a peak of 1 /|F(z)|2, which is impossible. This proves Jensen’s theorem. We nowturnour attentionbacktotheintegralinEq. (7.100). Saddle Point Method Since each saddle point z0necessarily lies above the complex plane, that is, |F(z0)|2>0, we write Fin exponential form, ef(z,s), in its vicinity without loss of generality. Note that having no zero in the complex plane is a characteristic property of the exponential function. Moreover, any saddle point with F(z)=0 becomes a trough of |F(z)|2because |F(z)|2≥0.A case in point is the function z2atz=0, where d(z2)/dz=2z=0.Here z2=(x+iy)2=x2−y2+2ixy,and2xyhasasaddlepointat z=0,andsohas x2−y2, but|z|4hasatroughthere. Atz0thetangentialplaneishorizontal;thatis,∂F ∂z|z=z0=0,orequivalently∂f ∂z|z=z0=0. This condition locates the saddle point. Our next goal is to determine the direction of steepestdescent. Atz0,fhasapowerseries f(z)=f(z0)+1 2f′′(z0)(z−z0)2+···, (7.102) 7.3 Method of Steepest Descents 491 FIGURE 7.18Asaddlepoint. or f(z)=f(z0)+1 2parenleftbig f′′(z0)+εparenrightbig (z−z0)2, (7.103) upon collecting all higher powers in the (small) ε. Let us take f′′(z0)/negationslash=0 for simplicity. Then f′′(z0)(z−z0)2=−t2,treal, (7.104) defines a line through z0(saddle point axisin Fig. 7.18). At z0,t=0. Along the axis ℑf′′(z0)(z−z0)2is zero and v=ℑf(z)≈ℑf(z0)is constant if εin Eq. (7.103) is ne- glected.Equation(7.104)canalsobeexpressedintermsof angles, arg(z−z0)=π 2−1 2argf′′(z0)=constant. (7.105) Since|F(z)|2=exp(2ℜf)varies monotonically with ℜf,|F(z)|2≈exp(−t2)falls off exponentiallyfromitsmaximumat t=0alongthisaxis.Hencethename steepestdescent . Thelinethrough z0definedby f′′(z0)(z−z0)2=+t2(7.106) isorthogonaltothisaxis( dashedinFig.7.18),whichis evidentfromits angle, arg(z−z0)=−1 2argf′′(z0)=constant, (7.107) whencomparedwithEq. (7.105).Here |F(z)|2growsexponentially. 492 Chapter 7 Functions of a Complex Variable II The curvesℜf(z)=ℜf(z0)go through z0,s oℜ[(f′′(z0)+ε)(z−z0)2]=0, or (f′′(z0)+ε)(z−z0)2=itfor realt. Expressingthisinanglesas arg(z−z0)=π 4−1 2argparenleftbig f′′(z0)+εparenrightbig ,t>0, (7.108a) arg(z−z0)=−π 4−1 2argparenleftbig f′′(z0)+εparenrightbig ,t<0, (7.108b) and comparing with Eqs. (7.105) and (7.107) we note that these curves ( dot-dashed in Fig. 7.18) divide the saddle point region into four sectors, two with ℜf(z)>ℜf(z0) (hence|F(z)|>|F(z0)|), shown shaded in Fig. 7.18, and two with ℜf(z)<ℜf(z0) (hence|F(z)|<|F(z0)|). They are at±π 4angles from the axis. Thus, the integration path hastoavoidtheshadedareas,where |F|rises.Ifapathischosentorunuptheslopesabove thesaddlepoint,thelargeimaginarypartof f(z)leadstorapidoscillationsof F(z)=ef(z) andcancellingcontributionstotheintegral. So far, our treatment has been general , except for f′′(z0)/negationslash=0, which can be relaxed. Nowwearereadyto specializetheintegrand Ffurtherinordertotieupthepathselection withtheasymptoticbehavioras s→∞. We assume that sappears linearly in the exponent, that is, we replace exp f(z,s)→ exp(sf(z)). This dependence on sensures that the saddle point contribution at z0grows withs→∞providingsteepslopes,asisthecaseinmostapplicationsinphysics.Inorder to account for the region far away from the saddle point that is not influenced by s,w e include another analytic function, g(z), which varies slowly near the saddle point and is independentof s. Altogether,then, ourintegralhasthemoreappropriateandspecificform I(s)=integraldisplay Cg(z)esf(z)dz. (7.109) The path of steepest descent is the saddle point axis when we neglect the higher-order terms,ε, in Eq. (7.103). With ε, the path of steepest descent is the curve close to the axis within the unshaded sectors, where v=ℑf(z)is strictly constant, while ℑf(z)is only approximately constant on the axis. We approximate I(s)by the integral along the piece oftheaxisinsidethepatchinFig.7.18, where(comparewithEq. (7.104)) z=z0+xeiα,α=π 2−1 2argf′′(z0), a≤x≤b. (7.110) Wefind I(s)≈eiαintegraldisplayb agparenleftbig z0+xeiαparenrightbig expbracketleftbig sfparenleftbig z0+xeiαparenrightbigbracketrightbig dx, (7.111a) and the omitted part is small and can be estimated because ℜ(f(z)−f(z0))has an upper negative bound, −Rsay, that depends on the size of the saddle point patch in Fig. 7.18 (that is, the values of a,bin Eq. (7.110)) that we choose. In Eq. (7.111) we use the power expansions fparenleftbig z0+xeiαparenrightbig =f(z0)+1 2f′′(z0)e2iαx2+···, (7.111b) gparenleftbig z0+xeiαparenrightbig =g(z0)+g′(z0)eiαx+···, 7.3 Method of Steepest Descents 493 andrecallfromEq. (7.110)that 1 2f′′(z0)e2iα=−1 2vextendsinglevextendsinglef′′(z0)vextendsinglevextendsingle<0. Wefindfortheleadingtermfor s→∞: I(s)=g(z0)esf(z0)+iαintegraldisplayb ae−1 2s|f′′(z0)|x2dx. (7.112) Since the integrand in Eq. (7.112) is essentially zero when xdeparts appreciably from the origin, we let b→∞anda→−∞. The small error involved is straightforward to estimate.Notingthattheremainingintegralis justaGausserror integral, integraldisplay∞ −∞e−1 2a2x2dx=1 aintegraldisplay∞ −∞e−1 2x2dx=√ 2π a, wefinallyobtain I(s)=√ 2πg(z0)esf(z0)eiα |sf′′(z0)|1/2, (7.113) wherethephase αwas introducedinEqs. (7.110)and(7.105). A note of warning: We assumed that the only significant contribution to the integral came from the immediate vicinity of the saddle point(s) z=z0. This condition must be checkedforeachnewproblem(Exercise7.3.5). Example 7.3.1 ASYMPTOTIC FORM OF THE HANKEL FUNCTION H(1) ν(s) InSection11.4itisshownthattheHankelfunctions,whichsatisfyBessel’sequation,may bedefinedby H(1) ν(s)=1 πiintegraldisplay∞eiπ C1,0e(s/2)(z−1/z)dz zν+1, (7.114) H(2) ν(s)=1 πiintegraldisplay0 C2,∞e−iπe(s/2)(z−1/z)dz zν+1. (7.115) The contour C1is the curve in the upper half-plane of Fig. 7.19. The contour C2is in the lower half-plane. We apply the method of steepest descents to the first Hankel function, H(1) ν(s), whichis convenientlyintheformspecifiedbyEq.(7.109), with f(z)givenby f(z)=1 2parenleftbigg z−1 zparenrightbigg . (7.116) Bydifferentiating,weobtain f′(z)=1 2+1 2z2. (7.117) 494 Chapter 7 Functions of a Complex Variable II FIGURE 7.19Hankelfunctioncontours. Settingf′(z)=0,weobtain z=i,−i. (7.118) Hence there are saddle points at z=+iandz=−i.A tz=i, f′′(i)=−i,or argf′′(i)= −π/2,so the saddle point direction is given by Eq. (7.110) as α=π 2+π 4=3 4π.For the integralfor H(1) ν(s)wemustchoosethecontourthroughthepoint z=+isothatitstartsat theorigin,movesouttangentiallytothepositiverealaxis,andthenmovesaroundthrough the saddle point at z=+iin the direction given by the angle α=3π/4 and then on out to minusinfinity,asymptoticwiththenegativerealaxis.Thepathofsteepestascent,whichwe must avoid, has the phase −1 2argf′′(i)=π 4,according to Eq. (7.107), and is orthogonal totheaxis, ourpathof steepestdescent. DirectsubstitutionintoEq. (7.113)with α=3π/4 nowyields H(1) ν(s)=1 πi√ 2πi−ν−1e(s/2)(i−1/i)e3πi/4 |(s/2)(−2/i3)|1/2 =radicalbigg 2 πse(iπ/2)(−ν−2)eisei(3π/4). (7.119) Bycombiningterms, weobtain H(1) ν(s)≈radicalbigg 2 πsei(s−ν(π/2)−π/4)(7.120) astheleadingtermoftheasymptoticexpansionoftheHankelfunction H(1) ν(s).Additional terms, if desired, may be picked up from the power series of fandgin Eq. (7.111b). The otherHankelfunctioncanbetreatedsimilarlyusingthesaddlepointat z=−i. /squaresolid Example 7.3.2 ASYMPTOTIC FORM OF THE FACTORIAL FUNCTION Ŵ(1+s) In many physical problems, particularly in the field of statistical mechanics, it is desir- able to have an accurate approximation of the gamma or factorial function of very large 7.3 Method of Steepest Descents 495 numbers. As developed in Section 8.1, the factorial function may be defined by the Euler integral Ŵ(1+s)=integraldisplay∞ 0ρse−ρdρ=ss+1integraldisplay∞ 0es(lnz−z)dz. (7.121) Here we have made the substitution ρ=zsin order to convert the integral to the form requiredbyEq.(7.109).Asbefore,weassumethat sisrealandpositive,fromwhichitfol- lowsthattheintegrandvanishesatthelimits0and ∞.Bydifferentiatingthe z-dependence appearingintheexponent,weobtain df(z) dz=d dz(lnz−z)=1 z−1,f′′(z)=−1 z2, (7.122) whichshowsthatthepoint z=1isasaddlepointandarg f′′(1)=arg(−1)=π.According toEq.(7.109) welet z−1=xeiα,α=π 2−1 2argf′′(1)=π 2−π 2=0, (7.123) withxsmall, to describe the contour in the vicinity of the saddle point. From this we see thatthedirectionofsteepestdescentisalongtherealaxis,aconclusionthatwecouldhave reachedmoreorless intuitively. DirectsubstitutionintoEq. (7.113)with α=0n o wg i v e s Ŵ(1+s)≈√ 2πss+1e−s |s(−1−2)|1/2. (7.124) Thusthefirsttermintheasymptoticexpansionofthefactorialfunctionis Ŵ(1+s)≈√ 2πssse−s. (7.125) This result is the first term in Stirling’s expansion of the factorial function. The method of steepest descent is probably the easiest way of obtaining this first term. If more terms in theexpansionaredesired,thenthemethodof Section8.3 ispreferable. /squaresolid In the foregoing example the calculation was carried out by assuming sto be real. This assumption is not necessary. We may show (Exercise 7.3.6) that Eq. (7.125) also holds whensis replaced by the complex variable w, provided only that the real part of wbe requiredtobelargeandpositive. Asymptotic limits of integral representations of functions are extremely important in manyapproximationsandapplicationsinphysics : integraldisplay Cg(z)esf(z)dz∼√ 2πg(z0)esf(z0)eiα radicalbig |sf′′(z0)|,f′(z0)=0. The saddle point method is one method of choice for deriving them and belongs in the toolkitofeveryphysicistandengineer. 496 Chapter 7 Functions of a Complex Variable II Exercises 7.3.1 Usingthemethodofsteepestdescents,evaluatethesecondHankelfunction,givenby H(2) ν(s)=1 πiintegraldisplay0 −∞C2e(s/2)(z−1/z)dz zν+1, withcontour C2asshowninFig.7.19. ANS.H(2) ν(s)≈radicalbigg 2 πse−i(s−π/4−νπ/2). 7.3.2 Find the steepest path and leading asymptotic expansion for the Fresnel integralsintegraltexts 0cosx2dx,integraltexts 0sinx2dx. Hint.Useintegraltext1 0eisz2dz. 7.3.3 (a) In applying the method of steepest descent to the Hankel function H(1) ν(s),s h o w that ℜbracketleftbig f(z)bracketrightbig <ℜbracketleftbig f(z0)bracketrightbig =0 forzonthecontour C1butawayfromthepoint z=z0=i. (b) Showthat ℜbracketleftbig f(z)bracketrightbig >0for0<r<1,  π 2<θ≤π −π≤θ<π 2 and ℜbracketleftbig f(z)bracketrightbig <0forr>1,−π 2<θ<π 2 (Fig. 7.20). This is why C1may not be deformed to pass through the second saddle point,z=−i. Comparewithandverifythedot-dashedlinesinFig.7.18forthis case. FIGURE 7.20 7.3 Additional Readings 497 7.3.4 DeterminetheasymptoticdependenceofthemodifiedBessel functions Iν(x),gi v en Iν(x)=1 2πiintegraldisplay Ce(x/2)(t+1/t)dt tν+1. The contour starts and ends at t=−∞, encircling the origin in a positive sense. There aretwosaddlepoints.Onlytheoneat z=+1contributessignificantlytotheasymptotic form. 7.3.5 Determine the asymptotic dependence of the modified Bessel function of the second kind,Kν(x),byus i ng Kν(x)=1 2integraldisplay∞ 0e(−x/2)(s+1/s)ds s1−ν. 7.3.6 ShowthatStirling’sformula, Ŵ(1+s)≈√ 2πssse−s, holdsforcomplexvaluesof s(withℜ(s)largeandpositive). Hint.Thisinvolvesassigningaphaseto sandthendemandingthat ℑ[sf(z)]=constant inthevicinityof thesaddlepoint. 7.3.7 AssumeH(1) ν(s)tohaveanegativepower-seriesexpansionof theform H(1) ν(s)=radicalbigg 2 πsei(s−ν(π/2)−π/4)∞summationdisplay n=0a−ns−n, with the coefficient of the summation obtainedby the method of steepest descent. Sub- stitute into Bessel’s equation and show that you reproduce the asymptotic series for H(1) ν(s)giveninSection11.6. AdditionalReadings N u s s e n z v e i g ,H .M . , Causality and Dispersion Relations , Mathematics in Science and Engineering Series, Vol. 95. New York: Academic Press (1972). This is an advanced text covering causality and dispersion re- lations in the first chapter and then moving on to develop the implications in a variety of areas of theoretical physics. Wyld, H. W., Mathematical Methods for Physics . Reading, MA: Benjamin/Cummings (1976), Perseus Books (1999). This is arelatively advancedtext that contains anextensive discussion of the dispersion relations. This page intentionally left blank CHAPTER 8 THEGAMMA FUNCTION (FACTORIAL FUNCTION ) The gamma function appears occasionally in physical problems such as the normalization of Coulomb wave functions and the computation of probabilities in statistical mechanics. In general, however, it has less direct physical application and interpretation than, say, the LegendreandBesselfunctionsofChapters11and12.Rather,itsimportancestemsfromits usefulnessindevelopingotherfunctionsthathavedirectphysicalapplication.Thegamma function,therefore,is includedhere. 8.1 D EFINITIONS ,SIMPLE PROPERTIES At least three different, convenient definitions of the gamma function are in common use. Ourfirsttaskistostatethesedefinitions,todevelopsomesimple,directconsequences,and toshowtheequivalenceof thethreeforms. Infinite Limit (Euler) Thefirstdefinition,namedafterEuler, is Ŵ(z)≡limn→∞1·2·3···n z(z+1)(z+2)···(z+n)nz,z/negationslash=0,−1,−2,−3,.... (8.1) This definition of Ŵ(z)is useful in developing the Weierstrass infinite-product form of Ŵ(z), Eq. (8.16), and in obtaining the derivative of ln Ŵ(z)(Section 8.2). Here and else- 499 500 Chapter 8 Gamma–Factorial Function whereinthischapter zmaybeeitherrealorcomplex.Replacing zwithz+1,wehave Ŵ(z+1)=limn→∞1·2·3···n (z+1)(z+2)(z+3)···(z+n+1)nz+1 =limn→∞nz z+n+1·1·2·3···n z(z+1)(z+2)···(z+n)nz =zŴ(z). (8.2) This is the basic functional relation for the gamma function. It should be noted that it is adifference equation. It has been shown that the gamma function is one of a general class of functions that do not satisfy any differential equation with rational coefficients. Specifically,thegammafunctionis oneof theveryfew functionsofmathematicalphysics that does not satisfy either the hypergeometric differential equation (Section 13.4) or the confluenthypergeometricequation(Section13.5). Also,from thedefinition, Ŵ(1)=limn→∞1·2·3···n 1·2·3···n(n+1)n=1. (8.3) Now,applicationofEq. (8.2) gives Ŵ(2)=1, Ŵ(3)=2Ŵ(2)=2,... (8.4) Ŵ(n)=1·2·3···(n−1)=(n−1)!. Definite Integral (Euler) Aseconddefinition,alsofrequentlycalledtheEulerintegral,is Ŵ(z)≡integraldisplay∞ 0e−ttz−1dt,ℜ(z)>0. (8.5) The restriction on zis necessary to avoid divergence of the integral. When the gamma function does appear in physical problems, it is often in this form or some variation, such as Ŵ(z)=2integraldisplay∞ 0e−t2t2z−1dt,ℜ(z)>0. (8.6) Ŵ(z)=integraldisplay1 0bracketleftbigg lnparenleftbigg1 tparenrightbiggbracketrightbiggz−1 dt,ℜ(z)>0. (8.7) Whenz=1 2, Eq. (8.6)is justtheGausserror integral,andwehavetheinterestingresult Ŵparenleftbig1 2parenrightbig =√π. (8.8) GeneralizationsofEq.(8.6),theGaussianintegrals,areconsideredinExercise8.1.11.This definiteintegralform of Ŵ(z), Eq. (8.5), leadstothebetafunction,Section8.4. 8.1 Definitions, Simple Properties 501 To show the equivalence of these two definitions, Eqs. (8.1) and (8.5), consider the functionoftwovariables F(z,n)=integraldisplayn 0parenleftbigg 1−t nparenrightbiggn tz−1dt,ℜ(z)>0, (8.9) withnapositiveinteger.1Since limn→∞parenleftbigg 1−t nparenrightbiggn ≡e−t, (8.10) fromthedefinitionof theexponential limn→∞F(z,n)=F(z,∞)=integraldisplay∞ 0e−ttz−1dt≡Ŵ(z) (8.11) byEq.(8.5). Returningto F(z,n),weevaluateitinsuccessiveintegrationsbyparts.Forconvenience letu=t/n. Then F(z,n)=nzintegraldisplay1 0(1−u)nuz−1du. (8.12) Integratingbyparts, weobtain F(z,n) nz=(1−u)nuz zvextendsinglevextendsinglevextendsinglevextendsingle1 0+n zintegraldisplay1 0(1−u)n−1uzdu. (8.13) Repeating this with the integrated part vanishing at both endpoints each time, we finally get F(z,n)=nzn(n−1)···1 z(z+1)···(z+n−1)integraldisplay1 0uz+n−1du =1·2·3···n z(z+1)(z+2)···(z+n)nz. (8.14) Thisis identicalwiththeexpressionontherightsideof Eq.(8.1). Hence limn→∞F(z,n)=F(z,∞)≡Ŵ(z), (8.15) byEq.(8.1), completingtheproof. Infinite Product (Weierstrass) Thethirddefinition(Weierstrass’ form)is 1 Ŵ(z)≡zeγz∞productdisplay n=1parenleftbigg 1+z nparenrightbigg e−z/n, (8.16) 1Theform of F(z,n)is suggested by the betafunction (compare Eq.(8.60)). 502 Chapter 8 Gamma–Factorial Function whereγis theEuler–Mascheroniconstant, γ=0.5772156619 .... (8.17) This infinite-product form may be used to develop the reflection identity, Eq. (8.23), and appliedintheexercises,suchasExercise8.1.17.Thisformcanbederivedfromtheoriginal definition(Eq. (8.1)) byrewritingitas Ŵ(z)=limn→∞1·2·3···n z(z+1)···(z+n)nz=limn→∞1 znproductdisplay m=1parenleftbigg 1+z mparenrightbigg−1 nz.(8.18) InvertingEq.(8.18) andusing n−z=e(−lnn)z, (8.19) weobtain 1 Ŵ(z)=zlimn→∞e(−lnn)znproductdisplay m=1parenleftbigg 1+z mparenrightbigg . (8.20) Multiplyinganddividingby expbracketleftbiggparenleftbigg 1+1 2+1 3+···+1 nparenrightbigg zbracketrightbigg =nproductdisplay m=1ez/m, (8.21) weget 1 Ŵ(z)=zbraceleftbigg limn→∞expbracketleftbiggparenleftbigg 1+1 2+1 3+···+1 n−lnnparenrightbigg zbracketrightbiggbracerightbigg ×bracketleftbigg limn→∞nproductdisplay m=1parenleftbigg 1+z mparenrightbigg e−z/mbracketrightbigg . (8.22) AsshowninSection5.2,theparenthesisintheexponentapproachesalimit,namely γ,the Euler–Mascheroniconstant.HenceEq. (8.16)follows. It was shown in Section 5.11 that the Weierstrass infinite-product definition of Ŵ(z)led directlytoanimportantidentity, Ŵ(z)Ŵ(1−z)=π sinzπ. (8.23) Alternatively,wecanstart fromtheproductofEuler integrals, Ŵ(z+1)Ŵ(1−z)=integraldisplay∞ 0sze−sdsintegraldisplay∞ 0t−ze−tdt =integraldisplay∞ 0vzdv (v+1)2integraldisplay∞ 0e−uudu=πz sinπz, transforming from the variables s,ttou=s+t,v=s/t, as suggested by combining the exponentialsandthepowersintheintegrands.TheJacobianis J=−vextendsinglevextendsinglevextendsinglevextendsingle11 1 t−s t2vextendsinglevextendsinglevextendsinglevextendsingle=s+t t2=(v+1)2 u, 8.1 Definitions, Simple Properties 503 where(v+1)t=u. The integralintegraltext∞ 0e−uudu=1, while that over vmay be derived by contourintegration,givingπz sinπz. This identity may also be derived by contour integration (Example 7.1.6 and Exer- cises 7.1.18 and 7.1.19) and the beta function, Section 8.4. Setting z=1 2in Eq. (8.23), weobtain Ŵparenleftbig1 2parenrightbig =√π (8.24a) (takingthepositivesquareroot),inagreementwithEq. (8.8). Similarlyonecanestablish Legendre’sduplicationformula , Ŵ(1+z)Ŵparenleftbig z+1 2parenrightbig =2−2z√πŴ(2z+1). (8.24b) The Weierstrass definition shows immediately that Ŵ(z)has simple poles at z= 0,−1,−2,−3,...andthat[Ŵ(z)]−1hasnopolesinthefinitecomplexplane,whichmeans thatŴ(z)hasnozeros.ThisbehaviormayalsobeseeninEq.(8.23),inwhichwenotethat π/(sinπz)is neverequaltozero. Actually the infinite-product definition of Ŵ(z)may be derived from the Weierstrass factorization theorem with the specification that [Ŵ(z)]−1have simple zeros at z= 0,−1,−2,−3,....The Euler–Mascheroni constant is fixed by requiring Ŵ(1)=1. See alsotheproductsexpansionsofentirefunctionsinSection7.1. In probabilitytheorythegammadistribution(probabilitydensity)isgivenby f(x)=  1 βαŴ(α)xα−1e−x/β,x>0 0,x ≤0.(8.24c) The constant[βαŴ(α)]−1is chosen so that the total (integrated) probability will be unity. Forx→E,kineticenergy, α→3 2,andβ→kT,Eq.(8.24c)yieldstheclassicalMaxwell– Boltzmannstatistics. Factorial Notation So far this discussion has been presented in terms of the classical notation. As pointed out by Jeffreys andothers, the −1o ft h ez−1 exponentin our seconddefinition(Eq. (8.5)) is acontinualnuisance.Accordingly,Eq. (8.5) issometimesrewrittenas integraldisplay∞ 0e−ttzdt≡z!,ℜ(z)>−1, (8.25) todefinea factorial function z!. Occasionally we may still encounter Gauss’ notation,producttext(z), for thefactorialfunction: productdisplay (z)=z!=Ŵ(z+1). (8.26) TheŴnotation is due to Legendre. The factorial function of Eq. (8.25) is related to the gammafunctionby Ŵ(z)=(z−1)!orŴ(z+1)=z!. (8.27) 504 Chapter 8 Gamma–Factorial Function FIGURE 8.1Thefactorial function—extensiontonegative arguments. Ifz=n,apositiveinteger(Eq. (8.4)) showsthat z!=n!=1·2·3···n, (8.28) thefamiliarfactorial.However,itshouldbenotedthatsince z!isnowdefinedbyEq.(8.25) (orequivalentlybyEq.(8.27))thefactorialfunctionisnolongerlimitedtopositiveintegral valuesof theargument(Fig. 8.1). Thedifferencerelation(Eq. (8.2)) becomes (z−1)!=z! z. (8.29) Thisshowsimmediatelythat 0!=1 (8.30) and n!=±∞ forn,anegative integer. (8.31) Intermsof thefactorial, Eq.(8.23) becomes z!(−z)!=πz sinπz. (8.32) Byrestrictingourselvestotherealvaluesoftheargument,wefindthat Ŵ(x+1)defines thecurvesshowninFigs.8.1and8.2.The minimumofthecurveis Ŵ(x+1)=x!=(0.46163...)!=0.88560.... (8.33a) 8.1 Definitions, Simple Properties 505 FIGURE 8.2Thefactorialfunctionandthefirsttwoderivativesof ln(Ŵ(x+1)). Double Factorial Notation Inmanyproblemsofmathematicalphysics,particularlyinconnectionwithLegendrepoly- nomials (Chapter 12), we encounter products of the odd positive integers and products of the even positive integers. For convenience these are given special labels as double facto- rials: 1·3·5···(2n+1)=(2n+1)!! 2·4·6···(2n)=(2n)!!.(8.33b) Clearly,thesearerelatedtotheregularfactorialfunctionsby (2n)!!=2nn!and(2n+1)!!=(2n+1)! 2nn!. (8.33c) Wealsodefine (−1)!!=1,aspecialcasethatdoesnotfollowfromEq. (8.33c). Integral Representation An integral representation that is useful in developing asymptotic series for the Bessel functionsisintegraldisplay Ce−zzνdz=parenleftbig e2πiν−1parenrightbig Ŵ(ν+1), (8.34) whereCis the contour shown in Fig. 8.3. This contour integral representation is only useful when νis not an integer, z=0 then being a branch point . Equation (8.34) may be 506 Chapter 8 Gamma–Factorial Function FIGURE 8.3Factorialfunctioncontour. FIGURE 8.4Thecontourof Fig.8.3deformed. readily verified for ν>−1 by deforming the contour as shown in Fig. 8.4. The integral from∞into the origin yields −(ν!), placing the phase of zat 0. The integral out to ∞(in the fourth quadrant) then yields e2πiνν!, the phase of zhaving increased to 2 π. Since the circlearoundtheorigincontributesnothingwhen ν>−1,Eq. (8.34)follows. It isoftenconvenienttocastthisresultintoa moresymmetricalform: integraldisplay Ce−z(−z)νdz=2iŴ(ν+1)sin(νπ). (8.35) This analysis establishes Eqs. (8.34) and (8.35) for ν>−1. It is relatively simple to extend the range to include all nonintegral ν. First, we note that the integral exists for ν<−1 as long as we stay away from the origin. Second, integrating by parts we find thatEq.(8.35)yieldsthefamiliardifferencerelation(Eq.(8.29)).Ifwetakethedifference relation to define the factorial function of ν<−1, then Eqs. (8.34) and (8.35) are verified forallν(exceptnegativeintegers). Exercises 8.1.1 Derivetherecurrencerelations Ŵ(z+1)=zŴ(z) from theEulerintegral(Eq. (8.5)), Ŵ(z)=integraldisplay∞ 0e−ttz−1dt. 8.1 Definitions, Simple Properties 507 8.1.2 In a power-series solution for the Legendre functions of the second kind we encounter theexpression (n+1)(n+2)(n+3)···(n+2s−1)(n+2s) 2·4·6·8···(2s−2)(2s)·(2n+3)(2n+5)(2n+7)···(2n+2s+1), inwhich sisa positiveinteger.Rewritethis expressionintermsoffactorials. 8.1.3 Showthat,as s−n→negativeinteger, (s−n)! (2s−2n)!→(−1)n−s(2n−2s)! (n−s)!. Heresandnare integers with s<n. This result can be used to avoid negative facto- rials, such as in the series representations of the spherical Neumann functions and the Legendrefunctionsof thesecondkind. 8.1.4 Showthat Ŵ(z)maybewritten Ŵ(z)=2integraldisplay∞ 0e−t2t2z−1dt,ℜ(z)>0, Ŵ(z)=integraldisplay1 0bracketleftbigg lnparenleftbigg1 tparenrightbiggbracketrightbiggz−1 dt,ℜ(z)>0. 8.1.5 In a Maxwellian distribution the fraction of particles with speed between vandv+dv is dN N=4πparenleftbiggm 2πkTparenrightbigg3/2 expparenleftbigg −mv2 2kTparenrightbigg v2dv, Nbeingthetotalnumberofparticles.Theaverageorexpectationvalueof vnisdefined as/angbracketleftvn/angbracketright=N−1integraltext vndN.Showthat angbracketleftbig vnangbracketrightbig =parenleftbigg2kT mparenrightbiggn/2Ŵparenleftbign+3 2parenrightbig Ŵ(3/2). 8.1.6 Bytransformingtheintegralintoagammafunction,showthat −integraldisplay1 0xklnxdx=1 (k+1)2,k>−1. 8.1.7 Showthat integraldisplay∞ 0e−x4dx=Ŵparenleftbigg5 4parenrightbigg . 8.1.8 Showthat lim x→0(ax−1)! (x−1)!=1 a. 8.1.9 Locatethepolesof Ŵ(z). Showthattheyaresimplepolesanddeterminetheresidues. 8.1.10 Showthattheequation x!=k,k/negationslash=0,hasaninfinitenumberof realroots. 8.1.11 Showthat 508 Chapter 8 Gamma–Factorial Function (a)integraldisplay∞ 0x2s+1expparenleftbig −ax2parenrightbig dx=s! 2as+1. (b)integraldisplay∞ 0x2sexpparenleftbig −ax2parenrightbig dx=(s−1 2)! 2as+1/2=(2s−1)!! 2s+1asradicalbiggπ a. TheseGaussianintegralsareofmajorimportanceinstatisticalmechanics. 8.1.12 (a) Developrecurrencerelationsfor (2n)!!andfor(2n+1)!!. (b) Usetheserecurrencerelationstocalculate(or todefine) 0 !!and(−1)!!. ANS. 0!!=1,(−1)!!=1. 8.1.13 Forsanonnegativeinteger,showthat (−2s−1)!!=(−1)s (2s−1)!!=(−1)s2ss! (2s)!. 8.1.14 Express thecoefficientofthe nthtermoftheexpansionof (1+x)1/2 (a) intermsoffactorials ofintegers, (b) intermsofthedoublefactorial( !!) functions. ANS.an=(−1)n+1(2n−3)! 22n−2n!(n−2)!=(−1)n+1(2n−3)!! (2n)!!,n=2,3,.... 8.1.15 Express thecoefficientofthe nthtermoftheexpansionof (1+x)−1/2 (a) intermsofthefactorialsof integers, (b) intermsofthedoublefactorial( !!) functions. ANS.an=(−1)n(2n)! 22n(n!)2=(−1)n(2n−1)!! (2n)!!,n=1,2,3,.... 8.1.16 TheLegendrepolynomialmaybewrittenas Pn(cosθ)=2(2n−1)!! (2n)!!braceleftbigg cosnθ+1 1·n 2n−1cos(n−2)θ +1·3 1·2n(n−1) (2n−1)(2n−3)cos(n−4)θ +1·3·5 1·2·3n(n−1)(n−2) (2n−1)(2n−3)(2n−5)cos(n−6)θ+···bracerightbigg . Letn=2s+1.Then Pn(cosθ)=P2s+1(cosθ)=ssummationdisplay m=0amcos(2m+1)θ. Findamintermsoffactorials anddoublefactorials. 8.1 Definitions, Simple Properties 509 8.1.17 (a) Showthat Ŵparenleftbig1 2−nparenrightbig Ŵparenleftbig1 2+nparenrightbig =(−1)nπ, wherenis aninteger. (b) Express Ŵ(1 2+n)andŴ(1 2−n)separatelyintermsof π1/2anda!!function. ANS.Ŵ(1 2+n)=(2n−1)!! 2nπ1/2. 8.1.18 Fromoneof thedefinitionsofthefactorialor gammafunction,showthat vextendsinglevextendsingle(ix)!vextendsinglevextendsingle2=πx sinhπx. 8.1.19 Provethat vextendsinglevextendsingleŴ(α+iβ)vextendsinglevextendsingle=vextendsinglevextendsingleŴ(α)vextendsinglevextendsingle∞productdisplay n=0bracketleftbigg 1+β2 (α+n)2bracketrightbigg−1/2 . Thisequationhasbeenusefulincalculationsofbetadecaytheory. 8.1.20 Showthat vextendsinglevextendsingle(n+ib)!vextendsinglevextendsingle=parenleftbiggπb sinhπbparenrightbigg1/2nproductdisplay s=1parenleftbig s2+b2parenrightbig1/2 forn,apositiveinteger. 8.1.21 Showthat |x!|≥vextendsinglevextendsingle(x+iy)!vextendsinglevextendsingle for allx.Thevariables xandyarereal. 8.1.22 Showthat vextendsingle vextendsingleŴparenleftbig1 2+iyparenrightbigvextendsinglevextendsingle2=π coshπy. 8.1.23 Theprobabilitydensityassociatedwiththenormaldistributionof statisticsisgivenby f(x)=1 σ(2π)1/2expbracketleftbigg −(x−µ)2 2σ2bracketrightbigg , with(−∞,∞)fortherangeof x.Showthat (a) themeanvalueof x,/angbracketleftx/angbracketrightisequalto µ, (b) thestandarddeviation (/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2)1/2is givenby σ. 8.1.24 Fromthegammadistribution f(x)=  1 βαŴ(α)xα−1e−x/β,x>0, 0,x ≤0, showthat (a)/angbracketleftx/angbracketright(mean)=αβ,(b)σ2(variance)≡/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2=αβ2. 510 Chapter 8 Gamma–Factorial Function 8.1.25 The wave function of a particle scattered by a Coulomb potential is ψ(r,θ).A tt h e originthewavefunctionbecomes ψ(0)=e−πγ/2Ŵ(1+iγ), whereγ=Z1Z2e2/¯hv.Showthat vextendsinglevextendsingleψ(0)vextendsinglevextendsingle2=2πγ e2πγ−1. 8.1.26 DerivethecontourintegralrepresentationofEq. (8.34), 2iν!sinνπ=integraldisplay Ce−z(−z)νdz. 8.1.27 Writeafunctionsubprogram FACT(N)(fixed-pointindependentvariable)thatwillcal- culateN!. Include provision for rejection and appropriate error message if Nis nega- tive. Note.For small integer N, direct multiplication is simplest. For large N, Eq. (8.55), Stirling’sseries wouldbeappropriate. 8.1.28 (a) Write a function subprogram to calculate the double factorial ratio (2N−1)!!/ (2N)!!.Includeprovisionfor N=0andforrejectionandanerrormessageif Nis negative.Calculateandtabulatethis ratiofor N=1(1)100. (b) Check your function subprogram calculation of 199 !!/200!!against the value ob- tainedfromStirling’sseries(Section8.3). ANS.199!! 200!!=0.056348. 8.1.29 Using either the FORTRAN-supplied GAMMA or a library-supplied subroutine for x!orŴ(x), determine the value of xfor which Ŵ(x)is a minimum (1≤x≤2)and this minimum value of Ŵ(x). Notice that although the minimum value of Ŵ(x)may be obtainedtoaboutsixsignificantfigures(singleprecision),thecorrespondingvalueof x ismuchlessaccurate.Whythisrelativelylowaccuracy? 8.1.30 The factorial function expressed in integral form can be evaluated by the Gauss– Laguerre quadrature. For a 10-point formula the resultant x!is theoretically exact for xan integer, 0 up through 19. What happens if xis not an integer? Use the Gauss– Laguerre quadrature to evaluate x!,x=0.0(0.1)2.0. Tabulate the absolute error as a functionof x. Checkvalue. x!exact−x!quadrature=0.00034 for x=1.3. 8.2 D IGAMMA AND POLYGAMMA FUNCTIONS Digamma Functions As may be noted from the three definitions in Section 8.1, it is inconvenient to deal with the derivatives of the gamma or factorial function directly. Instead, it is customary to take 8.2 Digamma and Polygamma Functions 511 thenaturallogarithmofthefactorialfunction(Eq.(8.1)),converttheproducttoasum,and thendifferentiate;thatis, Ŵ(z+1)=zŴ(z)=limn→∞n! (z+1)(z+2)···(z+n)nz(8.36) and lnŴ(z+1)=limn→∞bracketleftbig ln(n!)+zlnn−ln(z+1) −ln(z+2)−···−ln(z+n)bracketrightbig , (8.37) in which the logarithm of the limit is equal to the limit of the logarithm. Differentiating withrespectto z, weobtain d dzlnŴ(z+1)≡ψ(z+1)=limn→∞parenleftbigg lnn−1 z+1−1 z+2−···−1 z+nparenrightbigg ,(8.38) which defines ψ(z+1), the digamma function. From the definition of the Euler– Mascheroniconstant,2Eq. (8.38) mayberewrittenas ψ(z+1)=−γ−∞summationdisplay n=1parenleftbigg1 z+n−1 nparenrightbigg =−γ+∞summationdisplay n=1z n(n+z). (8.39) OneapplicationofEq.(8.39)isinthederivationoftheseriesformoftheNeumannfunction (Section11.3).Clearly, ψ(1)=−γ=−0.577215664901 ....3(8.40) Another,perhapsmoreuseful, expressionfor ψ(z)isderivedinSection8.3. Polygamma Function The digamma function may be differentiated repeatedly, giving rise to the polygamma function: ψ(m)(z+1)≡dm+1 dzm+1ln(z!) =(−1)m+1m!∞summationdisplay n=11 (z+n)m+1,m=1,2,3,.... (8.41) 2Compare Sections 5.2 and 5.9. Weadd and substractsummationtextn s=1s−1. 3γhas been computed to 1271 places by D. E. Knuth, Math. Comput. 16: 275 (1962), and to 3566 decimal places by D.W. Sweeney, ibid.17: 170 (1963). It may be of interest thatthe fraction 228/395 gives γaccuratetosix places. 512 Chapter 8 Gamma–Factorial Function Ap l o to f ψ(x+1)andψ′(x+1)is included in Fig. 8.2. Since the series in Eq. (8.41) definestheRiemannzetafunction4(withz=0), ζ(m)≡∞summationdisplay n=11 nm, (8.42) wehave ψ(m)(1)=(−1)m+1m!ζ(m+1), m=1,2,3,.... (8.43) The values of the polygamma functions of positive integral argument, ψ(m)(n+1),m a y becalculatedbyusingExercise8.2.6. In termsoftheperhapsmorecommon Ŵnotation, dn+1 dzn+1lnŴ(z)=dn dznψ(z)=ψ(n)(z). (8.44a) Maclaurin Expansion, Computation Itis nowpossibletowriteaMaclaurinexpansionfor ln Ŵ(z+1): lnŴ(z+1)=∞summationdisplay n=1zn n!ψ(n−1)(1)=−γz+∞summationdisplay n=2(−1)nzn nζ(n) (8.44b) convergent for |z|<1; forz=x, the range is−1<x≤1. Alternate forms of this series appearinExercise5.9.14.Equation(8.44b)isapossiblemeansofcomputing Ŵ(z+1)for real or complex z, but Stirling’s series (Section 8.3) is usually better, and in addition, an excellent table of values of the gamma function for complex arguments based on the use ofStirling’sseriesandtherecurrencerelation(Eq. (8.29)) is nowavailable.5 Series Summation Thedigammaandpolygammafunctionsmayalsobeusedinsummingseries.Ifthegeneral termoftheserieshastheformofarationalfraction(withthehighestpoweroftheindexin the numerator at least two less than the highest power of the index in the denominator), it maybetransformedbythemethodofpartialfractions(compareSection15.8).Theinfinite series may then be expressed as a finite sum of digamma and polygamma functions. The usefulnessofthismethoddependsontheavailabilityoftablesofdigammaandpolygamma functions. Such tables and examples of series summation are given in AMS-55, Chapter 6 (seeAdditionalReadingsforthereference). 4SeeSection5.9. For z/negationslash=0 this series maybe usedtodefine ageneralizedzetafunction. 5TableoftheGammaFunctionforComplexArguments ,AppliedMathematicsSeriesNo.34.Washington,DC:NationalBureau of Standards (1954). 8.2 Digamma and Polygamma Functions 513 Example 8.2.1 CATALAN ’SCONSTANT Catalan’sconstant,Exercise5.2.22, or β(2)ofSection5.9 isgivenby K=β(2)=∞summationdisplay k=0(−1)k (2k+1)2. (8.44c) Groupingthepositiveandnegativetermsseparatelyandstartingwithunitindex(tomatch theform of ψ(1), Eq. (8.41)), weobtain K=1+∞summationdisplay n=11 (4n+1)2−1 9−∞summationdisplay n=11 (4n+3)2. Now,quotingEq. (8.41), weget K=8 9+1 16ψ(1)parenleftbig 1+1 4parenrightbig −1 16ψ(1)parenleftbig 1+3 4parenrightbig . (8.44d) Using the values of ψ(1)from Table 6.1 of AMS-55 (see Additional Readings for the reference),weobtain K=0.91596559 .... Compare this calculation of Catalan’s constant with the calculations of Chapter 5, either directsummationoramodificationusingRiemannzetafunctionvalues. /squaresolid Exercises 8.2.1 Verifythatthefollowingtwoformsof thedigammafunction, ψ(x+1)=xsummationdisplay r=11 r−γ and ψ(x+1)=∞summationdisplay r=1x r(r+x)−γ, areequaltoeachother(for xapositiveinteger). 8.2.2 Showthat ψ(z+1)hastheseries expansion ψ(z+1)=−γ+∞summationdisplay n=2(−1)nζ(n)zn−1. 8.2.3 Forapower-seriesexpansionofln (z!),AMS-55(seeAdditionalReadingsforreference) lists ln(z!)=−ln(1+z)+z(1−γ)+∞summationdisplay n=2(−1)n[ζ(n)−1]zn n. 514 Chapter 8 Gamma–Factorial Function (a) ShowthatthisagreeswithEq. (8.44b)for |z|<1. (b) Whatistherangeofconvergenceofthisnewexpression? 8.2.4 Showthat 1 2lnparenleftbiggπz sinπzparenrightbigg =∞summationdisplay n=1ζ(2n) 2nz2n,|z|<1. Hint.TryEq. (8.32). 8.2.5 Write out a Weierstrass infinite-product definition of ln (z!). Without differentiating, showthatthis leadsdirectlytotheMaclaurinexpansionof ln (z!), Eq.(8.44b). 8.2.6 Derivethedifferencerelationfor thepolygammafunction ψ(m)(z+2)=ψ(m)(z+1)+(−1)mm! (z+1)m+1,m=0,1,2,.... 8.2.7 Showthatif Ŵ(x+iy)=u+iv, then Ŵ(x−iy)=u−iv. Thisis aspecialcaseof theSchwarzreflectionprinciple,Section6.5. 8.2.8 ThePochhammersymbol (a)nis definedas (a)n=a(a+1)···(a+n−1), (a) 0=1 (for integral n). (a) Express (a)nintermsof factorials. (b) Find (d/da)(a) nintermsof (a)nanddigammafunctions. ANS.d da(a)n=(a)nbracketleftbig ψ(a+n)−ψ(a)bracketrightbig . (c) Showthat (a)n+k=(a+n)k·(a)n. 8.2.9 Verifythefollowingspecialvaluesof the ψformof thedi-andpolygammafunctions: ψ(1)=−γ, ψ(1)(1)=ζ(2), ψ(2)(1)=−2ζ(3). 8.2.10 Derivethepolygammafunctionrecurrencerelation ψ(m)(1+z)=ψ(m)(z)+(−1)mm!/zm+1,m=0,1,2,.... 8.2.11 Verify (a)integraldisplay∞ 0e−rlnrdr=−γ. 8.2 Digamma and Polygamma Functions 515 (b)integraldisplay∞ 0re−rlnrdr=1−γ. (c)integraldisplay∞ 0rne−rlnrdr=(n−1)!+nintegraldisplay∞ 0rn−1e−rlnrdr , n=1,2,3,.... Hint.These may be verified by integration by parts, three parts, or differentiating the integralformof n!withrespectto n. 8.2.12 Diracrelativisticwavefunctionsforhydrogeninvolvefactorssuchas [2(1−α2Z2)1/2]! whereα, the fine structure constant, is1 137andZis the atomic number. Expand [2(1−α2Z2)1/2]!ina seriesof powersof α2Z2. 8.2.13 The quantummechanicaldescriptionof a particle in a Coulombfieldrequires a knowl- edge of the phase of the complex factorial function. Determine the phase of (1+ib)! for small b. 8.2.14 Thetotalenergyradiatedbyablackbodyis givenby u=8πk4T4 c3h3integraldisplay∞ 0x3 ex−1dx. Showthattheintegralinthis expressionisequalto 3 !ζ(4). [ζ(4)=π4/90=1.0823...]Thefinalresultis theStefan–Boltzmannlaw. 8.2.15 AsageneralizationoftheresultinExercise8.2.14,showthat integraldisplay∞ 0xsdx ex−1=s!ζ(s+1),ℜ(s)>0. 8.2.16 The neutrino energy density (Fermi distribution) in the early history of the universe is givenby ρν=4π h3integraldisplay∞ 0x3 exp(x/kT)+1dx. Showthat ρν=7π5 30h3(kT)4. 8.2.17 Provethat integraldisplay∞ 0xsdx ex+1=s!parenleftbig 1−2−sparenrightbig ζ(s+1),ℜ(s)>0. Exercises8.2.15and8.2.17actuallyconstituteMellinintegraltransforms(compareSec- tion15.1). 8.2.18 Provethat ψ(n)(z)=(−1)n+1integraldisplay∞ 0tne−zt 1−e−tdt,ℜ(z)>0. 516 Chapter 8 Gamma–Factorial Function 8.2.19 Usingdi-andpolygammafunctions,sumtheseries (a)∞summationdisplay n=11 n(n+1),(b)∞summationdisplay n=21 n2−1. Note.YoucanuseExercise8.2.6tocalculatetheneededdigammafunctions. 8.2.20 Showthat ∞summationdisplay n=11 (n+a)(n+b)=1 (b−a)braceleftbig ψ(1+b)−ψ(1+a)bracerightbig , wherea/negationslash=band neither anorbis a negative integer. It is of some interest to compare thissummationwiththecorrespondingintegral, integraldisplay∞ 1dx (x+a)(x+b)=1 b−abraceleftbig ln(1+b)−ln(1+a)bracerightbig . Therelationbetween ψ(x)and lnxismadeexplicitinEq. (8.51)inthenextsection. 8.2.21 Verifythecontourintegralrepresentationof ζ(s), ζ(s)=−(−s)! 2πiintegraldisplay C(−z)s−1 ez−1dz. Thecontour CisthesameasthatforEq.(8.35).Thepoints z=±2nπi, n=1,2,3,..., areallexcluded. 8.2.22 Show that ζ(s)is analytic in the entire finite complex plane except at s=1, where it hasasimplepolewitharesidueof +1. Hint.Thecontourintegralrepresentationwillbeuseful. 8.2.23 Using the complex variable capability of FORTRAN calculate ℜ(1+ib)!,ℑ(1+ib)!, |(1+ib)!|andphase (1+ib)!forb=0.0(0.1)1.0.Plotthephaseof (1+ib)!versusb. Hint.Exercise8.2.3 offers aconvenientapproach.Youwillneedtocalculate ζ(n). 8.3 S TIRLING ’SSERIES For computation of ln (z!)for very large z(statistical mechanics) and for numerical com- putations at nonintegral values of z, a series expansion of ln (z!)in negative powers of zis desirable.Perhapsthemostelegantwayofderivingsuchanexpansionisbythemethodof steepest descents (Section 7.3). The following method, starting with a numerical integra- tionformula,doesnotrequireknowledgeofcontourintegrationandis particularlydirect. 8.3 Stirling’s Series 517 Derivation from Euler–Maclaurin Integration Formula TheEuler–Maclaurinformulafor evaluatingadefiniteintegral6is integraldisplayn 0f(x)dx=1 2f(0)+f(1)+f(2)+···+1 2f(n) −b2bracketleftbig f′(n)−f′(0)bracketrightbig −b4bracketleftbig f′′′(n)−f′′′(0)bracketrightbig −···,(8.45) inwhichthe b2narerelatedtotheBernoullinumbers B2n(compareSection5.9) by (2n)!b2n=B2n, (8.46) B0=1,B 6=1 42, B2=1 6,B 8=−1 30, B4=−1 30,B 10=5 66,andsoon .(8.47) ByapplyingEq.(8.45) tothedefiniteintegral integraldisplay∞ 0dx (z+x)2=1 z(8.48) (forznotonthenegativerealaxis),weobtain 1 z=1 2z2+ψ(1)(z+1)−2!b2 z3−4!b4 z5−···. (8.49) ThisisthereasonforusingEq.(8.48).TheEuler–Maclaurinevaluationyields ψ(1)(z+1), whichisd2lnŴ(z+1)/dz2. UsingEq.(8.46) andsolvingfor ψ(1)(z+1),weha v e ψ(1)(z+1)=d dzψ(z+1)=1 z−1 2z2+B2 z3+B4 z5+··· =1 z−1 2z2+∞summationdisplay n=1B2n z2n+1. (8.50) Since the Bernoulli numbers diverge strongly, this series does not converge. It is a semi- convergent, or asymptotic, series, useful if one retains a small enough number of terms (compareSection5.10). Integratingonce,wegetthedigammafunction ψ(z+1)=C1+lnz+1 2z−B2 2z2−B4 4z4−··· =C1+lnz+1 2z−∞summationdisplay n=1B2n 2nz2n. (8.51) IntegratingEq.(8.51)withrespectto zfromz−1tozandthenletting zapproachinfinity, C1,theconstantofintegration,maybeshowntovanish.Thisgivesusasecondexpression forthedigammafunction,oftenmoreusefulthanEq. (8.38) or(8.44b). 6This is obtainedby repeatedintegration byparts, Section5.9. 518 Chapter 8 Gamma–Factorial Function Stirling’s Series Theindefiniteintegralofthedigammafunction(Eq. (8.51)) is lnŴ(z+1)=C2+parenleftbigg z+1 2parenrightbigg lnz−z+B2 2z+···+B2n 2n(2n−1)z2n−1+···,(8.52) in which C2is another constant of integration. To fix C2we find it convenient to use the doubling,orLegendreduplication,formuladerivedinSection8.4, Ŵ(z+1)Ŵparenleftbig z+1 2parenrightbig =2−2zπ1/2Ŵ(2z+1). (8.53) Thismaybeproveddirectlywhen zisapositiveintegerbywriting Ŵ(2z+1)asaproduct of even terms times a product of odd terms and extracting a factor of 2 from each term (Exercise 8.3.5). Substituting Eq. (8.52) into the logarithm of the doubling formula, we findthatC2is C2=1 2ln2π, (8.54) giving lnŴ(z+1)=1 2ln2π+parenleftbigg z+1 2parenrightbigg lnz−z+1 12z−1 360z3+1 1260z5−···.(8.55) This is Stirling’s series, an asymptotic expansion. The absolute value of the error is less thantheabsolutevalueof thefirst termomitted. The constants of integration C1andC2may also be evaluated by comparison with the first term of the series expansion obtained by the method of “steepest descent.” This is carriedoutinSection7.3. To help convey a feeling of the remarkable precision of Stirling’s series for Ŵ(s+1), the ratio of the first term of Stirling’s approximation to Ŵ(s+1)is plotted in Fig. 8.5. A tabulation gives the ratio of the first term in the expansion to Ŵ(s+1)and the ratio of the first two terms in the expansion to Ŵ(s+1)(Table 8.1). The derivation of these forms isExercise8.3.1. Exercises 8.3.1 RewriteStirling’sseries togive Ŵ(z+1)insteadof ln Ŵ(z+1). ANS.Ŵ(z+1)=√ 2πzz+1/2e−zparenleftbigg 1+1 12z+1 288z2−139 51,840z3+···parenrightbigg . 8.3.2 Use Stirling’s formula to estimate 52 !, the number of possible rearrangements of cards inastandarddeckofplayingcards. 8.3.3 ByintegratingEq.(8.51)from z−1t ozandthenletting z→∞,evaluatetheconstant C1intheasymptoticseries forthedigammafunction ψ(z). 8.3.4 Showthattheconstant C2inStirling’sformulaequals1 2ln2πbyusingthelogarithmof thedoublingformula. 8.3 Stirling’s Series 519 FIGURE 8.5Accuracyof Stirling’sformula. Table 8.1 s1 Ŵ(s+1)√ 2πss+1/2e−s1 Ŵ(s+1)√ 2πss+1/2e−sparenleftbigg 1+1 12sparenrightbigg 1 0.92213 0.99898 2 0.95950 0.99949 3 0.97270 0.99972 4 0.97942 0.99983 5 0.98349 0.99988 6 0.98621 0.99992 7 0.98817 0.99994 8 0.98964 0.99995 9 0.99078 0.99996 10 0.99170 0.99998 8.3.5 Bydirectexpansion,verifythedoublingformulafor z=n+1 2;nis aninteger. 8.3.6 WithoutusingStirling’sseriesshowthat (a) ln(n!)<integraldisplayn+1 1lnxdx,(b)ln(n!)>integraldisplayn 1lnxdx;nis aninteger≥2. Notice that the arithmetic mean of these two integrals gives a good approximation for Stirling’sseries. 8.3.7 Test forconvergence ∞summationdisplay p=0bracketleftbigg(p−1 2)! p!bracketrightbigg2 ×2p+1 2p+2=π∞summationdisplay p=0(2p−1)!!(2p+1)!! (2p)!!(2p+2)!!. 520 Chapter 8 Gamma–Factorial Function This series arises in an attempt to describe the magnetic field created by and enclosed byacurrentloop. 8.3.8 Showthat limx→∞xb−a(x+a)! (x+b)!=1. 8.3.9 Showthat limn→∞(2n−1)!! (2n)!!n1/2=π−1/2. 8.3.10 Calculatethebinomialcoefficientparenleftbig2n nparenrightbig tosixsignificantfiguresfor n=10,20,and30. Checkyourvaluesby (a) aStirlingseriesapproximationthroughtermsin n−1, (b) adoubleprecisioncalculation. ANS.parenleftbig20 10parenrightbig =1.84756×105,parenleftbig4020parenrightbig =1.37846×1011, parenleftbig60 30parenrightbig =1.18264×1017. 8.3.11 Write a program (or subprogram) that will calculate log10(x!)directly from Stirling’s series. Assume that x≥10. (Smaller values could be calculated via the factorial re- currence relation.) Tabulate log10(x!)versusxforx=10(10)300. Check your results againstAMS-55(seeAdditionalReadingsforthisreference)orbydirectmultiplication (forn=10,20,and30). Checkvalue .l o g10(100!)=157.97. 8.3.12 UsingthecomplexarithmeticcapabilityofFORTRAN,writeasubroutinethatwillcal- culate ln(z!)for complex zbased on Stirling’s series. Include a test and an appropriate error message if zis too close to a negative real integer. Check your subroutine against alternatecalculationsfor zreal,zpureimaginary,and z=1+ib(Exercise8.2.23). Checkvalues .|(i0.5)!|=0.82618 phase(i0.5)!=−0.24406. 8.4 T HEBETA FUNCTION Using the integral definition (Eq. (8.25)), we write the product of two factorials as the product of two integrals. To facilitate a change in variables, we take the integrals over a finiterange: m!n!=lim a2→∞integraldisplaya2 0e−uumduintegraldisplaya2 0e−vvndv,ℜ(m)>−1, ℜ(n)>−1.(8.56a) Replacing uwithx2andvwithy2, weobtain m!n!=lima→∞4integraldisplaya 0e−x2x2m+1dxintegraldisplaya 0e−y2y2n+1dy. (8.56b) 8.4 The Beta Function 521 FIGURE 8.6Transformationfrom Cartesiantopolarcoordinates. Transformingtopolarcoordinatesgivesus m!n!=lima→∞4integraldisplaya 0e−r2r2m+2n+3drintegraldisplayπ/2 0cos2m+1θsin2n+1θdθ =(m+n+1)!2integraldisplayπ/2 0cos2m+1θsin2n+1θdθ. (8.57) Here the Cartesian area element dxdyhas been replaced by rdrdθ(Fig. 8.6). The last equalityinEq. (8.57) followsfromExercise8.1.11. Thedefiniteintegral,togetherwiththefactor2, hasbeennamedthebetafunction: B(m+1,n+1)≡2integraldisplayπ/2 0cos2m+1θsin2n+1θdθ =m!n! (m+n+1)!. (8.58a) Equivalently,intermsofthegammafunctionandnotingitssymmetry, B(p,q)=Ŵ(p)Ŵ(q) Ŵ(p+q),B(q,p)=B(p,q). (8.58b) Theonlyreasonforchoosing m+1andn+1,ratherthan mandn,astheargumentsof B istobeinagreementwiththeconventional,historicalbetafunction. Definite Integrals, Alternate Forms The beta function is useful in the evaluation of a wide variety of definite integrals. The substitution t=cos2θconvertsEq.(8.58a) to7 B(m+1,n+1)=m!n! (m+n+1)!=integraldisplay1 0tm(1−t)ndt. (8.59a) 7TheLaplacetransform convolution theorem provides analternate derivation ofEq. (8.58a), compare Exercise 15.11.2. 522 Chapter 8 Gamma–Factorial Function Replacing tbyx2, weobtain m!n! 2(m+n+1)!=integraldisplay1 0x2m+1parenleftbig 1−x2parenrightbigndx. (8.59b) Thesubstitution t=u/(1+u)inEq. (8.59a)yieldsstillanotherusefulform, m!n! (m+n+1)!=integraldisplay∞ 0um (1+u)m+n+2du. (8.60) The beta function as a definite integral is useful in establishing integral representations of theBesselfunction(Exercise11.1.18)andthehypergeometricfunction(Exercise13.4.10). Verification of πα/sinπαRelation If wetake m=a,n=−a,−1<a<1,then integraldisplay∞ 0ua (1+u)2du=a!(−a)!. (8.61) By contour integration this integral may be shown to be equal to πa/sinπa(Exer- cise7.1.18),thusprovidinganothermethodofobtainingEq. (8.32). Derivation of Legendre Duplication Formula The form of Eq. (8.58a) suggests that the beta function may be useful in deriving the doubling formula used in the preceding section. From Eq. (8.59a) with m=n=zand ℜ(z)>−1, z!z! (2z+1)!=integraldisplay1 0tz(1−t)zdt. (8.62) Bysubstituting t=(1+s)/2,wehave z!z! (2z+1)!=2−2z−1integraldisplay1 −1parenleftbig 1−s2parenrightbigzds=2−2zintegraldisplay1 0parenleftbig 1−s2parenrightbigzds. (8.63) The last equality holds because the integrand is even. Evaluating this integral as a beta function(Eq. (8.59b)), weobtain z!z! (2z+1)!=2−2z−1z!(−1 2)! (z+1 2)!. (8.64) Rearranging terms and recalling that (−1 2)!=π1/2, we reduce this equation to one form oftheLegendreduplicationformula, z!parenleftbig z+1 2parenrightbig !=2−2z−1π1/2(2z+1)!. (8.65a) Dividingby (z+1 2), weobtainanalternateformoftheduplicationformula: z!parenleftbig z−1 2parenrightbig !=2−2zπ1/2(2z)!. (8.65b) 8.4 The Beta Function 523 Although the integrals used in this derivation are defined only for ℜ(z)>−1, the results (Eqs. (8.65a)and(8.65b)holdforallregularpoints zbyanalyticcontinuation.8 Using the double factorial notation (Section 8.1), we may rewrite Eq. (8.65a) (with z= n,aninteger)as parenleftbig n+1 2parenrightbig !=π1/2(2n+1)!!/2n+1. (8.65c) Thisis oftenconvenientfor eliminatingfactorialsoffractions. Incomplete Beta Function Just as there is an incomplete gamma function (Section 8.5), there is also an incomplete betafunction, Bx(p,q)=integraldisplayx 0tp−1(1−t)q−1dt,0≤x≤1,p>0,q>0(ifx=1).(8.66) Clearly,Bx=1(p,q)becomes the regular (complete) beta function, Eq. (8.59a). A power- series expansion of Bx(p,q)is the subject of Exercises 5.2.18 and 5.7.8. The relation to hypergeometricfunctionsappearsinSection13.4. The incomplete beta function makes an appearance in probability theory in calculating theprobabilityofatmost ksuccessesin nindependenttrials.9 Exercises 8.4.1 Derive the doubling formula for the factorial function by integrating (sin2θ)2n+1= (2sinθcosθ)2n+1(andusingthebetafunction). 8.4.2 Verifythefollowingbetafunctionidentities: (a)B(a,b)=B(a+1,b)+B(a,b+1), (b)B(a,b)=a+b bB(a,b+1), (c)B(a,b)=b−1 aB(a+1,b−1), (d)B(a,b)B(a+b,c)=B(b,c)B(a,b+c). 8.4.3 (a) Showthat integraldisplay1 −1parenleftbig 1−x2parenrightbig1/2x2ndx=  π/2,n =0 π(2n−1)!! (2n+2)!!,n=1,2,3,.... 8If 2zis anegative integer, weget the validbut unilluminating result ∞=∞. 9W. Feller, An Introduction to Probability Theory and Its Applications , 3rd ed. NewYork: Wiley (1968), Section VI.10. 524 Chapter 8 Gamma–Factorial Function (b) Showthat integraldisplay1 −1parenleftbig 1−x2parenrightbig−1/2x2ndx=  π, n =0 π(2n−1)!! (2n)!!,n=1,2,3,.... 8.4.4 Showthat integraldisplay1 −1parenleftbig 1−x2parenrightbigndx=  22n+1n!n! (2n+1)!,n>−1 2(2n)!! (2n+1)!!,n=0,1,2,.... 8.4.5 Evaluateintegraltext1 −1(1+x)a(1−x)bdxintermsofthebetafunction. ANS. 2a+b+1B(a+1,b+1). 8.4.6 Show,bymeansofthebetafunction,that integraldisplayz tdx (z−x)1−α(x−t)α=π sinπα,0<α<1. 8.4.7 ShowthattheDirichletintegral integraldisplayintegraldisplay xpyqdxdy=p!q! (p+q+2)!=B(p+1,q+1) p+q+2, wheretherangeofintegrationisthetriangleboundedbythepositive x-andy-axesand thelinex+y=1. 8.4.8 Showthatintegraldisplay∞ 0integraldisplay∞ 0e−(x2+y2+2xycosθ)dxdy=θ 2sinθ. Whatarethelimitson θ? Hint.Consideroblique xy-coordinates. ANS.−π<θ<π . 8.4.9 Evaluate(usingthebetafunction) (a) integraldisplayπ/2 0cos1/2θdθ=(2π)3/2 16[(1 4)!]2, (b) integraldisplayπ/2 0cosnθdθ=integraldisplayπ/2 0sinnθdθ=√π[(n−1)/2]! 2(n/2)! =  (n−1)!! n!!fornodd, π 2·(n−1)!! n!!forneven. 8.4 The Beta Function 525 8.4.10 Evaluateintegraltext1 0(1−x4)−1/2dxasa betafunction. ANS.[(1 4)!]2·4 (2π)1/2=1.311028777. 8.4.11 Given Jν(z)=2 π1/2(ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplayπ/2 0sin2νθcos(zcosθ)dθ,ℜ(ν)>−1 2, show,withtheaidofbetafunctions,thatthisreducestotheBesselseries Jν(z)=∞summationdisplay s=0(−1)s1 s!(s+ν)!parenleftbiggz 2parenrightbigg2s+ν , identifying the initial Jνas an integral representation of the Bessel function, Jν(Sec- tion11.1). 8.4.12 GiventheassociatedLegendrefunction Pm m(x)=(2m−1)!!parenleftbig 1−x2parenrightbigm/2, Section12.5,showthat (a)integraldisplay1 −1bracketleftbig Pm m(x)bracketrightbig2dx=2 2m+1(2m)!,m=0,1,2,..., (b)integraldisplay1 −1bracketleftbig Pm m(x)bracketrightbig2dx 1−x2=2·(2m−1)!,m=1,2,3,.... 8.4.13 Showthat (a)integraldisplay1 0x2s+1parenleftbig 1−x2parenrightbig−1/2dx=(2s)!! (2s+1)!!, (b)integraldisplay1 0x2pparenleftbig 1−x2parenrightbigqdx=1 2(p−1 2)!q! (p+q+1 2)!. 8.4.14 Aparticleofmass mmovinginasymmetricpotentialthatiswelldescribedby V(x)= A|x|nhas a total energy1 2m(dx/dt)2+V(x)=E. Solving for dx/dtand integrating wefindthattheperiodofmotionis τ=2√ 2mintegraldisplayxmax 0dx (E−Axn)1/2, wherexmaxisaclassicalturningpointgivenby Axn max=E.Showthat τ=2 nradicalbigg 2πm EparenleftbiggE Aparenrightbigg1/nŴ(1/n) Ŵ(1/n+1 2). 8.4.15 ReferringtoExercise8.4.14, 526 Chapter 8 Gamma–Factorial Function (a) Determinethelimitas n→∞of 2 nradicalbigg 2πm EparenleftbiggE Aparenrightbigg1/nŴ(1/n) Ŵ(1/n+1 2). (b) Find limn→∞τfromthebehavioroftheintegrand (E−Axn)−1/2. (c) Investigatethebehaviorofthephysicalsystem(potentialwell)as n→∞.Obtain theperiodfrominspectionofthis limitingphysicalsystem. 8.4.16 Showthat integraldisplay∞ 0sinhαx coshβxdx=1 2Bparenleftbiggα+1 2,β−α 2parenrightbigg ,−1<α<β. Hint.Let sinh2x=u. 8.4.17 Thebetadistributionof probabilitytheoryhasaprobabilitydensity f(x)=Ŵ(α+β) Ŵ(α)Ŵ(β)xα−1(1−x)β−1, withxrestrictedtotheinterval(0,1). Showthat (a)/angbracketleftx/angbracketright(mean)=α α+β. (b)σ2(variance)≡/angbracketleftx2/angbracketright−/angbracketleftx/angbracketright2=αβ (α+β)2(α+β+1). 8.4.18 From limn→∞integraltextπ/2 0sin2nθdθ integraltextπ/2 0sin2n+1θdθ=1 derivetheWallisformulafor π: π 2=2·2 1·3·4·4 3·5·6·6 5·7···. 8.4.19 Tabulatethebetafunction B(p,q)forpandq=1.0(0.1)2.0 independently. Checkvalue. B(1.3,1.7)=0.40774. 8.4.20 (a) Write a subroutine that will calculate the incomplete beta function Bx(p,q).F o r 0.5<x≤1 youwillfinditconvenienttousetherelation Bx(p,q)=B(p,q)−B1−x(q,p). (b) Tabulate Bx(3 2,3 2). Spot check your results by using the Gauss–Legendre quadra- ture. 8.5 Incomplete Gamma Function 527 8.5 T HEINCOMPLETE GAMMA FUNCTIONS AND RELATED FUNCTIONS Generalizing the Euler definition of the gamma function (Eq. (8.5)), we define the incom- pletegammafunctionsbythevariablelimitintegrals γ(a,x)=integraldisplayx 0e−tta−1dt,ℜ(a)>0 and Ŵ(a,x)=integraldisplay∞ xe−tta−1dt. (8.67) Clearly,thetwofunctionsarerelated,for γ(a,x)+Ŵ(a,x)=Ŵ(a). (8.68) Thechoiceofemploying γ(a,x)orŴ(a,x)ispurelyamatterofconvenience.Ifthepara- meterais apositiveinteger,Eq. (8.67) maybeintegratedcompletelytoyield γ(n,x)=(n−1)!parenleftbigg 1−e−xn−1summationdisplay s=0xs s!parenrightbigg Ŵ(n,x)=(n−1)!e−xn−1summationdisplay s=0xs s!,n=1,2,....(8.69) Fornonintegral a,apower-seriesexpansionof γ(a,x)forsmall xandanasymptoticex- pansionof Ŵ(a,x)(denotedas I(x,p))aredevelopedinExercise5.7.7andSection5.10: γ(a,x)=xa∞summationdisplay n=0(−1)nxn n!(a+n),|x|∼1(smallx), Ŵ(a,x)=xa−1e−x∞summationdisplay n=0(a−1)! (a−1−n)!·1 xn(8.70) =xa−1e−x∞summationdisplay n=0(−1)n(n−a)! (−a)!·1 xn,x≫1(largex). Theseincompletegammafunctionsmayalsobeexpressedquiteelegantlyintermsofcon- fluenthypergeometricfunctions(compareSection13.5). Exponential Integral Although the incomplete gamma function Ŵ(a,x)in its general form (Eq. (8.67)) is only infrequently encountered in physical problems, a special case is quite common and very 528 Chapter 8 Gamma–Factorial Function FIGURE 8.7Theexponentialintegral, E1(x)=−Ei(−x). useful.Wedefinetheexponentialintegralby10 −Ei(−x)≡integraldisplay∞ xe−t tdt=E1(x). (8.71) (SeeFig.8.7.)Cautionisneededhere,fortheintegralinEq.(8.71)divergeslogarithmically asx→0. Toobtainaseriesexpansionforsmall x,we startfrom E1(x)=Ŵ(0,x)=lim a→0bracketleftbig Ŵ(a)−γ(a,x)bracketrightbig . (8.72) Wemaysplitthedivergenttermintheseriesexpansionfor γ(a,x), E1(x)=lim a→0bracketleftbiggaŴ(a)−xa abracketrightbigg −∞summationdisplay n=1(−1)nxn n·n!. (8.73) Usingl’Hôpital’srule(Exercise5.6.8) and d dabraceleftbig aŴ(a)bracerightbig =d daa!=d daeln(a!)=a!ψ(a+1), (8.74) andthenEq.(8.40),11weobtaintherapidlyconvergingseries E1(x)=−γ−lnx−∞summationdisplay n=1(−1)nxn n·n!. (8.75) An asymptotic expansion E1(x)≈e−x[1 x−1! x2+···]forx→∞is developed in Sec- tion5.10. 10The appearance of the two minus signs in −Ei(−x)is a historical monstrosity. AMS-55, Chapter 5, denotes this integral as E1(x).SeeAdditional Readings for thereference. 11dxa/da=xalnx. 8.5 Incomplete Gamma Function 529 FIGURE 8.8Sineandcosineintegrals. Further special forms related to the exponential integral are the sine integral, cosine integral(Fig.8.8), andlogarithmicintegral,definedby12 si(x)=−integraldisplay∞ xsint tdt Ci(x)=−integraldisplay∞ xcost tdt (8.76) li(x)=integraldisplayx 0du lnu=Ei(lnx) fortheirprincipalbranch,withthebranchcutconventionallychosentobealongthenega- tive real axis from the branch point at zero. By transforming from real to imaginary argu- ment,wecanshowthat si(x)=1 2ibracketleftbig Ei(ix)−Ei(−ix)bracketrightbig =1 2ibracketleftbig E1(ix)−E1(−ix)bracketrightbig , (8.77) whereas Ci(x)=1 2bracketleftbig Ei(ix)+Ei(−ix)bracketrightbig =−1 2bracketleftbig E1(ix)+E1(−ix)bracketrightbig ,|argx|<π 2.(8.78) Addingthesetworelations,weobtain Ei(ix)=Ci(x)+isi(x), (8.79) to show that the relation among these integrals is exactly analogous to that among eix, cosx, and sin x. Reference to Eqs. (8.71) and (8.78) shows that Ci (x)agrees with the definitionsof AMS-55(see AdditionalReadingsfor thereference).In termsof E1, E1(ix)=−Ci(x)+isi(x). Asymptotic expansions of Ci (x)and si(x)are developed in Section 5.10. Power-series expansions about the origin for Ci (x),s i(x), and li(x)may be obtained from those for 12Another sine integral is given by Si (x)=si(x)+π/2. 530 Chapter 8 Gamma–Factorial Function FIGURE 8.9Error function, erf x. the exponentialintegral, E1(x), or by direct integration, Exercise 8.5.10. The exponential, sine, and cosine integrals are tabulated in AMS-55, Chapter 5, (see Additional Readings for the reference) and can also be accessed by symbolic software such as Mathematica, Maple,Mathcad,andReduce. Error Integrals Theerror integrals erfz=2√πintegraldisplayz 0e−t2dt,erfcz=1−erfz=2√πintegraldisplay∞ ze−t2dt (8.80a) (normalized so that erf ∞=1) are introduced in Exercise 5.10.4 (Fig. 8.9). Asymptotic forms are developed there. From the general form of the integrands and Eq. (8.6) we ex- pect that erf zand erfczmay be written as incomplete gamma functions with a=1 2.T h e relationsare erfz=π−1/2γparenleftbig1 2,z2parenrightbig ,erfcz=π−1/2Ŵparenleftbig1 2,z2parenrightbig . (8.80b) Thepower-seriesexpansionof erf zfollowsdirectlyfrom Eq.(8.70). Exercises 8.5.1 Showthat γ(a,x)=e−x∞summationdisplay n=0(a−1)! (a+n)!xa+n (a) byrepeatedlyintegratingbyparts. (b) DemonstratethisrelationbytransformingitintoEq.(8.70). 8.5.2 Showthat (a)dm dxmbracketleftbig x−aγ(a,x)bracketrightbig =(−1)mx−a−mγ(a+m,x), 8.5 Incomplete Gamma Function 531 (b)dm dxmbracketleftbig exγ(a,x)bracketrightbig =exŴ(a) Ŵ(a−m)γ(a−m,x). 8.5.3 Showthat γ(a,x)andŴ(a,x)satisfy therecurrencerelations (a)γ(a+1,x)=aγ(a,x)−xae−x, (b)Ŵ(a+1,x)=aŴ(a,x)+xae−x. 8.5.4 Thepotentialproducedbya 1 Shydrogenelectron(Exercise12.8.6) isgivenby V(r)=q 4πε0a0braceleftbigg1 2rγ(3,2r)+Ŵ(2,2r)bracerightbigg . (a) For r≪1,showthat V(r)=q 4πε0a0braceleftbigg 1−2 3r2+···bracerightbigg . (b) For r≫1,showthat V(r)=q 4πε0a0·1 r. Hererisexpressedinunitsof a0,theBohrradius. Note.Forcomputationatintermediatevaluesof r, Eqs. (8.69) areconvenient. 8.5.5 Thepotentialof a 2 Phydrogenelectronisfoundtobe(Exercise12.8.7) V(r)=1 4πε0·q 24a0braceleftbigg1 rγ(5,r)+Ŵ(4,r)bracerightbigg −1 4πε0·q 120a0braceleftbigg1 r3γ(7,r)+r2Ŵ(2,r)bracerightbigg P2(cosθ). Hereris expressed in units of a0, the Bohr radius. P2(cosθ)is a Legendre polynomial (Section12.1). (a) For r≪1,showthat V(r)=1 4πε0·q a0braceleftbigg1 4−1 120r2P2(cosθ)+···bracerightbigg . (b) For r≫1,showthat V(r)=1 4πε0·q a0rbraceleftbigg 1−6 r2P2(cosθ)+···bracerightbigg . 8.5.6 Provethattheexponentialintegralhastheexpansion integraldisplay∞ xe−t tdt=−γ−lnx−∞summationdisplay n=1(−1)nxn n·n!, whereγis theEuler–Mascheroniconstant. 532 Chapter 8 Gamma–Factorial Function 8.5.7 Showthat E1(z)maybewrittenas E1(z)=e−zintegraldisplay∞ 0e−zt 1+tdt. Showalsothatwemustimposethecondition |argz|≤π/2. 8.5.8 Related to the exponential integral (Eq. (8.71)) by a simple change of variable is the function En(x)=integraldisplay∞ 1e−xt tndt. Showthat En(x)satisfiestherecurrencerelation En+1(x)=1 ne−x−x nEn(x), n=1,2,3,.... 8.5.9 WithEn(x)asdefinedinExercise8.5.8,showthat En(0)=1/(n−1),n>1. 8.5.10 Developthefollowingpower-seriesexpansions: (a) si(x)=−π 2+∞summationdisplay n=0(−1)nx2n+1 (2n+1)(2n+1)!, (b) Ci(x)=γ+lnx+∞summationdisplay n=1(−1)nx2n 2n(2n)!. 8.5.11 Ananalysisof acenter-fedlinearantennaleadstotheexpression integraldisplayx 01−cost tdt. Showthatthisisequalto γ+lnx−Ci(x). 8.5.12 Usingtherelation Ŵ(a)=γ(a,x)+Ŵ(a,x), show that if γ(a,x)satisfies the relations of Exercise 8.5.2, then Ŵ(a,x)must satisfy thesamerelations. 8.5.13 (a) Writeasubroutinethatwillcalculatetheincompletegammafunctions γ(n,x)and Ŵ(n,x)fornapositiveinteger.Spotcheck Ŵ(n,x)byGauss–Laguerrequadratures. (b) Tabulate γ(n,x)andŴ(n,x)forx=0.0(0.1)1.0 andn=1,2, 3. 8.5.14 Calculatethepotentialproducedbya1 Shydrogenelectron(Exercise8.5.4)(Fig.8.10). TabulateV(r)/(q/ 4πε0a0)forx=0.0(0.1)4.0.Checkyourcalculationsfor r≪1 and forr≫1 bycalculatingthelimitingforms giveninExercise8.5.4. 8.5.15 UsingEqs.(5.182) and(8.75), calculatetheexponentialintegral E1(x)for (a)x=0.2(0.2)1.0, (b) x=6.0(2.0)10.0. Program your own calculationbut checkeach value, using a library subroutine if avail- able.AlsocheckyourcalculationsateachpointbyaGauss–Laguerrequadrature. 8.5 Additional Readings 533 FIGURE 8.10Distributedchargepotentialproduced bya 1Shydrogenelectron,Exercise8.5.14. You’ll find that the power-series converges rapidly and yields high precision for small x.Theasymptoticseries, evenfor x=10,yieldsrelativelypooraccuracy. Checkvalues. E1(1.0)=0.219384 E1(10.0)=4.15697×10−6. 8.5.16 Thetwoexpressionsfor E1(x),(1)Eq.(5.182),anasymptoticseriesand(2)Eq.(8.75), a convergent power series, provide a means of calculating the Euler–Mascheroni con- stantγto high accuracy. Using double precision, calculate γfrom Eq. (8.75), with E1(x)evaluatedbyEq.(5.182). Hint.As a convenient choice take xin the range 10 to 20. (Your choice of xwill set a limit on the accuracy of your result.) To minimize errors in the alternating series of Eq. (8.75),accumulatethepositiveandnegativetermsseparately. ANS.For x=10 and“doubleprecision,” γ=0.57721566. AdditionalReadings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Contains a wealth of information about gamma functions, incomplete gamma functions, exponential integrals, error functions, and related functions—Chapters 4 to 6. Artin,E., TheGammaFunction (translatedbyM.Butler).NewYork:Holt,RinehartandWinston(1964).Demon- strates that if a function f(x)is smooth (log convex) and equal to (n−1)!whenx=n=integer, it is the gamma function. Davis, H. T., Tables of the Higher Mathematical Functions . Bloomington, IN: Principia Press (1933). Volume I contains extensive information on the gamma function and the polygamma functions. Gradshteyn, I. S.,and I. M.Ryzhik, Table of Integrals, Series, and Products . NewYork: AcademicPress (1980). Lu k e,Y .L., The Special Functions and Their Approximations , Vol.1. NewYork: AcademicPress (1969). Luke, Y. L., Mathematical Functions and Their Approximations . New York: Academic Press (1975). This is an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables(AMS-55). Chapter1 deals with thegamma function. Chapter 4 treatsthe incomplete gamma function anda host of relatedfunctions. This page intentionally left blank CHAPTER 9 DIFFERENTIAL EQUATIONS 9.1 P ARTIAL DIFFERENTIAL EQUATIONS Introduction Inphysics theknowledgeof theforce inanequationof motionusuallyleadstoa differen- tial equation. Thus, almost all the elementary and numerous advanced parts of theoretical physics are formulated in terms of differential equations. Sometimes these are ordinary differential equations in one variable (abbreviated ODEs). More often the equations are partialdifferentialequations( PDEs) intwoormorevariables. Let us recall from calculus that the operation of taking an ordinary or partial derivative isalinearoperation (L),1 d(aϕ(x)+bψ(x)) dx=adϕ dx+bdψ dx, for ODEs involving derivatives in one variable xonly and no quadratic, (dψ/dx)2,o r higherpowers. Similarly,for partialderivations, ∂(aϕ(x,y)+bψ(x,y)) ∂x=a∂ϕ(x,y) ∂x+b∂ψ(x,y) ∂x. Ingeneral L(aϕ+bψ)=aL(ϕ)+bL(ψ). Thus,ODEsandPDEsappearaslinearoperatorequations, Lψ=F, (9.1) 1We are especially interested in linear operators because in quantum mechanics physical quantities are represented by linear operators operating in acomplex, infinite-dimensional Hilbert space. 535 536 Chapter 9 Differential Equations whereFisaknown(source)functionofone(forODEs)ormorevariables(forPDEs), Lis alinearcombinationofderivatives,and ψistheunknownfunctionorsolution.Anylinear combination of solutions is again a solution if F=0; this is the superposition principle forhomogeneousPDEs. Since the dynamics of many physical systems involve just two derivatives, for exam- ple, acceleration in classical mechanics and the kinetic energy operator, ∼∇2, in quan- tum mechanics, differential equations of second order occur most frequently in physics. (Maxwell’sandDirac’sequationsarefirstorderbutinvolvetwounknownfunctions.Elim- inating one unknown yields a second-order differential equation for the other (compare Section1.9).) Examples of PDEs AmongthemostfrequentlyencounteredPDEs arethefollowing: 1. Laplace’sequation, ∇2ψ=0. Thisverycommonandveryimportantequationoccursinstudiesof a. electromagneticphenomena,includingelectrostatics,dielectrics,steadycurrents, andmagnetostatics, b. hydrodynamics(irrotationalflowofperfectfluidandsurface waves), c. heatflow, d. gravitation. 2. Poisson’sequation, ∇2ψ=−ρ/ε0. IncontrasttothehomogeneousLaplaceequation,Poisson’sequationisnonhomo- geneouswithasourceterm −ρ/ε0. 3. Thewave(Helmholtz)andtime-independentdiffusionequations, ∇2ψ±k2ψ=0. Theseequationsappearinsuchdiversephenomenaas a. elasticwavesinsolids,includingvibratingstrings, bars, membranes, b. sound,oracoustics, c. electromagneticwaves, d. nuclearreactors. 4. Thetime-dependentdiffusionequation ∇2ψ=1 a2∂ψ ∂t and the corresponding four-dimensional forms involving the d’Alembertian, a four- dimensionalanalogof theLaplacianinMinkowskispace, ∂µ∂µ=∂2=1 c2∂2 ∂t2−∇2. 5. Thetime-dependentwaveequation, ∂2ψ=0. 6. Thescalarpotentialequation, ∂2ψ=ρ/ε0. Like Poisson’s equation, this equation is nonhomogeneous with a source term ρ/ε0. 9.1 Partial Differential Equations 537 7. TheKlein–Gordonequation, ∂2ψ=−µ2ψ,andthecorrespondingvectorequations, in which the scalar function ψis replaced by a vector function. Other, more compli- catedforms arecommon. 8. TheSchrödingerwaveequation, −¯h2 2m∇2ψ+Vψ=i¯h∂ψ ∂t and −¯h2 2m∇2ψ+Vψ=Eψ for thetime-independentcase. 9. Theequationsfor elasticwavesandviscousfluidsandthetelegraphyequation. 10. Maxwell’s coupled partial differential equations for electric and magnetic fields and those of Dirac for relativistic electron wave functions. For Maxwell’s equations see theIntroductionandalsoSection1.9. Somegeneraltechniquesfor solvingsecond-orderPDEsarediscussedinthissection. 1. Separation of variables, where the PDE is split into ODEs that are related by com- monconstantsthatappearaseigenvaluesoflinearoperators, Lψ=lψ,usuallyinone variable. This method is closely related to symmetries of the PDE and a group of transformations (seeSection4.2).TheHelmholtzequation,listedexample3,hasthis form, where the eigenvalue k2may arise by separation of the time tfrom the spatial variables. Likewise, in example 8 the energy Eis the eigenvalue that arises in the separation of tfromrin the Schrödinger equation. This is pursued in Chapter 10 in greaterdetail.Section9.2servesasintroduction.ODEsmaybeattackedbyFrobenius’ power-series method in Section 9.5. It does not always work but is often the simplest methodwhenitdoes. 2. Conversion of a PDE into an integral equation using Green’s functions applies to inhomogeneous PDEs, such as examples 2 and 6 given above. An introduction to the Green’sfunctiontechniqueis giveninSection9.7. 3. Other analytical methods, such as the use of integral transforms, are developed and appliedinChapter15. Occasionally, we encounter equations of higher order. In both the theory of the slow motionofaviscousfluidandthetheoryof anelasticbodywefindtheequation parenleftbig ∇2parenrightbig2ψ=0. Fortunately, these higher-order differential equations are relatively rare and are not dis- cussedhere. Although not so frequently encountered and perhaps not so important as second-order ODEs, first-order ODEs do appear in theoretical physics and are sometimes intermediate steps for second-order ODEs. The solutions of some more important types of first-order ODEs are developed in Section 9.2. First-order PDEs can always be reduced to ODEs. This is a straightforward but lengthy process and involves a search for characteristics that arebrieflyintroducedinwhatfollows;for moredetailswerefer totheliterature. 538 Chapter 9 Differential Equations Classes of PDEs and Characteristics Second-order PDEs form three classes: (i) Elliptic PDEs involve ∇2orc−2∂2/∂t2+∇2 (ii)parabolicPDEs, a∂/∂t+∇2;(iii)hyperbolicPDEs, c−2∂2/∂t2−∇2.Thesecanonical operatorscomeaboutbyachangeofvariables ξ=ξ(x,y),η=η(x,y)inalinearoperator (for twovariablesjustfor simplicity) L=a∂2 ∂x2+2b∂2 ∂x∂y+c∂2 ∂y2+d∂ ∂x+e∂ ∂y+f, (9.2) which can be reduced to the canonical forms (i), (ii), (iii) according to whether the dis- criminant D=ac−b2>0,=0, or<0. Ifξ(x,y)is determined from the first-order, but nonlinear,PDE aparenleftbigg∂ξ ∂xparenrightbigg2 +2bparenleftbigg∂ξ ∂xparenrightbiggparenleftbigg∂ξ ∂yparenrightbigg +cparenleftbigg∂ξ ∂yparenrightbigg2 =0, (9.3) thenthecoefficientof ∂2/∂ξ2inL(thatis,Eq.(9.3))iszero.If ηisanindependentsolution of the same Eq. (9.3), then the coefficient of ∂2/∂η2is also zero. The remaining operator, ∂2/∂ξ∂η,i nLis characteristic of the hyperbolic case (iii) with D<0(a=0=cleads to D=−b2<0),wherethequadraticform aλ2+2bλ+cfactorizesand,therefore,Eq.(9.3) has two independent solutions ξ(x,y),η(x,y). In the elliptic case (i) with D>0, the two solutions ξ,ηare complex conjugate, which, when substituted into Eq. (9.2), remove the mixed second-order derivative instead of the other second-order terms, yielding the canonical form (i). In the parabolic case (ii) with D=0, only∂2/∂ξ2remains in L, while thecoefficientsof theothertwosecond-orderderivativesvanish. If the coefficients a,b,cinLare functions of the coordinates, then this classificationis onlylocal;thatis, itstypemaychangeasthecoordinatesvary. Let us illustrate the physics underlying the hyperbolic case by looking at the wave equation,Eq.(9.2) (in 1 +1 dimensionsfor simplicity) parenleftbigg1 c2∂2 ∂t2−∂2 ∂x2parenrightbigg ψ=0. SinceEq. (9.3)nowbecomes parenleftbigg∂ξ ∂tparenrightbigg2 −c2parenleftbigg∂ξ ∂xparenrightbigg2 =parenleftbigg∂ξ ∂t−c∂ξ ∂xparenrightbiggparenleftbigg∂ξ ∂t+c∂ξ ∂xparenrightbigg =0 and factorizes, we determine the solution of ∂ξ/∂t−c∂ξ/∂x=0. This is an arbitrary functionξ=F(x+ct), andξ=G(x−ct)solves∂ξ/∂t+c∂ξ/∂x=0, which is readily verified.Bylinearsuperpositionageneralsolutionofthewaveequationis ψ=F(x+ct)+ G(x−ct). For periodic functions F,Gwe recognize the lines x+ctandx−ctas the phases of plane waves or wave fronts, where not all second-order derivatives of ψin the waveequationarewelldefined.Normaltothewavefrontsaretheraysofgeometricoptics. Thus, the lines that are solutions of Eq. (9.3) and are called characteristics or sometimes bicharacteristics (forsecond-orderPDEs)inthemathematicalliteraturecorrespondtothe wavefronts ofthegeometricopticssolutionof thewaveequation. 9.1 Partial Differential Equations 539 For theellipticcase letusconsiderLaplace’sequation, ∂2ψ ∂x2+∂2ψ ∂y2=0, fora potential ψoftwovariables.Herethecharacteristicsequation, parenleftbigg∂ξ ∂xparenrightbigg2 +parenleftbigg∂ξ ∂yparenrightbigg2 =parenleftbigg∂ξ ∂x+i∂ξ ∂yparenrightbiggparenleftbigg∂ξ ∂x−i∂ξ ∂yparenrightbigg =0, has complex conjugate solutions: ξ=F(x+iy)for∂ξ/∂x+i(∂ξ/∂y)=0 andξ= G(x−iy)for∂ξ/∂x−i(∂ξ/∂y)=0.AgeneralsolutionofLaplace’sequationistherefore ψ=F(x+iy)+G(x−iy),aswellastherealandimaginarypartsof ψ,whicharecalled harmonic functions,whilepolynomialsolutionsarecalled harmonicpolynomials . In quantum mechanics the Wentzel–Kramers–Brillouin (WKB) form ψ=exp(−iS/¯h) forthesolutionof theSchrödingerequation,acomplexparabolicPDE, parenleftbigg −¯h2 2m∇2+Vparenrightbigg ψ=i¯h∂ψ ∂t, leadstotheHamilton–Jacobiequationofclassicalmechanics, 1 2m(∇S)2+V=∂S ∂t, (9.4) inthelimit¯h→0.Theclassicalaction SobeystheHamilton–Jacobiequation,whichisthe analog of Eq. (9.3) of the Schrödinger equation. Substituting ∇ψ=−iψ∇S/¯h,∂ψ/∂t= −iψ(∂S/∂t)/¯hintotheSchrödingerequation,droppingtheoverallnonvanishingfactor ψ, andtakingthelimitof theresultingequationas ¯h→0,weindeedobtainEq.(9.4). Finding solutions of PDEs by solving for the characteristics is one of several general techniques. For more examples we refer to H. Bateman, Partial Differential Equations of Mathematical Physics , New York: Dover (1944); K. E. Gustafson, Partial Differential Equations and Hilbert Space Methods , 2nd ed., New York: Wiley (1987), reprinted Dover (1998). In order to derive and appreciate more the mathematical method behind these solutions of hyperbolic, parabolic, and elliptic PDEs let us reconsider the PDE (9.2) with constant coefficients and, at first, d=e=f=0 for simplicity. In accordance with the form of the wavefrontsolutions,weseekasolution ψ=F(ξ)ofEq.(9.2)withafunction ξ=ξ(t,x) usingthevariables t,xinsteadof x,y.Thenthepartialderivativesbecome ∂ψ ∂x=∂ξ ∂xdF dξ,∂ψ ∂t=∂ξ ∂tdF dξ,∂2ψ ∂x2=∂2ξ ∂x2dF dξ+parenleftbigg∂ξ ∂xparenrightbigg2d2F dξ2, and ∂2ψ ∂x∂t=∂2ξ ∂x∂tdF dξ+∂ξ ∂x∂ξ ∂td2F dξ2,∂2ψ ∂t2=∂2ξ ∂t2dF dξ+parenleftbigg∂ξ ∂tparenrightbigg2d2F dξ2, using the chain rule of differentiation. When ξdepends on xandtlinearly, these partial derivatives of ψyield a single term only and solve our PDE (9.2) as a consequence. From thelinear ξ=αx+βtweobtain ∂2ψ ∂x2=α2d2F dξ2,∂2ψ ∂x∂t=αβd2F dξ2,∂2ψ ∂t2=β2d2F dξ2, 540 Chapter 9 Differential Equations andourPDE(9.2) becomesequivalenttotheanalogof Eq.(9.3), parenleftbig α2a+2αβb+β2cparenrightbigd2F dξ2=0. (9.5) A solution ofd2F dξ2=0 only leads to the trivial ψ=k1x+k2t+k3with constant kithat is linear in the coordinates and for which all second derivatives vanish. From α2a+2αβb+ β2c=0,ontheotherhand,wegettheratios β α=1 cbracketleftbig −b±parenleftbig b2−acparenrightbig1/2bracketrightbig ≡r1,2 (9.6) as solutions of Eq. (9.5) withd2F dξ2/negationslash=0 in general. The lines ξ1=x+r1tandξ2=x+r2t willsolvethePDE(9.2),with ψ(x,t)=F(ξ1)+G(ξ2)correspondingtothegeneralization ofourprevioushyperbolicandellipticPDEexamples. For the parabolic case, where b2=ac, there is only one ratio from Eq. (9.6), β/α= r=−b/c, and one solution, ψ(x,t)=F(x−bt/c). In order to find the second gen- eral solution of our PDE (9.2) we make the Ansatz (trial solution) ψ(x,t)=ψ0(x,t)· G(x−bt/c).SubstitutingthisintoEq. (9.2) wefind a∂2ψ0 ∂x2+2b∂2ψ0 ∂x∂t+c∂2ψ0 ∂t2=0 forψ0since, upon replacing F→G,Gsolves Eq. (9.5) with d2G/dξ2/negationslash=0 in general. The solution ψ0can be any solution of our PDE (9.2), including the trivial ones such as ψ0=xandψ0=t. Thusweobtainthe generalparabolicsolution , ψ(x,t)=Fparenleftbigg x−b ctparenrightbigg +ψ0(x,t)Gparenleftbigg x−b ctparenrightbigg , withψ0=xorψ0=t,et c. With the same Ansatz one finds solutions of our PDE (9.2) with a source term, for example, f/negationslash=0,butstill d=e=0 andconstant a,b,c. Nextwedeterminethecharacteristics,thatis,curveswherethesecondorderderivatives ofthesolution ψarenotwelldefined.Thesearethewavefrontsalongwhichthesolutions of our hyperbolic PDE (9.2) propagate. We solve our PDE with a source term f/negationslash=0 and Cauchy boundary conditions (see Table 9.1) that are appropriate for hyperbolic PDEs, whereψanditsnormalderivative ∂ψ/∂narespecifiedonanopencurve C:x=x(s), t=t(s), with the parameter sthe length on C. Thendr=(dx,dt)is tangent and ˆnds=(dt,−dx) is normal to the curve C, and the first-order tangential and normal derivatives are given by thechainrule dψ ds=∇ψ·dr ds=∂ψ ∂xdx ds+∂ψ ∂tdt ds, dψ dn=∇ψ·ˆn=∂ψ ∂xdt ds−∂ψ ∂tdx ds. 9.1 Partial Differential Equations 541 Fromthesetwolinearequations, ∂ψ/∂tand∂ψ/∂xcanbedeterminedon C, provided vextendsinglevextendsinglevextendsinglevextendsinglevextendsingledx dsdt ds dt ds−dx dsvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−parenleftbiggdx dsparenrightbigg2 −parenleftbiggdt dsparenrightbigg2 /negationslash=0. Forthesecondderivativesweusethechainruleagain: d ds∂ψ ∂x=dx ds∂2ψ ∂x2+dt ds∂2ψ ∂x∂t, (9.7a) d ds∂ψ ∂t=dx ds∂2ψ ∂x∂t+dt ds∂2ψ ∂t2. (9.7b) From our PDE (9.2), and Eqs. (9.7a,b), which are linear in the second-order derivatives, theycannotbecalculatedwhenthedeterminantvanishes,thatis, vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglea2bc dx dsdt ds0 0dx dsdt dsvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=aparenleftbiggdt dsparenrightbigg2 −2bdx dsdt ds+cparenleftbiggdx dsparenrightbigg2 =0. (9.8) FromEq.(9.8),whichdefinesthecharacteristics,wefindthatthetangentratio dx/dtobeys cparenleftbiggdx dtparenrightbigg2 −2bdx dt+a=0, so dx dt=1 cbracketleftbig b±parenleftbig b2−acparenrightbig1/2bracketrightbig . (9.9) For the earlier hyperbolic wave (and elliptic potential )equation examples, b=0anda,c areconstants,sothesolutions ξi=x+trifromEq.(9.6)coincidewiththecharacteristics ofEq.(9.9). Nonlinear PDEs Nonlinear ODEs and PDEs are a rapidly growing and important field. We encountered earlierthesimplestlinearwaveequation, ∂ψ ∂t+c∂ψ ∂x=0, asthefirst-orderPDEofthewavefrontsofthewaveequation.Thesimplestnonlinearwave equation, ∂ψ ∂t+c(ψ)∂ψ ∂x=0, (9.10) results if the local speed of propagation, c, is not constant but depends on the wave ψ. When a nonlinear equation has a solution of the form ψ(x,t)=Acos(kx−ωt), where 542 Chapter 9 Differential Equations ω(k)varies with kso thatω′′(k)/negationslash=0, then it is called dispersive . Perhaps the best-known nonlineardispersiveequationisthe Korteweg–deVries equation, ∂ψ ∂t+ψ∂ψ ∂x+∂3ψ ∂x3=0, (9.11) which models the lossless propagation of shallow water waves and other phenomena. It is widely known for its solitonsolutions. A soliton is a traveling wave with the property of persisting through an interaction with another soliton: After they pass through each other, they emerge in the same shape and with the same velocity and acquire no more than a phase shift. Let ψ(ξ=x−ct)be such a traveling wave. When substituted into Eq. (9.11) thisyieldsthenonlinearODE (ψ−c)dψ dξ+d3ψ dξ3=0, (9.12) whichcanbeintegratedtoyield d2ψ dξ2=cψ−ψ2 2. (9.13) There is no additive integration constant in Eq. (9.13) to ensure that d2ψ/dξ2→0 with ψ→0f o rl a r g e ξ,s oψis localized at the characteristic ξ=0, orx=ct. Multiplying Eq.(9.13) by dψ/dξandintegratingagainyields parenleftbiggdψ dξparenrightbigg2 =cψ2−ψ3 3, (9.14) wheredψ/dξ→0f o rl a r g e ξ. Taking the root of Eq. (9.14) and integrating once more yieldsthesolitonsolution ψ(x−ct)=3c cosh2parenleftbig√cx−ct 2parenrightbig. (9.15) Some nonlinear topics, for example, the logistic equation and the onset of chaos, are re- viewedinChapter18.Formoredetailsandliterature,seeJ.Guckenheimer,P.Holmes,and F.John,NonlinearOscillations,DynamicalSystemsandBifurcationsofVectorFields ,re v . ed.,NewYork:Springer-Verlag(1990). Boundary Conditions Usually,whenweknowaphysicalsystematsometimeandthelawgoverningthephysical process, then we are able to predict the subsequent development. Such initial values are themostcommonboundaryconditionsassociatedwithODEsandPDE.Findingsolutions that match given points, curves, or surfaces corresponds to boundary value problems. So- lutionsusuallyarerequiredtosatisfycertainimposed(forexample,asymptotic)boundary conditions.Theseboundaryconditionsmaytakethreeforms: 1. Cauchyboundaryconditions .Thevalueofafunctionandnormalderivativespecified ontheboundary.Inelectrostaticsthiswouldmean ϕ,thepotential,and En,thenormal componentof theelectricfield. 9.2 First-Order Differential Equations 543 Table 9.1 Boundary Typeof partial differential equation conditions Elliptic Hyperbolic Parabolic Laplace,Poisson Waveequation in Diffusion equation in(x,y) (x,t) in(x,t) Cauchy Opensurface Unphysical results Unique,stable Toorestrictive (instability) solution Closedsurface Toorestrictive Toorestrictive Toorestrictive Dirichlet Opensurface Insufficient Insufficient Unique, stable solutionin one direction Closedsurface Unique,stable Solution not unique Too restrictive solution Neumann Opensurface Insufficient Insufficient Unique, stable solutionin one direction Closedsurface Unique,stable Solution not unique Too restrictive solution 2. Dirichletboundaryconditions .Thevalueof afunctionspecifiedontheboundary. 3. Neumann boundary conditions . The normal derivative (normal gradient) of a func- tion specified on the boundary. In the electrostatic case this would be Enand there- foreσ, thesurface chargedensity. Asummaryoftherelationofthesethreetypesofboundaryconditionstothethreetypes of two-dimensional partial differential equations is given in Table 9.1. For extended dis- cussionsofthesepartialdifferentialequationsthereadermayconsultMorseandFeshbach, Chapter6(see AdditionalReadings). PartsofTable9.1aresimplyamatterofmaintaininginternalconsistencyorofcommon sense.Forinstance,for Poisson’sequationwithaclosedsurface, Dirichletconditionslead to a unique, stable solution. Neumannconditions,independentof the Dirichletconditions, likewise lead to a unique stable solution independent of the Dirichlet solution. Therefore Cauchyboundaryconditions(meaningDirichletplusNeumann)couldleadtoaninconsis- tency. The term boundary conditions includes as a special case the concept of initial con- ditions. For instance, specifying the initial position x0and the initial velocity v0in some dynamicalproblemwouldcorrespondtotheCauchyboundaryconditions.Theonlydiffer- enceinthepresentusageofboundaryconditionsintheseone-dimensionalproblemsisthat weare goingtoapplytheconditionson bothendsoftheallowedrangeofthevariable. 9.2 F IRST -ORDER DIFFERENTIAL EQUATIONS Physics involves some first-order differential equations. For completeness (and review) it seems desirable to touch on them briefly. We consider here differential equations of the 544 Chapter 9 Differential Equations generalform dy dx=f(x,y)=−P(x,y) Q(x,y). (9.16) Equation (9.16) is clearly a first-order, ordinary differential equation. It is first order be- causeitcontainsthefirstandnohigherderivatives.Itis ordinary becausetheonlyderiva- tive,dy/dx, is an ordinary, or total, derivative. Equation (9.16) may or may not be linear, althoughweshalltreatthelinearcaseexplicitlylater,Eq. (9.25). Separable Variables FrequentlyEq. (9.16) willhavethespecialform dy dx=f(x,y)=−P(x) Q(y). (9.17) Thenit mayberewrittenas P(x)dx+Q(y)dy=0. Integratingfrom (x0,y0)to(x,y)yields integraldisplayx x0P(x)dx+integraldisplayy y0Q(y)dy=0. Since the lower limits, x0andy0, contribute constants, we may ignore the lower limits of integration and simply add a constant of integration. Note that this separation of variables techniquedoes notrequirethatthedifferentialequationbelinear. Example 9.2.1 PARACHUTIST We want to find the velocity of the falling parachutist as a function of time and are partic- ularly interested in the constant limiting velocity, v0, that comes about by air drag, taken, to be quadratic, −bv2, and opposing the force of the gravitational attraction, mg,o ft h e Earth. We choose a coordinate system in which the positive direction is downward so that the gravitational force is positive. For simplicity we assume that the parachute opens im- mediately, that is, at time t=0, where v(t=0)=0, our initial condition. Newton’s law appliedtothefallingparachutistgives m˙v=mg−bv2, wheremincludesthemassoftheparachute. The terminal velocity, v0, can be found from the equation of motion as t→∞; when thereis noacceleration, ˙v=0,so bv2 0=mg, orv0=radicalbiggmg b. 9.2 First-Order Differential Equations 545 Thevariables tandvseparate dv g−b mv2=dt, whichweintegratebydecomposingthedenominatorintopartialfractions.Therootsofthe denominatorareat v=±v0.Hence parenleftbigg g−b mv2parenrightbigg−1 =m 2v0bparenleftbigg1 v+v0−1 v−v0parenrightbigg . Integratingbothtermsyields integraldisplayvdV g−b mV2=1 2radicalbiggm gblnv0+v v0−v=t. Solvingfor thevelocityyields v=e2t/T−1 e2t/T+1v0=v0sinht T cosht T=v0tanht T, whereT=radicalBig m gbis the time constant governing the asymptotic approach of the velocity to thelimitingvelocity, v0. Putting in numerical values, g=9.8m/s2and taking b=700 kg/m,m=70 kg, gives v0=√9.8/10∼1m/s∼3.6k m/h∼2.23 mi/h, the walking speed of a pedestrian at landing, and T=radicalBig m bg=1/√ 10·9.8∼0.1 s. Thus, the constant speed v0is reached withinasecond.Finally,because it isalwaysimportantto checkthesolution ,wev erify thatoursolutionsatisfies ˙v=cosht/T cosht/Tv0 T−sinh2t/T cosh2t/Tv0 T=v0 T−v2 Tv0=g−b mv2, that is, Newton’s equation of motion. The more realistic case, where the parachutist is in free fall with an initial speed vi=v(0)>0 before the parachute opens, is addressed in Exercise9.2.18. /squaresolid Exact Differential Equations WerewriteEq. (9.16)as P(x,y)dx+Q(x,y)dy=0. (9.18) Thisequationissaidtobe exactifwecanmatchtheleft-handsideofittoadifferential dϕ, dϕ=∂ϕ ∂xdx+∂ϕ ∂ydy. (9.19) Since Eq. (9.18) has a zero on the right, we look for an unknown function ϕ(x,y)= constant and dϕ=0. 546 Chapter 9 Differential Equations We have(if suchafunction ϕ(x,y)exists) P(x,y)dx+Q(x,y)dy=∂ϕ ∂xdx+∂ϕ ∂ydy (9.20a) and ∂ϕ ∂x=P(x,y),∂ϕ ∂y=Q(x,y). (9.20b) The necessary and sufficient condition for our equation to be exact is that the second, mixed partial derivatives of ϕ(x,y)(assumed continuous) are independent of the order of differentiation: ∂2ϕ ∂y∂x=∂P(x,y) ∂y=∂Q(x,y) ∂x=∂2ϕ ∂x∂y. (9.21) Note the resemblance to Eqs. (1.133a) of Section 1.13, “Potential Theory.” If Eq. (9.18) correspondstoa curl(equaltozero), thenapotential, ϕ(x,y),mustexist. Ifϕ(x,y)exists,thenfromEqs. (9.18) and(9.20a) oursolutionis ϕ(x,y)=C. We may construct ϕ(x,y)from its partial derivatives just as we constructed a magnetic vectorpotentialinSection1.13fromits curl.SeeExercises9.2.7and9.2.8. It may well turn out that Eq. (9.18) is not exact and that Eq. (9.21) is not satisfied. However, there always exists at least one and perhaps an infinity of integrating factors α(x,y)suchthat α(x,y)P(x,y)dx +α(x,y)Q(x,y)dy =0 is exact. Unfortunately, an integrating factor is not always obvious or easy to find. Unlike the case of the linear first-order differential equation to be considered next, there is no systematicwaytodevelopanintegratingfactor for Eq.(9.18). Adifferentialequationinwhichthevariableshavebeenseparatedisautomaticallyexact. Anexactdifferentialequationis notnecessarilyseparable. ThewavefrontmethodofSection9.1alsoworks forafirst-order PDE: a(x,y)∂ψ ∂x+b(x,y)∂ψ ∂y=0. (9.22a) We look for a solution of the form ψ=F(ξ), whereξ(x,y)=constant for varying xand ydefinesthewavefront. Hence dξ=∂ξ ∂xdx+∂ξ ∂ydy=0, (9.22b) whilethePDEyields parenleftbigg a∂ξ ∂x+b∂ξ ∂yparenrightbiggdF dξ=0 (9.23a) withdF/dξ/negationslash=0 ingeneral.ComparingEqs. (9.22b)and(9.23a)yields dx a=dy b, (9.23b) 9.2 First-Order Differential Equations 547 which reduces the PDE to a first-order ODE for the tangent dy/dxof the wave front functionξ(x,y). Whenthereis anadditionalsourceterminthePDE, a∂ψ ∂x+b∂ψ ∂y+cψ=0, (9.23c) thenweusetheAnsatz ψ=ψ0(x,y)F(ξ) , whichconvertsourPDEto Fparenleftbigg a∂ψ0 ∂x+b∂ψ0 ∂y+cψ0parenrightbigg +ψ0dF dξparenleftbigg a∂ξ ∂x+b∂ξ ∂yparenrightbigg =0. (9.24) If we can guess a solution ψ0of Eq. (9.23c), then Eq. (9.24) reduces to our previous equation,Eq.(9.23a), fromwhichtheODEofEq. (9.23b)follows. Linear First-Order ODEs Iff(x,y)inEq.(9.16) hastheform −p(x)y+q(x),thenEq.(9.16) becomes dy dx+p(x)y=q(x). (9.25) Equation (9.25) is the most general linearfirst-order ODE. If q(x)=0, Eq. (9.25) is homogeneous (in y). A nonzero q(x)may represent a sourceor adriving term . Equa- tion (9.25) is linear; each term is linear in yordy/dx. There are no higher powers, that is,y2, and no products, y(dy/dx) . Note that the linearity refers to the yanddy/dx;p(x) andq(x)need not be linear in x. Equation (9.25), the most important of these first-order ODEsforphysics, maybesolvedexactly. L etusl ookf oran integratingfactor α(x)sothat α(x)dy dx+α(x)p(x)y=α(x)q(x) (9.26) mayberewrittenas d dxbracketleftbig α(x)ybracketrightbig =α(x)q(x). (9.27) The purpose of this is to make the left-hand side of Eq. (9.25) a derivative so that it can be integrated—by inspection. It also, incidentally, makes Eq. (9.25) exact. Expanding Eq.(9.27), weobtain α(x)dy dx+dα dxy=α(x)q(x). ComparisonwithEq. (9.26) showsthatwemustrequire dα dx=α(x)p(x). (9.28) Hereisadifferentialequationfor α(x),withthevariables αandxseparable .Weseparate variables,integrate,andobtain α(x)=expbracketleftbiggintegraldisplayx p(x)dxbracketrightbigg (9.29) 548 Chapter 9 Differential Equations asourintegratingfactor. Withα(x)known we proceed to integrate Eq. (9.27). This, of course, was the point of introducing αinthefirst place.Wehave integraldisplayxd dxbracketleftbig α(x)y(x)bracketrightbig dx=integraldisplayx α(x)q(x)dx. Nowintegratingbyinspection,wehave α(x)y(x)=integraldisplayx α(x)q(x)dx+C. The constants from a constant lower limit of integration are lumped into the constant C. Dividingby α(x), weobtain y(x)=bracketleftbig α(x)bracketrightbig−1braceleftbiggintegraldisplayx α(x)q(x)dx+Cbracerightbigg . Finally,substitutinginEq. (9.29)for αyields y(x)=expbracketleftbigg −integraldisplayx p(t)dtbracketrightbiggbraceleftbiggintegraldisplayx expbracketleftbiggintegraldisplays p(t)dtbracketrightbigg q(s)ds+Cbracerightbigg .(9.30) Here the (dummy) variables of integration have been rewritten to make them unambigu- ous. Equation (9.30) is the complete general solution of the linear, first-order differential equation,Eq.(9.25). Theportion y1(x)=Cexpbracketleftbigg −integraldisplayx p(t)dtbracketrightbigg (9.31) correspondstothecase q(x)=0andisageneralsolutionofthehomogeneousdifferential equation.TheotherterminEq. (9.30), y2(x)=expbracketleftbigg −integraldisplayx p(t)dtbracketrightbiggintegraldisplayx expbracketleftbiggintegraldisplays p(t)dtbracketrightbigg q(s)ds, (9.32) isaparticularsolutioncorrespondingto thespecificsourceterm q(x). Note that if our linear first-order differential equation is homogeneous (q=0), then it is separable. Otherwise, apart from special cases such as p=constant, q=constant, and q(x)=ap(x), Eq. (9.25)is notseparable. LetussummarizethissolutionoftheinhomogeneousODEintermsofa methodcalled variationoftheconstant asfollows.Inthefirststep,wesolvethehomogeneousODEby separationofvariablesasbefore, giving y′ y=−p,lny=−integraldisplayx p(X)dX+lnC, y(x)=Ce−integraltextxp(X)dX. Inthesecondstep,welettheintegrationconstantbecome x-dependent,thatis, C→C(x). This is the “variation of the constant” used to solve the inhomogeneous ODE. Differenti- atingy(x)weobtain y′=−pCe−integraltext p(x)dx+C′(x)e−integraltext p(x)dx=−py(x)+C′(x)e−integraltext p(x)dx. 9.2 First-Order Differential Equations 549 ComparingwiththeinhomogeneousODEwefindtheODEfor C: C′e−integraltext p(x)dx=q,orC(x)=integraldisplayx eintegraltextXp(Y)dYq(X)dX. Substitutingthis Cintoy=C(x)e−integraltextxp(X)dXreproducesEq. (9.32). Example 9.2.2 RL C IRCUIT Foraresistance-inductancecircuitKirchhoff’slawleadsto LdI(t) dt+RI(t)=V(t) forthecurrent I(t),whereListheinductanceand Ristheresistance,bothconstant. V(t) isthetime-dependentinputvoltage. FromEq. (9.29)ourintegratingfactor α(t)is α(t)=expintegraldisplaytR Ldt=eRt/L. ThenbyEq. (9.30), I(t)=e−Rt/Lbracketleftbiggintegraldisplayt eRt/LV(t) Ldt+Cbracketrightbigg , withtheconstant Ctobedeterminedbyaninitialcondition(aboundarycondition). Forthespecialcase V(t)=V0, aconstant, I(t)=e−Rt/LbracketleftbiggV0 L·L ReRt/L+Cbracketrightbigg =V0 R+Ce−Rt/L. If theinitialconditionis I(0)=0,thenC=−V0/Rand I(t)=V0 Rbracketleftbig 1−e−Rt/Lbracketrightbig . /squaresolid Nowwe provethe theoremthatthesolutionoftheinhomogeneousODEisuniqueup toanarbitrarymultipleof thesolutionof thehomogeneousODE . Toshowthis, suppose y1,y2bothsolvetheinhomogeneousODE,Eq. (9.25);then y′ 1−y′ 2+p(x)(y1−y2)=0 follows by subtracting the ODEs and says that y1−y2is a solution of the homogeneous ODE. The solution of the homogeneous ODE can always be multiplied by an arbitrary constant. Wealsoprovethe theoremthatafirst-orderhomogeneousODEhasonlyonelinearly independent solution . This is meant in the following sense. If two solutions are linearly dependent ,bydefinitiontheysatisfy ay1(x)+by2(x)=0withnonzeroconstants a, bfor all values of x.If the only solution of this linear relation is a=0=b, then our solutions y1andy2ar es aidtobe linearlyindependent . 550 Chapter 9 Differential Equations Toprovethistheorem,suppose y1,y2bothsolvethehomogeneousODE.Then y′ 1 y1=−p(x)=y′ 2 y2implies W(x)≡y′ 1y2−y1y′ 2≡0.(9.33) The functional determinant Wis called the Wronskian of the pair y1,y2.We now show thatW≡0istheconditionforthemtobelinearlydependent.Assuminglineardependence, thatis, ay1(x)+by2(x)=0 with nonzero constants a, bfor all values of x, we differentiate this linear relation to get anotherlinearrelation, ay′ 1(x)+by′ 2(x)=0. The conditionfor thesetwo homogeneouslinear equationsinthe unknowns a,bto havea nontrivialsolutionisthattheirdeterminantbezero,whichis W=0. Conversely, from W=0, there follows linear dependence, because we can find a non- trivialsolutionoftherelation y′ 1 y1=y′ 2 y2 byintegration,whichgives lny1=lny2+lnC,ory1=Cy2. Linear dependence and the Wronskian are generalized to three or more functions in Sec- tion9.6. Exercises 9.2.1 FromKirchhoff’slawthecurrent IinanRC(resistance–capacitance)circuit(Fig.9.1) obeystheequation RdI dt+1 CI=0. (a) Find I(t). (b) For a capacitanceof 10,000 µF charged to 100 V and discharging through a resis- tanceof1M /Omega1,findthecurrent Ifort=0 andfor t=100 seconds. Note.Theinitialvoltageis I0RorQ/C, whereQ=integraltext∞ 0I(t)dt. 9.2.2 TheLaplacetransformof Bessel’sequation (n=0)leadsto parenleftbig s2+1parenrightbig f′(s)+sf(s)=0. Solvefor f(s). 9.2 First-Order Differential Equations 551 FIGURE 9.1RCcircuit. 9.2.3 Thedecayofa populationbycatastrophictwo-bodycollisionsisdescribedby dN dt=−kN2. Thisis afirst-order, nonlinear differentialequation.Derivethesolution N(t)=N0parenleftbigg 1+t τ0parenrightbigg−1 , whereτ0=(kN0)−1. Thisimpliesaninfinitepopulationat t=−τ0. 9.2.4 The rate of a particular chemicalreaction A+B→Cis proportionalto the concentra- tionsofthereactants AandB: dC(t) dt=αbracketleftbig A(0)−C(t)bracketrightbigbracketleftbig B(0)−C(t)bracketrightbig . (a) Find C(t)forA(0)/negationslash=B(0). (b) Find C(t)forA(0)=B(0). Theinitialconditionis that C(0)=0. 9.2.5 A boat, coasting through the water, experiences a resisting force proportional to vn,v beingtheboat’sinstantaneousvelocity.Newton’ssecondlawleadsto mdv dt=−kvn. Withv(t=0)=v0,x(t=0)=0, integrate to find vas a function of time and vas a functionofdistance. 9.2.6 In the first-order differential equation dy/dx=f(x,y)the function f(x,y)is a func- tionof theratio y/x: dy dx=g(y/x). Showthatthesubstitutionof u=y/xleadstoaseparableequationin uandx. 552 Chapter 9 Differential Equations 9.2.7 Thedifferentialequation P(x,y)dx+Q(x,y)dy=0 isexact.Constructasolution ϕ(x,y)=integraldisplayx x0P(x,y)dx+integraldisplayy y0Q(x0,y)dy=constant. 9.2.8 Thedifferentialequation P(x,y)dx+Q(x,y)dy=0 isexact.If ϕ(x,y)=integraldisplayx x0P(x,y)dx+integraldisplayy y0Q(x0,y)dy, showthat ∂ϕ ∂x=P(x,y),∂ϕ ∂y=Q(x,y). Henceϕ(x,y)=constant isa solutionoftheoriginaldifferentialequation. 9.2.9 Prove that Eq. (9.26) is exact in the sense of Eq. (9.21), provided that α(x)satisfies Eq. (9.28). 9.2.10 Acertaindifferentialequationhas theform f(x)dx+g(x)h(y)dy=0, withnoneofthefunctions f(x),g(x),h(y)identicallyzero.Showthatanecessaryand sufficientconditionfor thisequationtobeexactisthat g(x)=constant. 9.2.11 Showthat y(x)=expbracketleftbigg −integraldisplayx p(t)dtbracketrightbiggbraceleftbiggintegraldisplayx expbracketleftbiggintegraldisplays p(t)dtbracketrightbigg q(s)ds+Cbracerightbigg isa solutionof dy dx+p(x)y(x)=q(x) bydifferentiatingtheexpressionfor y(x)andsubstitutingintothedifferentialequation. 9.2.12 Themotionofabodyfallinginaresistingmediummaybedescribedby mdv dt=mg−bv when the retarding force is proportional to the velocity, v. Find the velocity. Evaluate theconstantofintegrationbydemandingthat v(0)=0. 9.2 First-Order Differential Equations 553 9.2.13 Radioactivenucleidecayaccordingtothelaw dN dt=−λN, Nbeing the concentration of a given nuclide and λ, the particular decay constant. In a radioactiveseries of ndifferentnuclides,startingwith N1, dN1 dt=−λ1N1, dN2 dt=λ1N1−λ2N2,andsoon. FindN2(t)for theconditions N1(0)=N0andN2(0)=0. 9.2.14 The rate of evaporation from a particular spherical drop of liquid (constant density) is proportional to its surface area. Assuming this to be the sole mechanism of mass loss, findtheradiusof thedropas afunctionoftime. 9.2.15 Inthelinearhomogeneousdifferentialequation dv dt=−av the variables are separable. When the variables are separated, the equation is exact. Solvethisdifferentialequationsubjectto v(0)=v0bythefollowingthreemethods: (a) Separatingvariablesandintegrating. (b) Treatingtheseparatedvariableequationasexact. (c) Usingtheresultfor alinearhomogeneousdifferentialequation. ANS.v(t)=v0e−at. 9.2.16 Bernoulli’sequation, dy dx+f(x)y=g(x)yn, is nonlinear for n/negationslash=0 or 1. Show that the substitution u=y1−nreduces Bernoulli’s equationtoalinearequation.(SeeSection18.4.) ANS.du dx+(1−n)f(x)u=(1−n)g(x). 9.2.17 Solve the linear, first-order equation, Eq. (9.25), by assuming y(x)=u(x)v(x), where v(x)is a solution of the corresponding homogeneous equation [q(x)=0].T h i si st h e methodof variationofparameters duetoLagrange.Weapplyittosecond-orderequa- tionsinExercise9.6.25. 9.2.18 (a)SolveExample9.2.1foraninitialvelocity vi=60mi/h,whentheparachuteopens. Findv(t).(b) For a skydiver in free fall use the friction coefficient b=0.25 kg/m and massm=70 kg.Whatis thelimitingvelocityinthis case? 554 Chapter 9 Differential Equations 9.3 S EPARATION OF VARIABLES TheequationsofmathematicalphysicslistedinSection9.1areallpartialdifferentialequa- tions. Our first technique for their solution splits the partial differential equation of nvari- ables into nordinary differential equations. Each separation introduces an arbitrary con- stantofseparation.Ifwehave nvariables,wehavetointroduce n−1constants,determined bytheconditionsimposedintheproblembeingsolved. Cartesian Coordinates InCartesiancoordinatestheHelmholtzequationbecomes ∂2ψ ∂x2+∂2ψ ∂y2+∂2ψ ∂z2+k2ψ=0, (9.34) usingEq.(2.27)fortheLaplacian.Forthepresentlet k2beaconstant.Perhapsthesimplest way of treating a partial differential equation such as Eq. (9.34) is to split it into a set of ordinarydifferentialequations.This maybedoneas follows.Let ψ(x,y,z)=X(x)Y(y)Z(z) (9.35) andsubstitutebackintoEq.(9.34). HowdoweknowEq.(9.35) isvalid?Whenthediffer- entialoperatorsinvariousvariablesareadditiveinthePDE,thatis,whentherearenoprod- ucts of differential operators in different variables, the separation method usually works. Weareproceedinginthespiritoflet’stryandseeifitworks.Ifourattemptsucceeds,then Eq. (9.35) will be justified. If it does not succeed, we shall find out soon enough and then we shall try another attack, such as Green’s functions, integral transforms, or brute-force numericalanalysis.With ψassumedgivenbyEq. (9.35), Eq.(9.34) becomes YZd2X dx2+XZd2Y dy2+XYd2Z dz2+k2XYZ=0. (9.36) Dividingby ψ=XYZandrearrangingterms,weobtain 1 Xd2X dx2=−k2−1 Yd2Y dy2−1 Zd2Z dz2. (9.37) Equation (9.37) exhibits one separation of variables. The left-hand side is a function of x alone, whereas the right-hand side depends only on yandzand not on x.B u tx,y, andz areallindependentcoordinates.Theequalityofbothsidesdependingondifferentvariables means that the behavior of xas an independent variable is not determined by yandz. Therefore,eachsidemustbeequaltoa constant,aconstantof separation.We choose2 1 Xd2X dx2=−l2, (9.38) −k2−1 Yd2Y dy2−1 Zd2Z dz2=−l2. (9.39) 2The choice of sign, completely arbitrary here, will be fixed in specific problems by the need to satisfy specific boundary conditions. 9.3 Separation of Variables 555 Now,turningourattentiontoEq.(9.39), weobtain 1 Yd2Y dy2=−k2+l2−1 Zd2Z dz2, (9.40) and a second separation has been achieved. Here we have a function of yequated to a functionof z,asbefore.Weresolveit,asbefore,byequatingeachsidetoanotherconstant ofseparation ,2−m2, 1 Yd2Y dy2=−m2, (9.41) 1 Zd2Z dz2=−k2+l2+m2=−n2, (9.42) introducing a constant n2byk2=l2+m2+n2to produce a symmetric set of equations. NowwehavethreeODEs((9.38),(9.41),and(9.42))toreplaceEq.(9.34).Ourassumption (Eq. (9.35)) hassucceededandistherebyjustified. Oursolutionshouldbelabeledaccordingtothechoiceofourconstants l,m,andn;that is, ψlm(x,y,z)=Xl(x)Ym(y)Zn(z). (9.43) Subject to the conditions of the problem being solved and to the condition k2=l2+ m2+n2, we may choose l,m, andnas we like, and Eq. (9.43) will still be a solution of Eq.(9.34),provided Xl(x)isasolutionofEq.(9.38),andsoon.Wemaydevelop themost generalsolution ofEq. (9.34) bytakinga linearcombinationof solutions ψlm, /Psi1=summationdisplay l,malmψlm. (9.44) Theconstantcoefficients almarefinallychosentopermit /Psi1tosatisfytheboundarycondi- tionsoftheproblem,which,asarule, leadtoadiscretesetofvalues l,m. Circular Cylindrical Coordinates Withourunknownfunction ψdependenton ρ,ϕ,andz,theHelmholtzequationbecomes (seeSection2.4for ∇2) ∇2ψ(ρ,ϕ,z)+k2ψ(ρ,ϕ,z)=0, (9.45) or 1 ρ∂ ∂ρparenleftbigg ρ∂ψ ∂ρparenrightbigg +1 ρ2∂2ψ ∂ϕ2+∂2ψ ∂z2+k2ψ=0. (9.46) Asbefore,weassumeafactoredformfor ψ, ψ(ρ,ϕ,z)=P(ρ)/Phi1(ϕ)Z(z). (9.47) SubstitutingintoEq.(9.46), wehave /Phi1Z ρd dρparenleftbigg ρdP dρparenrightbigg +PZ ρ2d2/Phi1 dϕ2+P/Phi1d2Z dz2+k2P/Phi1Z=0. (9.48) 556 Chapter 9 Differential Equations Allthepartialderivativeshavebecomeordinaryderivatives.Dividingby P/Phi1Zandmoving thezderivativetotheright-handsideyields 1 ρPd dρparenleftbigg ρdP dρparenrightbigg +1 ρ2/Phi1d2/Phi1 dϕ2+k2=−1 Zd2Z dz2. (9.49) Again, a function of zon the right appears to depend on a function of ρandϕon the left. We resolve this by setting each side of Eq. (9.49) equal to the same constant. Let us choose3−l2. Then d2Z dz2=l2Z (9.50) and 1 ρPd dρparenleftbigg ρdP dρparenrightbigg +1 ρ2/Phi1d2/Phi1 dϕ2+k2=−l2. (9.51) Settingk2+l2=n2, multiplyingby ρ2,andrearrangingterms, weobtain ρ Pd dρparenleftbigg ρdP dρparenrightbigg +n2ρ2=−1 /Phi1d2/Phi1 dϕ2. (9.52) Wemayset theright-handsideto m2and d2/Phi1 dϕ2=−m2/Phi1. (9.53) Finally,for the ρdependencewehave ρd dρparenleftbigg ρdP dρparenrightbigg +parenleftbig n2ρ2−m2parenrightbig P=0. (9.54) This is Bessel’s differential equation. The solutions and their properties are presented in Chapter11.TheseparationofvariablesofLaplace’sequationinparaboliccoordinatesalso givesrisetoBessel’sequation.ItmaybenotedthattheBesselequationisnotoriousforthe varietyofdisguisesitmayassume.Foranextensivetabulationofpossibleformsthereader isreferred to TablesofFunctions byJahnkeandEmde.4 The original Helmholtz equation, a three-dimensional PDE, has been replaced by three ODEs,Eqs. (9.50), (9.53), and(9.54). AsolutionoftheHelmholtzequationis ψ(ρ,ϕ,z)=P(ρ)/Phi1(ϕ)Z(z). (9.55) Identifyingthespecific P,/Phi1,Zsolutionsbysubscripts,weseethatthemostgeneralsolu- tionof theHelmholtzequationis alinearcombinationof theproductsolutions: /Psi1(ρ,ϕ,z)=summationdisplay m,namnPmn(ρ)/Phi1m(ϕ)Zn(z). (9.56) 3The choice of sign of the separation constant is arbitrary. However, a minus sign is chosen for the axial coordinate zin expec- tation of a possible exponential dependence on z(from Eq. (9.50)). A positive sign is chosen for the azimuthal coordinate ϕin expectation of aperiodic dependenceon ϕ(from Eq.(9.53)). 4E. Jahnke and F. Emde, Tables of functions , 4th rev. ed., New York: Dover (1945), p. 146; also, E. Jahnke, F. Emde, and F. Lösch, Tables of Higher Functions , 6th ed.,NewYork: McGraw-Hill(1960). 9.3 Separation of Variables 557 Spherical Polar Coordinates Let us try to separate the Helmholtz equation, again with k2constant, in spherical polar coordinates.UsingEq.(2.48), weobtain 1 r2sinθbracketleftbigg sinθ∂ ∂rparenleftbigg r2∂ψ ∂rparenrightbigg +∂ ∂θparenleftbigg sinθ∂ψ ∂θparenrightbigg +1 sinθ∂2ψ ∂ϕ2bracketrightbigg =−k2ψ. (9.57) Now,inanalogywithEq.(9.35) wetry ψ(r,θ,ϕ)=R(r)/Theta1(θ)/Phi1(ϕ). (9.58) BysubstitutingbackintoEq. (9.57)anddividingby R/Theta1/Phi1,weha v e 1 Rr2d drparenleftbigg r2dR drparenrightbigg +1 /Theta1r2sinθd dθparenleftbigg sinθd/Theta1 dθparenrightbigg +1 /Phi1r2sin2θd2/Phi1 dϕ2=−k2.(9.59) Note that all derivatives are now ordinary derivatives rather than partials. By multiplying byr2sin2θ,wecanisolate (1//Phi1)(d2/Phi1/dϕ2)toobtain5 1 /Phi1d2/Phi1 dϕ2=r2sin2θbracketleftbigg −k2−1 r2Rd drparenleftbigg r2dR drparenrightbigg −1 r2sinθ/Theta1d dθparenleftbigg sinθd/Theta1 dθparenrightbiggbracketrightbigg .(9.60) Equation (9.60) relates a function of ϕalone to a function of randθalone. Since r,θ, andϕareindependentvariables,weequateeachsideofEq.(9.60)toaconstant.Inalmost all physical problems ϕwill appear as an azimuth angle. This suggests a periodic solution rather than an exponential. With this in mind, let us use −m2as the separation constant, which,then,mustbeanintegersquared.Then 1 /Phi1d2/Phi1(ϕ) dϕ2=−m2(9.61) and 1 r2Rd drparenleftbigg r2dR drparenrightbigg +1 r2sinθ/Theta1d dθparenleftbigg sinθd/Theta1 dθparenrightbigg −m2 r2sin2θ=−k2.(9.62) MultiplyingEq. (9.62)by r2andrearrangingterms, weobtain 1 Rd drparenleftbigg r2dR drparenrightbigg +r2k2=−1 sinθ/Theta1d dθparenleftbigg sinθd/Theta1 dθparenrightbigg +m2 sin2θ. (9.63) Again,thevariablesareseparated.Weequateeachsidetoaconstant, Q,andfinallyobtain 1 sinθd dθparenleftbigg sinθd/Theta1 dθparenrightbigg −m2 sin2θ/Theta1+Q/Theta1=0, (9.64) 1 r2d drparenleftbigg r2dR drparenrightbigg +k2R−QR r2=0. (9.65) 5The order in which the variables are separated here is not unique. Many quantum mechanics texts show the rdependence split off first. 558 Chapter 9 Differential Equations Once more we have replaced a partial differential equation of three variables by three ODEs.ThesolutionsoftheseODEsarediscussedinChapters11and12.InChapter12,for example,Eq.(9.64)isidentifiedastheassociatedLegendreequation,inwhichtheconstant Qbecomes l(l+1);lis a non-negative integer because θis an angular variable. If k2is a(positive)constant,Eq. (9.65)becomesthesphericalBessel equationof Section11.7. Again,ourmostgeneralsolutionmaybewritten ψQm(r,θ,ϕ)=summationdisplay Q,maQmRQ(r)/Theta1Qm(θ)/Phi1m(ϕ). (9.66) The restriction that k2be a constant is unnecessarily severe. The separation process will stillbepossiblefor k2as generalas k2=f(r)+1 r2g(θ)+1 r2sin2θh(ϕ)+k′2. (9.67) In the hydrogen atom problem, one of the most important examples of the Schrödinger wave equation with a closed form solution is k2=f(r), withk2independent of θ,ϕ. Equation(9.65)for thehydrogenatombecomestheassociatedLaguerreequation. Thegreatimportanceofthisseparationofvariablesinsphericalpolarcoordinatesstems fromthefactthatthecase k2=k2(r)coversatremendousamountofphysics:agreatdeal ofthetheoriesofgravitation,electrostatics,andatomic,nuclear,andparticlephysics.And withk2=k2(r), the angular dependence is isolated in Eqs. (9.61) and (9.64), which can besolvedexactly . Finally, as an illustration of how the constant min Eq. (9.61) is restricted, we note that ϕin cylindrical and spherical polar coordinates is an azimuth angle. If this is a classical problem,weshallcertainlyrequirethattheazimuthalsolution /Phi1(ϕ)besingle-valued;that is, /Phi1(ϕ+2π)=/Phi1(ϕ). (9.68) Thisisequivalenttorequiringtheazimuthalsolutiontohaveaperiodof2 π.6Therefore m mustbeaninteger.Whichintegeritisdependsonthedetailsoftheproblem.Iftheinteger |m|>1,then/Phi1willhavetheperiod2 π/m.Wheneveracoordinatecorrespondstoanaxis oftranslationor toanazimuthangle,theseparatedequationalways hastheform d2/Phi1(ϕ) dϕ2=−m2/Phi1(ϕ) forϕ, theazimuthangle,and d2Z(z) dz2=±a2Z(z) (9.69) forz, an axis of translation of the cylindrical coordinate system. The solutions, of course, are sinazand cosazfor−a2and the corresponding hyperbolic function (or exponentials) sinhazand coshazfor+a2. 6This also applies in most quantum mechanical problems, but the argument is much more involved. If mis not an integer, rotation group relations and ladder operator relations (Section 4.3) are disrupted. Compare E. Merzbacher,Single valuedness of wavefunctions. Am.J .Ph ys. 30: 237 (1962). 9.3 Separation of Variables 559 Table 9.2 SolutionsinSphericalPolarCoordinatesa ψ=summationdisplay l,malmψlm 1. ∇2ψ=0ψlm=braceleftBigg rl r−l−1bracerightBiggbraceleftBigg Pm l(cosθ) Qm l(cosθ)bracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBiggb 2.∇2ψ+k2ψ=0ψlm=braceleftBigg jl(kr) nl(kr)bracerightBiggbraceleftBigg Pm l(cosθ) Qml(cosθ)bracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBiggb 3.∇2ψ−k2ψ=0ψlm=braceleftBigg il(kr) kl(kr)bracerightBiggbraceleftBigg Pm l(cosθ) Qm l(cosθ)bracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBiggb aReferences for some of the functions are Pm l(cosθ),m=0, Section 12.1; m/negationslash=0, Sec- tion12.5; Qm l(cosθ),Section12.10; jl(kr),nl(kr),il(kr),a ndkl(kr), Section11.7. bcosmϕandsinmϕmay be replacedby e±imϕ. Other occasionally encountered ODEs include the Laguerre and associated Laguerre equationsfromthesupremelyimportanthydrogenatomprobleminquantummechanics: xd2y dx2+(1−x)dy dx+αy=0, (9.70) xd2y dx2+(1+k−x)dy dx+αy=0. (9.71) Fromthequantummechanicaltheoryof thelinearoscillatorwehaveHermite’sequation, d2y dx2−2xdy dx+2αy=0. (9.72) Finally,fromtimetotimewefindtheChebyshevdifferentialequation, parenleftbig 1−x2parenrightbigd2y dx2−xdy dx+n2y=0. (9.73) For convenient reference, the forms of the solutions of Laplace’s equation, Helmholtz’s equation, and the diffusion equation for spherical polar coordinates are collected in Ta- ble 9.2. The solutions of Laplace’s equation in circular cylindrical coordinates are pre- sentedinTable9.3. Generalpropertiesfollowingfromtheformofthedifferentialequationsarediscussedin Chapter10. TheindividualsolutionsaredevelopedandappliedinChapters11–13. Thepracticingphysicistmayandprobablywillmeetothersecond-orderODEs,someof which may possibly be transformed into the examples studied here. Some of these ODEs may be solved by the techniques of Sections 9.5 and 9.6. Others may require a computer fora numericalsolution. We refertothesecondeditionof thistextfor otherimportantcoordinatesystems. •ToputtheseparationmethodofsolvingPDEsinperspective,letusreviewitasaconse- quenceofasymmetryofthePDE.TakethestationarySchrödingerequation Hψ=Eψ asanexample,withapotential V(r)dependingonlyontheradialdistance r.Thenthis 560 Chapter 9 Differential Equations Table 9.3 SolutionsinCircularCylindricalCoordinatesa ψ=summationdisplay m,αamαψmα a.∇2ψ+α2ψ=0ψmα=braceleftBigg Jm(αρ) Nm(αρ)bracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBiggbraceleftBigg e−αz eαzbracerightBigg b.∇2ψ−α2ψ=0ψmα=braceleftBigg Im(αρ) Km(αρ)bracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBiggbraceleftBigg cosαz sinαzbracerightBigg c. ∇2ψ=0ψm=braceleftBigg ρm ρ−mbracerightBiggbraceleftBigg cosmϕ sinmϕbracerightBigg aReferencesfortheradialfunctionsare Jm(αρ),Section11.1; Nm(αρ),Section11.3; Im(αρ)andKm(αρ), Section11.5. PDE is invariant under rotations that comprise the group SO(3). Its diagonal genera- tor is the orbital angular momentum operator Lz=−i∂ ∂ϕ, and its quadratic (Casimir) invariant is L2. Since both commute with H(see Section 4.3), we end up with three separateeigenvalueequations: Hψ=Eψ, L2ψ=l(l+1)ψ, L zψ=mψ. Uponreplacing L2 zinL2byitseigenvalue m2,theL2PDEbecomesLegendre’sODE, andsimilarly Hψ=EψbecomestheradialODEoftheseparationmethodinspherical polarcoordinates. •For cylindrical coordinates the PDE is invariant under rotations about the z-axis only, which form a subgroup of SO(3). This invariance yields the generator Lz=−i∂/∂ϕ andseparateazimuthalODE Lzψ=mψ,asbefore.Ifthepotential Visinvariantunder translations along the z-axis, then the generator −i∂/∂zgives the separate ODE in the zvariable. •Ingeneral(seeSection4.3),thereare nmutuallycommutinggenerators Hiwitheigen- valuesmiof the (classical) Lie group Gof ranknand the corresponding Casimir in- variantsCiwitheigenvalues ci(Chapter4), whichyieldtheseparateODEs Hiψ=miψ, C iψ=ciψ inadditiontothe(bynow)radialODE Hψ=Eψ. Exercises 9.3.1 By letting the operator ∇2+k2act on the general form a1ψ1(x,y,z)+a2ψ2(x,y,z), show that it is linear, that is, that (∇2+k2)(a1ψ1+a2ψ2)=a1(∇2+k2)ψ1+ a2(∇2+k2)ψ2. 9.3.2 ShowthattheHelmholtzequation, ∇2ψ+k2ψ=0, 9.3 Separation of Variables 561 is still separable in circular cylindrical coordinates if k2is generalized to k2+f(ρ)+ (1/ρ2)g(ϕ)+h(z). 9.3.3 SeparatevariablesintheHelmholtzequationinsphericalpolarcoordinates,splittingoff the radial dependence first. Show that your separated equations have the same form as Eqs. (9.61), (9.64), and(9.65). 9.3.4 Verifythat ∇2ψ(r,θ,ϕ)+bracketleftbigg k2+f(r)+1 r2g(θ)+1 r2sin2θh(ϕ)bracketrightbigg ψ(r,θ,ϕ)=0 is separable (in spherical polar coordinates). The functions f,g, andhare functions onlyof thevariablesindicated; k2isaconstant. 9.3.5 An atomic (quantum mechanical) particle is confined inside a rectangular box of sides a,b,andc.Theparticleisdescribedbyawavefunction ψthatsatisfiestheSchrödinger waveequation −¯h2 2m∇2ψ=Eψ. The wave function is required to vanish at each surface of the box (but not to be identi- callyzero).Thisconditionimposesconstraintsontheseparationconstantsandtherefore on the energy E. What is the smallest value of Efor which such a solution can be ob- tained? ANS.E=π2¯h2 2mparenleftbigg1 a2+1 b2+1 c2parenrightbigg . 9.3.6 For a homogeneous spherical solid with constant thermal diffusivity, K, and no heat sources,theequationofheatconductionbecomes ∂T(r,t) ∂t=K∇2T(r,t). Assumeasolutionoftheform T=R(r)T(t) andseparatevariables.Showthattheradialequationmaytakeonthestandardform r2d2R dr2+2rdR dr+bracketleftbig α2r2−n(n+1)bracketrightbig R=0;n=integer. Thesolutionsof thisequationarecalled sphericalBesselfunctions . 9.3.7 Separate variables in the thermal diffusion equation of Exercise 9.3.6 in circular cylin- dricalcoordinates.Assumethatyoucanneglectendeffects andtake T=T(ρ,t). 9.3.8 Thequantummechanicalangularmomentumoperatorisgivenby L=−i(r×∇).Show that L·Lψ=l(l+1)ψ leadstotheassociatedLegendreequation. Hint.Exercises1.9.9and2.5.16maybehelpful. 562 Chapter 9 Differential Equations 9.3.9 The one-dimensional Schrödinger wave equation for a particle in a potential field V= 1 2kx2is −¯h2 2md2ψ dx2+1 2kx2ψ=Eψ(x). (a) Using ξ=axandaconstant λ,weha v e a=parenleftbiggmk ¯h2parenrightbigg1/4 ,λ=2E ¯hparenleftbiggm kparenrightbigg1/2 ; showthat d2ψ(ξ) dξ2+parenleftbig λ−ξ2parenrightbig ψ(ξ)=0. (b) Substituting ψ(ξ)=y(ξ)e−ξ2/2, showthat y(ξ)satisfiestheHermitedifferentialequation. 9.3.10 Verifythatthefollowingaresolutionsof Laplace’sequation: (a)ψ1=1/r,r/negationslash=0, (b) ψ2=1 2rlnr+z r−z. Note.T h ezderivatives of 1 /rgenerate the Legendre polynomials, Pn(cosθ),E x e r - cise12.1.7.The zderivativesof (1/2r)ln[(r+z)/(r−z)]generatetheLegendrefunc- tions,Qn(cosθ). 9.3.11 If/Psi1isasolutionof Laplace’sequation, ∇2/Psi1=0,showthat ∂/Psi1/∂zisalsoa solution. 9.4 S INGULAR POINTS In this section the concept of a singular point, or singularity (as applied to a differential equation), is introduced. The interest in this concept stems from its usefulness in (1) clas- sifyingODEsand(2)investigatingthefeasibilityofaseriessolution.Thisfeasibilityisthe topicofFuchs’theorem,Sections9.5and9.6. All the ODEs listed in Section 9.3 may be solved for d2y/dx2. Using the notation d2y/dx2=y′′,weha v e7 y′′=f(x,y,y′). (9.74) If wewriteoursecond-orderhomogeneousdifferentialequation(in y)as y′′+P(x)y′+Q(x)y=0, (9.75) wearereadytodefineordinaryandsingularpoints.Ifthefunctions P(x)andQ(x)remain finite atx=x0, pointx=x0is an ordinary point. However, if either P(x)orQ(x)(or 7This prime notation, y′=dy/dx, was introduced by Lagrange in the late 18th century as an abbreviation for Leibniz’s more explicit but more cumbersome dy/dx. 9.4 Singular Points 563 both)divergesas x→x0,pointx0isasingularpoint.UsingEq.(9.75),wemaydistinguish betweentwokindsofsingularpoints. 1. If either P(x)orQ(x)divergesas x→x0but(x−x0)P(x)and (x−x0)2Q(x)remain finite as x→x0, thenx=x0is called a regular, or nonessen- tial,singularpoint. 2. IfP(x)divergesfasterthan1 /(x−x0)sothat(x−x0)P(x)goestoinfinityas x→x0, orQ(x)diverges faster than 1 /(x−x0)2so that(x−x0)2Q(x)goes to infinity as x→x0, thenpoint x=x0is labeledan irregular ,oressential,singularity . Thesedefinitionsholdforallfinitevaluesof x0.Theanalysisofpoint x→∞issimilar tothetreatmentoffunctionsofacomplexvariable(Section6.6).Weset x=1/z,substitute intothedifferentialequation,andthenlet z→0.Bychangingvariablesinthederivatives, wehave dy(x) dx=dy(z−1) dzdz dx=−1 x2dy(z−1) dz=−z2dy(z−1) dz, (9.76) d2y(x) dx2=d dzbracketleftbiggdy(x) dxbracketrightbiggdz dx=parenleftbig −z2parenrightbigbracketleftbigg −2zdy(z−1) dz−z2d2y(z−1) dz2bracketrightbigg =2z3dy(z−1) dz+z4d2y(z−1) dz2. (9.77) Usingtheseresults, wetransform Eq. (9.75)into z4d2y dz2+bracketleftbig 2z3−z2Pparenleftbig z−1parenrightbigbracketrightbigdy dz+Qparenleftbig z−1parenrightbig y=0. (9.78) Thebehaviorat x=∞(z=0)thendependsonthebehaviorof thenewcoefficients, 2z−P(z−1) z2andQ(z−1) z4, asz→0. If these two expressions remain finite, point x=∞is an ordinary point. If they divergenomorerapidlythan 1 /zand 1/z2,respectively,point x=∞isaregularsingular point;otherwiseitis anirregularsingularpoint(anessentialsingularity). Example 9.4.1 Bessel’sequationis x2y′′+xy′+parenleftbig x2−n2parenrightbig y=0. (9.79) ComparingitwithEq. (9.75)wehave P(x)=1 x,Q(x)=1−n2 x2, which shows that point x=0 is a regular singularity. By inspection we see that there are no other singular points in the finite range. As x→∞(z→0), from Eq. (9.78) we have 564 Chapter 9 Differential Equations Table 9.4 Regular Irregular singularity singularity Equation x= x= 1. Hypergeometric 0,1,∞ – x(x−1)y′′+[(1+a+b)x−c]y′+aby=0. 2. Legendrea−1,1,∞ – (1−x2)y′′−2xy′+l(l+1)y=0. 3. Chebyshev −1,1,∞ – (1−x2)y′′−xy′+n2y=0. 4. Confluent hypergeometric 0 ∞ xy′′+(c−x)y′−ay=0. 5. Bessel 0 ∞ x2y′′+xy′+(x2−n2)y=0. 6. Laguerrea0 ∞ xy′′+(1−x)y′+ay=0. 7. Simple harmonic oscillator – ∞ y′′+ω2y=0. 8. Hermite – ∞ y′′−2xy′+2αy=0. aTheassociatedequationshavethesamesingularpoints. thecoefficients 2z−z z2and1−n2z2 z4. Since the latter expression diverges as z4, pointx=∞is an irregular, or essential, singu- larity. /squaresolid The ordinary differential equations of Section 9.3, plus two others, the hypergeometric andtheconfluenthypergeometric,havesingularpoints,asshowninTable9.4. It will be seen that the first three equations in Table 9.4, hypergeometric, Legendre, and Chebyshev,allhavethreeregularsingularpoints.Thehypergeometricequation,withregu- larsingularitiesat0,1,and ∞istakenasthestandard,thecanonicalform.Thesolutionsof theothertwomaythenbeexpressedintermsofitssolutions,thehypergeometricfunctions. Thisis doneinChapter13. In a similar manner, the confluent hypergeometric equation is taken as the canonical form of a linear second-order differential equation with one regular and one irregular sin- gularpoint. Exercises 9.4.1 ShowthatLegendre’sequationhasregularsingularitiesat x=−1,1,and∞. 9.4.2 Show that Laguerre’s equation, like the Bessel equation, has a regular singularity atx=0 andanirregularsingularityat x=∞. 9.5 Series Solutions — Frobenius’ Method 565 9.4.3 Showthatthesubstitution x→1−x 2,a=−l, b=l+1,c=1 convertsthehypergeometricequationintoLegendre’sequation. 9.5 S ERIES SOLUTIONS —F ROBENIUS ’M ETHOD In this section we develop a method of obtaining one solution of the linear, second-order, homogeneousODE.Themethod,aseriesexpansion,willalwayswork,providedthepoint ofexpansionisnoworsethanaregularsingularpoint.Inphysicsthisverygentlecondition isalmostalwayssatisfied. Alinear,second-order,homogeneous ODEmaybeputintheform d2y dx2+P(x)dy dx+Q(x)y=0. (9.80) The equation is homogeneous because each term contains y(x)or a derivative; linear because each y,dy/dx,o rd2y/dx2appears as the first power—and no products. In this section we develop (at least) one solution of Eq. (9.80). In Section 9.6 we develop the second, independent solution and prove that no third, independent solution exists . Thereforethe mostgeneralsolution ofEq. (9.80)maybewrittenas y(x)=c1y1(x)+c2y2(x). (9.81) Ourphysicalproblemmayleadtoa nonhomogeneous ,linear,second-orderODE, d2y dx2+P(x)dy dx+Q(x)y=F(x). (9.82) The function on the right, F(x), represents a source (such as electrostatic charge) or a driving force (as in a driven oscillator). Specific solutions of this nonhomogeneous equa- tion are touched on in Exercise 9.6.25. They are explored in some detail, using Green’s function techniques, in Sections 9.7 and 10.5, and with a Laplace transform technique in Section15.11.Callingthissolution yp,wemayaddtoitanysolutionofthecorresponding homogeneousequation(Eq. (9.80)). Hencethe most generalsolution of Eq.(9.82) is y(x)=c1y1(x)+c2y2(x)+yp(x). (9.83) Theconstants c1andc2willeventuallybefixedbyboundaryconditions. For the present, we assume that F(x)=0 and that our differential equation is homoge- neous. We shall attempt to develop a solution of our linear, second-order, homogeneous differentialequation,Eq.(9.80),bysubstitutinginapowerserieswithundeterminedcoef- ficients. Also available as a parameter is the power of the lowest nonvanishing term of the series. To illustrate, we apply the method to two important differential equations, first the 566 Chapter 9 Differential Equations linear(classical)oscillatorequation d2y dx2+ω2y=0, (9.84) withknownsolutions y=sinωx,cosωx. We try y(x)=xkparenleftbig a0+a1x+a2x2+a3x3+···parenrightbig =∞summationdisplay λ=0aλxk+λ,a 0/negationslash=0, (9.85) with the exponent kand all the coefficients aλstill undetermined. Note that kneed not be aninteger.Bydifferentiatingtwice,weobtain dy dx=∞summationdisplay λ=0aλ(k+λ)xk+λ−1, d2y dx2=∞summationdisplay λ=0aλ(k+λ)(k+λ−1)xk+λ−2. BysubstitutingintoEq.(9.84), wehave ∞summationdisplay λ=0aλ(k+λ)(k+λ−1)xk+λ−2+ω2∞summationdisplay λ=0aλxk+λ=0. (9.86) From our analysis of the uniqueness of power series (Chapter 5), the coefficients of each powerof xontheleft-handsideof Eq.(9.86) mustvanishindividually. Thelowestpowerof xappearinginEq.(9.86)is xk−2,forλ=0 inthefirstsummation. Therequirementthatthecoefficientvanish8yields a0k(k−1)=0. We had chosen a0as the coefficient of the lowest nonvanishing terms of the series (Eq. (9.85)), hence,bydefinition, a0/negationslash=0.Thereforewehave k(k−1)=0. (9.87) This equation, coming from the coefficient of the lowest power of x, we call the indicial equation. The indicial equation and its roots are of critical importance to our analysis. Ifk=1, the coefficient a1(k+1)kofxk−1must vanish so that a1=0. Clearly, in this examplewemustrequireeitherthat k=0o rk=1. Beforeconsideringthesetwopossibilitiesfor k,wereturntoEq.(9.86)anddemandthat theremainingnetcoefficients,say,thecoefficientof xk+j(j≥0),vanish.Weset λ=j+2 in the first summation and λ=jin the second. (They are independent summations and λ isadummyindex.)Thisresults in aj+2(k+j+2)(k+j+1)+ω2aj=0 8Seethe uniqueness of powerseries,Section 5.7. 9.5 Series Solutions — Frobenius’ Method 567 or aj+2=−ajω2 (k+j+2)(k+j+1). (9.88) This is a two-term recurrencerelation .9Givenaj, we may compute aj+2and thenaj+4, aj+6, and so on up as far as desired. Note that for this example, if we start with a0, Eq. (9.88) leads to the even coefficients a2,a4, and so on, and ignores a1,a3,a5, and so on. Since a1is arbitrary if k=0 and necessarily zero if k=1, let us set it equal to zero (compareExercises9.5.3and9.5.4) andthenbyEq. (9.88) a3=a5=a7=···=0, and all the odd-numbered coefficients vanish. The odd powers of xwill actually reappear whenthe secondrootoftheindicialequationis used. Returning to Eq. (9.87) our indicial equation, we first try the solution k=0. The recur- rencerelation(Eq. (9.88))becomes aj+2=−ajω2 (j+2)(j+1), (9.89) whichleadsto a2=−a0ω2 1·2=−ω2 2!a0, a4=−a2ω2 3·4=+ω4 4!a0, a6=−a4ω2 5·6=−ω6 6!a0,andso on. Byinspection(andmathematicalinduction), a2n=(−1)nω2n (2n)!a0, (9.90) andoursolutionis y(x)k=0=a0bracketleftbigg 1−(ωx)2 2!+(ωx)4 4!−(ωx)6 6!+···bracketrightbigg =a0cosωx. (9.91) Ifwechoosetheindicialequationroot k=1 (Eq.(9.88)),therecurrencerelationbecomes aj+2=−ajω2 (j+3)(j+2). (9.92) 9The recurrence relation may involve three terms, that is, aj+2, depending on ajandaj−2. Equation (13.2) for the Hermite functions provides anexample ofthis behavior. 568 Chapter 9 Differential Equations Substitutingin j=0,2,4,successively,weobtain a2=−a0ω2 2·3=−ω2 3!a0, a4=−a2ω2 4·5=+ω4 5!a0, a6=−a4ω2 6·7=−ω6 7!a0,andso on. Again,byinspectionandmathematicalinduction, a2n=(−1)nω2n (2n+1)!a0. (9.93) Forthischoice, k=1,weobtain y(x)k=1=a0xbracketleftbigg 1−(ωx)2 3!+(ωx)4 5!−(ωx)6 7!+···bracketrightbigg =a0 ωbracketleftbigg (ωx)−(ωx)3 3!+(ωx)5 5!−(ωx)7 7!+···bracketrightbigg =a0 ωsinωx. (9.94) To summarize this approach, we may write Eq . (9.86)schematically as shown in Fig .9 . 2 . From the uniqueness of power series (Section5.7),the total coefficient of each power of x must vanish—all by itself. The requirement that the first coefficient (1)vanish leads to the indicial equation, Eq . (9.87).The second coefficient is handled by setting a1=0. The vanishing of the coefficient of xk(and higher powers, taken one at a time )leads to the recurrencerelation,Eq . (9.88). This series substitution, known as Frobenius’ method, has given us two series solutions of the linear oscillator equation. However, there are two points about such series solutions thatmustbestronglyemphasized: 1. Theseriessolutionshouldalwaysbesubstitutedbackintothedifferentialequation,to see if it works, as a precaution against algebraic and logical errors. If it works, it is asolution. 2. Theacceptabilityofaseriessolutiondependsonitsconvergence(includingasymptotic convergence). It is quite possible for Frobenius’ method to give a series solution that satisfiestheoriginaldifferentialequationwhensubstitutedintheequationbutthatdoes FIGURE 9.2Recurrencerelationfrompowerseriesexpansion. 9.5 Series Solutions — Frobenius’ Method 569 notconvergeovertheregionofinterest.Legendre’sdifferentialequationillustratesthis situation. Expansion About x0 Equation (9.85) is an expansionabout the origin, x0=0. It is perfectly possible to replace Eq.(9.85) with y(x)=∞summationdisplay λ=0aλ(x−x0)k+λ,a 0/negationslash=0. (9.95) Indeed,fortheLegendre,Chebyshev,andhypergeometricequationsthechoice x0=1 has some advantages. The point x0should not be chosen at an essential singularity—or our Frobenius method will probably fail. The resultant series ( x0an ordinary point or regular singular point) will be valid where it converges. You can expect a divergence of some sort when|x−x0|=|zs−x0|,wherezsistheclosestsingularityto x0(inthecomplexplane). Symmetry of Solutions Let us note that we obtained one solution of even symmetry, y1(x)=y1(−x), and one of odd symmetry, y2(x)=−y2(−x). This is not just an accident but a direct consequence of theform oftheODE.WritingageneralODEas L(x)y(x)=0, (9.96) in which L(x)is the differential operator, we see that for the linear oscillator equation (Eq. (9.84)), L(x)isevenunderparity;thatis, L(x)=L(−x). (9.97) Wheneverthedifferentialoperatorhasaspecificparityorsymmetry,eitherevenorodd, wemayinterchange +xand−x,andEq. (9.96)becomes ±L(x)y(−x)=0, (9.98) +ifL(x)iseven,−ifL(x)isodd.Clearly,if y(x)isasolutionofthedifferentialequation, y(−x)is alsoasolution.Thenanysolutionmayberesolvedintoevenandoddparts, y(x)=1 2bracketleftbig y(x)+y(−x)bracketrightbig +1 2bracketleftbig y(x)−y(−x)bracketrightbig , (9.99) thefirst bracketontherightgivinganevensolution,thesecondanoddsolution. If we refer back to Section 9.4, we can see that Legendre, Chebyshev, Bessel, simple harmonic oscillator, and Hermite equations (or differential operators) all exhibit this even parity; that is, their P(x)in Eq. (9.80) is odd and Q(x)even. Solutions of all of them may be presented as series of even powers of xand separate series of odd powers of x. The Laguerre differential operator has neither even nor odd symmetry; hence its solutions cannot be expected to exhibit even or odd parity. Our emphasis on parity stems primarily from the importance of parity in quantum mechanics. We find that wave functions usually are either even or odd, meaning that they have a definite parity. Most interactions (beta decayisthebigexception)arealsoevenor odd,andtheresultis thatparityis conserved. 570 Chapter 9 Differential Equations Limitations of Series Approach — Bessel’s Equation This attack on the linear oscillator equation was perhaps a bit too easy. By substituting the power series (Eq. (9.85)) into the differential equation (Eq. (9.84)), we obtained two independentsolutionswithnotroubleatall. Togetsomeideaof whatcanhappenwetrytosolveBessel’sequation, x2y′′+xy′+parenleftbig x2−n2parenrightbig y=0, (9.100) usingy′fordy/dxandy′′ford2y/dx2. Again,assumingasolutionoftheform y(x)=∞summationdisplay λ=0aλxk+λ, wedifferentiateandsubstituteintoEq. (9.100). Theresultis ∞summationdisplay λ=0aλ(k+λ)(k+λ−1)xk+λ+∞summationdisplay λ=0aλ(k+λ)xk+λ +∞summationdisplay λ=0aλxk+λ+2−∞summationdisplay λ=0aλn2xk+λ=0. (9.101) By setting λ=0, we get the coefficient of xk, the lowest power of xappearing on the left-handside, a0bracketleftbig k(k−1)+k−n2bracketrightbig =0, (9.102) andagain a0/negationslash=0 bydefinition.Equation(9.102)thereforeyieldsthe indicialequation k2−n2=0 (9.103) withsolutions k=±n. It isof someinteresttoexaminethecoefficientof xk+1also.Hereweobtain a1bracketleftbig (k+1)k+k+1−n2bracketrightbig =0, or a1(k+1−n)(k+1+n)=0. (9.104) Fork=±n,neitherk+1−nnork+1+nvanishesandwe mustrequirea1=0.10 Proceeding to the coefficient of xk+jfork=n,w es e tλ=jin the first, second, and fourth terms of Eq. (9.101) and λ=j−2 in the third term. By requiring the resultant coefficientof xk+1tovanish,weobtain ajbracketleftbig (n+j)(n+j−1)+(n+j)−n2bracketrightbig +aj−2=0. Whenjis replacedby j+2,thiscanberewrittenfor j≥0a s aj+2=−aj1 (j+2)(2n+j+2), (9.105) 10k=±n=−1 2areexceptions. 9.5 Series Solutions — Frobenius’ Method 571 which is the desired recurrence relation. Repeated application of this recurrence relation leadsto a2=−a01 2(2n+2)=−a0n! 221!(n+1)!, a4=−a21 4(2n+4)=a0n! 242!(n+2)!, a6=−a41 6(2n+6)=−a0n! 263!(n+3)!,andsoon, andingeneral, a2p=(−1)pa0n! 22pp!(n+p)!. (9.106) Insertingthesecoefficientsinour assumedseries solution,wehave y(x)=a0xnbracketleftbigg 1−n!x2 221!(n+1)!+n!x4 242!(n+2)!−···bracketrightbigg . (9.107) Insummationform y(x)=a0∞summationdisplay j=0(−1)jn!xn+2j 22jj!(n+j)! =a02nn!∞summationdisplay j=0(−1)j1 j!(n+j)!parenleftbiggx 2parenrightbiggn+2j . (9.108) In Chapter 11 the final summation is identified as the Bessel function Jn(x). Notice that this solution, Jn(x), has either even or odd symmetry,11as might be expected from the formofBessel’s equation. Whenk=−nandnis not an integer, we may generate a second distinct series, to be labeledJ−n(x).However,when −nisanegativeinteger,troubledevelops.Therecurrence relation for the coefficients ajis still given by Eq. (9.105), but with 2 nreplaced by−2n. Then, when j+2=2norj=2(n−1), the coefficient aj+2blows up and we have no seriessolution.ThiscatastrophecanberemediedinEq.(9.108),asitisdoneinChapter11, withtheresultthat J−n(x)=(−1)nJn(x), n aninteger . (9.109) The second solution simply reproduces the first. We have failed to construct a second in- dependentsolutionfor Bessel’sequationbythisseries techniquewhen nis aninteger. By substituting in an infinite series, we have obtained two solutions for the linear oscil- lator equation and one for Bessel’s equation (two if nis not an integer). To the questions “Can we always do this? Will this method always work?” the answer is no, we cannot alwaysdothis.This methodofseries solutionwillnotalwayswork. 11Jn(x)is an even function if nis an even integer, an odd function if nis an odd integer. For nonintegral nthexnhas no such simple symmetry. 572 Chapter 9 Differential Equations Regular and Irregular Singularities Thesuccessoftheseriessubstitutionmethoddependsontherootsoftheindicialequation and the degree of singularity of the coefficients in the differential equation. To understand better the effect of the equation coefficients on this naive series substitution approach, considerfour simpleequations: y′′−6 x2y=0, (9.110a) y′′−6 x3y=0, (9.110b) y′′+1 xy′−a2 x2y=0, (9.110c) y′′+1 x2y′−a2 x2y=0. (9.110d) Thereadermayshoweasilythatfor Eq. (9.110a)theindicialequationis k2−k−6=0, givingk=3,−2.Sincetheequationishomogeneousin x(counting d2/dx2asx−2),there is no recurrence relation. However,we are left with two perfectly good solutions, x3and x−2. Equation (9.110b) differs from Eq. (9.110a) by only one power of x, but this sends the indicialequationto −6a0=0, with no solution at all, for we have agreed that a0/negationslash=0. Our series substitution worked for Eq. (9.110a), which had only a regular singularity, but broke down at Eq. (9.110b), which hasanirregularsingularpointattheorigin. ContinuingwithEq. (9.110c),wehaveaddedaterm y′/x.Theindicialequationis k2−a2=0, but again, there is no recurrence relation. The solutions are y=xa,x−a, both perfectly acceptableone-termseries. When we change the power of xin the coefficient of y′from−1t o−2, Eq. (9.110d), there is a drastic change in the solution. The indicial equation (with only the y′term con- tributing)becomes k=0. Thereisarecurrencerelation, aj+1=+aja2−j(j−1) j+1. 9.5 Series Solutions — Frobenius’ Method 573 Unlesstheparameter aisselectedtomaketheseriesterminate,wehave lim j→∞vextendsinglevextendsinglevextendsinglevextendsingleaj+1 ajvextendsinglevextendsinglevextendsinglevextendsingle=lim j→∞j(j+1) j+1 =lim j→∞j2 j=∞. Hence our series solution diverges for all x/negationslash=0. Again, our method worked for Eq. (9.110c) with a regular singularity but failed when we had the irregular singularity ofEq. (9.110d). Fuchs’ Theorem The answer to the basic question when the method of series substitution can be expected to work is given by Fuchs’ theorem, which asserts that we can always obtain at least one power-seriessolution,providedweareexpandingaboutapointthatisanordinarypointor atworstaregularsingularpoint. If we attempt an expansion about an irregular or essential singularity, our method may fail, as it did for Eqs. (9.110b) and (9.110d). Fortunately, the more important equations of mathematical physics, listed in Section 9.4, have no irregular singularities in the finite plane.FurtherdiscussionofFuchs’theoremappearsinSection9.6. From Table 9.4, Section 9.4, infinity is seen to be a singular point for all equations considered. As a further illustration of Fuchs’ theorem, Legendre’s equation (with infinity asaregularsingularity)hasaconvergent-seriessolutioninnegativepowersoftheargument (Section 12.10). In contrast, Bessel’s equation (with an irregular singularity at infinity) yieldsasymptoticseries(Sections5.10and11.6).Theseasymptoticsolutionsareextremely useful. Summary If we are expanding about an ordinary point or at worst about a regular singularity, the seriessubstitutionapproachwillyieldatleastonesolution(Fuchs’theorem). Whether we get one or two distinct solutions depends on the roots of the indicial equa- tion. 1. If the two roots of the indicial equation are equal, we can obtain only one solution by thisseries substitutionmethod. 2. If the two roots differ by a nonintegral number, two independent solutions may be obtained. 3. If thetworootsdiffer byaninteger,thelargerofthetwowillyieldasolution. The smaller may or may not give a solution, depending on the behavior of the coeffi- cients. In the linear oscillator equation we obtain two solutions; for Bessel’s equation, we getonlyonesolution. 574 Chapter 9 Differential Equations The usefulness of the series solution in terms of what the solution is (that is, numbers) dependsontherapidityofconvergenceoftheseriesandtheavailabilityofthecoefficients. ManyODEswillnotyieldnice,simplerecurrencerelationsforthecoefficients.Ingeneral, the available series will probably be useful for |x|(or|x−x0|) very small. Computers can be used to determine additional series coefficients using a symbolic language, such as Mathematica,12Maple,13or Reduce.14Often, however, for numerical work a direct numericalintegrationwillbepreferred. Exercises 9.5.1 Uniqueness theorem. The function y(x)satisfies a second-order, linear, homogeneous differential equation. At x=x0,y(x)=y0anddy/dx=y′ 0. Show that y(x)is unique, in that no other solution of this differential equation passes through the points (x0,y0) withaslopeof y′ 0. Hint. Assume a second solution satisfying these conditions and compare the Taylor series expansions. 9.5.2 A series solution of Eq. (9.80) is attempted, expanding about the point x=x0.I fx0is anordinarypoint,showthattheindicialequationhasroots k=0,1. 9.5.3 In the development of a series solution of the simple harmonic oscillator (SHO) equa- tion, the second series coefficient a1was neglected except to set it equal to zero. From the coefficient of the next-to-the-lowest power of x,xk−1, develop a second indicial- typeequation. (a) (SHO equation with k=0). Show that a1, may be assigned any finite value (in- cludingzero). (b) (SHOequationwith k=1).Showthat a1mustbeset equaltozero. 9.5.4 Analyzetheseriessolutionsofthefollowingdifferentialequationstoseewhen a1may besetequaltozerowithoutirrevocablylosinganythingandwhen a1mustbesetequal tozero. (a) Legendre,(b) Chebyshev,(c) Bessel, (d) Hermite. ANS.(a) Legendre,(b) Chebyshev,and(d) Hermite:For k=0,a1 maybesetequaltozero;for k=1,a1mustbesetequal tozero. (c) Bessel: a1mustbesetequaltozero(exceptfor k=±n=−1 2). 9.5.5 SolvetheLegendreequation parenleftbig 1−x2parenrightbig y′′−2xy′+n(n+1)y=0 bydirectseries substitution. 12S. Wolfram, Mathematica, ASystem for Doing Mathematics byComputer , NewYork: Addison Wesley(1991). 13A.Heck, Introduction to Maple ,NewYork: Springer (1993). 14G.Rayna, ReduceSoftware for Algebraic Computation , NewYork: Springer (1987). 9.5 Series Solutions — Frobenius’ Method 575 (a) Verifythattheindicialequationis k(k−1)=0. (b) Using k=0,obtainaseries ofevenpowersof x( a1=0). yeven=a0bracketleftbigg 1−n(n+1) 2!x2+n(n−2)(n+1)(n+3) 4!x4+···bracketrightbigg , where aj+2=j(j+1)−n(n+1) (j+1)(j+2)aj. (c) Using k=1,developaseriesof oddpowersof x( a1=1). yodd=a1bracketleftbigg x−(n−1)(n+2) 3!x3+(n−1)(n−3)(n+2)(n+4) 5!x5+···bracketrightbigg , where aj+2=(j+1)(j+2)−n(n+1) (j+2)(j+3)aj. (d) Showthatbothsolutions, yevenandyodd,divergefor x=±1iftheseriescontinue toinfinity . (e) Finally, show that by an appropriate choice of n, one series at a time may be con- vertedintoapolynomial,therebyavoidingthedivergencecatastrophe.Inquantum mechanics this restriction of nto integral values corresponds to quantization of angularmomentum . 9.5.6 Developseries solutionsfor Hermite’sdifferentialequation (a)y′′−2xy′+2αy=0. ANS.k(k−1)=0,indicialequation. Fork=0, aj+2=2ajj−α (j+1)(j+2)(jeven), yeven=a0bracketleftbigg 1+2(−α)x2 2!+22(−α)(2−α)x4 4!+···bracketrightbigg . Fork=1, aj+2=2ajj+1−α (j+2)(j+3)(jeven), yodd=a1bracketleftbigg x+2(1−α)x3 3!+22(1−α)(3−α)x5 5!+···bracketrightbigg . 576 Chapter 9 Differential Equations (b) Show that both series solutions are convergent for all x, the ratio of successive coefficientsbehaving,forlargeindex,likethecorrespondingratiointheexpansion of exp(x2). (c) Show that by appropriate choice of αthe series solutions may be cut off and con- verted to finite polynomials. (These polynomials, properly normalized, become theHermitepolynomialsinSection13.1.) 9.5.7 Laguerre’sODEis xL′′ n(x)+(1−x)L′n(x)+nLn(x)=0. Developaseries solutionselectingtheparameter ntomakeyourseriesapolynomial. 9.5.8 SolvetheChebyshevequation parenleftbig 1−x2parenrightbig T′′ n−xT′ n+n2Tn=0, by series substitution. What restrictions are imposedon nif you demandthat the series solutionconvergefor x=±1? ANS.Theinfiniteseries doesconvergefor x=±1andno restrictionon nexists(compareExercise5.2.16). 9.5.9 Solve parenleftbig 1−x2parenrightbig U′′ n(x)−3xU′ n(x)+n(n+2)Un(x)=0, choosing the root of the indicial equation to obtain a series of oddpowers of x. Since theserieswilldivergefor x=1,choose ntoconvertitintoa polynomial. k(k−1)=0. Fork=1, aj+2=(j+1)(j+3)−n(n+2) (j+2)(j+3)aj. 9.5.10 Obtainaseries solutionofthehypergeometricequation x(x−1)y′′+bracketleftbig (1+a+b)x−cbracketrightbig y′+aby=0. Test yoursolutionforconvergence. 9.5.11 Obtaintwoseriessolutionsoftheconfluenthypergeometricequation xy′′+(c−x)y′−ay=0. Test yoursolutionsfor convergence. 9.5.12 A quantum mechanical analysis of the Stark effect (parabolic coordinates) leads to the differentialequation d dξparenleftbigg ξdu dξparenrightbigg +parenleftbigg1 2Eξ+α−m2 4ξ−1 4Fξ2parenrightbigg u=0. Hereαis a separation constant, Eis the total energy, and Fis a constant, where Fzis thepotentialenergyaddedtothesystembytheintroductionof anelectricfield. 9.5 Series Solutions — Frobenius’ Method 577 Using the larger root of the indicial equation, develop a power-series solution about ξ=0.Evaluatethefirstthreecoefficientsintermsof ao. Indicialequation k2−m2 4=0, u(ξ)=a0ξm/2braceleftbigg 1−α m+1ξ+bracketleftbiggα2 2(m+1)(m+2)−E 4(m+2)bracketrightbigg ξ2+···bracerightbigg . Notethattheperturbation Fdoesnotappearuntil a3isincluded. 9.5.13 For the special case of no azimuthal dependence, the quantum mechanical analysis of thehydrogenmolecularionleadstotheequation d dηbracketleftbiggparenleftbig 1−η2parenrightbigdu dηbracketrightbigg +αu+βη2u=0. Develop a power-series solution for u(η). Evaluate the first three nonvanishing coeffi- cientsinterms of a0. Indicialequation k(k−1)=0, uk=1=a0ηbraceleftbigg 1+2−α 6η2+bracketleftbigg(2−α)(12−α) 120−β 20bracketrightbigg η4+···bracerightbigg . 9.5.14 To a good approximation, the interaction of two nucleons may be described by a mesonicpotential V=Ae−ax x, attractive for Anegative. Develop a series solution of the resultant Schrödinger wave equation ¯h2 2md2ψ dx2+(E−V)ψ=0 throughthefirstthreenonvanishingcoefficients. ψ=a0braceleftbig x+1 2A′x2+1 6bracketleftbig1 2A′2−E′−aA′bracketrightbig x3+···bracerightbig , wheretheprimeindicatesmultiplicationby 2 m/¯h2. 9.5.15 Nearthenucleusofacomplexatomthepotentialenergyof oneelectronis givenby V=Ze2 rparenleftbig 1+b1r+b2r2parenrightbig , wherethecoefficients b1andb2arisefromscreeningeffects.Forthecaseofzeroangu- larmomentumshowthatthefirstthreetermsofthesolutionoftheSchrödingerequation have the same form as those of Exercise 9.5.14. By appropriate translation of coeffi- cients or parameters, write out the first three terms in a series expansion of the wave function. 578 Chapter 9 Differential Equations 9.5.16 If theparameter a2inEq.(9.110d) isequalto2, Eq.(9.110d)becomes y′′+1 x2y′−2 x2y=0. From the indicial equation and the recurrence relation derivea solution y=1+2x+ 2x2. Verify that this is indeed a solution by substituting back into the differential equa- tion. 9.5.17 ThemodifiedBessel function I0(x)satisfiesthedifferentialequation x2d2 dx2I0(x)+xd dxI0(x)−x2I0(x)=0. FromExercise7.3.4 theleadingterminanasymptoticexpansionisfoundtobe I0(x)∼ex √ 2πx. Assumeaseriesof theform I0(x)∼ex √ 2πxbraceleftbig 1+b1x−1+b2x−2+···bracerightbig . Determinethecoefficients b1andb2. ANS.b1=1 8,b2=9 128. 9.5.18 Theevenpower-seriessolutionofLegendre’sequationisgivenbyExercise9.5.5.Take a0=1 andnnot an even integer, say n=0.5. Calculate the partial sums of the series throughx200,x400,x600,...,x2000forx=0.95(0.01)1.00. Also, write out the individ- ualtermcorrespondingtoeachofthesepowers. Note. This calculation does notconstitute proof of convergence at x=0.99 or diver- genceatx=1.00,butperhapsyoucanseethedifferenceinthebehaviorofthesequence ofpartialsumsfor thesetwovaluesof x. 9.5.19 (a) The odd power-series solution of Hermite’s equation is given by Exercise 9.5.6. Takea0=1. Evaluate this series for α=0,x=1,2,3. Cut off your calculation after the last term calculated has dropped below the maximum term by a factor of 106ormore.Setanupperboundtotheerrormadeinignoringtheremainingterms intheinfiniteseries. (b) Asacheckonthecalculationofpart(a),showthattheHermiteseries yodd(α=0) correspondstointegraltextx 0exp(x2)dx. (c) Calculatethisintegralfor x=1,2,3. 9.6 A S ECOND SOLUTION In Section 9.5 a solution of a second-order homogeneous ODE was developed by substi- tuting in a power series. By Fuchs’ theorem this is possible, provided the power series is anexpansionaboutanordinarypointoranonessentialsingularity.15Thereisnoguarantee 15This is whythe classification ofsingularities in Section9.4is of vital importance. 9.6 A Second Solution 579 thatthisapproachwillyieldthetwoindependentsolutionsweexpectfromalinearsecond- orderODE.Infact,weshallprovethatsuchanODEhasatmosttwolinearlyindependent solutions.Indeed,thetechniquegaveonlyonesolutionforBessel’sequation( naninteger). In this section we also develop two methods of obtaining a second independent solution: an integral method and a power series containing a logarithmic term. First, however, we considerthequestionofindependenceofaset offunctions. Linear Independence of Solutions Givenasetoffunctions ϕλ,thecriterionforlineardependenceistheexistenceofarelation oftheformsummationdisplay λkλϕλ=0, (9.111) in which not all the coefficients kλare zero. On the other hand, if the only solution of Eq.(9.111) is kλ=0 for allλ,thesetof functions ϕλis saidtobelinearly independent . It may be helpful to think of linear dependence of vectors. Consider A,B, andCin three-dimensionalspace,with A·B×C/negationslash=0.Thennonontrivialrelationoftheform aA+bB+cC=0 (9.112) exists.A,B, andCare linearly independent.On the other hand, any fourth vector, D,m a y beexpressedasalinearcombinationof A,B,andC(seeSection3.1).Wecanalwayswrite anequationof theform D−aA−bB−cC=0, (9.113) and the four vectors are notlinearly independent. The three noncoplanar vectors A,B, andCspanourrealthree-dimensionalspace. If a set of vectors or functions are mutually orthogonal, then they are automatically lin- early independent. Orthogonality implies linear independence. This can easily be demon- stratedbytakinginnerproducts(scalarordotproductforvectors,orthogonalityintegralof Section10.2for functions). Let us assume that the functions ϕλare differentiable as needed. Then, differentiating Eq.(9.111) repeatedly,wegenerateaset ofequations summationdisplay λkλϕ′ λ=0, (9.114) summationdisplay λkλϕ′′ λ=0, (9.115) and so on. This gives us a set of homogeneous linear equations in which kλare the un- known quantities. By Section 3.1 there is a solution kλ/negationslash=0 only if the determinant of the coefficientsofthe kλ’vanishes.This means vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleϕ1ϕ2···ϕn ϕ′ 1ϕ′ 2···ϕ′ n ··· ··· ··· ··· ϕ(n−1) 1ϕ(n−1) 2···ϕ(n−1) nvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (9.116) Thisdeterminantis calledthe Wronskian . 580 Chapter 9 Differential Equations 1. If the Wronskian is not equal to zero, then Eq. (9.111) has no solution other than kλ=0.Theset offunctions ϕλisthereforelinearlyindependent. 2. IftheWronskianvanishesatisolatedvaluesoftheargument,thisdoesnotnecessarily provelineardependence(unlessthesetoffunctionshasonlytwofunctions).However, if the Wronskian is zero over the entire range of the variable, the functions ϕλare linearly dependent over this range16(compare Exercise 9.5.2 for the simple case of twofunctions). Example 9.6.1 LINEAR INDEPENDENCE The solutions of the linear oscillator equation (9.84) are ϕ1=sinωx,ϕ2=cosωx.T h e Wronskianbecomes vextendsinglevextendsinglevextendsinglevextendsinglesinωxcosωx ωcosωx−ωsinωxvextendsinglevextendsinglevextendsinglevextendsingle=−ω/negationslash=0. These two solutions, ϕ1andϕ2, are therefore linearly independent. For just two functions thismeansthatoneisnotamultipleoftheother,whichisobviouslytrueinthiscase. Youknowthat sinωx=±parenleftbig 1−cos2ωxparenrightbig1/2, but this is notalinearrelation,oftheformof Eq.(9.111). /squaresolid Examples 9.6.2 LINEAR DEPENDENCE Foranillustrationoflineardependence,considerthesolutionsoftheone-dimensionaldif- fusion equation. We have ϕ1=exandϕ2=e−x, and we add ϕ3=coshx, also a solution. TheWronskianis vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingleexe−xcoshx ex−e−xsinhx exe−xcoshxvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. The determinant vanishes for all xbecause the first and third rows are identical. Hence ex,e−x, and cosh xare linearly dependent, and, indeed, we have a relation of the form of Eq.(9.111): ex+e−x−2coshx=0 with kλ/negationslash=0. /squaresolid Now we are ready to prove the theorem that a second-order homogeneous ODE has twolinearlyindependentsolutions . Suppose y1,y2,y3are three solutions of the homogeneous ODE (9.80). Then we form the Wronskian Wjk=yjy′ k−y′ jykof any pair yj,ykof them and recall that 16Compare H. Lass, Elements of Pure and Applied Mathematics , New York: McGraw-Hill (1957), p. 187, for proof of this assertion. It is assumed that the functions have continuous derivatives and that at least one of the minors of the bottom row of Eq.(9.116) (Laplaceexpansion) does not vanish in [a,b], the interval under consideration. 9.6 A Second Solution 581 W′ jk=yjy′′ k−y′′ jyk.We divide each ODE by y, getting−Qon their right-hand side, so y′′ j yj+Py′ j yj=−Q(x)=y′′ k yk+Py′ k yk. Multiplyingby yjyk, wefind (yjy′′ k−y′′ jyk)+P(yjy′ k−y′ jyk)=0,orW′ jk=−PWjk(9.117) foranypairofsolutions.FinallyweevaluatetheWronskianofallthreesolutions,expand- ingit alongthesecondrowandusingtheODEsfor the Wjk: W=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingley1y2y3 y′ 1y′ 2y′ 3 y′′ 1y′′ 2y′′ 3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=−y′ 1W′ 23+y′ 2W′ 13−y′ 3W′ 12 =P(y′ 1W23−y′ 2W13+y′ 3W12)=−Pvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingley1y2y3 y′ 1y′ 2y′ 3 y′ 1y′ 2y′ 3vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. The vanishing Wronskian, W=0, because of two identical rows, is just the condition for linear dependence of the solutions yj.Thus, there are at most two linearly independent solutions of the homogeneous ODE. Similarly one can prove that a linear homogeneous nth-order ODE has nlinearly independent solutions yj, so the general solution y(x)=summationtextcjyj(x)isalinearcombinationofthem. A Second Solution Returningtoourlinear,second-order,homogeneousODEof thegeneralform y′′+P(x)y′+Q(x)y=0, (9.118) lety1andy2betwoindependentsolutions.ThentheWronskian,bydefinition,is W=y1y′ 2−y′ 1y2. (9.119) BydifferentiatingtheWronskian,weobtain W′=y′ 1y′ 2+y1y′′ 2−y′′ 1y2−y′ 1y′ 2 =y1bracketleftbig −P(x)y′ 2−Q(x)y2bracketrightbig −y2bracketleftbig −P(x)y′ 1−Q(x)y1bracketrightbig =−P(x)(y 1y′ 2−y′ 1y2). Theexpressioninparenthesesis just W, theWronskian, andwehave W′=−P(x)W. (9.120) Inthespecialcasethat P(x)=0,thatis, y′′+Q(x)y=0, (9.121) theWronskian W=y1y′ 2−y′ 1y2=constant. (9.122) 582 Chapter 9 Differential Equations Sinceouroriginaldifferentialequationishomogeneous,wemaymultiplythesolutions y1 andy2by whatever constants we wish and arrange to have the Wronskian equal to unity (or−1).Thiscase, P(x)=0,appearsmorefrequentlythanmightbeexpected.Recallthat the portion of ∇2(ψ r)in spherical polar coordinates involving radial derivatives contains no first radial derivative. Finally, every linear second-order differential equation can be transformedintoanequationof theformofEq. (9.121)(compareExercise9.6.11). For the general case, let us now assume that we have one solution of Eq. (9.118) by a series substitution (or by guessing). We now proceed to develop a second, independent solutionfor which W/negationslash=0.RewritingEq.(9.120) as dW W=−Pdx, weintegrateoverthevariable x,fromatox,toobtain lnW(x) W(a)=−integraldisplayx aP(x1)dx1, or17 W(x)=W(a)expbracketleftbigg −integraldisplayx aP(x1)dx1bracketrightbigg . (9.123) But W(x)=y1y′ 2−y′ 1y2=y2 1d dxparenleftbiggy2 y1parenrightbigg . (9.124) BycombiningEqs. (9.123)and(9.124), wehave d dxparenleftbiggy2 y1parenrightbigg =W(a)exp[−integraltextx aP(x1)dx1] y2 1. (9.125) Finally,byintegratingEq. (9.125)from x2=btox2=xweget y2(x)=y1(x)W(a)integraldisplayx bexp[−integraltextx2 aP(x1)dx1] [y1(x2)]2dx2. (9.126) Hereaandbare arbitrary constants and a term y1(x)y2(b)/y1(b)has been dropped, for it leadstonothingnew.Since W(a),theWronskianevaluatedat x=a,isaconstantandour solutionsforthehomogeneousdifferentialequationalwayscontainanunknownnormaliz- ingfactor, weset W(a)=1 andwrite y2(x)=y1(x)integraldisplayxexp[−integraltextx2P(x1)dx1] [y1(x2)]2dx2. (9.127) Note that the lower limits x1=aandx2=bhave been omitted. If they are retained, they simply make a contribution equal to a constant times the known first solution, y1(x), and 17IfP(x)remains finite in the domain of interest, W(x)/negationslash=0 unlessW(a)=0. That is, the Wronskian of our two solutions is either identically zero or never zero. However, if P(x)does not remain finite in our interval, then W(x)can have isolated zeros in that domain and one must be careful to choose aso thatW(a)/negationslash=0. 9.6 A Second Solution 583 hence add nothing new. If we have the important special case of P(x)=0, Eq. (9.127) reducesto y2(x)=y1(x)integraldisplayxdx2 [y1(x2)]2. (9.128) This means that by using either Eq. (9.127) or Eq. (9.128) we can take one known solu- tion and by integrating can generate a second, independent solution of Eq. (9.118). This techniqueis used in Section 12.10 to generate a second solution of Legendre’s differential equation. Example 9.6.3 ASECOND SOLUTION FOR THE LINEAR OSCILLATOR EQUATION Fromd2y/dx2+y=0 withP(x)=0 let one solution be y1=sinx. By applying Eq.(9.128), weobtain y2(x)=sinxintegraldisplayxdx2 sin2x2=sinx(−cotx)=−cosx, whichisclearlyindependent(nota linearmultiple)of sin x. /squaresolid Series Form of the Second Solution Further insight into the nature of the second solution of our differential equation may be obtainedbythefollowingsequenceofoperations. 1. Express P(x)andQ(x)inEq.(9.118) as P(x)=∞summationdisplay i=−1pixi,Q(x)=∞summationdisplay j=−2qjxj. (9.129) The lower limits of the summations are selected to create the strongest possible reg- ularsingularity (at the origin). These conditions just satisfy Fuchs’ theorem and thus helpusgainabetterunderstandingofFuchs’theorem. 2. Developthefirst fewtermsof apower-seriessolution,asinSection9.5. 3. Using this solution as y1, obtain a second series type solution, y2, with Eq. (9.127), integratingtermbyterm. ProceedingwithStep1,wehave y′′+parenleftbig p−1x−1+p0+p1x+···parenrightbig y′+parenleftbig q−2x−2+q−1x−1+···parenrightbig y=0,(9.130) inwhichpoint x=0isatworstaregularsingularpoint.If p−1=q−1=q−2=0,itreduces toanordinarypoint.Substituting y=∞summationdisplay λ=0aλxk+λ 584 Chapter 9 Differential Equations (Step2), weobtain ∞summationdisplay λ=0(k+λ)(k+λ−1)aλxk+λ−2+∞summationdisplay i=−1pixi∞summationdisplay λ=0(k+λ)aλxk+λ−1 +∞summationdisplay j=−2qjxj∞summationdisplay λ=0aλxk+λ=0. (9.131) Assumingthat p−1/negationslash=0,q−2/negationslash=0,our indicialequationis k(k−1)+p−1k+q−2=0, whichsetsthenetcoefficientof xk−2equaltozero.Thisreducesto k2+(p−1−1)k+q−2=0. (9.132) We denote the two roots of this indicial equation by k=αandk=α−n, wherenis zero or a positive integer. (If nis not an integer, we expect two independent series solutions by themethodsofSection9.5andwearedone.)Then (k−α)(k−α+n)=0, (9.133) or k2+(n−2α)k+α(α−n)=0, andequatingcoefficientsof kinEqs. (9.132)and(9.133),wehave p−1−1=n−2α. (9.134) Theknownseriessolutioncorrespondingtothelargerroot k=αmaybewrittenas y1=xα∞summationdisplay λ=0aλxλ. Substitutingthisseries solutionintoEq.(9.127) (Step3), wearefacedwith y2(x)=y1(x)integraldisplayxexp(−integraltextx2 asummationtext∞ i=−1pixi 1dx1) x2α 2(summationtext∞ λ=0aλxλ 2)2dx2, (9.135) where the solutions y1andy2have been normalized so that the Wronskian W(a)=1. Tacklingtheexponentialfactorfirst, wehave integraldisplayx2 a∞summationdisplay i=−1pixi 1dx1=p−1lnx2+∞summationdisplay k=0pk k+1xk+1 2+f(a) (9.136) 9.6 A Second Solution 585 withf(a)anintegrationconstantthatmaydependon a.Hence, expparenleftbigg −integraldisplayx2 asummationdisplay ipixi 1dx1parenrightbigg =expbracketleftbig −f(a)bracketrightbig x−p−1 2expparenleftbigg −∞summationdisplay k=0pk k+1xk+1 2parenrightbigg =expbracketleftbig −f(a)bracketrightbig x−p−1 2bracketleftbigg 1−∞summationdisplay k=0pk k+1xk+1 2+1 2!parenleftbigg −∞summationdisplay k=0pk k+1xk+1 2parenrightbigg2 +···bracketrightbigg . (9.137) Thisfinalseriesexpansionoftheexponentialiscertainlyconvergentiftheoriginalexpan- sionofthecoefficient P(x)was uniformlyconvergent. ThedenominatorinEq. (9.135)maybehandledbywriting bracketleftbigg x2α 2parenleftbigg∞summationdisplay λ=0aλxλ 2parenrightbigg2bracketrightbigg−1 =x−2α 2parenleftbigg∞summationdisplay λ=0aλxλ 2parenrightbigg−2 =x−2α 2∞summationdisplay λ=0bλxλ 2.(9.138) Neglecting constant factors, which will be picked up anyway by the requirement that W(a)=1,weobtain y2(x)=y1(x)integraldisplayx x−p−1−2α 2parenleftbigg∞summationdisplay λ=0cλxλ 2parenrightbigg dx2. (9.139) ByEq. (9.134), x−p−1−2α 2=x−n−1 2, (9.140) andwehaveassumedherethat nisaninteger.SubstitutingthisresultintoEq.(9.139), we obtain y2(x)=y1(x)integraldisplayxparenleftbig c0x−n−1 2+c1x−n 2+c2x−n+1 2+···+cnx−1 2+···parenrightbig dx2.(9.141) The integration indicated in Eq. (9.141) leads to a coefficient of y1(x)consisting of two parts: 1. Apowerseries startingwith x−n. 2. Alogarithmtermfromtheintegrationof x−1(whenλ=n).Thistermalwaysappears whennisaninteger, unlesscnfortuitouslyhappenstovanish.18 18For parity considerations, ln xis takentobe ln |x|,e v en. 586 Chapter 9 Differential Equations Example 9.6.4 ASECOND SOLUTION OF BESSEL ’SEQUATION FromBessel’s equation,Eq. (9.100)(dividedby x2toagreewithEq.(9.118)), wehave P(x)=x−1Q(x)=1 forthecase n=0. Hencep−1=1,q0=1;allother piandqjvanish.TheBesselindicialequationis k2=0 (Eq. (9.103)with n=0).HenceweverifyEqs. (9.132)to(9.134)with nandα=0. Our first solution is available from Eq. (9.108). Relabeling it to agree with Chapter 11 (andusing a0=1),weobtain19 y1(x)=J0(x)=1−x2 4+x4 64−Oparenleftbig x6parenrightbig . (9.142a) Now, substituting all this into Eq. (9.127), we have the specific case corresponding to Eq.(9.135): y2(x)=J0(x)integraldisplayxexp[−integraltextx2x−1 1dx1] [1−x2 2/4+x4 2/64−···]2dx2. (9.142b) Fromthenumeratorof theintegrand, expbracketleftbigg −integraldisplayx2dx1 x1bracketrightbigg =exp[−lnx2]=1 x2. Thiscorrespondstothe x−p−1 2inEq.(9.137).Fromthedenominatoroftheintegrand,using abinomialexpansion,weobtain bracketleftbigg 1−x2 2 4+x4 2 64bracketrightbigg−2 =1+x2 2 2+5x4 2 32+···. CorrespondingtoEq. (9.139),wehave y2(x)=J0(x)integraldisplayx1 x2bracketleftbigg 1+x2 2 2+5x4 2 32+···bracketrightbigg dx2 =J0(x)braceleftbigg lnx+x2 4+5x4 128+···bracerightbigg . (9.142c) Letuscheckthisresult.FromEqs.(11.62)and(11.64),whichgivethestandardformof thesecondsolution(higher-ordertermsareneeded) N0(x)=2 π[lnx−ln2+γ]J0(x)+2 πbraceleftbiggx2 4−3x4 128+···bracerightbigg . (9.142d) Two points arise: (1) Since Bessel’s equation is homogeneous, we may multiply y2(x)by any constant. To match N0(x), we multiply our y2(x)by 2/π. (2) To our second solution, 19Thecapital O(order of) as written here means terms proportional to x6and possibly higher powers of x. 9.6 A Second Solution 587 (2/π)y2(x),wemayaddanyconstantmultipleofthefirstsolution.Again,tomatch N0(x) weadd 2 π[−ln2+γ]J0(x), whereγistheusualEuler–Mascheroniconstant(Section5.2).20Ournew,modifiedsecond solutionis y2(x)=2 π[lnx−ln2+γ]J0(x)+2 πJ0(x)braceleftbiggx2 4+5x4 128+···bracerightbigg . (9.142e) Now the comparison with N0(x)becomes a simple multiplication of J0(x)from Eq. (9.142a) and the curly bracket of Eq. (9.142c). The multiplication checks, through terms of order x2andx4, which is all we carried. Our second solution from Eqs. (9.127) and(9.135) agreeswiththestandardsecondsolution,theNeumannfunction, N0(x). From the preceding analysis, the second solution of Eq. (9.118), y2(x), may be written as y2(x)=y1(x)lnx+∞summationdisplay j=−ndjxj+α, (9.142f) the first solution times ln xand another power series, this one starting with xα−n, which means that we may look for a logarithmic term when the indicial equation of Sec- tion 9.5 gives only one series solution. With the form of the second solution specified by Eq. (9.142f), we can substitute Eq. (9.142f) into the original differential equation and determine the coefficients djexactly as in Section 9.5. It may be worth noting that no se- ries expansion of ln xis needed. In the substitution, ln xwill drop out; its derivatives will survive. /squaresolid The second solution will usually diverge at the origin because of the logarithmic factor and the negative powers of xin the series. For this reason y2(x)is often referred to as theirregular solution . The first series solution, y1(x), which usually converges at the origin,iscalledthe regularsolution .Thequestionofbehaviorattheoriginisdiscussedin more detail in Chapters 11 and 12, in which we take up Bessel functions, modified Bessel functions,andLegendrefunctions. Summary The two solutions of both sections (together with the exercises) provide a complete solu- tionof our linear, homogeneous, second-order ODE—assuming that the point of expan- sion is no worse than a regular singularity. At least one solution can always be obtained byseriessubstitution(Section9.5).A second,linearlyindependentsolution canbecon- structed by the Wronskian double integral, Eq. (9.127). This is all there are: No third, linearlyindependentsolutionexists (compareExercise9.6.10). Thenonhomogeneous , linear, second-order ODE will have an additional solution: the particular solution . This particular solution may be obtained by the method of variation ofparameters,Exercise9.6.25,orbytechniquessuchasGreen’sfunction,Section9.7. 20TheNeumann function N0is definedas it is in order to achieveconvenient asymptotic properties, Sections 11.3 and11.6. 588 Chapter 9 Differential Equations Exercises 9.6.1 Youknowthatthethreeunitvectors ˆx,ˆy,andˆzaremutuallyperpendicular(orthogonal). Showthatˆx,ˆy,andˆzarelinearlyindependent.Specifically,showthatnorelationofthe form of Eq.(9.111) existsfor ˆx,ˆy,andˆz. 9.6.2 Thecriterionforthelinear independence ofthreevectors A,B,andCisthattheequa- tion aA+bB+cC=0 (analogous to Eq. (9.111)) has no solution other than the trivial a=b=c=0. Using components A=(A1,A2,A3), and so on, set up the determinant criterion for the exis- tenceor nonexistenceof a nontrivialsolutionfor the coefficients a,b, andc. Show that yourcriterionis equivalenttothetriplescalarproduct A·B×C/negationslash=0. 9.6.3 UsingtheWronskiandeterminant,showthatthesetof functions braceleftbigg 1,xn n!(n=1,2,...,N)bracerightbigg islinearlyindependent. 9.6.4 If the Wronskian of two functions y1andy2is identicallyzero, show by direct integra- tionthat y1=cy2, thatis, that y1andy2are dependent.Assumethefunctionshavecontinuousderivatives andthatatleastoneofthefunctionsdoesnotvanishintheintervalunderconsideration. 9.6.5 TheWronskianoftwofunctionsisfoundtobezeroat x0−ε≤x≤x0+εforarbitrarily smallε>0.Show that this Wronskian vanishes for all xand that the functions are linearlydependent. 9.6.6 Thethreefunctions sin x,ex,ande−xarelinearlyindependent.Noonefunctioncanbe written as a linear combination of the other two. Show that the Wronskian of sin x,ex, ande−xvanishesbutonlyatisolatedpoints. ANS.W=4sinx, W=0f orx=±nπ, n=0,1,2,.... 9.6.7 Consider two functions ϕ1=xandϕ2=|x|=xsgnx(Fig. 9.3). The function sgn xis the sign of x. Sinceϕ′ 1=1 andϕ′ 2=sgnx,W(ϕ1,ϕ2)=0 for any interval, including [−1,+1].DoesthevanishingoftheWronskianover [−1,+1]provethat ϕ1andϕ2are linearlydependent?Clearly,theyare not.Whatiswrong? 9.6.8 Explainthat linearindependence doesnotmeantheabsenceofanydependence.Illus- trateyourargumentwith cosh xandex. 9.6.9 Legendre’sdifferentialequation parenleftbig 1−x2parenrightbig y′′−2xy′+n(n+1)y=0 9.6 A Second Solution 589 FIGURE 9.3xand|x|. has a regularsolution Pn(x)andan irregular solution Qn(x). Showthat theWronskian ofPnandQnis givenby Pn(x)Q′ n(x)−P′ n(x)Qn(x)=An 1−x2, withAnindependent ofx. 9.6.10 Show, by means of the Wronskian, that a linear, second-order, homogeneous ODE of theform y′′(x)+P(x)y′(x)+Q(x)y(x)=0 cannot have three independent solutions . (Assume a third solution and show that the Wronskianvanishesforall x.) 9.6.11 Transformourlinear,second-orderODE y′′+P(x)y′+Q(x)y=0 bythesubstitution y=zexpbracketleftbigg −1 2integraldisplayx P(t)dtbracketrightbigg andshowthattheresultingdifferentialequationfor zis z′′+q(x)z=0, where q(x)=Q(x)−1 2P′(x)−1 4P2(x). Note.This substitutioncanbederivedbythetechniqueofExercise9.6.24. 9.6.12 UsetheresultofExercise9.6.11toshowthatthereplacementof ϕ(r)byrϕ(r)maybe expected to eliminate the first derivative from the Laplacian in spherical polar coordi- nates.SeealsoExercise2.5.18(b). 590 Chapter 9 Differential Equations 9.6.13 Bydirectdifferentiationandsubstitutionshowthat y2(x)=y1(x)integraldisplayxexp[−integraltextsP(t)dt] [y1(s)]2ds satisfies(like y1(x))theODE y′′ 2(x)+P(x)y′ 2(x)+Q(x)y2(x)=0. Note.TheLeibnizformulafor thederivativeof anintegralis d dαintegraldisplayh(α) g(α)f(x,α)dx=integraldisplayh(α) g(α)∂f(x,α) ∂αdx+fbracketleftbig h(α),αbracketrightbigdh(α) dα−fbracketleftbig g(α),αbracketrightbigdg(α) dα. 9.6.14 Intheequation y2(x)=y1(x)integraldisplayxexp[−integraltextsP(t)dt] [y1(s)]2ds y1(x)satisfies y′′ 1+P(x)y′ 1+Q(x)y1=0. The function y2(x)is a linearly independent second solution of the same equation. Show that the inclusion of lower limits on the two integrals leads to nothing new, that is,thatitgeneratesonlyanoverallconstantfactorandaconstantmultipleoftheknown solutiony1(x). 9.6.15 Giventhatonesolutionof R′′+1 rR′−m2 r2R=0 isR=rm,showthatEq. (9.127)predictsasecondsolution, R=r−m. 9.6.16 Usingy1(x)=summationtext∞ n=0(−1)nx2n+1/(2n+1)!as a solution of the linear oscillator equa- tion, follow the analysis culminating in Eq. (9.142f) and show that c1=0 so that the secondsolutiondoesnot,inthis case,containalogarithmicterm. 9.6.17 Show that when nisnotan integer in Bessel’s ODE, Eq. (9.100), the second solution ofBessel’s equation,obtainedfrom Eq.(9.127), does notcontainalogarithmicterm. 9.6.18 (a) Onesolutionof Hermite’sdifferentialequation y′′−2xy′+2αy=0 forα=0isy1(x)=1.Findasecondsolution, y2(x),usingEq.(9.127).Showthat yoursecondsolutionis equivalentto yodd(Exercise 9.5.6). (b) Find a second solution for α=1, where y1(x)=x, using Eq. (9.127). Show that yoursecondsolutionis equivalentto yeven(Exercise 9.5.6). 9.6.19 OnesolutionofLaguerre’sdifferentialequation xy′′+(1−x)y′+ny=0 forn=0i sy1(x)=1. Using Eq. (9.127), developa second, linearly independentsolu- tion.Exhibitthelogarithmictermexplicitly. 9.6 A Second Solution 591 9.6.20 ForLaguerre’sequationwith n=0, y2(x)=integraldisplayxes sds. (a) Write y2(x)as alogarithmplusapowerseries. (b) Verifythattheintegralformof y2(x),previouslygiven,isasolutionofLaguerre’s equation (n=0)by direct differentiation of the integral and substitution into the differentialequation. (c) Verify that the series form of y2(x), part (a), is a solution by differentiating the seriesandsubstitutingbackintoLaguerre’sequation. 9.6.21 OnesolutionoftheChebyshevequation parenleftbig 1−x2parenrightbig y′′−xy′+n2y=0 forn=0i sy1=1. (a) UsingEq. (9.127), developasecond,linearlyindependentsolution. (b) FindasecondsolutionbydirectintegrationoftheChebyshevequation. Hint.L e tv=y′and integrate. Compare your result with the second solution given in Section13.3. ANS.(a) y2=sin−1x. (b)Thesecondsolution, Vn(x), is notdefinedfor n=0. 9.6.22 OnesolutionoftheChebyshevequation parenleftbig 1−x2parenrightbig y′′−xy′+n2y=0 forn=1i sy1(x)=x. Set up the Wronskian double integral solution and derive a secondsolution, y2(x). ANS.y2=−parenleftbig 1−x2parenrightbig1/2. 9.6.23 TheradialSchrödingerwaveequationhastheform braceleftbigg −¯h2 2md2 dr2+l(l+1)¯h2 2mr2+V(r)bracketrightbigg y(r)=Ey(r). Thepotentialenergy V(r)maybeexpandedabouttheoriginas V(r)=b−1 r+b0+b1r+···. (a) Showthatthereis one(regular)solutionstartingwith rl+1. (b) FromEq. (9.128)showthattheirregularsolutiondivergesattheoriginas r−l. 9.6.24 Show that if a second solution, y2,i sa s s u m e dt oh a v et h ef o r m y2(x)=y1(x)f(x), substitutionbackintotheoriginalequation y′′ 2+P(x)y′ 2+Q(x)y2=0 592 Chapter 9 Differential Equations leadsto f(x)=integraldisplayxexp[−integraltextsP(t)dt] [y1(s)]2ds, inagreementwithEq. (9.127). 9.6.25 If our linear, second-order ODE is nonhomogeneous, that is, of the form of Eq. (9.82), themostgeneralsolutionis y(x)=y1(x)+y2(x)+yp(x). (y1andy2areindependentsolutionsofthehomogeneousequation.) Showthat yp(x)=y2(x)integraldisplayxy1(s)F(s)ds W{y1(s),y2(s)}−y1(x)integraldisplayxy2(s)F(s)ds W{y1(s),y2(s)}, withW{y1(x),y2(x)}theWronskianof y1(s)andy2(s). Hint.AsinExercise9.6.24,let yp(x)=y1(x)v(x)anddevelopafirst-orderdifferential equationfor v′(x). 9.6.26 (a) Showthat y′′+1−α2 4x2y=0 hastwosolutions: y1(x)=a0x(1+α)/2, y2(x)=a0x(1−α)/2. (b) Forα=0 the two linearly independentsolutions of part (a) reduceto y10=a0x1/2. UsingEq.(9.128)deriveasecondsolution, y20(x)=a0x1/2lnx. Verifythat y20is indeedasolution. (c)Showthatthesecondsolutionfrompart(b)maybeobtainedasalimitingcasefrom thetwosolutionsof part(a): y20(x)=lim α→0parenleftbiggy1−y2 αparenrightbigg . 9.7 N ONHOMOGENEOUS EQUATION —G REEN ’SFUNCTION The series substitution of Section 9.5 and the Wronskian double integral of Section 9.6 provide the most general solution of the homogeneous , linear, second-order ODE. The specific solution, yp, linearly dependent on the source term ( F(x)of Eq. (9.82)) may be crankedoutbythevariationofparametersmethod,Exercise9.6.25.Inthissectionweturn toadifferentmethodofsolution—Green’sfunction. For a brief introduction to Green’s function method, as applied to the solution of a non- homogeneous PDE, it is helpful to use the electrostatic analog. In the presence of charges 9.7 Nonhomogeneous Equation — Green’s Function 593 the electrostatic potential ψsatisfies Poisson’s nonhomogeneous equation (compare Sec- tion1.14), ∇2ψ=−ρ ε0(mksunits) , (9.143) andLaplace’shomogeneousequation, ∇2ψ=0, (9.144) intheabsenceofelectriccharge (ρ=0).Ifthechargesarepointcharges qi,weknowthat thesolutionis ψ=1 4πε0summationdisplay iqi ri, (9.145) asuperpositionofsingle-pointchargesolutionsobtainedfromCoulomb’slawfortheforce betweentwopointcharges q1andq2, F=q1q2ˆr 4πε0r2. (9.146) Byreplacementofthediscretepointchargeswithasmeared-outdistributedcharge,charge densityρ,Eq.(9.145) becomes ψ(r=0)=1 4πε0integraldisplayρ(r) rdτ (9.147) or,for thepotentialat r=r1awayfromtheoriginandthechargeat r=r2, ψ(r1)=1 4πε0integraldisplayρ(r2) |r1−r2|dτ2. (9.148) We useψas the potential corresponding to the given distribution of charge and there- fore satisfying Poisson’s equation (9.143), whereas a function G, which we label Green’s function, is required to satisfy Poisson’s equation with a point source at the point defined byr2: ∇2G=−δ(r1−r2). (9.149) Physically, then, Gis the potential at r1corresponding to a unit source at r2. By Green’s theorem(Section1.11,Eq. (11.104)) integraldisplayparenleftbig ψ∇2G−G∇2ψparenrightbig dτ2=integraldisplay (ψ∇G−G∇ψ)·dσ. (9.150) Assuming that the integrand falls off faster than r−2we may simplify our problem by takingthevolumeso largethatthesurfaceintegralvanishes,leaving integraldisplay ψ∇2Gdτ2=integraldisplay G∇2ψdτ2, (9.151) or,bysubstitutinginEqs. (9.143)and(9.149), wehave −integraldisplay ψ(r2)δ(r1−r2)dτ2=−integraldisplayG(r1,r2)ρ(r2) ε0dτ2. (9.152) 594 Chapter 9 Differential Equations IntegrationbyemployingthedefiningpropertyoftheDiracdeltafunction(Eq.(1.171b)) produces ψ(r1)=1 ε0integraldisplay G(r1,r2)ρ(r2)dτ2. (9.153) Note that we have used Eq. (9.149) to eliminate ∇2Gbut that the function Gitself is stillunknown.InSection1.14,Gausslaw,wefoundthat integraldisplay ∇2parenleftbigg1 rparenrightbigg dτ=braceleftbigg0, −4π,(9.154) 0 if thevolumedidnot includethe originand −4πif theorigin wereincluded.This result from Section1.14mayberewrittenasinEq. (1.170),or ∇2parenleftbigg1 4πrparenrightbigg =−δ(r),or ∇2parenleftbigg1 4πr12parenrightbigg =−δ(r1−r2), (9.155) corresponding to a shift of the electrostatic charge from the origin to the position r=r2. Herer12=|r1−r2|, and the Dirac delta function δ(r1−r2)vanishes unless r1=r2. Therefore in a comparison of Eqs. (9.149) and (9.155) the function G(Green’s function) isgivenby G(r1,r2)=1 4π|r1−r2|. (9.156) Thesolutionofour differentialequation(Poisson’sequation)is ψ(r1)=1 4πε0integraldisplayρ(r2) |r1−r2|dτ2, (9.157) in complete agreement with Eq. (9.148). Actually ψ(r1), Eq. (9.157), is the particular solution of Poisson’s equation. We may add solutions of Laplace’s equation (compare Eq.(9.83)). Suchsolutionscoulddescribeanexternalfield. These results will be generalized to the second-order, linear, but nonhomogeneous dif- ferentialequation Ly(r1)=−f(r1), (9.158) where Lis alineardifferentialoperator.TheGreen’sfunctionistakentobeasolutionof LG(r1,r2)=−δ(r1−r2), (9.159) analogous to Eq. (9.149). The Green’s function depends on boundary conditions that may nolongerbethoseofelectrostaticsinaregionofinfiniteextent.Thentheparticularsolution y(r1)becomes y(r1)=integraldisplay G(r1,r2)f(r2)dτ2. (9.160) 9.7 Nonhomogeneous Equation — Green’s Function 595 (Theremayalsobeanintegraloveraboundingsurface,dependingontheconditionsspec- ified.) In summary, Green’s function, often written G(r1,r2)as a reminder of the name, is a solution of Eq. (9.149)or Eq.(9.159)more generally. It enters in an integral solution of our differential equation, as in Eqs. (9.148)and(9.153). For the simple, but impor- tant, electrostatic case we obtain Green’s function, G(r1,r2), by Gauss’ law, comparing Eqs.(9.149)and(9.155). Finally, from the final solution (Eq.(9.157))it is possible to develop a physical interpretation of Green’s function. It occurs as a weighting function or propagator function that enhances or reduces the effect of the charge element ρ(r2)dτ2 according to its distance from the field point r1. Green’s function, G(r1,r2), gives the ef- fectofaunitpointsourceat r2inproducingapotentialat r1.Thisishowitwasintroduced inEq.(9.149);this ishowitappearsinEq. (9.157). Symmetry of Green’s Function AnimportantpropertyofGreen’sfunctionis thesymmetryof itstwovariables;thatis, G(r1,r2)=G(r2,r1). (9.161) Although this is obvious in the electrostatic case just considered, it can be proved under moregeneralconditions.Inplaceof Eq.(9.149), letusrequirethat G(r,r1)satisfy21 ∇·bracketleftbig p(r)∇G(r,r1)bracketrightbig +λq(r)G(r,r1)=−δ(r−r1), (9.162) corresponding to a mathematical point source at r=r1. Here the functions p(r)andq(r) are well-behaved but otherwise arbitrary functions of r. The Green’s function, G(r,r2), satisfies the same equation, but the subscript 1 is replaced by subscript 2. The Green’s functions, G(r,r1)andG(r,r2), have the same values over a given surface Sof some volume of finite or infinite extent, and their normal derivatives have the same values over thesurface S,or theseGreen’sfunctionsvanishon S(Dirichletboundaryconditions,Sec- tion9.1).22ThenG(r,r2)is asortof potentialat r, createdbyaunitpointsourceat r2. We multiply the equation for G(r,r1)byG(r,r2)and the equation for G(r,r2)by G(r,r1)andthensubtractthetwo: G(r,r2)∇·bracketleftbig p(r)∇G(r,r1)bracketrightbig −G(r,r1)∇·bracketleftbig p(r)∇G(r,r2)bracketrightbig =−G(r,r2)δ(r−r1)+G(r,r1)δ(r−r2). (9.163) ThefirstterminEq. (9.163), G(r,r2)∇·bracketleftbig p(r)∇G(r,r1)bracketrightbig , maybereplacedby ∇·bracketleftbig G(r,r2)p(r)∇G(r,r1)bracketrightbig −∇G(r,r2)·p(r)∇G(r,r1). 21Equation (9.162) is a three-dimensional, inhomogeneous version of the self-adjoint eigenvalue equation, Eq. (10.8). 22Any attempt to demand that the normal derivatives vanish at the surface (Neumann’s conditions, Section 9.1) leads to trouble with Gauss’ law. It is like demanding thatintegraltext E·dσ=0 when you know perfectly well that there is some electric charge inside the surface. 596 Chapter 9 Differential Equations A similar transformation is carried out on the second term. Then integrating over the vol- umewhosesurfaceis SandusingGreen’stheorem,weobtaina surfaceintegral: integraldisplay Sbracketleftbig G(r,r2)p(r)∇G(r,r1)−G(r,r1)p(r)∇G(r,r2)bracketrightbig ·dσ =−G(r1,r2)+G(r2,r1). (9.164) The terms on the right-hand side appear when we use the Dirac delta functions in Eq. (9.163) and carry out the volume integration. With the boundary conditions earlier imposedontheGreen’sfunction,thesurface integralvanishesand G(r1,r2)=G(r2,r1), (9.165) whichshowsthatGreen’sfunctionissymmetric.Iftheeigenfunctionsarecomplex,bound- ary conditions corresponding to Eqs. (10.19) to (10.20) are appropriate. Equation (9.165) becomes G(r1,r2)=G∗(r2,r1). (9.166) Note that this symmetry property holds for Green’s functions in every equation in the form of Eq. (9.162). In Chapter 10 we shall call equations in this form self-adjoint .T h e symmetry is the basis of various reciprocity theorems; the effect of a charge at r2on the potentialat r1is thesameas theeffectofachargeat r1onthepotentialat r2. This use of Green’s functions is a powerful technique for solving many of the more difficultproblemsof mathematicalphysics. Form of Green’s Functions Letusassumethat Lisaself-adjointdifferentialoperatorofthegeneralform23 L1=∇1·bracketleftbig p(r1)∇1bracketrightbig +q(r1). (9.167) Herethe subscript 1 on Lemphasizesthat Loperateson r1. Then, as a simplegeneraliza- tionof Green’stheorem,Eq. (1.104),wehave integraldisplay (vL2u−uL2v)dτ2=integraldisplay p(v∇2u−u∇2v)·dσ2, (9.168) inwhichallquantitieshave r2astheirargument.(ToverifyEq.(9.168),takethedivergence of the integrand of the surface integral.) We let u(r2)=y(r2)so that Eq. (9.158) applies andv(r2)=G(r1,r2)so that Eq. (9.159) applies. (Remember, G(r1,r2)=G(r2,r1).) SubstitutingintoGreen’stheoremweget integraldisplaybraceleftbig −G(r1,r2)f(r2)+y(r2)δ(r1−r2)bracerightbig dτ2 =integraldisplay p(r2)braceleftbig G(r1,r2)∇2y(r2)−y(r2)∇2G(r1,r2)bracerightbig ·dσ2.(9.169) 23L1may be in 1, 2, or3 dimensions (with appropriate interpretation of ∇1). 9.7 Nonhomogeneous Equation — Green’s Function 597 WhenweintegrateovertheDiracdeltafunction y(r1)=integraldisplay G(r1,r2)f(r2)dτ2 +integraldisplay p(r2)braceleftbig G(r1,r2)∇2y(r2)−y(r2)∇2G(r1,r2)bracerightbig ·dσ2,(9.170) our solution to Eq. (9.158) appears as a volume integral plus a surface integral. If yand Gbothsatisfy Dirichlet boundary conditions or if both satisfy Neumann boundary con- ditions, the surface integral vanishes and we regain Eq. (9.160). The volume integral is a weighted integral over the source term f(r2)with our Green’s function G(r1,r2)as the weightingfunction. Forthespecialcaseof p(r1)=1andq(r1)=0,Lis∇2,theLaplacian.Letusintegrate ∇2 1G(r1,r2)=−δ(r1−r2) (9.171) overasmallvolumeincludingthepointsource.Then integraldisplay ∇1·∇1G(r1,r2)dτ1=−integraldisplay δ(r1−r2)dτ1=−1. (9.172) The volumeintegralon the left maybe transformed by Gauss’ theorem,as in the develop- mentofGauss’ law—Section1.14.Wefindthatintegraldisplay ∇1G(r1,r2)·dσ1=−1. (9.173) This shows, incidentally, that it may not be possible to impose a Neumann boundary con- dition, that the normal derivative of the Green’s function, ∂G/∂n, vanishes over the entire surface. If weareinthree-dimensionalspace,Eq. (9.173)is satisfiedbytaking ∂ ∂r12G(r1,r2)=−1 4π·1 |r1−r2|2,r 12=|r1−r2|. (9.174) Theintegrationisoverthesurfaceofaspherecenteredat r2.TheintegralofEq.(9.174)is G(r1,r2)=1 4π·1 |r1−r2|, (9.175) inagreementwithSection1.14. If weareintwo-dimensionalspace,Eq. (9.173)is satisfiedbytaking ∂ ∂ρ12G(ρ1,ρ2)=1 2π·1 |ρ1−ρ2|, (9.176) withrbeingreplacedby ρ,ρ=(x2+y2)1/2,andtheintegrationbeingoverthecircumfer- enceofa circlecenteredon ρ2.Her eρ12=|ρ1−ρ2|. IntegratingEq.(9.176), weobtain G(ρ1,ρ2)=−1 2πln|ρ1−ρ2|. (9.177) ToG(ρ1,ρ2)(and toG(r1,r2)) we may add any multiple of the regular solution of the homogeneous(Laplace’s)equationasneededtosatisfy boundaryconditions. 598 Chapter 9 Differential Equations Table 9.5 Green’sFunctionsa Laplace Helmholtz Modified Helmholtz ∇2∇2+k2∇2−k2 One-dimensional space No solutioni 2kexp(ik|x1−x2|)1 2kexp(−k|x1−x2|) for(−∞,∞) Two-dimensional space −1 2πln|ρ1−ρ2|i 4H(1) 0(k|ρ1−ρ2|)1 2πK0(k|ρ1−ρ2|) Three-dimensional space1 4π·1 |r1−r2|exp(ik|r1−r2|) 4π|r1−r2|exp(−k|r1−r2|) 4π|r1−r2| aThese are the Green’s functions satisfying the boundary condition G(r1,r2)=0asr1→∞for the Laplace and modified Helmholtz operators. For the Helmholtz operator, G(r1,r2)corresponds to an outgoing wave. H(1) 0is the Hankel function of Section11.4. K0ishemodifiedBessel functionofSection11.5. ThebehavioroftheLaplaceoperatorGreen’sfunctioninthevicinityofthesourcepoint r1=r2shown by Eqs. (9.175) and (9.177) facilitates the identification of the Green’s functionsfor theothercases, suchastheHelmholtzandmodifiedHelmholtzequations. 1. Forr1/negationslash=r2,G(r1,r2)mustsatisfy the homogeneous differentialequation L1G(r1,r2)=0,r1/negationslash=r2. (9.178) 2. Asr1→r2(orρ1→ρ2), G(ρ1,ρ2)≈−1 2πln|ρ1−ρ2|,two-dimensionalspace, (9.179) G(r1,r2)≈1 4π·1 |r1−r2|,three-dimensionalspace. (9.180) The term±k2in the operator does not affect the behavior of Gnear the singular point r1=r2. For convenience, the Green’s functions for the Laplace, Helmholtz, and modified HelmholtzoperatorsarelistedinTable9.5. Spherical Polar Coordinate Expansion24 AsanalternatedeterminationoftheGreen’sfunctionoftheLaplaceoperator,letusassume asphericalharmonicexpansionof theform G(r1,r2)=∞summationdisplay l=0lsummationdisplay m=−lgl(r1,r2)Ym l(θ1,ϕ1)Ym∗ l(θ2,ϕ2), (9.181) where the summation index lis the same for the spherical harmonics, as a consequence of the symmetry of the Green’s function. We will now determine the radial functions 24This section is optional here and may be postponed to Chapter 12. 9.7 Nonhomogeneous Equation — Green’s Function 599 gl(r1,r2). FromExercises1.15.11and12.6.6, δ(r1−r2)=1 r2 1δ(r1−r2)δ(cosθ1−cosθ2)δ(ϕ1−ϕ2) =1 r2 1δ(r1−r2)∞summationdisplay l=0lsummationdisplay m=−lYm l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). (9.182) Substituting Eqs. (9.181) and (9.182) into the Green’s function differential equation, Eq. (9.171), and making use of the orthogonality of the spherical harmonics, we obtain aradialequation: r1d2 dr2 1bracketleftbig r1gl(r1,r2)bracketrightbig −l(l+1)gl(r1,r2)=−δ(r1−r2). (9.183) This is now a one-dimensional problem. The solutions25of the corresponding homoge- neous equation are rl 1andr−l−1 1. If we demand that glremain finite as r1→0 and vanish asr1→∞,thetechniqueof Section10.5leadsto gl(r1,r2)=1 2l+1  rl 1 rl+1 2,r1<r2, rl 2 rl+1 1,r1>r2,(9.184) or gl(r1,r2)=1 2l+1·rl < rl+1>. (9.185) HenceourGreen’sfunctionis G(r1,r2)=∞summationdisplay l=0lsummationdisplay m=−l1 2l+1rl < rl+1>Ym l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). (9.186) Sincewealreadyhave G(r1,r2)inclosedform, Eq. (9.175),wemaywrite 1 4π·1 |r1−r2|=∞summationdisplay l=0lsummationdisplay m=−l1 2l+1rl < rl+1>Ym l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). (9.187) One immediate use for this spherical harmonic expansion of the Green’s function is in the development of an electrostatic multipole expansion. The potential for an arbitrary chargedistributionis ψ(r1)=1 4πε0integraldisplayρ(r2) |r1−r2|dτ2 25Compare Table9.2. 600 Chapter 9 Differential Equations (whichis Eq.(9.148)). SubstitutingEq.(9.187), weget ψ(r1)=1 ε0∞summationdisplay l=0lsummationdisplay m=−lbraceleftbigg1 2l+1Ym l(θ1,ϕ1) rl+1 1 ·integraldisplay ρ(r2)Ym∗ l(θ2,ϕ2)rl 2dϕ2sinθ2dθ2r2 2dr2bracerightbigg ,forr1>r2. Thisisthe multipoleexpansion .Therelativeimportanceofthevarioustermsinthedouble sumdependsontheform ofthesource, ρ(r2). Legendre Polynomial Addition Theorem26 Fromthegeneratingexpressionfor Legendrepolynomials,Eq. (12.4a), 1 4π·1 |r1−r2|=1 4π∞summationdisplay l=0rl < rl+1>Pl(cosγ), (9.188) whereγis the angle included between vectors r1andr2, Fig. 9.4. Equating Eqs. (9.187) and(9.188), wehavetheLegendrepolynomialadditiontheorem: Pl(cosγ)=4π 2l+1lsummationdisplay m=−lYm l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). (9.189) FIGURE 9.4Sphericalpolarcoordinates. 26This section is optional here and may be postponed to Chapter 12. 9.7 Nonhomogeneous Equation — Green’s Function 601 It is instructive to compare this derivation with the relatively cumbersome derivation of Section12.8leadingtoEq. (12.177). Circular Cylindrical Coordinate Expansion27 Inanalogywiththeprecedingsphericalpolarcoordinateexpansion,wewrite δ(r1−r2)=1 ρ1δ(ρ1−ρ2)δ(ϕ1−ϕ2)δ(z1−z2) =1 ρ1δ(ρ1−ρ2)1 4π2∞summationdisplay m=−∞eim(ϕ1−ϕ2)integraldisplay∞ −∞eik(z1−z2)dk,(9.190) usingExercise12.6.5andEq.(1.193c)andtheCauchyprincipalvalue.Butwhyasumma- tion for the ϕ-dependence and an integration for the z-dependence? The requirement that the azimuthal dependence be single-valued quantizes m, hence the summation. No such restrictionappliesto k. Toavoidproblemslaterwithnegativevaluesof k,werewriteEq. (9.190)as δ(r1−r2)=1 ρ1δ(ρ1−ρ2)1 2π∞summationdisplay m=−∞eim(ϕ1−ϕ2)1 πintegraldisplay∞ 0cosk(z1−z2)dk.(9.191) Weassumeasimilarexpansionof theGreen’sfunction, G(r1,r2)=1 2π2∞summationdisplay m=−∞gm(ρ1,ρ2)eim(ϕ1−ϕ2)integraldisplay∞ 0cosk(z1−z2)dk, (9.192) with the ρ-dependent coefficients gm(ρ1,ρ2)to be determined. Substituting into Eq.(9.171), nowincircularcylindricalcoordinates,wefindthatif g(ρ1,ρ2)satisfies d dρ1bracketleftbigg ρ1dgm dρ1bracketrightbigg −bracketleftbigg k2ρ1+m2 ρ1bracketrightbigg gm=−δ(ρ1−ρ2), (9.193) thenEq. (9.171)is satisfied. The operator in Eq. (9.193) is identified as the modified Bessel operator (in self- adjoint form). Hence the solutions of the corresponding homogeneous equation are u1= Im(kρ),u2=Km(kρ). As in the spherical polar coordinate case, we demand that Gbe finiteatρ1=0 andvanishas ρ1→∞.ThenthetechniqueofSection10.5yields gm(ρ1,ρ2)=−1 AIm(kρ<)Km(kρ>). (9.194) This corresponds to Eq. (9.155). The constant Acomes from the Wronskian (see Eq.(9.120)): Im(kρ)K′ m(kρ)−I′ m(kρ)Km(kρ)=A P(kρ). (9.195) 27This section is optional here and may be postponed to Chapter 11. 602 Chapter 9 Differential Equations FromExercise11.5.10, A=−1 and gm(ρ1,ρ2)=Im(kρ<)Km(kρ>). (9.196) ThereforeourcircularcylindricalcoordinateGreen’sfunctionis G(r1,r2)=1 4π·1 |r1−r2| =1 2π2∞summationdisplay m=−∞integraldisplay∞ 0Im(kρ<)Km(kρ>)eim(ϕ1−ϕ2)cosk(z1−z2)dk. (9.197) Exercise9.7.14is aspecialcaseof thisresult. Example 9.7.1 QUANTUM MECHANICAL SCATTERING —N EUMANN SERIES SOLUTION The quantum theory of scattering provides a nice illustration of integral equation tech- niques and an application of a Green’s function. Our physical picture of scattering is as follows. A beam of particles moves along the negative z-axis toward the origin. A small fraction of the particles is scattered by the potential V(r)and goes off as an outgoing spherical wave. Our wave function ψ(r)must satisfy the time-independent Schrödinger equation −¯h2 2m∇2ψ(r)+V(r)ψ(r)=Eψ(r), (9.198a) or ∇2ψ(r)+k2ψ(r)=−bracketleftbigg −2m ¯h2V(r)ψ(r)bracketrightbigg ,k2=2mE ¯h2. (9.198b) From the physical picture just presented we look for a solution having an asymptotic form ψ(r)∼eik0·r+fk(θ,ϕ)eikr r. (9.199) Hereeik0·ris the incident plane wave28withk0the propagation vector carrying the sub- script 0 to indicate that it is in the θ=0(z-axis) direction. The magnitudes k0andkare equal (ignoring recoil), and eikr/ris the outgoing spherical wave with an angular (and energy) dependent amplitude factor fk(θ,ϕ).29Vectorkhas the direction of the outgoing scattered wave. In quantum mechanics texts it is shown that the differential probability of scattering, dσ/d/Omega1,thescatteringcross sectionperunitsolidangle,is givenby |fk(θ,ϕ|2. Identifying[−(2m/¯h2)V(r)ψ(r)]withf(r)of Eq.(9.158), wehave ψ(r1)=−integraldisplay2m ¯h2V(r2)ψ(r2)G(r1,r2)d3r2 (9.200) 28Forsimplicityweassumeacontinuousincidentbeam.Inamoresophisticatedandmorerealistictreatment,Eq.(9.199)would be one component of aFourier wavepacket. 29IfV(r)represents a centralforce, fkwill be afunction of θonly, independent of azimuth. 9.7 Nonhomogeneous Equation — Green’s Function 603 byEq.(9.170).ThisdoesnothavethedesiredasymptoticformofEq.(9.199),butwemay add to Eq. (9.200) eik0·r1, a solution of the homogeneous equation, and put ψ(r)into the desiredform: ψ(r1)=eik0·r1−integraldisplay2m ¯h2V(r2)ψ(r2)G(r1,r2)d3r2. (9.201) Our Green’s function is the Green’s function of the operator L=∇2+k2(Eq. (9.198)), satisfyingtheboundaryconditionthatitdescribeanoutgoingwave.Then,fromTable9.5, G(r1,r2)=exp(ik|r1−r2|)/(4π|r1−r2|)and ψ(r1)=eik0·r1−integraldisplay2m ¯h2V(r2)ψ(r2)eik|r1−r2| 4π|r1−r2|d3r2. (9.202) ThisintegralequationanalogoftheoriginalSchrödingerwaveequationis exact.Employ- ing the Neumann series technique of Section 16.3 (remember, the scattering probability is verysmall), wehave ψ0(r1)=eik0·r1, (9.203a) whichhasthephysicalinterpretationof noscattering. Substituting ψ0(r2)=eik0·r2intotheintegral,weobtainthefirst correctionterm, ψ1(r1)=eik0·r1−integraldisplay2m ¯h2V(r2)eik|r1−r2| 4π|r1−r2|eik0·r2d3r2.(9.203b) Thisisthefamous Bornapproximation .Itisexpectedtobemostaccurateforweakpoten- tials and high incident energy. If a more accurate approximation is desired, the Neumann seriesmaybecontinued.30/squaresolid Example 9.7.2 QUANTUM MECHANICAL SCATTERING —G REEN ’SFUNCTION Again, we consider the Schrödinger wave equation (Eq. (9.198b)) for the scattering prob- lem. This time we use Fourier transform techniques and derive the desired form of the Green’s function by contour integration. Substituting the desired asymptotic form of the solution(with kreplacedby k0), ψ(r)∼eik0z+fk0(θ,ϕ)eik0r r=eik0z+/Phi1(r), (9.204) intotheSchrödingerwaveequation,Eq. (9.198b),yields parenleftbig ∇2+k2 0parenrightbig /Phi1(r)=U(r)eik0z+U(r)/Phi1(r). (9.205a) Here ¯h2 2mU(r)=V(r), 30ThisassumestheNeumannseriesisconvergent.Insomephysicalsituationsitisnotconvergentandthenothertechniquesare needed. 604 Chapter 9 Differential Equations thescattering(perturbing)potential.Sincetheprobabilityofscatteringismuchlessthan1, thesecondtermontheright-handsideofEq.(9.205a)isexpectedtobenegligible(relative tothefirsttermontheright-handside)andthuswedropit.Notethatweare approximat- ingour differentialequationwith parenleftbig ∇2+k2 0parenrightbig /Phi1(r)=U(r)eik0z. (9.205b) We now proceed to solve Eq. (9.205b), a nonhomogeneous PDE. The differential oper- ator∇2generatesa continuoussetof eigenfunctions ∇2ψk(r)=−k2ψk(r), (9.206) where ψk(r)=(2π)−3/2eik·r. Theseplane-waveeigenfunctionsform acontinuousbutorthonormalset,inthesensethat integraldisplay ψ∗ k1(r)ψk2(r)d3r=δ(k1−k2) (compareEq. (15.21d)).31Weusetheseeigenfunctionstoderivea Green’sfunction. We expandtheunknownfunction /Phi1(r1)intheseeigenfunctions, /Phi1(r1)=integraldisplay Ak1ψk1(r1)d3k1, (9.207) a Fourier integral with Ak1, the unknown coefficients. Substituting Eq. (9.207) into Eq.(9.205b)andusingEq.(9.206), weobtain integraldisplay Akparenleftbig k2 0−k2parenrightbig ψk(r)d3k=U(r)eik0z. (9.208) Using the now-familiar technique of multiplying by ψ∗ k2(r)and integrating over the space coordinates,wehave integraldisplay Ak1parenleftbig k2 0−k2 1parenrightbig d3k1integraldisplay ψ∗ k2(r)ψk1(r)d3r=Ak2parenleftbig k2 0−k2 2parenrightbig =integraldisplay ψ∗ k2(r)U(r)eik0zd3r.(9.209) Solvingfor Ak2andsubstitutingintoEq. (9.207)wehave /Phi1(r2)=integraldisplaybracketleftbiggparenleftbig k2 0−k2 2parenrightbig−1integraldisplay ψ∗ k2(r1)U(r1)eik0z1d3r1bracketrightbigg ψk2(r2)d3k2.(9.210) Hence /Phi1(r1)=integraldisplay ψk1(r1)parenleftbig k2 0−k2 1parenrightbig−1d3k1integraldisplay ψ∗ k1(r2)U(r2)eik0z2d3r2, (9.211) 31d3r=dxdydz, a(three-dimensional) volume element in r-space. 9.7 Nonhomogeneous Equation — Green’s Function 605 replacing k2byk1andr1byr2to agree with Eq. (9.207). Reversing the order of integra- tion,wehave /Phi1(r1)=−integraldisplay Gk0(r1,r2)U(r2)eik0z2d3r2, (9.212) whereGk0(r1,r2), ourGreen’sfunction,is givenby Gk0(r1,r2)=integraldisplayψ∗ k1(r2)ψk1(r1) k2 1−k2 0d3k1, (9.213) analogous to Eq. (10.90) of Section 10.5 for discrete eigenfunctions. Equation (9.212) shouldbecomparedwiththeGreen’sfunctionsolutionofPoisson’sequation(9.157). It is perhaps worth evaluatingthis integralto emphasizeoncemore the vitalrole played bytheboundaryconditions.UsingtheeigenfunctionsfromEq. (9.206)and d3k=k2dksinθdθdϕ, weobtain Gk0(r1,r2)=1 (2π)3integraldisplay∞ 0integraldisplayπ 0integraldisplay2π 0eikρcosθ k2−k2 0dϕsinθdθk2dk. (9.214) Herekρcosθhas replaced k·(r1−r2), with ρ=r1−r2indicating the polar axis in k- space. Integrating over ϕby inspection, we pick up a 2 π.T h eθ-integration then leads to Gk0(r1,r2)=1 4π2ρiintegraldisplay∞ 0eikρ−e−ikρ k2−k2 0kdk, (9.215) andsincetheintegrandis anevenfunctionof k,wemayset Gk0(r1,r2)=1 8π2ρiintegraldisplay∞ −∞(eiκ−e−iκ) κ2−σ2κdκ. (9.216) Thelatterstepistakeninanticipationoftheevaluationof Gk(r1,r2)asacontourintegral. Thesymbols κandσ(σ>0)represent kρandk0ρ,respectively. If the integral in Eq. (9.216) is interpreted as a Riemann integral, the integral does not exist. This implies that L−1does not exist, and in a literal sense it does not. L=∇2+k2 is singular since there exist nontrivial solutions ψfor which the homogeneous equation Lψ=0. We avoid this problem by introducing a parameter γ, defining a different opera- torL−1 γ, andtakingthelimitas γ→0. Splittingtheintegralintotwopartssothateachpartmaybewrittenasasuitablecontour integralgivesus G(r1,r2)=1 8π2ρicontintegraldisplay C1κeiκdκ κ2−σ2+1 8π2ρicontintegraldisplay C2κe−iκdκ κ2−σ2. (9.217) ContourC1is closed by a semicircle in the upper half-plane, C2by a semicircle in the lower half-plane. These integrals were evaluated in Chapter 7 by using appropriately choseninfinitesimalsemicirclestogoaroundthesingularpoints κ=±σ.Asanalternative procedure, let us first displace the singular points from the real axis by replacing σby σ+iγandthen,afterevaluation,takingthelimitas γ→0 (Fig.9.5). 606 Chapter 9 Differential Equations FIGURE 9.5PossibleGreen’s functioncontoursof integration. Forγpositive, contour C1encloses the singular point κ=σ+iγand the first integral contributes 2πi·1 2ei(σ+iγ). Fromthesecondintegralwealsoobtain 2πi·1 2ei(σ+iγ), theenclosedsingularitybeing κ=−(σ+iγ).ReturningtoEq.(9.217)andletting γ→0, wehave G(r1,r2)=1 4πρeiσ=eik0|r1−r2| 4π|r1−r2|, (9.218) infullagreementwithExercise9.7.16.Thisresultdependsonstartingwith γpositive.Had wechosen γnegative,ourGreen’sfunctionwouldhaveincluded e−iσ,whichcorresponds to anincoming wave. The choice of positive γis dictated by the boundary conditions we wishtosatisfy. Equations (9.212) and (9.218) reproduce the scattered wave in Eq. (9.203b) and consti- tuteanexactsolutionof theapproximateEq. (9.205b).Exercises9.7.18and9.7.20extend theseresults. /squaresolid 9.7 Nonhomogeneous Equation — Green’s Function 607 Exercises 9.7.1 VerifyEq. (9.168), integraldisplay (vL2u−uL2v)dτ2=integraldisplay p(v∇2u−u∇2v)·dσ2. 9.7.2 Showthattheterms +k2intheHelmholtzoperatorand −k2inthemodifiedHelmholtz operatordonotaffectthebehaviorof G(r1,r2)intheimmediatevicinityofthesingular pointr1=r2. Specifically,showthat lim |r1−r2|→0integraldisplay k2G(r1,r2)dτ2=1. 9.7.3 Showthat exp(ik|r1−r2|) 4π|r1−r2| satisfies the two appropriate criteria and therefore is a Green’s function for the Helmholtzequation. 9.7.4 (a) Find the Green’s function for the three-dimensional Helmholtz equation, Exer- cise9.7.3,whenthewaveisastandingwave. (b) Howis thisGreen’sfunctionrelatedtothesphericalBessel functions? 9.7.5 ThehomogeneousHelmholtzequation ∇2ϕ+λ2ϕ=0 haseigenvalues λ2 iandeigenfunctions ϕi.ShowthatthecorrespondingGreen’sfunction thatsatisfies ∇2G(r1,r2)+λ2G(r1,r2)=−δ(r1−r2) maybewrittenas G(r1,r2)=∞summationdisplay i=1ϕi(r1)ϕi(r2) λ2 i−λ2. An expansion of this form is called a bilinearexpansion. If the Green’s function is availablein closedform, thisprovidesameansofgeneratingfunctions. 9.7.6 Anelectrostaticpotential(mksunits)is ϕ(r)=Z 4πε0·e−ar r. Reconstruct the electrical charge distribution that will produce this potential. Note that ϕ(r)vanishesexponentiallyfor large r, showingthatthenetchargeis zero. ANS.ρ(r)=Zδ(r)−Za2 4πe−ar r. 608 Chapter 9 Differential Equations 9.7.7 TransformtheODE d2y(r) dr2−k2y(r)+V0e−r ry(r)=0 andtheboundaryconditions y(0)=y(∞)=0 intoaFredholmintegralequationofthe form y(r)=λintegraldisplay∞ 0G(r,t)e−t ty(t)dt. The quantities V0=λandk2are constants. The ODE is derived from the Schrödinger waveequationwithamesonicpotential: G(r,t)=  1 ke−ktsinhkr,0≤r<t, 1 ke−krsinhkt, t <r < ∞. 9.7.8 Achargedconductingringofradius a(Example12.3.3)maybedescribedby ρ(r)=q 2πa2δ(r−a)δ(cosθ). Using the known Green’s function for this system, Eq. (9.187) find the electrostatic potential. Hint.Exercise12.6.3willbehelpful. 9.7.9 Changingaseparationconstantfrom k2to−k2andputtingthediscontinuityofthefirst derivativeintothe z-dependence,showthat 1 4π|r1−r2|=1 4π∞summationdisplay m=−∞integraldisplay∞ 0eim(ϕ1−ϕ2)Jm(kρ1)Jm(kρ2)e−k|z1−z2|dk. Hint.Therequired δ(ρ1−ρ2)maybeobtainedfromExercise15.1.2. 9.7.10 Derivetheexpansion exp[ik|r1−r2|] 4π|r1−r2|=ik∞summationdisplay l=0  jl(kr1)h(1) l(kr2), r 1<r2 jl(kr2)h(1) l(kr1), r 1>r2   ×lsummationdisplay m=−lYm l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). Hint.TheleftsideisaknownGreen’sfunction.Assumeasphericalharmonicexpansion andworkontheremainingradialdependence.Thesphericalharmonicclosurerelation, Exercise12.6.6,coverstheangulardependence. 9.7.11 ShowthatthemodifiedHelmholtzoperatorGreen’sfunction exp(−k|r1−r2|) 4π|r1−r2| 9.7 Nonhomogeneous Equation — Green’s Function 609 hasthesphericalpolarcoordinateexpansion exp(−k|r1−r2|) 4π|r1−r2|=k∞summationdisplay l=0il(kr<)kl(kr>)lsummationdisplay m=−lYm l(θ1,ϕ1)Ym∗ l(θ2,ϕ2). Note. The modified spherical Bessel functions il(kr)andkl(kr)are defined in Exer- cise11.7.15. 9.7.12 From the spherical Green’s function of Exercise 9.7.10, derive the plane-wave expan- sion eik·r=∞summationdisplay l=0il(2l+1)jl(kr)Pl(cosγ), whereγis the angle included between kandr. This is the Rayleigh equation of Exer- cise12.4.7. Hint.T ak er2≫r1sothat |r1−r2|→r2−r20·r1=r2−k·r1 k. Letr2→∞andcancelafactor of eikr2/r2. 9.7.13 Fromtheresults ofExercises9.7.10and9.7.12,showthat eix=∞summationdisplay l=0il(2l+1)jl(x). 9.7.14 (a) FromthecircularcylindricalcoordinateexpansionoftheLaplaceGreen’sfunction (Eq. (9.197)), showthat 1 (ρ2+z2)1/2=2 πintegraldisplay∞ 0K0(kρ)coskzdk. Thissameresultis obtaineddirectlyinExercise15.3.11. (b) Asaspecialcaseof part(a) showthat integraldisplay∞ 0K0(k)dk=π 2. 9.7.15 Notingthat ψk(r)=1 (2π)3/2eik·r isaneigenfunctionof parenleftbig ∇2+k2parenrightbig ψk(r)=0 (Eq. (9.206)), showthattheGreen’sfunctionof L=∇2maybeexpandedas 1 4π|r1−r2|=1 (2π)3integraldisplay eik·(r1−r2)d3k k2. 610 Chapter 9 Differential Equations 9.7.16 Using Fourier transforms, show that the Green’s function satisfying the nonhomoge- neousHelmholtzequation parenleftbig ∇2+k2 0parenrightbig G(r1,r2)=−δ(r1−r2) is G(r1,r2)=1 (2π)3integraldisplayeik·(r1−r2) k2−k2 0d3k, inagreementwithEq. (9.213). 9.7.17 Thebasicequationof thescalarKirchhoffdiffractiontheoryis ψ(r1)=1 4πintegraldisplay S2bracketleftbiggeikr r∇ψ(r2)−ψ(r2)∇parenleftbiggeikr rparenrightbiggbracketrightbigg ·dσ2, whereψsatisfies the homogeneous Helmholtz equation and r=|r1−r2|. Derive this equation.Assumethat r1is interiortotheclosedsurface S2. Hint.UseGreen’stheorem. 9.7.18 The Born approximation for the scattered wave is given by Eq. (9.203b) (and Eq. (9.211)). Fromtheasymptoticform,Eq. (9.199), fk(θ,ϕ)eikr r=−2m ¯h2integraldisplay V(r2)eik|r−r2| 4π|r−r2|eik0·r2d3r2. Forascatteringpotential V(r2)thatis independentof anglesandfor r≫r2, showthat fk(θ,ϕ)=−2m ¯h2integraldisplay∞ 0r2V(r2)sin(|k0−k|r2) |k0−k|dr2. Herek0is in theθ=0 (original z-axis) direction, whereas kis in the(θ,ϕ)direction. Themagnitudesare equal: |k0|=|k|;misthereducedmass. Hint. You have Exercise 9.7.12 to simplify the exponential and Exercise 15.3.20 to transform the three-dimensional Fourier exponential transform into a one-dimensional Fouriersinetransform. 9.7.19 Calculatethescatteringamplitude fk(θ,ϕ)foramesonicpotential V(r)=V0(e−αr/αr). Hint.ThisparticularpotentialpermitstheBornintegral,Exercise9.7.18,tobeevaluated asaLaplacetransform. ANS.fk(θ,ϕ)=−2mV0 ¯h2α1 α2+(k0−k)2. 9.7.20 Themesonicpotential V(r)=V0(e−αr/αr)maybeusedtodescribetheCoulombscat- tering of two charges q1andq2.W el e tα→0 andV0→0 but take the ratio V0/αto beq1q2/4πε0.(ForGaussianunitsomitthe4 πε0.)Showthatthedifferentialscattering cross section dσ/d/Omega1=|fk(θ,ϕ)|2isgivenby dσ d/Omega1=parenleftbiggq1q2 4πε0parenrightbigg21 16E2sin4(θ/2),E=p2 2m=¯h2k2 2m. Ithappens(coincidentally)thatthisBornapproximationisinexactagreementwithboth theexactquantummechanicalcalculationsandtheclassicalRutherfordcalculation. 9.8 Heat Flow, or Diffusion, PDE 611 9.8 H EAT FLOW ,ORDIFFUSION ,P D E Here we return to a special PDE to developfairly general methods to adapt a special solu- tionofaPDEtoboundaryconditionsbyintroducingparametersthatapplytoothersecond- order PDEs with constant coefficients as well. To some extent, they are complementary to theearlierbasicseparationmethodfor findingsolutionsinasystematicway. We select the full time-dependent diffusion PDE for an isotropic medium. Assuming isotropy actually is not much of a restriction because, in case we have different (constant) ratesofdiffusionindifferentdirections,forexampleinwood,ourheatflowPDEtakesthe form ∂ψ ∂t=a2∂2ψ ∂x2+b2∂2ψ ∂y2+c2∂2ψ ∂z2, (9.219) if we put the coordinate axes along the principal directions of anisotropy. Now we sim- ply rescale the coordinates using the substitutions x=aξ,y=bη,z=cζto get back the originalisotropicform ofEq. (9.219), ∂/Phi1 ∂t=∂2/Phi1 ∂ξ2+∂2/Phi1 ∂η2+∂2/Phi1 ∂ζ2(9.220) forthetemperaturedistributionfunction /Phi1(ξ,η,ζ,t)=ψ(x,y,z,t) . For simplicity, we first solve the time-dependent PDE for a homogeneous one- dimensionalmedium,alongmetalrodinthe x-direction,say, ∂ψ ∂t=a2∂2ψ ∂x2, (9.221) where the constant ameasures the diffusivity, or heat conductivity, of the medium. We attempt to solve this linear PDE with constant coefficients with the relevant exponential product Ansatz ψ=eαx·eβt, which, when substituted into Eq. (9.221), solves the PDE withtheconstraint β=a2α2fortheparameters.Weseekexponentiallydecayingsolutions for large times, that is, solutions with negative βvalues, and therefore set α=iω,α2= −ω2for realωandhave ψ(x,t)=eiωxe−ω2a2t=(cosωx+isinωx)e−ω2a2t. Formingreal linearcombinationsweobtainthesolution ψ(x,t)=(Acosωx+Bsinωx)e−ω2a2t, foranychoiceof A,B,ω,whichareintroducedtosatisfyboundaryconditions.Uponsum- ming over multiples nωof the basic frequency for periodic boundary conditions or inte- grating over the parameter ωfor general (nonperiodic boundary conditions), we find a solution, ψ(x,t)=integraldisplaybracketleftbig A(ω)cosωx+B(ω)sinωxbracketrightbig e−a2ω2tdω, (9.222) thatisgeneralenoughtobeadaptedtoboundaryconditionsat t=0,say.Whenthebound- ary condition gives a nonzero temperature ψ0, as for our rod, then the summation method 612 Chapter 9 Differential Equations applies (Fourier expansion of the boundary condition). If the space is unrestricted (as for aninfinitelyextendedrod), theFourierintegralapplies. •Thissummationorintegrationoverparametersisoneofthestandardmethodsforgen- eralizingspecificPDEsolutionsinordertoadaptthemtoboundaryconditions. Example 9.8.1 ASPECIFIC BOUNDARY CONDITION Let us solve a one-dimensional case explicitly, where the temperature at time t=0i s ψ0(x)=1=const. in the interval between x=+1 andx=−1 and zero for x>1 and x<1.Attheends, x=±1,thetemperatureisalways heldatzero. Forafiniteintervalwechoosethecos (lπx/2)spatialsolutionsofEq.(9.221)forinteger l, becausetheyvanishat x=±1.Thus, at t=0 oursolutionis aFourierseries, ψ(x,0)=∞summationdisplay l=1alcosπlx 2=1,−1<x<1 withcoefficients(see Section14.1.) al=integraldisplay1 −11·cosπlx 2=2 lπsinπlx 2vextendsinglevextendsinglevextendsinglevextendsingle1 x=−1 =4 πlsinlπ 2=4(−1)m (2m+1)π,l=2m+1; al=0,l=2m. Includingitstimedependence,thefullsolutionisgivenbytheseries ψ(x,t)=4 π∞summationdisplay m=0(−1)m 2m+1cosbracketleftbigg (2m+1)πx 2bracketrightbigg e−t((2m+1)πa/2)2,(9.223) which converges absolutely for t>0 but only conditionally at t=0, as a result of the discontinuityat x=±1. Without the restriction to zero temperature at the endpoints of the given finite interval, the Fourier series is replaced by a Fourier integral. The general solution is then given by Eq. (9.222). At t=0 the given temperature distribution ψ0=1 gives the coefficients as (seeSection15.3) A(ω)=1 πintegraldisplay1 −1cosωxdx=1 πsinωx ωvextendsinglevextendsinglevextendsinglevextendsingle1 x=−1=2sinω πω,B(ω)=0. Therefore ψ(x,t)=2 πintegraldisplay∞ 0sinω ωcos(ωx)e−a2ω2tdω. (9.224) /squaresolid Inthree dimensions the corresponding exponential Ansatz ψ=eik·r/a+βtleads to a solution with the relation β=−k2=−k2for its parameter, and the three-dimensional 9.8 Heat Flow, or Diffusion, PDE 613 formofEq. (9.221)becomes ∂2ψ ∂x2+∂2ψ ∂y2+∂2ψ ∂z2+k2ψ=0, (9.225) which is the Helmholtz equation, which may be solved by the separation method just like the earlier Laplace equation in Cartesian, cylindrical, or spherical coordinates under appropriatelygeneralizedboundaryconditions. In Cartesian coordinates, with the product Ansatz of Eq. (9.35), the separated x- andy- ODEsfromEq.(9.221)arethesameasEqs.(9.38)and(9.41),whilethe z-ODE,Eq.(9.42), generalizesto 1 Zd2Z dz2=−k2+l2+m2=n2>0, (9.226) whereweintroduceanotherseparationconstant, n2, constrainedby k2=l2+m2−n2(9.227) to produce a symmetric set of equations. Now, our solution of Helmholtz’s Eq. (9.225) is labeled according to the choice of all three separation constants l,m,nsubject to the constraint Eq. (9.227). As before the z-ODE, Eq. (9.226), yields exponentially decaying solutions∼e−nz. The boundary conditionat z=0 fixes the expansioncoefficients alm,a s inEq.(9.44). Incylindricalcoordinates,wenowusetheseparationconstant l2forthez-ODE,withan exponentiallydecayingsolutioninmind, d2Z dz2=l2Z>0, (9.228) soZ∼e−lz, because the temperature goes to zero at large z.I fw es e t k2+l2=n2, Eqs.(9.53)to(9.54)staythesame,soweendupwiththesameFourier–Besselexpansion, Eq.(9.56), as before. Insphericalcoordinateswithradialboundaryconditions,theseparationmethodleadsto thesameangularODEsinEqs. (9.61) and(9.64), andtheradialODEnowbecomes 1 r2d drparenleftbigg r2dR drparenrightbigg +k2R−QR r2=0,Q=l(l+1), (9.229) that is, of Eq. (9.65), whose solutions are the spherical Bessel functions of Section 11.7. TheyarelistedinTable9.2. Therestrictionthat k2beaconstantisunnecessarilysevere.Theseparationprocesswill stillworkwithHelmholtz’sPDEfor k2asgeneralas k2=f(r)+1 r2g(θ)+1 r2sin2θh(ϕ)+k′2. (9.230) Inthehydrogenatomwehave k2=f(r)intheSchrödingerwaveequation,andthisleads toaclosed-formsolutioninvolvingLaguerrepolynomials. 614 Chapter 9 Differential Equations Alternate Solutions In a new approach to the heat flow PDE suggested by experiments, we now return to the one-dimensional PDE, Eq. (9.221), seeking solutions of a new functional form ψ(x,t)=u(x/√t), which is suggested by Example 15.1.1. Substituting u(ξ),ξ=x/√t, intoEq. (9.221)using ∂ψ ∂x=u′ √t,∂2ψ ∂x2=u′′ t,∂ψ ∂t=−x 2√ t3u′(9.231) withthenotation u′(ξ)≡du dξ, thePDEis reducedtotheODE 2a2u′′(ξ)+ξu′(ξ)=0. (9.232) WritingthisODEas u′′ u′=−ξ 2a2, we can integrate it once to get ln u′=−ξ2 4a2+lnC1, with an integration constant C1.E x - ponentiatingandintegratingagainwefindthesolution u(ξ)=C1integraldisplayξ 0e−ξ2 4a2dξ+C2, (9.233) involving two integration constants Ci. Normalizing this solution at time t=0 to temper- ature+1f o rx>0 and−1f o rx<0, our boundary conditions, fixes the constants Ci, so ψ=1 a√πintegraldisplayx√t 0e−ξ2 4a2dξ=2√πintegraldisplayx 2a√t 0e−v2dv=/Phi1parenleftbiggx 2a√tparenrightbigg ,(9.234) where/Phi1denotes Gauss’ error function (see Exercise 5.10.4). See Example 15.1.1 for a derivation using a Fourier transform. We need to generalize this specific solution to adapt ittoboundaryconditions. To this end we now generate new solutions of the PDE with constant coefficients by differentiating a special solution , Eq. (9.234). In other words, if ψ(x,t)solves the PDE in Eq. (9.221), so do∂ψ ∂tand∂ψ ∂x, because these derivatives and the differentiations of the PDE commute;that is, the order in whichthey are carried out does not matter. Note carefully that this method no longer works if any coefficient of the PDE depends on t orxexplicitly. However, PDEs with constant coefficients dominate in physics. Examples are Newton’s equations of motion (ODEs) in classical mechanics, the wave equations of electrodynamics,andPoisson’sandLaplace’sequationsinelectrostaticsandgravity.Even Einstein’s nonlinear field equations of general relativity take on this special form in local geodesiccoordinates. Therefore, by differentiating Eq. (9.234) with respect to x, we find the simpler, more basicsolution ψ1(x,t)=1 a√tπe−x2 4a2t, (9.235) 9.8 Heat Flow, or Diffusion, PDE 615 and,repeatingtheprocess, anotherbasicsolution, ψ2(x,t)=x 2a3√ t3πe−x2 4a2t. (9.236) Again, these solutions have to be generalized to adapt them to boundary conditions. And there is yet another method of generating new solutions of a PDE with constant coeffi- cients:We can translate a givensolution, for example, ψ1(x,t)→ψ1(x−α,t), and then integrateoverthetranslationparameter α.Therefore ψ(x,t)=1 2a√tπintegraldisplay∞ −∞C(α)e−(x−α)2 4a2tdα (9.237) isagainasolution,whichwerewriteusingthesubstitution ξ=x−α 2a√t,α=x−2aξ√ t, dα=−2adξ√ t. (9.238) Thus,wefindthat ψ(x,t)=1√πintegraldisplay∞ −∞C(x−2aξ√ t)e−ξ2dξ (9.239) isasolutionofourPDE.Inthisformwerecognizethesignificanceoftheweightfunction C(x)from the translation method because, at t=0,ψ ( x ,0)=C(x)=ψ0(x)is deter- mined by the boundary condition, andintegraltext∞ −∞e−ξ2dξ=√π. Therefore, we can also write thesolutionas ψ(x,t)=1√πintegraldisplay∞ −∞ψ0(x−2aξ√ t)e−ξ2dξ, (9.240) displaying the role of the boundary condition explicitly. From Eq. (9.240) we see that the initial temperature distribution, ψ0(x), spreads out over time and is damped by the Gaussianweightfunction. Example 9.8.2 SPECIAL BOUNDARY CONDITION AGAIN Let us express the solution of Example 9.8.1 in terms of the error function solution of Eq. (9.234). The boundary condition at t=0i sψ0(x)=1f o r−1<x<1 and zero for|x|>1. From Eq. (9.240) we find the limits on the integration variable ξby setting x−2aξ√t=±1. This yields the integration endpoints ξ=(±1+x)/2a√t. Therefore oursolutionbecomes ψ(x,t)=1√πintegraldisplayx+1 2a√t x−1 2a√te−ξ2dξ. Usingtheerror functiondefinedinEq.(9.234) wecanalsowritethissolutionasfollows ψ(x,t)=1 2bracketleftbigg erfparenleftbiggx+1 2a√tparenrightbigg −erfparenleftbiggx−1 2a√tparenrightbiggbracketrightbigg . (9.241) 616 Chapter 9 Differential Equations Comparing this form of our solution with that from Example 9.8.1 we see that we can express Eq. (9.241) as the Fourier integral of Example 9.8.1, an identity that gives the Fourierintegral,Eq.(9.224), inclosedform ofthetabulatederror function. /squaresolid Finally,weconsidertheheatflowcaseforanextended sphericallysymmetric medium centered at the origin, which prescribes polar coordinates r,θ,ϕ.We expect a solution of theformψ(r,t)=u(r,t). UsingEq.(2.48) wefindthePDE ∂u ∂t=a2parenleftbigg∂2u ∂r2+2 r∂u ∂rparenrightbigg , (9.242) whichwetransformtotheone-dimensionalheatflowPDEbythesubstitution u=v(r,t) r,∂u ∂r=1 r∂v ∂r−v r2,∂u ∂t=1 r∂v ∂t, ∂2u ∂r2=1 r∂2v ∂r2−2 r2∂v ∂r+2v r3. (9.243) ThisyieldsthePDE ∂v ∂t=a2∂2v ∂r2. (9.244) Example 9.8.3 SPHERICALLY SYMMETRIC HEATFLOW Let us apply the one-dimensional heat flow PDE with the solution Eq. (9.234) to a spheri- cally symmetric heat flow under fairly common boundary conditions, where xis released by the radial variable. Initially we have zero temperature everywhere. Then, at time t=0, afiniteamountofheatenergy Qisreleasedattheorigin,spreadingevenlyinalldirections. Whatistheresultingspatialandtemporaltemperaturedistribution? InspectingourspecialsolutioninEq. (9.236)weseethat,for t→0,thetemperature v(r,t) r=C√ t3e−r2 4a2t (9.245) goestozeroforall r/negationslash=0,sozeroinitialtemperatureisguaranteed.As t→∞,thetemper- aturev/r→0 for allrincluding the origin, which is implicit in our boundary conditions. Theconstant Ccanbedeterminedfrom energyconservation,whichgivestheconstraint Q=σρintegraldisplayv rd3r=4πσρC√ t3integraldisplay∞ 0r2e−r2 4a2tdr=8radicalbig π3σρa3C, (9.246) whereρis the constant density of the medium and σis its specific heat. Here we have rescaledtheintegrationvariableandintegratedbyparts toget integraldisplay∞ 0e−r2 4a2tr2dr=(2a√ t)3integraldisplay∞ 0e−ξ2ξ2dξ, integraldisplay∞ 0e−ξ2ξ2dξ=−ξ 2e−ξ2vextendsinglevextendsinglevextendsinglevextendsingle∞ 0+1 2integraldisplay∞ 0e−ξ2dξ=√π 4. 9.8 Heat Flow, or Diffusion, PDE 617 Thetemperature,asgivenbyEq.(9.245)atanymoment,whichisatfixed t,isaGaussian distribution that flattens out as time increases, because its width is proportional to√t.A s a function of time the temperature is proportional to t−3/2e−T/t, withT≡r2/4a2, which rises from zero to a maximum and then falls off to zero again for large times. To find the maximum,weset d dtparenleftbig t−3/2e−T/tparenrightbig =t−5/2e−T/tparenleftbiggT t−3 2parenrightbigg =0, (9.247) fromwhichwefind t=2T/3. /squaresolid In the case of cylindrical symmetry (in the plane z=0 in plane polar coordinates ρ=radicalbig x2+y2,ϕ)welookforatemperature ψ=u(ρ,t)thatthensatisfiestheODE(using Eq.(2.35) inthediffusionequation) ∂u ∂t=a2parenleftbigg∂2u ∂ρ2+1 ρ∂u ∂ρparenrightbigg , (9.248) whichistheplanaranalogofEq.(9.244).ThisODEalsohassolutionswiththefunctional dependence ρ/√t≡r. Uponsubstituting u=vparenleftbiggρ√tparenrightbigg ,∂u ∂t=−ρv′ 2t3/2,∂u ∂ρ=v′ √t,∂2u ∂ρ2=v′ t(9.249) intoEq. (9.248)withthenotation v′≡dv dr, wefindtheODE a2v′′+parenleftbigga2 r+r 2parenrightbigg v′=0. (9.250) This is a first-order ODE for v′, which we can integrate when we separate the variables v andras v′′ v′=−parenleftbigg1 r+r 2a2parenrightbigg . (9.251) Thisyields v(r)=C re−r2 4a2=C√t ρe−ρ2 4a2t. (9.252) Thisspecialsolutionforcylindricalsymmetrycanbesimilarlygeneralizedandadaptedto boundary conditions, as for the spherical case. Finally, the z-dependence can be factored in,because zseparatesfrom theplanepolarradialvariable ρ. Insummary,PDEscanbesolvedwithinitialconditions,justlikeODEs,orwithbound- aryconditionsprescribingthevalueofthesolutionoritsderivativeonboundarysurfaces, curves, or points. When the solution is prescribed on the boundary, the PDE is called a Dirichlet problem;if the normal derivative of the solution is prescribed on the boundary, thePDE iscalleda Neumann problem. Whentheinitialtemperatureisprescribedfortheone-dimensionalorthree-dimensional heatequation (withsphericalorcylindricalsymmetry )itbecomesaweightfunctionofthe solution,intermsofanintegraloverthegenericGaussiansolution.Thethree-dimensional 618 Chapter 9 Differential Equations heat equation, with spherical or cylindrical boundary conditions, is solved by separation of the variables, leading to eigenfunctions in each separated variable and eigenvalues as separationconstants.Forfiniteboundaryintervalsineachspatialcoordinate,thesumover separationconstantsleadstoaFourier-seriessolution,whileinfiniteboundaryconditions lead to a Fourier-integral solution. The separation of variables method attempts to solve a PDE by writing the solution as a product of functions of one variable each. General conditions for the separation method to work are provided by the symmetry properties of thePDE, towhichcontinuousgrouptheoryapplies. AdditionalReadings Bateman, H., Partial Differential Equations of Mathematical Physics . New York: Dover (1944), 1st ed. (1932). A wealth of applications of various partial differential equations in classical physics. Excellent examples of the use of different coordinate systems—ellipsoidal, paraboloidal, toroidal coordinates, and so on. Cohen, H., Mathematics for Scientists and Engineers .Englewood Cliffs, NJ: Prentice-Hall (1992). Courant, R., and D. Hilbert, Methods of Mathematical Physics , Vol. 1 (English edition). New York: Interscience (1953), Wiley (1989). This is one of the classic works of mathematical physics. Originally published in Ger- manin1924,therevisedEnglisheditionisanexcellentreferenceforarigoroustreatmentofGreen’sfunctions and for a widevariety of othertopics on mathematical physics. Davis,P.J.,andP.Rabinowitz, Numerical Integration . Waltham,MA:Blaisdell (1967). Thisbook covers agreat deal of material in a relatively easy-to-read form. Appendix 1 ( On the Practical Evaluation of Integrals by M.Abramowitz) is excellentas anoverall view. Garcia,A.L., Numerical Methods for Physics .Englewood Cliffs, NJ: Prentice-Hall(1994). Hamming, R. W., Numerical Methods for Scientists and Engineers , 2nd ed. New York: McGraw-Hill (1973), reprinted Dover (1987). This well-written text discusses a wide variety of numerical methods from zeros of functionstothefastFouriertransform.Alltopicsareselectedanddevelopedwithamoderncomputerinmind. Hubbard, J., andB. H.West, Differential Equations . Berlin: Springer (1995). Ince,E.L., OrdinaryDifferentialEquations .NewYork:Dover(1956).Theclassicworkinthetheoryofordinary differential equations. Lapidus, L., and J. H. Seinfeld, Numerical Solutions of Ordinary Differential Equations . New York: Academic Press(1971).Adetailedandcomprehensivediscussionofnumericaltechniques,withemphasisontheRunge– Kuttaandpredictor–correctormethods.Recentworkontheimprovementofcharacteristicssuchasstabilityis clearlypresented. Margenau, H., and G. M. Murhpy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nos- trand (1956). Chapter5 covers curvilinear coordinates and 13 specificcoordinate systems. Miller,R. K.,andA.N.Michel, Ordinary DifferentialEquations . NewYork: AcademicPress (1982). Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 includes a description of several different coordinate systems. Note that Morse and Feshbach are not above usingleft-handedcoordinatesystemsevenforCartesiancoordinates.Elsewhereinthisexcellent(anddifficult) bookaremanyexamplesoftheuseofthevariouscoordinatesystemsinsolvingphysicalproblems.Chapter7 is a particularly detailed, complete discussion of Green’s functions from the point of view of mathematical physics. Note, however, that Morse and Feshbach frequently choose a source of 4 πδ(r−r′)in place of our δ(r−r′).Considerable attention is devoted to bounded regions. Murphy,G.M., OrdinaryDifferentialEquationsandTheirSolutions .Princeton,NJ:VanNostrand(1960).Athor- ough, relatively readabletreatment of ordinary differential equations, both linear and nonlinear. Press, W. H., B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes , 2nd ed. Cambridge, UK: Cambridge University Press (1992). Ralston, A.,andH.Wilf, eds., Mathematical Methods for Digital Computers . NewYork: Wiley (1960). 9.8 Additional Readings 619 Ritger, P. D.,andN.J.Rose, Differential Equations with Applications . NewYork: McGraw-Hill(1968). Stakgold, I., Green’s Functions and Boundary ValueProblems , 2nd ed.NewYork: Wiley (1997). Stoer, J.,and R. Burlirsch, Introduction to Numerical Analysis . NewYork: Springer-Verlag (1992). Stroud, A. H., Numerical Quadrature and Solution of Ordinary Differential Equations , Applied Mathematics Series, Vol. 10. New York: Springer-Verlag (1974). A balanced, readable, and very helpful discussion of var- ious methods of integrating differential equations. Stroud is familiar with the work in this field and provides numerous references. This page intentionally left blank CHAPTER 10 STURM –LIOUVILLE THEORY —O RTHOGONAL FUNCTIONS In the preceding chapter we developed two linearly independent solutions of the second- order linear homogeneous differential equation and proved that no third, linearly inde- pendent solution existed. In this chapter the emphasis shifts from solving the differential equation to developing and understanding general properties of the solutions. There is a close analogy between the concepts in this chapter and those of linear algebra in Chap- ter 3. Functions here play the role of vectors there, and linear operators that of matri- ces in Chapter 3. The diagonalization of a real symmetric matrix in Chapter 3 corre- sponds here to the solution of an ODE defined by a self-adjoint operator Lin terms of its eigenfunctions, which are the “continuous” analog of the eigenvectors in Chap- ter 3. Examples for the corresponding analogy between Hermitian matrices and Her- mitian operators are Hamiltonians in quantum mechanics and their energy eigenfunc- tions. InSection10.1theconceptsofself-adjointoperator,eigenfunction,eigenvalue,andHer- mitian operator are presented. The concept of adjoint operator, given first in terms of dif- ferential equations, is then redefined in accordance with usage in quantum mechanics, where eigenfunctions take complex values. The vital properties of reality of eigenvalues and orthogonality of eigenfunctions are derived in Section 10.2. In Section 10.3 we dis- cuss the Gram–Schmidt procedure for systematically constructuring sets of orthogonal functions. Finally, the general property of the completeness of a set of eigenfunctions is explored in Section 10.4, and Green’s functions from Chapter 9 are continued in Sec- tion10.5. 621 622 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.1 S ELF-ADJOINT ODE S In Chapter 9 we studied, classified, and solved linear, second-order ODEs corresponding tolinear,second-orderdifferentialoperatorsofthegeneralform Lu(x)=p0(x)d2 dx2u(x)+p1(x)d dxu(x)+p2(x)u(x). (10.1) The coefficients p0(x),p1(x), andp2(x)are real functions of x, and over the region of interest, a≤x≤b, the first 2−iderivatives of pi(x)are continuous. Reference to Eq.(9.118)showsthat P(x)=p1(x)/p0(x)andQ(x)=p2(x)/p0(x).Hence,p0(x)must not vanish for a<x<b . Now, the zeros of p0(x)are singular points (Section 9.4), and the preceding statement means that our interval [a,b]must be given so that there are no singularpointsintheinterioroftheinterval.Theremaybeandoftenaresingularpointson theboundaries. For a linear operator L, the analog of a quadratic form for a matrix in Chapter 3 is the integral /angbracketleftu|L|u/angbracketright≡/angbracketleftu|Lu/angbracketright≡integraldisplayb au(x)Lu(x)dx =integraldisplayb au{p0u′′+p1u′+p2u}dx, (10.2) wheretheprimesontherealfunction u(x)denotederivatives,asusual,and,forsimplicity, u(x)is taken to be real. If we shift the derivatives to the first factor, u, in Eq. (10.2) by integratingbypartsonceor twice,weareledtotheequivalentexpression, /angbracketleftu|L|u/angbracketright=bracketleftbig u(x)(p1−p′ 0)u(x)bracketrightbigb x=a +integraldisplayb abraceleftbiggd2 dx2[p0u]−d dx[p1u]+p2ubracerightbigg udx. (10.3) If we require that the integrals in Eqs. (10.2) and (10.3) be identical for all (twice differ- entiable)functions u, thentheintegrandshavetobeequal.Thecomparisonthenyields u(p′′ 0−p′ 1)u+2u(p′ 0−p1)u′=0, or p′ 0(x)=p1(x), (10.4) and,asabonus,thetermsattheboundaries x=aandx=binEq.(10.3)thenalsovanish. BecauseoftheanalogywiththetransposedmatrixinChapter3,itisconvenienttodefine thelinearoperatorinEq.(10.3), ¯Lu=d2 dx2[p0u]−d dx[p1u]+p2u =p0d2u dx2+(2p′ 0−p1)du dx+(p′′ 0−p′ 1+p2)u, (10.5) 10.1 Self-Adjoint ODEs 623 astheadjoint1operator¯L.Wehavedefinedtheadjointoperator ¯Landhaveshownthatif Eq.(10.4)issatisfied, /angbracketleft¯Lu|u/angbracketright=/angbracketleftu|Lu/angbracketright.Followingthesameprocedurewecanshowmore generallythat/angbracketleftv|Lu/angbracketright=/angbracketleftLv|u/angbracketright.Whenthis conditionissatisfied, ¯Lu=Lu=d dxbracketleftbigg p(x)du(x) dxbracketrightbigg +q(x)u(x), (10.6) the operator Lis said to be self-adjoint . Here, for the self-adjoint case, p0(x)is replaced byp(x)andp2(x)byq(x)to avoid unnecessary subscripts. The form of Eq. (10.6) al- lows carrying out two integrations by parts in Eq. (10.3) (and Eq. (10.22) and following) withoutintegratedterms.2Notethatagivenoperatorisnotinherentlyself-adjoint;itsself- adjointness depends on the properties of the function space in which it acts and on the boundaryconditions. In a survey of the ODEs introduced in Section 9.3, Legendre’s equation and the linear oscillatorequationareself-adjoint,butothers,suchastheLaguerreandHermiteequations, are not. However, the theory of linear, second-order, self-adjoint differential equations is perfectly general because we can alwaystransform the non-self-adjoint operator into the requiredself-adjointform.ConsiderEq. (10.1) with p′ 0/negationslash=p1. If wemultiply Lby3 1 p0(x)expbracketleftbiggintegraldisplayxp1(t) p0(t)dtbracketrightbigg , weobtain 1 p0(x)expbracketleftbiggintegraldisplayxp1(t) p0(t)dtbracketrightbigg Lu(x)=d dxbraceleftbigg expbracketleftbiggintegraldisplayxp1(t) p0(t)dtbracketrightbiggdu(x) dxbracerightbigg +p2(x) p0(x)·expbracketleftbiggintegraldisplayxp1(t) p0(t)dtbracketrightbigg u,(10.7) which is clearly self-adjoint (see Eq. (10.6)). Notice the p0(x)in the denominator. This is whywerequire p0(x)/negationslash=0,a<x<b .Inthefollowingdevelopmentweassumethat Lhas beenputintoself-adjointform. 1Theadjointoperator bears a somewhat forced relationship to the adjointmatrix. A better justification for the nomenclature is found in a comparison of the self-adjoint operator (plus appropriate boundary conditions) with the self-adjoint matrix. The significant properties aredeveloped in Section10.2. Becauseofthese properties, weareinterested in self-adjoint operators. 2The full importance of the self-adjoint form (plus boundary conditions) will become apparent in Section 10.2. In addition, self-adjoint forms will be required for developing Green’s functions in Section 10.5. 3If wemultiply Lbyf(x)/p0(x)andthendemandthat f′(x)=fp1 p0, so that the newoperator will be self-adjoint, weobtain f(x)=expbracketleftbiggintegraldisplayxp1(t) p0(t)dtbracketrightbigg . 624 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Eigenfunctions, Eigenvalues Schrödinger’swaveequation Hψ(x)=Eψ(x) is the major example of an eigenvalue equation in physics; here the differential operator Lis defined by the Hamiltonian Hand may no longer be real, and the eigenvalue be- comes the total energy Eof the system. The eigenfunction ψ(x)may be complex and is usually called a wave function . A variational formulation of this Schrödinger equation appears in Section 17.7. Based on spherical, cylindrical, or some other symmetry prop- erties, a three- or four-dimensional PDE or eigenvalue equation such as the Schrödinger equation may separate into eigenvalue equations in a single variable each. Examples are Eqs. (9.41), (9.42), (9.50), and (9.53). However, sometimes an eigenvalue equation takes themoregeneralself-adjointform Lu(x)+λw(x)u(x)=0, (10.8) where the constant λis the eigenvalue4andw(x)is a known weight or density func- tion;w(x) >0 except possibly at isolated points at which w(x)=0. (In Section 10.1, w(x)≡1.) For a given choice of the parameter λ, a function uλ(x), which satisfies Eq.(10.8) andtheimposedboundaryconditions ,iscalledan eigenfunction correspond- ingtoλ.Theconstant λisthencalledan eigenvalue bymathematicians.Thereisnoguar- antee that an eigenfunction uλ(x)will exist for an arbitrary choice of the parameter λ. Indeed,therequirementthattherebeaneigenfunctionoftenrestrictstheacceptablevalues ofλto a discrete set. Examples of this for the Legendre, Hermite, and Chebyshev equa- tions appear in the exercises of Section 9.5. Here we have the mathematical approach to theprocess ofquantizationinquantummechanics. The inner product of two functions, /angbracketleftv|u/angbracketright=integraltextb av∗(x)w(x)u(x)dx, depends on the weight function and generalizes our previous definition, where w(x)≡1.The weight function also modifies the definition of orthogonality of two eigenfunctions: They are orthogonal if their inner product /angbracketleftuλ′|uλ/angbracketright=0.The extra weight function w(x)appears sometimes as an asymptotic wave function ψ∞that is a common factor in all solutions of a PDE such as the Schrödinger equation, for example, when the potential V(x)→0a s x→∞inH=T+V. We can find ψ∞when we set V=0 in the Schrödinger equa- tion. Another source for w(x)may be a nonzero angular momentum barrier l(l+1)/x2 in a PDE or separated ODE Eq. (9.65) that has a regular singularity and dominates at x→0. In such a case the indicial equation, such as Eq. (9.87) or (9.103), shows that the wave function has xlas an overall factor. Since the wave function enters twice in matrix elements and orthogonality relations, the weight functions in Table 10.1 come from these common factors in both radial wave functions. This is how the exp (−x)for Laguerre polynomials arises and xkexp(−x)for associated Laguerre polynomials in Ta- ble10.1. 4Notethat this mathematicaldefinition of the eigenvalue differs by a sign from the usage in physics. 10.1 Self-Adjoint ODEs 625 Table 10.1 Equation p(x) q(x) λ w(x) Legendrea1−x20 l(l+1) 1 Shifted Legendreax(1−x) 0 l(l+1) 1 AssociatedLegendrea1−x2−m2/(1−x2)l(l+1) 1 Chebyshev I (1−x2)1/20 n2(1−x2)−1/2 Shifted Chebyshev I [x(1−x)]1/20 n2[x(1−x)]−1/2 Chebyshev II (1−x2)3/20 n(n+2)(1−x2)1/2 Ultraspherical (Gegenbauer) (1−x2)α+1/20 n(n+2α) (1−x2)α−1/2 Besselb,0≤x≤ax −n2/x a2x Laguerre, 0≤x<∞ xe−x0 αe−x AssociatedLaguerrecxk+1e−x0 α−kxke−x Hermite, 0≤x<∞ e−x202 αe−x2 Simple harmonic oscillatord10 n21 al=0,1,...,−l≤m≤lareintegersand −1≤x≤1,0≤x≤1forshiftedLegendre. bOrthogonalityofBesselfunctionsisratherspecial.CompareSection11.2.fordetails.Asecondtypeoforthogonality isdevelopedinEq.(11.174). ckis anon-negativeinteger.Formore details,seeTable10.2. dThiswillformthebasisforChapter14,Fourierseries. Example 10.1.1 LEGENDRE ’SEQUATION Legendre’sequationis givenby parenleftbig 1−x2parenrightbig u′′−2xu′+n(n+1)u=0,−1≤x≤1. (10.9) FromEqs. (10.1), (10.8), and(10.9), p0(x)=1−x2=p, w(x)=1, p1(x)=−2x=p′,λ=n(n+1), p2(x)=0=q. RecallthatourseriessolutionsofLegendre’sequation(Exercise9.5.5)5divergedunless n wasrestrictedtooneoftheintegers.This representsaquantizationoftheeigenvalue λ./squaresolid When the equations of Chapter 9 are transformed into the self-adjoint form, we find the following values of the coefficients and parameters (Table 10.1). The coefficient p(x) is the coefficient of the second derivative of the eigenfunction. The eigenvalue λis the parameterthatis availablein a termof theform λw(x)u(x) ;a n yxdependenceapartfrom theeigenfunctionbecomestheweightingfunction w(x).Ifthereisanothertermcontaining theeigenfunction(notthederivatives),thecoefficientoftheeigenfunctioninthisadditional termisidentifiedas q(x). If nosuchtermispresent, q(x)is zero. 5Compare also Exercise5.2.15 and 12.10. 626 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Example 10.1.2 DEUTERON Further insight into the concepts of eigenfunction and eigenvalue may be provided by an extremely simple model of the deuteron, a bound state of a neutron and proton. From experiment,thebindingenergyofabout2MeV ≪Mc2,withM=Mp=Mn,thecommon neutron and proton mass whose small mass difference we neglect. Due to the short range of the nuclear force, the deuteron properties do not depend much on the detailed shape of the interactionpotential. Thus, the neutron–protonnuclear interaction may be modeled by asphericallysymmetricsquarewellpotential: V=V0<0f o r0≤r<a,V=0f o rr>a. TheSchrödingerwaveequationis −¯h2 M∇2ψ+Vψ=Eψ, (10.10) wheretheenergyeigenvalue E<0foraboundstate.Forthegroundstatetheorbitalangu- lar momentum l=0 because for l/negationslash=0 there is the additional positive angular momentum barrier. So, with ψ=ψ(r), we may write u(r)=rψ(r), and, using Exercise 2.5.18, the waveequationbecomes d2u dr2+k2 1u=0, (10.11) with k2 1=M ¯h2(E−V0)>0 (10.12) fortheinteriorrange, 0 ≤r<a.F ora<r<∞,weha v e d2u dr2−k2 2u=0, (10.13) with k2 2=−ME ¯h2>0. (10.14) Theboundaryconditionthat ψremainfiniteat r=0 implies u(0)=0 and u1(r)=sink1r,0≤r<a. (10.15) In the range outside the potential well, we have a linear combination of the two exponen- tials, u2(r)=Aexpk2r+Bexp(−k2r), a<r< ∞. (10.16) Continuity of particle and current density demand that u1(a)=u2(a)and thatu′ 1(a)= u′2(a). Thesejoining,ormatching,conditions give sink1a=Aexpk2a+Bexp(−k2a), k1cosk1a=k2Aexpk2a−k2Bexp(−k2a).(10.17) The condition that we actually have a bound proton–neutron combination is thatintegraltext∞ 0u2(r)dr=1. This constraint can be met if we impose a boundarycondition that ψ(r) 10.1 Self-Adjoint ODEs 627 FIGURE 10.1Adeuteroneigenfunction. remain finite as r→∞. And this, in turn, means that A=0. Dividing the preceding pair ofequations(tocancel B), weobtain tank1a=−k1 k2=−radicalbigg E−V0 −E, (10.18) atranscendentalequationfortheenergy Ewithonlycertaindiscretesolutions.If Eissuch that Eq. (10.18) can be satisfied, our solutions u1(r)andu2(r)can satisfy the boundary conditions. If Eq. (10.18) is not satisfied, no acceptable solution exists . The values of Efor which Eq. (10.18) is satisfied are the eigenvalues; the corresponding functions u1 andu2(orψ)aretheeigenfunctions.Forthedeuteron,problemthereisone(andonlyone) negativevalueof EsatisfyingEq.(10.18);thatis,thedeuteronhasoneandonlyonebound state. Now, what happens if Edoes not satisfy Eq. (10.18), that is, if E/negationslash=E0is not an eigenvalue? In graphical form, imagine that Eand therefore k1are varied slightly. For E=E1<E0,k1isreducedand sin k1ahasnotturneddownenoughtomatch exp (−k2a). The joining conditions, Eq. (10.17), require A>0 and the wave function goes to +∞ex- ponentially. For E=E2>E0,k1is larger, sin k1apeaks sooner and has descended more rapidlyat r=a.Thejoiningconditionsdemand A<0,andthewavefunctiongoesto −∞ exponentially. Only for E=E0, an eigenvalue, will the wave function have the required negativeexponentialasymptoticbehavior(see Fig.10.1). /squaresolid Boundary Conditions Intheforegoingdefinitionofeigenfunction,itwasnotedthattheeigenfunction uλ(x)was required to satisfy certain imposed boundary conditions. The term boundary conditions includes as a special case the concept of initial conditions . For instance, specifying the initialposition x0andtheinitialvelocity v0insomedynamicalproblemwouldcorrespond to the Cauchy boundary conditions. The only difference in the present usage of boundary 628 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions conditions in these one-dimensional problems is that we are going to apply the conditions onbothendsof theallowedrangeofthevariable. Usuallytheformofthedifferentialequationortheboundaryconditionsonthesolutions will guarantee that at the ends of our interval (that is, at the boundary, as suggested by Eq.(10.3)) thefollowingproductswillvanish: p(x)v∗(x)du(x) dxvextendsinglevextendsinglevextendsinglevextendsingle x=a=0 and p(x)v∗(x)du(x) dxvextendsinglevextendsinglevextendsinglevextendsingle x=b=0.(10.19) Hereu(x)andv(x)are solutions of the particular ODE (Eq. (10.8)) being considered. A reason for this particular form of Eq. (10.19) is suggested shortly. If we recall the radial wave function uof the hydrogen atom with u(0)=0 anddu/dr∼e−kr→0a sr→∞, then both boundary conditions are satisfied. Similarly in the deuteron Example 10.1.2, sink1r→0a sr→0 andd(e−k2r)/dr→0a sr→∞,both boundary conditions are obeyed. We can, however, work with a somewhat less restrictive set of boundary condi- tions, v∗pu′vextendsinglevextendsingle x=a=v∗pu′vextendsinglevextendsingle x=b, (10.20) inwhichu(x)andv(x)aresolutionsofthedifferentialequationcorrespondingtothesame ortodifferenteigenvalues.Equation(10.20)mightwellbesatisfiedifweweredealingwith aperiodicphysicalsystem, suchas acrystallattice. Equations (10.19) and (10.20) are written in terms of v∗, complex conjugate. When the solutionsarereal, v=v∗andtheasteriskmaybeignored.However,inFourierexponential expansions and in quantum mechanics the functions will be complex and the complex conjugatewillbeneeded. Example 10.1.3 INTEGRATION INTERVAL [a,b] ForL=d2/dx2, apossibleeigenvalueequationis d2 dx2u(x)+n2u(x)=0, (10.21) witheigenfunctions un=cosnx, v m=sinmx. Equation(10.20)becomes −nsinmxsinnxvextendsinglevextendsingleb a=0,ormcosmxcosnxvextendsinglevextendsingleb a=0, interchanging unandvm. Since sin mxand cosnxare periodic with period 2 π(fornand mintegral), Eq. (10.20) is clearly satisfied if a=x0andb=x0+2π. If a problem pre- scribes a different interval, the eigenfunctions and eigenvalues will change along with the boundary conditions. The functions must always be chosen so that the boundary condi- tions (Eq. (10.20) etc.) are satisfied. For this case (Fourier series) the usual choices are x0=0 leading to (0,2π)andx0=−πleading to (−π,π). Here and throughout the fol- lowing several chapters the orthogonality interval is so that the boundary conditions (Eq. (10.20)) will be satisfied . The interval[a,b]and the weighting factor w(x)for the mostcommonlyencounteredsecond-orderdifferentialequationsarelistedinTable10.2. /squaresolid 10.1 Self-Adjoint ODEs 629 Table 10.2 Equation ab w (x) Legendre −11 1 Shifted Legendre 0 1 1 AssociatedLegendre −11 1 Chebyshev I −11 (1−x2)−1/2 Shifted Chebyshev I 0 1 [x(1−x)]−1/2 Chebyshev II −11 (1−x2)1/2 Laguerre 0 ∞ e−x AssociatedLaguerre 0 ∞ xke−x Hermite −∞ ∞ e−x2 Simple harmonic oscillator 0 2 π 1 −ππ 1 1. The orthogonalityinterval [a,b]is determined by the boundary condi- tionsofSection10.1. 2. The weighting function is established by putting the ODE in self- adjointform. Hermitian Operators Wenowproveanimportantpropertyoftheself-adjoint,second-orderdifferentialoperator (Eq. (10.8)), in conjunction with solutions u(x)andv(x)that satisfy boundary conditions givenbyEq.(10.20). Thisis motivatedbyapplicationsinquantummechanics. By integrating v∗(complex conjugate) times the second-order self-adjoint differential operator L(operatingon u)overtherange a≤x≤b,weobtain integraldisplayb av∗Ludx=integraldisplayb av∗(pu′)′dx+integraldisplayb av∗qudx (10.22) usingEq. (10.6). Integratingbyparts,wehave integraldisplayb av∗(pu′)′dx=v∗pu′vextendsinglevextendsingleb a−integraldisplayb av∗′pu′dx. (10.23) Theintegratedpartvanishesonapplicationoftheboundaryconditions(Eq.(10.20)).Inte- gratingtheremainingintegralbyparts asecondtime,wehave −integraldisplayb av∗′pu′dx=−v∗′puvextendsinglevextendsingleb a+integraldisplayb au(pv∗′)′dx. (10.24) Again, the integrated part vanishes in an application of Eq. (10.20). A combination of Eqs. (10.22)to(10.24)givesus integraldisplayb av∗Ludx=integraldisplayb au(Lv)∗dx. (10.25) This property, given by Eq. (10.25), is expressed by saying that the operator Lis Her- mitian with respect to the functions u(x)andv(x), which satisfy the boundary conditions specifiedbyEq.(10.20).NotethatifthisHermitianpropertyfollowsfromself-adjointness in a Hilbert space, then it includes that boundary conditions are imposed on all functions ofthatspace. 630 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Hermitian Operators in Quantum Mechanics The proceeding development in this section has focused on the classical second-order dif- ferential operators of mathematical physics. Generalizing our Hermitian operator theory as required in quantum mechanics, we have an extension: The operators need be neither second-orderdifferentialoperatorsnorreal. px=−i¯h(∂/∂x)willbeaHermitianoperator. Wesimplyassume(asiscustomaryinquantummechanics)thatthewavefunctionssatisfy appropriate boundary conditions: vanishing sufficiently strongly at infinity or having peri- odicbehavior(asinacrystallattice,orunitintensityforscatteringproblems).Theoperator Lis calledHermitian if integraldisplay ψ∗ 1Lψ2dτ=integraldisplay (Lψ1)∗ψ2dτ. (10.26) Apart from the simple extension to complex quantities, this definition is identical with Eq.(10.25). TheadjointA†of anoperator Aisdefinedby integraldisplay ψ∗ 1A†ψ2dτ≡integraldisplay (Aψ1)∗ψ2dτ. (10.27) This generalizes our classical, second-derivative-operator–oriented definition, Eq. (10.5). Here the adjoint is defined in terms of the resultant integral, with the A†as part of the integrand. Clearly, if A=A†(self-adjoint ) and satisfies the aforementioned boundary conditions,then AisHermitian. Theexpectationvalue ofanoperator Lisdefinedas /angbracketleftL/angbracketright=integraldisplay ψ∗Lψdτ. (10.28a) In the framework of quantum mechanics /angbracketleftL/angbracketrightcorresponds to the result of a measurement of the physical quantity represented by Lwhen the physical system is in a state described by the wave function ψ. If we require Lto be Hermitian, it is easy to show that /angbracketleftL/angbracketrightis real (as would be expectedfrom a measurement in a physical theory). Taking the complex conjugateof Eq.(10.28a), weobtain /angbracketleftL/angbracketright∗=bracketleftbiggintegraldisplay ψ∗Lψdτbracketrightbigg∗ =integraldisplay ψL∗ψ∗dτ. Rearrangingthefactorsintheintegrand,wehave /angbracketleftL/angbracketright∗=integraldisplay (Lψ)∗ψdτ. Then,applyingour definitionofHermitianoperator,Eq.(10.26), weget /angbracketleftL/angbracketright∗=integraldisplay ψ∗Lψdτ=/angbracketleftL/angbracketright, (10.28b) or/angbracketleftL/angbracketrightisreal.It is worthnotingthat ψis notnecessarilyaneigenfunctionof L. 10.1 Self-Adjoint ODEs 631 Exercises 10.1.1 ShowthatLaguerre’sODE,Eq.(13.52),maybeputintoself-adjointformbymultiply- ingbye−xandthatw(x)=e−xis theweightingfunction. 10.1.2 Show that the Hermite ODE, Eq. (13.10), may be put into self-adjoint form by multi- plyingby e−x2andthatthisgives w(x)=e−x2astheappropriatedensityfunction. 10.1.3 ShowthattheChebyshev(typeI)ODE,Eq.(13.100),maybeputintoself-adjointform by multiplying by (1−x2)−1/2and that this gives w(x)=(1−x2)−1/2as the appro- priatedensityfunction. 10.1.4 Show the following when the linear second-order differential equation is expressed in self-adjointform: (a) TheWronskianisequaltoa constantdividedbytheinitialcoefficient p: W(x)=C p(x). (b) Asecondsolutionis givenby y2(x)=Cy1(x)integraldisplayxdt p(t)[y1(t)]2. 10.1.5 Un(x), theChebyshevpolynomial(typeII), satisfiestheODE,Eq.(13.101), parenleftbig 1−x2parenrightbig U′′ n(x)−3xU′ n(x)+n(n+2)Un(x)=0. (a) Locate the singular points that appear in the finite plane, and show whether they areregularor irregular. (b) Putthis equationinself-adjointform. (c) Identifythecompleteeigenvalue. (d) Identifytheweightingfunction. 10.1.6 For the very special case λ=0 andq(x)=0 the self-adjoint eigenvalue equation be- comes d dxbracketleftbigg p(x)du(x) dxbracketrightbigg =0, satisfiedby du dx=1 p(x). Usethistoobtaina“second”solutionof thefollowing: (a) Legendre’sequation, (b) Laguerre’sequation, (c) Hermite’sequation. 632 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions ANS.(a) u2(x)=1 2ln1+x 1−x, (b)u2(x)−u2(x0)=integraldisplayx x0etdt t, (c)u2(x)=integraldisplayx 0et2dt. Thesesecondsolutionsillustratethedivergentbehaviorusuallyfoundinasecondsolu- tion. Note.In allthreecases u1(x)=1. 10.1.7 Given that Lu=0 andgLuis self-adjoint, show that for the adjoint operator ¯L,¯L(gu)=0. 10.1.8 Forasecond-orderdifferentialoperator Lthatis self-adjointshowthat integraldisplayb a[y2Ly1−y1Ly2]dx=p(y′ 1y2−y1y′ 2)vextendsinglevextendsingleb a. 10.1.9 Show that if a function ψis required to satisfy Laplace’s equation in a finite region of space and to satisfy Dirichlet boundary conditions over the entire closed bounding surface, then ψis unique. Hint.Oneof theforms ofGreen’stheorem,Section1.11, willbehelpful. 10.1.10 ConsiderthesolutionsoftheLegendre,Chebyshev,Hermite,andLaguerreequationsto be polynomials. Show that the ranges of integration that guarantee that the Hermitian operatorboundaryconditionswillbesatisfiedare (a) Legendre [−1,1], (b) Chebyshev [−1,1], (c) Hermite (−∞,∞), (d) Laguerre [0,∞). 10.1.11 Within the framework of quantum mechanics (Eqs. (10.26) and following), show that thefollowingareHermitianoperators: (a) momentum p=−i¯h∇≡−ih 2π∇ (b) angularmomentum L=−i¯hr×∇≡−ih 2πr×∇. Hint. In Cartesian form Lis a linear combination of noncommuting Hermitian opera- tors. 10.1.12 (a)Ais anon-Hermitianoperator.In thesenseof Eqs. (10.26)and(10.27), showthat A+A†andi(A−A†) areHermitianoperators. (b) Usingtheprecedingresult,showthateverynon-Hermitianoperatormaybewritten asalinearcombinationoftwoHermitianoperators. 10.1 Self-Adjoint ODEs 633 10.1.13 UandVare two arbitrary operators, not necessarily Hermitian. In the sense of Eq. (10.27),showthat (UV)†=V†U†. NotetheresemblancetoHermitianadjointmatrices. Hint.Applythedefinitionofadjointoperator,Eq. (10.27). 10.1.14 ProvethattheproductoftwoHermitianoperatorsisHermitian(Eq.(10.26))ifandonly ifthetwooperatorscommute. 10.1.15 AandBarenoncommutingquantummechanicaloperators: AB−BA=iC. Showthat Cis Hermitian.Assumethatappropriateboundaryconditionsaresatisfied. 10.1.16 Theoperator LisHermitian.Showthat /angbracketleftL2/angbracketright≥0. 10.1.17 Aquantummechanicalexpectationvalueis definedby /angbracketleftA/angbracketright=integraldisplay ψ∗(x)Aψ(x)dx, whereAis a linear operator. Show that demanding that /angbracketleftA/angbracketrightbe real means that Amust beHermitian—withrespectto ψ(x). 10.1.18 From the definition of adjoint, Eq. (10.27), show that A††=Ain the sense thatintegraltext ψ∗ 1A††ψ2dτ=integraltext ψ∗ 1Aψ2dτ.Theadjointoftheadjointistheoriginaloperator. Hint. The functions ψ1andψ2of Eq. (10.27) represent a class of functions. The sub- scripts1and2maybeinterchangedorreplacedbyothersubscripts. 10.1.19 TheSchrödingerwaveequationforthedeuteron(withaWoods–Saxonpotential)is −¯h2 2M∇2ψ+V0 1+exp[(r−r0)/a]ψ=Eψ. HereE=−2.224 MeV, ais a “thickness parameter,” 0 .4×10−13cm. Expressing lengths in fermis (10−13cm) and energies in million electron volts (MeV), we may rewritethewaveequationas d2 dr2(rψ)+1 41.47bracketleftbigg E−V0 1+exp((r−r0)/a)bracketrightbigg (rψ)=0. Eis assumed known from experiment. The goal is to find V0for a specified value of r0(say,r0=2.1). If we let y(r)=rψ(r), theny(0)=0 and we take y′(0)=1. Find V0such that y(20.0)=0. (This should be y(∞),b u tr=20 is far enough beyond the rangeofnuclearforces toapproximateinfinity.) ANS.For a=0.4 andr0=2.1f m ,V0=−34.159 MeV. 10.1.20 Determine the nuclear potential well parameter V0of Exercise 10.1.19 as a function of r0forr=2.00(0.05)2.25 fermis. Express yourresultsas apowerlaw |V0|rν 0=k. Determine the exponent νand the constant k. This power-law formulation is useful for accurateinterpolation. 634 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.1.21 InExercise10.1.19itwasassumedthat20fermiswasagoodapproximationtoinfinity. Checkonthisbycalculating V0forrψ(r)=0a t( a )r=15,(b)r=20,(c)r=25,and (d)r=30.Sketchyourresults. Take r0=2.10 anda=0.4 (fermis). 10.1.22 For a quantum particle moving in a potential well, V(x)=1 2mω2x2, the Schrödinger waveequationis −¯h2 2md2ψ(x) dx2+1 2mω2x2ψ(x)=Eψ(x), or d2ψ(z) dz2−z2ψ(z)=−2E ¯hωψ(z), wherez=(mω/¯h)1/2x. Since this operator is even, we expect solutions of definite parity.Fortheinitialconditionsthatfollow,integrateoutfromtheoriginanddetermine the minimum constant 2 E/¯hωthat will lead to ψ(∞)=0 in each case. (You may take z=6 asanapproximationofinfinity.) (a) Foraneveneigenfunction, ψ(0)=1,ψ′(0)=0. (b) Foranoddeigenfunction, ψ(0)=0,ψ′(0)=1. Note.AnalyticalsolutionsappearinSection13.1. 10.2 H ERMITIAN OPERATORS Hermitian,orself-adjoint,operatorswithappropriateboundaryconditionshavethreeprop- ertiesthatare ofextremeimportanceinphysics, bothclassicalandquantum. 1. TheeigenvaluesofaHermitianoperatorare real. 2. AHermitianoperatorpossessesanorthogonalsetofeigenfunctions. 3. TheeigenfunctionsofaHermitianoperatorform acompleteset.6 Real Eigenvalues Weproceedtoprovethefirsttwo ofthesethreeproperties. Let Lui+λiwui=0. (10.29) 6This third property is not universal. It doeshold for our linear, second-order differential operators in Sturm–Liouville (self- adjoint)form.CompletenessisdefinedanddiscussedinSection10.4.Aproofthattheeigenfunctionsofourlinear,second-order, self-adjoint, differential equations form a complete setmay be developed from the calculus ofvariations of Section 17.8. 10.2 Hermitian Operators 635 Assumingtheexistenceof asecondeigenvalueandeigenfunction, Luj+λjwuj=0. (10.30) Then,takingthecomplexconjugate,weobtain L∗u∗ j+λ∗jwu∗j=0. (10.31) Herew(x)≥0 is a real function. But we permit λk, the eigenvalues, and uk, the eigen- functions, to be complex. Multiplying Eq. (10.29) by u∗ jand Eq. (10.31) by uiand then subtracting,wehave u∗ jLui−uiL∗u∗j=(λ∗j−λi)wuiu∗j. (10.32) Weintegrateovertherange a≤x≤b: integraldisplayb au∗jLuidx−integraldisplayb auiL∗u∗jdx=(λ∗j−λi)integraldisplayb auiu∗jwdx. (10.33) SinceLis Hermitian,theleft-handsidevanishesbyEq. (10.26)and (λ∗j−λi)integraldisplayb auiu∗jwdx=0. (10.34) Ifi=j, the integral cannot vanish [ w(x)>0, apart from isolated points], except in the trivialcase ui=0.Hencethecoefficient (λ∗i−λi)mustbezero, λ∗i=λi, (10.35) which says that the eigenvalue is real. Since λican represent any one of the eigenvalues, thisprovesthefirstproperty.Thisisanexactanalogofthenatureoftheeigenvaluesofreal symmetric(andofHermitian)matrices(compareSection3.5). The analog of the spectral decomposition of a real symmetric matrix in Section 3.5 for aHermitianoperator Lwithadiscreteset ofeigenvalues λitakestheform L=summationdisplay iλi|ui/angbracketright/angbracketleftui|,f(L)=summationdisplay if(λi)|ui/angbracketright/angbracketleftui| witheigenvectors |ui/angbracketrightandanyinfinitelydifferentiablefunction f. Real eigenvalues of Hermitian operators have a fundamental significance in quantum mechanics. In quantum mechanics the eigenvalues correspond to precisely measurable quantities, such as energy and angular momentum. With the theory formulated in terms of Hermitian operators, this proof of real eigenvalues guarantees that the theory will pre- dict real numbers for these measurable physical quantities. In Section 17.8 it will be seen thatthesetof realeigenvalueshasalowerbound(for nonrelativisticproblems). 636 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Orthogonal Eigenfunctions If we now take i/negationslash=jand ifλi/negationslash=λjin Eq. (10.34), the integral of the product of the two differenteigenfunctionsmustvanish: integraldisplayb auiu∗ jwdx=0. (10.36) This condition, called orthogonality , is the continuum analog of the vanishing of a scalar product of two vectors.7We say that the eigenfunctions ui(x)anduj(x)are orthogonal withrespecttotheweightingfunction w(x)overtheinterval [a,b].Equation(10.36)con- stitutesapartialproofofthesecondpropertyofourHermitianoperators.Again,theprecise analogywithmatrixanalysisshouldbenoted.Indeed,wecanestablishaone-to-onecorre- spondencebetweenthisSturm–Liouvilletheoryofdifferentialequationsandthetreatment ofHermitianmatrices.Historically,thiscorrespondencehasbeensignificantinestablishing themathematicalequivalenceofmatrixmechanicsdevelopedbyHeisenbergandwaveme- chanics developed by Schrödinger. Today, the two diverse approaches are merged into the theory of quantum mechanics, and the mathematical formulation that is more convenient foraparticularproblemisusedforthatproblem.Actuallythemathematicalalternativesdo not end here. Integral equations, Chapter 16, form a third equivalent and sometimes more convenientor morepowerfulapproach. This proof of orthogonality is not quite complete. There is a loophole, because we may haveui/negationslash=ujbut still have λi=λj. Such a case is labeled degenerate . Illustrations of degeneracy are given at the end of this section. If λi=λj, the integral in Eq. (10.34) need notvanish.Thismeansthatlinearlyindependenteigenfunctionscorrespondingtothesame eigenvalue are not automatically orthogonal and that some other method must be sought to obtain an orthogonal set. Although the eigenfunctions in this degenerate case may not be orthogonal, they can always be made orthogonal. One method is developed in the next section.SeealsoEq. (4.21)for degeneracyduetosymmetry. We shall see in succeeding chapters that it is just as desirable to have a given set of functions orthogonal as it is to have an orthogonal coordinate system. We can work with nonorthogonal functions, but they are likely to prove as messy as an oblique coordinate system. Example 10.2.1 FOURIER SERIES —O RTHOGONALITY TocontinueExample10.1.3, theeigenvalueequation,Eq. (10.21), d2 dx2y(x)+n2y(x)=0, 7From the definition of Riemann integral, integraldisplayb af(x)g(x)dx=lim N→∞parenleftbiggNsummationdisplay i=1f(xi)g(xi)/Delta1xparenrightbigg , wherex0=a,xN=b,a n dxi−xi−1=/Delta1x. If we interpret f(xi)andg(xi)as theith components of an N-component vector, then this sum (and therefore this integral) corresponds directly to a scalar product of vectors, Eq. (1.24). The vanishing of the scalarproduct is thecondition for orthogonality of thevectors—or functions. 10.2 Hermitian Operators 637 may describe a quantum mechanical particle in a box, or perhaps a vibrating violin string, a classical harmonic oscillator with degenerate eigenfunctions—cos nx,sinnx— andeigenvalues n2,naninteger. Withnreal(here takentobeintegral),theorthogonalityintegralsbecome (a)integraldisplayx0+2π x0sinmxsinnxdx=Cnδnm, (b)integraldisplayx0+2π x0cosmxcosnxdx=Dnδnm, (c)integraldisplayx0+2π x0sinmxcosnxdx=0. For an interval of 2 πthe preceding analysis guarantees the Kronecker delta in (a) and (b) but not the zero in (c) because (c) may involve degenerate eigenfunctions. However, inspectionshows that(c) alwaysvanishesfor allintegral mandn. OurSturm–Liouvilletheorysaysnothingaboutthevaluesof CnandDnbecausehomo- geneousODEshavesolutionswhosescalingis arbitrary.Actualcalculationyields Cn=braceleftBiggπ, n/negationslash=0, 0,n=0,Dn=braceleftBiggπ, n/negationslash=0, 2π, n=0. These orthogonality integrals form the basis of the Fourier series developed in Chap- ter14. /squaresolid Example 10.2.2 EXPANSION IN ORTHOGONAL EIGENFUNCTIONS —SQUARE WAVE Thepropertyofcompleteness(seeEq.(1.190)andSection10.4)meansthatcertainclasses of functions (for example, sectionally or piecewise continuous) may be represented by a seriesof orthogonaleigenfunctions.Considerthesquare-waveshape f(x)=  h 2,0<x<π, −h 2,−π<x<0.(10.37) Thisfunctionmaybeexpandedinanyofavarietyofeigenfunctions—Legendre,Hermite, Chebyshev,andsoon.Thechoiceofeigenfunctionismadeonthebasisofconvenienceor an application. To illustrate the expansion technique, let us choose the eigenfunctions of Example10.2.1, cos nxand sinnx. Theeigenfunctionseries isconveniently(andconventionally)writtenas f(x)=a0 2+∞summationdisplay m=1(amcosmx+bmsinmx). Upon multiplying f(t)by cosntor sinntand integrating, only the nth term survives, by theorthogonalityintegralsofExample10.2.1, thusyieldingthecoefficients an=1 πintegraldisplayπ −πf(t)cosntdt, b n=1 πintegraldisplayπ −πf(t)sinntdt, n=0,1,2.... 638 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Directsubstitutionof ±h/2f o rf(t)yields an=0, whichisexpectedherebecauseof theantisymmetry, f(−x)=−f(x), and bn=h nπ(1−cosnπ)=  0,n even, 2h nπ,nodd. Hencetheeigenfunction(Fourier)expansionofthesquarewaveis f(x)=2h π∞summationdisplay n=0sin(2n+1)x 2n+1. (10.38) Additionalexamples,usingothereigenfunctions,appearinChapters11and12. /squaresolid Degeneracy The concept of degeneracy was introduced earlier. If Nlinearly independent eigenfunc- tions correspond to the same eigenvalue, the eigenvalue is said to be N-fold degenerate. A particularly simple illustration is provided by the eigenvalues and eigenfunctions of the classical harmonic oscillator equation, Example 10.2.1. For each eigenvalue n2, there are two possible solutions: sin nxand cosnx(and any linear combination, nan integer). We saytheeigenfunctionsaredegenerateor theeigenvalueis degenerate. A more involved example is furnished by the physical system of an electron in an atom (nonrelativistictreatment,spinneglected).FromtheSchrödingerequation,Eq.(13.84)for hydrogen,thetotalenergyoftheelectronisoureigenvalue.Wemaylabelit EnLMbyusing thequantumnumbers n,L,andMassubscripts.Foreachdistinctsetofquantumnumbers (n,L,M) thereisadistinct,linearlyindependenteigenfunction ψnLM(r,θ,ϕ).Forhydro- gen, the energy EnLMis independent of LandM, reflecting the spherical (and SO(4)) symmetry of the Coulomb potential. With 0 ≤L≤n−1 and−L≤M≤L, the eigen- value isn2-fold degenerate (including the electron spin would raise this to 2 n2). In atoms withmorethanoneelectron,theelectrostaticpotentialisnolongerasimple r−1potential. Theenergydependson Laswellason n,although notonM;EnLMisstill(2L+1)-fold degenerate. This degeneracy—due to rotational invariance of the potential—may be re- movedbyapplyinganexternalmagneticfield,breakingsphericalsymmetryandgivingrise totheZeemaneffect.Asarule,theeigenfunctionsformaHilbertspace,thatis,acomplete vector space of functions with a metric defined by the inner product (see Section 10.4 for moredetailsandexamples). Often an underlying symmetry, such as rotational invariance, is causing the degenera- cies. States belonging to the same energy eigenvalue then will form a multiplet or repre- sentation of the symmetry group. The powerful group-theoretical methods are treated in Chapter4insomedetail. 10.2 Hermitian Operators 639 Exercises 10.2.1 The functions u1(x)andu2(x)are eigenfunctions of the same Hermitian operator but fordistincteigenvalues λ1andλ2.Provethat u1(x)andu2(x)arelinearlyindependent. 10.2.2 (a) Thevectors enareorthogonaltoeachother: en·em=0forn/negationslash=m.Showthatthey arelinearlyindependent. (b) Thefunctions ψn(x)areorthogonaltoeachotherovertheinterval [a,b]andwith respecttotheweightingfunction w(x).Showthatthe ψn(x)arelinearlyindepen- dent. 10.2.3 Giventhat P1(x)=xandQ0(x)=1 2lnparenleftbigg1+x 1−xparenrightbigg are solutions of Legendre’s differential equation corresponding to different eigenval- ues: (a) Evaluatetheirorthogonalityintegral integraldisplay1 −1x 2lnparenleftbigg1+x 1−xparenrightbigg dx. (b) Explain why these two functions are not orthogonal, that is, why the proof of orthogonalitydoesnotapply. 10.2.4 T0(x)=1andV1(x)=(1−x2)1/2aresolutionsoftheChebyshevdifferentialequation corresponding to different eigenvalues. Explain, in terms of the boundary conditions, whythesetwofunctionsarenotorthogonal. 10.2.5 (a) Show that the first derivatives of the Legendre polynomials satisfy a self-adjoint differentialequationwitheigenvalue λ=n(n+1)−2. (b) ShowthattheseLegendrepolynomialderivativessatisfyanorthogonalityrelation integraldisplay1 −1P′ m(x)P′ n(x)parenleftbig 1−x2parenrightbig dx=0,m/negationslash=n. Note.InSection12.5, (1−x2)1/2P′ n(x)willbelabeledanassociatedLegendrepolyno- mial,P1 n(x). 10.2.6 Asetof functions un(x)satisfiestheSturm–Liouvilleequation d dxbracketleftbigg p(x)d dxun(x)bracketrightbigg +λnw(x)un(x)=0. The functions um(x)andun(x)satisfy boundary conditions that lead to orthogonality. Thecorrespondingeigenvalues λmandλnaredistinct.Provethatforappropriatebound- aryconditions, u′ m(x)andu′n(x)areorthogonalwith p(x)asaweightingfunction. 640 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.2.7 A linear operator Ahasndistinct eigenvalues and ncorresponding eigenfunctions: Aψi=λiψi. Show that the neigenfunctions are linearly independent. Ais not neces- sarilyHermitian. Hint. Assume linear dependence—that ψn=summationtextn−1 i=1aiψi. Use this relation and the operator–eigenfunction equation first in one order and then in the reverse order. Show thata contradictionresults. 10.2.8 (a) ShowthattheLiouvillesubstitution u(x)=v(ξ)bracketleftbig p(x)w(x)bracketrightbig−1/4,ξ=integraldisplayx abracketleftbiggw(t) p(t)bracketrightbigg1/2 dt transforms d dxbracketleftbigg p(x)d dxubracketrightbigg +bracketleftbig λw(x)−q(x)bracketrightbig u(x)=0 into d2v dξ2+bracketleftbig λ−Q(ξ)bracketrightbig v(ξ)=0, where Q(ξ)=q(x(ξ)) w(x(ξ))+bracketleftbig pparenleftbig x(ξ)parenrightbig wparenleftbig x(ξ)parenrightbigbracketrightbig−1/4d2 dξ2(pw)1/4. (b) Ifv1(ξ)andv2(ξ)areobtainedfrom u1(x)andu2(x),respectively,byaLiouville substitution,showthatintegraltextb aw(x)u1u2dxistransformedintointegraltextc 0v1(ξ)v2(ξ)dξwith c=integraltextb a[w p]1/2dx. 10.2.9 Theultrasphericalpolynomials C(α) n(x)aresolutionsofthedifferentialequation braceleftbigg (1−x2)d2 dx2−(2α+1)xd dx+n(n+2α)bracerightbigg C(α) n(x)=0. (a) Transformthisdifferentialequationintoself-adjointform. (b) Show that the C(α) n(x)are orthogonal for different n. Specify the interval of inte- grationandtheweightingfactor. Note.Assumethatyoursolutionsarepolynomials. 10.2.10 WithLnotself-adjoint, Lui+λiwui=0 and ¯Lvj+λjwvj=0. 10.2 Hermitian Operators 641 (a) Showthat integraldisplayb avjLuidx=integraldisplayb aui¯Lvjdx, provided uip0v′ jvextendsinglevextendsingleb a=vjp0u′ ivextendsinglevextendsingleb a and ui(p1−p′ 0)vjvextendsinglevextendsingleb a=0. (b) Showthattheorthogonalityintegralfor theeigenfunctions uiandvjbecomes integraldisplayb auivjwdx=0(λi/negationslash=λj). 10.2.11 InExercise9.5.8theseriessolutionoftheChebyshevequationisfoundtobeconvergent foralleigenvalues n.Therefore nisnotquantizedbytheargumentusedforLegendre’s (Exercise 9.5.5). Calculate the sum of the indicial equation k=0 Chebyshev series for n=v=0.8,0.9,and1.0andfor x=0.0(0.1)0.9. Note.TheChebyshevseries recurrencerelationisgiveninExercise5.2.16. 10.2.12 (a) Evaluate the n=ν=0.9, indicial equation k=0 Chebyshev series for x= 0.98,0.99,and1.00.Theseries convergesveryslowlyat x=1.00.Youmaywish to use double precision. Upper bounds to the error in your calculation can be set bycomparisonwiththe ν=1.0 case,whichcorrespondsto (1−x2)1/2. (b) These series solutions for eigenvalue ν=0.9 and for ν=1.0 are obviously not orthogonal,despitethefactthattheysatisfyaself-adjointeigenvalueequationwith differenteigenvalues.Fromthebehaviorofthesolutionsinthevicinityof x=1.00 trytoformulateahypothesisastowhytheproofoforthogonalitydoesnotapply. 10.2.13 The Fourier expansion of the (asymmetric) square wave is given by Eq. (10.38). With h=2,evaluatethisseriesfor x=0(π/18)π/2,usingthefirst(a)10terms,(b)100terms oftheseries. Note.For10termsand x=π/18,or10◦,yourFourierrepresentationhasasharphump. ThisistheGibbsphenomenonofSection14.5.For100termsthishumphasbeenshifted overtoabout1◦. 10.2.14 Thesymmetric squarewave f(x)=  1,|x|<π 2 −1,π 2<|x|<π hasaFourierexpansion f(x)=4 π∞summationdisplay n=0(−1)ncos(2n+1)x 2n+1. Evaluatethisseries for x=0(π/18)π/2u s i n gt h efi r s t 642 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions (a) 10terms, (b)100termsof theseries. Note. As in Exercise 10.2.13, the Gibbs phenomenon appears at the discontinuity. This means that a Fourier series is not suitable for precise numerical work in the vicinity of adiscontinuity. 10.3 G RAM –SCHMIDT ORTHOGONALIZATION The Gram–Schmidt orthogonalization is a method that takes a nonorthogonal set of lin- early independent vectors (see Section 3.1) or functions8and constructs an orthogonal set of vectors or functions over an arbitrary interval and with respect to an arbitrary weight or density factor. In the language of linear algebra, the process is equivalent to a matrix transformation relating an orthogonal set of basis vectors (functions) to a nonorthogonal set.AspecificexampleofthismatrixtransformationappearsinExercise12.2.1. Next we apply the Gram–Schmidt procedure to a set of functions. The functions in- volved may be real or complex. Here for convenience they are assumed to be real. The generalizationtothecomplexcaseoffers nodifficulty. Before taking up orthogonalization, we should consider normalization of functions. So far nonormalizationhasbeenspecified.This meansthat integraldisplayb aϕ2 iwdx=N2 i, but no attention has been paid to the value of Ni. Since our basic equation (Eq. (10.8)) is linear and homogeneous, we may multiply our solution by any constant and it will still be asolution.Wenowdemandthateachsolution ϕi(x)bemultipliedby N−1 isothatthenew (normalized) ϕiwillsatisfy integraldisplayb aϕ2 i(x)w(x)dx=1 (10.39) and integraldisplayb aϕi(x)ϕj(x)w(x)dx=δij. (10.40) Equation (10.39) says that we have normalized to unity. Including the property of orthog- onality, we have Eq. (10.40). Functions satisfying this equation are said to be orthonor- mal(orthogonalplusunitnormalization).Othernormalizationsarecertainlypossible,and indeed, by historical convention, each of the special functions of mathematical physics treatedinChapters12and13willbenormalizeddifferently. We consider three sets of functions: an original, linearly independent given set un(x), n=0,1,2,...; an orthogonalized set ψn(x)to be constructed; and a final set 8Such a set of functions might well arise from the solutions of a PDE in which the eigenvalue was independent of one or more of the constants of separation. As an example, we have the hydrogen atom problem (Sections 10.2 and 13.2). The eigenvalue (energy) is independent of both the electron orbital angular momentum and its projection on the z-axis,m. Note, however, that the origin of the setof functions is irrelevant to the Gram–Schmidt orthogonalization procedure. 10.3 Gram–Schmidt Orthogonalization 643 of functions ϕn(x), which are the normalized ψn. The original unmay be degenerate eigenfunctions,butthisis notnecessary.We shallhavethefollowingproperties: un(x) ψ n(x) ϕ n(x) Linearlyindependent Linearlyindependent Linearlyindependent Nonorthogonal Orthogonal Orthogonal Unnormalized Unnormalized Normalized (orthonormal) The Gram–Schmidt procedure takes the nthψfunction(ψn)to beun(x)plus an un- knownlinearcombinationoftheprevious ϕ.Thepresenceofthenew un(x)willguarantee linear independence. The requirement that ψn(x)be orthogonal to each of the previous ϕyields just enough constraints to determine each of the unknown coefficients. Then the fully determined ψnwill be normalized to unity, yielding ϕn(x). Then the sequence of stepsis repeatedfor ψn+1(x). We startwith n=0,letting ψ0(x)=u0(x), (10.41) withno“previous” ϕtoworry about.Thenwenormalize ϕ0(x)=ψ0(x) [integraltext ψ2 0wdx]1/2. (10.42) Forn=1,let ψ1(x)=u1(x)+a1,0ϕ0(x). (10.43) We demand that ψ1(x)be orthogonal to ϕ0(x). (At this stage the normalization of ψ1(x) isirrelevant.) Thisorthogonalityleadsto integraldisplay ψ1ϕ0wdx=integraldisplay u1ϕ0wdx+a1,0integraldisplay ϕ2 0wdx=0. (10.44) Sinceϕ0is normalizedtounity(Eq. (10.42)), wehave a1,0=−integraldisplay u1ϕ0wdx, (10.45) fixingthevalueof a1,0.Normalizing,wedefine ϕ1(x)=ψ1(x) (integraltext ψ2 1wdx)1/2. (10.46) Finally,wegeneralizeso that ϕi(x)=ψi(x) (integraltext ψ2 i(x)w(x)dx)1/2, (10.47) where ψi(x)=ui+ai,0ϕ0+ai,1ϕ1+···+ai,i−1ϕi−1. (10.48) Thecoefficients ai,jaregivenby ai,j=−integraldisplay uiϕjwdx. (10.49) 644 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Equation(10.49)holdsfor unitnormalization.If someothernormalizationis selected, integraldisplayb abracketleftbig ϕj(x)bracketrightbig2w(x)dx=N2 j, thenEq. (10.47)is replacedby ϕi(x)=Niψi(x) (integraltext ψ2 iwdx)1/2. (10.47a) andai,jbecomes ai,j=−integraltext uiϕjwdx N2 j. (10.49a) Equations (10.48) and (10.49) may be rewritten in terms of projection operators, Pj.I f we consider the ϕn(x)to form a linear vector space, then the integral in Eq. (10.49) may beinterpretedastheprojectionof uiintotheϕj“coordinate,”orthe jthcomponentof ui. With Pjui(x)=braceleftbiggintegraldisplay ui(t)ϕj(t)w(t)dtbracerightbigg ϕj(x), Eq.(10.48) becomes ψi(x)=braceleftbigg 1−i−1summationdisplay j=1Pjbracerightbigg ui(x). (10.48a) Subtractingoff thecomponents, j=1t oi−1,leaves ψi(x)orthogonaltoallthe ϕj(x). It will be noticed that although this Gram–Schmidt procedure is one possible way of constructing an orthogonal or orthonormal set, the functions ϕi(x)are not unique. There is an infinite number of possible orthonormal sets for a given interval and a given density function. As an illustration of the freedom involved, consider two (nonparallel) vectors AandB in thexy-plane. We may normalize Ato unit magnitude and then form B′=aA+Bso thatB′is perpendicular to A. By normalizing B′we have completed the Gram–Schmidt orthogonalizationfortwovectors.Butanytwoperpendicularunitvectors,suchas ˆxandˆy, couldhavebeenchosenasourorthonormalset.Again,withaninfinitenumberofpossible rotations ofˆxandˆyabout the z-axis, we have an infinite number of possible orthonormal sets. Example 10.3.1 LEGENDRE POLYNOMIALS BY GRAM–SCHMIDT ORTHOGONALIZATION Let us form an orthonormal set from the set of functions un(x)=xn,n=0,1,2....T h e intervalis−1≤x≤1 andthedensityfunctionis w(x)=1. In accordancewiththeGram–Schmidtorthogonalizationprocessdescribed, u0=1,hence ϕ0=1√ 2. (10.50) 10.3 Gram–Schmidt Orthogonalization 645 Then ψ1(x)=x+a1,01√ 2(10.51) and a1,0=−integraldisplay1 −1x√ 2dx=0 (10.52) bysymmetry.We normalize ψ1toobtain ϕ1(x)=radicalbigg 3 2x. (10.53) ThenwecontinuetheGram–Schmidtprocedurewith ψ2(x)=x2+a2,01√ 2+a2,1radicalbigg 3 2x, (10.54) where a2,0=−integraldisplay1 −1x2 √ 2dx=−√ 2 3, (10.55) a2,1=−integraldisplay1 −1radicalbigg 3 2x3dx=0, (10.56) againbysymmetry.Therefore ψ2(x)=x2−1 3, (10.57) and,onnormalizingtounity,wehave ϕ2(x)=radicalbigg 5 2·1 2parenleftbig 3x2−1parenrightbig . (10.58) Thenextfunction, ϕ3(x), becomes ϕ3(x)=radicalbigg 7 2·1 2parenleftbig 5x3−3xparenrightbig . (10.59) ReferencetoChapter12willshowthat ϕn(x)=radicalbigg 2n+1 2Pn(x), (10.60) wherePn(x)is thenth-order Legendre polynomial. Our Gram–Schmidt process provides a possible but very cumbersome method of generating the Legendre polynomials. It il- lustrates how a power-series expansion in un(x)=xn, which is not orthogonal, can be convertedintoanorthogonalseries. /squaresolid 646 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions The equations for Gram–Schmidt orthogonalization tend to be ill-conditioned because ofthesubtractions,Eqs.(10.48)and(10.49).Atechniqueforavoidingthisdifficultyusing thepolynomialrecurrencerelationis discussedbyHamming.9 InExample10.3.1wehavespecifiedanorthogonalityinterval [−1,1],aunitweighting function, and a set of functions xnto be taken one at a time in increasing order. Given all these specifications, the Gram–Schmidt procedure is unique (to within a normaliza- tion factor and an overall sign, as discussed subsequently). Our resulting orthogonal set, the Legendre polynomials, P0up through Pn, form a complete set for the description of polynomials of order ≤nover[−1,1]. This concept of completeness is taken up in detail in Section 10.4. Expansions of functions in series of Legendre polynomials are found in Section12.3. Orthogonal Polynomials Example 10.3.1 has been chosen strictly to illustrate the Gram–Schmidt procedure. Al- though it has the advantage of introducing the Legendre polynomials, the initial functions un=xnare not degenerate eigenfunctions and are not solutions of Legendre’s equation. They are simply a set of functions that we have here rearranged to create an orthonor- mal set for the given interval and given weighting function. The fact that we obtained the Legendrepolynomialsisnotquiteblackmagicbutadirectconsequenceofthechoiceofin- tervalandweightingfunction.Theuseof un(x)=xnbutwithotherchoicesofintervaland Table 10.3 OrthogonalPolynomialsGeneratedbyGram–SchmidtOrthogonalization ofun(x)=xn,n=0,1,2,... Weighting Polynomials Interval function w(x) Standard normalization Legendre −1≤x≤11integraldisplay1 −1[Pn(x)]2dx=2 2n+1 Shifted Legendre 0 ≤x≤11integraldisplay1 0[P∗ n(x)]2dx=1 2n+1 Chebyshev I −1≤x≤1(1−x2)−1/2integraldisplay1 −1[Tn(x)]2 (1−x2)1/2dx=braceleftbiggπ/2,n/negationslash=0 π, n=0 Shifted Chebyshev I 0 ≤x≤1[x(1−x)]−1/2integraldisplay1 0[T∗n(x)]2 [x(1−x)]1/2dx=braceleftbiggπ/2,n>0 π, n=0 Chebyshev II −1≤x≤1(1−x2)1/2integraldisplay1 −1[Un(x)]2(1−x2)1/2dx=π 2 Laguerre 0 ≤x<∞ e−xintegraldisplay∞ 0[Ln(x)]2e−xdx=1 AssociatedLaguerre 0 ≤x<∞ xke−xintegraldisplay∞ 0[Lk n(x)]2xke−xdx=(n+k)! n! Hermite −∞<x<∞ e−x2integraldisplay∞ −∞[Hn(x)]2e−x2dx=2nπ1/2n! 9R. W. Hamming, Numerical Methods for Scientists and Engineers , 2nd ed.,NewYork: McGraw-Hill (1973). See Section 27.2 andreferences given there. 10.3 Gram–Schmidt Orthogonalization 647 weighting function leads to other sets of orthogonal polynomials, as shown in Table 10.3. We consider these polynomials in detail in Chapters 12 and 13 as solutions of particular differentialequations. Anexaminationofthisorthogonalizationprocesswillrevealtwoarbitraryfeatures.First, asemphasizedbefore,itisnotnecessarytonormalizethefunctionstounity.Intheexample justgivenwecouldhaverequired integraldisplay1 −1ϕn(x)ϕm(x)dx=2 2n+1δnm, (10.61) andtheresultingsetwouldhavebeentheactualLegendrepolynomials.Second,thesignof ϕnis always indeterminate. In the example we chose the sign by requiring the coefficient of the highest power of xin the polynomial to be positive. For the Laguerre polynomials, ontheotherhand,wewouldrequirethecoefficientof thehighestpowertobe (−1)n/n! Exercises 10.3.1 Rework Example 10.3.1 by replacing ϕn(x)by the conventional Legendre polynomial, Pn(x): integraldisplay1 −1bracketleftbig Pn(x)bracketrightbig2dx=2 2n+1. UsingEqs. (10.47a), and(10.49a), construct P0,P1(x), andP2(x). ANS.P0=1,P1=x,P2=3 2x2−1 2. 10.3.2 Following the Gram–Schmidt procedure, construct a set of polynomials P∗ n(x)orthog- onal (unit weighting factor) over the range [0,1]from the set[1,x]. Normalize so that P∗ n(1)=1. ANS.P∗ n(x)=1, P∗ 1(x)=2x−1, P∗ 2(x)=6x2−6x+1, P∗ 3(x)=20x3−30x2+12x−1. Thesearethefirstfour shiftedLegendrepolynomials. Note.The“*”isthestandardnotationfor“shifted”: [0,1]insteadof[−1,1].Itdoesnot meancomplexconjugate. 10.3.3 ApplytheGram–Schmidtproceduretoform thefirst threeLaguerrepolynomials un(x)=xn,n=0,1,2,..., 0≤x<∞,w(x)=e−x. Theconventionalnormalizationis integraldisplay∞ 0Lm(x)Ln(x)e−xdx=δmn. ANS.L0=1,L1=(1−x),L2=2−4x+x2 2. 648 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.3.4 You are given (a) asetof functions un(x)=xn,n=0,1,2,..., (b) aninterval (0,∞), (c) aweightingfunction w(x)=xe−x.UsetheGram–Schmidtproceduretoconstruct the firstthree orthonormal functions from the set un(x)for this interval and this weightingfunction. ANS.ϕ0(x)=1,ϕ1(x)=(x−2)/√ 2,ϕ2(x)=parenleftbig x2−6x+6parenrightbig /2√ 3. 10.3.5 Using the Gram–Schmidt orthogonalization procedure, construct the lowest three Her- mitepolynomials: un(x)=xn,n=0,1,2,...,−∞<x<∞,w(x)=e−x2. Forthissetof polynomialstheusualnormalizationis integraldisplay∞ −∞Hm(x)Hn(x)w(x)dx=δmn2mm!π1/2. ANS.H0=1,H1=2x,H2=4x2−2. 10.3.6 UsetheGram–SchmidtorthogonalizationschemetoconstructthefirstthreeChebyshev polynomials(typeI): un(x)=xn,n=0,1,2,...,−1≤x≤1,w(x)=parenleftbig 1−x2parenrightbig−1/2. Takethenormalization integraldisplay1 −1Tm(x)Tn(x)w(x)dx=δmn  π, m=n=0, π 2,m=n≥1. Hint.TheneededintegralsaregiveninExercise8.4.3. ANS.T0=1,T1=x,T2=2x2−1(T3=4x3−3x). 10.3.7 UsetheGram–SchmidtorthogonalizationschemetoconstructthefirstthreeChebyshev polynomials(typeII): un(x)=xn,n=0,1,2,...,−1≤x≤1,w(x)=parenleftbig 1−x2parenrightbig+1/2. Takethenormalizationtobe integraldisplay1 −1Um(x)Un(x)w(x)dx=δmnπ 2. Hint. integraldisplay1 −1parenleftbig 1−x2parenrightbig1/2x2ndx=π 2×1·3·5···(2n−1) 4·6·8···(2n+2),n=1,2,3,... =π 2,n=0. ANS.U0=1,U1=2x,U2=4x2−1. 10.4 Completeness of Eigenfunctions 649 10.3.8 As a modification of Exercise 10.3.5, apply the Gram–Schmidt orthogonalization pro- cedure to the set un(x)=xn,n=0,1,2,...,0≤x<∞.T a k ew(x)to be exp[−x2]. Find the first two nonvanishing polynomials. Normalize so that the coefficient of the highestpowerof xisunity.InExercise10.3.5theinterval (−∞,∞)ledtotheHermite polynomials.ThesearecertainlynottheHermitepolynomials. ANS.ϕ0=1,ϕ1=x−π−1/2. 10.3.9 Form an orthogonal set over the interval 0 ≤x<∞,u s i n gun(x)=e−nx,n= 1,2,3,....Take the weighting factor, w(x), to be unity. These functions are solutions ofu′′ n−n2un=0,whichisclearlyalreadyinSturm–Liouville(self-adjoint)form.Why doesn’ttheSturm–Liouvilletheoryguaranteetheorthogonalityof thesefunctions? 10.4 C OMPLETENESS OF EIGENFUNCTIONS The third important property of an Hermitian operator is that its eigenfunctions form a complete set. This completeness means that any well-behaved (at least piecewise continu- ous)function F(x)canbeapproximatedbyaseries F(x)=∞summationdisplay n=0anϕn(x) (10.62) to any desired degree of accuracy.10More precisely, the set ϕn(x)is calledcomplete11if thelimitofthemeansquareerror vanishes: limm→∞integraldisplayb abracketleftbigg F(x)−msummationdisplay n=0anϕn(x)bracketrightbigg2 w(x)dx=0. (10.63) Technically, the integral here is a Lebesgue integral. We have not required that the error vanishidenticallyin [a,b]butonlythattheintegralof theerror squaredgotozero. This convergence in the mean, Eq. (10.63), should be compared with uniform conver- gence (Section 5.5, Eq. (5.67)). Clearly, uniform convergence implies convergence in the mean, but the converse does not hold; convergence in the mean is less restrictive. Specifi- cally,Eq.(10.63)isnotupsetbypiecewisecontinuousfunctionswithonlyafinitenumber of finite discontinuities. A relevant example is the Gibbs phenomenon of discontinuous Fourier series discussed in Section 14.5, which occurs for other eigenfunction series as well. Equation (10.63) is perfectly adequate for our purposes and is far more convenient than Eq.(5.67).Indeed,sincewefrequentlyuseeigenfunctionstodescribediscontinuousfunc- tions,convergenceinthemeanisallwecanexpect. In Eq.(10.62) theexpansioncoefficients ammaybedeterminedby am=integraldisplayb aF(x)ϕ∗ m(x)w(x)dx. (10.64) 10If wehave afinite set,as with vectors, the summation is over the number of linearly independent members of the set. 11Many authors use the term closedhere. 650 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions This follows from multiplying Eq. (10.62) by ϕ∗ m(x)w(x)and integrating. From the or- thogonalityoftheeigenfunctions ϕn(x),onlythe mthtermsurvives.Hereweseethevalue of orthogonality. Equation (10.64) may be compared with the dot or inner product of vec- tors, Section 1.3, and aminterpreted as the mth projection of the function F(x).O f t e nt h e coefficient amiscalleda generalizedFouriercoefficient . For a known function F(x), Eq. (10.64) gives amas adefiniteintegral that can always beevaluated,bycomputerif notanalytically. In the language of linear algebra, we have a linear space, a function vector space. The linearly independent, orthonormal functions ϕn(x)form the basis for this (infinite- dimensional) space. Equation (10.62) is a statement that the functions ϕn(x)span this linear space. With an inner product defined by Eq. (10.64), our linear space is a Hilbert space. Setting the weight function w(x)=1 for simplicity, completeness in operator form for adiscreteset ofeigenfunctions |ϕi/angbracketrightbecomes summationdisplay i|ϕi/angbracketright/angbracketleftϕi|=1. Multiplyingthecompletenessrelationby |F/angbracketrightweobtaintheeigenfunctionexpansion |F/angbracketright=summationdisplay i|ϕi/angbracketright/angbracketleftϕi|F/angbracketright with the generalized Fourier coefficient ai=/angbracketleftϕi|F/angbracketright.Equivalently in coordinate represen- tation, summationdisplay iϕ∗ i(y)ϕi(x)=δ(x−y) implies F(x)=integraldisplay F(y)δ(x−y)dy=summationdisplay iϕi(x)integraldisplay ϕ∗ i(y)F(y)dy. Without proof, we state that the spectrum of a linear operator Athat maps a Hilbert spaceHintoitselfmaybedividedintoadiscrete(orpoint)spectrumwitheigenvectorsof finite length, a continuous spectrum so that the eigenvalue equation Av=λvwithvinH does not have a unique bounded inverse (A−λ)−1in a dense domain of Hand a residual spectrumwhere (A−λ)−1isunboundedina domainnotdensein H. The question of completeness of a set of functions is often determined by comparison with a Laurent series, Section 6.5. In Section 14.1 this is done for Fourier series, thus establishingthecompletenessofFourierseries.Forallorthogonalpolynomialsmentioned inSection10.3it ispossibletofindapolynomialexpansionofeachpowerof z, zn=nsummationdisplay i=0aiPi(z), (10.65) 10.4 Completeness of Eigenfunctions 651 wherePi(z)istheithpolynomial.Exercises12.4.6,13.1.6,13.2.5,and13.3.22arespecific examples of Eq. (10.65). Using Eq. (10.65), we may reexpress the Laurent expansion of f(z)in terms of the polynomials, showing that the polynomial expansion exists (when it exists, it is unique, Exercise 10.4.1). The limitation of this Laurent series development is thatitrequiresthefunctiontobeanalytic.Equations(10.62)and(10.63)aremoregeneral. F(x)maybeonlypiecewisecontinuous.Numerousexamplesoftherepresentationofsuch piecewise continuous functions appear in Chapter 14 (Fourier series). A proof that our Sturm–LiouvilleeigenfunctionsformcompletesetsappearsinCourantandHilbert.12 For examples of particular eigenfunction expansions, see the following: Fourier series, Section10.2 andChapter14; Bessel andFourier–Besselexpansions,Section11.2; Legen- dre series, Section 12.3; Laplace series, Section 12.6; Hermite series, Section 13.1; La- guerreseries, Section13.2;andChebyshevseries, Section13.3. Itmayalsohappenthattheeigenfunctionexpansion,Eq.(10.62),istheexpansionofan unknown F(x)in a series of known eigenfunctions ϕn(x)with unknown coefficients an. An example would be the quantum chemist’s attempt to describe an (unknown) mole- cular wave function as a linear combination of known atomic wave functions. The un- known coefficients anwould be determined by a variational technique—Rayleigh–Ritz, Section17.8. Bessel’s inequality If the set of functions ϕn(x)does not form a complete set, possibly because we simply have not included the required infinite number of members of an infinite set, we are led to Bessel’s inequality. First, consider the finite case from vector analysis. Let Abe ann componentvector, A=e1a1+e2a2+···+enan, (10.66) in which eiis a unit vector and aiis the corresponding component (projection) of A; that is, ai=A·ei. (10.67) Then parenleftbigg A−summationdisplay ieiaiparenrightbigg2 ≥0. (10.68) If we sum over all ncomponents, the summation clearly, equals Aby Eq. (10.66) and the equality holds. If, however, the summation does not include all ncomponents, the inequality results. By expanding Eq. (10.68) and choosing the unit vectors so as to satisfy anorthogonalityrelation, ei·ej=δij, (10.69) 12R. Courant and D. Hilbert, Methods of Mathematical Physics (English translation), Vol. 1, New York: Interscience (1953), reprinted, Wiley (1989), Chapter 6,Section 3. 652 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions wehave A2≥summationdisplay ia2 i. (10.70) Thisis Bessel’sinequality. Forrealfunctionsweconsidertheintegral integraldisplayb abracketleftbigg f(x)−summationdisplay iaiϕi(x)bracketrightbigg2 w(x)dx≥0. (10.71) This is the continuum analog of Eq. (10.68), letting n→∞and replacing the summation byanintegration.Again,withtheweightingfactor w(x)>0,theintegrandisnonnegative. The integral vanishes by Eq. (10.62) if we have a complete set. Otherwise it is positive. Expandingthesquaredterm,weobtain integraldisplayb abracketleftbig f(x)bracketrightbig2w(x)dx−2summationdisplay iaiintegraldisplayb af(x)ϕi(x)w(x)dx+summationdisplay ia2 i≥0.(10.72) ApplyingEq. (10.64),wehave integraldisplayb abracketleftbig f(x)bracketrightbig2w(x)dx≥summationdisplay ia2 i. (10.73) Hence the sum of the squares of the expansion coefficients aiis less than or equal to the weightedintegralof [f(x)]2,theequalityholdingifandonlyiftheexpansionisexact,that is, ifthesetof functions ϕn(x)is acompleteset. In later chapters, when we consider eigenfunctions that form complete sets (such as Legendre polynomials), Eq. (10.73) with the equal sign holding will be called a Parseval relation. Bessel’s inequality has a variety of uses, including proof of convergence of the Fourier series. Schwarz Inequality The frequently used Schwarz inequality is similar to the Bessel inequality. Consider the quadraticequationwithunknown x: nsummationdisplay i=1(aix+bi)2=nsummationdisplay i=1a2 iparenleftbigg x+bi aiparenrightbigg2 =0 (10.74) withreal ai,bi.Ifbi/ai=constant, c,thatis,independentoftheindex i,thenthesolution isx=−c.Ifbi/aiisnotaconstantin i,alltermscannotvanishsimultaneouslyforreal x. Sothesolutionmustbecomplex.Expanding,wefindthat x2nsummationdisplay ia2 i+2xnsummationdisplay iaibi+nsummationdisplay ib2 i=0, (10.75) 10.4 Completeness of Eigenfunctions 653 andsince xis complex(or =−bi/ai), thequadraticformula13forxleadsto parenleftbiggnsummationdisplay i=1aibiparenrightbigg2 ≤parenleftbiggnsummationdisplay i=1a2 iparenrightbiggparenleftbiggnsummationdisplay i=1b2 iparenrightbigg , (10.76) theequalityholdingwhen bi/aiequalsaconstant,independentof i. Oncemore, interms ofvectors,wehave (a·b)2=a2b2cos2θ≤a2b2, (10.77) whereθistheangleincludedbetween aandb. TheanalogousSchwarzinequalityforcomplexfunctionshastheform vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplayb af∗(x)g(x)dxvextendsinglevextendsinglevextendsinglevextendsingle2 ≤integraldisplayb af∗(x)f(x)dxintegraldisplayb ag∗(x)g(x)dx, (10.78) theequalityholdingifandonlyif g(x)=αf (x),αbeingaconstant.Toprovethisfunction formoftheSchwarzinequality,14consideracomplexfunction ψ(x)=f(x)+λg(x)with λa complex constant, where f(x)andg(x)are any two square integrable functions (for which the integrals on the right-hand side exist). Multiplying by the complex conjugate andintegrating,weobtain integraldisplayb aψ∗ψdx≡integraldisplayb af∗fdx+λintegraldisplayb af∗gdx+λ∗integraldisplayb ag∗fdx +λλ∗integraldisplayb ag∗gdx≥0. (10.79) The≥0appearssince ψ∗ψisnonnegative,theequal (=)signholdingonlyif ψ(x)isiden- ticallyzero.Notingthat λandλ∗arelinearlyindependent,wedifferentiatewithrespectto oneof themandsetthederivativeequaltozerotominimizeintegraltextb aψ∗ψdx: ∂ ∂λ∗integraldisplayb aψ∗ψdx=integraldisplayb ag∗fdx+λintegraldisplayb ag∗gdx=0. Thisyields λ=−integraltextb ag∗fdx integraltextb ag∗gdx. (10.80a) Takingthecomplexconjugate,weobtain λ∗=−integraltextb af∗gdx integraltextb ag∗gdx. (10.80b) Substituting these values of λandλ∗back into Eq. (10.79), we obtain Eq. (10.78), the Schwarzinequality. 13With negative (or zero) discriminant. 14An alternatederivation is provided by theinequalityintegraltextintegraltext [f(x)g(y)−f(y)g(x)]∗[f(x)g(y)−f(y)g(x)]dxdy≥0. 654 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions In quantum mechanics f(x)andg(x)might each represent a state or configuration of a physical system, that is, a linear combination of wave functions. Then the Schwarz in- equality gives an upper limit for the absolute value of the inner productintegraltextb af∗(x)g(x)dx . In some texts the Schwarz inequality is a key step in the derivation of the Heisenberg uncertaintyprinciple. ThefunctionnotationofEqs.(10.78)and(10.79)isrelativelycumbersome.Inadvanced mathematical physics and especially in quantum mechanics it is common to use the Dirac bra-ketnotation.Usingthisnotation,wesimplyunderstandtherangeofintegration, (a,b), andthepresenceoftheweightingfunction w(x)≥0.InthisnotationtheSchwarzinequal- itytakestheelegantform vextendsinglevextendsingle/angbracketleftf|g/angbracketrightvextendsinglevextendsingle2≤/angbracketleftf|f/angbracketright/angbracketleftg|g/angbracketright. (78a) Ifg(x)is anormalizedeigenfunction, ϕi(x), Eq. (10.78)yields(here w(x)=1) a∗ iai≤integraldisplayb af∗(x)f(x)dx, (10.81) aresultthatalsofollowsfrom Eq. (10.73). For useful representations of Dirac’s delta function in terms of orthogonal sets of func- tionsandtherelationbetweenclosureandcompletenesswerefertotherelevantsubsection of Section 1.15, including Exercise 1.15.16, and for coordinate versus momentum repre- sentationsinquantummechanicstoSection15.6. Summary — Vector Spaces, Completeness Here we summarize some properties of vector spaces, first with the vectors taken to be the familiar real vectors of Chapter 1 and then with the vectors taken to be ordinary func- tions.Theconceptof completeness hasbeendevelopedforfinitevectorspaces(Chapter1, Eq. (1.5)) and carries over into infinite vector spaces. For example, in three-dimensional Euclidean space every vector can be written in terms of a linear combination of the three coordinate unit vectors (representing a basis) involvingthe vector’s Cartesian components as the expansion coefficients. Or a periodic function of an infinite vector space can be ex- panded in terms of the set of periodic functions sin nx,cosnx,n=0,1,2,...,that form a basis of this space. Since any periodic function with reasonable properties (spelled out in Chapter14)canbeexpandedintermsofthesesineandcosinefunctions,theyarecomplete andform abasisofsuchalinearfunctionspace. 1v.We shall describe our vector space with a set of nlinearly independent vectors ei, i=1,2,...,n.I fn=3, thene1=ˆx,e2=ˆy, ande3=ˆz.T h eneispanthe linear vector space. 1f.We shall describe our vector (function) space with a set of nlinearly independent functions, ϕi(x),i=0,1,...,n−1. The index istarts with 0 to agree with the labeling of the classical polynomials. Here ϕi(x)is assumed to be a polynomial of degree i.T h e nϕi(x)spanthelinearvector(function)space. 2v.Thevectorsinourvectorspacesatisfythefollowingrelations(Section1.2;thevector componentsarenumbers): 10.4 Completeness of Eigenfunctions 655 a. Vectoradditionis commutative u+v=v+u b. Vectoradditionis associative [u+v]+w=u+[v+w] c. Thereisa nullvector 0+v=v d. Multiplicationbyascalar Distributive a[u+v]=au+av Distributive (a+b)u=au+bu Associative a[bu]=(ab)u e. Multiplication Byunitscalar 1 u=u Byzero 0 u=0 f. Negativevector (−1)u=−u. 2f.The functions in our linear function space satisfy the properties listed for vectors (substitute“function”for “vector”): f(x)+g(x)=g(x)+f(x) bracketleftbig f(x)+g(x)bracketrightbig +h(x)=f(x)+bracketleftbig g(x)+h(x)bracketrightbig 0+f(x)=f(x) abracketleftbig f(x)+g(x)bracketrightbig =af(x)+ag(x) (a+b)f(x)=af(x)+bf (x) abracketleftbig bf (x)bracketrightbig =(ab)f(x) 1·f(x)=f(x) 0·f(x)=0 (−1)·f(x)=−f(x). 3v.Inn-dimensionalvectorspaceanarbitraryvector cisdescribedbyits ncomponents (c1,c2,...,cn),or c=nsummationdisplay i=1ciei. Whennei(1) are linearly independent and (2) span the n-dimensional vector space, then theeiforma basisandconstitutea complete set. 3f.Inn-dimensionalfunctionspaceapolynomialof degree m≤n−1 isdescribedby f(x)=n−1summationdisplay i=0ciϕi(x). When the nϕi(x)(1) are linearly independent and (2) span the n-dimensional function space,thenthe ϕi(x)formabasisandconstitutea complete set(fordescribingpolynomi- alsofde gree m≤n−1). 4v.Aninnerproduct(scalar, dotproduct)of avectorspaceisdefinedby c·d=nsummationdisplay i=1cidi. 656 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Ifcanddhavecomplexcomponentsinanorthogonalcoordinatesystem,theinnerproduct isdefinedassummationtextn i=1c∗ idi.Theinnerproducthasthepropertiesof a. Distributivelawofaddition c·(d+e)=c·d+c·e b. Scalarmultiplication c·ad=ac·d c. Complexconjugation c·d=(d·c)∗. 4f.Aninnerproductof alinearspaceof functionsis definedby /angbracketleftf|g/angbracketright=integraldisplayb af∗(x)g(x)w(x)dx. The choice of the weighting function w(x)and the interval (a,b)follows from the dif- ferential equation satisfied by ϕi(x)and the boundary conditions—Section 10.1. In ma- trix terminology, Section 3.2, |g/angbracketrightis a column vector and /angbracketleftf|is a row vector, the adjoint of|f/angbracketright,where both may have infinitely many components. For example, if we expand g(x)=summationtext igiϕi(x),then|g/angbracketrighthas theith component giin a column vector and |f/angbracketrighthasf∗ i asitsithcomponentinarow vector. Theinnerproducthasthepropertieslistedforvectors: a./angbracketleftf|g+h/angbracketright=/angbracketleftf|g/angbracketright+/angbracketleftf|h/angbracketright b./angbracketleftf|ag/angbracketright=a/angbracketleftf|g/angbracketright c./angbracketleftf|g/angbracketright=/angbracketleftg|f/angbracketright∗. 5v. Orthogonality: ej·ej=0,i/negationslash=j. If theneiare not already orthogonal, the Gram–Schmidt process may be used to create anorthogonalset. 5f. Orthogonality: /angbracketleftϕi|ϕj/angbracketright=integraldisplayb aϕ∗ i(x)ϕj(x)w(x)dx=0,i/negationslash=j. Ifthenϕi(x)arenotalreadyorthogonal,theGram–Schmidtprocess(Section10.3)maybe usedtocreateanorthogonalset. 6v.Definitionof norm: |c|=(c·c)1/2=parenleftbiggnsummationdisplay i=1c2 iparenrightbigg1/2 . The basis vectors eiare taken to have unit norm (length) ei·ei=1. The components of c aregivenby ci=ei·c,i=1,2,...,n. 6f.Definitionof norm: /bardblf/bardbl=/angbracketleftf|f/angbracketright1/2=bracketleftbiggintegraldisplayb avextendsinglevextendsinglef(x)vextendsinglevextendsingle2w(x)dxbracketrightbigg1/2 =bracketleftbiggn−1summationdisplay i=0|ci|2bracketrightbigg1/2 , 10.4 Completeness of Eigenfunctions 657 Parseval’s identity. /bardblf/bardbl>0 unless f(x)is identically zero. The basis functions ϕi(x) maybetakentohaveunitnorm(unitnormalization), /bardblϕi/bardbl=1. Theexpansioncoefficientsof ourpolynomial f(x)aregivenby ci=/angbracketleftϕi|f/angbracketright,i=0,1,...,n−1. 7v. Bessel’sinequality: c·c≥summationdisplay ic2 i. If the equals sign holds for all c, it indicates that the eispan the vector space; that is, they arecomplete. 7f. Bessel’sinequality: /angbracketleftf|f/angbracketright=integraldisplayb avextendsinglevextendsinglef(x)vextendsinglevextendsingle2w(x)dx≥summationdisplay i|ci|2. If the equals sign holds for all allowable f, it indicates that the ϕi(x)span the function space;thatis, theyarecomplete. 8v. Schwarz’inequality: |c·d|≤|c|·|d|. The equals sign holds when cis a multiple of d. If the angle included between canddis θ,then|cosθ|≤1. 8f. Schwarz’inequality: vextendsingle vextendsingle/angbracketleftf|g/angbracketrightvextendsinglevextendsingle≤/angbracketleftf|f/angbracketright1/2/angbracketleftg|g/angbracketright1/2=/bardblf/bardbl·/bardblg/bardbl. The equals sign holds when f(x)andg(x)are linearly dependent, that is, when f(x)is a multipleof g(x). Now,letn→∞,forminganinfinite-dimensionallinearvectorspace, l2. 9v.Inaninfinite-dimensionalspaceourvector cis c=∞summationdisplay i=1ciei. Werequirethat ∞summationdisplay i=1c2 i<∞. Thecomponentsof caregivenby ci=ei·c,i=1,2,...,∞, exactlyas inafinite-dimensionalvectorspace. Then let n→∞, forming an infinite-dimensional vector (function) space L2. ThenL standsforLebesgue,thesuperscript2forthequadraticnorm,thatis,the2in |f(x)|2.Our functionsneednolongerbepolynomials,butwedorequirethat f(x)beatleastpiecewise 658 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions continuous (Dirichlet conditions for Fourier series) and that /angbracketleftf|f/angbracketright=integraltextb a|f(x)|2w(x)dx exist.Thislatterconditionis oftenstatedas arequirementthat f(x)besquareintegrable. 9f.Cauchy sequence (generalized Fourier expansion): Expand f(x)=summationtext∞ i=0fiϕi(x) andlet fn(x)=nsummationdisplay i=0fiϕi(x). If vextenddoublevextenddoublef(x)−fn(x)vextenddoublevextenddouble→0asn→∞ or limn→∞integraldisplayvextendsinglevextendsinglevextendsinglevextendsinglef(x)−nsummationdisplay i=0fiϕi(x)vextendsinglevextendsinglevextendsinglevextendsingle2 w(x)dx=0, then we have convergence in the mean. This is analogous to the partial sum–Cauchy se- quencecriterionfor theconvergenceof aninfiniteseries, Section5.1. IfeveryCauchysequenceofallowablevectors(squareintegrable,piecewisecontinuous functions) converges to a limit vector in our linear space, the space is said to be complete. Then f(x)=∞summationdisplay i=0ciϕi(x) (almosteverywhere ) inthesenseofconvergenceinthemean.Asnotedbefore,thisisaweakerrequirementthan pointwiseconvergence(fixedvalueof x)oruniformconvergence. Expansion Coefficients Forafunction fitsexpansioncoefficientsaredefinedas ci=/angbracketleftϕi|f/angbracketright,i=0,1,...,∞, exactlyas inafinite-dimensionalvectorspace.Hence f(x)=summationdisplay i/angbracketleftϕi|f/angbracketrightϕi(x). A linear space (finite- or infinite-dimensional) that (1) has an inner product defined (/angbracketleftf|g/angbracketright)and(2)is completeis a Hilbertspace . Infinite-dimensionalHilbertspaceprovidesanaturalmathematicalframe-workformod- ern quantummechanics.Away from quantummechanics,Hilbert space retains its abstract mathematicalpowerandbeautyandhasmanyuses. 10.4 Completeness of Eigenfunctions 659 Exercises 10.4.1 Afunction f(x)is expandedinaseriesof orthonormaleigenfunctions f(x)=∞summationdisplay n=0anϕn(x). Show that the series expansion is unique for a given set of ϕn(x). The functions ϕn(x) arebeingtakenhereasthe basisvectorsinaninfinite-dimensionalHilbertspace. 10.4.2 Afunction f(x)is representedbyafinitesetof basisfunctions ϕi(x), f(x)=Nsummationdisplay i=1ciϕi(x). Showthatthecomponents ciare unique,thatnodifferent set c′ iexists. Note. Your basis functions are automatically linearly independent. They are not neces- sarilyorthogonal. 10.4.3 A function f(x)is approximated by a power seriessummationtextn−1 i=0cixiover the interval [0,1]. Showthatminimizingthemeansquareerrorleadstoaset oflinearequations Ac=b, where Aij=integraldisplay1 0xi+jdx=1 i+j+1,i,j=0,1,2,...,n−1 and bi=integraldisplay1 0xif(x)dx, i =0,1,2,...,n−1. Note.TheAijaretheelementsoftheHilbertmatrixoforder n.Thedeterminantofthis Hilbertmatrixisarapidlydecreasingfunctionof n.Forn=5,detA=3.7×10−12and thesetofequations Ac=bisbecomingill-conditionedandunstable. 10.4.4 Inplaceof theexpansionofa function F(x)givenby F(x)=∞summationdisplay n=0anϕn(x), with an=integraldisplayb aF(x)ϕn(x)w(x)dx, takethefiniteseries approximation F(x)≈msummationdisplay n=0cnϕn(x). 660 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Showthatthemeansquareerror integraldisplayb abracketleftbigg F(x)−msummationdisplay n=0cnϕn(x)bracketrightbigg2 w(x)dx isminimizedbytaking cn=an. Note.Thevaluesofthecoefficientsareindependentofthenumberoftermsinthefinite series. This independence is a consequence of orthogonality and would not hold for a least-squaresfitusingpowersof x. 10.4.5 FromExample10.2.2, f(x)=  h 2,0<x<π −h 2,−π<x<0  =2h π∞summationdisplay n=0sin(2n+1)x 2n+1. (a) Showthat integraldisplayπ −πbracketleftbig f(x)bracketrightbig2dx=π 2h2=4h2 π∞summationdisplay n=0(2n+1)−2. For a finite upper limit this would be Bessel’s inequality. For the upper limit ∞, thisisParseval’sidentity. (b) Verifythat π 2h2=4h2 π∞summationdisplay n=0(2n+1)−2 byevaluatingtheseries. Hint.Theseriescanbeexpressedas theRiemannzetafunction. 10.4.6 DifferentiateEq. (10.79), /angbracketleftψ|ψ/angbracketright=/angbracketleftf|f/angbracketright+λ/angbracketleftf|g/angbracketright+λ∗/angbracketleftg|f/angbracketright+λλ∗/angbracketleftg|g/angbracketright, withrespectto λ∗andshowthatyougettheSchwarzinequality,Eq. (10.78). 10.4.7 DerivetheSchwarzinequalityfromtheidentity bracketleftbiggintegraldisplayb af(x)g(x)dxbracketrightbigg2 =integraldisplayb abracketleftbig f(x)bracketrightbig2dxintegraldisplayb abracketleftbig g(x)bracketrightbig2dx −1 2integraldisplayb aintegraldisplayb abracketleftbig f(x)g(y)−f(y)g(x)bracketrightbig2dxdy. 10.4.8 Ifthefunctions f(x)andg(x)oftheSchwarzinequality,Eq.(10.78),maybeexpanded in a series of eigenfunctions ϕi(x), show that Eq. (10.78) reduces to Eq. (10.76) (with npossiblyinfinite). 10.4 Completeness of Eigenfunctions 661 Notethedescriptionof f(x)asavectorinafunctionspaceinwhich ϕi(x)corresponds totheunitvector e1. 10.4.9 Theoperator HisHermitianandpositivedefinite;thatis, for all f: integraldisplayb af∗Hf dx > 0. ProvethegeneralizedSchwarzinequality: vextendsinglevextendsinglevextendsinglevextendsingleintegraldisplayb af∗Hgdxvextendsinglevextendsinglevextendsinglevextendsingle2 ≤integraldisplayb af∗Hf dxintegraldisplayb ag∗Hgdx. 10.4.10 A normalized wave function ψ(x)=summationtext∞ n=0anϕn(x). The expansion coefficients anare knownasprobabilityamplitudes.Wemaydefineadensitymatrix ρwithelements ρij= aia∗ j. Showthat parenleftbig ρ2parenrightbig ij=ρij, or ρ2=ρ. Thisresult, bydefinition,makes ρaprojectionoperator. Hint:Use integraldisplay ψ∗ψdx=1. 10.4.11 Showthat (a) theoperator vextendsinglevextendsingleϕi(x)angbracketrightbigangbracketleftbig ϕi(t)vextendsinglevextendsingle operatingon f(t)=summationdisplay jcjvextendsinglevextendsingleϕj(t)angbracketrightbig yields civextendsinglevextendsingleϕi(x)angbracketrightbig . (b)summationdisplay ivextendsinglevextendsingleϕi(x)angbracketrightbigangbracketleftbig ϕi(x)vextendsinglevextendsingle=1. This operator is a projection operator projecting f(x)onto theith coordinate, selectivelypickingoutthe ithcomponent ci|ϕi(x)/angbracketrightoff(x). Hint.Theoperatoroperatesviathewell-definedinnerproduct. 662 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.5 G REEN ’SFUNCTION —E IGENFUNCTION EXPANSION Aseriessomewhatsimilartothatrepresenting δ(x−t)resultswhenweexpandtheGreen’s functionintheeigenfunctionsofthecorrespondinghomogeneousequation.Intheinhomo- geneousHelmholtzequationwehave ∇2ψ(r)+k2ψ(r)=−ρ(r). (10.82) ThehomogeneousHelmholtzequationis satisfiedbyits orthonormaleigenfunctions ϕn, ∇2ϕn(r)+k2 nϕn(r)=0. (10.83) As outlined in Section 9.7, the Green’s function G(r1,r2)satisfies the point source equa- tion ∇2G(r1,r2)+k2G(r1,r2)=−δ(r1−r2) (10.84) and the boundary conditions imposed on the solutions of the homogeneous equation. Be- causeGis real, we expand the Green’s function in a series of real eigenfunctions of the homogeneousequation(10.83);thatis, G(r1,r2)=∞summationdisplay n=0an(r2)ϕn(r1), (10.85) andbysubstitutingintoEq.(10.84) weobtain −∞summationdisplay n=0an(r2)k2 nϕn(r1)+k2∞summationdisplay n=0an(r2)ϕn(r1)=−∞summationdisplay n=0ϕn(r1)ϕn(r2).(10.86) Hereδ(r1−r2)has been replaced by its eigenfunction expansion, Eq. (1.190). When we employtheorthogonalityof ϕn(r1)toisolate an,thisyields ∞summationdisplay m=0am(r2)parenleftbig k2−k2 mparenrightbigintegraldisplay ϕn(r1)ϕm(r1)d3r1=−∞summationdisplay m=0ϕm(r2)integraldisplay ϕn(r1)ϕm(r1)d3r1, or an(r2)parenleftbig k2−k2 nparenrightbig =−ϕn(r2). Thensubstitutingthis intoEq. (10.85), theGreen’sfunctionbecomes G(r1,r2)=∞summationdisplay n=0ϕn(r1)ϕn(r2) k2n−k2, (10.87) a bilinear expansion, symmetric with respect to r1andr2, as expected. Finally, ψ(r1),t h e desiredsolutionoftheinhomogeneousequation,is givenby ψ(r1)=integraldisplay G(r1,r2)ρ(r2)dτ2. (10.88) 10.5 Green’s Function — Eigenfunction Expansion 663 If wegeneralizeourinhomogeneousdifferentialequationto Lψ+λψ=−ρ, (10.89) where Lis aHermitianoperator,wefindthat G(r1,r2)=∞summationdisplay n=0ϕn(r1)ϕn(r2) λn−λ, (10.90) whereλnis thenth eigenvalue and ϕnis the corresponding orthonormal eigenfunction of thehomogeneousdifferentialequation Lψ+λψ=0. (10.91) The eigenfunction expansion of the Green’s function in Eq. (10.90) makes the symmetry propertyG(r1,r2)=G(r2,r1)explicitandisoftenusefulwhencomparingwithsolutions obtainedbyothermeans. Green’s Functions — One-Dimensional The development of the Green’s function for two- and three-dimensional systems was the topicdiscussedintheprecedingmaterialandinSection9.7.Here,forsimplicity,werestrict ourselvestoone-dimensionalcasesandfollowa somewhatdifferent approach. DefiningProperties Inourone-dimensionalanalysisweconsiderfirst theinhomogeneousequation Ly(x)+f(x)=0, (10.92) inwhich Listheself-adjoint differentialoperator L=d dxparenleftbigg p(x)d dxparenrightbigg +q(x). (10.93) AsinSection10.1, y(x)isrequiredtosatisfycertainboundaryconditionsattheendpoints aandbof ourinterval [a,b]. We now proceed to define a rather strange and arbitrary function Gover the interval [a,b]. At this stage the most that can be said in defense of Gis that the defining prop- erties are legitimate, or mathematically acceptable. Later, Gwill appear as a reasonable tool for obtaining solutions of the inhomogeneous ODE, Eq. (10.92); this role dictates its properties. 1. The interval a≤x≤bis divided by a parameter t. We label G(x)=G1(x)fora≤ x<tandG(x)=G2(x)fort<x≤b. 2. Thefunctions G1(x)andG2(x)eachsatisfy thehomogeneous15equation;thatis, LG1(x)=0,a≤x<t, LG2(x)=0,t<x≤b.(10.94) 15Homogeneous with respectto the unknown function. Thefunction f(x)in Eq.(10.92) is setequal to zero. 664 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 3. Atx=a,G1(x)satisfies the boundary conditions we impose on y(x),as o l u t i o no f the inhomogeneous ODE, Eq. (10.92). At x=b,G2(x)satisfies the boundary condi- tions imposed on y(x)at this endpoint of the interval. For convenience, the boundary conditionsaretakentobehomogeneous;thatis, at x=a, y(a)=0,ory′(a)=0,orαy(a)+βy′(a)=0 andsimilarlyat x=b. 4. We demandthat G(x)becontinuous ,16 limx→t−G1(x)=limx→t+G2(x). (10.95) 5. We requirethat G′(x)bediscontinuous ,specificallythat15 d dxG2(x)vextendsinglevextendsingle t−d dxG1(x)vextendsinglevextendsingle t=−1 p(t), (10.96) wherep(t)comes from the self-adjoint operator, Eq. (10.93). Note that with the first derivativediscontinuous,thesecondderivativedoesnotexist. These requirements, in effect, make G a function of two variables, G(x,t).A l s o ,w e notethat G(x,t)dependsonboththeformofthedifferentialoperator Landtheboundary conditions that y(x)must satisfy. Note that we have described the properties of Green’s functions for second-order differential equations. Note that for Green’s functions for first- orderdifferentialequations,thediscontinuitiesarisein Gitself. Now, assuming that we can find a function G(x,t)that has these properties, we label it aGreen’sfunctionandproceedtoshowthatasolutionof Eq.(10.92) is y(x)=integraldisplayb aG(x,t)f(t)dt. (10.97) To do this we first construct the Green’s function G(x,t).L e tu(x)be a solution of the homogeneous equation that satisfies the boundary conditions at x=a, and letv(x)be a solutionthatsatisfiestheboundaryconditionsat x=b.Thenwemaytake17 G(x,t)=braceleftBiggc1u(x), a≤x<t, c2v(x), t <x ≤b.(10.98) Continuityat x=t(Eq. (10.95))requires c2v(t)−c1u(t)=0. (10.99) Finally,thediscontinuityinthefirstderivative(Eq. (10.96)) becomes c2v′(t)−c1u′(t)=−1 p(t). (10.100) 16Strictly speaking, this is the limit as x→t. 17The“constants” c1andc2areindependent of x, but they may (and do) depend on the other variable, t. 10.5 Green’s Function — Eigenfunction Expansion 665 There will be a unique solution for our unknown coefficients c1andc2if the Wronskian determinantvextendsinglevextendsinglevextendsinglevextendsinglevextendsingleu(t) v(t) u′(t) v′(t)vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=u(t)v′(t)−v(t)u′(t) does not vanish. We have seen in Section 9.6 that the nonvanishing of this determinant is a necessary condition for linear independence. Let us assume u(x)andv(x)to be inde- pendent.(If u(x)andv(x)arelinearlydependent,thesituationbecomesmorecomplicated andisnotconsideredhere. SeeCourantandHilbertinAdditionalReadingsof Chapter9.) For independent u(x)andv(x)we have the Wronskian (again from Section 9.6 or Exer- cise10.1.4) u(t)v′(t)−v(t)u′(t)=A p(t), (10.101) inwhichAisaconstant.Equation(10.101)issometimescalled Abel’sformula .Numerous examples have appeared in connection with Bessel and Legendre functions. Now, from Eq.(10.100), weidentify c1=−v(t) A,c 2=−u(t) A. (10.102) Equation(10.99)isclearlysatisfied.SubstitutionintoEq.(10.98)yieldsourGreen’sfunc- tion G(x,t)=  −1 Au(x)v(t), a ≤x<t, −1 Au(t)v(x), t <x ≤b.(10.103) Notethat G(x,t)=G(t,x). Thisisthesymmetrypropertythatwas provedearlierinSec- tion9.7.Itsphysicalinterpretationisgivenbythereciprocityprinciple(viaourpropagator function)—a cause at tyields the same effect at xas a cause at xproduces at t.I nt e r m s of our electrostaticanalogythis is obvious,the propagatorfunctiondependingonly on the magnitudeofthedistancebetweenthetwopoints: |r1−r2|=|r2−r1|. Green’s Function Integral — Differential Equation We have constructed G(x,t), but there still remains the task of showing that the integral (Eq.(10.97))withournewGreen’sfunctionisindeedasolutionoftheoriginaldifferential equation(10.92). This we do by direct substitution. With G(x,t)given by Eq. (10.103),18 Eq.(10.97) becomes y(x)=−1 Aintegraldisplayx av(x)u(t)f(t)dt −1 Aintegraldisplayb xu(x)v(t)f(t)dt. (10.104) 18In thefirst integral, a≤t≤x. HenceG(x,t)=G2(x,t)=−(1/A)u(t)v(x) . Similarly, the second integral requires G=G1. 666 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Differentiating,weobtain y′(x)=−1 Aintegraldisplayx av′(x)u(t)f(t)dt −1 Aintegraldisplayb xu′(x)v(t)f(t)dt, (10.105) thederivativesofthelimitscanceling.Aseconddifferentiationyields y′′(x)=−1 Aintegraldisplayx av′′(x)u(t)f(t)dt −1 Aintegraldisplayb xu′′(x)v(t)f(t)dt −1 Abracketleftbig u(x)v′(x)−v(x)u′(x)bracketrightbig f(x). (10.106) ByEqs. (10.100)and(10.102)thismayberewrittenas y′′(x)=−v′′(x) Aintegraldisplayx au(t)f(t)dt−u′′(x) Aintegraldisplayb xv(t)f(t)dt−f(x) p(x).(10.107) Now,bysubstitutingintoEq. (10.93),wehave Ly(x)=−Lv(x) Aintegraldisplayx au(t)f(t)dt−Lu(x) Aintegraldisplayb xv(t)f(t)dt−f(x). (10.108) Sinceu(x)andv(x)were chosen to satisfy the homogeneous equation, the L-factors are zeroandtheintegraltermsvanish,andwesee thatEq. (10.92)is satisfied. Wemustalsocheckthat y(x)satisfiestherequiredboundaryconditions.Atpoint x=a, y(a)=−u(a) Aintegraldisplayb av(t)f(t)dt=cu(a), (10.109) y′(a)=−u′(a) Aintegraldisplayb av(t)f(t)dt=cu′(a), (10.110) sincethedefiniteintegralisa constant.Wechose u(x)tosatisfy αu(a)+βu′(a)=0. (10.111) Multiplying by the constant c, we verify that y(x)also satisfies Eq. (10.111). This illus- trates the utility of the homogeneous boundary conditions: The normalization does not matter. In quantum mechanical problems the boundary condition on the wave function is oftenexpressedintermsoftheratio ψ′(x) ψ(x)=d dxlnψ(x), comparedtod dxlnu(x)vextendsinglevextendsingle x=a=−α β, Eq.(10.111). Theadvantageis thatthewavefunctionneednotbenormalizedyet. Summarizing,wehaveEq. (10.97), y(x)=integraldisplayb aG(x,t)f(t)dt, whichsatisfiesthedifferentialequation(Eq. (10.92)), Ly(x)+f(x)=0, 10.5 Green’s Function — Eigenfunction Expansion 667 andtheboundaryconditions,theseboundaryconditionshavingbeenbuiltintotheGreen’s function, G(x,t). Basically, what we have done is to use the solutions of the homogeneous equation Eq.(10.94)toconstructasolutionoftheinhomogeneousequation.Again,Poisson’sequa- tionisanillustration.Thesolution(Eq.(9.148))representsaweighted [ρ(r2)]combination of solutions of the corresponding homogeneous Laplace’s equation. (We followed these samesteps earlyinthis section.) It should be noted that our y(x), Eq. (10.97), is actually the particular solution of the differential equation, Eq. (10.92). Our boundary conditions exclude the addition of so- lutions of the homogeneous equation. In an actual physical problem we may well have both types of solutions. In electrostatics, for instance (compare Section 9.7), the Green’s functionsolutionofPoisson’sequationgivesthepotentialcreatedbythegivenchargedis- tribution.Inaddition,theremaybeexternalfieldssuperimposed.Thesewouldbedescribed bysolutionsofthehomogeneousequation,Laplace’sequation. Eigenfunction, Eigenvalue Equation Theprecedinganalysisplacednospecialrestrictionsonour f(x).Letusnowassumethat f(x)=λρ(x)y(x) .19Thenwehave y(x)=λintegraldisplayb aG(x,t)ρ(t)y(t)dt (10.112) asasolutionof Ly(x)+λρ(x)y(x)=0 (10.113) anditsboundaryconditions.Equation(10.112)isahomogeneousFredholmintegralequa- tion of the second kind, and Eq. (10.113) is the homogeneous eigenvalue equation (with theweightingfunction w(x)replacedby ρ(x)). ThereisachangeintheinterpretationofourGreen’sfunction.Itstartedasapropagator function, a weighting function giving the importance of the charge ρ(r2)in producing the potential ϕ(r1). The charge ρwas the inhomogeneous term in the inhomogeneous differential equation (10.92). Now the differential equation and the integral equation are bothhomogeneous .G(x,t)has become a link relating the two equations, differential and integral. To complete the discussion of this differential equation–integral equation equivalence, let us now show that Eq. (10.113) implies Eq. (10.112), that is, that a solution of our differential equation (10.113) with its boundary conditions satisfies the integral equa- tion (10.112). We multiply Eq. (10.113) by G(x,t), the appropriate Green’s function, and integratefrom x=atox=btoobtain integraldisplayb aG(x,t) Ly(x)dx+λintegraldisplayb aG(x,t)ρ(x)y(x)dx =0. (10.114) 19Thefunction ρ(x)is some weighting function, not acharge density. 668 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions Thefirstintegralissplitintwo (x<t,x>t) ,accordingtotheconstructionofourGreen’s function,giving −integraldisplayt aG1(x,t)Ly(x)dx−integraldisplayb tG2(x,t)Ly(x)dx=λintegraldisplayb aG(x,t)ρ(x)y(x)dx. (10.115) Note that tis the upper limit for the G1integrals and the lower limit for the G2integrals. We are going to reduce the left-hand side of Eq. (10.115) to y(t). Then, with G(x,t)= G(t,x), wehaveEq. (10.112)(with xandtinterchanged). ApplyingGreen’stheoremtotheleft-handsideor,equivalently,integratingbyparts,we obtain −integraldisplayt aG1(t,x)bracketleftbiggd dxparenleftbigg p(x)d dxy(x)parenrightbigg +q(x)y(x)bracketrightbigg dx =−bracketleftbig G1(x,t)p(x)y′(x)bracketrightbigvextendsinglevextendsinglex=t x=a+integraldisplayt aparenleftbigg∂ ∂xG1(x,t)parenrightbigg p(x)y′(x)dx −integraldisplayt aG1(x,t)q(x)y(x)dx, (10.116) withanequivalentexpressionforthesecondintegral.Asecondintegrationbyparts yields −integraldisplayt aG1(x,t)Ly(x)dx=−integraldisplayt ay(x)LG1(x,t)dx −bracketleftbig G1(x,t)p(x)y′(x)bracketrightbigvextendsinglevextendsinglex=t x=a +bracketleftbig G′ 1(x,t)p(x)y(x)bracketrightbigvextendsinglevextendsinglex=t x=a. (10.117) The integral on the right vanishes because LG1=0. By combining the integrated terms withthosefromintegrating G2,weha v e −p(t)bracketleftbigg G1(t,t)y′(t)−y(t)∂ ∂xG1(x,t)vextendsinglevextendsingle x=t−G2(t,t)y′(t)+y(t)∂ ∂xG2(x,t)vextendsinglevextendsingle x=tbracketrightbigg +p(a)bracketleftbigg y′(a)G1(a,t)−y(a)∂ ∂xG1(x,t)vextendsinglevextendsingle x=abracketrightbigg −p(b)bracketleftbigg G2(b,t)y′(b)−y(b)∂ ∂xG2(x,t)vextendsinglevextendsingle x=bbracketrightbigg . (10.118) Each of the last two expressions vanishes, for G(x,t)andy(x)satisfy the same boundary conditions.Thefirstexpression,withthehelpofEqs.(10.95)and(10.96),reducesto y(t). Substituting into Eq. (10.115), we have Eq. (10.112), thus completing the demonstration of the equivalence of the integral equation and the differential equation plus boundary conditions. Example 10.5.1 LINEAR OSCILLATOR Asasimpleexample,considerthelinearoscillatorequation(for avibratingstring): y′′(x)+λy(x)=0. (10.119) 10.5 Green’s Function — Eigenfunction Expansion 669 We impose the conditions y(0)=y(1)=0, which correspond to a string clamped at both ends.Now,toconstructourGreen’sfunction,weneedsolutionsofthehomogeneousequa- tionLy(x)=0,whichis y′′(x)=0.Tosatisfytheboundaryconditions,wemusthaveone solutionvanishat x=0,theotherat x=1.Suchsolutions(unnormalized)are u(x)=x, v(x)=1−x. (10.120) Wefindthat uv′−vu′=−1 (10.121) or,byEq. (10.101)with p(x)=1,A=−1.OurGreen’sfunctionbecomes G(x,t)=braceleftBiggx(1−t),0≤x<t, t(1−x), t<x≤1.(10.122) HencebyEq.(10.112)ourclampedvibratingstringsatisfies y(x)=λintegraldisplay1 0G(x,t)y(t)dt. (10.123) YoumayshowthattheknownsolutionsofEq. (10.119), y=sinnπx, λ=n2π2, doindeedsatisfy Eq. (10.123).Notethatoureigenvalue λisnotthewavelength. /squaresolid Green’s Function and the Dirac Delta Function One more approach to the Green’s function may shed additional light on our formulation and particularly on its relation to physical problems. Let us refer once more to Poisson’s equation,thistimefor apointcharge: ∇2ϕ(r)=−ρpoint ε0. (10.124) The Green’s function solution of this equation was developed in Section 9.7. This time let ustakeaone-dimensionalanalog Ly(x)+f(x)point=0. (10.125) Heref(x)pointrefers to a unit point “charge,” or a point force. We may represent it by a numberofforms, butperhapsthemostconvenientis f(x)point=  1 2ε,t−ε<x<t+ε, 0,elsewhere ,(10.126) whichisessentiallythesameasEq. (1.172). Then,integratingEq. (10.125),wehave integraldisplayt+ε t−εLy(x)dx=−integraldisplayt+ε t−εf(x)pointdx=−1 (10.127) 670 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions fromthedefinitionof f(x). Letus examine Ly(x)moreclosely.We have integraldisplayt+ε t−εd dxbracketleftbig p(x)y′(x)bracketrightbig dx+integraldisplayt+ε t−εq(x)y(x)dx =vextendsinglevextendsinglep(x)y′(x)vextendsinglevextendsinglet+ε t−ε+integraldisplayt+ε t−εq(x)y(x)dx=−1. (10.128) In the limit ε→0 we may satisfy this relation by permitting y′(x)to have a discontinu- ity of−1/p(x)atx=t,y(x)itself remaining continuous.20These, however, are just the properties used to define our Green’s function, G(x,t). In addition, we note that in the limitε→0, f(x)point=δ(x−t), (10.129) inwhichδ(x−t)isourDiracdeltafunction,definedinthismannerinSection1.15.Hence Eq.(10.125)has become LG(x,t)=−δ(x−t). (10.130) Thisisaone-dimensionalversionofEq.(9.159),whichweexploitforthedevelopmentof Green’s functions in two and three dimensions—Section 9.7. It will be recalled that we usedthisrelationinSection9.7todetermineourGreen’sfunctions. Equation (10.130) could have been expected since it is actually a consequence of our differential equation, Eq. (10.92), and Green’s function integral solution, Eq. (10.97). If we let Lx(subscript to emphasize that it operates on the x-dependence) operate on both sidesofEq. (10.97), then Lxy(x)=Lxintegraldisplayb aG(x,t)f(t)dt. By Eq. (10.92) the left-hand side is just −f(x). On the right Lx, is independent of the variableofintegration t, so wemaywrite −f(x)=integraldisplayb abraceleftbig LxG(x,t)bracerightbig f(t)dt. BydefinitionofDirac’sdeltafunction,Eqs.(1.171b)and(1.183),wehaveEq.(10.130). Exercises 10.5.1 Showthat G(x,t)=braceleftBiggx,0≤x<t, t, t<x≤1, istheGreen’sfunctionfor theoperator L=d2/dx2andtheboundaryconditions y(0)=0,y′(1)=0. 20The functions p(x)andq(x)appearing in the operator Lare continuous functions. With y(x)remaining continuous,integraltext q(x)y(x)dx is certainly continuous. Hencethis integral over an interval 2 ε(Eq. (10.128)) vanishes as εvanishes. 10.5 Green’s Function — Eigenfunction Expansion 671 10.5.2 FindtheGreen’sfunctionfor (a)Ly(x)=d2y(x) dx2+y(x),braceleftBiggy(0)=0, y′(1)=0. (b)Ly(x)=d2y(x) dx2−y(x), y(x) finitefor−∞<x<∞. 10.5.3 FindtheGreen’sfunctionfor theoperators (a)Ly(x)=d dxparenleftbigg xdy(x) dxparenrightbigg . ANS.G(x,t)=braceleftBigg−lnt,0≤x<t, −lnx, t<x≤1. (b)Ly(x)=d dxparenleftbigg xdy(x) dxparenrightbigg −n2 xy(x), withy(0)finiteand y(1)=0. ANS.G(x,t)=  1 2nbracketleftbiggparenleftbiggx tparenrightbiggn −(xt)nbracketrightbigg ,0≤x<t, 1 2nbracketleftbiggparenleftbiggt xparenrightbiggn −(xt)nbracketrightbigg ,t<x≤1. ThecombinationofoperatorandintervalspecifiedinExercise10.5.3(a)ispathological, in that one of the endpoints of the interval (zero) is a singular point of the operator. As a consequence, the integrated part (the surface integral of Green’s theorem) does not vanish.Thenextfour exercisesexplorethissituation. 10.5.4 (a) Showthattheparticularsolutionof d dxbracketleftbigg xd dxy(x)bracketrightbigg =−1 isyP(x)=−x. (b) Showthat yP(x)=−x/negationslash=integraldisplay1 0G(x,t)(−1)dt, whereG(x,t)is theGreen’sfunctionof Exercise10.5.3(a). 10.5.5 Show that Green’s theorem, Eq. (1.104) in one dimension with a Sturm–Liouville-type operator(d/dt)p(t)(d/dt) replacing ∇·∇, mayberewrittenas integraldisplayb abracketleftbigg u(t)d dtparenleftbigg p(t)dv(t) dtparenrightbigg −v(t)d dtparenleftbigg p(t)du(t) dtparenrightbiggbracketrightbigg dt =bracketleftbigg u(t)p(t)dv(t) dt−v(t)p(t)du(t) dtbracketrightbiggvextendsinglevextendsinglevextendsinglevextendsingleb a. 672 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.5.6 Usingtheone-dimensionalformof Green’stheoremofExercise10.5.5,let v(t)=y(t)andd dtparenleftbigg p(t)dy(t) dtparenrightbigg =−f(t), u(t)=G(x,t) andd dtparenleftbigg p(t)∂G(x,t) ∂tparenrightbigg =−δ(x−t). ShowthatGreen’stheoremyields y(x)=integraldisplayb aG(x,t)f(t)dt +bracketleftbigg G(x,t)p(t)dy(t) dt−y(t)p(t)∂ ∂tG(x,t)bracketrightbiggvextendsinglevextendsinglevextendsinglevextendsinglet=b t=a. 10.5.7 Forp(t)=t,y(t)=−t, G(x,t)=braceleftBigg−lnt,0≤x<t −lnx, t<x≤1, verifythattheintegratedpartdoesnotvanish. 10.5.8 ConstructtheGreen’sfunctionfor x2d2y dx2+xdy dx+parenleftbig k2x2−1parenrightbig y=0, subjecttotheboundaryconditions y(0)=0,y(1)=0. 10.5.9 Giventhat L=parenleftbig 1−x2parenrightbigd2 dx2−2xd dx and G(±1,t)remainsfinite , showthatnoGreen’sfunctioncanbeconstructedbythetechniquesofthissection.( u(x) andv(x)are linearlydependent.) 10.5.10 Constructtheone-dimensionalGreen’sfunctionfor theHelmholtzequation parenleftbiggd2 dx2+k2parenrightbigg ψ(x)=g(x). The boundary conditions are those for a wave advancing in the positive x-direction— assumingatimedependence e−iwt. ANS.G(x1,x2)=i 2kexpparenleftbig ik|x1−x2|parenrightbig . 10.5.11 Constructtheone-dimensionalGreen’sfunctionfor themodifiedHelmholtzequation parenleftbiggd2 dx2−k2parenrightbigg ψ(x)=f(x). 10.5 Green’s Function — Eigenfunction Expansion 673 The boundary conditions are that the Green’s function must vanish for x→∞and x→−∞. ANS.G(x1,x2)=1 2kexpparenleftbig −k|x1−x2|parenrightbig . 10.5.12 FromtheeigenfunctionexpansionoftheGreen’sfunctionshowthat (a)2 π2∞summationdisplay n=1sinnπxsinnπt n2=braceleftBiggx(1−t),0≤x<t, t(1−x), t<x≤1. (b)2 π2∞summationdisplay n=0sin(n+1 2)πxsin(n+1 2)πt (n+1 2)2=braceleftBiggx,0≤x<t, t, t<x≤1. Note.InSection10.4theGreen’sfunctionof L+λisexpandedineigenfunctions.The λthereis anadjustableparameter,notaneigenvalue. 10.5.13 IntheFredholmequation, f(x)=λ2integraldisplayb aG(x,t)ϕ(t)dt, G(x,t)is aGreen’sfunctiongivenby G(x,t)=∞summationdisplay n=1ϕn(x)ϕn(t) λ2n−λ2. Showthatthesolutionis ϕ(x)=∞summationdisplay n=1λ2 n−λ2 λ2ϕn(x)integraldisplayb af(t)ϕn(t)dt. 10.5.14 ShowthattheGreen’sfunctionintegraltransform operator integraldisplayb aG(x,t)[]dt isequalto−L−1, inthesensethat (a)Lxintegraldisplayb aG(x,t)y(t)dt =−y(x), (b)integraldisplayb aG(x,t) Lty(t)dt=−y(x). Note.T ak e Ly(x)+f(x)=0,Eq. (10.92). 10.5.15 SubstituteEq.(10.87),theeigenfunctionexpansionofGreen’sfunction,intoEq.(10.88) and then show that Eq. (10.88) is indeed a solution of the inhomogeneous Helmholtz equation(10.82). 674 Chapter 10 Sturm–Liouville Theory — Orthogonal Functions 10.5.16 (a) Startingwithaone-dimensionalinhomogeneousdifferentialequation(Eq.(10.89)), assume that ψ(x)andρ(x)may be represented by eigenfunction expansions. WithoutanyuseoftheDiracdeltafunctionorits representations,showthat ψ(x)=∞summationdisplay n=0integraltextb aρ(t)ϕn(t)dt λn−λϕn(x). Note that (1) if ρ=0, no solution exists unless λ=λnand (2) if λ=λn,n o solutionexistsunless ρisorthogonalto ϕn.Thissamebehaviorwillreappearwith integralequationsinSection16.4. (b) Interchanging summation and integration, show that you have constructed the Green’sfunctioncorrespondingtoEq. (10.90). 10.5.17 The eigenfunctions of the Schrödinger equation are often complex. In this case the orthogonalityintegral,Eq. (10.40), isreplacedby integraldisplayb aϕ∗ i(x)ϕj(x)w(x)dx=δij. InsteadofEq. (1.189),wehave δ(r1−r2)=∞summationdisplay n=0ϕn(r1)ϕ∗ n(r2). ShowthattheGreen’sfunction,Eq.(10.87), becomes G(r1,r2)=∞summationdisplay n=0ϕn(r1)ϕ∗ n(r2) k2n−k2=G∗(r2,r1). AdditionalReadings Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics . Reading, MA: Addison- Wesley (1969). Dennery, P.,andA.Krzywicki, Mathematics for Physicists .Reprinted. NewYork: Dover (1996). Hirsch,M., DifferentialEquations,DynamicalSystems,andLinearAlgebra .SanDiego:AcademicPress(1974). Miller,K.S., Linear Differential Equations in the RealDomain . New York: Norton (1963). Titchmarsh, E. C., Eigenfunction Expansions Associated with Second-Order Differential Equations , 2nd ed., Vol. 1.London: Oxford University Press (1962), Vol.II (1958). CHAPTER 11 BESSEL FUNCTIONS 11.1 B ESSEL FUNCTIONS OF THE FIRST KIND,Jν(x) Bessel functions appear in a wide variety of physical problems. In Section 9.3, separa- tion of the Helmholtz, or wave, equation in circular cylindrical coordinates led to Bessel’s equation. In Section 11.7 we will see that the Helmholtz equation in spherical polar co- ordinates also leads to a form of Bessel’s equation. Bessel functions may also appear in integral form—integral representations. This may result from integral transforms (Chap- ter 15) or from the mathematical elegance of starting the study of Bessel functions with Hankelfunctions,Section11.4. Besselfunctionsandcloselyrelatedfunctionsformarichareaofmathematicalanalysis with many representations, many interesting and useful properties, and many interrela- tions. Some of the major interrelations are developed in Section 11.1 and in succeeding sections.NotethatBesselfunctionsarenotrestrictedtoChapter11.Theasymptoticforms are developed in Section 7.3 as well as in Section 11.6. The confluent hypergeometric representationsappearinSection13.5. Generating Function for Integral Order AlthoughBesselfunctionsareofinterestprimarilyassolutionsofdifferentialequations,it isinstructiveandconvenienttodevelopthemfromacompletelydifferentapproach,thatof thegeneratingfunction.1Thisapproachalsohastheadvantageoffocusingonthefunctions themselvesratherthanonthedifferentialequationstheysatisfy.Letusintroduceafunction oftwovariables, g(x,t)=e(x/2)(t−1/t). (11.1) 1Generating functions have already been used in Chapter 5. In Section 5.6 the generating function (1+x)nwas used to derive the binomial coefficients. InSection5.9the generating function x(ex−1)−1wasused to derive the Bernoulli numbers. 675 676 Chapter 11 Bessel Functions ExpandingthisfunctioninaLaurentseries (Section6.5), weobtain e(x/2)(t−1/t)=∞summationdisplay n=−∞Jn(x)tn. (11.2) Itis instructivetocompareEq.(11.2) withtheequivalentEqs. (11.23)and(11.25). Thecoefficientof tn,Jn(x),isdefinedtobeaBesselfunctionofthefirstkind,ofintegral ordern. Expanding the exponentials, we have a product of Maclaurin series in xt/2 and −x/2t,respectively, ext/2·e−x/2t=∞summationdisplay r=0parenleftbiggx 2parenrightbiggrtr r!∞summationdisplay s=0(−1)sparenleftbiggx 2parenrightbiggst−s s!. (11.3) Here,thesummationindex rischangedto n,withn=r−sandsummationlimits n=−s to∞, and the order of the summations is interchanged, which is justified by absolute convergence.Therangeofthesummationover nbecomes−∞to∞,whilethesummation oversextendsfrom max (−n,0)to∞.For a given swegettn(n≥0)fromr=n+s: parenleftbiggx 2parenrightbiggn+stn+s (n+s)!(−1)sparenleftbiggx 2parenrightbiggst−s s!. (11.4) Thecoefficientof tnisthen2 Jn(x)=∞summationdisplay s=0(−1)s s!(n+s)!parenleftbiggx 2parenrightbiggn+2s =xn 2nn!−xn+2 2n+2(n+1)!+···. (11.5) ThisseriesformexhibitsisbehavioroftheBesselfunction Jn(x)forsmall xandpermits numericalevaluationof Jn(x). The results for J0,J1, andJ2are shown in Fig. 11.1. From Section 5.3 the error in using only a finite number of terms of this alternating series in numerical evaluation is less than the first term omitted. For instance, if we want Jn(x) FIGURE 11.1Besselfunctions, J0(x),J1(x), andJ2(x). 2Fromthestepsleadingtothisseriesandfromitsconvergencecharacteristicsitshouldbeclearthatthisseriesmaybeusedwith xreplacedby zandwith zany point in the finite complex plane. 11.1 Bessel Functions of the First Kind, Jν(x) 677 to±1% accuracy, the first term alone of Eq. (11.5) will suffice, provided the ratio of the second term to the first is less than 1% (in magnitude) or x<0.2(n+1)1/2. The Bessel functionsoscillatebutare notperiodic—exceptinthelimitas x→∞(Section11.6).The amplitudeof Jn(x)isnotconstantbutdecreasesasymptoticallyas x−1/2.(SeeEq.(11.137) forthis envelope.) Forn<0,Eq.(11.5) gives J−n(x)=∞summationdisplay s=0(−1)s s!(s−n)!parenleftbiggx 2parenrightbigg2s−n . (11.6) Sincenisaninteger(here), (s−n)!→∞fors=0,...,(n−1).Hencetheseriesmaybe consideredtostartwith s=n.Replacing sbys+n,weobtain J−n(x)=∞summationdisplay s=0(−1)s+n s!(s+n)!parenleftbiggx 2parenrightbiggn+2s , (11.7) showingimmediatelythat Jn(x)andJ−n(x)arenotindependentbutarerelatedby J−n(x)=(−1)nJn(x) (integraln). (11.8) These series expressions (Eqs. (11.5) and (11.6)) may be used with nreplaced by νto defineJν(x)andJ−ν(x)for nonintegral ν(compareExercise11.1.7). Recurrence Relations The recurrence relations for Jn(x)and its derivatives may all be obtained by operating on the series, Eq. (11.5), although this requires a bit of clairvoyance (or a lot of trial and error). Verification of the known recurrence relations is straightforward, Exercise 11.1.7. Here it is convenient to obtain them from the generating function, g(x,t). Differentiating bothsidesof Eq. (11.1)withrespectto t, wefindthat ∂ ∂tg(x,t)=1 2xparenleftbigg 1+1 t2parenrightbigg e(x/2)(t−1/t) =∞summationdisplay n=−∞nJn(x)tn−1, (11.9) andsubstitutingEq.(11.2)fortheexponentialandequatingthecoefficientsoflikepowers oft,3weobtain Jn−1(x)+Jn+1(x)=2n xJn(x). (11.10) This is a three-term recurrence relation. Given J0andJ1, for example, J2(and any other integralorder Jn)maybecomputed. 3This depends onthe factthat thepower-series representation is unique (Sections 5.7 and6.5). 678 Chapter 11 Bessel Functions DifferentiatingEq. (11.1)withrespectto x,weha v e ∂ ∂xg(x,t)=1 2parenleftbigg t−1 tparenrightbigg e(x/2)(t−1/t)=∞summationdisplay n=−∞J′ n(x)tn. (11.11) Again,substitutinginEq.(11.2)andequatingthecoefficientsoflikepowersof t,weobtain theresult Jn−1(x)−Jn+1(x)=2J′ n(x). (11.12) Asaspecialcaseof thisgeneralrecurrencerelation, J′ 0(x)=−J1(x). (11.13) AddingEqs.(11.10) and(11.12) anddividingby2,wehave Jn−1(x)=n xJn(x)+J′ n(x). (11.14) Multiplyingby xnandrearrangingtermsproduces d dxbracketleftbig xnJn(x)bracketrightbig =xnJn−1(x). (11.15) SubtractingEq. (11.12)fromEq. (11.10)anddividingby2yields Jn+1(x)=n xJn(x)−J′ n(x). (11.16) Multiplyingby x−nandrearrangingterms, weobtain d dxbracketleftbig x−nJn(x)bracketrightbig =−x−nJn+1(x). (11.17) Bessel’s Differential Equation Suppose we consider a set of functions Zν(x)that satisfies the basic recurrence relations (Eqs. (11.10) and (11.12)), but with νnot necessarily an integer and Zνnot necessarily givenbytheseries (Eq. (11.5)). Equation(11.14)mayberewritten (n→ν)as xZ′ ν(x)=xZν−1(x)−νZν(x). (11.18) Ondifferentiatingwithrespectto x,weha v e xZ′′ ν(x)+(ν+1)Z′ ν−xZ′ ν−1−Zν−1=0. (11.19) Multiplyingby xandthensubtractingEq. (11.18)multipliedby νgivesus x2Z′′ ν+xZ′ ν−ν2Zν+(ν−1)xZν−1−x2Z′ ν−1=0. (11.20) NowwerewriteEq. (11.16)andreplace nbyν−1: xZ′ ν−1=(ν−1)Zν−1−xZν. (11.21) UsingEq. (11.21)toeliminate Zν−1andZ′ ν−1fromEq. (11.20), wefinallyget x2Z′′ ν+xZ′ ν+parenleftbig x2−ν2parenrightbig Zν=0, (11.22) 11.1 Bessel Functions of the First Kind, Jν(x) 679 which is Bessel’s ODE. Hence any functions Zν(x)that satisfy the recurrence relations (Eqs. (11.10) and (11.12), (11.14) and (11.16), or (11.15) and (11.17)) satisfy Bessel’s equation; that is, the unknown Zνare Bessel functions. In particular, we have shown that thefunctions Jn(x), definedbyourgeneratingfunction,satisfyBessel’sODE.If theargu- mentiskρratherthan x,Eq.(11.22) becomes ρ2d2 dρ2Zν(kρ)+ρd dρZν(kρ)+parenleftbig k2ρ2−ν2parenrightbig Zν(kρ)=0. (11.22a) Integral Representation A particularly useful and powerful way of treating Bessel functions employs integral rep- resentations.Ifwereturntothegeneratingfunction(Eq.(11.2)),andsubstitute t=eiθ,we get eixsinθ=J0(x)+2bracketleftbig J2(x)cos2θ+J4(x)cos4θ+···bracketrightbig +2ibracketleftbig J1(x)sinθ+J3(x)sin3θ+···bracketrightbig , (11.23) inwhichwehaveusedtherelations J1(x)eiθ+J−1(x)e−iθ=J1(x)parenleftbig eiθ−e−iθparenrightbig =2iJ1(x)sinθ, (11.24) J2(x)e2iθ+J−2(x)e−2iθ=2J2(x)cos2θ, andso on. In summationnotation, cos(xsinθ)=J0(x)+2∞summationdisplay n=1J2n(x)cos(2nθ), (11.25) sin(xsinθ)=2∞summationdisplay n=1J2n−1(x)sinbracketleftbig (2n−1)θbracketrightbig , equatingrealandimaginarypartsofEq. (11.23). Byemployingtheorthogonalitypropertiesofcosineandsine,4 integraldisplayπ 0cosnθcosmθdθ=π 2δnm, (11.26a) integraldisplayπ 0sinnθsinmθdθ=π 2δnm, (11.26b) 4They are eigenfunctions of a self-adjoint equation (linear oscillator equation) and satisfy appropriate boundary conditions (compare Sections 10.2 and 14.1). 680 Chapter 11 Bessel Functions inwhich nandmarepositiveintegers(zerois excluded),5weobtain 1 πintegraldisplayπ 0cos(xsinθ)cosnθdθ=braceleftbiggJn(x), n even, 0,n odd,(11.27) 1 πintegraldisplayπ 0sin(xsinθ)sinnθdθ=braceleftbigg0,n even, Jn(x), n odd.(11.28) If thesetwo equationsareaddedtogether, Jn(x)=1 πintegraldisplayπ 0bracketleftbig cos(xsinθ)cosnθ+sin(xsinθ)sinnθbracketrightbig dθ =1 πintegraldisplayπ 0cos(nθ−xsinθ)dθ, n=0,1,2,3,.... (11.29) Asaspecialcase(integrateEq. (11.25)over (0,π)toget) J0(x)=1 πintegraldisplayπ 0cos(xsinθ)dθ. (11.30) Noting that cos (xsinθ)repeats itself in all four quadrants, we may write Eq. (11.30) as J0(x)=1 2πintegraldisplay2π 0cos(xsinθ)dθ. (11.30a) Ontheotherhand, sin (xsinθ)reversesits signinthethirdandfourthquadrants,so 1 2πintegraldisplay2π 0sin(xsinθ)dθ=0. (11.30b) Adding Eq. (11.30a) and itimes Eq. (11.30b), we obtain the complex exponential repre- sentation J0(x)=1 2πintegraldisplay2π 0eixsinθdθ=1 2πintegraldisplay2π 0eixcosθdθ. (11.30c) This integral representation (Eq. (11.29)) may be obtained somewhat more directly by employing contour integration (compare Exercise 11.1.16).6Many other integral repre- sentationsexist(compareExercise11.1.18). Example 11.1.1 FRAUNHOFER DIFFRACTION ,CIRCULAR APERTURE Inthetheoryof diffractionthroughacircularapertureweencountertheintegral /Phi1∼integraldisplaya 0rdrintegraldisplay2π 0eibrcosθdθ (11.31) 5Equations (11.26a) and (11.26b) hold for either morn=0. If both mandn=0, the constant in (11.26a) becomes π;t h e constant in Eq.(11.26b) becomes 0. 6Forn=0 a simple integration over θfrom 0 to 2 πwillconvert Eq.(11.23) into Eq. (11.30c). 11.1 Bessel Functions of the First Kind, Jν(x) 681 FIGURE 11.2Fraunhoferdiffraction,circularaperture. for/Phi1,theamplitudeofthediffractedwave.7Hereθisanazimuthangleintheplaneofthe circular aperture of radius a, andαis the angle defined by a point on a screen below the circular aperture relative to the normal through the center point. The parameter bis given by b=2π λsinα, (11.32) withλthe wavelength of the incident wave. The other symbols are defined by Fig. 11.2. FromEq. (11.30c)weget8 /Phi1∼2πintegraldisplaya 0J0(br)rdr. (11.33) Equation(11.15)enablesus tointegrateEq. (11.33)immediatelytoobtain /Phi1∼2πab b2J1(ab)∼λa sinαJ1parenleftbigg2πa λsinαparenrightbigg . (11.34) Noteherethat J1(0)=0.Theintensityofthelightinthediffractionpatternisproportional to/Phi12and /Phi12∼braceleftbiggJ1[(2πa/λ)sinα] sinαbracerightbigg2 . (11.35) 7Theexponent ibrcosθgivesthephaseofthewaveonthedistantscreenatangle αrelativetothephaseofthewaveincidenton the aperture at the point (r,θ). The imaginary exponential form of this integrand means that the integral is technically a Fourier transform, Chapter 15. In general, theFraunhofer diffraction pattern is given by theFourier transform of the aperture. 8Wecould alsoreferto Exercise11.1.16(b). 682 Chapter 11 Bessel Functions Table 11.1 Zeros oftheBesselFunctionsandTheirFirstDerivatives Numberof zero J0(x) J 1(x) J 2(x) J 3(x) J 4(x) J 5(x) 12 .4048 3 .8317 5 .1356 6 .3802 7 .5883 8 .7715 25 .5201 7 .0156 8 .4172 9 .7610 11 .0647 12 .3386 38 .6537 10 .1735 11 .6198 13 .0152 14 .3725 15 .7002 41 1.7915 13 .3237 14 .7960 16 .2235 17 .6160 18 .9801 51 4.9309 16 .4706 17 .9598 19 .4094 20 .8269 22 .2178 J′ 0(x)aJ′ 1(x) J′ 2(x) J′ 3(x) 13 .8317 1 .8412 3 .0542 4 .2012 27 .0156 5 .3314 6 .7061 8 .0152 31 0.1735 8 .5363 9 .9695 11 .3459 aJ′ 0(x)=−J1(x). FromTable11.1,whichliststhezerosoftheBesselfunctionsandtheirfirstderivatives,9 Eq.(11.35) willhaveazeroat 2πa λsinα=3.8317..., (11.36) or sinα=3.8317λ 2πa. (11.37) Forgreenlight, λ=5.5×10−5cm.Hence,if a=0.5c m , α≈sinα=6.7×10−5(radian)≈14secondsofarc , (11.38) whichshowsthatthebendingorspreadingofthelightrayisextremelysmall.Ifthisanaly- sis had been known in the seventeenth century, the arguments against the wave theory of light would have collapsed. In mid-twentieth century this same diffraction pattern appears in the scattering of nuclear particles by atomic nuclei—a striking demonstration of the wavepropertiesofthenuclearparticles. /squaresolid A further example of the use of Bessel functions and their roots is provided by the electromagnetic resonant cavity (Example 11.1.2) and the example and exercises of Sec- tion11.2. Example 11.1.2 CYLINDRICAL RESONANT CAVITY The propagation of electromagnetic waves in hollow metallic cylinders is important in many practical devices. If the cylinder has end surfaces, it is called a cavity. Resonant cavitiesplayacrucialroleinmanyparticleaccelerators. 9Additional roots of the Bessel functions and their first derivatives may be found in C. L. Beattie, Table of first 700 zeros of Bessel functions. Bell Syst. Tech. J. 37: 689 (1958), and Bell Monogr. 3055. Roots may be accessed in Mathematica and other symbolic software and are on the Web. 11.1 Bessel Functions of the First Kind, Jν(x) 683 FIGURE 11.3Cylindricalresonant cavity. Wetakethe z-axisalongthecenterofthecavitywithendsurfacesat z=0andz=land use cylindricalcoordinates suggested by the geometry. Its walls are perfect conductors, so thetangentialelectricfieldvanishesonthem(as inFig.11.3): Ez=0=Eϕforρ=a, E ρ=0=Eϕforz=0,l. Inside the cavity we have a vacuum, so ε0µ0=1/c2.In the interior of a resonant cav- ity, electromagnetic waves oscillate with harmonic time dependence e−iωt,which follows from separating the time from the spatial variables in Maxwell’s equations (Section 1.9), so ∇×∇×E=−1 c2∂2E ∂t2=α2E,α=ω c. With∇·E=0 (vacuum, no charges) and Eq. (1.85), we obtain for the space part of the electricfield ∇2E+α2E=0, which is called the vector Helmholtz PDE .T h ez-component ( Ez, space part only) satis- fiesthescalarHelmholtzequation, ∇2Ez+α2Ez=0. (11.39) The transverse electric field components E⊥=(Eρ,Eϕ)obey the same PDE but different boundaryconditions,givenearlier.Once Ezisknown,Maxwell’sequationsdetermine Eϕ fully.SeeJackson, Electrodynamics inAdditionalReadingsfor details. 684 Chapter 11 Bessel Functions We separate the zvariable from ρandϕ,because there are no mixed derivatives∂2Ez ∂z∂ρ, etc.Theproductsolution, Ez=v(ρ,ϕ)w(z), issubstitutedintotheHelmholtzPDEfor Ez usingEq. (2.35) for ∇2incylindricalcoordinates,andthenwedivideby vw,yielding 1 w(z)d2w dz2+1 vparenleftbigg∂2v ∂ρ2+1 ρ∂v ∂ρ+1 ρ2∂2v ∂ϕ2+α2parenrightbigg v(ρ,ϕ)=0. Thisimplies −1 w(z)d2w dz2=1 v(ρ,ϕ)parenleftbigg∂2v ∂ρ2+1 ρ∂v ∂ρ+1 ρ2∂2v ∂ϕ2+α2vparenrightbigg =k2. Here,k2isaseparationconstant,becausetheleft-andright-handsidesdependondifferent variables.For w(z)wefindtheharmonicoscillatorODEwithstandingwavesolution(not transients)thatweseek, w(z)=Asinkz+Bcoskz, withA,Bconstants.For v(ρ,ϕ)weobtain ∂2v ∂ρ2+1 ρ∂v ∂ρ+1 ρ2∂2v ∂ϕ2+γ2v=0,γ2=α2−k2. In this PDE we can separate the ρandϕvariables, because there is no mixed term∂2v ∂ρ∂ϕ. Theproductform v=u(ρ)/Phi1(ϕ) yields ρ2 u(ρ)parenleftbiggd2u dρ2+1 ρdu dρ+γ2parenrightbigg =−1 /Phi1(ϕ)d2/Phi1 dϕ2=m2, wherethe separationconstant m2mustbeaninteger ,becausetheangularsolution /Phi1= eimϕoftheODE d2/Phi1 dϕ2+m2/Phi1=0 mustbeperiodicintheazimuthalangle. This leavesuswiththeradialODE d2u dρ2+1 ρdu dρ+parenleftbigg γ2−m2 ρ2parenrightbigg u=0. Dimensionalargumentssuggestrescaling ρ→r=γρanddividingby γ2,whichyields d2u dr2+1 rdu dr+parenleftbigg 1−m2 r2parenrightbigg u=0. ThisisBessel’sODEfor ν=m.Weusetheregularsolution Jm(γρ)becausethe(irregular) second independent solution is singular at the origin, which is unacceptable here. The completesolutionis Ez=Jm(γρ)eimϕ(Asinkz+Bcoskz), (11.40a) wheretheconstant γisdeterminedfromthe boundarycondition Ez=0onthecavitysur- faceρ=a,thatis,that γabearootoftheBesselfunction Jm(seeTable11.1).Thisgives adiscreteset ofvalues γ=γmn,wherendesignatesthe nthrootof Jm(see Table11.1). 11.1 Bessel Functions of the First Kind, Jν(x) 685 ForthetransversemagneticorTMmodeofoscillationwith Hz=0Maxwell’sequations imply. (See again Resonant Cavities in J. D. Jackson’s Electrodynamics in Additional Readings.) E⊥∼∇⊥∂Ez ∂z,∇⊥=parenleftbigg∂ ∂ρ,1 ρ∂ ∂ϕparenrightbigg . Theformofthisresultsuggests Ez∼coskz,thatis,setting A=0 sothatE⊥∼sinkz=0 atz=0,lcanbesatisfiedby k=pπ l,p=0,1,2,.... (11.41) Thus,the tangential electricfields EρandEϕvanishat z=0andl.Inotherwords, A=0 correspondsto dEz/dz=0a tz=0 andz=lfor theTM mode.Altogetherthen,wehave γ2=ω2 c2−k2=ω2 c2−p2π2 l2, (11.42) with γ=γmn=αmn a, (11.43) whereαmnis thenthzeroof Jm. Thegeneralsolution Ez=summationdisplay m,n,pJm(γmnρ)e±imϕBmnpcospπz l, (11.40b) withconstants Bmnp,nowfollowsfromthesuperpositionprinciple. The result of the two boundary conditions and the separation constant m2is that the angularfrequencyofouroscillationdependsonthreediscreteparameters: ωmnp=cradicalBigg α2mn a2+p2π2 l2,  m=0,1,2,..., n=1,2,3,..., p=0,1,2....(11.44) ThesearetheallowableresonantfrequenciesforourTMmode.TheTEmodeofoscillation isthetopicof Exercise11.1.26. /squaresolid Alternate Approaches Bessel functions are introduced here by means of a generating function, Eq. (11.2). Other approachesarepossible.Listingthevariouspossibilities,wehave: 1. Generatingfunction(magic),Eq. (11.2). 2. SeriessolutionofBessel’s differentialequation,Section9.5. 3. Contour integrals: Some writers prefer to start with contour integral definitions of the Hankel functions, Section 7.3 and 11.4, and develop the Bessel function Jν(x)from theHankelfunctions. 686 Chapter 11 Bessel Functions 4. Direct solution of physical problems: Example 11.1.1. Fraunhofer diffraction with a circular aperture, illustrates this. Incidentally, Eq. (11.31) can be treated by series ex- pansion,ifdesired.Feynman10developsBesselfunctionsfromaconsiderationofcav- ityresonators. In case the generating function seems too arbitrary, it can be derived from a contour inte- gral, Exercise 11.1.16, or from the Bessel function recurrence relations, Exercise 11.1.6. Notethatthecontourintegralisnotlimitedtointeger ν,thusprovidingastartingpointfor developingBesselfunctions. Bessel Functions of Nonintegral Order These different approaches are not exactly equivalent. The generating function approach is very convenient for deriving two recurrence relations, Bessel’s differential equation, integral representations, addition theorems (Exercise 11.1.2), and upper and lower bounds (Exercise 11.1.1). However, you will probably have noticed that the generating function defined only Bessel functions of integral order, J0,J1,J2, and so on. This is a limitation of the generating function approach that can be avoided by using the contour integral in Exercise11.1.16instead,thusleadingtoforegoingapproach(3).ButtheBesselfunctionof thefirstkind, Jν(x),mayeasilybedefinedfornonintegral νbyusingtheseries(Eq.(11.5)) asanewdefinition. Therecurrencerelationsmaybeverifiedbysubstitutingintheseriesformof Jν(x)(Ex- ercise11.1.7).FromtheserelationsBessel’sequationfollows.Infact,if νisnotaninteger, there is actually an important simplification. It is found that JνandJ−νare independent, fornorelationoftheformofEq.(11.8)exists.Ontheotherhand,for ν=n,aninteger,we need another solution. The development of this second solution and an investigation of its propertiesform thesubjectofSection11.3. Exercises 11.1.1 Fromtheproductofthegeneratingfunctions g(x,t)·g(x,−t)showthat 1=bracketleftbig J0(x)bracketrightbig2+2bracketleftbig J1(x)bracketrightbig2+2bracketleftbig J2(x)bracketrightbig2+··· andthereforethat |J0(x)|≤1 and|Jn(x)|≤1/√ 2,n=1,2,3,.... Hint.Useuniquenessofpowerseries, Section5.7. 11.1.2 Usingageneratingfunction g(x,t)=g(u+v,t)=g(u,t)·g(v,t), showthat (a)Jn(u+v)=∞summationdisplay s=−∞Js(u)·Jn−s(v), (b)J0(u+v)=J0(u)J0(v)+2∞summationdisplay s=1Js(u)J−s(v). 10R. P. Feynman, R. B. Leighton, and M. Sands, The Feynman Lectures on Physics , Vol. II. Reading, MA: Addison-Wesley (1964), Chapter 23. 11.1 Bessel Functions of the First Kind, Jν(x) 687 Theseareadditiontheoremsfor theBesselfunctions. 11.1.3 Usingonlythegeneratingfunction e(x/2)(t−1/t)=∞summationdisplay n=−∞Jn(x)tn andnottheexplicitseriesformof Jn(x),showthat Jn(x)hasoddorevenparityaccord- ingtowhether nisoddoreven,thatis,11 Jn(x)=(−1)nJn(−x). 11.1.4 DerivetheJacobi–Angerexpansion eizcosθ=∞summationdisplay m=−∞imJm(z)eimθ. Thisis anexpansionof aplanewaveinaseries ofcylindricalwaves. 11.1.5 Showthat (a) cosx=J0(x)+2∞summationdisplay n=1(−1)nJ2n(x), (b) sinx=2∞summationdisplay n=0(−1)nJ2n+1(x). 11.1.6 To help remove the generating function from the realm of magic, show that it can be derivedfrom therecurrencerelation,Eq.(11.10). Hint. (a) Assumeageneratingfunctionof theform g(x,t)=∞summationdisplay m=−∞Jm(x)tm. (b) MultiplyEq.(11.10) by tnandsumover n. (c) Rewritetheprecedingresultas parenleftbigg t+1 tparenrightbigg g(x,t)=2t x∂g(x,t) ∂t. (d) Integrate and adjust the “constant” of integration (a function of x) so that the coefficientof thezerothpower, t0,isJ0(x), as givenbyEq.(11.5). 11.1.7 Show,bydirectdifferentiation,that Jν(x)=∞summationdisplay s=0(−1)s s!(s+ν)!parenleftbiggx 2parenrightbiggν+2s 11This is easily seenfrom theseries form (Eq. (11.5)). 688 Chapter 11 Bessel Functions satisfiesthetworecurrencerelations Jν−1(x)+Jν+1(x)=2ν xJν(x), Jν−1(x)−Jν+1(x)=2J′ ν(x), andBessel’s differentialequation x2J′′ ν(x)+xJ′ ν(x)+parenleftbig x2−ν2parenrightbig Jν(x)=0. 11.1.8 Provethat sinx x=integraldisplayπ/2 0J0(xcosθ)cosθdθ,1−cosx x=integraldisplayπ/2 0J1(xcosθ)dθ. Hint.Thedefiniteintegral integraldisplayπ/2 0cos2s+1θdθ=2·4·6···(2s) 1·3·5···(2s+1) maybeuseful. 11.1.9 Showthat J0(x)=2 πintegraldisplay1 0cosxt√ 1−t2dt. This integral is a Fourier cosine transform (compare Section 15.3). The corresponding Fouriersinetransform, J0(x)=2 πintegraldisplay∞ 1sinxt√ t2−1dt, is established in Section 11.4 (Exercise 11.4.6) using a Hankel function integral repre- sentation. 11.1.10 Derive Jn(x)=(−1)nxnparenleftbigg1 xd dxparenrightbiggn J0(x). Hint.Trymathematicalinduction. 11.1.11 Show that between any two consecutive zeros of Jn(x)there is one and only one zero ofJn+1(x). Hint.Equations(11.15)and(11.17)maybeuseful. 11.1.12 An analysis of antenna radiation patterns for a system with a circular aperture involves theequation g(u)=integraldisplay1 0f(r)J0(ur)rdr. Iff(r)=1−r2, showthat g(u)=2 u2J2(u). 11.1 Bessel Functions of the First Kind, Jν(x) 689 11.1.13 The differential cross section in a nuclear scattering experiment is given by dσ/d/Omega1= |f(θ)|2. Anapproximatetreatmentleadsto f(θ)=−ik 2πintegraldisplay2π 0integraldisplayR 0exp[ikρsinθsinϕ]ρdρdϕ. Hereθis an angle through which the scattered particle is scattered. Ris the nuclear radius.Showthat dσ d/Omega1=parenleftbig πR2parenrightbig1 πbracketleftbiggJ1(kRsinθ) sinθbracketrightbigg2 . 11.1.14 Asetof functions Cn(x)satisfiestherecurrencerelations Cn−1(x)−Cn+1(x)=2n xCn(x), Cn−1(x)+Cn+1(x)=2C′ n(x). (a) Whatlinearsecond-orderODEdoesthe Cn(x)satisfy? (b) By a change of variable transform your ODE into Bessel’s equation. This sug- gests that Cn(x)may be expressed in terms of Bessel functions of transformed argument. 11.1.15 A particle (mass m) is contained in a right circular cylinder (pillbox) of radius Rand heightH.TheparticleisdescribedbyawavefunctionsatisfyingtheSchrödingerwave equation −¯h2 2m∇2ψ(ρ,ϕ,z)=Eψ(ρ,ϕ,z) andtheconditionthatthewavefunctiongotozerooverthesurfaceofthepillbox.Find thelowest(zeropoint)permittedenergy. ANS.E=¯h2 2mbracketleftbiggparenleftbiggzpq Rparenrightbigg2 +parenleftbiggnπ Hparenrightbigg2bracketrightbigg , Emin=¯h2 2mbracketleftbiggparenleftbigg2.405 Rparenrightbigg2 +parenleftbiggπ Hparenrightbigg2bracketrightbigg , wherezpqis theqthzeroof Jpandtheindex pis fixedbytheazimuthaldependence. 11.1.16 (a) Showbydirectdifferentiationandsubstitutionthat Jν(x)=1 2πiintegraldisplay Ce(x/2)(t−1/t)t−ν−1dt orthattheequivalentequation, Jν(x)=1 2πiparenleftbiggx 2parenrightbiggνintegraldisplay es−x2/4ss−ν−1ds, satisfies Bessel’s equation. Cis the contour shown in Fig. 11.4. The negative real axisisthecutline. 690 Chapter 11 Bessel Functions FIGURE 11.4Besselfunctioncontour. Hint.Showthatthetotalintegrand(aftersubstitutinginBessel’sdifferentialequa- tion)maybewrittenas atotalderivative: d dtbraceleftbigg expbracketleftbiggx 2parenleftbigg t−1 tparenrightbiggbracketrightbigg t−νbracketleftbigg ν+x 2parenleftbigg t+1 tparenrightbiggbracketrightbiggbracerightbigg . (b) Showthatthefirst integral(with naninteger)maybetransformedinto Jn(x)=1 2πintegraldisplay2π 0ei(xsinθ−nθ)dθ=i−n 2πintegraldisplay2π 0ei(xcosθ+nθ)dθ. 11.1.17 The contour Cin Exercise 11.1.16 is deformed to the path −∞to−1, unit circle e−iπ toeiπ,andfinally−1t o−∞. Showthat Jν(x)=1 πintegraldisplayπ 0cos(νθ−xsinθ)dθ−sinνπ πintegraldisplay∞ 0e−νθ−xsinhθdθ. Thisis Bessel’s integral. Hint.Thenegativevaluesofthevariableof integration umaybehandledbyusing u=te±ix. 11.1.18 (a) Showthat Jν(x)=2 π1/2(ν−1 2)!parenleftbiggx 2parenrightbiggνintegraldisplayπ/2 0cos(xsinθ)cos2νθdθ, whereν>−1 2. Hint. Here is a chance to use series expansion and term-by-term integration. The formulasofSection8.4willproveuseful. 11.1 Bessel Functions of the First Kind, Jν(x) 691 (b) Transform theintegralinpart(a)into Jν(x)=1 π1/2(ν−1 2)!parenleftbiggx 2parenrightbiggνintegraldisplayπ 0cos(xcosθ)sin2νθdθ =1 π1/2(ν−1 2)!parenleftbiggx 2parenrightbiggνintegraldisplayπ 0e±ixcosθsin2νθdθ =1 π1/2(ν−1 2)!parenleftbiggx 2parenrightbiggνintegraldisplay1 −1e±ipx(1−p2)ν−1/2dp. Thesearealternateintegralrepresentationsof Jν(x). 11.1.19 (a) From Jν(x)=1 2πiparenleftbiggx 2parenrightbiggνintegraldisplay t−ν−1et−x2/4tdt derivetherecurrencerelation J′ ν(x)=ν xJν(x)−Jν+1(x). (b) From Jν(x)=1 2πiintegraldisplay t−ν−1e(x/2)(t−1/t)dt derivetherecurrencerelation J′ ν(x)=1 2bracketleftbig Jν−1(x)−Jν+1(x)bracketrightbig . 11.1.20 Showthattherecurrencerelation J′ n(x)=1 2bracketleftbig Jn−1(x)−Jn+1(x)bracketrightbig followsdirectlyfromdifferentiationof Jn(x)=1 πintegraldisplayπ 0cos(nθ−xsinθ)dθ. 11.1.21 Evaluate integraldisplay∞ 0e−axJ0(bx)dx, a,b> 0. Actuallytheresults holdfor a≥0,−∞<b<∞.This isaLaplacetransformof J0. Hint.Eitheranintegralrepresentationof J0oraseries expansionwillbehelpful. 11.1.22 Usingtrigonometricforms, verifythat J0(br)=1 2πintegraldisplay2π 0eibrsinθdθ. 11.1.23 (a) Plot the intensity ( /Phi12of Eq. (11.35)) as a function of (sinα/λ)along a diameter ofthecirculardiffractionpattern.Locatethefirst twominima. 692 Chapter 11 Bessel Functions (b) Whatfractionofthetotallightintensityfalls withinthecentralmaximum? Hint.[J1(x)]2/xmay be written as a derivative and the area integral of the intensity integratedbyinspection. 11.1.24 Thefractionoflightincidentonacircularaperture(normalincidence)thatistransmitted isgivenby T=2integraldisplay2ka 0J2(x)dx x−1 2kaintegraldisplay2ka 0J2(x)dx. Hereaistheradiusoftheapertureand kisthewavenumber, 2 π/λ.Showthat (a)T=1−1 ka∞summationdisplay n=0J2n+1(2ka), (b)T=1−1 2kaintegraldisplay2ka 0J0(x)dx. 11.1.25 Theamplitude U(ρ,ϕ,t) ofavibratingcircularmembraneofradius asatisfiesthewave equation ∇2U−1 v2∂2U ∂t2=0. Herevis the phase velocity of the wave fixed by the elastic constants and whatever dampingis imposed. (a) Showthatasolutionis U(ρ,ϕ,t)=Jm(kρ)parenleftbig a1eimϕ+a2e−imϕparenrightbigparenleftbig b1eiωt+b2e−iωtparenrightbig . (b) From the Dirichlet boundary condition, Jm(ka)=0, find the allowable values of thewavelength λ(k=2π/λ). Note. There are other Bessel functions besides Jm, but they all diverge at ρ=0. This is shown explicitly in Section 11.3. The divergent behavior is actually implicit inEq. (11.6). 11.1.26 Example 11.1.2 describes the TM modes of electromagnetic cavity oscillation. The transverse electric (TE) modes differ, in that we work from the zcomponent of the magneticinduction B: ∇2Bz+α2Bz=0 withboundaryconditions Bz(0)=Bz(l)=0 and∂Bz ∂ρvextendsinglevextendsinglevextendsinglevextendsingle ρ=0=0. ShowthattheTEresonantfrequenciesaregivenby ωmnp=cradicalBigg β2mn a2+p2π2 l2,p=1,2,3,.... 11.1 Bessel Functions of the First Kind, Jν(x) 693 11.1.27 Plot the three lowest TM and the three lowest TE angular resonant frequencies, ωmnp, asafunctionoftheradius/length (a/l)ratiofor 0≤a/l≤1.5. Hint.Tryplotting ω2(inunitsof c2/a2)v e r s u s(a/l)2. Whythischoice? 11.1.28 A thin conducting disk of radius acarries a charge q. Show that the potential is de- scribedby ϕ(r,z)=q 4πε0aintegraldisplay∞ 0e−k|z|J0(kr)sinka kdk, whereJ0is the usual Bessel function and randzare the familiar cylindrical coordi- nates. Note. This is a difficult problem. One approach is through Fourier transforms such as Exercise15.3.11.ForadiscussionofthephysicalproblemseeJackson( ClassicalElec- trodynamics inAdditionalReadings). 11.1.29 Showthat integraldisplaya 0xmJn(x)dx, m ≥n≥0, (a) is integrable in terms of Bessel functions and powers of x(such asapJq(a))f o r m+nodd; (b) maybereducedtointegratedtermsplusintegraltexta 0J0(x)dxform+neven. 11.1.30 Showthat integraldisplayα0n 0parenleftbigg 1−y α0nparenrightbigg J0(y)ydy=1 α0nintegraldisplayα0n 0J0(y)dy. Hereα0nis thenth root of J0(y). This relation is useful (see Exercise 11.2.11): The expression on the right is easier and quicker to evaluate—and much more accurate. Taking the difference of two terms in the expression on the left leads to a large relative error. 11.1.31 The circular aperature diffraction amplitude /Phi1of Eq. (17.35) is proportional to f(z)= J1(z)/z. The corresponding single slit diffraction amplitude is proportional to g(z)= sinz/z. (a) Calculateandplot f(z)andg(z)forz=0.0(0.2)12.0. (b) Locatethetwolowestvaluesof z(z>0)forwhich f(z)takesonanextremevalue. Calculatethecorrespondingvaluesof f(z). (c) Locatethetwolowestvaluesof z(z>0)forwhich g(z)takesonanextremevalue. Calculatethecorrespondingvaluesof g(z). 11.1.32 Calculate the electrostatic potential of a charged disk ϕ(r,z)from the integral form of Exercise 11.1.28. Calculate the potential for r/a=0.0(0.5)2.0 andz/a= 0.25(0.25)1.25. Why is z/a=0 omitted? Exercise 12.3.17 is a spherical harmonic versionof thissameproblem. 694 Chapter 11 Bessel Functions 11.2 O RTHOGONALITY IfBessel’sequation,Eq.(11.22a),isdividedby ρ,weseethatitbecomesself-adjoint,and therefore, by the Sturm–Liouville theory, Section 10.2, the solutions are expected to be orthogonal—if we can arrange to have appropriate boundary conditions satisfied. To take care of the boundary conditions for a finite interval [0,a], we introduce parameters aand ανmintotheargumentof JνtogetJν(ανmρ/a).Hereaistheupperlimitofthecylindrical radialcoordinate ρ. FromEq. (11.22a), ρd2 dρ2Jνparenleftbigg ανmρ aparenrightbigg +d dρJνparenleftbigg ανmρ aparenrightbigg +parenleftbiggα2 νmρ a2−ν2 ρparenrightbigg Jνparenleftbigg ανmρ aparenrightbigg =0.(11.45) Changingtheparameter ανmtoανn,wefindthat Jν(ανnρ/a)satisfies ρd2 dρ2Jνparenleftbigg ανnρ aparenrightbigg +d dρJνparenleftbigg ανnρ aparenrightbigg +parenleftbiggα2 νnρ a2−ν2 ρparenrightbigg Jνparenleftbigg ανnρ aparenrightbigg =0.(11.45a) Proceeding as in Section 10.2, we multiply Eq. (11.45) by Jν(ανnρ/a)and Eq. (11.45a) byJν(ανmρ/a)andsubtract,obtaining Jνparenleftbigg ανnρ aparenrightbiggd dρbracketleftbigg ρd dρJνparenleftbigg ανmρ aparenrightbiggbracketrightbigg −Jνparenleftbigg ανmρ aparenrightbiggd dρbracketleftbigg ρd dρJνparenleftbigg ανnρ aparenrightbiggbracketrightbigg =α2 νn−α2 νm a2ρJνparenleftbigg ανmρ aparenrightbigg Jνparenleftbigg ανnρ aparenrightbigg . (11.46) Integratingfrom ρ=0t oρ=a, weobtain integraldisplaya 0Jνparenleftbigg ανnρ aparenrightbiggd dρbracketleftbigg ρd dρJνparenleftbigg ανmρ aparenrightbiggbracketrightbigg dρ−integraldisplaya 0Jνparenleftbigg ανmρ aparenrightbiggd dρbracketleftbigg ρd dρJνparenleftbigg ανnρ aparenrightbiggbracketrightbigg dρ =α2 νn−α2 νm a2integraldisplaya 0Jνparenleftbigg ανmρ aparenrightbigg Jνparenleftbigg ανnρ aparenrightbigg ρdρ. (11.47) Uponintegratingbyparts, wesee thattheleft-handsideof Eq.(11.47) becomes vextendsinglevextendsinglevextendsinglevextendsingleρJνparenleftbigg ανnρ aparenrightbiggd dρJνparenleftbigg ανmρ aparenrightbiggvextendsinglevextendsinglevextendsinglevextendsinglea 0−vextendsinglevextendsinglevextendsinglevextendsingleρJνparenleftbigg ανmρ aparenrightbiggd dρJνparenleftbigg ανnρ aparenrightbiggvextendsinglevextendsinglevextendsinglevextendsinglea 0.(11.48) Forν≥0 the factor ρguarantees a zero at the lower limit, ρ=0. Actually the lower limit on the index νmay be extended down to ν>−1, Exercise 11.2.4.12Atρ=a, each expression vanishes if we choose the parameters ανnandανmto be zeros, or roots of Jν; thatis,Jν(ανm)=0.Thesubscriptsnowbecomemeaningful: ανmis themthzeroof Jν. Withthischoiceofparameters,theleft-handsidevanishes(theSturm–Liouvillebound- aryconditionsaresatisfied)andfor m/negationslash=n, integraldisplaya 0Jνparenleftbigg ανmρ aparenrightbigg Jνparenleftbigg ανnρ aparenrightbigg ρdρ=0. (11.49) Thisgivesusorthogonalityovertheinterval [0,a]. 12Thecase ν=−1 reverts to ν=+1, Eq.(11.8). 11.2 Orthogonality 695 Normalization The normalization integral may be developed by returning to Eq. (11.48), setting ανn= ανm+ε, and taking the limit ε→0 (compare Exercise 11.2.2). With the aid of the recur- rencerelation,Eq. (11.16),theresultmaybewrittenas integraldisplaya 0bracketleftbigg Jνparenleftbigg ανmρ aparenrightbiggbracketrightbigg2 ρd ρ=a2 2bracketleftbig Jν+1(ανm)bracketrightbig2. (11.50) Bessel Series If we assume that the set of Bessel functions Jν(ανmρ/a))(νfixed,m=1,2,3,...)i s complete, then any well-behaved but otherwise arbitrary function f(ρ)may be expanded inaBesselseries (Bessel–FourierorFourier–Bessel) f(ρ)=∞summationdisplay m=1cνmJνparenleftbigg ανmρ aparenrightbigg ,0≤ρ≤a, ν>−1.(11.51) Thecoefficients cνmaredeterminedbyusingEq. (11.50), cνm=2 a2[Jν+1(ανm)]2integraldisplaya 0f(ρ)Jνparenleftbigg ανmρ aparenrightbigg ρdρ. (11.52) A similar series expansion involving Jν(βνmρ/a)with(d/dρ)J ν(βνmρ/a)|ρ=a=0i s includedinExercises11.2.3and11.2.6(b). Example 11.2.1 ELECTROSTATIC POTENTIAL IN A HOLLOW CYLINDER From Table 9.3 of Section 9.3 (with αreplaced by k), our solution of Laplace’s equation incircularcylindricalcoordinatesisalinearcombinationof ψkm(ρ,ϕ,z)=Jm(kρ)[amsinmϕ+bmcosmϕ]bracketleftbig c1ekz+c2e−kzbracketrightbig .(11.53) Theparticularlinearcombinationisdeterminedbytheboundaryconditionstobesatisfied. Ourcylinderherehasaradius aandaheight l.Thetopendsectionhasapotentialdistrib- utionψ(ρ,ϕ). Elsewhere on the surface the potential is zero.13The problem is to find the electrostaticpotential ψ(ρ,ϕ,z)=summationdisplay k,mψkm(ρ,ϕ,z) (11.54) everywhereintheinterior. For convenience, the circular cylindrical coordinates are placed as shown in Fig. 11.3. Sinceψ(ρ,ϕ,0)=0, we take c1=−c2=1 2.T h ezdependence becomes sinh kz, vanish- ing atz=0. The requirement that ψ=0 on the cylindrical sides is met by requiring the separationconstant ktobe k=kmn=αmn a, (11.55) 13Ifψ=0a tz=0,l,b u tψ/negationslash=0f o rρ=a, the modified Bessel functions, Section 11.5, are involved. 696 Chapter 11 Bessel Functions where the first subscript, m, gives the index of the Bessel function, whereas the second subscriptidentifiestheparticularzeroof Jm. Theelectrostaticpotentialbecomes ψ(ρ,ϕ,z)=∞summationdisplay m=0∞summationdisplay n=1Jmparenleftbigg αmnρ aparenrightbigg ·[amnsinmϕ+bmncosmϕ]·sinhparenleftbigg αmnz aparenrightbigg .(11.56) Equation(11.56)is adoubleseries:aBessel seriesin ρandaFourierseriesin ϕ. Atz=l,ψ=ψ(ρ,ϕ),aknownfunctionof ρandϕ. Therefore ψ(ρ,ϕ)=∞summationdisplay m=0∞summationdisplay n=1Jmparenleftbigg αmnρ aparenrightbigg ·[amnsinmϕ+bmncosmϕ]·sinhparenleftbigg αmnl aparenrightbigg . (11.57) The constants amnandbmnare evaluated by using Eqs. (11.49) and (11.50) and the corre- spondingequationsfor sin ϕand cosϕ(Example10.2.1andEqs.(14.2), (14.3),(14.15)to (14.17)). Wefind14 amn bmnbracerightbigg =2bracketleftbigg πa2sinhparenleftbigg αmnl aparenrightbigg J2 m+1(αmn)bracketrightbigg−1 ·integraldisplay2π 0integraldisplaya 0ψ(ρ,ϕ)J mparenleftbigg αmnρ aparenrightbiggbraceleftbiggsinmϕ cosmϕbracerightbigg ρdρdϕ. (11.58) These are definite integrals, that is, numbers. Substitutingback into Eq. (11.56), the series isspecifiedandthepotential ψ(ρ,ϕ,z) is determined. /squaresolid Continuum Form The Bessel series, Eq. (11.51), and Exercise 11.2.6 apply to expansions over the finite interval[0,a].I fa→∞, then the series forms may be expected to go over into integrals. Thediscreteroots ανmbecomeacontinuousvariable α.Asimilarsituationisencountered intheFourierseries,Section15.2.ThedevelopmentoftheBesselintegralfromtheBessel seriesis leftasExercise11.2.8. ForoperationswithacontinuumofBesselfunctions, Jν(αρ),akeyrelationistheBessel functionclosureequation, integraldisplay∞ 0Jν(αρ)Jν(α′ρ)ρdρ=1 αδ(α−α′), ν >−1 2. (11.59) ThismaybeprovedbytheuseofHankeltransforms,Section15.1.Analternateapproach, startingfromarelationsimilartoEq.(10.82),isgivenbyMorseandFeshbach,Section6.3. Asecondkindoforthogonality(varyingtheindex)isdevelopedforsphericalBesselfunc- tionsinSection11.7. 14Ifm=0,the factor 2 is omitted (compare Eq. (14.16)). 11.2 Orthogonality 697 Exercises 11.2.1 Showthat parenleftbig a2−b2parenrightbigintegraldisplayP 0Jν(ax)Jν(bx)xdx=Pbracketleftbig bJν(aP)J′ ν(bP)−aJ′ ν(aP)Jν(bP)bracketrightbig , with J′ ν(aP)=d d(ax)Jν(ax)vextendsinglevextendsingle x=P, integraldisplayP 0bracketleftbig Jν(ax)bracketrightbig2xdx=P2 2braceleftbiggbracketleftbig J′ ν(aP)bracketrightbig2+parenleftbigg 1−ν2 a2P2parenrightbiggbracketleftbig Jν(aP)bracketrightbig2bracerightbigg ,ν>−1. Thesetwointegralsareusuallycalledthe firstandsecondLommelintegrals . Hint. We have the development of the orthogonality of the Bessel functions as an anal- ogy. 11.2.2 Showthat integraldisplaya 0bracketleftbigg Jνparenleftbigg ανmρ aparenrightbiggbracketrightbigg2 ρdρ=a2 2bracketleftbig Jν+1(ανm)bracketrightbig2,ν>−1. Hereανmis themthzeroof Jν. Hint.Withανn=ανm+ε,expandJν[(ανm+ε)ρ/a]aboutανmρ/abyaTaylorexpan- sion. 11.2.3 (a) Ifβνmis themth zero of (d/dρ)J ν(βνmρ/a), show that the Bessel functions are orthogonalovertheinterval [0,a]withanorthogonalityintegral integraldisplaya 0Jνparenleftbigg βνmρ aparenrightbigg Jνparenleftbigg βνnρ aparenrightbigg ρdρ=0,m/negationslash=n, ν >−1. (b) Derivethecorrespondingnormalizationintegral (m=n). ANS.a2 2parenleftbigg 1−ν2 β2νmparenrightbiggbracketleftbig Jν(βνm)bracketrightbig2,ν>−1. 11.2.4 Verify that the orthogonality equation, Eq. (11.49), and the normalization equation, Eq. (11.50),holdfor ν>−1. Hint.Usingpower-seriesexpansions,examinethebehaviorof Eq.(11.48) as ρ→0. 11.2.5 From Eq. (11.49) develop a proof that Jν(z),ν >−1, has no complex roots (with nonzeroimaginarypart). Hint. (a) Usetheseries form of Jν(z)toexcludepureimaginaryroots. (b) Assume ανmtobecomplexandtake ανntobeα∗ νm. 11.2.6 (a) Intheseriesexpansion f(ρ)=∞summationdisplay m=1cνmJνparenleftbigg ανmρ aparenrightbigg ,0≤ρ≤a, ν>−1, 698 Chapter 11 Bessel Functions withJν(ανm)=0,showthatthecoefficientsare givenby cνm=2 a2[Jν+1(ανm)]2integraldisplaya 0f(ρ)Jνparenleftbigg ανmρ aparenrightbigg ρdρ. (b) Intheseriesexpansion f(ρ)=∞summationdisplay m=1dνmJνparenleftbigg βνmρ aparenrightbigg ,0≤ρ≤a, ν>−1, with(d/dρ)J ν(βνmρ/a)|ρ=a=0,showthatthecoefficientsaregivenby dνm=2 a2(1−ν2/β2νm)[Jν(βνm)]2integraldisplaya 0f(ρ)Jνparenleftbigg βνmρ aparenrightbigg ρdρ. 11.2.7 Arightcircularcylinderhasanelectrostaticpotentialof ψ(ρ,ϕ)onbothends.Thepo- tentialonthecurvedcylindricalsurface iszero.Findthepotentialatallinteriorpoints. Hint. Choose your coordinate system and adjust your zdependenceto exploit the sym- metryofyourpotential. 11.2.8 Forthecontinuumcase,showthatEqs. (11.51) and(11.52) arereplacedby f(ρ)=integraldisplay∞ 0a(α)Jν(αρ)dα, a(α)=αintegraldisplay∞ 0f(ρ)Jν(αρ)ρdρ. Hint.ThecorrespondingcaseforsinesandcosinesisworkedoutinSection15.2.These are Hankel transforms. A derivation for the special case ν=0 is the topic of Exer- cise15.1.1. 11.2.9 Afunction f(x)is expressedasaBessel series: f(x)=∞summationdisplay n=1anJm(αmnx), withαmnthenthrootof Jm. ProvetheParsevalrelation, integraldisplay1 0bracketleftbig f(x)bracketrightbig2xdx=1 2∞summationdisplay n=1a2 nbracketleftbig Jm+1(αmn)bracketrightbig2. 11.2.10 Provethat ∞summationdisplay n=1(αmn)−2=1 4(m+1). Hint.Expand xminaBesselseries andapplytheParsevalrelation. 11.3 Neumann Functions 699 11.2.11 Arightcircularcylinderof length lhasapotential ψparenleftbigg z=±l 2parenrightbigg =100parenleftbigg 1−ρ aparenrightbigg , whereais the radius. The potential over the curved surface (side) is zero. Using the Bessel series from Exercise 11.2.7, calculate the electrostatic potential for ρ/a= 0.0(0.2)1.0 andz/l=0.0(0.1)0.5.Takea/l=0.5. Hint.FromExercise11.1.30youhave integraldisplayα0n 0parenleftbigg 1−y α0nparenrightbigg J0(y)ydy. Showthatthisequals 1 α0nintegraldisplayα0n 0J0(y)dy. Numerical evaluation of this latter form rather than the former is both faster and more accurate. Note.F o rρ/a=0.0 andz/l=0.5 the convergence is slow, 20 terms giving only 98.4 ratherthan100. Checkvalue .F orρ/a=0.4andz/l=0.3, ψ=24.558. 11.3 N EUMANN FUNCTIONS ,BESSEL FUNCTIONS OF THE SECOND KIND FromthetheoryofODEsitisknownthatBessel’sequationhastwoindependentsolutions. Indeed, for nonintegral order νwe have already found two solutions and labeled them Jν(x)andJ−ν(x), using the infinite series (Eq. (11.5)). The trouble is that when νis integral, Eq. (11.8) holds and we have but one independent solution. A second solution may be developed by the methods of Section 9.6. This yields a perfectly good second solutionofBessel’s equationbutis notthestandardform. Definition and Series Form Asanalternateapproach,wetaketheparticularlinearcombinationof Jν(x)andJ−ν(x) Nν(x)=cosνπJν(x)−J−ν(x) sinνπ. (11.60) This is the Neumann function (Fig. 11.5).15For nonintegral ν,Nν(x)clearly satisfies Bessel’s equation, for it is a linear combination of known solutions Jν(x)andJ−ν(x). 15In AMS-55 (see footnote 4 in Chapter 5 or Additional Readings of Chapter 8 p. for this ref.) and in most mathematics tables, this is labeled Yν(x). 700 Chapter 11 Bessel Functions FIGURE 11.5Neumannfunctions N0(x),N1(x), andN2(x). Substitutingthepower-seriesEq. (11.6)for n→ν(giveninExercise11.1.7) yields Nν(x)=−(ν−1)! πparenleftbigg2 xparenrightbiggν +···,16(11.61) forν>0.However,forintegral ν,ν=n,Eq.(11.8)appliesandEq.(11.60)16becomesin- determinate. The definition of Nν(x)was chosen deliberately for this indeterminate prop- erty. Again substituting the power series and evaluating Nν(x)forν→0 by l’Hôpital’s ruleforindeterminateforms, weobtainthelimitingvalue N0(x)=2 π(lnx+γ−ln2)+Oparenleftbig x2parenrightbig (11.62) forn=0 andx→0,using ν!(−ν)!=πν sinπν(11.63) fromEq.(8.32).ThefirstandthirdtermsinEq.(11.62)comefromusing (d/dν)(x/ 2)ν= (x/2)νln(x/2),whileγcomesfrom (d/dν)ν!forν→0 usingEqs.(8.38)and(8.40).For n>0 weobtainsimilarly Nn(x)=−1 π(n−1)!parenleftbigg2 xparenrightbiggn +···+2 πparenleftbiggx 2parenrightbiggn1 n!lnparenleftbiggx 2parenrightbigg +···. (11.64) Equations(11.62)and(11.64)exhibitthelogarithmicdependencethatwastobeexpected. This, ofcourse,verifiestheindependenceof JnandNn. 16Note thatthis limiting form applies to both integral and nonintegral values of the index ν. 11.3 Neumann Functions 701 Other Forms As with all the other Bessel functions, Nν(x)has integral representations. For N0(x)we have N0(x)=−2 πintegraldisplay∞ 0cos(xcosht)dt=−2 πintegraldisplay∞ 1cos(xt) (t2−1)1/2dt, x> 0. These forms can be derived as the imaginary part of the Hankel representations of Exer- cise11.4.7.Thelatterformis aFouriercosinetransform. Toverifythat Nν(x),ourNeumannfunction(Fig.11.5)orBesselfunctionofthesecond kind, actually does satisfy Bessel’s equation for integral n, we may proceed as follows. L’Hôpital’sruleappliedtoEq. (11.60)yields Nn(x)=(d/dν)[cosνπJν(x)−J−ν(x)] (d/dν)sinνπvextendsinglevextendsinglevextendsinglevextendsingle ν=n =−πsinnπJn(x)+[cosnπ∂Jν/∂ν−∂J−ν/∂ν]|ν=n πcosnπ =1 πbracketleftbigg∂Jν(x) ∂ν−(−1)n∂J−ν(x) ∂νbracketrightbiggvextendsinglevextendsinglevextendsinglevextendsingle ν=n. (11.65) DifferentiatingBessel’s equationfor J±ν(x)withrespectto ν,weha v e x2d2 dx2parenleftbigg∂J±ν ∂νparenrightbigg +xd dxparenleftbigg∂J±ν ∂νparenrightbigg +parenleftbig x2−ν2parenrightbig∂J±ν ∂ν=2νJ±ν. (11.66) Multiplying the equation for J−νby(−1)ν, subtracting from the equation for Jν(as sug- gestedbyEq. (11.65)), andtakingthelimit ν→n,weobtain x2d2 dx2Nn+xd dxNn+parenleftbig x2−n2parenrightbig Nn=2n πbracketleftbig Jn−(−1)nJ−nbracketrightbig . (11.67) Forν=n, an integer, the right-hand side vanishes by Eq. (11.8) and Nn(x)is seen to be a solutionofBessel’sequation.Themostgeneralsolutionforany νcanthereforebewritten as y(x)=AJν(x)+BNν(x). (11.68) It is seen from Eqs. (11.62) and (11.64) that Nndiverges, at least logarithmically. Any boundary condition that requires the solution to be finite at the origin (as in our vibrat- ing circular membrane (Section 11.1)) automatically excludes Nn(x). Conversely, in the absenceofsucharequirement, Nn(x)mustbeconsidered. To a certain extent the definition of the Neumann function Nn(x)is arbitrary. Equa- tions (11.62) and (11.64) contain terms of the form anJn(x). Clearly, any finite value of the constant anwould still give us a second solution of Bessel’s equation. Why should an havetheparticularvalueimplicitin Eqs. (11.62)and(11.64)?Theanswerinvolvestheas- ymptotic dependence developed in Section 11.6. If Jncorresponds to a cosine wave, then Nncorresponds to a sine wave. This simple and convenient asymptotic phase relationship isaconsequenceof theparticularadmixtureof JninNn. 702 Chapter 11 Bessel Functions Recurrence Relations SubstitutingEq.(11.60)for Nν(x)(nonintegral ν)intotherecurrencerelations(Eqs.(11.10) and(11.12)for Jn(x),weseeimmediatelythat Nν(x)satisfiesthesesamerecurrencerela- tions. This actually constitutes another proof that Nνis a solution. Note that the converse is not necessarily true. All solutions need not satisfy the same recurrence relations. An exampleofthis sortof troubleappearsinSection11.5. Wronskian Formulas From Section 9.6 and Exercise 10.1.4 we have the Wronskian formula17for solutions of theBessel equation, uν(x)v′ ν(x)−u′ ν(x)vν(x)=Aν x, (11.69) in which Aνis a parameter that depends on the particular Bessel functions uν(x)and vν(x)being considered. Aνis a constant in the sense that it is independent of x. Consider thespecialcase uν(x)=Jν(x), v ν(x)=J−ν(x), (11.70) JνJ′ −ν−J′ νJ−ν=Aν x. (11.71) SinceAνis a constant, it may be identified at any convenient point, such as x=0. Using thefirst termsintheseriesexpansions(Eqs. (11.5) and(11.6)), weobtain Jν→xν 2νν!,J−ν→2νx−ν (−ν)! J′ ν→νxν−1 2νν!,J′ −ν→−ν2νx−ν−1 (−ν)!. (11.72) SubstitutionintoEq.(11.69) yields Jν(x)J′ −ν(x)−J′ ν(x)J−ν(x)=−2ν xν!(−ν)!=−2sinνπ πx, (11.73) usingEq.(8.32).Notethat Aνvanishesforintegral ν,asitmust,sincethenonvanishingof the Wronskian is a test of the independence of the two solutions. By Eq. (11.73), Jnand J−nare clearlylinearlydependent. Using our recurrence relations, we may readily develop a large number of alternate forms, amongwhichare JνJ−ν+1+J−νJν−1=2sinνπ πx, (11.74) 17This result depends on P(x)of Section 9.5 being equal to p′(x)/p(x), the corresponding coefficient of the self-adjoint form of Section 10.1. 11.3 Neumann Functions 703 JνJ−ν−1+J−νJν+1=−2sinνπ πx, (11.75) JνN′ ν−J′ νNν=2 πx, (11.76) JνNν+1−Jν+1Nν=−2 πx. (11.77) Manymorewillbefoundinthereferencesgivenatchapter’send. You will recall that in Chapter 9 Wronskians were of great value in two respects: (1) in establishingthelinearindependenceorlineardependenceofsolutionsofdifferentialequa- tions and (2) in developing an integral form of a second solution. Here the specific forms oftheWronskiansandWronskian-derivedcombinationsofBesselfunctionsareusefulpri- marilytoillustratethegeneralbehaviorofthevariousBesselfunctions.Wronskiansareof great use in checking tables of Bessel functions. In Section 10.5 Wronskians appeared in connectionwithGreen’sfunctions. Example 11.3.1 COAXIAL WAVEGUIDES Weareinterestedinanelectromagneticwaveconfinedbetweentheconcentric,conducting cylindricalsurfaces ρ=aandρ=b.MostofthemathematicsisworkedoutinSection9.3 andExample11.1.2.Togofromthestandingwaveoftheseexamplestothetravelingwave here,welet A=iB,A=amn,B=bmninEq. (11.40a)andobtain Ez=summationdisplay m,nbmnJm(γρ)e±imϕei(kz−ωt). (11.78) Additional properties of the components of the electromagnetic wave in the simple cylin- drical wave guide are explored in Exercises 11.3.8 and 11.3.9. For the coaxial wave guide one generalization is needed. The origin, ρ=0, is now excluded (0<a≤ρ≤b). Hence theNeumannfunction Nm(γρ)maynotbeexcluded. Ez(ρ,ϕ,z,t) becomes Ez=summationdisplay m,nbracketleftbig bmnJm(γρ)+cmnNm(γρ)bracketrightbig e±imϕei(kz−ωt). (11.79) Withthecondition Hz=0, (11.80) wehavethebasicequationsfora TM(transversemagnetic)wave. The(tangential)electricfieldmustvanishattheconductingsurfaces(Dirichletboundary condition),or bmnJm(γa)+cmnNm(γa)=0, (11.81) bmnJm(γb)+cmnNm(γb)=0. (11.82) 704 Chapter 11 Bessel Functions These transcendental equations may be solved for γ(γmn)and the ratio cmn/bmn.F r o m Example11.1.2, k2=ω2µ0ε0−γ2=ω2 c2−γ2. (11.83) Sincek2must be positive for a real wave, the minimum frequency that will be propagated (inthisTM mode)is ω=γc, (11.84) withγfixed by the boundary conditions, Eqs. (11.81) and (11.82). This is the cutoff fre- quencyof thewaveguide. ThereisalsoaTE(transverseelectric)mode,with Ez=0 andHzgivenbyEq.(11.79). ThenwehaveNeumannboundaryconditionsinplaceofEqs.(11.81)and(11.82).Finally, for the coaxial guide (not for the plain cylindrical guide, a=0), a TEM (transverse elec- tromagnetic)mode, Ez=Hz=0, is possible. This corresponds to a plane wave, as in free space. The simpler cases (no Neumann functions, simpler boundary conditions) of a circular waveguideareincludedasExercises11.3.8and11.3.9. ToconcludethisdiscussionofNeumannfunctions,weintroducetheNeumannfunction Nν(x)forthefollowingreasons: 1. Itisasecond,independentsolutionofBessel’sequation,whichcompletesthegeneral solution. 2. It is required for specific physical problems such as electromagnetic waves in coaxial cablesandquantummechanicalscatteringtheory. 3. It leadstoaGreen’sfunctionfor theBesselequation(Sections9.7 and10.5). 4. It leadsdirectlytothetwoHankelfunctions(Section11.4)./squaresolid Exercises 11.3.1 ProvethattheNeumannfunctions Nn(withnaninteger)satisfytherecurrencerelations Nn−1(x)+Nn+1(x)=2n xNn(x), Nn−1(x)−Nn+1(x)=2N′ n(x). Hint.Theserelationsmaybeprovedbydifferentiatingtherecurrencerelationsfor Jνor byusingthelimitformof Nνbutnotdividingeverythingbyzero. 11.3.2 Showthat N−n(x)=(−1)nNn(x). 11.3.3 Showthat N′ 0(x)=−N1(x). 11.3 Neumann Functions 705 11.3.4 IfYandZare anytwosolutionsof Bessel’sequation,showthat Yν(x)Z′ ν(x)−Y′ ν(x)Zν(x)=Aν x, in which Aνmay depend on νbut is independent of x. This is a special case of Exer- cise10.1.4. 11.3.5 VerifytheWronskianformulas Jν(x)J−ν+1(x)+J−ν(x)Jν−1(x)=2sinνπ πx, Jν(x)N′ ν(x)−J′ ν(x)Nν(x)=2 πx. 11.3.6 Asanalternativetoletting xapproachzerointheevaluationoftheWronskianconstant, we may invoke uniqueness of power series (Section 5.7). The coefficient of x−1in the seriesexpansionof uν(x)v′ ν(x)−u′ ν(x)vν(x)isthenAν.Showbyseriesexpansionthat thecoefficientsof x0andx1ofJν(x)J′ −ν(x)−J′ ν(x)J−ν(x)are eachzero. 11.3.7 (a) BydifferentiatingandsubstitutingintoBessel’s ODE,showthat integraldisplay∞ 0cos(xcosht)dt isasolution. Hint.Youcanrearrangethefinalintegralas integraldisplay∞ 0d dtbraceleftbig xsin(xcosht)sinhtbracerightbig dt. (b) Showthat N0(x)=−2 πintegraldisplay∞ 0cos(xcosht)dt islinearlyindependentof J0(x). 11.3.8 A cylindrical wave guide has radius r0. Find the nonvanishing components of the elec- tricandmagneticfieldsfor (a) TM 01, transversemagneticwave (Hz=Hρ=Eϕ=0), (b) TE 01, transverseelectricwave (Ez=Eρ=Hϕ=0). The subscripts 01 indicate that the longitudinal component ( EzorHz)i n v o l v e s J0and theboundaryconditionissatisfiedbythe firstzeroofJ0orJ′ 0. Hint.Allcomponentsofthewavehavethesamefactor: exp i(kz−ωt). 11.3.9 Foragivenmodeofoscillationthe minimum frequencythatwillbepassedbyacircular cylindricalwaveguide(radius r0)i s νmin=c λc, 706 Chapter 11 Bessel Functions inwhich λcis fixedbytheboundarycondition Jnparenleftbigg2πr0 λcparenrightbigg =0forTMnmmode, J′ nparenleftbigg2πr0 λcparenrightbigg =0forTEnmmode. The subscript ndenotes the order of the Bessel function and mindicates the zero used. Find this cutoff wavelength λcfor the three TM and three TE modes with the longest cutoff wavelengths. Explain your results in terms of the graph of J0,J1, andJ2 (Fig.11.1). 11.3.10 Write a program that will compute successive roots of the Neumann function Nn(x), that isαns, whereNn(αns)=0. Tabulate the first five roots of N0,N1, andN2. Check your values for the roots against those listed in AMS-55 (see Additional Readings of Chapter8forthefullref.). Checkvalue. α12=5.42968. 11.3.11 For the case m=0,a=1, andb=2, the coaxial wave guide boundary conditions lead to f(x)=J0(2x) N0(2x)−J0(x) N0(x) (Fig.11.6). (a) Calculate f(x)forx=0.0(0.1)10.0 and plot f(x)versusxto find the approxi- matelocationoftheroots. FIGURE 11.6f(x)of Exercise11.3.11. 11.4 Hankel Functions 707 (b) Callaroot-findingsubroutinetodeterminethefirstthreerootstohigherprecision. ANS.3.1230,6.2734,9.4182. Note. The higher roots can be expected to appear at intervals whose length approaches n. Why? AMS-55 (see Additional Readings of Chapter 8 for the reference), gives an approximateformulafortheroots.Thefunction g(x)=J0(x)N0(2x)−J0(2x)N0(x)is muchbetterbehavedthan f(x)previouslydiscussed. 11.4 H ANKEL FUNCTIONS ManyauthorsprefertointroducetheHankelfunctionsbymeansofintegralrepresentations andthentousethemtodefinetheNeumannfunction Nν(z).Anoutlineofthisapproachis givenattheendofthis section. Definitions Because we have already obtained the Neumann function by more elementary (and less powerful)techniques,wemayuseittodefinetheHankelfunctions H(1) ν(x)andH(2) ν(x): H(1) ν(x)=Jν(x)+iNν(x) (11.85) and H(2) ν(x)=Jν(x)−iNν(x). (11.86) Thisis exactlyanalogoustotaking e±iθ=cosθ±isinθ. (11.87) Forrealarguments, H(1) νandH(2) νare complexconjugates.The extentof theanalogywill be seen even better when the asymptotic forms are considered (Section 11.6). Indeed, it is theirasymptoticbehaviorthatmakestheHankelfunctionsuseful. Seriesexpansionof H(1) ν(x)andH(2) ν(x)maybeobtainedbycombiningEqs.(11.5)and (11.63).Oftenonlythefirst term isof interest;itis givenby H(1) 0(x)≈i2 πlnx+1+i2 π(γ−ln2)+···, (11.88) H(1) ν(x)≈−i(ν−1)! πparenleftbigg2 xparenrightbiggν +···,ν>0, (11.89) H(2) 0(x)≈−i2 πlnx+1−i2 π(γ−ln2)+···, (11.90) H(2) ν(x)≈i(ν−1)! πparenleftbigg2 xparenrightbiggν +···,ν>0. (11.91) 708 Chapter 11 Bessel Functions Since the Hankel functions are linear combinations (with constant coefficients) of Jν andNν,theysatisfy thesamerecurrencerelations(Eqs. (11.10)and(11.12)) Hν−1(x)+Hν+1(x)=2ν xHν(x), (11.92) Hν−1(x)−Hν+1(x)=2H′ ν(x), (11.93) forbothH(1) ν(x)andH(2) ν(x). AvarietyofWronskianformulascanbedeveloped: H(2) νH(1) ν+1−H(1) νH(2) ν+1=4 iπx, (11.94) Jν−1H(1) ν−JνH(1) ν−1=2 iπx, (11.95) JνH(2) ν−1−Jν−1H(2) ν=2 iπx. (11.96) Example 11.4.1 CYLINDRICAL TRAVELING WAVES AsanillustrationoftheuseofHankelfunctions,consideratwo-dimensionalwaveproblem similartothevibratingcircularmembraneofExercise11.1.25.Nowimaginethatthewaves are generated at r=0 and move outward to infinity. We replace our standing waves by traveling ones. The differential equation remains the same, but the boundary conditions change.Wenowdemandthatfor large rthewavebehavelike U∼ei(kr−ωt)(11.97) to describe an outgoing wave. As before, kis the wave number. This assumes, for sim- plicity, that there is no azimuthaldependence,that is, no angularmomentum,or m=0.In Sections7.3and11.6, H(1) 0(kr)isshowntohavetheasymptoticbehavior(for r→∞) H(1) 0(kr)∼eikr. (11.98) Thisboundaryconditionatinfinitythendeterminesourwavesolutionas U(r,t)=H(1) 0(kr)e−iωt. (11.99) This solution diverges as r→0, which is the behavior to be expected with a source at the origin. Thechoiceofatwo-dimensionalwaveproblemtoillustratetheHankelfunction H(1) 0(z) is not accidental. Bessel functions may appear in a variety of ways, such as in the sepa- ration of conical coordinates. However, they enter most commonly in the radial equations from the separation of variables in the Helmholtz equation in cylindrical and in spheri- cal polar coordinates. We have taken a degenerate form of cylindrical coordinates for this illustration. Had we used spherical polar coordinates (spherical waves), we should have encounteredindex ν=n+1 2,nan integer.These specialvalues yieldthe sphericalBessel functionstobediscussedinSection11.7. /squaresolid 11.4 Hankel Functions 709 Contour Integral Representation of the Hankel Functions Theintegralrepresentation(Schlaefliintegral) Jν(x)=1 2πicontintegraldisplay Ce(x/2)(t−1/t)dt tν+1(11.100) may easily be established as a Cauchy integral for ν=n, an integer (by recognizing that the numerator is the generating function (Eq. (11.1)) and integrating around the origin). Ifνis not an integer, the integrand is not single-valued and a cut line is needed in our complexplane.Choosingthenegativerealaxisasthecutlineandusingthecontourshown in Fig. 11.7, we can extend Eq. (11.100) to nonintegral ν. Substituting Eq. (11.100) into Bessel’s ODE, we can represent the combined integrand by an exact differential that van- ishesast→∞e±iπ(compareExercise11.1.16). We now deform the contour so that it approaches the origin along the positive real axis, as shown in Fig. 11.8. For x>0,this particular approach guarantees that the exact differ- entialmentionedwillvanishas t→0 becauseofthe e−x/2t→0 factor.Henceeachofthe separate portions ( ∞e−iπto 0) and (0 to ∞eiπ) is a solution of Bessel’s equation. We define H(1) ν(x)=1 πiintegraldisplay∞eiπ 0e(x/2)(t−1/t)dt tν+1, (11.101) H(2) ν(x)=1 πiintegraldisplay0 ∞e−iπe(x/2)(t−1/t)dt tν+1. (11.102) Theseexpressionsareparticularlyconvenientbecausetheymaybehandledbythemethod of steepest descents (Section 7.3). H(1) ν(x)has a saddle point at t=+i, whereas H(2) ν(x) hasasaddlepointat t=−i. FIGURE 11.7Besselfunctioncontour. 710 Chapter 11 Bessel Functions FIGURE 11.8Hankelfunctioncontours. TheproblemofrelatingEqs.(11.101)and(11.102)toourearlierdefinitionoftheHankel function (Eqs. (11.85) and (11.86)) remains. Since Eqs. (11.100) to (11.102) combined yield Jν(x)=1 2bracketleftbig H(1) ν(x)+H(2) ν(x)bracketrightbig (11.103) byinspection,weneedonlyshowthat Nν(x)=1 2ibracketleftbig H(1) ν(x)−H(2) ν(x)bracketrightbig . (11.104) Thismaybeaccomplishedbythefollowingsteps: 1. Withthesubstitutions t=eiπ/sforH(1) νandt=e−iπ/sforH(2) ν, weobtain H(1) ν(x)=e−iνπH(1) −ν(x), (11.105) H(2) ν(x)=eiνπH(2) −ν(x). (11.106) 2. FromEqs. (11.103) (ν→−ν), (11.105), and(11.106), J−ν(x)=1 2bracketleftbig eiνπH(1) ν(x)+e−iνπH(2) ν(x)bracketrightbig . (11.107) 3. Finally substitute Jν(Eq. (11.103)) and J−ν(Eq. (11.107)) into the defining equation forNν, Eq. (11.60). This leads to Eq. (11.104) and establishes the contour integrals Eqs. (11.101)and(11.102)astheHankelfunctions. Integral representations have appeared before: Eq. (8.35) for Ŵ(z)and various representa- tionsofJν(z)inSection11.1.WiththeseintegralrepresentationsoftheHankelfunctions, it is perhaps appropriate to ask why we are interested in integral representations. There are at least four reasons. The first is simply aesthetic appeal. Second, the integral repre- sentationshelptodistinguishbetweentwolinearlyindependentsolutions.InFig.11.6,the contoursC1andC2crossdifferent saddlepoints(Section7.3).FortheLegendrefunctions thecontourfor Pn(z)(Fig.12.11)andthatfor Qn(z)encircledifferent singularpoints. 11.4 Hankel Functions 711 Third, the integral representations facilitate manipulations, analysis, and the develop- ment of relations among the various special functions. Fourth, and probably most impor- tant of all, the integral representations are extremely useful in developing asymptotic ex- pansions.Oneapproach,themethodofsteepestdescents,appearsinSection7.3.Asecond approach,thedirectexpansionofanintegralrepresentationisgiveninSection11.6forthe modified Bessel function Kν(z). This same technique may be used to obtain asymptotic expansionsof theconfluenthypergeometricfunctions MandU—Exercise13.5.13. In conclusion,theHankelfunctionsareintroducedhereforthefollowingreasons: •Asanalogsof e±ixtheyare usefulfordescribingtravelingwaves. •Theyofferanalternate(contourintegral)andaratherelegantdefinitionofBesselfunc-tions. •H(1) νis usedtodefinethemodifiedBesselfunction KνofSection11.5. Exercises 11.4.1 VerifytheWronskianformulas (a)Jν(x)H(1)′ ν(x)−J′ ν(x)H(1) ν(x)=2i πx, (b)Jν(x)H(2)′ ν(x)−J′ ν(x)H(2) ν(x)=−2i πx, (c)Nν(x)H(1)′ ν(x)−N′ ν(x)H(1) ν(x)=−2 πx, (d)Nν(x)H(2)′ ν(x)−N′ ν(x)H(2) ν(x)=−2 πx, (e)H(1) ν(x)H(2)′ ν(x)−H(1)′ ν(x)H(2) ν(x)=−4i πx, (f)H(2) ν(x)H(1) ν+1(x)−H(1) ν(x)H(2) ν+1(x)=4 iπx, (g)Jν−1(x)H(1) ν(x)−Jν(x)H(1) ν−1(x)=2 iπx. 11.4.2 Showthattheintegralforms (a)1 iπintegraldisplay∞eiπ 0C1e(x/2)(t−1/t)dt tν+1=H(1) ν(x), (b)1 iπintegraldisplay0 ∞e−iπC2e(x/2)(t−1/t)dt tν+1=H(2) ν(x) satisfyBessel’s ODE.Thecontours C1andC2areshowninFig.11.8. 11.4.3 Usingtheintegralsandcontoursgiveninproblem11.4.2,showthat 1 2ibracketleftbig H(1) ν(x)−H(2) ν(x)bracketrightbig =Nν(x). 11.4.4 ShowthattheintegralsinExercise11.4.2maybetransformedtoyield (a)H(1) ν(x)=1 πiintegraldisplay C3exsinhγ−νγdγ,(b)H(2) ν(x)=1 πiintegraldisplay C4exsinhγ−νγdγ 712 Chapter 11 Bessel Functions FIGURE 11.9Hankelfunctioncontours. (seeFig.11.9). 11.4.5 (a) Transform H(1) 0(x), Eq. (11.101),into H(1) 0(x)=1 iπintegraldisplay Ceixcoshsds, where the contour Cruns from−∞−iπ/2 through the origin of the s-plane to ∞+iπ/2. (b) Justifyrewriting H(1) 0(x)as H(1) 0(x)=2 iπintegraldisplay∞+iπ/2 0eixcoshsds. (c) Verify that this integral representation actually satisfies Bessel’s differential equa- tion.(The iπ/2intheupperlimitisnotessential.Itservesasaconvergencefactor. Wecanreplaceitby iaπ/2 andtakethelimit.) 11.4.6 From H(1) 0(x)=2 iπintegraldisplay∞ 0eixcoshsds showthat (a)J0(x)=2 πintegraldisplay∞ 0sin(xcoshs)ds, (b)J0(x)=2 πintegraldisplay∞ 1sin(xt)√ t2−1dt. Thislastresult isaFouriersinetransform. 11.4.7 From (seeExercises11.4.4and11.4.5) H(1) 0(x)=2 iπintegraldisplay∞ 0eixcoshsds showthat (a)N0(x)=−2 πintegraldisplay∞ 0cos(xcoshs)ds. 11.5 Modified Bessel Functions, Iν(x)andKν(x) 713 (b)N0(x)=−2 πintegraldisplay∞ 1cos(xt)radicalbig t2−1)dt. ThesearetheintegralrepresentationsinSection11.3(OtherForms). Thislastresult isaFouriercosinetransform. 11.5 M ODIFIED BESSEL FUNCTIONS ,Iν(x) ANDKν(x) TheHelmholtzequation, ∇2ψ+k2ψ=0, separated in circular cylindrical coordinates, leads to Eq. (11.22a), the Bessel equation. Equation (11.22a) is satisfied by the Bessel and Neumann functions Jν(kρ)andNν(kρ) and any linear combination, such as the Hankel functions H(1) ν(kρ)andH(2) ν(kρ).N o w , the Helmholtz equation describes the space part of wave phenomena. If instead we have a diffusionproblem,thentheHelmholtzequationis replacedby ∇2ψ−k2ψ=0. (11.108) TheanalogtoEq.(11.22a)is ρ2d2 dρ2Yν(kρ)+ρd dρYν(kρ)−parenleftbig k2ρ2+ν2parenrightbig Yν(kρ)=0.(11.109) The Helmholtz equation may be transformed into the diffusion equation by the trans- formation k→ik. Similarly, k→ikchanges Eq. (11.22a) into Eq. (11.109) and shows that Yν(kρ)=Zν(ikρ). The solutions of Eq. (11.109) are Bessel functions of imaginary argument. To obtain a solution that is regular at the origin, we take Zνas the regular Bessel function Jν.I ti s customary(andconvenient)tochoosethenormalizationsothat Yν(x)=Iν(x)≡i−νJν(ix). (11.110) (Here the variable kρis being replaced by xfor simplicity.) The extra i−νnormalization cancelsthe iνfrom eachtermandleaves Iν(x)real.Oftenthisis writtenas Iν(x)=e−νπi/2Jνparenleftbig xeiπ/2parenrightbig . (11.111) I0andI1areshowninFig.11.10. 714 Chapter 11 Bessel Functions FIGURE 11.10ModifiedBessel functions. Series Form In terms of infinite series this is equivalent to removing the (−1)ssign in Eq. (11.5) and writing Iν(x)=∞summationdisplay s=01 s!(s+ν)!parenleftbiggx 2parenrightbigg2s+ν ,I−ν(x)=∞summationdisplay s=01 s!(s−ν)!parenleftbiggx 2parenrightbigg2s−ν .(11.112) For integral νthis yields In(x)=I−n(x). (11.113) Recurrence Relations The recurrence relations satisfied by Iν(x)may be developed from the series expansions, but it is perhaps easier to work from the existing recurrence relations for Jν(x). Let us replacexby−ixandrewriteEq.(11.110)as Jν(x)=iνIν(−ix). (11.114) ThenEq. (11.10)becomes iν−1Iν−1(−ix)+iν+1Iν+1(−ix)=2ν xiνIν(−ix). Replacing xbyix, wehavearecurrencerelationfor Iν(x), Iν−1(x)−Iν+1(x)=2ν xIν(x). (11.115) 11.5 Modified Bessel Functions, Iν(x)andKν(x) 715 Equation(11.12)transforms to Iν−1(x)+Iν+1(x)=2I′ ν(x). (11.116) ThesearetherecurrencerelationsusedinExercise11.1.14.Itisworthemphasizingthatal- thoughtworecurrencerelations,Eqs.(11.115)and(11.116)orExercise11.5.7,specifythe second-orderODE,theconverseisnottrue.TheODEdoesnotuniquelyfixtherecurrence relations.Equations(11.115)and(11.116)andExercise11.5.7provideanexample. From Eq. (11.113) it is seen that we have but one independent solution when νis an integer,exactlyasintheBesselfunctions Jν.Thechoiceofasecond,independentsolution of Eq. (11.108) is essentially a matter of convenience. The second solution given here is selected on the basis of its asymptotic behavior—as shown in the next section. The confusion of choice and notation for this solution is perhaps greater than anywhere else in this field.18Many authors19choose to define a second solution in terms of the Hankel functionH(1) ν(x)by Kν(x)≡π 2iν+1H(1) ν(ix)=π 2iν+1bracketleftbig Jν(ix)+iNν(ix)bracketrightbig . (11.117) Thefactor iν+1makesKν(x)realwhen xisreal.UsingEqs.(11.60)and(11.110),wemay transform Eq. (11.117)to20 Kν(x)=π 2I−ν(x)−Iν(x) sinνπ, (11.118) analogoustoEq.(11.60)for Nν(x).ThechoiceofEq.(11.117)asadefinitionissomewhat unfortunate in that the function Kν(x)does not satisfy the same recurrence relations as Iν(x)(compare Exercises 11.5.7 and 11.5.8). To avoid this annoyance, other authors21 haveincludedanadditionalfactorofcos νπ.Thispermits Kνtosatisfythesamerecurrence relationsas Iν, butithasthedisadvantageof making Kν=0f o rν=1 2,3 5,5 2,.... The series expansion of Kν(x)follows directly from the series form of H(1) ν(ix).T h e lowest-orderterms are(cf. Eqs. (11.61)and(11.62)) K0(x)=−lnx−γ+ln2+···, Kν(x)=2ν−1(ν−1)!x−ν+···. (11.119) BecausethemodifiedBessel function Iνis relatedtotheBessel function Jν, muchassinh is related to sine, Iνand the second solution Kνare sometimes referred to as hyperbolic Besselfunctions. K0andK1are showninFig.11.10. I0(x)andK0(x)havetheintegralrepresentations I0(x)=1 πintegraldisplayπ 0cosh(xcosθ)dθ, (11.120) K0(x)=integraldisplay∞ 0cos(xsinht)dt=integraldisplay∞ 0cos(xt)dt (t2+1)1/2,x>0. (11.121) 18Adiscussion and comparison of notations will befound in Math. Tables Aids Comput. 1: 207–308 (1944). 19Watson, Morse and Feshbach, Jeffreys and Jeffreys (without the π/2). 20For integral index nwetakethe limit as ν→n. 21Whittaker and Watson, seeAdditional Readings of Chapter 13. 716 Chapter 11 Bessel Functions Equation(11.120)maybederivedfromEq.(11.30)for J0(x)ormaybetakenasaspecial caseofExercise11.5.4, ν=0.Theintegralrepresentationof K0,Eq.(11.121),isaFourier transform and may best be derived with Fourier transforms, Chapter 15, or with Green’s functionsSection9.7.Avarietyofotherformsofintegralrepresentations(including ν/negationslash=0) appearintheexercises.Theseintegralrepresentationsareusefulindevelopingasymptotic forms(Section11.6) andinconnectionwithFouriertransforms, Chapter15. To put the modified Bessel functions Iν(x)andKν(x)in proper perspective, we intro- ducethemherebecause: •Thesefunctionsare solutionsofthefrequentlyencounteredmodifiedBessel equation. •Theyare neededfor specificphysicalproblems,suchas diffusionproblems. •Kν(x)providesa Green’sfunction,Section9.7. •Kν(x)leadstoaconvenientdeterminationof asymptoticbehavior(Section11.6). Exercises 11.5.1 Showthat e(x/2)(t+1/t)=∞summationdisplay n=−∞In(x)tn, thusgeneratingmodifiedBesselfunctions, In(x). 11.5.2 Verifythefollowingidentities (a) 1=I0(x)+2∞summationdisplay n=1(−1)nI2n(x), (b)ex=I0(x)+2∞summationdisplay n=1In(x), (c)e−x=I0(x)+2∞summationdisplay n=1(−1)nIn(x), (d) cosh x=I0(x)+2∞summationdisplay n=1I2n(x), (e) sinh x=2∞summationdisplay n=1I2n−1(x). 11.5.3 (a) FromthegeneratingfunctionofExercise11.5.1showthat In(x)=1 2πicontintegraldisplay expbracketleftbig (x/2)(t+1/t)bracketrightbigdt tn+1. 11.5 Modified Bessel Functions, Iν(x)andKν(x) 717 (b) For n=ν, not an integer, show that the preceding integral representation may be generalizedto Iν(x)=1 2πiintegraldisplay Cexpbracketleftbig (x/2)(t+1/t)bracketrightbigdt tν+1. Thecontour Cisthesameasthatfor Jν(x), Fig.11.7. 11.5.4 Forν>−1 2showthat Iν(z)mayberepresentedby Iν(z)=1 π1/2(ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplayπ 0e±zcosθsin2νθdθ =1 π1/2(ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplay1 −1e±zpparenleftbig 1−p2parenrightbigν−1/2dp =2 π1/2(ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplayπ/2 0cosh(zcosθ)sin2νθdθ. 11.5.5 A cylindrical cavity has a radius aand height l, Fig. 11.3. The ends, z=0 andl,a r ea t zeropotential.Thecylindricalwalls, ρ=a, haveapotential V=V(ϕ,z). (a) Showthattheelectrostaticpotential /Phi1(ρ,ϕ,z) has thefunctionalform /Phi1(ρ,ϕ,z)=∞summationdisplay m=0∞summationdisplay n=1Im(knρ)sinknz·(amnsinmϕ+bmncosmϕ), wherekn=nπ/l. (b) Showthatthecoefficients amnandbmnaregivenby22 amn bmnbracerightbigg =2 πlIm(kna)integraldisplay2π 0integraldisplayl 0V(ϕ,z)sinknz·braceleftbiggsinmϕ cosmϕbracerightbigg dzdϕ. Hint. Expand V(ϕ,z)as a double series and use the orthogonality of the trigonometric functions. 11.5.6 Verifythat Kν(x)is givenby Kν(x)=π 2I−ν(x)−Iν(x) sinνπ andfrom thisshowthat Kν(x)=K−ν(x). 11.5.7 Showthat Kν(x)satisfiestherecurrencerelations Kν−1(x)−Kν+1(x)=−2ν xKν(x), Kν−1(x)+Kν+1(x)=−2K′ ν(x). 22Whenm=0, the2in thecoefficient is replacedby 1. 718 Chapter 11 Bessel Functions 11.5.8 IfKν=eνπiKν, showthat Kνsatisfiesthesamerecurrencerelationsas Iν. 11.5.9 Forν>−1 2showthat Kν(z)mayberepresentedby Kν(z)=π1/2 (ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplay∞ 0e−zcoshtsinh2νtdt,−π 2<argz<π 2 =π1/2 (ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplay∞ 1e−zp(p2−1)ν−1/2dp. 11.5.10 Showthat Iν(x)andKν(x)satisfy theWronskianrelation Iν(x)K′ ν(x)−I′ ν(x)Kν(x)=−1 x. Thisresult isquotedinSection9.7inthedevelopmentof aGreen’sfunction. 11.5.11 Ifr=(x2+y2)1/2, prove that 1 r=2 πintegraldisplay∞ 0cos(xt)K0(yt)dt. Thisis aFouriercosinetransformof K0. 11.5.12 (a) Verifythat I0(x)=1 πintegraldisplayπ 0cosh(xcosθ)dθ satisfiesthemodifiedBessel equation, ν=0. (b) Showthatthisintegralcontainsnoadmixtureof K0(x),theirregularsecondsolu- tion. (c) Verifythenormalizationfactor 1 /π. 11.5.13 Verifythattheintegralrepresentations In(z)=1 πintegraldisplayπ 0ezcostcos(nt)dt, Kν(z)=integraldisplay∞ 0e−zcoshtcosh(νt)dt,ℜ(z)>0, satisfy the modified Bessel equation by direct substitution into that equation. How can you show that the first form does not contain an admixture of Knand that the second formdoesnotcontainanadmixtureof Iν?Howcanyoucheckthenormalization? 11.5.14 Derivetheintegralrepresentation In(x)=1 πintegraldisplayπ 0excosθcos(nθ)dθ. Hint. Start with the corresponding integral representation of Jn(x). Equation (11.120) isa specialcaseofthisrepresentation. 11.6 Asymptotic Expansions 719 11.5.15 Showthat K0(z)=integraldisplay∞ 0e−zcoshtdt satisfies the modified Bessel equation. How can you establish that this form is linearly independentof I0(z)? 11.5.16 Showthat eax=I0(a)T0(x)+2∞summationdisplay n=1In(a)Tn(x),−1≤x≤1. Tn(x)isthenth-orderChebyshevpolynomial,Section13.3. Hint.AssumeaChebyshevseriesexpansion.Usingtheorthogonalityandnormalization oftheTn(x), solvefor thecoefficientsof theChebyshevseries. 11.5.17 (a) Write a double precision subroutine to calculate In(x)to 12-decimal-place accu- racy forn=0,1,2,3,...and 0≤x≤1. Check your results against the 10-place valuesgiveninAMS-55,Table9.11,seeAdditionalReadingsofChapter8forthe reference. (b) Referring to Exercise 11.5.16, calculate the coefficients in the Chebyshev expan- sionsof cosh xandof sinh x. 11.5.18 Thecylindricalcavityof Exercise11.5.5hasa potentialalongthecylinderwalls: V(z)=braceleftBigg 100z l, 0≤z l≤1 2, 100parenleftbig 1−z lparenrightbig ,1 2≤z l≤1. Withtheradius–heightratio a/l=0.5,calculatethepotentialfor z/l=0.1(0.1)0.5and ρ/a=0.0(0.2)1.0. Checkvalue. Forz/l=0.3 andρ/a=0.8,V=26.396. 11.6 A SYMPTOTIC EXPANSIONS Frequently in physical problems there is a need to know how a given Bessel or modified Bessel function behaves for large values of the argument, that is, the asymptotic behavior. This is one occasion when computers are not very helpful. One possible approach is to developapower-seriessolutionofthedifferentialequation,asinSection9.5,butnowusing negative powers. This is Stokes’ method, Exercise 11.6.5. The limitation is that starting from some positive value of the argument (for convergence of the series), we do not know whatmixtureofsolutionsormultipleofagivensolutionwehave.Theproblemistorelate theasymptoticseries(usefulforlargevaluesofthevariable)tothepower-seriesorrelated definition (useful for small values of the variable). This relationship can be established by introducingasuitable integralrepresentation andthenusingeitherthemethodofsteepest descent,Section7.3, orthedirectexpansionasdevelopedinthissection. 720 Chapter 11 Bessel Functions Expansion of an Integral Representation Asadirectapproach,considertheintegralrepresentation(Exercise11.5.9) Kν(z)=π1/2 (ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplay∞ 1e−zxparenleftbig x2−1parenrightbigν−1/2dx, ν>−1 2.(11.122) For the present let us take zto be real, although Eq. (11.122) may be established for −π/2<argz<π/2(ℜ(z)>0). Wehavethreetasks: 1. To show that Kνas given in Eq. (11.122) actually satisfies the modified Bessel equa- tion(11.109). 2. Toshowthattheregularsolution Iνisabsent. 3. ToshowthatEq.(11.122) hasthepropernormalization. 1.ThefactthatEq.(11.122)isasolutionofthemodifiedBesselequationmaybeverified bydirectsubstitution.Weobtain zν+1integraldisplay∞ 1d dxbracketleftbig e−zxparenleftbig x2−1parenrightbigν+1/2bracketrightbig dx=0, which transforms the combined integrand into the derivative of a function that vanishes at bothendpoints.Hencetheintegralissomelinearcombinationof IνandKν. 2. The rejection of the possibility that this solution contains Iνconstitutes Exer- cise11.6.1. 3. The normalization may be verified by showing that, in the limit z→0,Kν(z)is in agreementwithEq.(11.119). Bysubstituting x=1+t/z, π1/2 (ν−1 2)!parenleftbiggz 2parenrightbiggνintegraldisplay∞ 1e−zxparenleftbig x2−1parenrightbigν−1/2dx =π1/2 (ν−1 2)!parenleftbiggz 2parenrightbiggν e−zintegraldisplay∞ 0e−tparenleftbiggt2 z2+2t zparenrightbiggν−1/2dt z(11.123a) =π1/2 (ν−1 2)!e−z 2νzνintegraldisplay∞ 0e−tt2ν−1parenleftbigg 1+2z tparenrightbiggν−1/2 dt, (11.123b) takingout t2/z2asafactor.Thissubstitutionhaschangedthelimitsofintegrationtoamore convenient range and has isolated the negative exponential dependence e−z. The integral inEq.(11.123b)maybeevaluatedfor z=0toyield (2ν−1)!.Then,usingtheduplication formula(Section8.4), wehave lim z→0Kν(z)=(ν−1)!2ν−1 zν,ν>0, (11.124) inagreementwithEq.(11.119), whichthuschecksthenormalization.23 23Forν→0 the integral diverges logarithmically, in agreement with the logarithmic divergence of K0(z)forz→0 (Sec- tion 11.5). 11.6 Asymptotic Expansions 721 Now,todevelopanasymptoticseries for Kν(z), wemayrewriteEq. (11.123a)as Kν(z)=radicalbiggπ 2ze−z (ν−1 2)!integraldisplay∞ 0e−ttν−1/2parenleftbigg 1+t 2zparenrightbiggν−1/2 dt (11.125) (takingout 2 t/zasafactor). We expand (1+t/2z)ν−1/2bythebinomialtheoremtoobtain Kν(z)=radicalbiggπ 2ze−z (ν−1 2)!∞summationdisplay r=0(ν−1 2)! r!(ν−r−1 2)!(2z)−rintegraldisplay∞ 0e−ttν+r−1/2dt.(11.126) Term-by-term integration (valid for asymptotic series) yields the desired asymptotic ex- pansionof Kν(z): Kν(z)∼radicalbiggπ 2ze−zbracketleftbigg 1+(4ν2−12) 1!8z+(4ν2−12)(4ν2−32) 2!(8z)2+···bracketrightbigg .(11.127) Althoughthe integralof Eq. (11.122), integratingalong the real axis, was convergentonly for−π/2<argz<π/2, Eq. (11.127) may be extended to −3π/2<argz<3π/2. Con- sidered as an infinite series, Eq. (11.127) is actually divergent.24However, this series is asymptotic, in the sense that for large enough z,Kν(z)may be approximated to any fixed degree of accuracy with a small number of terms. (Compare Section 5.10 for a definition anddiscussionofasymptoticseries.) It isconvenienttorewriteEq.(11.127)as Kν(z)=radicalbiggπ 2ze−zbracketleftbig Pν(iz)+iQν(iz)bracketrightbig , (11.128) where Pν(z)∼1−(µ−1)(µ−9) 2!(8z)2+(µ−1)(µ−9)(µ−25)(µ−49) 4!(8z)4−···,(11.129a) Qν(z)∼µ−1 1!(8z)−(µ−1)(µ−9)(µ−25) 3!(8z)3+···, (11.129b) and µ=4ν2. Itshouldbenotedthatalthough Pν(z)ofEq.(11.129a)and Qν(z)ofEq.(11.129b)have alternating signs, the series for Pν(iz)andQν(iz)of Eq. (11.128) have all signs positive. Finally,for zlarge,Pνdominates. Thenwiththeasymptoticformof Kν(z),Eq.(11.128),wecanobtainexpansionsforall otherBessel andhyperbolicBessel functionsbydefiningrelations: 24Our binomial expansion is valid only for t<2zand we have integrated tout to infinity. The exponential decrease of the integrandpreventsadisaster,buttheresultantseriesisstillonlyasymptotic,notconvergent.ByTable9.3, z=∞isanessential singularity of theBessel(and modified Bessel)equations. Fuchs’theoremdoes not guaranteeaconvergent series andwedo not get aconvergent series. 722 Chapter 11 Bessel Functions 1. From π 2iν+1H(1) ν(iz)=Kν(z) (11.130) wehave H(1) ν(z)=radicalbigg 2 πzexpbraceleftbigg ibracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbiggbracerightbigg ·bracketleftbig Pν(z)+iQν(z)bracketrightbig ,−π<argz<2π.(11.131) 2. The second Hankel function is just the complex conjugate of the first (for real argu- ment), H(2) ν(z)=radicalbigg 2 πzexpbraceleftbigg −ibracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbiggbracerightbigg ·bracketleftbig Pν(z)−iQν(z)bracketrightbig ,−2π<argz<π. (11.132) An alternate derivation of the asymptotic behavior of the Hankel functions appears in Section7.3as anapplicationofthemethodofsteepestdescents. 3. Since Jν(z)istherealpartof H(1) ν(z)for realz, Jν(z)=radicalbigg 2 πzbraceleftbigg Pν(z)cosbracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbigg −Qν(z)sinbracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbiggbracerightbigg ,−π<argz<π,(11.133) holds for real z,that is, arg z=0,π. Once Eq. (11.133) is established for real z,the relationisvalidforcomplex zinthegivenrangeof argument. 4. TheNeumannfunctionistheimaginarypartof H(1) ν(z)for realz,or Nν(z)=radicalbigg 2 πzbraceleftbigg Pν(z)sinbracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbigg +Qν(z)cosbracketleftbigg z−parenleftbigg ν+1 2parenrightbiggπ 2bracketrightbiggbracerightbigg ,−π<argz<π.(11.134) Initially, this relation is established for real z,but it may be extended to the complex domainas shown. 5. Finally,theregularhyperbolicormodifiedBesselfunction Iν(z)isgivenby Iν(z)=i−νJν(iz) (11.135) or Iν(z)=ez √ 2πzbracketleftbig Pν(iz)−iQν(iz)bracketrightbig ,−π 2<argz<π 2. (11.136) 11.6 Asymptotic Expansions 723 FIGURE 11.11Asymptoticapproximationof J0(x). This completes our determination of the asymptotic expansions. However, it is perhaps worth noting the primary characteristics. Apart from the ubiquitous z−1/2,JνandNνbe- have as cosine and sine, respectively. The zeros are almostevenly spaced at intervals of π;thespacingbecomesexactly πinthelimitas z→∞. TheHankelfunctionshavebeen defined to behave like the imaginary exponentials, and the modified Bessel functions Iν andKνgo into the positive and negative exponentials. This asymptotic behavior may be sufficienttoeliminateimmediatelyoneofthesefunctionsasasolutionforaphysicalprob- lem. We should also note that the asymptotic series Pν(z)andQν(z), Eqs. (11.129a) and (11.129b),terminatefor ν=±1/2,±3/2,...andbecomepolynomials(innegativepowers ofz).Forthesespecialvaluesof νtheasymptoticapproximationsbecomeexactsolutions. It is of some interest to consider the accuracy of the asymptotic forms, taking just the firstterm, forexample(Fig. 11.11), Jn(x)≈radicalbigg 2 πxcosbracketleftbigg x−parenleftbigg n+1 2parenrightbiggparenleftbiggπ 2parenrightbiggbracketrightbigg . (11.137) Clearly, the condition for the validity of Eq. (11.137) is that the sine term be negligible; thatis, 8x≫4n2−1. (11.138) Fornorν>1 theasymptoticregionmaybefar out. AspointedoutinSection11.3,theasymptoticformsmaybeusedtoevaluatethevarious Wronskianformulas(compareExercise11.6.3). Exercises 11.6.1 Incheckingthenormalizationoftheintegralrepresentationof Kν(z)(Eq.(11.122)),we assumed that Iν(z)was not present. How do we know that the integral representation (Eq. (11.122))doesnotyield Kν(z)+εIν(z)withε/negationslash=0? 724 Chapter 11 Bessel Functions FIGURE 11.12ModifiedBessel functioncontours. 11.6.2 (a) Showthat y(z)=zνintegraldisplay e−ztparenleftbig t2−1parenrightbigν−1/2dt satisfiesthemodifiedBessel equation,providedthecontouris chosenso that e−ztparenleftbig t2−1parenrightbigν+1/2 hasthesamevalueattheinitialandfinalpointsof thecontour. (b) VerifythatthecontoursshowninFig.11.12are suitablefor thisproblem. 11.6.3 UsetheasymptoticexpansionstoverifythefollowingWronskianformulas: (a)Jν(x)J−ν−1(x)+J−ν(x)Jν+1(x)=−2sinνπ/πx, (b)Jν(x)Nν+1(x)−Jν+1(x)Nν(x)=−2/πx, (c)Jν(x)H(2) ν−1(x)−Jν−1(x)H(2) ν(x)=2/iπx, (d)Iν(x)K′ ν(x)−I′ ν(x)Kν(x)=−1/x, (e)Iν(x)Kν+1(x)+Iν+1(x)Kν(x)=1/x. 11.6.4 From the asymptotic form of Kν(z), Eq. (11.127), derive the asymptotic form of H(1) ν(z), Eq.(11.131). Noteparticularlythephase, (ν+1 2)π/2. 11.6.5 Stokes’method. (a) ReplacetheBesselfunctioninBessel’sequationby x−1/2y(x)andshowthat y(x) satisfies y′′(x)+parenleftbigg 1−ν2−1 4 x2parenrightbigg y(x)=0. (b) Develop a power-series solution with negative powers of xstarting with the as- sumedform y(x)=eix∞summationdisplay n=0anx−n. Determine the recurrence relation giving an+1in terms of an. Check your result againsttheasymptoticseries, Eq.(11.131). (c) Fromtheresults ofSection7.4determinetheinitialcoefficient, a0. 11.7 Spherical Bessel Functions 725 11.6.6 Calculate the first 15 partial sums of P0(x)andQ0(x), Eqs. (11.129a) and (11.129b). Letxvary from 4 to 10 in unit steps. Determine the number of terms to be retained for maximum accuracy and the accuracy achieved as a function of x. Specifically, how smallmay xbewithoutraisingtheerrorabove 3 ×10−6? ANS.xmin=6. 11.6.7 (a) Using the asymptotic series (partial sums) P0(x)andQ0(x)determined in Exer- cise 11.6.6, write a function subprogram FCT(X) that will calculate J0(x),xreal, forx≥xmin. (b) Test your function by comparing it with the J0(x)(tables or computer library subroutine)for x=xmin(10)xmin+10. Note.Amoreaccurateandperhapssimplerasymptoticformfor J0(x)isgiveninAMS- 55,Eq. (9.4.3), seeAdditionalReadingsofChapter8for thereference. 11.7 S PHERICAL BESSEL FUNCTIONS WhentheHelmholtzequationisseparatedinsphericalcoordinates,theradialequationhas theform r2d2R dr2+2rdR dr+bracketleftbig k2r2−n(n+1)bracketrightbig R=0. (11.139) This is Eq. (9.65) of Section 9.3. The parameter kenters from the original Helmholtz equation, while n(n+1)is a separation constant. From the behavior of the polar angle function (Legendre’s equation, Sections 9.5 and 12.5), the separation constant must have this form, with na nonnegative integer. Equation (11.139) has the virtue of being self- adjoint,butclearlyitis notBessel’s equation.However,if wesubstitute R(kr)=Z(kr) (kr)1/2, Equation(11.139)becomes r2d2Z dr2+rdZ dr+bracketleftbigg k2r2−parenleftbigg n+1 2parenrightbigg2bracketrightbigg Z=0, (11.140) whichisBessel’s equation. Zis a Bessel function of order n+1 2(nan integer). Because oftheimportanceof sphericalcoordinates,thiscombination,thatis, Zn+1/2(kr) (kr)1/2, occursquiteoften. 726 Chapter 11 Bessel Functions Definitions ItisconvenienttolabelthesefunctionssphericalBesselfunctionswiththefollowingdefin- ingequations: jn(x)=radicalbiggπ 2xJn+1/2(x), nn(x)=radicalbiggπ 2xNn+1/2(x)=(−1)n+1radicalbiggπ 2xJ−n−1/2(x),25 h(1) n(x)=radicalbiggπ 2xH(1) n+1/2(x)=jn(x)+inn(x), h(2) n(x)=radicalbiggπ 2xH(2) n+1/2(x)=jn(x)−inn(x).(11.141) ThesesphericalBesselfunctions(Figs.11.13and11.14)canbeexpressedinseriesform byusingtheseries (Eq. (11.5))for Jn, replacing nwithn+1 2: Jn+1/2(x)=∞summationdisplay s=0(−1)s s!(s+n+1 2)!parenleftbiggx 2parenrightbigg2s+n+1/2 . (11.142) UsingtheLegendreduplicationformula, z!(z+1 2)!=2−2z−1π1/2(2z+1)!, (11.143) wehave jn(x)=radicalbiggπ 2x∞summationdisplay s=0(−1)s22s+2n+1(s+n)! π1/2(2s+2n+1)!s!parenleftbiggx 2parenrightbigg2s+n+1/2 =2nxn∞summationdisplay s=0(−1)s(s+n)! s!(2s+2n+1)!x2s. (11.144) Now,Nn+1/2(x)=(−1)n+1J−n−1/2(x)andfromEq. (11.5)wefindthat J−n−1/2(x)=∞summationdisplay s=0(−1)s s!(s−n−1 2)!parenleftbiggx 2parenrightbigg2s−n−1/2 . (11.145) This yields nn(x)=(−1)n+12nπ1/2 xn+1∞summationdisplay s=0(−1)s s!(s−n−1 2)!parenleftbiggx 2parenrightbigg2s . (11.146) 25This is possible because cos (n+1 2)π=0,seeEq. (11.60). 11.7 Spherical Bessel Functions 727 FIGURE 11.13SphericalBesselfunctions. FIGURE 11.14SphericalNeumannfunctions. 728 Chapter 11 Bessel Functions TheLegendreduplicationformulacanbeusedagaintogive nn(x)=(−1)n+1 2nxn+1∞summationdisplay s=0(−1)s(s−n)! s!(2s−2n)!x2s. (11.147) Theseseriesforms,Eqs.(11.144)and(11.147),areusefulinthreeways:(1)limitingvalues asx→0, (2) closed-form representations for n=0, and, as an extension of this, (3) an indicationthatthesphericalBesselfunctionsarecloselyrelatedtosineandcosine. Forthespecialcase n=0 wefindfrom Eq.(11.144) that j0(x)=∞summationdisplay s=0(−1)s (2s+1)!x2s=sinx x, (11.148) whereasfor n0,Eq. (11.147)yields n0(x)=−cosx x. (11.149) FromthedefinitionofthesphericalHankelfunctions(Eq. (11.141)), h(1) 0(x)=1 x(sinx−icosx)=−i xeix, h(2) 0(x)=1 x(sinx+icosx)=i xe−ix. (11.150) Equations (11.148) and (11.149) suggest expressing all spherical Bessel functions as combinationsofsineandcosine.Theappropriatecombinationscanbedevelopedfromthe power-seriessolutions,Eqs.(11.144)and(11.147),butthisapproachisawkward.Actually thetrigonometricformsarealreadyavailableastheasymptoticexpansionofSection11.6. FromEqs. (11.131)and(11.129a), h(1) n(x)=radicalbiggπ 2zH(1) n+1/2(z) =(−i)n+1eiz zbraceleftbig Pn+1/2(z)+iQn+1/2(z)bracerightbig . (11.151) Now,Pn+1/2andQn+1/2arepolynomials .ThismeansthatEq.(11.151)ismathematically exact,notsimplyanasymptoticapproximation.Weobtain h(1) n(z)=(−i)n+1eiz znsummationdisplay s=0is s!(8z)s(2n+2s)!! (2n−2s)!! =(−i)n+1eiz znsummationdisplay s=0is s!(2z)s(n+s)! (n−s)!. (11.152) Often a factor (−i)n=(e−iπ/2)nwill be combined with the eizto giveei(z−nπ/2).F o r zreal,jn(z)is the real part of this, nn(z)the imaginary part, and h(2) n(z)the complex conjugate.Specifically, h(1) 1(x)=eixparenleftbigg −1 x−i x2parenrightbigg , (11.153a) 11.7 Spherical Bessel Functions 729 h(1) 2(x)=eixparenleftbiggi x−3 x2−3i x3parenrightbigg , (11.153b) j1(x)=sinx x2−cosx x, (11.154) j2(x)=parenleftbigg3 x3−1 xparenrightbigg sinx−3 x2cosx, n1(x)=−cosx x2−sinx x, (11.155) n2(x)=−parenleftbigg3 x3−1 xparenrightbigg cosx−3 x2sinx, andso on. Limiting Values Forx≪1,26Eqs. (11.144)and(11.147)yield jn(x)≈2nn! (2n+1)!xn=xn (2n+1)!!, (11.156) nn(x)≈(−1)n+1 2n·(−n)! (−2n)!x−n−1 =−(2n)! 2nn!x−n−1=−(2n−1)!!x−n−1. (11.157) The transformation of factorials in the expressions for nn(x)employs Exercise 8.1.3. The limitingvaluesof thesphericalHankelfunctionsgoas ±inn(x). Theasymptoticvaluesof jn,nn,h(2) n,andh(1) nmaybeobtainedfromtheBesselasymp- toticforms, Section11.6.We find jn(x)∼1 xsinparenleftbigg x−nπ 2parenrightbigg , (11.158) nn(x)∼−1 xcosparenleftbigg x−nπ 2parenrightbigg , (11.159) h(1) n(x)∼(−i)n+1eix x=−iei(x−nπ/2) x, (11.160a) h(2) n(x)∼in+1e−ix x=ie−i(x−nπ/2) x. (11.160b) 26The condition that the second term in the series be negligible compared to the first is actually x≪2[(2n+2)(2n+3)/ (n+1)]1/2forjn(x). 730 Chapter 11 Bessel Functions The condition for these spherical Bessel forms is that x≫n(n+1)/2. From these as- ymptotic values we see that jn(x)andnn(x)are appropriate for a description of standing sphericalwaves ;h(1) n(x)andh(2) n(x)correspondto travelingsphericalwaves .Ifthetime dependence for the traveling waves is taken to be e−iωt, thenh(1) n(x)yields an outgoing travelingsphericalwave, h(2) n(x)anincomingwave.Radiationtheoryinelectromagnetism andscatteringtheoryinquantummechanicsprovidemanyapplications. Recurrence Relations Therecurrencerelationstowhichwenowturnprovideaconvenientwayofdevelopingthe higher-order spherical Bessel functions. These recurrence relations may be derived from theseries,but,aswiththemodifiedBesselfunctions,itiseasiertosubstituteintotheknown recurrencerelations(Eqs. (11.10)and(11.12)). Thisgives fn−1(x)+fn+1(x)=2n+1 xfn(x), (11.161) nfn−1(x)−(n+1)fn+1(x)=(2n+1)f′ n(x). (11.162) Rearrangingtheserelations(orsubstitutingintoEqs. (11.15)and(11.17)), weobtain d dxbracketleftbig xn+1fn(x)bracketrightbig =xn+1fn−1(x), (11.163) d dxbracketleftbig x−nfn(x)bracketrightbig =−x−nfn+1(x). (11.164) Herefnmayrepresent jn,nn,h(1) n,orh(2) n. The specific forms, Eqs. (11.154) and (11.155), may also be readily obtained from Eq.(11.164). BymathematicalinductionwemayestablishtheRayleighformulas jn(x)=(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbiggsinx xparenrightbigg , (11.165) nn(x)=−(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbiggcosx xparenrightbigg , (11.166) h(1) n(x)=−i(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbiggeix xparenrightbigg , (11.167) h(2) n(x)=i(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbigge−ix xparenrightbigg . 11.7 Spherical Bessel Functions 731 Orthogonality WemaytaketheorthogonalityintegralfortheordinaryBesselfunctions(Eqs.(11.49)and (11.50)), integraldisplaya 0Jνparenleftbigg ανpρ aparenrightbigg Jνparenleftbigg ανqρ aparenrightbigg ρdρ=a2 2bracketleftbig Jν+1(ανp)bracketrightbig2δpq, (11.168) andsubstituteintheexpressionfor jntoobtain integraldisplaya 0jnparenleftbigg αnpρ aparenrightbigg jnparenleftbigg αnqρ aparenrightbigg ρ2dρ=a3 2bracketleftbig jn+1(αnp)bracketrightbig2δpq. (11.169) Hereαnpandαnqarerootsof jn. This representsorthogonalitywithrespecttotherootsof theBesselfunctions.Anillus- trationofthissortoforthogonalityisprovidedinExample11.7.1,theproblemofaparticle in a sphere. Equation (11.169) guarantees orthogonality of the wave functions jn(r)for fixedn.(Ifnvaries,theaccompanyingsphericalharmonicwillprovideorthogonality.) Example 11.7.1 PARTICLE IN A SPHERE An illustration of the use of the spherical Bessel functions is provided by the problem of a quantum mechanical particle in a sphere of radius a. Quantum theory requires that the wavefunction ψ,describingourparticle,satisfy −¯h2 2m∇2ψ=Eψ, (11.170) and the boundary conditions (1) ψ(r≤a)remains finite, (2) ψ(a)=0. This corresponds to a square-well potential V=0,r≤a, andV=∞,r>a.H e r e¯his Planck’s constant divided by 2 π,mis the mass of our particle, and Eis, its energy. Let us determine the minimum value of the energy for which our wave equation has an acceptable solution. Equation (11.170) is Helmholtz’s equation with a radial part (compare Section 9.3 for separationofvariables): d2R dr2+2 rdR dr+bracketleftbigg k2−n(n+1) r2bracketrightbigg R=0, (11.171) withk2=2mE/¯h2. HencebyEq.(11.139), with n=0, R=Aj0(kr)+Bn0(kr). Wechoosetheorbitalangularmomentumindex n=0,for anyangulardependencewould raise the energy. The spherical Neumann function is rejected because of its divergent be- haviorattheorigin.Tosatisfythesecondboundarycondition(for allangles),werequire ka=√ 2mE ¯ha=α, (11.172) 732 Chapter 11 Bessel Functions whereαis a root of j0, that is, j0(α)=0. This has the effect of limiting the allowable energies to a certain discrete set, or, in other words, application of boundary condition (2) quantizestheenergy E. Thesmallest αisthefirst zeroof j0, α=π, and Emin=π2¯h2 2ma2=h2 8ma2, (11.173) which means that for any finite sphere the particle energy will have a positive minimum or zero-pointenergy.This is an illustrationof theHeisenberguncertaintyprinciplefor /Delta1p with/Delta1r≤a. In solid-state physics, astrophysics, and other areas of physics, we may wish to know how many different solutions (energy states) correspond to energies less than or equal to some fixed energy E0. For a cubic volume (Exercise 9.3.5) the problem is fairly simple. The considerably more difficult spherical case is worked out by R. H. Lambert, Am. J. Phys.36:417,1169(1968). Therelevantorthogonalityrelationforthe jn(kr)canbederivedfromtheintegralgiven inExercise11.7.23. /squaresolid Anotherform, orthogonalitywithrespecttotheindices,maybewrittenas integraldisplay∞ −∞jm(x)jn(x)dx=0,m/negationslash=n, m,n≥0. (11.174) Theproofisleft asExercise11.7.10.If m=n(compareExercise11.7.11),wehave integraldisplay∞ −∞bracketleftbig jn(x)bracketrightbig2dx=π 2n+1. (11.175) Most physical applications of orthogonal Bessel and spherical Bessel functions involve orthogonalitywithvaryingroots and an interval [0,a]and Eqs. (11.168) and (11.169) and Exercise11.7.23for continuous-energyeigenvalues. The spherical Bessel functions will enter again in connection with spherical waves, but furtherconsiderationispostponeduntilthecorrespondingangularfunctions,theLegendre functions,havebeenintroduced. Exercises 11.7.1 Showthatif nn(x)=radicalbiggπ 2xNn+1/2(x), itautomaticallyequals (−1)n+1radicalbiggπ 2xJ−n−1/2(x). 11.7 Spherical Bessel Functions 733 11.7.2 Derivethetrigonometric-polynomialforms of jn(z)andnn(z).27 jn(z)=1 zsinparenleftbigg z−nπ 2parenrightbigg[n/2]summationdisplay s=0(−1)s(n+2s)! (2s)!(2z)2s(n−2s)! +1 zcosparenleftbigg z−nπ 2parenrightbigg[(n−1)/2]summationdisplay s=0(−1)s(n+2s+1)! (2s+1)!(2z)2s(n−2s−1)!, nn(z)=(−1)n+1 zcosparenleftbigg z+nπ 2parenrightbigg[n/2]summationdisplay s=0(−1)s(n+2s)! (2s)!(2z)2s(n−2s)! +(−1)n+1 zsinparenleftbigg z+nπ 2parenrightbigg[(n−1)/2]summationdisplay s=0(−1)s(n+2s+1)! (2s+1)!(2z)2s+1(n−2s−1)!. 11.7.3 Usetheintegralrepresentationof Jν(x), Jν(x)=1 π1/2(ν−1 2)!parenleftbiggx 2parenrightbiggνintegraldisplay1 −1e±ixpparenleftbig 1−p2parenrightbigν−1/2dp, to show that the spherical Bessel functions jn(x)are expressible in terms of trigono- metricfunctions;thatis, for example, j0(x)=sinx x,j 1(x)=sinx x2−cosx x. 11.7.4 (a) Derivetherecurrencerelations fn−1(x)+fn+1(x)=2n+1 xfn(x), nfn−1(x)−(n+1)fn+1(x)=(2n+1)f′ n(x) satisfiedbythesphericalBessel functions jn(x),nn(x),h(1) n(x), andh(2) n(x). (b) Show,fromthesetworecurrencerelations,thatthesphericalBesselfunction fn(x) satisfiesthedifferentialequation x2f′′ n(x)+2xf′ n(x)+bracketleftbig x2−n(n+1)bracketrightbig fn(x)=0. 11.7.5 Provebymathematicalinductionthat jn(x)=(−1)nxnparenleftbigg1 xd dxparenrightbiggnparenleftbiggsinx xparenrightbigg fornanarbitrarynonnegativeinteger. 11.7.6 From the discussion of orthogonality of the spherical Bessel functions, show that a Wronskianrelationfor jn(x)andnn(x)is jn(x)n′ n(x)−j′ n(x)nn(x)=1 x2. 27Theupper limit on the summation [n/2]means the largest integerthat does not exceed n/2. 734 Chapter 11 Bessel Functions 11.7.7 Verify h(1) n(x)h(2)′ n(x)−h(1)′ n(x)h(2) n(x)=−2i x2. 11.7.8 VerifyPoisson’sintegralrepresentationofthesphericalBesselfunction, jn(z)=zn 2n+1n!integraldisplayπ 0cos(zcosθ)sin2n+1θdθ. 11.7.9 Showthatintegraldisplay∞ 0Jµ(x)Jν(x)dx x=2 πsin[(µ−ν)π/2] µ2−ν2,µ+ν>−1. 11.7.10 DeriveEq. (11.174): integraldisplay∞ −∞jm(x)jn(x)dx=0,m/negationslash=n m,n≥0. 11.7.11 DeriveEq. (11.175): integraldisplay∞ −∞bracketleftbig jn(x)bracketrightbig2dx=π 2n+1. 11.7.12 Set up the orthogonality integral for jL(kr)in a sphere of radius Rwith the boundary condition jL(kR)=0. The result is used in classifying electromagnetic radiation according to its angular mo- mentum. 11.7.13 The Fresnel integrals (Fig. 11.15 and Exercise 5.10.2) occurring in diffraction theory aregivenby x(t)=radicalbiggπ 2Cparenleftbiggradicalbiggπ 2tparenrightbigg =integraldisplayt 0cosparenleftbig v2parenrightbig dv, y(t)=radicalbiggπ 2Sparenleftbiggradicalbiggπ 2tparenrightbigg =integraldisplayt 0sinparenleftbig v2parenrightbig dv. Showthattheseintegralsmaybeexpandedinseries ofsphericalBesselfunctions x(s)=1 2integraldisplays 0j−1(u)u1/2du=s1/2∞summationdisplay n=0j2n(s), y(s)=1 2integraldisplays 0j0(u)u1/2du=s1/2∞summationdisplay n=0j2n+1(s). Hint. To establish the equality of the integral and the sum, you may wish to work with theirderivatives.ThesphericalBesselanalogsof Eqs. (11.12)and(11.14)arehelpful. 11.7.14 Ahollowsphereofradius a(Helmholtzresonator)containsstandingsoundwaves.Find theminimumfrequencyofoscillationintermsoftheradius aandthevelocityofsound v.Thesoundwavessatisfy thewaveequation ∇2ψ=1 v2∂2ψ ∂t2 11.7 Spherical Bessel Functions 735 FIGURE 11.15Fresnelintegrals. andtheboundarycondition ∂ψ ∂r=0,r=a. This is a Neumann boundary condition. Example 11.7.1 has the same PDE but with a Dirichletboundarycondition. ANS.νmin=0.3313v/a,λmax=3.018a. 11.7.15 DefiningthesphericalmodifiedBessel functions(Fig.11.16)by in(x)=radicalbiggπ 2xIn+1/2(x), k n(x)=radicalbigg 2 πxKn+1/2(x), showthat i0(x)=sinhx x,k 0(x)=e−x x. Notethatthenumericalfactors inthedefinitionsof inandknare notidentical. 11.7.16 (a) Showthattheparityof in(x)is(−1)n. (b) Showthat kn(x)hasnodefiniteparity. 736 Chapter 11 Bessel Functions FIGURE 11.16SphericalmodifiedBessel functions. 11.7.17 ShowthatthesphericalmodifiedBessel functionssatisfy thefollowingrelations: (a)in(x)=i−njn(ix), kn(x)=−inh(1) n(ix), (b)in+1(x)=xnd dxparenleftbig x−ninparenrightbig , kn+1(x)=−xnd dxparenleftbig x−nknparenrightbig , (c)in(x)=xnparenleftbigg1 xd dxparenrightbiggnsinhx x, kn(x)=(−1)nxnparenleftbigg1 xd dxparenrightbiggne−x x. 11.7.18 Showthattherecurrencerelationsfor in(x)andkn(x)are (a)in−1(x)−in+1(x)=2n+1 xin(x), nin−1(x)+(n+1)in+1(x)=(2n+1)i′ n(x), 11.7 Spherical Bessel Functions 737 (b)kn−1(x)−kn+1(x)=−2n+1 xkn(x), nkn−1(x)+(n+1)kn+1(x)=−(2n+1)k′ n(x). 11.7.19 Derivethelimitingvaluesfor thesphericalmodifiedBesselfunctions (a)in(x)≈xn (2n+1)!!,k n(x)≈(2n−1)!! xn+1,x≪1. (b)in(x)∼ex 2x,k n(x)∼e−x x,x≫1 2n(n+1). 11.7.20 ShowthattheWronskianofthesphericalmodifiedBesselfunctionsis givenby in(x)k′ n(x)−i′ n(x)kn(x)=−1 x2. 11.7.21 Aquantumparticleofmass Mistrappedina“square”wellofradius a.TheSchrödinger equationpotentialis V(r)=braceleftBigg −V0,0≤r<a 0,r>a. Theparticle’senergy Eis negative(aneigenvalue). (a) Show that the radial part of the wave function is given by jl(k1r)for 0≤r<a andkl(k2r)forr>a. (We require that ψ(0)be finite and ψ(∞)→0.) Here k2 1=2M(E+V0)/¯h2,k2 2=−2ME/¯h2, andlis the angular momentum ( nin Eq.(11.139)). (b) Theboundaryconditionat r=aisthatthewavefunction ψ(r)anditsfirstderiv- ativebecontinuous.Showthatthismeans (d/dr)j l(k1r) jl(k1r)vextendsinglevextendsinglevextendsinglevextendsingle r=a=(d/dr)k l(k2r) kl(k2r)vextendsinglevextendsinglevextendsinglevextendsingle r=a. Thisequationdeterminestheenergyeigenvalues. Note.This isageneralizationofExample10.1.2. 11.7.22 Thequantummechanicalradialwavefunctionfor ascatteredwaveisgivenby ψk=sin(kr+δ0) kr, wherekis the wave number, k=√2mE/¯h, andδ0is the scattering phase shift. Show thatthenormalizationintegralis integraldisplay∞ 0ψk(r)ψk′(r)r2dr=π 2kδ(k−k′). Hint.YoucanuseasinerepresentationoftheDiracdeltafunction.SeeExercise15.3.8. 738 Chapter 11 Bessel Functions 11.7.23 DerivethesphericalBessel functionclosurerelation 2a2 πintegraldisplay∞ 0jn(ar)jn(br)r2dr=δ(a−b). Note. An interesting derivation involving Fourier transforms, the Rayleigh plane-wave expansion, and spherical harmonics has been given by P. Ugincius, Am. J. Phys. 40: 1690(1972). 11.7.24 (a) Writea subroutinethatwillgeneratethesphericalBesselfunctions, jn(x), thatis, willgeneratethenumericalvalueof jn(x)givenxandn. Note.Onepossibilityistousetheexplicitknownformsof j0andj1andtodevelop thehigherindex jn, byrepeatedapplicationof therecurrencerelation. (b) Check your subroutine by an independent calculation, such as Eq. (11.154). If possible, compare the machine time needed for this check with the time required foryoursubroutine. 11.7.25 The wave function of a particle in a sphere (Example 11.7.1) with angular momen-tumlisψ(r,θ,ϕ)=Ajl((√ 2ME)r/¯h)Ym l(θ,ϕ).T h eYm l(θ,ϕ)is a spherical har- monic, described in Section 12.6. From the boundary condition ψ(a,θ,ϕ)=0o r jl((√ 2ME)a/¯h)=0 calculate the 10 lowest-energy states. Disregard the mdegen- eracy (2l+1 values of mfor each choice of l). Check your results against AMS-55, Table10.6, seeAdditionalReadingsfor Chapter8forthereference. Hint.YoucanuseyoursphericalBessel subroutineandaroot-findingsubroutine. Checkvalues. jl(αls)=0, α01=3.1416 α11=4.4934 α21=5.7635 α02=6.2832. 11.7.26 LetExample11.7.1bemodifiedso thatthepotentialisafinite V0outside(r >a). (a) For E<V0showthat ψout(r,θ,ϕ)∼klparenleftbiggr ¯hradicalbig 2M(V0−E)parenrightbigg . (b) Thenewboundaryconditionstobesatisfiedat r=aare ψin(a,θ,ϕ)=ψout(a,θ,ϕ), ∂ ∂rψin(a,θ,ϕ)=∂ ∂rψout(a,θ,ϕ) or 1 ψin∂ψin ∂rvextendsinglevextendsinglevextendsinglevextendsingle r=a=1 ψout∂ψout ∂rvextendsinglevextendsinglevextendsinglevextendsingle r=a. Forl=0 showthattheboundaryconditionat r=aleadsto f(E)=kbraceleftbigg cotka−1 kabracerightbigg +k′braceleftbigg 1+1 k′abracerightbigg =0, wherek=√ 2ME/¯handk′=√2M(V0−E)/¯h. 11.7 Additional Readings 739 (c) With a=4πε0¯h2/Me2(Bohr radius) and V0=4Me4/2¯h2, compute the possible boundstates (0<E<V 0). Hint. Call a root-finding subroutine after you know the approximate location of theroots of f(E)=0(0≤E≤V0). (d) Show that when a=4πε0¯h2/Me2the minimum value of V0for which a bound stateexistsis V0=2.4674Me4/2¯h2. 11.7.27 In some nuclear stripping reactions the differential cross section is proportional to jl(x)2,wherelistheangularmomentum.Thelocationofthemaximumonthecurveof experimentaldatapermitsadeterminationof l,ifthelocationofthe(first)maximumof jl(x)is known. Compute the location of the first maximum of j1(x),j2(x), andj3(x). Note.Forbetteraccuracylookforthefirstzeroof j′ l(x).Whyisthismoreaccuratethan directlocationof themaximum? AdditionalReadings Jackson, J. D., Classical Electrodynamics , 3rd ed.,NewYork: J. Wiley (1999). McBride,E.B., ObtainingGeneratingFunctions .NewYork:Springer-Verlag(1971).Anintroductiontomethods of obtaining generating functions. Watson, G. N., A Treatise on the Theory of Bessel Functions , 2nd ed. Cambridge, UK: Cambridge University Press (1952). This is the definitive text on Bessel functions and their properties. Although difficult reading, it is invaluable asthe ultimate reference. Watson,G.N., ATreatiseontheTheoryofBesselFunctions ,1sted.Cambridge,UK:CambridgeUniversityPress (1922). Seealso the references listed atthe endofChapter 13. This page intentionally left blank CHAPTER 12 LEGENDRE FUNCTIONS 12.1 G ENERATING FUNCTION Legendre polynomials appear in many different mathematical and physical situations. (1) They may originate as solutions of the Legendre ODE which we have already en- countered in the separation of variables (Section 9.3) for Laplace’s equation, Helmholtz’s equation,andsimilarODEsinsphericalpolarcoordinates.(2)Theyenterasaconsequence of a Rodrigues’ formula (Section 12.4). (3) They arise as a consequence of demanding a complete, orthogonal set of functions over the interval [−1,1](Gram–Schmidt orthogo- nalization, Section 10.3). (4) In quantum mechanics they (really the spherical harmonics, Sections 12.6 and 12.7) represent angular momentum eigenfunctions. (5) They are gen- erated by a generating function. We introduce Legendre polynomials here by way of a generatingfunction. Physical Basis — Electrostatics AswithBesselfunctions,itisconvenienttointroducetheLegendrepolynomialsbymeans of a generating function, which here appears in a physical context. Consider an electric chargeqplacedonthe z-axisatz=a.AsshowninFig.12.1,theelectrostaticpotentialof chargeqis ϕ=1 4πε0·q r1(SIunits). (12.1) We want to express the electrostatic potential in terms of the spherical polar coordinates r andθ(the coordinate ϕis absent because of symmetry about the z-axis). Using the law of cosinesinFig. 12.1,weobtain ϕ=q 4πε0parenleftbig r2+a2−2arcosθparenrightbig−1/2. (12.2) 741 742 Chapter 12 Legendre Functions FIGURE 12.1Electrostaticpotential. Chargeqdisplacedfrom origin. Legendre Polynomials Consider the case of r>aor, more precisely, r2>|a2−2arcosθ|. The radical in Eq. (12.2) may be expanded in a binomial series and then rearranged in powers of (a/r). The Legendre polynomial Pn(cosθ)(see Fig. 12.2) is defined as the coefficient of the nth powerin ϕ=q 4πε0r∞summationdisplay n=0Pn(cosθ)parenleftbigga rparenrightbiggn . (12.3) FIGURE 12.2Legendre polynomials P2(x),P3(x), P4(x), andP5(x). 12.1 Generating Function 743 Dropping the factor q/4πε0rand using xandtinstead of cos θanda/r, respectively, we have g(t,x)=parenleftbig 1−2xt+t2parenrightbig−1/2=∞summationdisplay n=0Pn(x)tn,|t|<1. (12.4) Equation (12.4) is our generating function formula. In the next section it is shown that |Pn(cosθ)|≤1, whichmeansthat theseries expansion(Eq. (12.4)) is convergentfor |t|< 1.1Indeed,theseries isconvergentfor |t|=1 exceptfor|x|=1. In physicalapplicationsEq. (12.4)oftenappearsinthevectorform(see Section9.7) 1 |r1−r2|=1 r>∞summationdisplay n=0parenleftbiggr< r>parenrightbiggn Pn(cosθ), (12.4a) where r>=|r1| r<=|r2|bracerightbigg for|r1|>|r2|, (12.4b) and r>=|r2| r<=|r1|bracerightbigg for|r2|>|r1|. (12.4c) Usingthebinomialtheorem(Section5.6)andExercise8.1.15,weexpandthegenerating functionas(compareEq.(12.33)) parenleftbig 1−2xt+t2parenrightbig−1/2=∞summationdisplay n=0(2n)! 22n(n!)2parenleftbig 2xt−t2parenrightbign =1+∞summationdisplay n=1(2n−1)!! (2n)!!parenleftbig 2xt−t2parenrightbign. (12.5) ForthefirstfewLegendrepolynomials,say, P0,P1,andP2,weneedthecoefficientsof t0, t1, andt2. These powers of tappear only in the terms n=0,1, and 2, and hence we may limitourattentiontothefirst threetermsoftheinfiniteseries: 0! 20(0!)2parenleftbig 2xt−t2parenrightbig0+2! 22(1!)2parenleftbig 2xt−t2parenrightbig1+4! 24(2!)2parenleftbig 2xt−t2parenrightbig2 =1t0+xt1+parenleftbigg3 2x2−1 2parenrightbigg t2+Oparenleftbig t3parenrightbig . Then,fromEq. (12.4)(anduniquenessof powerseries), P0(x)=1,P 1(x)=x, P 2(x)=3 2x2−1 2. Werepeatthislimiteddevelopmentinavectorframeworklaterinthissection. 1Note that the series in Eq. (12.3) is convergent for r>a, even though the binomial expansion involved is valid only for r>( a2+2ar)1/2and cosθ=−1,orr>a(1+√ 2). 744 Chapter 12 Legendre Functions Inemployingageneraltreatment,wefindthatthebinomialexpansionofthe (2xt−t2)n factoryieldsthedoubleseries parenleftbig 1−2xt+t2parenrightbig−1/2=∞summationdisplay n=0(2n)! 22n(n!)2tnnsummationdisplay k=0(−1)kn! k!(n−k)!(2x)n−ktk =∞summationdisplay n=0nsummationdisplay k=0(−1)k(2n)! 22nn!k!(n−k)!·(2x)n−ktn+k.(12.6) FromEq. (5.64) ofSection5.4(rearrangingtheorder ofsummation),Eq. (12.6) becomes parenleftbig 1−2xt+t2parenrightbig−1/2=∞summationdisplay n=0[n/2]summationdisplay k=0(−1)k(2n−2k)! 22n−2kk!(n−k)!(n−2k)!·(2x)n−2ktn,(12.7) with thetnindependent of the index k.2Now, equating our two power series (Eqs. (12.4) and(12.7)) termbyterm,wehave3 Pn(x)=[n/2]summationdisplay k=0(−1)k(2n−2k)! 2nk!(n−k)!(n−2k)!xn−2k. (12.8) Hence,for neven,Pnhasonlyevenpowersof xandevenparity(seeEq.(12.37)),andodd powersandoddparityfor odd n. Linear Electric Multipoles Returning to the electric charge on the z-axis, we demonstrate the usefulness and power of the generating function by adding a charge −qatz=−a, as shown in Fig. 12.3. The FIGURE 12.3Electricdipole. 2[n/2]=n/2f o rneven,(n−1)/2f o rnodd. 3Equation (12.8) starts with xn. By changing the index, we can transform it into a series that starts with x0forneven and x1 fornodd. Theseascending series aregiven as hypergeometric functions in Eqs. (13.138) and (13.139), Section 13.4. 12.1 Generating Function 745 potentialbecomes ϕ=q 4πε0parenleftbigg1 r1−1 r2parenrightbigg , (12.9) andbyusingthelawofcosines,wehave ϕ=q 4πε0rbraceleftbiggbracketleftbigg 1−2parenleftbigga rparenrightbigg cosθ+parenleftbigga rparenrightbigg2bracketrightbigg−1/2 −bracketleftbigg 1+2parenleftbigga rparenrightbigg cosθ+parenleftbigga rparenrightbigg2bracketrightbigg−1/2bracerightbigg ,(r>a). Clearly, the second radical is like the first, except that ahas been replaced by −a. Then, usingEq. (12.4), weobtain ϕ=q 4πε0rbracketleftbigg∞summationdisplay n=0Pn(cosθ)parenleftbigga rparenrightbiggn −∞summationdisplay n=0Pn(cosθ)(−1)nparenleftbigga rparenrightbiggnbracketrightbigg =2q 4πε0rbracketleftbigg P1(cosθ)parenleftbigga rparenrightbigg +P3(cosθ)parenleftbigga rparenrightbigg3 +···bracketrightbigg . (12.10) Thefirstterm(anddominanttermfor r≫a)i s ϕ=2aq 4πε0·P1(cosθ) r2, (12.11) which is the electric dipole potential, and 2 aqis the dipole moment (Fig. 12.3). This analysis may be extended by placing additional charges on the z-axis so that the P1term, as well as the P0(monopole) term, is canceled. For instance, charges of qatz=aand z=−a,−2qatz=0giverisetoapotentialwhoseseriesexpansionstartswith P2(cosθ). This is a linear electric quadrupole. Two linear quadrupoles may be placed so that the quadrupoletermis canceledbutthe P3,theoctupoleterm,survives. Vector Expansion Weconsidertheelectrostaticpotentialproducedbyadistributedcharge ρ(r2): ϕ(r1)=1 4πε0integraldisplayρ(r2) |r1−r2|d3r2. (12.12a) This expression has already appeared in Sections 1.16 and 9.7. Taking the denominator of the integrand, using first the law of cosines and then a binomial expansion, yields (see Fig.1.42) 1 |r1−r2|=parenleftbig r2 1−2r1·r2+r2 2parenrightbig−1/2(12.12b) =1 r1bracketleftbigg 1+parenleftbigg −2r1·r2 r2 1+r2 2 r2 1parenrightbiggbracketrightbigg−1/2 ,forr1>r2 =1 r1bracketleftbigg 1+r1·r2 r2 1−1 2r2 2 r2 1+3 2(r1·r2)2 r4 1+Oparenleftbiggr2 r1parenrightbigg3bracketrightbigg . 746 Chapter 12 Legendre Functions (Forr1=1,r2=t, andr1·r2=xt, Eq. (12.12b) reduces to the generating function, Eq.(12.4).) Thefirst terminthesquarebracket,1, yieldsa potential ϕ0(r1)=1 4πε01 r1integraldisplay ρ(r2)d3r2. (12.12c) Theintegralisjustthetotalcharge.Thispartofthetotalpotentialisanelectric monopole . Thesecondtermyields ϕ1(r1)=1 4πε0r1· r3 1integraldisplay r2ρ(r2)d3r2, (12.12d) where the integral is the dipole moment whose charge density ρ(r2)is weighted by a mo- ment arm r2. We have an electric dipole potential. For atomic or nuclear states of definite parity,ρ(r2)is anevenfunctionandthedipoleintegralis identicallyzero. The last two terms, both of order (r2/r1)2, may be handled by using Cartesian coordi- nates: (r1·r2)2=3summationdisplay i=1x1ix2i3summationdisplay j=1x1jx2j. Rearrangingvariablestotakethe x1componentsoutsidetheintegralyields ϕ2(r1)=1 4πε01 2r5 13summationdisplay i,j=1x1ix1jintegraldisplaybracketleftbig 3x2ix2j−δijr2 2bracketrightbig ρ(r2)d3r2. (12.12e) This is the electric quadrupole term. We note that the square bracket in the integrand formsasymmetric,zero-tracetensor. AgeneralelectrostaticmultipoleexpansioncanalsobedevelopedbyusingEq.(12.12a) forthepotential ϕ(r1)andreplacing1 /(4π|r1−r2|)byGreen’sfunction,Eq.(9.187).This yields the potential ϕ(r1)as a (double) series of the spherical harmonics Ym l(θ1,ϕ1)and Ym l(θ2,ϕ2). Before leavingmultipolefields,perhapsweshouldemphasizethreepoints. •First, an electric (or magnetic) multipole is isolated and well defined only if all lower- order multipoles vanish. For instance, the potential of one charge qatz=awas ex- panded in a series of Legendre polynomials. Although we refer to the P1(cosθ)term in this expansion as a dipole term, it should be remembered that this term exists only becauseofour choiceofcoordinates.We alsohaveamonopole, P0(cosθ). •Second, in physical systems we do not encounter pure multipoles. As an example, the potential of the finite dipole ( qatz=a,−qatz=−a) contained a P3(cosθ) term.Thesehigher-ordertermsmaybeeliminatedbyshrinkingthemultipoletoapoint multipole, in this case keeping the product qaconstant(a→0,q→∞)to maintain thesamedipolemoment. 12.1 Generating Function 747 •Third,themultipoletheoryisnotrestrictedtoelectricalphenomena.Planetaryconfigu- rationsaredescribedintermsofmassmultipoles,Sections12.3and12.6.Gravitational radiation depends on the time behavior of mass quadrupoles. (The gravitational radia- tion field is a tensorfield. The radiation quanta, gravitons, carry two units of angular momentum.) It might also be noted that a multipole expansion is actually a decomposition into the irreduciblerepresentationsoftherotationgroup(Section4.2). Extension to Ultraspherical Polynomials The generating function used here, g(t,x), is actually a special case of a more general generatingfunction, 1 (1−2xt+t2)α=∞summationdisplay n=0C(α) n(x)tn. (12.13) The coefficients C(α) n(x)are the ultraspherical polynomials (proportional to the Gegen- bauer polynomials). For α=1/2 this equation reduces to Eq. (12.4); that is, C(1/2) n(x)= Pn(x). The cases a=0 andα=1 are considered in Chapter 13 in connection with the Chebyshevpolynomials. Exercises 12.1.1 Develop the electrostatic potential for the array of charges shown. This is a linear elec- tricquadrupole(Fig. 12.4). 12.1.2 Calculate the electrostatic potential of the array of charges shown in Fig. 12.5. Here is an example of two equal but oppositely directed dipoles. The dipole contributions cancel.Theoctupoletermsdonotcancel. 12.1.3 Showthattheelectrostaticpotentialproducedbya charge qatz=aforr<ais ϕ(r)=q 4πε0a∞summationdisplay n=0parenleftbiggr aparenrightbiggn Pn(cosθ). FIGURE 12.4Linearelectricquadrupole. 748 Chapter 12 Legendre Functions FIGURE 12.5Linearelectricoctupole. FIGURE 12.6 12.1.4 UsingE=−∇ϕ, determine the components of the electric field corresponding to the (pure)electricdipolepotential ϕ(r)=2aqP1(cosθ) 4πε0r2. Hereitisassumedthat r≫a. ANS.Er=+4aqcosθ 4πε0r3,Eθ=+2aqsinθ 4πε0r3,Eϕ=0. 12.1.5 Apointelectricdipoleofstrength p(1)isplacedat z=a;asecondpointelectricdipole of equal but opposite strength is at the origin. Keeping the product p(1)aconstant, let a→0.Showthatthis resultsinapointelectricquadrupole. Hint.Exercise12.2.5(whenproved)willbehelpful. 12.1.6 A point charge qis in the interior of a hollow conducting sphere of radius r0.T h e chargeqis displaced a distance afrom the center of the sphere. If the conducting sphere is grounded, show that the potential in the interior produced by qand the dis- tributed induced charge is the same as that produced by qand its image charge q′.T h e image charge is at a distance a′=r2 0/afrom the center, collinear with qand the origin (Fig.12.6). Hint. Calculate the electrostatic potential for a<r0<a′. Show that the potential van- ishesforr=r0ifwetake q′=−qr0/a. 12.1.7 Provethat Pn(cosθ)=(−1)nrn+1 n!∂n ∂znparenleftbigg1 rparenrightbigg . Hint.ComparetheLegendrepolynomialexpansionofthegeneratingfunction( a→/Delta1z, Fig.12.1)withaTaylorseriesexpansionof1 /r,wherezdependenceof rchangesfrom ztoz−/Delta1z(Fig.12.7). 12.1.8 Bydifferentiationanddirectsubstitutionoftheseriesform,Eq.(12.8),showthat Pn(x) satisfies the Legendre ODE. Note that there is no restriction upon x. We may have any x,−∞<x<∞, andindeedany zintheentirefinitecomplexplane. 12.2 Recurrence Relations 749 FIGURE 12.7 12.1.9 TheChebyshevpolynomials(typeII) are generatedby(Eq. (13.93), Section13.3) 1 1−2xt+t2=∞summationdisplay n=0Un(x)tn. Using the techniques of Section 5.4 for transforming series, develop a series represen- tationofUn(x). ANS.Un(x)=[n/2]summationdisplay k=0(−1)k(n−k)! k!(n−2k)!(2x)n−2k. 12.2 R ECURRENCE RELATIONS AND SPECIAL PROPERTIES Recurrence Relations The Legendre polynomial generating function provides a convenient way of deriving the recurrence relations4and some special properties. If our generating function (Eq. (12.4)) isdifferentiatedwithrespectto t, weobtain ∂g(t,x) ∂t=x−t (1−2xt+t2)3/2=∞summationdisplay n=0nPn(x)tn−1. (12.14) BysubstitutingEq. (12.4)intothisandrearrangingterms,wehave parenleftbig 1−2xt+t2parenrightbig∞summationdisplay n=0nPn(x)tn−1+(t−x)∞summationdisplay n=0Pn(x)tn=0.(12.15) The left-hand side is a power series in t. Since this power series vanishes for all values of t, the coefficient of each power of tis equal to zero; that is, our power series is unique (Section 5.7). These coefficients are found by separating the individual summations and 4Wecanalso apply theexplicit series form Eq.(12.8) directly. 750 Chapter 12 Legendre Functions usingdistinctivesummationindices: ∞summationdisplay m=0mPm(x)tm−1−∞summationdisplay n=02nxPn(x)tn+∞summationdisplay s=0sPs(x)ts+1 +∞summationdisplay s=0Ps(x)ts+1−∞summationdisplay n=0xPn(x)tn=0. (12.16) Now,letting m=n+1,s=n−1,wefind (2n+1)xPn(x)=(n+1)Pn+1(x)+nPn−1(x), n=1,2,3,.... (12.17) This is another three-term recurrence relation, similar to (but not identicalwith) the recur- rence relation for Bessel functions. With this recurrence relation we may easily construct the higher Legendre polynomials. If we take n=1 and insert the easily found values of P0(x)andP1(x)(Exercise12.1.7orEq. (12.8)), weobtain 3xP1(x)=2P2(x)+P0(x), (12.18) or P2(x)=1 2parenleftbig 3x2−1parenrightbig . (12.19) This process may be continued indefinitely, the first few Legendre polynomials are listed inTable12.1. As cumbersome as it may appear at first, this technique is actually more efficient for adigitalcomputerthanisdirectevaluationoftheseries(Eq.(12.8)).Forgreaterstability(to avoid undue accumulation and magnification of round-off error), Eq. (12.17) is rewritten as Pn+1(x)=2xPn(x)−Pn−1(x)−1 n+1bracketleftbig xPn(x)−Pn−1(x)bracketrightbig . (12.17a) Onestartswith P0(x)=1,P1(x)=x,andcomputesthe numerical valuesofallthe Pn(x) for a given value of xup to the desired PN(x). The values of Pn(x),0≤n<N,a r e availableas afringebenefit. Table 12.1 LegendrePolynomials P0(x)=1 P1(x)=x P2(x)=1 2(3x2−1) P3(x)=1 2(5x3−3x) P4(x)=1 8(35x4−30x2+3) P5(x)=1 8(63x5−70x3+15x) P6(x)=1 16(231x6−315x4+105x2−5) P7(x)=1 16(429x7−693x5+315x3−35x) P8(x)=1 128(6435x8−12012x6+6930x4−1260x2+35) 12.2 Recurrence Relations 751 Differential Equations More information about the behavior of the Legendre polynomials can be obtained if we nowdifferentiateEq.(12.4) withrespectto x.Thisgives ∂g(t,x) ∂x=t (1−2xt+t2)3/2=∞summationdisplay n=0P′ n(x)tn, (12.20) or parenleftbig 1−2xt+t2parenrightbig∞summationdisplay n=0P′ n(x)tn−t∞summationdisplay n=0Pn(x)tn=0. (12.21) Asbefore,thecoefficientofeachpowerof tisset equaltozeroandweobtain P′ n+1(x)+P′ n−1(x)=2xP′ n(x)+Pn(x). (12.22) AmoreusefulrelationmaybefoundbydifferentiatingEq.(12.17)withrespectto xand multiplying by 2. To this we add (2n+1)times Eq. (12.22), canceling the P′ nterm. The resultis P′ n+1(x)−P′ n−1(x)=(2n+1)Pn(x). (12.23) From Eqs. (12.22) and (12.23) numerous additional equations may be developed,5in- cluding P′ n+1(x)=(n+1)Pn(x)+xP′ n(x), (12.24) P′ n−1(x)=−nPn(x)+xP′ n(x), (12.25) parenleftbig 1−x2parenrightbig P′ n(x)=nPn−1(x)−nxPn(x), (12.26) parenleftbig 1−x2parenrightbig P′ n(x)=(n+1)xPn(x)−(n+1)Pn+1(x). (12.27) By differentiating Eq. (12.26) and using Eq. (12.25) to eliminate P′ n−1(x), we find that Pn(x)satisfiesthelinearsecond-orderODE parenleftbig 1−x2parenrightbig P′′ n(x)−2xP′ n(x)+n(n+1)Pn(x)=0. (12.28) The previous equations, Eqs. (12.22) to (12.27), are all first-order ODEs, but with poly- nomials of two different indices. The price for having all indices alike is a second-order 5Usingthe equation number in parentheses to denote the left-hand side ofthe equation, wemay writethe derivatives as 2·d dx(12.17)+(2n+1)·(12.22)⇒(12.23), 1 2braceleftbig (12.22)+(12.23)bracerightbig ⇒(12.24), 1 2braceleftbig (12.22)−(12.23)bracerightbig ⇒(12.25), (12.24)n→n−1+x·(12.25)⇒(12.26), d dx(12.26)+n·(12.25)⇒(12.28). 752 Chapter 12 Legendre Functions differentialequation.Equation(12.28)is Legendre’s ODE.Wenowseethatthepolynomi- alsPn(x)generatedbythepowerseriesfor (1−2xt+t2)−1/2satisfyLegendre’sequation, which,of course,iswhytheyarecalledLegendrepolynomials. In Eq. (12.28) differentiation is with respect to x( x=cosθ). Frequently, we encounter Legendre’sequationexpressedintermsofdifferentiationwithrespectto θ: 1 sinθd dθparenleftbigg sinθdPn(cosθ) dθparenrightbigg +n(n+1)Pn(cosθ)=0. (12.29) Special Values Our generating function provides still more information about the Legendre polynomials. If weset x=1,Eq.(12.4) becomes 1 (1−2t+t2)1/2=1 1−t=∞summationdisplay n=0tn, (12.30) usingabinomialexpansionorthegeometricseries,Example5.1.1.ButEq.(12.4)for x=1 defines 1 (1−2t+t2)1/2=∞summationdisplay n=0Pn(1)tn. Comparingthetwoseries expansions(uniquenessof powerseries, Section5.7), wehave Pn(1)=1. (12.31) If weletx=−1 inEq. (12.4) anduse 1 (1+2t+t2)1/2=1 1+t, thisshowsthat Pn(−1)=(−1)n. (12.32) For obtaining these results, we find that the generating function is more convenient than theexplicitseries form,Eq. (12.8). If wetake x=0 inEq. (12.4), usingthebinomialexpansion parenleftbig 1+t2parenrightbig−1/2=1−1 2t2+3 8t4+···+(−1)n1·3···(2n−1) 2nn!t2n+···,(12.33) wehave6 P2n(0)=(−1)n1·3···(2n−1) 2nn!=(−1)n(2n−1)!! (2n)!!=(−1)n(2n)! 22n(n!)2(12.34) P2n+1(0)=0,n=0,1,2.... (12.35) TheseresultsalsofollowfromEq. (12.8)byinspection. 6Thedouble factorial notation is defined in Section8.1: (2n)!!=2·4·6···(2n), ( 2n−1)!!=1·3·5···(2n−1), (−1)!!=1. 12.2 Recurrence Relations 753 Parity SomeoftheseresultsarespecialcasesoftheparitypropertyoftheLegendrepolynomials. We refer once more to Eqs. (12.4) and (12.8). If we replace xby−xandtby−t,t h e generatingfunctionis unchanged.Hence g(t,x)=g(−t,−x)=bracketleftbig 1−2(−t)(−x)+(−t)2bracketrightbig−1/2 =∞summationdisplay n=0Pn(−x)(−t)n=∞summationdisplay n=0Pn(x)tn. (12.36) Comparingthesetwo series, wehave Pn(−x)=(−1)nPn(x); (12.37) thatis,thepolynomialfunctionsareoddoreven(withrespectto x=0,θ=π/2)according to whether the index nis odd or even. This is the parity,7or reflection, property that plays such an important role in quantum mechanics. For central forces the index nis a measure oftheorbitalangularmomentum,thuslinkingparityandorbitalangularmomentum. This parity property is confirmed by the series solution and for the special values tabu- lated in Table 12.1. It might also be noted that Eq. (12.37) may be predicted by inspection of Eq. (12.17), the recurrence relation. Specifically, if Pn−1(x)andxPn(x)are even, then Pn+1(x)mustbeeven. Upper and Lower Bounds for Pn(cosθ) Finally,inadditiontotheseresults,ourgeneratingfunctionenablesustosetanupperlimit on|Pn(cosθ)|.W eha v e parenleftbig 1−2tcosθ+t2parenrightbig−1/2=parenleftbig 1−teiθparenrightbig−1/2parenleftbig 1−te−iθparenrightbig−1/2 =parenleftbig 1+1 2teiθ+3 8t2e2iθ+···parenrightbig ·parenleftbig 1+1 2te−iθ+3 8t2e−2iθ+···parenrightbig ,(12.38) with all coefficients positive. Our Legendre polynomial, Pn(cosθ), still the coefficient of tn, maynowbewrittenasasumof termsoftheform 1 2amparenleftbig eimθ+e−imθparenrightbig =amcosmθ (12.39a) withallthe ampositiveandmandnbothevenor oddso that Pn(cosθ)=nsummationdisplay m=0o r1amcosmθ. (12.39b) 7In spherical polar coordinates the inversion of the point (r,θ,ϕ)through the origin is accomplished by the transformation [r→r,θ→π−θ,andϕ→ϕ±π].Then,cos θ→cos(π−θ)=−cosθ,correspondingto x→−x(compareExercise2.5.8). 754 Chapter 12 Legendre Functions This series, Eq. (12.39b), is clearly a maximum when θ=0 and cos mθ=1. But for x= cosθ=1,Eq.(12.31) showsthat Pn(1)=1.Therefore vextendsinglevextendsinglePn(cosθ)vextendsinglevextendsingle≤Pn(1)=1. (12.39c) A fringe benefit of Eq. (12.39b) is that it shows that our Legendre polynomial is a linear combination of cos mθ. This means that the Legendre polynomials form a complete set for any functions that may be expanded by a Fourier cosine series (Section 14.1) over the interval[0,π]. •In this section various useful properties of the Legendre polynomials are derived from thegeneratingfunction,Eq.(12.4). •The explicit series representation, Eq. (12.8), offers an alternate and sometimes supe- rior approach. Exercises 12.2.1 Giventheseries α0+α2cos2θ+α4cos4θ+α6cos6θ=a0P0+a2P2+a4P4+a6P6, express the coefficients αias a column vector αand the coefficients aias a column vectoraanddeterminethematrices AandBsuchthat Aα=aandBa=α. Checkyourcomputationbyshowingthat AB=1(unitmatrix).Repeatfortheoddcase α1cosθ+α3cos3θ+α5cos5θ+α7cos7θ=a1P1+a3P3+a5P5+a7P7. Note.Pn(cosθ)and cosnθare tabulated in terms of each other in AMS-55 (see Addi- tionalReadingsof Chapter8forthecompletereference). 12.2.2 By differentiating the generating function g(t,x)with respect to t, multiplying by 2 t, andthenadding g(t,x), showthat 1−t2 (1−2tx+t2)3/2=∞summationdisplay n=0(2n+1)Pn(x)tn. This result is useful in calculating the charge induced on a grounded metal sphere by a pointcharge q. 12.2.3 (a) DeriveEq. (12.27), parenleftbig 1−x2parenrightbig P′ n(x)=(n+1)xPn(x)−(n+1)Pn+1(x). (b) Write out the relation of Eq. (12.27) to preceding equations in symbolic form analogoustothesymbolicforms forEqs. (12.23) to(12.26). 12.2 Recurrence Relations 755 12.2.4 A point electric octupole may be constructed by placing a point electric quadrupole (pole strength p(2)in thez-direction) at z=aand an equal but opposite point elec- tric quadrupole at z=0 and then letting a→0, subject to p(2)a=constant. Find the electrostatic potential corresponding to a point electric octupole. Show from the con- structionofthepointelectricoctupolethatthecorrespondingpotentialmaybeobtained bydifferentiatingthepointquadrupolepotential. 12.2.5 Operatingin sphericalpolarcoordinates ,showthat ∂ ∂zbracketleftbiggPn(cosθ) rn+1bracketrightbigg =−(n+1)Pn+1(cosθ) rn+2. This is the key step in the mathematical argument that the derivative of one multipole leadstothenexthighermultipole. Hint.CompareExercise2.5.12. 12.2.6 From PL(cosθ)=1 L!∂L ∂tLparenleftbig 1−2tcosθ+t2parenrightbig−1/2vextendsinglevextendsingle t=0 showthat PL(1)=1,P L(−1)=(−1)L. 12.2.7 Provethat P′ n(1)=d dxPn(x)vextendsinglevextendsingle x=1=1 2n(n+1). 12.2.8 Show that Pn(cosθ)=(−1)nPn(−cosθ)by use of the recurrence relation relating Pn,Pn+1, andPn−1andyourknowledgeof P0andP1. 12.2.9 From Eq. (12.38) write out the coefficient of t2in terms of cos nθ,n≤2. This coeffi- cientisP2(cosθ). 12.2.10 Write a program that will generate the coefficients asin the polynomial form of the Legendrepolynomial Pn(x)=nsummationdisplay s=0asxs. 12.2.11 (a) Calculate P10(x)overtherange [0,1]andplotyourresults. (b) Calculateprecise(atleasttofivedecimalplaces)valuesofthefivepositiverootsof P10(x). Compare your values with the values listed in AMS-55, Table 25.4. (For thecompletereference,seeAdditionalReadingsofChapter8.) 12.2.12 (a) Calculatethe largestrootofPn(x)forn=2(1)50. (b) Develop an approximation for the largest root from the hypergeometric represen- tation of Pn(x)(Section 13.4) and compare your values from part (a) with your hypergeometric approximation. Compare also with the values listed in AMS-55, Table25.4.(For thecompletereference,seeAdditionalReadingsof Chapter8.) 756 Chapter 12 Legendre Functions 12.2.13 (a) From Exercise 12.2.1 and AMS-55, Table 22.9, develop the 6 ×6m a t r i xBthat will transform a series of even-order Legendre polynomials through P10(x)into a powerseriessummationtext5 n=0α2nx2n. (b) Calculate AasB−1.Checktheelementsof AagainstthevalueslistedinAMS-55, Table22.9.(For thecompletereference,seeAdditionalReadingsof Chapter8.) (c) By using matrix multiplication, transform some even power seriessummationtext5 n=0α2nx2n intoaLegendreseries. 12.2.14 Write a subroutinethatwill transform a finitepower seriessummationtextN n=0anxnintoa Legendre seriessummationtextN n=0bnPn(x).Usetherecurrencerelation,Eq.(12.17),andfollowthetechnique outlinedinSection13.3for aChebyshevseries. 12.3 O RTHOGONALITY Legendre’sODE(12.28)maybewrittenintheform d dxbracketleftbigparenleftbig 1−x2parenrightbig P′ n(x)bracketrightbig +n(n+1)Pn(x)=0, (12.40) showing clearly that it is self-adjoint. Subject to satisfying certain boundary condi- tions, then, it is known that the solutions Pn(x)will be orthogonal. Upon comparing Eq. (12.40) with Eqs. (10.6) and (10.8) we see that the weight function w(x)=1,L= (d/dx)(1−x2)(d/dx),p(x)=1−x2, and the eigenvalue λ=n(n+1). The integration limitson xare±1,where p(±1)=0.Thenfor m/negationslash=n,Eq. (10.34)becomes integraldisplay1 −1Pn(x)Pm(x)dx=0,8(12.41) integraldisplayπ 0Pn(cosθ)Pm(cosθ)sinθdθ=0, (12.42) showing that Pn(x)andPm(x)are orthogonal for the interval [−1,1]. This orthogonality mayalsobedemonstratedbyusingRodrigues’definitionof Pn(x)(compareSection12.4, Exercise12.4.2). Weshallneedtoevaluatetheintegral(Eq.(12.41))when n=m.Certainlyitisnolonger zero.Fromourgeneratingfunction, parenleftbig 1−2tx+t2parenrightbig−1=bracketleftbigg∞summationdisplay n=0Pn(x)tnbracketrightbigg2 . (12.43) Integratingfrom x=−1t ox=+1,wehave integraldisplay1 −1dx 1−2tx+t2=∞summationdisplay n=0t2nintegraldisplay1 −1bracketleftbig Pn(x)bracketrightbig2dx; (12.44) 8In Section 10.4 suchintegrals are interpreted as inner products in a linearvector(function) space.Alternate notations are integraldisplay1 −1bracketleftbig Pn(x)bracketrightbig∗Pm(x)dx≡/angbracketleftPn|Pm/angbracketright≡(Pn,Pm). The/angbracketleft/angbracketrightform, popularized by Dirac, is common in the physics literature. The () form is more common in the mathematics literature. 12.3 Orthogonality 757 the cross terms in the series vanish by means of Eq. (12.42). Using y=1−2tx+t2, dy=−2tdx,weobtain integraldisplay1 −1dx 1−2tx+t2=1 2tintegraldisplay(1+t)2 (1−t)2dy y=1 tlnparenleftbigg1+t 1−tparenrightbigg . (12.45) Expandingthisina powerseries (Exercise5.4.1)givesus 1 tlnparenleftbigg1+t 1−tparenrightbigg =2∞summationdisplay n=0t2n 2n+1. (12.46) Comparingpower-seriescoefficientsofEqs. (12.44) and(12.46), wemusthave integraldisplay1 −1bracketleftbig Pn(x)bracketrightbig2dx=2 2n+1. (12.47) CombiningEq. (12.42)withEq.(12.47) wehavetheorthonormalitycondition integraldisplay1 −1Pm(x)Pn(x)dx=2δmn 2n+1. (12.48) We shall return to this result in Section 12.6 when we construct the orthonormal spherical harmonics. Expansion of Functions, Legendre Series Inadditiontoorthogonality,theSturm–LiouvilletheoryimpliesthattheLegendrepolyno- mialsform acompleteset. Letusassume, then,thattheseries ∞summationdisplay n=0anPn(x)=f(x) (12.49) converges in the mean (Section 10.4) in the interval [−1,1]. This demands that f(x)and f′(x)be at least sectionally continuous in this interval. The coefficients anare found by multiplying the series by Pm(x)and integrating term by term. Using the orthogonality propertyexpressedinEqs. (12.42)and(12.48),weobtain 2 2m+1am=integraldisplay1 −1f(x)Pm(x)dx. (12.50) Wereplacethevariableof integration xbytandtheindex mbyn.Then,substitutinginto Eq.(12.49), wehave f(x)=∞summationdisplay n=02n+1 2parenleftbiggintegraldisplay1 −1f(t)Pn(t)dtparenrightbigg Pn(x). (12.51) This expansion in a series of Legendre polynomials is usually referred to as a Legendre series.9Its properties are quite similar to the more familiar Fourier series (Chapter 14). In 9Notethat Eq. (12.50) gives amasadefiniteintegral, that is, anumber for a given f(x). 758 Chapter 12 Legendre Functions particular, we can use the orthogonality property (Eq. (12.48)) to show that the series is unique. On a more abstract (and more powerful) level, Eq. (12.51) gives the representation of f(x)inthevectorspaceofLegendrepolynomials(a Hilbertspace,Section10.4). From the viewpoint of integral transforms (Chapter 15), Eq. (12.50) may be considered afiniteLegendretransformof f(x).Equation(12.51)isthentheinversetransform.Itmay also be interpreted in terms of the projection operators of quantum theory. We may take Pmin [Pmf](x)≡Pm(x)2m+1 2integraldisplay1 −1Pm(t)bracketleftbig f(t)bracketrightbig dt asan(integral)operator,readytooperateon f(t).(Thef(t)wouldgointhesquarebracket asafactor intheintegrand.)Then,fromEq. (12.50), [Pmf](x)=amPm(x).10 Theoperator Pmprojectsoutthe mthcomponentof thefunction f. Equation (12.3), which leads directly to the generating function definition of Legendre polynomials, is a Legendre expansion of 1 /r1. This Legendre expansion of 1 /r1or 1/r12 appears in several exercises of Section 12.8. Going beyond a Coulomb field, the 1 /r12is oftenreplacedbyapotential V(|r1−r2|),andthesolutionoftheproblemisagaineffected byaLegendreexpansion. The Legendre series, Eq. (12.49), has been treated as a knownfunctionf(x)that we arbitrarily chose to expandin a series of Legendrepolynomials.Sometimesthe origin and nature of the Legendre series are different. In the next examples we consider unknown functions we know can be represented by a Legendre series because of the differential equation the unknown functions satisfy. As before, the problem is to determine the un- known coefficients in the series expansion. Here, however, the coefficients are not found by Eq. (12.50). Rather, they are determined by demanding that the Legendre series match aknownsolutionataboundary.These areboundaryvalueproblems. Example 12.3.1 EARTH ’SGRAVITATIONAL FIELD AnexampleofaLegendreseriesisprovidedbythedescriptionoftheEarth’sgravitational potential U(for exteriorpoints),neglectingazimuthaleffects. With R=equatorialradius =6378.1±0.1km GM R=62.494±0.001km2/s2, wewrite U(r,θ)=GM RbracketleftbiggR r−∞summationdisplay n=2anparenleftbiggR rparenrightbiggn+1 Pn(cosθ)bracketrightbigg , (12.52) 10Thedependent variables arearbitrary. Here xcamefrom the xinPm. 12.3 Orthogonality 759 aLegendreseries. Artificialsatellitemotionshaveshownthat a2=(1,082,635±11)×10−9, a3=(−2,531±7)×10−9, a4=(−1,600±12)×10−9. This is the famous pear-shaped deformation of the Earth. Other coefficients have been computed through n=20. Note that P1is omitted because the origin from which ris measuredis theEarth’s centerof mass( P1wouldrepresentadisplacement). More recent satellite data permit a determination of the longitudinal dependence of the Earth’s gravitational field. Such dependence may be described by a Laplace series (Sec- tion12.6). /squaresolid Example 12.3.2 SPHERE IN A UNIFORM FIELD Another illustration of the use of Legendre polynomials is provided by the problem of a neutral conducting sphere (radius r0) placed in a (previously) uniform electric field (Fig. 12.8). The problem is to find the new, perturbed, electrostatic potential. If we call theelectrostaticpotential11V, itsatisfies ∇2V=0, (12.53) Laplace’sequation.Weselectsphericalpolarcoordinatesbecauseofthesphericalshapeof the conductor. (This will simplify the application of the boundary condition at the surface oftheconductor.)SeparatingvariablesandglancingatTable9.2,wecanwritetheunknown potential V(r,θ)intheregionoutsidethesphereasalinearcombinationofsolutions: V(r,θ)=∞summationdisplay n=0anrnPn(cosθ)+∞summationdisplay n=0bnPn(cosθ) rn+1. (12.54) FIGURE 12.8Conductingspherein auniformfield. 11It should beemphasizedthatthis isnot apresentationofaLegendre-series expansion ofaknown V(cosθ).Hereweareback toboundary value problems of PDEs. 760 Chapter 12 Legendre Functions Noϕ-dependence appears because of the axial symmetry of our problem. (The center of theconductingsphereistakenastheoriginandthe z-axisisorientedparalleltotheoriginal uniformfield.) It might be noted here that nis an integer, because only for integral nis theθdepen- dence well behaved at cos θ=±1. For nonintegral nthe solutions of Legendre’s equation divergeattheendsoftheinterval [−1,1],thepoles θ=0,πofthesphere(compareExam- ple5.2.4andExercises5.2.15and9.5.5).Itisforthissamereasonthatthesecondsolution ofLegendre’sequation, Qn,is alsoexcluded. Now we turn to our (Dirichlet) boundary conditions to determine the unknown anand bnof our series solution, Eq. (12.54). If the original, unperturbed electrostatic field is E0, werequire, asoneboundarycondition, V(r→∞)=−E0z=−E0rcosθ=−E0rP1(cosθ). (12.55) SinceourLegendreseriesisunique,wemayequatecoefficientsof Pn(cosθ)inEq.(12.54) (r→∞)andEq.(12.55) toobtain an=0,n>1 and n=0,a 1=−E0. (12.56) Ifan/negationslash=0f o rn>1, these terms would dominate at large rand the boundary condition (Eq. (12.55))couldnotbesatisfied. As a second boundary condition, we may choose the conducting sphere and the plane θ=π/2 tobeatzeropotential,whichmeansthatEq. (12.54)nowbecomes V(r=r0)=b0 r0+parenleftbiggb1 r2 0−E0r0parenrightbigg P1(cosθ)+∞summationdisplay n=2bnPn(cosθ) rn+1 0=0.(12.57) Inorderthatthismayholdforallvaluesof θ,eachcoefficientof Pn(cosθ)mustvanish.12 Hence b0=0,13bn=0,n≥2, (12.58) whereas b1=E0r3 0. (12.59) Theelectrostaticpotential(outsidethesphere) isthen V=−E0rP1(cosθ)+E0r3 0 r2P1(cosθ)=−E0rP1(cosθ)parenleftbigg 1−r3 0 r3parenrightbigg .(12.60) InSection1.16itwasshownthatasolutionofLaplace’sequationthatsatisfiedthebound- aryconditionsovertheentireboundarywasunique.Theelectrostaticpotential V,asgi ven byEq.(12.60),isasolutionofLaplace’sequation.Itsatisfiesourboundaryconditionsand thereforeisthesolutionofLaplace’sequationfor thisproblem. 12Again,thisisequivalenttosayingthataseriesexpansioninLegendrepolynomials(oranycompleteorthogonalset)isunique. 13The coefficient of P0isb0/r0.W es e tb0=0 because there is no net charge on the sphere. If there is a net charge q,t h e n b0/negationslash=0. 12.3 Orthogonality 761 It may further be shown (Exercise 12.3.13) that there is an induced surface charge den- sity σ=−ε0∂V ∂rvextendsinglevextendsinglevextendsinglevextendsingle r=r0=3ε0E0cosθ (12.61) onthesurface ofthesphereandaninducedelectricdipolemoment(Exercise12.3.13) P=4πr3 0ε0E0. (12.62) /squaresolid Example 12.3.3 ELECTROSTATIC POTENTIAL OF A RING OF CHARGE As a further example, consider the electrostatic potential produced by a conducting ring carrying a total electric charge q(Fig. 12.9). From electrostatics (and Section 1.14) the potential ψsatisfiesLaplace’sequation.Separatingvariablesinsphericalpolarcoordinates (compareTable9.2), weobtain ψ(r,θ)=∞summationdisplay n=0cnan rn+1Pn(cosθ), r>a. (12.63a) Hereais the radius of the ring that is assumed to be in the θ=π/2 plane. There is no ϕ(azimuthal) dependence because of the cylindrical symmetry of the system. The terms with positive exponent in the radial dependence have been rejected because the potential musthaveanasymptoticbehavior, ψ∼q 4πε0·1 r,r≫a. (12.63b) The problem is to determine the coefficients cnin Eq. (12.63a). This may be done by evaluating ψ(r,θ)atθ=0,r=z, and comparing with an independent calculation of the FIGURE 12.9Charged, conductingring. 762 Chapter 12 Legendre Functions potential from Coulomb’s law. In effect, we are using a boundary condition along the z- axis.FromCoulomb’slaw(withallchargeequidistant), ψ(r,θ)=q 4πε0·1 (z2+a2)1/2,braceleftbiggθ=0 r=z, =q 4πε0z∞summationdisplay s=0(−1)s(2s)! 22s(s!)2parenleftbigga zparenrightbigg2s ,z>a. (12.63c) ThelaststepusestheresultofExercise8.1.15.Now,Eq.(12.63a)evaluatedat θ=0,r=z (withPn(1)=1),yields ψ(r,θ)=∞summationdisplay n=0cnan zn+1,r=z. (12.63d) ComparingEqs. (12.63c)and(12.63d), weget cn=0f o rnodd.Setting n=2s,weha v e c2s=q 4πε0(−1)s(2s)! 22s(s!)2, (12.63e) andourelectrostaticpotential ψ(r,θ)is givenby ψ(r,θ)=q 4πε0r∞summationdisplay s=0(−1)s(2s)! 22s(s!)2parenleftbigga rparenrightbigg2s P2s(cosθ), r>a. (12.63f) Themagneticanalogofthis problemappearsinExample12.5.3. /squaresolid Exercises 12.3.1 YouhaveconstructedasetoforthogonalfunctionsbytheGram–Schmidtprocess(Sec- tion 10.3), taking un(x)=xn,n=0,1,2,...,in increasing order with w(x)=1 and an interval−1≤x≤1. Prove that the nth such function constructed is proportional to Pn(x). Hint.Usemathematicalinduction. 12.3.2 Expand the Dirac delta function in a series of Legendre polynomials using the interval−1≤x≤1. 12.3.3 VerifytheDiracdeltafunctionexpansions δ(1−x)=∞summationdisplay n=02n+1 2Pn(x) δ(1+x)=∞summationdisplay n=0(−1)n2n+1 2Pn(x). These expressions appear in a resolution of the Rayleigh plane-wave expansion (Exer- cise12.4.7)intoincomingandoutgoingsphericalwaves. Note. Assume that the entireDirac delta function is covered when integrating over [−1,1]. 12.3 Orthogonality 763 12.3.4 Neutrons (mass 1) are being scattered by a nucleus of mass A( A>1). In the center- of-masssystemthescatteringisisotropic.Then,inthelaboratorysystemtheaverageof thecosineof theangleof deflectionof theneutronis /angbracketleftcosψ/angbracketright=1 2integraldisplayπ 0Acosθ+1 (A2+2Acosθ+1)1/2sinθdθ. Show,byexpansionofthedenominator,that /angbracketleftcosψ/angbracketright=2/3A. 12.3.5 A particular function f(x)defined over the interval [−1,1]is expanded in a Legendre seriesoverthissameinterval.Showthattheexpansionisunique. 12.3.6 Afunction f(x)is expandedinaLegendreseries f(x)=summationtext∞ n=0anPn(x). Showthat integraldisplay1 −1bracketleftbig f(x)bracketrightbig2dx=∞summationdisplay n=02a2 n 2n+1. ThisistheLegendreformoftheFourierseriesParsevalidentity,Exercise14.4.2.Italso illustratesBessel’sinequality,Eq. (10.72),becominganequalityfor acompleteset. 12.3.7 Derivetherecurrencerelation parenleftbig 1−x2parenrightbig P′ n(x)=nPn−1(x)−nxPn(x) fromtheLegendrepolynomialgeneratingfunction. 12.3.8 Evaluateintegraltext1 0Pn(x)dx. ANS.n=2s;1 fors=0,0f ors>0, n=2s+1;P2s(0)/(2s+2)=(−1)s(2s−1)!!/1(2s+2)!! Hint. Use a recurrence relation to replace Pn(x)by derivatives and then integrate by inspection.Alternatively,youcanintegratethegeneratingfunction. 12.3.9 (a) For f(x)=braceleftbigg+1,0<x<1 −1,−1<x<0, showthat integraldisplay1 −1bracketleftbig f(x)bracketrightbig2dx=2∞summationdisplay n=0(4n+3)bracketleftbigg(2n−1)!! (2n+2)!!bracketrightbigg2 . (b) Bytestingtheseries, provethattheseries is convergent. 12.3.10 Provethat integraldisplay1 −1xparenleftbig 1−x2parenrightbig P′ nP′ mdx=0,unlessm=n±1, =2n(n2−1) 4n2−1δm,n−1,ifm<n. =2n(n+2)(n+1) (2n+1)(2n+3)δm,n+1,ifm>n. 764 Chapter 12 Legendre Functions 12.3.11 Theamplitudeof ascatteredwaveis givenby f(θ)=1 k∞summationdisplay l=0(2l+1)exp[iδl]sinδlPl(cosθ). Hereθis the angle of scattering, lis the angular momentum eigenvalue, ¯hkis the incident momentum, and δlis the phase shift produced by the central potential that is doingthescattering.Thetotalcross sectionis σtot=integraltext |f(θ)|2d/Omega1.Showthat σtot=4π k2∞summationdisplay l=0(2l+1)sin2δl. 12.3.12 The coincidence counting rate, W(θ), in a gamma–gamma angular correlation experi- menthastheform W(θ)=∞summationdisplay n=0a2nP2n(cosθ). Show that data in the range π/2≤θ≤πcan, in principle, define the function W(θ) (and permit a determination of the coefficients a2n). This means that although data in therange 0≤θ<π /2 maybeusefulasacheck,theyarenotessential. 12.3.13 A conducting sphere of radius r0is placed in an initially uniform electric field, E0. Showthefollowing: (a) Theinducedsurface chargedensityis σ=3ε0E0cosθ. (b) Theinducedelectricdipolemomentis P=4πr3 0ε0E0. The induced electric dipole moment can be calculated either from the surface charge [part (a)] or by noting that the final electric field Eis the result of su- perimposinga dipolefieldontheoriginaluniformfield. 12.3.14 A charge qis displaced a distance aalong the z-axis from the center of a spherical cavityofradius R. (a) Showthattheelectricfieldaveragedoverthevolume a≤r≤Riszero. (b) Showthattheelectricfieldaveragedoverthevolume 0 ≤r≤ais E=ˆzEz=−ˆzq 4πε0a2(SI units)=−ˆznqa 3ε0, wherenisthenumberofsuchdisplacedchargesperunitvolume.Thisisabasiccalcu- lationinthepolarizationofa dielectric. Hint.E=−∇ϕ. 12.3 Orthogonality 765 FIGURE 12.10Charged, conductingdisk. 12.3.15 Determine the electrostatic potential (Legendre expansion) of a circular ring of electric chargefor r<a. 12.3.16 Calculate the electric field produced by the charged conductingring of Example 12.3.3 for (a)r>a,(b)r<a. 12.3.17 As an extension of Example 12.3.3, find the potential ψ(r,θ)produced by a charged conducting disk, Fig. 12.10, for r>a, the radius of the disk. The charge density σ(on eachsideof thedisk) is σ(ρ)=q 4πa(a2−ρ2)1/2,ρ2=x2+y2. Hint. The definite integral you get can be evaluated as a beta function, Section 8.4. For moredetailsseeSection5.03ofSmytheinAdditionalReadings. ANS.ψ(r,θ)=q 4πε0r∞summationdisplay l=0(−1)l1 2l+1parenleftbigga rparenrightbigg2l P2l(cosθ). 12.3.18 From the result of Exercise 12.3.17 calculate the potential of the disk. Since you are violatingthecondition r>a, justifyyourcalculation. Hint.Youmayrunintotheseries giveninExercise5.2.9. 12.3.19 The hemisphere defined by r=a,0≤θ<π /2, has an electrostatic potential +V0. The hemisphere r=a,π/2<θ≤πhas an electrostatic potential −V0. Show that the potentialatinteriorpointsis V=V0∞summationdisplay n=04n+3 2n+2parenleftbiggr aparenrightbigg2n+1 P2n(0)P2n+1(cosθ) =V0∞summationdisplay n=0(−1)n(4n+3)(2n−1)!! (2n+2)!!parenleftbiggr aparenrightbigg2n+1 P2n+1(cosθ). Hint.YouneedExercise12.3.8. 12.3.20 Aconductingsphereofradius aisdividedintotwoelectricallyseparatehemispheresby a thin insulating barrier at its equator. The top hemisphere is maintained at a potential V0,thebottomhemisphereat −V0. 766 Chapter 12 Legendre Functions (a) Showthattheelectrostaticpotential exteriortothetwohemispheresis V(r,θ)=V0∞summationdisplay s=0(−1)s(4s+3)(2s−1)!! (2s+2)!!parenleftbigga rparenrightbigg2s+2 P2s+1(cosθ). (b) Calculatetheelectricchargedensity σontheoutsidesurface.Notethatyourseries diverges at cos θ=±1, as you expect from the infinite capacitance of this system (zerothicknessfor theinsulatingbarrier). ANS.σ=ε0En=−ε0∂V ∂rvextendsinglevextendsinglevextendsinglevextendsingle r=a =ε0V0∞summationdisplay s=0(−1)s(4s+3)(2s−1)!! (2s)!!P2s+1(cosθ). 12.3.21 In the notation of Section 10.4, ϕs(x)=√(2s+1)/2Ps(x), a Legendre polynomial is renormalized to unity. Explain how |ϕs/angbracketright/angbracketleftϕs|acts as a projection operator. In particular, showthatif|f/angbracketright=summationtext na′ n|ϕn/angbracketright,then |ϕs/angbracketright/angbracketleftϕs|f/angbracketright=a′ s|ϕs/angbracketright. 12.3.22 Expandx8asaLegendreseries.DeterminetheLegendrecoefficientsfromEq.(12.50), am=2m+1 2integraldisplay1 −1x8Pm(x)dx. CheckyourvaluesagainstAMS-55,Table22.9.(Forthecompletereference,seeAddi- tional Readingsin Chapter 8). This illustrates the expansionof a simple function f(x). Actually if f(x)is expressed as a power series, the technique of Exercise 12.2.14 is bothfaster andmoreaccurate. Hint.Gaussianquadraturecanbeusedtoevaluatetheintegral. 12.3.23 Calculate and tabulate the electrostatic potential created by a ring of charge, Exam- ple12.3.3,for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry termsthrough P22(cosθ). Note. The convergence of your series will be slow for r/a=1.5. Truncating the series atP22limitsyoutoaboutfour-significant-figureaccuracy. Checkvalue. Forr/a=2.5 andθ=60◦,ψ=0.40272(q/4πε0r). 12.3.24 Calculate and tabulate the electrostatic potential created by a charged disk, Ex- ercise 12.3.17, for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry terms through P22(cosθ). Checkvalue. Forr/a=2.0 andθ=15◦,ψ=0.46638(q/4πε0r). 12.3.25 Calculate the first five (nonvanishing) coefficients in the Legendre series expansion of f(x)=1−|x|using Eq. (12.51)—numerical integration. Actually these coefficients can be obtained in closed form. Compare your coefficients with those obtained from Exercise13.3.28. ANS.a0=0.5000,a2=−0.6250,a4=0.1875,a6=−0.1016,a8=0.0664. 12.4 Alternate Definitions 767 12.3.26 Calculate and tabulate the exterior electrostatic potential created by the two charged hemispheres of Exercise 12.3.20, for r/a=1.5(0.5)5.0 andθ=0◦(15◦)90◦. Carry termsthrough P23(cosθ). Checkvalue. Forr/a=2.0 andθ=45◦,V=0.27066V0. 12.3.27 (a) Given f(x)=2.0,|x|<0.5;f(x)=0,0.5<|x|<1.0, expand f(x)in a Legen- dreseriesandcalculatethecoefficients anthrougha80(analytically). (b) Evaluatesummationtext80 n=0anPn(x)forx=0.400(0.005)0.600.Plotyourresults. Note.ThisillustratestheGibbsphenomenonofSection14.5andthedangeroftryingto calculatewithaseriesexpansioninthevicinityof adiscontinuity. 12.4 A LTERNATE DEFINITIONS OF LEGENDRE POLYNOMIALS Rodrigues’ Formula The series form of the Legendre polynomials (Eq. (12.8)) of Section 12.1 may be trans- formedasfollows.FromEq. (12.8), Pn(x)=[n/2]summationdisplay r=0(−1)r(2n−2r)! 2nr!(n−2r)!(n−r)!xn−2r. (12.64) Fornaninteger, Pn(x)=[n/2]summationdisplay r=0(−1)r1 2nr!(n−r)!parenleftbiggd dxparenrightbiggn x2n−2r =1 2nn!parenleftbiggd dxparenrightbiggnnsummationdisplay r=0(−1)rn! r!(n−r)!x2n−2r. (12.64a) Note the extension of the upper limit. The reader is asked to show in Exercise 12.4.1 that the additional terms [n/2]+1t onin the summation contribute nothing. However, the effectoftheseextratermsistopermitthereplacementofthenewsummationby (x2−1)n (binomialtheoremonceagain)toobtain Pn(x)=1 2nn!parenleftbiggd dxparenrightbiggnparenleftbig x2−1parenrightbign. (12.65) This is Rodrigues’ formula. It is useful in proving many of the properties of the Legendre polynomials, such as orthogonality. A related application is seen in Exercise 12.4.3. The Rodrigues definition is extended in Section 12.5 to define the associated Legendre func- tions.InSection12.7itis usedtoidentifytheorbitalangularmomentumeigenfunctions. 768 Chapter 12 Legendre Functions Schlaefli Integral Rodrigues’ formula provides a means of developing an integral representation of Pn(z). UsingCauchy’sintegralformula(Section6.4) f(z)=1 2πicontintegraldisplayf(t) t−zdt (12.66) with f(z)=parenleftbig z2−1parenrightbign, (12.67) wehave parenleftbig z2−1parenrightbign=1 2πicontintegraldisplay(t2−1)n t−zdt. (12.68) Differentiating ntimeswithrespectto zandmultiplyingby 1 /2nn!gives Pn(z)=1 2nn!dn dznparenleftbig z2−1parenrightbign=2−n 2πicontintegraldisplay(t2−1)n (t−z)n+1dt, (12.69) withthecontourenclosingthepoint t=z. This is the Schlaefli integral. Margenau and Murphy14use this to derive the recurrence relationsweobtainedfromthegeneratingfunction. The Schlaefli integral may readily be shown to satisfy Legendre’s equation by differen- tiationanddirectsubstitution(Fig.12.11). Weobtain parenleftbig 1−z2parenrightbigd2Pn dz2−2zdPn dz+n(n+1)Pn=n+1 2n2πicontintegraldisplayd dtbracketleftbigg(t2−1)n+1 (t−z)n+2bracketrightbigg dt.(12.70) Forintegral nourfunction (t2−1)n+1/(t−z)n+2issingle-valued,andtheintegralaround the closed path vanishes. The Schlaefli integral may also be used to define Pν(z)for non- integralνintegrating around the points t=z,t=1, but not crossing the cut line −1t o −∞. We could equally well encircle the points t=zandt=−1, but this would lead to FIGURE 12.11Schlaefliintegralcontour. 14H. Margenau and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed., Princeton, NJ: Van Nostrand (1956), Section3.5. 12.4 Alternate Definitions 769 nothing new. A contour about t=+1 andt=−1 will lead to a second solution, Qν(z), Section12.10. Exercises 12.4.1 Showthat eachterminthesummation nsummationdisplay r=[n/2]+1parenleftbiggd dxparenrightbiggn(−1)rn! r!(n−r)!x2n−2r vanishes( randnintegral). 12.4.2 UsingRodrigues’ formula,showthatthe Pn(x)areorthogonalandthat integraldisplay1 −1bracketleftbig Pn(x)bracketrightbig2dx=2 2n+1. Hint.UseRodrigues’formulaandintegratebyparts. 12.4.3 Showthatintegraltext1 −1xmPn(x)dx=0 whenm<n. Hint.UseRodrigues’formulaor expand xminLegendrepolynomials. 12.4.4 Showthat integraldisplay1 −1xnPn(x)dx=2n+1n!n! (2n+1)!. Note.YouareexpectedtouseRodrigues’formulaandintegratebyparts,butalsoseeif youcangettheresult fromEq. (12.8)byinspection. 12.4.5 Showthat integraldisplay1 −1x2rP2n(x)dx=22n+1(2r)!(r+n!) (2r+2n+1)!(r−n)!,r≥n. 12.4.6 As a generalization of Exercises 12.4.4 and 12.4.5, show that the Legendre expansions ofxsare (a)x2r=rsummationdisplay n=022n(4n+1)(2r)!(r+n)! (2r+2n+1)!(r−n)!P2n(x), s=2r, (b)x2r+1=rsummationdisplay n=022n+1(4n+3)(2r+1)!(r+n+1)! (2r+2n+3)!(r−n)!P2n+1(x), s=2r+1. 12.4.7 AplanewavemaybeexpandedinaseriesofsphericalwavesbytheRayleighequation, eikrcosγ=∞summationdisplay n=0anjn(kr)Pn(cosγ). Showthat an=in(2n+1). 770 Chapter 12 Legendre Functions Hint. 1. Use theorthogonalityof the Pntosolvefor anjn(kr). 2. Differentiate ntimes with respect to (kr)and setr=0 to eliminate the r- dependence. 3. EvaluatetheremainingintegralbyExercise12.4.4. Note.Thisproblemmayalsobetreatedbynotingthatbothsidesoftheequationsatisfy the Heemholtz equation. The equality can be established by showing that the solutions have the same behavior at the origin and also behave alike at large distances. A “by inspection”typeofsolutionis developedinSection9.7usingGreen’sfunctions. 12.4.8 VerifytheRayleighequationofExercise12.4.7bystartingwiththefollowingsteps: (a) Differentiatewithrespectto (kr)toestablish summationdisplay nanj′ n(kr)Pn(cosγ)=isummationdisplay nanjn(kr)cosγPn(cosγ). (b) Use a recurrence relation to replace cos γPn(cosγ)by a linear combination of Pn−1andPn+1. (c) Usearecurrencerelationtoreplace j′ nbyalinearcombinationof jn−1andjn+1. 12.4.9 From Exercise12.4.7showthat jn(kr)=1 2inintegraldisplay1 −1eikrµPn(µ)dµ. This means that (apart from a constant factor) the spherical Bessel function jn(kr)is theFouriertransformof theLegendrepolynomial Pn(µ). 12.4.10 TheLegendrepolynomialsandthesphericalBesselfunctionsarerelatedby jn(z)=1 2(−i)nintegraldisplayπ 0eizcosθPn(cosθ)sinθdθ, n=0,1,2,.... Verifythis relationbytransformingtheright-handsideinto zn 2n+1n!integraldisplayπ 0cos(zcosθ)sin2n+1θdθ andusingExercise11.7.8. 12.4.11 Bydirectevaluationof theSchlaefliintegralshowthat Pn(1)=1. 12.4.12 Explain why the contour of the Schlaefli integral, Eq. (12.69), is chosen to enclose the pointst=zandt=1 whenn→ν, notaninteger. 12.4.13 Innumericalwork(forexample,theGauss–Legendrequadrature)itisusefultoestablish thatPn(x)hasnrealzerosintheinteriorof [−1,1].Showthatthisis so. Hint. Rolle’s theorem shows that the first derivative of (x2−1)2nhas one zero in the interior of[−1,1]. Extend this argument to the second, third, and ultimately the nth derivative. 12.5 Associated Legendre Functions 771 12.5 A SSOCIATED LEGENDRE FUNCTIONS When Helmholtz’s equation is separated in spherical polar coordinates (Section 9.3), one oftheseparatedODEsis theassociatedLegendreequation 1 sinθd dθparenleftbigg sinθdPm n(cosθ) dθparenrightbigg +bracketleftbigg n(n+1)−m2 sin2θbracketrightbigg Pm n(cosθ)=0.(12.71) Withx=cosθ,this becomes parenleftbig 1−x2parenrightbigd2 dx2Pm n(x)−2xd dxPm n(x)+bracketleftbigg n(n+1)−m2 1−x2bracketrightbigg Pm n(x)=0.(12.72) If the azimuthal separation constant m2=0, we have Legendre’s equation, Eq. (12.28). Theregularsolutions Pm n(x)(withmnotnecessarilyzero,butaninteger)are v≡Pm n(x)=parenleftbig 1−x2parenrightbigm/2dm dxmPn(x) (12.73a) withm≥0 aninteger. One way of developing the solution of the associated Legendre equation is to start with theregularLegendreequationandconvertitintotheassociatedLegendreequationbyusing multiple differentiation. These multiple differentiations are suggested by Eq. (12.73a), the generation of associated Legendre polynomials, and spherical harmonics of Section 12.6 moregenerally,inSection4.3usingraisingorloweringoperatorsofEq.(4.69)repeatedly. Fortheirderivativeform seeExercise12.6.8. WetakeLegendre’sequation parenleftbig 1−x2parenrightbig P′′ n−2xP′ n+n(n+1)Pn=0, (12.74) andwiththehelpofLeibniz’formula15differentiate mtimes.Theresultis parenleftbig 1−x2parenrightbig u′′−2x(m+1)u′+(n−m)(n+m+1)u=0, (12.75) where u≡dm dxmPn(x). (12.76) Equation(12.74)isnotself-adjoint.Toputitintoself-adjointformandconverttheweight- ingfunctionto1, wereplace u(x)by v(x)=parenleftbig 1−x2parenrightbigm/2u(x)=parenleftbig 1−x2parenrightbigm/2dmPn(x) dxm. (12.73b) 15Leibniz’ formula for the nth derivative of aproduct is dn dxnbracketleftbig A(x)B(x)bracketrightbig =nsummationdisplay s=0parenleftbiggn sparenrightbiggdn−s dxn−sA(x)ds dxsB(x),parenleftbiggn sparenrightbigg =n! (n−s)!s!, abinomial coefficient. 772 Chapter 12 Legendre Functions Solvingfor uanddifferentiating,weobtain u′=parenleftbigg v′+mxv 1−x2parenrightbiggparenleftbig 1−x2parenrightbig−m/2, (12.77) u′′=bracketleftbigg v′′+2mxv′ 1−x2+mv 1−x2+m(m+2)x2v (1−x2)2bracketrightbigg ·parenleftbig 1−x2parenrightbig−m/2.(12.78) Substituting into Eq. (12.74), we find that the new function vsatisfies the self-adjoint ODE parenleftbig 1−x2parenrightbig v′′−2xv′+bracketleftbigg n(n+1)−m2 1−x2bracketrightbigg v=0, (12.79) whichistheassociatedLegendreequation;itreducestoLegendre’sequationwhen misset equal to zero. Expressed in spherical polar coordinates, the associated Legendre equation is 1 sinθd dθparenleftbigg sinθdv dθparenrightbigg +bracketleftbigg n(n+1)−m2 sin2θbracketrightbigg v=0. (12.80) Associated Legendre Polynomials Theregularsolutions,relabeled Pm n(x),ar e v≡Pm n(x)=parenleftbig 1−x2parenrightbigm/2dm dxmPn(x). (12.73c) These are the associated Legendre functions.16Since the highest power of xinPn(x)is xn,w em u s th a v e m≤n(or them-fold differentiation will drive our function to zero). In quantum mechanics the requirement that m≤nhas the physical interpretation that the expectation value of the square of the zcomponent of the angular momentum is less than orequaltotheexpectationvalueofthesquareoftheangularmomentumvector L, angbracketleftbig L2 zangbracketrightbig ≤angbracketleftbig L2angbracketrightbig ≡integraldisplay ψ∗ lmL2ψlmd3r. FromtheformofEq.(12.73c)wemightexpect mtobenonnegative.However,if Pn(x) isexpressedbyRodrigues’formula,thislimitationon misrelaxedandwemayhave −n≤ m≤n,negativeaswellaspositivevaluesof mbeingpermitted.Theselimitsareconsistent withthoseobtainedbymeansofraisingandloweringoperatorsinChapter4.Inparticular, |m|>nis ruled out. This also follows from Eq. (12.73c). Using Leibniz’ differentiation formulaonceagain,wecanshow(Exercise12.5.1)that Pm n(x)andP−m n(x)arerelatedby P−m n(x)=(−1)m(n−m)! (n+m)!Pm n(x). (12.81) 16Occasionally (as in AMS-55; for the complete reference, see the Additional Readings of Chapter 8), one finds the associated Legendre functions defined with an additional factor of (−1)m.T h i s(−1)mseems an unnecessary complication at this point. It will be included in the definition of the spherical harmonics Ymn(θ,ϕ)in Section 12.6. Our definition agrees with Jackson’s Electrodynamics (seeAdditionalReadingsofChapter11forthisreference).Notealsothattheupperindex misnotanexponent. 12.5 Associated Legendre Functions 773 Fromour definitionoftheassociatedLegendrefunctions Pm n(x), P0 n(x)=Pn(x). (12.82) AgeneratingfunctionfortheassociatedLegendrefunctionsisobtained,viaEq.(12.71), fromthatoftheordinaryLegendrepolynomials: (2m)!(1−x2)m/2 2mm!(1−2tx+t2)m+1/2=∞summationdisplay s=0Pm s+m(x)ts. (12.83) If we drop the factor (1−x2)m/2=sinmθfrom this formula and define the polynomi- alsPm s+m(x)=Pm s+m(x)(1−x2)−m/2,then we obtain a practical form of the generating function, gm(x,t)≡(2m)! 2mm!(1−2tx+t2)m+1/2=∞summationdisplay s=0Pm s+m(x)ts. (12.84) WecanderivearecursionrelationforassociatedLegendrepolynomialsthatisanalogous toEqs. (12.14)and(12.17)bydifferentiationasfollows: parenleftbig 1−2tx+t2parenrightbig∂gm ∂t=(2m+1)(x−t)gm(x,t). Substitutingthedefiningexpansionsfor associatedLegendrepolynomialsweget parenleftbig 1−2tx+t2parenrightbigsummationdisplay ssPm s+m(x)ts−1=(2m+1)summationdisplay sbracketleftbig xPm s+mts−Pm s+mts+1bracketrightbig . Comparing coefficients of powers of tin these power series, we obtain the recurrence relation (s+1)Pm s+m+1−(2m+1+2s)xPm s+m+(s+2m)Pm s+m−1=0.(12.85) Form=0 ands=nthisrelationis Eq.(12.17). Before we can use this relation we need to initialize it, that is, relate the associated Legendre polynomials to ordinary Legendre polynomials. We can use Pm m=(2m−1)!! fromEq.(12.73c).Also,since |m|≤n,wemayset Pn+1 n=0andusethistoobtainstarting valuesfor variousrecursiveprocesses.We observethat parenleftbig 1−2xt+t2parenrightbig g1(x,t)=parenleftbig 1−2xt+t2parenrightbig−1/2=summationdisplay sPs(x)ts,(12.86) souponinsertingEq. (12.84)wegettherecursion P1 s+1−2xP1 s+P1 s−1=Ps(x). (12.87) Moregenerally,wealsohavetheidentity parenleftbig 1−2xt+t2parenrightbig gm+1(x,t)=(2m+1)gm(x,t), (12.88) fromwhichweextracttherecursion Pm+1 s+m+1−2xPm+1 s+m+Pm+1 s+m−1=(2m+1)Pm s+m(x), (12.89) whichrelatestheassociatedLegendrepolynomialswithsuperindex m+1tothosewith m. Form=0 werecovertheinitialrecursionEq. (12.87). 774 Chapter 12 Legendre Functions Table 12.2 AssociatedLegendreFunctions P1 1(x)=(1−x2)1/2=sinθ P1 2(x)=3x(1−x2)1/2=3cosθsinθ P2 2(x)=3(1−x2)=3sin2θ P1 3(x)=3 2(5x2−1)(1−x2)1/2=3 2(5cos2θ−1)sinθ P2 3(x)=15x(1−x2)=15cosθsin2θ P3 3(x)=15(1−x2)3/2=15sin3θ P1 4(x)=5 2(7x3−3x)(1−x2)1/2=5 2(7cos3θ−3cosθ)sinθ P2 4(x)=15 2(7x2−1)(1−x2)=15 2(7cos2θ−1)sin2θ P3 4(x)=105x(1−x2)3/2=105cosθsin3θ P4 4(x)=105(1−x2)2=105sin4θ Example 12.5.1 LOWEST ASSOCIATED LEGENDRE POLYNOMIALS Now we are ready to derive the entries of Table 12.2. For m=1 ands=0, Eq. (12.87) yields P1 1=1, because P1 0=0=P1 −1do not occur in the definition, Eq. (12.84), of the associated Legendre polynomials. Multiplying by (1−x2)1/2=sinθwe get the first line ofTable12.2. For s=1 wefind,from Eq.(12.87), P1 2(x)=P1+2xP1 1=x+2x=3x, from which the second line of Table 12.2, 3cos θsinθ, follows upon multiplying by sin θ. Fors=2 weget P1 3(x)=P2+2xP1 2−P1 1=1 2parenleftbig 3x2−1parenrightbig +6x2−1=15 2x2−3 2, inagreementwithline4ofTable12.2.Togetline3weuseEq.(12.88).For m=1,s=0, this gives P2 2(x)=3P1 1(x)=3, and multiplying by 1 −x2=sin2θreproduces line 3 of Table12.2.Forlines5,8,9,Eq.(12.84)maybeused,whichweleaveasanexercise.More generally, we use Eq. (12.89) instead of Eq. (12.87) to get a starting value of Pm m. Then Eq. (12.85) reduces to a two-term formula for Pm m,g i v i n g(2m−1)!!. Note that, if m=0, thisis(−1)!!=1. /squaresolid Example 12.5.2 SPECIAL VALUES Forx=1w eu s e parenleftbig 1−2t+t2parenrightbig−m−1/2=(1−t)−2m−1=∞summationdisplay s=0parenleftbigg−2m−1 sparenrightbigg ts inEq.(12.84) andfind Pm s+m(1)=(2m)! 2mm!parenleftbigg−2m−1 sparenrightbigg , (12.90) 12.5 Associated Legendre Functions 775 whereparenleftbigg−m sparenrightbigg =1 fors=0 andparenleftbigg−m sparenrightbigg =(−m)(−m−1)···(1−s−m) s! fors≥1. Form=1,s=0w eh a v e P1 1(1)=parenleftbig−3 0parenrightbig =1; fors=1,P1 2(1)=−parenleftbig−3 1parenrightbig =3; fors=2,P1 3(1)=parenleftbig−3 2parenrightbig =(−3)(−4) 2=6=3 2(5−1), which all agree with Table 12.2. For x=0 wecanalsousethebinomialexpansion,whichweleaveasanexercise. /squaresolid Recurrence Relations As expected and already seen, the associated Legendre functions satisfy recurrence rela- tions. Because of the existence of two indices instead of just one, we have a wide variety ofrecurrencerelations: Pm+1 n−2mx (1−x2)1/2Pm n+bracketleftbig n(n+1)−m(m−1)bracketrightbig Pm−1 n=0, (12.91) (2n+1)xPm n=(n+m)Pm n−1+(n−m+1)Pm n+1, (12.92) (2n+1)parenleftbig 1−x2parenrightbig1/2Pm n=Pm+1 n+1−Pm+1 n−1 =(n+m)(n+m−1)Pm−1 n−1 −(n−m+1)(n−m+2)Pm−1 n+1,(12.93) parenleftbig 1−x2parenrightbig1/2Pm′ n=1 2Pm+1 n−1 2(n+m)(n−m+1)Pm−1 n.(12.94) These relations, and many other similar ones, may be verified by use of the generat- ing function (Eq. (12.4)), by substitution of the series solution of the associated Legen- dre equation (12.79) or reduction to the Legendre polynomial recurrence relations, us- ing Eq. (12.73c). As an example of the last method, consider Eq. (12.93). It is similar to Eq.(12.23): (2n+1)Pn(x)=P′ n+1(x)−P′ n−1(x). (12.95) LetusdifferentiatethisLegendrepolynomialrecurrencerelation mtimestoobtain (2n+1)dm dxmPn(x)=dm dxmP′ n+1(x)−dm dxmP′ n−1(x) =dm+1 dxm+1Pn+1(x)−dm+1 dxm+1Pn−1(x). (12.96) Now multiplying by (1−x2)(m+1)/2and using the definition of Pn(x), we obtain the first partofEq. (12.93). 776 Chapter 12 Legendre Functions Parity The parity relation satisfied by the associated Legendre functions may be determined by examination of the defining equation (12.73c). As x→−x, we already know that Pn(x) contributesa (−1)n.T hem-fold differentiationyieldsafactor of (−1)m.Hencewehave Pm n(−x)=(−1)n+mPm n(x). (12.97) AglanceatTable12.2verifiesthis for 1 ≤m≤n≤4. Also,from thedefinitioninEq. (12.73c), Pm n(±1)=0,form/negationslash=0. (12.98) Orthogonality Theorthogonalityofthe Pm n(x)followsfromtheODE,justasforthe Pn(x)(Section12.3), ifmis the same for both functions. However, it is instructive to demonstrate the orthogo- nalitybyanothermethod,amethodthatwillalsoprovidethenormalizationconstant. UsingthedefinitioninEq.(12.73c)andRodrigues’formula(Eq.(12.65))for Pn(x),we find integraldisplay1 −1Pm p(x)Pm q(x)dx=(−1)m 2p+qp!q!integraldisplay1 −1Xmparenleftbiggdp+m dxp+mXpparenrightbiggdq+m dxq+mXqdx.(12.99) Thefunction Xisgivenby X≡(x2−1).Ifp/negationslash=q,letusassumethat p<q.Noticethatthe superscript mis the same for both functions. This is an essential condition. The technique is to integrate repeatedly by parts; all the integrated parts will vanish as long as there is a factorX=x2−1.Letusintegrate q+mtimestoobtain integraldisplay1 −1Pm p(x)Pm q(x)dx=(−1)m(−1)q+m 2p+qp!q!integraldisplay1 −1Xqdq+m dxq+mparenleftbigg Xmdp+m dxp+mXpparenrightbigg dx.(12.100) Theintegrandontheright-handsideis nowexpandedbyLeibniz’formulatogive Xqdq+m dxq+mparenleftbigg Xmdp+m dxp+mXpparenrightbigg =Xqq+msummationdisplay i=0(q+m)! i!(q+m−i)!parenleftbiggdq+m−i dxq+m−iXmparenrightbiggdp+m+i dxp+m+iXp.(12.101) Sincetheterm Xmcontainsnopowerof xgreaterthan x2m,wemustha v e q+m−i≤2m (12.102) orthederivativewillvanish.Similarly, p+m+i≤2p. (12.103) Addingbothinequalitiesyields q≤p, (12.104) 12.5 Associated Legendre Functions 777 which contradicts our assumption that p<q. Hence, there is no solution for iand the integralvanishes.Thesameresult obviouslywillfollowif p>q. For the remaining case, p=q, we have the single term corresponding to i=q−m. PuttingEq.(12.101)intoEq. (12.100),wehave integraldisplay1 −1bracketleftbig Pm q(x)bracketrightbig2dx=(−1)q+2m(q+m)! 22qq!q!(2m)!(q−m)!integraldisplay1 −1Xqparenleftbiggd2m dx2mXmparenrightbiggparenleftbiggd2q dx2qXqparenrightbigg dx. (12.105) Since Xm=parenleftbig x2−1parenrightbigm=x2m−mx2m−2+···, (12.106) d2m dx2mXm=(2m)!, (12.107) Eq.(12.105)reducesto integraldisplay1 −1bracketleftbig Pm q(x)bracketrightbig2dx=(−1)q+2m(2q)!(q+m)! 22qq!q!(q−m)!integraldisplay1 −1Xqdx. (12.108) Theintegralontherightis just (−1)qintegraldisplayπ 0sin2q+1θdθ=(−1)q22q+1q!q! (2q+1)!(12.109) (compare Exercise 8.4.9). Combining Eqs. (12.108) and (12.109), we have the orthogo- nalityintegral , integraldisplay1 −1Pm p(x)Pm q(x)dx=2 2q+1·(q+m)! (q−m)!δpq, (12.110) or,insphericalpolarcoordinates, integraldisplayπ 0Pm p(cosθ)Pm q(cosθ)sinθdθ=2 2q+1·(q+m)! (q−m)!δpq. (12.111) The orthogonality of the Legendre polynomials is a special case of this result, obtained by setting mequal to zero; that is, for m=0,Eq. (12.110) reduces to Eqs. (12.47) and (12.48). In both Eqs. (12.110) and (12.111), our Sturm–Liouville theory of Chapter 10 could provide the Kronecker delta. A special calculation, such as the analysis here, is re- quiredfor thenormalizationconstant. The orthogonality of the associated Legendre functions over the same interval and with the same weighting factor as the Legendre polynomials does not contradict the unique- ness of the Gram–Schmidt construction of the Legendre polynomials, Example 10.3.1. Table12.2suggests(andSection12.4verifies)thatintegraltext1 −1Pm p(x)Pm q(x)dxmaybewrittenas integraldisplay1 −1Pm p(x)Pm q(x)parenleftbig 1−x2parenrightbigmdx, wherewedefinedearlier Pm p(x)parenleftbig 1−x2parenrightbigm/2=Pm p(x). 778 Chapter 12 Legendre Functions Thefunctions Pm p(x)maybeconstructedbytheGram–Schmidtprocedurewiththeweight- ingfunction w(x)=(1−x2)m. It is possible to develop an orthogonality relation for associated Legendre functions of thesamelowerindexbutdifferentupperindex.Wefind integraldisplay1 −1Pm n(x)Pk n(x)parenleftbig 1−x2parenrightbig−1dx=(n+m)! m(n−m!)δm,k. (12.112) Notethatanewweightingfactor, (1−x2)−1,hasbeenintroduced.Thisrelationisamath- ematicalcuriosity.InphysicalproblemswithsphericalsymmetrysolutionsofEqs.(12.80) and (9.64) appear in conjunction with those of Eq. (9.61), and orthogonality of the az- imuthaldependencemakesthetwoupperindicesequalandalwaysleadstoEq.(12.111). Example 12.5.3 MAGNETIC INDUCTION FIELD OF A CURRENT LOOP Like the other ODEs of mathematical physics, the associated Legendre equation is likely to pop up quite unexpectedly. As an illustration, consider the magnetic induction field B and magnetic vector potential Acreated by a single circular current loop in the equatorial plane(Fig.12.12). We know from electromagnetic theory that the contribution of current element Idλto themagneticvectorpotentialis dA=µ0 4πIdλ r. (12.113) (ThisfollowsfromExercise1.14.4andSection9.7)Equation(12.113),plusthesymmetry ofoursystem,showsthat Ahasonlyaˆϕcomponentandthatthecomponentisindependent FIGURE 12.12Circularcurrent loop. 12.5 Associated Legendre Functions 779 ofϕ,17 A=ˆϕAϕ(r,θ). (12.114) ByMaxwell’sequations, ∇×H=J,∂D ∂t=0(SI units). (12.115) Since µ0H=B=∇×A, (12.116) wehave ∇×(∇×A)=µ0J, (12.117) whereJis the current density. In our problem Jis zero everywhere except in the current loop.Therefore,awayfrom theloop, ∇×∇׈ϕAϕ(r,θ)=0, (12.118) usingEq. (12.114). From the expression for the curl in spherical polar coordinates (Section 2.5), we obtain (Example2.5.2) ∇×bracketleftbig ∇׈ϕAϕ(r,θ)bracketrightbig =ˆϕbracketleftbigg −∂2Aϕ ∂r2−2 r∂Aϕ ∂r−1 r2∂2Aϕ ∂θ2−1 r2∂ ∂θ(cotθAϕ)bracketrightbigg =0. (12.119) LettingAϕ(r,θ)=R(r)/Theta1(θ) andseparatingvariables,wehave r2d2R dr2+2dR dr−n(n+1)R=0, (12.120) d2/Theta1 dθ2+cotθd/Theta1 dθ+n(n+1)/Theta1−/Theta1 sin2θ=0. (12.121) Thesecondequationis theassociatedLegendreequation(12.80) with m=1,andwemay immediatelywrite /Theta1(θ)=P1 n(cosθ). (12.122) Theseparationconstant n(n+1),nanonnegativeinteger,waschosentokeepthissolution wellbehaved. By trial, letting R(r)=rα, we find that α=n,o r−n−1. The first possibility is dis- carded,for oursolutionmustvanishas r→∞. Hence Aϕn=bn rn+1P1 n(cosθ)=cnparenleftbigga rparenrightbiggn+1 P1 n(cosθ) (12.123) 17Pair off corresponding current elements Idλ(ϕ1)andIdλ(ϕ2),wh er eϕ−ϕ1=ϕ2−ϕ. 780 Chapter 12 Legendre Functions and Aϕ(r,θ)=∞summationdisplay n=1cnparenleftbigga rparenrightbiggn+1 P1 n(cosθ) (r >a). (12.124) Hereaistheradiusofthecurrentloop. SinceAϕmust be invarianttoreflectionin the equatorialplane,by the symmetryof our problem, Aϕ(r,cosθ)=Aϕ(r,−cosθ), (12.125) theparitypropertyof Pm n(cosθ)(Eq. (12.97)) showsthat cn=0f o rneven. To complete the evaluation of the constants, we may use Eq. (12.124) to calculate Bz along the z-axis(Bz=Br(r,θ=0))and compare with the expression obtained from the Biot–Savartlaw.ThisisthesametechniqueasusedinExample12.3.3.Wehave(compare Eq.(2.47)) Br=∇×Avextendsinglevextendsingle r=1 rsinθbracketleftbigg∂ ∂θ(sinθAϕ)bracketrightbigg =cotθ rAϕ+1 r∂Aϕ ∂θ. (12.126) Using ∂P1 n(cosθ) ∂θ=−sinθdP1 n(cosθ) d(cosθ)=−1 2P2 n+n(n+1) 2P0 n (12.127) (Eq. (12.94))andthenEq. (12.91)with m=1, P2 n(cosθ)−2cosθ sinθP1 n(cosθ)+n(n+1)Pn(cosθ)=0, (12.128) weobtain Br(r,θ)=∞summationdisplay n=1cnn(n+1)an+1 rn+2Pn(cosθ), r>a, (12.129) (for allθ). Inparticular,for θ=0, Br(r,0)=∞summationdisplay n=1cnn(n+1)an+1 rn+2. (12.130) Wemayalso obtain Bθ(r,θ)=−1 r∂(rAϕ) ∂r=∞summationdisplay n=1cnnan+1 rn+2P1 n(cosθ), r>a, (12.131) TheBiot–Savartlawstatesthat dB=µ0 4πIdλ׈r r2(SI units). (12.132) We now integrate over the perimeter of our loop (radius a). The geometry is shown in Fig.12.13.Theresultingmagneticinductionfieldis ˆzBz, alongthe z-axis, with Bz=µ0I 2a2parenleftbig a2+z2parenrightbig−3/2=µ0I 2a2 z3parenleftbigg 1+a2 z2parenrightbigg−3/2 . (12.133) 12.5 Associated Legendre Functions 781 FIGURE 12.13Biot–Savartlawappliedtoacircularloop. Expandingbythebinomialtheorem,weobtain Bz=µ0I 2a2 z3bracketleftbigg 1−3 2parenleftbigga zparenrightbigg2 +15 8parenleftbigga zparenrightbigg4 −···bracketrightbigg =µ0I 2a2 z3∞summationdisplay s=0(−1)s(2s+1)!! (2s)!!parenleftbigga zparenrightbigg2s ,z>a. (12.134) EquatingEqs. (12.130)and(12.134)termbyterm(with r=z),18wefind c1=µ0I 4,c 3=−µ0I 16,c 2=c4=···=0. cn=(−1)(n−1)/2µ0I 2n(n+1)·(n/2)! [(n−1)/2]!(1 2)!,nodd.(12.135) Equivalently,wemaywrite c2n+1=(−1)nµ0I 22n+2·(2n)! n!(n+1)!=(−1)nµ0I 2·(2n−1)!! (2n+2)!!(12.136) 18Thedescending powerseries is alsounique. 782 Chapter 12 Legendre Functions and Aϕ(r,θ)=parenleftbigga rparenrightbigg2∞summationdisplay n=0c2n+1parenleftbigga rparenrightbigg2n P1 2n+1(cosθ), (12.137) Br(r,θ)=a2 r3∞summationdisplay n=0c2n+1(2n+1)(2n+2)parenleftbigga rparenrightbigg2n P2n+1(cosθ),(12.138) Bθ(r,θ)=a2 r3∞summationdisplay n=0c2n+1(2n+1)parenleftbigga rparenrightbigg2n P1 2n+1(cosθ). (12.139) These fields may be described in closed form by the use of elliptic integrals. Exer- cise 5.8.4 is an illustration of this approach. A third possibility is direct integration of Eq. (12.113) by expanding the denominator of the integral for Aϕin Exercise 5.8.4 as a Legendre polynomial generating function. The current is specified by Dirac delta func- tions.Thesemethodshavetheadvantageof yieldingtheconstants cndirectly. Acomparisonofmagneticcurrentloopdipolefieldsandfiniteelectricdipolefieldsmay beofinterest.Forthemagneticcurrentloopdipole,theprecedinganalysisgives Br(r,θ)=µ0I 2a2 r3bracketleftbigg P1−3 2parenleftbigga rparenrightbigg2 P3+···bracketrightbigg , (12.140) Bθ(r,θ)=µ0I 4a2 r3bracketleftbigg P1 1−3 4parenleftbigga rparenrightbigg2 P1 3+···bracketrightbigg . (12.141) Fromthefiniteelectricdipolepotentialof Section12.1wehave Er(r,θ)=qa πε0r3bracketleftbigg P1+2parenleftbigga rparenrightbigg2 P3+···bracketrightbigg , (12.142) Eθ(r,θ)=qa 2πε0r3bracketleftbigg P1 1+parenleftbigga rparenrightbigg2 P1 3+···bracketrightbigg . (12.143) The two fields agree in form as far as the leading term is concerned (r−3P1), and this is thebasis forcallingthembothdipolefields. As with electric multipoles, it is sometimes convenient to discuss pointmagnetic mul- tipoles (see Fig. 12.14). For the dipole case, Eqs. (12.140) and (12.141), the point dipole is formed by taking the limit a→0,I→∞, withIa2held constant. With na unit vector normal to the current loop (positive sense by right-hand rule, Section 1.10), the magnetic momentmisgivenby m=nIπa2. /squaresolid FIGURE 12.14Electricdipole. 12.5 Associated Legendre Functions 783 Exercises 12.5.1 Provethat P−m n(x)=(−1)m(n−m)! (n+m)!Pm n(x), wherePm n(x)is definedby Pm n(x)=1 2nn!parenleftbig 1−x2parenrightbigm/2dn+m dxn+mparenleftbig x2−1parenrightbign. Hint.Oneapproachis toapplyLeibniz’formulato (x+1)n(x−1)n. 12.5.2 Showthat P1 2n(0)=0, P1 2n+1(0)=(−1)n(2n+1)! (2nn!)2=(−1)n(2n+1)!! (2n)!!, byeachofthesethreemethods: (a) useofrecurrencerelations, (b) expansionof thegeneratingfunction, (c) Rodrigues’formula. 12.5.3 Evaluate Pm n(0). ANS.Pm n(0)=  (−1)(n−m)/2 (n+m)! 2n((n−m)/2)!((n+m)/2!),n+meven, 0,n +modd. Also,Pm n(0)=(−1)(n−m)/2(n+m−1)!! (n−m)!!,n+meven. 12.5.4 Showthat Pn n(cosθ)=(2n−1)!!sinnθ, n=0,1,2,.... 12.5.5 DerivetheassociatedLegendrerecurrencerelation Pm+1 n(x)−2mx (1−x2)1/2Pm n(x)+bracketleftbig n(n+1)−m(m−1)bracketrightbig Pm−1 n(x)=0. 12.5.6 Developarecurrencerelationthatwillyield P1 n(x)as P1 n(x)=f1(x,n)P n(x)+f2(x,n)P n−1(x). Followeither(a) or(b). (a) Derive a recurrence relation of the preceding form. Give f1(x,n)andf2(x,n) explicitly. (b) Findtheappropriaterecurrencerelationinprint. (1) Givethesource. 784 Chapter 12 Legendre Functions (2) Verifytherecurrencerelation. ANS.P1 n(x)=−nx (1−x2)1/2Pn+n (1−x2)1/2Pn−1. 12.5.7 Showthat sinθd dcosθPn(cosθ)=P1 n(cosθ). 12.5.8 Showthat (a)integraldisplayπ 0parenleftbiggdPm n dθdPm n′ dθ+m2Pm nPm n′ sin2θparenrightbigg sinθdθ=2n(n+1) 2n+1(n+m)! (n−m)!δnn′, (b)integraldisplayπ 0parenleftbiggP1 n sinθdP1 n′ dθ+P1 n′ sinθdP1 n dθparenrightbigg sinθdθ=0. Theseintegralsoccurinthetheoryofscatteringofelectromagneticwavesbyspheres. 12.5.9 AsarepeatofExercise12.3.10,show,usingassociatedLegendrefunctions,that integraldisplay1 −1xparenleftbig 1−x2parenrightbig P′ n(x)P′ m(x)dx=n+1 2n+1·2 2n−1·n! (n−2)!δm,n−1 +n 2n+1·2 2n+3·(n+2)! n!δm,n+1. 12.5.10 Evaluate integraldisplayπ 0sin2θP1 n(cosθ)dθ. 12.5.11 TheassociatedLegendrepolynomial Pm n(x)satisfiestheself-adjointODE parenleftbig 1−x2parenrightbigd2Pm n(x) dx2−2xdPm n(x) dx+bracketleftbigg n(n+1)−m2 1−x2bracketrightbigg Pm n(x)=0. Fromthedifferentialequationsfor Pm n(x)andPk n(x)showthat integraldisplay1 −1Pm n(x)Pk n(x)dx 1−x2=0, fork/negationslash=m. 12.5.12 Determinethevectorpotentialofamagneticquadrupolebydifferentiatingthemagnetic dipolepotential. ANS.AMQ=µ0 2parenleftbig Ia2parenrightbig (dz)ˆϕP1 2(cosθ) r3+higher-orderterms . BMQ=µ0parenleftbig Ia2parenrightbig (dz)bracketleftbigg ˆr3P2(cosθ) r4+ˆθP1 2(cosθ) r4bracketrightbigg . 12.5 Associated Legendre Functions 785 This corresponds to placing a current loop of radius aatz→dzand an oppositely directed current loop at z→−dzand letting a→0 subject to (dz)a(dipole strength) equalconstant. Anotherapproachtothisproblemwouldbetointegrate dA(Eq.(12.113),toexpandthe denominator in a series of Legendre polynomials, and to use the Legendre polynomial additiontheorem(Section12.8). 12.5.13 Asingleloopofwireof radius acarriesaconstantcurrent I. (a) Findthemagneticinduction Bforr<a,θ=π/2. (b) Calculate the integral of the magnetic flux (B·dσ)over the area of the current loop,thatis integraldisplaya 0integraldisplay2π 0Bzparenleftbigg r,θ=π 2parenrightbigg dϕrdr. ANS.∞. The Earth is within such a ring current, in which Iapproximates millions of amperes arisingfromthedriftof chargedparticlesintheVanAllenbelt. 12.5.14 (a) Showthatinthepointdipolelimitthemagneticinductionfieldofthecurrentloop becomes Br(r,θ)=µ0 2πm r3P1(cosθ), Bθ(r,θ)=µ0 2πm r3P1 1(cosθ) withm=Iπa2. (b) Comparetheseresultswiththemagneticinductionofthepointmagneticdipoleof Exercise1.8.17. Take m=ˆzm. 12.5.15 Auniformlychargedsphericalshellisrotatingwithconstantangularvelocity. (a) Calculatethemagneticinduction Balongtheaxisofrotationoutsidethesphere. (b) Using the vector potential series of Section 12.5, find Aand then Bfor all space outsidethesphere. 12.5.16 In the liquid drop model of the nucleus, the spherical nucleus is subjected to small deformations. Consider a sphere of radius r0that is deformed so that its new surface is givenby r=r0bracketleftbig 1+α2P2(cosθ)bracketrightbig . Findtheareaofthedeformedspherethroughtermsoforder α2 2. Hint. dA=bracketleftbigg r2+parenleftbiggdr dθparenrightbigg2bracketrightbigg1/2 rsinθdθdϕ. ANS.A=4πr2 0bracketleftbig 1+4 5α2 2+Oparenleftbig α3 2parenrightbigbracketrightbig . 786 Chapter 12 Legendre Functions Note. The area element dAfollows from noting that the line element dsfor fixed ϕis givenby ds=parenleftbig r2dθ2+dr2parenrightbig1/2=parenleftbigg r2+parenleftbiggdr dθparenrightbigg2parenrightbigg1/2 dθ. 12.5.17 A nuclear particle is in a spherical square well potential V(r,θ,ϕ)=0f o r0≤r<a and∞forr>a.Theparticleisdescribedbyawavefunction ψ(r,θ,ϕ) whichsatisfies thewaveequation −¯h2 2M∇2ψ+V0ψ=Eψ, r<a, andtheboundarycondition ψ(r=a)=0. Show that for the energy Eto be a minimum there must be no angular dependence in thewavefunction;thatis, ψ=ψ(r). Hint.Theproblemcentersontheboundaryconditionontheradialfunction. 12.5.18 (a) Write a subroutine to calculate the numerical value of the associated Legendre functionP1 N(x)for givenvaluesof Nandx. Hint. With the known forms of P1 1andP1 2you can use the recurrence relation Eq.(12.92) togenerate P1 N,N>2. (b) Checkyoursubroutinebyhavingitcalculate P1 N(x)forx=0.0(0.5)1.0andN= 1(1)10. Check these numerical values against the known values of P1 N(0)and P1 N(1)andagainstthetabulatedvaluesof P1 N(0.5). 12.5.19 Calculatethemagneticvectorpotentialofacurrentloop,Example12.5.1.Tabulateyour resultsfor r/a=1.5(0.5)5.0andθ=0◦(15◦)90◦.Includetermsintheseriesexpansion, Eq.(12.137),untiltheabsolutevaluesofthetermsdropbelowtheleadingtermbyafac- torof 105ormore. Note. This associated Legendre expansion can be checked by comparison with the el- lipticintegralsolution,Exercise5.8.4. Checkvalue. Forr/a=4.0andθ=20◦, Aϕ/µ0I=4.9398×10−3. 12.6 S PHERICAL HARMONICS In the separation of variables of (1) Laplace’s equation, (2) Helmholtz’s or the space- dependence of the electromagnetic wave equation, and (3) the Schrödinger wave equation forcentralforcefields, ∇2ψ+k2f(r)ψ=0, (12.144) 12.6 Spherical Harmonics 787 theangulardependence,comingentirelyfromtheLaplacianoperator,is19 /Phi1(ϕ) sinθd dθparenleftbigg sinθd/Theta1 dθparenrightbigg +/Theta1(θ) sin2θd2/Phi1(ϕ) dϕ2+n(n+1)/Theta1(θ)/Phi1(ϕ)=0.(12.145) Azimuthal Dependence — Orthogonality Theseparatedazimuthalequationis 1 /Phi1(ϕ)d2/Phi1(ϕ) dϕ2=−m2, (12.146) withsolutions /Phi1(ϕ)=e−imϕ,eimϕ, (12.147) withminteger,whichsatisfy theorthogonalcondition integraldisplay2π 0e−im1ϕeim2ϕdϕ=2πδm1m2. (12.148) Notice that it is the product /Phi1∗ m1(ϕ)/Phi1m2(ϕ)that is taken and that∗is used to indicate the complex conjugate function. This choice is not required, but it is convenient for quantum mechanicalcalculations.We couldhaveused /Phi1=sinmϕ,cosmϕ (12.149) and the conditions of orthogonality that form the basis for Fourier series (Chapter 14). For applications such as describing the Earth’s gravitational or magnetic field, sin mϕand cosmϕwouldbethepreferred choice(seeExample12.6.1). Inelectrostaticsandmostotherphysicalproblemswerequire mtobeanintegerinorder that/Phi1(ϕ)be a single-valued function of the azimuth angle. In quantum mechanics the questionismuchmoreinvolved:ComparethefootnoteinSection9.3. BymeansofEq. (12.148), /Phi1m=1√ 2πeimϕ(12.150) is orthonormal (orthogonal and normalized) with respect to integration over the azimuth angleϕ. 19For a separation constant of the form n(n+1)withnan integer, a Legendre-equation-series solution becomes a polynomial. Otherwise both series solutions diverge, Exercise 9.5.5. 788 Chapter 12 Legendre Functions Polar Angle Dependence Splitting off the azimuthal dependence, the polar angle dependence (θ)leads to the asso- ciatedLegendreequation(12.80), whichis satisfied bytheassociatedLegendrefunctions; that is,/Theta1(θ)=Pm n(cosθ). To include negative values of m, we use Rodrigues’ formula, Eq.(12.65), inthedefinitionof Pm n(cosθ).This leadsto Pm n(cosθ)=1 2nn!parenleftbig 1−x2parenrightbigm/2dm+n dxm+nparenleftbig x2−1parenrightbign,−n≤m≤n.(12.151) Pm n(cosθ)andP−m n(cosθ)arerelatedasindicatedinExercise12.5.1.Anadvantageofthis approach over simply defining Pm n(cosθ)for 0≤m≤nand requiring that P−m n=Pm nis thattherecurrencerelationsvalidfor 0 ≤m≤nremainvalidfor −n≤m<0. Normalizing the associated Legendre function by Eq. (12.110), we obtain the orthonor- malfunctionsradicalBigg 2n+1 2(n−m)! (n+m)!Pm n(cosθ),−n≤m≤n, (12.152) whichareorthonormalwithrespecttothepolarangle θ. Spherical Harmonics The function /Phi1m(ϕ)(Eq. (12.150)) is orthonormal with respect to the azimuthal an- gleϕ. We take the product of /Phi1m(ϕ)and the orthonormal function in polar angle from Eq.(12.152)anddefine Ym n(θ,ϕ)≡(−1)mradicalBigg 2n+1 4π(n−m)! (n+m)!Pm n(cosθ)eimϕ(12.153) to obtain functions of two angles (and two indices) that are orthonormal over the spheri- cal surface. These Ym n(θ,ϕ)are spherical harmonics, of which the first few are plotted in Fig.12.15.Thecompleteorthogonalityintegralbecomes integraldisplay2π ϕ=0integraldisplayπ θ=0Ym1∗ n1(θ,ϕ)Ym2n2(θ,ϕ)sinθdθdϕ=δn1n2δm1m2. (12.154) The extra (−1)mincluded in the defining equation of Ym n(θ,ϕ)deserves some comment. It is clearly legitimate, since Eq. (12.144) is linear and homogeneous. It is not necessary, but in moving on to certain quantum mechanical calculations, particularly in the quan- tum theory of angular momentum (Section 12.7), it is most convenient. The factor (−1)m is a phase factor, often called the Condon–Shortley phase, after the authors of a classic text on atomic spectroscopy. The effect of this (−1)m(Eq. (12.153)) and the (−1)mof Eq. (12.73c) for P−m n(cosθ)is to introduce an alternation of sign among the positive m sphericalharmonics.Thisis showninTable12.3. The functions Ym n(θ,ϕ)acquired the name spherical harmonics first because they are defined over the surface of a sphere with θthe polar angle and ϕthe azimuth. The har- monicwas included because solutions of Laplace’s equation were called harmonic func- tionsand Ym n(cos,ϕ)is theangularpartofsuchasolution. 12.6 Spherical Harmonics 789 FIGURE 12.15[ℜYm l(θ,ϕ)]2for 0≤l≤3,m=0,...,3. 790 Chapter 12 Legendre Functions Table 12.3 Spherical Harmon- ics(Condon–ShortleyPhase) Y0 0(θ,ϕ)=1√ 4π Y1 1(θ,ϕ)=−radicalbigg 3 8πsinθeiϕ Y0 1(θ,ϕ)=radicalbigg 3 4πcosθ Y−1 1(θ,ϕ)=+radicalbigg 3 8πsinθe−iϕ Y2 2(θ,ϕ)=radicalbigg 5 96π3sin2θe2iϕ Y1 2(θ,ϕ)=−radicalbigg 5 24π3sinθcosθeiϕ Y0 2(θ,ϕ)=radicalbigg 5 4πparenleftbigg3 2cos2θ−1 2parenrightbigg Y−1 2(θ,ϕ)=+radicalbigg 5 24π3sinθcosθe−iϕ Y−2 2(θ,ϕ)=radicalbigg 5 96π3sin2θe−2iϕ In the framework of quantum mechanics Eq. (12.145) becomes an orbital angular mo- mentum equation and the solution YM L(θ,ϕ)(nreplaced by L,mreplaced by M)i sa n angular momentum eigenfunction, Lbeing the angular momentum quantum number and Mthez-axis projection of L. These relationships are developed in more detail in Sec- tions4.3and12.7. Laplace Series, Expansion Theorem Part of the importance of spherical harmonics lies in the completeness property, a con- sequence of the Sturm–Liouville form of Laplace’s equation. This property, in this case, means that any function f(θ,ϕ)(with sufficient continuity properties) evaluated over the surfaceofthespherecanbeexpandedinauniformlyconvergentdoubleseriesofspherical harmonics20(Laplace’sseries): f(θ,ϕ)=summationdisplay m,namnYm n(θ,ϕ). (12.155) Iff(θ,ϕ)is known, the coefficients can be immediately found by the use of the orthogo- nalityintegral. 20For a proof of this fundamental theorem see E. W. Hobson, The Theory of Spherical and Ellipsoidal Harmonics ,N e wY o r k : Chelsea (1955), Chapter VII. If f(θ,ϕ)is discontinuous wemay still have convergence in the mean,Section10.4. 12.6 Spherical Harmonics 791 Table 12.4 GravityFieldCoefficients,Eq. (12.156) CoefficientaEarth Moon Mars C20 1.083×10−3(0.200±0.002)×10−3(1.96±0.01)×10−3 C22 0.16×10−5(2.4±0.5)×10−5(−5±1)×10−5 S22 −0.09×10−5(0.5±0.6)×10−5(3±1)×10−5 aC20represents an equatorial bulge, whereas C22andS22represent an azimuthal dependence of the gravitationalfield. Example 12.6.1 LAPLACE SERIES —G RAVITY FIELDS The gravity fields of the Earth, the Moon, and Mars have been described by a Laplace serieswithrealeigenfunctions: U(r,θ,ϕ)=GM RbracketleftbiggR r−∞summationdisplay n=2nsummationdisplay m=0parenleftbiggR rparenrightbiggn+1braceleftbig CnmYe mn(θ,ϕ)+SnmYo mn(θ,ϕ)bracerightbigbracketrightbigg .(12.156) HereMisthemassofthebodyand Ristheequatorialradius.Therealfunctions Ye mnand Yo mnare definedby Ye mn(θ,ϕ)=Pm n(cosθ)cosmϕ, Yo mn(θ,ϕ)=Pm n(cosθ)sinmϕ. For applications such as this, the real trigonometric forms are preferred to the imaginary exponential form of YM L(θ,ϕ). Satellite measurements have led to the numerical values showninTable12.4. /squaresolid Exercises 12.6.1 Show that the parity of YM L(θ,ϕ)is(−1)L. Note the disappearance of any Mdepen- dence. Hint.FortheparityoperationinsphericalpolarcoordinatesseeExercise2.5.8andfoot- note7inSection12.2. 12.6.2 Provethat YM L(0,ϕ)=parenleftbigg2L+1 4πparenrightbigg1/2 δM,0. 12.6.3 InthetheoryofCoulombexcitationofnucleiweencounter YM L(π/2,0). Showthat YM Lparenleftbiggπ 2,0parenrightbigg =parenleftbigg2L+1 4πparenrightbigg1/2[(L−M)!(L+M)!]1/2 (L−M)!!(L+M)!!(−1)(L+M)/2forL+Meven, =0forL+Modd. Here (2n)!!=2n(2n−2)···6·4·2, (2n+1)!!=(2n+1)(2n−1)···5·3·1. 792 Chapter 12 Legendre Functions 12.6.4 (a) Express the elements of the quadrupole moment tensor xixjas a linear combina- tionof thesphericalharmonics Ym 2(andY0 0). Note.Thetensor xixjisreducible.The Y0 0indicatesthepresenceofascalarcom- ponent. (b) Thequadrupolemomenttensorisusuallydefinedas Qij=integraldisplayparenleftbig 3xixj−r2δijparenrightbig ρ(r)dτ, withρ(r)the charge density. Express the components of (3xixj−r2δij)in terms ofr2YM 2. (c) Whatisthesignificanceof the −r2δijterm? Hint.CompareSections2.9and4.4. 12.6.5 TheorthogonalazimuthalfunctionsyieldausefulrepresentationoftheDiracdeltafunc- tion.Showthat δ(ϕ1−ϕ2)=1 2π∞summationdisplay m=−∞expbracketleftbig im(ϕ1−ϕ2)bracketrightbig . 12.6.6 Derivethesphericalharmonicclosurerelation ∞summationdisplay l=0+lsummationdisplay m=−lYm l(θ1,ϕ1)Ym∗ l(θ2,ϕ2)=1 sinθ1δ(θ1−θ2)δ(ϕ1−ϕ2) =δ(cosθ1−cosθ2)δ(ϕ1−ϕ2). 12.6.7 Thequantummechanicalangularmomentumoperators Lx±iLyaregivenby Lx+iLy=eiϕparenleftbigg∂ ∂θ+icotθ∂ ∂ϕparenrightbigg , Lx−iLy=−e−iϕparenleftbigg∂ ∂θ−icotθ∂ ∂ϕparenrightbigg . Showthat (a)(Lx+iLy)YM L(θ,ϕ)=radicalbig (L−M)(L+M+1)YM+1 L(θ,ϕ), (b)(Lx−iLy)YM L(θ,ϕ)=radicalbig (L+M)(L−M+1)YM−1 L(θ,ϕ). 12.6.8 WithL±givenby L±=Lx±iLy=±e±iϕbracketleftbigg∂ ∂θ±icotθ∂ ∂ϕbracketrightbigg , showthat (a)Ym l=radicalBigg (l+m)! (2l)!(l−m)!(L−)l−mYl l, 12.7 Orbital Angular Momentum Operators 793 (b)Ym l=radicalBigg (l−m)! (2l)!(l+m)!(L+)l+mY−l l. 12.6.9 Insomecircumstancesitisdesirabletoreplacetheimaginaryexponentialofourspher- ical harmonic by sine or cosine. Morse and Feshbach (see the General References at book’send)define Ye mn=Pm n(cosθ)cosmϕ, Yo mn=Pm n(cosθ)sinmϕ, where integraldisplay2π 0integraldisplayπ 0bracketleftbig Yeoro mn(θ,ϕ)bracketrightbig2sinθdθdϕ=4π 2(2n+1)(n+m)! (n−m)!,n=1,2,... =4πforn=0(Yo 00isundefined ). These spherical harmonics are often named according to the patterns of their positive and negative regions on the surface of a sphere—zonal harmonics for m=0, sectoral harmonics for m=n, and tesseral harmonics for 0 <m<n.F o rYe mn,n=4,m= 0,2,4,indicateonadiagramofahemisphere(onediagramforeachsphericalharmonic) theregionsinwhichthesphericalharmonicis positive. 12.6.10 Afunction f(r,θ,ϕ) maybeexpressedasa Laplaceseries f(r,θ,ϕ)=summationdisplay l,malmrlYm l(θ,ϕ). With/angbracketleft/angbracketrightsphereusedtomeantheaverageoverasphere(centeredontheorigin),showthat angbracketleftbig f(r,θ,ϕ)angbracketrightbig sphere=f(0,0,0). 12.7 O RBITAL ANGULAR MOMENTUM OPERATORS Now we return to the specific orbital angular momentum operators Lx,Ly, andLzof quantummechanicsintroducedinSection4.3. Equation(4.68)becomes LzψLM(θ,ϕ)=MψLM(θ,ϕ), andwewanttoshowthat ψLM(θ,ϕ)=YM L(θ,ϕ) aretheeigenfunctions |LM/angbracketrightofL2andLzofSection4.3insphericalpolarcoordinates,the sphericalharmonics.Theexplicitformof Lz=−i∂/∂ϕfromExercise2.5.13indicatesthat ψLMhas aϕdependence of exp (iMϕ)—withMan integer to keep ψLMsingle-valued. AndifMis aninteger,then Lisanintegeralso. To determine the θdependence of ψLM(θ,ϕ), we proceed in two main steps: (1) the determination of ψLL(θ,ϕ)and (2) the development of ψLM(θ,ϕ)in terms of ψLLwith thephasefixedby ψL0.L et ψLM(θ,ϕ)=/Theta1LM(θ)eiMϕ. (12.157) 794 Chapter 12 Legendre Functions FromL+ψLL=0,Lbeing the largest M, using the form of L+given in Exercises 2.5.14 and12.6.7, wehave ei(L+1)ϕbracketleftbiggd dθ−Lcotθbracketrightbigg /Theta1LL(θ)=0, (12.158) andthus ψLL(θ,ϕ)=cLsinLθeiLϕ. (12.159) Normalizing,weobtain c∗ LcLintegraldisplay2π 0integraldisplayπ 0sin2L+1θdθdϕ=1. (12.160) Theθintegralmaybeevaluatedasabetafunction(Exercise8.4.9) and |cL|=radicalBigg (2L+1)!! 4π(2L)!!=√(2L)! 2LL!radicalbigg 2L+1 4π. (12.161) Thiscompletesourfirst step. To obtain the ψLM,M/negationslash=±L, we return to the ladder operators. From Eqs. (4.83) and (4.84)andasshowninExercise12.7.2( J+replacedby L+andJ−replacedby L−), ψLM(θ,ϕ)=radicalBigg (L+M)! (2L)!(L−M)!(L−)L−MψLL(θ,ϕ), (12.162) ψLM(θ,ϕ)=radicalBigg (L−M)! (2L)!(L+M)!(L+)L+MψL,−L(θ,ϕ). Again, note that the relative phases are set by the ladder operators. L+andL−operating on/Theta1LM(θ)eiMϕmaybewrittenas L+/Theta1LM(θ)eiMϕ=ei(M+1)ϕbracketleftbiggd dθ−Mcotθbracketrightbigg /Theta1LM(θ) =ei(M+1)ϕsin1+Mθd d(cosθ)sin−M/Theta1LM(θ), (12.163) L−/Theta1LM(θ)eiMϕ=−ei(M−1)ϕbracketleftbiggd dθ+Mcotθbracketrightbigg /Theta1LM(θ) =ei(M−1)ϕsin1−Mθd d(cosθ)sinMθ/Theta1LM(θ). Repeatingtheseoperations ntimesyields (L+)n/Theta1LM(θ)eiMϕ=(−1)nei(M+n)ϕsinn+Mθdnsin−Mθ/Theta1LM(θ) d(cosθ)n, (12.164) (L−)n/Theta1LM(θ)eiMϕ=ei(M−n)ϕsinn−MθdnsinMθ/Theta1LM(θ) d(cosθ)n. 12.7 Orbital Angular Momentum Operators 795 FromEq. (12.163), ψLM(θ,ϕ)=cLradicalBigg (L+M)! (2L)!(L−M)!eiMϕsin−MθdL−M d(cosθ)L−Msin2Lθ,(12.165) andforM=−L: ψL,−L(θ,ϕ)=cL (2L)!e−iLϕsinLθd2L d(cosθ)2Lsin2Lθ =(−1)LcLsinLθe−iLϕ. (12.166) Notethecharacteristic (−1)LphaseofψL,−Lrelativeto ψL,L.This(−1)Lentersfrom sin2Lθ=parenleftbig 1−x2parenrightbigL=(−1)Lparenleftbig x2−1parenrightbigL. (12.167) CombiningEqs. (12.163), (12.163),and(12.166),weobtain ψLM(θ,ϕ)=(−1)LcLradicalBigg (L−M)! (2L)!(L+M)!(−1)L+MeiMϕsinMθdL+Msin2Lθ d(cosθ)L+M.(12.168) Equations(12.165)and(12.168)agreeif ψL0(θ,ϕ)=cL1√(2L)!dL (dcosθ)Lsin2Lθ. (12.169) UsingRodrigues’formula,Eq. (12.65), wehave ψL0(θ,ϕ)=(−1)LcL2LL!√(2L)!PL(cosθ) =(−1)LcL |cL|radicalbigg 2L+1 4πPL(cosθ). (12.170) The last equality follows from Eq. (12.161). We now demand that ψL0(0,0)be real and positive.Therefore cL=(−1)L|cL|=(−1)L√(2L)! 2LL!radicalbigg 2L+1 4π. (12.171) With(−1)LcL/|cL|=1,ψL0(θ,ϕ)in Eq. (12.170) may be identified with the spherical harmonic Y0 L(θ,ϕ)ofSection12.6. 796 Chapter 12 Legendre Functions Whenwesubstitutethevalueof (−1)LcLintoEq.(12.168), ψLM(θ,ϕ)=√(2L)! 2LL!radicalbigg 2L+1 4πradicalBigg (L−M)! (2L)!(L+M)!(−1)L+M ·eiMϕsinMθdL+M d(cosθ)L+Msin2Lθ =radicalbigg 2L+1 4πradicalBigg (L−M)! (L+M)!eiMϕ(−1)M ·braceleftbigg1 2LL!parenleftbig 1−x2parenrightbigM/2dL+M dxL+Mparenleftbig x2−1parenrightbigLbracerightbigg ,x=cosθ, M≥0. (12.172) The expression in the curly bracket is identified as the associated Legendre function (Eq. (12.151),andwehave ψLM(θ,ϕ)=YM L(θ,ϕ) =(−1)MradicalBigg 2L+1 4π·(L−M)! (L+M)!·PM L(cosθ)eiMϕ,M≥0, (12.173) in complete agreement with Section 12.6. Then by Eq. (12.73c), YM Lfor negative super- scriptisgivenby Y−M L(θ,ϕ)=(−1)Mbracketleftbig YM L(θ,ϕ)bracketrightbig∗. (12.174) •Our angular momentum eigenfunctions ψLM(θ,ϕ)are identified with the spherical harmonics. The phase factor (−1)Mis associated with the positive values of Mand is seentobeaconsequenceoftheladderoperators. •OurdevelopmentofsphericalharmonicsheremaybeconsideredaportionofLiealge- bra—relatedtogrouptheory,Section4.3. Exercises 12.7.1 Usingtheknownforms of L+andL−(Exercises2.5.14 and12.6.7), showthat integraldisplaybracketleftbig YM Lbracketrightbig∗L−parenleftbig L+YM Lparenrightbig d/Omega1=integraldisplayparenleftbig L+YM Lparenrightbig∗parenleftbig L+YM Lparenrightbig d/Omega1. 12.7.2 Derivetherelations (a)ψLM(θ,ϕ)=radicalBigg (L+M)! (2L)!(L−M)!(L−)L−MψLL(θ,ϕ), (b)ψLM(θ,ϕ)=radicalBigg (L−M)! (2L)!(L+M)!(L+)L+MψL,−L(θ,ϕ). 12.8 Addition Theorem for Spherical Harmonics 797 Hint.Equations(4.83) and(4.84) maybehelpful. 12.7.3 Derivethemultipleoperatorequations (L+)n/Theta1LM(θ)eiMϕ=(−1)nei(M+n)ϕsinn+Mθdnsin−Mθ/Theta1LM(θ) d(cosθ)n, (L−)n/Theta1LM(θ)eiMϕ=ei(M−n)ϕsinn−MθdnsinMθ/Theta1LM(θ) d(cosθ)n. Hint.Trymathematicalinduction. 12.7.4 Show,using (L−)n, that Y−M L(θ,ϕ)=(−1)MY∗M L(θ,ϕ). 12.7.5 Verifybyexplicitcalculationthat (a)L+Y0 1(θ,ϕ)=−radicalbigg 3 4πsinθeiϕ=√ 2Y1 1(θ,ϕ), (b)L−Y0 1(θ,ϕ)=+radicalbigg 3 4πsinθe−iϕ=√ 2Y−1 1(θ,ϕ). The signs (Condon–Shortley phase) are a consequence of the ladder operators L+and L−. 12.8 T HEADDITION THEOREM FOR SPHERICAL HARMONICS Trigonometric Identity In the following discussion, (θ1,ϕ1)and(θ2,ϕ2)denote two different directions in our spherical coordinate system (x1,y1,z1), separated by an angle γ(Fig. 12.16). The polar anglesθ1,θ2aremeasuredfromthe z1-axis.Theseanglessatisfythetrigonometricidentity cosγ=cosθ1cosθ2+sinθ1sinθ2cos(ϕ1−ϕ2), (12.175) whichisperhapsmosteasilyprovedbyvectormethods(compareChapter1). Theadditiontheorem,then,asserts that Pn(cosγ)=4π 2n+1nsummationdisplay m=−n(−1)mYm n(θ1,ϕ1)Y−m n(θ2,ϕ2), (12.176) orequivalently, Pn(cosγ)=4π 2n+1nsummationdisplay m=−nYm n(θ1,ϕ1)bracketleftbig Ym n(θ2,ϕ2)bracketrightbig∗.21(12.177) 21Theasterisk for complex conjugation may go on either spherical harmonic. 798 Chapter 12 Legendre Functions Intermsof theassociatedLegendrefunctions,theadditiontheoremis Pn(cosγ)=Pn(cosθ1)Pn(cosθ2) +2nsummationdisplay m=1(n−m)! (n+m)!Pm n(cosθ1)Pm n(cosθ2)cosm(ϕ1−ϕ2). (12.178) Equation(12.175)isaspecialcaseofEq. (12.178), n=1. Derivation of Addition Theorem WenowderiveEq.(12.177).Let (γ,ξ)betheanglesthatspecifythedirection (θ1,ϕ1)ina coordinatesystem (x2,y2,z2)whoseaxisis alignedwith (θ2,ϕ2).(Actually,thechoiceof the 0 azimuthangle ξinFig.12.16isirrelevant.)First,weexpand Ym n(θ1,ϕ1)inspherical harmonicsinthe (γ,ξ)angularvariables: Ym n(θ1,ϕ1)=nsummationdisplay σ=−nam nσYσ n(γ,ξ). (12.179) Wewritenosummationover ninEq.(12.179)becausetheangularmomentum nofYm nis conserved(seeSection4.3);asasphericalharmonic, Ym n(θ1,ϕ1)isaneigenfunctionof L2 witheigenvalue n(n+1). FIGURE 12.16Twodirectionsseparatedbyan angleγ. 12.8 Addition Theorem for Spherical Harmonics 799 Weneedforourproofonlythecoefficient am n0,whichwegetbymultiplyingEq.(12.179) by[Y0 n(γ,ξ)]∗andintegratingoverthesphere: am n0=integraldisplay Ym n(θ1,ϕ1)bracketleftbig Y0 n(γ,ξ)bracketrightbig∗d/Omega1γ,ξ. (12.180) Similarly,weexpand Pn(cosγ)intermsofsphericalharmonics Ym n(θ1,ϕ1): Pn(cosγ)=parenleftbigg4π 2n+1parenrightbigg1/2 Y0 n(γ,ξ)=nsummationdisplay m=−nbnmYm n(θ1,ϕ1), (12.181) where the bnmwill, of course, depend on θ2,ϕ2, that is, on the orientation of the z2-axis. Multiplyingby [Ym n(θ1,ϕ1)]∗andintegratingwithrespectto θ1andϕ1overthesphere,we have bnm=integraldisplay Pn(cosγ)Ym∗ n(θ1,ϕ1)d/Omega1θ1,ϕ1. (12.182) Intermsof sphericalharmonicsEq. (12.182)becomes parenleftbigg4π 2n+1parenrightbigg1/2integraldisplay Y0 n(γ,ψ)bracketleftbig Ym n(θ1,ϕ1)bracketrightbig∗d/Omega1=bnm. (12.183) Note that the subscripts have been dropped from the solid angle element d/Omega1. Since the range of integration is over all solid angles, the choice of polar axis is irrelevant. Then comparingEqs. (12.180)and(12.183),weseethat b∗ nm=am n0parenleftbigg4π 2n+1parenrightbigg1/2 . (12.184) Now we evaluate Ym n(θ2,ϕ2)using the expansion of Eq. (12.179) and noting that the valuesof (γ,ξ)correspondingto (θ1,ϕ1)=(θ2,ϕ2)are(0,0).Theresultis Ym n(θ2,ϕ2)=am n0Y0 n(0,0)=am n0parenleftbigg2n+1 4πparenrightbigg1/2 , (12.185) alltermswithnonzero σvanishing.SubstitutingthisbackintoEq. (12.184),weobtain bnm=4π 2n+1bracketleftbig Ym n(θ2,ϕ2)bracketrightbig∗. (12.186) Finally, substituting this expression for bnminto the summation, Eq. (12.181) yields Eq.(12.177), thusprovingouradditiontheorem. Those familiar with group theory will find a much more elegant proof of Eq. (12.177) byusingtherotationgroup.22Thisis Exercise4.4.5. One application of the addition theorem is in the construction of a Green’s function for the three-dimensional Laplace equation in spherical polar coordinates. If the source is on 22C o m p ar eM.E.R o s e, ElementaryTheory of Angular Momentum , NewYork: Wiley (1957). 800 Chapter 12 Legendre Functions thepolaraxisatthepoint (r=a,θ=0,ϕ=0),then,byEq. (12.4a), 1 R=1 |r−ˆza|=∞summationdisplay n=0Pn(cosγ)an rn+1,r>a =∞summationdisplay n=0Pn(cosγ)rn an+1,r<a. (12.187) Rotatingourcoordinatesystemtoputthesourceat (a,θ2,ϕ2)andthepointofobservation at(r,θ1,ϕ1), weobtain G(r,θ1,ϕ1,a,θ2,ϕ2)=1 R =∞summationdisplay n=0nsummationdisplay m=−n4π 2n+1bracketleftbig Ym n(θ1,ϕ1)bracketrightbig∗Ym n(θ2,ϕ2)an rn+1,r>a, =∞summationdisplay n=0nsummationdisplay m=−n4π 2n+1bracketleftbig Ym n(θ1,ϕ1)bracketrightbig∗Ym n(θ2,ϕ2)rn an+1,r<a.(12.188) In Section 9.7 this argument is reversed to provide another derivation of the Legendre polynomialadditiontheorem. Exercises 12.8.1 In proving the addition theorem, we assumed that Yk n(θ1,ϕ1)could be expanded in a series of Ym n(θ2,ϕ2), in which mvaried from−nto+nbutnwas held fixed. What arguments can you develop to justify summing only over the upper index, m, andnot overthelowerindex, n? Hints. One possibility is to examine the homogeneity of the Ym n, that is, Ym nmay be expressed entirely in terms of the form cosn−pθsinpθ,o rxn−p−sypzs/rn. Another possibility is to examine the behavior of the Legendre equation under rotation of the coordinatesystem. 12.8.2 Anatomicelectronwithangularmomentum Landmagneticquantumnumber Mhasa wavefunction ψ(r,θ,ϕ)=f(r)YM L(θ,ϕ). Show that the sum of the electron densities in a given complete shell is spherically symmetric;thatis,summationtextL M=−Lψ∗(r,θ,ϕ)ψ(r,θ,ϕ) isindependentof θandϕ. 12.8.3 Thepotentialof anelectronatpoint reinthefieldof Zprotonsatpoints rpis /Phi1=−e2 4πε0Zsummationdisplay p=11 |re−rp|. 12.8 Addition Theorem for Spherical Harmonics 801 Showthatthismaybewrittenas /Phi1=−e2 4πε0reZsummationdisplay p=1summationdisplay L,Mparenleftbiggrp reparenrightbiggL4π 2L+1bracketleftbig YM L(θp,ϕp)bracketrightbig∗YM L(θe,ϕe), wherere>rp.Howshould /Phi1bewrittenfor re<rp? 12.8.4 Two protons are uniformly distributed within the same spherical volume. If the coor- dinates of one element of charge are (r1,θ1,ϕ1)and the coordinates of the other are (r2,θ2,ϕ2)andr12is the distance between them, the element of energy of repulsion willbegivenby dψ=ρ2dτ1dτ2 r12=ρ2r2 1dr1sinθ1dθ1dϕ1r2 2dr2sinθ2dθ2dϕ2 r12. Here ρ=charge volume=3e 4πR3,chargedensity , r2 12=r2 1+r2 2−2r1r2cosγ. Calculatethetotalelectrostaticenergy(ofrepulsion)ofthetwoprotons.Thiscalculation isusedinaccountingfor themassdifferencein“mirror”nuclei,suchasO15andN15. ANS.6 5e2 R. This isdoublethat required to create a uniformly charged sphere because we have two separatecloudchargesinteracting,notonechargeinteractingwithitself(withpermuta- tionof pairs notconsidered). 12.8.5 Eachofthetwo1 Selectronsinheliummaybedescribedbyahydrogenicwavefunction ψ(r)=parenleftbiggZ3 πa3 0parenrightbigg1/2 e−Zr/a0 in the absence of the other electron. Here Z, the atomic number, is 2. The symbol a0is the Bohr radius, ¯h2/me2. Find the mutual potential energy of the two electrons, given by integraldisplay ψ∗(r1)ψ∗(r2)e2 r12ψ(r1)ψ(r2)d3r1d3r2. ANS.5e2Z 8a0. Note.d3r1=r2dr1sinθ1dθ1dϕ1≡dτ1,r12=|r1−r2|. 12.8.6 Theprobabilityoffindinga1 Shydrogenelectroninavolumeelement r2drsinθdθdϕ is 1 πa3 0exp[−2r/a0]r2drsinθdθdϕ. 802 Chapter 12 Legendre Functions Findthecorrespondingelectrostaticpotential.Calculatethepotentialfrom V(r1)=q 4πε0integraldisplayρ(r2) r12d3r2, withr1notonthez-axis.Expand r12.ApplytheLegendrepolynomialadditiontheorem andshowthattheangulardependenceof V(r1)dropsout. ANS.V(r1)=q 4πε0braceleftbigg1 2r1γparenleftbigg 3,2r1 a0parenrightbigg +1 a0Ŵparenleftbigg 2,2r1 a0parenrightbiggbracerightbigg . 12.8.7 Ahydrogenelectronina 2 Porbithas achargedistribution ρ=q 64πa5 0r2e−r/a0sin2θ, wherea0is the Bohr radius, ¯h2/me2. Find the electrostatic potential corresponding to thischargedistribution. 12.8.8 Theelectriccurrentdensityproducedbya 2 Pelectroninahydrogenatomis J=ˆϕq¯h 32ma5 0e−r/a0rsinθ. Using A(r1)=µ0 4πintegraldisplayJ(r2) |r1−r2|d3r2, findthemagneticvectorpotentialproducedbythishydrogenelectron. Hint. Resolve into Cartesian components. Use the addition theorem to eliminate γ,t h e angleincludedbetween r1andr2. 12.8.9 (a) As a Laplace series and as an example of Eq. (1.190) (now with complex func- tions),showthat δ(/Omega11−/Omega12)=∞summationdisplay n=0nsummationdisplay m=−nYm∗ n(θ2,ϕ2)Ym n(θ1,ϕ1). (b) Showalsothatthis sameDiracdeltafunctionmaybewrittenas δ(/Omega11−/Omega12)=∞summationdisplay n=02n+1 4πPn(cosγ). Now, if you can justify equating the summations over nterm by term , you have an alternatederivationofthesphericalharmonicadditiontheorem. 12.9 Integrals of Three Y’s 803 12.9 I NTEGRALS OF PRODUCTS OF THREE SPHERICAL HARMONICS Frequentlyinquantummechanicsweencounterintegralsofthegeneralform angbracketleftbig YM1 L1vextendsinglevextendsingleYM2 L2vextendsinglevextendsingleYM3 L3angbracketrightbig =integraldisplay2π 0integraldisplayπ 0bracketleftbig YM1 L1bracketrightbig∗YM2 L2YM3 L3sinθdθdϕ =radicalBigg (2L2+1)(2L3+1) 4π(2L1+1)C(L2L3L1|000)C(L2L3L1|M2M3M1), (12.189) inwhichallsphericalharmonicsdependon θ,ϕ.Thefirstfactorintheintegrandmaycome fromthewavefunctionofafinalstateandthethirdfactorfromaninitialstate,whereasthe middlefactormayrepresentanoperatorthatisbeingevaluatedorwhose“matrixelement” isbeingdetermined. By using group theoretical methods, as in the quantum theory of angular momentum, we may give a general expression for the forms listed. The analysis involves the vector– additionorClebsch–GordancoefficientsfromSection4.4,whicharetabulated.Threegen- eralrestrictionsappear. 1. The integral vanishes unless the triangle condition of the L’s (angular momentum) is zero,|L1−L3|≤L2≤L1+L3. 2. Theintegralvanishesunless M2+M3=M1.Herewehavethetheoreticalfoundation of thevectormodelofatomicspectroscopy. 3. Finally,theintegralvanishesunlesstheproduct [YM1 L1]∗YM2 L2YM3 L3iseven,thatis,unless L1+L2+L3isaneveninteger.This isaparityconservationlaw. ThekeytothedeterminationoftheintegralinEq.(12.189)istheexpansionoftheprod- uct of two spherical harmonics depending on the same angles (in contrast to the addition theorem),whicharecoupledbyClebsch–Gordancoefficientstoangularmomentum L,M, which, from its rotational transformation properties, must be proportional to YM L(θ,ϕ); thatis,summationdisplay M1,M2C(L2L3L1|M2M3M1)YM2 L2(θ,ϕ)YM3 L3(θ,ϕ)∼YM1 L1(θ,ϕ). Fordetailswerefer toEdmonds.23 LetusoutlinesomeofthestepsofthisgeneralandpowerfulapproachusingSection4.4. TheWigner–EckarttheoremappliedtothematrixelementinEq. (12.189)yields angbracketleftbig YM1 L1vextendsinglevextendsingleYM2 L2vextendsinglevextendsingleYM3 L3angbracketrightbig =(−1)L2−L3+L1C(L2L3L1|M2M3M1) ·/angbracketleftYL1/bardblYL2/bardblYL3/angbracketright √(2L1+1), (12.190) 23E. U. Condon and G. H. Shortley, The Theory of Atomic Spectra , Cambridge, UK: Cambridge University Press (1951); M. E. Rose,Elementary Theory of Angular Momentum , New York: Wiley (1957); A. Edmonds, Angular Momentum in Quantum Mechanics , Princeton, NJ: Princeton University Press (1957); E. P. Wigner, Group Theory and Its Applications to Quantum Mechanics (translated by J. J. Griffin), NewYork: AcademicPress (1959). 804 Chapter 12 Legendre Functions where the double bars denote the reduced matrix element, which no longer depends on theMi. Selection rules (1) and (2) mentioned earlier follow directly from the Clebsch– Gordan coefficient in Eq. (12.190). Next we use Eq. (12.190) for M1=M2=M3=0i n conjunctionwithEq. (12.153)for m=0,whichyields angbracketleftbig Y0 L1vextendsinglevextendsingleY0 L2vextendsinglevextendsingleY0 L3angbracketrightbig =(−1)L2−L3+L1 √2L1+1C(L2L3L1|000)·/angbracketleftYL1/bardblYL2/bardblYL3/angbracketright =radicalbigg (2L1+1)(2L2+1)(2L3+1) 4π ·1 2·integraldisplay1 −1PL1(x)PL2(x)PL3(x)dx, (12.191) wherex=cosθ.Byelementarymethodsitcanbeshownthat integraldisplay1 −1PL1(x)PL2(x)PL3(x)dx=2 2L1+1C(L2L3L1|000)2. (12.192) SubstitutingEq. (12.192)into(12.191)weobtain /angbracketleftYL1/bardblYL2/bardblYL3/angbracketright=(−1)L2−L3+L1C(L2L3L1|000)radicalbigg (2L2+1)(2L3+1) 4π.(12.193) The aforementioned parity selection rule (3) above follows from Eq. (12.193) in conjunc- tionwiththephaserelation C(L2L3L1|−M2,−M3,−M1)=(−1)L2+L3−L1C(L2L3L1|M2M3M1).(12.194) Note that the vector-addition coefficients are developed in terms of the Condon–Shortley phaseconvention,23inwhichthe (−1)mof Eq. (12.153)isassociatedwiththepositive m. It is possibleto evaluatemanyof the commonlyencounteredintegrals of this form with the techniques already developed. The integration over azimuth may be carried out by inspection: integraldisplay2π 0e−iM1ϕeiM2ϕeiM3ϕdϕ=2πδM2+M3−M1,0. (12.195) Physicallythiscorrespondstotheconservationofthe zcomponentofangularmomentum. Application of Recurrence Relations A glance at Table 12.3 will show that the θ-dependence of YM2 L2, that is,PM2 L2(θ), can be expressedintermsof cos θand sinθ.However,afactorof cos θor sinθmaybecombined withtheYM3 L3factorbyusingtheassociatedLegendrepolynomialrecurrencerelations.For 12.9 Integrals of Three Y’s 805 instance,fromEqs. (12.92) and(12.93) weget cosθYM L=+bracketleftbigg(L−M+1)(L+M+1) (2L+1)(2L+3)bracketrightbigg1/2 YM L+1 +bracketleftbigg(L−M)(L+M) (2L−1)(2L+1)bracketrightbigg1/2 YM L−1 (12.196) eiϕsinθYM L=−bracketleftbigg(L+M+1)(L+M+2) (2L+1)(2L+3)bracketrightbigg1/2 YM+1 L+1 +bracketleftbigg(L−M)(L−M−1) (2L−1)(2L+1)bracketrightbigg1/2 YM+1 L−1(12.197) e−iϕsinθYM L=+bracketleftbigg(L−M+1)(L−M+2) (2L+1)(2L+3)bracketrightbigg1/2 YM−1 L+1 −bracketleftbigg(L+M)(L+M−1) (2L−1)(2L+1)bracketrightbigg1/2 YM−1 L−1. (12.198) Usingtheseequations,weobtain integraldisplay YM1∗ L1cosθYM Ld/Omega1=bracketleftbigg(L−M+1)(L+M+1) (2L+1)(2L+3)bracketrightbigg1/2 δM1,MδL1,L+1 +bracketleftbigg(L−M)(L+M) (2L−1)(2L+1)bracketrightbigg1/2 δM1,MδL1,L−1.(12.199) The occurrence of the Kronecker delta (L1,L±1)is an aspect of the conservation of angular momentum. Physically, this integral arises in a consideration of ordinary atomic electromagneticradiation(electricdipole).Itleadstothefamiliarselectionrulethattransi- tionstoanatomiclevelwithorbitalangularmomentumquantumnumber L1canoriginate only from atomic levels with quantum numbers L1−1o rL1+1. The application to ex- pressionssuchas quadrupolemoment ∼integraldisplay YM∗ L(θ,ϕ)P 2(cosθ)YM L(θ,ϕ)d/Omega1 ismoreinvolvedbutperfectlystraightforward. Exercises 12.9.1 Verify (a)integraldisplay YM L(θ,ϕ)Y0 0(θ,ϕ)YM∗ L(θ,ϕ)d/Omega1=1√ 4π, (b)integraldisplay YM LY0 1YM∗ L+1d/Omega1=radicalbigg 3 4πradicalBigg (L+M+1)(L−M+1) (2L+1)(2L+3), 806 Chapter 12 Legendre Functions (c)integraldisplay YM LY1 1YM+1∗ L+1d/Omega1=radicalbigg 3 8πradicalBigg (L+M+1)(L+M+2) (2L+1)(2L+3), (d)integraldisplay YM LY1 1YM+1∗ L−1d/Omega1=−radicalbigg 3 8πradicalBigg (L−M)(L−M−1) (2L−1)(2L+1). These integrals were used in an investigation of the angular correlation of internal con- versionelectrons. 12.9.2 Showthat (a)integraldisplay1 −1xPL(x)PN(x)dx=  2(L+1) (2L+1)(2L+3),N=L+1, 2L (2L−1)(2L+1),N=L−1, (b)integraldisplay1 −1x2PL(x)PN(x)dx=  2(L+1)(L+2) (2L+1)(2L+3)(2L+5),N=L+2, 2(2L2+2L−1) (2L−1)(2L+1)(2L+3),N=L, 2L(L−1) (2L−3)(2L−1)(2L+1),N=L−2. 12.9.3 SincexPn(x)is a polynomial (degree n+1), it may be represented by the Legendre series xPn(x)=∞summationdisplay s=0asPs(x). (a) Showthat as=0f o rs<n−1 ands>n+1. (b) Calculate an−1,an, andan+1and show that you have reproduced the recurrence relation,Eq. 12.17. Note. This argument may be put in a general form to demonstrate the existence of a three-termrecurrencerelationfor anyof ourcompletesets oforthogonalpolynomials: xϕn=an+1ϕn+1+anϕn+an−1ϕn−1. 12.9.4 Show that Eq. (12.199) is a special case of Eq. (12.190) and derive the reduced matrix element/angbracketleftYL1/bardblY1/bardblYL/angbracketright. ANS./angbracketleftYL1/bardblY1/bardblYL/angbracketright=(−1)L1+1−LC(1LL1|000)√3(2L+1) 4π. 12.10 L EGENDRE FUNCTIONS OF THE SECOND KIND In all the analysis so far in this chapter we have been dealing with one solution of Legen- dre’sequation,thesolution Pn(cosθ),whichisregular(finite)atthetwosingularpointsof 12.10 Legendre Functions of the Second Kind 807 the differential equation, cos θ=±1. From the general theory of differential equations it isknownthatasecondsolutionexists.Wedevelopthissecondsolution, Qn,withnonneg- ative integer n(because Qnin applications will occur in conjunction with Pn) ,b yas e r i e s solutionofLegendre’sequation.Later aclosedform willbeobtained. Series Solutions of Legendre’s Equation To solve d dxbracketleftbigg (1−x2)dy dxbracketrightbigg +n(n+1)y=0 (12.200) weproceedasinChapter9,letting24 y=∞summationdisplay λ=0aλxk+λ, (12.201) with y′=∞summationdisplay λ=0(k+λ)aλxk+λ−1, (12.202) y′′=∞summationdisplay λ=0(k+λ)(k+λ−1)aλxk+λ−2. (12.203) Substitutionintotheoriginaldifferentialequationgives ∞summationdisplay λ=0(k+λ)(k+λ−1)aλxk+λ−2 +∞summationdisplay λ=0bracketleftbig n(n+1)−2(k+λ)−(k+λ)(k+λ−1)bracketrightbig aλxk+λ=0.(12.204) Theindicialequation is k(k−1)=0, (12.205) withsolutions k=0,1.Wetryfirst k=0witha0=1,a1=0.Thenourseriesisdescribed bytherecurrencerelation (λ+2)(λ+1)aλ+2+bracketleftbig n(n+1)−2λ−λ(λ−1)bracketrightbig aλ=0, (12.206) whichbecomes aλ+2=−(n+λ+1)(n−λ) (λ+1)(λ+2)aλ. (12.207) 24Notethat xmaybe replacedby the complex variable z. 808 Chapter 12 Legendre Functions Labelingthisseries, from Eq. (12.201), y(x)=pn(x),weha v e pn(x)=1−n(n+1) 2!x2+(n−2)n(n+1)(n+3) 4!x4+···. (12.208) The second solution of the indicial equation, k=1, witha0=0,a1=1, leads to the recurrencerelation aλ+2=−(n+λ+2)(n−λ−1) (λ+2)(λ+3)aλ. (12.209) Labelingthisseries, from Eq. (12.201), y(x)=qn(x), weobtain qn(x)=x−(n−1)(n+2) 3!x3+(n−3)(n−1)(n+2)(n+4) 5!x5−···.(12.210) Ourgeneralsolutionof Eq.(12.200), then,is yn(x)=Anpn(x)+Bnqn(x), (12.211) providedwehaveconvergence .FromGauss’test,Section5.2(seeExample5.2.4),wedo nothaveconvergenceat x=±1.Togetoutofthisdifficulty,wesettheseparationconstant nequaltoaninteger(Exercise9.5.5) andconverttheinfiniteseries intoapolynomial. Forna positive even integer (or zero), series pnterminates, and with a proper choice of a normalizing factor (selected to obtain agreement with the definition of Pn(x)in Sec- tion12.1) Pn(x)=(−1)n/2n! 2n[(n/2)!]2pn(x)=(−1)s(2s)! 22s(s!)2p2s(x) =(−1)s(2s−1)!! (2s)!!p2s(x), forn=2s. (12.212) Ifnis a positive odd integer, series qnterminates after a finite number of terms, and we write Pn(x)=(−1)n−1)/2 n! 2n−1{[n−1)/2]!}2qn(x) =(−1)s(2s+1)! 22s(s!)2q2s+1(x)=(−1)s(2s+1)!! (2s)!!q2s+1(x), forn=2s+1. (12.213) Note that these expressions hold for all real values of x,−∞<x<∞, and for complex values in the finite complex plane. The constants that multiply pnandqnare chosen to makePnagreewithLegendrepolynomialsgivenbythegeneratingfunction. Equations (12.208) and (12.210) may still be used with n=ν, not an integer, but now the series no longer terminates, and the range of convergence becomes −1<x<1. The endpoints, x=±1,are notincluded. It is sometimes convenient to reverse the order of the terms in the series. This may be donebyputting s=n 2−λ inthefirstform of Pn(x), n even, s=n−1 2−λinthesecondformof Pn(x), n odd, 12.10 Legendre Functions of the Second Kind 809 sothatEqs. (12.212)and(12.213)become Pn(x)=[n/2]summationdisplay s=0(−1)s(2n−2s)! 2ns!(n−s)!(n−2s)!xn−2s, (12.214) where the upper limit s=n/2 (forneven) or (n−1)/2( f o rnodd). This reproduces Eq. (12.8) of Section 12.1, which is obtained directly from the generating function. This agreement with Eq. (12.8) is the reason for the particular choice of normalization in Eqs. (12.212)and(12.213). Qn(x)Functions of the Second Kind Itwillbenoticedthatwehaveusedonly pnfornevenand qnfornodd(becausetheyter- minatedforthischoiceof n).WemaynowdefineasecondsolutionofLegendre’sequation (Fig.12.17)by Qn(x)=(−1)n/2[n/2]!22n n!qn(x) =(−1)s(2s)!! (2s−1)!!q2s(x), forneven,n=2s,(12.215) FIGURE 12.17SecondLegendrefunction, Qn(x), 0≤x<1. 810 Chapter 12 Legendre Functions FIGURE 12.18Second Legendrefunction, Qn(x), x>1. Qn(x)=(−1)(n+1)/2{[(n−1)/2]!}22n−1 n!pn(x) =(−1)s+1(2s)!! (2s+1)!!p2s+1(x), fornodd,n=2s+1.(12.216) This choice of normalizing factors forces Qnto satisfy the same recurrence relations as Pn. This may be verified by substituting Eqs. (12.215) and (12.216) into Eqs. (12.17) and (12.26).Inspectionofthe(series)recurrencerelations(Eqs.(12.207)and(12.209)),thatis, by theCauchy ratio test, shows that Qn(x)willconvergefor −1<x<1.If|x|≥1, these series forms of our second solution diverge. A solution in a series of negative powers of xcan be developed for the region |x|>1 (Fig. 12.18), but we proceed to a closed-form solution that can be used over the entire complex plane (apart from the singular points x=±1 andwithcareoncutlines). Closed-Form Solutions Frequently,aclosedformofthesecondsolution, Qn(z),isdesirable.Thismaybeobtained bythemethoddiscussedinSection9.6.We write Qn(z)=Pn(z)braceleftbigg An+Bnintegraldisplayzdx (1−x2)[Pn(x)]2bracerightbigg , (12.217) inwhichtheconstant Anreplacestheevaluationoftheintegralatthearbitrarylowerlimit. Bothconstants, AnandBn,maybedeterminedfor specialcases. 12.10 Legendre Functions of the Second Kind 811 Forn=0,Eq.(12.217)yields Q0(z)=P0(z)braceleftbigg A0+B0integraldisplayzdx (1−x2)[P0(x)]2bracerightbigg =A0+B01 2ln1+z 1−z =A0+B0parenleftbigg z+z3 3+z5 5+···+z2s+1 2s+1+···parenrightbigg , (12.218) thelastexpressionfollowingfromaMaclaurinexpansionofthelogarithm.Comparingthis withtheseriessolution(Eq. (12.210)),weobtain Q0(z)=q0(z)=z+z3 3+z5 5+···+z2s+1 2s+1+···, (12.219) wehaveA0=0,B0=1.Similarresultsfollowfor n=1.We obtain Q1(z)=zbracketleftbigg A1+B1integraldisplayzdx (1−x2)x2bracketrightbigg =A1z+B1zparenleftbigg1 2ln1+z 1−z−1 zparenrightbigg . (12.220) Expandingin a powerseries and comparingwith Q1(z)=−p1(z),w eh a v e A1=0,B1= 1.Thereforewemaywrite Q0(z)=1 2ln1+z 1−z,Q 1(z)=1 2zln1+z 1−z−1,|z|<1.(12.221) Perhaps the best way of determining the higher-order Qn(z)is to use the recurrence relation(Eq.(12.17)),whichmaybeverifiedforboth x2<1andx2>1bysubstitutingin theseries forms. Thisrecurrencerelationtechniqueyields Q2(z)=1 2P2(z)ln1+z 1−z−3 2P1(z). (12.222) Repeatedapplicationoftherecurrenceformulaleadsto Qn(z)=1 2Pn(z)ln1+z 1−z−2n−1 1·nPn−1(z)−2n−5 3(n−1)Pn−3(z)−···.(12.223) From the form ln [(1+z)/(1−z)]it will be seen that for real zthese expressions hold in the range−1<x<1. If we wish to have closed forms valid outside this range, we need onlyreplace ln1+x 1−xby lnz+1 z−1. When using the latter form, valid for large z, we take the line interval −1≤x≤1 as a cut line.Valuesof Qn(x)onthecutlineare customarilyassignedbytherelation Qn(x)=1 2bracketleftbig Qn(x+i0)+Qn(x−i0)bracketrightbig , (12.224) 812 Chapter 12 Legendre Functions the arithmetic average of approaches from the positive imaginary side and from the nega- tive imaginary side. Note that for z→x>1,z−1→(1−x)e±iπ. The result is that for allz, exceptontherealaxis, −1≤x≤1,wehave Q0(z)=1 2lnz+1 z−1, (12.225) Q1(z)=1 2zlnz+1 z−1−1, (12.226) andso on. Forconvenientreferencesomespecialvaluesof Qn(z)aregiven. 1.Qn(1)=∞,from thelogarithmicterm(Eq. (12.223)). 2.Qn(∞)=0. This is best obtained from a representation of Qn(x)as a series of nega- tivepowersof x,Exercise12.10.4. 3.Qn(−z)=(−1)n+1Qn(z). This follows from the series form. It may also be derived byusingQ0(z),Q1(z)andtherecurrencerelation(Eq. (12.17)). 4.Qn(0)=0,forneven,by(3). 5.Qn(0)=(−1)(n+1)/2{[(n−1)/2]!}2 n!2n−1 =(−1)s+1(2s)!! (2s+1)!!,fornodd,n=2s+1. Thislastresultcomesfrom theseries form(Eq. (12.216))with pn(0)=1. Exercises 12.10.1 Derivetheparityrelationfor Qn(x). 12.10.2 FromEqs. (12.212)and(12.213)showthat (a)P2n(x)=(−1)n 22n−1nsummationdisplay s=0(−1)s(2n+2s−1)! (2s)!(n+s−1)!(n−s)!x2s. (b)P2n+1(x)=(−1)n 22nnsummationdisplay s=0(−1)s(2n+2s+1)! (2s+1)!(n+s)!(n−s)!x2s+1. Checkthenormalizationbyshowingthatonetermofeachseriesagreeswiththecorre- spondingtermofEq. (12.8). 12.10.3 Showthat (a)Q2n(x)=(−1)n22nnsummationdisplay s=0(−1)s(n+s)!(n−s)! (2s+1)!(2n−2s)!x2s+1 +22n∞summationdisplay s=n+1(n+s)!(2s−2n)! (2s+1)!(s−n)!x2s+1,|x|<1. 12.11 Vector Spherical Harmonics 813 (b)Q2n+1(x)=(−1)n+122nnsummationdisplay s=0(−1)s(n+s)!(n−s)! (2s)!(2n−2s+1)!x2s +22n+1∞summationdisplay s=n+1(n+s)!(2s−2n−2)! (2s)!(s−n−1)!x2s,|x|<1. 12.10.4 (a) Startingwiththeassumedform Qn(x)=∞summationdisplay λ=0b−λxk−λ, showthat Qn(x)=b0x−n−1∞summationdisplay s=0(n+s)!(n+2s)!(2n+1)! s!(n!)2(2n+2s+1)!x−2s. (b) Thestandardchoiceof b0is b0=2n(n!)2 (2n+1)!. Show that this choice of bobrings this negative power-series form of Qn(x)into agreementwiththeclosed-formsolutions. 12.10.5 Verify that the Legendre functions of the second kind, Qn(x), satisfy the same recur- rencerelationsas Pn(x), bothfor|x|<1 andfor|x|>1: (2n+1)xQn(x)=(n+1)Qn+1(x)+nQn−1(x), (2n+1)Qn(x)=Q′ n+1(x)−Q′n−1(x). 12.10.6 (a) Usingtherecurrencerelations,prove(independentoftheWronskianrelation)that nbracketleftbig Pn(x)Qn−1(x)−Pn−1(x)Qn(x)bracketrightbig =P1(x)Q0(x)−P0(x)Q1(x). (b) Bydirectsubstitutionshowthattheright-handsideofthis equationequals1. 12.10.7 (a) Write a subroutine that will generate Qn(x)andQ0throughQn−1based on the recurrence relation for these Legendre functions of the second kind. Take xto be within(−1,1)—excludingtheendpoints. Hint.T ak eQ0(x)andQ1(x)tobeknown. (b) Test your subroutine for accuracy by computing Q10(x)and comparing with the valuestabulatedinAMS-55(foracompletereference,seeAdditionalReadingsin Chapter8). 12.11 V ECTOR SPHERICAL HARMONICS Most of our attention in this chapter has been directed toward solving the equations of scalarfields,suchastheelectrostaticpotential.Thiswasdoneprimarilybecausethescalar fields are easier to handle than vector fields. However, with scalar field problems under firmcontrol,moreandmoreattentionis beingpaidtovectorfieldproblems. 814 Chapter 12 Legendre Functions Maxwell’s equations for the vacuum, where the external current and charge densities vanish,leadtothewave(orvectorHelmholtz)equationforthevectorpotential A.Inapar- tialwaveexpansionof Ainsphericalpolarcoordinateswewanttouseangulareigenfunc- tions that are vectors. To this end we write the coordinate unit vectors ˆx,ˆy,ˆzin spherical notation(see Section4.4), ˆe+1=−ˆx+iˆy√ 2,ˆe0=ˆz,ˆe−1=ˆx−iˆy√ 2, (12.227) sothatˆemformasphericaltensorofrank1.Ifwecouplethesphericalharmonicswiththe ˆemto total angular momentum Jusing the relevant Clebsch–Gordan coefficients, we are ledtothevectorsphericalharmonics: YJLMJ(θ,ϕ)=summationdisplay m,MC(L1J|MmMJ)YM L(θ,ϕ)ˆem. (12.228) Itis obviousthattheyobeytheorthogonalityrelations integraldisplay Y∗ JLMJ(θ,ϕ)·YJ′L′M′ J(θ,ϕ)d/Omega1=δJJ′δLL′δMJM′ M. (12.229) GivenJ,theselectionrulesofangularmomentumcouplingtellusthat Lcanonlytakeon the values J+1,J, andJ−1. If we look up the Clebsch–Gordan coefficients and invert Eq.(12.228)weget ˆrYM L(θ,ϕ)=−bracketleftbiggL+1 2L+1bracketrightbigg1/2 YLL+1M+bracketleftbiggL 2L+1bracketrightbigg1/2 YLL−1M, (12.230) displayingthevectorcharacterofthe Yandtheorbitalangularmomentumcontents, L+1 andL−1,ofˆrYM L. Undertheparityoperations(coordinateinversion)thevectorsphericalharmonicstrans- form as YLL+1M(θ′,ϕ′)=(−1)L+1YLL+1M(θ,ϕ), YLL−1M(θ′,ϕ′)=(−1)L+1YLL−1M(θ,ϕ), (12.231) YLLM(θ′,ϕ′)=(−1)LYLLM(θ,ϕ), where θ′=π−θϕ′=π+ϕ. (12.232) The vector spherical harmonics are useful in a further development of the gradient (Eq. (2.46)), divergence (Eq. (2.47)) and curl (Eq. (2.49)) operators in spherical polar co- 12.11 Vector Spherical Harmonics 815 ordinates: ∇bracketleftbig F(r)YM L(θ,ϕ)bracketrightbig =−bracketleftbiggL+1 2L+1bracketrightbigg1/2bracketleftbiggd dr−L rbracketrightbigg FYLL+1M +bracketleftbiggL 2L+1bracketrightbigg1/2bracketleftbiggd dr+L+1 rbracketrightbigg FYLL−1M,(12.233) ∇·bracketleftbig F(r)YLL+1M(θ,ϕ)bracketrightbig =−parenleftbiggL+1 2L+1parenrightbigg1/2bracketleftbiggdF dr+L+2 rFbracketrightbigg YM L(θ,ϕ), (12.234) ∇·bracketleftbig F(r)YLL−1M(θ,ϕ)bracketrightbig =parenleftbiggL 2L+1parenrightbigg1/2bracketleftbiggdF dr−L−1 rFbracketrightbigg YM L(θ,ϕ), (12.235) ∇·bracketleftbig F(r)YLLM(θ,ϕ)bracketrightbig =0, (12.236) ∇×bracketleftbig F(r)YLL+1Mbracketrightbig =ibracketleftbiggL 2L+1bracketrightbigg1/2bracketleftbiggdF dr+L+2 rFbracketrightbigg YLLM,(12.237) ∇×bracketleftbig F(r)YLLMbracketrightbig =iparenleftbiggL 2L+1parenrightbigg1/2bracketleftbiggdF dr−L rFbracketrightbigg YLL+1M +iparenleftbiggL+1 2L+1parenrightbigg1/2bracketleftbiggdF dr+L+1 rFbracketrightbigg YLL−1M,(12.238) ∇×bracketleftbig F(r)YLL−1Mbracketrightbig =ibracketleftbiggL+1 2L+1bracketrightbigg1/2bracketleftbiggdF dr−L−1 rFbracketrightbigg YLLM.(12.239) If we substitute Eq. (12.230) into the radial component ˆr∂/∂rof the gradient operator, for example, we obtain both dF/drterms in Eq. (12.233). For a complete derivation of Eqs. (12.233) to (12.239) we refer to the literature.25These relations play an important roleinthepartialwaveexpansionofclassicalandquantumelectrodynamics. Thedefinitionsofthevectorsphericalharmonicsgivenherearedictatedbyconvenience, primarily in quantum mechanical calculations, in which the angular momentum is a sig- nificant parameter. Further examples of the usefulness and power of the vector spherical harmonics will be found in Blatt and Weisskopf,25in Morse and Feshbach (see General Referencesbook’send),andinJackson’s ClassicalElectrodynamics ,3rded.,NewYork:J. Wiley & Sons (1998), which use vector spherical harmonics in a description of multipole radiationandrelatedelectromagneticproblems. •Vector spherical harmonics are developed from coupling Lunits of orbital angular momentum and 1 unit of spin angular momentum. An extension, coupling Lunits of orbital angular momentum and 2 units of spin angular momentum to form tensor sphericalharmonics,ispresentedbyMathews.26 25E. H. Hill, Theory of vector spherical harmonics, A m .J .P h y s . 22: 211 (1954); also J. M. Blatt and V. Weisskopf, Theoret- ical Nuclear Physics , New York: Wiley (1952). Note that Hill assigns phases in accordance with the Condon–Shortley phase convention (Section 4.4). In Hill’s notation XLM=YLLM,VLM=YLL+1M,WLM=YLL−1M. 26J. Mathews, Gravitational multipole radiation, in In Memoriam (H.P. Robertson, ed.), Philadelphia: Society for Industrial and AppliedMathematics(1963). 816 Chapter 12 Legendre Functions •The major application of tensor spherical harmonics is in the investigation of gravita- tionalradiation. Exercises 12.11.1 Constructthe l=0,m=0 andl=1,m=0 vectorsphericalharmonics. ANS.Y010=−ˆr(4π)−1/2 Y000=0 Y120=−ˆr(2π)−1/2cosθ−ˆθ(8π)−1/2sinθ Y110=ˆϕi(3/8π)1/2sinθ Y100=ˆr(4π)−1/2cosθ−ˆθ(4π)−1/2sinθ. 12.11.2 Verify that the parity of YLL+1Mis(−1)L+1, that of YLLMis(−1)L, and that of YLL−1Mis(−1)L+1.What happenedtothe M-dependenceoftheparity? Hint.ˆrandˆϕhaveoddparity; ˆθhasevenparity(compareExercise2.5.8). 12.11.3 Verifytheorthonormalityof thevectorsphericalharmonics YJLMJ. 12.11.4 InJackson’s ClassicalElectrodynamics ,3rded.,(seeAdditionalReadingsofChapter11 for thereference)defines YLLMbytheequation YLLM(θ,ϕ)=1√L(L+1)LYM L(θ,ϕ), inwhichtheangularmomentumoperator Lis givenby L=−i(r×∇). ShowthatthisdefinitionagreeswithEq.(12.228). 12.11.5 Showthat Lsummationdisplay M=−LY∗ LLM(θ,ϕ)·YLLM(θ,ϕ)=2L+1 4π. Hint. One way is to use Exercise 12.11.4 with Lexpanded in Cartesian coordinates usingtheraisingandloweringoperatorsofSection4.3. 12.11.6 Showthatintegraldisplay YLLM·(ˆr×YLLM)d/Omega1=0. The integrand represents an interference term in electromagnetic radiation that con- tributestoangulardistributionsbutnottototalintensity. AdditionalReadings Hobson, E. W., The Theory of Spherical and Ellipsoidal Harmonics . New York: Chelsea (1955). This is a very complete reference and theclassictext on Legendre polynomials and allrelatedfunctions. Smythe, W. R., Static and DynamicElectricity ,3rd ed.NewYork: McGraw-Hill(1989). Seealsothe referenceslisted in Sections 4.4 and12.9 andatthe endofChapter 13. CHAPTER 13 MORESPECIAL FUNCTIONS Inthischapterweshallstudyfoursetsoforthogonalpolynomials,Hermite,Laguerre,and Chebyshev1of first and second kinds. Although these four sets are of less importance in mathematical physics than are the Bessel and Legendre functions of Chapters 11 and 12, they are used and therefore deserve attention. For example, Hermite polynomials occur in solutionsofthesimpleharmonicoscillatorofquantummechanicsandLaguerrepolynomi- als in wave functions of the hydrogen atom. Because the general mathematical techniques duplicate those of the preceding two chapters, the development of these functions is only outlined. Detailed proofs, along the lines of Chapters 11 and 12, are left to the reader. We express these polynomials and other functions in terms of hypergeometric and confluent hypergeometric functions. To conclude the chapter, we give an introduction to Mathieu functions,whichariseassolutionsofODEsandPDEswithellipticalboundaryconditions. 13.1 H ERMITE FUNCTIONS Generating Functions — Hermite Polynomials TheHermitepolynomials(Fig. 13.1), Hn(x), maybedefinedbythegeneratingfunction2 g(x,t)=e−t2+2tx=∞summationdisplay n=0Hn(x)tn n!. (13.1) 1This is the spelling choice of AMS-55 (for the complete reference see footnote 4 in Chapter 5). However, a variety of names, suchas Tschebyscheff, is encountered. 2Aderivation of this Hermite-generating function is outlined in Exercise 13.1.1. 817 818 Chapter 13 More Special Functions FIGURE 13.1Hermite polynomials. Recurrence Relations Notetheabsenceofasuperscript,whichdistinguishesHermitepolynomialsfromtheunre- latedHankelfunctions.FromthegeneratingfunctionwefindthattheHermitepolynomials satisfytherecurrencerelations Hn+1(x)=2xHn(x)−2nHn−1(x) (13.2) and H′ n(x)=2nHn−1(x). (13.3) Equation(13.2)is obtainedbydifferentiatingthegeneratingfunctionwithrespectto t: ∂g ∂t=(−2t+2x)e−t2+2tx=∞summationdisplay n=0Hn+1(x)tn n! =−2∞summationdisplay n=0Hn(x)tn+1 n!+2x∞summationdisplay n=0Hn(x)tn n!, whichcanberewrittenas ∞summationdisplay n=0tn n!bracketleftbig Hn+1(x)−2xHn(x)+2nHn−1(x)bracketrightbig =0. Becauseeachcoefficientofthispowerseriesvanishes,Eq.(13.2)isestablished.Similarly, differentiationwithrespectto xleadsto ∂g ∂x=2te−t2+2tx=∞summationdisplay n=0H′ n(x)tn n!=2∞summationdisplay n=0Hn(x)tn+1 n!, whichyieldsEq. (13.3) uponshiftingthesummationindex ninthelastsum n+1→n. 13.1 Hermite Functions 819 Table 13.1 HermitePolynomials H0(x)=1 H1(x)=2x H2(x)=4x2−2 H3(x)=8x3−12x H4(x)=16x4−48x2+12 H5(x)=32x5−160x3+120x H6(x)=64x6−480x4+720x2−120 TheMaclaurinexpansionofthegeneratingfunction e−t2+2tx=∞summationdisplay n=0(2tx−t2)n n!=1+parenleftbig 2tx−t2parenrightbig +··· (13.4) givesH0(x)=1 andH1(x)=2x, and then the recursion Eq. (13.2) permits the construc- tion of any Hn(x)desired (integral n). For convenient reference the first several Hermite polynomialsarelistedinTable13.1. Special values of the Hermite polynomials follow from the generating function for x=0: e−t2=∞summationdisplay n=0(−t2)n n!=∞summationdisplay n=0Hn(0)tn n!, thatis, H2n(0)=(−1)n(2n)! n!,H 2n+1(0)=0,n=0,1,.... (13.5) Wealsoobtainfromthegeneratingfunctiontheimportantparityrelation Hn(x)=(−1)nHn(−x) (13.6) bynotingthatEq. (13.1)yields g(−x,−t)=∞summationdisplay n=0Hn(−x)(−t)n n!=g(x,t)=∞summationdisplay n=0Hn(x)tn n!. Alternate Representations TheRodriguesrepresentationof Hn(x)is Hn(x)=(−1)nex2dn dxne−x2. (13.7) Letusshowthisusingmathematicalinductionasfollows. 820 Chapter 13 More Special Functions Example 13.1.1 RODRIGUES REPRESENTATION Werewritethegeneratingfunctionas g(x,t)=ex2e−(t−x)2andnotethat ∂ ∂te−(t−x)2=−∂ ∂xe−(t−x)2. Thisyields ∂g ∂tvextendsinglevextendsinglevextendsinglevextendsingle t=0=(2x−2t)gvextendsinglevextendsinglevextendsingle t=0=2x=H1(x)=−ex2d dxe−x2, whichistheinitial n=1 case.Assumingthecase nofEq.(13.7)asvalid,wenowusethe operatoridentityd dxex2=2xex2+ex2d dxin (−1)n+1ex2dn+1 dxn+1e−x2=(−1)n+1bracketleftbiggd dxex2−2xex2bracketrightbiggdn dxne−x2 =−d dxHn(x)+2xHn(x)=Hn+1(x) toestablishthe n+1 case,withthelastequalityfollowingfromEqs.(13.2) and(13.3). More directly, differentiation of the generating function ntimes with respect to tand thensetting tequaltozeroyields Hn(x)=∂n ∂tnparenleftbig e−t2+2txparenrightbigvextendsinglevextendsinglevextendsingle t=0=(−1)nex2∂n ∂xne−(t−x)2vextendsinglevextendsinglevextendsingle t=0=(−1)nex2dn dxne−x2. /squaresolid Asecondrepresentationmaybeobtainedbyusingthecalculusofresidues(Section7.1). IfwemultiplyEq.(13.1)by t−m−1andintegratearoundtheorigininthecomplex t-plane, onlythetermwith Hm(x)willsurvive: Hm(x)=m! 2πicontintegraldisplay t−m−1e−t2+2txdt. (13.8) Also, from the Maclaurin expansion, Eq. (13.4), we can derive our Hermite polynomial Hn(x)inseriesform:Usingthebinomialexpansionof (2x−t)νandtheindex N=s+ν, e−t2+2tx=∞summationdisplay ν=0tν ν!(2x−t)ν=∞summationdisplay ν=0tν ν!νsummationdisplay s=0parenleftbiggν sparenrightbigg (2x)ν−s(−t)s =∞summationdisplay N=0tN N![N/2]summationdisplay s=0(2x)N−2s(−1)sN! (N−s)!parenleftbiggN−s sparenrightbigg , where[N/2]is the largest integer less than or equal to N/2. Writing the binomial coeffi- cientintermsof factorialsandusingEq. (13.1)weobtain HN(x)=[N/2]summationdisplay s=0(2x)N−2s(−1)sN! s!(N−2s)!. 13.1 Hermite Functions 821 Moreexplicitly,replacing N→n,weha v e Hn(x)=(2x)n−2n! (n−2)!2!(2x)n−2+4n! (n−4)!4!(2x)n−41·3··· =[n/2]summationdisplay s=0(−2)s(2x)n−2sparenleftbiggn 2sparenrightbigg 1·3·5···(2s−1) =[n/2]summationdisplay s=0(−1)s(2x)n−2sn! (n−2s)!s!. (13.9) Thisseries terminatesfor integral nandyieldsourHermitepolynomial. Orthogonality If we substitute the recursion Eq. (13.3) into Eq. (13.2) we can eliminate the index n−1, obtaining Hn+1(x)=2xHn(x)−H′ n(x), which was used already in Example 13.1.1. If we differentiate this recursion relation and substituteEq. (13.3)for theindex n+1 wefind H′ n+1(x)=2(n+1)Hn(x)=2Hn(x)+2xH′ n(x)−H′′ n(x), which can be rearranged to the second-order ODE for Hermite polynomials. Thus, the recurrencerelations(Eqs. (13.2) and(13.3)) leadtothesecond-orderODE H′′ n(x)−2xH′ n(x)+2nHn(x)=0, (13.10) whichisclearly notself-adjoint. ToputtheODEinself-adjointform, followingSection10.1, wemultiplyby exp (−x2), Exercise10.1.2. Thisleadstotheorthogonalityintegral integraldisplay∞ −∞Hm(x)Hn(x)e−x2dx=0,m/negationslash=n, (13.11) withtheweightingfunctionexp (−x2),aconsequenceofputtingtheODEintoself-adjoint form. The interval (−∞,∞)is chosen to obtain the Hermitian operator boundary condi- tions, Section 10.1. It is sometimes convenient to absorb the weighting function into the Hermitepolynomials.We maydefine ϕn(x)=e−x2/2Hn(x), (13.12) withϕn(x)nolongerapolynomial. SubstitutionintoEq. (13.10)yieldsthedifferentialequationfor ϕn(x), ϕ′′ n(x)+parenleftbig 2n+1−x2parenrightbig ϕn(x)=0. (13.13) This is the differential equation for a quantum mechanical, simple harmonic oscillator, which is perhaps the most important physics application of the Hermite polynomials. 822 Chapter 13 More Special Functions Equation (13.13) is self-adjoint, and the solutions ϕn(x)are orthogonal for the interval (−∞<x<∞)witha unitweightingfunction. Theproblemofnormalizingthesefunctionsremains.ProceedingasinSection12.3,we multiplyEq. (13.1) byitself andthenby e−x2.This yields e−x2e−s2+2sxe−t2+2tx=∞summationdisplay m,n=0e−x2Hm(x)Hn(x)smtn m!n!. When we integrate this relation over xfrom−∞to+∞, the cross terms of the double sumdropoutbecauseof theorthogonalityproperty:3 ∞summationdisplay n=0(st)n n!n!integraldisplay∞ −∞e−x2bracketleftbig Hn(x)bracketrightbig2dx=integraldisplay∞ −∞e−x2−s2+2sx−t2+2txdx =integraldisplay∞ −∞e−(x−s−t)2e2stdx =π1/2e2st=π1/2∞summationdisplay n=02n(st)n n!,(13.14) usingEqs. (8.6) and(8.8). Byequatingcoefficientsoflikepowersof st, weobtain integraldisplay∞ −∞e−x2bracketleftbig Hn(x)bracketrightbig2dx=2nπ1/2n!. (13.15) Quantum Mechanical Simple Harmonic Oscillator The following development of Hermite polynomials via simple harmonic oscillator wave functions φn(x)is analogous to the use of the raising and lowering operators for angu- lar momentum operators presented in Section 4.3. This means that we derive the eigen- valuesn+1/2 and eigenfunctions (the Hn(x)) without assuming the development that led to Eq. (13.13). The key aspect of the eigenvalue Eq. (13.13), (d2 dx2−x2)ϕn(x)= −(2n+1)ϕn(x), isthattheHamiltonian −2H≡d2 dx2−x2=parenleftbiggd dx−xparenrightbiggparenleftbiggd dx+xparenrightbigg +bracketleftbigg x,d dxbracketrightbigg (13.16) almostfactorizes.Usingnaively a2−b2=(a−b)(a+b),thebasiccommutator [px,x]= ¯h/iof quantum mechanics (with momentum px=(¯h/i)d/dx ) enters as a correction in Eq. (13.16). (Because pxis Hermitian, d/dxis anti-Hermitian, (d/dx)†=−d/dx.) This commutator can be evaluated as follows. Imagine the differential operator d/dxacts on a wavefunction ϕ(x)totheright,asinEq. (13.13), so d dx(xϕ)=xd dxϕ+ϕ, (13.17) 3The cross terms (m/negationslash=n)may be left in, if desired. Then, when the coefficients of sαtβare equated, the orthogonality will be apparent. 13.1 Hermite Functions 823 bytheproductrule.Droppingthewavefunction ϕfromEq.(13.17),werewriteEq.(13.17) as d dxx−xd dx≡bracketleftbiggd dx,xbracketrightbigg =1, (13.18) a constant, and then verify Eq. (13.16) directly by expanding the product of operators. The product form of Eq. (13.16), up to the constant commutator, suggests introducing the non-Hermitianoperators ˆa†≡1√ 2parenleftbigg x−d dxparenrightbigg ,ˆa≡1√ 2parenleftbigg x+d dxparenrightbigg , (13.19) with(ˆa)†=ˆa†,whichareadjointsofeachother.Theyobeythecommutationrelations bracketleftbig ˆa,ˆa†bracketrightbig =bracketleftbiggd dx,xbracketrightbigg =1,[ˆa,ˆa]=0=bracketleftbig ˆa†,ˆa†bracketrightbig , (13.20) which are characteristic of these operators and straightforward to derive from Eq. (13.18) andbracketleftbiggd dx,d dxbracketrightbigg =0=[x,x]andbracketleftbigg x,d dxbracketrightbigg =−bracketleftbiggd dx,xbracketrightbigg . ReturningtoEq. (13.16)andusingEq. (13.19)werewritetheHamiltonianas H=ˆa†ˆa+1 2=ˆa†ˆa+1 2parenleftbig ˆa†ˆa+ˆaˆa†parenrightbig =1 2parenleftbig ˆa†ˆa+ˆaˆa†parenrightbig (13.21) and introduce the Hermitian number operator N=ˆa†ˆaso that H=N+1/2.Let|n/angbracketrightbe aneigenfunctionof H, H|n/angbracketright=λn|n/angbracketright, whose eigenvalue λnis unknown at this point. Now we prove the key property that Nhas nonnegativeintegereigenvalues N|n/angbracketright=parenleftbigg λn−1 2parenrightbigg |n/angbracketright=n|n/angbracketright,n=0,1,2..., (13.22) thatis,λn=n+1/2.Sinceˆa|n/angbracketrightiscomplexconjugateto /angbracketleftn|ˆa†,thenormalizationintegral /angbracketleftn|ˆa†ˆa|n/angbracketright≥0 andis finite.From parenleftbig ˆa|n/angbracketrightparenrightbig†ˆa|n/angbracketright=/angbracketleftn|ˆa†ˆa|n/angbracketright=parenleftbigg λn−1 2parenrightbigg ≥0 (13.23) wesee that Nhasnonnegativeeigenvalues. We now show that if ˆa|n/angbracketrightis nonzero it is an eigenfunction with eigenvalue λn−1= λn−1. After normalizing ˆa|n/angbracketright, this state is designated |n−1/angbracketright.T h i si sp r o v e db yt h e commutationrelations bracketleftbig N,ˆa†bracketrightbig =ˆa†,[N,ˆa]=−ˆa, (13.24) 824 Chapter 13 More Special Functions whichfollowfromEq.(13.20).Thesecommutationrelationscharacterize Nasthenumber operator.Toseethis,wedeterminetheeigenvalueof Nforthestatesˆa†|n/angbracketrightandˆa|n/angbracketright.Using ˆaˆa†=N+[ˆa,ˆa†]=N+1,wefindthat Nparenleftbig ˆa†|n/angbracketrightparenrightbig =ˆa†parenleftbig ˆaˆa†parenrightbig |n/angbracketright=ˆa†parenleftbigbracketleftbig ˆa,ˆa†bracketrightbig +Nparenrightbig |n/angbracketright =ˆa†(N+1)|n/angbracketright=parenleftbigg λn+1 2parenrightbigg ˆa†|n/angbracketright=(n+1)ˆa†|n/angbracketright,(13.25) Nparenleftbig ˆa|n/angbracketrightparenrightbig =parenleftbig ˆaˆa†−1parenrightbig ˆa|n/angbracketright=ˆa(N−1)|n/angbracketright=(n−1)ˆa|n/angbracketright. Inotherwords, Nactingonˆa†|n/angbracketrightshowsthatˆa†hasraisedtheeigenvalue ncorresponding to|n/angbracketrightbyoneunit,whenceitsname raising,orcreation,operator.Applyingˆa†repeatedly, we can reach all higher states. There is no upper limit to the sequence of eigenvalues. Similarly,ˆalowers the eigenvalue nby one unit; hence it is a lowering (orannihilation ) operator because.Therefore, ˆa†|n/angbracketright∼|n+1/angbracketright,ˆa|n/angbracketright∼|n−1/angbracketright. (13.26) Applyingˆarepeatedly,wecanreachthelowest,orground,state |0/angbracketrightwitheigenvalue λ0.W e cannotsteplowerbecause λ0≥1/2.Thereforeˆa|0/angbracketright≡0,suggestingweconstruct ψ0=|0/angbracketright from the(factored) first-order ODE √ 2ˆaψ0=parenleftbiggd dx+xparenrightbigg ψ0=0. (13.27) Integrating ψ′ 0 ψ0=−x, (13.28) weobtain lnψ0=−1 2x2+lnc0, (13.29) wherec0is anintegrationconstant.Thesolution, ψ0(x)=c0e−x2/2, (13.30) is a Gaussian that can be normalized, with c0=π−1/4using the error integral, Eqs. (8.6) and(8.8). Substituting ψ0intoEq.(13.13) wefind H|0/angbracketright=parenleftbigg ˆa†ˆa+1 2parenrightbigg |0/angbracketright=1 2|0/angbracketright, (13.31) so its energy eigenvalue is λ0=1/2 and its number eigenvalue is n=0, confirming the notation|0/angbracketright.Applyingˆa†repeatedlyto ψ0=|0/angbracketright,allothereigenvaluesareconfirmedtobe λn=n+1/2,provingEq.(13.13).ThenormalizationsneededforEq.(13.26)followfrom Eqs. (13.25)and(13.23)and /angbracketleftn|ˆaˆa†|n/angbracketright=/angbracketleftn|ˆa†ˆa+1|n/angbracketright=n+1, (13.32) 13.1 Hermite Functions 825 showing √ n+1|n+1/angbracketright=ˆa†|n/angbracketright,√n|n−1/angbracketright=ˆa|n/angbracketright. (13.33) Thus, the excited-state wave functions, ψ1,ψ2, and so on, are generated by the raising operator |1/angbracketright=ˆa†|0/angbracketright=1√ 2parenleftbigg x−d dxparenrightbigg ψ0(x)=x√ 2 π1/4e−x2/2, (13.34) yielding(andleadingtoupcomingEq. (13.38)) ψn(x)=NnHn(x)e−x2/2,N n≡π−1/4parenleftbig 2nn!parenrightbig−1/2, (13.35) whereHnaretheHermitepolynomials(Fig.13.2). Asshown,theHermitepolynomialsareusedinanalyzingthequantummechanicalsim- ple harmonic oscillator. For a potential energy V=1 2Kz2=1 2mω2z2(forceF=−∇V= −Kzˆz),theSchrödingerwaveequationis −¯h2 2m∇2/Psi1(z)+1 2Kz2/Psi1(z)=E/Psi1(z). (13.36) Ouroscillatingparticlehasmass mandtotalenergy E.By useoftheabbreviations x=αzwith α4=mK ¯h2=m2ω2 ¯h2, λ=2E ¯hparenleftbiggm Kparenrightbigg1/2 =2E ¯hω,(13.37) in which ωis the angular frequency of the corresponding classical oscillator, Eq. (13.36) becomes(with /Psi1(z)=/Psi1(x/α)=ψ(x)) d2ψ(x) dx2+parenleftbig λ−x2parenrightbig ψ(x)=0. (13.38) Thisis Eq.(13.13) with λ=2n+1.Hence(Fig.13.2), ψn(x)=2−n/2π−1/4(n!)−1/2e−x2/2Hn(x), (normalized ). (13.39) Alternatively, the requirement that nbe an integer is dictated by the boundary conditions ofthequantummechanicalsystem, lim z→±∞/Psi1(z)=0. Specifically,if n→ν,notaninteger,apower-seriessolutionofEq.(13.13)(Exercise9.5.6) shows that Hν(x)will behave as xνex2for large x. The functions ψν(x)and/Psi1ν(z)will thereforeblowupatinfinity,anditwillbeimpossibletonormalizethewavefunction /Psi1(z). Withthisrequirement,theenergy Ebecomes E=parenleftbigg n+1 2parenrightbigg ¯hω. (13.40) 826 Chapter 13 More Special Functions FIGURE 13.2Quantummechanical oscillatorwavefunctions:The heavybaronthe x-axisindicatesthe allowedrangeoftheclassical oscillatorwiththesametotalenergy. Asnrangesoverintegralvalues (n≥0),weseethattheenergyisquantizedandthatthere isaminimumorzeropointenergy Emin=1 2¯hω. (13.41) This zero point energy is an aspect of the uncertainty principle, a genuine quantum phe- nomenon. 13.1 Hermite Functions 827 In quantum mechanical problems, particularly in molecular spectroscopy, a number of integralsoftheform integraldisplay∞ −∞xre−x2Hn(x)Hm(x)dx areneeded.Examplesfor r=1andr=2(withn=m)areincludedintheexercisesatthe endofthissection.AlargenumberofotherexamplesarecontainedinWilson,Decius,and Cross.4 In the dynamics and spectroscopy of molecules in the Born–Oppenheimer approxima- tion, the motion of a molecule is separated into electronic, vibrational and rotational mo- tion. Each vibrating atom contributes to a matrix element two Hermite polynomials, its initial state and another one to its final state. Thus, integrals of products of Hermite poly- nomialsareneeded. Example 13.1.2 THREEFOLD HERMITE FORMULA Wewanttocalculatethefollowingintegralinvolving m=3 Hermitepolynomials: I3≡integraldisplay∞ −∞e−x2HN1(x)HN2(x)HN3(x)dx, (13.42) whereNi≥0 are integers. The formula (due to E. C. Titchmarsh, J. Lond. Math. Soc. 23: 15(1948),seeGradshteynandRyzhik,p.838,intheAdditionalReadings)generalizesthe m=2 case needed for the orthogonality and normalization of Hermite polynomials. To derive it, we start with the product of three generating functions of Hermite polynomials, multiplyby e−x2, andintegrateover xinordertogenerate I3: Z3≡integraldisplay∞ −∞e−x23productdisplay j=1e2xtj−t2 jdx=integraldisplay∞ −∞e−(summationtext3 j=1tj−x)2+2(t1t2+t1t3+t2t3)dx =√πe2(t1t2+t1t3+t2t3). (13.43) The last equality follows from substituting y=x−summationtext jtjand using the error integralintegraltext∞ −∞e−y2dy=√π, Eqs. (8.6) and (8.8). Expanding the generating functions in terms of Hermitepolynomialsweobtain Z3=∞summationdisplay N1,N2,N3=0tN1 1tN2 2tN3 3 N1!N2!N3!integraldisplay∞ −∞e−x2HN1(x)HN2(x)HN3(x)dx =√π∞summationdisplay N=02N N!(t1t2+t1t3+t2t3)N =√π∞summationdisplay N=02N N!summationdisplay 0≤ni≤N,summationtext ini=NN! n1!n2!n3!(t1t2)n1(t1t3)n2(t2t3)n3, 4E.B.Wilson,Jr.,J.C.Decius,andP.C.Cross, MolecularVibrations ,NewYork:McGraw-Hill(1955),reprintedDover(1980). 828 Chapter 13 More Special Functions usingthepolynomialexpansion parenleftbiggmsummationdisplay j=1ajparenrightbiggN =summationdisplay 0≤ni≤mN! n1!···nm!an1 1···anmm. Thepowersoftheforegoing tjtkbecome (t1t2)n1(t1t3)n2(t2t3)n3=tN1 1tN2 2tN3 3; N1=n1+n2,N 2=n1+n3,N 3=n2+n3. Thatis, from 2N=2(n1+n2+n3)=N1+N2+N3 therefollows 2N=2n1+2N3=2n2+2N2=2n3+2N1, soweobtain n1=N−N3,n 2=N−N2,n 3=N−N1. Theniare all fixed (making this case special and easy) because the Niare fixed, and 2N=3summationtext i=1Ni, withN≥0 an integer by parity. Hence, upon comparing the foregoing like t1t2t3powers, I3=√π2NN1!N2!N3! (N−N1)!(N−N2)!(N−N3)!, (13.44) which is the desired formula. If we orderN1≥N2≥N3≥0, thenn1≥n2≥n3≥0 follows,beingequivalentto N−N3≥N−N2≥N−N1≥0,whichoccurinthedenom- inatorsofthefactorialsof I3. /squaresolid Example 13.1.3 DIRECT EXPANSION OF PRODUCTS OF HERMITE POLYNOMIALS Inanalternativeapproach,wenowstartagainfromthegeneratingfunctionidentity ∞summationdisplay N1,N2=0HN1(x)HN2(x)tN1 1 N1!tN2 2 N2!=e2x(t1+t2)−t2 1−t2 2=e2x(t1+t2)−(t1+t2)2·e2t1t2 =∞summationdisplay N=0HN(x)(t1+t2)N N!∞summationdisplay ν=0(2t1t2)ν ν!. 13.1 Hermite Functions 829 Usingthebinomialexpansionandthencomparinglikepowersof t1t2weextractanidentity duetoE. Feldheim( J. Lond.Math.Soc. 13:22(1938)): HN1(x)HN2(x)=min(N1,N2)summationdisplay ν=0HN1+N2−2νN1!N2!2ν ν!(N1+N2−2ν)!parenleftbiggN1+N2−2ν N1−νparenrightbigg =summationdisplay 0≤ν≤min(N1,N2)HN1+N2−2ν2νν!parenleftbiggN1 νparenrightbiggparenleftbiggN2 νparenrightbigg . (13.45) Forν=0 thecoefficientof HN1+N2isobviouslyunity.Specialcases, suchas H2 1=H2+2,H 1H2=H3+4H1,H2 2=H4+8H2+8,H 1H3=H4+6H2, canbederivedfrom Table13.1andagreewiththegeneraltwofoldproductformula. This compact formula can be generalized to products of mHermite polynomials, and thisinturnyieldsa newclosedformresultfor theintegral Im. Let us beginwitha newresultfor I4containinga productof four Hermitepolynomials. InsertingtheFeldheimidentityfor HN1HN2andHN3HN4andusingorthogonality integraldisplay∞ −∞e−x2HN1HN2dx=√π2N1N1!δN1N2 fortheremainingproductof twoHermitepolynomialsyields I4=integraldisplay∞ −∞e−x2HN1HN2HN3HN4dx =summationdisplay 0≤µ≤min(N1,N2);0≤ν≤min(N3,N4)2µµ! ·parenleftbiggN1 µparenrightbiggparenleftbiggN2 µparenrightbigg 2νν!parenleftbiggN3 νparenrightbiggparenleftbiggN4 νparenrightbiggintegraldisplay∞ −∞e−x2HN1+N2−2µHN3+N4−2νdx =N4summationdisplay ν=0√π2M(N3+N4−2ν)!N1!N2!N3!N4! (M−N3−N4−ν)!(M−N1+ν)!(M−N2+ν)!(N3−ν)!(N4−ν)!ν!. (13.46) Hereweusethenotation M=(N1+N2+N3+N4)/2andwritethebinomialcoefficients explicitly,so 1 2(N1+N2−N3−N4)=M−N3−N4, 1 2(N1−N2+N3+N4)=M−N2, 1 2(N2−N1+N3+N4)=M−N1. From orthogonality we have µ=(N1+N2−N3−N4)/2+ν. The upper limit ofνis min(N3,N4,M−N1,M−N2)=min(N4,M−N1)and the lower limit is max(0,N3+N4−M)=0,if weorder N1≥N2≥N3≥N4. 830 Chapter 13 More Special Functions Nowwereturntotheproductexpansionof mHermitepolynomialsandthecorrespond- ingnewresultfrom itfor Im. Weprovea generalizedFeldheimidentity, HN1(x)···HNm(x)=summationdisplay ν1,...,νm−1HM(x)aν1,...,νm−1, (13.47) where M=m−1summationdisplay i=1(Ni−2νi)+Nm, by mathematical induction. Multiplying this by HNm+1and using the Feldheim identity, we end up with the same formula for m+1 Hermite polynomials, including the recursion relation aν1,...,νm=aν1,...,νm−12νmνm!parenleftbiggNm+1 νmparenrightbiggparenleftbiggsummationtextm−1 i=1(Ni−2νi)+Nm+1 νmparenrightbigg . Its solutionis aν1,...,νm−1=m−1productdisplay i=1parenleftbiggNi+1 νiparenrightbiggparenleftbiggsummationtexti−1 j=1(Nj−2νj)+Ni νiparenrightbigg 2νiνi!.(13.48) Thelimitsof thesummationindicesare 0≤ν1≤min(N1,N2),0≤ν2≤min(N3,N1+N2−2ν1),..., 0≤νm−1≤minparenleftbigg Nm,m−2summationdisplay i=1(Ni−2νi)+Nm−1parenrightbigg .(13.49) We now apply this generalized Feldheim identity, with indices ordered as N1≥N2≥ ···≥Nm,t oIm, grouping HN2···HNmtogether and using orthogonality for the re- maining product of two Hermite polynomials HN1Hsummationtextm−1 i=2(Ni−2νi)+Nm. This yields N1= summationtextm−1 i=2(Ni−2νi)+Nm,fixingνm−1, and Im=√π2N1N1!summationdisplay ν2,...,νm−1m−1productdisplay i=2parenleftbiggNi+1 νiparenrightbiggparenleftbiggsummationtexti−1 j=2(Nj−2νj)+Ni νiparenrightbigg νi!2νi,(13.50) wherethelimitsonthesummationindicesare 0≤ν2≤min(N3,N2),..., 0≤νm−1≤minparenleftbigg Nm,m−2summationdisplay i=2(Ni−2νi)+Nm−1parenrightbigg .(13.51) /squaresolid 13.1 Hermite Functions 831 Example 13.1.4 APPLICATIONS OF THE PRODUCT FORMULAS Tochecktheexpression Imform=3,wenotethatthesumsummationtexti−1 j=2withi−1=m−2=1 in the second binomial coefficient in Imis empty (that is, zero), so only Ni=Nm−1=N2 remains. Also, with νm−2=ν1the sum over the νiis that over ν2, which is fixed by the constraint on the summation index ν2:N1=N2−2ν2+N3. Henceν2=(N2+N3− N1)/2=N−N1, withN=(N1+N2+N3)/2. That is, only the product remains in Im. Thegeneralformulafor Imthereforegives I3=√π2N1N1!parenleftbiggN3 ν2parenrightbiggparenleftbiggN2 ν2parenrightbigg ν2!2ν2=√π2NN1!N2!N3! (N−N1)!(N−N2)!(N−N3)!, which agrees with our earlier result of Example 13.1.2. The last expression is based on the following observations. The power of 2 has the exponent N1+ν2=N. The factorials from the binomial coefficients are N3−ν2=(N1+N3−N2)/2=N−N2,N2−ν2= (N1+N2−N3)/2=N−N3. Next let us consider m=4, where we do not order the Hermite indices Nias yet. The reason is that the general Imexpression was derived with a different grouping of the Her- mite polynomials than the separate calculation of I4with which we compare. That is why we’ll have to permute the indices to get the earlier result for I4. That is a general conclu- sion: Different groupings of the Hermite polynomials just give different permutations of theHermiteindicesinthegeneralresult. We havetwosummationsover ν2andνm−1=ν3,whichisfixedbytheconstraint N1= N2−2ν2+N3−2ν3+N4. Hence ν3=1 2(N2+N3+N4−N1)−ν2=M−N1−ν2 withM=1 2(N1+N2+N3+N4).Theexponentof 2 is N1+ν2+ν3=M.Thereforefor m=4t h eImformulagives I4=√π2N1N1!summationdisplay ν2≥0parenleftbiggN3 ν2parenrightbiggparenleftbiggN4 ν3parenrightbiggparenleftbiggN2 ν2parenrightbiggparenleftbiggN2−2ν2+N3 ν3parenrightbigg ν2!2ν2ν3!2ν3 =summationdisplay ν2≥0√π2MN1!N2!N3!N4!(N2−2ν2+N3)! ν2!ν3!(N2−ν2)!(N3−ν2)!(N4−ν3)!(N2+N3−2ν2−ν3)! =summationdisplay ν2≥0√π2MN1!N2!N3!N4!(N2+N3−2ν2)! (N2−ν2)!(N3−ν2)!(N4−ν3)!ν2!ν3!(N2+N3−2ν2−ν3)! =summationdisplay ν2≥0√π2MN1!N2!N3!N4!(N2+N3−2ν2)! ν2!(M−N1−ν2)!(N3−ν2)!(N2−ν2)!(M−N2−N3+ν2)!(M−N4−ν2)!. Inthelastexpressionwehavesubstituted ν3andused N4−ν3=(N1−N2−N3+N4)+ν2=M−N2−N3+ν2, N2+N3−2ν2−ν3=N1+N2+N3−N4 2−ν2=M−N4−ν2. 832 Chapter 13 More Special Functions The upper limit is ν2≤min(N2,N3,M−N1,M−N4), and the lower limit is ν2≥ max(0,N2+N3−M). If we make the permutation N2↔N4,ν2→ν, then our previ- ousI4resultisobtainedwithupperlimit ν≤min(N4,M−N1)=1 2(N2+N3+N4−N1) and lower limit ν≥max(0,N3+N4−M)=0 because N3+N4−N1−N2≤0f o r N1≥N2≥N3≥N4≥0. /squaresolid The Hermite polynomial product formula also applies to products of simple harmonic oscillatorwavefunctions,integraltext∞ −∞e−mx2/2HN1(x)···HNm(x)dx,withadifferentexponential weight function. To evaluate such integrals we use the generalized Feldheim identity for HN2···HNmin conjunction with the integral (see Gradshteyn and Ryzhik, p. 845, in the AdditionalReadings), integraldisplay∞ −∞e−a2x2Hm(x)Hn(x)dx=1 2parenleftbigg2 aparenrightbiggm+n+1parenleftbig 1−a2parenrightbig(m+n)/2Ŵparenleftbiggm+n+1 2parenrightbigg ·2F1parenleftbigg −m,−n;1−m−n 2;a2 2(a2−1)parenrightbigg , instead of the standard orthogonality integral for the remaining product of two Hermite polynomials.Herethehypergeometricfunctionis thefinitesum 2F1parenleftbigg −m,−n;1−m−n 2;a2 2(a2−1)parenrightbigg =min(m,n)−1summationdisplay ν=0(−m)ν(−n)ν ν!(1−m−n 2)νparenleftbigga2 2(a2−1)parenrightbiggν with(−m)ν=(−m)(1−m)···(ν−1−m)and(−m)0≡1.Thisyieldsaresultsimilarto Imbutsomewhatmorecomplicated. The oscillator potential has also been employed extensively in calculations of nuclear structure(nuclearshellmodel)quarkmodelsof hadronsandthenuclearforce. There is a second independent solution of Eq. (13.13). This Hermite function of the second kind is an infinite series (Sections 9.5 and 9.6) and is of no physical interest, at leastnotyet. Exercises 13.1.1 Assume the Hermite polynomials are known to be solutions of the differential equa- tion (13.13). From this the recurrence relation, Eq. (13.3), and the values of Hn(0)are alsoknown. (a) Assumetheexistenceofageneratingfunction g(x,t)=∞summationdisplay n=0Hn(x)tn n!. (b) Differentiate g(x,t)with respect to xand using the recurrence relation develop a first-orderPDEfor g(x,t). (c) Integratewithrespectto x,holding tfixed. (d) Evaluate g(0,t)usingEq. (13.5). Finally,showthat g(x,t)=expparenleftbig −t2+2txparenrightbig . 13.1 Hermite Functions 833 13.1.2 In developing the properties of the Hermite polynomials, start at a number of different points,suchas: 1. Hermite’sODE,Eq. (13.13), 2. Rodrigues’formula, Eq.(13.7), 3. Integralrepresentation,Eq. (13.8), 4. Generatingfunction,Eq. (13.1), 5. Gram–Schmidt construction of a complete set of orthogonal polynomials over (−∞,∞)withaweightingfactor of exp (−x2),Section10.3. Outlinehowyoucangofrom anyoneofthesestartingpointstoalltheotherpoints. 13.1.3 Provethat parenleftbigg 2x−d dxparenrightbiggn 1=Hn(x). Hint.Checkoutthefirstcoupleof examplesandthenusemathematicalinduction. 13.1.4 Provethat vextendsinglevextendsingleHn(x)vextendsinglevextendsingle≤vextendsinglevextendsingleHn(ix)vextendsinglevextendsingle. 13.1.5 Rewritetheseries formof Hn(x), Eq. (13.9), asan ascending powerseries. ANS.H2n(x)=(−1)nnsummationdisplay s=0(−1)2s(2x)2s(2n)! (2s)!(n−s)!, H2n+1(x)=(−1)nnsummationdisplay s=0(−1)s(2x)2s+1(2n+1)! (2s+1)!(n−s)!. 13.1.6 (a) Expand x2rinaseriesof even-orderHermitepolynomials. (b) Expand x2r+1inaseries ofodd-orderHermitepolynomials. ANS.(a) x2r=(2r)! 22rrsummationdisplay n=0H2n(x) (2n)!(r−n)! (b)x2r+1=(2r+1)! 22r+1rsummationdisplay n=0H2n+1(x) (2n+1)!(r−n)!,r=0,1,2,.... Hint.UseaRodriguesrepresentationandintegratebyparts. 13.1.7 Showthat (a)integraldisplay∞ −∞Hn(x)expbracketleftbigg −x2 2bracketrightbigg dx=braceleftbigg2πn!/(n/2)!,neven 0,n odd. (b)integraldisplay∞ −∞xHn(x)expbracketleftbigg −x2 2bracketrightbigg dx=  0,n even 2π(n+1)! ((n+1)/2)!,nodd. 834 Chapter 13 More Special Functions 13.1.8 Showthat integraldisplay∞ −∞xme−x2Hn(x)dx=0formaninteger ,0≤m≤n−1. 13.1.9 Thetransitionprobabilitybetweentwooscillatorstates mandndependson integraldisplay∞ −∞xe−x2Hn(x)Hm(x)dx. Show that this integral equals π1/22n−1n!δm,n−1+π1/22n(n+1)!δm,n+1. This result shows that such transitions can occur only between states of adjacent energy levels, m=n±1. Hint. Multiply the generating function (Eq. (13.1)) by itself using two different sets of variables (x,s)and(x,t). Alternatively, the factor xmay be eliminated by the recur- rencerelation,Eq. (13.2). 13.1.10 Showthat integraldisplay∞ −∞x2e−x2Hn(x)Hn(x)dx=π1/22nn!parenleftbigg n+1 2parenrightbigg . Thisintegraloccursinthecalculationofthemean-squaredisplacementofourquantum oscillator. Hint.Usetherecurrencerelation,Eq.(13.2), andtheorthogonalityintegral. 13.1.11 Evaluate integraldisplay∞ −∞x2e−x2Hn(x)Hm(x)dx intermsof nandmandappropriateKroneckerdeltafunctions. ANS. 2n−1π1/2(2n+1)n!δnm+2nπ1/2(n+2)!δn+2,m+2n−2π1/2n!δn−2,m. 13.1.12 Showthat integraldisplay∞ −∞xre−x2Hn(x)Hn+p(x)dx=braceleftbigg0,p >r 2nπ1/2(n+r)!,p=r, withn,p, andrnonnegativeintegers. Hint.Usetherecurrencerelation,Eq.(13.2), ptimes. 13.1.13 (a) Using the Cauchy integral formula, develop an integral representation of Hn(x) basedonEq. (13.1)withthecontourenclosingthepoint z=−x. ANS.Hn(x)=n! 2πiex2contintegraldisplaye−z2 (z+x)n+1dz. (b) Showbydirectsubstitutionthatthisresultsatisfies theHermiteequation. 13.1 Hermite Functions 835 13.1.14 With ψn(x)=e−x2/2Hn(x) (2nn!π1/2)1/2, verifythat ˆanψn(x)=1√ 2parenleftbigg x+d dxparenrightbigg ψn(x)=n1/2ψn−1(x), ˆa† nψn(x)=1√ 2parenleftbigg x−d dxparenrightbigg ψn(x)=(n+1)1/2ψn+1(x). Note. The usual quantum mechanical operator approach establishes these raising and loweringpropertiesbeforetheform of ψn(x)is known. 13.1.15 (a) Verifytheoperatoridentity x−d dx=−expbracketleftbiggx2 2bracketrightbiggd dxexpbracketleftbigg −x2 2bracketrightbigg . (b) Thenormalizedsimpleharmonicoscillatorwavefunctionis ψn(x)=parenleftbig π1/22nn!parenrightbig−1/2expbracketleftbigg −x2 2bracketrightbigg Hn(x). Showthatthismaybewrittenas ψn(x)=parenleftbig π1/22nn!parenrightbig−1/2parenleftbigg x−d dxparenrightbiggn expbracketleftbigg −x2 2bracketrightbigg . Note. This corresponds to an n-fold application of the raising operator of Exer- cise13.1.14. 13.1.16 (a) ShowthatthesimpleoscillatorHamiltonian(from Eq.(13.38)) maybewrittenas H=−1 2d2 dx2+1 2x2=1 2parenleftbig ˆaˆa†+ˆa†ˆaparenrightbig . Hint.Express Einunitsof¯hω. (b) Usingthecreation–annihilationoperatorformulationofpart(a), showthat Hψ(x)=parenleftbigg n+1 2parenrightbigg ψ(x). This means the energy eigenvalues are E=(n+1 2)(¯hω), in agreement with Eq.(13.40). 13.1.17 Write a program that will generate the coefficients as, in the polynomial form of the Hermitepolynomial Hn(x)=summationtextn s=0asxs. 13.1.18 Afunction f(x)is expandedinaHermiteseries: f(x)=∞summationdisplay n=0anHn(x). 836 Chapter 13 More Special Functions From the orthogonality and normalization of the Hermite polynomials the coefficient anis givenby an=1 2nπ1/2n!integraldisplay∞ −∞f(x)Hn(x)e−x2dx. Forf(x)=x8determinetheHermitecoefficients anbytheGauss–Hermitequadrature. Check your coefficients against AMS-55, Table 22.12 (for the reference see footnote 4 inChapter5ortheGeneralReferencesatbook’send). 13.1.19 (a) In analogy with Exercise 12.2.13, set up the matrix of even Hermite polynomial coefficientsthatwilltransformanevenHermiteseries intoanevenpowerseries: B= 1−212··· 04−48··· 001 6··· .........··· . ExtendBtohandleanevenpolynomialseriesthrough H8(x). (b) Invert your matrix to obtain matrix A, which will transform an even power series (through x8) into a series of even Hermite polynomials. Check the elements of A against those listed in AMS-55 (Table 22.12, in the General References at book’s end). (c) Finally, using matrix multiplication, determine the Hermite series equivalent to f(x)=x8. 13.1.20 Write asubroutinethatwilltransform afinitepowerseries,summationtextN n=0anxn, intoaHermite series,summationtextN n=0bnHn(x). Usetherecurrencerelation,Eq.(13.2). Note.BothExercises13.1.19and13.1.20arefasterandmoreaccuratethantheGaussian quadrature,Exercise13.1.18,if f(x)isavailableasapowerseries. 13.1.21 Write asubroutineforevaluatingHermitepolynomialmatrixelementsof theform Mpqr=integraldisplay∞ −∞Hp(x)Hq(x)xre−x2dx, using the 10-point Gauss–Hermite quadrature (for p+q+r≤19). Include a parity check and set equal to zero the integrals with odd-parity integrand. Also, check to see ifris in the range |p−q|≤r. Otherwise Mpqr=0. Check your results against the specificcaseslistedinExercises13.1.9,13.1.10, 13.1.11,and13.1.12. 13.1.22 Calculateandtabulatethenormalizedlinearoscillatorwavefunctions ψn(x)=2−n/2π−1/4(n!)−1/2Hn(x)expparenleftbigg −x2 2parenrightbigg forx=0.0(0.1)5.0 andn=0(1)5.If aplottingroutineisavailable,plotyourresults. 13.1.23 Evaluateintegraltext∞ −∞e−2x2HN1(x)···HN4(x)dxinclosedform. Hint.integraltext∞ −∞e−2x2HN1(x)HN2(x)HN3(x)dx=1 π2(N1+N2+N3−1)/2·Ŵ(s−N1)Ŵ(s−N2) ·Ŵ(s−N3),s=(N1+N2+N3+1)/2o rintegraltext∞ −∞e−2x2HN1(x)HN2(x)dx= (−1)(N1+N2−1)/22(N1+N2−1)/2·Ŵ((N1+N2+1)/2)may be helpful. Prove these for- mulas(seeGradshteynandRyzhik,no.7.375onp.844,intheAdditionalReadings). 13.2 Laguerre Functions 837 13.2 L AGUERRE FUNCTIONS Differential Equation — Laguerre Polynomials If we start with the appropriate generating function, it is possible to develop the Laguerre polynomialsinanalogywiththeHermitepolynomials.Alternatively,aseriessolutionmay be developed by the methods of Section 9.5. Instead, to illustrate a different technique, let usstartwithLaguerre’sODEandobtainasolutionintheformofacontourintegral,aswe didwiththeintegralrepresentationforthemodifiedBesselfunction Kν(x)(Section11.6). Fromthis integralrepresentationageneratingfunctionwillbederived. Laguerre’s ODE (which derives from the radial ODE of Schrödinger’s PDE for the hy- drogenatom)is xy′′(x)+(1−x)y′(x)+ny(x)=0. (13.52) We shall attempt to represent y, or rather yn, sinceywill depend on the parameter n, anonnegativeinteger,bythecontourintegral yn(x)=1 2πicontintegraldisplaye−xz/(1−z) (1−z)zn+1dz (13.53a) anddemonstratethatitsatisfiesLaguerre’sODE.Thecontourincludestheoriginbutdoes notenclosethepoint z=1.By differentiatingtheexponentialinEq. (13.53a)weobtain y′ n(x)=−1 2πicontintegraldisplaye−xz/(1−z) (1−z)2zndz, (13.53b) y′′ n(x)=1 2πicontintegraldisplaye−xz/(1−z) (1−z)3zn−1dz. (13.53c) Substitutingintotheleft-handsideofEq. (13.52), weobtain 1 2πicontintegraldisplaybracketleftbiggx (1−z)3zn−1−1−x (1−z)2zn+n (1−z)zn+1bracketrightbigg e−xz/(1−z)dz, whichisequalto −1 2πicontintegraldisplayd dzbracketleftbigge−xz/(1−z) (1−z)znbracketrightbigg dz. (13.54) Ifweintegrateourperfectdifferentialaroundaclosedcontour(Fig.13.3),theintegralwill vanish,thusverifyingthat yn(x)(Eq. (13.53a)) isasolutionofLaguerre’sequation. It hasbecomecustomarytodefine Ln(x), theLaguerrepolynomial(Fig.13.4), by5 Ln(x)=1 2πicontintegraldisplaye−xz/(1−z) (1−z)zn+1dz. (13.55) 5Other definitions of Ln(x)are in use. The definitions here of the Laguerre polynomial Ln(x)and the associated Laguerre polynomial Lkn(x)agree with AMS-55, Chapter 22. (For the full ref. see footnote 4 in Chapter 5 or the General References at book’s end.) 838 Chapter 13 More Special Functions FIGURE 13.3Laguerre polynomialcontour. FIGURE 13.4Laguerre polynomials. Thisis exactlywhatwewouldobtainfromtheseries g(x,z)=e−xz/(1−z) 1−z=∞summationdisplay n=0Ln(x)zn,|z|<1, (13.56) ifwemultiplied g(x,z)byz−n−1andintegratedaroundtheorigin.Asinthedevelopment of the calculus of residues (Section 7.1), only the z−1term in the series survives. On this basisweidentify g(x,z)as thegeneratingfunctionfor theLaguerrepolynomials. 13.2 Laguerre Functions 839 Withthetransformation xz 1−z=s−x,orz=s−x s, (13.57) Ln(x)=ex 2πicontintegraldisplaysne−s (s−x)n+1ds, (13.58) the new contour enclosing the point s=xin thes-plane. By Cauchy’s integral formula (for derivatives), Ln(x)=ex n!dn dxnparenleftbig xne−xparenrightbig (integraln), (13.59) givingRodrigues’formulaforLaguerrepolynomials.Fromtheserepresentationsof Ln(x) wefindtheseriesform (for integral n), Ln(x)=(−1)n n!bracketleftbigg xn−n2 1!xn−1+n2(n−1)2 2!xn−2−···+(−1)nn!bracketrightbigg =nsummationdisplay m=0(−1)mn!xm (n−m)!m!m!=nsummationdisplay s=0(−1)n−sn!xn−s (n−s)!(n−s)!s!(13.60) and the specific polynomials listed in Table 13.2 (Exercise 13.2.1). Clearly, the defini- tion of Laguerre polynomials in Eqs. (13.55), (13.56), (13.59), and (13.60) are equivalent. Practical applications will decide which approach is used as one’s starting point. Equa- tion (13.59) is most convenient for generating Table 13.2, Eq. (13.56) for deriving recur- sionrelationsfrom whichtheODE(13.52)is recovered. By differentiating the generating function in Eq. (13.56) with respect to xandz,w e obtain recurrence relations for the Laguerre polynomials as follows. Using the product rulefordifferentiationweverifytheidentities (1−z)2∂g ∂z=(1−x−z)g(x,z), (z −1)∂g ∂x=zg(x,z). (13.61) Table 13.2 LaguerrePolynomials L0(x)=1 L1(x)=−x+1 2!L2(x)=x2−4x+2 3!L3(x)=−x3+9x2−18x+6 4!L4(x)=x4−16x3+72x2−96x+24 5!L5(x)=−x5+25x4−200x3+600x2−600x+120 6!L6(x)=x6−36x5+450x4−2400x3+5400x2−4320x+720 840 Chapter 13 More Special Functions Writingtheleft-handandright-handsidesofthefirstidentityintermsofLaguerrepolyno- mialsusingEq. (13.56)weobtain summationdisplay nbracketleftbig (n+1)Ln+1(x)−2nLn(x)+(n−1)Ln−1(x)bracketrightbig zn =summationdisplay nbracketleftbig (1−x)Ln(x)−Ln−1(x)bracketrightbig zn. Equatingcoefficientsof znyields (n+1)Ln+1(x)=(2n+1−x)Ln(x)−nLn−1(x). (13.62) To get the second recursion relation we use both identities of Eqs. (13.61) to verify the thirdidentity, x∂g ∂x=z∂g ∂z−z∂(zg) ∂z, (13.63) which, when written similarly in terms of Laguerre polynomials, is seen to be equivalent to xL′ n(x)=nLn(x)−nLn−1(x). (13.64) Equation(13.61),modifiedtoread Ln+1(x)=2Ln(x)−Ln−1(x)−1 n+1bracketleftbig (1+x)Ln(x)−Ln−1(x)bracketrightbig ,(13.65) for reasons of economy and numerical stability, is used for computation of numerical val- ues ofLn(x). The computer starts with known numerical values of L0(x)andL1(x),T a - ble 13.2, and works up step by step. This is the same technique discussed for computing Legendrepolynomials,Section12.2. Also,from Eq.(13.56) wefind g(0,z)=1 1−z=∞summationdisplay n=0zn=∞summationdisplay n=0Ln(0)zn, whichyieldsthespecialvaluesof Laguerrepolynomials Ln(0)=1. (13.66) As is seen from the form of the generating function, from the form of Laguerre’s ODE, or from Table13.2, theLaguerre polynomialshaveneitheroddnor evensymmetryunder the paritytransformation x→−x. The Laguerre ODE is not self-adjoint, and the Laguerre polynomials Ln(x)do not by themselves form an orthogonal set. However, following the method of Section 10.1, if we multiplyEq. (13.52)by e−x(Exercise10.1.1)weobtain integraldisplay∞ 0e−xLm(x)Ln(x)dx=δmn. (13.67) 13.2 Laguerre Functions 841 Thisorthogonalityisa consequenceof theSturm–Liouvilletheory,Section10.1. Thenor- malization follows from the generating function. It is sometimes convenient to define or- thogonalizedLaguerrefunctions(withunitweightingfunction)by ϕn(x)=e−x/2Ln(x). (13.68) Ourneworthonormalfunction, ϕn(x), satisfiestheODE xϕ′′ n(x)+ϕ′ n(x)+parenleftbigg n+1 2−x 4parenrightbigg ϕn(x)=0, (13.69) which is seen to have the (self-adjoint) Sturm–Liouville form. Note that the interval (0≤x<∞)was used because Sturm–Liouville boundary conditions are satisfied at its endpoints. Associated Laguerre Polynomials Inmanyapplications,particularlyinquantummechanics,weneedtheassociatedLaguerre polynomialsdefinedby6 Lk n(x)=(−1)kdk dxkLn+k(x). (13.70) From the series form of Ln(x)we verify that the lowest associated Laguerre polynomials aregivenby Lk 0(x)=1, Lk 1(x)=−x+k+1, Lk 2(x)=x2 2−(k+2)x+(k+2)(k+1) 2. (13.71) Ingeneral, Lk n(x)=nsummationdisplay m=0(−1)m(n+k)! (n−m)!(k+m)!m!xm,k>−1. (13.72) A generating function may be developed by differentiating the Laguerre generating func- tionktimestoyield (−1)kdk dxke−xz/(1−z) 1−z=(−1)k∞summationdisplay n=0dk dxkLn+k(x)zn+k=∞summationdisplay n=0Lk n(x)zn+k =parenleftbiggz 1−zparenrightbiggkexz/(1−z) 1−z. 6Some authors use Lk n+k(x)=(dk/dxk)[Ln+k(x)]. Henceour Lkn(x)=(−1)kLk n+k(x). 842 Chapter 13 More Special Functions Fromthelasttwomembersofthisequation,cancelingthecommonfactor zk, weobtain e−xz/(1−z) (1−z)k+1=∞summationdisplay n=0Lk n(x)zn,|z|<1. (13.73) Fromthis, for x=0,thebinomialexpansion 1 (1−z)k+1=∞summationdisplay n=0parenleftbigg−k−1 nparenrightbigg (−z)n=∞summationdisplay n=0Lk n(0)zn yields Lk n(0)=(n+k)! n!k!. (13.74) Recurrence relations can be derived from the generating function or by differentiating the Laguerrepolynomialrecurrencerelations.Amongthenumerouspossibilitiesare (n+1)Lk n+1(x)=(2n+k+1−x)Lk n(x)−(n+k)Lk n−1(x), (13.75) xdLk n(x) dx=nLk n(x)−(n+k)Lk n−1(x). (13.76) Thus,weobtainfrom differentiatingLaguerre’sODEonce xdL′′ n dx+L′′ n−L′n+(1−x)dL′ n dx+ndLn dx=0, andeventuallyfrom differentiatingLaguerre’sODE ktimes xdk dxkL′′ n+kdk−1 dxk−1L′′ n−kdk−1 dxk−1L′ n+(1−x)dk dxkL′ n+ndk dxkLn=0. Adjustingtheindex n→n+k,wehavetheassociatedLaguerreODE xd2Lk n(x) dx2+(k+1−x)dLk n(x) dx+nLk n(x)=0. (13.77) When associated Laguerre polynomials appear in a physical problem it is usually because that physical problem involves Eq. (13.77). The most important application is the bound statesofthehydrogenatom,whicharederivedinupcomingExample13.2.1. ARodriguesrepresentationoftheassociatedLaguerrepolynomial Lk n(x)=exx−k n!dn dxnparenleftbig e−xxn+kparenrightbig (13.78) maybeobtainedfromsubstitutingEq.(13.59)intoEq.(13.70).Notethatalltheseformulas for associated Legendre polynomials Lk n(x)reduce to the corresponding expressions for Ln(x)whenk=0. 13.2 Laguerre Functions 843 The associated Laguerre equation (13.77) is not self-adjoint, but it can be put in self- adjoint form by multiplying by e−xxk, which becomes the weighting function (Sec- tion10.1). Weobtain integraldisplay∞ 0e−xxkLk n(x)Lkm(x)dx=(n+k)! n!δmn. (13.79) Equation (13.79) shows the same orthogonality interval (0,∞)as that for the Laguerre polynomials, but with a new weighting function we have a new set of orthogonal polyno- mials,theassociatedLaguerrepolynomials. Byletting ψk n(x)=e−x/2xk/2Lk n(x),ψk n(x)satisfiestheself-adjointODE xd2ψk n(x) dx2+dψk n(x) dx+parenleftbigg −x 4+2n+k+1 2−k2 4xparenrightbigg ψk n(x)=0. (13.80) Theψk n(x)aresometimescalled Laguerrefunctions .Equation(13.67)isthespecialcase k=0 ofEq. (13.79). Afurther usefulformisgivenbydefining7 /Phi1k n(x)=e−x/2x(k+1)/2Lkn(x). (13.81) SubstitutionintotheassociatedLaguerreequationyields d2/Phi1k n(x) dx2+parenleftbigg −1 4+2n+k+1 2x−k2−1 4x2parenrightbigg /Phi1k n(x)=0. (13.82) Thecorrespondingnormalizationintegralintegraltext∞ 0|/Phi1k n(x)|2dxis integraldisplay∞ 0e−xxk+1bracketleftbig Lk n(x)bracketrightbig2dx=(n+k)! n!(2n+k+1). (13.83) Noticethatthe /Phi1k n(x)donotformanorthogonalset(exceptwith x−1asaweightingfunc- tion) because of the x−1in the term (2n+k+1)/2x. (The Laguerre functions Lµ ν(x)in which the indices νandµarenotintegers may be defined using the confluent hypergeo- metricfunctionsofSection13.5.) Example 13.2.1 THEHYDROGEN ATOM The most important application of the Laguerre polynomials is in the solution of the Schrödingerequationfor thehydrogenatom.This equationis −¯h2 2m∇2ψ−Ze2 4πǫ0rψ=Eψ, (13.84) in which Z=1 for hydrogen, 2 for ionized helium, and so on. Separating variables, we find that the angular dependence of ψis the spherical harmonics YM L(θ,ϕ). The radial part,R(r), satisfiestheequation −¯h2 2m1 r2d drparenleftbigg r2dR drparenrightbigg −Ze2 4πǫ0rR+¯h2 2mL(L+1) r2R=ER. (13.85) 7This corresponds to modifying the function ψin Eq.(13.80) to eliminate the first derivative (compare Exercise 9.6.11). 844 Chapter 13 More Special Functions Forboundstates, R→0asr→∞,andRisfiniteattheorigin, r=0.Wedonotconsider continuumstateswithpositiveenergy.Onlywhenthelatterareincludeddohydrogenwave functionsforma completeset. By use of the abbreviations (resulting from rescaling rto the dimensionless radial vari- ableρ) ρ=αrwithα2=−8mE ¯h2,E<0,λ=mZe2 2πǫ0α¯h2,(13.86) Eq.(13.85) becomes 1 ρ2d dρparenleftbigg ρ2dχ(ρ) dρparenrightbigg +parenleftbiggλ ρ−1 4−L(L+1) ρ2parenrightbigg χ(ρ)=0, (13.87) whereχ(ρ)=R(ρ/α). A comparison with Eq. (13.82) for /Phi1k n(x)shows that Eq. (13.87) issatisfiedby ρχ(ρ)=e−ρ/2ρL+1L2L+1 λ−L−1(ρ), (13.88) inwhich kis replacedby 2 L+1 andnbyλ−L−1,uponusing 1 ρ2d dρρ2dχ dρ=1 ρd2 dρ2(ρχ). Wemustrestricttheparameter λbyrequiringittobeaninteger n,n=1,2,3,....8This isnecessarybecausetheLaguerrefunctionofnonintegral nwoulddiverge9asρneρ,which isunacceptableforour physicalproblem,inwhich limr→∞R(r)=0. This restriction on λ, imposed by our boundary condition, has the effect of quantizing the energy, En=−Z2m 2n2¯h2parenleftbigge2 4πǫ0parenrightbigg2 . (13.89) The negativesign reflects the fact that we are dealing here with bound states ( E<0), cor- responding to an electron that is unable to escape to infinity, where the Coulomb potential goestozero.Usingthis resultfor En,weha v e α=me2 2πǫ0¯h2·Z n=2Z na0,ρ=2Z na0r, (13.90) with a0=4πǫ0¯h2 me2,theBohrradius. 8This is the conventional notation for λ.It is not the same nas the index nin/Phi1kn(x). 9This maybe shown, as in Exercise9.5.5. 13.2 Laguerre Functions 845 Thus,thefinalnormalizedhydrogenwavefunctioniswrittenas ψnLM(r,θ,ϕ)=bracketleftbiggparenleftbigg2Z na0parenrightbigg3(n−L−1)! 2n(n+L)!bracketrightbigg1/2 e−αr/2(αr)LL2L+1 n−L−1(αr)YM L(θ,ϕ). (13.91) Regular solutions exist for n≥L+1, so the lowest state with L=1 (called a 2P state) occursonlywith n=2. /squaresolid Exercises 13.2.1 Show with the aid of the Leibniz formula that the series expansion of Ln(x) (Eq. (13.60))followsfrom theRodriguesrepresentation(Eq. (13.59)). 13.2.2 (a) Usingtheexplicitseriesform (Eq. (13.60)) showthat L′ n(0)=−n, L′′n(0)=1 2n(n−1). (b) Repeatwithoutusingtheexplicitseriesform of Ln(x). 13.2.3 FromthegeneratingfunctionderivetheRodriguesrepresentation Lk n(x)=exx−k n!dn dxnparenleftbig e−xxn+kparenrightbig . 13.2.4 Derivethenormalizationrelation(Eq.(13.79))fortheassociatedLaguerrepolynomials. 13.2.5 Expandxrin aseries of associatedLaguerrepolynomials Lk n(x),kfixedand nranging from 0to r(or to∞ifris notaninteger). Hint.TheRodriguesform of Lkn(x)willbeuseful. ANS.xr=(r+k)!r!rsummationdisplay n=0(−1)nLk n(x) (n+k)!(r−n)!,0≤x<∞. 13.2.6 Expande−axinaseriesofassociatedLaguerrepolynomials Lk n(x),kfixedand nrang- ingfrom 0to∞. (a) Evaluatedirectlythecoefficientsinyourassumedexpansion. (b) Developthedesiredexpansionfromthegeneratingfunction. ANS.e−ax=1 (1+a)1+k∞summationdisplay n=0parenleftbigga 1+aparenrightbiggn Lk n(x), 0≤x<∞. 13.2.7 Showthatintegraldisplay∞ 0e−xxk+1Lkn(x)Lkn(x)dx=(n+k)! n!(2n+k+1). Hint.Notethat xLk n=(2n+k+1)Lkn−(n+k)Lk n−1−(n+1)Lkn+1. 846 Chapter 13 More Special Functions 13.2.8 AssumethataparticularprobleminquantummechanicshasledtotheODE d2y dx2−bracketleftbiggk2−1 4x2−2n+k+1 2x+1 4bracketrightbigg y=0 for nonnegativeintegers n,k.Writey(x)as y(x)=A(x)B(x)C(x), withtherequirementthat (a)A(x)be anegative exponential giving the required asymptotic behavior of y(x) and (b)B(x)beapositivepowerof xgivingthebehaviorof y(x)for 0≤x≪1. Determine A(x)andB(x).Findtherelationbetween C(x)andtheassociatedLaguerre polynomial. ANS.A(x)=e−x/2,B(x)=x(k+1)/2,C(x)=Lk n(x). 13.2.9 FromEq.(13.91) thenormalizedradialpartof thehydrogenicwavefunctionis RnL(r)=bracketleftbigg α3(n−L−1)! 2n(n+L)!bracketrightbigg1/2 e−αr(αr)LL2L+1 n−L−1(αr), inwhich α=2Z/na0=2Zme2/4πε0¯h2.Evaluate (a)/angbracketleftr/angbracketright=integraldisplay∞ 0rRnL(αr)RnL(αr)r2dr, (b)angbracketleftbig r−1angbracketrightbig =integraldisplay∞ 0r−1RnL(αr)RnL(αr)r2dr. The quantity/angbracketleftr/angbracketrightis the average displacement of the electron from the nucleus, whereas /angbracketleftr−1/angbracketrightistheaverageofthereciprocaldisplacement. ANS./angbracketleftr/angbracketright=a0 2bracketleftbig 3n2−L(L+1)bracketrightbig ,angbracketleftbig r−1angbracketrightbig =1 n2a0. 13.2.10 Derivetherecurrencerelationforthehydrogenwavefunctionexpectationvalues: s+2 n2angbracketleftbig rs+1angbracketrightbig −(2s+3)a0angbracketleftbig rsangbracketrightbig +s+1 4bracketleftbig (2L+1)2−(s+1)2bracketrightbig a2 0angbracketleftbig rs−1angbracketrightbig =0, withs≥−2L−1,/angbracketleftrs/angbracketright≡/overbarrs. Hint.TransformEq.(13.87)intoaformanalogoustoEq.(13.80).Multiplyby ρs+2u′− cρs+1u.Her eu=ρ/Phi1. Adjustctocanceltermsthatdonotyieldexpectationvalues. 13.2.11 The hydrogen wave functions, Eq. (13.91), are mutually orthogonal, as they should be, sincetheyareeigenfunctionsof theself-adjointSchrödingerequation integraldisplay ψ∗ n1L1M1ψn2L2M2r2drd/Omega1=δn1n2δL1L2δM1M2. 13.2 Laguerre Functions 847 Yettheradialintegralhas the(misleading)form integraldisplay∞ 0e−αr/2(αr)LL2L+1 n1−L−1(αr)e−αr/2(αr)LL2L+1 n2−L−1(αr)r2dr, whichappearsto match Eq. (13.83) and not the associated Laguerre orthogonality re- lation,Eq. (13.79). Howdoyouresolvethisparadox? ANS.The parameter αis dependent on n. The first three α,p r e v i - ously shown, are 2 Z/n1a0. The last three are 2 Z/n2a0.F o r n1=n2Eq.(13.83)applies.For n1/negationslash=n2neitherEq.(13.79) nor Eq.(13.83) isapplicable. 13.2.12 A quantum mechanical analysis of the Stark effect (parabolic coordinate) leads to the ODE d dξparenleftbigg ξdu dξparenrightbigg +parenleftbigg1 2Eξ+L−m2 4ξ−1 4Fξ2parenrightbigg u=0. HereFis a measure of the perturbation energy introduced by an external electric field. Find the unperturbed wave functions (F=0)in terms of associated Laguerre polyno- mials. ANS.u(ξ)=e−εξ/2ξm/2Lm p(εξ), withε=√ −2E>0, p=L/ε−(m+1)/2, anonnegativeinteger. 13.2.13 Thewaveequationforthethree-dimensionalharmonicoscillatoris −¯h2 2M∇2ψ+1 2Mω2r2ψ=Eψ. Hereωistheangularfrequencyofthecorrespondingclassicaloscillator.Showthatthe radial part of ψ(in spherical polar coordinates ) may be written in terms of associated Laguerrefunctionsof argument (βr2), whereβ=Mω/¯h. Hint. As in Exercise 13.2.8, split off radial factors of rlande−βr2/2. The associated Laguerrefunctionwillhavetheform Ll+1/2 1/2(n−l−1)(βr2). 13.2.14 Write acomputerprogramthatwillgeneratethecoefficients asinthepolynomialform oftheLaguerrepolynomial Ln(x)=summationtextn s=0asxs. 13.2.15 Write a computer program that will transform a finite power seriessummationtextNn=0anxninto a LaguerreseriessummationtextN n=0bnLn(x). Usetherecurrencerelation,Eq. (13.62). 13.2.16 Tabulate L10(x)forx=0.0(0.1)30.0. This will include the 10 roots of L10. Beyond x=30.0,L10(x)is monotonic increasing. If graphic software is available, plot your results. Checkvalue. Eighthroot=16.279. 13.2.17 Determinethe10rootsof L10(x)usingroot-findingsoftware.Youmayuseyourknowl- edgeoftheapproximatelocationoftherootsordevelopasearchroutinetolookforthe roots.The10rootsof L10(x)aretheevaluationpointsforthe10-pointGauss–Laguerre quadrature. Check your values by comparing with AMS-55, Table 25.9. (For the refer- enceseefootnote4 inChapter5or theGeneralReferencesatbook’send.) 848 Chapter 13 More Special Functions 13.2.18 Calculate the coefficients of a Laguerre series expansion (Ln(x),k=0)of the ex- ponential e−x. Evaluate the coefficients by the Gauss–Laguerre quadrature (compare Eq. (10.64)). CheckyourresultsagainstthevaluesgiveninExercise13.2.6. Note. Direct application of the Gauss–Laguerre quadrature with f(x)Ln(x)e−xgives poor accuracy because of the extra e−x. Try a change of variable, y=2x, so that the functionappearingintheintegrandwillbesimply Ln(y/2). 13.2.19 (a) WriteasubroutinetocalculatetheLaguerrematrixelements Mmnp=integraldisplay∞ 0Lm(x)Ln(x)xpe−xdx. Include a check of the condition |m−n|≤p≤m+n. (Ifpis outside this range, Mmnp=0.Why?) Note. A 10-point Gauss–Laguerre quadrature will give accurate results for m+n+p≤19. (b) Call your subroutine to calculate a variety of Laguerre matrix elements. Check Mmn1againstExercise13.2.7. 13.2.20 Writeasubroutinetocalculatethenumericalvalueof Lk n(x)forspecifiedvaluesof n,k, andx.Requirethat nandkbenonnegativeintegersand x≥0. Hint.Startingwithknownvaluesof Lk 0andLk1(x), we mayusethe recurrencerelation, Eq. (13.75),togenerate Lk n(x),n=2,3,4,.... 13.2.21 Showthatintegraltext∞ −∞xne−x2Hn(xy)dx=√ πn!Pn(y), wherePnisaLegendrepolynomial. 13.2.22 Write a program to calculate the normalized hydrogen radial wave function ψnL(r). This isψnLMof Eq. (13.91), omitting the spherical harmonic YM L(θ,ϕ).T a k eZ=1 anda0=1 (which means that ris being expressed in units of Bohr radii). Accept n andLas input data. Tabulate ψnL(r)forr=0.0(0.2)RwithRtaken large enough to exhibit the significant features of ψ. This means roughly R=5f o rn=1,R=10 for n=2,andR=30 forn=3. 13.3 C HEBYSHEV POLYNOMIALS In this section two types of Chebyshev polynomials are developed as special cases of ul- traspherical polynomials. Their properties follow from the ultraspherical polynomial gen- erating function. The primary importance of the Chebyshev polynomials is in numerical analysis. Generating Functions InSection12.1thegeneratingfunctionfor theultraspherical,orGegenbauer,polynomials 1 (1−2xt+t2)α=∞summationdisplay n=0C(α) n(x)tn,|x|<1,|t|<1 (13.92) wasmentioned,with α=1 2givingrisetotheLegendrepolynomials.Inthissectionwefirst takeα=1 and then α=0 to generate two sets of polynomials known as the Chebyshev polynomials. 13.3 Chebyshev Polynomials 849 Type II Withα=1 andC(1) n(x)=Un(x), Eq.(13.92) gives 1 1−2xt+t2=∞summationdisplay n=0Un(x)tn,|x|<1,|t|<1. (13.93) Thesefunctions Un(x)generatedby (1−2xt+t2)−1arelabeledChebyshevpolynomials type II. Although these polynomials have few applications in mathematical physics, one unusualapplicationisinthedevelopmentoffour-dimensionalsphericalharmonicsusedin angularmomentumtheory. Type I Withα=0 there is a difficulty. Indeed, our generating function reduces to the constant 1. WemayavoidthisproblembyfirstdifferentiatingEq.(13.92)withrespectto t.Thisyields −α(−2x+2t) (1−2xt+t2)α+1=∞summationdisplay n=1nC(α) n(x)tn−1, (13.94) or x−t (1−2xt+t2)α+1=∞summationdisplay n=1n 2bracketleftbiggC(α) n(x) αbracketrightbigg tn−1. (13.95) Wedefine C(0) n(x)by C(0) n(x)=lim α→0C(α) n(x) α. (13.96) The purpose of differentiating with respect to twas to get αin the denominator and to create an indeterminate form. Now multiplying Eq. (13.95) by 2 tand adding 1 = (1−2xt+t2)/(1−2xt+t2),weobtain 1−t2 1−2xt+t2=1+2∞summationdisplay n=1n 2C(0) n(x)tn. (13.97) Wedefine Tn(x)by Tn(x)=braceleftBigg1,n =0 n 2C(0) n(x), n> 0.(13.98) Noticethespecialtreatmentfor n=0.Thisissimilartothetreatmentofthe n=0t e r mi n theFourierseries.Also,notethat C(0) nisthelimitindicatedinEq.(13.96)andnotaliteral substitutionof α=0 intothegeneratingfunctionseries. Withthesenewlabels, 1−t2 1−2xt+t2=T0(x)+2∞summationdisplay n=1Tn(x)tn,|x|≤1,|t|<1. (13.99) 850 Chapter 13 More Special Functions WecallTn(x)thetypeIChebyshevpolynomials.Notethatthenotationandspellingofthe name for these functions differs from reference to reference. Here we follow the usage of AMS-55(for thefullreferenceseefootnote4inChapter5). Differentiating the generating function (Eqs. (13.99)) with respect to tand multiplying bythedenominator, 1 −2xt+t2, weobtain −t−(t−x)bracketleftbigg T0(x)+2∞summationdisplay n=1Tn(x)tnbracketrightbigg =parenleftbig 1−2xt+t2parenrightbig∞summationdisplay n=1nTn(x)tn−1 =∞summationdisplay n=1bracketleftbig nTntn−1−2xnTntn+nTntn+1bracketrightbig , fromwhichtherecurrencerelation Tn+1(x)−2xTn(x)+Tn−1(x)=0 (13.100) follows by shifting the summation index so as to get the same power, tn, in each term and thencomparingcoefficientsof tn. SimilarlytreatingEq.(13.93) wefind −2(t−x) 1−2xt+t2=parenleftbig 1−2xt+t2parenrightbig∞summationdisplay n=1nUn(x)tn−1 fromwhichtherecursionrelation Un+1(x)−2xUn(x)+Un−1(x)=0 (13.101) followsuponcomparingcoefficientsoflikepowersof t(seeTable13.3). Then, using the generating functions for the first few values of nand these recurrence relationsforthehigher-orderpolynomials,wegetTables13.4and13.5(seealsoFigs.13.5 and13.6). As with the Hermite polynomials, Section 13.1, the recurrence relations, Eqs. (13.100) and(13.101),togetherwiththeknownvaluesof T0(x),T1(x),U0(x),andU1(x),providea convenient—thatis,foracomputer—meansofgettingthenumericalvalueofany Tn(x0) orUn(x0),withx0agivennumber. Table 13.3 RecursionrelationaPn+1(x)= (Anx+Bn)Pn(x)−CnPn−1(x) Pn(x) A n BnCn Legendre Pn(x)2n+1 n+101 n+1 Chebyshev I Tn(x) 20 1 Shifted Chebyshev I T∗n(x) 4−21 Chebyshev II Un(x) 20 1 Shifted Chebyshev II U∗n(x) 4−21 AssociatedLaguerre L(k) n(x)−1 n+12n+k+1 n+1n+k n+1 Hermite Hn(x) 20 2 n aPndenotesany oftheorthogonalpolynomials. 13.3 Chebyshev Polynomials 851 Table 13.4 Chebyshev polynomials,typeI T0=1 T1=x T2=2x2−1 T3=4x3−3x T4=8x4−8x2+1 T5=16x5−20x3+5x T6=32x6−48x4+18x2−1Table 13.5 Chebyshev polynomials,typeII U0=1 U1=2x U2=4x2−1 U3=8x3−4x U4=16x4−12x2+1 U5=32x5−32x3+6x U6=64x6−80x4+24x2−1 FIGURE 13.5Chebyshevpolynomials Tn(x). FIGURE 13.6Chebyshevpolynomials Un(x). 852 Chapter 13 More Special Functions Again, from the generating functions, we can obtain the special values of various poly- nomials: Tn(1)=1,T n(−1)=(−1)n, (13.102) T2n(0)=(−1)n,T 2n+1(0)=0; Un(1)=n+1,U n(−1)=(−1)n(n+1), (13.103) U2n(0)=(−1)n,U 2n+1(0)=0. Forexample,comparingthepowerseries 1−t2 (1−t)2=1+t 1−t=∞summationdisplay n=0parenleftbig tn+tn+1parenrightbig withEq.(13.99)for x=1gi vesTn(1),andforx=−1asimilarexpansionof (1−t)/(1+ t)givesTn(−1),whilereplacing t→−t2inthefirstpowerseriesyields Tn(0).Thepower seriesfor (1±t)−2and(1+t2)−1generatethecorresponding Un(±1),Un(0). The parity relations for TnandUnfollow from their generating functions, with the sub- stitutions t→−t,x→−x,whichleavetheminvariant;theseare Tn(x)=(−1)nTn(−x), U n(x)=(−1)nUn(−x). (13.104) Rodrigues’representationsof Tn(x)andUn(x)are Tn(x)=(−1)nπ1/2(1−x2)1/2 2n(n−1 2)!dn dxnbracketleftbigparenleftbig 1−x2parenrightbign−1/2bracketrightbig (13.105) and Un(x)=(−1)n(n+1)π1/2 2n+1(n+1 2)!(1−x2)1/2dn dxnbracketleftbigparenleftbig 1−x2parenrightbign+1/2bracketrightbig .(13.106) Recurrence Relations — Derivatives Differentiationofthegeneratingfunctionsfor Tn(x)andUn(x)withrespecttothevariable xleads to a variety of recurrence relations involving derivatives. For example, from Eq. (13.99)wethusobtain parenleftbig 1−2xt+t2parenrightbig 2∞summationdisplay n=1T′ n(x)tn=2tbracketleftbigg T0(x)+2∞summationdisplay n=1Tn(x)tnbracketrightbigg , fromwhichweextracttherecursion 2Tn−1(x)=T′ n(x)−2xT′ n−1(x)+T′ n−2(x), (13.107) which is the derivative of Eq. (13.100) for n→n−1. Among the more useful recursions wethusfindare parenleftbig 1−x2parenrightbig T′ n(x)=−nxTn(x)+nTn−1(x) (13.108) 13.3 Chebyshev Polynomials 853 and parenleftbig 1−x2parenrightbig U′ n(x)=−nxUn(x)+(n+1)Un−1(x). (13.109) ManipulatingavarietyoftheserecursionsasinSection12.2forLegendrepolynomialsone can eliminate the index n−1a l s oi nf a v o ro f T′′ nand establish that Tn(x), the Chebyshev polynomialtypeI, satisfiestheODE parenleftbig 1−x2parenrightbig T′′ n(x)−xT′ n(x)+n2Tn(x)=0. (13.110) TheChebyshevpolynomialoftypeII, Un(x), satisfies parenleftbig 1−x2parenrightbig U′′ n(x)−3xU′ n(x)+n(n+2)Un(x)=0. (13.111) Chebyshev polynomials may be defined starting from these ODEs, but our emphasis has beenongeneratingfunctions. Theultrasphericalequation parenleftbig 1−x2parenrightbigd2 dx2C(α) n(x)−(2α+1)xd dxC(α) n(x)+n(n+2α)C(α) n(x)=0 (13.112) is a generalization of these differential equations, reducing to Eq. (13.110) for α=0 and Eq.(13.111)for α=1 (andtoLegendre’sequationfor α=1 2). Trigonometric Form AtthispointinthedevelopmentofthepropertiesoftheChebyshevsolutionsitisbeneficial to change variables, replacing xby cosθ. Withx=cosθandd/dx=(−1/sinθ)(d/dθ) , weverifythat parenleftbig 1−x2parenrightbigd2Tn dx2=d2Tn dθ2−cotθdTn dθ,xT′ n=−cotθdTn dθ. Addingtheseterms, Eq.(13.110) becomes d2Tn dθ2+n2Tn=0, (13.113) the simple harmonic oscillator equation with solutions cos nθand sinnθ. The special val- ues(boundaryconditionsat x=0,1)identify Tn=cosnθ=cosn(arccosx). (13.114a) Asecondlinearlyindependentsolutionof Eqs. (13.110)and(13.113)is labeled Vn=sinnθ=sinn(arccosx). (13.114b) ThecorrespondingsolutionsofthetypeII Chebyshevequation,Eq.(13.111), become Un=sin(n+1)θ sinθ, (13.115a) Wn=cos(n+1)θ sinθ. (13.115b) 854 Chapter 13 More Special Functions Thetwosetsof solutions,typeI andtypeII, arerelatedby Vn(x)=parenleftbig 1−x2parenrightbig1/2Un−1(x), (13.116a) Wn(x)=parenleftbig 1−x2parenrightbig−1/2Tn+1(x). (13.116b) As already seen from generating functions, Tn(x)andUn(x)are polynomials. Clearly, Vn(x)andWn(x)arenotpolynomials.From Tn(x)+iVn(x)=cosnθ+isinnθ =(cosθ+isinθ)n=bracketleftbig x+iparenleftbig 1−x2parenrightbig1/2bracketrightbign,|x|≤1 (13.117) weobtainexpansions Tn(x)=xn−parenleftbiggn 2parenrightbigg xn−2parenleftbig 1−x2parenrightbig +parenleftbiggn4parenrightbigg xn−4parenleftbig 1−x2parenrightbig2−··· (13.118a) and Vn(x)=radicalbig 1−x2bracketleftbiggparenleftbiggn 1parenrightbigg xn−1−parenleftbiggn3parenrightbigg xn−3parenleftbig 1−x2parenrightbig +···bracketrightbigg . (13.118b) Fromthegeneratingfunctions,orfrom theODEs, power-seriesrepresentationsare Tn(x)=n 2[n/2]summationdisplay m=0(−1)m(n−m−1)! m!(n−2m)!(2x)n−2m, (13.119a) forn≥1,with[n/2]thelargestintegerbelow n/2 and Un(x)=[n/2]summationdisplay m=0(−1)m(n−m)! m!(n−2m)!(2x)n−2m. (13.119b) Orthogonality IfEq.(13.110)isputintoself-adjointform(Section10.1),weobtain w(x)=(1−x2)−1/2 asaweightingfactor.ForEq.(13.111)thecorrespondingweightingfactoris (1−x2)+1/2. Theresultingorthogonalityintegrals, integraldisplay1 −1Tm(x)Tn(x)parenleftbig 1−x2parenrightbig−1/2dx=  0,m/negationslash=n,π 2,m=n/negationslash=0, π, m=n=0,(13.120) integraldisplay1 −1Vm(x)Vn(x)parenleftbig 1−x2parenrightbig−1/2dx=  0,m/negationslash=n, π 2,m=n/negationslash=0, 0,m=n=0,(13.121) integraldisplay1 −1Um(x)Un(x)parenleftbig 1−x2parenrightbig1/2dx=π 2δm,n, (13.122) 13.3 Chebyshev Polynomials 855 and integraldisplay1 −1Wm(x)Wn(x)parenleftbig 1−x2parenrightbig1/2dx=π 2δm,n, (13.123) are a direct consequence of the Sturm–Liouville theory, Chapter 10. The normalization values may best be obtained by using x=cosθand converting these four integrals into Fouriernormalizationintegrals(for thehalf-periodinterval [0,π]). Exercises 13.3.1 AnotherChebyshevgeneratingfunctionis 1−xt 1−2xt+t2=∞summationdisplay n=0Xn(x)tn,|t|<1. HowisXn(x)relatedto Tn(x)andUn(x)? 13.3.2 Given parenleftbig 1−x2parenrightbig U′′ n(x)−3xU′ n(x)+n(n+2)Un(x)=0, showthat Vn(x)(Eq. (13.116a))satisfies parenleftbig 1−x2parenrightbig V′′ n(x)−xV′ n(x)+n2Vn(x)=0, whichisChebyshev’sequation. 13.3.3 ShowthattheWronskianof Tn(x)andVn(x)is givenby Tn(x)V′ n(x)−T′ n(x)Vn(x)=−n (1−x2)1/2. This verifies that TnandVn(n/negationslash=0)are independent solutions of Eq. (13.110). Con- versely,for n=0,wedonothavelinearindependence.Whathappensat n=0?Where isthe“second”solution? 13.3.4 Showthat Wn(x)=(1−x2)−1/2Tn+1(x)isasolutionof parenleftbig 1−x2parenrightbig W′′ n(x)−3xW′ n(x)+n(n+2)Wn(x)=0. 13.3.5 EvaluatetheWronskianof Un(x)andWn(x)=(1−x2)−1/2Tn+1(x). 13.3.6 Vn(x)=(1−x2)1/2Un−1(x)is not defined for n=0. Show that a second and inde- pendent solution of the Chebyshev differential equation for Tn(x) (n=0)isV0(x)= arccosx(or arcsin x). 13.3.7 Show that Vn(x)satisfies the same three-term recurrence relation as Tn(x) (Eq. (13.100)). 13.3.8 Verifytheseries solutionsfor Tn(x)andUn(x)(Eqs. (13.109a)and(13.119b)). 856 Chapter 13 More Special Functions 13.3.9 Transform theseriesform of Tn(x), Eq.(13.119a),intoan ascending powerseries. ANS.T2n(x)=(−1)nnnsummationdisplay m=0(−1)m(n+m−1)! (n−m)!(2m)!(2x)2m,n≥1, T2n+1(x)=2n+1 2nsummationdisplay m=0(−1)m+n(n+m)! (n−m)!(2m+1)!(2x)2m+1. 13.3.10 Rewritetheseries formof Un(x), Eq. (13.119b),as anascendingpowerseries. ANS.U2n(x)=(−1)nnsummationdisplay m=0(−1)m(n+m)! (n−m)!(2m)!(2x)2m, U2n+1(x)=(−1)nnsummationdisplay m=0(−1)m(n+m+1)! (n−m)!(2m+1)!(2x)2m+1. 13.3.11 DerivetheRodriguesrepresentationof Tn(x), Tn(x)=(−1)nπ1/2(1−x2)1/2 2n(n−1 2)!dn dxnbracketleftbigparenleftbig 1−x2parenrightbign−1/2bracketrightbig . Hint.Onepossibilityis tousethehypergeometricfunctionrelation 2F1(a,b;c;z)=(1−z)−a2F1parenleftbigg a,c−b;c;−z 1−zparenrightbigg , withz=(1−x)/2.Analternateapproachistodevelopafirst-orderdifferentialequation fory=(1−x2)n−1/2.RepeateddifferentiationofthisequationleadstotheChebyshev equation. 13.3.12 (a) Fromthedifferentialequationfor Tn(in self-adjointform) showthat integraldisplay1 −1dTm(x) dxdTn(x) dxparenleftbig 1−x2parenrightbig1/2dx=0,m/negationslash=n. (b) Confirmtheprecedingresultbyshowingthat dTn(x) dx=nUn−1(x). 13.3.13 Theexpansionof apowerof xinaChebyshevseriesleadstotheintegral Imn=integraldisplay1 −1xmTn(x)dx√ 1−x2. (a) Showthatthisintegralvanishesfor m<n. (b) Showthatthisintegralvanishesfor m+nodd. 13.3.14 Evaluatetheintegral Imn=integraldisplay1 −1xmTn(x)dx√ 1−x2 form≥nandm+nevenbyeachof twomethods: 13.3 Chebyshev Polynomials 857 (a) Operatewith xasthevariablereplacing TnbyitsRodriguesrepresentation. (b) Using x=cosθtransform theintegraltoa formwith θasthevariable. ANS.Imn=πm! (m−n)!(m−n−1)!! (m+n)!!,m≥n, m+neven. 13.3.15 Establishthefollowingbounds, −1≤x≤1: (a)|Un(x)|≤n+1, (b)vextendsinglevextendsinglevextendsinglevextendsingled dxTn(x)vextendsinglevextendsinglevextendsinglevextendsingle≤n2. 13.3.16 (a) Establishthefollowingbound, −1≤x≤1:|Vn(x)|≤1. (b) Showthat Wn(x)is unboundedin −1≤x≤1. 13.3.17 Verifytheorthogonality-normalizationintegralsfor (a)Tm(x),Tn(x),(b)Vm(x),Vn(x), (c)Um(x),Un(x),(d)Wm(x),Wn(x). Hint.AllthesecanbeconvertedtoFourierorthogonality-normalizationintegrals. 13.3.18 Showwhether (a)Tm(x)andVn(x)are or are not orthogonal over the interval [−1,1]with respect totheweightingfactor (1−x2)−1/2. (b)Um(x)andWn(x)are or are not orthogonal over the interval [−1,1]with respect totheweightingfactor (1−x2)1/2. 13.3.19 Derive(a)Tn+1(x)+Tn−1(x)=2xTn(x), (b)Tm+n(x)+Tm−n(x)=2Tm(x)Tn(x), fromthe“corresponding”cosineidentities. 13.3.20 A number of equations relate the two types of Chebyshev polynomials. As examples showthat Tn(x)=Un(x)−xUn−1(x) and parenleftbig 1−x2parenrightbig Un(x)=xTn+1(x)−Tn+2(x). 13.3.21 Showthat dVn(x) dx=−nTn(x)√ 1−x2 (a) usingthetrigonometricformsof VnandTn, (b) usingtheRodriguesrepresentation. 858 Chapter 13 More Special Functions 13.3.22 Startingwith x=cosθandTn(cosθ)=cosnθ,expand xk=parenleftbiggeiθ+e−iθ 2parenrightbiggk andshowthat xk=1 2k−1bracketleftbigg Tk(x)+parenleftbiggk 1parenrightbigg Tk−2(x)+parenleftbiggk2parenrightbigg Tk−4+···bracketrightbigg , theseriesinbracketsterminatingwithparenleftbigk mparenrightbig T1(x)fork=2m+1or1 2parenleftbigk mparenrightbig T0fork=2m. 13.3.23 (a) Calculate and tabulate the Chebyshev functions V1(x),V2(x), andV3(x)forx= −1.0(0.1)1.0. (b) A second solution of the Chebyshev differential equation, Eq. (13.100), for n=0i sy(x)=sin−1x. Tabulate and plot this function over the same range: −1.0(0.1)1.0. 13.3.24 Write acomputerprogramthatwillgeneratethecoefficients asinthepolynomialform oftheChebyshevpolynomial Tn(x)=summationtextn s=0asxs. 13.3.25 Tabulate T10(x)for 0.00(0.01)1.00. This will include the five positive roots of T10.I f graphicssoftware isavailable,plotyourresults. 13.3.26 Determine the five positive roots of T10(x)by calling a root-finding subroutine. Use your knowledge of the approximate location of these roots from Exercise 13.3.25 or writeasearchroutinetolookfortheroots.Thesefivepositiveroots(andtheirnegatives) aretheevaluationpointsofthe10-pointGauss–Chebyshevquadraturemethod. Checkvalues. xk=cosbracketleftbig (2k−1)π/20bracketrightbig ,k=1,2,3,4,5. 13.3.27 DevelopthefollowingChebyshevexpansions(for [−1,1]): (a)parenleftbig 1−x2parenrightbig1/2=2 πbracketleftbigg 1−2∞summationdisplay s=1parenleftbig 4s2−1parenrightbig−1T2s(x)bracketrightbigg . (b)+1,0<x≤1 −1,−1≤x<0bracerightbigg =4 π∞summationdisplay s=0(−1)s(2s+1)−1T2s+1(x). 13.3.28 (a) Fortheinterval [−1,1]showthat |x|=1 2+∞summationdisplay s=1(−1)s+1(2s−3)!! (2s+2)!!(4s+1)P2s(x) =2 π+4 π∞summationdisplay s=1(−1)s+11 4s2−1T2s(x). (b) Showthattheratioofthecoefficientof T2s(x)tothatof P2s(x)approaches (πs)−1 ass→∞. This illustrates the relatively rapid convergence of the Chebyshev se- ries. 13.4 Hypergeometric Functions 859 Hint. With the Legendre recurrence relations, rewrite xPn(x)as a linear combination ofderivatives.Thetrigonometricsubstitution x=cosθ,Tn(x)=cosnθismosthelpful for theChebyshevpart. 13.3.29 Showthat π2 8=1+2∞summationdisplay s=1parenleftbig 4s2−1parenrightbig−2. Hint. Apply Parseval’s identity (or the completeness relation) to the results of Exer- cise13.3.28. 13.3.30 Showthat (a) cos−1x=π 2−4 π∞summationdisplay n=01 (2n+1)2T2n+1(x). (b) sin−1x=4 π∞summationdisplay n=01 (2n+1)2T2n+1(x). 13.4 H YPERGEOMETRIC FUNCTIONS InChapter9thehypergeometricequation10 x(1−x)y′′(x)+bracketleftbig c−(a+b+1)xbracketrightbig y′(x)−aby(x)=0 (13.124) wasintroducedasacanonicalformofalinearsecond-orderODEwithregularsingularities atx=0,1,and∞. Onesolutionis y(x)=2F1(a,b;c;x) =1+a·b cx 1!+a(a+1)b(b+1) c(c+1)x2 2!+···,c/negationslash=0,−1,−2,−3,..., (13.125) which is known as the hypergeometric function orhypergeometric series . The range of convergencefor c>a+bis|x|<1 andx=1,andisx=−1f o rc>a+b−1.Interms oftheoften-usedPochhammersymbol, (a)n=a(a+1)(a+2)···(a+n−1)=(a+n−1)! (a−1)!, (a)0=1, (13.126) thehypergeometricfunctionbecomes 2F1(a,b;c;x)=∞summationdisplay n=0(a)n(b)n (c)nxn n!. (13.127) 10This is sometimes calledGauss’ODE.Thesolutions thenbecome Gauss functions. 860 Chapter 13 More Special Functions In this form the subscripts 2 and 1 become clear. The leading subscript 2 indicates that two Pochhammer symbols appear in the numerator and the final subscript 1 indicates one Pochhammer symbol in the denominator.11(The confluent hypergeometric function 1F1 with one Pochhammer symbol in the numerator and one in the denominator appears in Section13.5.) FromtheformofEq.(13.125)weseethattheparameter cmaynotbezerooranegative integer.Ontheotherhand,if aorbequals0oranegativeinteger,theseriesterminatesand the hypergeometric function becomes a polynomial. Many more or less elementary func- tionscanberepresentedbythehypergeometricfunction.12Comparingthepowerserieswe verifythat ln(1+x)=x2F1(1,1;2;−x). (13.128) Forthecompleteellipticintegrals KandE, Kparenleftbig k2parenrightbig =integraldisplayπ/2 0parenleftbig 1−k2sin2θparenrightbig−1/2dθ=π 22F1parenleftbigg1 2,1 2;1;k2parenrightbigg ,(13.129) Eparenleftbig k2parenrightbig =integraldisplayπ/2 0parenleftbig 1−k2sinθparenrightbig1/2dθ=π 22F1parenleftbigg1 2,−1 2;1;k2parenrightbigg .(13.130) The explicit series forms and other properties of the elliptic integrals are developed in Section5.8. The hypergeometric equation as a second-order linear ODE has a second independent solution.Theusualform is y(x)=x1−c2F1(a+1−c,b+1−c;2−c;x), c/negationslash=2,3,4,.... (13.131) Ifcis an integer either the two solutions coincide or (barring a rescue by integral aor integralb) one of the solutions will blow up (see Exercise 13.4.1). In such a case the secondsolutionis expectedtoincludealogarithmicterm. Alternateforms ofthehypergeometricODEinclude parenleftbig 1−z2parenrightbigd2 dz2yparenleftbigg1−z 2parenrightbigg −bracketleftbig (a+b+1)z−(a+b+1−2c)bracketrightbigd dzyparenleftbigg1−z 2parenrightbigg −abyparenleftbigg1−z 2parenrightbigg =0, (13.132) parenleftbig 1−z2parenrightbigd2 dz2y(z2)−bracketleftbigg (2a+2b+1)z+1−2c zbracketrightbiggd dzyparenleftbig z2parenrightbig −4abyparenleftbig z2parenrightbig =0.(13.133) 11ThePochhammer symbol is often useful in other expressions involving factorials, for instance, (1−z)−a=∞summationdisplay n=0(a)nzn/n!,|z|<1. 12With threeparameters, a,b,a n dc,wecan represent almost anything. 13.4 Hypergeometric Functions 861 Contiguous Function Relations The parameters a,b, andcenter in the same way as the parameter nof Bessel, Legendre, and other special functions. As we found with these functions, we expect recurrence rela- tions involving unit changes in the parameters a,b, andc. The usual nomenclature for the hypergeometric functions, in which one parameter changes by +or−1, is a “contiguous function.” Generalizing this term to include simultaneous unit changes in more than one parameter, we find 26 functions contiguous to 2F1(a,b;c;x). Taking them two at a time, wecandeveloptheformidabletotalof325equationsamongthecontiguousfunctions.One typicalexampleis (a−b)braceleftbig c(a+b−1)+1−a2−b2+bracketleftbig (a−b)2−1bracketrightbig (1−x)bracerightbig 2F1(a,b;c;x) =(c−a)(a−b+1)b2F1(a−1,b+1;c;x) +(c−b)(a−b−1)a2F1(a+1,b−1;c;x). (13.134) AnothercontiguousfunctionrelationappearsinExercise13.4.10. Hypergeometric Representations Sincetheultrasphericalequation(13.112)inSection13.3isaspecialcaseofEq.(13.124), we see that ultraspherical functions (and Legendre and Chebyshev functions) may be ex- pressedashypergeometricfunctions.Fortheultrasphericalfunctionweobtain Cβ n(x)=(n+2β)! 2βn!β!2F1parenleftbigg −n,n+2β+1;1+β;1−x 2parenrightbigg (13.135) upon comparing its ODE with Eq. (13.124) and the power-series solutions. For Legendre andassociatedLegendrefunctionswefindsimilarly Pn(x)=2F1parenleftbigg −n,n+1;1;1−x 2parenrightbigg , (13.136) Pm n(x)=(n+m)! (n−m)!(1−x2)m/2 2mm!2F1parenleftbigg m−n,m+n+1;m+1;1−x 2parenrightbigg .(13.137) Alternateformsare P2n(x)=(−1)n(2n)! 22nn!n!2F1parenleftbigg −n,n+1 2;1 2;x2parenrightbigg =(−1)n(2n−1)!! (2n)!!2F1parenleftbigg −n,n+1 2;1 2;x2parenrightbigg , (13.138) P2n+1(x)=(−1)n(2n+1)! 22nn!n!2F1parenleftbigg −n,n+3 2;3 2;x2parenrightbigg x =(−1)n(2n+1)!! (2n)!!2F1parenleftbigg −n,n+3 2;3 2;x2parenrightbigg x. (13.139) 862 Chapter 13 More Special Functions Intermsof hypergeometricfunctions,theChebyshevfunctionsbecome Tn(x)=2F1parenleftbigg −n,n;1 2;1−x 2parenrightbigg , (13.140) Un(x)=(n+1)2F1parenleftbigg −n,n+2;3 2;1−x 2parenrightbigg , (13.141) Vn(x)=nradicalbig 1−x22F1parenleftbigg −n+1,n+1;3 2;1−x 2parenrightbigg . (13.142) The leading factors are determined by direct comparison of complete power series, com- parisonofcoefficientsofparticularpowersofthevariable,orevaluationat x=0 or1,and soon. Thehypergeometricseriesmaybeusedtodefinefunctionswithnonintegralindices.The physicalapplicationsare minimal. Exercises 13.4.1 (a) For c,aninteger,and aandbnonintegral,showthat 2F1(a,b;c;x)andx1−c2F1(a+1−c,b+1−c;2−c;x) yieldonlyonesolutiontothehypergeometricequation. (b) Whathappensif aisaninteger,say, a=−1,andc=−2? 13.4.2 Find the Legendre, Chebyshev I, and Chebyshev II recurrence relations corresponding tothecontiguoushypergeometricfunctionEq.(13.134). 13.4.3 Transformthefollowingpolynomialsintohypergeometricfunctionsofargument x2.(a) T2n(x);( b )x−1T2n+1(x);( c )U2n(x);( d)x−1U2n+1(x). ANS.(a) T2n(x)=(−1)n2F1(−n,n;1 2;x2). (b)x−1T2n+1(x)=(−1)n(2n+1)2F1(−n,n+1;3 2;x2). (c)U2n(x)=(−1)n2F1(−n,n+1;1 2;x2). (d)x−1U2n+1(x)=(−1)n(2n+2)2F1(−n,n+2;3 2;x2). 13.4.4 Derive or verify the leading factor in the hypergeometric representations of the Cheby- shevfunctions. 13.4.5 VerifythattheLegendrefunctionofthesecondkind, Qν(z), is givenby Qν(z)=π1/2ν! (ν+1 2)!(2z)ν+12F1parenleftbiggν 2+1 2,ν 2+1;ν 2+3 2;z−2parenrightbigg , |z|>1,|argz|<π, ν/negationslash=−1,−2,−3,.... 13.4.6 Analogoustotheincompletegammafunction,wemaydefineanincompletebetafunc- tionby Bx(a,b)=integraldisplayx 0ta−1(1−t)b−1dt. 13.5 Confluent Hypergeometric Functions 863 Showthat Bx(a,b)=a−1xa2F1(a,1−b;a+1;x). 13.4.7 Verifytheintegralrepresentation 2F1(a,b;c;z)=Ŵ(c) Ŵ(b)Ŵ(c−b)integraldisplay1 0tb−1(1−t)c−b−1(1−tz)−adt. Whatrestrictionsmustbeplacedontheparameters bandcandonthevariable z? Note. The restriction on |z|can be dropped—analytic continuation. For nonintegral a therealaxisinthe z-planefrom1to ∞isacutline. Hint.The integralis suspiciouslylikea betafunctionandcanbe expandedintoa series ofbetafunctions. ANS.ℜ(c)>ℜ(b)>0,and|z|<1. 13.4.8 Provethat 2F1(a,b;c;1)=Ŵ(c)Ŵ(c−a−b) Ŵ(c−a)Ŵ(c−b),c/negationslash=0,−1,−2,... c>a +b. Hint.Hereis achancetousetheintegralrepresentation,Exercise13.4.7. 13.4.9 Provethat 2F1(a,b;c;x)=(1−x)−a2F1parenleftbigg a,c−b;c;−x 1−xparenrightbigg . Hint.Tryanintegralrepresentation. Note.ThisrelationisusefulindevelopingaRodriguesrepresentationof Tn(x)(compare Exercise13.3.11). 13.4.10 Verifythat 2F1(−n,b;c;1)=(c−b)n (c)n. Hint. Here is a chance to use the contiguous function relation [2a−c+(b−a)x]· 2F1(a,b;c;x)=a(1−x)2F1(a+1,b;c;x)−(c−a)2F1(a−1,b;c;x)and mathe- maticalinduction.Alternatively,usetheintegralrepresentationandthebetafunction. 13.5 C ONFLUENT HYPERGEOMETRIC FUNCTIONS Theconfluenthypergeometricequation13 xy′′(x)+(c−x)y′(x)−ay(x)=0 (13.143) 13This is often called Kummer’s equation . The solutions, then, are Kummerfunctions . 864 Chapter 13 More Special Functions hasaregularsingularityat x=0andanirregularoneat x=∞.Itisobtainedfromthehy- pergeometricequationofSection13.4bymerging(byhand: x(1−x)→xinEq.(13.124)) two of the latter’s three singularities. One solution of the confluent hypergeometric equa- tionis y(x)=1F1(a;c;x)=M(a,c,x) =1+a cx 1!+a(a+1) c(c+1)x2 2!+···,c/negationslash=0,−1,−2,.... (13.144) This solution is convergent for all finite x(or complex z). In terms of the Pochhammer symbols,wehave M(a,c,x)=∞summationdisplay n=0(a)n (c)nxn n!. (13.145) Clearly,M(a,c,x) becomes a polynomial if the parameter ais 0 or a negative integer. Numerous more or less elementary functions may be represented by the confluent hyper- geometric function. Examples are the error function and the incomplete gamma function (from Eq. (8.69)): erf(x)=2 π1/2integraldisplayx 0e−t2dt=2 π1/2xMparenleftbigg1 2,3 2,−x2parenrightbigg , (13.146) γ(a,x)=integraldisplayx 0e−tta−1dt=a−1xaM(a,a+1,−x),ℜ(a)>0.(13.147) Clearly, this coincides with the first solution for c=a. The error function and the incom- pletegammafunctionarediscussedfurther inSection8.5. AsecondsolutionofEq. (13.143)isgivenby y(x)=x1−cM(a+1−c,2−c,x), c/negationslash=2,3,4,.... (13.148) The standard form of the second solution of Eq. (13.143) is a linear combination of Eqs. (13.144)and(13.148): U(a,c,x)=π sinπcbracketleftbiggM(a,c,x) (a−c)!(c−1)!−x1−cM(a+1−c,2−c,x) (a−1)!(1−c)!bracketrightbigg .(13.149) Note the resemblance to our definition of the Neumann function, Eq. (11.60). As with our Neumannfunction,Eq.(11.60),thisdefinitionof U(a,c,x) becomesindeterminateinthis caseforcaninteger. An alternate form of the confluent hypergeometric equation that will be useful later is obtainedbychangingtheindependentvariablefrom xtox2: d2 dx2yparenleftbig x2parenrightbig +bracketleftbigg2c−1 x−2xbracketrightbiggd dxyparenleftbig x2parenrightbig −4ayparenleftbig x2parenrightbig =0. (13.150) As with the hypergeometric functions, contiguous functions exist in which the para- metersaandcare changed by ±1. Including the cases of simultaneous changes in the 13.5 Confluent Hypergeometric Functions 865 twoparameters,14wehaveeightpossibilities.Takingtheoriginalfunctionandpairsofthe contiguousfunctions,wecandevelopatotalof 28equations.15 Integral Representations Itisfrequentlyconvenienttohavetheconfluenthypergeometricfunctionsinintegralform. Wefind(Exercise13.5.10) M(a,c,x)=Ŵ(c) Ŵ(a)Ŵ(c−a)integraldisplay1 0extta−1(1−t)c−a−1dt,ℜ(c)>ℜ(a)>0, (13.151) U(a,c,x)=1 Ŵ(a)integraldisplay∞ 0e−xtta−1(1+t)c−a−1dt,ℜ(x)>0,ℜ(a)>0. (13.152) Three important techniques for deriving or verifying integral representations are as fol- lows: 1. TransformationofgeneratingfunctionexpansionsandRodriguesrepresentations:The Bessel andLegendrefunctionsprovideexamplesof thisapproach. 2. Directintegrationtoyieldaseries:ThisdirecttechniqueisusefulforaBesselfunction representation(Exercise11.1.18)andahypergeometricintegral(Exercise13.4.7). 3. (a) Verification that the integral representation satisfies the ODE. (b) Exclusion of the other solution. (c) Verification of normalization. This is the method used in Sec- tion11.5toestablishanintegralrepresentationofthemodifiedBesselfunction Kν(z). It willwork heretoestablishEqs. (13.151)and(13.152). Bessel and Modified Bessel Functions Kummer’sfirstformula, M(a,c,x)=exM(c−a,c,−x), (13.153) isusefulinrepresentingtheBesselandmodifiedBesselfunctions.Theformulamaybever- ifiedbyseriesexpansionorbyuseofanintegralrepresentation(compareExercise13.5.10). As expected from the form of the confluent hypergeometric equation and the character of its singularities, the confluent hypergeometric functions are useful in representing a numberofthespecialfunctionsofmathematicalphysics.FortheBesselfunctions, Jν(x)=e−ix ν!parenleftbiggx 2parenrightbiggν Mparenleftbigg ν+1 2,2ν+1,2ixparenrightbigg , (13.154) whereasfor themodifiedBessel functionsof thefirst kind, Iν(x)=e−x ν!parenleftbiggx 2parenrightbiggν Mparenleftbigg ν+1 2,2ν+1,2xparenrightbigg . (13.155) 14Slaterrefers tothese as associated functions . 15Therecurrence relations for Bessel,Hermite,and Laguerrefunctions arespecial casesofthese equations. 866 Chapter 13 More Special Functions Hermite Functions TheHermitefunctionsaregivenby H2n(x)=(−1)n(2n)! n!Mparenleftbigg −n,1 2,x2parenrightbigg , (13.156) H2n+1(x)=(−1)n2(2n+1)! n!xMparenleftbigg −n,3 2,x2parenrightbigg , (13.157) usingEq. (13.150). ComparingtheLaguerreODEwiththeconfluenthypergeometricequation(13.143),we have Ln(x)=M(−n,1,x). (13.158) TheconstantisfixedasunitybynotingEq.(13.66)for x=0.FortheassociatedLaguerre functions, Lm n(x)=(−1)mdm dxmLn+m(x)=(n+m)! n!m!M(−n,m+1,x). (13.159) Alternate verification is obtained by comparing Eq. (13.159) with the power-series so- lution(Eq.(13.72)ofSection13.2).Notethatinthehypergeometricform,asdistinctfrom a Rodrigues representation, the indices nandmneed not be integers, and, if they are not integers,Lm n(x)willnotbeapolynomial. Miscellaneous Cases There are certain advantages in expressing our special functions in terms of hypergeomet- ricandconfluenthypergeometricfunctions.Ifthegeneralbehaviorofthelatterfunctionsis known,thebehaviorofthespecialfunctionswehaveinvestigatedfollowsasaseriesofspe- cialcases.Thismaybeusefulindeterminingasymptoticbehaviororevaluatingnormaliza- tion integrals. The asymptotic behavior of M(a,c,x) andU(a,c,x) may be conveniently obtained from integral representations of these functions, Eqs. (13.151) and (13.152). The further advantage is that the relations between the special functions are clarified. For in- stance,anexaminationofEqs.(13.156),(13.157),and(13.159)suggeststhattheLaguerre andHermitefunctionsarerelated. The confluent hypergeometric equation (13.143) is clearly not self-adjoint. For this and otherreasons itisconvenienttodefine Mkµ(x)=e−x/2xµ+1/2Mparenleftbigg µ−k+1 2,2µ+1,xparenrightbigg . (13.160) Thisnewfunction, Mkµ(x), is aWhittakerfunctionthatsatisfiestheself-adjointequation M′′ kµ(x)+parenleftbigg −1 4+k x+1 4−µ2 x2parenrightbigg Mkµ(x)=0. (13.161) Thecorrespondingsecondsolutionis Wkµ(x)=e−x/2xµ+1/2Uparenleftbigg µ−k+1 2,2µ+1,xparenrightbigg . (13.162) 13.5 Confluent Hypergeometric Functions 867 Exercises 13.5.1 Verifytheconfluenthypergeometricrepresentationoftheerrorfunction erf(x)=2x π1/2Mparenleftbigg1 2,3 2,−x2parenrightbigg . 13.5.2 Show that the Fresnel integrals C(x)andS(x)of Exercise 5.10.2 may be expressed in termsoftheconfluenthypergeometricfunctionas C(x)+iS(x)=xMparenleftbigg1 2,3 2,iπx2 2parenrightbigg . 13.5.3 Bydirectdifferentiationandsubstitutionverifythat y=ax−aintegraldisplayx 0e−tta−1dt=ax−aγ(a,x) satisfies xy′′+(a+1+x)y′+ay=0. 13.5.4 ShowthatthemodifiedBessel functionofthesecondkind, Kν(x), isgivenby Kν(x)=π1/2e−x(2x)νUparenleftbigg ν+1 2,2ν+1,2xparenrightbigg . 13.5.5 Show that the cosine and sine integrals of Section 8.5 may be expressed in terms of confluenthypergeometricfunctionsas Ci(x)+isi(x)=−eixU(1,1,−ix). ThisrelationisusefulinnumericalcomputationofCi (x)andsi(x)forlargevaluesof x. 13.5.6 Verify the confluent hypergeometric form of the Hermite polynomial H2n+1(x) (Eq. (13.157))byshowingthat (a)H2n+1(x)/xsatisfies the confluent hypergeometric equation with a=−n,c=3 2 andargument x2, (b) lim x→0H2n+1(x) x=(−1)n2(2n+1)! n!. 13.5.7 Showthatthecontiguousconfluenthypergeometricfunctionequation (c−a)M(a−1,c,x)+(2a−c+x)M(a,c,x)−aM(a+1,c,x)=0 leadstotheassociatedLaguerrefunctionrecurrencerelation(Eq. (13.75)). 13.5.8 VerifytheKummertransformations: (a)M(a,c,x)=exM(c−a,c,−x) (b)U(a,c,x)=x1−cU(a−c+1,2−c,x). 868 Chapter 13 More Special Functions 13.5.9 Provethat (a)dn dxnM(a,c,x)=(a)n (b)nM(a+n,b+n,x), (b)dn dxnU(a,c,x)=(−1)n(a)nU(a+n,c+n,x). 13.5.10 Verifythefollowingintegralrepresentations: (a)M(a,c,x)=Ŵ(c) Ŵ(a)Ŵ(c−a)integraldisplay1 0extta−1(1−t)c−a−1dt,ℜ(c)>ℜ(a)>0. (b)U(a,c,x)=1 Ŵ(a)integraldisplay∞ 0e−xtta−1(1+t)c−a−1dt,ℜ(x)>0,ℜ(a)>0. Underwhatconditionscanyouaccept ℜ(x)=0 inpart(b)? 13.5.11 Fromtheintegralrepresentationof M(a,c,x) , Exercise13.5.10(a),showthat M(a,c,x)=exM(c−a,c,−x). Hint. Replace the variable of integration tby 1−sto release a factor exfrom the integral. 13.5.12 Fromtheintegralrepresentationof U(a,c,x) ,Exercise13.5.10(b),showthattheexpo- nentialintegralisgivenby E1(x)=e−xU(1,1,x). Hint.Replacethevariableofintegration tinE1(x)byx(1+s). 13.5.13 From the integral representations of M(a,c,x) andU(a,c,x) in Exercise 13.5.10 de- velopasymptoticexpansionsof (a)M(a,c,x) ,(b)U(a,c,x) . Hint.Youcanusethetechniquethatwasemployedwith Kν(z), Section11.6. ANS.(a)Ŵ(c) Ŵ(a)ex xc−abraceleftbigg 1+(1−a)(c−a) 1!x+ (1−a)(2−a)(c−a)(c−a+1) 2!x2+···bracerightbigg (b)1 xabraceleftbigg 1+a(1+a−c) 1!(−x)+a(a+1)(1+a−c)(2+a−c) 2!(−x)2+···bracerightbigg . 13.5.14 ShowthattheWronskianofthetwoconfluenthypergeometricfunctions M(a,c,x) and U(a,c,x) is givenby MU′−M′U=−(c−1)! (a−1)!ex xc. Whathappensif ais0or anegativeinteger? 13.6 Mathieu Functions 869 13.5.15 The Coulomb wave equation (radial part of the Schrödinger equation with Coulomb potential)is d2y dρ2+bracketleftbigg 1−2η ρ−L(L+1) ρ2bracketrightbigg y=0. Showthataregularsolution y=FL(η,ρ)isgivenby FL(η,ρ)=CL(η)ρL+1e−iρM(L+1−iη,2L+2,2iρ). 13.5.16 (a) Showthattheradialpartofthehydrogenwavefunction,Eq.(13.81),maybewrit- tenas e−αr/2(αr)LL2L+1 n−L−1(αr) =(n+L)! (n−L−1)!(2L+1)!e−αr/2(αr)LM(L+1−n,2L+2,αr). (b) Itwasassumedpreviouslythatthetotal(kinetic +potential)energy Eoftheelec- tron was negative. Rewrite the (unnormalized) radial wave function for the free electron, E>0. ANS.eiαr/2(αr)LM(L+1−in,2L+2,−iαr), outgoing wave. This representation provides a powerful alternative tech- nique for the calculation of photoionization and recombina- tioncoefficients. 13.5.17 Evaluate (a)integraldisplay∞ 0bracketleftbig Mkµ(x)bracketrightbig2dx,(b)integraldisplay∞ 0bracketleftbig Mkµ(x)bracketrightbig2dx x, (c)integraldisplay∞ 0bracketleftbig Mkµ(x)bracketrightbig2dx x1−a, where 2µ=0,1,2,...,k−µ−1 2=0,1,2,...,a>−2µ−1. ANS.(a) (2µ)!2k.(b)(2µ)!.( c)(2µ)!(2k)a. 13.6 M ATHIEU FUNCTIONS When PDEs such as Laplace’s, Poisson’s, and the wave equation are solved with cylin- drical or spherical boundary conditions by separating variables in polar coordinates, we find radial solutions, which are the Bessel functions of Chapter 11, and angular solutions, which are sin mϕ,cosmϕin cylindrical cases and spherical harmonics in spherical cases. Examples are electromagnetic waves in resonant cavities, vibrating circular drumheads, andcoaxialwaveguides. When in such cylindrical problems the circular boundary condition becomes elliptical we are led to the angular and radial Mathieu functions, which, therefore, might be called elliptic cylinder functions. In fact, in 1868 Mathieu developed the leading terms of series solutionsof thevibratingellipticaldrumhead,andWhittaker andothersin theearly 1900s derivedhigher-orderterms aswell. 870 Chapter 13 More Special Functions Here our goal is to give an introduction to the rich and complex properties of Mathieu functions. Separation of Variables in Elliptical Coordinates Ellipticalcylinder coordinates ξ,η,z, whichare appropriate for ellipticalboundary condi- tions,areexpressedinrectangularcoordinatesas x=ccoshξcosη, y=csinhξsinη, z=z, (13.163) 0≤ξ<∞,0≤η≤2π, where the parameter 2 c>0 is the distance between the foci of the confocal ellipses de- scribed by these coordinates (Fig. 13.7). We want to show that in the limit c→0 the foci of the ellipses coalesce to the center of circles. We work at constant z-coordinate mostly, z=0,say.Indeedforfixedradialvariable ξ=const.wecaneliminatetheangularvariable ηtoobtainfromEq. (13.163) x2 c2cosh2ξ+y2 c2sinh2ξ=1, (13.164) describing confocal ellipses centered at the origin of the x,y-plane with major and minor half-axes a=ccoshξ, b=csinhξ, (13.165) respectively.Since b a=tanhξ=radicalBigg 1−1 cosh2ξ≡radicalbig 1−e2, (13.166) the eccentricity e=1/coshξof the ellipse with 0 ≤e≤1, and the distance between the foci 2ae=2c, providing a geometrical interpretation of the radial coordinate ξand the FIGURE 13.7Ellipticalcoordinates ξ,η. 13.6 Mathieu Functions 871 parameter c.A sξ→∞,e→0 and the ellipses become circles, which is indicated in Fig. 13.7. As ξ→0, the ellipse becomes more elongated until, at ξ=0, it has shrunk to thelinesegmentbetweenthefoci. Whenη=const.weeliminate ξtofindconfocalhyperbolas x2 c2cos2η−y2 c2sin2η=1, (13.167) whicharealsoplottedinFig.13.7.Differentiatingtheellipse,weobtain xdx cosh2ξ+ydy sinh2ξ=0, (13.168) which means that the tangent vector (dx,dy)of the ellipse is perpendicular to the vector (x cosh2ξ,y sinh2ξ). Forthehyperbolatheorthogonalityconditionis xdx cos2η−ydy sin2η=0, (13.169) so the scalar product of the ellipse and hyperbola tangent vectors at each of their intersec- tionpoints (x,y)of Eq.(13.163)obey x2 cosh2ξcos2η−y2 sinh2ξsin2η=c2−c2=0. (13.170) This means that these confocal ellipses and hyperbolas form an orthogonal coordinate system,inthesenseofSection2.1.Toextractthescalefactors hξ,hηfromthedifferentials oftheellipticalcoordinates dx=csinhξcosηdξ−ccoshξsinηdη, (13.171) dy=ccoshξsinηdξ+csinhξcosηdη, wesumtheirsquares, finding dx2+dy2=c2parenleftbig sinh2ξcos2η+cosh2ξsin2ηparenrightbigparenleftbig dξ2+dη2parenrightbig =c2parenleftbig cosh2ξ−cos2ηparenrightbigparenleftbig dξ2+dη2parenrightbig ≡h2 ξdξ2+h2 ηdη2(13.172) andyielding hξ=hη=cparenleftbig cosh2ξ−cos2ηparenrightbig1/2. (13.173) Note that there is no cross term involving dξdη, showing again that we are dealing with orthogonalcoordinates. Nowweare readytoderiveMathieu’sdifferentialequations. 872 Chapter 13 More Special Functions Example 13.6.1 ELLIPTICAL DRUM We consider vibrations of an elliptical drumhead with vertical displacement z=z(x,y,t) governedbythewaveequation ∂2z ∂x2+∂2z ∂y2=1 v2∂2z ∂t2, (13.174) wherethevelocitysquared v2=T/ρwithtension Tandmassdensity ρisaconstant.We firstseparatetheharmonictimedependence,writing z(x,y,t)=u(x,y)w(t), (13.175) wherew(t)=cos(ωt+δ), withωthe frequency and δa constant phase. Substituting this functionzintoEq.(13.174)yields 1 uparenleftbigg∂2u ∂x2+∂2u ∂y2parenrightbigg =1 v2w∂2w ∂t2=−ω2 v2=−k2=const., (13.176) that is, the two-dimensional Helmholtz equation for the displacement u. We now use Eq. (2.22) to convert the Laplacian ∇2to the elliptical coordinates, where we drop the z- coordinate.Thisgives ∂2u ∂x2+∂2u ∂y2+k2u=1 h2 ξparenleftbigg∂2u ∂ξ2+∂2u ∂η2parenrightbigg +k2u=0, (13.177) thatis, theHelmholtzequationinelliptical ξ,ηcoordinates, ∂2u ∂ξ2+∂2u ∂η2+c2k2parenleftbig cosh2ξ−cos2ηparenrightbig u=0. (13.178) Lastly,weseparate ξandη,writingu(ξ,η)=R(ξ)/Phi1(η) , whichyields 1 Rd2R dξ2+c2k2cosh2ξ=c2k2cos2η−1 /Phi1d2/Phi1 dη2=λ+1 2c2k2, (13.179) whereλ+c2k2/2 is the separation constant. Writing cosh2 ξ,cos2ηinstead of cosh2ξ, cos2η(which motivates the special form of the separation constant in Eq. (13.179)) we findthelinear,second-orderODE d2R dξ2−(λ−2qcosh2ξ)R(ξ)=0,q=1 4c2k2, (13.180) whichisalsocalledthe radialMathieuequation ,and d2/Phi1 dη2+(λ−2qcos2η)/Phi1(η)=0, (13.181) theangular,o rmodified,Mathieu equation . Note that the eigenvalue λ(q)is a function of the continuous parameter qin the Mathieu ODEs. It is this parameter dependence that complicates the analysis of Mathieu functions and makes them among the most difficult specialfunctionsusedinphysics. /squaresolid 13.6 Mathieu Functions 873 Clearly, all finite points are regular points of both ODEs, while infinity is an essential singularity for both ODEs, which are of the Sturm–Liouville type (Chapter 10) with coef- ficientfunctions p≡1 and q(ξ)=−λ+2qcosh2ξ, q(η)=λ−2qcos2η. (13.182) (These functions qmust not be confused with the parameter q.) As a consequence, their solutionsformorthogonalsetsoffunctions.Thesubstitution η→iξtransformstheangular totheradialMathieuODE,so theirsolutionsarecloselyrelated. Using the Lindemann–Stieltjes substitution z=cos2η, dz/dη=−sin2η, the angular Mathieu ODE is transformed into an ODE with coefficients that are algebraic in the vari- ablez(usingd dη=dz dηd dz=−sin2ηd dzandd2 dη2=−2cos2ηd dz+sin22ηd2 dz2): 4z(1−z)d2/Phi1 dz2+2(1−2z)d/Phi1 dz+bracketleftbig λ+2q(1−2z)bracketrightbig /Phi1=0. (13.183) This ODE has regular singularities at z=0 andz=1, whereas the point at infinity is an essential singularity (Chapter 9). By comparison, the hypergeometric ODE has three regular singularities. But not all ODEs with two regular singularities and one essential singularitycanbetransformedintoanODEof theMathieutype. Example 13.6.2 THEQUANTUM PENDULUM A planependulumof length landmass mwithgravitationalpotential V(θ)=−mglcosθ iscalleda quantumpendulum if itswavefunction /Psi1obeystheSchrödingerequation −¯h2 2ml2d2/Psi1 dθ2+bracketleftbig V(θ)−Ebracketrightbig /Psi1=0, (13.184) where the variable θis the angular displacement from the vertical direction. (For fur- ther details and illustrations we refer to Gutiérrez-Vega et al. in the Additional Readings.) A boundary condition applies to /Psi1so as to be single-valued; that is, /Psi1(θ+2π)=/Psi1(θ). Substituting θ=2η, λ=8Eml2 ¯h2,q=−4m2gl3 ¯h2(13.185) into the Schrödinger equation yields the angular Mathieu ODE for /Psi1(2(η+π))= /Psi1(2η). /squaresolid For many other applications involving Mathieu functions we refer to Ruby in the Addi- tionalReadings. Our main focus will be on the solutions of the angular Mathieu ODE, which has the importantpropertythatitscoefficientfunctionisperiodicwithperiod π. 874 Chapter 13 More Special Functions General Properties of Mathieu Functions InphysicsapplicationstheangularMathieufunctionsarerequiredtobesingle-valued,that is, periodic with period 2 π. Let us start with some nomenclature. Since Mathieu’s ODEs are invariant under parity ( η→−η), Mathieu functions have definite parity. Those of odd paritythathaveperiod2 πand,forsmall q,startwithsin (2n+1)ηarecalledse 2n+1(η,q), withnan integer, n=0,1,2,...(se is short for sine-elliptic). Mathieu functions of odd parity and period πthat start with sin2 nηfor small qare called se 2n(η,q), withn= 1,2,....Mathieu functions of even parity, period πthat start with cos2 nηfor small q are called ce 2n(η,q)(ce is short for cosine-elliptic), while those with period 2 πthat start withcos(2n+1)η,n=0,1,...,forsmall qarecalledce 2n+1(η,q).Inthelimitwherethe parameter q→0 (and theMathieuODEbecomestheclassicalharmonicoscillatorODE), Mathieufunctionsreducetothesetrigonometricfunctions. The periodicity condition /Phi1(η+2π)=/Phi1(η)is sufficient to determine a set of eigen- valuesλin terms of q. An elementaryanalogof this result is the fact that a solution of the classicalharmonicoscillatorODE u′′(η)+λu(η)=0hasperiod2 πif,andonlyif, λ=n2 is the square of an integer. Such problems will be pursued in Section 14.7 as applications ofFourierseries. Example 13.6.3 RADIAL MATHIEU FUNCTIONS Upon replacing the angular elliptic variable η→iξ, the angular Mathieu ODE, Eq. (13.181), becomes the radial ODE, Eq. (13.180). This motivates the definitions of radialMathieufunctionsas Ce2n+p(ξ,q)=ce2n+p(iξ,q), p =0,1;n=0,1,..., Se2n+p(ξ,q)=−ise2n+p(iξ,q), p =0,1;n=1,2,.... Because these functions are differentiable, they correspond to the regular solutions of the radialMathieuODE.Ofcourse,theyarenolongerperiodicbutareoscillatory(Fig.13.8). In physical problems involving elliptical coordinates, the radial Mathieu ODE, Eq. (13.180), plays a role corresponding to Bessel’s ODE in cylindrical geometry. Be- cause there are four families of independent Bessel functions—the regular solutions Jn and irregular Neumann functions Nn, along with the modified Bessel functions Inand Kn—we expect four kinds of radial Mathieu functions. Because of parity, the solutions splitintoevenandoddMathieufunctionsandsothereareeightkinds. For q>0, Je2n(ξ,q)=Ce2n(ξ,q), Je2n+1(ξ,q)=Ce2n+1(ξ,q), Jo2n(ξ,q)=Se2n(ξ,q), Jo2n+1(ξ,q)=Se2n+1(ξ,q), regularorfirst kind ; Nen(ξ,q),Non(ξ,q), irregularorsecondkind ; forq<0,thesolutionsoftheradialMathieuODEaredenotedby Ien(ξ,q),Ion(ξ,q), regularorfirst kind , Ken(ξ,q),Kon(ξ,q), irregularorsecondkind 13.6 Mathieu Functions 875 FIGURE 13.8RadialMathieufunctions: q=1 (solidline), q=2 (dashed line),q=3 (dottedline).(FromGutiérrez-Vega etal.,Am. J.Phys. 71: 233(2003).) and are known as the evanescent radial Mathieu functions . Mathieu functions corre- sponding to the Hankel functions can be similarly defined. In Fig. 13.8 some of them are plotted. In applications such as a vibrating drumhead with elliptical boundary conditions (see Example13.6.1),thesolutioncanbeexpandedinevenandoddMathieufunctions: zen≡Jen(ξ,q)cen(η,q)cos(ωnt), m≥0, zon≡Jon(ξ,q)sen(η,q)cos(ωnt), m≥1. TheyobeyDirichletboundaryconditions, zen(ξ0,η,t)=0=zon(ξ0,η,t),whichholdpro- vided the radial functions satisfy Je n(ξ0,q)=0=Jon(ξ0,q)at the elliptical boundary, whereξ=ξ0. Whenthefocaldistance c→0,theangularMathieufunctionsbecometheconventional trigonometricfunctions,whiletheradialMathieufunctionsbecomeBesselfunctions. In the case of oscillations of a confocal annular elliptic lake, the modes have to include theMathieufunctionsofthesecondkindandare thusgivenby zen≡bracketleftbig AJen(ξ,q)+BNen(ξ,q)bracketrightbig cen(η,q)cos(ωnt), m≥0, zon≡bracketleftbig AJon(ξ,q)+BNon(ξ,q)bracketrightbig sen(η,q)cos(ωnt), m≥1, withA,Bconstants. These standing wave solutions must obey Neumann boundary con- ditions at the inner ( ξ=ξ0) and outer ( ξ=ξ1) elliptical boundaries; that is, the normal 876 Chapter 13 More Special Functions derivatives (a prime denotes d/dξ)o fz enand zonvanish at each point of the boundaries. For even modes, we have ze′ n(ξ0,η,t)=0=ze′n(ξ1,η,t). The implied radial constraints are similar to Eqs. (11.81) and (11.82) of Example 11.3.1. Numerical examples and plots, alsofortravelingwaves,aregiveninGutiérrez-Vega etal.intheAdditionalReadings. /squaresolid ForzerosofMathieufunctions,theirasymptoticexpansions,andamorecompletelisting offormulaswerefertoAbramowitzandStegun(AMS-55)intheAdditionalReadings, Am. J.Phys.71,JahnkeandEmdeandGradshteynandRyzhikintheAdditionalReadings. To illustrate and support the nomenclature, we want to show16that there is an angular Mathieufunctionthatis •even inηandofperiod πifandonlyif /Phi1′ 1(π/2)=0; •oddandof period πifandonlyif /Phi12(π/2)=0; •evenandofperiod 2 πif andonlyif /Phi11(π/2)=0; •oddandof period 2 πifandonlyif /Phi1′2(π/2)=0, where/Phi11(η),/Phi12(η)are two linearly independent solutions of the angular Mathieu ODE sothat /Phi11(0)=1,/Phi1′1(0)=0;/Phi12(0)=0,/Phi1′2(0)=1. (13.186) Since the Mathieu ODE is a linear second-order ODE, we know (Chapter 9) that these initial conditions are realistic. The first case just given corresponds to ce 2n(η,q), with /Phi1′ 1(π/2)=−2nsin2nη|η=π/2+···=0f o rn=1,2,....The second is the se 2n(η,q), with/Phi12(π/2)=sin2nη|π/2+···=0.Thethirdcaseisthece 2n+1(η,q),with/Phi11(π/2)= cos(2n+1)π/2+···=0.Thefourthcaseisthese 2n+1(η,q). The key to the proof is Floquet’s approach to linear second-order ODEs with periodic coefficient functions, such as Mathieu’s angular ODE or the simple pendulum (Exercise 13.6.1). If /Phi11(η),/Phi12(η)are two linearly independent solutions of the ODE, any other solution/Phi1canbeexpressedas /Phi1(η)=c1/Phi11(η)+c2/Phi12(η), (13.187) withconstants c1,c2.Now ,/Phi1k(η+2π)arealsosolutionsbecausesuchanODEisinvariant underthetranslation η→η+2π, andinparticular /Phi11(η+2π)=a1/Phi11(η)+a2/Phi12(η), /Phi12(η+2π)=b1/Phi11(η)+b2/Phi12(η), (13.188) withconstants ai,bj. SubstitutingEq.(13.188)intoEq. (13.187)weget /Phi1(η+2π)=(c1a1+c2b1)/Phi11(η)+(c2b2+c1a2)/Phi12(η), (13.189) wheretheconstants cicanbechosenassolutionsoftheeigenvalueequations a1c1+b1c2=λc1, a2c1+b2c2=λc2. (13.190) 16SeeHochstadt in the Additional Readings. 13.6 Mathieu Functions 877 ThenFloquet’stheorem statesthat /Phi1(η+2π)=λ/Phi1(η), whereλisarootof vextendsinglevextendsinglevextendsinglevextendsinglea1−λb 1 a2b2−λvextendsinglevextendsinglevextendsinglevextendsingle=0. (13.191) Au s e f u l corollary is obtained if we define µandybyλ=exp(2πµ)andy(η)= exp(−µη)/Phi1(η),s o y(η+2π)=e−µηe−2πµ/Phi1(η+2π)=e−µη/Phi1(η)=y(η). (13.192) Thus,/Phi1(η)=eµηy(η), withyaperiodicfunctionof ηwithperiod 2 π. LetusapplyFloquet’sargumenttothe /Phi1k(η+π),whicharealsosolutionsofMathieu’s ODEbecausethelatteris invariantunderthetranslation η→η+π. UsingthespecialvaluesinEq. (13.186)weknowthat /Phi11(η+π)=/Phi11(π)/Phi11(η)+/Phi1′ 1(π)/Phi12(η), /Phi12(η+π)=/Phi12(π)/Phi11(η)+/Phi1′2(π)/Phi12(η), (13.193) because these linear combinations of /Phi1k(η)are solutions of Mathieu’s ODE with the cor- rectvalues /Phi1i(η+π),/Phi1′ i(η+π)forη=0.Therefore, /Phi1i(η+π)=λi/Phi1i(η), (13.194) wherethe λiaretheroots of vextendsinglevextendsinglevextendsinglevextendsingle/Phi11(π)−λ/Phi1 2(π) /Phi1′ 1(π) /Phi1′2(π)−λvextendsinglevextendsinglevextendsinglevextendsingle=0. (13.195) Theconstantterminthecharacteristicpolynomialis givenbytheWronskian Wparenleftbig /Phi11(η),/Phi12(η)parenrightbig =C, (13.196) aconstantbecausethecoefficientof d/Phi1/dηintheangularMathieuODEvanishes,imply- ingdW/dη=0.In fact, usingEq. (13.186), Wparenleftbig /Phi11(0),/Phi12(0)parenrightbig =/Phi11(0)/Phi1′ 2(0)−/Phi1′1(0)/Phi12(0)=1 =Wparenleftbig /Phi11(π),/Phi12(π)parenrightbig , (13.197) sotheeigenvalueEq. (13.195)for λbecomes parenleftbig /Phi11(π)−λparenrightbigparenleftbig /Phi1′2(π)−λparenrightbig −/Phi12(π)/Phi1′1(π)=0 =λ2−bracketleftbig /Phi11(π)+/Phi1′2(π)bracketrightbig λ+1, (13.198) withλ1·λ2=1 andλ1+λ2=/Phi11(π)+/Phi1′2(π). If|λ1|=|λ2|=1,thenλ1=exp(iφ)andλ2=exp(−iφ),soλ1+λ2=2cosφ.Forφ/negationslash= 0,π,2π,...this case corresponds to |/Phi11(π)+/Phi1′2(π)|<2, where both solutions remain bounded as η→∞in steps of πusing Eq. (13.194). These cases do not yield periodic Mathieu functions, and this is also the case when |/Phi11(π)+/Phi1′2(π)|>2. Ifφ=0, that is,λ1=1=λ2is a double root, then the /Phi1ihave period πand|/Phi11(π)+/Phi1′2(π)|=2. If φ=π, that is,λ1=−1=λ2is again a double root, then |/Phi11(π)+/Phi1′2(π)|=−2 and the /Phi1ihaveperiod 2 πwith/Phi1i(η+π)=−/Phi1i(η). 878 Chapter 13 More Special Functions Because the angular Mathieu ODE is invariant under a parity transformation η→−η, itisconvenienttoconsidersolutions /Phi1e(η)=1 2bracketleftbig /Phi1(η)+/Phi1(−η)bracketrightbig ,/Phi1 o(η)=1 2bracketleftbig /Phi1(η)−/Phi1(−η)bracketrightbig (13.199) of definite parity, which obey the same initial conditions as /Phi1i. We now relabel /Phi1e→ /Phi11,/Phi1o→/Phi12, taking/Phi11to be even and /Phi12to be odd under parity. These solutions of definite parity of Mathieu’s ODE are called Mathieu functions and are labeled according toournomenclaturediscussedearlier. If/Phi11(η)has period π, then/Phi1′ 1(η+π)=/Phi1′1(η)also has period πbut is odd under parity.Substituting η=−π/2 weobtain /Phi1′1parenleftbiggπ 2parenrightbigg =/Phi1′ 1parenleftbigg −π 2parenrightbigg =−/Phi1′ 1parenleftbiggπ 2parenrightbigg ,so/Phi1′ 1parenleftbiggπ 2parenrightbigg =0. (13.200) Conversely,if /Phi1′ 1(π/2)=0,then/Phi11(η)hasperiod π. Toseethis,weuse /Phi11(η+π)=c1/Phi11(η)+c2/Phi12(η). (13.201) This expansion is valid because /Phi11(η+π)is a solution of the angular Mathieu ODE. We now determine the coefficients ci, settingη=−π/2, and recall that /Phi11and/Phi1′2are even underparity,whereas /Phi12and/Phi1′1areodd.This yields /Phi11parenleftbiggπ 2parenrightbigg =c1/Phi11parenleftbiggπ 2parenrightbigg −c2/Phi12parenleftbiggπ 2parenrightbigg , (13.202) /Phi1′ 1parenleftbiggπ 2parenrightbigg =−c1/Phi1′ 1parenleftbiggπ 2parenrightbigg +c2/Phi1′ 2parenleftbiggπ 2parenrightbigg . Since/Phi1′ 1(π/2)=0,/Phi1′2(π/2)/negationslash=0, or the Wronskian would vanish and /Phi12∼/Phi11would follow. Hence c2=1 follows from the second equation and c1=1 from the first. Thus, /Phi11(η+π)=/Phi11(η). Theotherbulletedcaseslistedearliercanbeprovedsimilarly. Because the Mathieu ODEs are of the Sturm–Liouville type, Mathieu functions repre- sent orthogonal systems of functions. So, for m,nnonnegative integers, the orthogonality relationsandnormalizationsare integraldisplayπ −πcemcendη=integraldisplayπ −πsemsendη=0,ifm/negationslash=n; integraldisplayπ −πcemsendη=0; (13.203) integraldisplayπ −π[ce2n]2dη=integraldisplayπ −π[se2n]2dη=π,ifn≥1;integraldisplayπ 0bracketleftbig ce0(η,q)bracketrightbig2dη=π. If a function f(η)is periodic with period π, then it can be expanded in a series of orthog- onalMathieufunctionsas f(η)=1 2a0ce0(η,q)+∞summationdisplay n=1bracketleftbig ance2n(η,q)+bnse2n(η,q)bracketrightbig (13.204) 13.6 Additional Readings 879 with an=1 πintegraldisplayπ −πf(η)ce2n(η,q)dη, n ≥0; (13.205) bn=1 πintegraldisplayπ −πf(η)se2n(η,q)dη, n ≥1. Similarexpansionsexistfor functionsof period 2 πintermsof ce 2n+1andse2n+1. SeriesexpansionsofMathieufunctionswillbederivedinSection14.7. Exercises 13.6.1 For the simple pendulum ODE of Section 5.8, apply Floquet’s method and derive the propertiesofitssolutionssimilartothosemarkedbybulletsbeforeEq. (13.186). 13.6.2 Derive a Mathieu function analog for the Rayleigh expansion of a plane wave forcos(kcosηcosθ)and sin(kcosηcosθ). AdditionalReadings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions , Applied Mathematics Series- 55 (AMS-55). Washington, DC: National Bureau of Standards (1964). Paperback edition, New York: Dover (1974). Chapter 22 is a detailed summary of the properties and representations of orthogonal polynomials. Other chapters summarize properties of Bessel, Legendre, hypergeometric, and confluent hypergeometric functions and much more. Buchholz, H., The Confluent Hypergeometric Function . New York: Springer-Verlag (1953); translated (1969). BuchholzstronglyemphasizestheWhittakerratherthantheKummerforms.Applicationstoavarietyofother transcendental functions. Erdelyi, A., W. Magnus, F. Oberhettinger, and F. G. Tricomi, Higher Transcendental Functions ,3v o l s .N e w York: McGraw-Hill (1953). Reprinted Krieger (1981). A detailed, almost exhaustive listing of the properties of thespecial functions of mathematicalphysics. Fox,L.andI.B.Parker, ChebyshevPolynomialsinNumericalAnalysis .Oxford:OxfordUniversityPress(1968). Adetailed,thorough,butveryreadableaccountofChebyshevpolynomialsandtheirapplicationsinnumerical analysis. Gradshteyn, I. S.,and I. M.Ryzhik, Table of Integrals, Series and Products , NewYork: AcademicPress (1980). Gutiérrez-Vega, J. C., R. M. Rodríguez-Dagnino, M. A. Meneses-Nava and S. Chávez-Cerda, A m .J .P h y s . 71: 233 (2003). Hochstadt, H., Special Functions of Mathematical Physics . New York: Holt, Rinehart and Winston (1961), reprinted Dover (1986). Jahnke,E.,andF.Emde, Table of Functions . Leipzig:Teubner (1933); NewYork: Dover (1943). Lebedev,N.N., SpecialFunctionsandtheirApplications (translatedbyR.A.Silverman).EnglewoodCliffs,NJ: Prentice-Hall (1965). Paperback, NewYork: Dover (1972). Luke,Y .L., TheSpecialFunctionsandTheirApproximations .NewYork:AcademicPress(1969).Twovolumes: Volume 1 is a thorough theoretical treatment of gamma functions, hypergeometric functions, confluent hy- pergeometric functions, and related functions. Volume 2 develops approximations and other techniques for numerical work. Luke, Y. L., Mathematical Functions and Their Approximations . New York: Academic Press (1975). This is an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs and Mathematical Tables(AMS-55). 880 Chapter 13 More Special Functions Mathieu,E., J.de Math. Pures etAppl. 13: 137–203 (1868). McLachlan,N.W., Theory and Applications of Mathieu Functions . Oxford, UK:Clarendon Press (1947). Magnus,W.,F.Oberhettinger,andR.P.Soni, FormulasandTheoremsfortheSpecialFunctionsofMathematical Physics. NewYork: Springer (1966). An excellent summary of just what the title says, including the topics of Chapters 10 to 13. Rainville, E. D., Special Functions . New York: Macmillan (1960), reprinted Chelsea (1971). This book is a coherent,comprehensiveaccountofalmostallthespecialfunctionsofmathematicalphysicsthatthereaderis likelytoencounter. Rowland, D.R., Am.J .Ph ys. 72: 758–766 (2004). Ruby, L., Am.J .Ph ys. 64: 39–44 (1996). Sansone, G., Orthogonal Functions (translated by A. H. Diamond). New York: Interscience (1959). Reprinted Dover (1991). Slater, L. J., Confluent Hypergeometric Functions . Cambridge, UK: Cambridge University Press (1960). This is a clear and detailed development of the properties of the confluent hypergeometric functions and of relations of the confluent hypergeometric equation to other ODEsofmathematical physics. Sneddon, I. N., Special Functions of Mathematical Physicsand Chemistry , 3rd ed.NewYork: Longman (1980). Whittaker,E.T.,andG.N.Watson, ACourseofModernAnalysis .Cambridge,UK:CambridgeUniversityPress, reprinted (1997). Theclassic text on special functions andreal andcomplex analysis. CHAPTER 14 FOURIER SERIES 14.1 G ENERAL PROPERTIES Periodicphenomenainvolvingwaves,rotatingmachines(harmonicmotion),orotherrepet- itive driving forces are described by periodic functions. Fourier series are a basic tool for solving ordinary differential equations (ODEs) and partial differential equations (PDEs) with periodic boundary conditions. Fourier integrals for nonperiodic phenomena are de- velopedinChapter15.Thecommonnameforthefieldis Fourieranalysis . A Fourier series is defined as an expansion of a function or representation of a function inaseriesof sinesandcosines,suchas f(x)=a0 2+∞summationdisplay n=1ancosnx+∞summationdisplay n=1bnsinnx. (14.1) The coefficients a0,an, andbnare related to the periodic function f(x)by definite inte- grals: an=1 πintegraldisplay2π 0f(x)cosnxdx, (14.2) bn=1 πintegraldisplay2π 0f(x)sinnxdx, n=0,1,2,..., (14.3) which are subject to the requirement that the integrals exist. Notice that a0is singled out for special treatment by the inclusion of the factor1 2. This is done so that Eq. (14.2) will applytoall an,n=0a sw e l la s n>0. The conditions imposed on f(x)to make Eq. (14.1) valid are that f(x)have only a finitenumberoffinitediscontinuitiesandonlyafinitenumberofextremevalues,maxima, and minima in the interval [0,2π].1Functions satisfying these conditions may be called 1Theseconditions are sufficient but notnecessary . 881 882 Chapter 14 Fourier Series piecewise regular . The conditions themselves are known as the Dirichlet conditions. Al- thoughtherearesomefunctionsthatdonotobeytheseDirichletconditions,theymaywell be labeled pathological for purposes of Fourier expansions. In the vast majority of physi- cal problems involving a Fourier series these conditions will be satisfied. In most physical problemsweshallbeinterestedinfunctionsthataresquareintegrable(intheHilbertspace L2of Section 10.4). In this space the sines and cosines form a complete orthogonal set. AndthisinturnmeansthatEq. (14.1) isvalid,inthesenseofconvergenceinthemean. Expressing cos nxand sinnxinexponentialform, wemayrewriteEq.(14.1) as f(x)=∞summationdisplay n=−∞cneinx, (14.4) inwhich cn=1 2(an−ibn), c−n=1 2(an+ibn), n> 0, (14.5a) and c0=1 2a0. (14.5b) Complex Variables — Abel’s Theorem Considerafunction f(z)representedbyaconvergentpowerseries f(z)=∞summationdisplay n=0Cnzn=∞summationdisplay n=0Cnrneinθ. (14.6) This is our Fourier exponential series, Eq. (14.4). Separating real and imaginary parts we get u(r,θ)=∞summationdisplay n=0Cnrncosnθ, v(r,θ) =∞summationdisplay n=1Cnrnsinnθ, (14.7a) the Fourier cosine and sine series. Abel’s theorem asserts that if u(1,θ)andv(1,θ)are convergentfor agiven θ,then u(1,θ)+iv(1,θ)=lim r→1fparenleftbig reiθparenrightbig . (14.7b) Anapplicationof thisappearsasExercise14.1.9andinExample14.1.1. Example 14.1.1 SUMMATION OF A FOURIER SERIES Usually in this chapter we shall be concerned with finding the coefficients of the Fourier expansion of a known function. Occasionally, we may wish to reverse this process and determinethefunctionrepresentedbya givenFourierseries. 14.1 General Properties 883 Consider the seriessummationtext∞ n=1(1/n)cosnx,x∈(0,2π). Since this series is only condition- allyconvergent(anddivergesat x=0),wetake ∞summationdisplay n=1cosnx n=lim r→1∞summationdisplay n=1rncosnx n, (14.8) absolutely convergent for |r|<1. Our procedure is to try forming power series by trans- formingthetrigonometricfunctionsintoexponentialform: ∞summationdisplay n=1rncosnx n=1 2∞summationdisplay n=1rneinx n+1 2∞summationdisplay n=1rne−inx n. (14.9) Now, these power series may be identified as Maclaurin expansions of −ln(1−z),z= reix,re−ix(Eq. (5.95)), and ∞summationdisplay n=1rncosnx n=−1 2bracketleftbig lnparenleftbig 1−reixparenrightbig +lnparenleftbig 1−re−ixparenrightbigbracketrightbig =−lnbracketleftbigparenleftbig 1+r2parenrightbig −2rcosxbracketrightbig1/2. (14.10) Lettingr=1 andusingAbel’stheorem,weseethat ∞summationdisplay n=1cosnx n=−ln(2−2cosx)1/2 =−lnparenleftbigg 2sinx 2parenrightbigg ,x∈(0,2π).2(14.11) Bothsidesof thisexpressiondivergeas x→0 and 2π. /squaresolid Completeness The problem of establishing completeness may be approached in a number of different ways. One way is to transform the trigonometric Fourier series into exponential form and tocompareitwithaLaurentseries.Ifweexpand f(z)inaLaurentseries3(assuming f(z) isanalytic), f(z)=∞summationdisplay n=−∞dnzn. (14.12) Ontheunitcircle z=eiθand f(z)=f(eiθ)=∞summationdisplay n=−∞dneinθ. (14.13) 2Thelimits maybe shifted to [−π,π](andx/negationslash=0)using|x|on the right-hand side. 3Section6.5. 884 Chapter 14 Fourier Series FIGURE 14.1Fourier representationof sawtooth wave. TheLaurentexpansionontheunitcircle(Eq.(14.13))hasthesameformasthecomplex Fourier series (Eq. (14.12)), which shows the equivalence between the two expansions. SincetheLaurentseriesasapowerserieshasthepropertyofcompleteness,weseethatthe Fourier functions einxform a complete set. There is a significant limitation here. Laurent seriesandcomplexpowerseriescannothandlediscontinuitiessuchasasquarewaveorthe sawtoothwaveofFig.14.1, exceptonthecircleof convergence. Thetheoryofvectorspacesprovidesasecondapproachtothecompletenessofthesines andcosines.HerecompletenessisestablishedbytheWeierstrasstheoremfortwovariables. TheFourierexpansionandthecompletenesspropertymaybeexpected,forthefunctions sinnx,cosnx,einxare alleigenfunctionsofaself-adjointlinearODE, y′′+n2y=0. (14.14) Weobtainorthogonaleigenfunctionsfordifferentvaluesoftheeigenvalue nfortheinterval [0,2π]that satisfy the boundary conditions in the Sturm–Liouville theory (Chapter 10). Differenteigenfunctionsfor thesameeigenvalue nareorthogonal.We have integraldisplay2π 0sinmxsinnxdx=braceleftbiggπδmn,m/negationslash=0, 0,m=0,(14.15) integraldisplay2π 0cosmxcosnxdx=braceleftbiggπδmn,m/negationslash=0, 2π, m=n=0,(14.16) integraldisplay2π 0sinmxcosnxdx=0 forallintegral mandn. (14.17) Note that any interval x0≤x≤x0+2πwill be equally satisfactory. Frequently, we shall usex0=−πtoobtaintheinterval −π≤x≤π.Forthecomplexeigenfunctions e±inxor- thogonalityisusually definedintermsofthecomplexconjugateofoneofthetwofactors, integraldisplay2π 0parenleftbig eimxparenrightbig∗einxdx=2πδmn. (14.18) Thisagreeswiththetreatmentof thesphericalharmonics(Section12.6). 14.1 General Properties 885 Sturm–Liouville Theory The Sturm–Liouville theory guarantees the validity of Eq. (14.1) (for functions satis- fying the Dirichlet conditions) and, by use of the orthogonality relations, Eqs. (14.15), (14.16), and (14.17), allows us to compute the expansion coefficients an,bn,a ss h o w ni n Eqs. (14.2), and (14.3). Substituting Eqs. (14.2) and (14.3) into Eq. (14.1), we write our Fourierexpansionas f(x)=1 2πintegraldisplay2π 0f(t)dt +1 π∞summationdisplay n=1parenleftbigg cosnxintegraldisplay2π 0f(t)cosntdt+sinnxintegraldisplay2π 0f(t)sinntdtparenrightbigg =1 2πintegraldisplay2π 0f(t)dt+1 π∞summationdisplay n=1integraldisplay2π 0f(t)cosn(t−x)dt, (14.19) the first (constant) term being the average value of f(x)over the interval [0,2π]. Equa- tion (14.19) offers one approach to the development of the Fourier integral and Fourier transforms, Section15.1. Another way of describing what we are doing here is to say that f(x)is part of an infinite-dimensional Hilbert space, with the orthogonal cos nxand sinnxas the basis. (They can always be renormalized to unity if desired.) The statement that cos nxand sinnx (n=0,1,2,...)span this Hilbert space is equivalent to saying that they form a completeset.Finally,theexpansioncoefficients anandbncorrespondtotheprojectionsof f(x), with the integral inner products (Eqs. (14.2) and (14.3)) playing the role of the dot productofSection1.3. Thesepointsare outlinedinSection10.4. Example 14.1.2 SAWTOOTH WAVE An idea of the convergence of a Fourier series and the error in using only a finite number oftermsintheseries maybeobtainedbyconsideringtheexpansionof f(x)=braceleftbiggx, 0≤x<π, x−2π, π <x≤2π.(14.20) This is a sawtooth wave, and for convenience we shall shift our interval from [0,2π]to [−π,π]. In this interval we have f(x)=x. Using Eqs. (14.2) and (14.3), we show the expansiontobe f(x)=x=2bracketleftbigg sinx−sin2x 2+sin3x 3−···+(−1)n+1sinnx n+···bracketrightbigg .(14.21) Figure14.1shows f(x)for0≤x<πforthesumof4,6,and10termsoftheseries.Three featuresdeservecomment. 1. Thereisasteadyincreaseintheaccuracyoftherepresentationasthenumberofterms includedisincreased. 886 Chapter 14 Fourier Series 2. Allthecurvespass throughthemidpoint, f(x)=0,atx=π. 3. Inthevicinityof x=πthereisanovershootthatpersistsandshowsnosignofdimin- ishing. As a matter of incidental interest, setting x=π/2 in Eq. (14.21) provides an alternate derivationofLeibniz’formula,Exercise5.7.6. /squaresolid Behavior of Discontinuities The behavior of the sawtooth wave f(x)atx=πis an example of a general rule that at a finite discontinuity the series converges to the arithmetic mean. For a discontinuity at x=x0theseries yields f(x0)=1 2bracketleftbig f(x0+0)+f(x0−0)bracketrightbig , (14.22) the arithmetic mean of the right and left approaches to x=x0. A general proof using partial sums, as in Section 14.5, is given by Jeffreys and Jeffreys and by Carslaw (see the Additional Readings). The proof may be simplified by the use of Dirac delta functions— Exercise14.5.1. The overshoot of the sawtooth wave just before x=πin Fig. 14.1 is an example of the Gibbsphenomenon,discussedinSection14.5. Exercises 14.1.1 Afunction f(x)(quadraticallyintegrable)istoberepresentedbya finiteFourierseries. A convenient measure of the accuracy of the series is given by the integrated square of thedeviation, /Delta1p=integraldisplay2π 0bracketleftbigg f(x)−a0 2−psummationdisplay n=1(ancosnx+bnsinnx)bracketrightbigg2 dx. Showthattherequirementthat /Delta1pbeminimized,thatis, ∂/Delta1p ∂an=0,∂/Delta1p ∂bn=0, for alln,leadstochoosing anandbnas giveninEqs. (14.2)and(14.3). Note. Your coefficients anandbnare independent of p. This independence is a con- sequence of orthogonality and would not hold for powers of x, fitting a curve with polynomials. 14.1.2 Intheanalysisofacomplexwaveform(oceantides,earthquakes,musicaltones,etc.)it mightbemoreconvenienttohavetheFourierseries writtenas f(x)=a0 2+∞summationdisplay n=1αncos(nx−θn). 14.1 General Properties 887 ShowthatthisisequivalenttoEq. (14.1)with an=αncosθn,α2 n=a2 n+b2 n, bn=αnsinθn,tanθn=bn/an. Note. The coefficients α2 nas a function of ndefine what is called the power spectrum . Theimportanceof α2 nlies intheirinvarianceunderashiftinthephase θn. 14.1.3 Afunction f(x)is expandedinanexponentialFourierseries f(x)=∞summationdisplay n=−∞cneinx. Iff(x)isreal,f(x)=f∗(x), whatrestrictionis imposedonthecoefficients cn? 14.1.4 Assumingthatintegraltextπ −π[f(x)]2dxis finite,showthat limm→∞am=0,limm→∞bm=0. Hint.I n t e g r a t e[f(x)−sn(x)]2, wheresn(x)is thenth partial sum, and use Bessel’s inequality, Section 10.4. For our finite interval the assumption that f(x)is square inte- grable (integraltextπ −π|f(x)|2dxis finite) implies thatintegraltextπ −π|f(x)|dxis also finite. The converse doesnothold. 14.1.5 Applythesummationtechniqueof thissectiontoshowthat ∞summationdisplay n=1sinnx n=braceleftBigg1 2(π−x), 0<x≤π −1 2(π+x),−π≤x<0 (Fig.14.2). FIGURE 14.2Reversesawtoothwave. 888 Chapter 14 Fourier Series 14.1.6 Sumthetrigonometricseries ∞summationdisplay n=1(−1)n+1sinnx n andshowthatitequals x/2. 14.1.7 Sumthetrigonometricseries ∞summationdisplay n=0sin(2n+1)x 2n+1 andshowthatitequals braceleftbiggπ/4,0<x<π −π/4,−π<x<0. 14.1.8 Calculate the sum of the finite Fourier sine series for the sawtooth wave, f(x)= x,(−π,π), Eq. (14.21). Use 4-, 6-, 8-, and 10-term series and x/π=0.00(0.02)1.00. If aplottingroutineis available,plotyourresultsandcomparewithFig. 14.1. 14.1.9 Letf(z)=ln(1+z)=summationtext∞ n=1(−1)n+1zn/n. (This series converges to ln (1+z)for |z|≤1,exceptatthepoint z=−1.) (a) Fromtherealparts showthat lnparenleftbigg 2cosθ 2parenrightbigg =∞summationdisplay n=1(−1)n+1cosnθ n,−π<θ<π . (b) Usingachangeof variable,transform part(a)into −lnparenleftbigg 2sinθ 2parenrightbigg =∞summationdisplay n=1cosnθ n,0<θ<2π. 14.2 A DVANTAGES ,USES OF FOURIER SERIES Discontinuous Functions OneoftheadvantagesofaFourierrepresentationoversomeotherrepresentation,suchasa Taylorseries,isthatitcanrepresentadiscontinuousfunction.Anexampleisthesawtooth wave in the preceding section. Other examples are considered in Section 14.3 and in the exercises. Periodic Functions Related to this advantage is the usefulness of a Fourier series in representing a periodic function.If f(x)hasaperiodof 2 π,perhapsitisonlynaturalthatweexpanditinaseries of functions with period 2 π,2π/2,2π/3,.... This guarantees that if our periodic f(x)is representedoveroneinterval [0,2π]or[−π,π], therepresentationholdsforallfinite x. 14.2 Advantages, Uses of Fourier Series 889 At this point we may conveniently consider the properties of symmetry. Using the in- terval[−π,π],sinxis odd and cos xis an even function of x. Hence, by Eqs. (14.2) and (14.3),4iff(x)is odd,all an=0 andiff(x)iseven,all bn=0.In otherwords, f(x)=a0 2+∞summationdisplay n=1ancosnx, f(x) even, (14.23) f(x)=∞summationdisplay n=1bnsinnx, f(x) odd, (14.24) Frequentlythesepropertiesarehelpfulinexpandingagivenfunction. We have noted that the Fourier series is periodic. This is important in considering whetherEq.(14.1) holdsoutsidetheinitialinterval.Supposewearegivenonlythat f(x)=x,0≤x<π (14.25) and are asked to represent f(x)by a series expansion. Let us take three of the infinite numberofpossibleexpansions. 1. If weassumeaTaylorexpansion,wehave f(x)=x, (14.26) aone-termseries. This(one-term)series isdefinedfor allfinite x. 2. Using the Fourier cosine series (Eq. (14.23)), thereby assuming the function is repre- sented faithfully in the interval [0,π)and extended to neighboring intervals using the knownsymmetryproperties,wepredictthat f(x)=−x,−π<x≤0, f(x)=2π−x, π <x< 2π.(14.27) 3. Finally,fromtheFouriersineseries (Eq. (14.24)), wehave f(x)=x,−π<x≤0, f(x)=x−2π, π <x< 2π.(14.28) Thesethreepossibilities—Taylorseries,Fouriercosineseries,andFouriersineseries— are each perfectly valid in the original interval, [0,π]. Outside, however, their behavior is strikinglydifferent(compareFig.14.3).Whichofthethree,then,iscorrect?Thisquestion has no answer, unless we are given more information about f(x). It may be any of the three or none of them. Our Fourier expansionsare valid overthe basic interval. Unless the functionf(x)isknowntobeperiodicwithaperiodequaltoourbasicintervalorto (1/n)th ofourbasicinterval,thereisnoassurancewhateverthattherepresentation(Eq.(14.1))will haveanymeaningoutsidethebasicinterval. Inadditiontotheadvantagesofrepresentingdiscontinuousandperiodicfunctions,there is a third very real advantage in using a Fourier series. Suppose that we are solving the equationofmotionofanoscillatingparticlesubjecttoaperiodicdrivingforce.TheFourier 4With the range of integration −π≤x≤π. 890 Chapter 14 Fourier Series FIGURE 14.3ComparisonofFouriercosineseries, Fouriersineseries, andTaylorseries. expansion of the driving force then gives us the fundamental term and a series of harmon- ics. The (linear) ODE may be solved for each of these harmonics individually, a process that may be much easier than dealing with the original driving force. Then, as long as the ODEislinear,allthesolutionsmaybeaddedtogethertoobtainthefinalsolution.5This is morethanjustaclevermathematicaltrick. •It corresponds to finding the response of the system to the fundamental frequency and toeachof theharmonicfrequencies. Onequestionthatissometimesraisedis:“Weretheharmonicsthereallalong,orwerethey createdbyourFourieranalysis?”Oneanswercomparesthefunctionalresolutionintohar- monics with the resolution of a vector into rectangular components. The components may have been present, in the sense that they may be isolated and observed, but the resolution is certainly not unique. Hence many authors prefer to say that the harmonics were created by our choice of expansion. Other expansions in other sets of orthogonal functions would give different results. For further discussion we refer to a series of notes and letters in the AmericanJournalofPhysics .6 Change of Interval So far attention has been restricted to an interval of length 2 π. This restriction may easily berelaxed.If f(x)isperiodicwithaperiod 2 L,wemaywrite 5Oneof the nastier features of nonlinear differential equations is that this principle of superposition is not valid. 6B.L.Robinson,Concerningfrequenciesresultingfromdistortion. Am.J.Phys. 21:391(1953);F.W.VanName,Jr.,Concerning frequencies resulting from distortion. ibid.22: 94 (1954). 14.2 Advantages, Uses of Fourier Series 891 f(x)=a0 2+∞summationdisplay n=1bracketleftbigg ancosnπx L+bnsinnπx Lbracketrightbigg , (14.29) with an=1 LintegraldisplayL −Lf(t)cosnπt Ldt, n=0,1,2,3,..., (14.30) bn=1 LintegraldisplayL −Lf(t)sinnπt Ldt, n=1,2,3,..., (14.31) replacing xin Eq. (14.1) with πx/Landtin Eqs. (14.2) and (14.3) with πt/L.( F o r conveniencetheintervalinEqs.(14.2)and(14.3)isshiftedto −π≤t≤π.)Thechoiceof thesymmetricinterval (−L,L)isnotessential.For f(x)periodicwithaperiodof2 L,any interval(x0,x0+2L)will do. The choice is a matter of convenience or literally personal preference. Exercises 14.2.1 The boundary conditions (such as ψ(0)=ψ(l)=0) may suggest solutions of the form sin(nπx/l)andeliminatethecorrespondingcosines. (a) Verify that the boundary conditions used in the Sturm–Liouville theory are satis- fiedfor theinterval (0,l).Notethatthisis onlyhalftheusualFourierinterval. (b) Show that the set of functions ϕn(x)=sin(nπx/l), n=1,2,3,...,satisfies an orthogonalityrelation integraldisplayl 0ϕm(x)ϕn(x)dx=l 2δmn,n>0. 14.2.2 (a) Expand f(x)=xin the interval (0,2L). Sketch the series you have found (right- handsideofAns.) over (−2L,2L). ANS.x=L−2L π∞summationdisplay n=11 nsinparenleftbiggnπx Lparenrightbigg . (b) Expand f(x)=xas a sine series in the halfinterval(0,L). Sketch the series you havefound(right-handsideofAns.) over (−2L,2L). ANS.x=4L π∞summationdisplay n=01 2n+1sinparenleftbigg(2n+1)πx Lparenrightbigg . 14.2.3 In some problems it is convenient to approximate sin πxover the interval [0,1]by a parabola ax(1−x), whereais a constant. To get a feeling for the accuracy of this approximation,expand 4 x(1−x)inaFouriersineseries (−1≤x≤1): f(x)=braceleftbigg4x(1−x), 0≤x≤1 4x(1+x),−1≤x≤0bracerightbigg =∞summationdisplay n=1bnsinnπx. 892 Chapter 14 Fourier Series FIGURE 14.4Parabolicsinewave. ANS.bn=32 π3·1 n3,nodd bn=0, neven. (Fig.14.4.) 14.3 A PPLICATIONS OF FOURIER SERIES Example 14.3.1 SQUARE WAVE—H IGHFREQUENCIES OneapplicationofFourierseries,theanalysisofa“square”wave(Fig.14.5)intermsofits Fouriercomponents,occursinelectroniccircuitsdesignedtohandlesharplyrisingpulses. Supposethatour waveisdefinedby f(x)=0,−π<x<0, f(x)=h,0<x<π. (14.32) FromEqs. (14.2) and(14.3)wefind a0=1 πintegraldisplayπ 0hdt=h, (14.33) an=1 πintegraldisplayπ 0hcosntdt=0,n=1,2,3,..., (14.34) bn=1 πintegraldisplayπ 0hsinntdt=h nπ(1−cosnπ); (14.35) bn=2h nπ,nodd, (14.36) bn=0,neven. (14.37) Theresultingseries is f(x)=h 2+2h πparenleftbiggsinx 1+sin3x 3+sin5x 5+···parenrightbigg . (14.38) 14.3 Applications of Fourier Series 893 FIGURE 14.5Squarewave. Except for the first term, which represents an average of f(x)over the interval [−π,π], allthecosinetermshavevanished.Since f(x)−h/2isodd,wehaveaFouriersineseries. Althoughonlytheoddtermsinthesineseriesoccur,theyfallonlyas n−1.Thisconditional convergence is like that of the alternating harmonic series. Physically this means that our squarewavecontainsalotof high-frequencycomponents .Iftheelectronicapparatuswill not pass these components, our square-wave input will emerge more or less rounded off, perhapsas anamorphousblob. /squaresolid Example 14.3.2 FULL-WAVERECTIFIER As a second example, let us ask how well the output of a full-wave rectifier approaches puredirectcurrent(Fig.14.6).Ourrectifiermaybethoughtofashavingpassedthepositive peaksof anincomingsinewaveandinvertingthenegativepeaks.This yields f(t)=sinωt,0<ωt<π, (14.39) f(t)=−sinωt,−π<ω t< 0. Sincef(t)defined here is even, no terms of the form sin nωtwill appear. Again, from Eqs. (14.2)and(14.3), wehave a0=−1 πintegraldisplay0 −πsinωtd(ωt)+1 πintegraldisplayπ 0sinωtd(ωt) =2 πintegraldisplayπ 0sinωtd(ωt)=4 π, (14.40) an=2 πintegraldisplayπ 0sinωtcosnωtd(ωt) =−2 π2 n2−1,neven, =0,nodd. (14.41) 894 Chapter 14 Fourier Series FIGURE 14.6Full-waverectifier. Notethat[0,π]isnotanorthogonalityintervalforbothsinesandcosinestogetherandwe donotgetzerofor even n.Theresultingseriesis f(t)=2 π−4 π∞summationdisplay n=2,4,6,...cosnωt n2−1. (14.42) The original frequency, ω, has been eliminated. The lowest-frequency oscillation is 2 ω. The high-frequency components fall off as n−2, showing that the full-wave rectifier does a fairly good job of approximating direct current. Whether this good approximation is adequatedepends on the particular application.If the remainingac componentsare objec- tionable,theymaybefurthersuppressedbyappropriatefiltercircuits.Thesetwoexamples bringouttwofeaturescharacteristicof Fourierexpansions.7 •Iff(x)has discontinuities (as in the square wave in Example 14.3.1), we can expect thenthcoefficienttobedecreasingas O(1/n). Convergenceis conditionalonly. •Iff(x)is continuous (although possibly with discontinuous derivatives, as in the full- wave rectifier of Example 14.3.2), we can expect the nth coefficient to be decreasing as 1/n2, thatis, absoluteconvergence./squaresolid Example 14.3.3 INFINITE SERIES ,RIEMANN ZETAFUNCTION Asafinalexample,weconsidertheproblemof expanding x2.Let f(x)=x2,−π<x<π . (14.43) Sincef(x)iseven,all bn=0.Forthe anwehave a0=1 πintegraldisplayπ −πx2dx=2π2 3, (14.44) 7G.Raisbeek, Order of magnitude of Fourier coefficients. Am.Math. Mon. 62: 149–155 (1955). 14.3 Applications of Fourier Series 895 an=2 πintegraldisplayπ 0x2cosnxdx =2 π·(−1)n2π n2 =(−1)n4 n2. (14.45) Fromthis weobtain x2=π2 3+4∞summationdisplay n=1(−1)ncosnx n2. (14.46) Asitstands,Eq. (14.46)is ofnoparticularimportance.Butif weset x=π, cosnπ=(−1)n(14.47) andEq. (14.46)becomes8 π2=π2 3+4∞summationdisplay n=11 n2, (14.48) or π2 6=∞summationdisplay n=11 n2≡ζ(2), (14.49) thus yielding the Riemann zeta function, ζ(2), in closed form (in agreement with the BernoullinumberresultofSection5.9).Fromourexpansionof x2andexpansionsofother powersof x,numerousotherinfiniteseriescanbeevaluated.Afewareincludedinthislist ofexercises: Fourier series Reference 1.∞summationdisplay n=11 nsinnx=braceleftBigg −1 2(π+x),−π≤x<0 1 2(π−x), 0≤x<πExercise14 .1.5 Exercise14 .3.3 2.∞summationdisplay n=1(−1)n+11 nsinnx=1 2x,−π<x<πExercise14 .1.6 Exercise14 .3.2 3.∞summationdisplay n=01 2n+1sin(2n+1)x=braceleftbigg−π/4,−π<x<0 +π/4,0<x<πExercise14 .1.7 Eq.(14.38) 4.∞summationdisplay n=1cosnx n=−lnbracketleftbigg 2sinparenleftbigg|x| 2parenrightbiggbracketrightbigg ,−π<x<πEq.(14.11) Exercise14.1.9(b) 5.∞summationdisplay n=1(−1)n1 ncosnx=−lnbracketleftbigg 2cosparenleftbiggx 2parenrightbiggbracketrightbigg ,−π<x<π Exercise 14.1.9(a) 6.∞summationdisplay n=01 2n+1cos(2n+1)x=1 2lnbracketleftbigg cot|x| 2bracketrightbigg ,−π<x<π 8Notethat the point x=πis not a point of discontinuity. 896 Chapter 14 Fourier Series Thesquare-waveFourierseries fromEq. (14.38)anditem(3) inthetable, g(x)=∞summationdisplay n=0sin(2n+1)x 2n+1=(−1)mπ 4,mπ<x<(m+1)π, (14.50) can be used to derive Riemann’s functional equation for the zeta function . Its defining Dirichletseriescanbewritteninvariousforms: ζ(s)=∞summationdisplay n=1n−s=1+∞summationdisplay n=1(2n)−s+∞summationdisplay n=1(2n+1)−s =2−sζ(s)+∞summationdisplay n=0(2n+1)−s implyingthatthefunction λ(s)definedinSection5.9(alongwith η(s)) satisfies λ(s)≡∞summationdisplay n=0(2n+1)−s=parenleftbig 1−2−sparenrightbig ζ(s). (14.51) Heresisacomplexvariable.BothDirichletseriesconvergefor σ=ℜs>1.Alternatively, usingEq. (14.51), wehave η(s)≡∞summationdisplay n=1(−1)n−1n−s=∞summationdisplay n=0(2n+1)−s−∞summationdisplay n=1(2n)−s=parenleftbig 1−21−sparenrightbig ζ(s), (14.52) which converges already for ℜs>0 using the Leibniz convergence criterion (see Sec- tion5.3). AnotherapproachtoDirichletseriesstartsfromEuler’sintegralforthegammafunction, integraldisplay∞ 0ys−1e−nydy=n−sintegraldisplay∞ 0e−yys−1dy=n−sŴ(s), (14.53) whichmaybesummedusingthegeometricseries ∞summationdisplay n=1e−ny=e−y 1−e−y=1 ey−1 toyieldtheintegralrepresentationfor thezetafunction: integraldisplay∞ 0ys−1 ey−1dy=ζ(s)Ŵ(s). (14.54) If wecombinethealternativeforms ofEq. (14.53), integraldisplay∞ 0ys−1e−inydy=n−sŴ(s)e−iπs/2, integraldisplay∞ 0ys−1einydy=n−sŴ(s)eiπs/2, 14.3 Applications of Fourier Series 897 weobtain integraldisplay∞ 0ys−1sin(ny)dy=n−sŴ(s)sinπs 2. (14.55) DividingbothsidesofEq.(14.55)by nandsummingoverallodd nyields,for σ=ℜ(s)> 0, integraldisplay∞ 0g(y)ys−1dy=parenleftbig 1−2−s−1parenrightbig ζ(s+1)Ŵ(s)sinπs 2, (14.56) usingEqs. (14.50) and (14.51). Here, the interchangeof summationandintegrationis jus- tifiedbyuniformconvergence.Thisrelationisattheheartofthefunctionalequation.Ifwe divide the integration range into intervals mπ<y<(m +1)πand substitute Eq. (14.50) intoEq. (14.56)wefind integraldisplay∞ 0g(y)ys−1dy=π 4∞summationdisplay m=0(−1)mintegraldisplay(m+1)π mπys−1dy =πs+1 4sbraceleftbigg∞summationdisplay m=1(−1)mbracketleftbig (m+1)s−msbracketrightbig +1bracerightbigg =πs+1 2sparenleftbig 1−2s+1parenrightbig ζ(−s), (14.57) using Eq. (14.52). The series in Eq. (14.57) converges for ℜs<1 to an analytic func- tion. Comparing Eqs. (14.56) and (14.57) for the common area of convergence to analytic functions, 0 <σ=ℜs<1,wegetthe functionalequation πs+1 2sparenleftbig 1−2s+1parenrightbig ζ(−s)=parenleftbig 1−2−s−1parenrightbig ζ(s+1)Ŵ(s)sinπs 2, whichcanberewrittenas ζ(1−s)=2(2π)−sζ(s)Ŵ(s)cosπs 2. (14.58) This functional equation provides an analytic continuation of ζ(s)into the negative half- plane ofs.F o rs→1 the pole of ζ(s)and the zero of cos (πs/2)cancel in Eq. (14.58), so ζ(0)=−1/2results.Sincecos (πs/2)=0fors=2m+1=oddinteger,Eq.(14.58)gives ζ(−2m)=0,thetrivialzerosofthezetafunctionfor m=1,2,....Allotherzerosmustlie in the“criticalstrip” 0 <σ=ℜs<1. Theyare closely relatedto thedistributionof prime numbers because the prime number product for ζ(s)(see Section 5.9) can be converted into a Dirichlet series over prime powers for ζ′/ζ=dlnζ(s)/ds.F r o mh e r eo nw es k e t c h ideasonly,withoutproofs.UsingtheinverseMellintransform(seeSection16.2)yieldsthe relation summationdisplay pm<x,p=prime m=1,2,...lnp=−1 2πiintegraldisplayσ+i∞ σ−i∞ζ′(s) ζ(s)sxsds (14.59) forσ>1, which is a cornerstone of analytic number theory. Since zeros of ζ(s)become simple poles of ζ′/ζ, the asymptotic distribution of prime numbers is directly related by 898 Chapter 14 Fourier Series Eq. (14.59) to the zeros of the Riemann zeta function. Riemann conjectured that all zeros lieontheline σ=1/2,thatis,havetheform1 /2+itwithreal t.Ifso,onecouldshiftthe lineofintegrationtotheleftto σ=1/2+ε,thesimplepoleof ζ(s)ats=1 givingriseto the residue x, while the integral along the line σ=1/2+εis of order O(x1/2+ε). Hence, theremarkablysmallremainderintheasymptoticestimate summationdisplay p<xlnp∼x+Oparenleftbig x1/2+εparenrightbig ,x→∞ would result for arbitrarily small ε. This is equivalent to the estimate for the number of primesbelow x, π(x)=summationdisplay p<x1=integraldisplayx 2(lnt)−1dt+Oparenleftbig x1/2+εparenrightbig ,x→∞. In fact, numerical studies have shown that the first 300 ×109zeros are simple and lie all on the critical line σ=1/2. For more details the reader is referred to the classic text by E. C. Titchmarsh and D. R. Heath-Brown, The Theory of the Riemann Zeta Function , Oxford, UK: Clarendon Press (1986); H. M. Edwards, Riemann’s Zeta Function ,N e w York: Academic Press (1974) and Dover (2003); J. Van de Lune, H. J. J. Te Riele, and D. T. Winter, On the zeros of the Riemann zeta function in the critical strip. IV. Math. Comput.47:667(1986).PopularaccountscanbefoundinM.duSautoy, TheMusicofthe Primes:SearchingtoSolvetheGreatestMysteryinMathematics ,NewYork:HarperCollins (2003); J. Derbyshire, Prime Obsession: Bernhard Riemann and the Greatest Unsolved Problem in Mathematics . Washington, DC: Joseph Henry Press (2003); K. Sabbagh, The Riemann Hypothesis: The Greatest Unsolved Problem in Mathematics , New York: Farrar, StrausandGiroux(2003). More recently the statistics of the zeros ρof the Riemann zeta function on the critical line played a prominent role in the development of theories of chaos (see Chapter 18 for an introduction). Assuming that there is a quantum mechanical system whose energies are the imaginary parts of the ρ, then primes determine the primitive periodic orbits of the associated classically chaotic system. For this case Gutzwiller’s trace formula, which relates quantum energy levels and classical periodic orbits, plays a central role and can be better understood using properties of the zeta function and primes. For more details see Sections12.6and12.7byJ.Keating,in TheNatureofChaos (T.Mullin,ed.),Oxford,UK: ClarendonPress (1993),andreferencestherein. /squaresolid Exercises 14.3.1 DeveloptheFourierseriesrepresentationof f(t)=braceleftbigg0,−π≤ωt≤0, sinωt,0≤ωt≤π. Thisistheoutputofasimplehalf-waverectifier.Itisalsoanapproximationofthesolar thermaleffectthatproduces“tides”intheatmosphere. ANS.f(t)=1 π+1 2sinωt−2 π∞summationdisplay n=2,4,6,...cosnωt n2−1. 14.3 Applications of Fourier Series 899 FIGURE 14.7Triangularwave. 14.3.2 Asawtoothwaveisgivenby f(x)=x,−π<x<π . Showthat f(x)=2∞summationdisplay n=1(−1)n+1 nsinnx. 14.3.3 Adifferentsawtoothwaveis describedby f(x)=braceleftBigg −1 2(π+x),−π≤x<0 +1 2(π−x),0<x≤π. Showthat f(x)=summationtext∞ n=1(sinnx/n). 14.3.4 Atriangularwave(Fig. 14.7)is representedby f(x)=braceleftBiggx,0<x<π −x,−π<x<0. Represent f(x)bya Fourierseries. ANS.f(x)=π 2−4 πsummationdisplay n=1,3,5,...cosnx n2. 14.3.5 Expand f(x)=braceleftBigg1,x2<x2 0 0,x2>x2 0 intheinterval[−π,π]. Note.This variable-widthsquarewaveis ofsomeimportanceinelectronicmusic. 900 Chapter 14 Fourier Series FIGURE 14.8Cross section ofsplittube. 14.3.6 Ametalcylindricaltubeofradius aissplitlengthwiseintotwonontouchinghalves.The top half is maintained at a potential +V, the bottom half at a potential −V(Fig. 14.8). Separate the variables in Laplace’s equation and solve for the electrostatic potential for r≤a. Observe the resemblance between your solution for r=aand the Fourier series for asquarewave. 14.3.7 A metal cylinder is placed in a (previously) uniform electric field, E0, with the axis of thecylinderperpendiculartothatof theoriginalfield. (a) Findtheperturbedelectrostaticpotential. (b) Findtheinducedsurfacechargeonthecylinderasa functionofangularposition. 14.3.8 Transform the Fourier expansion of a square wave, Eq. (14.38), into a power series. Show that the coefficients of x1form adivergent series. Repeat for the coefficients ofx3. A power series cannot handle a discontinuity. These infinite coefficients are the result ofattemptingtobeatthisbasiclimitationonpowerseries. 14.3.9 (a) ShowthattheFourierexpansionof cos axis cosax=2asinaπ πbraceleftbigg1 2a2−cosx a2−12+cos2x a2−22−···bracerightbigg , an=(−1)n2asinaπ π(a2−n2). (b) Fromtheprecedingresultshowthat aπcotaπ=1−2∞summationdisplay p=1ζ(2p)a2p. ThisprovidesanalternatederivationoftherelationbetweentheRiemannzetafunction andtheBernoullinumbers,Eq. (5.152). 14.3 Applications of Fourier Series 901 14.3.10 DerivetheFourierseriesexpansionoftheDiracdeltafunction δ(x)intheinterval−π< x<π. (a) Whatsignificancecanbeattachedtotheconstantterm? (b) Inwhatregionisthisrepresentationvalid? (c) Withtheidentity Nsummationdisplay n=1cosnx=sin(Nx/2) sin(x/2)cosbracketleftbiggparenleftbigg N+1 2parenrightbiggx 2bracketrightbigg , showthatyourFourierrepresentationof δ(x)is consistentwithEq. (1.190). 14.3.11 Expandδ(x−t)in a Fourier series. Compare your result with the bilinear form of Eq. (1.190). ANS.δ( x−t)=1 2π+1 π∞summationdisplay n=1(cosnxcosnt+sinnxsinnt) =1 2π+1 π∞summationdisplay n=1cosn(x−t). 14.3.12 Verifythat δ(ϕ1−ϕ2)=1 2π∞summationdisplay m=−∞eim(ϕ1−ϕ2) isaDiracdeltafunctionbyshowingthatitsatisfiesthedefinitionofaDiracdeltafunc- tion: integraldisplayπ −πf(ϕ1)1 2π∞summationdisplay m=−∞eim(ϕ1−ϕ2)dϕ1=f(ϕ2). Hint.Represent f(ϕ1)byanexponentialFourierseries. Note. The continuum analog of this expression is developed in Section 15.2. The most important application of this expression is in the determination of Green’s functions, Section9.7. 14.3.13 (a) Using f(x)=x2,−π<x<π , showthat ∞summationdisplay n=1(−1)n+1 n2=π2 12=η(2). (b) Using the Fourier series for a triangular wave developed in Exercise 14.3.4, show that ∞summationdisplay n=11 (2n−1)2=π2 8=λ(2). 902 Chapter 14 Fourier Series (c) Using f(x)=x4,−π<x<π , showthat ∞summationdisplay n=11 n4=π4 90=ζ(4),∞summationdisplay n=1(−1)n+1 n4=7π4 720=η(4). (d) Using f(x)=braceleftbiggx(π−x),0<x<π, x(π+x), π <x< 0, derive f(x)=8 π∞summationdisplay n=1,3,5,...sinnx n3 andshowthat ∞summationdisplay n=1,3,5,...(−1)(n−1)/21 n3=1−1 33+1 53−1 73+···=π3 32=β(3). (e) UsingtheFourierseriesfor asquarewave,showthat ∞summationdisplay n=1,3,5,...(−1)(n−1)/21 n=1−1 3+1 5−1 7+···=π 4=β(1). ThisisLeibniz’formulafor π,obtainedbyadifferenttechniqueinExercise5.7.6. Note.T h eη(2),η(4),λ(2),β(1), andβ(3)functions are defined by the indicated series. GeneraldefinitionsappearinSection5.9. 14.3.14 (a) FindtheFourierseries representationof f(x)=braceleftbigg0,−π<x≤0 x,0≤x<π. (b) FromtheFourierexpansionshowthat π2 8=1+1 32+1 52+···. 14.3.15 Asymmetrictriangularpulseofadjustableheightandwidthisdescribedby f(x)=braceleftbigga(1−x/b), 0≤|x|≤b 0,b ≤|x|≤π. (a) ShowthattheFouriercoefficientsare a0=ab π,a n=2ab π(nb)2(1−cosnb). Sum the finite Fourier series through n=10 and through n=100 forx/π= 0(1/9)1.Takea=1 andb=π/2. 14.4 Properties of Fourier Series 903 (b) CallaFourieranalysissubroutine(ifavailable)tocalculatetheFouriercoefficients off(x),a 0througha10. 14.3.16 (a) Using a Fourier analysis subroutine, calculate the Fourier cosine coefficients a0 througha10of f(x)=bracketleftbigg 1−parenleftbiggx πparenrightbigg2bracketrightbigg1/2 ,x∈[−π,π]. (b) Spot-check by calculating some of the preceding coefficients by direct numerical quadrature. Checkvalues. a0=0.785,a2=0.284. 14.3.17 Using a Fourier analysis subroutine, calculate the Fourier coefficients through a10and b10for (a) afull-waverectifier,Example14.3.2, (b) ahalf-waverectifier,Exercise14.3.1.Checkyourresultsagainsttheanalyticforms given(Eq. (14.41)andExercise14.3.1). 14.4 P ROPERTIES OF FOURIER SERIES Convergence Itmightbenoted,first,thatourFourierseriesshouldnotbeexpectedtobeuniformlycon- vergentifitrepresentsadiscontinuousfunction.Auniformlyconvergentseriesofcontinu- ous functions (sinnx,cosnx)always yields a continuous function (compare Section 5.5). If, however, (a)f(x)iscontinuous,−π≤x≤π, (b)f(−π)=f(+π), and (c)f′(x)issectionallycontinuous, theFourierseries for f(x)willconvergeuniformly.These restrictions donotdemandthat f(x)beperiodic,buttheywillbesatisfiedbycontinuous,differentiable,periodicfunctions (period of 2 π). For a proof of uniform convergence we refer to the literature.9With or without a discontinuity in f(x), the Fourier series will yield convergence in the mean, Section10.4. 9See, for instance, R. V. Churchill, Fourier Series and Boundary Value Problems , 5th ed., New York: McGraw-Hill (1993), Section 38. 904 Chapter 14 Fourier Series Integration Term-by-termintegrationoftheseries f(x)=a0 2+∞summationdisplay n=1ancosnx+∞summationdisplay n=1bnsinnx (14.60) yields integraldisplayx x0f(x)dx=a0x 2vextendsinglevextendsinglevextendsinglevextendsinglex x0+∞summationdisplay n=1an nsinnxvextendsinglevextendsinglevextendsinglex x0−∞summationdisplay n=1bn ncosnxvextendsinglevextendsinglevextendsinglex x0. (14.61) Clearly, the effect of integration is to place an additional power of nin the denominator of each coefficient. This results in more rapid convergence than before. Consequently, a convergent Fourier series may always be integrated term by term, the resulting series con- verginguniformlytotheintegraloftheoriginalfunction.Indeed,term-by-termintegration may be valid even if the original series (Eq. (14.60)) is not itself convergent. The func- tionf(x)need only be integrable. A discussion will be found in Jeffreys and Jeffreys, Section14.06(see theAdditionalReadings). Strictlyspeaking,Eq.(14.61)maynotbeaFourierseries;thatis,if a0/negationslash=0,therewillbe ater m1 2a0x. However, integraldisplayx x0f(x)dx−1 2a0x (14.62) willstillbeaFourierseries. Differentiation The situation regarding differentiation is quite different from that of integration. Here the wordiscaution.Considertheseries for f(x)=x,−π<x<π . (14.63) Wereadilyfind(compareExercise14.3.2)thattheFourierseries is x=2∞summationdisplay n=1(−1)n+1sinnx n,−π<x<π . (14.64) Differentiatingtermbyterm, weobtain 1=2∞summationdisplay n=1(−1)n+1cosnx, (14.65) whichisnotconvergent. Warning :Checkyourderivativeforconvergence. For a triangular wave (Exercise 14.3.4), in which the convergence is more rapid (and uniform), f(x)=π 2−4 π∞summationdisplay n=1,oddcosnx n2. (14.66) 14.4 Properties of Fourier Series 905 Differentiatingtermbytermweget f′(x)=4 π∞summationdisplay n=1,oddsinnx n, (14.67) whichistheFourierexpansionofasquarewave, f′(x)=braceleftbigg1,0<x<π, −1,−π<x<0.(14.68) InspectionofFig.14.7verifiesthatthisis indeedthederivativeofourtriangularwave. •As the inverse of integration, the operation of differentiation has placed an additional factornin the numerator of each term. This reduces the rate of convergence and may, asinthefirst casementioned,renderthedifferentiatedseriesdivergent. •Ingeneral,term-by-termdifferentiationispermissibleunderthesameconditionslisted for uniformconvergence. Exercises 14.4.1 ShowthatintegrationoftheFourierexpansionof f(x)=x,−π<x<π ,leadsto π2 12=∞summationdisplay n=1(−1)n+1 n2=1−1 4+1 9−1 16+···. 14.4.2 Parseval’sidentity. (a) AssumingthattheFourierexpansionof f(x)isuniformlyconvergent,showthat 1 πintegraldisplayπ −πbracketleftbig f(x)bracketrightbig2dx=a2 0 2+∞summationdisplay n=1parenleftbig a2 n+b2 nparenrightbig . ThisisParseval’sidentity.Itisactuallyaspecialcaseofthecompletenessrelation, Eq.(10.73). (b) Given x2=π2 3+4∞summationdisplay n=1(−1)ncosnx n2,−π≤x≤π, applyParseval’sidentitytoobtain ζ(4)inclosedform. (c) Theconditionofuniformconvergenceisnotnecessary.Showthisbyapplyingthe Parsevalidentitytothesquarewave f(x)=braceleftbigg−1,−π<x<0 1,0<x<π =4 π∞summationdisplay n=1sin(2n−1)x 2n−1. 906 Chapter 14 Fourier Series FIGURE 14.9Rectangularpulse. 14.4.3 Show that integrating the Fourier expansion of the Dirac delta function (Exer- cise 14.3.10) leads to the Fourier representation of the square wave, Eq. (14.38), with h=1. Note. Integrating the constant term (1/2π)leads to a term x/2π. What are you going todowiththis? 14.4.4 IntegratetheFourierexpansionoftheunitstepfunction f(x)=braceleftbigg0,−π<x<0 x,0<x<π. Showthatyourintegratedseries agreeswithExercise14.3.14. 14.4.5 Intheinterval (−π,π), δn(x)=braceleftBiggn,for|x|<1 2n, 0,for|x|>1 2n (Fig.14.9). (a) Expand δn(x)as aFouriercosineseries. (b) Show that your Fourier series agrees with a Fourier expansion of δ(x)in the limit asn→∞. 14.4.6 Confirm the delta function nature of your Fourier series of Exercise 14.4.4 by showing thatfor any f(x)thatis finiteintheinterval [−π,π]andcontinuousat x=0, integraldisplayπ −πf(x)bracketleftbig Fourierexpansionof δ∞(x)bracketrightbig dx=f(0). 14.4.7 (a) Show that the Dirac delta function δ(x−a), expanded in a Fourier sine series in thehalf-interval (0,L)(0<a<L) , is givenby δ(x−a)=2 L∞summationdisplay n=1sinparenleftbiggnπa Lparenrightbigg sinparenleftbiggnπx Lparenrightbigg . Notethatthisseriesactuallydescribes −δ(x+a)+δ(x−a)intheinterval (−L,L). 14.4 Properties of Fourier Series 907 (b) By integrating both sides of the preceding equation from 0 to x, show that the cosineexpansionofthesquarewave f(x)=braceleftbigg0,0≤x<a 1, a<x<L, is f(x)=2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg −2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg cosparenleftbiggnπx Lparenrightbigg , for 0≤x<L. (c) Verifythattheterm 2 π∞summationdisplay n=11 nsinparenleftbiggnπa Lparenrightbigg is/angbracketleftf(x)/angbracketright. 14.4.8 Verify the Fourier cosine expansion of the square wave, Exercise 14.4.7(b), by direct calculationoftheFouriercoefficients. 14.4.9 (a) A string is clamped at both ends x=0 andx=L. Assuming small-amplitude vibrations,wefindthattheamplitude y(x,t)satisfiesthewaveequation ∂2y ∂x2=1 v2∂2y ∂t2. Herevisthewavevelocity.Thestringissetinvibrationbyasharpblowat x=a. Hencewehave y(x,0)=0,∂y(x,t) ∂t=Lv0δ(x−a)att=0. The constant Lis included to compensate for the dimensions (inverse length) of δ(x−a). Withδ(x−a)given by Exercise 14.4.7(a), solve the wave equation subjecttotheseinitialconditions. ANS.y(x,t)=2v0L πv∞summationdisplay n=11 nsinnπa Lsinnπx Lsinnπvt L. (b) Showthatthetransversevelocityofthestring ∂y(x,t)/∂t is givenby ∂y(x,t) ∂t=2v0∞summationdisplay n=1sinnπa Lsinnπx Lcosnπvt L. 14.4.10 A string, clamped at x=0 and atx=1, is vibrating freely. Its motion is described by thewaveequation ∂2u(x,t) ∂t2=v2∂2u(x,t) ∂x2. 908 Chapter 14 Fourier Series AssumeaFourierexpansionoftheform u(x,t)=∞summationdisplay n=1bn(t)sinnπx l anddeterminethecoefficients bn(t). Theinitialconditionsare u(x,0)=f(x) and∂ ∂tu(x,0)=g(x). Note.ThisisonlyhalftheconventionalFourierorthogonalityintegralinterval.However, aslongasonlythesinesareincludedhere,theSturm–Liouvilleboundaryconditionsare stillsatisfiedandthefunctionsare orthogonal. ANS.bn(t)=Ancosnπvt l+Bnsinnπvt l, An=2 lintegraldisplayl 0f(x)sinnπx ldx,Bn=2 nπvintegraldisplayl 0g(x)sinnπx ldx. 14.4.11 (a) Let us continue the vibrating string problem, Exercise 14.4.10. The presence of a resistingmediumwilldampthevibrationsaccordingtotheequation ∂2u(x,t) ∂t2=v2∂2u(x,t) ∂x2−k∂u(x,t) ∂t. Assumea Fourierexpansion u(x,t)=∞summationdisplay n=1bn(t)sinnπx l and again determine the coefficients bn(t). Take the initial and boundary condi- tionstobethesameasinExercise14.4.10.Assumethedampingtobesmall. (b) Repeat,butassumethedampingtobelarge. ANS.(a)bn(t)=e−kt/2{Ancosωnt+Bnsinωnt}, An=2 lintegraldisplayl 0f(x)sinnπx ldx, Bn=2 ωnlintegraldisplayl 0g(x)sinnπx ldx+k 2ωnAn, ω2 n=parenleftbiggnπv lparenrightbigg −parenleftbiggk 2parenrightbigg2 >0. 14.4 Properties of Fourier Series 909 (b)bn(t)=e−kt/2{Ancoshσnt+Bnsinhσnt}, An=2 lintegraldisplayl 0f(x)sinnπx ldx, Bn=2 σnlintegraldisplayl 0g(x)sinnπx ldx+k 2σnAn, σ2 n=parenleftbiggk 2parenrightbigg2 −parenleftbiggnπv lparenrightbigg2 >0. 14.4.12 Find the charge distribution over the interior surfaces of the semicircles of Exer- cise14.3.6. Note. You obtain a divergent series and this Fourier approach fails. Using conformal mapping techniques, we may show the charge density to be proportional to csc θ. Does cscθhaveaFourierexpansion? 14.4.13 Given ϕ1(x)=∞summationdisplay n=1sinnx n=braceleftBigg−1 2(π+x),−π≤x<0, 1 2(π−x), 0<x≤π, showbyintegratingthat ϕ2(x)≡∞summationdisplay n=1cosnx n2=  1 4(π+x)2−π2 12,−π≤x≤0 1 4(π−x)2−π2 12,0≤x≤π. 14.4.14 Given ψ2s(x)=∞summationdisplay n=1sinnx n2s,ψ 2s+1(x)=∞summationdisplay n=1cosnx n2s+1, developthefollowingrecurrencerelations: (a)ψ2s(x)=integraldisplayx 0ψ2s−1(x)dx (b)ψ2s+1(x)=ζ(2s+1)−integraldisplayx 0ψ2s(x)dx. Note. The functions ψs(x)and theϕs(x)of the preceding two exercises are known as Clausenfunctions .Intheorytheymaybeusedtoimprovetherateofconvergenceofa Fourierseries.AswiththeseriesofChapter5,thereisalwaysthequestionofhowmuch analyticalworkwedoandhowmucharithmeticworkwedemandthatthecomputerdo. As computers become steadily more powerful, the balance progressively shifts so that wearedoingless anddemandingthattheydomore. 14.4.15 Showthat f(x)=∞summationdisplay n=1cosnx n+1 910 Chapter 14 Fourier Series maybewrittenas f(x)=ψ1(x)−ϕ2(x)+∞summationdisplay n=1cosnx n2(n+1). Note.ψ1(x)andϕ2(x)aredefinedintheprecedingexercises. 14.5 G IBBS PHENOMENON TheGibbsphenomenonisanovershoot,apeculiarityoftheFourierseriesandothereigen- functionseriesata simplediscontinuity.AnexampleisseeninFig.14.1. Summation of Series In Section 14.1 the sum of the first several terms of the Fourier series for a sawtooth wave was plotted (Fig. 14.10). Now we develop analytic methods of summing the first rterms ofourFourierseries. FromEq. (14.19), ancosnx+bnsinnx=1 πintegraldisplayπ −πf(t)cosn(t−x)dt. (14.69) Thenthe rthpartialsumbecomes10 sr(x)=a0 2+rsummationdisplay n=1(ancosnx+bnsinnx) =ℜ1 πintegraldisplayπ −πf(t)bracketleftbigg1 2+rsummationdisplay n=1e−i(t−x)nbracketrightbigg dt. (14.70) Summingthefiniteseries ofexponentials(geometricprogression),11weobtain sr(x)=1 2πintegraldisplayπ −πf(t)sin[(r+1 2)(t−x)] sin1 2(t−x)dt. (14.71) Thisis convergentatallpoints,including t=x. Thefactor sin[(r+1 2)(t−x)] 2πsin1 2(t−x) istheDirichletkernelmentionedinSection1.15asaDiracdeltadistribution. 10It is of some interest to note that this series also occurs in theanalysis of the diffraction grating ( rslits). 11Compare Exercise 6.1.7 with initial value n=1. 14.5 Gibbs Phenomenon 911 Square Wave For convenience of numerical calculation we consider the behavior of the Fourier series thatrepresentstheperiodicsquarewave f(x)=braceleftBiggh 2,0<x<π, −h 2,−π<x<0.(14.72) This is essentially the square wave used in Section 14.3, and we immediately see that the solutionis f(x)=2h πparenleftbiggsinx 1+sin3x 3+sin5x 5+···parenrightbigg . (14.73) ApplyingEq.(14.71)tooursquarewave(Eq.(14.72)),wehavethesumofthefirst rterms (plus1 2a0, whichiszerohere): sr(x)=h 4πintegraldisplayπ 0sin[(r+1 2)(t−x)] sin1 2(t−x)dt−h 4πintegraldisplay0 −πsin[(r+1 2)(t−x)] sin1 2(t−x)dt =h 4πintegraldisplayπ 0sin[(r+1 2)(t−x)] sin1 2(t−x)dt−h 4πintegraldisplayπ 0sin[(r+1 2)(t+x)] sin1 2(t+x)dt.(14.74) Thislastresultfollowsfrom thetransformation /vectort−t inthesecondintegral.Replacing t−xinthefirsttermwith sandt+xinthesecondterm withs,weobtain sr(x)=h 4πintegraldisplayπ−x −xsin(r+1 2)s sin1 2sds−h 4πintegraldisplayπ+x xsin(r+1 2)s sin1 2sds. (14.75) The intervals of integration are shown in Fig. 14.10(top). Because the integrands have the same mathematical form, the integrals from xtoπ−xcancel, leaving the integral rangesshowninthebottomportionof Fig.14.10: sr(x)=h 4πintegraldisplayx −xsin(r+1 2)s sin1 2sds−h 4πintegraldisplayπ+x π−xsin(r+1 2)s sin1 2sds. (14.76) Considerthepartialsuminthevicinityofthediscontinuityat x=0.Asx→0,thesec- ond integral becomes negligible, and we associate the first integral with the discontinuity atx=0.Usingr+1 2=pandps=ξweobtain sr(x)=h 2πintegraldisplaypx 0sinξ sin(ξ/2p)dξ p. (14.77) 912 Chapter 14 Fourier Series FIGURE 14.10Intervalsof integration—Eq.(14.75). Calculation of Overshoot Our partial sum sr(x)starts at zero when x=0 (in agreement with Eq. (14.22)) and in- creases until ξ=ps=π, at which point the numerator, sin ξ, goes negative. For large r, andthereforeforlarge p,ourdenominatorremainspositive.Wegetthemaximumvalueof the partial sum by taking the upper limit px=π. Right here we see that x, the locationof theovershootmaximum,isinverselyproportionaltothenumberofterms taken: x=π p≈π r. Themaximumvalueofthepartialsumis then sr(x)max=h 2·1 πintegraldisplayπ 0sinξdξ sin(ξ/2p)p≈h 2·2 πintegraldisplayπ 0sinξ ξdξ. (14.78) Intermsof thesineintegral, si (x)of Section8.5, integraldisplayπ 0sinξ ξdξ=π 2+si(π). (14.79) Theintegralisclearlygreaterthan π/2,sinceitcanbewrittenas parenleftbiggintegraldisplay∞ 0−integraldisplay3π π−integraldisplay5π 3π−···parenrightbiggsinξ ξdξ=integraldisplayπ 0sinξ ξdξ. (14.80) WesawinExample7.1.4thattheintegralfrom0to ∞isπ/2.Fromthisintegralweare subtracting a series of negative terms. A Gaussian quadrature or a power-series expansion andterm-by-termintegrationyields 2 πintegraldisplayπ 0sinξ ξdξ=1.1789797..., (14.81) whichmeansthattheFourierseriestendstoovershootthepositivecornerbysome18per- centandtoundershootthenegativecornerbythesameamount,assuggestedinFig.14.11. 14.5 Gibbs Phenomenon 913 FIGURE 14.11Squarewave—Gibbsphenomenon. The inclusion of more terms (increasing r) does nothing to remove this overshoot but merely moves it closer to the point of discontinuity. The overshoot is the Gibbs phenom- enon, and because of it the Fourier series representation may be highly unreliable for pre- cisenumericalwork, especiallyinthevicinityof adiscontinuity. The Gibbs phenomenon is not limited to the Fourier series. It occurs with other eigen- functionexpansions.Exercise12.3.27isanexampleoftheGibbsphenomenonforaLegen- dreseries.Formoredetails,seeW.J.Thompson,FourierseriesandtheGibbsphenomenon, Am.J. Phys. 60:425(1992). Exercises 14.5.1 With the partial sum summation techniques of this section, show that at a discontinuityinf(x)the Fourier series for f(x)takes on the arithmetic mean of the right- and left- handlimits: f(x0)=1 2bracketleftbig f(x0+0)+f(x0−0)bracketrightbig . Inevaluatinglim r→∞sr(x0)youmayfinditconvenienttoidentifypartoftheintegrand asaDiracdeltafunction. 14.5.2 Determinethepartialsum, sn, oftheseries inEq.(14.73) byusing (a)sinmx m=integraldisplayx 0cosmy dy,(b)nsummationdisplay p=1cos(2p−1)y=sin2ny 2siny. DoyouagreewiththeresultgiveninEq. (14.79)? 14.5.3 Evaluate the finite step function series, Eq. (14.73), h=2, using 100, 200, 300, 400, and 500 terms for x=0.0000(0.0005)0.0200. Sketch your results (five curves) or, if a plottingroutineis available,plotyourresults. 914 Chapter 14 Fourier Series 14.5.4 (a) CalculatethevalueoftheGibbsphenomenonintegral I=2 πintegraldisplayπ 0sint tdt bynumericalquadratureaccurateto12significantfigures. (b) Check your result by (1) expanding the integrand as a series, (2) integrating term by term, and (3) evaluating the integrated series. This calls for double precision calculation. ANS.I=1.178979744472 . 14.6 D ISCRETE FOURIER TRANSFORM For many physicists the Fourier transform is automatically the continuous Fourier trans- form of Chapter 15. The use of digital computers, however, necessarily replaces a contin- uumofvaluesbyadiscreteset;anintegrationisreplacedbyasummation.Thecontinuous FouriertransformbecomesthediscreteFouriertransformandanappropriatetopicforthis chapter. Orthogonality over Discrete Points The orthogonality of the trigonometric functions and the imaginary exponentials is ex- pressedinEqs.(14.15)to(14.18).Thisistheusualorthogonalityforfunctions: integration ofaproductoffunctionsovertheorthogonalityinterval.Thesines,cosines,andimaginary exponentials have the remarkable property that they are also orthogonal over a series of discrete,equallyspacedpointsovertheperiod(theorthogonalityinterval). Consideraset of 2 Ntimevalues tk=0,T 2N,2T 2N,...,(2N−1)T 2N(14.82) forthetimeinterval (0,T). Then tk=kT 2N,k=0,1,2,...,2N−1. (14.83) We shall prove that the exponential functions exp (2πiptk/T)and exp(2πiqtk/T)satisfy anorthogonalityrelationoverthediscretepoints tk: 2N−1summationdisplay k=0bracketleftbigg expparenleftbigg2πiptk Tparenrightbiggbracketrightbigg∗ expparenleftbigg2πiqtk Tparenrightbigg =2Nδp,q±2nN. (14.84) Heren,p, andqareallintegers. Replacing q−pbys,wefindthattheleft-handsideofEq. (14.84)becomes 2N−1summationdisplay k=0expparenleftbigg2πistk Tparenrightbigg =2N−1summationdisplay k=0expparenleftbigg2πisk 2Nparenrightbigg . 14.6 Discrete Fourier Transform 915 Thisright-handsideisobtainedbyusingEq.(14.83)toreplace T.Thisisafinitegeometric serieswithaninitialterm1andaratio r=expparenleftbiggπis Nparenrightbigg . FromEq. (5.3), 2N−1summationdisplay k=0expparenleftbigg2πistk Tparenrightbigg =  1−r2N 1−r=0,r/negationslash=1 2N, r =1,(14.85) establishing Eq. (14.84), our basic orthogonality relation. The upper value, zero, is a con- sequenceof r2N=exp(2πis)=1 forsan integer. The lower value, 2 N,f o rr=1 corresponds to p=q. The orthogonality ofthecorrespondingtrigonometricfunctionsisleftas Exercise14.6.1. Discrete Fourier Transform To simplify the notation and to make more direct contact with physics, we introduce the (reciprocal) ω-space,or angularfrequency,with ωp=2πp T,p=0,1,2,...,2N−1. (14.86) We make prange over the same integers as k. The exponential exp (±2πiptk/T)of Eq. (14.84) becomes exp (±iωptk). The choice of whether to use the +or the−sign is amatterofconvenienceorconvention.Inquantummechanicsthenegativesignisselected whenexpressingthetimedependence. Consider a function of time defined (measured) at the discrete time values tk. Then we construct F(ωp)=1 2N2N−1summationdisplay k=0f(tk)eiωptk. (14.87) Employingtheorthogonalityrelation,weobtain 1 2N2N−1summationdisplay p=0parenleftbig eiωptmparenrightbig∗eiωptk=δmk, (14.88) andthenreplacingthesubscript mbyk, wefindthattheamplitudes f(tk)become f(tk)=2N−1summationdisplay p=0F(ωp)e−iωptk. (14.89) The time function f(tk),k=0,1,2,...,2N−1, and the frequency function F(ωp),p= 0,1,2,...,2N−1, arediscreteFouriertransforms ofeachother.12CompareEqs. (14.87) 12Thetwo transform equations may be symmetrized with aresulting (2N)−1/2ineachequation ifdesired. 916 Chapter 14 Fourier Series and (14.89) with the corresponding continuous Fourier transforms, Eqs. (15.22) and (15.23)ofChapter15. Limitations Taken as a pair of mathematical relations, the discrete Fourier transforms are exact. We can say that the 2 N2N-component vectors exp (−iωptk),k=0,1,2,...,2N−1, form a complete set13spanning the tk-space. Then f(tk)in Eq. (14.89) is simply a particular linear combination of these vectors. Alternatively, we may take the 2 Nmeasured compo- nentsf(tk)as defining a 2 N-component vector in tk-space. Then, Eq. (14.87) yields the 2N-component vector F(ωp)in thereciprocal ωp-space. Equations (14.87) and (14.89) becomematrixequations,with exp (iωptk)/(2N)1/2theelementsof aunitarymatrix. The limitations of the discrete Fourier transform arise when we apply Eqs. (14.87) and (14.89) to physical systems and attempt physical interpretation and the limit F(ωp)→ F(ω).Example14.6.1illustratestheproblemsthatcanoccur.Themostimportantprecau- tion to be taken to avoid trouble is to take Nsufficiently large so that there is no angular frequency component of a higher angular frequency than ωN=2πN/T. For details on errors and limitations in the use of the discrete Fourier transform we refer to Hamming in theAdditionalReadings. Example 14.6.1 DISCRETE FOURIER TRANSFORM —A LIASING Considerthesimplecaseof T=2π,N=2,andf(tk)=costk.From tk=kT 4=kπ 2,k=0,1,2,3, (14.90) f(tk)=cos(tk)isrepresentedbythefour-componentvector f(tk)=(1,0,−1,0). (14.91) Thefrequencies, ωp, aregivenbyEq. (14.86): ωp=2πp T=p. (14.92) Clearly, cos tkimpliesa p=1 componentandnootherfrequencycomponents. Thetransformationmatrix exp(iωptk) 2N=exp(ipkπ/2) 2N becomes 1 4 1 111 1i−1−i 1−11−1 1−i−1i . (14.93) 13By Eq. (14.85) thesevectors areorthogonal and aretherefore linearly independent. 14.6 Discrete Fourier Transform 917 Note that the 2 N×2Nmatrix has only 2 Nindependent components. It is the repetition ofvaluesthat makesthefast Fouriertransform techniquepossible . Operatingoncolumnvector f(tk), wefindthatthismatrixyieldsacolumnvector F(ωp)=parenleftbig 0,1 2,0,1 2parenrightbig . (14.94) Apparently, there is a p=3 frequency component present. We reconstruct f(tk)by Eq.(14.89), obtaining f(tk)=1 2e−itk+1 2e−3itk. (14.95) Takingrealparts, wecanrewritetheequationas ℜf(tk)=1 2costk+1 2cos3tk. (14.96) Obviously, this result, Eq. (14.96), is not identical with our original f(tk)costk.B u t costk=1 2costk+1 2cos3tkattk=0,π/2,π; and 3π/2. The cos tkand cos3 tkmimic each other because of the limited number of data points (and the particular choice of data points).Thiserrorofonefrequencymimickinganotherisknownas aliasing.Theproblem canbeminimizedbytakingmoredatapoints. /squaresolid Fast Fourier Transform The fast Fourier transform is a particular way of factoring and rearranging the terms in the sums of the discrete Fourier transform. Brought to the attention of the scientific com- munity by Cooley and Tukey,14its importance lies in the drastic reduction in the number of numerical operations required. Because of the tremendous increase in speed achieved (and reduction in cost), the fast Fourier transform has been hailed as one of the few really significantadvancesinnumericalanalysisinthepastfew decades. ForNtime values (measurements), a direct calculation of a discrete Fourier transform wouldmeanabout N2multiplications.For Napowerof2,thefastFouriertransformtech- nique of Cooley and Tukey cuts the number of multiplications required to (N/2)log2N. IfN=1024(=210), the fast Fourier transform achieves a computational reduction by a factor of over 200. This is why the fast Fourier transform is called fast and why it has rev- olutionized the digital processing of waveforms. Details on the internal operation will be foundinthepaperbyCooleyandTukeyandinthepaperbyBergland.15 14J. W.Cooley and J.W. Tukey, Math. Comput. 19: 297 (1965). 15G. D. Bergland, A guided tour of the fast Fourier transform, IEEE Spectrum , July, pp. 41–52 (1969); see also, W. H. Press, B.P.Flannery,S.A.Teukolsky,andW.T.Vetterling, NumericalRecipes ,2nded.,Cambridge,UK:CambridgeUniversityPress (1996), Section 12.3. 918 Chapter 14 Fourier Series Exercises 14.6.1 Derivethetrigonometricforms ofdiscreteorthogonalitycorrespondingtoEq. (14.84): 2N−1summationdisplay k=0cosparenleftbigg2πptk Tparenrightbigg sinparenleftbigg2πqtk Tparenrightbigg =0 2N−1summationdisplay k=0cosparenleftbigg2πptk Tparenrightbigg cosparenleftbigg2πqtk Tparenrightbigg =  0,p/negationslash=q N, p=q/negationslash=0,N 2N, p=q=0,N 2N−1summationdisplay k=0sinparenleftbigg2πptk Tparenrightbigg sinparenleftbigg2πqtk Tparenrightbigg =  0,p/negationslash=q N, p=q/negationslash=0,N 0,p=q=0,N. Hint.Trigonometricidentitiessuchas sinAcosB=1 2bracketleftbig sin(A+B)+sin(A−B)bracketrightbig areuseful. 14.6.2 Equation (14.84) exhibits orthogonality summing over time points. Show that we have thesameorthogonalitysummingoverfrequencypoints 1 2N2N−1summationdisplay p=0parenleftbig eiωptmparenrightbig∗eiωptk=δmk. 14.6.3 Showindetailhowtogofrom F(ωp)=1 2N2N−1summationdisplay k=0f(tk)eiωptk to f(tk)=2N−1summationdisplay p=0F(ωp)e−iωptk. 14.6.4 Thefunctions f(tk)andF(ωp)arediscreteFouriertransformsofeachother.Derivethe followingsymmetryrelations: (a) Iff(tk)is real,F(ωp)isHermitiansymmetric;thatis, F(ωp)=F∗parenleftbigg4πN T−ωpparenrightbigg . (b) Iff(tk)is pureimaginary, F(ωp)=−F∗parenleftbigg4πN T−ωpparenrightbigg . Note. The symmetry of part (a) is an illustration of aliasing. The frequency 4 πN/T− ωpmasqueradesasthefrequency ωp. 14.7 Fourier Expansions of Mathieu Functions 919 14.6.5 GivenN=2,T=2π,andf(tk)=sintk, (a) find F(ωp),p=0,1,2,3,and (b) reconstruct f(tk)fromF(ωp)andexhibitthealiasingof ω1=1 andω3=3. ANS.(a) F(ωp)=(0,i/2,0,−i/2) (b)f(tk)=1 2sintk−1 2sin3tk. 14.6.6 ShowthattheChebyshevpolynomials Tm(x)satisfy adiscreteorthogonalityrelation 1 2Tm(−1)Tn(−1)+N−1summationdisplay s=1Tm(xs)Tn(xs)+1 2Tm(1)Tn(1)=  0,m/negationslash=n N/2,m=n/negationslash=0 N, m=n=0. Here,xs=cosθs, wherethe (N+1)θsare equallyspacedalongthe θ-axis: θs=sπ N,s=0,1,2,...,N. 14.7 F OURIER EXPANSIONS OF MATHIEU FUNCTIONS As a realistic application of Fourier series we now derive first integral equations satisfied byMathieufunctions,from whichsubsequentlytheirFourierseriesare obtained. Integral Equations and Fourier Series for Mathieu Functions Our first goal is to establish Whittaker’s integral equations that Mathieu functions satisfy, fromwhichwethenobtaintheirFourierseriesrepresentations. We startfromanintegralrepresentation V(r)=integraldisplayπ −πf(z+ixcosθ+iysinθ,θ)dθ (14.97) of a solution Vof Laplace’s equation with a twice differentiable function f(v,θ).Apply- ing∇2toVweverifythatitobeysLaplace’sPDE.SeparatingvariablesinLaplace’sPDE suggestschoosingtheproductform f(v,θ)=ekvφ(θ).Substitutingtheellipticalvariables ofEq. (13.163)werewrite Vas R(ξ)/Phi1(η)ekz=integraldisplayπ −πφ(θ)ek(z+iccoshξcosηcosθ+icsinhξsinηsinθ)dθ (14.98) with normalization R(0)=1.Sinceξandηare independent variables we may set ξ=0, whichleadstoWhittaker’s integralrepresentation /Phi1(η)=integraldisplayπ −πφ(θ)exp(ickcosθcosη)dθ, (14.99) 920 Chapter 14 Fourier Series whereck=2√qfrom Eq. (13.180). Clearly, /Phi1is evenin the variable ηandperiodicwith periodπ.In order to prove that φ∼/Phi1we check how φ(θ)is constrained when /Phi1(η)is takentoobeytheangularMathieuODE d2/Phi1 dη2+(λ−2qcos2η)/Phi1(η) =integraldisplayπ −πφ(θ)exp(ickcosθcosη) ·bracketleftbig λ−2qcos2η+(ickcosθsinη)2−ickcosθcosηbracketrightbig dθ.(14.100) Hereweintegratethelasttermontheright-handsidebyparts, obtaining d2/Phi1 dη2+(λ−2qcos2η)/Phi1(η) =φ(θ)(−ickcosηsinθ)exp(ickcosθcosη)vextendsinglevextendsinglevextendsingleπ θ=−π +integraldisplayπ −πφ(θ)exp(ickcosθcosη)[λ−2qcos2η−ickcosθcosη]dθ +integraldisplayπ −πbracketleftbig −φ′(θ)(−ickcosηsinθ)+φ(θ)ickcosηcosθbracketrightbig exp(ickcosθcosη)dθ =integraldisplayπ −πexpickcosθcosηbracketleftbig φ(θ)(λ−2qcos2η)+φ′(θ)ickcosηsinθbracketrightbig dθ,(14.101) where the integrated term vanishes if φ(−π)=φ(π),which we assume to be the case. Integratingoncemorebyparts yields d2/Phi1 dη2+(λ−2qcos2η)/Phi1(η)=−φ′(θ)expickcosθcosηvextendsinglevextendsinglevextendsingleπ θ=−π +integraldisplayπ −πexp(ickcosθcosη)bracketleftbig φ(θ)(λ−2qcos2η)+φ′′(θ)bracketrightbig dθ,(14.102) where the integrated term vanishes if φ′is periodic with period π, which we assume is thecase.Therefore,if φ(θ)obeystheangularMathieuODE,sodoestheintegral /Phi1(η),in Eq. (14.99). As a consequence, φ(θ)∼/Phi1(θ), where the constantmay be a functionof the parameter q. Thus,wehavethemainresultthatasolution /Phi1(η)ofMathieu’sODEthatiseveninthe variableηsatisfies theintegralequation /Phi1(η)=/Lambda1n(q)integraldisplayπ −πe2i√qcosθcosη/Phi1(θ)dθ. (14.103) When these Mathieu functions are expanded in a Fourier cosine series and normalized so thattheleadingtermis cos nη,theyare denotedbyce n(η,q). 14.7 Fourier Expansions of Mathieu Functions 921 Similarly, solutions of Mathieu’s ODE that are odd in ηwith leading term sin nηin a Fourierseriesaredenotedbyse n(η,q),andtheycansimilarlybeshowntoobeytheintegral equation sen(η,q)=sn(q)integraldisplayπ −πsinparenleftbig 2i√qsinηsinθparenrightbig sen(θ,q)dθ. (14.104) We now come to the Fourier expansion for the angular Mathieu functions and start with se1(η,q)=sinη+∞summationdisplay ν=1βν(q)sin(2ν+1)η, βν(q)=∞summationdisplay µ=νβ(ν) µqµ(14.105) as a paradigm for the systematic construction of Mathieu functions of odd parity. Notice the key point that the coefficient βνof sin(2ν+1)ηin the Fourier series depends on the parameter qand is expanded in a power series. Moreover, se 1is normalized so that the coefficient of the leading term, sin η,is unity, that is, independent of q.This feature will become important when se 1is substituted into the angular Mathieu ODE to determine the eigenvalue λ(q). Thefactthatthe βνpowerseriesin qstartswithexponent νcanbeprovedbyasimpler butsimilarseries for se 1(η,q): se1(η,q)=∞summationdisplay ν=0γν(q)sin2ν+1η, γ ν(q)=∞summationdisplay µ=0γ(ν) µqµ,(14.106) whichisusefulfor thisdemonstrationalone.However,sinceweneedtoexpand sin2ν+1η=νsummationdisplay m=0Bνmsin(2m+1)η (14.107) withFouriercoefficients Bνm=1 πintegraldisplayπ −πsin2ν+1ηsin(2m+1)ηdη=(−1)m 22νparenleftbigg2ν+1 ν−mparenrightbigg (14.108) that we can look up in a table of integrals (see Gradshteyn and Ryzhik in the Additional Readings of Chapter 13), this proof gives us an opportunity to introduce the Bνmthat are nonzeroonlyif m≤νandareimportantingredientsoftherecursionrelationsforthelead- ing terms of se 1(and all other Mathieu functions of odd parity). Substituting Eq. (14.107) intoEq. (14.106)weobtain se1(η,q)=∞summationdisplay ν=0νsummationdisplay m=0Bνmγν(q)sin(2m+1)η. (14.109) Comparingthisexpressionfor se 1withEq.(14.105)wefind βν(q)=∞summationdisplay m=νBmνγm(q). (14.110) Here,thesumstartswith m=νbecauseBmν=0f o rm<ν. 922 Chapter 14 Fourier Series NextwesubstituteEq.(14.106)intotheintegralEq.(14.104)for n=1,whereweinsert thepowerseries for sin (2i√qsinηsinθ).This yields se1(η,q) 2πs1(q)=1 2πintegraldisplayπ −πsin(2i√qsinηsinθ)se1(θ,q)dθ =1 2πs1(q)∞summationdisplay m=0γm(q)sin2m+1η =i√q∞summationdisplay m,ν=0qmγν(q)sin2m+1η22m+1 (2m+1)!1 2πintegraldisplayπ −πsin2ν+2m+2θdθ, (14.111) fromwhichweobtaintherecursionrelations γm(q)=22m+1 (2m+1)!qmi√qs1(q)∞summationdisplay ν=0γν(q)integraldisplayπ −πsin2ν+2m+2θdθ, (14.112) uponcomparingcoefficientsofsin2m+1η.Thisshowsthatthepowerseriesfor γm(q)starts withqm.Using Eq. (14.110) proves that the power series for βm(q)also starts with qm, and this confirms Eq. (14.105). The integral in Eq. (14.112) can be evaluated analytically and expressed via the beta function (Chapter 8) in terms of ratios of factorials, but we do notneedthisformulahere. Our next goal is to establish a recursion relation for the leading term β(ν) νof se1,i n Eq.(14.105).WesubstituteEq.(14.105)intotheintegralEq.(14.104)for n=1,wherewe insertthepowerseries for sin (2i√qsinηsinθ)again,alongwiththeexpansion 1 2πs1(q)=i∞summationdisplay m=0αmqm+1/2. (14.113) Here, the extra factor, i√q, cancels the corresponding factor from the sine in the integral equation.Thisyields ∞summationdisplay ν=0β(λ) µqµ+νsin2ν+1η22ν+1 (2ν+1)!1 2πintegraldisplayπ −πsin2ν+2λ+2θdθ =∞summationdisplay m,µ=0αmqm+µβ(ν) µsin(2ν+1)η. (14.114) Here, we replace sin2ν+1ηby sin(2m+1)ηusing Eq. (14.107). Upon comparing the co- efficientsof qNsin(2ν+1)ηforN=µ+νweobtaintherecursionrelation Nsummationdisplay ν=nN−νsummationdisplay λ=0β(λ) N−ν22ν (2ν+1)!BνnBνλ=N−nsummationdisplay m=0αmβ(n) N−m. (14.115) 14.7 Fourier Expansions of Mathieu Functions 923 Now we substitute Eq. (14.108) to obtain the main recursion relation for the leading coefficients β(ν) νof se1: Nsummationdisplay ν=nN−νsummationdisplay λ=0β(λ) N−ν 22ν(2ν+1)!parenleftbigg2ν+1 ν−nparenrightbiggparenleftbigg2ν+1 ν−λparenrightbigg =N−nsummationdisplay m=0αmβ(n) N−m.(14.116) Example 14.7.1 LEADING COEFFICIENTS OF se1 We evaluate Eq. (14.116) starting with N=0,n=0.For this case we find β(0) 0=α0β(0) 0, orα0=1 because the coefficient of sin ηin se1,β(0) 0=1,by normalization. For N=1, n=0 Eq. (14.116)yields α0β(0) 1+α1β(0) 0=β(0) 1+1 4·3!parenleftbigg3 1parenrightbiggbracketleftbigg β(0) 0parenleftbigg31parenrightbigg +β(1) 0parenleftbigg30parenrightbiggbracketrightbigg ,(14.117) whereβ(1) 0=0 andβ(0) 1drops out, a general feature. Of course, β(0) 1=0 because sin ηin se1hascoefficientunity.Thisyields α1=3/8. Thecase N=1,n=1 yields −1 4·3!parenleftbigg3 0parenrightbigg β(0) 0parenleftbigg31parenrightbigg =α0β(1) 1, (14.118) orβ(1) 1=−1/8.Theleadingtermisobtainedfromthegeneralcase n=N, (−1)N 22N(2N+1)!parenleftbigg2N+1 0parenrightbigg β(0) 0parenleftbigg2N+1 Nparenrightbigg =α0β(N) N, (14.119) as β(N) N=(−1)N 22N(2N+1)!parenleftbigg2N+1 Nparenrightbigg , (14.120) which was first derived by Mathieu. For N=1 this formula reproduces our earlier result, β(1) 1=−1/8. /squaresolid In order to determine the first nonleading term β(N) N+1of se1,Eq. (14.105), and the eigenvalue λ1(q)we substitute se 1into the angular Mathieu ODE, Eq. (13.181), using thetrigonometricidentities 2cos2ηsin(2ν+1)η=sin(2ν+3)η+sin(2ν−1)η and d2sin(2ν+1)η dη2=−(2ν+1)2sin(2ν+1)η. 924 Chapter 14 Fourier Series Thisyields 0=d2se1 dη2+(λ1−2qcos2η)se1=q(sinη−sin3η)+λ1sinη−sinη +∞summationdisplay ν=1bracketleftbig λ1−(2ν+1)2bracketrightbigbracketleftbigg(−1)νqν 22νν!(ν+1)!+β(ν) ν+1qν+1+···bracketrightbigg sin(2ν+1)η −q∞summationdisplay ν=1bracketleftbigg(−q)ν 22νν!(ν+1)!+β(ν) ν+1qν+1+···bracketrightbiggparenleftbig sin(2ν+3)η+sin(2ν−1)ηparenrightbig =parenleftbigg λ1−1+q−qbracketleftbigg −q 222!+β(1) 2q2+···bracketrightbiggparenrightbigg sinη +sin3ηbracketleftbigg −q−qparenleftbiggq2 242!3!+β(2) 3q3parenrightbigg +parenleftbig λ1−32parenrightbigparenleftbigg −q 222!+β(1) 2q2parenrightbiggbracketrightbigg +sin(2ν+1)ηbracketleftbig λ1−(2ν+1)2bracketrightbigparenleftbigg(−q)ν 22νν!(ν+1)!+β(ν) ν+1qν+1parenrightbigg −qsin(2ν+1)ηparenleftbigg(−q)ν+1 22(ν+1)(ν+1)!(ν+2)!+β(ν+1) ν+2qν+2parenrightbigg −qsin(2ν+1)ηparenleftbigg(−q)ν−1 22(ν−1)(ν−1)!ν!+β(ν−1) νqνparenrightbigg +···. (14.121) In this series the coefficient of each power of qwithin different sine terms must vanish; thatof sin ηbeingzeroyieldstheeigenvalue λ1(q)=1−q−1 8q2+β(1) 2q3+···, (14.122) withβ(1) 2=1/26coming from the vanishing coefficient of q2in sin3η.Setting the coeffi- cientof(−q)νin sin(2ν+1)ηequaltozeroyieldstheidentity bracketleftbig 1−(2ν+1)2bracketrightbig1 22νν!(ν+1)!+1 22(ν−1)(ν−1)!ν!=0, (14.123) which verifies the correct determination of the leading terms β(ν) νin Eq. (14.120). The vanishingcoefficientof qν+1in sin(2ν+1)ηyields (−1)ν+1 22νν!(ν+1)!+bracketleftbig 1−(2ν+1)2bracketrightbig β(ν) ν+1−β(ν−1) ν=0, (14.124) whichimpliesthemain recursionrelationfornonleadingcoefficients , 4ν(ν+1)β(ν) ν+1=−β(ν−1) ν+(−1)ν+1 22νν!(ν+1)!, (14.125) forthefirst nonleadingterms.We verifythat β(ν) ν+1=(−1)ν+1ν 22ν+2(ν+1)!2(14.126) 14.7 Fourier Expansions of Mathieu Functions 925 satisfies this recursion relation. Higher nonleading terms may be obtained by setting to zerothecoefficientof qν+2, etc.AltogetherwehavederivedtheFourierseries for se1(η,q)=sinη+∞summationdisplay ν=1bracketleftbigg(−q)ν 22νν!(ν+1)!+(−q)ν+1ν 22ν+2(ν+1)!2+···bracketrightbigg ·sin(2ν+1)η. (14.127) AsimilartreatmentyieldstheFourierseriesforse 2n+1(η,q)andse2n(η,q).Aninvariance ofMathieu’sODEleadstothesymmetryrelation ce2n+1(η,q)=(−1)nse2n+1(η+π/2,−q), (14.128) which allows us to determine the ce 2n+1of period 2 πfrom se 2n+1.Similarly, ce2n(η+π/2,−q)=se2n(η,q)relatestheseMathieufunctionsofperiod πtoeachother. Finally,webrieflyoutlineaderivationoftheFourierseries for ce0(η,q)=1+∞summationdisplay n=1βn(q)cos2nη, β n(q)=∞summationdisplay m=nβ(n) mqm,(14.129) as a paradigm for the Mathieu functions of period π.Note that this normalization agrees with Whittaker and Watson and with Hochstadtin the AdditionalReadings of Chapter 13, whereas in AMS-55 (for the full reference see footnote 4 in Chapter 5) ce 0differs by a factorof 1 /√ 2.ThesymmetryrelationfromtheMathieuODE, ce0parenleftbiggπ 2−η,−qparenrightbigg =ce0(η,q), (14.130) implies βn(−q)=(−1)nβn(q); (14.131) thatis,β2ncontainsonlyevenpowersof qandβ2n+1onlyoddpowers. The fact that the power series for βn(q)in Eq. (14.129) starts with qncan be proved by thesimilarexpansion ce0(η,q)=∞summationdisplay n=0γn(q)cos2nη, γ n(q)=∞summationdisplay µ=0γ(n) µqµ, (14.132) asforse 1inEqs.(14.105)to(14.112).SubstitutingEq.(14.132)intotheintegralequation ce0(η,q)=c0(q)integraldisplayπ −πe2i√qcosθcosηce(θ,q)dθ, (14.133) inserting the power series for the exponential function (odd powers cos2m+1θdrop out) andequatingthecoefficientsof cos2mηyields γm(q)=c1(q)(−q)m22m (2m)!∞summationdisplay µ=0γµ(q)integraldisplayπ −πcos2m+2µθdθ. (14.134) 926 Chapter 14 Fourier Series Thisrecursionrelationshowsthatthepowerseries for γm(q)startswith qm.Weexpand cos2nη=nsummationdisplay m=0Anmcos2mη (14.135) withFouriercoefficients Anm=2 πintegraldisplayπ/2 −π/2cos2nηcos2mηdη=1 22n−1parenleftbigg2n n−mparenrightbigg , (14.136) which are nonzero only when m≤n.Using this result to replace the cosine powers in Eq.(14.132)by cos2 mηweobtain βn(q)=∞summationdisplay m=nAmnγm(q), (14.137) confirmingEq. (14.129). Proceeding as for se 1in Eqs. (14.113) to (14.120) we substitute Eq. (14.129) into the integralEq.(14.133)andobtain ∞summationdisplay m,µ,ν,λ=0(−1)m22m (2m)!qµ+mmsummationdisplay ν=0Amνcos(2νη)Amλ =∞summationdisplay m,µ,ν=0αmβ(ν) µqm+µcos(2νη), (14.138) with 1 2πc1(q)=∞summationdisplay m=0αmqm. (14.139) Uponcomparingthecoefficientof qNcos(2νη)withN=m+µ,weextractthe recursion relationforleadingcoefficients β(n) nof ce0 Nsummationdisplay m=ν(−1)m22m (2m)!Amνmsummationdisplay λ=0β(λ) N−mAmλ=Nsummationdisplay m=0αmβ(ν) N−m, (14.140) withAnminEq. (14.136). Example 14.7.2 LEADING COEFFICIENTS FOR ce0 Thecase N=0,ν=0 ofEq. (14.140)yields A2 00β(0) 0=α0β(0) 0, (14.141) withA00=1andβ(0) 0=1fromnormalizingtheleadingtermofce 0tounitysothat α0=1 results. Thecase N=1,ν=0 yields A00β(0) 1A00−2A10bracketleftbig β(0) 0A10+β(1) 0A11bracketrightbig =α0β(0) 1+α1β(0) 0,(14.142) 14.7 Fourier Expansions of Mathieu Functions 927 withβ(1) 0=0 byEq. (14.129).This simplifiesto β(0) 1−1 2β(0) 0=α1β(0) 0+β(0) 1, (14.143) whereβ(0) 1drops out. We know already that β(0) 1=0 from the leading term unity of ce 0. Therefore, α1=−1/2. Forthecase N=1,ν=1 weobtain −2A11β(0) 0A10=α0β(1) 1, (14.144) withA10=1/2=A11,from which β(1) 1=−1/2 follows. For the case N=2,ν=2w e find 24 4!23β(0) 0A20=α0β(2) 2, (14.145) withA20=3/8,from which β(2) 2=2−4follows.Thegeneralcase N,ν=Nyields (−1)N22N (2N)!ANNβ(0) 0AN0=α0β(N) N, (14.146) withANN=2−2N+1,AN0=1 22N−1parenleftbig2N Nparenrightbig , fromwhichtheleadingterm β(N) N=(−1)N 22N−1(2N)!parenleftbigg2N Nparenrightbigg =(−1)N 22N−1N!2(14.147) follows. /squaresolid The nonleading terms β(N) N+1of ce0are best determined from the angular Mathieu ODE by substitution of Eq. (14.129), in analogy with se 1,Eqs. (14.121) to (14.127). Using the identities 2cos(2nη)cos2η=cos(2n+2)η+cos(2n−2)η, d2 dη2cos(2nη)=−(2n)2cos(2nη), (14.148) weobtain d2ce0 dη2+parenleftbig λ0(q)−2qcos2ηparenrightbig ce0=0 =λ0(q)−qparenleftbigg −q 2+7 27q3parenrightbigg +··· +∞summationdisplay n=1parenleftbig λ0−4n2parenrightbig cos(2nη)bracketleftbigg(−q)n 22n−1n!2+β(n) n+2qn+2bracketrightbigg −q∞summationdisplay n=1bracketleftbigg(−q)n 22n−1n!2+β(n) n+2qn+2bracketrightbigg ×bracketleftbig cos(2n+2)η+cos(2n−2)ηbracketrightbig . (14.149) 928 Chapter 14 Fourier Series Settingthecoefficientof cos (2nη)forn=0 tozeroyieldstheeigenvalue λ0=−1 2q2+7 27q4+···. (14.150) Thecoefficientof cos (2nη)qnyieldsanidentity, (−1)n+14n2 22n−1n!2+(−1)n 22n−3(n−1)!2=0, (14.151) which shows that the leading term in Eq. (14.147) was correctly determined. The coeffi- cientofqn+2cos(2nη)yieldstherecursionrelation −4n2β(n) n+2−β(n−1) n+1+(−1)n+1 22nn!2+(−1)n 22n+1(n+1)!2=0. (14.152) Itis straightforwardtocheckthat β(n) n+2=(−1)n+1n(3n+4) 22n+3(n+1)!2(14.153) satisfiesthisrecursionrelation.Altogetherwehavederivedtheformula ce0(η,q)=1+cos2ηbracketleftbigg −1 2q2+7 27q3+···bracketrightbigg +cos4ηbracketleftbiggq2 25+···bracketrightbigg +cos6ηbracketleftbigg −q3 2732+···bracketrightbigg =1+∞summationdisplay n=1cos(2nη)bracketleftbigg(−q)n 22n−1n!2+(−1)n+1n(3n+4)qn+2 22n+3(n+1)!2+···bracketrightbigg . (14.154) Similarlyonecanderive ce1(η,q)=cosη+∞summationdisplay n=1cos(2n+1)ηbracketleftbigg(−q)n 22nn!(n+1)!−(−q)n+1n 22n+2(n+1)12+···bracketrightbigg , (14.155) whoseeigenvalueis givenbythepowerseries λ1(q)=1+q−1 8q2−1 26q3+···. (14.156) 14.7 Additional Readings 929 FIGURE 14.12AngularMathieufunctions.(FromGutiérrez-Vega etal., Am. J.Phys. 71:233(2003).) Exercises 14.7.1 Determinethenonleadingcoefficients β(n) n+2forse1.Deriveasuitablerecursionrelation. 14.7.2 Determinethenonleadingcoefficients β(n) n+4force0.Derivethecorrespondingrecursion relation. 14.7.3 Derivetheformulafor ce 1, Eq.(14.155), andits eigenvalue,Eq. (14.156). AdditionalReadings Carslaw, H. S., Introduction to the Theory of Fourier’s Series and Integrals , 2nd ed. London: Macmillan (1921); 3rd ed., paperback, New York: Dover (1952). This is a detailed and classic work; includes a considerable discussion of Gibbs phenomenon in Chapter IX. Hamming, R. W., Numerical Methods for Scientists and Engineers , 2nd ed. New York: McGraw-Hill (1973), reprinted Dover (1987). Chapter 33 provides an excellentdescription ofthe fast Fourier transform. Jeffreys,H.,andB.S.Jeffreys, MethodsofMathematicalPhysics ,3rded.Cambridge,UK:CambridgeUniversity Press (1972). Kufner, A., and J. Kadlec, Fourier Series . London: Iliffe (1971). This book is a clear account of Fourier series in the context ofHilbert space. 930 Chapter 14 Fourier Series Lanczos, C., Applied Analysis , Englewood Cliffs, NJ: Prentice-Hall (1956), reprinted Dover (1988). The book givesawell-writtenpresentationoftheLanczosconvergencetechnique(whichsuppressestheGibbsphenom- enon oscillations). This and several other topics are presented from the point of view of a mathematician who wantsuseful numerical results and not just abstractexistence theorems. Oberhettinger, F., Fourier Expansions, ACollection of Formulas . NewYork, AcademicPress (1973). Zygmund, A., Trigonometric Series . Cambridge, UK: Cambridge University Press (1988). The volume contains an extremely complete exposition, including relatively recent results in therealmof pure mathematics. CHAPTER 15 INTEGRAL TRANSFORMS 15.1 I NTEGRAL TRANSFORMS Frequently in mathematical physics we encounter pairs of functions related by an expres- sionoftheform g(α)=integraldisplayb af(t)K(α,t)dt. (15.1) The function g(α)is called the (integral) transform of f(t)by the kernel K(α,t). The operation may also be described as mapping a function f(t)int-space into another function, g(α),i nα-space. This interpretation takes on physical significance in the time- frequency relation of Fourier transforms, as in Example 15.3.1, and in the real space– momentumspacerelationsinquantumphysicsof Section15.6. Fourier Transform One of the most useful of the infinite number of possible transforms is the Fourier trans- form,givenby g(ω)=1√ 2πintegraldisplay∞ −∞f(t)eiωtdt. (15.2) Two modifications of this form, developed in Section 15.3, are the Fourier cosine and Fouriersinetransforms: gc(ω)=radicalbigg 2 πintegraldisplay∞ 0f(t)cosωtdt, (15.3) gs(ω)=radicalbigg 2 πintegraldisplay∞ 0f(t)sinωtdt. (15.4) 931 932 Chapter 15 Integral Transforms TheFouriertransformisbasedonthekernel eiωtanditsrealandimaginarypartstakensep- arately, cos ωtand sinωt. Because these kernels are the functions used to describe waves, Fouriertransformsappearfrequentlyinstudiesofwavesandtheextractionofinformation from waves, particularly when phase information is involved. The output of a stellar in- terferometer, for instance, involves a Fourier transform of the brightness across a stellar disk.TheelectrondistributioninanatommaybeobtainedfromaFouriertransformofthe amplitude of scattered X-rays. In quantum mechanics the physical origin of the Fourier relationsofSection15.6isthewavenatureofmatterandourdescriptionofmatterinterms ofwaves. Example 15.1.1 FOURIER TRANSFORM OF GAUSSIAN TheFouriertransform ofaGaussianfunction e−a2t2, g(ω)=1√ 2πintegraldisplay∞ −∞e−a2t2eiωtdt, canbedoneanalyticallybycompletingthesquareintheexponent, −a2t2+iωt=−a2parenleftbigg t−iω 2a2parenrightbigg2 −ω2 4a2, whichwecheckbyevaluatingthesquare.Substitutingthisidentityweobtain g(ω)=1√ 2πe−ω2/4a2integraldisplay∞ −∞e−a2t2dt, upon shifting the integration variable t→t+iω 2a2.This is justified by an application of Cauchy’s theorem to the rectangle with vertices −T, T, T+iω 2a2,−T+iω 2a2forT→ ∞,noting that the integrand has no singularities in this region and that the integrals over the sides from ±Tto±T+iω 2a2become negligible for T→∞.Finally we rescale the integrationvariableas ξ=atintheintegral(seeEqs. (8.6) and(8.8)): integraldisplay∞ −∞e−a2t2dt=1 aintegraldisplay∞ −∞e−ξ2dξ=√π a. Substitutingtheseresultswefind g(ω)=1 a√ 2expparenleftbigg −ω2 4a2parenrightbigg , againaGaussian,butin ω-space.Thebigger ais,thatis,thenarrowertheoriginalGaussian e−a2t2is, thewiderisits Fouriertransform ∼e−ω2/4a2. /squaresolid 15.1 Integral Transforms 933 Laplace, Mellin, and Hankel Transforms Threeotherusefulkernelsare e−αt,tJn(αt), tα−1. Thesegiverise tothefollowingtransforms g(α)=integraldisplay∞ 0f(t)e−αtdt,Laplacetransform (15.5) g(α)=integraldisplay∞ 0f(t)tJn(αt)dt, Hankeltransform (Fourier–Bessel) (15.6) g(α)=integraldisplay∞ 0f(t)tα−1dt,Mellintransform . (15.7) Clearly,thepossibletypesareunlimited.Thesetransformshavebeenusefulinmathemati- calanalysisandinphysicalapplications.WehaveactuallyusedtheMellintransformwith- out calling it by name; that is, g(α)=(α−1)!is the Mellin transform of f(t)=e−t. See E.C. Titchmarsh, IntroductiontotheTheoryofFourierIntegrals ,2nded.,NewYork:Ox- fordUniversityPress(1937),formoreMellintransforms.Ofcourse,wecouldjustaswell sayg(α)=n!/αn+1istheLaplacetransformof f(t)=tn.Ofthethree,theLaplacetrans- formisbyfarthemostused.ItisdiscussedatlengthinSections15.8to15.12.TheHankel transform, a Fourier transform for a Bessel function expansion, represents a limiting case of a Fourier–Bessel series. It occurs in potential problems in cylindrical coordinates and hasbeenappliedextensivelyinacoustics. Linearity Alltheseintegraltransforms arelinear;thatis, integraldisplayb abracketleftbig c1f1(t)+c2f2(t)bracketrightbig K(α,t)dt =c1integraldisplayb af1(t)K(α,t)dt+c2integraldisplayb af2(t)K(α,t)dt, (15.8) integraldisplayb acf(t)K(α,t)dt =cintegraldisplayb af(t)K(α,t)dt, (15.9) wherec1andc2are constants and f1(t)andf2(t)are functions for which the transform operationisdefined. Representingourlinearintegraltransformbytheoperator L,weobtain g(α)=Lf(t). (15.10) 934 Chapter 15 Integral Transforms FIGURE 15.1Schematicintegraltransforms. Weexpectaninverseoperator L−1existssuchthat1 f(t)=L−1g(α). (15.11) ForourthreeFouriertransforms L−1isgiveninSection15.3.Ingeneral,thedetermination of the inverse transform is the main problem in using integral transforms. The inverse Laplace transform is discussed in Section 15.12. For details of the inverse Hankel and inverseMellintransformswerefer totheAdditionalReadingsattheendof thechapter. Integral transforms have many special physical applications and interpretations that are noted in the remainder of this chapter. The most common application is outlined in Fig. 15.1. Perhaps an original problem can be solved only with difficulty, if at all, in the original coordinates (space). It often happens that the transform of the problem can be solved relatively easily. Then the inverse transform returns the solution from the trans- formcoordinatestotheoriginalsystem.Example15.4.1andExercise15.4.1illustratethis technique. Exercises 15.1.1 TheFouriertransforms for afunctionoftwovariablesare F(u,v)=1 2πintegraldisplay∞ −∞integraldisplay f(x,y)ei(ux+vy)dxdy, f(x,y)=1 2πintegraldisplay∞ −∞integraldisplay F(u,v)e−i(ux+vy)dudv. Usingf(x,y)=f([x2+y2]1/2),showthatthezero-orderHankeltransforms F(ρ)=integraldisplay∞ 0rf(r)J0(ρr)dr, f(r)=integraldisplay∞ 0ρF(ρ)J 0(ρr)dρ, areaspecialcaseof theFouriertransforms. 1Expectationisnot proof, andhereproofofexistenceiscomplicatedbecauseweareactuallyin an infinite-dimensional Hilbert space.Weshall prove existence inthe specialcasesof interest byactual construction. 15.1 Integral Transforms 935 This technique may be generalized to derive the Hankel transforms of order ν= 0,1 2,1,1 2,...(compare I. N. Sneddon, Fourier Transforms , New York: McGraw-Hill (1951)).Amoregeneralapproach,validfor ν>−1 2,ispresentedinSneddon’s TheUse of Integral Transforms (New York: McGraw-Hill (1972)). It might also be noted that the Hankel transforms of nonintegral order ν=±1 2reduce to Fourier sine and cosine transforms. 15.1.2 AssumingthevalidityoftheHankeltransform–inversetransformpairof equations g(α)=integraldisplay∞ 0f(t)Jn(αt)tdt, f(t)=integraldisplay∞ 0g(α)Jn(αt)αdα, showthattheDiracdeltafunctionhasaBessel integralrepresentation δ(t−t′)=tintegraldisplay∞ 0Jn(αt)Jn(αt′)αdα. This expression is useful in developing Green’s functions in cylindrical coordinates, wheretheeigenfunctionsareBessel functions. 15.1.3 FromtheFouriertransforms, Eqs. (15.22)and(15.23),showthatthetransformation t→lnx iω→α−γ leadsto G(α)=integraldisplay∞ 0F(x)xα−1dx and F(x)=1 2πiintegraldisplayγ+i∞ γ−i∞G(α)x−αdα. These are the Mellin transforms. A similar change of variables is employed in Sec- tion15.12toderivetheinverseLaplacetransform. 15.1.4 VerifythefollowingMellintransforms: (a)integraldisplay∞ 0xα−1sin(kx)dx=k−α(α−1)!sinπα 2,−1<α<1. (b)integraldisplay∞ 0xα−1cos(kx)dx=k−α(α−1)!cosπα 2,0<α<1. Hint.Youcanforcetheintegralsintoatractableformbyinsertingaconvergencefactor e−bxand(after integrating)letting b→0.Also, cos kx+isinkx=expikx. 936 Chapter 15 Integral Transforms 15.2 D EVELOPMENT OF THE FOURIER INTEGRAL In Chapter 14 it was shown that Fourier series are useful in representing certain func- tions (1) over a limited range [0,2π],[−L,L], and so on, or (2) for the infinite interval (−∞,∞),if the function is periodic . We now turn our attention to the problem of rep- resenting a nonperiodic function over the infinite range. Physically this means resolving a singlepulseor wavepacketintosinusoidalwaves. We have seen (Section 14.2) that for the interval [−L,L]the coefficients anandbn couldbewrittenas an=1 LintegraldisplayL −Lf(t)cosnπt Ldt, (15.12) bn=1 LintegraldisplayL −Lf(t)sinnπt Ldt. (15.13) TheresultingFourierseriesis f(x)=1 2LintegraldisplayL −Lf(t)dt+1 L∞summationdisplay n=1cosnπx LintegraldisplayL −Lf(t)cosnπt Ldt +1 L∞summationdisplay n=1sinnπx LintegraldisplayL −Lf(t)sinnπt Ldt, (15.14) or f(x)=1 2LintegraldisplayL −Lf(t)dt+1 L∞summationdisplay n=1integraldisplayL −Lf(t)cosnπ L(t−x)dt. (15.15) Wenowlettheparameter Lapproachinfinity,transformingthefiniteinterval [−L,L]into theinfiniteinterval (−∞,∞).W es e t nπ L=ω,π L=/Delta1ω, withL→∞. Thenwehave f(x)→1 π∞summationdisplay n=1/Delta1ωintegraldisplay∞ −∞f(t)cosω(t−x)dt, (15.16) or f(x)=1 πintegraldisplay∞ 0dωintegraldisplay∞ −∞f(t)cosω(t−x)dt, (15.17) replacing the infinite sum by the integral over ω. The first term (corresponding to a0) has vanished,assumingthatintegraltext∞ −∞f(t)dtexists. It must be emphasized that this result (Eq. (15.17)) is purely formal. It is not intended as a rigorous derivation, but it can be made rigorous (compare I. N. Sneddon, Fourier Transforms , Section 3.2). We take Eq. (15.17) as the Fourier integral. It is subject to the conditions that f(x)is (1) piecewise continuous, (2) piecewise differentiable, and (3) ab- solutelyintegrable—thatis,integraltext∞ −∞|f(x)|dxisfinite. 15.2 Development of the Fourier Integral 937 Fourier Integral — Exponential Form OurFourierintegral(Eq. (15.17)) maybeputintoexponentialformbynotingthat f(x)=1 2πintegraldisplay∞ −∞dωintegraldisplay∞ −∞f(t)cosω(t−x)dt, (15.18) whereas 1 2πintegraldisplay∞ −∞dωintegraldisplay∞ −∞f(t)sinω(t−x)dt=0; (15.19) cosω(t−x)is an even function of ωand sinω(t−x)is an odd function of ω. Adding Eqs. (15.18)and(15.19)(withafactor i), weobtainthe Fourierintegraltheorem f(x)=1 2πintegraldisplay∞ −∞e−iωxdωintegraldisplay∞ −∞f(t)eiωtdt. (15.20) The variable ωintroduced here is an arbitrary mathematical variable. In many physical problems, however, it corresponds to the angular frequency ω. We may then interpret Eq. (15.18) or (15.20) as a representation of f(x)in terms of a distribution of infinitely long sinusoidal wave trains of angular frequency ω, in which this frequency is a continu- ousvariable. Dirac Delta Function Derivation If theorder ofintegrationof Eq.(15.20) isreversed,wemayrewriteitas f(x)=integraldisplay∞ −∞f(t)braceleftbigg1 2πintegraldisplay∞ −∞eiω(t−x)dωbracerightbigg dt. (15.20a) Apparently the quantity in curly brackets behaves as a delta function δ(t−x). We might take Eq. (15.20a) as presenting us with a representation of the Dirac delta function. Alter- natively,wetakeitas acluetoanewderivationoftheFourierintegraltheorem. FromEq. (1.171b)(shiftingthesingularityfrom t=0t ot=x), f(x)=limn→∞integraldisplay∞ −∞f(t)δn(t−x)dt, (15.21a) whereδn(t−x)is a sequence defining the distribution δ(t−x). Note that Eq. (15.21a) assumesthat f(t)is continuousat t=x.W etak e δn(t−x)tobe δn(t−x)=sinn(t−x) π(t−x)=1 2πintegraldisplayn −neiω(t−x)dω, (15.21b) usingEq. (1.174). SubstitutingintoEq.(15.21a), wehave f(x)=limn→∞1 2πintegraldisplay∞ −∞f(t)integraldisplayn −neiω(t−x)dωdt. (15.21c) 938 Chapter 15 Integral Transforms Interchanging the order of integration and then taking the limit as n→∞,w eh a v e Eq.(15.20), theFourierintegraltheorem. With the understanding that it belongs under an integral sign, as in Eq. (15.21a), the identification δ(t−x)=1 2πintegraldisplay∞ −∞eiω(t−x)dω (15.21d) providesaveryusefulrepresentationofthedeltafunction. 15.3 F OURIER TRANSFORMS —I NVERSION THEOREM Letusdefineg(ω), theFouriertransformofthefunction f(t),by g(ω)≡1√ 2πintegraldisplay∞ −∞f(t)eiωtdt. (15.22) Exponential Transform Then,fromEq. (15.20), wehavetheinverserelation, f(t)=1√ 2πintegraldisplay∞ −∞g(ω)e−iωtdω. (15.23) Note that Eqs. (15.22) and (15.23) are almost but not quite symmetrical, differing in the signofi. Here two points deserve comment. First, the 1 /√ 2πsymmetry is a matter of choice, not of necessity. Many authors will attach the entire 1 /2πfactor of Eq. (15.20) to one of the two equations: Eq. (15.22) or Eq. (15.23). Second, although the Fourier integral, Eq.(15.20),hasreceivedmuchattentioninthemathematicsliterature,weshallbeprimar- ilyinterestedintheFouriertransformanditsinverse.Theyaretheequationswithphysical significance. WhenwemovetheFouriertransformpairtothree-dimensionalspace,itbecomes g(k)=1 (2π)3/2integraldisplay f(r)eik·rd3r, (15.23a) f(r)=1 (2π)3/2integraldisplay g(k)e−ik·rd3k. (15.23b) The integrals are over all space. Verification, if desired, follows immediately by substitut- ingtheleft-handsideofoneequationintotheintegrandoftheotherequationandusingthe three-dimensional delta function.2Equation (15.23b) may be interpreted as an expansion of a function f(r)in a continuum of plane wave eigenfunctions; g(k)then becomes the amplitudeofthewave, exp (−ik·r). 2δ(r1−r2)=δ(x1−x2)δ(y1−y2)δ(z1−z2)with Fourier integral δ(x1−x2)=1 2πintegraltext∞ −∞exp[ik1(x1−x2)]dk1,etc. 15.3 Fourier Transforms — Inversion Theorem 939 Cosine Transform Iff(x)is odd or even, these transforms may be expressed in a somewhat different form. Consider first an even function fcwithfc(x)=fc(−x). Writing the exponential ofEq. (15.22)intrigonometricform,wehave gc(ω)=1√ 2πintegraldisplay∞ −∞fc(t)(cosωt+isinωt)dt =radicalbigg 2 πintegraldisplay∞ 0fc(t)cosωtdt, (15.24) the sinωtdependence vanishing on integration over the symmetric interval (−∞,∞). Similarly,since cos ωtis even,Eqs. (15.23) transformsto fc(x)=radicalbigg 2 πintegraldisplay∞ 0gc(ω)cosωxdω. (15.25) Equations(15.24)and(15.25)areknownas Fouriercosinetransforms. Sine Transform The corresponding pair of Fourier sine transforms is obtained by assuming that fs(x)= −fs(−x),odd,andapplyingthesamesymmetryarguments.Theequationsare gs(ω)=radicalbigg 2 πintegraldisplay∞ 0fs(t)sinωtdt,3(15.26) fs(x)=radicalbigg 2 πintegraldisplay∞ 0gs(ω)sinωxdω. (15.27) From the last equation we may develop the physical interpretation that f(x)is being describedbyacontinuumofsinewaves.Theamplitudeof sin ωxisgivenby√2/πgs(ω), inwhich gs(ω)istheFouriersinetransformof f(x).ItwillbeseenthatEq.(15.27)isthe integral analog of the summation (Eq. (14.24)). Similar interpretations hold for the cosine andexponentialcases. If we take Eqs. (15.22), (15.24), and (15.26) as the direct integral transforms, de- scribed by Lin Eq. (15.10) (Section 15.1), the corresponding inverse transforms, L−1 ofEq. (15.11),are givenbyEqs.(15.23), (15.25), and(15.27). Note that the Fourier cosine transforms and the Fourier sine transforms each involve onlypositivevalues(andzero)ofthearguments.Weusetheparityof f(x)toestablishthe transforms; but once the transforms are established, the behavior of the functions fandg for negative argument is irrelevant. In effect, the transform equations themselves impose adefinite parity :even for the Fourier cosine transform and odd for the Fourier sine transform. 3Notethat afactor −ihas beenabsorbed into this g(ω). 940 Chapter 15 Integral Transforms FIGURE 15.2Finitewavetrain. Example 15.3.1 FINITE WAVETRAIN An important application of the Fourier transform is the resolution of a finite pulse into sinusoidal waves. Imagine that an infinite wave train sin ω0tis clipped by Kerr cell or saturabledyecellshutters sothatwehave f(t)=  sinω0t,|t|<Nπ ω0, 0,|t|>Nπ ω0.(15.28) This corresponds to Ncycles of our original wave train (Fig. 15.2). Since f(t)is odd, we mayusetheFouriersinetransform(Eq. (15.26))toobtain gs(ω)=radicalbigg 2 πintegraldisplayNπ/ω0 0sinω0tsinωtdt. (15.29) Integrating,wefindouramplitudefunction: gs(ω)=radicalbigg 2 πbracketleftbiggsin[(ω0−ω)(Nπ/ω 0)] 2(ω0−ω)−sin[(ω0+ω)(Nπ/ω 0)] 2(ω0+ω)bracketrightbigg .(15.30) It is of considerable interest to see how gs(ω)depends on frequency. For large ω0and ω≈ω0, only the first term will be of any importance because of the denominators. It is plottedinFig.15.3.Thisis theamplitudecurvefor thesingle-slitdiffractionpattern. Therearezeros at ω0−ω ω0=/Delta1ω ω0=±1 N,±2 N,andso on . (15.31) Forlarge N,gs(ω)mayalsobeinterpretedasaDiracdeltadistribution,asinSection1.15. Sincethecontributionsoutsidethecentralmaximumaresmallinthiscase,wemaytake /Delta1ω=ω0 N(15.32) as a good measure of the spread in frequency of our wave pulse. Clearly, if Nis large (alongpulse),thefrequencyspreadwillbesmall.Ontheotherhand,ifourpulseisclipped 15.3 Fourier Transforms — Inversion Theorem 941 FIGURE 15.3Fouriertransform of finitewavetrain. short,Nsmall,thefrequencydistributionwillbewiderandthesecondarymaximaaremore important. /squaresolid Uncertainty Principle Hereisaclassicalanalogofthefamousuncertaintyprincipleofquantummechanics.Ifwe aredealingwithelectromagneticwaves, hω 2π=E,energy(ofourphoton ) h/Delta1ω 2π=/Delta1E, (15.33) hbeing Planck’s constant. Here /Delta1Erepresents an uncertainty in the energy of our pulse. Thereisalsoanuncertaintyinthetime,forourwaveof Ncyclesrequires2 Nπ/ω0seconds topass.Taking /Delta1t=2Nπ ω0, (15.34) wehavetheproductof thesetwouncertainties: /Delta1E·/Delta1t=h/Delta1ω 2π·2πN ω0=hω0 2πN·2πN ω0=h. (15.35) TheHeisenberguncertaintyprincipleactuallystates /Delta1E·/Delta1t≥h 4π, (15.36) andthisis clearlysatisfiedinourexample. 942 Chapter 15 Integral Transforms Exercises 15.3.1 (a) Show that g(−ω)=g∗(ω)is a necessary and sufficient condition for f(x)to be real. (b) Showthat g(−ω)=−g∗(ω)isanecessaryandsufficientconditionfor f(x)tobe pureimaginary. Note. Theconditionof part(a) is used inthedevelopmentof thedispersionrelationsof Section7.2. 15.3.2 LetF(ω)betheFourier(exponential)transformof f(x)andG(ω)betheFouriertrans- form ofg(x)=f(x+a). Showthat G(ω)=e−iaωF(ω). 15.3.3 Thefunction f(x)=braceleftbigg1,|x|<1 0,|x|>1 isa symmetricalfinitestepfunction. (a) Findthe gc(ω), Fouriercosinetransformof f(x). (b) Takingtheinversecosinetransform, showthat f(x)=2 πintegraldisplay∞ 0sinωcosωx ωdω. (c) Frompart (b)showthat integraldisplay∞ 0sinωcosωx ωdω=  0,|x|>1, π 4,|x|=1, π 2,|x|<1. 15.3.4 (a) ShowthattheFouriersineandcosinetransforms of e−atare gs(ω)=radicalbigg 2 πω ω2+a2,g c(ω)=radicalbigg 2 πa ω2+a2. Hint.Eachof thetransforms canberelatedtotheotherbyintegrationbyparts. (b) Showthat integraldisplay∞ 0ωsinωx ω2+a2dω=π 2e−ax,x>0, integraldisplay∞ 0cosωx ω2+a2dω=π 2ae−ax,x>0. Theseresultsare alsoobtainedbycontourintegration(Exercise7.1.14). 15.3.5 FindtheFouriertransformofthetriangularpulse(Fig.15.4). f(x)=braceleftBigghparenleftbig 1−a|x|parenrightbig ,|x|<1 a, 0, |x|>1 a. Note.This functionprovidesanotherdeltasequencewith h=aanda→∞. 15.3 Fourier Transforms — Inversion Theorem 943 FIGURE 15.4Triangularpulse. 15.3.6 Defineasequence δn(x)=braceleftBiggn,|x|<1 2n, 0,|x|>1 2n. (This is Eq. (1.172).) Express δn(x)as a Fourier integral (via the Fourier integral theo- rem,inversetransform,etc.). Finally,showthatwemaywrite δ(x)=limn→∞δn(x)=1 2πintegraldisplay∞ −∞e−ikxdk. 15.3.7 Usingthesequence δn(x)=n√πexpparenleftbig −n2x2parenrightbig , showthat δ(x)=1 2πintegraldisplay∞ −∞e−ikxdk. Note. Remember that δ(x)is defined in terms of its behavior as part of an integrand (Section1.15), especiallyEqs. (1.178) and(1.179). 15.3.8 Derivesineandcosinerepresentationsof δ(t−x)thatarecomparabletotheexponential representation,Eq. (15.21d). ANS.2 πintegraldisplay∞ 0sinωtsinωxdω,2 πintegraldisplay∞ 0cosωtcosωxdω. 15.3.9 Inaresonantcavityanelectromagneticoscillationoffrequency ω0diesoutas A(t)=A0e−ω0t/2Qe−iω0t,t>0. (TakeA(t)=0fort<0.)Theparameter Qisameasureoftheratioofstoredenergyto energylosspercycle.Calculatethefrequencydistributionoftheoscillation, a∗(ω)a(ω), wherea(ω)is theFouriertransformof A(t). Note.Thelarger Qis, thesharperyourresonancelinewillbe. ANS.a∗(ω)a(ω)=A2 0 2π1 (ω−ω0)2+(ω0/2Q)2. 944 Chapter 15 Integral Transforms 15.3.10 Provethat ¯h 2πiintegraldisplay∞ −∞e−iωtdω E0−iŴ/2−¯hω=braceleftBigg expparenleftbig −Ŵt 2¯hparenrightbig expparenleftbig −iE0t ¯hparenrightbig ,t>0, 0,t <0. This Fourier integral appears in a variety of problems in quantum mechanics: WKB barrierpenetration,scattering,time-dependentperturbationtheory,andso on. Hint.Trycontourintegration. 15.3.11 VerifythatthefollowingareFourierintegraltransforms ofoneanother: (a)radicalbigg 2 π·1√ a2−x2,|x|<a, andJ0(ay), 0, |x|>a, (b)0, |x|<a, −radicalbigg 2 π1√ x2+a2,|x|>a, andN0(a|y|), (c)radicalbiggπ 2·1√ x2+a2andK0parenleftbig a|y|parenrightbig . (d) Canyousuggestwhy I0(ay)is notincludedinthislist? Hint.J0,N0, andK0may be transformed most easily by using an exponential repre- sentation, reversing the order of integration, and employing the Dirac delta function exponential representation (Section 15.2). These cases can be treated equally well as Fouriercosinetransforms. Note.T h eK0relation appears as a consequence of a Green’s function equation in Ex- ercise9.7.14. 15.3.12 A calculation of the magnetic field of a circular current loop in circular cylindrical coordinatesleadstotheintegral integraldisplay∞ 0coskzkK1(ka)dk. Showthatthisintegralisequalto πa 2(z2+a2)3/2. Hint.TrydifferentiatingExercise15.3.11(c). 15.3.13 AsanextensionofExercise15.3.11,showthat (a)integraldisplay∞ 0J0(y)dy=1, (b)integraldisplay∞ 0N0(y)dy=0, (c)integraldisplay∞ 0K0(y)dy=π 2. 15.3.14 The Fourier integral, Eq. (15.18), has been held meaningless for f(t)=cosαt. Show thattheFourierintegralcanbeextendedtocover f(t)=cosαtbyuseoftheDiracdelta function. 15.3 Fourier Transforms — Inversion Theorem 945 15.3.15 Showthat integraldisplay∞ 0sinkaJ0(kρ)dk=braceleftbiggparenleftbig a2−ρ2parenrightbig−1/2,ρ<a, 0,ρ >a. Hereaandρarepositive.Theequationcomesfromthedeterminationofthedistribution of charge on an isolated conducting disk, radius a. Note that the function on the right hasaninfinitediscontinuityat ρ=a. Note.ALaplacetransformapproachappearsinExercise15.10.8. 15.3.16 Thefunction f(r)hasaFourierexponentialtransform, g(k)=1 (2π)3/2integraldisplay f(r)eik·rd3r=1 (2π)3/2k2. Determine f(r). Hint.Usesphericalpolarcoordinatesin k-space. ANS.f(r)=1 4πr. 15.3.17 (a) CalculatetheFourierexponentialtransform of f(x)=e−a|x|. (b) Calculate the inverse transform by employing the calculus of residues (Sec- tion7.1). 15.3.18 ShowthatthefollowingareFouriertransforms ofeachother inJn(t)and  radicalbigg 2 πTn(x)parenleftbig 1−x2parenrightbig−1/2,|x|<1, 0, |x|>1. Tn(x)isthenth-orderChebyshevpolynomial. Hint. WithTn(cosθ)=cosnθ, the transform of Tn(x)(1−x2)−1/2leads to an integral representationof Jn(t). 15.3.19 ShowthattheFourierexponentialtransformof f(µ)=braceleftbiggPn(µ),|µ|≤1, 0,|µ|>1 is(2in/2π)jn(kr).H e r ePn(µ)is a Legendre polynomial and jn(kr)is a spherical Besselfunction. 15.3.20 Show that the three-dimensional Fourier exponential transform of a radially symmetric functionmayberewrittenasaFouriersinetransform: 1 (2π)3/2integraldisplay∞ −∞f(r)eik·rd3x=1 kradicalbigg 2 πintegraldisplay∞ 0bracketleftbig rf(r)bracketrightbig sinkrdr. 946 Chapter 15 Integral Transforms 15.3.21 (a) Show that f(x)=x−1/2is aself-reciprocal under both Fourier cosine and sine transforms;thatis, radicalbigg 2 πintegraldisplay∞ 0x−1/2cosxtdx=t−1/2, radicalbigg 2 πintegraldisplay∞ 0x−1/2sinxtds=t−1/2. (b) Use the preceding results to evaluate the Fresnel integralsintegraltext∞ 0cos(y2)dyandintegraltext∞ 0sin(y2)dy. 15.4 F OURIER TRANSFORM OF DERIVATIVES In Section 15.1, Fig. 15.1 outlines the overall technique of using Fourier transforms and inversetransforms tosolveaproblem.Herewetakeaninitialstepinsolvingadifferential equation—obtainingtheFouriertransformofaderivative. Usingtheexponentialform, wedeterminethattheFouriertransformof f(x)is g(ω)=1√ 2πintegraldisplay∞ −∞f(x)eiωxdx (15.37) andfordf(x)/dx g1(ω)=1√ 2πintegraldisplay∞ −∞df(x) dxeiωxdx. (15.38) IntegratingEq. (15.38)byparts, weobtain g1(ω)=eiωx √ 2πf(x)vextendsinglevextendsinglevextendsingle∞ −∞−iω√ 2πintegraldisplay∞ −∞f(x)eiωxdx. (15.39) Iff(x)vanishes4asx→±∞,weha v e g1(ω)=−iωg(ω); (15.40) thatis,thetransformofthederivativeis (−iω)timesthetransformoftheoriginalfunction. Thismayreadilybegeneralizedtothe nthderivativetoyield gn(ω)=(−iω)ng(ω), (15.41) provided all the integrated parts vanish as x→±∞. This is the power of the Fourier transform,thereasonitissousefulinsolving(partial)differentialequations.Theoperation ofdifferentiationhasbeenreplacedbya multiplicationin ω-space. 4Apart from casessuch asExercise 15.3.6, f(x)must vanish as x→±∞in order for the Fourier transform of f(x)to exist. 15.4 Fourier Transform of Derivatives 947 Example 15.4.1 WAVEEQUATION ThistechniquemaybeusedtoadvantageinhandlingPDEs.Toillustratethetechnique,let usderiveafamiliarexpressionofelementaryphysics.Aninfinitelylongstringisvibrating freely.Theamplitude yofthe(small)vibrationssatisfiesthewaveequation ∂2y ∂x2=1 v2∂2y ∂t2. (15.42) Weshallassumeaninitialcondition y(x,0)=f(x), (15.43) wherefis localized,thatis, approacheszeroatlarge x. Applying our Fourier transform in x, which means multiplying by eiαxand integrating overx,weobtain integraldisplay∞ −∞∂2y(x,t) ∂x2eiαxdx=1 v2integraldisplay∞ −∞∂2y(x,t) ∂t2eiαxdx (15.44) or (−iα)2Y(α,t)=1 v2∂2Y(α,t) ∂t2. (15.45) Herewehaveused Y(α,t)=1√ 2πintegraldisplay∞ −∞y(x,t)eiαxdx (15.46) and Eq. (15.41) for the second derivative. Note that the integrated part of Eq. (15.39) van- ishes: The wave has not yet gone to ±∞because it is propagating forward in time, and there is no source at infinity because f(±∞)=0.Since no derivatives with respect to αappear, Eq. (15.45) is actually an ODE—in fact, the linear oscillator equation. This transformation,fromaPDEtoanODE,isasignificantachievement.WesolveEq.(15.45) subject to the appropriate initial conditions. At t=0, applying Eq. (15.43), Eq. (15.46) reducesto Y(α,0)=1√ 2πintegraldisplay∞ −∞f(x)eiαxdx=F(α). (15.47) Thegeneralsolutionof Eq.(15.45) inexponentialformis Y(α,t)=F(α)e±ivαt. (15.48) Usingtheinversionformula(Eq. (15.23)), wehave y(x,t)=1√ 2πintegraldisplay∞ −∞Y(α,t)e−iαxdα, (15.49) and,byEq.(15.48), y(x,t)=1√ 2πintegraldisplay∞ −∞F(α)e−iα(x∓vt)dα. (15.50) 948 Chapter 15 Integral Transforms Sincef(x)istheFourierinversetransform of F(α), y(x,t)=f(x∓vt), (15.51) correspondingtowavesadvancinginthe +x-and−x-directions,respectively. The particular linear combinations of waves is given by the boundary condition of Eq.(15.43) andsomeotherboundarycondition,suchas arestrictionon ∂y/∂t. /squaresolid TheaccomplishmentoftheFouriertransform heredeservesspecialemphasis. •Our Fourier transform converted a PDE into an ODE, where the “degree of transcen- dence”oftheproblemwasreduced. In Section 15.9 Laplace transforms are used to convert ODEs (with constant coefficients) into algebraic equations. Again, the degree of transcendence is reduced. The problem is simplified—as outlinedinFig.15.1. Example 15.4.2 HEATFLOW PDE To illustrate another transformation of a PDE into an ODE, let us Fourier transform the heatflowpartialdifferentialequation ∂ψ ∂t=a2∂2ψ ∂x2, where the solution ψ(x,t)is the temperature in space as a function of time. By taking the Fourier transform of both sides of this equation (note that here only ωis the transform variableconjugateto xbecausetis thetimeintheheatflowPDE),where /Psi1(ω,t)=1√ 2πintegraldisplay∞ −∞ψ(x,t)eiωxdx, thisyieldsanODEfor theFouriertransform /Psi1ofψinthetimevariable t, ∂/Psi1(ω,t) ∂t=−a2ω2/Psi1(ω,t). Integratingweobtain ln/Psi1=−a2ω2t+lnC,or/Psi1=Ce−a2ω2t, where the integration constant Cmay still depend on ωand, in general, is determined by initial conditions. In fact, C=/Psi1(ω,0)is the initial spatial distribution of /Psi1,so it is given by the transform (in x) of the initial distribution of ψ,namely,ψ(x,0).Putting this solutionbackintoourinverseFouriertransform, thisyields ψ(x,t)=1√ 2πintegraldisplay∞ −∞C(ω)e−iωxe−a2ω2tdω. For simplicity, we here take Cω-independent (assuming a delta-function initial temper- ature distribution) and integrate by completing the square in ω,as in Example 15.1.1, 15.4 Fourier Transform of Derivatives 949 making appropriate changes of variables and parameters ( a2→a2t, ω→x,t→−ω). ThisyieldstheparticularsolutionoftheheatflowPDE, ψ(x,t)=C a√ 2texpparenleftbigg −x2 4a2tparenrightbigg , whichappearsasacleverguessinChapter8.Ineffect,wehaveshownthat ψistheinverse Fouriertransform of Cexp(−a2ω2t). /squaresolid Example 15.4.3 INVERSION OF PDE DeriveaFourierintegralfortheGreen’sfunction G0ofPoisson’sPDE,whichisasolution of ∇2G0(r,r′)=−δ(r−r′). OnceG0is known,thegeneralsolutionofPoisson’sPDE, ∇2/Phi1=−4πρ(r) ofelectrostatics,isgivenas /Phi1(r)=integraldisplay G0(r,r′)4πρ(r′)d3r′. Applying ∇2to/Phi1andusingthePDEtheGreen’sfunctionsatisfies,wecheckthat ∇2/Phi1(r)=integraldisplay ∇2G0(r,r′)4πρ(r′)d3r′=−integraldisplay δ(r−r′)4πρ(r′)d3r′=−4πρ(r). NowweusetheFouriertransformof G0,whichisg0,andofthatofthe δfunction,writing ∇2integraldisplay g0(p)eip·(r−r′)d3p (2π)3=−integraldisplay eip·(r−r′)d3p (2π)3. Because the integrands of equal Fourier integrals must be the same (almost) everywhere, whichfollowsfromtheinverseFouriertransform,andwith ∇eip·(r−r′)=ipeip·(r−r′), this yields−p2g0(p)=−1.Therefore, application of the Laplacian to a Fourier integral f(r)correspondstomultiplyingitsFouriertransform g(p)by−p2.Substitutingthissolu- tionintotheinverseFouriertransform for G0gives G0(r,r′)=integraldisplay eip·(r−r′)d3p (2π)3p2=1 4π|r−r′|. We can verify the last part of this result by applying ∇2toG0again and recalling from Chapter1that ∇21 |r−r′|=−4πδ(r−r′). 950 Chapter 15 Integral Transforms The inverse Fourier transform can be evaluated using polar coordinates, exploiting the sphericalsymmetryof p2.Forsimplicity,wewrite R=r−r′andcallθtheanglebetween Randp, integraldisplay eip·Rd3p p2=integraldisplay∞ 0dpintegraldisplay1 −1eipRcosθdcosθintegraldisplay2π 0dϕ =2π iRintegraldisplay∞ 0dp peipRcosθvextendsinglevextendsinglevextendsingle1 cosθ=−1=4π Rintegraldisplay∞ 0sinpR pdp =4π Rintegraldisplay∞ 0sinpR pRd(pR)=2π2 R, whereθandϕare the angles of pandintegraltext∞ 0sinx xdx=π 2, from Example 7.1.4. Dividing by (2π)3,we obtain G0(R)=1/(4πR),as claimed. An evaluation of this Fourier transform bycontourintegrationis giveninExample9.7.2. /squaresolid Exercises 15.4.1 Theone-dimensionalFermiageequationfor thediffusionofneutronsslowingdownin somemedium(suchas graphite)is ∂2q(x,τ) ∂x2=∂q(x,τ) ∂τ. Hereqis the number of neutrons that slow down, falling below some given energy per secondperunitvolume.TheFermiage, τ, isameasureoftheenergyloss. Ifq(x,0)=Sδ(x), corresponding to a plane source of neutrons at x=0, emitting S neutronsper unitareapersecond,derivethesolution q=Se−x2/4τ √ 4πτ. Hint.Replace q(x,τ)with p(k,τ)=1√ 2πintegraldisplay∞ −∞q(x,τ)eikxdx. Thisis analogoustothediffusionof heatinaninfinitemedium. 15.4.2 Equation(15.41)yields g2(ω)=−ω2g(ω) fortheFouriertransformofthesecondderivativeof f(x).Thecondition f(x)→0f o r x→±∞may be relaxed slightly. Find the least restrictive condition for the preceding equationfor g2(ω)tohold. ANS.bracketleftbiggdf(x) dx−iωf(x)bracketrightbigg eiωxvextendsinglevextendsinglevextendsingle∞ −∞=0. 15.5 Convolution Theorem 951 15.4.3 Theone-dimensionalneutrondiffusionequationwitha(plane)sourceis −Dd2ϕ(x) dx2+K2Dϕ(x)=Qδ(x), whereϕ(x)istheneutronflux, Qδ(x)isthe(plane)sourceat x=0,andDandK2are constants.ApplyaFouriertransform.Solvetheequationintransformspace.Transform yoursolutionbackinto x-space. ANS.ϕ(x)=Q 2KDe−|Kx|. 15.4.4 For a point source at the origin, the three-dimensional neutron diffusion equation be- comes −D∇2ϕ(r)+K2Dϕ(r)=Qδ(r). Apply a three-dimensional Fourier transform. Solve the transformed equation. Trans- formthesolutionbackinto r-space. 15.4.5 (a) Given that F(k)is the three-dimensional Fourier transform of f(r)andF1(k)is thethree-dimensionalFouriertransformof ∇f(r), showthat F1(k)=(−ik)F(k). Thisis athree-dimensionalgeneralizationofEq. (15.40). (b) Showthatthethree-dimensionalFouriertransformof ∇·∇f(r)is F2(k)=(−ik)2F(k). Note. Vectorkis a vector in the transform space. In Section 15.6 we shall have ¯hk=p,linearmomentum. 15.5 C ONVOLUTION THEOREM We shall employ convolutions to solve differential equations, to normalize momentum wavefunctions(Section15.6),andtoinvestigatetransfer functions(Section15.7). Let us consider two functions f(x)andg(x)with Fourier transforms F(t)andG(t), respectively.Wedefinetheoperation f∗g≡1√ 2πintegraldisplay∞ −∞g(y)f(x−y)dy (15.52) as theconvolution of the two functions fandgover the interval (−∞,∞). This form of an integral appears in probability theory in the determination of the probability density of two random, independent variables. Our solution of Poisson’s equation, Eq. (9.148), may be interpreted as a convolution of a charge distribution, ρ(r2), and a weighting function, (4πε0|r1−r2|)−1. In other works this is sometimes referred to as the Faltung,t ou s et h e 952 Chapter 15 Integral Transforms FIGURE 15.5 German term for “folding.”5We now transform the integral in Eq. (15.52) by introducing theFouriertransforms: integraldisplay∞ −∞g(y)f(x−y)dy=1√ 2πintegraldisplay∞ −∞g(y)integraldisplay∞ −∞F(t)e−it(x−y)dtdy =1√ 2πintegraldisplay∞ −∞F(t)bracketleftbiggintegraldisplay∞ −∞g(y)eitydybracketrightbigg e−itxdt =integraldisplay∞ −∞F(t)G(t)e−itxdt, (15.53) interchanging the order of integration and transforming g(y). This result may be inter- preted as follows: The Fourier inverse transform of a productof Fourier transforms is the convolutionof theoriginalfunctions, f∗g. Forthespecialcase x=0w eh a v e integraldisplay∞ −∞F(t)G(t)dt=integraldisplay∞ −∞f(−y)g(y)dy. (15.54) Theminussignin −ysuggeststhatmodificationsbetried.Wenowdothiswith g∗instead ofgusingadifferent technique. Parseval’s Relation Results analogous to Eqs. (15.53) and (15.54) may be derived for the Fourier sine and co- sinetransforms(Exercises15.5.1and15.5.3).Equation(15.54)andthecorrespondingsine and cosine convolutions are often labeled Parseval’s relations by analogy with Parseval’s theoremfor Fourierseries (Chapter14,Exercise14.4.2). 5Forf(y)=e−y,f(y)andf(x−y)are plotted in Fig. 15.5. Clearly, f(y)andf(x−y)are mirror images of each other in relationto thevertical line y=x/2,that is,wecould generate f(x−y)by folding over f(y)on the line y=x/2. 15.5 Convolution Theorem 953 TheParsevalrelation6,7 integraldisplay∞ −∞F(ω)G∗(ω)dω=integraldisplay∞ −∞f(t)g∗(t)dt (15.55) may be derived elegantly using the Dirac delta function representation, Eq. (15.21d). We have integraldisplay∞ −∞f(t)g∗(t)dt=integraldisplay∞ −∞1√ 2πintegraldisplay∞ −∞F(ω)e−iωtdω·1√ 2πintegraldisplay∞ −∞G∗(x)eixtdxdt, (15.56) withattentiontothecomplexconjugationinthe G∗(x)tog∗(t)transform.Integratingover tfirst, andusingEq.(15.21d), weobtain integraldisplay∞ −∞f(t)g∗(t)dt=integraldisplay∞ −∞F(ω)integraldisplay∞ −∞G∗(x)δ(x−ω)dxdω =integraldisplay∞ −∞F(ω)G∗(ω)dω, (15.57) our desired Parseval relation. If f(t)=g(t), then the integrals in the Parseval relation are normalization integrals (Section 10.4). Equation (15.57) guarantees that if a function f(t) isnormalizedtounity,itstransform F(ω)islikewisenormalizedtounity.Thisisextremely importantinquantummechanicsasdevelopedinthenextsection. ItmaybeshownthattheFouriertransformisaunitaryoperation(intheHilbertspace L2, squareintegrablefunctions).TheParsevalrelationisareflectionofthisunitaryproperty— analogoustoExercise3.4.26for matrices. In Fraunhofer diffraction optics the diffraction pattern (amplitude) appears as the trans- form of the function describing the aperture (compare Exercise 15.5.5). With intensity proportional to the square of the amplitude the Parseval relation implies that the energy passing through the aperture seems to be somewhere in the diffraction pattern—a state- ment of the conservation of energy. Parseval’s relations may be developed independently of the inverse Fourier transform and then used rigorously to derive the inverse transform. DetailsaregivenbyMorse andFeshbach,8Section4.8(see alsoExercise15.5.4). Exercises 15.5.1 WorkouttheconvolutionequationcorrespondingtoEq.(15.53) for (a) Fouriersinetransforms 1 2integraldisplay∞ 0g(y)bracketleftbig f(y+x)+f(y−x)bracketrightbig dy=integraldisplay∞ 0Fs(s)Gs(s)cossxds, wherefandgare oddfunctions. 6Notethat all arguments arepositive, in contrast to Eq.(15.54). 7Some authors prefer to restrict Parseval’s name to series and refer to Eq. (15.55) as Rayleigh’s theorem . 8P.M.Morse andH.Feshbach, Methods of Theoretical Physics ,NewYork: McGraw-Hill(1953). 954 Chapter 15 Integral Transforms (b) Fouriercosinetransforms 1 2integraldisplay∞ 0g(y)bracketleftbig f(y+x)+f(x−y)bracketrightbig dy=integraldisplay∞ 0Fc(s)Gc(s)cossxds, wherefandgare evenfunctions. 15.5.2 F(ρ)andG(ρ)are the Hankel transforms of f(r)andg(r), respectively (Exer- cise15.1.1). DerivetheHankeltransformParsevalrelation: integraldisplay∞ 0F∗(ρ)G(ρ)ρdρ=integraldisplay∞ 0f∗(r)g(r)rdr. 15.5.3 Show that for both Fourier sine and Fourier cosine transforms Parseval’s relation has theform integraldisplay∞ 0F(t)G(t)dt=integraldisplay∞ 0f(y)g(y)dy. 15.5.4 Starting from Parseval’s relation (Eq. (15.54)), let g(y)=1, 0≤y≤α, and zero else- where.FromthisderivetheFourierinversetransform (Eq. (15.23)). Hint.Differentiatewithrespectto α. 15.5.5 (a) Arectangularpulseis describedby f(x)=braceleftbigg1,|x|<a, 0,|x|>a. ShowthattheFourierexponentialtransform is F(t)=radicalbigg 2 πsinat t. This is the single-slit diffraction problem of physical optics. The slit is described byf(x).Thediffractionpattern amplitude isgivenbytheFouriertransform F(t). (b) UsetheParsevalrelationtoevaluate integraldisplay∞ −∞sin2t t2dt. This integral may also be evaluated by using the calculus of residues, Exer- cise7.1.12. ANS.(b)π. 15.5.6 Solve Poisson’s equation, ∇2ψ(r)=−ρ(r)/ε0, by the following sequence of opera- tions: (a) Take the Fourier transform of both sides of this equation. Solve for the Fourier transform of ψ(r). (b) CarryouttheFourierinversetransformbyusingathree-dimensionalanalogofthe convolutiontheorem,Eq. (15.53). 15.6 Momentum Representation 955 15.5.7 (a) Given f(x)=1−|x/2|,−2≤x≤2, and zero elsewhere, show that the Fourier transform of f(x)is F(t)=radicalbigg 2 πparenleftbiggsint tparenrightbigg2 . (b) UsingtheParsevalrelation,evaluate integraldisplay∞ −∞parenleftbiggsint tparenrightbigg4 dt. ANS.(b)2π 3. 15.5.8 WithF(t)andG(t)theFouriertransforms of f(x)andg(x), respectively,showthat integraldisplay∞ −∞vextendsinglevextendsinglef(x)−g(x)vextendsinglevextendsingle2dx=integraldisplay∞ −∞vextendsinglevextendsingleF(t)−G(t)vextendsinglevextendsingle2dt. Ifg(x)is an approximation to f(x), the preceding relation indicates that the mean squaredeviationin t-spaceis equaltothemeansquaredeviationin x-space. 15.5.9 UsetheParsevalrelationtoevaluate (a)integraldisplay∞ −∞dω (ω2+a2)2,(b)integraldisplay∞ −∞ω2dω (ω2+a2)2. Hint.CompareExercise15.3.4. ANS.(a)π 2a3,( b)π 2a. 15.6 M OMENTUM REPRESENTATION In advanced dynamics and in quantum mechanics, linear momentum and spatial position occur on an equal footing. In this section we shall start with the usual space distribution and derive the corresponding momentum distribution. For the one-dimensional case our wavefunction ψ(x)hasthefollowingproperties: 1.ψ∗(x)ψ(x)dx is the probability density of finding a quantum particle between xand x+dx,and 2.integraldisplay∞ −∞ψ∗(x)ψ(x)dx=1 (15.58) correspondstoprobabilityunity. 3. In addition,wehave /angbracketleftx/angbracketright=integraldisplay∞ −∞ψ∗(x)xψ(x)dx (15.59) fortheaveragepositionoftheparticlealongthe x-axis.Thisisoftencalledan expec- tationvalue . Wewantafunction g(p)thatwillgivethesameinformationaboutthemomentum: 956 Chapter 15 Integral Transforms 1.g∗(p)g(p)dp is the probability density that our quantum particle has a momentum betweenpandp+dp. 2.integraldisplay∞ −∞g∗(p)g(p)dp=1. (15.60) 3. /angbracketleftp/angbracketright=integraldisplay∞ −∞g∗(p)pg(p)dp. (15.61) As subsequently shown, such a function is given by the Fourier transform of our space functionψ(x). Specifically,9 g(p)=1√2π¯hintegraldisplay∞ −∞ψ(x)e−ipx/¯hdx, (15.62) g∗(p)=1√2π¯hintegraldisplay∞ −∞ψ∗(x)eipx/¯hdx. (15.63) Thecorrespondingthree-dimensionalmomentumfunctionis g(p)=1 (2π¯h)3/2integraldisplay∞integraldisplay −∞integraldisplay ψ(r)e−ir·p/¯hd3r. ToverifyEqs.(15.62) and(15.63), letuscheckonproperties2and3. Property 2, the normalization, is automatically satisfied as a Parseval relation, Eq. (15.55). If the space function ψ(x)is normalized to unity, the momentum function g(p)is alsonormalizedtounity. Tocheckonproperty3,wemustshowthat /angbracketleftp/angbracketright=integraldisplay∞ −∞g∗(p)pg(p)dp=integraldisplay∞ −∞ψ∗(x)¯h id dxψ(x)dx, (15.64) where(¯h/i)(d/dx) is the momentum operator in the space representation. We replace the momentum functions by Fourier-transformed space functions, and the first integral becomes 1 2π¯hintegraldisplay∞integraldisplay −∞integraldisplay pe−ip(x−x′)/¯hψ∗(x′)ψ(x)dpdx′dx. (15.65) Nowweusetheplane-waveidentity pe−ip(x−x′)/¯h=d dxbracketleftbigg −¯h ie−ip(x−x′)/¯hbracketrightbigg , (15.66) 9The¯hmay beavoided by using the wavenumber k,p=k¯h(andp=k¯h), so ϕ(k)=1 (2π)1/2integraldisplay ψ(x)e−ikxdx. An example of this notation appears in Section16.1. 15.6 Momentum Representation 957 withpa constant, not an operator. Substituting into Eq. (15.65) and integrating by parts, holdingx′andpconstant,weobtain /angbracketleftp/angbracketright=integraldisplay∞integraldisplay −∞bracketleftbigg1 2π¯hintegraldisplay∞ −∞e−ip(x−x′)/¯hdpbracketrightbigg ·ψ∗(x′)¯h id dxψ(x)dx′dx. (15.67) Here we assume ψ(x)vanishes as x→±∞, eliminating the integrated part. Using the Dirac delta function, Eq. (15.21c), Eq. (15.67) reduces to Eq. (15.64) to verify our mo- mentumrepresentation. Alternatively,if theintegrationover pis donefirst inEq.(15.65), leadingto integraldisplay∞ −∞pe−ip(x−x′)/¯hdp=2πi¯h2δ′(x−x′), andusingExercise1.15.9,wecandotheintegrationover x,whichcauses ψ(x)tobecome −dψ(x′)/dx′.Theremainingintegralover x′istheright-handsideofEq. (15.64). Example 15.6.1 HYDROGEN ATOM Thehydrogenatomgroundstate10maybedescribedbythespatialwavefunction ψ(r)=parenleftbigg1 πa3 0parenrightbigg1/2 e−r/a0, (15.68) a0being the Bohr radius, 4 πε0¯h2/me2. We now have a three-dimensional wave function. Thetransform correspondingtoEq.(15.62) is g(p)=1 (2π¯h)3/2integraldisplay ψ(r)e−ip·r/¯hd3r. (15.69) SubstitutingEq. (15.68)intoEq.(15.69) andusing integraldisplay e−ar+ib·rd3r=8πa (a2+b2)2, (15.70) weobtainthehydrogenicmomentumwavefunction, g(p)=23/2 πa3/2 0¯h5/2 (a2 0p2+¯h2)2. (15.71) Such momentum functions have been found useful in problems like Compton scattering fromatomicelectrons,thewavelengthdistributionofthescatteredradiation,dependingon themomentumdistributionof thetargetelectrons. The relation between the ordinary space representation and the momentum representa- tion may be clarified by considering the basic commutation relations of quantum mechan- ics.WegofromaclassicalHamiltoniantotheSchrödingerwaveequationbyrequiringthat momentum pandposition xnotcommute.Instead,werequirethat [p,x]≡px−xp=−i¯h. (15.72) 10See E. V. Ivash, A momentum representation treatment of the hydrogen atom problem. A m .J .P h y s . 40: 1095 (1972) for amomentum representation treatment of the hydrogen atom l=0 states. 958 Chapter 15 Integral Transforms Forthemultidimensionalcase,Eq.(15.72) isreplacedby [pi,xj]=−i¯hδij. (15.73) TheSchrödinger(space)representationis obtainedbyusing x→x:pi→−i¯h∂ ∂xi, replacingthemomentumbyapartialspacederivative.Wesee that [p,x]ψ(x)=−i¯hψ(x). (15.74) However,Eq. (15.72)canequallywellbesatisfiedbyusing p→p:xj→i¯h∂ ∂pj. Thisis themomentumrepresentation.Then [p,x]g(p)=−i¯hg(p). (15.75) Hencetherepresentation (x)is notunique; (p)is analternatepossibility. Ingeneral,theSchrödingerrepresentation (x)leadingtotheSchrödingerwaveequation is more convenient because the potential energy Vis generally given as a function of positionV(x,y,z) .Themomentumrepresentation (p)usuallyleadstoanintegralequation (compare Chapter 16 for the pros and cons of the integral equations). For an exception, considertheharmonicoscillator. /squaresolid Example 15.6.2 HARMONIC OSCILLATOR TheclassicalHamiltonian(kineticenergy +potentialenergy =totalenergy)is H(p,x)=p2 2m+1 2kx2=E, (15.76) wherekistheHooke’slawconstant. In theSchrödingerrepresentationweobtain −¯h2 2md2ψ(x) dx2+1 2kx2ψ(x)=Eψ(x). (15.77) Fortotalenergy Eequalto√(k/m)¯h/2 thereisanunnormalizedsolution(Section13.1), ψ(x)=e−(√ mk/2¯h)x2. (15.78) Themomentumrepresentationleadsto p2 2mg(p)−¯h2k 2d2g(p) dp2=Eg(p). (15.79) Again,for E=radicalbigg k m¯h 2(15.80) 15.6 Momentum Representation 959 themomentumwaveequation(15.79) issatisfiedbytheunnormalized g(p)=e−p2/(2¯h√ mk). (15.81) Either representation, space or momentum (and an infinite number of other possibilities), may be used, depending on which is more convenient for the particular problem under attack. The demonstration that g(p)is the momentum wave function corresponding to Eq. (15.78)—that it is the Fourier inverse transform of Eq. (15.78)—is left as Exer- cise15.6.3. /squaresolid Exercises 15.6.1 The function eik·rdescribes a plane wave of momentum p=¯hknormalized to unit density. (Time dependenceof e−iωtis assumed.) Show that these plane-wavefunctions satisfyanorthogonalityrelation integraldisplayparenleftbig eik·rparenrightbig∗eik′·rdxdydz=(2π)3δ(k−k′). 15.6.2 Aninfiniteplanewaveinquantummechanicsmayberepresentedbythefunction ψ(x)=eip′x/¯h. Findthecorrespondingmomentumdistributionfunction.Notethatithasaninfinityand thatψ(x)is notnormalized. 15.6.3 Alinearquantumoscillatorinits groundstatehasawavefunction ψ(x)=a−1/2π−1/4e−x2/2a2. Showthatthecorrespondingmomentumfunctionis g(p)=a1/2π−1/4¯h−1/2e−a2p2/2¯h2. 15.6.4 Thenthexcitedstateofthelinearquantumoscillatorisdescribedby ψn(x)=a−1/22−n/2π−1/4(n!)−1/2e−x2/2a2Hn(x/a), whereHn(x/a)is thenth Hermite polynomial, Section 13.1. As an extension of Exer- cise15.6.3,findthemomentumfunctioncorrespondingto ψn(x). Hint.ψn(x)may be represented by (ˆa†)nψ0(x), whereˆa†is the raising operator, Exer- cise13.1.14to13.1.16. 15.6.5 Afreeparticleinquantummechanicsisdescribedbyaplanewave ψk(x,t)=ei[kx−(¯hk2/2m)t]. Combiningwaves of adjacent momentumwith an amplitudeweightingfactor ϕ(k),w e formawavepacket /Psi1(x,t)=integraldisplay∞ −∞ϕ(k)ei[kx−(¯hk2/2m)t]dk. 960 Chapter 15 Integral Transforms (a) Solvefor ϕ(k)giventhat /Psi1(x,0)=e−x2/2a2. (b) Using the known value of ϕ(k), integrate to get the explicit form of /Psi1(x,t).N o t e thatthiswavepacketdiffuses orspreads outwithtime. ANS./Psi1(x,t)=e−{x2/2[a2+(i¯h/m)t]} [1+(i¯ht/ma2)]1/2. Note. An interesting discussion of this problem from the evolution operator point of view is given by S. M. Blinder, Evolution of a Gaussian wave packet, Am. J. Phys. 36: 525(1968). 15.6.6 Find the time-dependent momentum wave function g(k,t)corresponding to /Psi1(x,t)of Exercise 15.6.5. Show that the momentum wave packet g∗(k,t)g(k,t) isindependent oftime. 15.6.7 The deuteron, Example 10.1.2, may be described reasonably well with a Hulthén wave function ψ(r)=A[e−αr−e−βr]/r, withA,α,andβconstants.Find g(p), thecorrespondingmomentumfunction. Note. The Fourier transform may be rewritten as Fourier sine and cosine transforms or asaLaplacetransform,Section15.8. 15.6.8 The nuclear form factor F(k)and the charge distribution ρ(r)are three-dimensional Fouriertransforms ofeachother: F(k)=1 (2π)3/2integraldisplay ρ(r)eik·rd3r. If themeasuredform factor is F(k)=(2π)−3/2parenleftbigg 1+k2 a2parenrightbigg−1 , findthecorrespondingchargedistribution. ANS.ρ(r)=a2 4πe−ar r. 15.6.9 Checkthenormalizationofthehydrogenmomentumwavefunction g(p)=23/2 πa3/2 0¯h5/2 (a2 0p2+¯h2)2 bydirectevaluationoftheintegral integraldisplay g∗(p)g(p)d3p. 15.6.10 Withψ(r)a wave function in ordinary space and ϕ(p)the corresponding momentum function,showthat 15.7 Transfer Functions 961 (a)1 (2π¯h)3/2integraldisplay rψ(r)e−ir·p/¯hd3r=i¯h∇pϕ(p), (b)1 (2π¯h)3/2integraldisplay r2ψ(r)e−r·p/¯hd3r=(i¯h∇p)2ϕ(p). Note.∇pis thegradientinmomentumspace: ˆx∂ ∂px+ˆy∂ ∂py+ˆz∂ ∂pz. These results may be extended to any positive integer power of rand therefore to any (analytic)functionthatmaybeexpandedas aMaclaurinseriesin r. 15.6.11 The ordinary space wave function ψ(r,t)satisfies the time-dependent Schrödinger equation i¯h∂ψ(r,t) ∂t=−¯h2 2m∇2ψ+V(r)ψ. Show that the corresponding time-dependent momentum wave function satisfies the analogousequation, i¯h∂ϕ(p,t) ∂t=p2 2mϕ+V(i¯h∇p)ϕ. Note. Assume that V(r)may be expressed by a Maclaurin series and use Exer- cise 15.6.10. V(i¯h∇p)is the same function of the variable i¯h∇pthatV(r)is of the variabler. 15.6.12 Theone-dimensionaltime-independentSchrödingerwaveequationis −¯h2 2md2ψ(x) dx2+V(x)ψ(x)=Eψ(x). For the special case of V(x)an analytic function of x, show that the corresponding momentumwaveequationis Vparenleftbigg i¯hd dpparenrightbigg g(p)+p2 2mg(p)=Eg(p). Derive this momentum wave equation from the Fourier transform, Eq. (15.62), and its inverse.Donotusethesubstitution x→i¯h(d/dp)directly. 15.7 T RANSFER FUNCTIONS A time-dependent electrical pulse may be regarded as built-up as a superposition of plane wavesof manyfrequencies.Forangularfrequency ωwehaveacontribution F(ω)eiωt. Thenthecompletepulsemaybewrittenas f(t)=1 2πintegraldisplay∞ −∞F(ω)eiωtdω. (15.82) 962 Chapter 15 Integral Transforms FIGURE 15.6Servomechanismorastereoamplifier. Becausetheangularfrequency ωis relatedtothelinearfrequency νby ν=ω 2π, itiscustomarytoassociatetheentire 1 /2πfactorwiththisintegral. But ifωis a frequency, what about the negative frequencies? The negative ωmay be lookedonasamathematicaldevicetoavoiddealingwithtwofunctions(cos ωtandsinωt) separately(compareSection14.1). Because Eq. (15.82) has the form of a Fourier transform, we may solve for F(ω)by writingtheinversetransform, F(ω)=integraldisplay∞ −∞f(t)e−iωtdt. (15.83) Equation(15.83)representsa resolutionofthepulse f(t)intoitsangularfrequencycom- ponents.Equation(15.82) isa synthesis ofthepulse fromits components. Consider some device, such as a servomechanism or a stereo amplifier (Fig. 15.6), with an inputf(t)and an output g(t). For an input of a single frequency ω,fω(t)=eiωt,t h e amplifierwill alter the amplitudeand may also changethe phase. The changeswill proba- blydependonthefrequency.Hence gω(t)=ϕ(ω)fω(t). (15.84) This amplitudes-and phase-modifyingfunction ϕ(ω)is calleda transferfunction. It usu- allywillbecomplex: ϕ(ω)=u(ω)+iv(ω), (15.85) wherethefunctions u(ω)andv(ω)arereal. In Eq. (15.84) we assume that the transfer function ϕ(ω)is independentof input ampli- tude and of the presence or absence of any other frequency components. That is, we are assuming a linear mapping of f(t)ontog(t). Then the total output may be obtained by integratingovertheentireinput,as modifiedbytheamplifier g(t)=1 2πintegraldisplay∞ −∞ϕ(ω)F(ω)eiωtdω. (15.86) The transfer function is characteristic of the amplifier. Once the transfer function is known (measured or calculated), the output g(t)can be calculated for any input f(t). Letusconsider ϕ(ω)as theFourier(inverse)transformof somefunction /Phi1(t): ϕ(ω)=integraldisplay∞ −∞/Phi1(t)e−iωtdt. (15.87) 15.7 Transfer Functions 963 ThenEq.(15.86)istheFouriertransformoftwoinversetransforms.FromSection15.5we obtaintheconvolution g(t)=integraldisplay∞ −∞f(τ)/Phi1(t−τ)dτ. (15.88) Interpreting Eq. (15.88), we have an input—a “cause”— f(τ), modified by /Phi1(t−τ), producing an output—an “effect”— g(t). Adopting the concept of causality —that the causeprecedestheeffect—wemustrequire τ<t. Wedothisbyrequiring /Phi1(t−τ)=0,τ>t. (15.89) ThenEq. (15.88)becomes g(t)=integraldisplayt −∞f(τ)/Phi1(t−τ)dτ. (15.90) TheadoptionofEq.(15.89)hasprofoundconsequenceshereandequivalentlyindisper- siontheory,Section7.2. Significance of /Phi1(t) Toseethesignificanceof /Phi1,l etf(τ)bea suddenimpulsestartingat τ=0, f(τ)=δ(τ), whereδ(τ)isaDiracdeltadistributiononthepositivesideoftheorigin.ThenEq.(15.90) becomes g(t)=integraldisplayt −∞δ(τ)/Phi1(t−τ)dτ, (15.91) g(t)=braceleftbigg/Phi1(t), t > 0, 0,t <0. This identifies /Phi1(t)as the output function correspondingto a unit impulse at t=0. Equa- tion (15.91) also serves to establish that /Phi1(t)is real. Our original transfer function gives thesteady-stateoutputcorrespondingtoaunit-amplitudesingle-frequencyinput. /Phi1(t)and ϕ(ω)areFouriertransforms ofeachother. FromEq. (15.87)wenowhave ϕ(ω)=integraldisplay∞ 0/Phi1(t)e−iωtdt, (15.92) with the lower limit set equal to zero by causality (Eq. (15.89)). With /Phi1(t)real from Eq.(15.91) weseparaterealandimaginaryparts andwrite u(ω)=integraldisplay∞ 0/Phi1(t)cosωtdt, (15.93) v(ω)=−integraldisplay∞ 0/Phi1(t)sinωtdt, ω> 0. 964 Chapter 15 Integral Transforms From this we see that the real part of ϕ(ω),u(ω) , is even, whereas the imaginary part of ϕ(ω),v(ω) ,is odd: u(−ω)=u(ω), v(−ω)=−v(ω). ComparethisresultwithExercise15.3.1. InterpretingEq.(15.93)as Fouriercosineandsinetransforms, wehave /Phi1(t)=2 πintegraldisplay∞ 0u(ω)cosωtdω =−2 πintegraldisplay∞ 0v(ω)sinωtdω, t > 0. (15.94) CombiningEqs. (15.93) and(15.94), weobtain v(ω)=−integraldisplay∞ 0sinωtbraceleftbigg2 πintegraldisplay∞ 0u(ω′)cosω′tdω′bracerightbigg dt, (15.95) showingthatifourtransferfunctionhasarealpart,itwillalsohaveanimaginarypart(and viceversa).Ofcourse,thisassumesthattheFouriertransformsexist,thusexcludingcases suchas/Phi1(t)=1. The imposition of causality has led to a mutual interdependence of the real and imagi- nary parts of the transfer function. The reader should compare this with the results of the dispersiontheoryofSection7.2, alsoinvolvingcausality. It may be helpful to show that the parity properties of u(ω)andv(ω)require/Phi1(t)to vanishfor negative t. InvertingEq. (15.87),wehave /Phi1(t)=1 2πintegraldisplay∞ −∞bracketleftbig u(ω)+iv(ω)bracketrightbigbracketleftbig cosωt+isinωtbracketrightbig dω. (15.96) Withu(ω)evenand v(ω)odd,Eq. (15.96)becomes /Phi1(t)=1 πintegraldisplay∞ 0u(ω)cosωtdω−1 πintegraldisplay∞ 0v(ω)sinωtdω. (15.97) FromEq. (15.94), integraldisplay∞ 0u(ω)cosωtdω=−integraldisplay∞ 0v(ω)sinωtdω, t > 0. (15.98) If wereverse thesignof t,sinωtreverses signand,fromEq. (15.97), /Phi1(t)=0,t<0 (demonstratingtheinternalconsistencyofouranalysis). Exercise 15.7.1 Derivetheconvolution g(t)=integraldisplay∞ −∞f(τ)/Phi1(t−τ)dτ. 15.8 Laplace Transforms 965 15.8 L APLACE TRANSFORMS Definition TheLaplacetransform f(s)orLof afunction F(t)is definedby11 f(s)=Lbraceleftbig F(t)bracerightbig =lima→∞integraldisplaya 0e−stF(t)dt=integraldisplay∞ 0e−stF(t)dt. (15.99) Afewcommentsontheexistenceoftheintegralareinorder.Theinfiniteintegralof F(t), integraldisplay∞ 0F(t)dt, neednotexist .Forinstance, F(t)maydivergeexponentiallyforlarge t.However,ifthere issomeconstant s0suchthat vextendsinglevextendsinglee−s0tF(t)vextendsinglevextendsingle≤M, (15.100) a positive constant for sufficiently large t,t>t0, the Laplace transform (Eq. (15.99)) will exist fors>s0;F(t)is said to be of exponential order . As a counterexample, F(t)=et2 doesnotsatisfytheconditiongivenbyEq.(15.100)andis notofexponentialorder. L{et2} doesnotexist. The Laplace transform may also fail to exist because of a sufficiently strong singularity inthefunction F(t)ast→0;thatis, integraldisplay∞ 0e−sttndt divergesattheoriginfor n≤−1.TheLaplacetransform L{tn}doesnotexistfor n≤−1. Since,fortwo functions F(t)andG(t), for whichtheintegralsexist Lbraceleftbig aF(t)+bG(t)bracerightbig =aLbraceleftbig F(t)bracerightbig +bLbraceleftbig G(t)bracerightbig , (15.101) theoperationdenotedby Lislinear. Elementary Functions To introduce the Laplace transform, let us apply the operation to some of the elementary functions.In allcasesweassumethat F(t)=0f o rt<0.If F(t)=1,t>0, 11Thisissometimescalleda one-sidedLaplacetransform ;theintegralfrom −∞to+∞isreferredtoasa two-sidedLaplace transform . Some authors introduce an additional factor of s. This extra sappears to have little advantage and continually gets intheway(compareJeffreysandJeffreys,Section14.13—seetheAdditionalReadings—foradditionalcomments).Generally, wetakesto bereal and positive. It is possible to have scomplex, provided ℜ(s)>0. 966 Chapter 15 Integral Transforms then L{1}=integraldisplay∞ 0e−stdt=1 s,fors>0. (15.102) Again,let F(t)=ekt,t>0. TheLaplacetransformbecomes Lbraceleftbig ektbracerightbig =integraldisplay∞ 0e−stektdt=1 s−k,fors>k. (15.103) Usingthisrelation,weobtaintheLaplacetransformofcertainotherfunctions.Since coshkt=1 2parenleftbig ekt+e−ktparenrightbig ,sinhkt=1 2parenleftbig ekt−e−ktparenrightbig , (15.104) wehave L{coshkt}=1 2parenleftbigg1 s−k+1 s+kparenrightbigg =s s2−k2, (15.105) L{sinhkt}=1 2parenleftbigg1 s−k−1 s+kparenrightbigg =k s2−k2, bothvalidfor s>k.Wehavetherelations coskt=coshikt,sinkt=−isinhikt. (15.106) UsingEqs. (15.105)with kreplacedby ik,wefindthattheLaplacetransforms are L{coskt}=s s2+k2, (15.107) L{sinkt}=k s2+k2, both valid for s>0. Another derivation of this last transform is given in the next sec- tion. Note that lim s→0L{sinkt}=1/k. The Laplace transform assigns a value of 1 /ktointegraltext∞ 0sinktdt. Finally,for F(t)=tn,weha v e Lbraceleftbig tnbracerightbig =integraldisplay∞ 0e−sttndt, whichisjustthefactorialfunction.Hence Lbraceleftbig tnbracerightbig =n! sn+1,s>0,n>−1. (15.108) Note that in all these transforms we have the variable sin the denominator—negative powers of s. In particular, lim s→∞f(s)=0. The significance of this point is that if f(s) involvespositivepowersof s(lims→∞f(s)→∞), thennoinversetransformexists. 15.8 Laplace Transforms 967 Inverse Transform Thereislittleimportancetotheseoperationsunlesswecancarryouttheinversetransform, asinFouriertransforms. Thatis, with Lbraceleftbig F(t)bracerightbig =f(s), then L−1braceleftbig f(s)bracerightbig =F(t). (15.109) This inverse transform is notunique. Two functions F1(t)andF2(t)may have the same transform, f(s). However,inthis case F1(t)−F2(t)=N(t), whereN(t)is anullfunction(Fig.15.7), indicatingthat integraldisplayt0 0N(t)dt=0, for all positive t0. This result is known as Lerch’s theorem . Therefore to the physicist and engineer N(t)may almost always be taken as zero and the inverse operation becomes unique. The inverse transform can be determined in various ways. (1) A table of transforms can bebuiltupandusedtocarryouttheinversetransformation,exactlyasatableoflogarithms canbeusedtolookupantilogarithms.Theprecedingtransforms constitutethebeginnings of such a table. For a more complete set of Laplace transforms see upcoming Table 15.2 or AMS-55, Chapter 29 (see footnote 4 in Chapter 5 for the reference). Employing partial fractionexpansionsandvariousoperationaltheorems,whichareconsideredinsucceeding sections,facilitatesuse ofthetables. •There is some justification for suspecting that these tables are probably of more value insolvingtextbookexercisesthaninsolvingreal-worldproblems. •(2) A general technique for L−1will be developed in Section 15.12 by using the cal- culusofresidues. FIGURE 15.7Apossiblenull function. 968 Chapter 15 Integral Transforms •(3)Forthedifficultiesandthepossibilitiesofanumericalapproach—numericalinver- sion—werefer totheAdditionalReadings. Partial Fraction Expansion Utilizationofatableoftransforms(orinversetransforms)isfacilitatedbyexpanding f(s) inpartialfractions . Frequently f(s), our transform, occurs in the form g(s)/h(s) , whereg(s)andh(s)are polynomials with no common factors, g(s)being of lower degree than h(s). If the factors ofh(s)arealllinearanddistinct,thenbythemethodofpartialfractionswemaywrite f(s)=c1 s−a1+c2 s−a2+···+cn s−an, (15.110) wherethe ciareindependentof s.Theaiaretherootsof h(s).Ifanyoneoftheroots,say, a1, ismultiple(occurring mtimes), then f(s)hastheform f(s)=c1,m (s−a1)m+c1,m−1 (s−a1)m−1+···+c1,1 s−a1+nsummationdisplay i=2ci s−ai. (15.111) Finally, if one of the factors is quadratic, (s2+ps+q), then the numerator, instead of beingasimpleconstant,willhavetheform as+b s2+ps+q. There are various ways of determining the constants introduced. For instance, in Eq.(15.110)wemaymultiplythroughby (s−ai)andobtain ci=lims→ai(s−ai)f(s). (15.112) Inelementarycasesadirectsolutionisoftentheeasiest. Example 15.8.1 PARTIAL FRACTION EXPANSION Let f(s)=k2 s(s2+k2)=c s+as+b s2+k2. (15.113) Puttingtherightsideoftheequationoveracommondenominatorandequatinglikepowers ofsinthenumerator,weobtain k2 s(s2+k2)=c(s2+k2)+s(as+b) s(s2+k2), (15.114) c+a=0,s2;b=0,s1;ck2=k2,s0. Solvingthese (s/negationslash=0),weha v e c=1,b=0,a=−1, 15.8 Laplace Transforms 969 giving f(s)=1 s−s s2+k2, (15.115) and L−1braceleftbig f(s)bracerightbig =1−coskt (15.116) byEqs. (15.102)and(15.106). /squaresolid Example 15.8.2 ASTEPFUNCTION AsoneapplicationofLaplacetransforms, considertheevaluationof F(t)=integraldisplay∞ 0sintx xdx. (15.117) SupposewetaketheLaplacetransformofthis definite(andimproper)integral: Lbraceleftbiggintegraldisplay∞ 0sintx xdxbracerightbigg =integraldisplay∞ 0e−stintegraldisplay∞ 0sintx xdxdt. (15.118) Now,interchangingtheorderof integration(whichis justified),12weget integraldisplay∞ 01 xbracketleftbiggintegraldisplay∞ 0e−stsintxdtbracketrightbigg dx=integraldisplay∞ 0dx s2+x2, (15.119) sincethefactorinsquarebracketsisjusttheLaplacetransformofsin tx.Fromtheintegral tables, integraldisplay∞ 0dx s2+x2=1 stan−1parenleftbiggx sparenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle∞ 0=π 2s=f(s). (15.120) ByEq. (15.102)wecarry outtheinversetransformationtoobtain F(t)=π 2,t>0, (15.121) in agreement with an evaluation by the calculus of residues (Section 7.1). It has been as- sumed that t>0i nF(t).F o rF(−t)we need note only that sin (−tx)=−sintx,g i v i n g F(−t)=−F(t).Finally,if t=0,F(0)is clearlyzero. Therefore integraldisplay∞ 0sintx xdx=π 2bracketleftbig 2u(t)−1bracketrightbig =  π 2,t>0 0,t=0 −π 2,t <0.(15.122) Note thatintegraltext∞ 0(sintx/x)dx, taken as a function of t, describes a step function (Fig. 15.8), astepofheight πatt=0.Thisis consistentwithEq. (1.174). /squaresolid The technique in the preceding example was to (1) introduce a second integration— the Laplace transform, (2) reverse the order of integration and integrate, and (3) take the 12See—in the Additional Readings—Jeffreys and Jeffreys (1966), Chapter 1 (uniform convergence of integrals). 970 Chapter 15 Integral Transforms FIGURE 15.8F(t)=integraltext∞ 0sintx xdx, astepfunction. inverseLaplacetransform.Therearemanyopportunitieswherethistechniqueofreversing the order of integration can be applied and proved useful. Exercise 15.8.6 is a variation of this. Exercises 15.8.1 Provethat lims→∞sf(s)=lim t→+0F(t). Hint.Assumethat F(t)canbeexpressedas F(t)=summationtext∞ n=0antn. 15.8.2 Showthat 1 πlim s→0L{cosxt}=δ(x). 15.8.3 Verifythat Lbraceleftbiggcosat−cosbt b2−a2bracerightbigg =s (s2+a2)(s2+b2),a2/negationslash=b2. 15.8.4 Usingpartialfractionexpansions,showthat (a)L−1braceleftbigg1 (s+a)(s+b)bracerightbigg =e−at−e−bt b−a,a/negationslash=b. (b)L−1braceleftbiggs (s+a)(s+b)bracerightbigg =ae−at−be−bt a−b,a/negationslash=b. 15.8.5 Usingpartialfractionexpansions,showthatfor a2/negationslash=b2, (a)L−1braceleftbigg1 (s2+a2)(s2+b2)bracerightbigg =−1 a2−b2braceleftbiggsinat a−sinbt bbracerightbigg , (b)L−1braceleftbiggs2 (s2+a2)(s2+b2)bracerightbigg =1 a2−b2{asinat−bsinbt}. 15.9 Laplace Transform of Derivatives 971 15.8.6 The electrostatic potential of a charged conducting disk is known to have the general form(circularcylindricalcoordinates) /Phi1(ρ,z)=integraldisplay∞ 0e−k|z|J0(kρ)f(k)dk, withf(k)unknown. At large distances (z→∞)the potential must approach the Coulombpotential Q/4πε0z. Showthat lim k→0f(k)=q 4πε0. Hint. You may set ρ=0 and assume a Maclaurin expansion of f(k)or, using e−kz, constructadeltasequence. 15.8.7 Showthat (a)integraldisplay∞ 0coss sνds=π 2(ν−1)!cos(νπ/2),0<ν<1, (b)integraldisplay∞ 0sins sνds=π 2(ν−1)!sin(νπ/2),0<ν<2, Whyisνrestrictedto(0,1)for(a),to (0,2)for(b)?Theseintegralsmaybeinterpreted asFouriertransformsof s−νandasMellintransformsof sin sand coss. Hint. Replace s−νby a Laplace transform integral: L{tν−1}/(ν−1)!. Then integrate withrespectto s.Theresultingintegralcanbetreatedasabetafunction(Section8.4). 15.8.8 Afunction F(t)canbeexpandedinapowerseries (Maclaurin);thatis, F(t)=∞summationdisplay n=0antn. Then Lbraceleftbig F(t)bracerightbig =integraldisplay∞ 0e−st∞summationdisplay n=0antndt=∞summationdisplay n=0anintegraldisplay∞ 0e−sttndt. Show that f(s), the Laplace transform of F(t), contains no powers of sgreater than s−1. Checkyourresultbycalculating L{δ(t)},andcommentonthisfiasco. 15.8.9 ShowthattheLaplacetransformof M(a,c,x) is Lbraceleftbig M(a,c,x)bracerightbig =1 s2F1parenleftbigg a,1;c,1 sparenrightbigg . 15.9 L APLACE TRANSFORM OF DERIVATIVES Perhaps the main application of Laplace transforms is in converting differential equations intosimplerformsthatmaybesolvedmoreeasily.Itwillbeseen,forinstance,thatcoupled differentialequationswithconstantcoefficientstransformtosimultaneouslinearalgebraic equations. 972 Chapter 15 Integral Transforms Letus transform thefirstderivativeof F(t): Lbraceleftbig F′(t)bracerightbig =integraldisplay∞ 0e−stdF(t) dtdt. Integratingbyparts, weobtain Lbraceleftbig F′(t)bracerightbig =e−stF(t)vextendsinglevextendsinglevextendsingle∞ 0+sintegraldisplay∞ 0e−stF(t)dt =sLbraceleftbig F(t)bracerightbig −F(0). (15.123) Strictlyspeaking, F(0)=F(+0)13anddF/dtisrequiredtobeatleastpiecewisecontin- uous for 0≤t<∞. Naturally, both F(t)and its derivative must be such that the integrals do not diverge. Incidentally, Eq. (15.123) provides another proof of Exercise 15.8.8. An extensiongives Lbraceleftbig F(2)(t)bracerightbig =s2Lbraceleftbig F(t)bracerightbig −sF(+0)−F′(+0), (15.124) Lbraceleftbig F(n)(t)bracerightbig =snLbraceleftbig F(t)bracerightbig −sn−1F(+0)−···−F(n−1)(+0).(15.125) The Laplace transform, like the Fourier transform, replaces differentiation with multi- plication.InthefollowingexamplesODEsbecomealgebraicequations.Hereisthepower and the utility of the Laplace transform. But see Example 15.10.3 for what may happen if thecoefficientsare notconstant. Note how the initial conditions, F(+0),F′(+0), and so on, are incorporated into the transform.Equation(15.124)maybeusedtoderive L{sinkt}.We usetheidentity −k2sinkt=d2 dt2sinkt. (15.126) ThenapplyingtheLaplacetransformoperation,wehave −k2L{sinkt}=Lbraceleftbiggd2 dt2sinktbracerightbigg =s2L{sinkt}−ssin(0)−d dtsinktvextendsinglevextendsinglevextendsingle t=0. (15.127) Since sin(0)=0 andd/dtsinkt|t=0=k, L{sinkt}=k s2+k2, (15.128) verifyingEq.(15.107). 13Zero is approached from thepositive side. 15.9 Laplace Transform of Derivatives 973 Example 15.9.1 SIMPLE HARMONIC OSCILLATOR Asaphysicalexample,consideramass moscillatingundertheinfluenceofanidealspring, springconstant k. Asusual,frictionisneglected.ThenNewton’ssecondlawbecomes md2X(t) dt2+kX(t)=0; (15.129) also,wetakeasinitialconditions X(0)=X0,X′(0)=0. ApplyingtheLaplacetransform, weobtain mLbraceleftbiggd2X dt2bracerightbigg +kLbraceleftbig X(t)bracerightbig =0, (15.130) andbyuseof Eq.(15.124)this becomes ms2x(s)−msX0+kx(s)=0, (15.131) x(s)=X0s s2+ω2 0,withω2 0≡k m. (15.132) FromEq. (15.107)thisis seentobethetransform of cos ω0t, whichgives X(t)=X0cosω0t, (15.133) asexpected. /squaresolid Example 15.9.2 EARTH ’SNUTATION A somewhat more involved example is the nutation of the earth’s poles (force-free pre- cession). If we treat the Earth as a rigid (oblate) spheroid, the Euler equations of motion reduceto dX dt=−aY,dY dt=+aX, (15.134) wherea≡[(Iz−Ix)/Iz]ωz,X=ωx,Y=ωywith angular velocity vector ω= (ωx,ωy,ωz)(Fig. 15.9), Iz=moment of inertia about the z-axis and Iy=Ixmoment of inertia about the x-( o ry-)axis. The z-axis coincides with the axis of symmetry of the Earth.ItdiffersfromtheaxisfortheEarth’sdailyrotation, ω,bysome15meters,measured atthepoles. Transformationofthesecoupleddifferentialequationsyields sx(s)−X(0)=−ay(s), sy(s) −Y(0)=ax(s). (15.135) Combiningtoeliminate y(s),w eh a v e s2x(s)−sX(0)+aY(0)=−a2x(s), or x(s)=X(0)s s2+a2−Y(0)a s2+a2. (15.136) 974 Chapter 15 Integral Transforms FIGURE 15.9 Hence X(t)=X(0)cosat−Y(0)sinat. (15.137) Similarly, Y(t)=X(0)sinat+Y(0)cosat. (15.138) This is seen to be a rotation of the vector (X,Y)counterclockwise (for a>0) about the z-axiswithangle θ=atandangularvelocity a. Adirectinterpretationmaybefoundbychoosingthetimeaxisso that Y(0)=0.Then X(t)=X(0)cosat, Y(t)=X(0)sinat, (15.139) whicharetheparametricequationsforrotationof (X,Y)inacircularorbitofradius X(0), withangularvelocity ainthecounterclockwisesense. In the case of the Earth’s angular velocity, vector X(0)is about 15 meters, whereas a, as defined here, corresponds to a period (2π/a)of some 300 days. Actually because of departures from the idealized rigid body assumed in setting up Euler’s equations, the periodis about427days.14If inEq. (15.134)weset X(t)=Lx,Y(t)=Ly, whereLxandLyare thex- andy-components of the angular momentum L,a=−gLBz, gLis the gyromagnetic ratio, and Bzis the magnetic field (along the z-axis), then Eq. (15.134) describes the Larmor precession of charged bodies in a uniform magnetic fieldBz. /squaresolid 14D. Menzel, ed., Fundamental Formulas of Physics , Englewood Cliffs, NJ: Prentice-Hall (1955), reprinted, 2nd ed., Dover (1960), p. 695. 15.9 Laplace Transform of Derivatives 975 Dirac Delta Function Forusewithdifferentialequationsonefurthertransformishelpful—theDiracdeltafunc- tion:15 Lbraceleftbig δ(t−t0)bracerightbig =integraldisplay∞ 0e−stδ(t−t0)dt=e−st0,fort0≥0, (15.140) andfort0=0 Lbraceleftbig δ(t)bracerightbig =1, (15.141) whereitis assumedthatweareusingarepresentationof thedeltafunctionsuchthat integraldisplay∞ 0δ(t)dt=1,δ(t)=0,fort>0. (15.142) Asanalternatemethod, δ(t)maybeconsideredthelimitas ε→0o fF(t), where F(t)=  0,t <0, ε−1,0<t<ε, 0,t >ε.(15.143) Bydirectcalculation Lbraceleftbig F(t)bracerightbig =1−e−εs εs. (15.144) Takingthelimitoftheintegral(insteadoftheintegralofthelimit),wehave lim ε→0Lbraceleftbig F(t)bracerightbig =1, orEq. (15.141), Lbraceleftbig δ(t)bracerightbig =1. This delta function is frequently called the impulsefunction because it is so useful in describingimpulsiveforces,thatis, forceslastingonlyashort time. Example 15.9.3 IMPULSIVE FORCE Newton’ssecondlawfor impulsiveforceactingonaparticleof mass mbecomes md2X dt2=Pδ(t), (15.145) wherePisa constant.Transforming,weobtain ms2x(s)−msX(0)−mX′(0)=P. (15.146) 15Strictlyspeaking,theDiracdeltafunctionisundefined.However,theintegraloveritiswelldefined.Thisapproachisdeveloped inSection1.16 using deltasequences. 976 Chapter 15 Integral Transforms Foraparticlestartingfromrest, X′(0)=0.16We shallalsotake X(0)=0.Then x(s)=P ms2, (15.147) and X(t)=P mt, (15.148) dX(t) dt=P m,aconstant . (15.149) Theeffectoftheimpulse Pδ(t)istotransfer(instantaneously) Punitsoflinearmomentum totheparticle. Asimilaranalysisappliestotheballisticgalvanometer.Thetorqueonthegalvanometer is given initially by kι, in which ιis a pulse of current and kis a proportionality constant. Sinceιis ofshort duration,weset kι=kq δ(t), (15.150) whereqis thetotalchargecarriedbythecurrent ι.Then,with Ithemomentofinertia, Id2θ dt2=kq δ(t), (15.151) and, transforming as before, we find that the effect of the current pulse is a transfer of kq unitsofangularmomentumtothegalvanometer. /squaresolid Exercises 15.9.1 Use the expression for the transform of a second derivative to obtain the transform of coskt. 15.9.2 Amassmisattachedtooneendofanunstretchedspring,springconstant k(Fig.15.10). Attimet=0thefreeendofthespringexperiencesaconstantacceleration a,awayfrom themass.UsingLaplacetransforms, FIGURE 15.10Spring. 16This should be X′(+0). Toinclude the effect of the impulse, consider that the impulse will occurat t=εandletε→0. 15.9 Laplace Transform of Derivatives 977 (a) Findtheposition xofmasafunctionof time. (b) Determinethelimitingformof x(t)for small t. ANS.(a)x=1 2at2−a ω2(1−cosωt), ω2=k m, (b)x=aω2 4!t4,ωt≪1. 15.9.3 Radioactivenucleidecayaccordingtothelaw dN dt=−λN, Nbeingtheconcentrationofagivennuclideand λbeingtheparticulardecayconstant. This equation may be interpreted as stating that the rate of decay is proportional to the numberof theseradioactivenucleipresent.Theyalldecayindependently. Inaradioactiveseriesof ndifferentnuclides,startingwith N1, dN1 dt=−λ1N1, dN2 dt=λ1N1−λ2N2,andso on . dNn dt=λn−1Nn−1,stable. FindN1(t),N2(t),N3(t),n=3,withN1(0)=N0,N2(0)=N3(0)=0. ANS.N1(t)=N0e−λ1t,N2(t)=N0λ1 λ2−λ1parenleftbig e−λ1t−e−λ2tparenrightbig , N3(t)=N0parenleftbigg 1−λ2 λ2−λ1e−λ1t+λ1 λ2−λ1e−λ2tparenrightbigg . Findanapproximateexpressionfor N2andN3, validfor small twhenλ1≈λ2. ANS.N2≈N0λ1t,N3≈N0 2λ1λ2t2. Findapproximateexpressionsfor N2andN3, validfor large t,when (a)λ1≫λ2, (b)λ1≪λ2. ANS.(a) N2≈N0e−λ2t, N3≈N0parenleftbig 1−e−λ2tparenrightbig ,λ1t≫1. (b)N2≈N0λ1 λ2e−λ1t, N3≈N0parenleftbig 1−e−λ1tparenrightbig ,λ2t≫1. 15.9.4 Theformationof anisotopeinanuclearreactorisgivenby dN2 dt=nvσ1N10−λ2N2(t)−nvσ2N2(t). 978 Chapter 15 Integral Transforms Heretheproduct nvistheneutronflux,neutronspercubiccentimeter,timescentimeters per second mean velocity; σ1andσ2(cm2) are measures of the probability of neutron absorption by the original isotope, concentration N10, which is assumed constant and the newly formed isotope, concentration N2, respectively. The radioactive decay con- stantfortheisotopeis λ2. (a) Findtheconcentration N2ofthenewisotopeasa functionoftime. (b) If the original element is Eu153,σ1=400 barns=400×10−24cm2,σ2= 1000 barns=1000×10−24cm2, andλ2=1.4×10−9s−1.I fN10=1020and (nv)=109cm−2s−1, findN2, the concentration of Eu154after one year of con- tinuousirradiation.Is theassumptionthat N1isconstantjustified? 15.9.5 InanuclearreactorXe135isformedasbothadirectfissionproductandadecayproduct of I135, half-life, 6.7 hours. The half-life of Xe135is 9.2 hours. Because Xe135strongly absorbs thermalneutronsthereby“poisoning”the nuclearreactor, its concentrationis a matterof greatinterest.Therelevantequationsare dNI dt=γIϕσfNU−λINI, dNX dt=λINI+γXϕσfNU−λXNX−ϕσXNX. HereNI=concentrationof I135(Xe135,U235). Assume NU=constant, γI=yieldofI135per fission=0.060, γX=yieldofXe135directfromfission =0.003, λI=I135parenleftbig Xe135parenrightbig decayconstant =ln2 t1/2=0.693 t1/2, σf=thermalneutronfissioncross sectionfor U235, σX=thermalneutronabsorptioncross sectionfor Xe135 =3.5×106barns=3.5×10−18cm2. (σItheabsorptioncross sectionof I135,is negligible. ) ϕ=neutronflux=neutrons/cm3×meanvelocity(cm /s). (a) Find NX(t)intermsofneutronflux ϕandtheproduct σfNU. (b) Find NX(t→∞). (c) After NXhasreachedequilibrium,thereactorisshutdown, ϕ=0.FindNX(t)fol- lowing shutdown. Notice the increase in NX, which may for a few hours interfere withstartingthereactorupagain. 15.10 Other Properties 979 15.10 O THER PROPERTIES Substitution If we replace the parameter sbys−ain the definition of the Laplace transform (Eq. (15.99)), wehave f(s−a)=integraldisplay∞ 0e−(s−a)tF(t)dt=integraldisplay∞ 0e−steatF(t)dt =Lbraceleftbig eatF(t)bracerightbig . (15.152) Hence the replacement of swiths−acorresponds to multiplying F(t)byeat, and con- versely. This result can be used to good advantage in extending our table of transforms. FromEq. (15.107)wefindimmediatelythat Lbraceleftbig eatsinktbracerightbig =k (s−a)2+k2; (15.153) also, Lbraceleftbig eatcosktbracerightbig =s−a (s−a)2+k2,s>a. Example 15.10.1 DAMPED OSCILLATOR These expressions are useful when we consider an oscillating mass with damping propor- tionaltothevelocity.Equation(15.129), withsuchdampingadded,becomes mX′′(t)+bX′(t)+kX(t)=0, (15.154) in which bis a proportionality constant. Let us assume that the particle starts from rest at X(0)=X0,X′(0)=0.Thetransformedequationis mbracketleftbig s2x(s)−sX0bracketrightbig +bbracketleftbig sx(s)−X0bracketrightbig +kx(s)=0, (15.155) and x(s)=X0ms+b ms2+bs+k. (15.156) Thismaybehandledbycompletingthesquareofthedenominator: s2+b ms+k m=parenleftbigg s+b 2mparenrightbigg2 +parenleftbiggk m−b2 4m2parenrightbigg . (15.157) If thedampingissmall, b2<4km, thelasttermispositiveandwillbedenotedby ω2 1: x(s)=X0s+b/m (s+b/2m)2+ω2 1 =X0s+b/2m (s+b/2m)2+ω2 1+X0(b/2mω1)ω1 (s+b/2m)2+ω2 1. (15.158) 980 Chapter 15 Integral Transforms ByEq. (15.153), X(t)=X0e−(b/2m)tparenleftbigg cosω1t+b 2mω1sinω1tparenrightbigg =X0ω0 ω1e−(b/2m)tcos(ω1t−ϕ), (15.159) where tanϕ=b 2mω1,ω2 0=k m. Ofcourse,as b→0,thissolutiongoesovertotheundampedsolution(Section15.9). /squaresolid RLC Analog Itisworthnotingthesimilaritybetweenthisdampedsimpleharmonicoscillationofamass on a spring and an RLCcircuit (resistance, inductance, and capacitance) (Fig. 15.11). At any instant the sum of the potential differences around the loop must be zero (Kirchhoff’s law,conservationof energy).Thisgives LdI dt+RI+1 Cintegraldisplayt Idt=0. (15.160) Differentiatingthecurrent Iwithrespecttotime(to eliminatetheintegral),wehave Ld2I dt2+RdI dt+1 CI=0. (15.161) If we replace I(t)withX(t),Lwithm,Rwithb, andC−1withk, then Eq. (15.161) is identical with the mechanical problem. It is but one example of the unification of diverse branchesofphysicsbymathematics.AmorecompletediscussionwillbefoundinOlson’s book.17 FIGURE 15.11RLCcircuit. 17H.F.Olson, Dynamical Analogies , NewYork: VanNostrand (1943). 15.10 Other Properties 981 FIGURE 15.12Translation. Translation Thistimelet f(s)bemultipliedby e−bs,b>0: e−bsf(s)=e−bsintegraldisplay∞ 0e−stF(t)dt =integraldisplay∞ 0e−s(t+b)F(t)dt. (15.162) Nowlett+b=τ. Equation(15.162)becomes e−bsf(s)=integraldisplay∞ be−sτF(τ−b)dτ =integraldisplay∞ 0e−sτF(τ−b)u(τ−b)dτ, (15.163) whereu(τ−b)istheunitstepfunction.Thisrelationisoftencalledthe Heavisideshifting theorem (Fig. 15.12). SinceF(t)is assumed to be equal to zero for t<0,F(τ−b)=0f o r0≤τ<b. Thereforewecanextendthelowerlimittozerowithoutchangingthevalueoftheintegral. Then,notingthat τis onlyavariableof integration,weobtain e−bsf(s)=Lbraceleftbig F(t−b)bracerightbig . (15.164) Example 15.10.2 ELECTROMAGNETIC WAVES The electromagnetic wave equation with E=EyorEz, a transverse wave propagating alongthe x-axis,is ∂2E(x,t) ∂x2−1 v2∂2E(x,t) ∂t2=0. (15.165) Transformingthisequationwithrespectto t, weget ∂2 ∂x2Lbraceleftbig E(x,t)bracerightbig −s2 v2Lbraceleftbig E(x,t)bracerightbig +s v2E(x,0)+1 v2∂E(x,t) ∂tvextendsinglevextendsinglevextendsinglevextendsingle t=0=0.(15.166) 982 Chapter 15 Integral Transforms If wehavetheinitialcondition E(x,0)=0 and ∂E(x,t) ∂tvextendsinglevextendsinglevextendsinglevextendsingle t=0=0, then ∂2 ∂x2Lbraceleftbig E(x,t)bracerightbig =s2 v2Lbraceleftbig E(x,t)bracerightbig . (15.167) Thesolution(of this ODE)is Lbraceleftbig E(x,t)bracerightbig =c1e−(s/v)x+c2e+(s/v)x. (15.168) The “constants” c1andc2are obtained by additional boundary conditions. They are constant with respect to xbut may depend on s. If our wave remains finite as x→ ∞,L{E(x,t)}will also remain finite. Hence c2=0. IfE(0,t)is denoted by F(t), then c1=f(s)and Lbraceleftbig E(x,t)bracerightbig =e−(s/v)xf(s). (15.169) Fromthetranslationproperty(Eq. (15.164))wefindimmediatelythat E(x,t)=braceleftBiggFparenleftbig t−x vparenrightbig ,t≥x v, 0,t <x v.(15.170) Differentiation and substitution into Eq. (15.165) verifies Eq. (15.170). Our solution rep- resents a wave (or pulse) moving in the positive x-direction with velocity v. Note that for x>v tthe region remains undisturbed; the pulse has not had time to get there. If we had wanted a signal propagated along the negative x-axis,c1would have been set equal to 0 andwewouldhaveobtained E(x,t)=braceleftBigg Fparenleftbig t+x vparenrightbig ,t≥−x v, 0,t <−x v,(15.171) awavealongthenegative x-axis. /squaresolid Derivative of a Transform WhenF(t), which is at least piecewise continuous, and sare chosen so that e−stF(t) convergesexponentiallyforlarge s, theintegral integraldisplay∞ 0e−stF(t)dt is uniformly convergent and may be differentiated (under the integral sign) with respect tos. Then f′(s)=integraldisplay∞ 0(−t)e−stF(t)dt=Lbraceleftbig −tF(t)bracerightbig . (15.172) Continuingthisprocess, weobtain f(n)(s)=Lbraceleftbig (−t)nF(t)bracerightbig . (15.173) 15.10 Other Properties 983 Alltheintegralssoobtainedwillbeuniformlyconvergentbecauseofthedecreasingexpo- nentialbehaviorof e−stF(t). This sametechniquemaybeappliedtogeneratemoretransforms. Forexample, Lbraceleftbig ektbracerightbig =integraldisplay∞ 0e−stektdt=1 s−k,s>k. (15.174) Differentiatingwithrespectto s(orwithrespectto k), weobtain Lbraceleftbig tektbracerightbig =1 (s−k)2,s>k. (15.175) Example 15.10.3 BESSEL ’SEQUATION An interesting application of a differentiated Laplace transform appears in the solution of Bessel’sequationwith n=0.FromChapter11wehave x2y′′(x)+xy′(x)+x2y(x)=0. (15.176) Dividing by xand substituting t=xandF(t)=y(x)to agree with the present notation, wesee thattheBesselequationbecomes tF′′(t)+F′(t)+tF(t)=0. (15.177) We need a regular solution, in particular, F(0)=1. From Eq. (15.177) with t=0, F′(+0)=0. Also, we assume that our unknown F(t)has a transform. Transforming and usingEqs. (15.123), (15.124),and(15.172),wehave −d dsbracketleftbig s2f(s)−sbracketrightbig +sf(s)−1−d dsf(s)=0. (15.178) RearrangingEq. (15.178),weobtain parenleftbig s2+1parenrightbig f′(s)+sf(s)=0, (15.179) or df f=−sds s2+1, (15.180) afirst-order ODE.Byintegration, lnf(s)=−1 2lnparenleftbig s2+1parenrightbig +lnC, (15.181) whichmayberewrittenas f(s)=C√ s2+1. (15.182) To make use of Eq. (15.108), we expand f(s)in a series of negative powers of s, conver- gentfors>1: f(s)=C sparenleftbigg 1+1 s2parenrightbigg−1/2 =C sbracketleftbigg 1−1 2s2+1·3 22·2!s4−···+(−1)n(2n)! (2nn!)2s2n+···bracketrightbigg .(15.183) 984 Chapter 15 Integral Transforms Inverting,term byterm,weobtain F(t)=C∞summationdisplay n=0(−1)nt2n (2nn!)2. (15.184) WhenCis set equal to 1, as required by the initial condition F(0)=1,F(t)is justJ0(t), ourfamiliarBesselfunctionoforderzero.Hence Lbraceleftbig J0(t)bracerightbig =1√ s2+1. (15.185) Notethatweassumed s>1.Theprooffor s>0 is leftasaproblem. It is worth noting that this application was successful and relatively easy because we tookn=0 in Bessel’s equation. This made it possible to divide out a factor of x(ort). If this had not been done, the terms of the form t2F(t)would have introduced a second derivative of f(s). The resulting equation would have been no easier to solve than the originalone. WhenwegobeyondlinearODEswithconstantcoefficients,theLaplacetransformmay stillbeapplied,butthereis noguaranteethatitwillbehelpful. The application to Bessel’s equation, n/negationslash=0, will be found in the references. Alterna- tively,wecanshowthat Lbraceleftbig Jn(at)bracerightbig =a−n(√ s2+a2−s)n √ s2+a2(15.186) byexpressing Jn(t)as aninfiniteseriesandtransformingtermbyterm. /squaresolid Integration of Transforms Again, with F(t)at least piecewise continuous and xlarge enough so that e−xtF(t)de- creasesexponentially(as x→∞), theintegral f(x)=integraldisplay∞ 0e−xtF(t)dt (15.187) is uniformly convergent with respect to x. This justifies reversing the order of integration inthefollowingequation: integraldisplayb sf(x)dx=integraldisplayb sdxintegraldisplay∞ 0dte−xtF(t) =integraldisplay∞ 0F(t) tparenleftbig e−st−e−btparenrightbig dt, (15.188) on integrating with respect to x. The lower limit sis chosen large enough so that f(s)is withintheregionofuniformconvergence.Nowletting b→∞,weha v e integraldisplay∞ sf(x)dx=integraldisplay∞ 0F(t) te−stdt=LbraceleftbiggF(t) tbracerightbigg , (15.189) providedthat F(t)/tisfiniteat t=0 ordivergeslessstronglythan t−1(sothat L{F(t)/t} willexist). 15.10 Other Properties 985 Limits of Integration — Unit Step Function TheactuallimitsofintegrationfortheLaplacetransformmaybespecifiedwiththe(Heav- iside)unitstepfunction u(t−k)=braceleftbigg0,t<k 1,t>k. Forinstance, Lbraceleftbig u(t−k)bracerightbig =integraldisplay∞ ke−stdt=1 se−ks. A rectangular pulse of width kand unit height is described by F(t)=u(t)−u(t−k). TakingtheLaplacetransform,weobtain Lbraceleftbig u(t)−u(t−k)bracerightbig =integraldisplayk 0e−stdt=1 sparenleftbig 1−e−ksparenrightbig . The unit step function is also used in Eq. (15.163) and could be invoked in Exer- cise15.10.13. Exercises 15.10.1 Solve Eq. (15.154), which describes a damped simple harmonic oscillator for X(0)= X0,X′(0)=0,and (a)b2=4 km(criticallydamped), (b)b2>4 km(overdamped). ANS.(a)X( t)=X0e−(b/2m)tparenleftbigg 1+b 2mtparenrightbigg . 15.10.2 SolveEq.(15.154),whichdescribesadampedsimpleharmonicoscillatorfor X(0)=0, X′(0)=v0, and (a)b2<4 km(underdamped), (b)b2=4 km(criticallydamped), (c)b2>4 km(overdamped). ANS.(a) X(t)=v0 ω1e−(b/2m)tsinω1t, (b)X(t)=v0te−(b/2m)t. 15.10.3 Themotionofabodyfallinginaresistingmediummaybedescribedby md2X(t) dt2=mg−bdX(t) dt 986 Chapter 15 Integral Transforms FIGURE 15.13Ringingcircuit. whentheretardingforceisproportionaltothevelocity.Find X(t)anddX(t)/dt forthe initialconditions X(0)=dX dtvextendsinglevextendsinglevextendsinglevextendsingle t=0=0. 15.10.4 Ringing circuit . In certain electronic circuits, resistance, inductance, and capacitance are placed in the plate circuit in parallel (Fig. 15.13). A constant voltage is maintained across the parallel elements, keeping the capacitor charged. At time t=0 the circuit is disconnected from the voltage source. Find the voltages across the parallel elements R,L,andCasafunctionoftime.Assume Rtobelarge. Hint.ByKirchhoff’slaws IR+IC+IL=0 and ER=EC=EL, where ER=IRR, E L=LdIL dt EC=q0 C+1 Cintegraldisplayt 0ICdt, q0=initialchargeofcapacitor. WiththeDCimpedanceof L=0,letIL(0)=I0,EL(0)=0.Thismeans q0=0. 15.10.5 WithJ0(t)expressed as a contour integral, apply the Laplace transform operation, re- versetheorderofintegration,andthusshowthat Lbraceleftbig J0(t)bracerightbig =parenleftbig s2+1parenrightbig−1/2,fors>0. 15.10.6 Develop the Laplace transform of Jn(t)fromL{J0(t)}by using the Bessel function recurrencerelations. Hint.Hereis achancetousemathematicalinduction. 15.10.7 A calculation of the magnetic field of a circular current loop in circular cylindrical coordinatesleadstotheintegral integraldisplay∞ 0e−kzkJ1(ka)dk,ℜ(z)≥0. Showthatthisintegralisequalto a/(z2+a2)3/2. 15.10 Other Properties 987 15.10.8 The electrostatic potential of a point charge qat the origin in circular cylindrical coor- dinatesis q 4πε0integraldisplay∞ 0e−kzJ0(kρ)dk=q 4πε0·1 (ρ2+z2)1/2,ℜ(z)≥0. FromthisrelationshowthattheFouriercosineandsinetransformsof J0(kρ)are (a)radicalbiggπ 2Fcbraceleftbig J0(kρ)bracerightbig =integraldisplay∞ 0J0(kρ)coskζdk=braceleftbiggparenleftbig ρ2−ζ2parenrightbig−1/2,ρ>ζ, 0,ρ <ζ. (b)radicalbiggπ 2Fsbraceleftbig J0(kρ)bracerightbig =integraldisplay∞ 0J0(kρ)sinkζdk=braceleftbigg0,ρ >ζ,parenleftbig ρ2−ζ2parenrightbig−1/2,ρ<ζ. Hint.Replace zbyz+iζandtakethelimitas z→0. 15.10.9 Showthat Lbraceleftbig I0(at)bracerightbig =parenleftbig s2−a2parenrightbig−1/2,s>a. 15.10.10 VerifythefollowingLaplacetransforms: (a)Lbraceleftbig j0(at)bracerightbig =Lbraceleftbiggsinat atbracerightbigg =1 acot−1parenleftbiggs aparenrightbigg , (b)Lbraceleftbig n0(at)bracerightbig doesnotexist, (c)Lbraceleftbig i0(at)bracerightbig =Lbraceleftbiggsinhat atbracerightbigg =1 2alns+a s−a=1 acoth−1parenleftbiggs aparenrightbigg , (d)Lbraceleftbig k0(at)bracerightbig doesnotexist. 15.10.11 DevelopaLaplacetransformsolutionofLaguerre’sequation tF′′(t)+(1−t)F′(t)+nF(t)=0. Note that you need a derivative of a transform and a transform of derivatives. Go as far asyoucanwith n;then(andonlythen)set n=0. 15.10.12 ShowthattheLaplacetransformoftheLaguerrepolynomial Ln(at)is givenby Lbraceleftbig Ln(at)bracerightbig =(s−a)n sn+1,s>0. 15.10.13 Showthat Lbraceleftbig E1(t)bracerightbig =1 sln(s+1), s > 0, where E1(t)=integraldisplay∞ te−τ τdτ=integraldisplay∞ 1e−xt xdx. E1(t)is theexponential-integralfunction. 988 Chapter 15 Integral Transforms 15.10.14 (a) FromEq. (15.189)showthat integraldisplay∞ 0f(x)dx=integraldisplay∞ 0F(t) tdt, providedtheintegralsexist. (b) Fromtheprecedingresultshowthat integraldisplay∞ 0sint tdt=π 2, inagreementwithEqs. (15.122)and(7.56). 15.10.15 (a) Showthat Lbraceleftbiggsinkt tbracerightbigg =cot−1parenleftbiggs kparenrightbigg . (b) Usingthisresult (with k=1), prove that Lbraceleftbig si(t)bracerightbig =−1 stan−1s, where si(t)=−integraldisplay∞ tsinx xdx,thesineintegral . 15.10.16 IfF(t)is periodic (Fig. 15.14) with a period aso thatF(t+a)=F(t)for allt≥0, showthat Lbraceleftbig F(t)bracerightbig =integraltexta 0e−stF(t)dt 1−e−as, withtheintegrationnowoveronlythe first period ofF(t). 15.10.17 FindtheLaplacetransformof thesquarewave(period a) definedby F(t)=braceleftBigg 1,0<t<a 2 0,a 2<t<a. ANS.f(s)=1 s·1−e−as/2 1−e−as. FIGURE 15.14Periodicfunction. 15.10 Other Properties 989 15.10.18 Showthat (a)L{coshatcosat}=s3 s4+4a4,(c)L{sinhatcosat}=as2−2a3 s4+4a4, (b)L{coshatsinat}=as2+2a3 s4+4a4,(d)L{sinhatsinat}=2a2s s4+4a4. 15.10.19 Showthat (a)L−1braceleftbigparenleftbig s2+a2parenrightbig−2bracerightbig =1 2a3sinat−1 2a2tcosat, (b)L−1braceleftbig sparenleftbig s2+a2parenrightbig−2bracerightbig =1 2atsinat, (c)L−1braceleftbig s2parenleftbig s2+a2parenrightbig−2bracerightbig =1 2asinat+1 2tcosat, (d)L−1braceleftbig s3parenleftbig s2+a2parenrightbig−2bracerightbig =cosat−a 2tsinat. 15.10.20 Showthat Lbraceleftbigparenleftbig t2−k2parenrightbig−1/2u(t−k)bracerightbig =K0(ks). Hint. Try transforming an integral representation of K0(ks)into the Laplace transform integral. 15.10.21 TheLaplacetransform integraldisplay∞ 0e−xsxJ0(x)dx=s (s2+1)3/2 mayberewrittenas 1 s2integraldisplay∞ 0e−yyJ0parenleftbiggy sparenrightbigg dy=s (s2+1)3/2, whichisinGauss–Laguerrequadratureform.Evaluatethisintegralfor s=1.0,0.9,0.8, ...,decreasing sin steps of 0.1 until the relative error rises to 10 percent. (The effect ofdecreasing sistomaketheintegrandoscillatemorerapidlyperunitlengthof y,thus decreasingtheaccuracyof thenumericalquadrature.) 15.10.22 (a) Evaluate integraldisplay∞ 0e−kzkJ1(ka)dk bytheGauss–Laguerrequadrature.Take a=1 andz=0.1(0.1)1.0. (b) From the analytic form, Exercise 15.10.7, calculate the absolute error and the rel- ativeerror. 990 Chapter 15 Integral Transforms 15.11 C ONVOLUTION (FALTUNGS )THEOREM One of the most important properties of the Laplace transform is that given by the convo- lution,or Faltungs,theorem.18Wetaketwotransforms, f1(s)=Lbraceleftbig F1(t)bracerightbig andf2(s)=Lbraceleftbig F2(t)bracerightbig , (15.190) and multiplythem together. To avoid complicationswhen changingvariables, we hold the upperlimitsfinite: f1(s)f2(s)=lima→∞integraldisplaya 0e−sxF1(x)dxintegraldisplaya−x 0e−syF2(y)dy. (15.191) The upper limits are chosen so that the area of integration, shown in Fig. 15.15a, is the shaded triangle, not the square. If we integrate over a square in the xy-plane, we have a parallelogram in the tz-plane, which simply adds complications. This modification is permissiblebecausethetwointegrandsareassumedtodecreaseexponentially.Inthelimit a→∞, the integral over the unshaded triangle will give zero contribution. Substituting x=t−z,y=z,theregionofintegrationismappedintothetriangleshowninFig.15.15b. To verify the mapping, map the vertices: t=x+y,z=y. Using Jacobians to transform theelementof area,wehave dxdy=vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂t∂y ∂t ∂x ∂z∂y ∂zvextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingledtdz=vextendsinglevextendsinglevextendsinglevextendsingle10 −11vextendsinglevextendsinglevextendsinglevextendsingledtdz (15.192) ordxdy=dtdz.Withthis substitutionEq.(15.191)becomes f1(s)f2(s)=lima→∞integraldisplaya 0e−stintegraldisplayt 0F1(t−z)F2(z)dzdt =Lbraceleftbiggintegraldisplayt 0F1(t−z)F2(z)dzbracerightbigg . (15.193) a b FIGURE 15.15 Changeofvariables, (a)xy-plane(b) zt-plane. 18An alternatederivation employs the Bromwich integral (Section 15.12). This is Exercise 15.12.3. 15.11 Convolution (Faltungs) Theorem 991 Forconveniencethisintegralisrepresentedbythesymbol integraldisplayt 0F1(t−z)F2(z)dz≡F1∗F2 (15.194) and referred to as the convolution , closely analogous to the Fourier convolution (Sec- tion15.5). If wesubstitute w=t−z, wefind F1∗F2=F2∗F1, (15.195) showingthattherelationis symmetric. Carryingouttheinversetransform, wealsofind L−1braceleftbig f1(s)f2(s)bracerightbig =integraldisplayt 0F1(t−z)F2(z)dz. (15.196) This can be useful in the development of new transforms or as an alternative to a partial fractionexpansion.Oneimmediateapplicationisinthesolutionofintegralequations(Sec- tion16.2).Sincetheupperlimit, t,isvariable,thisLaplaceconvolutionisusefulintreating Volterraintegralequations.TheFourierconvolutionwithfixed(infinite)limitswouldapply toFredholmintegralequations. Example 15.11.1 DRIVEN OSCILLATOR WITH DAMPING As one illustration of the use of the convolution theorem, let us return to the mass mon a spring, with damping and a driving force F(t). The equation of motion ((15.129) or (15.154))nowbecomes mX′′(t)+bX′(t)+kX(t)=F(t). (15.197) Initial conditions X(0)=0,X′(0)=0 are used to simplify this illustration, and the trans- formedequationis ms2x(s)+bsx(s)+kx(s)=f(s), (15.198) or x(s)=f(s) m1 (s+b/2m)2+ω2 1, (15.199) whereω2 1≡k/m−b2/4m2, asbefore. Bytheconvolutiontheorem(Eq. (15.193)or(15.196)), X(t)=1 mω1integraldisplayt 0F(t−z)e−(b/2m)zsinω1zdz. (15.200) If theforce isimpulsive, F(t)=Pδ(t),19 X(t)=P mω1e−(b/2m)tsinω1t. (15.201) 19Notethat δ(t)liesinsidethe interval[0,t]. 992 Chapter 15 Integral Transforms Prepresents the momentum transferred by the impulse, and the constant P/mtakes the placeofaninitialvelocity X′(0). IfF(t)=F0sinωt,Eq.(15.200)maybeused,butapartialfractionexpansionisperhaps moreconvenient.With f(s)=F0ω s2+ω2 Eq.(15.199)becomes x(s)=F0ω m·1 s2+ω2·1 (s+b/2m)2+ω2 1 =F0ω mbracketleftbigga′s+b′ s2+ω2+c′s+d′ (s+b/2m)2+ω2 1bracketrightbigg . (15.202) Thecoefficients a′,b′,c′, andd′areindependentof s. Directcalculationshows −1 a′=b mω2+m bparenleftbig ω2 0−ω2parenrightbig2, −1 b′=−m bparenleftbig ω2 0−ω2parenrightbigbracketleftbiggb mω2+m bparenleftbig ω2 0−ω2parenrightbig2bracketrightbigg . Sincec′andd′will lead to exponentially decreasing terms (transients), they will be dis- cardedhere.Carryingouttheinverseoperation,wefindforthesteady-statesolution X(t)=F0 [b2ω2+m2(ω2 0−ω2)2]1/2sin(ωt−ϕ), (15.203) where tanϕ=bω m(ω2 0−ω2). Differentiatingthedenominator,wefindthattheamplitudehasamaximumwhen ω2=ω2 0−b2 2m2=ω2 1−b2 4m2. (15.204) This is the resonance condition.20At resonance the amplitude becomes F0/bω1, showing that the mass mgoes into infinite oscillationat resonance if dampingis neglected (b=0). Itis worthnotingthatwehavehadthreedifferentcharacteristicfrequencies: ω2 2=ω2 0−b2 2m2, resonancefor forcedoscillations,withdamping; ω2 1=ω2 0−b2 4m2, 20Theamplitude (squared) has thetypical resonance denominator, the Lorentz line shape,Exercise 15.3.9. 15.11 Convolution (Faltungs) Theorem 993 freeoscillationfrequency,withdamping;and ω2 0=k m, freeoscillationfrequency,nodamping.Theycoincideonlyif thedampingis zero. /squaresolid Returning to Eqs. (15.197) and (15.199), Eq. (15.197) is our ODE for the response of adynamicalsystemtoanarbitrarydrivingforce.Thefinalresponseclearlydependsonboth the driving force and the characteristics of our system. This dual dependence is separated in the transform space. In Eq. (15.199) the transform of the response (output) appears as theproductoftwofactors,onedescribingthedrivingforce(input)andtheotherdescribing the dynamical system. This latter part, which modifies the input and yields the output, is oftencalleda transferfunction .Specifically,[(s+b/2m)2+ω2 1]−1isthetransferfunction correspondingtothisdampedoscillator.Theconceptofatransferfunctionisofgreatusein thefieldofservomechanisms.Oftenthecharacteristicsofaparticularservomechanismare described by giving its transfer function. The convolution theorem then yields the output signalfor aparticularinputsignal. Exercises 15.11.1 Fromtheconvolutiontheoremshowthat 1 sf(s)=Lbraceleftbiggintegraldisplayt 0F(x)dxbracerightbigg , wheref(s)=L{F(t)}. 15.11.2 IfF(t)=taandG(t)=tb,a>−1,b>−1: (a) Showthattheconvolution F∗G=ta+b+1integraldisplay1 0ya(1−y)bdy. (b) Byusingtheconvolutiontheorem,showthat integraldisplay1 0ya(1−y)bdy=a!b! (a+b+1)!. Thisis theEulerformulafor thebetafunction(Eq. (8.59a)). 15.11.3 Usingtheconvolutionintegral,calculate L−1braceleftbiggs (s2+a2)(s2+b2)bracerightbigg ,a2/negationslash=b2. 15.11.4 Anundampedoscillatorisdrivenbyaforce F0sinωt.Findthedisplacementasafunc- tion of time. Notice that it is a linear combination of two simple harmonic motions, one with the frequency of the driving force and one with the frequency ω0of the free oscillator.(Assume X(0)=X′(0)=0.) ANS.X(t)=F0/m ω2−ω2 0parenleftbiggω ω0sinω0t−sinωtparenrightbigg . 994 Chapter 15 Integral Transforms OtherexercisesinvolvingtheLaplaceconvolutionappearinSection16.2. 15.12 I NVERSE LAPLACE TRANSFORM Bromwich Integral We now develop an expression for the inverse Laplace transform L−1appearing in the equation F(t)=L−1braceleftbig f(s)bracerightbig . (15.205) One approach lies in the Fourier transform, for which we know the inverse relation. There is a difficulty, however. Our Fourier transformable function had to satisfy the Dirichlet conditions.Inparticular,werequiredthat limω→∞G(ω)=0 (15.206) so that the infinite integral would be well defined.21Now we wish to treat functions F(t) that may diverge exponentially. To surmount this difficulty, we extract an exponential fac- tor,eγt, fromour(possibly) divergentLaplacefunctionandwrite F(t)=eγtG(t). (15.207) IfF(t)divergesas eαt,werequire γtobegreaterthan αsothatG(t)willbeconvergent . Now,with G(t)=0fort<0andotherwisesuitablyrestrictedsothatitmayberepresented byaFourierintegral(Eq. (15.20)), G(t)=1 2πintegraldisplay∞ −∞eiutduintegraldisplay∞ 0G(v)e−iuvdv. (15.208) UsingEq. (15.207),wemayrewrite(15.208)as F(t)=eγt 2πintegraldisplay∞ −∞eiutduintegraldisplay∞ 0F(v)e−γve−iuvdv. (15.209) Now,withthechangeofvariable, s=γ+iu, (15.210) the integral over visthrownintotheformof aLaplacetransform, integraldisplay∞ 0f(v)e−svdv=f(s); (15.211) 21If deltafunctions areincluded, G(ω)may be acosine. Although this does not satisfy Eq.(15.206), G(ω)is still bounded. 15.12 Inverse Laplace Transform 995 FIGURE 15.16Singularities ofestf(s). sis now a complex variable, and ℜ(s)≥γto guarantee convergence. Notice that the Laplace transform has mapped a function specified on the positive real axis onto the com- plexplane,ℜ(s)≥γ.22 Withγas aconstant, ds=idu.SubstitutingEq. (15.211)intoEq.(15.209), weobtain F(t)=1 2πiintegraldisplayγ+i∞ γ−i∞estf(s)ds. (15.212) Here is our inverse transform . We have rotated the line of integration through 90◦(by usingds=idu). The path has become an infinite vertical line in the complex plane, the constantγhavingbeenchosensothatallthesingularitiesof f(s)areontheleft-handside (Fig.15.16). Equation (15.212), our inverse transformation, is usually known as the Bromwich in- tegral, although sometimes it is referred to as the Fourier–Mellin theorem orFourier– Mellin integral . This integral may now be evaluated by the regular methods of contour integration(Chapter7).If t>0,thecontourmaybeclosedbyaninfinitesemicircleinthe lefthalf-plane.Thenbytheresiduetheorem(Section7.1) F(t)=/Sigma1(residuesincludedfor ℜ(s)<γ). (15.213) Possibly this means of evaluation with ℜ(s)ranging through negative values seems para- doxical in view of our previous requirement that ℜ(s)≥γ. The paradox disappears when we recall that the requirement ℜ(s)≥γwas imposed to guarantee convergence of the Laplace transform integral that defined f(s). Oncef(s)is obtained, we may then pro- ceed to exploit its properties as an analytical function in the complex plane wherever we choose.23In effect we are employing analytic continuation to get L{F(t)}in the left half- plane, exactly as the recurrence relation for the factorial function was used to extend the Eulerintegraldefinition(Eq. (8.5)) tothelefthalf-plane. PerhapsapairofexamplesmayclarifytheevaluationofEq. (15.212). 22For a derivation of the inverse Laplace transform using only real variables, see C. L. Bohn and R. W. Flynn, Real variable inversion of Laplacetransforms: Anapplication in plasma physics. Am.J.Phys. 46: 1250 (1978). 23In numerical work f(s)may well be available only for discrete real, positive values of s. Then numerical procedures are indicated. SeeKrylov and Skoblya in the Additional Reading. 996 Chapter 15 Integral Transforms Example 15.12.1 INVERSION VIA CALCULUS OF RESIDUES Iff(s)=a/(s2−a2), then estf(s)=aest s2−a2=aest (s+a)(s−a). (15.214) TheresiduesmaybefoundbyusingExercise6.6.1orvariousothermeans.Thefirststepis to identify the singularities, the poles. Here we have one simple pole at s=aand another simple pole at s=−a. By Exercise 6.6.1, the residue at s=ais(1 2)eatand the residue at s=−ais(−1 2)e−at. Then Residues=parenleftbig1 2parenrightbigparenleftbig eat−e−atparenrightbig =sinhat=F(t), (15.215) inagreementwithEq.(15.105). /squaresolid Example 15.12.2 If f(s)=1−e−as s, thenes(t−a)grows exponentially for t<aon the semicircle in the left-hand s-plane, so contour integration and the residue theorem are not applicable. However, we can evaluate theintegralexplicitlyas follows.Welet γ→0 andsubstitute s=iy,so F(t)=1 2πiintegraldisplayγ+i∞ γ−i∞estf(s)=1 2πintegraldisplay∞ −∞bracketleftbig eiyt−eiy(t−a)bracketrightbigdy y. (15.216) UsingtheEuleridentity,onlythesinessurvivethatareoddin yandweobtain F(t)=1 πintegraldisplay∞ −∞bracketleftbiggsinty y−sin(t−a)y ybracketrightbigg . (15.217) Ifk>0,thenintegraltext∞ 0sinky ydygivesπ/2, and it gives −π/2i fk<0.As a consequence, F(t)=0i ft>a>0 and if t<0.If 0<t<a, thenF(t)=1.This can be written compactlyintermsoftheHeavisideunitstepfunction u(t)asfollows: F(t)=u(t)−u(t−a)=  0,t<0, 1,0<t<a, 0,t>a,(15.218) astepfunctionofunitheightandlength a(Fig.15.17). /squaresolid Twogeneralcommentsmaybeinorder.First,thesetwoexampleshardlybegintoshow the usefulness and power of the Bromwich integral. It is always available for inverting a complicatedtransform whenthetablesproveinadequate. Second, this derivation is not presented as a rigorous one. Rather, it is given more as a plausibility argument, although it can be made rigorous. The determination of the in- verse transform is somewhat similar to the solution of a differential equation. It makes 15.12 Inverse Laplace Transform 997 FIGURE 15.17 Finite-lengthstepfunction u(t)−u(t−a). little difference how you get the solution. Guess at it if you want. The solution can al- ways be checked by substitution back into the original differential equation. Similarly, F(t)can(and,tocheckoncarelesserrors,should)becheckedbydeterminingwhether,by Eq.(15.99), Lbraceleftbig F(t)bracerightbig =f(s). Two alternate derivations of the Bromwich integral are the subjects of Exercises 15.12.1 and15.12.2. As a final illustration of the use of the Laplace inverse transform, we have some results fromtheworkofBrillouinandSommerfeld(1914)inelectromagnetictheory. Example 15.12.3 VELOCITY OF ELECTROMAGNETIC WAVES IN A DISPERSIVE MEDIUM Thegroupvelocity uof travelingwavesisrelatedtothephasevelocity vbytheequation u=v−λdv dλ. (15.219) Hereλis the wavelength. In the vicinity of an absorption line (resonance), dv/dλmay be sufficiently negative so that u>c(Fig. 15.18). The question immediately arises whether a signal can be transmitted faster than c, the velocity of light in vacuum. This question, which assumes that such a group velocity is meaningful, is of fundamental importance to thetheoryofspecialrelativity. We needasolutiontothewaveequation ∂2ψ ∂x2=1 v2∂2ψ ∂t2, (15.220) correspondingtoaharmonicvibrationstartingattheoriginattimezero.Sinceourmedium isdispersive, visafunctionoftheangularfrequency.Imagine,forinstance,aplanewave, angular frequency ω, incident on a shutter at the origin. At t=0 the shutter is (instanta- neously)opened,andthewaveis permittedtoadvancealongthepositive x-axis. 998 Chapter 15 Integral Transforms FIGURE 15.18Opticaldispersion. Let us then build up a solution starting at x=0. It is convenient to use the Cauchy integralformula,Eq.(6.43), ψ(0,t)=1 2πicontintegraldisplaye−izt z−z0dz=e−iz0t (for a contour encircling z=z0in the positive sense). Using s=−izandz0=ω,w e obtain ψ(0,t)=1 2πiintegraldisplayγ+i∞ γ−i∞est s+iωds=braceleftbigg0,t <0, e−iωt,t>0.(15.221) To be complete, the loop integral is along the vertical line ℜ(s)=γandan infinite semi- circle, as shown in Fig. 15.19. The location of the infinite semicircle is chosen so that the integral over it vanishes. This means a semicircle in the left half-plane for t>0 and the residue is enclosed. For t<0 we pick the right half-plane and no singularity is enclosed. Thefactthatthisisjust theBromwichintegralmaybeverifiedbynotingthat F(t)=braceleftbigg0,t <0, e−iωt,t>0(15.222) FIGURE 15.19Possibleclosedcontours. 15.12 Inverse Laplace Transform 999 andapplyingtheLaplacetransform.Thetransformedfunction f(s)becomes f(s)=1 s+iω. (15.223) Our Cauchy–Bromwich integral provides us with the time dependence of a signal leav- ingtheoriginat t=0.Toincludethespacedependence,wenotethat es(t−x/v) satisfiesthewaveequation.Withthisasaclue,wereplace tbyt−x/vandwriteasolution: ψ(x,t)=1 2πiintegraldisplayγ+i∞ γ−i∞es(t−x/v) s+iωds. (15.224) It was seen in the derivation of the Bromwich integral that our variable sreplaces the ω oftheFouriertransformation.Hencethewavevelocity vmaybecomeafunctionof s,that is,v(s).Itsparticularformneednotconcernushere.Weneedonlytheproperty v≤cand lim |s|→∞v(s)=constant,c . (15.225) ThisissuggestedbytheasymptoticbehaviorofthecurveontherightsideofFig.15.18.24 EvaluatingEq.(15.225)bythecalculusofresidues,wemayclosethepathofintegration byasemicircleintherighthalf-plane,provided t−x c<0. Hence ψ(x,t)=0,t−x c<0, (15.226) which means that the velocity of our signal cannot exceed the velocity of light in the vac- uum,c.Thissimplebutverysignificantresultwas extendedbySommerfeldandBrillouin toshowjusthowthewaveadvancedinthedispersivemedium. /squaresolid Summary — Inversion of Laplace Transform •Direct use of tables, Table 15.2, and references; use of partial fractions (Section 15.8) andtheoperationaltheoremsofTable15.1. •Bromwichintegral,Eq. (15.212),andthecalculusofresidues. •Numericalinversion,seetheAdditionalReadings. 24Equation (15.225) follows rigorously from the theory of anomalous dispersion. See also the Kronig–Kramers optical disper- sion relations of Section 7.2. 1000 Chapter 15 Integral Transforms Table 15.1 LaplaceTransformOperations Operations Equation 1. Laplacetransform f(s)=L{F(t)}=integraldisplay∞ 0e−stF(t)dt (15.99) 2. Transform of derivative sf(s)−F(+0)=L{F′(t)} (15.123) s2f(s)−sF(+0)−F′(+0)=L{F′′(t)}(15.124) 3. Transform of integral1 sf(s)=Lbraceleftbiggintegraldisplayt 0F(x)dxbracerightbigg (Exercise 15.11.1) 4. Substitution f(s−a)=L{eatF(t)} (15.152) 5. Translation e−bsf(s)=L{F(t−b)} (15.164) 6. Derivativeof transform f(n)(s)=L{(−t)nF(t)} (15.173) 7. Integral of transformintegraldisplay∞ sf(x)dx=LbraceleftbiggF(t) tbracerightbigg (15.189) 8. Convolution f1(s)f2(s)=Lbraceleftbiggintegraldisplayt 0F1(t−z)F2(z)dzbracerightbigg (15.193) 9. Inverse transform, Bromwich integral1 2πiintegraldisplayγ+i∞ γ−i∞estf(s)ds=F(t) (15.212) Exercises 15.12.1 DerivetheBromwichintegralfromCauchy’sintegralformula. Hint.Applytheinversetransform L−1to f(s)=1 2πilimα→∞integraldisplayγ+iα γ−iαf(z) s−zdz, wheref(z)is analyticforℜ(z)≥γ. 15.12.2 Startingwith 1 2πiintegraldisplayγ+i∞ γ−i∞estf(s)ds, showthatbyintroducing f(s)=integraldisplay∞ 0e−szF(z)dz, we can convert one integral into the Fourier representation of a Dirac delta function. FromthisderivetheinverseLaplacetransform. 15.12.3 Derive the Laplace transformation convolution theorem by use of the Bromwich inte- gral. 15.12.4 Find L−1braceleftbiggs s2−k2bracerightbigg (a) byapartialfractionexpansion. (b) Repeat,usingtheBromwichintegral. 15.12 Inverse Laplace Transform 1001 Table 15.2 LaplaceTransforms f(s) F(t) Limitation Equation 1. 1 δ(t) Singularity at+0 (15.141) 2.1 s1 s>0 (15.102) 3.n! sn+1tns>0 (15.108) n>−1 4.1 s−kekts>k (15.103) 5.1 (s−k)2tekts>k (15.175) 6.s s2−k2coshkt s>k (15.105) 7.k s2−k2sinhkt s>k (15.105) 8.s s2+k2coskt s> 0 (15.107) 9.k s2+k2sinkt s> 0 (15.107) 10.s−a (s−a)2+k2eatcoskt s>a (15.153) 11.k (s−a)2+k2eatsinkt s>a (15.153) 12.s2−k2 (s2+k2)2tcoskt s> 0 (Exercise 15.10.19) 13.2ks (s2+k2)2tsinkt s> 0 (Exercise 15.10.19) 14.(s2+a2)−1/2J0(at) s > 0 (15.185) 15.(s2−a2)−1/2I0(at) s >a (Exercise 15.10.9) 16.1 acot−1parenleftbiggs aparenrightbigg j0(at) s > 0 (Exercise 15.10.10) 17.1 2alns+a s−a 1 acoth−1parenleftbiggs aparenrightbigg  i0(at) s >a (Exercise 15.10.10) 18.(s−a)n sn+1Ln(at) s > 0 (Exercise 15.10.12) 19.1 sln(s+1)E 1(x)=−Ei(−x) s> 0 (Exercise 15.10.13) 20.lns s−lnt−γs >0 (Exercise 15.12.9) Amore extensivetableof LaplacetransformsappearsinChapter 29ofAMS-55(see footnote4 inChapter5 forthereference). 15.12.5 Find L−1braceleftbiggk2 s(s2+k2)bracerightbigg 1002 Chapter 15 Integral Transforms (a) byusingapartialfractionexpansion. (b) Repeatusingtheconvolutiontheorem. (c) RepeatusingtheBromwichintegral. ANS.F(t)=1−coskt. 15.12.6 Use the Bromwich integral to find the function whose transform is f(s)=s−1/2.N o t e thatf(s)hasabranchpointat s=0.Thenegative x-axismaybetakenasa cutline. ANS.F(t)=(πt)−1/2. 15.12.7 Showthat L−1braceleftbigparenleftbig s2+1parenrightbig−1/2bracerightbig =J0(t) byevaluationof theBromwichintegral. Hint. Convert your Bromwich integral into an integral representation of J0(t).F i g - ure15.20showsapossiblecontour. 15.12.8 EvaluatetheinverseLaplacetransform L−1braceleftbigparenleftbig s2−a2parenrightbig−1/2bracerightbig byeachofthefollowingmethods: (a) Expansioninaseries andterm-by-terminversion. (b) DirectevaluationoftheBromwichintegral. (c) ChangeofvariableintheBromwichintegral: s=(a/2)(z+z−1). FIGURE 15.20Apossible contourfor theinversionof J0(t). 15.12 Additional Readings 1003 15.12.9 Showthat L−1braceleftbigglns sbracerightbigg =−lnt−γ, whereγ=0.5772...,theEuler–Mascheroniconstant. 15.12.10 EvaluatetheBromwichintegralfor f(s)=s (s2+a2)2. 15.12.11 Heavisideexpansiontheorem .If thetransform f(s)maybewrittenas aratio f(s)=g(s) h(s), whereg(s)andh(s)areanalyticfunctions, h(s)havingsimple,isolatedzerosat s=si, showthat F(t)=L−1braceleftbiggg(s) h(s)bracerightbigg =summationdisplay ig(si) h′(si)esit. Hint.SeeExercise6.6.2. 15.12.12 Using the Bromwich integral, invert f(s)=s−2e−ks. Express F(t)=L−1{f(s)}in termsofthe(shifted) unitstepfunction u(t−k). ANS.F(t)=(t−k)u(t−k). 15.12.13 Youhavea Laplacetransform: f(s)=1 (s+a)(s+b),a/negationslash=b. Invertthistransform byeachof threemethods: (a) Partialfractionsanduse oftables. (b) Convolutiontheorem. (c) Bromwichintegral. ANS.F(t)=e−bt−e−at a−b,a/negationslash=b. AdditionalReadings Champeney, D. C., Fourier Transforms and Their Physical Applications. New York: Academic Press (1973). Fourier transforms are developed in a careful, easy-to-follow manner. Approximately 60% of the book is devoted to applications of interest in physics and engineering. Erdelyi, A.,W. Magnus, F. Oberhettinger, and F. G. Tricomi, Tables of Integral Transforms ,2v o l s .N e wY o r k : McGraw–Hill (1954). This text contains extensive tables of Fourier sine, cosine, and exponential transforms, Laplace and inverse Laplace transforms, Mellin and inverse Mellin transforms, Hankel transforms, and other, more specializedintegral transforms. 1004 Chapter 15 Integral Transforms Hanna, J. R., Fourier Series and Integrals of Boundary Value Problems . Somerset, NJ: Wiley (1990). This book is a broad treatment of the Fourier solution of boundary value problems. The concepts of convergence and completeness aregiven careful attention. Jeffreys,H.,andB.S.Jeffreys, MethodsofMathematicalPhysics ,3rded.Cambridge,UK:CambridgeUniversity Press (1972). Krylov, V. I., and N. S. Skoblya, Handbook of Numerical Inversion of Laplace Transform. Jerusalem: Israel Program for ScientificTranslations (1969). Lepage, W. R., Complex Variables and the Laplace Transform for Engineers . New York: McGraw-Hill (1961); New York: Dover (1980). A complex variable analysis that is carefully developed and then applied to Fourier and Laplacetransforms. It is written to be readby students, but intended for the serious student. McCollum, P. A., and B. F. Brown, Laplace Transform Tables and Theorems . New York: Holt, Rinehart and Winston (1965). Miles, J. W., Integral Transforms in Applied Mathematics . Cambridge, UK: Cambridge University Press (1971). Thisisabriefbutinterestingandusefultreatmentfortheadvancedundergraduate.Itemphasizesapplications ratherthan abstractmathematical theory. Papoulis, A., The Fourier Integral and Its Applications . New York: McGraw-Hill (1962). This is a rigorous development ofFourier and Laplacetransforms and has extensive applications in scienceand engineering. Roberts, G.E.,andH.Kaufman, Table of Laplace Transforms . Philadelphia: Saunders (1966). Sneddon, I. N., Fourier Transforms . New York: McGraw-Hill (1951), reprinted, Dover (1995). A detailed com- prehensive treatment, this book is loaded with applications to a wide variety of fields of modern and classical physics. Sneddon, I. H., The Useof Integral Transforms . NewYork: McGraw-Hill(1972). Written for students in science and engineering in terms they can understand, this book covers all the integral transforms mentioned in this chapteras wellasinseveral others. Manyapplications areincluded. Van der Pol, B., and H. Bremmer, Operational Calculus Based on the Two-sided Laplace Integral , 3rd ed. Cam- bridge, UK: Cambridge University Press (1987). Here is a development based on the integral range −∞to +∞, rather than the useful 0 to ∞. Chapter V contains a detailed study of the Dirac delta function (impulse function). W o l f ,K .B . , Integral Transforms in Science and Engineering . New York: Plenum Press (1979). This book is a very comprehensive treatment of integral transforms and their applications. CHAPTER 16 INTEGRAL EQUATIONS 16.1 I NTRODUCTION Withtheexceptionoftheintegraltransformsofthelastchapter,wehavebeenconsidering relations between the unknown function ϕ(x)and one or more of its derivatives. We now proceed to investigate equations containing the unknown function within an integral. As withdifferentialequations,weshallconfineourattentiontolinearrelations,linearintegral equations.Integralequationsareclassifiedintwoways: •Ifthelimitsofintegrationarefixed ,wecalltheequationa Fredholm equation;if one limitis variable ,itisaVolterra equation. •Iftheunknownfunction appearsonlyundertheintegral sign,welabelit firstkind . If itappearsboth insideandoutside theintegral,itis labeled secondkind . Definitions Symbolically,wehavea Fredholmequationofthefirst kind , f(x)=integraldisplayb aK(x,t)ϕ(t)dt; (16.1) theFredholmequationofthesecondkind ,withλbeingtheeigenvalue, ϕ(x)=f(x)+λintegraldisplayb aK(x,t)ϕ(t)dt; (16.2) theVolterraequationofthefirst kind , f(x)=integraldisplayx aK(x,t)ϕ(t)dt; (16.3) 1005 1006 Chapter 16 Integral Equations andtheVolterraequationof thesecondkind , ϕ(x)=f(x)+integraldisplayx aK(x,t)ϕ(t)dt. (16.4) In all four cases ϕ(t)is the unknown function. K(x,t), which we call the kernel, and f(x)areassumedtobeknown.When f(x)=0,theequationis saidtobe homogeneous . Why do we bother about integral equations? After all, the differential equations have done a rather good job of describing our physical world so far. There are several reasons forintroducingintegralequationshere. We have placed considerable emphasis on the solution of differential equations subject to particular boundary conditions . For instance, the boundary condition at r=0 deter- mines whether the Neumann function Nn(r)is present when Bessel’s equation is solved. The boundary condition for r→∞determines whether the In(r)is present in our solu- tion of the modified Bessel equation. The integral equation relates the unknown function notonly toits valuesat neighboringpoints(derivatives) but alsoto its values throughouta region, including the boundary. In a very real sense the boundary conditions are built into theintegralequationratherthanimposedatthefinalstageofthesolution.Itcanbeseenin Section10.5,wherekernelsareconstructed,thattheformofthekerneldependsontheval- uesontheboundary.Theintegralequation,then,iscompactandmayturnouttobeamore convenientorpowerfulformthanthedifferentialequation.Mathematicalproblemssuchas existence, uniqueness, and completeness may often be handled more easily and elegantly in integral form. Finally, whether or not we like it, there are some problems, such as some diffusion and transport phenomena, that cannot be represented by differential equations. If we wish to solve such problems, we are forced to handle integral equations. Finally, an integralequationmayalsoappearasamatterofdeliberatechoicebasedonconvenienceor theneedforthemathematicalpowerof anintegralequationformulation. Example 16.1.1 MOMENTUM REPRESENTATION IN QUANTUM MECHANICS TheSchrödingerequation(inordinaryspacerepresentation)is −¯h2 2m∇2ψ(r)+V(r)ψ(r)=Eψ(r), (16.5) or parenleftbig ∇2+a2parenrightbig ψ(r)=v(r)ψ(r), (16.6) where a2=2m ¯h2E, v( r)=2m ¯h2V(r). (16.7) If wegeneralizeEq. (16.6)to parenleftbig ∇2+a2parenrightbig ψ(r)=integraldisplay v(r,r′)ψ(r′)d3r′, (16.8) then,forthespecialcaseof v(r,r′)=v(r′)δ(r−r′), (16.9) 16.1 Introduction 1007 alocalinteraction,Eq.(16.8)reducestoEq.(16.6).ConsidertheFouriertransformpair ψ and/Psi1(comparefootnote9inSection15.6): /Psi1(k)=1 (2π)3/2integraldisplay ψ(r)e−ik·rd3r, ψ( r)=1 (2π)3/2integraldisplay /Psi1(k)eik·rd3k,(16.10) withtheabbreviation pfor momentumso that p ¯h=k(wavenumber). (16.11) MultiplyingEq. (16.8)bytheplane-wave e−ik·r, weobtain integraldisplay e−ik·rparenleftbig ∇2+a2parenrightbig ψ(r)d3r=integraldisplay d3re−ik·rintegraldisplay v(r,r′)ψ(r′)d3r′. (16.12) Note that the ∇2on the left operates only on the ψ(r). Integrating the left-hand side by partsandsubstitutingEq.(16.10) for ψ(r′)ontheright,weget integraldisplayparenleftbig −k2+a2parenrightbig ψ(r)e−ik·rd3r=(2π)3/2parenleftbig −k2+a2parenrightbig /Psi1(k) =1 (2π)3/2integraldisplayintegraldisplayintegraldisplay v(r,r′)/Psi1(k′)e−i(k·r−k′·r′)d3r′d3rd3k′.(16.13) If weuse f(k,k′)=1 (2π)3/2integraldisplayintegraldisplay v(r,r′)e−i(k·r−k′·r′)d3r′d3r, (16.14) Eq.(16.13) becomes parenleftbig −k2+a2parenrightbig /Psi1(k)=integraldisplay f(k,k′)/Psi1(k′)d3k′, (16.15) a Fredholm equation of the second kind in which the parameter a2corresponds to the eigenvalue. Forourspecialbutimportantcaseoflocalinteraction,applicationofEq.(16.9)leadsto f(k,k′)=f(k−k′). (16.16) Thisisourmomentumrepresentation,equivalenttoanordinarystaticinteractionpoten- tialincoordinatespace.Ourmomentumwavefunction /Psi1(k)satisfiestheintegralequation Eq.(16.15).Itmustbeemphasizedthatallthroughherewehaveassumedthattherequired Fourier integrals exist. For a harmonic oscillator potential, V(r)=r2, the required inte- grals would not exist. Equation (16.10) would lead to divergent oscillations and we would havenoEq.(16.15). /squaresolid 1008 Chapter 16 Integral Equations Transformation of a Differential Equation into an Integral Equation Often we find that we have a choice. The physical problem may be represented by a dif- ferential or an integral equation. Let us assume that we have the differential equation and wishtotransformitintoanintegralequation.Startingwitha linearsecond-orderODE y′′+A(x)y′+B(x)y=g(x) (16.17) withinitialconditions y(a)=y0,y′(a)=y′ 0, weintegratetoobtain y′(x)=−integraldisplayx aA(t)y′(t)dt−integraldisplayx aB(t)y(t)dt+integraldisplayx ag(t)dt+y′ 0. (16.18) Integratingthefirst integralontherightbypartsyields y′(x)=−Ay(x)−integraldisplayx a(B−A′)y(t)dt+integraldisplayx ag(t)dt+A(a)y0+y′ 0.(16.19) Notice how the initial conditions are being absorbed into our new version. Integrating a secondtime,weobtain y(x)=−integraldisplayx aAy dx−integraldisplayx aduintegraldisplayu abracketleftbig B(t)−A′(t)bracketrightbig y(t)dt +integraldisplayx aduintegraldisplayu ag(t)dt+bracketleftbig A(a)y0+y′ 0bracketrightbig (x−a)+y0.(16.20) Totransformthisequationintoaneaterform, weusetherelation integraldisplayx aduintegraldisplayu af(t)dt=integraldisplayx a(x−t)f(t)dt. (16.21) This may be verified by differentiating both sides. Since the derivatives are equal, the original expressions can differ only by a constant. Letting x→a, the constant vanishes andEq. (16.21)is established.ApplyingittoEq.(16.20), weobtain y(x)=−integraldisplayx abraceleftbig A(t)+(x−t)bracketleftbig B(t)−A′(t)bracketrightbigbracerightbig y(t)dt +integraldisplayx a(x−t)g(t)dt+bracketleftbig A(a)y0+y′ 0bracketrightbig (x−a)+y0.(16.22) If wenowintroducetheabbreviations K(x,t)=(t−x)bracketleftbig B(t)−A′(t)bracketrightbig −A(t), (16.23) f(x)=integraldisplayx a(x−t)g(t)dt+bracketleftbig A(a)y0+y′ 0bracketrightbig (x−a)+y0, 16.1 Introduction 1009 Eq.(16.22) becomes y(x)=f(x)+integraldisplayx aK(x,t)y(t)dt, (16.24) which is a Volterra equation of the second kind. This reformulation as a Volterra integral equationoffers certainadvantagesininvestigatingquestionsofexistenceanduniqueness. Example 16.1.2 LINEAR OSCILLATOR EQUATION Asanillustration,considerthelinearoscillatorequation y′′+ω2y=0 (16.25) with y(0)=0,y′(0)=1. Thisyields(comparewithEq. (16.17)) A(x)=0,B(x)=ω2,g(x)=0. Substituting into Eq. (16.22) (or Eqs. (16.23) and (16.24)), we find that the integral equa- tionbecomes y(x)=x+ω2integraldisplayx 0(t−x)y(t)dt. (16.26) •This integral equation, Eq. (16.26), is equivalent to the original differential equation plustheinitialconditions. Acheckshowsthateachformis indeedsatisfiedby y(x)=(1/ω)sinωx. /squaresolid Let us reconsider the linear oscillator equation (16.25) but now with the boundary con- ditions y(0)=0,y(b)=0. Sincey′(0)is notgiven,wemustmodifytheprocedure.Thefirst integrationgives y′=−ω2integraldisplayx 0ydx+y′(0). (16.27) Integratinga secondtimeandagainusingEq. (16.21),wehave y=−ω2integraldisplayx 0(x−t)y(t)dt+y′(0)x. (16.28) Toeliminatetheunknown y′(0), wenowimposethecondition y(b)=0.Thisgives ω2integraldisplayb 0(b−t)y(t)dt=by′(0). (16.29) 1010 Chapter 16 Integral Equations FIGURE 16.1 SubstitutingthisbackintoEq. (16.28), weobtain y(x)=−ω2integraldisplayx 0(x−t)y(t)dt+ω2x bintegraldisplayb 0(b−t)y(t)dt. (16.30) Nowletusbreaktheinterval [0,b]intotwointervals, [0,x]and[x,b]. Since x b(b−t)−(x−t)=t b(b−x), (16.31) wefind y(x)=ω2integraldisplayx 0t b(b−x)y(t)dt+ω2integraldisplayb xx b(b−t)y(t)dt. (16.32) Finally,if wedefineakernel(Fig. 16.1) K(x,t)=  t b(b−x), t<x, x b(b−t), t>x,(16.33) wehave y(x)=ω2integraldisplayb 0K(x,t)y(t)dt, (16.34) ahomogeneousFredholmequationof thesecondkind. Ournewkernel, K(x,t), has someinterestingproperties. 1. It issymmetric, K(x,t)=K(t,x). 2. It iscontinuous,inthesensethat t b(b−x)vextendsinglevextendsinglevextendsingle t=x=x b(b−t)vextendsinglevextendsinglevextendsingle t=x. 3. Itsderivativewithrespectto tisdiscontinuous .Astincreasesthroughthepoint t=x, thereis adiscontinuityof −1i n∂K(x,t)/∂t . AccordingtothesepropertiesinSection9.7weidentify K(x,t)asaGreen’sfunction. 1. Inthetransformationofalinear,second-orderODEintoanintegralequation,theinitial orboundaryconditionsplayadecisiverole.Ifwehave initialconditions(onlyoneend of our interval), the differential equation transforms into a Volterra integral equation. For the case of the linear oscillator equation with boundary conditions (both ends 16.1 Introduction 1011 of our interval), the differential equation leads to a Fredholm integral equation with akernelthatwillbeaGreen’sfunction. 2. Note that the reverse transformation (integral equation to differential equation) is not alwayspossible.Thereexistintegralequationsforwhichnocorrespondingdifferential equationisknown. Exercises 16.1.1 Starting with the ODE, integrate twice and derive the Volterra integral equation corre- spondingto (a)y′′(x)−y(x)=0;y(0)=0,y′(0)=1. ANS.y=integraldisplayx 0(x−t)y(t)dt+x. (b)y′′(x)−y(x)=0;y(0)=1,y′(0)=−1. ANS.y=integraldisplayx 0(x−t)y(t)dt−x+1. Checkyourresults withEq. (16.23). 16.1.2 DeriveaFredholmintegralequationcorrespondingto y′′(x)−y(x)=0,y(1)=1,y(−1)=1, (a) byintegratingtwice, (b) byformingtheGreen’sfunction. ANS.y(x)=1−integraldisplay1 −1K(x,t)y(t)dt , K(x,t)=braceleftBigg1 2(1−x)(t+1), x >t, 1 2(1−t)(x+1), x <t. 16.1.3 (a) Starting with the given answers of Exercise 16.1.1, differentiate and recover the originalODEs andtheboundaryconditions . (b) Repeatfor Exercise16.1.2. 16.1.4 Thegeneralsecond-orderlinearODEwithconstantcoefficientsis y′′(x)+a1y′(x)+a2y(x)=0. Giventheboundaryconditions y(0)=y(1)=0, integratetwiceanddeveloptheintegralequation y(x)=integraldisplay1 0K(x,t)y(t)dt, 1012 Chapter 16 Integral Equations with K(x,t)=braceleftBigg a2t(1−x)+a1(x−1), t <x, a2x(1−t)+a1x, x<t. Note that K(x,t)is symmetric and continuous if a1=0. How is this related to self- adjointnessoftheODE? 16.1.5 Verify thatintegraltextx aintegraltextx af(t)dtdx=integraltextx a(x−t)f(t)dt for allf(t)(for which the integrals exist). 16.1.6 Givenϕ(x)=x−integraltextx 0(t−x)ϕ(t)dt , solve this integral equation by converting it to an ODE(plusboundaryconditions)andsolvingtheODE(byinspection). 16.1.7 ShowthatthehomogeneousVolterraequationof thesecondkind ψ(x)=λintegraldisplayx 0K(x,t)ψ(t)dt hasnosolution(apartfrom thetrivial ψ=0). Hint.Develop a Maclaurin expansion of ψ(x). Assume ψ(x)andK(x,t)are differen- tiablewithrespectto xasneeded. 16.2 I NTEGRAL TRANSFORMS ,GENERATING FUNCTIONS Analogous to differentiation, linear ODEs are solved in Chapter 9. Analogous to integra- tion, there is no general method available for solving integral equations. However, certain special cases may be treated with our integral transforms (Chapter 15). For convenience theseare listedhere. If ψ(x)=1√ 2πintegraldisplay∞ −∞eixtϕ(t)dt, then ϕ(x)=1√ 2πintegraldisplay∞ −∞e−ixtψ(t)dt (Fourier). (16.35) If ψ(x)=integraldisplay∞ 0e−xtϕ(t)dt, then ϕ(x)=1 2πiintegraldisplayγ+i∞ γ−i∞extψ(t)dt (Laplace). (16.36) If ψ(x)=integraldisplay∞ 0tx−1ϕ(t)dt, 16.2 Integral Transforms, Generating Functions 1013 then ϕ(x)=1 2πiintegraldisplayγ+i∞ γ−i∞x−tψ(t)dt (Mellin). (16.37) If ψ(x)=integraldisplay∞ 0tϕ(t)Jν(xt)dt, then ϕ(x)=integraldisplay∞ 0tψ(t)Jν(xt)dt (Hankel). (16.38) Actually the usefulness of the integral transform technique extends a bit beyond these fourratherspecializedforms. Example 16.2.1 FOURIER TRANSFORM SOLUTION Let us consider a Fredholm equation of the first kind with a kernel of the general type k(x−t), f(x)=integraldisplay∞ −∞k(x−t)ϕ(t)dt, (16.39) in which ϕ(t)is our unknown function. Assuming that the needed transforms exist ,w e applytheFourierconvolutiontheorem(Section15.5)toobtain f(x)=integraldisplay∞ −∞K(ω)/Phi1(ω)e−iωxdω. (16.40) Thefunctions K(ω),/Phi1(ω) ,andF(ω)aretheFouriertransformsof k(x),ϕ(x) ,andf(x), respectively. Taking the Fourier transform of both sides of Eq. (16.40), by Eq. (16.35) we have K(ω)/Phi1(ω)=1 2πintegraldisplay∞ −∞f(x)eiωxdx=F(ω)√ 2π. (16.41) Then /Phi1(ω)=1√ 2π·F(ω) K(ω), (16.42) and,usingtheinverseFouriertransform, wehave ϕ(x)=1 2πintegraldisplay∞ −∞F(ω) K(ω)e−iωxdω. (16.43) For a rigorous justification of this result one can follow Morse and Feshbach (see the Additional Readings) (1953) across complex planes. An extension of this transformation solutionappearsasExercise16.2.1. /squaresolid 1014 Chapter 16 Integral Equations Example 16.2.2 GENERALIZED ABELEQUATION ,CONVOLUTION THEOREM ThegeneralizedAbelequationis f(x)=integraldisplayx 0ϕ(t) (x−t)αdt,0<α<1,withbraceleftbiggf(x)known, ϕ(t)unknown .(16.44) TakingtheLaplacetransformofbothsidesof thisequation,weobtain L{f(x)}=Lbraceleftbiggintegraldisplayx 0ϕ(t) (x−t)αdtbracerightbigg =Lbraceleftbig x−αbracerightbig Lbraceleftbig ϕ(x)bracerightbig , (16.45) thelast stepfollowingbytheLaplaceconvolutiontheorem(Section15.11).Then Lbraceleftbig ϕ(x)bracerightbig =s1−αL{f(x)} (−α)!. (16.46) Dividingby s,1weobtain 1 sLbraceleftbig ϕ(x)bracerightbig =s−αL{f(x)} (−α)!=L{xα−1}L{f(x)} (α−1)!(−α)!. (16.47) Combiningthefactorials(Eq.(8.32))andapplyingtheLaplaceconvolutiontheoremagain, wediscoverthat 1 sLbraceleftbig ϕ(x)bracerightbig =sinπα πLbraceleftbiggintegraldisplayx 0f(t) (x−t)1−αdtbracerightbigg . (16.48) InvertingwiththeaidofExercise15.11.1,weget integraldisplayx 0ϕ(t)dt=sinπα πintegraldisplayx 0f(t) (x−t)1−αdt, (16.49) andfinally,bydifferentiating, ϕ(x)=sinπα πd dxintegraldisplayx 0f(t) (x−t)1−αdt. (16.50) /squaresolid Generating Functions Occasionally, the reader may encounter integral equations that involve generating func- tions.Supposewehavetheadmittedlyspecialcase f(x)=integraldisplay1 −1ϕ(t) (1−2xt+x2)1/2dt,−1≤x≤1. (16.51) Wenoticetwoimportantfeatures: 1.(1−2xt+x2)−1/2generatestheLegendrepolynomials. 2.[−1,1]istheorthogonalityintervalfor theLegendrepolynomials. 1s1−αdoes not have an inverse for 0 <α<1. 16.2 Integral Transforms, Generating Functions 1015 Ifwenowexpandthedenominator(property1)andassumethatourunknown ϕ(t)may bewrittenasaseries ofthesesameLegendrepolynomials, f(x)=integraldisplay1 −1∞summationdisplay n=0anPn(t)∞summationdisplay r=0Pr(t)xrdt. (16.52) UtilizingtheorthogonalityoftheLegendrepolynomials(property2), weobtain f(x)=∞summationdisplay r=02ar 2r+1xr. (16.53) Wemayidentifythe anbydifferentiating ntimesandthensetting x=0: f(n)(0)=n!2 2n+1an. (16.54) Hence ϕ(t)=∞summationdisplay n=02n+1 2f(n)(0) n!Pn(t). (16.55) Similar results may be obtained with the other generating functions (compare Exer- cise7.1.6). •This technique of expanding in a series of special functions is always available. It is worth a try whenever the expansion is possible (and convenient) and the interval is appropriate. Exercises 16.2.1 ThekernelofaFredholmequationofthesecondkind, ϕ(x)=f(x)+λintegraldisplay∞ −∞K(x,t)ϕ(t)dt, isof theform k(x−t).2Assumingthattherequiredtransforms exist,showthat ϕ(x)=1√ 2πintegraldisplay∞ −∞F(t)e−ixtdt 1−√ 2πλK(t). F(t)andK(t)aretheFouriertransformsof f(x)andk(x), respectively. 16.2.2 ThekernelofaVolterraequationof thefirst kind, f(x)=integraldisplayx 0K(x,t)ϕ(t)dt, 2Thiskernelandarange 0 ≤x<∞arethecharacteristicsofintegralequationsoftheWiener–Hopftype.Detailswillbefound in Chapter 8 of Morse and Feshbach (1953); seethe Additional Readings. 1016 Chapter 16 Integral Equations hastheform k(x−t).Assumingthattherequiredtransforms exist,showthat ϕ(x)=1 2πiintegraldisplayγ+i∞ γ−i∞F(s) K(s)exsds. F(s)andK(s)aretheLaplacetransforms of f(x)andk(x), respectively. 16.2.3 ThekernelofaVolterraequationof thesecondkind, ϕ(x)=f(x)+λintegraldisplayx 0K(x,t)ϕ(t)dt, hastheform k(x−t).Assumingthattherequiredtransforms exist,showthat ϕ(x)=1 2πiintegraldisplayγ+i∞ γ−i∞F(s) 1−λK(s)exsds. 16.2.4 UsingtheLaplacetransform solution(Exercise16.2.3), solve (a)ϕ(x)=x+integraldisplayx 0(t−x)ϕ(t)dt . ANS.ϕ(x)=sinx. (b)ϕ(x)=x−integraldisplayx 0(t−x)ϕ(t)dt . ANS.ϕ(x)=sinhx. Checkyourresults bysubstitutingbackintotheoriginalintegralequations. 16.2.5 Reformulate the equations of Example 16.2.1 (Eqs. (16.39) to (16.43)), using Fourier cosinetransforms. 16.2.6 GiventheFredholmintegralequation, e−x2=integraldisplay∞ −∞e−(x−t)2ϕ(t)dt, applytheFourierconvolutiontechniqueof Example16.2.1tosolvefor ϕ(t). 16.2.7 SolveAbel’sequation, f(x)=integraldisplayx 0ϕ(t) (x−t)αdt,0<α<1, bythefollowingmethod: (a) Multiply both sides by (z−x)α−1and integrate with respect to xover the range 0≤x≤z. (b) Reverse the order of integration and evaluate the integral on the right-hand side (withrespectto x)bythebetafunction. 16.2 Integral Transforms, Generating Functions 1017 Note.integraldisplayz tdx (z−x)1−α(x−t)α=B(1−α,α)=(−α)!(α−1)!=π sinπα. 16.2.8 GiventhegeneralizedAbelequationwith f(x)=1, 1=integraldisplayx 0ϕ(t) (x−t)αdt,0<α<1, solvefor ϕ(t)andverifythat ϕ(t)is asolutionof thegivenequation. ANS.ϕ(t)=sinπα πtα−1. 16.2.9 AFredholmequationof thefirst kindhas akernel e−(x−t)2: f(x)=integraldisplay∞ −∞e−(x−t)2ϕ(t)dt. Showthatthesolutionis ϕ(x)=1√π∞summationdisplay π=0f(n)(0) 2nn!Hn(x), inwhich Hn(x)is annth-orderHermitepolynomial. 16.2.10 Solvetheintegralequation f(x)=integraldisplay1 −1ϕ(t) (1−2xt+x2)1/2dt,−1≤x≤1, for theunknownfunction ϕ(t)if (a)f(x)=x2s,(b)f( x)=x2s+1. ANS.(a) ϕ(t)=4s+1 2P2s(t),(b)ϕ(t)=4s+3 2P2s+1(t). 16.2.11 AKirchhoffdiffractiontheoryanalysisofalaserleadstotheintegralequation v(r2)=γintegraldisplayintegraldisplay K(r1,r2)v(r1)dA. The unknown, v(r1), gives the geometric distribution of the radiation field over one mirror surface; the range of integration is over the surface of that mirror. For square confocalsphericalmirrors theintegralequationbecomes v(x2,y2)=−iγeikb λbintegraldisplaya −aintegraldisplaya −ae−(ik/b)(x 1x2+y1y2)v(x1,y1)dx1dy1, in which bis the centerline distance between the laser mirrors. This can be put in a somewhatsimplerformbythesubstitutions kx2 i b=ξ2 i,ky2 i b=η2 i,andka2 b=2πa2 λb=α2. (a) Showthatthevariablesseparateandwegettwo integralequations. 1018 Chapter 16 Integral Equations (b) Show that the new limits, ±α, may be approximated by ±∞for a mirror dimen- siona≫λ. (c) Solvetheresultingintegralequations. 16.3 N EUMANN SERIES ,SEPARABLE (DEGENERATE ) KERNELS Many and probably most integral equations cannot be solved by the specialized integral transform techniques of the preceding section. Here we develop three rather general tech- niques for solving integral equations. The first, due largely to Neumann, Liouville, and Volterra, develops the unknown function ϕ(x)as a power series in λ, whereλis a given constant.Themethodis applicablewhenevertheseries converges. Thesecondmethodissomewhatrestrictedbecauseitrequiresthatthetwovariablesap- pearingin the kernel K(x,t)be separable. However,there are two major rewards: (1) The relation between an integral equation and a set of simultaneous linear algebraic equations is shown explicitly, and (2) the method leads to eigenvalues and eigenfunctions—in close analogytoSection3.5. Third, a technique for numerical solution of Fredholm equations of both the first and secondkindisoutlined.Theproblemposedbyill-conditionedmatricesis emphasized. Neumann Series We solve a linear integral equation of the second kind by successive approximations; our integralequationis theFredholmequation, ϕ(x)=f(x)+λintegraldisplayb aK(x,t)ϕ(t)dt, (16.56) in which f(x)/negationslash=0. If the upper limit of the integral is a variable (Volterra equation), the followingdevelopmentwillstillhold,butwithminormodifications.Letustry(thereisno guarantee thatitwillwork)toapproximateourunknownfunctionby ϕ(x)≈ϕ0(x)=f(x). (16.57) This choice is not mandatory. If you can make a better guess, go ahead and guess. The choice here is equivalent to saying that the integral or the constant λis small. To improve thisfirstcrudeapproximation,wefeed ϕ0(x)backintotheintegral,Eq. (16.56),andget ϕ1(x)=f(x)+λintegraldisplayb aK(x,t)f(t)dt. (16.58) Repeatingthisprocessofsubstitutingthenew ϕn(x)backintoEq.(16.56),wedevelopthe sequence ϕ2(x)=f(x)+λintegraldisplayb aK(x,t1)f(t1)dt1 +λ2integraldisplayb aintegraldisplayb aK(x,t1)K(t1,t2)f(t2)dt2dt1 (16.59) 16.3 Neumann Series, Separable (Degenerate) Kernels 1019 and ϕn(x)=nsummationdisplay i=0λiui(x), (16.60) where u0(x)=f(x), u1(x)=integraldisplayb aK(x,t1)f(t1)dt1, (16.61) u2(x)=integraldisplayb aintegraldisplayb aK(x,t1)K(t1,t2)f(t2)dt2dt1, un(x)=integraldisplayintegraldisplay ···integraldisplay K(x,t1)K(t1,t2)···K(tn−1,tn)·f(tn)dtn···dt1. Weexpectthatoursolution ϕ(x)willbe ϕ(x)=limn→∞ϕn(x)=limn→∞nsummationdisplay i=0λiui(x), (16.62) providedthatourinfiniteseriesconverges .Wemayconvenientlychecktheconvergence bytheCauchyratiotest,Section5.2, notingthat vextendsinglevextendsingleλnun(x)vextendsinglevextendsingle≤vextendsinglevextendsingleλnvextendsinglevextendsingle·|f|max·|K|n max·|b−a|n, (16.63) using|f|maxto represent the maximum value of|f(x)|in the interval [a,b]and|K|max to represent the maximum value of |K(x,t)|in its domain in the x,t-plane. We have convergenceif |λ|·|K|max·|b−a|<1. (16.64) Notethat λ|un(max)|isbeingusedasa comparison series.Ifitconverges,ouractualseries must converge. If this condition is not satisfied, we may or may not have convergence. A more sensitive test is required. Of course, even if the Neumann series diverges, there stillmaybeasolutionobtainablebyanothermethod. To see what has been done with this iterative manipulation, we may find it helpful to rewrite the Neumann series solution, Eq. (16.59), in operator form. We start by rewriting Eq.(16.56) as ϕ=λKϕ+f, whereKrepresents the integraloperatorintegraltextb aK(x,t)[]dt.Solvingfor ϕ, weobtain ϕ=(1−λK)−1f. Binomial expansion leads to Eq. (16.59). The convergence of the Neumann series is a demonstrationthattheinverseoperator (1−λK)−1exists. 1020 Chapter 16 Integral Equations Example 16.3.1 NEUMANN SERIES SOLUTION ToillustratetheNeumannmethod,weconsidertheintegralequation ϕ(x)=x+1 2integraldisplay1 −1(t−x)ϕ(t)dt. (16.65) TostarttheNeumannseries, wetake ϕ0(x)=x. (16.66) Then ϕ1(x)=x+1 2integraldisplay1 −1(t−x)tdt=x+1 2parenleftbigg1 3t3−1 2t2xparenrightbiggvextendsinglevextendsinglevextendsinglevextendsingle1 −1=x+1 3. Substituting ϕ1(x)backintoEq. (16.65), weget ϕ2(x)=x+1 2integraldisplay1 −1(t−x)tdt+1 2integraldisplay1 −1(t−x)1 3dt=x+1 3−x 3. Continuingthisprocess ofsubstitutingbackintoEq. (16.65),weobtain ϕ3(x)=x+1 3−x 3−1 32, andbyinduction ϕ2n(x)=x+nsummationdisplay s=1(−1)s−13−s−xnsummationdisplay s=1(−1)s−13−s. (16.67) Lettingn→∞, weget ϕ(x)=3 4x+1 4. (16.68) This solution can (and should) be checked by substituting back into the original equation, Eq.(16.65). /squaresolid It is interesting to note that our series converged easily even though Eq. (16.64) is not satisfied in this particular case. Actually Eq. (16.64) is a rather crude upper bound on λ. It can be shown that a necessary and sufficient condition for the convergence of our series solution is that |λ|<|λe|, whereλeis the eigenvalue of smallest magnitude of the cor- responding homogeneous equation [f(x)=0)]. For this particular example λe=√ 3/2. Clearly,λ=1 2<λe=√ 3/2. Oneapproachtothecalculationoftime-dependentperturbationsinquantummechanics startswiththeintegralequationfor theevolutionoperator U(t,t0)=1−i ¯hintegraldisplayt t0V(t1)U(t1,t0)dt1. (16.69a) Iterationleadsto U(t,t0)=1−i ¯hintegraldisplayt t0V(t1)dt1+parenleftbiggi ¯hparenrightbigg2integraldisplayt t0integraldisplayt1 t0V(t1)V(t2)dt2dt1+···.(16.69b) 16.3 Neumann Series, Separable (Degenerate) Kernels 1021 Theevolutionoperatorisobtainedasaseriesofmultipleintegralsoftheperturbingpoten- tialV(t), closely analogous to the Neumann series, Eq. (16.60). For V=V0, independent oft, the evolution operator becomes (see Exercise 3.4.13, replace t→/Delta1t, and construct Ufrom productsof T(t+/Delta1t,t)asinEq. (4.26)) U(t1,t0)=expbracketleftbigg −i ¯h(t−t0)V0bracketrightbigg . AsecondandsimilarrelationshipbetweentheNeumannseriesandquantummechanics appears when the Schrödinger wave equation for scattering is reformulated as an integral equation. The first term in a Neumann series solution is the incident (unperturbed) wave. Thesecondtermis thefirst-order Bornapproximation,Eq. (9.203b)of Section9.7. The Neumann method may also be applied to Volterra integral equations of the second kind, Eq. (16.4) or Eq. (16.56) with the fixed upper limit, b, replaced by a variable, x.I n the Volterra case the Neumann series converges for all λas long as the kernel is square integrable. Separable Kernel Thetechniqueofreplacingourintegralequationbysimultaneousalgebraicequationsmay alsobeusedwheneverourkernel K(x,t)is separable,inthesensethat K(x,t)=nsummationdisplay j=1Mj(x)Nj(t), (16.70) wheren, the upper limit of the sum, is finite. Such kernels are sometimes called degener- ate. Our class of separable kernels includes all polynomials and many of the elementary transcendentalfunctions;thatis, cos(t−x)=costcosx+sintsinx. (16.70a) If Eq. (16.70) is satisfied, substitution into the Fredholm equation of the second kind, Eq. (16.2), yields ϕ(x)=f(x)+λnsummationdisplay j=1Mj(x)integraldisplayb aNj(t)ϕ(t)dt, (16.71) interchangingintegrationandsummation.Now,theintegralwithrespectto tisaconstant, integraldisplayb aNj(t)ϕ(t)dt=cj. (16.72) HenceEq.(16.71) becomes ϕ(x)=f(x)+λnsummationdisplay j=1cjMj(x). (16.73) This gives us ϕ(x), our solution, once the constants cihave been determined. Equa- tion (16.73) further tells us the form of ϕ(x):f(x), plus a linear combination of the x-dependentfactors oftheseparablekernel. 1022 Chapter 16 Integral Equations We may find ciby multiplying Eq. (16.73) by Ni(x)and integrating to eliminate the x-dependence.UseofEq. (16.72)yields ci=bi+λnsummationdisplay j=1aijcj, (16.74) where bi=integraldisplayb aNi(x)f(x)dx, a ij=integraldisplayb aNi(x)Mj(x)dx. (16.75) It isperhapshelpfultowriteEq. (16.74)inmatrixform, with A=(aij): b=c−λAc=(1−λA)c, (16.76a) or3 c=(1−λA)−1b. (16.76b) Equation(16.76a)isequivalenttoasetof simultaneouslinearalgebraicequations (1−λa11)c1−λa12c2−λa13c3−···=b1, −λa21c1+(1−λa22)c2−λa23c3−···=b2, (16.77) −λa31c1−λa32c2+(1−λa33)c3−···=b3,andso on . If our integral equation is homogeneous, [f(x)=0], thenb=0. To get a solution, we set thedeterminantofthecoefficientsof ciequaltozero, |1−λA|=0, (16.78) exactly as in Section 3.5. The roots of Eq. (16.78) yield our eigenvalues. Substituting into (1−λA)c=0,wefindthe ci,andthenEq. (16.73)givesoursolution. Example 16.3.2 Toillustratethistechniquefordeterminingeigenvaluesandeigenfunctionsofthehomoge- neousFredholmequation,weconsiderthecase ϕ(x)=λintegraldisplay1 −1(t+x)ϕ(t)dt. (16.79) Here(comparewithEqs. (16.71)and(16.77)) M1=1,M 2(x)=x, N1(t)=t, N 2=1. Equation(16.75)yields a11=a22=0,a 12=2 3,a 21=2;b1=0=b2. 3Noticethe similarity to the operator form ofthe Neumann series. 16.3 Neumann Series, Separable (Degenerate) Kernels 1023 Equation(16.78),our secularequation,becomes vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle1−2λ 3 −2λ1vextendsinglevextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=0. (16.80) Expanding,weobtain 1−4λ2 3=0,λ=±√ 3 2. (16.81) Substitutingtheeigenvalues λ=±√ 3/2 intoEq.(16.76), wehave c1∓c2√ 3=0. (16.82) Finally,witha choiceof c1=1,Eq.(16.73) gives ϕ1(x)=√ 3 2(1+√ 3x), λ=√ 3 2, (16.83) ϕ2(x)=−√ 3 2(1−√ 3x), λ=−√ 3 2. (16.84) Sinceourequationishomogeneous,thenormalizationof ϕ(x)is arbitrary. /squaresolid If the kernel is not separable in the sense of Eq. (16.70), there is still the possibility that itmaybeapproximatedbyakernelthatisseparable.Thenwecangettheexactsolutionof anapproximateequation,anequationthatapproximatestheoriginalequation.Thesolution oftheseparableapproximatekernelproblemcanthenbecheckedbysubstitutingbackinto theoriginal,unseparablekernelproblem. Numerical Solution There is extensive literature on the numerical solution of integral equations, and much of it concerns special techniques for certain situations. One method of fair generality is the replacement of the single integral equation by a set of simultaneous algebraic equations. And again matrix techniques are invoked. This simultaneous algebraic equation–matrix approach is applied here to two different cases. For the homogeneous Fredholm equation ofthesecondkindthismethodworks well.FortheFredholmequationofthefirstkindthe methodis adisaster.First wedealwiththedisaster. We considertheFredholmintegralequationofthefirst kind, f(x)=integraldisplayb aK(x,t)ϕ(t)dt, (16.84a) withf(x)andK(x,t)known and ϕ(t)unknown. The integral can be evaluated (in prin- ciple) by quadrature techniques. For maximum accuracy the Gaussian method is recom- mended(ifthekerneliscontinuousandhascontinuousderivatives).Thenumericalquadra- turereplacestheintegralbyasummation, f(xi)=nsummationdisplay k=1AkK(xi,tk)ϕ(tk), (16.84b) 1024 Chapter 16 Integral Equations withAkthe quadrature coefficients. We abbreviate f(xi)asfi,ϕ(tk)asϕk, and AkK(xi,tk)asBik. In effect we are changing from a function description to a vector– matrix description, with the ncomponents of the vector (fi)defined as the values of the functionatthe ndiscretepoints [f(xi)]. Equation(16.84b)becomes fi=nsummationdisplay k=1Bikϕk, amatrixequation.Inverting (Bik), weobtain ϕ(xk)=ϕk=nsummationdisplay k=1B−1 kifi, (16.84c) andEq.(16.84a)issolved—inprinciple.Inpractice,thequadraturecoefficient–kernelma- trix is often “ill-conditioned” (with respect to inversion). This means that in the inversion process small (numerical) errors are multiplied by large factors. In the inversion process allsignificantfiguresmaybelost andEq. (16.84c)becomesnumericalnonsense. This disaster should not be entirely unexpected. Integration is essentially a smoothing operation. f(x)is relatively insensitive to local variation of ϕ(t). Conversely, ϕ(t)may be exceedingly sensitive to small changes in f(x). Small errors in f(x)or inB−1are magnified and accuracy disappears. This same behavior shows up in attempts to invert Laplacetransformsnumerically. When the quadrature–matrix technique is applied to the integral equation eigenvalue problem,thesymmetrickernel,homogeneousFredholmequationofthesecondkind,4 λϕ(x)=integraldisplayb aK(x,t)ϕ(t)dt, (16.84d) the technique is far more successful. Replacing the integral by a set of simultaneous alge- braicequations(numericalquadrature),wehave λϕi=nsummationdisplay k=1AkKikϕk, (16.84e) withϕi=ϕ(xi), as before. The points xi,i=1,2,...,n, are taken to be the same (nu- merically) as tk,k=1,2,...,n,s oKikwill be symmetric. The system is symmetrized by multiplyingby A1/2 isothat λparenleftbig A1/2 iϕiparenrightbig =nsummationdisplay k=1parenleftbig A1/2 iKikA1/2 kparenrightbigparenleftbig A1/2 kϕkparenrightbig . (16.84f) Replacing A1/2 iϕibyψiandA1/2 iKikA1/2 kbySik, weobtain λψ=Sψ, (16.84g) withSsymmetric (since the kernel K(x,t)was assumed symmetric). Of course, ψhas components ψi=ψ(xi).Equation(16.84g)isourmatrixeigenvalueequation,Eq.(3.136). 4The eigenvalue λhas been written on the left side, multiplying the eigenfunction, as is customary in matrix analysis (Section 3.5). In this form λwilltakeon a maximum value . 16.3 Neumann Series, Separable (Degenerate) Kernels 1025 Theeigenvaluesarereadilyobtainedbycallingacannedeigenroutine.5Forkernelssuchas those of Exercise 16.3.15 and using a 10-point Gauss–Legendre quadrature, the eigenrou- tine determines the largest eigenvalue to within about 0.5 percent for the cases where the kernel has discontinuities in its derivatives. If the derivatives are continuous, the accuracy ismuchbetter. Linz6hasdescribedaninterestingvariationalrefinementinthedeterminationof λmaxto highaccuracy.ThekeytohismethodisExercise17.8.7.Thecomponentsoftheeigenfunc- tionvectorareobtainedfromEq.(16.84d)with ϕ(tk)nowknownand ϕi=ϕ(xi)generated asrequired.(The xiarenolongertiedtothe tk.) Exercises 16.3.1 UsingtheNeumannseries, solve (a)ϕ(x)=1−2integraldisplayx 0tϕ(t)dt, (b)ϕ(x)=x+integraldisplayx 0(t−x)ϕ(t)dt , (c)ϕ(x)=x−integraldisplayx 0(t−x)ϕ(t)dt . ANS.(a) ϕ(x)=e−x2. 16.3.2 Solvetheequation ϕ(x)=x+1 2integraldisplay1 −1(t+x)ϕ(t)dt bytheseparablekernelmethod.ComparewiththeNeumannmethodsolutionofSection 16.3. ANS.ϕ(x)=1 2(3x−1). 16.3.3 Findtheeigenvaluesandeigenfunctionsof ϕ(x)=λintegraldisplay1 −1(t−x)ϕ(t)dt. 16.3.4 Findtheeigenvaluesandeigenfunctionsof ϕ(x)=λintegraldisplay2π 0cos(x−t)ϕ(t)dt. ANS.λ1=λ2=1 π,ϕ(x)=Acosx+Bsinx. 5SeeW.H.Press,B.P.Flannery,S.A.Teukolsky,andW.T.Vetterling, NumericalRecipes ,2nded.,Cambridge,UK:Cambridge UniversityPress(1992),Chapter11,fordetails,references,andcomputercodes.Thesymbolicsoftware Mathematica andMaple also include matrix functions for computing eigenvalues and eigenvectors. 6P. Linz, On the numerical computation of eigenvalues and eigenvectors of symmetric integral equations. Math. Comput. 24: 905 (1970). 1026 Chapter 16 Integral Equations 16.3.5 Findtheeigenvaluesandeigenfunctionsof y(x)=λintegraldisplay1 −1(x−t)2y(t)dt. Hint.This problem may be treated by the separable kernel method or by a Legendre expansion. 16.3.6 IftheseparablekerneltechniqueofthissectionisappliedtoaFredholmequationofthe firstkind(Eq. (16.1)), showthatEq.(16.76) isreplacedby c=A−1b. Ingeneralthesolutionfor theunknown ϕ(t)isnotunique. 16.3.7 Solve ψ(x)=x+integraldisplay1 0(1+xt)ψ(t)dt byeachofthefollowingmethods: (a) theNeumannseriestechnique, (b) theseparablekerneltechnique, (c) educatedguessing. 16.3.8 Usetheseparablekerneltechniquetoshowthat ψ(x)=λintegraldisplayπ 0cosxsintψ(t)dt hasnosolution(apartfromthetrivial ψ=0).Explainthisresultintermsofseparability andsymmetry. 16.3.9 Solve ϕ(x)=1+λ2integraldisplayx 0(x−t)ϕ(t)dt byeachofthefollowingmethods: (a) reductiontoanODE(findtheboundaryconditions), (b) theNeumannseries, (c) theuse ofLaplacetransforms. ANS.ϕ(x)=coshλx. 16.3.10 (a) In Eq. (16.69a) take V=V0, independent of t. Without using Eq. (16.69b), show thatEq. (16.69a)leadsdirectlyto U(t−t0)=expbracketleftbigg −i ¯h(t−t0)V0bracketrightbigg . (b) Repeatfor Eq.(16.69b) withoutusingEq. (16.69a). 16.3 Neumann Series, Separable (Degenerate) Kernels 1027 16.3.11 Givenϕ(x)=λintegraltext1 0(1+xt)ϕ(t)dt , solve for the eigenvalues and the eigenfunctions by theseparablekerneltechnique. 16.3.12 Knowingtheformof thesolutionscanbeagreatadvantage,for theintegralequation ϕ(x)=λintegraldisplay1 0(1+xt)ϕ(t)dt, assumeϕ(x)to have the form 1 +bx. Substitute into the integral equation. Integrate andsolvefor bandλ. 16.3.13 Theintegralequation ϕ(x)=λintegraldisplay1 0J0(αxt)ϕ(t)dt, J 0(α)=0, isapproximatedby ϕ(x)=λintegraldisplay1 0bracketleftbig 1−x2t2bracketrightbig ϕ(t)dt. Find the minimum eigenvalue λand the corresponding eigenfunction ϕ(t)of the ap- proximateequation. ANS.λmin=1.112486,ϕ(x)=1−0.303337x2. 16.3.14 Youare giventheintegralequation ϕ(x)=λintegraldisplay1 0sinπxtϕ(t)dt. Approximatethekernelby K(x,t)=4xt(1−xt)≈sinπxt. Find the positive eigenvalue and the corresponding eigenfunction for the approximate integralequation. Note.ForK(x,t)=sinπxt,λ=1.6334. ANS.λ=1.5678,ϕ(x)=x−0.6955x2 (λ+=√ 31−4,λ−=−√ 31−4). 16.3.15 Theequation f(x)=integraldisplayb aK(x,t)ϕ(t)dt hasadegeneratekernel K(x,t)=summationtextn i=1Mi(x)Ni(t). (a) Showthatthisintegralequationhasnosolutionunless f(x)canbewrittenas f(x)=nsummationdisplay i=1fiMi(x), withtheficonstants. 1028 Chapter 16 Integral Equations (b) Showthattoanysolution ϕ(x)wemayadd ψ(x),provided ψ(x)isorthogonalto allNi(x): integraldisplayb aNi(x)ψ(x)dx=0 for all i. 16.3.16 Usingnumericalquadrature,convert ϕ(x)=λintegraldisplay1 0J0(αxt)ϕ(t)dt, J 0(α)=0, toasetof simultaneouslinearequations. (a) Findtheminimumeigenvalue λ. (b) Determine ϕ(x)at discrete values of xand plotϕ(x)versusx. Compare with the approximateeigenfunctionofExercise16.3.13. ANS.(a) λmin=1.14502. 16.3.17 Usingnumericalquadrature,convert ϕ(x)=λintegraldisplay1 0sinπxtϕ(t)dt toasetof simultaneouslinearequations. (a) Findtheminimumeigenvalue λ. (b) Determine ϕ(x)at discrete values of xand plotϕ(x)versusx. Compare with the approximateeigenfunctionofExercise16.3.14. ANS.(a) λmin=1.6334. 16.3.18 GivenahomogeneousFredholmequationof thesecondkind λϕ(x)=integraldisplay1 0K(x,t)ϕ(t)dt. (a) Calculate the largest eigenvalue λ0. Use the 10-point Gauss–Legendre quadrature technique.ForcomparisontheeigenvalueslistedbyLinzaregivenas λexact. (b) Tabulate ϕ(xk), wherethe xkarethe10evaluationpointsin [0,1]. (c) Tabulatetheratio 1 λ0ϕ(x)integraldisplay1 0K(x,t)ϕ(t)dt forx=xk. Thisis thetestofwhetheror notyoureallyhaveasolution. (a)K(x,t)=ext. ANS.λexact=1.35303. 16.4 Hilbert–Schmidt Theory 1029 (b)K(x,t)=braceleftBigg1 2x(2−t), x<t, 1 2t(2−x), x>t. ANS.λexact=0.24296. (c)K(x,t)=|x−t|. ANS.λexact=0.34741. (d)K(x,t)=braceleftBiggx, x<t, t, x>t. ANS.λexact=0.40528. Note.(1) The evaluation points xiof Gauss–Legendre quadrature for [−1,1]may be linearlytransformedinto [0,1], xi[0,1]=1 2parenleftbig xi[−1,1]+1parenrightbig . Thentheweightingfactors Aiare reducedinproportiontothelengthoftheinterval: Ai[0,1]=1 2Ai[−1,1]. 16.3.19 UsingthematrixvariationaltechniqueofExercise17.8.7,refineyourcalculationofthe eigenvalueofExercise16.3.18(c) [K(x,t)=|x−t|].T rya40×40 matrix. Note.Your matrix should be symmetric so that the (unknown) eigenvectors will be or- thogonal. ANS.(40-pointGauss–Legendrequadrature)0.34727. 16.4 H ILBERT –SCHMIDT THEORY Symmetrization of Kernels Thisisthedevelopmentofthepropertiesoflinearintegralequations(Fredholmtype)with symmetrickernels: K(x,t)=K(t,x). (16.85) Beforeplungingintothetheory,wenotethatsomeimportantnonsymmetrickernelscanbe symmetrized.If wehavetheequation ϕ(x)=f(x)+λintegraldisplayb aK(x,t)ρ(t)ϕ(t)dt, (16.86) thetotalkernelisactually K(x,t)ρ(t) ,clearlynotsymmetricif K(x,t)aloneissymmetric. However,ifwemultiplyEq.(16.86) by√ρ(x)andsubstitute radicalbig ρ(x)ϕ(x)=ψ(x), (16.87) 1030 Chapter 16 Integral Equations weobtain ψ(x)=radicalbig ρ(x)f(x)+λintegraldisplayb abracketleftbig K(x,t)radicalbig ρ(x)ρ(t)bracketrightbig ψ(t)dt, (16.88) with a symmetric total kernel K(x,t)√ρ(x)ρ(t). We shall meet ρ(x)later as a positive weightingfactor inthis integralequationSturm–Liouvilletheory. Orthogonal Eigenfunctions WenowfocusonthehomogeneousFredholmequationofthesecondkind: ϕ(x)=λintegraldisplayb aK(x,t)ϕ(t)dt. (16.89) We assume that the kernel K(x,t)is symmetric and real. Perhaps one of the first ques- tions we might ask about the equation is: “Does it make sense?” or more precisely, “Does an eigenvalue λsatisfying this equation exist?” With the aid of the Schwarz and Bessel inequalities, Chapter 10 and Courant and Hilbert (Chapter III, Section 4—see the Addi- tional Readings) show that if K(x,t)is continuous, there is at least one such eigenvalue andpossiblyaninfinitenumberofthem. We show that the eigenvalues, λ, are real and that the corresponding eigenfunctions, ϕi(x), are orthogonal. Let λi,λjbe twodifferent eigenvalues and ϕi(x),ϕj(x)be the correspondingeigenfunctions.Equation(16.89)thenbecomes ϕi(x)=λiintegraldisplayb aK(x,t)ϕ i(t)dt, (16.90a) ϕj(x)=λjintegraldisplayb aK(x,t)ϕ j(t)dt. (16.90b) If we multiply Eq. (16.90a) by λjϕj(x)and Eq. (16.90b) by λiϕi(x)and then integrate withrespectto x, thetwoequationsbecome7 λjintegraldisplayb aϕi(x)ϕj(x)dx=λiλjintegraldisplayb aintegraldisplayb aK(x,t)ϕ i(t)ϕj(x)dtdx, (16.91a) λiintegraldisplayb aϕi(x)ϕj(x)dx=λiλjintegraldisplayb aintegraldisplayb aK(x,t)ϕ j(t)ϕi(x)dtdx, (16.91b) Sincewehavedemandedthat K(x,t)bysymmetric,Eq. (16.91b)mayberewrittenas λiintegraldisplayb aϕi(x)ϕj(x)dx=λiλjintegraldisplayb aintegraldisplayb aK(x,t)ϕ i(t)ϕj(x)dtdx. (16.92) SubtractingEq. (16.92)fromEq. (16.91a), weobtain (λj−λi)integraldisplayb aϕi(x)ϕj(x)dx=0. (16.93) 7Weassume that the necessaryintegrals exist. For anexample of asimple pathological case,seeExercise 16.4.3. 16.4 Hilbert–Schmidt Theory 1031 ThishasthesameformasEq. (10.34)intheSturm–Liouvilletheory.Since λi/negationslash=λj, integraldisplayb aϕi(x)ϕj(x)dx=0,i/negationslash=j, (16.94) proving orthogonality. Note that with a real symmetric kernel, no complex conjugates are involvedinEq.(16.94). Fortheself-adjointorHermitiankernel,see Exercise16.4.1. Iftheeigenvalue λiisdegenerate,8theeigenfunctionsforthatparticulareigenvaluemay be orthogonalized by the Gram–Schmidt method (Section 10.3). Our orthogonal eigen- functions may, of course, be normalized, and we assume that this has been done. The resultis integraldisplayb aϕi(x)ϕj(x)dx=δij. (16.95) To demonstrate that the λiare real, we need to admit complex conjugates. Taking the complexconjugateof Eq. (16.90a),wehave ϕ∗ i(x)=λ∗ iintegraldisplayb aK(x,t)ϕ∗ i(t)dt, (16.96) providedthekernel K(x,t)isreal.Now,usingEq.(16.96)insteadofEq.(16.90b),wesee thattheanalysisleadsto (λ∗ i−λi)integraldisplayb aϕ∗ i(x)ϕi(x)dx=0. (16.97) Thistimetheintegralcannotvanish(unlesswehavethetrivialsolution, ϕi(x)=0)and λ∗i=λi, (16.98) orλi, oureigenvalue,isreal. This is the thirdtime we have passed this way, first with Hermitian matrices, then with Sturm–Liouville (self-adjoint) ODEs, and now with Hilbert–Schmidt integral equations. The correspondence between the Hermitian matrices and the self-adjoint ODEs shows up in physics as the two outstanding formulations of quantum mechanics—the Heisenberg matrix approach and the Schrödinger differential operator approach. In Section 17.8 and Exercise 17.7.6 we shall explore further the correspondence between the Hilbert–Schmidt symmetrickernelintegralequationsandtheSturm–Liouvilleself-adjointdifferentialequa- tions. The eigenfunctions of our integral equations form a complete set,9in the sense that any functiong(x)thatcanbegeneratedbytheintegral g(x)=integraldisplay K(x,t)h(t)dt, (16.99) 8If more than one distinct eigenfunction corresponds to the same eigenvalue (satisfying Eq. (16.89)), that eigenvalue is said to be degenerate(see Chapters 3 and 4). 9For aproof of this statement, seeCourant and Hilbert (1953), Chapter III, Section 5, in the Additional Readings. 1032 Chapter 16 Integral Equations in which h(t)is any piecewise continuous function, can be represented by a series of eigenfunctions, g(x)=∞summationdisplay n=1anϕn(x). (16.100) Theseriesconvergesuniformlyandabsolutely. Letus extendthis tothekernel K(x,t)byassertingthat K(x,t)=∞summationdisplay n=1anϕn(t), (16.101) andan=an(x).Substitutingintotheoriginalintegralequation(Eq.(16.89))andusingthe orthogonalityintegral,weobtain ϕi(x)=λiai(x). (16.102) Therefore for our homogeneous Fredholm equation of the second kind, the kernel may be expressedintermsof theeigenfunctionsandeigenvaluesby K(x,t)=∞summationdisplay n=1ϕn(x)ϕn(t) λn(zero notaneigenvalue) . (16.103) Herewehaveabilinearexpansion,alinearexpansionin ϕn(x)andlinearin ϕn(t).Similar bilinear expansions appear in Section 9.7. It is possible that the expansion given by Eq. (16.101) may not exist. As an illustration of the sort of pathological behavior that may occur,youareinvitedtoapplythisanalysisto ϕ(x)=λintegraldisplay∞ 0e−xtϕ(t)dt (compareExercise16.4.3). ItshouldbeemphasizedthatthisHilbert–Schmidttheoryisconcernedwiththeestablish- ment of properties of the eigenvalues (real) and eigenfunctions (orthogonality, complete- ness), properties that may be of great interest and value. The Hilbert–Schmidttheory does not solve the homogeneous integral equation for us any more than the Sturm–Liouville theory of Chapter 10 solved the ODEs. The solutions of the integral equation come from Sections16.2and16.3(includingnumericalanalysis). Nonhomogeneous Integral Equation Weneedasolutionofthenonhomogeneousequation ϕ(x)=f(x)+λintegraldisplayb aK(x,t)ϕ(t)dt. (16.104) Let us assume that the solutions of the corresponding homogeneous integral equation are known: ϕn(x)=λnintegraldisplayb aK(x,t)ϕ n(t)dt, (16.105) 16.4 Hilbert–Schmidt Theory 1033 the solution ϕn(x)corresponding to the eigenvalue λn. We expand both ϕ(x)andf(x)in termsofthisset ofeigenfunctions: ϕ(x)=∞summationdisplay n=1anϕn(x)(anunknown) , (16.106) f(x)=∞summationdisplay n=1bnϕn(x)(bnknown). (16.107) SubstitutingintoEq.(16.104), weobtain ∞summationdisplay n=1anϕn(x)=∞summationdisplay n=1bnϕn(x)+λintegraldisplayb aK(x,t)∞summationdisplay n=1anϕn(t)dt. (16.108) By interchangingthe order of integrationand summation,we may evaluatethe integral by Eq.(16.105), andweget ∞summationdisplay n=1anϕn(x)=∞summationdisplay n=1bnϕn(x)+λ∞summationdisplay n=1anϕn(x) λn. (16.109) If wemultiplyby ϕi(x)andintegratefrom x=atox=b,the orthogonalityof our eigen- functionsleadsto ai=bi+λai λi. (16.110) Thiscanberewrittenas ai=bi+λ λi−λbi, (16.111) whichbringsustooursolution ϕ(x)=f(x)+λ∞summationdisplay i=1integraltextb af(t)ϕi(t)dt λi−λϕi(x). (16.112) Here it is assumed that the eigenfunctions ϕi(x)are normalized to unity. Note that if f(x)=0there is no solution unless λ=λi. This means that our homogeneous equation hasnosolution(exceptthetrivial ϕ(x)=0)unless λisaneigenvalue, λi. In the event that λfor the nonhomogeneous equation (16.104) is equal to one of the eigenvalues λpof the homogeneous equation, our solution (Eq. (16.112)) blows up. To repairthedamagewereturntoEq. (16.110)andgivethevalue ap=bp+λpap λp=bp+ap (16.113) specialattention.Clearly, apdropsoutandisnolongerdeterminedby bp,whereas bp=0. This implies thatintegraltext f(x)ϕp(x)dx=0; that is, f(x)is orthogonal to the eigenfunction ϕp(x).I fthisis notthecase,wehavenosolution. 1034 Chapter 16 Integral Equations Equation (16.111) still holds for i/negationslash=p, so we multiply by ϕi(x)and sum over i(i/negationslash=p) toobtain ϕ(x)=f(x)+apϕp+λp∞summationdisplay i=1 i/negationslash=pintegraltextb af(t)ϕi(t)dt λi−λpϕi(x). (16.114) Inthissolutionthe apremainsas anundeterminedconstant.10 Exercises 16.4.1 IntheFredholmequation ϕ(x)=λintegraldisplayb aK(x,t)ϕ(t)dt thekernel K(x,t)is self-adjointor Hermitian: K(x,t)=K∗(t,x). Showthat (a) theeigenfunctionsareorthogonal,inthesense that integraldisplayb aϕ∗ m(x)ϕn(x)dx=0,m/negationslash=n( λm/negationslash=λn), (b) theeigenvaluesarereal. 16.4.2 Solvetheintegralequation ϕ(x)=x+1 2integraldisplay1 −1(t+x)ϕ(t)dt (compareExercise16.3.2)bytheHilbert–Schmidtmethod. Note. The application of the Hilbert–Schmidt technique here is somewhat like using a shotgun to kill a mosquito, especially when the equation can be solved quickly by expandinginLegendrepolynomials. 16.4.3 SolvetheFredholmintegralequation ϕ(x)=λintegraldisplay∞ 0e−xtϕ(t)dt. Note.A series expansion of the kernel e−xtwould permit a separable kernel-type solu- tion(Section16.3),exceptthattheseriesisinfinite.Thissuggestsaninfinitenumberof eigenvaluesandeigenfunctions.If youstopwith ϕ(x)=x−1/2,λ=π−1/2, 10This is like the inhomogeneous linear ODE. We may add to its solution any constant times a solution of the corresponding homogeneous ODE. 16.4 Hilbert–Schmidt Theory 1035 you will have missed most of the solutions. Show that the normalization integrals of the eigenfunctions do notexist. A basic reason for this anomalous behavior is that the rangeofintegrationis infinite,makingthisa“singular”integralequation. 16.4.4 Given y(x)=x+λintegraldisplay1 0xty(t)dt. (a) Determine y(x)asaNeumannseries. (b) Find the range of λfor which your Neumann series solution is convergent. Com- parewiththevalueobtainedfrom |λ|·|K|max<1. (c) Find the eigenvalue and the eigenfunction of the corresponding homogeneous in- tegralequation. (d) Bytheseparablekernelmethodshowthatthesolutionis y(x)=3x 3−λ. (e) Find y(x)bytheHilbert–Schmidtmethod. 16.4.5 InExercise16.3.4, K(x,t)=cos(x−t). The(unnormalized)eigenfunctionsare cos xand sinx. (a) Show that there is a function h(t)such that K(x,s), considered as a function of s alone,maybewrittenas K(x,s)=integraldisplay2π 0K(s,t)h(t)dt. (b) Showthat K(x,t)maybeexpandedas K(x,t)=2summationdisplay n=1ϕn(x)ϕn(t) λn. 16.4.6 Theintegralequation ϕ(x)=λintegraltext1 0(1+xt)ϕ(t)dt haseigenvalues λ1=0.7889and λ2= 15.211 andeigenfunctions ϕ1=1+0.5352xandϕ2=1−1.8685x. (a) Showthattheseeigenfunctionsareorthogonalovertheinterval [0,1]. (b) Normalizetheeigenfunctionstounity. (c) Showthat K(x,t)=ϕ1(x)ϕ1(t) λ1+ϕ2(x)ϕ2(t) λ2. 1036 Chapter 16 Integral Equations ANS.(b) ϕ1(x)=0.7831+0.4191x ϕ2(x)=1.8403−3.4386x. 16.4.7 An alternate form of the solution to the nonhomogeneous integral equation, Eq. (16.104),is ϕ(x)=∞summationdisplay i=1biλi λi−λϕi(x). (a) DerivethisformwithoutusingEq. (16.112). (b) ShowthatthisformandEq.(16.112)are equivalent. 16.4.8 (a) Showthattheeigenfunctionsof Exercise16.3.5areorthogonal. (b) Showthattheeigenfunctionsof Exercise16.3.11are orthogonal. AdditionalReadings Bocher, M., An Introduction to the Study of Integral Equations , Cambridge Tracts in Mathematics and Mathe- maticalPhysics, No.10. NewYork: Hafner (1960). This is ahelpful introduction to integral equations. Cochran, J. A., The Analysis of Linear Integral Equations . New York: McGraw-Hill (1972). This is a compre- hensivetreatmentoflinearintegralequationswhichisintendedforappliedmathematiciansandmathematical physicists. It assumes amoderate to high levelof mathematicalcompetence onthe part of thereader. Courant, R., and D. Hilbert, Methods of Mathematical Physics , Vol.1 (English edition). New York: Interscience (1953). This is one of the classic works of mathematical physics. Originally published in German in 1924, the revised English edition is an excellent reference for a rigorous treatment of integral equations, Green’s functions, and awidevariety of other topics on mathematicalphysics. Golberg, M. A., ed., Solution Methods of Integral Equations . New York: Plenum Press (1979). This is a set of papers from a conference on integral equations. The initial chapter is excellent for up-to-date orientation and awealthofreferences. Kanval, R. P., Linear Integral Equations . New York: Academic Press (1971), reprinted, Birkhäuser (1996). This book is a detailedbut readabletreatment of avariety oftechniques for solving linearintegral equations. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 7 is a particularly detailed, complete discussion of Green’s functions from the point of view of mathematical physics. Note, however, that Morse and Feshbach frequently choose a source of 4 πδ(r−r′)in place of our δ(r−r′).Considerable attention is devoted to bounded regions. Muskhelishvili, N.I., Singular Integral Equations , 2nd ed.,NewYork: Dover (1992). Stakgold, I., Green’s Functions and Boundary ValueProblems . NewYork: Wiley (1979). CHAPTER 17 CALCULUS OF VARIATIONS Uses of the Calculus of Variations We now address problems where we search for a function or curve, rather than a value of somevariable,thatmakes agivenquantitystationary,usuallyan energyoractionintegral. Becauseafunctionisvaried,theseproblemsarecalled variational .Variationalprinciples, such as D’Alembert’s and Hamilton’s, have been developed in classical mechanics, and Lagrangiantechniquesoccurinquantummechanicsandfieldtheory,forexample,Fermat’s principle of the shortest optical path in electrodynamics. Before plunging into this rather differentbranchofmathematicalphysics,letussummarizesomeofitsusesinbothphysics andmathematics. 1. In existingphysicaltheories: a. Unificationofdiverseareasofphysicsusingenergyasakeyconcept. b. Convenienceinanalysis—Lagrangeequations,Section17.3. c. Eleganttreatmentofconstraints,Section17.7. 2. Starting point for new, complex areas of physics and engineering . In general rela- tivity the geodesic is taken as the minimum path of a light pulse or the free-fall path of a particle in curved Riemannian space (see geodesics in Section 2.10). Variational principles appear in quantum field theory. Variational principles have been applied extensivelyincontroltheory. 3. Mathematical unification . Variational analysis provides a proof of the completeness of the Sturm–Liouville eigenfunctions, Chapter 10, and establishes a lower bound for the eigenvalues. Similar results follow for the eigenvalues and eigenfunctions of the Hilbert–Schmidtintegralequation,Section16.4. 4. Calculationtechniques ,Section17.8.Calculationoftheeigenfunctionsandeigenval- uesoftheSturm–Liouvilleequation.Integralequationeigenfunctionsandeigenvalues maybecalculatedusingnumericalquadratureandmatrixtechniques,Section16.3. 1037 1038 Chapter 17 Calculus of Variations 17.1 A D EPENDENT AND AN INDEPENDENT VARIABLE Concept of Variation The calculus of variations involves problems in which the quantity to be minimized (or maximized)appearsasastationaryintegral,afunctional,becauseafunction y(x,α)needs to be determined from a class described by an infinitesimal parameter α. As the simplest case,let J=integraldisplayx2 x1f(y,yx,x)dx. (17.1) HereJis the quantity that takes on a stationary value. Under the integral sign, fis a knownfunctionoftheindicatedvariables xandα,asarey(x,α),y x(x,α)≡∂y(x,α)/∂x , but the dependence of yonx(andα)is not yet known; that is, y(x)isunknown .T h i s meansthatalthoughtheintegralisfrom x1tox2,theexactpathofintegrationisnotknown (Fig.17.1).Wearetochoosethepathofintegrationthroughpoints (x1,y1)and(x2,y2)to minimize J. Strictly speaking, we determine stationary values of J: minima, maxima, or saddle points. In most cases of physical interest the stationary value will be a minimum. This problem is considerably more difficult than the corresponding problem of a function y(x)in differential calculus. Indeed, there may be no solution. In differential calculus the minimum is determined by comparing y(x0)withy(x), wherexranges over neighboring points. Here we assume the existence of an optimum path, that is, an acceptable path for whichJis stationary, and then compare Jfor our (unknown) optimum path with that obtainedfromneighboringpaths.InFig.17.1twopossiblepathsareshown.(Therearean infinite number of possibilities.) The difference between these two for a given xis called the variation of y,δy, and is conveniently described by introducing a new function, η(x), to define the arbitrary deformation of the path and a scale factor, α, to give the magnitude ofthevariation.Thefunction η(x)is arbitraryexceptfortworestrictions. First, η(x1)=η(x2)=0, (17.2) FIGURE 17.1Avariedpath. 17.1 A Dependent and an Independent Variable 1039 which means that all varied paths must pass through the fixed endpoints. Second, as will beseenshortly, η(x)mustbedifferentiable;thatis, wemaynotuse η(x)=1,x=x0, (17.3) =0,x/negationslash=x0, but we can choose η(x)to have a form similar to the functions used to represent the Dirac deltafunction(Chapter1)sothat η(x)differsfromzeroonlyoveraninfinitesimalregion.1 Then,withthepathdescribedby αandη(x), y(x,α)=y(x,0)+αη(x) (17.4) and δy=y(x,α)−y(x,0)=αη(x). (17.5) Let us choose y(x,α=0)as the unknown path that will minimize J. Theny(x,α) for nonzero αdescribes a neighboring path. In Eq. (17.1), Jis now a function2of our parameter α: J(α)=integraldisplayx2 x1fbracketleftbig y(x,α),y x(x,α),xbracketrightbig dx, (17.6) andourconditionfor anextremevalueisthat bracketleftbigg∂J(α) ∂αbracketrightbigg α=0=0, (17.7) analogoustothevanishingof thederivative dy/dxindifferentialcalculus. Now, the α-dependence of the integral is contained in y(x,α)andyx(x,α)= (∂/∂x)y(x,α) . Therefore3 ∂J(α) ∂α=integraldisplayx2 x1bracketleftbigg∂f ∂y∂y ∂α+∂f ∂yx∂yx ∂αbracketrightbigg dx. (17.8) FromEq. (17.4), ∂y(x,α) ∂α=η(x), (17.9) ∂yx(x,α) ∂α=dη(x) dx, (17.10) soEq. (17.8)becomes ∂J(α) ∂α=integraldisplayx2 x1parenleftbigg∂f ∂yη(x)+∂f ∂yxdη(x) dxparenrightbigg dx. (17.11) 1Compare H. Jeffreys and B. S. Jeffreys, Methods of Mathematical Physics , 3rd ed., Cambridge, UK: Cambridge University Press (1966), Chapter 10, for amore complete discussion of this point. 2Technically, Jis afunctional of y,yx, but a function of αdepending on the functions y(x,α)andyx(x,α):J[y(x,α), yx(x,α)]. 3Notethat yandyxarebeingtreatedas independent variables. 1040 Chapter 17 Calculus of Variations Integrating the second term by parts to get η(x)as a common and arbitrary nonvanishing factor,weobtain integraldisplayx2 x1dη(x) dx∂f ∂yxdx=η(x)∂f ∂yxvextendsinglevextendsinglevextendsinglevextendsinglex2 x1−integraldisplayx2 x1η(x)d dx∂f ∂yxdx. (17.12) TheintegratedpartvanishesbyEq.(17.2), andEq. (17.11)becomes integraldisplayx2 x1bracketleftbigg∂f ∂y−d dx∂f ∂yxbracketrightbigg η(x)dx=0. (17.13) Inthisform αhasbeensetequaltozero,correspondingtothesolutionpath,and,ineffect, isnolongerpartoftheproblem. Occasionally we will see Eq. (17.13) multiplied by δα, which gives, upon using η(x)δα=δy, integraldisplayx2 x1parenleftbigg∂f ∂y−d dx∂f ∂yxparenrightbigg δydx=δαbracketleftbigg∂J ∂αbracketrightbigg α=0=δJ=0. (17.14) Sinceη(x)isarbitrary,wemaychooseittohavethesamesignasthebracketedexpression inEq.(17.13)wheneverthelatterdiffersfromzero.Hencetheintegrandisalwaysnonneg- ative. Equation (17.13), our condition for the existence of a stationary value, can then be satisfied only if the bracketed term itself is zero almost everywhere. The condition for our stationaryvalueis thusaPDE,4 ∂f ∂y−d dx∂f ∂yx=0, (17.15) known as the Euler equation, which can be expressed in various other forms. Sometimes solutionsaremissedwhentheyarenottwicedifferentiable,asrequiredbyEq.(17.15).An exampleisGoldschmidt’sdiscontinuoussolutionofSection17.2.ItisclearthatEq.(17.15) must be satisfied for Jto take on a stationary value, that is, for Eq. (17.14) to be satisfied. Equation(17.15)isnecessary,butitisbynomeanssufficient.5CourantandRobbins(1996; see the Additional Readings) illustrate this very nicely by considering the distance over a sphere between points on the sphere, AandB, Fig. 17.2. Path (1), a great circle, is found from Eq. (17.15). But path (2), the remainder of the great circle through points AandB, also satisfies the Euler equation. Path (2) is a maximum, but only if we demand that it be a great circle and then only if we make less than one circuit; that is, path (2) +ncomplete revolutions is also a solution. If the path is not required to be a great circle, any deviation from (2) will increase the length. This is hardly the property of a local maximum, and that is why it is important to check the properties of solutions of Eq. (17.15) to see if they satisfythephysicalconditionsofthegivenproblem. 4It is important to watchthemeaning of ∂/∂xandd/dxclosely. For example, if f=f[y(x),yx,x], df dx=∂f ∂x+∂f ∂ydy dx+∂f ∂yxd2y dx2. The first term on the right gives the explicitx-dependence. The second and third terms give the implicitx-dependence via y andyx. 5Foradiscussionofsufficiencyconditionsandthedevelopmentofthecalculusofvariationsasapartofmathematics,seeG.M. Ewing,Calculus of Variations with Applications , New York: Norton (1969). Sufficiency conditions are also covered by Sagan (in theAdditional Readings atthe end of this chapter). 17.1 A Dependent and an Independent Variable 1041 FIGURE 17.2Stationary pathsoverasphere. Example 17.1.1 OPTICAL PATHNEAREVENT HORIZON OF A BLACK HOLE Determine the optical path in an atmosphere where the velocity of light increases in pro- portion to the height, v(y)=y/b,withb>0 some parameter describing the light speed. Sov=0a ty=0,which simulates the conditions at the surface of a black hole, called its event horizon , where the gravitational force is so strong that the velocity of light goes to zero,thuseventrappinglight. Becauselighttakestheshortesttime,thevariationalproblemtakestheform /Delta1t=integraldisplayt2 t1dt=integraldisplayds v=bintegraldisplayradicalbig dx2+dy2 ydt=minimum . Herev=ds/dt=y/bis the velocity of light in this environment, the ycoordinate being the height. A look at the variational functional suggests choosing yas the independent variable because xdoes not appear in the integrand. We can bring dyoutside the radical and change the role of xandyinJof Eq. (17.1) and the resulting Euler equation. With x=x(y), x′=dx/dy,weobtain bintegraldisplay√ x′2+1 ydy=minimum , andtheEulerequationbecomes ∂f ∂x−d dy∂f ∂x′=0. Since∂f/∂x=0,thiscanbeintegrated,giving x′ y√ x′2+1=C1=const.,orx′2=C2 1y2parenleftbig x′2+1parenrightbig . Separating dxanddyinthisfirst-order ODEwefindtheintegral integraldisplayx dx=integraldisplayyC1ydyradicalBig 1−C2 1y2, 1042 Chapter 17 Calculus of Variations FIGURE 17.3Circularopticalpathinmedium. whichyields x+C2=−1 C1radicalBig 1−C2 1y2,or(x+C2)2+y2=1 C2 1. This is a circular light path with center on the x-axis along the event horizon. (See Fig. 17.3.) This example may be adapted to a mirage (Fata Morgana) in a desert with hot air near the ground and cooler air aloft (the index of refraction changes with height in cool versus hot air), thus changing the velocity law from v=y/b→v0−y/b.In this case, the circular light path is no longer convex with center on the x-axis, but becomes concave. /squaresolid Alternate Forms of Euler Equations Oneotherform(Exercise17.1.1), whichis oftenuseful, is ∂f ∂x−d dxparenleftbigg f−yx∂f ∂yxparenrightbigg =0. (17.16) In problems in which f=f(y,yx), that is, in which xdoes not appear explicitly, Eq.(17.16) reducesto d dxparenleftbigg f−yx∂f ∂yxparenrightbigg =0, (17.17) or f−yx∂f ∂yx=constant. (17.18) Example 17.1.2 Missing Dependent Variables Consider the variational problemintegraltext f(˙r)dt=minimum. Here ris absent from the inte- grand.ThereforetheEulerequationsbecome d dt∂f ∂˙x=0,d dt∂f ∂˙y=0,d dt∂f ∂˙z=0, 17.1 A Dependent and an Independent Variable 1043 withr=(x,y,z), sof˙r=c=const.Solvingthesethreeequationsforthethreeunknowns ˙x,˙y,˙zyields˙r=c1=const. Integrating this constant velocity gives r=c1t+c2.The solutionsarestraightlines,despitethegeneralnatureof thefunction f. A physical example illustrating this case is the propagation of light in a crystal, where thevelocityoflightdependsonthe(crystal)directionsbutnotonthelocationinthecrystal, becauseacrystalisananisotropic homogeneousmedium .The variationalproblem integraldisplayds v=integraldisplay√ ˙r2 v(˙r)dt=minimum hastheformofourexample.Notethat tneednotbethetime,butitparameterizesthelight path. /squaresolid Exercises 17.1.1 Fordy/dx≡yx/negationslash=0,showtheequivalenceofthetwoforms ofEuler’s equation: ∂f ∂x−d dx∂f ∂yx=0 and ∂f ∂y−d dxparenleftbigg f−yx∂f ∂yxparenrightbigg =0. 17.1.2 DeriveEuler’sequationbyexpandingtheintegrandof J(α)=integraldisplayx2 x1fbracketleftbig y(x,α),y x(x,α),xbracketrightbig dx inpowersof α,usingaTaylor(Maclaurin)expansionwith yandyxasthetwovariables (Section5.6). Note. The stationary condition is ∂J(α)/∂α=0, evaluated at α=0. The terms quadraticin αmaybeusefulinestablishingthenatureofthestationarysolution(maxi- mum,minimum,or saddlepoint). 17.1.3 FindtheEulerequationcorrespondingtoEq.(17.15) if f=f(yxx,yx,y,x). ANS.d2 dx2parenleftbigg∂f ∂yxxparenrightbigg −d dxparenleftbigg∂f ∂yxparenrightbigg +∂f ∂y=0, η(x1)=η(x2)=0,ηx(x1)=ηx(x2)=0. 17.1.4 Theintegrand f(y,yx,x)ofEq. (17.1)hastheform f(y,yx,x)=f1(x,y)+f2(x,y)y x. (a) ShowthattheEuler equationleadsto ∂f1 ∂y−∂f2 ∂x=0. (b) Whatdoesthisimplyforthedependenceoftheintegral Juponthechoiceofpath? 1044 Chapter 17 Calculus of Variations 17.1.5 Showthattheconditionthat J=integraldisplay f(x,y)dx hasastationaryvalue (a) leadsto f(x,y)independentof yand (b) yieldsnoinformationaboutany x-dependence. Wegetno(continuous,differentiable)solution.Tobeameaningfulvariationalproblem, dependenceon yorhigherderivativesis essential. Note. The situation will change when constraints are introduced (compare Exer- cise17.7.7). 17.2 A PPLICATIONS OF THE EULER EQUATION Example 17.2.1 STRAIGHT LINE PerhapsthesimplestapplicationoftheEulerequationisinthedeterminationoftheshortest distancebetweentwopointsintheEuclidean xy-plane.Sincetheelementofdistanceis ds=bracketleftbig (dx)2+(dy)2bracketrightbig1/2=bracketleftbig 1+y2 xbracketrightbig1/2dx, (17.19) thedistance Jmaybewrittenas J=integraldisplayx2,y2 x1,y1ds=integraldisplayx2 x1bracketleftbig 1+y2 xbracketrightbig1/2dx. (17.20) ComparisonwithEq. (17.1) showsthat f(y,yx,x)=parenleftbig 1+y2 xparenrightbig1/2. (17.21) SubstitutingintoEq.(17.16), weobtain −d dxbracketleftbigg1 (1+y2x)1/2bracketrightbigg =0, (17.22) or 1 (1+y2x)1/2=C,aconstant . (17.23) Thisis satisfiedby yx=a,asecondconstant , (17.24) and y=ax+b, (17.25) which is the familiar equation for a straight line. The constants aandbare chosen so that the line passes through the two points (x1,y1)and(x2,y2). Hence the Euler equation 17.2 Applications of the Euler Equation 1045 predictsthattheshortest6distancebetweentwofixedpointsinEuclideanspaceisastraight line. /squaresolid Thegeneralizationofthisincurvedfour-dimensionalspace–timeleadstotheimportant conceptof thegeodesicingeneralrelativity(seeSection2.10). Example 17.2.2 SOAPFILM As a second illustration (Fig. 17.4), consider two parallel coaxial wire circles to be con- nected by a surface of minimum area that is generated by revolving a curve y(x)about thex-axis.Thecurveisrequiredtopassthroughfixedendpoints (x1,y1)and(x2,y2).The variationalproblemistochoosethecurve y(x)sothattheareaoftheresultingsurfacewill beaminimum. Fortheelementof areashowninFig.17.4, dA=2πyds=2πyparenleftbig 1+y2 xparenrightbig1/2dx. (17.26) Thevariationalequationis then J=integraldisplayx2 x12πyparenleftbig 1+y2 xparenrightbig1/2dx. (17.27) Neglectingthe 2 π,weobtain f(y,yx,x)=yparenleftbig 1+y2 xparenrightbig1/2. (17.28) FIGURE 17.4Surfaceofrotation—soapfilm problem. 6Technically,wehave astationary value. From the α2terms it can beidentified asa minimum (Exercise 17.2.2). 1046 Chapter 17 Calculus of Variations Since∂f/∂x=0,wemayapplyEq. (17.18)directlyandget yparenleftbig 1+y2 xparenrightbig1/2−yy2 x1 (1+y2x)1/2=c1, (17.29) or y (1+y2x)1/2=c1. (17.30) Squaring,weget y2 1+y2x=c2 1withc2 1≤y2 min, (17.31) and (yx)−1=dx dy=c1radicalBig y2−c2 1. (17.32) Thismaybeintegratedtogive x=c1cosh−1y c1+c2. (17.33) Solvingfor y,weha v e y=c1coshparenleftbiggx−c2 c1parenrightbigg , (17.34) and again c1andc2are determined by requiring the hyperbolic cosine to pass through the points(x1,y1)and(x2,y2).Our“minimum”-areasurfaceisaspecialcaseofacatenaryof revolution,ora catenoid. /squaresolid Soap Film — Minimum Area This calculus of variations contains many pitfalls for the unwary. (Remember, the Euler equationisa necessary conditionassuminga differentiablesolution .Thesufficiencycon- ditions are quite involved. See the Additional Readings for details.) Respect for some of these hazards may be developed by considering a specific physical problem, for example, aminimum-areaproblemwith (x1,y1)=(−x0,1),(x2,y2)=(+x0,1).Theminimumsur- faceisasoapfilmstretchedbetweenthetworingsofunitradiusat x=±x0.Theproblem istopredictthecurve y(x)assumedbythesoapfilm. By referring to Eq. (17.34), we find that c2=0 by the symmetry of the problem about x=0.Then y=c1coshparenleftbiggx c1parenrightbigg ,c 1coshparenleftbiggx0 c1parenrightbigg =1. (17.34a) If wetake x0=1 2weobtainatranscendentalequationfor c1,vi z. 1=c1coshparenleftbigg1 2c1parenrightbigg . (17.35) 17.2 Applications of the Euler Equation 1047 We find that this equation has two solutions: c1=0.2350, leading to a “deep” curve, and c1=0.8483, leading to a “flat” curve. Which curve is assumed by the soap film? Before answering this question, consider the physical situation with the rings moved apart so that x0=1.ThenEq.(17.34a)becomes 1=c1coshparenleftbigg1 c1parenrightbigg , (17.36) whichhas norealsolutions .Thephysicalsignificanceisthatastheunit-radiusringswere moved out from the origin, a point was reached at which the soap film could no longer maintain the same horizontal force over each vertical section. Stable equilibrium was no longer possible. The soap film broke (irreversible process) and formed a circular film over each ring (with a total area of 2 π=6.2832...). This is the Goldschmidt discontinuous solution. Thenextquestionis:Howlargemay x0beandstillgivearealsolutionforEq.(17.34a)?7 Lettingc−1 1=p, Eq.(17.34a)becomes p=coshpx0. (17.37) To findx0maxwe could solve for x0(as in Eq. (17.33)) and then differentiate with respect top. Finally, with an eye on Fig. 17.5, dx0/dpwould be set equal to zero. Alternatively, directdifferentiationof Eq.(17.37) withrespectto pyields 1=bracketleftbigg x0+pdx0 dpbracketrightbigg sinhpx0. FIGURE 17.5Solutionsof Eq. (17.34a)forunit-radiusrings atx=±x0. 7Fromanumericalpointofviewitiseasiertoinverttheproblem.Pickavalueof c1andsolvefor x0.Equation(17.34a)becomes x0=c1cosh−1(1/c1).This has numerical solutions in the range 0 <c1≤1. 1048 Chapter 17 Calculus of Variations Therequirementthat dx0/dpvanishleadsto 1=x0sinhpx0. (17.38) Equations(17.37)and(17.38)maybecombinedtoform px0=cothpx0, (17.39) withtheroot px0=1.1997. (17.40) SubstitutingintoEq.(17.37) or(17.38), weobtain p=1.810,c 1=0.5524 (17.41) and x0max=0.6627. (17.42) Returning to the question of the solution of Eq. (17.35) that describes the soap film, let uscalculatetheareacorrespondingtoeachsolution.Wehave A=4πintegraldisplayx0 0yparenleftbig 1+y2 xparenrightbig1/2dx=4π c1integraldisplayx0 0y2dx (byEq.(17.30)) =4πc1integraldisplayx0 0parenleftbigg coshx c1parenrightbigg2 dx=πc2 1bracketleftbigg sinhparenleftbigg2x0 c1parenrightbigg +2x0 c1bracketrightbigg . (17.43) Forx0=1 2, Eq.(17.35) leadsto c1=0.2350→A=6.8456, c1=0.8483→A=5.9917, showing that the former can at most be only a local minimum. A more detailed investiga- tion(compareBliss, CalculusofVariations ,ChapterIV)showsthatthissurfaceisnoteven alocalminimum.For x0=1 2, thesoapfilmwillbedescribedbytheflatcurve y=0.8483coshparenleftbiggx 0.8483parenrightbigg . (17.44) This flat or shallow catenoid (catenary of revolution) will be an absolute minimum for 0≤x0<0.528.However,for 0 .528<x<0.6627 its areais greaterthanthatof theGold- schmidtdiscontinuoussolution(6.2832)anditis onlyarelativeminimum(Fig.17.6). For an excellent discussion of both the mathematical problems and experiments with soap films, we refer to Courant and Robbins (1996) in the Additional Readings at the end ofthechapter. 17.2 Applications of the Euler Equation 1049 FIGURE 17.6Catenoidarea(unit-radiusringsat x=±x0). Exercises 17.2.1 A soap film is stretched across the space between two rings of unit radius centered at ±x0on thex-axis and perpendicular to the x-axis. Using the solution developed in Section 17.2, set up the transcendental equations for the condition that x0is such that the area of the curved surface of rotation equals the area of the two rings (Goldschmidt discontinuoussolution).Solvefor x0(Fig.17.7). 17.2.2 In Example 17.2.1, expand J[y(x,α)]−J[y(x,0)]in powers of α. The term linear in αleads to the Euler equation and to the straight-line solution, Eq. (17.25). Investigate FIGURE 17.7Surfaceof rotation. 1050 Chapter 17 Calculus of Variations theα2term and show that the stationary value of J, the straight-line distance, is a minimum . 17.2.3 (a) Showthattheintegral J=integraldisplayx2 x1f(y,yx,x)dx, withf=y(x), hasnoextremevalues. (b) Iff(y,yx,x)=y2(x), find a discontinuous solution similar to the Goldschmidt solutionfor thesoapfilmproblem. 17.2.4 Fermat’sprincipleof opticsstates thatalightraywillfollowthepath y(x)for which integraldisplayx2,y2 x1,y1n(y,x)ds is a minimum when nis the index of refraction. For y2=y1=1,−x1=x2=1, find theraypathif (a)n=ey,(b)n=a(y−y0), y >y 0. 17.2.5 A frictionless particle moves from point Aon the surface of the Earth to point Bby sliding through a tunnel. Find the differential equation to be satisfied if the transit time istobeaminimum. Note.AssumetheEarthtobenonrotatingsphereofuniformdensity. ANS.(Eq. (17.15)): rϕϕ(r3−ra2)+r2 ϕ(2a2−r2)+a2r2=0, r(ϕ=0)=r0,rϕ(ϕ=0)=0,r ( ϕ=ϕA)=a, r(ϕ=ϕB)=a. Eq. (17.18): r2 ϕ=a2r2 r2 0·r2−r2 0 a2−r2. The solution of these equations is a hypocycloid, gener- ated by a circle of radius1 2(a−r0)rolling inside the circle of radius a. You might like toshowthatthetransittimeis t=π(a2−r2 0)1/2 (ag)1/2. For details see P. W. Cooper, A m .J .P h y s . 34: 68 (1966); G. Veneziano et al., ibid. , pp. 701–704. 17.2.6 A ray of light follows a straight-line path in a first homogeneous medium, is refracted at an interface, and then follows a new straight-line path in the second medium. Use Fermat’sprincipleof opticstoderiveSnell’slawofrefraction: n1sinθ1=n2sinθ2. Hint. Keep the points (x1,y1)and(x2,y2)fixed and vary x0to satisfy Fermat (Fig. 17.8). This is notan Euler equation problem. (The light path is not differentiable atx0.) 17.2 Applications of the Euler Equation 1051 FIGURE 17.8Snell’slaw. 17.2.7 A second soap film configuration for the unit-radius rings at x=±x0consists of a circular disk, radius a,i nt h ex=0 plane and two catenoids of revolution, one joining thediskandeachring.Onecatenoidmaybedescribedby y=c1coshparenleftbiggx c1+c3parenrightbigg . (a) Imposeboundaryconditionsat x=0 andx=x0. (b) Althoughnotnecessary,itisconvenienttorequirethatthecatenoidsformanangle of 120◦where they join the central disk. Express this third boundary condition in mathematicalterms. (c) Showthatthetotalareaofcatenoidspluscentraldiskis A=c2 1bracketleftbigg sinhparenleftbigg2x0 c1+2c3parenrightbigg +2x0 c1bracketrightbigg . Note. Although this soap film configuration is physically realizable and stable, the area is larger than that of the simple catenoid for all ring separations for which both films exist. ANS.(a)  1=c1coshparenleftbiggx0 c1+c3parenrightbigg a=c1coshc3,(b)dy dx=tan30◦=sinhc3. 17.2.8 For the soap film described in Exercise 17.2.7, find (numerically) the maximum value ofx0. Note.Thiscallsforapocketcalculatorwithhyperbolicfunctionsoratableofhyperbolic cotangents. ANS.x0max=0.4078. 1052 Chapter 17 Calculus of Variations 17.2.9 Find the root of px0=cothpx0(Eq. (17.39)) and determine the corresponding values ofpandx0(Eqs.(17.41)and(17.42)).Calculateyourvaluestofivesignificantfigures. 17.2.10 Forthetwo-ringsoapfilmproblemofthissectioncalculateandtabulate x0,p,p−1,and A,thesoapfilmareafor px0=0.00(0.02)1.30. 17.2.11 Findthevalueof x0(tofivesignificantfigures)thatleadstoasoapfilmarea,Eq.(17.43), equalto 2 π,theGoldschmidtdiscontinuoussolution. ANS.x0=0.52770. 17.2.12 Find the curve of quickest descent from (0,0)to(x0,y0)for a particle sliding under gravity and without friction. Show that the ratio of times taken by the particle along a straight line joining the two points compared to along the curve of quickest descent is (1+4/π2)1/2. Hint.T a k eyto increase downwards. Apply Eq. (17.18) to obtain y2 x=(1−c2y)/c2y, wherecis an integration constant. Then make the substitution y=(sin2ϕ/2)/c2to parametrizethecycloidandtake (x0,y0)=(π/2c2,1/c2). 17.3 S EVERAL DEPENDENT VARIABLES Our original variational problem, Eq. (17.1), may be generalized in several respects. In this section we consider the integrand fto be a function of several dependent vari- ablesy1(x),y2(x),y3(x),...,all of which depend on x, the independent variable. In Sec- tion 17.4 fagain will contain only one unknown function y,b u tywill be a function of several independent variables (over which we integrate). In Section 17.5 these two gener- alizations are combined. In Section 17.7 the stationary value is restricted by one or more constraints. Formorethanonedependentvariable,Eq.(17.1) becomes J=integraldisplayx2 x1fbracketleftbig y1(x),y2(x),...,y n(x),y1x(x),y2x(x),...,y nx(x),xbracketrightbig dx. (17.45) As in Section 17.1, we determine the extreme value of Jby comparing neighboring paths.Let yi(x,α)=yi(x,0)+αηi(x), i=1,2,...,n, (17.46) with the ηiindependent of one another but subject to the restrictions discussed in Sec- tion 17.1. By differentiating Eq. (17.45) with respect to αand setting α=0, since Eq.(17.7) stillapplies,weobtain integraldisplayx2 x1summationdisplay iparenleftbigg∂f ∂yiηi+∂f ∂yixηixparenrightbigg dx=0, (17.47) thesubscript xdenotingpartialdifferentiationwithrespectto x;thatis,yix=∂yi/∂x,and so on. Again, each of the terms (∂f/∂y ix)ηixis integrated by parts. The integrated part vanishesandEq. (17.47)becomes integraldisplayx2 x1summationdisplay iparenleftbigg∂f ∂yi−d dx∂f ∂yixparenrightbigg ηidx=0. (17.48) 17.3 Several Dependent Variables 1053 Since the ηiare arbitrary and independent of one another,8each of the terms in the sum mustvanish independently .W eha v e ∂f ∂yi−d dx∂f ∂(∂yi/∂x)=0,i=1,2,...,n, (17.49) awholesetof Eulerequations,eachof whichmustbesatisfiedfor anextremevalue. Hamilton’s Principle The most important application of Eq. (17.45) occurs when the integrand fis taken to be a Lagrangian L. The Langrangian (for nonrelativistic systems; see Exercise 17.3.5 for a relativistic particle) is defined as the difference of kinetic and potential energies of a system: L≡T−V. (17.50) Using time as an independent variable instead of xandxi(t)as the dependent variables, weget x→t, y i→xi(t), y ix→˙xi(t); xi(t)is the location and ˙xi=dxi/dtis the velocity of particle ias a function of time. The equation δJ=0 is then a mathematicalstatement of Hamilton’sprinciple of classical mechanics, δintegraldisplayt2 t1L(x1,x2,...,xn,˙x1,˙x2,...,˙xn;t)dt=0. (17.51) In words, Hamilton’s principle asserts that the motion of the system from time t1tot2 is such that the time integral of the Lagrangian L, or action, has a stationary value. The resultingEulerequationsareusuallycalledtheLagrangianequationsof motion, d dt∂L ∂˙xi−∂L ∂xi=0. (17.52) TheseLagrangianequationscanbederivedfromNewton’sequationsofmotion,andNew- ton’s equations can be derived from Lagrange’s. The two sets of equations are equally “fundamental.” The Lagrangian formulation has advantages over the conventional Newtonian laws. Whereas Newton’s equations are vector equations, we see that Lagrange’s equations in- volveonlyscalarquantities.Thecoordinates x1,x2,...neednotbeanystandardsetofco- ordinatesorlengths.Theycanbeselectedtomatchtheconditionsofthephysicalproblem. TheLagrangeequationsareinvariantwithrespecttothechoiceofcoordinatesystem.New- ton’s equations (in component form) are not manifestly invariant. Exercise 2.5.10 shows whathappensto F=maresolvedinsphericalpolarcoordinates. 8For example, we could set η2=η3=η4=···=0, eliminating all but one term of the sum, and then treat η1exactly as in Section 17.1. 1054 Chapter 17 Calculus of Variations Exploitingtheconceptofenergy,wemayeasilyextendtheLagrangianformulationfrom mechanicstodiversefields,suchaselectricalnetworksandacousticalsystems.Extensions to electromagnetism appear in the exercises. The result is a unity of otherwise-separate areas of physics. In the development of new areas, the quantization of Lagrangian particle mechanics provided a model for the quantization of electromagnetic fields and led to the gaugetheoryofquantumelectrodynamics. One of the most valuable advantages of the Hamilton principle—Lagrange equation formulation—istheeaseinseeingarelationbetweenasymmetryandaconservationlaw. As an example, let xi=ϕ, an azimuthal angle. If our Lagrangian is independent of ϕ (that is, if ϕis an ignorable coordinate), there are two consequences: (1) the conservation or invariance of a component of angular momentum and (2) from Eq. (17.52) ∂L/∂˙ϕ= constant.Similarly,invarianceundertranslationleadstoconservationoflinearmomentum. Noether’stheoremisageneralizationofthisinvariance(symmetry)—theconservationlaw relation. Example 17.3.1 MOVING PARTICLE —C ARTESIAN COORDINATES ConsiderEq.(17.50), whichdescribesoneparticlewithkineticenergy T=1 2m˙x2(17.53) and potential energy V(x), in which, as usual, the force is given by the negative gradient ofthepotential, F(x)=−dV(x) dx. (17.54) FromEq. (17.52), d dt(m˙x)−∂(T−V) ∂x=m¨x−F(x)=0, (17.55) whichisNewton’ssecondlawofmotion. /squaresolid Example 17.3.2 MOVING PARTICLE —C IRCULAR CYLINDRICAL COORDINATES Now let us describe a moving particle in cylindrical coordinates of the xy-plane, that is, z=0.Thekineticenergyis T=1 2mparenleftbig ˙x2+˙y2parenrightbig =1 2mparenleftbig ˙ρ2+ρ2˙ϕ2parenrightbig , (17.56) andwetake V=0 for simplicity. The transformation of ˙x2+˙y2into circular cylindrical coordinates could be carried out by taking x(ρ,ϕ)andy(ρ,ϕ), Eq. (2.28), and differentiating with respect to time and squaring. It is much easier to interpret ˙x2+˙y2asv2and just write down the components ofvasˆρ(dsρ/dt)=ˆρ˙ρ,andsoon.(The dsρisanincrementof length,ρchangingby dρ, ϕremainingconstant.SeeSections2.1and2.4.) TheLagrangianequationsyield d dt(m˙ρ)−mρ˙ϕ2=0,d dtparenleftbig mρ2˙ϕparenrightbig =0. (17.57) 17.3 Several Dependent Variables 1055 Thesecondequationisastatementofconservationofangularmomentum.Thefirstmaybe interpretedasradialacceleration9equatedtocentrifugalforce.Inthissensethecentrifugal force is a real force. It is of some interest that this interpretation of centrifugal force as a realforceissupportedbythegeneraltheoryofrelativity. /squaresolid Exercises 17.3.1 (a) Developtheequationsofmotioncorrespondingto L=1 2m(˙x2+˙y2). (b) Inwhatsense doyoursolutionsminimizetheintegralintegraltextt2 t1Ldt? Comparetheresultfor yoursolutionwith x=const.,y=const. 17.3.2 From the Lagrangian equations of motion, Eq. (17.52), show that a system in stable equilibriumhasa minimumpotentialenergy. 17.3.3 Write out the Lagrangian equations of motion of a particle in spherical coordinates forpotential Vequaltoaconstant.Identifythetermscorrespondingto(a)centrifugalforce and(b) Coriolisforce. 17.3.4 The spherical pendulum consists of a mass on a wire of length l, free to move in polar angleθandazimuthangle ϕ(Fig.17.9). (a) SetuptheLagrangianfor thisphysicalsystem. (b) DeveloptheLagrangianequationsofmotion. 17.3.5 ShowthattheLagrangian L=m0c2parenleftbigg 1−radicalBigg 1−v2 c2parenrightbigg −V(r) FIGURE 17.9Spherical pendulum. 9Hereis asecond method ofattackingExercise 2.4.8. 1056 Chapter 17 Calculus of Variations leadstoarelativisticformof Newton’ssecondlawof motion, d dtparenleftbiggm0viradicalbig 1−v2/c2parenrightbigg =Fi, inwhichtheforcecomponentsare Fi=−∂V/∂xi. 17.3.6 The Lagrangian for a particle with charge qin an electromagnetic field described by scalarpotential ϕandvectorpotential Ais L=1 2mv2−qϕ+qA·v. Findtheequationofmotionofthechargedparticle. Hint.(d/dt)A j=∂Aj/∂t+summationtext i(∂Aj/∂xi)˙xi.Thedependenceoftheforcefields Eand Buponthepotentials ϕandAisdevelopedinSection1.13(compareExercise1.13.10). ANS.m¨xi=q[E+v×B]i. 17.3.7 ConsiderasysteminwhichtheLagrangianisgivenby L(qi,˙qi)=T(qi,˙qi)−V(qi), whereqiand˙qirepresent sets of variables. The potential energy Vis independent of velocityandneither TnorVhasanyexplicittimedependence. (a) Showthat d dtparenleftbiggsummationdisplay j˙qj∂L ∂˙qj−Lparenrightbigg =0. (b) Theconstantquantity summationdisplay j˙qj∂L ∂˙qj−L defines the Hamiltonian H. Show that under the preceding assumed conditions, H=T+V, thetotalenergy. Note.Thekineticenergy Tis aquadraticfunctionofthe ˙qi. 17.4 S EVERAL INDEPENDENT VARIABLES Sometimes the integrand fof Eq. (17.1) will contain one unknown function, u, that is a function of several independent variables, u=u(x,y,z) , for the three-dimensional case, forexample.Equation(17.1) becomes J=integraldisplayintegraldisplayintegraldisplay f[u,ux,uy,uz,x,y,z]dxdydz, (17.58) ux=∂u/∂x,andsoon.Thevariationalproblemistofindthefunction u(x,y,z) forwhich Jisstationary, δJ=δα∂J ∂αvextendsinglevextendsinglevextendsinglevextendsingle α=0=0. (17.59) 17.4 Several Independent Variables 1057 GeneralizingSection17.1,welet u(x,y,z,α)=u(x,y,z, 0)+αη(x,y,z), (17.60) whereu(x,y,z,α=0)represents the (unknown) function for which Eq. (17.59) is satis- fied, whereas again η(x,y,z) is the arbitrary deviation that describes the varied function u(x,y,z,α) . This deviation η(x,y,z) is required to be differentiable and to vanish at the endpoints.ThenfromEq. (17.60), ux(x,y,z,α)=ux(x,y,z,0)+αηx, (17.61) andsimilarlyfor uyanduz. Differentiating the integral Eq. (17.58) with respect to the parameter αand then setting α=0,weobtain ∂J ∂αvextendsinglevextendsinglevextendsinglevextendsingle α=0=integraldisplayintegraldisplayintegraldisplayparenleftbigg∂f ∂uη+∂f ∂uxηx+∂f ∂uyηy+∂f ∂uzηzparenrightbigg dxdydz=0.(17.62) Again, we integrate each of the terms (∂f/∂u i)ηiby parts. The integrated part vanishes attheendpoints(becausethedeviation ηis requiredtogotozeroattheendpoints)and integraldisplayintegraldisplayintegraldisplayparenleftbigg∂f ∂u−∂ ∂x∂f ∂ux−∂ ∂y∂f ∂uy−∂ ∂z∂f ∂uzparenrightbigg η(x,y,z)dxdydz =0.10(17.63) Sincethevariation η(x,y,z) isarbitrary,theterminlargeparenthesesissetequaltozero. ThisyieldstheEuler equationfor (three)independentvariables, ∂f ∂y−∂ ∂x∂f ∂ux−∂ ∂y∂f ∂uy−∂ ∂z∂f ∂uz=0. (17.64) Example 17.4.1 LAPLACE ’SEQUATION Anexampleofthissortofvariationalproblemisprovidedbyelectrostatics.Theenergyof anelectrostaticfieldis energydensity =1 2εE2, (17.65) inwhichEis theusualelectrostaticforcefield.Interms ofthestaticpotential ϕ, energydensity =1 2ε(∇ϕ)2. (17.66) Now let us impose the requirement that the electrostatic energy (associated with the field) inagivenvolumebeaminimum.(Boundaryconditionson Eandϕmuststillbesatisfied.) Wehavethevolumeintegral11 J=integraldisplayintegraldisplayintegraldisplay (∇ϕ)2dxdydz=integraldisplayintegraldisplayintegraldisplayparenleftbig ϕ2 x+ϕ2 y+ϕ2 zparenrightbig dxdydz. (17.67) 10Recallthat ∂/∂xisapartialderivative,where yandzareheldconstant.But ∂/∂xalsoactson implicitx-dependenceaswell as onexplicitx-dependence. In this sense,for example, ∂ ∂xparenleftbigg∂f ∂uxparenrightbigg =∂2f ∂x∂ux+∂2f ∂u∂uxux+∂2f ∂u2xuxx+∂2f ∂uy∂uxuxy+∂2f ∂uz∂uxuxz. 11Thesubscript xindicatesthe x-partial derivative, not an x-component. 1058 Chapter 17 Calculus of Variations With f(ϕ,ϕx,ϕy,ϕz,x,y,z)=ϕ2 x+ϕ2 y+ϕ2 z, (17.68) thefunction ϕreplacingthe uofEq. (17.64), Euler’sequation(Eq. (17.64)) yields −2(ϕxx+ϕyy+ϕzz)=0, (17.69) or ∇2ϕ(x,y,z)=0, (17.70) whichisLaplace’sequationof electrostatics. Closer investigation shows that this stationary value is indeed a minimum. Thus the demandthatthefieldenergybeminimizedleadstoLaplace’sPDE. /squaresolid Exercises 17.4.1 TheLagrangianfor avibratingstring(small-amplitudevibrations)is L=integraldisplayparenleftbig1 2ρu2 t−1 2τu2 xparenrightbig dx, whereρis the (constant) linear mass density and τis the (constant) tension. The x- integrationisoverthelengthofthestring.ShowthatapplicationofHamilton’sprinciple totheLagrangiandensity(theintegrand),nowwithtwoindependentvariables,leadsto theclassicalwaveequation ∂2u ∂x2=ρ τ∂2u ∂t2. 17.4.2 Show that the stationary value of the total energy of the electrostatic field of Exam- ple17.4.1isa minimum . Hint.UseEq.(17.61) andinvestigatethe α2terms. 17.5 S EVERAL DEPENDENT AND INDEPENDENT VARIABLES In some cases our integrand fcontains more than one dependent variable and more than oneindependentvariable.Consider f=fbracketleftbig p(x,y,z),p x,py,pz,q(x,y,z),q x,qy,qz,r(x,y,z),r x,ry,rz,x,y,zbracketrightbig .(17.71) Weproceedas beforewith p(x,y,z,α)=p(x,y,z, 0)+αξ(x,y,z), q(x,y,z,α)=q(x,y,z, 0)+αη(x,y,z), (17.72) r(x,y,z,α)=r(x,y,z, 0)+αζ(x,y,z), andso on . 17.5 Several Dependent and Independent Variables 1059 Keeping in mind that ξ,η, andζare independent of one another, as were the ηiin Sec- tion17.3, thesamedifferentiationandthenintegrationbyparts leadsto ∂f ∂p−∂ ∂x∂f ∂px−∂ ∂y∂f ∂py−∂ ∂z∂f ∂pz=0, (17.73) with similar equations for functions qandr. Replacing p,q,r,... withyiandx,y,z,... withxy, wecanputEq. (17.73)inamorecompactform: ∂f ∂yi−summationdisplay j∂ ∂xjparenleftbigg∂f ∂yijparenrightbigg =0,i=1,2,..., (17.73a) inwhich yij≡∂yi ∂xj. Anapplicationof Eq.(17.73) appearsinSection17.7. Relation to Physics The calculus of variations as developed so far provides an elegant description of a wide variety of physical phenomena. The physics includes classical mechanics in Section 17.3; relativistic mechanics, Exercise 17.3.5; electrostatics, Example 17.4.1; and electromag- netictheoryinExercise17.5.1.Theconvenienceshouldnotbeminimized,butatthesame timeweshouldbeawarethatinthesecasesthecalculusofvariationshasonlyprovidedan alternate description of what was already known. The situation does change with incom- pletetheories. •If the basic physics is not yet known, a postulated variational principle can be a useful startingpoint. Exercise 17.5.1 TheLagrangian(perunitvolume)ofanelectromagneticfieldwithachargedensity ρis givenby L=1 2parenleftbigg ε0E2−1 µ0B2parenrightbigg −ρϕ+ρv·A. Show that Lagrange’s equations lead to two of Maxwell’s equations. (The remaining twoareaconsequenceofthedefinitionof EandBintermsof Aandϕ.)ThisLagrange densitycomesfromascalarexpressioninSection4.6. Hint.T a k eA1,A2,A3, andϕasdependent variables, x,y,z, andtasindependent variables. EandBare givenintermsof AandϕbyEq. (4.142)andEq. (1.88). 1060 Chapter 17 Calculus of Variations 17.6 L AGRANGIAN MULTIPLIERS In this section the concept of a constraint is introduced. To simplify the treatment, the constraint appears as a simple function rather than as an integral. In this section we are not concerned with the calculus of variations, but in Section 17.7 the constraints, with our newlydevelopedLagrangianmultipliers,are incorporatedintothecalculusofvariations. Consider a function of three independent variables, f(x,y,z) . For the function fto be amaximum(or extreme),12 df=0. (17.74) Thenecessaryandsufficientconditionfor thisis ∂f ∂x=∂f ∂y=∂f ∂z=0, (17.75) inwhich df=∂f ∂xdx+∂f ∂ydy+∂f ∂zdz. (17.76) Often in physical problems the variables x,y,zare subject to constraints so that they are no longer all independent. It is possible, at least in principle, to use each constraint to eliminateonevariableandtoproceedwithanewandsmallersetofindependentvariables. The use of Lagrangian multipliers is an alternate technique that may be applied when thiseliminationofvariablesisinconvenientorundesirable.Letour equationofconstraint be ϕ(x,y,z)=0, (17.77) from which z(x,y)may be extracted if x,yare taken as the independent coordinates. Returning to Eq. (17.74), Eq. (17.75) no longer follows because there are now only two independent variables, so dzis no longer arbitrary. From the total differential dϕ=0,we thenobtain −∂ϕ ∂zdz=∂ϕ ∂xdx+∂ϕ ∂ydy (17.78) andtherefore df=∂f ∂xdx+∂f ∂ydy+λparenleftbigg∂ϕ ∂xdx+∂ϕ ∂xdxparenrightbigg ,λ=−fz ϕz, assuming ϕz=∂ϕ ∂z/negationslash=0.Thus, we may add Eq. (17.76) and a multiple of Eq. (17.78) to obtain df+λdϕ=parenleftbigg∂f ∂x+λ∂ϕ ∂xparenrightbigg dx+parenleftbigg∂f ∂y+λ∂ϕ ∂yparenrightbigg dy+parenleftbigg∂f ∂z+λ∂ϕ ∂zparenrightbigg dz=0.(17.79) Inotherwords, ourLagrangianmultiplier λischosensothat ∂f ∂z+λ∂ϕ ∂z=0, (17.80) 12Including asaddle point. 17.6 Lagrangian Multipliers 1061 assumingthat ∂ϕ/∂z/negationslash=0.Equation(17.79)nowbecomes parenleftbigg∂f ∂x+λ∂ϕ ∂xparenrightbigg dx+parenleftbigg∂f ∂y+λ∂ϕ ∂yparenrightbigg dy=0. (17.81) However,now dxanddyarearbitraryandthequantitiesinparenthesesmustvanish: ∂f ∂x+λ∂ϕ ∂x=0,∂f ∂y+λ∂ϕ ∂y=0. (17.82) When Eqs. (17.80) and (17.82) are satisfied, df=0 andfis an extremum. Notice that there are now four unknowns: x,y,z, andλ. The fourth equation is, of course, the con- straintEq.(17.77).Wewantonly x,y,andz,soλneednotbedetermined.Forthisreason λis sometimes called Lagrange’s undetermined multiplier . This method will fail if all the coefficients of λvanish at the extremum, ∂ϕ/∂x,∂ϕ/∂y,∂ϕ/∂z=0. It is then impos- s ibletos olv ef or λ. NotethatfromtheformofEqs.(17.80)and(17.82),wecouldidentify fasthefunction takinganextremevaluesubjectto ϕ,theconstraint,orwecouldidentify fastheconstraint andϕas thefunction. If wehavea set ofconstraints ϕk, thenEqs. (17.80)and(17.82)become ∂f ∂xi+summationdisplay kλk∂ϕk ∂xi=0,i=1,2,...,n, withaseparateLagrangemultiplier λkfor eachϕk. Example 17.6.1 PARTICLE IN A BOX As an example of the use of Lagrangian multipliers, consider the quantum mechanical problemofaparticle(mass m)inabox.Theboxisarectangularparallelepipedwithsides a,b,andc.Theground-stateenergyof theparticleis givenby E=h2 8mparenleftbigg1 a2+1 b2+1 c2parenrightbigg . (17.83) We seektheshapeoftheboxthatwillminimizetheenergy E,subjecttoconstraintthat thevolumeis constant, V(a,b,c)=abc=k. (17.84) Withf(a,b,c)=E(a,b,c) andϕ(a,b,c)=abc−k=0,weobtain ∂E ∂a+λ∂ϕ ∂a=−h2 4ma3+λbc=0. (17.85) Also, −h2 4mb3+λac=0,−h2 4mc3+λab=0. 1062 Chapter 17 Calculus of Variations Multiplying the first of these expressions by a, the second by b, and the third by c,w e have λabc=h2 4ma2=h2 4mb2=h2 4mc2. (17.86) Thereforeoursolutionis a=b=c,acube. (17.87) Noticethat λhas notbeendeterminedbutfollowsfromEq. (17.86). /squaresolid Example 17.6.2 CYLINDRICAL NUCLEAR REACTOR A further example is provided by the nuclear reactor theory. Suppose a (thermal) nuclear reactor is to have the shape of a right circular cylinder of radius Rand height H. Neutron diffusiontheorysuppliesaconstraint: ϕ(R,H)=parenleftbigg2.4048 Rparenrightbigg2 +parenleftbiggπ Hparenrightbigg2 =constant.13(17.88) Wewishtominimizethevolumeofthereactorvessel, f(R,H)=πR2H. (17.89) ApplicationofEq. (17.82)leadsto ∂f ∂R+λ∂ϕ ∂R=2πRH−2λ(2.4048)2 R3=0, ∂f ∂H+λ∂ϕ ∂H=πR2−2λπ2 H3=0. (17.90) Bymultiplyingthefirstof theseequationsby R/2 andthesecondby H, weobtain πR2H=λ(2.4048)2 R2=λ2π2 H2, (17.91) orheight H=√ 2πR 2.4048=1.847R, (17.92) fortheminimum-volumeright-circularcylindricalreactor. Strictly speaking, we have found only an extremum. Its identification as a minimum followsfromaconsiderationoftheoriginalequations. /squaresolid 132.4048...is the lowest root of Besselfunction J0(R)(compare Section 11.1). 17.6 Lagrangian Multipliers 1063 Exercises ThefollowingproblemsaretobesolvedbyusingLagrangianmultipliers. 17.6.1 The ground-state energy of a quantum particle of mass min a pillbox (right-circular cylinder)is givenby E=¯h2 2mparenleftbigg(2.4048)2 R2+π2 H2parenrightbigg , inwhich Ristheradiusand Histheheightofthepillbox.Findtheratioof RtoHthat willminimizetheenergyfor afixedvolume. 17.6.2 Find the ratio of R(radius) to H(height) that will minimize the total surface area of a right-circularcylinderof fixedvolume. 17.6.3 TheU.S.PostOfficelimitsfirstclassmailtoCanadatoatotalof36inches,lengthplus girth. Using a Lagrange multiplier, find the maximum volume and the dimensions of a (rectangularparallelepiped)packagesubjecttothisconstraint. 17.6.4 Athermalnuclearreactoris subjecttotheconstraint ϕ(a,b,c)=parenleftbiggπ aparenrightbigg2 +parenleftbiggπ bparenrightbigg2 +parenleftbiggπ cparenrightbigg2 =B2,a constant . Findtheratiosofthesidesoftherectangularparallelepipedreactorofminimumvolume. ANS.a=b=c,cube. 17.6.5 For a lens of focal length f, the object distance pand the image distance qare related by 1/p+1/q=1/f. Find the minimum object–image distance (p+q)for fixed f. Assumerealobjectandimage( pandqbothpositive). 17.6.6 You have an ellipse (x/a)2+(y/b)2=1. Find the inscribed rectangle of maximum- area. Show that the ratio of the area of the maximum-area rectangle to the area of the ellipseis 2 /π=0.6366. 17.6.7 A rectangular parallelepiped is inscribed in an ellipsoid of semiaxes a,b, andc. Maxi- mize the volume of the inscribed rectangular parallelepiped. Show that the ratio of the maximumvolumetothevolumeof theellipsoidis 2 /π√ 3≈0.367. 17.6.8 Adeformed sphere has a radius given by r=r0{α0+α2P2(cosθ)}, whereα0≈1 and |α2|≪|α0|. FromExercise12.5.16theareaandvolumeare A=4πr2 0α2 0braceleftbigg 1+4 5parenleftbiggα2 α0parenrightbigg2bracerightbigg ,V=4πr3 0 3a3 0braceleftbigg 1+3 5parenleftbiggα2 α0parenrightbigg2bracerightbigg . Terms oforder α3 2havebeenneglected. (a) Withtheconstraintthattheenclosedvolumebeheldconstant,thatis, V=4πr3 0/3, showthattheboundingsurface ofminimumareaisasphere( α0=1,α2=0). (b) With the constraint that the area of the bounding surface be held constant, that is,A=4πr2 0, show that the enclosed volume is a maximum when the surface is asphere. 1064 Chapter 17 Calculus of Variations 17.6.9 Findthemaximumvalueofthedirectionalderivativeof ϕ(x,y,z) , dϕ ds=∂ϕ ∂xcosα+∂ϕ ∂ycosβ+∂ϕ ∂zcosγ, subjecttotheconstraint cos2α+cos2β+cos2γ=1. ANS.parenleftbiggdϕ dsparenrightbigg =|∇ϕ|. Note concerning the following exercises: In a quantum mechanical system there are gi distinct quantum states between energies EiandEi+dEi. The problem is to describe howniparticlesaredistributedamongthesestatessubjecttotwoconstraints: (a) fixednumberofparticles, summationdisplay ini=n. (b) fixedtotalenergy, summationdisplay iniEi=E. 17.6.10 For identical particles obeying the Pauli exclusion principle, the probability of a given arrangementis WFD=productdisplay igi! ni!(gi−ni)!. Show that maximizing WFD, subject to a fixed number of particles and fixed total en- ergy,leadsto ni=gi eλ1+λ2Ei+1. Withλ1=−E0/kTandλ2=1/kT, thisyieldsFermi–Diracstatistics. Hint.Tryworkingwithln WandusingStirling’sformula,Section8.3.Thejustification fordifferentiation withrespectto niisthatwearedealingherewithalargenumberof particles, /Delta1ni/ni≪1. 17.6.11 For identical particles but no restriction on the number in a given state, the probability ofagivenarrangementis WBE=productdisplay i(ni+gi−1)! ni!(gi−1)!. Show that maximizing WBE, subject to a fixed number of particles and fixed total en- ergy,leadsto ni=gi eλ1+λ2Ei−1. Withλ1=−E0/kTandλ2=1/kT, thisyieldsBose–Einsteinstatistics. Note.Assumethat gi≫1. 17.7 Variation with Constraints 1065 17.6.12 Photonssatisfy WBEandtheconstraintthattotalenergyisconstant.Theyclearlydo not satisfy the fixed-number constraint. Show that eliminating the fixed-number constraint leadstotheforegoingresultbutwith λ1=0. 17.7 V ARIATION WITH CONSTRAINTS Asintheprecedingsections,weseekthepaththatwillmaketheintegral J=integraldisplay fparenleftbigg yi,∂yi ∂xj,xjparenrightbigg dxj (17.93) stationary. This is the general case in which xjrepresents a set of independent variables andyiasetof dependentvariables.Again, δJ=0. (17.94) Now,however,we introduceone or moreconstraints. This meansthat the yiare no longer independent of each other. Not all the ηimay be varied arbitrarily, and Eqs. (17.62) and(17.73a)wouldnotapply.Theconstraintmayhavetheform ϕk(yi,xj)=0, (17.95) as in Section 17.6. In this case we may multiply by a function of xj,s a y ,λk(xj), and integrateoverthesamerangeasinEq. (17.93)toobtain integraldisplay λk(xj)ϕk(yi,xj)dxj=0. (17.96) Thenclearly δintegraldisplay λk(xj)ϕk(yi,xj)dxj=0. (17.97) Alternatively,theconstraintmayappearintheformofanintegral integraldisplay ϕk(yi,∂yi/∂xj,xj)dxj=constant. (17.98) We may introduce any constant Lagrangian multiplier, and again Eq. (17.97) follows— nowwith λaconstant. In either case, by adding Eqs. (17.94) and (17.97), possibly with more than one con- straint,weobtain δintegraldisplaybracketleftbigg fparenleftbigg yi,∂yi ∂xj,xjparenrightbigg +summationdisplay kλkϕk(yi,xj)bracketrightbigg dxj=0. (17.99) The Lagrangian multiplier λkmay depend on xjwhenϕ(yi,xj)is given in the form of Eq.(17.95). Treatingtheentireintegrandas anewfunction, gparenleftbigg yi,∂yi ∂xj,xjparenrightbigg , 1066 Chapter 17 Calculus of Variations weobtain gparenleftbigg yi,∂yi ∂xj,xjparenrightbigg =f+summationdisplay kλkϕk. (17.100) If we have Nyi(i=1,2,...,N)andmconstraints (k=1,2,...m), thenN−mof theηimay be taken as arbitrary. For the remaining mηi,t h eλmay, in principle, be cho- sen so that the remaining Euler–Lagrange equations are satisfied, completely analogous to Eq. (17.80). The result is that our composite function gmust satisfy the usual Euler– Lagrangeequations, ∂g ∂yi−summationdisplay j∂ ∂xj∂g (∂yi/∂xj)=0, (17.101) withonesuchequationforeachdependentvariable yi(compareEqs.(17.64)and(17.73)). These Euler equations and the equations of constraint are then solved simultaneously to findthefunctionyieldingastationaryvalue. Lagrangian Equations In the absence of constraints, Lagrange’s equations of motion (Eq. (17.52)) were found to be14 d dt∂L ∂˙qi−∂L ∂qi=0, witht(time)theoneindependentvariableand qi(t)(particlepositions)asetofdependent variables.Usuallythegeneralizedcoordinates qiarechosentoeliminatetheforcesofcon- straint, but this is not necessary and not always desirable. In the presence of (holonomic) constraints, ϕk=0,Hamilton’sprincipleis δintegraldisplaybracketleftbigg L(qi,˙qi,t)+summationdisplay kλk(t)ϕk(qi,t)bracketrightbigg dt=0, (17.102) andtheconstrainedLagrangianequationsofmotionare d dt∂L ∂˙qi−∂L ∂qi=summationdisplay kaikλk. (17.103) Usuallyϕk=ϕk(qi,t), independent of the generalized velocities ˙qi. In this case the coef- ficientaikis givenby aik=∂ϕk ∂qi. (17.104) Thenaikλk(no summation) represents the force of the kth constraint in the qi-direction, appearinginEq. (17.103)inexactlythesamewayas −∂V/∂qi. 14The symbol qis customary in classical mechanics. It serves to emphasize that the variable is not necessarily a Cartesian variable (and not necessarily alength). 17.7 Variation with Constraints 1067 FIGURE 17.10 Simplependulum. Example 17.7.1 SIMPLE PENDULUM To illustrate, consider the simple pendulum, a mass mconstrained by a wire of length lto swinginanarc(Fig. 17.10).Intheabsenceof theoneconstraint ϕ1=r−l=0 (17.105) there are two generalized coordinates randθ(motion in vertical plane). The Lagrangian is L=T−V=1 2mparenleftbig ˙r2+r2˙θ2parenrightbig +mgrcosθ, (17.106) taking the potential Vto be zero when the pendulum is horizontal, θ=π/2. By Eq.(17.103)theequationsofmotionare d dt∂L ∂˙r−∂L ∂r=λ1,d dt∂L ∂˙θ−∂L ∂θ=0(ar1=1,aθ1=0),(17.107) or d dt(m˙r)−mr˙θ2−mgcosθ=λ1, d dtparenleftbig mr2˙θparenrightbig +mgrsinθ=0. (17.108) Substitutingintheequationof constraint (r=l,˙r=0),weha v e ml˙θ2+mgcosθ=−λ1,ml2¨θ+mglsinθ=0. (17.109) The second equation may be solved for θ(t)to yield simple harmonic motion if the am- plitude is small (sinθ∼θ), whereas the first equation expresses the tension in the wire in termsofθand˙θ. Notethatsincetheequationofconstraint,Eq.(17.105),isintheformofEq.(17.95),the Lagrangemultiplier λmaybe(andhereis) afunctionof t(orofθ). /squaresolid 1068 Chapter 17 Calculus of Variations FIGURE 17.11Aparticle slidingonacylindrical surface. Example 17.7.2 SLIDING OFF A LOG Closely related to this is the problem of a particle sliding on a cylindrical surface. The object is to find the critical angle θcat which the particle flies off from the surface. This critical angle is the angle at which the radial force of constraint goes to zero (Fig. 17.11). We have L=T−V=1 2mparenleftbig ˙r2+r2˙θ2parenrightbig −mgrcosθ (17.110) andtheoneequationof constraint, ϕ1=r−l=0. (17.111) Proceedingas inExample17.7.1with ar1=1, m¨r−mr˙θ2+mgcosθ=λ1(θ), mr2¨θ+2mr˙r˙θ−mgrsinθ=0, (17.112) inwhichtheconstrainingforce λ1(θ)isafunctionoftheangle θ.15Sincer=l,¨r=˙r=0, Eq.(17.112)reducesto −ml˙θ2+mgcosθ=λ1(θ), (17.113a) ml2¨θ−mglsinθ=0. (17.113b) DifferentiatingEq.(17.113a)withrespecttotimeandrememberingthat df(θ) dt=df(θ) dθ˙θ, (17.114) weobtain −2ml¨θ−mgsinθ=dλ1(θ) dθ. (17.115) 15Note that λ1is theradialforce exerted by the cylinder on the particle. Consideration of the physical problem shows that λ1must depend on the angle θ. We permitted λ=λ(t). Now we are replacing the time dependence by an (unknown) angular dependence using θ=θ(t). 17.7 Variation with Constraints 1069 UsingEq. (17.113b)toeliminatethe ¨θtermandthenintegrating,wehave λ1(θ)=3mgcosθ+C. (17.116) Since λ1(0)=mg, (17.117) C=−2mg. (17.118) The particle mwill stay on the surface as long as the force of constraint is nonnegative, thatis, as longasthesurface hastopushoutwardontheparticle: λ1(θ)=3mgcosθ−2mg≥0. (17.119) The critical angle lies where λ1(θc)=0, the force of constraint going to zero. From Eq.(17.119), cosθc=2 3,orθc=48◦11′(17.120) fromthevertical.Atthisangle(neglectingallfriction)ourparticletakesoff. Itmustbeadmittedthatthisresultcanbeobtainedmoreeasilybyconsideringavarying centripetalforcefurnishedbytheradialcomponentofthegravitationalforce.Theexample was chosen to illustrate the use of Lagrange’s undetermined multiplier without confusing thereaderwithacomplicatedphysicalsystem. /squaresolid Example 17.7.3 THESCHRÖDINGER WAVEEQUATION Asafinalillustrationofaconstrainedminimum,letusfindtheEulerequationsforaquan- tummechanicalproblem δintegraldisplayintegraldisplayintegraldisplay ψ∗(x,y,z)Hψ(x,y,z)dxdydz =0, (17.121) withthenormalizationconstraintintegraldisplayintegraldisplayintegraldisplay ψ∗ψdxdydz=1. (17.122) Equation (17.121) is a statement that the energy of the system is stationary, Hbeing the quantummechanicalHamiltonianfor aparticleofmass m, adifferentialoperator, H=−¯h2 2m∇2+V(x,y,z). (17.123) Equation (17.122) is a bound-state constraint, ψis the usual wave function, a dependent variable,and ψ∗, itscomplexconjugate,istreatedas a second16dependentvariable. The integrand in Eq. (17.121) involves secondderivatives, which can be converted to firstderivativesbyintegratingbyparts: integraldisplay ψ∗∂2ψ ∂x2dx=ψ∗∂ψ ∂xvextendsinglevextendsinglevextendsinglevextendsingle−integraldisplay∂ψ∗ ∂x∂ψ ∂xdx. (17.124) We assume either periodic boundary conditions (as in the Sturm–Liouville theory, Chapter10) or that the volume of integration is so large that ψandψ∗vanish rapidly 16Compare Section 6.1. 1070 Chapter 17 Calculus of Variations enough17at the boundary. Then the integrated part vanishes and Eq. (17.121) may be rewrittenas δintegraldisplayintegraldisplayintegraldisplaybracketleftbigg ¯h2 2m∇ψ∗·∇ψ+Vψ∗ψbracketrightbigg dxdydz=0. (17.125) Thefunction gof Eq.(17.100)is g=¯h2 2m∇ψ∗·∇ψ+Vψ∗ψ−λψ∗ψ =¯h2 2m(ψ∗ xψx+ψ∗ yψy+ψ∗ zψz)+Vψ∗ψ−λψ∗ψ, (17.126) againusingthesubscript xtodenote ∂/∂x.F o ryi=ψ∗, Eq. (17.101)becomes ∂g ∂ψ∗−∂ ∂x∂g ∂ψ∗x−∂ ∂y∂g ∂ψ∗y−∂ ∂z∂g ∂ψ∗z=0. Thisyields Vψ−λψ−¯h2 2m(ψxx+ψyy+ψzz)=0, or −¯h2 2m∇2ψ+Vψ=λψ. (17.127) ReferencetoEq.(17.123)enablesustoidentify λphysicallyastheenergyofthequantum mechanical system. With this interpretation, Eq. (17.127) is the celebrated Schrödinger waveequation. /squaresolid This variational approach is more than just a matter of academic curiosity. It provides a verypowerfulmethodofobtainingapproximatesolutionsofthewaveequation(Rayleigh– Ritzvariationalmethod,Section17.8). Exercises 17.7.1 A particle, mass m, is on a frictionless horizontal surface. It is constrained to move so thatθ=ωt(rotatingradialarm,nofriction).Withtheinitialconditions t=0,r=r0,˙r=0, (a) findtheradialpositionsas afunctionoftime. ANS.r(t)=r0coshωt. (b) findtheforceexertedontheparticlebytheconstraint. ANS.F(c)=2m˙rω=2mr0ω2sinhωt. 17For example, lim r→∞rψ(r)=0. 17.7 Variation with Constraints 1071 17.7.2 A point mass mis moving over a flat, horizontal, frictionless plane. The mass is con- strained by a string to move radially inward at a constant rate. Using plane polar coor- dinates(ρ,ϕ),ρ=ρ0−kt, (a) SetuptheLagrangian. (b) ObtaintheconstrainedLagrangeequations. (c) Solve the ϕ-dependent Lagrange equation to obtain ω(t), the angular velocity. What is the physical significance of the constant of integration that you get from your“free”integration? (d) Usingthe ω(t)frompart(b),solvethe ρ-dependent(constrained)Lagrangeequa- tion to obtain λ(t). In other words, explain what is happening to the forceof con- straintas ρ→0. 17.7.3 A flexible cable is suspended from two fixed points. The length of the cable is fixed. Findthecurvethatwillminimizethetotalgravitationalpotentialenergyof thecable. ANS. Hyperboliccosine. 17.7.4 Afixedvolumeofwaterisrotatinginacylinderwithconstantangularvelocity ω.Find the curve of the water surface that will minimize the total potential energy of the water inthecombinedgravitational-centrifugalforcefield. ANS.Parabola. 17.7.5 (a) Showthatforafixed-lengthperimeterthefigurewithmaximumareaisacircle. (b) Showthatforafixedareathecurvewithminimumperimeterisacircle. Hint.Theradiusofcurvature Ris givenby R=(r2+r2 θ)3/2 rrθθ−2r2 θ−r2. Note. The problems of this section, variation subject to constraints, are often called isoperimetric . The term arose from problems of maximizing area subject to a fixed perimeter—asinExercise17.7.5(a). 17.7.6 Showthatrequiring J,givenby J=integraldisplayb abracketleftbig p(x)y2 x−q(x)y2bracketrightbig dx, tohaveastationaryvaluesubjecttothenormalizingcondition integraldisplayb ay2w(x)dx=1 leadstotheSturm–LiouvilleequationofChapter10: d dxparenleftbigg pdy dxparenrightbigg +qy+λwy=0. Note.Theboundarycondition pyxyvextendsinglevextendsinglevextendsingleb a=0 isusedinSection10.1inestablishingtheHermitianpropertyoftheoperator. 1072 Chapter 17 Calculus of Variations 17.7.7 Showthatrequiring J,givenby J=integraldisplayb aintegraldisplayb aK(x,t)ϕ(x)ϕ(t)dxdt, tohaveastationaryvaluesubjecttothenormalizingcondition integraldisplayb aϕ2(x)dx=1 leadstotheHilbert–Schmidtintegralequation,Eq.(16.89). Note.Thekernel K(x,t)issymmetric. 17.8 R AYLEIGH –RITZVARIATIONAL TECHNIQUE Exercise 17.7.6 opens up a relation between the calculus of variations and eigenfunction– eigenvalueproblems.Wemayrewritetheexpressionof Exercise17.7.6as Fbracketleftbig y(x)bracketrightbig =integraltextb a(py2 x−qy2)dx integraltextb ay2wdx, (17.128) in which the constraint appears in the denominator as a normalizing condition. After the unconstrained minimum of Fhas been found, ycan be normalized without changing the stationary value of Fbecause stationary values of Jcorrespond to stationary values of F. ThenfromExercise17.7.6,when y(x)issuchthat JandFtakeonastationaryvalue,the optimumfunction y(x)satisfiestheSturm–Liouvilleequation d dxparenleftbigg pdy dxparenrightbigg +qy+λwy=0, (17.129) withλtheeigenvalue( notaLagrangianmultiplier).Integratingthefirstterminthenumer- atorofEq. (17.128)bypartsandusingthe boundarycondition , pyxyvextendsinglevextendsinglevextendsingleb a=0, (17.130) weobtain Fbracketleftbig y(x)bracketrightbig =−integraldisplayb aybraceleftbiggd dxparenleftbigg pdy dxparenrightbigg +qybracerightbigg dxslashBigintegraldisplayb ay2wdx. (17.131) ThensubstitutinginEq. (17.129), thestationaryvaluesof F[y(x)]aregivenby Fbracketleftbig y(x)bracketrightbig =λn, (17.132) withλnthe eigenvalue corresponding to the eigenfunction yn. Equation (17.132) with F given by either Eq. (17.128) or Eq. (17.131) forms the basis of the Rayleigh–Ritz method forthecomputationofeigenfunctionsandeigenvalues. 17.8 Rayleigh–Ritz Variational Technique 1073 Ground State Eigenfunction Suppose that we seek to compute the ground-state eigenfunction y0and eigenvalue18λ0 of some complicated atomic or nuclear system. The classical example, for which no exact solutionexists,istheheliumatomproblem.Theeigenfunction y0isunknown ,butweshall assumewecanmakeaprettygoodguessatanapproximatefunction y,somathematically wemaywrite19 y=y0+∞summationdisplay i=1ciyi. (17.133) Theciare small quantities. (How small depends on how good our guess was.) The yiare orthonormalized eigenfunctions (also unknown), and therefore our trial function yis not normalized. Substitutingtheapproximatefunction yintoEq.(17.131)andnotingthat integraldisplayb ayibraceleftbiggd dxparenleftbigg pdyj dxparenrightbigg +qyibracerightbigg dx=−λiδij, (17.134) Fbracketleftbig y(x)bracketrightbig =λ0+summationtext∞ i=1c2 iλi 1+summationtext∞ i=1c2 i. (17.135) Herewehavetakentheeigenfunctionstobeorthonormal—sincetheyaresolutionsofthe Sturm–Liouvilleequation,Eq.(17.129).Wealsoassumethat y0isnondegenerate.Now,if wereplacesummationtext ic2 iλi→summationtext ic2 iλ0+summationtext ic2 i(λi−λ0)weobtain Fbracketleftbig y(x)bracketrightbig =λ0+summationtext∞ i=1c2 i(λi−λ0) 1+summationtext∞ i=1c2 i. (17.136) Equation(17.136)containstwoimportantresults. •Whereastheerrorintheeigenfunction ywasO(ci),theerrorin λisonly O(c2 i).Ev en a poor approximation of the eigenfunctions may yield an accurate calculation of the eigenvalue. •Ifλ0is thelowesteigenvalue(groundstate),thensince λi−λ0>0, Fbracketleftbig y(x)bracketrightbig =λ≥λ0, (17.137) or our approximation is always on the high side and becoming lower, converging on λ0as our approximate eigenfunction yimproves (ci→0). Note that Eq. (17.137) is a direct consequence of Eq. (17.135). More directly, F[y(x)]in Eq. (17.135) is a positively weighted average of the λiand, therefore, must be no smaller than the smallestλi, to wit,λ0. In practical problems in quantum mechanics, yoften depends on parameters that may be varied to minimize Fand thereby improve the estimate of the ground-state energy λ0. This is the “variational method” discussed in quantum mechanicstexts. 18Thismeansthat λ0isthelowesteigenvalue.ItisclearfromEq.(17.128) thatif p(x)≥0a n dq(x)≤0 (compareTable10.1), thenF[y(x)]has alowerbound and this lowerbound is nonnegative. Recallfrom Section 10.1 that w(x)≥0. 19Weareguessing atthe form of the function. Thenormalization is irrelevant. 1074 Chapter 17 Calculus of Variations Example 17.8.1 VIBRATING STRING Avibratingstring,clampedat x=0 and1,satisfiestheeigenvalueequation d2y dx2+λy=0 (17.138) andtheboundarycondition y(0)=y(1)=0.Forthissimpleexamplewerecognizeimme- diately that y0(x)=sinπx(unnormalized) and λ0=π2. But let us try out the Rayleigh– Ritztechnique. Withoneeyeontheboundaryconditions,wetry y(x)=x(1−x). (17.139) Thenwith p=1 andw=1,Eq. (17.128)yields F[y(x)]=integraltext1 0(1−2x)2dx integraltext1 0x2(1−x)2dx=1/3 1/30=10. (17.140) This result, λ=10, is a fairly good approximation (1.3% error)20ofλ0=π2=9.8696. You may have noted that y(x), Eq. (17.139), is not normalized to unity. The denominator inF[y(x)]compensatesforthelackofunitnormalization. Fmayalsobecalculatedfrom Eq.(17.131)sinceEq. (17.130)issatisfiedby yfrom Eq.(17.139). In the usual scientific calculation the eigenfunction would be improved by introducing moreterms andadjustableparameters,suchas y=x(1−x)+a2x2(1−x)2. (17.141) Itisconvenienttohavetheadditionaltermsorthogonal,butitisnotnecessary.Theparame- tera2isadjustedto minimize F[y(x)].Inthiscase,choosing a2=1.1353 drives F[y(x)] downto9.8697,veryclosetothecorrecteigenvaluevalue. /squaresolid Exercises 17.8.1 From Eq. (17.128) develop in detail the argument when λ≥0o rλ<0. Explain the circumstancesunderwhich λ=0,andillustratewithseveralexamples. 17.8.2 Anunknownfunctionsatisfies thedifferentialequation y′′+parenleftbiggπ 2parenrightbigg2 y=0 andtheboundaryconditions y(0)=1,y(1)=0. 20The closeness of the fit may be checked by a Fourier sine expansion (compare Exercise 14.2.3 over the half-interval [0,1]or, equivalently, over the interval [−1,1], withy(x)taken to be odd). Because of the even symmetry relative to x=1/2, only odd nterms appear: y(x)=x(1−x)=parenleftbigg8 π3parenrightbiggbracketleftbigg sinπx+sin3πx 33+sin5πx 53+···bracketrightbigg . 17.8 Rayleigh–Ritz Variational Technique 1075 (a) Calculatetheapproximation λ=F[ytrial] for ytrial=1−x2. (b) Comparewiththeexacteigenvalue. ANS.(a) λ=2.5, (b) λ/λexact=1.013. 17.8.3 InExercise17.8.2useatrialfunction y=1−xn. (a) Findthevalueof nthatwillminimize F[ytrial]. (b) Showthattheoptimumvalueof ndrivestheratio λ/λexactdownto1.003. ANS.(a)n=1.7247. 17.8.4 Aquantummechanicalparticleina sphere(Example11.7.1)satisfies ∇2ψ+k2ψ=0, withk2=2mE/¯h2.Theboundaryconditionisthat ψ(r=a)=0,whereaistheradius ofthesphere.Forthegroundstate[where ψ=ψ(r)]tryanapproximatewavefunction ψa(r)=1−parenleftbiggr aparenrightbigg2 andcalculateanapproximateeigenvalue k2 a. Hint. To determine p(r)andw(r), put your equation in self-adjoint form (in spherical polarcoordinates). ANS.k2 a=10.5 a2,k2 exact=π2 a2. 17.8.5 Thewaveequationforthequantummechanicaloscillatormaybewrittenas d2ψ(x) dx2+parenleftbig λ−x2parenrightbig ψ(x)=0, withλ=1 forthegroundstate(Eq. (13.18)). Take ψtrial=braceleftBigg 1−x2 a2,x2≤a2 0,x2>a2 for the ground-state wave function (with a2an adjustable parameter) and calculate the correspondingground-stateenergy.Howmucherror doyouhave? Note.YourparabolaisreallynotaverygoodapproximationtoaGaussianexponential. Whatimprovementscanyousuggest? 1076 Chapter 17 Calculus of Variations 17.8.6 TheSchrödingerequationfor acentralpotentialmaybewrittenas Lu(r)+¯h2l(l+1) 2Mr2u(r)=Eu(r). Thel(l+1)term, the angular momentum barrier, comes from splitting off the angu- lar dependence (Section 9.3). Treating this term as a perturbation, use your variational techniquetoshowthat E>E0,whereE0istheenergyeigenvalueof Lu0=E0u0cor- responding to l=0. This means that the minimum energy state will have l=0, zero angularmomentum. Hint.Youcanexpand u(r)asu0(r)+summationtext∞ i=1ciui, where Lui=Eiui,Ei>E0. 17.8.7 Inthematrixeigenvector,eigenvalueequation Ari=λiri, whereλisann×nHermitianmatrix.Forsimplicity,assumethatits nrealeigenvalues (Section3.5) aredistinct, λ1beingthelargest. If ris anapproximationto r1, r=r1+nsummationdisplay i=2δiri, showthat r†Ar r†r≤λ1 andthattheerror in λ1isof theorder|δi|2.T ak e|δi|≪1. Hint.T h enriform a complete orthogonal set spanning the n-dimensional (complex) space. 17.8.8 The variational solution of Example 17.8.1 may be refined by taking y=x(1−x)+ a2x2(1−x)2.Usingthenumericalquadrature,calculate λapprox=F[y(x)],Eq.(17.128), forafixedvalueof a2.V arya2tominimize λ.Calculatethevalueof a2thatminimizes λand calculate λitself, both to five significant figures. Compare your eigenvalue λ withπ2. AdditionalReadings Bliss, G. A., Calculus of Variations . The Mathematical Association of America. LaSalle, IL: Open Court Pub- lishing Co. (1925). As one of the older texts, this is still a valuable reference for details of problems such as minimum-area problems. Courant,R.,andH.Robbins, WhatIsMathematics? 2nded.NewYork:OxfordUniversityPress(1996).Chapter VII contains a fine discussion of the calculus of variations, including soap film solutions to minimum-area problems. Lanczos, C., The Variational Principles of Mechanics , 4th ed. Toronto: University of Toronto Press (1970), reprinted, Dover (1986). This book is a very complete treatment of variational principles and their applica- tions tothe development of classicalmechanics. Sagan, H., Boundary and Eigenvalue Problems in Mathematical Physics . New York: Wiley (1961), reprinted, Dover(1989).ThisdelightfultextcouldalsobelistedasareferenceforSturm–Liouvilletheory,Legendreand Besselfunctions,andFourierSeries.Chapter1isanintroductiontothecalculusofvariations,withapplications to mechanics. Chapter7 picks up the calculus ofvariations again andapplies it to eigenvalue problems. 17.8 Additional Readings 1077 Sagan, H., Introduction to the Calculus of Variations . New York: McGraw-Hill (1969), reprinted, Dover (1983). Thisisanexcellentintroductiontothemoderntheoryofthecalculusofvariations,whichismoresophisticated and complete than his 1961 text. Sagan covers sufficiency conditions and relates the calculus of variations to problems ofspace technology. Weinstock, R., Calculus of Variations . New York: McGraw-Hill (1952); New York: Dover (1974). A detailed, systematic development of the calculus of variations and applications to Sturm–Liouville theory and physical problems in elasticity, electrostatics,andquantum mechanics. Yourgrau,W.,andS.Mandelstam, VariationalPrinciplesinDynamicsandQuantumTheory ,3rded.Philadelphia: Saunders (1968); New York: Dover (1979). This is a comprehensive, authoritative treatment of variational principles. The discussions of the historical development and the many metaphysical pitfalls are of particular interest. This page intentionally left blank CHAPTER 18 NONLINEAR METHODS ANDCHAOS Our mind would lose itself in the complexity of the world if that complexity were not harmonious;liketheshort–sighted,itwouldonlyseethedetails, andwouldbeobliged toforgeteachofthesedetailsbeforeexaminingthenext,becauseitwouldbeincapable oftakinginthewhole.Theonlyfactsworthyofourattentionarethosewhichintroduce orderinto this complexity and somake it accessible to us. HENRIPOINCARÉ 18.1 I NTRODUCTION The origin of nonlinear dynamics goes back to the work of the renowned French mathe- matician Henri Poincaré on celestial mechanics at the turn of the twentieth century. Clas- sical mechanics is, in general, nonlinear in its dependence on the coordinates of the par- ticles and the velocities, one example being vibrations with a nonlinear restoring force. The Navier–Stokes equations are nonlinear, which makes hydrodynamics difficult to han- dle. For almost four centuries however, following the lead of Galileo, Newton, and others, physicists have focused on predictable, effectively linear responses of classical systems, whichusuallyhavelinearandnonlinearproperties. Poincaréwasthefirsttounderstandthepossibilityofcompletelyirregular,or“chaotic,” behavior of solutions of nonlinear differential equations that are characterized by an ex- treme sensitivity to initial conditions: Given slightly different initial conditions, from er- rors in measurements for example, solutions can grow exponentially apart with time, so the system soon becomes effectively unpredictable, or “chaotic.” This property of chaos, often called the “butterfly” effect, will be discussed in Section 18.3. Since the rediscovery ofthiseffectbyLorenzinmeteorologyintheearly1960s,thefieldofnonlineardynamics hasgrowntremendously.Thus,nonlineardynamicsandchaostheorynowhaveenteredthe mainstreamofphysics. 1079 1080 Chapter 18 Nonlinear Methods and Chaos Numerousexamplesofnonlinearsystemshavebeenfoundtodisplayirregularbehavior. Surprisingly,order,inthesenseofquantitativesimilaritiesasuniversalproperties,orother regularities may arise spontaneously in chaos; a first example. Feigenbaum’s universal numbers αandδwill appear in Section 18.2. Dynamical chaos is not a rare phenomenon but is ubiquitous in nature. It includes irregular shapes of clouds, coast lines, and other landscapes, which are examples of fractals, to be discussed in Section 18.3, and turbulent flowoffluids,waterdrippingfromafaucet,andtheweather,ofcourse.Thedamped,driven pendulumis amongthesimplestsystemsdisplayingchaoticmotion. Necessaryconditionsforchaoticmotionindynamicalsystemsdescribedby first-order differentialequationsare •atleastthreedynamicalvariables,and •oneormorenonlineartermscouplingtwoor severalof them. Asinclassicalmechanics,thespaceofthetime-dependentdynamicalvariablesofasystem of coupled differential equations is called its phase space . In such deterministic systems, trajectories in phase space are not allowed to cross. If they did, the system would have a choiceateachintersectionandwouldnotbedeterministic.Intwodimensionssuchnonlin- earsystemsallowonlyforfixedpoints.Anexampleisadampedpendulum,whosesecond derivative,¨θ=f(˙θ,θ), can be written as two first-order derivatives, ω=˙θ,˙ω=f(ω,θ), involving just two dynamic variables, ω(t)andθ(t). In the undamped case, there will onlybeperiodicmotionandequilibriumpoints.Withthreeormoredynamicvariables(for example, damped,drivenpendulum, written as first-order coupled ODEs again), more complicated nonintersecting trajectories are possible. These include chaotic motion and arecalled deterministicchaos . Acentralthemeinchaosistheevolutionof complex formsfromtherepetitionof simple butnonlinear operations;thisisbeingrecognizedasa fundamentalorganizingprinciple ofnature .Whilenonlineardifferentialequationsareanaturalplaceinphysicsforchaosto occur,themathematicallysimpleriterationofnonlinearfunctionsprovidesaquickerentry to chaos theory, which we will pursue first in Section 18.2. In this context, chaos already arisesincertainnonlinearfunctionsofa singlevariable. 18.2 T HELOGISTIC MAP Thenonlinearone-dimensionaliteration,ordifferenceequation, xn+1=µxn(1−xn), x n∈[0,1];1<µ<4, (18.1) is called the logistic map . It is patterned after the nonlinear differential equation dx/dt= µx(1−x),usedbyP.F.Verhulstin1845tomodelthedevelopmentofabreedingpopula- tion whose generations do not overlap. The density of the population at time nisxn.T h e lineartermsimulatesthebirthrateandthenonlineartermthedeathrateofthespeciesina constantenvironmentcontrolledbytheparameter µ. Thequadraticfunction fµ(x)=µx(1−x)ischosenbecauseithasonemaximuminthe interval[0,1]andiszeroattheendpoints, fµ(0)=0=fµ(1).Themaximumat xm=1/2 18.2 The Logistic Map 1081 FIGURE 18.1Cycle(x0,x1,...)forthelogisticmapfor µ=2, startingvalue x0=0.1 andattractor x∗=1/2. isdeterminedfrom f′(x)=0,thatis, f′ µ(xm)=µ(1−2xm)=0,x m=1 2, (18.2) wherefµ(1/2)=µ/4. •Varying the single parameter µcontrols a rich and complex behavior, including one- dimensionalchaos,asweshallsee.Moreparametersoradditionalvariablesarehardly necessaryatthispointtoincreasethecomplexity.Inaratherqualitativesensethesim- plelogisticmap of Eq. (18.1) is representative of many dynamical systems in biology, chemistry,andphysics. Figure (18.1) shows a plot of fµ(x)=µx(1−x)along with the diagonal and a series of points (x0,x1,...)called acycle. To construct a cycle for a fixed value of µ(=2i n Fig. 18.1), we choose some x0∈[0,1][x0=0.1 in Eq. (18.1)]. The vertical line through x0intersectsthecurve fµ(x)atx1=fµ(x0)(=0.18inFig.18.1).Proceedinghorizontally fromx1leads us to x1on the diagonal. Going vertically from the abscissa x1givesx2= fµ(x1)on the curve ( x2=0.2952 in Fig. 18.1), etc. That is, straight vertical lines show the intersections with the curve fµand horizontal lines convert fµ(xi)=xi+1to the next abscissa. For any initial value x0with 0<x0<1,thexiconverge toward the fixed point x∗,o r attractor [=(0.5,0.5)inFig.18.1]: fµ(x∗)=µx∗(1−x∗)=x∗,i.e.,x∗=1−1 µ. (18.3) The interval (0,1)defines a basin of attraction for the fixed point x∗. The attractor x∗is stableprovided the slope |f′ µ(x∗)|=|2−µ|<1, or 1<µ<3. This can be seen from a Taylorexpansionofaniterationneartheattractor: xn+1=fµ(xn)=fµ(x∗)+f′ µ(x∗)(xn−x∗)+···,i.e.,xn+1−x∗ xn−x∗=f′ µ(x∗), 1082 Chapter 18 Nonlinear Methods and Chaos FIGURE 18.2Partof thebifurcationplotfor thelogistic map:fixedpoints x∗versusµ. upon dropping all higher-order terms. Thus, if |f′ µ(x∗)|<1,the next iterate, xn+1, lies closer to x∗than does xn, implying convergence to and stability of the fixed point. How- ever, if|f′ µ(x∗)|>1,xn+1moves farther from x∗than does xnimplying divergence and instability. Given the continuity of f′ µinµ,the fixed point and its properties persist when theparameter(here µ)is slightlyvaried. Forµ>1 andx0<0o rx0>1, it is easy to verify graphically or analytically that thexi→−∞. The origin, x=0, is arepellent fixed point since f′ µ(0)=µ>1 and the iteratesmoveawayfrom it.Since f′ µ(1)=−µ,thepoint x=1 is arepellorfor µ>1. When f′ µ(x∗)=µ(1−2x∗)=2−µ=−1 isreachedfor µ=3,twofixedpointsoccur ,shownasthetwobranchesinFig.18.2,as µ increasesbeyondthevalue3. Theycanbelocatedbysolving x∗ 2=fµparenleftbig fµ(x∗ 2)parenrightbig =µ2x∗ 2(1−x∗ 2)bracketleftbig 1−µx∗ 2(1−x∗ 2)bracketrightbig forx∗ 2. Here it is convenient to abbreviate f(1)(x)=fµ(x),f(2)(x)=fµ(fµ(x))for the second iterate, etc. Now we drop the common x∗ 2and then reduce the remaining third- order polynomial to second-order by recalling that a fixed point of fµis also a fixed point off(2)becausefµ(fµ(x∗))=fµ(x∗)=x∗.S ox∗ 2=x∗is one solution. Factoring out the quadraticpolynomialweobtain 0=µ2bracketleftbig 1−(µ+1)x∗ 2+2(x∗ 2)2−µ(x∗ 2)3bracketrightbig −1 =(µ−1−µx∗ 2)bracketleftbig µ+1−µ(µ+1)x∗ 2+µ2(x∗ 2)2bracketrightbig . Therootsofthequadraticpolynomialare x∗ 2=1 2µparenleftbig µ+1±radicalbig (µ+1)(µ−3)parenrightbig , whicharethetwobranchesinFig.18.2for µ>3startingat x∗ 2=2/3.Thisshowsthatboth fixed points bifurcate at the same value of µ. Eachx∗ 2is a point of period 2 and invariant 18.2 The Logistic Map 1083 under two iterations of the map fµ. The iterates oscillate between both branches of fixed pointsx∗ 2. A point xnis defined as a periodic point of period nforfµiff(n)(x0)= x0,b u tf(i)(x0)/negationslash=x0for 0<i<n. Thus, for 3 <µ<3.45 (see Fig. 18.2) the stable attractorbifurcates ,orsplits,intotwofixedpoints x∗ 2.Thebifurcationfor µ=3,wherethe doublingoccurs,iscalleda pitchfork bifurcationbecauseofitscharacteristic(roundedY-) shape. A bifurcation is a sudden change in the evolution of the system, such as a splitting ofonecurveintotwocurves. Asµincreases beyond 3, the derivative df(2)/dxdecreases from unity to −1. Forµ= 1+√ 6∼3.44949,whichcanbederivedfrom df(2) dxvextendsinglevextendsinglevextendsinglevextendsingle x=x∗=−1,f(2)(x∗)=x∗, each branch of fixed points bifurcates again, so x∗ 4=f(4)(x∗ 4), that is, has period 4. For µ=1+√ 6 theseare x∗ 4=0.43996 and x∗ 4=0.849938. Withincreasingperioddoublingsitbecomesimpossibletoobtainanalyticsolutions.The iterations are better done numerically on a programmable pocket calculator or a personal computer, whose rapid improvements (computer-driven graphics, in particular) and wide distributionsincethe1970shasacceleratedthedevelopmentofchaostheory.Thesequence of bifurcations continues with ever longer periods until we reach µ∞=3.5699456..., where an infinite number of bifurcations occur. Near bifurcation points, fluctuations, roundingerrorsininitialconditions,etc.,playanincreasingrolebecausethesystemhasto choose between two possible branches and becomes much more sensitive to small pertur- bations.Inthepresentcasethe xnneverrepeat.Thebandsoffixedpoints x∗beginforming a continuum (shown dark in Fig. 18.2); this is where chaos starts. This increasing period doublingis the route to chaos for the logistic map that is characterizedby a universalcon- stantδ,calleda Feigenbaumnumber .Ifthefirstbifurcationoccursat µ1=3,thesecond atµ2=3.45,...,thentheratioof spacingsbetweenthe µnconvergesto δ: limn→∞µn−µn−1 µn+1−µn=δ=4.66920161 .... (18.4) FromthebifurcationplotinFig.18.2weobtain µ2−µ1 µ3−µ2=3.45−3.00 3.54−3.45=5.0 as a first approximation for the dimensionless δ. The corresponding critical-period-2 n pointsx∗ nleadtoanotheruniversalanddimensionlessquantity: limn→∞x∗ n−x∗ n−1 x∗ n+1−x∗n=α=2.5029.... (18.5) Againreadingoff Fig.18.2’sapproximatevaluesfor x∗ nweobtain 0.44−0.67 0.37−0.44=3.3 asafirst approximationfor α. 1084 Chapter 18 Nonlinear Methods and Chaos The Feigenbaum number δis universal for the route to chaos via period doublings for all maps with a quadratic maximum similar to the logistic map. It is an example of or- der in chaos. Experience shows that its validity is even wider, including two-dimensional (dissipative)systemsandtwicecontinuouslydifferentiablefunctionswithsubharmonicbi- furcations.1When the maps behave like |x−xm|1+εnear their maximum xmfor some ε between 0 and 1, the Feigenbaum number will depend on the exponent ε; thusδ(ε)varies betweenδ(1)giveninEq. (18.4)for quadraticmapsto δ(0)=2f o rε=0.2 Exercises 18.2.1 Show that x∗=1 is a nontrivial fixed point of the map xn+1=xnexp[r(1−xn)]with aslope 1−r, sothattheequilibriumisstableif 0 <r<2. 18.2.2 Drawabifurcationdiagramfor theexponentialmapofExercise18.2.1for r>1.9. 18.2.3 Determine fixed points of the cubic map xn+1=ax3 n+(1−a)xnfor 0<a<4 and 0<xn<1. 18.2.4 Writethetime-delayedlogisticalmap xn+1=µxn(1−xn−1)asatwo-dimensionalmap xn+1=µxn(1−yn),yn+1=xn, anddeterminesomeof itsfixedpoints. 18.2.5 Show that the second bifurcation for the logistical map that leads to cycles of period 4 islocatedat µ=1+√ 6. 18.2.6 Construct a nonlinear iteration function with Feigenbaum δin the interval 2 <δ< 4.6692.... 18.2.7 Determine the Feigenbaum δfor (a) the exponential map of Exercise 18.2.1, (b) some cubicmapofExercise18.2.3,(c) thetime-delayedlogisticmapof Exercise18.2.4. 18.2.8 RepeatExercise18.2.7for Feigenbaum’s αinsteadof δ. 18.2.9 Findnumericallythefirstfourpoints µforperioddoublingofthelogisticmap,andthen obtain the first two approximations to the Feigenbaum δ. Compare with Fig. 18.2 and Eq. (18.4). 18.2.10 Find numerically the values µwhere the cycle of period 1, 3, 4, 5, 6 begins and then whereit becomesunstable. Checkvalues. Forperiod3,µ=3.8284, 4,µ=3.9601, 5,µ=3.7382, 6,µ=3.6265. 18.2.11 RepeatExercise18.2.9for Feigenbaum’s α. 1More details and computer codes for the logistic map are given by G. L. Baker and J. P. Gollub, Chaotic Dynamics: An Introduction , Cambridge, UK: Cambridge University Press (1990). 2For other maps and a discussion of the fascinating history how chaos became again a hot research topic, see D. Holton and R. M. May in The Nature of Chaos (T. Mullin, ed.), Oxford, UK: Clarendon Press (1993), Section 5, p. 95; and Gleick’s Chaos (1987)—seethe Additional Readings. 18.3 Sensitivity to Initial Conditions and Parameters 1085 18.3 S ENSITIVITY TO INITIAL CONDITIONS AND PARAMETERS Lyapunov Exponents InSection18.2wedescribedhow,asweapproachtheperiod-doublingaccumulationpara- meter value µ∞=3.5699...from below, the period n+1 of cycles ( x0,x1,...,xn) with xn+1=x0getslonger.It isalsoeasytocheckthatthedistances dn=vextendsinglevextendsinglef(n)(x0+ε)−f(n)(x0)vextendsinglevextendsingle (18.6) grow as well for small ε>0. From experience with chaotic behavior we find that this distanceincreasesexponentiallywith n→∞;thatis,dn/ε=eλn,or λ=1 nlnparenleftbigg|f(n)(x0+ε)−f(n)(x0)| εparenrightbigg , (18.7) whereλis aLyapunov exponent for the cycle. For ε→0 we may rewrite Eq. (18.7) in termsofderivativesas λ=1 nlnvextendsinglevextendsinglevextendsinglevextendsingledf(n)(x0) dxvextendsinglevextendsinglevextendsinglevextendsingle=1 nnsummationdisplay i=0lnvextendsinglevextendsinglef′(xi)vextendsinglevextendsingle, (18.8) usingthechainruleof differentiationfor df(n)(x)/dx, where df(2)(x0) dx=dfµ dxvextendsinglevextendsinglevextendsinglevextendsingle x=fµ(x0)dfµ dxvextendsinglevextendsinglevextendsinglevextendsingle x=x0=f′ µ(x1)f′ µ(x0) (18.9) andf′ µ=dfµ/dx, etc. Our Lyapunov exponent has been calculated at the point x0, and Eq.(18.8) isexactfor one-dimensionalmaps. Asameasureofthesensitivityofthesystemtochangesininitialconditions,onepointis notenoughtodetermine λinhigher-dimensionaldynamicalsystemsingeneral,wherethe motionoftenisbounded,sothe dncannotgoto∞.Insuchcases,werepeattheprocedure forseveralpointsonthetrajectoryandaverageoverthem.Thisway,weobtainthe average Lyapunov exponent for the sample. This average value is often called and taken as the Lyapunovexponent. TheLyapunovexponent λisaquantitativemeasureofchaos:Aone-dimensionaliterated function similar to the logistic map has chaoticcycles (x0,x1,...) for the parameter µif the average Lyapunov exponent is positive for that value of µ. Any such initial point x0is called a strangeorchaotic attractor (the shaded region in Fig. 18.2). For cycles of finite period, λis negative. This is the case for µ<3, forµ<µ∞, and even in the periodicwindowat µ∼3.627insidethechaoticregionofFig. 18.2.Atbifurcationpoints, λ=0. Forµ>µ∞the Lyapunov exponent is positive, except in the periodic windows, whereλ<0,andλgrowswith µ.Inotherwords,thesystembecomesmorechaoticasthe controlparameter µincreases. In the chaos region of the logistic map there is a scaling law for the average Lyapunov exponent(wedonotderiveit), λ(µ)=λ0(µ−µ∞)ln2/lnδ, (18.10) 1086 Chapter 18 Nonlinear Methods and Chaos where ln2 /lnδ∼0.445,δis the universal Feigenbaum number of Section 18.2, and λ0 is a constant. This relation (18.10) is reminiscent of a physical observable at a (second- order) phase transition. The exponent in Eq. (18.10) is a universal number; the Lyapunov exponent plays the role of an order parameter , whileµ−µ∞is the analog of T−Tc, whereTcis thecriticaltemperatureatwhichthephasetransitionoccurs. Fractals In dissipative chaotic systems (but rarely in conservative Hamiltonian systems) often new geometric objects with intricate shapes appear that are called fractalsbecause of their noninteger dimension. Fractals are irregular geometric objects whose dimension is typi- cally not integral and that exist at many scales, so their smaller parts resemble their larger parts. Intuitively a fractal is a set which is (approximately) self-similar under magnifica- tion.Asetofattractingpointswithnonintegerdimensionis calleda strangeattractor. Weneedaquantitativemeasureofdimensionalityinordertodescribefractals.Unfortu- nately, there are several definitions with usually different numerical values, none of which has yet become a standard. For strictly self-similar, sets, one measure suffices. More com- plicated(forinstance,onlyapproximatelyself-similar)setsrequiremoremeasuresfortheir complete description. The simplest is the box-counting dimension , due to Kolmogorov and Hausdorff. For a one-dimensional set, we cover the curve by line segments of length R. In two dimensions the boxes are squares of area R2, in three dimensions cubes of vol- umeR3,etc.Thenwecountthenumber N(R)ofboxesneededtocovertheset.Letting R go to zero we expect Nto scale as N(R)∼R−d. Taking the logarithm the box-counting dimension isdefinedas d≡lim R→0lnN(R) lnR. (18.11) For example, in a two-dimensional space a single point is covered by one square, so lnN(R)=0 andd=0. A finite set of isolated points also has dimension d=0. For a differentiable curve of length L,N(R)∼L/RasR→0, sod=1 from Eq. (18.11), as expected. Let us now construct a more irregular set, the Kochcurve. We start with a line segment of unit length in Fig. 18.3 and remove the middle third. Then we replace it with two seg- ments of length 1 /3, which form a triangle in Fig. 18.3. We iterate this procedure with each segment ad infinitum. The resulting Koch curve is infinitely long and is nowhere dif- ferentiable because of the infinitely many discontinuous changes of slope. At the nth step each line segment has length Rn=3−nand there are N(Rn)=4nsegments. Hence its dimension is d=ln4/ln3=1.26...,which is more than a curve but less than a surface. BecausetheKochcurveresults fromiterationofthefirst step,itis strictlyself-similar. For the logistic map the box-counting dimension at a period–doubling accumulation pointµ∞is 0.5388...,which is a universal number for iterations of functions in one variable with a quadratic maximum. To see roughly how this comes about, consider the pairsoflinesegmentsoriginatingfromsuccessivebifurcationpointsforagivenparameter µinthechaosregime(seeFig.18.2).Imagineremovingtheinteriorspacefromthechaotic bands. When we go to the next bifurcation, the relevant scale parameter is α=2.5029... from Eq. (18.5). Suppose we need 2nline segments of length Rto cover 2nbands. In the 18.3 Sensitivity to Initial Conditions and Parameters 1087 FIGURE 18.3ConstructionoftheKochcurveby iterations. next stage then we need 2n+1segments of length R/αto cover the bands. This yields a dimension d=−ln(2n/2n+1)/lnα=0.4498.... This crude estimatecan be improvedby taking into account that the width between neighboring pairs of line segments differs by 1/α(see Fig. 18.2). The improved estimate, 0.543, is closer to 0.5388 .... This example suggests that when the fractal set does not have a simple self-similar structure, then the box-countingdimensiondependsonthebox-constructionmethod. Finally, we turn to the beautiful fractals that are surprisingly easy to generate and whose color pictures had considerable impact. For complex c=a+ib, the correspond- ingquadraticcomplexmapinvolvingthecomplexvariable z=x+iy, zn+1=z2 n+c, (18.12) looks deceptively simple, but the equivalent two-dimensional map in terms of the real variables xn+1=x2 n−y2 n+a, y n+1=2xnyn+b (18.13) revealsalreadymoreofitscomplexity.ThismapformsthebasisforsomeofMandelbrot’s beautifulmulticolorfractalpictures(wereferthereadertoMandelbrot(1988)andPeitgen and Richter (1986) in the Additional Readings), and it has been found to generate rather intricate shapes for various c/negationslash=0. For example, the Julia set of a map zn+1=F(zn)is defined as the set of all its repelling fixed or periodic points. Thus it forms the boundary betweeninitialconditionsofatwo-dimensionaliteratedmapleadingtoiteratesthatdiverge and those that stay within some finite region of the complex plane. For the case c=0 and F(z)=z2, the Julia set can be shown to be just a circle about the origin of the complex plane. Yet, just by adding a constant c/negationslash=0, the Julia set becomes fractal. For instance, for c=−1 one finds a fractal necklace with infinitely many loops (see Devaney (1989) in the AdditionalReadings). While the Julia set is drawn in the complex plane, the Mandelbrot set is constructed in the two-dimensional parameter space c=(a,b)=a+bi. It is constructed as follows. 1088 Chapter 18 Nonlinear Methods and Chaos Starting from the initial value z0=0=(0,0)one searches Eq. (18.12) for parameter val- uescsothattheiterated {zn}donotdivergeto∞.Eachcoloroutsidethefractalboundary of the Mandelbrot set represents a given number of iterations m, say, needed for the zn to go beyond a specified absolute (real) value R,|zm|>R>|zm−1|. For real parame- ter value c=a, the resulting map, xn+1=x2 n+a, is equivalent to the logistic map with period-doubling bifurcations (see Section 18.2) as aincreases on the real axis inside the Mandelbrotset. Exercises 18.3.1 Use a programmable pocket calculator (or a personal computer with BASIC or FOR- TRAN or symbolic software such as Mathematica or Maple) to obtain the iterates xi of an initial 0 <x0<1 andf′ µ(xi)for the logistic map. Then calculate the Lyapunov exponent for cycles of period 2 ,3,...of the logistic map for 2 <µ<3.7. Show that forµ<µ∞theLyapunovexponent λis0atbifurcationpointsandnegativeelsewhere, whilefor µ>µ∞itis positiveexceptinperiodicwindows. Hint.SeeFig.9.3 ofHilborn(1994)intheAdditionalReadings. 18.3.2 Considerthemap xn+1=F(xn)with F(x)=braceleftBigg a+bx, x < 1, c+dx, x> 1, forb>0andd<0.ShowthatitsLyapunovexponentispositivewhen b>1,d<−1. Plota fewiterationsinthe (xn+1,xn)plane. 18.4 N ONLINEAR DIFFERENTIAL EQUATIONS In Section 18.1 we mentioned nonlinear differential equations (abbreviated as NDEs) as the natural place in physics for chaos to occur, but continued with the simpler iteration of nonlinearfunctionsofonevariable(maps).Herewebrieflyaddressthemuchbroaderarea of NDEs and the far greater complexity in the behavior of their solutions. However, maps and systems of solutions of NDEs are closely related. The latter can often be analyzed in terms of discrete maps. One prescription is the so-called Poincaré section of a system of NDE solutions. Placing a plane transverse into a trajectory (of a solution of a NDE), it in- tersectstheplaneinaseriesofpointsatincreasingdiscretetimes,forexample,inFig.18.4 (x(t1),y(t1))=(x1,y1),(x2,y2),...,which are recorded and graphically or numerically analyzed for fixed points, period-doubling bifurcations, etc. This method is useful when solutions of NDEs are obtained numerically in computer simulations so that one can gen- erate Poincaré sections at various locations and with different orientations, with further analysisleadingtotwo-dimensionaliteratedmaps xn+1=F1(xn,yn), y n+1=F2(xn,yn) (18.14) stored by the computer. Extracting the functions Fjanalytically or graphically is not al- wayseasy,though. Let us start with a few classical examples of NDEs. In Chapter 9 we have already dis- cussedthesolitonsolutionofthenonlinearKorteweg–deVries PDE,Eq. (9.11). 18.4 Nonlinear Differential Equations 1089 FIGURE 18.4Schematicof aPoincarésection. Exercise 18.4.1 Forthedampedharmonicoscillator ¨x+2a˙x+x=0, consider the Poincaré section {x>0,y=˙x=0}.T a k e0<a≪1 and show that the mapis givenby xn+1=bxnwithb<1.Findanestimatefor b. Bernoulli and Riccati Equations Bernoulliequationsarealsononlinear,havingtheform y′(x)=p(x)y(x)+q(x)bracketleftbig y(x)bracketrightbign, (18.15) wherepandqare real functions and n/negationslash=0, 1 to exclude first-order linear ODEs. If we substitute u(x)=bracketleftbig y(x)bracketrightbig1−n, (18.16) thenEq. (18.15)becomesafirst-order linearODE, u′=(1−n)y−ny′=(1−n)bracketleftbig p(x)u(x)+q(x)bracketrightbig , (18.17) whichwecansolveasdescribedinSection9.2. Riccatiequationsare quadraticin y(x): y′=p(x)y2+q(x)y+r(x), (18.18) 1090 Chapter 18 Nonlinear Methods and Chaos wherep/negationslash=0toexcludelinearODEsand r/negationslash=0toexcludeBernoulliequations.Thereisno general method for solving Riccati equations. However, when a special solution y0(x)of Eq. (18.18) is known by a guess or inspection, then one can write the general solution in theformy=y0+u, withusatisfyingtheBernoulliequation u′=pu2+(2py0+q)u, (18.19) becausesubstitutionof y=y0+uintoEq. (18.18)removes r(x)fromEq. (18.18). JustasforRiccatiequationstherearenogeneralmethodsforobtainingexactsolutionsof other nonlinear ODEs. It is more important to develop methods for finding the qualitative behavior of solutions. In Chapter 9 we mentioned that power-series solutions of ODEs exist except (possibly) at regular or essential singularities, which are directly given by localanalysisofthecoefficientfunctionsoftheODE.Suchlocalanalysisprovidesuswith theasymptoticbehaviorofsolutionsaswell. Fixed and Movable Singularities, Special Solutions Solutions of NDEs also have such singular points, independent of the initial or boundary conditionsandcalled fixedsingularities .Inadditiontheymayhave spontaneous ,ormov- able, singularities that vary with the initial or boundary conditions. They complicate the (asymptotic)analysisofNDEs.ThispointisillustratedbyacomparisonofthelinearODE y′+y x−1=0, (18.20) which has the obvious regular singularity at x=1, with the NDE y′=y2. Both have the same solution with initial condition y(0)=1, namely, y(x)=1/(1−x).F o ry(0)=2, though, the pole in the (obvious, but check) solution y(x)=2/(1−2x)of the NDE has moved to x=1/2. For a second-order ODE we have a complete description of (the asymptotic behavior of) its solutions when (that of) two linearly independent solutions are known. For NDEs there may still be special solutions whose asymptotic behavior is not obtainable from two independent solutions. This is another characteristic property of NDEs, which we illustrateagainbyanexample. ThegeneralsolutionoftheNDE y′′=yy′/xis givenby y(x)=2c1tan(c1lnx+c2)−1, (18.21) whereciare integration constants. An obvious (check it) special solution is y=c3= constant, which cannot be obtained from Eq. (18.20) for any choice of the parameters c1,c2. Note that using the substitution x=et,Y(t)=y(et)so thatxdy/dx=dY/dt,w e obtain the ODE Y′′=Y′(Y+1). This ODE can be integrated once to give Y′=1 2Y2+ Y+cwithc=2(c2 1+1/4)an integration constant, and again according to Section 9.2 to leadtothesolutionofEq. (18.21). 18.4 Nonlinear Differential Equations 1091 Autonomous Differential Equations Differential equations that do not explicitly contain the independent variable, taken to be the timethere, are called autonomous . Verhulst’s NDE ˙y=dy/dt=µy(1−y), which weencounteredbrieflyin Section18.2 as motivationfor thelogistic map,is a specialcase of this wide and important class of ODEs.3For one dependent variable y(t)they can be writtenas ˙y=f(y), (18.22a) andfor severaldependentvariablesas asystem ˙yi=fi(y1,y2,...,yn), i=1,2,...,n, (18.22b) with sufficiently differentiable functions f,fi. A solution of Eq. (18.22b) is a curve or trajectory y(t)forn=1 and in general a trajectory (y1(t),y2(t),...,y n(t))in ann- dimensional(so-called) phasespace .AsdiscussedalreadyinSection18.1,twotrajectories cannot cross because of the uniqueness of the solutions of ODEs. Clearly, solutions of the algebraicsystem fi(y1,y2,...,yn)=0 (18.23) arespecialpointsinphasespace,wherethepositionvector( y1,y2,...,yn)doesnotmove onthetrajectory;theyarecalled critical(orfixed)points.Itturnsoutthatalocalanalysis of solutions near critical points leads to an understanding of the global behavior of the solutions.Firstletuslookatasimpleexample. ForVerhulst’sODE, f(y)=µy(1−y)=0g i v e sy=0andy=1 asthecriticalpoints. For the logistic map, y=0 andy=1 are repellent fixed points because df/dy(0)=µ aty=0 anddf/dy(1)=−µaty=1f o rµ>1. A local analysis near y=0 suggests neglecting the y2term and solving ˙y=µyinstead. Integratingintegraltext dy/y=µt+lncgives the solution y(t)=ceµt, which diverges as t→∞,s oy=0 is a repellent critical point. (Notethatfor µ<0ofthelogisticmapthecriticalpoint y=0wouldbeattracting,leading to a converging y∼eµtsolution.) Similarly at y=1,integraltext dy/(1−y)=µt−lncleads to y(t)=1−ce−µt→1f o rt→∞. Hencey=1 is an attracting critical point. Because the ODEis separable,itsgeneralsolutionis givenby integraldisplaydy y(1−y)=integraldisplay dybracketleftbigg1 y+1 1−ybracketrightbigg =lny 1−y=µt+lnc. Hencey(t)=ceµt/(1+ceµt)fort→∞convergesto 1, thus confirming the local analy- sis. This examplemotivatesus to look nextat the properties of fixedpoints in more detail. Foranarbitraryfunction f,itiseasytoseethat •inone dimension , fixed points yiwithf(yi)=0d i v i d et h e y-axis into dynamically separate intervals because, given an initial value in one of the intervals, the trajectory y(t)willstaythere,for itcannotgobeyondeitherfixedpointwhere ˙y=0. 3Solutions of nonautonomous equations canbe much more complicated. 1092 Chapter 18 Nonlinear Methods and Chaos FIGURE 18.5Fixedpoints:(a) repellor,(b) sink. Iff′(y0)>0 atthefixedpoint y0wheref(y0)=0,thenat y0+εforε>0 sufficiently small,˙y=f′(y0)ε+O(ε2)>0inaneighborhoodtotherightof y0,sothetrajectory y(t) keepsmovingtotheright,awayfromthefixedpoint y0.Totheleftof y0,˙y=−f′(y0)ε+ O(ε2)<0,sothetrajectorymovesawayfrom thefixedpointhereaswell.Hence, •a fixed point [with f(y0)=0] aty0withf′(y0)>0, as shown in Fig. 18.5a, repels trajectories;thatis,alltrajectoriesmoveawayfromthecriticalpoint: ···←·→··· ;it isarepellor.Similarly,weseethat •a fixed point at y0withf′(y0)<0, as shown in Fig. 18.5b, attracts trajectories; that is, all trajectories converge toward the critical point y0:···→·←··· ;i ti sasink or node. Letusnowconsidertheremainingcasewhenalso f′(y0)=0. Let us assume f′′(y0)>0. Then at y0+εto the right of fixed point y0,˙y= f′′(y0)ε2/2+O(ε3)>0, so the trajectory moves away from the fixed point there, while to the left it moves closer to y0. In other words, we have a saddle point .F o rf′′(y0)<0, the sign of˙yis reversed, so we deal again with a saddle point with the motion to the right ofy0toward the fixed point and at left away from it. Let us summarize the local behavior oftrajectoriesnearsuchafixedpoint y0:W eha v ea •a saddle point at y0whenf(y0)=0, andf′(y0)=0, as shown in Fig. 18.6a,b corre- sponding to the cases where (a) f′′(y0)>0 and trajectories on one side of the critical point, converge toward it and diverge from it on the other side: ···→·→··· ; and (b)f′′(y0)<0. Here the direction is simply reversed compared to (a). Figure 18.6(c) showsthecaseswhere f′′(y0)=0. So far we have ignored the additional dependence of f(y)on one or more parameters, such asµfor the logistic map. When a critical point maintains its properties qualitatively as we adjust a parameter slightly, we call it structurally stable . This is reasonable be- cause structurally unstable objects are unlikely to occur in reality because noise and other neglected degrees of freedom act as perturbations on the system that effectively prevent such unstable points from being observed. Let us now look at fixed points from this point of view. Upon varying such a control parameter slightly we deform the function f,o rw e 18.4 Nonlinear Differential Equations 1093 FIGURE 18.6Saddlepoints. may just shift fup or down or sideways in Fig. 18.5 a bit. This will move a little the locationy0of the fixed point with f(y0)=0, but maintain the sign of f′(y0). Thus, both sinks and repellors are stable , while a saddle point in general is not. For example, shift- ingfinFig. 18.6adownabitcreatestwo fixedpoints,oneasinkandtheotherarepellor, and removes the saddle point. Since two conditions must be satisfied at a saddle point, they are less common and important, being unstable with respect to variations of parame- ters. However, they mark the border between different types of dynamics and are useful and meaningful for the global analysis of the dynamics. We are now ready to consider the richer,butmorecomplicated,higher-dimensionalcases. Local and Global Behavior in Higher Dimensions In two or more dimensions we start the local analysis at a fixed point (y0 1,y0 2,...)with ˙yi=fi(y0 1,y0 2,...)=0 using the same Taylor expansion of the fiin Eq. (18.22b) as for the one-dimensional case. Retaining only the first-order derivatives, this approach lin- earizes the coupled NDEs of Eq. (18.22b) and reduces their solution to linear algebra as follows. We abbreviate the constant derivatives at the fixed point as a matrix Fwith ele- ments fij≡∂fi ∂yjvextendsinglevextendsinglevextendsinglevextendsingle (y0 1,y0 2,···). (18.24) 1094 Chapter 18 Nonlinear Methods and Chaos IncontrasttothestandardlinearalgebrainChapter3,however, Fisneithersymmetricnor Hermitianingeneral.Asaresult,itseigenvaluesmaynotbereal.Ifweshiftthefixedpoint to the origin and call the shifted coordinates xi=yi−y0 i, then the coupled NDEs of Eq. (18.22b)become ˙xi=summationdisplay jfijxj, (18.25) that is, coupled linear ODEs with constant coefficients. We solve Eq. (18.25) with the standardexponentialAnsatz, xi(t)=summationdisplay jcijeλjt, (18.26) with constant exponents λjand a constant matrix Cof coefficients cij,s ocj=(cij,i= 1,2,...)formsthe jthcolumnvectorof C.SubstitutingEq.(18.26)intoEq.(18.25)yields alinearcombinationof exponentialfunctions, summationdisplay jcijλjeλjt=summationdisplay j,kfikckjeλjt, (18.27) whichareindependentif λi/negationslash=λj.Thisisthegeneralcaseonwhichwefocus,whiledegen- eracieswheretwoormore λareequalrequirespecialtreatmentsimilartosaddlepointsin one dimension. Comparing coefficients of exponential functions with the same exponent yieldsthelineareigenvalueequations summationdisplay kfikckj=λjcij,orFcj=λjcj. (18.28) A nontrivial solution comprising the eigenvalue λjand eigenvector cjof the homoge- neous linear equations (18.28) requires λjto be a root of the secular equation (compare withSection3.5): det(F−λ·1)=0. (18.29) Equation (18.28) means that Cdiagonalizes F, so we can write Eq. (18.28) also as C−1FC=[λ1,λ2,...]. (18.30) In the new but in general nonorthogonal coordinates ξj, defined as Cξ=x,w eh a v e a fixed point for each direction ξj,a s˙ξj=λjξj, where the λjplay the role of f′(y0) in the one-dimensional case. The λarecharacteristic exponents and complex num- bers in general. This is seen by substituting x=Cξinto Eq. (18.25) in conjunction with Eqs. (18.28) and (18.30). Thus, this solution represents the independent combination of one-dimensional fixed points, one for each component of ξand each independent of the other components. In two dimensions for λ1<0 andλ2<0, then, we have a sink in all directions.Whenboth λaregreaterthan 0,wehaverepellorinalldirections. 18.4 Nonlinear Differential Equations 1095 Example 18.4.1 STABLE SINK ThecoupledODEs ˙x=−x,˙y=−x−3y haveanequilibriumpointattheorigin.Thesolutionshavetheform x(t)=c11eλ1t,y(t)=c21eλ1t+c22eλ2t, so the eigenvalue λ1=−1 results from λ1c11=−c11, and the solution is x=c11e−t.T h e determinantofEq. (18.29), vextendsinglevextendsinglevextendsinglevextendsingle−1−λ0 −1−3−λvextendsinglevextendsinglevextendsinglevextendsingle=(1+λ)(3+λ)=0, yieldstheeigenvalues λ1=−1,λ2=−3.Becausebotharenegativewehaveastablesink attheorigin.TheODEfor ygivesthelinearrelations λ1c21=−c11−3c21=−c21,λ 2c22=−3c22, from which we infer 2 c21=−c11,orc21=−c11/2. Because the general solution will containtwoconstants,itis givenby x(t)=c11e−t,y(t)=−c11 2e−t+c22e−3t. Asthetime t→∞,wehavey∼−x/2andx→0andy→0,whilefor t→−∞,y∼x3 andx,y→±∞. The motion toward the sink is indicated by arrows in Fig. 18.7. To find theorbit, weeliminatetheindependentvariable, t, andfindthecubics: y=−x 2+c22 c3 11x3./squaresolid When both λare greater than 0, we have repellor. In this case the motion is away from the fixed point. However, when the λhave different signs, we have a saddle point, that is, acombinationofasinkinonedimensionandarepellorintheother.Thistypeofbehavior generalizestohigherdimensions. Example 18.4.2 SADDLE POINT ThecoupledODEs ˙x=−2x−y,˙y=−x+2y haveafixedpointattheorigin.Thesolutionshavetheform x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t. Theeigenvalues λ=±√ 5 aredeterminedfrom vextendsinglevextendsinglevextendsinglevextendsingle−2−λ−1 −12−λvextendsinglevextendsinglevextendsinglevextendsingle=λ2−5=0. 1096 Chapter 18 Nonlinear Methods and Chaos FIGURE 18.7Stablesink. SubstitutingthegeneralsolutionsintotheODEsyieldsthelinearequations λ1c11=−2c11−c21=√ 5c11,λ 1c21=−c11+2c21=√ 5c21, λ2c12=−2c12−c22=−√ 5c12,λ 2c22=−c12+2c22=−√ 5c22, or (√ 5+2)c11=−c21,(√ 5−2)c12=c22, (√ 5−2)c21=−c11,(√ 5+2)c22=c12, soc21=−(2+√ 5)c11,c22=(√ 5−2)c12. The family of solutions depends on two parameters, c11,c12. For large time t→∞, the positive exponent prevails and y∼ −(√ 5+2)x,while for t→−∞we have y=(√ 5−2)x. These straight lines are the asymptotesoftheorbits.Because −(√ 5+2)(√ 5−2)=−1 theyareorthogonal.Wefind the orbits by eliminating the independent variable, t, as follows. Substituting the c2jwe write y=−2x−√ 5parenleftbig c11e√ 5t−c12e−√ 5tparenrightbig ,soy+2x√ 5=−c11e√ 5t+c12e−√ 5t. Nowweaddandsubtractthesolution x(t)toget 1√ 5(y+2x)+x=2c12e−√ 5t,1√ 5(y+2x)−x=−2c11e√ 5t, 18.4 Nonlinear Differential Equations 1097 FIGURE 18.8Saddlepoint. whichwemultiplytoobtain 1 5(y+2x)2−x2=−4c12c11=const. The resulting quadratic form, y2+4xy−x2=const., is a hyperbola because of the negative sign. The hyperbola is rotated in the sense that its asymptotes are not aligned with the x,y-axes (Fig. 18.8). Its orientation is given by the direction of the as- ymptotes that we found earlier. Alternatively we could find the direction of minimal distance from the origin, proceeding as follows. We set the xandyderivatives of f+/Lambda1g≡x2+y2+/Lambda1(y2+4xy−x2)equaltozero,where /Lambda1istheLagrangemultiplier for the hyperbolic constraint. The four branches of hyperbolas correspond to the differ- ent signs of the parameters c11andc12. Figure 18.8 is plotted for the cases c11=±1, c12=±2. /squaresolid However, a new kind of behavior arises for a pair of complex conjugate eigenvalues λ1,2=ρ±iκ. If we write the complex solutions ξ1,2=exp(ρt±iκt)in real variables ξ+=(ξ1+ξ2)/2,ξ−=(ξ1−ξ2)/2iuponusingtheEuleridentityexp (ix)=cosx+isinx (seeSection6.1), ξ+=exp(ρt)cos(κt), ξ −=exp(ρt)sin(κt) (18.31) describe a trajectory that spirals inward to the fixed point at the origin for ρ<0, aspiral node,andspiralsawayfromthefixedpointfor ρ>0,aspiralrepellor . 1098 Chapter 18 Nonlinear Methods and Chaos Example 18.4.3 SPIRAL FIXED POINT ThecoupledODEs ˙x=−x+3y,˙y=−3x+2y haveafixedpointattheoriginandsolutionsoftheform x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t. Theexponents λ1,2aresolutionsof vextendsinglevextendsinglevextendsinglevextendsingle−1−λ3 −32−λvextendsinglevextendsinglevextendsinglevextendsingle=(1+λ)(λ−2)+9=0, orλ2−λ+7=0.Theeigenvaluesarecomplexconjugate, λ=1/2±i√ 27/2,sowedeal withaspiralfixedpointattheorigin(arepellorbecause1 /2>0).Substitutingthegeneral solutionsintotheODEsyieldsthelinearequations λ1c11=−c11+3c21,λ 1c21=−3c11+2c21, λ2c12=−c12+3c22,λ 2c22=−3c12+2c22, or (λ1+1)c11=3c21,(λ1−2)c21=−3c11, (λ2+1)c12=3c22,(λ2−2)c22=−3c12, which,usingthevaluesof λ1,2, implythefamilyofcurves x(t)=et/2parenleftbig c11ei√ 27t/2+c12e−i√ 27t/2parenrightbig , y(t)=x 2+√ 27 6et/2iparenleftbig c11ei√ 27t/2−c12e−i√ 27t/2parenrightbig , whichdependsontwoparameters, c11,c12.Tosimplifywecanseparaterealandimaginary partsofx(t)andy(t)usingtheEuleridentity eix=cosx+isinx.Itisequivalent,butmore convenient, to choose c11=c12=c/2 and rescale t→2t,so with the Euler identity we have x(t)=cetcos(√ 27t), y(t)=x 2−√ 27 6cetsin(√ 27t). Herewecaneliminate tandfindtheorbit x2+4 3parenleftbigg y−x 2parenrightbigg2 =parenleftbig cetparenrightbig2. For fixed tthis is the positive definite quadratic form x2−xy+y2=const., that is, an ellipse. But there is no ellipse in the solutions because tis not fixed. Nonetheless, it is useful to find its orientation. We proceed as follows. With /Lambda1the Lagrange multiplier for the elliptical constraint we seek the directions of maximal and minimal distance from the origin,forming f(x,y)+/Lambda1g(x,y)≡x2+y2+/Lambda1parenleftbig x2−xy+y2parenrightbig 18.4 Nonlinear Differential Equations 1099 FIGURE 18.9Spiralpoint. andsetting ∂(f+/Lambda1g) ∂x=2x+2/Lambda1x−/Lambda1y=0,∂(f+/Lambda1g) ∂y=2y+2/Lambda1y−/Lambda1x=0. From 2(/Lambda1+1)x=/Lambda1y, 2(/Lambda1+1)y=/Lambda1x weobtainthedirections x y=/Lambda1 2(/Lambda1+1)=2(/Lambda1+1) /Lambda1, or/Lambda12+8 3/Lambda1+4 3=0. This yields the values /Lambda1=−2/3,−2 and the directions y=±x. In other words, our ellipse is centered at the origin and rotated by 45◦. As we vary the independent variable, t, the size of the ellipse changes, so we get the rotated spiral shown inFig.18.9for c=1. /squaresolid In the special case when ρ=0 in Eq. (18.31), the circular trajectory is called a cycle. Whentrajectoriesnearitareattractedastimegoeson,itiscalleda limitcycle ,representing periodicmotionforautonomoussystems. 1100 Chapter 18 Nonlinear Methods and Chaos FIGURE 18.10Center. Example 18.4.4 CENTER OR CYCLE TheundampedlinearharmonicoscillatorODE ¨x+ω2x=0canbewrittenastwocoupled ODEs: ˙x=−ωy,˙y=ωx. Integrating the resulting ODE ˙xx+˙yy=0 yields the circular orbits x2+y2=const., whichdefineacenterattheoriginandareshowninFig.18.10.Thesolutionscanbepara- meterizedas x=Rcost,y=Rsint,whereRistheradiusparameter.Theycorrespondto thecomplexconjugateeigenvalues λ1,2=±iω.Wecancheckthemifwewritethegeneral solutionas x(t)=c11eλ1t+c12eλ2t,y(t)=c21eλ1t+c22eλ2t. Thentheeigenvaluesfollowfrom vextendsinglevextendsinglevextendsinglevextendsingle−λ−ω ω−λvextendsinglevextendsinglevextendsinglevextendsingle=λ2+ω2=0=0. /squaresolid Anotherclassicattractorisquasiperiodicmotion,suchasthetrajectory x(t)=A1sin(ω1t+b1)+A2sin(ω2t+b2), (18.32) 18.4 Nonlinear Differential Equations 1101 where the ratio ω1/ω2is an irrational number. Such combined oscillations occur as solu- tionsofadampedanharmonicoscillator(VanderPolnonautonomoussystem) ¨x+2γ˙x+ω2 2x+βx3=fcos(ω1t). (18.33) In three dimensions, when there is a positive characteristic exponent and an attracting complex conjugate pair in the other directions, we have a spiral saddle point as a new feature.Conversely,anegativecharacteristicexponentinconjunctionwitharepellingpair also gives rise to a spiral saddle point, where trajectories spiral out in two dimensions but areattractedinathirddirection. In general, when some form of damping (or dissipation of energy) is present, the tran- sients decay and the system settles either in equilibrium, that is, a single point, or in pe- riodic or quasiperiodic motion. Chaotic motion in dissipative systems is now recognized as a fourth state, and its attractors are often called strange. In dissipative systems, initial conditions are not important because trajectories end up on some attractor. They are cru- cialinHamiltoniansystems.InnonintegrableHamiltoniansystems,chaosmayalsooccur, and then it is called conservative chaos. We refer to Chapter 8 of Hilborn (1994) in the AdditionalReadingsfor thismorecomplicatedtopic. For the driven damped pendulum when trajectories near a center (closed orbit) are at- tracted to it as time goes on, this closed orbit is defined as a limit cycle , representing periodicmotionforautonomoussystems.Adampedpendulumusuallyspiralsintotheori- gin (the position at rest); that is, the origin is a spiral fixed point in its phase space. When we turn on a driving force, then the system formally becomes nonautonomous,because of itsexplicittimedependence,butalsomoreinteresting.Inthiscase,wecancalltheexplicit timeinasinusoidaldrivingforceanewvariable, ϕ,whereω0isafixedrate,intheequation ofmotion ˙ω+γω+sinθ=fsinϕ, ω=˙θ, ϕ=ω0t. Then we increase the dimension of our phase space by 1 (adding one variable, ϕ) because ˙ϕ=ω0=const., but we keep the coupled ODEs autonomous. This driven damped pen- dulum has trajectories that cross a closed orbit in phase space and spiral back to it; it is calledlimit cycle . This happens for a range of strength fof the driving force, the control parameter of the system. As we increase f, the phase space trajectories go through sev- eral neighboring limit cycles and eventually become aperiodic and chaotic. Such closed limit cycles are called Hopf bifurcations of the pendulum on its road to chaos, after the mathematician E. Hopf, who generalized Poincaré’s results on such bifurcations to higher dimensionsofphasespace. Such spiral sinks or saddle points cannot occur in one dimension, but we might ask if theyarestablewhentheyoccurinhigherdimensions.AnanswerisgivenbythePoincaré– Bendixson theorem, which says that either trajectories (in the finite region to be specified in a moment) are attracted to a fixed point as time goes on, or they approach a limit cycle provided the relevant two-dimensional subsystem stays inside a finite region, that is, does not diverge there as t→∞. For a proof we refer to Hirsch and Smale (1974) and Jackson (1989)intheAdditionalReadings. In general, when some form of damping is present, the transients decay and the system settles either in equilibrium, that is, a single point, or in periodic or quasiperiodic motion. Chaoticmotionisnowrecognizedas afourthstate,anditsattractorsareoften strange. 1102 Chapter 18 Nonlinear Methods and Chaos Exercise 18.4.2 Showthatthe(Rössler) coupledODEs ˙x1=−x2−x3,˙x2=x1+a1x2,˙x3=a2+(x1−a3)x3 (a) havetwofixedpointsfor a2=2,a3=4,and 0<a1<2, (b) haveaspiralrepellorattheorigin,and (c) haveaspiralchaoticattractorfor a1=0.398. Dissipation in Dynamical Systems Dissipativeforcesofteninvolvevelocities,thatis,first-ordertimederivatives,suchasfric- tion (for example, for the damped oscillator). Let us look for a measure of dissipation, that is, how a small area A=c1,2/Delta1ξ1/Delta1ξ2at a fixed point shrinks or expands, first in two dimensions for simplicity. Here c1,2≡sin(ˆξ1,ˆξ2), involving the sine of the characteristic directions, is a time-independentangular factor that takes into accountthe nonorthogonal- ityofthecharacteristicdirections ˆξ1andˆξ2ofEq.(18.28).Ifwetakethetimederivativeof Aand use˙ξj=λjξjof the characteristic coordinates, implying ˙/Delta1ξj=λj/Delta1ξj, we obtain, tolowestorder inthe /Delta1ξj, ˙A=c1,2[/Delta1ξ1λ2/Delta1ξ2+/Delta1ξ2λ1/Delta1ξ1]=c1,2/Delta1ξ1/Delta1ξ2(λ1+λ2). (18.34) Inthelimit /Delta1ξj→0,wefindfromEq. (18.34)thattherate ˙A A=λ1+λ2=trace(F)=∇·f|y0, (18.35) withf=(f1,f2)thevectoroftimeevolutionfunctionsofEq.(18.22b).Notethatthetime- independent sine of the angle between ξ1andξ2drops out of the rate. The generalization tohigherdimensionsis obvious.Moreover,in ndimensions, trace(F)=summationdisplay iλi. (18.36) This trace formula follows from the invariance of the secular polynomial in Eq. (18.29) under a linear transformation, Cξ=xin particular, and it is a result of its determinental formusingtheproducttheoremfor determinants(see Section3.2), viz. det(F−λ·1)=bracketleftbig det(C)bracketrightbig−1det(F−λ·1)bracketleftbig det(C)bracketrightbig =detparenleftbig C−1(F−λ·1)Cparenrightbig =detparenleftbig C−1FC−λ·1parenrightbig =nproductdisplay i=1(λi−λ). (18.37) HeretheproductformcomesaboutbysubstitutingEq. (18.30). Now, trace (F)is thecoef- ficient of (−λ)n−1upon expanding det (F−λ·1)in powers of λ, while it issummationtext iλifrom theproductformproducttext i(λi−λ),whichprovesEq.(18.36).Clearly,accordingtoEqs.(18.35) and(18.36), 18.4 Nonlinear Differential Equations 1103 •itisthesign(andmorepreciselythetrace)ofthecharacteristicexponentsofthederiv- ative matrix at the fixed point that determines whether there is expansion or shrinkage ofareas andvolumesinhigherdimensionsnearacriticalpoint. In summary then, Eq. (18.35) states that dissipation requires ∇·f(y)/negationslash=0, where˙yj=fj, anddoesnotoccurinHamiltoniansystemswhere ∇·f=0. Moreover,intwoormoredimensions,therearethefollowingglobalpossibilities: •Thetrajectorymaydescribeaclosedorbit(cycle). •The trajectory may approach a closed orbit (spiraling inward or outward toward the orbit)ast→∞.Inthis casewehavealimitcycle. The local behavior of a trajectory near a critical point is also more varied in general than in one dimension: At a stable critical point all trajectories may approach the critical point along straight lines or spiral inward (toward the spiral node )o rm a yf o l l o wam o r e complicated path. If all time-reversed trajectories move toward the critical point in spirals ast→−∞, then the critical point is a divergent spiral point, or spiral repellor . When some trajectories approach the critical point while others move away from it, then it is calledasaddlepoint .Whenalltrajectoriesformclosedorbitsaboutthecriticalpoint,itis calledacenter. Bifurcations in Dynamical Systems A bifurcation is a sudden change in dynamics for specific parameter values, such as the birthofanode–repellorpairoffixedpointsortheirdisappearanceuponadjustingacontrol parameter; that is, the motions before and after the bifurcation are topologically different. At a bifurcation point, not only are solutions unstable when one or more parameters are changed slightly, but the character of the bifurcation in phase space or in the parameter manifold may change. Thus we are dealing with fairly sudden events of nonlinear dynam- ics. Rather sudden changes from regular to random behavior of trajectories are character- istic of bifurcations, as is sensitive dependence on initial conditions: Nearby initial condi- tions can lead to very different long-term behavior. If a bifurcation does not change quali- tatively with parameter adjustments, it is called structurally stable . Note that structurally unstable bifurcations are unlikely to occur in reality because noise and other neglected degrees of freedom act as perturbations on the system that effectively eliminate unstable bifurcations from our view. Bifurcations (such as doublings in maps) are important as one among many routes to chaos. Others are sudden changes in trajectories associated with several critical points called global bifurcations . Often they involve changes in basins of attractionand/orotherglobalstructures.Thetheoryofglobalbifurcationsisfairlycompli- catedandis stillinitsinfancyatpresent. Bifurcations that are linked to sudden changes in the qualitative behavior of dynamical systems at a single fixed point are called local bifurcations . More specifically, a change in stability occurs in parameter space where the real part of a characteristic exponent of thefixedpointaltersitssign,thatis,movesfromattractingtorepellingtrajectories,orvice versa. The center–manifold theorem says that at a local bifurcation only those degrees 1104 Chapter 18 Nonlinear Methods and Chaos of freedom matter that are involved with characteristic exponents going to zero: ℜλi=0. Locating the set of these points is the first step in a bifurcation analysis. Another step consists in cataloguing the types of bifurcations in dynamical systems, to which we turn next. The conventional normal forms of dynamical equations represent a start in classifying bifurcations. For systems with one parameter (that is, a one-dimensional center manifold) wewritethegeneralcaseofNDEas follows: ˙x=∞summationdisplay j=0a(0) jxj+c∞summationdisplay j=0a(1) jxj+c2∞summationdisplay j=0a(2) jxj+···, (18.38) wherethesuperscriptonthe a(m)denotesthepoweroftheparameter ctheyareassociated with. One-dimensional iterated nonlinear maps such as the logistic map of Section 18.2 (which occur in Poincaré sections) of nonlinear dynamical systems can be classified simi- larly,viz. xn+1=∞summationdisplay j=0a(0) jxj n+c∞summationdisplay j=0a(1) jxj n+c2∞summationdisplay j=0a(2) jxj n+···. (18.39) Thus,oneofthesimplestNDEswithabifurcationis ˙x=x2−c, (18.40) which corresponds to all a(m) j=0 except for a(1) 0=−1 anda(0) 2=1. Forc>0, there are two fixed points (recall, ˙x=0)x±=±√cwith characteristic exponents 2 x±,s ox−is a node and x+is a repellor. For c<0 there are no fixed points. Therefore, as c→0t h e fixed point pair disappears suddenly; that is, the parameter value c=0 is a repellor-node bifurcation that is structurally unstable. This complex map (with c→−c) generates the fractalJuliaandMandelbrotsets discussedinSection18.3. Apitchforkbifurcationoccursfortheundamped(nondissipativeandspecialcaseofthe Duffing)oscillatorwithacubicanharmonicity ¨x+ax+bx3=0,b>0. (18.41) It has a continuous frequency spectrum and is, among others, a model for a ball bouncing between two walls. When the control parameter a>0, there is only one fixed point, at x=0, a node, while for a<0 there are two more nodes, at x±=±√−a/b.T h u s ,w e havea pitchforkbifurcationof anodeattheorgin intoasaddlepointattheoriginandtwo nodes, at x±/negationslash=0. In terms of a potential formulation, V(x)=ax2/2+bx4/4 is a single wellfora>0 butadoublewell(withamaximumat x=0)fora<0. Whenapairofcomplexconjugatecharacteristicexponents ρ±iκcrossesfromaspiral node (ρ<0) to a repelling spiral ( ρ>0) and periodic motion (limit cycle) emerges, then we call the qualitative change a Hopf bifurcation . They occur in the quasiperiodic route tochaosthatwillbediscussedinthenextsection,onchaos. In a global analysis we piece together the motions near various critical points, such as nodes and bifurcations, to bundles of trajectories that flow more or less together in two dimensions.(Thisgeometricviewisthecurrentmodeofanalyzingsolutionsofdynamical systems.) But this flow is no longer collective in the case of three dimensions, where they diverge from each other in general, because chaotic motion is possible that typically fills theplaneofaPoincarésectionwithpoints. 18.4 Nonlinear Differential Equations 1105 Chaos in Dynamical Systems Our previous summaries of intricate and complicated features of dynamical systems due to nonlinearities in one and two dimensions do not include chaos, although some of them, such as bifurcations, sometimes are precursors to chaos. In three- or more-dimensional NDEs, chaoticmotionmayoccur, often whena constantof the motion(an energyintegral forNDEsdefinedbyaHamiltonian,forexample)restrictsthetrajectoriestoafinitevolume inphasespaceandwhentherearenocriticalpoints.Anothercharacteristicsignalforchaos iswhenforeachtrajectorytherearenearbyones,someofwhichmoveawayfromit,while others approach it with increasing time. The notion of exponential divergence of nearby trajectories is made quantitative by the Lyapunov exponent λ(see Section 18.3 for more details)ofiteratedmapsofPoincarésectionsassociatedwiththedynamicalsystem.Iftwo nearby trajectories are at a distance d0at timet=0 but diverge with a distance d(t)at a latertime t,thend(t)≈d0eλtholds.Thus,byanalyzingtheseriesofpoints,thatis,iterated maps generated on Poincaré sections, one can study routes to chaos of three-dimensional dynamical systems. This is the key method for studying chaos. As one varies the location and orientation of the Poincaré plane, a fixed point on it often is recognized to originate from a limit cycle in the three-dimensional phase space whose structural stability can be checked there. For example, attracting limit cycles show up as nodes in Poincaré sections, repelling limit cycles as repellors of Poincaré maps, and saddle cycles as saddle points of associatedPoincarémaps. Threeormoredimensionsofphasespacearerequiredforchaostooccurbecauseofthe interplayofthenecessaryconditionswejustdiscussed,viz. •boundedtrajectories(areoftenthecaseforHamiltoniansystems), •exponential divergence of nearby trajectories (is guaranteed by positive Lyapunov ex- ponentsof correspondingPoincarémaps), •nointersectionoftrajectories. The last condition is obeyed by deterministic systems in particular, as we discussed in Section 18.1. A surprising feature of chaos, mentioned in Section 18.1, is how prevalent it is and how universal the routes to chaos often are, despite the overwhelming variety of NDEs. An example for spatially complex patterns in classical mechanics is the planar pendu- lum,whoseone-dimensionalequationofmotion Idθ dt=L,dL dt=−lmgsinθ (18.42) isnonlinearinthedynamicvariable θ(t).HereIisthemomentofinertia, listhedistance tothecenterofmass, misthemass,and gisthegravitationalaccelerationconstant.When allparametersinEq.(18.42)areconstantintimeandspace,thenthesolutionsaregivenin termsofellipticintegrals(seeSection5.8)andnochaosexists.However,apendulumunder aperiodicexternalforcecanexhibitchaoticdynamics,for example,for theLagrangian L=m 2˙r2−mg(l−z), (x−x0)2+y2+z2=l2, (18.43) x0=εlcosωt. (18.44) 1106 Chapter 18 Nonlinear Methods and Chaos (SeeMoon(1992)intheAdditionalReadings.) Goodcandidatesfor chaosaremultiplewellpotentialproblems, d2r dt2+∇V(r)=Fparenleftbigg r,dr dt,tparenrightbigg , (18.45) whereFrepresentsdissipativeand/ordrivingforces.Anotherclassicexampleisrigid-body rotation,whosenonlinearthree-dimensionalEulerequationsarefamiliar,viz. d dtI1ω1=(I2−I3)ω2ω3+M1, d dtI2ω2=(I3−I1)ω1ω3+M2, (18.46) d dtI3ω3=(I1−I2)ω1ω2+M3. HeretheIjaretheprincipalmomentsofinertiaand ωistheangularvelocitywithcompo- nentsωjaboutthebody-fixedprincipalaxes.Evenfreerigid-bodyrotationcanbechaotic, foritsnonlinearcouplingsandthree-dimensionalformsatisfyallrequirementsforchaosto occur (see Section 18.1). A rigid-body exampleof chaos in our solar system is the chaotic tumbling of Hyperion, one of Saturn’s moons that is highly nonspherical. It is a world where the Saturn rise and set is so irregular as to be unpredictable. Another is Halley’s comet, whose orbit is perturbed by Jupiter and Saturn. In general, when three or more ce- lestial bodies interact gravitationally, stochastic dynamics are possible. Note, though, that computer simulations over large time intervals are required to ascertain chaotic dynamics in the solar system. For more details on chaos in such conservative Hamiltonian systems werefer toChapter8ofHilborn(1994)intheAdditionalReadings. Exercise 18.4.3 ConstructaPoincarémapfor theDuffingoscillatorinEq. (18.41). Routes to Chaos in Dynamical Systems Letusnowlookatsomeroutestochaos.Theperiod-doublingroutetochaosisexemplified by the logistic map in Section 18.2, and the universal Feigenbaum numbers α,δare its quantitativefeatures,alongwithLyapunovexponents.Itiscommonindynamicalsystems. Itmaybeginwithlimitcycle(periodic)motionthatshowsupasafixedpointinaPoincaré section. The limit cycle may have originated in a bifurcation from a node or some other fixedpoint.Asacontrolparameterchanges,thefixedpointofthePoincarémapsplitsinto two points; that is, the limit cycle has a characteristic exponent going through zero from attracting to repelling, say. The periodic motion now has a period twice as long as before, etc. We refer to Chapter 11 of Barger and Olsson (1995) in the Additional Readings for period-doublingplotsofPoincarésectionsfortheDuffingequation(18.41)withaperiodic externalforce.Anotherexampleforperioddoublingisaforcedoscillatorwithfriction(see HellemaninCvitanovic(1989)intheAdditionalReadings). 18.4 Additional Readings 1107 Thequasiperiodicroutetochaosisalsoquitecommonindynamicalsystems,forexam- ple,startingfrom atime-independentnode,a fixedpoint.If weadjustacontrolparameter, the system undergoes a Hopf bifurcation to the periodic motion corresponding to a limit cycle in phase space. With further change of the control parameter, a second frequency appears. If the frequency ratio is an irrational number, the trajectories are quasiperiodic, eventuallycoveringthesurfaceofatorusinphasespace;thatis,quasiperiodicorbitsnever close or repeat. Further changes of the control parameter may lead to a third frequency or directly to chaotic motion. Bands of chaotic motion can alternate with quasiperiodic mo- tion in parameter space. An example for such a dynamic system is a periodically driven pendulum. A third route to chaos goes via intermittency, where the dynamical system switches between two qualitatively different motions at fixed control parameters. For example, at thebeginning,periodicmotionalternateswithanoccasionalburstofchaoticmotion.With achangeofthecontrolparameter,thechaoticburststypicallylengthenuntil,eventually,no periodicmotionremains.Thechaoticpartsareirregularanddonotresembleeachother,but oneneedstocheckforapositiveLyapunovexponenttodemonstratechaos.Intermittencies of various types are common features of turbulent states in fluid dynamics. The Lorenz coupledNDEsalsoshowintermittency. Exercise 18.4.4 Plot the intermittency region of the logistic map at µ=3.8319. What is the period of thecycles?Whathappensat µ=1+2√ 2? ANS.Thereisatangentbifurcationtoperiod3cycles. AdditionalReadings Amann, H., Ordinary Differential Equations: An Introduction To Nonlinear Analysis . New York: de Gruyter (1990). Baker, G. L., and J. P. Gollub, Chaotic Dynamics: An Introduction , 2nd ed. Cambridge, UK: Cambridge Univer- sity Press (1996). B ar g er ,V .D.,an dM.G.Ol s s o n , Classical Mechanics , 2nd ed. NewYork: McGraw-Hill (1995). Bender, C. M., and S. A. Orszag, Advanced Mathematical Methods For Scientists and Engineers .N e wY o r k : McGraw-Hill(1978), Chapter 4inparticular. Bergé,P., Y.Pomeau, andC.Vidal, Order within Chaos . NewYork: Wiley (1987). Cvitanovic, P.,ed., Universalityin Chaos , 2nd ed.Bristol, UK:AdamHilger(1989). Devaney,R.L., AnIntroductiontoChaoticDynamicalSystems .MenloPark,CA:Benjamin/Cummings;2nded., Perseus (1989). Earnshaw, J.C., and D.Haughey, Lyapunov exponents for pedestrians. Am.J .Ph ys. 61: 401 (1993). Gleick,J., Chaos. New York: Penguin Books (1987). Hilborn, R.C., Chaos and Nonlinear Dynamics .NewYork: Oxford University Press (1994). Hirsch,M.W.,andS.Smale, DifferentialEquations, DynamicalSystems,and Linear Algebra .NewYork: Acad- emic Press (1974). Infeld,E.,andG.Rowlands, NonlinearWaves,SolitonsandChaos .Cambridge,UK:CambridgeUniversityPress (1990). 1108 Chapter 18 Nonlinear Methods and Chaos Jackson, E.A., Perspectivesof Nonlinear Dynamics .Cambridge, UK:Cambridge University Press (1989). Jordan,D.W.,andP.Smith, Nonlinear OrdinaryDifferentialEquations , 2nded.Oxford, UK:Oxford University Press (1987). Lyapunov, A.M., The General Problemof the Stability of Motion . Bristol, PA: Taylor & Francis (1992). Mandelbrot, B. B., The Fractal Geometryof Nature . San Francisco: W. H.Freeman, reprinted (1988). Moon, F.C., Chaotic and Fractal Dynamics .NewYork: Wiley (1992). Peitgen,H.-O.,andP. H.Richter, The Beautyof Fractals . NewYork: Springer (1986). Sachdev,P. L., Nonlinear Differential Equations and their Applications . NewYork: MarcelDekker(1991). Tufillaro,N.B.,T.Abbott,andJ.Reilly, AnExperimentalApproachtoNonlinearDynamicsandChaos .Redwood City, CA: Addison-Wesley (1992). CHAPTER 19 PROBABILITY Probabilitiesariseinmanyproblemsdealingwithrandomeventsorlargenumbersofparti- clesdefiningrandomvariables.Aneventiscalled randomifitispracticallyimpossibleto predict from the initial state. This includes those cases where we have merely incomplete information about initial states and/or the dynamics, as in statistical mechanics, where we may know the energy of the system that corresponds to very many possible microscopic configurations,preventingusfrompredictingindividualoutcomes.Oftentheaverageprop- ertiesofmanysimilareventsarepredictable,asinquantumtheory.Thisiswhyprobability theorycanbeandhasbeendeveloped. Randomvariablesareinvolvedwhendatadependonchance,suchasweatherreportsand stockprices.Thetheoryofprobabilitydescribesmathematicalmodelsofchanceprocesses in terms of probability distributions of random variables that describe how some “random events” are more likely than others. In this sense probability is a measure of our igno- rance, giving quantitative meaning to qualitative statements such as “It will probably rain tomorrow” and “I’m unlikely to draw the heart queen.” Probabilities are of fundamental importance in quantum mechanics and statistical mechanics and are applied in meteorol- ogy,economics,games,andmanyotherareasof dailylife. Toamathematician,probabilitiesarebasedonaxioms,butwewilldiscussherepractical ways of calculating probabilities for random events. Because experiments in the sciences are always subject to errors, theories of errors and their propagation involve probabilities. Instatisticswedealwiththeapplicationsofprobabilitytheorytoexperimentaldata. 19.1 D EFINITIONS ,SIMPLE PROPERTIES All possible mutually exclusive1outcomes of an experiment that is subject to chance represent the events (or points) of the sample space S. For example, each time we toss a coinwe givethe trial a number i=1,2,...andobserve the outcomes xi. Here the sample 1This means that given thatone particular event did occur, theothers could not have occurred. 1109 1110 Chapter 19 Probability consists of two events: heads and tails, and the xirepresent a discrete random variable that takes on one of two values, heads or tails. When two coins are tossed, the sample contains the events two heads, one head and one tail, two tails; the number of heads is a good value to assign to the random variable, so the possible values are 2, 1, and 0. There are four equally probable outcomes, of which one has value 2, two have value 1, and one hasvalue0.Sotheprobabilitiesofthethreevaluesoftherandomvariableare 1 /4f o rt w o heads(value2), 1 /4 for noheads(value0),and 1 /2 for value1.In otherwords, wedefine thetheoreticalprobability Pof aneventdenotedbythepoint xiofthesampleas P(xi)≡numberofoutcomesof event xi totalnumberofallevents. (19.1) An experimentaldefinition applies when the total number of events is not well defined (or isdifficulttoobtain)orequallylikelyoutcomesdonotalways occur.Then P(xi)≡numberof timesevent xioccurs totalnumberoftrials(19.2) is more appropriate. A large, thoroughly mixed pile of black and white sand grains of the samesizeandinequalproportionsisarelevantexample,becauseitisimpracticaltocount themall.Butwecancountthegrainsinasmallsamplevolumethatwepick.Thiswaywe cancheckthatwhiteandblackgrainsturnupwithroughlyequalprobability1 /2,provided we put back each sample and mix the pile again. It is found that the larger the sample volume, the smaller the spread about 1 /2 will be. The more trials we run, the closer the average occurrence of all trial counts will be to 1 /2.We could even pick single grains and check if the probability 1 /4 of picking two black grains in a row equals that of two white grains,etc.Therearelotsofstatisticsquestionswecanpursue.Thus,pilesofcoloredsand provideforinstructiveexperiments. Thefollowingaxiomsareself-evident. •Probabilities satisfy 0 ≤P≤1.Probability 1 means certainty; probability 0 means impossibility. •The entire sample has probability 1 .For example, drawing an arbitrary card has prob- ability 1. •The probabilities for mutually exclusive events add. The probability for getting one headintwocointossesis1 /4+1/4=1/2becauseitis1 /4forheadfirstandthentail, plus 1/4 for tailfirst andthenhead. Example 19.1.1 PROBABILITY FOR AORB Whatistheprobabilityfordrawing2acluborajackfromashuffleddeckofcards?Because there are 52 cards in a deck, each being equally likely, 13 cards for each suit and 4 jacks, there are 13 clubs including the club jack, and 3 other jacks; that is, there are 16 possible cardsoutof 52,givingtheprobability (13+3)/52=16/52=4/13. /squaresolid 2Theseareexamples of non-mutually exclusive events. 19.1 Definitions, Simple Properties 1111 Ifwerepresentthesamplespacebyaset Sofpoints,theneventsaresubsets A,B,...of S,denotedas A⊂S,etc.Twosets A,Bareequalif Aiscontainedin B,A⊂B,andBis containedin A,B⊂A.TheunionA∪Bconsistsofallpoints(events)thatarein AorB or both (see Fig. 19.1). The intersection A∩Bconsists of all points, that are in both A andB.IfAandBhavenocommonpoints,theirintersectionisthe emptyset ,A∩B=∅, whichhas noelements(events). The set of pointsin Athat are notin theintersectionof A andBisdenotedby A−A∩B,definingasubtractionofsets .Ifwetaketheclubsuitin Example 19.1.1 as set Aand the four jacks as set B,then their union comprises all clubs andjacks,andtheirintersectionistheclubjackonly. Each subset Ahas its probability P(A)≥0.In terms of these set theory concepts and notations,theprobabilitylawswejustdiscussedbecome 0≤P(A)≤1. The entire sample space has P(S)=1. The probability of the union A∪Bof mutually exclusiveeventsis thesum P(A∪B)=P(A)+P(B), A∩B=∅. Theadditionrule for probabilitiesofarbitrarysets isgivenbythefollowingtheorem. ADDITION RULE : P(A∪B)=P(A)+P(B)−P(A∩B). (19.3) To prove this, we decompose the union into two mutually exclusive sets A∪B=A∪ (B−B∩A),subtracting the intersection of AandBfromBbefore joining them. Their probabilitiesare P(A),P(B)−P(B∩A),whichweadd.Wecouldalsohavedecomposed A∪B=(A−A∩B)∪B,from which our theorem follows similarly by adding these probabilities, P(A∪B)=[P(A)−P(A∩B)]+P(B). Note that A∩B=B∩A.(See Fig.19.1.) Sometimestherulesanddefinitionsofprobabilitiesthatwehavediscussedsofararenot sufficient,however. FIGURE 19.1Theshadedareagivesthe intersection A∩B, correspondingtothe AandBevents,thedashedlineencloses A∪B, correspondingtothe AorB events. 1112 Chapter 19 Probability Example 19.1.2 CONDITIONAL PROBABILITY Asimpleexampleconsistsofaboxof10identicalredand20identicalbluepens,arranged in random order, from which we remove pens successively, that is, without putting them back.Supposewedrawaredpenfirst,event A.Thatwillhappenwithprobability P(A)= 10/30=1/3 if the pens are thoroughly mixed up. The conditional probability P(B|A) of drawing a blue pen in the next round, event B,however, will depend on the fact that we drew a red pen in the first round. It is given by 20 /29.There are 10·20 possible sample points (red/blue pen events) in two rounds, and the sample has 30 ·29 events, so thecombinedprobabilityis P(A,B)=10 3020 29=10·20 30·29=20 87. /squaresolid In general, the combined probability P(A,B) thatAandBhappen (in this order) is given by the product of the probability that Ahappens, P(A),and the probability that B happensif Adoes,P(B|A): P(A,B)=P(A)P(B|A). (19.4) Inotherwords, the conditionalprobability P(B|A)is givenbytheratio P(B|A)=P(A,B) P(A). (19.5) If the conditionalprobability P(B|A)=P(B)is independentof A,then the events Aand Barecalled independent ,andthecombinedprobability P(A∩B)=P(A)P(B) (19.6) issimplythe productofbothprobabilities . Example 19.1.3 SCHOLASTIC APTITUDE TESTS CollegesanduniversitiesrelyontheverbalandmathematicsSATscores,amongothers,as predictors of a student’s success in passing courses and graduating. A research university is known to admit mostly students with a combined verbal and mathematics score above 1400 points. The graduation rate is 95%; that is, 5% drop out or transfer elsewhere. Of thosewhograduate,97%haveanSATscoreofmorethan1400points,while80%ofthose who drop out have an SAT score below 1400 .Suppose a student has an SAT score below 1400.Whatis his/herprobabilityof graduating? LetAbethecaseshavinganSATtestscorebelow1400 ,Brepresentthoseabove1400 , mutuallyexclusiveeventswith P(A)+P(B)=1,andCbethosestudentswhograduate. That is, we want to know the conditional probabilities P(C|A)andP(C|B).To apply Eq. (19.5) we need P(A)andP(B).There are 3% of students with scores below 1400 amongthosewhograduate(95%)and80%ofthose5%whodonotgraduate,so P(A)=0.03·0.95+4 50.05=0.0685,P(B)=0.97·0.95+0.05 5=0.9315, 19.1 Definitions, Simple Properties 1113 andalso P(C∩A)=0.03·0.95=0.0285 and P(C∩B)=0.97·0.95=0.9215. Here the combined probabilities P(C,A)=P(C∩A),P(C,B)=P(C∩B)asCandA (andCandB) arepartsofthesamesamplespace.Therefore, P(C|A)=P(C∩A) P(A)=0.0285 0.0685∼41.6%, P(C|B)=P(C∩B) P(B)=0.9215 0.9315∼98.9%; that is, a little less than 42% is the probability for a student with a score below 1400 to graduateatthisparticularuniversity. /squaresolid As a corollary to the definition of a conditional probability, Eq. (19.5), we compare P(A|B)=P(A∩B)/P(B) andP(B|A)=P(A∩B)/P(A), whichleadstothefollowing theorem. BAYES THEOREM : P(A|B)=P(A) P(B)P(B|A). (19.7) This canbegeneralizedtothefollowing. THEOREM:If the random events Aiwith probabilities P(Ai)>0are mutually exclusive andtheirunionrepresentstheentiresample S,thenanarbitraryrandomevent B⊂Shas theprobability P(B)=nsummationdisplay i=1P(Ai)P(B|Ai). (19.8) FIGURE 19.2Theshadedarea Bis composedof mutuallyexclusivesubsetsof Bbelongingalsoto A1,A2,A3,wherethe Aiaremutuallyexclusive. 1114 Chapter 19 Probability This decomposition law resembles the expansion of a vector into a basis of unit vectors defining the components of the vector. This relation follows from the obvious decomposi- tionB=uniontext i(B∩Ai),Fig. 19.2, which implies P(B)=summationtext iP(B∩Ai)for the probabil- ities because the components B∩Aiare mutually exclusive. For each i, we know from Eq.(19.5) that P(B∩Ai)=P(Ai)P(B|Ai),whichprovesthetheorem. Counting of Permutations and Combinations Countingparticlesinsamplescanhelpusfindprobabilities,asinstatisticalmechanics. If we have ndifferent molecules, let us ask in how many ways we can arrange them in arow,thatis,permutethem.Thisnumberisdefinedasthenumberoftheir permutations . Thus, by definition, the order matters in permutations . There are nchoices of picking the first molecule, n−1 for the second, etc. Altogether there are n!permutations of n different moleculesor objects. Generalizingthis,supposethereare npeoplebutonly k<nchairstoseatthem.Inhow manywayscanweseat kpeopleinthechairs?Countingasbefore, weget n(n−1)···(n−k+1)=n! (n−k)! forthenumberof permutationsof ndifferentobjects, katatime. Wenowconsiderthenumberof combinations ofobjectswhentheir orderisirrelevant by definition. For example, three letters a,b,ccan be combined, two letters at a time, in 3=3! 2!ways:ab,ac,bc.Ifletterscanberepeated,thenweaddthepairs aa,bb,ccandhave sixcombinations.Thus,a combination ofdifferentparticlesdiffersfromapermutationin thattheir orderdoesnotmatter .Combinationsoccurwithrepetition(themathematician’s wayoftreatingindistinguishableobjects)andwithout,wherenotwosetscontainthesame particles. Thenumberofdifferentcombinationsof nparticles, katatimeandwithoutrepetitions, isgivenbythebinomialcoefficient n(n−1)···(n−k+1) k!=parenleftBign kparenrightBig . If repetitionisallowed,thenthenumberis parenleftbiggn+k−1 kparenrightbigg . In the number n!/(n−k)!of permutations of nparticles, kat a time, we have to divide out the number k!of permutations of the groups of kparticles because their order does not matter in a combination. This proves the first claim. The second one is shown by mathematicalinduction. In statistical mechanics, we ask in how many ways we can put nparticles in kboxes so that there will be ni(distinguishable) particles in the ith box, without regard to order in each box, withsummationtextk i=1ni=n.Counting as before, there are nchoices for selecting the first particle,n−1forpickingthesecond,etc.,butthe n1!permutationswithinthefirstboxare 19.1 Definitions, Simple Properties 1115 discounted,and n2!permutationswithinthesecondboxaredisregarded,etc.Thereforethe numberofcombinationsis n! n1!n2!···nk!,n 1+n2+···+nk=n. In statisticalmechanics,particlesthatobey •Maxwell–Boltzmann (MB) statistics are distinguishable, without restriction on their numberineachstate; •Bose–Einstein (BE) statistics are indistinguishable, with no restriction on the number ofparticlesineachquantumstate; •Fermi–Dirac(FD)statisticsareindistinguishable,withatmostoneparticleper state. Forexample,puttingthreeparticlesinfourboxes,thereare43equallylikelyarrangements fortheMBcase,becauseeachparticlecanbeputintoanyboxinfourways,givingatotalof 43choices.ForBEstatistics,thenumberofcombinationswithrepetitionsisparenleftbig3+4−1 3parenrightbig =parenleftbig6 3parenrightbig for the Bose–Einstein case. For FD statistics, it isparenleftbig3+1 3parenrightbig =parenleftbig43parenrightbig . More generally, for MB statistics the number of distinct arrangements of nparticles among kstates (boxes) is kn, forBE statisticsitisparenleftbign+k−1 nparenrightbig , andfor FDstatisticsitisparenleftbigk nparenrightbig . Exercises 19.1.1 A card is drawn from a shuffled deck. (a) What is the probability that it is black, (b) a rednine,(c) oraqueenof spades? 19.1.2 Find the probability of drawing two kings from a shuffled deck of cards (a) if the first cardisputbackbeforethesecondisdrawn,and(b)ifthefirstcardisnotputbackafter beingdrawn. 19.1.3 When two fair dice are thrown, what is the probability of (a) observing a number less than 4 or(b) anumbergreaterthanor equalto 4 butlessthan 6? 19.1.4 Rollingthreefair dice,whatis theprobabilityofobtainingsixpoints? 19.1.5 Determinetheprobability P(A∩B∩C)interms of P(A),P(B),P(C), etc. 19.1.6 Determine directly or by mathematical induction the probability of a distribution of N (Maxwell–Boltzmann) particles in kboxes with N1in box 1, N2in box 2,...,N kin thekthboxforanynumbers Nj≥1withN1+N2+···+Nk=N,k<N.Repeatthis for Fermi–DiracandBose–Einsteinparticles. 19.1.7 Showthat P(A∪B∪C)=P(A)+P(B)+P(C)−P(A∩B)−P(A∩C) −P(B∩C)+P(A∩B∩C). 19.1.8 Determinetheprobabilitythatapositiveinteger n≤100isdivisiblebyaprimenumber p≤100.Verifyyourresultfor p=3,5,7. 19.1.9 PuttwoparticlesobeyingMaxwell–Boltzmann(Fermi–Dirac,orBose–Einstein)statis- ticsinthreeboxes.Howmanywaysarethereineachcase? 1116 Chapter 19 Probability 19.2 R ANDOM VARIABLES Each time we toss a die, we give the trial a number i=1,2,...and observe the point xi=1,or 2, 3, 4, 5, 6 with probability 1 /6.Ifidenotes the trial number, then xiis a discreterandomvariablethattakesthediscretevaluesfrom1to6withadefiniteprobability P(xi)=1/6. Example 19.2.1 DISCRETE RANDOM VARIABLE If we toss two dice and record the sum of the points shown in each trial, then this sum is also a discrete random variable, which takes on the value 2 when both dice show 1 with probability (1/6)2; the value 3 when one die has 1 and the other 2 ,hence with proba- bility(1/6)2+(1/6)2=1/18; the value 4 when both dice have 2 or one has 1 and the other3,sowithprobability (1/6)2+(1/6)2+(1/6)2=1/12;thevalue5withprobability 4(1/6)2=1/9; the value 6 with probability 5 /36; the value 7 with the maximum proba- bility, 6(1/6)2=1/6; up to the value 12 when both dice show 6 points with probability (1/6)2.Thisprobabilitydistributionissymmetricabout7 .Thissymmetryisobviousfrom Fig. 19.3 and becomes visible algebraically when we write the rising and falling linear partsas P(x)=x−1 36=6−(7−x) 36,x=2,3,...,7, P(x)=13−x 36=6+(7−x) 36,x=7,8,...,12./squaresolid In summary,then, •The different values xithat a random variable Xassumes denote and distinguish the eventsinthesamplespaceofanexperiment;eacheventoccursbychancewithaprob- FIGURE 19.3Probabilitydistribution P(x)of thesumof pointswhentwo diceare tossed. 19.2 Random Variables 1117 abilityP(X=xi)=pi≥0 that is a function of the random variable X.A random variableX(ei)=xiis definedonthesamplespace,thatis, fortheevents ei∈S. •Wedefinetheprobabilitydensity f(x)ofacontinuousrandomvariable Xas P(x≤X≤x+dx)=f(x)dx; (19.9) that is,f(x)dxis the probability that Xlies in the interval x≤X≤x+dx.For f(x)to be a probability density, it has to satisfy f(x)≥0 andintegraltext f(x)dx=1.The generalization to probability distributions depending on several random variables is straightforward.Quantumphysicsaboundsinexamples. Example 19.2.2 CONTINUOUS RANDOM VARIABLE :HYDROGEN ATOM Quantum mechanics gives the probability |ψ|2d3rof finding a 1 selectron in a hydrogen atominvolume3d3r,whereψ=Ne−r/ais thewavefunctionthatis normalizedto 1=integraldisplay |ψ|2dV=4πN2integraldisplay∞ 0e−2r/ar2dr=πa3N2,dV=r2drdcosθdϕ being the volume element and athe Bohr radius. The radial integral is found by repeated integrationbypartsor byrescalingittothegammafunction integraldisplay∞ 0e−2r/ar2dr=parenleftbigga 2parenrightbigg3integraldisplay∞ 0e−xx2dx=a3 8Ŵ(3)=a3 4. Here all points in space constitute the sample and represent three random variables, but theprobabilitydensity |ψ|2in thiscase dependsonlyon theradial variablebecauseof the sphericalsymmetryof the 1 sstate. A measure for the size of the Hatom is givenby theaverageradial distanceof the elec- tron from the proton at the center, which in quantum mechanics is called the expectation value: /angbracketleft1s|r|1s/angbracketright=integraldisplay r|ψ|2dV=4πN2integraldisplay∞ 0re−2r/ar2dr=3 2a. Weshalldefinethisconceptforarbitraryprobabilitydistributionsshortly. /squaresolid •A random variable that takes only discrete values x1,x2,...,xnwith probabilities p1,p2,...,pn,respectively, is called a discrete random variable, sosummationtext ipi=1. If an “experiment”ortrialis performed,someoutcomemustoccur,withunitprobability. •If the values comprise a continuous range of values a≤x≤b,then we deal with a continuous random variable, whose probability distribution may or may not be a continuousfunctionaswell. 3Notethat|ψ|24πr2drgives the probability for the electronto befound between randr+dr, at anyangle. 1118 Chapter 19 Probability When we measure a quantity xntimes, obtaining the values xj,we define the average value ¯x=1 nnsummationdisplay j=1xj (19.10) of the trials, also called the meanorexpectation value , where this formula assumes that everyobservedvalue xiisequallylikelyandoccurswithprobability1 /n.Thisconnection isthekeylinkofexperimentaldatawithprobabilitytheory.Thisobservationandpractical experiencesuggestdefiningthe meanvaluefor adiscreterandomvariable Xas /angbracketleftX/angbracketright≡summationdisplay ixipi (19.11) andthatfor a continuousrandomvariable characterizedbyprobabilitydensity f(x)as /angbracketleftX/angbracketright=integraldisplay xf(x)dx. (19.12) Thesearelinearaverages.Othernotationsintheliteratureare ¯XandE(X). The use of the arithmetic mean ¯xofnmeasurements as the average value is suggested bysimplicityandplainexperience,assumingequalprobabilityforeach xiagain.Butwhy dowenotconsiderthegeometricmean xg=(x1·x2·····xn)1/n(19.13) ortheharmonicmean xhdeterminedbytherelation 1 xh=1 nparenleftbigg1 x1+1 x1+···+1 xnparenrightbigg (19.14) or that value˜xthat minimizes the sum of absolute deviations |xi−˜x|? Here the xiare taken to increase monotonically. When we plot O(x)=summationtext2n+1 i=1|xi−x|, as in Fig. 19.4a, for an odd number of points, we realize that it has a minimum at its central value i=n, FIGURE 19.4(a)summationtext3 i=1|xi−x|for anodd numberofpoints;(b)summationtext4 i=1|xi−x|foraneven numberof points. 19.2 Random Variables 1119 while for an even number of points E(x)=summationtext2n i=1|xi−x|is flat in its central region, as shown in Fig. 19.4b. These properties make these functions unacceptable for determining averagevalues.Instead, whenweminimizethesumofquadraticdeviations, nsummationdisplay i=1(x−xi)2=minimum , (19.15) settingthederivativeequaltozeroyields 2summationtext i(x−xi)=0,or x=1 nsummationdisplay ixi≡¯x, thatis,thearithmeticmean.Ithasanotherimportantproperty:Ifwedenoteby vi=xi−¯x the deviations, thensummationtext ivi=0,that is, the sum of positive deviations equals the sum of negative deviations. This principle of minimizing the quadratic sum of deviations, called themethodof leastsquares ,isduetoC. F.Gauss, amongothers. How close a fit of the mean value to a set of data points is depends on the spread of the individual measurements from this mean. Again, we reject the average sum of deviationssummationtextn i=1|xi−¯x|/nas a measure of the spread because it selects the central measurement as the best value for no good reason. A more appropriate definition of the spread is the averageof thedeviationsfrom themean,squared,or standarddeviation σ=radicaltpradicalvertexradicalvertexradicalbt1 nnsummationdisplay i=1(xi−¯x)2, wherethesquarerootismotivatedbydimensionalanalysis. Example 19.2.3 STANDARD DEVIATION OF MEASUREMENTS From the measurements x1=7,x2=9,x3=10,x4=11,x5=13 we extract ¯x=10 for the mean value and σ=√(9+1+1+9)/4=2.2361 for the standard deviation, or spread,usingtheexperimentalformula(19.2)becausetheprobabilitiesarenotknown. /squaresolid There is yet another interpretation of the standard variation, in terms of the sum of squaresofmeasurementdifferences summationdisplay i<k(xi−xk)2=1 2nsummationdisplay i=1nsummationdisplay k=1parenleftbig x2 i+x2 k−2xixkparenrightbig =1 2parenleftbig 2n2angbracketleftbig x2angbracketrightbig −2n2/angbracketleftx/angbracketright2parenrightbig =n2σ2, (19.16) becausebymultiplyingoutthesquareinthedefinitionof σ2weobtain σ2=1 nsummationdisplay iparenleftbig xi−/angbracketleftx/angbracketrightparenrightbig2=1 nsummationdisplay ix2 i−2/angbracketleftx/angbracketright nsummationdisplay ixi+/angbracketleftx/angbracketright2 =1 nsummationdisplay ix2 i−/angbracketleftx/angbracketright2=angbracketleftbig x2angbracketrightbig −/angbracketleftx/angbracketright2. (19.17) Thisformulaisoftenusedandwidelyappliedfor σ2. 1120 Chapter 19 Probability Now we are ready to generalize the spread in a set of nmeasurements with equal prob- ability 1/nto thevariance of an arbitrary probability distribution. For a discrete random variableXwithprobabilities piatX=xiwedefinethe variance σ2=summationdisplay jparenleftbig xj−/angbracketleftX/angbracketrightparenrightbig2pj, (19.18) andsimilarlyfor acontinuousprobabilitydistribution σ2=integraldisplay∞ −∞parenleftbig x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx. (19.19) Thesedefinitionsimplythefollowing. THEOREM:Ifarandomvariable Y=aX+bislinearlyrelatedto X,thenwecanimme- diately derive the mean value /angbracketleftY/angbracketright=a/angbracketleftX/angbracketright+band variance σ2(Y)=a2σ2(X)from these definitions. Weprovethistheoremonlyforacontinuousdistributionandleavethecaseofadiscrete random variable as an exercise for the reader. For the infinitesimal probability we know thatf(x)dx=g(y)dywithy=ax+b,becausethelineartransformationhastopreserve probability,so /angbracketleftY/angbracketright=integraldisplay∞ −∞yg(y)dy=integraldisplay∞ −∞(ax+b)f(x)dx=a/angbracketleftX/angbracketright+b, sinceintegraltext f(x)dx=1.Forthevariancewesimilarlyobtain σ2(Y)=integraldisplay∞ −∞parenleftbig y−/angbracketleftY/angbracketrightparenrightbig2g(y)dy=integraldisplay∞ −∞parenleftbig ax+b−a/angbracketleftX/angbracketright−bparenrightbig2f(x)dx =a2σ2(X) aftersubstitutingourresultfor themeanvalue /angbracketleftY/angbracketright. Finallyweprovethegeneral Chebychevinequality Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥kσparenrightbig ≤1 k2, (19.20) which demonstrates why the standard deviation serves as a measure of the spread of an arbitrary probability distribution from its mean value /angbracketleftX/angbracketrightand shows why experimental or other data are often characterized according to their spread in numbers of standard devia- tions.Wefirst showthesimplerinequality P(Y≥K)≤/angbracketleftY/angbracketright K for a continuous random variable Ywhose values y≥0.(The proof for a discrete random variablefollowsalongsimilarlines.) Thisinequalityfollowsfrom /angbracketleftY/angbracketright=integraldisplay∞ 0yf(y)dy=integraldisplayK 0yf(y)dy+integraldisplay∞ Kyf(y)dy ≥integraldisplay∞ Kyf(y)dy≥Kintegraldisplay∞ Kf(y)dy=KP(Y≥K). 19.2 Random Variables 1121 Nextweapplythesamemethodtothepositivevarianceintegral σ2=integraldisplayparenleftbig x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx≥integraldisplay |x−/angbracketleftX/angbracketright|≥kσparenleftbig x−/angbracketleftX/angbracketrightparenrightbig2f(x)dx ≥k2σ2integraldisplay |x−/angbracketleftX/angbracketright|≥kσf(x)dx=k2σ2Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥kσparenrightbig , decreasing the right-hand side first by omitting the part of the positive integral with |x−/angbracketleftX/angbracketright|≤kσand then again by replacing (x−/angbracketleftX/angbracketright)2in the remaining integral by its lowest limit, k2σ2.This proves the Chebychev inequality. For k=3 we have the conven- tionalthree-standard-deviationestimate Pparenleftbigvextendsinglevextendsinglex−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥3σparenrightbig ≤1 9. (19.21) It is straightforward to generalize the mean value to higher moments of probability dis- tributionsrelativetothemeanvalue /angbracketleftX/angbracketright: angbracketleftbigparenleftbig X−/angbracketleftX/angbracketrightparenrightbigkangbracketrightbig =summationdisplay jparenleftbig xj−/angbracketleftX/angbracketrightparenrightbigkpj,discretedistribution , (19.22) angbracketleftbigparenleftbig X−/angbracketleftX/angbracketrightparenrightbigkangbracketrightbig =integraldisplay∞ −∞parenleftbig x−/angbracketleftX/angbracketrightparenrightbigkf(x)dx, continuousdistribution . Themoment-generatingfunction angbracketleftbig etXangbracketrightbig =integraldisplay etxf(x)dx=1+t/angbracketleftX/angbracketright+t2 2!angbracketleftbig X2angbracketrightbig +··· (19.23) is a weighted sum of the moments of the continuous random variable Xupon substituting the Taylor expansion of the exponential functions. So /angbracketleftX/angbracketright=d/angbracketleftetX/angbracketright dtvextendsinglevextendsingle t=0.Notice that the moments here are not relative to the expectation value; they are called central moments . Thenthcentralmoment, /angbracketleftXn/angbracketright=dn/angbracketleftetX/angbracketright dtnvextendsinglevextendsingle t=0,isgivenbythe nthderivativeofthemoment- generating function at t=0.By a change of the parameter t→itthe moment-generating function is related to the characteristic function /angbracketlefteitX/angbracketright,which is the often-used Fourier transformoftheprobabilitydensity f(x). Moreover, mean values, moments, and variance can be defined similarly for probability distributions that depend on several random variables. For simplicity, let us restrict our attentiontotwocontinuousrandomvariables X, Yandlistthecorrespondingquantities: /angbracketleftX/angbracketright=integraldisplay∞ −∞integraldisplay∞ −∞xf(x,y)dx dy, /angbracketleftY/angbracketright=integraldisplay∞ −∞integraldisplay∞ −∞yf(x,y)dx dy, (19.24) σ2(X)=integraldisplay∞ −∞integraldisplay∞ −∞parenleftbig x−/angbracketleftX/angbracketrightparenrightbig2f(x,y)dxdy, σ2(Y)=integraldisplay∞ −∞integraldisplay∞ −∞parenleftbig y−/angbracketleftY/angbracketrightparenrightbig2f(x,y)dxdy. (19.25) 1122 Chapter 19 Probability Two random variables are said to be independent if the probability density f(x,y) factorizes into a product f(x)g(y) of probability distributions of one random variable each. Thecovariance,definedas cov(X,Y)=angbracketleftbigparenleftbig X−/angbracketleftX/angbracketrightparenrightbigparenleftbig Y−/angbracketleftY/angbracketrightparenrightbigangbracketrightbig , (19.26) is a measureof how muchthe randomvariables X, Yare correlated (or related): It is zero forindependentrandomvariablesbecause cov(X,Y)=integraldisplayparenleftbig x−/angbracketleftX/angbracketrightparenrightbigparenleftbig y−/angbracketleftY/angbracketrightparenrightbig f(x,y)dxdy =integraldisplayparenleftbig x−/angbracketleftX/angbracketrightparenrightbig f(x)dxintegraldisplayparenleftbig y−/angbracketleftY/angbracketrightparenrightbig g(y)dy=parenleftbig /angbracketleftX/angbracketright−/angbracketleftX/angbracketrightparenrightbigparenleftbig /angbracketleftY/angbracketright−/angbracketleftY/angbracketrightparenrightbig =0. Thenormalizedcovariancecov(X,Y) σ(X)σ(Y),whichhasvaluesbetween −1and+1,isoftencalled correlation . In ordertodemonstratethatthecorrelationisboundedby −1≤cov(X,Y) σ(X)σ(Y)≤1 weanalyzethepositivemeanvalue Q=angbracketleftbigbracketleftbig uparenleftbig X−/angbracketleftX/angbracketrightparenrightbig +vparenleftbig Y−/angbracketleftY/angbracketrightparenrightbigbracketrightbig2angbracketrightbig =u2angbracketleftbigbracketleftbig X−/angbracketleftX/angbracketrightbracketrightbig2angbracketrightbig +2uvangbracketleftbigbracketleftbig X−/angbracketleftX/angbracketrightbracketrightbigbracketleftbig Y−/angbracketleftY/angbracketrightbracketrightbigangbracketrightbig +v2angbracketleftbigbracketleftbig Y−/angbracketleftY/angbracketrightbracketrightbig2angbracketrightbig =u2σ(X)2+2uvcov(X,Y)+v2σ(Y)2≥0, (19.27) whereu, vare numbers, not functions. For this quadratic form to be nonnegative, its dis- criminantmustobeycov (X,Y)2−σ(X)2σ(Y)2≤0,whichprovesthedesiredinequality. Theusefulnessofthecorrelationasaquantitativemeasureisemphasizedbythefollow- ing. THEOREM:P(Y=aX+b)=1isvalidif,andonlyif,thecorrelationisequalto ±1. Thistheoremstatesthata ±100%correlationbetween X, Yimpliesnotonlysomefunc- tionalrelationbetweenbothrandomvariablesbutalsoa linearrelation betweenthem.We denoteby/angbracketleftB|A/angbracketrighttheexpectationvalueoftheconditionalprobabilitydistribution P(B|A). To prove this strong correlation property, we apply Bayes’ decomposition law (Eq. (19.8)) to the mean value and variance of the random variable Y, assuming first that P(Y=aX+b)=1,soP(Y/negationslash=aX+b)=0.Thisyields /angbracketleftY/angbracketright=P(Y=aX+b)/angbracketleftY|Y=aX+b/angbracketright +P(Y/negationslash=aX+b)/angbracketleftY|Y/negationslash=aX+b/angbracketright =/angbracketleftaX+b/angbracketright=a/angbracketleftX/angbracketright+b, σ(Y)2=P(Y=aX+b)angbracketleftbigbracketleftbig Y−/angbracketleftY/angbracketrightbracketrightbig2vextendsinglevextendsingleY=aX+bangbracketrightbig +P(Y/negationslash=aX+b)angbracketleftbigbracketleftbig Y−/angbracketleftY/angbracketrightbracketrightbig2vextendsinglevextendsingleY/negationslash=aX+bangbracketrightbig =angbracketleftbigbracketleftbig aX+b−/angbracketleftY/angbracketrightbracketrightbig2angbracketrightbig =angbracketleftbig a2bracketleftbig X−/angbracketleftX/angbracketrightbracketrightbig2angbracketrightbig =a2σ(X)2, 19.2 Random Variables 1123 substituting/angbracketleftY/angbracketright=a/angbracketleftX/angbracketright+b.Similarly we obtain cov (X,Y)=a2σ(X)2.These results showthatthecorrelationis ±1. Conversely, we start from cov (X,Y)2=σ(X)2σ(Y)2.Hence the quadratic form in Eq.(19.27) mustbezerofor (practically)all xfor some (u0,v0)/negationslash=(0,0): angbracketleftbigbracketleftbig u0parenleftbig X−/angbracketleftX/angbracketrightparenrightbig +v0parenleftbig Y−/angbracketleftY/angbracketrightparenrightbigbracketrightbig2angbracketrightbig =0. Because the argument of this mean value is positive definite, this relationship is satisfied only ifP(u0(X−/angbracketleftX/angbracketright)+v0(Y−/angbracketleftY/angbracketright)=0)=1,which means that YandXare linearly related. Whenweintegrateoutonerandomvariable,weareleftwiththeprobabilitydistribution oftheotherrandomvariable, F(x)=integraldisplay f(x,y)dy, orG(y)=integraldisplay f(x,y)dx, (19.28) andanalogouslyfordiscreteprobabilitydistributions.Whenoneormorerandomvariables are integrated out, the remaining probability distribution is called marginal, motivated by the geometric aspects of projection. It is straightforward to show that these marginal distributionssatisfyalltherequirementsofproperlynormalizedprobabilitydistributions. If we are interested in the distribution of the randomvariable Xfor a definitevalue y= y0oftheotherrandomvariable,thenwedealwitha conditionalprobabilitydistribution P(X=x|Y=y0).Thecorrespondingcontinuousprobabilitydensityis f(x,y0). Example 19.2.4 REPEATED DRAWS OF CARDS When we draw cards repeatedly, we shuffle the deck often because we want to make sure these events stay independent. So we draw the first card at random from a bridge deck containing 52 cardsandthenput itbackata randomplace.Nowwerepeattheprocess for asecondcard.Thenthedeckis reshuffled,etc.Wenowdefinetherandomvariables •X=number ofso-calledhonors,thatis,10s, jacks,queens,kings,oraces; •Y=number of 2sor 3s. In a single draw the probability of a 10 to an ace is a=5·4/52=5/13,andb= 2·4/52=2/13fortwoorthreetobedrawnand c=(13−5−2)/13=6/13foranything else,with a+b+c=1. In two drawings XandYcan bex=0=y,when no 10 to ace show up or 2 or 3 .This case has probability c2.In general X,Yhave the values 0, 1, and 2 ,so 0≤x+y≤2 because we will have drawn two cards. The probability function of ( X=x,Y=y)i s givenbytheproductoftheprobabilitiesof thethreepossibilities ax,by,c2−x−ytimesthe number of distributions (or permutations) of two cards over the three cases with probabil- itiesa,b,c, which is 2!/[x!y!(2−x−y)!].This number is the coefficient of the power axbyc2−x−yinthegeneralizedbinomialexpansionofallpossibilitiesintwodrawingswith 1124 Chapter 19 Probability probability 1: 1=(a+b+c)2=summationdisplay 0≤x+y≤22! x!y!(2−x−y)!axbyc2−x−y =a2+b2+c2+2(ab+ac+bc). (19.29) Hencetheprobabilitydistributionofourdiscreterandomvariablesisgivenby f(X=x,Y=y)=2! x!y!(2−x−y)!parenleftbigg5 13parenrightbiggxparenleftbigg2 13parenrightbiggyparenleftbigg6 13parenrightbigg2−x−y , x,y=0,1,2;0≤x+y≤2, (19.30) ormoreexplicitlyas f(0,0)=parenleftbigg6 13parenrightbigg2 ,f(1,0)=2·5 13·6 13=60 132, f(2,0)=parenleftbigg5 13parenrightbigg2 ,f(0,1)=22 13·6 13=24 132, f(0,2)=parenleftbigg2 13parenrightbigg2 ,f(1,1)=25 13·2 13=20 132. The probability distribution is properly normalized according to Eq. (19.29). Its expecta- tionvaluesaregivenby /angbracketleftX/angbracketright=summationdisplay 0≤x+y≤2xf(x,y)=f(1,0)+f(1,1)+2f(2,0) =60 132+20 132+2parenleftbigg5 13parenrightbigg2 =130 132=10 13=2a, and /angbracketleftY/angbracketright=summationdisplay 0≤x+y≤2yf(x,y)=f(0,1)+f(1,1)+2f(0,2) =24 132+20 132+2parenleftbigg2 13parenrightbigg2 =52 132=4 13=2b, asexpectedbecausewearedrawingacardtwotimes.Thevariancesare σ2(X)=summationdisplay 0≤x+y≤2parenleftbigg x−10 13parenrightbigg2 f(x,y) =parenleftbigg10 13parenrightbigg2bracketleftbig f(0,0)+f(0,1)+f(0,2)bracketrightbig +parenleftbigg3 13parenrightbigg2bracketleftbig f(1,0)+f(1,1)bracketrightbig +parenleftbigg16 13parenrightbigg2 f(2,0) =102·64+32·80+162·52 134=42·5·169 134=80 132, 19.2 Random Variables 1125 σ2(Y)=summationdisplay 0≤x+y≤2parenleftbigg y−4 13parenrightbigg2 f(x,y) =parenleftbigg4 13parenrightbigg2bracketleftbig f(0,0)+f(1,0)+f(2,0)bracketrightbig +parenleftbigg9 13parenrightbigg2bracketleftbig f(0,1)+f(1,1)bracketrightbig +parenleftbigg22 13parenrightbigg2 f(0,2) =42·112+92·44+222·22 134=11·4·169 134=44 132. It is reasonable that σ2(Y)<σ2(X),becauseYtakes only two values, 2 and 3, while X variesoverthefivehonors. Thecovarianceis givenby cov(X,Y)=summationdisplay 0≤x+y≤2parenleftbigg x−10 13parenrightbiggparenleftbigg y−4 13parenrightbigg f(x,y) =10·4 132·62 132−10·9 132·24 132−10·22 132·4 132−3·4 132·60 132 +3·9 132·20 132−16·4 132·52 132=−20·169 134=−20 132. Thereforethecorrelationoftherandomvariables X,Yisgivenby cov(X,Y) σ(X)σ(Y)=−20 8√ 5·11=−1 2radicalbigg 5 11=−0.3371, which means that there is a small (negative) correlation between these random variables, becauseifacardis anhonoritcannotb ea2ora3,andvicev er s a. Finally,letus determinethemarginaldistribution: F(X=x)=2summationdisplay y=0f(x,y), (19.31) orexplicitly F(0)=f(0,0)+f(0,1)+f(0,2)=parenleftbigg6 13parenrightbigg2 +24 132+parenleftbigg2 13parenrightbigg2 =parenleftbigg8 13parenrightbigg2 , F(1)=f(1,0)+f(1,1)=60 132+20 132=80 132, F(2)=f(2,0)=parenleftbigg5 13parenrightbigg2 , whichisproperlynormalizedbecause F(0)+F(1)+F(2)=64+80+25 132=169 132=1. 1126 Chapter 19 Probability Its meanvalueis givenby /angbracketleftX/angbracketrightF=2summationdisplay x=0xF(x)=F(1)+2F(2)=80+2·25 132=130 132=10 13=/angbracketleftX/angbracketright, andits variance σ2 F=2summationdisplay x=0parenleftbigg x−10 13parenrightbigg2 F(x)=parenleftbigg10 13parenrightbigg2 ·parenleftbigg8 13parenrightbigg2 +parenleftbigg3 13parenrightbigg280 132+parenleftbigg16 13parenrightbigg2 ·parenleftbigg5 13parenrightbigg2 =80·169 134=80 132=σ2(X). Fromthedefinitionsitfollowsthattheseresults holdgenerally. /squaresolid Finally we address the transformation of two random variables X,YintoU(X,Y), V(X,Y). We treatthecontinuouscase,leavingthediscretecase,as anexercise.If u=u(x,y), v =v(x,y);x=x(u,v), y =y(u,v) (19.32) describe the transformation and its inverse, then the probability stays invariant and the integral of the density transforms according to the rules of Jacobians of Chapter 2, so the transformedprobabilitydensitybecomes g(u,v)=fparenleftbig x(u,v),y(u,v)parenrightbig |J|, (19.33) withtheJacobian J=∂(x,y) ∂(u,v)=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle∂x ∂u∂x ∂v ∂y ∂u∂y ∂vvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle. (19.34) Example 19.2.5 SUM,PRODUCT ,AND RATIO OF RANDOM VARIABLES Let us consider three examples. (1) The sum Z=X+Y,where the transformation may betakentobe x=x, z=x+y, J=vextendsinglevextendsinglevextendsinglevextendsingle11 01vextendsinglevextendsinglevextendsinglevextendsingle, using ∂x ∂x=1,∂(z−y) ∂z=1,∂y ∂x=0,∂(z−x) ∂z=1, sotheprobabilityis givenby F(Z)=integraldisplayZ −∞integraldisplay∞ −∞f(x,z−x)dxdz. (19.35) If therandomvariables X,Yare independentwithdensities f1,f2,then F(Z)=integraldisplayZ −∞integraldisplay∞ −∞f1(x)f2(z−x)dxdz. (19.36) 19.2 Random Variables 1127 (2) Theproduct Z=XY,takingX,Zas thenewvariables,leadstotheJacobian J=vextendsinglevextendsinglevextendsinglevextendsinglevextendsingle11 y 01 xvextendsinglevextendsinglevextendsinglevextendsinglevextendsingle=1 x, using ∂x ∂x=1,∂(z y) ∂z=1 y,∂y ∂x=0,∂(z x) ∂z=1 x, sotheprobabilityis givenby F(Z)=integraldisplayZ −∞integraldisplay∞ −∞fparenleftbigg x,z xparenrightbiggdx |x|dz. (19.37) If therandomvariables X,Yare independentwithdensities f1,f2,then F(Z)=integraldisplayZ −∞integraldisplay∞ −∞f1(x)f2parenleftbiggz xparenrightbiggdx |x|dz. (19.38) (3) Theratio Z=X Y, takingY,Zasthenewvariables,hastheJacobian J=vextendsinglevextendsinglevextendsinglevextendsinglezy 10vextendsinglevextendsinglevextendsinglevextendsingle=−y, using ∂(yz) ∂y=z,∂(yz) ∂z=y,∂y ∂y=1,∂y ∂z=0, sotheprobabilityis givenby F(Z)=integraldisplayZ −∞integraldisplay∞ −∞f(yz,y)|y|dydz. (19.39) If therandomvariables X,Yare independentwithdensities f1,f2,then F(Z)=integraldisplayZ −∞integraldisplay∞ −∞f1(yz)f2(y)|y|dydz. (19.40) /squaresolid Exercises 19.2.1 Show that adding a constant cto a random variable Xchanges the expectation value /angbracketleftX/angbracketrightby that same constant but not the variance. Show also that multiplying a random variablebyaconstantmultipliesboththemeanandvariancebythatconstant.Showthat therandomvariable X−/angbracketleftX/angbracketrighthasmeanvaluezero. 19.2.2 If/angbracketleftX/angbracketright,/angbracketleftY/angbracketrightare the average values of two independent random variables X,Y,what is theexpectationvalueoftheproduct X·Y? 1128 Chapter 19 Probability 19.2.3 A velocity vj=xj/tjis measured by recording the distances xjat the corresponding timestj.Showthat¯x/¯tisagoodapproximationfortheaveragevelocity v,providedall the errors|xj−¯x|≪|¯x|and|tj−¯t|≪|¯t|are small. 19.2.4 Define the random variable Yin Example 19.2.4 as the number of 4s, 5s, 6s, 7s, 8s, or 9s.Thendeterminethecorrelationof the XandYrandomvariables. 19.2.5 IfXandYare two independent random variables with different probability densities andthefunction f(x,y)hasderivativesofanyorder,express /angbracketleftf(X,Y)/angbracketrightintermsof/angbracketleftX/angbracketright and/angbracketleftY/angbracketright.Developsimilarlythecovarianceandcorrelation. 19.2.6 Letf(x,y)be the joint probability density of two random variables X,Y.Find the varianceσ2(aX+bY),wherea,bare constants. What happens when X,Yare inde- pendent? 19.2.7 The probability that a particle of an ideal gas travels a small distance dxbetween col- lisions is∼e−x/fdx,wherefis the constant mean free path. Verify that fis the aver- age distance between collisions, and determine the probability of a free path of length l≥3f. 19.2.8 Determinethe probabilitydensityfor a particleinsimple harmonicmotionin theinter- val−A≤x≤A. Hint.The probability that the particle is between xandx+dxis proportional to the timeittakes totravelacrosstheinterval. 19.3 B INOMIAL DISTRIBUTION Example 19.3.1 REPEATED TOSSES OF DICE Whatistheprobabilityofthree 6sinfourtosses,alltrialsbeingindependent?Gettingone 6 in a single toss of a fair die has probability a=1/6,and anything else has probability b=5/6 witha+b=1.Let the random variable X=xbe the number of 6s. In four tosses, 0≤x≤4.The probability distribution f(X)is given by the product of the two possibilities, axandb4−x, times the number of combinations of four tosses over the two cases with probabilities a,b.This number is the coefficient of the power axb4−xin the binomialexpansionofallpossibilitiesinfour tosseswithprobability 1: 1=(a+b)4=4summationdisplay x=04! x!(4−x)!axb4−x =a4+b4+4a3b+4ab3+6a2b2. (19.41) Hencetheprobabilitydistributionofourdiscreterandomvariableisgivenby f(X=x)=4! x!(4−x)!axb4−x,0≤x≤4, ormoreexplicitly f(0)=b4,f(1)=4ab3,f(2)=6a2b2,f(3)=4a3b, f( 4)=a4. 19.3 Binomial Distribution 1129 The probability distribution is properly normalized according to Eq. (19.41). The proba- bilityofthree 6sinfourtosses is 4a3b=45/6 63=5 4·34, fairlysmall. /squaresolid This case dealt with repeated independent trials, each with two possible outcomes of constant probability pfor a hit and q=1−pfor a miss, and it is typical of many ap- plications, such as defective products, hits or misses of a target, and decays of radioactive atoms. The generalization to X=xsuccesses in ntrials is given by the binomial proba- bilitydistribution f(X=x)=n! x!(n−x)!pxqn−x=parenleftbiggn xparenrightbigg pxqn−x, (19.42) usingthebinomialcoefficients(seeChapter5).Thisdistributionisnormalizedtotheprob- ability 1 ofallpossibilitiesin ntrials,as canbeseenfromthebinomialexpansion 1=(p+q)n=pn+npn−1q+···+npqn−1+qn. (19.43) Figure19.5showstypicalhistograms.Therandomvariable Xtakesthevalues0 ,1,2,...,n indiscretestepsandcanalsobeviewedasacompositionsummationtext iXiofnindependentrandom variables Xi,one for each trial, that have the value 0 for a miss and 1 for a hit. This observationallowsus toemploythemoment-generatingfunctions angbracketleftbig etXiangbracketrightbig =P(Xi=0)+etP(Xi=1)=q+pet(19.44) and angbracketleftbig etXangbracketrightbig =productdisplay iangbracketleftbig etXiangbracketrightbig =parenleftbig pet+qparenrightbign, (19.45) FIGURE 19.5Binomialprobabilitydistributions forn=20 andp=0.1,0.3,0.5. 1130 Chapter 19 Probability from which the mean values and higher moments can be read off upon differentiating and settingt=0.Using ∂/angbracketleftetX/angbracketright ∂t=npetparenleftbig pet+qparenrightbign−1, ∂/angbracketleftetX/angbracketright ∂tvextendsinglevextendsinglevextendsinglevextendsingle t=0=/angbracketleftX/angbracketright=summationdisplay ixif(xi)=np, ∂2/angbracketleftetX/angbracketright ∂t2=npetparenleftbig pet+qparenrightbign−1+n(n−1)p2e2tparenleftbig pet+qparenrightbign−2, angbracketleftbig X2angbracketrightbig =∂2/angbracketleftetX/angbracketright ∂t2vextendsinglevextendsinglevextendsinglevextendsingle t=0=summationdisplay ix2 if(xi)=np+n(n−1)p2, weobtain,withEq. (19.17), σ2(X)=angbracketleftbig X2angbracketrightbig −/angbracketleftX/angbracketright2=np+n(n−1)p2−n2p2 =np(1−p)=npq. (19.46) Figure 19.5 illustrates these results with peaks at x=np=2,6,10,which widen with increasing p. Exercises 19.3.1 Show that the variable X=xnumber of heads in ncoin tosses is a random variable, anddetermineitsprobabilitydistribution.Describethesamplespace.Whatareitsmean value, the variance, and the standard deviation? Plot the probability function f(x)= n!/(x!(n−x)!2n)forn=10,20,30 usinggraphicalsoftware. 19.3.2 Plotthebinomialprobabilityfunctionfortheprobabilities p=1/6,q=5/6andn=6 throwsofa die. 19.3.3 A hardware company knows that the probability of mass-producing nails includes a small probability p=0.03 of defective nails (without a sharp tip usually). What is the probabilityoffindingmorethantwodefectivenailsinitscommercialboxof100nails? 19.3.4 Four cards are drawn from a shuffled bridge deck. What is the probability that they are all red? that they are all hearts? that they are honors? Compare the probabilities when thecardsareputbackatrandomplaces,or not. 19.3.5 Show that for the binomial distribution of Eq. (19.42) the most probable value of xis np. 19.4 P OISSON DISTRIBUTION The Poisson distribution typically occurs in situations involving an event repeated at a constant rate of probability, thereby depleting the population. The decay of a radioactive 19.4 Poisson Distribution 1131 sample is a case in point because, once a particle decays, it does not decay again. If the observation time dtis small enough so that the emission of two or more particles is neg- ligible, then the probability that one particle (He4inαdecay or an electron in βdecay) is emitted is µdtwith constant µandµdt≪1.We can set up a recursion relation for the probability Pn(t)of observing ncounts during a time interval t.Forn>0 the probability Pn(t+dt)iscomposedoftwomutuallyexclusiveeventsthat(i) nparticlesareemittedin thetimet,noneindt, and(ii)n−1 particlesareemittedintime t,oneindt.Therefore Pn(t+dt)=Pn(t)P0(dt)+Pn−1(t)P1(dt). Herewesubstitutetheprobabilityofobservingoneparticle, P1(dt)=µdt,andnoparticle, P0(dt)=1−P1(dt),intimedt.Thisyields Pn(t+dt)=Pn(t)(1−µdt)+Pn−1(t)µdt. So,afterrearranginganddividingby dt,weget dPn(t) dt=Pn(t+dt)−Pn(t) dt=µPn−1(t)−µPn(t). (19.47) Forn=0thisdifferentialrecursionrelationsimplifies,becausethereisnoparticleintimes tanddtgiving dP0(t) dt=−µP0(t). (19.48) The ODE says that particles have a constant decay probability and decay removes them fromthedistribution.ThisODEintegratesto P0(t)=e−µtiftheprobabilitythatnoparticle is emitted during a zero time interval P0(0)=1 is used. Here P0(0)=1 means no decay takesplaceat t≤0. NowwegobacktoEq.(19.47) for n=1, ˙P1=µparenleftbig e−µt−P1parenrightbig ,P 1(0)=0, (19.49) and solve the homogeneous equation, which is the same for P1as Eq. (19.48). This yields P1(t)=µ1e−µt.Then we solve the inhomogeneous ODE (Eq. (19.49)) by varying the constantµ1tofind˙µ1=µ,soP1(t)=µte−µt.Thegeneralsolutionis Pn(t)=(µt)n n!e−µt, (19.50) as may be confirmed by substitution into Eq. (19.47) and verifying the initial conditions, Pn(0)=0,n>0.This isanexampleof thePoissondistribution. ThePoissondistributionisdefinedwiththeprobabilities p(n)=µn n!e−µ,X=n=0,1,2,... (19.51) and is exhibited in Fig. 19.6. The random variable Xis discrete. The probabilities are properlynormalizedbecause e−µsummationtext∞ n=0µn n!=1.Themeanvalueandvariance, /angbracketleftX/angbracketright=e−µ∞summationdisplay n=1nµn n!=µe−µ∞summationdisplay n=0µn n!=µ, σ2=angbracketleftbig X2angbracketrightbig −/angbracketleftX/angbracketright2=µ(µ+1)−µ2=µ, (19.52) 1132 Chapter 19 Probability FIGURE 19.6Poissondistribution comparedwithbinomialdistribution. followfrom thecharacteristicfunction angbracketleftbig eitXangbracketrightbig =∞summationdisplay n=0eitn−µµn n!=e−µ∞summationdisplay n=0(µeit)n n!=eµ(eit−1) bydifferentiationandsetting t=0,usingEq.(19.17). APoissondistributionbecomesagoodapproximationofthebinomialdistributionfora largenumber noftrials andsmallprobability p∼µ/n,µaconstant. THEOREM:In the limit n→∞andp→0so that the mean value np→µstays finite, thebinomialdistributionbecomesaPoissondistribution. To prove this theorem, we apply Stirling’s formula (Chapter 8) n!∼√ 2πn(n/e)nfor largento the factorials in Eq. (19.42), keeping xfinite while n→∞.This yields for n→∞: n! (n−x)!∼parenleftbiggn eparenrightbiggnparenleftbigge n−xparenrightbiggn−x ∼parenleftbiggn eparenrightbiggxparenleftbiggn n−xparenrightbiggn−x ∼parenleftbiggn eparenrightbiggxparenleftbigg 1+x n−xparenrightbiggn−x ∼parenleftbiggn eparenrightbiggx ex∼nx, andforn→∞,p→0,withnp→µ: (1−p)n−x∼parenleftbigg 1−pn nparenrightbiggn ∼parenleftbigg 1−µ nparenrightbiggn ∼e−µ. 19.4 Poisson Distribution 1133 Table 19.1 i→0 12345678 9 1 0 ni→57 203 383 525 532 408 273 139 45 27 16 Finally,pxnx→µx,so altogether n! x!(n−x)!px(1−p)n−x→µx x!e−µ,n→∞, (19.53) which is a Poisson distribution for the random variable X=xwith 0≤x<∞.This limit theoremisaparticularexampleofthe lawsoflargenumbers . Exercises 19.4.1 Radioactive decays are governed by the Poisson distribution. In a Rutherford–Geiger experiment the number niof emitted αparticles is counted in n=2608 time intervals of7.5secondseach.InTable19.1 niisthenumberoftimeintervalsinwhich iparticles wereemitted.Determinetheaveragenumber λofemittedparticles,andcomparethe ni ofTable19.1with npicomputedfromthePoissondistributionwithmeanvalue λ. 19.4.2 DerivethestandarddeviationofaPoissondistributionof meanvalue µ. 19.4.3 Thenumberof αdecayparticlesofaradiumsampleiscountedperminutefor40hours. The total number is 5000 .How many 1-minute intervals are there expected to be with (a) 2,(b) 5 αparticles? 19.4.4 For a radioactive sample, 10 decays are counted on average in 100 seconds. Use the Poissondistributiontoestimatetheprobabilityofcounting 3 decaysin 10 seconds. 19.4.5238Uhasahalf-lifeof4 .51×109years.Itsdecayseriesendswiththestableleadisotope 206Pb. The ratio of the number of206Pb to238U atoms in a rock sample is measured as 0.0058. Estimate the age of the rock assuming that all the lead in the rock is from the initialdecayofthe238U,whichdeterminestherateoftheentiredecayprocess,because thesubsequentstepstakeplacefar morerapidly. Hint.The decay constant λin the decay law N(t)=Ne−λtis related to the half-life T byT=ln2/λ. ANS. 3.8×107years. 19.4.6 Theprobabilityofhittingatargetinoneshotisknowntobe 20% .Iffiveshotsarefired independently,whatistheprobabilityofstrikingthetargetatleastonce? 19.4.7 Apieceofuraniumisknowntocontaintheisotopes235 92Uand238 92Uaswellasfrom0 .80 gof206 82Pbpergramofuranium.Estimatetheageofthepiece(andthusEarth)inyears. Hint.Assume the lead comes only from the238 92U. Use the decay constant from Exer- cise19.4.5. 1134 Chapter 19 Probability 19.5 G AUSS ’NORMAL DISTRIBUTION Thebell-shapedGaussdistributionisdefinedbytheprobabilitydensity f(x)=1 σ√ 2πexpparenleftbigg −[x−µ]2 2σ2parenrightbigg ,−∞<x<∞, (19.54) with mean value µand variance σ2.It is by far the most importantcontinuous probability distributionandis displayedinFig.19.7. It isproperlynormalizedbecause,substituting y=x−µ σ√ 2,weobtain 1 σ√ 2πintegraldisplay∞ −∞e−(x−µ)2 2σ2dx=1√πintegraldisplay∞ −∞e−y2dy=2√πintegraldisplay∞ 0e−y2dy=1. Similarly,substituting y=x−µ,wesee that /angbracketleftX/angbracketright−µ=integraldisplay∞ −∞x−µ σ√ 2πe−(x−µ)2 2σ2dx=integraldisplay∞ −∞y σ√ 2πe−y2 2σ2dy=0, the integrand being odd in y,so the integral over y>0 cancels that over y<0.Similarly wecheckthatthestandarddeviationis σ. Fromthenormaldistribution(bythesubstitution y=x−/angbracketleftX/angbracketright σ) PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle>kσparenrightbig =Pparenleftbigg|X−/angbracketleftX/angbracketright| σ>kparenrightbigg =Pparenleftbig |Y|>kparenrightbig =radicalbigg 2 πintegraldisplay∞ ke−y2/2dy=radicalbigg 4 πintegraldisplay∞ k/√ 2e−z2dz=erfck√ 2, FIGURE 19.7NormalGaussdistributionformeanvaluezeroandvarious standarddeviations h=1/σ√ 2. 19.5 Gauss’ Normal Distribution 1135 we can evaluate the integral for k=1,2,3 and thus extract the following numerical rela- tionsforanormallydistributedrandomvariable: PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥σparenrightbig ∼0.3173,PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥2σparenrightbig ∼0.0455, PparenleftbigvextendsinglevextendsingleX−/angbracketleftX/angbracketrightvextendsinglevextendsingle≥3σparenrightbig ∼0.0027, (19.55) of which the last one is interesting to compare with Chebychev’s inequality (see Eq. (19.21).) giving ≤1/9 for anarbitrary probability distribution instead of ∼0.0027 forthe 3σ-ruleofthe normaldistribution. ADDITION THEOREM :If the random variables X,Yhave the same normal distributions, that is, the same mean value and variance, then Z=X+Yhas normal distribution with twicethemeanvalueandtwicethevarianceof XandY. Toprovethistheorem,wetaketheGaussdensityas f(x)=1 √ 2πe−x2/2,with1√ 2πintegraldisplay∞ −∞e−x2/2dx=1, withoutloss ofgenerality.Thentheprobabilitydensityof (X,Y)is theproduct f(x,y)=1√ 2πe−x2/21√ 2πe−y2/2=1 2πe−(x2+y2)/2. Also,Eq. (17.36)givesthedensityfor Z=X+Yas g(z)=integraldisplay∞ −∞1√ 2πe−x2/21√ 2πe−(x−z)2/2dx. Completingthesquareintheexponent, 2x2−2xz+z2=parenleftbigg x√ 2−z√ 2parenrightbigg2 +z2 2, weobtain g(z)=1 2πe−z2/4integraldisplay∞ −∞expparenleftbigg −1 2parenleftbigg x√ 2−z√ 2parenrightbigg2parenrightbigg dx. Usingthesubstitution u=x−z 2,wefindthattheintegraltransformsinto integraldisplay∞ −∞expparenleftbigg −1 2parenleftbigg x√ 2−z√ 2parenrightbigg2parenrightbigg dx=integraldisplay∞ −∞e−u2du=√π, sothedensityfor Z=X+Yis g(z)=1 2√πe−z2/4, (19.56) whichmeansithas meanvaluezeroandvariance 2 ,twicethatof XandY. In a special limit the discrete Poisson probability distribution is closely related to the continuous Gauss distribution. This limit theorem is another example of the laws of large numbers ,whichareoftendominatedbythebell-shapednormaldistribution. 1136 Chapter 19 Probability THEOREM:For large nand mean value µ, the Poisson distribution approaches a Gauss distribution. To prove this theorem for n→∞,we approximate the factorial in the Poisson’s proba- bilityp(n)ofEq. (19.51)byStirling’sasymptoticformula(seeChapter8), n!∼√ 2nπparenleftbiggn eparenrightbiggn ,n→∞, and choose the deviation v=n−µfrom the mean value as the new variable. We let the mean value µ→∞and treat v/µas small but v2/µas finite. Substituting n=µ+vand expandingthelogarithminaMacLaurinseries, keepingtwoterms, weobtain lnp(n)=−µ+nlnµ−nlnn+n−ln√ 2nπ =(µ+v)lnµ−(µ+v)ln(µ+v)+v−lnradicalbig 2π(µ+v) =(µ+v)lnparenleftbigg 1−v µ+vparenrightbigg +v−lnradicalbig 2πµ =(µ+v)parenleftbigg −v µ+v−v2 2(µ+v)2parenrightbigg +v−lnradicalbig 2πµ ∼−v2 2µ−lnradicalbig 2πµ, replacing µ+v→µbecause|v|≪µ.Exponentiating this result we find that for large n andµ p(n)→1√2πµe−v2/2µ, (19.57) whichis a Gauss distributionof the continuousvariable vwithmean value 0 and standard deviation σ=√µ. In a special limit the discrete binomial probability distribution is also closely related to the continuous Gauss distribution. This limit theorem is another example of the laws of largenumbers . THEOREM:In the limit n→∞, so that the mean value np→∞,the binomial distribu- tion becomes Gauss’ normal distribution. Recall from Section 19.4that, when np→µ< ∞,thebinomialdistributionbecomesaPoissondistribution. Instead of the large number xof successes in ntrials, we use the deviation v=x−pn fromthe(large)meanvalue pnasournewcontinuousrandomvariable,underthecondition that|v|≪pnbutv2/nis finite as n→∞.Thus, we replace xbyv+pnandn−xby qn−vinthefactorialsofEq.(19.42), f(x)→W(v)asn→∞,andthenapplyStirling’s formula.Thisyields W(v)=pxqn−xnn+1/2e−n+x+(n−x) √ 2π(v+pn)x+1/2(qn−v)n−x+1/2. 19.5 Gauss’ Normal Distribution 1137 Herewefactor outthedominantpowersof nandcancelpowersof pandqtofind W(v)=1√2πpqnparenleftbigg 1+v pnparenrightbigg−(v+pn+1/2)parenleftbigg 1−v qnparenrightbigg−(qn−v+1/2) . Intermsof thelogarithmwehave lnW(v)=ln1√2πpqn−(v+pn+1/2)lnparenleftbigg 1+v pnparenrightbigg −(qn−v+1/2)lnparenleftbigg 1−v qnparenrightbigg =ln1√2πpqn−(v+pn+1/2)parenleftbiggv pn−v2 2p2n2+···parenrightbigg −(qn−v+1/2)parenleftbigg −v qn−v2 2q2n2+···parenrightbigg =ln1√2πpqn−bracketleftbiggv nparenleftbigg1 2p−1 2qparenrightbigg +v2 nparenleftbigg1 2p+1 2qparenrightbigg +···bracketrightbigg , where v n→0,v2 n isfiniteand parenleftbigg1 2p+1 2qparenrightbigg =p+q 2pq=1 2pq. Neglectinghigherordersin v/n,suchasv2/p2n2andv2/q2n2,wefindthelarge nlimit W(v)=1√2πpqne−v2/2pqn, (19.58) whichisaGaussiandistributioninthedeviations x−pn,withmeanvalue 0 andstandard deviation σ=√npq.The large mean value pn(and the discarded terms) restricts the validityofthetheoremtothecentralpartoftheGaussianbellshape,excludingthetails. Exercises 19.5.1 What is the probability for a normally distributed random variable to differ by more than 4σfrom its mean value? Compare your result with the corresponding one from Chebychev’sinequality.Explainthedifferenceinyourownwords. 19.5.2 LetX1,X2,...,X nbeindependentnormalrandomvariableswiththesamemean ¯xand varianceσ2.Showthatsummationtext iXi/n−¯x√nσisnormalwithmeanzeroandvariance 1 . 1138 Chapter 19 Probability 19.5.3 An instructor grades a final exam of a large undergraduate class, obtaining the mean value of points Mand the variance σ2. Assuming a normal distribution for the number Mof points, he defines a grade F when M<m−3σ/2,D whenm−3σ/2<M< m−σ/2,C whenm−σ/2<M<m+σ/2,B whenm+σ/2<M<m+3σ/2, A whenM>m+3σ/2.What is the percentage of As, Fs; Bs, Ds; Cs? Redesign the cutoffs so that there are equal percentages of As and Fs (5%), 25% Bs and Ds, and 40% Cs. 19.5.4 If the random variable Xis normal with mean value 29 and standard deviation 3 ,what arethedistributionsof 2 X−1 and 3X+2? 19.5.5 Foranormaldistributionofmeanvalue mandvariance σ2,findthedistance rsuchthat halftheareaunderthebellshapeis between m−randm+r. 19.6 S TATISTICS In statistics, probability theory is applied to the evaluation of data from random experi- ments or to samples to test some hypothesis because the data have random fluctuations due to lack of complete control over the experimental conditions. Typically one attempts to estimate the mean value and variance of the distributions, from which the samples de- rive, and to generalize properties valid for a sample to the rest of the events at a prescised confidence level. Any assumption about an unknown probability distribution is called a statistical hypothesis . The concepts of tests and confidence intervals are among the most importantdevelopmentsof statistics. Error Propagation When we measure a quantity xrepeatedly, obtaining the values xjat random, or select a samplefor testing,wedeterminethemeanvalue(see Eq.(19.10)) andthevariance, ¯x=1 nnsummationdisplay j=1xj,σ2=1 nnsummationdisplay j=1(xj−¯x)2, as a measure for the error, or spread from the mean value ¯x.We can write xj=¯x+ej, wheretheerror ejisthedeviationfromthemeanvalue,andweknowthatsummationtext jej=0.(See thediscussionafterEq. (19.15).) Now suppose we want to determine a known function f(x)from these measurements; that is, we have a set fj=f(xj)from the measurements of x.Substituting xj=¯x+ej andformingthemeanvaluefrom ¯f=1 nsummationdisplay jf(xj)=1 nsummationdisplay jf(¯x+ej) =f(¯x)+1 nf′(¯x)summationdisplay jej+1 2nf′′(¯x)summationdisplay je2 j+··· =f(¯x)+1 2σ2f′′(¯x)+···, (19.59) 19.6 Statistics 1139 we obtain the average value ¯fasf(¯x)in lowest order, as expected. But in second order thereisacorrectiongivenbyhalfthevariancewithascalefactor f′′(¯x).Itisinterestingto compare this correction of the mean value with the average spread of individual fjfrom the mean value ¯f,the variance of f.To lowest order, this is given by the average of the sumofsquaresof thedeviations,inwhichweapproximate fj≈¯f+f′(¯x)ej,yielding σ2(f)≡1 nsummationdisplay j(fj−¯f)2=parenleftbig f′(¯x)parenrightbig21 nsummationdisplay je2 j=parenleftbig f′(¯x)parenrightbig2σ2.(19.60) Insummarywemayformulatesomewhatsymbolically f(¯x±σ)=f(¯x)±f′(¯x)σ asthesimplestform oferror propagationbyafunctionofonemeasuredvariable. Forafunction f(xj,yk)oftwomeasuredquantities xj=¯x+uj,yk=¯y+vk,weobtain similarly ¯f=1 rsrsummationdisplay j=1ssummationdisplay k=1fjk=1 rsrsummationdisplay j=1ssummationdisplay k=1f(¯x+uj,¯y+vk) =f(¯x,¯y)+1 rfxsummationdisplay juj+1 sfxsummationdisplay kvk+···, wheresummationtext juj=0=summationtext kvk,so again¯f=f(¯x,¯y)inlowestorder.Here fx=∂f ∂x(¯x,¯y), f y=∂f ∂y(¯x,¯y) (19.61) denote partial derivatives. The sum of squares of the deviations from the mean value is givenby rsummationdisplay j=1ssummationdisplay k+1(fjk−¯f)2=summationdisplay j,k(ujfx+vkfy)2=sf2 xsummationdisplay ju2 j+rf2 ysummationdisplay kv2 k, becausesummationtext j,kujvk=summationtext jujsummationtext kvk=0.Thereforethevarianceis σ2(f)=1 rssummationdisplay j,k(fjk−¯f)2=f2 xσ2 x+f2 yσ2 y, (19.62) withfx,fyfrom Eq.(19.61);and σ2 x=1 rsummationdisplay ju2 j,σ2 y=1 ssummationdisplay kv2 k are the variances of the xandydata points. Symbolically the error propagation for a functionoftwomeasuredvariablesmaybesummarizedas f(¯x±σx,¯y±σy)=f(¯x,¯y)±radicalBig f2xσ2x+f2yσ2y. As an application and generalization of the last result, we now calculate the error of the mean value ¯x=1 nsummationtextn j=1xjof a sample of nindividual measurements xj,each with 1140 Chapter 19 Probability spreadσ.Inthiscasethepartialderivativesaregivenby fx=1 n=fy=···andσx=σ= σy=···.Thus,ourlasterrorpropagationruletellsusthaterrorsofasumofvariablesadd quadratically,so theuncertaintyofthearithmeticmeanisgivenby ¯σ=1 nradicalbig nσ2=σ√n, (19.63) decreasingwiththenumberof measurements n. As the number nof measurements increases, we expect the arithmetic mean ¯xto con- vergetosometruevalue x.Let¯xdifferfrom xbyαandvj=xj−xbethetruedeviations; then summationdisplay j(xj−¯x)2=summationdisplay je2 j=summationdisplay jv2 j+nα2. Taking into account the error of the arithmetic mean, we determine the spread of the indi- vidualpointsabouttheunknowntruemeanvaluetobe σ2=1 nsummationdisplay je2 j=1 nsummationdisplay jv2 j+α2. AccordingtoourearlierdiscussionleadingtoEq.(19.63), α2=1 nσ2.Asaresult σ2=1 nsummationdisplay jv2 j+σ2 n, fromwhichthe standarddeviationofasampleinstatistics follows: σ=radicalBiggsummationtext jv2 j n−1=radicalBiggsummationtext j(xj−x)2 n−1, (19.64) withn−1 being the number of control measurements of the sample. This modified mean errorincludestheexpectederror inthearithmeticmean. Because the spread is not well defined when there is no comparison measurement, that is, whenn=1,the variance is sometimes defined by Eq. (19.64), in which we replace the numbernofmeasurementsbythenumber n−1 ofcontrolmeasurementsinstatistics. Fitting Curves to Data Supposewehaveasampleofmeasurements yj(forexample,aparticlemovingfreely,that is, no force) taken at known times tj(which are taken to be practically free of errors; that is, the time tis an ordinary independent variable) that we expect to be linearly related as y=at,ourhypothesis.Wewanttofitthislinetothedata. Firstweminimizethesumofdeviationssummationtext j(atj−yj)2todeterminetheslopeparame- tera,also called the regression coefficient , using the method of least squares. Differenti- atingwithrespectto aweobtain 2summationdisplay j(atj−yj)tj=0, 19.6 Statistics 1141 FIGURE 19.8Straightlinefitto datapoints (tj,yj)withtj known,yjmeasured. fromwhich a=summationtext jtjyjsummationtext jt2 j(19.65) follows.Notethatthenumeratorisbuiltlikeasamplecovariance,thescalarproductofthe variables t,yof the sample. As shown in Fig. 19.8, the measured values yjdo not lie on thelineasarule.Theyhavethespread(orrootmeansquaredeviationfromthefittedline) σ=radicalBiggsummationtext j(yj−atj)2 n−1. Alternatively,letthe yjvaluesbeknown(withouterror) while tjare measurements.As suggestedbyFig.19.9,inthiscaseweneedtointerchangetheroleof tandyandtofitthe linet=bytothedatapoints.Weminimizesummationtext j(byj−tj)2,setthederivativewithrespect FIGURE 19.9Straightlinefitto datapoints (tj,yj)withyj known,tjmeasured. 1142 Chapter 19 Probability FIGURE 19.10(a) Straightlinefittodatapoints (tj,yj).(b) Geometryof deviations uj,vj,dj. tobequaltozero,andfindsimilarlytheslopeparameter b=summationtext jtjyjsummationtext jy2 j. (19.66) In case both tjandyjhave errors (we take tandyto have the same units), we have to minimizethesumofsquaresofthedeviationsofbothvariablesandfittoaparameterization tsinα−ycosα=0,wheretandyoccuronanequalfooting.AsdisplayedinFig.(19.10a) thismeansgeometricallythatthelinehastobedrawnsothatthesumofthesquaresofthe distances djofthepoints (tj,yj)fromthelinebecomesaminimum.(SeeFig.19.10band Chapter 1.) Here dj=tjsinα−yjcosα,sosummationtext jd2 j=minimum must be solved for the angleα.Settingthederivativewithrespecttotheangleequaltozero, summationdisplay j(tjsinα−yjcosα)(tjcosα+yjsinα)=0, yields sinαcosαsummationdisplay jparenleftbig t2 j−y2 jparenrightbig −parenleftbig cos2α−sin2αparenrightbigsummationdisplay jtjyj=0. Thereforetheangleofthestraight-linefitis givenby tan2α=2summationtext jtjyjsummationtext j(t2 j−y2 j). (19.67) This least-squares fitting applies when the measurement errors are unknown. It allows as- signing at least some kind of error bar to the measured points. Recall that we did not use errors for the points. Our parameter a(orα) is most likely to reproduce the data under these circumstances. More precisely, the least-squares method is a maximum-likelihood estimate of the fitted parameters when it is reasonable to assume that the errors are in- dependent and normally distributed with the same deviation for all points . This fairly 19.6 Statistics 1143 strong assumption can be relaxed in “weighted” least-squares fits called chi square fits .4 (SeealsoExample19.6.1.) Theχ2Distribution This distribution is typically applied to fits of a curve y(t,a,...) with parameters a,... to datatjusing the method of least squares involving the weighted sum of squares of deviations;thatis, χ2=Nsummationdisplay j=1parenleftbiggyj−y(tj,a,...) /Delta1yjparenrightbigg2 isminimized,where Nisthenumberofpointsand risthenumberofadjustedparameters a,....This quadratic merit function gives more weight to points with small measurement uncertainties /Delta1yj. We represent each point by a normally distributed random variable Xwith zero mean value and variance σ2=1,the latter in view of the weights in the χ2function. In a first step, we determine the probability density for the random variable Y=X2of a single point that takes only positive values. Assuming a zero mean value is no loss of generality because,if/angbracketleftX/angbracketright=m/negationslash=0,wewouldconsidertheshiftedvariable Y=X−m,whosemean valueis zero.Weshowthatif XhasaGaussnormaldensity f(x)=1 σ√ 2πe−x2/2σ2,−∞<x<∞, thentheprobabilityof therandomvariable Yis zeroif y≤0,and P(Y<y)=P(X2<y)=P(−√y<X<√y)ify>0. From the continuous normal distribution P(y)=integraltexty −∞f(x)dx, we obtain the probability densityg(y)bydifferentiation: g(y)=d dybracketleftbig P(√y)−P(−√y)bracketrightbig =1 2√yparenleftbig f(√y)+f(−√y)parenrightbig =1 σ√2πye−y/2σ2,y>0. (19.68) This density,∼e−y/2σ2/√y,corresponds to the integrand of the Euler integral of the gammafunction.Suchaprobabilitydistribution g(y)=yp−1 Ŵ(p)(2σ2)pe−y/2σ2 is called a gamma distribution with parameters pandσ.Its characteristic function for ourcase, p=1/2,is proportionaltotheFouriertransform angbracketleftbig eitYangbracketrightbig =1 σ√ 2πintegraldisplay∞ 0e−y(1/2σ2−it)dy√y=1 σ√ 2π(1 2σ2−it)1/2integraldisplay∞ 0e−xdx√x =parenleftbig 1−2itσ2parenrightbig−1/2. 4For more details,seeChapter 14 of Press etal. in theAdditional Readings of Chapter9. 1144 Chapter 19 Probability Sincethe χ2samplefunctioncontainsasumofsquares,weneedthefollowingtheorem. ADDITION THEOREM :for the gamma distributions :If the independent random variables Y1andY2have a gamma distribution with p=1/2,and the same σthenY1+Y2has a gammadistributionwith p=1. SinceY1andY2are independent, the product of their densities (Eq. (19.36)) generates thecharacteristicfunction angbracketleftbig eit(Y1+Y2)angbracketrightbig =angbracketleftbig eitY1eitY2angbracketrightbig =angbracketleftbig eitY1angbracketrightbigangbracketleftbig eitY2angbracketrightbig =parenleftbig 1−2itσ2parenrightbig−1. (19.69) Now we come to the second step. We assess the quality of the fit by the random variable Y=summationtextn j=1X2 j,wheren=N−ris the number of degrees of freedom for Ndata points andrfitted parameters. The independent random variables Xjare taken to be normally distributed with the (sample) variance σ2.(In our case r=1 andσ=1.)T h eχ2analysis does not really test the assumptions of normality and independence, but if these are not approximately valid, there will be many outlying points in the fit. The addition theorem givestheprobabilitydensity(Fig.19.11)for Y, gn(y)=yn 2−1 2n/2σnŴ(n 2)e−y/2σ2,y>0, andgn(y)=0ify<0,whichisthe χ2distributioncorrespondingto ndegreesoffreedom. Its characteristicfunctionis angbracketleftbig eitYangbracketrightbig =parenleftbig 1−2itσ2parenrightbig−n/2. Differentiatingandsetting t=0 weobtainitsmeanvalueandvariance /angbracketleftY/angbracketright=nσ2,σ2(Y)=2nσ4. (19.70) FIGURE 19.11χ2probabilitydensity gn(y). 19.6 Statistics 1145 Table 19.2 χ2Distribution nv=0.8 v=0.7 v=0.5 v=0.3 v=0.2 v=0.1 1 0.064 0.148 0.455 1.074 1.642 2 .706 2 0.446 0.713 1.386 2.408 3.219 4 .605 3 1.005 1.424 2.366 3.665 4.642 6 .251 4 1.649 2.195 3.357 4.878 5.989 7 .779 5 2.343 3.000 4.351 6.064 7.289 9 .236 6 3.070 3.828 5.348 7.231 8.558 10 .645 Entries are χvfor the probabilities v=P(χ2≥χ2v)=1 2n/2Ŵ(n/2)integraltext∞ χ2ve−y/2y(n/2)−1dyforσ=1. Tablesgivevaluesfor the χ2probabilityfor ndegreesoffreedom, Pparenleftbig χ2≥y0parenrightbig =1 2n/2σnŴ(n 2)integraldisplay∞ y0yn/2−1e−y/2σ2dy forσ=1andy0>0.TouseTable19.2for σ/negationslash=1,rescale y0=v0σ2sothatP(χ2≥v0σ2) corresponds to P(χ2≥v0)of Table 19.2. The following example will illustrate the whole process. Example 19.6.1 Let us apply the χ2function to the fit in Fig. 19.8. The measured points (tj,yj±/Delta1yj) witherrors /Delta1yjare (1,0.8±0.1), (2,1.5±0.05), (3,3±0.2). Forcomparison,themaximum-likelihoodfit, Eq.(19.65), gives a=1·0.8+2·1.5+3·3 1+4+9=12.8 14=0.914. Minimizinginstead, χ2=summationdisplay jparenleftbiggyj−atj /Delta1yjparenrightbigg2 , gives 0=∂χ2 ∂a=−2summationdisplay jtj(yj−atj) (/Delta1yj)2, or a=summationtext jtjyj (/Delta1yj)2 summationtext jt2 j (/Delta1yj)2. 1146 Chapter 19 Probability Inourcase a=1·0.8 0.12+2·1.5 0.052+3·3 0.22 12 0.12+22 0.052+32 0.22=1505 1925=0.782 is dominated by the middle point with the smallest error, /Delta1y2=0.05.The error propaga- tionformula(Eq. (19.62)) givesus thevariance σ2 aof theestimateof a, σ2 a=summationdisplay j(/Delta1yj)2parenleftbigg∂a ∂yjparenrightbigg2 =summationdisplay jt2 j (/Delta1yj)2 parenleftbigsummationtext kt2 k (/Delta1yk)2parenrightbig2=1 summationtext jt2 j (/Delta1yj)2 using ∂a ∂yj=tj (/Delta1yj)2 summationtext kt2 k (/Delta1yk)2. Forourcase, σa=1/√ 1925=0.023;thatis, ourslopeparameteris a=0.782±0.023. To estimate the quality of this fit of a,we compute the χ2probability that the two independent(control)pointsmissthefitbytwostandarddeviations;thatis,onaverageeach point misses by one standard deviation. We apply the χ2distribution to the fit involving N=3datapointsand r=1parameter,thatis,for n=3−1=2degreesoffreedom.From Eq.(19.70)the χ2distributionhasameanvalue2andavariance4.Aruleofthumbisthat χ2≈nfor a reasonably good fit. Then P(χ2≥2)∼0.496 is read off Table 19.2, where weinterpolatebetween P(χ2≥1.3862)=0.50 andP(χ2≥2.4082)=0.30 asfollows: Pparenleftbig χ2≥2parenrightbig =Pparenleftbig χ2≥1.3862parenrightbig −2−1.3862 2.4082−1.3862bracketleftbig Pparenleftbig χ2≥1.3862parenrightbig −Pparenleftbig χ2≥2.4082parenrightbigbracketrightbig =0.5−0.02·0.2=0.496. Thus the χ2probability that, on average, each point misses by one standard deviation is nearly 50% andfairlylarge. /squaresolid Our next goal is to compute a confidence interval for the slope parameter of our fit. Aconfidenceintervalforanaprioriunknownparameterofsomedistribution(forexample, adetermined by our fit) is an interval that contains anot with certainty but with a high probability p,theconfidencelevel,whichwecanchoose.Suchanintervaliscomputedfor agivensample.SuchananalysisinvolvestheStudent tdistribution. The Student tDistribution Becausewealwayscomputethearithmeticmeanofmeasuredpoints,wenowconsiderthe samplefunction ¯X=1 nnsummationdisplay j=1Xj, 19.6 Statistics 1147 where the random variables Xjare assumed independent with a normal distribution of the same mean value mand variance σ2. The addition theorem for the Gauss distrib- ution tells us that X1+···+Xnhas the mean value nmand variance nσ2.Therefore (X1+···+Xn)/nisnormalwithmeanvalue mandvariance nσ2/n2=σ2/n.Theprob- abilitydensityof thevariable ¯X−mis theGaussdistribution ¯f(¯x−m)=√n σ√ 2πexpparenleftbigg −n(¯x−m)2 2σ2parenrightbigg . (19.71) The key problem solved by the Student tdistribution is to provide estimates for the mean valuem, whenσis not known , in terms of a sample function whose distribution is inde- pendentofσ.Tothis end,wedefinearescaledsamplefunction(traditionallycalled) t: t=¯X−m S√ n−1,S2=1 nnsummationdisplay j=1(Xj−¯X)2. (19.72) It can be shown that tandSare independent random variables. Following the arguments leading to the χ2distribution, the density of the denominator variable Sis given by the gammadistribution d(s)=n(n−1)/2sn−2e−ns2/2σ2 2n−3 2Ŵ(n−1 2)σn−1. (19.73) The probability for the ratio Z=X/Yof two independent random variables X,Ywith normaldensityfor ¯fanddasgivenbyEqs. (19.71)and(19.73)is (Eq. (19.40)) R(z)=integraldisplayz −∞integraldisplay∞ −∞f(yz)d(y)|y|dydz, (19.74) sothevariable V=(¯X−m)/Shasthedensity r(v)=integraldisplay∞ 0√n σ√ 2πexpparenleftbigg −nv2s2 2σ2parenrightbiggn(n−1)/2sn−2e−ns2/2σ2 2n−3 2Ŵ(n−1 2)σn−1sds =nn/2 σn√π2(n−2)/2Ŵ(n−1 2)integraldisplay∞ 0e−ns2(v2+1)/2σ2sn−1ds. Herewesubstitute z=s2andobtain r(v)=nn/2 σn√π2n/2Ŵ(n−1 2)integraldisplay∞ 0e−nz(v2+1)/2σ2z(n−2)/2dz. Nowwesubstitute Ŵ(1/2)=√π,definetheparameter a=n(v2+1) 2σ2, andtransform theintegralinto Ŵ(n/2)/an/2tofind r(v)=Ŵ(n/2) √πŴ(n−1 2)(v2+1)n/2,−∞<v<∞. 1148 Chapter 19 Probability FIGURE 19.12Studenttprobability densitygn(y)forn=3. Table 19.3 StudenttDistribution pn =1 n=2 n=3 n=4 n=5 0.8 1 .38 1 .06 0 .98 0.94 0.92 0.9 3 .08 1 .89 1 .64 1.53 1.48 0.95 6 .31 2 .92 2 .35 2.13 2.02 0.975 12 .74 .30 3 .18 2.78 2.57 0.99 31 .86 .96 4 .54 3.75 3.36 0.999 318 .32 2.31 0.2 7.17 5.89 Entries arethe values CinP(C)=KnintegraltextC −∞(1+t2 n)−(n+1)/2dt=p,nis the number of degrees of freedom. Finally we rescale this expression to the variable tin Eq. (19.72) with the density (Fig.19.12) g(t)=Ŵ(n/2) √π(n−1)Ŵ(n−1 2)(1+t2 n−1)n/2,−∞<t<∞,(19.75) fortheStudent tdistribution,whichmanifestlydoesnotdependon morσ.Theprobability fort1<t<t2isgivenbytheintegral P(t1,t2)=Ŵ(n/2)√π(n−1)Ŵ(n−1 2)integraldisplayt2 t1dt (1+t2 n−1)n/2, (19.76) andP(z)≡P(−∞,z)is tabulated. (See Table 19.3 for example.) Also, P(∞,−∞)=1 andP(−z)=1−P(z),becausetheintegrandinEq(19.76) isevenin t,so integraldisplay−z −∞dt (1+t2 n−1)n/2=integraldisplay∞ zdt (1+t2 n−1)n/2 and integraldisplay∞ zdt (1+t2 n−1)n/2=integraldisplay∞ −∞dt (1+t2 n−1)n/2−integraldisplayz −∞dt (1+t2 n−1)n/2. 19.6 Statistics 1149 Multiplying this by the factor preceding the integral in Eq. (19.76) yields P(−z)=1− P(z).In the following example we show how to apply the Student tdistribution to our fit ofExample19.6.1. Example 19.6.2 CONFIDENCE INTERVAL Here we want to determine a confidence interval for the slope ain the linear y=atfit of Fig.19.8.Weassume •firstthatthesamplepoints (tj,yj)are randomandindependent,and •secondthat,foreachfixedvalue t,therandomvariable Yisnormalwithmean µ(t)= atandvariance σ2independentof t. These values yjare measurements of the random variable Y,but we will regard them as single measurements of the independent random variables Yjwith the same normal distributionas Y(whosevariancewedonotknow). We chooseaconfidencelevel, p=95%,say.ThentheStudentprobabilityis P(−C,C)=P(C)−P(−C)=p=−1+2P(C), hence P(C)=1 2(1+p), usingP(−C)=1−P(C),and P(C)=1 2(1+p)=0.975=KnintegraldisplayC −∞parenleftbigg 1+t2 nparenrightbigg−(n+1)/2 dt, whereKn−1is the factor preceding the integral in Eq. (19.76). Now we determine a solu- tionC=4.3 from Table 19.3 of Student’s tdistribution, with n=N−r=3−1=2t h e numberofdegreesoffreedom,notingthat (1+p)/2 correspondsto pinTable19.3. Then we compute A=Cσa/√ Nfor sample size N=3.The confidence interval is givenby a−A≤a≤a+A,atp=95%confidencelevel . From the χ2analysis of Example 17.6.1 we use the slope a=0.782 and variance σ2 a= 0.0232,soA=4.30.023√ 3=0.057,and the confidence interval is determined by a−A= 0.782−0.057=0.725,a+A=0.839,or 0.725<a<0.839 at95%confidencelevel . Comparedto σa,theuncertaintyof ahasincreasedduetothehighconfidencelevel.Alook at Table 19.3 shows that a decrease in confidence level, p,reduces the uncertainty inter- val, and increasing the number of degrees of freedom, n,would also lower the range of uncertainty. /squaresolid 1150 Chapter 19 Probability Exercises 19.6.1 Let/Delta1Abetheerror ofameasurementof A,etc.Useerror propagationtoshowthat parenleftbiggσ(C) Cparenrightbigg2 =parenleftbiggσ(A) Aparenrightbigg2 +parenleftbiggσ(B) Bparenrightbigg2 holdsfortheproduct C=ABandtheratio C=A/B. 19.6.2 Find the mean value and standard deviation of the sample of measurements x1= 6.0,x2=6.5,x3=5.9,x5=6.2.If the point x6=6.1 is added to the sample, how doesthechangeaffect themeanvalueandstandarddeviation? 19.6.3 (a)Carryouta χ2analysisofthefitofcase binFig.19.9assumingthesameerrorsfor theti,/Delta1 ti=/Delta1yi,a sf o rt h e yiused in the χ2analysis of the fit in Fig. 19.8. (b) Deter- minetheconfidenceintervalat95% confidencelevel. 19.6.4 Ifx1,x2,...,xnareasampleofmeasurementswithmeanvaluegivenbythearithmetic mean¯xand the corresponding random variables Xjthat take the values xjwith the same probability are independent and have mean value µand variance σ2,then show that/angbracketleft¯x/angbracketright=µandσ2(¯x)=σ2/n.If¯σ2=1 nsummationtext j(xj−¯x)2is the sample variance, show that/angbracketleft¯σ2/angbracketright=n−1 nσ2. AdditionalReadings Kreyszig, E., Introductory Mathematical Statistics: Principles and Methods. NewYork: Wiley (1970). Suhir, E., Applied Probability for Engineers and Scientists. NewYork: McGraw-Hill(1997). Papoulis, A., Probability, Random Variables, and Stochastic Processes ,3rd ed.NewYork: McGraw-Hill(1991). Ross, S. M., FirstCourse in Probability , 5th ed.,Vol. A.NewYork: Prentice-Hall (1997). Ross, S. M., Introduction to Probability Models ,7th ed.NewYork: AcademicPress (2000). Ross,S.M., IntroductiontoProbabilityandStatisticsforEngineersandScientists ,2nded.NewYork:Academic Press (1999). Chung, K.L., ACourse in Probability Theory Revised ,3rd ed. NewYork: AcademicPress (2000). Devore,J.L., ProbabilityandStatisticsforEngineeringandtheSciences ,5thed.NewYork:DuxburyPr.(1999). Montgomery,D.C.,andG.C.Runger, AppliedStatisticsandProbabilityforEngineers ,2nded.NewYork:Wiley (1998). Degroot, M.H., Probability and Statistics , 2nd ed.NewYork: Addison-Wesley (1986). Bevington,P.R.,andD.K.Robinson, DataReductionandErrorAnalysisforthePhysicalSciences ,3rd.ed.New York: McGraw-Hill(2003). GeneralReferences Additional, more specializedreferences arelistedatthe endof eachchapter. 1. E. T. Whittaker and G. N. Watson, A Course of modern Analysis , 4th ed. Cambridge, UK: Cambridge Uni- versity Press (1962), paperback. Although this is the oldest of the references (original edition 1902), it still is the classicreference.It leansstrongly toward pure mathematics,as of1902, with full mathematicalrigor. 2. P.M.MorseandH.Feshbach, MethodsofTheoreticalPhysics ,2vols.NewYork:McGraw-Hill(1953).This work presents the mathematics of much of theoretical physics in detail but at a rather advanced level. It is recommended as the outstanding source of information for supplementary reading and advancedstudy. 19.6 General References 1151 3. H. S. Jeffreys and B. S. Jeffreys, Methods of Mathematical Physics , 3rd ed. Cambridge, UK: Cambridge University Press (1972). This is a scholarly treatment of a wide range of mathematical analysis, in which considerableattentionispaidtomathematicalrigor.Applicationsaretoclassicalphysicsandtogeophysics. 4. R. Courant and D. Hilbert, Methods of Mathematical Physics , Vol. 1 (1st English ed.). New York: Wiley (Interscience)(1953). Asareferencebookformathematicalphysics,itisparticularly valuableforexistence theoremsanddiscussionsofareassuchaseigenvalueproblems,integralequations,andcalculusofvariations. 5. F. W. Byron Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics , Reading, MA: Addison- Wesley, reprinted, Dover (1992). This is an advanced text that presupposes a moderate knowledge of math- ematicalphysics. 6. C. M. Bender and S. A. Orszag, Advanced Mathematical Methods for Scientists and Engineers .N e wY o r k : McGraw-Hill(1978). 7.Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables, Applied Math- ematics Series-55 (AMS-55). Washington, DC: National Bureau of Standards, U.S. Department of Com- merce;reprinted,Dover(1974).Asatremendouscompilationofjustwhatthetitlesays,thisisanextremely useful reference. This page intentionally left blank INDEX Numbers 1-forms, 304–5 2-forms, 305–6 3-forms, 306–7 A Abeliangroup, 242 Abel’sequation, 1016 Abel’stest, 351, 665, 882 Abel’stheorem, 882 absolute convergence, 340–42, 350, 363 addition of matrices, 178–79 of series, 324–25 of tensors, 136 addition rule, 1111 addition theorem forspherical harmonics, 797–802 Bessel functions, 636 derivation of addition theorem, 798–800 Legendre polynomials, 798 trigonometric identity, 797–98 adjoint operator property, 208 algebraicform, 405 aliasing, 916–17 alternating series, 339–42 absolute convergence, 340–42 exercises, 342 Leibniz criterion, 339–40 overview, 339 analyticcontinuation, 432–34 analyticfunctions, 415–18 z∗,416 z2,415 analyticlandscape, 489–90 angularMathieuequation, 872 angular momentum, 18, 215, 267, seealsoorbital angular momentum coupling, 266–78 Clebsch–Gordan coefficients SU(2)and SO(3), 267–70 exercises, 277–78 overview, 266–67 spherical tensors, 271–74 young tableaux for SU(n), 274–77angular momentum operators, 261 Clebsch–Gordan coefficients, 803 orbital, 793–96 spherical harmonics, 793 vector spherical harmonics, 813–16 annihilation operator, 824 anomalous dispersion, 998 anticommutation relation, 18 anti-Hermitian matrices,221–23 eigenvalues degenerate,223 andeigenvectors ofreal symmetric matrices, 221–22 overview, 221 antisymmetry antisymmetric matrices, 204 and determinants, 168–72 Gausselimination, 170–72 overview, 168–70 of tensors, 137, 147 arealawfor planetary motion, 116–19 Argand diagram, 405 associatedLegendre equation functions, see Legendre equation, functions polynomials associative, 2 asymptotic expansions, 719–25 asymptotic forms offactorial function Ŵ(1+s),494–95 of Hankelfunction, 493–94 asymptotic series, 389–96 Bessel functions, 722 confluent hypergeometric functions, 393 cosine and sine integrals, 392–93 definition of, 393–94 exercises, 394–96 incomplete gamma function, 389–92 overview, 389 overview: integral representation expansion, 719–23 steepestdescent, 489–95 Stokes’ method, 719, 724 asymptotic values, Besselfunctions, 722, 729 attractor, 1081 autonomous differential equations, 1091–93 average value,1117–18 axes,seerotations; symmetry 1153 1154 Index axialvector, 143–44 axis,coordinate, 4–5, 7–11, 195–99 azimuthal dependence—orthogonality, 787 B Baker–Hausdorff formula, 225 basin of attraction,1081 Bernoulli and Riccatiequations, 1089–90 Bernoulli numbers, 376–89, 473–74 Euler–Maclaurin integration formula, 380–82 exercises,385–89 improvement of convergence, 385 overview, 376–79 polynomials, 379–80 Riemann Zetafunction, 382–84 Bessel functions, 675–739, 865 asymptotic expansions, 719–25 exercises, 723–25 expansion ofan integral representation, 720–23 asymptotic values,722 closure equation, 696 of first kind, 675–93 alternateapproaches, 685–86 Bessel functions of nonintegral order, 686 Bessel’s differential equation, 678–79 Bessel’s differential equation: self-adjoint form, 694 confluent hypergeometric representation, 865 cylindrical resonant cavity, 682–85 cylindrical waveguide, 705 exercises, 686–93 Fourier transform, 933 Fraunhofer diffraction, circularaperture, 680–82 generatingfunction for integralorder,675–77 integral representation, 679–80 Laplacetransform solution, 983–84 orthogonality, 694 recurrence relations, 677–78 second kinds, 699–707 series solution, 570–72, 676–79 singularities, 564 spherical, 725–39 Wronskian, 702–5 Fraunhofer diffraction, 680–82 Hankelfunctions, 707–13 contour integral representation of, 709–11 cylindrical traveling waves,708 definitions, 707–8 exercises, 711–13 Helmholtz equation, 683–84, 725 Laplace’sequation, 695 modified, 713–19 asymptotic expansion, 711, 719exercises,716–19 Fourier transform, 716 generating function, 709 integral representation, 720–23 Laplacetransform, 933 recurrencerelations, 714–16 seriesform, 714 Neumann functions, Bessel functions of second kind, 699–707 coaxialwaveguides, 703–4 definition and series form, 699–700 exercises,704–7 other forms, 701 recurrencerelations, 702 Wronskian formulas, 702–3 of nonintegral order, 686 orthogonality, 694–99 Besselseries, 695 continuum form, 696 electrostaticpotential in a hollow cylinder, 695–96 exercises,697–99 normalization, 695 recurrence relations, 677 spherical, 725–39 asymptotic values,729 definitions, 726–29 exercises,732–39 limiting values, 729–30 orthogonality, 731 particlein a sphere, 731–32 recurrencerelations, 730 spherical waves,730 in wave guides, 703–4 zeros,682 Bessel’s differential equation, 678–79, 684 self-adjoint form, 694 Bessel’s equation, 983–84 Bessel series,695 Bessel’s inequality, 651–52 betafunction, 520–26 definite integrals, alternate forms, 521–22 derivation of Legendre duplication formula, 522–23 incomplete, 523 Laplaceconvolution, 993 verification of πα/sinπαrelation, 522 bifurcates, 1083 bifurcations in dynamical systems, 1103–4 Hopf, 1101, 1103–4, 1107 pitchfork, 1083, 1086, 1101, 1107 binomial coefficient,356 binomial distribution, 1128–30 binomial expansion, 1129 binomial probability distribution, 1129 Index 1155 binomial theorem, 356–57 Biot and Savart law, 780–82 blackhole, optical pathnearevent horizon of, 1041–42 Bohr radius, 844 Born approximation, 603 Bose–Einstein statistics, 1064, 1115 boundary conditions, 542–43 Cauchy, 542 Dirichlet, 543 hollow cylinder, 695 magnetic field of current loop, 778–82 Neumann, 543 ring of charge,761 sphere in uniform electricfield, 759 Sturm–Liouville theory, 627–29 waveguide, coaxial cable,703 bound state,627 box counting dimension, 1086 branchcut (cut line), 409 branch points, 440–42 and multivalent functions, 447–50 of order 2, 440–42 Bromwich integral, 994–95 Butterfly effect,1079 C calculusof residues, 455–82, seealsodefinite integrals Cauchy principal value, 457–60 exercises, 474–82 Jordan’s lemma, 466–68 overview, 455 pole expansion of meromorphic functions, 461 product expansion of entire functions, 462–63 residue theorem, 455–56 calculusof variations, 1037–77 applications of theEuler equation, 1044–52 exercises, 1049–52 soap film, 1045–46 soap film—minimum area,1046–49 straight line, 1044–45 dependent and an independent variable, 1038–44 alternateforms of Eulerequations, 1042 concept of variation, 1038–41 exercises, 1043–44 missing dependent variables, 1042–43 optical pathnearevent horizon ofablack hole, 1041–42 Lagrangian multipliers, 1060–65 constraints, 1060–72 cylindrical nuclearreactor, 1062 exercises, 1063–65 particle in abox, 1061–62Rayleigh–Ritz variational technique, 1072–76 exercises,1074–76 ground stateeigenfunction, 1073 Sturm–Liouville equation, 1072 vibrating string, 1074 several dependent and independent variables, 1058–59 exercises,1059 relation to physics, 1059 several dependent variables, 1052–58 exercises,1055–56 Hamilton’s principle, 1053–54 Laplace’sequation, 1057–58 moving particle—Cartesian coordinates, 1054 moving particle—circular cylindrical coordinates, 1054–55 several independent variables,exercises, 1058 surface of revolution, 1046 uses of, 1037 variation with constraints, 1065–72 exercises,1070–72 Lagrangian equations, 1066–67 Schrödinger wave equation, 1069–70 simple pendulum, 1067–68 sliding off a log, 1068–69 Cartesian components, 4 Cartesian coordinates, 554–55 unit vectors, 5 Casimir operators, 265 Catalan’s constant, 384, 513 catenoid,catenary ofrevolution, 1046 Cauchy (Maclaurin) integral test, 327–30, 418–30 contour integrals, 418–20 derivatives, 426–27 exercises, 424–25, 429–30 Goursat proof, 421–23 Morera’s theorem, 427–28 multiply connected regions, 423–24 overview, 418, 425–26 Stokes’ theorem proof, 420–21 Cauchy boundary conditions, 542 Cauchy criterion, 322 Cauchy inequality, 428 Cauchy principal value,457–60, 471 Cauchy–Riemann conditions, 413–18 analyticfunctions, 415–18 z∗,416 z2,415 exercises, 416–18 overview, 413–15 causality, 486–87 cavities,cylindrical, 682–85 Cayley–Klein parameters,252 centeror cycle,1100–1101 1156 Index centralforce, 117 centralforce field, 39–40, 44–46 centrifugal potentials, 72 chainrule, 34 chaos in dynamical systems, 1105–6 chaoticattractor, 1085 character,184, 293 characteristics,538–41 Chebyshev differential equation, 559 Chebyshev polynomials, 848–59 generating functions, 848 Gram–Schmidt construction, 646 hypergeometric representations, 862 orthogonality, 854–55 recurrencerelation, 850 recurrencerelations—derivatives, 852–53 shifted, 646, 850 trigonometric form, 853–54 type I, 849–52 type II, 849 chi-squared ( χ2) distribution, 1143–45 Christoffel symbols, 154–56, 314 circularcylinder coordinates, 115–23 arealaw for planetary motion, 116–19 exercises,120–23 Navier–Stokes term, 119 overview, 115–16 circularcylindrical coordinates, 555–56 expansion, 601–2 circularmembrane, Bessel functions, 693, 708 classesandcharacter,293 Clausen functions, 909 Clebsch–Gordan coefficients, 267–70, 803 Clifford algebra,211–12 closed-form solutions, 810–12 closure, of Bessel function, 89, 248, 696 closure, of spherical harmonics, 790, 792 coaxialwaveguides, 703–4 commutative, 1 commutator, 180, 225, 231, 249, 253, 262, 264 comparison tests, 325–26 completeness of eigenfunctions of Fourier series: of Sturm–Liouville eigenfunctions, 649–51 of Hilbert–Schmidt: of integral equations, 1031–33 complex variables,403–54, 455–97, seealso calculusof residues; Cauchy–Riemann conditions; functions; mapping; saddle points (steepest descent method); singularities algebra using, 404–13 calculus of residues, 455–82 complex conjugation, 407–8 exercises, 409–13overview, 404–5 permanenceof algebraicform, 405–7 Cauchy’s integral formula, 425–30 derivatives, 426–27 exercises,429–30 Morera’s theorem, 427–28 overview, 425–26 Cauchy’s integral theorem, 418–25 Cauchy–Goursat proof, 421–23 contour integrals, 418–20 exercises,424–25 multiply connected regions, 423–24 overview, 418 Stokes’ theorem proof, 420–21 dispersion relations, 482–89 causality,486–87 exercises,487–89 opticaldispersion, 484–85 overview, 482–83 Parseval relation, 485–86 symmetry relations, 484 Laurentexpansion, 430–38 analyticcontinuation, 432–34 exercises,437–38 Schwarzreflection principle, 431–32 Taylorexpansion, 430–31 overview, 403–4, 455 conditional convergence, 340 conditional probability, 1112 condition number, of ill-conditioned systems, 234 Condon–Shortley phase conventions, 270 confidence interval, 1146, 1149 confluent hypergeometric functions, 863–69 asymptotic expansions, 866 Bessel and modified Bessel functions, 865 Hermitefunctions, 866 integral representations, 865 Laguerrefunctions, 837–48 miscellaneous cases,866 Whittaker functions, 866 Wronskian, 868 conformal mapping, 451–54 conjugation, complex, 407–8 connected,simply or multiply, 60, 95, 420, 423, 426, 435 conservation theorem, 309 conservative force, 34, 69 constant 1-forms, 305 constantBfield, vector potentials of, 44 contiguous function relations, 861 continuation, analytic,432–34 continuity equation, 40–42 continuity of power series, 364 continuous random variable, 1117–19 continuum form, 696 Index 1157 contour integral representation, Hankelfunctions, 709–11 contour integrals, 418–20 contour of integration, simple pole on, 468–69 contraction, 139 contravariant tensor, 135, 156–58, 160, 162 contravariant vector, 134 convergence, rate of, 334, 345 convergence of infinite series,321, 903 absolute, 340–42, 350 improvement of, and Bernoulli numbers, 385 of infinite product, 397–98 of powerseries, 363 rate, 334, 345 and rational approximations, 345 tests, 325–39, seealsoCauchy(Maclaurin) integral test comparison, 325–26 exercises, 335–39 Gauss’, 332–33, 357 improvement of, 334–39 Kummer’s, 330–32 overview, 325 partial sum approximation, 390 Raabe’s, 332 uniform and nonuniform, 348–49 convolution (Faltungs) theorem, 951–55, 990–94 driven oscillator with damping, 991–93 Fourier transform, 931–32, 936–45 Laplacetransform, 965–1003 Parseval’s relation, 952–53 coordinates, seealsocircularcylinder coordinates; curved coordinates and vectors; orthogonal coordinates; spherical polarcoordinates axes, rotation of, seerotations curvilinear, 104, 105, 110, 111, 112 divergence of coordinate vector, 39 Laplacianin orthogonal, 316 rotation of, 199 correlation, 1122 cosets and subgroups, 293–94 cosines asymptotic expansion, 392–93 confluent hypergeometric representation, 867 cosine transform, 939 direction, 196–97 direction cosines (orthogonal matrices), 4, 196–97, 201 functions of infinite products, 398–99 infinite product, 398, 462 integral, 392 integral of in denominator, 464–65 integrals in asymptotic series,392–93 law of, 16–17, 118, 745 law of, theorem, 16theorem, 16 coupling, angular momentum, seeangular momentum covariance,1122 covariance ofMaxwell’sequations, Lorentz, see Lorentz covarianceof Maxwell’sequations covariant derivative, 151, 156 covariant vectors, 134, 152–53 tensor, 135, 156–58, 160, 162 Cramer’s rule, 166 creation operator, 824 criterion, Leibniz, 339–40 criticalpoint, 1091–93 criticalstrip, 897 criticaltemperature, 1086 crossing conditions, 484 cross product, 18–22, seealsotriple vector products exercises, 22–25 overview, 18–22 of vectors, 315 crystallographic point and spacegroups, 299–300 curl,∇×, 43–49 centralforce field, 44–46 asdifferential vector operator, 112–13 exercises, 47–49 gradient of dot product, 46 integral definitions of gradient, divergence and, 58–59 integration by parts of, 47 overview, 43 astensor derivative operator, 162–63 vector potential ofconstant Bfield,44 curl,∇× centralforce field in circularcylindrical coordinates, 118 in curvilinearcoordinates, 112–13 in spherical polarcoordinates, 126 irrotational, 45 curved coordinates and vectors, 103–33, seealso circularcylinder coordinates; orthogonal coordinates; spherical polar coordinates differential operators, 110–14 curl,112–13 divergence, 111–12 exercises,113–14 gradient, 110 overview, 110 overview, 103 specialcoordinate systems, 114–33 curves, fitting to data, 1140–43 curvilinear coordinates, 104, 105, 110, 111, 112 cut line (branch cut),409 cylinder coordinates, circular, seecircularcylinder coordinates 1158 Index cylindrical coordinates, 104 cylindrical symmetry, 617 cylindrical traveling waves,708 D d’Alembertian, 141 d’Alembert ratio test,326–27 damped oscillator, 979–80 damped simple harmonic oscillation, 980 decay,kaon, 282–83 definite integral (Euler), 500–501 definite integrals evaluation of, 463 exponential forms, 471–82 Bernoulli numbers, 473–74 factorial function, 472–73integraltext∞ −∞f(x)dx, 465–66integraltext∞ −∞f(x)eiaxdx, 466–71 quantum mechanicalscattering, 469–71 simplepoleoncontourofintegration,468–69integraltext2π 0f(sinθ,cosθ)dθ, 464–65 degeneracy, of Schrödinger’s waveequation, 638 degenerateeigenfunctions, 638 degenerateeigenvalues, 223, 638 Del(∇), 42, 43 forcentral force, 127 successiveapplications of, 49–54 electromagnetic waveequation, 51–53 exercises, 53–54 Laplacianofpotential, 50–51 overview, 49–50 deltafunction, Dirac,83–85, 669–70, 975 Besselrepresentation, 935 in circularcylindrical coordinates, 601 derivation, 937–38 eigenfunction expansion, 89, 650 exercises,91–95 Fourier integral, 90 Fourier representation, 90 Green’sfunction and, 592–610 impulse force, 975 integral representations for, 90 Laplacetransform, 975 overview, 83–87 phasespace,88 point source, 88, 593 quantum theory, 955–61 representation by orthogonal functions, 88–89 sequences,83, 86 sine,cosine representations, 943 in spherical polarcoordinates, 82, 599 theory of distributions, 86 totalcharge inside sphere, 88 DeMoivre’s formula, 408–13denominator, integral of cos in, 464–65 dependent and independent variables,1038–44 alternateforms of Euler equations, 1042 conceptof variation, 1038–41 missing dependent variables, 1042–43 optical pathnearevent horizon of ablackhole, 1041–42 derivative operators, tensor curl, 162–63 divergence, 160–61 exercises, 162–63 Laplacian,161–62 overview, 160 derivatives, 426–27, seealsoexterior derivative covariant, 156 gauge covariant, 76 descending power series solutions, 369, 781 descent,steepest, seesaddle points (steepest descentmethod) determinants, 165–239 antisymmetry, 168–72 Gausselimination, 170–72 Gauss–Jordan elimination, inversion, 185–86 overview, 168–70 exercises, 174–76 Gram–Schmidt procedure, 173–74 overview, 173–74 vectors by orthogonalization, 174 homogeneous linear equations, 165–66 inhomogeneous linearequations, 166–67 Laplaciandevelopment by minors, 167–68 lineardependenceof vectors, 172–73 overview, 165 product theorem, 181 representation ofavector product, 20 secularequation, 218 solution of aset of homogenous equations, 165 solution of a set of nonhomogenous equations, 166 deuteron, 626–27 diagonal matrices,182–83, 215–31, seealso anti-Hermitian matrices eigenvectors andeigenvalues, 216–19 exercises, 226–31 functions of, 224–26 Hermitian, 219–21 moment of inertia, 215–16 differential equations, 535–619, 751–52 first-order differential equations, 543–53 exactdifferential equations, 545–47 exercises,550–53 linearfirst-order ODEs,547–50 nonlinear, 1088–1102 parachutist, 544–49 Index 1159 RL circuit, 549–50 separable variables, 544–45 Fuchs’ theorem, 573 heat flow, or diffusion, PDE,611–18 alternatesolutions, 614–15 special boundary condition again,615–16 specific boundary condition, 612–13 spherically symmetric heatflow, 616–17 homogeneous, 536, 548–50, 565 linearindependence of solutions second solution, 581–83 series form of the second solution, 583–85 nonhomogeneous equation—Green’s function, 592–610 circularcylindrical coordinate expansion, 27, 601–2 exercises, 607–10 form of Green’s functions, 596–98 Legendre polynomial addition theorem, 600–601 quantum mechanicalscattering—Green’s function, 603–6 quantum mechanicalscattering—Neumann series solution, 602–3 spherical polar coordinate expansion, 598–600 symmetry of Green’sfunction, 595–96 partial differential equations, 535–43 boundary conditions, 542–43 classes ofPDEsand characteristics,538–41 examples of, 536–38 introduction, 535–36 nonlinear PDEs,541–42 particular solution, 548, 565 second solution, 573–78 exercises, 588–92 lineardependence,580–81 linearindependence, 580 linear independence of solutions, 579–80 second solution for the linear oscillator equation, 583 second solution of Bessel’s equation, 586–87 second solution logarithmic term, 585, 700 separation of variables,554–62 Cartesian coordinates, 554–55 circularcylindrical coordinates, 555–56 exercises, 560–62 spherical polar coordinates, 557–60 series solutions—Frobenius’ method, 565–78 exercises, 574–78 expansion about x0,569 Fuchs’ theorem, 573 limitations ofseries approach—Bessel’s equation, 570–71 regular and irregular singularities, 572–73symmetry of solutions, 569 singular points, 562–65 differential forms, 304–19, seealsopullbacks 1-forms, 304–5 2-forms, 305–6 3-forms, 306–7 exercises, 318–20 exterior derivative, 307–9 Hodge operator *, 314–19 cross product of vectors, 315 Laplacianin orthogonal coordinates, 316 Maxwell’sequations, 316–20 overview, 314–15 Legendre’s equation, 752 overview, 304 Stokes’ theorem on, 313–14 differentials, exact, seethermodynamics differential vector operators, 110–14 adjoint, 621, 623, 634–35 curl, 112–13 del,33 divergence, 111–12 exercises, 113–14 gradient, 110 overview, 110 differentiation, 904–5 differentiation ofpowerseries, 364 diffraction, 680–82, 1017 diffusion equation, seedifferential equations, heat flow PDE digamma and polygamma functions, 510–16 digamma functions, 510–11 Maclaurinexpansion, computation, 512 polygamma function, 511–12 series summation, 512 dihedral groups, Dn, 299 dimension box-counting, 1086 Hausdorff, 1086 Kolmogorov, 1086 dimensionality theorem, 298 dipoles, interactionenergy, magnetic dipole, radiation fields, seeelectricdipole Diracbra-ket notation, 177 Diracdeltafunction, seedeltafunction, Dirac Diracmatrices, 209–12 direction cosines (orthogonal matrices), 4, 196–97, 201 direct product, 139–41 exercises, 140–41 and matrix multiplication, 181–82 overview, 139–40 of tensors, 140 direct tensor, 181 direct tensor product, 139 1160 Index Dirichletconditions, 543, 882 kernel,910 Dirichletproblem, 617 Dirichletseries, 326 discontinuities, behavior of, 886 discontinuous functions, 888 discreteFourier transform, 914–19 aliasing,916–17 fastFourier transform, 917 limitations, 916 orthogonality over discretepoints, 914–15 discrete groups, 291–304 classesandcharacter,293 crystallographic point and spacegroups, 299–300 dihedral groups, Dn,299 exercises,300–304 subgroups and cosets,293–94 threefold symmetry axis, 296–99 twofold symmetry axis, 294–96 discreterandom variable,1116–17 dispersion relations, 482–89 causality,486–87 crossing relations, 484 exercises,487–89 Hilbert transform, 485 opticaldispersion, 484–85 overview, 482–83 Parseval relation, 485–86 sum rules, 484 symmetry relations, 484 displacement, 158 dissipation, 1101–3 distributive, 13 divergence,∇, 38–43 of central force field,39–40 circularcylindrical, 115–23 coordinates, Cartesian, 4–7 of coordinate vector, 39 curvilinear coordinates, 111 asdifferential vectoroperator, 111–12 exercises,42–43 integral definitions of gradient, curl, and,58–59 integration by parts of, 40 overview, 38–39 physical interpretation, 40–42 solenoidal, 42 spherical polar, 123–33 astensor derivative operator, 160–61 divergent series, 325–26 Doppler shift, 360 dot products, 12–17 exercises,17 gradient of, 46invariance ofscalarproduct under rotations, 15–17 overview, 12–15 double series,rearrangement of, 345–48 driven oscillator with damping, 991–93 dual tensors, 147–48 duplication formula for factorialfunctions, see Legendre duplication formula dynamical systems, dissipation in,1102–3 E E, Lorentztransformation of the electricfield, 287–88 Earth’s gravitational field, 758–59 Earth’s nutation, 973–74 eigenfunctions, 624, 626–27 Bessel’s inequality, 651–52 completeness of, 649–61 of Fourier series: of Sturm–Liouville eigenfunctions, 649–51 of Hilbert–Schmidt: of integral equations, 1031–33 eigenvalue equation, 667–68 expansion, Green’s function, 662–74 of Diracdelta function, 88–90 of Hermitian differential operators, 635, 649–51 of square wave,637 expansion coefficients, 658 orthogonal, 636, 637, 1030–32 Schwarzinequality, 652–54 summary—vector spaces,completeness, 654–58 variational calculation,1072–76 eigenvalues, 217–18, 223, 624–27, 634–35 of Hermitial differential operator, 634 of Hermitial matrices, 219–23 of Hilbert–Schmidt integral equation, 1030–36 of normal matrices, 231–32 of real symmetric matrices,215–19 variational principle for, 1072 eigenvectors, 216–19, 221–22 eightfold way(weight diagram), 257 Einstein’s energy relation, 281 Einstein’s summation convention, 136 velocity addition law,283, 290 electricalcharge inside spheres, 88 electricdipole potential, Legendre polynomial expansion, 745 electromagnetic invariants, 288 electromagnetic waveequation, 51–53 electromagnetic waves,981–82 electrostaticmultipole expansion, 599 electrostaticpotential, 593 in hollow cylinder, 695–96 Index 1161 of ring of charge,761–62 electrostatics,physical basis, 741–42 elimination, seedeterminants ellipticaldrum, 872–73 ellipticintegrals, 370–76 definitions of, 372 exercises, 374–76 of first kind, 372 hypergeometric representations, 373, 860 limiting values, 374 overview, 370 period of simple pendulum, 370–71 of second kind, 372 series expansion, 372–73 ellipticPDEs,538 empty set, 1111 energy potentials, 309 relativistic, 356–57 equality of matrices,178 equations, seealsolinearequations; Maxwell’s equations; Poisson’s equation of motion and field,142 error integrals, 530 asymptotic expansion, 393 confluent hypergeometric representation, 864 error propagation, 1138–40 essential(irregular) singular point, 563 Eulerangles, 202–3 Eulerequation, 1040 alternate forms of, 1042 applications of, 1044–52 soap film, 1045–46 soap film—minimum area,1046–49 straight line, 1044–45 Euleridentity, 224 product formula, 382–84 Eulerintegrals, 502 Euler–Maclaurin integration formula, 380–82, 517 Euler–Mascheroni constant, 330 Eulerproduct for Riemann Zetafunction, 382–84 event horizon, 1041 exactdifferential equations, 545–47 exactdifferentials, seethermodynamics expansion, seealsoLaurent expansion; Taylor’s expansion pole, of meromorphic functions, 461 product, of entire function, 462–63 of series, 372–73 expansion coefficients, 658 expansion of functions, Legendre series, 757–58 expectation value,630, 955–56, 1117–18 exponential forms, 471–82 Bernoulli numbers, 473–74 factorial function, 472–73exponential function, of Maclaurintheorem, 354–55 exponential integral, 527–30 exponential of diagonal matrix, 225–26 exponential transform, 938 exterior derivative, 307–9 extreme or stationary value,36, 881, 1038–40 F Factorial function Ŵ(1+s) asymptotic form of, 494–95 complex argument, 499–506 contour integrals, 505 digamma function, 510 double factorialnotation, 505 Gammafunctional relation, 503 infinite product, 501 integral (Euler) representation, 500 Legendre duplication formula, 503 Maclaurinexpansion, 512 polygamma functions, 511 steepestdescent asymptotic formula, 495 Stirling’s (formula) series, 516–18 factorial notation, 503–5 faithful group, 243 Faraday’s law,66–68 fast Fourier transform, 917 Feigenbaum number, 1083–84 Fermage equation, 950 Fermat’s principle, 157, 1041, 1050 Fermi–Dirac statistics, 1064, 1115 field equations, 142 finite wave train, 940–41 first-order differential equations, 543–53 exactdifferential equations, 545–47 linearfirst-order ODEs,547–50 separablevariables, 544–45 fixed andmovable singularities, specialsolutions, 1090 Floquet’s theorem, 877 force as gradient of potentials, 36 forced classicaloscillators, 457–60 force field, central, seecentral force field Fourier–Bessel series, 695 Fourier expansions of Mathieufunctions, 919–29 integral equations and Fourier series for Mathieu functions, 919–23 leadingcoefficients for ce 0,926–29 leadingcoefficients ofse 1,923–26 Fourier integral theorem, 937 development of, 936–38 exponential form, 937 Fourier–Mellin integral, 995 Fourier representation, of Diracdeltafunction, 90 Fourier series,881–930 1162 Index advantages,uses of, 888–92 change of interval, 890–91 completeness, seegeneral properties of Fourier series convergence, 893, 903–10 differentiation, seegeneral properties of Fourier series discontinuous functions, 888 exercises, 891–92 integration, seegeneral properties ofFourier series periodic functions, 888–90 applications of, 892–903 exercises, 898–903 full-wave rectifier, 893–94 infinite series,Riemann Zetafunction, 894–98 square wave—high frequencies, 892–93 discreteFourier transform, 914–19 discrete fourier transform, 915–16 discreteFouriertransform—aliasing,916–17 exercises, 918–19 fast Fourier transform, 917 limitations, 916 orthogonality over discrete points, 914–15 Fourier expansions of Mathieu functions, 919–29 exercises, 929 integral equations and Fourier seriesfor Mathieu functions, 919–23 leading coefficientsfor ce 0, 926–29 leading coefficientsof se 1,923–26 generalproperties, 881–88 behavior of discontinuities, 886 completeness, 883–84 complex variables—Abel’s theorem, 882 exercises, 886–88 sawtooth wave,885–86 square wave,892 Sturm–Liouville theory, 885 summation of aFourier series,882–83 Gibbs phenomenon, 910–14 calculation of overshoot, 912–13 exercises, 913–14 square wave,911–12 summation of series,910 orthogonality, 636–37 properties of, 903–10 convergence, 903 differentiation, 904–5 exercises, 905–10 integration, 904 summation of, 882–83 Fourier transform, of Gaussian,932 Fourier transform of derivatives, 946–51heatflow PDE,948–49 inversion of PDE,949–50 wave equation, 947–48 Fourier transforms, 486–87, 931–32 aliasing, 916 convolution (Faltungs) theorem, 951 deltafunction derivation, 937 transfer functions, 961 Fourier transforms—inversion theorem, 938–46 cosine transform, 939 exponential transform, 938 fastFourier transform, 917 finite wave train, 940–41 Fourier integral, 931, 936–39 momentum spacerepresentation, 955 sine transform, 939–40 uncertainty principle, 941 Fourier transform solution, 1013 fractals,1086–88 fractional order, 427 Fraunhofer diffraction, Bessel function, 680–82 Fredholm equation, 1005–6 Frobenius’ method, series solutions, 565–78 Fuchs’ theorem, 573 full-wave rectifier, 893–94 functional equation, Riemann Zetafunction, 896 Gamma function, 500 functions, 817–80, seealsoanalytic functions Chebyshev polynomials, 848–59 exercises,855–59 generating functions, 848 orthogonality, 854–55 recurrencerelations—derivatives, 852–53 trigonometric form, 853–54 type I, 849–52 type II, 849 of complex variable, 408–13 confluent hypergeometric functions, 863–69 Besseland modified Besselfunctions, 865 exercises,867–69 Hermitefunctions, 866 integral representations, 865 miscellaneous cases,866 entire,415, 451, 462–63 exponential, of Maclaurin theorem, 354–55 factorial,472–73 Hermitefunctions, 817–36 alternaterepresentations, 819 applications of theproduct formulas, 831–32 directexpansion ofproducts of Hermite polynomials, 828–30 exercises,832–36 generating functions—Hermitepolynomials, 817–18 Index 1163 orthogonality, 821–22 quantum mechanicalsimple harmonic oscillator, 822–27 recurrence relations, 818–19 Rodrigues’ representation, 820–21 threefold Hermite formula, 827–28 hypergeometric functions, 859–63 contiguous function relations, 861 exercises, 862–63 hypergeometric representations, 861–62 Laguerre functions, 837–48 associated Laguerrepolynomials, 841–43 differential equation—Laguerre polynomials, 837–41 exercises, 845–48 hydrogen atom, 843–45 Mathieu functions, 869–79 elliptical drum, 872–73 exercises, 879 general properties of Mathieufunctions, 874 quantum pendulum, 873 radial Mathieu functions, 874–79 separation ofvariables in elliptical coordinates, 870–71 of matrices, 224–26 meromorphic, 451, 461–62, 478 multivalent, and branch points, 447–50 rotation of, 251 series of, 348–52 Abel’s test,351 exercises, 352 overview, 348 uniform and nonuniform convergence, 348–49 Weierstrass Mtest,349–50 G gamma distribution, 1143 gamma function, seealsofactorialfunction, 499–533 beta function, 520–26 definite integrals, alternateforms, 521–22 derivation of Legendre duplication formula, 522–23 exercises, 523–26 incomplete betafunction, 523 verification of πα/sinπαrelation, 522 definitions, simple properties, 499–510 definite integral (Euler), 500–501 double factorial notation, 505 exercises, 506–10 factorial notation, 503–5 infinite limit (Euler), 499–500 infinite product (Weierstrass), 501–3 integral representation, 505–6digamma and polygamma functions, 510–16 Catalan’sconstant, 513 digamma functions, 510–11 exercises,513–16 Maclaurinexpansion, computation, 512 polygamma function, 511–12 seriessummation, 512 incomplete gamma functions andrelated functions, 527–33 error integrals, 530 exercises,530–33 exponential integral, 527–30 of infinite product, 398–99 Stirling’s series,516–20 derivation from Euler–Maclaurin integration formula, 517 exercises,518–20 Stirling’s series, 518 gauge covariant derivative, 76 theory, 259 transformation, 76 gauge theory, 76, 241 Gauss elimination, 170–72 Gauss’ differential equation, 312–13, 318, 614–15, 617–18 hypergeometric differential equation, 576, 859–62, 873 Gauss error integral, 500 asymptotic expansion, 530 Gauss’ fundamental theorem of algebra,428, 463 Gauss–Jordan matrix inversion, 185–87 Gauss’ law, 52, 79–83, 594 Gauss’ normal distribution, 1134–38 Gauss’ notation, 503 Gauss–Seidel iteration technique, 172 Gauss–Seidel method, 226 Gauss’ test, 332–33, 357 Legendre series,333 Gauss’ theorem, 60–64 alternateforms of, 62–64 exercises,62–64 overview, 62 overview, 60–61 pullbacks, 312–13 Gegenbauer polynomials, seeultraspherical polynomials general parabolic solution, 540 general properties, 881–88 behavior of discontinuities, 886 completeness, 883–84 complex variables—Abel’s theorem, 882 of Mathieu functions, 874 sawtooth wave,885–86 Sturm–Liouville theory, 885 summation of aFourier series, 882–83 1164 Index general tensors, 151–60, seealsoChristoffel symbols covariant derivative, 156 exercises,158–60 geodesics and paralleltransport, 157–60 metrictensor, 151–54 overview, 151 generating function, 741–49, 848, 1014–15 associatedLaguerre polynomials, 624–25, 841–45 associatedLegendre functions, 771–82, 788 associatedLegendre polynomials, 773 Bernoulli numbers, 376–89, 473–74, 517 Bernoulli polynomials, 379–80 Besselfunctions, modified, 711, 713–19, 723–24, 865 Chebyshev polynomials, 848–59 extension to ultraspherical polynomials, 747 Hermitepolynomials, 817–30 for integral order, 675–77 Laguerre polynomials, 647, 837–41, 843–45 Legendre polynomials, 742–44 linearelectricmultipoles, 744–45 physical basis—electrostatics,741–42 ultraspherical polynomials, 747 vector expansion, 745–47 generators of continuous groups, 246–61, seealso rotations;SU(2) exercises,260–61 overview, 246–50 geodesic equation, 157 geodesics, 157–60 geometrical interpretation of gradient, 35–38 integration by parts of, 36–37 of potential, force as, 36 geometric series, 322–23 Gibbs phenomenon, 886, 910–14 calculationof overshoot, 912–13 square wave,911–12 summation of series,910 global behavior, 1093–1101 Goursat proof of Cauchy’s integral theorem, 421–23 gradient in Cartesian coordinates, 37, 51, 113, 134 in circularcylindrical coordinates, 118 in curvilinearcoordinates, 110 in spherical polarcoordinates, 126 gradient, curvilinear coordinates, 110 gradient,∇,32–38 asdifferential vectoroperator, 110 of dot product, 46 exercises,37–38 geometrical interpretation, 35–38 integration by parts of, 36–37of potential, force as, 36 integral definitions of divergence, curl, and, 58–59 overview, 32–34 of potential, 34 Gram–Schmidt orthogonalization, 642–49 Gram–Schmidt procedure, 173–74 overview, 173–74 vectors by orthogonalization, 174 gravitational potentials, 72 great circle,1040 Green’s function, 662–74 construction of, one dimension, 598, 599, 663–65, 670 construction of, twodimension, 597–98 construction of, three dimension, 597–98 and Diracdeltafunction, 669–70 eigenfunction, eigenvalue equation, 667–68 eigenfunction expansion, 662–82 electrostaticanalog, 592, 665 form of, 596–98 Helmholtz, 598 Helmholtz equation, 662 integral—differential equation, 665–67 Laplaceoperator, 598 circularcylindrical expansion, 601–6 spherical polar expansion, 598–600 linear oscillator, 668–69 modified Helmholtz, 598 nonhomogeneous equation, 592–610 one-dimensional, 663–65 Poisson’s equation, 669–70 symmetry of, 595–96 Green’s theorem, 61–62, 593 Gregory series, 362, 368 ground stateeigenfunction, 1073 group theory, 241–320, seealsoangular momentum; differential forms; generators of continuous groups; homogeneous Lorentz group character,184 definition of, 242–43 discrete,291–93 classesandcharacter,293 crystallographic point and space,299–300 dihedral, Dn,299 exercises,300–304 subgroups and cosets,293–94 threefold symmetry axis, 296–99 twofold symmetry axis, 294–96 faithfulness, 243 homomorphic, 243 homomorphism and isomorphism, 243–45 homomorphism SU(2)–SO(3),252–6 isomorphic, 243 Index 1165 Lorentz covariance ofMaxwell’s equations, 283–91 electromagnetic invariants, 288 exercises, 289–91 overview, 283–86 transformation of EandB,287–88 overview, 241–42 permutation groups, 301–3 quantum chromodynamics (QCD), 258 reducible and irreducible representations, 245–46 special unitary group SU(2), 251 vierergruppe, 292–93, 296 Gutzwiller’strace formula, 898 H Hadamard product, 208 Hamilton–Jacobi equation, 539 Hamilton’s principle, 1053–54 Hankelfunctions and Lagrangeequations of motion, 493–94, 707–13 asymptotic forms, 493–94, 723 contour integral representation ofthe Hankel functions, 709–11 cylindrical traveling waves,708 definition, by Neumann function, 707–8 definitions, 707–8 series expansion, 707, 715 spherical, 728–29 Wronskian formulas, 708 Hankeltransforms, 933 harmonic functions, 539 harmonic oscillator, 822–27, 958–59 harmonics, seealsospherical harmonics, vector spherical harmonics harmonic series,323–24 Hausdorff, 225, 249, 1086 heatflow PDE,948–49 Heavisideexpansion theorem, 442, 1003 Heavisideshifting theorem, 981 Heavisideunit step function, 93, 996 Heisenberg uncertainty principle, 732, 941 Helmholtz diffusion equation, 536, 537 Helmholtz equation, 556, 557, 613 Bessel function, 683–84, 725 Green’s function, 662 spherical coordinates, 725 Helmholtz operators, 598 Helmholtz’s theorem, 95–101 exercises, 100–101 overview, 95–96 Hermitefunctions, 817–36, 866 alternate representations, 819 applications of theproduct formulas, 831–32 confluent hypergeometric representation, 866direct expansion of products of Hermite polynomials, 828–30 generating functions—Hermite polynomials, 817–18 Gram–Schmidt construction, 642 orthogonality, 821–22 quantum mechanicalsimple harmonic oscillator, 822–27 recurrence relations, 818–19 Rodrigues representation, 820–21 threefold Hermite formula, 827–28 Hermite polynomials direct expansion of products of, 828–30 generating functions, 817–18 orthogonality integral, 821 recurrence relations, 818–19 Rodrigues representation, 820–21 Hermitian matrices, 184, 209 and matrix diagonalization, 219–21 unitary and, 208–15 exercises,212–15 overview, 208–9 Pauli and Dirac,209–12 Hermitian matrices, anti-, 221–23 eigenvalues degenerate,223 andeigenvectors ofreal symmetric matrices, 221–22 overview, 221 Hermitian operators, 629–30, 634–42 completeness of eigenfunctions, 649–58 degeneracy, 638 expansion in orthogonal eigenfunctions—square wave,637 Fourier series—orthogonality, 636–37 integration interval, 628–29 orthogonal eigenfunctions, 636 properties of, 634–38 in quantum mechanics, 630 realeigenvalues, 634–35 Hilbert matrix, determinant, 235 Hilbert–Schmidt theory, 1029–36 nonhomogeneous integral equation, 1032–34 orthogonal eigenfunctions, 1030–32 symmetrization of kernels, 1029–30 Hilbert space,7, 535, 629, 638, 658, 885 Hilbert transforms, 483 Hodge * operator, 314–19 cross product of vectors, 315 Laplacianin orthogonal coordinates, 316 Maxwell’sequations, 316–20 overview, 314–15 holomorphic functions (analytic or regular functions), 415 homogeneous equations, 165–66, 565 1166 Index homogeneous Lorentz group, 278–83 exercises,283 kinematics and dynamics in Minkowski space–time, 280–83 overview, 278–80 homomorphic group, 243 homomorphism, 243–45 overview, 243 rotations, 244–45 SU(2)andSO(3), 252–56 hooks, 275–76 Hubble’s law, 7 hydrogen atom, 843–45, 957–58 Schrödinger’s wave equation, 843 hyperbolic PDEs,538 hypercharge, 257 hypergeometric equation alternateforms, 860, 864 second independent solution, 860 singularities, 564, 859, 864, 865, 873 hypergeometric functions, 859–63 contiguous function relations, 861 hypergeometric representations, 861–62 I ill-conditioned matrices,234–35 imaginary part, 407 impulsive force, 975–76 incomplete gamma function, 389–92 confluent hypergeometric representation, 500, 864 recurrencerelations, 399, 512 independence, linear,173, 579–81, 643, 665, 703 of solutions of ordinary differential equations, 579–81 of vectors, 173, 579 indicial equation, 566 inertia matrix, moment of, 215–16 infinite limit (Euler), 499–500 infinite product (Weierstrass), 499–500, 501–3 infinite products, 378, 383, 396–99, 499, 501–3 convergence, 397–98 cosine,398–99 entirefunctions, 462 gamma function, 398–99 sine,398–99 infinite series, 321–401, seealsoalternating series; powerseries; Taylor’s expansion algebra of, 342–48 alternating series,342–43 convergence, 342–45 convergence: absolute, 342–44 convergence: Cauchy integral, 327–29 convergence: Cauchy root, 326 convergence: comparison, 325–26convergence: conditional, Leibniz criterion, 344 convergence: D’Alembert ratio, 326–27 convergence: Gauss’, 332–33 convergence: improvement of, 345 convergence: Kummer’s, 330–32 convergence: Maclaurinintegral, 327–30 convergence: Raabe’s, 332 convergence: tests of, 325–35 convergence: uniform, 348–51, 363–64 divergence of squares, 344 double series, 345–47 exercises,347–48 overview, 342–44 rearrangement of double, 345–48 asymptotic series,389–96 cosine and sine integrals, 392–93 definition of, 393–94 exercises,394–96 incomplete gamma function, 389–92 overview, 389 Bernoulli numbers, 376–89 Euler–Maclaurin integration formula, 380–82 exercises,385–89 improvement of convergence, 385 overview, 376–79 polynomials, 379–80 Riemann Zetafunction, 382–84 ellipticintegrals, 370–76 definitions of, 372 exercises,374–76 limiting values, 374 overview, 370 period of simple pendulum, 370–71 seriesexpansion, 372–73 of functions, 348–52 Abel’stest, 351 exercises,352 overview, 348 uniform and nonuniform convergence, 348–49 Weierstrass Mtest, 349–50 fundamental concepts, 321–25 addition and subtraction of, 324–25 exercises,325 geometric, 322–23 harmonic, 323–24 overview, 321–22 power series,363–66, 578–79 products of, 396–401 convergence of, 397–98 exercises,399–401 overview, 396–97 sine,cosine, and gamma functions, 398–99 Riemann’s theorem, 894–98 Index 1167 infinity,seesingularity, pole, essentialsingularity inhomogeneous linearequations, 166–67 inhomogeneous ordinary differential equation (ODE),Green’s function solutions, 663–64 inner product and matrix multiplication, 179–81 integral definitions of gradient, divergence, and curl, 58–59 integral—differential equation, 665–67 integral equations, 1005–36 and Fourier series for Mathieu functions, 919–23 Fredholm equations, 1005, 1007, 1010, 1013, 1018, 1021–24, 1030–32 Hilbert–Schmidt theory, 1029–36 exercises, 1034–36 nonhomogeneous integral equation, 1032–34 orthogonal eigenfunctions, 1030–32 symmetrization of kernels, 1029–30 integral transforms, generating functions, 1012–18 exercises, 1015–18 Fourier transform solution, 1013 generalizedAbelequation, convolution theorem, 1014 generating functions, 1014–15 introduction, 1005–12 definitions, 1005–6 exercises, 1011 linear oscillator equation, 1009–11 momentum representation in quantum mechanics,1006–7 transformation of adifferential equation into an integral equation, 1008–9 Neumann series,separable (degenerate) kernels, 1018–29 exercises, 1025–29 Neumann series,1018–19 Neumann seriessolution, 1020–21 numerical solution, 1023–25 separable kernel,1021–22 Volterra equations, 991, 1005–6, 1009–11, 1021 integral form, Neumann functions, 701 integral representations, 505–6, 679–80 for Dirac deltafunction, 90 expansion of, 720–23 integrals, see alsoCauchy(Maclaurin) integral test;definite integrals; ellipticintegrals contour integration, 463, 471, 503, 522, 603 differentiation of, 590 evaluation of, 810 Lebesgue, 649, 657 line, 55–56, 65–67, 440 of products of three spherical harmonics, 803–6 Riemann, 55, 60–61, 605, 636 Stieltjes, 86, 873surface,56–57 volume, 57–58 integral test,Cauchy, seeCauchy(Maclaurin) integral test integral transforms, 931–1004 convolution (Faltungs) theorem, 990–94 driven oscillator withdamping, 991–93 exercises,993 convolution theorem, 951–55 exercises,953–55 Parseval’s relation, 952–53 development of theFourier integral, 936–38 Diracdeltafunction derivation, 937–38 Fourier integral—exponential form, 937 Fourier transform, 931–32 Fourier transform of derivatives, 946–51 heatflow PDE,948–49 inversion of PDE,949–50 wave equation, 947–48 Fourier transform of Gaussian, 932 Fourier transforms—inversion theorem, 938–46 cosinetransform, 939 exercises,942–46 exponential transform, 938 finite wave train, 940–41 sine transform, 939–40 uncertainty principle, 941 Fourier transforms of derivatives, 950–51 generating functions, 1012–18 Fourier transform solution, 1013 generalizedAbelequation, convolution theorem, 1014 generating functions, 1014–15 integral transforms, 931–35 exercises,934–35 Fourier transform, 931–32 Fourier transform of Gaussian,932 Laplace,Mellin, andHankeltransforms, 933 linearity, 933–34 inverse Laplacetransform, 994–1003 Bromwich integral, 994–95 exercises,1000–1003 inversion via calculusof residues, 996 summary—inversion of Laplacetransform, 999–1000 velocityof electromagneticwaves in a dispersive medium, 997–99 Laplace,Mellin,andHankeltransforms, 933 Laplacetransform of derivatives, 971–78 Diracdeltafunction, 975 Earth’s nutation, 973–74 exercises,976–78 impulsive force, 975–76 simple harmonic oscillator, 973 1168 Index Laplacetransforms, 965–71 definition, 965 elementary functions, 965–66 exercises, 970–71 inverse transform, 967–68 partial fraction expansion, 968–69 step function, 969–70 linearity, 933–34 momentum representation, 955–61 exercises, 959–61 harmonic oscillator, 958–59 hydrogen atom, 957–58 other properties, 979–89 Bessel’s equation, 983–84 damped oscillator, 979–80 derivative of a transform, 982–83 electromagnetic waves,981–82 exercises, 985–89 integration of transforms, 984 limits of integration—unit step function, 985 RLCanalog, 980–81 substitution, 979 translation, 981 transfer functions, 961–64 exercises, 964 significance of /Phi1(t),963–64 integration, 904, seealsopath-dependent work Euler–Maclaurin formula, 380–82 by parts of curl, 47 byparts of divergence, 40 by parts of gradient, 36 of powerseries, 364 simple pole on contour of, 468–69 of transforms, 984 of vectors, 54–60 exercises, 59–60 overview, 54–56 integration interval [a,b], 628–29 interpolating polynomials, 194 interpretation, seegeometrical interpretation of gradient; physical interpretation of divergence interitem, 1111 invarianceofscalarproductunderrotations,15–17 invariants, electromagnetic,288 inverse Laplacetransform, 994–1003 Bromwich integral, 994–95 inversion via calculusof residues, 996 summary—inversion of Laplacetransform, 999–1000 velocityof electromagneticwaves in a dispersive medium, 997–99 inverse matrix, 200 inverse operator, 934, 1019 inverse transform, 967–68inversion, 445–46 matrix, 184–87 Gauss–Jordan, 185–87 overview, 184–85 of PDE,949–50 of powerseries, 366 via calculusof residues, 996 irreducible representations, 245–46 irreducible tensors, 149–51 irregular (essential) singular point, 563 irregular sign changes, series with,341–42 irregular singularities, 572–73 irrotational, 45 isomorphic group, 243 isomorphism, 243–45 overview, 243 rotations, 244–45 isospin,SU(2),256–60 isospinI, 257 J Jacobian, 107–8 parity transformation, 146 Jacobians for polar coordinates, 108–10 Jacobi–Anger expansion, 687 Jacobi identity, 248 Jacobi technique, 226 Jordan’s lemma, 468 Julia set,1087 K kaon decay, 282–83 Kepler’s lawsof planetary motion, 116–17 kinematics and dynamics in Minkowski space–time,280–83 Kirchhoff diffraction theory, 426 Klein–Gordon equation, 537 Korteweg–deVries equation, 542 Kroneckerdelta,10, 136–37 Kroneckerproduct, 181 Kronig–Kramers optical dispersion relations, 484, 485 Kummer’s equation, seeconfluent hypergeometric equation Kummer’s test, 330–32 L ladderoperators, approach to orbital angular momentum, 262–64 Lagrange’s equations, 1066 Lagrangian, 1053–54 Lagrangian equations, 1066–67 Lagrangian multipliers, 1060–65 cylindrical nuclearreactor, 1062 particlein a box, 1061–62 Index 1169 Laguerrefunctions, 837–48 associated Laguerre polynomials, 841–43 differential equation—Laguerrepolynomials, 837–41 hydrogen atom, 843–45 Laguerrepolynomials, 624 associated, 624–25 confluent hypergeometric representation, 866 generating function, 841–42 integral representation, 843 orthogonality, 843 recurrence relations, 842 Rodrigues’ representation, 767, 842 confluent hypergeometric representation, 866 differential equation, 837–41 generating function, 837–39 Gram–Schmidt construction, 647 orthogonality, 624, 647 recurrence relations, 840, 842 Rodrigues’ formula, 839 Schrödinger’s wave equation, 843 self-adjoint form, 624, 840 singularities, 564 Laplace,Mellin, andHankeltransforms, 933 Laplacefunction, 598 Laplace’sequation, 536, 1057–58 Bessel functions, 695 Legendre polynomials, 760, 761 solutions, 50–51, 96, 443, 452, 539, 559–60 Laplaceseries expansion theorem, 790–91 gravity fields, 791 Laplacetransforms, 965–71 convolution theorem, 521, 1000, 1003, 1014 definition, 965 of derivatives, 971–78 Diracdelta function, 975 Earth’s nutation, 973–74 impulsive force, 975–76 simple harmonic oscillator, 973 elementary functions, 965–66 inverse transform, 967–68 partial fraction expansion, 968–69 step function, 969–70 table of transforms, 967–68, 979 translation, 1000 Laplacian in Cartesian coordinates, 51–52, 554–55 in circularcylindrical coordinates, 119 development by minors, 167–68 in orthogonal coordinates, 316 of potentials, 50–51 spherical polarcoordinates, 126 as tensor derivative operator, 161–62 Laurentexpansion, 430–38, 466, 472analyticcontinuation, 432–34 exercises, 437–38 Schwarzreflection principle, 431–32 Taylorexpansion, 430–31 lawof cosines, 16–17, 118, 745 leadingcoefficients for ce 0,926–29 leadingcoefficients of se 1, 923–26 leastsquares, method of, 1119 Legendre duplication formula, derivation of, 522–23 Legendre equation, Maxwell’sequation, 779 self-adjoint form, 623, 625 Legendre functions, 741–816 addition theorem, 798 addition theorem for spherical harmonics, 797–802 derivation of addition theorem, 798–800 exercises,800–802 trigonometric identity, 797–98 alternatedefinitions of Legendre polynomials, 767–70 exercises,769–70 Rodrigues’ formula, 767 Schlaefliintegral, 768–69 associated,772 associatedLegendre functions, 771–86 associatedLegendre polynomials, 772–74 equation, 558, 771–72, 778–82, 788 Fourier transform, 770 Gram–Schmidt construction, 789 hypergeometric representation, 861 lowestassociatedLegendre polynomials, 774 magneticinduction field of a current loop, 778–82 orthogonality, 776–78 parity, 776 poles, 760, 782 recurrencerelations, 775 Rodrigues’ formula, 772–73 Schlaefliintegral, 768 second kind, 806–12 self-adjoint form, 771–72 specialvalues,774–75 generating function, 741–49 exercises,747–49 extension to ultraspherical polynomials, 747 Legendre polynomials, 742–44 linearelectricmultipoles, 744–45 physical basis—electrostatics,741–42 vector expansion, 745–47 integrals ofproducts of three spherical harmonics, 803–6 application ofrecurrence relations, 804–5 exercises,805–6 orbital angularmomentum operators, 793–97 1170 Index orbital angular momentum operators, 796–97 orthogonality, 756–67 Earth’s gravitational field, 758–59 electrostaticpotential ofa ring ofcharge, 761–62 exercises, 762–66 expansion of functions, Legendre series, 757–58 polarization of dielectric,764 ring of electriccharge,761–62 sphere in a uniform field, 759–61 recurrencerelations andspecialproperties, 749–56 differential equations, 751–52 exercises, 754–56 parity, 753 recurrence relations, 749–50 special values,752 sphere in uniform electricfield, 759–61 upper and lower bounds for Pn(cosθ),753–4 of second kind, 806–13 closed-form solutions, 810–12 exercises, 812–13 Qn(x)functions of the second kind, 809–10 series solutions of Legendre’s equation, 807–9 spherical harmonics, 786–93 azimuthal dependence—orthogonality, 787 exercises, 791–93 Laplaceseries, expansion theorem, 790–91 Laplaceseries—gravity fields, 791 polar angle dependence,788 spherical harmonics, 788–90 vector spherical harmonics, 813–16 Legendre polynomials, 644–46, 741, 742–44 associated generating function, 773 recurrence relations, 775 generating function, 743 by Gram–Schmidt orthogonalization, 644–46 Laplace’sequation, 760, 761 orthogonality integral, 777 recurrencerelations, 749–50 Rodrigues’ formula, 761 Schlaefliintegral, 768 Legendre’s duplication formula, 503 Legendre’s equation, 625 differential form, 752 Legendre’s equation self-adjoint form, 625, 771 Legendre’s equation seriessolutions of, 807–9 Legendre series,333, 807–9 recurrencerelations, 807–8 Leibnizcriterion, 339–40formula for differentiating an integral, 590, 776 formula for differentiating aproduct, 771 Lerch’s theorem, 967 Levi-Civita symbol, 146–47 L’Hôpital’s rule,365 Lie groups and algebras, 243, 248, 264–66 limits of integration—unit step function, 985 limits to values of elliptic integrals, 374 linearelectricmultipoles, 744–45 linearequations homogeneous, 165–66 inhomogeneous, 165–66 linear independence ofsolutions, 581–83 linearity, 933–34 linearly dependent solutions, 549 linearly independent solutions, 549 linear operator, 87, 176, 208, 535, 622, 650 differential operator, 42–43, 249, 261, 285, 304, 307, 554, 569, 629, 634, 664 integral operator, 768, 1019 linearoscillator Green’sfunction, 668–69 linear oscillator equation, 1009–11 linear transformation law,152 line integrals, 55 Liouville’s theorem, 428 liquid drop model, 785 logistic map,1080–84 Lommel integrals, 697 Lorentz covarianceof Maxwell’sequations, 283–91 electromagneticinvariants, 288 exercises, 289–91 overview, 283–86 transformation of EandB,287–88 Lorentz–Fitzgerald contraction, 148 Lorentz gauge,52 Lorentz group, seehomogeneous Lorentzgroup lowering operator, 263 lowest associatedLegendre polynomials, 774 Lyapunov exponents, 1085–86 M Maclaurinexpansion, series computation, 512 Maclaurinintegral test, 327–30 Riemann Zetafunction, 329–30 Maclaurintheorem, 354–55 exponential function, 354–55 logarithm, 355 overview, 354 Madelung constant, 347 magnetic, 19, 46, 51, 66–67, 69, 74–76, 96, 100, 127–28, 144–45, 283–84, 288, 306, 311, 317, 447, 537, 638, 685, 703–5, 746, 778–82, 974 Index 1171 magnetic field constant ( Bfield), 44 magnetic fluxacross an oriented surface,306 magnetic induction field of current loop, 778–82 magnetic moments, 145 magnetic vector potentials, 44, 74–76, 127–28 Mandelbrot set, 1087–88 mapping complex variables, 443–51 branch points and multivalent functions, 447–50 exercises, 450–51 inversion, 445–46 overview, 443 rotation, 444 translation, 443–44 conformal, 451–54 exercises, 453–54 Mathieuequation angular, 872 modified, 872 radial, 872 Mathieufunctions, 869–79, 921 elliptical drum, 872–73 Fourier expansions of, 919–29 integral equations and Fourier seriesfor Mathieu functions, 919–23 leading coefficientsfor ce 0, 926–29 leading coefficientsof se 1,923–26 general properties ofMathieu functions, 874 quantum pendulum, 873 radial Mathieu functions, 874–79 separation of variables inellipticalcoordinates, 870–71 matrices,165–239, seealsodeterminants; diagonal matrices; orthogonal matrices addition and subtraction, 178–79 adjoint, 208 angular momentum matrices, 253, 273 anticommuting sets,236 antihermitian, 221–23, 231 antisymmetric, 204 definition, 176 diagonalization, 215–26 direct product, 181–82 equality, 178 Euler angle rotation, 202–3 exercises, 187–95 Gauss–Jordan matrix inversion technique, 185 Hermitian and unitary, 208–15 exercises, 212–15 overview, 208–9 Pauli and Dirac,209–12 inversion of, 184–87 Gauss–Jordan, 185–87 overview, 184–85ladder operators, 263 matrix multiplication, 176, 179–81 moment of intertia, 215–16, 220 multiplication, 179–80 directproduct, 181–82 inner product, 179–81 byscalar,179 normal, 231–39 exercises,236–39 ill-conditioned systems, 234–35 normal modes of vibration, 233–34 overview, 231–32 null matrix, 178 orthogonal matrix, 201, 206, 209 overview, 176–78 product theorem, 181 quaternions, 204, 212 rank, 178 relation to tensor, 206 representation, 177, 184, 205, 208, 212 self-adjoint, 209 similarity transformation, 205 skewsymmetric, 204 symmetric, 204 traces,139, 183–84 transposition, 177 unitary, 209 vector transformation law,198 Maxwell’sequations, 51, 284, 316–20 derivation of waveequations, 52 dual transformation, 290 Gauss’ law,52–53 Legendre equation, 779 Lorentzcovariance of, 283–91 electromagneticinvariants, 288 exercises,289–91 overview, 283–86 transformation of EandB, 287–88 Oersted’s law,52–53 mean valuetheorem, 353 Mellin transforms, 897, 933 meromorphic functions, 439, 461 integral of, 466 pole expansion of, 461 metric, curvilinearcoordinates, 105 metric tensor, 151–54 Christoffel symbols asderivatives of, 155–56 minimal substitution, 76 Minkowski space,136, 278–79 Minkowski space–time, seekinematics and dynamics in Minkowski space–time minor, 168 minors, Laplaciandevelopment by, 167–68 missing dependent variables, 1042–43 Mittag-Lefflertheorem, 461 1172 Index mixed tensor, 136, 139, 154 modes of vibration, normal, 233–34 modified Bessel functions, 713–19 asymptotic expansion, 711, 719 Fourier transform, 716 generating function, 709 integral representation, 720–23 Laplacetransform, 933 recurrencerelations, 714–16 seriesform, 714 modified Helmholtz operator, 598 modified Mathieuequation, 872 modulus, 406 moment of inertia matrix, 215–16 momentum, seeangular momentum; orbital angular momentum momentum representation, 955–61 harmonic oscillator, 958–59 hydrogen atom, 957–58 Schrödinger wave equation, 957–58 momentum representation in quantum mechanics, 1006–7 monopole, 745–46 Morera’s theorem, 427–28 motion arealaw for planetary, 116–19 equations of, 142 moving particle Cartesian coordinates, 1054 circularcylindrical coordinates, 1054–55 multiplet, 245 multiplication ofmatrices direct product, 181–82 of vectors, 182 inner product, 179–81 by scalar,179 multiply connected regions, 423–24 multipole expansion, electrostatic,599 multivalent functions, 447–50 multivalued function, 409 mutually exclusive, 1109–10 N Navier–Stokes equations, 119 Neumann boundary conditions, 543 Neumann functions, 511 asymptotic form, 602–3 Besselfunctions of second kind, 699–707 coaxial waveguides, 703–4 definition and series form, 699–700 other forms, 701 recurrence relations, 702 Wronskian formulas, 702–3 Fourier transform, 602Hankelfunction definition, 707–8 integral form, 701 recurrence relations, 702 spherical, 507, 727 Wronskian formulas, 702–3 Neumann functions, integral form, 701 Neumann problem, 617 Neumann series, 1018–19 Neumannseries, separable(degenerate) kernels, 1018–29 numerical solution, 1023–25 separablekernel, 1021–22 Neumann series solution, 1020–21 neutron diffusion theory, 369, 951 Newton’s second laws,130, 233, 370, 973, 1054 node, spiral, 1097, 1103–4 non-Cartesian tensors, 140 nonessential (regular) singular point, 563 nonhomogeneous equation—Green’s function, 592–610 circularcylindricalcoordinateexpansion, 601–2 form of Green’sfunctions, 596–98 Legendre polynomial addition theorem, 600–601 spherical polar coordinate expansion, 598–600 symmetry of Green’s function, 595–96 nonhomogeneous integral equation, 1032–34 nonlinear differential equations (NDEs),1088–89 autonomous differential equations, 1091–93 Bernoulli and Riccatiequations, 1089–90 bifurcations in dynamical systems, 1103–4 centeror cycle,1100–1101 chaos in dynamical systems, 1105–6 dissipation in dynamical systems, 1102–3 fixedandmovable singularities, special solutions, 1090 localandglobal behavior in higher dimensions, 1093–94 routes to chaos in dynamical systems, 1106–7 saddle point, 1095–97 spiral fixed point, 1098–1100 stablesink, 1095 nonlinear methods and chaos, 1079–1108 introduction, 1079–80 logistic map, 1080–84 exercises,1084 nonlinear differential equations (NDEs), 1088–89 autonomous differential equations, 1091–93 Bernoulli and Riccatiequations, 1089–90 bifurcations in dynamical systems, 1103–4 centerorcycle,1100–1101 chaosin dynamical systems, 1105–6 dissipation in dynamical systems, 1102–3 exercises,1089, 1102, 1106, 1107 Index 1173 fixed andmovable singularities, special solutions, 1090 local andglobal behavior inhigher dimensions, 1093–94 routestochaosindynamicalsystems,1106–7 saddle point, 1095–97 spiral fixed point, 1098–1100 stable sink, 1095 sensitivity to initial conditions and parameters, 1085–88 exercises, 1088 fractals, 1086–88 Lyapunov exponents, 1085–86 nonuniform convergence, 348–49 normalization, 695 normal matrices,231–39 exercises, 236–39 ill-conditioned, 234–35 normal modes of vibration, 233–34 overview, 231–32 null matrix, 178 number operator, 823 numbers, Bernoulli, seeBernoulli numbers numerical solution, 1023–25 nutation, 973–74 O Oersted’s law,52, 66–68 Olbers’ paradox, 337 operators, differential vector, seedifferential vector operators optical dispersion, 484–85 optical pathnearevent horizon of blackhole, 1041–42 orbital angular momentum, 251, 261–66 exercises, 266 ladder operator approach, 262–64 Liegroups and algebras, 264–66 Liegroups and operators, order of, 264–65 operators, 793–97 overview, 261 rotation of, 251 order 2 branch points, 440–42 order parameter, 1086 ordinary differential equations (ODEs), linear first-order, 547–50 orientedsurface, magnetic fluxacross, 306 orthogonal coordinates Laplacianin, 316 inR3, 103–10 exercises, 109–10 Jacobians for polar coordinates, 108 overview, 103–8 orthogonal eigenfunctions, 636, 637, 1030–32orthogonal functions, representation ofDiracdelta function by, 88–89 orthogonal groups, 243 orthogonal group SO(3), 254, 256 orthogonality, 694–99, 731, 776–78, 821–22, 854–55 Bessel series,695 continuum form, 696 curvilinear coordinates, 104 Earth’s gravitational field, 758–59 electrostaticpotential in hollow cylinder, 695–96 electrostaticpotential of ring ofcharge, 761–62 expansionoffunctions,Legendreseries,757–58 Fourier series, 636–37 Fourier series: Hilbert–Schmidt integral equations, 919–29 normalization, 695 over discrete points, 914–15 sphere in a uniform field, 759–61 Sturm–Liouville differential equations, 1031 of vectors, 14 orthogonality condition, 10, 198 orthogonality integral Hermitepolynomials, 821 Legendre polynomials, 777 spherical harmonics, 788 orthogonality relations, spherical harmonics, 814 orthogonalization, Gram–Schmidt, 173, 642–7 orthogonal matrices, 195–208, 209 applications to vectors, 197–98 direction cosines, 196–97 Eulerangles, 202–3 exercises, 206–8 inverse, 200 overview, 195–96 relation to tensors, 206 symmetry properties, 203–5 transpose matrix, ˜A,200–202 two-dimensional conditions for, 199–200 orthonormal functions, 642–44 polynomials, 646–47 vectors, 174 oscillator damping, 979–80, 991–93, 1101 driven, 457, 565, 991–93 forced classical,457–60 harmonic, 822–27 integral equation for, 991–93 Laplacetransform solution, 991–93 linear, 668–69 momentum spacewavefunction, 569 self-adjoint equation, 679 series solution of differential equation, 568–69 1174 Index singularities in harmonic oscillator differential equation, 564 oscillatory series,322 overshoot, calculation of, 912–13 P parabolic PDEs,538 parallelogram addition law,2–3 paralleltransport, 157–60 parity, 753, 776, 939 Besselfunctions, 687, 735 Chebyshev functions, 569 differential operator, 569 Fourier cosine, sine transforms, 939 Hermitefunctions, 569 Legendre functions, 569 Legendre functions, associated,776 Legendre functions, second kind, 812 spherical harmonics, 776, 791 vector spherical harmonics, 814 parity transformation, 143 Parseval relation, 485–86, 952–53 partial differential equations (PDEs), 535–43 bicharacteristicsof, 538 boundary conditions, 542–43 characteristicsof, 538–41 classesof, 538–41 elliptic,538 examples of, 536–38 harmonic functions, 539 hyperbolic, 538 introduction, 535–36 inversion of, 949–50 nonlinear, 541–42 parabolic,538 partial fraction expansion, 968–69 partial sum approximation, 390 particle in a box, 1061–62 in asphere, 731–32 particlemotion Cartesian coordinates, 1054 circularcylindrical coordinates, 1054–55 quantum mechanical,827 in rectangular box, 1061–62 in right circularcylinder, 1062 in sphere, 731–32 path-dependent work, 56–60 integral definitions ofgradient, divergence, and curl, 58–59 overview, 56 surfaceintegrals, 56–57 volume integrals, 57–58 Pauli matrices, 209–12 pendulums, period of simple, 370–71periodic functions, 888–90 boundary conditions, 636, 883–5 permutations and combinations, counting of, 1114–15 phase of a complex function, 409 phase of a complex number, 409 phase space,88, 1080, 1091 physical interpretation ofdivergence, 40–42 pi,π,219, 379, 396–400, 586 Leibniz formula, 590, 776, 886 Wallisformula, 399 pion photoproduction threshold, 282–83 pi-sine relation, verification of, 522 planetary motion, arealaw for, 116–19 Pochhammer symbol, 859–60, 864 Poincaré section, 1088–89, 1104–6 point and spacegroups, crystallographic, 299–300 point source equation, 662 Poisson distribution, 1130–33 Poisson’s equation, 536 and Gauss’ law,81–83 Green’sfunction, 669–70 polar angledependence, 788 polar coordinates, Jacobians for, 108–10, seealso spherical polar coordinates polar vectors, 143–44 pole expansion of meromorphic functions, 461 poles, 439 simple, on contour of integration, 468–69 polygamma functions, 511–12 Catalan’sconstant, 513 polynomials, Bernoulli, 379–80 potential energy, 309 potentials, 68–79, seealsothermodynamics gradient of, 34 force as,36 Laplacianof, 50–51 overview, 68 scalar,68–72 centrifugal, 72 gravitational, 72 overview, 68–71 vector, 73–79 of constant Bfield,44 exercises,77–79 magnetic,74–76 potential theory conservative force, 69 electrostaticpotential, 593 scalarpotential, 70 vector potential, 73, 311 power series, 363–70 continuity, 364 convergence, 363 uniform and absolute, 363 Index 1175 differentiation and integration, 364 exercises, 366–70 inversion of, 366 overview, 363 uniqueness theorem, 364–65 L’Hôpital’s rule, 365 prime numbers, 379, 382–83, 897 prime number theorem, asymptotic, 897–8 primitive, 898 principal axis, 216 principal value, 409 probability, 1109–51 binomial distribution, 1128–30 exercises, 1129 repeatedtosses of dice,1128–30 definitions, simple properties, 1109–15 conditional probability, 1112 counting of permutations and combinations, 1114–15 exercises, 1115 probability for AorB,1110–11 scholastic aptitude tests,1112–14 Gauss’ normal distribution, 1134–38 exercises, 1137–38 Poisson distribution, 1130–33 exercises, 1133 random variables, 1116–28 continuous random variable: hydrogen atom, 1117–19 discrete random variable, 1116–17 exercises, 1127–28 repeated draws of cards,1123–26 standarddeviationofmeasurements,1119–23 sum, product, and ratio of random variables, 1126–27 statistics, 1138 χ2distribution, 1143–45 confidence interval, 1149 error propagation, 1138–40 exercises, 1150 fitting curves to data,1140–43 studenttdistribution, 1146–49 probability for AorB,1110–11 product convergence theorem, 344 products, seealsocross product; direct product; dot products; scalars expansion of entirefunctions, 462–63 of infinite series,396–401 convergence of infinite, 397–98 exercises, 399–401 overview, 396–97 sine, cosine, and gamma functions, 398–99 product theorem, 181 projection operators, 644 projections, of vectors, 4, 12pseudoscalars, 146 pseudotensors, 142–51 dual tensors, 147–48 exercises, 149–51 irreducible tensors, 149–51 Levi-Civita symbol, 146–47 overview, 142–46 pseudovectors, 146 pullbacks, 309–13 Gauss’ theorem, differential form, 312–13 overview, 309–10 Stokes’ theorem, 310–11 differential form, 311 Q QCD(quantum chromodynamics), 259 Qn(x)functions of thesecond kind, 809–10 quadrupole, 149, 745–47, 805 quantization, 624–25, 1054 quantum chromodynamics (QCD), 259 quantum mechanicalscattering, 469–71 quantum mechanicalsimple harmonic oscillator, 822–27 quantum mechanics, sum rules, 484 angular momentum, 189–92, 251–6, 261–4, 266–70 configuration spacerepresentation, 957–58 expectation values, 263 hydrogen atom, 245–46 momentum representation, 1006–7 Schrödinger representation, 1006–7 quantum pendulum, 873 quasiperiodic, 1100–1101, 1107 quaternions, 189, 204, 212 quotient rule, 141–42 equations of motion and field equations, 142 exercises, 142 overview, 141–42 R R3,orthogonal coordinates in, seeorthogonal coordinates Raabe’stest, 332 radial Mathieuequation, 872 radial Mathieufunctions, 874–79 radioactive decay,553, 977–78, 1129, 1130–31, 1133 raising operator, 263 random variables, 1116–28 continuous random variable: hydrogen atom, 1117–19 discreterandom variable,1116–17 standard deviation of measurements, 1119–23 sum, product, and ratio of random variables, 1126–27 1176 Index rank, 264 of matrices, 178 of tensor, 133 rapidity, 280 rational approximations, 345 ratio test,Cauchy, d’Alembert, 326 Rayleigh formulas, 730 Rayleigh–Ritz variational technique, 1072–76 ground stateeigenfunction, 1073 vibrating string, 1074 realpart, 407 rearrangement of double series, 345–48 reciprocal,916 reciprocal lattice,27 reciprocity principle, 665 rectifier,full-wave, 893–94 recurrence relations, 567, 714–16, 749–50, 775, 818–19 application of, 804–5 associatedLegendre polynomials, 775 Bernoulli numbers, 377 Besselfunctions, 677 Besselfunctions, spherical, 730 Chebyshev polynomials, 850 confluent hypergeometric functions, 866 derivatives, 852–53 exponential integral, 529 factorialfunction, gamma, 504, 506, 529, 995 Hankelfunctions, 708 Hermitepolynomials, 818–19 hypergeometric functions, 861 Laguerrefunctions, associated,842 Legendre polynomials, 749–50 Legendre series,807–8 modified Bessel functions, 714–16 Neumann functions, 702 polygamma functions, 512 and specialproperties, 749–56 differential equations, 751–52 parity, 753 recurrence relations, 749–50 special values,752 upper and lower bounds for Pn(cosθ), 753–54 spherical Besselfunctions, 730 reducible representations, 245–46 reflection principle, Schwarz,431–32 regression coefficient,1140–41 regular (nonessential) singular point, 563 regular functions (holomorphic or analytic functions), 415 regular singularities, 572–73 relations, seedispersion relations relativistic energy, 356–57 repeateddraws ofcards, 1123–26repellant, 1082 repellor, 1092 representation fundamental, 259–60, 274–75 irreducible, 246, 263, 265, 268, 274, 276, 293, 298 reducible, 246 residue theorem, 455–56, 472; see alsocalculus of residues resonant cavity, 682–85 Riccatiequation, 1089–90 Riemann–Christoffel curvature tensor, 138 Riemann integral, 60–61 Riemann manifold, 313–14 Riemann’s theorem, 344 Riemann surface,448–49 Riemann Zetafunction, 329–30, 334–39 and Bernoulli numbers, 382–84 Fourier series evaluation, 329 tableof values,382 infinite series, 894–98 Riesz’ theorem, 177 RLC analog, 980–81 RL circuit,549–50 Rodrigues’ formula, 767 Laguerrepolynomials, 839 associated,767, 842 Rodrigues representation, 820–21 Hermitepolynomials, 820–21 root diagram, 264 root test,Cauchy, 326 rotations, 176, 444 of coordinate axes, 7–12 exercises,12 vectors andvector space,11–12 of coordinates, 199 offunctionsandorbitalangularmomentum,251 groupsSO(2)andSO(3), 250 invariance of scalar product under, 15–17 isomorphic and homomorphic, 244–45 Rouché’s theorem, 463 routes to chaos in dynamical systems, 1106–7 S saddle points (steepest descent method), 489–97, 1092, 1095–97, 1103 analyticlandscape, 489–90 asymptotic forms offactorial function Ŵ, 494–95 ofHankelfunction H(1) ν(s),493–94 exercises, 496–97 factorialfunction Ŵ(z),494–95 Hankelfunction H(1) ν,493–94 overview, 489 sample space,1109–10 Index 1177 sawtooth wave,885–86 scalarpotential, 33–34, 70, 536 scalarquantities, 1 scalars,7, 133 multiplication of matricesby, 179 potentials, 68–72 centrifugal, 72 gravitational, 72 overview, 68–71 products, 12–17 exercises, 17 invariance of under rotations, 15–17 overview, 12–15 triple, 25–27 scattering, quantum mechanical,469–71 Schlaefliintegral, 709, 768–69 Legendre polynomials, 768 Schmidt orthogonalonization, seeGram–Schmidt scholastic aptitude tests, 1112–14 Schrödinger’s wave equation, 624 degeneracy of, 638 hydrogen atom, 843 Schrödinger wave equation, 76, 537, 1069–70 momentum space representation, 957–58 variational derivations, 1069–70 Schur’s lemma,265 Schwarzinequality, 652–54 generalized, 661 Schwarzreflection principle, 431–32 second-rank tensors, 135–36 section,seePoincaré section secularequation, 218 selectionrules, 265 self-adjoint eigenvalue equations, 595 self-adjoint matrices, 209 self-adjoint ODEs,622–34 boundary conditions, 627–28 deuteron, 626–27 eigenfunctions, eigenvalues, 624–25 Hermitian operators, 629 Hermitian operators in quantum mechanics,630 integration interval [a,b], 628–29 Legendre’s equation, 625 self-adjoint operator, 623, 630 Sturm–Liouville theory, differential equations, 622–30 semiconvergent series,391 sensitivity to initial conditions and parameters, 1085–88 fractals, 1086–88 Lyapunov exponents, 1085–86 separablekernel, 1021–22 separablevariables, 544–45 separationof variables inellipticalcoordinates, 870–71series,seeinfinite series series approach, 570–71 Bessel’s equation, limitations of, 570–71 Chebyshev, 338, 569 Hermite,567 hypergeometric, 569, 859, 862 incomplete beta,523 Laguerre,569 Legendre,333, 507, 569, 757–62 shifted polynomials, 836 Chebyshev, 850 Legendre,629 ultraspherical, 338 series form, 714 of second solution, 583–85 series solutions—Frobenius’ method, 565–78 expansion about x0, 569 Fuchs’ theorem, 573 limitations of seriesapproach—Bessel’s equation, 570–71 regular and irregular singularities, 572–73 symmetry of solutions, 569 sign changes,series with alternating, irregular, 341–42 similarity transformation, 205 simple harmonic oscillator, 973 simple pendulum, 1067–68 simple pole on contour of integration, 468–69 sine confluent hypergeometric representation, 393 functions of infinite products, 398–99 integrals in asymptotic series, 392–93 sine transform, 939–40 singularities, 438–43 branch points, 440–42 of order 2, 440–42 exercises, 442–43 fixed, 1090 Laurentseries, 439 movable, 1090 overview, 438 poles, 439 singular points, 562–65 essential(irregular), 563 irregular (essential), 563 nonessential (regular), 563 regular (nonessential), 563 sink, 1092–96, 1101 sinx infinite product representation, 378, 392–93, 398, 439, 468–69, 583, 679–80, 728–29, 730, 885, 889, 892, 911, 950, 1021, 1097–98 power series,358, 406, 410 skewsymmetric matrices, 204 1178 Index Slaterdeterminant, 274 small oscillations, 370–71 SO(2)rotation groups, 250 SO(3) Clebsch–Gordan coefficients, 267–70 homomorphism, 252–56 rotation groups, 250 soap film, 1045–46 soap film—minimum area,1046–49 solenoidal, 42 soliton solutions, 542 space,vector, seevectors spaceand point groups, crystallographic, 299–300 space–time,Minkowski, seekinematics and dynamics in Minkowski space–time specialunitary groups SU(2), Pauli spin matrices,189, 203–4, 250–66, 267–70, 274–76 SU(3), Gell-Mannmatrices, 212, 256–60, 265–66, 274–76 SU(n), Young tableaux,274–76 specialunitary group SU(2), 252 specialvalues, 774–75 spectraldecomposition, 219, 225, 635 sphere in a uniform field, 759–61 spheres, total charge inside, 88 spherical Bessel functions, 725–39 asymptotic values,729 definitions, 726–29 limiting values, 729–30 recurrencerelations, 730 spherical components, 271 spherical coordinates, Helmholtz equation, 725 spherical harmonics, 264, 786–93 addition theorem for, 797–802 derivation of addition theorem, 798–800 trigonometric identity, 797–98 angular momentum operators, 793 azimuthal dependence—orthogonality, 787 Condon–Shortley phase conventions, 270 integrals of, 804 ladder operators, 796–97 Laplaceseries,expansion theorem, 790–91 orthogonality integral, 788 orthogonality relations, 814 polarangle dependence,788 spherical harmonics, 788–90 vector spherical harmonics, 813–16 spherical polar coordinates, 123–33, 557–60 exercises,128–33 expansion, 598–600 magneticvector potential, 127–28 ∇,∇·,∇×for centralforce, 127 overview, 123–26 unit vectors, 123spherical symmetry, 616 spherical tensor operator, 271 spherical tensors, 271–74 spherical waves,Bessel functions, 730 spinors, 138–39 exercises, 138–39 overview, 138 spinor wave functions, 212 spiral fixed point, 1098–1100 spiral node, 1097, 1103 spiral repellor, 1097, 1103 square integrable, 485, 487, 653, 658, 882 squares ofseries, divergent, 344 square wave,911–12 square wave—high frequencies, 892–93 stable sink, 1095 standard deviation, 1119, 1140 standard deviation of measurements, 1119–23 stark effect,576, 847 statisticalhypothesis, 1138 statistics, 1138 χ2distribution, 1143–45 confidence interval, 1149 error propagation, 1138–40 fitting curves to data,1140–43 studenttdistribution, 1146–49 steepestdescent, method of, 489–96 factorialfunction, 494–95 Hankelfunctions, 493–94 modified Bessel functions, 720 step function, 969–70 Stirling’s expansion, 495 Stirling’s series,516–20 derivation from Euler–Maclaurin integration formula, 517 Stokes’ theorem, 64–68 alternateforms of, 66–68 Oersted’s and Faraday’s laws, 66–68 overview, 66 on differential forms, 313–14 Riemann manifold, 313–14 exercises, 67–68 overview, 64–65 proof, 420–21 pullbacks, 310–11 straight line, 1044–45 strange attractor, 1085, 1086 string, Lagrangian of avibrational, 1058, 1071 structure constants, 248, 251, 259 studenttdistribution, 1146–49 Sturm–Liouville theory, 885 Sturm–Liouville theory—orthogonal functions, 621–74 completeness of egenfunctions, 649–61 Bessel’s inequality, 651–52 Index 1179 expansion coefficients,658 Schwarzinequality, 652–54 summary—vectorspaces,completeness, 654–58 Sturm–Liouville theory—orthogonal functions completeness of Eigenfunction, 659–61 Gram–Schmidt orthogonalization, 642–49 exercises, 647–49 Legendre polynomials by Gram–Schmidt orthogonalization, 644–46 Green’s function—eigenfunction expansion, 662–74 eigenfunction, eigenvalue equation, 667–68 exercises, 670–74 Green’sfunctionandtheDiracdeltafunction, 669–70 Green’s function integral—differential equation, 665–67 Green’s functions—one-dimensional, 663–65 linearocillator, 668–69 Hermitian operators, 634–42 degeneracy, 638 exercises, 639–42 expansion in orthogonal eigenfunctions—square wave,637 Fourier series—orthogonality, 636–37 orthogonal eigenfunctions, 636 real eigenvalues, 634–35 self-adjoint ODEs,622–34 boundary conditions, 627–28 deuteron, 626–27 eigenfunctions, eigenvalues, 624–25 exercises, 631–34 Hermitian operators, 629 Hermitian operators in quantum mechanics, 630 integration interval [a,b], 628–29 Legendre’s equation, 625 SU(2) Clebsch–Gordan coefficients, 267–70 isospin and SU(3)flavor symmetry, 256–60 andSO(3)homomorphism, 252–56 SU(3)flavor symmetry, 256–60 subgroups and cosets,293–94 substitution, 979 subtraction of matrices, 178–79 of series, 324–25 of sets, 1111 of tensors, 136 sum, product, and ratio of random variables, 1126–27 summation convention, 136–37, 139 summation of series, 910sum rules, 484 SU(n), young tableaux for, 274–78 superposition principle for homogenous ODEs, PDEs,536 surface integrals, 56–57 symmetric matrices, 204 symmetric tensor, 137 symmetrization of kernels, 1029–30 symmetry, 889 axes threefold, 296–99 twofold, 294–96 cylindrical, 617 properties of orthogonal matrices,203–5 relations, 484 of solutions, 569 spherical, 616 SU(3)flavor, 256–60 of tensors, 137 T tableauxfor SU(n), Young, 274–78 Taylor’s expansion, 352–63, 430–31 binomial theorem, 356–57 relativistic energy, 356–57 exercises, 358–63 Maclaurintheorem, 354–55 exponential function, 354–55 logarithm, 355 overview, 354 multiple variables,358 overview, 352–54 tensor analysis contravariant tensor, 135–36, 153 contravariant vector, 134–35, 139, 152–54, 158 covariant tensor, 135–36, 158 covariant vector, 134–35, 139–40, 152–54, 156, 158 definition, 133–36 displacement, 158 isotropic tensor, 137 non-Cartesian tensors, 140 paralleltransport, 158 scalarquantity, 8, 15, 57, 134, 149, 179 spherical components, 271 spherical tensor operator, 271 symmetry–asymmetry, 137 tensor density, seepseudotensors tensor derivative operators curl, 162–63 divergence, 160–61 exercises, 162–63 Laplacian,161–62 overview, 160 1180 Index tensors,seealsoderivative operators, tensor; direct product; general tensors; pseudotensors; quotient rule; spinors generaltensors, 151–60, seealsoChristoffel symbols covariant derivative, 156 exercises, 158–60 geodesics and paralleltransport, 157–60 metric tensor, 151–54 overview, 151 relation to orthogonal matrices, 206 spherical,271–74 vector analysis in, 133–63 addition and subtraction of, 136 contraction, 139 overview, 133–35 second-rank, 135–36 summation convention, 136–37 symmetry–antisymmetry, 137 thermodynamics, 72–79 exactdifferentials, 72–76 overview, 72–73 vector potential, 73–79 exercises, 77–79 magnetic, 74–76 Thomas precession, 280 threefold Hermite formula, 827–28 threefold symmetry axis,296–99 time-dependent diffusion equation, 536 time-independent diffusion equation, 536 Titchmarsh theorem, 487 trace,139 traceformula, 224–25 Gutzwiller’s,898 tracesof matrices,183–84 trajectory, 38, 1088, 1091–92, 1097, 1099, 1100, 1103, 1105 transfer function, 962 transfer functions, 961–64 significanceof /Phi1(t), 963–64 transform, derivative of, 982–83; seealsocosines; exponential; Fourier; Fourier–Bessel; Hankel;Laplace;Mellin; sine transformation law, 134 transformation of differential equation into integral equation, 1008–9 transformation of EandB, Lorentz,287–88 translation, 443–44, 981 transport, parallel,157–60 transpose matrix, ˜A,200–202 transposition, 177 triangle inequalities, 406 triangle rule, 268–69 trigonometric form, 853–54 trigonometric identity, 797–98triple scalar products, 25–27, 165–66 triple vector products, 27–29, 46 BAC–CABrule, 28, 46, 51 exercises, 27–32 overview, 27 Tschebyscheff, seeChebyshev two-dimensional conditions for orthogonal matrices,199–200 twofold symmetry axis,294–96 U ultraspherical polynomials, extension to, 747 equation, 853, 861 self-adjoint form, 854 uncertainty principle in quantum theory, 941 uniform convergence, 348–49, 363 union of sets,1111 uniqueness theorem, 364–65 descending power series,781 inverse operator, 181 Laurentexpansion, 433 of power series, 364–65 uniqueness theorem, L’Hôpital’s rule, 365 unitary groups, 243 unitary matrices, seealsoHermitian matrices algebra,210–12 Heaviside,93, 981, 985, 996 ring, 180, 187 unit element of group, 293 unit step function, 191 vectorspace,208 unit vectors, 123 Cartesian coordinates, 5 circularcylindrical, 116 orthogonality relation, 651 spherical polar, 201–2 spherical polarcoordinates, 126 upper and lower bounds for Pn(cosθ),753–4 V values, limiting, of elliptic integrals, 374 variables, seealsocomplex variables dependent, 1052–58 Hamilton’s Principle, 1053–54 Laplace’sequation, 1057–58 moving particle—Cartesian coordinates, 1054 moving particle—circular cylindrical coordinates, 1054–55 dependent and independent, 1038–44, 1058–59 alternateforms of Euler equations, 1042 conceptof variation, 1038–41 missing dependent variables,1042–43 opticalpath nearevent horizon ofa black hole, 1041–42 Index 1181 multiple, of Taylor’s expansion, 358 separation of, 554–62 variance,1120 variation, concept of, 1038–41 variation of the constant, 548 variations, seealsocalculus ofvariations variation with constraints, 1065–72 Lagrangian equations, 1066–67 Schrödinger wave equation, 1069–70 simple pendulum, 1067–68 sliding off a log, 1068–69 vectoranalysis parallelogram addition law,2–3 reciprocal lattice,27 rotation of coordinates, 205 transformation law, 134 vector definition, 8–9 representation of, 9 vector expansion, 745–47 vectorfield, 7 vectorintegrals line integrals, 55 surface integrals, 56–57 volume integrals, 57 vector potential, 73, 311 vector product, 5, 11, 18–22, 25–29, 43–44, 46, 147, 272 vectorquantities, 1 vectors, 1–101, seealsocurved coordinates and vectors; divergence, ∇;g r ad i en t ,∇; integration; potentials; rotations; Stokes’ theorem; tensors applications of orthogonal matrices to, 197–98 components, 8, 11, 153, 654 contravariant, 134–35, 139, 152–54, 158 covariant, 134–35, 139, 152–54, 156, 158 cross product of, 315 curl,∇×,43–49 of centralforce field, 44–46 exercises, 47–49 gradient of dot product, 46 integration by parts of, 47 overview, 43 potential ofconstant Bfield,44 definitions and elementary approach, 1–7 exercises, 6–7 differential vector operators, 110–14 curl, 112–13 divergence, 111–12 exercises, 113–14 gradient, 110 overview, 110 Dirac deltafunction, 83–95 exercises, 91–95 integral representations for, 90–95overview, 83–87 phasespace,88 representation by orthogonal functions, 88–89 totalcharge inside sphere, 88 direction, 5 direct product of, 182 elementary approach to, 1–7 elementary approach to, exercises, 6–7 Gauss’ law,79–83 exercises,82–83 Poisson’s equation, 81–83 Gauss’ theorem, 60–64 alternateforms of, 62 exercises,62–64 Green’stheorem, 61–62 overview, 60–61 by Gram–Schmidt orthogonalization, 174–76 Helmholtz’s theorem, 95–101 exercises,100–101 overview, 95–96 irrotational, 45–46 linear dependenceof, 172–73 normal, 15, 56 or cross product, 18–22 exercises,22–25 orthogonal, 108, 173 overview, 1 potentials, magnetic, 127–28 scalaror dot product, 12–17 exercises,17 invariance of under rotations, 15–16 serve as abasis, 5 space,11–12 exercises,12 successive applications of ∇, 49–54 electromagneticwaveequation, 51–53 exercises,53–54 Laplacianof potential, 50–51 overview, 49–50 triangle law of addition, 1–2 triple product of, 27–29 exercises,27–32 triple scalarproduct, 25–27 vectorspace,linearspace,7, 11–12, 173, 177, 208, 219, 247–48, 314, 638, 644, 650, 654–55, 657–58, 884 vector spaces, completeness, 654–58 vector spherical harmonics, 813–16 vector transformation law, 21 velocity ofelectromagnetic wavesin a dispersive medium, 997–99 vibrating string, 1074 vibration, normal modes of, 233–34 vierergruppe, 292–93, 296 1182 Index Volterra equation, 1005–6 volume integrals, 57–58 von Staudt–Clausen theorem, 379 W Wallis’ formula, 399 wavediffusion equation (Helmholtz diffusion equation), 536, 537 wave equation, 947–48 anomalous dispersion, 999 derivation from Maxwell’sequation, 52 Fourier transform solution, 947–48 Laplacetransform solution, 948 waveequation, electromagnetic, 51–53 wave functions, 624 spinor, 212 waveguides, coaxial,Bessel functions, 703–4 Weierstrass infinite-product form of Ŵ(z), 499–500 Weierstrass Mtest, 349–50 weight diagram, 257, 258, 260, 265 weight vectors, 265 Whittaker functions, 866, 919 Wigner–Eckart theorem, 273 WKBexpansion, 394work, potential, 34, 70–72 Wronskian, 550 Wronskian determinant, 579–80 Wronskian formulas, 702–3 absenceof third solution, 587, 589 Bessel functions, 702–4 Bessel functions, spherical, 601, 723 Chebyshev functions, 582–83, 591 confluent hypergeometric functions, 868 Green’sfunction, construction of, 592, 703 linear independence of functions, 579–80, 665, 703 second solution ofdifferential equations, 583–87 solutions of self-adjoint differential equation, 702 Y Young tableaux for SU(n),274–77 Z zero-point energy, 732, 826 zeros, Besselfunction, 682 zetafunction, seeRiemannZetafunction