Phil Lucht Math & Physics Archive
Home / Math and Physics Files / Physics / Physics Book Downloads / Math Methods in Physics Books

Mathematical_Methods_for_Physicists_7th

PDF · 1206 pages · 10.7 MB
Open PDF file

Commercial textbook (Academic Press/Elsevier, 2013) by George Arfken, Hans Weber and Frank Harris, kept in the archive's folder of downloaded math methods books. The contents list covers infinite series, determinants and matrices, vector analysis, tensors and differential forms, vector spaces, eigenvalue problems, ordinary differential equations and Sturm-Liouville theory, with further chapters beyond the part of the text seen here. No notes by Phil are evident in the extracted text.

AI-written summary; may contain errors.

Extracted text (machine-read; may contain errors)
MATHEMATICAL METHODS PHYSICISTS an A ARFKEN, WEBER, wwHARRIS ArfKen_FM-9780123846549.tex MATHEMATICAL METHODS FOR PHYSICISTS SEVENTH EDITION ArfKen_FM-9780123846549.tex MATHEMATICAL METHODS FOR PHYSICISTS A Comprehensive Guide SEVENTH EDITION George B. Arfken Miami University Oxford, OH Hans J. Weber University of Virginia Charlottesville, VA Frank E. Harris University of Utah, Salt Lake City, UT and University of Florida, Gainesville, FL AMSTERDAM •BOSTON •HEIDELBERG •LONDON NEW YORK •OXFORD •PARIS •SAN DIEGO SAN FRANCISCO •SINGAPORE •SYDNEY •TOKYO Academic Press is an imprint of Elsevier ArfKen_FM-9780123846549.tex Academic Press is an imprint of Elsevier 225 Wyman Street, Waltham, MA 02451, USA The Boulevard, Langford Lane, Kidlington, Oxford, OX5 1GB, UK © 2013 Elsevier Inc. All rights reserved. No part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or any information storage and retrieval system, without permission in writing from the publisher. Details on how to seek permission and further information about the Publisher’s permissions policies and our arrangements with organizations such as the Copyright Clearance Center and the Copyright Licensing Agency, can be found at our website: www.elsevier.com/permissions . This book and the individual contributions contained in it are protected under copyright by the Publisher (other than as may be noted herein). Notices Knowledge and best practice in this field are constantly changing. As new research and experience broaden our understanding, changes in research methods, professional practices, or medical treatment may become necessary. Practitioners and researchers must always rely on their own experience and knowledge in evaluating and using any information, methods, compounds, or experiments described herein. In using such information or methods they should be mindful of their own safety and the safety of others, including parties for whom they have a professional responsibility. To the fullest extent of the law, neither the Publisher nor the authors, contributors, or editors, assume any liability for any injury and/or damage to persons or property as a matter of products liability, negligence or otherwise, or from any use or operation of any methods, products, instructions, or ideas contained in the material herein. Library of Congress Cataloging-in-Publication Data Application submitted. British Library Cataloguing-in-Publication Data A catalogue record for this book is available from the British Library. ISBN: 978-0-12-384654-9 For information on all Academic Press publications, visit our website: www.elsevierdirect.com Typeset by : diacriTech, India Printed in the United States of America 12 13 14 9 8 7 6 5 4 3 2 1 v   CONTENTS             PREFACE ............................................................................................................................... ............ XI  1. MATHEMATICAL  PRELIMINARIES  ......................................................................................................  1  1.1.  Infinite Series ..................................................................................................................  1  1.2.  Series of Functions  .......................................................................................................  21  1.3.  Binomial Theorem ........................................................................................................  33  1.4.  Mathematical  Induction  ...............................................................................................  40  1.5.  Operations  of Series Expansions  of Functions  ..............................................................  41  1.6.  Some Important  Series .................................................................................................  45  1.7.  Vectors .........................................................................................................................  46  1.8.  Complex Numbers and Functions  .................................................................................  53  1.9.  Derivatives  and Extrema ..............................................................................................  62  1.10. Evaluation  of Integrals .................................................................................................  65  1.11. Dirac Delta Functions  ...................................................................................................  75  Additional  Readings ....................................................................................................  82  2. DETERMINANTS  AND MATRICES ....................................................................................................  83  2.1  Determinants  ...............................................................................................................  83  2.2  Matrices .......................................................................................................................  95  Additional  Readings ..................................................................................................  121  3. VECTOR ANALYSIS ....................................................................................................................  123  3.1  Review of Basics Properties  ........................................................................................  124  3.2  Vector in 3 ‐ D Spaces .................................................................................................  126  3.3  Coordinate  Transformations  ......................................................................................  133  vi   3.4  Rotations  in 3 ........................................................................................................  139  3.5  Differential  Vector Operators  .....................................................................................  143  3.6  Differential  Vector Operators:  Further Properties  ......................................................  153  3.7  Vector Integrations  ....................................................................................................  159  3.8  Integral Theorems  ......................................................................................................  164  3.9  Potential Theory .........................................................................................................  170  3.10  Curvilinear  Coordinates  ..............................................................................................  182  Additional  Readings ..................................................................................................  203  4. TENSOR AND DIFFERENTIAL  FORMS ..............................................................................................  205  4.1            Tensor Analysis ..........................................................................................................  205  4.2  Pseudotensors,  Dual Tensors .....................................................................................  215  4.3  Tensor in General Coordinates  ...................................................................................  218  4.4  Jacobians  ....................................................................................................................  227  4.5  Differential  Forms ......................................................................................................  232  4.6  Differentiating  Forms .................................................................................................  238  4.7  Integrating  Forms ......................................................................................................  243  Additional  Readings ..................................................................................................  249  5. VECTOR SPACES .......................................................................................................................  251  5.1  Vector in Function Spaces ..........................................................................................  251  5.2           Gram ‐ Schmidt Orthogonalization  .............................................................................  269  5.3           Operators  ...................................................................................................................  275  5.4           Self‐Adjoint Operators  ................................................................................................  283  5.5  Unitary Operators  ......................................................................................................  287  5.6  Transformations  of Operators....................................................................................  292  5.7  Invariants  ...................................................................................................................  294  5.8  Summary  – Vector Space Notations  ...........................................................................  296  Additional  Readings ..................................................................................................  297  6. EIGENVALUE  PROBLEMS .............................................................................................................  299  6.1  Eigenvalue  Equations  .................................................................................................  299  6.2  Matrix Eigenvalue  Problems  ......................................................................................  301  6.3  Hermitian  Eigenvalue  Problems  .................................................................................  310  6.4  Hermitian  Matrix Diagonalization  .............................................................................  311  6.5  Normal Matrices ........................................................................................................  319  Additional  Readings ..................................................................................................  328  7. ORDINARY  DIFFERENTIAL  EQUATIONS ...........................................................................................  329  7.1  Introduction  ...............................................................................................................  329  7.2  First ‐ Order Equations  ...............................................................................................  331  7.3  ODEs with Constant Coefficients  ................................................................................  342  7.4  Second‐Order Linear ODEs .........................................................................................  343  7.5  Series Solutions‐ Frobenius‘  Method ..........................................................................  346  7.6  Other Solutions ..........................................................................................................  358  vii   7.7  Inhomogeneous  Linear ODEs .....................................................................................  375  7.8  Nonlinear  Differential  Equations  ................................................................................  377  Additional  Readings ..................................................................................................  380  8. STURM – LIOUVILLE  THEORY .......................................................................................................  381  8.1  Introduction  ...............................................................................................................  381  8.2  Hermitian  Operators  ..................................................................................................  384  8.3  ODE Eigenvalue  Problems  ..........................................................................................  389  8.4  Variation  Methods .....................................................................................................  395  8.5  Summary,  Eigenvalue  Problems  .................................................................................  398  Additional  Readings ..................................................................................................  399  9. PARTIAL DIFFERENTIAL  EQUATIONS ..............................................................................................  401  9.1  Introduction  ...............................................................................................................  401  9.2  First ‐ Order Equations  ...............................................................................................  403  9.3  Second – Order Equations  ..........................................................................................  409  9.4  Separation  of  Variables  .............................................................................................  414  9.5  Laplace and Poisson Equations  ..................................................................................  433  9.6  Wave Equations  .........................................................................................................  435  9.7  Heat – Flow, or Diffution PDE .....................................................................................  437  9.8  Summary  ....................................................................................................................  444  Additional  Readings ..................................................................................................  445  10. GREEN’ FUNCTIONS ..................................................................................................................  447  10.1  One – Dimensional   Problems  ....................................................................................  448  10.2  Problems  in Two and Three Dimensions  ....................................................................  459  Additional  Readings ..................................................................................................  467  11. COMPLEX VARIABLE THEORY ......................................................................................................  469  11.1  Complex Variables  and Functions  ..............................................................................  470  11.2  Cauchy – Riemann Conditions  ....................................................................................  471  11.3  Cauchy’s Integral Theorem ........................................................................................  477  11.4  Cauchy’s Integral Formula .........................................................................................  486  11.5  Laurent Expansion  ......................................................................................................  492  11.6  Singularities  ...............................................................................................................  497  11.7  Calculus of Residues ...................................................................................................  509  11.8  Evaluation  of Definite Integrals ..................................................................................  522  11.9  Evaluation  of Sums .....................................................................................................  544  11.10     Miscellaneous  Topics ..................................................................................................  547  Additional  Readings ..................................................................................................  550  12. FURTHER TOPICS IN ANALYSIS .....................................................................................................  551  12.1  Orthogonal  Polynomials  .............................................................................................  551  12.2  Bernoulli Numbers .....................................................................................................  560  12.3  Euler – Maclaurin  Integration  Formula ......................................................................  567  12.4  Dirichlet Series ...........................................................................................................  571  viii   12.5  Infinite Products .........................................................................................................  574  12.6  Asymptotic  Series .......................................................................................................  577  12.7  Method of Steepest Descents .....................................................................................  585  12.8  Dispertion  Relations ...................................................................................................  591  Additional  Readings ..................................................................................................  598  13. GAMMA FUNCTION ...................................................................................................................  599  13.1  Definitions,  Properties  ................................................................................................  599  13.2  Digamma  and Polygamma  Functions  ........................................................................  610  13.3  The Beta Function ......................................................................................................  617  13.4  Stirling’s Series ...........................................................................................................  622  13.5  Riemann Zeta Function ..............................................................................................  626  13.6  Other Ralated Function ..............................................................................................  633  Additional  Readings ..................................................................................................  641  14. BESSEL FUNCTIONS ...................................................................................................................  643  14.1  Bessel Functions  of the First kind, Jν(x) .......................................................................  643  14.2  Orthogonality  .............................................................................................................  661  14.3  Neumann  Functions,  Bessel Functions  of  the Second kind ........................................  667  14.4  Hankel Functions  ........................................................................................................  674  14.5  Modified Bessel Functions,    Iν(x) and  Kν(x) ................................................................  680  14.6  Asymptotic  Expansions  ..............................................................................................  688  14.7  Spherical Bessel Functions  .........................................................................................  698  Additional  Readings ..................................................................................................  713  15. LEGENDRE  FUNCTIONS ...............................................................................................................  715  15.1  Legendre Polynomials  ................................................................................................  716  15.2  Orthogonality  .............................................................................................................  724  15.3  Physical Interpretation  of Generating  Function .........................................................  736  15.4  Associated  Legendre Equation ...................................................................................  741  15.5  Spherical Harmonics...................................................................................................  756  15.6  Legendre Functions  of the Second Kind ......................................................................  766  Additional  Readings ..................................................................................................  771  16. ANGULAR MOMENTUM .............................................................................................................  773  16.1  Angular Momentum  Operators  ..................................................................................  774  16.2  Angular Momentum  Coupling ....................................................................................  784  16.3  Spherical Tensors .......................................................................................................  796  16.4  Vector Spherical Harmonics  .......................................................................................  809  Additional  Readings ..................................................................................................  814  17. GROUP THEORY .......................................................................................................................  815  17.1  Introduction  to Group Theory ....................................................................................  815  17.2  Representation  of Groups ..........................................................................................  821  17.3  Symmetry  and Physics ................................................................................................  826  17.4  Discrete Groups ..........................................................................................................  830  ix   17.5  Direct Products ...........................................................................................................  837  17.6  Simmetric  Group ........................................................................................................  840  17.7  Continous  Groups .......................................................................................................  845  17.8  Lorentz Group ............................................................................................................  862  17.9  Lorentz Covariance  of Maxwell’s  Equantions  .............................................................  866  17.10      Space Groups .............................................................................................................  869  Additional  Readings ..................................................................................................  870  18. MORE SPECIAL FUNCTIONS .........................................................................................................  871  18.1  Hermite Functions  ......................................................................................................  871  18.2  Applications  of Hermite Functions  .............................................................................  878  18.3  Laguerre Functions  .....................................................................................................  889  18.4  Chebyshev  Polynomials  ..............................................................................................  899  18.5  Hypergeometric  Functions  .........................................................................................  911  18.6  Confluent  Hypergeometric  Functions  .........................................................................  917  18.7  Dilogarithm  ................................................................................................................  923  18.8  Elliptic Integrals ..........................................................................................................  927  Additional  Readings ..................................................................................................  932  19. FOURIER SERIES ........................................................................................................................  935  19.1  General Properties  .....................................................................................................  935  19.2  Application  of Fourier Series ......................................................................................  949  19.3  Gibbs Phenomenon  ....................................................................................................  957  Additional  Readings ..................................................................................................  962  20. INTEGRAL  TRANSFORMS  .............................................................................................................  963  20.1  Introduction  ...............................................................................................................  963  20.2  Fourier Transforms  .....................................................................................................  966  20.3  Properties  of Fourier Transforms  ...............................................................................  980  20.4  Fourier Convolution  Theorem .....................................................................................  985  20.5  Signal – Proccesing  Applications  ................................................................................  997  20.6  Discrete Fourier Transforms  .....................................................................................  1002  20.7  Laplace Transforms  ..................................................................................................  1008  20.8  Properties  of Laplace Transforms  .............................................................................  1016  20.9  Laplace Convolution  Transforms  ..............................................................................  1034  20.10      Inverse Laplace Transforms  ......................................................................................  1038  Additional  Readings ................................................................................................  1045  21. INTEGRAL  EQUATIONS .............................................................................................................  1047  21.1  Introduction  .............................................................................................................  1047  21.2  Some Special Methods .............................................................................................  1053  21.3  Neumann  Series .......................................................................................................  1064  21.4  Hilbert – Schmidt Theory ..........................................................................................  1069  Additional  Readings ................................................................................................  1079    x   22. CALCULUS OF VARIATIONS ........................................................................................................  1081  22.1  Euler Equation ..........................................................................................................  1081  22.2  More General Variations  ..........................................................................................  1096  22.3  Constrained  Minima/Maxima  ..................................................................................  1107  22.4  Variation  with Constraints  .......................................................................................  1111  Additional  Readings ................................................................................................  1124  23. PROBABILITY  AND STATISTICS ....................................................................................................  1125  23.1  Probability:  Definitions,  Simple Properties  ...............................................................  1126  23.2  Random Variables  ....................................................................................................  1134  23.3  Binomial Distribution  ...............................................................................................  1148  23.4  Poisson Distribution  .................................................................................................  1151  23.5  Gauss’ Nomal Distribution  .......................................................................................  1155  23.6  Transformation  of Random Variables  ......................................................................  1159  23.7  Statistics ...................................................................................................................  1165  Additional  Readings ................................................................................................  1179    INDEX  ............................................................................................................................... ............ 1181                        ArfKen_Preface-9780123846549.tex PREFACE This, the seventh edition of Mathematical Methods for Physicists , maintains the tradition set by the six previous editions and continues to have as its objective the presentation of all the mathematical methods that aspiring scientists and engineers are likely to encounter as students and beginning researchers. While the organization of this edition differs in some respects from that of its predecessors, the presentation style remains the same: Proofs are sketched for almost all the mathematical relations introduced in the book, and they are accompanied by examples that illustrate how the mathematics applies to real-world physics problems. Large numbers of exercises provide opportunities for the student to develop skill in the use of the mathematical concepts and also show a wide variety of contexts in which the mathematics is of practical use in physics. As in the previous editions, the mathematical proofs are not what a mathematician would consider rigorous, but they nevertheless convey the essence of the ideas involved, and also provide some understanding of the conditions and limitations associated with the rela- tionships under study. No attempt has been made to maximize generality or minimize the conditions necessary to establish the mathematical formulas, but in general the reader is warned of limitations that are likely to be relevant to use of the mathematics in physics contexts. TO THE STUDENT The mathematics presented in this book is of no use if it cannot be applied with some skill, and the development of that skill cannot be acquired passively, e.g., by simply reading the text and understanding what is written, or even by listening attentively to presentations by your instructor. Your passive understanding needs to be supplemented by experience in using the concepts, in deciding how to convert expressions into useful forms, and in developing strategies for solving problems. A considerable body of background knowledge xi ArfKen_Preface-9780123846549.tex xii Preface needs to be built up so as to have relevant mathematical tools at hand and to gain experi- ence in their use. This can only happen through the solving of problems, and it is for this reason that the text includes nearly 1400 exercises, many with answers (but not methods of solution). If you are using this book for self-study, or if your instructor does not assign a considerable number of problems, you would be well advised to work on the exercises until you are able to solve a reasonable fraction of them. This book can help you to learn about mathematical methods that are important in physics, as well as serve as a reference throughout and beyond your time as a student. It has been updated to make it relevant for many years to come. WHAT’SNEW This seventh edition is a substantial and detailed revision of its predecessor; every word of the text has been examined and its appropriacy and that of its placement has been consid- ered. The main features of the revision are: (1) An improved order of topics so as to reduce the need to use concepts before they have been presented and discussed. (2) An introduc- tory chapter containing material that well-prepared students might be presumed to know and which will be relied on (without much comment) in later chapters, thereby reducing redundancy in the text; this organizational feature also permits students with weaker back- grounds to get themselves ready for the rest of the book. (3) A strengthened presentation of topics whose importance and relevance has increased in recent years; in this category are the chapters on vector spaces, Green’s functions, and angular momentum, and the inclu- sion of the dilogarithm among the special functions treated. (4) More detailed discussion of complex integration to enable the development of increased skill in using this extremely important tool. (5) Improvement in the correlation of exercises with the exposition in the text, and the addition of 271 new exercises where they were deemed needed. (6) Addition of a few steps to derivations that students found difficult to follow. We do not subscribe to the precept that “advanced” means “compressed” or “difficult.” Wherever the need has been recognized, material has been rewritten to enhance clarity and ease of understanding. In order to accommodate new and expanded features, it was necessary to remove or reduce in emphasis some topics with significant constituencies. For the most part, the material thereby deleted remains available to instructors and their students by virtue of its inclusion in the on-line supplementary material for this text. On-line only are chapters on Mathieu functions, on nonlinear methods and chaos, and a new chapter on periodic sys- tems. These are complete and newly revised chapters, with examples and exercises, and are fully ready for use by students and their instuctors. Because there seems to be a sig- nificant population of instructors who wish to use material on infinite series in much the same organizational pattern as in the sixth edition, that material (largely the same as in the print edition, but not all in one place) has been collected into an on-line infinite series chapter that provides this material in a single unit. The on-line material can be accessed at www.elsevierdirect.com. ArfKen_Preface-9780123846549.tex Preface xiii PATHWAYS THROUGH THE MATERIAL This book contains more material than an instructor can expect to cover, even in a two-semester course. The material not used for instruction remains available for reference purposes or when needed for specific projects. For use with less fully prepared students, a typical semester course might use Chapters 1 to 3, maybe part of Chapter 4, certainly Chapters 5 to 7, and at least part of Chapter 11. A standard graduate one-semester course might have the material in Chapters 1 to 3 as prerequisite, would cover at least part of Chapter 4, all of Chapters 5 through 9, Chapter 11, and as much of Chapters 12 through 16 and/or 18 as time permits. A full-year course at the graduate level might supplement the foregoing with several additional chapters, almost certainly including Chapter 20 (and Chapter 19 if not already familiar to the students), with the actual choice dependent on the institution’s overall graduate curriculum. Once Chapters 1 to 3, 5 to 9, and 11 have been covered or their contents are known to the students, most selections from the remain- ing chapters should be reasonably accessible to students. It would be wise, however, to include Chapters 15 and 16 if Chapter 17 is selected. ACKNOWLEDGMENTS This seventh edition has benefited from the advice and help of many people; valuable advice was provided both by anonymous reviewers and from interaction with students at the University of Utah. At Elsevier, we received substantial assistance from our Acqui- sitions Editor Patricia Osborn and from Editorial Project Manager Kathryn Morrissey; production was overseen skillfully by Publishing Services Manager Jeff Freeland. FEH gratefully acknowledges the support and encouragement of his friend and partner Sharon Carlson. Without her, he might not have had the energy and sense of purpose needed to help bring this project to a timely fruition. ArfKen_Ch01-9780123846549.tex CHAPTER 1 MATHEMATICAL PRELIMINARIES This introductory chapter surveys a number of mathematical techniques that are needed throughout the book. Some of the topics (e.g., complex variables) are treated in more detail in later chapters, and the short survey of special functions in this chapter is supplemented by extensive later discussion of those of particular importance in physics (e.g., Bessel func- tions). A later chapter on miscellaneous mathematical topics deals with material requiring more background than is assumed at this point. The reader may note that the Additional Readings at the end of this chapter include a number of general references on mathemati- cal methods, some of which are more advanced or comprehensive than the material to be found in this book. 1.1 I NFINITE SERIES Perhaps the most widely used technique in the physicist’s toolbox is the use of infinite series (i.e., sums consisting formally of an infinite number of terms) to represent functions, to bring them to forms facilitating further analysis, or even as a prelude to numerical eval- uation. The acquisition of skill in creating and manipulating series expansions is therefore an absolutely essential part of the training of one who seeks competence in the mathemat- ical methods of physics, and it is therefore the first topic in this text. An important part of this skill set is the ability to recognize the functions represented by commonly encountered expansions, and it is also of importance to understand issues related to the convergence of infinite series. 1 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch01-9780123846549.tex 2 Chapter 1 Mathematical Preliminaries Fundamental Concepts The usual way of assigning a meaning to the sum of an infinite number of terms is by introducing the notion of partial sums. If we have an infinite sequence of terms u1,u2,u3, u4,u5, . . . , we define the ith partial sum as siDiX nD1un: (1.1) This is a finite summation and offers no difficulties. If the partial sums siconverge to a finite limit as i!1 , lim i!1siDS; (1.2) the infinite seriesP1 nD1unis said to be convergent and to have the value S. Note that wedefine the infinite series as equal to Sand that a necessary condition for convergence to a limit is that limn!1unD0. This condition, however, is not sufficient to guarantee convergence. Sometimes it is convenient to apply the condition in Eq. (1.2) in a form called the Cauchy criterion, namely that for each " >0there is a fixed number Nsuch that jsjsij<"for all iandjgreater than N. This means that the partial sums must cluster together as we move far out in the sequence. Some series diverge, meaning that the sequence of partial sums approaches 1; others may have partial sums that oscillate between two values, as for example, 1X nD1unD11C11C1.1/nC: This series does not converge to a limit, and can be called oscillatory. Often the term divergent is extended to include oscillatory series as well. It is important to be able to determine whether, or under what conditions, a series we would like to use is convergent. Example 1.1.1 THEGEOMETRIC SERIES The geometric series, starting with u0D1and with a ratio of successive terms rD unC1=un, has the form 1CrCr2Cr3CC rn1C: Itsnth partial sum sn(that of the first nterms) is1 snD1rn 1r: (1.3) Restricting attention to jrj<1, so that for large n,rnapproaches zero, and snpossesses the limit limn!1snD1 1r; (1.4) 1Multiply and divide snDPn1 mD0rmby1r. ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 3 showing that forjrj<1, the geometric series converges. It clearly diverges (or is oscilla- tory) forjrj1, as the individual terms do not then approach zero at large n.  Example 1.1.2 THEHARMONIC SERIES As a second and more involved example, we consider the harmonic series 1X nD11 nD1C1 2C1 3C1 4CC1 nC: (1.5) The terms approach zero for large n, i.e., limn!11=nD0, but this is not sufficient to guarantee convergence. If we group the terms (without changing their order) as 1C1 2C1 3C1 4 C1 5C1 6C1 7C1 8 C1 9CC1 16 C; each pair of parentheses encloses pterms of the form 1 pC1C1 pC2CC1 pCp>p 2pD1 2: Forming partial sums by adding the parenthetical groups one by one, we obtain s1D1;s2D3 2;s3>4 2;s4>5 2;:::; sn>nC1 2; and we are forced to the conclusion that the harmonic series diverges. Although the harmonic series diverges, its partial sums have relevance among other places in number theory, where HnDPn mD1m1are sometimes referred to as harmonic numbers.  We now turn to a more detailed study of the convergence and divergence of series, considering here series of positive terms. Series with terms of both signs are treated later. Comparison Test If term by term a series of terms unsatisfies 0unan, where the anform a convergent series, then the seriesP nunis also convergent. Letting siandsjbe partial sums of the useries, with j>i, the difference sjsiisPj nDiC1un, and this is smaller than the corresponding quantity for the aseries, thereby proving convergence. A similar argument shows that if term by term a series of terms vnsatisfies 0bnvn, where the bnform a divergent series, thenP nvnis also divergent. For the convergent series anwe already have the geometric series, whereas the harmonic series will serve as the divergent comparison series bn. As other series are identified as either convergent or divergent, they may also be used as the known series for comparison tests. ArfKen_Ch01-9780123846549.tex 4 Chapter 1 Mathematical Preliminaries Example 1.1.3 A DIVERGENT SERIES TestP1 nD1np,pD0:999 , for convergence. Since n0:999>n1andbnDn1forms the divergent harmonic series, the comparison test shows thatP nn0:999is divergent. Generalizing,P nnpis seen to be divergent for all p1.  Cauchy Root Test If.an/1=nr<1for all sufficiently large n, with rindependent of n, thenP nanis convergent. If .an/1=n1for all sufficiently large n, thenP nanis divergent. The language of this test emphasizes an important point: The convergence or divergence of a series depends entirely on what happens for large n. Relative to convergence, it is the behavior in the large- nlimit that matters. The first part of this test is verified easily by raising .an/1=nto the nth power. We get anrn<1: Since rnis just the nth term in a convergent geometric series,P nanis convergent by the comparison test. Conversely, if .an/1=n1, then an1and the series must diverge. This root test is particularly useful in establishing the properties of power series (Section 1.2). D’Alembert (or Cauchy) Ratio Test IfanC1=anr<1for all sufficiently large nandris independent of n, thenP nanis convergent. If anC1=an1for all sufficiently large n, thenP nanis divergent. This test is established by direct comparison with the geometric series .1CrCr2C/. In the second part, anC1anand divergence should be reasonably obvious. Although not quite as sensitive as the Cauchy root test, this D’Alembert ratio test is one of the easiest to apply and is widely used. An alternate statement of the ratio test is in the form of a limit: If limn!1anC1 an8 >< >:<1; convergence, >1; divergence, D1; indeterminate.(1.6) Because of this final indeterminate possibility, the ratio test is likely to fail at crucial points, and more delicate, sensitive tests then become necessary. The alert reader may wonder how this indeterminacy arose. Actually it was concealed in the first statement, anC1=anr< 1. We might encounter anC1=an<1for all finite nbut be unable to choose an r<1 and independent of n such that anC1=anrfor all sufficiently large n. An example is provided by the harmonic series, for which anC1 anDn nC1<1: Since limn!1anC1 anD1; no fixed ratio r<1exists and the test fails. ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 5 Example 1.1.4 D’ALEMBERT RATIO TEST TestP nn=2nfor convergence. Applying the ratio test, anC1 anD.nC1/=2nC1 n=2nD1 2nC1 n: Since anC1 an3 4forn2; we have convergence.  Cauchy (or Maclaurin) Integral Test This is another sort of comparison test, in which we compare a series with an integral. Geometrically, we compare the area of a series of unit-width rectangles with the area under a curve. Letf.x/be a continuous, monotonic decreasing function in which f.n/Dan. ThenP nanconverges ifR1 1f.x/dx is finite and diverges if the integral is infinite. The ith partial sum is siDiX nD1anDiX nD1f.n/: But, because f.x/is monotonic decreasing, see Fig. 1.1(a), siiC1Z 1f.x/dx: On the other hand, as shown in Fig. 1.1(b), sia1iZ 1f.x/dx: Taking the limit as i!1 , we have 1Z 1f.x/dx1X nD1an1Z 1f.x/dxCa1: (1.7) Hence the infinite series converges or diverges as the corresponding integral converges or diverges. This integral test is particularly useful in setting upper and lower bounds on the remain- der of a series after some number of initial terms have been summed. That is, 1X nD1anDNX nD1anC1X nDNC1an; (1.8) ArfKen_Ch01-9780123846549.tex 6 Chapter 1 Mathematical Preliminaries 4 3 2 1f(x) f(x) (a)xf(1)=a1 f(1)=a1 f(2)=a2 4 3 2 1 (b)xi= FIGURE 1.1 (a) Comparison of integral and sum-blocks leading. (b) Comparison of integral and sum-blocks lagging. and 1Z NC1f.x/dx1X nDNC1an1Z NC1f.x/dxCaNC1: (1.9) To free the integral test from the quite restrictive requirement that the interpolating func- tion f.x/be positive and monotonic, we shall show that for any function f.x/with a continuous derivative, the infinite series is exactly represented as a sum of two integrals: N2X nDN1C1f.n/DN2Z N1f.x/dxCN2Z N1.xTxU/f0.x/dx: (1.10) HereTxUis the integral part of x, i.e., the largest integer x, soxTxUvaries sawtoothlike between 0 and 1. Equation (1.10) is useful because if both integrals in Eq. (1.10) converge, the infinite series also converges, while if one integral converges and the other does not, the infinite series diverges. If both integrals diverge, the test fails unless it can be shown whether the divergences of the integrals cancel against each other. We need now to establish Eq. (1.10). We manipulate the contributions to the second integral as follows: 1. Using integration by parts, we observe that N2Z N1x f0.x/dxDN2f.N2/N1f.N1/N2Z N1f.x/dx: 2. We evaluate N2Z N1TxUf0.x/dxDN21X nDN1nnC1Z nf0.x/dxDN21X nDN1nh f.nC1/f.n/i DN2X nDN1C1f.n/N1f.N1/CN2f.N2/: Subtracting the second of these equations from the first, we arrive at Eq. (1.10). ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 7 An alternative to Eq. (1.10) in which the second integral has its sawtooth shifted to be symmetrical about zero (and therefore perhaps smaller) can be derived by methods similar to those used above. The resulting formula is N2X nDN1C1f.n/DN2Z N1f.x/dxCN2Z N1.xTxU1 2/f0.x/dx C1 2h f.N2/f.N1/i :(1.11) Because they do not use a monotonicity requirement, Eqs. (1.10) and(1.11) can be applied to alternating series, and even those with irregular sign sequences. Example 1.1.5 RIEMANN ZETA FUNCTION The Riemann zeta function is defined by .p/D1X nD1np; (1.12) providing the series converges. We may take f.x/Dxp, and then 1Z 1xpdxDxpC1 pC1 1 xD1;p6D1; Dlnx 1 xD1; pD1: The integral and therefore the series are divergent for p1, and convergent for p>1. Hence Eq. (1.12) should carry the condition p>1. This, incidentally, is an independent proof that the harmonic series ( pD1) diverges logarithmically. The sum of the first million termsP1;000;000 nD1n1is only 14:392 726:  While the harmonic series diverges, the combination Dlimn!1 nX mD1m1lnn! (1.13) converges, approaching a limit known as the Euler-Mascheroni constant. Example 1.1.6 A SLOWLY DIVERGING SERIES Consider now the series SD1X nD21 nlnn: ArfKen_Ch01-9780123846549.tex 8 Chapter 1 Mathematical Preliminaries We form the integral 1Z 21 xlnxdxD1Z xD2dlnx lnxDln lnx 1 xD2; which diverges, indicating that Sis divergent. Note that the lower limit of the integral is in fact unimportant so long as it does not introduce any spurious singularities, as it is the large- xbehavior that determines the convergence. Because nlnn>n, the divergence is slower than that of the harmonic series. But because lnnincreases more slowly than n", where"can have an arbitrarily small positive value, we have divergence even though the seriesP nn.1C"/converges.  More Sensitive Tests Several tests more sensitive than those already examined are consequences of a theorem by Kummer. Kummer’s theorem, which deals with two series of finite positive terms, un andan, states: 1. The seriesP nunconverges if limn!1 anun unC1anC1 C>0; (1.14) where Cis a constant. This statement is equivalent to a simple comparison test if the seriesP na1 nconverges, and imparts new information only if that sum diverges. The more weaklyP na1 ndiverges, the more powerful the Kummer test will be. 2. IfP na1 ndiverges and limn!1 anun unC1anC1 0; (1.15) thenP nundiverges. The proof of this powerful test is remarkably simple. Part 2 follows immediately from the comparison test. To prove Part 1, write cases of Eq. (1.14) fornDNC1through any larger n, in the following form: uNC1.aNuNaNC1uNC1/=C; uNC2.aNC1uNC1aNC2uNC2/=C; :::::::::::::::::::::::::::; un.an1un1anun/=C: ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 9 Adding, we get nX iDNC1uiaNuN Canun C(1.16) <aNuN C: (1.17) This shows that the tail of the seriesP nunis bounded, and that series is therefore proved convergent when Eq. (1.14) is satisfied for all sufficiently large n. Gauss’ test is an application of Kummer’s theorem to series un>0when the ratios of successive unapproach unity and the tests previously discussed yield indeterminate results. If for large n un unC1D1Ch nCB.n/ n2; (1.18) where B.n/is bounded for nsufficiently large, then the Gauss test states thatP nuncon- verges for h>1and diverges for h1: There is no indeterminate case here. The Gauss test is extremely sensitive, and will work for all troublesome series the physi- cist is likely to encounter. To confirm it using Kummer’s theorem, we take anDnlnn. The seriesP na1 nis weakly divergent, as already established in Example 1.1.6. Taking the limit on the left side of Eq. (1.14), we have limn!1 nlnn 1Ch nCB.n/ n2 .nC1/ln.nC1/ Dlimn!1 .nC1/lnnC.h1/lnnCB.n/lnn n.nC1/ln.nC1/ Dlimn!1 .nC1/lnnC1 n C.h1/lnn : (1.19) Forh<1, both terms of Eq. (1.19) are negative, thereby signaling a divergent case of Kummer’s theorem; for h>1, the second term of Eq. (1.19) dominates the first and is pos- itive, indicating convergence. At hD1, the second term vanishes, and the first is inherently negative, thereby indicating divergence. Example 1.1.7 LEGENDRE SERIES The series solution for the Legendre equation (encountered in Chapter 7) has successive terms whose ratio under certain conditions is a2jC2 a2jD2j.2jC1/ .2jC1/.2jC2/: To place this in the form now being used, we define ujDa2jand write uj ujC1D.2jC1/.2jC2/ 2j.2jC1/: ArfKen_Ch01-9780123846549.tex 10 Chapter 1 Mathematical Preliminaries In the limit of large j, the constant becomes negligible (in the language of the Gauss test, it contributes to an extent B.j/=j2, where B.j/is bounded). We therefore have uj ujC1!2jC2 2jCB.j/ j2D1C1 jCB.j/ j2: (1.20) The Gauss test tells us that this series is divergent.  Exercises 1.1.1 (a) Prove that if limn!1npunDA<1;p>1, the seriesP1 nD1unconverges. (b) Prove that if limn!1nunDA>0, the series diverges. (The test fails for AD0.) These two tests, known as limit tests, are often convenient for establishing the convergence of a series. They may be treated as comparison tests, comparing with X nnq;1q<p: 1.1.2 Iflimn!1bn anDK, a constant with 0<K<1, show that6nbnconverges or diverges with6an. Hint. If6anconverges, rescale bntob0 nDbn 2K. If6nandiverges, rescale to b00 nD2bn K. 1.1.3 (a) Show that the seriesP1 nD21 n.lnn/2converges. (b) By direct additionP100;000 nD2Tn.lnn/2U1D2:02288 . Use Eq. (1.9) to make a five- significant-figure estimate of the sum of this series. 1.1.4 Gauss’ test is often given in the form of a test of the ratio un unC1Dn2Ca1nCa0 n2Cb1nCb0: For what values of the parameters a1andb1is there convergence? divergence? ANS: Convergent for a1b1>1, divergent for a1b11. 1.1.5 Test for convergence (a)1X nD2.lnn/1(d)1X nD1Tn.nC1/U1=2 (b)1X nD1nW 10n(e)1X nD01 2nC1 (c)1X nD11 2n.2nC1/ ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 11 1.1.6 Test for convergence (a)1X nD11 n.nC1/(d)1X nD1ln 1C1 n (b)1X nD21 nlnn(e)1X nD11 nn1=n (c)1X nD11 n2n 1.1.7 For what values of pandqwillP1 nD21 np.lnn/qconverge? ANS: Convergent for(p>1;allq; pD1;q>1;divergent for(p<1; allq; pD1; q1: 1.1.8 GivenP1;000 nD1n1D7:485 470:::set upper and lower bounds on the Euler-Mascheroni constant. ANS: 0:5767< < 0:5778 . 1.1.9 (From Olbers’ paradox.) Assume a static universe in which the stars are uniformly distributed. Divide all space into shells of constant thickness; the stars in any one shell by themselves subtend a solid angle of !0.Allowing for the blocking out of distant stars by nearer stars, show that the total net solid angle subtended by all stars, shells extending to infinity, is exactly 4. [Therefore the night sky should be ablaze with light. For more details, see E. Harrison, Darkness at Night: A Riddle of the Universe. Cambridge, MA: Harvard University Press (1987).] 1.1.10 Test for convergence 1X nD1135.2n1/ 246.2n/2 D1 4C9 64C25 256C: Alternating Series In previous subsections we limited ourselves to series of positive terms. Now, in contrast, we consider infinite series in which the signs alternate. The partial cancellation due to alternating signs makes convergence more rapid and much easier to identify. We shall prove the Leibniz criterion, a general condition for the convergence of an alternating series. For series with more irregular sign changes, the integral test of Eq. (1.10) is often helpful. TheLeibniz criterion applies to series of the formP1 nD1.1/nC1anwith an>0, and states that if anismonotonically decreasing (for sufficiently large n) and limn!1anD0, then the series converges. To prove this theorem, note that the remainder R2nof the series beyond s2n, the partial sum after 2nterms, can be written in two alternate ways: R2nD.a2nC1a2nC2/C.a2nC3a2nC4/C Da2nC1.a2nC2a2nC3/.a2nC4a2nC5/: ArfKen_Ch01-9780123846549.tex 12 Chapter 1 Mathematical Preliminaries Since the anare decreasing, the first of these equations implies R2n>0, while the second implies R2n<a2nC1, so 0<R2n<a2nC1: Thus, R2nis positive but bounded, and the bound can be made arbitrarily small by taking larger values of n. This demonstration also shows that the error from truncating an alter- nating series after a2nresults in an error that is negative (the omitted terms were shown to combine to a positive result) and bounded in magnitude by a2nC1. An argument similar to that made above for the remainder after an odd number of terms, R2nC1, would show that the error from truncation after a2nC1is positive and bounded by a2nC2. Thus, it is generally true that the error in truncating an alternating series with monotonically decreasing terms is of the same sign as the last term kept and smaller than the first term dropped. The Leibniz criterion depends for its applicability on the presence of strict sign alternation. Less regular sign changes present more challenging problems for convergence determination. Example 1.1.8 SERIES WITH IRREGULAR SIGN CHANGES For0<x<2, the series SD1X nD1cos.nx/ nDln 2 sinx 2 (1.21) converges, having coefficients that change sign often, but not so that the Leibniz criterion applies easily. To verify the convergence, we apply the integral test of Eq. (1.10), inserting the explicit form for the derivative of cos.nx/=n(with respect to n) in the second integral: SD1Z 1cos.nx/ ndnC1Z 1 nTnU x nsin.nx/cos.nx/ n2 dn: (1.22) Using integration by parts, the first integral in Eq. (1.22) is rearranged to 1Z 1cos.nx/ ndnDsin.nx/ nx1 1C1 x1Z 1sin.nx/ n2dn; and this integral converges because 1Z 1sin.nx/ n2dn <1Z 1dn n2D1: Looking now at the second integral in Eq. (1.22), we note that its term cos.nx/=n2also leads to a convergent integral, so we need only to examine the convergence of 1Z 1 nTnUsin.nx/ ndn: ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 13 Next, setting .nTnU/sin.nx/Dg0.n/, which is equivalent to defining g.N/DRN 1.n TnU/sin.nx/dn, we write 1Z 1 nTnUsin.nx/ ndnD1Z 1g0.n/ ndnDg.n/ n1 nD1C1Z 1g.n/ n2dn; where the last equality was obtained using once again an integration by parts. We do not have an explicit expression for g.n/, but we do know that it is bounded because sinx oscillates with a period incommensurate with that of the sawtooth periodicity of .nTnU/. This boundedness enables us to determine that the second integral in Eq. (1.22) converges, thus establishing the convergence of S.  Absolute and Conditional Convergence An infinite series is absolutely convergent if the absolute values of its terms form a con- vergent series. If it converges, but not absolutely, it is termed conditionally convergent. An example of a conditionally convergent series is the alternating harmonic series, 1X nD1.1/n1n1D11 2C1 31 4CC.1/n1 nC: (1.23) This series is convergent, based on the Leibniz criterion. It is clearly not absolutely con- vergent; if all terms are taken with + signs, we have the harmonic series, which we already know to be divergent. The tests described earlier in this section for series of positive terms are, then, tests for absolute convergence. Exercises 1.1.11 Determine whether each of these series is convergent, and if so, whether it is absolutely convergent: (a)ln 2 2ln 3 3Cln 4 4ln 5 5Cln 6 6; (b)1 1C1 21 31 4C1 5C1 61 71 8C; (c) 11 21 3C1 4C1 5C1 61 71 81 91 10C1 11C1 151 161 21C: 1.1.12 Catalan’s constant .2/ is defined by .2/D1X kD0.1/k.2kC1/2D1 121 32C1 52: Calculate .2/ to six-digit accuracy. ArfKen_Ch01-9780123846549.tex 14 Chapter 1 Mathematical Preliminaries Hint. The rate of convergence is enhanced by pairing the terms, .4k1/2.4kC1/2D16k .16k21/2: If you have carried enough digits in your summation,P 1kN16k=.16k21/2, addi- tional significant figures may be obtained by setting upper and lower bounds on the tail of the series,P1 kDNC1. These bounds may be set by comparison with integrals, as in the Maclaurin integral test. ANS: .2/D0:9159 6559 4177. Operations on Series We now investigate the operations that may be performed on infinite series. In this connec- tion the establishment of absolute convergence is important, because it can be proved that the terms of an absolutely convergent series may be reordered according to the familiar rules of algebra or arithmetic: If an infinite series is absolutely convergent, the series sum is independent of the order in which the terms are added. An absolutely convergent series may be added termwise to, or subtracted termwise from, or multiplied termwise with another absolutely convergent series, and the result- ing series will also be absolutely convergent. The series (as a whole) may be multiplied with another absolutely convergent series. The limit of the product will be the product of the individual series limits. The product series, a double series, will also converge absolutely. No such guarantees can be given for conditionally convergent series, though some of the above properties remain true if only one of the series to be combined is conditionally convergent. Example 1.1.9 REARRANGEMENT OF ALTERNATING HARMONIC SERIES Writing the alternating harmonic series as 11 2C1 31 4CD 11 21 3 1 41 5 ; (1.24) it is clear thatP1 nD1.1/n1n1<1. However, if we rearrange the order of the terms, we can make this series converge to3 2. We regroup the terms of Eq. (1.24), as  1C1 3C1 5 1 2 C1 7C1 9C1 11C1 13C1 15 1 4 C1 17CC1 25 1 6 C1 27CC1 35 1 8 C:(1.25) ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 15 Partial sum, sn1.500 1.400 1.3001.200 1.100 Number of terms in sum, n10 9 8 7 6 5 4 3 2 1 FIGURE 1.2 Alternating harmonic series. Terms are rearranged to give convergence to 1.5. Treating the terms grouped in parentheses as single terms for convenience, we obtain the partial sums s1D1:5333 s2D1:0333 s3D1:5218 s4D1:2718 s5D1:5143 s6D1:3476 s7D1:5103 s8D1:3853 s9D1:5078 s10D1:4078: From this tabulation of snand the plot of snversus nin Fig. 1.2, the convergence to3 2is fairly clear. Our rearrangement was to take positive terms until the partial sum was equal to or greater than3 2and then to add negative terms until the partial sum just fell below3 2 and so on. As the series extends to infinity, all original terms will eventually appear, but the partial sums of this rearranged alternating harmonic series converge to3 2.  As the example shows, by a suitable rearrangement of terms, a conditionally convergent series may be made to converge to any desired value or even to diverge. This statement is sometimes called Riemann’s theorem. Another example shows the danger of multiplying conditionally convergent series. Example 1.1.10 SQUARE OF A CONDITIONALLY CONVERGENT SERIES MAY DIVERGE The seriesP1 nD1.1/n1 pnconverges by the Leibniz criterion. Its square, "1X nD1.1/n1 pn#2 DX n.1/n1p 11pn1C1p 21pn2CC1pn11p 1 ; ArfKen_Ch01-9780123846549.tex 16 Chapter 1 Mathematical Preliminaries has a general term, in T:::U, consisting of n1additive terms, each of which is bigger than 1pn1pn1, so the entireT:::Uterm is greater thann1 n1and does not go to zero. Hence the general term of this product series does not approach zero in the limit of large nand the series diverges.  These examples show that conditionally convergent series must be treated with caution. Improvement of Convergence This section so far has been concerned with establishing convergence as an abstract math- ematical property. In practice, the rate of convergence may be of considerable importance. A method for improving convergence, due to Kummer, is to form a linear combination of our slowly converging series and one or more series whose sum is known. For the known series the following collection is particularly useful: 1D1X nD11 n.nC1/D1; 2D1X nD11 n.nC1/.nC2/D1 4; 3D1X nD11 n.nC1/.nC2/.nC3/D1 18; :::::::::::::::::::::::::::::: pD1X nD11 n.nC1/.nCp/D1 p pW: (1.26) These sums can be evaluated via partial fraction expansions, and are the subject of Exercise 1.5.3. The series we wish to sum and one or more known series (multiplied by coefficients) are combined term by term. The coefficients in the linear combination are chosen to cancel the most slowly converging terms. Example 1.1.11 RIEMANN ZETA FUNCTION(3) From the definition in Eq. (1.12), we identify .3/ asP1 nD1n3. Noting that 2of Eq. (1.26) has a large- ndependencen3, we consider the linear combination 1X nD1n3Ca 2D.3/Ca 4: (1.27) We did not use 1because it converges more slowly than .3/. Combining the two series on the left-hand side termwise, we obtain 1X nD11 n3Ca n.nC1/.nC2/ D1X nD1n2.1Ca/C3nC2 n3.nC1/.nC2/: ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 17 Table 1.1 Riemann Zeta Function s .s/ 2 1:64493 40668 3 1:20205 69032 4 1:08232 32337 5 1:03692 77551 6 1:01734 30620 7 1:00834 92774 8 1:00407 73562 9 1:00200 83928 10 1:00099 45751 If we choose aD1 , we remove the leading term from the numerator; then, setting this equal to the right-hand side of Eq. (1.27) and solving for .3/, .3/D1 4C1X nD13nC2 n3.nC1/.nC2/: (1.28) The resulting series may not be beautiful but it does converge as n4, faster than n3. A more convenient form with even faster convergence is introduced in Exercise 1.1.16. There, the symmetry leads to convergence as n5.  Sometimes it is helpful to use the Riemann zeta function in a way similar to that illustrated for the pin the foregoing example. That approach is practical because the zeta function has been tabulated (see Table 1.1). Example 1.1.12 CONVERGENCE IMPROVEMENT The problem is to evaluate the seriesP1 nD11=.1Cn2/. Expanding.1Cn2/1Dn2.1C n2/1by direct division, we have .1Cn2/1Dn2 1n2Cn4n6 1Cn2 D1 n21 n4C1 n61 n8Cn6: Therefore 1X nD11 1Cn2D.2/.4/C.6/1X nD11 n8Cn6: The remainder series converges as n8. Clearly, the process can be continued as desired. You make a choice between how much algebra you will do and how much arithmetic the computer will do.  ArfKen_Ch01-9780123846549.tex 18 Chapter 1 Mathematical Preliminaries Rearrangement of Double Series An absolutely convergent double series (one whose terms are identified by two summation indices) presents interesting rearrangement opportunities. Consider SD1X mD01X nD0an;m: (1.29) In addition to the obvious possibility of reversing the order of summation (i.e., doing the m sum first), we can make rearrangements that are more innovative. One reason for doing this is that we may be able to reduce the double sum to a single summation, or even evaluate the entire double sum in closed form. As an example, suppose we make the following index substitutions in our double series: mDq,nDpq. Then we will cover all n0,m0by assigning pthe range.0;1/, andqthe range.0;p/, so our double series can be written SD1X pD0pX qD0apq;q: (1.30) In the nmplane our region of summation is the entire quadrant m0,n0; in the pq plane our summation is over the triangular region sketched in Fig. 1.3. This same pqregion can be covered when the summations are carried out in the reverse order, but with limits SD1X qD01X pDqapq;q: The important thing to note here is that these schemes all have in common that, by allowing the indices to run over their designated ranges, every an;mis eventually encountered, and is encountered exactly once. 4q 2 0 024p FIGURE 1.3 Thepqindex space. ArfKen_Ch01-9780123846549.tex 1.1 In/f_inite Series 19 Another possible index substitution is to set nDs,mDr2s. If we sum over sfirst, its range must be .0;Tr=2U/, whereTr=2Uis the integer part of r=2, i.e.,Tr=2UD r=2forr even and.r1/=2 forrodd. The range of ris.0;1/. This situation corresponds to SD1X rD0Tr=2UX sD0as;r2s: (1.31) The sketches in Figs. 1.4 to1.6show the order in which the an;mare summed when using the forms given in Eqs. (1.29), (1.30), and (1.31), respectively. If the double series introduced originally as Eq. (1.29) is absolutely convergent, then all these rearrangements will give the same ultimate result. m 23 1 0 024 n FIGURE 1.4 Order in which terms are summed with m;nindex set, Eq. (1.29). 4m 23 1 0 02 3 14 n FIGURE 1.5 Order in which terms are summed with p;qindex set, Eq. (1.30). ArfKen_Ch01-9780123846549.tex 20 Chapter 1 Mathematical Preliminaries 46m 2 0 02 3 1 n FIGURE 1.6 Order in which terms are summed with r;sindex set, Eq. (1.31). Exercises 1.1.13 Show how to combine .2/DP1 nD1n2with 1and 2to obtain a series converging asn4. Note..2/ has the known value 2=6. See Eq. (12.66). 1.1.14 Give a method of computing .3/D1X nD01 .2nC1/3 that converges at least as fast as n8and obtain a result good to six decimal places. ANS: .3/D1:051800: 1.1.15 Show that (a)P1 nD2T.n/1UD 1, (b)P1 nD2.1/nT.n/1UD1 2, where.n/is the Riemann zeta function. 1.1.16 The convergence improvement of 1.1.11 may be carried out more expediently (in this special case) by putting 2, from Eq. (1.26), into a more symmetric form: Replacing n byn1, we have 0 2D1X nD21 .n1/n.nC1/D1 4: (a) Combine .3/ and 0 2to obtain convergence as n5. (b) Let 0 4be 4with n!n2. Combine.3/; 0 2, and 0 4to obtain convergence asn7. ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 21 (c) If.3/ is to be calculated to six-decimal place accuracy (error 5107), how many terms are required for .3/ alone? combined as in part (a)? combined as in part (b)? Note. The error may be estimated using the corresponding integral. ANS: .a/ .3/D5 41X nD21 n3.n21/. 1.2 S ERIES OF FUNCTIONS We extend our concept of infinite series to include the possibility that each term unmay be a function of some variable, unDun.x/. The partial sums become functions of the variable x, sn.x/Du1.x/Cu2.x/CC un.x/; (1.32) as does the series sum, defined as the limit of the partial sums: 1X nD1un.x/DS.x/Dlimn!1sn.x/: (1.33) So far we have concerned ourselves with the behavior of the partial sums as a function of n. Now we consider how the foregoing quantities depend on x. The key concept here is that of uniform convergence. Uniform Convergence If for any small ">0there exists a number N,independent of xin the intervalTa;bU (that is, axb) such that jS.x/sn.x/j<"; for all nN; (1.34) then the series is said to be uniformly convergent in the intervalTa;bU. This says that for our series to be uniformly convergent, it must be possible to find a finite Nso that the absolute value of the tail of the infinite series, P1 iDNC1ui.x/ , will be less than an arbitrary small "for all xin the given interval, including the endpoints. Example 1.2.1 NONUNIFORM CONVERGENCE Consider on the interval T0;1Uthe series S.x/D1X nD0.1x/xn: ArfKen_Ch01-9780123846549.tex 22 Chapter 1 Mathematical Preliminaries For0x<1, the geometric seriesP nxnis convergent, with value 1=.1x/, soS.x/D 1for these xvalues. But at xD1, every term of the series will be zero, and therefore S.1/D0. That is, 1X nD0.1x/xnD1;0x<1; D0;xD1: (1.35) SoS.x/is convergent for the entire interval T0;1U, and because each term is nonnegative, it is also absolutely convergent. If x6D0, this is a series for which the partial sum sN is1xN, as can be seen by comparison with Eq. (1.3). Since S.x/D1, the uniform convergence criterion is 1.1xN/ DxN<": No matter what the values of Nand a sufficiently small "may be, there will be an xvalue (close to 1) where this criterion is violated. The underlying problem is that xD1is the convergence limit of the geometric series, and it is not possible to have a convergence rate that is bounded independently of xin a range that includes xD1. We note also from this example that absolute and uniform convergence are independent concepts. The series in this example has absolute, but not uniform convergence. We will shortly present examples of series that are uniformly, but only conditionally convergent. And there are series that have neither or both of these properties.  Weierstrass M(Majorant) Test The most commonly encountered test for uniform convergence is the Weierstrass Mtest. If we can construct a series of numbersP1 iD1Mi, in which Miju i.x/jfor all xin the intervalTa;bUandP1 iD1Miis convergent, our series ui.x/will be uniformly convergent inTa;bU. The proof of this Weierstrass Mtest is direct and simple. SinceP iMiconverges, some number Nexists such that for nC1N, 1X iDnC1Mi<": This follows from our definition of convergence. Then, with jui.x/jMifor all xin the interval axb, 1X iDnC1ui.x/<": Hence S.x/DP1 nD1ui.x/satisfies jS.x/sn.x/jD 1X iDnC1ui.x/ <"; (1.36) ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 23 we see thatP1 nD1ui.x/is uniformly convergent in Ta;bU. Since we have specified absolute values in the statement of the Weierstrass Mtest, the seriesP1 nD1ui.x/is also seen to be absolutely convergent. As we have already observed in Example 1.2.1, absolute and uniform convergence are different concepts, and one of the limitations of the Weierstrass Mtest is that it can only establish uniform convergence for series that are also absolutely convergent. To further underscore the difference between absolute and uniform convergence, we provide another example. Example 1.2.2 UNIFORMLY CONVERGENT ALTERNATING SERIES Consider the series S.x/D1X nD1.1/n nCx2;1<x<1: (1.37) Applying the Leibniz criterion, this series is easily proven convergent for the entire inter- val1<x<1, but it is notabsolutely convergent, as the absolute values of its terms approach for large nthose of the divergent harmonic series. The divergence of the absolute value series is obvious at xD0, where we then exactly have the harmonic series. Never- theless, this series is uniformly convergent on 1<x<1, as its convergence is for all xat least as fast as it is for xD0. More formally, jS.x/sn.x/j<junC1.x/jjunC1.0/j: Since unC1.0/is independent of x, uniform convergence is confirmed.  Abel’s Test A somewhat more delicate test for uniform convergence has been given by Abel. If un.x/ can be written in the form anfn.x/, and 1. The anform a convergent series,P nanDA, 2. For all xinTa;bUthe functions fn.x/are monotonically decreasing in n, i.e., fnC1.x/ fn.x/, 3. For all xinTa;bUall the f.n/are bounded in the range 0fn.x/M, where Mis independent of x, thenP nun.x/converges uniformly in Ta;bU. This test is especially useful in analyzing the convergence of power series. Details of the proof of Abel’s test and other tests for uniform convergence are given in the works by Knopp and by Whittaker and Watson (see Additional Readings listed at the end of this chapter). ArfKen_Ch01-9780123846549.tex 24 Chapter 1 Mathematical Preliminaries Properties of Uniformly Convergent Series Uniformly convergent series have three particularly useful properties. If a seriesP nun.x/ is uniformly convergent in Ta;bUand the individual terms un.x/are continuous, 1. The series sum S.x/DP1 nD1un.x/is also continuous. 2. The series may be integrated term by term. The sum of the integrals is equal to the integral of the sum: bZ aS.x/dxD1X nD1bZ aun.x/dx: (1.38) 3. The derivative of the series sum S.x/equals the sum of the individual-term deriva- tives: d dxS.x/D1X nD1d dxun.x/; (1.39) provided the following additional conditions are satisfied: dun.x/ dxis continuous inTa;bU; 1X nD1dun.x/ dxis uniformly convergent in Ta;bU: Term-by-term integration of a uniformly convergent series requires only continuity of the individual terms. This condition is almost always satisfied in physical applications. Term-by-term differentiation of a series is often not valid because more restrictive condi- tions must be satisfied. Exercises 1.2.1 Find the range of uniform convergence of the series (a).x/D1X nD1.1/n1 nx, (b).x/D1X nD11 nx. ANS: (a)0<sx<1. (b)1<sx<1. 1.2.2 For what range of xis the geometric seriesP1 nD0xnuniformly convergent? ANS:1<sxs<1. 1.2.3 For what range of positive values of xisP1 nD01=.1Cxn/ (a) convergent? (b) uniformly convergent? ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 25 1.2.4 If the series of the coefficientsPanandPbnare absolutely convergent, show that the Fourier series X .ancosnxCbnsinnx/ isuniformly convergent for1<x<1. 1.2.5 The Legendre seriesP jevenuj.x/satisfies the recurrence relations ujC2.x/D.jC1/.jC2/l.lC1/ .jC2/.jC3/x2uj.x/; in which the index jis even and lis some constant (but, in this problem, nota non- negative odd integer). Find the range of values of xfor which this Legendre series is convergent. Test the endpoints. ANS:1<x<1: 1.2.6 A series solution of the Chebyshev equation leads to successive terms having the ratio ujC2.x/ uj.x/D.kCj/2n2 .kCjC1/.kCjC2/x2; with kD0andkD1. Test for convergence at xD1 . ANS: Convergent. 1.2.7 A series solution for the ultraspherical (Gegenbauer) function C n.x/leads to the recurrence ajC2Daj.kCj/.kCjC2 /n.nC2 / .kCjC1/.kCjC2/: Investigate the convergence of each of these series at xD1 as a function of the parameter . ANS: Convergent for <1, divergent for 1. Taylor’s Expansion Taylor’s expansion is a powerful tool for the generation of power series representations of functions. The derivation presented here provides not only the possibility of an expansion into a finite number of terms plus a remainder that may or may not be easy to evaluate, but also the possibility of the expression of a function as an infinite series of powers. ArfKen_Ch01-9780123846549.tex 26 Chapter 1 Mathematical Preliminaries We assume that our function f.x/has a continuous nth derivative2in the interval a xb. We integrate this nth derivative ntimes; the first three integrations yield xZ af.n/.x1/dx1Df.n1/.x1/ x aDf.n1/.x/f.n1/.a/; xZ adx2x2Z af.n/.x1/dx1DxZ adx2h f.n1/.x2/f.n1/.a/i Df.n2/.x/f.n2/.a/.xa/f.n1/.a/; xZ adx3x3Z adx2x2Z af.n/.x1/dx1Df.n3/.x/f.n3/.a/ .xa/f.n2/.a/.xa/2 2Wf.n1/.a/: Finally, after integrating for the nth time, xZ adxnx2Z af.n/.x1/dx1Df.x/f.a/.xa/f0.a/.xa/2 2Wf00.a/ .xa/n1 .n1/Wfn1.a/: Note that this expression is exact. No terms have been dropped, no approximations made. Now, solving for f.x/, we have f.x/Df.a/C.xa/f0.a/ C.xa/2 2Wf00.a/CC.xa/n1 .n1/Wf.n1/.a/CRn; (1.40) where the remainder, Rn, is given by the n-fold integral RnDxZ adxnx2Z adx1f.n/.x1/: (1.41) We may convert Rninto a perhaps more practical form by using the mean value theorem of integral calculus: xZ ag.x/dxD.xa/g./; (1.42) 2Taylor’s expansion may be derived under slightly less restrictive conditions; compare H. Jeffreys and B. S. Jeffreys, in the Additional Readings, Section 1.133. ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 27 with ax. By integrating ntimes we get the Lagrangian form3of the remainder: RnD.xa/n nWf.n/./: (1.43) With Taylor’s expansion in this form there are no questions of infinite series convergence. The series contains a finite number of terms, and the only questions concern the magnitude of the remainder. When the function f.x/is such that limn!1RnD0, Eq. (1.40) becomes Taylor’s series: f.x/Df.a/C.xa/f0.a/C.xa/2 2Wf00.a/C D1X nD0.xa/n nWf.n/.a/: (1.44) Here we encounter for the first time nWwith nD0. Note that we define 0WD1. Our Taylor series specifies the value of a function at one point, x, in terms of the value of the function and its derivatives at a reference point a. It is an expansion in powers of thechange in the variable, namely xa. This idea can be emphasized by writing Taylor’s series in an alternate form in which we replace xbyxChandabyx: f.xCh/D1X nD0hn nWf.n/.x/: (1.45) Power Series Taylor series are often used in situations where the reference point, a, is assigned the value zero. In that case the expansion is referred to as a Maclaurin series, and Eq. (1.40) becomes f.x/Df.0/Cx f0.0/Cx2 2Wf00.0/CD1X nD0xn nWf.n/.0/: (1.46) An immediate application of the Maclaurin series is in the expansion of various transcen- dental functions into infinite (power) series. Example 1.2.3 EXPONENTIAL FUNCTION Letf.x/Dex. Differentiating, then setting xD0, we have f.n/.0/D1 for all n,nD1;2;3;::: . Then, with Eq. (1.46), we have exD1CxCx2 2WCx3 3WCD1X nD0xn nW: (1.47) 3An alternate form derived by Cauchy is RnD.x/n1.xa/ .n1/Wf.n/./. ArfKen_Ch01-9780123846549.tex 28 Chapter 1 Mathematical Preliminaries This is the series expansion of the exponential function. Some authors use this series to define the exponential function. Although this series is clearly convergent for all x, as may be verified using the d’Alembert ratio test, it is instructive to check the remainder term, Rn. By Eq. (1.43) we have RnDxn nWf.n/./Dxn nWe; whereis between 0andx. Irrespective of the sign of x, jRnjjxjnejxj nW: No matter how large jxjmay be, a sufficient increase in nwill cause the denominator of this form for Rnto dominate over the numerator, and limn!1RnD0. Thus, the Maclaurin expansion of exconverges absolutely over the entire range 1<x<1.  Now that we have an expansion for exp.x/, we can return to Eq. (1.45), and rewrite that equation in a form that focuses on its differential operator characteristics. Defining Das theoperator d=dx, we have f.xCh/D1X nD0hnDn nWf.x/DehDf.x/: (1.48) Example 1.2.4 LOGARITHM For a second Maclaurin expansion, let f.x/Dln.1Cx/. By differentiating, we obtain f0.x/D.1Cx/1; f.n/.x/D.1/n1.n1/W.1Cx/n: (1.49) Equation (1.46) yields ln.1Cx/Dxx2 2Cx3 3x4 4CC Rn DnX pD1.1/p1xp pCRn: (1.50) In this case, for x>0our remainder is given by RnDxn nWf.n/./; 0x xn n; 0x1: (1.51) This result shows that the remainder approaches zero as nis increased indefinitely, pro- viding that 0x1. For x<0, the mean value theorem is too crude a tool to establish a ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 29 meaningful limit for Rn. As an infinite series, ln.1Cx/D1X nD1.1/n1xn n(1.52) converges for1<x1. The range1<x<1is easily established by the d’Alembert ratio test. Convergence at xD1follows by the Leibniz criterion. In particular, at xD1we have the conditionally convergent alternating harmonic series, to which we can now put a value: ln 2D11 2C1 31 4C1 5D1X nD1.1/n1n1: (1.53) AtxD1 , the expansion becomes the harmonic series, which we well know to be divergent.  Properties of Power Series The power series is a special and extremely useful type of infinite series, and as illustrated in the preceding subsection, may be constructed by the Maclaurin formula, Eq. (1.44). However obtained, it will be of the general form f.x/Da0Ca1xCa2x2Ca3x3CD1X nD0anxn; (1.54) where the coefficients aiare constants, independent of x. Equation (1.54) may readily be tested for convergence either by the Cauchy root test or the d’Alembert ratio test. If limn!1 anC1 an DR1; the series converges for R<x<R. This is the interval or radius of convergence. Since the root and ratio tests fail when xis at the limit points R, these points require special attention. For instance, if anDn1, then RD1and from Section 1.1 we can conclude that the series converges for xD1 but diverges for xDC1 . IfanDnW, then RD0and the series diverges for all x6D0. Suppose our power series has been found convergent for R<x<R; then it will be uniformly and absolutely convergent in any interior intervalSxS, where 0<S< R. This may be proved directly by the Weierstrass Mtest. Since each of the terms un.x/Danxnis a continuous function of xandf.x/DPanxn converges uniformly for SxS,f.x/must be a continuous function in the inter- val of uniform convergence. This behavior is to be contrasted with the strikingly different behavior of series in trigonometric functions, which are used frequently to represent dis- continuous functions such as sawtooth and square waves. ArfKen_Ch01-9780123846549.tex 30 Chapter 1 Mathematical Preliminaries With un.x/continuous andPanxnuniformly convergent, we find that term by term dif- ferentiation or integration of a power series will yield a new power series with continuous functions and the same radius of convergence as the original series. The new factors in- troduced by differentiation or integration do not affect either the root or the ratio test. Therefore our power series may be differentiated or integrated as often as desired within the interval of uniform convergence (Exercise 1.2.16). In view of the rather severe restric- tion placed on differentiation of infinite series in general, this is a remarkable and valuable result. Uniqueness Theorem We have already used the Maclaurin series to expand exandln.1Cx/into power series. Throughout this book, we will encounter many situations in which functions are repre- sented, or even defined by power series. We now establish that the power-series represen- tation is unique. We proceed by assuming we have two expansions of the same function whose intervals of convergence overlap in a region that includes the origin: f.x/D1X nD0anxn;Ra<x<Ra D1X nD0bnxn;Rb<x<Rb: (1.55) What we need to prove is that anDbnfor all n. Starting from 1X nD0anxnD1X nD0bnxn;R<x<R; (1.56) where Ris the smaller of RaandRb, we set xD0to eliminate all but the constant term of each series, obtaining a0Db0: Now, exploiting the differentiability of our power series, we differentiate Eq. (1.56), getting 1X nD1nanxn1D1X nD1nbnxn1: (1.57) We again set xD0, to isolate the new constant terms, and find a1Db1: By repeating this process ntimes, we get anDbn; ArfKen_Ch01-9780123846549.tex 1.2 Series of Functions 31 which shows that the two series coincide. Therefore our power series representation is unique. This theorem will be a crucial point in our study of differential equations, in which we develop power series solutions. The uniqueness of power series appears frequently in theoretical physics. The establishment of perturbation theory in quantum mechanics is one example. Indeterminate Forms The power-series representation of functions is often useful in evaluating indeterminate forms, and is the basis of l’Hôpital’s rule, which states that if the ratio of two differentiable functions f.x/andg.x/becomes indeterminate, of the form 0=0, atxDx0, then limx!x0f.x/ g.x/Dlimx!x0f0.x/ g0.x/: (1.58) Proof of Eq. (1.58) is the subject of Exercise 1.2.12. Sometimes it is easier just to introduce power-series expansions than to evaluate the derivatives that enter l’Hôpital’s rule. For examples of this strategy, see the following Example and Exercise 1.2.15. Example 1.2.5 ALTERNATIVE TO L’HÔPITAL’S RULE Evaluate lim x!01cosx x2: (1.59) Replacing cosxby its Maclaurin-series expansion, Exercise 1.2.8, we obtain 1cosx x2D1.11 2Wx2C1 4Wx4/ x2D1 2Wx2 4WC: Letting x!0, we have lim x!01cosx x2D1 2: (1.60)  The uniqueness of power series means that the coefficients anmay be identified with the derivatives in a Maclaurin series. From f.x/D1X nD0anxnD1X mD01 nWf.n/.0/xn; we have anD1 nWf.n/.0/: ArfKen_Ch01-9780123846549.tex 32 Chapter 1 Mathematical Preliminaries Inversion of Power Series Suppose we are given a series yy0Da1.xx0/Ca2.xx0/2CD1X nD1an.xx0/n: (1.61) This gives.yy0/in terms of.xx0/. However, it may be desirable to have an explicit expression for .xx0/in terms of.yy0/. That is, we want an expression of the form xx0D1X nD1bn.yy0/n; (1.62) with the bnto be determined in terms of the assumed known an. A brute-force approach, which is perfectly adequate for the first few coefficients, is simply to substitute Eq. (1.61) into Eq. (1.62). By equating coefficients of .xx0/non both sides of Eq. (1.62), and using the fact that the power series is unique, we find b1D1 a1; b2Da2 a3 1; b3D1 a5 1 2a2 2a1a3 ; b4D1 a7 1 5a1a2a3a2 1a45a3 2 ;and so on.(1.63) Some of the higher coefficients are listed by Dwight.4A more general and much more elegant approach is developed by the use of complex variables in the first and second editions of Mathematical Methods for Physicists. Exercises 1.2.8 Show that (a) sinxD1X nD0.1/nx2nC1 .2nC1/W; (b) cosxD1X nD0.1/nx2n .2n/W: 4H. B. Dwight, Tables of Integrals and Other Mathematical Data, 4th ed. New York: Macmillan (1961). (Compare formula no. 50.) ArfKen_Ch01-9780123846549.tex 1.3 Binomial Theorem 33 1.2.9 Derive a series expansion of cot xin increasing powers of xby dividing the power series for cosxby that for sinx. Note. The resultant series that starts with 1=xis known as a Laurent series (cotxdoes not have a Taylor expansion about xD0, although cot.x/x1does). Although the two series for sinxandcosxwere valid for all x, the convergence of the series for cot x is limited by the zeros of the denominator, sinx. 1.2.10 Show by series expansion that 1 2ln0C1 01Dcoth10;j0j>1: This identity may be used to obtain a second solution for Legendre’s equation. 1.2.11 Show that f.x/Dx1=2(a) has no Maclaurin expansion but (b) has a Taylor expansion about any point x06D0. Find the range of convergence of the Taylor expansion about xDx0. 1.2.12 Prove l’Hôpital’s rule, Eq. (1.58). 1.2.13 With n>1, show that .a/1 nlnn n1 <0; .b/1 nlnnC1 n >0: Use these inequalities to show that the limit defining the Euler-Mascheroni constant, Eq. (1.13), is finite. 1.2.14 In numerical analysis it is often convenient to approximate d2 .x/=dx2by d2 dx2 .x/1 h2T .xCh/2 .x/C .xh/U: Find the error in this approximation. ANS: ErrorDh2 12 .4/.x/. 1.2.15 Evaluate lim x!0sin.tan x/tan.sin x/ x7 . ANS:1 30. 1.2.16 A power series converges for R<x<R. Show that the differentiated series and the integrated series have the same interval of convergence. (Do not bother about the endpoints xDR.) 1.3 B INOMIAL THEOREM An extremely important application of the Maclaurin expansion is the derivation of the binomial theorem. ArfKen_Ch01-9780123846549.tex 34 Chapter 1 Mathematical Preliminaries Letf.x/D.1Cx/m, in which mmay be either positive or negative and is not limited to integral values. Direct application of Eq. (1.46) gives .1Cx/mD1CmxCm.m1/ 2Wx2CC Rn: (1.64) For this function the remainder is RnDxn nW.1C/mnm.m1/.mnC1/; (1.65) withbetween 0 and x. Restricting attention for now to x0, we note that for n>m, .1C/mnis a maximum for D0, so for positive x, jRnjxn nWjm.m1/.mnC1/j; (1.66) withlimn!1RnD0when 0x<1. Because the radius of convergence of a power series is the same for positive and for negative x, the binomial series converges for 1<x<1. Convergence at the limit points 1is not addressed by the present analysis, and depends onm. Summarizing, we have established the binomial expansion, .1Cx/mD1CmxCm.m1/ 2Wx2Cm.m1/.m2/ 3Wx3C; (1.67) convergent for1<x<1. It is important to note that Eq. (1.67) applies whether or not mis integral, and for both positive and negative m. Ifmis a nonnegative integer, Rnfor n>mvanishes for all x, corresponding to the fact that under those conditions .1Cx/mis a finite sum. Because the binomial expansion is of frequent occurrence, the coefficients appearing in it, which are called binomial coefficients, are given the special symbol m n Dm.m1/.mnC1/ nW; (1.68) and the binomial expansion assumes the general form .1Cx/mD1X nD0m n xn: (1.69) In evaluating Eq. (1.68), note that when nD0, the product in its numerator is empty (start- ing from manddescending tomC1); in that case the convention is to assign the product the value unity. We also remind the reader that 0Wis defined to be unity. In the special case that mis a positive integer, we may write our binomial coefficient in terms of factorials: m n DmW nW.mn/W: (1.70) Since nWis undefined for negative integer n, the binomial expansion for positive integer mis understood to end with the term nDm, and will correspond to the coefficients in the polynomial resulting from the (finite) expansion of .1Cx/m. ArfKen_Ch01-9780123846549.tex 1.3 Binomial Theorem 35 For positive integer m, them n also arise in combinatorial theory, being the number of different ways nout of mobjects can be selected. That, of course, is consistent with the coefficient set if .1Cx/mis expanded. The term containing xnhas a coefficient that corresponds to the number of ways one can choose the “ x” from nof the factors .1Cx/ and the 1from the mnother.1Cx/factors. For negative integer m, we can still use the special notation for binomial coefficients, but their evaluation is more easily accomplished if we set mDp, with pa positive integer, and write p n D.1/np.pC1/.pCn1/ nWD.1/n.pCn1/W nW.p1/W: (1.71) For nonintegral m, it is convenient to use the Pochhammer symbol, defined for general aand nonnegative integer nand given the notation .a/n, as .a/0D1; .a/1Da; .a/nC1Da.aC1/.aCn/; .n1/: (1.72) For both integral and nonintegral m, the binomial coefficient formula can be written m n D.mnC1/n nW: (1.73) There is a rich literature on binomial coefficients and relationships between them and on summations involving them. We mention here only one such formula that arises if we evaluate 1=p1Cx, i.e.,.1Cx/1=2. The binomial coefficient 1 2 n D1 nW 1 2 3 2  2n1 2 D.1/n13.2n1/ 2nnWD.1/n.2n1/WW .2n/WW; (1.74) where the “double factorial” notation indicates products of even or odd positive integers as follows: 135.2n1/D.2n1/WW 246.2n/D.2n/WW:(1.75) These are related to the regular factorials by .2n/WWD 2nnWand.2n1/WWD.2n/W 2nnW: (1.76) Note that these relations include the special cases 0WWD.1/WWD 1. Example 1.3.1 RELATIVISTIC ENERGY The total relativistic energy of a particle of mass mand velocityvis EDmc2 1v2 c21=2 ; (1.77) ArfKen_Ch01-9780123846549.tex 36 Chapter 1 Mathematical Preliminaries where cis the velocity of light. Using Eq. (1.69) with mD1=2 andxDv2=c2, and evaluating the binomial coefficients using Eq. (1.74), we have EDmc2" 11 2 v2 c2 C3 8 v2 c22 5 16 v2 c23 C# Dmc2C1 2mv2C3 8mv2v2 c2 C5 16mv2 v2 c22 C: (1.78) The first term, mc2, is identified as the rest-mass energy. Then EkineticD1 2mv2" 1C3 4v2 c2C5 8 v2 c22 C# : (1.79) For particle velocity vc, the expression in the brackets reduces to unity and we see that the kinetic portion of the total relativistic energy agrees with the classical result.  The binomial expansion can be generalized for positive integer nto polynomials: .a1Ca2CC am/nDX nW n1Wn2WnmWan1 1an2 2anmm; (1.80) where the summation includes all different combinations of nonnegative integers n1;n2;:::; nmwithPm iD1niDn. This generalization finds considerable use in statisti- cal mechanics. In everyday analysis, the combinatorial properties of the binomial coefficients make them appear often. For example, Leibniz’s formula for the nth derivative of a product of two functions, u.x/v.x/, can be written d dxn u.x/v.x/ DnX iD0n idiu.x/ dxidniv.x/ dxni : (1.81) Exercises 1.3.1 The classical Langevin theory of paramagnetism leads to an expression for the magnetic polarization, P.x/Dccosh x sinhx1 x : Expand P.x/as a power series for small x(low fields, high temperature). 1.3.2 Given that 1Z 0dx 1Cx2Dtan1x 1 0D 4; ArfKen_e9780123846549.tex 1.3 Binomial Theorem 37 expand the integrand into a series and integrate term by term obtaining5  4D11 3C1 51 7C1 9C.1/n1 2nC1C; which is Leibniz’s formula for . Compare the convergence of the integrand series and the integrated series at xD1. Leibniz’s formula converges so slowly that it is quite useless for numerical work. 1.3.3 Expand the incomplete gamma function .nC1;x/xZ 0ettndtin a series of powers ofx. What is the range of convergence of the resulting series? ANS:xZ 0ettndtDxnC11 nC1x nC2Cx2 2W.nC3/ .1/pxp pW.nCpC1/C : 1.3.4 Develop a series expansion of yDsinh1x(that is, sinhyDx) in powers of xby (a) inversion of the series for sinhy, (b) a direct Maclaurin expansion. 1.3.5 Show that for integral n0,1 .1x/nC1D1X mDnm n xmn: 1.3.6 Show that.1Cx/m=2D1X nD0.1/n.mC2n2/WW 2nnW.m2/WWxn, formD1;2;3;::: . 1.3.7 Using binomial expansions, compare the three Doppler shift formulas: .a/ 0D 1v c1 moving sourceI .b/ 0D 1v c moving observerI .c/ 0D 1v c 1v2 c21=2 relativistic: Note. The relativistic formula agrees with the classical formulas if terms of order v2=c2 can be neglected. 1.3.8 In the theory of general relativity there are various ways of relating (defining) a velocity of recession of a galaxy to its red shift, . Milne’s model (kinematic relativity) gives 5The series expansion of tan1x(upper limit 1 replaced by x) was discovered by James Gregory in 1671, 3 years before Leibniz. See Peter Beckmann’s entertaining book, A History of Pi, 2nd ed., Boulder, CO: Golem Press (1971), and L. Berggren, J. Borwein, and P. Borwein, Pi: A Source Book, New York: Springer (1997). ArfKen_Ch01-9780123846549.tex 38 Chapter 1 Mathematical Preliminaries (a)v1Dc 1C1 2 , (b)v2Dc 1C1 2 .1C/2, (c) 1CD1Cv3=c 1v3=c1=2 . 1. Show that for 1(andv3=c1), all three formulas reduce to vDc. 2. Compare the three velocities through terms of order 2. Note. In special relativity (with replaced by z), the ratio of observed wavelength to emitted wavelength 0is given by  0D1CzDcCv cv1=2 : 1.3.9 The relativistic sum wof two velocities uandvin the same direction is given by w cDu=cCv=c 1Cuv=c2: If v cDu cD1 ; where 0 1, findw=c in powers of through terms in 3. 1.3.10 The displacement xof a particle of rest mass m0, resulting from a constant force m0g along the x-axis, is xDc2 g8 < :" 1C gt c2#1=2 19 = ;; including relativistic effects. Find the displacement xas a power series in time t. Compare with the classical result, xD1 2gt2: 1.3.11 By use of Dirac’s relativistic theory, the fine structure formula of atomic spectroscopy is given by EDmc2 1C 2 .sCnjkj/21=2 ; where sD.jkj2 2/1=2;kD1;2;3;:::: Expand in powers of 2through order 4. 2DZe2=4 0Nhc, with Zthe atomic num- ber). This expansion is useful in comparing the predictions of the Dirac electron theory with those of a relativistic Schrödinger electron theory. Experimental results support the Dirac theory. ArfKen_Ch01-9780123846549.tex 1.3 Binomial Theorem 39 1.3.12 In a head-on proton-proton collision, the ratio of the kinetic energy in the center of mass system to the incident kinetic energy is RDTp 2mc2.EkC2mc2/2mc2U=Ek: Find the value of this ratio of kinetic energies for (a) Ekmc2(nonrelativistic), (b) Ekmc2(extreme-relativistic). ANS: (a)1 2, (b) 0. The latter answer is a sort of law of diminish- ing returns for high-energy particle accelerators (with stationary targets). 1.3.13 With binomial expansions x 1xD1X nD1xn;x x1D1 1x1D1X nD0xn: Adding these two series yieldsP1 nD1 xnD0. Hopefully, we can agree that this is nonsense, but what has gone wrong? 1.3.14 (a) Planck’s theory of quantized oscillators leads to an average energy h"iD1P nD1n"0exp.n"0=kT/ 1P nD0exp.n"0=kT/; where"0is a fixed energy. Identify the numerator and denominator as binomial expansions and show that the ratio is h"iD"0 exp." 0=kT/1: (b) Show that theh"iof part (a) reduces to kT, the classical result, for kT"0. 1.3.15 Expand by the binomial theorem and integrate term by term to obtain the Gregory series foryDtan1x(note tanyDx): tan1xDxZ 0dt 1Ct2DxZ 0f1t2Ct4t6Cg dt D1X nD0.1/nx2nC1 2nC1;1x1: 1.3.16 The Klein-Nishina formula for the scattering of photons by electrons contains a term of the form f."/D.1C"/ "22C2" 1C2"ln.1C2"/ " : ArfKen_Ch01-9780123846549.tex 40 Chapter 1 Mathematical Preliminaries Here"Dh=mc2, the ratio of the photon energy to the electron rest mass energy. Find lim "!0f."/: ANS:4 3: 1.3.17 The behavior of a neutron losing energy by colliding elastically with nuclei of mass A is described by a parameter 1, 1D1C.A1/2 2AlnA1 AC1: An approximation, good for large A, is 2D2 AC2 3: Expand1and2in powers of A1. Show that2agrees with1through.A1/2. Find the difference in the coefficients of the .A1/3term. 1.3.18 Show that each of these two integrals equals Catalan’s constant: (a)1Z 0arctan tdt t;(b)1Z 0lnxdx 1Cx2: Note. The definition and numerical computation of Catalan’s constant was addressed in Exercise 1.1.12. 1.4 M ATHEMATICAL INDUCTION We are occasionally faced with the need to establish a relation which is valid for a set of integer values, in situations where it may not initially be obvious how to proceed. However, it may be possible to show that if the relation is valid for an arbitrary value of some index n, then it is also valid if nis replaced by nC1. If we can also show that the relation is unconditionally satisfied for some initial value n0, we may then conclude (unconditionally) that the relation is also satisfied for n0C1,n0C2,:::. This method of proof is known asmathematical induction. It is ordinarily most useful when we know (or suspect) the validity of a relation, but lack a more direct method of proof. Example 1.4.1 SUM OF INTEGERS The sum of the integers from 1 through n, here denoted S.n/, is given by the formula S.n/Dn.nC1/=2. An inductive proof of this formula proceeds as follows: 1. Given the formula for S.n/, we calculate S.nC1/DS.n/C.nC1/Dn.nC1/ 2C.nC1/Dhn 2C1i .nC1/D.nC1/.nC2/ 2: Thus, given S.n/, we can establish the validity of S.nC1/. ArfKen_Ch01-9780123846549.tex 1.5 Operations on Series Expansions of Functions 41 2. It is obvious that S.1/D1.2/=2D1, so our formula for S.n/is valid for nD1. 3. The formula for S.n/is therefore valid for all integers n1.  Exercises 1.4.1 Show thatnX jD1j4Dn 30.2nC1/.nC1/.3n2C3n1/. 1.4.2 Prove the Leibniz formula for the repeated differentiation of a product: d dxn f.x/g.x/ DnX jD0n j"d dxj f.x/#"d dxnj g.x/# : 1.5 O PERATIONS ON SERIES EXPANSIONS OF FUNCTIONS There are a number of manipulations (tricks) that can be used to obtain series that represent a function or to manipulate such series to improve convergence. In addition to the proce- dures introduced in Section 1.1, there are others that to varying degrees make use of the fact that the expansion depends on a variable. A simple example of this is the expansion off.x/Dln.1Cx/, which we obtained in 1.2.4 by direct use of the Maclaurin expansion and evaluation of the derivatives of f.x/. An even easier way to obtain this series would have been to integrate the power series for 1=.1Cx/term by term from 0tox: 1 1CxD1xCx2x3C D) ln.1Cx/Dxx2 2Cx3 3x4 4C: A problem requiring somewhat more deviousness is given by the following example, in which we use the binomial theorem on a series that represents the derivative of the function whose expansion is sought. Example 1.5.1 APPLICATION OF BINOMIAL EXPANSION Sometimes the binomial expansion provides a convenient indirect route to the Maclaurin series when direct methods are difficult. We consider here the power series expansion sin1xD1X nD0.2n1/WW .2n/WWx2nC1 .2nC1/DxCx3 6C3x5 40C: (1.82) Starting from sinyDx, we find dy=dxD1=p 1x2, and write the integral sin1xDyDxZ 0dt .1t2/1=2: ArfKen_Ch01-9780123846549.tex 42 Chapter 1 Mathematical Preliminaries We now introduce the binomial expansion of .1t2/1=2and integrate term by term. The result is Eq. (1.82).  Another way of improving the convergence of a series is to multiply it by a polynomial in the variable, choosing the polynomial’s coefficients to remove the least rapidly convergent part of the resulting series. Here is a simple example of this. Example 1.5.2 MULTIPLY SERIES BY POLYNOMIAL Returning to the series for ln.1Cx/, we form .1Ca1x/ln.1Cx/D1X nD1.1/n1xn nCa11X nD1.1/n1xnC1 n DxC1X nD2.1/n11 na1 n1 xn DxC1X nD2.1/n1n.1a1/1 n.n1/xn: If we take a1D1, the nin the numerator disappears and our combined series converges as n2; the resulting series for ln.1Cx/is ln.1Cx/Dx 1Cx 11X nD1.1/n n.nC1/xn! :  Another useful trick is to employ partial fraction expansions, which may convert a seemingly difficult series into others about which more may be known. Ifg.x/andh.x/are polynomials in x, with g.x/of lower degree than h.x/, and h.x/ has the factorization h.x/D.xa1/.xa2/:::. xan/, in the case that the factors of h.x/are distinct (i.e., hhas no multiple roots), then g.x/=h.x/can be written in the form g.x/ h.x/Dc1 xa1Cc2 xa2CCcn xan: (1.83) If we wish to leave one or more quadratic factors in h.x/, perhaps to avoid the introduction of imaginary quantities, the corresponding partial-fraction term will be of the form axCb x2CpxCq: Ifh.x/has repeated linear factors, such as .xa1/m, the partial fraction expansion for this power of xa1takes the form c1;m .xa1/mCc1;m1 .xa1/m1CCc1;1 xa1: ArfKen_Ch01-9780123846549.tex 1.5 Operations on Series Expansions of Functions 43 The coefficients in partial fraction expansions are usually found easily; sometimes it is useful to express them as limits, such as ciDlimx!ai.xai/g.x/=h.x/: (1.84) Example 1.5.3 PARTIAL FRACTION EXPANSION Let f.x/Dk2 x.x2Ck2/Dc xCaxCb x2Ck2: We have written the form of the partial fraction expansion, but have not yet determined the values of a,b, and c. Putting the right side of the equation over a common denominator, we have k2 x.x2Ck2/Dc.x2Ck2/Cx.axCb/ x.x2Ck2/: Expanding the right-side numerator and equating it to the left-side numerator, we get 0.x2/C0.x/Ck2D.cCa/x2CbxCck2; which we solve by requiring the coefficient of each power of xto have the same value on both sides of this equation. We get bD0,cD1, and then aD1 . The final result is therefore f.x/D1 xx x2Ck2: (1.85)  Still more cleverness is illustrated by the following procedure, due to Euler, for changing the expansion variable so as to improve the range over which an expansion converges. Euler’s transformation, the proof of which (with hints) is deferred to Exercise 1.5.4, makes the conversion: f.x/D1X nD0.1/ncnxn(1.86) D1 1Cx1X nD0.1/nanx 1Cxn : (1.87) The coefficients anare repeated differences of the cn: a0Dc0;a1Dc1c0;a2Dc22c1Cc0;a3Dc33c2C3c1c0;:::I their general formula is anDnX jD0.1/jn j cnj: (1.88) The series to which the Euler transformation is applied need not be alternating. The coef- ficients cncan have a sign factor which cancels that in the definition. ArfKen_Ch01-9780123846549.tex 44 Chapter 1 Mathematical Preliminaries Example 1.5.4 EULER TRANSFORMATION The Maclaurin series for ln.1Cx/converges extremely slowly, with convergence only for jxj<1. We consider the Euler transformation on the related series ln.1Cx/ xD1x 2Cx2 3; (1.89) so, in Eq. (1.86), cnD1=.nC1/. The first few anare:a0D1,a1D1 21D1 2,a2D 1 321 2 C1D1 3,a3D1 431 3 C31 2 1D1 4, or in general anD.1/n nC1: The converted series is then ln.1Cx/ xD1 1Cx" 1C1 2x 1Cx C1 3x 1Cx2 C# ; which rearranges to ln.1Cx/Dx 1Cx C1 2x 1Cx2 C1 3x 1Cx3 C: (1.90) This new series converges nicely at xD1, and in fact is convergent for all x<1. Exercises 1.5.1 Using a partial fraction expansion, show that for 0<x<1, xZ xdt 1t2Dln1Cx 1x : 1.5.2 Prove the partial fraction expansion 1 n.nC1/.nCp/ D1 pWp 01 np 11 nC1Cp 21 nC2C.1/pp p1 nCp ; where pis a positive integer. Hint. Use mathematical induction. Two binomial coefficient formulas of use here are pC1 pC1jp j DpC1 j ;pC1X jD1.1/j1pC1 j D1: ArfKen_Ch01-9780123846549.tex 1.6 Some Important Series 45 1.5.3 The formula for p, Eq. (1.26), is a summation of the formP1 nD1un.p/, with un.p/D1 n.nC1/.nCp/: Applying a partial fraction decomposition to the first and last factors of the denominator, i.e., 1 n.nCp/D1 p1 n1 nCp ; show that un.p/Dun.p1/u nC1.p1/ pand thatP1 nD1un.p/D1 p pW: Hint. It is useful to note that u1.p1/D1=pW. 1.5.4 Proof of Euler transformation: By substituting Eq. (1.88) into Eq. (1.87), verify that Eq. (1.86) is recovered. Hint. It may help to rearrange the resultant double series so that both indices are summed on the range .0;1/. Then the summation not containing the coefficients cjcan be recognized as a binomial expansion. 1.5.5 Carry out the Euler transformation on the series for arctan. x/: arctan. x/Dxx3 3Cx5 5x7 7Cx9 9: Check your work by computing arctan.1/D=4 andarctan.31=2/D=6. 1.6 S OME IMPORTANT SERIES There are a few series that arise so often that all physicists should recognize them. Here is a short list that is worth committing to memory. exp.x/D1X nD0xn nWD1CxCx2 2WCx3 3WCx4 4WC;1<x<1; (1.91) sin.x/D1X nD0.1/nx2nC1 .2nC1/WDxx3 3WCx5 5Wx7 7WC;1<x<1; (1.92) cos.x/D1X nD0.1/nx2n .2n/WD1x2 2WCx4 4Wx6 6WC;1<x<1; (1.93) sinh. x/D1X nD0x2nC1 .2nC1/WDxCx3 3WCx5 5WCx7 7WC;1<x<1; (1.94) cosh. x/D1X nD0x2n .2n/WD1Cx2 2WCx4 4WCx6 6WC;1<x<1; (1.95) ArfKen_Ch01-9780123846549.tex 46 Chapter 1 Mathematical Preliminaries 1 1xD1X nD0xnD1CxCx2Cx3Cx4C;1x<1; (1.96) ln.1Cx/D1X nD1.1/n1xn nDxx2 2Cx3 3x4 4C;1<x1; (1.97) .1Cx/pD1X nD0p n xnD1X nD0.pnC1/n nWxn;1<x<1: (1.98) Reminder. The notation .a/nis the Pochhammer symbol: .a/0D1,.a/1Da, and for inte- gersn>1,.a/nDa.aC1/.aCn1/. It is not required that a, orpin Eq. (1.98), be positive or integral. Exercises 1.6.1 Show that ln1Cx 1x D2 xCx3 3Cx5 5C ;1<x<1: 1.7 V ECTORS In science and engineering we frequently encounter quantities that have algebraic magni- tude only (i.e., magnitude and possibly a sign): mass, time, and temperature. These we label scalar quantities, which remain the same no matter what coordinates we may use. In con- trast, many interesting physical quantities have magnitude and, in addition, an associated direction. This second group includes displacement, velocity, acceleration, force, momen- tum, and angular momentum. Quantities with magnitude and direction are labeled vector quantities. To distinguish vectors from scalars, we usually identify vector quantities with boldface type, as in Vorx. This section deals only with properties of vectors that are not specific to three- dimensional (3-D) space (thereby excluding the notion of the vector cross product and the use of vectors to describe rotational motion). We also restrict the present discussion to vectors that describe a physical quantity at a single point, in contrast to the situation where a vector is defined over an extended region, with its magnitude and/or direction a function of the position with which it is associated. Vectors defined over a region are called vector fields; a familiar example is the electric field, which describes the direction and magnitude of the electrical force on a test charge throughout a region of space. We return to these important topics in a later chapter. The key items of the present discussion are (1) geometric and algebraic descriptions of vectors; (2) linear combinations of vectors; and (3) the dot product of two vectors and its use in determining the angle between their directions and the decomposition of a vector into contributions in the coordinate directions. ArfKen_Ch01-9780123846549.tex 1.7 Vectors 47 Basic Properties We define a vector in a way that makes it correspond to an arrow from a starting point to another point in two-dimensional (2-D) or 3-D space, with vector addition identified as the result of placing the tail (starting point) of a second vector at the head (endpoint) of the first vector, as shown in Fig. 1.7. As seen in the figure, the result of addition is the same if the vectors are added in either order; vector addition is a commutative operation. Vector addition is also associative; if we add three vectors, the result is independent of the order in which the additions take place. Formally, this means .ACB/CCDAC.BCC/: It is also useful to define an operation in which a vector Ais multiplied by an ordinary number k(ascalar). The result will be a vector that is still in the original direction, but with its length multiplied by k. Ifkis negative, the vector’s length is multiplied by jkjbut its direction is reversed. This means we can interpret subtraction as illustrated here: ABAC.1/B; and we can form polynomials such as AC2B3C. Up to this point we are describing our vectors as quantities that do not depend on any coordinate system that we may wish to use, and we are focusing on their geometric prop- erties. For example, consider the principle of mechanics that an object will remain in static equilibrium if the vector sum of the forces on it is zero. The net force at the point Oof Fig. 1.8 will be the vector sum of the forces labeled F1,F2, and F3. The sum of the forces at static equilibrium is illustrated in the right-hand panel of the figure. It is also important to develop an algebraic description for vectors. We can do so by placing a vector Aso that its tail is at the origin of a Cartesian coordinate system and by noting the coordinates of its head. Giving these coordinates (in 3-D space) the names Ax, Ay,Az, we have a component description of A. From these components we can use the Pythagorean theorem to compute the length or magnitude ofA, denoted AorjAj, as AD.A2 xCA2 yCA2 z/1=2: (1.99) The components Ax;::: are also useful for computing the result when vectors are added or multiplied by scalars. From the geometry in Cartesian coordinates, it is obvious that if CDkACk0B, then Cwill have components CxDk AxCk0Bx;CyDk AyCk0By;CzDk AzCk0Bz: At this stage it is convenient to introduce vectors of unit length (called unit vectors) in the directions of the coordinate axes. Letting Oexbe a unit vector in the xdirection, we can B CA BA FIGURE 1.7 Addition of two vectors. ArfKen_Ch01-9780123846549.tex 48 Chapter 1 Mathematical Preliminaries F1F2 F3F1 F2 O F3wt FIGURE 1.8 Equilibrium of forces at the point O. now identify AxOexas a vector of signed magnitude Axin the xdirection, and we see that Acan be represented as the vector sum ADAxOexCAyOeyCAzOez: (1.100) IfAis itself the displacement from the origin to the point .x;y;z/, we denote it by the special symbol r(sometimes called the radius vector), and Eq. (1.100) becomes rDxOexCyOeyCzOez: (1.101) The unit vectors are said to span the space in which our vectors reside, or to form a basis for the space. Either of these statements means that any vector in the space can be constructed as a linear combination of the basis vectors. Since a vector Ahas specific values of Ax,Ay, and Az, this linear combination will be unique. Sometimes a vector will be specified by its magnitude Aand by the angles it makes with the Cartesian coordinate axes. Letting , , be the respective angles our vector makes with the x,y, and zaxes, the components of Aare given by AxDAcos ; AyDAcos ; AzDAcos : (1.102) The quantities cos ,cos ,cos (seeFig. 1.9) are known as the direction cosines ofA. Since we already know that A2 xCA2 yCA2 zDA2, we see that the direction cosines are not entirely independent, but must satisfy the relation cos2 Ccos2 Ccos2 D1: (1.103) While the formalism of Eq. (1.100) could be developed with complex values for the components Ax,Ay,Az, the geometric situation being described makes it natural to restrict these coefficients to real values; the space with all possible real values of two coordinates ArfKen_Ch01-9780123846549.tex 1.7 Vectors 49 Az Ay (Ax, Ay, 0)(Ax, Ay, Az) A y Ax xz γ β α FIGURE 1.9 Cartesian components and direction cosines of A. y AxexxA ^Ayey^ FIGURE 1.10 Projections of Aon the xandyaxes. is denoted by mathematicians (and occasionally by us) I R2; the complete 3-D space is named I R3. Dot (Scalar) Product When we write a vector in terms of its component vectors in the coordinate directions, as in ADAxOexCAyOeyCAzOez; we can think of AxOexas its projection in the xdirection. Stated another way, it is the portion of Athat is in the subspace spanned by Oexalone. The term projection corresponds to the idea that it is the result of collapsing (projecting) a vector onto one of the coordinate axes. See Fig. 1.10. It is useful to define a quantity known as the dot product, with the property that it produces the coefficients, e.g., Ax, in projections onto the coordinate axes according to AOexDAxDAcos ; AOeyDAyDAcos ; AOezDAzDAcos ; (1.104) where cos ,cos ,cos are the direction cosines of A. ArfKen_Ch01-9780123846549.tex 50 Chapter 1 Mathematical Preliminaries We want to generalize the notion of the dot product so that it will apply to arbitrary vectors AandB, requiring that it, like projections, be linear and obey the distributive and associative laws A.BCC/DABCAC; (1.105) A.kB/D.kA/BDkAB; (1.106) with ka scalar. Now we can use the decomposition of Binto Cartesian components as in Eq. (1.100), BDBxOexCByOeyCBzOez, to construct the dot product of the vectors Aand Bas ABDA.BxOexCByOeyCBzOez/ DBxAOexCByAOeyCBzAOez DBxAxCByAyCBzAz: (1.107) This leads to the general formula ABDX iBiAiDX iAiBiDBA; (1.108) which is also applicable when the number of dimensions in the space is other than three. Note that the dot product is commutative, with ABDBA. An important property of the dot product is that AAis the square of the magnitude ofA: AADA2 xCA2 yCDjAj2: (1.109) Applying this observation to CDACB, we have jCj2DCCD.ACB/.ACB/DAACBBC2AB; which can be rearranged to ABD1 2h jCj2jAj2jBj2i : (1.110) From the geometry of the vector sum CDACB, as shown in Fig. 1.11, and recalling the law of cosines and its similarity to Eq. (1.110), we obtain the well-known formula ABDjAjjBj cos; (1.111) y xB AC θ FIGURE 1.11 Vector sum, CDACB. ArfKen_Ch01-9780123846549.tex 1.7 Vectors 51 whereis the angle between the directions of AandB. In contrast with the algebraic formula Eq. (1.108), Eq. (1.111) is ageometric formula for the dot product, and shows clearly that it depends only on the relative directions of AandBand is therefore indepen- dent of the coordinate system. For that reason the dot product is sometimes also identified as ascalar product. Equation (1.111) also permits an interpretation in terms of the projection of a vector A in the direction of Bor the reverse. IfObis a unit vector in the direction of B, the projection ofAin that direction is given by AbObD.ObA/ObD.Acos/Ob; (1.112) whereis the angle between AandB. Moreover, the dot product ABcan then be identi- fied asjBjtimes the magnitude of the projection of Ain the Bdirection, so ABDAbB. Equivalently, ABis equal tojAjtimes the magnitude of the projection of Bin the A direction, so we also have ABDBaA. Finally, we observe that since jcosj1,Eq. (1.111) leads to the inequality jABjjAjjBj: (1.113) The equality in Eq. (1.113) holds only if AandBare collinear (in either the same or opposite directions). This is the specialization to physical space of the Schwarz inequality, which we will later develop in a more general context. Orthogonality Equation (1.111) shows that ABbecomes zero when cosD0, which occurs at D=2 (i.e., atD90). These values of correspond to AandBbeing perpendicular, the technical term for which is orthogonal. Thus, AandBare orthogonal if and only if ABD0. Checking this result for two dimensions, we note that AandBare perpendicular if the slope of B,By=Bx, is the negative of the reciprocal of Ay=Ax, or By BxDAx Ay: This result expands to AxBxCAyByD0, the condition that AandBbe orthogonal. In terms of projections, ABD0means that the projection of Ain the Bdirection vanishes (and vice versa). That is of course just another way of saying that AandBare orthogonal. The fact that the Cartesian unit vectors are mutually orthogonal makes it possible to simplify many dot product computations. Because OexOeyDOexOezDOeyOezD0;OexOexDOeyOeyDOezOezD1; (1.114) ArfKen_Ch01-9780123846549.tex 52 Chapter 1 Mathematical Preliminaries we can evaluate ABas .AxOexCAyOeyCAzOez/.BxOexCByOeyCBzOez/DAxBxOexOexCAyByOeyOeyCAzBzOezOez C.AxByCAyBx/OexOeyC.AxBzCAzBx/OexOezC.AyBzCAzBy/OeyOez DAxBxCAyByCAzBz: See Chapter 3: Vector Analysis, Section 3.2: Vectors in 3-D Space for an introduction of the cross product of vectors, needed early in Chapter 2. Exercises 1.7.1 The vector Awhose magnitude is 1:732 units makes equal angles with the coordinate axes. Find Ax;Ay, and Az. 1.7.2 A triangle is defined by the vertices of three vectors A;BandCthat extend from the origin. In terms of A;B, and Cshow that the vector sum of the successive sides of the triangle.ABCBCCC A/is zero, where the side ABis from AtoB;etc. 1.7.3 A sphere of radius ais centered at a point r1. (a) Write out the algebraic equation for the sphere. (b) Write out a vector equation for the sphere. ANS. (a).xx1/2C.yy1/2C.zz1/2Da2. (b) rDr1Ca, where atakes on all directions but has a fixed magnitude a. 1.7.4 Hubble’s law. Hubble found that distant galaxies are receding with a velocity propor- tional to their distance from where we are on Earth. For the ith galaxy, viDH0ri with us at the origin. Show that this recession of the galaxies from us does notimply that we are at the center of the universe. Specifically, take the galaxy at r1as a new origin and show that Hubble’s law is still obeyed. 1.7.5 Find the diagonal vectors of a unit cube with one corner at the origin and its three sides lying along Cartesian coordinates axes. Show that there are four diagonals with lengthp 3:Representing these as vectors, what are their components? Show that the diagonals of the cube’s faces have lengthp 2and determine their components. 1.7.6 The vector r, starting at the origin, terminates at and specifies the point in space .x;y;z/. Find the surface swept out by the tip of rif (a).ra/aD0:Characterize ageometrically. (b).ra/rD0:Describe the geometric role of a: The vector ais constant (in magnitude and direction). ArfKen_Ch01-9780123846549.tex 1.8 Complex Numbers and Functions 53 1.7.7 A pipe comes diagonally down the south wall of a building, making an angle of 45with the horizontal. Coming into a corner, the pipe turns and continues diagonally down a west-facing wall, still making an angle of 45with the horizontal. What is the angle between the south-wall and west-wall sections of the pipe? ANS. 120. 1.7.8 Find the shortest distance of an observer at the point .2;1;3/from a rocket in free flight with velocity .1;2;3/km/s. The rocket was launched at time tD0from.1;1;1/: Lengths are in kilometers. 1.7.9 Show that the medians of a triangle intersect in the center which is 2=3of the median’s length from each vertex. Construct a numerical example and plot it. 1.7.10 Prove the law of cosines starting from A2D.BC/2. 1.7.11 Given the three vectors, PD3OexC2OeyOez; QD6Oex4OeyC2Oez; RDOex2OeyOez; find two that are perpendicular and two that are parallel or antiparallel. 1.8 C OMPLEX NUMBERS AND FUNCTIONS Complex numbers and analysis based on complex variable theory have become extremely important and valuable tools for the mathematical analysis of physical theory. Though the results of the measurement of physical quantities must, we firmly believe, ultimately be described by real numbers, there is ample evidence that successful theories predicting the results of those measurements require the use of complex numbers and analysis. In a later chapter we explore the fundamentals of complex variable theory. Here we introduce complex numbers and identify some of their more elementary properties. Basic Properties A complex number is nothing more than an ordered pair of two real numbers, .a;b/. Sim- ilarly, a complex variable is an ordered pair of two real variables, z.x;y/: (1.115) The ordering is significant. In general .a;b/is not equal to .b;a/and.x;y/is not equal to.y;x/. As usual, we continue writing a real number .x;0/simply as x, and we call i.0;1/the imaginary unit. All of complex analysis can be developed in terms of ordered pairs of numbers, variables, and functions .u.x;y/;v.x;y//. We now define addition of complex numbers in terms of their Cartesian components as z1Cz2D.x1;y1/C.x2;y2/D.x1Cx2;y1Cy2/: (1.116) ArfKen_Ch01-9780123846549.tex 54 Chapter 1 Mathematical Preliminaries Multiplication of complex numbers is defined as z1z2D.x1;y1/.x2;y2/D.x1x2y1y2;x1y2Cx2y1/: (1.117) It is obvious that multiplication is not just the multiplication of corresponding components. Using Eq. (1.117) we verify that i2D.0;1/.0;1/D.1; 0/D1 , so we can also identify iDp1as usual, and further rewrite Eq. (1.115) as zD.x;y/D.x;0/C.0;y/DxC.0;1/.y;0/DxCiy: (1.118) Clearly, introduction of the symbol iis not necessary here, but it is convenient, in large part because the addition and multiplication rules for complex numbers are consistent with those for ordinary arithmetic with the additional property that i2D1 : .x1Ciy1/.x2Ciy2/Dx1x2Ci2y1y2Ci.x1y2Cy1x2/D.x1x2y1y2/Ci.x1y2Cy1x2/; in agreement with Eq. (1.117). For historical reasons, iand its multiples are known as imaginary numbers. The space of complex numbers, sometimes denoted Zby mathematicians, has the fol- lowing formal properties: It is closed under addition and multiplication, meaning that if two complex numbers are added or multiplied, the result is also a complex number. It has a unique zero number, which when added to any complex number leaves it unchanged and which, when multiplied with any complex number yields zero. It has a unique unit number, 1, which when multiplied with any complex number leaves it unchanged. Every complex number zhas an inverse under addition (known as z), and every nonzero zhas an inverse under multiplication, denoted z1or1=z. It is closed under exponentiation: if uandvare complex numbers uvis also a complex number. From a rigorous mathematical viewpoint, the last statement above is somewhat loose, as it does not really define exponentiation, but we will find it adequate for our purposes. Some additional definitions and properties include the following: Complex conjugation: Like all complex numbers, ihas an inverse under addition, denotedi, in two-component form, .0;1/. Given a complex number zDxCiy, it is useful to define another complex number, zDxiy, which we call the complex con- jugate ofz.6Forming zzD.xCiy/.xiy/Dx2Cy2; (1.119) we see that zzis real; we define the absolute value of z, denotedjzj, aspzz. 6The complex conjugate of zis often denoted zin the mathematical literature. ArfKen_Ch01-9780123846549.tex 1.8 Complex Numbers and Functions 55 Division: Consider now the division of two complex numbers: z0=z. We need to manipulate this quantity to bring it to the complex number form uCiv(with uandvreal). We may do so as follows: z0 zDz0z zzD.x0Ciy0/.xiy/ x2Cy2; or x0Ciy0 xCiyDxx0Cyy0 x2Cy2Cixy0x0y x2Cy2: (1.120) Functions in the Complex Domain Since the fundamental operations in the complex domain obey the same rules as those for arithmetic in the space of real numbers, it is natural to define functions so that their real and complex incarnations are similar, and specifically so that the complex and real definitions agree when both are applicable. This means, among other things, that if a function is repre- sented by a power series, we should, within the region of convergence of the power series, be able to use such series with complex values of the expansion variable. This notion is called permanence of the algebraic form. Applying this concept to the exponential, we define ezD1CzC1 2Wz2C1 3Wz3C1 4Wz4C: (1.121) Now, replacing zbyiz, we have eizD1CizC1 2W.iz/2C1 3W.iz/3C1 4W.iz/4C D 11 2Wz2C1 4Wz4 Ci z1 3Wz3C1 5Wz5 : (1.122) It was permissible to regroup the terms in the series of Eq. (1.122) because that series is absolutely convergent for all z; the d’Alembert ratio test succeeds for all z, real or complex. If we now identify the bracketed expansions in the last line of Eq. (1.122) ascoszandsinz, we have the extremely valuable result eizDcoszCisinz: (1.123) This result is valid for all z, real, imaginary, or complex, but is particularly useful when z is real. Any function w.z/of a complex variable zDxCiycan in principle be divided into its real and imaginary parts, just as we did when we added, multiplied, or divided complex numbers. That is, we can write w.z/Du.x;y/Civ.x;y/; (1.124) in which the separate functions u.x;y/andv.x;y/are pure real. For example, if f.z/Dz2, we have f.z/D.zCiy/2D.x2y2/Ci.2xy/: ArfKen_Ch01-9780123846549.tex 56 Chapter 1 Mathematical Preliminaries The real part of a function f.z/will be labeled Ref.z/, whereas the imaginary part will be labeled Imf.z/. In Eq. (1.124), Rew.z/Du.x;y/;Imw.z/Dv.x;y/: The complex conjugate of our function w.z/isu.x;y/iv.x;y/, and depending on w, may or may not be equal to w.z/. Polar Representation We may visualize complex numbers by assigning them locations on a planar graph, called anArgand diagram or, more colloquially, the complex plane. Traditionally the real com- ponent is plotted horizontally, on what is called the real axis, with the imaginary axis in the vertical direction. See Fig. 1.12. An alternative to identifying points by their Cartesian coordinates.x;y/is to use polar coordinates .r;/, with xDrcos;yDrsin; orrDq x2Cy2; Dtan1y=x: (1.125) The arctan function tan1.y=x/is multiple valued; the correct location on an Argand dia- gram needs to be consistent with the individual values of xandy. The Cartesian and polar representations of a complex number can also be related by writing xCiyDr.cosCisin/Drei; (1.126) where we have used Eq. (1.123) to introduce the complex exponential. Note that ris alsojzj, so the magnitude of zis given by its distance from the origin in an Argand di- agram. In complex variable theory, ris also called the modulus ofzandis termed the argument or the phase ofz. If we have two complex numbers, zandz0, in polar form, their product zz0can be written zz0D.rei/.r0ei0/D.rr0/ei.C0/; (1.127) showing that the location of the product in an Argand diagram will have argument (polar angle) at the sum of the polar angles of the factors, and with a magnitude that is the product xyz r θJm Re FIGURE 1.12 Argand diagram, showing location of zDxCiyDrei. ArfKen_Ch01-9780123846549.tex 1.8 Complex Numbers and Functions 57 Re RezJm Jm z−z* z−z* z+z*z* z*θ θ FIGURE 1.13 Left: Relation of zandz. Right: zCzandzz. of their magnitudes. Conversely, the quotient z=z0will have magnitude r=r0and argument 0. These relationships should aid in getting a qualitative understanding of complex multiplication and division. This discussion also shows that multiplication and division are easier in the polar representation, whereas addition and subtraction have simpler forms in Cartesian coordinates. The plotting of complex numbers on an Argand diagram makes obvious some other properties. Since addition on an Argand diagram is analogous to 2-D vector addition, it can be seen that jzjjz0j jzz0jjzjCjz0j: (1.128) Also, since zDreihas the same magnitude as zbut an argument that differs only in sign, zCzwill be real and equal to 2Rez, while zzwill be pure imaginary and equal to2iImz. See Fig. 1.13 for an illustration of this discussion. We can use an Argand diagram to plot values of a function w.z/as well as just zitself, in which case we could label the axes uandv, referring to the real and imaginary parts of w. In that case, we can think of the function w.z/as providing a mapping from the xy plane to the uvplane, with the effect that any curve in the xy(sometimes called z) plane is mapped into a corresponding curve in the uv(Dw) plane. In addition, the statements of the preceding paragraph can be extended to functions: jw.z/jjw0.z/j jw. z/w0.z/jjw. z/jCjw0.z/j; Rew.z/Dw.z/CTw. z/U 2;Imw.z/Dw.z/Tw. z/U 2: (1.129) Complex Numbers of Unit Magnitude Complex numbers of the form eiDcosCisin; (1.130) where we have given the variable the name to emphasize the fact that we plan to restrict it to real values, correspond on an Argand diagram to points for which xDcos,yDsin, ArfKen_Ch01-9780123846549.tex 58 Chapter 1 Mathematical Preliminaries Jm Rexy 1i eiθ θθ=π −i−1 FIGURE 1.14 Some values of zon the unit circle. and whose magnitude is therefore cos2Csin2D1. The points exp.i/therefore lie on the unit circle, at polar angle . This observation makes obvious a number of relations that could in principle also be deduced from Eq. (1.130). For example, if has the special values=2,, or3=2 , we have the interesting relationships ei=2Di;eiD1; e3i=2Di: (1.131) We also see that exp.i/is periodic, with period 2, so e2iDe4iDD 1; e3i=2Dei=2Di;etc. (1.132) A few relevant values of zon the unit circle are illustrated in Fig. 1.14. These relation- ships cause the real part of exp.i!t/to describe oscillation at angular frequency !, with exp.iT!tCU/describing an oscillation displaced from that first mentioned by a phase difference. Circular and Hyperbolic Functions The relationship encapsulated in Eq. (1.130) enables us to obtain convenient formulas for the sine and cosine. Taking the sum and difference of exp.C i/andexp. i/, we have cosDeiCei 2;sinDeiei 2i: (1.133) These formulas place the definitions of the hyperbolic functions in perspective: coshDeCe 2;sinhDee 2: (1.134) Comparing these two sets of equations, it is possible to establish the formulas cosh izDcosz;sinhizDisinz: (1.135) Proof is left to Exercise 1.8.5. The fact that exp.in/can be written in the two equivalent forms cosnCisinnD.cosCisin/n(1.136) ArfKen_Ch01-9780123846549.tex 1.8 Complex Numbers and Functions 59 establishes a relationship known as de Moivre’s Theorem. By expanding the right mem- ber of Eq. (1.136), we easily obtain trigonometric multiple-angle formulas, of which the simplest examples are the well-known results sin.2/D2 sincos;cos.2/Dcos2sin2: If we solve the sinformula of Eq. (1.133) forexp.i/, we get (choosing the plus sign for the radical) eiDisinCq 1sin2: Setting sinDzandDsin1.z/, and taking the logarithm of both sides of the above equation, we express the inverse trigonometric function in terms of logarithms. sin1.z/Dilnh izCp 1z2i : The set of formulas that can be generated in this way includes: sin1.z/Dilnh izCp 1z2i ;tan1.z/Di 2h ln.1iz/ln.1Ciz/i ; sinh1.z/Dlnh zCp 1Cz2i ;tanh1.z/D1 2h ln.1Cz/ln.1z/i :(1.137) Powers and Roots The polar form is very convenient for expressing powers and roots of complex numbers. For integer powers, the result is obvious and unique: zDrei';znDrnein': For roots (fractional powers), we also have zDrei';z1=nDr1=nei'=n; but the result is not unique. If we write zin the alternate but equivalent form zDrei.'C2m/; where mis an integer, we now get additional values for the root: z1=nDr1=nei.'C2m/=n;(any integer m). IfnD2(corresponding to the square root), different choices of mwill lead to two distinct values of z1=2, both of the same modulus but differing in argument by . This corresponds to the well-known result that the square root is double-valued and can be written with either sign. In general, z1=nisn-valued, with successive values having arguments that differ by 2=n.Figure 1.15 illustrates the multiple values of 11=3,i1=3, and.1/1=3. ArfKen_Ch01-9780123846549.tex 60 Chapter 1 Mathematical Preliminaries (a) (b) (c)−1 −i1(1+ 3i)1 2 (−1+3i)1 2 (−1−3i)1 2(1− 3i)1 2(3+i)1 2 (3+i)1 2− FIGURE 1.15 Cube roots: (a) 11=3; (b)i1=3; (c).1/1=3. Logarithm Another multivalued complex function is the logarithm, which in the polar representation takes the form lnzDln.rei/DlnrCi: However, it is also true that lnzDln rei.C2n/ DlnrCi.C2n/; (1.138) foranypositive or negative integer n. Thus, lnzhas, for a given z, the infinite number of values corresponding to all possible choices of nin Eq. (1.138). Exercises 1.8.1 Find the reciprocal of xCiy, working in polar form but expressing the final result in Cartesian form. 1.8.2 Show that complex numbers have square roots and that the square roots are contained in the complex plane. What are the square roots of i? 1.8.3 Show that (a) cosnDcosnn 2 cosn2sin2Cn 4 cosn4sin4 , (b) sinnDn 1 cosn1sinn 3 cosn3sin3C . 1.8.4 Prove that (a)N1X nD0cosnxDsin.N x=2/ sinx=2cos.N1/x 2; (b)N1X nD0sinnxDsin.N x=2/ sinx=2sin.N1/x 2: These series occur in the analysis of the multiple-slit diffraction pattern. ArfKen_Ch01-9780123846549.tex 1.8 Complex Numbers and Functions 61 1.8.5 Assume that the trigonometric functions and the hyperbolic functions are defined for complex argument by the appropriate power series. Show that isinzDsinhiz;sinizDisinhz; coszDcosh iz;cosizDcosh z: 1.8.6 Using the identities coszDeizCeiz 2;sinzDeizeiz 2i; established from comparison of power series, show that (a) sin.xCiy/Dsinxcosh yCicosxsinhy; cos.xCiy/Dcosxcosh yisinxsinhy, (b)jsinzj2Dsin2xCsinh2y,jcoszj2Dcos2xCsinh2y. This demonstrates that we may have jsinzj;jcoszj>1in the complex plane. 1.8.7 From the identities in Exercises 1.8.5 and1.8.6 show that (a) sinh. xCiy/DsinhxcosyCicosh xsiny, cosh. xCiy/Dcosh xcosyCisinhxsiny, (b)jsinhzj2Dsinh2xCsin2y,jcosh zj2Dcosh2xCsin2y. 1.8.8 Show that (a) tanhz 2DsinhxCisiny cosh xCcosy; (b) cothz 2Dsinhxisiny cosh xcosy: 1.8.9 By comparing series expansions, show that tan1xDi 2ln1ix 1Cix . 1.8.10 Find the Cartesian form for all values of (a).8/1=3; (b) i1=4; (c) ei=4: 1.8.11 Find the polar form for all values of (a).1Ci/3; (b).1/1=5: ArfKen_Ch01-9780123846549.tex 62 Chapter 1 Mathematical Preliminaries 1.9 D ERIVATIVES AND EXTREMA We recall the familiar limit identified as the derivative, d f.x/=dx , of a function f.x/at a point x: d f.x/ dxDlim "D0f.xC"/f.x/ "I (1.139) the derivative is only defined if the limit exists and is independent of the direction from which"approaches zero. The variation ordifferential off.x/associated with a change dxin its independent variable from the reference value xassumes the form d fDf.xCdx/f.x/Dd f dxdx; (1.140) in the limit that dxis small enough that terms dependent on dx2and higher powers of dx become negligible. The mean value theorem (based on the continuity of f) tells us that here, d f=dxis evaluated at some point between xandxCdx, but as dx!0,!x. When a quantity of interest is a function of two or more independent variables, the generalization of Eq. (1.140) is (illustrating for the physically important three-variable case): d fDh f.xCdx;yCdy;zCdz/f.x;yCdy;zCdz/i Ch .f.x;yCdy;zCdz/f.x;y;zCdz/i Ch f.x;y;zCdz/f.x;y;z/i D@f @xdxC@f @ydyC@f @zdz; (1.141) where the partial derivatives indicate differentiation in which the independent variables not being differentiated are kept fixed. The fact that @f=@xis evaluated at yCdyand zCdzinstead of at yandzalters the derivative by amounts that are of order dyand dz, and therefore the change becomes negligible in the limit of small variations. It is thus consistent to interpret Eq. (1.141) as involving partial derivatives that are all evaluated at the reference point x;y;z. Further analysis of the same sort as led to Eq. (1.141) can be used to define higher derivatives and to establish the useful result that cross derivatives (e.g.,@2=@x@y) are independent of the order in which the differentiations are performed: @ @y@f @x @2f @y@xD@2f @x@y: (1.142) Sometimes it is not clear from the context which variables other than that being dif- ferentiated are independent, and it is then advisable to attach subscripts to the derivative notation to avoid ambiguity. For example, if x,y, and zhave been defined in a problem, but only two of them are independent, one might write @f @x yor@f @x z; whichever is actually meant. ArfKen_Ch01-9780123846549.tex 1.9 Derivatives and Extrema 63 For working with functions of several variables, we note two useful formulas that follow from Eq. (1.141): 1. The chain rule, d f dsD@f @xdx dsC@f @ydy dsC@f @zdz ds; (1.143) which applies when x,y, and zare functions of another variable, s, 2. A formula obtained by setting d fD0(here shown for the case where there are only two independent variables and the dzterm of Eq. (1.141) is absent): @y @x fD@f @x y@f @y x: (1.144) In Lagrangian mechanics, one occasionally encounters expressions such as7 d dtL.x;Px;t/D@L @xPxC@L @PxRxC@L @t ; an example of use of the chain rule. Here it is necessary to distinguish between the formal dependence of Lon its three arguments and the overall dependence of Lon time. Note the use of the ordinary ( d=dt) and partial ( @=@t) derivative notation. Stationary Points Whether or not a set of independent variables (e.g., x,y,zof our previous discussion) represents directions in space, one can ask how a function fchanges if we move in various directions in the space of the independent variables; the answer is provided by Eq. (1.143), where the “direction” is defined by the values of dx=ds,dy=ds, etc. It is often desired to find the minimum of a function fofnvariables xi,iD1;:::; n, and a necessary but not sufficient condition on its position is that d f dsD0for all directions of ds. This is equivalent to requiring @f @xiD0; iD1;:::; n: (1.145) All points in thefxigspace that satisfy Eq. (1.145) are termed stationary; for a stationary point of fto be a minimum, it is also necessary that the second derivatives d2f=ds2be positive for all directions of s. Conversely, if the second derivatives in all directions are negative, the stationary point is a maximum. If neither of these conditions are satisfied, the stationary point is neither a maximum nor a minimum, and is often called a saddle point because of the appearance of the surface of fwhen there are two independent variables 7Here dots indicate time derivatives. ArfKen_Ch01-9780123846549.tex 64 Chapter 1 Mathematical Preliminaries f(x, y) y x FIGURE 1.16 A stationary point that is neither a maximum nor minimum (a saddle point). (see Fig. 1.16). It is often obvious whether a stationary point is a minimum or maximum, but a complete discussion of the issue is nontrivial. Exercises 1.9.1 Derive the following formula for the Maclaurin expansion of a function of two variables: f.x;y/Df.0;0/Cx@f @xCy@f @y C1 2W2 0 x2@2f @x2C2 1 xy@2f @x@yC2 2 y2@2f @y2 C1 3W3 0 x3@3f @x3C3 1 x2y@3f @x2@yC3 2 xy2@3f @x@y2C3 3 y3@3f @y3 C; where all the partial derivatives are to be evaluated at the point .0;0/. 1.9.2 The result in Exercise 1.9.1 can be generalized to larger numbers of independent vari- ables. Prove that for an m-variable system, the Maclaurin expansion can be written in ArfKen_Ch01-9780123846549.tex 1.10 Evaluation of Integrals 65 the symbolic form f.x1;:::; xm/D1X nD0tn nW mX iD1 i@ @xi!n f.0;:::; 0/; where in the right-hand side we have made the substitutions xjD jt. 1.10 E VALUATION OF INTEGRALS Proficiency in the evaluation of integrals involves a mixture of experience, skill in pat- tern recognition, and a few tricks. The most familiar include the technique of integration by parts, and the strategy of changing the variable of integration. We review here some methods for integrals in one and multiple dimensions. Integration by Parts The technique of integration by parts is part of every elementary calculus course, but its use is so frequent and ubiquitous that it bears inclusion here. It is based on the obvious relation, for uandvarbitrary functions of x, d.uv/Du dvCvdu: Integrating both sides of this equation over an interval .a;b/, we reach uv b aDbZ au dvCbZ avdu; which is usually rearranged to the well-known form bZ au dvDuv b abZ avdu: (1.146) Example 1.10.1 INTEGRATION BY PARTS Consider the integralZb axsinx dx. We identify uDxanddvDsinx dx. Differentiating and integrating, we find duDdxandvDcosx, so Eq. (1.146) becomes bZ axsinx dxD.x/.cosx/ b abZ a.cosx/dxDacosabcosbCsinbsina:  The key to the effective use of this technique is to see how to partition an integrand into uanddvin a way that makes it easy to form duandvand also to integrateR vdu. ArfKen_Ch01-9780123846549.tex 66 Chapter 1 Mathematical Preliminaries Special Functions A number of special functions have become important in physics because they arise in fre- quently encountered situations. Identifying a one-dimensional (1-D) integral as one yield- ing a special function is almost as good as a straight-out evaluation, in part because it prevents the waste of time that otherwise might be spent trying to carry out the integration. But of perhaps more importance, it connects the integral to the full body of knowledge regarding its properties and evaluation. It is not necessary for every physicist to know everything about all known special functions, but it is desirable to have an overview per- mitting the recognition of special functions which can then be studied in more detail if necessary. It is common for a special function to be defined in terms of an integral over the range for which that integral converges, but to have its definition extended to a larger domain Table 1.2 Special Functions of Importance in Physics Gamma function 0.x/D1Z 0tx1etdt See Chap. 13. Factorial (n integral) nWD1Z 0tnetdt nWD0.nC1/ Riemann zeta function .x/D1 0.x/1Z 0tx1dt et1See Chaps. 1 and 12. Exponential integrals En.x/D1Z 1tnetdt E 1.x/Ei. x/ Sine integral si.x/D1Z xsint tdt Cosine integral Ci. x/D1Z xcost tdt Error functions erf. x/D2pxZ 0et2dt erf.1/D1 erfc.x/D2p1Z xet2dt erfc.x/D1erf.x/ Dilogarithm Li2.x/DxZ 0ln.1t/ tdt ArfKen_Ch01-9780123846549.tex 1.10 Evaluation of Integrals 67 by analytic continuation in the complex plane (cf. Chapter 11) or by the establishment of suitable functional relations. We present in Table 1.2 only the most useful integral repre- sentations of a few functions of frequent occurrence. More detail is provided by a variety of on-line sources and in material listed under Additional Readings at the end of this chapter, particularly the compilations by Abramowitz and Stegun and by Gradshteyn and Ryzhik. A conspicuous omission from the list in Table 1.2 is the extensive family of Bessel func- tions. A short table cannot suffice to summarize their numerous integral representations; a survey of this topic is in Chapter 14. Other important functions in more than one variable or with indices in addition to arguments have also been omitted from the table. Other Methods An extremely powerful method for the evaluation of definite integrals is that of contour integration in the complex plane. This method is presented in Chapter 11 and will not be discussed here. Integrals can often be evaluated by methods that involve integration or differentiation with respect to parameters, thereby obtaining relations between known integrals and those whose values are being sought. Example 1.10.2 DIFFERENTIATE PARAMETER We wish to evaluate the integral ID1Z 0ex2 x2Ca2dx: We introduce a parameter, t, to facilitate further manipulations, and consider the related integral J.t/D1Z 0et.x2Ca2/ x2Ca2dxI we note that IDea2J.1/. We now differentiate J.t/with respect to tand evaluate the resulting integral, which is a scaled version of Eq. (1.148): d J.t/ dtD1Z 0et.x2Ca2/dxDeta21Z 0etx2dxD1 2r teta2: (1.147) To recover J.t/we integrate Eq. (1.147) between tand1, making use of the fact that J.1/D0. To carry out the integration it is convenient to make the substitution u2Da2t, ArfKen_Ch01-9780123846549.tex 68 Chapter 1 Mathematical Preliminaries so we get J.t/Dp 21Z teta2 t1=2dtDp a1Z at1=2eu2du; which we now recognize as J.t/D.=2 a/erfc. at1=2/. Thus, our final result is ID 2aea2erfc.a/:  Many integrals can be evaluated by first converting them into infinite series, then manipulating the resulting series, and finally either evaluating the series or recognizing it as a special function. Example 1.10.3 EXPAND, THEN INTEGRATE Consider IDR1 0dx xln 1Cx 1x :Using Eq. (1.120) for the logarithm, ID1Z 0dx2 1Cx2 3Cx4 5C D2 1C1 32C1 52C : Noting that 1 22.2/D1 22C1 42C1 62C; we see that .2/1 4.2/D1C1 32C1 52C; soID3 2.2/.  Simply using complex numbers aids in the evaluation of some integrals. Take, for example, the elementary integral IDZdx 1Cx2: Making a partial fraction decomposition of .1Cx2/1and integrating, we easily get IDZ1 21 1CixC1 1ix dxDi 2h ln.1ix/ln.1Cix/i : From Eq. (1.137), we recognize this as tan1.x/. The complex exponential forms of the trigonometric functions provide interesting approaches to the evaluation of certain integrals. Here is an example. ArfKen_Ch01-9780123846549.tex 1.10 Evaluation of Integrals 69 Example 1.10.4 A TRIGONOMETRIC INTEGRAL Consider ID1Z 0eatcosbt dt; where aandbare real and positive. Because cosbtDReeibt, we note that IDRe1Z 0e.aCib/tdt: The integral is now just that of an exponential, and is easily evaluated, leading to IDRe1 aibDReaCib a2Cb2; which yields IDa=.a2Cb2/. As a bonus, the imaginary part of the same integral gives us 1Z 0eatsinbt dtDb a2Cb2:  Recursive methods are often useful in obtaining formulas for a set of related integrals. Example 1.10.5 RECURSION Consider InD1Z 0tnsint dt for positive integer n. Integrating Inby parts twice, taking uDtnanddvDsint dt, we have InD1 n.n1/ 2In2; with starting values I0D2= andI1D1=. There is often no practical need to obtain a general, nonrecursive formula, as repeated application of the recursion is frequently more efficient that a closed formula, even when one can be found.  ArfKen_Ch01-9780123846549.tex 70 Chapter 1 Mathematical Preliminaries Multiple Integrals An expression that corresponds to integration in two variables, say xandy, may be written with two integral signs, as in x f.x;y/dxdy orx2Z x1dxy2.x/Z y1.x/dy f.x;y/; where the right-hand form can be more specific as to the integration limits, and also gives an explicit indication that the yintegration is to be performed first, or with a single integral sign, as in Z Sf.x;y/d A; where S(if explicitly shown) is a 2-D integration region and d Ais an element of “area” (in Cartesian coordinates, equal to dxdy ). In this form we are leaving open both the choice of coordinate system to be used for evaluating the integral, and the order in which the variables are to be integrated. In three dimensions, we may either use three integral signs or a single integral with a symbol dindicating a 3-D “volume” element in an unspecified coordinate system. In addition to the techniques available for integration in a single variable, multiple in- tegrals provide further opportunities for evaluation based on changes in the order of inte- gration and in the coordinate system used in the integral. Sometimes simply reversing the order of integration may be helpful. If, before the reversal, the range of the inner integral depends on the outer integration variable, care must be exercised in determining the inte- gration ranges after reversal. It may be helpful to draw a diagram identifying the range of integration. Example 1.10.6 REVERSING INTEGRATION ORDER Consider 1Z 0erdr1Z res sds; in which the inner integral can be identified as an exponential integral, suggesting diffi- culty if the integration is approached in a straightforward manner. Suppose we proceed by reversing the order of integration. To identify the proper coordinate ranges, we draw on a.r;s/plane, as in Fig. 1.17, the region s>r0, which is covered in the original inte- gration order as a succession of vertical strips, for each rextending from sDrtosD1 . See the left-hand panel of the figure. If the outer integration is changed from rtos, this same region is covered by taking, for each s, a horizontal range of rthat runs from rD0to rDs. See the right-hand panel of the figure. The transformed double integral then assumes ArfKen_Ch01-9780123846549.tex 1.10 Evaluation of Integrals 71 s rs r FIGURE 1.17 2-D integration region for Example 1.10.6. Left panel: inner integration over s; right panel: inner integration over r. the form 1Z 0es sdssZ 0erdr; where the inner integral over ris now elementary, evaluating to 1es. This leaves us with a 1-D integral, 1Z 0es s.1es/ds: Introducing a power series expansion for 1es, this integral becomes 1Z 0es s1X nD1.1/n1sn nWD1X nD1.1/n1 nW1Z 0sn1esdsD1X nD1.1/n1 nW.n1/W; where in the last step we have identified the sintegral (cf. Table 1.2) as .n1/W. We complete the evaluation by noting that .n1/W=nWD1=n, so that the summation can be recognized as ln 2, thereby giving the final result 1Z 0erdr1Z res sdsDln 2:  A significant change in the form of 2-D or 3-D integrals can sometimes be accomplished by changing between Cartesian and polar coordinate systems. ArfKen_Ch01-9780123846549.tex 72 Chapter 1 Mathematical Preliminaries Example 1.10.7 EVALUATION IN POLAR COORDINATES In many calculus texts, the evaluation ofR1 0exp. x2/dxis carried out by first converting it into a 2-D integral by taking its square, which is then written and evaluated in polar coordinates. Using the fact that dxdyDr drd', we have 1Z 0dx ex21Z 0dy ey2D=2Z 0d'1Z 0r dr er2D 21Z 01 2du euD 4: This yields the well-known result 1Z 0ex2dxD1 2p: (1.148)  Example 1.10.8 ATOMIC INTERACTION INTEGRAL For study of the interaction of a small atom with an electromagnetic field, one of the in- tegrals that arises in a simple approximate treatment using Gaussian-type orbitals is (in dimensionless Cartesian coordinates) IDZ dz2 .x2Cy2Cz2/3=2e.x2Cy2Cz2/; where the range of the integration is the entire 3-D physical space ( I R3). Of course, this is a problem better addressed in spherical polar coordinates .r;;'/ , where ris the distance from the origin of the coordinate system, is the polar angle (for the Earth, known as colatitude), and 'is the azimuthal angle ( longitude). The relevant conversion formulas are:x2Cy2Cz2Dr2andz=rDcos. The volume element is dDr2sindrdd', and the ranges of the new coordinates are 0r<1,0, and 0'<2. In the spherical coordinates, our integral becomes IDZ dcos2 rer2D1Z 0dr rer2Z 0dcos2sin2Z 0d' D1 22 3 2 D2 3:  Remarks: Changes of Integration Variables In a 1-D integration, a change in the integration variable from, say, xtoyDy.x/involves two adjustments: (1) the differential dxmust be replaced by .dx=dy/dy, and (2) the in- tegration limits must be changed from x1;x2toy.x1/;y.x2/. Ify.x/is not single-valued ArfKen_Ch01-9780123846549.tex 1.10 Evaluation of Integrals 73 over the entire range .x1;x2/, the process becomes more complicated and we do not con- sider it further at this point. For multiple integrals, the situation is considerably more complicated and demands some discussion. Illustrating for a double integral, initially in variables x;y, but trans- formed to an integration in variables u;v, the differential dx dy must be transformed to J du dv, where J, called the Jacobian of the transformation and sometimes symbolically represented as [email protected];y/ @.u;v/ may depend on the variables. For example, the conversion from 2-D Cartesian coordinates x;yto plane polar coordinates r;involves the Jacobian [email protected];y/ @.r;/Dr;sodx dyDr dr d: For some coordinate transformations the Jacobian is simple and of a well-known form, as in the foregoing example. We can confirm the value assigned to Jby noticing that the area (in xyspace) enclosed by boundaries at r,rCdr,, andCdis an infinitesimally distorted rectangle with two sides of length drand two of length rd. See Fig. 1.18. For other transformations we may need general methods for obtaining Jacobians. Computation of Jacobians will be treated in detail in Section 4.4. Of interest here is the determination of the transformed region of integration. In prin- ciple this issue is straightforward, but all too frequently one encounters situations (both in other texts and in research articles) where misleading and potentially incorrect argu- ments are presented. The confusion normally arises in cases for which at least a part of the boundary is at infinity. We illustrate with the conversion from 2-D Cartesian to plane polar coordinates. Figure 1.19 shows that if one integrates for 0<2and0r<a, there are regions in the corners of a square (of side 2a) that are not included. If the integral is to be evaluated in the limit a!1 , it is both incorrect and meaningless to advance ar- guments about the “neglect” of contributions from these corner regions, as every point in these corners is ultimately included as ais increased. A similar, but slightly less obvious situation arises if we transform an integration over Cartesian coordinates 0x<1,0y<1, into one involving coordinates uDxCy, rdθ dr FIGURE 1.18 Element of area in plane polar coordinates. ArfKen_Ch01-9780123846549.tex 74 Chapter 1 Mathematical Preliminaries FIGURE 1.19 2-D integration, Cartesian and plane polar coordinates. V=a→ u=a AB V=o→ FIGURE 1.20 Integral in transformed coordinates. vDy, with integration limits 0u<1,0vu. See Fig. 1.20. Again it is incorrect and meaningless to make arguments justifying the “neglect” of the outer triangle (labeled Bin the figure). The relevant observation here is that ultimately, as the value of uis increased, any arbitrary point in the quarter-plane becomes included in the region being integrated. Exercises 1.10.1 Use a recursive method to show that, for all positive integers n,0.n/D.n1/W. Evaluate the integrals in Exercises 1.10.2 through 1.10.9. 1.10.21Z 0sinx xdx: Hint. Multiply integrand by eaxand take the limit a!0. 1.10.31Z 0dx cosh x: Hint. Expand the denominator is a way that converges for all relevant x. ArfKen_Ch01-9780123846549.tex 1.11 Dirac Delta Function 75 1.10.41Z 0dx eaxC1, for a>0. 1.10.51Z sinx x2dx: 1.10.61Z 0exsinx xdx: 1.10.7xZ 0erf.t/dt: The result can be expressed in terms of special functions in Table 1.2. 1.10.8xZ 1E1.t/dt: Obtain a result in which the only special function is E1. 1.10.91Z 0ex xC1dx: 1.10.10 Show that1Z 0tan1x x2 dxDln 2: Hint. Integrate by parts, to linearize in tan1. Then replace tan1xbytan1axand evaluate for aD1. 1.10.11 By direct integration in Cartesian coordinates, find the area of the ellipse defined by x2 a2Cy2 b2D1: 1.10.12 A unit circle is divided into two pieces by a straight line whose distance of closest approach to the center is 1/2 unit. By evaluating a suitable integral, find the area of the smaller piece thereby produced. Then use simple geometric considerations to verify your answer. 1.11 D IRAC DELTA FUNCTION Frequently we are faced with the problem of describing a quantity that is zero everywhere except at a single point, while at that point it is infinite in such a way that its integral over ArfKen_Ch01-9780123846549.tex 76 Chapter 1 Mathematical Preliminaries any interval containing that point has a finite value. For this purpose it is useful to introduce theDirac delta function, which is defined to have the properties .x/D0; x6D0; (1.149) f.0/DbZ af.x/.x/dx; (1.150) where f.x/is any well-behaved function and the integration includes the origin. As a special case of Eq. (1.150), 1Z 1.x/dxD1: (1.151) From Eq. (1.150),.x/must be an infinitely high, thin spike at xD0, as in the description of an impulsive force or the charge density for a point charge. The problem is that no such function exists, in the usual sense of function. However, the crucial property in Eq. (1.150) can be developed rigorously as the limit of a sequence of functions, a distribution. For example, the delta function may be approximated by any of the sequences of functions, Eqs. (1.152) to(1.155) andFigs. 1.21 and1.22: n.x/D8 >< >:0; x<1 2n n;1 2n<x<1 2n 0; x>1 2n;(1.152) n.x/Dnpexp.n2x2/; (1.153) xy=δn(x) xn πe−n2x2 FIGURE 1.21-Sequence function: left, Eq. (1.152); right, Eq. (1.153). ArfKen_Ch01-9780123846549.tex 1.11 Dirac Delta Function 77 xn π⋅1 1+n2x2 xsin nx πx FIGURE 1.22-Sequence function: left, Eq. (1.154); right, Eq. (1.155). n.x/Dn 1 1Cn2x2; (1.154) n.x/Dsinnx xD1 2nZ neixtdt: (1.155) While all these sequences (and others) cause .x/to have the same properties, they dif- fer somewhat in ease of use for various purposes. Equation (1.152) is useful in providing a simple derivation of the integral property, Eq. (1.150). Equation (1.153) is convenient to differentiate. Its derivatives lead to the Hermite polynomials. Equation (1.155) is particu- larly useful in Fourier analysis and in applications to quantum mechanics. In the theory of Fourier series, Eq. (1.155) often appears (modified) as the Dirichlet kernel: n.x/D1 2sinT.nC1 2/xU sin1 2x: (1.156) In using these approximations in Eq. (1.150) and elsewhere, we assume that f.x/is well behaved—that it offers no problems at large x. The forms for n.x/given in Eqs. (1.152) to(1.155) all obviously peak strongly for large natxD0. They must also be scaled in agreement with Eq. (1.151). For the forms inEqs. (1.152) and(1.154), verification of the scale is the topic of Exercises 1.11.1 and 1.11.2. To check the scales of Eqs. (1.153) and(1.155), we need values of the integrals 1Z 1en2x2dxDr nand1Z 1sinnx xdxD: These results are respectively trivial extensions of Eqs. (1.148) and (11.107) (the latter of which we derive later). ArfKen_Ch01-9780123846549.tex 78 Chapter 1 Mathematical Preliminaries For most physical purposes the forms describing delta functions are quite adequate. However, from a mathematical point of view the situation is still unsatisfactory. The limits limn!1n.x/ do not exist. A way out of this difficulty is provided by the theory of distributions. Recognizing that Eq. (1.150) is the fundamental property, we focus our attention on it rather than on .x/ itself. Equations (1.152) to(1.155) with nD1;2;3:::may be interpreted as sequences of normalized functions, and we may consistently write 1Z 1.x/f.x/dxlimn!1Z n.x/f.x/dx: (1.157) Thus,.x/is labeled a distribution (not a function) and is regarded as defined by Eq. (1.157). We might emphasize that the integral on the left-hand side of Eq. (1.157) is not a Riemann integral.8 Properties of .x/ From any of Eqs. (1.152) through (1.155) we see that Dirac’s delta function must be even in x,.x/D.x/. Ifa>0, .ax/D1 a.x/;a>0: (1.158) Equation (1.158) can be proved by making the substitution xDy=a: 1Z 1f.x/.ax/dxD1 a1Z 1f.y=a/.y/dyD1 af.0/: Ifa<0, Eq. (1.158) becomes .ax/D.x/=jaj. Shift of origin: 1Z 1.xx0/f.x/dxDf.x0/; (1.159) which can be proved by making the substitution yDxx0and noting that when yD0, xDx0. 8It can be treated as a Stieltjes integral if desired; .x/dxis replaced by du.x/, where u.x/is the Heaviside step function (compare Exercise 1.11.9). ArfKen_Ch01-9780123846549.tex 1.11 Dirac Delta Function 79 If the argument of .x/is a function g.x/with simple zeros at points aion the real axis (and therefore g0.ai/6D0),  g.x/ DX i.xai/ jg0.ai/j: (1.160) To prove Eq. (1.160), we write 1Z 1f.x/.x/dxDX iaiC"Z ai"f.x/ .xai/g0.ai/ dx; where we have decomposed the original integral into a sum of integrals over small in- tervals containing the zeros of g.x/. In these intervals, we replaced g.x/by the leading term in its Taylor series. Applying Eqs. (1.158) and(1.159) to each term of the sum, we confirm Eq. (1.160). Derivative of delta function: 1Z 1f.x/0.xx0/dxD1Z 1f0.x/.xx0/dxD f0.x0/: (1.161) Equation (1.161) can be taken as defining the derivative 0.x/; it is evaluated by per- forming an integration by parts on any of the sequences defining the delta function. In three dimensions, the delta function .r/ is interpreted as .x/.y/.z/, so it de- scribes a function localized at the origin and with unit integrated weight, irrespective of the coordinate system in use. Thus, in spherical polar coordinates, y f.r2/.r 2r1/r2 2dr2sin2d2d2Df.r1/: (1.162) Equation (1.155) corresponds in the limit to .tx/D1 21Z 1exp i!.tx/ d!; (1.163) with the understanding that this has meaning only when under an integral sign. In that context it is extremely useful for the simplification of Fourier integrals (Chapter 20). Expansions of .x/are addressed in Chapter 5. See Example 5.1.7. Kronecker Delta It is sometimes useful to have a symbol that is the discrete analog of the Dirac delta func- tion, with the property that it is unity when the discrete variable has a certain value, and zero otherwise. A quantity with these properties is known as the Kronecker delta, defined for indices iandjas i jD( 1;iDj; 0;i6Dj:(1.164) ArfKen_Ch01-9780123846549.tex 80 Chapter 1 Mathematical Preliminaries Frequent uses of this symbol are to select a special term from a summation, or to have one functional form for all nonzero values of an index, but a different form when the index is zero. Examples: X i jfi ji jDX ifii;CnD1 1Cn02 L: Exercises 1.11.1 Let n.x/D8 >>>>< >>>>:0;x<1 2n; n;1 2n<x<1 2n; 0;1 2n<x: Show that limn!11Z 1f.x/n.x/dxDf.0/; assuming that f.x/is continuous at xD0. 1.11.2 For n.x/Dn 1 1Cn2x2; show that 1Z 1n.x/dxD1: 1.11.3 Fejer’s method of summing series is associated with the function n.t/D1 2nsin.nt=2/ sin.t=2/2 : Show thatn.t/is a delta distribution, in the sense that limn!11 2n1Z 1f.t/sin.nt=2/ sin.t=2/2 dtDf.0/: 1.11.4 Prove that Ta.xx1/UD1 a.xx1/: Note. IfTa.xx1/Uis considered even, relative to x1, the relation holds for negative aand1=amay be replaced by 1=jaj. 1.11.5 Show that T.xx1/.xx2/UDT. xx1/C.xx2/U=jx1x2j: Hint. Try using Exercise 1.11.4. ArfKen_Ch01-9780123846549.tex 1.11 Dirac Delta Function 81 xn large n small1n → ∞un(x) FIGURE 1.23 Heaviside unit step function. 1.11.6 Using the Gauss error curve delta sequence nDnpen2x2, show that xd dx.x/D. x/; treating.x/and its derivative as in Eq. (1.157). 1.11.7 Show that 1Z 10.x/f.x/dxD f0.0/: Here we assume that f0.x/is continuous at xD0. 1.11.8 Prove that .f.x//D d f.x/ dx 1 xDx0.xx0/; where x0is chosen so that f.x0/D0. Hint. Note that .f/d fD.x/dx: 1.11.9 (a) If we define a sequence n.x/Dn=.2 cosh2nx/, show that 1Z 1n.x/dxD1; independent of n. (b) Continuing this analysis, show that9 xZ 1n.x/dxD1 2[1Ctanhnx]un.x/ and limn!1un.x/D0;x<0; 1;x>0: This is the Heaviside unit step function (Fig. 1.23). 9Many other symbols are used for this function. This is the AMS-55 notation (in Additional Readings, see Abramowitz and Stegun): ufor unit. ArfKen_Ch01-9780123846549.tex 82 Chapter 1 Mathematical Preliminaries Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Contains a wealth of information about a large number of special functions. Bender, C. M., and S. Orszag, Advanced Mathematical Methods for Scientists and Engineers. New York: McGraw-Hill (1978). Particularly recommended for methods of accelerating convergence. Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969), reprinted, Dover (1992). This is an advanced text that presupposes moderate knowledge of mathematical physics. Courant, R., and D. Hilbert, Methods of Mathematical Physics, Vol. 1 (1st English ed.). New York: Wiley (Interscience) (1953). As a reference book for mathematical physics, it is particularly valuable for existence theorems and discussion of areas such as eigenvalue problems, integral equations, and calculus of variations. Galambos, J., Representations of Real Numbers by Infinite Series. Berlin: Springer (1976). Gradshteyn, I. S., and I. M. Ryzhik, Table of Integrals, Series, and Products. Corrected and enlarged 7th ed., edited by A. Jeffrey and D. Zwillinger. New York: Academic Press (2007). Hansen, E., A Table of Series and Products. Englewood Cliffs, NJ: Prentice-Hall (1975). A tremendous compi- lation of series and products. Hardy, G. H., Divergent Series. Oxford: Clarendon Press (1956), 2nd ed., Chelsea (1992). The standard, com- prehensive work on methods of treating divergent series. Hardy includes instructive accounts of the gradual development of the concepts of convergence and divergence. Jeffrey, A., Handbook of Mathematical Formulas and Integrals. San Diego: Academic Press (1995). Jeffreys, H. S., and B. S. Jeffreys, Methods of Mathematical Physics, 3rd ed. Cambridge, UK: Cambridge Univer- sity Press (1972). This is a scholarly treatment of a wide range of mathematical analysis, in which considerable attention is paid to mathematical rigor. Applications are to classical physics and geophysics. Knopp, K., Theory and Application of Infinite Series. London: Blackie and Son, 2nd ed. New York: Hafner (1971), reprinted A. K. Peters Classics (1997). This is a thorough, comprehensive, and authoritative work that covers infinite series and products. Proofs of almost all the statements about series not proved in this chapter will be found in this book. Mangulis, V., Handbook of Series for Scientists and Engineers. New York: Academic Press (1965). A most convenient and useful collection of series. Includes algebraic functions, Fourier series, and series of the special functions: Bessel, Legendre, and others. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics, 2 vols. New York: McGraw-Hill (1953). This work presents the mathematics of much of theoretical physics in detail but at a rather advanced level. It is recommended as the outstanding source of information for supplementary reading and advanced study. Rainville, E. D., Infinite Series. New York: Macmillan (1967). A readable and useful account of series constants and functions. Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering, 2nd ed. New York: McGraw-Hill (1966). A long chapter 2 (101 pages) presents infinite series in a thorough but very read- able form. Extensions to the solutions of differential equations, to complex series, and to Fourier series are included. Spiegel, M. R., Complex Variables, in Schaum’s Outline Series. New York: McGraw-Hill (1964, reprinted 1995). Clear, to the point, and with very large numbers of examples, many solved step by step. Answers are provided for all others. Highly recommended. Whittaker, E. T., and G. N. Watson, A Course of Modern Analysis, 4th ed. Cambridge, UK: Cambridge University Press (1962), paperback. Although this is the oldest of the general references (original edition 1902), it still is the classic reference. It leans strongly towards pure mathematics, as of 1902, with full mathematical rigor. ArfKen_Ch02-9780123846549.tex CHAPTER 2 DETERMINANTS AND MATRICES 2.1 D ETERMINANTS We begin the study of matrices by solving linear equations that will lead us to determi- nants and matrices. The concept of determinant and the notation were introduced by the renowned German mathematician and philosopher Gottfried Wilhelm von Leibniz. Homogeneous Linear Equations One of the major applications of determinants is in the establishment of a condition for the existence of a nontrivial solution for a set of linear homogeneous algebraic equations. Suppose we have three unknowns x1;x2;x3(ornequations with nunknowns): a1x1Ca2x2Ca3x3D0; b1x1Cb2x2Cb3x3D0; (2.1) c1x1Cc2x2Cc3x3D0: The problem is to determine under what conditions there is any solution, apart from the trivial one x1D0;x2D0;x3D0. If we use vector notation xD.x1;x2;x3/for the solution and three rows aD.a1;a2;a3/;bD.b1;b2;b3/;cD.c1;c2;c3/of coefficients, then the three equations, Eqs. (2.1), become axD0; bxD0; cxD0: (2.2) These three vector equations have the geometrical interpretation that xis orthogonal toa;b;andc. If the volume spanned by a;b;cgiven by the determinant (or triple scalar 83 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch02-9780123846549.tex 84 Chapter 2 Determinants and Matrices product, see Eq. (3.12) of Section 3.2) D3D.ab/cDdet.a; b;c/D a1a2a3 b1b2b3 c1c2c3 (2.3) is not zero, then there is only the trivial solution xD0. For an introduction to the cross product of vectors, see Chapter 3: Vector Analysis, Section 3.2: Vectors in 3-D Space. Conversely, if the aforementioned determinant of coefficients vanishes, then one of the row vectors is a linear combination of the other two. Let us assume that clies in the plane spanned by aandb, that is, that the third equation is a linear combination of the first two and not independent. Then xis orthogonal to that plane so that xab. Since homogeneous equations can be multiplied by arbitrary numbers, only ratios of the xiare relevant, for which we then obtain ratios of 22determinants x1 x3Da2b3a3b2 a1b2a2b1;x2 x3Da1b3a3b1 a1b2a2b1(2.4) from the components of the cross product ab;provided x3a1b2a2b16D0. This is Cramer’s rule for three homogeneous linear equations. Inhomogeneous Linear Equations The simplest case of two equations with two unknowns, a1x1Ca2x2Da3;b1x1Cb2x2Db3; (2.5) can be reduced to the previous case by imbedding it in three-dimensional (3-D) space with a solution vector xD.x1;x2;1/ and row vectors aD.a1;a2;a3/;bD.b1;b2;b3/. As before, Eqs. (2.5) in vector notation, axD0andbxD0, imply that xab;so the analog of Eq. (2.4) holds. For this to apply, though, the third component of abmust not be zero, that is, a1b2a2b16D0, because the third component of xis16D0. This yields thexias x1Da3b2b3a2 a1b2a2b1D a3a2 b3b2 a1a2 b1b2 ; (2.6) x2Da1b3a3b1 a1b2a2b1D a1a3 b1b3 a1a2 b1b2 : (2.7) The determinant in the numerator of x1.x2/is obtained from the determinant of the coef- ficients a1a2 b1b2 by replacing the first (second) column vector by the vectora3 b3 of the inhomogeneous side of Eq. (2.5). This is Cramer’s rule for a set of two inhomogeneous linear equations with two unknowns. A full understanding of the above exposition requires now that we introduce a formal definition of the determinant and show how it relates to the foregoing. ArfKen_Ch02-9780123846549.tex 2.1 Determinants 85 Definitions Before defining a determinant, we need to introduce some related concepts and definitions. When we write two-dimensional (2-D) arrays of items, we identify the item in the nth horizontal row and the mth vertical column by the index set n;m; note that the row index is conventionally written first. Starting from a set of nobjects in some reference order (e.g., the number sequence 1;2;3;:::; n), we can make a permutation of them to some other order; the total number of distinct permutations that are possible is nW(choose the first object nways, then choose the second in n1ways, etc.). Every permutation of nobjects can be reached from the reference order by a succession of pairwise interchanges (e.g., 1234!4132 can be reached by the successive steps 1234!1432!4132 ). Although the number of pairwise interchanges needed for a given permutation depends on the path (compare the above example with 1234! 1243!1423!4123!4132 ), for a given permutation the number of interchanges will always either be even orodd. Thus a permutation can be identified as having either even or odd parity. It is convenient to introduce the Levi-Civita symbol, which for an n-object system is denoted by"i j:::, where"hasnsubscripts, each of which identifies one of the objects. This Levi-Civita symbol is defined to be C1ifi j:::represents an even permutation of the objects from a reference order; it is defined to be 1ifi j:::represents an odd permutation of the objects, and zero if i j:::does not represent a permutation of the objects (i.e., contains an entry duplication). Since this is an important definition, we set it out in a display format: "i j:::DC1; i j:::an even permutation, D1; i j:::an odd permutation, D0;i j:::not a permutation. (2.8) We now define a determinant of order nto be an nnsquare array of numbers (or func- tions), with the array conventionally written within vertical bars (not parentheses, braces, or any other type of brackets), as follows: DnD a11a12::: a1n a21a22::: a2n a31a32::: a3n ::: ::: ::: ::: an1an2::: ann : (2.9) The determinant Dnhas a value that is obtained by 1. Forming all nWproducts that can be formed by choosing one entry from each row in such a way that one entry comes from each column, 2. Assigning each product a sign that corresponds to the parity of the sequence in which the columns were used (assuming the rows were used in an ascending sequence), 3. Adding (with the assigned signs) the products. ArfKen_Ch02-9780123846549.tex 86 Chapter 2 Determinants and Matrices More formally, the determinant in Eq. (2.9) is defined to have the value DnDX i j:::"i j:::a1ia2j: (2.10) The summations in Eq. (2.10) need not be restricted to permutations, but can be assumed to range independently from 1 through n; the presence of the Levi-Civita symbol will cause only the index combinations corresponding to permutations to actually contribute to the sum. Example 2.1.1 DETERMINANTS OF ORDERS 2AND3 To make the definition more concrete, we illustrate first with a determinant of order 2. The Levi-Civita symbols needed for this determinant are "12DC1 and"21D1 (note that "11D"22D0), leading to D2D a11a12 a21a22 D"12a11a22C"21a12a21Da11a22a12a21: We see that this determinant expands into 2WD 2terms. A specific example of a determi- nant of order 2 is a1a2 b1b2 Da1b2b1a2: Determinants of order 3 expand into 3WD 6terms. The relevant Levi-Civita symbols are"123D"231D"312DC1 ,"213D"321D"132D1 ; all other index combinations have "i jkD0, so D3D a11a12a13 a21a22a23 a31a32a33 DX i jk"i jka1ia2ja3k Da11a22a33a11a23a32a13a22a31a12a21a33Ca12a23a31Ca13a21a32: The expression in Eq. (2.3) is the determinant of order 3 a1a2a3 b1b2b3 c1c2c3 Da1b2c3a1b3c2a2b1c3Ca2b3c1Ca3b1c2a3b2c1: Note that half of the terms in the expansion of a determinant bear negative signs. It is quite possible that a determinant of large elements will have a very small value. Here is one example: 8 11 7 9 11 5 8 12 9 D1:  ArfKen_Ch02-9780123846549.tex 2.1 Determinants 87 Properties of Determinants The symmetry properties of the Levi-Civita symbol translate into a number of symme- tries exhibited by determinants. For simplicity, we illustrate with determinants of order 3. The interchange of two columns of a determinant causes the Levi-Civita symbol multi- plying each term of the expansion to change sign; the same is true if two rows are inter- changed. Moreover, the roles of rows and columns may be interchanged; if a determinant with elements ai jis replaced by one with elements bi jDaji, we call the bi jdetermi- nant the transpose of the ai jdeterminant. Both these determinants have the same value. Summarizing: Interchanging two rows (or two columns) changes the sign of the value of a determi- nant. Transposition does not alter its value. Thus, a11a12a13 a21a22a23 a31a32a33 D a12a11a13 a22a21a23 a32a31a33 D a11a21a31 a12a22a32 a13a23a33 : (2.11) Further consequences of the definition in Eq. (2.10) are: (1) Multiplication of all members of a single column (or a single row) by a constant k causes the value of the determinant to be multiplied by k, (2) If the elements of a column (or row) are actually sums of two quantities, the deter- minant can be decomposed into a sum of two determinants. Thus, k a11a12a13 a21a22a23 a31a32a33 D ka11a12a13 ka21a22a23 ka31a32a33 D ka11ka12ka13 a21 a22 a23 a31 a32 a33 ; (2.12) a11Cb1a12a13 a21Cb2a22a23 a31Cb3a32a33 D a11a12a13 a21a22a23 a31a32a33 C b1a12a13 b2a22a23 b3a32a33 : (2.13) These basic properties and/or the basic definition mean that Any determinant with two rows equal, or two columns equal, has the value zero. To prove this, interchange the two identical rows or columns; the determinant both remains the same and changes sign, and therefore must have the value zero. An extension of the above is that if two rows (or columns) are proportional, the deter- minant is zero. The value of a determinant is unchanged if a multiple of one row is added (column by column) to another row or if a multiple of one column is added (row by row) to another column. Applying Eq. (2.13), the addition does not contribute to the value of the determinant. If each element in a row or each element in a column is zero, the determinant has the value zero. ArfKen_Ch02-9780123846549.tex 88 Chapter 2 Determinants and Matrices Laplacian Development by Minors The fact that a determinant of order nexpands into nWterms means that it is important to identify efficient means for determinant evaluation. One approach is to expand in terms of minors. The minor corresponding to ai j, denoted Mi j, orMi j.a/if we need to identify M as coming from the ai j, is the determinant (of order n1) produced by striking out row i and column jof the original determinant. When we expand into minors, the quantities to be used are the cofactors of the.i j/elements, defined as .1/iCjMi j. The expansion can be made for any row or column of the original determinant. If, for example, we expand the determinant of Eq. (2.9) using row i, we have DnDnX jD1ai j.1/iCjMi j: (2.14) This expansion reduces the work involved in evaluation if the row or column selected for the expansion contains zeros, as the corresponding minors need not be evaluated. Example 2.1.2 EXPANSION IN MINORS Consider the determinant (arising in Dirac’s relativistic electron theory) D a11a12a13a14 a21a22a23a24 a31a32a33a34 a41a42a43a44 D 0 1 0 0 1 0 0 0 0 0 0 1 0 01 0 : Expanding across the top row, only one 33matrix survives: DD.1/1C2a12M12.a/D.1/.1/ 1 0 0 0 0 1 01 0 .1/ b11b12b13 b21b22b23 b31b32b33 : Expanding now across the second row, we get DD.1/.1/2C3b23M23.b/D 1 0 01 D1: When we finally reached a 22determinant, it was simple to evaluate it without further expansion.  Linear Equation Systems We are now ready to apply our knowledge of determinants to the solution of systems of linear equations. Suppose we have the simultaneous equations a1x1Ca2x2Ca3x3Dh1; b1x1Cb2x2Cb3x3Dh2; c1x1Cc2x2Cc3x3Dh3: (2.15) ArfKen_Ch02-9780123846549.tex 2.1 Determinants 89 To use determinants to help solve this equation system, we define DD a1a2a3 b1b2b3 c1c2c3 : (2.16) Starting from x1D, we manipulate it by (1) moving x1to multiply the entries of the first column of D, then (2) adding to the first column x2times the second column and x3times the third column (neither of these operations change the value). We then reach the second line of Eq. (2.17) by substituting the right-hand sides of Eqs. (2.15). These operations are illustrated here: x1DD a1x1a2a3 b1x1b2b3 c1x1c2c3 D a1x1Ca2x2Ca3x3a2a3 b1x1Cb2x2Cb3x3b2b3 c1x1Cc2x2Cc3x3c2c3 D h1a2a3 h2b2b3 h3c2c3 : (2.17) IfD6D0,Eq. (2.17) may now be solved for x1: x1D1 D h1a2a3 h2b2b3 h3c2c3 : (2.18) Analogous procedures starting from x2Dandx3Dgive the parallel results x2D1 D a1h1a3 b1h2b3 c1h3c3 ;x3D1 D a1a2h2 b1b2h2 c1c2h3 : We see that the solution for xiis1=Dtimes a numerator obtained by replacing the ith column of Dby the right-hand-side coefficients, a result that can be generalized to an arbi- trary number nof simultaneous equations. This scheme for the solution of linear equation systems is known as Cramer’s rule. IfDis nonzero, the above construction of the xiis definitive and unique, so that there will be exactly one solution to the equation set. If D6D0and the equations are homoge- neous (i.e., all the hiare zero), then the unique solution is that all the xiare zero. Determinants and Linear Dependence The preceding subsections go a long way toward identifying the role of the determi- nant with respect to linear dependence. If nlinear equations in nvariables, written as inEq. (2.15), have coefficients that form a nonzero determinant, the variables are uniquely determined, meaning that the forms constituting the left-hand sides of the equations must in fact be linearly independent. However, we would still like to prove the property illus- trated in the introduction to this chapter, namely that if a set of forms is linearly depen- dent, the determinant of their coefficients will be zero. But this result is nearly immediate. The existence of linear dependence means that there exists one equation whose coefficients ArfKen_Ch02-9780123846549.tex 90 Chapter 2 Determinants and Matrices are linear combinations of the coefficients of the other equations, and we may use that fact to reduce to zero the row of the determinant corresponding to that equation. In summary, we have therefore established the following important result: If the coefficients of nlinear forms in nvariables form a nonzero determinant, the forms are linearly independent; if the determinant of the coefficients is zero, the forms exhibit linear dependence. Linearly Dependent Equations If a set of linear forms is linearly dependent, we can distinguish three distinct situations when we consider equation systems based on these forms. First, and of most importance for physics, is the case in which all the equations are homogeneous, meaning that the right-hand side quantities hiin equations of the type Eq. (2.15) are all zero. Then, one or more of the equations in the set will be equivalent to linear combinations of others, and we will have less than nequations in our nvariables. We can then assign one (or in some cases, more than one) variable an arbitrary value, obtaining the others as functions of the assigned variables. We thus have a manifold (i.e., a parameterized set) of solutions to our equation system. Combining the above analysis with our earlier observation that if a set of homogeneous linear equations has a nonvanishing determinant it has the unique solution that all the xi are zero, we have the following important result: A system of nhomogeneous linear equations in nunknowns has solutions that are not identically zero only if the determinant of its coefficients vanishes. If that determinant vanishes, there will be one or more solutions that are not identically zero and are arbitrary as to scale. A second case is where we have (or combine equations so that we have) the same linear form in two equations, but with different values of the right-hand quantities hi. In that case the equations are mutually inconsistent, and the equation system has no solution. A third, related case, is where we have a duplicated linear form, but with a common value of hi. This also leads to a solution manifold. Example 2.1.3 LINEARLY DEPENDENT HOMOGENEOUS EQUATIONS Consider the equation set x1Cx2Cx3D0; x1C3x2C5x3D0; x1C2x2C3x3D0: Here DD 1 1 1 1 3 5 1 2 3 D1.3/.3/1.5/.2/1.3/.1/1.1/.3/C1.5/.1/C1.1/.2/D0: ArfKen_Ch02-9780123846549.tex 2.1 Determinants 91 The third equation is half the sum of the other two, so we drop it. Then, second equation minus first: 2x2C4x3D0!x2D2 x3; (3first equation) minus second: 2x12x3D0!x1Dx3: Since x3can have any value, there is an infinite number of solutions, all of the form .x1;x2;x3/Dconstant.1;2;1/. Our solution illustrates an important property of homogeneous linear equations, namely that any multiple of a solution is also a solution. The solution only becomes less arbitrary if we impose a scale condition. For example, in the present case we could require the squares of the xito add to unity. Even then, the solution would still be arbitrary as to overall sign.  Numerical Evaluation There is extensive literature on determinant evaluation. Computer codes and many refer- ences are given, for example, by Press et al.1We present here a straightforward method due to Gauss that illustrates the principles involved in all the modern evaluation methods. Gauss elimination is a versatile procedure that can be used for evaluating determinants, for solving linear equation systems, and (as we will see later) even for matrix inversion. Example 2.1.4 GAUSS ELIMINATION Our example, a 33linear equation system, can easily be done in other ways, but it is used here to provide an understanding of the Gauss elimination procedure. We wish to solve 3xC2yCzD11 2xC3yCzD13 xCyC4zD12: (2.19) For convenience and for the optimum numerical accuracy, the equations are rearranged so that, to the extent possible, the largest coefficients run along the main diagonal (upper left to lower right). The Gauss technique is to use the first equation to eliminate the first unknown, x, from the remaining equations. Then the (new) second equation is used to eliminate yfrom the last equation. In general, we work down through the set of equations, and then, with one un- known determined, we work back up to solve for each of the other unknowns in succession. 1W. H. Press, B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes, 2nd ed. Cambridge, UK: Cambridge University Press (1992), Chapter 2. ArfKen_Ch02-9780123846549.tex 92 Chapter 2 Determinants and Matrices It is convenient to start by dividing each row by its initial coefficient, converting Eq. (2.19) to xC2 3yC1 3zD11 3 xC3 2yC1 2zD13 2 xCyC4zD12: (2.20) Now, using the first equation, we eliminate xfrom the second and third equations by subtracting the first equation from each of the others: xC2 3yC1 3zD11 3 5 6yC1 6zD17 6 1 3yC11 3zD25 3: (2.21) Then we divide the second and third rows by their initial coefficients: xC2 3yC1 3zD11 3 yC1 5zD17 5 yC11zD25: (2.22) Repeating the technique, we use the new second equation to eliminate yfrom the third equation, which can then be solved for z: xC2 3yC1 3zD11 3 yC1 5zD17 5 54 5zD108 5! zD2: (2.23) Now that zhas been determined, we can return to the second equation, finding yC1 52D17 5! yD3; and finally, continuing to the first equation, xC2 33C1 32D11 3! xD1: The technique may not seem as elegant as the use of Cramer’s rule, but it is well adapted to computers and is far faster than the time spent with determinants. If we had not kept the right-hand sides of the equation system, the Gauss elimination process would have simply brought the original determinant into triangular form (but note ArfKen_Ch02-9780123846549.tex 2.1 Determinants 93 that our processes for making the leading coefficients unity cause corresponding changes in the value of the determinant). In the present problem, the original determinant DD 3 2 1 2 3 1 1 1 4 was divided by 3 and by 2 going from Eq. (2.19) to(2.20), and multiplied by 6/5 and by 3 going from Eq. (2.21) to(2.22), so that Dand the determinant represented by the left-hand side of Eq. (2.23) are related by DD.3/.2/5 61 3 12 31 3 0 11 5 0 054 5 D5 354 5D18: (2.24) Because all the entries in the lower triangle of the determinant explicitly shown in Eq. (2.24) are zero, the only term that contributes to it is the product of the diagonal elements: To get a nonzero term, we must use the first element of the first row, then the second element of the second row, etc. It is easy to verify that the final result obtained in Eq. (2.24) agrees with the result of evaluating the original form of D.  Exercises 2.1.1 Evaluate the following determinants: .a/ 1 0 1 0 1 0 1 0 0 ; .b/ 1 2 0 3 1 2 0 3 1 ; .c/1p 2 0p 3 0 0p 3 0 2 0 0 2 0p 3 0 0p 3 0 : 2.1.2 Test the set of linear homogeneous equations xC3yC3zD0; xyCzD0;2xCyC3zD0 to see if it possesses a nontrivial solution. In any case, find a solution to this equation set. 2.1.3 Given the pair of equations xC2yD3;2xC4yD6; (a) Show that the determinant of the coefficients vanishes. (b) Show that the numerator determinants, see Eq. (2.18), also vanish. (c) Find at least two solutions. 2.1.4 IfCi jis the cofactor of element ai j, formed by striking out the ith row and jth column and including a sign .1/iCj, show that ArfKen_Ch02-9780123846549.tex 94 Chapter 2 Determinants and Matrices (a)P iai jCi jDP iajiCjiDjAj, wherejAjis the determinant with the elements ai j, (b)P iai jCikDP iajiCkiD0;j6Dk. 2.1.5 A determinant with all elements of order unity may be surprisingly small. The Hilbert determinant Hi jD.iCj1/1,i;jD1;2;:::; nis notorious for its small values. (a) Calculate the value of the Hilbert determinants of order nfornD1;2;and 3. (b) If an appropriate subroutine is available, find the Hilbert determinants of order n fornD4;5, and 6. ANS. n Det.Hn/ 1 1. 2 8:33333102 3 4:62963104 4 1:65344107 5 3:749301012 6 5:367301018. 2.1.6 Prove that the determinant consisting of the coefficients from a set of linearly dependent forms has the value zero. 2.1.7 Solve the following set of linear simultaneous equations. Give the results to five decimal places. 1:0x1C0:9x2C0:8x3C0:4x4C0:1x5D1:0 0:9x1C1:0x2C0:8x3C0:5x4C0:2x5C0:1x6D0:9 0:8x1C0:8x2C1:0x3C0:7x4C0:4x5C0:2x6D0:8 0:4x1C0:5x2C0:7x3C1:0x4C0:6x5C0:3x6D0:7 0:1x1C0:2x2C0:4x3C0:6x4C1:0x5C0:5x6D0:6 0:1x2C0:2x3C0:3x4C0:5x5C1:0x6D0:5: Note. These equations may also be solved by matrix inversion, as discussed in Section 2.2. 2.1.8 Show that (in 3-D space) (a)X iiiD3; (b)X i ji j"i jkD0; (c)X pq"ipq"jpqD2i j; (d)X i jk"i jk"i jkD6: ArfKen_Ch02-9780123846549.tex 2.2 Matrices 95 Note. The symboli jis the Kronecker delta, defined in Eq. (1.164), and "i jkis the Levi-Civita symbol, Eq. (2.8). 2.1.9 Show that (in 3-D space) X k"i jk"pqkDipjqiqjp: Note. See Exercise 2.1.8 for definitions of i jand"i jk. 2.2 M ATRICES Matrices are 2-D arrays of numbers or functions that obey the laws that define matrix algebra. The subject is important for physics because it facilitates the description of linear transformations such as changes of coordinate systems, provides a useful formu- lation of quantum mechanics, and facilitates a variety of analyses in classical and rel- ativistic mechanics, particle theory, and other areas. Note also that the development of a mathematics of two-dimensionally ordered arrays is a natural and logical extension of concepts involving ordered pairs of numbers (complex numbers) or ordinary vectors (one- dimensional arrays). The most distinctive feature of matrix algebra is the rule for the multiplication of matrices. As we will see in more detail later, the algebra is defined so that a set of lin- ear equations such as a1x1Ca2x2Dh1 b1x1Cb2x2Dh2 can be written as a single matrix equation of the form a1a2 b1b2x1 x2 Dh1 h2 : In order for this equation to be valid, the multiplication indicated by writing the two matrices next to each other on the left-hand side has to produce the result a1x1Ca2x2 b1x1Cb2x2 and the statement of equality in the equation has to mean element-by-element agreement of its left-hand and right-hand sides. Let’s move now to a more formal and precise description of matrix algebra. Basic Definitions Amatrix is a set of numbers or functions in a 2-D square or rectangular array. There are no inherent limitations on the number of rows or columns. A matrix with m(horizontal) rows and n(vertical) columns is known as an mnmatrix, and the element of a matrix A in row iand column jis known as its i;jelement, often labeled ai j. As already observed ArfKen_Ch02-9780123846549.tex 96 Chapter 2 Determinants and Matrices 0 BB@u1 u2 u3 u41 CCA0 @4 2 1 3 0 11 A6 7 0 1 4 3 0 1 1 0a11a12 FIGURE 2.1 From left to right, matrices of dimension 41(column vector), 32,23,22(square), 12(row vector). when we introduced determinants, when row and column indices or dimensions are men- tioned together, it is customary to write the row indicator first. Note also that order matters, in general the i;jandj;ielements of a matrix are different, and (if m6Dn)nmand mnmatrices even have different shapes. A matrix for which nDmis termed square; one consisting of a single column (an m1matrix) is often called a column vector, while a matrix with only one row (therefore 1n) is a row vector. We will find that identi- fying these matrices as vectors is consistent with the properties identified for vectors in Section 1.7. The arrays constituting matrices are conventionally enclosed in parentheses (not vertical lines, which indicate determinants, or square brackets). A few examples of matrices are shown in Fig. 2.1. We will usually write the symbols denoting matrices as upper-case letters in a sans-serif font (as we did when introducing A); when a matrix is known to be a column vector we often denote it by a lower-case boldface letter in a Roman font (e.g., x). Perhaps the most important fact to note is that the elements of a matrix are not combined with one another. A matrix is not a determinant. It is an ordered array of numbers, not a single number. To refer to the determinant whose elements are those of a square matrix A (more simply, “the determinant of A”), we can write det.A/ . Matrices, so far just arrays of numbers, have the properties we assign to them. These properties must be specified to complete the definition of matrix algebra. Equality IfAandBare matrices, ADBonly if ai jDbi jfor all values of iandj. A necessary but not sufficient condition for equality is that both matrices have the same dimensions. Addition, Subtraction Addition and subtraction are defined only for matrices AandBof the same dimensions, in which case ABDC, with ci jDai jbi jfor all values of iandj, the elements combining according to the law of ordinary algebra (or arithmetic if they are simple numbers). This means that Cwill be a matrix of the same dimensions as AandB. Moreover, we see that addition is commutative: ACBDBCA. It is also associative, meaning that .ACB/C CDAC.BCC/. A matrix with all elements zero, called a null matrix orzero matrix, can either be written as Oor as a simple zero, with its matrix character and dimensions determined from the context. Thus, for all A, AC0D0CADA: (2.25) ArfKen_Ch02-9780123846549.tex 2.2 Matrices 97 Multiplication (by a Scalar) Here what we mean by a scalar is an ordinary number or function (not another matrix). The multiplication of matrix Aby the scalar quantity produces BD A, with bi jD ai j for all values of iandj. This operation is commutative, with ADA . Note that the definition of multiplication by a scalar causes each element of matrix Ato be multiplied by the scalar factor. This is in striking contrast to the behavior of determinants in which det.A/is a determinant in which the factor multiplies only one column or one row of det.A/and not every element of the entire determinant. If Ais an nnsquare matrix, then det. A/D ndet.A/: Matrix Multiplication (Inner Product) Matrix multiplication is not an element-by-element operation like addition or multiplica- tion by a scalar. Instead, it is a more complicated operation in which each element of the product is formed by combining elements of a row of the first operand with correspond- ing elements of a column of the second operand. This mode of combination proves to be that which is needed for many purposes, and gives matrix algebra its power for solving important problems. This inner product of matrices AandBis defined as A BDC; with ci jDX kaikbkj: (2.26) This definition causes the i jelement of Cto be formed from the entire ith row of Aand the entire jth column of B. Obviously this definition requires that Ahave the same number of columns ( n) asBhas rows. Note that the product will have the same number of rows asAand the same number of columns as B. Matrix multiplication is defined only if these conditions are met. The summation in Eq. (2.26) is over the range of kfrom 1 to n, and, more explicitly, corresponds to ci jDai1b1jCai2b2jCC a1nbnj: This combination rule is of a form similar to that of the dot product of the vectors .ai1;ai2;:::; ain/and.b1j;b2j;:::; bnj/. Because the roles of the two operands in a matrix multiplication are different (the first is processed by rows, the second by columns), the operation is in general not commutative, that is, A B6DB A. In fact, A Bmay even have a different shape than B A. IfAandBare square, it is useful to define the commutator of AandB, TA;BUD A BB A; (2.27) which, as stated above, will in many cases be nonzero. Matrix multiplication is associative, meaning that (AB)CDA(BC). Proof of this state- ment is the topic of Exercise 2.2.26. ArfKen_Ch02-9780123846549.tex 98 Chapter 2 Determinants and Matrices Example 2.2.1 MULTIPLICATION, PAULI MATRICES These three 22matrices, which occurred in early work in quantum mechanics by Pauli, are encountered frequently in physics contexts, so a familiarity with them is highly advis- able. They are 1D0 1 1 0 ;  2D0i i 0 ;  3D1 0 01 : (2.28) Let’s form12. The 1;1element of the product involves the first row of1and the first column of2; these are shaded and lead to the indicated computation: 0 1 1 00i i0 !0.0/C1.i/Di: Continuing, we have 12D0.0/C1.i/0.i/C1.0/ 1.0/C0.i/1.i/C0.0/ Di 0 0i : (2.29) In a similar fashion, we can compute 21D0i i 00 1 1 0 Di0 0i : (2.30) It is clear that 1and2do not commute. We can construct their commutator: T1;2UD1221Di 0 0i i0 0i D2i1 0 01 D2i3: (2.31) Note that not only have we verified that 1and2do not commute, we have even evaluated and simplified their commutator.  Example 2.2.2 MULTIPLICATION, ROW AND COLUMN MATRICES As a second example, consider AD0 @1 2 31 A;BD4 5 6 : Let us form A BandB A: A BD0 @4 5 6 8 10 12 12 15 181 A;B AD.41C52C63/D.32/: The results speak for themselves. Often when a matrix operation leads to a 11ma- trix, the parentheses are dropped and the result is treated as an ordinary number or function.  ArfKen_Ch02-9780123846549.tex 2.2 Matrices 99 Unit Matrix By direct matrix multiplication, it is possible to show that a square matrix with elements of value unity on its principal diagonal (the elements .i;j/with iDj), and zeros every- where else, will leave unchanged any matrix with which it can be multiplied. For example, the33unit matrix has the form 0 @1 0 0 0 1 0 0 0 11 AI note that it is nota matrix all of whose elements are unity. Giving such a matrix the name 1, 1ADA1DA: (2.32) In interpreting this equation, we must keep in mind that unit matrices, which are square and therefore of dimensions nn, exist for all n; the nvalues for use in Eq. (2.32) must be those consistent with the applicable dimension of A. So if Aismn, the unit matrix in 1Amust be mm, while that in A1must be nn. The previously introduced null matrices have only zero elements, so it is also obvious that for all A, O ADA ODO: (2.33) Diagonal Matrices If a matrix Dhas nonzero elements di jonly for iDj, it is said to be diagonal; a 33 example is DD0 @1 0 0 0 2 0 0 0 31 A: The rules of matrix multiplication cause all diagonal matrices (of the same size) to com- mute with each other. However, unless proportional to a unit matrix, diagonal matrices will not commute with nondiagonal matrices containing arbitrary elements. Matrix Inverse It will often be the case that given a square matrix A, there will be a square matrix Bsuch thatA BDB AD1. A matrix Bwith this property is called the inverse ofAand is given the name A1. IfA1exists, it must be unique. The proof of this statement is simple: If B andCare both inverses of A, then A BDB ADA CDC AD1: We now look at C A BD(C A)BDB;but also C A BDC(A B)DC: This shows that BDC. ArfKen_Ch02-9780123846549.tex 100 Chapter 2 Determinants and Matrices Every nonzero real (or complex) number has a nonzero multiplicative inverse, often written 1= . But the corresponding property does not hold for matrices; there exist nonzero matrices that do not have inverses. To demonstrate this, consider the following: AD1 1 0 0 ;BD1 0 1 0 ;so A BD0 0 0 0 : IfAhas an inverse, we can multiply the equation A BDOon the left byA1, thereby obtaining ABDO! A1ABDA1O! BDO: Since we started with a matrix Bthat was nonzero, this is an inconsistency, and we are forced to conclude that A1does not exist. A matrix without an inverse is said to be singu- lar, so our conclusion is that Ais singular. Note that in our derivation, we had to be careful to multiply both members of A BDOfrom the left, because multiplication is noncommu- tative. Alternatively, assuming B1to exist, we could multiply this equation on the right byB1, obtaining ABDO! ABB1DOB1! ADO: This is inconsistent with the nonzero Awith which we started; we conclude that Bis also singular. Summarizing, there are nonzero matrices that do not have inverses and are iden- tified as singular. The algebraic properties of real and complex numbers (including the existence of inverses for all nonzero numbers) define what mathematicians call a field. The properties we have identified for matrices are different; they form what is called a ring. The numerical inversion of matrices is another topic that has been given much attention, and computer programs for matrix inversion are widely available. A closed, but cumber- some formula for the inverse of a matrix exists; it expresses the elements of A1in terms of the determinants that are the minors of det.A/ ; recall that minors were defined in the para- graph immediately before Eq. (2.14). That formula, the derivation of which is in several of the Additional Readings, is .A1/i jD.1/iCjMji det.A/: (2.34) We describe here a well-known method that is computationally more efficient than Eq. (2.34), namely the Gauss-Jordan procedure. Example 2.2.3 GAUSS-JORDAN MATRIX INVERSION The Gauss-Jordan method is based on the fact that there exist matrices MLsuch that the product MLAwill leave an arbitrary matrix Aunchanged, except with (a) one row multiplied by a constant, or (b) one row replaced by the original row minus a multiple of another row, or (c) the interchange of two rows. The actual matrices MLthat carry out these transformations are the subject of Exercise 2.2.21. ArfKen_Ch02-9780123846549.tex 2.2 Matrices 101 By using these transformations, the rows of a matrix can be altered (by matrix multipli- cation) in the same ways we were able to change the elements of determinants, so we can proceed in ways similar to those employed for the reduction of determinants by Gauss elim- ination. If Ais nonsingular, the application of a succession of ML, i.e., MD.:::M00 LM0 LML/, can reduce Ato a unit matrix: M AD1; orMDA1: Thus, what we need to do is apply successive transformations to Auntil these transforma- tions have reduced Ato1, keeping track of the product of these transformations. The way in which we keep track is to successively apply the transformations to a unit matrix. Here is a concrete example. We want to invert the matrix AD0 @3 2 1 2 3 1 1 1 41 A: Our strategy will be to write, side by side, the matrix Aand a unit matrix of the same size, and to perform the same operations on each until Ahas been converted to a unit matrix, which means that the unit matrix will have been changed to A1. We start with 0 @3 2 1 2 3 1 1 1 41 A and0 @1 0 0 0 1 0 0 0 11 A: We multiply the rows as necessary to set to unity all elements of the first column of the left matrix: 0 BBBB@12 31 3 13 21 2 1 1 41 CCCCAand0 BBBB@1 30 0 01 20 0 0 11 CCCCA: Subtracting the first row from the second and third rows, we obtain 0 BBBBBB@12 31 3 05 61 6 01 311 31 CCCCCCAand0 BBBBBB@1 30 0 1 31 20 1 30 11 CCCCCCA: ArfKen_Ch02-9780123846549.tex 102 Chapter 2 Determinants and Matrices Then we divide the second row (of both matrices) by5 6and subtract2 3times it from the first row and1 3times it from the third row. The results for both matrices are 0 BBBBBB@1 01 5 0 11 5 0 018 51 CCCCCCAand0 BBBBBB@3 52 50 2 53 50 1 51 511 CCCCCCA: We divide the third row (of both matrices) by18 5. Then as the last step,1 5times the third row is subtracted from each of the first two rows (of both matrices). Our final pair is 0 @1 0 0 0 1 0 0 0 11 A and A1D0 BBBBBB@11 187 181 18 7 1811 181 18 1 181 185 181 CCCCCCA: We can check our work by multiplying the original Aby the calculated A1to see if we really do get the unit matrix 1.  Derivatives of Determinants The formula giving the inverse of a matrix in terms of its minors enables us to write a compact formula for the derivative of a determinant det.A/ where the matrix Ahas ele- ments that depend on some variable x. To carry out the differentiation with respect to the xdependence of its element ai j, we write det.A/ as its expansion in minors Mi jabout the elements of row i, as in Eq. (2.14), so, appealing also to Eq. (2.34), we have @det.A/ @ai jD.1/iCjMi jD.A1/jidet.A/: Applying now the chain rule to allow for the xdependence of all elements of A, we get ddet.A/ dxDdet.A/X i j.A1/jidai j dx: (2.35) Systems of Linear Equations Using the matrix inverse, we can write down formal solutions to linear equation systems. To start, we note that if Ais annsquare matrix, and xandharen1column vectors, the matrix equation AxDhis, by the rule for matrix multiplication, AxD0 BB@a11x1Ca12x2CC a1nxn a21x1Ca22x2CC a2nxn  an1x1Can2x2CC annxn1 CCADhD0 BB@h1 h2  hn1 CCA; ArfKen_Ch02-9780123846549.tex 2.2 Matrices 103 which is entirely equivalent to a system of nlinear equations with the elements of Aas coefficients. If Ais nonsingular, we can multiply AxDhon the left by A1, obtaining the result xDA1h. This result tells us two things: (1) that if we can evaluate A1, we can compute the solution x; and (2) that the existence of A1means that this equation system has a unique solution. In our study of determinants we found that a linear equation system had a unique solution if and only if the determinant of its coefficients was nonzero. We therefore see that the condition that A1exists, i.e., that Ais nonsingular, is the same as the condi- tion that the determinant of A, which we write det.A/ , be nonzero. This result is important enough to be emphasized: A square matrix Ais singular if and only if det.A/D0: (2.36) Determinant Product Theorem The connection between matrices and their determinants can be made deeper by estab- lishing a product theorem which states that the determinant of a product of two nn matrices AandBis equal to the products of the determinants of the individual matrices: det.A B/Ddet.A/ det.B/: (2.37) As an initial step toward proving this theorem, let us look at det.A B/ with the elements of the matrix product written out. Showing the first two columns explicitly, we have det.A B/D a11b11Ca12b21CC a1nbn1a11b12Ca12b22CC a1nbn2 a21b11Ca22b21CC a2nbn1a21b12Ca22b22CC a2nbn2     an1b11Can2b21CC annbn1an1b12Can2b22CC annbn2 : Introducing the notation ajD0 BB@a1j a2j  anj1 CCA;this becomes det.A B/D X j1aj1bj1;1X j2aj2bj2;2 ; where the summations over j1,j2, . . . , jnrun independently from 1 though n. Now, calling upon Eqs. (2.12) and(2.13), we can move the summations and the factors boutside the determinant, reaching det.A B/DX j1X j2X jnbj1;1bj2;2bjn;ndet.a j1aj2ajn/: (2.38) The determinant on the right-hand side of Eq. (2.38) will vanish if any of the indices j are equal; if all are unequal, that determinant will be det.A/ , with the sign corresponding to the parity of the column permutation needed to put the ajin numerical order. Both ArfKen_Ch02-9780123846549.tex 104 Chapter 2 Determinants and Matrices of these conditions are met by writing det.a j1aj2ajn/D"j1:::jndet.A/ , where"is the Levi-Civita symbol defined in Eq. (2.8). The above manipulations bring us to det.A B/Ddet.A/X j1:::jn"j1:::jnbj1;1bj2;2bjn;nDdet.A/ det.B/; where the final step was to invoke the definition of the determinant, Eq. (2.10). This result proves the determinant product theorem. From the determinant product theorem, we can gain additional insight regarding singular matrices. Noting first that a special case of the theorem is that det.A A1/Ddet.1/D1Ddet.A/ det.A1/; we see that det.A1/D1 det.A/: (2.39) It is now obvious that if det.A/D0, then det.A1/cannot exist, meaning that A1cannot exist either. This is a direct proof that a matrix is singular if and only if it has a vanishing determinant. Rank of a Matrix The concept of matrix singularity can be refined by introducing the notion of the rank of a matrix. If the elements of a matrix are viewed as the coefficients of a set of linear forms, as in Eq. (2.1) and its generalization to nvariables, a square matrix is assigned a rank equal to the number of linearly independent forms that its elements describe. Thus, a nonsingular nnmatrix will have rank n, while a nnsingular matrix will have a rank rless than n. The rank provides a measure of the extent of the singularity; if rDn1, the matrix describes one linear form that is dependent on the others; rDn2describes a situation in which there are two forms that are linearly dependent on the others, etc. We will in Chapter 6 take up methods for systematically determining the rank of a matrix. Transpose, Adjoint, Trace In addition to the operations we have already discussed, there are further operations that depend on the fact that matrices are arrays. One such operation is transposition. The transpose of a matrix is the matrix that results from interchanging its row and column indices. This operation corresponds to subjecting the array to reflection about its principal diagonal. If a matrix is not square, its transpose will not even have the same shape as the original matrix. The transpose of A, denotedQAor sometimes AT, thus has elements .QA/i jDaji: (2.40) ArfKen_Ch02-9780123846549.tex 2.2 Matrices 105 Note that transposition will convert a column vector into a row vector, so ifxD0 BB@x1 x2 ::: xn1 CCA;thenQxD.x1x2:::xn/: A matrix that is unchanged by transposition (i.e., QADA) is called symmetric. For matrices that may have complex elements, the complex conjugate of a matrix is defined as the matrix resulting if all elements of the original matrix are complex conju- gated. Note that this does not change the shape or move any elements to new positions. The notation for the complex conjugate of AisA. The adjoint of a matrix A, denoted A†, is obtained by both complex conjugating and transposing it (the same result is obtained if these operations are performed in either order). Thus, .A†/i jDa ji: (2.41) The trace, a quantity defined for square matrices, is the sum of the elements on the principal diagonal. Thus, for an nnmatrix A, trace.A/DnX iD1aii: (2.42) From the rule for matrix addition, is is obvious that trace.ACB/Dtrace.A/Ctrace.B/: (2.43) Another property of the trace is that its value for a product of two matrices AandBis independent of the order of multiplication: trace.AB/DX i.AB/ iiDX iX jai jbjiDX jX ibjiai j DX j.BA/ j jDtrace.BA/: (2.44) This holds even if AB6DBA. Equation (2.44) means that the trace of any commutator TA;BUD ABBAis zero. Considering now the trace of the matrix product ABC, if we group the factors as A(BC) , we easily see that trace.ABC/Dtrace.BCA/: Repeating this process, we also find trace.ABC/ Dtrace.CAB/ . Note, however, that we cannot equate any of these quantities to trace.CBA/ or to the trace of any other noncyclic permutation of these matrices. ArfKen_Ch02-9780123846549.tex 106 Chapter 2 Determinants and Matrices Operations on Matrix Products We have already seen that the determinant and the trace satisfy the relations det.AB/Ddet.A/ det.B/Ddet.BA/; trace.AB/Dtrace.BA/; whether or not AandBcommute. We also found that trace.A CB/Dtrace.A/Ctrace.B/ and can easily show that trace. A/D trace.A/ , establishing that the trace is a linear operator (as defined in Chapter 5). Since similar relations do not exist for the determinant, it isnota linear operator. We consider now the effect of other operations on matrix products. The transpose of a product,.AB/T, can be shown to satisfy .AB/TDQBQA; (2.45) showing that a product is transposed by taking, in reverse order, the transposes of its fac- tors. Note that if the respective dimensions of AandBare such as to make ABdefined, it will also be true that QBQAis defined. Since complex conjugation of a product simply amounts to conjugation of its individual factors, the formula for the adjoint of a matrix product follows a rule similar to Eq. (2.45): .AB/†DB†A†: (2.46) Finally, consider .AB/1. In order for ABto be nonsingular, neither AnorBcan be singular (to see this, consider their determinants). Assuming this nonsingularity, we have .AB/1DB1A1: (2.47) The validity of Eq. (2.47) can be demonstrated by substituting it into the obvious equation .AB/.AB/1D1. Matrix Representation of Vectors The reader may have already noted that the operations of addition and multiplication by a scalar are defined in identical ways for vectors (Section 1.7) and the matrices we are calling column vectors. We can also use the matrix formalism to generate scalar products, but in order to do so we must convert one of the column vectors into a row vector. The operation of transposition provides a way to do this. Thus, letting aandbstand for vectors in I R3, ab!.a1a2a3/0 @b1 b2 b31 ADa1b1Ca2b2Ca3b3: If in a matrix context we regard aandbas column vectors, the above equation assumes the form ab! aTb: (2.48) This notation does not really lead to significant ambiguity if we note that when dealing with matrices, we are using lower-case boldface symbols to denote column vectors. Note also that because aTbis a11matrix, it is synonymous with its transpose, which is bTa. The ArfKen_Ch02-9780123846549.tex 2.2 Matrices 107 matrix notation preserves the symmetry of the dot product. As in Section 1.7, the square of the magnitude of the vector corresponding to awill be aTa. If the elements of our column vectors aandbare real, then an alternate way of writing aTbisa†b. But these quantities are not equal if the vectors have complex elements, as will be the case in some situations in which the column vectors do not represent displacements in physical space. In that situation, the dagger notation is the more useful because then a†a will be real and can play the role of a magnitude squared. Orthogonal Matrices A real matrix (one whose elements are real) is termed orthogonal if its transpose is equal to its inverse. Thus, if Sis orthogonal, we may write S1DST;orSSTD1(Sorthogonal). (2.49) Since, for Sorthogonal, det.SST/Ddet.S/ det.ST/DTdet.S/U2D1, we see that det.S/D1 (Sorthogonal). (2.50) It is easy to prove that if SandS0are each orthogonal, then so also are SS0andS0S. Unitary Matrices Another important class of matrices consists of matrices Uwith the property that U†D U1, i.e., matrices for which the adjoint is also the inverse. Such matrices are identified as unitary. One way of expressing this relationship is U U†DU†UD1(Uunitary). (2.51) If all the elements of a unitary matrix are real, the matrix is also orthogonal. Since for any matrix det.AT/Ddet.A/ , and therefore det.A†/Ddet.A/, application of the determinant product theorem to a unitary matrix Uleads to det.U/ det.U†/Djdet.U/j2D1; (2.52) showing that det.U/ is a possibly complex number of magnitude unity. Since such numbers can be written in the form exp.i/, withreal, the determinants of UandU†will, for some, satisfy det.U/Dei;det.U†/Dei: Part of the significance of the term unitary is associated with the fact that the determinant has unit magnitude. A special case of this relationship is our earlier observation that if Uis real, and therefore also an orthogonal matrix, its determinant must be either C1or1. Finally, we observe that if UandVare both unitary, then UVandVUwill be unitary as well. This is a generalization of our earlier result that the matrix product of two orthogonal matrices is also orthogonal. ArfKen_Ch02-9780123846549.tex 108 Chapter 2 Determinants and Matrices Hermitian Matrices There are additional classes of matrices with useful characteristics. A matrix is identified as Hermitian, or, synonymously, self-adjoint, if it is equal to its adjoint. To be self-adjoint, a matrix Hmust be square, and in addition, its elements must satisfy .H†/i jD.H/i j! h jiDhi j(His Hermitian). (2.53) This condition means that the array of elements in a self-adjoint matrix exhibits a reflection symmetry about the principal diagonal: elements whose positions are connected by reflec- tion must be complex conjugates. As a corollary to this observation, or by direct reference to Eq. (2.53), we see that the diagonal elements of a self-adjoint matrix must be real. If all the elements of a self-adjoint matrix are real, then the condition of self-adjointness will cause the matrix also to be symmetric, so all real, symmetric matrices are self-adjoint (Hermitian). Note that if two matrices AandBare Hermitian, it is not necessarily true that ABorBA is Hermitian; however, ABCBA, if nonzero, will be Hermitian, and ABBA, if nonzero, will be anti-Hermitian, meaning that .ABBA/†D.ABBA/. Extraction of a Row or Column It is useful to define column vectors Oeiwhich are zero except for the .i;1/element, which is unity; examples are Oe1D0 BBBB@1 0 0  01 CCCCA;Oe2D0 BBBB@0 1 0  01 CCCCA;etc. (2.54) One use of these vectors is to extract a single column from a matrix. For example, if Ais a 33matrix, then AOe2D0 @a11a12a13 a21a22a23 a31a32a331 A0 @0 1 01 AD0 @a12 a22 a321 A: The row vectorOeT ican be used in a similar fashion to extract a row from an arbitrary matrix, as in OeT iAD.ai1ai2ai3/: These unit vectors will also have many uses in other contexts. Direct Product A second procedure for multiplying matrices, known as the direct tensor or Kronecker product, combines a mnmatrix Aand a m0n0matrix Bto make the direct product ArfKen_Ch02-9780123846549.tex 2.2 Matrices 109 matrix CDA B, which is of dimension mm0nn0and has elements C DAi jBkl; (2.55) with Dm0.i1/Ck, Dn0.j1/Cl. The direct product matrix uses the indices of the first factor as major and those of the second factor as minor; it is therefore a noncom- mutative process. It is, however, associative. Example 2.2.4 DIRECT PRODUCTS We give some specific examples. If AandBare both 22matrices, we may write, first in a somewhat symbolic and then in a completely expanded form, A BDa11Ba12B a21Ba22B D0 BB@a11b11a11b12a12b11a12b12 a11b21a11b22a12b21a12b22 a21b11a21b12a22b11a22b12 a21b21a21b22a22b21a22b221 CCA: Another example is the direct product of two two-element column vectors, xandy. Again writing first in symbolic, and then expanded form, x1 x2 y1 y2 Dx1y x2y D0 BB@x1y1 x1y2 x2y1 x2y21 CCA: A third example is the quantity ABfrom Example 2.2.2. It is an instance of the special case (column vector times row vector) in which the direct and inner products coincide: ABDA B.  IfCandC0are direct products of the respective forms CDA Band C0DA0 B0; (2.56) and these matrices are of dimensions such that the matrix inner products AA0andBB0are defined, then CC0D.AA0/ .BB0/: (2.57) Moreover, if matrices AandBare of the same dimensions, then C .ACB/DC ACC Band.ACB/ CDA CCB C: (2.58) Example 2.2.5 DIRAC MATRICES In the original, nonrelativistic formulation of quantum mechanics, agreement between theory and experiment for electronic systems required the introduction of the concept of electron spin (intrinsic angular momentum), both to provide a doubling in the number of available states and to explain phenomena involving the electron’s magnetic moment. The concept was introduced in a relatively ad hoc fashion; the electron needed to be given spin quantum number 1/2, and that could be done by assigning it a two-component wave ArfKen_Ch02-9780123846549.tex 110 Chapter 2 Determinants and Matrices function, with the spin-related properties described using the Pauli matrices, which were introduced in Example 2.2.1: 1D0 1 1 0 ;  2D0i i 0 ;  3D1 0 01 : Of relevance here is the fact that these matrices anticommute and have squares that are unit matrices: 2 iD12;andijCjiD0; i6Dj: (2.59) In 1927, P. A. M. Dirac developed a relativistic formulation of quantum mechanics applicable to spin-1/2 particles. To do this it was necessary to place the spatial and time variables on an equal footing, and Dirac proceeded by converting the relativistic expression for the kinetic energy to an expression that was first order in both the energy and the momentum (parallel quantities in relativistic mechanics). He started from the relativistic equation for the energy of a free particle, E2D.p2 1Cp2 2Cp2 3/c2Cm2c4Dp2c2Cm2c4; (2.60) where piare the components of the momentum in the coordinate directions, mis the particle mass, and cis the velocity of light. In the passage to quantum mechanics, the quantities piare to be replaced by the differential operators iNh@=@xi, and the entire equation is applied to a wave function. It was desirable to have a formulation that would yield a two-component wave function in the nonrelativistic limit and therefore might be expected to contain the i. Dirac made the observation that a key to the solution of his problem was to exploit the fact that the Pauli matrices, taken together as a vector D1Oe1C2Oe2C3Oe3; (2.61) could be combined with the vector pto yield the identity .p/2Dp212; (2.62) where 12denotes a 22unit matrix. The importance of Eq. (2.62) is that, at the price of going to 22matrices, we can linearize the quadratic occurrences of Eandpin Eq. (2.60) as follows. We first write E212c2.p/2Dm2c412: (2.63) We then factor the left-hand side of Eq. (2.63) and apply both sides of the resulting equation (which is a 22matrix equation) to a two-component wave function that we will call 1: .E12Ccp/.E12cp/ 1Dm2c4 1: (2.64) The meaning of this equation becomes clearer if we make the additional definition .E12cp/ 1Dmc2 2: (2.65) ArfKen_Ch02-9780123846549.tex 2.2 Matrices 111 Substituting Eq. (2.65) intoEq. (2.64), we can then write the modified Eq. (2.64) and the (unchanged) Eq. (2.65) as the equation set .E12Ccp/ 2Dmc2 1; .E12cp/ 1Dmc2 2I(2.66) both these equations will need to be satisfied simultaneously. To bring Eqs. (2.66) to the form actually used by Dirac, we now make the substitution 1D AC B, 2D A B, and then add and subtract the two equations from each other, reaching a set of coupled equations in Aand B: E Acp BDmc2 A; cp AE BDmc2 B: In anticipation of what we will do next, we write these equations in the matrix form E12 0 0E12 0 cp cp 0 A B Dmc2 A B : (2.67) We can now use the direct product notation to condense Eq. (2.67) into the simpler form [.3 12/E c.p/]9Dmc29; (2.68) where9is the four-component wave function built from the two-component wave functions: 9D A B ; and the terms on the left-hand side have the indicated structure because 3D1 0 01 and we define D0 1 1 0 : (2.69) It has become customary to identify the matrices in Eq. (2.68) as and to refer to them asDirac matrices, with 0D3 12D12 0 012 D0 BB@1 0 0 0 0 1 0 0 0 01 0 0 0 011 CCA: (2.70) The matrices resulting from the individual components of inEq. (2.68) are (for iD1;2;3) iD iD0i i0 : (2.71) ArfKen_Ch02-9780123846549.tex 112 Chapter 2 Determinants and Matrices Expanding Eq. (2.71), we have 1D0 BB@0 0 0 1 0 0 1 0 01 0 0 1 0 0 01 CCA; 2D0 BB@0 0 0i 0 0 i 0 0i0 0 i0 0 01 CCA; 3D0 BB@0 0 1 0 0 0 01 1 0 0 0 0 1 0 01 CCA: (2.72) Now that the have been defined, we can rewrite Eq. (2.68), expanding pinto components: h 0Ec. 1p1C 2p2C 3p3/i 9Dmc29: To put this matrix equation into the specific form known as the Dirac equation we multiply both sides of it (on the left) by 0. Noting that . 0/2D1and giving 0 ithe new name i, we reach h 0mc2Cc. 1p1C 2p2C 3p3/i 9DE9: (2.73) Equation (2.73) is in the notation used by Dirac with the exception that he used as the name for the matrix here called 0. The Dirac gamma matrices have an algebra that is a generalization of that exhibited by the Pauli matrices, where we found that the 2 iD1and that if i6Dj, theniand janticommute. Either by further analysis or by direct evaluation, it is found that, for D0;1;2;3andiD1;2;3, . 0/2D1; . i/2D1; (2.74)  iC i D0; 6Di: (2.75) In the nonrelativistic limit, the four-component Dirac equation for an electron reduces to a two-component equation in which each component satisfies the Schrödinger equation, with the Pauli and Dirac matrices having completely disappeared. See Exercise 2.2.48. In this limit, the Pauli matrices reappear if we add to the Schrödinger equation an addi- tional term arising from the intrinsic magnetic moment of the electron. The passage to the nonrelativistic limit provides justification for the seemingly arbitrary introduction of a two-component wavefunction and use of the Pauli matrices for discussions of spin angular momentum. The Pauli matrices (and the unit matrix 12) form what is known as a Clifford algebra,2 with the properties shown in Eq. (2.59). Since the algebra is based on 22matrices, it can have only four members (the number of linearly independent such matrices), and is said to be of dimension 4. The Dirac matrices are members of a Clifford algebra of dimen- sion 16. A complete basis for this Clifford algebra with convenient Lorentz transformation 2D. Hestenes, Am. J. Phys. 39: 1013 (1971); and J. Math. Phys. 16: 556 (1975). ArfKen_Ch02-9780123846549.tex 2.2 Matrices 113 properties consists of the 16 matrices 14; 5Di 0 1 2 3D012 120 ; .D0;1;2;3/; 5 .D0;1;2;3/; Di  .0<3/: (2.76)  Functions of Matrices Polynomials with one or more matrix arguments are well defined and occur often. Power series of a matrix may also be defined, provided the series converges for each matrix ele- ment. For example, if Ais any nnmatrix, then the power series exp.A/D1X jD01 jWAj; (2.77) sin.A/D1X jD0.1/j .2jC1/WA2jC1; (2.78) cos.A/D1X jD0.1/j .2j/WA2j(2.79) are well-defined nnmatrices. For the Pauli matrices k, the Euler identity for real andkD1;2;or 3, exp.ik/D12cosCiksin; (2.80) follows from collecting all even and odd powers of in separate series using 2 kD1. For the44Dirac matrices , defined in Eq. (2.76), we have for 1<3, exp.i/D14cosCisin; (2.81) while exp.i0k/D14coshCi0ksinh (2.82) holds for real because.i0k/2D1forkD1, 2, or 3. Hermitian and unitary matrices are related in that U, given as UDexp.iH/; (2.83) is unitary if His Hermitian. To see this, just take the adjoint: U†Dexp. iH†/D exp. iH/DTexp. iH/U1DU1. Another result which is important to identify here is that any Hermitian matrix Hsatisfies a relation known as the trace formula, det.exp.H//Dexp.trace.H//: (2.84) This formula is derived at Eq. (6.27). ArfKen_Ch02-9780123846549.tex 114 Chapter 2 Determinants and Matrices Finally, we note that the multiplication of two diagonal matrices produces a matrix that is also diagonal, with elements that are the products of the corresponding elements of the multiplicands. This result implies that an arbitrary function of a diagonal matrix will also be diagonal, with diagonal elements that are that function of the diagonal elements of the original matrix. Example 2.2.6 EXPONENTIAL OF A DIAGONAL MATRIX If a matrix Ais diagonal, then its nth power is also diagonal, with the original diagonal matrix elements raised to the nth power. For example, given 3D1 0 01 ; then .3/nD1 0 0.1/n : We can now compute e3D0 BBBB@1X nD01 nW0 01X nD0.1/n nW1 CCCCADe0 0e1 :  A final and important result is the Baker-Hausdorff formula, which, among other places is used in the coupled-cluster expansions that yield highly accurate electronic struc- ture calculations on atoms and molecules3: exp.T/A exp.T/DACTA,TUC1 2WTTA,TU; TUC1 3WTTTA,TU; TU;TUC: (2.85) Exercises 2.2.1 Show that matrix multiplication is associative, .AB/CDA.BC/ . 2.2.2 Show that .ACB/.AB/DA2B2 if and only if AandBcommute, TA, BUD 0: 3F. E. Harris, H. J. Monkhorst, and D. L. Freeman, Algebraic and Diagrammatic Methods in Many-Fermion Theory. New York: Oxford University Press (1992). ArfKen_Ch02-9780123846549.tex 2.2 Matrices 115 2.2.3 (a) Complex numbers, aCib, with aandbreal, may be represented by (or are iso- morphic with) 22matrices: aCib !a b b a : Show that this matrix representation is valid for (i) addition and (ii) multiplication. (b) Find the matrix corresponding to .aCib/1. 2.2.4 IfAis an nnmatrix, show that det.A/D.1/ndetA: 2.2.5 (a) The matrix equation A2D0does not imply AD0. Show that the most general 22matrix whose square is zero may be written as ab b2 a2ab ; where aandbare real or complex numbers. (b) If CDACB, in general detC6DdetACdetB: Construct a specific numerical example to illustrate this inequality. 2.2.6 Given KD0 @0 0 i i 0 0 01 01 A; show that KnDKKK.nfactors/D1 (with the proper choice of n;n6D0/. 2.2.7 Verify the Jacobi identity, TA;TB, CUUDTB;TA, CUUTC;TA, BUU: 2.2.8 Show that the matrices AD0 @0 1 0 0 0 0 0 0 01 A;BD0 @0 0 0 0 0 1 0 0 01 A;CD0 @0 0 1 0 0 0 0 0 01 A satisfy the commutation relations TA, BUD C;TA, CUD 0; andTB, CUD 0: ArfKen_Ch02-9780123846549.tex 116 Chapter 2 Determinants and Matrices 2.2.9 Let iD0 BB@0 1 0 0 1 0 0 0 0 0 0 1 0 01 01 CCA;jD0 BB@0 0 01 0 01 0 0 1 0 0 1 0 0 01 CCA; and kD0 BB@0 01 0 0 0 0 1 1 0 0 0 01 0 01 CCA: Show that (a) i2Dj2Dk2D1 , where 1is the unit matrix. (b) ijDjiDk, jkDkjDi, kiDikDj. These three matrices ( i, j, and k) plus the unit matrix 1form a basis for quaternions. An alternate basis is provided by the four 22matrices, i1;i2;i3, and 1, where the iare the Pauli spin matrices of Example 2.2.1. 2.2.10 A matrix with elements ai jD0forj<imay be called upper right triangular. The elements in the lower left (below and to the left of the main diagonal) vanish. Show that the product of two upper right triangular matrices is an upper right triangular matrix. 2.2.11 The three Pauli spin matrices are 1D0 1 1 0 ;  2D0i i 0 ;and3D1 0 01 : Show that (a).i/2D12, (b)ijDik,.i;j;k/D.1;2;3/or a cyclic permutation thereof, (c)ijCjiD2i j12;12is the 22unit matrix. 2.2.12 One description of spin-1 particles uses the matrices MxD1p 20 @0 1 0 1 0 1 0 1 01 A;MyD1p 20 @0i 0 i 0i 0 i 01 A; and MzD0 @1 0 0 0 0 0 0 011 A: ArfKen_Ch02-9780123846549.tex 2.2 Matrices 117 Show that (a)TMx;MyUDiMz, and so on (cyclic permutation of indices). Using the Levi-Civita symbol, we may write TMi;MjUDiX k"i jkMk: (b) M2M2 xCM2 yCM2 zD213, where 13is the 33unit matrix. (c)TM2;MiUD0, TMz;LCUDLC, TLC;LUD2Mz, where LCMxCiMyandLMxiMy. 2.2.13 Repeat Exercise 2.2.12, using the matrices for a spin of 3=2, MxD1 20 BB@0p 3 0 0p 3 0 2 0 0 2 0p 3 0 0p 3 01 CCA;MyDi 20 BB@0p 3 0 0p 3 02 0 0 2 0p 3 0 0p 3 01 CCA; and MzD1 20 BB@3 0 0 0 0 1 0 0 0 01 0 0 0 031 CCA: 2.2.14 IfAis a diagonal matrix, with all diagonal elements different, and AandBcommute, show that Bis diagonal. 2.2.15 IfAandBare diagonal, show that AandBcommute. 2.2.16 Show that trace.ABC/ Dtrace.CBA/ if any two of the three matrices commute. 2.2.17 Angular momentum matrices satisfy a commutation relation TMj;MkUDiMl;j;k;lcyclic: Show that the trace of each angular momentum matrix vanishes. 2.2.18 Aand Banticommute: ABDBA . Also, A2D1,B2D1. Show that trace.A/ D trace.B/D0. Note. The Pauli and Dirac matrices are specific examples. 2.2.19 (a) If two nonsingular matrices anticommute, show that the trace of each one is zero. (Nonsingular means that the determinant of the matrix is nonzero.) (b) For the conditions of part (a) to hold, AandBmust be nnmatrices with neven. Show that if nisodd, a contradiction results. 2.2.20 IfA1has elements .A1/i jDa.1/ i jDCji jAj; ArfKen_Ch02-9780123846549.tex 118 Chapter 2 Determinants and Matrices where Cjiis the jith cofactor ofjAj, show that A1AD1: Hence A1is the inverse of A(ifjAj6D 0). 2.2.21 Find the matrices MLsuch that the product MLAwill be Abut with: (a) The ith row multiplied by a constant k.ai j!kai j,jD1, 2, 3,:::/; (b) The ith row replaced by the original ith row minus a multiple of the mth row .ai j!ai jK amj,iD1, 2, 3,:::/; (c) The ith and mth rows interchanged .ai j!amj,amj!ai j,jD1;2;3;:::/ . 2.2.22 Find the matrices MRsuch that the product AMRwill be Abut with: (a) The ith column multiplied by a constant k.aji!kaji;jD1;2;3;:::/ ; (b) The ith column replaced by the original ith column minus a multiple of the mth column.aji!ajikajm;jD1;2;3;:::/ ; (c) The ith and mth columns interchanged .aji!ajm,ajm!aji,jD1, 2, 3,:::/. 2.2.23 Find the inverse of AD0 @3 2 1 2 2 1 1 1 41 A: 2.2.24 Matrices are far too useful to remain the exclusive property of physicists. They may appear wherever there are linear relations. For instance, in a study of population move- ment the initial fraction of a fixed population in each of nareas (or industries or religions, etc.) is represented by an n-component column vector P. The movement of people from one area to another in a given time is described by an nn(stochastic) matrix T. Here Ti jis the fraction of the population in the jth area that moves to the ith area. (Those not moving are covered by iDj.) With Pdescribing the initial population distribution, the final population distribution is given by the matrix equation TPDQ. From its definition,Pn iD1PiD1. (a) Show that conservation of people requires that nX iD1Ti jD1; jD1;2;:::; n: (b) Prove that nX iD1QiD1 continues the conservation of people. ArfKen_Ch02-9780123846549.tex 2.2 Matrices 119 2.2.25 Given a 66matrix Awith elements ai jD0:5jijj,i;jD0;1;2;:::; 5, find A1. ANS. A1D1 30 BBBBBB@42 0 0 0 0 2 52 0 0 0 02 52 0 0 0 02 52 0 0 0 02 52 0 0 0 0 2 41 CCCCCCA: 2.2.26 Show that the product of two orthogonal matrices is orthogonal. 2.2.27 IfAis orthogonal, show that its determinant D1 . 2.2.28 Show that the trace of the product of a symmetric and an antisymmetric matrix is zero. 2.2.29 Ais22and orthogonal. Find the most general form of ADa b c d : 2.2.30 Show that det.A/D.detA/Ddet.A†/: 2.2.31 Three angular momentum matrices satisfy the basic commutation relation TJx;JyUDiJz (and cyclic permutation of indices). If two of the matrices have real elements, show that the elements of the third must be pure imaginary. 2.2.32 Show that.AB/†DB†A†. 2.2.33 A matrix CDS†S. Show that the trace is positive definite unless Sis the null matrix, in which case trace .C/D0. 2.2.34 IfAandBare Hermitian matrices, show that .ABCBA/andi.ABBA/are also Her- mitian. 2.2.35 The matrix CisnotHermitian. Show that then CCC†andi.CC†/are Hermitian. This means that a non-Hermitian matrix may be resolved into two Hermitian parts, CD1 2.CCC†/C1 2ii.CC†/: This decomposition of a matrix into two Hermitian matrix parts parallels the decompo- sition of a complex number zintoxCiy, where xD.zCz/=2andyD.zz/=2i. 2.2.36 AandBare two noncommuting Hermitian matrices: ABBADiC: Prove that Cis Hermitian. 2.2.37 Two matrices AandBare each Hermitian. Find a necessary and sufficient condition for their product ABto be Hermitian. ANS.TA;BUD 0: ArfKen_Ch02-9780123846549.tex 120 Chapter 2 Determinants and Matrices 2.2.38 Show that the reciprocal (that is, inverse) of a unitary matrix is unitary. 2.2.39 Prove that the direct product of two unitary matrices is unitary. 2.2.40 Ifis the vector with the ias components given in Eq. (2.61), and pis an ordinary vector, show that .p/2Dp212; where 12is a22unit matrix. 2.2.41 Use the equations for the properties of direct products, Eqs. (2.57) and(2.58), to show that the four matrices ,D0;1;2;3, satisfy the conditions listed in Eqs. (2.74) and (2.75). 2.2.42 Show that 5,Eq. (2.76), anticommutes with all four . 2.2.43 In this problem, the summations are over D0;1;2;3. Define gDgby the relations g00D1I gkkD1; kD1;2;3I gD0; 6DI and define asPg . Using these definitions, show that (a)P  D2 , (b)P  D4g , (c)P   D2  . 2.2.44 IfMD1 2.1C 5/, where 5is given in Eq. (2.76), show that M2DM: Note that this equation is still satisfied if is replaced by any other Dirac matrix listed in Eq. (2.76). 2.2.45 Prove that the 16 Dirac matrices form a linearly independent set. 2.2.46 If we assume that a given 44matrix A(with constant elements) can be written as a linear combination of the 16 Dirac matrices (which we denote here as 0i) AD16X iD1ci0i; show that citrace.A0 i/: 2.2.47 The matrix CDi 2 0is sometimes called the charge conjugation matrix. Show that C C1D. /T. 2.2.48 (a) Show that, by substitution of the definitions of the matrices from Eqs. (2.70) and(2.72), that the Dirac equation, Eq. (2.73), takes the following form when written as 22blocks (with Land Scolumn vectors of dimension 2). Here ArfKen_Ch02-9780123846549.tex Additional Readings 121 LandSstand, respectively, for “large” and “small” because of their relative size in the nonrelativistic limit): mc2E c .1p1C2p2C3p3/ c.1p1C2p2C3p3/mc2E L S D0: (b) To reach the nonrelativistic limit, make the substitution EDmc2C"and approx- imate2mc2"by2mc2. Then write the matrix equation as two simultaneous two-component equations and show that they can be rearranged to yield 1 2m p2 1Cp2 2Cp2 3 LD" L; which is just the Schrödinger equation for a free particle. (c) Explain why is it reasonable to call Land S“large” and “small.” 2.2.49 Show that it is consistent with the requirements that they must satisfy to take the Dirac gamma matrices to be (in 22block form) 0D012 120 ; iD0i i0 ; .iD1;2;3/: This choice for the gamma matrices is called the Weyl representation. 2.2.50 Show that the Dirac equation separates into independent 22blocks in the Weyl rep- resentation (see Exercise 2.2.49) in the limit that the mass mapproaches zero. This observation is important in the ultra relativistic regime where the rest mass is inconse- quential, or for particles of negligible mass (e.g., neutrinos). 2.2.51 (a) Given r0DUr, with Ua unitary matrix and ra (column) vector with complex elements, show that the magnitude of ris invariant under this operation. (b) The matrix Utransforms any column vector rwith complex elements into r0; leaving the magnitude invariant: r†rDr0†r0. Show that Uis unitary. Additional Readings Aitken, A. C., Determinants and Matrices. New York: Interscience (1956), reprinted, Greenwood (1983). A read- able introduction to determinants and matrices. Barnett, S., Matrices: Methods and Applications. Oxford: Clarendon Press (1990). Bickley, W. G., and R. S. H. G. Thompson, Matrices—Their Meaning and Manipulation. Princeton, NJ: Van Nostrand (1964). A comprehensive account of matrices in physical problems, their analytic properties, and numerical techniques. Brown, W. C., Matrices and Vector Spaces. New York: Dekker (1991). Gilbert, J., and L. Gilbert, Linear Algebra and Matrix Theory. San Diego: Academic Press (1995). Golub, G. H., and C. F. Van Loan, Matrix Computations, 3rd ed. Baltimore: JHU Press (1996). Detailed mathe- matical background and algorithms for the production of numerical software, including methods for parallel computation. A classic computer science text. Heading, J., Matrix Theory for Physicists. London: Longmans, Green and Co. (1958). A readable introduction to determinants and matrices, with applications to mechanics, electromagnetism, special relativity, and quantum mechanics. Vein, R., and P. Dale, Determinants and Their Applications in Mathematical Physics. Berlin: Springer (1998). Watkins, D.S., Fundamentals of Matrix Computations. New York: Wiley (1991). ArfKen_Ch03-9780123846549.tex CHAPTER 3 VECTOR ANALYSIS The introductory section on vectors, Section 1.7, identified some basic properties that are universal, in the sense that they occur in a similar fashion in spaces of different dimension. In summary, these properties are (1) vectors can be represented as linear forms, with oper- ations that include addition and multiplication by a scalar, (2) vectors have a commutative and distributive dot product operation that associates a scalar with a pair of vectors and depends on their relative orientations and hence is independent of the coordinate system, and (3) vectors can be decomposed into components that can be identified as projections onto the coordinate directions. In Section 2.2 we found that the components of vectors could be identified as the elements of a column vector and that the scalar product of two vectors corresponded to the matrix multiplication of the transpose of one (the transposition makes it a row vector) with the column vector of the other. The current chapter builds on these ideas, mainly in ways that are specific to three- dimensional (3-D) physical space, by (1) introducing a quantity called a vector cross product to permit the use of vectors to represent rotational phenomena and volumes in 3-D space, (2) studying the transformational properties of vectors when the coordinate system used to describe them is rotated or subjected to a reflection operation, (3) developing math- ematical methods for treating vectors that are defined over a spatial region (vector fields), with particular attention to quantities that depend on the spatial variation of the vector field, including vector differential operators and integrals of vector quantities, and (4) extending vector concepts to curvilinear coordinate systems, which are very useful when the sym- metry of the coordinate system corresponds to a symmetry of the problem under study (an example is the use of spherical polar coordinates for systems with spherical symmetry). A key idea of the present chapter is that a quantity that is properly called a vector must have the transformation properties that preserve its essential features under coordinate transformation; there exist quantities with direction and magnitude that do not transform appropriately and hence are not vectors. This study of transformation properties will, in a subsequent chapter, ultimately enable us to generalize to related quantities such as tensors. 123 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch03-9780123846549.tex 124 Chapter 3 Vector Analysis Finally, we note that the methods developed in this chapter have direct application in electromagnetic theory as well as in mechanics, and these connections are explored through the study of examples. 3.1 R EVIEW OF BASIC PROPERTIES In Section 1.7 we established the following properties of vectors: 1. Vectors satisfy an addition law that corresponds to successive displacements that can be represented by arrows in the underlying space. Vector addition is commutative and associative: ACBDBCAand.ACB/CCDAC.BCC/. 2. A vector Acan be multiplied by a scalar k; ifk>0the result will be a vector in the direction of Abut with its length multiplied by k; ifk<0the result will be in the direction opposite to Abut with its length mutiplied by jkj. 3. The vector ABis interpreted as AC.1/B , so vector polynomials, e.g., A2BC 3C, are well-defined. 4. A vector of unit length in the coordinate direction xiis denotedOei. An arbitrary vector Acan be written as a sum of vectors along the coordinate directions, as ADA1Oe1CA2Oe2C: TheAiare called the components of A, and the operations in Properties 1 to 3 cor- respond to the component formulas GDA2BC3CD) GiDAi2BiC3Ci;(each i): 5. The magnitude or length of a vector A, denotedjAjorA, is given in terms of its components as jAjD A2 1CA2 2C1=2: 6. The dot product of two vectors is given by the formula ABDA1B1CA2B2CI consequences are jAj2DAA; ABDjAjjBj cos; whereis the angle between AandB. 7. If two vectors are perpendicular to each other, their dot product vanishes and they are termed orthogonal. The unit vectors of a Cartesian coordinate system are orthogonal: OeiOejDi j; (3.1) wherei jis the Kronecker delta, Eq. (1.164). 8. The projection of a vector in any direction has an algebraic magnitude given by its dot product with a unit vector in that direction. In particular, the projection of Aon theOeidirection is AiOei, with AiDOeiA: ArfKen_Ch03-9780123846549.tex 3.1 Review of Basic Properties 125 9. The components of AinR3are related to its direction cosines (cosines of the angles thatAmakes with the coordinate axes) by the formulas AxDAcos ; AyDAcos ; AzDAcos ; andcos2 Ccos2 Ccos2 D1. In Section 2.2 we noted that matrices consisting of a single column could be used to represent vectors. In particular, we found, illustrating for the 3-D space R3, the following properties. 10. A vector Acan be represented by a single-column matrix awhose elements are the components of A, as in AD) aD0 @A1 A2 A31 A: The rows (i.e., individual elements Ai) of aare the coefficients of the individual members of the basis used to represent A, so the element Aiis associated with the basis unit vectorOei. 11. The vector operations of addition and multiplication by a scalar correspond exactly to the operations of the same names applied to the single-column matrices represent- ing the vectors, as illustrated here: GDA2BC3CD)0 @G1 G2 G31 AD0 @A1 A2 A31 A20 @B1 B2 B31 AC30 @C1 C2 C31 A D0 @A12B1C3C1 A22B2C3C2 A32B3C3C31 A;orgDa2bC3c: It is therefore appropriate to call these single-column matrices column vectors. 12. The transpose of the matrix representing a vector Ais a single-row matrix, called a row vector: aTD.A1A2A3/: The operations illustrated in Property 11 also apply to row vectors. 13. The dot product ABcan be evaluated as aTb, or alternatively, because aandbare real, as a†b. Moreover, aTbDbTa. ABDaTbD.A1A2A3/0 @B1 B2 B31 ADA1B1CA2B2CA3B3: ArfKen_Ch03-9780123846549.tex 126 Chapter 3 Vector Analysis 3.2 V ECTORS IN 3-D S PACE We now proceed to develop additional properties for vectors, most of which are applicable only for vectors in 3-D space. Vector or Cross Product A number of quantities in physics are related to angular motion or the torque required to cause angular acceleration. For example, angular momentum about a point is defined as having a magnitude equal to the distance rfrom the point times the component of the linear momentum pperpendicular to r—the component of pcausing angular motion (see Fig. 3.1). The direction assigned to the angular momentum is that perpendicular to both randp, and corresponds to the axis about which angular motion is taking place. The mathematical construction needed to describe angular momentum is the cross product, defined as CDABD.ABsin/Oec: (3.2) Note that C, the result of the cross product, is stated to be a vector, with a magnitude that is the product of the magnitudes of A,Band the sine of the angle between AandB. The direction of C, i.e., that ofOec, is perpendicular to the plane of AandB, such that A,B, andCform a right-handed system.1This causes Cto be aligned with the rotational axis, with a sign that indicates the sense of the rotation. From Fig. 3.2, we also see that ABhas a magnitude equal to the area of the parallel- ogram formed by AandB, and with a direction normal to the parallelogram. Other places the cross product is encountered include the formulas vD!rand FMDqvB: The first of these equations is the relation between linear velocity vand and angular veloc- ity!, and the second equation gives the force FMon a particle of charge qand velocity v in the magnetic induction field B(in SI units). y p sin θ xrθp FIGURE 3.1 Angular momentum about the origin, LDrp. Lhas magnitude rpsinand is directed out of the plane of the paper. 1The inherent ambiguity in this statement can be resolved by the following anthropomorphic prescription: Point the right hand in the direction A, and then bend the fingers through the smaller of the two angles that can cause the fingers to point in the direction B; the thumb will then point in the direction of C. ArfKen_Ch03-9780123846549.tex 3.2 Vectors in 3-D Space 127 xB Ay B sin θ θ FIGURE 3.2 Parallelogram of AB. We can get our right hands out of the analysis by compiling some algebraic properties of the cross product. If the roles of AandBare reversed, the cross product changes sign, so BADAB(anticommutation): (3.3) The cross product also obeys the distributive laws A.BCC/DABCAC; k.AB/D.kA/B; (3.4) and when applied to unit vectors in the coordinate directions, we get OeiOejDX k"i jkOek: (3.5) Here"i jkis the Levi-Civita symbol defined in Eq. (2.8); Eq. (3.5) therefore indicates, for example, thatOexOexD0,OexOeyDOez, butOeyOexDOez. Using Eq. (3.5) and writing AandBin component form, we can expand ABto obtain CDABD.AxOexCAyOeyCAzOez/.BxOexCByOeyCBzOez/ D.AxByAyBx/.OexOey/C.AxBzAzBx/.OexOez/ C.AyBzAzBy/.OeyOez/ D.AxByAyBx/OezC.AxBzAzBx/.Oey/C.AyBzAzBy/Oex: (3.6) The components of Care important enough to be displayed prominently: CxDAyBzAzBy;CyDAzBxAxBz;CzDAxByAyBx; (3.7) equivalent to CiDX jk"i jkAjBk: (3.8) ArfKen_Ch03-9780123846549.tex 128 Chapter 3 Vector Analysis Yet another way of expressing the cross product is to write it as a determinant. It is straightforward to verify that Eqs. (3.7) are reproduced by the determinantal equation CD OexOeyOez AxAyAz BxByBz : (3.9) when the determinant is expanded in minors of its top row. The anticommutation of the cross product now clearly follows if the rows for the components of AandBare inter- changed. We need to reconcile the geometric form of the cross product, Eq. (3.2), with the alge- braic form in Eq. (3.6). We can confirm the magnitude of ABby evaluating (from the component form of C) .AB/.AB/DA2B2.AB/2DA2B2A2B2cos2 DA2B2sin2: (3.10) The first step in Eq. (3.10) can be verified by expanding its left-hand side in component form, then collecting the result into the terms constituting the central member of the first line of the equation. To confirm the direction of CDAB, we can check that ACDBCD0, showing thatC(in component form) is perpendicular to both AandB. We illustrate for AC: ACDAx.AyBzAzBy/CAy.AzBxAxBz/CAz.AxByAyBx/D0: (3.11) To verify the sign of C, it suffices to check special cases (e.g., ADOex,BDOey, orAxD ByD1, all other components zero). Next, we observe that it is obvious from Eq. (3.2) that if CDABin a given coordinate system, then that equation will also be satisfied if we rotate the coordinates, even though the individual components of all three vectors will thereby be changed. In other words, the cross product, like the dot product, is a rotationally invariant relationship. Finally, note that the cross product is a quantity specifically defined for 3-D space. It is possible to make analogous definitions for spaces of other dimensionality, but they do not share the interpretation or utility of the cross product in R3. Scalar Triple Product While the various vector operations can be combined in many ways, there are two combi- nations involving three operands that are of particular importance. We call attention first to the scalar triple product, of the form A.BC/. Taking.BC/in the determinantal form, Eq. (3.9), one can see that taking the dot product with Awill cause the unit vector Oex to be replaced by Ax, with corresponding replacements to OeyandOez. The overall result is A.BC/D AxAyAz BxByBz CxCyCz : (3.12) We can draw a number of conclusions from this highly symmetric determinantal form. To start, we see that the determinant contains no vector quantities, so it must evaluate ArfKen_Ch03-9780123846549.tex 3.2 Vectors in 3-D Space 129 y xB×C BC Az FIGURE 3.3 A.BC/parallelepiped. to an ordinary number. Because the left-hand side of Eq. (3.12) is a rotational invariant, the number represented by the determinant must also be rotationally invariant, and can therefore be identified as a scalar. Since we can permute the rows of the determinant (with a sign change for an odd permutation, and with no sign change for an even permutation), we can permute the vectors A,B, and Cto obtain ABCDBCADCABDACB;etc. (3.13) Here we have followed common practice and dropped the parentheses surrounding the cross product, on the basis that they must be understood to be present in order for the expressions to have meaning. Finally, noting that BChas a magnitude equal to the area of the BCparallelog ram and a direction perpendicular to it, and that the dot product with Awill multiply that area by the projection of AonBC, we see that the scalar triple product gives us ./the volume of the parallelepiped defined by A,B, and C; see Fig. 3.3. Example 3.2.1 RECIPROCAL LATTICE Leta,b, and c(not necessarily mutually perpendicular) represent the vectors that define a crystal lattice. The displacements from one lattice point to another may then be written RDnaaCnbbCncc; (3.14) ArfKen_Ch03-9780123846549.tex 130 Chapter 3 Vector Analysis with na,nb, and nctaking integral values. In the band theory of solids,2it is useful to introduce what is called a reciprocal lattice a0,b0,c0such that aa0Dbb0Dcc0D1; (3.15) and with ab0Dac0Dba0Dbc0Dca0Dcb0D0: (3.16) The reciprocal-lattice vectors are easily constructed by calling on the fact that for any u andv,uvis perpendicular to both uandv; we have a0Dbc abc;b0Dca abc;c0Dab abc: (3.17) The scalar triple product causes these expressions to satisfy the scale condition of Eq. (3.15).  Vector Triple Product The other triple product of importance is the vector triple product, of the form A .BC/. Here the parentheses are essential since, for example, .OexOex/OeyD0, while Oex.OexOey/DOexOezDOey. Our interest is in reducing this triple product to a simpler form; the result we seek is A.BC/DB.AC/C.AB/: (3.18) Equation (3.18), which for convenience we will sometimes refer to as the BAC–CAB rule, can be proved by inserting components for all vectors and evaluating all the products, but it is instructive to proceed in a more elegant fashion. Using the formula for the cross product in terms of the Levi-Civita symbol, Eq. (3.8), we write A.BC/DX iOeiX jk"i jkAj X pq"kpqBpCq! DX i jX pqOeiAjBpCqX k"i jk"kpq: (3.19) The summation over kof the product of Levi-Civita symbols reduces, as shown in Exercise 2.1.9, to ipjqiqjp; we are left with A.BC/DX i jOeiAj.BiCjBjCi/DX iOei0 @BiX jAjCjCiX jAjBj1 A; which is equivalent to Eq. (3.18). 2It is often chosen to require aa0, etc., to be 2rather than unity, because when Bloch states for a crystal (labeled by k) are set up, a constituent atomic function in cell Renters with coefficient exp.ikR/, and if kis changed by a reciprocal lattice step (in, say, the a0direction), the coefficient becomes exp.iTkCa0UR/, which reduces to exp.2 ina/exp.ikR/and therefore, because exp.2 ina/D1, to its original value. Thus, the reciprocal lattice identifies the periodicity in k. The unit cell of the k vectors is called the Brillouin zone ArfKen_Ch03-9780123846549.tex 3.2 Vectors in 3-D Space 131 Exercises 3.2.1 IfPDOexPxCOeyPyandQDOexQxCOeyQyare any two nonparallel (also nonantiparal- lel) vectors in the xy-plane, show that PQis in the z-direction. 3.2.2 Prove that.AB/.AB/D.AB/2.AB/2: 3.2.3 Using the vectors PDOexcosCOeysin; QDOexcos'Oeysin'; RDOexcos'COeysin'; prove the familiar trigonometric identities sin.C'/Dsincos'Ccossin'; cos.C'/Dcoscos'sinsin': 3.2.4 (a) Find a vector Athat is perpendicular to UD2OexCOeyOez; VDOexOeyCOez: (b) What is Aif, in addition to this requirement, we demand that it have unit magnitude? 3.2.5 If four vectors a;b;c, and dall lie in the same plane, show that .ab/.cd/D0: Hint. Consider the directions of the cross-product vectors. 3.2.6 Derive the law of sines (see Fig. 3.4): sin jAjDsin jBjDsin jCj: 3.2.7 The magnetic induction Bisdefined by the Lorentz force equation, FDq.vB/: AC B βα γ FIGURE 3.4 Plane triangle. ArfKen_Ch03-9780123846549.tex 132 Chapter 3 Vector Analysis Carrying out three experiments, we find that if vDOex;F qD2Oez4Oey; vDOey;F qD4OexOez; vDOez;F qDOey2Oex: From the results of these three separate experiments calculate the magnetic induction B. 3.2.8 You are given the three vectors A,B, and C, ADOexCOey; BDOeyCOez; CDOexOez: (a) Compute the scalar triple product, ABC. Noting that ADBCC, give a geometric interpretation of your result for the scalar triple product. (b) Compute A.BC/. 3.2.9 Prove Jacobi’s identity for vector products: a.bc/Cb.ca/Cc.ab/D0: 3.2.10 A vector Ais decomposed into a radial vector Arand a tangential vector At:IfOris a unit vector in the radial direction, show that (a) ArDOr.AOr/and (b) AtDOr.OrA/. 3.2.11 Prove that a necessary and sufficient condition for the three (nonvanishing) vectors A, B, and Cto be coplanar is the vanishing of the scalar triple product ABCD0: 3.2.12 Three vectors A,B, and Care given by AD3Oex2OeyC2Oz; BD6OexC4Oey2Oz; CD3Oex2Oey4Oz: Compute the values of ABCandA.BC/;C.AB/andB.CA/: 3.2.13 Show that .AB/.CD/D.AC/.BD/.AD/.BC/: 3.2.14 Show that .AB/.CD/D.ABD/C.ABC/D: ArfKen_Ch03-9780123846549.tex 3.3 Coordinate Transformations 133 3.2.15 An electric charge q1moving with velocity v1produces a magnetic induction B given by BD0 4q1v1Or r2(mks units), whereOris a unit vector that points from q1to the point at which Bis measured (Biot and Savart law). (a) Show that the magnetic force exerted by q1on a second charge q2, velocity v2, is given by the vector triple product F2D0 4q1q2 r2v2 v1Or : (b) Write out the corresponding magnetic force F1thatq2exerts on q1. Define your unit radial vector. How do F1andF2compare? (c) Calculate F1andF2for the case of q1andq2moving along parallel trajectories side by side. ANS. (b) F1D0 4q1q2 r2v1 v2Or : In general, there is no simple relation between F1andF2. Specifically, Newton’s third law, F1DF 2, does not hold. (c) F1D0 4q1q2 r2v2OrDF 2: Mutual attraction. 3.3 C OORDINATE TRANSFORMATIONS As indicated in the chapter introduction, an object classified as a vector must have specific transformation properties under rotation of the coordinate system; in particular, the com- ponents of a vector must transform in a way that describes the same object in the rotated system. Rotations Considering initially R2, and a rotation of the coordinate axes as shown in Fig. 3.5, we wish to find how the components AxandAyof a vector Ain the unrotated system are related to A0 xandA0 y, its components in the rotated coordinate system. Perhaps the easiest way to answer this question is by first asking how the unit vectors OexandOeyare represented in the new coordinates, after which we can perform vector addition on the new incarnations ofAxOexandAyOey. From the right-hand part of Fig. 3.5, we see that OexDcos'Oe0 xsin'Oe0 y;andOeyDsin'Oe0 xCcos'Oe0 y; (3.20) ArfKen_Ch03-9780123846549.tex 134 Chapter 3 Vector Analysis exˆ exˆeyˆeyˆ ϕ ϕϕ ϕˆe/prime y ˆe/prime x sinϕ −ˆe/prime ycosϕ ˆe/prime xsinϕ ˆe/prime x cosϕ ˆe/prime y FIGURE 3.5 Left: Rotation of two-dimensional (2-D) coordinate axes through angle '. Center and right: Decomposition of OexandOeyinto their components in the rotated system. so the unchanged vector Anow takes the changed form ADAxOexCAyOeyDAx.cos'Oe0 xsin'Oe0 y/CAy.sin'Oe0 xCcos'Oe0 y/ D.Axcos'CAysin'/Oe0 xC.Axsin'CAycos'/Oe0 y: (3.21) If we write the vector Ain the rotated (primed) coordinate system as ADA0 xOe0 xCA0 yOe0 y; we then have A0 xDAxcos'CAysin'; A0 yDAxsin'CAycos'; (3.22) which is equivalent to the matrix equation A0DA0 x A0 y Dcos'sin' sin'cos'Ax Ay : (3.23) Suppose now that we start from Aas given by its components in the rotated system, .A0 x;A0 y/, and rotate the coordinate system back to its original orientation. This will entail a rotaton in the amount ', and corresponds to the matrix equation Ax Ay Dcos.'/ sin.'/ sin.'/ cos.'/A0 x A0 y Dcos'sin' sin' cos'A0 x A0 y : (3.24) Assigning the 22matrices in Eqs. (3.23) and(3.24) the respective names SandS0, we see that these two equations are equivalent to A0DSAandADS0A0, with SDcos'sin' sin'cos' and S0Dcos'sin' sin' cos' : (3.25) Now, applying StoAand then S0toSA(corresponding to first rotating the coordinate system an amountC'and then an amount'), we recover A, or ADS0SA: Since this result must be valid for any A, we conclude that S0DS1. We also see that S0DST. We can check that SS0D1by matrix multiplication: SS0Dcos'sin' sin'cos'cos'sin' sin' cos' D1 0 0 1 : ArfKen_Ch03-9780123846549.tex 3.3 Coordinate Transformations 135 Since Sis real, the fact that S1DSTmeans that it is orthogonal. In summary, we have found that the transformation connecting AandA0(the same vector, but represented in the rotated coordinate system) is A0DSA; (3.26) withSan orthogonal matrix. Orthogonal Transformations It was no accident that the transformation describing a rotation in R2wasorthogonal, by which we mean that the matrix effecting the transformation was an orthogonal matrix. An instructive way of writing the transformation Sis, returning to Eq. (3.20), to rewrite those equations as OexD.Oe0 xOex/Oe0 xC.Oe0 yOex/Oe0 y;OeyD.Oe0 xOey/Oe0 xC.Oe0 yOey/Oe0 y: (3.27) This corresponds to writing OexandOeyas the sum of their projections on the orthogonal vectorsOe0 xandOe0 y. Now we can rewrite Sas SDOe0 xOexOe0 xOey Oe0 yOexOe0 yOey : (3.28) This means that each row of Scontains the components (in the unprimed coordinates) of a unit vector (eitherOe0 xorOe0 y) that is orthogonal to the vector whose components are in the other row. In turn, this means that the dot products of different row vectors will be zero, while the dot product of any row vector with itself (because it is a unit vector) will be unity. That is the deeper significance of an orthogonal matrix S; theelement of SSTis the dot product formed from the th row of Sand theth column of ST(which is the same as theth row of S). Since these row vectors are orthogonal, we will get zero if 6D, and because they are unit vectors, we will get unity if D. In other words, SSTwill be a unit matrix. Before leaving Eq. (3.28), note that its columns also have a simple interpretation: Each contains the components (in the primed coordinates) of one of the unit vectors of the unprimed set. Thus the dot product formed from two different columns ofSwill van- ish, while the dot product of any column with itself will be unity. This corresponds to the fact that, for an orthogonal matrix, we also have STSD1. Summarizing part of the above, The transformation from one orthogonal Cartesian coordinate system to another Carte- sian system is described by an orthogonal matrix. In Chapter 2 we found that an orthogonal matrix must have a determinant that is real and of magnitude unity, i.e., 1. However, for rotations in ordinary space the value of the determinant will always be C1. One way to understand this is to consider the fact that any rotation can be built up from a large number of small rotations, and that the determinant must vary continuously as the amount of rotation is changed. The identity rotation (i.e., no rotation at all) has determinant C1. Since no value close to C1exceptC1itself is a permitted value for the determinant, rotations cannot change the value of the determinant. ArfKen_Ch03-9780123846549.tex 136 Chapter 3 Vector Analysis Reflections Another possibility for changing a coordinate system is to subject it to a reflection operation. For simplicity, consider first the inversion operation, in which the sign of each coordinate is reversed. In R3, the transformation matrix Swill be the 33analog of Eq. (3.28), and the transformation under discussion is to set Oe0 DOe, withDx,y, andz. This will lead to SD0 @1 0 0 01 0 0 011 A; which clearly results in detSD1 . The change in sign of the determinant corresponds to the change from a right-handed to a left-handed coordinate system (which obviously cannot be accomplished by a rotation). Reflection about a plane (as in the image produced by a plane mirror) also changes the sign of the determinant and the handedness of the coordinate system; for example, reflection in the xy-plane changes the sign of Oez, leaving the other two unit vectors unchanged; the transformation matrix Sfor this transformation is SD0 @1 0 0 0 1 0 0 011 A: Its determinant is also 1. The formulas for vector addition, multiplication by a scalar, and the dot product are unaffected by a reflection transformation of the coordinates, but this is not true of the cross product. To see this, look at the formula for any one of the components of AB, and how it would change under inversion (where the same, unchanged vectors in physical space now have sign changes to all their components): Cx:AyBzAzBy!.Ay/.Bz/.Az/.By/DAyBzAzBy: Note that this formula says that the sign of Cxshould not change, even though it must in order to describe the unchanged physical situation. The conclusion is that our transforma- tion law fails for the result of a cross-product operation. However, the mathematics can be salvaged if we classify BCas a different type of quantity than BandC. Many texts on vector analysis call vectors whose components change sign under coordinate reflec- tionpolar vectors, and those whose components do not then change sign axial vectors. The term axial doubtless arises from the fact that cross products frequently describe phe- nomena associated with rotation about the axis defined by the axial vector. Nowadays, it is becoming more usual to call polar vectors justvectors, because we want that term to describe objects that obey for all Sthe transformation law A0DSA (vectors); (3.29) (and specifically without a restriction to Swhose determinants are C1). Axial vectors, for which the vector transformation law fails for coordinate reflections, are then referred to aspseudovectors, and their transformation law can be expressed in the somewhat more complicated form C0Ddet.S/ SC (pseudovectors): (3.30) ArfKen_Ch03-9780123846549.tex 3.3 Coordinate Transformations 137 A zxB By Ax AA yz x FIGURE 3.6 Inversion (right) of original coordinates (left) and the effect on a vector Aand a pseudovector B. The effect of an inversion operation on a coordinate system and on a vector and a pseu- dovector are shown in Fig. 3.6. Since vectors and pseudovectors have different transformation laws, it is in general with- out physical meaning to add them together.3It is also usually meaningless to equate quan- tities of different transformational properties: in ADB, both quantities must be either vectors or pseudovectors. Pseudovectors, of course, enter into more complicated expressions, of which an example is the scalar triple product ABC. Under coordinate reflection, the components of BC do not change (as observed earlier), but those of Aare reversed, with the result that ABCchanges sign. We therefore need to reclassify it as a pseudoscalar. On the other hand, the vector triple product, A.BC/, which contains two cross products, evaluates, as shown in Eq. (3.18), to an expression containing only legitimate scalars and (polar) vectors. It is therefore proper to identify A.BC/as a vector. These cases illustrate the general principle that a product with an odd number of pseudo quantities is “pseudo,” while those with even numbers of pseudo quantities are not. Successive Operations One can carry out a succession of coordinate rotations and/or reflections by applying the relevant orthogonal transformations. In fact, we already did this in our introductory discus- sion for R2where we applied a rotation and then its inverse. In general, if RandR0refer to such operations, the application to AofRfollowed by the application of R0corresponds to A0DS.R0/S.R/A; (3.31) and the overall result of the two transformations can be identified as a single transformation whose matrix S.R0R/is the matrix product S.R0/S.R/. 3The big exception to this is in beta-decay weak interactions. Here the universe distinguishes between right- and left-handed systems, and we add polar and axial vector interactions. ArfKen_Ch03-9780123846549.tex 138 Chapter 3 Vector Analysis Two points should be noted: 1. The operations take place in right-to-left order: The rightmost operator is the one applied to the original A; that to its left then applies to the result of the first opera- tion, etc. 2. The combined operation R0Ris a transformation between two orthogonal coordinate systems and therefore can be described by an orthogonal matrix: The product of two orthogonal matrices is orthogonal. Exercises 3.3.1 A rotation'1C'2about the z-axis is carried out as two successive rotations '1and '2, each about the z-axis. Use the matrix representation of the rotations to derive the trigonometric identities cos.' 1C'2/Dcos'1cos'2sin'1sin'2; sin.' 1C'2/Dsin'1cos'2Ccos'1sin'2: 3.3.2 A corner reflector is formed by three mutually perpendicular reflecting surfaces. Show that a ray of light incident upon the corner reflector (striking all three surfaces) is reflected back along a line parallel to the line of incidence. Hint. Consider the effect of a reflection on the components of a vector describing the direction of the light ray. 3.3.3 Letxandybe column vectors. Under an orthogonal transformation S, they become x0DSxandy0DSy. Show that.x0/Ty0DxTy, a result equivalent to the invariance of the dot product under a rotational transformation. 3.3.4 Given the orthogonal transformation matrix Sand vectors aandb, SD0 @0:80 0:60 0:00 0:48 0:64 0:60 0:360:48 0:801 A;aD0 @1 0 11 A;bD0 @0 2 11 A; (a) Calculate det.S/ . (b) Verify that abis invariant under application of Stoaandb. (c) Determine what happens to abunder application of Stoaandb. Is this what is expected? 3.3.5 Using aandbas defined in Exercise 3.3.4, but with SD0 @0:60 0:00 0:80 0:640:60 0:48 0:48 0:80 0:361 A and cD0 @2 1 31 A; (a) Calculate det.S/ . Apply Stoa,b, and c, and determine what happens to (b) ab; ArfKen_Ch03-9780123846549.tex 3.4 Rotations in R3139 (c).ab/c; (d) a.bc/: (e) Classify the expressions in (b) through (d) as scalar, vector, pseudovector, or pseu- doscalar. 3.4 R OTATIONS IN R3 Because of its practical importance, we discuss now in some detail the treatment of rotations in R3. An obvious starting point, based on our experience in R2, would be to write the 33matrix Sof Eq. (3.28), with rows that describe the orientations of a rotated (primed) set of unit vectors in terms of the original (unprimed) unit vectors: SD0 B@Oe0 1Oe1Oe0 1Oe2Oe0 1Oe3 Oe0 2Oe1Oe0 2Oe2Oe0 2Oe3 Oe0 3Oe1Oe0 3Oe2Oe0 3Oe31 CA (3.32) We have switched the coordinate labels from x,y,zto 1, 2, 3 for convenience in some of the formulas that use Eq. (3.32). It is useful to make one observation about the elements ofS, namely sDOe0 Oe. This dot product is the projection of Oe0 onto theOedirection, and is therefore the change in xthat is produced by a unit change in x0 . Since the relation between the coordinates is linear, we can identify Oe0 Oeas@x=@x0 , so our transformation matrix Scan be written in the alternate form SD0 B@@x1=@x0 1@x2=@x0 1@x3=@x0 1 @x1=@x0 2@x2=@x0 2@x3=@x0 2 @x1=@x0 3@x2=@x0 3@x3=@x0 31 CA: (3.33) The argument we made to evaluate Oe0 Oecould as easily have been made with the roles of the two unit vectors reversed, yielding instead of @x=@x0 the derivative @x0 =@x. We then have what at first may seem to be a surprising result: @x @x0D@x0  @x: (3.34) A superficial look at this equation suggests that its two sides would be reciprocals. The problem is that we have not been notationally careful enough to avoid ambiguity: the derivative on the left-hand side is to be taken with the other x0coordinates fixed, while that on the right-hand side is with the other unprimed coordinates fixed. In fact, the equality in Eq. (3.34) is needed to make San orthogonal matrix. We note in passing that the observation that the coordinates are related linearly restricts the current discussion to Cartesian coordinate systems. Curvilinear coordinates are treated later. Neither Eq. (3.32) norEq. (3.33) makes obvious the possibility of relations among the elements of S. InR2, we found that all the elements of Sdepended on a single variable, the rotation angle. In R3, the number of independent variables needed to specify a general rotation is three: Two parameters (usually angles) are needed to specify the direction of Oe0 3; then one angle is needed to specify the direction of Oe0 1in the plane perpendicular to Oe0 3; ArfKen_Ch03-9780123846549.tex 140 Chapter 3 Vector Analysis at this point the orientation of Oe0 2is completely determined. Therefore, of the nine elements ofS, only three are in fact independent. The usual parameters used to specify R3rotations are the Euler angles.4It is useful to have Sgiven explicitly in terms of them, as the Lagrangian formulation of mechanics requires the use of a set of independent variables. The Euler angles describe an R3rotation in three steps, the first two of which have the effect of fixing the orientation of the new Oe3axis (the polar direction in spherical coordinates), while the third Euler angle indicates the amount of subsequent rotation about that axis. The first two steps do more than identify a new polar direction; they describe rotations that cause the realignment. As a result, we can obtain the matrix representations of these (and the third rotation), and apply them sequentially (i.e., as a matrix product) to obtain the overall effect of the rotation. The three steps describing rotation of the coordinate axes are the following (also illus- trated in Fig. 3.7): 1. The coordinates are rotated about the Oe3axis counterclockwise (as viewed from posi- tiveOe3) through an angle in the range 0 <2, into new axes denoted Oe0 1,Oe0 2,Oe0 3. (The polar direction is not changed; the Oe3andOe0 3axes coincide.) 2. The coordinates are rotated about the Oe0 2axis counterclockwise (as viewed from posi- tiveOe0 2) through an angle in the range 0 , into new axes denoted Oe00 1,Oe00 2,Oe00 3. (This tilts the polar direction toward the Oe0 1direction, but leaves Oe0 2unchanged.) 3. The coordinates are now rotated about the Oe00 3axis counterclockwise (as viewed from positiveOe00 3) through an angle in the range 0 <2, into the final axes, denoted Oe000 1,Oe000 2,Oe000 3. (This rotation leaves the polar direction, Oe00 3, unchanged.) In terms of the usual spherical polar coordinates .r;;'/ , the final polar axis is at the orientationD ,'D . The final orientations of the other axes depend on all three Euler angles. We now need the transformation matrices. The first rotation causes Oe0 1andOe0 2to remain in the xy-plane, and has in its first two rows and columns exactly the same form (b) (c)x/prime/prime 3 x/prime/prime 1x2x2x1x/prime1 x/prime1 x/prime 2ββ β αααx3= x/prime 3 x3= x/prime3x3= x/prime3 x/prime 2 = (a)γ γ=x/prime/prime 3 x/prime/prime 1x/prime/prime 2x/prime/prime/prime 3 =x/prime 2x/prime/prime 2x/prime/prime/prime 1 FIGURE 3.7 Euler angle rotations: (a) about Oe3through angle ; (b) aboutOe0 2through angle ; (c) aboutOe00 3through angle . 4There are almost as many definitions of the Euler angles as there are authors. Here we follow the choice generally made by workers in the area of group theory and the quantum theory of angular momentum. ArfKen_Ch03-9780123846549.tex 3.4 Rotations in R3141 asSin Eq. (3.25): S1. /D0 @cos sin 0 sin cos 0 0 0 11 A: (3.35) The third row and column of S1indicate that this rotation leaves unchanged the Oe3com- ponent of any vector on which it operates. The second rotation (applied to the coordinate system as it exists after the first rotation) is in the Oe0 3Oe0 1plane; note that the signs of sin have to be consistent with a cyclic permutation of the axis numbering: S2. /D0 @cos 0sin 0 1 0 sin 0 cos 1 A: The third rotation is like the first, but with rotation amount : S3. /D0 @cos sin 0 sin cos 0 0 0 11 A: The total rotation is described by the triple matrix product S. ; ; /DS3. /S 2. /S 1. /: (3.36) Note the order: S1. /operates first, then S2. /, and finally S3. /. Direct multiplication gives S. ; ; /D 0 @cos cos cos sin sin cos cos sin Csin cos cos sin sin cos cos cos sin sin cos sin Ccos cos sin sin sin cos sin sin cos 1 A: (3.37) In case they are wanted, note that the elements si jin Eq. (3.37) give the explicit forms of the dot productsOe000 iOej(and therefore also the partial derivatives @xi=@x000 j). Note that each of S1,S2, andS3are orthogonal, with determinant C1, so that the overall Swill also be orthogonal with determinant C1. Example 3.4.1 ANR3ROTATION Consider a vector originally with components .2;1;3/. We want its components in a coordinate system reached by Euler angle rotations D D D=2. Evaluating S. ; ; / : S. ; ; /D0 @1 0 0 0 0 1 0 1 01 A: A partial check on this value of Sis obtained by verifying that det.S/DC1 . ArfKen_Ch03-9780123846549.tex 142 Chapter 3 Vector Analysis Then, in the new coordinates, our vector has components 0 @1 0 0 0 0 1 0 1 01 A0 @2 1 31 AD0 @2 3 11 A: The reader should check this result by visualizing the rotations involved.  Exercises 3.4.1 Another set of Euler rotations in common use is (1) a rotation about the x3-axis through an angle ', counterclockwise, (2) a rotation about the x0 1-axis through an angle , counterclockwise, (3) a rotation about the x00 3-axis through an angle , counterclockwise. If D'=2 'D C=2 D orD D C=2 D =2; show that the final systems are identical. 3.4.2 Suppose the Earth is moved (rotated) so that the north pole goes to 30north, 20west (original latitude and longitude system) and the 10west meridian points due south (also in the original system). (a) What are the Euler angles describing this rotation? (b) Find the corresponding direction cosines. ANS..b/SD0 @0:95510:25520:1504 0:0052 0:5221 0:8529 0:2962 0:8138 0:50001 A: 3.4.3 Verify that the Euler angle rotation matrix, Eq. (3.37), is invariant under the transfor- mation ! C; ! ; ! : 3.4.4 Show that the Euler angle rotation matrix S. ; ; / satisfies the following relations: (a)S1. ; ; /DQS. ; ; /; (b)S1. ; ; /DS. ; ; /. 3.4.5 The coordinate system .x;y;z/is rotated through an angle 8counterclockwise about an axis defined by the unit vector Oninto system.x0;y0;z0/. In terms of the new coordinates ArfKen_Ch03-9780123846549.tex 3.5 Differential Vector Operators 143 the radius vector becomes r0Drcos8Crnsin8COn.Onr/.1cos8/: (a) Derive this expression from geometric considerations. (b) Show that it reduces as expected for OnDOez. The answer, in matrix form, appears in Eq. (3.35). (c) Verify that r02Dr2. 3.5 D IFFERENTIAL VECTOR OPERATORS We move now to the important situation in which a vector is associated with each point in space, and therefore has a value (its set of components) that depends on the coordinates specifying its position. A typical example in physics is the electric field E.x;y;z/, which describes the direction and magnitude of the electric force if a unit “test charge” was placed atx;y;z. The term field refers to a quantity that has values at all points of a region; if the quantity is a vector, its distribution is described as a vector field. While we already have a standard name for a simple algebraic quantity which is assigned a value at all points of a spatial region (it is called a function), in physics contexts it may also be referred to as a scalar field. Physicists need to be able to characterize the rate at which the values of vectors (and also scalars) change with position, and this is most effectively done by introducing differential vector operator concepts. It turns out that there are a large number of relations between these differential operators, and it is our current objective to identify such relations and learn how to use them. Gradient, r Our first differential operator is that known as the gradient, which characterizes the change of a scalar quantity, here ', with position. Working in R3, and labeling the coordinates x1, x2,x3, we write'.r/ as the value of 'at the point rDx1Oe1Cx2Oe2Cx3Oe3, and consider the effect of small changes dx1,dx2,dx3, respectively, in x1,x2, and x3. This situation corresponds to that discussed in Section 1.9, where we introduced partial derivatives to describe how a function of several variables (there x,y, and z) changes its value when these variables are changed by respective amounts dx,dy, and dz. The equation governing this process is Eq. (1.141). To first order in the differentials dxi,'in our present problem changes by an amount d'D@' @x1 dx1C@' @x2 dx2C@' @x3 dx3; (3.38) which is of the form corresponding to the dot product of r'D0 @@'=@ x1 @'=@ x2 @'=@ x31 A and drD0 @dx1 dx2 dx31 A: ArfKen_Ch03-9780123846549.tex 144 Chapter 3 Vector Analysis These quantities can also be written r'D@' @x1 Oe1C@' @x2 Oe2C@' @x3 Oe3; (3.39) drDdx1Oe1Cdx2Oe2Cdx3Oe3; (3.40) in terms of which we have d'D.r'/dr: (3.41) We have given the 31matrix of derivatives the name r'(often referred to in speech as “del phi” or “grad phi”); we give the differential of position its customary name dr. The notation of Eqs. (3.39) and (3.41) is really only appropriate if r'is actually a vector, because the utility of the present approach depends on our ability to use it in coor- dinate systems of arbitrary orientation. To prove that r'is a vector, we must show that it transforms under rotation of the coordinate system according to .r'/0DS.r'/: (3.42) Taking Sin the form given in Eq. (3.33), we examine S.r'/. We have S.r'/D0 B@@x1=@x0 1@x2=@x0 1@x3=@x0 1 @x1=@x0 2@x2=@x0 2@x3=@x0 2 @x1=@x0 3@x2=@x0 3@x3=@x0 31 CA0 @@'=@ x1 @'=@ x2 @'=@ x31 A D0 BBBBBBBBBB@3X D1@x @x0 1@' @x 3X D1@x @x0 2@' @x 3X D1@x @x0 3@' @x1 CCCCCCCCCCA: (3.43) Each of the elements in the final expression in Eq. (3.43) is a chain-rule expression for @'=@ x0 ,D1;2;3, showing that the transformation did produce .r'/0, the representa- tion of r'in the rotated coordinates. Having now established the legitimacy of the form r', we proceed to give ra life of its own. We therefore define (calling the coordinates x,y,z) rDOex@ @xCOey@ @yCOez@ @z: (3.44) We note that ris avector differential operator, capable of operating on a scalar (such as') to produce a vector as the result of the operation. Because a differential operator only operates on what is to its right, we have to be careful to maintain the correct order in expressions involving r, and we have to use parentheses when necessary to avoid ambi- guity as to what is to be differentiated. ArfKen_Ch03-9780123846549.tex 3.5 Differential Vector Operators 145 The gradient of a scalar is extremely important in physics and engineering, as it expresses the relation between a force field F.r/ experienced by an object at rand the related potential V.r/, F.r/Dr V.r/: (3.45) The minus sign in Eq. (3.45) is important; it causes the force exerted by the field to be in a direction that lowers the potential. We consider later (in Section 3.9) the conditions that must be satisfied if a potential corresponding to a given force can exit. The gradient has a simple geometric interpretation. From Eq. (3.41), we see that, if dris constrained to have a fixed magnitude, the direction of drthat maximizes d'will be when r'anddrare collinear. So, the direction of most rapid increase in 'is the gradient direction, and the magnitude of the gradient is the directional derivative of 'in that direction. We now see that rV, inEq. (3.45), is the direction of most rapid decrease inV, and is the direction of the force associated with the potential V. Example 3.5.1 GRADIENT OF rn As a first step toward computation of rrn, let’s look at the even simpler rr. We begin by writing rD.x2Cy2Cz2/1=2, from which we get @r @xDx .x2Cy2Cz2/1=2Dx r;@r @yDy r;@r @zDz r: (3.46) From these formulas we construct rrDx rOexCy rOeyCz rOezD1 r.xOexCyOeyCzOez/Dr r: (3.47) The result is a unit vector in the direction of r, denotedOr. For future reference, we note that OrDx rOexCy rOeyCz rOez (3.48) and that Eq. (3.47) takes the form rrDOr: (3.49) The geometry of randOris illustrated in Fig. 3.8. y ϕxryrˆ xx/ry/r FIGURE 3.8 Unit vectorOr(inxy-plane). ArfKen_Ch03-9780123846549.tex 146 Chapter 3 Vector Analysis Continuing now to rrn, we have @rn @xDnrn1@r @x; with corresponding results for the yandzderivatives. We get rrnDnrn1rrDnrn1Or: (3.50)  Example 3.5.2 COULOMB’S LAW In electrostatics, it is well known that a point charge produces a potential proportional to1=r, where ris the distance from the charge. To check that this is consistent with the Coulomb force law, we compute FDr1 r : This is a case of Eq. (3.50) with nD1 , and we get the expected result FD1 r2Or:  Example 3.5.3 GENERAL RADIAL POTENTIAL Another situation of frequent occurrence is that the potential may be a function only of the radial distance from the origin, i.e., 'Df.r/. We then calculate @' @xDd f.r/ dr@r @x;etc., which leads, invoking Eq. (3.49), to r'Dd f.r/ drrrDd f.r/ drOr: (3.51) This result is in accord with intuition; the direction of maximum increase in 'must be radial, and numerically equal to d'=dr .  Divergence, r Thedivergence of a vector Ais defined as the operation rAD@Ax @xC@Ay @yC@Az @z: (3.52) The above formula is exactly what one might expect given both the vector and differential- operator character of r. After looking at some examples of the calculation of the divergence, we will discuss its physical significance. ArfKen_Ch03-9780123846549.tex 3.5 Differential Vector Operators 147 Example 3.5.4 DIVERGENCE OF COORDINATE VECTOR Calculate rr: rrD Oex@ @xCOey@ @yCOez@ @z OexxCOeyyCOezz D@x @xC@y @yC@z @z; which reduces to rrD3.  Example 3.5.5 DIVERGENCE OF CENTRAL FORCE FIELD Consider next rf.r/Or. Using Eq. (3.48), we write rf.r/OrD Oex@ @xCOey@ @yCOez@ @z x f.r/ rOexCy f.r/ rOeyCz f.r/ rOez : D@ @xx f.r/ r C@ @yy f.r/ r C@ @zz f.r/ r : Using @ @xx f.r/ r Df.r/ rx f.r/ r2@r @xCx rd f.r/ dr@r @xDf.r/1 rx2 r3 Cx2 r2d f.r/ dr and corresponding formulas for the yandzderivatives, we obtain after simplification rf.r/OrD2f.r/ rCd f.r/ dr: (3.53) In the special case f.r/Drn, Eq. (3.53) reduces to rrnOrD.nC2/rn1: (3.54) FornD1, this reduces to the result of Example 3.5.4. For nD2 , corresponding to the Coulomb field, the divergence vanishes, except at rD0, where the differentiations we performed are not defined.  If a vector field represents the flow of some quantity that is distributed in space, its divergence provides information as to the accumulation or depletion of that quantity at the point at which the divergence is evaluated. To gain a clearer picture of the concept, let us suppose that a vector field v.r/ represents the velocity of a fluid5at the spatial points r, and that.r/ represents the fluid density at rat a given time t. Then the direction and magnitude of the flow rate at any point will be given by the product .r/v.r/ . Our objective is to calculate the net rate of change of the fluid density in a volume element at the point r. To do so, we set up a parallelepiped of dimensions dx,dy,dz centered at rand with sides parallel to the xy,xz, and yzplanes. See Fig. 3.9. To first order (infinitesimal dranddt), the density of fluid exiting the parallelepiped per unit time 5It may be helpful to think of the fluid as a collection of molecules, so the number per unit volume (the density) at any point is affected by the flow in and out of a volume element at the point. ArfKen_Ch03-9780123846549.tex 148 Chapter 3 Vector Analysis −ρvx⏐x−dx/2+ρvx⏐x+dx/2 dxdydz FIGURE 3.9 Outward flow of vfrom a volume element in the xdirections. The quantitiesv xmust be multiplied by dy dz to represent the total flux through the bounding surfaces at xdx=2. through the yzface located at x.dx=2/will be Flow out, face at xdx 2:.vx/ .xdx=2;y;z/dy dz: Note that only the velocity component vxis relevant here. The other components of v will not cause motion through a yzface of the parallelepiped. Also, note the following: dy dz is the area of the yzface; the average of vxover the face is to first order its value at .xdx=2;y;z/, as indicated, and the amount of fluid leaving per unit time can be identified as that in a column of area dy dz and heightvx. Finally, keep in mind that outward flow corresponds to that in the xdirection, explaining the presence of the minus sign. We next compute the outward flow through the yzplanar face at xCdx=2. The result is Flow out, face at xCdx 2:C.vx/ .xCdx=2;y;z/dy dz: Combining these, we have for both yzfaces  .v x/ xdx=2C.vx/ xCdx=2 dy dzD@.v x/ @x dx dy dz: Note that in combining terms at xdx=2andxCdx=2we used the partial derivative notation, because all the quantities appearing here are also functions of yandz. Finally, adding corresponding contributions from the other four faces of the parallelepiped, we reach Net flow outper unit timeD@ @x.vx/C@ @y.v y/C@ @z.vz/ dx dy dz Dr.v/dx dy dz: (3.55) We now see that the name divergence is aptly chosen. As shown in Eq. (3.55), the divergence of the vector vrepresents the net outflow per unit volume, per unit time. If the physical problem being described is one in which fluid (molecules) are neither created or destroyed, we will also have an equation of continuity, of the form @ @tCr.v/D0: (3.56) This equation quantifies the obvious statement that a net outflow from a volume element results in a smaller density inside the volume. When a vector quantity is divergenceless (has zero divergence) in a spatial region, we can interpret it as describing a steady-state “fluid-conserving” flow (flux) within that region ArfKen_Ch03-9780123846549.tex 3.5 Differential Vector Operators 149 A BC (b) (a) FIGURE 3.10 Flow diagrams: (a) with source and sink; (b) solenoidal. The divergence vanishes at volume elements A and C, but is negative at B. (even if the vector field does not represent material that is moving). This is a situation that arises frequently in physics, applying in general to the magnetic field, and, in charge-free regions, also to the electric field. If we draw a diagram with lines that follow the flow paths, the lines (depending on the context) may be called stream lines orlines of force. Within a region of zero divergence, these lines must exit any volume element they enter; they cannot terminate there. However, lines will begin at points of positive divergence (sources) and end at points where the divergence is negative (sinks). Possible patterns for a vector field are shown in Fig. 3.10. If the divergence of a vector field is zero everywhere, its lines of force will consist entirely of closed loops, as in Fig. 3.10(b); such vector fields are termed solenoidal. For emphasis, we write rBD0everywhere! Bis solenoidal. (3.57) Curl,r Another possible operation with the vector operator ris to take its cross product with a vector. Using the established formula for the cross product, and being careful to write the derivatives to the left of the vector on which they are to act, we obtain rVDOex@ @yVz@ @zVy COey@ @zVx@ @xVz COez@ @xVy@ @yVx D OexOeyOez @=@x@=@y@=@z VxVyVz : (3.58) This vector operation is called the curl ofV. Note that when the determinant in Eq. (3.58) is evaluated, it must be expanded in a way that causes the derivatives in the second row to be applied to the functions in the third row (and not to anything in the top row); we will encounter this situation repeatedly, and will identify the evaluation as being from the top down. Example 3.5.6 CURL OF A CENTRAL FORCE FIELD Calculate rTf.r/OrU. Writing OrDx rOexCy rOeyCz rOez; ArfKen_Ch03-9780123846549.tex 150 Chapter 3 Vector Analysis and remembering that @r=@yDy=rand@r=@zDz=r, the x-component of the result is found to be  rTf.r/OrU xD@ @yz f.r/ r@ @zy f.r/ r Dzd drf.r/ r@r @yyd drf.r/ r@r @z Dzd drf.r/ ry ryd drf.r/ rz rD0: By symmetry, the other components are also zero, yielding the final result rTf.r/OrUD0: (3.59)  Example 3.5.7 A NONZERO CURL Calculate FDr.yOexCxOey/, which is of the form rb, where bxDy,byDx, bzD0. We have FxD@bz @y@by @zD0; FyD@bx @z@vz @xD0; FzD@by @x@bx @yD2; soFD2Oez.  The results of these two examples can be better understood from a geometric interpreta- tion of the curl operator. We proceed as follows: Given a vector field B, consider the line integralH Bdsfor a small closed path. The circle through the integral sign is a signal that the path is closed. For simplicity in the computations, we take a rectangular path in thexy-plane, centered at a point .x0;y0/, of dimensions 1x1y, as shown in Fig. 3.11. We will traverse this path in the counterclockwise direction, passing through the four seg- ments labeled 1 through 4 in the figure. Since everywhere in this discussion zD0, we do not show it explicitly. xy 43 2 1 22x0−Δx , y0−Δy 22x0+Δx , y0−Δy22x0+Δx , y0+Δy 22x0−Δx , y0+Δy FIGURE 3.11 Path for computing circulation at .x0;y0/. ArfKen_Ch03-9780123846549.tex 3.5 Differential Vector Operators 151 Segment 1 of the path contributes to the integral Segment 1Dx0C1x=2Z x01x=2Bx.x;y01y=2/dxBx.x0;y01y=2/1 x; where the approximation, replacing Bxby its value at the middle of the segment, is good to first order. In a similar fashion, we have Segment 2Dy0C1y=2Z y01y=2By.x0C1x=2;y/dyBy.x0C1x=2;y0/1y; Segment 3Dx01x=2Z x0C1x=2Bx.x;y0C1y=2/dxBx.x0;y0C1y=2/1 x; Segment 4Dy01y=2Z y0C1y=2By.x01x=2;y/dyBy.x01x=2;y0/1y: Note that because the paths of segments 3 and 4 are in the direction of decrease in the value of the integration variable, we obtain minus signs in the contributions of these segments. Combining the contributions of Segments 1 and 3, and those of Segments 2 and 4, we have Segments 1C3D Bx.x0;y01y=2/Bx.x0;y0C1y=2/ 1x@Bx @y1y1x; Segments 2C4D By.x0C1x=2;y0/By.x01x=2;y0/ 1yC@By @x1x1y: Combining these contributions to obtain the value of the entire line integral, we have I Bds@By @x@Bx @y 1x1yTrBUz1x1y: (3.60) The thing to note is that a nonzero closed-loop line integral of Bcorresponds to a nonzero value of the component of rBnormal to the loop. In the limit of a small loop, the line integral will have a value proportional to the loop area; the value of the line integral per unit area is called the circulation (in fluid dynamics, it is also known as the vorticity). A nonzero circulation corresponds to a pattern of stream lines that form closed loops. Obviously, to form a closed loop, a stream line must curl; hence the name of the r operator. Returning now to Example 3.5.6, we have a situation in which the lines of force must be entirely radial; there is no possibility to form closed loops. Accordingly, we found this example to have a zero curl. But, looking next at Example 3.5.7, we have a situation in which the stream lines of yOexCxOeyform counterclockwise circles about the origin, and the curl is nonzero. ArfKen_Ch03-9780123846549.tex 152 Chapter 3 Vector Analysis We close the discussion by noting that a vector whose curl is zero everywhere is termed irrotational. This property is in a sense the opposite of solenoidal, and deserves a parallel degree of emphasis: rBD0everywhere! Bis irrotational. (3.61) Exercises 3.5.1 IfS.x;y;z/D x2Cy2Cz23=2, find (a)rS at the point .1;2;3/, (b) the magnitude of the gradient of S,jrSjat.1;2;3/, and (c) the direction cosines of rSat.1;2;3/. 3.5.2 (a) Find a unit vector perpendicular to the surface x2Cy2Cz2D3 at the point.1;1;1/. (b) Derive the equation of the plane tangent to the surface at .1;1;1/. ANS. (a)OexCOeyCOez =p 3, (b) xCyCzD3: 3.5.3 Given a vector r12DOex.x1x2/COey.y1y2/COez.z1z2/, show that r1r12(gradient with respect to x1,y1, and z1of the magnitude r12) is a unit vector in the direction of r12. 3.5.4 If a vector function Fdepends on both space coordinates .x;y;z/and time t, show that dFD.drr/FC@F @tdt: 3.5.5 Show that r.uv/DvruCurv, where uandvare differentiable scalar functions of x;y;andz. 3.5.6 For a particle moving in a circular orbit rDOexrcos!tCOeyrsin!t: (a) Evaluate rPr;withPrDdr=dtDv: (b) Show thatRrC!2rD0withRrDdv=dt . Hint. The radius rand the angular velocity !are constant. ANS. (a)Oez!r2: 3.5.7 Vector Asatisfies the vector transformation law, Eq. (3.26). Show directly that its time derivative dA=dt also satisfies Eq. (3.26) and is therefore a vector. 3.5.8 Show, by differentiating components, that (a)d dt.AB/DdA dtBCAdB dt; ArfKen_Ch03-9780123846549.tex 3.6 Differential Vector Operators: Further Properties 153 (b)d dt.AB/DdA dtBCAdB dt; just like the derivative of the product of two algebraic functions. 3.5.9 Prover.ab/Db.ra/a.rb/. Hint. Treat as a scalar triple product. 3.5.10 Classically, orbital angular momentum is given by LDrp, where pis the lin- ear momentum. To go from classical mechanics to quantum mechanics, pis replaced (in units withNhD1) by the operatorir. Show that the quantum mechanical angular momentum operator has Cartesian components LxDi y@ @zz@ @y ; LyDi z@ @xx@ @z ; LzDi x@ @yy@ @x : 3.5.11 Using the angular momentum operators previously given, show that they satisfy com- mutation relations of the form TLx;LyULxLyLyLxDi Lz and hence LLDiL: These commutation relations will be taken later as the defining relations of an angular momentum operator. 3.5.12 With the aid of the results of Exercise 3.5.11, show that if two vectors aandbcommute with each other and with L, that is,Ta;bUDTa; LUDTb; LUD 0, show that TaL;bLUD i.ab/L: 3.5.13 Prove that the stream lines of bin of Example 3.5.7 are counterclockwise circles. 3.6 D IFFERENTIAL VECTOR OPERATORS : FURTHER PROPERTIES Successive Applications of r Interesting results are obtained when we operate with ron the differential vector operator forms we have already introduced. The possible results include the following: (a)rr' (b)rr' (c)r.rV/ (d)r.rV/ (e)r.rV/: ArfKen_Ch03-9780123846549.tex 154 Chapter 3 Vector Analysis All five of these expressions involve second derivatives, and all five appear in the second-order differential equations of mathematical physics, particularly in electromag- netic theory. Laplacian The first of these expressions, rr', the divergence of the gradient, is named the Laplacian of '. We have rr'D Oex@ @xCOey@ @yCOez@ @z  Oex@' @xCOey@' @yCOez@' @z D@2' @x2C@2' @y2C@2' @z2: (3.62) When'is the electrostatic potential, we have rr'D0 (3.63) at points where the charge density vanishes, which is Laplace’s equation of electrostatics. Often the combination rris written r2, or1in the older European literature. Example 3.6.1 LAPLACIAN OF A CENTRAL FIELD POTENTIAL Calculate r2'.r/. Using Eq. (3.51) to evaluate r'and then Eq. (3.53) for the divergence, we have r2'.r/Drr'.r/Drd'.r/ drOerD2 rd'.r/ drCd2'.r/ dr2: We get a term in addition to d2'=dr2becauseOerhas a direction that depends on r. In the special case '.r/Drn, this reduces to r2rnDn.nC1/rn2: This vanishes for nD0('Dconstant) and for nD1 (Coulomb potential). For nD1 , our derivation fails for rD0, where the derivatives are undefined.  Irrotational and Solenoidal Vector Fields Expression (b), the second of our five forms involving two roperators, may be written as a determinant: rr'D OexOeyOez @=@x@=@y@=@z @'=@ x@'=@ y@'=@ z D OexOeyOez @=@x@=@y@=@z @=@x@=@y@=@z 'D0: Because the determinant is to be evaluated from the top down, it is meaningful to move 'outside and to its right, leaving a determinant with two identical rows and yielding the indicated value of zero. We are thereby actually assuming that the order of the partial ArfKen_Ch03-9780123846549.tex 3.6 Differential Vector Operators: Further Properties 155 differentiations can be reversed, which is true so long as these second derivatives of 'are continuous. Expression (d) is a scalar triple product that may be written r.rV/D @=@x@=@y@=@z @=@x@=@y@=@z Vx Vy Vz D0: This determinant also has two identical rows and yields zero if Vhas sufficient continuity. These two vanishing results tell us that any gradient has a vanishing curl and is therefore irrotational, and that any curl has a vanishing divergence, and is therefore solenoidal. These properties are of such importance that we set them out here in display form: rr'D0;all', (3.64) r.rV/D0;allV. (3.65) Maxwell’s Equations The unification of electric and magnetic phenomena that is encapsulated in Maxwell’s equations provides an excellent example of the use of differential vector operators. In SI units, these equations take the form rBD0; (3.66) rED "0; (3.67) rBD"00@E @tC0J; (3.68) rED@B @t: (3.69) Here Eis the electric field, Bis the magnetic induction field, is the charge density, Jis the current density, "0is the electric permittivity, and 0is the magnetic permeability, so "00D1=c2, where cis the velocity of light. Vector Laplacian Expressions (c) and (e) in the list at the beginning of this section satisfy the relation r.rV/Dr.rV/rrV: (3.70) The term rrV, which is called the vector Laplacian and sometimes written r2V, has prior to this point not been defined; Eq. (3.70) (solved for r2V) can be taken to be its definition. In Cartesian coordinates, r2Vis a vector whose icomponent is r2Vi, and that fact can be confirmed either by direct component expansion or by applying the BAC–CAB rule, Eq. (3.18), with care always to place Vso that the differential operators act on it. While Eq. (3.70) is general, r2Vseparates into Laplacians for the components of Vonly in Cartesian coordinates. ArfKen_Ch03-9780123846549.tex 156 Chapter 3 Vector Analysis Example 3.6.2 ELECTROMAGNETIC WAVE EQUATION Even in vacuum, Maxwell’s equations can describe electromagnetic waves. To derive an electromagnetic wave equation, we start by taking the time derivative of Eq. (3.68) for the case JD0, and the curl of Eq. (3.69). We then have @ @trBD00@2E @t2; r.rE/D@ @trBD 00@2E @t2: We now have an equation that involves only E; it can be brought to a more convenient form by applying Eq. (3.70), dropping the first term on the right of that equation because, in vacuum, rED0. The result is the vector electromagnetic wave equation for E, r2ED00@2E @t2D1 c2@2E @t2: (3.71) Equation (3.71) separates into three scalar wave equations, each involving the (scalar) Laplacian. There is a separate equation for each Cartesian component of E.  Miscellaneous Vector Identities Our introduction of differential vector operators is now formally complete, but we present two further examples to illustrate how the relationships between these operators can be manipulated to obtain useful vector identities. Example 3.6.3 DIVERGENCE AND CURL OF A PRODUCT First, simplify r.fV/, where fandVare, respectively, scalar and vector functions. Working with the components, r.fV/D@ @x.f Vx/C@ @y.f Vy/C@ @z.f Vz/ D@f @xVxCf@Vx @xC@f @yVyCf@Vy @yC@f @zVzCf@Vz @z D.rf/VCfrV: (3.72) Now simplify r.fV/. Consider the x-component: @ @y.f Vz/@ @z.f Vy/Df@Vz @y@Vy @z C@f @yVz@f @zVy : This is the x-component of f.rV/C.rf/V, so we have r.fV/Df.rV/C.rf/V: (3.73)  ArfKen_Ch03-9780123846549.tex 3.6 Differential Vector Operators: Further Properties 157 Example 3.6.4 GRADIENT OF A DOT PRODUCT Verify that r.AB/D.Br/AC.Ar/BCB.rA/CA.rB/: (3.74) This problem is easier to solve if we recognize that r.AB/is a type of term that appears in the BAC–CAB expansion of a vector triple product, Eq. (3.18). From that equation, we have A.rB/DrB.AB/.Ar/B; where we placed Bat the end of the final term because rmust act on it. We write rBto indicate an operation our notation is not really equipped to handle. In this term, racts only onB, because Aappeared to its left on the left-hand side of the equation. Interchanging the roles of AandB, we also have B.rA/DrA.AB/.Br/A; whererAacts only on A. Adding these two equations together, noting that rBCrAis simply an unrestricted r, we recover Eq. (3.74).  Exercises 3.6.1 Show that uvis solenoidal if uandvare each irrotational. 3.6.2 IfAis irrotational, show that Aris solenoidal. 3.6.3 A rigid body is rotating with constant angular velocity !. Show that the linear velocity vis solenoidal. 3.6.4 If a vector function V.x;y;z/is not irrotational, show that if there exists a scalar func- tiong.x;y;z/such that gVis irrotational, then VrVD0: 3.6.5 Verify the vector identity r.AB/D.Br/A.Ar/BB.rA/CA.rB/: 3.6.6 As an alternative to the vector identity of Example 3.6.4 show that r.AB/D.Ar/BC.Br/ACA.rB/CB.rA/: 3.6.7 Verify the identity A.rA/D1 2r.A2/.Ar/A: 3.6.8 IfAandBare constant vectors, show that r.ABr/DAB: ArfKen_Ch03-9780123846549.tex 158 Chapter 3 Vector Analysis 3.6.9 Verify Eq. (3.70), r.rV/Dr.rV/rrV; by direct expansion in Cartesian coordinates. 3.6.10 Prove that r.'r'/D0: 3.6.11 You are given that the curl of Fequals the curl of G. Show that FandGmay differ by (a) a constant and (b) a gradient of a scalar function. 3.6.12 The Navier-Stokes equation of hydrodynamics contains a nonlinear term of the form .vr/v. Show that the curl of this term may be written as rTv.rv/U: 3.6.13 Prove that.ru/.rv/is solenoidal, where uandvare differentiable scalar functions. 3.6.14 The function 'is a scalar satisfying Laplace’s equation, r2'D0. Show that r'is both solenoidal and irrotational. 3.6.15 Show that any solution of the equation r.rA/k2AD0 automatically satisfies the vector Helmholtz equation r2ACk2AD0 andthe solenoidal condition rAD0: Hint. Let roperate on the first equation. 3.6.16 The theory of heat conduction leads to an equation r29Dkjr8j2; where8is a potential satisfying Laplace’s equation: r28D0. Show that a solution of this equation is 9Dk82=2. 3.6.17 Given the three matrices MxD0 @0 0 0 0 0i 0i01 A;MyD0 @0 0 i 0 0 0 i0 01 A; and MzD0 @0i0 i 0 0 0 0 01 A; show that the matrix-vector equation  MrC131 c@ @t D0 ArfKen_Ch03-9780123846549.tex 3.7 Vector Integration 159 reproduces Maxwell’s equations in vacuum. Here is a column vector with compo- nents jDBji Ej=c;jDx;y;z. Note that"00D1=c2and that 13is the 33unit matrix. 3.6.18 Using the Pauli matrices iof Eq. (2.28), show that .a/.b/D.ab/12Ci.ab/: Here Oex1COey2COez3; aandbare ordinary vectors, and 12is the 22unit matrix. 3.7 V ECTOR INTEGRATION In physics, vectors occur in line, surface, and volume integrals. At least in principle, these integrals can be decomposed into scalar integrals involving the vector components; there are some useful general observations to make at this time. Line Integrals Possible forms for line integrals include the following: Z C'dr;Z CFdr;Z CVdr: (3.75) In each of these the integral is over some path Cthat may be open (with starting and endpoints distinct) or closed (forming a loop). Inserting the form of dr, the first of these integrals reduces immediately to Z C'drDOexZ C'.x;y;z/dxCOeyZ C'.x;y;z/dyCOezZ C'.x;y;z/dz: (3.76) The unit vectors need not remain within the integral beause they are constant in both mag- nitude and direction. The integrals in Eq. (3.76) are one-dimensional scalar integrals. Note, however, that the integral over xcannot be evaluated unless yandzare known in terms of x; similar observations apply for the integrals over yandz. This means that the path Cmust be specified. Unless 'has special properties, the value of the integral will depend on the path. The other integrals in Eq. (3.75) can be handled similarly. For the second integral, which is of common occurrence, being that which evaluates the work associated with displace- ment on the path C, we have: WDZ CFdrDZ CFx.x;y;z/dxCZ CFy.x;y;z/dyCZ CFz.x;y;z/dz: (3.77) ArfKen_Ch03-9780123846549.tex 160 Chapter 3 Vector Analysis Example 3.7.1 LINE INTEGRALS We consider two integrals in 2-D space: ICDZ C'.x;y/dr;with'.x;y/D1, JCDZ CF.x;y/dr;with F.x;y/DyOexCxOey. We perform integrations in the xy-plane from (0,0) to (1,1) by the two different paths shown in Fig. 3.12: Path C1is.0;0/!.1;0/!.1;1/, Path C2is the straight line .0;0/!.1;1/. For the first segment of C1,xranges from 0 to 1 while yis fixed at zero. For the second segment, yranges from 0 to 1 while xD1. Thus, IC1DOex1Z 0dx'.x;0/COey1Z 0dy'.1; y/DOex1Z 0dxCOey1Z 0dyDOexCOey; JC1D1Z 0dx F x.x;0/C1Z 0dy F y.1;y/D1Z 0D1Z 0dx.0/C1Z 0dy.1/D1: On Path 2, both dxanddyrange from 0 to 1, with xDyat all points of the path. Thus, IC2DOex1Z 0dx'.x;x/COey1Z 0dy'.y;y/DOexCOey; JC2D1Z 0dx F x.x;x/C1Z 0dyF y.y;y/D1Z 0dx.x/C1Z 0dy.y/D1 2C1 2D0: We see that integral Iis independent of the path from (0,0) to (1,1), a nearly trivial special case, while the integral Jis not.  y 1 1xC2 C1C1 FIGURE 3.12 Line integration paths. ArfKen_Ch03-9780123846549.tex 3.7 Vector Integration 161 FIGURE 3.13 Positive normal directions: left, disk; right, spherical surface with hole. Surface Integrals Surface integrals appear in the same forms as line integrals, the element of area being a vector, d, normal to the surface: Z 'd;Z Vd;Z Vd: Often dis writtenOndA, whereOnis a unit vector indicating the normal direction. There are two conventions for choosing the positive direction. First, if the surface is closed (has no boundary), we agree to take the outward normal as positive. Second, for an open surface, the positive normal depends on the direction in which the perimeter of the surface is tra- versed. Starting from an arbitrary point on the perimeter, we define a vector uto be in the direction of travel along the perimeter, and define a second vector vat our perimeter point but tangent to and lying on the surface. We then take uvas the positive normal direction. This corresponds to a right-hand rule, and is illustrated in Fig. 3.13. It is necessary to define the orientation carefully so as to deal with cases such as that of Fig. 3.13, right. The dot-product form is by far the most commonly encountered surface integral, as it corresponds to a flow or flux through the given surface. Example 3.7.2 A SURFACE INTEGRAL Consider a surface integral of the form IDR SBdover the surface of a tetrahe- dron whose vertices are at the origin and at the points (1,0,0), (0,1,0), and (0,0,1), with BD.xC1/OexCyOeyzOez. See Fig. 3.14. The surface consists of four triangles, which can be identified and their contributions evaluated, as follows: 1. On the xy-plane ( zD0), vertices at .x;y/D(0,0), (1,0), and (0,1); direction of out- ward normal isOez, sodDOezd A(d ADelement of area on this triangle). Here, BD.xC1/OexCyOey, and BdD0. So there is no contribution to I. 2. On the xzplane ( yD0), vertices at .x;z/D(0,0), (1,0), and (0,1); direction of out- ward normal isOey, sodDOeyd A. On this triangle, BD.xC1/OexzOez, Again, BdD0. There is no contribution to I. 3. On the yzplane ( xD0), vertices at .y;z/D(0,0), (1,0), and (0,1); direction of outward normal is Oex, so dDOexd A. Here, BDOexCyOeyzOez, and ArfKen_Ch03-9780123846549.tex 162 Chapter 3 Vector Analysis B O1 11 A xCCz yz=1 z=0AB2 2 3 2 2 FIGURE 3.14 Tetrahedron, and detail of the oblique face. BdD.1/d A ; the contribution to Iis1times the area of the triangle ( D1/2), orI3D1=2 . 4. Obliquely oriented, vertices at .x;y;z/D(1,0,0), (0,1,0), (0,0,1); direction of out- ward normal isOnD.OexCOeyCOez/=p 3, and dDOnd A. Using also BD.xC1/OexC yOeyzOez, this contribution to Ibecomes I4DZ 14xC1Cyzp 3d ADZ 142.1z/p 3d A; where we have used the fact that on this triangle, xCyCzD1. To complete the evaluation, we note that the geometry of the triangle is as shown in Fig. 3.14, that the width of the triangle at height zisp 2.1z/, and a change dzin zproduces a displacementp3=2dz on the triangle. I4therefore can be written I4D1Z 02.1z/2dzD2 3: Combining the nonzero contributions I3andI4, we obtain the final result ID1 2C2 3D1 6:  Volume Integrals Volume integrals are somewhat simpler, because the volume element dis a scalar quantity. Sometimes dis written d3r, ord3xwhen the coordinates were designated .x1;x2;x3/. In the literature, the form dris frequently encountered, but in contexts that usually reveal that it is a synonym for d, and not a vector quantity. The volume integrals under consideration here are of the formZ VdDOexZ VxdCOeyZ VydCOezZ Vzd: The integral reduces to a vector sum of scalar integrals. ArfKen_Ch03-9780123846549.tex 3.7 Vector Integration 163 Some volume integrals contain vector quantities in combinations that are actually scalar. Often these can be rearranged by applying techniques such as integration by parts. Example 3.7.3 INTEGRATION BY PARTS Consider an integral over all space of the formR A.r/rf.r/d3rin the frequently occur- ring special case in which either forAvanish sufficiently strongly at infinity. Expanding the integrand into components, Z A.r/rf.r/d3rDx dy dz Axf 1 xD1Z f@Ax @xdx C Dy f@Ax @xdx dy dzy f@Ay @ydx dy dzy f@Az @zdx dy dz DZ f.r/rA.r/d3r: (3.78) For example, if ADeikzOpdescribes a photon with a constant polarization vector in the directionOpand .r/ is a bound-state wave function (so it vanishes at infinity), then Z eikzOpr .r/d3rD.OpOez/Z .r/deikz dzd3rDik.OpOez/Z .r/eikzd3r: Only the z-component of the gradient contributes to the integral. Analogous rearrangements (assuming the integrated terms vanish at infinity) include Z f.r/rA.r/d3rDZ A.r/rf.r/d3r; (3.79) Z C.r/ rA.r d3rDZ A.r/ rC.r/ d3r: (3.80) In the cross-product example, the sign change from the integration by parts combines with the signs from the cross product to give the result shown.  Exercises 3.7.1 The origin and the three vectors A,B, and C(all of which start at the origin) define a tetrahedron. Taking the outward direction as positive, calculate the total vector area of the four tetrahedral surfaces. 3.7.2 Find the workH Fdrdone moving on a unit circle in the xy-plane, doing work against a force field given by FDOexy x2Cy2COeyx x2Cy2V (a) Counterclockwise from 0to, (b) Clockwise from 0to. Note that the work done depends on the path. ArfKen_Ch03-9780123846549.tex 164 Chapter 3 Vector Analysis 3.7.3 Calculate the work you do in going from point .1;1/to point.3;3/. The force you exert is given by FDOex.xy/COey.xCy/: Specify clearly the path you choose. Note that this force field is nonconservative. 3.7.4 EvaluateH rdrfor a closed path of your choosing. 3.7.5 Evaluate 1 3Z srd over the unit cube defined by the point .0;0;0/and the unit intercepts on the positive x-,y-, and z-axes. Note that rdis zero for three of the surfaces and that each of the three remaining surfaces contributes the same amount to the integral. 3.8 I NTEGRAL THEOREMS The formulas in this section relate a volume integration to a surface integral on its boundary (Gauss’ theorem), or relate a surface integral to the line defining its perimeter (Stokes’ theorem). These formulas are important tools in vector analysis, particularly when the functions involved are known to vanish on the boundary surface or perimeter. Gauss’ Theorem Here we derive a useful relation between a surface integral of a vector and the volume integral of the divergence of that vector. Let us assume that a vector Aand its first deriva- tives are continuous over a simply connected region of R3(regions that contain holes, like a donut, are not simply connected). Then Gauss’ theorem states that I @VAdDZ VrAd: (3.81) Here the notations Vand@Vrespectively denote a volume of interest and the closed sur- face that bounds it. The circle on the surface integral is an additional indication that the surface is closed. To prove the theorem, consider the volume Vto be subdivided into an arbitrary large number of tiny (differential) parallelepipeds, and look at the behavior of rAfor each. See Fig. 3.15. For any given parallelepiped, this quantity is a measure of the net outward flow (of whatever Adescribes) through its boundary. If that boundary is interior (i.e., is shared by another parallelepiped), outflow from one parallelepiped is inflow to its neighbor; in a summation of all the outflows, all the contributions of interior boundaries cancel. Thus, the sum of all the outflows in the volume will just be the sum of those through the exterior boundary. In the limit of infinite subdivision, these sums become integrals: The left-hand side of Eq. (3.81) becomes the total outflow to the exterior, while its right-hand side is the sum of the outflows of the differential elements (the parallelepipeds). ArfKen_Ch03-9780123846549.tex 3.8 Integral Theorems 165 FIGURE 3.15 Subdivision for Gauss’ theorem. A simple alternate explanation of Gauss’ theorem is that the volume integral sums the outflows rAfrom all elements of the volume; the surface integral computes the same thing, by directly summing the flow through all elements of the boundary. If the region of interest is the complete R3, and the volume integral converges, the surface integral in Eq. (3.81) must vanish, giving the useful result Z rAdD0;integration over R3and convergent. (3.82) Example 3.8.1 TETRAHEDRON We check Gauss’ theorem for a vector BD.xC1/OexCyOeyzOez, comparing Z VrBdvs.Z @VBd; where Vis the tetrahedron of Example 3.7.2. In that example we computed the surface integral needed here, obtaining the value 1=6. For the integral over V, we take the diver- gence, obtaining rBD1. The volume integral therefore reduces to the volume of the tetrahedron that, with base of area 1/2 and height 1, has volume 1=31=21D1=6. This instance of Gauss’ theorem is confirmed.  Green’s Theorem A frequently useful corollary of Gauss’ theorem is a relation known as Green’s theorem. Ifuandvare two scalar functions, we have the identities r.urv/Dur2vC.ru/.rv/; (3.83) r.urv/Dur2vC.ru/.rv/: (3.84) ArfKen_Ch03-9780123846549.tex 166 Chapter 3 Vector Analysis Subtracting Eq. (3.84) from Eq. (3.83), integrating over a volume Von which u,v, and their derivatives are continuous, and applying Gauss’ theorem, Eq. (3.81), we obtain Z V.ur2vvr2u/dDI @V.urvvru/d: (3.85) This is Green’s theorem. An alternate form of Green’s theorem, obtained from Eq. (3.83) alone, is I @VurvdDZ Vur2vdCZ Vrurvd: (3.86) While the results already obtained are by far the most important forms of Gauss’ theo- rem, volume integrals involving the gradient or the curl may also appear. To derive these, we consider a vector of the form B.x;y;z/DB.x;y;z/a; (3.87) in which ais a vector with constant magnitude and constant but arbitrary direction. Then Eq. (3.81) becomes, applying Eq. (3.72), aI @VB dDZ Vr.Ba/dDaZ VrB d: This may be rewritten a2 4I @VB dZ VrB d3 5D0: (3.88) Since the direction of ais arbitrary, Eq. (3.88) cannot always be satisfied unless the quan- tity in the square brackets evaluates to zero.6The result is I @VB dDZ VrB d: (3.89) In a similar manner, using BDaPin which ais a constant vector, we may show I @VdPDZ VrPd: (3.90) These last two forms of Gauss’ theorem are used in the vector form of Kirchoff diffraction theory. 6This exploitation of the arbitrary nature of a part of a problem is a valuable and widely used technique. ArfKen_Ch03-9780123846549.tex 3.8 Integral Theorems 167 Stokes’ Theorem Stokes’ theorem is the analog of Gauss’ theorem that relates a surface integral of a deriva- tive of a function to the line integral of the function, with the path of integration being the perimeter bounding the surface. Let us take the surface and subdivide it into a network of arbitrarily small rectangles. In Eq. (3.60) we saw that the circulation of a vector Babout such a differential rectan- gles (in the xy-plane) is rBzOezdx dy . Identifying dx dyOezas the element of area d, Eq. (3.60) generalizes to X four sidesBdrDrBd: (3.91) We now sum over all the little rectangles; the surface contributions, from the right-hand side of Eq. (3.91), are added together. The line integrals (left-hand side) of all interior line segments cancel identically. See Fig. 3.16. Only the line integral around the perimeter survives. Taking the limit as the number of rectangles approaches infinity, we have I @SBdrDZ SrBd: (3.92) Here@Sis the perimeter of S. This is Stokes’ theorem. Note that both the sign of the line integral and the direction of ddepend on the direction the perimeter is traversed, so consistent results will always be obtained. For the area and the line-integral direction shown in Fig. 3.16, the direction of for the shaded rectangle will be outof the plane of the paper. Finally, consider what happens if we apply Stokes’ theorem to a closed surface. Since it has no perimeter, the line integral vanishes, so Z SrBdD0; forSa closed surface. (3.93) As with Gauss’ theorem, we can derive additional relations connecting surface integrals with line integrals on their perimeter. Using the arbitrary-vector technique employed to FIGURE 3.16 Direction of normal for the shaded rectangle when perimeter of the surface is traversed as indicated. ArfKen_Ch03-9780123846549.tex 168 Chapter 3 Vector Analysis reach Eqs. (3.89) and (3.90), we can obtain Z Sdr'DI @S'dr; (3.94) Z S.dr/PDI @SdrP: (3.95) Example 3.8.2 OERSTED’S AND FARADAY’S LAWS Consider the magnetic field generated by a long wire that carries a time-independent cur- rent I(meaning that @E=@tD@B=@tD0). The relevant Maxwell equation, Eq. (3.68), then takes the form rBD0J. Integrating this equation over a disk Sperpendicular to and surrounding the wire (see Fig. 3.17), we have IDZ SJdD1 0Z S.rB/d: Now we apply Stokes’ theorem, obtaining the result ID.1= 0/H @SBdr, which is Oersted’s law. Similarly, we can integrate Maxwell’s equation for rE, Eq. (3.69). Imagine moving a closed loop .@S/of wire (of area S) across a magnetic induction field B. We have Z S.rE/dDd dtZ SBdDd8 dt; where8is the magnetic flux through the area S. By Stokes’ theorem, we have Z @SEdrDd8 dt: This is Faraday’s law. The line integral represents the voltage induced in the wire loop; it is equal in magnitude to the rate of change of the magnetic flux through the loop. There is no sign ambiguity; if the direction of @Sis reversed, that causes a reversal of the direction of dand thereby of 8.  BI FIGURE 3.17 Direction of Bgiven by Oersted’s law. ArfKen_Ch03-9780123846549.tex 3.8 Integral Theorems 169 Exercises 3.8.1 Using Gauss’ theorem, prove that I SdD0 ifSD@Vis a closed surface. 3.8.2 Show that 1 3I SrdDV; where Vis the volume enclosed by the closed surface SD@V. Note. This is a generalization of Exercise 3.7.5. 3.8.3 IfBDrA, show that I SBdD0 for any closed surface S. 3.8.4 From Eq. (3.72), with Vthe electric field Eand fthe electrostatic potential ';show that, for integration over all space, Z 'dD"0Z E2d: This corresponds to a 3-D integration by parts. Hint. EDr';rED=" 0:You may assume that 'vanishes at large rat least as fast as r1. 3.8.5 A particular steady-state electric current distribution is localized in space. Choosing a bounding surface far enough out so that the current density Jis zero everywhere on the surface, show that Z JdD0: Hint. Take one component of Jat a time. With rJD0, show that JiDr.xiJ/and apply Gauss’ theorem. 3.8.6 Given a vector tDOexyCOeyx, show, with the help of Stokes’ theorem, that the integral oftaround a continuous closed curve in the xy-plane satisfies 1 2I tdD1 2I .x dyy dx/DA; where Ais the area enclosed by the curve. ArfKen_Ch03-9780123846549.tex 170 Chapter 3 Vector Analysis 3.8.7 The calculation of the magnetic moment of a current loop leads to the line integral I rdr: (a) Integrate around the perimeter of a current loop (in the xy-plane) and show that the scalar magnitude of this line integral is twice the area of the enclosed surface. (b) The perimeter of an ellipse is described by rDOexacosCOeybsin. From part (a) show that the area of the ellipse is ab. 3.8.8 EvaluateH rdrby using the alternate form of Stokes’ theorem given by Eq. (3.95): Z S.dr/PDI dP: Take the loop to be entirely in the xy-plane. 3.8.9 Prove that I urvdDI vrud: 3.8.10 Prove that I urvdDZ S.ru/.rv/d: 3.8.11 Prove that I @VdPDZ VrPd: 3.8.12 Prove that Z Sdr'DI @S'dr: 3.8.13 Prove that Z S.dr/PDI @SdrP: 3.9 P OTENTIAL THEORY Much of physics, particularly electromagnetic theory, can be treated more simply by intro- ducing potentials from which forces can be derived. This section deals with the definition and use of such potentials. ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 171 Scalar Potential If, over a given simply connected region of space (one with no holes), a force can be expressed as the negative gradient of a scalar function ', FDr'; (3.96) we call'ascalar potential, and we benefit from the feature that the force can be described in terms of one function instead of three. Since the force is a derivative of the scalar poten- tial, the potential is only determined up to an additive constant, which can be used to adjust its value at infinity (usually zero) or at some other reference point. We want to know what conditions Fmust satisfy in order for a scalar potential to exist. First, consider the result of computing the work done against a force given by r' when an object subject to the force is moved from a point Ato a point B. This is a line integral of the form BZ AFdrDBZ Ar'dr: (3.97) But, as pointed out in Eq. (3.41), r'drDd', so the integral is in fact independent of the path, depending only on the endpoints AandB. So we have BZ AFdrD'.rB/'.r A/; (3.98) which also means that if AandBare the same point, forming a closed loop, I FdrD0: (3.99) We conclude that a force (on an object) described by a scalar potential is a conservative force, meaning that the work needed to move the object between any two points is inde- pendent of the path taken, and that '.r/ is the work needed to move to the point rfrom a reference point where the potential has been assigned the value zero. Another property of a force given by a scalar potential is that rFDrr'D0 (3.100) as prescribed by Eq. (3.64). This observation is consistent with the notion that the lines of force of a conservative Fcannot form closed loops. The three conditions, Eqs. (3.96), (3.99), and (3.100), are all equivalent. If we take Eq. (3.99) for a differential loop, its left side and that of Eq. (3.100) must, according to Stokes’ theorem, be equal. We already showed both these equations followed from Eq. (3.96). To complete the establishment of full equivalence, we need only to derive Eq. (3.96) from Eq. (3.99). Going backward to Eq. (3.97), we rewrite it as BZ A.FCr'/drD0; ArfKen_Ch03-9780123846549.tex 172 Chapter 3 Vector Analysis which must be satisfied for all AandB. This means its integrand must be identically zero, thereby recovering Eq. (3.96). Example 3.9.1 GRAVITATIONAL POTENTIAL We have previously, in Example 3.5.2, illustrated the generation of a force from a scalar potential. To perform the reverse process, we must integrate. Let us find the scalar potential for the gravitational force FGDGm 1m2Or r2DkOr r2; radially inward. Setting the zero of scalar potential at infinity, we obtain by integrating (radially) from infinity to position r, 'G.r/'G.1/DrZ 1FGdrDC1Z rFGdr: The minus sign in the central member of this equation arises because we are calculating the work done against the gravitational force. Evaluating the integral, 'G.r/D1Z rkdr r2Dk rDGm 1m2 r: The final negative sign corresponds to the fact that gravity is an attractive force.  Vector Potential In some branches of physics, especially electrodynamics, it is convenient to introduce a vector potential A such that a (force) field Bis given by BDrA: (3.101) An obvious reason for introducing Ais that it causes Bto be solenoidal; if Bis the mag- netic induction field, this property is required by Maxwell’s equations. Here we want to develop a converse, namely to show that when Bis solenoidal, a vector potential Aexists. We demonstrate the existence of Aby actually writing it. Our construction is ADOeyxZ x0Bz.x;y;z/dxCOez2 4yZ y0Bx.x0;y;z/dyxZ x0By.x;y;z/dx3 5: (3.102) ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 173 Checking the y- and z-components of rAfirst, noting that AxD0, .rA/yD@Az @xDC@ @xxZ x0By.x;y;z/dxDBy; .rA/zDC@Ay @xDD@ @xxZ x0Bz.x;y;z/dxDBz: Thex-component of rAis a bit more complicated. We have .rA/xD@Az @y@Ay @z D@ @y2 4yZ y0Bx.x0;y;z/dyxZ x0By.x;y;z/dx3 5@ @zxZ x0Bz.x;y;z/dx DBx.x0;y;z/xZ x0@By.x;y;z/ @[email protected];y;z/ @z dx: To go further, we must use the fact that Bis solenoidal, which means rBD0. We can therefore make the replacement @By.x;y;z/ @[email protected];y;z/ @zD@Bx.x;y;z/ @x; after which the xintegration becomes trivial, yielding CZx [email protected];y;z/ @xdxDBx.x;y;z/Bx.x0;y;z/; leading to the desired final result .rA/xDBx. While we have shown that there exists a vector potential Asuch that rADBsubject only to the condition that Bbe solenoidal, we have in no way established that Ais unique. In fact, Ais far from unique, as we can add to it not only an arbitrary constant, but also the gradient of anyscalar function, r', without affecting Bat all. Moreover, our verification ofAwas independent of the values of x0andy0, so these can be assigned arbitrarily without affecting B. In addition, we can derive another formula for Ain which the roles of xandyare interchanged: ADOexyZ y0Bz.x;y;z/dyOez2 4xZ x0By.x;y0;z/dxyZ y0Bx.x;y;z/dy3 5: (3.103) ArfKen_Ch03-9780123846549.tex 174 Chapter 3 Vector Analysis Example 3.9.2 MAGNETIC VECTOR POTENTIAL We consider the construction of the vector potential for a constant magnetic induction field BDBzOez: (3.104) Using Eq. (3.102), we have (choosing the arbitrary value of x0to be zero) ADOeyxZ 0BzdxDOeyx Bz: (3.105) Alternatively, we could use Eq. (3.103) for A, leading to A0DOexyBz: (3.106) Neither of these is the form for Afound in many elementary texts, which for Bfrom Eq. (3.104) is A00D1 2.Br/DBz 2.xOeyyOex/: (3.107) These disparate forms can be reconciled if we use the freedom to add to Aany expression of the form r'. Taking'DCxy, the quantity that can be added to Awill be of the form r'DC.yOexCxOey/: We now see that ABz 2.yOexCxOey/DA0CBz 2.yOexCxOey/DA00; showing that all these formulas predict the same value of B.  Example 3.9.3 POTENTIALS IN ELECTROMAGNETISM If we introduce suitably defined scalar and vector potentials 'andAinto Maxwell’s equations, we can obtain equations giving these potentials in terms of the sources of the electromagnetic field (charges and currents). We start with BDrA, thereby assuring satisfaction of the Maxwell’s equation rBD0. Substitution into the equation for rE yields rEDr@A @t! r EC@A @t D0; showing that EC@A=@tis a gradient and can be written as r', thereby defining '. This preserves the notion of an electrostatic potential in the absence of time dependence, and means that Aand'have now been defined to give BDrA; EDr'@A @t: (3.108) At this point Ais still arbitrary to the extent of adding any gradient, which is equivalent to making an arbitrary choice of rA. A convenient choice is to require 1 c2@' @tCrAD0: (3.109) ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 175 This gauge condition is called the Lorentz gauge, and transformations of Aand'to satisfy it or any other legitimate gauge condition are called gauge transformations . The invariance of electromagnetic theory under gauge transformation is an important precursor of contemporary directions in fundamental physical theory. From Maxwell’s equation for rEand the Lorentz gauge condition, we get  "0DrEDr2E@ @trADr2'C1 c2@2' @t2; (3.110) showing that the Lorentz gauge permitted us to decouple Aand'to the extent that we have an equation for 'in terms only of the charge density ; neither Anor the current density Jenters this equation. Finally, from the equation for rB, we obtain 1 c2@2A @t2r2AD0J: (3.111) Proof of this formula is the subject of Exercise 3.9.11.  Gauss’ Law Consider a point charge qat the origin of our coordinate system. It produces an electric field E, given by EDqOr 4" 0r2: (3.112) Gauss’ law states that for an arbitrary volume V, I @VEdD(q "0if@Vencloses q, 0 if@Vdoes not enclose q.(3.113) The case that @Vdoes not enclose qis easily handled. From Eq. (3.54), the r2central force Eis divergenceless everywhere except at rD0, and for this case, throughout the entire volume V. Thus, we have, invoking Gauss’ theorem, Eq. (3.81), Z VrED0! EdD0: Ifqis within the volume V, we must be more devious. We surround rD0by a small spherical hole (of radius ), with a surface we designate S0, and connect the hole with the boundary of Vvia a small tube, thereby creating a simply connected region V0to which Gauss’ theorem will apply. See Fig. 3.18. We now considerH Edon the surface of this modified volume. The contribution from the connecting tube will become negligible in the limit that it shrinks toward zero cross section, as Eis finite everywhere on the tube’s surface. The integral over the modified @Vwill thus be that of the original @V(over the outer boundary, which we designate S), plus that of the inner spherical surface ( S0). ArfKen_Ch03-9780123846549.tex 176 Chapter 3 Vector Analysis FIGURE 3.18 Making a multiply connected region simply connected. But note that the “outward” direction for S0is toward smaller r, sod0DOrd A. Because the modified volume contains no charge, we have I @V0EdDI SEdCq 4" 0I S0Ord0 2D0; (3.114) where we have inserted the explicit form of Ein the S0integral. Because S0is a sphere of radius, this integral can be evaluated. Writing das the element of solid angle, so d AD 2d, I S0Ord0 2DZOr 2.Or2d/DZ dD4; independent of the value of . Returning now to Eq. (3.114), it can be rearranged into I SEdDq 4" 0.4/DCq "0; the result needed to confirm the second case of Gauss’ law, Eq. (3.113). Because the equations of electrostatics are linear, Gauss’ law can be extended to collec- tions of charges, or even to continuous charge distributions. In that case, qcan be replaced byR Vd, and Gauss’ law becomes Z @VEdDZ V "0d: (3.115) If we apply Gauss’ theorem to the left side of Eq. (3.115), we have Z VrEdDZ V "0d: Since our volume is completely arbitrary, the integrands of this equation must be equal, so rED "0: (3.116) We thus see that Gauss’ law is the integral form of one of Maxwell’s equations. ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 177 Poisson’s Equation If we return to Eq. (3.116) and, assuming a situation independent of time, write EDr', we obtain r2'D "0: (3.117) This equation, applicable to electrostatics,7is called Poisson’s equation. If, in addition, D0, we have an even more famous equation, r2'D0; (3.118) Laplace’s equation. To make Poisson’s equation apply to a point charge q, we need to replace by a con- centration of charge that is localized at a point and adds up to q. The Dirac delta function is what we need for this purpose. Thus, for a point charge qat the origin, we write r2'Dq "0.r/; (charge qatrD0). (3.119) If we rewrite this equation, inserting the point-charge potential for ', we have q 4" 0r21 r Dq "0.r/; which reduces to r21 r D4.r/: (3.120) This equation circumvents the problem that the derivatives of 1=rdo not exist at rD0, and gives appropriate and correct results for systems containing point charges. Like the definition of the delta function itself, Eq. (3.120) is only meaningful when inserted into an integral. It is an important result that is used repeatedly in physics, often in the form r2 11 r12 D4.r 1r2/: (3.121) Here r12Djr 1r2j, and the subscript in r1indicates that the derivatives apply to r1. Helmholtz’s Theorem We now turn to two theorems that are of great formal importance, in that they establish conditions for the existence and uniqueness of solutions to time-independent problems in electromagnetic theory. The first of these theorems is: A vector field is uniquely specified by giving its divergence and its curl within a simply connected region and its normal component on the boundary. 7For general time dependence, see Eq. (3.110). ArfKen_Ch03-9780123846549.tex 178 Chapter 3 Vector Analysis Note that both for this theorem and the next (Helmholtz’s theorem), even if there are points in the simply connected region where the divergence or the curl is only defined in terms of delta functions, these points are not to be removed from the region. LetPbe a vector field satisfying the conditions rPDs;rPDc; (3.122) where smay be interpreted as a given source (charge) density and cas a given circulation (current) density. Assuming that the normal component Pnon the boundary is also given, we want to show that Pis unique. We proceed by assuming the existence of a second vector, P0, which satisfies Eq. (3.122) and has the same value of Pn. We form QDPP0, which must have rQ,rQ, and Qnall identically zero. Because Qis irrotational, there must exist a potential 'such that QDr', and because rQD0, we also have r2'D0: Now we draw on Green’s theorem in the form given in Eq. (3.86), letting uandveach equal'. Because QnD0on the boundary, Green’s theorem reduces to Z V.r'/.r'/dDZ VQQdD0: This equation can only be satisfied if Qis identically zero, showing that P0DP, thereby proving the theorem. The second theorem we shall prove, Helmholtz’s theorem, is A vector Pwith both source and circulation densities vanishing at infinity may be writ- ten as the sum of two parts, one of which is irrotational, the other of which is solenoidal. Helmholtz’s theorem will clearly be satisfied if Pcan be written in the form PDr'CrA; (3.123) sincer'is irrotational, while rAis solenoidal. Because Pis known, so are also s andc, defined as sDrP; cDrP: We proceed by exhibiting expressions for 'andAthat enable the recovery of sandc. Because the region here under study is simply connected and the vector involved vanishes at infinity (so that the first theorem of this subsection applies), having the correct sandc guarantees that we have properly reproduced P. The formulas proposed for 'andAare the following, written in terms of the spatial variable r1: '.r1/D1 4Zs.r2/ r12d2; (3.124) A.r 1/D1 4Zc.r2/ r12d2: (3.125) Here r12Djr 1r2j. ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 179 If Eq. (3.123) is to be satisfied with the proposed values of 'andA, it is necessary that rPDrr'Cr.rA/Dr2'Ds; rPDrr'Cr.rA/Dr.rA/Dc: To check thatr2'Ds, we examine r2 1'.r1/D1 4Z r2 11 r12 s.r2/d2 D1 4Z 4.r 1r2/ s.r2/d2Ds.r1/: (3.126) We have written r1to make clear that it operates on r1and not r2, and we have used the delta-function property given in Eq. (3.121). So shas been recovered. We now check that r.rA/Dc. We start by using Eq. (3.70) to convert this condition to a more easily utilized form: r.rA/Dr.rA/r2ADc: Taking r1as the free variable, we look first at r1 r1A.r 1/ D1 4r1Z r1c.r2/ r12 d2 D1 4r1Z c.r2/r11 r12 d2 D1 4r1Z c.r2/ r21 r12 d2: To reach the second line of this equation, we used Eq. (3.72) for the special case that the vector in that equation is not a function of the variable being differentiated. Then, to obtain the third line, we note that because the r1within the integral acts on a function of r1r2, we can change r1intor2and introduce a sign change. Now we integrate by parts, as in Example 3.7.3, reaching r1 r1A.r 1/ D1 4r1Z r2c.r2/1 r12 d2: At last we have the result we need: r2c.r2/vanishes, because cis a curl, so the entire r.rA/term is zero and may be dropped. This reduces the condition we are checking to r2ADc. The quantityr2Ais a vector Laplacian and we may individually evaluate its Cartesian components. For component j, r2 1Aj.r1/D1 4Z cj.r2/r2 11 r12 d2 D1 4Z cj.r2/ 4.r 1r2/ d2Dcj.r1/: This completes the proof of Helmholtz’s theorem. Helmholtz’s theorem legitimizes the division of the quantities appearing in electromag- netic theory into an irrotational vector field Eand a solenoidal vector field B, together ArfKen_Ch03-9780123846549.tex 180 Chapter 3 Vector Analysis with their respective representations using scalar and vector potentials. As we have seen in numerous examples, the source sis identified as the charge density (divided by "0) and the circulation cis the current density (multiplied by 0). Exercises 3.9.1 If a force Fis given by FD.x2Cy2Cz2/n.OexxCOeyyCOezz/; find (a)rF: (b)rF. (c) A scalar potential '.x;y;z/so that FDr'. (d) For what value of the exponent ndoes the scalar potential diverge at both the origin and infinity? ANS. (a).2nC3/r2n(b) 0 (c)r2nC2=.2nC2/;n6D1 (d) nD1; 'Dlnr: 3.9.2 A sphere of radius ais uniformly charged (throughout its volume). Construct the elec- trostatic potential '.r/for0r<1. 3.9.3 The origin of the Cartesian coordinates is at the Earth’s center. The moon is on the z-axis, a fixed distance Raway (center-to-center distance). The tidal force exerted by the moon on a particle at the Earth’s surface (point x;y;z) is given by FxDG Mmx R3;FyDG Mmy R3;FzDC2 G Mmz R3: Find the potential that yields this tidal force. ANS.G Mm R3 z21 2x21 2y2 . 3.9.4 A long, straight wire carrying a current Iproduces a magnetic induction Bwith com- ponents BD0I 2 y x2Cy2;x x2Cy2;0 : Find a magnetic vector potential A. ANS. ADOz.0I=4/ ln.x2Cy2/:(This solution is not unique.) 3.9.5 If BDOr r2Dx r3;y r3;z r3 ; find a vector Asuch that rADB. ANS. One possible solution is ADOexyz r.x2Cy2/Oeyxz r.x2Cy2/: ArfKen_Ch03-9780123846549.tex 3.9 Potential Theory 181 3.9.6 Show that the pair of equations AD1 2.Br/; BDrA; is satisfied by any constant magnetic induction B. 3.9.7 Vector Bis formed by the product of two gradients BD.ru/.rv/; where uandvare scalar functions. (a) Show that Bis solenoidal. (b) Show that AD1 2.urvvru/ is a vector potential for B;in that BDrA: 3.9.8 The magnetic induction Bis related to the magnetic vector potential AbyBDrA: By Stokes’ theorem Z BdDI Adr: Show that each side of this equation is invariant under the gauge transformation, A! ACr'. Note. Take the function 'to be single-valued. 3.9.9 Show that the value of the electrostatic potential 'at any point Pis equal to the average of the potential over any spherical surface centered on P, provided that there are no electric charges on or within the sphere. Hint. Use Green’s theorem, Eq. (3.85), with uDr1, the distance from P, andvD'. Equation (3.120) will also be useful. 3.9.10 Using Maxwell’s equations, show that for a system (steady current) the magnetic vector potential Asatisfies a vector Poisson equation, r2ADJ; provided we require rAD0. 3.9.11 Derive, assuming the Lorentz gauge, Eq. (3.109): 1 c2@2A @t2r2AD0J: Hint. Eq. (3.70) will be helpful. ArfKen_Ch03-9780123846549.tex 182 Chapter 3 Vector Analysis 3.9.12 Prove that an arbitrary solenoidal vector Bcan be described as BDrA, with ADOexyZ y0Bz.x;y;z/dyOez2 4xZ x0By.x;y0;z/dxyZ y0Bx.x;y;z/dy3 5: 3.10 C URVILINEAR COORDINATES Up to this point we have treated vectors essentially entirely in Cartesian coordinates; when ror a function of it was encountered, we wrote rasp x2Cy2Cz2, so that Cartesian coordinates could continue to be used. Such an approach ignores the simplifications that can result if one uses a coordinate system that is appropriate to the symmetry of a problem. Central force problems are frequently easiest to deal with in spherical polar coordinates. Problems involving geometrical elements such as straight wires may be best handled in cylindrical coordinates. Yet other coordinate systems (of use too infrequent to be described here) may be appropriate for other problems. Naturally, there is a price that must be paid for the use of a non-Cartesian coordinate sys- tem. Vector operators become different in form, and their specific forms may be position- dependent. We proceed here to examine these questions and derive the necessary formulas. Orthogonal Coordinates in R3 In Cartesian coordinates the point .x0;y0;z0/can be identified as the intersection of three planes: (1) the plane xDx0(a surface of constant x), (2) the plane yDy0(constant y), and (3) the plane zDz0(constant z). A change in xcorresponds to a displacement normal to the surface of constant x; similar remarks apply to changes in yorz. The planes of constant coordinate value are mutually perpendicular, and have the obvious feature that the normal to any given one of them is in the same direction, no matter where on the plane it is constructed (a plane of constant xhas a normal that is, of course, everywhere in the direction of Oex). Consider now, as an example of a curvilinear coordinate system, spherical polar coor- dinates (see Fig. 3.19). A point ris identified by r(distance from the origin), (angle of rrelative to the polar axis, which is conventionally in the zdirection), and '(dihedral angle between the zxplane and the plane containing Oezandr). The point ris therefore at the intersection of (1) a sphere of radius r, (2) a cone of opening angle , and (3) a half- plane through equatorial angle '. This example provides several observations: (1) general θ ϕyrz x FIGURE 3.19 Spherical polar coordinates. ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 183 θkeθˆ r r/prime FIGURE 3.20 Effect of a “large” displacement in the direction Oe. Note that r06Dr. coordinates need not be lengths, (2) a surface of constant coordinate value may have a normal whose direction depends on position, (3) surfaces with different constant values of the same coordinate need not be parallel, and therefore also (4) changes in the value of a coordinate may move rin both an amount and a direction that depends on position. It is convenient to define unit vectors Oer,Oe,Oe'in the directions of the normals to the surfaces, respectively, of constant r,, and'. The spherical polar coordinate system has the feature that these unit vectors are mutually perpendicular, meaning that, for example, Oe will be tangent to both the constant- rand constant- 'surfaces, so that a small displacement in theOedirection will not change the values of either the ror the'coordinate. The reason for the restriction to “small” displacements is that the directions of the normals are position-dependent; a “large” displacement in the Oedirection would change r(see Fig. 3.20). If the coordinate unit vectors are mutually perpendicular, the coordinate system is said to be orthogonal. If we have a vector field V(so we associate a value of Vwith each point in a region of R3), we can write V.r/ in terms of the orthogonal set of unit vectors that are defined for the point r; symbolically, the result is V.r/DVrOerCVOeCV'Oe': It is important to realize that the unit vectors Oeihave directions that depend on the value ofr. If we have another vector field W.r/ for the same point r, we can perform algebraic processes8onVandWby the same rules as for Cartesian coordinates. For example, at the point r, VWDVrWrCVWCV'W': However, if VandWare not associated with the same r, we cannot carry out such opera- tions in this way, and it is important to realize that r6DrOerCOeC'Oe': Summarizing, the component formulas for VorWdescribe component decompositions applicable to the point at which the vector is specified; an attempt to decompose ras illustrated above is incorrect because it uses fixed unit-vector orientations where they do not apply. Dealing for the moment with an arbitrary curvilinear system, with coordinates labeled .q1;q2;q3/, we consider how changes in the qiare related to changes in the Cartesian coordinates. Since xcan be thought of as a function of the qi, namely x.q1;q2;q3/, we have dxD@x @q1dq1C@x @q2dq2C@x @q3dq3; (3.127) with similar formulas for dyanddz. 8Addition, multiplication by a scalar, dot and cross products (but not application of differential or integral operators). ArfKen_Ch03-9780123846549.tex 184 Chapter 3 Vector Analysis We next form a measure of the differential displacement, dr, associated with changes dqi. We actually examine .dr/2D.dx/2C.dy/2C.dz/2: Taking the square of Eq. (3.127), we get .dx/2DX i j@x @qi@x @qjdqidqj and similar expressions for .dy/2and.dz/2. Combining these and collecting terms with the same dqidqj, we reach the result .dr/2DX i jgi jdqidqj; (3.128) where gi j.q1;q2;q3/D@x @qi@x @qjC@y @qi@y @qjC@z @qi@z @qj: (3.129) Spaces with a measure of distance given by Eq. (3.128) are called metric orRiemannian. Equation (3.129) can be interpreted as the dot product of a vector in the dqidirection, of components.@x=@qi; @y=@qi; @z=@qi/, with a similar vector in the dqjdirection. If the qicoordinates are perpendicular, the coefficients gi jwill vanish when i6Dj. Since it is our objective to discuss orthogonal coordinate systems, we specialize Eqs. (3.128) and(3.129) to .dr/2D.h1dq1/2C.h2dq2/2C.h3dq3/2; (3.130) h2 iD@x @qi2 C@y @qi2 C@y @qi2 : (3.131) If we consider Eq. (3.130) for a case dq2Ddq3D0, we see that we can identify h1dq1 asdr1, meaning that the element of displacement in the q1direction is h1dq1. Thus, in general, driDhidqi;or@r @qiDhiOei: (3.132) HereOeiis a unit vector in the qidirection, and the overall drtakes the form drDh1dq1Oe1Ch2dq2Oe2Ch3dq3Oe3: (3.133) Note that himay be position-dependent and must have the dimension needed to cause hidqito be a length. Integrals in Curvilinear Coordinates Given the scale factors hifor a set of coordinates, either because they have been tabulated or because we have evaluated them via Eq. (3.131), we can use them to set up formulas for integration in the curvilinear coordinates. Line integrals will take the form Z CVdrDX iZ CVihidqi: (3.134) ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 185 Surface integrals take the same form as in Cartesian coordinates, with the exception that instead of expressions like dx dy we have.h1dq1/.h2dq2/Dh1h2dq1dq2etc. This means that Z SVdDZ SV1h2h3dq2dq3CZ SV2h3h1dq3dq1CZ SV3h1h2dq1dq2: (3.135) The element of volume in orthogonal curvilinear coordinates is dDh1h2h3dq1dq2dq3; (3.136) so volume integrals take the form Z V'.q1;q2;q3/h1h2h3dq1dq2dq3; (3.137) or the analogous expression with 'replaced by a vector V.q1;q2;q3/. Differential Operators in Curvilinear Coordinates We continue with a restriction to orthogonal coordinate systems. Gradient—Because our curvilinear coordinates are orthogonal, the gradient takes the same form as for Cartesian coordinates, providing we use the differential displacements driDhidqiin the formula. Thus, we have r'.q1;q2;q3/DOe11 h1@' @q1COe21 h2@' @q2COe31 h3@' @q3; (3.138) this corresponds to writing ras rDOe11 h1@ @q1COe21 h2@ @q2COe31 h3@ @q3: (3.139) Divergence—This operator must have the same meaning as in Cartesian coordinates, sorVmust give the net outward flux of Vper unit volume at the point of evaluation. The key difference from the Cartesian case is that an element of volume will no longer be a parallelepiped, as the scale factors hiare in general functions of position. See Fig. 3.21. To compute the net outflow of Vin the q1direction from a volume element defined by +B1h2dq2h3dq3|q1+dq1/2h2dq2 −B1h2dq2h3dq3|q1−dq1/2 h3dq3 FIGURE 3.21 Outflow of B1in the q1direction from a curvilinear volume element. ArfKen_Ch03-9780123846549.tex 186 Chapter 3 Vector Analysis dq1,dq2,dq3and centered at .q1;q2;q3/, we must form Netq1outflowDV1h2h3dq2dq3 q1dq 1=2;q2;q3CV1h2h3dq2dq3 q1Cdq 1=2;q2;q3: (3.140) Note that not only V1, but also h2h3must be evaluated at the displaced values of q1; this product may have different values at q1Cdq1=2andq1dq1=2. Rewriting Eq. (3.140) in terms of a derivative with respect to q1, we have Netq1outflowD@ @q1.V1h2h3/dq1dq2dq3: Combining this with the q2andq3outflows and dividing by the differential volume h1h2h3dq1dq2dq3, we get the formula rV.q1;q2;q3/D1 h1h2h3@ @q1.V1h2h3/C@ @q2.V2h3h1/C@ @q3.V3h1h2/ :(3.141) Laplacian—From the formulas for the gradient and divergence, we can form the Laplacian in curvilinear coordinates: r2'.q1;q2;q3/Drr'D 1 h1h2h3@ @q1h2h3 h1@' @q1 C@ @q2h3h1 h2@' @q2 C@ @q3h1h2 h3@' @q3 : (3.142) Note that the Laplacian contains no cross derivatives, such as @2=@q1@q2. They do not appear because the coordinate system is orthogonal. Curl—In the same spirit as our treatment of the divergence, we calculate the circulation around an element of area in the q1q2plane, and therefore associated with a vector in theq3direction. Referring to Fig. 3.22, the line integralH Bdrconsists of four segment −B1h1dq1|q2+dq 2/2 B2h2dq2|q1+dq 1/2 −B2h2dq2|q1−dq 1/2 +B1h1dq1|q2−dq 2/2 FIGURE 3.22 CirculationH Bdraround curvilinear element of area on a surface of constant q3. ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 187 contributions, which to first order are Segment 1D.h1B1/ q1;q2dq 2=2;q3dq1; Segment 2D.h2B2/ q1Cdq 1=2;q2;q3dq2; Segment 3D. h1B1/ q1;q2Cdq 2=2;q3dq1; Segment 4D. h2B2/ q1dq 1=2;q2;q3dq2: Keeping in mind that the hiare functions of position, and that the loop has area h1h2dq1dq2, these contributions combine into a circulation per unit area .rB/3D1 h1h2 @ @q2.h1B1/C@ @q1.h2B2/ : The generalization of this result to arbitrary orientation of the circulation loop can be brought to the determinantal form rBD1 h1h2h3 Oe1h1Oe2h2Oe3h3 @ @q1@ @q2@ @q3 h1B1h2B2h3B3 : (3.143) Just as for Cartesian coordinates, this determinant is to be evaluated from the top down, so that the derivatives will act on its bottom row. Circular Cylindrical Coordinates Although there are at least 11 coordinate systems that are appropriate for use in solving physics problems, the evolution of computers and efficient programming techniques have greatly reduced the need for most of these coordinate systems, with the result that the dis- cussion in this book is limited to (1) Cartesian coordinates, (2) spherical polar coordinates (treated in the next subsection), and (3) circular cylindrical coordinates, which we discuss here. Specifications and details of other coordinate systems will be found in the first two editions of this work and in Additional Readings at the end of this chapter (Morse and Feshbach, Margenau and Murphy). In the circular cylindrical coordinate system the three curvilinear coordinates are labeled .;'; z/. We usefor the perpendicular distance from the z-axis because we reserve rfor the distance from the origin. The ranges of ,', and zare 0<1; 0'<2;1<z<1: ForD0,'is not well defined. The coordinate surfaces, shown in Fig. 3.23, follow: 1. Right circular cylinders having the z-axis as a common axis, D x2Cy21=2 Dconstant. ArfKen_Ch03-9780123846549.tex 188 Chapter 3 Vector Analysis ρϕy xz FIGURE 3.23 Cylindrical coordinates ,',z. 2. Half-planes through the z-axis, at an angle 'measured from the xdirection, 'Dtan1y x Dconstant. The arctangent is double valued on the range of ', and the correct value of 'must be determined by the individual signs of xandy. 3. Planes parallel to the xy-plane, as in the Cartesian system, zDconstant. Inverting the preceding equations, we can obtain xDcos'; yDsin'; zDz: (3.144) This is essentially a 2-D curvilinear system with a Cartesian z-axis added on to form a 3-D system. The coordinate vector rand a general vector Vare expressed as rDOeCzOez;VDVOeCV'Oe'CVzOez: From Eq. (3.131), the scale factors for these coordinates are hD1; h'D; hzD1; (3.145) ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 189 so the elements of displacement, area, and volume are drDOedCOe'd'COezdz; dDOed'dzCOe'ddzCOezdd'; (3.146) dDdd'dz: It is perhaps worth emphasizing that the unit vectors OeandOe'have directions that vary with'; if expressions containing these unit vectors are differentiated with respect to ', the derivatives of these unit vectors must be included in the computations. Example 3.10.1 KEPLER’S AREA LAW FOR PLANETARY MOTION One of Kepler’s laws states that the radius vector of a planet, relative to an origin at the sun, sweeps out equal areas in equal time. It is instructive to derive this relationship using cylindrical coordinates. For simplicity we consider a planet of unit mass and motion in the plane zD0. The gravitational force Fis of the form f.r/Oer, and hence the torque about the origin, rF, vanishes, so angular momentum LDrdr=dt is conserved. To evaluate dr=dt, we start from dras given in Eq. (3.146), writing dr dtDOePCOe'P'; where we have used the dot notation (invented by Newton) to indicate time derivatives. We now form LDOeOePCOe'P' D2P'Oez: We conclude that 2P'is constant. Making the identification 2P'D2d A=dt, where Ais the area swept out, we confirm Kepler’s law.  Continuing now to the vector differential operators, using Eqs. (3.138), (3.141), (3.142), and(3.143), we have r .;'; z/DOe@ @bCOe'1 @ @'COez@ @z; (3.147) rVD1 @ @.V/C1 @V' @'C@Vz @z; (3.148) r2 D1 @ @ @ @ C1 2@2 @'2C@2 @z2; (3.149) rVD1  OeOe'Oez @ @@ @'@ @z VV'Vz : (3.150) ArfKen_Ch03-9780123846549.tex 190 Chapter 3 Vector Analysis Finally, for problems such as circular wave guides and cylindrical cavity resonators, one needs the vector Laplacian r2V. From Eq. (3.70), its components in cylindrical coordi- nates can be shown to be r2V Dr2V1 2V2 2@V' @'; r2V 'Dr2V'1 2V'C2 2@V @'; (3.151) r2V zDr2Vz: Example 3.10.2 A NAVIER-STOKES TERM The Navier-Stokes equations of hydrodynamics contain a nonlinear term r v.rv/ ; where vis the fluid velocity. For fluid flowing through a cylindrical pipe in the zdirection, vDOezv./: From Eq. (3.150), rvD1  OeOe'Oez @ @@ @'@ @z 0 0v./ DOe'@v @; v.rv/D OeOe'Oez 0 0v 0@v @0 DOev./@v @: Finally, r v.rv/ D1  OeOe'Oez @ @@ @'@ @z v@v @0 0 D0: For this particular case, the nonlinear term vanishes.  Spherical Polar Coordinates Spherical polar coordinates were introduced as an initial example of a curvilinear coordi- nate system, and were illustrated in Fig. 3.19. We reiterate: The coordinates are labeled .r;;'/ . Their ranges are 0r<1; 0; 0'<2: ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 191 ForrD0, neithernor'is well defined. Additionally, 'is ill-defined for D0and D. The coordinate surfaces follow: 1. Concentric spheres centered at the origin, rD x2Cy2Cz21=2 Dconstant. 2. Right circular cones centered on the z(polar) axis with vertices at the origin, Darccosz rDconstant. 3. Half-planes through the z(polar) axis, at an angle 'measured from the xdirection, 'Darctany xDconstant. The arctangent is double valued on the range of ', and the correct value of 'must be determined by the individual signs of xandy. Inverting the preceding equations, we can obtain xDrsincos'; yDrsinsin'; zDrcos: (3.152) The coordinate vector rand a general vector Vare expressed as rDrOer;VDVrOerCVOeCV'Oe': From Eq. (3.131), the scale factors for these coordinates are hrD1; hDr;h'Drsin; (3.153) so the elements of displacement, area, and volume are drDOerdrCrOedCrsinOe'd'; dDr2sinOerdd'CrsinOedr d'CrOe'dr d; (3.154) dDr2sinddd': Frequently one encounters a need to perform a surface integration over the angles, in which case the angular dependence of dreduces to dDsindd'; (3.155) where dis called an element of solid angle, and has the property that its integral over all angles has the value Z dD4: Note that for spherical polar coordinates, all three of the unit vectors have directions that depend on position, and this fact must be taken into account when expressions containing the unit vectors are differentiated. ArfKen_Ch03-9780123846549.tex 192 Chapter 3 Vector Analysis The vector differential operators may now be evaluated, using Eqs. (3.138), (3.141), (3.142), and (3.143): r .r;;'/DOer@ @rCOe1 r@ @COe'1 rsin@ @'; (3.156) rVD1 r2sin sin@ @r.r2Vr/Cr@ @.sinV/Cr@V' @' ; (3.157) r2 D1 r2sin sin@ @r r2@ @r C@ @ sin@ @ C1 sin@2 @'2 ;(3.158) rVD1 r2sin Oer rOersinOe' @ @r@ @@ @' Vr rVrsinV' : (3.159) Finally, again using Eq. (3.70), the components of the vector Laplacian r2Vin spherical polar coordinates can be shown to be r2V rDr2Vr2 r2Vr2 r2cotV2 r2@V @2 r2sin@V' @'; r2V Dr2V1 r2sin2VC2 r2@Vr @2 cos r2sin2@V' @'; (3.160) r2V 'Dr2V'1 r2sin2V'C2 r2sin@Vr @'C2 cos r2sin2@V' @': Example 3.10.3 r,r,rFOR A CENTRAL FORCE We can now easily derive some of the results previously obtained more laboriously in Cartesian coordinates: From Eq. (3.156), rf.r/DOerd f dr;rrnDOernrn1: (3.161) Specializing to the Coulomb potential of a point charge at the origin, VDZe=.4" 0r/, so the electric field has the expected value EDr VD.Ze=4" 0r2/Oer. Taking next the divergence of a radial function, we have from Eq. (3.157), rOerf.r/ D2 rf.r/Cd f dr;r.Oerrn/D.nC2/rn1: (3.162) Specializing the above to the Coulomb force ( nD2 ), we have (except for rD0) rr2D0, which is consistent with Gauss’ law. Continuing now to the Laplacian, from Eq. (3.158) we have r2f.r/D2 rd f drCd2f dr2;r2rnDn.nC1/rn2; (3.163) in contrast to the ordinary second derivative of rninvolving n1. ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 193 Finally, from Eq. (3.159), rOerf.r/ D0; (3.164) which confirms that central forces are irrotational.  Example 3.10.4 MAGNETIC VECTOR POTENTIAL A single current loop in the xy-plane has a vector potential Athat is a function only of r and, is entirely in theOe'direction and is related to the current density Jby the equation 0JDrBDr rOe'A'.r;/ : In spherical polar coordinates this reduces to 0JDr1 r2sin Oer rOe rsinOe' @ @r@ @@ @' 0 0 rsinA' Dr1 r2sin Oer@ @.rsinA'/rOe@ @r.rsinA'/ : Taking the curl a second time, we obtain 0JD1 r2sin Oer rOe rsinOe' @ @r@ @@ @' 1 rsin@ @.sinA'/1 r@ @r.r A'/ 0 : Expanding this determinant from the top down, we reach 0JDOe'@2A' @r2C2 r@A' @rC1 r2sin@ @ sin@A' @ 1 r2sin2A' : (3.165) Note that we get, in addition to r2A', one more term:A'=r2sin2.  Example 3.10.5 STOKES’ THEOREM As a final example, let’s computeH Bdrfor a closed loop, comparing the result with integralsR .rB/dfor two different surfaces having the same perimeter. We use spherical polar coordinates, taking BDerOe'. The loop will be a unit circle about the origin in the xy-plane; the line integral about it will be taken in a counterclockwise sense as viewed from positive z, so the normal to the surfaces it bounds will pass through the xy-plane in the direction of positive z. The surfaces we consider are (1) a circular disk bounded by the loop, and (3) a hemisphere bounded by the loop, with its surface in the region z<0. See Fig. 3.24. ArfKen_Ch03-9780123846549.tex 194 Chapter 3 Vector Analysis n nn n FIGURE 3.24 Surfaces for Example 3.10.5: (left) S1, disk; (right) S2, hemisphere. For the line integral, drDrsinOe'd', which reduces to drDOe'd'sinceD=2and rD1on the entire loop. We then have I BdrD2Z 'D0e1Oe'Oe'd'D2 e: For the surface integrals, we need rB: rBD1 r2sin@ @.rsiner/Oerr@ @r.rsiner/Oe Dercos rsinOer.1r/erOe: Taking first the disk, at all points of which D=2, with integration range 0r1, and0'<2, we note that dDOersindr d'DOer dr d'. The minus sign arises because the positive normal is in the direction of decreasing. Then, Z S1.rB/Oer dr d'D2Z 0d'1Z 0dr.1r/erD2 e: For the hemisphere, defined by rD1,=2< , and 0'<2, we have dD Oerr2sindd'DOersindd'(the normal is in the direction of decreasing r), and Z S2.rB/Oersindd'DZ =2de1cos2Z 0d'D2 e: The results for both surfaces agree with that from the line integral of their common perimeter. Because rBis solenoidal, all the flux that passes through the disk in the xy-plane must continue through the hemispherical surface, and for that matter, through anysurface with the same perimeter. That is why Stokes’ theorem is indifferent to features of the surface other than its perimeter.  ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 195 Rotation and Reflection in Spherical Coordinates It is infrequent that rotational coordinate transformations need be applied in curvilinear coordinate systems, and they usually arise only in contexts that are compatible with the symmetry of the coordinate system. We limit the current discussion to rotations (and reflections) in spherical polar coordinates. Rotation—Suppose a coordinate rotation identified by Euler angles . ; ; / converts the coordinates of a point from .r;;'/ to.r;0;'0/. It is obvious that rretains its original value. Two questions arise: (1) How are 0and'0related toand'? and (2) How do the components of a vector A, namely ( Ar;A;A'/, transform? It is simplest to proceed, as we did for Cartesian coordinates, by analyzing the three consecutive rotations implied by the Euler angles. The first rotation, by an angle about thez-axis, leavesunchanged, and converts 'into' . However, it causes no change in any of the components of A. The second rotation, which inclines the polar direction by an angle toward the (new) x-axis, does change the values of both and'and, in addition, changes the directions ofOeandOe'. Referring to Fig. 3.25, we see that these two unit vectors are subjected to a rotationin the plane tangent to the sphere of constant r, thereby yielding new unit vectorsOe0 andOe0 'such that OeDcosOe0 sinOe0 ';Oe'DsinOe0 CcosOe0 ': This transformation corresponds to S2Dcossin sincos : Carrying out the spherical trigonometry corresponding to Fig. 3.25, we have the new coordinates cos0Dcos cosCsin sincos.' /; cos'0Dcos cos0cos sin sin0;(3.166) P x xz z/primeϕ−αθ θ/prime ϕ/primeβ π−ϕ/primeˆeϕ/prime ˆeϕˆeθ/prime eθ FIGURE 3.25 Rotation and unit vectors in spherical polar coordinates, shown on a sphere of radius r. The original polar direction is marked z; it is moved to the direction z0, at an inclination given by the Euler angle . The unit vectorsOeandOe'at the point Pare thereby rotated through the angle . ArfKen_Ch03-9780123846549.tex 196 Chapter 3 Vector Analysis and cosDcos coscos0 sinsin0: (3.167) The third rotation, by an angle about the new z-axis, leaves the components of A unchanged but requires the replacement of '0by'0 . Summarizing, 0 @A0 r A0  A0 '1 AD0 @1 0 0 0 cos sin 0sincos1 A0 @Ar A A'1 A: (3.168) This equation specifies the components of Ain the rotated coordinates at the point .r;0;'0 /in terms of the original components at the same physical point, .r;;'/ . Reflection—Inversion of the coordinate system reverses the sign of each Cartesian coor- dinate. Taking the angle 'as that which moves the new Cxcoordinate toward the new Cy coordinate, the system (which was originally right-handed) now becomes left-handed. The coordinates.r;;'/ of a (fixed) point become, in the new system, .r;;C'/. The unit vectorsOerandOe'are invariant under inversion, but Oechanges sign, so 0 @A0 r A0  A0 '1 AD0 @Ar A A'1 A;coordinate inversion. (3.169) Exercises 3.10.1 Theu-,v-,z-coordinate system frequently used in electrostatics and in hydrodynamics is defined by xyDu;x2y2Dv; zDz: This u-,v-,z-system is orthogonal. (a) In words, describe briefly the nature of each of the three families of coordinate surfaces. (b) Sketch the system in the xy-plane showing the intersections of surfaces of constant uand surfaces of constant vwith the xy-plane. (c) Indicate the directions of the unit vectors OeuandOevin all four quadrants. (d) Finally, is this u-,v-,z-system right-handed .OeuOevDCOez/or left-handed .Oeu OevDOez/? 3.10.2 The elliptic cylindrical coordinate system consists of three families of surfaces: (1)x2 a2cosh2uCy2 a2sinh2uD1; (2)x2 a2cos2vy2 a2sin2vD1; (3) zDz. Sketch the coordinate surfaces uDconstant and vDconstant as they intersect the first quadrant of the xy-plane. Show the unit vectors OeuandOev. The range of uis0u<1. The range of vis0v2. ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 197 3.10.3 Develop arguments to show that dot and cross products (not involving r) in orthogonal curvilinear coordinates in R3proceed, as in Cartesian coordinates, with no involvement of scale factors. 3.10.4 WithOe1a unit vector in the direction of increasing q1, show that (a)rOe1D1 [email protected]/ @q1 (b)rOe1D1 h1 Oe21 h3@h1 @q3Oe31 h2@h1 @q2 . Note that even though Oe1is a unit vector, its divergence and curl do not necessarily vanish. 3.10.5 Show that a set of orthogonal unit vectors Oeimay be defined by OeiD1 hi@r @qi: In particular, show that OeiOeiD1leads to an expression for hiin agreement with Eq. (3.131). The above equation for Oeimay be taken as a starting point for deriving @Oei @qjDOej1 hi@hj @qi;i6Dj and @Oei @qiDX j6DiOej1 hj@hi @qj: 3.10.6 Resolve the circular cylindrical unit vectors into their Cartesian components (see Fig. 3.23). ANS.OeDOexcos'COeysin'; Oe'DOexsin'COeycos'; OezDOez: 3.10.7 Resolve the Cartesian unit vectors into their circular cylindrical components (see Fig. 3.23). ANS.OexDOecos'Oe'sin'; OeyDOesin'COe'cos'; OezDOez: 3.10.8 From the results of Exercise 3.10.6, show that @Oe @'DOe';@Oe' @'DOe and that all other first derivatives of the circular cylindrical unit vectors with respect to the circular cylindrical coordinates vanish. ArfKen_Ch03-9780123846549.tex 198 Chapter 3 Vector Analysis 3.10.9 Compare rVas given for cylindrical coordinates in Eq. (3.148) with the result of its computation by applying to Vthe operator rDOe@ @COe'1 @ @'COez@ @z Note that racts both on the unit vectors and on the components of V. 3.10.10 (a) Show that rDOeCOezz. (b) Working entirely in circular cylindrical coordinates, show that rrD3andrrD0: 3.10.11 (a) Show that the parity operation (reflection through the origin) on a point .;'; z/ relative to fixed x-,y-,z-axes consists of the transformation !; '!'; z! z: (b) Show thatOeandOe'have odd parity (reversal of direction) and that Oezhas even parity. Note. The Cartesian unit vectors Oex,Oey;andOezremain constant. 3.10.12 A rigid body is rotating about a fixed axis with a constant angular velocity !. Take! to lie along the z-axis. Express the position vector rin circular cylindrical coordinates and using circular cylindrical coordinates, (a) calculate vD!r, (b) calculate rv. ANS..a/ vDOe'! .b/rvD2!: 3.10.13 Find the circular cylindrical components of the velocity and acceleration of a moving particle, vDP; aDRP'2; v'DP';a'DR'C2PP'; vzDPz; azDRz: Hint. r.t/DOe.t/.t/COezz.t/ DTOexcos'.t/COeysin'.t/U.t/COezz.t/: Note.PDd=dt ,RDd2=dt2, and so on. 3.10.14 In right circular cylindrical coordinates, a particular vector function is given by V.;'/DOeV.;'/COe'V'.;'/: Show that rVhas only a z-component. Note that this result will hold for any vector confined to a surface q3Dconstant as long as the products h1V1andh2V2are each independent of q3. ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 199 3.10.15 A conducting wire along the z-axis carries a current I. The resulting magnetic vector potential is given by ADOezI 2ln1  : Show that the magnetic induction Bis given by BDOe'I 2: 3.10.16 A force is described by FDOexy x2Cy2COeyx x2Cy2: (a) Express Fin circular cylindrical coordinates. Operating entirely in circular cylindrical coordinates for (b) and (c), (b) Calculate the curl of Fand (c) Calculate the work done by Fin encircling the unit circle once counter-clockwise. (d) How do you reconcile the results of (b) and (c)? 3.10.17 A calculation of the magnetohydrodynamic pinch effect involves the evaluation of .Br/B. If the magnetic induction Bis taken to be BDOe'B'./, show that .Br/BDOeB2 '=: 3.10.18 Express the spherical polar unit vectors in terms of Cartesian unit vectors. ANS.OerDOexsincos'COeysinsin'COezcos; OeDOexcoscos'COeycossin'Oezsin; Oe'DOexsin'COeycos': 3.10.19 Resolve the Cartesian unit vectors into their spherical polar components: OexDOersincos'COecoscos'Oe'sin'; OeyDOersinsin'COecossin'COe'cos'; OezDOercosOesin: 3.10.20 (a) Explain why it is not possible to relate a column vector r(with components x, y,z) to another column vector r0(with components r,,'), via a matrix equation of the form r0DBr. (b) One can write a matrix equation relating the Cartesian components of a vector to its components in spherical polar coordinates. Find the transformation matrix and determine whether it is orthogonal. 3.10.21 Find the transformation matrix that converts the components of a vector in spherical polar coordinates into its components in circular cylindrical coordinates. Then find the matrix of the inverse transformation. 3.10.22 (a) From the results of Exercise 3.10.18, calculate the partial derivatives of Oer,Oe;and Oe'with respect to r,, and'. ArfKen_Ch03-9780123846549.tex 200 Chapter 3 Vector Analysis (b) With rgiven by Oer@ @rCOe1 r@ @COe'1 rsin@ @' (greatest space rate of change), use the results of part (a) to calculate rr . This is an alternate derivation of the Laplacian. Note. The derivatives of the left-hand roperate on the unit vectors of the right-hand r before the dot product is evaluated. 3.10.23 A rigid body is rotating about a fixed axis with a constant angular velocity !. Take!to be along the z-axis. Using spherical polar coordinates, (a) calculate vD!r: (b) calculate rv: ANS..a/ vDOe'!rsin: .b/rvD2!: 3.10.24 A certain vector Vhas no radial component. Its curl has no tangential components. What does this imply about the radial dependence of the tangential components of V? 3.10.25 Modern physics lays great stress on the property of parity (whether a quantity remains invariant or changes sign under an inversion of the coordinate system). In Cartesian coordinates this means x! x,y! y, and z! z. (a) Show that the inversion (reflection through the origin) of a point .r;;'/ relative tofixed x-,y-,z-axes consists of the transformation r!r; !; '!': (b) Show thatOerandOe'have odd parity (reversal of direction) and that Oehas even parity. 3.10.26 With Aany vector, ArrDA: (a) Verify this result in Cartesian coordinates. (b) Verify this result using spherical polar coordinates. Equation (3.156) provides r. 3.10.27 Find the spherical coordinate components of the velocity and acceleration of a moving particle: vrDPr; arDRrrP2rsin2P'2; vDrP; aDrRC2PrPrsincosP'2; v'DrsinP';a'DrsinR'C2PrsinP'C2rcosPP': ArfKen_Ch03-9780123846549.tex 3.10 Curvilinear Coordinates 201 Hint. r.t/DOer.t/r.t/ DTOexsin.t/cos'.t/COeysin.t/sin'.t/COezcos.t/Ur.t/: Note. The dot inPr;P;P'means time derivative: PrDdr=dt;PDd=dt; P'Dd'=dt . 3.10.28 Express@=@x,@=@y,@=@zin spherical polar coordinates. ANS.@ @xDsincos'@ @rCcoscos'1 r@ @sin' rsin@ @'; @ @yDsinsin'@ @rCcossin'1 r@ @Ccos' rsin@ @'; @ @zDcos@ @rsin1 r@ @: Hint. Equate rxyzandrr'. 3.10.29 Using results from Exercise 3.10.28, show that i x@ @yy@ @x Di@ @': This is the quantum mechanical operator corresponding to the z-component of orbital angular momentum. 3.10.30 With the quantum mechanical orbital angular momentum operator defined as LD i.rr/, show that (a) LxCi LyDei'@ @Cicot@ @' ; (b) Lxi LyDei'@ @icot@ @' : 3.10.31 Verify that LLDiLin spherical polar coordinates. LDi.rr/, the quantum mechanical orbital angular momentum operator. Written in component form, this relation is LyLzLzLyDi Lx;LzLxLxLzDLy;LxLyLyLxDi Lz: Using the commutator notation, TA;BUDABB A, and the definition of the Levi- Civita symbol "i jk, the above can also be written TLi;LjUDi"i jkLk; where i,j,karex;y;zin any order. Hint. Use spherical polar coordinates for Lbut Cartesian components for the cross product. ArfKen_Ch03-9780123846549.tex 202 Chapter 3 Vector Analysis 3.10.32 (a) Using Eq. (3.156) show that LDi.rr/Di Oe1 sin@ @'Oe'@ @ : (b) ResolvingOeandOe'into Cartesian components, determine Lx,Ly, and Lzin terms of,', and their derivatives. (c) From L2DL2 xCL2 yCL2 zshow that L2D1 sin@ @ sin@ @ 1 sin2@2 @'2 Dr2r2C@ @r r2@ @r : 3.10.33 With LDirr, verify the operator identities (a)rDOer@ @rirL r2, (b) rr2r 1Cr@ @r DirL. 3.10.34 Show that the following three forms (spherical coordinates) of r2 .r/are equivalent: (a)1 r2d dr r2d .r/ dr ; (b)1 rd2 dr2Tr .r/U; (c)d2 .r/ dr2C2 rd .r/ dr. The second form is particularly convenient in establishing a correspondence between spherical polar and Cartesian descriptions of a problem. 3.10.35 A certain force field is given in spherical polar coordinates by FDOer2Pcos r3COeP r3sin; rP=2: (a) Examine rFto see if a potential exists. (b) CalculateH Fdrfor a unit circle in the plane D=2. What does this indicate about the force being conservative or nonconservative? (c) If you believe that Fmay be described by FDr , find . Otherwise simply state that no acceptable potential exists. 3.10.36 (a) Show that ADOe'cot=ris a solution of rADOer=r2. (b) Show that this spherical polar coordinate solution agrees with the solution given for Exercise 3.9.5: ADOexyz r.x2Cy2/Oeyxz r.x2Cy2/: Note that the solution diverges for D0,corresponding to x,yD0. (c) Finally, show that ADOe'sin=ris a solution. Note that although this solution does not diverge .r6D0/;it is no longer single-valued for all possible azimuth angles. ArfKen_Ch03-9780123846549.tex Additional Readings 203 3.10.37 An electric dipole of moment pis located at the origin. The dipole creates an electric potential at rgiven by .r/Dpr 4" 0r3: Find the electric field, EDr atr. Additional Readings Borisenko, A. I., and I. E. Tarpov, Vector and Tensor Analysis with Applications. Englewood Cliffs, NJ: Prentice- Hall (1968), reprinting, Dover (1980). Davis, H. F., and A. D. Snider, Introduction to Vector Analysis, 7th ed. Boston: Allyn & Bacon (1995). Kellogg, O. D., Foundations of Potential Theory. Berlin: Springer (1929), reprinted, Dover (1953). The classic text on potential theory. Lewis, P. E., and J. P. Ward, Vector Analysis for Engineers and Scientists. Reading, MA: Addison-Wesley (1989). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry, 2nd ed. Princeton NJ: Van Nos- trand (1956). Chapter 5 covers curvilinear coordinates and 13 specific coordinate systems. Marion, J. B., Principles of Vector Analysis. New York: Academic Press (1965). A moderately advanced presen- tation of vector analysis oriented toward tensor analysis. Rotations and other transformations are described with the appropriate matrices. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 includes a description of several different coordinate systems. Note that Morse and Feshbach are not above using left-handed coordinate systems even for Cartesian coordinates. Elsewhere in this excellent (and diffi- cult) book there are many examples of the use of the various coordinate systems in solving physical problems. Eleven additional fascinating but seldom-encountered orthogonal coordinate systems are discussed in the sec- ond (1970) edition of Mathematical Methods for Physicists. Spiegel, M. R., Vector Analysis. New York: McGraw-Hill (1989). Tai, C.-T., Generalized Vector and Dyadic Analysis. Oxford: Oxford University Press (1966). Wrede, R. C., Introduction to Vector and Tensor Analysis. New York: Wiley (1963), reprinting, Dover (1972). Fine historical introduction. Excellent discussion of differentiation of vectors and applications to mechanics. ArfKen_Ch04-9780123846549.tex CHAPTER 4 TENSORS AND DIFFERENTIAL FORMS 4.1 T ENSOR ANALYSIS Introduction, Properties Tensors are important in many areas of physics, ranging from topics such as general relativ- ity and electrodynamics to descriptions of the properties of bulk matter such as stress (the pattern of force applied to a sample) and strain (its response to the force), or the moment of inertia (the relation between a torsional force applied to an object and its resultant angu- lar acceleration). Tensors constitute a generalization of quantities previously introduced: scalars and vectors. We identified a scalar as an quantity that remained invariant under rotations of the coordinate system and which could be specified by the value of a sin- gle real number. Vectors were identified as quantities that had a number of real compo- nents equal to the dimension of the coordinate system, with the components transforming like the coordinates of a fixed point when a coordinate system is rotated. Calling scalars tensors of rank 0 and vectors tensors of rank 1, we identify a tensor of rank nin a d-dimensional space as an object with the following properties: It has components labeled by nindices, with each index assigned values from 1 through d, and therefore having a total of dncomponents; The components transform in a specified manner under coordinate transformations. The behavior under coordinate transformation is of central importance for tensor anal- ysis and conforms both with the way in which mathematicians define linear spaces and 205 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch04-9780123846549.tex 206 Chapter 4 Tensors and Differential Forms with the physicist’s notion that physical observables must not depend on the choice of coordinate frames. Covariant and Contravariant Tensors In Chapter 3, we considered the rotational transformation of a vector ADA1Oe1CA2Oe2C A3Oe3from the Cartesian system defined by Oei(iD1;2;3) into a rotated coordinate system defined byOe0 i, with the same vector Athen represented as A0DA0 1Oe0 1CA0 2Oe0 2CA0 3e0 3. The components of AandA0are related by A0 iDX j.Oe0 iOej/Aj; (4.1) where the coefficients .Oe0 iOej/are the projections of Oe0 iin theOejdirections. Because the Oei and theOejare linearly related, we can also write A0 iDX j@x0 i @xjAj: (4.2) The formula of Eq. (4.2) corresponds to the application of the chain rule to convert the set Ajinto the set A0 i, and is valid for AjandA0 iof arbitrary magnitude because both vectors depend linearly on their components. We have also previously noted that the gradient of a scalar 'has in the unrotated Carte- sian coordinates the components .r'/jD.@'=@ xj/Oej, meaning that in a rotated system we would have .r'/0 i@' @x0 iDX j@xj @x0 i@' @xj; (4.3) showing that the gradient has a transformation law that differs from that of Eq. (4.2) in that@x0 i=@xjhas been replaced by @xj=@x0 i. Remembering that these two expressions, if written in detail, correspond, respectively, to .@x0 i=@xj/xkand.@xj=@x0 i/x0 k, where kruns over the index values other than that already in the denominator, and also noting that (in Cartesian coordinates) they are two different ways of computing the same quantity (the magnitude and sign of the projection of one of these unit vectors upon the other), we see that it was legitimate to identify both Aandr'asvectors, as we did in Chapter 3. However, as the alert reader may note from the repeated insertion of the word “Cartesian,” the partial derivatives in Eqs. (4.2) and(4.3) are only guaranteed to be equal in Cartesian coordinate systems, and since there is sometimes a need to use non-Cartesian systems it becomes necessary to distinguish these two different transformation rules. Quan- tities transforming according to Eq. (4.2) are called contravariant vectors, while those transforming according to Eq. (4.3) are termed covariant. When non-Cartesian systems may be in play, it is therefore customary to distinguish these transformation properties by writing the index of a contravariant vector as a superscript and that of a covariant vector as a subscript. This means, among other things, that the components of the position vector r, ArfKen_Ch04-9780123846549.tex 4.1 Tensor Analysis 207 which is contravariant, must now be written .x1;x2;x3/. Thus, summarizing, .A0/iDX [email protected]/i @xjAjA, a contravariant vector, (4.4) A0 iDX j@xj @.x0/iAjA, a covariant vector. (4.5) It is useful to note that the occurrence of subscripts and superscripts is systematic; the free (i.e., unsummed) index ioccurs as a superscript on both sides of Eq. (4.4), while it appears as a subscript on both sides of Eq. (4.5), if we interpret an upper index in the denominator as equivalent to a lower index. The summed index occurs once as upper and once as lower (again treating an upper index in the denominator as a lower index). A frequently used shorthand (the Einstein convention) is to omit the summation sign in formulas like Eqs. (4.4) and(4.5) and to understand that when the same symbol occurs both as an upper and a lower index in the same expression, it is to be summed. We will gradually back into the use of the Einstein convention, giving the reader warnings as we start to do so. Tensors of Rank 2 Now we proceed to define contravariant, mixed, and covariant tensors of rank 2 by the following equations for their components under coordinate transformations: .A0/i jDX [email protected]/i @[email protected]/j @xlAkl; .B0/i jDX [email protected]/i @xk@xl @.x0/jBk l; (4.6) .C0/i jDX kl@xk @.x0/i@xl @.x0/jCkl: Clearly, the rank goes as the number of partial derivatives (or direction cosines) in the definition: 0 for a scalar, 1 for a vector, 2 for a second-rank tensor, and so on. Each index (subscript or superscript) ranges over the number of dimensions of the space. The number of indices (equal to the rank of tensor) is not limited by the dimensionality of the space. We see that Aklis contravariant with respect to both indices, Cklis covariant with respect to both indices, and Bk ltransforms contravariantly with respect to the index kbut covariantly with respect to the index l. Once again, if we are using Cartesian coordinates, all three forms of the tensors of second rank, contravariant, mixed, and covariant are the same. As with the components of a vector, the transformation laws for the components of a tensor, Eq. (4.6), cause its physically relevant properties to be independent of the choice of reference frame. This is what makes tensor analysis important in physics. The inde- pendence relative to reference frame (invariance) is ideal for expressing and investigating universal physical laws. ArfKen_Ch04-9780123846549.tex 208 Chapter 4 Tensors and Differential Forms The second-rank tensor A(with components Akl) may be conveniently represented by writing out its components in a square array ( 33if we are in three-dimensional (3-D) space): AD0 @A11A12A13 A21A22A23 A31A32A331 A: (4.7) This does not mean that any square array of numbers or functions forms a tensor. The essential condition is that the components transform according to Eq. (4.6). We can view each of Eq. (4.6) as a matrix equation. For A, it takes the form .A0/i jDX klSikAkl.ST/l j;orA0DSAST; (4.8) a construction that is known as a similarity transformation and is discussed in Section 5.6. In summary, tensors are systems of components organized by one or more indices that transform according to specific rules under a set of transformations. The number of indices is called the rank of the tensor. Addition and Subtraction of Tensors The addition and subtraction of tensors is defined in terms of the individual elements, just as for vectors. If A + B = C; (4.9) then, taking as an example A,B, and Cto be contravariant tensors of rank 2, Ai jCBi jDCi j: (4.10) In general, of course, AandBmust be tensors of the same rank (of both contra- and co-variance) and in the same space. Symmetry The order in which the indices appear in our description of a tensor is important. In general, Amnis independent of Anm, but there are some cases of special interest. If, for all mandn, AmnDAnm;Aissymmetric. (4.11) If, on the other hand, AmnDAnm;Aisantisymmetric. (4.12) Clearly, every (second-rank) tensor can be resolved into symmetric and antisymmetric parts by the identity AmnD1 2.AmnCAnm/C1 2.AmnAnm/; (4.13) the first term on the right being a symmetric tensor, the second, an antisymmetric tensor. ArfKen_Ch04-9780123846549.tex 4.1 Tensor Analysis 209 Isotropic Tensors To illustrate some of the techniques of tensor analysis, let us show that the now-familiar Kronecker delta, kl, is really a mixed tensor of rank 2, k l.1The question is: Does k ltrans- form according to Eq. (4.6)? This is our criterion for calling it a tensor. If k lis the mixed tensor corresponding to this notation, it must satisfy (using the summation convention, meaning that the indices kandlare to be summed) .0/i [email protected]/i @xk@xl @.x0/jk [email protected]/i @xk@xk @.x0/j; where we have performed the lsum and used the definition of the Kronecker delta. Next, @.x0/i @xk@xk @.x0/[email protected]/i @.x0/j; where we have identified the ksummation on the left-hand side as an instance of the chain rule for differentiation. However, .x0/iand.x0/jare independent coordinates, and therefore the variation of one with respect to the other must be zero if they are different, unity if they coincide; that is, @.x0/i @.x0/jD.0/i j: (4.14) Hence .0/i [email protected]/i @xk@xl @.x0/jk l; (4.15) showing that the k lare indeed the components of a mixed second-rank tensor. Note that this result is independent of the number of dimensions of our space. The Kronecker delta has one further interesting property. It has the same components in all of our rotated coordinate systems and is therefore called isotropic. In Section 4.2 and Exercise 4.2.4 we shall meet a third-rank isotropic tensor and three fourth-rank isotropic tensors. No isotropic first-rank tensor (vector) exists. Contraction When dealing with vectors, we formed a scalar product by summing products of corre- sponding components: ABDX iAiBi: The generalization of this expression in tensor analysis is a process known as contraction. Two indices, one covariant and the other contravariant, are set equal to each other, and then (as implied by the summation convention) we sum over this repeated index. For example, 1It is common practice to refer to a tensor Aby specifying a typical component, such as Ai j, thereby also conveying information as to its covariant vs. contravariant nature. As long as you refrain from writing nonsense such as ADAi j, no harm is done. ArfKen_Ch04-9780123846549.tex 210 Chapter 4 Tensors and Differential Forms let us contract the second-rank mixed tensor Bi jby setting jtoi, then summing over i. To see what happens, let’s look at the transformation formula that converts BintoB0. Using the summation convention, .B0/i [email protected]/i @xk@xl @.x0/iBk lD@xl @xkBk l; where we recognized the isummation as an instance of the chain rule for differentiation. Then, because the xiare independent, we may use Eq. (4.14) to reach .B0/i iDl kBk lDBk k: (4.16) Remembering that the repeated index ( iork) is summed, we see that the contracted B is invariant under transformation and is therefore a scalar.2In general, the operation of contraction reduces the rank of a tensor by 2. Direct Product The components of two tensors (of any ranks and covariant/contravariant characters) can be multiplied, component by component, to make an object with all the indices of both factors. The new quantity, termed the direct product of the two tensors, can be shown to be a tensor whose rank is the sum of the ranks of the factors, and with covariant/contravariant character that is the sum of those of the factors. We illustrate: Ci j klmDAi kBj lm;Fi j klDAjBi lk: Note that the index order in the direct product can be defined as desired, but the covari- ance/contravariance of the factors must be maintained in the direct product. Example 4.1.1 DIRECT PRODUCT OF TWO VECTORS Let’s form the direct product of a covariant vector ai(rank-1 tensor) and a contravariant vector bj(also a rank-1 tensor) to form a mixed tensor of rank 2, with components Cj iD aibj. To verify that Cj iis a tensor, we consider what happens to it under transformation: .C0/j iD.a0/i.b0/jD@xk @.x0/[email protected]/j @xlblD@xk @.x0/[email protected]/j @xlCl k; (4.17) confirming that Cj iis the mixed tensor indicated by its notation. If we now form the contraction Ci i(remember that iis summed), we obtain the scalar product aibi. From Eq. (4.17) it is easy to see that aibiD.a0/i.b0/i, indicating the invari- ance required of a scalar product.  Note that the direct product concept gives a meaning to quantities such as rE, which was not defined within the framework of vector analysis. However, this and other tensor- like quantities involving differential operators must be used with caution, because their 2In matrix analysis this scalar is the trace of the matrix whose elements are the Bi j. ArfKen_Ch04-9780123846549.tex 4.1 Tensor Analysis 211 transformation rules are simple only in Cartesian coordinate systems. In non-Cartesian systems, operators @=@xiact also on the partial derivatives in the transformation expres- sions and alter the tensor transformation rules. We summarize the key idea of this subsection: The direct product is a technique for creating new, higher-rank tensors. Inverse Transformation If we have a contravariant vector Ai, which must have the transformation rule (using sum- mation convention) .A0/[email protected]/j @xiAi; the inverse transformation (which can be obtained simply by interchanging the roles of the primed and unprimed quantities) is AiD@xi @.x0/j.A0/j; (4.18) as may also be verified by applying @.x0/k=@xi(and summing i) to Aias given by Eq. (4.18): @.x0/k @[email protected]/k @xi@xi @.x0/j.A0/jDk j.A0/jD.A0/k: (4.19) We see that.A0/kis recovered. Incidentally, note that @xi @.x0/j6D@.x0/j @xi1 I as we have previously pointed out, these derivatives have different other variables held fixed. The cancellation in Eq. (4.19) only occurs because the product of derivatives is summed. In Cartesian systems, we do have @xi @.x0/[email protected]/j @xi; both equal to the direction cosine connecting the xiand.x0/jaxes, but this equality does not extend to non-Cartesian systems. Quotient Rule If, for example, Ai jandBklare tensors, we have already observed that their direct product, Ai jBkl, is also a tensor. Here we are concerned with the inverse problem, illustrated by ArfKen_Ch04-9780123846549.tex 212 Chapter 4 Tensors and Differential Forms equations such as KiAiDB; Kj iAjDBi; Kj iAjkDBik; (4.20) Ki jklAi jDBkl; Ki jAkDBi jk: In each of these expressions AandBare known tensors of ranks indicated by the number of indices, Ais arbitrary, and the summation convention is in use. In each case Kis an unknown quantity. We wish to establish the transformation properties of K. The quotient rule asserts: If the equation of interest holds in all transformed coordinate systems, then Kis a tensor of the indicated rank and covariant/contravariant character. Part of the importance of this rule in physical theory is that it can establish the tensor nature of quantities. For example, the equation giving the dipole moment minduced in an anisotropic medium by an electric field Eis miDPi jEj: Since presumably we know that mandEare vectors, the general validity of this equation tells us that the polarization matrix Pis a tensor of rank 2. Let’s prove the quotient rule for a typical case, which we choose to be the second of Eqs. (4.20). If we apply a transformation to that equation, we have Kj iAjDBi!.K0/j iA0 jDB0 i: (4.21) We now evaluate B0 i, reaching the last member of the equation below by using Eq. (4.18) to convert Ajinto components of A0(note that this is the inverse of the transformation to the primed quantities): B0 iD@xm @.x0/iBmD@xm @.x0/iKj mAjD@xm @.x0/iKj [email protected]/n @xjA0 n: (4.22) It may lessen possible confusion if we rename the dummy indices in Eq. (4.22), so we interchange nandj, causing that equation to then read B0 iD@xm @.x0/[email protected]/j @xnKn mA0 j: (4.23) It has now become clear that if we subtract the expression for B0 iin Eq. (4.23) from that in Eq. (4.21) we will get  .K0/j i@xm @.x0/[email protected]/j @xnKn m A0 jD0: (4.24) Since A0is arbitrary, the coefficient of A0 jin Eq. (4.24) must vanish, showing that Khas the transformation properties of the tensor corresponding to its index configuration. ArfKen_Ch04-9780123846549.tex 4.1 Tensor Analysis 213 Other cases may be treated similarly. One minor pitfall should be noted: The quotient rule does not necessarily apply if Bis zero. The transformation properties of zero are indeterminate. Example 4.1.2 EQUATIONS OF MOTION AND FIELD EQUATIONS In classical mechanics, Newton’s equations of motion mPvDFtell us on the basis of the quotient rule that, if the mass is a scalar and the force a vector, then the acceleration aPv is a vector. In other words, the vector character of the force as the driving term imposes its vector character on the acceleration, provided the scale factor mis scalar. The wave equation of electrodynamics can be written in relativistic four-vector form as 1 c2@2 @t2r2 ADJ; where Jis the external charge/current density (a four-vector) and Ais the four- component vector potential. The second-derivative expression in square brackets can be shown to be a scalar. From the quotient rule, we may then infer that Amust be a tensor of rank 1, i.e., also a four-vector.  The quotient rule is a substitute for the illegal division of tensors. Spinors It was once thought that the system of scalars, vectors, tensors (second-rank), and so on formed a complete mathematical system, one that is adequate for describing a physics independent of the choice of reference frame. But the universe and mathematical physics are not that simple. In the realm of elementary particles, for example, spin-zero particles3 (mesons, particles) may be described with scalars, spin 1 particles (deuterons) by vectors, and spin 2 particles (gravitons) by tensors. This listing omits the most common particles: electrons, protons, and neutrons, all with spin1 2. These particles are properly described by spinors. A spinor does not have the properties under rotation consistent with being a scalar, vector, or tensor of any rank. A brief introduction to spinors in the context of group theory appears in Chapter 17. Exercises 4.1.1 Show that if all the components of any tensor of any rank vanish in one particular coordinate system, they vanish in all coordinate systems. Note. This point takes on special importance in the four-dimensional (4-D) curved space of general relativity. If a quantity, expressed as a tensor, exists in one coordinate sys- tem, it exists in all coordinate systems and is not just a consequence of a choice of a coordinate system (as are centrifugal and Coriolis forces in Newtonian mechanics). 3The particle spin is intrinsic angular momentum (in units of Nh). It is distinct from classical (often called orbital) angular momentum that arises from the motion of the particle. ArfKen_Ch04-9780123846549.tex 214 Chapter 4 Tensors and Differential Forms 4.1.2 The components of tensor Aare equal to the corresponding components of tensor Bin one particular coordinate system denoted, by the superscript 0; that is, A0 i jDB0 i j: Show that tensor Ais equal to tensor B,Ai jDBi j, in all coordinate systems. 4.1.3 The last three components of a 4-D vector vanish in each of two reference frames. If the second reference frame is not merely a rotation of the first about the x0axis, meaning that at least one of the coefficients @.x0/[email protected];2;3/is nonzero, show that the zeroth component vanishes in all reference frames. Translated into relativistic mechan- ics, this means that if momentum is conserved in two Lorentz frames, then energy is conserved in all Lorentz frames. 4.1.4 From an analysis of the behavior of a general second-rank tensor under 90and180 rotations about the coordinate axes, show that an isotropic second-rank tensor in 3-D space must be a multiple of i j. 4.1.5 The 4-D fourth-rank Riemann-Christoffel curvature tensor of general relativity, Riklm, satisfies the symmetry relations RiklmDRikmlDRkilm: With the indices running from 0 to 3, show that the number of independent components is reduced from 256 to 36 and that the condition RiklmDRlmik further reduces the number of independent components to 21. Finally, if the components satisfy an identity RiklmCRilmkCRimklD0, show that the number of independent components is reduced to 20. Note. The final three-term identity furnishes new information only if all four indices are different. 4.1.6 Tiklmis antisymmetric with respect to all pairs of indices. How many independent com- ponents has it (in 3-D space)? 4.1.7 IfT:::iis a tensor of rank n, show that@T:::i=@xjis a tensor of rank nC1(Cartesian coordinates). Note. In non-Cartesian coordinate systems the coefficients ai jare, in general, functions of the coordinates, and the derivatives the components of a tensor of rank ndo not form a tensor except in the special case nD0. In this case the derivative does yield a covariant vector (tensor of rank 1). 4.1.8 IfTi jk:::is a tensor of rank n, show thatP j@Ti jk:::=@xjis a tensor of rank n1 (Cartesian coordinates). 4.1.9 The operator r21 c2@2 @t2 ArfKen_Ch04-9780123846549.tex 4.2 Pseudsotensors, Dual Tensors 215 may be written as 4X iD1@2 @x2 i; using x4Dict. This is the 4-D Laplacian, sometimes called the d’Alembertian and denoted by 2. Show that it is a scalar operator, that is, invariant under Lorentz trans- formations, i.e., under rotations in the space of vectors ( x1;x2;x3;x4). 4.1.10 The double summation Ki jAiBjis invariant for any two vectors AiandBj. Prove that Ki jis a second-rank tensor. Note. In the form ds2(invariant)Dgi jdxidxj, this result shows that the matrix gi jis a tensor. 4.1.11 The equation Ki jAjkDBk iholds for all orientations of the coordinate system. If Aand Bare arbitrary second-rank tensors, show that Kis a second-rank tensor also. 4.2 P SEUDOTENSORS , DUALTENSORS The topics of this section will be treated for tensors restricted for practical reasons to Carte- sian coordinate systems. This restriction is not conceptually necessary but simplifies the discussion and makes the essential points easy to identify. Pseudotensors So far the coordinate transformations in this chapter have been restricted to passive rota- tions, by which we mean rotation of the coordinate system, keeping vectors and tensors at fixed orientations. We now consider the effect of reflections or inversions of the coordinate system (sometimes also called improper rotations). In Section 3.3, where attention was restricted to orthogonal systems of Cartesian coor- dinates, we saw that the effect of a coordinate rotation on a fixed vector could be described by a transformation of its components according to the formula A0DSA; (4.25) where Swas an orthogonal matrix with determinant C1. If the coordinate transformation included a reflection (or inversion), the transformation matrix was still orthogonal, but had determinant1. While the transformation rule of Eq. (4.25) was obeyed by vectors describing quantities such as position in space or velocity, it produced the wrong sign when vectors describing angular velocity, torque, and angular momentum were subject to improper rotations. These quantities, called axial vectors, or nowadays pseudovectors, obeyed the transformation rule A0Ddet.S/SA (pseudovector). (4.26) The extension of this concept to tensors is straightforward. We insist that the designation tensor refer to objects that transform as in Eq. (4.6) and its generalization to arbitrary ArfKen_Ch04-9780123846549.tex 216 Chapter 4 Tensors and Differential Forms rank, but we also accommodate the possibility of having, at arbitrary rank, objects whose transformation requires an additional sign factor to adjust for the effect associated with improper rotations. These objects are called pseudotensors, and constitute a generalization of the objects already identified as pseudoscalars and pseudovectors. If we form a tensor or pseudotensor as a direct product or identify one via the quotient rule, we can determine its pseudo status by what amounts to a sign rule. Letting Tbe a tensor and Pa pseudotensor, then, symbolically, T TDP PDT; T PDP TDP: (4.27) Example 4.2.1 LEVI-CIVITA SYMBOL The three-index version of the Levi-Civita symbol, introduced in Eq. (2.8), has the values "123D"231D"312DC1; "132D"213D"321D1; (4.28) all other"i jkD0: Suppose now that we have a rank-3 pseudotensor i jk, which in one particular Cartesian coordinate system is equal to "i jk. Then, letting Astand for the matrix of coefficients in an orthogonal transformation of R3, we have in the transformed coordinate system 0 i jkDdet.A/X pqraipajqakr"pqr; (4.29) by definition of pseudotensor. All terms of the pqr sum will vanish except those where pqris a permutation of 123, and when pqris such a permutation the sum will correspond to the determinant of Aexcept that its rows will have been permuted from 123 to i jk. This means that the pqrsum will have the value "i jkdet.A/ , and 0 i jkD"i jk[det.A/ ]2D"i jk; (4.30) where the final result depends on the fact that jdet.A/jD 1. If the reader is uncomfortable with the above analysis, the result can be checked by enumeration of the contributions of the six permutations that correspond to nonzero values of 0 i jk. Equation (4.30) not only shows that "is a rank-3 pseudotensor, but that it is also isotropic. In other words, it has the same components in all rotated Cartesian coordinate systems, and1times those component values in all Cartesian systems that are reached by improper rotations.  Dual Tensors With any antisymmetric second-rank tensor C(in 3-D space) we may associate a pseu- dovector Cwith components defined by CiD1 2"i jkCjk: (4.31) ArfKen_Ch04-9780123846549.tex 4.2 Pseudsotensors, Dual Tensors 217 In matrix form the antisymmetric Cmay be written CD0 B@0 C12C31 C120 C23 C31C2301 CA: (4.32) We know that Cimust transform as a vector under rotations because it was obtained from the double contraction of "i jkCjk, but that it is really a pseudovector because of the pseudo nature of"i jk. Specifically, the components of Care given by .C1;C2;C3/D.C23;C31;C12/: (4.33) Note the cyclic order of the indices that comes from the cyclic order of the components of"i jk. We identify the pseudovector of Eq. (4.33) and the antisymmetric tensor of Eq. (4.32) asdual tensors; they are simply different representations of the same information. Which of the dual pair we choose to use is a matter of convenience. Here is another example of duality. If we take three vectors A,B, and C, we may define the direct product Vi jkDAiBjCk: (4.34) Vi jkis evidently a rank-3 tensor. The dual quantity VD"i jkVi jk(4.35) is clearly a pseudoscalar. By expansion it is seen that VD A1B1C1 A2B2C2 A3B3C3 (4.36) is our familiar scalar triple product. Exercises 4.2.1 An antisymmetric square array is given by 0 @0C3C2 C30C1 C2C101 AD0 @0 C12C13 C120C23 C13C2301 A; where.C1;C2;C3/form a pseudovector. Assuming that the relation CiD1 2W"i jkCjk holds in all coordinate systems, prove that Cjkis a tensor. (This is another form of the quotient theorem.) 4.2.2 Show that the vector product is unique to 3-D space, that is, only in three dimensions can we establish a one-to-one correspondence between the components of an antisymmetric tensor (second-rank) and the components of a vector. ArfKen_Ch04-9780123846549.tex 218 Chapter 4 Tensors and Differential Forms 4.2.3 WriterrAandrr'in tensor (index) notation in IR3so that it becomes obvious that each expression vanishes. AN S:rrAD"i jk@ @xi@ @xjAk .rr'/iD"i jk@ @xj@ @xk': 4.2.4 Verify that each of the following fourth-rank tensors is isotropic, that is, that it has the same form independent of any rotation of the coordinate systems. (a) Aik jlDi jk l, (b) Bi j klDi kj lCi lj k, (c) Ci j klDi kj li lj k. 4.2.5 Show that the two-index Levi-Civita symbol "i jis a second-rank pseudotensor (in two- dimensional [2-D] space). Does this contradict the uniqueness of i j(Exercise 4.1.4)? 4.2.6 Represent"i jby a22matrix, and using the 22rotation matrix of Eq. (3.23), show that"i jis invariant under orthogonal similarity transformations. 4.2.7 Given AkD1 2"i jkBi jwith Bi jDBji, antisymmetric, show that BmnD"mnkAk: 4.3 T ENSORS IN GENERAL COORDINATES Metric Tensor The distinction between contravariant and covariant transformations was established in Section 4.1, where we also observed that it only became meaningful when working with coordinate systems that are not Cartesian. We now want to examine relationships that can systematize the use of more general metric spaces (also called Riemannian spaces). Our initial illustrations will be for spaces with three dimensions. Letting qidenote coordinates in a general coordinate system, writing the index as a superscript to reflect the fact that coordinates transform contravariantly, we define covari- ant basis vectors "ithat describe the displacement (in Euclidean space) per unit change inqi, keeping the other qjconstant. For the situations of interest here, both the direction and magnitude of "imay be functions of position, so it is defined as the derivative "iD@x @qiOexC@y @qiOeyC@z @qiOez: (4.37) An arbitrary vector Acan now be formed as a linear combination of the basis vectors, multiplied by coefficients: ADA1"1CA2"2CA3"3: (4.38) ArfKen_Ch04-9780123846549.tex 4.3 Tensors in General Coordinates 219 At this point we have a linguistic ambiguity: Ais a fixed object (usually called a vector) that may be described in various c oordinate systems. But it is also customary to call the collection of coefficients Aia vector (more specifically, a contravariant vector), while we have already called "ia covariant basis vector. The important thing to observe here is thatAis a fixed object that is not changed by our transformations, while its representation (theAi) and the basis used for the representation (the "i) change in mutually inverse ways (as the coordinate system is changed) so as to keep Afixed. Given our basis vectors, we can compute the displacement (change in position) associ- ated with changes in the qi. Because the basis vectors depend on position, our computation needs to be for small (infinitesimal) displacements ds. We have .ds/2DX i j."idqi/."jdqj/; which, using t he summation c onvention, can be written .ds/2Dgi jdqidqj; (4.39) with gi jD"i"j: (4.40) Since.ds/2is an invariant under rotational (and reflection) transformations, it is a scalar, and the quotient rule permits us to identify gi jas a covariant tensor. Because of its role in defining displacement, gi jis called the covariant metric tensor. Note that the basis vectors can be defined by their Cartesian components, but they are, in general, neither unit vectors nor mutually orthogonal. Because they are often notunit vec- tors we have identified them by the symbol ", notOe. The lack of both a normalization and an orthogonality requirement means that gi j, though manifestly symmetric, is not required to be diagonal, and its elements (including those on the diagonal) may be of either sign. It is convenient to define a contravariant metric tensor that satisfies gikgkjDgjkgkiDi j; (4.41) and is therefore the inverse of the covariant metric tensor. We will use gi jandgi jto make conversions between contravariant and covariant vectors that we then regard as related. Thus, we write gi jFjDFiand gi jFjDFi: (4.42) Returning now to Eq. (4.38), we can manipulate it as follows: ADAi"iDAik i"kD Aigi j gjk"k DAj"j; (4.43) showing that the same vector can be represented either by contravariant or covariant com- ponents, with the two sets of components related by the transformation in Eq. (4.42). ArfKen_Ch04-9780123846549.tex 220 Chapter 4 Tensors and Differential Forms Covariant and Contravariant Bases We now define the contravariant basis vectors "iD@qi @xOexC@qi @yOeyC@qi @zOez; (4.44) giving them this name in anticipation of the fact that we can prove them to be the con- travariant versions of the "i. Our first step in this direction is to verify that "i"jD@qi @x@x @qjC@qi @y@y @qjC@qi @z@z @qjDi j; (4.45) a consequence of the chain rule and the fact that qiandqjare independent variables. We next note that ."i"j/."j"k/Di k; (4.46) also proved using the chain rule; the terms can be collected so that groups of them corre- spond to the identities in Eq. (4.45). Equation (4.46) shows that gi jD"i"j: (4.47) Multiplying both sides of Eq. (4.47) on the right by "jand performing the implied sum- mation, the left-hand side of that equation, gi j"j, becomes the formula for "i, while the right-hand side simplifies to the expression in Eq. (4.44), thereby proving that the con- travariant vector in that equation was appropriately named. We illustrate now some typical metric tensors and basis vectors in both covariant and contravariant form. Example 4.3.1 SOME METRIC TENSORS In spherical polar coordinates, .q1;q2;q3/.r;;'/ , and xDrsincos',yDrsinsin', zDrcos. The covariant basis vectors are "rDsincos'OexCsinsin'OeyCcosOez; "Drcoscos'OexCrcossin'OeyrsinOez; "'Drsinsin'OexCrsincos'Oey; and the contravariant basis vectors, which can be obtained in many ways, one of which is to start from r2Dx2Cy2Cz2,cosDz=r,tan'Dy=x, are "rDsincos'OexCsinsin'OeyCcosOez; "Dr1coscos'OexCr1cossin'Oeyr1sinOez; "'Dsin' rsinOexCcos' rsinOey; ArfKen_Ch04-9780123846549.tex 4.3 Tensors in General Coordinates 221 leading to g11D"r"rD1; g22D""Dr2; g33D"'"'Dr2sin2I all other gi jvanish. Combining these to make gi jand taking the inverse (to make gi j), we have .gi j/D0 @1 0 0 0r20 0 0 r2sin21 A; .gi j/D0 @1 0 0 0r20 0 0.rsin/21 A: We can check that we have inverted gi jcorrectly by comparing the expression given for gi jfrom that built directly from "i"j. This check is left for the reader. The Minkowski metric of special relativity has the form .gi j/D.gi j/D0 BB@1 0 0 0 01 0 0 0 01 0 0 0 011 CCA: The motivation for including it in this example is to emphasize that for some met- rics important in physics, distances ds2need not be positive (meaning that dscan be imaginary).  The relation between the covariant and contravariant basis vectors is useful for writing relationships between vectors. Let AandBbe vectors with contravariant representations .Ai/and.Bi/. We may convert the representation of BtoBiDgi jBj, after which the scalar product ABtakes the form ABD.Ai"i/.Bj"j/DAiBj."i"j/DAiBi: (4.48) Another application is in writing the gradient in general coordinates. If a function is given in a general coordinate system .qi/, its gradient r is a vector with Cartesian com- ponents .r /jD@ @qi@qi @xj: (4.49) In vector notation, Eq. (4.49) becomes r D@ @qi"i; (4.50) showing that the covariant representation of r is the set of derivatives @ =@ qi. If we have reason to use a contravariant representation of the gradient, we can convert its com- ponents using Eq. (4.42). ArfKen_Ch04-9780123846549.tex 222 Chapter 4 Tensors and Differential Forms Covariant Derivatives Moving on to the derivatives of a vector, we find that the situation is much more compli- cated because the basis vectors "iare in general not constant, and the derivative will not be a tensor whose components are the derivatives of the vector components. Starting from the transformation rule for a contravariant vector, .V0/iD@xi @qkVk; and differentiating with respect to qj, we get (for each i) @.V0/i @qjD@xi @qk@Vk @qjC@2xi @qj@qkVk; (4.51) which appears to differ from the transformation law for a second-rank tensor because it contains a second derivative. To see what to do next, let’s write Eq. (4.51) as a single vector equation in the xicoor- dinates, which we take to be Cartesian. The result is @V0 @qjD@Vk @qj"kCVk@"k @qj: (4.52) We now recognize that @"k=@qjmust be some vector in the space spanned by the set of all"iand we therefore write @"k @qjD0 jk": (4.53) The quantities 0 jkare known as Christoffel symbols of the second kind (those of the first kind will be encountered shortly). Using the orthogonality property of the ", Eq. (4.45), we can solve Eq. (4.53) by taking its dot product with any "m, reaching 0m jkD"m@"k @qj: (4.54) Moreover, we note that 0m kjD0m jk, which can be demonstrated by writing out the compo- nents of@"k=@qj. Returning now to Eq. (4.52) and inserting Eq. (4.53), we initially get @V0 @qjD@Vk @qj"kCVk0 jk": (4.55) Interchanging the dummy indices kandin the last term of Eq. (4.55), we get the final result @V0 @qjD@Vk @qjCV0k j "k: (4.56) The parenthesized quantity in Eq. (4.56) is known as the covariant derivative ofV, and it has (unfortunately) become standard to identify it by the awkward notation Vk IjD@Vk @qjCV0k j;so@V0 @qjDVk Ij"k: (4.57) ArfKen_Ch04-9780123846549.tex 4.3 Tensors in General Coordinates 223 If we rewrite Eq. (4.56) in the form dV0Dh Vk Ijdqji "k; and take note that dqjis a contravariant vector, while "kis covariant, we see that the covariant derivative, Vk Ijis a mixed second-rank tensor.4However, it is important to realize that although they bristle with indices, neither@Vk=@qjnor0k jhave individually the correct transformation properties to be tensors. It is only the combination in Eq. (4.57) that has the requisite transformational attributes. It can be shown (see Exercise 4.3.6) that the covariant derivative of a covariant vector Viis given by ViIjD@Vi @qjVk0k i j: (4.58) Like Vi Ij,ViIjis a second-rank tensor. The physical importance of the covariant derivative is that it includes the changes in the basis vectors pursuant to a general dqi, and is therefore more appropriate for describing physical phenomena than a formulation that considers only the changes in the coefficients multiplying the basis vectors. Evaluating Christoffel Symbols It may be more convenient to evaluate the Christoffel symbols by relating them to the metric tensor than simply to use Eq. (4.54). As an initial step in this direction, we define the Christoffel symbol of the first kindTi j;kUby Ti j;kUgmk0m i j; (4.59) from which the symmetry Ti j;kUDT ji;kUfollows. Again, this Ti j;kUis not a third-rank tensor. Inserting Eq. (4.54) and applying the index-lowering transformation, Eq. (4.42), we have Ti j;kUDgmk"m@"i @qj D"k@"i @qj: (4.60) Next, we write gi jD"i"jas in Eq. (4.40) and differentiate it, identifying the result with the aid of Eq. (4.60): @gi j @qkD@"i @qk"jC"i@"j @qk DTik;jUCT jk;iU: 4V0does not contribute to the covariant/contravariant character of the equation as its implicit index labels the Cartesian coordi- nates, as is also the case for "k. ArfKen_Ch04-9780123846549.tex 224 Chapter 4 Tensors and Differential Forms We then note that we can combine three of these derivatives with different index sets, with a result that simplifies to give 1 2@gik @qjC@gjk @qi@gi j @qk DTi j;kU: (4.61) We now return to Eq. (4.59), which we solve for 0m i jby multiplying both sides by gnk, summing over k, and using the fact that .g/and.g/are mutually inverse, see Eq. (4.41): 0n i jDX kgnkTi j;kU: (4.62) Finally, substituting for Ti j;kUfrom Eq. (4.61), and once again using the summation con- vention, we get: 0n i jDgnkTi j;kUD1 2gnk@gik @qjC@gjk @qi@gi j @qk : (4.63) The apparatus of this subsection becomes unnecessary in Cartesian coordinates, because the basis vectors have vanishing derivatives, and the covariant and ordinary partial deriva- tives then coincide. Tensor Derivative Operators With covariant differentiation now available, we are ready to derive the vector differential operators in general tensor form. Gradient—We have already discussed it, with the result from Eq. (4.50): r D@ @qi"i: (4.64) Divergence—A vector Vwhose contravariant representation is Vi"ihas divergence rVD"j@.Vi"i/ @qjD"j@Vi @qjCVk0i jk "iD@Vi @qiCVk0i ik: (4.65) Note that the covariant derivative has appeared here. Expressing 0i ikby Eq. (4.63), we have 0i ikD1 2gim@gim @qkC@gkm @qi@gik @qm D1 2gim@gim @qk; (4.66) where we have recognized that the last two terms in the bracket will cancel because by changing the names of their dummy indices they can be identified as identical except in sign. Because.gim/is the matrix inverse to .gim/, we note that the combination of matrix elements on the right-hand side of Eq. (4.66) is similar to those in the formula for the derivative of a determinant, Eq. (2.35); remember that gis symmetric: gimDgmi. In the present notation, the relevant formula is ddet.g/ dqkDdet.g/gim@gim @qk; (4.67) ArfKen_Ch04-9780123846549.tex 4.3 Tensors in General Coordinates 225 where det.g/is the determinant of the covariant metric tensor .g/. Using Eq. (4.67), Eq. (4.66) becomes 0i ikD1 2 det. g/ddet.g/ dqkD1 Tdet.g/[email protected]/U1=2 @qk: (4.68) Combining the result in Eq. (4.68) with Eq. (4.65), we obtain a maximally compact formula for the divergence of a contravariant vector V: rVDVi IiD1 Tdet.g/U1=2@ @qk Tdet.g/U1=2Vk : (4.69) To compare this result with that for an orthogonal coordinate system, Eq. (3.141), note that det.g/D.h1h2h3/2and that the kcomponent of the vector represented by Vin Eq. (3.141) is, in the present notation, equal to Vkj"kjDhkVk(no summation). Laplacian—We can form the Laplacian r2 by inserting an expression for the gradi- entr into the formula for the divergence, Eq. (4.69). However, that equation uses the contravariant coefficients Vk, so we must describe the gradient in its contravariant rep- resentation. Since Eq. (4.64) shows that the covariant coefficients of the gradient are the derivatives@ =@ qi, its contravariant coefficients have to be gki@ @qi: Insertion into Eq. (4.69) then yields r2 D1 Tdet.g/U1=2@ @qk Tdet.g/U1=2gki@ @qi : (4.70) Fororthogonal systems the metric tensor is diagonal and the contravariant giiD.hi/2 (no summation). Equation (4.70) then reduces to rr D1 h1h2h3@ @qi h1h2h3 h2 i@ @qi! ; in agreement with Eq. (3.142). Curl—The difference of derivatives that appears in the curl has components that can be written @Vi @qj@Vj @qiD@Vi @qjVk0k i j@Vj @qiCVk0k jiDViIjVjIi; (4.71) where we used the symmetry of the Christoffel symbols to obtain a cancellation. The rea- son for the manipulation in Eq. (4.71) is to bring all the terms on its right-hand side to tensor form. In using Eq. (4.71), it is necessary to remember that the quantities Viare coef- ficients of the possibly nonunit "iand are therefore notcomponents of Vin the orthonor- mal basisOei. ArfKen_Ch04-9780123846549.tex 226 Chapter 4 Tensors and Differential Forms Exercises 4.3.1 For the special case of 3-D space ( "1,"2,"3defining a right-handed coordinate system, not necessarily orthogonal), show that "iD"j"k "j"k"i;i;j;kD1, 2, 3 and cyclic permutations. Note. These contravariant basis vectors "idefine the reciprocal lattice space of Example 3.2.1. 4.3.2 If the covariant vectors "iare orthogonal, show that (a) gi jis diagonal, (b) giiD1=gii(no summation), (c)j"ijD1=j" ij. 4.3.3 Prove that."i"j/."j"k/Di k. 4.3.4 Show that0m jkD0m kj. 4.3.5 Derive the covariant and contravariant metric tensors for circular cylindrical coordi- nates. 4.3.6 Show that the covariant derivative of a covariant vector is given by ViIj@Vi @qjVk0k i j: Hint. Differentiate "i"jDi j: 4.3.7 Verify that ViIjDgikVk Ijby showing that @Vi @qjVk0k i jDgik@Vk @qjCVm0k mj : 4.3.8 From the circular cylindrical metric tensor gi j, calculate the 0k i jfor circular cylindrical coordinates. Note. There are only three nonvanishing 0. 4.3.9 Using the0k i jfrom Exercise 4.3.8, write out the covariant derivatives Vi Ijof a vector V in circular cylindrical coordinates. 4.3.10 Show that for the metric tensor gi jIkDgi j IkD0. 4.3.11 Starting with the divergence in tensor notation, Eq. (4.70), develop the divergence of a vector in spherical polar coordinates, Eq. (3.157). 4.3.12 The covariant vector Aiis the gradient of a scalar. Show that the difference of covariant derivatives AiIjAjIivanishes. ArfKen_Ch04-9780123846549.tex 4.4 Jacobians 227 4.4 J ACOBIANS In the preceding chapters we have considered the use of curvilinear coordinates, but have not placed much focus on transformations between coordinate systems, and in particular on the way in which multidimensional integrals must transform when the coordinate sys- tem is changed. To provide formulas that will be useful in spaces with arbitrary numbers of dimensions, and with transformations involving coordinate systems that are not orthog- onal, we now return to the notion of the Jacobian, introduced but not fully developed in Chapter 1. As already mentioned in Chapter 1, changes of variables in multiple integrations, say from variables x1;x2;:::tou1;u2;:::requires that we replace the differential dx1dx2::: with J du 1du2:::, where J, called the Jacobian, is the quantity (usually dependent on the variables) needed to make these expressions mutually consistent. More specifically, we identify dDJ du 1du2:::as the “volume” of a region of width du1inu1,du2in u2, . . . , where the “volume” is to be computed in the x1,x2, . . . space, treated as Cartesian coordinates. To obtain a formula for Jwe start by identifying the displacement (in the Cartesian system defined by the xi) that corresponds to a change in each variable ui. Letting ds.ui/ be that displacement (which is a vector), we can decompose it into Cartesian components as follows: ds.u1/D@x1 @u1 Oe1C@x2 @u1 Oe2C du1; ds.u2/D@x1 @u2 Oe1C@x2 @u2 Oe2C du2; (4.72) ds.u3/D@x1 @u3 Oe1C@x2 @u3 Oe2C du3; ::::::D ::::::::: The partial derivatives .@xi=@uj/in Eq. (4.72) must be understood to be evaluated with the other ukheld constant. It would clutter the formula an unreasonable amount to indicate this explicitly. If we had only two variables, u1andu2, the differential area would simply be jds.u1/j times the component of ds.u2/that is perpendicular to ds.u1/. If there were a third vari- able, u3, we would further multiply by the component of ds.u3/that was perpendicular to both ds.u1/andds.u2/. Extension to arbitrary numbers of dimensions is obvious. What is less obvious is an explicit formula for the “volume” for an arbitrary number of dimensions. Let’s start by writing Eq. (4.72) in matrix form: 0 [email protected]/ du1 ds.u2/ du2 ds.u3/ du3 :::1 CCCCCCCCCAD0 BBBBBBBBB@@x1 @u1@x2 @u1@x3 @u1 @x1 @u2@x2 @u2@x3 @u2 @x1 @u3@x2 @u3@x3 @u3 ::: ::: ::: :::1 CCCCCCCCCA0 BB@Oe1 Oe2 Oe3 :::1 CCA: (4.73) ArfKen_Ch04-9780123846549.tex 228 Chapter 4 Tensors and Differential Forms We now proceed to make changes to the second and succeeding rows of the square matrix in Eq. (4.73) that may destroy the relation to the ds.ui/=du i, but which will leave the “volume” unchanged. In particular, we subtract from the second row of the derivative matrix that multiple of the first row which will cause the first element of the modified second row to vanish. This will not change the “volume” because it modifies ds.u2/=du 2 by adding or subtracting a vector in the ds.u1/=du 1direction, and therefore does not affect the component of ds.u2/=du 2perpendicular to ds.u1/=du 1. See Fig. 4.1. The alert reader will recall that this modification of the second row of our matrix is an operation that was used when evaluating determinants, and was there justified because it did not change the value of the determinant. We have a similar situation here; the operation will not change the value of the differential “volume” because we are changing only the component of ds.u2/=du 2that is in the ds.u1/=du 1direction. In a similar fashion, we can carry out further operations of the same kind that will lead to a matrix in which all the elements below the principal diagonal have been reduced to zero. The situation at this point is indicated schematically for an 4-D space as the transition from the first to the second matrix in Fig. 4.2. These modified ds.ui/=du iwill lead to the same differential volume as the original ds.ui/=du i. This modified matrix will no longer provide a faithful representation of the differential region in the uispace, but that is irrelevant since our only objective is to evaluate the differential “volume.” We next take the final ( nth) row of our modified matrix, which will be entirely zero except for its last element, and subtract a suitable multiple of it from all the other rows to introduce zeros in the last element of every row above the principal diagonal. These operations correspond to changes in which we modify only the components of the other ds.ui/=du ithat are in the direction of ds.un/, and therefore will not change the differential “volume.” Then, using the next-to-last row (which now has only a diagonal element), we can in a similar fashion introduce zeros in the next-to-last column of all the preceding rows. Continuing this process, we will ultimately have a set of modified ds.ui/=du ithat will have the structure shown as the last matrix in Fig. 4.2. Because our modified matrix is diagonal, with each nonzero element associated with a single different Oei, the “volume” is u1u2ku1 h u1u2h FIGURE 4.1 Area remains unchanged when vector proportional to u1is added to u2. 0 BB@a11a12a13a14 a21a22a23a24 a31a32a33a34 a41a42a43a441 CCA!0 BB@a11a12a13a14 0b22b23b24 0 0 b33b34 0 0 0 b441 CCA!0 BB@a11 0 0 0 0b22 0 0 0 0 b33 0 0 0 0 b441 CCA FIGURE 4.2 Manipulation of Jacobian matrix. Here ai jD.@xj=@ui/, and bi jare formed by combining rows (see text). ArfKen_Ch04-9780123846549.tex 4.4 Jacobians 229 then easily computed as the product of the diagonal elements. This product of the diagonal elements of a diagonal matrix is an evaluation of its determinant. Reviewing what we have done, we see that we have identified the differential “volume” as a quantity which is equal to the determinant of the original derivative set. This must be so, because we obtained our final result by carrying out operations each of which leaves a determinant unchanged. The final result can be expressed as the well-known formula for the Jacobian: dDJ du 1du2:::; JD @x1 @u1@x2 @u1@x3 @u1 @x1 @u2@x2 @u2@x3 @u2 @x1 @u3@x2 @u3@x3 @u3 ::: ::: ::: ::: @.x1;x2;:::/ @.u1;u2;:::/: (4.74) The standard notation for the Jacobian, shown as the last member of Eq. (4.74), is a conve- nient reminder of the way in which the partial derivatives appear in it. Note also that when the standard notation for Jis inserted in the expression for d, the overall expression has du1du2:::in the numerator, while @.u1;u2;:::/ appears in the denominator. This feature can help the user to make a proper identification of the Jacobian. A few words about nomenclature: The matrix in Eq. (4.73) is sometimes called the Jacobian matrix, with the determinant in Eq. (4.74) then distinguished by calling it theJacobian determinant. Unless within a discussion in which both these quantities appear and need to be separately identified, most authors simply call J, the determinant in Eq. (4.74), the Jacobian. That is the usage we follow in this book. We close with one final observation. Since Jis a determinant, it will have a sign that depends on the order in which the xianduiare specified. This ambiguity corresponds to our freedom to choose either right- or left-handed coordinates. In typical applications involving a Jacobian, it is usual to take its absolute value and to choose the ranges of the individual uiintegrals in a way that gives the correct sign for the overall integral. Example 4.4.1 2-D and 3-D JACOBIANS In two dimensions, with Cartesian coordinates x;yand transformed coordinates u;v, the element of area d Ahas, following Eq. (4.74), the form d ADdu dv@x @u@y @v @x @v@y @u : This is the expected result, as the quantity in square brackets is the formula for the z component of the cross product of the two vectors @x @u OexC@y @u Oeyand@x @v OexC@y @v Oey; and it is well known that the magnitude of the cross product of two vectors is a measure of the area of the parallelogram with sides formed by the vectors. ArfKen_Ch04-9780123846549.tex 230 Chapter 4 Tensors and Differential Forms In three dimensions, the determinant in the Jacobian corresponds exactly with the for- mula for the scalar triple product, Eq. (3.12). Letting Ax,Ay,Azin that formula refer to the derivatives.@x=@u/,.@y=@u/,.@z=@u/, with the components of BandCsimilarly related to derivatives with respect to vandw, we recover the formula for the volume within the parallelepiped defined by three vectors.  Inverse of Jacobian Since the xiand the uiare arbitrary sets of coordinates, we could have carried out the entire analysis of the preceding subsection regarding the uias the fundamental coordinate system, with the xias coordinates reached by a change of variables. In that case, our Jacobian (which we choose to label J1), would be J[email protected];u2;:::/ @.x1;x2;:::/: (4.75) It is clear that if dx1dx2:::DJ du 1du2:::, then it must also be true that du1du2:::D .1=J/dx1dx2:::. Let’s verify that the quantity we have called J1is in fact 1=J. Let’s represent the two Jacobian matrices involved here as AD0 BBBBBBBBB@@x1 @u1@x2 @u1@x3 @u1 @x1 @u2@x2 @u2@x3 @u2 @x1 @u3@x2 @u3@x3 @u3 ::: ::: ::: :::1 CCCCCCCCCA;BD0 BBBBBBBBB@@u1 @x1@u2 @x1@u3 @x1 @u1 @x2@u2 @x2@u3 @x2 @u1 @x3@u2 @x3@u3 @x3 ::: ::: ::: :::1 CCCCCCCCCA: We then have JDdet.A/ and J1Ddet.B/ . We would like to show that J J1D det.A/ det.B/D1. The proof is fairly simple if we use the determinant product theorem. Thus, we write det.A/ det.B/Ddet.AB/; and now all we need show is that the matrix product ABis a unit matrix. Carrying out the matrix multiplication, we find, as a result of the chain rule, .AB/ i jDX k@xk @ui@uj @xk D@uj @ui Di j; (4.76) verifying that ABis indeed a unit matrix. The relation between the Jacobian and its inverse is of practical interest. It may turn out that the derivatives @ui=@xjare easier to compute than @xi=@uj, making it convenient to obtain Jby first constructing and evaluating the determinant for J1. ArfKen_Ch04-9780123846549.tex 4.4 Jacobians 231 Example 4.4.2 DIRECT AND INVERSE APPROACHES TO JACOBIAN Suppose we need the [email protected];;'/ @.x;y;z/, where x,y, and zare Cartesian coordinates and r,,'are spherical polar coordinates. Using Eq. (4.74) and the relations rDq x2Cy2Cz2; Dcos1 zp x2Cy2Cz2! ; 'Dtan1y x ; we find after significant effort (letting 2Dx2Cy2), [email protected];;'/ @.x;y;z/D x ry rz r xz r2yz r2 r2 y 2x 20 D1 rD1 r2sin: It is much less effort to use the relations xDrsincos'; yDrsinsin'; zDrcos; and then to evaluate (easily), J[email protected];y;z/ @.r;;'/D sincos'sinsin'cos rcoscos'rcossin'rsin rsinsin'rsincos' 0 Dr2sin: We finish by writing JD1=J1D1=r2sin.  Exercises 4.4.1 Assuming the functions uandvto be differentiable, (a) Show that a necessary and sufficient condition that u.x;y;z/andv.x;y;z/are related by some function f.u;v/D0is that.ru/.rv/D0; (b) If uDu.x;y/andvDv.x;y/, show that the condition .ru/.rv/D0leads to the 2-D Jacobian [email protected];v/ @.x;y/D @u @x@u @y @v @x@v @y D0: 4.4.2 A 2-D orthogonal system is described by the coordinates q1andq2. Show that the Jacobian Jsatisfies the equation J@.x;y/ @.q1;q2/@x @q1@y @q2@x @q2@y @q1Dh1h2: Hint. It’s easier to work with the square of each side of this equation. ArfKen_Ch04-9780123846549.tex 232 Chapter 4 Tensors and Differential Forms 4.4.3 For the transformation uDxCy,vDx=y, with x0andy0, find the Jacobian @.x;y/ @.u;v/ (a) By direct computation, (b) By first computing J1. 4.5 D IFFERENTIAL FORMS Our study of tensors has indicated that significant complications arise when we leave Carte- sian coordinate systems, even in traditional contexts such as the introduction of spherical or cylindrical coordinates. Much of the difficulty arises from the fact that the metric (as expressed in a coordinate system) becomes position-dependent, and that the lines or sur- faces of constant coordinate values become curved. Many of the most vexing problems can be avoided if we work in a geometry that deals with infinitesimal displacements, because the situations of most importance in physics then become locally similar to the simpler and more familiar conditions based on Cartesian coordinates. The calculus of differential forms, of which the leading developer was Elie Cartan, has become recognized as a natural and very powerful tool for the treatment of curved coordi- nates, both in classical settings and in contemporary studies of curved space-time. Cartan’s calculus leads to a remarkable unification of concepts and theorems of vector analysis that is worth pursuing, with the result that in differential geometry and in theoretical physics the use of differential forms is now widespread. Differential forms provide an important entry to the role of geometry in physics, and the connectivity of the spaces under discussion (technically, referred to as their topology) has physical implications. Illustrations are provided already by situations as simple as the fact that a coordinate defined on a circle cannot be single-valued and continuous at all angles. More sophisticated consequences of topology in physics, largely beyond the scope of the present text, include gauge transformations, flux quantization, the Bohm-Aharanov effect, emerging theories of elementary particles, and phenomena of general relativity. Introduction For simplicity we begin our discussion of differential forms in a notation appropriate for ordinary 3-D space, though the real power of the methods under study is that they are not limited either by the dimensionality of the space or by its metric properties (and are therefore also relevant to the curved space-time of general relativity). The basic quantities under consideration are the differentials dx,dy,dz(identified with linearly independent directions in the space), linear combinations thereof, and more complicated quantities built from these by combination rules we will shortly discuss in detail. Taking for example dx, it is essential to understand that in our current context it is not just an infinitesimal number describing a change in the xcoordinate, but is to be viewed as a mathematical object with certain operational properties (which, admittedly, may include its eventual use in contexts such as the evaluation of line, surface, or volume integrals). The rules by which dxand ArfKen_Ch04-9780123846549.tex 4.5 Differential Forms 233 related quantities can be manipulated have been designed to permit expressions such as !DA.x;y;z/dxCB.x;y;z/dyCC.x;y;z/dz; (4.77) which are called 1-forms, to be related to quantities that occur as the integrands of line integrals, to permit expressions of the type !DF.x;y;z/dx^dyCG.x;y;z/dx^dzCH.x;y;z/dy^dz; (4.78) which are called 2-forms, to be related to the integrands of surface integrals, and to permit expressions like !DK.x;y;z/dx^dy^dz; (4.79) known as 3-forms, to be related to the integrands of volume integrals. The^symbol (called “wedge”) indicates that the individual differentials are to be com- bined to form more complicated objects using the rules of exterior algebra (sometimes called Grassmann algebra), so more is being implied by Eqs. (4.77) to(4.79) than the somewhat similar formulas that might appear in the conventional notation for various kinds of integrals. To maintain contact with other presentations on differential forms, we note that some authors omit the wedge symbol, thereby assuming that the reader knows that the differentials are to be combined according to the rules of exterior algebra. In order to minimize potential confusion, we will continue to write the wedge symbol for these combinations of differentials (which are called exterior, or wedge products). To write differential forms in ways that do not presuppose the dimension of the under- lying space, we sometimes write the differentials as dxi, designating a form as a p-form if it contains pfactors dxi. Ordinary functions (containing no dxi) can be identified as 0-forms. The mathematics of differential forms was developed with the aim of systematizing the application of calculus to differentiable manifolds, loosely defined as sets of points that can be identified by coordinates that locally vary “smoothly” (meaning that they are differ- entiable to whatever degree is needed for analysis).5We are presently focusing attention on the differentials that appear in the forms; one could also consider the behavior of the coefficients. For example, when we write the 1-form !DAxdxCAydyCAzdz; Ax,Ay,Azwill behave under a coordinate transformation like the components of a vec- tor, and in the older differential-forms literature the differentials and the coefficients were referred to as contravariant and covariant vector components, since these two sets of quan- tities must transform in mutually inverse ways under rotations of the coordinate system. What is relevant for us at this point is that relationships we develop for differential forms can be translated into related relationships for their vector coefficients, yielding not only various well-known formulas of vector analysis but also showing how they can be gener- alized to spaces of higher dimension. 5A manifold defined on a circle or sphere must have a coordinate that cannot be globally smooth (in the usual coordinate systems it will jump somewhere by 2). This and related issues connect topology and physics, and are for the most part outside the scope of this text. ArfKen_Ch04-9780123846549.tex 234 Chapter 4 Tensors and Differential Forms Exterior Algebra The central idea in exterior algebra is that the operations are designed to create permuta- tional antisymmetry. Assuming the basis 1-forms are dxi, that!jare arbitrary p-forms (of respective orders pj), and that aandbare ordinary numbers or functions, the wedge product is defined to have the properties .a!1Cb!2/^!3Da!1^!3Cb!2^!3.p1Dp2/; .!1^!2/^!3D!1^.!2^!3/;a.!1^!2/D.a!1/^!2; (4.80) dxi^dxjDdx j^dxi: We thus have the usual associative and distributive laws, and each term of an arbitrary differential form can be reduced to a coefficient multiplying a dxior a wedge product of the generic form dxi^dxj^^ dxp: Moreover, the properties in Eq. (4.80) permit all the coefficient functions to be collected at the beginning of a form. For example, a dx 1^b dx 2Da.b dx 2^dx1/Dab.dx2^dx1/Dab.dx1^dx2/: We therefore generally do not need parentheses to indicate the order in which products are to be carried out. We can use the last of Eqs. (4.80) to bring the index set into any desired order. If any two of the dxiare the same, the expression will vanish because dxi^dxiDdx i^dxiD0; otherwise, the ordered-index form will have a sign determined by the parity of the index permutation needed to obtain the ordering. It is nota coincidence that this is the sign rule for the terms of a determinant, compare Eq. (2.10). Letting "Pstand for the Levi-Civita symbol for the permutation to ascending index order, an arbitrary wedge product of dxi can, for example, be brought to the form "Pdxh1^dxh2^^ dxhp;1h1<h2<<hp: If any of the dxiin a differential form is linearly dependent on the others, then its expansion into linearly independent terms will produce a duplicated dxjand cause the form to vanish. Since the number of linearly independent dxjcannot be larger than the dimension of the underlying space, we see that in a space of dimension dwe only need to consider p-forms with pd. Thus, in 3-D space, only up through 3-forms are relevant; for Minkowski space .ct;x;y;z/, we will also have 4-forms. Example 4.5.1 SIMPLIFYING DIFFERENTIAL FORMS Consider the wedge product !D.3dxC4dydz/^.dxdyC2dz/D3dx^dx3dx^dyC6dx^dz C4dy^dx4dy^dyC8dy^dzdz^dxCdz^dy2dz^dz: ArfKen_Ch04-9780123846549.tex 4.5 Differential Forms 235 The terms with duplicate differentials, e.g., dx^dx, vanish, and products that differ only in the order of the 1-forms can be combined, changing the sign of the product when we interchange its factors. We get !D7 dx^dyC7dx^dzC7dy^dzD7.dy^dzdz^dxdx^dy/: We will shortly see that in three dimensions there are some advantages to bringing the 1-forms into cyclic order (rather than ascending or descending order) in the wedge prod- ucts, and we did so in the final simplification of !.  The antisymmetry built into the exterior algebra has an important purpose: It causes p-forms to depend on the differentials in ways appropriate (in three dimensions) for the description of elements of length, area, and volume, in part because the fact that dxi^dxiD0prevents the appearance of duplicated differentials. In particular, 1-forms can be associated with elements of length, 2-forms with area, and 3-forms with volume. This feature carries forward to spaces of arbitrary dimensionality, thereby resolving poten- tially difficult questions that would otherwise have to be handled on a case-by-case basis. In fact, one of the virtues of the differential-forms approach is that there now exists a con- siderable body of general mathematical results that is pretty much completely absent from tensor analysis. For example, we will shortly find that the rules for differentiation in the exterior algebra cause the derivative of a p-form to be a .pC1/-form, thereby avoiding a pitfall that arises in tensor calculus: When the transformation coefficients are position- dependent, simply differentiating the coefficients representing a tensor of rank pdoes not yield another tensor. As we have seen, this dilemma is resolved in tensor analysis by intro- ducing the notion of covariant derivative. Another consequence of the antisymmetry is that lengths, areas, volumes, and (at higher dimensionality) hypervolumes are oriented (meaning that they have signs that depend on the way the p-forms defining them are writ- ten), and the orientation must be taken into account when making computations based on differential forms. Complementary Differential Forms Associated with each differential form is a complementary (or dual) form that contains the differentials notincluded in the original form. Thus, if our underlying space has dimension d, the form dual to a p-form will be a .dp/-form. In three dimensions, the complement to a 1-form will be a 2-form (and vice versa), while the complement to a 3-form will be a 0-form (a scalar). It is useful to work with these complementary forms, and this is done by introducing an operator known as the Hodge operator; it is usually designated nota- tionally as an asterisk (preceding the quantity to which it is applied, not as a superscript), and is therefore also referred to either as the Hodge star operator or simply as the star operator. Formally, its definition requires the introduction of a metric and the selection of an orientation (chosen by specifying the standard order of the differentials comprising the 1-form basis), and if the 1-form basis is not orthogonal there result complications we shall not discuss. For orthogonal bases, the dual forms depend on the index positions of the factors and on the metric tensor.6 6In the current discussion, restricted to Euclidean and Minkowski metrics, the metric tensor is diagonal, with diagonal elements 1, and the relevant quantities are the signs of the diagonal elements. ArfKen_Ch04-9780123846549.tex 236 Chapter 4 Tensors and Differential Forms To find!, where!is ap-form, we start by writing the wedge product !0of all mem- bers of the 1-form basis not represented in !, with the sign corresponding to the permuta- tion that is needed to bring the index set (indices of!) followed by (indices of !0) to standard order. Then !consists of!0(with the sign we just found), but also multi- plied by.1/, whereis the number of differentials in !0whose metric-tensor diagonal element is1. ForR3, ordinary 3-D space, the metric tensor is a unit matrix, so this final multiplication can be omitted, but it becomes relevant for our other case of current interest, the Minkowski metric. For Euclidean 3-D space, we have 1Ddx1^dx2^dx3; dx 1Ddx2^dx3;dx 2Ddx3^dx1;dx 3Ddx1^dx2; (4.81) .dx 1^dx2/Ddx3;.dx 3^dx1/Ddx2;.dx 2^dx3/Ddx1; .dx 1^dx2^dx3/D1: Cases not shown above are linearly dependent on those that were shown and can be obtained by permuting the differentials in the above formulas and taking the resulting sign changes into account. At this point, two observations are in order. First, note that by writing the indices 1, 2, 3 in cyclic order, we have caused all the starred quantities to have positive signs. This choice makes the symmetry more evident. Second, it can be seen that all the formulas inEq. (4.81) are consistent with.!/D!. However, this is not universally true; com- pare with the formulas for Minkowski space, which are in the example we next consider. See also Exercise 4.5.1. Example 4.5.2 HODGE OPERATOR IN MINKOWSKI SPACE Taking the oriented 1-form basis .dt;dx1;dx2;dx3/, and the metric tensor 0 BB@1 0 0 0 01 0 0 0 01 0 0 0 011 CCA; let’s determine the effect of the Hodge operator on the various possible differential forms. Consider initially *1, for which the complementary form contains dt^dx1^dx2^dx3. Since we took these differentials in the basis order, they are assigned a plus sign. Since !D 1contains no differentials, its number of negative metric-tensor diagonal elements is zero, so.1/D.1/0D1and there is no sign change arising from the metric. Therefore, 1Ddt^dx1^dx2^dx3: Next, take.dt^dx1^dx2^dx3/. The complementary form is just unity, with no sign change due to the index ordering, as the differentials are already in standard order. ArfKen_Ch04-9780123846549.tex 4.5 Differential Forms 237 However, this time we have three entries in the quantity being starred with negative metric- tensor diagonal elements; this generates .1/3D1 , so .dt^dx1^dx2^dx3/D1: Moving next todx 1, the complementary form is dt^dx2^dx3, and the index order- ing (based on dx1;dt;dx2;dx3) requires one pair interchange to reach the standard order (thereby yielding a minus sign). But the quantity being starred contains one differential that generates a minus sign, namely dx1, so dx 1Ddt^dx2^dx3: Looking explicitly at one more case, consider .dt^dx1/, for which the complementary form is dx2^dx3. This time the indices are in standard order, but the dx1being starred generates a minus sign, so .dt^dx1/Ddx 2^dx3: Development of the remaining possibilities is left to Exercise 4.5.1; the results are summa- rized below, where i;j;kdenotes any cyclic permutation of 1,2,3. 1Ddt^dx1^dx2^dx3; dx iDdt^dxj^dxk;dtDdx1^dx2^dx3; .dx j^dxk/Ddt^dxi;.dt^dxi/Ddx j^dxk; (4.82) .dx 1^dx2^dx3/Ddt;.dt^dxi^dxj/Ddxk; .dt^dx1^dx2^dx3/D1: Note that all the starred forms in Eq. (4.82) with an even number of differentials have the property that.!/D! , confirming our earlier statement that complementing twice does not always restore the original form with its original sign.  We now consider some examples illustrating the utility of the star operator. Example 4.5.3 MISCELLANEOUS DIFFERENTIAL FORMS In the Euclidean space R3, consider the wedge product A^Bof the two 1-forms AD AxdxCAydyCAzdzandBDBxdxCBydyCBzdz. Simplifying using the rules for exterior products, A^BD.AyBzAzBy/dy^dzC.AzBxAxBz/dz^dxC.AxByAyBx/dx^dy: If we now apply the star operator and use the formulas in Eq. (4.81) we get .A^B/D.AyBzAzBy/dxC.AzBxAxBz/dyC.AxByAyBx/dz; showing that in R3,.A^B/forms an expression that is analogous to the cross product ABof vectors AxOexCAyOeyCAzOezandBxOexCByOeyCBzOez. In fact, we can write .A^B/D.AB/xdxC.AB/ydyC.AB/zdz: (4.83) ArfKen_Ch04-9780123846549.tex 238 Chapter 4 Tensors and Differential Forms Note that the sign of .A^B/is determined by our implicit choice that the standard ordering of the basis differentials is .dx;dy;dz/. Next, consider the exterior product A^B^C, where Cis a 1-form with coefficients Cx;Cy;Cz. Applying the evaluation rules, we find that every surviving term in the product is proportional to dx^dy^dz, and we obtain A^B^CD.AxByCzAxBzCyAyBxCz CAyBzCxCAzBxCyAzByCx/dx^dy^dz; which we recognize can be written in the form A^B^CD AxAyAz BxByBz CxCyCz dx^dy^dz: Applying now the star operator, we reach .A^B^C/D AxAyAz BxByBz CxCyCz DA.BC/: (4.84) Not only were the results in Eqs. (4.83) and(4.84) easily obtained, they also generalize nicely to spaces of arbitrary dimension and metric, while the traditional vector notation, which uses the cross product, is applicable only to R3.  Exercises 4.5.1 Using the rules for the application of the Hodge star operator, verify the results given in Eq. (4.82) for its application to all linearly independent differential forms in Minkowski space. 4.5.2 If the force field is constant and moving a particle from the origin to .3;0;0/requires a units of work, from .1;1;0/to.1; 1;0/takes bunits of work, and from .0;0;4/ to.0;0;5/cunits of work, find the 1-form of the work. 4.6 D IFFERENTIATING FORMS Exterior Derivatives Having introduced differential forms and their exterior algebra, we next develop their prop- erties under differentiation. To accomplish this, we define the exterior derivative, which we consider to be an operator identified by the traditional symbol d. We have, in fact, already introduced that operator when we wrote dxi, stating at the time that we intended to interpret dxias a mathematical object with specified properties and not just as a small ArfKen_Ch04-9780123846549.tex 4.6 Differentiating Forms 239 change in xi. We are now refining that statement to interpret dxias the result of applying the operator dto the quantity xi. We complete our definition of the operator dby requir- ing it to have the following properties, where !is ap-form,!0is ap0-form, and fis an ordinary function (a 0-form): d.!C!0/Dd!Cd!0.pDp0/; d.f!/D.d f/^!Cf d!; d.!^!0/Dd!^!0C.1/p!^d!0; (4.85) d.d!/D0; d fDX j@f @xjdxj; where the sum over jspans the underlying space. The formula for the derivative of the wedge product is sometime called by mathematicians an antiderivation, referring to the fact that when applied to the right-hand factor an antisymmetry-motivated minus sign appears. Example 4.6.1 EXTERIOR DERIVATIVE Equations (4.85) areaxioms, so they are not subject to proof, though they arerequired to be consistent. It is of interest to verify that the sign for the derivative of the second term in a wedge product is needed. Taking !and!0to be monomials, we first bring their coefficients to the left and then apply the differentiation operator (which, irrespective of the choice of sign, gives zero when applied to any of the differentials). Thus, d.!^!0/Dd.AB/h dx1^^ dxpi ^h dx1^^ dxp0i DX @A @xBCA@B @x dx^h dx1^^ dxpi ^h dx1^^ dxp0i : On expanding the sum, the first term is clearly d!^!0; to make the second term look like !^d!0, it is necessary to permute dxthrough the pdifferentials in !, yielding the sign factor.1/p. Extension to general polynomial forms is trivial. One might also ask whether the fourth of the above axioms, d.d!/D0, sometimes referred to as Poincaré’s lemma, is necessary or consistent with the others. First, it pro- vides new information, as otherwise we have no way of reducing d.dxi/. Next, to see why the axiom set is consistent, we illustrate by examining (in R2) d fD@f @xdxC@f @ydy; ArfKen_Ch04-9780123846549.tex 240 Chapter 4 Tensors and Differential Forms from which we form d.d f/D@ @x@f @x dx^dxC@ @y@f @x dy^dx C@ @x@f @y dx^dyC@ @y@f @y dy^dyD0: We obtain the zero result because of the antisymmetry of the wedge product and because the mixed second derivatives are equal. We see that the central reason for the validity of Poincaré’s lemma is that the mixed derivatives of a sufficiently differentiable function are invariant with respect to the order in which the differentiations are carried out.  To catalog the possibilities for the action of the doperator in ordinary 3-D space, we first note that the derivative of an ordinary function (a 0-form) is d fD@f @xdxC@f @ydyC@f @zdzD.rf/xdxC.rf/ydyC.rf/zdz: (4.86) We next differentiate the 1-form !DAxdxCAydyCAzdz. After simplification, d!D@Az @y@Ay @z dy^dzC@Ax @z@Az @x dz^dxC@Ay @x@Ax @y dx^dy: We recognize this as d.AxdxCAydyCAzdz/D .rA/xdy^dzC.rA/ydz^dxC.rA/zdx^dy;(4.87) which is equivalent to d AxdxCAydyCAzdz D.rA/xdxC.rA/ydyC.rA/zdz:(4.88) Finally we differentiate the 2-form Bxdy^dzCBydz^dxCBzdx^dy, obtaining the three-form d Bxdy^dzCBydz^dxCBzdx^dy D@Bx @xC@By @yC@Bz @z dx^dy^dz; equivalent to d Bxdy^dzCBydz^dxCBzdx^dy D.rB/dx^dy^dz (4.89) and d Bxdy^dzCBydz^dxCBzdx^dy DrB: (4.90) We see that application of the doperator directly generates all the differential operators of traditional vector analysis. ArfKen_Ch04-9780123846549.tex 4.6 Differentiating Forms 241 If now we return to Eq. (4.87) and take the 1-form on its left-hand side to be d f, so that ADrf, we have, inserting Eq. (4.86), d.d f/D r.rf/ xdy^dzC r.rf/ ydz^dxC r.rf/ zdx^dyD0: (4.91) We have invoked Poincaré’s lemma to set this expression to zero. The result is equivalent to the well-known identity r.rf/D0. Another identity is obtained if we start from Eq. (4.89) and take the 2-form on its left- hand side to be d.AxdxCAydyCAzdz/. Then, with the aid of Eq. (4.88), we have d d.AxdxCAydyCAzdz/ Dr.rA/dx^dy^dzD0; (4.92) where once again the zero result follows from Poincaré’s lemma and we have established the well-known formula r.rA/D0. Part of the importance of the derivation of these formulas using differential-forms methods is that these are merely the first members of hierarchies of identities that can be derived for spaces with higher numbers of dimensions and with different metric properties. Example 4.6.2 MAXWELL’S EQUATIONS Maxwell’s equations of electromagnetic theory can be written in an extremely compact and elegant way using differential forms notation. In that notation, the independent elements of the electromagnetic field tensor can be written as the coefficients of a 2-form in Minkowski space with oriented basis .dt;dx;dy;dz/: FDExdt^dxEydt^dyEzdt^dz CBxdy^dzCBydz^dxCBzdx^dy: (4.93) Here EandBare respectively the electric field and the magnetic induction. The sources of the field, namely the charge density and the components of the current density J, become the coefficients of the 3-form JDdx^dy^dzJxdt^dy^dzJydt^dz^dxJzdt^dx^dy:(4.94) For simplicity we work in units with the permitivity, magnetic permeability, and velocity of light all set to unity ( "DDcD1). Note that it is natural that the charge and current densities occur in a 3-form; although they have together the number of components needed to constitute a four-vector, they are of dimension inverse volume. Note also that some of the signs in the formulas of this example depend on the details of the metric, and are chosen to be correct for the Minkowski metric as given in Example 4.5.2. This Minkowski metric hassignature (1,3), meaning that it has one positive and three negative diagonal elements. Some workers define the Minkowski metric to have signature (3,1), reversing all its signs. Either choice will give correct results to problems of physics if used consistently; trouble only arises if material from inconsistent sources is combined. The two homogeneous Maxwell equations are obtained from the simple formula dFD0. This equation is not a mathematical requirement on F; it is a statement of the ArfKen_Ch04-9780123846549.tex 242 Chapter 4 Tensors and Differential Forms physical properties of electric and magnetic fields. To relate our new formula to the more usual vector equations, we simply apply the doperator to F: dFD@Ex @ydyC@Ex @zdz ^dt^dx@Ey @xdxC@Ey @zdz ^dt^dy @Ez @xdxC@Ez @ydy ^dt^dzC@Bx @tdtC@Bx @xdx ^dy^dz C@By @tdtC@By @ydy ^dz^dxC@Bz @tdtC@Bz @zdz ^dx^dyD0:(4.95) Equation (4.95) is easily simplified to dFD@Ez @y@Ey @zC@Bx @t dt^dy^dzC@Ex @z@Ez @xC@By @t dt^dz^dx C@Ey @x@Ex @yC@Bz @t dt^dx^dyC@Bx @xC@By @yC@Bz @z dx^dy^dzD0: (4.96) Since the coefficient of each 3-form monomial must individually vanish, we obtain from Eq. (4.96) the vector equations rEC@B @tD0andrBD0: We now go on to obtain the two inhomogeneous Maxwell equations from the almost equally simple formula d.F/DJ. To verify this, we first form F, evaluating the starred quantities using the formulas in Eqs. (4.82): FDExdy^dzCEydz^dxCEzdx^dyCBxdt^dxCBydt^dyCBzdt^dz: We now apply the doperator, reaching after steps similar to those taken while obtaining Eq. (4.96): d.F/DrEdx^dy^dzC@Ex @t.rBx dt^dy^dz C@Ey @t.rBy dt^dz^dxC@Ez @t.rBz dt^dx^dy:(4.97) Setting d.F/from Eq. (4.97) equal to Jas given in Eq. (4.94), we obtain the remaining Maxwell equations rEDandrB@E @tDJ: We close this example by applying the doperator to J. The result must vanish because d JDd.d.F//. We get, starting from Eq. (4.94), d JD@ @tC@Jx @xC@Jy @yC@Jz @z dt^dx^dy^dzD0; ArfKen_Ch04-9780123846549.tex 4.7 Integrating Forms 243 showing that @ @tCrJD0: (4.98) Summarizing, the differential-forms approach has reduced Maxwell’s equations to the two simple formulas dFD0and d.F/DJ; (4.99) and we have also shown that Jmust satisfy an equation of continuity.  Exercises 4.6.1 Given the two 1-forms !1Dx dyCy dx and!2Dx dyy dx, calculate (a) d!1; (b) d!2: (c) For each of your answers to (a) or (b) that is nonzero, apply the operator da second time and verify that d.d!i/D0. 4.6.2 Apply the operator dtwice to!3Dxy dzCxz dyyz dx . Verify that the second application of dyields a zero result. 4.6.3 For!2and!3the 1-forms with these names in Exercises 4.6.1 and4.6.2, evaluate d.!2^!3/: (a) By forming the exterior product and then differentiating, and (b) Using the formula for differentiating a product of two forms. Verify that both approaches give the same result. 4.7 I NTEGRATING FORMS It is natural to define the integrals of differential forms in a way that preserves our usual notions of integration. The integrals with which we are concerned are over regions of the manifolds on which our differential forms are defined; this fact and the antisymmetry of the wedge product need to be taken into account in developing definitions and properties of integrals. For convenience, we illustrate in two or three dimensions; the notions extend to spaces of arbitrary dimensionality. Consider first the integral of a 1-form !in 2-D space, integrated over a curve Cfrom a start-point Pto an endpoint Q: Z C!DZ Ch AxdxCAydyi : We interpret the integration as a conventional line integral. If the curve is described para- metrically by x.t/;y.t/astincreases monotonically from tPtotQ, our integral takes the ArfKen_Ch04-9780123846549.tex 244 Chapter 4 Tensors and Differential Forms elementary form Z C!DtQZ tPh Ax.t/dx dtCAy.t/dy dti dt; and (at least in principle) the integral can be evaluated by the usual methods. Sometimes the integral will have a value that will be independent of the path from Pto Q; in physics this situation arises when a 1-form with coefficients AD.Ax;Ay/describes what is known as a conservative force (i.e., one that can be written as the gradient of a potential). In our present language, we then call !exact, meaning that there exists some function fsuch that !Dd f.x;y/ (4.100) for a region that includes the points P,Q, and all other points through which the path may pass. To check the significance of Eq. (4.100), note that it implies !D@f @xdxC@f @ydy; showing that !has as coefficients the components of the gradient of f. Given Eq. (4.100), we also see that if!Dd f;QZ P!Df.Q/f.P/: (4.101) This admittedly obvious result is independent of the dimension of the space, and is of importance to the remainder of this section. Looking next at 2-forms, we have (in 2-D space) integrals such as Z S!DZ SB.x;y/dx^dy: (4.102) We interpret dx^dyas the element of area corresponding to displacements dxanddy in mutually orthogonal directions, so in the usual notation of integral calculus we would write dx dy . Let’s now return to the wedge product notation and consider what happens if we make a change of variables from x;ytou;v, with xDauCbv,yDeuCfv. Then dxD a duCb dv,dyDe duCf dv, and dx^dyD.a duCb dv/^.e duCf dv/D.a fbe/du^dv: (4.103) We note that the coefficient of du^dvis just the Jacobian of the transformation from x;y tou;v, which becomes clear if we write aD@x=@u, etc., after which we have a fbeD @x @u@x @v @y @u@y @v D a b e f : (4.104) ArfKen_Ch04-9780123846549.tex 4.7 Integrating Forms 245 We now see a fundamental reason why the wedge product has been introduced; it has the algebraic properties needed to generate in a natural fashion the relations between elements of area (or its higher-dimension analogs) in different coordinate systems. To emphasize that observation, note that the Jacobian occurred as a natural consequence of the transfor- mation; we did not have to take additional steps to insert it, and it was generated simply by evaluating the relevant differential forms. In addition, the present formulation has one new feature: because dx^dyanddy^dxare opposite in sign, areas must be assigned algebraic signs, and it is necessary to retain the sign of the Jacobian if we make a change of variables. We therefore take as the element of area corresponding to dx^dythe ordi- nary productdxdy , with a choice of sign known as the orientation of the area. Then, Eq. (4.102) becomes Z S!DZ SB.x;y/.dxdy/; (4.105) and if elsewhere in the same computation we had dy^dx, we must convert it to dxdy using the sign opposite to that used for dx^dy. For p-forms with p>2, a corresponding analysis applies: If we transform from .x;y;:::/ to.u;v;:::/ , the wedge product dx^dy^ becomes J du^dv^ , where Jis the (signed) Jacobian of the transformation. Since the p-space volumes are oriented, the sign of the Jacobian is relevant and must be retained. Exercise 4.7.1 shows that the change of variables from the 3-form dx^dy^dztodu^dv^dwyields the determinant which is the (signed) Jacobian of the transformation. Stokes’ Theorem A key result regarding the integration of differential forms is a formula known as Stokes’ theorem, a restricted form of which we encountered in our study of vector analysis in Chapter 3. Stokes’ theorem, in its simplest form, states that if Ris a simply-connected region (i.e., one with no holes) of a p-dimensional differen- tiable manifold in a n-dimensional space ( np); Rhas a boundary denoted @R, of dimension p1; !is a.p1/-form defined on Rand its boundary, with derivative d!; then Z Rd!DZ @R!: (4.106) This is the generalization, to pdimensions, of Eq. (4.101). Note that because d!results from applying the doperator to!, the differentials in d!consist of all those in !, in the same order, but preceded by that produced by the differentiation. This observation is relevant for identifying the signs to be associated with the integrations. ArfKen_Ch04-9780123846549.tex 246 Chapter 4 Tensors and Differential Forms A rigorous proof of Stokes’ theorem is somewhat complicated, but an indication of its validity is not too involved. It is sufficient to consider the case that !is a monomial: !DA.x1;:::; xp/dx2^ dxp;d!D@A @x1dx1^dx2dxp: (4.107) We start by approximating the portion of Radjacent to the boundary by a set of small p-dimensional parallelepipeds whose thickness in the x1direction is, withhaving for each parallelepiped the sign that makes x1!x1in the interior of R. For each such parallepiped (symbolically denoted 1, with faces of constant x1denoted@1), we integrate d!inx1from x1tox1and over the full range of the other xi, obtaining Z 1d!DZ @1x1Z x1@A @x1 dx1^dx2^ dxp DZ @1A.x1;x2;:::/ dx2^ dxpZ @1A.x1;x2;:::/ dx2^ dxp:(4.108) Equation (4.108) indicates the validity of Stokes’ theorem for a laminar region whose exterior boundary is @R; if we perform the same process repeatedly, we can collapse the inner boundary to a region of zero volume, thereby reaching Eq. (4.106). Stokes’ theorem applies for manifolds of any dimension; different cases of this single theorem in two and three dimensions correspond to results originally identified as distinct theorems. Some examples follow. Example 4.7.1 GREEN’S THEOREM IN THE PLANE Consider in a 2-D space the 1-form !and its derivative: !DP.x;y/dxCQ.x;y/dy; (4.109) d!D@P @ydy^dxC@Q @xdx^dyD@Q @x@P @y dx^dy; (4.110) where we have without comment discarded terms containing dx^dxordy^dy. We apply Stokes’ theorem for this !to a region Swith boundary C, obtaining Z S@Q @x@P @y dx^dyDZ C.P dxCQ dy/: With orientation such that dx^dyDdS(ordinary element of area), we have the formula usually identified as Green’s theorem in the plane: Z C P dxCQ dy DZ S@Q @x@P @y dS: (4.111) ArfKen_Ch04-9780123846549.tex 4.7 Integrating Forms 247 Some cases of this theorem: taking PD0,QDx, we have the well-known formula Z Cx dyDZ SdSDA; where Ais the area enclosed by Cwith the line integral evaluated in the mathematically positive (counterclockwise) direction. If we take PDy,QD0, we get instead another familiar formula: Z Cy dxDZ S.1/dSDA:  When working Example 4.7.1, we assumed (without comment) that the line integral on the closed curve Cwas to be evaluated for travel in the counterclockwise direction, and we also related area to the conversion from dx^dytoCdxdy . These are choices that were not dictated by the theory of differential forms but by our intention to make its results correspond to computation in the usual system of planar Cartesian coordinates. What is certainly true is that the differential forms calculus gives a different sign for the integral ofy dx than it gave for the integral of x dy; the user of the calculus has the responsibility to make definitions corresponding to the situation for which the results are claimed to be relevant. Example 4.7.2 STOKES’ THEOREM (USUAL 3-D CASE) Let the vector potential Abe represented by the differential form !, with it and its deriva- tive of the forms !DAxdxCAydyCAzdz; (4.112) d!D@Az @y@Ay @z dy^dzC@Ax @z@Az @x dz^dxC@Ay @x@Ax @y dx^dy D.rA/xdy^dzC.rA/ydz^dxC.rA/zdx^dy: (4.113) Applying Stokes’ theorem to a region Swith boundary Cand noting that if the standard order for orienting the differentials is dx;dy;dz, then dy^dz!dx,dz^dz!dy, dx^dy!dz, and Stokes’ theorem takes the familiar form Z C AxdxCAydyCAzdz DZ CAdrDZ S.rA/d: (4.114)  Once again we have results whose interpretation depends on how we have chosen to define the quantities involved. The differential forms calculus does not know whether we intend to use a right-handed coordinate system, and that choice is implicit in our identifi- cation of the elements of area dj. In fact, the mathematics does not even tell us that the quantities we identified as components of rAactually correspond to anything physical ArfKen_Ch04-9780123846549.tex 248 Chapter 4 Tensors and Differential Forms in their indicated directions. So, once again, we emphasize that the mathematics of differ- ential forms provides a structure appropriate to the physics to which we apply it, but part of what the physicist brings to the table is the correlation between mathematical objects and the physical quantities they represent. Example 4.7.3 GAUSS’ THEOREM As a final example, consider a 3-D region Vwith boundary @V, containing an electric field given on@Vas the 2-form !, with !DExdy^dzCEydz^dzCEzdx^dy; (4.115) d!D@Ex @xC@Ey @yC@Ez @z dx^dy^dzD.rE/dx^dy^dz: (4.116) For this case, Stokes’ theorem is Z Vd!DZ V.rE/dx^dy^dzDZ V.rE/dDZ @VEd; (4.117) where dx^dy^dz!dand, just as in Example 4.7.2, dy^dz!dx, etc. We have recovered Gauss’ theorem.  Exercises 4.7.1 Use differential-forms relations to transform the integral A.x;y;z/dx^dy^dzto the equivalent expression in du^dv^dw, where u;v;w is a linear transformation of x;y;z, and thereby find the determinant that can be identified as the Jacobian of the transformation. 4.7.2 Write Oersted’s law, Z @SHdrDZ SrHdaI; in differential form notation. 4.7.3 A 1-form AdxCBdy is defined as closed if@A @yD@B @x:It is called exact if there is a function fsuch that@f @xDAand@f @yDB:Determine which of the following 1-forms are closed, or exact, and find the corresponding functions ffor those that are exact: y dxCx dy;y dxCx dy x2Cy2;Tln.xy/C1UdxCx ydy; y dx x2Cy2Cx dy x2Cy2;f.z/dzwith zDxCiy: ArfKen_Ch04-9780123846549.tex Additional Readings 249 Additional Readings Dirac, P. A. M., General Theory of Relativity. Princeton, NJ: Princeton University Press (1996). Edwards, H. M., Advanced Calculus: A Differential Forms Approach. Boston, MA: Birkhäuser (1994). Flanders, H., Differential Forms with Applications to the Physical Sciences. New York: Dover (1989). Hartle, J. B., Gravity. San Francisco: Addison-Wesley (2003). This text uses a minimum of tensor analysis. Hassani, S., Foundations of Mathematical Physics. Boston, MA: Allyn and Bacon (1991). Jeffreys, H., Cartesian Tensors. Cambridge: Cambridge University Press (1952). This is an excellent discussion of Cartesian tensors and their application to a wide variety of fields of classical physics. Lawden, D. F., An Introduction to Tensor Calculus, Relativity and Cosmology, 3rd ed. New York: Wiley (1982). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nostrand (1956). Chapter 5 covers curvilinear coordinates and 13 specific coordinate systems. Misner, C. W., K. S. Thorne, and J. A. Wheeler, Gravitation. San Francisco: W. H. Freeman (1973). A leading text on general relativity and cosmology. Moller, C., The Theory of Relativity. Oxford: Oxford University Press (1955), reprinting, (1972). Most texts on general relativity include a discussion of tensor analysis. Chapter 4 develops tensor calculus, including the topic of dual tensors. The extension to non-Cartesian systems, as required by general relativity, is presented in Chapter 9. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 includes a description of several different coordinate systems. Note that Morse and Feshbach are not above using left-handed coordinate systems even for Cartesian coordinates. Elsewhere in this excellent (and difficult) book there are many examples of the use of the various coordinate systems in solving physical problems. Eleven additional fascinating but seldom encountered orthogonal coordinate systems are discussed in the second (1970) edition of Mathematical Methods for Physicists. Ohanian, H. C., and R. Ruffini, Gravitation and Spacetime, 2nd ed. New York: Norton & Co. (1994). A well- written introduction to Riemannian geometry. Sokolnikoff, I. S., Tensor Analysis—Theory and Applications, 2nd ed. New York: Wiley (1964). Particularly useful for its extension of tensor analysis to non-Euclidean geometries. Weinberg, S., Gravitation and Cosmology. Principles and Applications of the General Theory of Relativity. New York: Wiley (1972). This book and the one by Misner, Thorne, and Wheeler are the two leading texts on general relativity and cosmology (with tensors in non-Cartesian space). Young, E. C., Vector and Tensor Analysis, 2nd ed. New York: Dekker (1993). ArfKen_Ch05-9780123846549.tex CHAPTER 5 VECTOR SPACES A large body of physical theory can be cast within the mathematical framework of vector spaces. Vector spaces are far more general than vectors in ordinary space, and the analogy may to the uninitiated seem somewhat strained. Basically, this subject deals with quantities that can be represented by expansions in a series of functions, and includes the methods by which such expansions can be generated and used for various purposes. A key aspect of the subject is the notion that a more or less arbitrary function can be represented by such an expansion, and that the coefficients in these expansions have transformation properties similar to those exhibited by vector components in ordinary space. Moreover, operators can be introduced to describe the application of various processes to a function, thereby converting it (and also the coefficients defining it) into other functions within our vector space. The concepts presented in this chapter are crucial to an understanding of quan- tum mechanics, to classical systems involving oscillatory motion, transport of material or energy, even to fundamental particle theory. Indeed, it is not excessive to claim that vector spaces are one of the most fundamental mathematical structures in physical theory. 5.1 V ECTORS IN FUNCTION SPACES We now seek to extend the concepts of classical vector analysis (from Chapter 3) to more general situations. Suppose that we have a two-dimensional (2-D) space in which the two coordinates, which are real (or in the most general case, complex) numbers that we will calla1anda2, are, respectively, associated with the two functions '1.s/and'2.s/. It is important at the outset to understand that our new 2-D space has nothing whatsoever to do with the physical xyspace. It is a space in which the coordinate point .a1;a2/corresponds to the function f.s/Da1'1.s/Ca2'2.s/: (5.1) The analogy with a physical 2-D vector space with vectors ADA1Oe1CA2Oe2is that'i.s/ corresponds toOei, while ai ! Ai, and f.s/ ! A. In other words, the coordinate 251 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch05-9780123846549.tex 252 Chapter 5 Vector Spaces values are the coefficients of the'i.s/, so each point in the space identifies a different function f.s/. Both fand'are shown above as dependent on an independent variable we calls. We choose the name sto emphasize the fact that the formulation is not restricted to the spatial variables x,y,z, but can be whatever variable, or set of variables, is needed for the problem at hand. Note further that the variable sis not a continuous analog of the discrete variables xiof an ordinary vector space. It is a parameter reminding the reader that the'ithat correspond to the dimensions of our vector space are usually not just numbers, but are functions of one or more variables. The variable(s) denoted by smay sometimes correspond to physical displacements, but that is not always the case. What should be clear is that shas nothing to do with the coordinates in our vector space; that is the role of the ai. Equation (5.1) defines a set of functions (a function space) that can be built from the basis'1,'2; we call this space a linear vector space because its members are linear com- binations of the basis functions and the addition of its members corresponds to component (coefficient) addition. If f.s/is given by Eq. (5.1) andg.s/is given by another linear combination of the same basis functions, g.s/Db1'1.s/Cb2'2.s/; with b1andb2the coefficients defining g.s/, then h.s/Df.s/Cg.s/D.a1Cb1/'1.s/C.a2Cb2/'2.s/ (5.2) defines h.s/, the member of our space (i.e., the function), which is the sum of the members f.s/andg.s/. In order for our vector space to be useful, we consider only spaces in which the sum of any two members of the space is also a member. In addition, the notion of linearity includes the requirement that if f.s/is a member of our vector space, then u.s/Dk f.s/, where kis a real or complex number, is also a member, and we can write u.s/Dk f.s/Dka1'1.s/Cka2'2.s/: (5.3) Vector spaces for which addition of two members or multiplication of a member by scalar always produces a result that is also a member are termed closed under these operations. We can summarize our findings up to this point as follows: addition of two members of our vector space causes the coefficients of the sum, h.s/inEq. (5.2), to be the sum of the coefficients of the addends, namely f.s/andg.s/; multiplication of f.s/by a ordinary number k(which, by analogy with ordinary vectors, we call a scalar), results in the multiplication of the coefficients by k. These are exactly the operations we would carry out to form the sum of two ordinary vectors, ACB, or the multiplication of a vector by a scalar, as in kA. However, here we have the coefficients aiandbi, which combine under vector addition and multiplication by a scalar in exactly the same way that we would combine the ordinary vector components AiandBi. The functions that form the basis of our vector space can be ordinary functions, and may be as simple as powers of s, or more complicated, as for example '1D.1C3sC3s2/es, '2D.13sC3s2/es, or compound quantities such as the Pauli matrices i, or even completely abstract quantities that are defined only by certain properties they may possess. The number of basis functions (i.e., the dimension of our basis) may be a small number such as 2 or 3, a larger but finite integer, or even denumerably infinite (as would arise in an ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 253 untruncated power series). The main universal restriction on the form of a basis is that the basis members be linearly independent, so that any function (member) of our vector space will be described by a unique linear combination of the basis functions. We illustrate the possibilities with some simple examples. Example 5.1.1 SOME VECTOR SPACES 1. We consider first a vector space of dimension 3, which is spanned by (meaning that it has a basis that consists of) the three functions P0.s/D1,P1.s/Ds,P2.s/D3 2s21 2. Some members of this vector space include the functions sC3D3P0.s/CP1.s/;s2D1 3P0.s/C2 3P2.s/;43sD4P0.s/3P1.s/: In fact, because we can write 1, s, and s2in terms of our basis, we can see that any quadratic form in swill be a member of our vector space, and that our space includes only functions of sthat can be written in the form c0Cc1sCc2s2. To illustrate our vector-space operations, we can form s22.sC3/D1 3P0.s/C2 3P2.s/ 2h 3P0.s/CP1.s/i D1 36 P0.s/2P1.s/C2 3P2.s/: This calculation involves only operations on the coefficients; we do not need to refer to the definitions of the Pnto carry it out. Note that we are free to define our basis any way we want, so long as its members are linearly independent. We could have chosen as our basis for this same vector space '0D1,'1Ds,'2Ds2, but we chose not to do so. 2. The set of functions 'n.s/Dsn(nD0;1;2;::: ) is a basis for a vector space whose members consist of functions that can be represented by a Maclaurin series. To avoid difficulties with this infinite-dimensional basis, we will usually need to restrict consid- eration to functions and ranges of sfor which the Maclaurin series converges. Conver- gence and related issues are of great interest in pure mathematics; in physics problems we usually proceed in ways such that convergence is assured. The members of our vector space will have representations f.s/Da0Ca1sCa2s2CD1X nD0ansn; and we can (at least in principle) use the rules for making power series expansions to find the coefficients that correspond to a given f.s/. 3. The spin space of an electron is spanned by a basis that consists of a linearly indepen- dent set of possible spin states. It is well known that an electron can have two linearly independent spin states, and they are often denoted by the symbols and . One pos- sible spin state is fDa1 Ca2 , and another is gDb1 Cb2 . We do not even need ArfKen_Ch05-9780123846549.tex 254 Chapter 5 Vector Spaces to know what and really stand for to discuss the 2-D vector space spanned by these functions, nor do we need to know the role of any parametric variable such as s. We can, however, state that the particular spin state corresponding to fCigmust have the form fCigD.a1Cib1/ C.a2Cib2/ :  Scalar Product To make the vector space concept useful and parallel to that of vector algebra in ordinary space, we need to introduce the concept of a scalar product in our function space. We shall write the scalar product of two members of our vector space, fandg, ashfjgi. This is the notation that is almost universally used in physics; various other notations can be found in the mathematics literature; examples include Tf;gUand.f;g/. The scalar product has two main features, the full meaning of which may only become clear as we proceed. They are: 1. The scalar product of a member with itself, e.g., hfjfi, must evaluate to a numeri- cal value (not a function) that plays the role of the square of the magnitude of that member, corresponding to the dot product of an ordinary vector with itself, and 2. The scalar product must be linear in each of the two members.1 There exists an extremely wide range of possibilities for defining scalar products that meet these criteria. The situation that arises most often in physics is that the members of our vector space are ordinary functions of the variable s(as in the first vector space of Example 5.1.1), and the scalar product of the two members f.s/andg.s/is computed as an integral of the type hfjgiDbZ af.s/g.s/w.s/ds; (5.4) with the choice of a,b, andw.s/dependent on the particular definition we wish to adopt for our scalar product. In the special case hfjfi, the scalar product is to be interpreted as the square of a “length,” and this scalar product must therefore be positive for any fthat is not itself identically zero. Since the integrand in the scalar product is then f.s/f.s/w.s/ and f.s/f.s/0for all s(even if f.s/is complex), we can see that w.s/must be positive over the entire range Ta;bUexcept possibly for zeros at isolated points. Let’s review some of the implications of Eq. (5.4). It is not appropriate to interpret that equation as a continuum analog of the ordinary dot product, with the variable sthought of as the continuum limit of an index labeling vector components. The integral actually arises pursuant to a decision to compute a “squared length” as a possibly weighted average over the range of values of the parameter s. We can illustrate this point by considering the 1If the members of the vector space are complex, this statement will need adjustment; see the formal definitions in the next subsection. ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 255 other situation that arises occasionally in physics, and illustrated by the third vector space in Example 5.1.1. Here we simply define the scalar products of the individual and to have values h j iDh j iD1;h j iDh j iD0; and then, taking the simple one-electron functions fDa1 Ca2 ;gDb1 Cb2 ; and assuming aiandbito be real, we expand hfjgi(using its linearity property) to reach hfjgiDa1b1h j iCa1b2h j iCa2b1h j iCa2b2h j iDa1b1Ca2b2: (5.5) These equations show that the introduction of an integral is not an indispensible step toward generalization of the scalar product; they also show that the final formula in Eq. (5.5), which isanalogous to ordinary vector algebra, arises from the expansion of hfjgiin a basis whose two members, and , are orthogonal (i.e., have a zero scalar product). Thus, the analogy to ordinary vector algebra is that the “unit vectors” of this spin system define an orthogonal “coordinate system” and that the “dot product” then has the expected form. Vector spaces that are closed under addition and multiplication by a scalar and which have a scalar product that exists for all pairs of its members are termed Hilbert spaces; these are the vector spaces of primary importance in physics. Hilbert Space Proceeding now somewhat more formally (but still without complete rigor), and includ- ing the possibility that our function space may require more than two basis functions, we identify a Hilbert space Has having the following properties: Elements (members) f,g, orhofHare subject to two operations, addition, and multiplication by a scalar (here k,k1, ork2). These operations produce quantities that are also members of the space. Addition is commutative and associative: f.s/Cg.s/Dg.s/Cf.s/;Tf.s/Cg.s/UCh.s/Df.s/CTg.s/Ch.s/U: Multiplication by a scalar is commutative, associative, and distributive: k f.s/Df.s/k;kTf.s/Cg.s/UDk f.s/Ckg.s/; .k1Ck2/f.s/Dk1f.s/Ck2f.s/;k1Tk2f.s/UDk1k2f.s/: Hisspanned by a set of basis functions 'i, where for the purposes of this book the number of such basis functions (the range of i) can either be finite or denumerably infi- nite (like the positive integers). This means that every function in Hcan be represented by the linear form f.s/DP nan'n.s/. This property is also known as completeness. We require that the basis functions be linearly independent, so that each function in the space will be a unique linear combination of the basis functions. ArfKen_Ch05-9780123846549.tex 256 Chapter 5 Vector Spaces For all functions f.s/andg.s/inH, there exists a scalar product, denoted as hfjgi, which evaluates to a finite real or complex numerical value (i.e., does not contain s) and which has the properties that 1.hfjfi0, with the equality holding only if fis identically zero.2The quantity hfjfi1=2is called the norm offand is writtenjjfjj. 2.hgjfiDhfjgi,hfjgChiDh fjgiCh fjhi;andhfjkgiDkhfjgi. Consequences of these properties are that hfjk1gCk2hiDk1hfjgiCk2hfjhi, but hk fjgiDkhfjgiandhk1fCk2gjhiDk 1hfjhiCk 2hgjhi. Example 5.1.2 SOME SCALAR PRODUCTS Continuing with the first vector space of Example 5.1.1, let’s assume that our scalar product of any two functions f.s/andg.s/takes the form hfjgiD1Z 1f.s/g.s/ds; (5.6) i.e., the formula given as Eq. (5.4) with aD1 ,bD1, andw.s/D1. Since all the mem- bers of this vector space are quadratic forms and the integral in Eq. (5.6) is over the finite range from1toC1, the scalar product will always exist and our three basis functions indeed define a Hilbert space. Before we make a few sample computations, let’s note that the brackets in the left member of Eq. (5.6) do not show the detailed form of the scalar product, thereby concealing information about the integration limits, the number of vari- ables (here we have only one, s), the nature of the space involved, the presence or absence of a weight factor w.s/, and even the exact operation that forms the product. All these features must be inferred from the context or by a previously provided definition. Now let’s evaluate two scalar products: hP0js2iD1Z 1P 0.s/s2dsD1Z 1.1/.s2/dxDs3 31 1D2 3; hP0jP2iD1Z 1.1/3 2s21 2 dsD3 2s3 31 2s1 1D0: (5.7) Looking further at the scalar product definition of the present example, we note that it is consistent with the general requirements for a scalar product, as (1) hfjfiis formed as the integral of an inherently nonnegative integrand, and will be positive for all nonzero 2To be rigorous, the phrase “identically zero” needs to be replaced by “zero except on a set of measure zero,” and other conditions need to be more tightly specified. These are niceties that are important for a precise formulation of the mathematics but are not often of practical importance to the working physicist. We note, however, that discontinuous functions do arise in applications of Fourier series, with consequences that are discussed in Chapter 19. ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 257 f; and (2) the placement of the complex-conjugate asterisk makes it obvious that hgjfiDhfjgi.  Schwarz Inequality Any scalar product that meets the Hilbert space conditions will satisfy the Schwarz inequality, which can be stated as jhfjgij2hfjfihgjgi: (5.8) Here there is equality only if fandgare proportional. In ordinary vector space, the equiv- alent result is, referring to Eq. (1.113), .AB/2DjAj2jBj2cos2jAj2jBj2; (5.9) whereis the angle between the directions of AandB. As observed previously, the equal- ity only holds if AandBare collinear. If we also require Ato be of unit length, we have the intuitively obvious result that the projection of Bonto a noncollinear Adirection will have a magnitude less than that of B. The Schwarz inequality extends this property to functions; their norms shrink on nontrivial projection. The Schwarz inequality can be proved by considering IDhfgjfgi0; (5.10) whereis an as yet undetermined constant. Treating andas linearly independent,3we differentiate Iwith respect to (remember that the left member of the product is complex conjugated) and set the result to zero, to find the value for which Iis a minimum: hgjfgiD0D)Dhgjfi hgjgi: Substituting this value into Eq. (5.10), we get (using properties of the scalar product) hfjfihfjgihgjfi hgjgi0: Noting thathgjgimust be positive, and rewriting hgjfiashfjgi, we confirm the Schwarz inequality, Eq. (5.8). Orthogonal Expansions With now a well-behaved scalar product in hand, we can make the definition that two func- tions fandgareorthogonal ifhfjgiD0, which means thathgjfiwill also vanish. An example of two functions that are orthogonal under the then-applicable definition of the scalar product are P0.s/andP2.s/, where the scalar product is that defined in Eq. (5.6) andP0,P2are the functions from Example 5.1.1; the orthogonality is shown by Eq. (5.7). We further define a function fasnormalized if the scalar product hfjfiD1; this is the 3It is not obvious that one can do this, but consider DCi,Di, withandreal. Then1 2[@=@Ci@=@ ]is equivalent to taking @=@keepingconstant. ArfKen_Ch05-9780123846549.tex 258 Chapter 5 Vector Spaces function-space equivalent of a unit vector. We will find that great convenience results if the basis functions for our function space are normalized and mutually orthogonal, cor- responding to the description of a 2-D or three-dimensional (3-D) physical vector space based on orthogonal unit vectors. A set of functions that is both normalized and mutually orthogonal is called an orthonormal set. If a member fof an orthogonal set is not nor- malized, it can be made so without disturbing the orthogonality: we simply rescale it to fDf=hfjfi1=2, so any orthogonal set can easily be made orthonormal if desired. If our basis is orthonormal, the coefficients for the expansion of an arbitrary function in that basis take a simple form. We return to our 2-D example, with the assumption that the 'iare orthonormal, and consider the result of taking the scalar product of f.s/, as given byEq. (5.1), with '1.s/: h'1jfiDh' 1j.a1'1Ca2'2/iDa1h'1j'1iCa2h'1j'2i: (5.11) The orthonormality of the 'now comes into play; the scalar product multiplying a1is unity, while that multiplying a2is zero, so we have the simple and useful result h'1jfiD a1. Thus, we have a rather mechanical means of identifying the components of f. The general result corresponding to Eq. (5.11) follows: Ifh'ij'jiDi jand fDnX iD1ai'i;then aiDh' ijfi: (5.12) Here the Kronecker delta, i j, is unity if iDjand zero otherwise. Looking once again atEq. (5.11), we consider what happens if the 'iare orthogonal but not normalized. Then instead of Eq. (5.12) we would have: If the'iare orthogonal and fDnX iD1ai'i;then aiDh'ijfi h'ij'ii: (5.13) This form of the expansion will be convenient when normalization of the basis introduces unpleasant factors. Example 5.1.3 EXPANSION IN ORTHONORMAL FUNCTIONS Consider the set of functions n.x/Dsinnx, for nD1;2;::: , to be used for xin the interval 0xwith scalar product hfjgiDZ 0f.x/g.x/dx: (5.14) We wish to use these functions for the expansion of the function x2.x/. First, we check that they are orthogonal: SnmDZ 0 n.x/m.x/dxDZ 0sinnxsinmx dx: ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 259 Forn6Dmthis integral can be shown to vanish, either by symmetry considerations or by consulting a table of integrals. To determine normalization, we need Snn; from symmetry considerations, the integrand, sin2nxD1 2.1cos 2nx/, can be seen to have average value 1/2 over the range .0;/ , leading to SnnD=2for all integer n. This means the nare not normalized, but can be made so if we multiply byp2=. So our orthonormal basis will be 'n.x/D2 1=2 sinnx;nD1;2;3;:::: (5.15) To expand x2.x/, we apply Eq. (5.2), which requires the evaluation of anDh' njx2.x/iD2 1=2Z 0.sinnx/x2.x/dx; (5.16) for use in the expansion x2.x/D2 1=21X nD0ansinnx: (5.17) Evaluating cases of Eq. (5.16) by hand or using a computer for symbolic computation, we have for the first few an:a1D5:0132 ,a2D1:8300 ,a3D0:1857 ,a4D0:2350 . The convergence is not very fast.  Example 5.1.4 SPIN SPACE A system of four spin-1 2particles in a triplet state has the following three linearly indepen- dent spin functions: 1D ; 2D ; 3D C : The four symbols in each term of these expressions refer to the spin assignments of the four particles, in numerical order. The scalar product in the spin space has the form, for monomials, habcdjwxyziDawbxcydz; meaning that the scalar product is unity if the two monomials are identical, and is zero if they are not. Scalar products involving polynomials can be evaluated by expanding them into sums of monomial products. It is easy to confirm that this definition meets the require- ments for a valid scalar product. Our mission will be (1) verify that the iare orthogonal; (2) convert them, if neces- sary, to normalized form to make an orthonormal basis for the spin space; and (3) expand the following triplet spin function as a linear combination of the orthonormal spin basis functions: 0D : The functions 1and2are orthogonal, as they have no terms in common. Although 1 and3have two terms in common, they occur in sign combinations leading to a vanishing ArfKen_Ch05-9780123846549.tex 260 Chapter 5 Vector Spaces scalar product. The same observation applies to h2j3i. However, none of the iare nor- malized. We findh1j1iDh2j2iD2,h3j3iD4, so an orthonormal basis would be '1D21=21; ' 2D21=22; ' 3D1 23: Finally, we obtain the coefficients for the expansion of 0by forming a1Dh' 1j0iD 1=p 2,a2Dh' 2j0iD1=p 2, and a3Dh' 3j0iD1. Thus, the desired expansion is 0D1p 2'11p 2'2C'3:  Expansions and Scalar Products If we have found the expansions of two functions, fDX a'and gDX b'; then their scalar product can be written hfjgiDX a bh'j'i: If the'set is orthonormal, the above reduces to hfjgiDX a b: (5.18) In the special case gDf, this reduces to hfjfiDX jaj2; (5.19) consistent with the requirement that hfjfi0, with equality only if fis zero “almost everywhere.” If we regard the set of expansion coefficients aas the elements of a column vector arepresenting f, with column vector bsimilarly representing g, Eqs. (5.18) and(5.19) correspond to the matrix equations hfjgiDa†b;hfjfiDa†a: (5.20) Note that by taking the adjoint of a, we both complex conjugate it and convert it into a row vector, so that the matrix products in Eq. (5.20) collapse to scalars, as required. ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 261 Example 5.1.5 COEFFICIENT VECTORS A set of functions that is orthonormal on 0xis 'n.x/Dr 2n0 cosnx;nD0;1;2;:::: First, let us expand in terms of this basis the two functions 1Dcos3xCsin2xCcosxC1and 2Dcos2xcosx: We write the expansions as vectors a1anda2with components nD0;:::; 3: a1D0 BB@h'0j 1i h'1j 1i h'2j 1i h'3j 1i1 CCA;a2D0 BB@h'0j 2i h'1j 2i h'2j 2i h'3j 2i1 CCA: All components beyond nD3vanish and need not be shown. It is straightforward to evaluate these scalar products. Alternatively, we can rewrite the iusing trigonometric identities, reaching the forms 1Dcos 3 x 4cos 2 x 2C7 4cosxC3 2; 2Dcos 2 x 2cosxC1 2: These expressions are now easily recognized as equivalent to 1Dr 2 '3 4'2 2C7'1 4C3p 2'0 2! ; 2Dr 2 '2 2'1Cp 2'0 2! ; so a1Dr 20 BB@3p 2=2 7=4 1=2 1=41 CCA;a2Dr 20 BB@p 2=2 1 1=2 01 CCA: We see from the above that the general formula for finding the coefficients in an orthonor- mal expansion, Eq. (5.12), is a systematic way of doing what sometimes can be carried out in other ways. We can now evaluate the scalar products h ij ji. Identifying these first as matrix prod- ucts that we then evaluate, h 1j 1iDa† 1a1D63 16;h 1j 2iDa† 1a2D 4;h 2j 2iDa† 2a2D7 8:  ArfKen_Ch05-9780123846549.tex 262 Chapter 5 Vector Spaces Bessel’s Inequality Given a set of basis functions and the definition of a space, it is not necessarily assured that the basis functions span the space (a property sometimes referred to as completeness). For example, we might have a space defined to be that containing all functions possessing a scalar product of a given definition, while the basis functions have been specified by giving their functional form. This issue is of some importance, because we need to know whether an attempt to expand a function in a given basis can be guaranteed to converge to the correct result. Totally general criteria are not available, but useful results have been obtained if the function being expanded has, at worst, a finite number of finite discontinuities, and results are accepted as “accurate” if deviations from the correct value occur only at isolated points. Power series and trigonometric series have been proved complete for the expansion of square integrable functions f(those for whichhfjfias defined in Eq. (5.7) exists; mathematicians identify such spaces by the designation L2). Also proved complete are the orthonormal sets of functions that arise as the solutions to Hermitian eigenvalue problems.4 A not too practical test for completeness is provided by Bessel’s inequality, which states that if a function fhas been expanded in an orthonormal basis asP nan'n, then hfjfiX njanj2; (5.21) with the inequality occurring if the expansion of fis incomplete. The impracticality of this as a completeness test is that one needs to apply it for all fbefore using it to claim completeness of the space. We establish Bessel’s inequality by considering ID* fX iai'i fX jaj'j+ 0; (5.22) where ID0represents what is termed convergence in the mean, a criterion that per- mits the integrand to deviate from zero at isolated points. Expanding the scalar product, and eliminating terms that vanish because the 'are orthonormal, we arrive at Eq. (5.21), with equality only resulting if the expansion converges to f. We note in passing that con- vergence in the mean is a less stringent requirement than uniform convergence, but is adequate for almost all physical applications of basis-set expansions. Example 5.1.6 EXPANSION OF A DISCONTINUOUS FUNCTION The functions cosnx.nD0;1;2;:::/ andsinnx.nD1;2;:::/ have (together) been shown to form a complete set on the interval  < x<. Since this determination is obtained subject to convergence in the mean, there is the possibility of deviation at iso- lated points, thereby permitting the description of functions with isolated discontinuities. 4See R. Courant and D. Hilbert, Methods of Mathematical Physics (English translation), Vol. 1, New York: Interscience (1953), reprinting, Wiley (1989), chapter 6, section 3. ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 263 We illustrate with the square-wave function f.x/D8 >< >:h 2; 0<x< h 2;< x<0:(5.23) The functions cosnxandsinnxare orthogonal on the expansion interval (with unit weight in the scalar product), and the expansion of f.x/takes the form f.x/Da0C1X nD1.ancosnxCbnsinnx/: Because f.x/is an odd function of x, all the anvanish, and we only need to compute bnD1 Z f.t/sinnt dt: The factor 1= preceding the integral arises because the expansion functions are not normalized. Upon substitution of h=2forf.t/, we find bnDh n.1cosn/D8 < :0; neven; 2h n;nodd: Thus, the expansion of the square wave is f.x/D2h 1X nD0sin.2nC1/x 2nC1: (5.24) To give an idea of the rate at which the series in Eq. (5.24) converges, some of its partial sums are plotted in Fig. 5.1.  Expansions of Dirac Delta Function Orthogonal expansions provide opportunities to develop additional representations of the Dirac delta function. In fact, such a representation can be built from any complete set of functions'n.x/. For simplicity we assume the 'nto be orthonormal with unit weight on the interval.a;b/, and consider the expansion .xt/D1X nD0cn.t/'n.x/; (5.25) where, as indicated, the coefficients must be functions of t. From the rule for determining the coefficients, we have, for talso in the interval .a;b/, cn.t/DbZ a' n.x/.xt/dxD' n.t/; (5.26) ArfKen_Ch05-9780123846549.tex 264 Chapter 5 Vector Spaces FIGURE 5.1 Expansion of square wave. Computed using Eq. (5.24) with summation terminated after nD4, 8, 12, and 20. Curves are at different vertical scales to enhance visibility. where the evaluation has used the defining property of the delta function. Substituting this result back into Eq. (5.25), we have .xt/D1X nD0' n.t/'n.x/: (5.27) This result is clearly not uniformly convergent at xDt. However, remember that it is not to be used by itself, but has meaning only when it appears as part of an integrand. Note also that Eq. (5.27) is only valid when xandtare within the range .a;b/. Equation (5.27) is called the closure relation for the Dirac delta function (with respect to the'n) and obviously depends on the completeness of the 'set. If we apply Eq. (5.27) to an arbitrary function F.t/that we assume to have the expansion F.t/DP pcp'p.t/, we have bZ aF.t/.xt/dtDbZ adt1X pD0cp'p.t/1X nD0' n.t/'n.x/ D1X pD0cp'p.x/DF.x/; (5.28) which is the expected result. However, if we replace the integration limits .a;b/by.t1;t2/ such that at1<t2b, we get a more general result that reflects the fact that our ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 265 020406080 0.2 1 0.8 0.6 0.4 FIGURE 5.2 Approximation at ND80to.tx/, Eq. (5.30), for tD0:4. representation of .xt/is negligible except when xt: t2Z t1F.t/.xt/dtD(F.x/;t1<x<t2; 0; x<t1orx>t2.(5.29) Example 5.1.7 DELTA FUNCTION REPRESENTATION To illustrate an expansion of the Dirac delta function in an orthonormal basis, take 'n.x/Dp 2 sin nx, which are orthonormal and complete on xD.0;1/fornD1;2;::: . Then the Dirac delta function has representation, valid for 0<x<1,0<t<1, .xt/Dlim N!1NX nD12 sinntsinnx: (5.30) Plotting this with ND80fortD0:4and0<x<1gives the result shown in Fig. 5.2.  Dirac Notation Much of what we have discussed can be brought to a form that promotes clarity and suggests possibilities for additional analysis by using a notational device invented by P. A. M. Dirac. Dirac suggested that instead of just writing a function f, it be written enclosed in the right half of an angle-bracket pair, which he named a ket. Thus f!jfi, 'i!j' ii, etc. Then he suggested that the complex conjugates of functions be enclosed in left half-brackets, which he named bras. An example of a bra is ' i!h' ij. Finally, he ArfKen_Ch05-9780123846549.tex 266 Chapter 5 Vector Spaces suggested that when the sequence (bra followed by ket DbraCketbracket) is encoun- tered, the pair should be interpreted as a scalar product (with the dropping of one of the two adjacent vertical lines). As an initial example of the use of this notation, take Eq. (5.12), which we now write as jfiDX jajj'jiDX jj'jih'jjfiD0 @X jj'jih'jj1 Ajfi: (5.31) This notational rearrangement shows that we can view the expansion in the 'basis as the insertion of a set of basis members in a way which, in sum, has no effect. If the sum is over a complete set of 'j, the ket-bra sum in Eq. (5.31) will have no net effect when inserted before any ket in the space, and therefore we can view the sum as a resolution of the identity. To emphasize this, we write 1DX jj'jih'jj: (5.32) Many expressions involving expansions in orthonormal sets can be derived by the insertion of resolutions of the identity. Dirac notation can also be applied to expressions involving vectors and matrices, where it illuminates the parallelism between physical vector spaces and the function spaces here under study. If aandbare column vectors and Mis a matrix, then we can write jbias a synonym for b, we can writehajto mean a†, and thenhajbi is interpreted as equivalent to a†b, which (when the vectors are real) is matrix notation for the (scalar) dot product ab. Other examples are expressions such as aDMb$jaiDjMbiD Mjbi ora†MbD.M†a/†b$hajMbiDhM†ajbi: Exercises 5.1.1 A function f.x/is expanded in a series of orthonormal functions f.x/D1X nD0an'n.x/: Show that the series expansion is unique for a given set of 'n.x/. The functions 'n.x/ are being taken here as the basis vectors in an infinite-dimensional Hilbert space. 5.1.2 A function f.x/is represented by a finite set of basis functions 'i.x/, f.x/DNX iD1ci'i.x/: Show that the components ciare unique, that no different set c0 iexists. Note. Your basis functions are automatically linearly independent. They are not neces- sarily orthogonal. ArfKen_Ch05-9780123846549.tex 5.1 Vectors in Function Spaces 267 5.1.3 A function f.x/is approximated by a power seriesPn1 iD0cixiover the intervalT0;1U. Show that minimizing the mean square error leads to a set of linear equations AcDb; where Ai jD1Z 0xiCjdxD1 iCjC1;i;jD0;1;2;:::; n1 and biD1Z 0xif.x/dx;iD0;1;2;:::; n1: Note. The Ai jare the elements of the Hilbert matrix of order n. The determinant of this Hilbert matrix is a rapidly decreasing function of n. For nD5,detAD3:71012and the set of equations AcDbis becoming ill-conditioned and unstable. 5.1.4 In place of the expansion of a function F.x/given by F.x/D1X nD0an'n.x/; with anDbZ aF.x/'n.x/w.x/dx; take the finite series approximation F.x/mX nD0cn'n.x/: Show that the mean square error bZ a" F.x/mX nD0cn'n.x/#2 w.x/dx is minimized by taking cnDan. Note. The values of the coefficients are independent of the number of terms in the finite series. This independence is a consequence of orthogonality and would not hold for a least-squares fit using powers of x. 5.1.5 From Example 5.1.6, f.x/D8 >>< >>:h 2; 0<x< h 2;< x<09 >>= >>;D2h 1X nD0sin.2nC1/x 2nC1: ArfKen_Ch05-9780123846549.tex 268 Chapter 5 Vector Spaces (a) Show that Z h f.x/i2 dxD 2h2D4h2 1X nD0.2nC1/2: For a finite upper limit this would be Bessel’s inequality. For the upper limit 1, this is Parseval’s identity. (b) Verify that  2h2D4h2 1X nD0.2nC1/2 by evaluating the series. Hint. The series can be expressed in terms of the Riemann zeta function .2/D2=6. 5.1.6 Derive the Schwarz inequality from the identity 2 4bZ af.x/g.x/dx3 52 DbZ ah f.x/i2 dxbZ aTh g.x/i2 dx 1 2bZ adxbZ adyh f.x/g.y/f.y/g.x/i2 : 5.1.7 Starting from ID* fX iai'i fX jaj'j+ 0; derive Bessel’s inequality, hfjfiX njanj2. 5.1.8 Expand the function sinxin a series of functions 'ithat are orthogonal (but not nor- malized) on the range 0x1when the scalar product has definition hfjgiD1Z 0f.x/g.x/dx: Keep the first four terms of the expansion. The first four 'iare: '0D1; ' 1D2x1; ' 2D6x26xC1; ' 3D20x330x2C12x1: Note. The integrals that are needed are the subject of Example 1.10.5. 5.1.9 Expand the function exin Laguerre polynomials Ln.x/, which are orthonormal on the range 0x<1with scalar product hfjgiD1Z 0f.x/g.x/exdx: ArfKen_Ch05-9780123846549.tex 5.2 Gram-Schmidt Orthogonalization 269 Keep the first four terms of the expansion. The first four Ln.x/are L0D1; L1D1x;L2D24xCx2 2;L3D618xC9x2x3 6: 5.1.10 The explicit form of a function fis not known, but the coefficients anof its expan- sion in the orthonormal set 'nare available. Assuming that the 'nand the members of another orthonormal set, n, are available, use Dirac notation to obtain a formula for the coefficients for the expansion of fin thenset. 5.1.11 Using conventional vector notation, evaluateX jjOejihOejjai, where ais an arbitrary vec- tor in the space spanned by the Oej. 5.1.12 Letting aDa1Oe1Ca2Oe2andbDb1Oe1Cb2Oe2be vectors in R2, for what values of k, if any, is hajbiD a1b1a1b2a2b1Cka2b2 a valid definition of a scalar product? 5.2 G RAM-SCHMIDT ORTHOGONALIZATION Crucial to carrying out the expansions and transformations under discussion is the avail- ability of useful orthonormal sets of functions. We therefore proceed to the description of a process whereby a set of functions that is neither orthogonal or normalized can be used to construct an orthonormal set that spans the same function space. There are many ways to accomplish this task. We present here the method called the Gram-Schmidt orthogo- nalization process. The Gram-Schmidt process assumes the availability of a set of functions and an appropriately defined scalar product hfjgi. We orthonormalize sequentially to form the orthonormal functions ', meaning we make the first orthonormal function, '0, from0, the next,'1, from0and1, etc. If, for example, the are powers x, the orthonormal function'will be a polynomial of degree inx. Because the Gram-Schmidt process is often applied to powers, we have chosen to number both the and the'sets starting from zero (rather than 1). Thus, our first orthonormal function will simply be a normalized version of 0. Specifically, '0D0 h0j0i1=2: (5.33) To check that Eq. (5.33) is correct, we form h'0j'0iD* 0 h0j0i1=2 0 h0j0i1=2+ D1: Next, starting from '0and1, we form a function that is orthogonal to '0. We use'0rather than0to be consistent with what we will do in later steps of the process. Thus, we write 1D1a1;0'0: (5.34) ArfKen_Ch05-9780123846549.tex 270 Chapter 5 Vector Spaces What we are doing here is the removal from 1of its projection onto '0, leaving a remain- der that will be orthogonal to '0. Remembering that '0is normalized (of “unit length”), that projection is identified as h'0j1i'0, so that a1;0Dh' 0j1i: (5.35) In case Eq. (5.35) is not intuitively obvious, we can confirm it by writing the requirement that 1be orthogonal to '0: h'0j 1iDD '0  1a1;0'0E Dh' 0j1ia1;0h'0j'0iD0; which, because '0is normalized, reduces to Eq. (5.35). The function 1is not in general normalized. To normalize it and thereby obtain '1, we form '1D 1 h 1j 1i1=2: (5.36) To continue further, we need to make, from '0,'1, and2, a function that is orthogonal to both'0and'1. It will have the form 2D2a0;2'0a1;2'1: (5.37) The last two terms of Eq. (5.37), respectively, remove from 2its projections on '0and'1; these projections are independent because '0and'1are orthogonal. Thus, either from our knowledge of projections or by setting to zero the scalar products h'ij 2i(iD0and 1), we establish a0;2Dh' 0j2i;a1;2Dh' 1j2i: (5.38) Finally, we make '2D 2=h 2j 2i1=2. The generalization for which the above is the first few terms is that, given the prior formation of 'i,iD0;:::; n1, the orthonormal function 'nis obtained from nby the following two steps: nDnn1X D0h'jni'; 'nD n h nj ni1=2: (5.39) Reviewing the above process, we note that different results would have been obtained if we used the same set of i, but simply took them in a different order. For example, if we had started with 3, one of our orthonormal functions would have been a multiple of 3, while the set we constructed yielded '3as a linear combination of ,D0;1;2;3. Example 5.2.1 LEGENDRE POLYNOMIALS Let us form an orthonormal set, taking the asx, and making the definition hfjgiD1Z 1f.x/g.x/dx: (5.40) ArfKen_Ch05-9780123846549.tex 5.2 Gram-Schmidt Orthogonalization 271 This scalar product definition will cause the members of our set to be orthogonal, with unit weight, on the range .1; 1/. Moreover, since the are real, the complex conjugate asterisk has no operational significance here. The first orthonormal function, '0, is '0.x/D1 h1j1i1=2D1 " 1R 1dx#1=2D1p 2: To obtain'1, we first obtain 1by evaluating 1.x/Dxh'0jxi'0.x/Dx; where the scalar product vanishes because '0is an even function of x, whereas xis odd, and the range of integration is even. We then find '1.x/Dx " 1R 1x2dx#1=2Dr 3 2x: The next step is less trivial. We form 2.x/Dx2h' 0jx2i'0.x/h' 1jx2i'1.x/Dx21p 2 x21p 2 Dx21 3; where we have used symmetry to set h'1jx2ito zero and evaluated the scalar product 1p 2 x2 D1p 21Z 1x2dxDp 2 3: Then, '2.x/Dx21 3" 1R 1 x21 32dx#1=2Dr 5 23 2x21 2 : Continuation to one more orthonormal function yields '3.x/Dr 7 25 2x33 2x : Reference to Chapter 15 will show that 'n.x/Dr 2nC1 2Pn.x/; (5.41) where Pn.x/is the nth degree Legendre polynomial. Our Gram-Schmidt process provides a possible but very cumbersome method of generating the Legendre polynomials; other, more efficient approaches exist.  ArfKen_Ch05-9780123846549.tex 272 Chapter 5 Vector Spaces Table 5.1 Orthogonal Polynomials Generated by Gram-Schmidt Orthogonalization of un.x/Dxn,nD0;1;2;::: . Polynomials Scalar Products Table Legendre1Z 1Pn.x/Pm.x/dxD2mn=.2nC1/ Table 15.1 Shifted Legendre1Z 0P n.x/P m.x/dxDmn=.2nC1/ Table 15.2 Chebyshev I1Z 1Tn.x/Tm.x/ 1x21=2dxDmn=.2n0/ Table 18.4 Shifted Chebyshev I1Z 0T n.x/T m.x/Tx.1x/U1=2dxDmn=.2n0/ Table 18.5 Chebyshev II1Z 1Un.x/Um.x/ 1x21=2dxDmn=2 Table 18.4 Laguerre1Z 0Ln.x/Lm.x/exdxDmn Table 18.2 Associated Laguerre1Z 0Lk n.x/Lk m.x/exdxDmn.nCk/W=nW Table 18.3 Hermite1Z 1Hn.x/Hm.x/ex2dxD2nmn1=2nW Table 18.1 The intervals, weights, and conventional normalization can be deduced from the forms of the scalar products. Tables of explicit formulas for the first few polynomials of each type are included in the indicated tables appearing in Chapters 15 and 18 of this book. The Legendre polynomials are, except for sign and scale, uniquely defined by the Gram- Schmidt process, the use of successive powers of x, and the definition adopted for the scalar product. By changing the scalar product definition (different weight or range), we can generate other useful sets of orthogonal polynomials. A number of these are presented in Table 5.1. For various reasons most of these polynomial sets are not normalized to unity. The scalar product formulas in the table give the conventional normalizations, and are those of the explicit formulas referenced in the table. Orthonormalizing Physical Vectors The Gram-Schmidt process also works for ordinary vectors that are simply given by their components, it being understood that the scalar product is just the ordinary dot product. ArfKen_Ch05-9780123846549.tex 5.2 Gram-Schmidt Orthogonalization 273 Example 5.2.2 ORTHONORMALIZING A 2-D MANIFOLD A 2-D manifold (subspace) in 3-D space is defined by the two vectors a1DOe1COe22Oe3 anda2DOe1C2Oe23Oe3. In Dirac notation, these vectors (written as column matrices) are ja1iD0 @1 1 21 A;ja2iD0 @1 2 31 A: Our task is to span this manifold with an orthonormal basis. We proceed exactly as for functions: Our first orthonormal basis vector, which we call b1, will be a normalized version of a1, and therefore formed as jb1iDa1 ha1ja1i1=2D1 61=2ja1iD1 61=20 @1 1 21 A: An unnormalized version of a second orthonormal function will have the form jb0 2iDja 2ihb 1ja2ijb1iDja 2i9 61=2jb1iD0 @1=2 1=2 01 A: Normalizing, we reach jb2iDb0 2 hb0 2jb0 2i1=2D1p 20 @1 1 01 A:  Exercises For the Gram-Schmidt constructions in Exercises 5.2.1 through 5.2.6, use a scalar prod- uct of the form given in Eq. (5.7) with the specified interval and weight. 5.2.1 Following the Gram-Schmidt procedure, construct a set of polynomials P n.x/orthog- onal (unit weighting factor) over the range T0;1Ufrom the setT1;x;x2; :::U. Scale so thatP n.1/D1. ANS. P n.x/D1, P 1.x/D2x1, P 2.x/D6x26xC1, P 3.x/D20x330x2C12x1. These are the first four shifted Legendre polynomials. Note. The “*” is the standard notation for “shifted”: T0;1Uinstead ofT1; 1U. It does not mean complex conjugate. 5.2.2 Apply the Gram-Schmidt procedure to form the first three Laguerre polynomials: un.x/Dxn;nD0;1;2;:::; 0x<1; w. x/Dex: ArfKen_Ch05-9780123846549.tex 274 Chapter 5 Vector Spaces The conventional normalization is 1Z 0Lm.x/Ln.x/exdxDmn: ANS. L0D1,L1D.1x/,L2D24xCx2 2. 5.2.3 You are given (a) a set of functions un.x/Dxn,nD0;1;2;::: , (b) an interval .0;1/, (c) a weighting function w.x/Dxex. Use the Gram-Schmidt procedure to construct the first three orthonormal functions from the set un.x/for this interval and this weighting function. ANS.'0.x/D1,'1.x/D.x2/=p 2,'2.x/D.x26xC6/=2p 3. 5.2.4 Using the Gram-Schmidt orthogonalization procedure, construct the lowest three Hermite polynomials: un.x/Dxn;nD0;1;2;:::;1<x<1; w. x/Dex2: For this set of polynomials the usual normalization is 1Z 1Hm.x/Hn.x/w.x/dxDmn2mmW1=2: ANS. H0D1,H1D2x,H2D4x22. 5.2.5 Use the Gram-Schmidt orthogonalization scheme to construct the first three Chebyshev polynomials (type I): un.x/Dxn;nD0;1;2;:::;1x1; w. x/D.1x2/1=2: Take the normalization 1Z 1Tm.x/Tn.x/w.x/dxDmn8 < :; mDnD0;  2;mDn1: Hint. The needed integrals are given in Exercise 13.3.2. ANS. T0D1,T1Dx,T2D2x21,.T3D4x33x/. 5.2.6 Use the Gram-Schmidt orthogonalization scheme to construct the first three Chebyshev polynomials (type II): un.x/Dxn;nD0;1;2;:::;1x1; w. x/D.1x2/C1=2: Take the normalization to be 1Z 1Um.x/Un.x/w.x/dxDmn 2: ArfKen_Ch05-9780123846549.tex 5.3 Operators 275 Hint. 1Z 1.1x2/1=2x2ndxD 2135.2n1/ 468.2nC2/;nD1;2;3;::: D 2;nD0: ANS. U0D1,U1D2x,U2D4x21. 5.2.7 As a modification of Exercise 5.2.5, apply the Gram-Schmidt orthogonalization proce- dure to the set un.x/Dxn,nD0;1;2;::: ,0x<1. Takew.x/to be exp. x2/. Find the first two nonvanishing polynomials. Normalize so that the coefficient of the highest power of xis unity. In Exercise 5.2.5, the interval .1;1/led to the Hermite polynomials. The functions found here are certainly not the Hermite polynomials. ANS.'0D1,'1Dx1=2. 5.2.8 Form a set of three orthonormal vectors by the Gram-Schmidt process using these input vectors in the order given: c1D0 @1 1 11 A;c2D0 @1 1 21 A;c3D0 @1 0 21 A: 5.3 O PERATORS An operator is a mapping between functions in its domain (those to which it can be applied) and functions in its range (those it can produce). While the domain and the range need not be in the same space, our concern here is for operators whose domain and range are both in all or part of the same Hilbert space. To make this discussion more concrete, here are a few examples of operators: Multiplication by 2: Converts finto2f; For a space containing algebraic functions of a variable x,d=dx: Converts f.x/into d f=dx; An integral operator Adefined by A f.x/DR G.x;x0/f.x0/dx0: A special case of this is a projection operator j'iih'ij, which converts fintoh'ijfi'i. In addition to the abovementioned restriction on domain and range, we also for our present purposes restrict attention to operators that are linear, meaning that if AandBare linear operators, fandgfunctions, and ka constant, then .ACB/fDA fCB f;A.fCg/DA fCAg;AkDk A: For both electromagnetic theory and quantum mechanics, an important class of operators aredifferential operators, those that include differentiation of the functions to which they are applied. These operators arise when differential equations are written in operator form; ArfKen_Ch05-9780123846549.tex 276 Chapter 5 Vector Spaces for example, the operator L.x/D 1x2d2 dx22xd dx enables us to write Legendre’s differential equation, 1x2d2y dx22xdy dxCyD0; in the form L.y/yD y. When no confusion thereby results, this can be shortened to LyD y. Commutation of Operators Because differential operators act on the function(s) to their right, they do not necessarily commute with other operators containing the same independent variable. This fact makes it useful to consider the commutator of operators AandB, TA;BUDABB A: (5.42) We can often reduce ABB Ato a simpler operator expression. When we write an operator equation, its meaning is that the operator on the left-hand side of the equation produces the same effect on every function in its domain as is produced by the opera- tor on the right-hand side. Let’s illustrate this point by evaluating the commutator Tx;pU, where pDi d=dx. The imaginary unit iand the name pappear because this operator is that corresponding in quantum mechanics to momentum (in a system of units such that NhDh=2D1). The operator xstands for multiplication by x. To carry out the evaluation, we apply Tx;pUto an arbitrary function f.x/. Inserting the explicit form of p, we have Tx;pUf.x/D.xppx/f.x/Dixd f.x/ dx id dx x f.x/ Dix f0.x/Ci f.x/Cx f0.x/ Di f.x/; indicating that Tx;pUDi: (5.43) As indicated before, this means Tx;pUf.x/Di f.x/for all f. We can carry out various algebraic manipulations on commutators. In general, if A,B, Care operators and kis a constant, TA;BUDT B;AU;TA;BCCUDT A;BUCTA;CU; kTA;BUDTk A;BUDT A;k BU: (5.44) ArfKen_Ch05-9780123846549.tex 5.3 Operators 277 Example 5.3.1 OPERATOR MANIPULATION GivenTx;pU, we can simplify the commutator Tx;p2U. We write, being careful about the operator ordering and using Eq. (5.43), Tx;p2UDxp2pxpCpxpp2xDTx;pUpCpTx;pUD2i p; (5.45) a result also obtainable from x d2 dx2 f.x/ d2 dx2 x f.x/D2f0.x/D2i id dx f.x/: However, note that Eq. (5.45) follows solely from the validity of Eq. (5.43), and will apply to any quantities xandpthat satisfy that commutation relation, whether or not we are operating with ordinary functions and their derivatives. Put another way, if xandpare operators in some abstract Hilbert space and all we know about them is Eq. (5.43), we may still conclude that Eq. (5.45) is also valid.  Identity, Inverse, Adjoint An operator that is generally available is the identity operator, namely one that leaves functions unchanged. Depending on the context, this operator will be denoted either Ior simply 1. Some, but not all operators will have an inverse, namely an operator that will “undo” its effect. Letting A1denote the inverse of A, ifA1exists, it will have the property A1ADAA1D1: (5.46) Associated with many operators will be another operator, called its adjoint and denoted A†, which will be such that for all functions fandgin the Hilbert space, hfjAgiDh A†fjgi: (5.47) Thus, we see that A†is an operator that, applied to the left member of anyscalar product, produces the same result as is obtained if Ais applied to the right member of the same scalar product. Equation (5.47) is, in essence, the defining equation for A†. Depending on the specific operator A, and the definitions in use of the Hilbert space and the scalar product, A†may or may not be equal to A. IfADA†,Ais referred to as self-adjoint, or equivalently, Hermitian. If A†DA,Ais called anti-Hermitian. This definition is worth emphasis: IfH†DH;His Hermitian. (5.48) Another situation of frequent occurrence is that the adjoint of an operator is equal to its inverse, in which case the operator is called unitary. A unitary operator Uis therefore defined by the following statement: IfU†DU1;Uis unitary. (5.49) In the special case that Uis both real and unitary, it is called orthogonal . ArfKen_Ch05-9780123846549.tex 278 Chapter 5 Vector Spaces The reader will doubtless note that the nomenclature for operators is similar to that previously introduced for matrices. This is not accidental; we shall shortly develop corre- spondences between operator and matrix expressions. Example 5.3.2 FINDING THE ADJOINT Consider an operator ADx.d=dx/whose domain is the Hilbert space whose members f have a finite value of hfjfiwhen the scalar product has definition hfjgiD1Z 1f.x/g.x/dx: This space is often referred to as L2on.1;1/. Starting fromhfjA gi, we integrate by parts as needed to move the operator out of the right half of the scalar product. Because f andgmust vanish at1, the integrated terms vanish, and we get hfjA giD1Z 1fxdg dxdxD1Z 1 x fdg dxdxD1Z 1d x f dxg dx D d dx x f g : We see from the above that A†D.d=dx/x, from which we can find A†DA1. This Ais clearly neither Hermitian nor unitary (with the specified definition of the scalar product).  Example 5.3.3 ADJOINT DEPENDS ON SCALAR PRODUCT For the Hilbert space and scalar product of Example 5.3.2, an integration by parts easily establishes that an operator ADi.d=dx/is self-adjoint, i.e., A†DA. But now let’s consider the same operator A, but for the L2space with1x1(and with a scalar product of the same form, but with integration limits 1). In this space, the integrated terms from the integration by parts do not vanish, but we can incorporate them into an operator on the left half of the scalar product by adding delta-function terms:  f id dx g Di fg 1 1C1Z 1 id f dx g dx D1Z 1h i.x1/i.xC1/id dxi f.x/ g.x/dx: In this truncated space the operator Aisnotself-adjoint.  ArfKen_Ch05-9780123846549.tex 5.3 Operators 279 Basis Expansions of Operators Because we are dealing only with linear operators, we can write the effect of an operator on an arbitrary function if we know the result of its action on all members of a basis spanning our Hilbert space. In particular, assume that the action of an operator Aon member'of an orthonormal basis has the result, also expanded in that basis, A'DX a': (5.50) Assuming this form for the result of operation with Ais not a major restriction; all it says is that the result is in our Hilbert space. Formally, the coefficients acan be obtained by taking scalar products: aDh'jA'iDh'jAj'i: (5.51) Following common usage, we have inserted an optional (operationally meaningless) verti- cal line between Aand'. This notation has the aesthetic effect of separating the operator from the two functions entering the scalar product, and also emphasizes the possibility that instead of evaluating the scalar product as written, we can without changing its value evaluate it using the adjoint of A, ashA†'j'i. We now apply Eq. (5.50) to a function whose expansion in the 'basis is DX c';cDh'j i: (5.52) The result is A DX cA'DX cX a'DX  X ac! ': (5.53) If we think of A as a function in our Hilbert space, with expansion DX b'; (5.54) we then see from Eq. (5.53) that the coefficients bare related to candain a way corresponding to matrix multiplication. To make this more concrete, Define cas a column vector with elements ci, representing the function , Define bas a column vector with elements bi, representing the function , Define Aas a matrix with elements ai j, representing the operator A, The operator equation DA then corresponds to the matrix equation bDAc. In other words, the expansion of the result of applying Ato any function can be com- puted (by matrix multiplication) from the expansions of Aand . In effect, that means that the operator Acan be thought of as completely defined by its matrix elements, while andDA are completely characterized by their coefficients. ArfKen_Ch05-9780123846549.tex 280 Chapter 5 Vector Spaces We obtain an interesting expression if we introduce Dirac notation for all the quantities entering Eq. (5.53). We then have, moving the ket representing 'to the left, A DX j'ih'jAj'ih'j i; (5.55) which leads us to identify Aas ADX j'ih'jAj'ih'j; (5.56) which we note is nothing other than A, multiplied on each side by a resolution of the identity, of the form given in Eq. (5.32). Another interesting observation results if we reintroduce into Eq. (5.56) the coefficient a, bringing us to ADX j'iah'j: (5.57) Here we have the general form for an operator A, with a specific behavior that is deter- mined entirely by the set of coefficients a. The special case AD1has already been seen to be of the form of Eq. (5.57) with aD. Example 5.3.4 MATRIX ELEMENTS OF AN OPERATOR Consider the expansion of the operator xin a basis consisting of functions 'n.x/D CnHn.x/ex2=2,nD0;1;::: , where the Hnare Hermite polynomials, with scalar product hfjgiD1Z 1f.x/g.x/dx: From Table 5.1, we can see that the 'nare orthogonal and that they will also be normalized ifCnD.2nnWp/1=2. The matrix elements of x, which we denote xand are written collectively as a matrix denoted x, are given by xDh'jxj'iDCC1Z 1H.x/x H.x/ex2dx: The integral leading to xcan be evaluated in general by using the properties of the Hermite polynomials, but our present purposes are adequately served by a straightfor- ward case-by-case computation. From the table of Hermite polynomials in Table 18.1, we identify H0D1; H1D2x;H2D4x22; H3D8x312x; :::; and we take note of the integration formula InD1Z 1x2nex2dxD.2n1/WWp 2n: ArfKen_Ch05-9780123846549.tex 5.3 Operators 281 Making use of the parity (even/odd symmetry) of the Hnand the fact that the matrix x is symmetric, we note that many matrix elements are either zero or equal to others. We illustrate with the explicit computation of one matrix element, x12: x12DC1C21Z 1.2x/x.4x22/ex2dxDC1C21Z 1 8x44x2 ex2dx DC1C2h 8I24I1i D1: Evaluating other matrix elements, we find that x, the matrix of x, has the form xD0 BBBBBBB@0p 2=2 0 0  p 2=2 0 1 0  0 1 0p 6=2 0 0p 6=2 0     1 CCCCCCCA: (5.58)  Basis Expansion of Adjoint We now look at the adjoint of our operator Aas an expansion in the same basis. Our starting point is the definition of the adjoint. For arbitrary functions and, h jAjiDh A† jiDhjA†j i; where we reached the last member of the equation by using the complex conjugation prop- erty of the scalar product. This is equivalent to hjA†j iDh jAjiD" h j X j'iah'j! ji# DX h j'ia h'ji DX hj'ia h'j i; (5.59) where in the last line we have again used the scalar product complex conjugation property and have reordered the factors in the sum. We are now in a position to note that Eq. (5.59) corresponds to A†DX j'ia h'j: (5.60) In writing Eq. (5.60) we have changed the dummy indices to make the formula as similar as possible to Eq. (5.57). It is important to note the differences: The coefficient aof Eq. (5.57) has been replaced by a , so we see that the index order has been reversed and ArfKen_Ch05-9780123846549.tex 282 Chapter 5 Vector Spaces the complex conjugate taken. This is the general recipe for forming the basis set expansion of the adjoint of an operator. The relation between the matrix elements of Aand of A† is exactly that which relates a matrix Ato its adjoint A†, showing that the similarity in nomenclature is purposeful. We thus have the important and general result: IfAis the matrix representing an operator A, then the operator A†, the adjoint of A, is represented by the matrix A†. Example 5.3.5 ADJOINT OF SPIN OPERATOR Consider a spin space spanned by functions we call and , with a scalar product com- pletely defined by the equations h j iDh j iD1,h j iD0. An operator Bis such that B D0; B D : Taking all possible linearly independent scalar products, this means that h jB iD0;h jB iD0;h jB iD1;h jB iD0: It is therefore necessary that hB† j iD0;hB† j iD0;hB† j iD1;hB† j iD0; which means that B†is an operator such that B† D ; B† D0: The above equations correspond to the matrices BD0 1 0 0 ;B†D0 0 1 0 : We see that B†is the adjoint of B, as required.  Functions of Operators Our ability to represent operators by matrices also implies that the observations made in Chapter 3 regarding functions of matrices also apply to linear operators. Thus, we have definite meanings for quantities such as exp.A/,sin.A/, orcos.A/, and can also apply to operators various identities involving matrix commutators. Important examples include the Jacobi identity (Exercise 2.2.7), and the Baker-Hausdorff formula, Eq. (2.85). Exercises 5.3.1 Show (without introducing matrix representations) that the adjoint of the adjoint of an operator restores the original operator, i.e., that .A†/†DA. ArfKen_Ch05-9780123846549.tex 5.4 Self-Adjoint Operators 283 5.3.2 UandVare two arbitrary operators. Without introducing matrix representations of these operators, show that .U V/†DV†U†: Note the resemblance to adjoint matrices. 5.3.3 Consider a Hilbert space spanned by the three functions '1Dx1,'2Dx2,'3Dx3, and a scalar product defined by hxjxiD. (a) Form the 33matrix of each of the following operators: A1D3X iD1xi@ @xi ;A2Dx1@ @x2 x2@ @x1 : (b) Form the column vector representing Dx12x2C3x3. (c) Form the matrix equation corresponding to D.A1A2/ and verify that the matrix equation reproduces the result obtained by direct application of A1A2 to . 5.3.4 (a) Obtain the matrix representation of ADx.d=dx/in a basis of Legendre polyno- mials, keeping terms through P3. Use the orthonormal forms of these polynomials as given in 5.2.1 and the scalar product defined there. (b) Expand x3in the orthonormal Legendre polynomial basis. (c) Verify that Ax3is given correctly by its matrix representation. 5.4 S ELF-ADJOINT OPERATORS Operators that are self-adjoint (Hermitian) are of particular importance in quantum mechanics because observable quantities are associated with Hermitian operators. In particular, the average value of an observable Ain a quantum mechanical state described by any normalized wave function is given by the expectation value ofA, defined as hAiDh jAj i: (5.61) This, of course, only makes sense if it can be assured that hAiis real, even if and/or Ais complex. Using the fact that Ais postulated to be Hermitian, we take the complex conjugate ofhAi: hAiDh jAj iDhA j i; which reduces tohAibecause Ais self-adjoint. We have already seen that if AandA†are expanded in a basis, the matrix A†must be the matrix adjoint of the matrix A. This means that the coefficients in its expansion must satisfy aDa  (coefficients of self-adjoint A): (5.62) ArfKen_Ch05-9780123846549.tex 284 Chapter 5 Vector Spaces Thus, we have the nearly self-evident result: A matrix representing a Hermitian operator is a Hermitian matrix. It is also obvious from Eq. (5.62) that the diagonal elements of a Hermitian matrix (which are expectation values for the basis functions) are real. We can easily verify from basis expansions that hAimust be real. Letting cbe the vector of expansion coefficients of in the basis for which aare the matrix elements of A, then hAiDh jAj iD*X c' A X c'+ DX c h'jAj'ic DX c acDc†Ac; which reduces, as it must, to a scalar. Because Ais a self-adjoint matrix, c†Acis easily seen to be a self-adjoint 11matrix, i.e., a real scalar (use the facts that .BAC/†DC†A†B†and thatA†DA). Example 5.4.1 SOME SELF-ADJOINT OPERATORS Consider the operators xandpintroduced earlier, with a scalar product of definition hfjgiD1Z 1f.x/g.x/dx; (5.63) where our Hilbert space is the set of all functions ffor whichhfjfiexists (i.e.,hfjfiis finite). This is the L2space on the interval .1;1/. To test whether xis self-adjoint, we comparehfjxgiandhx fjgi. Writing these out as integrals, we consider 1Z 1f.x/x g.x/dx vs.1Z 1Tx f.x/Ug.x/dx: Because the order of ordinary functions (including x) can be changed without affecting the value of an integral, and because xis inherently real, these two expressions are equal and xis self-adjoint. Turning next to pDi.d=dx/, the comparison we must make is 1Z 1f.x/ id g.x/ dx dx vs.1Z 1 id f.x/ dx g.x/dx: (5.64) We can bring these expressions into better correspondence if we integrate the first by parts, differentiating f.x/and integrating dg.x/=dx . Doing so, the first expression above be- comes 1Z 1f.x/ id g.x/ dx dxDi f.x/g.x/ 1 11Z 1d f.x/ dxh ig.x/i dx: ArfKen_Ch05-9780123846549.tex 5.4 Self-Adjoint Operators 285 The boundary terms to be evaluated at 1 must vanish because hfjfiandhgjgiare finite, which assures also (from the Schwarz inequality) that hfjgiis finite as well. Upon moving iwithin the complex conjugate in the remaining integral, we verify agreement with the second expression in Eq. (5.64). Thus both xandpare self-adjoint. Note that if phad not contained the factor i, it would nothave been self-adjoint, as we obtained a needed sign change when iwas moved within the scope of the complex conjugate.  Example 5.4.2 EXPECTATION VALUE OF p Because p, though Hermitian, is also imaginary, consider what happens when we compute its expectation value for a wave function of the form .x/Deif.x/, where f.x/is a realL2wave function and is a real phase angle. Using the scalar product as defined in Eq. (5.63), and remembering that pDi.d=dx/, we have hpiD i1Z 1f.x/d f.x/ dxdxDi 21Z 1d dxh f.x/i2 dx Di 2h f.C1/2f.1/2i D0: As shown, this integral vanishes because f.x/D0at1 (this is fortunate because expec- tation values must be real). This result corresponds to the well-known property that wave functions that describe time-dependent phenomena (nonzero momentum) cannot either be real or real except for a constant (complex) phase factor.  The relations between operators and their adjoints provide opportunities for rearrange- ments of operator expressions that may facilitate their evaluation. Some examples follow. Example 5.4.3 OPERATOR EXPRESSIONS (a) Suppose we wish to evaluate h.x2Cp2/ j'i, with of a complicated functional form that might be unpleasant to differentiate (as required to apply p2), whereas'is simple. Because xis self-adjoint, so also is x2: hx2 j'iDhx jx'iDh jx2'i: The same is true of p2, soh.x2Cp2/ j'iDh j.x2Cp2/'i. (b) Look next ath.xCip/ j.xCip/ i, which is the expression to be evaluated if we want the norm of .xCip/ . Note that xCipisnotself-adjoint, but has adjoint xip. Our norm rearranges to h.xCip/ j.xCip/ iDh j.xip/.xCip/j i Dh jx2Cp2Ci.xppx/j i Dh jx2Cp2Ci.i/j iDh jx2Cp21j i: ArfKen_Ch05-9780123846549.tex 286 Chapter 5 Vector Spaces To reach the last line of the above equation, we recognized the commutator Tx;pUDi, as established in Eq. (5.43). (c) Suppose that AandBare self-adjoint. What can we say about the self-adjointness of AB? Consider h jABj'iDh A jBj'iDh B A j'i: Note that because we moved Ato the left first (with no dagger needed because it is self-adjoint), it is part of what the subsequently moved Bmust operate on. So we see that the adjoint of ABisB A. We conclude that ABis only self-adjoint if AandB commute (so that B ADAB). Note that if AandBwere not individually self-adjoint, their commutation would not be sufficient to make ABself-adjoint.  Exercises 5.4.1 (a) Ais a non-Hermitian operator. Show that the operators ACA†andi.AA†/are Hermitian. (b) Using the preceding result, show that every non-Hermitian operator may be written as a linear combination of two Hermitian operators. 5.4.2 Prove that the product of two Hermitian operators is Hermitian if and only if the two operators commute. 5.4.3 AandBare noncommuting quantum mechanical operators, and Cis given by the formula ABB ADiC: Show that Cis Hermitian. Assume that appropriate boundary conditions are satisfied. 5.4.4 The operator Lis Hermitian. Show that hL2i0, meaning that for all in the space in which Lis defined,h jL2j i0. 5.4.5 Consider a Hilbert space whose members are functions defined on the surface of the unit sphere, with a scalar product of the form hfjgiDZ dfg; where dis the element of solid angle. Note that the total solid angle of the sphere is 4. We work here with the three functions '1DCx=r,'2DCy=r,'3DCz=r, with C assigned a value that makes the 'inormalized. (a) Find C, and show that the 'iare also mutually orthogonal. (b) Form the 33matrices of the angular momentum operators LxDi y@ @zz@ @y ;LyDi z@ @xx@ @z ; LzDi x@ @yy@ @x : ArfKen_Ch05-9780123846549.tex 5.5 Unitary Operators 287 (c) Verify that the matrix representations of the components of Lsatisfy the angular momentum commutator TLx;LyUDi Lz. 5.5 U NITARY OPERATORS One of the reasons unitary operators are important in physics is that they can be used to describe transformations between orthonormal bases. This property is the generalization to the complex domain of the rotational transformations of ordinary (physical) vectors that we analyzed in Chapter 3. Unitary Transformations Suppose we have a function that has been expanded in the orthonormal basis ': DX c'D X j'ih'j! j i: (5.65) We now wish to convert this expansion to a different orthonormal basis, with functions '0 . A possible starting point is to recognize that each of the original basis functions can be expanded in the primed basis. We can obtain the expansion by inserting a resolution of the identity in the primed basis: 'DX u'0 D X j'0 ih'0 j! j'iDX h'0 j'i'0 : (5.66) Comparing the second and fourth members of this equation, we identify uas the ele- ments of a matrix U: uDh'0 j'i: (5.67) Note how the use of resolutions of the identity makes these formulas obvious, and that Eqs. (5.65) to(5.67) are only valid because the 'and the'0 are complete orthonormal sets. Inserting the expansion for 'from Eq. (5.66) intoEq. (5.65), we reach DX cX u'0 DX  X uc! '0 Dc0 '0 ; (5.68) where the coefficients c0 of the expansion in the primed basis form a column vector c0that is related to the coefficient vector cin the unprimed basis by the matrix equation c0DUc; (5.69) with Uthe matrix whose elements are given in Eq. (5.67). If we now consider the reverse transformation, from an expansion in the primed basis toone in the unprimed basis, starting from '0 DX v'DX h'j'0 i'; (5.70) ArfKen_Ch05-9780123846549.tex 288 Chapter 5 Vector Spaces we see that V, the matrix of the transformation inverse to U, has elements vDh'j'0 iD.U/D.U†/: (5.71) In other words, VDU†: (5.72) If we now transform the expansion of , given in the unprimed basis by the coefficient vector c, first to the primed basis and then back to the original unprimed basis, the coeffi- cients will transform, first to c0and then back to c, according to cDV UcDU†Uc: (5.73) In order for Eq. (5.73) to be consistent it is necessary that U†Ube a unit matrix, meaning thatUmust be unitary. We thus have the important following result: The transformation that converts the expansion of a vector cin any orthonormal basis f'gto its expansion c0in any other orthonormal basis f'0gis described by the matrix equation c0DUc, where the transformation matrix Uisunitary and has ele- ments uDh'0j'i. A transformation between orthonormal bases is called a unitary transformation. Equation (5.69) is a direct generalization of the ordinary 2-D vector rotational transfor- mation equation, Eq. (3.26), A0DSA: For further emphasis, we compare the transformation matrix Uintroduced here (at right, below) with the matrix S(at left) from Eq. (3.28), for rotations in ordinary 2-D space: SD Oe0 1Oe1Oe0 1Oe2 Oe0 2Oe1Oe0 2Oe2! UD0 @h'0 1j'1i h'0 1j'2i  h'0 2j'1i h'0 2j'2i    1 A: The resemblance becomes even more striking if we recognize that in Dirac notation, the quantitiesOe0 iOejassume the formhOe0 ijOeji. As for ordinary vectors (except that the quantities involved here are complex), the ith row of Ucontains the (complex conjugated) components (a.k.a. coefficients) of '0 iin terms of the unprimed basis; the orthonormality of the primed 'is consistent with the fact that U U†is a unit matrix. The columns of Ucontain the components of the 'jin terms of the primed basis; that also is analogous to our earlier observations. The matrix Sis orthogonal; Uis unitary, which is the generalization to a complex space of the orthogonality condition. Summarizing, we see that unitary transformations are analogous, in vector spaces, to the orthogonal transformations that describe rotations (or reflections) in ordinary space. ArfKen_Ch05-9780123846549.tex 5.5 Unitary Operators 289 Example 5.5.1 A UNITARY TRANSFORMATION A Hilbert space is spanned by five functions defined on the surface of a unit sphere and expressed in spherical polar coordinates ;': 1Dr 15 4sincoscos'; 2Dr 15 4sincossin'; 3Dr 15 4sin2sin'cos'; 4Dr 15 16sin2.cos2'sin2'/; 5Dr 5 16.3 cos21/: These are orthonormal when the scalar product is defined as hfjgiDZ 0sind2Z 0d'f.;'/ g.;'/: This Hilbert space can alternatively be spanned by the orthonormal set of functions 0 1Dr 15 8sincosei'; 0 2Dr 15 8sincosei'; 0 3Dr 15 32sin2e2i'; 0 4Dr 15 32sin2e2i'; 0 5D5: The matrix Udescribing the transformation from the unprimed to the primed basis has elements uDh0 ji. Working out a representative matrix element, u22Dh0 2j2iD15 4p 2Z 0sind2Z 0d'sin2cos2eCi'sin' D15 4p 2Z 0sin3cos2d2Z 0d'eCi'ei'ei' 2i D15 4p 24 152 2iDip 2: ArfKen_Ch05-9780123846549.tex 290 Chapter 5 Vector Spaces We obtained this result by using the formulaR2 0eni'd'D2 n0and by looking up a tabulated value for the integral. We evaluate explicitly one more matrix element: u21Dh0 2j1iD15 4p 2Z 0sin3cos2d2Z 0d'eCi'ei'Cei' 2 D15 4p 24 152 2D1p 2: Evaluating the remaining elements of U, we reach UD0 BBBBB@1=p 2i=p 2 0 0 0 1=p 2i=p 2 0 0 0 0 0 i=p 2 1=p 2 0 0 0i=p 2 1=p 2 0 0 0 0 0 11 CCCCCA: As a check, note that the ith column of Ushould yield the components of iin the primed basis. For the first column, we have r 15 4sincoscos'D1p 2 r 15 8sincosei'! C1p 2 r 15 8sincosei'! ; which simplifies easily to an identity. Further checks are left as Exercise 5.5.1.  Successive Transformations It is possible to make two or more successive unitary transformations, each of which will convert an input orthonormal basis to an output basis that is also orthonormal. Just as for ordinary vectors, the successive transformations are applied in right-to-left order, and the product of the transformations can be viewed as a resultant unitary transformation. Exercises 5.5.1 Show that the matrix Uof Example 5.5.1 correctly transforms the vector f.;'/D 31C2i23C5to thef0 igbasis by (a) (1) Making a column vector cthat represents f.;'/ in thefigbasis, (2) forming c0DUc, and (3) comparing the expansionP ic0 i0 i.;'/ with f.;'/ ; (b) Verifying that Uis unitary. 5.5.2 (a) Given (in R3) the basis'1Dx,'2Dy,'3Dz, consider the basis transformation x!z,y!y,z! x. Find the 33matrix Ufor this transformation. ArfKen_Ch05-9780123846549.tex 5.5 Unitary Operators 291 (b) This transformation corresponds to a rotation of the coordinate axes. Identify the rotation and reconcile your transformation matrix with an appropriate matrix S. ; ; / of the form given in Eq. (3.37). (c) Form the column vector crepresenting (in the original basis) fD2x3yCz, find the result of applying Utoc, and show that this is consistent with the basis transformation of part (a). Note. You do not need to be able to form scalar products to handle this exercise; a knowledge of the linear relationship between the original and transformed functions is sufficient. 5.5.3 Construct the matrix representing the inverse of the transformation in Exercise 5.5.2, and show that this matrix and the transformation matrix of that exercise are matrix inverses of each other. 5.5.4 The unitary transformation Uthat converts an orthonormal basis f'iginto the basisf'0 ig and the unitary transformation Vthat converts the basis f'0 iginto the basisfighave matrix representations UD0 @isin cos0 cosisin0 0 0 11 A;VD0 @1 0 0 0 cosisin 0 cosisin1 A: Given the function f.x/D3'1.x/'2.x/2'3.x/, (a) By applying U, form the vector representing f.x/in thef'0 igbasis and then by applying Vform the vector representing f.x/in thefigbasis. Use this result to write f.x/as a linear combination of the i. (b) Form the matrix products UVandVUand then apply each to the vector represent- ingf.x/in thef'igbasis. Verify that the results of these applications differ and that only one of them gives the result corresponding to part (a). 5.5.5 Three functions which are orthogonal with unit weight on the range 1x1are P0D1,P1Dx, and P2D3 2x21 2. Another set of functions that are orthogonal and span the same space are F0Dx2,F1Dx,F2D5x23. Although much of this exercise can be done by inspection, write down and evaluate all the integrals that lead to the results when they are obtained in terms of scalar products. (a) Normalize each of the PiandFi. (b) Find the unitary matrix Uthat transforms from the normalized Pibasis to the normalized Fibasis. (c) Find the unitary matrix Vthat transforms from the normalized Fibasis to the normalized Pibasis. (d) Show that UandVare unitary, and that VDU1. (e) Expand f.x/D5x23xC1in terms of the normalized versions of both bases, and verify that the transformation matrix Uconverts the P-basis expansion of f.x/into its F-basis expansion. ArfKen_Ch05-9780123846549.tex 292 Chapter 5 Vector Spaces 5.6 T RANSFORMATIONS OF OPERATORS We have seen how unitary transformations can be used to transform the expansion of a function from one orthonormal basis set to another. We now consider the corresponding transformation for operators. Given an operator A, which when expanded in the 'basis has the form ADX j'iah'j; we convert it to the '0basis by the simple expedient of inserting resolutions of the identity (written in terms of the primed basis) on both sides of the above expression. This is an excellent example of the benefits of using Dirac notation. Remembering that this does not change A(but of course does change its appearance), we get ADX j'0 ih'0 j'iah'j'0 ih'0 j; which we simplify by identifying h'0 j'iDu, as defined in Eq. (5.67), andh'j'0 iD u . Thus, ADX j'0 iuau h'0 jDX j'0 ia0 h'0 j; (5.74) where a0 is thematrix element of Ain the primed basis, related to the unprimed values by a0 DX uau : (5.75) If we now note that u D.U†/, we can write Eq. (5.75) as the matrix equation A0DU A U†DU A U1; (5.76) where in the final member of the equation we used the fact that Uis unitary. Another way of getting at Eq. (5.76) is to consider the operator equation A D, where initially A, , andare all regarded as expanded in the orthonormal set ', with Ahaving matrix elements a, and with andhaving the forms DP c'andDP b'. This state of affairs corresponds to the matrix equation AcDb: Now we simply insert U1Ubetween Aandc, and multiply both sides of the equation on the left by U. The result is  U A U1 Uc DUb! A0c0Db0; (5.77) showing that the operator and the functions are properly related when the functions have been transformed by applying Uand the operator has been transformed as required by Eq. (5.76). Since this relationship is valid for any choice of candU, it confirms the trans- formation equation for A. ArfKen_Ch05-9780123846549.tex 5.6 Transformations of Operators 293 Nonunitary Transformations It is possible to consider transformations similar to that illustrated by Eq. (5.77), but using a transformation matrix Gthat must be nonsingular, but is not required to be unitary. Such more general transformations occasionally appear in physics applications, are called simi- larity transformations, and lead to an equation deceptively similar to Eq. (5.77):  G A G1 Gc DGb: (5.78) There is one important difference: Although a general similarity transformation preserves the original operator equation, corresponding items do not describe the same quantity in a different basis. Instead, they describe quantities that have been systematically (but consis- tently) altered by the transformation. Sometimes we encounter a need for transformations that are not even similarity trans- formations. For example, we may have an operator whose matrix elements are given in a nonorthogonal basis, and we consider the transformation to an orthonormal basis generated by use of the Gram-Schmidt procedure. Example 5.6.1 GRAM-SCHMIDT TRANSFORMATION The Gram-Schmidt process describes the transformation from an initial function set ito an orthonormal set 'according to equations that can be brought to the form 'DX iD1tii; D1;2;:::: Because the Gram-Schmidt process only generates coefficients tiwith i, the transfor- mation matrix Tcan be described as upper triangular, i.e., a square matrix with nonzero elements tionly on and above its principal diagonal. Defining Sas a matrix with elements si jDhijji(often called an overlap matrix), the orthonormality of the 'is evidenced by the equation h'j'iDX i jhtiijtjjiDX i jt ihijjitjD.T†ST/D: (5.79) Note that because Tis upper triangular, T†must be lower triangular. In writing Eq. (5.79) we did not have to restrict the iandjsummations, as the coefficients outside the contribut- ing ranges of iandjare present, but set to zero. From Eq. (5.79) we can obtain a representation of S: SD.T†/1T1D.TT†/1: (5.80) Moreover, if we replace Sfrom Eq. (5.79) by the matrix of a general operator A(in thei basis), we find that in the orthonormal 'basis its representation A0is A0DT†AT: (5.81) In general, T†will not be equal to T1, so this equation does not define a similarity trans- formation.  ArfKen_Ch05-9780123846549.tex 294 Chapter 5 Vector Spaces Exercises 5.6.1 (a) Using the two spin functions '1D and'2D as an orthonormal basis (so h j iDh j iD1,h j iD0), and the relations Sx D1 2 ;Sx D1 2 ;Sy D1 2i ;Sy D1 2i ;Sz D1 2 ;Sz D1 2 ; construct the 22matrices of Sx,Sy, and Sz. (b) Taking now the basis '0 1DC. C /,'0 2DC. /: (i) Verify that '0 1and'0 2are orthogonal, (ii) Assign Ca value that makes '0 1and'0 2normalized, (iii) Find the unitary matrix for the transformation f'ig!f'0 ig. (c) Find the matrices of Sx,Sy, and Szin thef'0 igbasis. 5.6.2 For the basis '1DCxer2,'2DCyer2,'3DCzer2, where r2Dx2Cy2Cz2, with the scalar product defined as an unweighted integral over R3and with Cchosen to make the'inormalized: (a) Find the 33matrix of LxDi y@ @zz@ @y ; (b) Using the transformation matrix UD0 B@1 0 0 0 1=p 2i=p 2 0 1=p 2i=p 21 CA, find the trans- formed matrix of Lx; (c) Find the new basis functions '0 idefined by the transformation U, and write explic- itly (in terms of x,y, and z) the functional forms of Lx'0 i,iD1;2;3. Hint. UseR er2d3rD3=2,R x2er2d3rD1 23=2; the integrals are over R3. 5.6.3 The Gram-Schmidt process for converting an arbitrary basis into an orthonormal set'is described in Section 5.2 in a way that introduces coefficients of the form h'ji. For bases consisting of three functions, convert the formulation so that ' is expressed entirely in terms of the , thereby obtaining an expression for the upper- triangular matrix Tappearing in Eq. (5.81). 5.7 I NVARIANTS Just as coordinate rotations leave invariant the essential properties of physical vectors, we can expect unitary transformations to preserve essential features of our vector spaces. These invariances are most directly observed in the basis-set expansions of operators and functions. Consider first a matrix equation of the form bDAc, where all quantities have been evaluated using a particular orthonormal basis 'i. Now suppose that we wish to use a basisiwhich can be reached from the original basis by applying a unitary transformation ArfKen_Ch05-9780123846549.tex 5.7 Invariants 295 such that c0DUcand b0DUb: In the new basis, the matrix Abecomes A0DUAU1, and the invariance we seek corre- sponds to b0DA0c0. In other words, all the quantities must change coherently so that their relationship is unaltered. It is easy to verify that this is the case. Substituting for the primed quantities, UbD.UAU1/.Uc/! UbDU Ac; from which we can recover bDAcby multiplying from the left by U1. Scalar quantities should remain invariant under unitary transformation; the prime example here is the scalar product. If fandgare represented in some orthonormal basis, respectively, by aandb, their scalar product is given by a†b. Under a unitary transfor- mation whose matrix representation is U,abecomes a0DUaandbbecomes b0DUb, and hfjgiD.a0/†b0D.Ua/†.Ub/D.a†U†/.Ub/Da†b: (5.82) The fact that U†DU1enables us to confirm the invariance. Another scalar that should remain invariant under basis transformation is the expectation value of an operator. Example 5.7.1 EXPECTATION VALUE IN TRANSFORMED BASIS Suppose that DX ici'i, and that we wish to compute the expectation value of Afor this , where A, the matrix corresponding to A, has elements aDh'jAj'i. We have hAiDh jAj i ! c†Ac: If we now choose to use a basis obtained from the 'iby a unitary transformation U, the expression forhAibecomes .Uc/†.UAU1/.Uc/Dc†U†UAU1Uc; which, because Uis unitary and therefore U†DU1, reduces, as it must, to the previously obtained value ofhAi.  Vector spaces have additional useful matrix invariants. The trace of a matrix is invariant under unitary transformation. If A0DU A U1, then trace.A0/DX .U A U1/DX ua.U1/DX  X .U1/u! a DX aDX aDtrace.A/: (5.83) Here we simply used the property U1UD1. Another matrix invariant is the determinant. From the determinant product theorem, det.U A U1/Ddet.U1U A/Ddet.A/ . Further invariants will be identified when we study matrix eigenvalue problems in Chapter 6. ArfKen_Ch05-9780123846549.tex 296 Chapter 5 Vector Spaces Exercises 5.7.1 Using the formal properties of unitary transformations, show that the commutator Tx;pUDiis invariant under unitary transformation of the matrices representing xandp. 5.7.2 The Pauli matrices 1D0 1 1 0 ;2D0i i0 ;3D1 0 01 ; have commutatorT1;2UD2i3. Show that this relationship continues to be valid if these matrices are transformed by UD cossin sincos! : 5.7.3 (a) The operator Lxis defined as LxDi y@ @zz@ @y : Verify that the basis '1DCxer2,'2DCyer2,'3DCzer2, where r2Dx2C y2Cz2, forms a closed set under the operation of Lx, meaning that when Lxis applied to any member of this basis the result is a function within the basis space, and construct the 33matrix of Lxin this basis from the result of the application ofLxto each basis function. (b) Verify that Lxh .xCiy/er2i Dzer2, and note that this result, using the f'ig basis, can be written Lx.'1Ci'2/D' 3. (c) Express the equation of part (b) in matrix form, and write the matrix equation that results when each of the quantities is transformed using the transformation matrix UD0 BB@1 0 0 0 1=p 2i=p 2 0 1=p 2i=p 21 CCA: (d) Regarding the transformation Uas producing a new basis f'0 ig, find the explicit form (in x,y,z) of the'0 i. (e) Using the operator form of Lxand the explicit forms of the '0 i, verify the validity of the transformed equation found in part (c). Hint. The results of Exercise 5.6.2 may be useful. 5.8 S UMMARY —V ECTOR SPACE NOTATION It may be useful to summarize some of the relationships found in this chapter, highlighting the essentially complete mathematical parallelism between the properties of vectors and those of basis expansions in vector spaces. We do so here, using Dirac notation wherever appropriate. ArfKen_Ch05-9780123846549.tex 5.8 Summary—Vector Space Notation 297 1.Scalar product: h'j iDbZ a'.t/ .t/w.t/dt()hujviD u†vDuv: (5.84) The result of the scalar product operation is a scalar (i.e., a real or complex number). Here u†vrepresents the product of a row and a column vector; it is equivalent to the dot-product notation also shown. 2.Expectation value: h'jAj'iDbZ a'.t/A'.t/w.t/dt()hujAjuiD u†Au: (5.85) 3.Adjoint: h'jAj iDh A†'j i()hujAjviDhA†ujviDTA†uU†vDu†Av: (5.86) Note that the simplification of TA†uU†vshows that the matrix A†has the property expected of an operator adjoint. 4.Unitary transformation: DA'!U D.U AU1/.U'/() wDAv! UwD.UAU1/.Uv/: (5.87) 5.Resolution of identity: 1DX ij'iih'ij() 1DX ijOeiihOeij; (5.88) where the'iare orthonormal and the Oeiare orthogonal unit vectors. Applying Eq. (5.88) to a function (or vector): DX ij'iih'ij iDX iai'i() wDX ijOeiihOeijwiDX iwiOei; (5.89) where aiDh' ij iandwiDhOeijwiDOeiw. Additional Readings Brown, W. A., Matrices and Vector Spaces. New York: M. Dekker (1991). Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969), reprinting, Dover (1992). Dennery, P., and A. Krzywicki, Mathematics for Physicists. New York: Harper & Row, reprinting, Dover (1996). Halmos, P. R., Finite-Dimensional Vector Spaces, 2nd ed. Princeton, NJ: Van Nostrand (1958), reprinting, Springer (1993). Jain, M. C., Vector Spaces and Matrices in Physics, 2nd ed. Oxford: Alpha Science International (2007). Kreyszig, E., Advanced Engineering Mathematics, 6th ed. New York: Wiley (1988). Lang, S., Linear Algebra. Berlin: Springer (1987). Roman, S., Advanced Linear Algebra, Graduate Texts in Mathematics 135, 2nd ed. Berlin: Springer (2005). ArfKen_Ch06-9780123846549.tex CHAPTER 6 EIGENVALUE PROBLEMS 6.1 E IGENVALUE EQUATIONS Many important problems in physics can be cast as equations of the generic form A D ; (6.1) where Ais a linear operator whose domain and range is a Hilbert space, is a function in the space, and is a constant. The operator Ais known, but both andare unknown, and the task at hand is to solve Eq. (6.1). Because the solutions to an equation of this type yield functions that are unchanged by the operator (except for multiplication by a scale factor), they are termed eigenvalue equations: Eigen is German for “[its] own.” A function that solves an eigenvalue equation is called an eigenfunction, and the value of that goes with an eigenfunction is called an eigenvalue. The formal definition of an eigenvalue equation may not make its essential content totally apparent. The requirement that the operator Aleaves unchanged except for a scale factor constitutes a severe restriction upon . The possibility that Eq. (6.1) has any solutions at all is in many cases not intuitively obvious. To see why eigenvalue equations are common in physics, let’s cite a few examples: 1. The resonant standing waves of a vibrating string will be those in which the restor- ing force on the elements of the string (represented by A ) are proportional to their displacements from equilibrium. 2. The angular momentum Land the angular velocity !of a rigid body are three- dimensional (3-D) vectors that are related by the equation LDI!; where Iis the 33moment of inertia matrix. Here the direction of !defines the axis of rotation, while the direction of Ldefines the axis about which angular momentum is 299 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch06-9780123846549.tex 300 Chapter 6 Eigenvalue Problems generated. The condition that these two axes be in the same direction (thereby defin- ing what are known as the principal axes of inertia) is that LD!, whereis a proportionality constant. Combining with the formula for L, we obtain I!D!; which is an eigenvalue equation in which the operator is the matrix Iand the eigen- function (then usually called an eigenvector) is the vector !. 3. The time-independent Schrödinger equation in quantum mechanics is an eigenvalue equation, with Athe Hamiltonian operator H, a wave function and DEthe energy of the state represented by . Basis Expansions A powerful approach to eigenvalue problems is to express them in terms of an orthonormal basis whose members we designate 'i, using the formulas developed in Chapter 5. Then the operator Aand the function are represented by a matrix Aand a vector cwhose elements are obtained, according to Eqs. (5.51) and (5.52), as the scalar products ai jDh' ijAj'ji;ciDh' ij i: Our original eigenvalue equation has now been reduced to a matrix equation: AcDc: (6.2) When an eigenvalue equation is presented in this form, we can call it a matrix eigenvalue equation and call the vectors cthat solve it eigenvectors. As we shall see in later sections of this chapter, there is a well-developed technology for the solution of matrix eigenvalue equations, so a route always available for solving eigenvalue equations is to cast them in matrix form. Once a matrix eigenvalue problem has been solved, we can recover the eigenfunctions of the original problem from their expansion: DX ici'i: Sometimes, as in the moment of inertia example mentioned above, our eigenvalue prob- lem originates as a matrix problem. Then, of course, we do not have to begin its solution process by introducing a basis and converting it into matrix form, and our solutions will be vectors that do not need to be interpreted as expansions in a basis. Equivalence of Operator and Matrix Forms It is important to note that we are dealing with eigenvalue equations in which the operator involved is linear and that it operates on elements of a Hilbert space. Once these conditions are met, the operator and function involved can always be expanded in a basis, leading to a matrix eigenvalue equation that is totally equivalent to our original problem. Among other things, this means that any theorems about properties of eigenvectors or eigenvalues that are developed from basis-set expansions of an eigenvalue problem must apply also to the original problem, and that solution of the matrix eigenvalue equation provides also a ArfKen_Ch06-9780123846549.tex 6.2 Matrix Eigenvalue Problems 301 solution to the original problem. These facts, plus the practical observation that we know how to solve matrix eigenvalue problems, strongly suggest that the detailed investigation of the matrix problems should be on our agenda. When we explore matrix eigenvalue problems, we will find that certain properties of the matrix influence the nature of the solutions, and that in particular significant simplifications become available when the matrix is Hermitian. Many eigenvalue equations of interest in physics involve differential operators, so it is of importance to understand whether (or under what conditions) these operators are Hermitian. That issue is taken up in Chapter 8. Finally, we note that the introduction of a basis-set expansion is not the only possibility for solving an eigenvalue equation. Eigenvalue equations involving differential operators can also be approached by the general methods for solving differential equations. That topic is also discussed in Chapter 8. 6.2 M ATRIX EIGENVALUE PROBLEMS While in principle the notion of an eigenvalue problem is already fully defined, we open this section with a simple example that may help to make it clearer how such problems are set up and solved. A Preliminary Example We consider here a simple problem of two-dimensional (2-D) motion in which a particle slides frictionlessly in an ellipsoidal basin (see Fig. 6.1). If we release the particle (initially at rest) at an arbitrary point in the basin, it will start to move downhill in the (negative) gradient direction, which in general will not aim directly at the potential minimum at the bottom of the basin. The particle’s overall trajectory will then be a complicated path, as sketched in the bottom panel of Fig. 6.1. Our objective is to find the positions, if any, from which the trajectories will aim at the potential minimum, and will therefore represent simple one-dimensional oscillatory motion. This problem is sufficiently elementary that we can analyze it without great difficulty. We take a potential of the form V.x;y/Dax2CbxyCcy2; with parameters a,b,cin ranges that describe an ellipsoidal basin with a minimum in V atxDyD0. We then calculate the xandycomponents of the force on our particle when at.x;y/: FxD@V @xD2 axby;FyD@V @yDbx2cy: It is pretty clear that, for most values of xandy,Fx=Fy6Dx=y, so the force will not be directed toward the minimum at xDyD0. To search for directions in which the force is directed toward xDyD0, we begin by writing the equations for the force in matrix form: Fx Fy D2ab b2cx y ;orfDHr; ArfKen_Ch06-9780123846549.tex 302 Chapter 6 Eigenvalue Problems –505 y −10 0 10 x −505 y −10 0 10 x FIGURE 6.1Top: Contour lines of basin potential VDx2p 5xyC3y2. Bottom: Trajectory of sliding particle of unit mass starting from rest at .8:0;1:92/ . where f,H, and rare defined as indicated. Now the condition Fx=FyDx=yis equivalent to the statement that fandrare proportional, and therefore we can write HrDr; (6.3) where, as already suggested, His a known matrix, while andrare to be determined. This is an eigenvalue equation, and the column vectors rthat are its solutions are its eigenvec- tors, while the corresponding values of are its eigenvalues. Equation (6.3) is a homogeneous linear equation system, as becomes more obvious if written as .H1/rD0; (6.4) and we know from Chapter 2 that it will have the unique solution rD0unless det.H 1/D0. However, the value of is at our disposal, so we can search for values of that cause this determinant to vanish. Proceeding symbolically, we look for such that det.H1/D h11 h12 h21 h22 D0: Expanding the determinant, which is sometimes called a secular determinant (the name arising from early applications in celestial mechanics), we have an algebraic equation, the secular equation, .h11/.h22/h12h21D0; (6.5) ArfKen_Ch06-9780123846549.tex 6.2 Matrix Eigenvalue Problems 303 which can be solved for . The left hand side of Eq. (6.5) is also called the characteristic polynomial (in) ofH, and Eq. (6.5) is for that reason also known as the characteristic equation ofH. Once a value of that solves Eq. (6.5) has been obtained, we can return to the homo- geneous equation system, Eq. (6.4), and solve it for the vector r. This can be repeated for allthat are solutions to the secular equation, thereby giving a set of eigenvalues and the associated eigenvectors. Example 6.2.1 2-D ELLIPSOIDAL BASIN Let’s continue with our ellipsoidal basin example, with the specific parameter values aD1, bDp 5,cD3. Then our matrix Hhas the form HD2p 5p 56 ; and the secular equation takes the form det.H1/D 2p 5p 56 D2C8C7D0: Since2C8C7D.C1/.C7/, we see that the secular equation has as solutions the eigenvaluesD1 andD7 . To get the eigenvector corresponding to D1 , we return to Eq. (6.4), which, written in great detail, is .H1/rD2.1/p 5p 56.1/x y D1p 5p 55x y D0; which expands into a linearly dependent pair of equations: xCp 5yD0 p 5x5yD0: This is, of course, the intention associated with the secular equation, because if these equa- tions were linearly independent they would inexorably lead to the solution xDyD0. Instead, from either equation, we have xDp 5y, so we have the eigenvalue/eigenvector pair 1D1; r1DCp 5 1 ; where Cis a constant that can assume any value. Thus, there is an infinite number of x;y pairs that define a direction in the 2-D space, with the magnitude of the displacement in that direction arbitrary. The arbitrariness of scale is a natural consequence of the fact that the equation system was homogeneous; any multiple of a solution of a linear homogeneous equation set will also be a solution. This eigenvector corresponds to trajectories that start from the particle at rest anywhere on the line defined by r1. A trajectory of this sort is illustrated in the top panel of Fig. 6.2. ArfKen_Ch06-9780123846549.tex 304 Chapter 6 Eigenvalue Problems −505 y −10 0 10 x −505 y −10 0 10 x FIGURE 6.2 Trajectories starting at rest. Top: At a point on the line xDyp 5. Bottom: At a point on the line yDxp 5. Wehave not yet considered the possibility that D7 . This leads to a different eigen- vector, obtained by solving .H1/rD2C7p 5p 56C7x y D5p 5p 51x y D0; corresponding to yDxp 5.This defines the eigenvalue/eigenvector pair 2D7; r2DC01p 5 : Atrajectory of this sort is shown in the bottom panel of Fig. 6.2. We thus have two directions in which the force is directed toward the minimum, and they are mutually perpendicular: the first direction has dy=dxD1=p 5;for the second, dy=dxDp 5. Wecan easily check our eigenvectors and eigenvalues. For 1andr1, Hr1D2p 5p 56 Cp 5 C DC p 5 1 D.1/ Cp 5 C D1r1: Itis often useful to normalize eigenvectors, which we can do by choosing the constant (CorC0) to make rof magnitude unity. In the present example, r1D p5=6 p1=6! ;r2D p1=6 p5=6! : (6.6) ArfKen_Ch06-9780123846549.tex 6.2 Matrix Eigenvalue Problems 305 Each of these normalized eigenvectors is still arbitrary as to overall sign (or if we accept complex coefficients, as to an arbitrary complex factor of magnitude unity). Before leaving this example, we make three further observations: (1) the number of eigenvalues was equal to the dimension of the matrix H. This is a consequence of the fundamental theorem of algebra, namely that an equation of degree nwill have nroots; (2) although the secular equation was of degree 2 and quadratic equations can have com- plex roots, our eigenvalues were real; and (3) our two eigenvectors are orthogonal.  Our 2-D example is easily understood physically. The directions in which the displace- ment and the force are collinear are the symmetry directions of the elliptical potential field, and they are associated with different eigenvalues (the proportionality constant between position and force) because the ellipses have axes of different lengths. We have, in fact, identified the principal axes of our basin. With the parameters of Example 6.2.1, the poten- tial could have been written (using the normalized eigenvectors) VD1 2 p 5xCyp 6!2 C7 2 xp 5yp 6!2 D1 2.x0/2C7 2.y0/2; which shows that Vdivides into two quadratic terms, each dependent on a parenthesized quantity (a new coordinate) proportional to one of our eigenvectors. The new coordinates are related to the original x;yby a rotation with unitary transformation U: UrDp5=6p1=6p1=6p5=6x y D .p 5xCy/=p 6 .xp 5y/=p 6! Dx0 y0 : Finally, we note that when we calculate the force in the primed coordinate system, we get Fx0Dx0;Fy0D7 y0; corresponding to the eigenvalues we found. Another Eigenproblem Example 6.2.1 is not complicated enough to provide a full illustration of the matrix eigen- value problem. Consider next the following example. Example 6.2.2 BLOCK-DIAGONAL MATRIX Find the eigenvalues and eigenvectors of HD0 @0 1 0 1 0 0 0 0 21 A: (6.7) Writing the secular equation and expanding in minors using the third row, we have  1 0 1 0 0 0 2 D.2/  1 1 D.2/.21/D0: (6.8) We see that the eigenvalues are 2, C1, and1. ArfKen_Ch06-9780123846549.tex 306 Chapter 6 Eigenvalue Problems To obtain the eigenvector corresponding to D2, we examine the equation set TH2.1/U cD0: 2c 1Cc2D0; c12c2D0; 0D0: The first two equations of this set lead to c1Dc2D0. The third obviously conveys no information, and we are led to the conclusion that c3is arbitrary. Thus, at this point we have 1D2; c1D0 @0 0 C1 A: (6.9) Taking nextDC1 , our matrix equation is TH1.1/U cD0, which is equivalent to the ordinary equations c1Cc2D0; c1c2D0; c3D0: We clearly have c1Dc2andc3D0, so 2DC1; c2D0 @C C 01 A: (6.10) Similar operations for D1 yield 3D1; c3D0 @C C 01 A: (6.11) Collecting our results, and normalizing the eigenvectors (often useful, but not in general necessary), we have 1D2; c1D0 @0 0 11 A;  2D1; c2D0 @21=2 21=2 01 A;  3D1; c3D0 @21=2 21=2 01 A: Note that because Hwas block-diagonal, with an upper-left 22block and a lower- right 11block, the secular equation separated into a product of the determinants for the two blocks, and its solutions corresponded to those of an individual block, with coeffi- cients of value zero for the other block(s). Thus, D2was a solution for the 11block in row/column 3, and its eigenvector involved only the coefficient c3. Thevalues1 came from the 22block in rows/columns 1 and 2, with eigenvectors involving only coefficients c1andc2.  In the case of a 11block in row/column i, we saw, for iD3inExample 6.2.2, that its only element was the eigenvalue, and that the corresponding eigenvector is proportional ArfKen_Ch06-9780123846549.tex 6.2 Matrix Eigenvalue Problems 307 toOei(a unit vector whose only nonzero element is ciD1). A generalization of this obser- vation is that if a matrix His diagonal, its diagonal elements hiiwill be the eigenvalues i, and that the eigenvectors ciwill be the unit vectors Oei. Degeneracy If the secular equation has a multiple root, the eigensystem is said to be degenerate or to exhibit degeneracy. Here is an example. Example 6.2.3 DEGENERATE EIGENPROBLEM Let’s find the eigenvalues and eigenvectors of HD0 @0 0 1 0 1 0 1 0 01 A: (6.12) The secular equation for this problem is  0 1 0 1 0 1 0 D2.1/.1/D.21/.1/D0 (6.13) with the three rootsC1,C1, and1. Let’s consider first D1 . Then we have c1Cc3D0; 2c2D0; c1Cc3D0: Thus, 1D1; c1DC0 @1 0 11 A: (6.14) For the double root DC1 , c1Cc3D0; 0D0; c1c3D0: Note that of the three equations, only one is now linearly independent; the double root sig- nalstwolinear dependencies, and we have solutions for anyvalues of c1andc2, with only the condition that c3Dc1. The eigenvectors for DC1 thus span a 2-D manifold (Dsub- space), in contrast to the trivial one-dimensional manifold characteristic of nondegenerate ArfKen_Ch06-9780123846549.tex 308 Chapter 6 Eigenvalue Problems solutions. The general form for these eigenvectors is DC1; cD0 @C C0 C1 A: (6.15) It is convenient to describe the degenerate eigenspace for D1by identifying two mutu- ally orthogonal vectors that span it. We can pick the first vector by choosing arbitrary val- ues of CandC0(an obvious choice is to set one of these, say C0, to zero). Then, using the Gram-Schmidt process (or in this case by simple inspection), we find a second eigenvector orthogonal to the first. Here, this leads to 2D3DC1; c2DC0 @1 0 11 A;c3DC00 @0 1 01 A: (6.16) Normalizing, our eigenvalues and eigenvectors become 1D1; c1D0 @21=2 0 21=21 A;  2D3D1; c2D0 @21=2 0 21=21 A;c3D0 @0 1 01 A:  The eigenvalue problems we have used as examples all led to secular equations with simple solutions; realistic applications frequently involve matrices of large dimension and secular equations of high degree. The solution of matrix eigenvalue problems has been an active field in numerical analysis and very sophisticated computer programs for this purpose are now available. Discussion of the details of such programs is outside the scope of this book, but the ability to use such programs should be part of the technology available to the working physicist. Exercises Find the eigenvalues and corresponding normalized eigenvectors of the matrices in Exercises 6.2.1 through 6.2.14. Orthogonalize any degenerate eigenvectors. 6.2.1 AD0 @1 0 1 0 1 0 1 0 11 A: ANS.D0;1;2: 6.2.2 AD0 @1p 2 0p 2 0 0 0 0 01 A: ANS.D1; 0;2. 6.2.3 AD0 @1 1 0 1 0 1 0 1 11 A: ANS.D1; 1;2. ArfKen_Ch06-9780123846549.tex 6.2 Matrix Eigenvalue Problems 309 6.2.4 AD0 @1p 8 0p 8 1p 8 0p 8 11 A: ANS.D3; 1;5. 6.2.5 AD0 @1 0 0 0 1 1 0 1 11 A: ANS.D0;1;2. 6.2.6 AD0 @1 0 0 0 1p 2 0p 2 01 A: ANS.D1; 1;2. 6.2.7 AD0 @0 1 0 1 0 1 0 1 01 A: ANS.Dp 2;0;p 2. 6.2.8 AD0 @2 0 0 0 1 1 0 1 11 A: ANS.D0;2;2. 6.2.9 AD0 @0 1 1 1 0 1 1 1 01 A: ANS.D1;1;2. 6.2.10 AD0 @111 1 11 11 11 A: ANS.D1; 2;2. 6.2.11 AD0 @1 1 1 1 1 1 1 1 11 A: ANS.D0;0;3. 6.2.12 AD0 @5 0 2 0 1 0 2 0 21 A: ANS.D1;1;6. 6.2.13 AD0 @1 1 0 1 1 0 0 0 01 A: ANS.D0;0;2. ArfKen_Ch06-9780123846549.tex 310 Chapter 6 Eigenvalue Problems 6.2.14 AD0 @5 0p 3 0 3 0p 3 0 31 A: ANS.D2;3;6. 6.2.15 Describe the geometric properties of the surface x2C2xyC2y2C2yzCz2D1: How is it oriented in 3-D space? Is it a conic section? If so, which kind? 6.3 H ERMITIAN EIGENVALUE PROBLEMS All the illustrative problems we have thus far examined have turned out to have real eigen- values; this was also true of all the exercises at the end of Section 6.2. We also found, whenever we bothered to check, that the eigenvectors corresponding to different eigenval- ues were orthogonal. The purpose of the present section is to show that these properties are consequences of the fact that all the eigenvalue problems we have considered were for Hermitian matrices. We remind the reader that the check for Hermiticity is simple: We simply verify that H is equal to its adjoint, H†; if a matrix is real, this condition is simply that it be symmetric. All the matrices to which we referred are clearly Hermitian. We now proceed to characterize the eigenvalues and eigenvectors of Hermitian matri- ces. Let Hbe a Hermitian matrix, with ciandcjtwo of its eigenvectors corresponding, respectively, to the eigenvalues iandj. Then, using Dirac notation, HjciiDijcii;Hjc jiDjjcji: (6.17) Multiplying on the left the first of these by c† j, which in Dirac notation is hcjj, and the second byhcij, hcjjHjc iiDihcjjcii;hcijHjc jiDjhcijcji: (6.18) We next take the complex conjugate of the second of these equations, noting that hcijcjiD hcjjcii, that we must complex conjugate the occurrence of j, and that hcijHjc jiDhHc jjciiDhc jjHjc ii: (6.19) Note that the first member of Eq. (6.19) contains the scalar product of ciwith Hcj. Com- plex conjugating this scalar product yields the second member of that equation. The final member of the equation follows because His Hermitian. The complex conjugation therefore converts Eqs. (6.18) into hcjjHjc iiDihcjjcii;hcjjHjc iiD jhcjjcii: (6.20) Equations (6.20) permit us to obtain two important results: First, if iDj, the scalar product hcjjciibecomeshcijcii, which is an inherently positive quantity. This means that the two equations are only consistent if iD i, meaning that imust be real. Thus, The eigenvalues of a Hermitian matrix are real. ArfKen_Ch06-9780123846549.tex 6.4 Hermitian Matrix Diagonalization 311 Next, if i6Dj, combining the two equations of Eq. (6.20), and remembering that the i are real, .ij/hcjjciiD0; (6.21) so that either iDjorhcjjciiD0. This tells us that Eigenvectors of a Hermitian matrix corresponding to different eigenvalues are orthogonal. Note, however, that if iDj, which will occur if iandjrefer to two degenerate eigen- vectors, we know nothing about their orthogonality. In fact, in Example 6.2.3 we examined a pair of degenerate eigenvectors, noting that they spanned a two-dimensional manifold and were not required to be orthogonal. However, we also noted in that context that we could choose them to be orthogonal. Sometimes (as in Example 6.2.3), it is obvious how to choose orthogonal degenerate eigenvectors. When it is not obvious, we can start from any linearly independent set of degenerate eigenvectors and orthonormalize them by the Gram-Schmidt process. Since the total number of eigenvectors of a Hermitian matrix is equal to its dimension, and since (whether or not there is degeneracy) we can make from them an orthonormal set of eigenvectors, we have the following important result: It is possible to choose the eigenvectors of a Hermitian matrix in such a way that they form an orthonormal set that spans the space of the matrix basis. This situation is often referred to by the statement, “The eigenvectors of a Hermitian matrix form a complete set.” This means that if the matrix is of order n, any vector of dimension ncan be written as a linear combination of the orthonormal eigenvectors, with coefficients determined by the rules for orthogonal expansions. We close this section by reminding the reader that theorems which have been established for an arbitrary basis-set expansion of a Hermitian eigenvalue equation apply also to that eigenvalue equation in its original form. Therefore, this section has also shown that: If H is a linear Hermitian operator on an arbitrary Hilbert space, 1. The eigenvalues of H are real. 2. Eigenfunctions corresponding to different eigenvalues of H are orthogonal. 3. It is possible to choose the eigenfunctions of H in a way such that they form a orthonormal basis for the Hilbert space. In general, the eigenfunctions of a Her- mitian operator form a complete set (i.e., a complete basis for the Hilbert space). 6.4 H ERMITIAN MATRIX DIAGONALIZATION InSection 6.2 we observed that if a matrix is diagonal, the diagonal elements are its eigen- values. This observation opens an alternative approach to the matrix eigenvalue problem. Given the matrix eigenvalue equation HcDc; (6.22) ArfKen_Ch06-9780123846549.tex 312 Chapter 6 Eigenvalue Problems where His a Hermitian matrix, consider what happens if we insert unity between Handc, as follows, with Ua unitary matrix, and then left-multiply the resulting equation by U: HU1UcDc! UHU1 Uc D Uc : (6.23) Equation (6.23) shows that our original eigenvalue equation has been converted into one in which Hhas been replaced by its unitary transformation (by U) and the eigenvector c has also been transformed by U, but the value of remains unchanged. We thus have the important result: The eigenvalues of a matrix remain unchanged when the matrix is subjected to a unitary transformation. Next, suppose that we choose Uin such a way that the transformed matrix UHU1is in the eigenvector basis. While we may or may not know how to construct this U, we know that such a unitary matrix exists because the eigenvectors form a complete orthog- onal set, and can be specified to be normalized. If we transform with the chosen U, the matrix UHU1will be diagonal, with the eigenvalues as diagonal elements. Moreover, the eigenvector UcofUHU1corresponding to the eigenvalue iD.UHU1/iiisOei(a column vector with all elements zero except for unity in the ith row). We may find the eigenvector ciof Eq. (6.22) by solving the equation UciDOei, obtaining ciDU1Oei. These observations correspond to the following: For any Hermitian matrix H, there exists a unitary transformation Uthat will cause UHU1to be diagonal, with the eigenvalues of Has its diagonal elements. This is an extremely important result. Another way of stating it is: A Hermitian matrix can be diagonalized by a unitary transformation, with its eigenval- ues as the diagonal elements. Looking next at the ith eigenvector U1Oei, we have 0 [email protected]1/11::: .U1/1i::: .U1/1n .U1/21::: .U1/2i::: .U1/2n           .U1/n1::: .U1/ni::: .U1/nn1 CCCCA0 BBBB@0  1  01 CCCCAD0 [email protected]1/1i .U1/2i   .U1/ni1 CCCCA: (6.24) We see that the columns of U1are the eigenvectors of H, normalized because U1is a unitary matrix. It is also clear from Eq. (6.24) thatU1is not entirely unique; if its columns are permuted, all that will happen is that the order of the eigenvectors are changed, with a corresponding permutation of the diagonal elements of the diagonal matrix UHU1. Summarizing, If a unitary matrix Uis such that, for a Hermitian matrix H,UHU1is diagonal, the normalized eigenvector of Hcorresponding to the eigenvalue .UHU1/iiwill be the ith column of U1. IfHis not degenerate, U1(and also U) will be unique except for a possible permutation of the columns of U1(and a corresponding permutation of the rows of U). However, if His degenerate (has a repeated eigenvalue), then the columns of U1corresponding to the same ArfKen_Ch06-9780123846549.tex 6.4 Hermitian Matrix Diagonalization 313 eigenvalue can be transformed among themselves, thereby giving additional flexibility to UandU1. Finally, calling on the fact that both the determinant and the trace of a matrix are unchanged when the matrix is subjected to a unitary transformation (shown in Section 5.7), we see that the determinant of a Hermitian matrix can be identified as the product of its eigenvalues, and its trace will be their sum. Apart from the individual eigenvalues them- selves, these are the most useful of the invariants that a matrix has with respect to unitary transformation. We illustrate the ideas thus far introduced in this section in the next example. Example 6.4.1 TRANSFORMING A MATRIX TO DIAGONAL FORM We return to the matrix HofExample 6.2.2: HD0 @0 1 0 1 0 0 0 0 21 A: We note that it is Hermitian, so there exists a unitary transformation Uthat will diagonalize it. Since we already know the eigenvectors of H, we can use them to construct U. Noting that we need normalized eigenvectors, and consulting Eqs. (6.9) to(6.11), we have D2;0 @0 0 11 AID1;0 @1=p 2 1=p 2 01 AID1;0 @1=p 2 1=p 2 01 A: Combining these as columns into U1, U1D0 @0 1=p 2 1=p 2 0 1=p 21=p 2 1 0 01 A: Since UD.U1/†, we easily form UD0 B@0 0 1 1=p 2 1=p 2 0 1=p 21=p 2 01 CA and UHU1D0 B@2 0 0 0 1 0 0 011 CA: The trace of His 2, as is the sum of the eigenvalues; det.H/ is2, equal to 21.1/ .  Finding a Diagonalizing Transformation As Example 6.4.1 shows, a knowledge of the eigenvectors of a Hermitian matrix Henables the direct construction of a unitary matrix Uthat transforms Hinto diagonal form. But we are interested in diagonalizing matrices for the purpose of finding their eigenvectors and eigenvalues, so the construction illustrated in Example 6.4.1 does not meet our present needs. Applied mathematicians (and even theoretical chemists!) have over many years ArfKen_Ch06-9780123846549.tex 314 Chapter 6 Eigenvalue Problems given attention to numerical methods for diagonalizing matrices of order large enough that direct, exact solution of the secular equation is not possible, and computer programs for carrying out these methods have reached a high degree of sophistication and efficiency. In varying ways, such programs involve processes that approach diagonalization via suc- cessive approximations. That is to be expected, since explicit formulas for the solution of high-degree algebraic equations (including, of course, secular equations) do not exist. To give the reader a sense of the level that has been reached in matrix diagonalization technol- ogy, we identify a computation1that determined some of the eigenvalues and eigenvectors of a matrix whose dimension exceeded 109. One of the older techniques for diagonalizing a matrix is due to Jacobi. It has now been supplanted by more efficient (but less transparent) methods, but we discuss it briefly here to illustrate the ideas involved. The essence of the Jacobi method is that if a Hermitian matrix Hhas a nonzero value of some off-diagonal hi j(and thus also hji), a unitary trans- formation that alters only rows/columns iand jcan reduce hi jandhjito zero. While this transformation may cause other, previously zeroed elements to become nonzero, it can be shown that the resulting matrix is closer to being diagonal (meaning that the sum of the squared magnitudes of its off-diagonal elements has been reduced). One may therefore apply Jacobi-type transformations repeatedly to reduce individual off-diagonal elements to zero, continuing until there is no off-diagonal element larger than an acceptable tolerance. If one constructs the unitary matrix that is the product of the individual transformations, one obtains thereby the overall diagonalizing transformation. Alternatively, one can use the Jacobi method only for retrieval of the eigenvalues, after which the method presented previously can be used to obtain the eigenvectors. Simultaneous Diagonalization It is of interest to know whether two Hermitian matrices AandBcan have a common set of eigenvectors. It turns out that this is possible if and only if they commute. The proof is simple if the eigenvectors of either AorBare nondegenerate. Assume that ciare a set of eigenvectors of both AandBwith respective eigenvalues ai andbi. Then form, for any i, BAc iDBaiciDbiaici; ABc iDAbiciDaibici: These equations show that BAc iDABc ifor every ci. Since any vector vcan be written as a linear combination of the ci, we find that .BAAB/vD0for all v, which means that BADAB. We have found that the existence of a common set of eigenvectors implies com- mutation. It remains to prove the converse, namely that commutation permits construction of a common set of eigenvectors. For the converse, we assume that AandBcommute, that ciis an eigenvector of Awith eigenvalue ai, and that this eigenvector of Ais nondegenerate. Then we form ABc iDBAc iDBaici;orA.Bc i/Dai.Bci/: 1J. Olsen, P. Jørgensen, and J. Simons, Passing the one-billion limit in full configuration-interaction calculations, Chem. Phys. Lett. 169: 463 (1990). ArfKen_Ch06-9780123846549.tex 6.4 Hermitian Matrix Diagonalization 315 This equation shows that Bciis also an eigenvector of Awith eigenvalue ai. Since the eigenvector of Awas assumed nondegenerate, Bcimust be proportional to ci, meaning thatciis also an eigenvector of B. This completes the proof that if AandBcommute, they have common eigenvectors. The proof of this theorem can be extended to include the case in which both opera- tors have degenerate eigenvectors. Including that extension, we summarize by stating the general result: Hermitian matrices have a complete set of eigenvectors in common if and only if they commute. It may happen that we have three matrices A,B, and C, and thatTA;BUD 0andTA;CUD 0, butTB;CU6D 0. In that case, which is actually quite common in atomic physics, we have a choice. We can insist upon a set of cithat are simultaneous eigenvectors of Aand B, in which case not all the cican be eigenvectors of C, or we can have simultaneous eigenvectors of AandC, but not B. In atomic physics these choices typically correspond to descriptions in which different angular momenta are required to have definite values. Spectral Decomposition Once the eigenvalues and eigenvectors of a Hermitian matrix Hhave been found, we can express Hin terms of these quantities. Since mathematicians call the set of eigen- values of Hitsspectrum, the expression we now derive for His referred to as its spectral decomposition. As previously noted, in the orthonormal eigenvector basis the matrix Hwill be diagonal. Then, instead of the general form for the basis expansion of an operator, Hwill be of the diagonal form HDX jcihcj;each csatisfies HcDcandhcjciD1: (6.25) This result, the spectral decomposition ofH, is easily checked by applying it to any eigen- vector c. Another result related to the spectral decomposition of Hcan be obtained if we multiply both sides of the equation HcDcon the left by H, reaching H2cD./2cI further applications of Hshow that all positive powers of Hhave the same eigenvectors asH, so if f.H/is any function of Hthat has a power-series expansion, it has spectral decomposition f.H/DX jcif./hcj: (6.26) Equation (6.26) can be extended to include negative powers if His nonsingular; to do so, multiply HcDcon the left by H1and rearrange, to obtain H1cD1 c; showing that negative powers of Halso have the same eigenvectors as H. ArfKen_Ch06-9780123846549.tex 316 Chapter 6 Eigenvalue Problems Finally, we can now easily prove the trace formula, Eq. (2.84). In the eigenvector basis, det exp.A/ DY eDexp X ! Dexp trace.A/ : (6.27) Since the determinant and trace are basis-independent, this proves the trace formula. Expectation Values The expectation value of a Hermitian operator Hassociated with the normalized function was defined in Eq. (5.61) as hHiDh jHj i; (6.28) where it was shown that if an orthonormal basis was introduced, with Hthen represented by a matrix Hand represented by a vector a, this expectation value assumed the form hHiDa†HaDhajHjaiDX a ha: (6.29) If these quantities are expressed in the orthonormal eigenvector basis, Eq. (6.29) becomes hHiDX a aDX jaj2; (6.30) where ais the coefficient of the eigenvector c(with eigenvalue ) in the expansion of . We note that the expectation value is a weighted sum of the eigenvalues, with the weights nonnegative, and adding to unity because hajaiDX a aDX jaj2D1: (6.31) An obvious implication of Eq. (6.30) is that the expectation value hHicannot be smaller than the smallest eigenvalue nor larger than the largest eigenvalue. The quantum- mechanical interpretation of this observation is that if Hcorresponds to a physical quantity, measurements of that quantity will yield the values with relative probabilities given by jaj2, and hence with an average value corresponding to the weighted sum, which is the expectation value. Hermitian operators arising in physical problems often have finite smallest eigenvalues. This, in turn, means that the expectation value of the physical quantity associated with the operator has a finite lower bound. We thus have the frequently useful relation If the algebraically smallest eigenvalue of His finite, then, for any ,h jHj iwill be greater than or equal to this eigenvalue, with the equality occurring only if is an eigenfunction corresponding to the smallest eigenvalue. ArfKen_Ch06-9780123846549.tex 6.4 Hermitian Matrix Diagonalization 317 Positive Definite and Singular Operators If all the eigenvalues of an operator Aare positive, it is termed positive definite. If and only if Ais positive definite, its expectation value for any nonzero , namelyh jAj i, will also be positive, since (when is normalized) it must be equal to or larger than the smallest eigenvalue. Example 6.4.2 OVERLAP MATRIX LetSbe an overlap matrix of elements sDhji, where theare members of a linearly independent, but nonorthogonal basis. If we assume an arbitrary nonzero function to be expanded in terms of the , according to DP b, the scalar producth j i will be given by h j iDX b sb; which is of the form of an expectation value for the matrix S. Sinceh j iis an inherently positive quantity, we conclude that Sis positive definite.  If, on the other hand, the rows (or the columns) of a square matrix represent linearly dependent forms, either as coefficients in a basis-set expansion or as the coefficients of a linear expression in a set of variables, the matrix will be singular, and that fact will be signaled by the presence of eigenvalues that are zero. The number of zero eigenvalues provides an indication of the extent of the linear dependence; if an nnmatrix has mzero eigenvalues, its rank will be nm. Exercises 6.4.1 Show that the eigenvalues of a matrix are unaltered if the matrix is transformed by a similarity transformation—a transformation that need not be unitary, but of the form given in Eq. (5.78). This property is not limited to symmetric or Hermitian matrices. It holds for any matrix satisfying an eigenvalue equation of the type AxDx. If our matrix can be brought into diagonal form by a similarity transformation, then two immediate conse- quences are that: 1. The trace (sum of eigenvalues) is invariant under a similarity transformation. 2. The determinant (product of eigenvalues) is invariant under a similarity transformation. Note. The invariance of the trace and determinant are often demonstrated by using the Cayley-Hamilton theorem, which states that a matrix satisfies its own characteristic (secular) equation. 6.4.2 As a converse of the theorem that Hermitian matrices have real eigenvalues and that eigenvectors corresponding to distinct eigenvalues are orthogonal, show that if ArfKen_Ch06-9780123846549.tex 318 Chapter 6 Eigenvalue Problems (a) the eigenvalues of a matrix are real and (b) the eigenvectors satisfy x† ixjDi j, then the matrix is Hermitian. 6.4.3 Show that a real matrix that is not symmetric cannot be diagonalized by an orthogonal or unitary transformation. Hint. Assume that the nonsymmetric real matrix can be diagonalized and develop a contradiction. 6.4.4 The matrices representing the angular momentum components Lx;Ly, and Lzare all Hermitian. Show that the eigenvalues of L2;where L2DL2 xCL2 yCL2 z;are real and nonnegative. 6.4.5 Ahas eigenvalues iand corresponding eigenvectors jxii. Show that A1has the same eigenvectors but with eigenvalues 1 i. 6.4.6 A square matrix with zero determinant is labeled singular. (a) If Ais singular, show that there is at least one nonzero column vector vsuch that AjviD 0: (b) If there is a nonzero vector jvisuch that AjviD 0; show that Ais a singular matrix. This means that if a matrix (or operator) has zero as an eigenvalue, the matrix (or operator) has no inverse and its determinant is zero. 6.4.7 Two Hermitian matrices AandBhave the same eigenvalues. Show that AandBare related by a unitary transformation. 6.4.8 Find the eigenvalues and an orthonormal set of eigenvectors for each of the matrices of Exercise 2.2.12. 6.4.9 The unit process in the iterative matrix diagonalization procedure known as the Jacobi method is a unitary transformation that operates on rows/columns iand jof a real symmetric matrix Ato make ai jDajiD0. If this transformation (from basis functions 'iand'jto'0 iand'0 j) is written '0 iD'icos'jsin; '0 jD'isinC'jcos; (a) Show that ai jis transformed to zero if tan 2D2ai j aj jaii, (b) Show that aremains unchanged if neither norisiorj, (c) Find a0 iianda0 j jand show that the trace of Ais not changed by the transformation, (d) Find a0 ianda0 j(whereis neither inorj) and show that the sum of the squares of the off-diagonal elements of Ais reduced by the amount 2a2 i j. ArfKen_Ch06-9780123846549.tex 6.5 Normal Matrices 319 6.5 N ORMAL MATRICES Thus far the discussion has been centered on Hermitian eigenvalue problems, which we showed to have real eigenvalues and orthogonal eigenvectors, and therefore capable of being diagonalized by a unitary transformation. However, the class of matrices which can be diagonalized by a unitary transformation contains, in addition to Hermitian matrices, all other matrices that commute with their adjoints; a matrix Awith this property, namely TA;A†UD0; is termed normal.2Clearly Hermitian matrices are normal, as H†DH. Unitary matrices are also normal, as Ucommutes with its inverse. Anti-Hermitian matrices (with A†DA ) are also normal. And there exist normal matrices that are not in any of these categories. To show that normal matrices can be diagonalized by a unitary transformation, it suf- fices to prove that their eigenvectors can form an orthonormal set, which reduces to the requirement that eigenvectors of different eigenvalues be orthogonal. The proof proceeds in two steps, of which the first is to demonstrate that a normal matrix Aand its adjoint have the same eigenvectors. Assumingjxito be an eigenvector of Awith eigenvalue , we have the equation .A1/jxiD 0: Multiplying this equation on its left by hxj.A†1/, we have hxj.A†1/.A1/jxiD 0; after which we use the normal property to interchange the two parenthesized quantities, bringing us to hxj.A1/.A†1/jxiD 0: Moving the first parenthesized quantity into the left half-bracket, we have h.A†1/xj.A†1/jxiD 0; which we identify as a scalar product of the form hfjfi. The only way this scalar product can vanish is if .A†1/jxiD 0; showing thatjxiis an eigenvector of A†in addition to being an eigenvector of A. However, the eigenvalues of AandA†are complex conjugates; for general normal matrices need not be real. A demonstration that the eigenvectors are orthogonal proceeds along the same lines are for Hermitian matrices. Letting jxiiandjxjibe two eigenvectors (of both AandA†), we form hxjjAjx iiDihxjjxii;hxijA†jxjiD jhxijxji: (6.32) 2Normal matrices are the largest class of matrices that can be diagonalized by unitary transformations. For an extensive discus- sion of normal matrices, see P. A. Macklin, Normal matrices for physicists, Am. J. Phys. 52: 513 (1984). ArfKen_Ch06-9780123846549.tex 320 Chapter 6 Eigenvalue Problems We now take the complex conjugate of the second of these equations, noting that hxijxjiDhx jjxii. To form the complex conjugate of hxijA†jxji, we convert it first to hAxijxjiand then interchange the two half-brackets. Equations (6.32) then become hxjjAjx iiDihxjjxii;hxjjAjx iiDjhxjjxii: (6.33) These equations indicate that if i6Dj, we must havehxjjxiiD0, thus proving orthogonality. The fact that the eigenvalues of a normal matrix A†are complex conjugates of the eigen- values of Aenables us to conclude that the eigenvalues of an anti-Hermitian matrix are pure imaginary (because A†DA , D ), and the eigenvalues of a unitary matrix are of unit magnitude (because D1=, equivalent toD1). Example 6.5.1 A NORMAL EIGENSYSTEM Consider the unitary matrix UD0 @0 0 1 1 0 0 0 1 01 A: This matrix describes a rotational transformation in which z!x,x!y, and y!z. Because it is unitary, it is also normal, and we may find its eigenvalues from the secular equation det.U1/D  0 1 1 0 0 1 D3C1D0; which has solutions D1,!, and!, where!De2i=3. (Note that!3D1, so!D 1=!D!2.) Because Uis real, unitary and describes a rotation, its eigenvalues must fall on the unit circle, their sum (the trace) must be real, and their product (the determinant) must beC1. This means that one of the eigenvalues must be C1, and the remaining two may be real (bothC1or both1) or form a complex conjugate pair. We see that the eigenvalues we have found satisfy these criteria. The trace of Uis zero, as is the sum 1C!C!(this may be verified graphically; see Fig. 6.3). Proceeding to the eigenvectors, substitution into the equation .U1/cD0 ArfKen_Ch06-9780123846549.tex 6.5 Normal Matrices 321 ω ω*1 FIGURE 6.3 Eigenvalues of the matrix U,Example 6.5.1. yields (in unnormalized form) 1D1; c1D0 @1 1 11 A;  2D!; c2D0 @1 ! !1 A;  3D!2;c3D0 @1 ! !1 A: The interpretation of this result is interesting. The eigenvector c1is unchanged by U (application of Umultiplies it by unity), so it must lie in the direction of the axis of the rotation described by U. The other two eigenvectors are complex linear combinations of the coordinates that are invariant in “direction,” but not in phase under application of U. We write “direction” in quotes, since the complex coefficients in the eigenvectors cause them not to identify directions in physical space. Nevertheless, they do form quantities that are invariant except for multiplication by the eigenvalue (which we identify as a phase, since it is of magnitude unity). The argument of !,2=3 , identifies the amount of the rotation about the c1axis. Coming back to physical reality, we note that we have found that U corresponds to a rotation of amount 2=3 about an axis in the (1,1,1) direction; the reader can verify that this indeed takes xintoy,yintoz, and zintox. Because Uis normal, its eigenvectors must be orthogonal. Since we now have complex quantities, in order to check this we must compute the scalar product of two vectors aand bfrom the formula a†b. Our eigenvectors pass this test. Finally, let’s verify that UandU†have the same eigenvectors, and that corresponding eigenvalues are complex conjugates. Taking the adjoint of U, we have U†D0 @0 1 0 0 0 1 1 0 01 A: Using the eigenvectors we have already found to form U†ci, the verification is easily estab- lished. We illustrate with c2: 0 @0 1 0 0 0 1 1 0 01 A0 @1 ! !1 AD0 @! ! 11 AD!0 @1 ! !1 A; as required.  ArfKen_Ch06-9780123846549.tex 322 Chapter 6 Eigenvalue Problems Nonnormal Matrices Matrices that are not even normal sometimes enter problems of importance in physics. Such a matrix, A, still has the property that the eigenvalues of A†are the complex conju- gates of the eigenvalues of A, because det.A†/DTdet.A/U, so det.A1/D0! det.A†1/D0; for the same, but it is no longer true that the eigenvectors are orthogonal or that AandA† have common eigenvectors. Here is an example arising from the analysis of vibrations in mechanical systems. We consider the vibrations of a classical model of the CO 2molecule. Even though the model is classical, it is a good representation of the actual quantum-mechanical system, as to good approximation the nuclei execute small (classical) oscillations in the Hooke’s-law potential generated by the electron distribution. This problem is an illustration of the application of matrix techniques to a problem that does not start as a matrix problem. It also provides an example of the eigenvalues and eigenvectors of an asymmetric real matrix. Example 6.5.2 NORMAL MODES Consider three masses on the x-axis joined by springs as shown in Fig. 6.4. The spring forces are assumed to be linear in the displacements from equilibrium (small displace- ments, Hooke’s law), and the masses are constrained to stay on the x-axis. Using a different coordinate for the displacement of each mass from its equilibrium position, Newton’s second law yields the set of equations Rx1Dk M.x1x2/ Rx2Dk m.x2x1/k m.x2x3/ (6.34) Rx3Dk M.x3x2/; whereRxstands for d2x=dt2. We seek the frequencies, !, such that all the masses vibrate at the same frequency. These are called the normal modes of vibration,3and are solutions to Eqs. (6.34) with xi.t/Dxisin!t;iD1;2;3: C k x1kO O x2x3 FIGURE 6.4 The three-mass spring system representing the CO 2molecule. 3For detailed discussion of normal modes of vibration, see E. B. Wilson, Jr., J. C. Decius, and P. C. Cross, Molecular Vibrations—The Theory of Infrared and Raman Vibrational Spectra. New York: Dover (1980). ArfKen_Ch06-9780123846549.tex 6.5 Normal Matrices 323 Substituting this solution set into Eqs. (6.34), these equations, after cancellation of the common factor sin!t, become equivalent to the matrix equation Ax0 BBBBB@k Mk M0 k m2k mk m 0k Mk M1 CCCCCA0 @x1 x2 x31 ADC!20 @x1 x2 x31 A: (6.35) We can find the eigenvalues of Aby solving the secular equation k M!2k M0 k m2k m!2k m 0k Mk M!2 D0; (6.36) which expands to !2k M!2 !22k mk M D0: The eigenvalues are !2D0;k M;k MC2k m: For!2D0, substitution back into Eq. (6.35) yields x1x2D0;x1C2x2x3D0;x2Cx3D0; which corresponds to x1Dx2Dx3. This describes pure translation with no relative motion of the masses and no vibration. For!2Dk=M,Eq. (6.35) yields x1Dx3;x2D0: The two outer masses are moving in opposite directions. The central mass is stationary. In CO2this is called the symmetric stretching mode. Finally, for!2Dk=MC2k=m, the eigenvector components are x1Dx3;x2D2M mx1: In this antisymmetric stretching mode, the two outer masses are moving, together, in a direction opposite to that of the central mass, so one CO bond stretches while the other contracts the same amount. In both of these stretching modes, the net momentum of the motion is zero. Any displacement of the three masses along the x-axis can be described as a linear combination of these three types of motion: translation plus two forms of vibration. ArfKen_Ch06-9780123846549.tex 324 Chapter 6 Eigenvalue Problems The matrix Aof Eq. (6.35) is not normal; the reader can check that AA†6DA†A. As a result, the eigenvectors we have found are not orthogonal, as is obvious by examination of the unnormalized eigenvectors: !2D0; xD0 @1 1 11 A; !2Dk M;xD0 @1 0 11 A; !2Dk MC2k m;xD0 @1 2M=m 11 A: Using the same values, we can solve the simultaneous equations  A†1 yD0: The resulting eigenvectors are !2D0; xD0 @1 m=M 11 A; !2Dk M;xD0 @1 0 11 A; !2Dk MC2k m;xD0 @1 2 11 A: These vectors are neither orthogonal nor the same as the eigenvectors of A.  Defective Matrices If a matrix is not normal, it may not even have a full complement of eigenvectors. Such matrices are termed defective. By the fundamental theorem of algebra, a matrix of dimen- sion Nwill have Neigenvalues (when their multiplicity is taken into account). It can also be shown that any matrix will have at least one eigenvector corresponding to each of its distinct eigenvalues. But it is notalways true that that an eigenvalue of multiplicity k>1 will have keigenvectors. We give as a simple example a matrix with the doubly degenerate eigenvalueD1: 1 1 0 1 has only the single eigenvector1 0 : Exercises 6.5.1 Find the eigenvalues and corresponding eigenvectors for 2 4 1 2 : Note that the eigenvectors are notorthogonal. ANS.1D0;c1D.2;1/; 2D4;c2D.2;1/. 6.5.2 IfAis a22matrix, show that its eigenvalues satisfy the secular equation 2trace.A/Cdet.A/D0: 6.5.3 Assuming a unitary matrix Uto satisfy an eigenvalue equation UcDc, show that the eigenvalues of the unitary matrix have unit magnitude. This same result holds for real orthogonal matrices. ArfKen_Ch06-9780123846549.tex 6.5 Normal Matrices 325 6.5.4 Since an orthogonal matrix describing a rotation in real 3-D space is a special case of a unitary matrix, such an orthogonal matrix can be diagonalized by a unitary transformation. (a) Show that the sum of the three eigenvalues is 1C2 cos', where'is the net angle of rotation about a single fixed axis. (b) Given that one eigenvalue is 1, show that the other two eigenvalues must be ei' andei'. Our orthogonal rotation matrix (real elements) has complex eigenvalues. 6.5.5 Ais an nth-order Hermitian matrix with orthonormal eigenvectors jxiiand real eigen- values123n. Show that for a unit magnitude vector jyi, 1hyjAjyin: 6.5.6 A particular matrix is both Hermitian and unitary. Show that its eigenvalues are all 1. Note. The Pauli and Dirac matrices are specific examples. 6.5.7 For his relativistic electron theory Dirac required a set of four anticommuting matrices. Assume that these matrices are to be Hermitian and unitary. If these are nnmatrices, show that nmust be even. With 22matrices inadequate (why?), this demonstrates that the smallest possible matrices forming a set of four anticommuting, Hermitian, unitary matrices are 44. 6.5.8 Ais a normal matrix with eigenvalues nand orthonormal eigenvectors jxni. Show that Amay be written as ADX nnjxnihxnj: Hint. Show that both this eigenvector form of Aand the original Agive the same result acting on an arbitrary vector jyi. 6.5.9 Ahas eigenvalues 1 and 1 and corresponding eigenvectors1 0 and0 1 . Construct A. ANS. AD1 0 01 : 6.5.10 A non-Hermitian matrix Ahas eigenvalues iand corresponding eigenvectors juii. The adjoint matrix A†has the same set of eigenvalues but different corresponding eigen- vectors,jvii. Show that the eigenvectors form a biorthogonal set in the sense that hvijujiD0for i6Dj: 6.5.11 You are given a pair of equations: AjfniDnjgni QAjgniDnjfniwith Areal. ArfKen_Ch06-9780123846549.tex 326 Chapter 6 Eigenvalue Problems (a) Prove thatjfniis an eigenvector of .QAA)with eigenvalue 2 n. (b) Prove thatjgniis an eigenvector of .AQA/with eigenvalue 2 n. (c) State how you know that (1) Thejfniform an orthogonal set. (2) Thejgniform an orthogonal set. (3)2 nis real. 6.5.12 Prove that Aof the preceding exercise may be written as ADX nnjgnihfnj; with thejgniandhfnjnormalized to unity. Hint. Expand an arbitrary vector as a linear combination of jfni. 6.5.13 Given AD1p 52 2 14 ; (a) Construct the transpose QAand the symmetric forms QAAandAQA. (b) From AQAjgniD2 njgni, findnandjgni. Normalize thejgni. (c) FromQAAjf niD2 njgni, findn[same as (b)] andjfni. Normalize thejfni. (d) Verify that AjfniDnjgniandQAjgniDnjfni. (e) Verify that ADP nnjgnihfnj. 6.5.14 Given the eigenvalues 1D1; 2D1 and the corresponding eigenvectors jf1iD1 0 ;jg1iD1p 21 1 ;jf2iD0 1 ;and jg2iD1p 21 1 ; (a) construct A; (b) verify that AjfniDnjgni; (c) verify thatQAjgniDnjfni. ANS. AD1p 211 1 1 : 6.5.15 Two matrices UandHare related by UDeiaH; with areal. ArfKen_Ch06-9780123846549.tex 6.5 Normal Matrices 327 (a) If His Hermitian, show that Uis unitary. (b) If Uis unitary, show that His Hermitian. ( His independent of a.) (c) If trace HD0, show that detUDC1 . (d) If detUDC1 , show that trace HD0. Hint. Hmay be diagonalized by a similarity transformation. Then Uis also diagonal. The corresponding eigenvalues are given by ujDexp.iah j). 6.5.16 Annnmatrix Ahasneigenvalues Ai. IfBDeA;show that Bhas the same eigen- vectors as Awith the corresponding eigenvalues Bigiven by BiDexp.Ai/. 6.5.17 A matrix Pis a projection operator satisfying the condition P2DP: Show that the corresponding eigenvalues .2/andsatisfy the relation .2/D./2D: This means that the eigenvalues of Pare 0 and 1. 6.5.18 In the matrix eigenvector-eigenvalue equation AjxiiDijxii; Ais an nnHermitian matrix. For simplicity assume that its nreal eigenvalues are distinct,1being the largest. Ifjxiis an approximation to jx1i, jxiDjx 1iCnX iD2ijxii; show that hxjAjxi hxjxi1 and that the error in 1is of the orderjij2. Takejij<<1. Hint. Thenvectorsjxiiform a complete orthogonal set spanning the n-dimensional (complex) space. 6.5.19 Two equal masses are connected to each other and to walls by springs as shown in Fig. 6.5. The masses are constrained to stay on a horizontal line. (a) Set up the Newtonian acceleration equation for each mass. (b) Solve the secular equation for the eigenvectors. (c) Determine the eigenvectors and thus the normal modes of motion. 6.5.20 Given a normal matrix Awith eigenvalues j;show that A†has eigenvalues  j;its real part.ACA†/=2has eigenvalues<e.j/;and its imaginary part .AA†/=2ihas eigenvalues=m.j/: ArfKen_Ch06-9780123846549.tex 328 Chapter 6 Eigenvalue Problems kkk mm FIGURE 6.5 Triple oscillator. 6.5.21 Consider a rotation given by Euler angles D=4, D=2, D5=4 . (a) Using the formula of Eq. (3.37), construct the matrix Urepresenting this rotation. (b) Find the eigenvalues and eigenvectors of U, and from them describe this rotation by specifying a single rotation axis and an angle of rotation about that axis. Note. This technique provides a representation of rotations alternative to the Euler angles. Additional Readings Bickley, W. G., and R. S. H. G. Thompson, Matrices—Their Meaning and Manipulation. Princeton, NJ: Van Nostrand (1964). A comprehensive account of matrices in physical problems, and their analytic properties and numerical techniques. Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969), reprinting, Dover (1992). Gilbert, J. and L. Gilbert, Linear Algebra and Matrix Theory. San Diego: Academic Press (1995). Golub, G. H., and C. F. Van Loan, Matrix Computations, 3rd ed. Baltimore: JHU Press (1996). Detailed mathe- matical background and algorithms for the production of numerical software, including methods for parallel computation. A classic computer science text. Halmos, P. R., Finite-Dimensional Vector Spaces, 2nd ed. Princeton, NJ: Van Nostrand (1958), reprinting, Springer (1993). Hirsch, M., Differential Equations, Dynamical Systems, and Linear Algebra. San Diego: Academic Press (1974). Heading, J., Matrix Theory for Physicists. London: Longmans, Green and Co. (1958). A readable introduction to determinants and matrices, with applications to mechanics, electromagnetism, special relativity, and quantum mechanics. Jain, M. C., Vector Spaces and Matrices in Physics, 2nd ed. Oxford: Alpha Science International (2007). Watkins, D. S., Fundamentals of Matrix Computations. New York: Wiley (1991). Wilkinson, J. H., The Algebraic Eigenvalue Problem. London: Oxford University Press (1965), reprinting (2004). Classic treatise on numerical computation of eigenvalue problems. Perhaps the most widely read book in the field of numerical analysis. ArfKen_Ch07-9780123846549.tex CHAPTER 7 ORDINARY DIFFERENTIAL EQUATIONS Much of theoretical physics is originally formulated in terms of differential equations in the three-dimensional physical space (and sometimes also time). These variables (e.g., x,y,z, t) are usually referred to as independent variables, while the function or functions being differentiated are referred to as dependent variable(s). A differential equation involving more than one independent variable is called a partial differential equation, often abbre- viated PDE. The simpler situation considered in the present chapter is that of an equation in a single independent variable, known as an ordinary differential equation, abbreviated ODE. As we shall see in a later chapter, some of the most frequently used methods for solv- ing PDEs involve their expression in terms of the solutions to ODEs, so it is appropriate to begin our study of differential equations with ODEs. 7.1 I NTRODUCTION To start, we note that the taking of a derivative is a linear operation, meaning that d dx a'.x/Cb .x/ Dad' dxCbd dx; and the derivative operation can be viewed as defining a linear operator: LDd=dx. Higher derivatives are also linear operators, as for example d2 dx2 a'.x/Cb .x/ Dad2' dx2Cbd2 dx2: 329 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch07-9780123846549.tex 330 Chapter 7 Ordinary Differential Equations Note that the linearity under discussion is that of the operator. For example, if we define LDp.x/d dxCq.x/; it is identified as linear because L a'.x/Cb .x/ Da p.x/d' dxCq.x/' Cb p.x/d dxCq.x/  DaL'CbL : We see that the linearity of Limposes no requirement that either p.x/orq.x/be a linear function of x. Linear differential operators therefore include those of the form LnX D0p.x/d dx ; where the functions p.x/are arbitrary. An ODE is termed homogeneous if the dependent variable (here ') occurs to the same power in all its terms, and inhomogeneous otherwise; it is termed linear if it can be written in the form L'.x/DF.x/; (7.1) where Lis a linear differential operator and F.x/is an algebraic function of x(i.e., not a differential operator). An important class of ODEs are those that are both linear and homogeneous, and thereby of the form L'D0. The solutions to ODEs are in general not unique, and if there are multiple solutions it is useful to identify those that are linearly independent (linear dependence is discussed in Section 2.1). Homogeneous linear ODEs have the general property that any multiple of a solution is also a solution, and that if there are multiple linearly independent solutions, any linear combination of those solutions will also solve the ODE. This statement is equivalent to noting that if Lis linear, then, for all aandb, L'D0andL D0! L.a'Cb /D0: The Schrödinger equation of quantum mechanics is a homogeneous linear ODE (or if in more than one dimension, a homogeneous linear PDE), and the property that any linear combination of its solutions is also a solution is the conceptual basis for the well-known superposition principle in electrodynamics, wave optics and quantum theory. Notationally, it is often convenient to use the symbols xandyto refer, respectively, to independent and dependent variables, and a typical linear ODE then takes the form LyDF.x/. It is also customary to use primes to indicate derivatives: y0dy=dx. In terms of this notation, the superposition property of solutions y1andy2of a homogeneous linear ODE tells us that the ODE also has as solutions c1y1,c2y2, and c1y1Cc2y2, with theciarbitrary constants. Some physically important problems (particularly in fluid mechanics and in chaos the- ory) give rise to nonlinear differential equations. A well-studied example is the Bernoulli equation y0Dp.x/yCq.x/yn;n6D0;1; which cannot be written in terms of a linear operator applied to y. ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 331 Further terms used to classify ODEs include their order (highest derivative appear- ing therein), and degree (power to which the highest derivative appears after the ODE is rationalized if that is necessary). For many applications, the concept of linearity is more relevant than that of degree. 7.2 F IRST-ORDER EQUATIONS Physics involves some first-order differential equations. For completeness it seems desir- able to touch upon them briefly. We consider the general form dy dxDf.x;y/DP.x;y/ Q.x;y/: (7.2) While there is no systematic way to solve the most general first-order ODE, there are a number of techniques that are often useful. After reviewing some of these techniques, we proceed to a more detailed treatment of linear first-order ODEs, for which systematic procedures are available. Separable Equations Frequently Eq. (7.2) will have the special form dy dxDP.x/ Q.y/: (7.3) Then it may be rewritten as P.x/dxCQ.y/dyD0: Integrating from .x0;y0/to.x;y/yields xZ x0P.x/dxCyZ y0Q.y/dyD0: Since the lower limits, x0andy0, contribute constants, we may ignore them and simply add a constant of integration. Note that this separation of variables technique does notrequire that the differential equation be linear. Example 7.2.1 PARACHUTIST We want to find the velocity of a falling parachutist as a function of time and are partic- ularly interested in the constant limiting velocity, v0, that comes about by air drag, taken to be quadratic,bv2, and opposing the force of the gravitational attraction, mg, of the Earth on the parachutist. We choose a coordinate system in which the positive direction is downward so that the gravitational force is positive. For simplicity we assume that the parachute opens immediately, that is, at time tD0, wherev.t/D0, our initial condition. Newton’s law applied to the falling parachutist gives mPvDmgbv2; (7.4) where mincludes the mass of the parachute. ArfKen_Ch07-9780123846549.tex 332 Chapter 7 Ordinary Differential Equations The terminal velocity, v0, can be found from the equation of motion as t!1 ; when there is no acceleration, PvD0, and bv2 0Dmg;orv0Drmg b: It simplifies further work to rewrite Eq. (7.4) as m bPvDv2 0v2: This equation is separable, and we write it in the form dv v2 0v2Db mdt: (7.5) Using partial fractions to write 1 v2 0v2D1 2v01 vCv01 vv0 ; it is straightforward to integrate both sides of Eq. (7.5) (the left-hand side from vD0tov, the right-hand side from tD0tot), yielding 1 2v0lnv0Cv v0vDb mt: Solving for the velocity, we have vDe2t=T1 e2t=TC1v0Dv0sinh.t=T/ cosh.t=T/Dv0tanht T; where TDpm=gbis the time constant governing the asymptotic approach of the velocity to its limiting value, v0. Inserting numerical values, gD9:8m/s2, and taking bD700kg/m, mD70kg, gives v0Dp9:8=101m/s3:6km/h2:234 mi/h, the walking speed of a pedestrian at landing, and TDpm=bgD1=p 109:80:1s. Thus, the constant speed v0is reached within a second. Finally, because it is always important to check the solution, we verify that our solution satisfies the original differential equation: PvDcosh.t=T/ cosh.t=T/v0 Tsinh2.t=T/ cosh2.t=T/v0 TDv0 Tv2 Tv0Dgb mv2: A more realistic case, where the parachutist is in free fall with an initial speed v.0/> 0 before the parachute opens, is addressed in Exercise 7.2.16.  ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 333 Exact Differentials Again we rewrite Eq. (7.2) as P.x;y/dxCQ.x;y/dyD0: (7.6) This equation is said to be exact if we can match the left-hand side of it to a differential d', and thereby reach d'D@' @xdxC@' @ydyD0: (7.7) Exactness therefore implies that there exists a function '.x;y/such that @' @xDP.x;y/and@' @yDQ.x;y/; (7.8) because then our ODE corresponds to an instance of Eq. (7.7), and its solution will be '.x;y/Dconstant. Before seeking to find a function 'satisfying Eq. (7.8), it is useful to determine whether such a function exists. Taking the two formulas from Eq. (7.8), differentiating the first with respect to yand the second with respect to x, we find @2' @y@[email protected];y/ @yand@2' @x@[email protected];y/ @x; and these are consistent if and only if @P.x;y/ @[email protected];y/ @x: (7.9) We therefore conclude that Eq. (7.6) is exact only if Eq. (7.9) is satisfied. Once exactness has been verified, we can integrate Eqs. (7.8) to obtain'and therewith a solution to the ODE. The solution takes the form '.x;y/DxZ x0P.x;y/dxCyZ y0Q.x0;y/dyDconstant. (7.10) Proof of Eq. (7.10) is left to Exercise 7.2.7. We note that separability and exactness are independent attributes. All separable ODEs are automatically exact, but not all exact ODEs are separable. Example 7.2.2 A NONSEPARABLE EXACT ODE Consider the ODE y0C 1Cy x D0: Multiplying by x dx, this ODE becomes .xCy/dxCx dyD0; ArfKen_Ch07-9780123846549.tex 334 Chapter 7 Ordinary Differential Equations which is of the form P.x;y/dxCQ.x;y/dyD0; with P.x;y/DxCyandQ.x;y/Dx. The equation is not separable. To check if it is exact, we compute @P @[email protected]/ @yD1;@Q @xD@x @xD1: These partial derivatives are equal; the equation is exact, and can be written in the form d'DP dxCQ dyD0: The solution to the ODE will be 'DC, with'computed according to Eq. (7.10): 'DxZ x0.xCy/dxCyZ y0x0dyD x2 2Cxyx2 0 2x0y! C.xoyx0y0/ Dx2 2CxyCconstant terms: Thus, the solution is x2 2CxyDC; which if desired can be solved to give yas a function of x. We can also check to make sure that our solution actually solves the ODE.  It may well turn out that Eq. (7.6) is not exact and that Eq. (7.9) is not satisfied. However, there always exists at least one and perhaps an infinity of integrating factors .x;y/such that .x;y/P.x;y/dxC .x;y/Q.x;y/dyD0 is exact. Unfortunately, an integrating factor is not always obvious or easy to find. A sys- tematic way to develop an integrating factor is known only when a first-order ODE is linear; this will be discussed in the subsection on linear first-order ODEs. Equations Homogeneous in xand y An ODE is said to be homogeneous (of order n) inxandyif the combined powers of xandyadd to nin all the terms of P.x;y/andQ.x;y/when the ODE is written as in Eq. (7.6). Note that this use of the term “homogeneous” has a different meaning than when it was used to describe a linear ODE as given in Eq. (7.1) with the term F.x/equal to zero, because it now applies to the combined power of xandy. A first-order ODE, which is homogeneous of order nin the present sense (and not nec- essarily linear), can be made separable by the substitution yDxv, with dyDx dvCvdx. This substitution causes the xdependence of all the terms of the equation containing dvto bexnC1, with all the terms containing dxhaving x-dependence xn. The variables xandv can then be separated. ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 335 Example 7.2.3 ANODE HOMOGENEOUS IN xAND y Consider the ODE .2xCy/dxCx dyD0; which is homogeneous in xandy. Making the substitution yDxv, with dyDx dvCvdx, the ODE becomes .2vC2/dxCx dvD0; which is separable, with solution lnxC1 2ln.vC1/DC, which is equivalent to x2.vC1/D C. Forming yDxv, the solution can be rearranged into yDC xx:  Isobaric Equations A generalization of the preceding subsection is to modify the definition of homogeneity by assigning different weights to xandy(note that corresponding weights must then also be assigned to dxanddy). If assigning unit weight to each instance of xordxand a weight mto each instance of yordymakes the ODE homogeneous as defined here, then the substitution yDxmvwill make the equation separable. We illustrate with an example. Example 7.2.4 ANISOBARIC ODE Here is an isobaric ODE: .x2y/dxCx dyD0: Assigning xweight 1, and yweight m, the term x2dxhas weight 3; the other two terms have weight 1Cm. Setting 3D1Cm, we find that all terms can be assigned equal weight if we take mD2. This means that we should make the substitution y Dx2v. Doing so, we get .1v/dxCx dvD0; which separates into dx xCdv vC1D0! lnxCln.vC1/DlnC;orx.vC1/DC: From this, we get vDC x1. Since yDx2v, the ODE has solution yDCxx2. ArfKen_Ch07-9780123846549.tex 336 Chapter 7 Ordinary Differential Equations Linear First-Order ODEs While nonlinear first-order ODEs can often (but not always) be solved using the strategies already presented, the situation is different for the linear first-order ODE because proce- dures exist for solving the most general equation of this type, which we write in the form dy dxCp.x/yDq.x/: (7.11) If our linear first-order ODE is exact, its solution is straightforward. If it is not exact, we make it exact by introducing an integrating factor .x/, so that the ODE becomes .x/dy dxC .x/p.x/yD .x/q.x/: (7.12) The reason for multiplication by .x/is to cause the left-hand side of Eq. (7.12) to become a perfect differential, so we require that .x/be such that d dx .x/y D .x/dy dxC .x/p.x/y: (7.13) Expanding the left-hand side of Eq. (7.13), that equation becomes .x/dy dxCd dxyD .x/dy dxC .x/p.x/y; so must satisfy d dxD .x/p.x/: (7.14) This is a separable equation and therefore soluble. Separating the variables and integrat- ing, we obtain Zd DxZ p.x/dx: We need not consider the lower limits of these integrals because they combine to yield a constant that does not affect the performance of the integrating factor and can be set to zero. Completing the evaluation, we reach .x/Dexp2 4xZ p.x/dx3 5: (7.15) With now known we proceed to integrate Eq. (7.12), which, because of Eq. (7.13), assumes the form d dxT .x/y.x/UD .x/q.x/; which can be integrated (and divided through by ) to yield y.x/D1 .x/2 4xZ .x/q.x/dxCC3 5y2.x/Cy1.x/: (7.16) ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 337 The two terms of Eq. (7.16) have an interesting interpretation. The term y1DC= .x/ is the general solution of the homogeneous equation obtained by replacing q.x/with zero. To see this, write the homogeneous equation as dy yDp.x/dx; which integrates to lnyDxZ p.x/dxCCDln CC: Taking the exponential of both sides and renaming eCasC, we get just yDC= .x/. The other term of Eq. (7.16), y2D1 .x/xZ .x/q.x/dx (7.17) corresponds to the right-hand side (source) term q.x/, and is a solution of the original inhomogeneous equation (as is obvious because Ccan be set to zero). We thus have the general solution to the inhomogeneous equation presented as a particular solution (or, in ODE parlance, a particular integral) plus the general solution to the corresponding homogeneous equation. The above observations illustrate the following theorem: The solution of an inhomogeneous first-order linear ODE is unique except for an arbi- trary multiple of the solution of the corresponding homogeneous ODE. To show this, suppose y1andy2both solve the inhomogeneous ODE, Eq. (7.11). Then, subtracting the equation for y2from that for y1, we have y0 1y0 2Cp.x/.y1y2/D0: This shows that y1y2is (at some scale) a solution of the homogeneous ODE. Remember that any solution of the homogeneous ODE remains a solution when multiplied by an arbitrary constant. We also have the theorem: A first-order linear homogeneous ODE has only one linearly independent solution. Two functions y1.x/andy2.x/are linearly dependent if there exist two constants aand b, both nonzero, that cause ay1Cby2to vanish for all x. In the present situation, this is equivalent to the statement that y1andy2are linearly dependent if they are proportional to each other. To prove the theorem, assume that the homogeneous ODE has the linearly independent solutions y1andy2. Then, from the homogeneous ODE, we have y0 1 y1Dp.x/Dy0 2 y2: Integrating the first and last members of this equation, we obtain lny1Dlny2CC;equivalent to y1DCy2; contradicting our assumption that y1andy2are linearly independent. ArfKen_Ch07-9780123846549.tex 338 Chapter 7 Ordinary Differential Equations Example 7.2.5 RL CIRCUIT For a resistance-inductance circuit Kirchoff’s law leads to Ld I.t/ dtCRI.t/DV.t/; where I.t/is the current, LandRare, respectively, constant values of the inductance and the resistance, and V.t/is the time-dependent input voltage. From Eq. (7.15), our integrating factor .t/is .t/DexptZR LdtDeRt=L: Then, by Eq. (7.16), I.t/DeRt=L2 4tZ eRt=LV.t/ LdtCC3 5; with the constant Cto be determined by an initial condition. For the special case V.t/DV0, a constant, I.t/DeRt=LV0 L:L ReRt=LCC DV0 RCCeRt=L: If the initial condition is I.0/D0, then CDV0=Rand I.t/DV0 Rh 1eRt=Li :  We close this section by pointing out that the inhomogeneous linear first-order ODE can also be solved by a method called variation of the constant, or alternatively variation of parameters, as follows. First, we solve the homogeneous ODE y0CpyD0by separation of variables as before, giving y0 yDp;lnyDxZ p.X/d XClnC;y.x/DCexp0 @xZ p.X/d X1 A: Next we allow the integration constant to become x-dependent, that is, C!C.x/. This is the reason the method is called “variation of the constant.” To prepare for substitution into the inhomogeneous ODE, we calculate y0: y0Dexp0 @xZ p.X/d X1 A pC.x/CC0.x/ Dpy.x/CC0.x/exp0 @xZ p.X/d X1 A: Making the substitution for y0into the inhomogeneous ODE y0CpyDq, some cancella- tion occurs, and we are left with C0.x/exp0 @xZ p.X/d X1 ADq; ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 339 which is a separable ODE for C.x/that integrates to yield C.x/DxZ exp0 @XZ p.Y/dY1 Aq.X/d X and yDC.x/exp0 @xZ p.X/d X1 A: This particular solution of the inhomogeneous ODE is in agreement with that called y2in Eq. (7.17). Exercises 7.2.1 From Kirchhoff’s law the current Iin an RC(resistance-capacitance) circuit (Fig. 7.1) obeys the equation Rd I dtC1 CID0: (a) Find I.t/. (b) For a capacitance of 10,000 mF charged to 100 V and discharging through a resis- tance of 1 M, find the current IfortD0and for tD100seconds. Note. The initial voltage is I0RorQ=C, where QDR1 0I.t/dt. 7.2.2 The Laplace transform of Bessel’s equation .nD0/leads to .s2C1/f0.s/Cs f.s/D0: Solve for f.s/. 7.2.3 The decay of a population by catastrophic two-body collisions is described by d N dtDkN2: This is a first-order, nonlinear differential equation. Derive the solution N.t/DN0 1Ct 01 ; where0D.kN 0/1. This implies an infinite population at tD 0. C + −R FIGURE 7.1 RC circuit. ArfKen_Ch07-9780123846549.tex 340 Chapter 7 Ordinary Differential Equations 7.2.4 The rate of a particular chemical reaction ACB!Cis proportional to the concentra- tions of the reactants AandB: dC.t/ dtD TA.0/C.t/UTB.0/C.t/U: (a) Find C.t/forA.0/6DB.0/. (b) Find C.t/forA.0/DB.0/. The initial condition is that C.0/D0. 7.2.5 A boat, coasting through the water, experiences a resisting force proportional to vn;v being the boat’s instantaneous velocity. Newton’s second law leads to mdv dtDkvn: Withv.tD0/Dv0;x.tD0/D0, integrate to find vas a function of time and vas a function of distance. 7.2.6 In the first-order differential equation dy=dxDf.x;y/, the function f.x;y/is a func- tion of the ratio y=x: dy dxDg.y=x/: Show that the substitution of uDy=xleads to a separable equation in uandx. 7.2.7 The differential equation P.x;y/dxCQ.x;y/dyD0 isexact. Show that its solution is of the form '.x;y/DxZ x0P.x;y/dxCyZ y0Q.x0;y/dyDconstant: 7.2.8 The differential equation P.x;y/dxCQ.x;y/dyD0 isexact. If '.x;y/DxZ x0P.x;y/dxCyZ y0Q.x0;y/dy; show that @' @xDP.x;y/;@' @yDQ.x;y/: Hence,'.x;y/Dconstant is a solution of the original differential equation. 7.2.9 Prove that Eq. (7.12) is exact in the sense of Eq. (7.9), provided that .x/satisfies Eq. (7.14). ArfKen_Ch07-9780123846549.tex 7.2 First-Order Equations 341 7.2.10 A certain differential equation has the form f.x/dxCg.x/h.y/dyD0; with none of the functions f.x/;g.x/;h.y/identically zero. Show that a necessary and sufficient condition for this equation to be exact is that g.x/Dconstant. 7.2.11 Show that y.x/Dexp2 4xZ p.t/dt3 58 < :xZ exp2 4sZ p.t/dt3 5q.s/dsCC9 = ; is a solution of dy dxCp.x/y.x/Dq.x/ by differentiating the expression for y.x/and substituting into the differential equation. 7.2.12 The motion of a body falling in a resisting medium may be described by mdv dtDmgbv when the retarding force is proportional to the velocity, v. Find the velocity. Evaluate the constant of integration by demanding that v.0/D0. 7.2.13 Radioactive nuclei decay according to the law d N dtD N; Nbeing the concentration of a given nuclide and , the particular decay constant. In a radioactive series of two different nuclides, with concentrations N1.t/andN2.t/, we have d N1 dtD 1N1; d N2 dtD1N12N2: Find N2.t/for the conditions N1.0/DN0andN2.0/D0. 7.2.14 The rate of evaporation from a particular spherical drop of liquid (constant density) is proportional to its surface area. Assuming this to be the sole mechanism of mass loss, find the radius of the drop as a function of time. 7.2.15 In the linear homogeneous differential equation dv dtDav the variables are separable. When the variables are separated, the equation is exact. Solve this differential equation subject to v.0/Dv0by the following three methods: (a) Separating variables and integrating. (b) Treating the separated variable equation as exact. ArfKen_Ch07-9780123846549.tex 342 Chapter 7 Ordinary Differential Equations (c) Using the result for a linear homogeneous differential equation. ANS.v.t/Dv0eat. 7.2.16 (a) Solve Example 7.2.1, assuming that the parachute opens when the parachutist’s velocity has reached viD60mi/h (regard this time as tD0). Findv.t/. (b) For a skydiver in free fall use the friction coefficient bD0:25 kg/m and mass mD70kg. What is the limiting velocity in this case? 7.2.17 Solve the ODE .xy2y/dxCx dyD0: 7.2.18 Solve the ODE .x2y2ey=x/dxC.x2Cxy/ey=xdyD0: Hint. Note that the quantity y=xin the exponents is of combined degree zero and does not affect the determination of homogeneity. 7.3 ODE S WITH CONSTANT COEFFICIENTS Before addressing second-order ODEs, the main topic of this chapter, we discuss a special- ized, but frequently occurring class of ODEs that are not constrained to be of specific order, namely those that are linear and whose homogeneous terms have constant coefficients. The generic equation of this type is dny dxnCan1dn1y dxn1CC a1dy dxCa0yDF.x/: (7.18) The homogeneous equation corresponding to Eq. (7.18) has solutions of the form yDemx, where mis a solution of the algebraic equation mnCan1mn1CC a1mCa0D0; as may be verified by substitution of the assumed form of the solution. In the case that the mequation has a multiple root, the above prescription will not yield the full set of nlinearly independent solutions for the original nth order ODE. If one then considers the limiting process in which two roots approach each other, it is possible to conclude that if emxis a solution, then so is d emx=dmDxemx. A triple root would have solutions emx,xemx,x2emx, etc. Example 7.3.1 HOOKE’S LAW SPRING A mass Mattached to a Hooke’s Law spring (of spring constant k) is in oscillatory motion. Letting ybe the displacement of the mass from its equilibrium position, Newton’s law of motion takes the form Md2y dt2Dky; ArfKen_Ch07-9780123846549.tex 7.4 Second-Order Linear ODEs 343 which is an ODE of the form y00Ca0yD0, with a0DCk=M. The general solution to this ODE is of the form C1em1tCC2em2t, where m1andm2are the solutions of the algebraic equation m2Ca0D0. The values of m1andm2arei!, where!Dpk=M, so the ODE has solution y.t/DC1eCi!tCC2ei!t: Since the ODE is homogeneous, we may alternatively describe its general solution using arbitrary linear combinations of the above two terms. This permits us to combine them to obtain forms that are real and therefore appropriate to the current problem. Noting that ei!tCei!t 2Dcos!tandei!tei!t 2iDsin!t; a convenient alternate form is y.t/DC1cos!tCC2sin!t: The solution to a specific oscillation problem will now involve fitting the coefficients C1andC2to the initial conditions, as for example y.0/andy0.0/.  Exercises Find the general solutions to the following ODEs. Write the solutions in forms that are entirely real (i.e., that contain no complex quantities). 7.3.1 y0002y00y0C2yD0: 7.3.2 y0002y00Cy02yD0: 7.3.3 y0003y0C2yD0: 7.3.4 y00C2y0C2yD0: 7.4 S ECOND -ORDER LINEAR ODE S We now turn to the main topic of this chapter, second-order linear ODEs. These are of particular importance because they arise in the most frequently used methods for solving PDEs in quantum mechanics, electromagnetic theory, and other areas in physics. Unlike the first-order linear ODE, we do not have a universally applicable closed-form solution, and in general it is found advisable to use methods that produce solutions in the form of power series. As a precursor to the general discussion of series-solution methods, we begin by examining the notion of singularity as applied to ODEs. Singular Points The concept of singularity of an ODE is important to us for two reasons: (1) it is useful for classifying ODEs and identifying those that can be transformed into common forms (discussed later in this subsection), and (2) it bears on the feasibility of finding series ArfKen_Ch07-9780123846549.tex 344 Chapter 7 Ordinary Differential Equations solutions to the ODE. This feasibility is the topic of Fuchs’ theorem (to be discussed shortly). When a linear homogeneous second-order ODE is written in the form y00CP.x/y0CQ.x/yD0; (7.19) points x0for which P.x/andQ.x/are finite are termed ordinary points of the ODE. However, if either P.x/orQ.x/diverge as x!x0, the point x0is called a singular point. Singular points are further classified as regular orirregular (the latter also sometimes called essential singularities): A singular point x0isregular if either P.x/orQ.x/diverges there, but .xx0/P.x/ and.xx0/2Q.x/remain finite. A singular point x0isirregular ifP.x/diverges faster than 1=.xx0/so that.x x0/P.x/goes to infinity as x!x0, or if Q.x/diverges faster than 1=.xx0/2so that .xx0/2Q.x/goes to infinity as x!x0. These definitions hold for all finite values of x0. To analyze the behavior at x!1 , we setxD1=z, substitute into the differential equation, and examine the behavior in the limit z!0. The ODE, originally in the dependent variable y.x/, will now be written in terms ofw.z/, defined asw.z/Dy.z1/. Converting the derivatives, y0Ddy.x/ dxDdy.z1/ dzdz dxDdw.z/ dz 1 x2 Dz2w0; (7.20) y00Ddy0 dzdz dxD.z2/d dz z2w0 Dz4w00C2z3w0: (7.21) Using Eqs. (7.20) and (7.21), we transform Eq. (7.19) into z4w00C 2z3z2P.z1/ w0CQ.z1/wD0: (7.22) Dividing through by z4to place the ODE in standard form, we see that the possibility of a singularity at zD0depends on the behavior of 2zP.z1/ z2andQ.z1/ z4: If these two expressions remain finite at zD0, the point xD1 is an ordinary point. If they diverge no more rapidly than 1=zand1=z2, respectively, xD1 is a regular singular point; otherwise it is an irregular singular point (an essential singularity). Example 7.4.1 BESSEL’S EQUATION Bessel’s equation is x2y00Cxy0C.x2n2/yD0: ArfKen_Ch07-9780123846549.tex 7.4 Second-Order Linear ODEs 345 Comparing it with Eq. (7.19), we have P.x/D1 x;Q.x/D1n2 x2; which shows that the point xD0is a regular singularity. By inspection we see that there are no other singularities in the finite range. As x!1 (z!0), from Eq. (7.22) we have the coefficients 2zz z2and1n2z2 z4: Since the latter expression diverges as 1=z4, the point xD1 is an irregular, or essential, singularity.  Table 7.1 lists the singular points of a number of ODEs of importance in physics. It will be seen that the first three equations in Table 7.1, the hypergeometric, Legendre, and Chebyshev, all have three regular singular points. The hypergeometric equation, with reg- ular singularities at 0, 1, and 1, is taken as the standard, the canonical form. The solutions of the other two may then be expressed in terms of its solutions, the hypergeometric func- tions. This is done in Chapter 18. In a similar manner, the confluent hypergeometric equation is taken as the canonical form of a linear second-order differential equation with one regular and one irregular sin- gular point. Table 7.1 Singularities of Some Important ODEs. Equation Regular Irregular Singularity Singularity xD xD 1. Hypergeometric 0;1;1  x.x1/y00CT.1CaCb/xCcUy0CabyD0 2. Legendrea1;1;1 .1x2/y002xy0Cl.lC1/yD0 3. Chebyshev 1;1;1 .1x2/y00xy0Cn2yD0 4. Confluent hypergeometric 0 1 xy00C.cx/y0ayD0 5. Bessel 0 1 x2y00Cxy0C.x2n2/yD0 6. Laguerrea0 1 xy00C.1x/y0CayD0 7. Simple harmonic oscillator  1 y00C!2yD0 8. Hermite  1 y002xy0C2 yD0 aThe associated equations have the same singular points. ArfKen_Ch07-9780123846549.tex 346 Chapter 7 Ordinary Differential Equations Exercises 7.4.1 Show that Legendre’s equation has regular singularities at xD1; 1, and1. 7.4.2 Show that Laguerre’s equation, like the Bessel equation, has a regular singularity at xD0and an irregular singularity at xD1 . 7.4.3 Show that Chebyshev’s equation, like the Legendre equation, has regular singularities atxD1; 1, and1. 7.4.4 Show that Hermite’s equation has no singularity other than an irregular singularity at xD1 . 7.4.5 Show that the substitution x!1x 2;aDl;bDlC1; cD1 converts the hypergeometric equation into Legendre’s equation. 7.5 S ERIES SOLUTIONS —FROBENIUS ’ METHOD In this section we develop a method of obtaining solution(s) of the linear, second-order, homogeneous ODE. For the moment, we develop the mechanics of the method. After studying examples, we return to discuss the conditions under which we can expect these series solutions to exist. Consider a linear, second-order, homogeneous ODE, in the form d2y dx2CP.x/dy dxCQ.x/yD0: (7.23) In this section we develop (at least) one solution of Eq. (7.23) by expansion about the point xD0. In the next section we develop the second, independent solution and prove that no third, independent solution exists. Therefore the most general solution ofEq. (7.23) may be written in terms of the two independent solutions as y.x/Dc1y1.x/Cc2y2.x/: (7.24) Our physical problem may lead to a nonhomogeneous, linear, second-order ODE, d2y dx2CP.x/dy dxCQ.x/yDF.x/: (7.25) The function on the right, F.x/, typically represents a source (such as electrostatic charge) or a driving force (as in a driven oscillator). Methods for solving this inhomogeneous ODE are also discussed later in this chapter and, using Laplace transform techniques, in Chapter 20. Assuming a single particular integral (i.e., specific solution), yp, of the in- homogeneous ODE to be available, we may add to it any solution of the corresponding homogeneous equation, Eq. (7.23), and write the most general solution of Eq. (7.25) as y.x/Dc1y1.x/Cc2y2.x/Cyp.x/: (7.26) In many problems, the constants c1andc2will be fixed by boundary conditions. ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 347 For the present, we assume that F.x/D0, and that therefore our differential equation is homogeneous. We shall attempt to develop a solution of our linear, second-order, homoge- neous differential equation, Eq. (7.23), by substituting into it a power series with undeter- mined coefficients. Also available as a parameter is the power of the lowest nonvanishing term of the series. To illustrate, we apply the method to two important differential equa- tions. First Example—Linear Oscillator Consider the linear (classical) oscillator equation d2y dx2C!2yD0; (7.27) which we have already solved by another method in Example 7.3.1. The solutions we found there were yDsin!xandcos!x. We try y.x/Dxs.a0Ca1xCa2x2Ca3x3C/ D1X jD0ajxsCj;a06D0; (7.28) with the exponent sand all the coefficients ajstill undetermined. Note that sneed not be an integer. By differentiating twice, we obtain dy dxD1X jD0aj.sCj/xsCj1; d2y dx2D1X jD0aj.sCj/.sCj1/xsCj2: By substituting into Eq. (7.27), we have 1X jD0aj.sCj/.sCj1/xsCj2C!21X jD0ajxsCjD0: (7.29) From our analysis of the uniqueness of power series (Chapter 1), we know that the coef- ficient of each power of xon the left-hand side of Eq. (7.29) must vanish individually, xs being an overall factor. The lowest power of xappearing in Eq. (7.29) isxs2, occurring only for jD0in the first summation. The requirement that this coefficient vanish yields a0s.s1/D0: Recall that we chose a0as the coefficient of the lowest nonvanishing term of the series in Eq. (7.28), so that, by definition, a06D0. Therefore we have s.s1/D0: (7.30) ArfKen_Ch07-9780123846549.tex 348 Chapter 7 Ordinary Differential Equations This equation, coming from the coefficient of the lowest power of x, is called the indicial equation. The indicial equation and its roots are of critical importance to our analysis. Clearly, in this example it informs us that either sD0orsD1, so that our series solution must start either with an x0or an x1term. Looking further at Eq. (7.29), we see that the next lowest power of x, namely xs1, also occurs uniquely (for jD1in the first summation). Setting the coefficient of xs1to zero, we have a1.sC1/sD0: This shows that if sD1, we must have a1D0. However, if sD0, this equation imposes no requirement on the coefficient set. Before considering further the two possibilities for s, we return to Eq. (7.29) and demand that the remaining net coefficients vanish. The contributions to the coefficient of xsCj, (j0/, come from the term containing ajC2in the first summation and from that with aj in the second. Because we have already dealt with jD0andjD1in the first summation, when we have used all j0, we will have used all the terms of both series. For each value ofj, the vanishing of the net coefficient of xsCjresults in ajC2.sCjC2/.sCjC1/C!2ajD0; equivalent to ajC2Daj!2 .sCjC2/.sCjC1/: (7.31) This is a two-term recurrence relation.1In the present problem, given aj, Eq. (7.31) permits us to compute ajC2and then ajC4;ajC6, and so on, continuing as far as desired. Thus, if we start with a0, we can make the even coefficients a2,a4, . . . , but we obtain no information about the odd coefficients a1,a3,a5, . . . . But because a1is arbitrary if sD0 and necessarily zero if sD1, let us set it equal to zero, and then, by Eq. (7.31), a3Da5Da7DD 0I the result is that all the odd-numbered coefficients vanish. Returning now to Eq. (7.30), our indicial equation, we first try the solution sD0. The recurrence relation, Eq. (7.31), becomes ajC2Daj!2 .jC2/.jC1/; (7.32) 1In some problems, the recurrence relation may involve more than two terms; its exact form will depend on the functions P.x/ andQ.x/of the ODE. ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 349 which leads to a2Da0!2 12D!2 2Wa0; a4Da2!2 34DC!4 4Wa0; a6Da4!2 56D!6 6Wa0;and so on. By inspection (and mathematical induction, see Section 1.4), a2nD.1/n!2n .2n/Wa0; (7.33) and our solution is y.x/sD0Da0 1.!x/2 2WC.!x/4 4W.!x/6 6WC Da0cos!x: (7.34) If we choose the indicial equation root sD1from Eq. (7.30), the recurrence relation of Eq. (7.31) becomes ajC2Daj!2 .jC3/.jC2/: (7.35) Evaluating this successively for jD0, 2, 4, . . . , we obtain a2Da0!2 23D!2 3Wa0; a4Da2!2 45DC!4 5Wa0; a6Da4!2 67D!6 7Wa0;and so on. Again, by inspection and mathematical induction, a2nD.1/n!2n .2nC1/Wa0: (7.36) For this choice, sD1, we obtain y.x/sD1Da0x 1.!x/2 3WC.!x/4 5W.!x/6 7WC Da0 ! .!x/.!x/3 3WC.!x/5 5W.!x/7 7WC Da0 !sin!x: (7.37) ArfKen_Ch07-9780123846549.tex 350 Chapter 7 Ordinary Differential Equations I II III IV a0k(k−1) a1(k+1)k a2(k+2)(k+1) a0ω2a3(k+3)(k+2) a1ω2xk+ xk+xk+1+…xk−2+ xk+1+ + xk+1+…=0 =0 =0=0 =0 FIGURE 7.2 Schematics of series solution. For future reference we note that the ODE solution from the indicial equation root sD0 consisted only of even powers of x, while the solution from the root sD1contained only odd powers. To summarize this approach, we may write Eq. (7.29) schematically as shown in Fig. 7.2. From the uniqueness of power series (Section 1.2), the total coefficient of each power of xmust vanish—all by itself. The requirement that the first coefficient vanish (I) leads to the indicial equation, Eq. (7.30). The second coefficient is han- dled by setting a1D0 (II). The vanishing of the coefficients of xs(and higher pow- ers, taken one at a time) is ensured by imposing the recurrence relation, Eq. (7.31), (III), (IV). This expansion in power series, known as Frobenius’ method, has given us two series solutions of the linear oscillator equation. However, there are two points about such series solutions that must be strongly emphasized: 1. The series solution should always be substituted back into the differential equation, to see if it works, as a precaution against algebraic and logical errors. If it works, it is a solution. 2. The acceptability of a series solution depends on its convergence (including asymp- totic convergence). It is quite possible for Frobenius’ method to give a series solution that satisfies the original differential equation when substituted in the equation but that does notconverge over the region of interest. Legendre’s differential equation (examined in Section 8.3) illustrates this situation. Expansion about x0 Equation (7.28) is an expansion about the origin, x0D0. It is perfectly possible to replace Eq. (7.28) with y.x/D1X jD0aj.xx0/sCj;a06D0: (7.38) Indeed, for the Legendre, Chebyshev, and hypergeometric equations, the choice x0D1 has some advantages. The point x0should not be chosen at an essential singularity, or Frobenius’ method will probably fail. The resultant series ( x0an ordinary point or regular singular point) will be valid where it converges. You can expect a divergence of some sort whenjxx0jDjz1x0j, where z1is the ODE’s closest singularity to x0(in the complex plane). ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 351 Symmetry of Solutions Let us note that for the classical oscillator problem we obtained one solution of even sym- metry, y1.x/Dy1.x/, and one of odd symmetry, y2.x/Dy2.x/. This is not just an accident but a direct consequence of the form of the ODE. Writing a general homogeneous ODE as L.x/y.x/D0; (7.39) in which L.x/is the differential operator, we see that for the linear oscillator equation, Eq. (7.27), L.x/is even under parity; that is, L.x/DL.x/: Whenever the differential operator has a specific parity or symmetry, either even or odd, we may interchange Cxandx, and Eq. (7.39) becomes L.x/y.x/D0: Clearly, if y.x/is a solution of the differential equation, y.x/is also a solution. Then, either y.x/andy.x/are linearly dependent (i.e., proportional), meaning that yis either even or odd, or they are linearly independent solutions that can be combined into a pair of solutions, one even, and one odd, by forming yevenDy.x/Cy.x/;yoddDy.x/y.x/: For the classical oscillator example, we obtained two solutions; our method for finding them caused one to be even, the other odd. If we refer back to Section 7.4 we can see that Legendre, Chebyshev, Bessel, simple har- monic oscillator, and Hermite equations are all based on differential operators with even parity; that is, their P.x/inEq. (7.19) is odd and Q.x/even. Solutions of all of them may be presented as series of even powers of xor separate series of odd powers of x. The Laguerre differential operator has neither even nor odd symmetry; hence its solutions cannot be expected to exhibit even or odd parity. Our emphasis on parity stems primarily from the importance of parity in quantum mechanics. We find that in many problems wave functions are either even or odd, meaning that they have a definite parity. Most interac- tions (beta decay is the big exception) are also even or odd, and the result is that parity is conserved. A Second Example—Bessel’s Equation This attack on the linear oscillator was perhaps a bit too easy. By substituting the power series, Eq. (7.28), into the differential equation, Eq. (7.27), we obtained two independent solutions with no trouble at all. To get some idea of other things that can happen, we try to solve Bessel’s equation, x2y00Cxy0C.x2n2/yD0: (7.40) ArfKen_Ch07-9780123846549.tex 352 Chapter 7 Ordinary Differential Equations Again, assuming a solution of the form y.x/D1X jD0ajxsCj; we differentiate and substitute into Eq. (7.40). The result is 1X jD0aj.sCj/.sCj1/xsCjC1X jD0aj.sCj/xsCj C1X jD0ajxsCjC21X jD0ajn2xsCjD0: (7.41) By setting jD0, we get the coefficient of xs, the lowest power of xappearing on the left-hand side, a0 s.s1/Csn2 D0; (7.42) and again a06D0by definition. Equation (7.42) therefore yields the indicial equation s2n2D0; (7.43) with solutions sDn . We need also to examine the coefficient of xsC1. Here we obtain a1T.sC1/sCsC1n2UD0; or a1.sC1n/.sC1Cn/D0: (7.44) ForsDn , neither sC1nnorsC1Cnvanishes and we must require a1D0. Proceeding to the coefficient of xsCjforsDn, we see that it is the term containing aj in the first, second, and fourth terms of Eq. (7.41), but is that containing aj2in the third term. By requiring the overall coefficient of xsCjto vanish, we obtain ajT.nCj/.nCj1/C.nCj/n2UCaj2D0: When jis replaced by jC2, this can be rewritten for j0as ajC2Daj1 .jC2/.2nCjC2/; (7.45) which is the desired recurrence relation. Repeated application of this recurrence relation leads to a2Da01 2.2nC2/Da0nW 221W.nC1/W; a4Da21 4.2nC4/Da0nW 242W.nC2/W; a6Da41 6.2nC6/Da0nW 263W.nC3/W;and so on, ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 353 and in general, a2pD.1/p a0nW 22ppW.nCp/W: (7.46) Inserting these coefficients in our assumed series solution, we have y.x/Da0xn 1nWx2 221W.nC1/WCnWx4 242W.nC2/W : (7.47) In summation form, y.x/Da01X jD0.1/jnWxnC2j 22jjW.nCj/W Da02nnW1X jD0.1/j 1 jW.nCj/Wx 2nC2j : (7.48) In Chapter 14 the final summation (with a0D1=2nnW) is identified as the Bessel function Jn.x/: Jn.x/D1X jD0.1/j 1 jW.nCj/Wx 2nC2j : (7.49) Note that this solution, Jn.x/;has either even or odd symmetry,2as might be expected from the form of Bessel’s equation. When sDn andnis not an integer, we may generate a second distinct series, to be labeled Jn.x/. However, whennis a negative integer, trouble develops. The recurrence relation for the coefficients ajis still given by Eq. (7.45), but with 2nreplaced by2n. Then, when jC2D2norjD2.n1/, the coefficient ajC2blows up and Frobenius’ method does not produce a solution consistent with our assumption that the series starts with xn. By substituting in an infinite series, we have obtained two solutions for the linear oscil- lator equation and one for Bessel’s equation (two if nis not an integer). To the questions “Can we always do this? Will this method always work?” the answer is “No, we cannot always do this. This method of series solution will not always work.’ ’ Regular and Irregular Singularities The success of the series substitution method depends on the roots of the indicial equation and the degree of singularity of the coefficients in the differential equation. To understand better the effect of the equation coefficients on this naive series substitution approach, 2Jn.x/is an even function if nis an even integer, and an odd function if nis an odd integer. For nonintegral n,Jnhas no such simple symmetry. ArfKen_Ch07-9780123846549.tex 354 Chapter 7 Ordinary Differential Equations consider four simple equations: y006 x2yD0; (7.50) y006 x3yD0; (7.51) y00C1 xy0b2 x2yD0; (7.52) y00C1 x2y0b2 x2yD0: (7.53) The reader may show easily that for Eq. (7.50) the indicial equation is s2s6D0; giving sD3andsD2 . Since the equation is homogeneous in x(counting d2=dx2as x2), there is no recurrence relation. However, we are left with two perfectly good solu- tions, x3andx2. Equation (7.51) differs from Eq. (7.50) by only one power of x, but this sends the indicial equation to 6a0D0; with no solution at all, for we have agreed that a06D0. Our series substitution worked for Eq. (7.50), which had only a regular singularity, but broke down at Eq. (7.51), which has an irregular singular point at the origin. Continuing with Eq. (7.52), we have added a term y0=x. The indicial equation is s2b2D0; but again, there is no recurrence relation. The solutions are yDxbandxb, both perfectly acceptable one-term series. When we change the power of xin the coefficient of y0from1to2, inEq. (7.53), there is a drastic change in the solution. The indicial equation (with only the y0term con- tributing) becomes sD0: There is a recurrence relation, ajC1DCajb2j.j1/ jC1: Unless the parameter bis selected to make the series terminate, we have lim j!1 ajC1 aj Dlim j!1j.jC1/ jC1 Dlim j!1j2 jD1: ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 355 Hence our series solution diverges for all x6D0. Again, our method worked for Eq. (7.52) with a regular singularity but failed when we had the irregular singularity of Eq. (7.53). Fuchs’ Theorem The answer to the basic question as to when the method of series substitution can be expected to work is given by Fuchs’ theorem, which asserts that we can always obtain at least one power-series solution, provided that we are expanding about a point which is an ordinary point or at worst a regular singular point. If we attempt an expansion about an irregular or essential singularity, our method may fail as it did for Eqs. (7.51) and(7.53). Fortunately, the more important equations of mathe- matical physics, listed in Section 7.4, have no irregular singularities in the finite plane. Further discussion of Fuchs’ theorem appears in Section 7.6. From Table 7.1, Section 7.4, infinity is seen to be a singular point for all the equations considered. As a further illustration of Fuchs’ theorem, Legendre’s equation (with infinity as a regular singularity) has a convergent series solution in negative powers of the argu- ment (Section 15.6). In contrast, Bessel’s equation (with an irregular singularity at infinity) yields asymptotic series (Sections 12.6 and 14.6). Although only asymptotic, these solu- tions are nevertheless extremely useful. Summary If we are expanding about an ordinary point or at worst about a regular singularity, the series substitution approach will yield at least one solution (Fuchs’ theorem). Whether we get one or two distinct solutions depends on the roots of the indicial equation. 1. If the two roots of the indicial equation are equal, we can obtain only one solution by this series substitution method. 2. If the two roots differ by a nonintegral number, two independent solutions may be obtained. 3. If the two roots differ by an integer, the larger of the two will yield a solution, while the smaller may or may not give a solution, depending on the behavior of the coefficients. The usefulness of a series solution for numerical work depends on the rapidity of con- vergence of the series and the availability of the coefficients. Many ODEs will not yield nice, simple recurrence relations for the coefficients. In general, the available series will probably be useful for very small jxj(orjxx0j). Computers can be used to determine additional series coefficients using a symbolic language, such as Mathematica3or Maple.4 Often, however, for numerical work a direct numerical integration will be preferred. 3S. Wolfram, Mathematica: A System for Doing Mathematics by Computer. Reading, MA. Addison Wesley (1991). 4A. Heck, Introduction to Maple. New York: Springer (1993). ArfKen_Ch07-9780123846549.tex 356 Chapter 7 Ordinary Differential Equations Exercises 7.5.1 Uniqueness theorem. The function y.x/satisfies a second-order, linear, homogeneous differential equation. At xDx0;y.x/Dy0anddy=dxDy0 0. Show that y.x/is unique, in that no other solution of this differential equation passes through the points .x0;y0/ with a slope of y0 0. Hint. Assume a second solution satisfying these conditions and compare the Taylor series expansions. 7.5.2 A series solution of Eq. (7.23) is attempted, expanding about the point xDx0. Ifx0is an ordinary point, show that the indicial equation has roots sD0;1. 7.5.3 In the development of a series solution of the simple harmonic oscillator (SHO) equa- tion, the second series coefficient a1was neglected except to set it equal to zero. From the coefficient of the next-to-the-lowest power of x;xs1, develop a second-indicial type equation. (a) (SHO equation with sD0). Show that a1, may be assigned any finite value (including zero). (b) (SHO equation with sD1). Show that a1must be set equal to zero. 7.5.4 Analyze the series solutions of the following differential equations to see when a1may be set equal to zero without irrevocably losing anything and when a1must be set equal to zero. (a) Legendre, (b) Chebyshev, (c) Bessel, (d) Hermite. ANS. (a) Legendre, (b) Chebyshev, and (d) Hermite: For sD0,a1 may be set equal to zero; for sD1,a1must be set equal to zero. (c) Bessel: a1must be set equal to zero (except for sDnD1 2). 7.5.5 Obtain a series solution of the hypergeometric equation x.x1/y00CT.1CaCb/xcUy0CabyD0: Test your solution for convergence. 7.5.6 Obtain two series solutions of the confluent hypergeometric equation xy00C.cx/y0ayD0: Test your solutions for convergence. 7.5.7 A quantum mechanical analysis of the Stark effect (parabolic coordinates) leads to the differential equation d d du d C1 2EC m2 41 4F2 uD0: Here is a constant, Eis the total energy, and Fis a constant such that Fzis the potential energy added to the system by the introduction of an electric field. ArfKen_Ch07-9780123846549.tex 7.5 Series Solutions—Frobenius’ Method 357 Using the larger root of the indicial equation, develop a power-series solution about D0. Evaluate the first three coefficients in terms of ao. ANS. Indicial equation s2m2 4D0; u./Da0m=2 1 mC1C 2 2.mC1/.mC2/E 4.mC2/ 2C : Note that the perturbation Fdoes not appear until a3is included. 7.5.8 For the special case of no azimuthal dependence, the quantum mechanical analysis of the hydrogen molecular ion leads to the equation d d .12/du d C uC 2uD0: Develop a power-series solution for u./. Evaluate the first three nonvanishing coeffi- cients in terms of a0. ANS. Indicial equation s.s1/D0; ukD1Da0 1C2 62C.2 /.12 / 120 20 4C : 7.5.9 To a good approximation, the interaction of two nucleons may be described by a mesonic potential VDAeax x; attractive for Anegative. Show that the resultant Schrödinger wave equation Nh2 2md2 dx2C.EV/ D0 has the following series solution through the first three nonvanishing coefficients: Da0 xC1 2A0x2C1 61 2A02E0a A0 x3C ; where the prime indicates multiplication by 2m=Nh2. 7.5.10 If the parameter b2inEq. (7.53) is equal to 2, Eq. (7.53) becomes y00C1 x2y02 x2yD0: From the indicial equation and the recurrence relation, derive a solution yD1C2xC 2x2. Verify that this is indeed a solution by substituting back into the differential equation. ArfKen_Ch07-9780123846549.tex 358 Chapter 7 Ordinary Differential Equations 7.5.11 The modified Bessel function I0.x/satisfies the differential equation x2d2 dx2I0.x/Cxd dxI0.x/x2I0.x/D0: Given that the leading term in an asymptotic expansion is known to be I0.x/ex p 2x; assume a series of the form I0.x/ex p 2xn 1Cb1x1Cb2x2Co : Determine the coefficients b1andb2. ANS. b1D1 8,b2D9 128. 7.5.12 The even power-series solution of Legendre’s equation is given by Exercise 8.3.1. Take a0D1andnnot an even integer, say nD0:5. Calculate the partial sums of the series through x200,x400,x600,:::,x2000forxD0:95.0:01/1:00 . Also, write out the individ- ual term corresponding to each of these powers. Note. This calculation does notconstitute proof of convergence at xD0:99 or diver- gence at xD1:00, but perhaps you can see the difference in the behavior of the sequence of partial sums for these two values of x. 7.5.13 (a) The odd power-series solution of Hermite’s equation is given by Exercise 8.3.3. Take a0D1. Evaluate this series for D0;xD1;2;3. Cut off your calculation after the last term calculated has dropped below the maximum term by a factor of 106or more. Set an upper bound to the error made in ignoring the remaining terms in the infinite series. (b) As a check on the calculation of part (a), show that the Hermite series yodd. D0/ corresponds toRx 0exp.x2/dx. (c) Calculate this integral for xD1;2;3. 7.6 O THER SOLUTIONS InSection 7.5 a solution of a second-order homogeneous ODE was developed by substi- tuting in a power series. By Fuchs’ theorem this is possible, provided the power series is an expansion about an ordinary point or a nonessential singularity.5There is no guarantee that this approach will yield the two independent solutions we expect from a linear second- order ODE. In fact, we shall prove that such an ODE has at most two linearly independent solutions. Indeed, the technique gave only one solution for Bessel’s equation ( nan integer). In this section we also develop two methods of obtaining a second independent solution: an integral method and a power series containing a logarithmic term. First, however, we consider the question of independence of a set of functions. 5This is why the classification of singularities in Section 7.4 is of vital importance. ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 359 Linear Independence of Solutions In Chapter 2 we introduced the concept of linear dependence for forms of the type a1x1C a2x2C:::, and identified a set of such forms as linearly dependent if any one of the forms could be written as a linear combination of others. We need now to extend the concept to a set of functions '. The criterion for linear dependence of a set of functions of a variable xis the existence of a relation of the form X k'.x/D0; (7.54) in which not all the coefficients kare zero. The interpretation we attach to Eq. (7.54) is that it indicates linear dependence if it is satisfied for all relevant values of x. Isolated points or partial ranges of satisfaction of Eq. (7.54) do not suffice to indicate linear dependence. The essential idea being conveyed here is that if there is linear dependence, the function space spanned by the '.x/can be spanned using less than all of them. On the other hand, if the only global solution of Eq. (7.54) iskD0for all;the set of functions '.x/is said to be linearly independent. If the members of a set of functions are mutually orthogonal, then they are automatically linearly independent. To establish this, consider the evaluation of SD*X k' X k'+ for a set of orthonormal 'and with arbitrary values of the coefficients k. Because of the orthonormality, Sevaluates toP jkj2, and will be nonzero (showing thatP k'6D0) unless all the kvanish. We now proceed to consider the ramifications of linear dependence for solutions of ODEs, and for that purpose it is appropriate to assume that the functions '.x/are differ- entiable as needed. Then, differentiating Eq. (7.54) repeatedly, with the assumption that it is valid for all x, we generate a set of equations X k'0 .x/D0; X k'00 .x/D0; continuing until we have generated as many equations as the number of values. This gives us a set of homogeneous linear equations in which kare the unknown quantities. By Section 2.1 there is a solution other than all kD0only if the determinant of the coefficients of the kvanishes. This means that the linear dependence we have assumed by accepting Eq. (7.54) implies that '1'2::: ' n '0 1'0 2::: '0 n ::: ::: ::: ::: '.n1/ 1'.n1/ 2::: '.n1/ n D0: (7.55) ArfKen_Ch07-9780123846549.tex 360 Chapter 7 Ordinary Differential Equations This determinant is called the Wronskian, and the analysis leading to Eq. (7.55) shows that: 1. If the Wronskian is not equal to zero, then Eq. (7.54) has no solution other than kD0. The set of functions 'is therefore linearly independent. 2. If the Wronskian vanishes at isolated values of the argument, this does not prove linear dependence. However, if the Wronskian is zero over the entire range of the variable, the functions 'are linearly dependent over this range.6 Example 7.6.1 LINEAR INDEPENDENCE The solutions of the linear oscillator equation, Eq. (7.27), are '1Dsin!x,'2Dcos!x. The Wronskian becomes sin!x cos!x !cos!x!sin!x D!6D0: These two solutions, '1and'2, are therefore linearly independent. For just two functions this means that one is not a multiple of the other, which is obviously true here. Incidentally, you know that sin!xD.1cos2!x/1=2; but this is notalinear relation, of the form of Eq. (7.54).  Example 7.6.2 LINEAR DEPENDENCE For an illustration of linear dependence, consider the solutions of the ODE d2'.x/ dx2D'.x/: This equation has solutions '1Dexand'2Dex, and we add'3Dcosh x, also a solution. The Wronskian is exexcosh x exexsinhx exexcosh x D0: The determinant vanishes for all xbecause the first and third rows are identical. Hence ex,ex, and cosh xare linearly dependent, and, indeed, we have a relation of the form of Eq. (7.54): exCex2 cosh xD0with k6D0:  6Compare H. Lass, Elements of Pure and Applied Mathematics, New York: McGraw-Hill (1957), p. 187, for proof of this assertion. It is assumed that the functions have continuous derivatives and that at least one of the minors of the bottom row of Eq. (7.55) (Laplace expansion) does not vanish in Ta;bU, the interval under consideration. ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 361 Number of Solutions Now we are ready to prove the theorem that a second-order homogeneous ODE has two linearly independent solutions. Suppose y1,y2,y3are three solutions of the homogeneous ODE, Eq. (7.23). Then we form the Wronskian WjkDyjy0 ky0 jykof any pair yj,ykof them and note also that W0 jkD.y0 jy0 kCyjy00 k/.y00 jykCy0 jy0 k/ Dyjy00 ky00 jyk: (7.56) Next we divide the ODE by yand move Q.x/to its right-hand side (where it becomes Q.x/), so, for solutions yjandyk: y00 j yjCP.x/y0 j yjDQ.x/Dy00 k ykCP.x/y0k yk: Taking now the first and third members of this equation, multiplying by yjykand rearrang- ing, we find that .yjy00 ky00 jyk/CP.x/.yjy0 ky0 jyk/D0; which simplifies for any pair of solutions to W0 jkDP.x/Wjk: (7.57) Finally, we evaluate the Wronskian of all three solutions, expanding it by minors along the second row and identifying each term as containing a W0 i jas given by Eq. (7.56): WD y1y2y3 y0 1y0 2y0 3 y00 1y00 2y00 3 Dy0 1W0 23Cy0 2W0 13y0 3W0 12: We now use Eq. (7.57) to replace each W0 i jbyP.x/Wi jand then reassemble the minors into a 33determinant, which vanishes because it contains two identical rows: WDP.x/ y0 1W23y0 2W13Cy0 3W12 DP.x/ y1y2y3 y0 1y0 2y0 3 y0 1y0 2y0 3 D0: We therefore have WD0, which is just the condition for linear dependence of the solutions yj. Thus, we have proved the following: A linear second-order homogeneous ODE has at most two linearly independent solu- tions. Generalizing, a linear homogeneous nth-order ODE has at most nlinearly inde- pendent solutions yj, and its general solution will be of the form y.x/DPn jD1cjyj.x/. ArfKen_Ch07-9780123846549.tex 362 Chapter 7 Ordinary Differential Equations Finding a Second Solution Returning to our linear, second-order, homogeneous ODE of the general form y00CP.x/y0CQ.x/yD0; (7.58) lety1andy2be two independent solutions. Then the Wronskian, by definition, is WDy1y0 2y0 1y2: (7.59) By differentiating the Wronskian, we obtain, as already demonstrated in Eq. (7.57), W0DP.x/W: (7.60) In the special case that P.x/D0, that is, y00CQ.x/yD0; (7.61) the Wronskian WDy1y0 2y0 1y2Dconstant: (7.62) Since our original differential equation is homogeneous, we may multiply the solutions y1 andy2by whatever constants we wish and arrange to have the Wronskian equal to unity (or1). This case, P.x/D0, appears more frequently than might be expected. Recall that r2. =r/in spherical polar coordinates contains no first radial derivative. Finally, every linear second-order differential equation can be transformed into an equation of the form ofEq. (7.61) (compare Exercise 7.6.12). For the general case, let us now assume that we have one solution of Eq. (7.58) by a series substitution (or by guessing). We now proceed to develop a second, independent solution for which W6D0. Rewriting Eq. (7.60) as dW WDPdx; we integrate over the variable x, from atox;to obtain lnW.x/ W.a/DxZ aP.x1/dx1; or7 W.x/DW.a/exp2 4xZ aP.x1/dx13 5: (7.63) 7IfP.x/remains finite in the domain of interest, W.x/6D0unless W.a/D0. That is, the Wronskian of our two solutions is either identically zero or never zero. However, if P.x/does not remain finite in our interval, then W.x/can have isolated zeros in that domain and one must be careful to choose aso that W.a/6D0: ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 363 Now we make the observation that W.x/Dy1y0 2y0 1y2Dy2 1d dxy2 y1 ; (7.64) and, by combining Eqs. (7.63) and (7.64), we have d dxy2 y1 DW.a/expTRx aP.x1/dx1U y2 1: (7.65) Finally, by integrating Eq. (7.65) from x2Dbtox2Dxwe get y2.x/Dy1.x/W.a/xZ bexp Rx2 aP.x1/dx1 Ty1.x2/U2dx2: (7.66) Here aandbare arbitrary constants and a term y1.x/y2.b/=y1.b/has been dropped, because it is a multiple of the previously found first solution y1. Since W.a/, the Wronskian evaluated at xDa, is a constant and our solutions for the homogeneous differential equa- tion always contain an arbitrary normalizing factor, we set W.a/D1and write y2.x/Dy1.x/xZexpTRx2P.x1/dx1U Ty1.x2/U2dx2: (7.67) Note that the lower limits x1Daandx2Dbhave been omitted. If they are retained, they simply make a contribution equal to a constant times the known first solution, y1.x/, and hence add nothing new. If we have the important special case P.x/D0, Eq. (7.67) reduces to y2.x/Dy1.x/xZdx2 Ty1.x2/U2: (7.68) This means that by using either Eq. (7.67) orEq. (7.68) we can take one known solution and by integrating can generate a second, independent solution of Eq. (7.58). This technique is used in Section 15.6 to generate a second solution of Legendre’s differential equation. Example 7.6.3 A SECOND SOLUTION FOR THE LINEAR OSCILLATOR EQUATION From d2y=dx2CyD0with P.x/D0let one solution be y1Dsinx. By applying Eq. (7.68), we obtain y2.x/DsinxxZdx2 sin2x2Dsinx.cot x/Dcos x; which is clearly independent (not a linear multiple) of sinx.  ArfKen_Ch07-9780123846549.tex 364 Chapter 7 Ordinary Differential Equations Series Form of the Second Solution Further insight into the nature of the second solution of our differential equation may be obtained by the following sequence of operations. 1. Express P.x/andQ.x/in Eq. (7.58) as P.x/D1X iD1pixi;Q.x/D1X jD2qjxj: (7.69) The leading terms of the summations are selected to create the strongest possible regular singularity (at the origin). These conditions just satisfy Fuchs’ theorem and thus help us gain a better understanding of that theorem. 2. Develop the first few terms of a power-series solution, as in Section 7.5. 3. Using this solution as y1, obtain a second series-type solution, y2, from Eq. (7.67), by integrating it term by term. Proceeding with Step 1, we have y00C.p1x1Cp0Cp1xC/y0C.q2x2Cq1x1C/yD0; (7.70) where xD0is at worst a regular singular point. If p1Dq1Dq2D0, it reduces to an ordinary point. Substituting yD1X D0axsC (Step 2), we obtain 1X D0.sC/.sC1/axsC2C1X iD1pixi1X D0.sC/axsC1 C1X jD2qjxj1X D0axsCD0: (7.71) Assuming that p16D0, our indicial equation is s.s1/Cp1kCq2D0; which sets the net coefficient of xs2equal to zero. This reduces to s2C.p11/sCq2D0: (7.72) We denote the two roots of this indicial equation by sD andsD n, where nis zero or a positive integer. (If nis not an integer, we expect two independent series solutions by the methods of Section 7.5 and we are done.) Then .s /.s Cn/D0; (7.73) or s2C.n2 /sC . n/D0; ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 365 and equating coefficients of sin Eqs. (7.72) and (7.73), we have p11Dn2 : (7.74) The known series solution corresponding to the larger root sD may be written as y1Dx 1X D0ax: Substituting this series solution into Eq. (7.67) (Step 3), we are faced with y2.x/Dy1.x/xZ exp Rx2 aP1 iD1pixi 1dx1 x2 2P1 D0ax 22! dx2; (7.75) where the solutions y1andy2have been normalized so that the Wronskian W.a/D1. Tackling the exponential factor first, we have x2Z a1X iD1pixi 1dx1Dp1lnx2C1X kD0pk kC1xkC1 2Cf.a/; (7.76) with f.a/an integration constant that may depend on a:Hence, exp0 @x2Z aX ipixi 1dx11 ADexpT f.a/Uxp1 2exp 1X kD0pk kC1xkC1 2! DexpT f.a/Uxp1 22 411X kD0pk kC1xkC1 2C1 2W 1X kD0pk kC1xkC1 2!2 C3 5:(7.77) This final series expansion of the exponential is certainly convergent if the original expan- sion of the coefficient P.x/was uniformly convergent. The denominator in Eq. (7.75) may be handled by writing 2 4x2 2 1X D0ax 2!23 51 Dx2 2 1X D0ax 2!2 Dx2 21X D0bx 2: (7.78) Neglecting constant factors, which will be picked up anyway by the requirement that W.a/D1, we obtain y2.x/Dy1.x/xZ xp12 2 1X D0cx 2! dx2: (7.79) Applying Eq. (7.74), xp12 2Dxn1 2; (7.80) ArfKen_Ch07-9780123846549.tex 366 Chapter 7 Ordinary Differential Equations and we have assumed here that nis an integer. Substituting this result into Eq. (7.79), we obtain y2.x/Dy1.x/xZ c0xn1 2Cc1xn 2Cc2xnC1 2CC cnx1 2C dx2: (7.81) The integration indicated in Eq. (7.81) leads to a coefficient of y1.x/consisting of two parts: 1. A power series starting with xn. 2. A logarithm term from the integration of x1(whenDn). This term always appears when nis an integer, unless cnfortuitously happens to vanish.8 If we choose to combine y1and the power series starting with xn, our second solution will assume the form y2.x/Dy1.x/lnjxjC1X jDndjxjC : (7.82) Example 7.6.4 A SECOND SOLUTION OF BESSEL’S EQUATION From Bessel’s equation, Eq. (7.40), divided by x2to agree with Eq. (7.59), we have P.x/Dx1Q.x/D1for the case nD0: Hence p1D1,q0D1; all other piandqjvanish. The Bessel indicial equation, Eq. (7.43) with nD0, is s2D0: Hence we verify Eqs. (7.72) to(7.74) with nand set to zero. Our first solution is available from Eq. (7.49). It is9 y1.x/DJ0.x/D1x2 4Cx4 64O.x6/: (7.83) Now, substituting all this into Eq. (7.67), we have the specific case corresponding to Eq. (7.75): y2.x/DJ0.x/xZ0 BBB@exph Rx2x1 1dx1i  1x2 2 4Cx4 2 6421 CCCAdx2: (7.84) 8For parity considerations, ln xis taken to be lnjxj, even. 9The capital O(order of) as written here means terms proportional to x6and possibly higher powers of x. ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 367 From the numerator of the integrand, exp2 4x2Zdx1 x13 5DexpT lnx2UD1 x2: This corresponds to the xp1 2in Eq. (7.77). From the denominator of the integrand, using a binomial expansion, we obtain " 1x2 2 4Cx4 2 64#2 D1Cx2 2 2C5x4 2 32C: Corresponding to Eq. (7.79), we have y2.x/DJ0.x/xZ1 x2" 1Cx2 2 2C5x4 2 32C# dx2 DJ0.x/ lnxCx2 4C5x4 128C : (7.85) Let us check this result. From Eq. (14.62), which gives the standard form of the second solution, which is called a Neumann function and designated Y0, Y0.x/D2  lnxln 2C  J0.x/C2 x2 43x4 128C : (7.86) Two points arise: (1) Since Bessel’s equation is homogeneous, we may multiply y2.x/by any constant. To match Y0.x/, we multiply our y2.x/by2=. (2) To our second solution, .2=/ y2.x/, we may add any constant multiple of the first solution. Again, to match Y0.x/ we add 2  ln 2C  J0.x/; where is the Euler-Mascheroni constant, defined in Eq. (1.13).10Our new, modified second solution is y2.x/D2  lnxln 2C  J0.x/C2 J0.x/x2 4C5x4 128C : (7.87) Now the comparison with Y0.x/requires only a simple multiplication of the series for J0.x/from Eq. (7.83) and the curly bracket of Eq. (7.87). The multiplication checks, through terms of order x2andx4, which is all we carried. Our second solution from Eqs. (7.67) and(7.75) agrees with the standard second solution, the Neumann function Y0.x/.  The analysis that indicated the second solution of Eq. (7.58) to have the form given in Eq. (7.82) suggests the possibility of just substituting Eq. (7.82) into the original differen- tial equation and determining the coefficients dj. However, the process has some features different from that of Section 7.5, and is illustrated by the following example. 10The Neumann function Y0is defined as it is in order to achieve convenient asymptotic properties; see Sections 14.3 and 14.6. ArfKen_Ch07-9780123846549.tex 368 Chapter 7 Ordinary Differential Equations Example 7.6.5 MORE NEUMANN FUNCTIONS We consider here second solutions to Bessel’s ODE of integer orders n>0, using the expansion given in Eq. (7.82). The first solution, designated Jnand presented in Eq. (7.49), arises from the value Dnfrom the indicial equation, while the quantity called nin Eq. (7.82), the separation of the two roots of the indicial equation, has in the current context the value 2n. Thus, Eq. (7.82) takes the form y2.x/DJn.x/lnjxjC1X jD2ndjxjCn; (7.88) where y2must, apart from scale and a possible multiple of Jn, be the second solution Ynof the Bessel equation. Substituting this form into Bessel’s equation, carrying out the indicated differentiations and using the fact that Jn.x/is a solution of our ODE, we get after combining similar terms x2y00 2Cxy0 2C.x2n2/y2D 2x J0 n.x/CX j2nj.jC2n/djxjCnCX j2ndjxjCnC2D0: (7.89) We next insert the power-series expansion 2x J0 n.x/DX j0ajxjCn; (7.90) where the coefficients can be obtained by differentiation of the expansion of Jn, see Eq. (7.49), and have the values (for j0) a2jD.1/j.nC2j/ jW.nCj/W2nC2j1; a2jC1D0: (7.91) This, and a redefinition of the index jin the last term, bring Eq. (7.89) to the form X j0ajxjCnCX j2nj.jC2n/djxjCnCX j2nC2dj2xjCnD0: (7.92) Considering first the coefficient of xnC1(corresponding to jD2nC1), we note that its vanishing requires that d2nC1vanish, as the only contribution comes from the middle summation. Since all ajof odd jvanish, the vanishing of d2nC1implies that all other dj of odd jmust also vanish. We therefore only need to give further consideration to even j. We next note that the coefficient d0is arbitrary, and may without loss of generality be set to zero. This is true because we may bring d0to any value by adding to y2an appropriate multiple of the solution Jn, whose expansion has an xnleading term. We have then exhausted all freedom in specifying y2; its scale is determined by our choice of its logarithmic term. Now, taking the coefficient of xn(terms with jD0), and remembering that d0D0, we have d2Da0; ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 369 and we may recur downward in steps of 2, using formulas based on the coefficients of xn2,xn4, . . . , corresponding to dj2Dj.2nCj/dj;jD2;4;:::;2nC2: To obtain djwith positive j, we recur upward, obtaining from the coefficient of xnCj djDajdj2 j.2nCj/;jD2;4;:::; again remembering that d0D0. Proceeding to nD1as a specific example, we have from Eq. (7.91) a0D1,a2D3=8 , anda4D5=192 , so d2D1; d2Da2 8D3 64;d4Da4d2 24D7 2304I thus y2.x/DJ1.x/lnjxj1 xC3 64x37 2304x5C; in agreement (except for a multiple of J1and a scale factor) with the standard form of the Neumann function Y1: Y1.x/D2  ln x 2 C 1 2 J1.x/C2  1 xC3 64x37 2304x5C : (7.93)  As shown in the examples, the second solution will usually diverge at the origin because of the logarithmic factor and the negative powers of xin the series. For this reason y2.x/is often referred to as the irregular solution. The first series solution, y1.x/, which usually converges at the origin, is called the regular solution. The question of behavior at the origin is discussed in more detail in Chapters 14 and 15, in which we take up Bessel functions, modified Bessel functions, and Legendre functions. Summary The two solutions of both sections (together with the exercises) provide a complete solu- tion of our linear, homogeneous, second-order ODE, assuming that the point of expansion is no worse than a regular singularity. At least one solution can always be obtained by series substitution (Section 7.5). A second, linearly independent solution can be con- structed by the Wronskian double integral, Eq. (7.67). This is all there are: No third, linearly independent solution exists (compare Exercise 7.6.10). Theinhomogeneous, linear, second-order ODE will have a general solution formed by adding a particular solution to the complete inhomogeneous equation to the general solu- tion of the corresponding homogeneous ODE. Techniques for finding particular solutions of linear but inhomogeneous ODEs are the topic of the next section. ArfKen_Ch07-9780123846549.tex 370 Chapter 7 Ordinary Differential Equations Exercises 7.6.1 You know that the three unit vectors Oex,Oey, andOezare mutually perpendicular (orthogonal). Show that Oex,Oey, andOezare linearly independent. Specifically, show that no relation of the form of Eq. (7.54) exists for Oex,Oey, andOez. 7.6.2 The criterion for the linear independence of three vectors A, B, and Cis that the equation aACbBCcCD0; analogous to Eq. (7.54), has no solution other than the trivial aDbDcD0. Using components AD.A1;A2;A3/, and so on, set up the determinant criterion for the exis- tence or nonexistence of a nontrivial solution for the coefficients a;b, and c. Show that your criterion is equivalent to the scalar triple product ABC6D0. 7.6.3 Using the Wronskian determinant, show that the set of functions  1;xn nW.nD1;2;:::; N/ is linearly independent. 7.6.4 If the Wronskian of two functions y1andy2is identically zero, show by direct integra- tion that y1Dcy2; that is, that y1andy2are linearly dependent. Assume the functions have continuous derivatives and that at least one of the functions does not vanish in the interval under consideration. 7.6.5 The Wronskian of two functions is found to be zero at x0"xx0C"for arbitrarily small">0:Show that this Wronskian vanishes for all xand that the functions are linearly dependent. 7.6.6 The three functions sinx;ex, and exare linearly independent. No one function can be written as a linear combination of the other two. Show that the Wronskian of sinx;ex, andexvanishes but only at isolated points. ANS. WD4 sin x; WD0forxDn,nD0;1;2;::: . 7.6.7 Consider two functions '1Dxand'2Djxj. Since'0 1D1and'0 2Dx=jxj,W.'1;'2/D 0for any interval, including T1;C1U. Does the vanishing of the Wronskian over T1;C1U prove that'1and'2are linearly dependent? Clearly, they are not. What is wrong? 7.6.8 Explain that linear independence does not mean the absence of any dependence. Illus- trate your argument with cosh xandex. 7.6.9 Legendre’s differential equation .1x2/y002xy0Cn.nC1/yD0 ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 371 has a regular solution Pn.x/and an irregular solution Qn.x/. Show that the Wronskian ofPnandQnis given by Pn.x/Q0 n.x/P0 n.x/Qn.x/DAn 1x2; with Anindependent ofx. 7.6.10 Show, by means of the Wronskian, that a linear, second-order, homogeneous ODE of the form y00.x/CP.x/y0.x/CQ.x/y.x/D0 cannot have three independent solutions. Hint. Assume a third solution and show that the Wronskian vanishes for all x. 7.6.11 Show the following when the linear second-order differential equation py00Cqy0CryD 0is expressed in self-adjoint form: (a) The Wronskian is equal to a constant divided by p: W.x/DC p.x/: (b) A second solution y2.x/is obtained from a first solution y1.x/as y2.x/DCy1.x/xZdt p.t/Ty1.t/U2: 7.6.12 Transform our linear, second-order ODE y00CP.x/y0CQ.x/yD0 by the substitution yDzexp2 41 2xZ P.t/dt3 5 and show that the resulting differential equation for zis z00Cq.x/zD0; where q.x/DQ.x/1 2P0.x/1 4P2.x/: Note. This substitution can be derived by the technique of Exercise 7.6.25. 7.6.13 Use the result of Exercise 7.6.12 to show that the replacement of '.r/byr'.r/may be expected to eliminate the first derivative from the Laplacian in spherical polar coordi- nates. See also Exercise 3.10.34. ArfKen_Ch07-9780123846549.tex 372 Chapter 7 Ordinary Differential Equations 7.6.14 By direct differentiation and substitution show that y2.x/Dy1.x/xZexpTRsP.t/dtU Ty1.s/U2ds satisfies, like y1.x/, the ODE y00 2.x/CP.x/y0 2.x/CQ.x/y2.x/D0: Note. The Leibniz formula for the derivative of an integral is d d h. /Z g. /f.x; /dxDh. /Z g. /@f.x; / @ dxCfTh. /; Udh. / d fTg. /; Udg. / d : 7.6.15 In the equation y2.x/Dy1.x/xZexpTRsP.t/dtU Ty1.s/U2ds; y1.x/satisfies y00 1CP.x/y0 1CQ.x/y1D0: The function y2.x/is a linearly independent second solution of the same equation. Show that the inclusion of lower limits on the two integrals leads to nothing new, that is, that it generates only an overall constant factor and a constant multiple of the known solution y1.x/. 7.6.16 Given that one solution of R00C1 rR0m2 r2RD0 isRDrm, show that Eq. (7.67) predicts a second solution, RDrm. 7.6.17 Using y1.x/D1X nD0.1/n .2nC1/Wx2nC1 as a solution of the linear oscillator equation, follow the analysis that proceeds through Eq. (7.81) and show that in that equation cnD0, so that in this case the second solution does not contain a logarithmic term. 7.6.18 Show that when nisnotan integer in Bessel’s ODE, Eq. (7.40), the second solution of Bessel’s equation, obtained from Eq. (7.67), does notcontain a logarithmic term. ArfKen_Ch07-9780123846549.tex 7.6 Other Solutions 373 7.6.19 (a) One solution of Hermite’s differential equation y002xy0C2 yD0 for D0isy1.x/D1. Find a second solution, y2.x/, using Eq. (7.67). Show that your second solution is equivalent to yodd(Exercise 8.3.3). (b) Find a second solution for D1, where y1.x/Dx, using Eq. (7.67). Show that your second solution is equivalent to yeven(Exercise 8.3.3). 7.6.20 One solution of Laguerre’s differential equation xy00C.1x/y0CnyD0 fornD0isy1.x/D1. Using Eq. (7.67), develop a second, linearly independent solu- tion. Exhibit the logarithmic term explicitly. 7.6.21 For Laguerre’s equation with nD0; y2.x/DxZes sds: (a) Write y2.x/as a logarithm plus a power series. (b) Verify that the integral form of y2.x/, previously given, is a solution of Laguerre’s equation.nD0/by direct differentiation of the integral and substitution into the differential equation. (c) Verify that the series form of y2.x/, part (a), is a solution by differentiating the series and substituting back into Laguerre’s equation. 7.6.22 One solution of the Chebyshev equation .1x2/y00xy0Cn2yD0 fornD0isy1D1. (a) Using Eq. (7.67), develop a second, linearly independent solution. (b) Find a second solution by direct integration of the Chebyshev equation. Hint. LetvDy0and integrate. Compare your result with the second solution given in Section 18.4. ANS. (a) y2Dsin1x. (b) The second solution, Vn.x/, is not defined for nD0. 7.6.23 One solution of the Chebyshev equation .1x2/y00xy0Cn2yD0 fornD1isy1.x/Dx. Set up the Wronskian double integral solution and derive a second solution, y2.x/. ANS. y2D.1x2/1=2. ArfKen_Ch07-9780123846549.tex 374 Chapter 7 Ordinary Differential Equations 7.6.24 The radial Schrödinger wave equation for a spherically symmetric potential can be writ- ten in the form " Nh2 2md2 dr2Cl.lC1/Nh2 2mr2CV.r/# y.r/DEy.r/: The potential energy V.r/may be expanded about the origin as V.r/Db1 rCb0Cb1rC: (a) Show that there is one (regular) solution y1.r/starting with rlC1. (b) From Eq. (7.69) show that the irregular solution y2.r/diverges at the origin as rl. 7.6.25 Show that if a second solution, y2, is assumed to be related to the first solution, y1, according to y2.x/Dy1.x/f.x/, substitution back into the original equation y00 2CP.x/y0 2CQ.x/y2D0 leads to f.x/DxZexpTRsP.t/dtU Ty1.s/U2ds; in agreement with Eq. (7.67). 7.6.26 (a) Show that y00C1 2 4x2yD0 has two solutions: y1.x/Da0x.1C /=2; y2.x/Da0x.1 /=2: (b) For D0the two linearly independent solutions of part (a) reduce to the single solution y10Da0x1=2. Using Eq. (7.68) derive a second solution, y20.x/Da0x1=2lnx: Verify that y20is indeed a solution. (c) Show that the second solution from part (b) may be obtained as a limiting case from the two solutions of part (a): y20.x/Dlim !0y1y2  : ArfKen_Ch07-9780123846549.tex 7.7 Inhomogeneous Linear ODEs 375 7.7 I NHOMOGENEOUS LINEAR ODE S We frame the discussion in terms of second-order ODEs, although the methods can be extended to equations of higher order. We thus consider ODEs of the general form y00CP.x/y0CQ.x/yDF.x/; (7.94) and proceed under the assumption that the corresponding homogeneous equation, with F.x/D0, has been solved, thereby obtaining two independent solutions designated y1.x/ andy2.x/. Variation of Parameters The method of variation of parameters (variation of the constant) starts by writing a par- ticular solution of the inhomogeneous ODE, Eq. (7.94), in the form y.x/Du1.x/y1.x/Cu2.x/y2.x/: (7.95) We have specifically written u1.x/andu2.x/to emphasize that these are functions of the independent variable, and notconstant coefficients. This, of course, means that Eq. (7.95) does not constitute a restriction to the functional form of y.x/. For clarity and compactness, we will usually write these functions just as u1andu2. In preparation for inserting y.x/, from Eq. (7.95), into the inhomogeneous ODE, we compute its derivative: y0Du1y0 1Cu2y0 2C.y1u0 1Cy2u0 2/; and take advantage of the redundancy in the form assumed for yby choosing u1andu2in such a way that y1u0 1Cy2u0 2D0; (7.96) where Eq. (7.96) is assumed to be an identity (i.e., to apply for all x). We will shortly show that requiring Eq. (7.96) does not lead to an inconsistency. After applying Eq. (7.96), y0, and its derivative y00, are found to be y0Du1y0 1Cu2y0 2; y00Du1y00 1Cu2y00 2Cu0 1y0 1Cu0 2y0 2; and substitution into Eq. (7.94) yields .u1y00 1Cu2y00 2Cu0 1y0 1Cu0 2y0 2/CP.x/.u1y0 1Cu2y0 2/CQ.x/.u1y1Cu2y2/DF.x/; which, because y1andy2are solutions of the homogeneous equation, reduces to u0 1y0 1Cu0 2y0 2DF.x/: (7.97) Equations (7.96) and(7.97) are, for each value of x, a set of two simultaneous algebraic equations in the variables u0 1andu0 2; to emphasize this point we repeat them here: y1u0 1Cy2u0 2D0; y0 1u0 1Cy0 2u0 2DF.x/:(7.98) ArfKen_Ch07-9780123846549.tex 376 Chapter 7 Ordinary Differential Equations The determinant of the coefficients of these equations is y1y2 y0 1y0 2 ; which we recognize as the Wronskian of the linearly independent solutions to the homo- geneous equation. That means this determinant is nonzero, so there will, for each x, be a unique solution to Eqs. (7.98), i.e., unique functions u0 1andu0 2. We conclude that the restriction implied by Eq. (7.96) is permissible. Once u0 1andu0 2have been identified, each can be integrated, respectively yielding u1 andu2, and, via Eq. (7.95), a particular solution of our inhomogeneous ODE. Example 7.7.1 ANINHOMOGENEOUS ODE Consider the ODE .1x/y00Cxy0yD.1x/2: (7.99) The corresponding homogeneous ODE has solutions y1Dxandy2Dex. Thus, y0 1D1, y0 2Dex, and the simultaneous equations for u0 1andu0 2are x u0 1Cexu0 2D0; u0 1Cexu0 2DF.x/:(7.100) Here F.x/is the inhomogeneous term when the ODE has been written in the standard form, Eq. (7.94). This means that we must divide Eq. (7.99) through by 1x(the coeffi- cient of y00), after which we see that F.x/D1x. With the above choice of F.x/, we solve Eqs. (7.100), obtaining u0 1D1; u0 2Dxex; which integrate to u1Dx;u2D.xC1/ex: Now forming a particular solution to the inhomogeneous ODE, we have yp.x/Du1y1Cu2y2Dx.x/C .xC1/ex exDx2CxC1: Because xis a solution to the homogeneous equation, we may remove it from the above expression, leaving the more compact formula ypDx2C1. The general solution to our ODE therefore takes the final form y.x/DC1xCC2exCx2C1:  ArfKen_Ch07-9780123846549.tex 7.8 Nonlinear Differential Equations 377 Exercises 7.7.1 If our linear, second-order ODE is inhomogeneous, that is, of the form of Eq. (7.94), themost general solution is y.x/Dy1.x/Cy2.x/Cyp.x/; where y1andy2are independent solutions of the homogeneous equation. Show that yp.x/Dy2.x/xZy1.s/F.s/ds Wfy1.s/;y2.s/gy1.x/xZy2.s/F.s/ds Wfy1.s/;y2.s/g; with Wfy1.x/;y2.x/gthe Wronskian of y1.s/andy2.s/. Find the general solutions to the following inhomogeneous ODEs: 7.7.2 y00CyD1: 7.7.3 y00C4yDex. 7.7.4 y003y0C2yDsinx: 7.7.5 xy00.1Cx/y0CyDx2: 7.8 N ONLINEAR DIFFERENTIAL EQUATIONS The main outlines of large parts of physical theory have been developed using mathe- matics in which the objects of concern possessed some sort of linearity property. As a result, linear algebra (matrix theory) and solution methods for linear differential equations were appropriate mathematical tools, and the development of these mathematical topics has progressed in the directions illustrated by most of this book. However, there is some physics that requires the use of nonlinear differential equations (NDEs). The hydrodynam- ics of viscous, compressible media is described by the Navier-Stokes equations, which are nonlinear. The nonlinearity evidences itself in phenomena such as turbulent flow, which cannot be described using linear equations. Nonlinear equations are also at the heart of the description of behavior known as chaotic, in which the evolution of a system is so sensitive to its initial conditions that it effectively becomes unpredictable. The mathematics of nonlinear ODEs is both more difficult and less developed than that of linear ODEs, and accordingly we provide here only an extremely brief survey. Much of the recent progress in this area has been in the development of computational methods for nonlinear problems; that is also outside the scope of this text. In this final section of the present chapter we discuss briefly some specific NDEs, the classical Bernoulli and Riccati equations. ArfKen_Ch07-9780123846549.tex 378 Chapter 7 Ordinary Differential Equations Bernoulli and Riccati Equations Bernoulli equations are nonlinear, having the form y0.x/Dp.x/y.x/Cq.x/Ty.x/Un; (7.101) where pandqare real functions and n6D0, 1 to exclude first-order linear ODEs. However, if we substitute u.x/DTy.x/U1n; then Eq. (7.101) becomes a first-order linear ODE, u0D.1n/yny0D.1n/ p.x/u.x/Cq.x/ ; (7.102) which we can solve (using an integrating factor) as described in Section 7.2. Riccati equations are quadratic in y.x/: y0Dp.x/y2Cq.x/yCr.x/; (7.103) where we require p6D0to exclude linear ODEs and r6D0to exclude Bernoulli equations. There is no known general method for solving Riccati equations. However, when a special solution y0.x/ofEq. (7.103) is known by a guess or inspection, then one can write the general solution in the form yDy0Cu, with usatisfying the Bernoulli equation u0Dpu2C.2py0Cq/u; (7.104) because substitution of yDy0Cuinto Eq. (7.103) removes r.x/from the resulting equation. There are no general methods for obtaining exact solutions of most nonlinear ODEs. This fact makes it more important to develop methods for finding the qualitative behavior of solutions. In Section 7.5 of this chapter we mentioned that power-series solutions of ODEs exist except (possibly) at essential singularities of the ODE. The coefficients in the power-series expansions provide us with the asymptotic behavior of the solutions. By making expansions of solutions to NDEs and retaining only the linear terms, it will often be possible to understand the qualitative behavior of the solutions in the neighborhood of the expansion point. Fixed and Movable Singularities, Special Solutions A first step in analyzing the solutions of NDEs is to identify their singularity structures. Solutions of NDEs may have singular points that are independent of the initial or bound- ary conditions; these are called fixed singularities. But in addition they may have spon- taneous, or movable, singularities that vary with the initial or boundary conditions. This feature complicates the asymptotic analysis of NDEs. ArfKen_Ch07-9780123846549.tex 7.8 Nonlinear Differential Equations 379 Example 7.8.1 MOVEABLE SINGULARITY Compare the linear ODE y0Cy x1D0; (which has an obvious regular singularity at xD1), with the NDE y0Dy2. Both have the same solution with initial condition y.0/D1, namely y.x/D1=.1x/. But for y.0/D2, the linear ODE has solution yD1C1=.1x/, while the NDE now has solution y.x/D 2=.12x/. The singularity in the solution of the NDE has moved to xD1=2.  For a linear second-order ODE we have a complete description of its solutions and their asymptotic behavior when two linearly independent solutions are known. But for NDEs there may still be special solutions whose asymptotic behavior is not obtainable from two independent solutions. This is another characteristic property of NDEs, which we illustrate again by an example. Example 7.8.2 SPECIAL SOLUTION The NDE y00Dyy0=xhas two linearly independent solutions that define the two-parameter family of curves y.x/D2c1tan.c 1lnxCc2/1; (7.105) where the ciare integration constants. However, this NDE also has the special solution yD c3Dconstant, which cannot be obtained from Eq. (7.105) by any choice of the parameters c1,c2. The “general solution” in Eq. (7.105) can be obtained by making the substitution xDet, and then defining Y.t/y.et/so that x.dy=dx/DdY=dt, thereby obtaining the ODE Y00DY0.YC1/. This ODE can be integrated once to give Y0D1 2Y2CYCcwith cD 2.c2 1C1=4/ an integration constant. The equation for Y0is separable and can be integrated again to yield Eq. (7.105).  Exercises 7.8.1 Consider the Riccati equation y0Dy2y2. A particular solution to this equation is yD2. Find a more general solution. 7.8.2 A particular solution to y0Dy2=x3y=xC2xisyDx2. Find a more general solution. 7.8.3 Solve the Bernoulli equation y0CxyDxy3. 7.8.4 ODEs of the form yDxy0Cf.y0/are known as Clairaut equations. The first step in solving an equation of this type is to differentiate it, yielding y0Dy0Cxy00Cf0.y0/y00;ory00 xCf0.y0/ D0: Solutions may therefore be obtained both from y00D0and from f0.y0/Dx. The so-called general solution comes from y00D0. ArfKen_Ch07-9780123846549.tex 380 Chapter 7 Ordinary Differential Equations Forf.y0/D.y0/2, (a) Obtain the general solution (note that it contains a single constant). (b) Obtain the so-called singular solution from f0.y0/Dx. By substituting back into the original ODE show that this singular solution contains no adjustable constants. Note. The singular solution is the envelope of the general solutions. Additional Readings Cohen, H., Mathematics for Scientists and Engineers. Englewood Cliffs, NJ: Prentice-Hall (1992). Golomb, M., and M. Shanks, Elements of Ordinary Differential Equations. New York: McGraw-Hill (1965). Hubbard, J., and B. H. West, Differential Equations. Berlin: Springer (1995). Ince, E. L., Ordinary Differential Equations. New York: Dover (1956). The classic work in the theory of ordinary differential equations. Jackson, E. A., Perspectives of Nonlinear Dynamics. Cambridge: Cambridge University Press (1989). Jordan, D. W., and P. Smith, Nonlinear Ordinary Differential Equations, 2nd ed. Oxford: Oxford University Press (1987). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nostrand (1956). Miller, R. K., and A. N. Michel, Ordinary Differential Equations. New York: Academic Press (1982). Murphy, G. M., Ordinary Differential Equations and Their Solutions. Princeton, NJ: Van Nostrand (1960). A thorough, relatively readable treatment of ordinary differential equations, both linear and nonlinear. Ritger, P. D., and N. J. Rose, Differential Equations with Applications. New York: McGraw-Hill (1968). Sachdev, P. L., Nonlinear Differential Equations and their Applications. New York: Marcel Dekker (1991). Tenenbaum, M., and H. Pollard, Ordinary Differential Equations. New York: Dover (1985). Detailed and read- able (over 800 pages). This is a reprint of a work originally published in 1963, and stresses formal manipula- tions. Its references to numerical methods are somewhat dated. ArfKen_Ch08-9780123846549.tex CHAPTER 8 STURM-LIOUVILLE THEORY 8.1 I NTRODUCTION Chapter 7 examined methods for solving ordinary differential equations (ODEs), with emphasis on techniques that can generate the solutions. In the present chapter we shift the focus to the general properties that solutions must have to be appropriate for specific physics problems, and to discuss the solutions using the notions of vector spaces and eigen- value problems that were developed in Chapters 5 and 6. A typical physics problem controlled by an ODE has two important properties: (1) Its solution must satisfy boundary conditions, and (2) It contains a parameter whose value must be set in a way that satisfies the boundary conditions. From a vector-space perspec- tive, the boundary conditions (plus continuity and differentiability requirements) define the Hilbert space of our problem, while the parameter normally occurs in a way that permits the ODE to be written as an eigenvalue equation within that Hilbert space. These ideas can be made clearer by examining a specific example. The standing waves of a vibrating string clamped at its ends are governed by the ODE d2 dx2Ck2 D0; (8.1) where .x/is the amplitude of the transverse displacement at the point xalong the string, andkis a parameter. This ODE has solutions for any value of k, but the solutions of relevance to the string problem must have .x/D0for the values of xat the ends of the string. The boundary conditions of this problem can be interpreted as defining a Hilbert space whose members are differentiable functions with zeros at the boundary values of x; the ODE itself can be written as the eigenvalue equation L Dk2 ;LDd2 dx2: (8.2) 381 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch08-9780123846549.tex 382 Chapter 8 Sturm-Liouville Theory For practical reasons the eigenvalue is given the name k2. It is required to find functions .x/that solve Eq. (8.2) subject to the boundary conditions, i.e., to find members .x/of our Hilbert space that solve the eigenvalue equation. We could now follow the procedures developed in Chapter 5, namely (1) choose a basis for our Hilbert space (a set of functions with zeros at the boundary values of x), (2) define a scalar product for our space, (3) expand Land in terms of our basis, and (4) solve the resulting matrix equation. However, that procedure makes no use of any specific features of the current ODE, and in particular ignores the fact that it is easily solved. Instead, we continue with the example defined by Eq. (8.1), using our ability to solve the ODE involved. Example 8.1.1 STANDING WAVES, VIBRATING STRING We consider a string clamped at xD0andxDland undergoing transverse vibrations. As already indicated, its standing wave amplitudes .x/are solutions of the differential equation d2 .x/ dx2Ck2 .x/D0; (8.3) where kis not initially known and .x/is subject to the boundary conditions that the ends of the string be fixed in position: .0/D .l/D0. This is the eigenvalue problem defined in Eq. (8.2). The general solution to this differential equation is .x/DAsinkxCBcoskx, and in the absence of the boundary conditions solutions would exist for all values of k,A, and B. However, the boundary condition at xD0requires us to set BD0, leaving .x/D Asinkx. We have yet to satisfy the boundary condition at xDl. The fact that Ais as yet unspecified is not helpful for this purpose, as AD0leaves us with only the trivial solution D0. We must, instead, require sinklD0, which is accomplished by setting klDn, where nis a nonzero integer, leading to n.x/DAsinnx l ;k2Dn22 l2;nD1;2;:::: (8.4) Because Eq. (8.3) is homogeneous, it will have solutions of arbitrary scale, so Acan have any value. Since our purpose is usually to identify linearly independent solutions, we disre- gard changes in the sign or magnitude of A. In the vibrating string problem, these quantities control the amplitude and phase of the standing waves. Since changing the sign of nsim- ply changes the sign of ,Cnandnin Eq. (8.4) are regarded here as equivalent, so we restricted nto positive values. The first few nare shown in Fig. 8.1. Note that the number of nodes increases with n: nhasnC1nodes (including the two nodes at the ends of the string). The fact that our problem has solutions only for discrete values of kis typical of eigen- value problems, and in this problem the discreteness in kcan be traced directly to the presence of the boundary conditions. Figure 8.2 shows what happens when kis varied in either direction from the acceptable value =l, with the boundary condition at xD0 maintained for all k. It is obvious that the eigenvalues (here k2) lie at separated points, and ArfKen_Ch08-9780123846549.tex 8.1 Introduction 383 FIGURE 8.1 Standing wave patterns of a vibrating string. a b c d ex 0 FIGURE 8.2 Solutions to Eq. (8.3) on the range 0xlfor: (a) kD0:9= l, (b)kD=l, (c)kD1:2= l, (d)kD1:5= l, (e)kD1:9= l. that the boundary condition at xDlcannot be satisfied for k<= l. Moreover, the first acceptable kvalue larger than =lis clearly larger than 1:9= l(it is actually 2=l). As already noted, the solution to this eigenvalue problem is undetermined as to scale because the underlying equation (together with its boundary conditions) is homogeneous. However, if we introduce a scalar product of definition hfjgiDlZ 0f.x/g.x/dx; (8.5) we can define solutions that are normalized; requiring h nj niD1, we have, with arbitrary sign, n.x/Dr 2 lsinnx l : (8.6) Although we did not solve Eq. (8.2) by an expansion technique, the solutions (the eigen- functions) will still have properties that depend on whether the operator Lis Hermitian. As we saw in Chapter 5, the Hermitian property depends both on Land the definition of the scalar product, and a topic for discussion in the present chapter is the identification of con- ditions making an operator Hermitian. This issue is important because Hermiticity implies real eigenvalues as well as orthogonality and completeness of the eigenfunctions.  ArfKen_Ch08-9780123846549.tex 384 Chapter 8 Sturm-Liouville Theory Summarizing, the matters of interest here, and the subject matter of the current chapter, include: 1. The conditions under which an ODE can be written as an eigenvalue equation with a self-adjoint (Hermitian) operator, 2. Methods for the solution of ODEs subject to boundary conditions, and 3. The properties of the solutions to ODE eigenvalue equations. 8.2 H ERMITIAN OPERATORS Characterization of the general features of eigenproblems arising from second-order dif- ferential equations is known as Sturm-Liouville theory. It therefore deals with eigenvalue problems of the form L .x/D .x/; (8.7) where Lis a linear second-order differential operator, of the general form L.x/Dp0.x/d2 dx2Cp1.x/d dxCp2.x/: (8.8) The key matter at issue here is to identify the conditions under which Lis a Hermitian operator. Self-Adjoint ODEs Lis known in differential equation theory as self-adjoint if p0 0.x/Dp1.x/: (8.9) This feature enables L.x/to be written L.x/Dd dx p0.x/d dx Cp2.x/; (8.10) and the operation of Lon a function u.x/then takes the form LuD.p0u0/0Cp2u: (8.11) Inserting Eq. (8.11) into an integral of the formRb av.x/Lu.x/dx, we proceed by applying an integration by parts to the p0term (assuming that p0is real): bZ av.x/Lu.x/dxDbZ ah v p0u00Cvp2ui dx Dh vp0u0ib aCbZ a .v/0p0u0Cvp2u dx: ArfKen_Ch08-9780123846549.tex 8.2 Hermitian Operators 385 Another integration by parts leads to bZ av.x/Lu.x/dxDh vp0u0.v/0p0uib aCbZ ah p0.v/00uCvp2ui dx Dh vp0u0.v/0p0uib aCbZ a.Lv/u dx: (8.12) Equation (8.12) shows that, if the boundary terms b avanish and the scalar product is an unweighted integral from atob, then the operator Lis self-adjoint, as that term was defined for operators. In passing, we observe that the notion of self-adjointness in differential equation theory is weaker than the corresponding concept for operators in our Hilbert spaces, due to the lack of a requirement on the boundary terms. We again stress that the Hilbert-space definition of self-adjoint depends not only on the form of Lbut also on the definition of the scalar product and the boundary conditions. Looking further at the boundary terms, we see that they are surely zero if uandvboth vanish at the endpoints xDaandxDb(a case of what are termed Dirichlet boundary conditions). The boundary terms are also zero if both u0andv0vanish at aandb(Neu- mann boundary conditions). Even if neither Dirichlet nor Neumann boundary conditions apply, it may happen (particularly in a periodic system, such as a crystal lattice) that the boundary terms vanish because vp0u0 aDvp0u0 bfor all uandv. Specializing Eq. (8.12) to the case that uandvare eigenfunctions of Lwith respective real eigenvalues uandv, that equation reduces to .uv/bZ avu dxDh p0.vu0.v/0u/ib a: (8.13) It is thus apparent that if the boundary terms vanish and u6Dv, then uandvmust be orthogonal on the interval .a;b/. This is a specific illustration of the orthogonality requirement for eigenfunctions of a Hermitian operator in a Hilbert space. Making an ODE Self-Adjoint Some of the differential equations that are important in physics involve operators Lthat are self-adjoint in the differential-equation sense, meaning that they satisfy Eq. (8.9); others are not. However, if an operator does not satisfy Eq. (8.9), it is known how to multiply it by a quantity that converts it into self-adjoint form. Letting such a quantity be designated w.x/, the Sturm-Liouville eigenvalue problem of Eq. (8.7) becomes w.x/L.x/ .x/Dw.x/ . x/; (8.14) an equation that has the same eigenvalues and eigenfunctions .x/as the original prob- lem in Eq. (8.7). If now w.x/is chosen to be w.x/Dp1 0expZp1.x/ p0.x/dx ; (8.15) ArfKen_Ch08-9780123846549.tex 386 Chapter 8 Sturm-Liouville Theory where p0andp1are the quantities in Las given in Eq. (8.8), we can by direct evaluation find that w.x/L.x/Dp0d2 dx2Cp1d dxCw.x/p2.x/; (8.16) where p0DexpZp1.x/ p0.x/dx ;p1Dp1 p0expZp1.x/ p0.x/dx : (8.17) It is then straightforward to show that p0 0Dp1, sowLsatisfies the self-adjoint condition. If we now apply the process represented by Eq. (8.12) to wL, we get bZ av.x/w.x/Lu.x/dxDh vp0u0 v0p0uib aCbZ aw.x/.Lv/u dx: (8.18) If the boundary terms vanish, Eq. (8.18) is equivalent tohvjLjuiDhLvjuiwhen the scalar product is defined to be hvjuiDbZ av.x/u.x/w.x/dx: (8.19) Again considering the case that uandvare eigenfunctions of L, with respective eigen- valuesuandv,Eq. (8.18) reduces to .uv/bZ avuwdxDh wp0 vu0.v/0uib a; (8.20) where p0is the coefficient of y00in the original ODE. We thus see that if the right-hand side of Eq. (8.20) vanishes, then uandvare orthogonal on .a;b/with weight factor w whenu6Dv. In other words, our choice of scalar product definition and boundary con- ditions have made La self-adjoint operator in our Hilbert space, thereby producing an eigenfunction orthogonality condition. Summarizing, we have the useful and important result: If a second-order differential operator Lhas coefficients p0.x/and p1.x/that sat- isfy the self-adjoint condition, Eq. (8.9), then it is Hermitian, given (a) a scalar prod- uct of uniform weight and (b) boundary conditions that remove the endpoint terms of Eq. (8.12). IfEq. (8.9) is not satisfied, then Lis Hermitian if (a) the scalar product is defined to include the weight factor given in Eq. (8.15), and (b) boundary conditions cause removal of the endpoint terms in Eq. (8.18). Note that once the problem has been defined such that Lis Hermitian, then the general properties proved for Hermitian problems apply: the eigenvalues are real; the eigenfunc- tions are (or if degenerate can be made) orthogonal, using the relevant scalar product definition. ArfKen_Ch08-9780123846549.tex 8.2 Hermitian Operators 387 Example 8.2.1 LAGUERRE FUNCTIONS Consider the eigenvalue problem L D , with LDxd2 dx2C.1x/d dx; (8.21) subject to (a) nonsingular on 0x<1, and (b) limx!1 .x/D0. Condition (a) is simply a requirement that we use the solution of the differential equation that is regular at xD0; and condition (b) is a typical Dirichlet boundary condition. The operator Lis not self-adjoint, with p0Dxandp1D1x. But we can form w.x/D1 xexpZ1x xdx D1 xelnxxDex: (8.22) The boundary terms, for arbitrary eigenfunctions uandv, are of the form h xex vu0.v/0ui1 0I their contributions at xD1 vanish because uandvgo to zero; the common factor x causes the xD0contribution to vanish also. We therefore have a self-adjoint problem, with uandvof different eigenvalues orthogonal under the definition hvjuiD1Z 0v.x/u.x/exdx: The eigenvalue equation of this example is that whose solutions are the Laguerre polynomials; what we have shown here is that they are orthogonal on .0;1/ with weight ex.  Exercises 8.2.1 Show that Laguerre’s ODE, Table 7.1, may be put into self-adjoint form by multiplying byexand thatw.x/Dexis the weighting function. 8.2.2 Show that the Hermite ODE, Table 7.1, may be put into self-adjoint form by multiplying byex2and that this gives w.x/Dex2as the appropriate weighting function. 8.2.3 Show that the Chebyshev ODE, Table 7.1, may be put into self-adjoint form by mul- tiplying by.1x2/1=2and that this gives w.x/D.1x2/1=2as the appropriate weighting function. 8.2.4 The Legendre, Chebyshev, Hermite, and Laguerre equations, given in Table 7.1, have solutions that are polynomials. Show that ranges of integration that guarantee that the Hermitian operator boundary conditions will be satisfied are (a) LegendreT1; 1U, (b) Chebyshev T1; 1U, (c) Hermite .1;1/, (d) Laguerre T0;1/. ArfKen_Ch08-9780123846549.tex 388 Chapter 8 Sturm-Liouville Theory 8.2.5 The functions u1.x/andu2.x/are eigenfunctions of the same Hermitian operator but for distinct eigenvalues 1and2. Prove that u1.x/andu2.x/are linearly independent. 8.2.6 Given that P1.x/Dxand Q0.x/D1 2ln1Cx 1x are solutions of Legendre’s differential equation (Table 7.1) corresponding to different eigenvalues: (a) Evaluate their orthogonality integral 1Z 1x 2ln1Cx 1x dx: (b) Explain why these two functions are not orthogonal, that is, why the proof of orthogonality does not apply. 8.2.7 T0.x/D1andV1.x/D.1x2/1=2are solutions of the Chebyshev differential equation corresponding to different eigenvalues. Explain, in terms of the boundary conditions, why these two functions are not orthogonal on the range .1; 1/with the weighting function found in Exercise 8.2.3. 8.2.8 A set of functions un.x/satisfies the Sturm-Liouville equation d dx p.x/d dxun.x/ Cnw.x/un.x/D0: The functions um.x/andun.x/satisfy boundary conditions that lead to orthogonality. The corresponding eigenvalues mandnare distinct. Prove that for appropriate bound- ary conditions, u0 m.x/andu0 n.x/are orthogonal with p.x/as a weighting function. 8.2.9 Linear operator Ahasndistinct eigenvalues and ncorresponding eigenfunctions: A iDi i. Show that the neigenfunctions are linearly independent. Do not assume Ato be Hermitian. Hint. Assume linear dependence, i.e., that nDPn1 iD1ai i. Use this relation and the operator-eigenfunction equation first in one order and then in the reverse order. Show that a contradiction results. 8.2.10 The ultraspherical polynomials C. / n.x/are solutions of the differential equation  .1x2/d2 dx2.2 C1/xd dxCn.nC2 / C. / n.x/D0: (a) Transform this differential equation into self-adjoint form. (b) Find an interval of integration and weighting factor that make C. / n.x/of the same but different northogonal. Note. Assume that your solutions are polynomials. ArfKen_Ch08-9780123846549.tex 8.3 ODE Eigenvalue Problems 389 8.3 ODE E IGENVALUE PROBLEMS Now that we have identified the conditions that make a second-order ODE eigenvalue problem Hermitian, let’s examine several such problems to gain further understanding of the processes involved and to illustrate techniques for finding solutions. Example 8.3.1 LEGENDRE EQUATION The Legendre equation, Ly.x/D.1x2/y00.x/C2xy0.x/Dy.x/; (8.23) defines an eigenvalue problem that arises when r2is written in spherical polar coordinates, with xidentified as cos, whereis the polar angle of the coordinate system. The range ofxin this context is1x1, and in typical circumstances one needs solutions to Eq. (8.23) that are nonsingular on the entire range of x. It turns out that this is a nontrivial requirement, mainly because xD1 are singular points of the Legendre ODE. If we regard nonsingularity of yatxD1 as a set of boundary conditions, we shall find that this requirement is sufficient to define eigenfunctions of the Legendre operator. This eigenvalue problem, namely Eq. (8.23) plus nonsingularity at xD1 , is conve- niently handled by the method of Frobenius. We assume solutions of the form yD1X jD0ajxsCj; (8.24) with indicial equation s.s1/D0, whose solutions are sD0andsD1. For sD0, we obtain the following recurrence relation for the coefficients aj: ajC2Dj.jC1/ .jC1/.jC2/aj: (8.25) We may set a1D0, thereby causing all ajof odd jto vanish, so (for sD0) our series will contain only even powers of x. The boundary condition comes into play because Eq. (8.24) diverges at xD1 for allexcept those that actually cause the series to terminate after a finite number of terms. To see how the divergence arises, note that for large jandjxjD1the ratio of successive terms of the series approaches ajxj ajC2xjC2!j.jC1/ .jC1/.jC2/!1; so the ratio test is indeterminate. However, application of the Gauss test shows that this series diverges, as was discussed in more detail in Example 1.1.7. The series in Eq. (8.24) can be made to terminate after alfor some even lby choosing Dl.lC1/, a value that makes alC2D0. Then alC4,alC6;::: will also vanish, and our solution will be a polynomial, which is clearly nonsingular for all jxj1. Summarizing, we have, for even l, solutions that are polynomials of degree las eigenfunctions, and the corresponding eigenvalues are l.lC1/. ArfKen_Ch08-9780123846549.tex 390 Chapter 8 Sturm-Liouville Theory ForsD1we must set a1D0and the recurrence relation is ajC2D.jC1/.jC2/ .jC2/.jC3/aj; (8.26) which also leads to divergence at jxjD1. However, the divergence can now be avoided by settingD.lC1/.lC2/for some even value of l, thereby causing alC2,alC4;::: to vanish. The result will be a polynomial of degree lCs, i.e., of an odd degree lC1. These solutions can be described equivalently as, for odd l, polynomials of degree lwith eigenvaluesDl.lC1/, so the overall set of eigenfunctions consists of polynomials of all integer degrees l, with respective eigenvalues l.lC1/. When given the conventional scal- ing, these polynomials are called Legendre polynomials. Verification of these properties of solutions to the Legendre equation is left to Exercise 8.3.1. Before leaving the Legendre equation, note that its ODE is self-adjoint, and that the coefficient of d2=dx2in the Legendre operator is p0D.1x2/, which vanishes at xD 1. Comparing with Eq. (8.12), we see that this value of p0causes the vanishing of the boundary terms when we take the adjoint of L, so the Legendre operator on the range 1x1is Hermitian, and therefore has orthogonal eigenfunctions. In other words, the Legendre polynomials are orthogonal with unit weight on .1; 1/.  Let’s examine one more ODE that leads to an interesting eigenvalue problem. Example 8.3.2 HERMITE EQUATION Consider the Hermite differential equation, LyDy00C2xy0Dy; (8.27) which we wish to regard as an eigenvalue problem on the range 1<x<1. To make LHermitian, we define a scalar product with a weight factor as given by Eq. (8.15), hfjgiD1Z 1f.x/g.x/ex2dx; (8.28) and demand (as a boundary condition) that our eigenfunctions ynhave finite norms using this scalar product, meaning that hynjyni<1. Again we obtain a solution by the method of Frobenius, as a series of the form given inEq. (8.24). Again the indicial equation is s.s1/D0, and for sD0we can develop a series of even powers of xwith coefficients satisfying the recurrence relation ajC2D2j .jC1/.jC2/aj: (8.29) This series converges for all x, but (assuming it does not terminate) it behaves asymptoti- cally for largejxjasex2and therefore does not describe a function of finite norm, even with theex2weight factor in the scalar product. Thus, even though the series solution always converges, our boundary conditions require that we arrange to terminate the series, thereby producing polynomial solutions. From Eq. (8.29) we see that the condition for obtaining an even polynomial of degree jis thatD2j. Odd polynomial solutions can be obtained ArfKen_Ch08-9780123846549.tex 8.3 ODE Eigenvalue Problems 391 using the indicial equation solution sD1. Details of both the solutions and the asymptotic properties are the subject of Exercise 8.3.3. Since we have established that this is a Hermitian eigenvalue problem with the scalar product as defined in Eq. (8.28), its solutions (when scaled conventionally they are called Hermite polynomials) are orthogonal using that scalar product.  Some ODE eigenvalue problems can be attacked by dividing the space in which they reside into regions that are most naturally treated in different ways. The following example illustrates this situation, with a potential that is assumed nonzero only within a finite region. Example 8.3.3 DEUTERON GROUND STATE The deuteron is a bound state of a neutron and a proton. Due to the short range of the nuclear force, the deuteron properties do not depend much on the detailed shape of the interaction potential. Thus, this system may be modeled by a spherically symmetric square well potential with the value VDV0<0when the nucleons are within a distance aof each other, but with VD0when the internucleon distance is greater than a. The Schrödinger equation for the relative motion of the two nucleons assumes the form Nh2 2r2 CV DE ; whereis the reduced mass of the system (approximately half the mass of either particle). This eigenvalue equation must be solved subject to the boundary conditions that be finite atrD0and approach zero at rD1 sufficiently rapidly to be a member of an L2Hilbert space. The eigenfunctions must also be continuous and differentiable for all r, including rDa. It can be shown that if there is to be a bound state, Ewill have to have a negative value in the range V0<E<0, and the lowest state (the ground state) will be described by a wave function that is spherically symmetric (thereby having no angular momentum). Thus, taking D .r/and using a result from Exercise 3.10.34 to write r2 D1 rd2u dr2;with u.r/Dr .r/; the Schrödinger equation reduces to an ODE that assumes the form, for r<a, d2u1 dr2Ck2 1u1D0; with k2 1D2 Nh2.EV0/>0; while, for r>a, d2u2 dr2k2 2u2D0; with k2 2D2E Nh2>0: The solutions for these two ranges of rmust connect smoothly, meaning that both u anddu=dr must be continuous across rDa, and therefore must satisfy the matching conditions u1.a/Du2.a/,u0 1.a/Du0 2.a/. In addition, the requirement that be finite atrD0dictates that u1.0/D0, and the boundary condition at rD1 requires that limr!1u2.r/D0. ArfKen_Ch08-9780123846549.tex 392 Chapter 8 Sturm-Liouville Theory Forr<a, our Schrödinger equation has the general solution u1.r/DAsink1rCCcosk1r; and the boundary condition at rD0is only met if we set CD0. The Schrödinger equation forr>ahas the general solution u2.r/DC0exp.k 2r/CBexp.k 2r/; (8.30) and the boundary condition at rD1 requires us to set C0D0. The matching conditions atrDathen take the form Asink1aDBexp.k 2a/and Ak1cosk1aDk 2Bexp.k 2a/: Using the second of these equations to eliminate Bexp.k 2a/from the first, we reach Asink1aDAk1 k2cosk1a; (8.31) showing that the overall scale of the solution (i.e., A) is arbitrary, which is of course a consequence of the fact that the Schrödinger equation is homogeneous. Rearranging Eq. (8.31), and inserting values for k1andk2, our matching conditions become tank1aDk1 k2;ortan2a2 Nh2.EV0/1=2 Dr EV0 E: (8.32) This is an admittedly unpleasant implicit equation for E; if it has solutions with Ein the range V0<E<0, our model predicts deuteron bound state(s). One way to search for solutions to Eq. (8.32) is to plot its left- and right-hand sides as a function of E, identifying the Evalues, if any, for which they are equal. Taking V0D4:0461012J,aD2:5fermi,1D0:8351027kg, andNhD1:051034J-s (joule-seconds), the two sides of Eq. (8.32) are plotted in Fig. 8.3 for the range of Ein which a bound state is possible. The Evalues have been plotted in MeV (mega electron volts), the energy unit most frequently used in nuclear physics (1 MeV 1:61013J). The curves cross at only one point, indicating that the model predicts just one bound state. Its energy is at approximately ED2:2 MeV. It is instructive to see what happens if we take Evalues that may or may not solve Eq. (8.32), using u.r/DAsink1rforr<a(thereby satisfying the rD0boundary condi- tion) but for r>ausing the general form of u.r/as given in Eq. (8.30), with the coefficient values BandC0that are required by the matching conditions for the chosen Evalue. Let- tingEandEC, respectively, denote values of Eless than and greater than the eigenvalue E, we find that by forcing a smooth connection at rDawe lose the required asymptotic behavior except at the eigenvalue. See Fig. 8.4.  11 fermiD1015m. ArfKen_Ch08-9780123846549.tex 8.3 ODE Eigenvalue Problems 393 0LHS RHS LHS –20 –25 –10 –15 0 5 E (MeV) FIGURE 8.3 Left- and right-hand sides of Eq. (8.32) as a function of Efor the model parameters given in the text. –0.500.51 u(r) 4 81 2 1 6E– E E+ r (fermi ) FIGURE 8.4 Wavefunctions for the deuteron problem when the energy is chosen to be less than the eigenvalue E(E<E) or greater than E(EC>E). Exercises 8.3.1 Solve the Legendre equation .1x2/y002xy0Cn.nC1/yD0 by direct series substitution. (a) Verify that the indicial equation is s.s1/D0: (b) Using sD0and setting the coefficient a1D0, obtain a series of even powers of x: yevenDa0 1n.nC1/ 2Wx2C.n2/n.nC1/.nC3/ 4Wx4C ; ArfKen_Ch08-9780123846549.tex 394 Chapter 8 Sturm-Liouville Theory where ajC2Dj.jC1/n.nC1/ .jC1/.jC2/aj: (c) Using sD1and noting that the coefficient a1must be zero, develop a series of odd powers of x: yoddDa0 x.n1/.nC2/ 3Wx3 C.n3/.n1/.nC2/.nC4/ 5Wx5C ; where ajC2D.jC1/.jC2/n.nC1/ .jC2/.jC3/aj: (d) Show that both solutions, yevenandyodd, diverge for xD1 if the series continue to infinity. (Compare with Exercise 1.2.5.) (e) Finally, show that by an appropriate choice of n, one series at a time may be con- verted into a polynomial, thereby avoiding the divergence catastrophe. In quantum mechanics this restriction of nto integral values corresponds to quantization of angular momentum. 8.3.2 Show that with the weight factor exp. x2/and the interval1<x<1for the scalar product, the Hermite ODE eigenvalue problem is Hermitian. 8.3.3 (a) Develop series solutions for Hermite’s differential equation y002xy0C2 yD0: ANS. s.s1/D0, indicial equation. ForsD0; ajC2D2ajj .jC1/.jC2/.jeven/; yevenDa0 1C2. / x2 2WC22. /.2 /x4 4WC : ForsD1; ajC2D2ajjC1 .jC2/.jC3/.jeven/; yoddDa1 xC2.1 /x3 3WC22.1 /.3 /x5 5WC : (b) Show that both series solutions are convergent for all x, the ratio of successive coefficients behaving, for a large index, like the corresponding ratio in the expan- sion of exp.x2/. ArfKen_Ch08-9780123846549.tex 8.4 Variation Method 395 (c) Show that by appropriate choice of , the series solutions may be cut off and converted to finite polynomials. (These polynomials, properly normalized, become the Hermite polynomials in Section 18.1.) 8.3.4 Laguerre’s ODE is x L00 n.x/C.1x/L0 n.x/CnLn.x/D0: Develop a series solution and select the parameter nto make your series a polynomial. 8.3.5 Solve the Chebyshev equation .1x2/T00 nxT0 nCn2TnD0; by series substitution. What restrictions are imposed on nif you demand that the series solution converge for xD1 ? ANS. The infinite series does converge for xD1 and no restriction on nexists (compare with Exercise 1.2.6). 8.3.6 Solve .1x2/U00 n.x/3xU0 n.x/Cn.nC2/Un.x/D0; choosing the root of the indicial equation to obtain a series of odd powers of x. Since the series will diverge for xD1, choose nto convert it into a polynomial. 8.4 V ARIATION METHOD We saw in Chapter 6 that the expectation value of a Hermitian operator Hfor the normal- ized function can be written as hHih jHj i; and that the expansion of this quantity in a basis consisting of the orthonormal eigenfunc- tions of Hhad the form given in Eq. (6.30): hHiDX jaj2; where ais the coefficient of the th eigenfunction of Handiis the corresponding eigenvalue. As we noted when we obtained this result, one of its consequences is that hHi is a weighted average of the eigenvalues of H, and therefore is at least as large as the small- est eigenvalue, and equal to the smallest eigenvalue only if is actually an eigenfunction to which that eigenvalue corresponds. The observations of the foregoing paragraph hold true even if we do not actually make an expansion of and even if we do not actually know or have available the eigenfunctions or eigenvalues of H. The knowledge that hHiis an upper limit to the smallest eigenvalue ofHis sufficient to enable us to devise a method for approximating that eigenvalue and the associated eigenfunction. This eigenfunction will be the member of the Hilbert space of our problem that yields the smallest expectation value of H, and a strategy for finding it is to search for the minimum in hHiwithin our Hilbert space. This is the essential idea ArfKen_Ch08-9780123846549.tex 396 Chapter 8 Sturm-Liouville Theory behind what is known as the variation method for the approximate solution of eigenvalue problems. Since in many problems (including most that arise in quantum mechanics) it is imprac- tical to computehHifor all members of a Hilbert space, the actual approach is to define a portion of the Hilbert space by introducing an assumed functional form for that contains parameters, and then to minimize hHiwith respect to the parameters; this is the source of the name “variation method.” The success of the method will depend on whether the functional form that is chosen is capable of representing functions that are “close” to the desired eigenfunction (meaning that its coefficient in the expansion is relatively large, with other coefficients much smaller). The great advantage of the variation method is that we do not need to know anything about the exact eigenfunction and we do not actually have to make an expansion; we simply choose a suitable functional form and minimize hHi. Since eigenvalue equations for energies and related quantities in quantum mechanics usually have finite smallest eigenvalues (e.g., ground energy levels), the variation method is frequently applicable. We point out that it is not a method having only academic inter- est; it is at the heart of some of the most powerful methods for solving the Schrödinger eigenvalue equation for complex quantum systems. Example 8.4.1 VARIATION METHOD Given a single-electron wave function (in three-dimensional space) of the form D3 1=2 er; (8.33) where the factor .=/3=2makes normalized, it can be shown that, in units with the electron mass, its charge, and Nh(Planck’s constant divided by 2) all set to unity (so-called Hartree atomic units), the quantum-mechanical kinetic energy operator has expectation valueh jTj iD2=2, and the potential energy of interaction between the electron and a fixed nucleus of charge CZ hash jVj iD Z. For a one-electron atom with a nucleus of chargeCZatrD0, the total energy will be less than or equal to the expectation value of the Hamiltonian HDTCV, given for the ofEq. (8.33) as hHiDhTiChViD2 2Z: (8.34) As is customary when the meaning is clear, we no longer explicitly show within all the angle brackets. We can now optimize our upper bound to the lowest eigenvalue of H by minimizing the expectation value hHiwith respect to the parameter in . To do so, we set d d2 2Z D0; leading toZD0, orDZ. This tells us that the wave function yielding the energy closest to the smallest eigenvalue is that with DZ, and the energy expectation value for this value ofisZ2=2Z2DZ2=2. ArfKen_Ch08-9780123846549.tex 8.4 Variation Method 397 The result we have just found is exact, because, with malice aforethought and with appropriate knowledge, we chose a functional form that included the exact wave function. But now let us continue to a two-electron atom, taking a wave function of the form 9D .1/ .2/ , with both of the samevalue. For this two-electron atom, the scalar product is defined as integration over the coordinates of both electrons, and the Hamiltonian is now HDT.1/CT.2/CV.1/CV.2/CU.1;2/, where T.i/andV.i/denote the kinetic energy and the electron-nuclear potential energy for electron i;U.1;2/is the electron- electron repulsion energy operator, equal in Hartree units to 1=r12, where r12is the distance between the positions of the two electrons. For the wave function in use here, the electron- electron repulsion has expectation value hUiD5=8 and the expectation value hHi(for ZD2, thereby representing the He atom) is hHiD2 2C2 2ZZC5 8D227 8: MinimizinghHiwith respect to , we obtain the optimum value D27=16 , and for this value ofwe havehHiD.27=16/2D2:8477 hartree. This is the best approximation available using a wave function of the form we chose. It cannot be exact, as the exact solu- tion for this system with two interacting electrons cannot be a product of two one-electron functions. We have therefore not included in our variational search the exact ground-state eigenfunction. A highly precise value of the smallest eigenvalue for this problem can only be obtained numerically, and in fact was produced by using the variation method with a trial function containing thousands of parameters and yielding a result accurate to about 40 decimal places.2The value found here by very simple means is higher than the exact value,2:9037hartree, by only about 2%, and already conveys much physically rele- vant information. If the two electrons did not interact, they would each have had an opti- mum wave function with D2; the fact that the optimum is somewhat smaller shows that each electron partially screens the nucleus from the other electron. From the viewpoint of the mathematical method in use here, it is desirable to note that we did not need to assume any relation between the trial wave function and the exact form of the eigenfunction; the variational optimization adjusts the trial function to give an ener- getically optimum fit. The quality of the final result of course depends on the degree to which the trial function can mimic the actual eigenfunction, and trial functions are ordi- narily chosen in a way that balances inherent quality against convenience of use.  Exercises 8.4.1 A function that is normalized on the interval 0x<1with an unweighted scalar product is D2 3=2xe x: (a) Verify the normalization. (b) Verify that for this ,hx1iD . 2C. Schwartz, Experiment and theory in computations of the He atom ground state, Int. J. Mod. Phys. E: Nuclear Physics 15: 877 (2006). ArfKen_Ch08-9780123846549.tex 398 Chapter 8 Sturm-Liouville Theory (c) Verify that for this ,hd2=dx2iD 2. (d) Use the variation method to find the value of that minimizes  1 2d2 dx21 x  ; and find the minimum value of this expectation value. 8.5 S UMMARY , EIGENVALUE PROBLEMS Because any Hermitian operator on a Hilbert space can be expanded in a basis and is there- fore mathematically equivalent to a matrix, all the properties derived for matrix eigenvalue problems automatically apply whether or not a basis-set expansion is actually carried out. It may be helpful to summarize some of those results, along with some that were developed in the present chapter. 1. A second-order differential operator is Hermitian if it is self-adjoint in the differential- equation sense and the functions on which it operates are required to satisfy appropri- ate boundary conditions. In that event, the scalar product consistent with Hermiticity is an unweighted integral over the range between its boundaries. 2. If a second-order differential operator is not self-adjoint in the differential-equation sense, it will nevertheless be Hermitian if it satisfies appropriate boundary condi- tions and if the scalar product includes the weight function that makes the original differential equation self-adjoint. 3. A Hermitian operator on a Hilbert space has a complete set of eigenfunctions. Thus, they span the space and can be used as basis for an expansion. 4. The eigenvalues of a Hermitian operator are real. 5. The eigenfunctions of a Hermitian operator corresponding to different eigenvalues are orthogonal, using the appropriate scalar product. 6. Degenerate eigenfunctions of a Hermitian operator can be orthogonalized using the Gram-Schmidt or any other orthogonalization process. 7. Two operators have a common set of eigenfunctions if and only if they commute. 8. An algebraic function of an operator has the same eigenfunctions as the original operator, and its eigenvalues are the corresponding function of the eigenvalues of the original operator. 9. Eigenvalue problems involving a differential operator may be solved either by expressing the problem in any basis and solving the resulting matrix problem or by using relevant properties of the differential equation. 10. The matrix representation of a Hermitian operator can be brought to diagonal form by a unitary transformation. In diagonal form, the diagonal elements are the eigenvalues, and the eigenvectors are the basis functions. The orthonormal eigenvectors are the columns of the unitary matrix U1when a Hermitian matrix His transformed to the diagonal matrix UHU1. ArfKen_Ch08-9780123846549.tex Additional Readings 399 11. Hermitian-operator eigenvalue problems which have a finite smallest eigenvalue may have their solutions approximated by the variation method, which is based on the theorem that for all members of the relevant Hilbert space, the expectation value of the operator will be larger than its smallest eigenvalue (or equal to it only if the Hilbert space member is actually a corresponding eigenfunction). Additional Readings Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969). Dennery, P., and A. Krzywicki, Mathematics for Physicists. Reprinted. New York: Dover (1996). Hirsch, M., Differential Equations, Dynamical Systems, and Linear Algebra. San Diego: Academic Press (1974). Miller, K. S., Linear Differential Equations in the Real Domain. New York: Norton (1963). Titchmarsh, E. C., Eigenfunction Expansions Associated with Second-Order Differential Equations, Part 1. 2nd ed. London: Oxford University Press (1962). Titchmarsh, E. C., Eigenfunction Expansions Associated with Second-Order Differential Equations. Part 2. London: Oxford University Press (1958). ArfKen_Ch09-9780123846549.tex CHAPTER 9 PARTIAL DIFFERENTIAL EQUATIONS 9.1 I NTRODUCTION As mentioned in Chapter 7, partial differential equations (PDEs) involve derivatives with respect to more than one independent variable; if the independent variables are xandy, a PDE in a dependent variable '.x;y/will contain partial derivatives, with the mean- ing discussed in Eq. (1.141). Thus, @'=@ ximplies an xderivative with yheld constant, @2'=@x2is the second derivative with respect to x(again keeping yconstant), and we may also have mixed derivatives @2' @x@yD@ @x@' @y : Like ordinary derivatives, partial derivatives (of any order, including mixed derivatives) are linear operators, since they satisfy equations of the type @T'.x;y/Cb'.x;y/U @xDa@'.x;y/ @xCb@'.x;y/ @x: Similar to the situation for ODEs, general differential operators, L, which may contain partial derivatives of any order, pure or mixed, multiplied by arbitrary functions of the independent variables, are linear operators, and equations of the form L'.x;y/DF.x;y/ are linear PDEs. If the source term F.x;y/vanishes, the PDE is termed homogeneous; ifF.x;y/is nonzero, it is inhomogeneous. 401 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch09-9780123846549.tex 402 Chapter 9 Partial Differential Equations Homogeneous PDEs have the property, previously noted in other contexts, that any linear combination of solutions will also be a solution to the PDE. This is the superposition principle that is fundamental in electrodynamics and quantum mechanics, and which also permits us to build specific solutions by the linear combination of suitable members of the set of functions constituting the general solution to the homogeneous PDE. Example 9.1.1 VARIOUS TYPES OF PDES Laplace r2 D0; linear, homogeneous Poisson r2 Df.r/; linear, inhomogeneous Euler (inviscid flow)@u @tCuruDrP nonlinear, inhomogeneous  Since the dynamics of many physical systems involve just two derivatives, for exam- ple, acceleration in classical mechanics, and the kinetic energy operator r2in quantum mechanics, differential equations of second order occur most frequently in physics. Even when the defining equations are first order, they may, as in Maxwell’s equations, involve two coupled unknown vector functions (they are the electric and magnetic fields), and the elimination of one unknown vector yields a second-order PDE for the other (compare Example 3.6.2). Examples of PDEs Among the most frequently encountered PDEs are the following: 1. Laplace’s equation, r2 D0. This very common and very important equation occurs in studies of (a) electromagnetic phenomena, including electrostatics, dielectrics, steady currents, and magnetostatics, (b) hydrodynamics (irrotational flow of perfect fluid and surface waves), (c) heat flow, (d) gravitation. 2. Poisson’s equation, r2 D=" 0. This inhomogeneous equation describes electrostatics with a source term =" 0. 3. Helmholtz and time-independent diffusion equations, r2 k2 D0. These equations appear in such diverse phenomena as (a) elastic waves in solids, including vibrating strings, bars, membranes, (b) acoustics (sound waves), (c) electromagnetic waves, (d) nuclear reactors. ArfKen_Ch09-9780123846549.tex 9.2 First-Order Equations 403 4. The time-dependent diffusion equation, r2 D1 a2@ @t. 5. The time-dependent classical wave equation,1 c2@2 @t2Dr2 . 6. The Klein-Gordon equation, @2 D2 , and the corresponding vector equations in which the scalar function is replaced by a vector function. Other, more complicated forms are also common. 7. The time-dependent Schrödinger wave equation, Nh2 2mr2 CV DiNh@ @t and its time-independent form Nh2 2mr2 CV DE : 8. The equations for elastic waves and viscous fluids and the telegraphy equation. 9. Maxwell’s coupled partial differential equations for electric and magnetic fields and those of Dirac for relativistic electron wave functions. We begin our study of PDEs by considering first-order equations, which illustrate some of the most important principles involved. We then continue to classification and prop- erties of second-order PDEs, and a preliminary discussion of prototypical homogeneous equations of the different classes. Finally, we examine a very useful and powerful method for obtaining solutions to homogeneous PDEs, namely the method of separation of variables. This chapter is mainly devoted to general properties of homogeneous PDEs; full detail on specific equations is for the most part postponed to chapters that discuss the spe- cial functions involved. Questions arising from the extension to inhomogeneous PDEs (i.e., problems involving sources ordriving terms) are also deferred, mainly to later chap- ters on Green’s functions and integral transforms. Occasionally, we encounter equations of higher order. In both the theory of the slow motion of a viscous fluid and the theory of an elastic body we find the equation .r2/2 D0: Fortunately, these higher-order differential equations are relatively rare and are not dis- cussed here. Sometimes, particularly in fluid mechanics, we encounter nonlinear PDEs. 9.2 F IRST-ORDER EQUATIONS While the most important PDEs arising in physics are linear and second order, many involving three spatial variables plus possibly a time variable, first-order PDEs do arise (e.g., the Cauchy-Riemann equations of complex variable theory). Part of the motivation for studying these easily solved equations is that the study provides insights that apply also to higher-order problems. ArfKen_Ch09-9780123846549.tex 404 Chapter 9 Partial Differential Equations Characteristics Let us start by considering the following homogeneous linear first-order equation in two independent variables xandy, with constant coefficients aandb, and with dependent variable'.x;y/: L'Da@' @xCb@' @yD0: (9.1) This equation would be easier to solve if we could rearrange it so that it contained only one derivative; one way to do this is would be to rewrite our PDE in terms of new coordinates .s;t/such that one of them, say s, is such that.@=@s/twould expand into the linear combi- nation of@=@xand@=@yin the original PDE, while the other new coordinate, t, is such that .@=@t/sdoes not occur in the PDE. It is easily verified that definitions of sandtconsistent with these objectives for the PDE in Eq. (9.1) aresDaxCbyandtDbxay. To check this, write'.x;y/D'.x.s;t/;y.s;t//DO'.s;t/, and we can verify that @' @x yDa@' @s tCb@' @t sand@' @y xDb@' @s ta@' @t s; so a@' @xCb@' @yD.a2Cb2/@O' @s: We see that the PDE does not contain a derivative with respect to t. Since our PDE now has the simple form .a2Cb2/@O' @sD0; it clearly has solution O'.s;t/Df.t/;with f.t/completely arbitrary. (9.2) In terms of the original variables, '.x;y/Df.bxay/; (9.3) where we again stress that f.t/is an arbitrary function of its argument. Checking our work to this point, we note that L'[email protected]ay/ @[email protected]ay/ @yDabf0.bxay/Cb a f0.bxay/ D0: Since the satisfaction of this equation does not depend on the properties of the function f, we verify that '.x;y/as given in Eq. (9.3) is a solution of our PDE, irrespective of the choice of the function f. In fact, it is the general solution of our PDE. It is useful to visualize the significance of what we have just observed. Note that holding tDbxayto a fixed value defines a line in the xyplane on which our solution 'is con- stant, with individual points on this line corresponding to different values of sDaxCby. In addition, we observe that the lines of constant sare orthogonal to those of constant t, and thatshas the same coefficients as the derivatives in the PDE. The general solution to our PDE can thus be characterized as independent of s and with arbitrary dependence on t. ArfKen_Ch09-9780123846549.tex 9.2 First-Order Equations 405 The curves of constant tare called characteristic curves, or more frequently just char- acteristics of our PDE. An alternative and insightful way of describing the characteristic curves is to observe that they are the stream lines (flow lines) of s. Put another way, they are the lines that are traced out as the value of sis changed, keeping tconstant. The char- acteristic can also be characterized by its slope, dy dxDb a;forLin Eq. (9.1). (9.4) For our present first-order PDE, the solution 'is constant along each characteristic. We shall shortly see that more general PDEs can be solved using ODE methods on charac- teristic lines, a feature that causes it to be said that PDE solutions propagate along the characteristics, giving further significance to the notion that in some sense these are lines of flow. In the present problem this translates into the statement that if we know 'at any point on a characteristic, we know it on the entire characteristic line. The characteristics have one additional (but related) property of importance. Ordinarily, if a PDE solution '.x;y/is specified on a curve segment (a boundary condition), one can deduce from it the values of the solution at nearby points that are not on the curve. If one introduces a Taylor expansion about some point .x0;y0/on the curve (thereby tacitly assuming that there are no singularities that invalidate the expansion), the value of 'at a nearby point .x;y/will be given by '.x;y/D'.x0;y0/C@'.x0;y0/ @x.xx0/C@'.x0;y0/ @y.yy0/C: (9.5) To use Eq. (9.5), we need values of the derivatives of '. To obtain these derivatives, note the following: The specification of 'on a given curve, with the curve parametrically described by x.l/;y.l/, means that the curve direction, i.e., dx=dlanddy=dl, is known, as is the derivative of 'along the curve, namely d' dlD@' @xdx dlC@' @ydy dl: (9.6) Equation (9.6) therefore provides us with a linear equation satisfied by the two deriva- tives@'=@ xand@'=@ y. The PDE supplies a second linear equation, in this case a@' @xCb@' @yD0: (9.7) Providing that the determinant of their coefficients is not zero, we can solve Eqs. (9.6) and(9.7) for@'=@ xand@'=@ yat.x0;y0/and therefore evaluate the leading terms of the Taylor series for '.x;y/.1The determinant of coefficients of Eqs. (9.6) and(9.7) takes the form DD dx dldy dl a b Dbdx dlady dl: 1The linear terms are all that are necessary; one can choose xandyclose enough to .x0;y0/that second- and higher-order terms can be made negligible relative to those retained. ArfKen_Ch09-9780123846549.tex 406 Chapter 9 Partial Differential Equations Now we make the observation that if 'was specified along a characteristic (for which tDbxayDconstant), we have b dxa dyD0; orbdx dlady dlD0; so that DD0and we cannot solve for the derivatives of '. Our conclusions relative to characteristics, which can be extended to more general equations, are: 1.If the dependent variable 'of the PDE in Eq. (9.1) is specified along a curve (i.e., ' has a boundary condition specified on a boundary curve), this fixes the value of ' at a point of each characteristic that intersects the boundary curve, and hence at all points of each such characteristic; 2.If the boundary curve is along a characteristic, the boundary condition on it will ordi- narily lead to inconsistency, and therefore, unless the boundary condition is redundant (i.e., coincidentally equal everywhere to the solution constructed from the value of ' at any one point on the characteristic), the PDE will not have a solution; 3.If the boundary curve has more than one intersection with the same characteristic, this will usually lead to an inconsistency, as the PDE may not have a solution that is simultaneously consistent with the values of 'at both intersections; and 4.Only if the boundary curve is nota characteristic can a boundary condition fix the value of'at points not on the curve. Values of 'specified only on a character- istic of the PDE provide no information as to the value of 'at points not on that characteristic. In the above example, the argument tof the arbitrary function fwas a linear combina- tion of xandy, which worked because the coefficients of the derivatives in the PDE were constants. If these coefficients were more general functions of xandy, the foregoing type of analysis could still be carried out, but the form of twould have to be different. This more complicated case is illustrated in Exercises 9.2.5 and9.2.6. More General PDEs Consider now a first-order PDE of a form more general than Eq. (9.1), L'Da@' @xCb@' @yCq.x;y/'DF.x;y/: (9.8) We may identify its characteristic curves just as before, which amounts to making a trans- formation to new variables sDaxCby,tDbxay, in terms of which our PDE becomes, compare Eq. (9.5), .a2Cb2/@' @s COq.s;t/O'DOF.s;t/: (9.9) HereOq.s;t/is obtained by converting q.x;y/to the new coordinates: Oq.s;t/DqasCbt a2Cb2;bsat a2Cb2 ; ArfKen_Ch09-9780123846549.tex 9.2 First-Order Equations 407 andOFis related in a similar fashion to F. Equation (9.9) is really an ODE in s(containing what can be viewed as a parameter, t), and its general solution can be obtained by the usual procedures for solving ODEs. Example 9.2.1 ANOTHER FIRST-ORDER PDE Consider the PDE @' @xC@' @yC.xCy/'D0: Applying a transformation to the characteristic direction tDxyand the direction orthogonal thereto sDxCy, our PDE becomes 2@' @sCs'D0: This equation separates into 2d' 'Cs dsD0; with general solution ln'Ds2 4CC.t/;or'Des2=4f.t/; where f.t/, originally expTC.t/U, is completely arbitrary. One can simplify the result slightly by noting that s2=4Dt2=4Cxy; then exp.t2=4/can be absorbed into f.t/, leaving the compact result (in terms of xandy) '.x;y/Dexyf.xy/;(farbitrary).  More Than Two Independent Variables It is useful to consider how the concept of characteristic can be generalized to PDEs with more than two independent variables. Given the three-dimensional (3-D) differential form a@' @xCb@' @yCc@' @z; we apply a transformation to convert our PDE to the new variables sDaxCbyCcz, tD 1xC 2yC 3z,uD 1xC 2yC 3z, with iand isuch that.s;t;u/form an orthogonal coordinate system. Then our 3-D differential form is found equivalent to .a2Cb2Cc2/@' @s; and the stream lines of s(those with tanduconstant) are our characteristics, along which we can propagate a solution 'by solving an ODE. Each characteristic can be identified by ArfKen_Ch09-9780123846549.tex 408 Chapter 9 Partial Differential Equations its fixed values of tandu. For the 3-D analog of Eq. (9.1), a@' @xCb@' @yCc@' @zD0; (9.10) we have .a2Cb2Cc2/@' @sD0; with solution 'Df.t;u/, with fa completely arbitrary function of its two arguments. Consider next an attempt to solve our 3-D PDE subject to a boundary condition fixing the values of the PDE solution 'on a surface. If the characteristic through a point on the surface lies in the surface, we have a potential inconsistency between the boundary condition and the solution propagated along the characteristic. We are then also unable to extend'away from the boundary surface because the data on the surface is insufficient to yield values of the derivatives that are needed for a Taylor expansion. To see this, note that the derivatives @'=@ x,@'=@ y, and@'=@ zcan only be determined if we can find two directions (parametrically designated landl0) such that we can solve simultaneously Eq. (9.10) and @' @lD@' @xdx dlC@' @ydy dlC@' @zdz dl; @' @l0D@' @xdx dl0C@' @ydy dl0C@' @zdz dl0: A solution can be obtained only if DD dx dldy dldz dl dx dl0dy dl0dz dl0 a b c 6D0: If a characteristic, with dx=dl00Da,dy=dl00Db, and dz=dl00Dc, lies in the two- dimensional (2-D) surface, there will only be one further linearly independent direction l, and Dwill necessarily be zero. Summarizing, our earlier observations extend to the 3-D case: A boundary condition is effective in determining a unique solution to a first-order PDE only if the boundary does not include a characteristic, and inconsistencies may arise if a characteristic intersects a boundary more than once. Exercises Find the general solutions of the PDEs in Exercises 9.2.1 to9.2.4. 9.2.1@ @xC2@ @yC.2xy/ D0. 9.2.2@ @x2@ @yCxCyD0. ArfKen_Ch09-9780123846549.tex 9.3 Second-Order Equations 409 9.2.3@ @xC@ @yD@ @z. 9.2.4@ @xC@ @yC@ @zDxy. 9.2.5 (a) Show that the PDE y@ @xCx@ @yD0 can be transformed into a readily soluble form by writing it in the new variables uDxy,vDx2y2, and find its general solution. (b) Discuss this result in terms of characteristics. 9.2.6 Find the general solution to the PDE x@ @xy@ @yD0: Hint. The solution to Exercise 9.2.5 may provide a suggestion as to how to proceed. 9.3 S ECOND -ORDER EQUATIONS Classes of PDEs We consider here extending the notion of characteristics to second-order PDEs. This can sometimes be done in a useful fashion. As a preliminary example, consider the following homogeneous second-order equation a2@2'.x;y/ @x2c2@2'.x;y/ @y2D0; (9.11) where aandcare assumed to be real. This equation can be written in the factored form  a@ @xCc@ @y a@ @xc@ @y 'D0; (9.12) and, since the two operator factors commute, we see that Eq. (9.12) will be satisfied if 'is a solution to either of the first-order equations a@' @xCc@' @yD0ora@' @xc@' @yD0: (9.13) However, these first-order equations are of just the type discussed in the preceding subsec- tion, so we can identify their respective general solutions as '1.x;y/Df.cxay/; ' 2.x;y/Dg.cxCay/; (9.14) where fandgare arbitrary (and totally unrelated) functions. Moreover, we can iden- tify the stream lines of axCcyandaxcyas characteristics, with implications as to the effectiveness and possible consistency of boundary conditions. For some PDEs with ArfKen_Ch09-9780123846549.tex 410 Chapter 9 Partial Differential Equations second derivatives as given in Eq. (9.11), it will also be practical to propagate solutions along the characteristics. Look next at the superficially similar equation a2@2'.x;y/ @x2Cc2@2'.x;y/ @y2D0; (9.15) with aandcagain assumed to be real. If we factor this, we get  a@ @xCic@ @y a@ @xic@ @y 'D0: (9.16) This factorization is of less practical value, as it leads to complex characteristics, which do not have an obvious relevance to boundary conditions. In addition, propagation along such characteristics does not provide a solution to the PDE for physically relevant (i.e., real) coordinate values. It is customary to identify second-order PDEs as hyperbolic if they are of (or can be transformed into) the form given in Eq. (9.11), with real values of aandc. PDEs that are of (or can be transformed into) the form given in Eq. (9.15) are called elliptic. The designation is useful because it correlates with the existence (or nonexistence) of real characteristics, and therefore with the behavior of the PDE relative to boundary conditions, with further implications as to convenient methods for solving the PDE. The terms elliptic andhyper- bolic have been introduced based on an analogy to quadratic forms, where a2x2Cc2y2Dd is the equation of an ellipse, while a2x2c2y2Ddis that of a hyperbola. More general PDEs will have second derivatives of the differential form LDa@2' @x2C2b@2' @x@yCc@2' @y2: (9.17) The form in Eq. (9.17) has the following factorization: LD bCp b2ac c1=2@ @xCc1=2@ @y! bp b2ac c1=2@ @xCc1=2@ @y! : (9.18) Equation (9.18) is easily verified by expanding the product. The equation also shows that the characteristics of Eq. (9.17) are real if and only if b2ac0. This quantity is well known from elementary algebra, being the discriminant of the quadratic form at2C2btCc. Ifb2ac>0, the two factors identify two linearly independent real charac- teristics, as were found for the prototype hyperbolic PDE discussed in Eqs. (9.11) to(9.14). Ifb2ac<0, the characteristics will, as for the prototype elliptic PDE in Eqs. (9.15) and(9.16), form a complex conjugate pair. We now have, however, one new possibility: Ifb2acD0(a case that for quadratic forms is that of a parabola), we have a PDE that has exactly one linearly independent characteristic; such PDEs are termed parabolic, and the canonical form adopted for them is a@' @xD@2' @y2: (9.19) If the original PDE lacked a @=@xterm, it would in effect be an ODE in ythat depends onxonly parametrically and need not be considered further in the context of methods for PDEs. ArfKen_Ch09-9780123846549.tex 9.3 Second-Order Equations 411 To complete our discussion of the second-order form in Eq. (9.17), we need to show that it can be transformed into the canonical form for the PDE of its classification. For this purpose we consider the transformation to new variables ,, defined as Dc1=2xc1=2by; Dc1=2y: (9.20) By systematic application of the chain rule to evaluate @2=@x2,@2=@x@y, and@2=@y2, it can be shown that LD.acb2/@2' @2C@2' @2: (9.21) Verification of Eq. (9.21) is the subject of Exercise 9.3.1. Equation (9.21) shows that the classification of our PDE remains invariant under trans- formation, and is hyperbolic if b2ac>0, elliptic if b2ac<0, and parabolic if b2acD0. Perhaps better seen from Eq. (9.18), we see that the stream lines of the char- acteristics have slope dy dxDc bp b2ac: (9.22) More than Two Independent Variables While we will not carry out a full analysis, it is important to note that many problems in physics involve more than two dimensions (often, three spatial dimensions or several spatial dimensions plus time). Often, the behavior in the multiple spatial dimensions is similar, and we apply the terms hyperbolic, elliptic, and parabolic in a way that relates the spatial to the time derivatives when the latter occur. Thus, these equations are classified as indicated: Laplace equation r2 D0 elliptic Poisson equation r2 D=" 0 elliptic Wave equation r2 D1 c2@2 @t2hyperbolic Diffusion equation a@ @tDr2 parabolic The specific equations mentioned here are very important in physics and will be further discussed in later sections of this chapter. These examples, of course, do not represent the full range of second-order PDEs, and do not include cases where the coefficients in the differential operator are functions of the coordinates. In that case, the classification into elliptic, hyperbolic, and parabolic is only local; the class may change as the coordinates vary. Boundary Conditions Usually, when we know a physical system at some time and the law governing the phys- ical process, then we are able to predict the subsequent development. Such initial val- ues are the most common boundary conditions associated with ODEs and PDEs. Finding ArfKen_Ch09-9780123846549.tex 412 Chapter 9 Partial Differential Equations solutions that match given points, curves, or surfaces corresponds to boundary value prob- lems. Solutions usually are required to satisfy certain imposed (for example, asymptotic) boundary conditions. These boundary conditions ordinarily take one of three forms: 1.Cauchy boundary conditions. The value of a function and normal derivative speci- fied on the boundary. In electrostatics this would mean ', the potential, and En, the normal component of the electric field. 2.Dirichlet boundary conditions. The value of a function specified on the boundary. In electrostatics, this would mean the potential '. 3.Neumann boundary conditions. The normal derivative (normal gradient) of a func- tion specified on the boundary. In the electrostatic case this would be Enand therefore , the surface charge density. Because the three classes of second-order PDEs have different patterns of character- istics, the boundary conditions needed to specify (in a consistent way) a unique solution will depend on the equation class. An exact analysis of the role of boundary conditions is complicated and beyond the scope of the present text. However, a summary of the relation of these three types of boundary conditions to the three classes of 2-D partial differential equations is given in Table 9.1. For a more extended discussion of these partial differ- ential equations the reader may consult Morse and Feshbach, Chapter 6 (see Additional Readings). Parts of Table 9.1 are simply a matter of maintaining internal consistency or of common sense. For instance, for Poisson’s equation with a closed surface, Dirichlet conditions lead Table 9.1 Relation between PDE and Boundary Conditions Boundary Class of Partial Differential Equation Conditions Elliptic Hyperbolic Parabolic Laplace, Poisson Wave equation in Diffusion equation in.x;y/ .x;t/ in.x;t/ Cauchy Open surface Unphysical results Unique, stable Too restrictive (instability) solution Closed surface Too restrictive Too restrictive Too restrictive Dirichlet Open surface Insufficient Insufficient Unique, stable solution in one direction Closed surface Unique, stable Solution not unique Too restrictive solution Neumann Open surface Insufficient Insufficient Unique, stable solution in one direction Closed surface Unique, stable Solution not unique Too restrictive solution ArfKen_Ch09-9780123846549.tex 9.3 Second-Order Equations 413 to a unique, stable solution. Neumann conditions, independent of the Dirichlet conditions, likewise lead to a unique stable solution independent of the Dirichlet solution. There- fore, Cauchy boundary conditions (meaning Dirichlet plus Neumann) could lead to an inconsistency. The term boundary conditions includes as a special case the concept of initial condi- tions. For instance, specifying the initial position x0and the initial velocity v0in some dynamical problem would correspond to the Cauchy boundary conditions. Note, how- ever, that an initial condition corresponds to applying the condition at only one end of the allowed range of the (time) variable. Finally, we note that Table 9.1 oversimplifies the situation in various ways. For example, the Helmholtz PDE, r2 k2 D0; (which could be thought of as the reduction of a parabolic time-dependent equation to its spatial part) has solution(s) for Dirichlet conditions on a closed boundary only for certain values of its parameter k. The determination of kand the characterization of these solutions is an eigenvalue problem and is important for physics. Nonlinear PDEs Nonlinear ODEs and PDEs are a rapidly growing and important field. We encountered earlier the simplest linear wave equation, @ @tCc@ @xD0; as the first-order PDE of the wavefronts of the wave equation. The simplest nonlinear wave equation, @ @tCc. /@ @xD0; (9.23) results if the local speed of propagation, c, is not constant but depends on the wave . When a nonlinear equation has a solution of the form .x;t/DAcos.kx!t/;where !.k/varies with kso that!00.k/6D0, then it is called dispersive. Perhaps the best-known nonlinear dispersive equation is the Korteweg-deVries equation, @ @tC @ @xC@3 @x3D0; (9.24) which models the lossless propagation of shallow water waves and other phenomena. It is widely known for its soliton solutions. A soliton is a traveling wave with the property of persisting through an interaction with another soliton: After they pass through each other, they emerge in the same shape and with the same velocity and acquire no more than a phase shift. Let .Dxct/be such a traveling wave. When substituted into Eq. (9.24) this yields the nonlinear ODE . c/d dCd3 d3D0; (9.25) ArfKen_Ch09-9780123846549.tex 414 Chapter 9 Partial Differential Equations which can be integrated to yield d2 d2Dc 2 2: (9.26) There is no additive integration constant in Eq. (9.26), because the solution must be such thatd2 =d2!0with !0for large. This causes to be localized at the character- isticD0, orxDct. Multiplying Eq. (9.26) byd =dand integrating again yields d d2 Dc 2 3 3; (9.27) where d =d!0for large. Taking the root of Eq. (9.27) and integrating again yields the soliton solution .xct/D3c cosh21 2pc.xct/: (9.28) Exercises 9.3.1 Show that by making a change of variables to Dc1=2xc1=2by,Dc1=2y, the operator Lof Eq. (9.18) can be brought to the form LD.acb2/@2 @2C@2 @2: 9.4 S EPARATION OF VARIABLES Partial differential equations are clearly important in physics, as evidenced by the PDEs listed in Section 9.1, and of equal importance is the development of methods for their solution. Our discussion of characteristics has suggested an approach that will be useful for some problems. Other general techniques for solving PDEs can be found, for example, in the books by Bateman and by Gustafson listed in the Additional Readings at the end of this chapter. However, the technique described in the present section is probably that most widely used. The method developed in this section for solution of a PDE splits a partial differential equation of nvariables into nordinary differential equations, with the intent that an overall solution to the PDE will be a product of single-variable functions which are solutions to the individual ODEs. In problems amenable to this method, the boundary conditions are usually such that they separate at least partially into conditions that can be applied to the separate ODEs. Further discussion of the method depends on the nature of the problem we seek to solve, so we now make the observation that PDEs occur in physics in two contexts, either as An equation with no unknown parameters for which there is expected to be a unique solution consistent with the boundary conditions (typical example: Laplace equation for the electrostatic potential with the potential specified on the boundary), or ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 415 An eigenvalue problem which will have solutions consistent with the boundary con- ditions only for certain values of an embedded but initially unknown parameter (the eigenvalue). In the first of these two cases, the unique solution is typically approached by first applying boundary conditions to the separate ODEs to specialize their solutions as much as possible. The solution is at this point normally not unique, and we have a (usually infinite) number of product solutions that satisfy the boundary conditions thus far applied. We then regard these product solutions as a basis that can be used to form an expansion that satisfies the remaining boundary condition(s). We illustrate with the first and fourth examples of this section. In the second case identified above, we typically have homogeneous boundary condi- tions (solution equal to zero on the boundary), and in favorable situations can satisfy all the boundary conditions by imposing them on the separate ODEs. At this point we usually find that each product solves our PDE with a different value of its embedded parameter, so that we are obtaining eigenfunctions and eigenvalues. This process is illustrated in the second and third examples of the present section. The method of separation of variables proceeds by dividing the PDE into pieces each of which can be set equal to a constant of separation. If our PDE has nindependent variables, there will be n1independent separation constants (though we often prefer a more symmetric formulation with nseparation constants plus an equation connecting them). The separation constants may have values that are restricted by invoking boundary conditions. To get a broad understanding of the method of separation of variables, it is useful to see how it is carried out in a variety of coordinate systems. Here we examine the process in Cartesian, cylindrical, and spherical polar coordinates. For application to other coordinate systems we refer the reader to the second edition of this text. Cartesian Coordinates In Cartesian coordinates the Helmholtz equation becomes @2 @x2C@2 @y2C@2 @z2Ck2 D0; (9.29) using Eq. (3.62) for the Laplacian. For the present, let k2be a constant. As stated in the introductory paragraphs of this section, our strategy will be to split Eq. (9.29) into a set of ordinary differential equations. To do so, let .x;y;z/DX.x/Y.y/Z.z/ (9.30) and substitute back into Eq. (9.29). How do we know Eq. (9.30) is valid? When the dif- ferential operators in various variables are additive in the PDE, that is, when there are no products of differential operators in different variables, the separation method has a chance to succeed. For success, it is usually also necessary that at least some of the boundary con- ditions separate into conditions on the separate factors. At any rate, we are proceeding in the spirit of let’s try and see if it works. If our attempt succeeds, then Eq. (9.30) will be ArfKen_Ch09-9780123846549.tex 416 Chapter 9 Partial Differential Equations justified. If it does not succeed, we shall find out soon enough and then we can try another attack, such as Green’s functions, integral transforms, or brute-force numerical analysis. With assumed given by Eq. (9.30), Eq. (9.29) becomes Y Zd2X dx2CX Zd2Y dy2CXYd2Z dz2Ck2XY ZD0: (9.31) Dividing by DXY Z and rearranging terms, we obtain 1 Xd2X dx2Dk21 Yd2Y dy21 Zd2Z dz2: (9.32) Equation (9.32) exhibits one separation of variables. The left-hand side is a function of xalone, whereas the right-hand side depends only on yandzand not on x. But x;y, andzare all independent coordinates. The equality of two sides that depend on different variables can only be attained if each side must be equal to the same constant, a constant of separation. We choose2 1 Xd2X dx2Dl2; (9.33) k21 Yd2Y dy21 Zd2Z dz2Dl2: (9.34) Now, turning our attention to Eq. (9.34), we obtain 1 Yd2Y dy2Dk2Cl21 Zd2Z dz2; (9.35) and a second separation has been achieved. Here we have a function of yequated to a function of z. We resolve it, as before, by equating each side to another constant of sepa- ration,m2, 1 Yd2Y dy2Dm2; (9.36) k2Cl21 Zd2Z dz2Dm2: (9.37) The separation is now complete, but to make the formulation more symmetrical, we will set 1 Zd2Z dz2Dn2; (9.38) and then consistency with Eq. (9.37) leads to the condition l2Cm2Cn2Dk2: (9.39) Now we have three ODEs, Eqs. (9.33), (9.36), and (9.38), to replace Eq. (9.29). Our assumption, Eq. (9.30), has succeeded in splitting the PDE; if we can also use the fac- tored form to satisfy the boundary conditions, our solution of the PDE will be complete. 2The choice of sign for separation constants is completely arbitrary, and will be fixed in specific problems by the need to satisfy specific boundary conditions, and particularly to avoid the unnecessary introduction of complex numbers. ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 417 It is convenient to label the solution according to the choice of our constants l;m, and n; that is, lmn.x;y;z/DXl.x/Ym.y/Zn.z/: (9.40) Subject to the boundary conditions of the problem being solved and to the condition k2Dl2Cm2Cn2, we may choose l,m, and nas we like, and Eq. (9.40) will still be a solution of Eq. (9.29), provided only that Xl.x/is a solution of Eq. (9.33), and so on. Because our original PDE is homogeneous and linear, we may develop the most general solution ofEq. (9.29) by taking a linear combination of solutions lmn, 9DX l;malm lmn; (9.41) where it is understood that nwill be given a value consistent with Eq. (9.39) and with the values of landm. Finally, the constant coefficients almmust be chosen to permit 9to satisfy the boundary conditions of the problem, leading usually to a discrete set of values l;m. Reviewing what we have done, it can be seen that the separation into ODEs could still have been achieved if k2were replaced by any function that depended additively on the variables, i.e., if k2! f.x/Cg.y/Ch.z/: A case of practical importance would be the choice k2!C.x2Cy2Cz2/, leading to the problem of a 3-D quantum harmonic oscillator. Replacing the constant term k2by a separable function of the variables will, of course, change the ODEs we obtain in the separation process and may have implications relative to the boundary conditions. Example 9.4.1 LAPLACE EQUATION FOR A PARALLELEPIPED As a concrete example we take Eq. (9.29) with kD0, which makes it a Laplace equation, and ask for its solution in a parallelepiped defined by the planar surfaces xD0,xDc, yD0,yDc,zD0,zDL, with the Dirichlet boundary condition D0on all the bound- aries except that at zDL; on that boundary is given the constant value V. See Fig. 9.1. This is a problem in which the PDE contains no unknown parameters and should have a unique solution. We expect a solution of the generic form given by Eq. (9.41), with lmngiven by Eq. (9.40). To proceed further, we need to develop the actual functional forms of X.x/, Y.y/, and Z.z/. For XandY, the ODEs, written in conventional form, are X00Dl2X;Y00Dm2Y; with general solutions XDAsinlxCBcoslx;YDA0sinmyCB0cosmy: We could have written XandYas complex exponentials, but that choice would be less convenient when we consider the boundary conditions. To satisfy the boundary condition atxD0, we set X.0/D0, which can be accomplished by choosing BD0; to satisfy ArfKen_Ch09-9780123846549.tex 418 Chapter 9 Partial Differential Equations x c cL yz FIGURE 9.1 Parallelepiped for solution of Laplace equation. the boundary condition at xDc, we set X.c/D0, which causes us to choose lsuch that lcD, wheremust be a nonzero integer. Without loss of generality, we can restrict to positive values, asXandXare linearly dependent. Moreover, we can include whatever scale factor is ultimately needed in our solution for Z.z/, so we may set AD1. Similar remarks apply to the solution Y.y/, so our solutions for XandYtake the final form X.x/Dsinx c ;Y.y/Dsiny c ; (9.42) withD1;2;3;::: andD1;2;3;::: . Next we consider the ODE for Z. It must be solved with a value of n2, calculated from Eq. (9.39) with kD0as n2D2 c2.2C2/: This equation suggests that nwill be imaginary, but that is unimportant here. Returning to the ODE for Z, we now see that it becomes Z00DC2 c2.2C2/Z; and the general solution for Z.z/for givenandis then easily identified as Z.z/DA ezCB ez;withD cp 2C2: (9.43) We now specialize Eq. (9.43) in a way that makes Z.0/D0andZ.L/DV. Noting thatsinh.z/is a linear combination of ezandez, we write Z.z/DVsinh.z/ sinh.L/: (9.44) At this point, we have made choices that cause all the boundary conditions to be satisfied except that at zDL, and we are now ready to select the coefficients aas required by the ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 419 remaining boundary condition, which because of Eq. (9.44) corresponds to 1 V9.x;y;L/DX asinx c siny c D1: (9.45) The symmetry of this expression suggests that we write aDbb, and find the coeffi- cients bfrom the equation X bsinx c D1: (9.46) Because the sine functions in Eq. (9.46) are the eigenfunctions of the one-dimensional (1-D) equation for X, which is a Hermitian eigenproblem, they form an orthogonal set on the interval.0;c/, so the bcan be computed by the following formulas: bD* sinx c 1+ * sinx c sinx c+DcZ 0sin. x=c/dx cZ 0sin2.x=c/dx D4 ;  odd; D0;  even; and our complete solution for the potential in the parallelepiped becomes 9.x;y;z/DVX bbsinx c siny csinh.z/ sinh.L/: (9.47)  As briefly mentioned earlier, PDEs also occur as eigenvalue problems. Here is a simple example. Example 9.4.2 QUANTUM PARTICLE IN A BOX We consider a particle of mass mtrapped in a box with planar faces at xD0,xDa,yD0, yDb,zD0,zDc. The quantum stationary states of this system are the eigenfunctions of the Schrödinger equation 1 2r2 .x;y;z/DE .x;y;z/; (9.48) where this PDE is subject to the Dirichlet boundary condition D0on the walls of the box. We identify Eas the stationary-state energy (the eigenvalue), in a system of units with mDNhD1. This is a Helmholtz equation with the new wrinkle that Eis not initially known. The boundary conditions are such that this PDE has no solution except for a set of discrete values of E. We want to find both those values and the corresponding eigenfunctions. ArfKen_Ch09-9780123846549.tex 420 Chapter 9 Partial Differential Equations Separating the variables in Eq. (9.48) by assuming a solution of the form Eq. (9.30), the PDE becomes X00 XCY00 YCZ00 Z D2E; (9.49) and the separation yields X00 XDl2;with solution XDAsinlxCBcoslx: After applying the boundary conditions at xD0andxDawe get (scaling to AD1) XDsinx a ; D1;2;3;:::; solD=a: (9.50) Because the Xequation is a 1-D Hermitian eigenvalue problem, these functions X.x/are orthogonal on 0xa. Similar processing of the YandZequations, with separation constants m2andn2, yields YDsiny b ; D1;2;3;:::; somD=b; ZDsinz c ; D1;2;3;:::; sonD=c;(9.51) yielding two additional 1-D eigenvalue problems. Replacing X00=X,Y00=Y,Z00=ZinEq. (9.49), respectively, by l2,m2,n2, and then evaluating these quantities from Eqs. (9.50) and(9.51), we have l2Cm2Cn2D2E;orED2 22 a2C2 b2C2 c2 ; (9.52) with,, andarbitrary positive integers. The situation is quite different from our solution, Example 9.4.1, of the Laplace equation. Instead of a unique solution we have an infinite set of solutions, corresponding to all positive integer triples .;;/ , each with its own value of E. Making the observation that the differential operator on the left-hand side of Eq. (9.47) is Hermitian in the presence of the chosen boundary conditions, we have found a complete orthogonal set of its eigenfunctions. The orthogonality is obvious, as it can be confirmed from the orthogonality of the X,Y, and Zon their respective 1-D intervals. Because we set the coefficients of all the sine functions to unity, our overall eigenfunctions are not normalized, but we can easily normalize them if we so choose. We close this example with the observation that this boundary-value problem will not have a solution for arbitrarily chosen values of E, as the Evalues must satisfy Eq. (9.52) with integer values of,, and. This will cause the Evalues of the problem solutions to be a discrete set; using terminology introduced in a previous chapter, our boundary-value problem can be said to have a discrete spectrum.  ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 421 Circular Cylindrical Coordinates Curvilinear coordinate systems introduce additional nuances into the process for separating variables. Again we consider the Helmholtz equation, now in circular cylindrical coordi- nates. With our unknown function dependent on ;', and z, that equation becomes, using Eq. (3.149) for r2: r2 .;'; z/Ck2 .;'; z/D0; (9.53) or 1 @ @ @ @ C1 2@2 @'2C@2 @z2Ck2 D0: (9.54) As before, we assume a factored form3for , .;'; z/DP./8.'/ Z.z/: (9.55) Substituting into Eq. (9.46), we have 8Z d d d P d CP Z 2d28 d'2CP8d2Z dz2Ck2P8ZD0: (9.56) All the partial derivatives have become ordinary derivatives. Dividing by P8Zand mov- ing the zderivative to the right-hand side yields 1 Pd d d P d C1 28d28 d'2Ck2D1 Zd2Z dz2: (9.57) Again, a function of zon the right appears to depend on a function of and'on the left. We resolve this by setting each side of Eq. (9.57) equal to the same constant. Let us choose4l2. Then d2Z dz2Dl2Z (9.58) and 1 Pd d d P d C1 28d28 d'2Ck2Dl2: (9.59) Setting k2Cl2Dn2; (9.60) multiplying by 2, and rearranging terms, we obtain  Pd d d P d Cn22D1 8d28 d'2: (9.61) 3For those with limited familiarity with the Greek alphabet, we point out that the symbol Pis the upper-case form of . 4Again, the choice of sign of the separation constant is arbitrary. However, the minus sign chosen for the axial coordinate zis optimum if we expect exponential dependence on z, from Eq. (9.58). A positive sign is chosen for the azimuthal coordinate 'in expectation of a periodic dependence on ', from Eq. (9.62). ArfKen_Ch09-9780123846549.tex 422 Chapter 9 Partial Differential Equations We set the right-hand side equal to m2, so d28 d'2Dm28; (9.62) and the left-hand side of Eq. (9.61) rearranges into a separate equation for : d d d P d C.n22m2/PD0: (9.63) Typically, Eq. (9.62) will be subject to the boundary condition that 8have periodicity 2 and will therefore have solutions eim'or, equivalently sinm';cosm';with integer m. Theequation, Eq. (9.63), is Bessel’s differential equation (in the independent variable n), originally encountered in Chapter 7. Because of its occurrence here (and in many other places relevant to physics), it warrants extensive study and is the topic of Chapter 14. The separation of variables of Laplace’s equation in parabolic coordinates also gives rise to Bessel’s equation. It may be noted that the Bessel equation is notorious for the variety of disguises it may assume. For an extensive tabulation of possible forms the reader is referred to Tables of Functions by Jahnke and Emde.5 Summarizing, we have found that the original Helmholtz equation, a 3-D PDE, can be replaced by three ODEs, Eqs. (9.58), (9.62), and (9.63). Noting that the ODE for contains the separation constants from the zand'equations, the solutions we have obtained for the Helmholtz equation can be written, with labels, as lm.;'; z/DPlm./8 m.'/Zl.z/; (9.64) where we probably should recall that the nin Eq. (9.63) for Pis a function of l(specif- ically, n2Dl2Ck2). The most general solution of the Helmholtz equation can now be constructed as a linear combination of the product solutions: 9.;'; z/DX l;malmPlm./8 m.'/Zl.z/: (9.65) Reviewing what we have done, we note that the separation could still have been achieved ifk2had been replaced by any additive function of the form k2! f.r/Cg.'/ 2Ch.z/: Example 9.4.3 CYLINDRICAL EIGENVALUE PROBLEM In this example we regard Eq. (9.53) as an eigenvalue problem, with Dirichlet boundary conditions D0on all boundaries of a finite cylinder, with k2initially unknown and to be determined. Our region of interest will be a cylinder with curved boundaries at DRand with end caps at zDL=2, as shown in Fig. 9.2. To emphasize that k2is an eigenvalue, 5E. Jahnke and F. Emde, Tables of Functions, 4th rev. ed., New York: Dover (1945), p. 146; also, E. Jahnke, F. Emde, and F. Lösch, Tables of Higher Functions, 6th ed., New York: McGraw-Hill (1960). ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 423 R+L/2 −L/2z FIGURE 9.2 Cylindrical region for solution of the Helmholtz equation. we rename it , and our eigenvalue equation is, symbolically, r2 D ; (9.66) with boundary conditions D0atDRand at zDL=2. Apart from constants, this is the time-independent Schrödinger equation for a particle in a cylindrical cavity. We limit the present example to the determination of the smallest eigenvalue (the ground state). This will be the solution to the PDE with the smallest number of oscillations, so we seek a solution without zeros (nodes) in the interior of the cylindrical region. Again, we seek separated solutions of the form given in Eq. (9.55). The ODEs for Zand 8,Eqs. (9.58) and(9.62), have the simple forms Z00Dl2Z; 800Dm28; with general solutions ZDA elzCB elz; 8DA0sinm'CB0cosm': We now need to specialize these solutions to satisfy the boundary conditions. The condition on8is simply that it be periodic in 'with period 2; this result will be obtained if mis any integer (including mD0, which corresponds to the simple solution 8Dconstant). Since our objective here is to obtain the least oscillatory solution, we choose that form, 8Dconstant, for8. Looking next at Z, we note that the arbitrary choice of sign for the separation constant l2has led to a form of solution that appears not to be optimum for fulfilling conditions requiring ZD0at the boundaries. But, writing l2D!2,lDi!,Zbecomes a linear combination of sin!zandcos!z; the least oscillatory solution with Z.L=2/D0isZD cos. z=L/, so!D=L, and l2D2=L2. The functions Z.z/and8.'/ that we have found satisfy the boundary conditions in z and'but it remains to choose P./in a way that produces PD0atDRwith the least oscillation in P. The equation governing P,Eq. (9.63), is 2P00CP0Cn22PD0; (9.67) ArfKen_Ch09-9780123846549.tex 424 Chapter 9 Partial Differential Equations where nwas introduced as satisfying (in the current notation) n2DCl2, see Eq. (9.60). Continuing now with Eq. (9.67), we identify as the Bessel equation of order zero in xDn. As we learned in Chapter 7, this ODE has two linearly independent solutions, of which only the one designated J0is nonsingular at the origin. Since we need here a solution that is regular over the entire range 0xnR, the solution we must choose is J0.n/. We can now see what is necessary to satisfy the boundary condition at DR, namely thatJ0.nR/vanish. This is a condition on the parameter n. Remembering that we want the least oscillatory function P, we need for nto be such that nRwill be the location of the smallest zero of J0. Giving this point the name (which by numerical methods can be found to be approximately 2.4048), our boundary condition takes the form nRD , or nD =R, and our complete solution to the Helmholtz equation can be written .;'; z/DJ0  R cosz L : (9.68) To complete our analysis, we must figure out how to arrange that nD =R. Since the condition connecting n,l, andrearranges to Dn2l2; (9.69) we see that the condition on ntranslates into one on . Our PDE has a unique ground- state solution consistent with the boundary conditions, namely an eigenfunction whose eigenvalue can be computed from Eq. (9.69), yielding D 2 R2C2 L2: If we had not restricted consideration to the ground state (by choosing the least oscillatory solution), we would have (in principle) been able to obtain a complete set of eigenfunctions, each with its own eigenvalue.  Spherical Polar Coordinates As a final exercise in the separation of variables in PDEs, let us try to separate the Helmholtz equation, again with k2constant, in spherical polar coordinates. Using Eq. (3.158), our PDE is 1 r2sin sin@ @r r2@ @r C@ @ sin@ @ C1 sin@2 @'2 Dk2 : (9.70) Now, in analogy with Eq. (9.30) we try .r;;'/DR.r/2./8.'/: (9.71) By substituting back into Eq. (9.70) and dividing by R28, we have 1 R r2d dr r2d R dr C1 2r2sind d sind2 d C1 8r2sin2d28 d'2Dk2: (9.72) ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 425 Note that all derivatives are now ordinary derivatives rather than partials. By multiplying byr2sin2, we can isolate .1=8/.d28=d'2/to obtain 1 8d28 d'2Dr2sin2 k21 R r2d dr r2d R dr 1 2r2sind d sind2 d :(9.73) Equation (9.73) relates a function of 'alone to a function of randalone. Since r,, and'are independent variables, we equate each side of Eq. (9.73) to a constant. In almost all physical problems, 'will appear as an azimuth angle. This suggests a periodic solution rather than an exponential. With this in mind, let us use m2as the separation constant, which then must be an integer squared. Then 1 8d28.'/ d'2Dm2(9.74) and 1 R r2d dr r2d R dr C1 2r2sind d sind2 d m2 r2sin2Dk2: (9.75) Multiplying Eq. (9.75) byr2and rearranging terms, we obtain 1 Rd dr r2d R dr Cr2k2D1 2sind d sind2 d Cm2 sin2: (9.76) Again, the variables are separated. We equate each side to a constant, , and finally obtain 1 sind d sind2 d m2 sin22C2D0; (9.77) 1 r2d dr r2d R dr Ck2RR r2D0: (9.78) Once more we have replaced a partial differential equation of three variables by three ODEs. The ODE for 8is the same as that encountered in cylindrical coordinates, with solutions exp. im'/orsinm',cosm'. The2ODE can be made less forbidding by changing the independent variable from totDcos, after which Eq. (9.77), with 2./ now written asP.cos/DP.t/, becomes .1t2/P00.t/2t P0.t/m2 1t2P.t/CP.t/D0: (9.79) This is the associated Legendre equation (called the Legendre equation ifmD0), and is discussed in detail in Chapter 15. We normally require solutions for P.t/that do not have singularities in the region within the range of the spherical polar coordinate (namely that it be nonsingular for the entire range 0, equivalent to1tC1 ). The solutions satisfying these conditions, called associated Legendre functions, are tradition- ally denoted Pm l, with la nonnegative integer. In Section 8.3 we discussed the Legendre equation as a 1-D eigenvalue problem, finding that the requirement of nonsingularity attD1 is a sufficient boundary condition to make its solutions well defined. We found also that its eigenfunctions are the Legendre polynomials and that its eigenvalues ArfKen_Ch09-9780123846549.tex 426 Chapter 9 Partial Differential Equations (in the present notation) have the values l.lC1/, where lis an integer. The generalization of these findings to the associated Legendre equation (that with nonzero m) shows that continues to be given as l.lC1/, but with the additional restriction that ljmj. Details are deferred to Chapter 15. Before continuing to the Requation, Eq. (9.78), let us observe that in deriving the 8and 2equations we have assumed that k2was a constant. However, if k2was not a constant, but an additive expression of the form k2! f.r/Cg./ r2Ch.'/ r2sin2; we could still carry out the separation of variables, but the relatively familiar 8and2 equations we have identified will be changed in ways that make them different, and prob- ably less tractable. However, if the departure of k2from a constant value is restricted to the form k2Dk2.r/, then the angular parts of the separation will remain as presented in Eqs. (9.74) and (9.79), and we only need to deal with increased generality in the R equation. It is worth stressing that the great importance of this separation of variables in spherical polar coordinates stems from the fact that the case k2Dk2.r/covers a tremendous amount of physics, such as a great deal of the theories of gravitation, electrostatics and atomic, nuclear, and particle physics. Problems with k2Dk2.r/can be characterized as central force problems, and the use of spherical polar coordinates is natural in such problems. From both a practical and a theoretical point of view, it is a key observation that the angu- lar dependence is isolated in Eqs. (9.74) and(9.77), or its equivalent, Eq. (9.79), that these equations are the same for all central force problems, and that they can be solved exactly. A detailed discussion of the angular properties of central force problems in quantum me- chanics is deferred to Chapter 16. Returning now to the remaining separated ODE, namely the Requation, we consider in some depth two special cases: (1) The case k2D0, corresponding to the Laplace equation, and (2) k2a nonzero constant, corresponding to the Helmholtz equation. For both cases we assume that the 8and2equations have been solved subject to the boundary conditions already discussed, so that the separation constant must have the value l.lC1/for some nonnegative integer l. Continuing on the assumption that k2is a (possibly zero) constant, Eq. (9.79) expands into r2R00C2r R0Ch k2r2l.lC1/i RD0: (9.80) Taking first the case of the Laplace equation, for which k2D0,Eq. (9.80) is easy to solve. Either by inspection or by attempting to carry out a series solution by the method of Frobenius, it is found that the initial term of the series, a0rs, is by itself a complete solution to Eq. (9.80). In fact, substituting the assumed solution RDrsinto Eq. (9.80), that equation reduces to s.s1/rsC2s rsl.lC1/rsD0; showing that s.sC1/Dl.lC1/, which has two solutions, sDl(obviously), and sD l1. In other words, given the value lfrom the choice of solution to the 2equation, ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 427 we find that the Requation (for the Laplace equation) has the two solutions rlandrl1, so its general solution takes the form R.r/DA rlCB rl1: (9.81) Combining the solutions to the separated ODEs, and summing over all choices of the separation constants, we see that the most general solution of the Laplace equation that has a nonsingular angular dependence can be written .r;;'/DX l;m.AlmrlCBlmrl1/Pm l.cos/.A0 lmsinm'CB0 lmcosm'/: (9.82) If our problem now has Dirichlet or Neumann boundary conditions on a spherical surface (with the region under study either within or outside the sphere), we may be able (by meth- ods more fully articulated in later chapters) to choose the coefficients in Eq. (9.82) so that the boundary conditions are satisfied. Note that if the region in which we are to solve the Laplace equation includes the origin, rD0, then only the rlterm should be retained and we set Blmto zero. If our region for the Laplace equation is, say, external to a sphere of some finite radius, then we must avoid the large- rdivergence of rland set Almto zero, retaining only rl1. More complicated cases, e.g., where we study the annular region between two concentric spheres, will require the retention of both AlmandBlmand will in general be somewhat more difficult. We continue now to the case of nonzero but constant k2.Equation (9.80) looks a lot like a Bessel equation, but differs therefrom by the coefficient “2” in the R0term and the factor k2that multiplies r2in the coefficient of R. Both these differences can be resolved by rewriting R.r/as R.r/DZ.kr/ .kr/1=2; (9.83) which will then give us a differential equation for Z. Carrying out the differentiations to obtain R0andR00in terms of Z, and changing the independent variable from rtoxDkr, Eq. (9.83) becomes x2Z00Cx Z0Ch x2 lC1 22i ZD0; (9.84) showing that Zis a Bessel function, of order lC1 2. Returning to Eq. (9.83), we can now identify R.r/in terms of quantities known as spherical Bessel functions, where jl.x/, the spherical Bessel functions that are regular at xD0, have definition jl.x/Dr 2xJlC1=2.x/: Since the status of R.r/as the solution to a homogeneous ODE is not affected by the scale factor in the definition of jl.x/, we see that Eq. (9.83) is equivalent to the observation that Eq. (9.80) has a solution jl.kr/. The spherical Bessel function that is the second solution of Eq. (9.80) is designated yl, so that solution is yl.kr/, and the general solution of Eq. (9.80) can be written R.r/DAjl.kr/CByl.kr/: (9.85) ArfKen_Ch09-9780123846549.tex 428 Chapter 9 Partial Differential Equations We note here that the properties of spherical Bessel functions are discussed more fully in Chapter 14. With the solutions to the radial ODE in hand, we can now write that the general solution to the Helmholtz equation in spherical polar coordinates takes the form .r;;'/DX l;m Almjl.kr/CBlmyl.kr/ Pm l.cos/.A0 lmsinm'CB0 lmcosm'/: (9.86) The above discussion assumes that k2>0; negative values of k2(and therefore imaginary values of k) simply correspond to our identifying an equation of the form .r2k2/ D0as a somewhat peculiar case of .r2Ck2/ D0. For negative k2, we can see we then get solutions that involve jl.kr/oryl.kr/with imaginary k. In order to avoid notations that unnecessarily involve imaginary quantities, it is usual to define a new set of functions il.x/that are proportional to jl.ix/, and are called modified spher- ical Bessel functions. The modified solutions parallel to yl.ix/are denoted kl.x/. These functions are also discussed in Chapter 14. The cases we have just surveyed do not, of course, cover all possibilities, and various other choices of k2.r/lead to problems that are of importance in physics. Without pro- ceeding to a detailed analysis here, we cite a couple: Taking k2DA=rCyields (with boundary condition that vanish in the limit r!1 ) the time-independent Schrödinger equation for the hydrogen atom; the R equation can then be identified as the associated Laguerre differential equation, dis- cussed in Chapter 18. Taking k2DAr2Cyields (with boundary condition at rD1 ) the equation for the 3-D quantum harmonic oscillator, for which the Requation can be reduced to the Hermite ODE, also discussed in Chapter 18. Some other boundary-value problems lead to well-studied ODEs. However, sometimes the practicing physicist will encounter a radial equation that may have to be solved using the techniques presented in Chapter 7, or if all else fails, by numerical methods. We close this subsection with an example that is a simple boundary-value problem in spherical coordinates. Example 9.4.4 SPHERE WITH BOUNDARY CONDITION In this example we solve the Laplace equation for the electrostatic potential .r/ in a region interior to a sphere of radius a, using spherical polar coordinates .r;;'/ with origin at the center of the sphere. Our solution is to be subject to the Neumann boundary condition d =dnDV0coson the spherical surface. See Fig. 9.3. To start, we note that totally arbitrary Neumann boundary conditions will not be consis- tent with our assumption of a charge-free sphere, as the integral of the normal derivative on the spherical surface gives, according to Gauss’ law, a measure of the total charge within. ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 429 z zero zero FIGURE 9.3 Arrows indicate sign and relative magnitude of the (inward) normal derivative of the electrostatic potential on a spherical surface (boundary condition for Example 9.4.4). The present example is internally consistent, as Z ScosdDZ 0d2Z 0d'cosD0: Next, we need to take the general solution for the Laplace equation within a sphere, as given by Eq. (9.82), and calculate therefrom the inward normal derivative at rDa. Since the normal is in the rdirection, we need only compute @ =@ r, evaluated at rDa. Noting that for the present problem BlmD0, our boundary condition becomes VcosDX l;ml Almal1Pm l.cos/.A0 lmsinm'CB0 lmcosm'/: Since the left-hand side of this equation is independent of ', its right-hand side has nonzero coefficients only for mD0, for which we only have the term originally containing B0 l0, because sin.0/D0. Thus, consolidating the constants, the boundary condition becomes the simpler form VcosDX ll Alal1Pl.cos/; (9.87) Without having made a detailed study of the properties of Legendre functions, the solution of an equation of this type might need to be deferred to Chapter 15, but this one is easy to solve because P1.cos/Dcos(see Legendre polynomials in Table 15.1) Thus, from Eq. (9.87), l Alal1DVl1; soA1DVand all the other coefficients except A0vanish. The coefficient A0is not deter- mined by the boundary conditions and represents an arbitrary constant that may be added ArfKen_Ch09-9780123846549.tex 430 Chapter 9 Partial Differential Equations to the potential. Thus, the potential within the sphere has the form DV r P 1.cos/CA0DV rcosCA0DV zCA0; corresponding to a uniform electric field within the sphere, in the zdirection and of magnitude V. The electric field is, of course, unaffected by the arbitrary value of the constant A0.  Summary: Separated-Variable Solutions For convenient reference, the forms of the solutions of Laplace’s and Helmholtz’s equa- tions for spherical polar coordinates are collected in Table 9.2. Although the ODEs obtained from the separation of variables are the same irrespective of the boundary con- ditions, the ODE solutions to be used, and the constants of separation, do depend on the boundaries. Boundaries with less than spherical symmetry may lead to values of mand lthat are not integral, and may also require use of the second solution of the Legendre equation (quantities normally denoted Qm l). Engineering applications frequently require solutions to PDEs for regions of low symmetry, but such problems are nowadays almost universally approached using numerical, rather than analytical methods. Consequently, Table 9.2 only contains data that are relevant for problems inside or outside a spherical boundary, or between two concentric spherical boundaries. This restriction to spherical symmetry causes the angular portion of the solutions to be uniquely of the form we have already identified. In contrast to the unique angular solution, both linearly independent solutions to the radial ODE are relevant, with the choice of solution dependent on the geometry. Solutions within a sphere must employ only the radial functions that are regular at the origin, i.e., rl,jl, oril. Solutions external to a sphere may employ rl1,kl(defined so that it will decay exponentially to zero at large r), or a linear combination of jlandyl(both of which are oscillatory and decay as r1=2). Solutions between concentric spheres can use both the radial functions appropriate to the PDE. It is also possible to summarize the forms of solution to the Laplace and Helmholtz equations in circular cylindrical coordinates, if we restrict attention to problems that have circular symmetry about the axial direction of the coordinate system. However, the situa- tion is considerably more complicated than for spherical coordinates, as we now have two Table 9.2 Solutions of PDEs in Spherical Polar Co- ordinatesa DX l;mfl.r/Pm l.cos/8 < :almcosm'Cblmsinm'/ or clmeim'9 = ; r2 D0 fl.r/Drl;rl1 r2 Ck2 D0 fl.r/Djl.kr/;yl.kr/ r2 k2 D0 fl.r/Dil.kr/;kl.kr/ aForil,jl,kl,yl, see Chapter 14; for Pm l, see Chapters 15 and 16. ArfKen_Ch09-9780123846549.tex 9.4 Separation of Variables 431 Table 9.3 Solutions of PDEs in Circular Cylindrical Coordinatesa DX m; fm ./g .z/8 < :am cosm'Cbm sinm'/ or cm eim'9 = ; r2 D0 fm ./DJm. /; Ym. / g .z/De z;e z or fm ./DIm. /; Km. / g .z/Dsin. z/;cos. z/orei z or fm ./Dm; mg .z/D1 r2 C D0 fm ./DJm. /; Ym. / if 2D 2>0, g .z/De z;e z if 2D 2>0, g .z/Dsin. z/;cos. z/orei z ifD 2, g .z/D1 or fm ./DIm. /; Km. / if 2D 2>0, g .z/De z;e z if 2DC 2>0, g .z/Dsin. z/;cos. z/orei z ifD 2, g .z/D1 or fm ./Dm; m if 2D> 0, g .z/De z;e z if 2D>0, g .z/Dsin. z/;cos. z/orei z aThe parameter can have any real values consistent with the boundary conditions. For Im, Jm,Km,Ym, see Chapter 14. coordinates ( andz) that can have a variety of boundary conditions, in contrast to the single such coordinate ( r) in the spherical system. In spherical coordinates the form of the radial function is completely determined by the PDE, and specific problems differ only in the choice (or relative weight) of the two linearly independent radial solutions. But in cylindrical coordinates the forms of the andzsolutions, as well as their coefficients, are determined by the boundary conditions, and not entirely by the value of the constant in the Helmholtz equation. Choices of the andzsolutions, though coupled, can vary widely. For details, the reader is referred to Table 9.3. Our final observations of this section deal with the functions we encountered in the course of the separations in cylindrical and spherical coordinates. For the purpose of this discussion, it is useful to think of our PDE as an operator equation subject to boundary conditions. If, in cylindrical coordinates, we restrict attention to PDEs in which the param- eterk2is independent of '(and with boundary conditions that do not depend upon '), we have chosen our operator equation as one that has circular symmetry. Moreover, we will then always get the same 8equation, with (of course) the same solutions. In these cir- cumstances, the solutions will have symmetry properties derived from those of our overall boundary-value problem.6The8equation can also be thought of as an operator equa- tion, and we can go further and identify the operator as L2 zD@2=@'2, where Lzis the zcomponent of the angular momentum. The solutions of the 8equation are eigenfunc- tions of this operator; the reason they can occur as part of the PDE solution is because 6Note that the solutions to a boundary-value problem need not have the full problem symmetry (a point that will be elaborated in great detail when we develop group-theoretical methods). An obvious example is that the Sun-Earth gravitational potential is spherically symmetric, while the most familiar solution (the Earth’s orbit) is planar. The dilemma is resolved by noting that the spherical symmetry manifests itself in the possible existence of Earth orbits at all angular orientations. ArfKen_Ch09-9780123846549.tex 432 Chapter 9 Partial Differential Equations L2 zcommutes with the operator defining the PDE (clearly so, because the PDE operator does not contain '). In other words, because L2 zand the PDE operator commute, they will have simultaneous eigenfunctions, and the overall solutions of the PDE can be labeled to identify the L2 zeigenfunction that was chosen. Looking now at the situation in spherical polar coordinates, we note that if k2is inde- pendent of the angles, i.e., k2Dk2.r/, then our PDE always has the same angular solutions 2lm./8 m.'/. Looking further at the angular terms of our PDE, we can identify them as the operator L2, and we see that the angular solutions we have found are eigenfunctions of this operator. When the PDE operator is independent of the angles, it will commute with L2and the solutions to the PDE can be labeled accordingly. These symmetry features are very important and are discussed in great detail in Chapter 16. Exercises 9.4.1 By letting the operator r2Ck2act on the general form a1 1.x;y;z/Ca2 2.x;y;z/, show that it is linear, i.e., that .r2Ck2/.a1 1Ca2 2/Da1.r2Ck2/ 1Ca2.r2C k2/ 2. 9.4.2 Show that the Helmholtz equation, r2 Ck2 D0; is still separable in circular cylindrical coordinates if k2is generalized to k2Cf./C .1=2/g.'/Ch.z/. 9.4.3 Separate variables in the Helmholtz equation in spherical polar coordinates, splitting off the radial dependence first. Show that your separated equations have the same form as Eqs. (9.74), (9.77), and (9.78). 9.4.4 Verify that r2 .r;;'/C k2Cf.r/C1 r2g./C1 r2sin2h.'/ .r;;'/D0 is separable (in spherical polar coordinates). The functions f,g, and hare functions only of the variables indicated; k2is a constant. 9.4.5 An atomic (quantum mechanical) particle is confined inside a rectangular box of sides a;b, and c. The particle is described by a wave function that satisfies the Schrödinger wave equation Nh2 2mr2 DE : The wave function is required to vanish at each surface of the box (but not to be identi- cally zero). This condition imposes constraints on the separation constants and therefore on the energy E. What is the smallest value of Efor which such a solution can be obtained? ANS. ED2Nh2 2m1 a2C1 b2C1 c2 . ArfKen_Ch09-9780123846549.tex 9.5 Laplace and Poisson Equations 433 9.4.6 The quantum mechanical angular momentum operator is given by LD i.rr/. Show that LL Dl.lC1/ leads to the associated Legendre equation. Hint. Section 8.3 and Exercise 8.3.1 may be helpful. 9.4.7 The 1-D Schrödinger wave equation for a particle in a potential field VD1 2kx2is Nh2 2md2 dx2C1 2kx2 DE .x/: (a) Defining aDmk Nh21=4 ; D2E Nhm k1=2 ; and settingDax, show that d2 ./ d2C.2/ ./D0: (b) Substituting ./Dy./e2=2; show that y./satisfies the Hermite differential equation. 9.5 L APLACE AND POISSON EQUATIONS The Laplace equation can be considered the prototypical elliptic PDE. At this point we supplement the discussion motivated by the method of separation of variables with some additional observations. The importance of Laplace’s equation for electrostatics has stim- ulated the development of a great variety of methods for its solution in the presence of boundary conditions ranging from simple and symmetrical to complicated and convoluted. Techniques for present-day engineering problems tend to rely heavily on computational methods. The thrust of this section, however, will be on general properties of the Laplace equation and its solutions. The basic properties of the Laplace equation are independent of the coordinate system in which it is expressed; we assume for the moment that we will use Cartesian coordinates. Then, because the PDE sets the sum of the second derivatives, @2 =@x2 i, to zero, it is obvious that if any of the second derivatives has a positive sign, at least one of the others must be negative. This point is illustrated in Example 9.4.1, where the xandydependence of a solution to the Laplace equation was sinusoidal, and as a result, the zdependence was exponential (corresponding to different signs for the second derivative). Since the second derivative is a measure of curvature, we conclude that if has positive curvature in any coordinate direction, it must have negative curvature in some other coordinate direction. That observation, in turn, means that all the stationary points of (points where its first derivatives in all directions vanish) must be saddle points, not maxima or minima. ArfKen_Ch09-9780123846549.tex 434 Chapter 9 Partial Differential Equations Since the Laplace equation describes the static electric potential in charge-free regions, we conclude that the potential cannot have an extremum at a point where there is no charge. A corollary to this observation is that the extrema of the electrostatic potential in a charge- free region must be on the boundary of the region. A related property of the Laplace equation is that its solution, subject to Dirichlet bound- ary conditions for the entire closed boundary of its region, is unique. This property applies also to its inhomogeneous generalization, the Poisson equation. The proof is simple: Sup- pose there are two distinct solutions 1and 2for the same boundary conditions. Then, their difference D 1 2(for either the Laplace or Poisson equation) will be a solution to the Laplace equation with D0on the boundary. Since cannot have extrema within the bounded region, it must be zero everywhere, meaning that 1D 2. If we have a Laplace or Poisson equation subject to Neumann boundary conditions on the entire closed boundary of its region, then the difference D 1 2of two solutions will also be a solution to the Laplace equation with a zero Neumann boundary condition. To analyze this situation, we invoke Green’s Theorem, in the form provided by Eq. (3.86), taking both uandvof that equation to be . Equation (3.86) then becomes Z S @ @ndSDZ V r2 dCZ Vr r d: (9.88) The boundary condition causes the left-hand side of Eq. (9.88) to vanish, the first integral on the right-hand side vanishes because is a solution of the Laplace equation, and the remaining integral on the right-hand side must therefore also vanish. But that integral can only vanish if r is zero everywhere, which can only be true if is constant. Thus, solutions to the Laplace equation with Neumann boundary conditions are also unique, except for an additive constant to the potential. An oft-cited application of this uniqueness theorem is the solution of electrostatics prob- lems by the method of images, which replaces a problem containing boundaries by one without a boundary but with additional charge added in such a way that the potential at the boundary location has the desired value. For example, a positive charge in front of a grounded boundary (one with D0) can be augmented by a negative charge at the mirror- image position behind the boundary. Then the two-charge system (ignoring the boundary) will yield the desired zero potential at the boundary location, and the uniqueness theorem tells us that the potential calculated for the two-charge system must be the same (within the original region) as that for the original system. Exercises 9.5.1 Verify that the following are solutions of Laplace’s equation: (a) 1D1=r;r6D0, (b) 2D1 2rlnrCz rz. 9.5.2 If9is a solution of Laplace’s equation, r29D0, show that@9=@ zis also a solution. 9.5.3 Show that an argument based on Eq. (9.88) can be used to prove that the Laplace and Poisson equations with Dirichlet boundary conditions have unique solutions. ArfKen_Ch09-9780123846549.tex 9.6 Wave Equation 435 9.6 W AVEEQUATION The wave equation is the prototype hyperbolic PDE. As we have seen earlier in this chap- ter, hyperbolic PDEs have two characteristics, and for the equation 1 c2@2 @t2D@2 @x2; (9.89) the characteristics are lines of constant xctand those of constant xCct. This means that the general solution to Eq. (9.89) takes the form .x;t/Df.xct/Cg.xCct/; (9.90) with fandgcompletely arbitrary. Viewing xas a position variable and tas the time, we can interpret f.xct/as a wave, moving with velocity c, in theCxdirection. By this we mean that the entire profile of f, as a function of xattD0, will be shifted uniformly toward positive xby an amount c when tD1. See Fig. 9.4. Similarly, g.xCct/describes a wave moving at velocity cin the xdirection. Because fandgare arbitrary, the traveling waves they describe need not be sinusoidal or periodic, but may be entirely irregular; moreover, there is no requirement that fandghave any particular relationship to each other. An obvious special case of the general situation described above is that when f.xct/ is chosen to be sinusoidal, fDsin.xct/. For simplicity we have taken fto have unit amplitude and wavelength 2. We also take g.xCct/to be gDsin.xCct/, a sinusoidal wave of the same wavelength and amplitude traveling in the direction opposite to f. At a point xand time t, these two waves add to produce a resultant .x;t/Dsin.xct/Csin.xCct/; which, using trigonometric identities, can be rearranged to .x;t/D.sinxcosctcosxsinct/C.sinxcosctCcosxsinct/D2 sin xcosct: This form for can be identified as a standing wave distribution, meaning that the time evolution of the wave’s profile in xis an oscillation in amplitude, with the wave pattern not moving in either direction. An obvious point of difference from a traveling wave is that for a standing wave, the nodes (points where D0) are stationary in time, while in a traveling wave they are moving in time at velocity c. Our current interest in traveling vs. standing waves is their relation to solutions to the wave equation that we might find using the method of separation of variables. That method would obviously lead us to standing-wave solutions. However, it is useful to note that the totality of the solution set from the separated variables has the same content as x⇒ ⇒ FIGURE 9.4 Traveling wave f.xct/. Dashed line is profile at tD0; full line is profile at a time t>0. ArfKen_Ch09-9780123846549.tex 436 Chapter 9 Partial Differential Equations the traveling-wave solutions. For example, the products sinxcosctandcosxsinctare solutions we would get by separating the variables, and linear combinations of these yield sin.xct/. d’Alembert’s Solution While all ways of writing the general solution to the wave equation are mathemati- cally equivalent, diverse forms differ in their convenience of use for various purposes. To illustrate this, we consider how we might construct a solution to the wave equation, given, as an initial condition, (1) the entire spatial distribution of the wave amplitude at tD0and (2) the time derivative of the wave amplitude at tD0for the entire spatial distri- bution. The solution to this problem is generally referred to as d’Alembert’s solution of the wave equation; it was also (and slightly earlier) found by Euler. We start by using Eq. (9.90) to write our initial conditions in terms of the presently unknown functions fandg: .x;0/Df.x/Cg.x/; (9.91) @ .x;t/ @t tD0Dcf0.x/Ccg0.x/: (9.92) We now integrate Eq. (9.92) between the limits xctandxCct(and divide the result by2c), obtaining 1 2cxCctZ xct@ .x;0/ @tdxD1 2 f.xCct/Cf.xct/Cg.xCct/g.xct/ :(9.93) From Eq. (9.91), we also have 1 2 .xCct;0/C .xct;0/ D 1 2 f.xCct/Cg.xCct/Cf.xct/Cg.xct/ : (9.94) Adding together the right-hand sides of Eqs. (9.93) and(9.94), half the terms cancel, and those that survive combine to give the result f.xct/Cg.xCct/;which is .x;t/: Therefore, from the left-hand sides of Eqs. (9.93) and(9.94), we obtain the final result .x;t/D1 2 .xCct;0/C .xct;0/ C1 2cxCctZ xct@ .x;0/ @tdx: (9.95) This equation gives .x;t/in terms of data at tD0that are within the distance ctof the point x. This is a reasonable result, since ctis the distance that waves in this problem can move between times tD0andtDt. More specifically, Eq. (9.95) contains terms that represent half the tD0amplitude at distances ctfrom x(half, because a disturbance that starts at these points is split between propagation in both directions), plus an additional integral that accumulates the effect of the initial amplitude derivative over the region of influence. ArfKen_Ch09-9780123846549.tex 9.7 Heat-Flow, or Diffusion PDE 437 Exercises Solve the wave equation, Eq. (9.89), subject to the indicated conditions. 9.6.1 Determine .x;t/given that at tD0 0.x/Dsinxand@ .x/=@tDcosx. 9.6.2 Determine .x;t/given that at tD0 0.x/D.x/(Dirac delta function) and the initial time derivative of is zero. 9.6.3 Determine .x;t/given that at tD0 0.x/is a single square-wave pulse as defined below, and the initial time derivative of is zero. 0.x/D0;jxj>a=2; 0.x/D1=a;jxj<a=2: 9.6.4 Determine .x;t/given that at tD0 0D0for all x, but@ =@ tDsin.x/. 9.7 H EAT-FLOW,ORDIFFUSION PDE Here we return to a parabolic PDE to develop methods that adapt a special solution of a PDE to boundary conditions by introducing parameters. The methods are fairly general and apply to other second-order PDEs with constant coefficients as well. To some extent, they are complementary to the earlier basic separation method for finding solutions in a systematic way. We consider the 3-D time-dependent diffusion PDE for an isotropic medium, using it to describe heat flow subject to given boundary conditions. Assuming isotropy actually is not much of a restriction because, in case we have different (constant) rates of diffusion in different directions, for example, in wood, our heat-flow PDE takes the form @ @tDa2@2 @x2Cb2@2 @y2Cc2@2 @z2; (9.96) if we put the coordinate axes along the principal directions of anisotropy. Now we sim- ply rescale the coordinates using the substitutions xDa;yDb;zDcto get back the original isotropic form of Eq. (9.96), @8 @tD@28 @2C@28 @2C@28 @2; (9.97) for the temperature distribution function 8.;;; t/D .x;y;z;t/. For simplicity, we first solve the time-dependent PDE for a homogeneous one- dimensional medium, a long metal rod in the x-direction, for which the PDE is @ @tDa2@2 @x2; (9.98) where the constant ameasures the diffusivity, or heat conductivity, of the medium. We obtain solutions to this linear PDE with constant coefficients by the method of separation of variables, for which we set .x;t/DX.x/T.t/, leading to the separate equations 1 TdT dtD ;1 Xd2X dx2D a2: These equations have, for any nonzero value of , solutions TDe tandXDe x, with 2D =a2. We seek solutions whose time dependence decays exponentially at large t, ArfKen_Ch09-9780123846549.tex 438 Chapter 9 Partial Differential Equations that is, solutions with negative values of , and therefore set Di!,a2D!2for real!, and have .x;t/Dei!xe!2a2tD.cos!xisin!x/e!2a2t: (9.99) Note that D0, for which .x;t/DC0 0xCC0; (9.100) is also included in the solution set for the PDE. If we use this solution for a rod of infinite length, we must set C0 0D0to avoid a nonphysical divergence; in any case, the value of C0 is then the constant value that the temperature approaches at long times. Forming real linear combinations of sin!xandcos!xwith arbitrary coefficients, and keeping the D0solution, we obtain from Eq. (9.99) for any choice of A,B,!,C0 0, and C0, a solution .x;t/D.Acos!xCBsin!x/e!2a2tCC0 0xCC0: (9.101) Solutions for different values of these parameters can now be combined as needed to form an overall solution consistent with the required boundary conditions. If the rod we are studying is finite in length, it may be that the boundary conditions can be satisfied if we restrict !to discrete nonzero values that are multiples of a basic value !0. For a rod of infinite length, it may be better to let !assume a continuous range of values, so that .x;t/will have the general form .x;t/DZ TA.!/cos!xCB.!/sin!xUea2!2td!CC0: (9.102) We call specific attention to the fact that Forming linear combinations of solutions by summation or integration over parameters is a powerful and standard method for generalizing specific PDE solutions in order to adapt them to boundary conditions. Example 9.7.1 A SPECIFIC BOUNDARY CONDITION Let us solve a 1-D case explicitly, where the temperature at time tD0is 0.x/D1D constant in the interval between xDC1 andxD1 and zero for x>1andx<1. At the ends, xD1; the temperature is always held at zero. Note that this problem, including its initial conditions, has even parity, 0.x/D 0.x/, so .x;t/must also be even. We choose the spatial solutions of Eq. (9.98) to be of the form given in Eq. (9.101), but restricted to C0 0DC0D0(since the t!1 limit of .x;t/is zero for the entire range1x1), and to cos.lx=2/for odd integer l, because these functions are the even-parity members of an orthonormal basis for the interval 1x1that satisfy the boundary condition D0atxD1 . Then, at tD0our solution takes the form .x;0/D1X lD1alcoslx 2;1<x<1; and we need to choose the coefficients also that .x;0/D1. ArfKen_Ch09-9780123846549.tex 9.7 Heat-Flow, or Diffusion PDE 439 Using the orthonormality, we compute alD1Z 11coslx 2D2 lsinlx 2 1 xD1 D4 lsinl 2D4.1/m .2mC1/;lD2mC1: Including its time dependence, the full solution is given by the series .x;t/D4 1X mD0.1/m 2mC1cosh .2mC1/x 2i et..2mC1/ a=2/2; (9.103) which converges absolutely for t>0but only conditionally at tD0, as a result of the discontinuity at xD1 .  We are now ready to consider the diffusion equation in three dimensions. We start by assuming a solution of the form Df.x;y;z/T.t/, and separate the spatial from the time dependence. As in the 1-D case, T.t/will have exponentials as solutions, and we can choose the solution that decays exponentially at large t. Assigning the separation constant the valuek2, so that the time dependence is exp.k2t/, the separated equation in the spatial coordinates takes the form @2f @x2C@2f @y2C@2f @z2Ck2fD0; (9.104) which we recognize as the Helmholtz equation. Assuming that we can solve this equation for various values of k2by further separations of variables or by other means, we can form whatever sum or integral of individual solutions that may be needed to satisfy the boundary conditions. Alternate Solutions In an alternative approach to the heat flow equation, we now return to the one-dimensional PDE, Eq. (9.98), seeking solutions of a new functional form .x;t/Du.x=pt/, which is suggested by dimensional considerations and experimental data. Substituting u./,D x=pt, into Eq. (9.98) using @ @xDu0 pt;@2 @x2Du00 t;@ @tDx 2p t3u0(9.105) with the notation u0./du=d, the PDE is reduced to the ODE 2a2u00./Cu0./D0: (9.106) Writing this ODE as u00 u0D 2a2; ArfKen_Ch09-9780123846549.tex 440 Chapter 9 Partial Differential Equations we can integrate it once to get lnu0D2=4a2ClnC1, where C1is an integration constant. Exponentiating and integrating again we find the general solution u./DC1Z 0e2=4a2dCC2; (9.107) which contains two integration constants Ci. We initialize this solution at time tD0to temperatureC1forx>0and1forx<0, corresponding to u.1/DC1 andu.1/D 1. Noting that 1Z 0e2=4a2dDap; a case of the integral evaluated in Eq. (1.148), we obtain u.1/DapC1CC2D1; u.1/DapC1CC2D1; which fixes the constants C1D1=ap,C2D0. We therefore have the specific solution D1 apx=ptZ 0e2=4a2dD2px=2aptZ 0ev2dvDerfx 2apt ; (9.108) where erfis the standard name for Gauss’ error function (one of the special functions listed in Table 1.2). We need to generalize this specific solution to adapt it to boundary conditions. To this end we now generate new solutions of the PDE with constant coefficients by differentiating the special solution given in Eq. (9.108). In other words, if .x;t/ solves the PDE in Eq. (9.98), so do @ =@ tand@ =@ x, because these derivatives and the differentiations of the PDE commute; that is, the order in which they are carried out does not matter. Note carefully that this method no longer works if any coefficient of the PDE depends on torxexplicitly. However, PDEs with constant coefficients dominate in physics. Examples are Newton’s equations of motion in classical mechanics, the wave equations of electrodynamics, and Poisson’s and Laplace’s equations in electrostatics and gravity. Even Einstein’s nonlinear field equations of general relativity take on this special form in local geodesic coordinates. Therefore, by differentiating Eq. (9.108) with respect to x;we find the simpler, more basic solution, 1.x;t/D1 aptex2=4a2t; (9.109) and, repeating the process, another basic solution 2.x;t/Dx 2a3p t3ex2=4a2t: (9.110) Again, these solutions have to be generalized to adapt them to boundary conditions. And there is yet another method of generating new solutions of a PDE with constant coeffi- cients: We can translate a given solution, for example, 1.x;t/! 1.x ;t/;and then ArfKen_Ch09-9780123846549.tex 9.7 Heat-Flow, or Diffusion PDE 441 integrate over the translation parameter :Therefore, .x;t/D1 2apt1Z 1C. /e.x /2=4a2td (9.111) is again a solution, which we rewrite using the substitution Dx 2apt; Dx2ap t;d D2 ap td: (9.112) These substitutions lead to .x;t/D1p1Z 1C.x2ap t/e2d; (9.113) a solution of our PDE. Equation (9.113) is in a form permitting us to understand the sig- nificance of the weight function C.x/from the translation method. If we set tD0in that equation, the function Cin the integrand then becomes independent of , and the integral can then be recognized as 1Z 1e2dDp; a well-known result equivalent to Eq. (1.148). Equation (9.113) then becomes the simpler form .x;0/DC.x/;orC.x/D 0.x/; where 0is the initial spatial distribution of . Using this notation, we can write the solution to our PDE as .x;t/D1p1Z 1 0.x2ap t/e2d; (9.114) a form that explicitly displays the role of the boundary (initial) condition. From Eq. (9.114) we see that the initial temperature distribution, 0.x/;spreads out over time and is damped by the Gaussian weight function. Example 9.7.2 SPECIAL BOUNDARY CONDITION AGAIN We consider now a problem similar to Example 9.7.1, but instead of keeping D0at all times at xD1 , we regard the system as infinite in length, with 0D0everywhere except forjxj<1, where 0D1. This change makes Eq. (9.114) usable, because our PDE now applies over the range .1;1/, and heat will flow (and temporarily increase the temperature) at and beyond jxjD1. The range of 0.x/corresponds to a range of with endpoints found from x2aptD 1, so our solution becomes .x;t/D1p.xC1/=2 aptZ .x1/=2 apte2d: ArfKen_Ch09-9780123846549.tex 442 Chapter 9 Partial Differential Equations In terms of the error function, we can also write this solution as .x;t/D1 2 erfxC1 2apt erfx1 2apt : (9.115) Equation (9.115) applies for all x, includingjxj>1.  Next we consider the problem of heat flow for an extended spherically symmetric medium centered at the origin, suggesting that we should use polar coordinates r;;' . We expect a solution of the form u.r;t/. Using Eq. (3.158) for the Laplacian, we find the PDE @u @tDa2@2u @r2C2 r@u @r ; (9.116) which we transform to the 1-D heat-flow PDE by the substitution uDv.r;t/ r;@u @rD1 r@v @rv r2;@u @tD1 r@v @t; @2u @r2D1 r@2v @r22 r2@v @rC2v r3: (9.117) This yields the PDE @v @tDa2@2v @r2: (9.118) Example 9.7.3 SPHERICALLY SYMMETRIC HEAT FLOW Let us apply the 1-D heat-flow PDE to a spherically symmetric heat flow under fairly common boundary conditions, where xis replaced by the radial variable. Initially we have zero temperature everywhere. Then, at time tD0, a finite amount of heat energy Qis released at the origin, spreading evenly in all directions. What is the resulting spatial and temporal temperature distribution? Inspecting our special solution in Eq. (9.110) we see that, for t!0, the temperature v.r;t/ rDCp t3er2=4a2t(9.119) goes to zero for all r6D0, so zero initial temperature is guaranteed. As t!1 , the temper- aturev=r!0for all rincluding the origin, which is implicit in our boundary conditions. The constant Ccan be determined from energy conservation, which gives (for arbitrary t) the constraint QDZv rdD4 Cp t31Z 0r2er2=4a2tdrD8p 3a3C; (9.120) ArfKen_Ch09-9780123846549.tex 9.7 Heat-Flow, or Diffusion PDE 443 whereis the constant density of the medium and is its specific heat. The final result in Eq. (9.120) is obtained by first making a change of variable from rtoDr=2apt, obtaining 1Z 0er2=4a2tr2drD.2ap t/31Z 0e22d; then evaluating the integral via an integration by parts: 1Z 0e22dD 2e2 1 0C1 21Z 0e2dDp 4: The temperature, as given by Eq. (9.119) at any moment, i.e., at fixed t, is a Gaussian distribution that flattens out as time increases, because its width is proportional topt. As a function of time the temperature at any fixed point is proportional to t3=2eT=t, with Tr2=4a2. This functional form shows that the temperature rises from zero to a maximum and then falls off to zero again for large times. To find the maximum, we set d dt t3=2eT=t Dt5=2eT=tT t3 2 D0; (9.121) from which we find tmaxD2T=3Dr2=6a2. The temperature maximum arrives at later times at larger distances from the origin.  In the case of cylindrical symmetry (in the plane zD0in plane polar coordinates D p x2Cy2;'), we look for a temperature Du.;t/that then satisfies the ODE (using Eq. (2.35) in the diffusion equation) @u @tDa2@2u @2C1 @u @ ; (9.122) which is the planar analog of Eq. (9.118). This ODE also has solutions with the functional dependence=ptr. Upon substituting uDvpt ;@u @tDv0 2t3=2;@u @Dv0 pt;@2u @2Dv0 t(9.123) into Eq. (9.122) with the notation v0dv=dr , we find the ODE a2v00Ca2 rCr 2 v0D0: (9.124) This is a first-order ODE for v0, which we can integrate when we separate the variables v andras v00 v0D1 rCr 2a2 : (9.125) This yields v.r/DC rer2=4a2DCpt e2=4a2t: (9.126) ArfKen_Ch09-9780123846549.tex 444 Chapter 9 Partial Differential Equations This special solution for cylindrical symmetry can be similarly generalized and adapted to boundary conditions, as for the spherical case. Finally, the z-dependence can be factored in, because zseparates from the plane polar radial variable . Exercises 9.7.1 For a homogeneous spherical solid with constant thermal diffusivity, K, and no heat sources, the equation of heat conduction becomes @T.r;t/ @tDKr2T.r;t/: Assume a solution of the form TDR.r/T.t/ and separate variables. Show that the radial equation may take on the standard form r2d2R dr2C2rd R drC 2r2RD0; and that sin r=randcos r=rare its solutions. 9.7.2 Separate variables in the thermal diffusion equation of Exercise 9.7.1 in circular cylin- drical coordinates. Assume that you can neglect end effects and take TDT.;t/. 9.7.3 Solve the PDE @ @tDa2@2 @x2; to obtain .x;t/for a rod of infinite extent (in both the Cxandxdirections), with a heat pulse at time tD0that corresponds to 0.x/DA.x/. 9.7.4 Solve the same PDE as in Exercise 9.7.3 for a rod of length L, with position on the rod given by the variable x, with the two ends of the rod at xD0andxDLkept (at all times t) at the respective temperatures TD1andTD0, and with the rod initially at T.x/D0, for0<xL. 9.8 S UMMARY This chapter has provided an overview of methods for the solution of first- and second- order linear PDEs, with emphasis on homogeneous second-order PDEs subject to bound- ary conditions that either determine unique solutions or define eigenvalue problems. We found that the usual boundary conditions are identified as of Dirichlet type (solution spec- ified on boundary), Neumann type (normal derivative of solution specified on boundary), or Cauchy type (both solution and its normal derivative specified). Applicable types of boundary conditions depend on the classification of the PDE; second-order PDEs are clas- sified as hyperbolic (e.g., wave equation), elliptic (e.g., Laplace equation), or parabolic (e.g., heat/diffusion equation). ArfKen_Ch09-9780123846549.tex Additional Readings 445 The method of widest applicability to the solution of PDEs is the method of separation of variables, which, when effective, reduces a PDE to a set of ODEs. The chapter has pre- sented a very small number of complete PDE solutions to illustrate the technique. A wider variety of examples only becomes possible when we are prepared to exploit the proper- ties of the special functions that are the solutions of various ODEs, and, as a result, fuller illustration of PDE solutions will be provided in the chapters that discuss these special functions. We point out, in particular, that general PDEs with spherical symmetry all have the same angular solutions, known as spherical harmonics. These, and the functions from which they are constructed (Legendre polynomials and associated Legendre functions), are the subject matter of Chapters 15 and 16. Some spherically symmetric problems have radial solutions that can be identified as spherical Bessel functions; these are treated in the Bessel function chapter (Chapter 14). PDE problems with cylindrical symmetry usually involve Bessel functions, often in ways more complex than in the examples of the present chapter. Further illustrations appear in Chapter 14. This chapter has not attempted to discuss methods for the solution of inhomogeneous PDEs. That topic deserves its own chapter, and will be developed in Chapter 10. Finally, we repeat an earlier observation: Fourier expansions (Chapter 19) and integral transforms (Chapter 20) can also have a role in the solution of PDEs, and applications of these techniques to PDEs are included in the appropriate chapters of this book. Additional Readings Bateman, H., Partial Differential Equations of Mathematical Physics. New York: Dover (1944), 1st ed. (1932). A wealth of applications of various partial differential equations in classical physics. Excellent examples of the use of different coordinate systems, including ellipsoidal, paraboloidal, toroidal coordinates, and so on. Cohen, H., Mathematics for Scientists and Engineers. Englewood Cliffs, NJ: Prentice-Hall (1992). Folland, G. B., Introduction to Partial Differential Equations, 2nd ed. Princeton, NJ: Princeton University Press (1995). Guckenheimer, J., P. Holmes, and F. John, Nonlinear Oscillations, Dynamical Systems and Bifurcations of Vector Fields, revised ed. New York: Springer-Verlag (1990). Gustafson, K. E., Partial Differential Equations and Hilbert Space Methods, 2nd ed., New York: Wiley (1987), reprinting Dover (1998). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry, 2nd ed. Princeton, NJ: Van Nostrand (1956). Chapter 5 covers curvilinear coordinates and 13 specific coordinate systems. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics . New York: McGraw-Hill (1953). Chapter 5 includes a description of several different coordinate systems. Note that Morse and Feshbach are not above using left-handed coordinate systems even for Cartesian coordinates. Elsewhere in this excellent (and difficult) book are many examples of the use of the various coordinate systems in solving physical problems. Chapter 6 discusses characteristics in detail. ArfKen_Ch10-9780123846549.tex CHAPTER 10 GREEN’S FUNCTIONS In contrast to the linear differential operators that have been our main concern when formulating problems as differential equations, we now turn to methods involving inte- gral operators, and in particular to those known as Green’s functions. Green’s-function methods enable the solution of a differential equation containing an inhomogeneous term (often called a source term) to be related to an integral operator containing the source. As a preliminary and elementary example, consider the problem of determining the potential .r/ generated by a charge distribution whose charge density is .r/. From the Poisson equation, we know that .r/ satisfies r2 .r/D1 "0.r/: (10.1) We also know, applying Coulomb’s law to the potential at r1produced by each element of charge.r2/d3r2, and assuming the space is empty except for the charge distribution, that .r 1/D1 4" 0Z d3r2.r2/ jr1r2j: (10.2) Here the integral is over the entire region where .r2/6D0. We can view the right-hand side of Eq. (10.2) as an integral operator that converts into , and identify the kernel (the function of two variables, one of which is to be integrated) as the Green’s function for this problem. Thus, we write G.r1;r2/D1 4"1 jr1r2j; (10.3) .r 1/DZ d3r2G.r1;r2/.r 2/; (10.4) assigning our Green’s function the symbol G(for “Green”). 447 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch10-9780123846549.tex 448 Chapter 10 Green’s Functions This example is preliminary because the response of more general problems to an inhomogeneous term will depend on the boundary conditions. For example, an electro- statics problem may include conductors whose surfaces will contain charge layers with magnitudes that depend on and which will also contribute to the potential at general r. It is elementary because the form of the Green’s function will also depend on the differen- tial equation to be solved, and often it will not be possible to obtain a Green’s function in a simple, closed form. The essential feature of any Green’s function is that it provides a way to describe the response of the differential-equation solution to an arbitrary source term (in the presence of the boundary conditions). In our present example, G.r1;r2/gives us the contribution to at the point r1produced by a point source of unit magnitude (a delta function) at the point r2. The fact that we can determine everywhere by an integration is a consequence of the fact that our differential equation is linear, so each element of the source contributes additively. In the more general context of a PDE that depends on both spatial and time coordinates, Green’s functions also appear as responses of the PDE solution to impulses at given positions and times. The aim of this chapter is to identify some general properties of Green’s functions, to survey methods for finding them, and to begin building connections between differential- operator and integral-operator methods for the description of physics problems. We start by considering problems in one dimension. 10.1 O NE-DIMENSIONAL PROBLEMS Let’s consider the second-order self-adjoint inhomogeneous ODE Lyd dx p.x/dy dx Cq.x/yDf.x/; (10.5) which is to be satisfied on the range axbsubject to homogeneous boundary condi- tions at xDaandxDbthat will cause Lto be Hermitian.1Our Green’s function for this problem needs to satisfy the boundary conditions and the ODE LG.x;t/D.xt/; (10.6) so that y.x/, the solution to Eq. (10.5) with its boundary conditions, can be obtained as y.x/DbZ aG.x;t/f.t/dt: (10.7) To verify Eq. (10.7), simply apply L: Ly.x/DbZ aLG.x;t/f.t/dtDbZ a.xt/f.t/dtDf.x/: 1Ahomogeneous boundary condition is one that continues to be satisfied if the function satisfying it is multiplied by a scale factor. Most of the more commonly encountered types of boundary conditions are homogeneous, e.g., yD0,y0D0, even c1yCc2y0D0. However, yDcwith ca nonzero constant is not homogeneous. ArfKen_Ch10-9780123846549.tex 10.1 One-Dimensional Problems 449 General Properties To gain an understanding of the properties G.x;t/must have, we first consider the result of integrating Eq. (10.6) over a small range of xthat includes xDt. We have tC"Z t"d dx p.x/dG.x;t/ dx dxCtC"Z t"q.x/G.x;t/dxDtC"Z t".tx/dx; which, carrying out some of the integrations, simplifies to p.x/dG.x;t/ dx tC" t"CtC"Z t"q.x/G.x;t/dxD1: (10.8) It is clear that Eq. (10.8) cannot be satisfied in the limit of small "ifG.x;t/and dG.x;t/=dx are both continuous (in x) atxDt, but we can satisfy that equation if we require G.x;t/to be continuous but accept a discontinuity in dG.x;t/=dx atxDt. In particular, continuity in Gwill cause the integral containing q.x/to vanish in the limit "!0, and we are left with the requirement lim "!0C" dG.x;t/ dx xDtC"dG.x;t/ dx xDt"# D1 p.t/: (10.9) Thus, the discontinuous impulse at xDtleads to a discontinuity in the xderivative of G.x;t/at that xvalue. Note, however, that because of the integration in Eq. (10.7), the singularity in dG=dx does not lead to a similar singularity in the overall solution y.x/in the usual case that f.x/is continuous. As a next step toward reaching understanding of the properties of Green’s functions, let’s expand G.x;t/in the eigenfunctions of our operator L, obtained subject to the boundary conditions already identified. Since Lis Hermitian, its eigenfunctions can be chosen to be orthonormal on .a;b/, with L'n.x/Dn'n.x/;h'nj'miDnm: (10.10) Expanding both the xand the tdependence of G.x;t/in this orthonormal set (using the complex conjugates of the 'nfor the texpansion), G.x;t/DX nmgnm'n.x/' m.t/: (10.11) We also expand .xt/in the same orthonormal set, according to Eq. (5.27): .xt/DX m'm.x/' m.t/: (10.12) Inserting both these expansions into Eq. (10.6), we have before any simplification LX nmgnm'n.x/' m.t/DX m'm.x/' m.t/: (10.13) ArfKen_Ch10-9780123846549.tex 450 Chapter 10 Green’s Functions Applying L, which operates only on 'n.x/, Eq. (10.13) reduces to X nmngnm'n.x/' m.t/DX m'm.x/' m.t/: Taking scalar products in the xandtdomains, we find that gnmDnm=n, soG.x;t/must have the expansion G.x;t/DX n' n.t/'n.x/ n: (10.14) The above analysis fails in the case that any nis zero, but we shall not pursue that special case further. The importance of Eq. (10.14) does not lie in its dubious value as a computational tool, but in the fact that it reveals the symmetry of G: G.x;t/DG.t;x/: (10.15) Form of Green’s Function The properties we have identified for Gare sufficient to enable its more complete identifi- cation, given a Hermitian operator Land its boundary conditions. We continue with the study of problems on an interval .a;b/with one homogeneous boundary condition at each endpoint of the interval. Given a value of t, it is necessary for xin the range ax<tthatG.x;t/have an x dependence y1.x/that is a solution to the homogeneous equation LD0and that also satis- fies the boundary condition at xDa. The most general G.x;t/satisfying these conditions must have the form G.x;t/Dy1.x/h1.t/; . x<t/; (10.16) where h1.t/is presently unknown. Conversely, in the range t<xb, it is necessary that G.x;t/have the form G.x;t/Dy2.x/h2.t/; . x>t/; (10.17) where y2is a solution of LD0that satisfies the boundary condition at xDb. The sym- metry condition, Eq. (10.15), permits Eqs. (10.16) and(10.17) to be consistent only if h 2DA y1andh 1DA y2, with Aa constant that is still to be determined. Assuming that y1andy2can be chosen to be real, we are led to the conclusion that G.x;t/D(A y1.x/y2.t/; x<t; A y2.x/y1.t/; x>t;(10.18) where LyiD0, with y1satisfying the boundary condition at xDaandy2satisfying that at xDb. The value of AinEq. (10.18) depends, of course, on the scale at which the yihave been specified, and must be set to a value that is consistent with Eq. (10.9). As applied here, that condition reduces to Ah y0 2.t/y1.t/y0 1.t/y2.t/i D1 p.t/; ArfKen_Ch10-9780123846549.tex 10.1 One-Dimensional Problems 451 equivalent to AD p.t/ Ty0 2.t/y1.t/y0 1.t/y2.t/1: (10.19) Despite its appearance, Adoes not depend on t. The expression involving the yiis their Wronskian, and it has a value proportional to 1=p.t/. See Exercise 7.6.11. It is instructive to verify that the form for G.x;t/given by Eq. (10.18) causes Eq. (10.7) to generate the desired solution to the ODE LyDf. To this end, we obtain an explicit form fory.x/: y.x/DA y2.x/xZ ay1.t/f.t/dtCA y1.x/bZ xy2.t/f.t/dt: (10.20) From Eq. (10.20) it is easy to verify that the boundary conditions on y.x/are satisfied; if xDathe first of the two integrals vanishes, and the second is proportional to y1; corre- sponding remarks apply at xDb. It remains to show that Eq. (10.20) yields LyDf. Differentiating with respect to x, we first have y0.x/DA y0 2.x/xZ ay1.t/f.t/dtCA y2.x/y1.x/f.x/ CA y0 1.x/bZ xy2.t/f.t/dtA y1.x/y2.x/f.x/ DA y0 2.x/xZ ay1.t/f.t/dtCA y0 1.x/bZ xy2.t/f.t/dt: (10.21) Proceeding to .py0/0: h p.x/y0.x/i0 DAh p.x/y0 2.x/i0xZ ay1.t/f.t/dtCAh p.x/y0 2.x/i y1.x/f.x/ CAh p.x/y0 1.x/i0bZ xy2.t/f.t/dtAh p.x/y0 1.x/i y2.x/f.x/:(10.22) Combining Eq. (10.22) andq.x/times Eq. (10.20), many terms drop because Ly1D Ly2D0, leaving Ly.x/DA p.x/h y0 2.x/y1.x/y0 1.x/y2.x/i f.x/Df.x/; (10.23) where the final simplification took place using Eq. (10.19). ArfKen_Ch10-9780123846549.tex 452 Chapter 10 Green’s Functions Example 10.1.1 SIMPLE SECOND-ORDER ODE Consider the ODE y00Df.x/; with boundary conditions y.0/Dy.1/D0. The corresponding homogeneous equation y00D0has general solution y0Dc0Cc1x; from these we construct the solution y1Dx that satisfies y1.0/D0and the solution y2D1x, satisfying y2.1/D0. For this ODE, the coefficient p.x/D1 ,y0 1.x/D1,y0 2.x/D1 , and the constant Ain the Green’s function is ADh .1/T.1/. x/.1/.1x/Ui1 D1: Our Green’s function is therefore G.x;t/D(x.1t/;0x<t; t.1x/;t<x1: Assuming we can perform the integral, we can now solve this ODE with boundary condi- tions for any function f.x/. For example, if f.x/Dsinx, our solution would be y.x/D1Z 0G.x;t/sint dtD.1x/xZ 0tsint dtCx1Z x.1t/sint dt D1 2sinx: The correctness of this result is easily checked. One advantage of the Green’s function formalism is that we do not need to repeat most of our work if we change the function f.x/. If we now take f.x/Dcosx, we get y.x/D1 2 2x1Ccosx : Note that our solution takes full account of the boundary conditions.  Other Boundary Conditions Occasionally one encounters problems other than the Hermitian second-order ODEs we have been considering. Some, but not always all of the Green’s-function properties we have identified, carry over to such problems. Consider first the possibility that we may have nonhomogeneous boundary conditions, such as the problem LyDfwith y.a/Dc1andy.b/Dc2, with one or both cinonzero. This problem can be converted into one with homogeneous boundary conditions by making a change of the dependent variable from yto uDyc1.bx/Cc2.xa/ ba: ArfKen_Ch10-9780123846549.tex 10.1 One-Dimensional Problems 453 In terms of u, the boundary conditions are homogeneous: u.a/Du.b/D0. A nonhomo- geneous condition on the derivative, e.g., y0.a/Dc, can be treated analogously. Another possibility for a second-order ODE is that we may have two boundary condi- tions at one endpoint and none at the other; this situation corresponds to an initial-value problem, and has lost the close connection to Sturm-Liouville eigenvalue problems. The result is that Green’s functions can still be constructed by invoking the condition of conti- nuity in G.x;t/atxDtand the prescribed discontinuity in @G=@x, but they will no longer be symmetric. Example 10.1.2 INITIAL VALUE PROBLEM Consider LyDd2y dx2CyDf.x/; (10.24) with the initial conditions y.0/D0andy0.0/D0. This operator Lhasp.x/D1. We start by noting that the homogeneous equation LyD0has the two linearly indepen- dent solutions y1Dsinxandy2Dcosx. However, the only linear combination of these solutions that satisfies the boundary condition at xD0is the trivial solution yD0, so our Green’s function for x<tcan only be G.x;t/D0. On the other hand, for the region x>t there are no boundary conditions to serve as constraints, and in that region we are free to write G.x;t/DC1.t/y1CC2.t/y2;orG.x;t/DC1.t/sinxCC2.t/cosx;x>t: We now impose the requirements G.t;t/DG.tC;t/! 0DC1.t/sintCC2.t/cost; @G @x.tC;t/@G @x.t;t/D1 p.t/D1! C1.t/costC2.t/sint.0/D1: These equations can now be solved, yielding C1.t/Dcost,C2.t/Dsint, so for x>t G.x;t/DcostsinxsintcosxDsin.xt/: Thus, the complete specification of G.x;t/is G.x;t/D(0; x<t; sin.xt/; x>t:(10.25) The lack of correspondence to a Sturm-Liouville problem is reflected in the lack of sym- metry of the Green’s function. Nevertheless, the Green’s function can be used to construct ArfKen_Ch10-9780123846549.tex 454 Chapter 10 Green’s Functions the solution to Eq. (10.24) subject to its initial conditions: y.x/D1Z 0G.x;t/f.t/dt DxZ 0sin.xt/f.t/dt: (10.26) Note that if we regard xas a time variable, our solution at “time” xis only influenced by source contributions from times tprior to x, so Eq. (10.24) obeys causality. We conclude this example by observing that we can verify that y.x/as given by Eq. (10.26) is the correct solution to our problem. Details are left as Exercise 10.1.3.  Example 10.1.3 BOUNDARY AT INFINITY Consider d2 dx2Ck2 .x/Dg.x/; (10.27) an equation essentially similar to one we have already studied several times, but now with boundary conditions that correspond (when multiplied by ei!t) to an outgoing wave. The general solution to Eq. (10.27) with gD0is spanned by the two functions y1Deikxand y2DeCikx: The outgoing wave boundary condition means that for large positive xwe must have the solution y2, while for large negative xthe solution must be y1. This information suffices to indicate that the Green’s function for this problem must have the form G.x;x0/D(Ay1.x0/y2.x/; x>x0; Ay2.x0/y1.x/; x<x0: We find the coefficient Afrom Eq. (10.19), in which p.x/D1: AD1 y0 2.x/y1.x/y0 1.x/y2.x/D1 ikCikDi 2k: Combining these results, we reach G.x;x0/Di 2kexp ijxx0j : (10.28) This result is yet another illustration that the Green’s function depends on boundary con- ditions as well as on the differential equation. Verification that this Green’s function yields the desired problem solution is the topic of Exercise 10.1.8.  ArfKen_Ch10-9780123846549.tex 10.1 One-Dimensional Problems 455 Relation to Integral Equations Consider now an eigenvalue equation of the form Ly.x/Dy.x/; (10.29) where we assume Lto be self-adjoint and subject to the boundary conditions y.a/D y.b/D0. We can proceed formally by treating Eq. (10.29) as an inhomogeneous equa- tion whose right-hand side is the particular function y.x/. To do so, we would first find the Green’s function G.x;t/for the operator Land the given boundary conditions, after which, as in Eq. (10.7), we could write y.x/DbZ aG.x;t/y.t/dt: (10.30) Equation (10.30) is not a solution to our eigenvalue problem, since the unknown function y.x/appears on both sides and, moreover, it does not tell us the possible values of the eigenvalue. What we have accomplished, however, is to convert our eigenvalue ODE and its boundary conditions into an integral equation which we can regard as an alternate starting point for solution of our eigenvalue problem. Our generation of Eq. (10.30) shows that it is implied by Eq. (10.29). If we can also show that we can connect these equations in the reverse order, namely that Eq. (10.30) implies Eq. (10.29), we can then conclude that they are equivalent formulations of the same eigenvalue problem. We proceed by applying LtoEq. (10.30), labeling it Lxto make clear that it is an operator on x, not t: Lxy.x/DLxbZ aG.x;t/y.t/dt DbZ aLxG.x;t/y.t/dtDbZ a.xt/y.t/dt Dy.x/: (10.31) The above analysis shows that under rather general circumstances we will be able to convert an eigenvalue equation based on an ODE into an entirely equivalent eigenvalue equation based on an integral equation. Note that to specify completely the ODE eigen- value equation we had to make an explicit identification of the accompanying boundary conditions, while the corresponding integral equation appears to be entirely self-contained. Of course, what has happened is that the effect of the boundary conditions has influenced the specification of the Green’s function that is the kernel of the integral equation. Conversion to an integral equation may be useful for two reasons, the more practical of which is that the integral equation may suggest different computational procedures for solution of our eigenvalue problem. There is also a fundamental mathematical reason why an integral-equation formulation may be preferred: It is that integral operators, such as that inEq. (10.30), are bounded operators (meaning that their application to a function yof ArfKen_Ch10-9780123846549.tex 456 Chapter 10 Green’s Functions finite norm produces a result whose norm is also finite). On the other hand, differential operators are unbounded; their application to a function of finite norm can produce a result of unbounded norm. Stronger theorems can be developed for operators that are bounded. We close by making the now obvious observation that Green’s functions provide the link between differential-operator and integral-operator formulations of the same problem. Example 10.1.4 DIFFERENTIAL VS. INTEGRAL FORMULATION Here we return to an eigenvalue problem we have already treated several times in various contexts, namely y00.x/Dy.x/; subject to boundary conditions y.0/Dy.1/D0. In Example 10.1.1 we found the Green’s function for this problem to be G.x;t/D(x.1t/;0x<t; t.1x/;t<x1; and, following Eq. (10.30), our eigenvalue problem can be rewritten as y.x/D1Z 0G.x;t/y.t/dt: (10.32) Methods for solution of integral equations will not be discussed until Chapter 21, but we can easily verify that the well-known solution set for this problem, yDsinnx;  nDn22;nD1;2;:::; also solves Eq. (10.32).  Exercises 10.1.1 Show that G.x;t/D(x;0x<t; t;t<x1; is the Green’s function for the operator LDd2=dx2and the boundary conditions y.0/D0,y0.1/D0. 10.1.2 Find the Green’s function for (a)Ly.x/Dd2y.x/ dx2Cy.x/;( y.0/D0; y0.1/D0: (b)Ly.x/Dd2y.x/ dx2y.x/;y.x/finite for1<x<1. ArfKen_Ch10-9780123846549.tex 10.1 One-Dimensional Problems 457 10.1.3 Show that the function y.x/defined by Eq. (10.26) satisfies the initial-value problem defined by Eq. (10.24) and its initial conditions y.0/Dy0.0/D0. 10.1.4 Find the Green’s function for the equation d2y dx2y 4Df.x/; with boundary conditions y.0/Dy./D0. ANS. G.x;t/D(2 sin. x=2/cos.t=2/; 0x<t; 2 cos. x=2/sin.t=2/; t<x: 10.1.5 Construct the Green’s function for x2d2y dx2Cxdy dxC.k2x21/yD0; subject to the boundary conditions y.0/D0,y.1/D0. 10.1.6 Given that LD.1x2/d2 dx22xd dx and that G.1; t/remains finite, show that no Green’s function can be constructed by the techniques of this section. Note. The solutions to LD0needed for the regions x<tandx>tare linearly depen- dent. 10.1.7 Find the Green’s function for d2 dt2Ckd dtDf.t/; subject to the initial conditions .0/D 0.0/D0, and solve this ODE for t>0given f.t/Dexp.t/. 10.1.8 Verify that the Green’s function G.x;x0/Di 2kexp ikjxx0j yields an outgoing wave solution to the ODE d2 dx2Ck2 .x/Dg.x/: Note. Compare with Example 10.1.3. 10.1.9 Construct the 1-D Green’s function for the modified Helmholtz equation, d2 dx2k2 .x/Df.x/: ArfKen_Ch10-9780123846549.tex 458 Chapter 10 Green’s Functions The boundary conditions are that the Green’s function must vanish for x!1 and x!1 . ANS. G.x1;x2/D1 2kexp kjx1x2j . 10.1.10 From the eigenfunction expansion of the Green’s function show that (a)2 21X nD1sinnxsinnt n2D(x.1t/;0x<t; t.1x/;t<x1: (b)2 21X nD0sin.nC1 2/xsin.nC1 2/t .nC1 2/2D(x;0x<t; t;t<x1: 10.1.11 Derive an integral equation corresponding to y00.x/y.x/D0; y.1/D1; y.1/D1; (a) by integrating twice. (b) by forming the Green’s function. ANS. y.x/D11Z 1K.x;t/y.t/dt, K.x;t/D(1 2.1x/.tC1/; x>t; 1 2.1t/.xC1/; x<t: 10.1.12 The general second-order linear ODE with constant coefficients is y00.x/Ca1y0.x/Ca2y.x/D0: Given the boundary conditions y.0/Dy.1/D0, integrate twice and develop the inte- gral equation y.x/D1Z 0K.x;t/y.t/dt; with K.x;t/D(a2t.1x/Ca1.x1/; t<x; a2x.1t/Ca1x; x<t: Note that K.x;t/is symmetric and continuous if a1D0. How is this related to self- adjointness of the ODE? 10.1.13 Transform the ODE d2y.r/ dr2k2y.r/CV0er ry.r/D0 ArfKen_Ch10-9780123846549.tex 10.2 Problems in Two and Three Dimensions 459 and the boundary conditions y.0/Dy.1/D0into an integral equation of the form y.r/DV01Z 0G.r;t/et ty.t/dt: The quantities V0andk2are constants. The ODE is derived from the Schrödinger wave equation with a mesonic potential: G.r;t/D8 >>< >>:1 kektsinhkr;0r<t; 1 kekrsinhkt;t<r<1: 10.2 P ROBLEMS IN TWO AND THREE DIMENSIONS Basic Features The principles, but unfortunately not all the details of our analysis of Green’s functions in one dimension, extend to problems of higher dimensionality. We summarize here proper- ties of general validity for the case where Lis a linear second-order differential operator in two or three dimensions. 1. A homogeneous PDE L .r 1/D0and its boundary conditions define a Green’s function G.r1;r2/, which is the solution of the PDE LG.r1;r2/D.r1r2/ subject to the relevant boundary conditions. 2. The inhomogeneous PDE L .r/Df.r/has, subject to the boundary conditions of Item 1, the solution .r 1/DZ G.r1;r2/f.r2/d3r2; where the integral is over the entire space relevant to the problem. 3. When Land its boundary conditions define the Hermitian eigenvalue problem L D with eigenfunctions 'n.r/and corresponding eigenvalues n, then G.r1;r2/is symmetric, in the sense that G.r1;r2/DG.r2;r1/;and G.r1;r2/has the eigenfunction expansion G.r1;r2/DX n' n.r2/'n.r1/ n: ArfKen_Ch10-9780123846549.tex 460 Chapter 10 Green’s Functions 4.G.r1;r2/will be continuous and differentiable at all points such that r16Dr2. We cannot even require continuity in a strict sense at r1Dr2(because our Green’s func- tion may become infinite there), but we can have the weaker condition that Gremain continuous in regions that surround, but do not include r1Dr2.Gmust have more serious singularities in its first derivatives, so that the second-order derivatives in L will generate the delta-function singularity characteristic of Gand specified in Item 1. What does not carry over from the 1-D case are the explicit formulas we used to con- struct Green’s functions for a variety of problems. Self-Adjoint Problems In more than one dimension, a second-order differential equation is self-adjoint if it has the form L .r/Drh p.r/r .r/i Cq.r/ .r/Df.r/; (10.33) with p.r/andq.r/real. This operator will define a Hermitian problem if its boundary conditions are such that h'jL iDhL'j i. See Exercise 10.2.2. Assuming we have a Hermitian problem, consider the scalar product D G.r;r1/ LG.r;r2/E DD LG.r;r1/ G.r;r2/E : (10.34) Here the scalar product and Lboth refer to the variable r, and the Hermitian property is responsible for this equality. The points r1andr2are arbitrary. Noting that LGresults in a delta function, we have, from the left-hand side of Eq. (10.34), D G.r;r1/ LG.r;r2/E DD G.r;r1/ .rr2/E DG.r2;r1/: (10.35) But, from the right-hand side of Eq. (10.34), D LG.r;r1/ G.r;r2/E DD .rr1/ G.r;r2/E DG.r1;r2/: (10.36) Substituting Eqs. (10.35) and(10.36) intoEq. (10.34), we recover the symmetry condition G.r1;r2/DG.r2;r1/. Eigenfunction Expansions We already saw, in 1-D Hermitian problems, that the Green’s function of a Hermitian problem can be written as an eigenfunction expansion. If L, with its boundary conditions, has normalized eigenfunctions 'n.r/and corresponding eigenvalues n, our expansion took the form G.r1;r2/DX n' n.r2/'n.r1/ n: (10.37) It turns out to be useful to consider the somewhat more general equation L .r 1/ .r 1/D.r2r1/; (10.38) ArfKen_Ch10-9780123846549.tex 10.2 Problems in Two and Three Dimensions 461 whereis a parameter (not an eigenvalue of L). In this more general case, an expansion in the'nyields for the Green’s function of the entire left-hand side of Eq. (10.38) the formula G.r1;r2/DX n' n.r2/'n.r1/ n: (10.39) Note that Eq. (10.39) will be well-defined only if the parameter is not equal to any of the eigenvalues of L. Form of Green’s Functions In spaces of more than one dimension, we cannot divide the region under consideration into two intervals, one on each side of a point (here designated r2), then choosing for each interval a solution to the homogeneous equation appropriate to its outer boundary. A more fruitful approach will often be to obtain a Green’s function for an operator L subject to some particularly convenient boundary conditions, with a subsequent plan to add to it whatever solution to the homogeneous equation L .r/D0that may be needed to adapt to the boundary conditions actually under consideration. This approach is clearly legitimate, as the addition of any solution to the homogeneous equation will not affect the (dis)continuity properties of the Green’s function. We consider first the Laplace operator in three dimensions, with the boundary condition thatGvanish at infinity. We therefore seek a solution to the inhomogeneous PDE r2 1G.r1;r2/D.r1r2/ (10.40) with limr1!1G.r1;r2/D0. We have added a subscript “1” to rto remind the reader that it operates on r1and not on r2. Since our boundary conditions are spherically symmetric and at an infinite distance from r1andr2, we may make the simplifying assumption that G.r1;r2/is a function only of r12Djr 1r2j. Our first step in processing Eq. (10.40) is to integrate it over a spherical volume of radius acentered at r2: Z r12<ar1r1G.r1;r2/d3r1D1; (10.41) where we have reduced the right-hand side using the properties of the delta function and written the left-hand side in a form making it ready for the application of Gauss’ theorem. We now apply that theorem to the left-hand side of Eq. (10.41), reaching Z r12Dar1G.r1;r2/d1D4a2dG dr12 r12DaD1: (10.42) Since Eq. (10.42) must be satisfied for all values of a, it is necessary that d dr12G.r1;r2/D1 4r2 12; ArfKen_Ch10-9780123846549.tex 462 Chapter 10 Green’s Functions which can be integrated to yield G.r1;r2/D1 41 jr1r2j: (10.43) We do not need to add a constant of integration because this form for Gvanishes at infinity. At this point it may be useful to note that the sign of G.r1;r2/depends on the sign asso- ciated with the differential operator of which it is a Green’s function. Some texts (including previous editions of this book) have defined Gas produced by a negative delta function so that Eq. (10.43) when associated with Cr2would not need a minus sign. There is, of course, no ambiguity in any physical results because a change in the sign of Gmust be accompanied by a change in the sign of the integral in which Gis combined with the inhomogeneous term of a differential equation. The Green’s function of Eq. (10.43) is only going to be appropriate for an infinite system with GD0at infinity but, as mentioned already, it can be converted into the Green’s func- tions of another problem by addition of a suitable solution to the homogeneous equation (in this case, Laplace’s equation). Since that is a reasonable starting point for a variety of problems, the form given in Eq. (10.43) is sometimes called the fundamental Green’s function of Laplace’s equation (in three dimensions). Let’s now repeat our analysis for the Laplace operator in two dimensions for a region of infinite extent, using circular coordinates D.;'/ . The integral in Eq. (10.41) is then over a circular area, and the 2-D analog of Eq. (10.42) becomes Z 12Dar1G.1;2/d1D2adG d12 12DaD1; leading to d d12G.1;2/D1 2 12; which has the indefinite integral G.1;2/D1 2lnj12j: (10.44) The form given in Eq. (10.44) becomes infinite at infinity, but it nevertheless can be regarded as a fundamental 2-D Green’s function. However, note that we will generally need to add to it a suitable solution to the 2-D Laplace equation to obtain the form needed for specific problems. The above analysis indicates that the Green’s function for the Laplace equation in 2-D space is rather different than the 3-D result. This observation illustrates the fact that there is a real difference between flatland (2-D) physics and actual (3-D) physics, even when the latter is applied to problems with translational symmetry in one direction. This is also a good time to note that the symmetry in the Green’s function corresponds to the notion that a source at r2produces a result (a potential) at r1that is the same as the potential at r2from a similar source at r1. This property will persist in more complicated problems so long as their definition makes them Hermitian. ArfKen_Ch10-9780123846549.tex 10.2 Problems in Two and Three Dimensions 463 Table 10.1 Fundamental Green’s Functionsa Laplace HelmholtzbModified r2r2Ck2Helmholtzc r2k2 1-D1 2jx1x2j i 2kexp.ikjx1x2j/1 2kexp.kjx1x2j/ 2-D1 2lnj12j i 4H.1/ 0.kj12j/1 2K0.kj12j/ 3-D1 41 jr1r2jexp.ikjr1r2j/ 4jr1r2jexp.kjr1r2j/ 4jr1r2j aBoundary conditions: For the Helmholtz equation, outgoing wave; for modified Helmholtz and 3-D Laplace equations, G!0at infinity; for 1-D and 2-D Laplace equation, arbitrary. bH1 0is a Hankel function, Section 14.4. cK0is a modified Bessel function, Section 14.5. Because they occur rather frequently, it is useful to have Green’s functions for the Helmholtz and modified Helmholtz equations in two and three dimensions (for one dimen- sion these Green’s functions were introduced in Example 10.1.3 andExercise 10.1.9). For the Helmholtz equation, a convenient fundamental form results if we take boundary con- ditions corresponding to an outgoing wave, meaning that the asymptotic rdependence must be of the form exp.C ikr/. For the modified Helmholtz equation, the most convenient boundary condition (for one, two, and three dimensions) is that Gdecay to zero in all direc- tions at large r. The one-, two-, and three-dimensional (3-D) fundamental Green’s functions for the Laplace, Helmholtz, and modified Helmholtz operators are listed in Table 10.1. We shall not derive here the forms of the Green’s functions for the Helmholtz equations; in fact, for two dimensions, they involve Bessel functions and are best treated in detail in a later chapter. However, for three dimensions, the Green’s functions are of relatively simple form, and the verification that they return correct results is the topic of Exercises 10.2.4 and10.2.6 . The fundamental Green’s function for the 1-D Laplace equation may not be instantly recognizable in comparison to the formulas we derived in Section 10.1, but con- sistency with our earlier analysis is the topic of Example 10.2.1 Sometimes it is useful to represent Green’s functions as expansions that take advantage of the specific properties of various coordinate systems. The so-called spherical Green’s function is the radial part of such an expansion in spherical polar coordinates. For the Laplace operator, it takes a form developed in Eqs. (16.65) and (16.66). We write it here only to show that it exhibits the two-region character that provides a convenient represen- tation of the discontinuity in the derivative: 1 41 jr1r2jD1X lD02lC1 4g.r1;r2/Pl.cos/; ArfKen_Ch10-9780123846549.tex 464 Chapter 10 Green’s Functions whereis the angle between r1andr2,Plis a Legendre polynomial, and the spherical Green’s function g.r1;r2/is gl.r1;r2/D8 >>>>< >>>>:1 2lC1rl 1 rlC1 2;r1<r2; 1 2lC1rl 2 rlC1 1;r1>r2: An explicit derivation of the formula for glis given in Example 16.3.2. In cylindrical coordinates .;'; z/one encounters an axial Green’s function gm.1;2/, in terms of which the fundamental Green’s function for the Laplace operator takes the form (also involving a continuous parameter k) G.r1;r2/D1 41 jr1r2j D1 221X mD1eim.'1'2/1Z 0gm.k1;k2/cosk.z1z2/dk: Here gm.k1;k2/DIm.k</Km.k>/; where<and>are, respectively, the smaller and larger of 1and2. The quantities Im andKmare modified Bessel functions, defined in Chapter 14. This expansion is discussed in more detail in Example 14.5.1. Again we note the two-region character. Example 10.2.1 ACCOMMODATING BOUNDARY CONDITIONS Let’s use the fundamental Green’s function of the 1-D Laplace equation, d2 .x/ dx2D0; namely G.x1;x2/D1 2jx1x2j; to illustrate how we can modify it to accommodate specific boundary conditions. We return to the oft-used example with Dirichlet conditions D0atxD0andxD1. The continu- ity of Gand the discontinuity in its derivative are unaffected if we add to the above Gone or more terms of the form f.x1/g.x2/, where fandgare solutions of the 1-D Laplace equation, i.e., any functions of the form axCb. For the boundary conditions we have specified, the Green’s function we require has the form G.x1;x2/D1 2.x1Cx2/Cx1x2C1 2jx1x2j: The continuous and differentiable terms we have added to the fundamental form bring us to the result G.x1;x2/D( 1 2.x1Cx2/Cx1x2C1 2.x2x1/Dx1.1x2/;x1<x2; 1 2.x1Cx2/Cx1x2C1 2.x1x2/Dx2.1x1/;x2<x1: This result is consistent with what we found in Example 10.1.1.  ArfKen_Ch10-9780123846549.tex 10.2 Problems in Two and Three Dimensions 465 Example 10.2.2 QUANTUM MECHANICAL SCATTERING: BORN APPROXIMATION The quantum theory of scattering provides a nice illustration of Green’s function tech- niques and the use of the Green’s function to obtain an integral equation. Our physical picture of scattering is as follows. A beam of particles moves along the negative z-axis toward the origin. A small fraction of the particles is scattered by the potential V.r/and goes off as an outgoing spherical wave. Our wave function .r/ must satisfy the time- independent Schrödinger equation Nh2 2mr2 .r/CV.r/ .r/DE .r/; (10.45) or r2 .r/Ck2 .r/D2m Nh2V.r/ .r/ ;k2D2mE Nh2: (10.46) From the physical picture just presented we look for a solution having the asymptotic form .r/eik0rCfk.;'/eikr r; (10.47) where eik0ris an incident plane wave2with the propagation vector k0carrying the sub- script 0to indicate that it is in the D0.z-axis) direction. The eikr=rterm describes an outgoing spherical wave with an angular and energy-dependent amplitude factor fk.;'/ ,3 and its 1=rradial dependence causes its asymptotic total flux to be independent of r. This is a consequence of the fact that the scattering potential V.r/becomes negligible at large r. Equation (10.45) contains nothing describing the internal structure or possible motion of the scattering center and therefore can only represent elastic scattering, so the propagation vector of the incoming wave, k0, must have the same magnitude, k, as the scattered wave. In quantum mechanics texts it is shown that the differential probability of scattering, called thescattering cross section, is given by jfk.;'j2. We now need to solve Eq. (10.46) to obtain .r/and the scattering cross section. Our approach starts by writing the solution in terms of the Green’s function for the operator on the left-hand side of Eq. (10.46), obtaining an integral equation because the inhomoge- neous term of that equation has the form .2m=Nh2/V.r/ .r/ : .r 1/DZ2m Nh2V.r2/ .r 2/G.r1;r2/d3r2: (10.48) We intend to take the Green’s function to be the fundamental form given for the Helmholtz equation in Table 10.1. We then recover the exp.ikr/=rpart of the desired asymptotic form, but the incident-wave term will be absent. We therefore modify our tentative for- mula, Eq. (10.48), by adding to its right-hand side the term exp.ik0r/, which is legiti- mate because this quantity is a solution to the homogeneous (Helmholtz) equation. That 2For simplicity we assume a continuous incident beam. In a more sophisticated and more realistic treatment, Eq. (10.47) would be one component of a wave packet. 3IfV.r/represents a central force, fkwill be a function of only, independent of the azimuthal angle '. ArfKen_Ch10-9780123846549.tex 466 Chapter 10 Green’s Functions approach leads us to .r 1/Deik0r1Z2m Nh2V.r2/ .r 2/eikjr1r2j 4jr1r2jd3r2: (10.49) This integral equation analog of the original Schrödinger wave equation is exact. It is called the Lippmann-Schwinger equation, and is an important starting point for studies of quantum-mechanical scattering phenomena. We will later study methods for solving integral equations such as that in Eq. (10.49). However, in the special case that the unscattered amplitude 0.r1/Deik0r1 (10.50) dominates the solution, it is a satisfactory approximation to replace .r 2/by 0.r2/within the integral, obtaining 1.r1/Deik0r1Z2m Nh2V.r2/eikjr1r2j 4jr1r2jeik0r2d3r2: (10.51) This is the famous Born approximation. It is expected to be most accurate for weak potentials and high incident energy.  Exercises 10.2.1 Show that the fundamental Green’s function for the 1-D Laplace equation, jx1x2j=2, is consistent with the form found in Example 10.1.1. 10.2.2 Show that if L .r/rh p.r/r .r/i Cq.r/ .r/; thenLis Hermitian for p.r/andq.r/real, assuming Dirichlet boundary conditions on the boundary of a region and that the scalar product is an integral over that region with unit weight. 10.2.3 Show that the termsCk2in the Helmholtz operator and k2in the modified Helmholtz operator do not affect the behavior of G.r1;r2/in the immediate vicinity of the singular point r1Dr2. Specifically, show that lim jr1r2j!0Z k2G.r1;r2/d3r2D1: 10.2.4 Show that exp.ikjr1r2j/ 4jr1r2j satisfies the appropriate criteria and therefore is a Green’s function for the Helmholtz equation. ArfKen_Ch10-9780123846549.tex Additional Readings 467 10.2.5 Find the Green’s function for the 3-D Helmholtz equation, Exercise 10.2.4, when the wave is a standing wave. 10.2.6 Verify that the formula given for the 3-D Green’s function of the modified Helmholtz equation in Table 10.1 is correct when the boundary conditions of the problem are that Gvanish at infinity. 10.2.7 An electrostatic potential (mks units) is '.r/DZ 4" 0ear r: Reconstruct the electrical charge distribution that will produce this potential. Note that '.r/vanishes exponentially for large r, showing that the net charge is zero. ANS..r/DZ.r/Za2 4ear r. Additional Readings Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969), reprinting, Dover (1992). This book contains nearly 100 pages on Green’s functions, starting with some good introductory material. Courant, R., and D. Hilbert, Methods of Mathematical Physics, Vol. 1 (English edition). New York: Interscience (1953). This is one of the classic works of mathematical physics. Originally published in German in 1924, the revised English edition is an excellent reference for a rigorous treatment of integral equations, Green’s functions, and a wide variety of other topics on mathematical physics. Jackson, J. D., Classical Electrodynamics, 3rd ed. New York: Wiley (1999). Contains applications to electro- magnetic theory. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics, 2 vols. New York: McGraw-Hill (1953). Chapter 7is a particularly detailed, complete discussion of Green’s functions from the point of view of mathematical physics. Note, however, that Morse and Feshbach frequently choose a source of 4.rr0/in place of our .rr0/. Considerable attention is devoted to bounded regions. Stakgold, I., Green’s Functions and Boundary Value Problems. New York: Wiley (1979). ArfKen_Ch11-9780123846549.tex CHAPTER 11 COMPLEX VARIABLE THEORY The imaginary numbers are a wonderful flight of God’s spirit; they are almost an amphibian between being and not being. GOTTFRIED WILHELM VON LEIBNIZ ,1702 We turn now to a study of complex variable theory. In this area we develop some of the most powerful and widely useful tools in all of analysis. To indicate, at least partly, why complex variables are important, we mention briefly several areas of application. 1. In two dimensions, the electric potential, viewed as a solution of Laplace’s equation, can be written as the real (or the imaginary) part of a complex-valued function, and this identification enables the use of various features of complex variable theory (specifi- cally, conformal mapping) to obtain formal solutions to a wide variety of electrostatics problems. 2. The time-dependent Schrödinger equation of quantum mechanics contains the imagi- nary unit i, and its solutions are complex. 3. In Chapter 9 we saw that the second-order differential equations of interest in physics may be solved by power series. The same power series may be used in the complex plane to replace xby the complex variable z. The dependence of the solution f.z/at a given z0on the behavior of f.z/elsewhere gives us greater insight into the behavior of our solution and a powerful tool (analytic continuation) for extending the region in which the solution is valid. 4. The change of a parameter kfrom real to imaginary, k!ik, transforms the Helmholtz equation into the time-independent diffusion equation. The same change connects the spherical and hyperbolic trigonometric functions, transforms Bessel functions into their modified counterparts, and provides similar connections between other super- ficially dissimilar functions. 469 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch11-9780123846549.tex 470 Chapter 11 Complex Variable Theory 5. Integrals in the complex plane have a wide variety of useful applications: Evaluating definite integrals and infinite series, Inverting power series, Forming infinite products, Obtaining solutions of differential equations for large values of the variable (asymptotic solutions), Investigating the stability of potentially oscillatory systems, Inverting integral transforms. 6. Many physical quantities that were originally real become complex as a simple physi- cal theory is made more general. The real index of refraction of light becomes a com- plex quantity when absorption is included. The real energy associated with an energy level becomes complex when the finite lifetime of the level is considered. 11.1 C OMPLEX VARIABLES AND FUNCTIONS We have already seen (in Chapter 1) the definition of complex numbers zDxCiyas ordered pairs of two real numbers, xandy. We reviewed there the rules for their arithmetic operations, identified the complex conjugate zof the complex number z, and discussed both the Cartesian and polar representations of complex numbers, introducing for that pur- pose the Argand diagram (complex plane). In the polar representation zDrei;we noted thatr(the magnitude of the complex number) is also called its modulus, and the angle is known as its argument. We proved that eisatisfies the important equation eiDcosCisin: (11.1) This equation shows that for real ,eiis of unit magnitude and is therefore situated on the unit circle, at an angle from the real axis. Our focus in the present chapter is on functions of a complex variable and on their analytical properties. We have already noted that by defining complex functions f.z/to have the same power-series expansion (in z) as the expansion (in x) of the correspond- ing real function f.x/, the real and complex definitions coincide when zis real. We also showed that by use of the polar representation, zDrei;it becomes clear how to com- pute powers and roots of complex quantities. In particular, we noted that roots, viewed as fractional powers, become multivalued functions in the complex domain, due to the fact thatexp.2ni/D1for all positive and negative integers n. We thus found z1=2to have two values (not a surprise, since for positive real x, we havepx). But we also noted thatz1=mwill have mdifferent complex values. We also noted that the logarithm becomes multivalued when extended to complex values, with lnzDln.rei/DlnrCi.C2n/; (11.2) with nany positive or negative integer (including zero). If necessary, the reader should review the topics mentioned above by rereading Section 1.8. ArfKen_Ch11-9780123846549.tex 11.2 Cauchy-Riemann Conditions 471 11.2 C AUCHY -RIEMANN CONDITIONS Having established complex functions of a complex variable, we now proceed to differen- tiate them. The derivative of f.z/, like that of a real function, is defined by lim z!0f.zCz/f.z/ .zCz/zDlim z!0f.z/ zDd f dzDf0.z/; (11.3) provided that the limit is independent of the particular approach to the point z. For real variables we require that the right-hand limit ( x!x0from above) and the left-hand limit (x!x0from below) be equal for the derivative d f.x/=dx to exist at xDx0. Now, with z (orz0) some point in a plane, our requirement that the limit be independent of the direction of approach is very restrictive. Consider increments xandyof the variables xandy, respectively. Then zDxCiy: (11.4) Also, writing fDuCiv, fDuCiv; (11.5) so that f zDuCiv xCiy: (11.6) Let us take the limit indicated by Eq. (11.3) by two different approaches, as shown in Fig. 11.1. First, with yD0, we letx!0.Equation (11.3) yields lim z!0f zDlim x!0u xCiv x D@u @xCi@v @x; (11.7) assuming that the partial derivatives exist. For a second approach, we set xD0and then lety!0. This leads to lim z!0f zDlim y!0 iu yCv y Di@u @yC@v @y: (11.8) If we are to have a derivative d f=dz, Eqs. (11.7) and(11.8) must be identical. Equating real parts to real parts and imaginary parts to imaginary parts (like components of vectors), we obtain @u @xD@v @y;@u @yD@v @x: (11.9) y xdx → 0z0 dy = 0 dx = 0 dy → 0 FIGURE 11.1 Alternate approaches to z0. ArfKen_Ch11-9780123846549.tex 472 Chapter 11 Complex Variable Theory These are the famous Cauchy-Riemann conditions. They were discovered by Cauchy and used extensively by Riemann in his development of complex variable theory. These Cauchy-Riemann conditions are necessary for the existence of a derivative of f.z/. That is, in order for d f=dzto exist, the Cauchy-Riemann conditions must hold. Conversely, if the Cauchy-Riemann conditions are satisfied and the partial derivatives ofu.x;y/andv.x;y/are continuous, the derivative d f=dzexists. To show this, we start by writing fD@u @xCi@v @x xC@u @yCi@v @y y; (11.10) where the justification for this expression depends on the continuity of the partial deriva- tives of uandv. Using the Cauchy-Riemann equations, Eq. (11.9), we convert Eq. (11.10) to the form fD@u @xCi@v @x xC @v @xCi@u @x y D@u @xCi@v @x .xCiy/: (11.11) ReplacingxCiybyzand bringing it to the left-hand side of Eq. (11.11), we reach f zD@u @xCi@v @x; (11.12) an equation whose right-hand side is independent of the direction of z(i.e., the relative values ofxandy). This independence of directionality meets the condition for the exis- tence of the derivative, d f=dz. Analytic Functions Iff.z/is differentiable and single-valued in a region of the complex plane, it is said to be an analytic function in that region.1Multivalued functions can also be analytic under certain restrictions that make them single-valued in specific regions; this case, which is of great importance, is taken up in detail in Section 11.6. If f.z/is analytic everywhere in the (finite) complex plane, we call it an entire function. Our theory of complex vari- ables here is one of analytic functions of a complex variable, which points up the crucial importance of the Cauchy-Riemann conditions. The concept of analyticity carried on in advanced theories of modern physics plays a crucial role in the dispersion theory (of ele- mentary particles). If f0.z/does not exist at zDz0, then z0is labeled a singular point; singular points and their implications will be discussed shortly. To illustrate the Cauchy-Riemann conditions, consider two very simple examples. 1Some writers use the term holomorphic orregular. ArfKen_Ch11-9780123846549.tex 11.2 Cauchy-Riemann Conditions 473 Example 11.2.1 z2ISANALYTIC Letf.z/Dz2. Multiplying out .xiy/.xiy/Dx2y2C2ixy, we identify the real part ofz2asu.x;y/Dx2y2and its imaginary part as v.x;y/D2xy. Following Eq. (11.9), @u @xD2xD@v @y;@u @yD2 yD@v @x: We see that f.z/Dz2satisfies the Cauchy-Riemann conditions throughout the complex plane. Since the partial derivatives are clearly continuous, we conclude that f.z/Dz2is analytic, and is an entire function.  Example 11.2.2 zISNOT ANALYTIC Letf.z/Dz, the complex conjugate of z. Now uDxandvDy. Applying the Cauchy- Riemann conditions, we obtain @u @xD16D@v @yD1: The Cauchy-Riemann conditions are not satisfied for any values of xoryandf.z/Dz is nowhere an analytic function of z. It is interesting to note that f.z/Dzis continu- ous, thus providing an example of a function that is everywhere continuous but nowhere differentiable in the complex plane.  The derivative of a real function of a real variable is essentially a local characteristic, in that it provides information about the function only in a local neighborhood, for instance, as a truncated Taylor expansion. The existence of a derivative of a function of a com- plex variable has much more far-reaching implications, one of which is that the real and imaginary parts of our analytic function must separately satisfy Laplace’s equation in two dimensions, namely @2 @x2C@2 @y2D0: To verify the above statement, we differentiate the first Cauchy-Riemann equation in Eq. (11.9) with respect to xand the second with respect to y, obtaining @2u @x2D@2v @x@y;@2u @y2D@2v @y@x: Combining these two equations, we easily reach @2u @x2C@2u @y2D0; (11.13) confirming that u.x;y/, the real part of a differentiable complex function, satisfies the Laplace equation. Either by recognizing that if f.z/is differentiable, so is i f.z/D v.x;y/iu.x;y/, or by steps similar to those leading to Eq. (11.13), we can confirm thatv.x;y/also satisfies the two-dimensional (2-D) Laplace equation. Sometimes uand vare referred to as harmonic functions (not to be confused with spherical harmonics, which we will later encounter as the angular solutions to central force problems). ArfKen_Ch11-9780123846549.tex 474 Chapter 11 Complex Variable Theory The solutions u.x;y/andv.x;y/are complementary in that the curves of constant u.x;y/make orthogonal intersections with the curves of constant v.x;y/. To confirm this, note that if.x0;y0/is on the curve u.x;y/Dc, then x0Cdx;y0Cdyis also on that curve if @u @xdxC@u @ydyD0; meaning that the slope of the curve of constant uat.x0;y0/is dy dx uD@u=@x @u=@y; (11.14) where the derivatives are to be evaluated at .x0;y0/. Similarly, we can find that the slope of the curve of constant vat.x0;y0/is dy dx vD@v=@ x @v=@ yD@u=@y @u=@x; (11.15) where the last member of Eq. (11.15) was reached using the Cauchy-Riemann equations. Comparing Eqs. (11.14) and (11.15), we note that at the same point, the slopes they describe are orthogonal (to check, verify that dxudxvCdyudyvD0). The properties we have just examined are important for the solution of 2-D electrostatics problems (governed by the Laplace equation). If we have identified (by methods outside the scope of the present text) an appropriate analytic function, its lines of constant uwill describe electrostatic equipotentials, while those of constant vwill be the stream lines of the electric field. Finally, the global nature of our analytic function is also illustrated by the fact that it has not only a first derivative, but in addition, derivatives of all higher orders, a property which is not shared by functions of a real variable. This property will be demonstrated in Section 11.4. Derivatives of Analytic Functions Working with the real and imaginary parts of an analytic function f.z/is one way to take its derivative; an example of that approach is to use Eq. (11.12). However, it is usually easier to use the fact that complex differentiation follows the same rules as those for real variables. As a first step in establishing this correspondence, note that, iff.z/is analytic, then, from Eq. (11.12), f0.z/D@f @x; and that h f.z/g.z/i0 Dd dzh f.z/g.z/i D@ @xh f.z/g.z/i D@f @x g.z/Cf.z/@g @x Df0.z/g.z/Cf.z/g0.z/; ArfKen_Ch11-9780123846549.tex 11.2 Cauchy-Riemann Conditions 475 the familiar rule for differentiating a product. Given also that dz dzD@z @xD1; we can easily establish that dz2 dzD2z;and, by induction,dzn dzDnzn1: Functions defined by power series will then have differentiation rules identical to those for the real domain. Functions not ordinarily defined by power series also have the same differentiation rules as for the real domain, but that will need to be demonstrated case by case. Here is an example that illustrates the establishment of a derivative formula. Example 11.2.3 DERIVATIVE OF LOGARITHM We want to verify that dlnz=dzD1=z. Writing, as in Eq. (1.138), lnzDlnrCiC2ni; we note that if we write lnzDuCiv, we have uDlnr,vDC2n. To check whether lnzsatisfies the Cauchy-Riemann equations, we evaluate @u @xD1 r@r @xDx r2;@u @yD1 r@r @yDy r2; @v @xD@ @xDy r2;@v @yD@ @yDx r2: The derivatives of randwith respect to xandyare obtained from the equations connect- ing Cartesian and polar coordinates. Except at rD0, where the derivatives are undefined, the Cauchy-Riemann equations can be confirmed. Then, to obtain the derivative, we can simply apply Eq. (11.12), dlnz dzD@u @xCi@v @xDxiy r2D1 xCiyD1 z: Because lnzis multivalued, it will not be analytic except under conditions restricting it to single-valuedness in a specific region. This topic will be taken up in Section 11.6.  Point at Infinity In complex variable theory, infinity is regarded as a single point, and behavior in its neigh- borhood is discussed after making a change of variable from ztowD1=z. This transfor- mation has the effect that, for example, zDR, with Rlarge, lies in the wplane close tozDCR, thereby among other things influencing the values computed for derivatives. An elementary consequence is that entire functions, such as zorez, have singular points atzD1 . As a trivial example, note that at infinity the behavior of zis identified as that of 1=w asw!0, leading to the conclusion that zis singular there. ArfKen_Ch11-9780123846549.tex 476 Chapter 11 Complex Variable Theory Exercises 11.2.1 Show whether or not the function f.z/D<. z/Dxis analytic. 11.2.2 Having shown that the real part u.x;y/and the imaginary part v.x;y/of an analytic functionw.z/each satisfy Laplace’s equation, show that neither u.x;y/norv.x;y/can have either a maximum or a minimum in the interior of any region in which w.z/is analytic. (They can have saddle points only.) 11.2.3 Find the analytic function w.z/Du.x;y/Civ.x;y/ (a) if u.x;y/Dx33xy2, (b) ifv.x;y/Deysinx. 11.2.4 If there is some common region in which w1Du.x;y/Civ.x;y/andw2Dw 1D u.x;y/iv.x;y/are both analytic, prove that u.x;y/andv.x;y/are constants. 11.2.5 Starting from f.z/D1=.xCiy/, show that 1=zis analytic in the entire finite zplane except at the point zD0. This extends our discussion of the analyticity of znto negative integer powers n. 11.2.6 Show that given the Cauchy-Riemann equations, the derivative f0.z/has the same value fordzDa dxCib dy (with neither anorbzero) as it has for dzDdx. 11.2.7 Using f.rei/DR.r;/ei2.r;/, in which R.r;/and2.r;/are differentiable real functions of rand, show that the Cauchy-Riemann conditions in polar coordinates become .a/@R @rDR r@2 @; .b/1 r@R @DR@2 @r: Hint. Set up the derivative first with zradial and then with ztangential. 11.2.8 As an extension of Exercise 11.2.7 show that2.r;/satisfies the 2-D Laplace equation in polar coordinates, @22 @r2C1 r@2 @rC1 r2@22 @2D0: 11.2.9 For each of the following functions f.z/, find f0.z/and identify the maximal region within which f.z/is analytic. (a) f.z/Dsinz z; (d) f.z/De1=z; (b) f.z/D1 z2C1;(e) f.z/Dz23zC2; (c) f.z/D1 z.zC1/;(f) f.z/Dtan.z/; (g) f.z/Dtanh. z/: ArfKen_Ch11-9780123846549.tex 11.3 Cauchy’s Integral Theorem 477 11.2.10 For what complex values do each of the following functions f.z/have a derivative? (a) f.z/Dz3=2; (b) f.z/Dz3=2; (c) f.z/Dtan1.z/; (d) f.z/Dtanh1.z/: 11.2.11 Two-dimensional irrotational fluid flow is conveniently described by a complex poten- tialf.z/Du.x;v/Civ.x;y/. We label the real part, u.x;y/, the velocity potential, and the imaginary part, v.x;y/, the stream function. The fluid velocity Vis given by VDru. Iff.z/is analytic: (a) Show that d f=dzDVxiVy. (b) Show that rVD0(no sources or sinks). (c) Show that rVD0(irrotational, nonturbulent flow). 11.2.12 The function f.z/is analytic. Show that the derivative of f.z/with respect to zdoes not exist unless f.z/is a constant. Hint. Use the chain rule and take xD.zCz/=2;yD.zz/=2i. Note. This result emphasizes that our analytic function f.z/is not just a complex func- tion of two real variables xandy. It is a function of the complex variable xCiy. 11.3 C AUCHY ’SINTEGRAL THEOREM Contour Integrals With differentiation under control, we turn to integration. The integral of a complex vari- able over a path in the complex plane (known as a contour) may be defined in close analogy to the (Riemann) integral of a real function integrated along the real x-axis. We divide the contour, from z0toz0 0, designated C, into nintervals by picking n1 intermediate points z1;z2;:::on the contour (Fig. 11.2). Consider the sum SnDnX jD1f.j/.zjzj1/; wherejis a point on the curve between zjandzj1. Now let n!1 with jzjzj1j!0 for all j. Iflimn!1Snexists, then limn!1nX jD1f.j/.zjzj1/Dz0 0Z z0f.z/dzDZ Cf.z/dz: (11.16) The right-hand side of Eq. (11.16) is called the contour integral of f.z/(along the specified contour Cfrom zDz0tozDz0 0). ArfKen_Ch11-9780123846549.tex 478 Chapter 11 Complex Variable Theory y xz3 z2 z1 z0ζ1 ζ0z¢0=zn FIGURE 11.2 Integration path. As an alternative to the above, the contour integral may be defined by z2Z z1f.z/dzDx2;y2Z x1;y1Tu.x;y/Civ.x;y/UTdxCi dyU Dx2;y2Z x1;y1Tu.x;y/dxv.x;y/dyUCix2;y2Z x1;y1Tv.x;y/dxCu.x;y/dyU; (11.17) with the path joining .x1;y1/and.x2;y2/specified. This reduces the complex integral to the complex sum of real integrals. It is somewhat analogous to the replacement of a vector integral by the vector sum of scalar integrals. Often we are interested in contours that are closed, meaning that the start and end of the contour are at the same point, so that the contour forms a closed loop. We normally define the region enclosed by a contour as that which lies to the left when the contour is traversed in the indicated direction; thus a contour intended to surround a finite area will normally be deemed to be traversed in the counterclockwise direction. If the origin of a polar coordinate system is within the contour, this convention will cause the normal direction of travel on the contour to be that in which the polar angle increases. Statement of Theorem Cauchy’s integral theorem states that: If f(z) is an analytic function at all points of a simply connected region in the complex plane and if C is a closed contour within that region, then I Cf.z/dzD0: (11.18) ArfKen_Ch11-9780123846549.tex 11.3 Cauchy’s Integral Theorem 479 To clarify the above, we need the following definition: A region is simply connected if every closed curve within it can be shrunk continu- ously to a point that is within the region. In everyday language, a simply connected region is one that has no holes. We also need to explain that the symbolH will be used from now on to indicate an integral over a closed contour; a subscript (such as C) is attached when further specification of the contour is desired. Note also that for the theorem to apply, the contour must be “within” the region of analyticity. That means it cannot be on the boundary of the region. Before proving Cauchy’s integral theorem, we look at some examples that do (and do not) meet its conditions. Example 11.3.1 znONCIRCULAR CONTOUR Let’s examine the contour integralH Czndz, where Cis a circle of radius r>0around the origin zD0in the positive mathematical sense (counterclockwise). In polar coordinates, cf. Eq. (1.125), we parameterize the circle as zDreianddzDireid. For n6D1; nan integer, we then obtain I CzndzDi rnC12Z 0expTi.nC1/Ud Di rnC1" ei.nC1/ i.nC1/#2 0D0 (11.19) because 2is a period of ei.nC1/. However, for nD1 I Cdz zDi2Z 0dD2i; (11.20) independent of rbut nonzero. The fact that Eq. (11.19) is satisfied for all integers n0is required by Cauchy’s the- orem, because for these nvalues znis analytic for all finite z, and certainly for all points within a circle of radius r. Cauchy’s theorem does not apply for any negative integer n because, for these n,znis singular at zD0. The theorem therefore does not prescribe any particular values for the integrals of negative n. We see that one such integral (that for nD1 ) has a nonzero value, and that others (for integral n6D1 ) do vanish.  Example 11.3.2 znONSQUARE CONTOUR We next examine the integration of znfor a different contour, a square with vertices at 1 21 2i. It is somewhat tedious to perform this integration for general integer n, so we illustrate only with nD2andnD1 . ArfKen_Ch11-9780123846549.tex 480 Chapter 11 Complex Variable Theory y x−1+i 21+i 2 1−i 2−1−i 2 FIGURE 11.3 Square integration contour. FornD2, we have z2Dx2y2C2ixy. Referring to Fig. 11.3, we identify the con- tour as consisting of four line segments. On Segment 1, dzDdx(yD1 2anddyD0); on Segment 2, dzDi dy,xD1 2,dxD0; on Segment 3, dzDdx,yD1 2,dyD0; and on Segment 4, dzDi dy,xD1 2,dxD0. Note that for Segments 3 and 4 the integration is in the direction of decreasing value of the integration variable. These segments therefore contribute as follows to the integral: Segment 1:1 2Z 1 2dx.x21 4ix/D1 31 8 1 8 1 4i 2.0/D1 6; Segment 2:1 2Z 1 2i dy.1 4y2Ciy/Di 4i 31 8 1 8 1 2.0/Di 6; Segment 3:1 2Z 1 2.dx/.x21 4Cix/D1 31 8 1 8 C1 4i 2.0/D1 6; Segment 4:1 2Z 1 2.i dy/.1 4y2iy/Di 4Ci 31 8 1 8 1 2.0/Di 6: We find that the integral of z2over the square vanishes, just as it did over the circle. This is required by Cauchy’s theorem. FornD1 , we have, in Cartesian coordinates, z1Dxiy x2Cy2; ArfKen_Ch11-9780123846549.tex 11.3 Cauchy’s Integral Theorem 481 and the integral over the four segments of the square contour takes the form 1 2Z 1 2xCi=2 x2C1 4dxC1 2Z 1 21 2iy y2C1 4.i dy/C1 2Z 1 2xi=2 x2C1 4dxC1 2Z 1 21 2Ciy y2C1 4.i dy/: Several of the terms vanish because they involve the integration of an odd integrand over an even interval, and others simply cancel. All that remains is Z z1dzDi1 2Z 1 2dx x2C1 4D2i1Z 1du u2C1D2ih 2  2i D2i; the same result as was obtained for the integration of z1around a circle of any radius. Cauchy’s theorem does not apply here, so the nonzero result is not problematic.  Cauchy’s Theorem: Proof We now proceed to a proof of Cauchy’s integral theorem. The proof we offer is subject to a restriction originally accepted by Cauchy but later shown unnecessary by Goursat. What we need to show is that I Cf.z/dzD0; subject to the requirement that Cis a closed contour within a simply connected region R where f.z/is analytic. See Fig. 11.4. The restriction needed for Cauchy’s (and the present) proof is that if we write f.z/Du.x;y/Civ.x;y/, the partial derivatives of uandvare continuous. y xCR FIGURE 11.4 A closed-contour Cwithin a simply connected region R. ArfKen_Ch11-9780123846549.tex 482 Chapter 11 Complex Variable Theory We intend to prove the theorem by direct application of Stokes’ theorem (Section 3.8). Writing dzDdxCi dy, I Cf.z/dzDI C.uCiv/.dxCi dy/ DI C.u dxvdy/CiI C.vdxCu dy/: (11.21) These two line integrals may be converted to surface integrals by Stokes’ theorem, a pro- cedure that is justified because we have assumed the partial derivatives to be continuous within the area enclosed by C. In applying Stokes’ theorem, note that the final two integrals ofEq. (11.21) are real. To proceed further, we note that all the integrals involved here can be identified as having integrands of the form .VxOexCVyOey/dr, the integration is around a loop in the xyplane, and the value of the integral will be the surface integral, over the enclosed area, of the zcomponent of r.VxOexCVyOey/. Thus, Stokes’ theorem says that I C.VxdxCVydy/DZ A@Vy @x@Vx @y dx dy; (11.22) with Abeing the 2-D region enclosed by C. For the first integral in the second line of Eq. (11.21), let uDVxandvDVy.2Then I C.u dxvdy/DI C.VxdxCVydy/ DZ A@Vy @x@Vx @y dx dyDZ A@v @xC@u @y dx dy: (11.23) For the second integral on the right side of Eq. (11.21) we let uDVyandvDVx. Using Stokes’ theorem again, we obtain I C.vdxCu dy/DZ A@u @x@v @y dx dy: (11.24) Inserting Eqs. (11.23) and (11.24) into Eq. (11.21), we now have I Cf.z/dzDZ A@v @xC@u @y dx dyCiZ A@u @x@v @y dx dyD0: (11.25) Remembering that f.z/has been assumed analytic, we find that both the surface integrals in Eq. (11.25) are zero because application of the Cauchy-Riemann equations causes their integrands to vanish. This establishes the theorem. 2For Stokes’ theorem, VxandVyare any two functions with continuous partial derivatives, and they need not be connected by any relations stemming from complex variable theory. ArfKen_Ch11-9780123846549.tex 11.3 Cauchy’s Integral Theorem 483 Multiply Connected Regions The original statement of Cauchy’s integral theorem demanded a simply connected region of analyticity. This restriction may be relaxed by the creation of a barrier, a narrow region we choose to exclude from the region identified as analytic. The purpose of the barrier construction is to permit, within a multiply connected region, the identification of curves that can be shrunk to a point within the region, that is, the construction of a subregion that is simply connected. Consider the multiply connected region of Fig. 11.5, in which f.z/is only analytic in the unshaded area labeled R. Cauchy’s integral theorem is not valid for the contour C, as shown, but we can construct a contour C0for which the theorem holds. We draw a barrier from the interior forbidden region, R0, to the forbidden region exterior to Rand then run a new contour, C0, as shown in Fig. 11.6. The new contour, C0;through ABDEFGA , never crosses the barrier that converts Rinto a simply connected region. Incidentally, the three-dimensional analog of this technique was used in Section 3.9 to prove Gauss’ law. Because f.z/is in fact continuous across the barrier dividing DEfrom G Aand the line segments DEandG Acan be arbitrarily close together, we have AZ Gf.z/dzDDZ Ef.z/dz: (11.26) y xCR R′ FIGURE 11.5 A closed contour Cin a multiply connected region. y xDA FB C′2 C′1 EG FIGURE 11.6 Conversion of a multiply connected region into a simply connected region. ArfKen_Ch11-9780123846549.tex 484 Chapter 11 Complex Variable Theory Then, invoking Cauchy’s integral theorem, because the contour is now within a simply connected region, and using Eq. (11.26) to cancel the contributions of the segments along the barrier, I C0f.z/dzDZ ABDf.z/dzCZ EFGf.z/dzD0: (11.27) Now that we have established Eq. (11.27), we note that AandDare only infinitesimally separated and that f.z/is actually continuous across the barrier. Hence, integration on the path ABD will yield the same result as a truly closed contour ABDA . Similar remarks apply to the path EFG , which can be replaced by EFGE . Renaming ABDA asC0 landEFGE as C0 2, we have the simple result I C0 1f.z/dzDI C0 2f.z/dz; (11.28) in which C0 1andC0 2are both traversed in the same (counterclockwise, that is, positive) direction. This result calls for some interpretation. What we have shown is that the integral of an analytic function over a closed contour surrounding an “island” of nonanalyticity can be subjected to any continuous deformation within the region of analyticity without changing the value of the integral. The notion of continuous deformation means that the change in contour must be able to be carried out via a series of small steps, which precludes processes whereby we “jump over” a point or region of nonanalyticity. Since we already know that the integral of an analytic function over a contour in a simply connected region of analyticity has the value zero, we can make the more general statement The integral of an analytic function over a closed path has a value that remains unchanged over all possible continuous deformations of the contour within the region of analyticity. Looking back at the two examples of this section, we see that the integrals of z2vanished for both the circular and square contours, as prescribed by Cauchy’s integral theorem for an analytic function. The integrals of z1did not vanish, and vanishing was not required because there was a point of nonanalyticity within the contours. However, the integrals of z1for the two contours had the same value, as either contour can be reached by continuous deformation of the other. We close this section with an extremely important observation. By a trivial extension to Example 11.3.1 plus the fact that closed contours in a region of analyticity can be deformed continuously without altering the value of the integral, we have the valuable and useful result: The integral of .zz0/naround any counterclockwise closed path Cthat encloses z0 has, for any integer n, the values I C.zz0/ndzD0; n6D1; 2i;nD1:(11.29) ArfKen_Ch11-9780123846549.tex 11.3 Cauchy’s Integral Theorem 485 Exercises 11.3.1 Show thatz2Z z1f.z/dzDz1Z z2f.z/dz. 11.3.2 Prove that Z Cf.z/dz jfjmaxL, wherejfjmaxis the maximum value of jf.z/jalong the contour CandLis the length of the contour. 11.3.3 Show that the integral 43iZ 3C4i.4z23iz/dz has the same value on the two paths: (a) the straight line connecting the integration limits, and (b) an arc on the circle jzjD5. 11.3.4 LetF.z/DzZ .1C i/cos 2 d: Show that F.z/is independent of the path connecting the limits of integration, and evaluate F.i/. 11.3.5 EvaluateH C.x2iy2/dz, where the integration is (a) clockwise around the unit circle, (b) on a square with vertices at 1i. Explain why the results of parts (a) and (b) are or are not identical. 11.3.6 Verify that 1CiZ 0zdz depends on the path by evaluating the integral for the two paths shown in Fig. 11.7. Recall that f.z/Dzis not an analytic function of zand that Cauchy’s integral theorem therefore does not apply. 11.3.7 Show that I Cdz z2CzD0; in which the contour Cis a circle defined by jzjDR>1. Hint. Direct use of the Cauchy integral theorem is illegal. The integral may be evaluated by expanding into partial fractions and then treating the two terms individually. This yields 0forR>1and2iforR<1. ArfKen_Ch11-9780123846549.tex 486 Chapter 11 Complex Variable Theory y x2 121(1,1) FIGURE 11.7 Contours for Exercise 11.3.6. 11.4 C AUCHY ’SINTEGRAL FORMULA As in the preceding section, we consider a function f.z/that is analytic on a closed contour Cand within the interior region bounded by C. This means that the contour Cis to be traversed in the counterclockwise direction. We seek to prove the following result, known asCauchy’s integral formula: 1 2iI Cf.z/ zz0dzDf.z0/; (11.30) in which z0is any point in the interior region bounded by C. Note that since zis on the contour Cwhile z0is in the interior, zz06D0and the integral Eq. (11.30) is well defined. Although f.z/is assumed analytic, the integrand is f.z/=.zz0/and is not analytic at zDz0unless f.z0/D0. We now deform the contour, to make it a circle of small radius rabout zDz0, traversed, like the original contour, in the counterclockwise direction. As shown in the preceding section, this does not change the value of the integral. We therefore write zDz0Crei, sodzDireid, the integration is from D0toD2, and I Cf.z/ zz0dzD2Z 0f.z0Crei/ reiireid: Taking the limit r!0, we obtain I Cf.z/ zz0dzDi f.z0/2Z 0dD2i f.z0/; (11.31) where we have replaced f.z/by its limit f.z0/because it is analytic and therefore contin- uous at zDz0. This proves the Cauchy integral formula. Here is a remarkable result. The value of an analytic function f.z/is given at an arbitrary interior point zDz0once the values on the boundary Care specified. ArfKen_Ch11-9780123846549.tex 11.4 Cauchy’s Integral Formula 487 It has been emphasized that z0is an interior point. What happens if z0is exterior to C? In this case the entire integrand is analytic on and within C. Cauchy’s integral theorem, Section 11.3, applies and the integral vanishes. Summarizing, we have 1 2iI Cf.z/dz zz0Df.z0/;z0within the contour, 0; z0exterior to the contour. Example 11.4.1 ANINTEGRAL Consider IDI Cdz z.zC2/; where the integration is counterclockwise over the unit circle. The factor 1=.zC2/is analytic within the region enclosed by the contour, so this is a case of Cauchy’s integral formula, Eq. (11.30), with f.z/D1=.zC2/andz0D0. The result is immediate: ID2i1 zC2 zD0Di:  Example 11.4.2 INTEGRAL WITH TWO SINGULAR FACTORS Consider now IDI Cdz 4z21; also integrated counterclockwise over the unit circle. The denominator factors into 4 z1 2 zC1 2 , and it is apparent that the region of integration contains two singular fac- tors. However, we may still use Cauchy’s integral formula if we make the partial fraction expansion 1 4z21D1 4 1 z1 21 zC1 2! ; after which we integrate the two terms individually. We have ID1 42 4I Cdz z1 2I Cdz zC1 23 5: Each integral is a case of Cauchy’s formula with f.z/D1, and for both integrals the point z0D1 2is within the contour, so each evaluates to 2i, and their sum is zero. So ID0.  ArfKen_Ch11-9780123846549.tex 488 Chapter 11 Complex Variable Theory Derivatives Cauchy’s integral formula may be used to obtain an expression for the derivative of f.z/. Differentiating Eq. (11.30) with respect to z0, and interchanging the differentiation and the zintegration,3 f0.z0/D1 2iIf.z/ .zz0/2dz: (11.32) Differentiating again, f00.z0/D2 2iIf.z/dz .zz0/3: Continuing, we get4 f.n/.z0/DnW 2iIf.z/dz .zz0/nC1I (11.33) that is, the requirement that f.z/be analytic guarantees not only a first derivative but derivatives of allorders as well! The derivatives of f.z/are automatically analytic. As indicated in a footnote, this statement assumes the Goursat version of the Cauchy integral theorem. This is a reason why Goursat’s contribution is so significant in the development of the theory of complex variables. Example 11.4.3 USE OF DERIVATIVE FORMULA Consider IDI Csin2z dz .za/4; where the integral is counterclockwise on a contour that encircles the point zDa. This is a case of Eq. (11.33) with nD3andf.z/Dsin2z. Therefore, ID2i 3Wd3 dz3sin2z zDaDi 3h 8 sin zcoszi zDaD8i 3sinacosa:  3The interchange can be proved legitimate, but the proof requires that Cauchy’s integral theorem not be subject to the continuous derivative restriction in Cauchy’s original proof. We are therefore now depending on Goursat’s proof of the integral theorem. 4This expression is a starting point for defining derivatives of fractional order. See A. Erdelyi, ed., Tables of Integral Trans- forms, Vol. 2. New York: McGraw-Hill (1954). For more recent applications to mathematical analysis, see T. J. Osler, An inte- gral analogue of Taylor’s series and its use in computing Fourier transforms, Math. Comput. 26: 449 (1972), and references therein. ArfKen_Ch11-9780123846549.tex 11.4 Cauchy’s Integral Formula 489 Morera’s Theorem A further application of Cauchy’s integral formula is in the proof of Morera’s theorem, which is the converse of Cauchy’s integral theorem. The theorem states the following: If a function f.z/is continuous in a simply connected region RandH Cf.z/dzD0for every closed contour Cwithin R, then f.z/is analytic throughout R. To prove the theorem, let us integrate f.z/from z1toz2. Since every closed-path inte- gral of f.z/vanishes, this integral is independent of path and depends only on its end- points. We may therefore write F.z2/F.z1/Dz2Z z1f.z/dz; (11.34) where F.z/, presently unknown, can be called the indefinite integral of f.z/. We then construct the identity F.z2/F.z1/ z2z1f.z1/D1 z2z1z2Z z1h f.t/f.z1/i dt; (11.35) where we have introduced another complex variable, t. Next, using the fact that f.t/is continuous, we write, keeping only terms to first order in tz1, f.t/f.z1/Df0.z1/.tz1/C; which implies that z2Z z1h f.t/f.z1/i dtDz2Z z1h f0.z1/.tz1/Ci dtDf0.z1/ 2.z2z1/2C: It is thus apparent that the right-hand side of Eq. (11.35) approaches zero in the limit z2!z1, so f.z1/Dlimz2!z1F.z2/F.z1/ z2z1DF0.z1/: (11.36) Equation (11.36) shows that F.z/, which by construction is single-valued, has a derivative at all points within Rand is therefore analytic in that region. Since F.z/is analytic, then so also must be its derivative, f.z/, thereby proving Morera’s theorem. At this point, one comment might be in order. Morera’s theorem, which establishes the analyticity of F.z/in a simply connected region, cannot be extended to prove that F.z/, as well as f.z/, is analytic throughout a multiply connected region via the device of introducing a barrier. It is not possible to show that F.z/will have the same value on both sides of the barrier, and in fact it does not always have that property. Thus, if extended to a multiply connected region, F.z/may fail to have the single-valuedness that is one of the requirements for analyticity. Put another way, a function which is analytic in a ArfKen_Ch11-9780123846549.tex 490 Chapter 11 Complex Variable Theory multiply connected region will have analytic derivatives of all orders in that region, but its integral is not guaranteed to be analytic in the entire multiply connected region. This issue is elaborated in Section 11.6. The proof of Morera’s theorem has given us something additional, namely that the indefinite integral of f.z/is its antiderivative, showing that: The rules for integration of complex functions are the same as those for real functions. Further Applications An important application of Cauchy’s integral formula is the following Cauchy inequal- ity. If f.z/DPanznis analytic and bounded, jf.z/j Mon a circle of radius rabout the origin, then janjrnM (Cauchy’s inequality) (11.37) gives upper bounds for the coefficients of its Taylor expansion. To prove Eq. (11.37) let us define M.r/DmaxjzjDrjf.z/jand use the Cauchy integral for anDf.n/.z/=nW, janjD1 2 I jzjDrf.z/ znC1dz M.r/2r 2rnC1: An immediate consequence of the inequality, Eq. (11.37), is Liouville’s theorem: If f.z/is analytic and bounded in the entire complex plane it is a constant. In fact, if jf.z/jMfor all z, then Cauchy’s inequality Eq. (11.37), applied for jzjDr, gives janjMrn. If now we choose to let rapproach1, we may conclude that for all n>0, janjD0. Hence f.z/Da0. Conversely, the slightest deviation of an analytic function from a constant value implies that there must be at least one singularity somewhere in the infinite complex plane. Apart from the trivial constant functions then, singularities are a fact of life, and we must learn to live with them. As pointed out when introducing the concept of the point at infinity, even innocuous functions such as f.z/Dzhave singularities at infinity; we now know that this is a property of every entire function that is not simply a constant. But we shall do more than just tolerate the existence of singularities. In the next section, we show how to expand a function in a Laurent series at a singularity, and we go on to use singularities to develop the powerful and useful calculus of residues in a later section of this chapter. A famous application of Liouville’s theorem yields the fundamental theorem of alge- bra(due to C. F. Gauss), which says that any polynomial P.z/DPn D0azwith n>0 andan6D0hasnroots. To prove this, suppose P.z/has no zero. Then 1=P.z/is analytic and bounded asjzj!1 , and, because of Liouville’s theorem, P.z/would have to be a constant. To resolve this contradiction, it must be the case that P.z/has at least one root  that we can divide out, forming P.z/=.z/, a polynomial of degree n1. We can repeat this process until the polynomial has been reduced to degree zero, thereby finding exactly nroots. ArfKen_Ch11-9780123846549.tex 11.4 Cauchy’s Integral Formula 491 Exercises Unless explicitly stated otherwise, closed contours occurring in these exercises are to be understood as traversed in the mathematically positive (counterclockwise) direction. 11.4.1 Show that 1 2iI zmn1dz;mandnintegers (with the contour encircling the origin once), is a representation of the Kronecker mn. 11.4.2 Evaluate I Cdz z21; where Cis the circlejz1jD1. 11.4.3 Assuming that f.z/is analytic on and within a closed contour Cand that the point z0 is within C, show that I Cf0.z/ zz0dzDI Cf.z/ .zz0/2dz: 11.4.4 You know that f.z/is analytic on and within a closed contour C. You suspect that the nth derivative f.n/.z0/is given by f.n/.z0/DnW 2iI Cf.z/ .zz0/nC1dz: Using mathematical induction (Section 1.4), prove that this expression is correct. 11.4.5 (a) A function f.z/is analytic within a closed contour C(and continuous on C). If f.z/6D0within Candjf.z/jMonC, show that jf.z/jM for all points within C. Hint. Consider w.z/D1=f.z/. (b) If f.z/D0within the contour C, show that the foregoing result does not hold and that it is possible to have jf.z/jD0at one or more points in the interior with jf.z/j>0over the entire bounding contour. Cite a specific example of an analytic function that behaves this way. 11.4.6 Evaluate I Ceiz z3dz; for the contour a square with sides of length a>1, centered at zD0. ArfKen_Ch11-9780123846549.tex 492 Chapter 11 Complex Variable Theory 11.4.7 Evaluate I Csin2zz2 .za/3dz; where the contour encircles the point zDa. 11.4.8 Evaluate I Cdz z.2zC1/; for the contour the unit circle. 11.4.9 Evaluate I Cf.z/ z.2zC1/2dz; for the contour the unit circle. Hint. Make a partial fraction expansion. 11.5 L AURENT EXPANSION Taylor Expansion The Cauchy integral formula of the preceding section opens up the way for another deriva- tion of Taylor’s series (Section 1.2), but this time for functions of a complex variable. Suppose we are trying to expand f.z/about zDz0and we have zDz1as the nearest point on the Argand diagram for which f.z/is not analytic. We construct a circle Ccen- tered at zDz0with radius less than jz1z0j(Fig. 11.8). Since z1was assumed to be the nearest point at which f.z/was not analytic, f.z/is necessarily analytic on and within C. From the Cauchy integral formula, Eq. (11.30), f.z/D1 2iI Cf.z0/dz0 z0z D1 2iI Cf.z0/dz0 .z0z0/.zz0/ D1 2iI Cf.z0/dz0 .z0z0/T1.zz0/=.z0z0/U: (11.38) Here z0is a point on the contour Candzis any point interior to C. It is not legal yet to expand the denominator of the integrand in Eq. (11.38) by the binomial theorem, for ArfKen_Ch11-9780123846549.tex 11.5 Laurent Expansion 493 z′Cz1 zz |z1−z0| |z′−z0|z0 FIGURE 11.8 Circular domains for Taylor expansion. we have not yet proved the binomial theorem for complex variables. Instead, we note the identity 1 1tD1CtCt2Ct3CD1X nD0tn; (11.39) which may easily be verified by multiplying both sides by 1t. The infinite series, fol- lowing the methods of Section 1.2, is convergent for jtj<1. Now, for a point zinterior to C;jzz0j<jz0z0j, and, using Eq. (11.39), Eq. (11.38) becomes f.z/D1 2iI C1X nD0.zz0/nf.z0/dz0 .z0z0/nC1: (11.40) Interchanging the order of integration and summation, which is valid because Eq. (11.39) is uniformly convergent for jtj<1", with 0<"< 1, we obtain f.z/D1 2i1X nD0.zz0/nI Cf.z0/dz0 .z0z0/nC1: (11.41) Referring to Eq. (11.33), we get f.z/D1X nD0f.n/.z0/ nW.zz0/n; (11.42) which is our desired Taylor expansion. It is important to note that our derivation not only produces the expansion given in Eq. (11.41); it also shows that this expansion converges when jzz0j<jz1z0j. For this reason the circle defined by jzz0jDjz1z0jis called the circle of convergence of our ArfKen_Ch11-9780123846549.tex 494 Chapter 11 Complex Variable Theory z R C2C1z¢(C1) z¢(C2)z0 rContour line FIGURE 11.9 Annular region for Laurent series. jz0z0jC1>jzz0jIjz0z0jC2<jzz0j. Taylor series. Alternatively, the distance jz1z0jis sometimes referred to as the radius of convergence of the Taylor series. In view of the earlier definition of z1, we can say that: The Taylor series of a function f.z/about any interior point z0of a region in which f.z/is analytic is a unique expansion that will have a radius of convergence equal to the distance from z0to the singularity of f.z/closest to z0, meaning that the Taylor series will converge within this circle of convergence. The Taylor series may or may not converge at individual points onthe circle of convergence. From the Taylor expansion for f.z/a binomial theorem may be derived. That task is left to Exercise 11.5.2. Laurent Series We frequently encounter functions that are analytic in an annular region, say, between circles of inner radius rand outer radius Rabout a point z0, as shown in Fig. 11.9. We assume f.z/to be such a function, with za typical point in the annular region. Draw- ing an imaginary barrier to convert our region into a simply connected region, we apply Cauchy’s integral formula to evaluate f.z/, using the contour shown in the figure. Note that the contour consists of the two circles centered at z0, labeled C1andC2(which can be considered closed since the barrier is fictitious), plus segments on either side of the barrier whose contributions will cancel. We assign C2andC1the radii r2andr1, respectively, where r<r2<r1<R. Then, from Cauchy’s integral formula, f.z/D1 2iI C1f.z0/dz0 z0z1 2iI C2f.z0/dz0 z0z: (11.43) Note that in Eq. (11.43)) an explicit minus sign has been introduced so that the contour C2(like C1) is to be traversed in the positive (counterclockwise) sense. The treatment of ArfKen_Ch11-9780123846549.tex 11.5 Laurent Expansion 495 Eq. (11.43) now proceeds exactly like that of Eq. (11.38) in the development of the Taylor series. Each denominator is written as .z0z0/.zz0/and expanded by the binomial theorem, which is now regarded as proven (see Exercise 11.5.2). Noting that for C1,jz0z0j>jzz0j, while for C2,jz0z0j<jzz0j, we find f.z/D1 2i1X nD0.zz0/nI C1f.z0/dz0 .z0z0/nC1C1 2i1X nD1.zz0/nI C2.z0z0/n1f.z0/dz0: (11.44) The minus sign of Eq. (11.43) has been absorbed by the binomial expansion. Labeling the first series S1and the second S2we have S1D1 2i1X nD0.zz0/nI C1f.z0/dz0 .z0z0/nC1; (11.45) which has the same form as the regular Taylor expansion, convergent for jzz0j<jz0 z0jDr1, that is, for all zinterior to the larger circle, C1. For the second series in Eq. (6.65) we have S2D1 2i1X nD1.zz0/nI C2.z0z0/n1f.z0/dz0; (11.46) convergent forjzz0j>jz0z0jDr2, that is, for all zexterior to the smaller circle, C2. Remember, C2now goes counterclockwise. These two series are combined into one series,5known as a Laurent series, of the form f.z/D1X nD1an.zz0/n; (11.47) where anD1 2iI Cf.z0/dz0 .z0z0/nC1: (11.48) Since convergence of a binomial expansion is not relevant to the evaluation of Eq. (11.48), Cin that equation may be any contour within the annular region r<jzz0j<Rthat encircles z0once in a counterclockwise sense. If such an annular region of analyticity does exist, then Eq. (11.47) is the Laurent series, or Laurent expansion, of f.z/. The Laurent series differs from the Taylor series by the obvious feature of negative powers of.zz0/. For this reason the Laurent series will always diverge at least at zDz0 and perhaps as far out as some distance r. In addition, note that Laurent series coefficients need not come from evaluation of contour integrals (which may be very intractable). Other techniques, such as ordinary series expansions, may provide the coefficients. Numerous examples of Laurent series appear later in this book. We limit ourselves here to one simple example to illustrate the application of Eq. (11.47). 5Replace nbyninS2and add. ArfKen_Ch11-9780123846549.tex 496 Chapter 11 Complex Variable Theory Example 11.5.1 LAURENT EXPANSION Letf.z/DTz.z1/U1. If we choose to make the Laurent expansion about z0D0, then r>0andR<1. These limitations arise because f.z/diverges both at zD0andzD1. A partial fraction expansion, followed by the binomial expansion of .1z/1, yields the Laurent series 1 z.z1/D1 1z1 zD1 z1zz2z3D1X nD1zn: (11.49) From Eqs. (11.49), (11.47), and (11.48), we then have anD1 2iIdz0 .z0/nC2.z01/D1 forn1; 0forn<1;(11.50) where the contour for Eq. (11.50) is counterclockwise in the annular region between z0D0 andjz0jD1. The integrals in Eq. (11.50) can also be directly evaluated by insertion of the geometric- series expansion of .1z0/1: anD1 2iI1X mD0.z0/mdz0 .z0/nC2: (11.51) Upon interchanging the order of summation and integration (permitted because the series is uniformly convergent), we have anD1 2i1X mD0I .z0/mn2dz0: (11.52) The integral in Eq. (11.52) (including the initial factor 1=2 i, but not the minus sign) was shown in Exercise 11.4.1 to be an integral representation of the Kronecker delta, and is therefore equal to m;nC1. The expression for anthen reduces to anD1X mD0m;nC1D1; n1; 0;n<1; in agreement with Eq. (11.50).  Exercises 11.5.1 Develop the Taylor expansion of ln.1Cz/. ANS.1X nD1.1/n1zn n. ArfKen_Ch11-9780123846549.tex 11.6 Singularities 497 11.5.2 Derive the binomial expansion .1Cz/mD1CmzCm.m1/ 12z2CD1X nD0m n zn form, any real number. The expansion is convergent for jzj<1. Why? 11.5.3 A function f.z/is analytic on and within the unit circle. Also, jf.z/j<1forjzj1 andf.0/D0. Show thatjf.z/j<jzjforjzj1. Hint. One approach is to show that f.z/=zis analytic and then to express Tf.z0/=z0Un by the Cauchy integral formula. Finally, consider absolute magnitudes and take the nth root. This exercise is sometimes called Schwarz’s theorem. 11.5.4 Iff.z/is a real function of the complex variable zDxCiy, that is, f.x/Df.x/, and the Laurent expansion about the origin, f.z/DPanzn, has anD0forn<N, show that all of the coefficients anare real. Hint. Show that zNf.z/is analytic (via Morera’s theorem, Section 11.4). 11.5.5 Prove that the Laurent expansion of a given function about a given point is unique; that is, if f.z/D1X nDNan.zz0/nD1X nDNbn.zz0/n; show that anDbnfor all n. Hint. Use the Cauchy integral formula. 11.5.6 Obtain the Laurent expansion of ez=z2about zD0. 11.5.7 Obtain the Laurent expansion of zez=.z1/about zD1. 11.5.8 Obtain the Laurent expansion of .z1/e1=zabout zD0. 11.6 S INGULARITIES Poles We define a point z0as an isolated singular point of the function f.z/iff.z/is not analytic at zDz0but is analytic at all neighboring points. There will therefore be a Laurent expansion about an isolated singular point, and one of the following statements will be true: 1. The most negative power of zz0in the Laurent expansion of f.z/about zDz0will be some finite power, .zz0/n, where nis an integer, or 2. The Laurent expansion of f.z/about zz0will continue to negatively infinite powers ofzz0. ArfKen_Ch11-9780123846549.tex 498 Chapter 11 Complex Variable Theory In the first case, the singularity is called a pole, and is more specifically identified as a pole of order n. A pole of order 1 is also called a simple pole. The second case is not referred to as a “pole of infinite order,” but is called an essential singularity. One way to identify a pole of f.z/without having available its Laurent expansion is to examine limz!z0.zz0/nf.z0/ for various integers n. The smallest integer nfor which this limit exists (i.e., is finite) gives the order of the pole at zDz0. This rule follows directly from the form of the Laurent expansion. Essential singularities are often identified directly from their Laurent expansions. For example, e1=zD1C1 zC1 2W1 z2 C D1X nD01 nW1 zn clearly has an essential singularity at zD0. Essential singularities have many pathologi- cal features. For instance, we can show that in any small neighborhood of an essential singularity of f.z/the function f.z/comes arbitrarily close to any (and therefore every) preselected complex quantity w0.6Here, the entire w-plane is mapped by finto the neigh- borhood of the point z0. The behavior of f.z/asz!1 is defined in terms of the behavior of f.1=t/ast!0. Consider the function sinzD1X nD0.1/nz2nC1 .2nC1/W: (11.53) Asz!1 , we replace the zby1=tto obtain sin1 t D1X nD0.1/n .2nC1/Wt2nC1: (11.54) It is clear that sin.1= t/has an essential singularity at tD0, from which we conclude that sinzhas an essential singularity at zD1 . Note that although the absolute value of sinx for all real xis equal to or less than unity, the absolute value of siniyDisinhyincreases exponentially without limit as yincreases. A function that is analytic throughout the finite complex plane except for isolated poles is called meromorphic. Examples are ratios of two polynomials, also tanzandcotz. As previously mentioned, functions that have no singularities in the finite complex plane are called entire functions. Examples are expz,sinz, and cosz. 6This theorem is due to Picard. A proof is given by E. C. Titchmarsh, The Theory of Functions, 2nd ed. New York: Oxford University Press (1939). ArfKen_Ch11-9780123846549.tex 11.6 Singularities 499 Branch Points In addition to the isolated singularities identified as poles or essential singularities, there are singularities uniquely associated with multivalued functions. It is useful to work with these functions in ways that to the maximum possible extent remove ambiguity as to the function values. Thus, if at a point z0(at which f.z/has a derivative) we have chosen a specific value of the multivalued function f.z/, then we can assign to f.z/values at nearby points in a way that causes continuity in f.z/. If we think of a succession of closely spaced points as in the limit of zero spacing defining a path, our current observation is that a given value of f.z0/then leads to a unique definition of the value of f.z/to be assigned to each point on the path. This scheme creates no ambiguity so long as the path is entirely open, meaning that the path does not return to any point previously passed. But if the path returns to z0, thereby forming a closed loop, our prescription might lead, upon the return, to a different one of the multiple values of f.z0/. Example 11.6.1 VALUE OF z1=2ON A CLOSED LOOP We consider f.z/Dz1=2on the path consisting of counterclockwise passage around the unit circle, starting and ending at zDC1 . At the start point, where z1=2has the multiple valuesC1and1, let us choose f.z/DC1 . See Fig. 11.10. Writing f.z/Dei=2, we note that this form (with D0) is consistent with the desired starting value of f.z/,C1. In the figure, the start point is labeled A. Next, we note that passage counterclockwise on the unit circle corresponds to an increase in , so that at the points marked B, C, and D in the figure, the respective values of are=2,, and 3=2 . Note that because of the path we have decided to take, we cannot assign to point C the valueor to point D the  value=2 . Continuing further along the path, when we return to point A the value of  has become 2(not zero). Now that we have identified the behavior of , let’s examine what happens to f.z/. At the points B, C, and D, we have f.zB/DeiB=2Dei=4D1Cip 2; f.zC/Dei=2DCi; f.zD/De3i=4D1Cip 2: y B CA Dxθ FIGURE 11.10 Path encircling zD0for evaluation of z1=2. ArfKen_Ch11-9780123846549.tex 500 Chapter 11 Complex Variable Theory y B CA Dx θ FIGURE 11.11 Path not encircling zD0for evaluation of z1=2. When we return to point A, we have f.C1/DeiD1 , which is the other value of the multivalued function z1=2. If we continue for a second counterclockwise circuit of the unit circle, the value of  would continue to increase, from 2to4(reached when we arrive at point A after the second loop). We now have f.C1/De.4i/=2De2iD1, so a second circuit has brought us back to the original value. It should now be clear that we are only going to be able to obtain two different values of z1=2for the same point z.  Example 11.6.2 ANOTHER CLOSED LOOP Let’s now see what happens to the function z1=2as we pass counterclockwise around a circle of unit radius centered at zDC2 , starting and ending at zDC3 . See Fig. 11.11. AtzD3, the values of f.z/areCp 3andp 3; let’s start with f.zA/DCp 3. As we move from point A through point B to point C, note from the figure that the value of  first increases (actually, to 30) and then decreases again to zero; further passage from C to D and back to A causes first to decrease (to30) and then to return to zero at A. So in this example the closed loop does not bring us to a different value of the multivalued function z1=2.  The essential difference between these two examples is that in the first, the path encircled zD0; in the second it did not. What is special about zD0is that (from a complex-variable viewpoint) it is singular; the function z1=2does not have a derivative there. The lack of a well-defined derivative means that ambiguity in the function value will result from paths that circle such a singular point, which we call a branch point. The order of a branch point is defined as the number of paths around it that must be taken before the function involved returns to its original value; in the case of z1=2, we saw that the branch point at zD0is of order 2. We are now ready to see what must be done to cause a multivalued function to be restricted to single-valuedness on a portion of the complex plane. We simply need to pre- vent its evaluation on paths that encircle a branch point. We do so by drawing a line (known as abranch line, or more commonly, a branch cut) that the evaluation path cannot cross; the branch cut must start from our branch point and continue to infinity (or if consistent with maintaining single-valuedness) to another finite branch point. The precise path of a branch cut can be chosen freely; what must be chosen appropriately are its endpoints. Once appropriate branch cut(s) have been drawn, the originally multivalued function has been restricted to being single-valued in the region bounded by the branch cut(s); we call the function as made single-valued in this way a branch of our original function. Since we ArfKen_Ch11-9780123846549.tex 11.6 Singularities 501 could construct such a branch starting from any one of the values of the original function at a single arbitrary point in our region, we identify our multivalued function as having multiple branches. In the case of z1=2, which is double-valued, the number of branches is two. Note that a function with a branch point and a corresponding branch cut will not be continuous across the cut line. Hence line integrals in opposite directions on the two sides of the branch cut will not generally cancel each other. Branch cuts, therefore, are real boundaries to a region of analyticity, in contrast to the artificial barriers we introduced in extending Cauchy’s integral theorem to multiply connected regions. While from a fundamental viewpoint all branches of a multivalued function f.z/are equally legitimate, it is often convenient to agree on the branch to be used, and such a branch is sometimes called the principal branch, with the value of f.z/on that branch called its principal value. It is common to take the branch of z1=2which is positive for real, positive zas its principal branch. An observation that is important for complex analysis is that by drawing appropriate branch cut(s), we have restricted a multivalued function to single-valuedness, so that it can be an analytic function within the region bounded by the branch cut(s), and we can therefore apply Cauchy’s two theorems to contour integrals within the region of analyticity. Example 11.6.3 lnzHAS AN INFINITE NUMBER OF BRANCHES Here we examine the singularity structure of lnz. As we already saw in Eq. (1.138), the logarithm is multivalued, with the polar representation lnzDln rei.C2n/ DlnrCi.C2n/; (11.55) where ncan have anypositive or negative integer value. Noting that lnzis singular at zD0(it has no derivative there), we now identify zD0 as a branch point. Let’s consider what happens if we encircle it by a counterclockwise path on a circle of radius r, starting from the initial value lnr, atzDrDreiwithD0. Every passage around the circle will add 2to, and after ncomplete circuits the value we have for lnzwill be lnrC2ni. The branch point of lnzatzD0is of infinite order, corresponding to the infinite number of its multiple values. (By encircling zD0repeatedly in the clockwise direction, we can also reach all negative integer values of n.) We can make lnzsingle-valued by drawing a branch cut from zD0tozD1 inany way (though there is ordinarily no reason to use cuts that are not straight lines). It is typical to identify the branch with nD0as the principal branch of the logarithm. Incidentally, we note that the inverse trigonometric functions, which can be written in terms of logarithms, as in Eq. (1.137), will also be infinitely multivalued, with principal values that are usually chosen on a branch that will yield real values for real z. Compare with the usual choices of the values assigned the real-variable forms of sin1xDarcsin x, etc.  Using the logarithm, we are now in a position to look at the singularity structures of expressions of the form zp, where both zandpmay be complex. To do so, we write zDelnz;sozpDeplnz; (11.56) which is single-valued if pis an integer, t-valued if pis a real rational fraction (in lowest terms) of the form s=t, and infinitely multivalued otherwise. ArfKen_Ch11-9780123846549.tex 502 Chapter 11 Complex Variable Theory Example 11.6.4 MULTIPLE BRANCH POINTS Consider the function f.z/D.z21/1=2D.zC1/1=2.z1/1=2: The first factor on the right-hand side, .zC1/1=2, has a branch point at zD1 . The second factor has a branch point at zDC1 . At infinity f.z/has a simple pole. This is best seen by substituting zD1=tand making a binomial expansion at tD0: .z21/1=2D1 t.1t2/1=2D1 t1X nD01=2 n .1/nt2nD1 t1 2t1 8t3C: We want to make f.z/single-valued by making appropriate branch cut(s). There are many ways to accomplish this, but one we wish to investigate is the possibility of making a branch cut from zD1 tozDC1 , as shown in Fig. 11.12. To determine whether this branch cut makes our f.z/single-valued, we need to see what happens to each of the multivalent factors in f.z/as we move around on its Argand dia- gram. Figure 11.12 also identifies the quantities that are relevant for this purpose, namely those that relate a point Pto the branch points. In particular, we have written the position relative to the branch point at zD1asz1Dei', with the position relative to zD1 denoted zC1Drei. With these definitions, we have f.z/Dr1=21=2e.C'/=2: Our mission is to note how 'andchange as we move along the path, so that we can use the correct value of each for evaluating f.z/. We consider a closed path starting at point AinFig. 11.13, proceeding via points B through F, then back to A. At the start point, we choose D'D0, thereby causing the multivalued f.zA/to have the specific value Cp 3. As we pass above zDC1 on the way to point B,remains essentially zero, but 'increases from zero to . These angles do not change as we pass from BtoC, but on going to point D,increases to, and then, passing below zD1 on the way to point E, it further increases to 2(not zero!). Meanwhile, ' remains essentially at . Finally, returning to point Abelow zDC1 ,'increases to 2, so that upon the return to point Aboth'andhave become 2. The behavior of these angles and the values of .C'/=2 (the argument of f.z/) are tabulated in Table 11.1. yP pt x θϕ −1 +1 FIGURE 11.12 Possible branch cut for Example 11.6.4 and the quantities relating a point Pto the branch points. ArfKen_Ch11-9780123846549.tex 11.6 Singularities 503 A DB C F E FIGURE 11.13 Path around the branch cut in Example 11.6.4. Table 11.1 Phase Angles, Path in Fig. 11.13 Point  ' . C'/=2 A 0 0 0 B 0  =2 C 0  =2 D    E 2  /2 F 2  3/2 A 2 2 2 Two features emerge from this analysis: 1. The phase of f.z/at points BandCis not the same as that at points EandF. This behavior can be expected at a branch cut. 2. The phase of f.z/at point A0(the return to A) exceeds that at point Aby2, meaning that the function f.z/D.z21/1=2issingle-valued for the contour shown, encircling both branch points. What actually happened is that each of the two multivalued factors contributed a sign change upon passage around the closed loop, so the two factors together restored the origi- nal sign of f.z/. Another way we could have made f.z/single-valued would have been to make a sepa- rate branch cut from each branch point to infinity; a reasonable way to do this would be to make cuts on the real axis for all x>1and for all x<1. This alternative is explored in Exercises 11.6.2 and11.6.4.  Analytic Continuation We saw in Section 11.5 that a function f.z/which is analytic within a region can be uniquely expanded in a Taylor series about any interior point z0of the region of analyti- city, and that the resulting expansion will be convergent within a circle of convergence extending to the singularity of f.z/closest to z0. Since ArfKen_Ch11-9780123846549.tex 504 Chapter 11 Complex Variable Theory The coefficients in the Taylor series are proportional to the derivatives of f.z/, An analytic function has derivatives of all orders that are independent of direction, and therefore The values of f.z/on a single finite line segment with z0as a interior point will suffice to determine all derivatives of f.z/atzDz0, we conclude that if two apparently different analytic functions (e.g., a closed expression vs. an integral representation or a power series) have values that coincide on a range as restricted as a single finite line segment, then they are actually the same function within the region where both functional forms are defined. The above conclusion will provide us with a technique for extending the definition of an analytic function beyond the range of any particular functional form initially used to define it. All we will need to do is to find another functional form whose range of definition is not entirely included in that of the initial form and which yields the same function values on at least a finite line segment within the area where both functional forms are defined. To make the approach more concrete, consider the situation illustrated in Fig. 11.14, where a function f.z/is defined by its Taylor expansion about a point z0with a circle of convergence C0defined by the singularity nearest to z0, labeled zs. If we now make a Taylor expansion about some point z1within C0(which we can do because f.z/has known values in the neighborhood of z1), this new expansion may have a circle of con- vergence C1that is not entirely within C0, thereby defining a function that is analytic in the region that is the union of C1andC2. Note that if we need to obtain actual values of f.z/forzwithin the intersection of C0andC1we may use either Taylor expansion, but in the region within only one circle we must use the expansion that is valid there (the other expansion will not converge). A generalization of the above analysis leads to the beautiful and valuable result that if two analytic functions coincide in any region, or even on any finite line segment, they are the same function, and therefore defined over the entire range of both function definitions. After Weierstrass this process of enlarging the region in which we have the specification of an analytic function is called analytic continuation, and the process may be carried out repeatedly to maximize the region in which the function is defined. Consider the situation pictured in Fig. 11.15, where the only singularity of f.z/is at zsand f.z/is originally defined by its Taylor expansion about z0, with circle of convergence C0. By making ana- lytic continuations as shown by the series of circles C1;:::, we can cover the entire annular region of analyticity shown in the figure, and can use the original Taylor series to generate new expansions that apply to regions within the other circles. C0C1z0z1zs FIGURE 11.14 Analytic continuation. One step. ArfKen_Ch11-9780123846549.tex 11.6 Singularities 505 C1 C2 C3 C4 C5C0 z0zs FIGURE 11.15 Analytic continuation. Many steps. iP 1 FIGURE 11.16 Radii of convergence of power-series expansions for Example 11.6.5. Example 11.6.5 ANALYTIC CONTINUATION Consider these two power-series expansions: f1.z/D1X nD0.1/n.z1/n; (11.57) f2.z/D1X nD0in1.zi/n: (11.58) Each has a unit radius of convergence; the circles of convergence overlap, as can be seen from Fig. 11.16. To determine whether these expansions represent the same analytic function in over- lapping domains, we can check to see if f1.z/Df2.z/for at least a line segment in the region of overlap. A suitable line is the diagonal that connects the origin with 1Ci, pass- ing through the intermediate point .1Ci/=2. Setting zD. C1 2/.1Ci/(chosen to make D0an interior point of the overlap region), we expand f1andf2about D0to find out ArfKen_Ch11-9780123846549.tex 506 Chapter 11 Complex Variable Theory whether their power series coincide. Initially we have (as functions of ) f1D1X nD0.1/n .1Ci/ 1i 2n ; f2D1X nD0in1 .1Ci/ C1i 2n : Applying the binomial theorem to obtain power series in , and interchanging the order of the two sums, f1D1X jD0.1/j.1Ci/j j1X nDjn j1i 2nj ; f2D1X jD0ij1.1Ci/j j1X nDjinjn j1i 2nj D1X jD01 i.1/j.1i/j j1X nDjn j1Ci 2nj : To proceed further we need to evaluate the summations over n. Referring to Exercise 1.3.5, where it was shown that 1X nDjn j xnjD1 .1x/jC1; we get f1D1X jD0.1/j.1Ci/j j2 1CijC1 D1X jD0.1/j2jC1 j 1Ci; f2D1X jD01 i.1/j.1i/j j2 1ijC1 D1X jD0.1/j2jC1 j i.1i/Df1; confirming that f1and f2are the same analytic function, now defined over the union of the two circles in Fig. 11.16. Incidentally, both f1andf2are expansions of 1=z(about the respective points 1 and i), so1=zcould also be regarded as an analytic continuation of f1,f2, or both to the entire complex plane except the singular point at zD0. The expansion in powers of is also a representation of 1=z, but its range of validity is only a circle of radius 1=p 2about .1Ci/=2and it does not analytically continue f.z/outside the union of C1andC2. The use of power series is not the only mechanism for carrying out analytic continu- ations; an alternative and powerful method is the use of functional relations , which are formulas that relate values of the same analytic function f.z/at different z. As an exam- ple of a functional relation, the integral representation of the gamma function, given in ArfKen_Ch11-9780123846549.tex 11.6 Singularities 507 Table 1.2, can be manipulated (see Chapter 13) to show that 0.zC1/Dz0.z/, consis- tent with the elementary result that nWDn.n1/W. This functional relation can be used to analytically continue 0.z/to values of zfor which the integral representation does not converge. Exercises 11.6.1 As an example of an essential singularity consider e1=zaszapproaches zero. For any complex number z0;z06D0, show that e1=zDz0 has an infinite number of solutions. 11.6.2 Show that the function w.z/D.z21/1=2 is single-valued if we make branch cuts on the real axis for x>1and for x<1. 11.6.3 A function f.z/can be represented by f.z/Df1.z/ f2.z/; in which f1.z/and f2.z/are analytic. The denominator, f2.z/;vanishes at zDz0; showing that f.z/has a pole at zDz0. However, f1.z0/6D0;f0 2.z0/6D0. Show that a1, the coefficient of .zz0/1in a Laurent expansion of f.z/atzDz0, is given by a1Df1.z0/ f0 2.z0/: 11.6.4 Determine a unique branch for the function of Exercise 11.6.2 that will cause the value it yields for f.i/to be the same as that found for f.i/inExample 11.6.4. Although Exercise 11.6.2 andExample 11.6.4 describe the same multivalued function, the specific values assigned for various zwill not agree everywhere, due to the difference in the location of the branch cuts. Identify the portions of the complex plane where both these descriptions do and do not agree, and characterize the differences. 11.6.5 Find all singularities of z1=3Cz1=4 .z3/3C.z2/1=2; and identify their types (e.g., second-order branch point, fifth-order pole, . . . ). Include any singularities at the point at infinity. Note. A branch point is of nth order if it requires n, but no fewer, circuits around the point to restore the original value. 11.6.6 The function F.z/Dln.z2C1/is made single-valued by straight-line branch cuts from.x;y/D.0;1/ to.1;1/ and from.0;C1/ to.0;C1/ . See Fig. 11.17. If F.0/D2 i, find the value of F.i2/. ArfKen_Ch11-9780123846549.tex 508 Chapter 11 Complex Variable Theory i−2i o −i FIGURE 11.17 Branch cuts for Exercise 11.6.6. 11.6.7 Show that negative numbers have logarithms in the complex plane. In particular, find ln.1/ . ANS. ln.1/Di. 11.6.8 For noninteger m, show that the binomial expansion of Exercise 11.5.2 holds only for a suitably defined branch of the function .1Cz/m. Show how the z-plane is cut. Explain whyjzj<1may be taken as the circle of convergence for the expansion of this branch, in light of the cut you have chosen. 11.6.9 The Taylor expansion of Exercises 11.5.2 and11.6.8 isnotsuitable for branches other than the one suitably defined branch of the function .1Cz/mfor noninteger m. (Note that other branches cannot have the same Taylor expansion since they must be distin- guishable.) Using the same branch cut of the earlier exercises for all other branches, find the corresponding Taylor expansions, detailing the phase assignments and Taylor coefficients. 11.6.10 (a) Develop a Laurent expansion of f.z/DTz.z1/U1about the point zD1valid for small values ofjz1j. Specify the exact range over which your expansion holds. This is an analytic continuation of the infinite series in Eq. (11.49). (b) Determine the Laurent expansion of f.z/about zD1but forjz1jlarge. Hint. Make a partial fraction decomposition of this function and use the geometric series. 11.6.11 (a) Given f1.z/DR1 0eztdt(with treal), show that the domain in which f1.z/exists (and is analytic) is Re.z/>0. (b) Show that f2.z/D1=zequals f1.z/overRe.z/>0and is therefore an analytic continuation of f1.z/over the entire z-plane except for zD0. (c) Expand 1=zabout the point zDi. You will have f3.z/D1X nD0an.zCi/n: What is the domain of this formula for f3.z/? ANS.1 zDi1X nD0in.zCi/n,jzCij<1. ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 509 11.7 C ALCULUS OF RESIDUES Residue Theorem If the Laurent expansion of a function, f.z/D1X nD1an.zz0/n; is integrated term by term by using a closed contour that encircles one isolated singular point z0once in a counterclockwise sense, we obtain, applying Eq. (11.29), anI .zz0/ndzD0; n6D1: (11.59) However, for nD1 ,Eq. (11.29) yields a1I .zz0/1dzD2ia1: (11.60) Summarizing Eqs. (11.59) and(11.60), we have I f.z/dzD2ia1: (11.61) The constant a1;the coefficient of .zz0/1in the Laurent expansion, is called the residue off.z/atzDz0. Now consider the evaluation of the integral, over a closed contour C, of a function that has isolated singularities at points z1,z2, . . . . We can handle this integral by deforming our contour as shown in Fig. 11.18. Cauchy’s integral theorem (Section 11.3) then leads to I Cf.z/dzCI C1f.z/dzCI C2f.z/dzCD 0; (11.62) C2 C1 C0Cℑz=y ℜz=x FIGURE 11.18 Excluding isolated singularities. ArfKen_Ch11-9780123846549.tex 510 Chapter 11 Complex Variable Theory where Cis in the positive, counterclockwise direction, but the contours C1;C2;:::, that, respectively, encircle z1;z2;::: are all clockwise. Thus, referring to Eq. (11.61), the inte- grals Ciabout the individual isolated singularities have the values I Cif.z/dzD2 ia1;i; (11.63) where a1;iis the residue obtained from the Laurent expansion about the singular point zDzi. The negative sign comes from the clockwise integration. Combining Eqs. (11.62) and(11.63), we have I Cf.z/dzD2i.a1;1Ca1;2C/ D2i.sum of the enclosed residues). (11.64) This is the residue theorem. The problem of evaluating a set of contour integrals is replaced by the algebraic problem of computing residues at the enclosed singular points. Computing Residues It is, of course, not necessary to obtain an entire Laurent expansion of f.z/about zDz0 to identify a1, the coefficient of .zz0/1in the expansion. If f.z/has a simple pole at zz0, then, with anthe coefficients in the expansion of f.z/, .zz0/f.z/Da1Ca0.zz0/Ca1.zz0/2C; (11.65) and, recognizing that .zz0/f.z/may not have a form permitting an obvious cancellation of the factor zz0, we take the limit of Eq. (11.65) asz!z0: a1Dlimz!z0 .zz0/f.z/ : (11.66) If there is a pole of order n>1atzz0, then.zz0/nf.z/must have the expansion .zz0/nf.z/DanCC a1.zz0/n1Ca0.zz0/nC: (11.67) We see that a1is the coefficient of .zz0/n1in the Taylor expansion of .zz0/nf.z/, and therefore we can identify it as satisfying a1D1 .n1/Wlimz!z0dn1 dzn1 .zz0/nf.z/ ; (11.68) where a limit is indicated to take account of the fact that the expression involved may be indeterminate. Sometimes the general formula, Eq. (11.68), is found to be more com- plicated than the judicious use of power-series expansions. See items 4 and 5 in Exam- ple 11.7.1 below. Essential singularities will also have well-defined residues, but finding them may be more difficult. In principle, one can use Eq. (11.48) with nD1 , but the integral involved may seem intractable. Sometimes the easiest route to the residue is by first finding the Laurent expansion. ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 511 Example 11.7.1 COMPUTING RESIDUES Here are some examples: 1. Residue of1 4zC1atzD1 4islimzD1 4 zC1 4 4zC1! D1 4; 2. Residue of1 sinzatzD0islimz!0 z sinz D1; 3. Residue oflnz z2C4atzD2eiis lim z!2ei.z2ei/lnz z2C4 D.ln 2Ci/ 4iD 4iln 2 4; 4. Residue ofz sin2zatzD; the pole is second order, and the residue is given by 1 1Wlimz!d dzz.z/ sin2z : However, it may be easier to make the substitution wDz, to note that sin2zD sin2w, and to identify the residue as the coefficient of 1=w in the expansion of .wC /=sin2waboutwD0. This expansion can be written wC  ww3 3WC2DwC w2w4 3C: The denominator expands entirely into even powers of w, so thein the numerator cannot contribute to the residue. Then, from the win the numerator and the leading term of the denominator, we find the residue to be 1. 5. Residue of f.z/Dcotz z.zC2/atzD0. The pole at zD0is second-order, and direct application of Eq. (11.48) leads to a complicated indeterminate expression requiring multiple applications of l’Hôpital’s rule. Perhaps easier is to introduce the initial terms of the expansions about zD0: cotzD.z/1CO.z/,1=.zC2/D1 2T1.z=2/CO.z2/U, reaching f.z/D1 z1 zCO.z/1 2h 1z 2CO.z2/i ; from which we can read out the residue as the coefficient of z1, namely1=4 . 6. Residue of e1=zatzD0. This is at an essential singularity; from the Taylor series of ewwithwD1= z, we have e1=zD11 zC1 2W 1 z2 C; from which we read out the value of the residue, 1.  ArfKen_Ch11-9780123846549.tex 512 Chapter 11 Complex Variable Theory Cauchy Principal Value Occasionally an isolated pole will be directly on the contour of an integration, causing the integral to diverge. A simple example is provided by an attempt to evaluate the real integral bZ adx x; (11.69) which is divergent because of the logarithmic singularity at xD0; note that the indefinite integral of x1islnx. However, the integral in Eq. (11.69) can be given a meaning if we obtain a convergent form when replaced by a limit of the form lim !0CZ adx xCbZ dx x: (11.70) To avoid issues with the logarithm of negative values of x, we change the variable in the first integral to yDx, and the two integrals are then seen to have the respective values lnlnaandlnbln, with sum lnblna. What has happened is that the increase towardC1 as1=xapproaches zero from positive values of xis compensated by a decrease toward1 as1=xapproaches zero from negative x. This situation is illustrated graphically in Fig. 11.19. Note that the procedure we have described does notmake the original integral of Eq. (11.69) convergent. In order for that integral to be convergent, it would be necessary F(x) ≈a−1 x−x0 x0−d x0+d x x0 FIGURE 11.19 Cauchy principal value cancellation, integral of 1=z. ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 513 that lim 1;2!0C2 641Z adx xCbZ 2dx x3 75 exist (meaning that the limit has a unique value) when 1and2approach zero indepen- dently. However, different rates of approach to zero by 1and2will cause a change in value of the integral. For example, if 2D21, then an evaluation like that of Eq. (11.70) would yield the result .ln1lna/C.lnbln2/Dlnblnaln 2. The limit then has no definite value, confirming our original statement that the integral diverges. Generalizing from the above example, we define the Cauchy principal value of the real integral of a function f.x/with an isolated singularity on the integration path at the point x0as the limit lim !0Cx0Z f.x/dxCZ x0Cf.x/dx: (11.71) The Cauchy principal value is sometimes indicated by preceding the integral sign by Por by drawing a horizontal line through the integration sign, as in PZ f.x/dx orZ f.x/dx: This notation, of course, presumes that the location of the singularity is known. Example 11.7.2 A CAUCHY PRINCIPAL VALUE Consider the integral ID1Z 0sinx xdx: (11.72) If we substitute for sinxthe equivalent formula sinxDeixeix 2i; we then have ID1Z 0eixeix 2ixdx: (11.73) We would like to separate this expression for Iinto two terms, but if we do so, each will become a logarithmically divergent integral. However, if we change the integration range in Eq. (11.72), originally .0;1/, to.;1/, that integral remains unchanged in the limit ArfKen_Ch11-9780123846549.tex 514 Chapter 11 Complex Variable Theory of small, and the integrals in Eq. (11.73) remain convergent so long as is not precisely zero. Then, rewriting the second of the two integrals in Eq. (11.73), to reach 1Z eix 2ixdxDZ 1eix 2ixdx; we see that the two integrals which together form Ican be written (in the limit !0C) as the Cauchy principal value integral ID1Z 1eix 2ixdx: (11.74)  The Cauchy principal value has implications for complex variable theory. Suppose now that, instead of having a break in the integration path from x0tox0C, we connect the two parts of the path by a circular arc passing, in the complex plane, either above or below the singularity at x0. Let’s continue the discussion in conventional complex-variable notation, denoting the singular point as z0, so our arc will be a half circle (of radius ) passing either counterclockwise below the singularity at z0or clockwise above z0. We restrict further analysis to singularities no stronger than 1=.zz0/, so we are dealing with a simple pole. Looking at the Laurent expansion of the function f.z/to be integrated, it will have initial terms a1 zz0Ca0C; and the integration over a semicircle of radius will take (in the limit !0C) one of the two forms (in the polar representation zz0Drei, with dzDireidandrD): IoverD0Z dieiha1 eiCa0Ci D0Z  ia1Cieia0C d! ia1; (11.75) IunderD2Z dieiha1 eiCa0Ci D2Z  ia1Cieia0C d!ia1: (11.76) Note that all but the first term of each of Eqs. (11.75) and(11.76) vanishes in the limit !0C, and that each of these equations yields a result that is in magnitude half the value that would have been obtained by a full circuit around the pole. The signs associated with the semicircles correspond as expected to the direction of travel, and the two semicircular integrals average to zero. We occasionally will want to evaluate a contour integral of a function f.z/on a closed path that includes the two pieces of a Cauchy principal value integralR f.z/dzwith a simple pole at z0, a semicircular arc connecting them at the singularity, and whatever other curve Cis needed to close the contour (see Fig. 11.20). ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 515 C1C2 xr R FIGURE 11.20 A contour including a Cauchy principal value integral. These contributions combine as follows, noting that in the figure the contour passes over the point z0: Z f.z/dzCIoverCZ C2f.z/dzD2iX residues (other than at z0); which rearranges to give Z f.z/dzDIoverZ C2f.z/dzC2iX residues (other than at z0): (11.77) On the other hand, we could have chosen the contour to pass under z0, in which case, instead of Eq. (11.77) we would get Z f.z/dzDIunderZ C2f.z/dzC2iX residues (other than at z0)C2ia1; (11.78) where the residue denoted a1is from the pole at z0.Equations (11.77) and(11.78) are in agreement because 2ia1IunderDIover, so for the purpose of evaluating the Cauchy principal value integral, it makes no difference whether we go below or above the singu- larity on the original integration path. Pole Expansion of Meromorphic Functions Analytic functions f.z/that have only isolated poles as singularities are called meromor- phic. Mittag-Leffler showed that, instead of making an expansion about a single regu- lar point (a Taylor expansion) or about an isolated singular point (a Laurent expansion), it was also possible to make an expansion each of whose terms arises from a different pole of f.z/. Mittag-Leffler’s theorem assumes that f.z/is analytic at zD0and at all other points (excluding infinity) with the exception of discrete simple poles at points z1, z2,:::, with respective residues b1,b2,:::. We choose to order the poles in a way such that0<jz1jjz2j , and we assume that in the limit of large z,jf.z/=zj!0. Then, ArfKen_Ch11-9780123846549.tex 516 Chapter 11 Complex Variable Theory Mittag-Leffler’s theorem states that f.z/Df.0/C1X nD1bn1 zznC1 zn : (11.79) To prove the theorem, we make the preliminary observation that the quantity being summed in Eq. (11.79) can be written z bn zn.znz/; suggesting that it might be useful to consider a contour integral of the form INDI CNf.w/dw w.wz/; wherewis another complex variable and CNis a circle enclosing the first Npoles of f.z/. Since CN, which has a radius we denote RN, has total arc length 2RN, and the absolute value of the integrand asymptotically approaches jf.RN/j=R2 N, the large- zbehavior of f.z/guarantees that limRN!1IND0. We now obtain an alternate expression for INusing the residue theorem. Recognizing thatCNencircles simple poles at wD0,wDz, andwDzn,nD1:::N, that f.w/is nonsingular at wD0andwDz, and that the residue of f.z/=w.wz/atznis just bn=zn.znz/, we have IND2if.0/ zC2if.z/ zCNX nD12ibn zn.znz/: Taking the large- Nlimit, in which IND0, we recover Mittag-Leffler’s theorem, Eq. (11.79). The pole expansion converges when the condition limz!1jf.z/=zjD0is satisfied. Mittag-Leffler’s theorem leads to a number of interesting pole expansions. Consider the following examples. Example 11.7.3 POLE EXPANSION OF tanz Writing tanzDeizeiz i.eizCeiz/; we easily see that the only singularities of tanzare for real values of z, and they occur at the zeros of cosx, namely at=2 ,3=2 , . . . , or in general at znD.2nC1/=2 . ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 517 To obtain the residues at these points, we take the limit (using l’Hôpital’s rule) bnD lim z!.2nC1/ 2.z.2nC1/=2/ sinz cosz DsinzC.z.2nC1/=2/ cosz sinz zD.2nC1/ 2D1; the same value for every pole. Noting that tan.0/D0, and that the poles within a circle of radius .NC1/ will be those (of both signs) referred to here by nvalues 0through N,Eq. (11.79) for the current case (but only through N) yields tanzDNX nD0.1/1 z.2nC1/=2C1 .2nC1/=2 CNX nD0.1/1 zC.2nC1/=2C1 .2nC1/=2 DNX nD0.1/1 z.2nC1/=2C1 zC.2nC1/=2 : Combining terms over a common denominator, and taking the limit N!1 , we reach the usual form of the expansion: tanzD2z1 .=2/2z2C1 .3=2/2z2C1 .5=2/2z2C : (11.80)  Example 11.7.4 POLE EXPANSION OF cotz This example proceeds much as the preceding one, except that cotzhas a simple pole at zD0, with residueC1. We therefore consider instead cotz1=z, thereby removing the singularity. The singular points are now simple poles at n(n6D0), with residues (again obtained via l’Hôpital’s rule) bnDlimz!n.zn/cotzDlimz!n.zn/.zcoszsinz/ zsinz DzcoszsinzC.zn/. zsinz/ sinzCzcosz zDnDC1: Noting that cotz1=zis zero at zD0(the second term in the expansion of cotzisz=3), we have cotz1 zDNX nD11 znC1 nC1 zCnC1 n ; ArfKen_Ch11-9780123846549.tex 518 Chapter 11 Complex Variable Theory which rearranges to cotzD1 zC2z1 z22C1 z2.2/2C1 z2.3/2C : (11.81)  In addition to Eqs. (11.80) and (11.81), two other pole expansions of importance are seczD1 .=2/2z23 .3=2/2z2C5 .5=2/2z2 ; (11.82) csczD1 z2z1 z221 z2.2/2C1 z2.3/2C : (11.83) Counting Poles and Zeros It is possible to obtain information about the numbers of poles and zeros of a function f.z/that is otherwise analytic within a closed region by consideration of its logarithmic derivative, namely f0.z/=f.z/. The starting point for this analysis is to write an expression forf.z/relative to a point z0where there is either a zero or a pole in the form f.z/D.zz0/g.z/; with g.z/finite and nonzero at zDz0. That requirement identifies the limiting behavior of f.z/near z0as proportional to .zz0/, and also causes f0=fto assume near zDz0the form f0.z/ f.z/D.zz0/1g.z/C.zz0/g0.z/ .zz0/g.z/D zz0Cg0.z/ g.z/: (11.84) Equation (11.84) shows that, for all nonzero (i.e., if z0is either a zero or a pole), f0=f has a simple pole at zDz0with residue. Note that because g.z/is required to be nonzero and finite, the second term of Eq. (11.84) cannot be singular. Applying now the residue theorem to Eq. (11.84) for a closed region within which f.z/ is analytic except possibly at poles, we see that the integral of f0=faround a closed contour yields the result I Cf0.z/ f.z/dzD2i NfPf ; (11.85) where Pfis the number of poles of f.z/within the region enclosed by C, each multiplied by its order, and Nis the number of zeros of f.z/enclosed by C, each multiplied by its multiplicity. The counting of zeros is often facilitated by using Rouché’s theorem, which states Iff.z/andg.z/are analytic in the region bounded by a curve Candjf.z/j>jg.z/j onC, then f.z/and f.z/Cg.z/have the same number of zeros in the region bounded byC. ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 519 To prove Rouché’s theorem, we first write, from Eq. (11.85), I Cf0.z/ f.z/dzD2i NfandI Cf0.z/Cg0.z/ f.z/Cg.z/dzD2i NfCg; where Nfdesignates the number of zeros of fwithin C. Then we observe that because the indefinite integral of f0=fislnf,Nfis the number of times the argument of fcycles through 2when Cis traversed once in the counterclockwise direction. Similarly, we note thatNfCgis the number of times the argument of fCgcycles through 2on traversal of the contour C. We next write fCgDf 1Cg f and arg.fCg/Darg.f/Carg 1Cg f ; (11.86) using the fact that the argument of a product is the sum of the arguments of its factors. It is then clear that the number of cycles through 2ofarg.fCg/is equal to the number of cycles of arg.f/plus the number of cycles of arg.1Cg=f/. But becausejg=fj<1, the real part of 1Cg=fnever becomes negative, and its argument is therefore restricted to the range=2<arg.1Cg=f/<=2 . Therefore arg.1Cg=f/cannot cycle through 2, the number of cycles of arg.fCg/must be equal to the number of cycles of argf, and fCgand fmust have the same number of zeros within C. This completes the proof of Rouché’s theorem. Example 11.7.5 COUNTING ZEROS Our problem is to determine the number of zeros of F.z/Dz32zC11with moduli between 1 and 3. Since F.z/is analytic for all finite z, we could in principle simply apply Eq. (11.85) for the contour consisting of the circles jzjD1(clockwise) andjzjD3(coun- terclockwise), setting PFD0and solving for NF. However, that approach will in practice prove difficult. Instead, we simplify the problem by using Rouché’s theorem. We first compute the number of zeros within jzjD1, writing F.z/Df.z/Cg.z/, with f.z/D11andg.z/Dz32z. It is clear thatjf.z/j>jg.z/jwhenjzjD1, so, by Rouché’s theorem, fandfCghave the same number of zeros within this circle. Since f.z/D11 has no zeros, we conclude that all the zeros of F.z/are outsidejzjD1. Next we compute the number of zeros within jzjD3, taking for this purpose f.z/Dz3, g.z/D112z. WhenjzjD3, we havejf.z/jD27>jg.z/j, soFand fhave the same number of zeros, namely three (the three-fold zero of fatzD0). Thus, the answer to our problem is that Fhas three zeros, all with moduli between 1 and 3.  Product Expansion of Entire Functions We remind the reader that a function f.z/that is analytic for all finite zis called an entire function. Referring to Eq. (11.84), we see that if f.z/is an entire function, then f0.z/=f.z/ will be meromorphic, with all its poles simple. Assuming for simplicity that the zeros of f ArfKen_Ch11-9780123846549.tex 520 Chapter 11 Complex Variable Theory are simple and at points zn, so thatin Eq. (11.84) is1, we can invoke the Mittag-Leffler theorem to write f0=fas the pole expansion f0.z/ f.z/Df0.0/ f.0/C1X nD11 zznC1 zn : (11.87) Integrating Eq. (11.87) yields zZ 0f0.z/ f.z/dzDlnf.z/lnf.0/ Dz f0.0/ f.0/C1X nD1 ln.zzn/ln.zn/Cz zn : Exponentiating, we obtain the product expansion f.z/Df.0/expz f0.0/ f.0/1Y nD1 1z zn ez=zn: (11.88) Examples are the product expansions for sinzDz1Y nD1 n6D0 1z n ez=nDz1Y nD1 1z2 n22 ; (11.89) coszD1Y nD1 1z2 .n1=2/22 : (11.90) The expansion of sinzcannot be obtained directly from Eq. (11.88), but its derivation is the subject of Exercise 11.7.5. We also point out here that the gamma function has a product expansion, discussed in Chapter 13. Exercises 11.7.1 Determine the nature of the singularities of each of the following functions and evaluate the residues.a>0/. (a)1 z2Ca2: (b)1 .z2Ca2/2: (c)z2 .z2Ca2/2: (d)sin 1= z z2Ca2: (e)zeCiz z2Ca2: (f)zeCiz z2a2: (g)eCiz z2a2: (h)zk zC1;0<k<1: ArfKen_Ch11-9780123846549.tex 11.7 Calculus of Residues 521 Hint. For the point at infinity, use the transformation wD1=zforjzj!0. For the residue, transform f.z/dzintog.w/dwand look at the behavior of g.w/. 11.7.2 Evaluate the residues at zD0andzD1 ofcotz=z.zC1/. 11.7.3 The classical definition of the exponential integral Ei .x/forx>0is the Cauchy prin- cipal value integral Ei.x/DxZ 1et tdt; where the integration range is cut at xD0. Show that this definition yields a convergent result for positive x. 11.7.4 Writing a Cauchy principal value integral to deal with the singularity at xD1, show that, if 0<p<1, 1Z 0xp x1dxD cotp: 11.7.5 Explain why Eq. (11.88) is not directly applicable to the product expansion of sinz. Show how the expansion, Eq. (11.89), can be obtained by expanding instead sinz=z. 11.7.6 Starting from the observations 1. f.z/Danznhasnzeros, and 2. for sufficiently large jRj,jPn1 mD0amRmj<janRnj, use Rouché’s theorem to prove the fundamental theorem of algebra (namely that every polynomial of degree nhasnroots). 11.7.7 Using Rouché’s theorem, show that all the zeros of F.z/Dz64z3C10lie between the circlesjzjD1andjzjD2. 11.7.8 Derive the pole expansions of seczandcsczgiven in Eqs. (11.82) and (11.83). 11.7.9 Given that f.z/D.z23zC2/=z, apply a partial fraction decomposition to f0=fand show directly thatH Cf0.z/=f.z/dzD2i.NfPf/, where NfandPfare, respec- tively, the numbers of zeros and poles encircled by C(including their multiplicities). 11.7.10 The statement that the integral halfway around a singular point is equal to one-half the integral all the way around was limited to simple poles. Show, by a specific example, that Z Semicirclef.z/dzD1 2I Circlef.z/dz does not necessarily hold if the integral encircles a pole of higher order. Hint. Try f.z/Dz2. ArfKen_Ch11-9780123846549.tex 522 Chapter 11 Complex Variable Theory 11.7.11 A function f.z/is analytic along the real axis except for a third-order pole at zDx0. The Laurent expansion about zDx0has the form f.z/Da3 .zx0/3Ca1 zx0Cg.z/; with g.z/analytic at zDx0. Show that the Cauchy principal value technique is appli- cable, in the sense that (a) lim!0nRx0 1f.x/dxCR1 x0Cf.x/dxo is finite. (b)R Cx0f.z/dzDia1, where Cx0denotes a small semicircle about zDx0. 11.7.12 The unit step function is defined as (compare Exercise 1.15.13) u.sa/D0; s<a 1; s>a: Show that u.s/has the integral representations (a) u.s/Dlim"!0C1 2i1Z 1eixs xi"dx. (b) u.s/D1 2C1 2i1Z 1eixs xdx. Note. The parameter sis real. 11.8 E VALUATION OF DEFINITE INTEGRALS Definite integrals appear repeatedly in problems of mathematical physics as well as in pure mathematics. In Chapter 1 we reviewed several methods for integral evaluation, there noting that contour integration methods were powerful and deserved detailed study. We have now reached a point where we can explore these methods, which are applicable to a wide variety of definite integrals with physically relevant integration limits. We start with applications to integrals containing trigonometric functions, which we can often convert to forms in which the variable of integration (originally an angle) is converted into a complex variable z, with the integration integral becoming a contour integral over the unit circle. Trigonometric Integrals, Range (0,2 ) We consider here integrals of the form ID2Z 0f.sin;cos/d; (11.91) ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 523 where fis finite for all values of . We also require fto be a rational function of sin andcosso that it will be single-valued. We make a change of variable to zDei;dzDieid; with the range in , namely.0;2/, corresponding to eimoving counterclockwise around the unit circle to form a closed contour. Then we make the substitutions dDidz z;sinDzz1 2i;cosDzCz1 2; (11.92) where we have used Eq. (1.133) to represent sinandcos. Our integral then becomes IDiI fzz1 2i;zCz1 2dz z; (11.93) with the path of integration the unit circle. By the residue theorem, Eq. (11.64), ID.i/2iX residues within the unit circle. (11.94) Note that we must use the residues of f=z. Here are two preliminary examples. Example 11.8.1 INTEGRAL OF cos INDENOMINATOR Our problem is to evaluate the definite integral ID2Z 0d 1Cacos;jaj<1: ByEq. (11.93) this becomes IDiI unit circledz zT1C.a=2/.zCz1/U Di2 aIdz z2C.2=a/zC1: The denominator has roots z1D1Cp 1a2 aand z2D1p 1a2 a: Noting that z1z2D1, it is easy to see that z2is within the unit circle and z1is outside. Writing the integral in the form Idz .zz1/.zz2/; we see that the residue of the integrand at zDz2is1=.z2z1/, so application of the residue theorem yields IDi2 a2i1 z2z1: ArfKen_Ch11-9780123846549.tex 524 Chapter 11 Complex Variable Theory Inserting the values of z1andz2, we obtain the final result 2Z 0d 1CacosD2p 1a2;jaj<1:  Example 11.8.2 ANOTHER TRIGONOMETRIC INTEGRAL Consider ID2Z 0cos 2 d 54 cos: Making the substitutions identified in Eqs. (11.92) and(11.93), the integral Iassumes the form IDI 1 2.z2Cz2/ 52.zCz1/i dz z Di 4I.z4C1/dz z2 z1 2 .z2/; where the integration is around the unit circle. Note that we identified cos 2 as.z2C z2/=2, which is simpler than reducing it first to its equivalent in terms of sinzandcosz. We see that the integrand has poles at zD0(of order 2), and simple poles at zD1=2and zD2. Only the poles at zD0andzD1=2are within the contour. AtzD0the residue of the integrand is d dz" z4C1 z1 2 .z2/# zD0D5 2; while its residue at zD1=2is z4C1 z2.z2/ zD1=2D17 6: Applying the residue theorem, we have IDi 4.2i/5 217 6 D 6:  We stress that integrals of the type now under consideration are evaluated after trans- forming them so that they can be identified as exactly equivalent to contour integrals to which we can apply the residue theorem. Further examples are in the exercises. ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 525 Integrals, Range1 to1 Consider now definite integrals of the form ID1Z 1f.x/dx; (11.95) where it is assumed that f.z/is analytic in the upper half-plane except for a finite number of poles. For the moment will be assumed that there are no poles on the real axis. Cases not satisfying this condition will be considered later. In the limitjzj!1 in the upper half-plane ( 0argz),f.z/vanishes more strongly than 1=z. Note that there is nothing unique about the upper half-plane. The method described here can be applied, with obvious modifications, if f.z/vanishes sufficiently strongly on the lower half-plane. The second assumption stated above makes it useful to evaluate the contour integralH f.z/dzon the contour shown in Fig. 11.21, because the integral Iis given by the inte- gration along the real axis, while the arc, of radius R, with R!1 , gives a negligible contribution to the contour integral. Thus, IDI f.z/dz; and the contour integral can be evaluated by applying the residue theorem. Situations of this sort are of frequent occurrence, and we therefore formalize the condi- tions under which the integral over a large arc becomes negligible: IflimR!1 z f.z/D0for all zDReiwithin the range12, then lim R!1Z Cf.z/dzD0; (11.96) where C is the arc over the angular range 1to2on a circle of radius R with center at the origin. Poles FIGURE 11.21 A contour closed by a large semicircle in the upper half-plane. ArfKen_Ch11-9780123846549.tex 526 Chapter 11 Complex Variable Theory To prove Eq. (11.96), simply write the integral over Cin polar form: lim R!1 Z Cf.z/dz 2Z 1lim R!1 f.Rei/i Rei d .21/lim R!1 f.Rei/Rei D0: Now, using the contour of Fig. 11.21, letting Cdenote the semicircular arc from D0 toD, I f.z/dzDlim R!1RZ Rf.x/dxClim R!1Z Cf.z/dz D2iX residues (upper half-plane), (11.97) where our second assumption has caused the vanishing of the integral over C. Example 11.8.3 INTEGRAL OF MEROMORPHIC FUNCTION Evaluate ID1Z 0dx 1Cx2: This is not in the form we require, but it can be made so by noting that the integrand is even and we can write ID1 21Z 1dx 1Cx2: (11.98) We note that f.z/D1=.1Cz2/is meromorphic; all its singularities for finite zare poles, and it also has the property that z f.z/vanishes in the limit of large jzj. Therefore, we may apply Eq. (11.97), so 1 21Z 1dx 1Cx2D1 2.2i/X residues of1 1Cz2(upper half-plane): Here and in every other similar problem we have the question: Where are the poles? Rewriting the integrand as 1 z2C1D1 .zCi/.zi/; we see that there are simple poles (order 1) at zDiandzDi. The residues are atzDi:1 zCi zDiD1 2i;and at zDi:1 zi zDiD1 2i: ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 527 However, only the pole at zDCiis enclosed by the contour, so our result is 1Z 0dx 1Cx2D1 2.2i/1 2iD 2: (11.99) This result is hardly a surprise, as we presumably already know that 1Z 0dx 1Cx2Dtan1x 1 0Darctan x 1 0D 2; but, as shown in later examples, the techniques illustrated here are also easy to apply when more elementary methods are difficult or impossible. Before leaving this example, note that we could equally well have closed the contour with a semicircle in the lower half-plane, as z f.z/vanishes on that arc as well as that in the upper half-plane. Then, taking the contour so the real axis is traversed from 1 toC1, the path would be clockwise (see Fig. 11.22), so we would need to take 2 itimes the residue of the pole that is now encircled (at zDi). Thus, we have ID1 2.2i/.1=2 i/, which (as it must) evaluates to the same result we obtained previously, namely =2. Integrals with Complex Exponentials Consider the definite integral ID1Z 1f.x/eiaxdx; (11.100) with areal and positive. (This is a Fourier transform; see Chapter 19.) We assume the following two conditions: f.z/is analytic in the upper half-plane except for a finite number of poles. limjzj!1 f.z/D0;0argz. Note that this is a less restrictive condition than the second condition imposed on f.z/for our previous integration ofR1 1f.x/dx. Pole FIGURE 11.22 A contour closed by a large semicircle in the lower half-plane. ArfKen_Ch11-9780123846549.tex 528 Chapter 11 Complex Variable Theory We again employ the half-circle contour shown in Fig. 11.21. The application of the calculus of residues is the same as the example just considered, but here we have to work harder to show that the integral over the (infinite) semicircle goes to zero. This integral becomes, for a semicircle of radius R, IRDZ 0f.Rei/eiaRcosaRsini Reid; where theintegration is over the upper half-plane, 0. Let Rbe sufficiently large thatjf.z/jDj f.Rei/j<"for allwithin the integration range. Our second assumption onf.z/tells us that as R!1 ,"!0. Then jIRj"RZ 0eaRsindD2"R=2Z 0eaRsind: (11.101) We now note that in the range T0;=2U; 2 sin; as is easily seen from Fig. 11.23. Substituting this inequality into Eq. (11.101), we have jIRj2"R=2Z 0e2aR=dD2"R1eaR 2aR=< a"; showing that lim R!1IRD0: This result is also important enough to commemorate; it is sometimes known as Jordan’s lemma. Its formal statement is IflimRD1f.z/D0for all zDReiin the range 0, then lim R!1Z Ceiazf.z/dzD0; (11.102) where a>0andCis a semicircle of radius Rin the upper half-plane with center at the origin. Note that for Jordan’s lemma the upper and lower half-planes are not equivalent, because the condition a>0causes the exponent aRsinonly to be negative and yield a neg- ligible result in the upper half-plane. In the lower half-plane, the exponential is positive and the integral on a large semicircle there would diverge. Of course, we could extend the theorem by considering the case a<0, in which event the contour to be used would then be a semicircle in the lower half-plane. ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 529 (a) (b)1y 2π θ FIGURE 11.23 (a)yD.2=/ , (b) yDsin. Returning now to integrals of the type represented by Eq. (11.100), and using the contour shown in Fig. 11.21, application of the residue theorem yields the general result (for a>0), 1Z 1f.x/eiaxdxD2iX residues of eiazf.z/(upper half-plane); (11.103) where we have used Jordan’s lemma to set to zero the contribution to the contour integral from the large semicircle. Example 11.8.4 OSCILLATORY INTEGRAL Consider ID1Z 0cosx x2C1dx; which we initially manipulate, introducing cosxD.eixCeix/=2, as follows: ID1 21Z 0eixdx x2C1C1 21Z 0eixdx x2C1 D1 21Z 0eixdx x2C1C1 21Z 0eixd.x/ .x/2C1D1 21Z 1eixdx x2C1; thereby bringing Ito the form presently under discussion. We now note that in this problem f.z/D1=.z2C1/, which certainly approaches zero for largejzj, and the exponential factor is of the form eiaz, with aDC1 . We may therefore evaluate the integral using Eq. (11.103), with the contour shown in Fig. 11.21. The quantity whose residues are needed is eiz z2C1Deiz .zCi/.zi/; ArfKen_Ch11-9780123846549.tex 530 Chapter 11 Complex Variable Theory and we note that the exponential, an entire function, contributes no singularities. So our singularities are simple poles at zDi. Only the pole at zDCiis within the contour, and its residue is ei2=2i, which reduces to 1=2ie. Our integral therefore has the value ID1 2.2i/1 2ieD 2e:  Our next example is an important integral, the evaluation of which involves the principal-value concept and a contour that apparently needs to go through a pole. Example 11.8.5 SINGULARITY ON CONTOUR OF INTEGRATION We now consider the evaluation of ID1Z 0sinx xdx: (11.104) Writing the integrand as .eizeiz/=2iz, an attempt to do as we did in Example 11.8.4 leads to the problem that each of the two integrals into which Ican be separated is individ- ually divergent. This is a problem we have already encountered in discussing the Cauchy principal value of this integral. Referring to (11.74), we write Ias ID1Z 1eixdx 2ix; (11.105) suggesting that we consider the integral of eiz=2izover a suitable closed contour. We now note that although the gap at xD0is infinitesimal, that point is a pole of eiz=2iz, and we must draw a contour which avoids it, using a small semicircle to con- nect the points atandC. Compare with the discussion at Eqs. (11.75) and(11.76). Choosing the small semicircle above the pole, as in Fig. 11.20, we then have a contour that encloses nosingularities. The integral around this contour can now be identified as consisting of (1) the two semi- infinite segments constituting the principal value integral in Eq. (11.105), (2) the large semicircle CRof radius R(R!1 ), and (3) a semicircle Crof radius r(r!0), traversed clockwise, so Ieiz 2izdzDICZ Creiz 2izdzCZ CReiz 2izdzD0: (11.106) By Jordan’s lemma, the integral over CRvanishes. As discussed at Eq. (11.75), the clock- wise path Crhalf-way around the pole at zD0contributes half the value of a full circuit, namely (allowing for the clockwise direction of travel) itimes the residue of eiz=2izat zD0. This residue has value 1=2i, soR CrD i.1=2i/D=2 , and, solving Eq. (11.106) ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 531 forI, we then obtain ID1Z 0sinx xdxD 2: (11.107) Note that it was necessary to close the contour in the upper half-plane. On a large circle in the lower half-plane, eizbecomes infinite and Jordan’s lemma cannot be applied.  Another Integration Technique Sometimes we have an integral on the real range .0;1/that lacks the symmetry needed to extend the integration range to .1;1/. However, it may be possible to identify a direction in the complex plane on which the integrand has a value identical to or conve- niently related to that of the original integral, thereby permitting construction of a contour facilitating the evaluation. Example 11.8.6 EVALUATION ON A CIRCULAR SECTOR Our problem is to evaluate the integral ID1Z 0dx x3C1; which we cannot convert easily into an integral on the range .1;1/. However, we note that along a line with argument D2=3 ,z3will have the same values as at corresponding points on the real line; note that .re2i=3/3Dr3e2iDr3. We therefore consider Idz z3C1 on the contour shown in Fig. 11.24. The part of the contour along the positive real axis, labeled A, simply yields our integral I. The integrand approaches zero sufficiently rapidly for largejzjthat the integral on the large circular arc, labeled Cin the figure, vanishes. On B AC 2π/3 FIGURE 11.24 Contour for Example 11.8.6. ArfKen_Ch11-9780123846549.tex 532 Chapter 11 Complex Variable Theory the remaining segment of the contour, labeled B, we note that dzDe2i=3dr,z3Dr3, and Z Bdz z3C1D0Z 1e2i=3dr r3C1De2i=31Z 0dr r3C1De2i=3I: Therefore, Idz z3C1D 1e2i=3 I: (11.108) We now need to evaluate our complete contour integral using the residue theorem. The integrand has simple poles at the three roots of z3C1, which are at z1Dei=3,z2Dei, andz3De5i=3, as marked in Fig. 11.24. Only the pole at z1is enclosed by our contour. The residue at zDz1is limzDz1zz1 z3C1D1 3z2 zDz1D1 3e2i=3: Equating 2itimes this result to the value of the contour integral as given in Eq. (11.108), we have  1e2i=3 ID2i1 3e2i=3 : Solution for Iis facilitated if we multiply through by ei=3, obtaining initially  ei=3ei=3 ID2i 1 3 ; which is easily rearranged to ID 3 sin=3D 3p 3=2D2 3p 3:  Avoidance of Branch Points Sometimes we must deal with integrals whose integrands have branch points. In order to use contour integration methods for such integrals we must choose contours that avoid the branch points, enclosing only point singularities. Example 11.8.7 INTEGRAL CONTAINING LOGARITHM We now look at ID1Z 0lnx dx x3C1: (11.109) The integrand in Eq. (11.109) is singular at xD0, but the integration converges (the indef- inite integral of lnxisxlnxx). However, in the complex plane this singularity manifests ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 533 ABC FIGURE 11.25 Contour for Example 11.8.7. itself as a branch point, so if we are to recast this problem in a way involving a contour integral, we must avoid zD0and a branch cut from that point to zD1 . It turns out to be convenient to use a contour similar to that for Example 11.8.6, except that we must make a small circular detour about zD0and then draw the branch cut in a direction that remains outside our chosen contour. Noting also that the integrand has poles at the same points as those of Example 11.8.6, we consider a contour integral Ilnz dz z3C1; where the contour and the locations of the singularities of the integrand are as illustrated in Fig. 11.25. The integral over the large circular arc, labeled C, vanishes, as the factor z3in the denominator dominates over the weakly divergent factor lnzin the numerator (which diverges more weakly than any positive power of z). We also get no contribution to the contour integral from the arc at small r, since we have there lim r!02=3Z 0ln.rei/ 1Cr3e3iireid; which vanishes because rlnr!0. The integrals over the segments labeled AandBdo not vanish. To evaluate the integral over these segments, we need to make an appropriate choice of the branch of the multi- valued function lnz. It is natural to choose the branch so that on the real axis we have lnzDlnx(and not lnxC2niwith some nonzero n). Then the integral over the segment labeled Awill have the value I.7 To compute the integral over B, we note that on this segment z3Dr3anddzDe2i=3dr (as in Example 11.8.6), also but note that lnzDlnrC2i=3. There is little temptation here to use a different one of the multiple values of the logarithm, but for future reference note that we must use the value that is reached continuously from the value we already chose on the positive real axis, moving in a way that does not cross the branch cut. Thus, we cannot reach segment Aby clockwise travel from the positive real axis (thereby getting 7Because the integration converges at xD0, the value is not affected by the fact that this segment terminates infinitesimally before reaching that point. ArfKen_Ch11-9780123846549.tex 534 Chapter 11 Complex Variable Theory lnzDlnr4i=3) or any other value that would require multiple circuits around the branch point zD0. Based on the foregoing, we have Z Blnz dz z3C1D0Z 1lnrC2i=3 r3C1e2i=3drDe2i=3I2i 3e2i=31Z 0dr r3C1:(11.110) Referring to Example 11.8.6 for the value of the integral in the final term of Eq. (11.110), and combining the contributions to the overall contour integral, Ilnz dz z3C1D 1e2i=3 I2i 3e2i=32 3p 3 : (11.111) Our next step is to use the residue theorem to evaluate the contour integral. Only the pole at zDz1lies within the contour. The residue we must compute is limzDz1.zz1/lnz z3C1Dlnz 3z2 zDz1Di=3 3e2i=3Di 9e2 i=3; and application of the residue theorem to Eq. (11.111) yields  1e2i=3 I2i 3e2i=32 3p 3 D.2i/i 9 e2 i=3: (11.112) Solving for I, we get ID22 27: (11.113) Verification of the passage from Eq. (11.112) to (11.113) is left to Exercise 11.8.6.  Exploiting Branch Cuts Sometimes, rather than being an annoyance, a branch cut provides an opportunity for a creative way of evaluating difficult integrals. Example 11.8.8 USING A BRANCH CUT Let’s evaluate ID1Z 0xpdx x2C1;0<p<1: Consider the contour integral Izpdz z2C1; where the contour is that shown in Fig. 11.26. Note that zD0is a branch point, and we have taken the cut along the positive real axis. We assign zpits usual principal value ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 535 A B FIGURE 11.26 Contour for Example 11.8.8. (which is xp) just above the cut, so that the segment of the contour labeled A, which actually extends from "to1, converges in the limit of small "to the integral I. Neither the circle of radius "nor that at R!1 contributes to the value of the contour integral. On the remaining segment of the contour, labeled B, we have zDre2i, written this way so we can see that zpDrpe2pi. We use this value for zpon segment Bbecause we must get to Bby encircling zD0in the counterclockwise, mathematically positive direction. The contribution of segment Bto the contour integral is then seen to be 0Z 1rpe2pidr r2C1De2piI; so Izpdz z2C1D 1e2pi I: (11.114) To apply the residue theorem, we note that there are simple poles at z1Diandz2Di; to use these for evaluation of zpwe need to identify these as z1Dei=2andz2De3i=2. It would be a serious mistake to use z2Dei=2when evaluating zp 2. We now find the residues to be: Residue at z1:epi=2 2i;Residue at z2:e3pi=2 2i; and we have, referring to Eq. (11.114),  1e2pi ID.2i/1 2i epi=2e3pi=2 : (11.115) This equation simplifies to IDsin.p=2/ sinpD 2 cos. p=2/: (11.116) The details of the evaluation are left to Exercise 11.8.7.  The use of a branch cut, as illustrated in Example 11.8.8, is so helpful that sometimes it is advisable to insert a factor into a contour integral to create one that would not otherwise exist. To illustrate this, we return to an integral we evaluated earlier by another method. ArfKen_Ch11-9780123846549.tex 536 Chapter 11 Complex Variable Theory Example 11.8.9 INTRODUCING A BRANCH POINT Let’s evaluate once again the integral ID1Z 0dx x3C1; which we previously considered in Example 11.8.6. This time, we proceed by setting up the contour integral Ilnz dz z3C1; taking the contour to be that depicted in Fig. 11.26. Note that in the present problem the poles of the integrand are not those shown in Fig. 11.26, which was originally drawn to illustrate a different problem; for the locations of the poles of the present integrand, see Fig. 11.24. The virtue of the introduction of the factor lnzis that its presence causes the integral segments above and below the positive real axis not to cancel completely, but to yield a net contribution corresponding to an integral of interest. In the present problem (using the labeling in Fig. 11.26), we again have vanishing contributions from the small and large circles, and (taking the usual principal value for the logarithm on segment A), that segment contributes to the contour integral the expected value Z Alnz dz z3C1D1Z 0lnx dx x3C1: (11.117) However, segment Bmake the contribution Z Blnz dz z3C1D0Z 1.lnxC2i/dx x3C1; (11.118) and when Eqs. (11.117) and (11.118) are combined, the logarithmic terms cancel, and we are left with Ilnz dz z3C1DZ ACBlnz dz z3C1D2 i1Z 0dx x3C1D2 i I: (11.119) Note that what has happened is that the logarithm has disappeared (its contributions can- celed), but its presence caused the integral of current interest to be proportional to the value of the contour integral we introduced. To complete the evaluation, we need to evaluate the contour integral using the residue theorem. Note that the residues are those of the integrand, including the logarithmic factor, and this factor must be computed taking account of the branch cut. In the present problem, we identify poles at z1Dei=3,z2Dei, and z3De5i=3(not ei=3). The contour now ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 537 in use encircles all three poles. Their respective residues (denoted Ri) are R1Di 31 3e2i=3;R2D.i/1 3e6i=3;and R3D5i 31 3e10i=3; where the first parenthesized factor of each residue comes from the logarithm. Continuing, we have, referring to Eq. (11.119), 2 i ID2i.R1CR2CR3/I ID.R1CR2CR3/Di 9h e2 i=3C3C5e2i=3i D2 3p 3: More robust examples involving the introduction of lnzappear in the exercises.  Exploiting Periodicity The periodicity of the trigonometric functions (and that, in the complex plane, of the hyper- bolic functions) creates opportunities to devise contours in which multiple contributions corresponding to an integral of interest can be used to encircle singularities and enable use of the residue theorem. We illustrate with one example. Example 11.8.10 INTEGRAND PERIODIC ON IMAGINARY AXIS We wish to evaluate ID1Z 0x dx sinhx: Taking account of the sinusoidal behavior of the hyperbolic sine in the imaginary direction, we considerIz dz sinhz(11.120) on the contour shown in Fig.11.27. In drawing the contour we needed to be mindful of the singularities of the integrand, which are poles associated with the zeros of sinhz. Recog- nizing that sinh. xCiy/Dsinhxcosh iyCcosh xsinhiyDsinhxcosyCicosh xsiny;(11.121) B A OC B′πi−R −Rπi+R R FIGURE 11.27 Contour for Example 11.8.10. ArfKen_Ch11-9780123846549.tex 538 Chapter 11 Complex Variable Theory and that for all x,cosh x1, we see that sinhzis zero only for zDni, with nan integer. Moreover, because limz!0z=sinhzD1, the integrand of our present contour integral will not have a pole at zD0, but will have poles at zDnifor all nonzero integral n. For that reason, the lower horizontal line of the contour in Fig. 11.27, marked A, continues through zD0as a straight line on the real axis, but the upper horizontal line (for which yD), marked BandB0, has an infinitesimal semicircular detour, marked C, around the pole at zDi. Because the integrand in Eq. (11.120) is an even function of z, the integral on segment A, which extends from 1 toC1, has the value 2I. To evaluate the integral on segments BandB0, we first note, using Eq. (11.121), that sinh. xCi/Dsinh x, and that the integral on these segments is in the direction of negative x. Recognizing the integral on these segments as a Cauchy principal value, we write Z BCB0z dz sinhzD1Z 1xCi sinhxdx: Because x=sinhxis even and nonsingular at zD0, while i=sinhxis odd, this integral reduces to 1Z 1xCi sinhxdxD2I: Combining what we have up to this point, invoking the residue theorem, and noting that the integrand is negligible on the vertical connections at xD1 . We have Iz dz sinhzD4ICZ Cz dz sinhzD2i(residue of z=sinhzatzDi). (11.122) To complete the evaluation, we now note that the residue we need is lim z!iz.zi/ sinhzDi coshiD i; and, cf. Eqs. (11.75) and (11.76), the counterclockwise semicircle Cevaluates toitimes this residue. We have then 4IC.i/. i/D.2i/. i/;soID2 4:  Exercises 11.8.1 Generalizing Example 11.8.1, show that 2Z 0d abcosD2Z 0d absinD2 .a2b2/1=2;fora>jbj: What happens ifjbj>jaj? ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 539 11.8.2 Show thatZ 0d .aCcos/2Da .a21/3=2;a>1: 11.8.3 Show that2Z 0d 12tcosCt2D2 1t2;forjtj<1: What happens ifjtj>1? What happens if jtjD1? 11.8.4 Evaluate2Z 0cos 3 d 54 cos: ANS.=12 . 11.8.5 With the calculus of residues, show that Z 0cos2ndD.2n/W 22n.nW/2D.2n1/WW .2n/WW;nD0;1;2;:::: The double factorial notation is defined in Eq. (1.76). Hint. cosD1 2.eiCei/D1 2.zCz1/;jzjD1. 11.8.6 Verify that simplification of the expression in Eq. (11.112) yields the result given in Eq. (11.113). 11.8.7 Complete the details of Example 11.8.8 by verifying that there is no contribution to the contour integral from either the small or the large circles of the contour, and that Eq. (11.115) simplifies to the result given as (11.116). 11.8.8 Evaluate1Z 1cosbxcosax x2dx;a>b>0: ANS..ab/: 11.8.9 Prove that1Z 1sin2x x2dxD 2: Hint. sin2xD1 2.1cos 2 x/. 11.8.10 Show that1Z 0xsinx x2C1dxD 2e: ArfKen_Ch11-9780123846549.tex 540 Chapter 11 Complex Variable Theory 11.8.11 A quantum mechanical calculation of a transition probability leads to the function f.t;!/D2.1cos!t/=!2. Show that 1Z 1f.t;!/d!D2t: 11.8.12 Show that.a>0/: (a)1Z 1cosx x2Ca2dxD aea: How is the right side modified if cosxis replaced by coskx? (b)1Z 1xsinx x2Ca2dxDea: How is the right side modified if sinxis replaced by sinkx? 11.8.13 Use the contour shown (Fig. 11.28) with R!1 to prove that 1Z 1sinx xdxD: 11.8.14 In the quantum theory of atomic collisions, we encounter the integral ID1Z 1sint teiptdt; R R.R R FIGURE 11.28 Contour for Exercise 11.8.13. ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 541 in which pis real. Show that ID0;jpj>1 ID;jpj<1: What happens if pD1 ? 11.8.15 Show that1Z 0dx .x2Ca2/2D 4a3;a>0: 11.8.16 Evaluate1Z 1x2 1Cx4dx: ANS.=p 2. 11.8.17 Evaluate1Z 0xplnx x2C1dx;0<p<1: ANS.2 4sin. p=2/ cos2.p=2/: 11.8.18 Evaluate1Z 0.lnx/2 1Cx2dx; (a) by appropriate series expansion of the integrand to obtain 41X nD0.1/n.2nC1/3; (b) and by contour integration to obtain3 8: Hint. x!zDet. Try the contour shown in Fig. 11.29, letting R!1 . −R+iπ −RR+iπ y Rx FIGURE 11.29 Contour for Exercise 11.8.18. ArfKen_Ch11-9780123846549.tex 542 Chapter 11 Complex Variable Theory 11.8.19 Prove that1Z 0ln.1Cx2/ 1Cx2dxDln 2: 11.8.20 Show that 1Z 0xa .xC1/2dxDa sina; where1<a<1. Hint. Use the contour shown in Fig. 11.26, noting that zD0is a branch point and the positive x-axis can be chosen to be a cut line. 11.8.21 Show that 1Z 1x2dx x42x2cos 2C1D 2 sinD 21=2.1cos 2/1=2: Exercise 11.8.16 is a special case of this result. 11.8.22 Show that 1Z 0dx 1CxnD=n sin.= n/: Hint. Try the contour shown in Fig. 11.30, with D2=n. 11.8.23 (a) Show that f.z/Dz42z2cos 2C1 has zeros at ei;ei;ei, andei. (b) Show that 1Z 1dx x42x2cos 2C1D 2 sinD 21=2.1cos 2/1=2: Exercise 11.8.22 .nD4/is a special case of this result. R Rθ FIGURE 11.30 Sector contour. ArfKen_Ch11-9780123846549.tex 11.8 Evaluation of De/f_inite Integrals 543 11.8.24 Show that 1Z 0xa xC1dxD sina; where 0<a<1. Hint. You have a branch point and you will need a cut line. Try the contour shown in Fig. 11.26. 11.8.25 Show that1Z 0cosh bx cosh xdxD 2 cos. b=2/;jbj<1: Hint. Choose a contour that encloses one pole of cosh z. 11.8.26 Show that 1Z 0cos.t2/dtD1Z 0sin.t2/dtDp 2p 2: Hint. Try the contour shown in Fig. 11.30, with D=4. Note. These are the Fresnel integrals for the special case of infinity as the upper limit. For the general case of a varying upper limit, asymptotic expansions of the Fresnel integrals are the topic of Exercise 12.6.1. 11.8.27 Show that1Z 01 .x2x3/1=3dxD2=p 3. Hint. Try the contour shown in Fig. 11.31. 11.8.28 Evaluate1Z 1tan1ax dx x.x2Cb2/, foraandbpositive, with ab<1. Explain why the integrand does not have a singularity at xD0. 01 FIGURE 11.31 Contour for Exercise 11.8.27. ArfKen_Ch11-9780123846549.tex 544 Chapter 11 Complex Variable Theory Hint. Try the contour shown in Fig. 11.32, and use Eq. (1.137) to represent tan1az. After cancellation, the integrals on segments BandB0combine to give an elementary integral. 11.9 E VALUATION OF SUMS The fact that the cotangent is a meromorphic function with regularly spaced poles, all with the same residue, enables us to use it to write a wide variety of infinite summations in terms of contour integrals. To start, note that cotzhas simple poles at all integers on the real axis, each with residue limz!ncosz sinzD1: Suppose that we now evaluate the integral INDI CNf.z/cotz dz; where the contour is a circle about zD0of radius NC1 2(thereby not passing close to the singularities of cotz). Assuming also that f.z/has only isolated singularities, at points zjother than real integers, we get by application of the residue theorem (see also Exercise 11.9.1), IND2iNX nDNf.n/C2iX j(residues of f.z/cotzat singularities zjoff). This integral over the circular contour CNwill be negligible for large jzjifz f.z/!0at largejzj.8When that condition is met, limN!1IND0, and we have the useful result 1X nD1f.n/DX j(residues of f.z/cotzat singularities zjoff). (11.123) The condition required of f.z/will usually be satisfied if the summation of Eq. (11.123) converges. AibB′B xy i a FIGURE 11.32 Contour for Exercise 11.8.28. 8See also Exercise 11.9.2. ArfKen_Ch11-9780123846549.tex 11.9 Evaluation of Sums 545 Example 11.9.1 EVALUATING A SUM Consider the summation SD1X nD11 n2Ca2; where, for simplicity, we assume that ais nonintegral. To bring our problem to the form we know how to treat, we note that also 1X nD11 n2Ca2DS; so that 1X nD11 n2Ca2D2SC1 a2; (11.124) where we have added on the right-hand side the contribution from nD0that was not included in S. The summation is now identified as of the form of Eq. (11.123), with f.z/D1=.z2C a2/;f.z/approaches zero at large zrapidly enough to make Eq. (11.123) applicable. We therefore proceed to the observation that the only singularities of f.z/are simple poles at zDia. The residues we need are those of cot. z/=.z2Ca2/; they are cotia 2iaDcotha 2aandcot. ia/ 2iaDcoth. a/ 2a: These are equal, so from Eqs. (11.123) and(11.124), 2SC1 a2Dcotha a; which we easily solve to reach SDcotha 2a1 2a2:  Additional types of summations can be performed if we replace cotzby functions with other regularly repeating patterns of residues. For example, csczhas residues for integer zthat alternate in sign between C1and1;tanzhas residues that are all C1, but occur at the points nC1 2. Andseczhas residues1at the half-integers with a sign alternation. For convenience, we list in Table 11.2 the contour-integral formulas for the four types of summations we have just discussed. We close this section with another example, this time illustrating what can be done if f.z/has a pole at an integer value of z. Example 11.9.2 ANOTHER SUM Consider now the summation SD1X nD11 n.nC1/: ArfKen_Ch11-9780123846549.tex 546 Chapter 11 Complex Variable Theory Table 11.2 Contour-Integral-Based Formulas for Summations Summation Formula 1X nD1f.n/X (residues of f.z/cotzat singularities of f). 1X nD1.1/nf.n/nX (residues of f.z/csczat singularities of f). 1X nD1f nC1 2 X (residues of f.z/tanzat singularities of f). 1X nD1.1/nf nC1 2X (residues of f.z/seczat singularities of f). To extend the summation to nD1 , we note that SD2X nD11 n.nC1/;so that 2SD1X nD101 n.nC1/; (11.125) where the prime on the sum indicates that the terms for nD0andnD1 are to be omit- ted. The derivation of Eq. (11.123) indicates that this equation will apply if we omit the (singular) nD0andnD1 terms from the sum and include the points zD0andzD1 as points where the residues of f.z/cotzare to be included. Based on that insight, we find that in the present problem, 2SD(sum of residues of cotz=z.zC1/atzD0andzD1 ). The singularities at zD0andzD1 are second-order poles, at which the residues are most easily computed by the method illustrated in item 5 of Example 11.7.1. In Exer- cise 11.7.2 it is shown that the residue at each pole has value 1. Completing the problem, 2SD.11/D2; soSD1: In this instance the result is easily verified by making the partial fraction expansion 1 n.nC1/D1 n1 nC1: When inserted in the summation S, all terms cancel except the initial term of the 1=n summation, yielding SD1.  Exercises 11.9.1 Show that if f.z/is analytic at zDz0andg.z/has a simple pole at zDz0with residue b0, then f.z/g.z/also has a simple pole at zDz0, with residue f.z0/b0. 11.9.2 Show that cotzhas magnitude of order 1 for large jzjwhen not extremely close to one of its poles and does not affect the limiting behavior of IN. ArfKen_Ch11-9780123846549.tex 11.10 Miscellaneous Topics 547 11.9.3 Evaluate1 131 33C1 53: 11.9.4 EvaluateP1 nD11 n.nC2/: 11.9.5 EvaluateP1 nD1.1/n .nCa/2;where ais real and not an integer. 11.9.6 (a) Using a method based on contour integration, evaluateP1 nD01 .2nC1/2: (b) Check your work by relating your answer to an appropriate expression involving zeta functions. 11.9.7 Show that1 cosh.=2/1 3 cosh.3=2/C1 5 cosh.5=2/D 8: 11.9.8 For'C , show thatP1 nD1.1/nsinn' n3D' 12.'22/: 11.10 M ISCELLANEOUS TOPICS Schwarz Reflection Principle Our starting point for this topic is the observation that g.z/D.zx0/nfor integral nand realx0satisfies g.z/DT.zx0/nUD.zx0/nDg.z/: (11.126) A generalization of the result in Eq. (11.126) is the Schwarz reflection principle: If a function f.z/is (1) analytic over some region including a portion of the real axis and (2) real when zis real, then f.z/Df.z/: (11.127) Expanding f.z/about some point x0within the region of analyticity on the real axis, f.z/D1X nD0.zx0/nf.n/.x0/ nW: Since f.z/is analytic at zDx0, this Taylor expansion exists. Since f.z/is real when zis real, f.n/.x0/must be real for all n. Then, invoking Eq. (11.126), the Schwarz reflection principle, Eq. (11.127), follows immediately. This completes the proof within a circle of convergence. Analytic continuation then permits the extension of this result to the entire region of analyticity. Note that the reflection principle can also be derived by the consideration of Laurent expansions. See Exercise 11.10.2. Mapping An analytic function w.z/Du.x;y/Civ.x;y/can be regarded as a mapping in which points or curves in an xyplane can be associated with the corresponding points or curves in auvplane. As a relatively simple example, consider the transformation wD1=z. From an ArfKen_Ch11-9780123846549.tex 548 Chapter 11 Complex Variable Theory examination of its polar form, with zDrei,wDei', we see that D1=rand'D , leading to the conclusion that the interior of the unit circle maps into its exterior (see Fig. 11.33). Circles in other locations in the zplane are transformed by wD1=zinto other circles (or straight lines, which can be thought of as circles of infinite radius). This statement is the subject of Exercise 11.10.6. The transformation of two such circles are shown in the four panels of Fig. 11.34. Compare the way in which the interiors of the circles transform in Figs. 11.33 and11.34. Note that the transformation does not preserve lengths, as can be seen in the figure from the labeling of various points and their locations when mapped. 1/rr -θθ FIGURE 11.33 MappingwD1=z. The shaded areas transform into each other. y x xa b3 1v ub a1 b=i da cy d d b=−iuv a c1 3 FIGURE 11.34 Left panels: circles in zplane. Right panels: their transformations inwplane underwD1=z. ArfKen_Ch11-9780123846549.tex 11.10 Miscellaneous Topics 549 Historically, the notion of mapping was useful for identifying and carrying out transfor- mations that would facilitate the solution of 2-D problems in electrostatics, fluid dynamics, and other areas of classical physics. An important aspect of such mappings is that they are conformal, meaning that (except at singularities of the transformation) the angles at which curves intersect remain unchanged when transformed. This feature preserves relations, e.g., between equipotentials and lines of force (stream lines). With the nearly universal use of high-speed computers, procedures based on conformal mapping are no longer central to the practical solution of most physics and engineering problems, and as a consequence will not be explored here in further detail. For problems where these techniques are still relevant, we refer the reader to earlier editions of this book and to sources identified under Additional Readings. In that connection, we call particular attention to the book by Spiegel, which contains (in chapter 8) descriptions of a large number of mappings and (in chapter 9) many applications to problems of fluid flow, electrostatics, and heat conduction. Exercises 11.10.1 A function f.z/Du.x;y/Civ.x;y/satisfies the conditions for the Schwarz reflection principle. Show that (a) uis an even function of y. (b)vis an odd function of y. 11.10.2 A function f.z/can be expanded in a Laurent series about the origin with the coeffi- cients anreal. Show that the complex conjugate of this function of zis the same function of the complex conjugate of z; that is, f.z/Df.z/: Verify this explicitly for (a) f.z/Dzn;nan integer. (b) f.z/Dsinz. Iff.z/Diz.a1Di/, show that the foregoing statement does not hold. 11.10.3 The function f.z/is analytic in a domain that includes the real axis. When zis real .zDx/,f.x/is pure imaginary. (a) Show that f.z/DT f.z/U: (b) For the specific case f.z/Diz, develop the Cartesian forms of f.z/,f.z/, and f.z/. Do not quote the general result of part (a). 11.10.4 How do circles centered on the origin in the z-plane transform for .a/ w 1.z/DzC1 z; .b/ w 2.z/Dz1 z;forz6D0? What happens when jzj!1? ArfKen_Ch11-9780123846549.tex 550 Chapter 11 Complex Variable Theory 11.10.5 What part of the z-plane corresponds to the interior of the unit circle in the w-plane if (a)wDz1 zC1?(b)wDzi zCi? 11.10.6 (a) Writing zDxCiy,wDuCiv, show that if wD1=z, the circle in the xyplane defined by.xa/2C.yb/2Dr2transforms into .uA/2C.vB/2DR2. (b) Does the center of the circle in the zplane transform into the center of the corre- sponding circle in the wplane? 11.10.7 Assume that a curve in the xyplane passes through point z0in the direction dzDeids, where sindicates arc length on the curve. Then, if wDf.z/, with f.z/analytic at zD z0, we have dwD.dw=dz/dzDf0.z/eids, where dwis in the direction the mapping of the xycurve passes through w0Df.z0/in thewplane. Use this observation to prove that if f0.z0/6D0, the angle at which two curves intersect in the zplane is the same (both in magnitude and direction) as the angle of intersection of their mappings in thewplane. Additional Readings Ahlfors, L. V., Complex Analysis, 3rd ed. New York: McGraw-Hill (1979). This text is detailed, thorough, rigor- ous, and extensive. Churchill, R. V., J. W. Brown, and R. F. Verkey, Complex Variables and Applications, 5th ed. New York: McGraw-Hill (1989). This is an excellent text for both the beginning and advanced student. It is readable and quite complete. A detailed proof of the Cauchy-Goursat theorem is given in Chapter 5. Greenleaf, F. P., Introduction to Complex Variables. Philadelphia: Saunders (1972). This very readable book has detailed, careful explanations. Kurala, A., Applied Functions of a Complex Variable. New York: Wiley (Interscience) (1972). An intermediate- level text designed for scientists and engineers. Includes many physical applications. Levinson, N., and R. M. Redheffer, Complex Variables. San Francisco: Holden-Day (1970). This text is written for scientists and engineers who are interested in applications. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics. New York: McGraw-Hill (1953). Chapter 4 is a presentation of portions of the theory of functions of a complex variable of interest to theoretical physicists. Remmert, R., Theory of Complex Functions. New York: Springer (1991). Sokolnikoff, I. S., and R. M. Redheffer, Mathematics of Physics and Modern Engineering, 2nd ed. New York: McGraw-Hill (1966). Chapter 7 covers complex variables. Spiegel, M. R., Complex Variables, in Schaum’s Outline Series. New York: McGraw-Hill (original 1964, reprinted 1995). An excellent summary of the theory of complex variables for scientists. Titchmarsh, E. C., The Theory of Functions, 2nd ed. New York: Oxford University Press (1958). A classic. Watson, G. N., Complex Integration and Cauchy’s Theorem. New York: Hafner (original 1917, reprinted 1960). A short work containing a rigorous development of the Cauchy integral theorem and integral formula. Appli- cations to the calculus of residues are included. Cambridge Tracts in Mathematics, and Mathematical Physics, No. 15. ArfKen_15-ch12-0551-0598- 9780123846549.tex CHAPTER 12 FURTHER TOPICS IN ANALYSIS The broader perspective and additional tools made available through complex variable theory enable us to consider fruitfully a number of topics in analysis that have wide appli- cation in areas of relevance to physics. In this chapter we survey several such topics. 12.1 O RTHOGONAL POLYNOMIALS Many physical problems lead to second-order differential equations corresponding to Sturm-Liouville problems, and often the solutions of interest in physics are polynomials, defined on a range and with weighting factors that make them eigenfunctions of Hermitian problems. A number of interesting features of such problems can be approached with the aid of complex variable theory. Rodrigues Formulas Odile Rodrigues showed that a large class of second-order Sturm-Liouville ordinary dif- ferential equations (ODEs) had polynomial solutions which could be put in a compact and useful form now generally called a Rodrigues formula. While such formulas could be pre- sented case by case with an aura of coincidence or mystery, the approach we take here is to develop them from a general viewpoint, after which we can proceed to more detailed discussion of well-known special cases. Consider a second-order Sturm-Liouville ODE of the general form p.x/y00Cq.x/y0CyD0; (12.1) 551 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_15-ch12-0551-0598- 9780123846549.tex 552 Chapter 12 Further Topics in Analysis with p.x/andq.x/restricted to the polynomial forms p.x/D x2C xC ; q.x/DxC: (12.2) The forms of pandqare sufficiently general to include most of the ODEs with classi- cal sets of polynomials as solutions (the Legendre, Hermite, and Laguerre ODEs, among others). When Eq. (12.1) has as a solution a polynomial of degree n, we can write yn.x/DnX jD0gjxj; (12.3) with coefficient gnnonzero. Setting to zero the coefficient of xnwhen ynis inserted into the ODE, we have n.n1/ gnCngnCgnD0; (12.4) showing that the eigenvalue nwhich corresponds to ynmust have the value nDn.n1/ n: (12.5) In Chapter 7 we identified an ODE of the form of Eq. (12.1) as self-adjoint if p0.x/D q.x/, and also showed that if an ODE was not already self-adjoint as written, it could be converted to self-adjoint form by multiplying all its terms by a weight factor w.x/, which must be such that .wp/0Dwq;orw0Dwqp0 p: (12.6) As shown previously, this equation is separable and has solution w.x/Dp1exp0 @xZq.x/ p.x/dx1 A: (12.7) The introduction of wenables the ODE to assume the form d dx w.x/p.x/y0 Cw.x/yD0; (12.8) which was useful for discussing orthogonality properties of its solutions. Our current interest in w.x/, however, is in the observation by Rodrigues that its par- ticular form permits the solutions yn.x/to be written in the compact and interesting form that is now called its Rodrigues formula: yn.x/D1 w.x/d dxn wp.x/n : (12.9) The proof of Eq. (12.9) is both simple and ingenious. Using the defining condition for w.x/,Eq. (12.6), we first obtain p wpn0Dwpn .n1/p0Cq : (12.10) ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.1 Orthogonal Polynomials 553 We then differentiate this equation nC1times and divide by w. Because pis only quadratic in xandqis linear, application of Leibniz’s formula to the multiple differen- tiations leads to only three terms on the left-hand side and two on the right: p wd dxnC2 wpn C.nC1/p0 wd dxnC1 wpn Cn.nC1/p00 2wd dxn wpn D.n1/p0Cq wd dxnC1 wpn C.nC1/T.n1/p00Cq0U wd dxn wpn : (12.11) Our objective is to manipulate Eq. (12.11) into a form showing that ynas given in Eq. (12.9) is a solution to the ODE of Eq. (12.1). We start by identifying the terms with yn where that is possible, and then, combining or canceling similar terms, we reach p wd dxnC2 wpn C2p0q wd dxnC1 wpn n2n2 2p00C.nC1/q0 ynD0: (12.12) To complete our analysis we now need to move the factors 1=w so that only ndifferenti- ations appear to their right, enabling identification of the remaining terms of the equation with ynor its derivatives. We note the identity p wd dxnC2 wpn Dp1 wd dxn wpn00 2pdw1 dxd dxnC1 wpn pd2w1 dx2d dxn wpn ; which reduces, using Eq. (12.6), to p wd dxnC2 wpn Dpy00 nC2.qp0/ wd dxnC1 wpn  p00q0qp0 p yn: (12.13) Substituting Eq. (12.13) into Eq. (12.12), some further simplification results: py00 nCq wd dxnC1 wpn n2n 2p00Cnq0q.qp0/ p ynD0: (12.14) Our final step is to use the identity q wd dxnC1 wpn Dqy0 nq.qp0/ pyn; (12.15) which brings us to py00 nCqy0 nn2n 2p00Cnq0 ynD0: (12.16) ArfKen_15-ch12-0551-0598- 9780123846549.tex 554 Chapter 12 Further Topics in Analysis Noting that p00D2 andq0D, we confirm that ynis a solution of Eq. (12.1) with the eigenvalue given in Eq. (12.5). Finally, we need to show that Rodrigues’ formula, Eq. (12.9), results in an expression that is a polynomial of degree n. We note that a typical term of that formula will in- volve a j-fold differentiation of wand an.nj/-fold differentiation of pn. After the differentiation of pn, we are left with pjtimes a polynomial. The differentiation of w will, applying Eq. (12.6), leave .w=pj/times a polynomial, and the numerator and de- nominator factors pjcancel. In addition, the wfrom the differentiation cancels against the initial factor w1, leaving each term of ynin polynomial form. When all terms of yn are combined, the resulting polynomial must have the degree consistent with Eq. (12.5), namely n. Example 12.1.1 RODRIGUES FORMULA FOR HERMITE ODE The Hermite ODE is y002xy0CyD0; orpy00Cqy0CyD0 with pD1,qD2 x. We easily find wDexp0 @xZ .2x/dx1 ADex2: The Rodrigues formula is therefore (with a factor .1/nto obtain the Hermite polynomials with their conventional signs) yn.x/D.1/n wd dxn wpn D.1/nex2d dxn ex2: (12.17)  Schlaefli Integral One of the nice features of the Rodrigues formulas is that the multiple differentiations can be converted to a convenient form by use of Cauchy’s integral formula. Using Eq. (11.33), we have yn.x/D1 w.x/nW 2iI Cw.z/Tp.z/Un .zx/nC1dz; (12.18) where the contour Cencloses the point x, and must be such that w.z/Tp.z/Unis analytic everywhere on and within C. This formula is known as the Schlaefli integral foryn.x/. It is possible to introduce the Schlaefli integral as the definition of a set of functions yn and, from that definition, prove that ynis a solution to the corresponding ODE. Because we created the Schaefli integral to represent a function already known to be a solution, verification that it solves the ODE becomes redundant. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.1 Orthogonal Polynomials 555 Generating Functions Many sets of functions arising in mathematical physics can be defined in terms of gener- ating functions. Such functions include, but are not limited to the orthogonal polynomials ynthat have been the subject of our discussion of Rodrigues formulas. For now, we make no assumptions as to the source of the functions involved. Iffn.x/is a set of functions, defined for integer values of the index n, it may be the case that the fn.x/can be described as the coefficients of the powers of an auxiliary variable, t, in the expansion of a function g.x;t/, which is called a generating function: g.x;t/DX ncnfn.x/tn: (12.19) The range of nmay be semi-infinite, with n0, thereby describing a Taylor series, or it may extend from1 toC1, thus describing a Laurent series. The additional coefficient, cn, permits adjustment of the function set to an agreed-upon scaling. Different choices of cnwill also lead to different generating functions g.x;t/for the same set of fn. Applying the residue theorem, we can see that the generating function expansion is closely related to contour integral representations of the functions fn: cnfn.x/D1 2iIg.x;t/ tnC1dt; (12.20) where the contour encircles tD0but no other singularities of the integrand (with respect tot). A generating function may be regarded as providing the definition of a function set fn.x/, or alternatively it may have been obtained as the encapsulation of the fnwhich were already defined in some other way (e.g., as polynomial solutions of a Sturm-Liouville ODE). We shall later take up the issue of obtaining generating functions for previously specified fn, focusing for now only on ways in which they can be used. It is obvious that by explicitly evaluating the implied expansion one can extract the members of a function set from its generating function. However, a more important feature of generating functions is that they can be very useful in deriving relationships between members of the set fn. For example, @g.x;t/ @tDX nncnfn.x/tn1DX n.nC1/cnC1fnC1.x/tn; and if we can relate gand@g=@twe have a corresponding relation between fnandfnC1. Relations between the fn.x/and their derivatives f0 n.x/can be deduced by differentiating g.x;t/with respect to x. Example 12.1.2 HERMITE POLYNOMIALS A generating function formula for the Hermite polynomials Hn.x/(at their conventional scaling) is et2C2txD1X nD0Hn.x/tn nW: (12.21) ArfKen_15-ch12-0551-0598- 9780123846549.tex 556 Chapter 12 Further Topics in Analysis To develop a recurrence formula connecting Hnof contiguous index values, we compute @ @tet2C2txD.2x2t/et2C2txD1X nD0nHn.x/tn1 nW: (12.22) Expanding the exponential in the central member of Eq. (12.22) (and suppressing tem- porarily the argument of Hn), 1X nD02x Hntn nW1X nD02HntnC1 nWD1X nD0nHntn1 nW: Extracting the coefficient of tnfrom each of these summations, we reach (for each n) 2x Hn nW2Hn1 .n1/WD.nC1/HnC1 .nC1/W; which reduces to 2x Hn.x/2nH n1.x/DHnC1.x/: (12.23) Equation (12.23) is called a recurrence formula; it permits the construction of the en- tire series of Hnfrom starting values (typically H0andH1, which are easily computed directly). A derivative formula can be obtained by differentiating Eq. (12.21) with respect to x. We have @ @xet2C2txD2tet2C2txD1X nD0H0 n.x/tn nW: Substituting Eq. (12.21) into the central member of this equation, we get 1X nD02Hn.x/tnC1 nWD1X nD0H0 n.x/tn nW; which leads directly to 2nH n1.x/DH0 n.x/: (12.24)  In later chapters we illustrate the application of these ideas to a variety of special func- tions; in the next section of this chapter we apply them to a generating function that leads to quantities known as Bernoulli numbers. Finding Generating Functions To take generating functions out of the realm of magic, we next consider how they might be obtained. For a more or less arbitrary function set, this question has been a topic of current interest in mathematical research, with methods of several sorts devised during the past century by Rainville, Weisner, Truesdell, and others. See the works by McBride and Talman in Additional Readings. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.1 Orthogonal Polynomials 557 For sets of polynomials arising in Sturm-Liouville problems and described by Rodrigues formulas, we can be more explicit. Using the Schlaefli integral, Eq. (12.18), we can form g.x;t/D1 w.x/1X nD0cntnnW 2iI Cw.z/Tp.z/Un .zx/nC1dz: (12.25) Recall that Cencloses xand thatwpnmust be analytic throughout the region within the contour. In principle Eq. (12.25) can be evaluated to obtain g.x;t/, for example by choosing C to be such that the summation can be brought inside the zintegral and (after specifying cn) evaluating first the sum and then the contour integral. In practice the difficulty of doing this may depend on the problem, including the choice of cn. We provide one example of the process. Example 12.1.3 LEGENDRE POLYNOMIALS We use the formal process described above to obtain a generating function for the Legendre polynomials. The Legendre ODE is of the form discussed in Eq. (12.1), .1x2/y002xy0CyD0; implying that p.x/D1x2;q.x/D2 x; and the equation is, as written, self-adjoint, so w.x/D1. From the generating-function formula based on the Schlaefli integral, Eq. (12.25), we choose cnD.1/n=2nnW, thereby reaching g.x;t/D1X nD0.1/ntn 2nnWnW 2iI C.1z2/n .zx/nC1dz: Interchanging the summation and integration (which we will justify later), the factors de- pendent on nform a geometric series, which we can sum: 1X nD0.z21/t 2.zx/n1 zxD1 zx1 2.z21/t D2 t z22z tC2xt t1 : Inserting this result into the formula for g.x;t/, we now have g.x;t/D2 t1 2iI C z22z tC2xt t1 dz D2 t1 2iI Cdz .zz1/.zz2/; (12.26) ArfKen_15-ch12-0551-0598- 9780123846549.tex 558 Chapter 12 Further Topics in Analysis where z1andz2are the roots of the quadratic form in the first line of the equation: z1D1 tp 12xtCt2 t;z2D1 tCp 12xtCt2 t: In order for Eq. (12.26) to be valid, it must have been legitimate to interchange the sum- mation and integration, which is the case only if the summation is uniformly convergent (with respect to z) for all points at which it is used (i.e., everywhere on the contour C). It is convenient to analyze the convergence for small tandxand for a contour with jzjD1. Once a final formula has been obtained, its range of validity can be extended by appeal to analytic continuation. On the assumed contour and for small x, there will be a range of jtj1for which .z21/t 2.zx/ <1; guaranteeing convergence of the geometric series. We now return to the evaluation of the contour integral in Eq. (12.26). It has two poles, at zDz1andzDz2. For small xandjtj, z2will be approximately 2=tand will be exterior to the contour, while z1will be close to the origin of z. Thus, only the residue of the integrand at zDz1will contribute to the contour integral, which will have the value g.x;t/D2 t1 z1z2: Since z1z2D2 tp 12xtCt2; we obtain the Legendre polynomial generating function as g.x;t/D1p 12xtCt2: (12.27)  Summary—Orthogonal Polynomials For five classical sets of orthogonal polynomials, we summarize in Table 12.1 their ODEs, Rodrigues formulas, and generating functions. Omitted from the list are important sub- sidiary polynomial sets (e.g., those connected with the associated Legendre and associated Laguerre ODEs). Exercises 12.1.1 Starting from the Rodrigues formula in Table 12.1 for the Hermite polynomials Hn, derive the generating function for the Hngiven in that table. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.1 Orthogonal Polynomials 559 Table 12.1 Orthogonal Polynomials: ODEs, Rodrigues Formulas, and Generating Functions Rodrigues Formula Generating Function Legendre:.1x2/y002xy0Cn.nC1/yD0 Pn.x/D1 2nnWd dxn .x21/n.12xtCt2/1=2D1X nD0Pn.x/tn Hermite: y002xy0C2nyD0 Hn.x/D.1/nex2d dxn ex2et2C2xtD1X nD01 nWHn.x/tn Laguerre: xy00C.1x/y0CnyD0 Ln.x/Dex nWd dxn xnex ext=.1t/ 1tD1X nD0Ln.x/tn Chebyshev I: .1x2/y00xy0Cn2yD0 Tn.x/D.1/n.1x2/1=2 .2n1/WWd dxn .1x2/n1=2 1t2 12xtCt2DT0.x/C21X nD1Tn.x/tn Chebyshev II: .1x2/y003xy0Cn.nC2/yD0 Un.x/D.1/n.nC1/ .2nC1/WW.1x2/1=2d dxn .1x2/nC1=2 1 12xtCt2D1X nD0Un.x/tn 12.1.2 (a) Starting from the Laguerre ODE, xy00C.1x/y0CyD0; obtain the Rodrigues formula for its polynomial solutions Ln.x/. (b) From the Rodrigues formula, scaled as in Table 12.1, derive the generating func- tion for the Ln.x/given in that table. 12.1.3 Carry out in detail the steps needed to confirm that the .nC1/-fold differentiation of Eq. (12.10) leads to Eq. (12.12). 12.1.4 Confirm the algebraic steps that convert Eq. (12.12) intoEq. (12.16). 12.1.5 Given the following integral representations, in which the contours encircle the origin but no other singular points, derive the corresponding generating functions: (a) Bessel functions: Jn.x/D1 2iI e.x=2/.t1=t/tn1dt: (b) Modified Bessel functions: In.x/D1 2iI e.x=2/.tC1=t/tn1dt: ArfKen_15-ch12-0551-0598- 9780123846549.tex 560 Chapter 12 Further Topics in Analysis 12.1.6 Expand the generating function for the Legendre polynomials, .12tzCt2/1=2, in powers of t. Assume that tis small. Collect the coefficients of t0;t1, and t2. ANS:a0DP0.z/D1; a1DP1.z/Dz; a2DP2.z/D1 2.3z21/: 12.1.7 The set of Chebyshev polynomials usually denoted Un.x/has the generating-function formula 1 12xtCt2D1X nD0Un.x/tn: Derive a recurrence formula (for integer n0) connecting three Unof consecutive n. 12.2 B ERNOULLI NUMBERS A generating-function approach is a convenient way to introduce the set of numbers first used in mathematics by Jacques (James, Jacob) Bernoulli. These quantities have been de- fined in a number of different ways, so extreme care must be taken in combining formulas from works by different authors. Our definition corresponds to that used in the reference work Handbook of Mathematical Functions (AMS-55). See Additional Readings. Since the Bernoulli numbers, denoted Bn, do not depend on a variable, their generating function depends only on a single (complex) variable, and the generating-function formula has the specific form t et1D1X nD0Bntn nW: (12.28) The inclusion of the factor 1=nWin the definition is just one of the ways some definitions of Bernoulli numbers differ. We defer for the moment the important question as to the circle of convergence of the expansion in Eq. (12.28). Since Eq. (12.28) is a Taylor series, we may identify the Bnas successive derivatives of the generating function: BnDdn dtnt et1 tD0: (12.29) To obtain B0, we must take the limit of t=.et1/ast!0, easily finding B0D1. Applying Eq. (12.29), we also have B1Dd dtt et1 tD0Dlim t!01 et1tet .et1/2 D1 2: (12.30) ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.2 Bernoulli Numbers 561 In principle we could continue to obtain further Bn, but it is more convenient to proceed in a more sophisticated fashion. Our starting point is to examine 1X nD2Bntn nWDt et1B0B1tDt et11Ct 2 Dt et11t 2; (12.31) where we have used the fact that t et1Dt et1t: (12.32) Equation (12.31) shows that the summation on its left-hand side is an even function of t, leading to the conclusion that all Bnof odd n(other than B1) must vanish. We next use the generating function to obtain a recursion relation for the Bernoulli numbers. We form et1 tt et1D1D"1X mD0tm .mC1/W#" 1t 2C1X nD1B2nt2n .2n/W# D1C1X mD1tm1 .mC1/W1 2mW C1X ND2tNN=2X nD1B2n .2n/W.N2nC1/W D1C1X ND2tN .NC1/W2 4N1 2CN=2X nD1NC1 2n B2n3 5: (12.33) Since the coefficient of each power of tin the final summation of Eq. (12.33) must vanish, we may set to zero for each Nthe expression in its square brackets. Changing N, if even, to2Nand if odd, to 2N1,Eq. (12.33) leads to the pair of equations N1 2DNX nD12NC1 2n B2n; N1DN1X nD12N 2n B2n:(12.34) Either of these equations can be used to obtain the B2nsequentially, starting from B2. The first few Bnare listed in Table 12.2. To obtain additional relations involving the Bernoulli numbers, we next consider the following representation of cott: cottDcost sintDieitCeit eiteit Die2itC1 e2it1 Di 1C2 e2it1 : ArfKen_15-ch12-0551-0598- 9780123846549.tex 562 Chapter 12 Further Topics in Analysis Table 12.2 Bernoulli Numbers n B n Bn 0 1 1:000000000 11 20:500000000 21 60:166666667 41 300:033333333 61 420:023809524 81 300:033333333 105 660:0757 57576 Note. Further values are given in AMS-55; see Abramowitz in Additional Readings. Multiplying by tand rearranging slightly, tcottD2it 2C2it e2it1D1X nD0B2n.2it/2n .2n/W D1X nD0.1/nB2n.2t/2n .2n/W; (12.35) where the term 2it=2has canceled the B1term that would otherwise appear in the expan- sion. Now that we have our Bernoulli-number expansion identified with tcott, we can see that it represents a function with singularities (poles) at tDm, where mD1 ,2;::: . There is no singularity at tD0(due to the presence of the factor t), so the singularity nearest the expansion point (the origin) is at jtjD. Since the argument in the expansion is2t, we conclude that the generating series for the Bernoulli numbers, Eq. (12.28), will have the radius of convergence j2tjD2. This observation is, of course, consistent with the fact that the zeros of et1are for tat integer multiples of 2i. To obtain another representation of the Bernoulli numbers, we write Bnusing the contour-integration formula, Eq. (12.20). Noting that for use in this equation cnfn.x/D Bn=nW, we have BnDnW 2iIt et1dt tnC1; (12.36) where the integral is a circle within the radius of convergence of the generating series. We can, at least in principle, evaluate the integral using the residue theorem. For nD0we have a simple pole with a residue of C1, and B0D0W 2i2i.C1/D1: ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.2 Bernoulli Numbers 563 −2πi −4πi −8πi−6πi2πi4πi6πi8πi RC xy A A¢ FIGURE 12.1 Contour of integration for Bernoulli numbers. FornD1the singularity at tD0becomes a second-order pole, and the limiting process prescribed by Eq. (11.68) yields the residue 1 2, so B1D1W 2i2i 1 2 D1 2; consistent with our previous result. For n2the poles at tD0are of increasing order and this procedure becomes rather tedious, so we resort to a different approach. We deform the contour of our integral representation as shown in Fig. 12.1, which differs from the original circular contour in that it surrounds all the poles of the integrand at tD2 mi, mD1;2;::: , while avoiding the inclusion of the pole at tD0. In contrast to the high-order pole at tD0, the other poles are all first-order, with residues that are easily evaluated. To use the new contour, we need to identify the contributions from its constituent parts. The direction of travel around the contour causes the small circle about tD0to contribute C2 itimes the residue of the integrand at tD0, i.e., the result that when multiplied by nW=2iis equal to Bn. The remainder of the contour makes no contribution to the integral: (1) Because the integrand is analytic along the real axis and there is no branch cut there, the segments AandA0, which are in opposite directions of travel, cancel; and (2) the large circle contributes negligibly (for n2) because at largejtjthe integrand behaves asymptotically as 1=jtjn. Noting that the poles at nonzero tare encircled in a clockwise sense, we have the following relatively simple result (for n2): BnDnW 2iX 2i residues oftn et1at poles t6D0 : (12.37) Since the residue at tD2miis simply.2mi/n, Eq. (12.37) becomes BnDnW .2i/n1X mD11 mnC1 .m/n ; ArfKen_15-ch12-0551-0598- 9780123846549.tex 564 Chapter 12 Further Topics in Analysis which further reduces, for 2n2, to B2nD.1/nC1.2n/W .2/2n1X mD12 m2nD.1/nC12.2n/W .2/2n.2n/; B2nC1D0:(12.38) Note that the Bnof odd n>1are correctly shown to vanish, and that the Bernoulli numbers of even n>0are identified as proportional to Riemann zeta functions, which first appeared in this book at Eq. (1.12). We repeat the definition: .z/D1X mD11 mz: Equation (12.38) is an important result because we already have a straightforward way to obtain values of the Bn, via Eq. (12.34), and Eq. (12.38) can be inverted to give a closed expression for .2n/, which otherwise was known only as a summation. This representa- tion of the Bernoulli numbers was discovered by Euler. It is readily seen from Eq. (12.38) thatjB2njincreases without limit as n!1 . Numeri- cal values have been calculated by Glaisher.1Illustrating the divergent behavior of the Bernoulli numbers, we have B20D5:291102 B200D3:64710215: Some authors prefer to define the Bernoulli numbers with a modified version of Eq. (12.38) by using BnD2.2n/W .2/2n.2n/; (12.39) the subscript being just half of our subscript and all signs positive. Again, when using other texts or references, you must check to see exactly how the Bernoulli numbers are defined. The Bernoulli numbers occur frequently in number theory. The von Staudt-Clausen the- orem states that B2nDAn1 p11 p21 p31 pk; (12.40) in which Anis an integer and p1;p2;:::; pkare all the prime numbers such that pi1is a divisor of 2n. It may readily be verified that this holds for B6.A3D1;pD2;3;7/; B8.A4D1;pD2;3;5/; B10.A5D1;pD2;3;11/; and other special cases. 1J. W. L. Glaisher, table of the first 250 Bernoulli numbers (to nine figures) and their logarithms (to ten figures). Trans. Cam- bridge Philos. Soc. 12: 390 (1871-1879). ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.2 Bernoulli Numbers 565 The Bernoulli numbers appear in the summation of integral powers of the integers, NX jD1jp;pintegral; and in numerous series expansions of the transcendental functions, including tanx,cotx, lnjsinxj,.sinx/1,lnjcosxj,lnjtanxj,.cosh x/1,tanhx, and cothx. For example, tanxDxCx3 3C2 15x5CC.1/n122n.22n1/B2n .2n/Wx2n1C: (12.41) The Bernoulli numbers are likely to appear in such series expansions because of the defi- nition, Eq. (12.28), the form of Eq. (12.35), and the relation to the Riemann zeta function, Eq. (12.38). Bernoulli Polynomials IfEq. (12.28) is generalized slightly, we have tets et1D1X nD0Bn.s/tn nW(12.42) defining the Bernoulli polynomials, Bn.s/. It is clear that Bn.s/will be a polynomial of degree n, since the Taylor expansion of the generating function will contain contributions in which each instance of tmay (or may not) be accompanied by a factor s. The first seven Bernoulli polynomials are given in Table 12.3. If we set sD0in the generating function formula, Eq. (12.42), we have Bn.0/DBn;nD0;1;2;:::; (12.43) showing that the Bernoulli polynomial evaluated at zero equals the corresponding Bernoulli number. Table 12.3 Bernoulli Polynomials B0D1 B1Dx1 2 B2Dx2xC1 6 B3Dx33 2x2C1 2x B4Dx42x3Cx21 30 B5Dx55 2x4C5 3x31 6x B6Dx63x5C5 2x41 2x2C1 42 ArfKen_15-ch12-0551-0598- 9780123846549.tex 566 Chapter 12 Further Topics in Analysis Two other important properties of the Bernoulli polynomials follow from the defining relation, Eq. (12.42). If we differentiate both sides of that equation with respect to s, we have t2ets et1D1X nD0B0 n.s/tn nW D1X nD0Bn.s/tnC1 nWD1X nD1Bn1.s/tn .n1/W; (12.44) where the second line of Eq. (12.44) is obtained by rewriting its left-hand side using the generating-function formula. Equating the coefficients of equal powers of tin the two lines of Eq. (12.44), we obtain the differentiation formula d dsBn.s/DnBn1.s/;nD1;2;3;:::: (12.45) We also have a symmetry relation, which we can obtain by setting sD1in Eq. (12.42). The left-hand side of that equation then becomes tet et1Dt et1: (12.46) Thus, equating Eq. (12.42) forsD1with the Bernoulli-number expansion (in t) of the right-hand side of Eq. (12.46), we reach 1X nD0Bn.1/tn nWD1X nD0Bn.t/n nW; which is equivalent to Bn.1/D.1/nBn.0/: (12.47) These relations are used in the development of the Euler-Maclaurin integration formula. Exercises 12.2.1 Verify the identities, Eqs. (12.32) and (12.46). 12.2.2 Show that the first Bernoulli polynomials are B0.s/D1 B1.s/Ds1 2 B2.s/Ds2sC1 6: Note that Bn.0/DBn, the Bernoulli number. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.3 Euler-Maclaurin Integration Formula 567 12.2.3 Show that tanxD1X nD1.1/n122n.22n1/B2n .2n/Wx2n1; 2<x< 2: Hint. tanxDcotx2 cot 2 x. 12.3 E ULER -M ACLAURIN INTEGRATION FORMULA One use of the Bernoulli polynomials is in the derivation of the Euler-Maclaurin integra- tion formula. This formula is used both to develop asymptotic expansions (treated later in this chapter) and to obtain approximate values for summations. An important application of the Euler-Maclaurin formula, presented in Chapter 13, is its use to derive Stirling’s formula, an asymptotic expression for the gamma function. The technique we use to develop the Euler-Maclaurin formula is repeated integration by parts, using Eq. (12.45) to create new derivatives. We start with 1Z 0f.x/dxD1Z 0f.x/B0.x/dx; (12.48) where we have, for reasons that will shortly become apparent, inserted the redundant factor B0.x/D1. From Eq. (12.45), we note that B0.x/DB0 1.x/; and we substitute B0 1.x/forB0.x/inEq. (12.48), integrate by parts, and identify B1.1/D B1.0/D1 2, thereby obtaining 1Z 0f.x/dxDf.1/B1.1/f.0/B1.0/1Z 0f0.x/B1.x/dx D1 2 f.1/Cf.0/ 1Z 0f0.x/B1.x/dx: (12.49) Again using Eq. (12.45), we have B1.x/D1 2B0 2.x/: Inserting B0 2.x/and integrating by parts again, we get 1Z 0f.x/dxD1 2h f.1/Cf.0/i 1 2h f0.1/B2.1/f0.0/B2.0/i C1 21Z 0f.2/.x/B2.x/dx: (12.50) ArfKen_15-ch12-0551-0598- 9780123846549.tex 568 Chapter 12 Further Topics in Analysis Using the relation B2n.1/DB2n.0/DB2n;nD0;1;2;:::; (12.51) Eq. (12.50) simplifies to 1Z 0f.x/dxD1 2h f.1/Cf.0/i B2 2h f0.1/f0.0/i C1 21Z 0f.2/.x/B2.x/dx:(12.52) Continuing, we replace B2.x/byB0 3.x/=3and once again integrate by parts. Because B2nC1.1/DB2nC1.0/D0; nD1;2;3;:::; (12.53) the integration by parts produces no integrated terms, and 1 21Z 0f.2/.x/B2.x/dxD1 231Z 0f.2/.x/B0 3.x/dxD1 3W1Z 0f.3/.x/B3.x/dx:(12.54) Substituting B3.x/DB0 4.x/=4and carrying out one more partial integration, we get inte- grated terms containing B4.x/, which simplify according to Eq. (12.51). The result is 1 3W1Z 0f.3/.x/B3.x/dxDB4 4Wh f.3/.1/f.3/.0/i C1 4W1Z 0f.4/.x/B4.x/dx:(12.55) We may continue this process, with steps that are entirely analogous to those that led to Eqs. (12.54) and(12.55). After steps leading to derivatives of fof order 2q1, we have 1Z 0f.x/dxD1 2 f.1/Cf.0/ qX pD11 .2p/WB2p f.2p1/.1/f.2p1/.0/ C1 .2q/W1Z 0f.2q/.x/B2q.x/dx: (12.56) This is the Euler-Maclaurin integration formula. It assumes that the function f.x/has the required derivatives. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.3 Euler-Maclaurin Integration Formula 569 The range of integration in Eq. (12.56) may be shifted from T0;1UtoT1;2Uby replacing f.x/byf.xC1/. Adding such results up to Tn1;nU, we obtain nZ 0f.x/dxD1 2f.0/Cf.1/Cf.2/CC f.n1/C1 2f.n/ qX pD11 .2p/WB2ph f.2p1/.n/f.2p1/.0/i C1 .2q/W1Z 0B2q.x/n1X D0f.2q/.xC/dx: (12.57) Note that the derivative terms at the intermediate integer arguments all cancel. However, the intermediate terms f.j/do not, and1 2f.0/Cf.1/CC1 2f.n/appear exactly as in trapezoidal integration, or quadrature, so the summation over pmay be interpreted as a correction to the trapezoidal approximation. Equation (12.57) may therefore be seen as a generalization of Eq. (1.10). In many applications of Eq. (12.57) the final integral containing f.2q/, though small, will not approach zero as qis increased without limit, and the Euler-Maclaurin formula then has an asymptotic, rather than convergent character. Such series, and the implications regarding their use, are the topic of a later section of this chapter. One of the most important uses of the Euler-Maclaurin formula is in summing series by converting them to integrals plus correction terms.2Here is an illustration of the process. Example 12.3.1 ESTIMATION OF .3/ A straightforward application of Eq. (12.57) to.3/ proceeds as follows (noting that all derivatives of f.x/D1=x3vanish in the limit x!1 ): .3/D1X nD11 n3D1 2f.1/C1Z 1dx x3qX pD1B2p .2p/Wf.2p1/.1/Cremainder: (12.58) Evaluating the integral, setting f.1/D1, and inserting f.2n1/.x/D.2nC1/W 2x2nC2 with xD1,Eq. (12.58) becomes .3/D1 2C1 2CqX pD1.2pC1/B2p 2x2pC2Cremainder: (12.59) 2See R. P. Boas and C. Stutz, Estimating sums with integrals. Am. J. Phys. 39: 745 (1971), for a number of examples. ArfKen_15-ch12-0551-0598- 9780123846549.tex 570 Chapter 12 Further Topics in Analysis Table 12.4 Contributions to .3/ of Terms in Euler-Maclaurin Formula n0D1 n0D2 n0D4 Explicit terms 0:500000 1:062500 1:169849R1 n0x3dx 0:500000 0:125000 0:031250 B2term 0:250000 0:015615 0:000977 B4term0:0833330:0013020:000020 B6term 0:083333 0:000326 0:000001 B8term0:1500000:0001460:000000 B10term 0:416667 0:000102 0:000000 B12term1:6452380:0001000:000000 B14term 8:750000 0:000134 0:000000 Suma1:166667 1:201995 1:202057 aSums only include data above horizontal marker. Left column: formula applied to entire summation; central column: formula applied starting from second term; right column: formula starting from fourth term. To assess the quality of this result, we list, in the first data column of Table 12.4, the con- tributions to it. The line marked “explicit terms” consists presently of only the term1 2f.1/. We note that the individual terms start to increase after the B4term; since it is our inten- tion not to evaluate the remainder, the accuracy of the expansion is limited. As discussed more extensively in the section on asymptotic expansions, the best result available from these data is obtained by truncating the expansion before the terms start to increase; adding the contributions above the marker line in the table, we get the value listed as “Sum.” For reference, the accurate value of .3/ is 1.202057. We can improve the result available from the Euler-Maclaurin formula by explicitly calculating some initial terms and applying the formula only to those that remain. This stratagem causes the derivatives entering the formula to be smaller and diminishes the correction from the trapezoid-rule estimate. Simply starting the formula at nD2instead ofnD1reduces the error markedly; see the second data column of Table 12.4. Now the “explicit terms” consist of f.1/C1 2f.2/. Starting the Euler-Maclaurin formula at nD4 further improves the result, then reaching better than seven-figure accuracy.  When the Euler-Maclaurin formula is applied to sums whose summands have a finite number of nonzero derivatives, it can evaluate them exactly. See Exercise 12.3.1. Exercises 12.3.1 The Euler-Maclaurin integration formula may be used for the evaluation of finite series: nX mD1f.m/DnZ 1f.x/dxC1 2f.1/C1 2f.n/CB2 2Wh f0.n/f0.1/i C: ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.4 Dirichlet Series 571 Show that (a)nX mD1mD1 2n.nC1/. (b)nX mD1m2D1 6n.nC1/.2nC1/. (c)nX mD1m3D1 4n2.nC1/2. (d)nX mD1m4D1 30n.nC1/.2nC1/.3n2C3n1/. 12.3.2 The Euler-Maclaurin integration formula provides a way of calculating the Euler- Mascheroni constant to high accuracy. Using f.x/D1=xin Eq. (12.57) (with interval T1;nU) and the definition of , Eq. (1.13), we obtain DnX sD1s1lnn1 2nCNX kD1B2k .2k/n2k: Using double-precision arithmetic, calculate forND1;2;::: . Note. See D. E. Knuth, Euler’s constant to 1271 places. Math. Comput. 16: 275 (1962). ANS. For nD1000; ND2 D0:5772 1566 4901. 12.4 D IRICHLET SERIES Series expansions of the general form S.s/DX nan ns are known as Dirichlet series , and our knowledge of contour integration methods and Bernoulli numbers enables us to evaluate a variety of expressions of this type. One of the most important Dirichlet series is that of the Riemann zeta function, .s/D1X nD11 ns: (12.60) We have already evaluated a sum from which .2/ can be extracted. ArfKen_15-ch12-0551-0598- 9780123846549.tex 572 Chapter 12 Further Topics in Analysis Example 12.4.1 EVALUATION OF .2/ From Example 11.9.1, we have S.a/D1X nD11 n2Ca2Dcotha 2a1 2a2: Simply by taking the limit a!0, we have .2/Dlim a!0S.a/Dlim a!0 2a1 aCa 3C 1 2a2 D2 6: (12.61)  From the relation with the Bernoulli numbers, or alternatively (and perhaps less conve- niently) by contour-integration methods, we find .4/D4 90: Values of.2n/through.10/ are listed in Exercise 12.4.1. The zeta functions of odd integer argument seem unamenable to evaluation in closed form, but are easy to compute numerically (see Example 12.3.1). Other useful Dirichlet series, in the notation of AMS-55 (see Additional Readings), include .s/D1X nD1.1/n1nsD.121s/.s/; (12.62) .s/D1X nD0.2n1/sD.12s/.s/; (12.63) .s/D1X nD0.1/n.2nC1/s: (12.64) Closed expressions are available (for integer n1) for.2n/,.2n/, and.2n/, and for .2n1/. The sums with exponents of opposite parity cannot be reduced to .2n/or per- formed by the contour-integral methods we discussed in Chapter 11. An important series that can only be evaluated numerically is that whose result is Catalan’s constant , which is .2/D11 32C1 52D 0:91596559:::: (12.65) ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.4 Dirichlet Series 573 For reference, we list a few of these summable Dirichlet series: .2/D1C1 22C1 32CD2 6; (12.66) .4/D1C1 24C1 34CD4 90; (12.67) .2/D11 22C1 32CD2 12; (12.68) .4/D11 24C1 34CD74 720; (12.69) .2/D1C1 32C1 52CD2 8; (12.70) .4/D1C1 34C1 54CD4 96; (12.71) .1/D11 3C1 5D 4; (12.72) .3/D11 33C1 53D3 32: (12.73) Exercises 12.4.1 From B2nD.1/n12.2n/W .2/2n.2n/;show that (a).2/D2 6; (d).8/D8 9450; (b).4/D4 90; (e).10/D10 93;555: (c).6/D6 945; 12.4.2 The integral 1Z 0Tln.1x/U2dx x appears in the fourth-order correction to the magnetic moment of the electron. Show that it equals 2.3/ . Hint. Let 1xDet. ArfKen_15-ch12-0551-0598- 9780123846549.tex 574 Chapter 12 Further Topics in Analysis 12.4.3 (a) Show that 1Z 0.lnz/2 1Cz2dzD4 11 33C1 531 73C : (b) By contour integration show that this series evaluates to 3=8. 12.4.4 Show that Catalan’s constant, .2/ , may be written as .2/D21X kD1.4k3/22 8: Hint.2D6.2/ . 12.4.5 Show that (a)Z1 0ln.1Cx/ xdxD1 2.2/, (b) lim a!1Za 0ln.1x/ xdxD.2/. Note that the integrand in part (b) diverges for aD1but that the integral is convergent. 12.4.6 (a) Show that the equation ln 2DP1 sD1.1/sC1s1, Eq. (1.53), may be rewritten as ln 2DnX sD22s.s/C1X pD1.2p/n1 11 2p1 : Hint. Take the terms in pairs. (b) Calculate ln 2to six significant figures. 12.4.7 (a) Show that the equation =4DP1 sD1.1/sC1.2s1/1,Eq. (12.72), may be rewritten as  4D12nX sD142s.2s/21X pD1.4p/2n2 11 .4p/21 : (b) Calculate =4 to six significant figures. 12.5 I NFINITE PRODUCTS We saw in Chapter 11 that complex variable theory can be used to generate infinite-product representations of analytic functions. Here we develop some of their properties. For that purpose it is convenient to write these products in the form PD1Y nD1.1Can/: The infinite product may be related to an infinite series by the obvious method of taking the logarithm: ln1Y nD1.1Can/D1X nD1ln.1Can/: (12.74) ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.5 In/f_inite Products 575 The main theorem regarding convergence of infinite products is the following: If0an<1, the infinite productsQ1 nD1.1Can/andQ1 nD1.1an/converge if P1 nD1anconverges and diverge ifP1 nD1andiverges. For the infinite productQ.1Can/, note that 1Canean; which means that the partial product consisting of the first nfactors satisfies pnesn; where snis the sum of the first n an. Letting n!1 , 1Y nD1.1Can/exp1X mD1an; (12.75) thereby giving an upper bound for the infinite product. To develop a lower bound, we note that, because all ai>0, pnD1CnX iD1aiCnX iD1nX jD1aiajC sn: Hence 1Y nD1.1Can/1X nD1an: (12.76) If the infinite sum remains finite, the infinite product will also. But if the infinite sum diverges, so will the infinite product. The caseQ.1an/is complicated by the negative signs, but a proof similar to the foregoing may be developed by noting that for an<1 2, .1an/.1Can/1and.1an/.1C2an/1: Example 12.5.1 CONVERGENCE OF INFINITE PRODUCTS FOR sinzANDcosz These products, developed in Eqs. (11.89) and (11.90), are sinzDz1Y nD1 1z2 n22 ;coszD1Y nD1 1z2 .n1=2/22 : (12.77) The product expansion of sinzconverges for all z, because, writing the factors as .1an/, 1X nD1anDz2 21X nD1n2Dz2 2.2/Dz2 6; ArfKen_15-ch12-0551-0598- 9780123846549.tex 576 Chapter 12 Further Topics in Analysis a convergent result. For the expansion of cosz, we have 1X nD1anD4z2 21X nD1.2n1/2D4z2 2.2/Dz2 2; also convergent for all z. Note, however, that if zis large, many terms of the product will have to be taken before either of these series approaches convergence. In fact, the main use of these series is in establishing mathematical results rather than for precise numerical work in physics.  We close this section with one further example illustrating a technique for working with infinite products. Example 12.5.2 ANINTERESTING PRODUCT We wish to evaluate the infinite product PD1Y nD2 11 n2 : We note that the product we seek is equivalent to all but the first term of the product expansion of sinzwith zDas given in Eq. (12.77). In fact, the missing first term, which is zero, guarantees that we will get the correct result for sin. For general z, we move the first term (and the prefactor z) to the left-hand side of the product formula for sinz, reaching sinz z.1z2=2/D1Y nD2 1z2 n22 : We now take the limits of the two sides of this equation as z!, applying l’Hôpital’s rule to evaluate the left-hand side and recognizing the right-hand side as P. Thus, PDlimz!sinz z.1z2=2/Dcosz 13z2=2 zDD1 13DC1 2:  Exercises 12.5.1 Using ln1Y nD1.1an/D1X nD1ln.1an/ and the Maclaurin expansion of ln.1an/, show that the infinite productQ1 nD1.1an/ converges or diverges with the infinite seriesP1 nD1an. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.6 Asymptotic Series 577 12.5.2 An infinite product appears in the form 1Y nD11Ca=n 1Cb=n ; where aandbare constants. Show that this infinite product converges only if aDb. 12.5.3 Show that the infinite product representations of sinxandcosxare consistent with the identity 2 sin xcosxDsin 2x. 12.5.4 Determine the limit to whichQ1 nD2 1C.1/n n converges. 12.5.5 Show thatQ1 nD2h 12 n.nC1/i D1 3: 12.5.6 Prove thatQ1 nD2 11 n2 D1 2: 12.5.7 Verify the Euler identityQ1 pD1.1Czp/DQ1 qD1.1z2q1/1;jzj<1: 12.5.8 Show thatQ1 rD1.1Cx=r/ex=rconverges for all finite x(except for the zeros of 1C x=r). Hint. Write the nth factor as 1Can. 12.5.9 Derive the formula, valid for small x, ln sin xDlnxCX anxn; giving the explicit form for the coefficients an. Hint. d.ln sin x/=dxDcotx. 12.5.10 Using the infinite product representations of sinz, show that zcotzD121X m;nD1z n2m ; and hence that the Bernoulli numbers are given by the formula B2nD.1/n12.2n/W .2/2n.2n/: This is an alternate route to Eq. (12.38). Hint. The result of Exercise 12.5.9 will be helpful. 12.6 A SYMPTOTIC SERIES Asymptotic series frequently occur in physics. In fact, one of the earliest and still impor- tant approximations of quantum mechanics, the WKB expansion (the initials stand for its originators, Wenzel, Kramers, and Brillouin), is an asymptotic series. In numerical com- putations, these series are employed for the accurate computation of a variety of functions. ArfKen_15-ch12-0551-0598- 9780123846549.tex 578 Chapter 12 Further Topics in Analysis We consider here two types of integrals that lead to asymptotic series: first, integrals of the form I1.x/D1Z xeuf.u/du; where the variable xappears as the lower limit of an integral. Second, we consider the form I2.x/D1Z 0eufu x du; with the function fto be expanded as a Taylor series (binomial series). Asymptotic series often occur as solutions of differential equations; we encounter many examples in later chapters of this book. Exponential Integral The nature of an asymptotic series is perhaps best illustrated by a specific example. Sup- pose that we have the exponential integral function3 Ei.x/DxZ 1eu udu; (12.78) which we find more convenient to write in the form Ei. x/D1Z xeu uduDE1.x/; (12.79) to be evaluated for large values of x. This function has a series expansion that converges for all x, namely E1.x/D lnx1X nD1.1/nxn nnW; (12.80) which we derive in Chapter 13, but the series is totally useless for numerical evalua- tion when xis large. We need another approach, for which it is convenient to generalize Eq. (12.79) to I.x;p/D1Z xeu updu; (12.81) where we restrict consideration to cases in which xandpare positive. As already stated, we seek an evaluation for large values of x. 3This function occurs frequently in astrophysical problems involving gas with a Maxwell-Boltzmann energy distribution. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.6 Asymptotic Series 579 Integrating by parts, we obtain I.x;p/Dex xpp1Z xeu upC1duDex xppex xpC1Cp.pC1/1Z xeu upC2du: Continuing to integrate by parts, we develop the series I.x;p/Dex1 xpp xpC1Cp.pC1/ xpC2C.1/n1.pCn2/W .p1/WxpCn1 C.1/n.pCn1/W .p1/W1Z xeu upCndu: (12.82) This is a remarkable series. Checking the convergence by the d’Alembert ratio test, we find limn!1junC1j junjDlimn!1.pCn/W .pCn1/W1 xDlimn!1pCn xD1 (12.83) for all finite values of x. Therefore our series as an infinite series diverges everywhere! Before discarding Eq. (12.83) as worthless, let us see how well a given partial sum approxi- mates our function I.x;p/. Taking snas the partial sum of the series through nterms and Rnas the corresponding remainder, I.x;p/sn.x;p/D.1/nC1.pCn/W .p1/W1Z xeu upCnC1duDRn.x;p/: In absolute value jRn.x;p/j.pCn/W .p1/W1Z xeu upCnC1du: When we substitute uDvCx;the integral becomes 1Z xeu upCnC1duDex1Z 0ev .vCx/pCnC1dv Dex xpCnC11Z 0ev 1Cv xpn1 dv: For large xthe final integral approaches 1 and jRn.x;p/j.pCn/W .p1/Wex xpCnC1: (12.84) ArfKen_15-ch12-0551-0598- 9780123846549.tex 580 Chapter 12 Further Topics in Analysis This means that if we take xlarge enough, our partial sum snwill be an arbitrarily good approximation to the function I.x;p/. Our divergent series, Eq. (12.82), therefore is per- fectly good for computations of partial sums. For this reason it is sometimes called a semi- convergent series. Note that the power of xin the denominator of the remainder, namely pCnC1, is higher than the power of xin the last term included in sn.x;p/, namely pCn. Thus, our asymptotic series for E1.x/assumes the form exE1.x/Dex1Z xeu udu sn.x/D1 x1W x2C2W x23W x4CC.1/nnW xnC1; (12.85) where we must choose to terminate the series after some n. Since the remainder Rn.x;p/alternates in sign, the successive partial sums give alter- nately upper and lower bounds for I.x;p/. The behavior of the series (with pD1) as a function of the number of terms included is shown in Fig. 12.2, where we have plotted partial sums of exE1.x/for the value xD5. The optimum determination of exE1.x/is given by the closest approach of the upper and lower bounds, that is, for xD5, between s6D0:1664 ands5D0:1741 . Therefore 0:1664exE1.x/ xD50:1741: (12.86) Actually, from tables, exE1.x/ xD5D0:1704; (12.87) 0.17040.1714 0.16640.20 0.190.180.170.160.150.14s n(x=5) n 10 9 8 7 6 5 4 3 2 1 FIGURE 12.2 Partial sums of exE1.x/jxD5. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.6 Asymptotic Series 581 within the limits established by our asymptotic expansion. Note that inclusion of addi- tional terms in the series expansion beyond the optimum point reduces the accuracy of the representation. As xis increased, the spread between the lowest upper bound and the highest lower bound will diminish. By taking xlarge enough, one may compute exE1.x/ to any desired degree of accuracy. Other properties of E1.x/are derived and discussed in Section 13.6. Cosine and Sine Integrals Asymptotic series may also be developed from definite integrals, provided that the inte- grand has the required behavior. As an example, the cosine and sine integrals (in Table 1.2) are defined by Ci.u/D1Z ucost tdt; (12.88) si.u/D1Z usint tdt: (12.89) Combining these, using the formula for eit, Ci.u/Cisi.u/D1Z ueit tdt; and then changing the integration variable from ttoz, we reach F.u/DCi.u/Cisi.u/Deiu1Z 0eizdz uCz: (12.90) To further process F.u/, we now consider the contour integral eiuI Ceizdz uCz; where the contour Cis that shown in Fig. 12.3. Since we are interested in evaluation for large positive (and real) u, our integrand has as its only singularity a pole on the negative real axis, so the region enclosed by the contour is entirely analytic and the contour integral therefore vanishes. The exponential and the denominator cause the arc at infinity (labeled B) not to contribute to the contour integral, so the integral we seek is obtained from seg- ment Aand must be equal to the negative of the integral on segment D. Therefore, we have F.u/Deiu1Z 0eyidy uCiy; (12.91) ArfKen_15-ch12-0551-0598- 9780123846549.tex 582 Chapter 12 Further Topics in Analysis ADB FIGURE 12.3 Contour for sine and cosine integrals. which is already helpful since we have converted an oscillatory integral into one with a monotonically and exponentially decreasing integrand. To obtain an asymptotic expansion, we continue by expanding the denominator of the integrand using the binomial theorem, writing 1 uCiyD1 u" 1iy uCiy u2 # : We plan to integrate in yfrom zero to infinity, and the proposed expansion will be diver- gent when y>u, but we proceed anyway, because the terms of the series will initially be decreasing and will be satisfactory as an asymptotic expansion. Formally, we take the viewpoint that we are writing 1=.uCiy/as a finite series plus a remainder, and we will abandon the expansion at or before the point that the remainder is a minimum. Inserting the expansion, and integrating termwise using the formula 1Z 0yneydyDnW; we get F.u/ieiu u 1i1W u 2W u2 Ci3W u3 C4W u4  : (12.92) As for our earlier example, the exponential integral, this series will diverge for all u, but if uis sufficiently large the terms will initially decrease to very small values before increasing again toward divergence. To go from the expansion of F.u/to those of Ci and si, we need to separate it into real and imaginary parts. Writing eiuDcosuCisinuand collecting terms appropriately, we ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.6 Asymptotic Series 583 get as the desired asymptotic expansions Ci.u/sinu uNX nD0.1/n.2n/W u2ncosu uNX nD0.1/n.2nC1/W u2nC1; (12.93) si.u/cosu uNX nD0.1/n.2n/W u2nsinu uNX nD0.1/n.2nC1/W u2nC1: (12.94) Definition of Asymptotic Series Poincaré has introduced a formal definition for an asymptotic series.4Following Poincaré, we consider a function f.x/whose asymptotic expansion is sought, the partial sums snin its expansion, and the corresponding remainders Rn.x/. Though the expansion need not be a power series, we assume that form for simplicity in the present discussion. Thus, xnRn.x/DxnTf.x/sn.x/U; (12.95) where sn.x/Da0Ca1 xCa2 x2CCan xn: (12.96) The asymptotic expansion of f.x/is defined to have the properties that limx!1xnRn.x/D0;for fixed n; (12.97) and limn!1xnRn.x/D1; for fixed x: (12.98) These conditions were met for our examples, Eqs. (12.85), (12.93), and (12.94).5 For power series, as assumed in the form of sn.x/;Rn.x/xn1. With the conditions ofEqs. (12.97) and(12.98) satisfied, we write f.x/1X nD0anxn: (12.99) Note the use ofin place ofD. The function f.x/is equal to the series only in the limit asx!1 and with the restriction to a finite number of terms in the series. Asymptotic expansions of two functions may be multiplied together, and the result will be an asymptotic expansion of the product of the two functions. The asymptotic expansion of a given function f.t/may be integrated term by term (just as in a uniformly convergent series of continuous functions) from xt<1;and the result will be an asymptotic 4Poincaré’s definition allows (or neglects) exponentially decreasing functions. The refinement of his definition is of considerable importance for the advanced theory of asymptotic expansions, particularly for extensions into the complex plane. However, for purposes of an introductory treatment and especially for numerical computation of expansions for which the variable is real and positive, Poincaré’s approach is perfectly satisfactory. 5Some writers feel that the requirement of Eq. (12.98), which excludes convergent series of inverse powers of x, is artificial and unnecessary. ArfKen_15-ch12-0551-0598- 9780123846549.tex 584 Chapter 12 Further Topics in Analysis expansion ofR1 xf.t/dt. Term-by-term differentiation, however, is valid only under very special conditions. Some functions do not possess an asymptotic expansion; exis an example of such a function. However, if a function has an asymptotic expansion of the power-series form in Eq. (12.99), it has only one. The correspondence is not one to one; many functions may have the same asymptotic expansion. One of the most useful and powerful methods of generating asymptotic expansions, the method of steepest descents, is developed in the next section of this text. Exercises 12.6.1 Integrating by parts, develop asymptotic expansions of the Fresnel integrals (a)C.x/DZx 0cosu2 2du, (b) s.x/DZx 0sinu2 2du. These integrals appear in the analysis of a knife-edge diffraction pattern. 12.6.2 Rederive the asymptotic expansions of Ci .x/and si.x/by repeated integration by parts. Hint. Ci. x/Cisi.x/DZ1 xeit tdt: 12.6.3 Derive the asymptotic expansion of the Gauss error function erf.x/D2pxZ 0et2dt 1ex2 px 11 2x2C13 22x4135 23x6CC.1/n.2n1/WW 2nx2n : Hint. erf. x/D1erfc.x/D12pZ1 xet2dt. Normalized so that erf.1/ D1, this function plays an important role in probability theory. It may be expressed in terms of the Fresnel integrals (Exercise 12.6.1), the in- complete gamma functions (Section 13.6), or the confluent hypergeometric functions (Section 18.5). 12.6.4 The asymptotic expressions for the various Bessel functions, Section 14.6, contain the series P.z/1C1X nD1.1/nQ2n sD1T42.2s1/2U .2n/W.8z/2n; Q.z/1X nD1.1/nC1Q2n1 sD1T42.2s1/2U .2n1/W.8 z/2n1: Show that these two series are indeed asymptotic series. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.7 Method of Steepest Descents 585 12.6.5 Forx>1; 1 1CxD1X nD0.1/n1 xnC1: Test this series to see if it is an asymptotic series. 12.6.6 Derive the following Bernoulli-number asymptotic series for the Euler-Mascheroni con- stant, defined in Eq. (1.13): nX sD1s1lnn1 2nC1X kD1B2k .2k/n2k: Here nplays the role of x. Hint. Apply the Euler-Maclaurin integration formula to f.x/Dx1over the interval T1;nUforND1;2;::: . 12.6.7 Develop an asymptotic series for 1Z 0exv .1Cv2/2dv: Take xto be real and positive. ANS.1 x2W x3C4W x5C.1/n.2n/W x2nC1. 12.7 M ETHOD OF STEEPEST DESCENTS In this section we consider the frequently occurring situation that we require the asymptotic behavior (for large t, assumed real) of a function f.t/, where f.t/is represented by an integral of the generic form f.t/DZ CF.z;t/dz; with F.z;t/analytic in z, but also parametrically dependent on t; The integration path Cis, or can be deformed to be, such that for large tthe dominant contribution to the integral arises from a small range of zin the neighborhood of the point z0wherejF.z0;t/jis a maximum on the path; The integration path will pass through z0in the orientation that causes the most rapid decrease injFjon departure from z0in either direction along the path (hence the name steepest descents); and In the limit of large tthe contribution to the integral from the neighborhood of z0 asymptotically approaches the exact value of f.t/. ArfKen_15-ch12-0551-0598- 9780123846549.tex 586 Chapter 12 Further Topics in Analysis While the above conditions seem rather restrictive, they can in fact be met for many of the important special functions of mathematical physics, including, among others, the gamma function and various Bessel functions. Saddle Points The integration path supplied with the original definition of an integral representation defining a function f.t/will not usually meet the conditions outlined above, and we need to consider the features of the integrand F.z;t/that will be useful in defining a more suit- able path which, even if the original formulation is entirely real, may be a more general contour in the complex plane. We already know (Exercise 11.2.2) that neither the real nor the imaginary part of an analytic function can have an extremum (either a minimum or maximum) within the region of analyticity, and the same is also true of its modulus (this result is Jensen’s theorem; see Exercise 12.7.1). To better understand that, let us represent F.z;t/(in a region where it is assumed nonzero) in the form F.z;t/Dew.z;t/Deu.z;t/Civ.z;t/; (12.100) where uandvare the real and imaginary parts of an analytic function w; this representation permits us to identify uaslnjFj; the fact that ucannot have an extremum makes Jensen’s theorem obvious. Although ucannot have an extremum, it can have a saddle point (a point at which w0D0; then also du=dsD0for all directions ds, but with higher derivatives that are positive in some directions and negative in others (see Fig. 12.4). Let us examine some general features of wand its components uandvin the neighborhood of a saddle point of u, which we designate z0. We proceed by expanding w.z;t/in a Taylor series about z0. Becausew0D0there, the first two nonzero terms of the expansion are w.z;t/Dw.z0;t/Cw00.z0;t/ 2W.zz0/2C: (12.101) FIGURE 12.4 Saddle point of u.DjFj/; see Eq. (12.100). ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.7 Method of Steepest Descents 587 It could be that w00.z0;t/D0, but that possibility makes the analysis more complicated without changing it in a fundamental way, so we proceed under the assumption that w00.z0;t/6D0. Using the abbreviated notations w0Dw.z0;t/,w00.z0;t/Dw00 0, and in- troducing the polar forms w00 0Djw00 0jei ,zz0Drei, Eq. (12.101) becomes w.z;t/Dw0C1 2jw00 0jei. C2/r2C (12.102) Dw0C1 2jw00 0jr2h cos. C2/Cisin. C2/i C: (12.103) For later reference we note that is the argument ofw00.z0;t/. We see that, in general at a saddle point, u(the real part of w) will increase most rapidly when C2D2n, corresponding to the opposite directions D =2 andD =2C. On the other hand, uwill decrease most rapidly when C2D.2nC1/, i.e.,D =2C.1 2or3 2/, the two directions perpendicular to those of maximum increase. And uwill (to second order) remain constant (so-called level lines) in the directions D =2C.1 4;3 4;5 4;7 4/. See the left panel of Fig. 12.5. The behavior of v(the imaginary part of w) will be similar to that of u, but displaced in angle by 45. The level lines of vwill be in the directions D sC.0; =2; ; 3=2/ , and therefore will coincide with the directions of maximum increase or decrease in u. See the right panel of Fig. 12.5. We are now ready to identify an optimum contour for evaluating the integral repre- sentation of f.t/, namely one that passes through the saddle point z0in the directions of maximum rate of decrease in uwith distance from z0, and therefore also in jFj. These directions have the additional advantage that they are level lines of v, so that the factor eiv will not produce changes of phase (oscillatory behavior and therefore numerical instabil- ity) in Fas we leave the saddle point. If we had chosen z0to be a point other than a saddle point, the expansion of wwould have contained a nonzero linear term in r, and it would not have been possible to construct a curve through z0that would causejFjto decrease in both path directions, or to keep the phase of Fconstant. LevelLevel uv++ ++− − −− Level Level FIGURE 12.5 Near a saddle point in wDuCiv: When features of uare oriented as in the left panel, those of vare as shown in the right panel. Arrows indicate ascending directions. ArfKen_15-ch12-0551-0598- 9780123846549.tex 588 Chapter 12 Further Topics in Analysis Saddle Point Method Now that we have identified z0and the directions of steepest descent in jF.z;t/j, we com- plete the specification of the method of steepest descents, also called the saddle point method of asymptotic approximation, by assuming that the significant contributions to the integral are from a small range of 0rain each of the two directions along the path. Before obtaining a final result, we must make one more observation. Looking at the way in which the contour had to be deformed to pass through z0, we need to determine the sense of the path (i.e., we must decide whether the direction of travel is at D =2C1 2 or atD =2C3 2). Assuming that this has been decided, we can then identify, for the portion of the path in which we descend from F.z0/,dzDeidr. The contribution in which we ascend to F.z0/will have the opposite sign for dzbut we can handle it simply by multiplying the descending contribution by two. Then, noting that ei. C2/D1 , our approximation to f.t/is f.t/2ew0CiaZ 0ejw00 0jr2=2dr; (12.104) where the initial “2” causes inclusion of the ascent to z0. We now make the key assumption of the method, namely that jw00 0j, the measure of the rate of decrease in jFjas we leave z0, is large enough that the bulk of the value of the integral has already been attained for small a, and that the exponential decrease in the value of the integrand enables us to replace a by infinity without making significant error. In problems where the saddle point method is applicable, this condition is met when tis sufficiently large. We complete the present analysis by remembering that ew0DF.z0;t/and by evaluating the integral for aD1 , where, cf. Eq. (1.148), it has the valueq =2jw00 0j. We get f.t/F.z0;t/eis 2 jw00.z0;t/j: (12.105) We remind the reader that Darg.w00.z0;t// 2C 2or3 2 ; (12.106) with the choice (which affects only the sign of the final result) determined from the sense in which the contour passes through the saddle point z0. Sometimes it is sufficient to apply the method of steepest descents only to the rapidly varying part of an integral. This corresponds to assuming that we may make the approxi- mation f.t/DZ Cg.z;t/F.z;t/dzg.z0;t/Z CF.z;t/dz; (12.107) ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.7 Method of Steepest Descents 589 after which we proceed as before. Note that this causes gnot to be considered when we defineworw00, and our final formula is replaced by f.t/g.z0;t/F.z0;t/eis 2 jw00.z0;t/j: (12.108) A final note of warning: We assumed that the only significant contribution to the inte- gral came from the immediate vicinity of the saddle point zDz0. This condition must be checked for each new problem. Example 12.7.1 ASYMPTOTIC FORM OF THE GAMMA FUNCTION In many physical problems, particularly in the field of statistical mechanics, it is desir- able to have an accurate approximation of the gamma or factorial function of very large numbers. As listed in Table 1.2, the factorial function may be defined by the Euler integral tWD0.tC1/D1Z 0tedDttC11Z 0et.lnzz/dz: (12.109) Here we have made the substitution Dztin order to convert the integral to the form given in Eq. (12.108). As before, we assume that tis real and positive, from which it follows that the integrand vanishes at the limits 0 and 1. By differentiating the exponent, which we callw.z;t/, we obtain dw dzDtd dz.lnzz/Dt zt; w00Dt z2; which shows that the point zD1is a saddle point and argw00.1;t/Darg.t/D:Apply- ing Eq. (12.106), we see that the direction of travel through the saddle point is Dargw00 2C 2or3 2 D0orI the choiceD0is that consistent with deformation from a path that was originally along the real axis. In fact, what we have found is that the direction of steepest descent is along the real axis, a conclusion that we might have reached more or less intuitively. Direct substitution into Eq. (12.108) with gDttC1,FDet,D0, andjw00jDt yields tWD0.tC1/r 2 tttC1etDp 2ttC1=2et: (12.110) This result is the leading term in Stirling’s expansion of the gamma function. The method of steepest descents is probably the easiest way of obtaining this term. Further terms in the asymptotic expansion are developed in Section 13.4. In this example the calculation was carried out assuming tto be real. This assumption is not necessary. We may show (Exercise 12.7.3) that Eq. (12.110) also holds when tis complex, provided only that its real part be required to be large and positive.  ArfKen_15-ch12-0551-0598- 9780123846549.tex 590 Chapter 12 Further Topics in Analysis Sometimes the application of the saddle point method to a real integral results in a con- tour that goes through a saddle point that is not on the real axis. Here is a relatively simple example. A more complicated case of practical importance appears in the chapter on Bessel functions (see Section 14.6). Example 12.7.2 SADDLE POINT METHOD AVOIDS OSCILLATIONS As a second example of the method of steepest descents, consider the integral H.t/D1Z 1et.z21=4/costz 1Cz2dz; (12.111) which we wish to evaluate for large positive t. When tis large, the integrand oscillates very rapidly, and ordinary quadrature methods become difficult. We proceed by bringing H.t/ to a form appropriate for applying the saddle point method, replacing costzbycostzC isintzDeitz(a replacement that does not change the value of the integral because we added an odd term to the previously even integrand). We then have H.t/DZ Cg.z/et.z2iz1=4/dz; (12.112) with g.z/D1=.1Cz2/. This form corresponds to w.z/Dt.z2iz1 4/, so we have w0.z/Dt.2zi/;which has a zero at z0Di=2. (12.113) Then, at z0, which is a saddle point, w0D0; w00.z0/D2t;g.z0/D4 3: (12.114) We also need the phase of the steepest-descent direction. Noting that arg.w00.z0//D and applying Eq. (12.106), we find D0(or). We are now ready to apply Eq. (12.108). The result is H.t/p 2.4=3/.e0/ j2tjD4 3r t: (12.115) As a check, we compare this approximate formula for H.t/with the result of a tedious numerical integration: For tD100,HexactD0:23284 , and HsaddleD0:23633 .  Exercises We present here a rather small number of exercises on the method of steepest descents. Several additional exercises appear elsewhere in this book, in particular in Section 14.6, where the technique is applied to the contour integral representations of Bessel func- tions. 12.7.1 Prove Jensen’s theorem (that jF.z/j2can have no extremum in the interior of a region in which Fis analytic) by showing that the mean value of jFj2on a circle about any ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.8 Dispersion Relations 591 point z0is equal tojF.z0/j2. Explain why you can then conclude that there cannot be an extremum ofjFjatz0. 12.7.2 Find the steepest path and leading asymptotic expansion for the Fresnel integralsZs 0cosx2dx;Zs 0sinx2dx. Hint. UseZ1 0eitz2dz. 12.7.3 Show that the formula 0.1Cs/p 2ssses holds for complex values of s(with<e.s/large and positive). Hint. This involves assigning a phase to sand then demanding that ImTs f.z/Ube con- stant in the vicinity of the saddle point. 12.8 D ISPERSION RELATIONS The concept of dispersion relations entered physics with the work of Kronig and Kramers in optics. The name dispersion comes from optical dispersion, a result of the dependence of the index of refraction on wavelength, or angular frequency. As we shall soon see, the index of refraction nmay have a real part determined by the phase velocity and a (negative) imaginary part determined by the absorption. Kronig and Kramers showed in 1926–1927 that the real part of .n21/could be expressed as an integral of the imaginary part. Generalizing this, we shall apply the label dispersion relations to any pair of equations giving the real part of a function as an integral of its imaginary part and the imaginary part as an integral of its real part (we develop this in more detail below). The existence of such integral relations might be suspected as an integral analog of the Cauchy-Riemann differential equations, Eq. (11.9). The applications in modern physics are widespread. For instance, the real part of the function might describe the forward scattering of a gamma ray in a nuclear Coulomb field (a dispersive process). Then the imaginary part would describe the electron-positron pair production in that same Coulomb field (the absorptive process). As will be seen later, the dispersion relations may be taken as a consequence of causality and therefore are indepen- dent of the details of the particular interaction. We consider a complex function f.z/that is analytic in the upper half-plane and on the real axis. We also require that f.z/approach zero for large jzjin the upper half-plane sufficiently rapidly that its integral over the semicircular part of the contour in Fig. 12.6 will be negligible. The point of these conditions is that we may express f.z/by the Cauchy integral formula, Eq. (11.30), using this contour, obtaining f.z0/D1 2i1Z 1f.x/ xz0dx: (12.116) The integral over the contour shown in Fig. 12.6 has become an integral along the x-axis. ArfKen_15-ch12-0551-0598- 9780123846549.tex 592 Chapter 12 Further Topics in Analysis xzy −R −∞ R ∞ FIGURE 12.6 Contour for dispersion integral. Equation (12.116) assumes that z0is in the upper half-plane, interior to the closed con- tour. If z0were in the lower half-plane, the integral would yield zero by the Cauchy in- tegral theorem, Section 11.3. Now, if we move z0onto the real axis (then calling it x0) and pass it via a small clockwise semicircle sin the upper half-plane, the contour integral (which would contain no singularities) would have nonzero contributions corresponding to a Cauchy principal value integral minus half the usual contribution from the pole at x0, or 0DZf.x/ xx0dxCZ sf.z/ zx0dz DZf.x/ xx0dxi f.x0/; equivalent to the final formula f.x0/D1 i1Z 1f.x/ xx0dx: (12.117) Note that the cut integral sign denotes the Cauchy principal value. Splitting Eq. (12.117) into real and imaginary parts6yields f.x0/Du.x0/Civ.x0/ D1 1Z 1v.x/ xx0dxi 1Z 1u.x/ xx0dx: Finally, equating real part to real part and imaginary part to imaginary part, we obtain u.x0/D1 1Z 1v.x/ xx0dx; v.x0/D1 1Z 1u.x/ xx0dx:(12.118) 6The second argument, yD0, is dropped: u.x0;0/!u.x0/. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.8 Dispersion Relations 593 These are the dispersion relations. The real part of our complex function is expressed as an integral over the imaginary part. The imaginary part is expressed as an integral over the real part. Alternatively, the real part can be called an integral transform of the imaginary part (and vice versa); the particular transform involved is known as a Hilbert transform, and we note that (apart from a minus sign) the Hilbert transform is its own inverse. Note that these relations are meaningful only when f.x/is a complex function of the real variable x. Compare Exercise 12.8.1. From a physical point of view u.x/and/orv.x/may represent some physical measure- ments. Then f.z/Du.z/Civ.z/is an analytic continuation over the upper half-plane, with the value on the real axis serving as a boundary condition. Symmetry Relations On occasion f.x/will satisfy a symmetry relation and the integral from 1 toC1 may be replaced by an integral over positive values only. This is of considerable physical importance because the variable xmight represent a frequency and only zero and positive frequencies are available for physical measurements. Suppose7 f.x/Df.x/: (12.119) Then u.x/Civ.x/Du.x/iv.x/: (12.120) The real part of f.x/is even and the imaginary part is odd.8In quantum mechanical scat- tering problems these relations, Eq. (12.120), are called crossing conditions. To exploit these crossing conditions, we rewrite the first of Eqs. (12.118) as u.x0/D1 0Z 1v.x/ xx0dxC1 1Z 0v.x/ xx0dx: (12.121) Letting x! xin the first integral on the right-hand side of Eq. (12.121) and substituting v.x/Dv. x/from Eq. (12.120), we obtain u.x0/D1 1Z 0v.x/1 xCx0C1 xx0 dx D2 1Z 0xv.x/ x2x2 0dx: (12.122) 7This is not just a curiosity. It ensures that the Fourier transform of f.x/will be real. Or conversely, Eq. (12.119) is a conse- quence when f.x/is obtained as the Fourier transform of a real function. 8u.x;0/Du.x;0/;v. x;0/Dv. x;0/. Compare these symmetry conditions with those that follow from the Schwarz reflec- tion principle, Section 11.10. ArfKen_15-ch12-0551-0598- 9780123846549.tex 594 Chapter 12 Further Topics in Analysis Similarly, v.x0/D2 1Z 0x0u.x/ x2x2 0dx: (12.123) The original Kronig-Kramers optical dispersion relations were in this form. The asymptotic behavior.x0!1/ of Eqs. (12.122) and (12.123) lead to quantum mechanical sum rules. See Exercise 12.8.4. Optical Dispersion The function expTi.kx!t/Ucan describe an electromagnetic wave moving along the x- axis in the positive direction with velocity vD!=k;!is the angular frequency, kthe wave number or propagation vector, and nDck=!, the index of refraction. From Maxwell’s equations with electric permittivity "and magnetic permeability unity, and using Ohm’s law with conductivity , the propagation vector kfor a dielectric becomes9 k2D"!2 c2 1Ci4 !" : (12.124) The presence of the conductivity (which means absorption) causes k2to have an imagi- nary part. The propagation vector k(and therefore the index of refraction n) have become complex. For poor conductivity ( 4=!"1) a binomial expansion yields kDp"! cCi2 cp" and ei.kx!t/Dei!.xp"=ct/e2 x=cp"; an attenuated wave. Returning to the general expression for k2;Eq. (12.124), we find that the index of re- fraction becomes n2Dc2k2 !2D"Ci4 !: (12.125) We take n2to be a function of the complex variable!(with"anddepending on !). However, n2does not vanish as !!1 but instead approaches unity. It therefore does not satisfy the condition needed for a dispersion relation, but this difficulty can be cir- cumvented by working with f.!/Dn2.!/1. The Kronig-Kramers relations then take 9See J. D. Jackson, Classical Electrodynamics, 3rd ed. New York: Wiley (1999), Sections 7.7 and 7.10. Equation (12.124) is in Gaussian units. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.8 Dispersion Relations 595 the form <eTn2.!0/1UD2 1Z 0!ImTn2.!/1U !2!2 0d!; ImTn2.!0/1UD2 1Z 0!0<eTn2.!/1U !2!2 0d!:(12.126) Knowledge of the absorption coefficient at all frequencies specifies the real part of the index of refraction, and vice versa. The Parseval Relation When the functions u.x/andv.x/are Hilbert transforms of each other, given by Eqs. (12.118), and each is square integrable,10the two functions satisfy the scaling condi- tion 1Z 1ju.x/j2dxD1Z 1jv.x/j2dx: (12.127) This is the Parseval relation. To derive Eq. (12.127), we start with 1Z 1ju.x/j2dxD1Z 1dx2 41 1Z 1v.s/ds sx3 52 41 1Z 1v.t/dt tx3 5; using the formula for u.x/from Eq. (12.118) twice. Integrating first with respect to x, we have 1Z 1ju.x/j2dxD1Z 1v.s/ds1Z 1v.t/dt1 21Z 1dx .sx/.tx/; (12.128) where both principal-value limits at the singularities of the integrand must now be taken for thexintegration. As shown in Exercise 12.8.8, that integration yields a delta function11: 1 21Z 1dx .sx/.tx/D.st/: Thus, 1Z 1ju.x/j2dxD1Z 1v.t/dt1Z 1v.s/.st/ds: (12.129) 10This means thatR1 1ju.x/j2dxandR1 1jv.x/j2dxare finite. 11Note that when sDt, the integrand has the same sign (for small ") atxDs"and at xDsC", so the limit defining the principal value then does not exist. The singularity in the integration is that which is needed to represent a delta function. ArfKen_15-ch12-0551-0598- 9780123846549.tex 596 Chapter 12 Further Topics in Analysis Then the sintegration is carried out by inspection, using the defining property of the delta function: 1Z 1v.s/.st/dsDv.t/: (12.130) Substituting Eq. (12.130) intoEq. (12.129), we have Eq. (12.127), the Parseval relation. Again, in terms of optics, the presence of refraction over some frequency range .n6D1/ implies the existence of absorption, and vice versa. Exercises 12.8.1 Assume that the function f.z/satisfies the conditions for the dispersion relations. In addition, assume that f.z/Df.z/, i.e., that it meets the conditions of the Schwarz reflection principle, Eq. (11.127). Show that f.z/is identically zero. 12.8.2 Forf.z/such that we may replace the closed contour of the Cauchy integral formula by an integral over the real axis we have f.x0/D1 2i8 >< >:x0Z 1f.x/ xx0dxC1Z x0Cf.x/ xx0dx9 >= >;C1 2iZ Cf.x/ xx0dx: Here we take Cto be a small semicircle about x0in the lower half-plane. Show that the formula for f.x0/reduces to f.x0/D1 i1Z 1f.x/ xx0dx; which is Eq. (12.117). 12.8.3 (a) The function f.z/Deizdoes not vanish at the endpoints of the range of argz, a and. Show, with the help of Jordan’s lemma, Eq. (11.102), that Eq. (12.116) still holds. (b) For f.z/Deizverify by direct integration the dispersion relations, Eq. (12.117) orEqs. (12.118). 12.8.4 With f.x/Du.x/Civ.x/andf.x/Df.x/, show that as x0!1 , (a) u.x0/2 x2 0Z1 0xv.x/dx, (b)v.x0/2 x0Z1 0u.x/dx. In quantum mechanics relations of this form are often called sum rules. ArfKen_15-ch12-0551-0598- 9780123846549.tex 12.8 Dispersion Relations 597 12.8.5 (a) Given the integral equation (valid for all real x0) 1 1Cx2 0D1 1Z 1u.x/ xx0dx; use Hilbert transforms to determine u.x0/. (b) Verify that the u.x0/found as your answer to part (a) actually satisfies the integral equation. (c) From f.z/jyD0Du.x/Civ.x/, replace xbyzand determine f.z/. Verify that the conditions for the Hilbert transforms are satisfied. (d) Are the crossing conditions satisfied? ANS. (a)u.x0/Dx0 1Cx2 0, (c) f.z/D.zCi/1. 12.8.6 (a) If the real part of the complex index of refraction (squared) is constant (no optical dispersion), show that the imaginary part is zero (no absorption). (b) Conversely, if there is absorption, show that there must be dispersion. In other words, if the imaginary part of n21is not zero, show that the real part of n21 is not constant. 12.8.7 Given u.x/Dx=.x2C1/andv.x/D1=. x2C1/, show by direct evaluation of each integral that 1Z 1ju.x/j2dxD1Z 1jv.x/j2dx: ANS.Z1 1ju.x/j2dxDZ1 1jv.x/j2dxD 2. 12.8.8 Take u.x/D.x/, a delta function, and assume that the Hilbert transform equations hold. (a) Show that .w/D1 21Z 1dy y.yw/: (b) With changes of variables wDstandxDsy, transform the representation of part (a) into .st/D1 21Z 1dx .xs/.xt/: Note. Thefunction is discussed in Section 1.11. ArfKen_15-ch12-0551-0598- 9780123846549.tex 598 Chapter 12 Further Topics in Analysis Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Lewin, L., Polylogarithms and Associated Functions . New York: North-Holland (1981). This is a definitive resource for the dilogarithm and its generalizations up through its publication date. It is clear and no more difficult than necessary. McBride, E. B., Obtaining Generating Functions. New York: Springer-Verlag (1971). An introduction to meth- ods of obtaining generating functions, both for sets of functions arising from ODEs and for those that do not. Nussenzveig, H. M., Causality and Dispersion Relations, Mathematics in Science and Engineering Series, Vol. 95. New York: Academic Press (1972). This is an advanced text covering causality and dispersion relations in the first chapter and then moving on to develop implications in a variety of areas of theoretical physics. Talman, J. D., Special Functions. New York: W. A. Benjamin (1968). Develops the theory of a number of special functions using their underlying group-theoretical properties, including presentation of their generating functions. Wyld, H. W., Mathematical Methods for Physics. Reading, MA: Benjamin/Cummings (1976), Perseus Books (1999). This is a relatively advanced text that contains an extensive discussion of dispersion relations. ArfKen_Ch13-9780123846549.tex CHAPTER 13 GAMMA FUNCTION The gamma function is probably the special function that occurs most frequently in the discussion of problems in physics. For integer values, as the factorial function, it appears in every Taylor expansion. As we shall later see, it also occurs frequently with half-integer arguments, and is needed for general nonintegral values in the expansion of many func- tions, e.g., Bessel functions of noninteger order. It has been shown that the gamma function is one of a general class of functions that do not satisfy any differential equation with rational coefficients. Specifically, the gamma function is one of very few functions of mathematical physics that do not satisfy either the hypergeometric differential equation (Section 18.5) or the confluent hypergeomet- ric equation (Section 18.6). Since most physical theories involve quantities governed by differential equations, the gamma function (by itself) does not usually describe a physi- cal quantity of interest, but rather tends to appear as a factor in expansions of physically relevant quantities. 13.1 D EFINITIONS , PROPERTIES At least three different convenient definitions of the gamma function are in common use. Our first task is to state these definitions, to develop some simple, direct consequences, and to show the equivalence of the three forms. Infinite Limit (Euler) The first definition, named after Euler, is 0.z/limn!1123n z.zC1/.zC2/.zCn/nz;z6D0;1;2;3;:::: (13.1) 599 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch13-9780123846549.tex 600 Chapter 13 Gamma Function This definition of 0.z/is useful in developing the Weierstrass infinite-product form of 0.z/, Eq. (13.16), and in obtaining the derivative of ln0.z/(Section 13.2). Here and else- where in this chapter zmay be either real or complex. Replacing zwith zC1, we have 0.zC1/Dlimn!1123n .zC1/.zC2/.zC3/.zCnC1/nzC1 Dlimn!1nz zCnC1123n z.zC1/.zC2/.zCn/nz Dz0.z/: (13.2) This is the basic functional relation for the gamma function. It should be noted that it is a difference equation. Also, from the definition, 0.1/Dlimn!1123n 123n.nC1/nD1: (13.3) Now, repeated application of Eq. (13.2) gives 0.2/D1; 0.3/D20.2/D2; 0.4/D30.3/D23; etc., so 0.n/D123.n1/D.n1/W: (13.4) Definite Integral (Euler) A second definition, also frequently called the Euler integral, and already presented in Table 1.2, is 0.z/1Z 0ettz1dt;<e.z/>0: (13.5) The restriction on zis necessary to avoid divergence of the integral. When the gamma function does appear in physical problems, it is often in this form or some variation, such as 0.z/D21Z 0et2t2z1dt;<e.z/>0; (13.6) or 0.z/D1Z 0 ln1 tz1 dt;<e.z/>0: (13.7) ArfKen_Ch13-9780123846549.tex 13.1 De/f_initions, Properties 601 When zD1 2, Eq. (13.6) is just the Gauss error integral, and, cf. Eq. (1.148), we have the interesting result 01 2 Dp: (13.8) Generalizations of Eq. (13.6), the Gaussian integrals, are considered in Exercise 13.1.10. To show the equivalence of these two definitions, Eqs. (13.1) and(13.5), consider the function of two variables F.z;n/DnZ 0 1t nn tz1dt;<e.z/>0; (13.9) with na positive integer. This form was chosen because the exponential has the definition limn!1 1t nn et: (13.10) Inserting Eq. (13.10) into Eq. (13.9), we see that the infinite- nlimit of F.z;n/corresponds to0.z/as given by Eq. (13.5): limn!1F.z;n/DF.z;1/D1Z 0ettz1dt0.z/: (13.11) Our remaining task is to identify this limit also with Eq. (13.1). Returning to F.z;n/, we evaluate it by carrying out successive integrations by parts. For convenience we make the substitution uDt=n. Then F.z;n/Dnz1Z 0.1u/nuz1du: (13.12) The first integration by parts yields F.z;n/ nzD.1u/nuz z 1 0Cn z1Z 0.1u/n1uzduI (13.13) note that (because z6D0) the integrated part vanishes at both endpoints. Repeating this n times, with the integrated part vanishing at both endpoints each time, we finally get F.z;n/Dnz n.n1/1 z.zC1/.zCn1/1Z 0uzCn1du D123n z.zC1/.zC2/.zCn/nz: (13.14) This is identical with the expression on the right side of Eq. (13.1). Hence limn!1F.z;n/DF.z;1/0.z/; where0.z/is in the form given by Eq. (13.1), thereby completing the proof. ArfKen_Ch13-9780123846549.tex 602 Chapter 13 Gamma Function Infinite Product (Weierstrass) The third definition (Weierstrass’ form) is the infinite product 1 0.z/ze z1Y nD1 1Cz n ez=n; (13.15) where is the Euler-Mascheroni constant D0:5772156619; (13.16) which was introduced as a limit in Eq. (1.13). Existence of the limit was the topic of Exercise 1.2.13. This infinite-product form is useful for proving various properties of 0.z/. It can be derived from the original definition, Eq. (13.1), by rewriting it as 0.z/Dlimn!1123n z.zC1/.zCn/nzDlimn!11 znY mD1 1Cz m1 nz: (13.17) Taking the reciprocal of Eq. (13.17) and using nzDe.lnn/z; (13.18) we obtain 1 0.z/Dzlimn!1e.lnn/znY mD1 1Cz m : (13.19) Multiplying and dividing the right-hand side of Eq. (13.19) by exp 1C1 2C1 3CC1 n z DnY mD1ez=m; (13.20) we get 1 0.z/Dz limn!1exp 1C1 2C1 3CC1 nlnn z " limn!1nY mD1 1Cz m ez=m# :(13.21) Comparing with Eq. (1.13), we see that the parenthesized quantity in the exponent approaches as a limit the Euler-Mascheroni constant, thereby confirming Eq. (13.15). ArfKen_Ch13-9780123846549.tex 13.1 De/f_initions, Properties 603 Functional Relations In Eq. (13.2) we already obtained the most important functional relation for the gamma function, 0.zC1/Dz0.z/: (13.22) Viewed as a complex-valued function, this formula permits the extension to negative zof values obtained via numerical evaluation of the integral representation, Eq. (13.5). While the Euler limit formula already tells us that 0.z/is an analytic function for all zexcept 0, 1;::: , stepwise extrapolation from the integral is a more efficient numerical approach. The gamma function satisfies several other functional relations, of which one of the most interesting is the reflection formula, 0.z/0.1z/D sinz: (13.23) This relation connects (for nonintegral z) values of0.z/that are related by reflection about the line zD1=2. One way to prove the reflection formula starts from the product of Euler integrals, 0.zC1/0.1z/D1Z 0szesds1Z 0tzetdt D1Z 0vzdv .vC1/21Z 0u eudu: (13.24) In obtaining the second line of Eq. (13.24) we transformed from the variables s;tto uDsCt;vDs=t;as suggested by combining the exponentials and the powers in the integrands. We also needed to insert the Jacobian of this transformation, J1D 1 1 1 ts t2 DsCt t2D.vC1/2 uI the final substitution becomes obvious if we note that vC1Du=t. Returning to Eq. (13.24), the uintegration is elementary, being equal to 1!, while thevintegration can be evaluated by contour-integration methods; it was the topic of Exercise 11.8.20, and has the value 1Z 0vzdv .vC1/2Dz sinz: (13.25) Using these results, and then replacing 0.zC1/inEq. (13.24) byz0.z/and canceling z from the two sides of the resulting equation, we complete the demonstration of Eq. (13.23). A special case of Eq. (13.23) results if we set zD1=2. Then (taking the positive square root), we get 01 2 Dp; (13.26) in agreement with Eq. (13.8). ArfKen_Ch13-9780123846549.tex 604 Chapter 13 Gamma Function Another functional relation is Legendre’s duplication formula, 0.1Cz/0 zC1 2 D22zp 0.2 zC1/; (13.27) which we prove for general zin Section 13.3. However, it is instructive to prove it now for integer values of z. Assuming zto be a nonnegative integer n, we start the proof by writing 0.nC1/DnW,0.2nC1/D.2n/W, and 0 nC1 2 D01 2 1 23 22n1 2 Dp13.2n1/ 2nDp.2n1/WW 2n; (13.28) where we have used Eq. (13.26) and the double factorial notation first introduced in Eqs. (1.75) and (1.76). The double factorial notation is used frequently enough in physics applications that a familiarity with it is essential, and will from here on be used without comment. Making the further observation that nWD2n.2n/WW,Eq. (13.27) follows directly. Incidentally, we call attention to the fact that gamma functions with half-integer argu- ments appear frequently in physics problems, and Eq. (13.28) shows how to write them in closed form. Analytic Properties The Weierstrass definition shows immediately that 0.z/has simple poles at zD0,1,2, 3,:::and thatT0.z/U1has no poles in the finite complex plane, which means that 0.z/ has no zeros. This behavior may also be seen in Eq. (13.23), if we note that =.sinz/is never equal to zero. A plot of 0.z/for real zis shown in Fig. 13.1. We note sign changes for each unit interval of negative z, that0.1/D0.2/D1, and that the gamma function has a minimum between zD1andzD2, atz0D0:46143:::, with0.z0/D0:88560:::. The residues Rnat the poles zDn (nan integer0) are RnDlim "!0 "0.nC"/ Dlim "!0"0.nC1C"/ nC"Dlim "!0"0.nC2C"/ .nC"/.nC1C"/ Dlim "!0"0.1C"/ .nC"/."/D.1/n nW; (13.29) showing that the residues alternate in sign, with that at zDn having magnitude 1=nW. Schlaefli Integral A contour integral representation of the gamma function that we will find useful in devel- oping asymptotic series for the Bessel functions is the Schlaefli integral Z CettdtD.e2i1/0.C1/; (13.30) ArfKen_Ch13-9780123846549.tex 13.1 De/f_initions, Properties 605 5G(x+1) 4 3 2 1 12 345 −5−4−3−2−1−1 −3 −5x FIGURE 13.1 Gamma function 0.xC1/for real x. 0Cut line A BDxy ε +∞ FIGURE 13.2 Gamma function contour. where Cis the contour shown in Fig. 13.2. This contour integral representation is only useful when is not an integer. For integer , the integrand is an entire function; both sides of Eq. (13.30) vanish and it yields no information. However, for noninteger ,tD0 is a branch point of the integrand and the right-hand side of Eq. (13.30) then evaluates to a nonzero result. Note that, unlike the contour representations we considered in earlier chapters, the present contour is open; we cannot close it at zDC1 because of the branch cut, nor can we close it with a large circle, as etbecomes infinite in the limit of large negative t. To verify Eq. (13.30), we proceed (for C1>0) by evaluating the contributions from the various parts of the integration path. The integral from 1toC"on the real axis yields 0.C1/, choosing arg.z/D0. The integralC"to1(in the fourth quadrant) then yields e2i0.C1/, the argument of zhaving increased to 2. Since the circle around the origin contributes nothing when >1,Eq. (13.30) follows. Now that this equation is established, we can deform the contour as desired (providing that we avoid the branch point and cut), since there are no other singularities we must avoid. ArfKen_Ch13-9780123846549.tex 606 Chapter 13 Gamma Function It is often convenient to cast Eq. (13.30) into the more symmetrical form Z CettdtD2iei0.C1/sin./; (13.31) where Ccan be the contour of Fig. 13.2 or any deformation thereof that encircles the origin, does not cross the branch cut, and begins and ends at any points respectively above and below the cut for which xDC1 . The above analysis establishes Eqs. (13.30) and(13.31) for>1. However, we note that the integral exists for <1as long as we stay away from the origin, and therefore it remains valid for all nonintegral . What we have found is that this contour integral representation provides an analytic continuation of the Euler integral, Eq. (13.5), to all nonintegral. Factorial Notation Our discussion of the gamma function has been presented in terms of the classical notation, which was first introduced by Legendre. In an attempt to make a closer correspondence to the factorial notation (traditionally used for integers), and to simplify the Euler integral representation of the gamma function, Eq. (13.5), some authors have chosen to use the notation zWas a synonym for 0.zC1/even when zhas an arbitrary complex value. Occa- sionally one even encounters Gauss’ notation,Q.z/, for the factorial function: Y .z/DzWD0.zC1/: Neither the factorial (for nonintegral arguments) nor the Gauss notation are currently favored by most serious investigators, and we will not use them in this book. Example 13.1.1 MAXWELL-BOLTZMANN DISTRIBUTION In classical statistical mechanics, a state of energy Eis occupied, according to the equa- tion of Maxwell-Boltzmann statistics, with a probability proportional to eE=kT, where k is Boltzmann’s constant and Tis the absolute temperature; it is usual to define D1=kT and to write the probability of occupancy of a state of energy Easp.E/DCe E. If the number of states in a small energy interval d Eat energy Eis given, using a density distri- bution function n.E/, asn.E/d E, then the total probability of states at energy Eassumes the form C n.E/e Ed E. Under those conditions, the total probability of occupancy in anystate (namely, unity) must be 1DCZ n.E/e Ed E; (13.32) which enables us to set the normalization constant C, and the average energy hEiof such a classical system will be hEiDCZ E n.E/e Ed E: (13.33) ArfKen_Ch13-9780123846549.tex 13.1 De/f_initions, Properties 607 For a structureless ideal gas, it can be shown that n.E/is proportional to E1=2, with E, the kinetic energy of a gas molecule, in the range .0;1/. Then we may find Cfrom 1DC1Z 0E1=2e Ed EDC0.3 2/ 3=2DCp 2 3=2;orCD2 3=2 p; and hEiDC1Z 0E3=2e Ed EDC0.5 2/ 5=2D2 3=2 pp 5=21 23 2 D3 2kT; the known value of the average kinetic energy per molecule for a structureless classical gas at temperature T. In probability theory, the distribution used here is known as a gamma distribution; it is further discussed in Chapter 23. Exercises 13.1.1 Derive the recurrence relations 0.zC1/Dz0.z/ from the Euler integral, Eq. (13.5), 0.z/D1Z 0ettz1dt: 13.1.2 In a power-series solution for the Legendre functions of the second kind we encounter the expression .nC1/.nC2/.nC3/.nC2s1/.nC2s/ 2468.2s2/.2s/.2nC3/.2nC5/.2nC7/.2nC2sC1/; in which sis a positive integer. (a) Rewrite this expression in terms of factorials. (b) Rewrite this expression using Pochhammer symbols; see Eq. (1.72). 13.1.3 Show that0.z/may be written 0.z/D21Z 0et2t2z1dt;<e.z/>0; 0.z/D1Z 0 ln1 tz1 dt;<e.z/>0: ArfKen_Ch13-9780123846549.tex 608 Chapter 13 Gamma Function 13.1.4 In a Maxwellian distribution the fraction of particles of mass mwith speed between v andvCdvis d N ND4m 2kT3=2 exp mv2 2kT v2dv; where Nis the total number of particles, kis Boltzmann’s constant, and Tis the absolute temperature. The average or expectation value of vnis defined ashvniD N1R vnd N. Show that hvniD2kT mn=20.nC3 2/ 0.3 2/: This is an extension of Example 13.1.1, in which the distribution was in kinetic energy EDmv2=2, with d EDmvdv. 13.1.5 By transforming the integral into a gamma function, show that 1Z 0xklnx dxD1 .kC1/2;k>1: 13.1.6 Show that 1Z 0ex4dxD05 4 : 13.1.7 Show that lim x!00.ax/ 0.x/D1 a: 13.1.8 Locate the poles of 0.z/. Show that they are simple poles and determine the residues. 13.1.9 Show that the equation 0.x/Dk,k6D0, has an infinite number of real roots. 13.1.10 Show that, for integer s, (a)1Z 0x2sC1exp. ax2/dxDsW 2asC1. (b)1Z 0x2sexp. ax2/dxD0.sC1 2/ 2asC1=2D.2s1/WW 2sC1asr a. These Gaussian integrals are of major importance in statistical mechanics. 13.1.11 Express the coefficient of the nth term of the expansion of .1Cx/1=2in powers of x (a) in terms of factorials of integers, (b) in terms of the double factorial ( WW) functions. ArfKen_Ch13-9780123846549.tex 13.1 De/f_initions, Properties 609 ANS. anD.1/nC1.2n3/W 22n2nW.n2/WD.1/nC1.2n3/WW .2n/WW,nD2;3;:::: 13.1.12 Express the coefficient of the nth term of the expansion of .1Cx/1=2in powers of x (a) in terms of the factorials of integers, (b) in terms of the double factorial ( WW) functions. ANS. anD.1/n.2n/W 22n.nW/2D.1/n.2n1/WW .2n/WW;nD1;2;3:::: 13.1.13 The Legendre polynomial Pnmay be written as Pn.cos/D2.2n1/WW .2n/WW cosnC1 1n 2n1cos.n2/ C13 12n.n1/ .2n1/.2n3/cos.n4/ C135 123n.n1/.n2/ .2n1/.2n3/.2n5/cos.n6/C : LetnD2sC1. Then the above can be written Pn.cos/DP2sC1.cos/DsX mD0amcos.2mC1/: Find amin terms of factorials and double factorials. 13.1.14 (a) Show that 01 2n 01 2Cn D.1/n;where nis an integer. (b) Express 01 2Cn and0.1 2n/separately in terms of 1=2and a double factorial function. ANS.0.1 2Cn/D.2n1/WW 2n1=2. 13.1.15 Show that if0.xCiy/DuCiv;then0.xiy/Duiv: This is a special case of the Schwarz reflection principle, Section 11.10. 13.1.16 Prove thatj0. Ci /jDj0. /j1Y nD0 1C 2 . Cn/21=2 : This equation has been useful in calculations of beta decay theory. 13.1.17 Show that for n, a positive integer, j0.nCibC1/jDb sinhb1=2 nY sD1.s2Cb2/1=2: 13.1.18 Show that for all real values of xandy,j0.x/jj0. xCiy/j: ArfKen_Ch13-9780123846549.tex 610 Chapter 13 Gamma Function 13.1.19 Show thatj.0.1 2Ciy/j2D coshy: 13.1.20 The probability density associated with the normal distribution of statistics is given by f.x/D1 .2/1=2exp .x/2 22 ; with.1;1/for the range of x. Show that (a)hxi, the mean value of x, is equal to, (b) the standard deviation .hx2ihxi2/1=2is given by. 13.1.21 For the gamma distribution f.x/D8 >< >:1 0. /x 1ex= ;x>0; 0; x0; show that (a)hxi, the mean value of x, is equal to , (b)2, its variance, defined as hx2ihxi2, has the value 2. 13.1.22 The wave function of a particle scattered by a Coulomb potential is .r;/. Given that at the origin the wave function becomes .0/De =20.1Ci /; where >0is a dimensionless parameter, show that j .0/j2D2 e2 1: 13.1.23 Derive the contour integral representation of Eq. (13.31), 2i0.C1/sinDZ Cet.t/dt: 13.2 D IGAMMA AND POLYGAMMA FUNCTIONS Digamma Function As may be noted from the three definitions in Section 13.1, it is inconvenient to deal with the derivatives of the gamma function directly. It is more productive to take the natural logarithm of the gamma function as given by Eq. (13.1), thereby converting the product to a sum, and then to differentiate. The most useful results are obtained if we start with 0.zC1/: ArfKen_Ch13-9780123846549.tex 13.2 Digamma and Polygamma Functions 611 0.zC1/Dz0.z/Dlimn!1nW .zC1/.zC2/.zCn/nz; (13.34) ln0.zC1/Dlimn!1h ln.nW/Czlnnln.zC1/ ln.zC2/ ln.zCn/i ; (13.35) in which the logarithm of the limit is equal to the limit of the logarithm. Differentiating with respect to z, we obtain d dzln0.zC1/ .zC1/Dlimn!1 lnn1 zC11 zC21 zCn ;(13.36) which defines .zC1/, the digamma function. Note that this definition also corre- sponds to .zC1/DT0.zC1/U0 0.zC1/: (13.37) To bring Eq. (13.36) to a better form, we add and subtract the harmonic number HnDnX mD11 m; thereby obtaining .zC1/Dlimn!1" .lnnHn/nX mD11 zCm1 m# D C1X mD1z m.mCz/: (13.38) We have now arranged the contributions in a way that causes each group of terms to approach a finite limit as n!1 : in that limit lnnHnbecame (minus) the Euler- Mascheroni constant, defined in Eq. (1.13), and the summation is convergent. Setting zD0, we find1 .1/D D0:577 215 664 901 : (13.39) For integer n>0, Eq. (13.38) reduces to a form that is good for revealing its structure but less desirable for actual computation: .nC1/D CHnD CnX mD11 m: (13.40) 1 has been computed to 1271 places by D. E. Knuth, Math. Comput. 16: 275 (1962), and to 3566 decimal places by D. W. Sweeney, ibid. 17: 170 (1963). It may be of interest that the fraction 228=395 gives accurate to six places. ArfKen_Ch13-9780123846549.tex 612 Chapter 13 Gamma Function Polygamma Function The digamma function may be differentiated repeatedly, giving rise to the polygamma function: .m/.zC1/dmC1 dzmC1ln0.zC1/ D.1/mC1mW1X nD11 .zCn/mC1;mD1;2;3;:::: (13.41) Plots of0.x/, .x/, and 0.x/are presented in Fig. 13.3. If we set zD0inEq. (13.41), the series in that equation is that defining the Riemann zeta function,2 .m/1X nD11 nm; (13.42) and we have .m/.1/D.1/mC1mW.mC1/; mD1;2;3;:::: (13.43) The values of polygamma functions of the positive integral argument, .m/.nC1/, may be calculated recursively; see Exercise 13.2.8. 6 5 4 3 2 1 0 −1.0 −1.0 0 1.0 2.0 3.0 4.0 xIn Γ(x+1)d dx In Γ(x+1)d dxIn Γ(x+1)d2 dx2In Γ(x+1)d2 dx2Γ(x+1) FIGURE 13.3 Gamma function and its first two logarithmic derivatives. 2Forz6D0this series has been used to define a generalization of .m/known as the Hurwitz zeta function. ArfKen_Ch13-9780123846549.tex 13.2 Digamma and Polygamma Functions 613 Maclaurin Expansion It is now possible to write a Maclaurin expansion for ln0.zC1/: ln0.zC1/D1X nD1zn nW .n1/.1/D zC1X nD2.1/nzn n.n/: (13.44) This expansion is convergent for jzj<1; for zDx, the range is1<x1. Alternate forms of this series appear in Exercise 13.2.2. Equation (13.44) is a possible means of computing0.zC1/for real or complex z, but Stirling’s series (Section 13.4) is usually better, and in addition, an excellent table of values of the gamma function for complex arguments based on the use of Stirling’s series and the functional relation, Eq. (13.22), is now available.3 Series Summation The digamma and polygamma functions may also be used in summing series. If the general term of the series has the form of a rational fraction (with the highest power of the index in the numerator at least two less than the highest power of the index in the denominator), it may be transformed by the method of partial fractions; see Eq. (1.83). This transformation permits the infinite series to be expressed as a finite sum of digamma and polygamma functions. The usefulness of this method depends on the availability of tables of digamma and polygamma functions. Such tables and examples of series summation are given in AMS-55, chapter 6 (see Additional Readings for the reference). Example 13.2.1 CATALAN’S CONSTANT Catalan’s constant, .2/ , Eq. (12.65), is given by KD .2/D1X kD0.1/k .2kC1/2: Grouping the positive and negative terms separately and starting with the unit index, to match the form of .1/,Eq. (13.41), we obtain KD1C1X nD11 .4nC1/21 91X nD11 .4nC3/2: Now, identifying the summations in terms of .1/, we get KD8 9C1 16 .1/ 1C1 4 1 16 .1/ 1C3 4 : 3Table of the Gamma Function for Complex Arguments, Applied Mathematics Series No. 34. Washington, DC: National Bureau of Standards (1954). ArfKen_Ch13-9780123846549.tex 614 Chapter 13 Gamma Function Using the values of .1/from Table 6.1 of AMS-55 (see Additional Readings for the reference), we obtain KD0:91596559:::: Compare this calculation of Catalan’s constant with those carried out in earlier chapters (Exercises 1.1.12 and 12.4.4). Exercises 13.2.1 For “small” values of x; ln0.xC1/D xC1X nD2.1/n.n/ nxn; where is the Euler-Mascheroni constant and .n/the Riemann zeta function. For what values of xdoes this series converge? ANS.1<x1. Note that if xD1, we obtain D1X nD2.1/n.n/ n; a series for the Euler-Mascheroni constant. The convergence of this series is exceed- ingly slow. For actual computation of , other, indirect, approaches are far superior (see Exercise 12.3.2). 13.2.2 Show that the series expansion of ln0.xC1/(Exercise 13.2.1) may be written as (a) ln0.xC1/D1 2lnx sinx x1X nD1.2nC1/ 2nC1x2nC1, (b) ln0.xC1/D1 2lnx sinx 1 2ln1Cx 1x C.1 /x 1X nD1h .2nC1/1ix2nC1 2nC1: Determine the range of convergence of each of these expressions. 13.2.3 Verify that for n, a positive integer, the following two forms of the digamma function are equal to each other: .nC1/DnX jD11 j and .nC1/D1X jD1n j.nCj/ : ArfKen_Ch13-9780123846549.tex 13.2 Digamma and Polygamma Functions 615 13.2.4 Show that .zC1/has the series expansion .zC1/D C1X nD2.1/n.n/zn1: 13.2.5 For a power-series expansion of ln0.zC1/, AMS-55 (see Additional Readings for the reference) lists ln0.zC1/Dln.1Cz/Cz.1 /C1X nD2.1/nh .n/1izn n: (a) Show that this agrees with Eq. (13.44) for jzj<1. (b) What is the range of convergence of this new expression? 13.2.6 Show that 1 2lnz sinz D1X nD1.2n/ 2nz2n;jzj<1: Hint. UseEqs. (13.23) and(13.35). 13.2.7 Write out a Weierstrass infinite-product definition of ln0.zC1/. Without differentiat- ing, show that this leads directly to the Maclaurin expansion of ln0.zC1/,Eq. (13.44). 13.2.8 Derive the difference relation for the polygamma function, .m/.zC2/D .m/.zC1/C.1/m mW .zC1/mC1;mD0;1;2;:::: 13.2.9 The Pochhammer symbol .a/nis defined (for integral n) as .a/nDa.aC1/.aCn1/; . a/0D1: (a) Express .a/nin terms of factorials. (b) Find.d=da/.a/nin terms of.a/nand digamma functions. ANS.d da.a/nD.a/nT .aCn/ .a/U: (c) Show that .a/nCkD.aCn/k.a/n: ArfKen_Ch13-9780123846549.tex 616 Chapter 13 Gamma Function 13.2.10 Verify the following special values of the form of the digamma and polygamma functions: .1/D ; .1/.1/D.2/; .2/.1/D2.3/: 13.2.11 Verify: (a)1Z 0erlnr drD . (b)1Z 0rerlnr drD1 . (c)1Z 0rnerlnr drD.n1/WC n1Z 0rn1erlnr dr;nD1;2;3;:::. Hint. These may be verified by integration by parts, or by differentiating the Euler integral formula for 0.nC1/with respect to n. 13.2.12 Dirac relativistic wave functions for hydrogen involve factors such as 0T2.1 2Z2/1=2C1Uwhere , the fine structure constant, is 1/137 and Zis the atomic number. Expand0T2.1 2Z2/1=2C1Uin a series of powers of 2Z2. 13.2.13 The quantum mechanical description of a particle in a Coulomb field requires a knowl- edge of the argument of 0.z/when zis complex. Determine the argument of 0.1Cib/ for small, real b. 13.2.14 Using digamma and polygamma functions, sum the series (a)1X nD11 n.nC1/;(b)1X nD21 n21: Note. You can use Exercise 13.2.8 to calculate the needed digamma functions. 13.2.15 Show that 1X nD11 .nCa/.nCb/D1 .ba/h .1Cb/ .1Ca/i ; where a6Db, and neither anorbis a negative integer. It is of some interest to compare this summation with the corresponding integral, 1Z 1dx .xCa/.xCb/D1 bah ln.1Cb/ln.1Ca/i : The relation between .x/andlnxis made explicit in the analysis leading to Stirling’s formula. ArfKen_Ch13-9780123846549.tex 13.3 The Beta Function 617 13.3 T HEBETA FUNCTION Products of gamma functions can be identified as describing an important class of definite integrals involving powers of sine and cosine functions, and these integrals, in turn, can be further manipulated to evaluate a large number of algebraic definite integrals. These properties make it useful to define the beta function, defined as B.p;q/D0.p/0.q/ 0.pCq/: (13.45) For whatever it is worth, note that the Bin Eq. (13.45) is an upper-case beta. To understand the virtue of this definition, let us write the product 0.p/0.q/using the integral representation given as Eq. (13.6), valid for <e.p/;<e.q/>0: 0.p/0.q/D41Z 0s2p1es2ds1Z 0t2q1et2dt: (13.46) The reason for using this integral representation is that the quadratic terms in the exponent, s2andt2, combine in a convenient way if we change the integration variables from s;t to polar coordinates r;, with sDrcos,tDrsin,r2Ds2Ct2, and ds dtDr dr d. Equation (13.46) becomes 0.p/0.q/D41Z 0r2pC2q1er2dr=2Z 0cos2p1sin2q1d D20.pCq/=2Z 0cos2p1sin2q1d; where we have used Eq. (13.6) to recognize the rintegration as 0.pCq/. This gives us our first integral evaluation based on the beta function: B.p;q/D2=2Z 0cos2p1sin2q1d: (13.47) Because Eq. (13.47) is often used when pandqare integers, we rewrite for the case pDmC1,qDnC1, mWnW .mCnC1/WD2=2Z 0cos2mC1sin2nC1d: (13.48) Because gamma functions of a half-integral argument are available in closed form, Eq. (13.47) also provides a route to these trigonometric integrals for even powers of the sine and/or cosine. Note also that from its definition it is obvious that B.p;q/DB.q;p/, showing that the integral in Eq. (13.47) does not change in value if the powers of the sine and cosine are interchanged. ArfKen_Ch13-9780123846549.tex 618 Chapter 13 Gamma Function Alternate Forms, Definite Integrals The substitution tDcos2converts Eq. (13.47) to B.pC1;qC1/D1Z 0tp.1t/qdt: (13.49) Replacing tbyx2, we obtain B.pD1;qC1/D21Z 0x2pC1.1x2/qdx: (13.50) The substitution tDu=.1Cu/inEq. (13.49) yields still another useful form, B.pC1;qC1/D1Z 0up .1Cu/pCqC2du: (13.51) The beta function as a definite integral is useful in establishing integral representations of the Bessel function (Exercise 14.1.17) and the hypergeometric function (Exercise 18.5.12). Derivation of Legendre Duplication Formula The Legendre duplication formula involves products of gamma functions, which suggests that the beta function may provide a useful route to its proof. We start by using Eq. (13.49) forB zC1 2;zC1 2 : B zC1 2;zC1 2 D1Z 0tz1=2.1t/z1=2dt: (13.52) Making the substitution tD.1Cs/=2, we have B zC1 2;zC1 2 D22z1Z 1.1s2/z1=2ds D22zC11Z 0.1s2/z1=2dsD22zB1 2;zC1 2 ; (13.53) where we used the fact that the sintegrand was even to change the integration range to .0;1/, and then used Eq. (13.50) to evaluate the resulting integral. Now, inserting the definition, Eq. (13.45), for both instances of Bin Eq. (13.53), we reach 0.zC1 2/0.zC1 2/ 0.2zC1/D22z0.1 2/0.zC1 2/ 0.zC1/; ArfKen_Ch13-9780123846549.tex 13.3 The Beta Function 619 which is easily rearranged into 0.zC1/0 zC1 2 Dp 22z0.2zC1/; (13.54) the Legendre duplication formula, originally introduced as Eq. (13.27), but proved at that time only for integer values of z. Although the integrals used in this derivation are defined only for <e.z/>1, the result, Eq. (13.54), holds, by analytic continuation, for all zwhere the gamma functions are analytic. Exercises 13.3.1 Verify the following beta function identities: (a) B.a;b/DB.aC1;b/CB.a;bC1/, (b) B.a;b/DaCb bB.a;bC1/, (c) B.a;b/Db1 aB.aC1;b1/, (d) B.a;b/B.aCb;c/DB.b;c/B.a;bCc/. 13.3.2 (a) Show that 1Z 1.1x2/1=2x2ndxD8 >< >:=2; nD0 .2n1/WW .2nC2/WW;nD1;2;3;:::: (b) Show that 1Z 1.1x2/1=2x2ndxD8 >< >:; nD0; .2n1/WW .2n/WW;nD1;2;3;:::: 13.3.3 Show that 1Z 1.1x2/ndxD2.2n/WW .2nC1/WW;nD0;1;2;:::: 13.3.4 Evaluate1Z 1.1Cx/a.1x/bdxin terms of the beta function. ANS. 2aCbC1B.aC1;bC1/. ArfKen_Ch13-9780123846549.tex 620 Chapter 13 Gamma Function 13.3.5 Show, by means of the beta function, that zZ tdx .zx/1 .xt/ D sin ;0< < 1: 13.3.6 Show that the Dirichlet integral Z Z xpyqdx dyDpWqW .pCqC2/WDB.pC1;qC1/ pCqC2; where the range of integration is the triangle bounded by the positive x- and y-axes and the line xCyD1. 13.3.7 Show that 1Z 01Z 0e.x2Cy2C2xycos/dx dyD 2 sin: What are the limits on ? Hint. Consider oblique xy-coordinates. ANS.<< . 13.3.8 Evaluate (using the beta function) (a)=2Z 0cos1=2dD.2/3=2 16T0.5=4/U2; (b)=2Z 0cosndD=2Z 0sinndDpT.n1/=2UW 2.n=2/W D8 >>< >>:.n1/WW nWWfornodd;  2.n1/WW nWWforneven: 13.3.9 Evaluate1Z 0.1x4/1=2dxas a beta function. ANS.T0.5=4/U24 .2/1=2D1:311028777 . 13.3.10 Using beta functions, show that the integral representation J.z/D2 1=20.C1 2/z 2=2Z 0sin2cos.zcos/d;<./>1 2; ArfKen_Ch13-9780123846549.tex 13.3 The Beta Function 621 reduces to the Bessel series J.z/D1X sD0.1/s 1 sW0.sCC1/z 22sC ; thereby confirming its validity. 13.3.11 Given the associated Legendre function, defined in Chapter 15, Pm m.x/D.2m1/WW.1x2/m=2; show that (a)1Z 1TPm m.x/U2dxD2 2mC1.2m/W; mD0;1;2;::: , (b)1Z 1TPm m.x/U2dx 1x2D2.2m1/W; mD1;2;3;::: . 13.3.12 Show that, for integers pandq, (a)1Z 0x2pC1.1x2/1=2dxD.2p/WW .2pC1/WW, (b)1Z 0x2p.1x2/qdxD.2p1/WW.2q/WW .2pC2qC1/WW: 13.3.13 A particle of mass mmoving in a symmetric potential that is well described by V.x/D Ajxjnhas a total energy1 2m.dx=dt/2CV.x/DE. Solving for dx=dtand integrating we find that the period of motion is D2p 2mxmaxZ 0dx .EAxn/1=2; where xmaxis a classical turning point given by Axn maxDE. Show that D2 nr 2m EE A1=n0.1= n/ 0.1= nC1 2/: 13.3.14 Referring to Exercise 13.3.13, (a) Determine the limit as n!1 of 2 nr 2m EE A1=n0.1= n/ 0.1= nC1 2/: ArfKen_Ch13-9780123846549.tex 622 Chapter 13 Gamma Function (b) Find limn!1from the behavior of the integrand .EAxn/1=2. (c) Investigate the behavior of the physical system (potential well) as n!1 . Obtain the period from inspection of this limiting physical system. 13.3.15 Show that 1Z 0sinh x cosh xdxD1 2B C1 2; 2 ;1< < : Hint. Letsinh2xDu. 13.3.16 The beta distribution of probability theory has a probability density f.x/D0. C / 0. /0. /x 1.1x/ 1; with xrestricted to the interval (0, 1). Show that (a)hxi;the mean value, is C . (b)2;its variance, ishx2ihxi2D . C /2. C C1/: 13.3.17 From limn!1Z=2 0sin2nd Z=2 0sin2nC1dD1; derive the Wallis formula for :  2D22 1344 3566 57: 13.4 S TIRLING ’SSERIES In statistical mechanics we encounter the need to evaluate ln.nW/for very large values of n, and we occasionally need ln0.z/for nonintegral zwhenjzjis large enough that it is inconvenient or impractical to use the Maclaurin series, Eq. (13.44), possibly followed by repeated use of the functional relation 0.zC1/Dz0.z/. These needs can be met by the asymptotic expansion for ln0.z/known as Stirling’s series orStirling’s formula. While it is in principle possible to develop such an asymptotic formula by the method of steepest descents (and in fact we have already obtained the leading term of the expansion in this way; see Example 12.7.1), a relatively simple way of obtaining the full asymptotic expansion is by use of the Euler-Maclaurin integration formula in Section 12.3. ArfKen_Ch13-9780123846549.tex 13.4 Stirling’s Series 623 Derivation from Euler-Maclaurin Integration Formula The Euler-Maclaurin formula for evaluating a definite integral on the range .0;1/, obtained by specializing Eq. (12.57) and ignoring the remainder, is 1Z 0f.x/dxD1 2f.0/Cf.1/Cf.2/Cf.3/C CB2 2Wf0.0/CB4 4Wf.3/.0/CB6 6Wf.5/.x/C; (13.55) where Bnare Bernoulli numbers: B2D1 6;B4D1 30;B6D1 42;B8D1 30;: We proceed by applying Eq. (13.55) to the definite integral 1Z 0dx .zCx/2D1 z (forznot on the negative real axis). We note, by comparing with Eq. (13.41), that f.1/Cf.2/CD1X nD11 .zCn/2D .1/.zC1/I this makes a connection to the gamma function and is the reason for our current strategy. We also note that f.2n1/.0/Dd dx2n11 .zCx/2 xD0D.2n/W z2nC1; so the expansion yields 1 zD1Z 0dx .zCx/2D1 2z2C .1/.zC1/B2 z3B4 z5: Solving for .1/.zC1/, we have .1/.zC1/Dd dz .zC1/D1 z1 2z2CB2 z3CB4 z5C D1 z1 2z2C1X nD1B2n z2nC1: (13.56) Since the Bernoulli numbers diverge strongly, this series does not converge. It is a semi- convergent, or asymptotic, series, useful if one retains a small number of terms (compare with Section 12.6). ArfKen_Ch13-9780123846549.tex 624 Chapter 13 Gamma Function Integrating once, we get the digamma function .zC1/DC1ClnzC1 2zB2 2z2B4 4z4 DC1ClnzC1 2z1X nD1B2n 2nz2n; (13.57) where C1has a value still to be determined. In the next subsection we will show that C1D0. Equation (13.57), then, gives us another expression for the digamma function, often more useful than Eq. (13.38) or Eq. (13.44). Stirling’s Formula The indefinite integral of the digamma function, obtained by integrating Eq. (13.57), is ln0.zC1/DC2C zC1 2 lnzC.C11/zCB2 2zCCB2n 2n.2n1/z2n1C; (13.58) in which C2is another constant of integration. We are now ready to determine C1and C2, which we can do by requiring that the asymptotic expansion be consistent with the Legendre duplication formula, Eq. (13.54). Substituting Eq. (13.58) into the logarithm of the duplication formula, we find that satisfaction of that formula dictates that C1D0and thatC2must have the value C2D1 2ln 2: (13.59) Thus, inserting also values of the B2n, our final result is ln0.zC1/D1 2ln 2C zC1 2 lnzzC1 12z1 360z3C1 1260 z5:(13.60) This is Stirling’s series, an asymptotic expansion. The absolute value of the error is less than the absolute value of the first term neglected. The leading term in the asymptotic behavior of the gamma function was one of the ex- amples used to illustrate the method of steepest descents. In Example 12.7.1, we found that 0.zC1/p 2zzC1=2ez; corresponding to ln0.zC1/1 2ln 2C zC1 2 lnzz; yielding all the terms of Eq. (13.60) that do not vanish in the limit of large jzj. To help convey a feeling of the remarkable precision of Stirling’s series for 0.sC1/; the ratio of the first term of Stirling’s approximation to 0.sC1/is plotted in Fig. 13.4. In Table 13.1 we give the ratio of the first term in the expansion to 0.sC1/and a similar ratio when two terms are kept in the expansion to 0.sC1/. The derivation of these forms isExercise 13.4.1. ArfKen_Ch13-9780123846549.tex 13.4 Stirling’s Series 625 1.02 1.00 0.99 0.98 0.97 0.96 0.95 0.94 0.93 0.92 123456789S2π ss+1/2 e−s(1+1 12s) 0.83% low Γ(s+1 ) 2π ss+1/2 e−s Γ(s+1 ) FIGURE 13.4 Accuracy of Stirling’s formula. Table 13.1 Ratios of One- and Two-Term Stirling Series to Exact Values of0.sC1/ s1 0.sC1/p 2ssC1=2es 1 0.sC1/p 2ssC1=2es 1C1 12s 1 0.92213 0.99898 2 0.95950 0.99949 3 0.97270 0.99972 4 0.97942 0.99983 5 0.98349 0.99988 6 0.98621 0.99992 7 0.98817 0.99994 8 0.98964 0.99995 9 0.99078 0.99996 10 0.99170 0.99998 Exercises 13.4.1 Rewrite Stirling’s series to give 0.zC1/instead of ln0.zC1/. ANS.0.zC1/Dp 2zzC1=2ez 1C1 12zC1 288z2139 51;840 z3C : 13.4.2 Use Stirling’s formula to estimate 52W, the number of possible rearrangements of cards in a standard deck of playing cards. 13.4.3 Show that the constants C1andC2in Stirling’s formula have the respective values zero and1 2ln 2 by using the logarithm of the Legendre duplication formula (see Fig. 3.4). ArfKen_Ch13-9780123846549.tex 626 Chapter 13 Gamma Function 13.4.4 Without using Stirling’s series show that (a) ln.nW/<nC1Z 1lnxdx, (b) ln.nW/>nZ 1lnxdx;nis an integer2. Note that the arithmetic mean of these two integrals gives a good approximation for Stirling’s series. 13.4.5 Test for convergence 1X pD0" 0.pC1 2/ pW#22pC1 2pC2D1X pD0.2p1/WW.2pC1/WW .2p/WW.2pC2/WW: This series arises in an attempt to describe the magnetic field created by and enclosed by a current loop. 13.4.6 Show that limx!1xba0.xCaC1/ 0.xCbC1/D1: 13.4.7 Show that limn!1.2n1/WW .2n/WWn1=2D1=2: 13.4.8 A set of Ndistinguishable particles is assigned to states i,iD1;2;:::; M:If the numbers of particles in the various states are n1,n2;:::; nM(with MN), the number of ways this can be done is WDNW n1Wn2WnMW: The entropy associated with this assignment is SDklnW, where kis Boltzmann’s constant. In the limit N!1 , with niDpiN(sopiis the fraction of the particles in state i), find Sas a function of Nand the pi. (a) In the limit of large N, find the entropy associated with an arbitrary set of ni. Is the entropy an extensive function of the system size (i.e., is it proportional to N)? (b) Find the set of pithat maximize S. Hint. Remember thatP ipiD1and that this is a constrained maximization (see Section 22.3). Note. These formulas correspond to classical, or Boltzmann, statistics. 13.5 R IEMANN ZETA FUNCTION We are now in a position to broaden our earlier survey of .z/, the Riemann zeta function. In so doing, we note an interesting degree of parallelism between some of the properties of .z/and corresponding properties of the gamma function. ArfKen_Ch13-9780123846549.tex 13.5 Riemann Zeta Function 627 We open this section by repeating the definition of .z/, which is valid when the series converges: .z/1X nD1nz: (13.61) The values of .n/for integral nfrom 2 to 10 were listed in Table 1.1 on page 17. We now want to consider the possibility of analytically continuing .z/beyond the range of convergence of Eq. (13.61). As a first step toward doing so, we prove the integral representation that was given in Table 1.1: .z/D1 0.z/1Z 0tz1dt et1: (13.62) Equation (13.62) has a range of validity that is limited by the behavior of its integrand at small t; since the denominator then approaches t, the overall small- tdependence is tz2. Writing zDxCiyandtz2Dtx2eiylnt, we see that, like Eq. (13.61), Eq. (13.62) will only converge when <e z>1. We start from the right-hand side of Eq. (13.62), denoted I, by multiplying the numer- ator and denominator of its integrand by etand expanding the denominator in powers of et, reaching ID1 0.z/1Z 0tz1etdt 1etD1 0.z/1Z 01X mD1tz1emtdt: We next change the variable of integration for the individual terms so that all terms contain an identical factor et: ID1 0.z/1Z 01X mD1t mz1 etdt m D1 0.z/ 1X mD11 mz!1Z 0tz1etdt D.z/1 0.z/1Z 0tz1etdtD.z/: (13.63) In the second line of Eq. (13.63) we recognize the summation as a zeta function and the integral as the Euler integral representation of 0.z/, Eq. (13.5). It then cancels against the initial factor 1=0. z/, leaving the desired final result, Eq. (13.62). In passing, we note that the only difference between the integral of Eq. (13.62) and the Euler integral for the gamma function is that we now have a denominator et1instead of simply et. The next step toward the analytic continuation we seek is to introduce a contour integral with the same integrand as Eq. (13.62), using the same open contour that was found useful for the gamma function, shown in Fig. 13.2. Just as for the gamma function, we do not wish to restrict zto integral values, so the integrand will in general have a branch point attD0, and again we have placed the branch cut on the positive real axis. Restricting consideration for now to zwith<e z>1, we evaluate the contour integral, denoted I, as the sum of its contributions from the sections of the contour, respectively, labeled A,B, ArfKen_Ch13-9780123846549.tex 628 Chapter 13 Gamma Function andDin Fig. 13.2. For<e z>1, the small circle Dmakes no contribution to the integral, while IAD1 0.z/"Z 1tz1dt et1D. z/; IBD1 0.z/1Z "tz1e2i.z1/dt et1De2i.z1/.z/De2iz.z/: Combining the above, we get ID1 0.z/Z Ctz1dt et1D e2iz1 .z/: (13.64) Note that Eq. (13.64) is useful as a relation involving .z/only if zis not an integer. We now wish to deform the contour of Eq. (13.64) in a way that will remove the restriction<e z>1, which we originally needed to obtain that equation. The deforma- tion corresponds to an analytic continuation of .z/to a larger range of z, and will be effective because the deformation can avoid the divergence in the neighborhood of tD0. When we consider possible deformations, we need to make the observation that, unlike the gamma function, the integrand of Eq. (13.64) has simple poles at the points tD2ni, nD1;2;::: , so that if we deform the contour in a way that encloses any of these poles, we must allow for the change thereby produced in the value of the contour integral. If we initially deform the contour by expanding the circle Dto some finite radius less than 2i, we do not change the value of the integral Ibut extend its range of validity to negative z. If, for z<0, we further expand Duntil it becomes an open circle of infinite radius (but not through any of the poles), the value of the contour integral is reduced to zero, with the change caused by the inclusion of the contribution from the poles that are then encircled. We therefore have the interesting result that the original contour integral had a value that was the negative of 2itimes the sum of the residues that were newly enclosed. Thus, ID e2iz1 .z/D2i 0.z/1X nD1(residues of tz1=.et1/attD2ni). At the pole tDC2 ni, the residue is 2nei=2z1, while at tD2 niit is 2ne3i=2z1. Note that we must evaluate the residues taking cognizance of the branch cut. Inserting these values and rearranging a bit,  e2iz1 .z/D 1X nD11 nzC1! .2/zi 0.z/ ei.z1/=2Ce3i.z1/=2 D.1z/.2/z 0.z/ e3iz=2eiz=2 : (13.65) Note that because z<0, the summation over nconverges and can be identified as .1z/. Equation (13.65) can be simplified, but we already see its essential feature, namely that it ArfKen_Ch13-9780123846549.tex 13.5 Riemann Zeta Function 629 provides a functional relation connecting .z/and.1z/, parallel to but more compli- cated than the reflection formula for the gamma function, Eq. (13.23). The derivation of Eq. (13.65) was carried out for z<0, but now that we have obtained it, we can, appeal- ing to analytic continuation, assert its validity for all zsuch that its constituent factors are nonsingular. This formula, in the simplified form we shall shortly obtain, was first found by Riemann. The simplification of Eq. (13.65) can be accomplished by recognizing, with the aid of the gamma-function reflection formula, Eq. (13.23), that e3iz=2eiz=2 e2iz1Dsin. z=2/ sinzD0.z/0.1z/ 0.z=2/0.1z=2/; so .z/D.1z/z2z0.1z/ 0.z=2/0.1z=2/D.1z/z1=20..1z/=2/ 0.z=2/; (13.66) where the final member of Eq. (13.66) was obtained by using the duplication formula, Eq. (13.27), with the value of zin the duplication formula set to the present z=2. Equa- tion (13.66) can now be rearranged to the more symmetrical form 0z 2 z=2.z/D01z 2 .1 z/=2.1z/: (13.67) Equation (13.67), the zeta-function reflection formula, enables generation of .z/in the half-plane<e z<0from values in the region <e z>1, where the series definition con- verges. It is possible to show that .z/has no zeros in the region where the series defi- nition converges, and, from Eq. (13.67), this implies that .z/is also nonzero for all zin the half-plane<e z<0except at points where 0.z=2/is singular, namely zD 2;4;:::;2n;:::: 0. z=2/is also singular at zD0but, as we shall see shortly, the singularity at .1/ compensates the singularity at 0.0/ , with the result that .0/ is nonzero. The zeros of .z/at the negative even integers are called its trivial zeros, as they arise from the singularities of the gamma function. Any other zeros of .z/(and there are an infinite number of them) must lie in the region 0<e z1, which has been called the critical strip of the Riemann zeta function. To obtain values of .z/in the critical strip, we proceed by analytically continuing toward<e zD0the formula from Eq. (12.62) that defines the Dirichlet series .z/(clearly valid for<e z>1), .z/D.z/ 121zD1 121z1X nD1.1/n1 nz: (13.68) This alternating series converges for all <e z>0, thereby providing a formula for .z/ throughout the critical strip, but it is best used where the convergence is relatively rapid, namely for<e z1 2. Values of.z/for<e z<1 2may be more conveniently obtained from those for<e z1 2using the reflection formula, Eq. (13.67). ArfKen_Ch13-9780123846549.tex 630 Chapter 13 Gamma Function Equation (13.68) may be used to verify that the singularity of .z/atzD1is a simple pole and to find its residue. We proceed as follows: (Residue at zD1)Dlim z!1.z1/.z/Dlim z!1z1 121z1X nD1.1/n1 n D1 ln 2 .ln 2/D1; (13.69) where we used l’Hôpital’s rule, recognized that d21z=dzD21zln 2, and identified the summation as that of Eq. (1.53). Returning now to Eq. (13.67), noting that lim z!0.1z/ 0.z=2/Dresidue of.s/atsD1 2(residue of0.s/atsD0)D1 2; we obtain the nonzero result .0/D0.1=2/1=2 1 2 D1 2: (13.70) In addition to the practical utility we have already noted for the Riemann zeta function, it plays a major role in current developments in analytic number theory. A starting point for such investigations is the celebrated Euler prime number product formula, which can be developed by forming .s/.12s/D1C1 2sC1 3sC1 2sC1 4sC1 6sC ; (13.71) eliminating all the ns, where nis a multiple of 2. Then we write .s/.12s/.13s/D1C1 3sC1 5sC1 7sC1 9sC 1 3sC1 9sC1 15sC ; eliminating all the remaining terms in which nis a multiple of 3. Continuing, we have .s/.12s/.13s/.15s/.1Ps/, where Pis a prime number, and all terms ns, in which nis a multiple of any integer up through P, are canceled out. In the limit P!1 , we reach .s/.12s/.13s/.1Ps/!.s/1Y P.prime/D2.1Ps/D1: Therefore .s/D1Y P.prime/D2.1Ps/1; (13.72) ArfKen_Ch13-9780123846549.tex 13.5 Riemann Zeta Function 631 giving.s/as an infinite product.4Incidentally, the cancellation procedure in the above derivation has a clear application in numerical computation. For example, Eq. (13.71) will give.s/.12s/to the same accuracy as Eq. (13.61) gives.s/, but with only half as many terms. The asymptotic distribution of prime numbers can be related to the poles of 0=, and in particular to the nontrivial zeros of the zeta function. Riemann conjectured that all the nontrivial zeros were on the critical line<e zD1 2, and there are potentially important results that can be proved if Riemann’s conjecture is correct. Numerical work has verified that the first 300109nontrivial zeros of .z/are simple and indeed fall on the critical line. See J. Van de Lune, H. J. J. Te Riele, and D. T. Winter, “On the zeros of the Riemann zeta function in the critical strip. IV,” Math. Comput. 47, 667 (1986). Although many gifted mathematicians have attempted to establish what has come to be known as the Riemann hypothesis, it has for about 150 years remained unproven and is considered one of the premier unsolved problems in modern mathematics. Pop- ular accounts of this fascinating problem can be found in M. du Santoy, The Music of the Primes: Searching to Solve the Greatest Mystery in Mathematics, New York: Harper- Collins (2003); J. Derbyshire, Prime Obsession: Bernhard Riemann and the Greatest Un- solved Problem in Mathematics, Washington, DC: Joseph Henry Press (2003); and K. Sab- bagh, The Riemann Hypothesis: The Greatest Unsolved Problem in Mathematics, New York: Farrar, Straus and Giroux (2003). Exercises 13.5.1 Show that the symmetrical functional relation 0z 2 z=2.z/D01z 2 .1 z/=2.1z/ follows from the equation  e2iz1 .z/D.1z/.2/z 0.z/ e3iz=2eiz=2 : 13.5.2 Prove that 1Z 0xnexdx .ex1/2DnW.n/: Assuming nto be real, show that each side of the equation diverges if nD1. Hence the preceding equation carries the condition n>1. Integrals such as this appear in the quantum theory of transport effects: thermal and electrical conductivity. 4For further discussion, the reader is referred to the works by Edwards, Ivíc, Patterson, and Titchmarsh in Additional Readings. ArfKen_Ch13-9780123846549.tex 632 Chapter 13 Gamma Function 13.5.3 The Bloch-Grüneisen approximation for the resistance in a monovalent metal at abso- lute temperature Tis DCT5 262=TZ 0x5dx .ex1/.1ex/; where2is the Debye temperature characteristic of the metal. (a) For T!1; show that C 4T 22: (b) For T!0, show that 5W.5/CT5 26: 13.5.4 Derive the following expansion of the Debye function for n1: xZ 0tndt et1Dxn" 1 nx 2.nC1/C1X kD1B2kx2k .2kCn/.2k/W# ,jxj<2: The complete integral .0;1/equals nW.nC1/(Exercise 13.5.6). 13.5.5 The total energy radiated by a blackbody is given by uD8k4T4 c3h31Z 0x3 ex1dx: Show that the integral in this expression is equal to 3W.4/. The final result is the Stefan- Boltzmann law. 13.5.6 As a generalization of the result in Exercise 13.5.5, show that 1Z 0xsdx ex1DsW.sC1/;<e.s/>0: 13.5.7 Prove that 1Z 0xsdx exC1DsW.12s/.sC1/;<e.s/>0: Exercises 13.5.6 and13.5.7 give the Mellin integral transform of 1=.ex1/; this trans- form is defined in Eq. (20.9). ArfKen_Ch13-9780123846549.tex 13.6 Other Related Functions 633 13.5.8 The neutrino energy density (Fermi distribution) in the early history of the universe is given by D4 h31Z 0x3 exp.x=kT/C1dx: Show that D75 30h3.kT/4: 13.5.9 Prove that .n/.z/D.1/nC11Z 0tnezt 1etdt;<e.z/>0: 13.5.10 Show that.s/is analytic in the entire finite complex plane except at sD1;where it has a simple pole with a residue of C1. Hint. The contour integral representation will be useful. 13.6 OTHER RELATED FUNCTIONS Incomplete Gamma Functions Generalizing the Euler-integral definition of the gamma function, Eq. (13.5), we define incomplete gamma functions by the variable-limit integrals .a;x/DxZ 0etta1dt;<.a/>0; (13.73) 0.a;x/D1Z xetta1dt: Clearly, these two functions are related, for .a;x/C0.a;x/D0.a/: (13.74) The choice of employing .a;x/or0.a;x/is purely a matter of convenience. If the parameter ais a positive integer, Eqs. (13.73) may be integrated completely to yield .n;x/D.n1/W 1exn1X sD0xs sW! ; (13.75) 0.n;x/D.n1/Wexn1X sD0xs sW: ArfKen_Ch13-9780123846549.tex 634 Chapter 13 Gamma Function While the above expressions are valid only for positive integer n, the function 0.n;x/is well defined (providing x>0) for nD0and corresponds to an exponential integral (see later subsection). For nonintegral a;a power-series expansion of .a;x/for small xand an asymptotic expansion of 0.a;x/are developed in Exercises 1.3.3 and 13.6.4: .a;x/Dxa1X nD0.1/n xn nW.aCn/;small x, 0.a;x/xa1ex1X nD00.a/ 0.an/1 xn(13.76) xa1ex1X nD0.an/n1 xn;large x, where.an/nis a Pochhammer symbol. The final expression in Eq. (13.76) makes it clear how to obtain an asymptotic expansion for 0.0; x/. Noting that .n/nD.1/nnW, we have 0.0; x/ex x1X nD0.1/nnW xn: (13.77) These incomplete gamma functions may also be expressed quite elegantly in terms of confluent hypergeometric functions (compare Section 18.6). Incomplete Beta Function Just as there are incomplete gamma functions, there is also an incomplete beta function, customarily defined for 0x1,p>0(and, if xD1, also q>0) as Bx.p;q/DxZ 0tp1.1t/q1dt: (13.78) Clearly, BxD1.p;q/becomes the regular (complete) beta function, Eq. (13.49). A power- series expansion of Bx.p;q/is the subject of Exercise 13.6.5. The relation to hypergeo- metric functions appears in Section 18.5. The incomplete beta function makes an appearance in probability theory in calculating the probability of at most ksuccesses in nindependent trials.5 Exponential Integral Although the incomplete gamma function 0.a;x/in its general form, Eq. (13.73), is only infrequently encountered in physical problems, a special case is quite common and very 5W. Feller, An Introduction to Probability Theory and Its Applications, 3rd ed. New York: Wiley (1968), Section VI.10. ArfKen_Ch13-9780123846549.tex 13.6 Other Related Functions 635 E1(x) 3 2 1 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6x FIGURE 13.5 The exponential integral, E1.x/DEi. x/. useful. We define the exponential integral by6 Ei. x/1Z xet tdtE1.x/: (13.79) For a graph of this function, see Fig. 13.5. To obtain a series expansion of E1.x/for small x, we will need to proceed with caution, because the integral in Eq. (13.78) diverges logarithmically as x!0. We start from E1.x/D0.0; x/Dlim a!0h 0.a/ .a;x/i : (13.80) Setting aD0in the convergent terms (those with n1) in the expansion of .a;x/and moving them outside the scope of the limiting process, we rearrange Eq. (13.80) to E1.x/Dlim a!0a0.a/xa a 1X nD1.1/nxn nnW: (13.81) Using l’Hôpital’s rule, Eq. (1.58), writing a0.a/D0.aC1/, and noting that d xa=daD xalnx, the limit in Eq. (13.81) reduces to hd da0.aC1/d daxai aD1D0.1/ .1/lnxD lnx; (13.82) where (without arguments) is the Euler-Mascheroni constant.7From Eqs. (13.81) and (13.82) we obtain the rapidly converging series E1.x/D lnx1X nD1.1/nxn nnW: (13.83) 6The appearance of the two minus signs in Ei. x/is a historical monstrosity. AMS-55, chapter 5, denotes this integral as E1.x/. See Additional Readings for the reference. 7Having the notations .a;x/and in the same discussion and with different meanings may seem unfortunate, but these are the traditional notations and should not lead to confusion if the reader is alert. ArfKen_Ch13-9780123846549.tex 636 Chapter 13 Gamma Function 1.0 10Ci(x) si(x) x −1.0 FIGURE 13.6 Sine and cosine integrals. The asymptotic expansion for E1.x/is simply that given in Eq. (13.77) for 0.0; x/. We repeat it here: E1.x/ex1 x1W x2C2W x33W x4C : (13.84) Further special forms related to the exponential integral are the sine integral, cosine integral (for both see Fig. 13.6), and the logarithmic integral, defined by8 si.x/D1Z xsint tdt; Ci.x/D1Z xcost tdt; (13.85) li.x/DxZ 0dt lntDEi.ln x/: Viewed as functions of a complex variable, Ci. z/and li. z/are multivalued, with a branch cut conventionally chosen to be along the negative real axis from the branch point at zD0. By transforming from real to imaginary argument, we can show that si.x/D1 2ih Ei.ix/Ei. ix/i D1 2ih E1.ix/E1.ix/i ; (13.86) whereas Ci.x/D1 2h Ei.ix/CEi. ix/i D1 2h E1.ix/CE1.ix/i ;jargxj< 2:(13.87) Adding these two relations, we obtain Ei.ix/DCi.x/Cisi.x/; (13.88) 8Another sine integral is denoted Si .x/Dsi.x/C=2. ArfKen_Ch13-9780123846549.tex 13.6 Other Related Functions 637 showing that the relation among these integrals is exactly analogous to that among eix, cosx, and sinx. In terms of E1, E1.ix/DCi. x/Cisi.x/: (13.89) Asymptotic expansions of Ci .x/and si.x/were developed in Section 12.6, with explicit formulas in Eqs. (12.93) and (12.94). Power-series expansions about the origin for Ci .x/, si.x/, and li.x/may be obtained from those for the exponential integral, E1.x/, or by direct integration, Exercise 13.6.13. The exponential, sine, and cosine integrals are tabulated in AMS-55, chapter 5 (see Additional Readings for the reference), and can also be accessed by symbolic software such as Mathematica, Maple, Mathcad, and Reduce. Error Function Theerror function erf.z/and the complementary error function erfc.z/are defined by the integrals erfzD2pzZ 0et2dt;erfczD1erfzD2p1Z zet2dt: (13.90) The factors 2=pcause these functions to be scaled so that erf 1D 1. For a plot of erf x, see Fig. 13.7. The power-series expansion of erf xfollows directly from the expansion of the expo- nential in the integrand: erfxD2p1X nD0.1/nx2nC1 .2nC1/nW: (13.91) Its asymptotic expansion, the subject of Exercise 12.6.3, is erfx1ex2 px 11 2x2C13 22x4135 23x6CC.1/n.2n1/WW 2nx2n :(13.92) xerf x −2 −1 −1121 FIGURE 13.7 Error function, erf x. ArfKen_Ch13-9780123846549.tex 638 Chapter 13 Gamma Function From the general form of the integrands and Eq. (13.6) we expect that erf zand erfc zmay be written as incomplete gamma functions with aD1 2. The relations are erfzD1=2 .1 2;z2/;erfczD1=20.1 2;z2/: (13.93) Exercises 13.6.1 Show that .a;x/Dex1X nD0.a1/W .aCn/WxaCn (a) by repeatedly integrating by parts, (b) by transforming it into Eq. (13.76). 13.6.2 Show that (a)dm dxmTxa .a;x/UD.1/mxam .aCm;x/, (b)dm dxmTex .a;x/UDex0.a/ 0.am/ .am;x/. 13.6.3 Show that .a;x/and0.a;x/satisfy the recurrence relations (a) .aC1;x/Da .a;x/xaex, (b)0.aC1;x/Da0.a;x/Cxaex. 13.6.4 Show that the asymptotic expansion (for large x) of the incomplete gamma function 0.a;x/has the form 0.a;x/xa1ex1X nD00.a/ 0.an/1 xn; and that the above expression is equivalent to 0.a;x/xa1ex1X nD0.an/n1 xn: 13.6.5 A series expansion of the incomplete beta function yields Bx.p;q/Dxp1 pC1q pC1xC.1q/.2q/ 2W.pC2/x2C C.1q/.2q/.nq/ nW.pCn/xnC : ArfKen_Ch13-9780123846549.tex 13.6 Other Related Functions 639 Given that 0x1;p>0, and q>0, test this series for convergence. What happens atxD1? 13.6.6 Using the definitions of the various functions, show that (a) si. x/D1 2iTE1.ix/E1.ix/U; (b) Ci. x/D1 2TE1.ix/CE1.ix/U; (c) E1.ix/DCi. x/Cisi.x/: 13.6.7 The potential produced by a 1shydrogen electron is given by V.r/Dq 4" 0a01 2r .3;2r/C0.2; 2r/ : (a) For r1;show that V.r/Dq 4" 0a0 12 3r2C : (b) For r1;show that V.r/Dq 4" 0a01 r: Here ris expressed in units of a0, the Bohr radius. Note. V.r/is illustrated in Fig. 13.8. Distributed charge potentialPoint charge potential 1/r r FIGURE 13.8 Distributed charge potential produced by a 1shydrogen electron, Exercise 13.6.7. ArfKen_Ch13-9780123846549.tex 640 Chapter 13 Gamma Function 13.6.8 The potential produced by a 2phydrogen electron can be shown to be V.r/D1 4" 0q 24a01 r .5; r/C0.4; r/ 1 4" 0q 120a01 r3 .7; r/Cr20.2; r/ P2.cos/: Here ris expressed in units of a0, the Bohr radius. P2.cos/is a Legendre polynomial (Section 15.1). (a) For r1, show that V.r/D1 4" 0q a01 41 120r2P2.cos/C : (b) For r1, show that V.r/D1 4" 0q a0r 16 r2P2.cos/C : 13.6.9 Prove that the exponential integral has the expansion 1Z xet tdtD lnx1X nD1.1/nxn nnW; where is the Euler-Mascheroni constant. 13.6.10 Show that E1.z/may be written as E1.z/Dez1Z 0ezt 1Ctdt: Show also that we must impose the condition jargzj=2. 13.6.11 Related to the exponential integral by a simple change of variable is the function En.x/D1Z 1ext tndt: Show that En.x/satisfies the recurrence relation EnC1.x/D1 nexx nEn.x/;nD1;2;3;: 13.6.12 With En.x/as defined in Exercise 13.6.11, show that for n>1, En.0/D1=.n1/. 13.6.13 Develop the following power-series expansions: ArfKen_Ch13-9780123846549.tex 13.6 Other Related Functions 641 (a) si. x/D 2C1X nD0.1/nx2nC1 .2nC1/.2nC1/W; (b) Ci. x/D ClnxC1X nD1.1/nx2n 2n.2n/W: 13.6.14 An analysis of a center-fed linear antenna leads to the expression xZ 01cost tdt: Show that this is equal to ClnxCi.x/. 13.6.15 Using the relation 0.a/D .a;x/C0.a;x/; show that if .a;x/satisfies the relations of Exercise 13.6.2, then 0.a;x/must satisfy the same relations. 13.6.16 Forx>0, show that 1Z xtndt et1D1X kD1ekxxn kCnxn1 k2Cn.n1/xn2 k3CCnW knC1 . Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Contains a wealth of information about gamma functions, incomplete gamma functions, exponential integrals, error functions, and related functions in chapters 4 to 6. Artin, E., The Gamma Function (translated by M. Butler). New York: Holt, Rinehart and Winston (1964). Demon- strates that if a function f.x/is smooth (log convex) and equal to .n1/Wwhen xDnDinteger, it is the gamma function. Davis, H. T., Tables of the Higher Mathematical Functions. Bloomington, IN: Principia Press (1933). Volume I contains extensive information on the gamma function and the polygamma functions. Edwards, H. M., Riemann’s Zeta Function. New York: Academic Press (1974) and Dover (2003). Gradshteyn, I. S., and I. M. Ryzhik, Table of Integrals, Series, and Products. New York: Academic Press (1980). Ivi´c, A., The Riemann Zeta Function. New York: Wiley (1985). Luke, Y. L., The Special Functions and Their Approximations, Vol. 1. New York: Academic Press (1969). Luke, Y. L., Mathematical Functions and Their Approximations. New York: Academic Press (1975). This is an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Chapter 1 deals with the gamma function. Chapter 4 treats the incomplete gamma function and a host of related functions. Patterson, S. J., Introduction to the Theory of the Reimann Zeta Function. Cambridge: Cambridge University Press (1988). Titchmarsh, E. C., and D. R. Heath-Brown, The Theory of the Riemann Zeta-Function. Oxford: Clarendon Press (1986). A detailed, classic work. ArfKen_17-ch14-0643-0714- 9780123846549.tex CHAPTER 14 BESSEL FUNCTIONS Bessel functions appear in a wide variety of physical problems. In Section 9.4 we saw that separation of the Helmholtz, or wave, equation in circular cylindrical coordinates led to Bessel’s equation in the coordinate describing distance from the axis of the cylindrical system. In that same section, we also identified spherical Bessel functions (closely related to Bessel functions of half-integral order) in Helmholtz equations in spherical coordinates. In summarizing the forms of solutions to partial differential equations (PDEs) in these coordinate systems, we not only identified the original and spherical Bessel functions, but also those of imaginary argument (usually expressed as modified Bessel functions to avoid the explicit use of imaginary quantities). Since these PDEs can describe many types of problems ranging from stationary problems in quantum mechanics to those of spherical or cylindrical wave propagation, a good familiarity with Bessel functions is important to the practicing physicist. Often problems in physics involve integrals that can be identified as Bessel functions, even when the original problem did not explicitly involve cylindrical or spherical geom- etry. Moreover, Bessel and closely related functions form a rich area of mathematical analysis with many representations, many interesting and useful properties, and many interrelations. Some of the major interrelations are developed in the present chapter. In addition to the material presented here, we call attention to further relations in terms of confluent hypergeometric functions; see Section 18.6. 14.1 B ESSEL FUNCTIONS OF THE FIRST KIND,J.x/ Bessel functions of the first kind, normally labeled J, are those obtained by the Frobenius method for solution of the Bessel ODE, x2J00 Cx J0 C.x22/JD0: (14.1) The term “first kind” reflects the fact that J.x/includes the functions that, for non- negative integer , are regular at xD0. All solutions to the Bessel ordinary differential 643 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_17-ch14-0643-0714- 9780123846549.tex 644 Chapter 14 Bessel Functions equation (ODE) that are linearly independent of J.x/are irregular at xD0for all; a specific choice for a second solution is denoted Y.x/and is called a Bessel function of the second kind.1 Generating Function for Integral Order We start our detailed study of Bessel functions by introducing a generating function yield- ing the Jnfor integer n(of either sign). Because the Jnare not polynomials, the generating function cannot be found by the methods of Section 12.1, but we will be able to show that the functions defined by the generating function are indeed the solutions of the Bessel ODE obtained by the Frobenius method. Our generating function formula, a Laurent series, is g.x;t/De.x=2/.t1=t/D1X nD1Jn.x/tn: (14.2) Although the Bessel ODE is homogeneous and its solutions are of arbitrary scale, Eq. (14.2) fixes a specific scale for Jn.x/. To relate Eq. (14.2) to the Frobenius solution, Eq. (7.48), we manipulate the exponential as follows: g.x;t/Dext=2ex=2tD1X rD0x 2rtr rW1X sD0.1/sx 2sts sW D1X rD01X sD0.1/sx 2rCstrs rWsW: We now change the summation index rtonDrs, yielding g.x;t/D1X nD1"X s.1/s .nCs/WsWx 2nC2s# tn; (14.3) where the ssummation starts at max.0;n/. For n0, the coefficient of tnis seen to be Jn.x/D1X sD0.1/s sW.nCs/Wx 2nC2s : (14.4) Comparing with Eq. (7.48), we confirm that for n0,Jnas given by Eq. (14.4) is the Frobenius solution, at the specific scale given here. If now we replace nbyn, the summation in Eq. (14.3) becomes Jn.x/D1X sDn.1/s sW.sn/Wx 2nC2s I 1We use the notation of AMS-55, also used by Watson in his definitive treatise (for both sources, see Additional Readings). The Yare sometimes also called Neumann functions; for that reason some workers write them as N. They were denoted Nin previous editions of this book. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 645 changing stosCn, we reach Jn.x/D1X sD0.1/sCn sW.sCn/Wx 2nC2s D.1/nJn.x/(integral n); (14.5) confirming both that Jn.x/is a solution to the Bessel ODE and that it is linearly dependent onJn. If we now consider Jwithnonintegral, we get no information from the generating function, but the Frobenius method then gives linearly independent solutions for both C and, which are both solutions of the Bessel ODE, Eq. (14.1), for the same value of 2. Looking at the details of the development of Eqs. (7.46) to (7.48), we see that the generali- zation of Eq. (14.4) to noninteger is J.x/D1X sD0.1/s sW0.CsC1/x 2C2s ; .6D1;2;:::/; (14.6) and that J.x/as given in Eq. (14.6) is a solution to the Bessel ODE. For0the series of Eq. (14.6) is convergent for all x, and for small xis a practi- cal way to evaluate J.x/. Graphs of J0,J1, and J2are shown in Fig. 14.1. The Bessel functions oscillate but are notperiodic, except in the limit x!1 , with the amplitude of the oscillation decreasing asymptotically as x1=2. This behavior is discussed further in Section 14.6. Recurrence Relations The Bessel functions Jn.x/satisfy recurrence relations connecting functions of contigu- ousn, as well as some connecting the derivative J0 nto various Jn. Such recurrence rela- tions may all be obtained by operating on the series, Eq. (14.6), although this requires a bit of clairvoyance (or a lot of trial and error). However, if the recurrence relations are already known, their verification is straightforward; see Exercise 14.1.8. Our approach here will be to obtain them from the generating function g.x;t/, using a process similar to that illustrated in Example 12.1.2. x1.0J0(x) J1(x) J2(x) 123 67J 089 5 4 FIGURE 14.1 Bessel functions J0.x/,J1.x/, and J2.x/. ArfKen_17-ch14-0643-0714- 9780123846549.tex 646 Chapter 14 Bessel Functions We start by differentiating g.x;t/: @ @tg.x;t/Dx 2 1C1 t2 e.x=2/.t1=t/D1X nD1n Jn.x/tn1; @ @xg.x;t/D1 2 t1 t e.x=2/.t1=t/D1X nD1J0 n.x/tn: Inserting the right-hand side of Eq. (14.2) in place of the exponentials and equating the coefficients of equal powers of t(as illustrated in Example 12.1.2), we obtain the two basic Bessel-function recurrence formulas: Jn1.x/CJnC1.x/D2n xJn.x/; (14.7) Jn1.x/JnC1.x/D2J0 n.x/: (14.8) Because Eq. (14.7) is a three-term recurrence relation, its use to generate Jnwill require two starting values. For example, given J0andJ1, then J2(and any other integral order Jn including those for n<0) may be computed. An important special case of Eq. (14.8) is J0 0.x/DJ1.x/: (14.9) Equations (14.7) and (14.8) can also be combined (Exercise 14.1.4) to form the useful additional formulas d dx xnJn.x/ DxnJn1.x/; (14.10) d dx xnJn.x/ DxnJnC1.x/; (14.11) Jn.x/DJ0 n1Cn1 xJn1.x/: (14.12) Bessel’s Differential Equation Suppose we consider a set of functions Z.x/that satisfies the basic recurrence relations, Eqs. (14.7) and (14.8), but with not necessarily an integer and Znot necessarily given by the series in Eq. (14.6). It is our objective to show that any functions that satisfy these recur- rence relations must also be solutions to Bessel’s ODE. We start by forming (1) x2Z00.x/ from x2=2times the derivative of Eq. (14.8), (2) x Z0 .x/from Eq. (14.8) multiplied by x=2, and (3)2Z.x/from Eq. (14.7) multiplied by x=2. Putting these together we obtain x2Z00.x/Cx Z0 .x/2Z.x/ Dx2 2 Z0 1.x/Z0 C1.x/1 xZ1.x/C1 xZC1.x/ : (14.13) ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 647 The terms within square brackets in Eq. (14.13) can now by use of Eq. (14.12) be simplified to2Z.x/, soEq. (14.13) can be rewritten x2Z00.x/Cx Z0 .x/C.x22/Z.x/D0; (14.14) which is Bessel’s ODE. Reiterating, we have shown that any functions Z.x/that satisfy the basic recurrence formulas, Eqs. (14.7) and(14.8), also satisfy Bessel’s equation; that is, the Zare Bessel functions. For later use, we note that if the argument of Ziskrather than x,Eq. (14.14) becomes 2d2 d2Z.k/Cd dZ.k/C.k222/Z.k/D0: (14.15) Integral Representation It is of great value to have integral representations of Bessel functions. Starting from the generating-function formula, we can apply the residue theorem to evaluate the contour integral I Ce.x=2/.tC1=t/ tnC1dtDI CX mJm.x/tmn1dtD2i Jn.x/; (14.16) where the contour Cencircles the singularity at tD0. The integral on the left-hand side of Eq. (14.16) can now be brought to a convenient form by taking the contour to be the unit circle and changing the integration variable by making the substitution tDei. Then dtDieid,e.x=2/.t1=t/Deixsin, and we have 2i Jn.x/D2Z 0eixsin e.nC1/iieidD2Z 0ei.xsinn/id: (14.17) Assuming xto be real and taking the imaginary parts of both sides of Eq. (14.17), we find Jn.x/D1 22Z 0cos.xsinn/dD1 Z 0cos.xsinn/d; (14.18) where the last equality only holds because we are assuming nto be an integer. Though we will not need it now, the real part of this equation also gives an interesting formula: 2Z 0sin.xsinn/dD0: (14.19) An oft-occurring special case of Eq. (14.18) is J0.x/D1 22Z 0eixcosdD1 Z 0cos.xsin/d: (14.20) ArfKen_17-ch14-0643-0714- 9780123846549.tex 648 Chapter 14 Bessel Functions Table 14.1 Zeros of the Bessel Functions and Their First Derivatives Number of zeros J0.x/ J1.x/ J2.x/ J3.x/ J4.x/ J5.x/ 1 2:4048 3:8317 5:1356 6:3802 7:5883 8:7715 2 5:5201 7:0156 8:4172 9:7610 11:0647 12:3386 3 8:6537 10:1735 11:6198 13:0152 14:3725 15:7002 4 11:7915 13:3237 14:7960 16:2235 17:6160 18:9801 5 14:9309 16:4706 17:9598 19:4094 20:8269 22:2178 J0 0.x/ J0 1.x/ J0 2.x/ J0 3.x/ J0 4.x/ J0 5.x/ 1 3:8317 1:8412 3:0542 4:2012 5:3176 6:4156 2 7:0156 5:3314 6:7061 8:0152 9:2824 10:5199 3 10:1735 8:5363 9:9695 11:3459 12:6819 13:9872 4 13:3237 11:7060 13:1704 14:5858 15:9641 17:3128 5 16:4706 14:8636 16:3475 17:7887 19:1960 20:5755 Equation (14.18) is only one of many integral representations of Jn, and some of these can be derived (using an appropriately modified contour) for Jof a nonintegral order. This topic is explored in the subsection below entitled “Bessel Functions of Nonintegral Order”. Zeros of Bessel Functions In many physical problems in which phenomena are described by Bessel functions, we are interested in the points where these functions (which have oscillatory character) are zero. For example, in a problem involving standing waves, these zeros identify the positions of thenodes. And in boundary value problems, we may need to choose the argument of our Bessel function to put a zero at an appropriate point. There are no closed formulas for the zeros of Bessel functions; they must be found by numerical methods. Because the need for them arises frequently, tables of the zeros are available, both in compilations such as AMS-55 (see Additional Readings) and at a variety of sources online.2Table 14.1 lists the first few zeros of Jn.x/for integer nfrom nD0 through nD5, giving also the positions of the zeros of J0 n. Example 14.1.1 FRAUNHOFER DIFFRACTION, CIRCULAR APERTURE In the theory of diffraction of radiation of wavelength , incident normal to a circular aperture of radius a, we encounter the integral 8aZ 0r dr2Z 0eibrcosd; (14.21) 2Additional roots of the Bessel functions and those of their first derivatives may be found in C. L. Beattie, Table of first 700 zeros of Bessel functions, Bell Syst. Tech. J. 37, 689 (1958), and Bell Monogr. 3055. Roots may be also be accessed in Mathematica, Maple, and other symbolic software. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 649 xθ αy rIncident waves FIGURE 14.2 Geometry for Fraunhofer diffraction, circular aperture. where8is the amplitude of the diffracted wave and .r;/identifies points in the aperture. The exponent brcosis the phase of the radiation through .r;/that is diffracted to an angle from the incident direction, with bD2 sin : (14.22) The geometry is illustrated in Fig. 14.2. Fraunhofer diffraction, for which the above are the relevant formulas, applies in the limit that the outgoing radiation is detected at large distances from the aperture. The behavior of the complex exponential will cause the amplitude to oscillate as is increased, creating (for each wavelength) a diffraction pattern. To understand the patterns more fully, we need to evaluate the integral in Eq. (14.21). From Eq. (14.20) we may immediately reduce Eq. (14.21) to 82aZ 0J0.br/rdr; (14.23) which can be integrated in rusing Eq. (14.10): 82aZ 01 b2d dr .br/J1.br/ drD2 b2 br J1.br/a 0D2a bJ1.ab/; (14.24) where we have used the fact that J1.0/D0. The intensity of the light in the diffraction ArfKen_17-ch14-0643-0714- 9780123846549.tex 650 Chapter 14 Bessel Functions 0100002000030000 0.0002 Radians0.0004 FIGURE 14.3 Amplitude of Fraunhofer diffraction vs. deflection angle (green light, aperture of radius 0.5 cm). pattern is proportional to 82and, substituting for bfrom Eq. (14.22), 82J1T.2a=/sin U sin 2 : (14.25) For visible light and apertures of reasonable size, 2a=is quite small: for green light (D5:5105cm) and an aperture with aD0:5cm,2a=D57120 , and these parame- ter values lead to the pattern for 8shown in Fig. 14.3. Note that the figure plots 8(a plot of82would make the oscillations too small to be observable on the same graph as the maximum at D0). We see that 8exhibits a central maximum at D0of amplitude 30,000, with subsidiary extrema that by D0:001 radian have decreased in magnitude to less than 1% of the central maximum. Remembering that the intensity is 82, we see that the diffraction spreading of the incident light is exceedingly small. To make a quantitative analysis of the diffraction pattern, we need to identify the positions of its minima. They correspond to the zeros of J1; for example, from Table 14.1 we find the first minimum to be where.2a=/sin D3:8317 , or 14seconds of arc. If this analysis had been known in the 17th century, the arguments against the wave theory of light would have collapsed. In mid-20th century this same diffraction pattern appears in the scattering of nuclear particles by atomic nuclei, a striking demonstration of the wave properties of the nuclear particles.  Further examples of the use of Bessel functions and their roots are provided by the following example and by the exercises of this section and Section 14.2. Example 14.1.2 CYLINDRICAL RESONANT CAVITY The propagation of electromagnetic waves in hollow metallic cylinders is important in many practical devices. If the cylinder has end surfaces, it is called a cavity. Resonant cavities play a crucial role in many particle accelerators. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 651 The resonant frequencies of a cavity are those of the oscillatory solutions to Maxwell’s equations that correspond to standing wave patterns. By combining Maxwell’s equations, we derived in Example 3.6.2 the vector Laplace equation for the electric field Ein a region free of electric charges and currents. Taking the z-axis along the axis of the cavity, our concern here is the equation for Ez, which from Eq. (3.71) we found to have the form r2EzD1 c2@2Ez @t2; (14.26) which has standing-wave solutions Ez.x;y;z;t/DEz.x;y;z/f.t/, where f.t/has real solutions sin!tandcos!t, corresponding to sinusoidal oscillations at angular frequency !. We are implicitly assuming that our solution has a nonzero component Ez, and we will also setBzD0, so we intend to obtain solutions that are usually called the TM (for ‘‘transverse magnetic”) modes of oscillation. Additional solutions, with EzD0andBznonzero, corre- spond to TE (transverse electric) modes and are the subject of Exercise 14.1.25. Thus, for the present problem, in which our cavity is that shown in Fig. 14.4, we seek solutions to the spatial PDE: r2EzCk2EzD0; kD! c: (14.27) The aim of the present example is to find the values of !for which Eq. (14.27) has solutions consistent with the boundary conditions at the cavity walls. Assuming the metallic walls to be perfect conductors, the boundary conditions are that the tangential components of the electric field vanish there. Taking the cavity to have planar end caps at zD0andzDh, and (in cylindrical coordinates ;') to be bounded by a curved surface at Da, our boundary conditions are ExDEyD0on the end caps, and E'DEzD0on the boundary at Da. xyz ha FIGURE 14.4 Resonant cavity. ArfKen_17-ch14-0643-0714- 9780123846549.tex 652 Chapter 14 Bessel Functions Once a solution (with BzD0) has been found for Ez, then the remaining components of BandEhave definite values. For further details, see J. D. Jackson, Electrodynamics in Additional Readings. Equation (14.27) can be solved by the method of separation of variables, with solutions of the form given in Eq. (9.64): Ez.;; z/DPlm./8 m.'/Zl.z/; (14.28) with8m./Deim'or its equivalent in terms of sines and cosines, while Zl.z/and Plm./are solutions of the ODEs d2Zl dz2Dl2Zl; (14.29) d d d Plm d C .k2l2/2m2 PlmD0: (14.30) Equation (14.29) corresponds to Eq. (9.58), but with a different choice of the sign for the separation constant in anticipation of the fact that Zlwill turn out to be oscillatory. This change causes n2in Eq. (9.60) to become k2l2, and Eq. (14.30) is then seen to correspond exactly with Eq. (9.63). Recognizing now Eq. (14.30) as Bessel’s ODE and Eq. (14.29) as the ODE for a classical harmonic oscillator, we find, before imposing boundary conditions, EzDJm.n/eim' AsinlzCBcoslz ; (14.31) and the general solution will be an arbitrary linear combination of the above for different values of n,m, and l. We have chosen the solution to Bessel’s ODE to be of the first kind to maintain regularity at D0, since thisvalue is inside the cavity. We have written the 'dependence of the solution as a complex exponential for notational convenience. The physically relevant solutions will be arbitrary mixtures of the corresponding real quanti- ties,sinm'andcosm'. Continuity and single-valuedness in 'dictate that mhave integer values. The condition that EzD0on the curved boundary translates into the requirement Jm.na/D0. Letting mjstand for the jth positive zero of Jm, we find that naD mj;ork2l2D mj a2 : (14.32) To complete the solution we need to identify the boundary condition on Z. Because @Ex=@xD@Ey=@yD0on the end caps, we have from the Maxwell equation for rE: @Ex @xC@Ey @yC@Ez @zD0!@Ez @zD0; (14.33) so we have the requirement Z0.0/DZ0.h/D0, and we must choose ZDBcoslz;with lDp h;pD0;1;2;:::: (14.34) Combining Eqs. (14.32) and (14.34), we find k2D mj a2 Cp h2 D!2 c2; (14.35) ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 653 thereby providing an equation for the resonant frequencies: !mjpDcs 2 mj a2Cp22 h2;8 < :mD0;1;2;:::; jD1;2;3;:::; pD0;1;2;::::(14.36) Recapitulating, the functions we have found, labeled by the indices m,j, and p, are the spatial parts of standing-wave solutions of TM character whose time dependence and overall amplitude are of the form Cei!mjpt.  Bessel Functions of Nonintegral Order While Jof noninteger are not produced from a generating-function approach, they are readily identified from the Taylor series expansion, and they are conventionally given a scale consistent with that of the Jnof integer n. They then satisfy the same recurrence relations as those derived from the generating function. Ifis not an integer, there is actually an important simplification. The functions J andJare then independent solutions of the same ODE, and a relation of the form of Eq. (14.5) does not exist. On the other hand, for Dn, an integer, we need another solution. The development of this second solution and an investigation of its properties form the subject of Section 14.3. Schlaefli Integral It is useful to modify the integral representation, Eq. (14.16), so that it can be applied for Bessel functions of nonintegral order. Our first step in doing so is to deform the circular contour by stretching it to infinity on the negative real axis and opening the contour there, as shown in Fig. 14.5. Our integral, written F.x/D1 2iZ Ce.x=2/.t1=t/ tC1dt; (14.37) −∞ (t)(t) FIGURE 14.5 Contour, Schlaefli integral for J. ArfKen_17-ch14-0643-0714- 9780123846549.tex 654 Chapter 14 Bessel Functions now has a branch point at tD0, and because we have opened the contour we can place the branch cut along the negative real axis. We might anticipate that this procedure will not affect our integral representation, as the integrand vanishes at tD1 on both sides of the cut. However, that remains to be proved. Our first step toward a proof that Fis actually Jis to verify that Fstill satisfies Bessel’s ODE. If we substitute Fand its xderivatives into the ODE, we can, after some manipulation, reach the expression 1 2iZ Cd dt( e.x=2/.t1=t/ t Cx 2 tC1 t) dt; (14.38) and because the integration is within a region of analyticity of the integrand, the integral reduces to( e.x=2/.t1=t/ t Cx 2 tC1 t) end( e.x=2/.t1=t/ t Cx 2 tC1 t) start: We therefore conclude that the ODE is satisfied if the above expression vanishes; in our present situation each of the quantities in braces is zero for large negative tand positive x, confirming that Fsatisfies Bessel’s ODE. We still need to show that Fis the solution designated J; to accomplish this we con- sider its value for small x>0. Deforming the contour to a large open circle and making a change of variable to uDeixt=2, we get (to lowest order in x) F.x/1 2ix 2 eiZ C0eu uC1du: (14.39) Because of the change of variable, the contour C0becomes that which we introduced when developing a Schlaefli integral representation of the gamma function, and, using Eq. (13.31), we reduce Eq. (14.39) to F.x/x 2sinT.C1/U0./ D1 0.C1/x 2 ; (14.40) where the last step used the reflection formula for the gamma function, Eq. (13.23). Since this is the leading term of the expansion for J, our proof is complete. Exercises 14.1.1 From the product of the generating functions g.x;t/g.x;t/, show that 1DTJ0.x/U2C2TJ1.x/U2C2TJ2.x/U2C and therefore thatjJ0.x/j1andjJn.x/j1=p 2;nD1;2;3;::: . Hint. Use uniqueness of power series, (Section 1.2). 14.1.2 Using a generating function g.x;t/Dg.uCv;t/Dg.u;t/g.v;t/, show that (a) Jn.uCv/DP1 sD1 Js.u/Jns.v/, (b) J0.uCv/DJ0.u/J0.v/C2P1 sD1Js.u/Js.v/. These are addition theorems for the Bessel functions. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 655 14.1.3 Using only the generating function e.x=2/.t1=t/D1X nD1Jn.x/tn and not the explicit series form of Jn.x/, show that Jn.x/has odd or even parity accord- ing to whether nis odd or even, that is, Jn.x/D.1/nJn.x/: 14.1.4 Use the basic recurrence formulas, Eqs. (14.7) and (14.8), to prove the following formulas: (a)d dxTxnJn.x/UDxnJn1.x/; (b)d dxTxnJn.x/UD xnJnC1.x/; (c) Jn.x/DJ0 nC1CnC1 xJnC1.x/: 14.1.5 Derive the Jacobi-Anger expansion eicos'D1X mD1imJm./eim': This is an expansion of a plane wave in a series of cylindrical waves. 14.1.6 Show that (a) cosxDJ0.x/C2P1 nD1.1/nJ2n.x/; (b) sinxD2P1 nD0.1/nJ2nC1.x/: 14.1.7 To help remove the generating function from the realm of magic, show that it can be derived from the recurrence relation, Eq. (14.7). Hint. (a) Assume a generating function of the form g.x;t/D1X mD1Jm.x/tm: (b) Multiply Eq. (14.7) bytnand sum over n. (c) Rewrite the preceding result as  tC1 t g.x;t/D2t [email protected];t/ @t: (d) Integrate and adjust the “constant” of integration (a function of x) so that the coefficient of the zeroth power, t0;isJ0.x/as given by Eq. (14.6). ArfKen_17-ch14-0643-0714- 9780123846549.tex 656 Chapter 14 Bessel Functions 14.1.8 Show, by direct differentiation, that J.x/D1X sD0.1/s sW0.sCC1/x 2C2s satisfies the two recurrence relations J1.x/CJC1.x/D2 xJ.x/; J1.x/JC1.x/D2J0 .x/; and Bessel’s differential equation x2J00 .x/Cx J0 .x/C.x22/J.x/D0: 14.1.9 Prove that sinx xD=2Z 0J0.xcos/cosd;1cosx xD=2Z 0J1.xcos/d: Hint. The definite integral =2Z 0cos2sC1dD246.2s/ 135.2sC1/ may be useful. 14.1.10 Derive Jn.x/D.1/nxn1 xd dxn J0.x/: Hint. Try mathematical induction (Section 1.4). 14.1.11 Show that between any two consecutive zeros of Jn.x/there is one and only one zero ofJnC1.x/. Hint. Equations (14.10) and(14.11) may be useful. 14.1.12 An analysis of antenna radiation patterns for a system with a circular aperture involves the equation g.u/D1Z 0f.r/J0.ur/rdr: Iff.r/D1r2, show that g.u/D2 u2J2.u/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 657 14.1.13 The differential cross section in a nuclear scattering experiment is given by d=dD jf./j2. An approximate treatment leads to f./Dik 22Z 0RZ 0expTiksinsin'Udd': Hereis an angle through which the scattered particle is scattered. Ris the nuclear radius. Show that d dD.R2/1 J1.k Rsin/ sin2 : 14.1.14 A set of functions Cn.x/satisfies the recurrence relations Cn1.x/CnC1.x/D2n xCn.x/; Cn1.x/CCnC1.x/D2C0 n.x/: (a) What linear second-order ODE does the Cn.x/satisfy? (b) By a change of variable transform your ODE into Bessel’s equation. This sug- gests that Cn.x/may be expressed in terms of Bessel functions of transformed argument. 14.1.15 (a) Show by direct differentiation and substitution that J.x/D1 2iZ Ce.x=2/.t1=t/t1dt (this is the Schlaefli integral representation of J), and that the equivalent equation, J.x/D1 2ix 2Z Cesx2=4ss1ds; both satisfy Bessel’s equation. Cis the contour shown in Fig. 14.5. The negative real axis is the cut line. Hint. This exercise is aimed at providing details of the discussion that starts at Eq. (14.38). (b) Show that the first integral (with nan integer) may be transformed into Jn.x/D1 22Z 0ei.xsinn/dDin 22Z 0ei.xcosCn/d: 14.1.16 The contour Cin Exercise 14.1.15 is deformed to the path 1 to1, unit circle ei toei, and finally1to1. Show that J.x/D1 Z 0cos.xsin/dsin 1Z 0exsinhd: This is Bessel’s integral. ArfKen_17-ch14-0643-0714- 9780123846549.tex 658 Chapter 14 Bessel Functions Hint. The negative values of the variable of integration umust be represented in a manner consistent with the presence of the branch cut, for example, by writing uDteix. 14.1.17 (a) Show that J.x/D2 1=20.C1 2/x 2=2Z 0cos.xsin/cos2d; where>1 2. Hint. Here is a chance to use series expansion and term-by-term integration. The formulas of Section 13.3 will prove useful. (b) Transform the integral in part (a) into J.x/D1 1=20.C1 2/x 2Z 0cos.xcos/sin2d D1 1=20.C1 2/x 2Z 0eixcossin2d D1 1=20.C1 2/x 21Z 1eipx.1p2/1=2dp: These are alternate integral representations of J.x/. 14.1.18 Given that Cis the contour in Fig. 14.5, (a) From J.x/D1 2ix 2Z Ct1etx2=4tdt derive the recurrence relation J0 .x/D xJ.x/JC1.x/: (b) From J.x/D1 2iZ Ct1e.x=2/.t1=t/dt derive the recurrence relation J0 .x/D1 2 J1.x/JC1.x/ : ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.1 Bessel Functions of the First Kind, J.x/ 659 14.1.19 Show that the recurrence relation J0 n.x/D1 2 Jn1.x/JnC1.x/ follows directly from differentiation of Jn.x/D1 Z 0cos.nxsin/d: 14.1.20 Evaluate 1Z 0eaxJ0.bx/dx;a;b>0: Actually the results hold for a0,1<b<1. This is a Laplace transform of J0. Hint. Either an integral representation of J0or a series expansion will be helpful. 14.1.21 Using the symmetries of the trigonometric functions, confirm that for integer n, 1 22Z 0cos.xsinn/dD1 Z 0cos.xsinn/d: 14.1.22 (a) Plot the intensity, 82of Eq. (14.25), as a function of .sin =/ along a diameter of the circular diffraction pattern. Locate the first two minima. (b) Estimate the fraction of the total light intensity that falls within the central maximum. Hint.TJ1.x/U2=xmay be written as a derivative and the area integral of the intensity integrated by inspection. 14.1.23 The fraction of light incident on a circular aperture (normal incidence) that is transmitted is given by TD22kaZ 0J2.x/dx x1 2ka2kaZ 0J2.x/dx: Here ais the radius of the aperture and kis the wave number, 2= . Show that (a) TD11 ka1X nD0J2nC1.2ka/, (b) TD11 2ka2kaZ 0J0.x/dx. 14.1.24 The amplitude U.;'; t/of a vibrating circular membrane of radius asatisfies the wave equation r2U@2U @2C1 @U @C1 2@2U @'2D1 v2@2U @t2: Herevis the phase velocity of the wave, determined by the properties of the membrane. ArfKen_17-ch14-0643-0714- 9780123846549.tex 660 Chapter 14 Bessel Functions (a) Show that a physically relevant solution is U.;'; t/DJm.k/ c1eim'Cc2eim' b1ei!tCb2ei!t : (b) From the Dirichlet boundary condition Jm.ka/D0, find the allowable values of k. 14.1.25 Example 14.1.2 describes the TM modes of electromagnetic cavity oscillation. To obtain the transverse electric (TE) modes, we set EzD0and work from the zcom- ponent of the magnetic induction B: r2BzC 2BzD0 with boundary conditions Bz.0/DBz.l/D0and@Bz @ DaD0: Show that the TE resonant frequencies are given by !mnpDcs 2mn a2Cp22 l2;pD1;2;3;:::; and identify the quantities mn. 14.1.26 A conducting cylinder can accommodate traveling electromagnetic waves; when used for this purpose it is called a wave guide. The equations describing traveling waves are the same as those of Example 14.1.2, but there is no boundary condition on EzatzD0 orzDhother than that its zdependence be oscillatory. For each TM mode (values ofmandjofExample 14.1.2), there is a minimum frequency that can be transmitted through a wave guide of radius a. Explain why this is so, and give a formula for the cutoff frequencies. 14.1.27 Plot the three lowest TM and the three lowest TE angular resonant frequencies, !mnp, as a function of the ratio radius/length .a=l/for0a=l1:5. Hint. Try plotting !2(in units of c2=a2) vs..a=l/2. Why this choice? 14.1.28 Show that the integral aZ 0xmJn.x/dx;mn0; (a) is integrable for mCnodd in terms of Bessel functions and powers of x, i.e., is expressible as linear combinations of apJq.a/; (b) may be reduced for mCneven to integrated terms plusRa 0J0.x/dx. 14.1.29 Show that 0nZ 0 1y 0n J0.y/ydyD1 0n 0nZ 0J0.y/dy: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.2 Orthogonality 661 Here 0nis the nth zero of J0.y/. This relation is useful (see Exercise 14.2.9): The expres- sion on the right is easier and quicker to evaluate, and is much more accurate. Taking the difference of two terms in the expression on the left leads to a large relative error. 14.2 O RTHOGONALITY To identify the orthogonality properties of Bessel functions, it is convenient to start by writing Bessel’s ODE in a form that we can recognize as a Sturm-Liouville eigenvalue problem, the general properties of which were discussed in detail starting from Eq. (8.15). If we divide Eq. (14.15) through by2and rearrange slightly, we have d2 d2C1 d d2 2 Z.k/Dk2Z.k/; (14.41) showing that Z.k/is an eigenfunction of the operator LDd2 d2C1 d d2 2 (14.42) with eigenvalue k2. Since we are most often interested in problems whose solutions in cylindrical coordinates .;'; z/separate into products P./8.'/ Z.z/and which are for the region within a cylindrical boundary at some Da, we usually have 8.'/Deim'with man integer (thereby causing 2!m2), and find that P./DJm.k/. We choose Pto be a Bessel function of the first kind because D0is interior to our region and we want a solution that is nonsingular there. From Sturm-Liouville theory, we find that the weight factor needed to make Lof Eq. (14.42) self-adjoint (as an ODE) is w./D, and the orthogonality integral for the two eigenfunctions J.k/andJ.k0/, a case of Eq. (8.20), is (whether or not is an integer) a k0J.ka/J0 .k0a/k J0 .ka/J.k0a/ k2k02DaZ 0J.k/J.k0/d: (14.43) In writing Eq. (14.43) we have used the fact that the presence of a factor in the boundary terms causes there to be no contribution from the lower limit D0.3 Equation (14.43) shows us that the J.k/of different kwill be orthogonal (with weight factor) if we can cause the left-hand side of that equation to vanish. We may do so by choosing kandk0in such a way that J.ka/DJ.k0a/D0. In other words, we can require thatkandk0be such that kaandk0aare zeros of J, and our Bessel functions will then satisfy Dirichlet boundary conditions. If now we let idenote the ith zero of J, the above analysis corresponds to the fol- lowing orthogonality formula for the interval T0;aU: aZ 0J i a J j a dD0; i6Dj: (14.44) 3This will be true for all 1 , as will become more evident when we discuss Bessel functions of the second kind. ArfKen_17-ch14-0643-0714- 9780123846549.tex 662 Chapter 14 Bessel Functions −0.200.20.40.6 x0.2 0.6 0.8 1 0.4 FIGURE 14.6 Bessel functions J1. 1n/,nD1;2;3on range 01. Note that all members of our orthogonal set of Bessel functions have the same value of the index, differing only in the scale of the argument of J. Successive members of the orthogonal set will have increasing numbers of oscillations in the interval .0;a/. Note also that the weight factor, , is just that which corresponds to unweighted orthogonality over the region within a circle of radius a. We show in Fig. 14.6 the first three Bessel functions of orderD1that are orthogonal within the unit circle. An alternative to the foregoing analysis would be to ensure the vanishing of the bound- ary term of Eq. (14.43) at Daby choosing values of kcorresponding to the Neumann boundary condition J0 .ka/D0. The functions obtained in this way would also form an orthogonal set. Normalization Our orthogonal sets of Bessel functions are not normalized, and to use them in expansions we need their normalization integrals. These integrals may be developed by returning to Eq. (14.43), which is valid for all kandk0, whether or not the boundary terms vanish. We take the limits of both sides of that equation as k0!k, evaluating the limit on the left-hand side using l’Hôpital’s rule, which here corresponds to taking the derivatives of numerator and denominator with respect to k0: aZ 0[J.k/]2dDlim k0!ka J.ka/d dk0 k0J0 .k0a/ k J0 .ka/d dk0 J.k0a/ d dk0.k2k02/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.2 Orthogonality 663 We now simplify this equation for the case that kaD i, so we set J.ka/D0and reach aZ 0h J i ai2 dDa2k J0 .ka/2 2kDa2 2 J0 . i/2: (14.45) Now, because iis a zero of J, Eq. (14.12) permits us to recognize that J0 . i/D JC1. i/. We then obtain from Eq. (14.45) the desired result, aZ 0h J i ai2 dDa2 2 JC1. i/2: (14.46) Bessel Series If we assume that the set of Bessel functions J. j=a/for fixedand for jD1;2;3;::: is complete, then any well-behaved but otherwise arbitrary function f./may be expanded in a Bessel series f./D1X jD1cjJ j a ;0a; >1: (14.47) The coefficients cjare determined by the usual rules for orthogonal expansions. With the aid of Eq. (14.46) we have cjD2 a2TJC1. j/U2aZ 0f./J j a d: (14.48) As pointed out earlier, it is also possible to obtain an orthogonal set of Bessel functions of given order by imposing the Neumann boundary condition J0 .k/D0atDa, corresponding to kD j=a, where jis the jth zero of J0 . These functions can also be used for orthogonal expansions. This approach is explored in Exercises 14.2.2 and 14.2.5. The following example illustrates the usefulness of Bessel series. Example 14.2.1 ELECTROSTATIC POTENTIAL IN A HOLLOW CYLINDER We consider a hollow cylinder, which in cylindrical coordinates .;'; z/is bounded by a curved surface at Daand end caps at zD0andzDh. The base ( zD0) and curved surface are assumed to be grounded, and therefore at potential D0, while the end cap atzDhhas a known potential distribution V.;'; h/. Our problem is to determine the potential V.;'; z/throughout the interior of the cylinder. We proceed by finding separated-variable solutions to the Laplace equation in cylindri- cal coordinates, along the lines discussed in Section 9.4. Our first step is to identify product ArfKen_17-ch14-0643-0714- 9780123846549.tex 664 Chapter 14 Bessel Functions solutions, which, as in Eq. (9.64), must take the form4 lm.;'; z/DPlm./8 m.'/Zl.z/; (14.49) with8mDeim', and d2 dz2Zl.z/Dl2Zl.z/; (14.50) 2d2 d2PlmCd dPlmC.l22m2/PlmD0: (14.51) The equation for Plmis Bessel’s ODE, with solutions of relevance here Jm.l/. To satisfy the boundary condition at Dawe need to choose lD mj=a, where jcan be any positive integer and mjis the jth zero of Jm. The equation for Zlhas solutions elz; to satisfy the boundary condition at zD0we need to take the linear combination of these solutions that is equivalent to sinhlz. Combin- ing these observations, we see that possible solutions to the Laplace equation that satisfy all the boundary conditions other than that at zDhcan be written mjDcmjJm mj a eim'sinh mjz a : (14.52) Since Laplace’s equation is homogeneous, any linear combination of the mjwith arbi- trary values of the cmjwill be a solution, and our remaining task is to find the linear combination of such solutions that satisfies the boundary condition at zDh. Therefore, V.;'; z/D1X mD11X jD1 mj; (14.53) with the boundary condition at zDhexpressed as 1X mD11X jD1cmjJm mj a eim'sinh mjh a DV.;'; h/: (14.54) Our solution is both a trigonometric series and a Bessel series, each with orthogonality properties that can be used to determine the coefficients. From Eq. (14.48) and the formula 2Z 0eim'eim0'D2 mm0; (14.55) we find cmjD a2sinh mjh a J2 mC1. mj/1 2Z 0d'aZ 0V.;'; h/Jm mj a eim'd: (14.56) 4Note that here Zlis a function of zarising from the separation of variables; the notation is not intended to identify it as a Bessel function. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.2 Orthogonality 665 These are definite integrals, that is, numbers. Substituting back into Eq. (14.52), the series inEq. (14.53) is specified and the potential V.;'; z/is determined.  Exercises 14.2.1 Show that .k2k02/aZ 0J.kx/J.k0x/xdxDaTk0J.ka/J0 .k0a/k J0 .ka/J.k0a/U; where J0 .ka/Dd d.kx/J.kx/jxDa, and that aZ 0TJ.kx/U2xdxDa2 2 TJ0 .ka/U2C 12 k2a2 TJ.ka/U2 ; >1: These two integrals are usually called the first and second Lommel integrals. 14.2.2 (a) If mis the mth zero of.d=d/J. m=a/, show that the Bessel functions are orthogonal over the interval T0;aUwith an orthogonality integral aZ 0J m a J n a dD0; m6Dn; >1: (b) Derive the corresponding normalization integral .mDn/. ANS: .b/a2 2 12 2m TJ. m/U2; >1: 14.2.3 Verify that the orthogonality equation, Eq. (14.44), and the normalization equation, Eq. (14.46), hold for >1. Hint. Using power-series expansions, examine the behavior of Eq. (14.43) as!0. 14.2.4 From Eq. (11.49), develop a proof that J.z/,>1has no complex roots (with a nonzero imaginary part). Hint. (a) Use the series form of J.z/to exclude pure imaginary roots. (b) Assume mto be complex and take nto be  m. 14.2.5 (a) In the series expansion f./D1X mD1cmJ m a ;0a; >1; with J. m/D0, show that the coefficients are given by cmD2 a2TJC1. m/U2aZ 0f./J m a d: ArfKen_17-ch14-0643-0714- 9780123846549.tex 666 Chapter 14 Bessel Functions (b) In the series expansion f./D1X mD1dmJ m a ;0a; >1; with.d=d/J. m=a/jDaD0, show that the coefficients are given by dmD2 a2.12= 2m/TJ. m/U2aZ 0f./J m a d: 14.2.6 A right circular cylinder has an electrostatic potential of .;'/ on both ends. The potential on the curved cylindrical surface is zero. Find the potential at all interior points. Hint. Choose your coordinate system and adjust your zdependence to exploit the symmetry of your potential. 14.2.7 A function f.x/is expressed as a Bessel series: f.x/D1X nD1anJm. mnx/; with mnthenth root of Jm. Prove the Parseval relation, 1Z 0Tf.x/U2x dxD1 21X nD1a2 nTJmC1. mn/U2: 14.2.8 Prove that 1X nD1. mn/2D1 4.mC1/: Hint. Expand xmin a Bessel series and apply the Parseval relation. 14.2.9 A right circular cylinder of length land radius ahas on its end caps a potential  zDl 2 D100 1 a : The potential on the curved surface (the side) is zero. Using the Bessel series from Exercise 14.2.6, calculate the electrostatic potential for =aD0:0.0:2/1:0 andz=lD 0:0.0:1/0:5 . Take a=lD0:5. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.3 Neumann Functions, Bessel Functions of the Second Kind 667 Hint. From Exercise 14.1.29 you have 0nZ 0 1y 0n J0.y/ydy: Show that this equals 1 0n 0nZ 0J0.y/dy: Numerical evaluation of this latter form rather than the former is both faster and more accurate. Note. For=aD0:0andz=lD0:5the convergence is slow, 20 terms giving only 98.4 rather than 100. Check value: For=aD0:4andz=lD0:3; D24:558: 14.3 N EUMANN FUNCTIONS , BESSEL FUNCTIONS OF THE SECOND KIND From the theory of ODEs, it is known that Bessel’s equation has two independent solutions. Indeed, for nonintegral order we have already found two solutions and labeled them J.x/andJ.x/using the infinite series, Eq. (14.6). The trouble is that when is integral, Eq. (14.5) holds and we have but one independent solution. A second solution may be developed by the methods of Section 7.6. This yields a perfectly good second solution of Bessel’s equation. However, that solution is not the standard form, which is called a Bessel function of the second kind or alternatively, a Neumann function. Definition and Series Form The standard definition of the Neumann functions is the following linear combination of J.x/andJ.x/: Y.x/DcosJ.x/J.x/ sin: (14.57) For nonintegral ;Y.x/clearly satisfies Bessel’s equation, for it is a linear combination of known solutions, J.x/andJ.x/. The behavior of Y.x/for small x(and nonintegral ) can be determined from the power-series expansion of J, Eq. (14.6); we may write, ArfKen_17-ch14-0643-0714- 9780123846549.tex 668 Chapter 14 Bessel Functions calling upon Eq. (13.23), Y.x/D1 sin1 0.1/x 2  D0./0.1/ 1 0.1/x 2  D0./ x 2 C: (14.58) However, for integral , Eq. (14.57) becomes indeterminate; in fact, Yn.x/for integral nis defined as Yn.x/Dlim!nY.x/: (14.59) To determine that the limit represented by Eq. (14.59) exists and is not identically zero (so that Yn.x/has a meaningful definition), we apply l’Hôpital’s rule to Eq. (14.57), obtaining initially Yn.x/D1 d J d.1/nd J d Dn: (14.60) Inserting the expansions of JandJfrom Eq. (14.6), the differentiations of .x=2/2s combine to yield .2=/ Jn.x/ln.x=2/, while the derivatives of 1=0.snC1/yield terms containing .snC1/=0.snC1/, where is the digamma function (Section 13.2). The final result, whose verification is the topic of Exercise 14.3.8, is Yn.x/D2 Jn.x/lnx 2 1 n1X kD0.nk1/W kWx 22kn 1 1X kD0.1/k kW.nCk/W .kC1/C .nCkC1/x 22kCn ; (14.61) An explicit form for .n/for integer nis given in Eq. (13.40). Equation (14.61) shows that for n>0, the most divergent term for small xis in agree- ment with the result for noninteger ngiven in Eq. (14.58). We also see that all solutions for integer ncontain a logarithmic term with the regular function Jnmultiplying the log- arithm. In our earlier study of ODEs, we found that a second solution will usually have a contribution of this type when the indicial equation causes the exponents of the power- series expansion to be integers. We may also conclude from Eq. (14.61) thatYnis linearly independent of Jn, confirming that we indeed have a second solution to Bessel’s ODE. It is of some interest to obtain the expansion of Y0.x/in a more explicit form. Returning toEq. (14.61), we note that its first summation is vacant, and we have the relatively simple expansion Y0.x/D2 J0.x/lnx 2 2 1X kD0.1/k kWkW[ CHk]x 22k D2 J0.x/h Clnx 2i 2 1X kD1.1/k kWkWHkx 22k ; (14.62) where Hkis the harmonic numberPk mD1m1and is the Euler-Mascheroni constant. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.3 Neumann Functions, Bessel Functions of the Second Kind 669 −1.00.4 24 6 8Y0(x) Y1(x) Y2(x) x10 FIGURE 14.7 Neumann functions Y0.x/,Y1.x/, and Y2.x/. The Neumann functions Yn.x/are irregular at xD0, but with increasing xbecome oscillatory, as may be seen from the graphs of Y0,Y1, and Y2in Fig. 14.7. The definition of Eq. (14.57) was specifically chosen to cause the oscillatory behavior to be at the same scale as that of Jnand displaced asymptotically in phase by =2, similarly to the relative behavior of the sine and cosine. However, unlike the sine and cosine, JnandYnonly exhibit exact periodicity in the asymptotic limit. This point is covered in detail in Section 14.6. Figure 14.8 compares J0.x/andY0.x/over a large range of x. Integral Representations As with all the other Bessel functions, Y.x/has integral representations. For Y0.x/we have Y0.x/D2 1Z 0cos.xcosh t/dtD2 1Z 1cos.xt/ .t21/1=2dt;x>0: (14.63) See Exercise 14.3.7, which shows that the above integral is a solution to Bessel’s ODE that is linearly independent of J0.x/. Specific identification as Y0is the topic of Exercise 14.4.8. Recurrence Relations Substituting Eq. (14.57) forY.x/(nonintegral) into the recurrence relations for Jn.x/, Eqs. (14.7) and(14.8), we see immediately that Y.x/satisfies these same recurrence rela- tions. This actually constitutes a proof that Yis a solution to the Bessel ODE. Note that the converse is not necessarily true. All solutions need not satisfy the same recurrence relations, as the relations depend on the scales assigned to the solutions of different . An example of this sort of trouble appears in Section 14.5. ArfKen_17-ch14-0643-0714- 9780123846549.tex 670 Chapter 14 Bessel Functions −0.4−0.200.20.40.6 51 0 2 0 2 5 3 0 15 x FIGURE 14.8 Oscillatory behavior of J0.x/(solid line) and Y0.x/(dashed line) for 1x30. Wronskian Formulas An ODE p.x/y00Cq.x/y0Cr.x/yD0in self-adjoint form (so qDp0) was found in Exercise 7.6.1 to have the following Wronskian formula connecting its solutions uandv: u.x/v0.x/u0.x/v.x/DA p.x/: (14.64) To bring Bessel’s equation to self-adjoint form, we need to write it as xy00Cy0C .x2=x/yD0, thereby showing that for our present purposes p.x/Dx, and we therefore have for each noninteger  JJ0 J0 JDA x: (14.65) Since Ais a constant but can be expected to depend on , it may be identified for each at any convenient point, such as xD0. From the power-series expansion, Eq. (14.6), we obtain the following limiting behaviors for small x: J!1 0.1C/x 2 ; J0 ! 20.1C/x 21 ; (14.66) J!1 0.1/x 2 ;J0 ! 20.1/x 21 : Substitution into Eq. (14.65) yields J.x/J0 .x/J0 .x/J.x/D2 x0.1C/0.1/D2 sin x; (14.67) ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.3 Neumann Functions, Bessel Functions of the Second Kind 671 using Eq. (13.23). Although Eq. (14.67) was obtained for x!0, comparison with Eq. (14.65) shows that it must be true for all x, and that AD.2=/ sin. Note that A vanishes for integral , showing that the Wronskian of JnandJnvanishes and that these Bessel functions are linearly dependent. Using our recurrence relations, we may readily develop a large number of alternate forms, among which are JJC1CJJ1D2 sin x; (14.68) JJ1CJJC1D2 sin x; (14.69) JY0 J0 YD2 x; (14.70) JYC1JC1YD2 x: (14.71) Many more will be found in the Additional Readings. You will recall that in Chapter 7, Wronskians were of great value in two respects: (1) in establishing the linear independence or linear dependence of solutions of differential equations, and (2) in developing an integral form of a second solution. Here the specific forms of the Wronskians and Wronskian-derived combinations of Bessel functions are useful primarily in development of the general behavior of the various Bessel functions. Wronskians are also of great use in checking tables of Bessel functions. Uses of Neumann Functions The Neumann functions Y.x/are of importance for a number of reasons: 1. They are second, independent solutions of Bessel’s equation, thereby completing the general solution. 2. They are needed for physical problems in which they are not excluded by a require- ment of regularity at xD0. Specific examples include electromagnetic waves in coax- ial cables and quantum mechanical scattering theory. 3. They lead directly to the two Hankel functions, whose definition and use, particularly in studies of wave propagation, are discussed in Section 14.4. We close with one example in which Neumann functions play a vital role. Example 14.3.1 COAXIAL WAVE GUIDES We are interested in an electromagnetic wave confined between the concentric, conducting cylindrical surfaces DaandDb. The equations governing the wave propagation are the same as those discussed in Example 14.1.2, but the boundary conditions are now dif- ferent, and our interest is in solutions that are traveling waves (compare Exercise 14.1.26). ArfKen_17-ch14-0643-0714- 9780123846549.tex 672 Chapter 14 Bessel Functions For wave propagation problems, it is convenient to write the solution in terms of com- plex exponentials, with the actual physical quantities involved ultimately identified as their real (or imaginary) parts. Thus, in place of Eq. (14.31) (the solution for standing waves in a cylindrical cavity), we now have for Ezsolutions in which the dependence must involve both JmandYm(as the latter is not ruled out by a requirement for regularity at D0). Including the time dependence, we have for the TM (transverse magnetic) solutions the separated-variable forms EzD cmnJm. mn/CdmnYm. mn/ eim'ei.lz!t/; (14.72) with lnow permitted to have any real value (there is no boundary condition on z). The index nidentifies different possible values of mn. As in Eq. (14.30), the relation between mn,l, and!is !2 c2D 2 mnCl2: (14.73) The most general TM traveling-wave solution will be an arbitrary linear combination of all functions of the form given by Eq. (14.72) with mn,cmn, and dmnchosen so that Ezwill vanish at DaandDb. A main difference between this problem and that of Example 14.1.2 is that the condition on Ezis not given by the zeros of the Bessel functions Jm, but by zeros of linear combinations of JmandYm. Specifically, we require that cmnJm. mna/CdmnYm. mna/D0; (14.74) cmnJm. mnb/CdmnYm. mnb/D0: (14.75) These transcendental equations may be solved, for each relevant m, to yield an infinite set of solutions (indexed by n) for mnand the ratio dmn=cmn. An example of this process is inExercise 14.3.10. Returning now to the equation for !, we observe that the smallest value it can attain for the solution indexed by mandnisc mn, showing that TM waves can only propagate if the angular frequency !of the electromagnetic radiation is equal to or larger than this cutoff. In general, larger values of mncorrespond to higher degrees of transverse oscillation, and modes with greater transverse oscillation will therefore have higher cutoff frequencies. As for the circular wave guide (the subject of Exercise 14.1.26, there will also be TE modes of propagation, also with mode-dependent cutoffs. However, the coaxial guide can also support traveling waves in TEM (transverse electric and magnetic) modes. These modes, not possible for a circular waveguide, do not exhibit a cutoff, are the confined equivalent of plane waves, and correspond to the flow of current (in opposite directions) on the coaxial conductors.  Exercises 14.3.1 Prove that the Neumann functions Yn(with nan integer) satisfy the recurrence relations Yn1.x/CYnC1.x/D2n xYn.x/; Yn1.x/YnC1.x/D2Y0 n.x/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.3 Neumann Functions, Bessel Functions of the Second Kind 673 Hint. These relations may be proved by differentiating the recurrence relations for J or by using the limit form of Ybutnotdividing everything by zero. 14.3.2 Show that for integer n Yn.x/D.1/nYn.x/: 14.3.3 Show that Y0 0.x/DY1.x/: 14.3.4 IfXandZare any two solutions of Bessel’s equation, show that X.x/Z0 .x/X0 .x/Z.x/DA x; in which Amay depend on but is independent of x. This is a special case of Exer- cise 7.6.11. 14.3.5 Verify the Wronskian formulas J.x/JC1.x/CJ.x/J1.x/D2 sin x; J.x/Y0 .x/J0 .x/Y.x/D2 x: 14.3.6 As an alternative to letting xapproach zero in the evaluation of the Wronskian constant, we may invoke the uniqueness of power-series expansions. The coefficient of x1in the series expansion of u.x/v0 .x/u0 .x/v.x/is then A. Show by series expansion that the coefficients of x0andx1ofJ.x/J0 .x/J0 .x/J.x/are each zero. 14.3.7 (a) By differentiating and substituting into Bessel’s ODE for D0, show thatR1 0cos.xcosh t/dtis a solution. Hint. Rearrange the final integral toR1 0d dt xsin.xcosh t/sinht dt: (b) Show that Y0.x/D2 R1 0cos.xcosh t/dtis linearly independent of J0.x/. 14.3.8 Verify the expansion formula for Yn.x/given in Eq. (14.61). Hint. Start from Eq. (14.60) and perform the indicated differentiations on the power- series expansions of JandJ. The digamma functions arise from the differen- tiation of the gamma function. You will need the identity (not derived in this book) lim z!n .z/=0. z/D.1/n1nW, where nis a positive integer. 14.3.9 If Bessel’s ODE (with solution J) is differentiated with respect to , one obtains x2d2 dx2@J @ Cxd dx@J @ C.x22/@J @D2J: Use the above equation to show that Yn.x/is a solution to Bessel’s ODE. Hint. Equation (14.60) will be useful. 14.3.10 For the case mD0,aD1, and bD2;the coaxial wave-guide TM boundary conditions become f./D0, with f.x/DJ0.2x/ Y0.2x/J0.x/ Y0.x/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 674 Chapter 14 Bessel Functions 5 −5 xJ0 (2x) J0 (x) Y0 (x) Y0 (2x) 2468 1 0 FIGURE 14.9 The function f.x/of Exercise 14.3.10. This function is plotted in Fig. 14.9. (a) Calculate f.x/forxD0:0.0:1/10:0 and plot f.x/vs.xto find the approximate location of the roots. (b) Call a root-finding program to determine the first three roots to higher precision. ANS. (b) 3.1230, 6.2734, 9.4182. Note. The higher roots can be expected to appear at intervals whose length approaches . Why? AMS-55 (see Additional Readings) gives an approximate formula for the roots. The function g.x/DJ0.x/Y0.2x/J0.2x/Y0.x/is much better behaved than the f.x/ previously discussed. 14.4 H ANKEL FUNCTIONS Hankel functions are solutions of Bessel’s ODE with asymptotic properties that make them particularly useful in problems involving the propagation of spherical or cylindrical waves. Since the functions JandYform the complete solution of this ODE, the Hankel functions cannot be anything completely new; they must be linear combinations of the solutions we have already found. We introduce them here via straightforward algebraic definitions; later in this section we identify integral representations that some authors have used as a starting point. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.4 Hankel Functions 675 Definitions Starting from the Bessel functions of the first and second kinds, namely J.x/andY.x/, we define the two Hankel functions H.1/ .x/and H.2/ .x/(sometimes, but nowadays infrequently referred to as Bessel functions of the third kind) as follows: H.1/ .x/DJ.x/CiY.x/; (14.76) H.2/ .x/DJ.x/iY.x/: (14.77) This is exactly analogous to taking eiDcosisin: (14.78) For real arguments, H.1/ andH.2/ are complex conjugates. The extent of the analogy will be seen even better when their asymptotic forms are considered. Indeed, it is their asymptotic behavior that makes the Hankel functions useful. This behavior is discussed in Section 14.6, and in that section we provide an illustrative example in which the asymptotic properties play a key role. Series expansion of H.1/ .x/andH.2/ .x/may be obtained by combining Eqs. (14.6) and (14.62). Often only the first term is of interest; it is given by H.1/ 0.x/i2 lnxC1Ci2 . ln 2/C; (14.79) H.1/ .x/i0./ 2 x C; > 0; (14.80) H.2/ 0.x/i2 lnxC1i2 . ln 2/C; (14.81) H.2/ .x/i0./ 2 x C; > 0: (14.82) In these equations is the Euler-Mascheroni constant, defined in Eq. (1.13). Since the Hankel functions are linear combinations (with constant coefficients) of J andY, they satisfy the same recurrence relations, Eqs. (14.7) and(14.8). For both H.1/ .x/ andH.2/ .x/, H1.x/CHC1.x/D2 xH.x/; (14.83) H1.x/HC1.x/D2H0 .x/: (14.84) A variety of Wronskian formulas can be developed, including: H.2/ H.1/ C1H.1/ H.2/ C1D4 ix; (14.85) J1H.1/ JH.1/ 1D2 ix; (14.86) J1H.2/ JH.2/ 1D2 ix: (14.87) ArfKen_17-ch14-0643-0714- 9780123846549.tex 676 Chapter 14 Bessel Functions Contour Integral Representation of the Hankel Functions The integral representation (Schlaefli integral) for J.x/was introduced in Section 14.1, where we established that J.x/D1 2iZ Ce.x=2/.t1=t/dt tC1; (14.88) with Cthe contour shown in Fig. 14.5. Recall that when is nonintegral, the integrand has a branch point at tD0and the contour had to avoid a cut line that was drawn along the negative real axis. In developing the Schlaefli integral for general , we began by showing that Bessel’s ODE was satisfied for any open contour for which an expression of the form e.x=2/.t1=t/ t Cx 2 tC1 t (14.89) vanished at both endpoints of the contour. We now make further use of those observations by noting that the expression in Eq. (14.89) not only vanishes at tD1 on the real axis both below and above the cut, but that it also vanishes at tD0when that point is approached from positive t. We therefore consider the contour shown in Fig. 14.10, calling attention to the fact that the upper half of the contour (from tD0CtotD1ei), labeled C1, meets the conditions necessary to yield a solution to Bessel’s ODE, and that the remaining (lower) half of the contour, labeled C2, also yields a solution. What remains to be determined is the identi- fication of these solutions: We will show that they are the Hankel functions. For x>0, we assert that H.1/ .x/D1 iZ C1e.x=2/.t1=t/dt tC1; (14.90) H.2/ .x/D1 iZ C2e.x=2/.t1=t/dt tC1: (14.91) C1C1 C2C2∞eiπ ∞e−iπt=i t=−i (t)(t) FIGURE 14.10 Hankel function contours. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.4 Hankel Functions 677 These expressions are particularly convenient because they may be handled by the method of steepest descents (Section 12.7). H.1/ .x/has a saddle point at tDCi, whereas H.2/ .x/ has a saddle point at tDi. There remains the problem of relating Eqs. (14.90) and(14.91) to our earlier definition of the Hankel functions, Eqs. (14.76) and(14.77). Since the contours of Eqs. (14.90) and (14.91) combine to produce a contour yielding J,Eq. (14.88), we have, from the integral representations, J.x/D1 2h H.1/ .x/CH.2/ .x/i : (14.92) If we can show (also from the integral representations) that Y.x/D1 2ih H.1/ .x/H.2/ .x/i ; (14.93) we will be able to recover the original definitions of the H.i/ . We therefore rewrite Eq. (14.90) by replacing the integration variable tbyei=s, so the integrand of that equation becomes e.x=2/.s1=s/eis1. After the substitution the contour (in s) is found to be the same as C1, but traversed in the opposite direction (thereby compensating the initial minus sign in the transformed integrand). The result, with details left as Exercise 14.4.3, is that the contour integral representation of H.1/is consistent with the identification H.1/ .x/DeiH.1/ .x/: (14.94) Similar processing of Eq. (14.91), with tDei=s, leads to H.2/ .x/DeiH.2/ .x/: (14.95) We now combine Eqs. (14.94) and(14.95) to reach J.x/D1 2h eiH.1/ .x/CeiH.2/ .x/i ; (14.96) where again the H.i/ refer to the contour integral representations. Substituting Eqs. (14.92) and(14.96) into the defining equation for Y,Eq. (14.57), we confirm that Yis described properly when the H.i/ stand for their contour integral representations. This completes the proof that Eqs. (14.90) and(14.91) are consistent with the original definitions of the Hankel functions. The reader may wonder why so much stress is placed on the development of integral representations. There are several reasons. The first is simply aesthetic appeal. Second, the integral representations facilitate manipulations, analysis, and the development of rela- tions among the various special functions. We have already seen an example of this in the development of Eqs. (14.94) to (14.96). And, probably most important of all, integral rep- resentations are extremely useful in developing asymptotic expansions. Such expansions can often be obtained using the method of steepest descents (Section 12.7), or by methods involving expansion in negative powers of the expansion variable, as in Section 12.6. ArfKen_17-ch14-0643-0714- 9780123846549.tex 678 Chapter 14 Bessel Functions In conclusion, the Hankel functions are introduced here for the following reasons: As analogs of eixthey are useful for describing traveling waves. These applications are best studied when the asymptotic properties of the functions are in hand, and there- fore are postponed to Section 14.6. They offer an alternate (contour integral) and rather elegant definition of Bessel functions. We will see in Section 14.5 that they offer a route to the definition of the quantities known as modified Bessel functions, and that in Section 14.6 they are useful for the development of the asymptotic properties of Bessel functions. Exercises 14.4.1 Verify the Wronskian formulas (a) J.x/H.1/0 .x/J0 .x/H.1/ .x/D2i x; (b) J.x/H.2/0 .x/J0 .x/H.2/ .x/D2i x; (c) Y.x/H.1/0 .x/Y0 .x/H.1/ .x/D2 x; (d) Y.x/H.2/0 .x/Y0 .x/H.2/ .x/D2 x; (e) H.1/ .x/H.2/0 .x/H.1/0 .x/H.2/ .x/D4i x; (f) H.2/ .x/H.1/ C1.x/H.1/ .x/H.2/ C1.x/D4 ix; (g) J1.x/H.1/ .x/J.x/H.1/ 1.x/D2 ix: 14.4.2 Show that the integral forms (a)1 i1eiZ 0C1e.x=2/.t1=t/dt tC1DH.1/ .x/; (b)1 i0Z 1eiC2e.x=2/.t1=t/dt tC1DH.2/ .x/ satisfy Bessel’s ODE. The contours C1andC2are shown in Fig. 14.10. 14.4.3 Show that the substitution tDei=sinto Eq. (14.90) for H.1/ .x/not only produces the integrand for the similar integral representation of H.1/ .x/but that the contour in sis identical to the original contour in t. 14.4.4 Using the integrals and contours given in Exercise 14.4.2, show that 1 2iTH.1/ .x/H.2/ .x/UDY.x/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.4 Hankel Functions 679 iπC3 C4∞−i π∞+i π −iπ(γ)(γ) FIGURE 14.11 Hankel function contours for Exercise 14.4.5. 14.4.5 Show that the integrals in Exercise 14.4.2 may be transformed to yield (a) H.1/ .x/D1 iR C3exsinh  d ; (b) H.2/ .x/D1 iR C4exsinh  d ; where C3andC4are the contours in Fig. 14.11. 14.4.6 (a) Transform H.1/ 0.x/, Eq. (14.90), into H.1/ 0.x/D1 iZ Ceixcosh sds; where the contour Cruns from1 i=2 through the origin of the s-plane to 1C i=2. (b) Justify rewriting H.1/ 0.x/as H.1/ 0.x/D2 i1Ci=2Z 0eixcosh sds: (c) Verify that this integral representation actually satisfies Bessel’s differential equa- tion. (The i=2in the upper limit is not essential. It serves as a convergence factor. We can replace it by ia=2 and take the limit a!0.) 14.4.7 From H.1/ 0.x/D2 i1Z 0eixcosh sds show that (a)J0.x/D2 R1 0sin.xcosh s/ds;(b)J0.x/D2 R1 1sin.xt/p t21dt: This last result is a Fourier sine transform. ArfKen_17-ch14-0643-0714- 9780123846549.tex 680 Chapter 14 Bessel Functions 14.4.8 From H.1/ 0.x/D2 i1Z 0eixcosh sds(see Exercises 14.4.5 and 14.4.6), show that (a) Y0.x/D2 1Z 0cos.xcosh s/ds; (b) Y0.x/D2 1Z 1cos.xt/p t21/dt: These are the integral representations in Eq. (14.63). This last result is a Fourier cosine transform. 14.5 M ODIFIED BESSEL FUNCTIONS ,I.x/AND K.x/ The Laplace and Helmholtz equations, when separated in circular cylindrical coordinates, may lead to Bessel’s ODE in the coordinate that describes distance from the cylindrical axis. When that is the case, the behavior of the solutions as a function of is inherently oscillatory; as we have already seen, the Bessel functions J.k/, and also Y.k/, have for any value of an infinite number of zeros, and this property may be useful in causing satisfaction of boundary conditions. However, as already shown in Section 9.4, the con- nection constants arising when the variables are separated may have a sign opposite to that required to yield Bessel’s ODE, and the equation in the coordinate then assumes the form 2d2 d2P.k/Cd dP.k/.k22C2/P.k/D0: (14.97) Equation (14.97), known as the modified Bessel equation, differs from the Bessel ODE only in the sign of the quantity k22, but this small change is sufficient to alter the nature of the solutions. As we shall shortly discuss in more detail, the solutions to Eq. (14.97), called modified Bessel functions, are notoscillatory and have behavior that is exponential (rather than trigonometric) in character. Fortunately, the knowledge we have developed regarding the Bessel ODE can be put to good use for the modified Bessel equation, since the substitution k!ikconverts the conventional Bessel ODE to its modified form, and shows that if P.k/is a solution to the Bessel ODE, then P.ik/must be a solution to the modified Bessel equation. One way of stating this fact is to note that the solutions of Eq. (14.97) are Bessel functions of imaginary argument. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.5 Modi/f_ied Bessel Functions, I.x/andK.x/ 681 Series Solution Since any solution of Bessel’s ODE can be converted into a solution of the modified ODE by insertion of iinto its argument, let’s start by looking at the series expansion J.ix/D1X sD0.1/s sW0.sCC1/ix 2C2s Di1X sD01 sW0.sCC1/x 2C2s :(14.98) Since all the terms of the summation have the same sign, it is evident that J.ix/cannot exhibit oscillatory behavior. It is convenient to choose the solutions of the modified Bessel equation in a way that causes them to be real, and we accordingly defined the modified Bessel functions of the first kind, denoted I.x/, as I.x/DiJ.ix/Dei=2J.xei=2/D1X sD01 sW0.sCC1/x 2C2s : (14.99) Like Jfor0,Iis finite at the origin, with a power-series expansion that is convergent for all x. At small x, its limiting behavior will be of the form I.x/Dx 20.C1/C: (14.100) From the relation between JandJ, we may also conclude that IandIare linearly independent unless is an integer n; taking cognizance of the factor inin the definition ofIn, the linear dependence takes the form In.x/DIn.x/: (14.101) Graphs of I0andI1are shown in Fig. 14.12. Recurrence Relations for I The recurrence relations satisfied by I.x/may be developed from the series expansions, but it is perhaps easier to work from the existing recurrence relations for J.x/. Our starting point is Eq. (14.7), written for ix: J1.ix/CJC1.ix/D2n ixJn.ix/: (14.102) We change JtoI, related according to Eq. (14.99) by J.ix/DiI.x/; (14.103) thereby obtaining i1I1.x/CiC1IC1.x/D2 ixiI.x/; which simplifies to I1.x/IC1.x/D2 xI.x/: (14.104) ArfKen_17-ch14-0643-0714- 9780123846549.tex 682 Chapter 14 Bessel Functions 2.4 2.0 1.6 1.2 0.8 0.4 123xK0K1 I1 I0 FIGURE 14.12 Modified Bessel functions. In a similar fashion, Eq. (14.8) transforms into I1.x/CIC1.x/D2I0 .x/: (14.105) The above analysis is also the topic of Exercise 14.1.14. Second Solution K As already pointed out we have but one independent solution when is an integer, exactly as for the Bessel functions J. The choice of a second, independent solution of Eq. (14.97) is essentially a matter of convenience. The second solution given here is selected on the basis of its asymptotic behavior, which we examine in the next section. The confusion of choice and notation for this solution is perhaps greater than anywhere else in this field.5 There is also no universal nomenclature; the Kare sometimes referred to as Whittaker functions. Following AMS-55 (see Additional Readings for reference), we here define a second solution in terms of the Hankel function H.1/ .x/as K.x/ 2iC1H.1/ .ix/D 2iC1 J.ix/CiY.ix/ : (14.106) 5Discussion and comparison of notations will be found in Math. Tables Aids Comput. 1: 207–308 (1944) and in AMS-55 (see Additional Readings). ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.5 Modi/f_ied Bessel Functions, I.x/andK.x/ 683 The factor iC1makes K.x/real when xis real.6Using Eqs. (14.57) and(14.99), we may transform Eq. (14.106) to7 K.x/D 2I.x/I.x/ sin; (14.107) somewhat analogous to Eq. (14.57) for Y.x/. The choice of Eq. (14.106) as a definition is somewhat unfortunate in that the function K.x/does not satisfy the same recurrence relations as I.x/. The recurrence formulas for the Kare K1.x/KC1.x/D2 xK.x/; (14.108) K1.x/CKC1.x/D2 K0 .x/: (14.109) To avoid this discrepancy in the recurrence relations, some authors8have included an additional factor of cosin the definition of K. This would permit Kto satisfy the same recurrence relations as I(see Exercise 14.5.8), but it has the disadvantage of making KD0forD1 2;3 2;5 2;:::. The series expansion of K.x/follows directly from the series form of H.1/ .ix/, pro- viding that we choose the branch of lnixappropriately (see Exercise 14.5.9). Using Eqs. (14.79) and(14.80), the lowest-order terms are then found to be K0.x/Dlnx Cln 2C; (14.110) K.x/D210./ xC: (14.111) Because the modified Bessel function Iis related to the Bessel function J, much as sinh is related to sine, the modified Bessel functions IandKare sometimes referred to as hyperbolic Bessel functions. K0andK1are shown in Fig. 14.12. Integral Representations I0.x/andK0.x/have the integral representations I0.x/D1 Z 0cosh. xcos/d; (14.112) K0.x/D1Z 0cos.xsinht/dtD1Z 0cos.xt/dt .t2C1/1=2;x>0: (14.113) Equation (14.112) may be derived from Eq. (14.20) forJ0.x/or may be taken as a special case of Exercise 14.5.14. The integral representation of K0,Eq. (14.113), is derived in Section 14.6. A variety of other forms of integral representations (including 6D0) appear 6Ifis not an integer, K.z/has a branch point at zD0due to the presence of a fractional power; if Dn, an integer, Kn.z/ has a branch point at zD0due to the term lnz. We normally identify Kn.z/as the branch that is real for real z. 7For integral index nwe take the limit as !n. 8For example, Whittaker and Watson (see Additional Readings). ArfKen_17-ch14-0643-0714- 9780123846549.tex 684 Chapter 14 Bessel Functions in the exercises. These integral representations are useful in developing asymptotic forms (Section 14.6) and in connection with Fourier transforms (Chapter 19). Example 14.5.1 A GREEN’S FUNCTION We wish to develop an expansion for the fundamental Green’s function for the Laplace equation in cylindrical coordinates .;'; z/. The defining equation is " @2 @2 1C1 1@ @1C1 2 1@2 @'2 1C@2 @z2 1# G.r1;r2/D.12/1 2 1.'1'2/.z1z2/: (14.114) We now write the Dirac delta function for the 'coordinate in the form corresponding to Eq. (5.27): .'1'2/D1 21X mD1eim.'1'2/: For the zcoordinate, we use the continuum limit of the above formula, or, equivalently, the large- nlimit of Eq. (1.155), .z1z2/D1 21Z 1eik.z1z2/dkD1 1Z 0cosk.z1z2/dk: We use the last form of the above equation so that kwill never be negative. We now expand G.r1;r2/as G.r1;r2/D1 22X m1Z 0dkg m.k;1;2/eim.'1'2/cosk.z1z2/: (14.115) For'1and'2, this is simply an expansion in orthogonal functions; the dependence on z1,z2, and kis actually an integral transform that will be more completely justified in Chapter 20. For our present purposes, what is significant is that we can apply the orthog- onality properties of the expansion to find that Eq. (14.114) will be satisfied if (for all relevant values of kandm) " @2 @2 1C1 1@ @1m2 2 1k2# gm.k;1;2/D.12/: (14.116) We now have a one-dimensional (1-D) Green’s function problem for which the homoge- neous equation can be identified as the modified Bessel equation, with solutions Im.k/ andKm.k/. Keeping in mind that Imis regular at the origin, that Kmis regular at infinity, and that the Green’s function we seek must be regular at both these limits, we write our 1-Daxial Green’s function in the more explicit form gm.k1;k2/DIm.k</Km.k>/; (14.117) ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.5 Modi/f_ied Bessel Functions, I.x/andK.x/ 685 where<and>are, respectively, the smaller and larger of 1and2. The coefficient in the above equation,1, is evaluated according to Eq. (10.19), from  p.k/ K0 m.k/Im.k/I0 m.k/Km.k/1 : The coefficient pis from the differential equation, and has here the value k; the form involving modified Bessel functions is their Wronskian, and has the value 1=k; that is the topic of Exercise 14.5.11. Given our explicit formula for gm, Eq. (14.115) assumes the final form G.r1;r2/D1 22X m1Z 0dkg m.k1;k2/eim.'1'2/cosk.z1z2/: (14.118) This is the form quoted in Section 10.2.  Summary To put the modified Bessel functions I.x/andK.x/in proper perspective, note that we have introduced them here because: These functions are solutions of the frequently encountered modified Bessel equation, which arises in a variety of physically important problems, K.x/will be found useful in determining the asymptotic behavior of all the Bessel and modified Bessel functions (Section 14.6), and I.x/andK.x/arise in our discussion of Green’s functions (Example 14.5.1). Exercises 14.5.1 Show that e.x=2/.tC1=t/D1X nD1In.x/tn;thus generating modified Bessel functions, In.x/. 14.5.2 Verify the following identities (a) 1DI0.x/C2P1 nD1.1/nI2n.x/; (b) exDI0.x/C2P1 nD1In.x/; (c) exDI0.x/C2P1 nD1.1/nIn.x/; (d) cosh xDI0.x/C2P1 nD1I2n.x/; (e) sinhxD2P1 nD1I2n1.x/: 14.5.3 (a) From the generating function of Exercise 14.5.1 show that In.x/D1 2iI e.x=2/.tC1=t/dt tnC1: ArfKen_17-ch14-0643-0714- 9780123846549.tex 686 Chapter 14 Bessel Functions (b) For nD, not an integer, show that the preceding integral representation may be generalized to I.x/D1 2iZ Ce.x=2/.tC1=t/dt tC1: The contour Cis the same as that for J.x/(Fig. 14.5). 14.5.4 For>1 2show that I.z/may be represented by I.z/D1 1=20.C1 2/z 2Z 0ezcossin2d D1 1=20.C1 2/z 21Z 1ezp.1p2/1=2dp D2 1=20.C1 2/z 2=2Z 0cosh. zcos/sin2d: 14.5.5 The cylindrical cavity depicted in Fig. 14.4 has radius aand height h. For this exercise, the end caps zD0andhare at zero potential, while the cylindrical wall Dahas a potential of functional form VDV.';z/. (a) Show that the electrostatic potential 8.;'; z/has the functional form 8.;'; z/D1X mD01X nD1Im.kn/.amnsinm'Cbmncosm'/sinknz; where knDn=h. (b) Show that the coefficients amnandbmnare given by amn bmn D2m0 l Im.kna/2Z 0lZ 0V.';z/sinm' cosm' sinknzdzd': Hint. Expand V.';z/as a double series and use the orthogonality of the trigonometric functions. 14.5.6 Verify that K.x/as defined in Eq. (14.106) is equivalent to K.x/D 2I.x/I.x/ sin and from this show that K.x/DK.x/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.5 Modi/f_ied Bessel Functions, I.x/andK.x/ 687 14.5.7 Show that K.x/satisfies the following recurrence relations: K1.x/KC1.x/D2 xK.x/; K1.x/CKC1.x/D2 K0 .x/: Note. These differ from the recurrence relations for I. 14.5.8 IfKDeiK, show that Ksatisfies the same recurrence relations as I. 14.5.9 Show that when K0is evaluated from its series expansion about xD0, the formula given as Eq. (14.110) only follows if a specific branch of its logarithmic term is chosen. 14.5.10 For>1 2show that K.z/may be represented by K.z/D1=2 0.C1 2/z 21Z 0ezcosh tsinh2t dt; 2<argz< 2 D1=2 0.C1 2/z 21Z 1ezp.p21/1=2dp: 14.5.11 Show that I.x/andK.x/satisfy the Wronskian relation I.x/K0 .x/I0 .x/K.x/D1 x: 14.5.12 Verify that the coefficient in the axial Green’s function of Eq. (14.117) is 1. 14.5.13 IfrD.x2Cy2/1=2, prove that 1 rD2 1Z 0cos.xt/K0.yt/dt: This is a Fourier cosine transform of K0. 14.5.14 Derive the integral representation In.x/D1 Z 0excoscos.n/d: Hint. Start with the corresponding integral representation of Jn.x/. Equation (14.112) is a special case of this representation. 14.5.15 Show that K0.z/D1Z 0ezcosh tdt satisfies the modified Bessel equation. How can you establish that this form is linearly independent of I0.z/? ArfKen_17-ch14-0643-0714- 9780123846549.tex 688 Chapter 14 Bessel Functions 14.5.16 The cylindrical cavity of Exercise 14.5.5 has along the cylinder walls the potential walls: V.z/D8 >< >:100z h; 100 1z h ;0z h1=2; 1=2z h1: With the radius-height ratio a=hD0:5, calculate the potential for z=hD0:1.0:1/0:5 and=aD0:0.0:2/1:0: Check value. Forz=hD0:3and=aD0:8,VD26:396 . 14.6 A SYMPTOTIC EXPANSIONS Frequently in physical problems there is a need to know how a given Bessel or modified Bessel function behaves for large values of the argument, that is, its asymptotic behavior. This is one occasion when computers are not very helpful. One possible approach is to develop a power-series solution of the differential equation, but now using negative pow- ers. This is Stokes’ method, illustrated in Exercise 14.6.10. The limitation is that starting from some positive value of the argument (for convergence of the series), we do not know what mixture of solutions or multiple of a given solution we have. The problem is to relate the asymptotic series (useful for large values of the variable) to the power-series or related definition (useful for small values of the variable). This relationship can be established is various ways, one of which is to introduce a suitable integral representation whose asymptotic behavior can be studied by application of the method of steepest descents, Section 12.7. We start this process with a study of the Hankel functions, for which a contour integral representation was introduced in Section 14.4. Asymptotic Forms of Hankel Functions InSection 14.4 it was shown that the Hankel functions, which satisfy Bessel’s equation, may be defined by the contour integrals H.1/ .t/D1 iZ C1e.t=2/.z1=z/dz zC1; (14.119) H.2/ .t/D1 iZ C2e.t=2/.z1=z/dz zC1; (14.120) where C1andC2are the contours shown in Fig. 14.10. We desire formulas based on these representations for the asymptotic behavior of the Hankel functions at large positive t. The direct and exact evaluation of these integrals appears to be nearly impossible, but the situation does have features permitting us to use the method of steepest descents to make an asymptotic evaluation. Referring to the exposition of that method in Section 12.7, ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.6 Asymptotic Expansions 689 we have the approximate evaluation Z Cg.z;t/ew.z;t/dzg.z0;t/ew.z0;t/eis 2 jw00.z0;t/j; (14.121) where the contour Cpasses through a saddle point at zDz0and Darg.w00.z0;t// 2C 2or3 2 (14.122) is a phase arising from the direction of passage through the saddle point. We regard the common integrand of Eqs. (14.119) and (14.120) as possessing a slowly varying factor g.z/Dz1and an exponential ewwithwD.t=2/.zz1/, and seek saddle points by finding the zeros of w0Dt 2 1C1 z2 : (14.123) Solving the above equation, we identify the two saddle points z0DCiandz0Di. Limiting attention to H.1/ .t/, we see that we can deform the contour C1so that it passes through the saddle point at z0Di; there is neither the need nor the possibility to deform this contour to pass through z0Di. Thus, at the saddle point, we have w.Ci/Dit; w00.Ci/Dt z3 0 z0DiDit: (14.124) The argument of w00.z0/is=2 , so the possible values of the phase (the direction of descent from the saddle point) are 3=4 and7=4 . We must choose D3=4 since we cannot get into position to cross the saddle point in the direction D7=4D=4 without first crossing a region where the integrand is larger in absolute value than its value at the saddle point. We now have all the information needed to use Eq. (14.121) to estimate the integral. The result is H.1/ .t/1 ie.i=2/.1/e3i=4eitr 2 t r 2 tei.t=2=4/: (14.125) This is the leading term of the asymptotic expansion of the Hankel function H.1/ .t/for large t. The other Hankel function can be treated similarly, but using the saddle point at zDi, with result H.2/ .t/r 2 tei.t=2=4/: (14.126) Equations (14.125) and (14.126) permit us to obtain the leading terms in the asymp- totic behavior of all the Bessel and modified Bessel functions. In particular, inserting the ArfKen_17-ch14-0643-0714- 9780123846549.tex 690 Chapter 14 Bessel Functions asymptotic form for H.1/.ix/intoEq. (14.106), which defines K.x/, we find K.x/ 2iC1r 2 ixexet=2=4/; r 2xex: (14.127) Another solution to the modified Bessel equation can be obtained from H.2/.ix/; its asymptotic behavior will be proportional to eCx. Combining the present observations with Eqs. (14.100), (14.110), and (14.111), we can conclude that: 1. The modified Bessel function K.x/will be irregular at xD0as given by Eqs. (14.110) or(14.111), and will decay exponentially at large x; 2. The modified Bessel function I.x/will (for0) be finite at the origin, as given by Eq. (14.100), and will increase exponentially at large x. Rather than developing additional asymptotic forms from Eq. (14.127), we find it more interesting to obtain more complete asymptotic expansions by use of a particular integral representation of K. Expansion of an Integral Representation for K Here we start from the integral representation K.z/D1=2 0.C1 2/z 21Z 1ezx.x21/1=2dx; >1 2: (14.128) For the present let us take zto be real, although Eq. (14.128) may be established for =2<argz<=2 (i.e., for<e.z/>0). Before using Eq. (14.128) we need to verify that (1) the form claimed to be K.z/ satisfies the modified Bessel equation, (2) that it has the small- zbehavior required for K, and (3) that it has the required exponentially decaying asymptotic value. These three features suffice to establish the validity of Eq. (14.128). The fact that Eq. (14.128) is a solution of the modified Bessel equation may be verified by direct substitution into Eq. (14.97). After some manipulation, we obtain zC11Z 1d dx ezx.x21/C1=2 dxD0; which transforms the combined integrand into the derivative of a function that vanishes at both endpoints. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.6 Asymptotic Expansions 691 We next consider how Eq. (14.128) behaves for small z. We proceed by substituting xD1Ct=z: 1=2 0.C1 2/z 21Z 1ezx.x21/1=2dx D1=2 0.C1 2/z 2 ez1Z 0ett2 z2C2t z1=2dt z D1=2 0.C1 2/ez 2z1Z 0ett21 1C2z t1=2 dt: (14.129) This substitution has changed the limits of integration to a more convenient range and has isolated the negative exponential dependence ez. The integral in Eq. (14.129) may now (for>0) be evaluated for zD0to yield0.2/ . Then, using the duplication formula, Eq. (13.27), we have lim z!0K.z/D0./21 z; > 0: (14.130) Equation (14.130) agrees with Eq. (14.111), showing that Eq. (14.128) has the proper small- zbehavior to represent K. Note that for D0,Eq. (14.128) diverges logarithmi- cally at zD0and the verification of its scale requires a different approach, which is the topic of Exercise 14.6.4. Finally, to complete the identification of Eq. (14.128) with K, we need to verify that it decays exponentially at large z. That feature will be a by-product of our main interest here, which is to develop an asymptotic series for K.z/. We do so by rewriting Eq. (14.129) as K.z/Dr 2zez 0.C1 2/1Z 0ett1=2 1Ct 2z1=2 dt: (14.131) We next expand .1Ct=2z/1=2by the binomial theorem and interchange the summation and integration (valid for the asymptotic series we plan to obtain), reaching K.z/Dr 2zez 0.C1 2/1X rD01 2 r .2z/r1Z 0ettCr1=2dt Dr 2zez1X rD00.CrC1 2/ rW0.rC1 2/.2z/r: (14.132) Equation (14.132) can now be rearranged to K.z/r 2zez 1C.4212/ 1W8zC.4212/.4232/ 2W.8z/2C : (14.133) ArfKen_17-ch14-0643-0714- 9780123846549.tex 692 Chapter 14 Bessel Functions Equation (14.133) yields the anticipated exponential dependence, confirming that Eq. (14.128) actually represents K. Although the integral of Eq. (14.128), integrating along the real axis, was convergent only for=2<argz<=2 ,Eq. (14.133) may be extended to3=2<argz<3=2 . Consid- ered as an infinite series, Eq. (14.133) is actually divergent. However, this series is asymp- totic, in the sense that for large enough z;K.z/may be approximated to any fixed degree of accuracy with a small number of terms. Compare Section 12.6 for a definition and discus- sion of asymptotic series. The asymptotic character arises because our binomial expansion was valid only for t<2zbut we integrated tout to infinity. The exponential decrease of the integrand has prevented a disaster, but the series is only asymptotic and not convergent. By Table 7.1, zD1 is an essential singularity of the Bessel (and modified Bessel) equations. Fuchs’ theorem does not guarantee a convergent series and we did not get one. It is convenient to rewrite Eq. (14.133) as K.z/Dr 2zez P.iz/Ci Q.iz/ ; (14.134) where P.z/1.1/.9/ 2W.8z/2C.1/.9/.25/.49/ 4W.8z/4; (14.135) Q.z/1 1W.8z/.1/.9/.25/ 3W.8z/3C; (14.136) andD42. It should be noted that although P.z/ofEq. (14.135) and Q.z/of Eq. (14.136) have alternating signs, the series for P.iz/andQ.iz/inEq. (14.134) have all positive signs. Finally, note that for zlarge, Pdominates. Additional Asymptotic Forms We started our detailed study of asymptotic behavior with Kbecause, with its properties in hand, we can deduce the asymptotic expansions of the other members of the family of Bessel-related functions. 1. Rearranging the definition of Kto H.1/ .x/D2 e.i=2/.C1/K.ix/; (14.137) we have H.1/ .z/Dr 2 zexp i z C1 2 2 P.z/Ci Q.z/ ; (14.138) which although originally derived for real values of ix, can be analytically continued into the larger range < argz<2. 2. The second Hankel function is just (for real arguments) the complex conjugate of the first, and therefore H.2/ .z/Dr 2 zexp i z C1 2 2 P.z/i Q.z/ ; (14.139) valid for2< argz<. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.6 Asymptotic Expansions 693 3. Since J.z/is the real part of H.1/ .z/for real z, J.z/Dr 2 z P.z/cos z C1 2 2 Q.z/sin z C1 2 2 ; (14.140) valid for< argz<. 4. The Neumann function is the imaginary part of H.1/ .z/for real z, or Y.z/Dr 2 z P.z/sin z C1 2 2 CQ.z/cos z C1 2 2 ; (14.141) also valid for< argz<. 5. Finally, the modified Bessel function I.z/is given by I.z/DiJ.iz/; (14.142) so I.z/Dez p2z P.iz/i Q.iz/ ; (14.143) valid for=2<argz<=2 . Properties of the Asymptotic Forms Having derived the asymptotic forms of the various Bessel functions, it is opportune to note their essential characteristics. Remembering that in the limit of large z,Papproaches unity while Q1=z, we see that at large z, all the Bessel functions have leading terms with a 1=z1=2dependence, multiplied by either a real or complex exponential. The modified functions KandI, respectively, contain decreasing and increasing exponentials, while the ordinary Bessel functions JandYhave leading terms with sinusoidal oscillation (damped by the z1=2factor). When multiplied by a time factor ei!t, the Hankel functions can describe incoming and outgoing traveling waves. Looking at the oscillatory functions J,Y,H.i/ in more detail, we see that exact sinu- soidal behavior is only reached in the limit of large z, as for finite zthe terms involving Q will to some extent alter the periodicity. The reader may wish to compare the positions of the zeros of Jnin Table 14.1 with those predicted by its leading term, namely the zeros of cos z nC1 2 2 : We see that Jnbehaves asymptotically like a phase-shifted cosine function, with the phase shift a function of n. The asymptotic form of Ynwill be that of a sine function, with (for the same n) the same phase shift. This causes the zeros of JnandYnfor large zto alternate, as we saw for J0andY0in Fig. 14.8. The asymptotic behavior of the two solutions to a problem described by ordinary or modified Bessel functions may be sufficient to eliminate immediately one of these func- tions as a solution for a physical problem. This observation may enable us to use the behav- ior at zD1 as well as that at zD0to restrict the functional forms we need to consider. ArfKen_17-ch14-0643-0714- 9780123846549.tex 694 Chapter 14 Bessel Functions 5.60 4.80 4.00 3.20 2.40 1.60 0.80 0.00−1.20−0.80J0(x) x −0.400.000.400.801.20 2cos (x −πx 4π) FIGURE 14.13 Asymptotic approximation of J0.x/: Finally, we note that the asymptotic series P.z/andQ.z/, Eqs. (14.135) and (14.136), terminate for D1=2 ,3=2;::: and become polynomials (in negative powers of z). For these special values of the asymptotic approximations become exact solutions. It is of some interest to consider the accuracy of the asymptotic forms, taking for exam- ple just the first term Jn.x/r 2 xcos x nC1 2 2 : (14.144) Clearly, the condition for Eq. (14.144) to be accurate is that the sine term of Eq. (14.140) be negligible; that is, 8x4n21: (14.145) InFig. 14.13 we plot J0.x/and the leading term of its asymptotic approximation. The agreement is nearly quantitative for x>5. However, for nor>1the asymptotic region may be far out. Another use of the asymptotic formulas is to establish the constants in Wronskian for- mulas, where we know the Wronskian of any two Bessel functions of argument xhas a 1=xfunctional dependence but with a premultiplying constant that depends on the Bessel functions involved. Example 14.6.1 CYLINDRICAL TRAVELING WAVES As an illustration of a problem in which we have chosen a specific Bessel function because of its asymptotic properties, consider a two-dimensional (2-D) wave problem similar to the vibrating circular membrane of Exercise 14.1.24. Now imagine that the waves are gener- ated at rD0and move outward to infinity. We replace our standing waves by traveling ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.6 Asymptotic Expansions 695 ones. The differential equation remains the same, but the boundary conditions change. We now demand that for large rthe wave behave like Uei.kr!t/; (14.146) to describe an outgoing wave with wavelength 2=k. We assume, for simplicity, that there is no azimuthal dependence, so we have circular symmetry, implying mD0. The Bessel function of order zero with this asymptotic dependence is H.1/ 0.kr/, as can be seen from Eq. (14.138). This boundary condition at infinity then determines our wave solution as U.r;t/DH.1/ 0.kr/ei!t: (14.147) This solution diverges as r!0, which is the behavior to be expected with a source at the origin.  Exercises 14.6.1 Determine the asymptotic dependence of the modified Bessel functions I.x/, given I.x/D1 2iZ Ce.x=2/.tC1=t/dt tC1: The contour starts and ends at tD1 , encircling the origin in a positive sense. There are two saddle points. Only the one at zDC1 contributes significantly to the asymptotic form. 14.6.2 Determine the asymptotic dependence of the modified Bessel function of the second kind, K.x/, by using K.x/D1 21Z 0e.x=2/.sC1=s/ds s1: 14.6.3 Verify that the integral representations In.z/D1 Z 0ezcostcos.nt/dt; K.z/D1Z 0ezcosh tcosh. t/dt;<e.z/>0; satisfy the modified Bessel equation by direct substitution into that equation. How can you check the normalization? 14.6.4 (a) Show that when Kis defined by Eq. (14.128), d K0.z/ dzDK1.z/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 696 Chapter 14 Bessel Functions t plane (2) (1) −1 1 FIGURE 14.14 Modified Bessel function contours. (b) Show that the indefinite integral of K1.x/as defined by Eq. (14.128) has in the limit of small zthe valuelnzCC, and therefore, by comparison with Eq. (14.110), that K0as defined by Eq. (14.128) has the correct normalization. 14.6.5 Verify that Eq. (14.132) can be rearranged to the form given as Eq. (14.133). 14.6.6 (a) Show that y.z/DzZ ezt.t21/1=2dt satisfies the modified Bessel equation, provided the contour is chosen so that ezt.t21/C1=2 has the same value at the initial and final points of the contour. (b) Verify that the contours shown in Fig. 14.14 are suitable for this problem. 14.6.7 Use the asymptotic expansions to verify the following Wronskian formulas: (a) J.x/J1.x/CJ.x/JC1.x/D2 sin= x, (b) J.x/NC1.x/JC1.x/N.x/D2= x, (c) J.x/H.2/ 1.x/J1.x/H.2/ .x/D2=ix, (d) I.x/K0 .x/I0 .x/K.x/D1= x, (e) I.x/KC1.x/CIC1.x/K.x/D1=x. 14.6.8 Verify that the Green’s function for the 2-D Helmholtz equation (operator r2Ck2) with outgoing-wave boundary conditions is G.1;2/Di 4H.1/ 0.kj12j/: Hint. H.1/ 0.k/is known to be an outgoing-wave solution to the homogeneous Helmholtz equation. 14.6.9 From the asymptotic form of K.z/,Eq. (14.134), derive the asymptotic form of H.1/ .z/,Eq. (14.138). Note particularly the phase, .C1 2/=2 . 14.6.10 Apply Stokes’ method for obtaining an asymptotic expansion for the Hankel function H.1/ as follows: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.6 Asymptotic Expansions 697 (a) Replace the Bessel function in Bessel’s equation by x1=2y.x/and show that y.x/ satisfies y00.x/C 121 4 x2! y.x/D0: (b) Develop a power-series solution with negative powers of xstarting with the assumed form y.x/Deix1X nD0anxn: Obtain the recurrence relation giving anC1in terms of an. Check your result against the asymptotic series, Eq. (14.138). (c) From Eq. (14.125), determine the initial coefficient, a0. 14.6.11 Using the method of steepest descents, evaluate the second Hankel function given by H.2/ .t/D1 iZ C2e.t=2/.z1=z/dz zC1; with contour C2as shown in Fig. 14.10. ANS. H.2/ .t/r 2 tei.t=4=2/. 14.6.12 (a) In applying the method of steepest descents to the Hankel function H.1/ .t/, show thatw.z;t/, which appears in Eq. (14.121), satisfies <eTw.z;t/U<<eTw.z0;t/UD0 forzon the contour C1(Fig. 14.10) but away from the point zDz0Di. (b) For general values of zDrei, show that <eTw.z;t/U>0for 0<r<1;8 < : 2< < 2 and <eTw.z;t/U<0for r>1; 2<< 2: Your demonstration verifies that the distribution of the sign of wis as shown schematically in Fig. 14.15. (c) Explain why the contour C1(Fig. 14.10) cannot be deformed to go through both saddle points, and why it may not go through the saddle point at iif it is to end atzD1 with argumentC. 14.6.13 Calculate the first 15 partial sums of P0.x/andQ0.x/, Eqs. (14.135) and (14.136). Letxvary from 4 to 10 in unit steps. Determine the number of terms to be retained for maximum accuracy and the accuracy achieved as a function of x. Specifically, how small may xbe without raising the error above 3106? ANS. xminD6. ArfKen_17-ch14-0643-0714- 9780123846549.tex 698 Chapter 14 Bessel Functions + + + + −− − −y x FIGURE 14.15 Sign ofw.z;t/, occurring in Eq. (14.121), for integral representation of Hankel functions. 14.7 S PHERICAL BESSEL FUNCTIONS In Section 9.4 we discussed the separation of the Helmholtz equation in spherical coordi- nates. We showed there that in the oft-occurring case that the boundary conditions of the problem have spherical symmetry, the radial equation has the form given in Eq. (9.80), namely, r2d2R dr2C2rd R drC k2r2l.lC1/ RD0: (14.148) We remind the reader that the parameter kis that from the original Helmholtz equation, while l.lC1/is the separation constant associated with solutions of the angular equations identified by the index l(which is required by the boundary conditions to be an integer). In Section 9.4 we went on to discuss the fact that the substitution R.kr/DZ.kr/ .kr/1=2(14.149) permits us to rewrite Eq. (14.148) as r2d2Z dr2Crd Z drC" k2r2 lC1 22# ZD0; (14.150) which we identified in Eq. (9.84) as Bessel’s equation of order lC1 2. We can now identify the general solution Z.kr/as a linear combination of JlC1=2.kr/ andYlC1=2.kr/, which in turn means that we can write R.kr/in terms of these Bessel functions of half-integral order, illustrated (for JlC1=2) by R.kr/DCp krJlC1=2.kr/: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 699 Since the R.kr/describe radial functions in spherical coordinates, they are termed spheri- cal Bessel functions. Note also that since Eq. (14.148) is homogeneous, we are free to define our spherical Bessel functions at any scale; the scale ordinarily used is that intro- duced in the next subsection. Definitions We define our spherical Bessel functions by the following equations. It is not ordinarily useful to introduce spherical Bessel functions with indices that are not integers, so we assume the index nto be integral (but not necessarily nonnegative). jn.x/Dr 2xJnC1=2.x/; yn.x/Dr 2xYnC1=2.x/; (14.151) h.1/ n.x/Dr 2xH.1/ nC1=2.x/Djn.x/Ciyn.x/; h.2/ n.x/Dr 2xH.2/ nC1=2.x/Djn.x/iyn.x/: Referring to the definition of YnC1=2, we see that YnC1=2.x/Dcos.nC1 2/JnC1=2.x/Jn1=2.x/ sin.nC1 2/D.1/nC1Jn1 2.x/; which means that yn.x/D.1/nC1jn1.x/: (14.152) These spherical Bessel functions (Figs. 14.16 and14.17) can be expressed in series form. Using Eq. (14.6), we have initially jn.x/Dr 2x1X sD0.1/s sW0.sCnC3 2/x 22sCnC1=2 : (14.153) Writing 0.sCnC3 2/D0.nC3 2/.nC3 2/s; (14.154) where.::/sis a Pochhammer symbol, defined in Eq. (1.72), we can bring Eq. (14.153) to the form jn.x/Dr 2xx 2nC1=2 1 0.nC3 2/1X sD0.1/s sW.nC3 2/sx 22s Dxn .2nC1/WW1X sD0.1/s sW.nC3 2/sx 22s : (14.155) ArfKen_17-ch14-0643-0714- 9780123846549.tex 700 Chapter 14 Bessel Functions x0.6 0.50.40.3 0.2 0.1 0 2 4 6 8 10 12 14 −0.1 −0.2−0.31.0 j 0(x) j1(x) j2(x) FIGURE 14.16 Spherical Bessel functions. x0.3 0.2 0.1 0135 71 1 13 −0.1 −0.2 −0.3−0.4y1(x)y0(x) y2(x) 9 FIGURE 14.17 Spherical Neumann functions. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 701 We reached the last line of Eq. (14.155) by writing0.nC3 2/using the double factorial notation (compare with Exercise 13.1.14). If we now develop a series expansion for yn.x/by the same method that was used for jn.x/, but starting from Eq. (14.152), we get yn.x/D.2n1/WW xnC11X sD0.1/s sW.1 2n/sx 22s : (14.156) The spherical Bessel functions are oscillatory, as can be seen from the graphs in Figs. 14.16 and 14.17. Note that jn.x/are regular at xD0, with limiting behavior there proportional to xn. The ynare all irregular at xD0, approaching that point as xn1. The infinite series in Eqs. (14.155) and(14.156) can be evaluated in closed form (but with increasing difficulty as nincreases). For the special case nD0, we can substitute into Eq. (14.155) sWD2s.2s/WWand.3=2/ sD2s.2sC1/WW, reaching j0.x/D1X sD0.1/s22s .2s/WW.2sC1/WWx 22s D1X sD0.1/s .2sC1/Wx2s Dsinx x: (14.157) A similar treatment of the expansion for y0yields y0.x/Dcosx x: (14.158) From the definition of the spherical Hankel functions, Eq. (14.151), we also have h.1/ 0.x/D1 x.sinxicosx/Di xeix; (14.159) h.2/ 0.x/D1 x.sinxCicosx/Di xeix: (14.160) Since we anticipate the availability of recurrence formulas for the spherical Bessel func- tions, and since y0is justj1, we expect all the jnandynto be linear combinations of sines and cosines. In fact, the recurrence formulas are good ways of getting these functions for small n. However, we identify here an alternate approach, which depends on the fact, noted in Section 14.6, that the asymptotic expansion for the Hankel functions actually ter- minates when the order is a half-integer, thereby yielding exact, closed expressions. We start from h.1/ n.x/Dr 2xH.1/ nC1=2.x/ D.i/nC1eix x PnC1=2.x/Ci QnC1=2.x/ ; (14.161) ArfKen_17-ch14-0643-0714- 9780123846549.tex 702 Chapter 14 Bessel Functions where PandQare given by Eqs. (14.135) and(14.136). Now, PnC1=2 andQnC1=2 are polynomials, and we can bring Eq. (14.161) to the form h.1/ n.x/D.i/nC1eix xnX sD0is sW.8x/s.2nC2s/WW .2n2s/WW D.i/nC1eix xnX sD0is sW.2x/s.nCs/W .ns/W: (14.162) For real x,jn.x/is the real part of this, yn.x/the imaginary part, and h.2/ n.x/the complex conjugate. Specifically, h.1/ 1.x/Deix 1 xi x2 ; (14.163) h.1/ 2.x/Deixi x3 x23i x3 ; (14.164) j1.x/Dsinx x2cosx x; (14.165) j2.x/D3 x31 x sinx3 x2cosx; (14.166) y1.x/Dcosx x2sinx x; (14.167) y2.x/D3 x31 x cosx3 x2sinx: (14.168) Recurrence Relations The recurrence relations to which we now turn provide a convenient way of developing the higher-order spherical Bessel functions. These recurrence relations may be derived from the power-series expansions, but it is easier to substitute into the known recurrence relations, Eqs. (14.8) and (14.9). This gives fn1.x/CfnC1.x/D2nC1 xfn.x/; (14.169) n fn1.x/.nC1/fnC1.x/D.2nC1/f0 n.x/: (14.170) Rearranging these relations, or substituting into Eqs. (14.10) and(14.11), we obtain d dxTxnC1fn.x/UDxnC1fn1.x/; (14.171) d dxTxnfn.x/UD xnfnC1.x/: (14.172) In these equations fnmay represent jn,yn,h.1/ n, orh.2/ n. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 703 By mathematical induction (Section 1.4) we may establish the Rayleigh formulas: jn.x/D.1/nxn1 xd dxnsinx x ; (14.173) yn.x/D.1/nxn1 xd dxncosx x ; (14.174) h.1/ n.x/Di.1/nxn1 xd dxneix x ; (14.175) h.2/ n.x/Di.1/nxn1 xd dxneix x : (14.176) Limiting Values Forx1,9Eqs. (14.155) and (14.156) yield jn.x/xn .2nC1/WW; (14.177) yn.x/.2n1/WW xnC1: (14.178) The limiting values of the spherical Hankel functions for small xgo asiyn.x/. The asymptotic values of jn,yn,h.1/ n, and h.2/ nmay be obtained from the asymptotic forms of the corresponding Bessel functions, as given in Section 14.6. We find jn.x/1 xsin xn 2 ; (14.179) yn.x/1 xcos xn 2 ; (14.180) h.1/ n.x/.i/nC1eix xDiei.xn=2/ x; (14.181) h.2/ n.x/inC1eix xDiei.xn=2/ x: (14.182) The condition for these spherical Bessel forms is that xn.nC1/=2. From these asymp- totic values we see that jn.x/andyn.x/are appropriate for a description of standing spherical waves; h.1/ n.x/andh.2/ n.x/correspond to traveling spherical waves. If the time dependence for the traveling waves is taken to be ei!t, then h.1/ n.x/yields an outgoing traveling spherical wave, and h.2/ n.x/an incoming wave. Radiation theory in electromag- netism and scattering theory in quantum mechanics provide many applications. 9The condition that the second term in the series be negligible compared to the first is actually x2T.2nC2/.2nC3/= .nC1/U1=2forjn.x/. ArfKen_17-ch14-0643-0714- 9780123846549.tex 704 Chapter 14 Bessel Functions Orthogonality and Zeros We may take the orthogonality integral for the ordinary Bessel functions, Eqs. (11.49) and (11.50), aZ 0J p a J q a dDa2 2 JC1. p/2pq; and rewrite it in terms of jnto obtain aZ 0jn npr a jn nqr a r2drDa3 2 jnC1. np/2pq: (14.183) Here npis the p-th positive zero of jn. Note that in contrast to the formula for the orthogonality of the J, Eq. (14.183) has the weight factor r2, not r. This of course comes from the factors x1=2in the definition of jn.x/, but also has the effect that if the integration is construed as being over a spherical volume rather than a linear interval, it is the factor corresponding to uniform weight of all volume elements. (Remember that the weight for the Jintegral produces uniform weight if we construe the integration in that case as over the area within a circle.) As for the ordinary Bessel functions, the functions that are orthogonal on .0;a/all sat- isfy a Dirichlet boundary condition, with zeros at rDa. We therefore find it useful to know the values of the zeros of the jn. The first few zeros for small n, and also the locations of the zeros of j0 n, are listed in Table 14.2. The following example illustrates a problem in which the zeros of the jnplay an essential role. Table 14.2 Zeros of the Spherical Bessel Functions and Their First Derivatives Number of zero j0.x/ j1.x/ j2.x/ j3.x/ j4.x/ j5.x/ 1 3:1416 4:4934 5:7635 6:9879 8:1826 9:3558 2 6:2832 7:7253 9:0950 10:4171 11:7049 12:9665 3 9:4248 10:9041 12:3229 13:6980 15:0397 16:3547 4 12:5664 14:0662 15:5146 16:9236 18:3013 19:6532 5 15:7080 17:2208 18:6890 20:1218 21:5254 22:9046 j0 0.x/ j0 1.x/ j0 2.x/ j0 3.x/ j0 4.x/ j0 5.x/ 1 4:4934 2:0816 3:3421 4:5141 5:6467 6:7565 2 7:7253 5:9404 7:2899 8:5838 9:8404 11:0702 3 10:9041 9:2058 10:6139 11:9727 13:2956 14:5906 4 14:0662 12:4044 13:8461 15:2445 16:6093 17:9472 5 17:2208 15:5792 17:0429 18:4681 19:8624 21:2311 ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 705 Example 14.7.1 PARTICLE IN A SPHERE An illustration of the use of the spherical Bessel functions is provided by the problem of a quantum mechanical particle of mass min a sphere of radius a. Quantum theory requires that the wave function , describing our particle, satisfy the Schrödinger equation Nh2 2mr2 DE ; (14.184) subject to the conditions that (1) .r/is finite for all 0ra, and (2) .a/D0. This corresponds to a square-well potential VD0forra,VD1 forr>a. HereNhis Planck’s constant divided by 2. Equation (14.184) with its boundary conditions is an eigenvalue equation; its eigenvalues Eare the possible values of the particle’s energy. Let us determine the minimum value of the energy for which our wave equation has an acceptable solution. Equation (14.184) is Helmholtz’s equation, which after separation of variables leads to the radial equation previously presented as Eq. (14.148): d2R dr2C2 rd R drC k2l.lC1/ r2 RD0; (14.185) with k2D2mE=Nh2(14.186) andl(determined from the angular equation) a nonnegative integer. Comparing with Eq. (14.150) and the definitions of the spherical Bessel functions, Eq. (14.151), we see that the general solution to Eq. (14.185) is RDAjl.kr/CByl.kr/: (14.187) To satisfy the boundary conditions of the present problem, we must reject the solution yl because it is singular at rD0, and we must choose ksuch that jl.ka/D0. This boundary condition at rDacan be satisfied if kkliD li a; (14.188) where liis the ith positive zero of jl. From Eq. (14.186) we see that the smallest E will correspond to the smallest acceptable k, which in turn corresponds to the smallest li. Thus, scanning Table 14.2, we identify the smallest lias the first zero of j0, a result which we would expect after we have learned that the value lD0is associated with an angular function with no kinetic energy. We conclude this example by solving Eq. (14.186) forEwith kassigned the value 01=aD=a10: EminD2Nh2 2ma2Dh2 8ma2: (14.189) 10Most of the entries in Table 14.2 are only accessible numerically, but the zeros of j0are readily identified due to their simple form, 0mDm. ArfKen_17-ch14-0643-0714- 9780123846549.tex 706 Chapter 14 Bessel Functions This example illustrates several features common to bound-state problems in quantum mechanics. First, we see that for any finite sphere the particle will have a positive minimum or zero-point energy. Second, we note that the particle cannot have a continuous range of energy values; the energy is restricted to discrete values corresponding to the eigenvalues of the Schrödinger equation. Third, the possible energies in this spherically symmetric problem depend on l; as is evident from the table of zeros of jl, the minimum energy for a given lincreases with l. Finally, note that the orthogonality of the jlunder the conditions of this problem shows us that the eigenfunctions corresponding to the same lbut different iare orthogonal (with the weight factor corresponding to spherical polar coordinates).  We close this subsection with the observation that, in addition to orthogonality with respect to the scaling (to bring zeros to a specified rvalue), the spherical Bessel functions also possess orthogonality with respect to the indices: 1Z 1jm.x/jn.x/dxD0; m6Dn;m;n0: (14.190) The proof is left as Exercise 14.7.12. If mDn(compare Exercise 14.7.13), we have 1Z 1Tjn.x/U2dxD 2nC1: (14.191) The spherical Bessel functions will enter again in connection with spherical waves, but further consideration is postponed until the corresponding angular functions, the Legendre functions, have been more thoroughly discussed. Modifed Spherical Bessel Functions Problems involving the radial equation r2d2R dr2C2rd R dr k2r2Cl.lC1/ RD0; (14.192) which differs from Eq. (14.148) only in the sign of k2, also arise frequently in physics. The solutions to this equation are spherical Bessel functions with imaginary arguments, leading us to define modified spherical Bessel functions (Fig. 14.18) as follows: in.x/Dr 2xInC1=2.x/; (14.193) kn.x/Dr 2 xKnC1=2.x/: (14.194) Note that the scale factor in the definition of kndiffers from that of the other spherical Bessel functions. ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 707 x6 54321 012345k0(x) k1(x) i1(x)i0(x) FIGURE 14.18 Modified spherical Bessel functions. With the above definitions, these functions have the following recurrence relations: in1.x/inC1.x/D2nC1 xin.x/; nin1.x/C.nC1/inC1.x/D.2nC1/i0 n.x/;(14.195) kn1.x/knC1.x/D2nC1 xkn.x/; nkn1.x/C.nC1/knC1.x/D.2nC1/k0 n.x/: The first few of these functions are i0.x/Dsinhx x; k0.x/Dex x; i1.x/Dcosh x xsinhx x2; k1.x/Dex1 xC1 x2 ; (14.196) i2.x/Dsinhx1 xC3 x3 3 cosh x x2;k2.x/Dex1 xC3 x2C3 x3 : ArfKen_17-ch14-0643-0714- 9780123846549.tex 708 Chapter 14 Bessel Functions Limiting values of the modified spherical Bessel functions are, for small x, in.x/xn .2nC1/WW;kn.x/.2n1/WW xnC1: (14.197) For large z, the asymptotic behavior of these functions is in.x/ex 2x;kn.x/ex x: (14.198) Example 14.7.2 PARTICLE IN A FINITE SPHERICAL WELL As a final example, we return to the problem of a particle trapped in a spherical potential well of radius a(Example 14.7.1), but instead of confining the particle by a wall at potential VD1 (equivalent to requiring that its wave function vanish at rDa), we now consider a well of finite depth, corresponding to V.r/DV0<0;0ra; 0; r>a: If the particle can have an energy E<0, it will be localized in and near the potential well, with a wave function that decays to zero as rincreases to values greater than a. A simple case of this problem was one of our examples of an eigenvalue problem (Example 8.3.3), but in that case we did not proceed with enough generality to identify its solutions as Bessel functions. This problem is governed by the Schrödinger equation, which now has the form Nh2 2mr2 CV.r/ DE : This is an eigenvalue equation, to be solved for andEover the full three-dimensional space, subject to the condition that be continuous and differentiable for all r, and that it be normalizable (thus approaching zero asymptotically at large r). Here mis the mass of the particle andNhis Planck’s constant divided by 2. While this problem is more difficult than that of Example 14.7.1, it becomes manageable if we realize that it is equivalent to two separate problems for the respective regions 0 raandr>a, within each of which the potential has a constant value, but constrained to (1) have the same eigenvalue E, and (2) connect smoothly (so the rderivative will exist) atrDa. When our Schrödinger equation is processed by the method of separation of variables, we obtain as its radial component d2R dr2C2 rd R drC 2m Nh2 EV.r/ l.lC1/ r2! RD0; which is either the spherical Bessel equation, Eq. (14.150), or the modified spherical Bessel equation, Eq. (14.192), depending on the sign of EV.r/. We see that if V0<E<0, then ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 709 forrawe will have EV.r/>0, yielding a Bessel ODE with an acceptable solution involving jl, while for r>awe have EV.r/<0, leading to a modified Bessel ODE for which we can choose the klsolution to obtain the necessary asymptotic behavior. Summarizing the above, we have, for the two regions: Rin.r/DAjl.kr/; k2D2m Nh2.EV0/ra; Rout.r/DBkl.k0r/;k02D2m Nh2E r>a: Smooth connection at rDathen corresponds to the equations Rin.a/DRout.a/! Ajl.ka/DBkl.k0r/; (14.199) d Rin dr rDaDd Rout dr rDa! k Aj0 l.ka/Dk0Bk0 l.k0a/: (14.200) ForlD0this problem reduces to that considered in Example 8.3.3, where we indicate a numerical procedure of solving it, but we are now in a position to obtain solutions for alll.  Exercises 14.7.1 Show how one can obtain Eq. (14.162) starting from Eq. (14.161). 14.7.2 Show that if yn.x/Dr 2xYnC1=2.x/; it automatically equals .1/nC1r 2xJn1=2.x/: 14.7.3 Derive the trigonometric-polynomial forms of jn.z/andyn.z/11: jn.z/D1 zsin zn 2Tn=2UX sD0.1/s.nC2s/W .2s/W.2z/2s.n2s/W C1 zcos zn 2T.n1/=2UX sD0.1/s.nC2sC1/W .2sC1/W.2 z/2s.n2s1/W; 11The upper summation limit Tn=2Umeans the largest integer that does not exceed n=2. ArfKen_17-ch14-0643-0714- 9780123846549.tex 710 Chapter 14 Bessel Functions yn.z/D.1/nC1 zcos zCn 2Tn=2UX sD0.1/s.nC2s/W .2s/W.2z/2s.n2s/W C.1/nC1 zsin zCn 2T.n1/=2UX sD0.1/s.nC2sC1/W .2sC1/W.2 z/2sC1.n2s1/W: 14.7.4 Use the integral representation of J.x/, J.x/D1 1=20.C1 2/x 21Z 1eixp.1p2/1=2dp; to show that the spherical Bessel functions jn.x/are expressible in terms of trigono- metric functions; that is, for example, j0.x/Dsinx x;j1.x/Dsinx x2cosx x: 14.7.5 (a) Derive the recurrence relations fn1.x/CfnC1.x/D2nC1 xfn.x/; n fn1.x/.nC1/fnC1.x/D.2nC1/f0 n.x/ satisfied by the spherical Bessel functions jn.x/;yn.x/;h.1/ n.x/, and h.2/ n.x/. (b) Show, from these two recurrence relations, that the spherical Bessel function fn.x/satisfies the differential equation x2f00 n.x/C2x f0 n.x/C x2n.nC1/ fn.x/D0: 14.7.6 Prove by mathematical induction (Section 1.4) that jn.x/D.1/nxn1 xd dxnsinx x forn, an arbitrary nonnegative integer. 14.7.7 From the discussion of orthogonality of the spherical Bessel functions, show that a Wronskian relation for jn.x/andnn.x/is jn.x/y0 n.x/j0 n.x/yn.x/D1 x2: 14.7.8 Verify h.1/ n.x/h.2/0 n.x/h.1/0 n.x/h.2/ n.x/D2i x2: 14.7.9 Verify Poisson’s integral representation of the spherical Bessel function, jn.z/Dzn 2nC1nWZ 0cos.zcos/sin2nC1d: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 711 14.7.10 A well-known integral representation for K.x/has the form K.x/D20.C1 2/px1Z 0cosxt .t2C1/C1=2dt: Starting from this formula, show that kn.x/D2nC2.nC1/W xnC11Z 0k2j0.kx/ .k2C1/nC2dk: 14.7.11 Show that1Z 0J.x/J.x/dx xD2 sinT./=2U 22; C>0: 14.7.12 Derive Eq. (14.190):1Z 1jm.x/jn.x/dxD0;m6Dn; m;n0: 14.7.13 Derive Eq. (14.191):1Z 1 jn.x/2dxD 2nC1: 14.7.14 The Fresnel integrals (Fig. 14.19 and Exercise 12.7.2) occurring in diffraction theory are given by x.t/Dr 2Cr 2t DtZ 0cos.v2/dv; y.t/Dr 2sr 2t DtZ 0sin.v2/dv: Show that these integrals may be expanded in series of spherical Bessel functions as follows: x.s/D1 2sZ 0j1.u/u1=2duDs1=21X nD0j2n.s/; y.s/D1 2sZ 0j0.u/u1=2xduDs1=21X nD0j2nC1.s/: Hint. To establish the equality of the integral and the sum, you may wish to work with their derivatives. The spherical Bessel analogs of Eqs. (14.8) and(14.12) may be helpful. ArfKen_17-ch14-0643-0714- 9780123846549.tex 712 Chapter 14 Bessel Functions xx(t) y(t) yx π 21 2⋅1.0 0.5 1.0 2 3 4 FIGURE 14.19 Fresnel integrals. 14.7.15 A hollow sphere of radius a(Helmholtz resonator) contains standing sound waves. Find the minimum frequency of oscillation in terms of the radius aand the velocity of soundv. The sound waves satisfy the wave equation r2 D1 v2@2 @t2 and the boundary condition@ @rD0;rDa: The spatial part of this PDE is the same as the PDE discussed in Example 14.7.1, but here we have a Neumann boundary condition, in contrast to the Dirichlet boundary condition of that example. ANS.minD0:3313v= a,maxD3:018 a. 14.7.16 (a) Show that the parity of in.x/(the behavior under x! x) is.1/n. (b) Show that kn.x/has no definite parity. 14.7.17 Show that the Wronskian of the spherical modified Bessel functions is given by in.x/k0 n.x/i0 n.x/kn.x/D1 x2: ArfKen_17-ch14-0643-0714- 9780123846549.tex 14.7 Spherical Bessel Functions 713 Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Jackson, J. D., Classical Electrodynamics, 3rd ed. New York: Wiley (1999). Morse, P. M., and H. Feshbach, Methods of Theoretical Physics, 2 vols. New York: McGraw-Hill (1953). This work presents the mathematics of much of theoretical physics in detail but at a rather advanced level. Watson, G. N., A Treatise on the Theory of Bessel Functions, 1st ed. Cambridge: Cambridge University Press (1922). Watson, G. N., A Treatise on the Theory of Bessel Functions, 2nd ed. Cambridge: Cambridge University Press (1952). This is the definitive text on Bessel functions and their properties. Although difficult reading, it is invaluable as the ultimate reference. Whittaker, E. T., and G. N. Watson, A Course of Modern Analysis, 4th ed. Cambridge: Cambridge University Press (1962), paperback. ArfKen_Ch15-9780123846549.tex CHAPTER 15 LEGENDRE FUNCTIONS Legendre functions are important in physics because they arise when the Laplace or Helmholtz equations (or their generalizations) for central force problems are separated in spherical coordinates. They therefore appear in the descriptions of wave functions for atoms, in a variety of electrostatics problems, and in many other contexts. In addition, the Legendre polynomials provide a convenient set of functions that is orthogonal (with unit weight) on the interval .1;C1/that is the range of the sine and cosine functions. And from a pedagogical viewpoint, they provide a set of functions that are easy to work with and form an excellent illustration of the general properties of orthogonal polynomials. Several of these properties were discussed in a general way in Chapter 12. We collect here those results, expanding them with additional material that is of great utility and importance. As indicated above, Legendre functions are encountered when an equation written in spherical polar coordinates .r;;'/ , such as r2 CV.r/ D ; is solved by the method of separation of variables. Note that we are assuming that this equation is to be solved for a spherically symmetric region and that V.r/is a func- tion of the distance from the origin of the coordinate system (and therefore not a func- tion of the three-component position vector r). As in Eqs. (9.77) and (9.78), we write DR.r/2./8.'/ and decompose our original partial differential equation (PDE) into the three one-dimensional ordinary differential equations (ODEs): d28 d'2Dm28; (15.1) 1 sind d sind2 d m22 sin2Cl.lC1/2D0; (15.2) 1 r2d dr r2d R dr Ch V.r/i Rl.lC1/R r2D0: (15.3) 715 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch15-9780123846549.tex 716 Chapter 15 Legendre Functions The quantities m2andl.lC1/are constants that occur when the variables are separated; the ODE in'is easy to solve and has natural boundary conditions (cf. Section 9.4), which dictate that mmust be an integer and that the functions 8can be written as eim'or as sin.m'/,cos.m'/. The2equation can now be transformed by the substitution xDcos, cf. Eq. (9.79), reaching .1x2/P00.x/2x P0.x/m2 1x2P.x/Cl.lC1/P.x/D0: (15.4) This is the associated Legendre equation; the special case with mD0, which we will treat first, is the Legendre ODE. 15.1 L EGENDRE POLYNOMIALS The Legendre equation, .1x2/P00.x/2x P0.x/CP.x/D0; (15.5) has regular singular points at xD1 andxD1 (see Table 7.1), and therefore has a series solution about xD0that has a unit radius of convergence, i.e., the series solution will (for all values of the parameter ) converge forjxj<1. In Section 8.3 we found that for most values of , the series solutions will diverge at xD1 (corresponding to D0and D), making the solutions inappropriate for use in central force problems. However, if has the value l.lC1/, with lan integer, the series become truncated after xl, leaving a polynomial of degree l. Now that we have identified the desired solutions to the Legendre equations as polyno- mials of successive degrees, called Legendre polynomials and designated Pl, let us use the machinery of Chapter 12 to develop them from a generating-function approach. This course of action will set a scale for the Pland provide a good starting point for deriving recurrence relations and related formulas. We found in Example 12.1.3 that the generating function for the polynomial solutions of the Legendre ODE is given by Eq. (12.27): g.x;t/D1p 12xtCt2D1X nD0Pn.x/tn: (15.6) To identify the scale that is given to Pnby Eq. (15.6), we simply set xD1in that equation, bringing its left-hand side to the form g.1;t/D1p 12tCt2D1 1tD1X nD0tn; (15.7) where the last step in Eq. (15.7) was to expand 1=.1t/using the binomial theorem. Comparing with Eq. (15.6), we see that the scaling it predicts is Pn.1/D1. ArfKen_Ch15-9780123846549.tex 15.1 Legendre Polynomials 717 Next, consider what happens if we replace xbyxandtbyt. The value of g.x;t/in Eq. (15.6) is unaffected by this substitution, but the right-hand side takes a different form: 1X nD0Pn.x/tnDg.x;t/Dg.x;t/D1X nD0Pn.x/.t/n; (15.8) showing that Pn.x/D.1/nPn.x/: (15.9) From this result it is obvious that Pn.1/D.1/n, and that Pn.x/will have the same parity as xn. Another useful special value is Pn.0/. Writing P2nandP2nC1to distinguish even and odd index values, we note first that because P2nC1is odd under parity, i.e., x! x, we must have P2nC1.0/D0. To obtain P2n.0/, we again resort to the binomial expansion: g.0;t/D.1Ct2/1=2D1X nD01=2 n t2nD1X nD0P2n.0/t2n: (15.10) Then, using Eq. (1.74) to evaluate the binomial coefficient, we get P2n.0/D.1/n.2n1/WW .2n/WW: (15.11) It is also useful to characterize the leading terms of the Legendre polynomials. Applying the binomial theorem to the generating function, .12xtCt2/1=2D1X nD01=2 n .2xtCt2/n; (15.12) from which we see that the maximum power of xthat can multiply tnwill be xn, and is obtained from the term .2xt/nin the expansion of the final factor. Thus, the coefficient of xninPn.x/is1=2 n .2/nD.2n1/WW nW: (15.13) These results are important, so we summarize: Pn.x/has sign and scaling such that Pn.1/D1and Pn.1/D.1/n. P2n.x/is an even function of x;P2nC1.x/is odd. P2nC1.0/D0, and P2n.0/ is given by Eq. (15.11). Pn.x/is a polynomial of degree ninx, with the coefficient of xngiven by Eq. (15.13); Pn.x/contains alternate powers of x:xn;xn2;;.x0orx1/. From the fact that Pnis of degree nwith alternate powers, it is clear that P0.x/D constant and that P1.x/D(constant) x. From the scaling requirements these must reduce to P0.x/D1andP1.x/Dx. Returning to Eq. (15.12), we can get explicit closed expressions for the Legendre poly- nomials. All we need to do is expand the quantity .2xtCt2/nand rearrange the summa- tions to identify the xdependence associated with each power of t. The result, which is in ArfKen_Ch15-9780123846549.tex 718 Chapter 15 Legendre Functions general less useful than the recurrence formulas to be developed in the next subsection, is Pn.x/DTn=2UX kD0.1/k.2n2k/W 2nkW.nk/W.n2k/Wxn2k: (15.14) HereTn=2Ustands for the largest integer n=2. This formula is consistent with the requirement that for neven, Pn.x/has only even powers of xand even parity, while fornodd, it has only odd powers of xand odd parity. Proof of Eq. (15.14) is the topic of Exercise 15.1.2. Recurrence Formulas From the generating function equation we can generate recurrence formulas by differenti- ating g.x;t/with respect to xort. We start from @g.x;t/ @tDxt .12xtCt2/3=2D1X nD0n Pn.x/tn1; (15.15) which we rearrange to .12xtCt2/1X nD0n Pn.x/tn1C.tx/1X nD0Pn.x/tnD0; (15.16) and then expand, reaching 1X nD0n Pn.x/tn121X nD0nx P n.x/tnC1X nD0n Pn.x/tnC1 C1X nD0Pn.x/tnC11X nD0x Pn.x/tnD0: (15.17) Collecting the coefficients of tnfrom the various terms and setting the result to zero, Eq. (15.17) is seen to be equivalent to .2nC1/x Pn.x/D.nC1/PnC1.x/Cn Pn1.x/;nD1;2;3;:::: (15.18) Equation (15.18) permits us to generate successive Pnfrom the starting values P0andP1 that we have previously identified. For example, 2P2.x/D3x P1.x/P0.x/! P2.x/D1 2 3x21 : (15.19) Continuing this process, we can build the list of Legendre polynomials given in Table 15.1. We can also obtain a recurrence formula involving P0 nby differentiating g.x;t/with respect to x. This gives @g.x;t/ @xDt .12xtCt2/3=2D1X nD0P0 n.x/tn; ArfKen_Ch15-9780123846549.tex 15.1 Legendre Polynomials 719 Table 15.1 Legendre Polynomials P0.x/D1 P1.x/Dx P2.x/D1 2.3x21/ P3.x/D1 2.5x33x/ P4.x/D1 8.35x430x2C3/ P5.x/D1 8.63x570x3C15x/ P6.x/D1 16.231x6315x4C105x25/ P7.x/D1 16.429x7693x5C315x335x/ P8.x/D1 128.6435 x812012 x6C6930 x41260 x2C35/ or .12xtCt2/1X nD0P0 n.x/tnt1X nD0Pn.x/tnD0: (15.20) As before, the coefficient of each power of tis set to zero and we obtain P0 nC1.x/CP0 n1.x/D2x P0 n.x/CPn.x/: (15.21) A more useful relation may be found by differentiating Eq. (15.18) with respect to xand multiplying by 2. To this we add .2nC1/times Eq. (15.21), canceling the P0 nterm. The result is P0 nC1.x/P0 n1.x/D.2nC1/Pn.x/: (15.22) Starting from Eqs. (15.21) and(15.22), numerous additional relations can be developed,1 including P0 nC1.x/D.nC1/Pn.x/Cx P0 n.x/; (15.23) P0 n1.x/Dn P n.x/Cx P0 n.x/; (15.24) .1x2/P0 n.x/Dn Pn1.x/nx P n.x/; (15.25) .1x2/P0 n.x/D.nC1/x Pn.x/.nC1/PnC1.x/: (15.26) Because we derived the generating function g.x;t/from the Legendre ODE and then obtained the recurrence formulas using g.x;t/, that ODE will automatically be consis- tent with these recurrence relations. It is nevertheless of interest to verify this consistency, because then we can conclude that anyset of functions satisfying the recurrence formulas will be a set of solutions to the Legendre ODE, and that observation will be relevant to 1Using the equation numbers in parentheses to indicate how they are to be combined, we may obtain some of these derivative formulas as follows: 2d dx.15:18/C.2nC1/.15:21/).15:22/;1 2f.15:21/C.15:22/g).15:23/; 1 2f.15:21/.15:22/g).15:24/; .15:23/ n!n1Cx.15:24/).15:25/: ArfKen_Ch15-9780123846549.tex 720 Chapter 15 Legendre Functions the Legendre functions of the second kind (solutions linearly independent of the polyno- mials Pl). A demonstration that functions satisfying the recurrence formulas also satisfy the Legendre ODE is the topic of Exercise 15.1.1. Upper and Lower Bounds for Pn.cos/ Our generating function can be used to set an upper limit on jPn.cos/j. We have .12tcosCt2/1=2D.1tei/1=2.1tei/1=2 D 1C1 2teiC3 8t2e2iC 1C1 2teiC3 8t2e2iC :(15.27) We may make two immediate observations from Eq. (15.27). First, when any term within the first set of parentheses is multiplied by any term from the second set of parentheses, the power of tin the product will be even if and only if min the net exponential eimis even. Second, for every term of the form tneim, there will be another term of the form tneim, and the two terms will occur with the same coefficient, which must be positive (since all the terms in both summations are individually positive). These two observations mean that: (1) Taking the terms of the expansion two at a time, we can write the coefficient of tnas a linear combination of forms 1 2anm.eimCeim/Danmcosm with all the anmpositive, and (2) The parity of nandmmust be the same (either they are both even, or both odd). This, in turn, means that Pn.cos/DnX mD0or1anmcosm: (15.28) This expression is clearly a maximum when D0, where we already know, from the Sum- mary following Eq. (15.11), that Pn.1/D1. Thus, The Legendre polynomial Pn.x/has a global maximum on the inter- val.1;C1/ atxD1, with value Pn.1/D1, and if nis even, also at xD1 . Ifnis odd, xD1 will be a global minimum on this interval with Pn.1/D1 . The maxima and minima of the Legendre polynomials can be seen from the graphs of P2through P5, in which are plotted in Fig. 15.1. Rodrigues Formula In Section 12.1 we showed that orthogonal polynomials could be described by Rodrigues formulas, and that the repeated differentiations occurring therein were good ArfKen_Ch15-9780123846549.tex 15.1 Legendre Polynomials 721 x1Pn(x) 0.5P2 P3P4P5 1 0 −0.5−1 FIGURE 15.1 Legendre polynomials P2.x/through P5.x/. starting points for developing properties of these functions. Applying Eq. (12.9), we find that the Rodrigues formula for the Legendre polynomials must be proportional to d dxn .1x2/n: (15.29) Equation (12.9) is not sufficient to set the scale of the orthogonal polynomials, and to bring Eq. (15.29) to the scaling already adopted via Eq. (15.6) we multiply Eq. (15.29) by .1/n=2nnW, so Pn.x/D1 2nnWd dxn .x21/n: (15.30) To establish that Eq. (15.30) has a scaling in agreement with our earlier analyses, it suffices to check the coefficient of a single power of x; we choose xn. From the Rodrigues formula, this power of xcan only arise from the term x2nin the expansion of .x21/n, and the coefficient of xninPn.x/(Rodrigues) is1 2nnW.2n/W nWD.2n1/WW nW; in agreement with Eq. (15.13). This confirms the scale of Eq. (15.30). ArfKen_Ch15-9780123846549.tex 722 Chapter 15 Legendre Functions Exercises 15.1.1 Derive the Legendre ODE by manipulation of the Legendre polynomial recurrence relations. Suggested starting point: Eqs. (15.24) and (15.25). 15.1.2 Derive the following closed formula for the Legendre polynomials Pn.x/. Pn.x/DTn=2UX kD0.1/k.2n2k/W 2nkW.nk/W.n2k/Wxn2k; whereTn=2Ustands for the integer part of n=2. Hint. Further expand Eq. (15.12) and rearrange the resulting double sum. 15.1.3 By differentiation and direct substitution of the series form given in Exercise 15.1.2, show that Pn.x/satisfies the Legendre ODE. Note that there is no restriction on x. We may have any x,1<x<1, and indeed any zin the entire finite complex plane. 15.1.4 The shifted Legendre polynomials, designated by the symbol P n.x/(where the as- terisk does notmean complex conjugate) are orthogonal with unit weight on T0;1U, with normalization integral hP njP niD1=.2nC1/. The P nthrough nD6are shown in Table 15.2. (a) Find the recurrence relation satisfied by the P n. (b) Show that all the coefficients of the P nare integers. Hint. Look at the closed formula in Exercise 15.1.2. 15.1.5 Given the series 0C 2cos2C 4cos4C 6cos6Da0P0Ca2P2Ca4P4Ca6P6; where the arguments of the Pnarecos, express the coefficients ias a column vector and the coefficients aias a column vector aand determine the matrices AandBsuch that A DaandBaD : Table 15.2 Shifted Legendre Polynomials P 0.x/D1 P 1.x/D2x1 P 2.x/D6x26xC1 P 3.x/D20x330x2C12x1 P 4.x/D70x4140x3C90x220xC1 P 5.x/D252x5630x4C560x3210x2C30x1 P 6.x/D924x62772 x5C3150 x41680 x3C420x242xC1 ArfKen_Ch15-9780123846549.tex 15.1 Legendre Polynomials 723 Check your computation by showing that ABD1(unit matrix). Repeat for the odd case 1cosC 3cos3C 5cos5C 7cos7Da1P1Ca3P3Ca5P5Ca7P7: Note. Pn.cos/andcosnare tabulated in terms of each other in AMS-55 (see Addi- tional Readings for the complete reference). 15.1.6 By differentiating the generating function g.x;t/with respect to t, multiplying by 2t, and then adding g.x;t/, show that 1t2 .12txCt2/3=2D1X nD0.2nC1/Pn.x/tn: This result is useful in calculating the charge induced on a grounded metal sphere by a nearby point charge. 15.1.7 (a) Derive Eq. (15.26), .1x2/P0 n.x/D.nC1/x Pn.x/.nC1/PnC1.x/: (b) Write out the relation of Eq. (15.26) to preceding equations in symbolic form analogous to the symbolic forms for Eqs. (15.22) to(15.25). 15.1.8 Prove that P0 n.1/Dd dxPn.x/jxD1D1 2n.nC1/: 15.1.9 Show that Pn.cos/D.1/nPn.cos/by use of the recurrence relation relating Pn, PnC1, and Pn1and your knowledge of P0andP1. 15.1.10 From Eq. (15.27) write out the coefficient of t2in terms of cosn,n2. This coeffi- cient is P2.cos/. 15.1.11 Derive the recurrence relation .1x2/P0 n.x/Dn Pn1.x/nx P n.x/ from the Legendre polynomial generating function. 15.1.12 Evaluate1Z 0Pn.x/dx. ANS. nD2s, 1 for sD0;0fors>0; nD2sC1, P2s.0/=.2sC2/D.1/s.2s1/WW=1.2sC2/WW. Hint. Use a recurrence relation to replace Pn.x/by derivatives and then integrate by inspection. Alternatively, you can integrate the generating function. 15.1.13 Show that each term in the summation nX rDTn=2UC1d dxn.1/rnW rW.nr/Wx2n2r vanishes ( randnintegral). HereTn=2Uis the largest integer n=2. ArfKen_Ch15-9780123846549.tex 724 Chapter 15 Legendre Functions 15.1.14 Show thatR1 1xmPn.x/dxD0when m<n. Hint. Use Rodrigues formula or expand xmin Legendre polynomials. 15.1.15 Show that 1Z 1xnPn.x/dxD2nW .2nC1/WW: Note. You are expected to use the Rodrigues formula and integrate by parts, but also see if you can get the result from Eq. (15.14) by inspection. 15.1.16 Show that 1Z 1x2rP2n.x/dxD22nC1.2r/W.rCn/W .2rC2nC1/W.rn/W;rn: 15.1.17 As a generalization of Exercises 15.1.15 and15.1.16, show that the Legendre expan- sions of xsare (a) x2rDrX nD022n.4nC1/.2r/W.rCn/W .2rC2nC1/W.rn/WP2n.x/;sD2r, (b) x2rC1DrX nD022nC1.4nC3/.2rC1/W.rCnC1/W .2rC2nC3/W.rn/WP2nC1.x/,sD2rC1. 15.1.18 In numerical work (for e.g., the Gauss-Legendre quadrature), it is useful to establish thatPn.x/hasnreal zeros in the interior of T1; 1U. Show that this is so. Hint. Rolle’s theorem shows that the first derivative of .x21/2nhas one zero in the interior ofT1; 1U:Extend this argument to the second, third, and ultimately the nth derivative. 15.2 O RTHOGONALITY Because the Legendre ODE is self-adjoint and the coefficient of P00.x/, namely.1x2/, vanishes at xD1 , its solutions of different nwill automatically be orthogonal with unit weight on the interval .1; 1/, 1Z 1Pn.x/Pm.x/dxD0; .n6Dm/: (15.31) Because the Pnare real, no complex conjugation needs to be indicated in the orthogonality integral. Since Pnis often used with argument cos, we note that Eq. (15.31) is equivalent ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 725 to Z 0Pn.cos/Pm.cos/sindD0; .n6Dm/: (15.32) The definition of the Pndoes not guarantee that they are normalized, and in fact they are not. One way to establish the normalization starts by squaring the generating-function formula, yielding initially .12xtCt2/1D"1X nD0Pn.x/tn#2 : (15.33) Integrating from xD1 toxD1and dropping the cross terms because they vanish due to orthogonality, Eq. (15.31), we have 1Z 1dx 12txCt2D1X nD0t2n1Z 1h Pn.x/i2 dx: (15.34) Making now the substitution yD12txCt2, with dyD2t dx , we obtain 1Z 1dx 12txCt2D1 2t.1Ct/2Z .1t/2dy yD1 tln1Ct 1t : (15.35) Expanding this result in a power series (Exercise 1.6.1), 1 tln1Ct 1t D21X nD0t2n 2nC1; (15.36) and equating the coefficients of powers of tinEqs. (15.34) and(15.36), we must have 1Z 1h Pn.x/i2 dxD2 2nC1: (15.37) Combining Eqs. (15.31) and (15.37), we have the orthonormality condition 1Z 1Pn.x/Pm.x/dxD2nm 2nC1: (15.38) This result can also be obtained using the Rodrigues formulas for PnandPm. See Exer- cise 15.2.1. ArfKen_Ch15-9780123846549.tex 726 Chapter 15 Legendre Functions Legendre Series The orthogonality of the Legendre polynomials makes it natural to use them as a basis for expansions. Given a function f.x/defined on the range .1; 1/, the coefficients in the expansion f.x/D1X nD0anPn.x/ (15.39) are given by the formula anD2nC1 21Z 1f.x/Pn.x/dx: (15.40) The orthogonality property guarantees that this expansion is unique. Since we can (but perhaps will not wish to) convert our expansion into a power series by inserting the expan- sion of Eq. (15.14) and collecting the coefficients of each power of x, we can also obtain a power series, which we thereby know must be unique. An important application of Legendre series is to solutions of the Laplace equation. We saw in Section 9.4 that when the Laplace equation is separated in spherical polar coordi- nates, its general solution (for spherical symmetry) takes the form .r;;'/DX l;m.AlmrlCBlmrl1/Pm l.cos/.A0 lmsinm'CB0 lmcosm'/; (15.41) with lrequired to be an integer to avoid a solution that diverges in the polar directions. Here we consider solutions with no azimuthal dependence (i.e., with mD0), soEq. (15.41) reduces to .r;/D1X lD0.alrlCblrl1/Pl.cos/: (15.42) Often our problem is further restricted to a region either within or external to a boundary sphere, and if the problem is such that must remain finite, the solution will have one of the two following forms: .r;/D1X lD0alrlPl.cos/ . rr0/; (15.43) .r;/D1X lD0alrl1Pl.cos/ . rr0/: (15.44) Note that this simplification is not always appropriate; see Example 15.2.2. Sometimes the coefficients ( al) are determined from the boundary conditions of a problem rather than from the expansion of a known function. See the examples to follow. ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 727 Example 15.2.1 EARTH’S GRAVITATIONAL FIELD An example of a Legendre series is provided by the description of the Earth’s gravitational potential Uat points exterior to the Earth’s surface. Because gravitation is an inverse- square force, its potential in mass-free regions satisfies the Laplace equation, and therefore (if we neglect azimuthal effects, i.e., those dependent on longitude) it has the form given inEq. (15.44). To specialize to the current example, we define Rto be the Earth’s radius at the equator, and take as the expansion variable the dimensionless quantity R=r. In terms of the total mass of the Earth Mand the gravitational constant G, we have RD6378:10:1km; G M RD62:4940:001 km2=s2; and we write U.r;/DG M R" R r1X lD2alR rlC1 Pl.cos/# : (15.45) The leading term of this expansion describes the result that would be obtained if the Earth were spherically symmetric; the higher terms describe distortions. The P1term is absent because the origin from which ris measured is the Earth’s center of mass. Artificial satellite motions have shown that a2D.1;082;63511/109; a3D.2; 5317/109; a4D.1; 60012/109: This is the famous pear-shaped deformation of the Earth. Other coefficients have been computed through a20. More recent satellite data permit a determination of the longitudinal dependence of the Earth’s gravitational field. Such dependence may be described by a Laplace series (see Section 15.5).  Example 15.2.2 SPHERE IN A UNIFORM FIELD Another illustration of the use of a Legendre series is provided by the problem of a neutral conducting sphere (radius r0) placed in a (previously) uniform electric field of magnitude E0(Fig. 15.2). The problem is to find the new, perturbed electrostatic potential that satisfies Laplace’s equation, r2 D0: We select spherical polar coordinates with origin at the center of the conducting sphere and the polar ( z) axis oriented parallel to the original uniform field, a choice that will simplify ArfKen_Ch15-9780123846549.tex 728 Chapter 15 Legendre Functions E V=0z FIGURE 15.2 Conducting sphere in a uniform field. the application of the boundary condition at the surface of the conductor. Separating vari- ables, we note that because we require a solution to Laplace’s equation, the potential for rr0will be of the form of Eq. (15.42). Our solution will be independent of 'because of the axial symmetry of the problem. Because the insertion of the conducting sphere will have an effect that is local, the asymptotic behavior of must be of the form .r!1/DE0zDE0rcosDE0r P1.cos/; (15.46) equivalent to anD0; n>1; a1DE0: (15.47) Note that if an6D0for any n>1, that term would dominate at large rand the boundary condition, Eq. (15.46), could not be satisfied. In addition, the neutrality of the conducting sphere requires that not contain a contribution proportional to 1=r, so we also must have b0D0. As a second boundary condition, the conducting sphere must be an equipotential, and without loss of generality we can set its potential to zero. Then, on the sphere rDr0we have .r0;/Da0C b1 r2 0E0r0! P1.cos/C1X nD2bnPn.cos/ rnC1 0D0: (15.48) In order that Eq. (15.48) may hold for all values of , we set a0D0; b1DE0r3 0bnD0; n2: (15.49) ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 729 The electrostatic potential (outside the sphere) is then completely determined: .r;/DE0r P1.cos/CE0r3 0 r2P1.cos/ DE0r P1.cos/ 1r3 0 r3! DE0z 1r3 0 r3! : (15.50) In Section 9.5 we showed that Laplace’s equation with Dirichlet boundary conditions on a closed boundary (parts of which may be at infinity) had a unique solution. Since we have now found a solution to our current problem, it must (apart from an additive constant) be the only solution. It may further be shown that there is an induced surface charge density D" 0@ @r rDr0D3"0E0cos (15.51) on the surface of the sphere and an induced electric dipole moment of magnitude PD4r3 0"0E0: (15.52) See Exercise 15.2.11.  Example 15.2.3 ELECTROSTATIC POTENTIAL FOR A RING OF CHARGE As a further example, consider the electrostatic potential produced by a thin conducting ring of radius aplaced symmetrically in the equatorial plane of a spherical polar coordinate system and carrying a total electric charge q(Fig. 15.3). Again we rely on the fact that the potential satisfies Laplace’s equation. Separating the variables and recognizing that a solution for the region r>amust go to zero as r!1 , we use the form given by (r, θ) y xqθ r az FIGURE 15.3 Charged, conducting ring. ArfKen_Ch15-9780123846549.tex 730 Chapter 15 Legendre Functions Eq. (15.44), obtaining .r;/D1X nD0cnan rnC1Pn.cos/; r>a: (15.53) There is no'(azimuthal) dependence because of the cylindrical symmetry of the system. Note also that by including an explicit factor anwe cause all the coefficients cnto have the same dimensionality; this choice simply modifies the definition of cnand was, of course, not required. Our problem is to determine the coefficients cnin Eq. (15.53). This may be done by evaluating .r;/atD0,rDz, and comparing with an independent calculation of the potential from Coulomb’s law. In effect, we are using a boundary condition along the z-axis. From Coulomb’s law (using the fact that all the charge is equidistant from any point on the zaxis), .z;0/Dq 4" 01 .z2Ca2/1=2Dq 4" 0z1X sD01=2 sa2 z2s Dq 4" 0z1X sD0.1/s.2s1/WW .2s/WWa z2s ;z>a; (15.54) where we have evaluated the binomial coefficient using Eq. (1.74). Now, evaluating .z;0/from Eq. (15.53), remembering that Pn.1/D1for all n, we have .z;0/D1X nD0cnan znC1: (15.55) Since the power series expansion in zis unique, we may equate the coefficients of corre- sponding powers of zfrom Eqs. (15.54) and(15.55), reaching the conclusion that cnD0 fornodd, while for neven and equal to 2s, c2sDq 4" 0z.1/s.2s1/WW .2s/WW; (15.56) and our electrostatic potential .r;/is given by .r;/Dq 4" 0r1X sD0.1/s.2s1/WW .2s/WWa r2s P2s.cos/; r>a: (15.57) The magnetic analog of this problem appears in Example 15.4.2.  Exercises 15.2.1 Using a Rodrigues formula, show that the Pn.x/are orthogonal and that 1Z 1TPn.x/U2dxD2 2nC1: Hint. Integrate by parts. ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 731 15.2.2 You have constructed a set of orthogonal functions by the Gram-Schmidt process (Section 5.2), taking un.x/Dxn,nD0;1;2;::: , in increasing order with w.x/D1 and an interval1x1. Prove that the nth such function constructed in this way is proportional to Pn.x/. Hint. Use mathematical induction (Section 1.4). 15.2.3 Expand the Dirac delta function .x/in a series of Legendre polynomials using the interval1x1. 15.2.4 Verify the Dirac delta function expansions .1x/D1X nD02nC1 2Pn.x/; .1Cx/D1X nD0.1/n2nC1 2Pn.x/: These expressions appear in a resolution of the Rayleigh plane wave expansion (Exercise 15.2.24) into incoming and outgoing spherical waves. Note. Assume that the entire Dirac delta function is covered when integrating over T1; 1U. 15.2.5 Neutrons (mass 1) are being scattered by a nucleus of mass A.A>1). In the center- of-mass system the scattering is isotropic. Then, in the laboratory system the average of the cosine of the angle of deflection of the neutron is hcos iD1 2Z 0AcosC1 .A2C2AcosC1/1=2sind: Show, by expansion of the denominator, that hcos iD2=.3A/. 15.2.6 A particular function f.x/defined over the interval T1; 1Uis expanded in a Legendre series over this same interval. Show that the expansion is unique. 15.2.7 A function f.x/is expanded in a Legendre series f.x/DP1 nD0anPn.x/. Show that 1Z 1Tf.x/U2dxD1X nD02a2 n 2nC1: This is a statement that the Legendre polynomials form a complete set. 15.2.8 (a) For f.x/DC1; 0<x<1; 1;1<x<0; ArfKen_Ch15-9780123846549.tex 732 Chapter 15 Legendre Functions show that 1Z 1h f.x/i2 dxD21X nD0.4nC3/.2n1/WW .2nC2/WW2 : (b) By testing the series, prove that it is convergent. (c) The value of the integral in part (a) is 2. Check the rate at which the series con- verges by summing its first 10 terms. 15.2.9 Prove that 1Z 1x.1x2/P0 nP0 mdxD2n.n21/ 4n21m;n1C2n.nC2/.nC1/ .2nC1/.2nC3/m;nC1: 15.2.10 The coincidence counting rate, W./, in a gamma-gamma angular correlation experi- ment has the form W./D1X nD0a2nP2n.cos/: Show that data in the range =2can, in principle, define the function W./ (and permit a determination of the coefficients a2n). This means that although data in the range 0<=2 may be useful as a check, they are not essential. 15.2.11 A conducting sphere of radius r0is placed in an initially uniform electric field, E0. Show the following: (a) The induced surface charge density is D3"0E0cos, (b) The induced electric dipole moment is PD4r3 0"0E0. Note. The induced electric dipole moment can be calculated either from the surface charge [part (a)] or by noting that the final electric field Eis the result of superimposing a dipole field on the original uniform field. 15.2.12 Obtain as a Legendre expansion the electrostatic potential of the circular ring of Example 15.2.3, for points .r;/with r<a. 15.2.13 Calculate the electric field produced by the charged conducting ring of Example 15.2.3 for (a) r>a, (b) r<a. 15.2.14 As an extension of Example 15.2.3, find the potential .r;/produced by a charged conducting disk, Fig. 15.4, for r>a, where ais the radius of the disk. The charge density(on each side of the disk) is ./Dq 4a.a22/1=2; 2Dx2Cy2: ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 733 yz xa FIGURE 15.4 Charged conducting disk. Hint. The definite integral you get can be evaluated as a beta function, Section 13.3. For more details see section 5.03 of Smythe in Additional Readings. ANS: . r;/Dq 4" 0r1X lD0.1/l1 2lC1a r2l P2l.cos/: 15.2.15 The hemisphere defined by rDa,0 <=2 , has an electrostatic potential CV0. The hemisphere rDa,=2<has an electrostatic potential V0. Show that the potential at interior points is VDV01X nD04nC3 2nC2r a2nC1 P2n.0/P2nC1.cos/ DV01X nD0.1/n.4nC3/.2n1/WW .2nC2/WWr a2nC1 P2nC1.cos/: Hint. You need Exercise 15.1.12. 15.2.16 A conducting sphere of radius ais divided into two electrically separate hemispheres by a thin insulating barrier at its equator. The top hemisphere is maintained at a potential V0, and the bottom hemisphere at V0. (a) Show that the electrostatic potential exterior to the two hemispheres is V.r;/DV01X sD0.1/s.4sC3/.2s1/WW .2sC2/WWa r2sC2 P2sC1.cos/: (b) Calculate the electric charge density on the outside surface. Note that your series diverges at cosD1; as you expect from the infinite capacitance of this system (zero thickness for the insulating barrier). ANS: .b/D"0EnD" 0@V @r rDa D"0V01X sD0.1/s.4sC3/.2s1/WW .2s/WWP2sC1.cos/: ArfKen_Ch15-9780123846549.tex 734 Chapter 15 Legendre Functions 15.2.17 By writing's.x/Dp.2sC1/=2 Ps.x/, a Legendre polynomial is renormalized to unity. Explain how j'sih'sjacts as a projection operator. In particular, show that if jfiDP na0 nj'ni, then j'sih'sjfiDa0 sj'si: 15.2.18 Expand x8as a Legendre series. Determine the Legendre coefficients from Eq. (15.40), amD2mC1 21Z 1x8Pm.x/dx: Check your values against AMS-55, Table 22.9. (For the complete reference, see Addi- tional Readings.) This illustrates the expansion of a simple function f.x/. Hint. Gaussian quadrature can be used to evaluate the integral. 15.2.19 Calculate and tabulate the electrostatic potential created by a ring of charge, Example 15.2.3, for r=aD1:5.0:5/5:0 andD0.15/90. Carry terms through P22.cos/. Note. The convergence of your series will be slow for r=aD1:5. Truncating the series atP22limits you to about four-significant-figure accuracy. Check value. Forr=aD2:5andD60, D0:40272. q=4" 0r/. 15.2.20 Calculate and tabulate the electrostatic potential created by a charged disk (Exercise 15.2.14), for r=aD1:5.0:5/5:0 andD0.15/90. Carry terms through P22.cos/. Check value. Forr=aD2:0andD15, D0:46638. q=4" 0r/. 15.2.21 Calculate the first five (nonvanishing) coefficients in the Legendre series expansion off.x/D1jxj, evaluating the coefficients in the series by numerical integration. Actually these coefficients can be obtained in closed form. Compare your coefficients with those listed in Exercise 18.4.26. ANS. a0D0:5000 ,a2D0:6250 ,a4D0:1875 ,a6D0:1016 ,a8D0:0664 . 15.2.22 Calculate and tabulate the exterior electrostatic potential created by the two charged hemispheres of Exercise 15.2.16, for r=aD1:5.0:5/5:0 andD0.15/90. Carry terms through P23.cos/. Check value. Forr=aD2:0andD45,VD0:27066 V0. 15.2.23 (a) Given f.x/D2:0,jxj<0:5and f.x/D0,0:5<jxj<1:0, expand f.x/in a Legendre series and calculate the coefficients anthrough a80(analytically). (b) EvaluateP80 nD0anPn.x/forxD0:400.0:005/0:600 . Plot your results. Note. This illustrates the Gibbs phenomenon of Section 19.3 and the danger of trying to calculate with a series expansion in the vicinity of a discontinuity. ArfKen_Ch15-9780123846549.tex 15.2 Orthogonality 735 15.2.24 A plane wave may be expanded in a series of spherical waves by the Rayleigh equation, eikrcos D1X nD0anjn.kr/Pn.cos /: Show that anDin.2nC1/. Hint. 1. Use the orthogonality of the Pnto solve for anjn.kr/. 2. Differentiate ntimes with respect to .kr/and set rD0to eliminate the r-dependence. 3. Evaluate the remaining integral by Exercise 15.1.15. Note. This problem may also be treated by noting that both sides of the equation satisfy the Helmholtz equation. The equality can be established by showing that the solutions have the same behavior at the origin and also behave alike at large distances. 15.2.25 Verify the Rayleigh equation of Exercise 15.2.24 by starting with the following steps: (a) Differentiate with respect to .kr/to establish X nanj0 n.kr/Pn.cos /DiX nanjn.kr/cos Pn.cos /: (b) Use a recurrence relation to replace cos Pn.cos /by a linear combination of Pn1andPnC1. (c) Use a recurrence relation to replace j0 nby a linear combination of jn1andjnC1. 15.2.26 From Exercise 15.2.24 show that jn.kr/D1 2in1Z 1eikrPn./d: This means that (apart from a constant factor) the spherical Bessel function jn.kr/is an integral transform of the Legendre polynomial Pn./. 15.2.27 Rewriting the formula of Exercise 15.2.26 as jn.z/D1 2.i/nZ 0eizcosPn.cos/sind;nD0;1;2;:::; verify it by transforming the right-hand side into zn 2nC1nWZ 0cos.zcos/sin2nC1d and using Exercise 14.7.9. ArfKen_Ch15-9780123846549.tex 736 Chapter 15 Legendre Functions 15.3 P HYSICAL INTERPRETATION OF GENERATING FUNCTION The generating function for the Legendre polynomials has an interesting and important interpretation. If we introduce spherical polar coordinates .r;;'/ and place a charge q at the point aon the positive zaxis (see Fig. 15.5), the potential at a point .r;/(it is independent of ') can be calculated, using the law of cosines, as .r;/Dq 4" 01 r1Dq 4" 0.r2Ca22arcos/1=2: (15.58) The expression in Eq. (15.58) is essentially that appearing in the generating function; to identify the correspondence we rewrite that equation as .r;/Dq 4" 0r 12a rcosCa2 r21=2 Dq 4" 0rg cos;a r (15.59) Dq 4" 0r1X nD0Pn.cos/a rn ; (15.60) where we reached Eq. (15.60) by inserting the generating-function expansion. The series in Eq. (15.60) only converges for r>a, with a rate of convergence that improves as r=aincreases. If, on the other hand, we desire an expression for .r;/when r<a, we can perform a different rearrangement of Eq. (15.58), to .r;/Dq 4" 0a 12r acosCr2 a21=2 ; (15.61) which we again recognize as the generating-function expansion, but this time with the result .r;/Dq 4" 0a1X nD0Pn.cos/r an ; (15.62) valid when r<a. r1qq 4πε0r1 z=ar θϕ = z FIGURE 15.5 Electrostatic potential, charge qdisplaced from origin. ArfKen_Ch15-9780123846549.tex 15.3 Physical Interpretation of Generating Function 737 Expansion of 1=jr 1r2j Equations (15.60) and(15.62) describe the interaction of a charge qat position aDaOez with a unit charge at position r. Dropping the factors needed for an electrostatics calcula- tion, these equations yield formulas for 1=jraj. The fact that ais aligned with the z-axis is actually of no importance for the computation of 1=jraj; the relevant quantities are r, a, and the angle between randa. Thus, we may rewrite either Eq. (15.60) or(15.62) in a more neutral notation, to give the value of 1=jr 1r2jin terms of the magnitudes r1,r2and the angle between r1andr2, which we now call . If we define r>andr<to be respec- tively the larger and the smaller of r1andr2,Eqs. (15.60) and(15.62) can be combined into the single equation 1 jr1r2jD1 r>1X nD0r< r>n Pn.cos/; (15.63) which will converge everywhere except when r1Dr2. Electric Multipoles Returning to Eq. (15.60) and restricting consideration to r>a, we may note that its initial term (with nD0) gives the potential we would get if the charge qwere at the origin, and that further terms must describe corrections arising from the actual position of the charge. One way to obtain further understanding of the second and later terms in the expansion is to consider what would happen if we added a second charge, q, atzDa, as shown in Fig. 15.6. The potential due to the second charge will be given by an expression similar to that in Eq. (15.58), except that the signs of qandcosmust be reversed (the angle opposite r2in the figure is ). We now have Dq 4" 01 r11 r2 r1r2 z = a z = −a−q qr θzq 4πε0ϕ =1 r11 r2− FIGURE 15.6 Electric dipole. ArfKen_Ch15-9780123846549.tex 738 Chapter 15 Legendre Functions Dq 4" 0r" 12a rcosCa2 r21=2  1C2a rcosCa2 r21=2# Dq 4" 0r"1X nD0Pn.cos/a rn 1X nD0Pn.cos/ a rn# : (15.64) If we combine the two summations in Eq. (15.64), alternate terms cancel, and we get D2q 4" 0ra rP1.cos/Ca3 r3P3.cos/C : (15.65) This configuration of charges is called an electric dipole, and we note that its leading dependence on rgoes as r2. The strength of the dipole (called the dipole moment) can be identified as 2qa, equal to the magnitude of each charge multiplied by their separation (2a). If we let a!0while keeping the product 2qaconstant at a value , all but the first term becomes negligible, and we have D 4" 0P1.cos/ r2; (15.66) the potential of a point dipole of dipole moment , located at the origin of the coordinate system (at rD0). Note that because we have limited the discussion to situations of cylin- drical symmetry, our dipole is oriented in the polar direction; more general orientations can be considered after we have developed formulas for solutions of the associated Legendre equation (cases where the parameter min Eq. (15.4) is nonzero). We can extend the above analysis by combining a pair of dipoles of opposite orienta- tion, for example, in the configuration shown in Fig. 15.7, thereby causing cancellation of their leading terms, leaving a potential whose leading contribution will be proportional to r3P2.cos/. A charge configuration of this sort is called an electric quadrupole, and theP2term of the generating function expansion can be identified as the contribution of apoint quadrupole, also located at rD0. Further extensions, to 2n-poles, with contri- butions proportional to Pn.cos/=rnC1, permit us to identify each term of the generating expansion with the potential of a point multipole. We thus have a multipole expansion. Again we observe that because we have limited discussion to situations with cylindrical symmetry our multipoles are presently required to be linear; that restriction will be elimi- nated when this topic is revisited in Chapter 16. We look next at more general charge distributions, for simplicity limiting consideration to charges qiplaced at respective positions aion the polar axis of our coordinate system. z = a z = −aqq −2q z FIGURE 15.7 Linear electric quadrupole. ArfKen_Ch15-9780123846549.tex 15.3 Physical Interpretation of Generating Function 739 Adding together the generating-function expansions of the individual charges, our com- bined expansion takes the form D1 4" 0r"X iqiCX iqiai rP1.cos/CX iqia2 i r2P2.cos/C# D1 4" 0rh 0C1 rP1.cos/C2 r2P2.cos/Ci ; (15.67) where theiare called the multipole moments of the charge distribution; 0is the 20-pole, or monopole moment, with a value equal to the total net charge of the distribution; 1is the 21-pole, or dipole moment, equal toP iqiai;2is the 22-pole, or quadrupole moment, given asP iqia2 i, etc. Our general (linear) multipole expansion will converge for values of rthat are larger than all the aivalues of the individual charges. Put another way, the expansion will converge at points further from the coordinate origin than all parts of the charge distribution. We next ask: What happens if we move the origin of our coordinate system? Or, equiv- alently, consider replacing rbyjrrpj. For r>rp, the binomial expansion of 1=jrrpjn will have the generic form 1 jrrpjnD1 rnCCrp rnC1C; with the result that only the leading nonzero term of Eq. (15.67) will be unaffected by the change of expansion center. Translated into everyday language, this means that the lowest nonzero moment of the expansion will be independent of the choice of origin, but all higher moments will change when the expansion center is moved. Specifically, the total net charge (monopole moment) will always be independent of the choice of expansion center. The dipole moment will be independent of the expansion point only when the net charge is zero; the quadrupole moment will have such independence only if both the net charge and dipole moments vanish, etc. We close this section with three observations. First, while we have illustrated our discussion with discrete arrays of point charges, we could have reached the same conclusions using continuous charge distributions, with the result that the summations over charges would become integrals over the charge density. Second, if we remove our restriction to linear arrays, our expansion would involve components of the multipole moments in different directions. In three-dimensional space, the dipole moment would have three components: ageneralizes to .ax;ay;az/, while the higher-order multipoles will have larger numbers of components ( a2! axax;axay;:::). The details of that analysis will be taken up when the necessary back- ground is in place. Third, the multipole expansion is not restricted to electrical phenomena, but applies anywhere we have an inverse-square force. For example, planetary configurations are described in terms of mass multipoles. And gravitational radiation depends on the time behavior of mass quadrupoles. ArfKen_Ch15-9780123846549.tex 740 Chapter 15 Legendre Functions Exercises 15.3.1 Develop the electrostatic potential for the array of charges shown in Fig. 15.7. This is a linear electric quadrupole. 15.3.2 Calculate the electrostatic potential of the array of charges shown in Fig. 15.8. Here is an example of two equal but oppositely directed quadrupoles. The quadrupole contri- butions cancel. The octopole terms do not cancel. 15.3.3 Show that the electrostatic potential produced by a charge qatzDaforr<ais '.r/Dq 4" 0a1X nD0r an Pn.cos/: 15.3.4 Using EDr', determine the components of the electric field corresponding to the (pure) electric dipole potential, '.r/D2aq P 1.cos/ 4" 0r2: Here it is assumed that ra. ANS. ErDC4aqcos 4" 0r3,EDC2aqsin 4" 0r3,E'D0. 15.3.5 Operating in spherical polar coordinates, show that @ @zPl.cos/ rlC1 D. lC1/PlC1.cos/ rlC2: This is the key step in the mathematical argument that the derivative of one multipole leads to the next higher multipole. Hint. Compare with Exercise 3.10.28. 15.3.6 A point electric dipole of strength p.1/is placed at zDa; a second point electric dipole of equal but opposite strength is at the origin. Keeping the product p.1/aconstant, let a!0. Show that this results in a point electric quadrupole. Hint. Exercise 15.3.5 (when proved) will be helpful. 15.3.7 A point electric octupole may be constructed by placing a point electric quadrupole (pole strength p.2/in the z-direction) at zDaand an equal but opposite point elec- tric quadrupole at zD0and then letting a!0, subject to p.2/aDconstant. Find the electrostatic potential corresponding to a point electric octupole. Show from the con- struction of the point electric octupole that the corresponding potential may be obtained by differentiating the point quadrupole potential. q −q −2q +2q z = −2a −a a2 az FIGURE 15.8 Linear electric octopole. ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 741 qq ′ a′az FIGURE 15.9 Image charges for Exercise 15.3.8. 15.3.8 A point charge qis in the interior of a hollow conducting sphere of radius r0. The charge qis displaced a distance afrom the center of the sphere. If the conducting sphere is grounded, show that the potential in the interior produced by qand the dis- tributed induced charge is the same as that produced by qand its image charge q0. The image charge is at a distance a0Dr2 0=afrom the center, collinear with qand the origin (Fig. 15.9). Hint. Calculate the electrostatic potential for a<r0<a0. Show that the potential vani- shes for rDr0if we take q0Dqr0=a. 15.4 A SSOCIATED LEGENDRE EQUATION We need to extend our analysis to the associated Legendre equation because it is important to be able to remove the restriction to azimuthal symmetry that pervaded the discussion of the previous sections of this chapter. We therefore return to Eq. (15.4), which, before determining what its eigenvalue should be, assumed the form .1x2/P00.x/2x P0.x/C m2 1x2 P.x/D0: (15.68) Trial and error (or great insight) suggests that the troublesome factor 1x2in the denominator of this equation can be eliminated by making a substitution of the form PD.1x2/pP, and further experimentation shows that a suitable choice for the exponent pism=2. By straightforward differentiation, we find PD.1x2/m=2P; (15.69) P0D.1x2/m=2P0mx.1x2/m=21P; (15.70) P00D.1x2/m=2P002mx.1x2/m=21P0 Ch m.1x2/m=21C.m22m/x2.1x2/m=22i P: (15.71) Substitution of Eqs. (15.69)–(15.71) intoEq. (15.68), we obtain an equation that is poten- tially easier to solve, namely, .1x2/P002x.mC1/P0Ch m.mC1/i PD0: (15.72) We continue by seeking to solve Eq. (15.72) by the method of Frobenius, assuming a solution in the series formP jajxkCj. The indicial equation for this ODE has solutions ArfKen_Ch15-9780123846549.tex 742 Chapter 15 Legendre Functions kD0andkD1. For kD0, substitution into the series solution leads to the recurrence formula ajC2Dajj2C.2mC1/jCm.mC1/ .jC1/.jC2/ : (15.73) Just as for the original Legendre equation, we need solutions P.cos/that are nonsingular for the range1cosC1 , but the recurrence formula leads to a power series that in general is divergent at 1.2 To avoid the divergence, we must cause the numerator of the fraction in Eq. (15.73) to become zero for some nonnegative even integer j, thereby causing Pto be a polynomial. By direct substitution into Eq. (15.73), we can verify that a zero numerator is obtained for jDlmwhenis assigned the value l.lC1/, a condition that can only be met if lis an integer at least as large as mand of the same parity. Further analysis for the other indicial equation solution, kD1, extends our present result to values of lthat are larger than mand of opposite parity. Summarizing our results to this point, we have found that the regular solutions to the associated Legendre equation depend on integer indices landm. Letting Pm l, called an associated Legendre function, denote such a solution (note that the superscript misnot an exponent), we define Pm l.x/D.1x2/m=2Pm l.x/; (15.74) wherePm lis a polynomial of degree lm(consistent with our earlier observation that lmust be at least as large as m), and with an explicit form and scale that we will now address. A convenient explicit formula for Pm lcan be obtained by repeated differentiation of the regular Legendre equation. Admittedly, this strategy would have been difficult to devise without prior knowledge of the solution, but there are certain advantages to using the experience of those who have gone before. So, without apology, we apply Leibniz’s for- mula for the mth derivative of a product (proved in Exercise 1.4.2), dm dxmh A.x/B.x/i DmX sD0m sdmsA.x/ dxmsdsB.x/ dxs; (15.75) to the Legendre equation, .1x2/P00 l2x P0 lCl.lC1/PlD0; reaching .1x2/u002x.mC1/u0Ch l.lC1/m.mC1/i uD0; (15.76) where udm dxmPl.x/: (15.77) 2The solution to the associated Legendre equation is .1x2/m=2P.x/, suggesting the possibility that the .1x2/m=2factor might compensate the divergence in P.x/, yielding a convergent limit. It can be shown that such a compensation does not occur. ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 743 Comparing Eq. (15.76) with Eq. (15.72), we see that when Dl.lC1/they are identical, meaning that the polynomial solutions PofEq. (15.72) for given lcan be identified with the corresponding u. Specifically, Pm lD.1/mdm dxmPl.x/; (15.78) where the factor .1/mis inserted to maintain agreement with AMS-55 (see Additional Readings), which has become the most widely accepted notational standard.3 We can now write a complete, explicit form for the associated Legendre functions: Pm l.x/D.1/m.1x2/m=2dm dxmPl.x/: (15.79) Since the Pm lwith mD0are just the original Legendre functions, it is customary to omit the upper index when it is zero, so, for example, P0 lPl. Note that the condition on landmcan be stated in two ways: (1) For each m, there are an infinite number of acceptable solutions to the associated Legendre ODE with lvalues ranging from mto infinity, or (2) For each l, there are acceptable solutions with mvalues ranging from lD0tolDm. Because menters the associated Legendre equation only as m2, we have up to this point tacitly considered only values m0. However, if we insert the Rodrigues formula for Pl into Eq. (15.73), we get the formula Pm l.x/D.1/m 2llW.1x2/m=2dlCm dxlCm.x21/l; (15.80) which gives results for mthat do not appear similar to those for Cm. However, it can be shown that if we apply Eq. (15.75) for mvalues between zero and l, we get Pm l.x/D.1/m.lm/W .lCm/WPm l.x/: (15.81) Equation (15.81) shows that Pm land Pm lare proportional; its proof is the topic of Exercise 15.4.3. The main reason for discussing both is that recurrence formulas we will develop for Pm lwith contiguous values of mwill give results for m<0that can best be understood if we remember the relative scaling of Pm landPm l. Associated Legendre Polynomials For further development of properties of the Pm l, it is useful to develop a generating func- tion for the polynomials Pm l.x/, which we can do by differentiating the Legendre generat- ing function with respect to x. The result is gm.x;t/.1/m.2m1/WW .12xtCt2/mC1=2D1X sD0Pm sCm.x/ts: (15.82) 3However, we note that the popular text, Jackson’s Electrodynamics (see Additional Readings), does not include this phase factor. The factor is introduced to cause the definition of spherical harmonics (Section 15.5) to have the usual phase convention. ArfKen_Ch15-9780123846549.tex 744 Chapter 15 Legendre Functions The factors tthat result from differentiating the generating function have been used to change the powers of tthat multiply the Pon the right-hand side. If we now differentiate Eq. (15.82) with respect to t, we obtain initially .12txCt2/@gm @tD.2mC1/.xt/gm.x;t/; which we can use together with Eq. (15.82) in a now-familiar way to obtain the recurrence formula, .sC1/Pm sCmC1.x/.2mC1C2s/xPm sCm.x/C.sC2m/Pm sCm1D0: (15.83) Making the substitution lDsCm, we bring Eq. (15.83) to the more useful form, .lmC1/Pm lC1.2lC1/xPm lC.lCm/Pm l1D0: (15.84) FormD0this relation agrees with Eq. (15.18). From the form of gm.x;t/, it is also clear that .12xtCt2/gmC1.x;t/D.2mC1/gm.x;t/: (15.85) From Eqs. (15.85) and(15.82) we may extract the recursion formula PmC1 sCmC1.x/2xPmC1 sCm.x/CPmC1 sCm1.x/D.2mC1/Pm sCm.x/; which relates the associated Legendre polynomials with upper index mC1to those with upper index m. Again we may simplify by making the substitution lDsCm: PmC1 lC1.x/2xPmC1 l.x/CPmC1 l1.x/D.2mC1/Pm l.x/: (15.86) Associated Legendre Functions The recurrence relations for the associated Legendre polynomials or alternatively, differ- entiation of formulas for the original Legendre polynomials, enable the construction of recurrence formulas for the associated Legendre functions. The number of such formulas is extensive because these functions have two indices, and there exists a wide variety of formulas with different index combinations. Results of importance include the following: PmC1 l.x/C2mx .1x2/1=2Pm l.x/C.lCm/.lmC1/Pm1 l.x/D0; (15.87) .2lC1/x Pm l.x/D.lCm/Pm l1.x/C.lmC1/Pm lC1.x/; (15.88) .2lC1/.1x2/1=2Pm l.x/DPmC1 l1.x/PmC1 lC1.x/ (15.89) D.lmC1/.lmC2/Pm1 lC1.x/ .lCm/.lCm1/Pm1 l1.x/; (15.90) .1x2/1=2 Pm l.x/0 D1 2.lCm/.lmC1/Pm1 l.x/1 2PmC1 l.x/; (15.91) D.lCm/.lmC1/Pm1 l.x/Cmx .1x2/1=2Pm l.x/: (15.92) ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 745 Table 15.3 Associated Legendre Functions P1 1.x/D.1x2/1=2Dsin P1 2.x/D3 x.1x2/1=2D3 cossin P2 2.x/D3.1x2/D3 sin2 P1 3.x/D3 2.5x21/.1x2/1=2D3 2.5 cos21/sin P2 3.x/D15x.1x2/D15 cossin2 P3 3.x/D15.1x2/3=2D15 sin3 P1 4.x/D5 2.7x33x/.1x2/1=2D5 2.7 cos33 cos/sin P2 4.x/D15 2.7x21/.1x2/D15 2.7 cos21/sin2 P3 4.x/D105 x.1x2/3=2D105 cossin3 P4 4.x/D105.1x2/2D105 sin4 It is obvious that, using Eq. (15.90), all the Pm lwith m>0can be generated from those with mD0(the Legendre polynomials), and that these, in turn, can be built recursively from P0.x/D1andP1.x/Dx. In this fashion (or in other ways as suggested below), we can build a table of associated Legendre functions, the first members of which are listed in Table 15.3. The table shows the Pm l.x/both as functions of xand as functions of , where xDcos. It is often easier to use recurrence formulas other than Eq. (15.90) to obtain the Pm l, keeping in mind that when a formula contains Pm m1form>0, that quantity can be set to zero. It is also easy to obtain explicit formulas for certain values of landmwhich can then be alternate starting points for recursion. See the example that follows. Example 15.4.1 RECURRENCE STARTING FROM Pm m The associated Legendre function Pm m.x/is easily evaluated: Pm m.x/D.1/m 2mmW.1x2/m=2d2m dx2m.x21/mD.1/m 2mmW.2m/W.1x2/m=2 D.1/m.2m1/WW.1x2/m=2: (15.93) We can now use Eq. (15.88) with lDmto obtain Pm mC1, dropping the term containing Pm m1because it is zero. We get Pm mC1.x/D.2mC1/x Pm m.x/D.1/m.2mC1/WWx.1x2/m=2: (15.94) Further increases in lcan now be obtained by straightforward application of Eq. (15.88). Illustrating for a series of Pm lwith mD2:P2 2.x/D.1/2.3WW/.1x2/D3.1x2/, in agreement with the table value. P2 3can be computed from Eq. (15.94) asP2 3.x/D .1/2.5WW/x.1x2/, which simplifies to the tabulated result. Finally, P2 4is obtained from ArfKen_Ch15-9780123846549.tex 746 Chapter 15 Legendre Functions the following case of Eq. (15.88): 7x P2 3.x/D5P2 2.x/C2P2 4.x/; the solution of which for P2 4.x/is again in agreement with the tabulated value.  Parity and Special Values We have already established that Plhas even parity if lis even and odd parity if lis odd. Since we can form Pm lby differentiating Plmtimes, with each differentiation changing the parity, and thereafter multiplying by .1x2/m=2, which has even parity, Pm lmust have a parity that depends on lCm, namely, Pm l.x/D.1/lCmPm l.x/: (15.95) We occasionally encounter a need for the value of Pm l.x/atxD1 or at xD0. At xD1 the result is simple: The factor .1x2/m=2causes Pm l.1/ to vanish unless mD 0, in which case we recover the values Pl.1/D1,Pl.1/D.1/l. AtxD0, the value ofPm ldepends on whether lCmis even or odd. The result, proof of which is left to Exercises 15.4.4 and15.4.5, is Pm l.0/D8 < :.1/.lCm/=2.lCm1/WW .lm/WW;lCmeven; 0; lCmodd:(15.96) Orthogonality For each m, the Pm lof different lcan be proved orthogonal by identifying them as eigenfunctions of a Sturm-Liouville system. However, it is instructive to demonstrate the orthogonality explicitly, and to do so by a method that also yields their normalization. We start by writing the orthgonality integral, with the Pm lgiven by the Rodrigues for- mula in Eq. (15.80). For compactness and clarity, we introduce the abbreviated notation RDx21, thereby getting 1Z 1Pm p.x/Pm q.x/dxD.1/m 2pCqpWqW1Z 1RmdpCmRp dxpCmdqCmRq dxqCm dx: (15.97) We consider first the case p<q, for which we plan to prove the integral in Eq. (15.97) vanishes. We proceed by carrying out repeated integrations by parts, in which we differ- entiate uDRmdpCmRp dxpCm (15.98) pCmC1times while integrating a like number of times the remainder of the integrand, dvDdqCmRq dxqCm dx: (15.99) ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 747 For each of these pCmC1qCmpartial integrations the integrated ( uv) terms will vanish because there will be at least one factor Rthat is not differentiated and will therefore vanish at xD1 . After the repeated differentiation, we will have dpCmC1 dxpCmC1uDdpCmC1 dxpCmC1 RmdpCmRp dxpCm ; (15.100) in which a quantity whose largest power of xisx2pC2mcontains also a .2pC2mC1/-fold differentiation. There is no way these components can yield a nonzero result. Since both the integrated terms and the transformed integral vanish, we get an overall vanishing result, confirming the orthogonality. Note that the orthogonality is with unit weight, independent of the value of m. We now examine Eq. (15.97) forpDq, repeating the process we just carried out, but this time performing pCmpartial integrations. Again all the integrated terms vanish, but now there is a nonvanishing contribution from the repeated differentiation of u, see Eq. (15.98). Since the overall power of xis still x2pC2mand the total number of differ- entiations is also 2pC2m, the only contributing terms are those in which the factor Rm is differentiated 2mtimes and the factor Rpis differentiated 2ptimes. Thus, applying Leibniz’s formula, Eq. (15.75), to the pCm-fold differentiation of u, but keeping only the contributing term, we have dpCm dxpCm RmdpCmRp dxpCm DpCm 2md2mRm dx2md2pRp dx2p D.pCm/W .2m/W.pm/W.2m/W.2p/WD.pCm/W .pm/W.2p/W:(15.101) Inserting this result into the integration by parts, remembering that the transformed integration is accompanied by the sign factor .1/pCm, and recognizing that the repeated integration of dv, Eq. (15.99) with qDp, just yields Rp, we have, returning to Eq. (15.97), 1Z 1h Pm p.x/i2 dxD.1/2mCp 22ppWpW.pCm/W .pm/W.2p/W1Z 1Rpdx: (15.102) To complete the evaluation, we identify the integral of Rpas a beta function, with an evaluation given in Exercise 13.3.3 as 1Z 1RpdxD.1/p2.2p/WW .2pC1/WWD.1/p22pC1pWpW .2pC1/W: (15.103) Inserting this result, and combining with the previously established orthogonality relation, we have 1Z 1Pm p.x/Pm q.x/dxD2 2pC1.pCm/W .pm/Wpq: (15.104) ArfKen_Ch15-9780123846549.tex 748 Chapter 15 Legendre Functions Making the substitution xDcos, we obtain this formula in spherical polar coordinates: Z 0Pm p.cos/Pm q.cos/sindD2 2pC1.pCm/W .pm/Wpq: (15.105) Another way to look at the orthogonality of the associated Legendre functions is to rewrite Eq. (15.104) in terms of the associated Legendre polynomials Pm l. Invoking Eq. (15.74), Eq. (15.104) becomes 1Z 1Pm pPm q.1x2/mdxD2 2pC1.pCm/W .pm/Wpq; (15.106) showing that these polynomials are, for each m, orthogonal with the weight factor .1 x2/m. From that viewpoint, we can observe that each value of mcorresponds to a set of polynomials that are orthogonal with a different weight. However, since our main interest is in the functions that are in general notpolynomials but are solutions of the associated Legendre equation, it is usually more relevant to us to note that these functions, which include the factor .1x2/m=2, are orthogonal with unit weight. It is possible, but not particularly useful, to note that we can also have orthogonality of thePm lwith respect to the upper index when the lower index is held constant: 1Z 1Pm l.x/Pn l.x/.1x2/1dxD.lCm/W m.lm/Wmn: (15.107) This equation is not very useful because in spherical polar coordinates the boundary con- dition on the azimuthal coordinate 'causes there already to be orthogonality with respect tom, and we are not usually concerned with orthogonality of the Pm lwith respect to m. Example 15.4.2 CURRENT LOOP—MAGNETIC DIPOLE An important problem in which we encounter associated Legendre functions is in the mag- netic field of a circular current loop, a situation that may at first seem surprising since this problem has azimuthal symmetry. Our starting point is the formula relating a current element I dsto the vector potential Athat it produces (this is discussed in the chapter on Green’s functions, and also in texts such as Jackson’s Classical Electrodynamics; see Additional Readings). This formula is dA.r/D0 4I ds jrrsj; (15.108) where ris the point at which Ais to be evaluated and rsis the position of element dsof the current loop. We place our current loop, of radius a, in the equatorial plane of a spherical polar coordinate system, as shown in Fig. 15.10. Our task is to determine Aas a function of position, and therefrom to obtain the components of the magnetic induction field B. ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 749 yxz r dsrsa FIGURE 15.10 Circular current loop. It is in principle possible to figure out the geometry and integrate Eq. (15.108) for the present problem, but a more practical approach will be to determine from general consid- erations the functional form of an expansion describing the solution, and then to determine the coefficients in the expansion by requiring correct results for points of high symmetry, where the calculation is not too difficult. This is an approach similar to that employed in Example 15.2.3, where we first identified the functional form of an expansion giving the potential generated by a circular ring of charge, after which we found the coefficients in the expansion from the easily computed potential on the axis of the ring. From the form of Eq. (15.108) and the symmetry of the problem, we see immediately that for all r,Amust lie in a plane of constant z, and in fact it must be in the Oe'direction, with A'independent of ', i.e., ADA'.r;/Oe': (15.109) IfAhad a component other than A', it would have a nonzero divergence, as then Awould have a nonzero inward or outward flux, resulting in a singularity on the axis of the loop. Since everywhere except on the current loop itself there is no current, Maxwell’s equa- tion for the curl of Breduces to rBDr.rA/D0; and, since Ahas only a'component, it further reduces to rh rA'.r;/Oe'i D0: (15.110) The left-hand side of Eq. (15.110) was the subject of Example 3.10.4, and its evaluation was presented as Eq. (3.165). Setting that result to zero gives the equation that must be satisfied by A'.r;/: @2A' @r2C2 r@A' @rC1 r2sin@ @ sin@A' @ 1 r2sin2A'D0: (15.111) ArfKen_Ch15-9780123846549.tex 750 Chapter 15 Legendre Functions Equation (15.111) may now be solved by the method of separation of variables; setting A'.r;/DR.r/2./ , we have r2d2R dr2C2d R drl.lC1/RD0; (15.112) 1 sind d sind2 d Cl.lC1/22 sin2D0: (15.113) Because the second of these equations can be recognized as the associated Legendre equa- tion, in the form given as Eq. (15.2), we have set the separation constant to the value it must have, namely l.lC1/, with lintegral. The first equation is also familiar, with solutions for a given lbeing rlandrl1. The second equation has solutions P1 l.cos/, i.e., its specific form dictates that the associated Legendre functions which solve it must have upper index mD1. Since our main interest is in the pattern of Batrvalues larger than a, the radius of the current loop, we retain only the radial solution rl1, and write A'.r;/D1X lD1cla rlC1 P1 l.cos/: (15.114) When we obtain a more detailed solution, we will find that it converges only for r>a, so Eq. (15.114) and the value of Bderived therefrom will only be valid outside a sphere containing the current loop. If we were also interested in solving this problem for r<a, we would need to construct a series solution using only the powers rl. From Eq. (15.114) we can compute the components of B. Clearly, B'D0. And, using Eq. (3.159), we have Br.r;/DrA'Oe' rDcot rA'C1 r@A' @; (15.115) B.r;/DrA'Oe' D1 [email protected] A'/ @r: (15.116) To evaluate the derivative in Eq. (15.115), we need d P1 l.cos/ dDsind P1 l.cos/ dcosDl.lC1/Pl.cos/cotP1 l.cos/; (15.117) a special case of Eq. (15.92) with mD1andxDcos. It is now straightforward to insert the expansion for A'into Eqs. (15.115) and(15.116); because of Eq. (15.117) thecot term of Eq. (15.115) cancels, and we reach Br.r;/D1 r1X lD1l.lC1/cla rlC1 Pl.cos/; (15.118) B.r;/D1 r1X lD1l cla rlC1 P1 l.cos/: (15.119) To complete our analysis, we must determine the values of the cl, which we do by using the Biot-Savart law to calculate Brat points along the polar axis, where Bris synonymous ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 751 with Bz. SinceD0on the positive polar axis and Pl.cos/D1, Eq. (15.118) reduces to Br.z;0/D1 z1X lD1l.lC1/cla zlC1 Da2 z31X sD0.sC1/.sC2/csC1a zs :(15.120) The symmetry of the problem permits one more simplification; the value of Bzmust be the same atzas at z, from which we conclude that the coefficients c2,c4,. . . must all vanish, and we can rewrite Eq. (15.120) as Br.z;0/Da2 z31X sD02.sC1/.2sC1/c2sC1a z2s : (15.121) The Biot-Savart law (in SI units) gives the contribution from the current element I dsto Bat a point whose displacement from the current element is rsas dBD0 4IdsOrs r2s: (15.122) We now compute Bby integration of dsaround the current loop. The geometry is shown in Fig. 15.11. Note that d Bz, which will be the same for all current elements I ds, has the value d BzD0I 4r2ssinds; z aIˆr dB ds→rs=a2+z2 FIGURE 15.11 Biot-Savart law applied to a circular loop. ArfKen_Ch15-9780123846549.tex 752 Chapter 15 Legendre Functions whereis the labeled angle in Fig. 15.11 and rshas the value indicated in the figure. The integration over ssimply yields a factor 2a, and we see that sinDa=.a2Cz2/1=2, so BzD0Ia2 2.a2Cz2/3=2D0Ia2 2z3 1Ca2 z23=2 D0Ia2 2z31X sD0.1/s.2sC1/WW .2s/WWa z2s : (15.123) The binomial expansion in the second line of Eq. (15.123) is convergent for z>a. We are now ready to reconcile Eqs. (15.121) and (15.123), finding that 2.sC1/.2sC1/c2sC1D0I 2.1/s.2sC1/WW .2s/WW; which reduces to c2sC1D0I 2.1/sC1.2s1/WW .2sC2/WW: (15.124) We write final formulas for AandBin a form that recognizes that c2sD0, applicable forr>a: A'.r;/Da2 r21X sD0c2sC1a r2s P1 2sC1.cos/; (15.125) Br.r;/Da2 r21X sD0.2sC1/.2sC2/c2sC1a r2s P2sC1.cos/; (15.126) B.r;/Da2 r31X sD0.2sC1/c2sC1a r2s P1 2sC1.cos/: (15.127) These formulas can also be written in terms of complete elliptic integrals. See Smythe (Additional Readings) and Section 18.8 of this book. A comparison of magnetic current loop and finite electric dipole fields may be of interest. For the magnetic loop dipole, the preceding analysis gives Br.r;/D0Ia2 2r3 P13 2a r2 P3C ; (15.128) B.r;/D0Ia2 4r3 P1 1C3 4a r2 P1 3C : (15.129) From the finite electric dipole potential, Eq. (15.65), one can find Er.r;/Dqa "0r3 P1C2a r2 P3C ; (15.130) E.r;/Dqa 2" 0r3 P1 1a r2 P1 3C : (15.131) ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 753 The leading terms of both fields agree, and this is the basis for identifying both as dipole fields. As with electric multipoles, it is sometimes convenient to discuss point magnetic mul- tipoles. A point dipole can be formed by taking the limit a!0,I!1 , with Ia2held constant. The magnetic moment m is taken to be Ia2n, where nis a unit vector perpen- dicular to the plane of the current loop and in the sense given by the right-hand rule.  Exercises 15.4.1 Apply the Frobenius method to Eq. (15.72) to obtain Eq. (15.73) and verify that the numerator of that equation becomes zero if Dl.lC1/andjDlm. 15.4.2 Starting from the entries for P2 2andP1 2inTable 15.3, apply a recurrence formula to obtain P0 2(which is P2),P1 2, and P2 2. Compare your results with the value of P2from Table 15.1 and with values of P1 2andP2 2obtained by applying Eq. (15.81) to entries from Table 15.3. 15.4.3 Prove that Pm l.x/D.1/m.lm/W .lCm/WPm l.x/; where Pm l.x/is defined by Pm l.x/D.1/m 2llW.1x2/m=2dlCm dxlCm.x21/l: Hint. One approach is to apply Leibniz’s formula to .xC1/l.x1/l. 15.4.4 Show that P1 2l.0/D0; P1 2lC1.0/D.1/lC1.2lC1/WW .2l/WW; by each of these three methods: (a) Use of recurrence relations, (b) Expansion of the generating function, (c) Rodrigues formula. 15.4.5 Evaluate Pm l.0/form>0. ANS. Pm l.0/D8 < :.1/.lCm/=2.lCm1/WW .lm/WW;lCmeven; 0; lCmodd: 15.4.6 Starting from the potential of a finite dipole, Eq. (15.65), verify the formulas for the electric field components given as Eqs. (15.130) and(15.131). ArfKen_Ch15-9780123846549.tex 754 Chapter 15 Legendre Functions 15.4.7 Show that Pl l.cos/D.1/l.2l1/WWsinl; lD0;1;2;:::: 15.4.8 Derive the associated Legendre recurrence relation, PmC1 l.x/C2mx .1x2/1=2Pm l.x/Ch l.lC1/m.m1/i Pm1 l.x/D0: 15.4.9 Develop a recurrence relation that will yield P1 l.x/as P1 l.x/Df1.x;l/Pl.x/Cf2.x;l/Pl1.x/: Follow either of the procedures (a) or (b): (a) Derive a recurrence relation of the preceding form. Give f1.x;l/and f2.x;l/ explicitly. (b) Find the appropriate recurrence relation in print. (1) Give the source. (2) Verify the recurrence relation. ANS..a/P1 l.x/Dlx .1x2/1=2Pll .1x2/1=2Pl1. 15.4.10 Show that sind dcosPn.cos/DP1 n.cos/: 15.4.11 Show that (a)Z 0 d Pm l dd Pm l0 dCm2Pm lPm l0 sin2! sindD2l.lC1/ 2lC1.lCm/W .lm/Wl l0, (b)Z 0 P1 l sind P1 l0 dCP1 l0 sind P1 l d! sindD0. These integrals occur in the theory of scattering of electromagnetic waves by spheres. 15.4.12 As a repeat of Exercise 15.2.9, show, using associated Legendre functions, that 1Z 1x.1x2/P0 n.x/P0 m.x/dxDnC1 2nC12 2n1nW .n2/Wm;n1 Cn 2nC12 2nC3.nC2/W nWm;nC1: 15.4.13 EvaluateZ 0sin2P1 n.cos/d: ArfKen_Ch15-9780123846549.tex 15.4 Associated Legendre Equation 755 15.4.14 The associated Legendre function Pm l.x/satisfies the self-adjoint ODE .1x2/d2Pm l.x/ dx22xd Pm l.x/ dxC l.lC1/m2 1x2 Pm l.x/D0: From the differential equations for Pm l.x/andPk l.x/show that for k6Dm, 1Z 1Pm l.x/Pk l.x/dx 1x2D0: 15.4.15 Determine the vector potential and the magnetic induction field of a magnetic quadrupole by differentiating the magnetic dipole potential. ANS. AM QD0 2.Ia2/.dz/P1 2.cos/ r3Oe'Chigher-order terms, BM QD0.Ia2/.dz/" 3P2.cos/ r4OerP1 2.cos/ r4Oe# C . This corresponds to placing a current loop of radius aatz!dzand an oppositely directed current loop at z!dz . The vector potential and magnetic induction field of a point dipole are given by the leading terms in these expansions if we take the limit dz!0,a!0, and I!1 subject to Ia2dzDconstant. 15.4.16 A single circular wire loop of radius acarries a constant current I: (a) Find the magnetic induction Bforr<a; D=2. (b) Calculate the integral of the magnetic flux .Bd/over the area of the current loop, that is, aZ 0r dr2Z 0d'Bz r;D 2 : ANS.1. The Earth is within such a ring current, in which Iapproximates millions of amperes arising from the drift of charged particles in the Van Allen belt. 15.4.17 The vector potential Aof a magnetic dipole, dipole moment m, is given by A.r/D .0=4/.mr=r3/. Show by direct computation that the magnetic induction BDr Ais given by BD0 43OrOrm m r3: ArfKen_Ch15-9780123846549.tex 756 Chapter 15 Legendre Functions 15.4.18 (a) Show that in the point dipole limit the magnetic induction field of the current loop becomes Br.r;/D0 2m r3P1.cos/; B.r;/D0 2m r3P1 1.cos/; with mDIa2. (b) Compare these results with the magnetic induction of the point magnetic dipole of Exercise 15.4.17. Take mDOzm. 15.4.19 A uniformly charged spherical shell is rotating with constant angular velocity. (a) Calculate the magnetic induction Balong the axis of rotation outside the sphere. (b) Using the vector potential series of Example 15.4.2, find Aand then Bfor all points outside the sphere. 15.4.20 In the liquid-drop model of the nucleus, a spherical nucleus is subjected to small de- formations. Consider a sphere of radius r0that is deformed so that its new surface is given by rDr0h 1C 2P2.cos/i : Find the area of the deformed sphere through terms of order 2 2. Hint. dAD" r2Cdr d2#1=2 rsindd': ANS. AD4r2 0 1C4 5 2 2CO 3 2 . Note. The area element dAfollows from noting that the line element dsfor fixed'is given by dsD.r2d2Cdr2/1=2D" r2Cdr d2#1=2 d: 15.5 S PHERICAL HARMONICS Our earlier discussion of separated-variable methods for solving the Laplace, Helmholtz, or Schrödinger equations in spherical polar coordinates showed that the possible angular solutions2./8.'/ are always the same in spherically symmetric problems; in particular we found that the solutions for 8depended on the single integer index m, and can be written in the form 8m.'/D1p 2eim';mD:::;2;1;0;1;2;:::; (15.132) ArfKen_Ch15-9780123846549.tex 15.5 Spherical Harmonics 757 or, equivalently, 8m.'/D8 >>>>>>>< >>>>>>>:1p 2; mD0; 1pcosm'; m>0; 1psinjmj'; m<0:(15.133) The above equations contain the constant factors needed to make 8mnormalized, and those of different m2are automatically orthogonal because they are eigenfunctions of a Sturm-Liouville problem. It is straightforward to verify that in either Eq. (15.132) or Eq. (15.133) our choices of the functions for Cmandmmake8mand8morthogonal. Formally, our definitions are such that 2Z 0h 8m.'/i 8m0.'/d'Dmm0: (15.134) In Section 15.4 we found that the solutions 2./ could be identified as associated Leg- endre functions that can be labeled by the two integer indices landm, withlml. From the orthonormality integral for these functions, Eq. (15.105), we can define the nor- malized solutions 2lm.cos/Ds 2lC1 2.lm/W .lCm/WPm l.cos/; (15.135) satisfying the relation Z 0h 2lm.cos/i 2l0m.cos/sindDll0: (15.136) We have previously noted that an orthonormality condition of this type only applies if both functions2have the same value of the index m. The complex conjugate is not really nec- essary in Eq. (15.136) because the 2are real, but we write it anyway to maintain consistent notation. Note also that when the argument of Pm lisxDcos, then.1x2/1=2Dsin, so the Pm lare polynomials of overall degree lincosandsin. The product2lm8mis called a spherical harmonic, with that name usually implying that8mis taken with the definition as a complex exponential; see Eq. (15.132). Therefore we define Ym l.;'/s 2lC1 4.lm/W .lCm/WPm l.cos/eim': (15.137) These functions, being normalized solutions of a Sturm-Liouville problem, are orthonor- mal over the spherical surface, with 2Z 0d'Z 0sindh Ym1 l1.;'/i Ym2 l2.;'/Dl1l2m1m2: (15.138) ArfKen_Ch15-9780123846549.tex 758 Chapter 15 Legendre Functions The definition we introduced for the associated Legendre functions leads to specific signs for the Ym lthat are sometimes identified as the Condon-Shortley phase, after the authors of a classic text on atomic spectroscopy. This sign convention has been found to simplify various calculations, particularly in the quantum theory of angular momentum. One of the effects of this phase factor is to introduce an alternation of sign with mamong the positive- mspherical harmonics. The word “harmonic” enters the name of Ym lbecause solutions of Laplace’s equation are sometimes called harmonic functions. The squares of the real parts of the first few spherical harmonics are sketched in Figure 15.12; their functional forms are given in Table 15.4. Cartesian Representations For some purposes it is useful to express the spherical harmonics using Cartesian coordi- nates, which can be done by writing exp. i'/ascos'isin'and using the formulas forx;y;zin spherical polar coordinates (retaining, however, an overall dependence on r, necessary because the angular quantities must be independent of scale). For example, cosDz=r;sinexp. i'/Dsincos'isinsin'Dx riy rI (15.139) these quantities are all homogeneous (of degree zero) in the coordinates. Continuing to higher values of l, we obtain fractions in which the numerators are homo- geneous products of x,y,zof overall degree l, divided by a common factor rl. Table 15.4 includes the Cartesian expression for each of its entries. Overall Solutions As we have already seen in Section 9.4, the separation of a Laplace, Helmholtz, or even a Schrödinger equation in spherical polar coordinates can be written in terms of equations of the generic form R00C2 rR0Ch f.r/l.lC1/i RD0; (15.140) 1 sind d sind d C1 sin2d2 d'2Cl.lC1/ Ym l.;'/D0: (15.141) The function f.r/inEq. (15.140) is zero for the Laplace equation, k2for the Helmholtz equation, and EV.r/(VDpotential energy, EDtotal energy, an eigenvalue) for the Schrödinger equation. We have combined the and'equations into Eq. (15.141) and identified one of its solutions as Ym l. What is important to note right now is that the com- bined angular equation (and its boundary conditions and therefore its solutions) will be the same for all spherically symmetric problems, and that the angular solution affects the radial equation only through the separation constant l.lC1/. Thus, the radial equation will have solutions that depend on lbut are independent of the index m. In Section 9.4 we solved the radial equation for the Laplace and Helmholtz equations, with the results given in Table 9.2. For the Laplace equation r2 D0, the general solution ArfKen_Ch15-9780123846549.tex 15.5 Spherical Harmonics 759 m=0, l=1 m=1, l=1 m=0, l=2 m=1, l=2 m=2, l=2 m=0, l=3 m=1, l=3 m=2, l=3 m=3, l=3m=0, l=0 FIGURE 15.12 Shapes ofjReYm l.;'/j2for0l3,mD0:::l. in spherical polar coordinates is a sum, with arbitrary coefficients, of the solutions for the various possible values of landm: .r;;'/D1X lD0lX mDl almrlCblmrl1 Ym l.;'/I (15.142) ArfKen_Ch15-9780123846549.tex 760 Chapter 15 Legendre Functions Table 15.4 Spherical Harmonics (Condon-Shortley Phase) Y0 0.;'/D1p 4 Y1 1.;'/Dq 3 8sinei'Dq 3 8.xCiy/=r Y0 1.;'/Dq 3 4cosDq 3 4z=r Y1 1.;'/DCq 3 8sinei'Dq 3 8.xiy/=r Y2 2.;'/Dq 5 963 sin2e2i'D3q 5 96.x2y2C2ixy/=r2 Y1 2.;'/Dq 5 243 sincosei'Dq 5 243z.xCiy/=r2 Y0 2.;'/Dq 5 4 3 2cos21 2 Dq 5 4 3 2z21 2r2 =r2 Y1 2.;'/Dq 5 243 sincosei'DCq 5 243z.xiy/=r2 Y2 2.;'/Dq 5 963 sin2e2i'D3q 5 96.x2y22ixy/=r2 Y3 3.;'/Dq 7 288015 sin3e3i'Dq 7 288015Tx33xy2Ci.3x2yy3/U=r3 Y2 3.;'/Dq 7 48015 cossin2e2i'Dq 7 48015z.x2y2C2ixy/=r3 Y1 3.;'/Dq 7 48 15 2cos23 2 sinei'Dq 7 48 15 2z23 2r2 .xCiy/=r3 Y0 3.;'/Dq 7 4 5 2cos33 2cos Dq 7 4z 5 2z23 2r2 =r3 Y1 3.;'/DCq 7 48 15 2cos23 2 sinei'Dq 7 48 15 2z23 2r2 .xiy/=r3 Y2 3.;'/Dq 7 48015 cossin2e2i'Dq 7 48015z.x2y22ixy/=r3 Y3 3.;'/DCq 7 288015 sin3e3i'Dq 7 288015Tx33xy2i.3x2yy3/U=r3 for the Helmholtz equation .r2Ck2/ D0, the radial equation has the form given in Eq. (14.148), so the general solution assumes the form .r;;'/D1X lD0lX mDl almjl.kr/Cblmyl.kr/ Ym l.;'/: (15.143) Laplace Expansion Part of the importance of spherical harmonics lies in the completeness property, a conse- quence of the Sturm-Liouville form of Laplace’s equation. Here this property means that any function f.;'/ (with sufficient continuity properties) evaluated over the surface of a sphere can be expanded in a uniformly convergent double series of spherical harmonics.4 4For a proof of this fundamental theorem, see E. W. Hobson (Additional Readings), chapter VII. ArfKen_Ch15-9780123846549.tex 15.5 Spherical Harmonics 761 This expansion, known as a Laplace series, takes the form f.;'/D1X lD0lX mDlclmYm l.;'/; (15.144) with clmDD Ym l f.;'/E D2Z 0d'Z 0sindYm l.;'/f.;'/: (15.145) A frequent use of the Laplace expansion is in specializing the general solution of the Laplace equation to satisfy boundary conditions on a spherical surface. This situation is illustrated in the following example. Example 15.5.1 SPHERICAL HARMONIC EXPANSION Consider the problem of determining the electrostatic potential within a charge-free spheri- cal region of radius r0, with the potential on the spherical bounding surface specified as an arbitrary function V.r0;;'/ of the angular coordinates and'. The potential V.r;;'/ is the solution of the Laplace equation satisfying the boundary condition at rDr0and regular for all rr0. This means it must be of the form of Eq. (15.142), with the coefficients blm set to zero to ensure a solution that is nonsingular at rD0. We proceed by obtaining the spherical harmonic expansion of V.r0;;'/ , namely Eq. (15.144), with coefficients clmDD Ym l.;'/ V.r0;;'/E : Then, comparing Eq. (15.142), evaluated for rDr0, V.r0;;'/D1X lD0lX mDlalmrl 0Ym l.;'/; with the expression from Eq. (15.144), V.r0;;'/D1X lD0lX mDlclmYm l.;'/; we see that almDclm=rl 0, so V.r;;'/D1X lD0lX mDlclmr r0l Ym l.;'/:  ArfKen_Ch15-9780123846549.tex 762 Chapter 15 Legendre Functions Example 15.5.2 LAPLACE SERIES—GRAVITY FIELDS This example illustrates the notion that sometimes it is appropriate to replace the spherical harmonics by their real counterparts (in terms of sine and cosine functions). The gravity fields of the Earth, the Moon, and Mars have been described by a Laplace series of the form U.r;;'/DG M R" R r1X lD2lX mD0R rlC1 ClmYe ml.;'/CSlmYo ml.;'/# : (15.146) Here Mis the mass of the body, Ris its equatorial radius, and Gis the gravitational con- stant. The real functions Ye mlandYo mlare defined by Morse and Feshbach (see Additional Readings) as the unnormalized forms Ye ml.;'/DPm l.cos/cos m'; Yo ml.;'/DPm l.cos/sin m': Note that Morse and Feshbach place the mindex before l. The normalization integrals for YeandYoare the topic of Exercise 15.5.6. Satellite measurements have led to the numerical values for C20,C22, and S22shown in Table 15.5. Table 15.5 Gravity Field Coefficients, Eq. (15.145). CoefficientaEarth Moon Mars C20 1:083103.0:2000:002/103.1:960:01/103 C22 0:16105.2:40:5/105.51/105 S220:09105.0:50:6/105.31/105 aC20represents an equatorial bulge, whereas C22andS22represent an azimuthal dependence of the gravitational field.  Symmetry of Solutions The angular solutions of given lbut different mare closely related in that they lead to the same solution for the radial equation. Except when lD0, the individual solutions Ym lare not spherically symmetric, and we must recognize that a spherically symmetric problem can have solutions with less than the full spherical symmetry. A classical example of this phenomenon is provided by the Earth-Sun system, which has a spherically symmetric grav- itational potential. However, the actual orbit of the Earth is planar. This apparent dilemma is resolved by noting that a solution exists for any orientation of the Earth’s orbital plane; that actually occurring was determined by “initial conditions.” Returning now to the Laplace equation, we see that a radial solution for given l, i.e., rl orrl1, is associated with 2lC1different angular solutions Ym l(lml), no one of which (for l6D0) has spherical symmetry. The most general solution for this lmust be a linear combination of these 2lC1mutually orthogonal functions. Put another way, ArfKen_Ch15-9780123846549.tex 15.5 Spherical Harmonics 763 the solution space of the angular solution of the Laplace equation for given lis a Hilbert space containing the 2lC1members Yl l.;'/;:::; Yl l.;'/ . Now, if we write the Laplace equation in a coordinate system .0;'0/oriented differently than the original coordinates, we must still have the same angular solution set, meaning that Ym l.0;'0/must be a linear combination of the original Ym l. Thus, we may write Ym l.0;'0/DlX m0DlDl m0mYm0 l.;'/; (15.147) where the coefficients Ddepend on the coordinate rotation involved. Note that a coordi- nate rotation cannot change the rdependence of our solution to the Laplace equation, so Eq. (15.147) does not need to include a sum over all values of l. As a specific example, we see (Fig. 15.12) that for lD1we have three solutions that appear similar, but with differ- ent orientations. Alternatively, from Table 15.4 we see that the angular solutions Ym 1have forms proportional to z=r,.xCiy/=r, and.xiy/=r, meaning that they can be combined to form arbitrary combinations of x=r,y=r, and z=r. Since a rotation of the coordinate axes converts x,y, and zinto linear combinations of each other, we can understand why the set of three functions Ym 1(mD0;1;1) is closed under coordinate rotations. ForlD2, there are five possible mvalues, so the angular functions of this lvalue form a closed space containing five independent members. A fuller discussion of these spaces spanned by angular functions is part of what will be considered in Chapter 16. Applying the preceding analysis to solutions of the Schrödinger equation, the eigenval- ues of which are determined by solving its radial ODE for various values of the separation constant l.lC1/, we see that all solutions for the same lbut different mwill have the same eigenvalues Eand radial functions, but will differ in the orientation of their angular parts. States of the same energy are called degenerate, and the independence of Ewith respect tomwill cause a.2lC1/-fold degeneracy of the eigenstates of given l. Example 15.5.3 SOLUTIONS FOR lD1ATARBITRARY ORIENTATION Let’s do this problem in Cartesian coordinates. The angular solution Y0 1to Laplace’s equa- tion is shown in Table 15.4 to be proportional to z=r, which for our present purposes we write.rOez/=r, whereOezis a unit vector in the zdirection. We seek a similar solution, withOezreplaced by an arbitrary unit vector OeuDcos OexCcos OeyCcos Oez, where cos ,cos , and cos are the direction cosines of Oeu. We get immediately .rOeu/ rDx rcos Cy rcos Cz rcos : Consulting the Cartesian-coordinate expressions for the spherical harmonics in Table 15.4, we see that the above expression can be written .rOu/ rDr 8 3 Y1 1Y1 1 2! cos Cr 8 3 Y1 1Y1 1 2i! cos Cr 4 3Y0 1cos : This shows that all three Ym 1are needed to reproduce Y0 1at an arbitrary orientation. Similar manipulations can be carried out for other landmvalues.  ArfKen_Ch15-9780123846549.tex 764 Chapter 15 Legendre Functions Further Properties The main properties of the spherical harmonics follow directly from those of the functions 2lmand8m. We summarize briefly: Special values. At D0, the polar direction in the spherical coordinates, the value of ' becomes immaterial, and all Ym lthat have'dependence must vanish. Using also the fact thatPl.1/D1, we find in general Ym l.0;'/Dr 2lC1 4m0: (15.148) A similar argument for Dleads to Ym l.;'/D.1/lr 2lC1 4m0: (15.149) Recurrence formulas. Using the recurrence formulas developed for the associated Leg- endre functions, we get for the spherical harmonics with arguments .;'/ : cosYm lD.lmC1/.lCmC1/ .2lC1/.2lC3/1=2 Ym lC1 C.lm/.lCm/ .2l1/.2lC1/1=2 Ym l1; (15.150) ei'sinYm lD.lmC1/.lmC2/ .2lC1/.2lC3/1=2 Ym1 lC1 .lm/.lm1/ .2l1/.2lC1/1=2 Ym1 l1: (15.151) Some integrals. These recurrence relations permit the ready evaluation of some inte- grals of practical importance. Our starting point is the orthonormalization condition, Eq. (15.138). For example, the matrix elements describing the dominant (electric dipole) mode of interaction of an electromagnetic field with a charged system in a spherical har- monic state are proportional to Zh Ym0 l0i cosYm ld: Using Eq. (15.150) and invoking the orthonormality of the Ym l, we find Zh Ym0 l0i cosYm ldD.lmC1/.lCmC1/ .2lC1/.2lC3/1=2 m0ml0;lC1 C.lm/.lCm/ .2l1/.2lC1/1=2 m0ml0;l1: (15.152) Equation (15.152) provides a basis for the well-known selection rule for dipole radiation. ArfKen_Ch15-9780123846549.tex 15.5 Spherical Harmonics 765 Additional formulas involving products of three spherical harmonics and the detailed behavior of these quantities under coordinate rotations are more appropriately discussed in connection with a study of angular momentum and are therefore deferred to Chapter 16. Exercises 15.5.1 Show that the parity of Ym l.;'/ is.1/l. Note the disappearance of any mdependence. Hint. For the parity operation in spherical polar coordinates, see Exercise 3.10.25. 15.5.2 Prove that Ym l.0;'/D2lC1 41=2 m0: 15.5.3 In the theory of Coulomb excitation of nuclei we encounter Ym l.=2; 0/. Show that Ym l 2;0 D2lC1 41=2T.lm/W.lCm/WU1=2 .lm/WW.lCm/WW.1/.lm/=2;lCmeven, D0;lCmodd: 15.5.4 The orthogonal azimuthal functions yield a useful representation of the Dirac delta function. Show that .'1'2/D1 21X mD1eim.'1'2/: Note. This formula assumes that '1and'2are restricted to 0'<2. Without this restriction there will be additional delta-function contributions at intervals of 2in '1'2. 15.5.5 Derive the spherical harmonic closure relation 1X lD0ClX mDlh Ym l.1;'1/i Ym l.2;'2/D1 sin1.12/.' 1'2/ D.cos1cos2/.' 1'2/: 15.5.6 In some circumstances it is desirable to replace the imaginary exponential of our spher- ical harmonic by sine or cosine. Morse and Feshbach (see Additional Readings) define Ye mlDPm l.cos/cosm'; m0; Yo mlDPm l.cos/sinm'; m>0; and their normalization integrals are 2Z 0Z 0TYeoro mn.;'/U2sindd'D4 2.2nC1/.nCm/W .nm/W;nD1;2;::: D4; nD0: ArfKen_Ch15-9780123846549.tex 766 Chapter 15 Legendre Functions These spherical harmonics are often named according to the patterns of their positive and negative regions on the surface of a sphere: zonal harmonics for mD0, sectoral har- monics for mDn, and tesseral harmonics for 0<m<n. For Ye mn,nD4,mD0;2;4, indicate on a diagram of a hemisphere (one diagram for each spherical harmonic) the regions in which the spherical harmonic is positive. 15.5.7 A function f.r;;'/ may be expressed as a Laplace series f.r;;'/DX l;malmrlYm l.;'/: Lettinghi sphere denote the average over a sphere centered on the origin, show that D f.r;;'/E sphereDf.0;0;0/: 15.6 L EGENDRE FUNCTIONS OF THE SECOND KIND The Legendre equation, a linear second-order ODE, has two independent solutions. Writing this equation in the form y002x 1x2y0l.lC1/ 1x2yD0; (15.153) and restricting consideration to integer l0, our objective is to find a second solution that is linearly independent from the Legendre polynomials Pl.x/. Using the procedure of Section 7.6, and denoting the second solution Ql.x/, we have Ql.x/DPl.x/xZexp2 4xZ 2x=.1x2/dx3 5 TPl.x/U2dx DPl.x/xZdx .1x2/TPl.x/U2dx: (15.154) Since any linear combination of Pland the right-hand side of Eq. (15.154) is equally valid as a second solution of the Legendre ODE, we note that Eq. (15.154) defines both the scale and the specific functional form of Ql. Using Eq. (15.154), we can obtain explicit formulas for the Ql. We find (remembering thatP0D1and expanding the denominator in partial fractions): Q0.x/DxZ1 1x2dxD1 2Z1 1CxC1 1x dxD1 2ln1Cx 1x : (15.155) ArfKen_Ch15-9780123846549.tex 15.6 Legendre Functions of the Second Kind 767 Continuing to Q1, the partial fraction expansion is a bit more involved, but leads to a simple result. Noting that P1.x/Dx, we have Q1.z/DxxZdx .1x2/x2dxDx 2ln1Cx 1x 1: (15.156) With significantly more work, we can obtain Q2: Q2.x/D1 2P2.x/ln1Cx 1x 3x 2: (15.157) This process can in principle be repeated for larger l, but it is easier and more instructive to verify that the forms of Q0,Q1, and Q2are consistent with the Legendre-function recur- rence relations,5and then to obtain Qlof larger lby recurrence. The recurrence formulas, originally written for Plin Eq. (15.18), are .lC1/QlC1.x/.2lC1/x Ql.x/Cl Ql1.x/D0; (15.158) .2lC1/Ql.x/DQ0 lC1.x/Q0 n1.x/: (15.159) Verification that Q0,Q1, and Q2satisfy these recurrence formulas is straightforward and is left as a exercise. Extension to higher lleads to the formula Ql.x/D1 2Pl.x/ln1Cx 1x 2l1 1lPl1.x/2l5 3.l1/Pl3.x/:(15.160) Many applications using the functions Ql.x/involve values of xoutside the range 1<x<1. IfEq. (15.160) is extended, say, beyond C1, then 1xwill become neg- ative and make a contribution ito the logarithm, thereby making a contribution iPl toQl. Our solution will still remain a solution if this contribution is removed, and it is therefore convenient to define the second solution for xoutside the range .1;C1/ with ln1Cx 1x replaced by lnxC1 x1 : From a complex-variable perspective, the logarithmic term in the solutions Qlis related to the singularity in the ODE at zD1 , reflecting the fact that to make the solutions single- valued it will be necessary to make a branch cut, traditionally taken on the real axis from 1toC1. Then the Qlwith the.1Cx/=.1x/logarithm are recovered on 1<x<1 if we average the results from the .zC1/=.z1/form on the two sides of the branch cut. The behavior of the Qlis illustrated by plots for x<1in Fig. 15.13 and for x>1in Fig. 15.14. Note that there is no singularity at xD0but all the Qlexhibit a logarithmic singularity at xD1. 5In Section 15.1 we showed that any set of functions that satisfies the recurrence relations reproduced here also satisfies the Legendre ODE. ArfKen_Ch15-9780123846549.tex 768 Chapter 15 Legendre Functions 1.5 1.00.5 0.2 0.4 0.6 0.8 1.00 −0.5 −1.0Q0(x) Q2(x) Q1(x)x FIGURE 15.13 Legendre functions Ql.x/,0x<1. 100 10−1 10−2 10−3 10−4x 2468 1 0Q0(x) Q1(x) Q2(x) FIGURE 15.14 Legendre functions Ql.x/,x>1. ArfKen_Ch15-9780123846549.tex 15.6 Legendre Functions of the Second Kind 769 Properties 1. An examination of the formulas for Ql.x/reveals that if lis even, then Ql.x/is an odd function of x, while Ql.x/of odd lare even functions of x. More succinctly, Ql.x/D.1/lC1Ql.x/. 2. The presence of the logarithmic term causes Ql.1/D1 for all l. 3. Because xD0is a regular point of the Legendre ODE, Ql.0/must for all lbe finite. The symmetry of Qlcauses Ql.0/to vanish for even l; it is shown in the next subsec- tion that for odd l, Q2sC1.0/D.1/sC1.2s/WW .2sC1/WW: (15.161) 4. From the result of Exercise 15.6.3, it can be shown that Ql.1/D0. Alternate Formulations Because the singular points of the Legendre ODE nearest to the origin are at the points 1, it should be possible to describe Ql.x/as a power series about the origin, with convergence forjxj<1. Moreover, since the only other singular point of the Legendre equation is a regular singular point at infinity, it should also be possible to express one of its solutions as a power series in 1=x, i.e., a series about the point at infinity, which must converge for jxj>1. To obtain a power series about xD0, we return to the discussion of the Legendre ODE presented in Section 8.3, where we saw that an expansion of the form y.x/D1X jD0ajxsCj(15.162) led to an indicial equation with solutions sD0andsD1, and with the ajsatisfying the recurrence formula, for eigenvalue l.lC1/, ajC2Daj.sCj/.sCjC1/l.lC1/ .sCjC2/.sCjC1/;jD0;2;:::: (15.163) When lis even, we found that Pl.x/was obtained as the solution y.x/from the indicial- equation solution sD0, and we did not make use (for even l) of the solution from sD1 because that solution was not a polynomial and did not converge at xD1. However, we are now seeking a second solution and are no longer restricting attention to those that converge atxD1 . Thus, a second solution linearly independent of Plmust be that produced (again, for even l) as the series obtained when sD1. This second solution will have odd parity, and therefore must be proportional to Ql.x/. Continuing, for even l, with sD1,Eq. (15.163) becomes ajC2Daj.lCjC2/.lj1/ .jC2/.jC3/; ArfKen_Ch15-9780123846549.tex 770 Chapter 15 Legendre Functions corresponding to Ql.x/Dbl x.l1/.lC2/ 3Wx3C.l3/.l1/.lC2/.lC4/ 5Wx5 :(15.164) Here blis the value of the coefficient of the expansion needed to give the formula for Ql the proper scaling. For odd l, the corresponding formula, with sD0, is an even function ofx, and must therefore be proportional to Ql: Ql.x/Dbl 1l.lC1/ 2Wx2C.l2/l.lC1/.lC3/ 4Wx4 : (15.165) To find the values of the scale factors bl, we turn now to the explicit forms for Q0and Q1,Eqs. (15.155) and(15.156). Expanding the logarithm, we find (again keeping only the lowest-order terms) Q0.x/DxC;Q1.x/D1C: From the recurrence formula, Eq. (15.158), keeping only the lowest-order contributions, we find 2Q2D3x Q1Q0! Q2D2 xC 3Q3D5x Q22Q1! Q3D2=3C 4Q4D7x Q33Q2! Q4D8x=3C D: These results generalize to blD8 >>>< >>>:.1/p.2p/WW .2p1/WWleven, lD2p, .1/pC1.2p/WW .2pC1/WWlodd, lD2pC1.(15.166) One may now combine the values of the coefficients blwith the expansions in Eqs. (15.164) and (15.165) to obtain entirely explicit series expansions of Ql.x/about xD0. This is the topic of Exercise 15.6.2. As mentioned earlier, the point xD1 is a regular singular point, and expansion about this point yields an expansion of Ql.x/in inverse powers of x. That expansion is consid- ered in Exercise 15.6.3. Exercises 15.6.1 Show that if lis even, Ql.x/DQl.x/, and that if lis odd, Ql.x/DQl.x/. 15.6.2 Show that ArfKen_Ch15-9780123846549.tex Additional Readings 771 (a) Q2p.x/D.1/p22ppX sD0.1/s.pCs/W.ps/W .2sC1/W.2p2s/Wx2sC1 C22p1X sDpC1.pCs/W.2s2p/W .2sC1/W.sp/Wx2sC1;jxj<1; (b) Q2pC1.x/D.1/pC122ppX sD0.1/s.pCs/W.ps/W .2s/W.2p2sC1/Wx2s C22pC11X sDpC1.pCs/W.2s2p2/W .2s/W.sp1/Wx2s;jxj<1: 15.6.3 (a) Starting with the assumed form Ql.x/D1X jD0bl jxkj; show that Ql.x/Dbl0xl11X sD0.lCs/W.lC2s/W.2lC1/W sW.lW/2.2lC2sC1/Wx2s: (b) The standard choice of bl0is bl0D2l.lW/2 .2lC1/W; leading to the final result Ql.x/Dxl11X sD0.lC2s/W .2s/WW.2lC2sC1/WWx2s: Show that this choice of bl0brings this negative power-series form of Qn.x/into agreement with the closed-form solutions. 15.6.4 (a) Using the recurrence relations, prove (independent of the Wronskian relation) that nh Pn.x/Qn1.x/Pn1.x/Qn.x/i DP1.x/Q0.x/P0.x/Q1.x/: (b) By direct substitution show that the right-hand side of this equation equals 1. Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Hobson, E. W., The Theory of Spherical and Ellipsoidal Harmonics. New York: Chelsea (1955). This is a very complete reference, which is the classic text on Legendre polynomials and all related functions. ArfKen_Ch15-9780123846549.tex 772 Chapter 15 Legendre Functions Jackson, J. D., Classical Electrodynamics, 3rd ed. New York: Wiley (1999). Margenau, H., and G. M. Murphy, The Mathematics of Physics and Chemistry , 2nd ed. Princeton, NJ: Van Nostrand (1956). Morse, P. M., and H. Feshbach, Methods of Theoretical Physics, 2 vols. New York: McGraw-Hill (1953). This work is detailed but at a rather advanced level. Smythe, W. R., Static and Dynamic Electricity , 3rd ed. New York: McGraw-Hill (1968), reprinted, Taylor & Francis (1989), paperback. Advanced, detailed, and difficult. Includes use of elliptic integrals to obtain closed formulas. Whittaker, E. T., and G. N. Watson, A Course of Modern Analysis, 4th ed. Cambridge, UK: Cambridge University Press (1962), paperback. ArfKen_Ch16-9780123846549.tex CHAPTER 16 ANGULAR MOMENTUM The traditional quantum mechanical treatment of central force problems starts from solutions to the time-independent Schrödinger equation, which, for a single particle of mass mmoving subject to a potential V.r/, is an eigenvalue problem of the general form Nh2 2mr2 .r/CV.r/ .r/DE .r/: (16.1) HereNhis Planck’s constant divided by 2, in SI units approximately 1:051034J-s (joule-seconds); the very small value of this constant causes quantum behavior to be per- ceptible under most circumstances only at small distances and for particles of small mass; the relevant ranges are typically at atomic scales of mass and length. The basic interpretation of the Schrödinger equation is that if the energy Eof the particle is measured, the result will be one of the eigenvalues of Eq. (16.1), and (subsequent to the measurement) the location of the particle will be described by a probability distribution P.r/d3rDj .r/j2d3r; where .r/is an eigenfunction corresponding to E. As we have seen in Chapters 9 and 15, will in general have angular as well as radial dependence, and its angular part can be written in terms of the spherical harmonics Ym l.;'/ . A more detailed interpretation of Eq. (16.1) is to identify it as an operator equation in which the momentum pis identified with the operator iNhr, while functions of position, such as the potential energy V.r/, are identified as multiplicative operators. Viewed in this way, the operator.Nh2=2m/r2is seen to represent p2=2m (i.e., the kinetic energy T), and Eq. (16.1) then becomes equivalent to H .TCV/ DE ; (16.2) where H, the Hamiltonian, is an operator whose eigenvalues are the possible values of the total energy. The Hamiltonian His a special operator in quantum mechanics because its eigenfunc- tions yield stationary probability distributions (they do not evolve into different distribu- tions over time). However, His just like any other quantum operator Krepresenting a 773 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch16-9780123846549.tex 774 Chapter 16 Angular Momentum dynamical quantity (with eigenvalues kthat can be the result of measurement of K). If is simultaneously an eigenfunction of Hand of K, then we can have definite values of both Eandkthat will not evolve as a function of time, and the measuring of either will not disturb the definite value of the other. This state of affairs can only be achieved if Hand Kcommute, because (see Section 6.4) TH;KUD0is a necessary and sufficient condition thatHandKhave a set of simultaneous eigenfunctions. In earlier chapters, we examined commutators such as Tx;pxUDi(here and except when noted, we use a unit system with Nhset to unity to avoid unnecessary notational complex- ity). The nonzero commutator of xandpxtells us that we cannot simultaneously obtain unambiguous measurement of both these quantities (i.e., we do not have a complete set of states that are simultaneously eigenfunctions of xandpx). This is the mathematical basis of the Heisenberg uncertainty principle in quantum mechanics. The notion of simultaneous eigenfunctions and therefore commutation plays a key role in the study of angular momentum in quantum mechanics. Angular momentum is con- served in the classical central force problem, and one of the focal points of the present chap- ter is to understand the properties of angular momentum operators in quantum mechanics. 16.1 A NGULAR MOMENTUM OPERATORS In classical physics, the kinetic energy of a particle of mass can be written in terms of its momentum pasTclassDp2=2. Note that we are using for the particle mass to avoid confusion with the usual notation of the azimuthal wave functions 'm. Most of the literature uses mfor both quantities. Introducing spherical polar coordinates, Tclass can be divided into radial and angular parts, with the angular kinetic energy of the form L2 class=2r2. Here Lclassis the angular momentum, defined as LclassDrp. Following the usual Schrödinger representation of quantum mechanics, the classical linear momentum p is replaced (in a unit system with NhD1) by the operatorir. The quantum-mechanical kinetic energy operator is TQMDr2=2, which in spherical polar coordinates can be written TQMD1 2@2 @r2C2 r@ @r 1 2r21 sin@ @ sin@ @ C1 sin2@2 @'2 : (16.3) Like the classical kinetic energy, TQMcan also be divided into radial and angular parts, with the angular part identified in terms of the angular momentum: TQMDTradial;QMC1 2r2L2 QM; (16.4) Tradial;QMD1 2@2 @r2C2 r@ @r ; (16.5) L2 QMD1 sin@ @ sin@ @ 1 sin2@2 @'2: (16.6) Since our focus here is on the quantum-mechanical operators, we drop the notation “QM” from now on. ArfKen_Ch16-9780123846549.tex 16.1 Angular Momentum Operators 775 The notation L2in Eq. (16.6) is only really appropriate if it is consistent with the defini- tion of the quantum-mechanical angular momentum operator, which must have the form LDrpDirr: (16.7) One way to confirm Eq. (16.6) is to start from the expression for Lin spherical polar coor- dinates, which can be deduced by applying the operator rpto an arbitrary function : L Dirr DirOer Oer@ @rCOe1 r@ @COe'1 rsin@ @' ; from which we extract the formula LDi Oe1 sin@ @'Oe'@ @ : (16.8) We then rewrite Lin Cartesian components Lx,Ly,Lz(but still expressed in polar coor- dinates) and evaluate L2DLLDL2 xCL2 yCL2 z: (16.9) This process is the topic of Exercise 3.10.32, and leads, as expected, to Eq. (16.6). In Section 15.5 we identified the solutions of the angular part of the Laplace and Schrödinger equations for central force problems as the spherical harmonics, denoted Ym l.;'/ . Now that we have also written the angular part of these equations in terms of L2, we see that the Ym lcan be identified as eigenfunctions of L2, i.e., that they are angular momentum eigenfunctions, satisfying an eigenvalue equation of the form L2Ym l.;'/Dl.lC1/Ym l.;'/: (16.10) Summarizing the discussion to this point, and drawing on previously established properties of the spherical harmonics: The spherical harmonics Ym lare eigenfunctions of L2with eigenvalue l.lC1/. The eigen- functions for a given lare.2lC1/-fold degenerate and can be indexed by their mvalues, which range in unit steps from ltol. We now strive for a deeper understanding of the role of angular momentum. The solu- tions to the time-independent Schrödinger equation are the eigenfunctions of its total energy operator, the Hamiltonian H. We have just observed that for central force prob- lems the angular solutions are eigenfunctions of the angular momentum operator L2. In order for these two statements to be mutually consistent, it is necessary that HandL2 commute. For the systems under consideration here, this is clearly true, since we have assumed that His of the form TCV.r/, so HDTradial.r/C1 2r2L2.;'/CV.r/: Since the only angle-dependent quantity in His the operator L2, and since L2obviously commutes with itself and is independent of r, we have TH;L2UD0: ArfKen_Ch16-9780123846549.tex 776 Chapter 16 Angular Momentum The fact that HandL2have simultaneous eigenfunctions in central force problems means that the stationary states of such systems can be characterized by definite values of both the energy and the angular momentum quantum number l. States of different lwere ultimately identified with series of lines in the emission and absorption spectra of the hydrogen atom that had previously been labeled “sharp,” “diffuse,” “principal,” and “fundamental.” This identification caused physicists to use the initial letters of these names as synonyms for lvalues; hence it has become essential to know that the code letters for lD0;1;2, and 3 are respectively s;p;d, and f. For l>3, the code letters run alphabetically: g;h:::. Turning now to the components of L, we have (cf. Exercise 3.10.31) TLj;LkUDi"jknLnandTL2;LjUD0; (16.11) where j;k;nare different members of the set (1,2,3) and "jknis a Levi-Civita symbol. Although the Ljdo not commute with each other, all commute with L2and hence also with H, soH,L2, and any one component of Lmutually commute. We conclude that there exists a set of simultaneous eigenfunctions of H,L2, and any one component of L. For this purpose we usually pick Lz, motivated by the fact that, in spherical polar coordinates, it is, as found in Exercise 3.10.29, LzDi@ @': (16.12) For reference, we copy here the far more complicated results for LxandLy, obtained from Exercise 3.10.30: LxDisin'@ @Cicotcos'@ @'; LyDicos'@ @Cicotsin'@ @':(16.13) The spherical harmonics are, in fact, eigenfunctions of Lz. Since Lzeim'Di@ @'eim'Dmeim'; (16.14) we see that Ym lis an eigenfunction of Lzwith eigenvalue m. This is one of the reasons why the complex exponentials, rather than the trigonometric functions, were chosen in the definitions of the spherical harmonics. It is obvious that cosm'is not an eigenfunction of Lz:i.@=@'/ cosm'Dimsinm'. Note, however, that exp. im'/,cosm', and sinm' are all eigenfunctions of the operator L2 zD@2=@'2with eigenvalue m2. Ladder Operators The commutators of the angular momentum components permit the development of some useful algebraic relationships. While these relationships can be found from the specific forms of the operators (cf. Exercise 3.10.30), more general and valuable results are ArfKen_Ch16-9780123846549.tex 16.1 Angular Momentum Operators 777 obtained by derivations based only on the commutators given in Eq. (16.11). We define the operators LCDLxCi Ly;LDLxi Ly; (16.15) and consider the commutators TLz;LCUDTLz;LxUCiTLz;LyUDi LyCi.i Lx/DLC; (16.16) TLz;LUDTLz;LxUiTLz;LyUDi Lyi.i Lx/DL: (16.17) We start by applying Eq. (16.16) to a function m l, which is assumed to be a normalized simultaneous eigenfunction of L2, with eigenvalue l, and of Lz, with eigenvalue m; the form of m l(and even the space within which it resides) need not be specified to carry out the present discussion. Moreover, at this point we introduce no information about the possible values of landm. However, to visualize what we are doing, the reader can keep in mind that one possible interpretation of m lis the spherical harmonic Ym l. We have TLz;LCU m lDLzLC m lLCLz m lDLC m l: Since Lz m lDm m l, we can rewrite the central and right-hand members of the above equation as Lz.LC m l/m.LC m l/D.LC m l/; which rearranges to Lz.LC m l/D.mC1/.LC m l/: (16.18) This tells us that if LC m lis nonzero, it is an eigenfunction of Lzwith eigenvalue mC1; for that reason LCcan be called a raising operator. By itself, this analysis tells us nothing about the value(s) of m, but only that LCincreases min unit steps. A similar development shows that Lis alowering operator, corresponding to the equation Lz.L m l/D.m1/.L m l/: (16.19) Raising and lowering operators are collectively referred to as ladder operators. Next, we recall that TL2;LiUD0for all components Li. This means also that TL2;LCUD0, so L2.LC m l/DLCL2 m lDl.LC m l/; showing that .LC m l/is still an eigenfunction of L2with the same eigenvalue, l, as m l. Note that we did not need to know the value of lto draw this conclusion. Summarizing, the operators Lconvert m linto quantities proportional to m1 l, with the conversion failing only if L m lD0. While Eqs. (16.18) and(16.19) tell us that Lare ladder operators, they do not tell us whether the quantities L m lare normalized. To address this problem, we write the normalization expression for LC m lin the form hLC m ljLC m liDh m ljLLCj m li; where we have used the fact that, because LxandLyare Hermitian, .LC/†DL. ArfKen_Ch16-9780123846549.tex 778 Chapter 16 Angular Momentum To obtain more information about LLC, we rearrange L2as follows: L2D1 2.LCLCLLC/CL2 zDLLCC1 2TLC;LUCL2 z; (16.20) a result that can be easily verified by expanding LC,L, and the commutator. Then we introduce TLC;LUDTLxCi Ly;Lxi LyUD iTLx;LyUCiTLy;LxUD2Lz; (16.21) and solve Eq. (16.20) for LLC, obtaining LLCDL2L2 zLz: (16.22) Using the fact that m lis a normalized eigenfunction of both L2andLz, we can now perform the evaluation hLC m ljLC m liDh m ljLLCj m liDh m ljL2L2 zLzj m li Dlm2m: (16.23) A parallel analysis leads to the companion result hL m ljL m liDh m ljLCLj m liDlm2Cm: (16.24) If we use the expressions in Eqs. (16.23) and (16.24) to account for the scale factors gen- erated by the ladder operators, we can summarize their action as LC m lDp lm.mC1/ mC1 l; (16.25) L m lDp lm.m1/ m1 l; where, the reader may recall, lis the eigenvalue of L2corresponding to quantum number l; the current analysis has not yet determined its value. The expressions in Eq. (16.25) have also incorporated the assumption that the signs of the m lare related as shown. That is a matter of definition, and when the m lare taken to be the spherical harmonics Ym l, the Condon-Shortley phase assignment was deliberately designed to make Eq. (16.25) consis- tent with the signs given the Ym lin Table 15.4. Next, we return to Eq. (16.23) and note that since it describes a normalization integral, it is inherently nonnegative, and can be zero only if LC m lis identically zero. The right- hand side of Eq. (16.23), however, will be become negative if mis permitted to get too large, so for any fixed l(and therefore a fixed l), there must be some largest m, which we callmmax, for which there exists a mmax l. But if we use Eq. (16.23) to evaluate LC mmax l, we will, unless lmmax.mmaxC1/D0, generate a function with mDmmaxC1, thereby creating an inconsistency. Giving mmaxthe name l(permitted because within the current derivation we have not yet assigned a meaning to l), what we have found so far is that lDl.lC1/and that the maximum value of mismDl. Remember that we still know nothing about possible values for l. Turning now to Eq. (16.24), and inserting l.lC1/forl, we note that if mis permitted to become too negative we will again have an inconsistent situation, and that it is necessary that for some mminthe right-hand side of Eq. (16.24) must vanish. Thus, we require l.lC 1/mmin.mmin1/D0, an equation that is satisfied for mminDlC1(which is clearly irrelevant), and for mminDl(the solution we want). ArfKen_Ch16-9780123846549.tex 16.1 Angular Momentum Operators 779 Finally, we observe that, starting from some m lwith mDl, we have the severe limita- tion that application of the lowering operator Lwill decrease the mvalue in unit steps, but must ultimately reach mDlto avoid the generation of an inconsistency. This state of affairs is possible if lis a nonnegative integer, in which case there are 2lC1possible mvalues, ranging in unit steps from ltol. However, it is also possible to assign la half-integer value, as mDlandmDlare then still connected by a series of unit steps. In this case, also, there will be 2lC1different mvalues. This quantity, 2lC1, is sometimes called the multiplicity of the angular momentum states. The fact that it is mathematically possible to have a series (multiplet) of states corre- sponding to either integral or half-integral land satisfying the angular momentum commu- tation rules does not prove that such states are realizable in a particular algebraic system (such as that describing ordinary three-dimensional [3-D] space), or that such states are indeed relevant for physics. However, by solving Laplace’s equation, we have already found that the angular momentum states of integral lcan be described in ordinary space and that they can be identified as states of ordinary (so-called orbital) angular momentum. It is not possible to describe states of half-integral las ordinary functions in 3-D space, so orbital angular momentum will only involve integral l. Example 16.1.1 SPHERICAL HARMONICS LADDER From Exercise 3.10.30, or alternatively by combining the formulas for LxandLyfrom Eq. (16.13), the orbital angular momentum ladder operator LCis found to be LCDei'@ @Cicot@ @' : Starting from Y0 1.;'/Dp3=4 cos, we can apply Eq. (16.25): LCY0 1.;'/Dr 3 4ei'@cos @Dr 3 4ei'.sin/Dp 2Y1 1.;'/; (16.26) which when solved for Y1 1gives (with proper scale and sign) the value tabulated in Table 15.4. The reader can verify that the application of LCtoY1 1gives zero.  Spinors It turns out that half-integral angular momentum states are needed to describe the intrinsic angular momentum of the electron and many other particles. Since these particles also have magnetic moments, an intuitive interpretation is that their charge distributions are spinning about some axis; hence the term spin. It is now understood that the spin phenomena can- not be explained consistently by describing these particles as ordinary charge distributions undergoing rotational motion, but are better treated by assigning these particles to states in an abstract space that, for the electron, has the lvalue1 2(but in this context, we normally usesand write sD1 2), which means that the possible mvalues (often written ms) are msDC1 2andmsD1 2. It is not productive to try to think of this situation in terms of ArfKen_Ch16-9780123846549.tex 780 Chapter 16 Angular Momentum ordinary functions, but to accept an abstract formulation in which spin states are repre- sented by symbols; popular choices are orj"ifor the state mDC1 2and orj#ifor that with mD1 2. These spin states can also be represented by two-component column vectors, with the angular momentum operators given in terms of the Pauli matrices as1 2i. The quantities forming a basis for the multiplets for half-integer angular momentum are called spinors. In addition to their manipulation using ladder operators, they have rotational properties that are discussed in more detail in Chapter 17. Example 16.1.2 SPINOR LADDER Calling the angular momentum operator S, we write Sx,Sy,Szas the 22matrices1 2i, whereiare defined in Eq. (2.28): SxD1 2 0 1 1 0! ;SyD1 2 0i i0! ;SzD1 2 1 0 01! : (16.27) By carrying out matrix operations we can verify that these matrices satisfy the angular momentum commutation rules. For example, SxSySySxD1 4 0 1 1 0! 0i i0! 1 4 0i i0! 0 1 1 0! Di 2 1 0 01! Di Sz: We also find that S2 xDS2 yDS2 zD.1=4/1; and we therefore have S2DS2 xCS2 yCS2 zD1 4h 1C1C1i D3 41: Note that 3=4isS.SC1/forSD1=2. The interpretation of these matrix relationships is that we have an abstract space spanned by the two functions 1=2 1=2 j"iD 1 0! ; 1=2 1=2 j#iD 0 1! ; ArfKen_Ch16-9780123846549.tex 16.1 Angular Momentum Operators 781 and that the operators S2andSzoperate on these functions as follows: S2 DS2 1=2 1=2D3 4 1 0 0 1! 1 0! D3 4 1 0! D3 4 1=2 1=2D3 4 ; Sz DSz 1=2 1=2D1 2 1 0 01! 1 0! D1 2 1 0! D1 2 1=2 1=2D1 2 ; S2 DS2 1=2 1=2D3 4 1 0 0 1! 0 1! D3 4 0 1! D3 4 1=2 1=2D3 4 ; Sz DSz 1=2 1=2D1 2 1 0 01! 0 1! D1 2 0 1! D1 2 1=2 1=2D1 2 : The above formulas show that 1=2 1=2(also denoted and ) are simultaneous eigen- functions of S2andSz. To illustrate that they are not also eigenfunctions of SxorSy, we compute Sx DSx 1=2 1=2D1 2 0 1 1 0! 1 0! D1 2 0 1! D1 2 1=2 1=2D1 2 : To make ladders, we now form SCDSxCi SyD 0 1 0 0! ;SDSxi SyD 0 0 1 0! : Applying these operators to D 1=2 1=2, SC D 0 1 0 0! 1 0! D0; S D 0 0 1 0! 1 0! D 0 1! D : These results are in agreement with Eq. (16.25), for which, with the current parameters D3=4,mD1=2, its coefficients are p m.mC1/D0;p m.m1/D1:  Summary, Angular Momentum Formulas The analysis of the preceding subsection applies to any system of operators satisfying the angular momentum commutation rules. Possible areas of application include orbital angu- lar momentum (for which the eigenfunctions are the spherical harmonics), the intrinsic (spin) angular momentum we now know is associated with most fundamental particles, and even the overall angular momenta that result either from considering both the orbital and spin angular momenta of the same particle, or the total angular momentum of a collec- tion of particles (as in a many-electron atom or even a nucleus). ArfKen_Ch16-9780123846549.tex 782 Chapter 16 Angular Momentum It is useful to summarize the key results; we do so giving the operators the name J, to emphasize the fact that the results are not restricted to orbital angular momentum (for which the symbol Lis nearly universally used), or to spin angular momentum (traditionally denoted S). Thus: 1. We assume that there exists a Hermitian operator Jwith components Jx,Jy,Jzsuch thatJ2 xCJ2 yCJ2 zDJ2and that these quantities satisfy the commutation relations TJk;JlUDi"klnJn;TJ2;JkUD0; (16.28) where k;l;narex;y;zin any order and "klnis a Levi-Civita symbol. Other than the requirement of Eq. (16.28), Jis arbitrary. 2. Because the operators J2andJzcommute, there can exist functions (in some abstract space), generically denoted M J, with M Jsimultaneously a normalized eigenfunction ofJzwith eigenvalue Mand an eigenfunction of J2with eigenvalue J.JC1/: Jz M JDM M J;J2 M JDJ.JC1/ M J;D M Jj M JE D1: (16.29) 3. Operators satisfying the above conditions can be called angular momentum opera- tors; those which were used as examples of angular momentum in ordinary space (orbital angular momentum) are clearly relevant for physics; similar operators in more abstract spaces are relevant only to the extent that they can be identified with physical phenomena. We have already seen that these assumptions are sufficient to enable the introduction of ladder operators, and to reach the following conclusions: 1. The possible values of Jare integral and half-integral; in ordinary 3-D space only functions of integral Jcan be realized. 2. For a given J, the possible values of Mrange in unit steps from MDJtoMDJ; this produces 2JC1different Mvalues. 3. Given any one M J, we can generate others by use of the operators JCDJxCi Jy;JDJxi Jy: The result of applying these operators to M Jis, see Eq. (16.25), JC M JDp .JM/.JCMC1/ MC1 J; (16.30) J M JDp .JCM/.JMC1/ M1 J: (16.31) These formulas give zero results when JCis applied to J Jand when Jis applied to J J. Exercises 16.1.1 The quantum mechanical angular momentum operators Lxi Lyin 3-D physical space are given by ArfKen_Ch16-9780123846549.tex 16.1 Angular Momentum Operators 783 LxCi LyDei'@ @Cicot@ @' ; Lxi LyDei'@ @icot@ @' : Show that (a).LxCi Ly/YM L.;'/Dp .LM/.LCMC1/YMC1 L.;'/; (b).Lxi Ly/YM L.;'/Dp .LCM/.LMC1/YM1 L.;'/: 16.1.2 With Lgiven by LDLxi LyDei'@ @icot@ @' ; show that (a) Ym lDs .lCm/W .2l/W.lm/W.L/lmYl l, (b) Ym lDs .lm/W .2l/W.lCm/W.LC/lCmYl l. 16.1.3 Using the known forms of LCandL(Exercise 16.1.2), show that Z TYM LUL.LCYM L/dDZ .LCYM L/.LCYM L/d: Here dis the element of solid angle .sindd'/, and the integration is over the entire angular space. 16.1.4 (a) Show that J2D1 2h JCJCJJCi CJ2 z. (b) Use the result from part (a) and the explicit formulas for LCandLfrom Exer- cise 16.1.2 to verify that all the spherical harmonics with lD2are eigenfunctions ofL2with eigenvalue l.lC1/D6. 16.1.5 Derive the following relations without assuming anything about M Lother than that they are angular momentum eigenfunctions: (a) M L.;'/Ds .LCM/W .2L/W.LM/W.L/LM L L.;'/ , (b) M L.;'/Ds .LM/W .2L/W.LCM/W.LC/LCM L L.;'/ . ArfKen_Ch16-9780123846549.tex 784 Chapter 16 Angular Momentum 16.1.6 Derive the operator equations .LC/nYM L.;'/D.1/nein'sinnCMdnsinMYM L.;'/ .dcos/n; .L/nYM L.;'/Dein'sinnMdnsinMYM L.;'/ .dcos/n: Hint. Try mathematical induction (Section 1.4). 16.1.7 Show, using.L/n, that YM L.;'/D.1/Mh YM L.;'/i : 16.1.8 Verify by explicit calculation that (a) LCY0 1.;'/Dr 3 4sinei'Dp 2Y1 1.;'/ , (b) LY0 1.;'/DCr 3 4sinei'Dp 2Y1 1.;'/ . The signs have the indicated values because the spherical harmonics were defined to be consistent with the results obtained using the ladder operators LCandL(Condon- Shortley phase). 16.2 A NGULAR MOMENTUM COUPLING An important application of ladder operators is to systems in which a resultant angular momentum is the sum of two individual angular momenta. Because the angular momenta have directional properties, we anticipate a result that has some properties in common with vector addition, but because these are quantum mechanical quantities involving non- commuting operators, we need to study the problem in more detail. Ifj1andj2are two individual angular momentum operators that act on different coor- dinate sets (as, e.g., the coordinates of two different particles), then they are unrelated and all components of each must commute with every component of the other. This will enable us to carry out a detailed analysis of operators of the combined system, for which the total angular momentum operator is JDj1Cj2, with components JxDj1xCj2x, JyDj1yCj2y,JzDj1zCj2z, with the overall operator J2DJ2 xCJ2 yCJ2 z. To discuss the problem, we will need the commutators Tj1k;j1lUDi"klnj1n;Tj2k;j2lUDi"klnj2n;Tj1k;j2lUD0: (16.32) For the first two commutators, k;l;narex;y;zin any order; the third commutator vanishes for all k;lincluding kDl. From the commutators in Eq. (16.32), it is easily established that the overall angular momentum components obey the commutation rules TJx;JyUDi Jz;TJy;JzUDi Jx;TJz;JxUDi Jy; (16.33) ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 785 so these overall components satisfy the generic angular momentum commutation relations, meaning also that TJ2;JiUD0: (16.34) In addition, TJ2;j2 1UDTJ2;j2 2UD0: (16.35) However, it is nottrue that the components of j1orj2, namely j1iorj2i, commute with J2, even though J2and the sum j1iCj2ido commute. Example 16.2.1 COMMUTATION RULES FOR JCOMPONENTS To find the commutator TJ2;j1zU, write J2D.j1Cj2/2Dj2 1Cj2 2C2j1j2 Dj2 1Cj2 2C2 j1xj2xCj1yj2yCj1zj2z ; so we have TJ2;j1zUDTj2 1;j1zUCTj2 2;j1zUC2 Tj1xj2x;j1zUCT j1yj2y;j1zUCT j1zj2z;j1zU D2 j2xTj1x;j1zUCj2yTj1y;j1zU D2i.j1xj2yj1yj2x/; (16.36) where we have dropped terms in which the commutators involve different particles and those, e.g.,Tj2 1;j1zU, which vanish because the individual-particle operators are angular momenta. Equation (16.36) clearly shows thatTJ2;j1zUis nonzero. However, its contributions are equal and opposite to those of TJ2;j2zU, explaining whyTJ2;JzUdoes vanish. Consider nextTJ2;j2 1U. Again expanding J2, we get TJ2;j2 1UDTj2 1;j2 1UCTj2 2;j2 1UC2 Tj1xj2x;j2 1UCT j1yj2y;j2 1UCT j1zj2z;j2 1U : Every term of this equation vanishes, so J2andj2 1commute.  We have noted that j2 1,j2 2, and Jzall commute with each other and with J2,j1z, and j2z, but that the last three of these operators do not all commute with each other. There are therefore different ways of selecting maximal sets of mutually commuting operators for which we can construct simultaneous eigenfunctions. One possibility is to select j2 1, j2 2,j1z,j2z, and Jz, which has the advantage that the simultaneous eigenfunctions are just products of the eigenstates for individual ji, but has the disadvantage that we will not have states of definite total angular momentum J2. This is a big disadvantage, because in reality different angular momenta in the same system actually interact to some extent. If we add to the Hamiltonian of our problem a small term (a perturbation) that causes the individual angular momenta not quite to be independent, our system will still strictly have conservation of J2(i.e., HandJ2will still commute), but the perturbation added to the Hamiltonian will not commute with j1andj2. ArfKen_Ch16-9780123846549.tex 786 Chapter 16 Angular Momentum Alternatively, and for most purposes better, we could choose the mutually commut- ing operator set J2,j2 1,j2 2, and Jz, which would describe states of definite total angular momentum, but these states would be mixtures of the individual angular-momentum states and would not have definite values of j1zorj2z. It is the purpose of this section to relate these two descriptions by finding the equations that connect (i.e., couple) the individual angular momenta to form states of definite J2. To simplify future discussion, we can refer to the product basis of the preceding para- graph as the m1;m2basis, and call the alternative basis of definite JtheJ;Mbasis. The m1;m2basis members also have definite values of M, but not J; most members of the J;Mbasis will not have definite values of either m1orm2. If we stick with problems in which j1andj2are fixed, all members of both bases will have the same definite values of these quantum numbers. Before getting into the details, let’s make two observations. First, since we have raising and lowering operators that we can apply to the J;Mbasis, the functions in this basis must include all the Mvalues for any Jthat is present at all. Second, since both bases have definite values of M, the transition from one basis to the other cannot mix functions of different M. Vector Model We begin with some qualitative observations. Since JzDj1zCj2z(with eigenvalues we callM) is part of both our commuting operator sets, we can conclude, from looking at them1;m2basis, that the maximum eigenvalue MmaxofJzwill occur when m1Dj1and m2Dj2, soMmaxDj1Cj2. Moving now to the J;Mbasis, which of course spans the same function space, we see that because Mmaxis the maximum Mvalue, it must be a member of a multiplet with JDMmax, and this must be the largest possible J. Thus, JmaxDj1Cj2. To establish the minimum value possible for the quantum number Jis a little trickier, and we will come back to that shortly. The result, which is simple, is that JminDjj1j2j. These maximum and minimum values of Jcorrespond to the notion that the classical vec- tor sum j1Cj2has a maximum length equal to the sum of the lengths of these vectors and a minimum length equal to the absolute value of their difference; the quantum analog of this notion is not quantitatively exact because the magnitude of each jis actuallypj.jC1/. Further qualitative observations follow if we tabulate the various possible m1;m2func- tions of various Mvalues. The concept can be understood from a simple example. Suppose j1D2,j2D1. Then the members of the m1;m2basis can be grouped as shown here. The kets in the table are labeled in more detail than usual to avoid potential confusion; those labeled m1have jvalue j1, those labeled m2have jDj2. MDC3jm1DC2ijm 2DC1i MDC2jm1DC2ijm 2D0i jm 1DC1ijm 2DC1i MDC1jm1DC2ijm 2D1i jm 1DC1ijm 2D0i jm 1D0ijm2DC1i MD0jm1DC1ijm 2D1i jm 1D0ijm2D0i jm 1D1ijm 2DC1i MD1jm1D2ijm 2DC1i jm 1D1ijm 2D0i jm 1D0ijm2D1i MD2jm1D2ijm 2D0i jm 1D1ijm 2D1i MD3jm1D2ijm 2D1i ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 787 Because the basis transformations we are discussing only mix basis functions of the same M, a transition to the J;Mbasis will have the same number of functions of each M as are in the row of our table for that M. If there is only one function in the row, it must (without change) be a member of the J;Mbasis, and we may use it as a starting point for getting all the other members of the multiplet for the same Jby application of the ladder operators. So in our current example, we can start from jm1DC2ijm 2DC1i , and make one member of the multiplet for each Mvalue in the table. Once this has been done, we will have constructed as many J;Mfunctions as there are entries in the first column of the table (but remember that in most cases they will not be the specific functions sitting in that column). But that observation does tell us that the numbers of functions that are still unused (but not their exact forms) will correspond with the numbers of functions in the remainder of the table. In particular, we see that there will in our example be one function left over with MD2. Because it cannot have a JD3 component, it must be a jJD2;MD2ieigenfunction and therefore must be orthogonal to thejJD3;MD2ifunction we have already found. That means that we can obtain it by Gram-Schmidt orthogonalization within the function space for MDC2 . From thejJD2;MD2ifunction, we can apply a ladder operator to find jJD2;Mi basis members with other Mvalues, the number of which will correspond to the number of entries in the second column of our table. To continue to a third column, we would need to find a MDC1 function orthogonal to both thejJD3;MDC1i andjJD2;MDC1i functions. This process can be continued until the m1;m2basis has been exhausted. Taking now a further look at our table, we see that the number of columns with entries increases as we decrease Mfrom its maximum value until Mhas reachedjj1j2j; for smallerjMjthan that, the number of columns in use stays constant, because of limitations in the way the individual mvalues can be chosen to add up to M. That gives us a graphical indication that the smallest Jvalue will bejj1j2j. A more algebraic way of determining the smallest resultant Jis based on a computa- tion of the total number of J;Mstates generated if the possible Jvalues run from an as yet undetermined value Jminto our previously determined maximum value Jmax. Since the number of states for each Jis2JC1, the total number of J;Mstates we will have produced is JDJmaxX JDJmin.2JC1/D.JmaxJminC1/.JmaxCJminC1/ D.2j1C1/.2j2C1/; (16.37) where the second line of this equation reflects the fact that the total number of states is readily counted in the m1;m2basis. Inserting the value JmaxDj1Cj2and solving for Jmin, we find JminDjj1j2j: Another way of stating this result is to observe that the possible values of Jsatisfy a triangle rule, meaning that they occur in unit steps from a maximum of j1Cj2to a minimum ofjj1j2j. ArfKen_Ch16-9780123846549.tex 788 Chapter 16 Angular Momentum Ladder Operator Construction To develop a quantitative description of angular momentum coupling, we consider the case of general j1and j2, and start from the lone member of the m1;m2basis with MDj1Cj2. In line with our earlier discussion, this m1;m2basis member must also be a function of the definite Jvalue JmaxDj1Cj2. Using a notation in which the lower entry in each ket is its Jvalue and the upper entry gives the value of M, we indicate this by writing Jmax Jmax D j1 j1 j2 j2 : (16.38) We now generate additional states of the same Jbut with different Mby applying the lowering operator Jto Eq. (16.38); when we apply it to the right-hand side, we do so in the form JDj1Cj2. The result, for the left side of Eq. (16.38), is J Jmax Jmax Dp 2Jmax Jmax1 Jmax : (16.39) The coefficientp2Jmaxis that given by Eq. (16.31) forJDMDJmax. For the right side ofEq. (16.38), we get .j1Cj2/2 4 j1 j1 j2 j23 5D2 4j1 j1 j13 5 j2 j2 C j1 j12 4j2 j2 j23 5 Dp 2j1 j11 j1 j2 j2 Cp 2j2 j1 j1 j21 j2 ; (16.40) where we have again obtained the coefficients from Eq. (16.31), but now evaluating them for the first term with .J;M/D.j1;j1/and for the second term with .J;M/D.j2;j2/. Combining these results, and solving for Jmax1 Jmax , Jmax1 Jmax Ds j1 Jmax j11 j1 j2 j2 Cs j2 Jmax j1 j1 j21 j2 : (16.41) With escalating complexity, we could continue this process to smaller values of M. As indicated in our earlier, more qualitative discussion, we can reach functions with JDJmax1by starting from the unused member of the set of two functions j11 j1 j2 j2 ; j1 j1 j21 j2 : The quantity we seek, Jmax1 Jmax1 ; ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 789 will be the function in the above-defined subspace that is orthogonal to Jmax1 Jmax as given in Eq. (16.41), and therefore will be Jmax1 Jmax1 Ds j2 Jmax j11 j1 j2 j2 Cs j1 Jmax j1 j1 j21 j2 : (16.42) At this point we note that the function produced by Eq. (16.42) could have been written with all its signs changed, as the orthogonalization process does not determine the sign of the orthogonal function. This only matters if we wish to correlate the signs of our J;M constructions with work by others. Irrespective of our choice of signs, we can apply Jto reach the full set of Mvalues, and then continue to states of smaller Juntil the m1;m2 space is exhausted. The general result of the above-described processes is to obtain each J;Meigenstate as a linear combination of m1;m2states of the same M, in a fashion summarized by the following equation (written in a less cumbersome notation now that the need for detail has disappeared): jJ;MiDX m1;m2C.j1;j2;Jjm1;m2;M/jj1;m1Ij2;m2i: (16.43) Herejj1;m1Ij2;m2istands forjj1;m1ijj2;m2iand we have given over to the coeffi- cient C.j1;j2;Jjm1;m2;M/the responsibility to vanish when m1Cm26DM. Thus, the apparent double summation in Eq. (16.43) is actually a single sum. The coefficients in Eq. (16.43) are called Clebsch-Gordan coefficients. To resolve the sign ambiguity result- ing from the orthogonalization processes, they are defined to have signs specified by the Condon-Shortley phase convention. It is important to realize that all the results of this section remain valid irrespective of whether j1,j2, or both are integral or half-integral. For example, if j1D1and j2D1 2 (corresponding to the coupling of the orbital and spin angular momenta of an electron), the possible J;Mstates will be a quartet forJD3=2(with Mvalues +3/2, +1/2,1/2, 3/2), and a doublet forJD1=2(with Mvalues +1/2 and1/2). A second way to look at the Clebsch-Gordan coefficients is to identify them as the scalar products C.j1;j2;Jjm1;m2;M/DhJ;Mjj1;m1Ij2;m2i: (16.44) Because of the method used for the construction of the jJ;Mi, we can make one additional observation: The Clebsch-Gordan coefficients will all be real, even if the jj1;m1iand jj2;m2iused for their construction are not. The Clebsch-Gordan expansion can be interpreted in yet another way. We can view the Clebsch-Gordan coefficients as the elements of a transformation matrix converting func- tions of the m1;m2basis into those of the J;Mbasis; since both basis sets are orthonormal, the transformation must be unitary (and because it is real, orthogonal). This means that the inverse transformation, .J;M/!.m1;m2/, must have a transformation matrix that is the ArfKen_Ch16-9780123846549.tex 790 Chapter 16 Angular Momentum transpose of that for the forward transformation .m1;m2/!.J;M/. That means that we also have the equation jj1;m1Ij2;m2iDX J MC.j1;j2;Jjm1;m2;M/jJ;Mi: (16.45) This equation is correct and corresponds to our discussion. Note that instead of reversing the index order of the transformation matrix we have interchanged the index sets identify- ing the functions. In passing, we make one further comment. While the Clebsch-Gordan coefficients can be identified as forming a transformation matrix, note that their row/column indexing differs from the pattern to which we are accustomed, since, instead of labels running from 1 to n (the dimension of the transformation), we are using in one dimension the compound index .m1;m2/, and in the other dimension the compound quantity .J;M/. This Clebsch-Gordan matrix will be somewhat sparse (containing many zero elements). The zeros occur because the coefficients vanish unless MDm1Cm2. There is a significant literature on the practical computation of Clebsch-Gordan coeffi- cients,1but to make the present discussion complete we simply give here a closed general formula: C.j1;j2;Jjm1;m2;M/DF1F2F3; (16.46) where F1Ds .j1Cj2J/W.JCj1j2/W.JCj2j1/W.2JC1/ .j1Cj2CJC1/W F2Dq .JCM/W.JM/W.j1Cm1/W.j1m1/W.j2Cm2/W.j2m2/W; F3DX s.1/s .j1m1s/W.j2Cm2s/W.Jj2Cm1Cs/W 1 .Jj1m2Cs/W.j1Cj2Js/WsW: The F3summation is over all integer values of sfor which the factorials all have non- negative arguments (which will be integral). The sum is therefore finite in extent and F3 is a closed form. Equation (16.46) is only to be used for parameter values that satisfy the angular momentum and coupling conditions: j1,j2,Jmust satisfy the triangle condition, miis to be from the sequence li,li1;:::;li(iD1;2),Mto be from J;J1;:::;J, andMDm1Cm2. Finally, we call attention to the fact that Clebsch-Gordan coefficients have symmetries that are not obvious from the foregoing development. To expose the symmetries, it is convenient to convert them to the Wigner 3 j-symbols, defined as j1j2j3 m1m2m3 D.1/j1j2m3 .2j3C1/1=2C.j1;j2;j3jm1;m2;m 3/: (16.47) 1See Biedenharn and Louck, Brink and Satchler, Edmonds, Rose, and Wigner in Additional Readings. Clebsch-Gordan coeffi- cients are also tabulated in many places, and can easily be found online by a Web search. ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 791 Extensive discussion of 3j-symbols and related quantities is beyond the scope of this book. This important, but advanced, topic is presented in most of the sources listed under Addi- tional Readings. The3j-symbols are invariant under even permutations of the indices (1,2,3), but under odd permutations .1;2;3/!.k;l;n/transform as follows: j1j2j3 m1m2m3 D.1/j1Cj2Cj3jkjljn mkmlmn : (16.48) They also have the following symmetry under change of sign of their lower indices: j1j2j3 m1m2m3 D.1/j1Cj2Cj3j1 j2 j3 m 1m 2m 3 : (16.49) Even though some of the jimay be half-integral, remember that j3must be equal to j1Cj2 or differ therefrom by an integer. This fact causes the powers of 1inEqs. (16.47) through (16.49) to be integral, so these factors are not multiple-valued and the sign assignments of the3j-symbols are unambiguous. These symmetry relations make a table of 3j-symbols more compact than one of Clebsch-Gordan coefficients; such tables can be found in the literature,2and a short list is included here, as Table 16.1. We close this section with two examples. Table 16.1 Wigner 3 j-Symbols 0 B@1 21 21 1 21 201 CAD1p 60 B@1 21 21 1 21 211 CAD1p 30 B@1 21 20 1 21 201 CAD1p 2 0 B@11 23 2 11 23 21 CAD1 20 B@11 23 2 11 21 21 CAD1p 120 B@11 23 2 01 21 21 CAD1p 6 1 1 0 0 0 0! D1p 3 1 1 0 11 0! D1p 3 1 1 1 0 0 0! D0 1 1 1 11 0! D1p 6 1 1 2 0 0 0! Dr 2 15 1 1 2 11 0! D1p 30 1 1 2 1 01! D1p 10 1 1 2 1 12! D1p 5 1 2 2 11 0! D1p 10 1 2 2 12 1! D1p 15 1 2 2 0 11! D1p 30 1 2 2 0 22! Dr 2 15 1 2 2 0 0 0! D0 1 2 3 0 0 0! Dr 3 35 1 2 3 11 0! D1p 35 1 2 3 1 01! Dr 2 35 1 2 3 12 1! D1p 105 1 2 3 1 12! Dr 2 21 1 2 3 1 23! D1p 7 1 2 3 0 11! Dr 8 105 1 2 3 0 22! D1p 21 2See, for example, M. Rotenberg, R. Bivins, N. Metropolis, and J. K. Wooten, Jr., The 3j- and 6j-Symbols. Cambridge, MA: Massachusetts Institute of Technology Press (1959). ArfKen_Ch16-9780123846549.tex 792 Chapter 16 Angular Momentum Example 16.2.2 TWO SPINORS This example describes a problem that exists entirely in a abstract space, namely the cou- pling of two spin-1 2particles (e.g., electrons) to form combined states of definite J. Letting stand for a normalized single-particle state with jD1 2,mDC1 2, with a normalized state with jD1 2,mD1 2, we have the following four states in the m1;m2basis: MD1V MD0V MD1V For all these two-particle states, the first symbol refers to particle 1, the second to particle 2. From Eq. (16.31) andExample 16.1.2, we have j D ,j D0, and we can use the following rearrangement of Eq. (16.31) to deal with thejJ;Mistates. Again we use a notation in which the lower entry in the ket is J; the upper entry is M: M1 J D1p.JCM/.JMC1/J M J : (16.50) The maximum Mvalue in this system is MDC1, so the one state of this Mvalue must have JD1. showing that 1 1 D . Starting from it, we lower M: 1 1 D ; 0 1 D1p 2J 1 1 D1p 2 C1p 2 ; 1 1 D1p 2J 0 1 D1p 21p 2 C1p 2  D : These are the well-known members of the SD1spin multiplet, which is known as a triplet. At MD0, where there were two m1;m2states, the state orthogonal to 0 1 must be the 0 0 state. Even though we do not have an entirely explicit representation of the states and , we do know that they are normalized eigenstates of a Hermitian operator ( Jz) with different eigenvalues, and therefore they must be orthogonal. Thus, we can apply the Gram-Schmidt process to the MD0subspace, using the relations h j iDh j iD1;h j iD0: We easily find that the normalized function orthogonal to . C /=p 2is. /=p 2. It is the only member of the SD0multiplet, and is therefore known as a singlet. Note that we didn’t have to know anything specific about spin operators to carry out this analysis. ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 793 Our tableau of states can now be written in the J;Mbasis: JD1 JD0 MD1 MD0. C /=p 2. /=p 2 MD1 From the J;Mtableau, we can read out the Clebsch-Gordan coefficients: C1 2;1 2;1 1 2;1 2;1 D1 C1 2;1 2;1 1 2;1 2;0 DC1 2;1 2;1 1 2;1 2;0 D1p 2 C1 2;1 2;0 1 2;1 2;0 DC1 2;1 2;1 1 2;1 2;0 D1p 2 C1 2;1 2;1 1 2;1 2;1 D1 These coefficients can also be obtained from our table of 3 j-symbols. Using Eq. (16.47), we find the coefficients for jJD1;MD0ito be C1 2;1 2;1 1 2;1 2;0 Dp 3 1 21 21 1 21 20! ; C1 2;1 2;1 1 2;1 2;0 Dp 3 1 21 21 1 21 20! : Both these 3 j-symbols correspond to the same entry in Table 16.1, and the symmetry rules give each the valueC1=p 6. Therefore, both these Clebsch-Gordan coefficients evaluate top 3=p 6, or, as expected, 1=p 2. ForjJD0;MD0i, we have, again calling on Eq. (16.47), C1 2;1 2;0 1 2;1 2;0 D 1 21 20 1 21 20! ; C1 2;1 2;0 1 2;1 2;0 D 1 21 20 1 21 20! : Again these 3 j-symbols both correspond to the same tabulated entry (with value 1=p 2), but this time the symmetry rules cause them to have the respective values C1=p 2and 1=p 2, in agreement with our explicit evaluation.  Example 16.2.3 COUPLING OF pANDdELECTRONS As most physics students know, a pstate is an angular momentum eigenstate with lD1(so mcan be 1, 0, or1). The three normalized functions constituting its multiplet are often denoted pC,p0, and p. Adstate has lD2; we denote the five normalized members of its ArfKen_Ch16-9780123846549.tex 794 Chapter 16 Angular Momentum multiplet dC2,dC,d0,d, and d2. The m1;m2basis has 15 members; grouped according to their Mvalues, they consist of MDC3 pCdC2 MDC2 pCdCp0dC2 MDC1 pCd0p0dCpdC2 MD0 pCdp0d0pdC MD1 pCd2p0dpd0 MD2 p0d2 pd MD3 pd2 This is the same coupling of angular momenta jD1andjD2that was introduced at the beginning of the subsection entitled Vector Model, but we are now illustrating how to carry out the coupling computations using Clebsch-Gordan coefficients and 3j-symbols. From this diagram, we expect one multiplet with JD3, which in atomic spectroscopy is denoted F(multiparticle orbital angular momentum states are designated using upper-case letters); one with JD2(called D), and one with JD1(called P). Our plan is to construct these using the 3 j-symbols given in Table 16.1. We start by writing, in the notation jJ;Mi, the members of the Fmultiplet with M1 in terms of Clebsch-Gordan coefficients (those for M<1do not raise important new points): j3;3iD C.1;2;3j1;2;3/pCdC2; j3;2iD C.1;2;3j1;1;2/pCdCCC.1;2;3j0;2;2/p0dC2; j3;1iD C.1;2;3j1;0;1/pCd0CC.1;2;3j0;1;1/p0dCCC.1;2;3j1;2;1/pdC2: TheDandPmultiplet members for M1are j2;2iD C.1;2;2j1;1;2/pCdCCC.1;2;2j0;2;2/p0dC2; j2;1iD C.1;2;2j1;0;1/pCd0CC.1;2;2j0;1;1/p0dCCC.1;2;2j1;2;1/pdC2; j1;1iD C.1;2;1j1;0;1/pCd0CC.1;2;1j0;1;1/p0dCCC.1;2;1j1;2;1/pdC2: We then express the Clebsch-Gordan coefficients in terms of 3 j-symbols. Doing just a rep- resentative few, using Eq. (16.47) and then the symmetry rules, Eqs. (16.48) and (16.49), C.1;2;3j1;2;3/DCp 7 1 2 3 1 23! D1; C.1;2;2j1;1;2/Dp 5 1 2 2 1 12! DCp 5 1 2 2 12 1! Dr 1 3; C.1;2;1j1;2;1/Dp 3 1 2 1 1 21! Dp 3 1 1 2 1 12! Dr 3 5: ArfKen_Ch16-9780123846549.tex 16.2 Angular Momentum Coupling 795 Substituting these and other Clebsch-Gordan coefficients into the formulas for jJ;Mi, we obtain the final results: j3;3iD pCdC2; j3;2iDr 1 3p0dC2Cr 2 3pCdC; j3;1iDr 1 15pdC2Cr 8 15p0dCCr 2 5pCd0; j2;2iDr 2 3p0dC2Cr 1 3pCdC: j2;1iDr 1 3pdC2r 1 6p0dCCr 1 2pCd0; j1;1iDr 3 5pdC2r 3 10p0dCCr 1 10pCd0: The reader may verify that states of the same Mbut different Jhave the required orthog- onality. It is also easy to check that all these jJ;Mistates are normalized.  Exercises 16.2.1 Derive recursion relations for Clebsch-Gordan coefficients. Use them to calculate C.11Jjm1m2M/forJD0;1;2. Hint. Use the known matrix elements of JCDJ1CCJ2C;JiC, and J2D.J1CJ2/2, etc. 16.2.2 Defining.Yl/M Jby the formula .Yl/M JDX C.l1 2JjmlmsM/Ylmlms; where1=2 are the spin up and down eigenfunctions of 3Dz, show that.Yl/M Jis aJ;Meigenfunction. 16.2.3 Find the.j;m/states of a pelectron ( lD1), in which the orbital angular momen- tum of the electron is coupled to its spin angular momentum ( sD1=2) to form states whose conventional labelings are2p1=2and2p3=2. The notation is of the general form 2sC1.symbol/ j, where “symbol” is that indicating the lvalue (i.e., s,p;:::). 16.2.4 Repeat Exercise 16.2.3 for lD1,sD3=2. Apply the conventional labels to the j;m states. 16.2.5 A deuterium atom consists of a proton, a neutron, and an electron. Each of these par- ticles has spin 1/2. The coupling of these three spins can produce Jvalues of 3/2 and 1/2. We consider here only states with no orbital angular momentum. (a) Show that these J;Mstates consist of one quartet ( JD3=2) and two linearly independent doublets ( JD1=2). ArfKen_Ch16-9780123846549.tex 796 Chapter 16 Angular Momentum Hint. Make a vector-model diagram. (b) One way to analyze this problem is to couple the spins of the proton and neutron to form a nuclear triplet or singlet, and then to couple the resultant nuclear spin to the electron spin. Find the states that are obtained in this way (designate the single-particle states p ,p ,n ,n ,e ,e ). (c) Another way to analyze this problem is to couple the spins of the proton and elec- tron to form an atomic triplet or singlet, and then to couple that resultant to the neutron spin. Find the states that result from this coupling scheme. (d) Show that the coupling schemes of parts (b) and (c) span the same Hilbert space. Note. The actual interaction energies among these angular momenta cause the scheme of part (b) to be the better way of treating this problem (the triplet nuclear state is substantially the more stable), and the system actually looks like a spin-1 deuterium nucleus plus an electron. 16.3 S PHERICAL TENSORS We have already seen that the set of spherical harmonics of given ltransforms within itself under rotations. We now pursue this idea more formally. In Chapter 3 we saw that rota- tions could be characterized by the 33unitary transformation matrices that transform a set of coordinates (their basis) into the new set corresponding to the rotation. These matri- ces could be viewed as second-rank tensors, but because they are restricted to rotational transformations, they are also known as spherical tensors. We now wish to consider spherical tensors that transform more general sets of objects under rotation, and in particular those spherical tensors that have spherical harmonics as bases. Our new spherical tensors will then have dimensions other than 33; in fact, they must exist at all the sizes that correspond to sets of angular momentum eigenfunctions. Because we have already observed that a set of angular momentum eigenfunctions of a given Jcannot be decomposed into subsets that transform only among themselves under rotation, we go one step further and call our spherical tensors irreducible. Continuing for general angular momentum eigenfunctions jL;Mi, which we assume are representable in 3-D space as spherical harmonics or objects built from them by angular momentum coupling, we write the following defining equation for the spherical tensor describing the effect of a coordinate rotation RonjL;Mi: RjL;MiDX M0DL M0M.R/jL;M0i: (16.51) If thejL;Miare actually spherical harmonics (and not more complicated objects that resulted from angular momentum coupling), Eq. (16.51) can also be written as Ym l.R/DX m0Dl m0m.R/Ym0 l./: (16.52) Because we do not need to become embroiled in the details of the action of Ron the coordinates, we have simply replaced .;'/ by the generic symbol and have written Rto indicate the coordinates .0;'0/that describe the point that was labeled .;'/ in the unrotated system. For any given l,Dl m0m.R/can be regarded as an element of a square ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 797 matrix of dimension 2lC1with rows and columns labeled by indices m0andmwhose ranges are.l;:::;Cl/, not the more customary sequence starting from 1. The Dl m0m.R/ are unitary, since they describe a transformation between two orthonormal sets. Because of their early exploitation by Eugene Wigner, they are sometimes called Wigner matrices. There is an extensive literature (see Additional Readings) on relationships satisfied by the Dl m0m.R/and on formulas for their evaluation. A related topic included in this book is the formula, Eq. (3.37), giving the transformation of the basis x;y;zby a rotation through Euler angles , , . Addition Theorem Equation (16.52) can be used to establish important rotational invariance properties. For example, consider a quantity Adefined as ADX mYm l.1/Ym l.2/; (16.53) where1and2are two unrelated sets of angular coordinates. We apply a rotation Rto the coordinate system, denoting the result RA, and evaluating the right-hand side using Eq. (16.52): RADX m X Dl m.R/Y l.1/! X Dl m.R/Y l.2/! : (16.54) We now reorder the summations in Eq. (16.54), and, in the second line of Eq. (16.55), use the fact that Dis unitary to change Dto the transpose of D1, thereby leading to the simplification in the third line. We have RADX  X mDl m.R/Dl m.R/! Y l.1/Y l.2/ DX  X mh Dl.R/1i mh Dl.R/i m! Y l.1/Y l.2/ DX Y l.1/Y l.2/DX Y l.1/Y l.2/DA: (16.55) This shows that Ais rotationally invariant, and is the starting point for an explanation of why a totally occupied atomic subshell (particles occupying all mvalues for a given l) leads to a spherically symmetric overall distribution. The rotational invariance of Amakes it easier for us to actually evaluate it, because we can choose to do so at a coordinate orientation for which the computation is relatively simple. Let’s rotate the coordinates to place 1in the polar direction (so now 1D0), and thevalue of2in the rotated coordinates will be equal to the angle between the 1and2directions, which is not affected by a coordinate rotation. In this new set of coordinates, Ym l.1/isYm l.0;'/ and is given, according to Eq. (15.148), as Ym l.1/Dr 2lC1 4m0: ArfKen_Ch16-9780123846549.tex 798 Chapter 16 Angular Momentum The summation in Eq. (16.53) therefore reduces to its mD0term, and the only 2contri- bution we need is Y0 l.;' 2/. But because mD0, this Ydoes not actually depend on '2, and has the unambiguous value, from Eq. (15.137), Y0 l.;' 2/Dr 2lC1 4Pl.cos/: These results enable us to obtain AD2lC1 4Pl.cos/; (16.56) which, because of the rotational invariance, remains true whether or not the coordinate sys- tem was rotated. Inserting the original formula for A, and solving Eq. (16.56) forPl.cos/, we obtain the spherical harmonic addition theorem, Pl.cos/D4 2lC1X mYm l.1/Ym l.2/; (16.57) whereis the angle between the directions 1and2. Example 16.3.1 ANGLE BETWEEN TWO VECTORS A useful special case of the addition theorem is for lD1, for which P1.cos/Dcos. Then, writing ii;'i, and evaluating all the spherical harmonics on the right-hand side ofEq. (16.57), we have cosD1 2 sin1ei'1 sin2ei'2 Ccos1cos2 C1 2 sin1ei'1 sin2ei'2 Dcos1cos2C1 2sin1sin2 ei.'1'2/Cei.'2'1/ : (16.58) This reduces to the standard formula for the angle between directions .1;'1/and .2;'2/: cosDcos1cos2Csin1sin2cos.' 2'1/: (16.59)  Spherical Wave Expansion An important application of the addition theorem is the spherical wave expansion, which states eikrD41X lD0lX mDliljl.kr/Ym l.k/Ym l.r/ (16.60) D41X lD0lX mDliljl.kr/Ym l.k/Ym l.r/: (16.61) ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 799 Here kandrare the magnitudes of kandr, andk,rdenote their respective angu- lar coordinates. The two forms shown are equivalent because a change in the sign of m changes each harmonic to its complex conjugate (possibly with both harmonics under- going a sign change). The quantity jl.kr/is a spherical Bessel function. This formula is particularly useful because it expresses the plane wave on its left-hand side as a series of spherical waves. This conversion is useful in scattering problems in which a plane wave, incident upon a scattering center, produces outgoing spherical waves with different spherical-harmonic (called partial-wave) components. To establish Eq. (16.61), we write kraskrcos, whereis the angle between kand r, and then expand exp.ikrcos/as a series of Legendre polynomials: eikrcosD1X lD0clPl.cos/; (16.62) with the coefficients clgiven by clD2lC1 21Z 1eikrtPl.t/dt: (16.63) We now recognize the integral in Eq. (16.63) as proportional to an integral representation ofjlthat was the topic of Exercise 15.2.26 and which we repeat here: jl.x/Dil 21Z 1eixtPl.t/dt: (16.64) This permits us to evaluate cl, obtaining clD.2lC1/iljl.kr/: Inserting this expression for clintoEq. (16.62) and replacing Pl.cos/in that equation by its equivalent as given by the addition theorem, Eq. (16.57), we have the desired verifica- tion of Eq. (16.61). Laplace Spherical Harmonic Expansion Another application of the addition theorem is to the Laplace expansion, where in Chapter 15 we found that the inverse distance between points r1andr2could be expanded in Legendre polynomials: 1 jr1r2jD1X lD0rl < rlC1>Pl.cos/: (16.65) Here r1andr2are measured from a common origin, with respective magnitudes r1and r2;is the angle between r1andr2. We define r>andr<as, respectively, the larger and ArfKen_Ch16-9780123846549.tex 800 Chapter 16 Angular Momentum the smaller of r1andr2. If we now insert the addition theorem, we bring this expansion to the form 1 jr1r2jD1X lD04 2lC1rl < rlC1>lX mDlYm l.1/Ym l.2/; (16.66) where1and2are the angular coordinates of r1andr2in a coordinate system of arbi- trary orientation. Example 16.3.2 SPHERICAL GREEN’S FUNCTION An explicit expansion of the Green’s function for the 3-D Laplace equation may be obtained by considering its defining equation r2 1G.r1;r2/D.r1r2/. 12/ r2 1; (16.67) where we have written r1to remind the reader that it acts only on r1. Also, note that on the right-hand side the factor 1=r2 1is inserted to adjust the angular delta function to unit scale; it could equally well have been written 1=r2 2because of the presence also of .r1r2/. We now insert into Eq. (16.67), the following general expansion for G.r1;r2/: G.r1;r2/DX lmX l0m0gll0mm0.r1;r2/Ym0 l0.1/Ym l.2/; and the expansion of Exercise 16.3.9 for the angular delta function: . 12/DX lmYm l.1/Ym l.2/: We also write the Laplacian in the form r2 1D@2 @r2 1C2 r1@ @r1L2 1 r2 1; where L1operates only on functions of 1. We next take scalar products of the resulting expanded equation with all possible spher- ical harmonics of both 1and2, in addition taking note that Ym l.1/is an eigenfunction ofL2 1with eigenvalue l.lC1/. We find that many terms cancel, so the scalar products lead, for each landm, to the following result: " d2 dr2 1C2 r1d dr1l.lC1/# gl.r1;r2/D.r1r2/: (16.68) We have collapsed the original four indices of gll0mm0.r1;r2/into the single index l because all instances of Eq. (16.68) with l6Dl0orm6Dm0vanish, and ghas the same value for all m. Equation (16.68) is for each lan ODE which, with boundary conditions gD0atrD0 andrD1 , defines the spherical Green’s functions we identified in Section 10.2. Since ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 801 the homogeneous equation corresponding to Eq. (16.68) has solutions rlandrl1, its Green’s function must have the form g.r1;r2/DAlrl < rlC1>; (16.69) with AlD1=.2 lC1/, a result that can be obtained by application of Eq. (10.19). Comparing Eq. (16.66) with the result for G.r1;r2/obtained by using Eq. (16.69), we now have yet another way of verifying the result that is familiar from Coulomb’s law: G.r1;r2/D1X lD01 2lC1rl < rlC1>lX mDlYm l.1/Ym l.2/ (16.70) D1 41 jr1r2j: (16.71)  General Multipoles We are now ready to return to the multipole expansion. Given a set of charges qiat respec- tive points ri, all located within a sphere of radius acentered at the origin of a spherical polar coordinate system, we now consider the calculation of the electrostatic potential .r/ at points outside the sphere, i.e., at points rsuch that r>a. Our starting point is the Laplace expansion of 1=jr 1r2jin the form presented as Eq. (16.66). Since for all riwe have ri<r, we can write .r/D1 4 0X iqi1X lD04 2lC1rl i rlC1lX mDlYm l.i;'i/Ym l.;'/ D1 4" 01X lD0lX mDl4 2lC1"X iqirl iYm l.i;'i/# Ym l.;'/ rlC1: (16.72) We see that this substitution has caused the entire effect of the charges qito be localized into the expressions Mm lD4 2lC1X iqirl iYm l.i;'i/; (16.73) so that the potential due to the qi, for points farther from rD0than all the charges, assumes the compact form, .r/D1 4" 01X lD0lX mDlMm lYm l.;'/ rlC1: (16.74) Equation (16.74) is called the multipole expansion, and the Mm lare known as the multi- pole moments of the charge distribution. At this point we note that different authors define the multipole moments with different scalings, making up the difference by the inclusion of an appropriate factor in their formulas correponding to Eq. (16.74). One reason for the ArfKen_Ch16-9780123846549.tex 802 Chapter 16 Angular Momentum variety of notations is that Mm las defined in Eq. (16.73), which leads to the simplest for- mulas, does not yield the low-order moments at their “traditional” scalings. For example, the monopole moment, M0 0, evaluates to .4/1=2times the total charge, while M0 1, the z-component of the dipole moment, comes out as .4=3/1=2P iqizi. Of more fundamental interest is the relation between the multipole moments and the Cartesian forms that can represent them. We proceed by considering the Mm lthat result from a unit charge placed at .x;y;z/. Using the Cartesian representations of the spherical harmonics given in Table 15.4, the first few Mm lhave the forms given here: M2 2D3 101=2 .x2y2C2ixy/ M1 1D2 31=2 .xCiy/M1 2D3 401=2 z.xCiy/ M0 0D.4/1=2M0 1D4 31=2 z M0 2D4 51=2 2z2x2y2 2 M1 1D2 31=2 .xiy/ M1 2D3 401=2 z.xiy/ M2 2D3 101=2 .x2y22ixy/ The first point to note is that for any lvalue, the Cartesian representation of each Mm l involves a homogeneous polynomial of combined degree linx,y, and z. It is obviously necessary that the Mm lof different mbe linearly independent, and we see that for lD0and lD1, the number of independent monomials is equal to 2lC1, the number of mvalues. Specifically, for lD0we have only the monomial 1, while for lD1we have x,y, and z. But for lD2, there are six independent monomials ( x2,y2,z2,xy,xz,yz), but only five values of m. The discrepancy is resolved by observing that one linear combination of these monomials, namely r2Dx2Cy2Cz2, remains invariant under all rotations of the coor- dinates, and it therefore has different symmetry properties than the five-dimensional space orthogonal to r2. In fact, r2has the same symmetry as M0 0, but has the wrong rdepen- dence to contribute to a solution to the Laplace equation (and therefore to the potential of a charge distribution). If we were to continue to lD3, we would find that there are 10 linearly independent monomials of degree 3, but they divide into a group of seven functions (the space spanned byMm 3) with an orthogonal complement (functions orthogonal to the first seven) of dimen- sion 3. These three remaining functions have a rotational symmetry similar to Mm 1, but again with the wrong rdependence to contribute to the potential. This type of pattern con- tinues to higher l, making logical the observation that a multipole moment of degree l(a “2l-moment”) has only 2lC1components, despite the fact that in general the space of homogeneous polynomials of degree lhas a larger dimension. The multipole expansion is useful for continuous distributions of charge in addition to the discrete charge sets we have considered up to this point. The generalization of Eq. (16.73) is Mm lD4 2lC1Z .r0/.r0/lC2Ym l.0;'0/sin0dr0d0d'0; (16.75) ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 803 where.r/ is the charge density. This expression will yield valid results when .r/ is computed via Eq. (16.74) forrvalues greater than the largest r0for which.r/ is nonzero. Integrals of Three Spherical Harmonics Our final spherical tensor application is to the integrals of three spherical harmonics (all of the same argument). These integrals arise in the evaluation of matrix elements of angle- dependent operators which themselves can be written in terms of spherical harmonics. While it is possible to evaluate some such integrals using the techniques illustrated in Eq. (15.152), a more general result is available. Note that this is not an angular momentum coupling problem of the type we considered in Section 16.2, because that section treated angular momenta with independent arguments that depended on different variables. Here we have a different and more specialized situation in which all three angular momentum functions have the same argument. The formula we seek is most easily derived if we have access to values of some of the rotation coefficients Dl m0m(a.k.a. Wigner matrices) defined in Eq. (16.52). The coefficients we need can be easily deduced with the aid of the spherical harmonic addition theorem, so we start by establishing the following lemma (a lemma is a mathematical result needed to prove something else): Lemma: Evaluation of Dl m0.R/: Writing first the spherical harmonic addition theorem, Eq. (16.57), Pl.cos/D4 2lC1X mYm l.1/Ym l.2/; whereis the angle between the directions 1and2, we replace its left-hand side by the equivalent form Pl.cos/Dr 4 2lC1Y0 l.;0/; thereby reaching Y0 l.;0/Dr 4 2lC1X mYm l.1/Ym l.2/: (16.76) We now compare this expression with Eq. (16.52), which we write here in a notation designed to make the comparison more obvious: Y0 l.R 2/DX mDl m0.R/Ym l.2/: (16.77) If we select Rto be a rotation that converts 1to the polar direction, then R2will be .;0/; note that Y0 l.R 2/is independent of 'so we can set its 'coordinate to zero. Thus, the comparison of Eqs. (16.76) and(16.77) yields Dl m0.R/Dr 4 2lC1Ym l.1/: (16.78) ArfKen_Ch16-9780123846549.tex 804 Chapter 16 Angular Momentum We remind the reader that Ris a rotation that converts 1to the polar direction. Equation (16.78) has been derived under the assumption that the quantities being rotationally transformed are spherical harmonics (and not more complicated angular- momentum functions such as might be obtained via angular-momentum coupling). How- ever, it is possible to show that the result generalizes, without change, to any angular momentum functions of integer l.  We now continue toward the goal of this subsection, namely the evaluation of integrals involving three spherical harmonics. The result we seek involves products of spherical harmonics with the same argument, but our method of obtaining that result proceeds by considering the rotational behavior of an angular momentum coupling formula (i.e., a prod- uct involving spherical harmonics of different arguments). So we now look at a special case of Eq. (16.45), Y0 l1.1/Y0 l2.2/DX LC.l1;l2;Lj0;0;0/jL;0i; (16.79) wherejj1;m1Ij2;m2iofEq. (16.45) is the product of spherical harmonics with m1D m2D0shown on the left-hand side of Eq. (16.79); thejJ;Mistate of Eq. (16.45) is now jL;0i. We next apply a rotation RtoEq. (16.79), using Eqs. (16.51) and(16.52) to get X m1m2Dl1 m10.R/Dl2 m20.R/Ym1 l1.1/Ym2 l2.2/DX L;C.l1;l2;Lj0;0;0/DL 0.R/jL;i: (16.80) Finally, we convertjL;iback to the m1;m2basis, using Eq. (16.43): X m1m2Dl1 m10.R/Dl2 m20.R/Ym1 l1.1/Ym2 l2.2/DX L;C.l1;l2;Lj0;0;0/DL 0.R/ X m1m2C.l1;l2;Ljm1;m2;/Ym1 l1.1/Ym2 l2.2/: (16.81) This relatively complicated equation must be satisfied for all values of 1and2, which will only be possible if its two sides are equal for each set of m1;m2values. We therefore have the set of simpler equations, Dl1 m10.R/Dl2 m20.R/DX LC.l1;l2;Lj0;0;0/C.l1;l2;Ljm1;m2;/DL 0.R/; (16.82) satisfied separately for all values of the free parameters. We are now ready to replace all the DlinEq. (16.82) by the result obtained in our lemma, Eq. (16.78). Since the rotation Ris arbitrary, both in the lemma and in the present work, our use of Eq. (16.78) will produce some angular coordinates that have nothing to do with the iwe were previously using; the point that is important here is that because the same Roccurs throughout Eq. (16.82), the application of Eq. (16.78) will everywhere ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 805 produce the same . Substitution of the lemma result yields 4p.2l1C1/.2l2C1/Ym1 l1./Ym2 l2./D X LC.l1;l2;Lj0;0;0/C.l1;l2;Ljm1;m2;/r 4 2LC1Y L./: Since the Ym lare the only potentially complex quantities appearing here, we may remove the complex conjugate signs by complex conjugating the entire equation. After other minor rearrangements and recognition of the fact that the only contributing value is Dm1Cm2, we reach the final form Ym1 l1./Ym2 l2./DX Ls .2l1C1/.2l2C1/ 4.2 LC1/ C.l1;l2;Lj0;0;0/C.l1;l2;Ljm1;m2;m1Cm2/Ym1Cm2 L./: (16.83) At last we can meet the objective of this subsection. Multiplying both sides of Eq. (16.83) by some Ym3 l3./and integrating in over the angular space, we get D Ym3 l3 Ym1 l1 Ym2 l2E D2Z 0d'Z 0sindYm3 l3.;'/Ym1 l1.;'/ Ym2 l2.;'/ Ds .2l1C1/.2l2C1/ 4.2 LC1/C.l1;l2;l3j0;0;0/C.l1;l2;l3jm1;m2;m3/: (16.84) We do not have to include a Kronecker delta because the condition m3Dm1Cm2is taken care of by the fact that the Clebsch-Gordan coefficients vanish in the absence of this or any other condition needed for a nonzero result. Some further insight can be obtained by considering the special case m1Dm2Dm3D0 and writing the spherical harmonics in terms of Legendre polynomials. This brings us (after the substitution tDcos) to 1Z 1Pl3.t/Pl1.t/Pl2.t/dtD2 2l3C1C.l1;l2;l3j0;0;0/2: (16.85) Since we know that the Legendre polynomial Pl.t/of even lis an even function of t, while that of odd lis odd in t, we see from Eq. (16.85) that unless l1Cl2Cl3is even, the integral will vanish, telling us that C.l1;l2;l3j0;0;0/will only be nonzero if l1Cl2Cl3is even. In addition, if the product of any two of the Pl.t/does not contain a power of tas large as the index of the third Pl, the integral will vanish due to the orthogonality of the Legendre functions. This observation translates into a triangle condition, namely that the integral will vanish unlessjl1l2jl3l1Cl2. Since these are conditions on the Clebsch-Gordan coefficient C.l1;l2;l3j0;0;0/, they apply also to the general integral formula, Eq. (16.84). Summarizing, integrals of products of three spherical harmonics, evaluated in Eq. (16.84), will only be nonzero if the three following conditions are satisfied: ArfKen_Ch16-9780123846549.tex 806 Chapter 16 Angular Momentum 1.Thelvalues satisfy the triangle condition jl1l2jl3l1Cl2, 2.Themvalues satisfy the condition m3Dm1Cm2, 3.The sum of the lvalues, l1Cl2Cl3, is even. Exercises 16.3.1 ForlD1, Eq. (16.52) becomes Ym 1.0;'0/D1X m0D1D1 m0m. ; ; / Ym0 1.;'/: Rewrite these spherical harmonics in Cartesian form. Show that the resulting Cartesian coordinate equations are equivalent to the Euler rotation matrix A. ; ; / , Eq. (3.37). 16.3.2 In proving the addition theorem, we assumed that Yk l.1;'1/could be expanded in a series of Ym l.2;'2/;in which mvaried fromltoClbutlwas held fixed. What argu- ments can you develop to justify summing only over the upper index, m;andnotover the lower index, l? Hints. One possibility is to examine the homogeneity of the Ym l;that is, Ym lmay be expressed entirely in terms of the form coslpsinp;orxlpsypzs=rl. Another possibility is to examine the behavior of the Legendre equation under rotation of the coordinate system. 16.3.3 An atomic electron with angular momentum land magnetic quantum number mhas a wave function .r;;'/Df.r/Ym l.;'/: Show that the sum of the electron densities in a given complete shell is spherically symmetric; that is,Pl mDl .r;;'/ . r;;'/ is independent of and'. 16.3.4 The potential of an electron at point rein the field of Zprotons at points rpis 8De2 4" 0ZX pD11 jrerpj: Show that for relarger than all rp, this may be written as 8De2 4" 0reZX pD1X L;Mrp reL4 2LC1YM L.p;'p/YM L.e;'e/: How should8be written for re<rp? 16.3.5 Two protons are uniformly distributed within the same spherical volume. If the coor- dinates of one element of charge are .r1;1;'1/and the coordinates of the other are .r2;2;'2/andr12is the distance between them, the element of repulsion energy will be given by d D2d1d2 r12D2r2 1dr1sin1d1d'1r2 2dr2sin2d2d'2 r12; ArfKen_Ch16-9780123846549.tex 16.3 Spherical Tensors 807 where Dcharge volumeD3e 4R3and r2 12Dr2 1Cr2 22r1r2cos : Hereis the charge density and is the angle between r1and r2. Calculate the total electrostatic energy (of repulsion) of the two protons. This calculation is used in accounting for the mass difference in “mirror” nuclei, such as O15and N15. ANS.6 5e2 R. 16.3.6 Each of the two 1selectrons in helium may be described by a hydrogenic wave function .r/D Z3 a3 0!1=2 eZr=a0 in the absence of the other electron. Here Z, the atomic number, is 2. The symbol a0 is the Bohr radius, Nh2=me2. Find the mutual potential energy of the two electrons, given by Z .r1/ .r2/e2 jr1r2j .r 1/ .r 2/d3r1d3r2: ANS.5e2Z 8a0. 16.3.7 The probability of finding a 1shydrogen electron in a volume element r2drsindd'is 1 a3 0e2r=a0r2drsindd'; where ris the distance of the electron from the nucleus. Find the electrostatic potential of this charge distribution at points r1, where you may notassume that r1is on the polar axis of your coordinate system. Calculate the potential from V.r1/De 4" 0Z.r2/ r12d3r2; where r12Djr 1r2j. Expand r12. Apply the Legendre polynomial addition theorem and show that the angular dependence of V.r1/drops out. ANS. V.r1/De 4" 01 2r1  3;2r1 a0 C1 a00 2;2r1 a0 , where and0are incomplete gamma functions, Eq. (13.73). 16.3.8 A hydrogen electron in a 2porbital has a charge distribution De 64a5 0r2er=a0sin2; where a0DNh2=me2is the Bohr radius, and ris the distance between the electron and the nucleus. Find the electrostatic potential energy for this atomic state. ArfKen_Ch16-9780123846549.tex 808 Chapter 16 Angular Momentum 16.3.9 (a) As a Laplace series and as an example of Eq. (5.27), show that . 12/D1X lD0lX mDlYm l.2;'2/Ym l.1;'1/: (b) Show also that this same representation of the Dirac delta function may be written as . 12/D1X lD02lC1 4Pl.cos /; and identify . Now, if you can justify equating the summations over lterm by term, you have an alternate derivation of the spherical harmonic addition theorem. 16.3.10 Verify (a)Z YM L.;'/ Y0 0.;'/ YM L.;'/dD1p 4, (b)Z YM LY0 1YM LC1dDr 3 4s .LCMC1/.LMC1/ .2LC1/.2LC3/, (c)Z YM LY1 1YMC1 LC1dDr 3 8s .LCMC1/.LCMC2/ .2LC1/.2LC3/, (d)Z YM LY1 1YMC1 L1dDr 3 8s .LM/.LM1/ .2L1/.2LC1/. These integrals were used in an investigation of the angular correlation of internal conversion electrons. 16.3.11 Show that (a)1Z 1x PL.x/PN.x/dxD8 >>< >>:2.LC1/ .2LC1/.2LC3/;NDLC1; 2L .2L1/.2LC1/;NDL1; (b)1Z 1x2PL.x/PN.x/dxD8 >>>>>>>>< >>>>>>>>:2.LC1/.LC2/ .2LC1/.2LC3/.2LC5/;NDLC2; 2.2L2C2L1/ .2L1/.2LC1/.2LC3/;NDL; 2L.L1/ .2L3/.2L1/.2LC1/;NDL2: ArfKen_Ch16-9780123846549.tex 16.4 Vector Spherical Harmonics 809 16.3.12 Since x Pn.x/is a polynomial (of degree nC1), it may be represented by the Legendre series x Pn.x/D1X sD0asPs.x/: (a) Show that asD0fors<n1ands>nC1. (b) Calculate an1,an, and anC1and show that you have reproduced the recurrence relation, Eq. (15.18). Note. This argument may be put in a general form to demonstrate the existence of a three-term recurrence relation for any of our complete sets of orthogonal polynomials: x'nDanC1'nC1Can'nCan1'n1: 16.4 V ECTOR SPHERICAL HARMONICS Maxwell’s equations lead naturally to applications involving a vector Helmholtz equation for the vector potential A, and various classical and quantum-mechanical problems in this area are usefully attacked by introducing vector spherical harmonics. Our first step in this direction will be to recognize that a set of unit vectors can be thought of as a spherical tensor of rank 1 and can be discussed in terms of the angular momentum formalism. We will later (in Chapter 17) pursue rotational symmetry in greater depth; for our present purposes it suffices to confirm the relationship between rotations in 3-D space and angular momentum operators. A Spherical Tensor We consider here vectors in 3-D space, of the form uDuxOexCuyOeyCuzOez, but, unlike our practice in Chapter 3, we will permit the ujto be complex, and use the complex scalar producthujui1=2as a measure of the magnitude of u. If we restrict the vectors uto be of unit length, they satisfy the conditions necessary to be identified as spherical tensors of rank 1. We now introduce operators Kidefined by the following matrices: K1D0 @0 0 0 0 0i 0i01 A;K2D0 @0 0 i 0 0 0 i0 01 A;K3D0 @0i0 i0 0 0 0 01 A: (16.86) The reader can easily verify that these matrices satisfy the angular momentum commu- tation rules, and in fact describe the result of applying the angular momentum operator LDrp, where pDir, to the basis x,y,z. We next calculate K2DK2 1CK2 2CK2 3D20 @1 0 0 0 1 0 0 0 11 A; ArfKen_Ch16-9780123846549.tex 810 Chapter 16 Angular Momentum showing that all members of the basis are eigenvectors of K2, with eigenvalue 2, which is k.kC1/with kD1. All members of our basis therefore have one unit of some abstract sort of angular momentum (often referred to as spin), and we can obtain a set of eigenvectors with values of an index mthat can have valuesC1, 0, and1. By diagonalizing the matrix K3, we find its eigenvectors to be k1D0 @1p 2 i=p 2 01 A;k0D0 @0 0 11 A;k1D0 @1p 2 i=p 2 01 A: (16.87) While in principle the signs of these eigenvectors are arbitrary, they have been chosen here to agree with the Condon-Shortley phase convention. Vector Coupling The vector spherical harmonics are now defined as the quantities that result from the cou- pling of ordinary spherical harmonics and the vectors emto form states of definite J(the resultant of the orbital angular momentum of the spherical harmonic and the one unit pos- sessed by the em). It is customary to label the vector spherical harmonics to show both theLvalue from the ordinary (scalar) harmonic and the Mvalue (the eigenvalue of Jz). Thus, the vector spherical harmonic will have three indices: J,L, and M. From the general formula for angular-momentum coupling, Eq. (16.43), we have YJ L M.;'/DX mm0C.L;1;Jjmm0M/Ym L.;'/Oem0: (16.88) Remember that MisMJ, not the mvalue of Ym L, and thatOem0are the angular momentum eigenfunctions given in Eq. (16.87). Because Eq. (16.88) couples an angular momentum Lwith one of magnitude kD1, the Lvalues in a vector spherical harmonic of given Jare restricted to JC1,J, and J1, a condition enforced by the values of the Clebsch-Gordan coefficients. Moreover, because the Clebsch-Gordan coefficients describe a unitary transformation, the obvious orthogo- nality of the states in the m;m0basis ( Ym lOem0) will cause the vector spherical harmonics also to be orthonormal: Z YJ L M.;'/YJ0L0M0.;'/dDJ J0L L0M M0: (16.89) In addition, we can invert Eq. (16.89) using Eq. (16.45), reaching Ym L.;'/Oem0DX J MC.L;1;Jjmm0M/YJ L M: (16.90) The manipulation of expressions involving the vector spherical harmonics depends crucially on a few identities, of which perhaps the most important is the formula OrYM L.;'/DLC1 2LC11=2 YL;LC1;MCL 2LC11=2 YL;L1;M: (16.91) ArfKen_Ch16-9780123846549.tex 16.4 Vector Spherical Harmonics 811 To establish this formula, and at the same time to make its meaning more obvious, we start by noting thatOrhas a form that depends on the angular coordinates; specifically, it is OrDr sincos'OexCsinsin'OeyCcosOez: For our present purposes, it is more convenient to rearrange this to the form OrDsinei'Oe1ei'OeC1p 2 CcosOe0: (16.92) It is now clear that in order to prove Eq. (16.91) we must show that each Oemhas the same coefficient on both sides of the equation. Taking first the coefficient of Oe0, the left-hand side of Eq. (16.91) yields, after use of Eq. (16.92), cosYM L.;'/D.lmC1/.lCmC1/ .2lC1/.2lC3/1=2 Ym lC1 C.lm/.lCm/ .2l1/.2lC1/1=2 Ym l1; (16.93) a result previously exhibited as Eq. (15.150). The Oe0terms from the right-hand side of Eq. (16.91) consist of LC1 2LC11=2 C.LC1;1;LjM;0;M/YM LC1Oe0 CL 2LC11=2 C.L1;1;LjM;0;M/YM L1Oe0: The Clebsch-Gordan coefficients appearing here have the values C.LC1;1;LjM;0;M/D.LCMC1/.LMC1/ .LC1/.2LC3/1=2 ; C.L1;1;LjM;0;M/DL2M2 L.2L1/1=2 : These data permit confirmation of the e0terms of Eq. (16.91). The terms in eC1ande1 can also be shown consistent; the formulas needed for that purpose are Eqs. (15.151) and (15.152). Another useful formula, which can be obtained by using Eq. (16.91) to simplify the radial component when the gradient operator is applied to the form f.r/YM L.;'/ , is rh f.r/YM K.;'/i DLC1 2LC11=2@ @rL r f.r/YL;LC1;M.;'/ CL 2LC11=2@ @rCLC1 r f.r/YL;L1;M.;'/: (16.94) ArfKen_Ch16-9780123846549.tex 812 Chapter 16 Angular Momentum Under coordinate inversion the vector spherical harmonics transform as YL;LC1;M.0;'0/D.1/LC1YL;LC1;M.;'/; YL;L1;M.0;'0/D.1/LC1YL;L1;M.;'/; (16.95) YL L M.0;'0/D.1/LYL L M.;'/; where 0D '0DC': Starting from Eqs. (16.91) and (16.94), a number of formulas can be derived for the divergence and curl of vector spherical harmonics. These formulas include the following: rh f.r/YL;LC1;M.;'/i DLC1 2LC11=2d f.r/ drCLC2 rf.r/ YM L.;'/; (16.96) rh f.r/YL;L1;M.;'/i DL 2LC11=2d f.r/ drL1 rf.r/ YM L.;'/; (16.97) rh f.r/YL L M.;'/i D0; (16.98) rh f.r/YL;LC1;M.;'/i DiL 2LC11=2d f.r/ drCLC2 rf.r/ YL L M;(16.99) rh f.r/YL L M.;'i DiL 2LC11=2d f.r/ drL rf.r/ YL;LC1;M.;'/; CiLC1 2LC11=2d f.r/ drCLC1 rf.r/ YL;L1;M; (16.100) rh f.r/YL;L1;M.;'/i DiLC1 2LC11=2d f.r/ drL1 rf.r/ YL L M.;'/: (16.101) For a complete derivation of Eqs. (16.96) to(16.101) we refer to the literature.3These relations play an important role in the partial wave expansion of classical and quantum electrodynamics. The definitions of the vector spherical harmonics given here are dictated by convenience, primarily in quantum mechanical calculations, in which the angular momentum is a sig- nificant parameter. Further examples of the usefulness and power of the vector spherical harmonics will be found in Blatt and Weisskopf, in Morse and Feshbach, and in Jackson (all in Additional Readings). In closing, we note that 3E. H. Hill, Theory of vector spherical harmonics, Am. J. Phys. 22: 211 (1954). Note that Hill assigns phases in accordance with the Condon-Shortley phase convention. In Hill’s notation, XL MDYL L M ,VL MDYL;LC1;M,WL MDYL;L1;M. ArfKen_Ch16-9780123846549.tex 16.4 Vector Spherical Harmonics 813 Vector spherical harmonics are developed from coupling Lunits of orbital angular momentum and one unit of spin angular momentum. An extension, coupling Lunits of orbital angular momentum and two units of spin angular momentum to form tensor spherical harmonics, is presented by Mathews.4 The major application of tensor spherical harmonics is in the investigation of gravita- tional radiation. Exercises 16.4.1 Construct the lD0;mD0andlD1;mD0vector spherical harmonics. ANS. Y010DOr.4/1=2 Y000D0 Y120DOr.2/1=2cosO.8/1=2sin Y110DO'i.3=8/1=2sin Y100DOr.4/1=2cosO.4/1=2sin. 16.4.2 Verify that the parity of YL LC1Mis.1/LC1, that of YL L M is.1/L, and that of YL L1Mis.1/LC1. What happened to the M-dependence of the parity? Hint. OrandO'have odd parity; Ohas even parity (compare with Exercise 3.10.25). 16.4.3 Verify the orthonormality of the vector spherical harmonics YJ L M J. 16.4.4 Jackson’s Classical Electrodynamics (see Additional Readings) defines YL L M by the equation YL L M.;'/D1pL.LC1/LYM L.;'/; in which the angular momentum operator Lis given by LDi.rr/: Show that this definition agrees with Eq. (16.88). 16.4.5 Show that LX MDLY L L M.;'/YL L M.;'/D2LC1 4: Hint. One way is to use Exercise 16.4.4 with Lexpanded in Cartesian coordinates and to apply raising and lowering operators. 4J. Mathews, Gravitational multipole radiation, J. Soc. Ind. Appl. Math. 10: 768 (1963). ArfKen_Ch16-9780123846549.tex 814 Chapter 16 Angular Momentum 16.4.6 Show thatZ YL L M.OrYL L M/dD0: The integrand represents an interference term in electromagnetic radiation that contributes to angular distributions but not to total intensity. Additional Readings Biedenharn, L. C., and J. D. Louck, Angular Momentum in Quantum Physics: Theory and Application. Ency- clopedia of Mathematics and Its Applications, vol. 8. Reading, MA: Addison-Wesley (1981). An extremely detailed account, containing much material not easily found elsewhere. Blatt, J. M., and V. Weisskopf, Theoretical Nuclear Physics. New York: Wiley (1952). Treats vector spherical harmonics. Brink, D. M., and G. R. Satchler, Angular Momentum. New York: Oxford (1993). Contains a good presentation of graphical methods for the manipulation of 3 j, 6j, and even 9 jsymbols. The 6 jand 9 jsymbols are useful in dealing with the coupling of more than two angular momenta. Condon, E. U., and G. H. Shortley, Theory of Atomic Spectra. Cambridge: Cambridge University Press (1935). This is the original and standard work on spin-orbit coupling in atomic states. It is extremely thorough and not for the beginner. Edmonds, A. R., Angular Momentum in Quantum Mechanics. Princeton, NJ: Princeton University Press (1957). A good introductory text, with detailed discussion of the symmetries of 3j,6j, and 9jsymbols. Jackson, J. D., Classical Electrodynamics, 3rd ed. New York: Wiley (1999). Applies vector spherical harmonics to multipole radiation and related problems. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics, 2 vols. New York: McGraw-Hill (1953). Includes material on vector spherical harmonics. Rose, M. E., Elementary Theory of Angular Momentum. New York: Wiley (1957), reprinted, Dover (1995). As part of the development of the quantum theory of angular momentum, Rose includes a detailed and readable account of the rotation group. Wigner, E. P., Group Theory and Its Application to the Quantum Mechanics of Atomic Spectra (translated by J. J. Griffin). New York: Academic Press (1959). This is the classic reference on group theory for the physi- cist. The rotation group is treated in considerable detail. There is a wealth of applications to atomic physics. The translation from the original German edition included a conversion from a left-handed to a right-handed coordinate system. This conversion introduced a few errors that can be resolved by comparison with the untranslated book. ArfKen_Ch17-9780123846549.tex CHAPTER 17 GROUP THEORY Disciplined judgment, about what is neat and symmetrical and elegant, has time and time again proved an excellent guide to how nature works. MURRAY GELL-MANN 17.1 I NTRODUCTION TO GROUP THEORY Symmetry has long been important in the study of physical systems. Connections between the geometric symmetry of crystalline systems and their x-ray diffraction spectra were found to be crucial to the interpretation of the diffraction patterns and the extraction therefrom of information locating the atoms in the crystal. The geometric symmetries of molecules determine which vibrational modes will be active in absorbing or emitting radi- ation; the symmetries of periodic systems have implications as to their energy bands, their ability to conduct electricity, and even their superconductivity. The invariance of physi- cal laws with respect to position or orientation (i.e., the symmetry of space) gives rise to conservation laws for linear and angular momentum. Sometimes the implications of sym- metry invariance are far more complicated or sophisticated than might at first be supposed; the invariance of the forces predicted by electromagnetic theory when measurements are made in observation frames moving uniformly at different speeds (inertial frames) was an important clue leading Einstein to the discovery of special relativity. With the advent of quantum mechanics, considerations of angular momentum and spin introduced new sym- metry concepts into physics. These ideas have since catalyzed the modern development of particle theory. Central to all these symmetry notions is the fact that complete sets of symmetry oper- ations form what in mathematics are known as groups. The elements of a group may be finite in number, in which case the group is then termed finite ordiscrete, as for example the symmetry operations shown for the object depicted in Fig. 17.2 . But alternatively, the 815 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch17-9780123846549.tex 816 Chapter 17 Group Theory symmetry operations may be infinite in number and described by continuously variable parameter(s); such groups are termed continuous. An example of a continuous group is the set of possible rotational displacements of a circular object about its axis (in which case the parameter is the rotation angle). Definition of a Group A group Gis defined as a set of objects or operations (e.g., rotations or other transfor- mations), called the elements of G, that may be combined, by a procedure to be called multiplication and denoted by *, to form a well-defined product, subject to the following four conditions: 1. If aandbare any two elements of G, then the product abis also an element of G; more formally, abassociates an element of Gwith the ordered pair .a;b/of elements of G. In other words, Gisclosed under multiplication of its own elements. 2. This multiplication is associative: .ab/cDa.bc/. 3. There is a unique identity element1IinG, such that IaDaIDafor every element ainG. 4. Each element aofGhas an inverse, denoted a1, such that aa1Da1aDI. The above simple rules have a number of direct consequences, including the following: It can be shown that the inverse of any element ais unique: If a1andOa1are both inverses of a, thenOa1DOa1.aa1/D.Oa1a/a1Da1. The products ga, where ais fixed and granges over all elements of the group, consist (in some order) of all the elements of the group. If gandg0produce the same element, thengaDg0a. Multiplying on the right by a1, we get.ga/a1D.g0a/a1, which reduces to gDg0. Here are some useful conventions and further definitions: The * for multiplication is tedious to write; when no ambiguity will result it is custom- ary to drop it, and instead of abwe write ab. When aandbare operations, and abis to be applied to an object appearing to their right, bis deemed to act first, with athen applied to the result of operation with b. If a discrete group possesses nelements (including I), its order isn; a continuous group of order nhas elements that are defined by nparameters. IfabDbafor all a,bofG, the multiplication is commutative, and the group is called abelian. If a group possesses an element asuch that the sequence I,a,a2.Daa/,a3, includes all elements of the group, it is termed cyclic. If a group is cyclic, it must also be abelian. However, not all abelian groups are cyclic. 1Following E. Wigner, the identity element of a group is often labeled E, from the German Einheit, that is, unit; some other authors just write 1. ArfKen_Ch17-9780123846549.tex 17.1 Introduction to Group Theory 817 Two groupsfI;a;b;g andfI0;a0;b0;g areisomorphic if their elements can be put into one-to-one correspondence such that for all aandb,abDc() a0b0Dc0. If the correspondence is many-to-one, the groups are homomorphic. If a subset G0ofGis closed under the multiplication defined for G, it is also a group and called a subgroup ofG. The identity IofGalways forms a subgroup of G. Examples of Groups Example 17.1.1 D3, SYMMETRY OF AN EQUILATERAL TRIANGLE The symmetry operations of an equilateral triangle form a finite group with six elements; our triangle can be placed either side up, and with any vertex in the top position. The six operations that convert the initial orientation into symmetry equivalents are I(the identity operation that makes no orientation change), C3, an operation which rotates the triangle counterclockwise by 1/3 of a revolution, C2 3(two successive C3operations), C2, rotation by 1/2 a revolution (for this group the rotation is about an axis in the plane of the trian- gle), and C0 2andC00 2(180rotations about additional axes in the plane of the triangle). Figure 17.1 is a schematic diagram indicating these symmetry operations, and Fig. 17.2 shows their result, with the vertices of the triangle numbered to show the effect of each operation. The multiplication table for the group is shown in Table 17.1, where the prod- uctab(which describes the result of first applying operation b, and then operation a) is (0, 1)y C″2C′2 xC31 C22 3) )( (23−, −2123, −21 FIGURE 17.1 Diagram identifying symmetry operations of an equilateral triangle. Iis the identity operation (the diagram as shown here). C3andC2 3are counterclockwise rotations, by, respectively, 120and240;C2,C0 2,C00 2are operations that turn the triangle over by rotation about the indicated axes. ArfKen_Ch17-9780123846549.tex 818 Chapter 17 Group Theory 1 11 1 13 33 I C3 C2 C2 3C′2 C″23 31 32 2 2 22 2 FIGURE 17.2 Result of applying the symmetry operations identified in Fig. 17.1 to an equilateral triangle. One side of the triangle is shaded to make it obvious when that side is up. Table 17.1 Multiplication Table for Group D3 I C 3 C2 3C2 C0 2C00 2 I I C 3 C2 3C2 C0 2C00 2 C3 C3 C2 3I C00 2C2 C0 2 C2 3C2 3I C 3 C0 2C00 2C2 C2 C2 C0 2C00 2I C 3 C2 3 C0 2C0 2C00 2C2 C2 3I C 3 C00 2C00 2C2 C0 2C3 C2 3I Operations are pictured in Fig. 17.2. The table entry for row aand column b is the product element ab. For example, C2C3DC0 2. the group element listed in row aand column bof the table. This group has several names, of which one is D3(“D” for dihedral, referring to a 180rotation axis lying in a plane perpendicular to the main symmetry axis). From the multiplication table or by examination of the symmetry operations themselves, we can see that the inverse of IisI, the inverse ofC3isC2 3(so the inverse of C2 3isC3), and each C2is its own inverse. This group is not abelian; C3C26DC2C3(C3C2DC00 2, while C2C3DC0 2).  Example 17.1.2 ROTATION OF A CIRCULAR DISK The rotations of a circular disk about its symmetry axis form a continuous group of order 1 whose elements consist of rotations through angles '. The group elements R.'/ are infinite in number, with 'any angle in the range .0;2/. The identity element is clearly R.0/ ; the inverse of R.'/ isR.2'/. The multiplication rule for this group is R.'/R./D R.'C/(reduced to a value between 0 and 2), soR.'/R./DR./R.'/ , and this group is abelian. It will be useful to figure out what happens to a point on the disk that before the rotation was at .x;y/. The rotation is by an angle 'about the zaxis, clockwise, looking down from positive z, a choice made to be consistent with the counterclockwise rotations of the coordinate axes used elsewhere in this book. The final location of this point, ( x0;y0), ArfKen_Ch17-9780123846549.tex 17.1 Introduction to Group Theory 819 is given by the matrix equation x0 y0 D cos' sin' sin'cos'!x y : (17.1)  Example 17.1.3 ANABSTRACT GROUP Groups do not need to represent geometric operations. Consider a set of four quantities (elements) I,A,B,C, with our knowledge about them only that when any two are multi- plied, the result is an element of the set. The multiplication table of this four-element set is shown in Table 17.2. These elements form a group, because each has an inverse (itself), there is an identity element ( I), and the set is closed under multiplication. Table 17.2 Multiplication Table for the Vierergruppe I A B C I I A B C A A I C B B B C I A C C B A I The table entry for row aand column bis the product element ab. Example 17.1.4 ISOMORPHISM AND HOMOMORPHISM: C4GROUP The symmetry operations of a square that cannot be turned over form a four-membered group sometimes called C4whose elements can be named I,C4(90rotation), C2(180 rotation), C0 4(270rotation). The four complex numbers 1, i,1,ialso form a group when the group operation is ordinary multiplication. These groups are isomorphic, and can be put into correspondence in two different ways: I$1;C4$i;C2$1; C0 4$ iorI$1;C4$ i;C2$1; C0 4$i: This group is also cyclic, as C2 4DC2,C3 4DC0 4, or equivalently i2D1 ,i3Di. The group C4has a two-to-one correspondence with the ordinary multiplicative group containing only 1 and 1:IandC2$1, while C4andC0 4$1 . This is a homomor- phism. A more trivial homomorphism, possessed by all groups, is obtained when every element is assigned to correspond to the identity.  ArfKen_Ch17-9780123846549.tex 820 Chapter 17 Group Theory Exercises 17.1.1 TheVierergruppe (German: four-membered group) is a group different from the C4 group introduced in Example 17.1.4. The Vierergruppe has the multiplication table shown in Table 17.2. Determine whether this group is cyclic and whether it is abelian. 17.1.2 (a) Show that the permutations of ndistinct objects satisfy the group postulates. (b) Construct the multiplication table for the permutations of three objects, giving each permutation a name of some sort. (Suggestion: Use Ifor the permutation that leaves the order unchanged.) (c) Show that this permutation group (named S3) is isomorphic with D3and identify corresponding operations. Is your identification unique? 17.1.3 Rearrangement theorem: Given a group of distinct elements .I;a;b;:::; n/, show that the set of products .aI;a2;ab;ac;:::; an/reproduces all the group elements in a new order. 17.1.4 A group Ghas a subgroup Hwith elements hi. Let xbe a fixed element of the original group Gandnota member of H. The transform xhix1;iD1;2;::: generates a conjugate subgroup x H x1. Show that this conjugate subgroup satisfies each of the four group postulates and therefore is a group. 17.1.5 (a) A particular group is abelian. A second group is created by replacing gibyg1 i for each element in the original group. Show that the two groups are isomorphic. Note. This means showing that if abDc, then a1b1Dc1. (b) Continuing part (a), show that the second group is also abelian. 17.1.6 Consider a cubic crystal consisting of identical atoms at rD.la;ma;na/, with l;m, andntaking on all integral values. (a) Show that each Cartesian axis is a fourfold symmetry axis. (b) The cubic point group will consist of all operations (rotations, reflections, inver- sion) that leave the simple cubic crystal invariant and that do not move the atom at lDmDnD0. From a consideration of the permutation of the positive and nega- tive coordinate axes, predict how many elements this cubic group will contain. 17.1.7 A plane is covered with regular hexagons, as shown in Fig. 17.3. (a) Determine the rotational symmetry of an axis perpendicular to the plane through the common vertex of three hexagons .A/. That is, if the axis has n-fold symmetry, show (with careful explanation) what nis. (b) Repeat part (a) for an axis perpendicular to the plane through the geometric center of one hexagon .B/. (c) Find all the different kinds of axes within the plane of hexagons about which a 180rotation is a symmetry element (this corresponds to turning the plane over by rotation about that axis). ArfKen_Ch17-9780123846549.tex 17.2 Representation of Groups 821 A B FIGURE 17.3 Plane covered by hexagons. 17.2 R EPRESENTATION OF GROUPS All discrete groups and the continuous groups we study here can be represented by square matrices. By this we mean that to each element of the group we can associate a matrix, and that if U.a/is the matrix associated with aandU.b/the matrix associated with b, then the matrix product U.a/U.b/will be the matrix associated with ab. In other words, the matrices have the same multiplication table as the group. We call these matrices U because they can be chosen to be unitary. It is not necessary that Uhave a dimension equal to the order of the group. Sometimes we need to identify representations with a label. For specific representations we can use their generally adopted names; when we need a generic label, we will use Kor K0. Thus, we can refer to representation K, consisting of matrices UK.a/. Example 17.2.1 A UNITARY REPRESENTATION Here is a unitary representation of the group D3illustrated in Fig. 17.2: U.I/D 1 0 0 1! ; U.C3/D 1 21 2p 3 1 2p 31 2! ; U.C2 3/D 1 21 2p 3 1 2p 31 2! ; U.C2/D 1 0 01! ; U.C0 2/D 1 21 2p 3 1 2p 31 2! ; U.C00 2/D 1 21 2p 3 1 2p 31 2! : (17.2) Several features of this representation are apparent: The unit operation is represented by a unit matrix. The inverse of an operation is represented by the inverse of its matrix. ArfKen_Ch17-9780123846549.tex 822 Chapter 17 Group Theory We can check that the Uform a representation: From the multiplication table, we have C2C3DC0 2. Now we evaluate U.C2/U.C3/D 1 0 01! 1 21 2p 3 1 2p 31 2! D 1 21 2p 3 1 2p 31 2! ; which is indeed U.C0 2/. The reader can easily verify that other products of group elements correspond to the products of the representation matrices. Matrix multiplication is in gen- eral not commutative, and gives results that are consistent with the lack of commutativity of the group operations. The22representation shown above is faithful, meaning that each group element cor- responds to a different matrix. In other words, our 22representation is isomorphic with the original group. Not all representations are faithful; consider the relatively trivial repre- sentation in which every group element is represented by the 11matrix.1/. Every group will possess this representation. A somewhat less trivial, but still unfaithful, representation ofD3is one in which U.I/DU.C3/DU.C2 3/D1;U.C2/DU.C0 2/DU.C00 2/D1: (17.3) This representation distinguishes elements according to whether they involve turning the triangle over. Not all groups will possess this 11representation; if we had not permit- ted the triangle to be turned over, this representation would have been excluded. These unfaithful representations are homomorphic with the original group.  An important feature of a representation of a group Gis that its essential fea- tures are invariant if we make the same unitary transformation on the matrices repre- senting all the group elements. To see this, consider what happens when we replace eachU.g/byVU.g/V1. Then the product U.g/U.g0/, which is some U.g00/, becomes .VU. g/V1/.VU. g0/V1/DVU.g/U.g0/V1DVU.g00/V1, so the transformed matrices still form a representation of G. Representations that can be transformed into each other by application of a unitary transformation are termed equivalent. The possibility of unitary transformation also enables us to consider whether a repre- sentation of Gisreducible. An irreducible representation of Gis defined as one that cannot be broken into a direct sum of representations of smaller dimension by application of the same unitary transformation to all members of the representation. What we mean by a direct sum of representations is that each matrix will be block diagonal (all with the same sequence of blocks). Since different blocks will not mix under matrix multiplication, cor- responding blocks of the representation members will themselves define representations (see Fig. 17.4). If a representation named Kis a direct sum of smaller representations K1 andK2, that fact can be indicated by the notation KDK1K2: It is not always obvious whether a representation is reducible. We will shortly encounter theorems that provide (for discrete groups) ways of determining what irreducible repre- sentations are present in a representation that may be reducible. Moreover, if a group is abelian, then the fact that all its elements commute means that the matrices represent- ing them can all be diagonalized simultaneously. From that fact we can conclude that all irreducible representations of abelian groups are 11. ArfKen_Ch17-9780123846549.tex 17.2 Representation of Groups 823 UK(a)=UK′(a) UK″(a) UK′″(a) FIGURE 17.4 A member of a reducible representation in direct-sum form. All members will have the same block structure, so individual blocks define representations of smaller dimension. It is important to understand that reducibility implies the existence of a unitary trans- formation that brings all members of a representation to the same block-diagonal form; a reducible representation may not exhibit the block-diagonal form if it has not been sub- jected to a suitable unitary transformation. Here is an example illustrating that point. Example 17.2.2 A REDUCIBLE REPRESENTATION Here is a reducible representation for our equilateral triangle: U.I/D0 @1 0 0 0 1 0 0 0 11 A;U.C3/D0 @0 1 0 0 0 1 1 0 01 A;U.C2 3/D0 @0 0 1 1 0 0 0 1 01 A; U.C2/D0 @0 0 1 0 1 0 1 0 01 A;U.C0 2/D0 @1 0 0 0 0 1 0 1 01 A;U.C00 2/D0 @0 1 0 1 0 0 0 0 11 A:(17.4) Note that some of these matrices are not in any direct-sum form. To show that the repre- sentation of Eq. (17.4) is reducible, we transform all the UtoU0DVUV1, using VD0 BB@1=p 3 1=p 3 1=p 3 1=p 6p2=3 1=p 6 1=p 2 01=p 21 CCA; ArfKen_Ch17-9780123846549.tex 824 Chapter 17 Group Theory which brings us to U0.I/D0 B@1 0 0 0 1 0 0 0 11 CA; U0.C3/D0 B@1 0 0 01 21 2p 3 01 2p 31 21 CA; U0.C2 3/D0 B@1 0 0 01 21 2p 3 01 2p 31 21 CA;U0.C2/D0 B@1 0 0 0 1 0 0 011 CA; U0.C0 2/D0 B@1 0 0 01 21 2p 3 01 2p 31 21 CA;U0.C00 2/D0 B@1 0 0 01 21 2p 3 01 2p 31 21 CA: (17.5) All the matrices of this representation are block diagonal, and are direct sums that consist of an upper 11block that is the trivial representation, all of whose elements are (1), and a lower 22block that is exactly the 22representation illustrated in Eq. (17.2). There exists no unitary transformation that will simultaneously reduce the 22blocks of all members of the representation to direct sums of 11blocks, so we have reduced the representation of Eq. (17.4) to its irreducible components.2 Example 17.2.3 REPRESENTATIONS OF A CONTINUOUS GROUP Example 17.1.2 presented a continuous group of order 1 whose elements are rotations R.'/ about the symmetry axis of a circular disk. These rotations were taken to be defined by the matrix equation presented as Eq. (17.1). The 22matrix in that equation can also be viewed as a representation of R.'/: U.'/D cos'sin' sin'cos'! : Because this group is abelian (two successive rotations yield the same result if applied in either order), we know that this representation is reducible. If we apply the unitary transformation U0.'/DVU.'/V1;with VD0 @1=p 2i=p 2 1=p 2i=p 21 A; the result is U0.'/D0 @cos'Cisin' 0 0 cos 'isin'1 AD ei'0 0ei'! : (17.6) 2We know this because some of these 22matrices do not commute with each other and therefore cannot be diagonalized simultaneously. ArfKen_Ch17-9780123846549.tex 17.2 Representation of Groups 825 Equation (17.6) applies to every element of our rotation group after transforming with V, and we see that every rotation now has a diagonal representation. In other words, U.'/has been transformed into a direct sum of two one-dimensional (1-D) representations, U0D U1U.1/, with U1.'/Dei'andU.1/.'/Dei'. In fact, these are only two of an infinite number of irreducible representations, all of dimension 1: Un.'/Dein'; where ncan have any positive or negative integer value, including zero. The reason nis limited to integer values is to assure that U.2/DU.0/ . Note that only the nvalues1 lead to faithful representations.  Exercises 17.2.1 For any representation Kof a group, and for any group element a, show that h UK.a/i1 DUK.a1/: 17.2.2 Show that these four matrices form a representation of the Vierergruppe, whose multi- plication table is in Table 17.2. ID1 0 0 1 ;AD1 0 01 ;BD0 1 1 0 ;CD01 1 0 : 17.2.3 Show that the matrices 1;A;B, andCofExercise 17.2.2 are reducible. Reduce them. Note. This means transforming BandCto diagonal form (by the same unitary transfor- mation). 17.2.4 (a) Once you have a matrix representation of any group, a 1-D representation can be obtained by taking the determinants of the matrices. Show that the multiplicative relations are preserved in this determinant representation. (b) Use determinants to obtain a 1-D representation of D3from the 22representa- tion in Eq. (17.2). 17.2.5 Show that the cyclic group of nobjects, Cn, may be represented by rm,mD 0;1;2;:::; n1. Here ris a generator given by rDexp.2 is=n/: The parameter stakes on the values sD1;2;3;:::; n, with each value of syielding a different 1-D (irreducible) representation of Cn. 17.2.6 Develop the irreducible 22matrix representation of the group of rotations (including those that turn it over) that transform a square into itself. Give the group multiplication table. Note. This group has the name D4(seeFig. 17.5). ArfKen_Ch17-9780123846549.tex 826 Chapter 17 Group Theory C4 C2C2z xC2′ C2′y FIGURE 17.5 D4symmetry group. 17.3 S YMMETRY AND PHYSICS Representations of groups provide a key connection between group theory and the sym- metry properties of physical systems. Our discussion will be directed mainly at quantum systems, but much of it will also apply to systems that can be described using classical physics. Consider a quantum system whose Hamiltonian Hpossesses certain geometric sym- metries. If we write HDTCV, the symmetries will be those of the potential energy V, since the kinetic energy operator Tis invariant with respect to rotations and displace- ments of the coordinate axes. A concrete example that illustrates the concept would be the determination of the wave function of an electron in the presence of nuclei in some fixed configuration possessing symmetry, such as the equilibrium locations of the nuclei in a symmetric molecule. The symmetry of Hcorresponds to a requirement that Hbe invariant with respect to the application of any element of its symmetry group. Letting Rdenote such a symmetry element, the invariance of Hmeans that if 'is a solution of the Schrödinger equation with energy E, then R'must also be a solution with the same energy eigenvalue: H.R'/DE.R'/: By successively applying the elements of our symmetry group to ', we can generate a set of eigenfunctions, all with the same eigenvalue. If 'happened to have the full sym- metry of H, this set would contain only one member and the situation would be easy to understand. But if 'had less symmetry,3our eigenfunction set would have more than one member, with its maximum possible size being the number of elements in our symmetry group. When the eigenfunction set has more than one member, the eigenfunctions do not individually have the complete symmetry of the Hamiltonian, but they form a closed set that permits the partial symmetry to be expressed in all symmetry-equivalent ways. For example, the hydrogenic eigenfunctions known as pstates form a three-membered set; 3This is possible; an example is a hydrogen-atom pstate. ArfKen_Ch17-9780123846549.tex 17.3 Symmetry and Physics 827 none has the full spherical symmetry of the hydrogen atom Hamiltonian, but linear com- binations of the three pstates can describe a porbital at an arbitrary orientation (obvious because a vector in an arbitrary direction can be written as a linear combination of vectors in the coordinate directions). So let’s assume that, starting from some chosen ', we have found a full set of symmetry- related eigenfunctions, have eliminated from them any linear dependence, and have formed an orthonormal eigenfunction set, denoted 'i,iD1:::N. Because of the way in which the 'iwere constructed, they will transform linearly among themselves if we apply to them any operation Rfrom our symmetry group, so we may write R'iDX jUji.R/'j: (17.7) If we apply two symmetry operations ( Rfollowed by S), the transformation rule for the result will be SR'iDX jkUkj.S/Uji.R/'k: (17.8) Equations (17.7) and (17.8) show that the transformation for the group element SRis the matrix product of those for SandR, so the matrices U.S/andU.R/have properties that make them members of a representation of our symmetry group. What is new here is that we have identified Uas a representation associated with the basis f'ig. At this point we do not know whether the representation formed from our f'igbasis is reducible; its reducibility depends on the quantum system under study and the particular choice made for the initial function '. If our Uare reducible, let’s assume we now apply a transformation that will convert them into the direct-sum form. The transformation to obtain the direct-sum separation corresponds to a division of the basis into smaller sets of functions that transform only among themselves. Our overall conclusion from the above analysis is: If a Hamiltonian H is fully symmetric under the operations of a symmetry group, all its eigenfunctions can be classified into sets, with each set forming a basis for an irreducible representation of the symmetry group. The members of a symmetry- related set of eigenfunctions will be degenerate and are referred to as a multi- plet. Ordinarily different multiplets will correspond to different eigenvalues; any degeneracy between eigenfunctions of different irreducible representations arises from sources other than the symmetry under study. Because the eigenfunctions of a Hamiltonian possessing geometric symmetry can be identified with irreducible representations of its symmetry group, it is natural to use approximate eigenfunctions with similar symmetry restrictions. Example 17.3.1 ANEVEN HAMILTONIAN Consider a Hamiltonian H.x/, which is even inx, meaning that H.x/DH.x/, but has no other symmetry. Letting stand for the reflection operator x! x(is the usual ArfKen_Ch17-9780123846549.tex 828 Chapter 17 Group Theory notation for a reflection operation), our symmetry group, called Cs, consists only of the two operations Iand, and its multiplication table is I ID DI;IDID: This group is abelian, and has two irreducible representations of dimension 1: one ( A1) that is completely symmetric, U.I/DU./D1, and one ( A2) with sign alternation, U.I/D1, U./D1 . The eigenfunctions of Hwill therefore be even or odd, and there is no inherent symmetry requirement that even and odd states be degenerate with each other. If we start with a function '.x/that is even, we will have I'D'D', so our basis will consist only of ', and U.I/DU./D1, indicating that the representation constructed using this basis will be the fully symmetric A1. On the other hand, if our starting function '.x/was odd, then I'D'but'D' ; again our basis will consist only of '.x/, but now the representation constructed from it will consist of U.I/D1,U./D1 , and will be the alternating-sign representation A2. But if we start with a function '.x/that is neither even nor odd, then I'.x/D'.x/, but'.x/D'.x/. Our assumption that '.x/is neither even nor odd means that '.x/ and'.x/are linearly independent, so our basis will consist of two members (and there- fore be of dimension 2). Since the symmetry group has only A1andA2as irreducible representations, the representation built from our two-membered basis will be reducible, and will reduce to A1A2. The basis will separate into the two members '.x/C'.x/ (a 1-D A1basis) and'.x/'.x/(anA2basis). Given a problem with an even Hamiltonian, one may use the above-identified symmetry analysis to seach for solutions that are restricted to have either even or odd symmetry. This strategy may greatly simplify the process of finding solutions. The notion can be extended to problems with different or greater degrees of symmetry.  It is important to note that all geometric symmetry groups (other than the trivial group, which has only the element I) will possess representations other than A1, which means that they will have bases of less symmetry than the original group. In Example 17.3.1, our Hamiltonian was even, but could have eigenfunctions that are either even ( A1) or odd ( A2). A Hamiltonian with D3symmetry (which we have already seen has irreducible represen- tations of dimensions 1 and 2) can have A1eigenfunctions of the full three-dimensional (3-D) symmetry or A2eigenfunctions with alternating-sign symmetry. It can also have sets of two degenerate eigenfunctions corresponding to the representation in Eq. (17.2), where (as indicated by the 22matrices) the symmetry operations can convert either of the basis members into linear combinations of both. The irreducibility means that there exists no single function built from this two-member basis that will remain the same (except for a possible sign or phase factor) under all the group operations. The existence of an irre- ducible basis with more than one member is a consequence of the fact that the symmetry group is not abelian. Although the elements of a symmetry group may not all commute with each other, they all commute with a Hamiltonian (or other operator) having the full group symmetry. To show this, note that for any eigenfunction and any group element R, H DE ! H.R /DE.R /DR.E /DRH ! HRDRH: The last step follows because the previous steps are valid for all members of a complete set of eigenfunctions . ArfKen_Ch17-9780123846549.tex 17.3 Symmetry and Physics 829 Sometimes, especially for continuous groups, we will know in advance how to construct bases for irreducible representations. For example, the spherical harmonics of a given l value form a basis for representation of the 3-D rotation group. From Chapter 16, we know that these spherical harmonics form a closed set under rotation, but only if the set includes allmvalues. This information, together with the orthonormality of the Ym l, tells us that Ym l, mDl;:::; lis an orthonormal basis of dimension 2lC1for an irreducible representation of the 3-D rotation group, which is named SO.3/ . In contrast to the situation for discrete groups, continuous groups (even of low order) may possess an infinite number of finite- dimensional irreducible representations. An experienced investigator can often find bases for irreducible representations by inspection or educated insight. However, if simple methods for finding a basis prove insuf- ficient, general methods can be used to construct basis functions if the matrices defining the relevant irreducible representation are available. Details of the process can be found in the works by Falicov, Hamermesh, and Tinkham (see Additional Readings). Example 17.3.2 QUANTUM MECHANICS, TRIANGULAR SYMMETRY Let’s consider a Hamiltonian that has the D3symmetry of an equilateral triangle that can be turned over, and our problem is such that its solution can be approximated as a wave function that is distributed over orbitals centered at the three vertices Riof the triangle, of the form .r/Da1'.r1/Ca2'.r2/Ca3'.r3/, where riis the distancejrRij, and'is a spherically symmetric orbital. The function 0D'.r1/C'.r2/C'.r3/ is a basis for the trivial ( A1) representation of the D3group. But because we have three orbitals, there will be two other linear combinations of them that are linearly independent of 0, and one way to choose them is 1D1p 2 '.r1/'.r3/ ; 2D1p 6 '.r1/C2'.r2/'.r3/ : Neither of these functions (nor any linear combination of them) has enough symmetry to be either A1orA2basis functions, and they therefore must (together) form a basis for a 22irreducible representation of the D3symmetry group that is called E. Knowing that this would be the case, we chose these functions in a way that makes them orthogonal and at a consistent normalization, and they are in fact a basis for the irreducible representation given in Eq. (17.2). We can check this by applying group operations to 1and 2, verifying that the result corresponds to the appropriate column of the matrix for the operation. We make one such check here: Applying C3to 1, we get C3 1DT'. r3/'.r2/U=p 2, while the first column ofU.C3/in Eq. (17.2) yields C3 1D1 2 1p 3 2 2D1 2'.r1/'.r3/p 2 p 3 2'.r1/C2'.r2/'.r3/p 6 : The reader can verify that these two expressions for C3 1are equal, and can make further checks if desired. ArfKen_Ch17-9780123846549.tex 830 Chapter 17 Group Theory One might think that because of the triangular symmetry there would be an irreducible representation of dimension 3. But mathematics is not that simple; all D3representations of dimension 3 are reducible!  The symmetries required of solutions to Schrödinger equations have implications beyond their role in causing or explaining degeneracy. The dominant interaction between an electromagnetic field and a molecule can occur only if the molecule has an electric dipole moment, and the presence of a dipole moment depends on the symmetry of the electronic wave function. Another context in which symmetry is important is in the evalu- ation of the expectation values of quantum operators. These expectation values will vanish unless the integrals that define them have integrands with a fully symmetric part. In addi- tion, it is worth mentioning that many quantum calculations are simplified by limiting them to contributions that do not vanish by reason of symmetry. All these issues can be framed in terms of the irreducible representations for which our wave functions are bases. In the next sections, we develop some key results of group representation theory, first for discrete groups because the analysis is simpler, and then (in less detail) for continuous groups that have become important in particle theory and relativity. Exercises 17.3.1 Consider a quantum mechanics problem with D3symmetry, with the threefold symme- try axis taken as the zdirection, and with orbitals '.rRj/located at the vertices of an equilateral triangle. This is the same system geometry as in Example 17.3.2, but in the present problem 'will no longer be chosen to have spherical symmetry. Given that'.r/D.z=r/f.r/(so'has the symmetry of a porbital oriented along the symmetry axis), construct linear combinations of the 'that are bases for irreducible representations of D3, for each basis indicating its representation. 17.4 D ISCRETE GROUPS Classes It has been found useful to divide the elements of a finite group Ginto sets called classes. Starting from a group element a1, one can apply similarity transformations of the form ga1g1, where gcan be any member of G. If we let a1be transformed in this way, using all the elements gofG, the result will be a set of elements that we can denote a1;:::; ak, where kmay or may not be larger than 1. Certainly this set will include a1itself, as that result is obtained when gDIand also when gDa1orgDa1 1. The set of elements obtained in this way is called a class ofG, and can be identified by specifying one of its members. If we choose a1DI, we find that Iis in a class all by itself; often classes will have larger numbers of members. A class will have the same members no matter which of its elements is assigned the role ofa1. This is clear, since if aiDga1g1then also a1Dg1aig, showing that we can get a1from any other element of the class, and therefrom all the elements reachable from a1. ArfKen_Ch17-9780123846549.tex 17.4 Discrete Groups 831 Example 17.4.1 CLASSES OF THE TRIANGULAR GROUP D3 As observed already in general, one class of D3will consist solely of I. The class including C3contains also C2 3(the result of C2C3C1 2). Finally, C2,C0 2, and C00 2constitute a third class.  Classes are important because: For a given representation (whether or not reducible), all matrices of the same class will have the same value of their trace—obvious because trace( gag1/Dtrace( ag1g/D trace( a). In the group theory world, the trace is also known as the character, custom- arily identified with the symbol 0. It can be shown that the number of inequivalent irreducible representations of a finite group is equal to its number of classes. (For proof and fuller discussion, see Additional Readings at the end of this chapter.) It can be shown (again, see Additional Readings) that the set of characters for all elements and irreducible representations of a finite group defines an orthogonal finite-dimensional vector space. Writing 0K.g/as the character of group element gin irreducible representa- tionK, we have the key relations, for a group of order n: ngX K0K.g/0K.g0/Dngg0;X g0K.g/0K0.g/DnK K0: (17.9) Here ngis the number of elements in the class containing g. These relations enable any reducible representation to be decomposed into a direct sum of irreducible representations, and can also be of aid in finding the characters of irreducible representations if they were not already known. Another theorem of great importance in the theory of finite groups, sometimes called thedimensionality theorem, is that the sum of the squares of the dimensions nKof the inequivalent irreducible representations is equal to the order, n, of the group: X Kn2 KDn: (17.10) This theorem, together with the theorem that the number of irreducible Kequals the num- ber of classes, imposes stringent limits on the number and size of the irreducible represen- tations of a group. These two requirements are often enough to determine completely the inventory of irreducible representations. Since the finite groups of interest in physics have been well studied, the most frequent use of these orthogonality relations is to extract from a basis that may be reducible (i.e., a basis for a possibly reducible representation) the irreducible bases that may be included therein. This task is usually carried out with a table of irreducible representations at hand. Example 17.4.2 ORTHOGONALITY RELATIONS, GROUP D3 The usual scheme for tabulating discrete group characters is called a character table; that for our triangle group D3is shown in Table 17.3. The rows of the table are labeled with ArfKen_Ch17-9780123846549.tex 832 Chapter 17 Group Theory Table 17.3 Character Table for Group D3 I 2C3 3C2 A1 1 1 1 A2 1 1 1 E 21 0 9 3 0 1 Each row corresponds to an irreducible representation, and each column corresponds to a class. The table entry is the character for each element of that irreducible representation and class. The row below the boxed table (labeled 9) is not part of the table but is used in connection with Example 17.4.4. the usual names assigned the irreducible representations: The labels AandB(the latter not used for this group) are reserved for 11representations. Representations of dimension 2 are normally assigned a label E, and those of dimension 3 (also not occurring here) are called T. Each column of the character table is labeled with a typical member of the class, preceded by a number indicating the number of group elements in the class. This number is omitted if the class contains only one element. Because the representation of group element Iis a unit matrix, the characters (traces) in column Idirectly indicate the dimensions of the representations. We see that A1is a 11representation, so each A1matrix contains a single number equal to the character shown, meaning that A1is the trivial totally symmetric representation. We see that A2is also11, but the three group elements for which the triangle was turned over are now represented by1. Finally, representation Eis seen to be 22, and is the representation we found long ago in Eq. (17.2). Checking the first orthogonality relation for gDg0DI, for which ngD1, we have 1.12C12C22/D6, as expected. For gDI,g0DC3, we have 1T1.1/C1.1/C2.1/UD 0, and for gDg0DC3, we note that ngD2and we have 2T12C12C.1/2UD6. The reader can check other cases of this orthogonality relation. Moving to the second orthogonality relation, we take KDK0DE, finding 1.22/C 2.1/2C3.02/D6; the 1, 2, and 3 multiplying individual terms allow for the fact that the sum is over all elements, not just over classes. Other cases follow similarly.  Example 17.4.3 COUNTING IRREDUCIBLE REPRESENTATIONS We consider two cases, first the group C4, which was the subject of Example 17.1.4. This group is cyclic, with elements I,a,a2,a3; those are all the elements, because a4DI. As already indicated, a faithful representation of this group consists of 1, i,1,i, with the group operation being ordinary multiplication. Another realization of C4is an object that is symmetric under 90rotation about a single axis. This group is abelian, as apaqDaqap. Then gag1Dafor any group elements aandg, so each element is in a class by itself. So we have four classes, and hence four irreducible representations. We also have, from the ArfKen_Ch17-9780123846549.tex 17.4 Discrete Groups 833 dimension theorem, 4X KD1n2 KD4: The only way to satisfy this equation is to have four irreducible representations, each of dimension 1. This result should have been expected, since C4is abelian. Our irreducible representations can be built from the four following choices of U.a/: 1,i,1,i, leading to the following character table. I a a2a3 A11 1 1 1 A21 i1i A311 11 A41i1 i Our second case is D3, which has six elements and the three classes identified in Example 17.4.1. This means that it has three irreducible representations with dimensions whose squares add to six. The only set of dimensions satisfying this requirements is 1, 1, and 2.  Example 17.4.2 can be generalized to deal with reducible representations; any repre- sentation whose characters do not match any row of the character table must be reducible (unless just wrong!). If we were to transform a reducible representation to direct-sum form, it would then be obvious that its trace will be the sum of the traces of its blocks, and that property will hold even if we do not know how to make the block-diagonalizing trans- formation. In group-theory lingo we would say that the characters of a reducible repre- sentation will be the sum of the characters of the irreducible representations it contains. Note that if a given irreducible representation occurs more than once, its characters must be added a corresponding number of times. Now suppose that we have a reducible representation 9of a group of order n. Even if we do not yet know its decomposition into irreducible components, we can write its characters for group elements gin the form 09.g/DX KcK0K.g/; (17.11) where cKis the number of times irreducible representation Kis contained in 9. If we multiply both sides of this equation by 0K0.g/and sum over g, the orthogonality kicks in, and X g0K0.g/09.g/DX gX KcK0K0.g/0K.g/DncK0: (17.12) Evaluating the left-hand side of Eq. (17.12), we easily solve for cK0. We can repeat this sequence of steps with different K0until all the irreducible representations in 9have been found. ArfKen_Ch17-9780123846549.tex 834 Chapter 17 Group Theory Example 17.4.4 DECOMPOSING A REDUCIBLE REPRESENTATION Suppose we start from the following set of three basis functions for the triangular group D34: 1Dx2; 2Dy2; 3Dp 2xy; (17.13) where x;y;zare Cartesian coordinates with origin at the center of the triangle, and the axes are in the directions shown in Fig. 17.1. Since C3xD1 2xC1 2p 3y,C3yD1 2p 3x 1 2y, we can (somewhat tediously) determine that C3x2D1 4x2C3 4y2r 3 8.p 2xy/; C3y2D3 4x2C1 4y2Cr 3 8.p 2xy/; C3.p 2xy/Dr 3 8x2r 3 8y21 2.p 2xy/; so in the basis, U9.C3/D0 BBB@1 43 4q 3 8 3 41 4q 3 8 q 3 8q 3 81 21 CCCA: (17.14) Similar analysis can be used to obtain the matrix of C2, which is easier because the opera- tion involved is just x! x, with yremaining unchanged. We get U9.C2/D0 B@1 0 0 0 1 0 0 011 CA: (17.15) The representation of Iis, of course, just the 33unit matrix. Since the only data we need right now are the traces of one representative of each class, we are ready to proceed, and we see that 09.I/D3; 09.C3/D0; 09.C2/D1: We are labeling the characters with superscript 9as a reminder that the representation is that associated with the i. These characters have been appended below their respective columns in Table 17.3. Now we use the fact that the 9representation must decompose into 9Dc1A1c2A2c3E; (17.16) 4These basis functions have been chosen in a way that makes the reducible representation unitary. The factorp 2in 3is needed to make all the iat the same scale. ArfKen_Ch17-9780123846549.tex 17.4 Discrete Groups 835 and we find the ciby applying Eq. (17.12). Using the data in Table 17.3, and taking K0to be in turn A1,A2, and E, A1V.1/.3/C2.1/.0/C3.1/.1/D6D6c1;soc1D1; A2V.1/.3/C2.1/.0/C3.1/1/D0D6c2;soc2D0; EV.2/.3/C2.1/.0/C3.0/.1/D6D6c3;soc3D1: Thus,9DA1E. We can check our work by summing the A1andEentries from the character table. As they must, they add to give the entries for 9.  For some purposes it is insufficient just to know which irreducible representations are included in a reducible basis for a group G. We may also need to know how to transform the basis so that each basis member will be associated with a specific irreducible represen- tation of G. Sometimes it is easy to see how to do this. For the above example, the basis function for A1must have the full group symmetry, while the Ebasis functions must be orthogonal to the A1basis. These considerations lead us to A1V'D 1C 2Dx2Cy2; (17.17) EV'1D 1 2Dx2y2; ' 2Dp 2 3D2xy: (17.18) However, if finding the irreducible basis functions by inspection proves difficult, there are formulas that can be used to find them. See Additional Readings. Other Discrete Groups Most of the examples we have used have been for one group, D3, in which we have considered symmetry operations that involve rotations about axes through the center of the system. Groups keeping a central point fixed are called point groups, and they arise, among other places, when studying phenomena that depend on the geometric symmetries of molecules. Some point groups have additional symmetries associated with inversion or reflection. It is possible for a point group to have a single n-fold axis for any positive inte- gern(meaning that a symmetry element is a rotation through an angle 2=n). However, the number of point groups having multiple symmetry axes with n3is very limited; they correspond to the Platonic regular polyhedra, and therefore can only be tetrahedral, cubic/octahedral, and dodecahedral/icosahedral. Other discrete groups arise when we consider permutational symmetry; the symmetric group is important in many-body physics and is the subject of a separate section of this chapter. Exercises 17.4.1 The Vierergruppe has the multiplication table shown in Table 17.2. (a) Divide its elements into classes. (b) Using the class information, determine for the Vierergruppe its number of inequiv- alent irreducible representations and their dimensions. (c) Construct a character table for the Vierergruppe. ArfKen_Ch17-9780123846549.tex 836 Chapter 17 Group Theory 17.4.2 The group D3may be discussed as a permutation group of three objects. Operation C3, for instance, moves vertex 1 to the position formerly occupied by vertex 2; like- wise vertex 2 moves to the original position of vertex 3 and vertex 3 moves to the original position of vertex 1. So this shuffling could be described as the permutation of (1,2,3) to (2,3,1). Using now letters a;b;cto avoid notational confusion, this permuta- tion.abc/!.bca/corresponds to the matrix equation C30 BB@a b c1 CCAD0 BB@0 1 0 0 0 1 1 0 01 CCA0 BB@a b c1 CCAD0 BB@b c a1 CCA; thereby identifying a 33representation of the operation C3. (a) Develop analogous 33representations for the other elements of D3. (b) Reduce your 33representation to the direct sum of a 11and a 22repre- sentation. Note: This 33representation must be reducible or Eq. (17.10) would be violated. 17.4.3 The group named D4has a fourfold axis of symmetry, and twofold axes in four direc- tions perpendicular to the fourfold axis. See Fig. 17.5. D4has the following classes (the numbers preceding the class descriptors indicate the number of elements in the class): I,2C4,C2,2C0 2,2C00 2. The twofold axes marked with primes are in the plane of fourfold symmetry. (a) Find the number and dimensions of the irreducible representations. (b) Given that all the characters of the representations of dimension 1 are 1and that C2DC2 4, use the orthogonality conditions to construct a complete character table forD4. 17.4.4 The eight functions x3,x2y,xy2,y3form a reducible basis for D4, with C4 a90counterclockwise rotation in the xyplane, C2DC2 4,C0 2D.x! x;y!y/, C00 2D.x!y;y!x/, and the remaining members of D4are additional members of the classes containing the above operations. Find the characters of the reducible repre- sentation for which these functions form a basis, and find the direct sum of irreducible representations of which it consists. 17.4.5 The group C4vhas a fourfold symmetry axis in the zdirection, reflection symmetries (v) about the xzandyzplanes, and additional reflection symmetries ( d,dDdihedral) about planes that contain the zaxis but are 45from the xandyaxes. See Fig. 17.6. The character table for C4vfollows. I 2C4 C2 2v2d A11 1 1 1 1 A21 1 1 11 B111 1 1 1 B211 11 1 E2 02 0 0 ArfKen_Ch17-9780123846549.tex 17.5 Direct Products 837 σv′σd′ σd σvy x C4 FIGURE 17.6 C4vsymmetry group. At left, a molecule with this symmetry. At right, a diagram identifying the reflection planes, which are perpendicular to the plane of the diagram. (a) Construct the matrix representing one member of each class of C4vusing as a basis a pzorbital at each of the points .x;y/D.a;0/,.0;a/,.a;0/,.0;a/, and therefrom extract the characters of the reducible representation for which these pz orbitals form a basis. A pzorbital has functional form .z=r/f.r/. (b) Determine the irreducible representations contained in our reducible pzrepresen- tation. (c) Form those linear combinations of our pzfunctions that are bases for each of the irreducible representations found in part (a). 17.4.6 Using the notation and geometry of Exercise 17.4.5, repeat that exercise for the eight- member basis consisting of a pxand a pyorbital at each of the points .x;y/D.a;0/, .0;a/,.a;0/,.0;a/. 17.5 D IRECT PRODUCTS Many multiparticle quantum-mechanical systems are described using wave functions that are products of individual-particle states. This approach is that of an independent-particle model, which at a higher degree of approximation can include interparticle interactions. The single-particle states can then be chosen to reflect the symmetry of the system, mean- ing that each one-particle state will be a basis member of some irreducible representation of the system’s symmetry group. This idea is obvious, for example, in atomic structure, where we encounter notations such as 1s22s22p3(the ground-state electron configuration of the N atom). ArfKen_Ch17-9780123846549.tex 838 Chapter 17 Group Theory When a multiparticle system with symmetry group Gis subjected to one of its sym- metry operations, each single-particle factor in its wave function transforms according to its individual irreducible representation of G, so the overall wave function may contain products of arbitrary components of each particle’s representation. Thus, the multiparticle basis consists of all the products that can be formed by taking one member of each single- particle basis. This is what is termed a direct product. This multiparticle basis will also constitute a representation of G. The notation KDK1 K2 indicates that the representation KofGis the direct product of the representations K1and K2. This means also that the representation matrix UK.a/of any element aofGcan be formed as the direct product (see Eq. 2.55) of the matrices UK1.a/andUK2.a/. The representation of a group Gformed as a direct product of two (or more) of its irre- ducible representations may or may not be irreducible. For finite groups, a useful theorem is that the characters for a direct product of representations are, for each class, the product of the individual characters for that class. Once the characters for the direct product have been constructed, the methods of the previous section can be used to find the irreducible components of the product states. Example 17.5.1 EVEN-ODD SYMMETRY Sometimes the analysis of a direct product is simple. Consider a system of nindependent particles subject to a potential whose only symmetry element (other than I) is inversion (denoted i) through the origin of the coordinate system, so V.r/DV.r/. In this case, G(conventionally named Ci) has the two elements Iandi, with the following character table. I i Ag1 1 Au11 Individual particles with A1wave functions, which remain unchanged under inversion, are conventionally labeled g(from the German word gerade). Particles with A2wave functions, which change sign on inversion, are labeled u, for ungerade. In fact, the usual notation for the character table of the Cigroup writes AgandAuin place of A1and A2, thereby conveying more information about the symmetries of the corresponding basis functions. Now suppose that this system is in a state with jof the particles in ustates and njof the particles in gstates. Intuitively, we know that if jis an odd number, the overall wave function will change sign on inversion, but will not change sign if jis even. Formally, we examine the direct product representation K: KDu.1/ u.2/  u.j/ g.jC1/  g.n/: ArfKen_Ch17-9780123846549.tex 17.5 Direct Products 839 Using the theorem that the characters of representation Kcan be obtained by multiplying those of its constituent factors, we find 0K.I/D1,0K.i/D.1/j. Irrespective of the value of j,Kwill be irreducible: It is Agifjis even, and Auifjis odd.  Example 17.5.2 TWO QUANTUM PARTICLES IN D3SYMMETRY This case is not as simple. Suppose both particles are in states of Esymmetry, a situa- tion spectroscopists would identify with the notation e2; they use lower-case symbols to identify individual-particle states, reserving capital letters for the overall symmetry desig- nation. For definiteness, let’s further suppose5that each particle has a wave function of the form found in Eq. (17.18), so particle iwill have the two-member basis 'a.i/D x2 iy2 i ; ' b.i/D2xiyi; and the product basis will therefore have the four members 8aaD'a.1/' a.2/; 8 abD'a.1/' b.2/; (17.19) 8baD'b.1/' a.2/; 8 bbD'b.1/' b.2/: The matters at issue are (1) to find the overall symmetries this system can exhibit, and (2) to identify the basis functions for each symmetry. Consulting Table 17.3, we compute the products for e e: I2C33C2 e eV4 1 0: Since this representation has dimension 4 while the largest irreducible representation has dimension 2, it must be reducible. Applying the technique of Example 17.4.4, we can find that it decomposes into e eDA1A2E, a result that is easily checked by adding entries in the D3character table. A set of basis functions corresponding to the decomposition into irreducible representa- tions are A1D.x2 1y2 1/.x2 2y2 2/C4x1y1x2y2; (17.20) A2D2h .x2 1y2 1/x2y2x1y1.x2 2y2 2/i ; (17.21) E 1D.x2 1y2 1/.x2 2y2 2/4x1x1y1x2y2; (17.22) E 2D2h .x2 1y2 1/x2y2Cx1y1.x2 2y2 2/i : Finding these could be challenging; verifying them is less so.  For continuous groups, it is usually simpler to decompose direct-product representations in other ways. For example, in Chapter 16 we used ladder operators to identify overall 5An actual problem will have a wave function that, in addition to the functional dependence shown here, will have a completely symmetric additional factor that is not relevant for the present group-theoretic discussion. ArfKen_Ch17-9780123846549.tex 840 Chapter 17 Group Theory Table 17.4 Character Table, Group C4v I 2C4 C2 2v 2d A1 1 1 1 1 1 A2 1 1 1 11 B1 11 1 1 1 B2 11 1 1 1 E 2 02 0 0 angular-momentum states (irreducible representations) formed from products of individual angular momenta. The resulting multiplets correspond to the irreducible representations, and the angular momentum functions that we found are their bases. Exercises 17.5.1 The group C4vhas eight elements, corresponding to the rotational and reflection sym- metries of a square that cannot be turned over. See Fig. 17.6. Symmetry rotations about thezaxis are denoted C4,C2,C0 4. Reflections relative to the xzandyzplanes are named vand0 v; those at 45relative to the xzandyzplanes are called dand0 d(d indicates “dihedral”). The character table for C4vis in Table 17.4. (a) Find the direct sum of irreducible representations of C4vcorresponding to the direct product E E. (b) A basis for E(in the context of Fig. 17.5) consists of the two functions '1Dx, '2Dy. Apply a few of the group operations to this basis and verify the entries for Ein the character table. (c) Assume now that we have two sets of variables, x1,y1andx2,y2, and we form the direct-product basis x1x2,x1y2,y1x2,y1y2. Determine how the direct-product basis functions can be combined to form bases for each of the irreducible repre- sentations in the direct sum corresponding to E E. 17.6 S YMMETRIC GROUP Thesymmetric group S nis the group of permutations of ndistinguishable objects, and is therefore of order nW. To see this, note that to make a permutation, we may choose the first object in ndifferent ways, then the second in n1ways, etc., until we reach the nth object, which can be chosen in only one way. The total number of possible permutations is therefore n.n1/:::.1/DnW. This group is important in the physics of identical-particle systems, whose wave functions must be either symmetric with respect to particle inter- changes (particles with this symmetry are called bosons), or antisymmetric under pairwise particle interchanges (these particles are called fermions). This means that an n-boson wave function 9B.1;2;:::; n/must satisfy P9B.1;:::; n/D9B.1;:::; n/; (17.23) ArfKen_Ch17-9780123846549.tex 17.6 Symmetric Group 841 where Pis any permutation of the particle numbers. From a group-theoretical viewpoint this means that 9Bis a sole basis function for the trivial A1representation of Sn:11, with all members of the representation equal to (1). Many-fermion wave functions 9F.1;:::; n/ satisfy P9F.1;:::; n/DP9F.1;:::; n/; (17.24) wherePis the n-particle Levi-Civita symbol with an index string corresponding to P; in simple language this means PD1ifPis an even permutation of the particle numbers (one requiring an even number of pairwise interchanges), and PD1 ifPisodd. This means that9Fis the sole basis function for the 11totally antisymmetric representation ofSnwith members ( P), which we will call A2. Since the representations needed for either bosons or fermions are simple and of dimension 11, it might seem that sophisticated group-theoretic considerations would be unnecessary. But that is an oversimplification, because many-fermion systems (and some boson systems) consist of direct products of spatial and spin functions, and the spin func- tions may form a basis of Snof dimension larger than one. Example 17.6.1 TWO AND THREE IDENTICAL FERMIONS In elementary quantum mechanics, the ground state of a two-fermion system such as the two electrons of the He atom can be treated using a simple wave function of the form 9FD f.1/g.2/Cg.1/f.2/ .1/ .2/ .1/ .2/ : Here fandgare single-particle spatial functions, and , describe single-particle spin states. We continue, using a streamlined notation in which the particle numbers are suppressed, understanding that they always occur in ascending numerical order, so 9F will henceforth be written .f gCg f/. /. It is obvious that 9Fhas the fermion (anti)symmetry; we note that it is an A2basis function, which is the product of a symmet- ricA1spatial function and an antisymmetric A2spin function. The physics of this problem demands that the overall ground-state wave function 9Fcontain spin function because it is a two-particle spin eigenstate. The two-particle example shows that the A2 overall representation was obtained as A1 A2. For three particles, things are different. To treat the ground state of the Li atom, we cannot form a completely antisymmetric spin function using only the two single-particle spin functions and . The actual spin functions relevant for the ground state form a 22 representation of Sn, which we will call E: 1D1p 6.2 /;  2D1p 2. /: (17.25) Since permutations mix 1and2, the overall wave function for this three-particle system must be of the form 9FD11C22; ArfKen_Ch17-9780123846549.tex 842 Chapter 17 Group Theory where1and2are three-body spatial functions such that 9Fhas the required A2sym- metry. If theiare built from spatial orbitals f,g, and h, one possible set of iare 1D1 2.gh fh f ghg fCf hg/; (17.26) 2D1p 3 f ghCg f h1 2gh f1 2h f g1 2hg f1 2f hg ; a result that it is difficult to find by trial and error. Since the spin functions, and therefore also the spatial functions, become more complicated as the system size increases, the value of a group-theoretic description clearly becomes more urgent.  We consider now, from a formal viewpoint, only the many-fermion case. As illustrated in Example 17.6.1, we deal with space-spin functions in which the spin function has, for reasons we will not discuss here, been chosen to be built from an irreducible representation Kof the symmetric group, whose member for permutation Pis a unitary matrix designated UK.P/, and whose basis is a set of spin functions i,iD1;:::; nK, where nKis the dimension of the spin representation. This means that PiDnKX jD1UK ji.P/j: (17.27) We shall now show that an antisymmetric overall space-spin function can result if we form 9FDnKX iD1ii; (17.28) where theiare basis functions for a representation K0, of the same dimension as K, meaning that PiDnKX kD1UK0 ki.P/k: (17.29) The representation K0is assumed to have members that satisfy UK0.P/DPUK.P/: (17.30) The representation K0must exist, since it is (apart from a complex conjugate) the direct product of representations KandA2. Because A2only imparts sign changes to various UK, the representation K0will be irreducible because representation Kis. The representation K0is termed dual to representation K. ArfKen_Ch17-9780123846549.tex 17.6 Symmetric Group 843 To verify that the assumed form of 9Fhas the required A2symmetry, we take it, as given in Eq. (17.28), and apply to it an arbitrary permutation P: P9FDnKX iD1.Pi/.Pi/DX i X kUK0 ki.P/k!0 @X jUK ji.P/j1 A DX jk X iUK0 ki.P/UK ji.P/! kj DX jk X iPUK ki.P/UK ji.P/! kj: (17.31) The steps taken in the processing of Eq. (17.31) are substitutions of Eqs. (17.27) and (17.29) forPiandPi, followed by a conversion from UK0toUKthrough the use ofEq. (17.30). We complete our analysis by recognizing that because Uis unitary, Uki.P/D.U1/ik.P/, so X iPUK ki.P/UK ji.P/DPjk; leading to the final result P9FDX jkPjkkjDPX kkkDP9F: (17.32) Equation (17.32) shows that the overall wave function 9Fhas the required fermion anti- symmetry. Our only remaining problem is to construct spatial functions k, which are bases for rep- resentation K0. We state without proof (see Additional Readings) that this can be accom- plished using the formula j iDX PUK0 i j.P/P0; (17.33) where0is a single spatial function whose permutations will be used to construct the i. The index jidentifies an entire set of i; if0has no permutational symmetry, we can create sets of iinnK0in different ways, each corresponding to a different value of j. Example 17.6.2 CONSTRUCTION OF MANY-BODY SPATIAL FUNCTIONS We consider a three-electron problem in which the spin states are given by Eq. (17.25). We need the representation of S3for which these iare a basis. We are fortunate to already have this representation, as S3is isomorphic (in 1–1 correspondence) with D3, so we can use the set of 22representation matrices given in Eq. (17.2), if we make the identification C2$P.12/,C0 2$P.13/,C00 2$P.23/, where P.i j/denotes the permutation that inter- changes the ith and jth items in the ordered list to which the permutation is applied. The permutation P.123!312/ corresponds to C3, and P.123!231/ corresponds to C2 3. ArfKen_Ch17-9780123846549.tex 844 Chapter 17 Group Theory We now apply Eq. (17.33); an easy way to do this is to start by generating the matrix T that results from keeping all iandj. In the present case, that means forming the matrix sum TDU.I/0U.C2/P.12/ 0U.C0 2/P.13/ 0U.C00 2/P.23/ 0 CU.C3/P.123!312/ 0CU.C2 3/P.123!231/ 0: The minus signs for the U.C2/terms arise from the Pwhich is needed to convert UK intoUK0. Taking0as the product f.1/g.2/h.3/, hereafter written f gh, and inserting numerical values for the U, we reach TD0 BBBB@f ghg f h1 2.gh f1 2p 3.gh fh f g Ch f ghg ff hg/hg fCf hg/ 1 2p 3.gh fCh f g f ghCg f h1 2.gh f hg fCf hg/Ch f gChg fCf hg/1 CCCCA: (17.34) Each column of Eq. (17.34) defines a set of i, in a form that is not guaranteed to be normalized. From the second column, dividing through byp 3for normalization, we obtain theithat were listed as a possible wave function in Example 17.6.1 atEq. (17.26). The first column of Eq. (17.34) shows that there is a second possibility for an antisymmetric wave function built from the spatial product f gh, namely one that can be written 90 FD0 11C0 22; with the normalized spatial functions 0 1D1p 3 f ghg f h1 2gh f1 2h f gC1 2hg fC1 2f hg ; 0 2D1 2.gh fCh f ghg fCf hg/:  Exercises 17.6.1 (a) The objects .abcd/are permuted to .dacb/. Write out a 44matrix representa- tion of this one permutation. Hint: Compare with Exercise 17.4.2. (b) Is the permutation .abdc/!.dacb/odd or even? (c) Is this permutation a possible member of the D4group, which was the subject of Exercise 17.4.3? Why or why not? 17.6.2 (a) The permutation group of four objects, S4, has 4WD 24elements. Treating the four elements of the cyclic group, C4, as permutations, set up a 44matrix representation of C4. Note that C4is a subgroup of P4. (b) How do you know that this 44matrix representation of C4must be reducible? ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 845 17.6.3 The permutation group of four objects, S4, has five classes. (a) Determine the number of elements in each class of S4and identify one element of each class as a product of cycles. (b) Two of the irreducible representations of S4are of dimension 1 (and are usually denoted A1andA2). Noting that permutations can be classified as even or odd, find the characters of A1andA2. Hint. Set up a character table and fill in the A1andA2lines. (c) One irreducible representation of S4(usually denoted E) is of dimension 2. Deter- mine the dimensions of all the irreducible representations of S4other than A1,A2, andE. (d) Complete the character table of S4. Hint. Only the even permutations have nonzero characters in the Erepresentation. 17.7 C ONTINUOUS GROUPS Several continuous groups whose importance in physics was recognized long ago corre- spond to rotational symmetry in two- or three-dimensional space. Here the group elements are the rotations, the angles of which can vary continuously and thereby assume an infinite number of values. For rotations, the group multiplication rule corresponds to the applica- tion of successive rotations, which we have seen can be described by matrix multiplication. Rotations clearly form a group since they contain an identity element (no rotation), suc- cessive rotations are equivalent to a single rotation, and every rotation has an inverse (its reverse). Rotations in two-dimensional (2-D) space can be described by 22orthogonal matrices with determinantC1; the group consisting of these rotations is named SO(2) (SO stands for “special orthogonal”). If we also include reflections, so that the determinant can be 1, the group is named O(2). Since a 2-D rotation is completely specified by a single angle, SO(2) is a one-parameter group. A matrix representation of SO(2) was introduced inEq. (17.1); the group parameter is the rotation angle '. Rotations in 3-D space are described by 33orthogonal matrices. The resulting groups are designated O(3) and SO(3); for SO(3), three angles (e.g., the Euler angles) are group parameters. Generalizing to nnmatrices, the groups are named O.n/andSO.n/; the number of parameters needed to specify fully an nnreal orthogonal matrix is n.n1/=2, and that is the number of independent parameters (generalizations of the Euler angles) needed in SO.n/. If we further generalize to unitary matrices, we have the groups SU.n/ andU.n/. Proof that these sets of unitary matrices form groups is left as an exercise. Let’s introduce some nomenclature. The nnmatrices referred to above can be thought of as the defining, or fundamental, representations of the groups involved. The order of a continuous group is defined as the number of independent parameters needed to specify its fundamental representation, so the order of SO(n) is the previously stated n.n1/=2; the order of the group SU.n/isn21. In addition to their use for the treatment of rotational symmetry, continuous groups are also relevant to the classification of elementary (and not so elementary) particles. It has been experimentally observed that regularities in the masses and charges of sets of particles ArfKen_Ch17-9780123846549.tex 846 Chapter 17 Group Theory can be explained if their wave functions are identified as basis members of an irreducible representation of an appropriate group. Note that now the group does not describe rotations in ordinary space, but refers to a more abstract space relevant to an understanding of the physics involved. The earliest example of this idea was electron spin; spin wave functions are objects in an abstract SU.2/ space, together with rules to unravel their observational properties. A further abstraction began with the notion that the proton and neutron might form a basis for an abstract SU.2/ representation, and has since blossomed with the intro- duction of SU(3) and other continuous groups into particle physics. A brief survey of these ideas is presented in our specific discussion of SU.3/ . Lie Groups and Their Generators It is extremely useful to manage groups such as SO.n/orSU.n/in ways that do not explicitly involve an infinite number of elements; a formalism for doing so was devised by the Norwegian mathematician Sophus Lie. Groups for which Lie’s analysis is appli- cable, called Lie groups, have elements that depend continuously on parameters that vary over closed intervals (meaning that the parameter set includes the limit of any converging sequence of parameters). The groups SO.n/andSU.n/are Lie groups. Lie’s essential idea was to describe a group in terms of its generators, a minimal set of quantities that could be used in a specific way (multiplied by parameters) to produce any element of the group. Our starting point is, for each parameter 'controlling a group operation, to introduce a generator Swith the property that when 'is infinitesimal (and therefore written ') the group element with parameter '(which must be close to the identity element of the group) can be represented by U.'/D1Ci'S: (17.35) The factor iinEq. (17.35) could have been included in Sbut it is more convenient not to do so. Group operations corresponding to larger values of 'can now be generated from repeated operation ( Ntimes) by'=N, where'=Nis small. We therefore identify U.'/ as the limit U.'/Dlim N!1 1Ci'S NN I This large- Nlimit defines the exponential, so we have the general result U.'/Dexp.i'S/: (17.36) Given any representation Uof our continuous group, we can find the generator Scor- responding to the parameter 'for that representation by differentiation of Eq. (17.36), evaluated at the identity element of our group. In particular, idU.'/ d' 'D0DS; (17.37) revealing that the entire behavior of a representation Ucan be deduced from its behavior in an infinitesimal parameter-space neighborhood of the identity operation. However, to obtain complete knowledge of the structure of a Lie group we need to study the behavior ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 847 of its generators for a representation that is faithful; for that purpose it is desirable to use the fundamental representation. Example 17.7.1 SO(2) GENERATOR SO(2) involves rotational symmetry about a single axis, and its operations are counter- clockwise rotations of the coordinate axes through angles '. Working with the 22 fundamental representation of SO(2), an infinitesimal rotation 'causes (to first order) .x0;y0/D.xCy';yx'/, or x0 y0! D 1' ' 1! x y! D" 1 0 0 1! C' 0 1 1 0!# x y! D1Ci'S; with iSD 0 1 1 0! ; or SD 0i i0! D2; (17.38) where 2is a Pauli matrix. A general rotation is then represented by Eq. (17.36) as U.'/Dei'SD12cos'Ci2sin'D cos' sin' sin'cos'! ; (17.39) where we have evaluated the exponential of the matrix in Eq. (17.39) using the Euler identity, Eq. (2.80). This equation can be recognized as the transformation law for a 2-D coordinate rotation, Eq. (3.23), verifying that the generator formalism works as expected. If we had started from the final expression for U.'/ given in Eq. (17.39), we could have generated Sfrom it by applying the differentiation formula, Eq. (17.37).  The generator form, Eq. (17.36), has some nice features: 1. For the groups SO.n/andSU.n/, anyUwill be unitary (remember, “orthogonal” is a special case of “unitary”). This means that U1Dexp. i'S/DU†Dexp. i'S†/; (17.40) soSDS†, showing that Sis Hermitian. That is the proximate reason for inclusion ofiin the defining equation for S. 2. Because for both SO.n/andSU.n/,det.U/D1, we also have, invoking the trace formula, Eq. (2.84), det.U/Dexp trace.ln U/ Dexp i'trace.S/ D1: (17.41) This condition is satisfied for general 'only if trace.S/D0. SoSis not only Hermi- tian, but traceless. 3. It can be shown (but is not proved here) that the number of independent generators of a Lie group is equal to the order of the group. ArfKen_Ch17-9780123846549.tex 848 Chapter 17 Group Theory One of Lie’s key observations was that by focusing on infinitesimal group elements, various properties of the generators could be deduced. We have already seen that if the form of Uin terms of its parameters is known, the generators Scan be obtained by differ- entiation of Eq. (17.36) in the limit corresponding to the identity group element. Second, relations between the generators can be developed, as follows: Let us consider two operations Uj.j/andUk.k/of a group G, that respectively correspond to the gen- erators SjandSk. The values of jandkare assumed small, so the resulting Ujand Ukdiffer, but only slightly, from the identity element. Expanding the exponentials and keeping terms through second order in , UjDexp.ijSj/D1CijSj1 22 jS2 jC; UkDexp.ikSk/D1CikSk1 22 kS2 kC; we evaluate the leading term (in ) of the matrix product U1 kU1 jUkUj. The linear terms all cancel, as do several of the quadratic terms. The remaining quadratic terms can be grouped so as to reach the result U1 kU1 jUkUjD1CjkTSj;SkUC D1CijkX lfjklSlC: (17.42) The last line of Eq. (17.42) reflects the fact that the left-hand side of the equation must correspond to some group element, and that element must, to first order in the generators, be of the form shown. Note that the premultipliers ijkare not a form restriction, as their presence simply changes the value of fjkl. Comparing the two lines of Eq. (17.42), we obtain the important closure relation among the generators of the group G: TSj;SkUDiX lfjklSl: (17.43) The coefficients fjklare called the structure constants ofG. It can be shown that fjklis antisymmetric with respect to index permutations, so fjklDfkl jDfl jkDfkjlDflkjD fjlkj. The structure constants provide a representation-independent characterization of a Lie group, but as already mentioned, to determine them we will need to work with a faithful representation, such as the group’s fundamental representation. We will shortly do so for the groups we study in detail. As is obvious from the foregoing analysis, Lie group generators will not in general com- mute. In 3-D, rotations about different axes do not commute, and therefore their generators cannot commute either. An additional indicator for group classification is the maximum number of independent generators that all mutually commute. This number is called the rank of the group; it is significant because the generators can be subjected to unitary trans- formations without changing the ultimate group structure, and the mutually commuting generators can therefore be brought simultaneously to diagonal form. Once this is done, the basis members of the generator set can be labeled using the diagonal elements (the ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 849 eigenvalues) of the commuting generators. The values of the labels (and the physical phe- nomena related thereto) depend on the representation in use. For the orthogonal groups SO.n/and unitary groups SU.n/the commutation relations, Eq. (17.43), can be developed along the lines of angular momentum, leading to generalized ladder operators (and selection rules) in conjunction with the mutually commuting opera- tors. For these central aspects of (the so-called classical) Lie groups we refer to the work by Greiner and Mueller (see Additional Readings). Summarizing, the rank of a group indicates the number of indices needed to label the basis. In applications to quantum mechanics, these indices are often referred to as quantum numbers. For example, in SO.3/ , which is of rank 1, the index is usually taken to be ML, usually identified physically as the zcomponent of an angular momentum; when SU.2/ , also of rank 1, is used for the description of electron spin, the index is usually called MS. The possible values of MLorMSdepend on the representation, and we saw in Chapter 16 that the values range, in unit steps, between CLandL(orCSandS), so that diagrams identifying these basis members can be plotted on a line. In contrast, we will see that SU.3/ is of rank 2, so its basis members are labeled with two quantum numbers. Diagrams identifying the label assignments will in that case need to be 2-D. It is also possible to label entire representations. One way to label them is to use the eigenvalues of operators that commute with all the generators of the group; such operators are called Casimir operators; the number of independent Casimir operators is equal to the rank of the group. SO.3/ has therefore one Casimir operator; it is the operator usually known in angular-momentum applications as L2orJ2. Groups SO.2/ andSO.3/ SO.2/ andSO.3/ are rotation groups; SO.2/ corresponds to rotational symmetry about one axis, which we will take to be the zaxis when the symmetry is for a 3-D system. SO.2/ will therefore have only one generator, that already found in Eq. (17.38): SzD2D 0i i0! : (17.44) To use Szas one of the generators of SO.3/ , we extend to a 33basis, calling the generator S3, obtaining S3D0 B@0i0 i 0 0 0 0 01 CA: (17.45) SO.3/ has two other generators, S1andS2. To obtain S1, the generator corresponding to Ux. /D0 B@1 0 0 0 cos sin 0sin cos 1 CA; (17.46) ArfKen_Ch17-9780123846549.tex 850 Chapter 17 Group Theory we apply Eq. (17.37): S1Did Rx./ d  D0Di0 B@0 0 0 0sin cos 0cos sin 1 CA D0D0 B@0 0 0 0 0i 0i01 CA:(17.47) In a similar fashion, starting from Uy./D0 B@cos0sin 0 1 0 sin0 cos1 CA; (17.48) we find S2D0 B@0 0 i 0 0 0 i0 01 CA: (17.49) Summarizing, the structure of SO.2/ is trivial, as it has only a single generator, and has order 1 and rank 1. However, the structure of SO.3/ is not entirely trivial. Because no two ofS1,S2, andS3commute, SO.3/ will have order 3, but rank 1. By matrix multiplication, we may compute its structure constants. It is easily verified that TSj;SkUDijklSl; (17.50) wherejklis a Levi-Civita symbol. Thus, the Levi-Civita symbols are the structure con- stants for SO.3/ . Note also that the Sjobey the angular momentum commutation rules. In fact, these are the same matrices that were called Kiin Eq. (16.86) in Chapter 16, and they were identified there as matrices describing the components of angular momentum in a basis consisting of x,y, and z. This observation can be generalized to reach the conclu- sion that for any representation of SO.3/ , the generators can be taken to be the angular momentum components Lj(jD1;2;3) as expressed in any basis for that representation. Example 17.7.2 GENERATORS DEPEND ON BASIS To show that the generators indeed have a form that depends on the choice of basis, con- sider a basis for SO.3/ proportional to the spherical harmonics for lD1with standard phases, 1D1p 2.xCiy/; 2Dz; 3D1p 2.xiy/: (17.51) We now apply LxDiTy@=@zz@=@yUto the basis members, getting the result Lx 1D z=p 2D 2=p 2,Lx 2DiyD. 1C 3/=p 2,Lx 3Dz=p 2D 2=p 2, meaning that the matrix representation of Lx, and therefore of a generator we will call Sx, is SxD1p 20 @0 1 0 1 0 1 0 1 01 A: (17.52) ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 851 Applying LyandLzto the spherical harmonic basis, we obtain generators SyandSz: SyD1p 20 B@0i0 i0i 0i01 CA;SzD0 B@1 0 0 0 0 0 0 011 CA: (17.53) These generators, though different from those given in Eqs. (17.45), (17.47), and (17.49), are equivalent to them in the sense that they define the same irreducible representation ofSO3.  Group SU(2) and SU(2)–SO(3) Homomorphism A complete set of generators for the fundamental representation of SU.2/ must span the space of traceless 22Hermitian matrices; since there is only one off-diagonal element above the diagonal that can have an arbitrary complex value, it can, if nonzero, be assigned in two linearly independent ways (such as 1 and i). The below-diagonal element is then completely determined by Hermiticity. There is only one independent way to assign the diagonal elements, as there are two and they must be real and sum to zero. Thus, a simple set of matrices satisfying the necessary conditions consists of the three Pauli matrices j, jD1;2;3. Noting also that there would be advantages to having the generators scaled so that they would satisfy the angular momentum commutation relations, we choose the definition SjD1 2j;jD1;2;3: (17.54) Then, based on our many previous encounters or by performing the matrix multiplications, we can confirm TSj;SkUDijklSl: (17.55) In addition, for rotation parameters denoted as jin connection with generators Sj, we have, calling the corresponding SU.2/ members Uj, Uj. j/Dexp.i jj=2/; jD1;2;3: (17.56) Invoking the Euler identity, Eq. (2.80), we can rewrite Eq. (17.56) as Uj. j/D12cos j 2 Cijsin j 2 : (17.57) The group SU.2/ was first recognized as relevant for physics when it was observed that spin states of the electron form a basis for its fundamental representation. We already know, from Chapter 16, that orbital angular momentum multiplets come in sets with odd numbers of members ( 2LC1, with Lintegral). But we also observed that abstract quanti- ties that obey the angular momentum commutation rules with half-integer Lvalues come in multiplets with even numbers of members. The multiplet with two members is the fun- damental basis for the group SU.2/ . These basis functions are conventionally written j"i andj#i, (or just and ), and in matrix notation are j"iD1 0 ;j#iD0 1 : (17.58) ArfKen_Ch17-9780123846549.tex 852 Chapter 17 Group Theory Since the structure constants for SU(2) show that its generators satisfy the angular momen- tum commutation rules, we may conclude that all angular momentum multiplets define representations of SU(2); in Chapter 16 we found that the multiplets of odd dimension (2LC1with Lintegral) can be chosen to be the spherical harmonics of angular momen- tumLand are therefore also a basis for a representation of SO(3). Angular momentum multiplets of even dimension do not have a 3-D spatial representation and cannot corre- spond to a representation of SO(3). They are the more abstract quantities we call spinors, have half-integer angular-momentum quantum numbers, and are bases only for represen- tations of SU(2). Further understanding of the situation can be obtained by applying Ux.'/, a synonym forU1.'/, to the spin function j"i. Taking'D, this corresponds to a 180rotation about thexaxis, which we might expect would convert j"iintoj#i. Applying Eq. (17.57), which for the current case assumes the form UxDi1, we have Uxj"iDi0 1 1 01 0 Di0 1 Dij#i: (17.59) So far, so good. But let’s now try a similar rotation with 'D2. We then have UxD1 2, meaning that a complete 360rotation does not restore j"i, but gives insteadj"i, namely the expected state, but with a change of sign. To recover j"iwith its original (+) sign would require a rotation 'D4, i.e., two revolutions. Each rotation between 'D2and'D4 is, with opposite sign, equivalent to one in the .0;2/range. We now see the essential difference between SU.2/ andSO.3/ : The angular range of the rotation parameters in SU.2/ is twice that in SO.3/ , so each SO.3/ element is gen- erated twice in each dimension (with different signs) in SU.2/ . Thus the correspondence between the two groups is not one-to-one (an isomorphism), but is two-to-one, a homo- morphism. The existence of this homomorphism is not important for irreducible represen- tations of odd dimension (corresponding to integer Lor, in more general contexts, J), since thenU.2/DU.0/ and the range .2;4/simply duplicates .0;2/. But the homomor- phism remains important for even-dimension representations of SU.2/ , which correspond to half-integer Jand are not representations of SO.3/ . However, the fact that all rep- resentations of SO.3/ are also representations of SU.2/ means that we can form within SU.2/ direct products that include representations of both even and odd dimension. This observation validates our analysis of states with both orbital and spin angular momentum. In summary, we observe that half-integer angular momentum basis functions, which in earlier discussion we have already labeled as spinors, not only are objects that cannot be represented as functions in ordinary 3-D space, but are also objects whose rotational properties are unusual in that their angular periodicity is 4, not the value 2that would ordinarily be expected. They are thus somewhat abstract quantities whose relevance to physics rests on their ability to explain the “spin” properties of electrons and other fermions. Group SU(3) Starting in the 1930s, physicists began to give considerable attention to the symmetries of baryons, particles that, as the prefix “bary” implies, are heavy in comparison to electrons, and that interact subject to a force called the strong interaction. The earliest conjecture, ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 853 by Heisenberg, was to the effect that the approximate charge independence of the nuclear forces involving protons and neutrons suggested that they could be viewed as different quantum states of the same particle (called the nucleon), with the nucleon having a sym- metry appropriate to the existence of a doublet of states. The nucleon was postulated to have the same symmetry as electron spin, namely that of the continuous group SU.2/ . Although the nucleon symmetry has nothing to do with spin, it is referred to as isospin, with the isospin symmetry described by the matrices i,iD1;2;3(equal to the corre- sponding Pauli spin matrices i), and the isospin states can be classified by the eigenvalue of3(designated I3), with I3DC1=2 corresponding to the proton, I3D1=2 correspond- ing to the neutron. By the early 1960s, a large number of additional baryons with strong interactions had been identified, of which eight (proton, neutron, and six others) were rather similar in mass. The masses of the baryons discussed in this section are listed in Table 17.5. In 1961, Gell-Mann, and independently Ne’eman, suggested that these eight baryons might be symmetry-related, and proposed that they be identified with an irreducible rep- resentation of the group SU.3/ , with the relatively small mass differences attributed to forces weaker than the strong interaction and with different symmetry. The states describ- ing these eight particles would be a basis for the generators of an SU3representation of dimension 8. Subsequently, it was proposed that all eight of these particles were actually formed from combinations of three smaller, and presumably more fundamental, particles called quarks, and the three types of quarks initially postulated, given the names up(u), down (d), and strange (s), were ultimately identified as forming a basis for the generators ofSU.3/ . This original insight then led to the identification of a set of mesons involved with strong interaction as species consisting of one quark and one antiquark, thereby also corresponding to basis members of representations of SU.3/ . The situation described in the preceding paragraph can be more fully understood by pro- ceeding to a somewhat detailed discussion of the group SU.3/ . This group is defined by its generators, of which there are eight. The maximum number that commute with each other is two, so the group is of order 321D8and rank 2. The simplest useful way to specify the Table 17.5 Baryon Octet Mass Y I 3 4V41321:3211 2 401314:91 +1 2 6V61197:43 0 1 601192:55 0 0 6C1189:37 0 C1 3V3 1115:63 0 0 NV n 939:566 1 1 2 p 938:272 1 C1 2 Masses are given as rest-mass energies, in MeV (1 MeV D 106eV). ArfKen_Ch17-9780123846549.tex 854 Chapter 17 Group Theory generators is to write them as 33matrices in the SU.3/ fundamental representation. Like other continuous groups, SU.3/ has an infinite number of other irreducible representations of various sizes, but the key properties of the generators (specifically, their commutation rules) will be the same as those of the fundamental representation. We accordingly write the eight SU.3/ generators in terms of zero-trace Hermitian matrices 1through 8, with SiD1 2i; (17.60) where the i, known as the Gell-Mann matrices, are 1D0 @0 1 0 1 0 0 0 0 01 A; 2D0 @0i0 i 0 0 0 0 01 A; 3D0 @1 0 0 01 0 0 0 01 A;4D0 @0 0 1 0 0 0 1 0 01 A; (17.61) 5D0 @0 0i 0 0 0 i0 01 A;6D0 @0 0 0 0 0 1 0 1 01 A; 7D0 @0 0 0 0 0i 0i 01 A;8D1p 30 @1 0 0 0 1 0 0 021 A: In our use of SU.3/ , we will associate the rows and columns of this representation (in order) to the quarks u,d, and s. Note that 1,2, and 3are block diagonal with the upper block being the SU.2/ isospin matrices, signaling the presence of an SU.2/ subgroup with generators 1=2,2=2, and 3=2. If we combine 3and8so as to choose the generators in different ways, we can replace 3with one of the following: 0 3Dp 383D0 @0 0 0 0 1 0 0 011 A; (17.62) 00 3Dp 38C3D0 @1 0 0 0 0 0 0 011 A; (17.63) indicating the existence of another SU.2/ subgroup with generators S0 1D6=2,S0 2D 7=2,S0 3D0 3=2, and a third SU.2/ subgroup, with generators S00 1D4=2,S00 2D5=2, S00 3D00 3=2. These observations support the notion that isospin multiplets can exist within anSU.3/ basis. Because SU.3/ is of rank 2, the members of its representations can be labeled according to the eigenvalues of two commuting generators, in contrast to the single label, SzorIz, that we employed to label SU.2/ members. It is customary to use for this purpose the two generators ( 3and8) already in diagonal form. Continuing with the notation introduced for the nucleon, the eigenvalue of the SU.3/ generator S3is identified as I3, while S8is used to construct the identifier Y(known as hypercharge), defined as the eigenvalue of 2S8=p 3. An oft-used alternative to Yis the strangeness SY1. ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 855 Example 17.7.3 QUANTUM NUMBERS OF QUARKS From S3D1 20 @1 0 0 01 0 0 0 01 A; we can read out the quark I3valuesC1 2foru,1 2ford, and 0 for s. From 2S8=p 3D8=p 3D1 30 @1 0 0 0 1 0 0 021 A; we find the Yvalues1 3foruandd, and2 3fors.  From the definitions of the Siin Eq. (17.60), one can readily carry out the matrix opera- tions needed to establish their commutation rules. Note that even though the commutation rules will be obtained by examining the specific representation introduced in Eq. (17.60), they apply to all representations of the SU.3/ generators. We will use the commutation rules in a ladder-operator approach to the analysis of the symmetry properties of the three-quark multiplets. It is helpful to systematize the work by temporarily renaming S1,S2asI1,I2;S6,S7asU1,U2; andS4,S5asV1,V2. Then we introduce ICDI1CiI2; IDI1iI2; UCDU1CiU2;UDU1iU2; (17.64) VCDV1CiV2;VDV1iV2; and write some relevant commutators as TS3;IUDI;TS3;UUD1 2U;TS3;VUD1 2V; (17.65) TS8;IUD0;TS8;UUD1 2p 3U;TS8;VUD1 2p 3V: Using the logic of ladder operators (described in detail for applications to angular momen- tum operators in Section 16.1), the above commutators can be used to show that, starting from a basis function .I3;Y/, we can apply I,U, orVto obtain basis functions with other label sets. For example, TS8;UCU .I3;Y/DS8UC .I3;Y/UCS8 .I3;Y/D1 2p 3UC .I3;Y/: Replacing S8 .I3;Y/by1 2p 3Y .I3;Y/, this equation can be rearranged to S8 UC .I3;Y/ D1 2p 3.YC1/ UC .I3;Y/ ; which shows that if it does not vanish, UC .I3;Y/is an eigenvector of S8with an eigenvalue corresponding to an increase of one unit in Y. Similarly, from the relation TS3;UCU .I3;Y/D1 2UC .I3;Y/, we find that UC .I3;Y/, if nonvanishing, is an ArfKen_Ch17-9780123846549.tex 856 Chapter 17 Group Theory eigenvector of S3with an eigenvalue less by 1/2 than that of .I3;Y/. These observa- tions correspond to the equation UC .I3;Y/DC .I31 2;YC1/. This and other ladder identities are summarized in the following equations: I .I3;Y/DCI .I31;Y/; U .I3;Y/DCU .I31 2;Y1/; (17.66) V .I3;Y/DCV .I31 2;Y1/: The constants Cwill depend on the representation under study and on the values of I3and Y; if the result of an operation according to any of these equations leads to an .I3;Y/set that is not part of the representation’s basis, the Cassociated with that equation will vanish and the ladder construction will terminate. It is important to stress that the operators in Eq. (17.66) only move within the represen- tation under study, so if we start with a basis member of an irreducible representation, all the functions we will be able to reach will also be members of the same representation. Example 17.7.4 QUARK LADDERS As a preliminary to our study of baryon and meson symmetries, let’s see how the ladder operators work, with the quarks, symbolically .I3;Y/, represented by uD 1 2;1 3 D0 @1 0 01 A;dD  1 2;1 3 D0 @0 1 01 A;sD  0;2 3 D0 @0 0 11 A: As explained in Example 17.7.3, the values of I3andYare obtained from the diagonal ele- ments (the eigenvalues) of S3andS8. The 33matrices representing the ladder operators in this example are ICD0 @0 1 0 0 0 0 0 0 01 A;UCD0 @0 0 1 0 0 0 0 0 01 A;VCD0 @0 0 0 0 0 1 0 0 01 A; (17.67) ID0 @0 0 0 1 0 0 0 0 01 A;UD0 @0 0 0 0 0 0 1 0 01 A;VD0 @0 0 0 0 0 0 0 1 01 A: By straightforward matrix multiplication, we find IuDd,ICdDu,UdDs,UCsDd, VuDs,VCsDu; all other operations yield vanishing results. These relationships can be represented in the 2-D graph shown as Fig. 17.7 with Yin the vertical direction and I3 horizontal. The arrows in the graph are labeled to indicate the results of application of the ladder operators.  Continuing now to the baryons, we consider representations appropriate to three quarks, which we can form as the direct product of three single-quark representations. Using the notation 3as shorthand for the fundamental representation (which is of dimension 3), the ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 857 + −+− + −I Y Iu v s (0, − )2 3d (− , )1 213u ( , )1 213 FIGURE 17.7 Conversions between u,d, and squarks by application of ladder operators I,U, and V. The coordinates of each particle are its .I;Y/. I− IYI+v+ u− v−u+ (−1, 0) (+1, 0)(− , +1)1 2 (− , −1)1 2(+ , −1)12(+ , +1)1 2 FIGURE 17.8 Root diagram of SU.3/ . Each operator is labeled by the changes it causes: (1I,1Y). direct product we need is 3 3 3. This direct product is a reducible representation, which decomposes into the direct sum 3 3 3D10881; (17.68) where 10,8, and 1refer to irreducible representations of the indicated dimensions. A standard way to decompose product representations such as we have here uses dia- grams known as Young tableaux. Because development of the rules for construction and use of Young tableaux would take us beyond the scope of this text, we pursue here an alter- nate route that uses the ladder operators of Eq. (17.66). Use of the ladder operators also has the advantage that it yields explicit expressions for the I3;Yeigenfunctions. Since the direction in which ladder operators connect states in an I3;Ydiagram is general, we can draw a picture that summarizes their properties. Such a picture is called a root diagram; that for SU.3/ is shown in Fig. 17.8. Example 17.7.5 GENERATORS FOR DIRECT PRODUCTS If we apply an operation Rdepending on a parameter 'to a product of basis functions for different particles, each function will transform according to its representation, which we ArfKen_Ch17-9780123846549.tex 858 Chapter 17 Group Theory presently assume to be the fundamental representation: R i.1/ j.2/ D U.R/ i.1/ U.R/ j.2/ D ei'S.1/ i.1/ ei'S.2/ j.2/ Dei'TS.1/CS.2/U i.1/ j.2/; where the notation is supposed to indicate that S.1/ acts only on particle 1 and S.2/ acts only on particle 2 (this can be arranged by an appropriate definition of the direct-product matrices and the operators to which they correspond). The important point here is that because generators appear in an exponent, a product of single-particle operations can be obtained using a sum of single-particle generators. This observation is a generalization of our earlier writing of resultant multiparticle angular momenta as sums of individual contributions, and enables us to write, for three-quark products, expressions such as IDI.1/CI.2/CI.3/I so, for example (dropping the proportionality constant CI), Iu.1/u.2/u.3/Dd.1/u.2/u.3/Cu.1/d.2/u.3/Cu.1/u.2/d.3/: Suppressing the explicit particle numbers, this can be shortened to IuuuDduuCuduC uud. Corresponding results apply to all the other ladder operators and to all three-quark products, and to the application of the diagonal generators, such as S3u.1/u.2/u.3/D S3.1/u.1/ u.2/u.3/Cu.1/ S3.2/u.2/ u.3/ Cu.1/u.2/ S3.3/u.3/ D3 2u.1/u.2/u.3/; orS3uuuD3 2uuu, equivalent to assigning I3D3 2touuu. Similar analysis can yield results such as I3D1 2foruud, or.2S 8=p 3/dssDdss , showing that dsshasYD1 . We are now ready to return to the verification of Eq. (17.68). Example 17.7.6 DECOMPOSITION OF BARYON MULTIPLETS There are 27 three-quark products, which, using the analysis of Example 17.7.5, have the .I3;Y/values shown here. C3 2;1 uuu C1 2;1 uud;udu;duu 1 2;1 udd;dud;ddu 3 2;1 ddd .C1; 0/ uus;usu;suu.0;0/ uds;dus;usd;dsu;sud;sdu .1; 0/ dds;dsd;sdd C1 2;1 uss;sus;ssu 1 2;1 dss;sds;ssd .0;2/ sss ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 859 We can find the irreducible representations in our direct product in a relatively mechani- cal way. We start by placing the 27 quark products at their coordinate positions in an I3;Y diagram. We note that the point .3 2;1/is occupied by only one product, uuu, so it must, by itself, be a member of some irreducible representation of SU3. Starting there, we may take steps in any of the directions indicated in the root diagram, providing there is a function at each point to which we move. Since all we are doing is identifying possible states, we need not make any sophisticated computations as we proceed. Since uuu is completely symmetric under permutations, the basis function at each point will be a symmetric sum of the products at each point reached. When we have reached all the points, we will have identified a total of 10 basis functions, all members of the same irreducible representation, the one we called 10. This set of 10 basis functions is called a decuplet. The graph for these basis functions, called a weight diagram, is shown in Fig. 17.9. At the points where there was more than one quark product, there will be products left over after accounting for 10; if we want to be quantitative, they will be linear combinations that are orthogonal to the symmetric forms used in 10. Continuing with either of the two leftover functions at .1 2;1/, we may construct another set of basis functions from the left- overs; these sets will contain eight members, with the weight diagram shown in Fig. 17.10. (There are only seven points still occupied in the diagram, but the one at (0,0) yields two different functions when approached from different directions; the function obtained when (0,0) is reached horizontally can, via a subgroup analysis, be related to the members of its representation at .1; 0/. These points are elaborated in Exercise 17.7.4.) After account- ing for these two octets, corresponding to representations 8, there will be one completely antisymmetric function left at (0,0); it is a basis for 1.  Both the representations 8and10are relevant for particle physics. The rationalization of the similar-mass baryon octet was based on assignment of those particles to members of8, with the small mass differences associated with the breaking of the strong-interaction symmetry by a weaker force which retained some of the SU(2) subgroup symmetries, and by the (weaker still) electromagnetic forces that also broke the SU(2) symmetries. The identification of the octet members with the basis functions of 8is included in Fig. 17.10, and the energetics of the overall situation is indicated schematically in Fig. 17.11. −101 −2 −1 01I3Ω−Ξ∗−Ξ∗0Σ∗+Σ∗−Δ−Δ+Δ++Δ0 Σ∗0Y FIGURE 17.9 Weight diagram, baryon decuplet. The symbols at the various points are the names of particles assigned to the basis. ArfKen_Ch17-9780123846549.tex 860 Chapter 17 Group Theory −1/2 1/2 −11 −101I3 Ξ0Ξ−Σ−Σ+ Σ0, ΛY np FIGURE 17.10 Weight diagram, baryon octet. Mass NΞ− ΛΣΞ Hstrong Hstrong +HmediumHstrong +Hmedium +HelectromagneticΛ n pΣ− Σ+Σ0Ξ0 FIGURE 17.11 Baryon mass splitting. The representation 10provides an explanation for the set of 10 excited-state baryons whose weight diagram is shown in Fig. 17.9. When Gell-Mann fitted the then existent data to the decuplet representation, the particle had not yet been discovered, and its prediction and subsequent detection provided a strong indication of the relevance of SU.3/ ArfKen_Ch17-9780123846549.tex 17.7 Continuous Groups 861 to physics. Yet another instance of the importance of SU.3/ is provided by the existence of a meson octet (displaced by one unit in Yrelative to the primary baryon octet). Finally, we caution the reader that the foregoing discussion is by no means complete. It does not take full account of fermion antisymmetry requirements, the consideration of which led to the SU.3/ -color gauge theory of the strong interaction called quantum chro- modynamics (QCD). QCD also, at a minimum, involves the group SU.3/ . We have also left much unsaid about subgroup decompositions of the overall symmetry group, qualita- tively alluded to in the discussion supporting Fig. 17.9. To keep group theory and its very real value in proper perspective, we should empha- size that group theory identifies and formalizes symmetries. It classifies (and sometimes predicts) particles. But apart from saying, e.g., that one part of the Hamiltonian has SU.2/ symmetry and another part has SU.3/ symmetry, group theory says nothing about the par- ticle interaction. Likewise, a spherically symmetric Hamiltonian has (in ordinary space) SO.3/ symmetry, but this fact tells us nothing about the radial dependence of either the potential or the wave function. Exercises 17.7.1 Determine three SU(2) subgroups of SU(3). 17.7.2 Prove that the matrices U.n/(unitary matrices of order n) form a group, and that SU.n/ (those with determinant unity) form a subgroup of U.n/. 17.7.3 Using Eq. (17.56) for the matrix elements of SU(2) corresponding to rotations about the coordinate axes, find the matrix corresponding to a rotation defined by Euler angles . ; ; / . The Euler angles are defined in Section 3.4. 17.7.4 For a product of three quarks, the member of SU(3) representation 10with.I3;Y/D .C3 2;1/isuuu. (a) Apply operators in the root diagram for SU.3/ , Fig. 17.8, to obtain all the remain- ing members of the decuplet comprising the representation 10. (b) The two representations 8can be chosen to have for I3D1 2,YD1the respective members 11 2;1 D.uddu/uand 21 2;1 D2uudududuu. Briefly explain why this choice is possible. (c) Using the operators in the root diagram and the above 11 2;1 , find expressions for 1 1 2;1 , 1.1; 0/, 1.1;0/, 1 1 2;1 , and 11 2;1 . (d) Taking each of the six 1functions you now have, apply an operator that will convert it into 1.0;0/. Show that you obtain exactly two linearly independent 1.0;0/, thereby justifying the claim that the 1are an octet at the points shown inFig. 17.10. (e) Show that the octet built starting from 2.1 2;1/is linearly independent from that built from 1. (f) Find the wave function .0; 0/that is linearly independent of all the .0; 0/func- tions found in parts (a)–(e). It is the sole member of the representation 1. ArfKen_Ch17-9780123846549.tex 862 Chapter 17 Group Theory 17.8 L ORENTZ GROUP It has long been accepted that the laws of physics should be covariant, meaning that they should have forms that are (1) independent of the origin of the coordinates used to describe them (leading from an isolated system to the law of conservation of linear momentum); (2) independent of the orientation of our coordinates (leading to a conservation law for angular momentum); and (3) independent of the zero from which time is measured. Most of our experience suggests that velocities should add like ordinary vectors; for example, a person walking toward the front of a moving train would, as viewed by a stationary observer, have a net velocity equal to the sum of that of the train and the walker’s veloc- ity relative to the train. This rule for velocity addition is identified as Galilean, and is correct in the limit of small velocities. However, it is now known that transformations between coordinate systems with a constant nonzero relative velocity must lead to a non- intuitive velocity addition law that causes the velocity of light to be the same as measured by observers in all coordinate systems (reference frames). As Einstein showed in 1905, the necessary velocity addition law could be obtained if coordinate-system changes were described by Lorentz transformations. Einstein’s theory, now known as special relativ- ity(its extension to curved space-time to describe gravitation is called general relativity), also helped to complete an understanding of the way in which electric and magnetic phe- nomena become interconverted when charges at rest in one coordinate system are viewed as moving in another. The transformations that are consistent with the symmetry of space-time form a group known as the inhomogeneous Lorentz group or the Poincaré group. The Poincaré group consists of space and time displacements and all Lorentz transformations; here we shall only discuss the Lorentz transformations, which by themselves form the Lorentz group, sometimes for clarity referred to as the homogeneous Lorentz group. Homogeneous Lorentz Group Lorentz transformations can be likened to rotations that affect both the spatial and the time coordinates. An ordinary spatial rotation about the origin, in which .x1;x2/!.x0 1;x0 2/, has the property that the length of the associated vector is unchanged by the rotation, so that x2 1Cx2 2Dx02 1Cx02 2. But we now consider transformations involving a spatial coordinate (let’s choose z) and a time coordinate t, but with z2c2t2Dz02c2t02, so that the velocity of light, c, computed for travel from the origin ( 0;0) to.z;t/will be the same as that for travel from the origin to .z0;t0/. We are therefore abandoning the notion that the time variable is universal, assuming instead that it changes together with changes in the spatial variable(s) in a way that keeps the velocity of light constant. We also see that it is natural to rescale the tcoordinate to x0Dct, so that the invariant of the transformation becomes z2x2 0. Let’s now examine a situation in which the coordinate system is moving in the Cz direction at an infinitesimal velocity c(so that a Galilean transformation applies to z): z0Dzc./tDz./x0: But we assume that talso changes, to t0Dta./z;orx0 0Dx0ac./z; ArfKen_Ch17-9780123846549.tex 17.8 Lorentz Group 863 with achosen to keep z2x2 0constant to first order in . The value of athat satisfies this requirement is aDC1=c , so our infinitesimal Lorentz transformation is x0 0 z0! D 1  1! x0 z! D" 12 0 1 1 0!# x0 z! : To identify this equation in terms of a generator, we note that  0 1 1 0! Di./S; orSDi 0 1 1 0! Di1; (17.69) where 1is a Pauli matrix. Extending now to a finite velocity just as we did for ordinary rotations in the passage from Eq. (17.35) toEq. (17.36), we have an expression that is similar to Eq. (17.39), except that we now have 1instead of 2, while in place of 'we now have i. The result is U./Dexp.iTi1U/Dcos.i/Ci1sin.i/Dcosh./1sinh./ D coshsinh sinhcosh! : (17.70) Whilewas an infinitesimal velocity (in units of c), it does not follow that , the result of repeatedtransformations, is proportional to the resultant velocity in the final trans- formed coordinates. However, from the equation z0Dzcoshx0sinh, we identify the resultant velocity as vDcsinh=coshDctanh. Summarizing, and introducing the symbols usually used in relativistic mechanics, we identify v c;tanhD ; coshD1p 1 2 ; sinhD : (17.71) The range of(sometimes called the rapidity) is unlimited, but tanh<1, thereby show- ing that cis an upper limit to v(which cannot be reached for finite ). A Lorentz transformation that does not also involve a spatial rotation is known as a boost or a pure Lorentz transformation. Successive boosts can be analyzed using the group property of the Lorentz transformations: A boost of rapidity followed by another, of rapidity0, both in the zdirection, must have transformation matrix U.0/U./D cosh0sinh0 sinh0cosh0! coshsinh sinhcosh! D cosh0coshCsinh0sinhcosh0sinhsinh0cosh sinh0coshcosh0sinh sinh0sinhCcosh0cosh! D cosh.C0/sinh.C0/ sinh.C0/cosh.C0/! DU.C0/; ArfKen_Ch17-9780123846549.tex 864 Chapter 17 Group Theory showing that the rapidity (not the velocity) is the additive parameter for successive boosts in the same direction. The result we have just obtained is obvious if we write it in the generator notation; it is U.0/U./Dexp.01/exp. 1/Dexp..0C/ 1/DU.0C/: (17.72) Because of the group property, successive boosts in different spatial directions must yield a resultant Lorentz transformation, but the result is not equivalent to any single boost, and corresponds to a boost plus a spatial rotation. This rotation is the origin of the Thomas precession that arises in the treatment of spin-orbit coupling terms in atomic and nuclear physics. A good discussion of the Thomas precession frequency is in the work by Goldstein (Additional Readings). Example 17.8.1 ADDITION OF COLLINEAR VELOCITIES Let’s now apply Eq. (17.72) to two successive boosts in the zdirection, identifying each by its individual velocity ( v0for the first boost, v00for the second), or equivalently 0Dv0=c, 00Dv00=c. The corresponding rapidities will be denoted 0and00, so tanh0D 0Dv0 c;tanh00D 00Dv00 c: The resultant of the two successive boosts will have rapidity D0C00, and will therefore be associated with a resultant velocity vsatisfying tanh.0C00/Dv=cD . From the summation formula for the hyperbolic tangent, we have v cD Dtanh.0C00/Dtanh0Ctanh00 1Ctanh0tanh00Dv0 cCv00 c 1Cv0v00 c2D 0C 00 1C 0 00: (17.73) Equation (17.73) shows that when v0andv00are both small compared to c, the velocity addition is approximately Galilean, becoming exactly Galilean in the small-velocity limit. But as the individual velocities increase, their resultant decreases relative to their arithmetic sum, and never exceeds c. This behavior is to be expected, since (for real arguments) the hyperbolic tangent cannot exceed unity.  Minkowski Space If we make the definition x4Dict, the formulas we have just obtained, and many others as well, can be written in a systematic form that does not have minus signs explicitly present for the time coordinate. Then Lorentz transformations act like rotations in a space with basis.x1;x2;x3;x4/, and the conserved quantity is x2 1Cx2 2Cx2 3Cx2 4. This approach is appealing and is widely used. An alternative way to proceed, which has the disadvantage of being a bit more cum- bersome, but with the advantage of providing a framework suitable for the extension to general relativity, is to use real coordinates (as was done in the preceding subsection), but to handle the difference in behavior of the spatial and time coordinates by introducing a suitably defined metric tensor. One possibility (for basis x0Dct,x1,x2,x3), where xi ArfKen_Ch17-9780123846549.tex 17.8 Lorentz Group 865 (iD1;2;3) are Cartesian spatial coordinates, is to use the Minkowski metric tensor, first introduced in Example 4.5.2, .g/D.g/D0 BB@1 0 0 0 01 0 0 0 01 0 0 0 011 CCA; (17.74) where it is understood that Greek indices run over the four-index set 0 to 3, and that displacements are rendered as scalar products of the form xgx0orxgx0 , where the repeated indices are understood to be summed (the Einstein summation convention). Note that because all the analysis in this section is in Cartesian coordinates, the distinction between contravariant and covariant indices is limited to the insertion of minus signs in some elements of products that involve the metric tensor. As was pointed out in Example 4.6.2, this metric tensor sometimes appears with the signs of all its diagonal elements reversed. Either choice of signs is valid and yields proper results for problems of physics if used consistently, but trouble can arise if material from inconsistent sources is combined. The cited example also indicates how Maxwell’s equa- tions can be written in a manifestly covariant form. Note that the transformation matrices SandUmust be mixed tensors, since they convert a vector (whether covariant or contravariant) into another vector of the same variance status. Since for a pure boost these matrices are symmetric, either index can be deemed to be covariant (the other then being contravariant). Exercises 17.8.1 Show that in 3C1dimensions (this means three spatial dimension plus time), a boost in the xyplane at an angle from the xdirection has, in coordinates .x0;x1;x2;x3/, the generator SDi0 BB@0 cossin0 cos 0 0 0 sin 0 0 0 0 0 0 01 CCA: 17.8.2 (a) Show that the generator in Exercise 17.8.1 produces a Lorentz transformation matrix for rapidity given by U.I/D0 BBBB@coshcossinhsinsinh 0 cossinhsin2Ccos2coshcossin.cosh1/0 sinsinhcossin.cosh1/ cos2Csin2cosh0 0 0 0 11 CCCCA: Note. This transformation matrix is symmetric. All single boosts (in any spatial direction) have symmetric transformation matrices. (b) Verify that the transformation matrix of part (a) is consistent with (1) rotating the spatial coordinates to align the boost direction with a coordinate axis, (2) performing a boost in the direction of that axis using Eq. (17.70), and (3) rotat- ing back to the original coordinate system. ArfKen_Ch17-9780123846549.tex 866 Chapter 17 Group Theory 17.8.3 Obtain the Lorentz transformation matrix for a boost of finite amount 0in the xdirec- tion followed by a finite boost 00in the ydirection. Show that there are no values of  andthat can bring this transformation to the form given in Exercise 17.8.2. 17.9 L ORENTZ COVARIANCE OF MAXWELL ’SEQUATIONS We start our discussion of Lorentz covariance by recalling how the magnetic and electric fields BandEdepend on the vector and scalar potentials Aand': BDrA; (17.75) ED@A @tr': Restricting consideration to situations where "andhave their free-space values "0and 0(with"00D1=c2), it can be shown that Aand'form a four-vector whose components A(in contravariant form) are AiDc"0Ai;iD1;2;3; (17.76) A0D"0': We now form the tensor Fwith elements FD@A @x@A @x; (17.77) which we evaluate (consistently with our choice of Minkowski metric) using @ @x0D@ c@t;@ @x1D@ @x;@ @x2D@ @y;@ @x3D@ @z: (17.78) The resulting form for F, known as the electromagnetic field tensor, is FD"00 BBBB@0ExEyEz Ex 0cB zcBy EycBz 0cB x EzcB ycBx 01 CCCCA: (17.79) The quantity Fis, as its name implies, a second-order tensor that must have the transformation properties associated with the Lorentz group. We know this to be the case because we constructed Fas a linear combination of terms, each of which was the derivative of a four-vector; differentiation of a vector (in a Cartesian system) generates a second-order tensor. An interesting aside to the above analysis is provided by the discussion of Maxwell’s equations in the language of differential forms. In Example 4.6.2 we showed that the dif- ferential form FDExdt^dxEydt^dyEzdt^dzCBxdy^dzCBydz^dxCBzdx^dy ArfKen_Ch17-9780123846549.tex 17.9 Lorentz Covariance of Maxwell’s Equations 867 was a starting point from which Maxwell’s equations could be derived; we now observe that the individual terms of this differential form correspond to the elements of the tensor under discussion here. Lorentz Transformation of E and B Returning to the main matter of present concern, we now apply a Lorentz transformation toF. For simplicity we take a pure boost in the zdirection, which will have matrix elements similar to those of Eq. (17.70); using the notations introduced in Eq. (17.71), our transformation matrix can be written UD0 BBBB@ 0 0 0 1 0 0 0 0 1 0 0 0 1 CCCCA: (17.80) Noting that we must apply our Lorentz transformation to both indices of F, and keeping in mind that Uis symmetric and, as pointed out in Section 17.8, a mixed tensor, we can write F0DUFU; (17.81) where FandF0are both contravariant matrices. If we now compare the individual elements ofF0with those of F, we obtain formulas for the components of E0andB0in terms of the components of EandB. For the transformation at issue here, the results are (where vis the velocity of the transformed coordinate system, in the zdirection, relative to the original coordinates): E0 xD Ex cBy D ExvBy ; E0 yD EyC cBx D EyCvBx ; (17.82) E0 zDEz; B0 xD  BxC cEy D  BxCv c2Ey ; B0 yD  By cEx D  Byv c2Ex ; (17.83) B0 zDBz: We can generalize the above to a boost vin an arbitrary direction: E0D .ECvB/C.1 /Ev; B0D  BvE c2 C.1 /Bv;(17.84) ArfKen_Ch17-9780123846549.tex 868 Chapter 17 Group Theory where EvD.EOv/OvandBvD.BOv/Ovare the projections of EandBin the direction of v. In the limitvc, these equations reduce to E0DECvB; B0DBvE c2:(17.85) Note that the coordinate transformation changes the velocity with which charges move and therefore changes the magnetic force. It is now clear that the Lorentz transformation explains how the total force (electric plus magnetic) can be independent of the reference frame (i.e., the relative velocities of the coordinate systems). In fact, the need to make the total electromagnetic force independent of the reference frame was first noted by Lorentz and Poincaré. This was where Lorentz transformations were first recognized as relevant for physics, and that may have provided Einstein with a clue as he developed his formulation of special relativity. Example 17.9.1 TRANSFORMATION TO BRING CHARGE TO REST Consider a charge qmoving at a velocity v, withvc. By giving the coordinate system a boost v, we transform to a frame in which the charge is at rest and experiences only an electric force qE0. But since the total force is independent of the reference frame, it is also given, according to Eq. (17.86), as FDq.ECvB/; (17.86) which is just the classical Lorentz force.  The ability to write Maxwell’s equations in a tensor form that gives the experimentally observed results under Lorentz transformation is an important achievement because it guar- antees that the formulation is consistent with special relativity. This is one of the reasons that modern theories of quantum electrodynamics and elementary particles are often writ- ten in this manifestly covariant form. Conversely, the insistence on such a tensor form has been a useful guide in the construction of these theories. We close with the following general observations: The Lorentz group is the symmetry group of electrodynamics, of the electroweak gauge theory, and of the strong interactions described by quantum chromo- dynamics. It appears necessary that mechanics in general have the symmetry of the Lorentz group, and that requirement corresponds to the general applicability of special relativity. With respect to electrodynamics, the Lorentz symmetry explains the fact that the velocity of light is the same in all inertial frames, and it explains how electric and magnetic forces are interrelated and yield physical results that are frame-independent. While a detailed study of relativistic mechanics is beyond the scope of this book, the extension to special relativity of Newton’s equations of motion is straightforward and leads to a variety of results, some of which challenge human intuition. ArfKen_Ch17-9780123846549.tex 17.10 Space Groups 869 Exercises 17.9.1 Apply the Lorentz transformation of Eq. (17.80) to Fas given in Eq. (17.79). Verify that the result is a matrix F0whose elements confirm the results given in Eqs. (17.82) and (17.83). 17.9.2 Confirm that the generalization of Eqs. (17.82) and(17.83) to a boost corresponding to an arbitrary velocity vis properly given by Eq. (17.84). 17.10 S PACE GROUPS Perfect crystals exhibit translational symmetry, meaning that they can be considered as a space-filling array of parallelepipeds stacked end-to-end and side-to-side, with each con- taining an identical set of identically placed atoms. A single parallelepiped is referred to as theunit cell of the crystal; a unit cell can be specified by giving the vectors that define its edges. Calling these vectors h1,h2,h3, equivalent points in any two unit cells are separated from each other by vectors bDn1h1Cn2h2Cn3h3; where n1,n2,n3can be any integers (positive, negative, or zero). The set of these equiva- lent points is called the Bravais lattice of the crystal. A Bravais lattice will have a symmetry that depends on the angles and relative lengths of the lattice vectors; in three dimensions there are 14 different symmetries possible for Bravais lattices. There are 32 3-D point groups that are symmetry-compatible with at least one Bravais lattice; these are called crystallographic point groups to distinguish them from the infinite number of point groups that can exist in the absence of any compatibility requirement. Example 17.10.1 TILING A FLOOR To understand the notion of crystallographic point group, consider what would happen (in two dimensions) if we try to tile a floor with identical tiles in the shape of a regular polygon. We will have success with squares and triangles, and even with hexagons. These work because an integer number of tiles can be placed so that they have vertices at the same point. A triangle has an internal angle of 60, so six of them can meet at a point; similarly, four squares can meet at a point, as can three hexagons (internal angle 120). But we cannot tile with regular pentagons (internal angle 108) or any regular polygon with more than six sides.  Combining Bravais lattices and compatible point groups, there is a total of 230 different groups in 3-D that exhibit translational symmetry and some sort of point-group symmetry. These 230 groups are called space groups. Their study and use in crystallography (e.g., to determine the detailed structure of a crystal from its x-ray scattering) is the topic of several of the larger books in the Additional Readings. Systems with periodicity in only one or two dimensions also exist in nature; some lin- ear polymers are 1-D periodic systems; surface systems and single-layer arrays such as graphene (a macroscopic hexagonal array of carbon atoms) exhibit periodicity in two ArfKen_Ch17-9780123846549.tex 870 Chapter 17 Group Theory dimensions. There is even a kind of translational symmetry that involves elements that form helical structures. The recognition of this type of symmetry in crystallographic stud- ies of DNA was the key contribution leading to the discovery that DNA existed as a dou- ble helix. Additional Readings Buerger, M. J., Elementary Crystallography . New York: Wiley (1956). A comprehensive discussion of crystal symmetries. Buerger develops all 32 point groups and all 230 space groups. Related books by this author include Contemporary Crystallography . New York: McGraw-Hill (1970); Crystal Structure Analysis. New York: Krieger (1979) (reprint, 1960); and Introduction to Crystal Geometry. New York: Krieger (1977) (reprint, 1971). Burns, G., and A. M. Glazer, Space Groups for Solid-State Scientists. New York: Academic Press (1978). A well-organized, readable treatment of groups and their application to the solid state. de-Shalit, A., and I. Talmi, Nuclear Shell Model. New York: Academic Press (1963). We adopt the Condon- Shortley phase conventions of this text. Falicov, L. M., Group Theory and Its Physical Applications. Notes compiled by A. Luehrmann. Chicago: Uni- versity of Chicago Press (1966). Group theory, with an emphasis on applications to crystal symmetries and solid-state physics. Gell-Mann, M., and Y. Ne’eman, The Eightfold Way. New York: Benjamin (1965). A collection of reprints of significant papers on SU(3) and the particles of high-energy physics. Several introductory sections by Gell- Mann and Ne’eman are especially helpful. Goldstein, H., Classical Mechanics, 2nd ed. Reading, MA: Addison-Wesley (1980). Chapter 7 contains a short but readable introduction to relativity from a viewpoint consonant with that presented here. Greiner, W., and B. Müller, Quantum Mechanics Symmetries. Berlin: Springer (1989). We refer to this textbook for more details and numerous exercises that are worked out in detail. Hamermesh, M., Group Theory and Its Application to Physical Problems. Reading, MA: Addison-Wesley (1962). A detailed, rigorous account of both finite and continuous groups. The 32 point groups are developed. The continuous groups are treated, with Lie algebra included. A wealth of applications to atomic and nuclear physics. Hassani, S., Foundations of Mathematical Physics. Boston: Allyn and Bacon (1991). Heitler, W., The Quantum Theory of Radiation, 2nd ed. Oxford: Oxford University Press (1947), reprinting, Dover (1983). Higman, B., Applied Group-Theoretic and Matrix Methods. Oxford: Clarendon Press (1955). A rather complete and unusually intelligible development of matrix analysis and group theory. Jackson, J. D., Classical Electrodynamics, 3rd ed. New York: Wiley (1998). Messiah, A., Quantum Mechanics, vol. II. Amsterdam: North-Holland (1961). Panofsky, W. K. H., and M. Phillips, Classical Electricity and Magnetism, 2nd ed. Reading, MA: Addison- Wesley (1962). The Lorentz covariance of Maxwell’s equations is developed for both vacuum and material media. Panofsky and Phillips use contravariant and covariant tensors. Park, D., Resource letter SP-1 on symmetry in physics. Am. J. Phys. 36: 577–584 (1968). Includes a large selec- tion of basic references on group theory and its applications to physics: atoms, molecules, nuclei, solids, and elementary particles. Ram, B., Physics of the SU(3) symmetry model. Am. J. Phys. 35: 16 (1967). An excellent discussion of the applications of SU(3) to the strongly interacting particles (baryons). For a sequel to this see R. D. Young, Physics of the quark model. Am. J. Phys. 41: 472 (1973). Tinkham, M., Group Theory and Quantum Mechanics. New York: McGraw-Hill (1964), reprinting, Dover (2003). Clear and readable. Wigner, E. P., Group Theory and Its Application to the Quantum Mechanics of Atomic Spectra (translated by J. J. Griffin). New York: Academic Press (1959). This is the classic reference on group theory for the physicist. The rotation group is treated in considerable detail. There is a wealth of applications to atomic physics. ArfKen_Ch18-9780123846549.tex CHAPTER 18 MORE SPECIAL FUNCTIONS In this chapter we shall study four sets of orthogonal polynomials: Hermite, Laguerre, and Chebyshev1of the first and second kinds. Although these four sets are of less importance in mathematical physics than are the Bessel and Legendre functions of Chapters 14 and 15, they are used and therefore deserve attention. For example, Hermite polynomials occur in solutions of the simple harmonic oscillator of quantum mechanics and Laguerre polynomi- als in wave functions of the hydrogen atom. Because the general mathematical techniques duplicate those used for Bessel and Legendre functions, the development of these functions is only outlined. Detailed proofs are for the most part left to the reader. The sets of polynomials treated in this chapter can be related to the more general quan- tities known as hypergeometric andconfluent hypergeometric functions (solutions of the hypergeometric ODE). For practical reasons we defer most discussion of these rela- tionships until we have had an opportunity to define the hypergeometric functions and the associated nomenclature. The benefit accruing from the connection to hypergeomet- ric functions is that the hypergeometic recurrence formulas and other general properties translate into useful relationships for the polynomial sets that we are presently studying. We conclude the chapter with a short section on elliptic integrals. Although the impor- tance of this subject has declined as the power of computers has increased, there are some physical problems for which they are useful and it is not yet time to eliminate them from this text. 18.1 H ERMITE FUNCTIONS We start by identifying Hermite functions as solutions of the Hermite ODE, H00 n.x/2x H0 n.x/C2nH n.x/D0: (18.1) 1This is the spelling choice of AMS-55 (for the complete reference, see Abramowitz in Additional Readings). However, various names, such as Tschebyscheff, are encountered in the literature. 871 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch18-9780123846549.tex 872 Chapter 18 More Special Functions Here nis a parameter. When n0is integral, this ODE will have a solution Hn.x/which is a polynomial of degree n; these solutions are known as Hermite polynomials. In the presence of appropriate boundary conditions, the Hermite ODE is a Sturm- Liouville system; polynomial solutions to such ODEs was the topic of Section 12.1. We showed there, in Example 12.1.1, that the Hermite polynomials could be generated from their Rodrigues formula, Eq. (12.17), and that, in turn, a Rodrigues formula can be obtained from the underlying ODE. We also showed in that same section how we can go from the Rodrigues formula to a generating function for a given polynomial set, pre- senting in Table 12.1 a list of generating functions that could be found in this way. That list included the following generating function for the Hermite polynomials: g.x;t/Det2C2txD1X nD0Hn.x/tn nW: (18.2) Here we elect not to depend on the analysis of Section 12.1 but rather to take the view- point that Eq. (18.2) can be regarded as a definition of the Hermite polynomials, thereby making the present analysis completely self-contained. Accordingly, we proceed by verify- ing that these polynomials satisfy the Hermite ODE, have the expected Rodrigues formula, and exhibit the other properties that can be developed starting from the generating function. Recurrence Relations Note the absence of a superscript, which distinguishes Hermite polynomials from the unre- lated Hankel functions. From the generating function we find that the Hermite polynomials satisfy the recurrence relations HnC1.x/D2x Hn.x/2nH n1.x/ (18.3) and H0 n.x/D2nH n1.x/: (18.4) The Hermite polynomials were used in Example 12.1.2 as a detailed illustration of the method for obtaining recurrence formulas from generating functions; we summarize the process here. By differentiating the generating function formula with respect to twe obtain @g @tD.2tC2x/et2C2txD1X nD0HnC1.x/tn nW;or 21X nD0Hn.x/tnC1 nWC2x1X nD0Hn.x/tn nWD1X nD0HnC1.x/tn nW: Because this equation must be satisfied separately for each power of t, we arrive at Eq. (18.3). Similarly, differentiation with respect to xleads to @g @xD2tet2C2txD1X nD0H0 n.x/tn nWD21X nD0Hn.x/tnC1 nW; from which we can obtain Eq. (18.4). ArfKen_Ch18-9780123846549.tex 18.1 Hermite Functions 873 The Maclaurin expansion of the generating function et2C2txD1X nD0.2txt2/n nWD1C.2txt2/C (18.5) gives H0.x/D1andH1.x/D2x, and then the recursion formula, Eq. (18.3), permits the construction of any Hn.x/desired. For convenient reference the first several Hermite polynomials are listed in Table 18.1 and presented graphically in Fig. 18.1. Special Values Special values of the Hermite polynomials follow from the generating function for xD0: et2D1X nD0.t2/n nWD1X nD0Hn.0/tn nW; Table 18.1 Hermite Polynomials H0.x/D1 H1.x/D2x H2.x/D4x22 H3.x/D8x312x H4.x/D16x448x2C12 H5.x/D32x5160x3C120x H6.x/D64x6480x4C720x2120 x10 86 42 1 20 −2H 2(x) H1(x) H0(x) FIGURE 18.1 Hermite polynomials. ArfKen_Ch18-9780123846549.tex 874 Chapter 18 More Special Functions that is, H2n.0/D.1/n.2n/W nW;H2nC1.0/D0; nD0;1;: (18.6) We also obtain from the generating function the important parity relation Hn.x/D.1/nHn.x/ (18.7) by noting that Eq. (18.3) yields g.x;t/D1X nD0Hn.x/.t/n nWDg.x;t/D1X nD0Hn.x/tn nW: Hermite ODE If we substitute the recursion formula Eq. (18.4) intoEq. (18.3), we can eliminate the index n1, obtaining HnC1.x/D2x Hn.x/H0 n.x/: If we differentiate this recurrence relation and substitute Eq. (18.4) for the index nC1, we find H0 nC1.x/D2.nC1/Hn.x/D2Hn.x/C2x H0 n.x/H00 n.x/; which can be rearranged to the second-order Hermite ODE, Eq. (18.1). This completes the process of establishing the identification of the Hermite polynomials obtained from the generating function as solutions of the Hermite ODE. Rodrigues Formula A simple way to generate the Rodrigues formula for the Hermite polynomials starts from the observations that g.x;t/Det2C2txDex2e.tx/2and@ @te.tx/2D@ @xe.tx/2: We note that n-fold differentiation of the generating function formula, Eq. (18.2), followed by setting tD0, yields @n @tng.x;t/ tD0DHn.x/; and we can therefore obtain the Rodrigues formula as Hn.x/D@n @tng.x;t/ tD0Dex2@n @tne.tx/2 tD0D.1/nex2@n @xne.tx/2 tD0 D.1/nex2@n @xnex2: (18.8) ArfKen_Ch18-9780123846549.tex 18.1 Hermite Functions 875 Series Expansion Starting from the Maclaurin expansion, Eq. (18.5), we can derive our Hermite polynomial Hn.x/in series form: Using the binomial expansion of .2xt/, we initially get et2C2txD1X D0t W.2xt/D1X D0t WX sD0 s .2x/s.t/s D1X D0X sD0tCs .Cs/W.1/s.Cs/W.2x/s .s/WsW: Changing the first summation index from tonDCs, and noting that this change causes thessummation to range from zero to Tn=2U, the largest integer less than or equal to n=2, our expansion takes the form et2C2txD1X nD0tn nWTn=2UX sD0.1/snW .n2s/WsW.2x/n2s; from which we can read out the formula for Hn: Hn.x/DTn=2UX sD0.1/snW .n2s/WsW.2x/n2s: (18.9) Finally, we note that Hn.x/can be written as a Schlaefli integral. Comparing with Eq. (12.18), Hn.x/DnW 2iI tn1et2C2txdt: (18.10) Orthogonality and Normalization The orthogonality of the Hermite polynomials is demonstrated by identifying them as aris- ing in a Sturm-Liouville system. The Hermite ODE, however, is clearly notself-adjoint, but can be made so by multiplying it by exp. x2/(see Exercise 8.2.2). With exp. x2/as a weighting factor, we obtain the orthogonality integral 1Z 1Hm.x/Hn.x/ex2dxD0; m6Dn: (18.11) The interval.1;1/is chosen to obtain the Hermitian operator boundary conditions (see Section 8.2). It is sometimes convenient to absorb the weighting function into the Hermite polynomi- als. We may define 'n.x/Dex2=2Hn.x/; (18.12) ArfKen_Ch18-9780123846549.tex 876 Chapter 18 More Special Functions with'n.x/no longer a polynomial. Substitution into Eq. (18.1) yields the differential equa- tion for'n.x/, '00 n.x/C.2nC1x2/'n.x/D0: (18.13) Equation (18.13) is self-adjoint, and the solutions 'n.x/are orthogonal on the interval 1<x<1with a unit weighting function. We still need to normalize these functions. One approach is to combine two instances of the generating function formula (using variables sandt), after which we multiply by ex2 and integrate over xfrom1 to1. These steps yield 1Z 1ex2es2C2sxet2C2txdxD1X m;nD0smtn mWnW1Z 1ex2Hm.x/Hn.x/dx: (18.14) We next note that the exponentials on the left-hand side of Eq. (18.14) can be combined intoe2ste.xst/2, after which the integral can be evaluated: 1Z 1ex2es2C2sxet2C2txdxDe2st1Z 1e.xst/2dxD1=2e2st: Inserting this result into Eq. (18.14) after expanding it in a power series, we get 1=2e2stD1=21X nD02nsntn nWD1X m;nD0smtn mWnW1Z 1ex2Hm.x/Hn.x/dx: By equating coefficients of equal powers of sandt, we both confirm the orthogonality and obtain the normalization integral 1Z 1ex2h Hn.x/i2 dxD2n1=2nW: (18.15) Exercises 18.1.1 Assume that the Hermite polynomials are known to be solutions of the Hermite ODE, Eq. (18.1). Assume further that the recurrence relation, Eq. (18.3), and the values of Hn.0/are also known. Given the existence of a generating function g.x;t/D1X nD0Hn.x/tn nW; (a) Differentiate g.x;t/with respect to xand using the recurrence relation develop a first-order PDE for g.x;t/. ArfKen_Ch18-9780123846549.tex 18.1 Hermite Functions 877 (b) Integrate with respect to x, holding tfixed. (c) Evaluate g.0;t/using the known values of Hn.0/. (d) Finally, show that g.x;t/Dexp.t2C2tx/. 18.1.2 In developing the properties of the Hermite polynomials, start at a number of different points, such as: 1. Hermite’s ODE, Eq. (18.1), 2. Rodrigues’s formula, Eq. (18.8), 3. Integral representation, Eq. (18.10), 4. Generating function, Eq. (18.2), 5. Gram-Schmidt construction of a complete set of orthogonal polynomials over .1;1/with a weighting factor of exp. x2/(Section 5.2). Outline how you can go from any one of these starting points to all the other points. 18.1.3 Prove thatjHn.x/jj Hn.ix/j. 18.1.4 Rewrite the series form of Hn.x/,Eq. (18.9), as an ascending power series. ANS. H2n.x/D.1/nnX sD0.1/2s.2x/2s.2n/W .2s/W.ns/W, H2nC1.x/D.1/nnX sD0.1/s.2x/2sC1.2nC1/W .2sC1/W.ns/W: 18.1.5 (a) Expand x2rin a series of even-order Hermite polynomials. (b) Expand x2rC1in a series of odd-order Hermite polynomials. ANS. (a)x2rD.2r/W 22rrX nD0H2n.x/ .2n/W.rn/W (b)x2rC1D.2rC1/W 22rC1rX nD0H2nC1.x/ .2nC1/W.rn/W,rD0;1;2;::: . Hint. Use a Rodrigues representation and integrate by parts. 18.1.6 Show that (a)1Z 1Hn.x/exp x2 2 dxD(2nW=.n=2/W;neven 0;nodd: (b)1Z 1x Hn.x/exp x2 2 dxD8 >< >:0;neven 2.nC1/W ..nC1/=2/W;nodd: ArfKen_Ch18-9780123846549.tex 878 Chapter 18 More Special Functions 18.1.7 (a) Using the Cauchy integral formula, develop an integral representation of Hn.x/ based on Eq. (18.2) with the contour enclosing the point zDx. ANS. Hn.x/DnW 2iex2Iez2 .zCx/nC1dz. (b) Show by direct substitution that this result satisfies the Hermite equation. 18.2 A PPLICATIONS OF HERMITE FUNCTIONS One of the most important applications of Hermite functions in physics arises from the fact that the functions 'n.x/ofEq. (18.12) are the eigenstates of the quantum-mechanical simple harmonic oscillator, which describes motion subject to a quadratic (also known as aharmonic or aHooke’s-law) potential. This fact causes Hermite polynomials not only to appear in elementary quantum-mechanics problems, but also in analyses of the vibra- tional states of molecules, where the lowest-order description of the interatomic potential is harmonic. In view of the importance of these topics, we now proceed to examine them in some detail. Simple Harmonic Oscillator The quantum mechanical simple harmonic oscillator is governed by a Schrödinger equa- tion of the form Nh2 2md2 .z/ dz2Ck 2z2 .z/DE .z/; (18.16) where mis the mass of the oscillator, kis the force constant for its Hooke’s law force directed toward zD0,Nhis Planck’s constant divided by 2, and Eis an eigenvalue giv- ing the energy of the oscillator. Equation (18.16) is to be solved subject to the boundary condition that .z/vanish at zD1 . It is convenient to make a change of variable that eliminates the various constants from the equation, and we therefore make the substitutions zDNh1=2x .km/1=4;k 2z2DNh 2r k mx2;Nh2 2md2 dz2DNh 2r k md2 dx2; which converts Eq. (18.16) into 1 2d2'.x/ dx2Cx2 2'.x/D'.x/; (18.17) with boundary conditions at xD1 . The eigenvalue in this equation is related to Eby EDNhr k m: (18.18) The solutions of Eq. (18.17) that satisfy the boundary conditions can now be identified as given by Eq. (18.13), and we can identify n, the eigenvalue of Eq. (18.17) corresponding to'n.x/, as having the value nC1 2. Turning to Eq. (18.12), and expressing xin terms ArfKen_Ch18-9780123846549.tex 18.2 Applications of Hermite Functions 879 of the original variable z, the eigenstates of Eq. (18.16) can be characterized (including a normalization constant Nn) as n.z/DNne. z/2=2Hn. z/;EnD.nC1 2/Nhr k m; DNh1=2 .km/1=4; (18.19) with nrestricted to the integer values 0;1;2;. The normalization constant can be deduced from Eq. (18.15). Noting that the normalization integral is to be over the vari- ablez, we find it to be NnD 2n1=2nW1=2 : (18.20) It is of interest to examine a few of the eigenstates of this oscillator problem. For refer- ence, a classical oscillator of mass mand force constant kwill have the angular oscillation frequency !classDr k m; and can have an arbitrary energy of oscillation, while our quantum oscillator is restricted to oscillation energies .nC1 2/Nh!class, with na nonnegative integer. We note that the quantum oscillator must have at least the total energy1 2Nh!class; this is usually referred to as its zero-point energy and is a consequence of the fact that its spatial distribution must be described by a wave function of finite extent. The three lowest-energy eigenfunctions of the quantum oscillator are shown in Fig. 18.2. We note that these wave functions predict a position distribution that extends to 1, albeit with exponentially decaying amplitude for larger jzj. The corresponding classical oscillator will have excursions in zthat are strictly bounded by kz2 max=2DE, where E can be assigned any value greater than or equal to zero. We have marked in Fig. 18.2 the excursion range of a classical oscillator with an energy equal to the eigenvalue of the quantum oscillator; note that the exponential decay of the quantum wave function begins at the ends of the classical range. Operator Approach While the analysis of the preceding subsection is straightforward and provides a complete set of eigenstates for the simple quantum oscillator, additional insight can be obtained by an alternative approach that uses the commutation and other algebraic properties of the quantum-mechanical operators. Our starting point for this development is the recogni- tion that the differential operator d2=dx2ofEq. (18.17) arose as a representation of the dynamical quantity p2, where (in units with NhD1)p ! i d=dx. Then our Schrödinger equation of Eq. (18.17) can be written H'Dp2Cx2 2'D'; (18.21) whereHis the Hamiltonian operator, with eigenvalues . ArfKen_Ch18-9780123846549.tex 880 Chapter 18 More Special Functions 0.5 5xψ2(x) 0.5 5 5ψ1(x) ψ0(x) xx FIGURE 18.2 Quantum mechanical oscillator wave functions. The heavy bar on the x-axis indicates the allowed range of the classical oscillator with the same total energy. The key to an approach starting from Eq. (18.21) is that xandpsatisfy the basic com- mutation relation Tx;pUDxppxDi; (18.22) a result discussed in detail in the analysis leading to Eq. (5.43). In fact, if we proceed under the assumption that Eq. (18.22) is all that we know about xandp, there is the additional advantage that any results we obtain will be more general than those from our original oscillator problem in ordinary space. This observation underlies much recent work in which physical theory has evolved in more abstract directions. With a knowledge of the way in which angular momentum theory was developed in terms of raising and lowering operators, one can easily motivate a somewhat similar ArfKen_Ch18-9780123846549.tex 18.2 Applications of Hermite Functions 881 procedure here, by defining the two operators aD1p 2.xCip/;a†D1p 2.xip/: (18.23) Since we typically use ato denote a constant, we remind the reader that in the present development it is an operator (involving xandd=dx). With suitable Sturm-Liouville boundary conditions, xandpare both Hermitian. But the presence of the imaginary unit icauses anot to be Hermitian, and (as indicated by the notation) changing the sign of the term ipconverts ainto its adjoint, a†. Our first use of Eq. (18.23) is to form a†aandaa†: a†aD1 2.xip/.xCip/D1 2.x2Cp2/Ci 2.xppx/DHCi 2Tx;pUDH1 2; aa†D1 2.xCip/.xip/D1 2.x2Cp2/i 2.xppx/DHi 2Tx;pUDHC1 2: From these equations we obtain the useful formulas HDa†aC1 2; (18.24) Ta;a†UDaa†a†aD1; (18.25) and therefrom TH;aUDTa†aC1 2;aUDTa†a;aUDa†aaaa†aD.a†aaa†/aDa: (18.26) ApplyingTH;aUto an eigenfunction 'nwith eigenvalue n(assumed not yet known), we write TH;aU'nDH.a'n/aH'nDH.a'n/n.a'n/D. a'n/; which we easily rearrange to the form H.a'n/D.n1/.a'n/: (18.27) Equation (18.27) shows that we can interpret aas a lowering operator that converts an eigenfunction with eigenvalue ninto another eigenfunction that has eigenvalue n1. A similar analysis, left to the reader, shows that from the commutator TH;a†Uwe find a† to be a raising operator, according to H.a†'n/D.nC1/.a†'n/: (18.28) These formulas show that, given any eigenfunction 'n, we can construct a ladder of eigen- states whose eigenvalues differ by unit steps. The only limitation that would terminate the construction of an infinite ladder would be the possibility that for some 'n, either a'nor a†'nmight be zero. To investigate the circumstances under which a'nmight vanish, let’s form the scalar productha'nja'ni. We find ha'nja'niDh' nja†aj'niDh' njH1 2j'niDh' njn1 2j'ni: (18.29) ArfKen_Ch18-9780123846549.tex 882 Chapter 18 More Special Functions Equation (18.29) shows that only if nD1 2will we have a'nD0. That equation also shows that ifn<1 2we have the mathematical inconsistency that the norm of a'nis predicted to be negative. These observations together imply that the only possible values of nare positive half-integers, as otherwise by repeated application of the lowering operator awe can move to a value prohibited by Eq. (18.29). We leave to the reader the verification that the application of the raising operator a†to any valid'nproduces a new eigenfunction a†'nwith a positive norm. Our overall conclusion is that any system with a Hamiltonian of the form given by Eq. (18.21), whether or not represented by an ODE in ordinary space, will have eigenstates whose eigenvalues form a ladder of unit spacing, with the smallest eigenvalue equal to1 2. This makes it natural to label the states 'nby integers n0, and therefore to write H'nDn'n;  nDnC1 2;nD0;1;2; (18.30) in agreement with what we found from our original approach; compare Eq. (18.19). Before leaving this exercise in operator algebra, it may be worth noting that the notion of raising and lowering operators also arises in contexts where the states thereby reached can be interpreted as those containing different numbers of particles (or quasiparticles, a physics jargon that refers to objects, such as photons, whose population is easily changed by interaction with their surroundings). In such contexts, a raising operator is then often referred to as a creation operator, with a lowering operator then called an annihilation (or sometimes a destruction) operator. Obviously these terms have to be interpreted with an understanding of the underlying physics. Returning to the description of pas a differential operator, the equation a'0D0can be identified as a differential equation satisfied by the ground (lowest-energy) state of our oscillator. More specifically, p 2a'0D.xCip/'0D xCi id dx '0D xCd dx '0D0; (18.31) which has the advantage of being a first-order ODE. This ODE is separable, and can be integrated: d'0 '0Dx dx;ln'0Dx2 2Clnc0; ' 0Dc0ex2=2; in agreement with our previous analysis. Eigenstates for arbitrary ncan now be generated by repeated application of a†to'0. Doing so is left as an exercise. Molecular Vibrations In the dynamics and spectroscopy of molecules in the Born-Oppenheimer approximation, the motion of a molecule is separated into electronic, vibrational, and rotational motion. In ArfKen_Ch18-9780123846549.tex 18.2 Applications of Hermite Functions 883 treating the vibrational motion, the departure of nuclei from their equilibrium positions is to lowest order described by a quadratic potential, and the resulting oscillations are identified asharmonic. These harmonic motions can be treated as coupled simple harmonic oscil- lators, and we can decouple the individual nuclear motions by making a transformation to normal coordinates, as was illustrated in Example 6.5.2. In this harmonic oscillation limit, the vibrational wave functions have the form given in the preceding subsection, and the computation of properties associated with these wave functions then involve integrals in which products of Hermite functions appear. The simplest integrals occurring in vibrational problems are of the form 1Z 1xrex2Hn.x/Hm.x/dx: Examples for rD1andrD2(with nDm) are included in the exercises at the end of this section. A large number of other examples can be found in the work by Wilson, Decius, and Cross.2Some of the vibrational properties of molecules require the evaluation of integrals containing as many as four Hermite functions. In the remainder of this subsection we illustrate some of the possibilities and the associated mathematical procedures. Example 18.2.1 THREEFOLD HERMITE FORMULA Consider the following integral involving three Hermite polynomials I31Z 1ex2Hm1.x/Hm2.x/Hm3.x/dx; (18.32) where Ni0are integers. The formula (due to E. C. Titchmarsh, J. Lond. Math. Soc. 23: 15 (1948); see Gradshteyn and Ryzhik, p. 804, in Additional Readings) generalizes the I2 case needed for the orthogonality and normalization of Hermite polynomials. To start, we note that the integrand of I3will be even if the index sum m1Cm2Cm3is even, and odd if that index sum is odd, so I3will vanish unless m1Cm2Cm3is even. In addition, we see that if the product Hm1Hm2is expanded and written as a sum of Hermite polynomials, the resulting polynomial of largest index will be Hm1Cm2, soI3will vanish due to orthog- onality unless m1Cm2is at least as large as m3. This condition must continue to hold if the roles of the miare permuted; a convenient way of summarizing these observations is to state that the mimust satisfy a triangle condition. Both the even index sum and the triangle condition parallel similar conditions on integrals of Legendre polynomials which we encountered in Section 16.3 and discussed in detail at Eq. (16.85). 2E. B. Wilson, Jr., J. C. Decius, and P. C. Cross, Molecular Vibrations, New York: McGraw-Hill (1955), reprinted, Dover (1980). ArfKen_Ch18-9780123846549.tex 884 Chapter 18 More Special Functions To derive I3, we start with the product of three generating functions of Hermite polyno- mials, multiply by ex2;and integrate over x: Z31Z 1ex23Y jD1e2xtjt2 jdxD1Z 1e.t1Ct2Ct3x/2C2.t 1t2Ct1t3Ct2t3/dx Dpe2.t1t2Ct1t3Ct2t3/Dp1X ND02N NWX n1;n2;n30 n1Cn2Cn3DNNW n1Wn2Wn3Wtn2Cn3 1tn1Cn3 2tn1Cn2 3: (18.33) In reaching Eq. (18.33), we recognized the xintegration as an error integral, Eq. (1.148), and then expanded the resulting exponential, first as a power series in wD2.t1t2Ct1t3C t2t3/, and then expanding the powers of wby the generalization of the binomial theorem given as Eq. (1.80). Note that the index for the power of titjin the polynomial expansion was designated nk, where i;j;kare (in some order) 1, 2, 3. We next expand the generating functions in terms of Hermite polynomials and set the result equal to a slightly simplified version of the expression just obtained for Z3: Z3D1X m1;m2;m3D0tm1 1tm2 2tm3 3 m1Wm2Wm3W1Z 1ex2Hm1.x/Hm2.x/Hm3.x/dx Dp1X n1;n2;n3D02Ntn2Cn3 1tn1Cn3 2tn1Cn2 3 n1Wn2Wn3W; (18.34) with NDn1Cn2Cn3. In Eq. (18.34) we now equate the coefficients of equal powers of thetj, finding that m1Dn2Cn3,m2Dn1Cn3,m3Dn1Cn2, that NDm1Cm2Cm3 2; and that n1DNm1,n2DNm2,n3DNm3. From the coefficients of tm1 1tm2 2tm3 3, we obtain the final result I3Dp2Nm1Wm2Wm3W .Nm1/W.Nm2/W.Nm3/W: (18.35) Equation (18.35) explicitly reflects the necessity of the triangle condition. If it is not satis- fied but the sum of the miis even, at least one of the factorials in the denominator of Eq. (18.35) will have a negative integer argument, thereby causing I3to be zero. The requirement that the sum of the mibe even is not explicit in the form of Eq. (18.35), but the formula for I3is restricted to that case because the right-hand side of Eq. (18.34) only contains terms in which the sum of the powers of the tiis even.  Hermite Product Formula The integrals Imwith m>3can be obtained in closed form, but as finite sums. The starting point for that analysis is a formula for the product of two Hermite polynomials due to ArfKen_Ch18-9780123846549.tex 18.2 Applications of Hermite Functions 885 E. Feldheim, J. Lond. Math. Soc. 13: 22 (1938). To derive Feldheim’s formula, we can start from a product of two generating functions, written as e2x.t1Ct2/t2 1t2 2D1X m1;m2D0Hm1.x/Hm2.x/tm1 1 m1Wtm2 2 m2W De2x.t1Ct2/.t 1Ct2/2e2t1t2D1X nD0Hn.x/.t1Ct2/n nW1X D0.2t1t2/ W: Applying the binomial expansion to .t1Ct2/nand then comparing like powers of t1and t2in the two lines of the above equation, we find Hm1.x/Hm2.x/Dmin.m 1;m2/X D0Hm1Cm22.x/m1Wm2W2 W.m1Cm22/Wm1Cm22 m1 Dmin.m 1;m2/X D0Hm1Cm22.x/2Wm1 m2  : (18.36) ForD0the coefficient of HN1CN2is obviously unity. Special cases, such as H2 1DH2C2;H1H2DH3C4H1;H2 2DH4C8H2C8;H1H3DH4C6H2 can be derived from Table 13.1 and agree with the general twofold product formula. The product formula has been generalized to products of m>2Hermite polynomials, thereby providing a new way of evaluating the integrals Im. For details we refer the reader to work by Liang, Weber, Hayashi, and Lin.3 Example 18.2.2 FOURFOLD HERMITE FORMULA An important application of the Hermite product formula is a newly reported evaluation of the integral I4containing a product of four Hermite polynomials. The analysis is that of one of the present authors and his colleagues.3 The integral we are about to study is of the form I4D1Z 1ex2Hm1.x/Hm2.x/Hm3.x/Hm4.x/dx: (18.37) It is convenient to order the indices of the Hermite polynomials so that m1m2m3 m4. Our approach will be to apply the product formula to Hm1Hm2and to Hm3Hm4, thereby 3K. K. Liang, H. J. Weber, M. Hayashi, and S. H. Lin, Computational aspects of Franck-Condon overlap intervals. In Pandalai, S. G., ed., Recent Research Developments in Physical Chemistry, Vol. 8, Transworld Research Network (2005). ArfKen_Ch18-9780123846549.tex 886 Chapter 18 More Special Functions initially obtaining I4Dmin.m 1;m2/X D02Wm1 m2 min.m 3;m4/X D02Wm3 m4  1Z 1ex2Hm1Cm22.x/Hm3Cm42.x/dx: (18.38) Invoking the orthogonality of the Hmwith the weighting factor shown, the integral in Eq. (18.38) can be evaluated, yielding 1Z 1ex2Hm1Cm22.x/Hm3Cm42.x/dx Dp2m3Cm42.m3Cm42/Wm1Cm22;m 3Cm42: (18.39) The Kronecker delta in Eq. (18.39) limits the value of to the single value, if any, that satisfies Dm1Cm2m3m4 2C; (18.40) so the double summation collapses to a single sum over . Moreover, when the powers of 2 in Eqs. (18.38) and (18.39) are combined, their resultant is 2M, where MDm1Cm2Cm3Cm4 2: (18.41) We now rewrite Eq. (18.38), removing the summation and assigning the value from Eq. (18.40), writing the binomial coefficients in terms of their constituent factorials, and introducing Mwherever it will result in simplification. We reach I4DX p2M.m3Cm42/Wm1Wm2Wm3Wm4W .Mm3m4C/W.Mm1/W.Mm2/W.m3/W.m4/WW: (18.42) This formula for I4will only be valid when the sum of the miis even, equivalent to the requirement that M(and therefore also ) be integral. If the sum of the miis odd, then I4 will have an odd integrand and will vanish by symmetry. The summation in Eq. (18.42) will be over the nonnegative integral values of for which none of the factorials in the denominator of that summation has a negative argument. Note that there will be no value of that satisfies this condition if m1>m2Cm3Cm4, because Mm1will then be negative, and then I4D0. Thus we have a generalization of the triangle condition that applied to the threefold Hermite formula: If the largest of the miis greater than the sum of the others, the Hmof smaller mcannot combine to yield a Hermite polynomial of sufficiently large index to avoid an orthogonality zero. Further examination of the factorials in the denominator of Eq. (18.42) reveals that the lower limit of the summation will (if m1m2Cm3Cm4) always beD0; note that ArfKen_Ch18-9780123846549.tex 18.2 Applications of Hermite Functions 887 Mm3m4will always be nonnegative. The upper limit of the summation will be the smaller of m4andMm1.  The Hermite polynomial product formula can also be applied to products of Hermite polynomials with a different exponential weighting function than in the examples we have presented. To evaluate such integrals we use the generalized product formula in conjunc- tion with the integral (see Gradshteyn and Ryzhik, p. 803, in the Additional Readings), 1Z 1ea2x2Hm.x/Hn.x/dxD2mCn amCnC1.1a2/.mCn/=20mCnC1 2 min.m;n/X D0.m/.n/ W1mn 2 a2 2.a21/ ; (18.43) instead of the standard orthogonality integral for the product of two Hermite polynomials. The quantity .m/is a Pochhammer symbol, and causes the summation in Eq. (18.43) to be a finite sum. The summation can also be identified as a hypergeometric function; see Exercise 18.5.11. The process we have sketched yields a result that is similar to Imbut somewhat more complicated. We omit details. The oscillator potential has also been employed extensively in calculations of nuclear structure (nuclear shell model), as well as in quark models of hadrons and the nuclear force. Exercises 18.2.1 Prove that 2xd dxn 1DHn.x/: Hint. Check out the cases nD0andnD1and then use mathematical induction (Section 1.4). 18.2.2 Show thatZ1 1xmex2Hn.x/dxD0forman integer; 0mn1: 18.2.3 The transition probability between two oscillator states mandndepends on 1Z 1xex2Hn.x/Hm.x/dx: Show that this integral equals 1=22n1nWm;n1C1=22n.nC1/Wm;nC1. This result shows that such transitions can occur only between states of adjacent energy levels, mDn1. Hint. Multiply the generating function, Eq. (18.2), by itself using two different sets of variables .x;s/and.x;t/. Alternatively, the factor xmay be eliminated by the recurrence relation, Eq. (18.3). ArfKen_Ch18-9780123846549.tex 888 Chapter 18 More Special Functions 18.2.4 Show thatZ1 1x2ex2Hn.x/Hn.x/dxD1=22nnW nC1 2 : This integral occurs in the calculation of the mean-square displacement of our quantum oscillator. Hint. Use the recurrence relation, Eq. (18.3), and the orthogonality integral. 18.2.5 Evaluate 1Z 1x2ex2Hn.x/Hm.x/dx in terms of nandmand appropriate Kronecker delta functions. ANS. 2n11=2.2nC1/nWnmC2n1=2.nC2/WnC2;mC2n21=2nWn2;m . 18.2.6 Show thatZ1 1xrex2Hn.x/HnCp.x/dxD(0; p>r 2n1=2.nCr/W; pDr; with n;p, and rnonnegative integers. Hint. Use the recurrence relation, Eq. (18.3), ptimes. 18.2.7 With n.x/Dex2=2 Hn.x/ .2nnW1=2/1=2;verify that a n.x/Dxipp 2D1p 2 xCd dx n.x/Dn1=2 n1.x/; a† n.x/DxCipp 2D1p 2 xd dx n.x/D.nC1/1=2 nC1.x/: Note. The usual quantum mechanical operator approach establishes these raising and lowering properties before the form of n.x/is known. 18.2.8 (a) Verify the operator identity xCipDxd dxDexpx2 2d dxexp x2 2 : (b) The normalized simple harmonic oscillator wave function is n.x/D.1=22nnW/1=2exp x2 2 Hn.x/: Show that this may be written as n.x/D.1=22nnW/1=2 xd dxn exp x2 2 : ArfKen_Ch18-9780123846549.tex 18.3 Laguerre Functions 889 Note. This corresponds to an n-fold application of the raising operator of Exer- cise 18.2.7. 18.3 L AGUERRE FUNCTIONS Rodrigues Formula and Generating Function Let’s start from the Laguerre ODE, xy00.x/C.1x/y0.x/Cny.x/D0: (18.44) This ODE is not self-adjoint, but the weighting factor needed to make it self-adjoint can be computed from the usual formula, w.x/D1 xexpZ1x xdx D1 xexp.ln xx/Dex: (18.45) Givenw.x/, we may now use the method developed in Section 12.1 to obtain a Rodrigues formula and generating function for the Laguerre polynomials. Letting Ln.x/denote the nth Laguerre polynomial, the Rodrigues formula is (apart from a scale factor) given by Eq. (12.9): Ln.x/D1 w.x/d dxn w.x/p.x/n ; where p.x/is the coefficient of y00in the ODE. Inserting the expressions for w.x/and p.x/, and inserting a factor 1=nWto bring the Laguerre polynomials to their conventional scaling, the Rodrigues formula takes the more complete and explicit form, Ln.x/Dex nWd dxn xnex : (18.46) A generating function can now be written as a sum of contour integrals of the Schlaefli type, as in Eq. (12.25): g.x;t/D1X tD0Ln.x/tnD1 w.x/1X nD0cntnnW 2iI Cw.z/Tp.z/Un .zx/nC1dz; where the contour surrounds the point xand no other singularities. Specializing to our current problem, and noting that the coefficient cnhas the value 1=nW, this formula becomes g.x;t/Dex 2i1X nD0I Cez.tz/n .zx/nC1dzDex 2iI Cezdz .zx/1X nD0tz zxn : (18.47) We now recognize the nsummation as a geometric series, so our generating function becomes g.x;t/Dex 2iI Cezdz zxtz: (18.48) ArfKen_Ch18-9780123846549.tex 890 Chapter 18 More Special Functions Our integrand has a simple pole at zDx=.1t/, with residue ex=.1t/=.1t/, and g.x;t/ reduces to g.x;t/Dexex=.1t/ 1tDext=.1t/ 1tD1X nD0Ln.x/tn; (18.49) the form given in Table 12.1. Not all workers define Laguerre polynomials at the scale chosen here and represented by the specific formulas in Eq. (18.46) and(18.49). However, our choice is probably the most common, and is consistent with that in AMS-55 (see Abramowitz in Additional Readings). Properties of Laguerre Polynomials By differentiating the generating function in Eq. (18.45) with respect to xandt, we obtain recurrence relations for the Laguerre polynomials as follows. Using the product rule for differentiation we verify the identities .1t/2@g @tD.1xt/g.x;t/; .t1/@g @xDtg.x;t/: (18.50) Writing the left-hand and right-hand sides of the first identity in terms of Laguerre polyno- mials using the expansion given in Eq. (18.49), we obtain X n .nC1/LnC1.x/2nL n.x/C.n1/Ln1.x/ tn DX n .1x/Ln.x/Ln1.x/ tn: Equating coefficients of znyields .nC1/LnC1.x/D.2nC1x/Ln.x/nLn1.x/: (18.51) To get the second recursion relation we use both identities of Eqs. (18.50) to verify a third identity, x@g @xDt@g @t[email protected]/ @t; which, when written similarly in terms of Laguerre polynomials, is seen to be equivalent to x L0 n.x/DnLn.x/nLn1.x/: (18.52) To use these recurrence formulas we need starting values. From the Rodrigues formula, we easily find L0.x/D1andL1.x/D1x. Applying Eq. (18.51) we continue to Ln.x/ with n>1, obtaining the results given in Table 18.2. The first three Laguerre polynomials are plotted in Fig. 18.3. ArfKen_Ch18-9780123846549.tex 18.3 Laguerre Functions 891 Table 18.2 Laguerre Polynomials L0.x/D1 L1.x/DxC1 2WL2.x/Dx24xC2 3WL3.x/Dx3C9x218xC6 4WL4.x/Dx416x3C72x296xC24 5WL5.x/Dx5C25x4200x3C600x2600xC120 6WL6.x/Dx636x5C450x42400 x3C5400 x24320 xC720 1 −1 −2 −3L0(x) 1 24 L2(x) L1(x)x 3 FIGURE 18.3 Laguerre polynomials. From the recurrence relations or the Rodrigues formula, we find the the power series expansion of Ln.x/: Ln.x/D.1/n nW xnn2 1Wxn1Cn2.n1/2 2Wxn2C.1/nnW DnX mD0.1/mnWxm .nm/WmWmWDnX sD0.1/nsnWxns .ns/W.ns/WsW: (18.53) Also, from Eq. (18.49) we find g.0;t/D1 1tD1X nD0tnD1X nD0Ln.0/tn; which shows that at xD0the Laguerre polynomials have the special value Ln.0/D1: (18.54) ArfKen_Ch18-9780123846549.tex 892 Chapter 18 More Special Functions The form of the generating function, that of Laguerre’s ODE, and Table 18.2 all show that the Laguerre polynomials have neither odd nor even symmetry under the parity transfor- mation x! x. As we already observed at the beginning of this section, the Laguerre ODE is not self- adjoint but can be made so by appending the weighting factor ex. Noting also that with this weighting factor, the Laguerre polynomials satisfy Sturm-Liouville boundary condi- tions at xD0andxD1 , we see that the Ln.x/must satisfy an orthogonality condition of the form 1Z 0exLm.x/Ln.x/dxDmn: (18.55) Equation (18.55) indicates that for this interval and weighting factor the Laguerre polyno- mials are normalized. Proof is the topic of Exercise 18.3.3. It is sometimes convenient to define orthogonalized Laguerre functions (with unit weighting factor) by 'n.x/Dex=2Ln.x/: (18.56) Our new orthonormal functions, 'n.x/;satisfy the self-adjoint ODE x'00 n.x/C'0 n.x/C nC1 2x 4 'n.x/D0; (18.57) and are eigenfunctions of a Sturm-Liouville system on the range .0x<1/. Associated Laguerre Polynomials In many applications, particularly in quantum mechanics, we need the associated Laguerre polynomials defined by4 Lk n.x/D.1/kdk dxkLnCk.x/: (18.58) By differentiating the power series for Ln.x/given in Eq. (18.53) (compare Table 18.2), we can get the explicit forms shown in Table 18.3. In general, Lk n.x/DnX mD0.1/m.nCk/W .nm/W.kCm/WmWxm;k0: (18.59) One of the present authors5has recently found a new generating function for the associ- ated Laguerre polynomials with the remarkably simple form gl.x;t/Detx.1Ct/lD1X nD0Lln n.x/tn: (18.60) 4Some authors use Lk nCk.x/D.dk=dxk/TLnCk.x/U. Hence our Lkn.x/D.1/kLk nCk.x/. 5H. J. Weber, Connections between real polynomial solutions of hypergeometric-type differential equations with Rodrigues formula, Cent. Eur. J. Math. 5: 415–427 (2007). ArfKen_Ch18-9780123846549.tex 18.3 Laguerre Functions 893 Table 18.3 Associated Laguerre Polynomials Lk 0D1 1WLk 1DxC.kC1/ 2WLk 2Dx22.kC2/xC.kC1/2 3WLk 3Dx3C3.kC3/x23.kC2/2xC.kC1/3 4WLk 4Dx44.kC4/x3C6.kC3/24.kC2/3C.kC1/4 5WLk 5Dx5C5.kC5/x410.kC4/2x3C10.kC3/3x25.kC2/4xC.kC1/5 6WLk 6Dx66.kC6/x5C15.kC5/2x420.kC4/3x3C15.kC3/4x2 6.kC2/5xC.kC1/6 7WLk 7Dx7C7.kC7/x621.kC6/2x5C35.kC5/3x435.kC4/4x3 C21.kC3/5x27.kC2/6xC.kC1/7 The notations .kCn/mare Pochhammer symbols, defined in Eq. (1.72). Rather than deriving this formula, we verify it by showing that it produces the defining relation for the Lk n, Eq. (18.58), and is consistent with the previously presented formulas for the ordinary Laguerre polynomials (i.e., the Lk nwith kD0). If we multiply both members of Eq. (18.60) by 1t, the coefficients of tnyield the recurrence formula Lln nCLlnC1 n1DLlnC1 n;orLk nLkC1 nDLkC1 n1: (18.61) On the other hand, differentiation of Eq. (18.60) with respect to xand writing @gl.x;t/ @xDX nd Lln n.x/ dxtnDtetx.1Ct/lDetx.1Ct/letx.1Ct/lC1; the coefficients of tnyield a formula for d Lln n.x/=dx , namely (with kDln) d Lk n.x/ dxDLk n.x/LkC1 n.x/; (18.62) and substituting the result from Eq. (18.61), we reach d Lk n.x/ dxDLkC1 n1; (18.63) thereby confirming that our generating function yields Eq. (18.58). The verification that our generating function is correct is now completed by using it to findL0 n.x/, which is the coefficient of tninetx.1Ct/n. Using the binomial expansion of .1Ct/nand the Maclaurin series for the exponential, we get L0 n.x/DnX mD0n nm.x/m mWDnX mD0.1/mnW .nm/WmWmWxm; in agreement with Eq. (18.53). ArfKen_Ch18-9780123846549.tex 894 Chapter 18 More Special Functions We can also confirm the series expansion given as Eq. (18.59) forLk n. It is the coefficient oftninetx.1Ct/kCn, obtained in a manner similar to the procedure we just carried out forL0 n. The generating function provides convenient routes to other properties of the associated Laguerre polynomials. Special values for xD0can be obtained from X nLln n.0/tnD.1Ct/lDlX nD0l n tn: We therefore have Lk n.0/DnCk n : (18.64) A formula for recurrence in the index nofLk n.x/can be obtained by differentiating the generating function formula with respect to t. Doing so, from the coefficient of tnand setting lDkCn, .nC1/Lk1 nC1.x/D.kCn/Lk1 n.x/x Lk n.x/: (18.65) Using Eq. (18.61) to raise the upper index in the two terms for which it is k1, we find after collecting similar terms, .nC1/Lk nC1.x/.2nCkC1x/Lk n.x/C.nCk/Lk n1.x/D0; (18.66) a lower-index recurrence formula. Finally, returning to Eq. (18.65), differentiating it once with respect to x, and identifyingh Lk1 nC1i0 DLk n, we get .nCk/h Lk1 ni0 Dxh Lk ni0 CLk n.nC1/Lk nDxh Lk ni0 nLk n: (18.67) A second differentiation brings us to xh Lk ni00 C.1n/h Lk ni0 D.nCk/h Lk1 ni00 D.nCk/h Lk1 ni0 .nk/h Lk ni0 ; (18.68) where the final member of Eq. (18.68) was the result of substituting the derivative of Eq. (18.62) with k!k1. Using Eq. (18.67) to replace.nCk/ Lk1 n0by a form in which the upper index is k, we reach an ODE for Lk n: xd2Lk n.x/ dx2C.kC1x/d Lk n.x/ dxCnLk n.x/D0: (18.69) This ODE is known as the associated Laguerre equation. When associated Laguerre polynomials appear in a physical problem it is usually because that physical problem involves Eq. (18.69). The most important application is their use to describe the bound states of the hydrogen atom, which are derived in upcoming Example 18.3.1. ArfKen_Ch18-9780123846549.tex 18.3 Laguerre Functions 895 The associated Laguerre equation, Eq. (18.69), is not self-adjoint, but the weighting function needed to bring it to self-adjoint form (for upper index k) can be found in the usual way: wk.x/D1 xexpZkC1x xdx Dxkex: (18.70) When we also note that Sturm-Liouville boundary conditions are satisfied at xD0and xD1 , we see that the associated Laguerre polynomials are orthogonal according to the equation 1Z 0exxkLk n.x/Lk m.x/dxD.nCk/W nWmn: (18.71) The value of the integral in Eq. (18.71) formDncan be established using the generating function, Eq. (18.58). Doing so is left as an exercise. Equation (18.71) shows the same orthogonality interval .0;1/as that for the Laguerre polynomials, but with a different weighting function for each k. We see that for each kthe associated Laguerre polynomials define a new set of orthogonal polynomials. A Rodrigues representation of the associated Laguerre polynomials is useful and can be found in various ways. A fairly direct approach is simply to use Eq. (12.9) with p.x/Dx, the coefficient of the second-derivative term in Eq. (18.69) and the value of wk.x/given inEq. (18.70). The result is Lk n.x/Dexxk nWdn dxn.exxnCk/: (18.72) Note that this and all our earlier formulas involving the Lk n.x/reduce properly to corre- sponding expressions involving Ln.x/when kD0. By letting k n.x/Dex=2xk=2Lk n.x/, we find that k n.x/satisfies the self-adjoint ODE, xd2 k n.x/ dx2Cd k n.x/ dxC x 4C2nCkC1 2k2 4x k n.x/D0: (18.73) The k n.x/are sometimes called Laguerre functions . Equation (18.57) is the special case kD0of Eq. (18.73). A further useful form is given by defining6 8k n.x/Dex=2x.kC1/=2Lk n.x/: (18.74) Substitution into the associated Laguerre equation yields d28k n.x/ dx2C 1 4C2nCkC1 2xk21 4x2 8k n.x/D0: (18.75) The8k n.x/are orthogonal with weighting function x1. The associated Laguerre ODE, Eq. (18.69), has solutions even if nis not an integer, but they are then not polynomials and diverge proportionally to xkexasx!1 . This fact is useful in the following example. 6This corresponds to modifying the function in Eq. (18.73) to eliminate the first derivative. ArfKen_Ch18-9780123846549.tex 896 Chapter 18 More Special Functions Example 18.3.1 THEHYDROGEN ATOM The most important application of the Laguerre polynomials is in the solution of the Schrödinger equation for the hydrogen-like atom (H, HeC, Li2C, etc.). For a system con- sisting of a nucleus of charge Zefixed at the origin and one electron whose distribution is described by a wave function , this equation is Nh2 2mr2 Ze2 4 0r DE ; (18.76) in which ZD1for hydrogen, ZD2for HeC, and so on. Separating variables in spherical polar coordinates and recognizing that the angular part of the solution to this equation must be a spherical harmonic, we set .r/DR.r/YM L.;'/ with R.r/satisfying the ODE Nh2 2m1 r2d dr r2d R dr Ze2 4 0rRCNh2 2mL.LC1/ r2RDE: (18.77) For bound states, R!0asr!1 , and it can be shown that these conditions can only be met if E<0. In addition, Rmust be finite at rD0. We do not consider unbound (continuum) states with positive energy. Only when the latter are included do hydrogenic wave functions form a complete set. By use of the abbreviations (resulting from rescaling rto the dimensionless radial vari- able) D 8mE Nh21=2 ; D r; Dm Ze2 2 0 Nh2; ./DR.r/; (18.78) Eq. (13.85) becomes 1 2d d 2d./ d C 1 4L.LC1/ 2 ./D0: (18.79) For our present purposes, it is useful to rewrite the first term of Eq. (18.79) using the identity 1 2d d 2d d D1 d2 d2./ and then multiply the resulting equation by , reaching d d2./C 1 4L.LC1/ 2 ./D0: (18.80) A comparison with Eq. (18.75) for8k n.x/shows that Eq. (18.80) is satisfied by ./De=2LC1L2LC1 L1./; (18.81) where kandnofEq. (18.75) have been, respectively, replaced by 2LC1andL1. The parameter must be restricted to values such that L1is both integral and nonnegative. If this requirement is violated, L2LC1 L1will diverge too rapidly to permit ./ to go to zero at large r, which is required for a bound-state electron distribution. ArfKen_Ch18-9780123846549.tex 18.3 Laguerre Functions 897 Since we already know that L, a spherical harmonic index, must be integral and nonnega- tive, we see that the possible values of are integers nat least as large as LC1.7 This restriction on , imposed by our boundary condition, has the effect of quantizing the energy. Inserting Dn, the definitions in Eqs. (18.78) lead to EnDZ2m 2n2Nh2e2 4 02 : (18.82) Since our Schrödinger equation implicitly set the potential energy to zero when the electron is at an infinite separation from the nucleus, the negative sign reflects the fact that we are dealing here with bound states in which the electron cannot escape to infinity. The other quantities introduced in Eq. (18.78) can also be expressed in terms of n: Dme2 2 0Nh2Z nD2Z na0; D2Z na0r;with a0D4 0Nh2 me2: (18.83) The quantity a0, of dimension length, is known as the Bohr radius, and its appearance as a scale factor causes the potential energy (for nD1, the smallest possible value) to have an average value corresponding to this electron-nuclear separation. Summarizing, the final normalized hydrogen wave function is nL M.r;;'/D"2Z na03.nL1/W 2n.nCL/W#1=2 e r=2. r/LL2LC1 nL1. r/YM L.;'/: (18.84) Note that the energy corresponding to nL M depends only on n, which is called the principal quantum number of this system. Note also that if nis assigned a specific inte- gral value, the condition on requires that Ln1, thereby explaining the well-known pattern of possible hydrogenic energy states: If nD1,Lcan only be zero; for nD2, we can have LD0orLD1, etc.  Exercises 18.3.1 Show with the aid of the Leibniz formula that the series expansion of Ln.x/,Eq. (18.53), follows from the Rodrigues representation, Eq. (18.72). 18.3.2 (a) Using the explicit series form, Eq. (18.53), show that L0 n.0/Dn;L00 n.0/D1 2n.n1/: (b) Repeat without using the explicit series form of Ln.x/. 18.3.3 Derive the normalization relation, Eq. (18.71) for the associated Laguerre polynomials, thereby also confirming Eq. (18.55) for the Ln. 7This is the conventional notation for . It is not the same nas the index nin8kn.x/. ArfKen_Ch18-9780123846549.tex 898 Chapter 18 More Special Functions 18.3.4 Expand xrin a series of associated Laguerre polynomials Lk n.x/, with kfixed and n ranging from 0 to r(or to1ifris not an integer). Hint. The Rodrigues form of Lk n.x/will be useful. ANS. xrD.rCk/WrWrX nD0.1/nLk n.x/ .nCk/W.rn/W;0x<1. 18.3.5 Expand eaxin a series of associated Laguerre polynomials Lk n.x/;with kfixed and n ranging from 0 to1. (a) Evaluate directly the coefficients in your assumed expansion. (b) Develop the desired expansion from the generating function. ANS. eaxD1 .1Ca/1Ck1X nD0a 1Can Lk n.x/;0x<1. 18.3.6 Show thatZ1 0exxkC1Lk n.x/Lk n.x/dxD.nCk/W nW.2nCkC1/: Hint. Note that x Lk nD.2nCkC1/Lk n.nCk/Lk n1.nC1/Lk nC1: 18.3.7 Assume that a particular problem in quantum mechanics has led to the ODE d2y dx2k21 4x22nCkC1 2xC1 4 yD0 for nonnegative integers n;k:Write y.x/asy.x/DA.x/B.x/C.x/, with the require- ment that (a) A.x/be a negative exponential giving the required asymptotic behavior of y.x/, and (b) B.x/be a positive power of x giving the behavior of y.x/for0x1. Determine A.x/andB.x/. Find the relation between C.x/and the associated Laguerre polynomial. ANS. A.x/Dex=2;B.x/Dx.kC1/=2,C.x/DLk n.x/. 18.3.8 From Eq. (18.84) the normalized radial part of the hydrogenic wave function is RnL.r/D 3.nL1/W 2n.nCL/W1=2 e r. r/LL2LC1 nL1. r/; in which D2Z=na0D2Zme2=4 0Nh2. Evaluate .a/hriD1Z 0r RnL. r/RnL. r/r2dr; ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 899 .b/hr1iD1Z 0r1RnL. r/RnL. r/r2dr: The quantityhriis the average displacement of the electron from the nucleus, whereas hr1iis the average of the reciprocal displacement. ANS.hriDa0 2h 3n2L.LC1/i ,hr1iD1 n2a0. 18.3.9 Derive a recurrence formula for the hydrogen wave function expectation values: sC2 n2hrsC1i.2sC3/a0hrsiCsC1 4h .2LC1/2.sC1/2i a2 0hrs1iD0; with s2 L1. Hint. Transform Eq. (18.80) into a form analogous to Eq. (18.73). Multiply by sC2u0csC1u, with uD8. Adjust cto cancel terms that do not yield expecta- tion values. 18.3.10 Show that1Z 1xnex2Hn.xy/dxDpnWPn.y/;where Pnis a Legendre polynomial. 18.4 C HEBYSHEV POLYNOMIALS The generating function for the Legendre polynomials can be generalized to the following form: 1 .12xtCt2/ D1X nD0C. / n.x/tn: (18.85) The coefficients C. / n.x/are known as the ultraspherical polynomials (also called Gegenbauer polynomials). For D1=2, we recover the Legendre polynomials; the special cases D0and D1yield two types of Chebyshev polynomials that are the sub- ject of this section. The primary importance of the Chebyshev polynomials is in numerical analysis. Type II Polynomials With D1andC.1/ n.x/written as Un.x/,Eq. (18.85) gives 1 12xtCt2D1X nD0Un.x/tn;jxj<1;jtj<1: (18.86) These functions are called type II Chebyshev polynomials. Although these polynomials have few applications in mathematical physics, one unusual application is in the develop- ment of four-dimensional spherical harmonics used in angular momentum theory. ArfKen_Ch18-9780123846549.tex 900 Chapter 18 More Special Functions Type I Polynomials With D0there is a difficulty. Indeed, our generating function reduces to the constant 1. We may avoid this problem by first differentiating Eq. (18.85) with respect to t. This yields .2 xC2t/ .12xtCt2/ C1D1X nD1nC. / n.x/tn1; or xt .12xtCt2/ C1D1X nD1n 2" C. / n.x/ # tn1: (18.87) We define C.0/ n.x/as C.0/ n.x/Dlim !0C. / n.x/ : (18.88) The purpose of differentiating with respect to twas to get in the denominator and to create an indeterminate form. Now multiplying Eq. (18.87) by2tand adding 1 in the form .12xtCt2/=.12xtCt2/, we obtain 1t2 12xtCt2D1C21X nD1n 2C.0/ n.x/tn: (18.89) We define Tn.x/as Tn.x/D8 < :1; nD0; n 2C.0/ n.x/;n>0:(18.90) Note the special treatment for nD0. We will encounter a similar treatment of the nD0 term when we study Fourier series in Chapter 19. Also, note that C.0/ nis the limit indicated in Eq. (18.88) and not a literal substitution of D0into the generating function series. With these new labels, 1t2 12xtCt2DT0.x/C21X nD1Tn.x/tn;jxj1;jtj<1: (18.91) We call Tn.x/the type I Chebyshev polynomials. Note that the notation and spelling of the name for these functions differ from reference to reference. Here we follow the usage of AMS-55 (Additional Readings). ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 901 Recurrence Relations Differentiating the generating function, Eq. (18.91), with respect to tand multiplying by the denominator, 12xtCt2, we obtain t.tx/" T0.x/C21X nD1Tn.x/tn# D.12xtCt2/1X nD1nTn.x/tn1 D1X nD1h nTntn12xnT ntnCnTntnC1i ; from which after several simplification steps we reach the recurrence relation TnC1.x/2xTn.x/CTn1.x/D0; n>0: (18.92) A similar treatment of Eq. (18.86) yields the corresponding recursion relation for Un: UnC1.x/2xUn.x/CUn1.x/D0; n>0: (18.93) Using the generating functions directly for nD0and 1, and then applying these recur- rence relations for the higher-order polynomials, we get Table 18.4. Plots of the TnandUn are presented in Figs. 18.4 and18.5. Differentiation of the generating functions for Tn.x/andUn.x/with respect to the vari- able xleads to a variety of recurrence relations involving derivatives. For example, from Eq. (18.89) we thus obtain .12xtCt2/21X nD1T0 n.x/tnD2t" T0.x/C21X nD1Tn.x/tn# ; from which we extract the recursion formula 2Tn.x/DT0 nC1.x/2xT0 n.x/CT0 n1.x/: (18.94) Other useful recurrence formulas we can find in this way are .1x2/T0 n.x/DnxT n.x/CnTn1.x/ (18.95) Table 18.4 Chebyshev Polynomials: Type I (Left), Type II (Right) T0D1 U0D1 T1Dx U 1D2x T2D2x21 U2D4x21 T3D4x33x U 3D8x34x T4D8x48x2C1 U4D16x412x2C1 T5D16x520x3C5x U 5D32x532x3C6x T6D32x648x4C18x21 U6D64x680x4C24x21 ArfKen_Ch18-9780123846549.tex 902 Chapter 18 More Special Functions T1(x)1 −1−11 T3(x) T2(x)x FIGURE 18.4 The Chebyshev polynomials T1,T2, and T3. U1(x)5 4321 −1 1 −1 −2−3xU2(x)U3(x) FIGURE 18.5 The Chebyshev polynomials U1,U2, and U3. and .1x2/U0 n.x/DnxU n.x/C.nC1/Un1.x/: (18.96) Manipulating a variety of these formulas as in Section 15.1 for Legendre polynomials one can eliminate the index n1in favor of T00 nand establish that Tn.x/, the Chebyshev ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 903 polynomial type I, satisfies the ODE .1x2/T00 n.x/xT0 n.x/Cn2Tn.x/D0: (18.97) The Chebyshev polynomial of type II, Un.x/, satisfies .1x2/U00 n.x/3xU0 n.x/Cn.nC2/Un.x/D0: (18.98) We could have defined the Chebyshev polynomials starting from these ODEs, but we chose instead a development based on generating functions. Processes similar to those used for the Chebyshev polynomials can be applied to the general ultraspherical polynomials; the result is the ultraspherical ODE .1x2/d2 dx2C. / n.x/.2 C1/xd dxC. / n.x/Cn.nC2 /C. / n.x/D0: (18.99) Special Values Again, from the generating functions, we can obtain the special values of various polyno- mials: Tn.1/D1; Tn.1/D.1/n; T2n.0/D.1/n;T2nC1.0/D0I Un.1/DnC1;Un.1/D.1/n.nC1/; U2n.0/D.1/n;U2nC1.0/D0:(18.100) Verification of Eq. (18.100) is left to the exercises. The polynomials TnandUnsatisfy parity relations that follow from their generating functions with the substitutions t!t;x! x, which leave them invariant; these are Tn.x/D.1/nTn.x/;Un.x/D.1/nUn.x/: (18.101) Rodrigues representations of Tn.x/andUn.x/are Tn.x/D.1/n1=2.1x2/1=2 2n0.nC1 2/dn dxnh .1x2/n1=2i (18.102) and Un.x/D.1/n.nC1/1=2 2nC10.nC3 2/.1x2/1=2dn dxnh .1x2/nC1=2i : (18.103) ArfKen_Ch18-9780123846549.tex 904 Chapter 18 More Special Functions Trigonometric Form At this point in the development of the properties of the Chebyshev polynomials it is beneficial to change variables, replacing xbycos. With xDcosandd=dxD .1= sin/.d=d/;we verify that .1x2/d2Tn dx2Dd2Tn d2cotdTn d;xT0 nDcotdTn d: Adding these terms, Eq. (18.97) becomes d2Tn d2Cn2TnD0; (18.104) the simple harmonic oscillator equation with solutions cosnandsinn. The special val- ues (boundary conditions at xD0and 1) identify TnDcosnDcos.n arccos x/: (18.105) Forn6D0a second linearly independent solution of Eq. (18.104) is labeled VnDsinnDsin.n arccos x/: (18.106) The corresponding solutions of the type II Chebyshev equation, Eq. (18.98), become UnDsin.nC1/ sin; (18.107) WnDcos.nC1/ sin: (18.108) The two sets of solutions, type I and type II, are related by Vn.x/D.1x2/1=2Un1.x/; (18.109) Wn.x/D.1x2/1=2TnC1.x/: (18.110) As already seen from the generating functions, Tn.x/andUn.x/are polynomials. Clearly, Vn.x/andWn.x/arenotpolynomials. From Tn.x/CiVn.x/DcosnCisinn D.cosCisin/nDh xCi.1x2/1=2in ;jxj1 (18.111) we can apply the binomial theorem to obtain expansions Tn.x/Dxnn 2 xn2.1x2/Cn 4 xn4.1x2/2 (18.112) and, for n>0 Vn.x/Dp 1x2n 1 xn1n 3 xn3.1x2/C : (18.113) ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 905 From the generating functions, or from the ODEs, power-series representations are Tn.x/Dn 2Tn=2UX mD0.1/m.nm1/W mW.n2m/W.2x/n2m(18.114) forn1;withTn=2Uthe integer part of n=2and Un.x/DTn=2UX mD0.1/m.nm/W mW.n2m/W.2x/n2m: (18.115) Application to Numerical Analysis An important feature of the Chebyshev polynomials Tn.x/with n>0is that as xis varied, they oscillate between the extreme values TnDC1 andTnD1 . This behavior is readily seen from Eq. (18.105) and is illustrated for T12inFig. 18.6. If a function is expanded in theTnand the expansion is extended sufficiently that the contributions of successive Tnare decreasing rapidly, a good approximation to the truncation error will be proportional to the firstTnnot included in the expansion. In this approximation, there will be negligible error at the nvalues of xwhere Tnis zero, and there will be maximum errors (all of the same magnitude but alternating in sign) at the extrema of Tnthat fall between the zeros. In that sense, the errors satisfy a minimax principle, meaning that the maximum of the error has been minimized by distributing it evenly into the regions between the points of negligible error. −1−0.50.51 1 5 . 0 5 . 0− 1− FIGURE 18.6 The Chebyshev polynomial T12. ArfKen_Ch18-9780123846549.tex 906 Chapter 18 More Special Functions Example 18.4.1 MINIMIZING THE MAXIMUM ERROR Figure 18.7 shows the errors in four-term expansions of exon the rangeT1; 1Ucarried out in various ways: (a) Maclaurin series, (b) Legendre expansion, and (c) Chebyshev expansion. The power series is optimum at the point xD0and the error increases with increasing values of jxj. The orthogonal expansions produce a fit over the region T1; 1U, with the maximum errors occurring at xD1 and three intermediate values of x. How- ever, the Legendre expansion has larger errors at 1than it has at the interior points, while the Chebyshev expansion yields smaller errors at 1(with a concomitant increase in the error at the other maxima) with the result that all the error maxima are comparable. This choice approximately minimizes the maximum error. 0.006 (a)(c)(b) −0.006−1 1 FIGURE 18.7 Error in four-term approximations to ex: (a) Power series; (b) Legendre expansion; and (c) Chebyshev expansion.  Orthogonality If Eq. (18.97) is put into self-adjoint form (Section 8.2), we obtain w.x/D.1x2/1=2 as a weighting factor. For Eq. (18.98) the corresponding weighting factor is .1x2/C1=2. ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 907 The resulting orthogonality integrals, 1Z 1Tm.x/Tn.x/.1x2/1=2dxD8 >>< >>:0;m6Dn;  2;mDn6D0; ;mDnD0;(18.116) 1Z 1Vm.x/Vn.x/.1x2/1=2dxD8 >>< >>:0;m6Dn;  2;mDn6D0; 0;mDnD0;(18.117) 1Z 1Um.x/Un.x/.1x2/1=2dxD 2mn; (18.118) and 1Z 1Wm.x/Wn.x/.1x2/1=2dxD 2mn; (18.119) are a direct consequence of the Sturm-Liouville theory. The normalization values may best be obtained by making the substitution xDcos. Exercises 18.4.1 By evaluating the generating function for special values of x, verify the special values Tn.1/D1; Tn.1/D.1/n;T2n.0/D.1/n;T2nC1.0/D0: 18.4.2 By evaluating the generating function for special values of x, verify the special values Un.1/DnC1;Un.1/D.1/n.nC1/; U2n.0/D.1/n;U2nC1.0/D0: 18.4.3 Another Chebyshev generating function is 1xt 12xtCt2D1X nD0Xn.x/tn;jtj<1: How is Xn.x/related to Tn.x/andUn.x/? 18.4.4 Given .1x2/U00 n.x/3xU0 n.x/Cn.nC2/Un.x/D0; show that Vn.x/,Eq. (18.106), satisfies .1x2/V00 n.x/xV0 n.x/Cn2Vn.x/D0; which is Chebyshev’s equation. ArfKen_Ch18-9780123846549.tex 908 Chapter 18 More Special Functions 18.4.5 Show that the Wronskian of Tn.x/andVn.x/is given by Tn.x/V0 n.x/T0 n.x/Vn.x/Dn .1x2/1=2: This verifies that TnandVn.n6D0/are independent solutions of Eq. (18.97). Con- versely, for nD0, we do not have linear independence. What happens at nD0? Where is the “second” solution? 18.4.6 Show that Wn.x/D.1x2/1=2TnC1.x/is a solution of .1x2/W00 n.x/3xW0 n.x/Cn.nC2/Wn.x/D0: 18.4.7 Evaluate the Wronskian of Un.x/andWn.x/D.1x2/1=2TnC1.x/. 18.4.8 Vn.x/D.1x2/1=2Un1.x/is not defined for nD0. Show that a second and inde- pendent solution of the Chebyshev differential equation for Tn.x/.nD0/isV0.x/D arccos x(or arcsin x). 18.4.9 Show that Vn.x/satisfies the same three-term recurrence relation as Tn.x/,Eq. (18.92). 18.4.10 Verify the series solutions for Tn.x/andUn.x/,Eqs. (18.114) and(18.115). 18.4.11 Transform the series form of Tn.x/,Eq. (18.114), into an ascending power series. ANS. T2n.x/D.1/nnnX mD0.1/m.nCm1/W .nm/W.2m/W.2x/2m;n1, T2nC1.x/D2nC1 2nX mD0.1/mCn.nCm/W .nm/W.2mC1/W.2x/2mC1. 18.4.12 Rewrite the series form of Un.x/, Eq. (18.115), as an ascending power series. ANS. U2n.x/D.1/nnX mD0.1/m.nCm/W .nm/W.2m/W.2x/2m, U2nC1.x/D.1/nnX mD0.1/m.nCmC1/W .nm/W.2mC1/W.2x/2mC1. 18.4.13 (a) From the differential equation for Tn(in self-adjoint form) show that 1Z 1dTm.x/ dxdTn.x/ dx.1x2/1=2dxD0; m6Dn: (b) Confirm the preceding result by showing that dTn.x/ dxDnUn1.x/: 18.4.14 The substitution xD2x01converts Tn.x/into the shifted Chebyshev polynomials T n.x0/. Verify that this produces the shifted polynomials shown in Table 18.5 and that ArfKen_Ch18-9780123846549.tex 18.4 Chebyshev Polynomials 909 Table 18.5 Shifted Type I Chebyshev Polynomials T 0D1 T 1D2x1 T 2D8x28xC1 T 3D32x348x2C18x1 T 4D128x4256x3C160x232xC1 T 5D512x51280 x4C120x3400x2C50x1 T 6D2048 x66144 x5C6912 x43584 x3C840x272xC1 they satisfy the orthonormality condition 1Z 0T n.x0/Tm.x0/Tx.1x/U1=2dxDmn 2n0: 18.4.15 The expansion of a power of xin a Chebyshev series leads to the integral ImnD1Z 1xmTn.x/dxp 1x2: (a) Show that this integral vanishes for m<n. (b) Show that this integral vanishes for mCnodd. 18.4.16 Evaluate the integral ImnD1Z 1xmTn.x/dxp 1x2 formnandmCneven by each of two methods: (a) Replacing Tn.x/by its Rodrigues representation. (b) Using xDcosto transform the integral to a form with as the variable. ANS. ImnDmW .mn/W.mn1/WW .mCn/WW;mn;mCneven. 18.4.17 Establish the following bounds, 1x1: (a)jUn.x/jnC1, (b) d dxTn.x/ n2. 18.4.18 (a) Show that for1x1,jVn.x/j1: (b) Show that Wn.x/is unbounded in1x1. ArfKen_Ch18-9780123846549.tex 910 Chapter 18 More Special Functions 18.4.19 Verify the orthogonality-normalization integrals for (a) Tm.x/;Tn.x/; (b) Vm.x/,Vn.x/, (c) Um.x/;Un.x/;(d) Wm.x/;Wn.x/. Hint. All these can be converted to trigonometric integrals. 18.4.20 Show whether (a) Tm.x/andVn.x/are or are not orthogonal over the interval T1; 1Uwith respect to the weighting factor .1x2/1=2. (b) Um.x/andWn.x/are or are not orthogonal over the interval T1; 1Uwith respect to the weighting factor .1x2/1=2. 18.4.21 Derive (a) TnC1.x/CTn1.x/D2xTn.x/, (b) TmCn.x/CTmn.x/D2Tm.x/Tn.x/, from the “corresponding” cosine identities. 18.4.22 A number of equations relate the two types of Chebyshev polynomials. As examples show that Tn.x/DUn.x/xUn1.x/ and .1x2/Un.x/DxTnC1.x/TnC2.x/: 18.4.23 Show that dVn.x/ dxDnTn.x/p 1x2 (a) using the trigonometric forms of VnandTn, (b) using the Rodrigues representation. 18.4.24 Starting with xDcosandTn.cos/Dcosn, expand xkDeiCei 2k and show that xkD1 2k1 Tk.x/Ck 1 Tk2.x/Ck 2 Tk4C ; the series in brackets terminating after the term containing T1orT0. ArfKen_Ch18-9780123846549.tex 18.5 Hypergeometric Functions 911 18.4.25 Develop the following Chebyshev expansions (for T1; 1U): (a).1x2/1=2D2 " 121X sD1.4s21/1T2s.x/# , (b)C1; 0<x1 1;1x<0) D4 1X sD0.1/s.2sC1/1T2sC1.x/. 18.4.26 (a) For the interval T1; 1Ushow that jxjD1 2C1X sD1.1/sC1.2s3/WW .2sC2/WW.4sC1/P2s.x/ D2 C4 1X sD1.1/sC1 1 4s21T2s.x/: (b) Show that the ratio of the coefficient of T2s.x/to that of P2s.x/approaches.s/1 ass!1 . This illustrates the relatively rapid convergence of the Chebyshev series. Hint. With the Legendre recurrence relations, rewrite x Pn.x/as a linear combination of derivatives. The trigonometric substitution xDcos;Tn.x/Dcosnis most helpful for the Chebyshev part. 18.4.27 Show that 2 8D1C21X sD1.4s21/2: Hint. Apply Parseval’s identity (or the completeness relation) to the results of Exercise 18.4.26. 18.4.28 Show that (a) cos1xD 24 1X nD01 .2nC1/2T2nC1.x/: (b) sin1xD4 1X nD01 .2nC1/2T2nC1.x/: 18.5 H YPERGEOMETRIC FUNCTIONS In Chapter 7 the hypergeometric equation8 x.1x/y00.x/CTc.aCbC1/xUy0.x/ab y.x/D0 (18.120) 8This is sometimes called Gauss’ ODE. The solutions are then referred to as Gauss functions. ArfKen_Ch18-9780123846549.tex 912 Chapter 18 More Special Functions was introduced as a canonical form of a linear second-order ODE with regular singularities atxD0;1, and1. One solution, designated 2F1, is y.x/D2F1.a;bIcIx/ D1Ca b cx 1WCa.aC1/b.bC1/ c.cC1/x2 2WC; c6D0;1;2;3;:::; which is known as the hypergeometric function orhypergeometric series. For real a,b, andc(the only case considered here), the range of convergence for c>aCbis1x 1, while for aCb1<caCbthe convergence range is 1x<1. For caCb1 the hypergeometric series diverges. The terms of the hypergeometric series are conveniently written in terms of the Pochhammer symbol, introduced at Eq. (1.72); we repeat the definition here: .a/nDa.aC1/.aC2/.aCn1/; . a/0D1: Using this notation, the hypergeometric function becomes 2F1.a;bIcIx/D1X nD0.a/n.b/n .c/nxn nW: (18.121) In this form the significance of the subscripts 2 and 1 becomes clear. The leading sub- script 2 indicates that two Pochhammer symbols appear in the numerator and the trail- ing subscript 1 indicates one Pochhammer symbol in the denominator. The subscripts 2 and 1 are only useful if one intends to discuss analogs of the “standard” hypergeometric function that involve different numbers of Pochhammer symbols. We retain the subscripts because we will shortly identify confluent hypergeometric functions with forms similar to Eq. (18.121) but with only one Pochhammer symbol in the numerator, therefore of the form 1F1.aIcIz/. Note also that the numerator and denominator parameters are set off by semicolons (actually making the subscripts unnecessary). We retain them to conform to the most widely used notations for these functions. Looking further at Eq. (18.121), we note that the series will reduce to zero (for all x) ifcis either zero or a negative integer (unless the denominator is fortuitously cancelled by a particular choice of aorb). On the other hand, if aorbequals 0 or a negative inte- ger, the series terminates and the hypergeometric function becomes a polynomial. Many more or less elementary functions can be represented by the hypergeometric function.9For example, ln.1Cx/Dx2F1.1;1I2Ix/: (18.122) The hypergeometric equation as a second-order linear ODE has a second independent solution. The usual form is y.x/Dx1c2F1.aC1c;bC1cI2cIx/;c6D2;3;4;:::: (18.123) Ifcis an integer either the two solutions coincide or (barring a rescue by integral aor integral b) one of the solutions will blow up (see Exercise 18.5.1). In such a case the second solution is expected to include a logarithmic term. 9With three parameters, a;b, and c, we can represent almost anything. ArfKen_Ch18-9780123846549.tex 18.5 Hypergeometric Functions 913 Alternate forms of the hypergeometric ODE include .1z2/d2 dz21z 2 y h .aCbC1/z.aCbC12c/id dz1z 2 y ab1z 2 y D0; (18.124) .1z2/d2 dz2y.z2/ .2aC2bC1/zC12c zd dzy.z2/4ab y.z2/D0: (18.125) Contiguous Function Relations The parameters a;b, and center in the same way as the parameter nof Bessel, Legen- dre, and other special functions. As we found with these functions, we expect recurrence relations involving unit changes in the parameters a;b, and c. Hypergeometric functions that differ by1in a parameter are referred to as contiguous functions. Generalizing this term to include simultaneous unit changes in more than one parameter, we find 26 functions contiguous to 2F1.a;bIcIx/. Taking them two at a time, we can develop the formidable total of 325 equations among the contiguous functions. Two typical examples are .ab/n c.aCb1/C1a2b2CT.ab/21U.1x/o 2F1.a;bIcIx/ D.ca/.abC1/b 2F1.a1;bC1IcIx/ C.cb/.ab1/a2F1.aC1;b1IcIx/; (18.126) T2acC.ba/xU2F1.a;bIcIx/Da.1x/2F1.aC1;bIcIx/ .ca/2F1.a1;bIcIx/: (18.127) Many more contiguous relations can be found in AMS-55 or in Olver et al. (Additional Readings). Hypergeometric Representations A number of the special functions introduced in this book can be expressed in terms of hypergeometric functions. The identification can usually be made by noting that these functions are solutions of ODEs that are special cases of the hypergeometric ODE. It is also necessary to determine the factors needed to express the functions at the agreed-upon scale. We cite several examples. 1. The ultraspherical functions C. / n.x/satisfy the ODE given as Eq. (18.99), and since that equation is a special case of the hypergeometric equation, Eq. (18.120), we see that ultraspherical functions (and Legendre and Chebyshev functions) may be expressed as ArfKen_Ch18-9780123846549.tex 914 Chapter 18 More Special Functions hypergeometric functions. For the ultraspherical function we obtain C. / n.x/D.nC2 /W 2 nW0. C1/2F1 n;nC2 C1I1C I1x 2 ; (18.128) with the factor preceding the 2F1function determined by requiring C. / nto have the proper scale. 2. For Legendre and associated Legendre functions we find Pn.x/D2F1 n;nC1I1I1x 2 ; (18.129) Pm n.x/D.nCm/W .nm/W.1x2/m=2 2mmW2F1 mn;mCnC1ImC1I1x 2 :(18.130) Alternate forms for the Legendre functions are P2n.x/D.1/n.2n/W 22nnWnW2F1 n;nC1 2I1 2Ix2 D.1/n.2n1/WW .2n/WW2F1 n;nC1 2I1 2Ix2 ; (18.131) P2nC1.x/D.1/n.2nC1/W 22nnWnWx2F1 n;nC3 2I3 2Ix2 D.1/n.2nC1/WW .2n/WWx2F1 n;nC3 2I3 2Ix2 : (18.132) 3. The Chebyshev functions have representations Tn.x/D2F1 n;nI1 2I1x 2 ; (18.133) Un.x/D.nC1/2F1 n;nC2I3 2I1x 2 ; (18.134) Vn.x/Dnp 1x22F1 nC1;nC1I3 2I1x 2 : (18.135) The leading factors are determined by direct comparison of complete power series, comparison of coefficients of particular powers of the variable, or evaluation at xD0or 1. The hypergeometric series may be used to define functions with nonintegral indices. The physical applications are minimal. ArfKen_Ch18-9780123846549.tex 18.5 Hypergeometric Functions 915 Exercises 18.5.1 (a) For c, an integer, and aandbnonintegral, show that 2F1.a;bIcIx/and x1c2F1.aC1c;bC1cI2cIx/ yield only one solution to the hypergeometric equation. (b) What happens if ais an integer, say, aD1 , and cD2 ? 18.5.2 Find the Legendre, Chebyshev I, and Chebyshev II recurrence relations corresponding to the hypergeometric contiguous function relation given as Eq. (18.126). 18.5.3 Transform the following polynomials into hypergeometric functions of argument x2: (a) T2n.x/; (b) x1T2nC1.x/; (c) U2n.x/; (d) x1U2nC1.x/. ANS. (a) T2n.x/D.1/n2F1.n;nI1 2Ix2/. (b)x1T2nC1.x/D.1/n.2nC1/2F1 n;nC1I3 2Ix2 . (c)U2n.x/D.1/n2F1 n;nC1I1 2Ix2 . (d)x1U2nC1.x/D.1/n.2nC2/2F1 n;nC2I3 2Ix2 . 18.5.4 Derive or verify the leading factor in the hypergeometric representations of the Chebyshev functions. 18.5.5 Verify that the Legendre function of the second kind, Q.z/, is given by Q.z/D1=2W 0.C3 2/.2z/C12F1 2C1 2; 2C1I 2C3 2Iz2 ; wherejzj>1,jargzj<, and6D1;2;3;. 18.5.6 The incomplete beta function was defined in Eq. (13.78) as Bx.p;q/DxZ 0tp1.1t/q1dt: Show that Bx.p;q/Dp1xp2F1.p;1qIpC1Ix/: 18.5.7 Verify the integral representation 2F1.a;bIcIz/D0.c/ 0.b/0.cb/1Z 0tb1.1t/cb1.1tz/adt: What restrictions must be placed on the parameters bandc? ArfKen_Ch18-9780123846549.tex 916 Chapter 18 More Special Functions Note. Although the power series used to establish this integral representation is only valid forjzj<1, the representation is valid for general z, as can be established by analytic continuation. For nonintegral athe real axis in the z-plane from 1 to1is a cut line. Hint. The integral is suspiciously like a beta function and can be expanded into a series of beta functions. ANS. c>b>0. 18.5.8 Prove that 2F1.a;bIcI1/D0.c/0.cab/ 0.ca/0.cb/;c6D0;1;2;:::; c>aCb: Hint. Here is a chance to use the integral representation in Exercise 18.5.7. 18.5.9 Prove that 2F1.a;bIcIx/D.1x/a2F1 a;cbIcIx 1x : Hint. Try an integral representation. Note. This relation is useful in developing a Rodrigues representation of Tn.x/(see Exercise 18.5.10). 18.5.10 Derive the Rodrigues representation of Tn.x/, Tn.x/D.1/n1=2.1x2/1=2 2n.n1 2/Wdn dxnh .1x2/n1=2i : Hint. One possibility is to use the hypergeometric function relation 2F1.a;bIcIz/D.1z/a2F1 a;cbIcIz 1z ; with zD.1x/=2. An alternate approach is to develop a first-order differential equation foryD.1x2/n1=2. Repeated differentiation of this equation leads to the Chebyshev equation. 18.5.11 Show that the summation in Eq. (18.43), min.m;n/X D0.m/.n/ W1mn 2 a2 2.a21/ ; can be written as a hypergeometric function. 18.5.12 Verify that 2F1.n;bIcI1/D.cb/n .c/n: Hint. Here is a chance to use the contiguous function relation Eq. (18.127) and math- ematical induction (Section 1.4). Alternatively, use the integral representation and the beta function. ArfKen_Ch18-9780123846549.tex 18.6 Con/f_luent Hypergeometric Functions 917 18.6 C ONFLUENT HYPERGEOMETRIC FUNCTIONS The confluent hypergeometric equation,10 xy00.x/C.cx/y0.x/ay.x/D0; (18.136) has a regular singularity at xD0and an irregular one at xD1 . It is obtained from the hypergeometric equation of Section 18.5 in the limit that one of the singularities at finite x is merged with that at infinity, causing that singularity to become irregular. One solution of the confluent hypergeometric equation is y.x/D1F1.aIcIx/DM.a;c;x/ D1Ca cx 1WCa.aC1/ c.cC1/x2 2WC; c6D0;1;2;: (18.137) The notation M.a;c;x/(with commas, not semicolons) has become standard for this solu- tion. It is convergent for all finite x(or complex z). In terms of the Pochhammer symbols, we have M.a;c;x/D1X nD0.a/n .c/nxn nW: (18.138) Clearly, M.a;c;x/becomes a polynomial if the parameter ais 0 or a negative integer. Numerous more or less elementary functions may be represented by the confluent hyper- geometric function. Examples are the error function and the incomplete gamma function: erf.x/D2 1=2xZ 0et2dtD2 1=2x M1 2;3 2;x2 ; (18.139) .a;x/DxZ 0etta1dtDa1xaM.a;aC1;x/;<e.a/>0: (18.140) A second solution of Eq. (18.136) is given by y.x/Dx1cM.aC1c;2c;x/;c6D2;3;4;: (18.141) Clearly, this coincides with the first solution for cD1. The standard form of the second solution of Eq. (18.136) is a linear combination of Eqs. (18.137) and(18.141): U.a;c;x/D sincM.a;c;x/ 0.acC1/0.c/x1cM.aC1c;2c;x/ 0.a/0.c/ :(18.142) Note the resemblance to our definition of the Neumann function, Eq. (14.57). As with the Neumann function, this definition of U.a;c;x/becomes indeterminate for certain param- eter values, namely when cis an integer. 10This is often called Kummer’s equation. The solutions, then, are Kummer functions. ArfKen_Ch18-9780123846549.tex 918 Chapter 18 More Special Functions An alternate form of the confluent hypergeometric equation is obtained by changing the independent variable from xtox2: d2 dx2y.x2/C2c1 x2xd dxy.x2/4ay.x2/D0: (18.143) As with the hypergeometric functions, contiguous functions exist in which the param- eters aandcare changed by1. Including the cases of simultaneous changes in the two parameters, we have eight possibilities. Taking the original function and pairs of the con- tiguous functions, we can develop a total of 28 equations. The recurrence relations for Bessel, Hermite, and Laguerre functions are special cases of these equations. Integral Representations It is frequently convenient to have the confluent hypergeometric functions in integral form. We find (Exercise 18.6.10) M.a;c;x/D0.c/ 0.a/0.ca/1Z 0extta1.1t/ca1dt;c>a>0; (18.144) U.a;c;x/D1 0.a/1Z 0extta1.1Ct/ca1dt;<e.x/>0;a>0: (18.145) Three important techniques for deriving or verifying integral representations are as fol- lows: 1. Transformation of generating function expansions and Rodrigues representations: The Bessel and Legendre functions provide examples of this approach. 2. Direct integration to yield a series: This direct technique is useful for a Bessel function representation (Exercise 14.1.17) and a hypergeometric integral (Exercise 18.5.7). 3. (a) Verification that the integral representation satisfies the ODE. (b) Exclusion of the other solution. (c) Verification of normalization. This is the method used in Sec- tion 14.6 to establish an integral representation of the modified Bessel function K.z/. It will work here to establish Eqs. (18.144) and(18.145). Confluent Hypergeometric Representations Special functions that can be represented in terms of confluent hypergeometric functions include the following: 1.Bessel functions: J.x/Deix 0.C1/x 2 M C1 2;2C1;2ix ; (18.146) ArfKen_Ch18-9780123846549.tex 18.6 Con/f_luent Hypergeometric Functions 919 whereas for the modified Bessel functions of the first kind, I.x/Dex 0.C1/x 2 M C1 2;2C1;2x : (18.147) 2.Hermite functions: H2n.x/D.1/n.2n/W nWM n;1 2;x2 : (18.148) H2nC1.x/D.1/n2.2nC1/W nWx M n;3 2;x2 ; (18.149) using Eq. (13.150). 3.Laguerre functions: Ln.x/DM.n;1;x/: (18.150) The constant is fixed as unity by noting Eq. (18.54) forxD0. For the associated Laguerre functions, Lm n.x/D.1/mdm dxmLnCm.x/D.nCm/W nWmWM.n;mC1;x/: (18.151) Alternate verification is obtained by comparing Eq. (18.151) with the power-series solu- tion, Eq. (18.59). Note that in the hypergeometric form, as distinct from a Rodrigues rep- resentation, the indices nandmneed not be integers, but if they are not integers, Lm n.x/ will not be a polynomial. Further Observations There are certain advantages in expressing our special functions in terms of hypergeo- metric and confluent hypergeometric functions. If the general behavior of the latter func- tions is known, the behavior of the special functions we have investigated follows as a series of special cases. This may be useful in determining asymptotic behavior or evalu- ating normalization integrals. The asymptotic behavior of M.a;c;x/andU.a;c;x/may be conveniently obtained from integral representations of these functions, Eqs. (18.144) and(18.145). The further advantage is that the relations between the special functions are clarified. For instance, an examination of Eqs. (18.148), (18.149), and (18.151) suggests that the Laguerre and Hermite functions are related. The confluent hypergeometric equation, Eq. (18.136), is clearly not self-adjoint. For this and other reasons it is convenient to define Mk.x/Dex=2xC1=2M.kC1 2;2C1;x/: (18.152) This new function, Mk.x/;is called a Whittaker function; it satisfies the self-adjoint equa- tion M00 k.x/C 1 4Ck xC1 42 x2! Mk.x/D0: (18.153) ArfKen_Ch18-9780123846549.tex 920 Chapter 18 More Special Functions The corresponding second solution is Wk.x/Dex=2xC1=2U.kC1 2;2C1;x/: (18.154) Exercises 18.6.1 Verify the confluent hypergeometric representation of the error function erf.x/D2x 1=2M1 2;3 2;x2 : 18.6.2 Show that the Fresnel integrals C.x/ands.x/of Exercise 12.6.1 may be expressed in terms of the confluent hypergeometric function as C.x/Cis.x/Dx M1 2;3 2;ix2 2 : 18.6.3 By direct differentiation and substitution verify that yDaxaxZ 0etta1dtDaxa .a;x/ satisfies xy00C.aC1Cx/y0CayD0: 18.6.4 Show that the modified Bessel function of the second kind, K.x/;is given by K.x/D1=2ex.2x/U.C1 2;2C1;2x/: 18.6.5 Show that the cosine and sine integrals of Section 13.6 may be expressed in terms of confluent hypergeometric functions as Ci.x/Cisi.x/DeixU.1;1;ix/: This relation is useful in numerical computation of Ci .x/and si.x/for large values of x. 18.6.6 Verify the confluent hypergeometric form of the Hermite polynomial H2nC1.x/, Eq. (18.149), by showing that (a) H2nC1.x/=xsatisfies the confluent hypergeometric equation with aDn ,cD3=2 and argument x2, (b) lim x!0H2nC1.x/ xD.1/n2.2nC1/W nW. 18.6.7 Show that the contiguous confluent hypergeometric function equation .ca/M.a1;c;x/C.2acCx/M.a;c;x/aM.aC1;c;x/D0 leads to the associated Laguerre function recurrence relation, Eq. (18.66). ArfKen_Ch18-9780123846549.tex 18.6 Con/f_luent Hypergeometric Functions 921 18.6.8 Verify the Kummer transformations: (a) M.a;c;x/DexM.ca;c;x/; (b) U.a;c;x/Dx1cU.acC1;2c;x/. 18.6.9 Prove that (a)dn dxnM.a;c;x/D.a/n .b/nM.aCn;bCn;x/, (b)dn dxnU.a;c;x/D.1/n.a/nU.aCn;cCn;x/. 18.6.10 Verify the following integral representations: (a) M.a;c;x/D0.c/ 0.a/0.ca/Z1 0extta1.1t/ca1dt,c>a>0, (b) U.a;c;x/D1 0.a/Z1 0extta1.1Ct/ca1dt;<e.x/>0,a>0. Under what conditions can you accept <e.x/D0in part (b)? 18.6.11 From the integral representation of M.a;c;x/, Exercise 18.6.10(a), show that M.a;c;x/DexM.ca;c;x/: Hint. Replace the variable of integration tby1sto release a factor exfrom the integral. 18.6.12 From the integral representation of U.a;c;x/in Exercise 18.6.10(b), show that the exponential integral is given by E1.x/DexU.1;1;x/: Hint. Replace the variable of integration tinE1.x/byx.1Cs/. 18.6.13 From the integral representations of M.a;c;x/andU.a;c;x/in Exercise 18.6.10, develop asymptotic expansions of (a)M.a;c;x/, (b) U.a;c;x/. Hint. You can use the technique that was employed with K.z/in Section 14.6. ANS. (a)0.c/ 0.a/ex xca 1C.1a/.ca/ 1WxC.1a/.2a/.ca/.caC1/ 2Wx2C ; (b)1 xa 1Ca.1Cac/ 1W.x/Ca.aC1/.1Cac/.2Cac/ 2W.x/2C . ArfKen_Ch18-9780123846549.tex 922 Chapter 18 More Special Functions 18.6.14 Show that the Wronskian of the two confluent hypergeometric functions M.a;c;x/and U.a;c;x/is given by MU0M0UD.c1/W .a1/Wex xc: What happens if ais 0 or a negative integer? 18.6.15 The Coulomb wave equation (radial part of the Schrödinger equation with Coulomb potential) is d2y dr2C 12 rL.LC1/ r2 yD0: Show that a regular solution yDFL.;r/is given by FL.;r/DCL./rLC1eirM.LC1i;2LC2;2ir/: 18.6.16 (a) Show that the radial part of the hydrogen-atom wave function, Eq. (18.81), may be written as e r=2. r/LL2LC1 nL1. r/D .nCL/W .nL1/W.2 LC1/We r=2. r/LM.LC1n;2LC2; r/: (b) It was assumed previously that the total (kinetic Cpotential) energy Eof the electron was negative. Rewrite the (unnormalized) radial wave function for an unbound hydrogenic electron, E>0. ANS. ei r=2. r/LM.LC1in;2LC2;i r/, outgoing wave. This repre- sentation provides a powerful alternative technique for the calculation of photoionization and recombination coefficients. 18.6.17 Evaluate (a)Z1 0TMk.x/U2dx, (b)Z1 0TMk.x/U2dx x, (c)Z1 0TMk.x/U2dx x1a, where 2D0;1;2;::: ,k1 2D0;1;2;::: ,a>21. ANS. (a)2k.2/W , (b).2/W , (c).2k/a.2/W . ArfKen_Ch18-9780123846549.tex 18.7 Dilogarithm 923 18.7 D ILOGARITHM Thedilogarithm, defined as Li2.z/DzZ 0ln.1t/ tdt (18.155) and its analytic continuation beyond the range of convergence of the above integral, arises in the evaluation of matrix elements in few-body problems of atomic physics and in vari- ous perturbation-theoretic contributions to quantum electrodynamics. Because of a historic lack of familiarity with this special function among physicists, many places of its occur- rence have only been recognized in recent years. Expansion and Analytic Properties Expanding the logarithm in Eq. (18.155), using the series in Eq. (1.97), we directly obtain the series expansion Li2.z/D1X nD1zn n2: (18.156) Note that we have inserted the logarithm without an additional multiple of 2i, thereby obtaining the branch of Li 2that is nonsingular at zD0. Further applications of the operator that converts ln.1z/into Li 2.z/produce poly- logarithms, which also occur in physics, albeit less frequently: Lip.z/DzZ 0Lip1.t/dt tD1X nD1zn nppD3;4;:::: (18.157) However, in this text we limit consideration to the first member of this sequence, Li 2. The series expansion of Li 2,Eq. (18.156), has circle of convergence jzjD1, with con- vergence for all zon this circle. The singularity limiting the radius of convergence is not apparent from the form of the expansion, but, looking at Eq. (18.155), we identify it as a branch point located at zD1. It is customary to draw a branch cut from zD1tozD1 along, and just below the positive real axis, and to define the principal value of Li 2as that which corresponds to Eq. (18.156) and its analytic continuation. From the form of Eq. (18.156), it is apparent that for real zin the interval1zC1 , Li2.z/will also be real. For z>1, we see from Eq. (18.155) that for part of the range of integration, the factor ln.1t/will necessarily be complex, with the result that Li 2.z/will no longer be real, even for real z. However, there is no similar problem for negative real z, as the principal value of ln.1t/remains real for all negative real values of t. Analyzing further the behavior of the integral in Eq. (18.155), we note that if we reach a point zby carrying out the integral, along a path (in t) that goes first from tD0to just above the branch point at tD1, and then in a straight line to z, we will for the last segment of the path alter the argument of 1tby some amount in the clockwise direction, thereby adding an amountito the numerator of the integrand. See Fig. 18.8. ArfKen_Ch18-9780123846549.tex 924 Chapter 18 More Special Functions 01z z . 2π−θ 01.θ FIGURE 18.8 Contours for integral representation of dilogarithm. This addition to the numerator means that the evaluation of Li 2.z/will have the form Li2.z/D1Z 0ln.1t/ tdtzZ 1ln.j1tj/ tdtCizZ 1dt t DLi2.1/zZ 1ln.j1tj/ tdtCilnz(path above zD1). (18.158) If we repeat the above analysis to reach the same point zby a path (in t) that passes around zD1below the real axis, the argument of 1twill be changed by an amount 2in the counterclockwise direction, and Li2.z/DLi2.1/zZ 1ln.j1tj/ tdti.2/lnz(path below zD1). (18.159) Comparing Eqs. (18.158) and(18.159), we see that the values of Li 2.z/, for the same z, but on these two different branches the values, will differ by an amount 2ilnz. Ifzis complex, the difference will affect both the real and imaginary parts of Li 2.z/, in ways more complicated than either changing the phase or adding a multiple of to the imagi- nary part. When working with the dilogarithm, it is therefore essential to make a careful determination of the branch on which it is to be evaluated. In fact, whenever possible formulas involving the dilogarithm and (because of the context) known to be real-valued should be manipulated (using formulas such as those in the next subsection) to cause each dilogarithm in the formula to be for a value of zthat is real and with z<1. Properties and Special Values From Eq. (18.156), we see that Li 2.0/D0:Setting zD1, we note that we get the series for.2/, so Li 2.1/D.2/D2=6. We also have Li 2.1/D.2/ , where.2/ is the Dirichlet series in Eq. (12.62), so Li 2.1/D2=12. ArfKen_Ch18-9780123846549.tex 18.7 Dilogarithm 925 The dilogarithm has a derivative that follows directly from Eq. (18.155), dLi2.z/ dzDln.1z/ z; (18.160) and possesses several functional relations enabling an easy analytic continuation beyond the convergence range of Eq. (18.156). Some of these are the following: Li2.z/CLi2.1z/D2 6lnzln.1z/ (18.161) Li2.z/CLi2.z1/D2 61 2ln2.z/ (18.162) Li2.z/CLi2z z1 D1 2ln2.1z/: (18.163) These relationships are most easily established by showing that the derivatives of both sides of the equations are equal and that the values of the two sides correspond for some convenient value of z. These functional relations enable the determination of Li 2.z/for all realzfrom values on the real line in the range jzj1 2, for which the series in Eq. (18.155) converges rapidly. From the functional relations it is possible to identify a few more specific values of z for which the principal value of Li 2.z/can be expressed in terms of elementary functions. For example, Li 2.1=2/D1 2ln2.2/C2=12. But for most z, closed expressions are not available. Example 18.7.1 CHECK USEFULNESS OF FORMULA The integral ID1 82ZZ d3r1d3r2e r1 r2 r12 r2 1r2 2r12 arises in computations of the electronic structure of the He atom. Here riare the positions of two electrons relative to the nucleus (which is at the origin of our coordinate system), the integration is over the full three-dimensional spaces of r1andr2,riDjr ij, and r12D jr1r2j. This integral is found to have the value ID1 2 6CLi2 C  CLi2 C  C1 2ln2 C C  : We now ask: Are its individual terms real? We note from the definition of Ithat it will be convergent only if C , C , and C are all positive. If that is not the case, in the portion of the space in which some particle is far from the other two, the overall exponential will increase without limit. Looking now ArfKen_Ch18-9780123846549.tex 926 Chapter 18 More Special Functions at the formula for the integral, we see immediately that the ln2term will be real, as its argument is the quotient of two positive numbers. The first Li 2term can be written Li2 C  DLi2 1 C C  ; showing that the argument of Li 2is real and less than C1, meaning that this Li 2will evaluate to a real result. Similar observations apply to the second instance of Li 2. We conclude that our formula is in a proper form for unambiguous computation using principal values of its multivalued functions.  Exercises 18.7.1 Prove that the expansion of Li 2.z/,Eq. (18.156), converges everywhere on the circle jzjD1. 18.7.2 Use the functional relations, Eqs. (18.161) to(18.163), to find the principal value of Li2.1=2/ . 18.7.3 Find all the multiple values of Li 2.1=2/ . 18.7.4 Explain why Eq. (18.161) gives the expected result for zD0when on the principal branch of the dilogarithm. 18.7.5 Show that Li21Cz1 2 DLi 21Cz 1z 1 2ln21z1 2 : 18.7.6 The following integral arises in the computation of the electronic energy of the Li atom using a correlated wave function (one that explicitly includes the electron-electron dis- tances as well as the distances of electrons from the nucleus): IDZZZ d3r1d3r2d3r3e 1r1 2r2 3r3 r1r2r3r12r13r23; where riDjr ij,ri jDjr irjj, and the integrations are over the entire three-dimensional space of each ri. For convergence of I, we require all j>0, but there are no restric- tions on their relative magnitudes. In terms of the auxiliary quantities 1D 1 2C 3;  2D 2 1C 3;  3D 3 1C 2; this integral has the value ID323 1 2 30 @2 2C3X jD1 Li2.j/Li2. j/Clnjln1j 1Cj1 A ArfKen_Ch18-9780123846549.tex 18.8 Elliptic Integrals 927 Rearrange Ito a form (first found by Remiddi11), in which all terms in the final expres- sion are guaranteed to evaluate to real quantities and can be evaluated as principal val- ues. 18.8 E LLIPTIC INTEGRALS Elliptic integrals occasionally arise in physical problems and therefore it is worthwhile to summarize their definitions and properties. Before the advent of computers, it was also important for physicists and engineers to be familiar with methods for hand computation of elliptic integrals, but that need has diminished with time and expansion methods for these functions will not be emphasized here. We do, however, illustrate problems in which elliptic integrals arise; the following example is a case in point. Example 18.8.1 PERIOD OF A SIMPLE PENDULUM For small-amplitude oscillations, a pendulum (Fig. 18.9) has simple harmonic motion with a period TD2.l=g/1=2. But for a maximum amplitude Mlarge enough that sinMcan- not be approximated by M, a direct application of Newton’s second law of motion and solution of the resulting ODE becomes difficult. In that situation a good way to proceed is to write the equation for conservation of energy. Setting the zero of potential energy at the point from which the pendulum is suspended, the potential energy of a pendulum of mass mand length lat angleismgl cos, and its total energy (the potential energy at angleM) ismgl cosM. The pendulum has kinetic energy ml2.d=dt/2=2, so energy conservation requires 1 2ml2d dt2 mglcosDmgl cosM: (18.164) Solving for d=dt we obtain d dtD2g l1=2 .coscosM/1=2; (18.165) θ m FIGURE 18.9 Simple pendulum. 11E. Remiddi, Analytic value of the atomic three-electron correlation integral with Slater wave functions. Phys. Rev. A 44: 5492 (1991). ArfKen_Ch18-9780123846549.tex 928 Chapter 18 More Special Functions with the mass mcanceling out. At tD0we choose as initial conditions D0andd=dt> 0. An integration from D0toDMyields MZ 0.coscosM/1=2dD2g l1=2tZ 0dtD2g l1=2 t: (18.166) This is1 4of a cycle, and therefore the time tis1 4of the period T. We note that M, and with a bit of clairvoyance we try the half-angle substitution sin 2 DsinM 2 sin': (18.167) With this, Eq. (18.166) becomes TD4l g1=2=2Z 0 1sin2M 2 sin2'1=2 d': (18.168) The integral in Eq. (18.168) does not reduce to an elementary function; in fact, it is an ellip- tic integral of a standard type. Further examples of elliptic integrals in physical problems can be found in the exercises.  Definitions Theelliptic integral of the first kind is defined as F.'n /D'Z 0.1sin2 sin2/1=2d; (18.169) or F.xjm/DxZ 0h .1t2/.1mt2/i1=2 dt;0m<1: (18.170) This is the notation of AMS-55 (Additional Readings). Note the use of the separators nand jto identify the specific functional forms. When the upper limit in these integrals is set to 'D=2 orxD1, we have the complete elliptic integral of the first kind, K.m/D=2Z 0.1msin2/1=2d D1Z 0h .1t2/.1mt2/i1=2 dt;(18.171) with mDsin2 ,0m<1. ArfKen_Ch18-9780123846549.tex 18.8 Elliptic Integrals 929 Theelliptic integral of the second kind is defined by E.'n /D'Z 0.1sin2 sin2/1=2d (18.172) or E.xjm/DxZ 01mt2 1t21=2 dt;0m1: (18.173) Again, for the case 'D=2; xD1, we have the complete elliptic integral of the second kind: E.m/D=2Z 0.1msin2/1=2d D1Z 01mt2 1t21=2 dt;0m1:(18.174) Series Expansions For our range 0m<1, the denominator of K.m/may be expanded by the binomial series in Eq. (1.74): .1msin2/1=2D1X nD0.2n1/WW .2n/WWmnsin2n; after which the resulting series is then integrated term by term. The integrals of the indi- vidual terms are beta functions (see Exercise 13.3.8), and we get K.m/D 2( 1C1X nD1.2n1/WW .2n/WW2 mn) : (18.175) Similarly (see Exercise 18.8.2), E.m/D 2( 11X nD1.2n1/WW .2n/WW2mn 2n1) : (18.176) These series can be identified as hypergeometric functions. Comparing with the general definitions in Section 18.5, we have K.m/D 22F1.1 2;1 2I1Im/; E.m/D 22F1.1 2;1 2I1Im/:(18.177) The complete elliptic integrals are plotted in Fig. 18.10. ArfKen_Ch18-9780123846549.tex 930 Chapter 18 More Special Functions 3.0 2.0 1.0 0π/2 1.0 1.0 0.5 mE(m)K(m) FIGURE 18.10 Complete elliptic integrals, K.m/andE.m/. Limiting Values From the series Eqs. (18.175) and(18.176), or from the defining integrals, lim m!0K.m/D 2;lim m!0E.m/D 2: (18.178) Form!1the series expansions are of little use. However, the integrals yield lim m!1K.m/D1; lim m!1E.m/D1: (18.179) The divergence in K.m/is logarithmic. Elliptic integrals have been used extensively in the past for evaluating integrals. For instance, general integrals of the form IDxZ 0R t;p a4t4Ca3t3Ca2t2Ca1t1Ca0 dt; where Ris a rational function of its arguments, may be expressed in terms of elliptic integrals. Jahnke and Emde (Additional Readings) give pages of such transformations. With computers available for direct numerical evaluation, interest in these elliptic integral techniques has declined. A more extensive account of elliptic functions, integrals, and the related Jacobi theta functions can be found in Whittaker and Watson’s treatise. Many ArfKen_Ch18-9780123846549.tex 18.8 Elliptic Integrals 931 formulas and tables of elliptic integrals are in AMS-55 and even more formulas are in Olver et al. (all of these sources are in the Additional Readings). Exercises 18.8.1 The ellipse x2=a2Cy2=b2D1may be represented parametrically by xDasin,yD bcos. Show that the length of arc within the first quadrant is a=2Z 0.1msin2/1=2dDaE.m/: Here 0mD.a2b2/=a21. 18.8.2 Derive the series expansion E.m/D 2( 11 22m 113 242m2 3) : 18.8.3 Show that lim m!0.KE/ mD 4: 18.8.4 A circular loop of wire in the xy-plane, as shown in Fig. 18.11, carries a current I. Given that the vector potential is A'.;'; z/Da0I 2Z 0cos d .a2C2Cz22acos /1=2; xa yϕ (ρ, ϕ, z) (ρ, ϕ, 0)ρz FIGURE 18.11 Circular wire loop. ArfKen_Ch18-9780123846549.tex 932 Chapter 18 More Special Functions show that A'.;'; z/D0I ka 1=2 1k2 2 K.k2/E.k2/ ; where k2D4a .aC/2Cz2: Note. For extension of this exercise to B, see Smythe.12 18.8.5 An analysis of the magnetic vector potential of a circular current loop leads to the expression f.k2/Dk2h .2k2/K.k2/2E.k2/i ; where K.k2/andE.k2/are the complete elliptic integrals of the first and second kinds. Show that for k21.rradius of loop) f.k2/k2 16: 18.8.6 Show that (a)d E.k2/ dkD1 k.EK/, (b)d K.k2/ dkDE k.1k2/K k. Hint. For part (b) show that E.k2/D.1k2/=2Z 0.1ksin2/3=2d by comparing series expansions. Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions, Applied Mathematics Series-55 (AMS-55). Washington, DC: National Bureau of Standards (1964), paperback edition, Dover (1974). Chapter 22 is a detailed summary of the properties and representations of orthogonal polynomials. Other chapters summarize properties of Bessel, Legendre, hypergeometric, and confluent hypergeometric functions and much more. See also Olver et al., below. Buchholz, H., The Confluent Hypergeometric Function. New York: Springer Verlag (1953), translated (1969). Buchholz strongly emphasizes the Whittaker rather than the Kummer forms. Applications to a variety of other transcendental functions. 12W. R. Smythe, Static and Dynamic Electricity, 3rd ed. New York: McGraw-Hill (1969), p. 270. ArfKen_Ch18-9780123846549.tex Additional Readings 933 Erdelyi, A., W. Magnus, F. Oberhettinger, and F. G. Tricomi, Higher Transcendental Functions, 3 vols. New York: McGraw-Hill (1953), reprinted, Krieger (1981). A detailed, almost exhaustive listing of the properties of the special functions of mathematical physics. Fox, L., and I. B. Parker, Chebyshev Polynomials in Numerical Analysis. Oxford: Oxford University Press (1968). A detailed, thorough, but very readable account of Chebyshev polynomials and their applications in numerical analysis. Gradshteyn, I. S., and I. M. Ryzhik, Table of Integrals, Series and Products (A. Jeffrey and D. Zwillinger, eds.), 7th ed. New York: Academic Press (2007). Jahnke, E., and F. Emde, Tables of Functions with Formulae and Curves. Leipzig: Teubner (1933), Dover (1945). Jahnke, E., F. Emde, and F. Lösch, Tables of Higher Functions, 6th ed. New York: McGraw-Hill (1960). An enlarged update of the work by Jahnke and Emde. Lebedev, N. N., Special Functions and Their Applications (translated by R. A. Silverman). Englewood Cliffs, NJ: Prentice-Hall (1965), paperback, Dover (1972). Luke, Y. L., The Special Functions and Their Approximations, 2 vols. New York: Academic Press (1969). Volume 1 is a thorough theoretical treatment of gamma functions, hypergeometric functions, confluent hyper- geometric functions, and related functions. Volume 2 develops approximations and other techniques for numerical work. Luke, Y. L., Mathematical Functions and Their Approximations. New York: Academic Press (1975). This is an updated supplement to Handbook of Mathematical Functions with Formulas, Graphs and Mathematical Tables (AMS-55). Magnus, W., F. Oberhettinger, and R. P. Soni, Formulas and Theorems for the Special Functions of Mathematical Physics. New York: Springer (1966). An excellent summary of just what the title says. Olver, F. W. J., D. W. Lozier, R. F. Boisvert, and C. W. Clark, eds., NIST Handbook of Mathematical Functions. Cambridge: Cambridge University Press (2010). Update of AMS-55 (Abramowitz and Stegun, above), but links to computer programs are provided instead of tables of data. Rainville, E. D., Special Functions. New York: Macmillan (1960), reprinted, Chelsea (1971). This book is a coherent, comprehensive account of almost all the special functions of mathematical physics that the reader is likely to encounter. Sansone, G., Orthogonal Functions (translated by A. H. Diamond). New York: Interscience (1959), reprinted, Dover (1991). Slater, L. J., Confluent Hypergeometric Functions. Cambridge: Cambridge University Press (1960). This is a clear and detailed development of the properties of the confluent hypergeometric functions and of relations of the confluent hypergeometric equation to other ODEs of mathematical physics. Sneddon, I. N., Special Functions of Mathematical Physics and Chemistry, 3rd ed. New York: Longman (1980). Whittaker, E. T., and G. N. Watson, A Course of Modern Analysis. Cambridge: Cambridge University Press, reprinted (1997). The classic text on special functions and real and complex analysis. ArfKen_Ch19-9780123846549.tex CHAPTER 19 FOURIER SERIES Periodic phenomena involving waves, rotating machines (harmonic motion), or other repetitive driving forces are described by periodic functions. Fourier series are a basic tool for solving ordinary differential equations (ODEs) and partial differential equations (PDEs) with periodic boundary conditions. Fourier integrals for nonperiodic phenomena are developed in Chapter 20. The common name for the field is Fourier analysis. 19.1 G ENERAL PROPERTIES A Fourier series is defined as an expansion of a function or representation of a function in a series of sines and cosines, such as f.x/Da0 2C1X nD1ancosnxC1X nD1bnsinnx: (19.1) The coefficients a0;an, and bnare related to f.x/by definite integrals: anD1 2Z 0f.s/cosnsds; nD0;1;2;:::; (19.2) bnD1 2Z 0f.s/sinnsds; nD1;2;:::; (19.3) which are subject to the requirement that the integrals exist. Note that a0is singled out for special treatment by the inclusion of the factor1 2. This is done so that Eq. (19.2) will apply to all an,nD0as well as n>0. The conditions imposed on f.x/to make Eq. (19.1) valid are that f.x/have only a finite number of finite discontinuities and only a finite number of extreme values (max- ima and minima) in the interval T0;2U.1Functions satisfying these conditions may be 1These conditions are sufficient but not necessary. 935 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch19-9780123846549.tex 936 Chapter 19 Fourier Series called piecewise regular. The conditions themselves are known as the Dirichlet condi- tions. Although there are some functions that do not obey these conditions, they can be considered pathological for purposes of Fourier expansions. In the vast majority of physi- cal problems involving a Fourier series, the Dirichlet conditions will be satisfied. Expressing cosnxandsinnxin exponential form, we may rewrite Eq. (19.1) as f.x/D1X nD1cneinx; (19.4) in which cnD1 2.anibn/;cnD1 2.anCibn/;n>0; (19.5) and c0D1 2a0: (19.6) Sturm-Liouville Theory The ODE y00.x/Dy.x/ on the intervalT0;2Uwith boundary conditions y.0/Dy.2/ ,y0.0/Dy0.2/ is a Sturm-Liouville problem, and these boundary conditions make it Hermitian. Therefore its eigenfunctions, either cosnx.nD0;1;:::/ andsinnx.nD1;2;:::/ , orexp.inx/ .nD:::;1;0;1;:::/ , form a complete set, with eigenfunctions of different eigenval- ues orthogonal. Since the eigenfunctions have respective values n2, those of differentjnj will automatically be orthogonal, while those of the same jnjcan be orthogonalized if necessary. Defining the scalar product for this problem as hfjgiD2Z 0f.x/g.x/dx; it is easy to check that heinxjeinxiD0forn6D0, and if we write cosnxandsinnx as complex exponentials, it is also easy to see that hsinnxjcosnxiD0. To make the eigenfunctions normalized, a simple approach is to note that the average value of sin2nx orcos2nxover an integer number of oscillations is 1/2 (again for n6D0), so 2Z 0sin2nx dxD2Z 0cos2nx dxD .n6D0/; andheinxjeinxiD2. The relationships identified above indicate that the eigenfunctions 'nDeinx=p 2,.nD :::;1;0;1;:::/ form an orthonormal set, as do '0D1p 2; ' nDcosnxp; 'nDsinnxp; .nD1;2;:::/; ArfKen_Ch19-9780123846549.tex 19.1 General Properties 937 so expansions in these functions have the forms given in Eqs. (19.1) to(19.3) orEqs. (19.4) to(19.6). Since we know that the eigenfunctions of a Sturm-Liouville operator form a complete set, we know that our Fourier-series expansions of L2functions will at least converge in the mean. Discontinuous Functions There are significant differences between the behavior of Fourier- and power-series expansions. A power series is essentially an expansion about a point, using only infor- mation from that point about the function to be expanded (including, of course, the values of its derivatives). We already know that such expansions only converge within a radius of convergence defined by the position of the nearest singularity. However, a Fourier series (or any expansion in orthogonal functions) uses information from the entire expansion interval, and therefore can describe functions that have “nonpathological” singularities within that interval. However, we also know that the representation of a function by an orthogonal expansion is only guaranteed to converge in the mean. This feature comes into play for the expansion of functions with discontinuities, where there is no unique value to which the expansion must converge. However, for Fourier series, it can be shown that if a function f.x/satisfying the Dirichlet conditions is discontinuous at a point x0, its Fourier series evaluated at that point will be the arithmetic average of the limits of the left and right approaches: fFourier series.x0/Dlim "!0f.x0C"/Cf.x0"/ 2 : (19.7) For proof of Eq. (19.7), see Jeffreys and Jeffreys or Carslaw (Additional Readings). It can also be shown that if the function to be expanded is continuous but has a finite dis- continuity in its first derivative, its Fourier series will then exhibit uniform convergence (see Churchill, Additional Readings). These features make Fourier expansions useful for functions with a variety of types of discontinuities. Example 19.1.1 SAWTOOTH WAVE An idea of the convergence of a Fourier series and the error in using only a finite number of terms in the series may be obtained by considering the expansion of f.x/D( x; 0x<; x2; < x2:(19.8) This is a sawtooth wave form, as shown in Fig. 19.1. Using Eqs. (19.2) and(19.3), we find the expansion to be f.x/D2 sinxsin 2x 2Csin 3x 3C.1/nC1sinnx nC : (19.9) Figure 19.2 shows f.x/for0x<2for the sum of 4, 6, and 10 terms of the series. Three features deserve comment. ArfKen_Ch19-9780123846549.tex 938 Chapter 19 Fourier Series 2π π −ππ FIGURE 19.1 Sawtooth wave form. 24610terms 0 −2π FIGURE 19.2 Expansion of sawtooth wave form, range T0;2U. 1. There is a steady increase in the accuracy of the representation as the number of terms included is increased. 2. At xD, where f.x/changes discontinuously from Cto, all the curves pass through the average of these two values, namely f./D0. 3. In the vicinity of the discontinuity at xD, there is an overshoot that persists and shows no sign of diminishing. As a matter of incidental interest, setting xD=2 in Eq. (19.9) leads to f 2 D 2D2 101 30C1 501 7C ; ArfKen_Ch19-9780123846549.tex 19.1 General Properties 939 thereby yielding an alternate derivation of Leibniz’s formula for =4, which was obtained by another method in Exercise 1.3.2.  Periodic Functions Fourier series are used extensively to represent periodic functions, especially wave forms for signal processing. The form of the series is inherently periodic; the expansions in Eqs. (19.1) and(19.4) are periodic with period 2, with sinnx,cosnx, and exp.inx/, each completing ncycles of oscillation in that interval. Thus, while the coefficients in a Fourier expansion are determined from an interval of length 2, the expansion itself (if the function involved is actually periodic) applies for an indefinite range of x. The period- icity also means that the interval used for determining the coefficients need not be T0;2U but may be any other interval of that length. Often one encounters situations in which the formulas in Eqs. (19.2) and(19.3) are changed so that their integrations run between  and. In fact, it would have been natural to have restated Example 19.1.1 as dealing with f.x/Dx, for< x<. This of course does not remove the discontinuity or change the form of the Fourier series. The discontinuity has simply been moved to the ends of the interval in x. In actual situations, the natural interval for a Fourier expansion will be the wavelength of our wave form, so it may make sense to redefine our Fourier series so that Eq. (19.1) becomes f.x/Da0 2C1X nD1ancosnx LC1X nD1bnsinnx L; (19.10) with anD1 LLZ Lf.s/cosns Lds; nD0;1;2;:::; (19.11) bnD1 LLZ Lf.s/sinns Lds; nD1;2;:::: (19.12) In many problems the xdependence of a Fourier expansion describes the spatial depen- dence of a wave distribution that is moving (say, toward Cx) with phase velocity v. This means that in place of xwe need to write xvt, and this substitution carries the implicit assumption that the wave form retains the same shape as it moves forward.2The individual terms of the Fourier expansion can now be given an interesting interpretation. Taking as an example the term coshn L.xvt/i ; 2For waves in physical media, this assumption is by no means always true, as it depends on the time-dependent response properties of the medium. ArfKen_Ch19-9780123846549.tex 940 Chapter 19 Fourier Series we note that it describes a contribution of wavelength 2L=n(when xincreases this much at constant t, the argument of the cosine function increases by 2). We also note that the period of the oscillation (the change in tat constant xfor one cycle of the cosine function) is TD2L=nv, corresponding to the oscillation frequency Dnv=2L. If we call the frequency for nD1thefundamental frequency and denote it 0Dv=2L, we identify the terms for each n>1in the Fourier series as describing overtones, or harmonics of the fundamental frequency, with individual frequencies n0. A typical problem for which Fourier analysis is suitable is one in which a particle under- going oscillatory motion is subject to a periodic driving force. If the problem is described by a linear ODE, we may make a Fourier expansion of the driving force and solve for each harmonic individually. This makes the Fourier expansion a practical tool as well as a nice analytical device. We stress, however, that its utility depends crucially on the linearity of our problem; in nonlinear problems an overall solution is not a superposition of component solutions. As suggested earlier, we have proceeded on the assumption that v, the phase velocity, is the same for all terms of the Fourier series. We now see that this assumption corresponds to the notion that the medium supporting the wave motion can respond equally well to forces at all frequencies. If, for example, the medium consists of particles too massive to respond quickly at high frequency, those components of the wave form will become attenuated and damped out of a propagating wave. Conversely, if the system contains components that resonate at certain frequencies, the response at those frequencies will be enhanced. Fourier expansions give physicists (and engineers) a powerful tool for analyzing wave forms and for designing media (e.g., circuits) that yield desired behaviors. One question that is sometimes raised is: “Were the harmonics there all along, or were they created by our Fourier analysis?” One answer compares the functional resolution into harmonics with the resolution of a vector into rectangular components. The components may have been present, in the sense that they may be isolated and observed, but the reso- lution is certainly not unique. Hence many authors prefer to say that the harmonics were created by our choice of expansion. Other expansions in other sets of orthogonal functions would produce a different decomposition. For further discussion, we refer to a series of notes and letters in the American Journal of Physics.3 What if a function is not periodic? We can still obtain its Fourier expansion, but (a) the results will of course depend on how the expansion interval is chosen (both as to posi- tion and length), and (b) because no information outside the expansion interval was used in obtaining the expansion, we can have no realistic expectation that the expansion will produce there a reasonable approximation to our function. Symmetry Suppose we have a function f.x/that is either an even or an odd function of x. If it is even, then its Fourier expansion cannot contain any odd terms (since all terms are linearly independent, no odd term can be removed by retaining others). Our expansion, developed 3B. L. Robinson, Concerning frequencies resulting from distortion. Am. J. Phys. 21: 391 (1953); F. W. Van Name, Jr., Concern- ing frequencies resulting from distortion. Am J. Phys. 22: 94 (1954). ArfKen_Ch19-9780123846549.tex 19.1 General Properties 941 for the intervalT;U, then must take the form f.x/Da0 2C1X nD1ancosnx;f.x/even: (19.13) On the other hand, if f.x/is odd, we must have f.x/D1X nD1bnsinnx;f.x/odd: (19.14) In both cases, when determining the coefficients we only need consider the interval T0;U, referring to Eqs. (19.2) and(19.3), as the adjoining interval of length will make a con- tribution identical to that considered. The series in Eqs. (19.13) and(19.14) are sometimes called Fourier cosine andFourier sine series. If we have a function defined on the interval T0;U, we can represent it either as a Fourier sine series or as a Fourier cosine series (or, if it has no interfering singularities, as a power series), with similar results on the interval of definition. However, the results outside that interval may differ markedly because these expansions carry different assumptions as to symmetry and periodicity. Example 19.1.2 DIFFERENT EXPANSIONS OF f.x/Dx We consider three possible ways to expand f.x/Dxbased on its values on the range T0;U: Its power-series expansion will (obviously) have the power-series expansion f.x/Dx. Comparing with Example 19.1.1, its Fourier sine series will have the form given in Eq. (19.9). Its Fourier cosine series will have coefficients determined from anD2 Z 0xcosnx dxD8 >>>< >>>:; nD0; 4 n2;nD1;3;5;:::; 0;nD2;4;6;:::; corresponding to the expansion f.x/D 21X nD04 cos.2nC1/x .2nC1/2: All three of these expansions represent f.x/well in the range of definition, T0;U, but their behavior becomes strikingly different outside that range. We compare the three expansions for a range larger than T0;UinFig. 19.3.  ArfKen_Ch19-9780123846549.tex 942 Chapter 19 Fourier Series (b) (a)(c)(a) (c) (b)0(abc) (ab)−π π 2π FIGURE 19.3 Expansions of f.x/DxonT0;U: (a) power series, (b) Fourier sine series, (c) Fourier cosine series. Operations on Fourier Series Term-by-term integration of the series f.x/Da0 2C1X nD1ancosnxC1X nD1bnsinnx (19.15) yields xZ x0f.x/dxDa0x 2 x x0C1X nD1an nsinnx x x01X nD1bn ncosnx x x0: (19.16) Clearly, the effect of integration is to place an additional power of nin the denomina- tor of each coefficient. This results in more rapid convergence than before. Consequently, a convergent Fourier series may always be integrated term by term, the resulting series converging uniformly to the integral of the original function. Indeed, term-by-term inte- gration may be valid even if the original series, Eq. (19.15), is not itself convergent. The function f.x/need only be integrable. A discussion will be found in Jeffreys and Jeffreys (Additional Readings). Strictly speaking, Eq. (19.16) may not be a Fourier series; that is, if a06D0, there will be a term1 2a0x. However, xZ x0f.x/dx1 2a0x (19.17) will still be a Fourier series. ArfKen_Ch19-9780123846549.tex 19.1 General Properties 943 The situation regarding differentiation is quite different from that of integration. Here the word is caution. Consider the series for f.x/Dx;< x<: (19.18) We readily found (in Example 19.1.1) that the Fourier series is xD21X nD1.1/nC1sinnx n;< x<: (19.19) Differentiating term by term, we obtain 1D21X nD1.1/nC1cosnx; (19.20) which is not convergent. Warning: Check your derivative for convergence. For the triangular wave shown in Fig. 19.4 (and treated in Exercise 19.2.9), the Fourier expansion is f.x/D 24 1X nD1;oddcosnx n2; (19.21) which converges more rapidly than the expansion of Eq. (19.19); in fact, it exhibits uniform convergence. Differentiating term by term we get f0.x/D4 1X nD1;oddsinnx n; (19.22) which is the Fourier expansion of a square wave, f0.x/D( 1; 0<x<; 1;< x<0:(19.23) Inspection of Fig. 19.3 verifies that this is indeed the derivative of our triangular wave. f(x) xπ 2π 3π 4ππ −4π −3π −2π −π FIGURE 19.4 Triangular wave. ArfKen_Ch19-9780123846549.tex 944 Chapter 19 Fourier Series As the inverse of integration, the operation of differentiation has placed an additional factor nin the numerator of each term. This reduces the rate of convergence and may, as in the first case mentioned, render the differentiated series divergent. In general, term-by-term differentiation is permissible if the series to be differentiated is uniformly convergent. Summing Fourier Series Often the most efficient way to identify the function represented by a Fourier series is simply to identify the expansion in a table. But if it is our desire to sum the series ourselves, a useful approach is to replace the trigonometric functions by their complex exponential forms, and then identifying the Fourier series as one or more power series in eix. Example 19.1.3 SUMMATION OF A FOURIER SERIES Consider the seriesP1 nD1.1=n/cosnx;x2.0;2/. Since this series is only conditionally convergent (and diverges at xD0), we take 1X nD1cosnx nDlim r!11X nD1rncosnx n; absolutely convergent for jrj<1. Our procedure is to try forming power series by trans- forming the trigonometric functions into exponential form: 1X nD1rncosnx nD1 21X nD1rneinx nC1 21X nD1rneinx n: Now, these power series may be identified as Maclaurin expansions of ln.1z/, with zDreixorreix. From Eq. (1.97), 1X nD1rncosnx nD1 2Tln.1reix/Cln.1reix/U DlnT.1Cr2/2rcosxU1=2: Setting rD1, we see that 1X nD1cosnx nDln.22 cos x/1=2 Dln 2 sinx 2 ; .0<x<2/: (19.24) Both sides of this expression diverge as x!0and as x!2.4 4Note that the range of validity of Eq. (19.24) may be shifted toT;U(excluding xD0) if we replace xbyjxjon the right-hand side. ArfKen_Ch19-9780123846549.tex 19.1 General Properties 945 Exercises 19.1.1 A function f.x/(quadratically integrable) is to be represented by a finite Fourier series. A convenient measure of the accuracy of the series is given by the integrated square of the deviation, 1pD2Z 0" f.x/a0 2pX nD1.ancosnxCbnsinnx/#2 dx: Show that the requirement that 1pbe minimized, that is, @1p @anD0;@1p @bnD0; for all n, leads to choosing anandbnas given in Eqs. (19.2) and (19.3). Note. Your coefficients anandbnare independent of p. This independence is a conse- quence of orthogonality and would not hold if we expanded f.x/in a power series. 19.1.2 In the analysis of a complex waveform (ocean tides, earthquakes, musical tones, etc.), it might be more convenient to have the Fourier series written as f.x/Da0 2C1X nD1 ncos.nxn/: Show that this is equivalent to Eq. (19.1) with anD ncosn; 2 nDa2 nCb2 n; bnD nsinn;tannDbn=an: Note. The coefficients 2 nas a function of ndefine what is called the power spectrum. The importance of 2 nlies in their invariance under a shift in the phase n. 19.1.3 A function f.x/is expanded in an exponential Fourier series f.x/D1X nD1cneinx: Iff.x/is real, f.x/Df.x/, what restriction is imposed on the coefficients cn? 19.1.4 Assuming thatR Tf.x/U2dxis finite, show that limm!1amD0; limm!1bmD0: Hint. IntegrateTf.x/sn.x/U2, where sn.x/is the nth partial sum, and use Bessel’s inequality (Section 5.1). For our finite interval the assumption that f.x/is square inte- grable (R jf.x/j2dxis finite) implies thatR jf.x/jdxis also finite. The converse does not hold. ArfKen_Ch19-9780123846549.tex 946 Chapter 19 Fourier Series 3π π −π2π 2π− FIGURE 19.5 Reverse sawtooth wave. 19.1.5 Apply the summation technique of this section to show that 1X nD1sinnx nD(1 2.x/; 0<x; 1 2.Cx/;x<0: This is the reverse sawtooth wave shown in Fig. 19.5. 19.1.6 Sum the seriesP1 nD1.1/nC1sinnx nand show that it equals x=2. 19.1.7 Sum the trigonometric seriesP1 nD0sin.2nC1/x 2nC1and show that it equals ( =4; 0<x<; =4;<x<0: 19.1.8 Letf.z/Dln.1Cz/DP1 nD1.1/nC1zn n. This series converges to ln.1Cz/forjzj1, except at the point zD1 . (a) From the real parts show that ln 2 cos 2 D1X nD1.1/nC1cosn n;<<: (b) Using a change of variable, transform part (a) into ln 2 sin 2 D1X nD1cosn n;0<< 2: 19.1.9 (a) Expand f.x/Dxin the interval .0;2L/. Sketch the series you have found (right- hand side of ANS.) over .2L;2L/. ANS. xDL2L 1X nD11 nsinnx L . ArfKen_Ch19-9780123846549.tex 19.1 General Properties 947 (b) Expand f.x/Dxas a sine series in the half interval.0;L/. Sketch the series you have found (right-hand side of Ans.) over .2L;2L/. ANS. xD4L 1X nD01 2nC1sin.2nC1/x L . 19.1.10 In some problems it is convenient to approximate sinxover the intervalT0;1Uby a parabola ax.1x/, where ais a constant. To get a feeling for the accuracy of this approximation, expand 4x.1x/in a Fourier sine series ( 1x1): f.x/D(4x.1x/;0x1 4x.1Cx/;1x0) D1X nD1bnsinnx: ANS. bnD32 31 n3,nodd, bnD0, neven. This approximation is shown in Fig. 19.6. 19.1.11 Verify that.'1'2/D1 2P1 mD1 eim.'1'2/is a Dirac delta function by showing that it satisfies the definition, Z f.'1/1 21X mD1eim.'1'2/d'1Df.'2/: Hint. Represent f.'1/by an exponential Fourier series. 19.1.12 Show that integration of the Fourier expansion of f.x/Dx,< x<, leads to 2 12D1X nD1.1/nC1 n2D11 4C1 91 16C: Note. The series for f.x/Dxwas the subject of Example 19.1.1. Confirm that the change in the defined range from T0;2UtoT;Uhas no effect on the expansion. f(x) x1 −1 FIGURE 19.6 Parabolic approximation to sine wave. ArfKen_Ch19-9780123846549.tex 948 Chapter 19 Fourier Series 19.1.13 (a) Assuming that the Fourier expansion of f.x/is uniformly convergent, show that 1 Z  f.x/2dxDa2 0 2C1X nD1.a2 nCb2 n/: This is Parseval’s identity. Note that it is a completeness relation for the Fourier expansion. (b) Given x2D2 3C41X nD1.1/ncosnx n2;x; apply Parseval’s identity to obtain .4/ in closed form. (c) The condition of uniform convergence is not necessary. Show this by applying the Parseval identity to the square wave f.x/D( 1;< x<0 1; 0<x< D4 1X nD1sin.2n1/x 2n1: 19.1.14 Given '1.x/1X nD1sinnx nD8 >< >:1 2.Cx/;x<0; 1 2.x/; 0<x; show by integrating that '2.x/1X nD1cosnx n2D8 >>< >>:1 4.Cx/22 12;x0; 1 4.x/22 12; 0x: 19.1.15 Given 2s.x/D1X nD1sinnx n2s; 2sC1.x/D1X nD1cosnx n2sC1; develop the following recurrence relations: (a) 2s.x/DxZ 0 2s1.x/dx, (b) 2sC1.x/D.2sC1/xZ 0 2s.x/dx. ArfKen_Ch19-9780123846549.tex 19.2 Applications of Fourier Series 949 Note. The functions s.x/and's.x/of this and the preceding exercise are known as Clausen functions. In theory they may be used to improve the rate of convergence of a Fourier series. As is often the case, there is the question of how much analytical work we do and how much arithmetic work we demand that a computer do. As computers become steadily more powerful, the balance progressively shifts so that we are doing less and demanding that computers do more. 19.1.16 Show that f.x/DP1 nD1cosnx nC1may be written as f.x/D 1.x/'2.x/C1X nD1cosnx n2.nC1/; where 1.x/and'2.x/are the Clausen functions defined in Exercises 19.1.14 and 19.1.15. 19.2 A PPLICATIONS OF FOURIER SERIES We present in this section two typical problems and a short table of useful Fourier series, followed by a substantial number of exercises that illustrate some of the techniques that arise in applications. Example 19.2.1 SQUARE WAVE One application of Fourier series, the analysis of a “square” wave (Fig. 19.7) in terms of its Fourier components, occurs in electronic circuits designed to handle sharply rising pulses. Suppose that our wave is defined by f.x/D0;< x<0; (19.25) f.x/Dh;0<x<: f(x) xh π −3π 3π −2π 2π −π FIGURE 19.7 Square wave. ArfKen_Ch19-9780123846549.tex 950 Chapter 19 Fourier Series From Eqs. (19.2) and (19.3), we find a0D1 Z 0h dtDh; anD1 Z 0hcosnt dtD0; nD1;2;3;:::; bnD1 Z 0hsinnt dtDh n.1cosn/ D8 < :2h n;nodd; 0; neven: The resulting series is f.x/Dh 2C2h sinx 1Csin 3x 3Csin 5x 5C : (19.26) Except for the first term, which represents an average of f.x/over the intervalT;U, all the cosine terms have vanished. Since f.x/h=2is odd, we have a Fourier sine series. Although only the odd terms in the sine series occur, they fall only as n1. This conditional convergence is like that of the alternating harmonic series. Physically this means that our square wave contains a lot of high-frequency components. If the electronic apparatus will not pass these components, our square-wave input will emerge more or less rounded off, perhaps as an amorphous blob.  Example 19.2.2 FULL-WAVE RECTIFIER As a second example, let us ask how well the output of a full-wave rectifier approaches pure direct current. Our rectifier may be thought of as passing the positive peaks of an incoming sine wave and inverting the negative peaks, as shown in Fig. 19.8. This yields f.t/D(sin!t;0<!t<; sin!t;<! t<0:(19.27) Since f.t/as defined here is even, no terms of the form sinn!twill appear. Again, from Eqs. (14.2) and (14.3), we have a0D1 0Z sin!t d.!t/C1 Z 0sin!t d.!t/ D2 Z 0sin!t d.!t/D4 ; ArfKen_Ch19-9780123846549.tex 19.2 Applications of Fourier Series 951 f(t) ωt π 2π −2π −π FIGURE 19.8 Full wave rectifier. anD2 Z 0sin!tcosn!t d.!t/ D8 < :2 2 n21;neven; 0;nodd: Note thatT0;Uis not an orthogonality interval for both sines and cosines together and we do not get zero when nis even. The resulting series is f.t/D2 4 1X nD2;4;6;:::cosn!t n21: (19.28) The original frequency, !;has been eliminated; in fact, all its odd harmonics are also absent. The lowest-frequency oscillation is 2!. The high-frequency components fall off as n2;showing that the full-wave rectifier does a fairly good job of approximating direct current. Whether this good approximation is adequate depends on the particular applica- tion. If the remaining alternating current components are objectionable, they may be further suppressed by appropriate filter circuits.  These examples bring out two features characteristic of Fourier expansions:5 Iff.x/has discontinuities, as in the square wave in Example 19.2.1, we can expect the nth coefficient to be decreasing as O.1=n/. Convergence is conditional only. Iff.x/is continuous (although possibly with discontinuous derivatives as in the full- wave rectifier of Example 19.2.2), we can expect the nth coefficient to be decreasing as1=n2, that is, absolute convergence. We close this section by providing, in Table 19.1, a list of Fourier series that have been introduced either as examples or in the exercises of this chapter. More extensive lists can be found in the Additional Readings, particularly in the work by Oberhettinger, but also in the texts by Carslaw, Churchill, and Zygmund. 5G. Raisbeek, Order of magnitude of Fourier coefficients. Am. Math. Mon. 62: 149 (1955). ArfKen_Ch19-9780123846549.tex 952 Chapter 19 Fourier Series Table 19.1 Some Fourier Series Used in This Text Fourier Series Reference 1.1X nD1sinnx nD( 1 2.Cx/;x<0 1 2.x/;0x<Exercise 19.1.5 Exercise 19.2.8 2.1X nD1.1/nC1sinnx nDx 2;< x<Exercise 19.1.6 Exercise 19.2.7 3.1X nD0sin.2nC1/x 2nC1D( =4;< x<0 C=4; 0<x<Exercise 19.1.7 Eq. (19.26) 4.1X nD1cosnx nDln 2 sinjxj 2 ;< x<Exercise 19.1.8(b) Eq. (19.24) 5.1X nD1.1/ncosnx nDlnh 2 cosx 2i ;< x< Exercise 19.1.8(a) 6.1X nD0cos.2nC1/x 2nC1D1 2ln cotjxj 2 ;< x< Exercise 19.2.5 Exercises 19.2.1 Transform the Fourier expansion of a square wave, Eq. (19.26), into a power series. Show that the coefficients of x1form a divergent series. Repeat for the coefficients ofx3. Note. A power series cannot handle a discontinuity. These infinite coefficients are the result of attempting to beat this basic limitation on power series. 19.2.2 Derive the Fourier series expansion of the Dirac delta function .x/in the interval < x<. (a) What significance can be attached to the constant term? (b) In what region is this representation valid? (c) With the identity NX nD1cosnxDsin.N x=2/ sin.x=2/cos NC1 2x 2 ; show that your Fourier representation of .x/is consistent with Eq. (5.27). 19.2.3 Expand.xt/in a Fourier series. Compare your result with the bilinear form of Eq. (5.27). ArfKen_Ch19-9780123846549.tex 19.2 Applications of Fourier Series 953 ANS..xt/D1 2C1 1X nD1.cosnxcosntCsinnxsinnt/ D1 2C1 1X nD1cosn.xt/. 19.2.4 Show that integrating the Fourier expansion of the Dirac delta function (Exercise 19.2.2) leads to the Fourier representation of the square wave, Eq. (19.26), with hD1. Note. Integrating the constant term .1=2/ leads to a term x=2. What are you going to do with this? 19.2.5 Starting from the Fourier series given as lines 4 and 5 of Table 19.1, show that: 1X nD0cos.2nC1/x 2nC1D1 2ln cotjxj 2 : 19.2.6 Develop the Fourier series representation of f.t/D(0;!t0; sin!t; 0!t: This is the output of a simple half-wave rectifier. It is also an approximation of the solar thermal effect that produces “tides” in the atmosphere. ANS. f.t/D1 C1 2sin!t2 1X nD2;4;6;:::cosn!t n21: 19.2.7 A sawtooth wave is given by f.x/Dx;< x<: Show that f.x/D21X nD1.1/nC1 nsinnx: 19.2.8 A different sawtooth wave is described by f.x/D(1 2.Cx/;x<0 C1 2.x/; 0<x: Show that f.x/D1X nD1.sinnx=n/. 19.2.9 A triangular wave (Fig. 19.4) is represented by f.x/D( x; 0<x< x;< x<0: ArfKen_Ch19-9780123846549.tex 954 Chapter 19 Fourier Series Represent f.x/by a Fourier series. ANS. f.x/D 24 X nD1;3;5;:::cosnx n2. 19.2.10 Expand f.x/D( 1;x2<x2 0 0;x2>x2 0 in the intervalT;U. Note. This variable-width square wave is of some importance in electronic music. 19.2.11 A metal cylindrical tube of radius ais split lengthwise into two nontouching halves. The top half is maintained at a potential CV, the bottom half at a potential V. See Fig. 19.9. Separate the variables in Laplace’s equation and solve for the electrostatic potential for ra. Observe the resemblance between your solution for rDaand the Fourier series for a square wave. 19.2.12 A metal cylinder is placed in a (previously) uniform electric field, E0;with the axis of the cylinder perpendicular to that of the original field. (a) Find the perturbed electrostatic potential. (b) Find the induced surface charge on the cylinder as a function of angular position. 19.2.13 (a) Find the Fourier series representation of f.x/D( 0;< x0 x;0x<: (b) From the Fourier expansion show that 2 8D1C1 32C1 52C: +V −V FIGURE 19.9 Cross section of split tube. ArfKen_Ch19-9780123846549.tex 19.2 Applications of Fourier Series 955 δn(x) n −π π −1 2n1 2nx FIGURE 19.10 Rectangular pulse. 19.2.14 Integrate the Fourier expansion of the unit step function f.x/D( 0;< x<0 1; 0<x<: Show that your integrated series agrees with Exercise 19.2.13. 19.2.15 In the interval .;/ ,n.x/D( n;jxj<1=2n; 0;jxj>1=2n: This wave form is the pulse shown in Fig. 19.10. (a) Expand n.x/as a Fourier cosine series. (b) Show that your Fourier series agrees with a Fourier expansion of .x/in the limit asn!1 . 19.2.16 Confirm the delta function nature of your Fourier series of Exercise 19.2.15 by showing that for any f.x/that is finite in the interval T;Uand continuous at xD0, Z f.x/TFourier expansion of 1.x/UdxDf.0/: 19.2.17 (a) Show that the Dirac delta function .xa/, expanded in a Fourier sine series in the half-interval .0;L/.0<a<L/is given by .xa/D2 L1X nD1sinna L sinnx L : Note that this series actually describes .xCa/C.xa/in the interval .L;L/. (b) By integrating both sides of the preceding equation from 0 to x, show that the cosine expansion of the square wave f.x/D(0;0x<a 1;a<x<L; ArfKen_Ch19-9780123846549.tex 956 Chapter 19 Fourier Series is f.x/D2 1X nD11 nsinna L 2 1X nD11 nsinna L cosnx L ; for0x<L. (c) Show that the term2 1X nD11 nsinna L is the average of f.x/on.0;L/. 19.2.18 Verify the Fourier cosine expansion of the square wave, Exercise 19.2.17(b), by direct calculation of the Fourier coefficients. 19.2.19 (a) A string is clamped at both ends xD0andxDL. Assuming small-amplitude vibrations, we find that the amplitude y.x;t/satisfies the wave equation @2y @x2D1 v2@2y @t2: Herevis the wave velocity. The string is set in vibration by a sharp blow at xDa. Hence we have y.x;0/D0;@y.x;t/ @tDLv0.xa/attD0: The constant Lis included to compensate for the dimensions (inverse length) of .xa/. With.xa/given by Exercise 19.2.17(a), solve the wave equation subject to these initial conditions. ANS. y.x;t/D2v0L v1X nD11 nsinna Lsinnx Lsinnvt L. (b) Show that the transverse velocity of the string @y.x;t/=@tis given by @y.x;t/ @tD2v01X nD1sinna Lsinnx Lcosnvt L: 19.2.20 A string, clamped at xD0and at xDL, is vibrating freely. Its motion is described by the wave equation @2u.x;t/ @[email protected];t/ @x2: Assume a Fourier expansion of the form u.x;t/D1X nD1bn.t/sinnx L and determine the coefficients bn.t/. The initial conditions are u.x;0/Df.x/and@ @tu.x;0/Dg.x/: ArfKen_Ch19-9780123846549.tex 19.3 Gibbs Phenomenon 957 Note. This is only half the conventional Fourier orthogonality integral interval. How- ever, as long as only the sines are included here, the Sturm-Liouville boundary condi- tions are still satisfied and the functions are orthogonal. ANS. bn.t/DAncosnvt LCBnsinnvt L; AnD2 LLZ 0f.x/sinnx Ldx;BnD2 nvLZ 0g.x/sinnx Ldx. 19.2.21 (a) Let us continue the vibrating string problem in Exercise 19.2.20. We assume now that the presence of a resisting medium will damp the vibrations according to the equation @2u.x;t/ @[email protected];t/ @x2[email protected];t/ @t: Introduce a Fourier expansion u.x;t/D1X nD1bn.t/sinnx L and again determine the coefficients bn.t/. Take the initial and boundary condi- tions to be the same as in Exercise 19.2.20. Assume the damping to be small. (b) Repeat, but assume the damping to be large. ANS. (a) bn.t/Dekt=2TAncos!ntCBnsin!ntU; !2 nDnv L k 22 >0; AnD2 LLZ 0f.x/sinnx Ldx;BnD2 !nLLZ 0g.x/sinnx LdxCk 2!nAn: (b) bn.t/Dekt=2TAncoshntCBnsinhntU; 2 nDk 22 nv L2 >0; AnD2 LLZ 0f.x/sinnx Ldx;BnD2 nLLZ 0g.x/sinnx LdxCk 2nAn: 19.3 G IBBS PHENOMENON The Gibbs phenomenon is an overshoot, a peculiarity of the Fourier series and other eigen- function series at a simple discontinuity. An example is seen in Fig. 19.2. Partial Summation of Fourier Series To better understand the Gibbs phenomenon we examine methods for the partial summa- tion of Fourier series. This procedure is unlikely to lead to convenient solutions of practical ArfKen_Ch19-9780123846549.tex 958 Chapter 19 Fourier Series problems for which Fourier series are ideal, but it may provide insight that is needed for our present study. We start from the Fourier series of a function f.x/in exponential form, truncating it to retain terms only for njrjand labeling the truncated expansion fr.x/: fr.x/DrX nDrcneinx;cnD1 2Z f.t/eintdt: Combining these equations in a way useful for the present discussion, we have fr.x/D1 2Z f.t/rX nDrei.xt/dt: (19.29) The summation in Eq. (19.29) is a geometric series. Using a result easily obtained from Eq. (1.96), rX nDrynDyryrC1 1yDyrC1 2y.rC1 2/ y1=2y1=2; we set yDei.xt/, after which we can identify the resulting expression as a quotient of sine functions:6 rX nDrein.xt/Dei.rC1 2/.xt/ei.rC1 2/.xt/ ei.xt/=2ei.xt/=2DsinT.rC1 2/.xt/U sin1 2.xt/: (19.30) Inserting Eq. (19.30) into Eq. (19.29), we reach fr.x/D1 2Z f.t/sinT.rC1 2/.xt/U sin1 2.xt/dt: (19.31) This is convergent at all points, including tDx. Equation (19.31) shows that the quantity 1 2sinT.rC1 2/.xt/U sin1 2.xt/ is in the large- rlimit a Dirac delta distribution. Square Wave For convenience of numerical calculation we consider the behavior of the Fourier series that represents the periodic square wave f.x/D8 >< >:h 2; 0<x<; h 2;< x<0:(19.32) 6This series also occurs in the analysis of a diffraction grating consisting of rslits. ArfKen_Ch19-9780123846549.tex 19.3 Gibbs Phenomenon 959 This is essentially the square wave used in Example 19.2.1, and we immediately see that its Fourier expansion is f.x/D2h sinx 1Csin 3x 3Csin 5x 5C : (19.33) Applying Eq. (19.31) to our square wave, we have fr.x/Dh 4Z 0sinT.rC1 2/.xt/U sin1 2.xt/dth 40Z sinT.rC1 2/.xt/U sin1 2.xt/dt: Making the substitution xtDsin the first integral and xtDs in the second, we obtain fr.x/Dh 4xZ Cxsin.rC1 2/s sin1 2sdsh 4xZ xsin.rC1 2/s sin1 2sds: (19.34) It is important to note that both the integrals in Eq. (19.34) have the same integrand, and therefore have the same indefinite integral, which we denote 8.t/. We may therefore write f.r/Dh 4h 8.x/8.Cx/i h 4h 8. x/8.x/i Dh 4h 8.x/8. x/i h 4h 8.Cx/8.x/i ; (19.35) where the second line of Eq. (19.35) is an obvious rearrangement of the first. However, this second line is useful because it shows that we can also write fr.x/as fr.x/Dh 4xZ xsin.rC1 2/s sin1 2sdsh 4CxZ xsin.rC1 2/s sin1 2sds: (19.36) We are now ready to consider the partial sums in the vicinity of the discontinuity, xD0. For small x, the denominator of the second integrand approaches 1, and the second inte- gral therefore becomes negligible in the limit x!0. On the other hand, the first integrand becomes large near sD0, and the value of the first integral depends on the magnitudes of randx. If we now introduce the new variables pDrC1 2andDps, we have (noting that the integrand is an even function of s) fr.x/h 2pxZ 0sin sin.=2 p/d p: (19.37) Calculation of Overshoot We are now prepared to make a computation of the Fourier series overshoot. From Eq. (19.37), we see that for any finite r,fr.0/will be zero, giving at xD0the average of the two square-wave values ( Ch=2andh=2). However (keeping rfixed), Eq. (19.37) also tells us that fr.x/will increase as pxbecomes nonzero, reaching a maximum when ArfKen_Ch19-9780123846549.tex 960 Chapter 19 Fourier Series pxD. This maximum, which we will shortly show constitutes an overshoot, will there- fore occur at xD=p, which is approximately xD=r. We thus see that the location of the overshoot maximum will differ from xD0in a manner approximately inversely proportional to the number of terms taken in the Fourier expansion. To estimate the maximum value of fr.x/, we substitute pxDinto Eq. (19.37), which we then simplify by making the good approximation sin.=2 p/=2p: fr.xmax/Dh 2Z 0sind psin.=2 p/h Z 0sin d: (19.38) If the upper limit of the final integral of Eq. (19.38) were replaced by infinity, we would have 1Z 0sin dD 2; (19.39) a result found in Example 11.8.5. Note that this replacement would cause fr.x/to have the value h=2, which is the exact value of f.x/forx>0. The integral we would have to add to that of Eq. (19.38) to obtain the infinite range is 1Z sin dDsi./I (19.40) we have identified this integral as the sine integral function si. x/introduced in Table 1.2 and plotted in Fig. 13.6. Thus, Z 0sin dD 2Csi./: (19.41) The graph of si. x/shows that si./> 0, indicating an overshoot. A direct demonstration that our integral is larger than =2 can also be deduced by writing 0 @1Z 03Z 5Z 31 Asin dDZ 0sin d: (19.42) The first integral on the left-hand side has value =2, while each of those to be subtracted is negative (and therefore makes a further positive contribution). A Gaussian quadrature or a power-series expansion and term-by-term integration yields 2 Z 0sin dD1:1789797:::; (19.43) which means that the Fourier series tends to overshoot the positive corner of the square wave by some 18% and to undershoot the negative corner by the same amount. This behav- ior is illustrated in Fig. 19.11. The inclusion of more terms (increasing r) does nothing to ArfKen_Ch19-9780123846549.tex 19.3 Gibbs Phenomenon 961 x 0.04 0.101.2 1.00.8 0.6 0.4 0.2 0.120 terms100 terms 80 0.02 0.06 0.0840 60 FIGURE 19.11 Square wave: Gibbs phenomenon. remove this overshoot but merely moves it closer to the point of discontinuity. The over- shoot is the Gibbs phenomenon, and because of it the Fourier series representation may be highly unreliable for precise numerical work, especially in the vicinity of a discontinuity. The Gibbs phenomenon is not limited to the Fourier series. It occurs with other eigen- function expansions. For more details, see W. J. Thompson, Fourier series and the Gibbs phenomenon, Am. J. Phys. 60: 425 (1992). Exercises 19.3.1 With the partial-sum summation techniques of this section, show that at a discontinuity inf.x/the Fourier series for f.x/takes on the arithmetic mean of the right- and left- hand limits: f.x0/D1 2Tf.x0C0/Cf.x00/U: In evaluating limr!1sr.x0/, you may find it convenient to identify part of the integrand as a Dirac delta function. 19.3.2 Determine the partial sum, sn, of the series in Eq. (19.33) by using (a)sinmx mDxZ 0cosmy dy , (b)nX pD1cos.2 p1/yDsin 2ny 2 sin y. Do you agree with the result given in Eq. (19.40)? 19.3.3 (a) Calculate the value of the Gibbs phenomenon integral ID2 Z 0sint tdt by numerical quadrature accurate to 12 significant figures. ArfKen_Ch19-9780123846549.tex 962 Chapter 19 Fourier Series (b) Check your result by (1) expanding the integrand as a series, (2) integrating term by term, and (3) evaluating the integrated series. This calls for double-precision calculation. ANS. ID1:178979744472: Additional Readings Carslaw, H. S., Introduction to the Theory of Fourier’s Series and Integrals, 2nd ed. London: Macmillan (1921); 3rd ed., paperback, Dover (1952). This is a detailed and classic work; includes considerable discussion of Gibbs phenomenon in chapter IX. Churchill, R. V., Fourier Series and Boundary Value Problems, 5th ed., New York: McGraw-Hill (1993). Discusses uniform convergence in Section 38. Jeffreys, H., and B. S. Jeffreys, Methods of Mathematical Physics, 3rd ed. Cambridge: Cambridge University Press (1972). Termwise integration of Fourier series is treated in section 14.06. Kufner, A., and J. Kadlec, Fourier Series. London: Iliffe (1971). This book is a clear account of Fourier series in the context of Hilbert space. Lanczos, C., Applied Analysis. Englewood Cliffs, NJ: Prentice-Hall (1956), reprinted, Dover (1988). The book gives a well-written presentation of the Lanczos convergence technique (which suppresses the Gibbs phe- nomenon oscillations). This and several other topics are presented from the point of view of a mathematician who wants useful numerical results and not just abstract existence theorems. Oberhettinger, F., Fourier Expansions; A Collection of Formulas. New York: Academic Press (1973). Zygmund, A., Trigonometric Series. Cambridge: Cambridge University Press (1988). The volume contains an extremely complete exposition, including relatively recent results in the realm of pure mathematics. ArfKen_Ch20-9780123846549.tex CHAPTER 20 INTEGRAL TRANSFORMS 20.1 I NTRODUCTION Frequently in mathematical physics we encounter pairs of functions related by an expres- sion of the form g.x/DbZ af.t/K.x;t/dt; (20.1) where it is understood that a,b, and K.x;t/(called the kernel) will be the same for all function pairs fandg. We can write the relationship expressed in Eq. (20.1) in the more symbolic form g.x/DLf.t/; (20.2) thereby emphasizing the fact that Eq. (20.1) can be interpreted as an operator equation. The function g.x/is called the integral transform of f.t/by the operator L, with the specific transform determined by the choice of a,b, and K.x;t/. The operator defined by Eq. (20.1) will be linear: bZ aTf1.t/Cf2.t/UK.x;t/dtDbZ af1.t/K.x;t/dtCbZ af2.t/K.x;t/dt; (20.3) bZ acf.t/K. ;t/dtDcbZ af.t/K. ;t/dt: (20.4) In order for transforms to be useful, we will shortly see that we need to be able to “undo” their effect. From a practical viewpoint, this means that not only must there exist 963 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch20-9780123846549.tex 964 Chapter 20 Integral Transforms an operator L1, but also that we have a reasonably convenient and powerful method of evaluating L1g.x/Df.t/ (20.5) for an acceptably broad range of g.x/. The procedure for inverting a transform takes a wide variety of forms that depend on the specific properties of K.x;t/, so we cannot write a formula that is as general as that for LinEq. (20.1). Not all superficially reasonable choices for the kernel K.x;t/will lead to operators L that have inverses, and even for strategically chosen kernels it may be the case that Land L1will only exist for substantially restricted classes of functions. Thus, the entire devel- opment of the present chapter is restricted (for any given integral transform) to functions for which the indicated operations can be carried out. Before embarking on a study of integral transforms, we may well ask, “Why are integral transforms useful?” Their most common applications are in situations illustrated schemat- ically in Fig. 20.1, where we have a problem that can be solved only with difficulty, if at all, in its original formulation (usually in ordinary space, sometimes called direct or physical space). However, it may happen that the transform of the problem can be solved relatively easily. Our strategy, then, will be to formulate and solve our problem in the trans- form space, after which we transform the solution back to direct space. This strategy often works because the most popular integral transforms are changed in simple ways by differ- entiation and integration operators, with the result that differential and integral equations assume relatively simple forms. This feature will be discussed and illustrated at length later in this chapter. Another frequent use of integral transforms is to use one, together with its inverse, to form an integral representation of a function that we originally had in an explicit form. This move (which appears to be in the direction of generating greater complexity) has value that arises from the relatively simple behavior of the transforms of differentiation and integration operators. Procedures involving integral representations are also presented in later sections of this chapter. Problem in transform spaceSolution in transform space Solution of original problemInverse transform Original problemIntegral transform Difficult solutionRelatively easy solution FIGURE 20.1 Schematic: use of integral transforms. ArfKen_Ch20-9780123846549.tex 20.1 Introduction 965 Some Important Transforms The integral transform that has seen the widest use is the Fourier transform, defined as g.!/D1p 21Z 1f.t/ei!tdt: (20.6) The notation for this transform is not entirely universal; some writers omit the prefactor 1=p 2; we keep it because it causes the transform and its inverse to have formulas that are more symmetrical. In applications involving periodic systems, one occasionally encounters a definition with kernel exp.2 i!t=a0/, where a0is a lattice constant. These differences in notation do not change the mathematics, but cause formulas to differ by powers of 2 ora0. Caution is therefore advised when combining material from different sources. We have defined the Fourier transform in a notation that assigns the symbol !to the transform variable. We did so because, in studying signal processing (an important use of Fourier transforms), the function f.t/usually represents the time behavior of a signal (typically a wave distribution of some kind). Its Fourier transform, g.!/, can then be iden- tified as the corresponding frequency distribution. However, it is worth pointing out that Fourier transforms turn up in contexts far removed from signal-processing problems; they can be used to advantage in evaluating integrals, in alternative formulations of quantum mechanics, and in a wide range of other mathematical procedures. A second transform that has historically been of great importance is the Laplace transform, F.s/D1Z 0etsf.t/dt: (20.7) One of its useful features is the fact that under transformation, differential equations become algebraic equations (as we shall see in detail in Section 20.8). Since algebraic equations are usually easier to solve than differential equations, this feature lends itself to the strategy illustrated in Fig. 20.1. A disadvantage of the Laplace transform is that the for- mula for its inverse is relatively difficult to use. Historically, this difficulty was dealt with by developing tables of Laplace transforms (which can be used to identify inverses). As digital computers have become more powerful, the use of Laplace transforms has declined, but they remain sufficiently useful that we treat them in some detail in the present chapter. Among other transforms that have seen significant use, we mention here two: 1. The Hankel transform, g. /D1Z 0f.t/t Jn. t/dt: (20.8) This transform represents the continuum limit of the Bessel series we studied in Eqs. (14.47) and (14.48). ArfKen_Ch20-9780123846549.tex 966 Chapter 20 Integral Transforms 2. The Mellin transform, g. /D1Z 0f.t/t 1dt: (20.9) We have actually used the Mellin transform without calling it by name; for example, g. /D0. / is the Mellin transform of f.t/Det. Many Mellin transforms are given in a text by Titchmarsh (see Additional Readings). 20.2 F OURIER TRANSFORM We proceed now to a more detailed discussion of the Fourier transform, g.!/D1p 21Z 1f.t/ei!tdt: (20.10) If we rewrite the exponential in Eq. (20.10) in terms of the sine and cosine, and then restrict consideration to functions that are assumed to be either even or odd functions of x, we obtain variants of the original form that are also useful integral transforms: gc.!/Dr 2 1Z 0f.t/cos!t dt; (20.11) gs.!/Dr 2 1Z 0f.t/sin!t dt: (20.12) These formulas define the Fourier cosine andFourier sine transforms. Their kernels, which are real, are natural for use in studies of wave motion and for extracting informa- tion from waves, particularly when phase information is involved. The output of a stellar interferometer, for instance, involves a Fourier transform of the brightness across a stellar disk. The electron distribution in an atom may be obtained from a Fourier transform of the amplitude of scattered x-rays. Example 20.2.1 SOME FOURIER TRANSFORMS 1. f.t/De jtj, with >0. To deal with the absolute value, we break the transform integral into two regions: g.!/Dr 1 20Z 1e tCi!tdtCr 1 21Z 0e tCi!tdt Dr 1 21 Ci!C1 i! Dr 1 22 2C!2: (20.13) ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 967 We note two features of this result: (1) It is real; from the form of the transform, we can see that if f.t/is even, its transform will be real. (2) The more localized is f.t/, the less localized will be g.!/. The transform will have an appreciable value until ! ; larger corresponds to greater localization of f.t/. 2. f.t/D.t/. We easily find g.!/Dr 1 21Z 1.t/ei!tdtDr 1 2: (20.14) This is the ultimately localized f.t/, and we see that g.!/is completely delocalized; it has the same value for all !. 3. f.t/D2 p1=2=. 2Ct2/, with >0. One way to evaluate this transform is by contour integration. It is convenient to start by writing initially g.!/D1 21Z 12 ei!t .ti /.tCi /dt: The integrand has two poles: one at tDi with residue e !=iand one at tDi with residue eC !=.i/. If!> 0, our integrand will become negligible on a large semicircle in the upper half-plane, so an integral over the contour shown in Fig. 20.2(a) will be that needed for g.!/. This contour encloses only the pole at tDi , so we get g.!/D1 2.2i/e ! i.!> 0/: (20.15) However, if ! < 0, we must close the contour in the lower half-plane, as in Fig. 20.2(b), circling the pole at tDi in a clockwise sense (thereby generating a minus sign). This procedure yields g.!/D1 2.2 i/eC ! i.!< 0/: (20.16) If!D0, we cannot perform a contour integration on either of the paths shown in Fig. 20.2, but we then do not need this sophisticated an approach, as we have the iα (a) (b)−iαiα −iα FIGURE 20.2 Contours for third transform in Example 20.2.1. ArfKen_Ch20-9780123846549.tex 968 Chapter 20 Integral Transforms elementary integral g.0/D1 21Z 12 t2C 2dtD1: (20.17) Combining Eqs. (20.15)–(20.17) and simplifying, we have g.!/De j!j: Here we Fourier transformed the transform from our first example, recovering the original untransformed function. This provides an interesting clue as to the form to be expected for the inverse Fourier transform. It is only a clue, because our example involved a transform that was real (i.e., not complex).  An important Fourier transform follows. Example 20.2.2 FOURIER TRANSFORM OF GAUSSIAN The Fourier transform of a Gaussian function eat2, with a>0, g.!/D1p 21Z 1eat2ei!tdt; can be evaluated analytically by completing the square in the exponent, at2Ci!tDa ti! 2a2 !2 4a; which we can check by evaluating the square. Substituting this identity and changing the integration variable from ttosDti!=2a, we obtain (in the limit of large T) g.!/D1p 2e!2=4aTi!=2aZ Ti!=2aeas2ds: (20.18) Thesintegration, shown in Fig. 20.3, is on a path parallel to, but below the real axis by an amount i!=2a. But because connections from that path to the real axis at Tmake negligible contributions to a contour integral and since the contours in Fig. 20.3 enclose no −T T −T−iω/2aT −iω/2aO FIGURE 20.3 Contour for transform of Gaussian in Example 20.2.2. ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 969 singularities, the integral in Eq. (20.18) is equivalent to one along the real axis. Changing the integration limits to 1 and rescaling to the new variable Ds=pa, we reach 1Z 1eas2dtD1pa1Z 1e2dDr a; where we have used Eq. (1.148) to evaluate the error-function integral. Substituting these results we find g.!/D1p 2aexp !2 4a ; (20.19) again a Gaussian, but in !-space. An increase in amakes the original Gaussian eat2 narrower, while making wider its Fourier transform, the behavior of which is dominated by the exponential e!2=4a.  Fourier Integral When we first encountered the delta function, its representation which is the large- nlimit of n.t/D1 2nZ nei!td!; (20.20) was identified as particularly useful in Fourier analysis. We now use that representation to obtain an important result known as the Fourier integral. We write the fairly obvious equation, f.x/Dlimn!11Z 1f.t/n.tx/dt Dlimn!11 21Z 1f.t/2 4nZ nei!.tx/d!3 5dt: (20.21) We now interchange the order of integration and take the limit n!1 , reaching f.x/D1 21Z 1d!1Z 1dt f.t/ei!.tx/: Finally, we rearrange this equation to the form f.x/D1 21Z 1ei!xd!1Z 1f.t/ei!tdt: (20.22) Equation (20.22), the Fourier integral , is an integral representation of f.x/, and will be more obviously recognized as such if the inner integration (over t) is performed, leaving ArfKen_Ch20-9780123846549.tex 970 Chapter 20 Integral Transforms unevaluated that over !. In fact, if we identify the inner integration as (apart from a factorp1=2 ) the Fourier transform of f.t/, and label it g.!/as in Eq. (20.10), then Eq. (20.22) can be rewritten f.t/Dr 1 21Z 1g.!/ei!td!; (20.23) showing that whenever we have the Fourier transform of a function f.t/we can use it to make a Fourier integral representation of that function. The Fourier integral formula, written as in Eq. (20.23), illustrates the value of Fourier analysis in signal processing. If f.t/is an arbitrary signal, Eq. (20.23) describes the signal as composed of a superposition of waves ei!tat angular frequencies1!, with respec- tive amplitudes g.!/. Thus, the Fourier integral is the underlying justification that one can express a signal either by its time dependence f.t/or by its (angular) frequency distribu- tiong.!/. Before leaving the Fourier integral, we should remark that our derivation of it did not provide a rigorous justification for the reversal of the order of integration and the passage to the infinite- nlimit. The interested reader can find a more rigorous treatment in, for example, the work Fourier Transforms by I. N. Sneddon (Additional Readings). Example 20.2.3 FOURIER INTEGRAL REPRESENTATION From the first transform of Example 20.2.1, we found that f.t/De jtjhas Fourier trans- form g.!/Dp1=2 2 =. 2C!2/. If we substitute these data into Eq. (20.23), we obtain e jtjDf.t/D1 21Z 12 ei!t 2C!2d!D 1Z 1ei!t 2C!2d!: (20.24) Equation (20.24) provides an integral representation for exp. jtj/that contains no absolute value signs and may constitute a useful starting point for various analytical mani- pulations. We will shortly encounter some more substantive examples with immediate applications for physics.  Inverse Fourier Transform As the reader may have noticed, Eq. (20.23) is a formula for the inverse Fourier trans- form. Note that the regular (“direct”) and inverse Fourier transforms are given by very similar (but not quite identical) formulas. The only difference is in the sign of the complex exponential. This change of sign causes two successive applications of the Fourier trans- form not to be identical with applying the transform and then its inverse, and the difference shows up when g.!/is not real.2 1The wave ei!thas period 2=! , thus frequency D!=2 . Its angular frequency (radians per unit time rather than cycles) is 2D!. 2Even functions have real Fourier transforms; the transforms of odd functions are imaginary. A function that is neither even nor odd will have a Fourier transform that is complex. ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 971 The analysis of the preceding subsection can also be applied to the Fourier cosine and sine transforms. For convenience, we summarize the formulas for all three varieties of the Fourier transform and their respective inverses. g.!/D1p 21Z 1f.t/ei!tdt; (20.25) f.t/D1p 21Z 1g.!/ei!td!; (20.26) gc.!/Dr 2 1Z 0f.t/cos!t dt; (20.27) fc.t/Dr 2 1Z 0g.!/cos!t d!; (20.28) gs.!/Dr 2 1Z 0f.t/sin!t dt; (20.29) fs.t/Dr 2 1Z 0g.!/sin!t d!: (20.30) Note that the Fourier sine and cosine transforms only use data for 0t<1. Therefore, even though it is possible to evaluate the corresponding inverse transforms for negative t, the results may be irrelevant to the actual situation at those tvalues. But if our function f.t/is even, then the cosine transform will reproduce it faithfully for negative t; odd functions will be properly described for negative tby the sine transform. Example 20.2.4 FINITE WAVE TRAIN An important application of the Fourier transform is the resolution of a finite pulse into sinusoidal waves. Imagine that an infinite wave train sin!0tis clipped by Kerr cell or saturable dye cell shutters so that we have f.t/D8 >>< >>:sin!0t;jtj<N !0; 0;jtj>N !0:(20.31) ArfKen_Ch20-9780123846549.tex 972 Chapter 20 Integral Transforms This corresponds to Ncycles of our original wave train (Fig. 20.4). Since f.t/is odd, we use the Fourier sine transform, Eq. (20.28), to obtain gs.!/Dr 2 N=! 0Z 0sin!0tsin!t dt: (20.32) Integrating, we find our amplitude function: gs.!/Dr 2 sinT.! 0!/.N=! 0/U 2.! 0!/sinT.! 0C!/.N=! 0/U 2.! 0C!/ : (20.33) It is of considerable interest to see how gs.!/depends on frequency. For large !0and !!0;only the first term will be of any importance because of the denominators. It is plotted in Fig. 20.5. This is the amplitude curve for the single-slit diffraction pattern. It has zeros at !0! !0D1! !0D1 N;2 N;and so on: (20.34) For large N,gs.!/may also be interpreted as proportional to a Dirac delta distribution. f(t) t t = 5π ω0 FIGURE 20.4 Finite wave train. ω=ω0ωgs(ω) 1 2πω0ω0Nπ FIGURE 20.5 Fourier transform of finite wave train. ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 973 Since a large fraction of the frequency distribution falls within its central maximum, the half-width of that maximum, 1!D!0 N; (20.35) is a good measure of the spread in angular frequency of our wave pulse. Clearly, if Nis large (a long pulse), the frequency spread will be small. On the other hand, if our pulse is clipped short, Nsmall, the frequency distribution will be wider. The inverse relationship between frequency spread and pulse length is a fundamental property of finite wave distributions; the precision with which a signal can be identified as a specific frequency depends on the pulse length. This same principle finds expression as the Heisenberg uncertainty principle of quantum mechanics, in which position uncertainty (the quantum variable corresponding to pulse length) is inversely related to momentum uncertainty (quantum analog of frequency). It is worth noting that the uncertainty principle in quantum mechanics is a consequence of the wave nature of matter and does not depend on additional ad hoc postulates.  Transforms in 3-D Space Applying the Fourier transform operator in each dimension of a three-dimensional (3-D) space, we obtain the extremely useful formulas g.k/D1 .2/3=2Z f.r/eikrd3r; (20.36) f.r/D1 .2/3=2Z g.k/eikrd3k: (20.37) These integrals are over all space. Verification, if desired, follows immediately by sub- stituting the left-hand side of one equation into the integrand of the other equation and choosing the integration order that permits the complex exponentials to be identified as delta functions in each of the three dimensions. Equation (20.37) may be interpreted as an expansion of a function f.r/in a continuum of plane waves; g.k/then becomes the amplitude of the wave exp. ikr/. Example 20.2.5 SOME 3-D TRANSFORMS 1. Let’s find the Fourier transform of the Yukawa potential, e r=r. Using the notation TUTto denote the Fourier transform of the included object, we seek e r rT .k/D1 .2/3=2Ze r reikrd3r: (20.38) ArfKen_Ch20-9780123846549.tex 974 Chapter 20 Integral Transforms Perhaps the simplest way to proceed is to introduce the spherical wave expansion for exp.ikr/, Eq. (16.61). Equation (20.38), written in spherical polar coordinates, then assumes the form e r rT .k/D4 .2/3=21Z 0r drZ drX lmile rjl.kr/Ym l.k/Ym l.r/:(20.39) All terms of the angular integration vanish except that with lDmD0. For that term, each Y0 0has the constant value 1=p 4, and Eq. (20.39) reduces to e r rT .k/D4 .2/3=21Z 0r e rj0.kr/dr: (20.40) Inserting j0.kr/Dsinkr=kr, the rintegration becomes elementary, and we reach e r rT .k/D1 .2/3=24 k2C 2: (20.41) We wrote Eq. (20.41) as we did to make obvious that if the transform were scaled without the factor 1=.2/3=2, we would have the well-known result 4=.k2C 2/. 2. Even more important than the Fourier transform of the Yukawa potential is that of the Coulomb potential, 1=r. An attempt to evaluate this transform directly leads to convergence problems, but it is easy to evaluate it as the limiting case of the Yukawa potential with D0. Thus, we have the extremely important result, 1 rT .k/D1 .2/3=24 k2: (20.42) 3. From the relation between the Fourier transform and its inverse, Eq. (20.42) can effec- tively be inverted to yield 1 r2T .k/D 21=21 k: (20.43) 4. Another useful Fourier transform is that of the hydrogenic 1sorbital, which (in unnor- malized form) is exp. Zr/. A simple way to evaluate this transform is to differentiate the transform for the Yukawa potential with respect to its parameter, in Eq. (20.41). Noting that differentiation with respect to this parameter commutes with the transform operator (which involves integration with respect to other variables), we have @ @ZeZr rT .k/Dh eZriT .k/D1 .2/3=28Z .k2CZ2/2: (20.44) ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 975 5. Consider next an arbitrary function whose angular dependence is a spherical harmonic (i.e., an angular momentum eigenfunction). Using spherical polar coordinates, we look at h f.r/Ym l.r/iT .k/D1 .2/3=21Z 0f.r/r2drZ drYm l.r/eikr D4 .2/3=21Z 0f.r/r2drZ drYm l.r/ X l0m0il0jl0.kr/Ym0 l0.k/Ym0 l0.r/; where we have inserted the spherical wave expansion, Eq. (16.61), for exp.ikr/. Because the Ym lare orthonormal, the summation reduces to a single term, and we have h f.r/Ym l.r/iT .k/D4il .2/3=2Ym l.k/1Z 0f.r/jl.kr/r2dr: (20.45) Equation (20.45) shows that a function with spherical harmonic angular dependence has a transform containing the same spherical harmonic and that the radial dependence of the transform is essentially a Hankel transform. Compare with Eq. (20.8). 6. As a final example, consider the Fourier transform of a 3-D Gaussian. Again using spherical polar coordinates and the spherical wave expansion (a procedure generally applicable for transforms of spherically symmetric functions), we get h ear2iT .k/D4 .2/3=21Z 0r2ear2j0.kr/dr: (20.46) Using methods similar to those of Example 20.2.2, we find h ear2iT .k/D1 .2a/3=2ek2=4a: (20.47) This result could also be obtained using Cartesian coordinates and using the result of Example 20.2.2 in each of the three dimensions.  Exercises 20.2.1 (a) Show that g.!/Dg.!/is a necessary and sufficient condition for f.x/to be real. (b) Show that g.!/Dg.!/is a necessary and sufficient condition for f.x/to be pure imaginary. Note. The condition of part (a) is used in the development of the dispersion relations of Section 12.8. ArfKen_Ch20-9780123846549.tex 976 Chapter 20 Integral Transforms 20.2.2 The function f.x/D( 1;jxj<1 0;jxj>1 is a symmetrical finite step function. (a) Find gc.!/, Fourier cosine transform of f.x/. (b) Taking the inverse cosine transform, show that f.x/D2 1Z 0sin!cos!x !d!: (c) From part (b) show that 1Z 0sin!cos!x !d!D8 >>>>< >>>>:0;jxj>1;  4;jxjD1;  2;jxj<1: 20.2.3 (a) Show that the Fourier sine and cosine transforms of eatare gs.!/Dr 2 ! !2Ca2;gc.!/Dr 2 a !2Ca2: Hint. Each of the transforms can be related to the other by integration by parts. (b) Show that 1Z 0!sin!x !2Ca2d!D 2eax;x>0; 1Z 0cos!x !2Ca2d!D 2aeax;x>0: These results can also be obtained by contour integration (Exercise 11.8.12). 20.2.4 Find the Fourier transform of the triangular pulse (Fig. 20.6), f.x/D(h.1ajxj/;jxj<1=a; 0;jxj>1=a: Note. This function provides another delta sequence with hDaanda!1 . 20.2.5 Consider the sequence n.x/D( n;jxj<1=2n; 0;jxj>1=2n: ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 977 f(x) h −1/a 1/ax FIGURE 20.6 Triangular pulse. This is Eq. (1.152). Express n.x/as a Fourier integral (via the Fourier integral theorem, inverse transform, etc.). Finally, show that we may write .x/Dlimn!1n.x/D1 21Z 1eikxdk: 20.2.6 Using the sequence n.x/Dnpexp.n2x2/; show that .x/D1 21Z 1eikxdk: Hint. Remember that .x/is defined in terms of its behavior as part of an integrand. 20.2.7 The formula .tx/D1 21Z 1ei!.tx/d!D1 21Z 1ei!tei!xd! can be identified as the continuum limit of an eigenfunction expansion. Derive sine and cosine representations of .tx/that are comparable to the exponential representation just given. ANS.2 1Z 0sin!tsin!x d!,2 1Z 0cos!tcos!x d!. 20.2.8 In a resonant cavity, an electromagnetic oscillation of frequency !0dies out as A.t/D( A0e!0t=2Qei!0t;t>0; 0; t<0: The parameter Qis a measure of the ratio of stored energy to energy loss per cycle. Calculate the frequency distribution of the oscillation, a.!/a.!/, where a.!/is the Fourier transform of A.t/. ArfKen_Ch20-9780123846549.tex 978 Chapter 20 Integral Transforms Note. The larger Qis, the sharper your resonance line will be. ANS. a.!/a.!/DA2 0 21 .!!0/2C.!0=2Q/2: 20.2.9 Prove that Nh 2i1Z 1ei!td! E0i0=2Nh!D8 < :exp 0t 2Nh exp iE0t Nh ;t>0; 0; t<0: This Fourier integral appears in a variety of problems in quantum mechanics: barrier penetration, scattering, time-dependent perturbation theory, and so on. Hint. Try contour integration. 20.2.10 Verify that the following are Fourier integral transforms of one another: (a)8 < :r 2 1p a2x2;jxj<a; 0;jxj>a;9 = ;andJ0.ay/, (b)8 >< >:0;jxj<a; r 2 1p x2Ca2;jxj>a;9 >= >;andY0.ajyj/, (c)r 21p x2Ca2and K0.ajyj/. (d) Can you suggest why I0.ay/is not included in this list? Hint. J0,Y0, and K0may be transformed most easily by using an exponential represen- tation, reversing the order of integration, and employing the Dirac delta function expo- nential representation, Eq. (20.20). These cases can be treated equally well as Fourier cosine transforms. 20.2.11 Show that the following are Fourier transforms of each other: inJn.t/and8 >< >:r 2 Tn.x/.1x2/1=2;jxj<1; 0; jxj>1: Tn.x/is the nth-order Chebyshev polynomial. Hint. With Tn.cos/Dcosn, the transform of Tn.x/.1x2/1=2leads to an integral representation of Jn.t/. ArfKen_Ch20-9780123846549.tex 20.2 Fourier Transform 979 20.2.12 Show that the Fourier exponential transform of f./D(Pn./;jj 1 0;jj>1 is.2in=2/ jn.kr/. Here Pn./is a Legendre polynomial and jn.kr/is a spherical Bessel function. 20.2.13 (a) Show that f.x/Dx1=2is aself-reciprocal under both Fourier cosine and sine transforms; that is, r 2 1Z 0x1=2cosxt dxDt1=2; r 2 1Z 0x1=2sinxt dsDt1=2: (b) Use the preceding results to evaluate the Fresnel integrals 1Z 0cos.y2/dy and1Z 0sin.y2/dy. 20.2.14 Show that1 r2T .k/D 21=21 k: 20.2.15 The Fourier transform formulas for a function of two variables are F.u;v/D1 2ZZ f.x;y/ei.uxCvy/dx dy; f.x;y/D1 2ZZ F.u;v/ei.uxCvy/du dv; where the integrations are over the entire xyoruvplane. For f.x;y/Df.Tx2C y2U1=2/Df.r/, show that the zero-order Hankel transforms F./D1Z 0r f.r/J0.r/dr; f.r/D1Z 0F./J0.r/d; are a special case of the Fourier transforms. Note. This technique may be generalized to derive the Hankel transforms of order D0, 1 2, 1,3 2;:::: See the two texts by Sneddon (Additional Readings). It might also be noted that the Hankel transforms of half-integral orders D1 2reduce to Fourier sine and cosine transforms. ArfKen_Ch20-9780123846549.tex 980 Chapter 20 Integral Transforms 20.2.16 Show that the 3-D Fourier exponential transform of a radially symmetric function may be rewritten as a Fourier sine transform: 1 .2/3=21Z 1f.r/eikrd3xD1 kr 2 1Z 0r f.r/sinkr dr: 20.3 P ROPERTIES OF FOURIER TRANSFORMS Fourier transforms have a number of useful properties, many of which follow directly from the transform definition. Using the 3-D transform as an illustration, and letting g.k/be the Fourier transform of f.r/: h f.rR/iT .k/DeikRg.k/; (translation), (20.48) h f. r/iT .k/D1 3g. 1k/; (change of scale), (20.49) h f.r/iT .k/Dg.k/; (sign change), (20.50) h f.r/iT .k/Dg.k/; (complex conjugation), (20.51) h rf.r/iT .k/Dikg.k/; (gradient), (20.52) h r2f.r/iT .k/Dk2g.k/; (Laplacian). (20.53) The first four of the above formulas can be obtained by carrying out appropriate operations on the defining equation of the transform; details are left to the exercises. Equations (20.52) and(20.53) are easily established from the inverse transform formula. For example, from Eq. (20.37), rf.r/D1 .2/3=2Z g.k/h rreikri dk D1 .2/3=2Z g.k/h .ik/eikri dk D1 .2/3=2Zh ikg.k/i eikrdk; (20.54) showing thatikg.k/is indeed the Fourier transform of rf.r/. It should be noted that this demonstration requires the existence of the integrals involved. The translation formula is of considerable practical value, as it enables a function that is most conveniently described relative to an origin at Rto have a transform whose nat- ural representation is about the origin in the kspace, albeit with a complex phase factor, exp.ikR/. This feature will become important, for example, in problems involving atoms centered at different spatial points, because the transforms of atomic orbitals on such atoms can all be written as centered at a single point in the transform space. Thus, the translation ArfKen_Ch20-9780123846549.tex 20.3 Properties of Fourier Transforms 981 formula can convert a spatially complex problem into a single-center problem (though now with oscillatory character due to the phase factors). The formulas for the gradient and Laplacian, as well as their one-dimensional (1-D) variants, h f0.t/iT .!/Di!g.!/; (first derivative), (20.55) hdn dtnf.t/iT .!/D.i!/ng.!/; (nth derivative), (20.56) make the application of these differential operators have simple forms in the transform space. As we see from Eq. (20.55), the operation of differentiation corresponds in the transform space to multiplication by i!. Example 20.3.1 WAVE EQUATION Fourier transform techniques may be used to advantage in handling partial differential equations (PDEs). To illustrate the technique, let us derive a familiar expression from elementary physics. An infinitely long string is vibrating freely. The amplitude yof the (small) vibrations satisfies the wave equation @2y @x2D1 v2@2y @t2; (20.57) wherevis the phase velocity of the wave propagation. We take as initial conditions y.x;0/Df.x/;@y.x;t/ @t tD0D0; (20.58) where fis assumed localized, meaning that limxD1 f.x/D0. Our method for solving the PDE of Eq. (20.57) will be to take the Fourier transforms (in x) of its two members, using as the transform variable. This is equivalent to multiplying Eq. (20.57) byei xand integrating over x. Before simplifying, we have 1Z [email protected];t/ @x2ei xdxD1 v21Z [email protected];t/ @t2ei xdx: (20.59) If we recognize Y. ;t/D1p 21Z 1y.x;t/ei xdx (20.60) as the transform (from our initial variable xto our transform variable ) of the solution y.x;t/of our PDE, we can rewrite Eq. (20.59) as .i /2Y. ;t/D1 v2@2Y. ;t/ @t2: (20.61) ArfKen_Ch20-9780123846549.tex 982 Chapter 20 Integral Transforms Here we have used Eq. (20.56) for the transform of @2y=@x2and moved the operator @2=@t2, which is irrelevant to the transform operator, outside the integral, leaving behind justY. ;t/. Our original problem has now been converted into Eq. (20.61), but this new equation has the important simplifying feature that the only derivative appearing in it is that with respect to t; we have therefore succeeded in replacing our original PDE (in xandt) with an ordinary differential equation (ODE) (in tonly). The dependence of our problem on (the variable to which xwas converted) is only algebraic. This transformation, from a PDE to an ODE, is a significant achievement. We are now ready to solve Eq. (20.61), subject to the initial conditions, which we need to express in terms of Y. Taking transforms of the quantities in Eq. (20.58), we have 8 >>>>< >>>>:Y. ;0/D1p 21Z 1f.x/ei xdxDF. /; @Y. ;t/ @t tD0D0:(20.62) It is important to recognize that F. /is (in principle) known; it is the Fourier transform of the known initial amplitude f.x/. Solving Eq. (20.61) subject to the initial conditions on Ygiven in Eq. (20.62), we obtain Y. ;t/DF. /ei vtCei vt 2: (20.63) We could have written the tdependence as cos. v t/, but the exponential form is better suited to what we will do next. Since we really want our solution in terms of xrather than , our final step will be to apply inverse Fourier transforms to both sides of Eq. (20.63): 1p 21Z 1Y. ;t/ei xd D1p 21Z 1F. /ei vti xCei vti x 2d : (20.64) The left-hand side of Eq. (20.64) is clearly y.x;t/; each term on the right-hand side is an inverse transform of F(and is therefore f), but the first exponential, if written ei .xvt/, can be seen to lead to an inverse transform of argument xvt, while the second expo- nential leads to an inverse transform of argument xCvt. Thus, our final simplification of Eq. (20.64) takes the form y.x;t/D1 2h f.xvt/Cf.xCvt/i : (20.65) Our solution thus consists of a superposition in which half the amplitude of the original wave form is moving toward Cx(at velocityv) while the other half of the original wave form is moving (also at velocity v) in thexdirection.  ArfKen_Ch20-9780123846549.tex 20.3 Properties of Fourier Transforms 983 Example 20.3.2 HEAT FLOW PDE To illustrate another transformation of a PDE into an ODE, let us Fourier transform the 1-D heat-flow PDE, @ @tDa2@2 @x2; where the solution .x;t/is the temperature at position xand time t. We transform the xdependence, with the transform variable denoted y, writing the transform of .x;t/as9.y;t/, and identifying the transform of @2 .x;t/=@x2as y29.y;t/. Our heat flow equation then takes the form @9.y;t/ @tDa2y29.y;t/; with general solution ln9.y;t/Da2y2tClnC.y/;or9DC.y/ea2y2t: The physical significance of C.y/is that it is the initial spatial distribution of 9or, in other words, the Fourier transform of the initial temperature profile .x;0/. Thus, if we assume the initial temperature distribution is known, then so also is C.y/, and our PDE solution, the inverse transform of 9, assumes the form .x;t/D1 21Z 1C.y/ea2y2teiyxdy: (20.66) Further progress depends on the specific form of C.y/. If we assume the initial tem- perature to be a delta-function spike at xD0, corresponding to an instantaneous pulse of thermal energy at xDtD0, we then have as its Fourier transform C.y/Dconstant, see Eq. (20.14). We can now evaluate the integral in Eq. (20.66) to obtain an explicit form for .x;t/. With Cconstant, the functional form of Eq. (20.66) is (apart from the sign of i) just that encountered in Example 20.2.2 for the Fourier transform of a Gaussian, and we can evaluate the integral to obtain .x;t/DC ap 2texp x2 4a2t : This form for was obtained in Section 9.7, but it arose there as a clever guess that was ultimately justified because it led to a solution of the diffusion PDE.  Example 20.3.3 COULOMB GREEN’S FUNCTION The Green’s function associated with the Poisson equation satisfies the PDE r2 rG.r;r0/D.rr0/: (20.67) We take the Fourier transform of both sides of this equation with respect to r, desig- nating g.k;r0/as the transform of G. Note that r0is unaffected by the transformation. ArfKen_Ch20-9780123846549.tex 984 Chapter 20 Integral Transforms Using Eq. (20.53), the left-hand side of Eq. (20.67) becomes k2g.k;r0/, while the right- hand side, in which the delta function has been translated an amount r0, has according to Eq. (20.48) the transform eikr0T.k/. Thus, Eq. (20.67) transforms into k2g.k;r0/D1 .2/3=2eikr0; where the transform of the delta function has been evaluated as the 3-D equivalent of Eq. (20.14). We may now solve for g: g.k;r0/D1 .2/3=2eikr0 k2; and recover Gby taking the inverse transform, G.r;r0/D1 .2/3Zeikr0 k2eikrd3kD1 .2/3Zd3k k2eik.rr0/: We see that the evaluation is proportional to that of the inverse transform of 1=k2, but for argument rr0. Using Eq. (20.43) (which applies also for the inverse transform because it is real), we reach G.r;r0/D1 .2/3=2 21=2 1 jrr0jD1 41 jrr0j; a result we have previously obtained by other methods (cf. Section 10.2). Note that we did not assume Gto be a function of rr0; we found it to have that form.  Successes and Limitations Some of the above examples illustrate an important role played by the Fourier transform: Use of the Fourier transform can convert a PDE into an ODE, thereby reducing the “degree of transcendence” of the problem. All the examples also illustrate the procedure sketched schematically in Fig. 20.1: Fourier transformation can often convert a difficult problem into one which we are able to solve. A useful form for our solution can then be obtained by transforming it back to physical space. Despite these successes, it is worth noting that not all problems posed as differential equations are amenable to Fourier-transform solution methods. Some of the limitations arise from the implicit requirement that the necessary transforms and their inverses exist. We can also expect Fourier methods to work only when the solution is unique, as the process of taking a transform and then solving an algebraic equation produces a single result, and not a set of two or more linearly independent solutions. Usually the boundary conditions are the proximate reason that a differential equa- tion solution is unique, and the requirement that an (exponential) Fourier transform exist ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 985 imposes Dirichlet boundary conditions at infinity. For 1-D systems on the semi-infinite range 0x<1, use of the Fourier sine transform imposes a Dirichlet condition at the finite boundary xD0, while use of the cosine transform corresponds to a Neumann bound- ary condition there. Additional opportunities for solving differential equations by transform methods are provided by use of the Laplace transform, for which it is more natural to introduce bound- ary data. See the later sections of this chapter. Exercises 20.3.1 Write the 1-D equivalents of the equations for translation, scale change, sign change, and complex conjugation that were given for 3-D transforms in Eqs. (20.48) to (20.51). 20.3.2 (a) Show that by replacement of rbyrRin the formula for the Fourier transform off.r/, one can derive the translation formula, Eq. (20.48). (b) Using methods similar to those for part (a), establish the formulas for scale change, sign change, and complex conjugation, Eqs. (20.49) to (20.51). 20.3.3 Derive Eq. (20.53), the formula for the Fourier transform of r2f.r/. 20.3.4 Verify Eqs. (20.55) and(20.56), the formulas for the derivatives of 1-D Fourier trans- forms. 20.3.5 Derive the inverse of Eq. (20.56), namely that  tnf.t/T.!/Dindn d!ng.!/: 20.3.6 The 1-D neutron diffusion equation with a (plane) source is Dd2'.x/ dx2CK2D'.x/DQ.x/; where'.x/is the neutron flux, Q.x/is the (plane) source at xD0, and DandK2are constants. Apply a Fourier transform. Solve the equation in transform space. Transform your solution back into x-space. ANS.'.x/DQ 2K DejK xj: 20.4 F OURIER CONVOLUTION THEOREM An important relationship satisfied by Fourier transforms is that known as the convolu- tion theorem. As we shall soon see, this theorem is useful in the solution of differential equations, in establishing the normalization of momentum wave functions, in the evalua- tion of integrals arising in many branches of physics, and in a variety of signal-processing applications. ArfKen_Ch20-9780123846549.tex 986 Chapter 20 Integral Transforms We define the convolution of two functions f.x/andg.x/, understood here to be over the interval.1;1/, as the following operation designated fg: .fg/.x/1p 21Z 1g.y/f.xy/dy: (20.68) The corresponding definition in three dimensions is .fg/.r/1 .2/3=2Z g.r0/f.rr0/d3r0; (20.69) where the integral is over the full 3-D space. This operation is sometimes referred to as Faltung, the German term for “folding.” To better understand the origin of this name, look at Fig. 20.7, where we have plotted f.y/Deyandf.xy/De.xy/. Clearly, f.y/andf.xy/are related by reflection relative to the vertical line yDx=2; that is, we could generate f.xy/by folding over f.y/on the line yDx=2. Our interest here is not primarily in the nomenclature, but rather to understand what hap- pens if we take the Fourier transform of a convolution. Letting F.t/andG.t/, respectively, be the Fourier transforms of fandg, we find .fg/T.t/D1p 21Z 1dx2 41p 21Z 1dy g.y/f.xy/3 5eitx D2 41p 21Z 1dy g.y/eity3 52 41p 21Z 1dx f.xy/eit.xy/3 5 D2 41p 21Z 1dy g.y/eity3 52 41p 21Z 1dz f.z/eitz3 5 DG.t/F.t/: (20.70) e−ye−(x−y) xy FIGURE 20.7 Factors in a Faltung. ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 987 In the second line of the above equation set we simply divided eitxinto the two factors eity andeit.xy/; the third line was reached by changing the integration variable of the second integral from xtozDxy. After this change, yonly appears in the first set of square brackets and zonly appears in the second bracket set. We are then able to continue to the fourth line where we identify the integrals as Fourier transforms. We often encounter integrals that have the form of a convolution fg. The convolution theorem then enables the construction of the Fourier transform of the integral, and the integral itself will then be given by taking the inverse transform of .fg/T:This process corresponds to 1Z 1g.y/f.xy/dyDp 2.fg/.x/Dp 21p 21Z 1.fg/T.t/eixtdt D1Z 1G.t/F.t/eixtdt: (20.71) Once again we see an appealing feature inherent to Fourier analysis. While the two func- tions in our original integral, g.y/andf.xy/, had different arguments, their transforms, G.t/andF.t/, have the same argument. We still have an integral to evaluate after using the convolution theorem, but (as just observed) the integrand consists of a product of quan- tities both of which are evaluated at the same point. The cost of the transformation is the presence of a complex exponential, which imparts oscillatory character to the integral. We have thus traded geometric complexity for oscillational complexity. Often this will be an advantageous trade-off. For the record, here is the 3-D equivalent of Eq. (20.71): Z g.r0/f.rr0/d3r0DZ F.k/G.k/eikrd3k: (20.72) Parseval Relation If we specialize Eq. (20.71) to xD0, we get the relatively simple result 1Z 1f.y/g.y/dyD1Z 1F.t/G.t/dt: (20.73) This equation becomes more easily interpreted if we change f.y/tof.y/. Then we must replace f.y/in Eq. (20.73) by f.y/, while F.t/becomesTf.y/UT, which, invoking Eq. (20.51), can be written F.t/. With these changes, we have 1Z 1f.y/g.y/dyD1Z 1F.t/G.t/dt: (20.74) This equation is known as the Parseval relation; some authors prefer to call it Rayleigh’s theorem. ArfKen_Ch20-9780123846549.tex 988 Chapter 20 Integral Transforms The integrals in Eq. (20.74) are of the form of scalar products, and will exist if fandg (and therefore also FandG) are quadratically integrable (i.e., members of an L2space). Letting Fdenote the Fourier transform operator, we can rewrite Eq. (20.74) in the compact form hfjgiDhF fjFgi: (20.75) If we now move the Fout of the left half-bracket, writing instead its adjoint in the right half-bracket, we reach hfjgiDh fjF†Fgi: (20.76) Since this equation must hold for all fandgin our Hilbert space, it is necessary that F†F reduce to the identity operator, meaning that F†DF1: (20.77) Our conclusion is that the Fourier transform operator is unitary. If, next, we consider the special case gDf,Eq. (20.75) takes the form hfjfiDh FjFi; (20.78) showing that fand its transform, F, have the same norm, a result that is hardly surprising since we already know that transforming ftwice brings us back to at worst fmultiplied by a complex phase factor. An interesting consequence of the unitarity property is illustrated by the formulas gov- erning Fraunhofer diffraction optics. The amplitude of the diffraction pattern appears as the Fourier transform of the function describing the aperture (compare Exercise 20.4.3). With intensity proportional to the square of the amplitude, the Parseval relation implies that the energy passing through the aperture (the integral of jfj2) is equal to that in the diffraction pattern, whose total energy is the integral of jFj2. In this problem the Parseval relation corresponds to energy conservation. We close this topic with two observations. First, note how the clarity and simplicity of our discussion of the Parseval relation was greatly enhanced by introducing appropriate notation. Much of our insight and intuition regarding mathematical concepts flows directly from the use of good notations for their description. Secondly, we call attention to the fact that Parseval’s relation can be developed independently of the inverse Fourier transform and then used rigorously to derive the inverse transform. Details can be found in the text by Morse and Feshbach (Additional Readings). Here are some examples illustrating use of the convolution theorem. Example 20.4.1 POTENTIAL OF CHARGE DISTRIBUTION We require the potential at all points rproduced by a charge distribution .r0/. From Coulomb’s law, or equivalently from the Green’s function for Poisson’s equation, we have .r/D1 4Z.r0/ jrr0jd3r: (20.79) ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 989 The integral for is of the convolution form, and its presence in this problem suggests that convolutions will arise in a wide variety of problems in which there is a distributed source of almost any kind and an effect therefrom that depends on relative position. Taking f.r/D1=r, so that f.rr0/D1=jrr0j, and g.r/D.r/, application of the convolution formula Eq. (20.72) yields .r/D1 4Z fT.k/gT.k/eikrd3k: Since fT.k/D1 .2/3=24 k2and gT.k/DT.k/; we have .r/D1 .2/3=2ZT.k/ k2eikrd3k: (20.80) Depending on the functional form of , Eq. (20.80) may or may not be easier to evaluate than the original equation for ,Eq. (20.79).  Example 20.4.2 TWO-CENTER OVERLAP INTEGRAL In quantum mechanics problems involving molecules, one often encounters the so-called overlap integral, which is the scalar product of two atomic orbitals, one, 'a, centered at a point A, and another, 'b, centered at a different point B. This overlap integral, denoted Sab, can be written SabDZ ' a.rA/' b.rB/d3r: (20.81) The integral is over the full 3-D space. One way to evaluate Sabstarts by changing to coordinates in which the origin is at A; this amounts to the substitution r0DrA, in terms of which rBDr0.BA/, so SabDZ ' a.r0/'b.r0R/d3r0; where RDBA. We note the physically expected feature that the value of Sabdoes not depend on AandBseparately but only on the vector Rdescribing their relative position. This integral for Sabis almost in the standard form for a convolution (it differs there- from by having r0Rinstead of Rr0). This discrepancy can be handled by invoking Eq. (20.50); the net effect is to change the sign of the transform variable kwhen we eval- uate'T b: Again using Eq. (20.72), we write SabDZh ' aiT .k/'T b.k/eikRd3k: ArfKen_Ch20-9780123846549.tex 990 Chapter 20 Integral Transforms We continue with the specific case that 'aand'bare Slater-type orbitals (STOs), both with the same screening parameter . These STOs, and their Fourier transforms (which can be obtained by differentiating Eq. (20.41) with respect to its parameter ), are 'D'Der; 'TD1 .2/3=28 .k2C2/2: Inserting the formula for 'Tinto the integral for Sab, we get SabD.8/2 .2/3ZeikR .k2C2/4d3k: At this point we already see an advantage of the convolution-based procedure. This integral (whether or not we can easily evaluate it) has assumed a single-center character, with the interorbital spacing relegated to the complex exponential factor. To complete the evaluation, we now insert the spherical wave expansion for exp. ikR/, Eq. (16.61), and we note the further simplification that the only term sur- viving the integration over the angular coordinates of kis the lD0term of the expansion. Keeping in mind that Y0 0D1=p 4, that term is seen to be just j0.k R/, so our formula for Sabbecomes SabD.8/2 .2/31Z 0j0.k R/ .k2C2/44k2dk: We now have a known 1-D integral, which in fact we encountered in Exercise 14.7.10: kn.x/D2nC2.nC1/W xnC11Z 0k2j0.kx/ .k2C1/nC2dk: Changing xin this formula to Rand replacing kbyk=, we reach SabDR3 3k2.R/DeR 33 2R2C3RC3 : (20.82) Note that when we insert the explicit form for k2, we obtain a relatively simple final result. There are other ways to obtain this formula (one of which is to use prolate ellip- soidal coordinates with AandBas foci), but the method we have chosen here provides a good illustration of the issues and formulas that arise when the convolution method is applicable.  Multiple Convolutions Some important problems take the form of multiple convolutions, which we illustrate in one dimension by the convolution of a function hfirst with a function gfollowed by the ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 991 convolution of that result with f, i.e., f.gh/. Thus, h f.gh/i .x/D1p 21Z 1dy f.y/.gh/.xy/ D1 21Z 1dy1Z 1dt f.y/g.t/h.xyt/; which after making the substitution tDzy(and therefore xytDxz) becomes h f.gh/i .x/D1 21Z 1dy1Z 1dz f.y/g.zy/h.xz/: (20.83) Letting F,G, and Hbe the Fourier transforms of f,g, and h, this case of the convolution theorem is h fghiT .!/DF.!/G.!/H.!/: (20.84) We have now omitted the parentheses surrounding ghsince we would have gotten the same result if we convoluted f,g, and hin any order. Then, taking the inverse transform, we have 1Z 1dy1Z 1dz f.y/g.zy/h.xz/D.2/1=21Z 1F.!/G.!/H.!/ei!xd!:(20.85) In three dimensions, the corresponding formulas are h f.gh/i .r/D1 .2/3Z d3r0Z d3r00f.r0/g.r00r0/h.rr00/; (20.86) h f.gh/iT .k/DF.k/G.k/H.k/; (20.87) Z d3r0Z d3r00f.r0/g.r00r0/h.rr00/D.2/3=2Z F.k/G.k/H.k/eikrd3k: (20.88) Example 20.4.3 INTERACTION OF TWO CHARGE DISTRIBUTIONS The electrostatic interaction of two charge distributions 1.r/and2.r/is given by the integral VDZ d3r0Z d3r001.r0/2.r00/ jr00r0j; (20.89) which is a double convolution, as in Eq. (20.88), but with the free argument rset to zero and with a sign discrepancy in the argument of h(which is2of the present example). ArfKen_Ch20-9780123846549.tex 992 Chapter 20 Integral Transforms Taking the above into account and applying Eq. (20.88), we have VD.2/3=2Z d3kT 1.k/1 rT .k/T 2.k/ D4Zd3k k2T 1.k/T 2.k/; (20.90) where we have inserted the value of .1=r/Tfrom Eq. (20.42). This expression has the obvious advantage that it is a 3-D integral in place of the original six-fold integration in Eq. (20.89). The price we have to pay for this simplification is the cost of taking the Fourier transforms of 1and2.  Transform of a Product The similarity between the formulas for the direct and inverse Fourier transforms sug- gest that we may be able to identify the Fourier transform of a product as a convolution. Accordingly, we rewrite Eq. (20.71) with xreplaced byxand also change the variable of integration in that equation from ytoy. We then have (multiplying the equation by 1=p 2) 1p 21Z 1g.y/f.yx/dyD1p 21Z 1G.t/F.t/eixtdtDh G.t/F.t/iT .x/:(20.91) If we now make the further identifications h G.t/iT .y/Dg.y/andh F.t/iT .xy/Df.yx/; we have .FTGT/.x/Dh G.t/F.t/iT .x/: (20.92) Rewriting Eq. (20.92) with the functions renamed fandgand their respective transforms denoted FandG, we have our desired final result: h f giT DFG: (20.93) Equation (20.93) will be useful only if fandgindividually have Fourier transforms. It is possible that this condition is not satisfied despite the fact that f gpossesses a trans- form. We therefore proceed to consider the case that fnot have a transform, but instead possesses a Maclaurin expansion, and therefore can be represented by a series in positive integer powers of x. Then, starting from the relation h xng.x/iT .t/Dindn dtnG.t/; the topic of Exercise 20.3.5, we can write h f giT .t/Df id dt G.t/; (20.94) ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 993 where the expression i.d=dt/is the argument off(and not a multiplicative factor). Unless fis quite simple, this expression may be of limited practical value. Momentum Space Hamilton’s equations of classical mechanics formalize a symmetry between position vari- ables qand the corresponding (conjugate) momentum variables p. This same corre- spondence carries over into quantum mechanics, where (in one dimension, in units with NhD1), the fundamental relationship is the commutator Tx;pUDi. The time-independent Schrödinger equation (for a particle of mass m) is H h1 2mp2CV.x/i DE ; and it is usually made more explicit by taking pDi.d=dx/, in which case the wave function is a function of x: D .x/. In principle we could have chosen pas the fundamental variable, in which case the proper value of the commutator is recovered if we take xDCi.d=dp/, and (which we will now give the name ') will be a function ofp:'D'.p/. These two representations of the Schrödinger equation in one dimension correspond, respectively, to the two ODEs: 1 2md2 dx2 .x/CV.x/ .x/DE .x/; (20.95) p2 2m'.p/CV id dp '.p/DE'.p/: (20.96) Note that in the second of these two equations, the argument of Vis a differential operator, and unless the form of Vis relatively simple, the momentum-space ODE will be quite complicated and correspondingly difficult to solve. In the coordinate representation .x;id=dx/, a wave function exp.ikx/is an eigen- function of momentum with eigenvalue k: p eikxDid dxeikxDi.ik/eikxDk eikx; and this fact suggests that momentum wave functions will be Fourier transforms of their coordinate counterparts. We therefore seek to verify the consistency of Eqs. (20.95) and (20.96) by Fourier transforming the first of these two equations, letting g.t/represent the transform of and using Eq. (20.56) to take the transform of the second derivative.3In the case that Vhas a Maclaurin expansion, we then use Eq. (20.94), obtaining t2 2mg.t/CV id dt g.t/DEg.t/: This equation can be brought into agreement with Eq. (20.96) if we take its complex con- jugate (assuming Vto be real), so we can make the identification '.p/ ! g.t/. 3Here tis the transform variable; in the present context it has nothing to do with time. ArfKen_Ch20-9780123846549.tex 994 Chapter 20 Integral Transforms On the other hand, if Vhas a transform we can use the convolution formula, Eq. (20.93), thereby converting Eq. (20.95) into an integral equation: p2 2m'.p/C1p 21Z 1VT.pp0/'.p0/dp0DE'.p/: (20.97) Example 20.4.4 MOMENTUM-SPACE SCHRÖDINGER EQUATION The time-independent Schrödinger equation for the hydrogen atom has (in hartree atomic unitsNhDmDeD1) the coordinate representation 1 2r2 .r/1 r .r/DE .r/: Taking the Fourier transform of this equation, we get for the momentum-space wave function'.k/ k2 2'.k/1 .2/3Z4 jkk0j2'.k0/d3k0DE'.k/: (20.98) In reaching Eq. (20.98), we have used the 3-D version of Eq. (20.97), inserting for the transform of Vthe result from Eq. (20.42). In principle one can solve Eq. (20.98) for'.k/ and the corresponding eigenvalues E, and the results should be equivalent to the original equation. That is a more difficult task than we will undertake now, but it is straightforward to verify that the Fourier transform of the known solution for the hydrogen ground state is a solution to Eq. (20.98). From Eq. (20.44), the hydrogen 1swave function eris seen to have Fourier transform '.k/DC .k2C1/2; where Cis independent of kand has a value that is irrelevant here. Inserting this result into Eq. (20.98), we find 1 2Ck2 .k2C1/2C 22Zd3k0 jkk0j2.k02C1/2DEC .k2C1/2: (20.99) Writingjkk0j2Dk2C2kk0cosCk02, the integral, though a bit tedious, is found to be elementary. Inserting its value, Eq. (20.99) becomes (canceling the common factor C), 1 2k2 .k2C1/21 21 k2C1DE1 .k2C1/2: This equation is satisfied if ED1=2 , the correct energy (in hartree atomic units) for the hydrogen 1sstate.  ArfKen_Ch20-9780123846549.tex 20.4 Fourier Convolution Theorem 995 Exercises 20.4.1 Work out the convolution equation corresponding to Eq. (20.71) for (a) Fourier sine transforms 1 21Z 0g.y/h f.yCx/Cf.yx/i dyD1Z 0Fs.s/Gs.s/cossx ds; where fandgare odd functions. (b) Fourier cosine transforms 1 21Z 0g.y/h f.yCx/Cf.xy/i dyD1Z 0Fc.s/Gc.s/cossx ds; where fandgare even functions. 20.4.2 Show that for both Fourier sine and Fourier cosine transforms Parseval’s relation has the form 1Z 0F.t/G.t/dtD1Z 0f.y/g.y/dy: 20.4.3 (a) A rectangular pulse is described by f.x/D( 1;jxj<a; 0;jxj>a: Show that the Fourier exponential transform is F.t/Dr 2 sinat t: This is the single-slit diffraction problem of physical optics. The slit is described byf.x/. The diffraction pattern amplitude is given by the Fourier transform F.t/. (b) Use the Parseval relation to evaluate 1Z 1sin2t t2dt: This integral may also be evaluated by using the calculus of residues (Exercise 11.8.9). ANS..b/: ArfKen_Ch20-9780123846549.tex 996 Chapter 20 Integral Transforms 20.4.4 Solve Poisson’s equation, r2 .r/D.r/=" 0;by the following sequence of opera- tions: (a) Take the Fourier transform of both sides of this equation. Solve for the Fourier transform of .r/ . (b) Carry out the Fourier inverse transform. 20.4.5 (a) Given f.x/D1jx=2jfor2x2, with f.x/D0elsewhere, show that the Fourier transform of f.x/is F.t/Dr 2 sint t2 : (b) Using the Parseval relation, evaluate 1Z 1sint t4 dt: ANS..b/2 3: 20.4.6 With F.t/andG.t/the Fourier transforms of f.x/andg.x/, respectively, show that 1Z 1 f.x/g.x/ 2 dxD1Z 1 F.t/G.t/ 2 dt: Ifg.x/is an approximation to f.x/, the preceding relation indicates that the mean square deviation in t-space is equal to the mean square deviation in x-space. 20.4.7 Use the Parseval relation to evaluate .a/1Z 1d! .!2Ca2/2; .b/1Z 1!2d! .!2Ca2/2: Hint. Compare Exercise 20.2.3. ANS. (a) 2a3, (b) 2a: 20.4.8 The nuclear form factor F.k/and the charge distribution .r/ are 3-D Fourier trans- forms of each other: F.k/D1 .2/3=2Z .r/eikrd3r: If the measured form factor is F.k/D.2/3=2 1Ck2 a21 ; find the corresponding charge distribution. ANS..r/Da2 4ear r: ArfKen_Ch20-9780123846549.tex 20.5 Signal-Processing Applications 997 20.4.9 Using convolution methods, find an integral whose value is the electrostatic interaction energy between a charge distribution .rA/and a unit point charge at C. 20.4.10 With .r/ a wave function in ordinary space and '.p/ the corresponding momentum function, show that (a)1 .2Nh/3=2Z r .r/eirp=Nhd3rDiNhrp'.p/; (b)1 .2Nh/3=2Z r2 .r/erp=Nhd3rD.iNhrp/2'.p/: Note. rpis the gradient in momentum space: Oex@ @pxCOey@ @pyCOez@ @pz: These results may be extended to any positive integer power of rand therefore to any (analytic) function that may be expanded as a Maclaurin series in r. 20.4.11 The ordinary space wave function .r; t/satisfies the time-dependent Schrödinger equation, iNh@ .r; t/ @tDNh2 2mr2 CV.r/ : Show that the corresponding time-dependent momentum wave function satisfies the analogous equation iNh@'.p; t/ @tDp2 2m'CV.iNhrp/': Note. Assume that V.r/may be expressed by a Maclaurin series and use Exer- cise 20.4.10. V.iNhrp/is the same function of the variable iNhrpthatV.r/is of the variable r. 20.5 S IGNAL -PROCESSING APPLICATIONS A time-dependent electrical pulse f.t/may be regarded as a superposition of waves of many frequencies. For angular frequency !, we have a contribution F.!/ei!t: Then the complete pulse may be written as f.t/D1 21Z 1F.!/ei!td!: (20.100) Because the angular frequency !is related to the linear frequency by D! 2; ArfKen_Ch20-9780123846549.tex 998 Chapter 20 Integral Transforms most physicists associate the entire 1=2 factor with this integral, so this formula differs by a factor.2/1=2from the definition we have adopted for the Fourier transform. But if!is a frequency, what about the negative frequencies? The negative !may be looked on as a mathematical device to avoid dealing with two functions ( cos!tandsin!t) separately. Because Eq. (20.100) has the form of a Fourier transform, we may solve for F.!/by taking the inverse transform. Keeping in mind the scale at which we wrote Eq. (20.100), we get F.!/D1Z 1f.t/ei!tdt: (20.101) Equation (20.101) represents a resolution of the pulse f.t/into its angular frequency components. Equation (20.100) is a synthesis of the pulse from its components. Now consider some device, such as a servomechanism or a stereo amplifier, with an input f.t/and an output g.t/. For an input of a single frequency f!with input f!.t/D F.!/ei!t, the device will alter the amplitude and may also change the phase. For the situations we discuss here, we assume a linear response, which means that we are assuming thatg!(the output corresponding to f!) will be a signal at the same frequency as f!, will scale linearly with f!, and be independent of the simultaneous presence of signals at other frequencies. However, the responses of interesting devices will depend on the frequency. Hence, our assumption is that g!andf!are related by an equation of the form g!.t/D'.!/ f!.t/: (20.102) This amplitude- and phase-modifying function, '.!/ , is called a transfer function. When making schematic diagrams of electronic circuits, it is customary to designate a device characterized by a transfer function by a suitably labeled box with input and output con- ductors, as shown in (Fig. 20.8). Because we have assumed the operation corresponding to the transfer function to be linear, the total output from a pulse containing many frequencies may be obtained by inte- grating over the entire input, as modified by the transfer function, g.t/D1 21Z 1'.!/ F.!/ei!td!: (20.103) The transfer function is characteristic of the device to which it applies. Once it is known (either by calculation or measurement), the output g.t/can be calculated for any input f.t/. ϕ(ω) Input Outputf(t) g(t) FIGURE 20.8 Schematic for device described by transfer function. ArfKen_Ch20-9780123846549.tex 20.5 Signal-Processing Applications 999 Equation (20.103) can be brought to a convenient form if we recognize that it is simply the formula for the Fourier transform of the product '.!/ F.!/. We already know that F.!/has transform f.t/. Letting8.t/be (at the scaling of this section) the transform of'.!/ , we may then use Eq. (20.93) to rewrite Eq. (20.103) as the convolution of the transforms fand8: g.t/D1Z 1f.t0/8.tt0/dt0: (20.104) Interpreting Eq. (20.104), we have an input (a “cause”), namely f.t0/, modified by 8.tt0/, producing an output (an “effect”), namely g.t/. Adopting the concept of causal- ity(that the cause precedes the effect), we must obtain contributions to g.t/only from times t0such that t0<t. We do this by requiring 8.tt0/D0; t0>t: (20.105) Then Eq. (20.104) becomes g.t/DtZ 1f.t0/8.tt0/dt0: (20.106) Since Eq. (20.106) must yield real output g.t/for arbitrary real input f.t/, we see that in addition to the requirement in Eq. (20.105), we also know that 8.t/must be real. The adoption of Eq. (20.106) and the reality of 8have profound consequences here and equivalently in dispersion theory (Section 12.8). Example 20.5.1 TRANSFER FUNCTION: HIGH-PASS FILTER Ahigh-pass filter permits almost complete transmission of high-frequency electrical sig- nals but strongly attenuates those at lower frequencies. A very simple high-pass filter is shown in Fig. 20.9. Its transfer function describes the steady-state behavior of the filter in the absence of loading (meaning that the output terminals are not connected to anything), so we can assume that, for a signal at frequency !, the input, output, and current are the real parts of the respective quantities Vinei!t,Voutei!t,I ei!t. Possible phase differences in these quantities are allowed for by permitting Vin,Vout, and Ito be complex. I IRC Vin VoutI FIGURE 20.9 Simple high-pass filter. ArfKen_Ch20-9780123846549.tex 1000 Chapter 20 Integral Transforms Following the usual procedure for electrical circuit analysis, we solve Kirchhoff’s equa- tion (the condition that the net change in potential around any loop of the circuit vanishes): Vinei!tDtZI Cei!tdtCR I ei!t: (20.107) Differentiating with respect to t(to eliminate the integral), we have Vind dtei!tDI Cei!tCR Id dtei!t; which, evaluating the derivatives, reduces to i!VinDI CCi!RI;with solution IDi!CVin 1Ci!RC: (20.108) Since VoutDI R, we easily find the transfer function '.!/DVout VinDi!RC 1Ci!RC: (20.109) To confirm the behavior of the filter, note that in the limit of large !,'.!/!1, while at small!,'.!/!i!RC, which vanishes in the limit of small !. The transition between these two limiting behaviors is a function of the product RC.  Limitations on Transfer Functions Let us write the transfer function '.!/ as the inverse Fourier transform of 8.t/(still using the scaling of this section), keeping in mind that 8.t/vanishes for t<0, '.!/D1Z 08.t/ei!tdt: (20.110) Now, separating 'into its real and imaginary parts: '.!/Du.!/Civ.!/ , and making the same separation for the right-hand side of Eq. (20.110), we have u.!/D1Z 08.t/cos!t dt; v.!/D1Z 08.t/sin!t dt:(20.111) These formulas tell us that u.!/is even, and that v.!/ is odd. Since Eqs. (20.111) are cosine and sine transforms, they can be inverted to give two alternative formulas for 8.t/in the range of applicability of these transforms, namely for ArfKen_Ch20-9780123846549.tex 20.5 Signal-Processing Applications 1001 t>0. Continuing to use the transform scaling of this section, 8.t/D2 1Z 0u.!/cos!t d!; .t>0/ (20.112) D2 1Z 0v.!/sin!t d!: The present significance of these results is that 1Z 0u.!/cos!t d!D1Z 0v.!/sin!t d!; .t>0/: (20.113) The imposition of causality has led to a mutual interdependence of the real and imagi- nary parts of the transfer function. The present result is similar to those involving causality that were discussed in Section 12.8. We close this subsection by verifying that the conditions on uandvare consistent with the properties required of 8. Writing 8.t/D1 21Z 1'.!/ei!tdt; then inserting ei!tDcos!tCisin!tand'DuCiv, we have 8.t/D1 21Z 1h u.!/cos!tv.!/sin!ti d! Ci 21Z 1h u.!/sin!tCv.!/cos!ti d!: (20.114) The imaginary part of Eq. (20.114) vanishes because its integrand is an odd function of !. Ift>0, we know from Eq. (20.113) that the two terms of the real part of Eq. (20.114) are equal, and we get the expected nonzero result. But if t<0, the sign of the second term of the real part is changed and they then add to zero. Exercises 20.5.1 Find the transfer function '.!/ for the circuit shown in the left panel of Fig. 20.10. Is this a high-pass, a low-pass, or a more complicated filter? 20.5.2 Find the transfer function '.!/ for the circuit shown in the right panel of Fig. 20.11. Hint. The potential difference across an inductor is given by L d I=dt. ArfKen_Ch20-9780123846549.tex 1002 Chapter 20 Integral Transforms VinVout Vin VoutR RL C FIGURE 20.10 Circuits for Exercise 20.5.1 (left) and Exercise 20.5.2 (right). Vin Vout C2 I2I2R2 R1I1C1 I1 + I2I1 + I2 FIGURE 20.11 Circuit for Exercise 20.5.3. Vin (Fig. 20.9)(Fig. 20.10) leftVout FIGURE 20.12 Representation of the circuit in Fig. 20.11 in terms of successive transfer functions. 20.5.3 Find the transfer function '.!/ for the circuit shown in Fig. 20.11. This is a band-pass filter. Hint. Assume the currents in the various parts of the circuit to have the values shown in the figure. 20.5.4 The circuit elements for Exercise 20.5.3 correspond to the successive transfer functions shown in Fig. 20.12. Explain why the transfer function for this exercise is only the product of the individual transfer functions in the limit R2R1. 20.6 D ISCRETE FOURIER TRANSFORM For many physicists the Fourier transform is automatically the continuous Fourier trans- form whose analytical properties we have been discussing in previous sections of this chapter. The use of digital computers, however, presents an opportunity to work with numerically determined Fourier transforms, which consist of values given at a discrete set of points. Integrations are therefore converted into finite summations. Transforms defined on discrete point sets have properties worth pursuing, and analysis in that area is the topic of this section. ArfKen_Ch20-9780123846549.tex 20.6 Discrete Fourier Transform 1003 Orthogonality on Discrete Point Sets Throughout the earlier chapters of this book we have introduced and made use of the prop- erties of orthogonal functions, where orthogonality has been defined as the vanishing of an integral whose integrand contains a product of the functions under study. The alternative, to be discussed here, is to define orthogonality over a discrete point set as the vanishing of a sum of products computed at the individual points. It turns out that sines, cosines, and imaginary exponentials have the remarkable property that they are also orthogonal over a series of discrete, equally spaced points on an orthogonality interval. To analyze this situation, we take a set of Nequally spaced points xk, on the interval .0;2/: xkD2k N;kD0;1;2;:::; N1; (20.115) and we consider functions 'p.x/, defined only on the points xkand for integer p, as 'p.x/Deipx: (20.116) In line with our introductory discussion, we define the scalar products of these functions as h'pj'qiDN1X kD0' p.xk/'q.xk/: (20.117) Inserting Eq. (20.115) for the xk, the scalar product takes the form h'pj'qiDN1X kD0e2ik.qp/=NDN1X kD0rk; (20.118) where rDe2i.qp/=N. This is a finite geometric series; if rD1its sum has the value N; otherwise the sum evaluates to .1rN/=.1r/. But rNDe2i.qp/, and because pand qwere restricted to integer values, we have rND1, so the sum vanishes. To complete our understanding of the situation, we need to determine the conditions under which rD1. We clearly have rD1when qDp. Note that we also have rD1when qpis any integer multiple of N. Thus, a formal statement relative to this scalar product is h'pj'qiDN1X nD1qp;nN: (20.119) Note that at most only one of the infinite sum of Kronecker deltas will be nonzero, and all will be zero unless qpis a multiple of N(one of which is qpD0). Equation (20.119) is more complicated than necessary. Because the functions 'pare defined by their values at Npoints, only Nof them are linearly independent. In fact, 'pCN.xk/De2i.pCN/k=NDe2ipk=ND'p.xk/: ArfKen_Ch20-9780123846549.tex 1004 Chapter 20 Integral Transforms We can therefore restrict pandqin Eq. (20.119) to the range .0;N1/, and our orthog- onality relation then becomes h'pj'qiDNpq;0p;qN1: (20.120) Obviously, if function values on a discrete point set are to be used to represent a contin- uous function, the amount of detail that is retained in the analysis will depend on the size of the point set. We will come back to this issue in a later part of the present section. Discrete Fourier Transform By analogy with the definitions introduced for the conventional Fourier transform, we define the discrete transform gp(pD0;:::; N1) of a function fdefined only on the points xkby the formula gpDN1=2N1X kD0e2ikp=Nfk: (20.121) We are now writing fkas a shorthand for f.xk/, and that substitution has pretty much decoupled the problem from the original interval of definition 0x2. In essence, we are now discussing transformations between two N-member sets of function values. The transformation inverse to Eq. (20.121) is fjDN1=2N1X pD0e2 i jp=NgpI (20.122) Eq. (20.122) can be verified by substituting into it the formula for gp, yielding fjDN1N1X pD0N1X kD0e2i.kj/p=NfkDN1X kD0kjfkDfj; as required. These discrete transforms have properties similar to those of their continuous cousins. For example, the transform of fkj, where jis an integer, corresponding to translation by jsteps in the farray, is TfkjUT pDN1=2N1X kD0e2ikp=NfkjDe2i jp=NN1=2N1X kD0e2i.kj/p=Nfkj: Because of the periodicity of the fk, we note that N1=2N1X kD0e2i.kj/p=NfkjDN1=2Nj1X k0Dje2ik0p=Nfk0DN1=2N1X k0D0e2ik0p=Nfk0; ArfKen_Ch20-9780123846549.tex 20.6 Discrete Fourier Transform 1005 which is the formula for the pcoefficient in the transform of f. We therefore have the translation formula TfkjUT pDe2i jp=Ngp: (20.123) We examine next the convolution theorem, where the discrete convolution of two point sets fandgis defined as TfgUkDN1=2N1X jD0fjgkj: (20.124) Taking the transform of this convolution, we have N1N1X kD0e2ikp=NN1X jD0fjgkj D2 4N1=2N1X jD0e2i jp=Nfj3 52 4N1=2N1X kD0e2i.kj/p=Ngkj3 5: As in the continuous case, we have split the complex exponential into two factors. We now redefine the index of the second summation from ktolDkj, thereby making the two square brackets completely independent. Each can then be recognized as a transform (for the second, we need to use the fact that the gkare periodic). The final result is TfgUT pDFpGp; (20.125) where FandGare the respective discrete transforms of fandg. This result is completely analogous with the convolution theorem for the continuous transform. We close this discussion with the observation that the discrete transform and its inverse are linear transformations on coefficient arrays (vectors) of finite dimension N. Therefore, each transform operator can be represented as an NNmatrix whose rows and columns correspond to the points korp. The fact that the transform and its inverse are complex conjugates means that the transformation matrices are unitary. Moreover, from the forms of the transform and its inverse, we see that all the elements of these matrices are proportional to complex exponentials. Limitations As mentioned earlier, the ability of discrete transforms to reproduce phenomena that are actually based on continuous functions will depend on the size of the point set in use. A large amount of detail on errors and limitations in the use of the discrete Fourier transform is provided by Hamming (see Additional Readings). We illustrate the potential problems in the following example. ArfKen_Ch20-9780123846549.tex 1006 Chapter 20 Integral Transforms Example 20.6.1 DISCRETE FOURIER TRANSFORM: ALIASING Let’s consider the simple case f.x/Dcos 3 xon the interval 0x2, which we (ill-advisedly) attempt to treat by the discrete Fourier transform method with ND4. Our four points are at xD0,=2,, and 3=2 , and the four corresponding values of fk are.1;0;1;0/. The problem is that these same four values would be produced from g.x/Dcosx, so neither our discrete transform nor any information derived therefrom can properly reflect any difference in behavior between f.x/andg.x/. If all that we are given are the four values .1;0;1;0/, the most straightforward thing to do is take the discrete transform, yielding .0;1;0;1/, which (from the formula for the inverse transform) corre- sponds to 1 2.0;1;0;1/!eix=2Ce3ix=2 2: If evaluated only at the chosen points, this expression is correct, but if used as an approx- imation over the continuous range .0;2/ it cannot distinguish between cosx,cos 3 x, or any linear combination of the two with unit overall weight. Situations in which the behavior at one wavelength or frequency is mistaken for that at another is called aliasing. The best way to avoid aliasing errors is to use point sets of sufficient size to accommodate the expected extent of oscillatory character in our problem.  Fast Fourier Transform The fast Fourier transform (FFT) is a particular way of factoring and rearranging the terms in the sums of the discrete Fourier transform. Brought to the attention of the scientific com- munity by Cooley and Tukey,4its importance lies in the drastic reduction in the number of numerical operations required. The reduction is possible because the transformation matrix contains large numbers of duplicate entries, and the FFT procedure organizes the compu- tation in a way permitting identical sets of coefficients to be computed only once. Because of the tremendous increase in speed achieved (and reduction in cost), the fast Fourier trans- form has been hailed as one of the few really significant advances in numerical analysis in the past few decades. ForNdata points, a direct calculation of a discrete Fourier transform would require about N2multiplications. For Na power of 2, the fast Fourier transform technique of Cooley and Tukey cuts the number of multiplications required to .N=2/log2N. IfND 1024 (210), the fast Fourier transform achieves a computational reduction by a factor of over 200. This is why the fast Fourier transform is called fast and why it has revolutionized 4J. W. Cooley and J. W. Tukey, Math. Comput. 19: 297 (1965). ArfKen_Ch20-9780123846549.tex 20.6 Discrete Fourier Transform 1007 the digital processing of waveforms. Details on its internal operation will be found in the paper by Cooley and Tukey and in other sources.5 Exercises 20.6.1 Derive the trigonometric forms of discrete orthogonality corresponding to Eq. (20.120): N1X kD0cos.2 pk=N/sin.2 qk=N/D0 N1X kD0cos.2 pk=N/cos.2 qk=N/D8 >< >:0; p6Dq N=2; pDq6D0;N=2 N; pDqD0;N=2 N1X kD0sin.2 pk=N/sin.2 qk=N/D8 >< >:0; p6Dq N=2; pDq6D0;N=2 0; pDqD0;N=2: Note. If Nis odd, pandqwill never have the value N=2. Hint. Consider the use of trigonometric identities such as sinAcosBD1 2h sin.ACB/Csin.AB/i : 20.6.2 Show in detail how to go from FpD1 N1=2N1X kD0fke2ipkto fkD1 N1=2N1X pD0Fpe2 ipk: 20.6.3 TheN-membered point sets fkandFpare discrete Fourier transforms of each other. Derive the following symmetry relations: (a) If fkis real, Fpis Hermitian symmetric; that is, FpDF Np: (b) If fkis pure imaginary, then FpDF Np: Note. The symmetry of part (a) is an illustration of aliasing. If Fpdescribes an ampli- tude at a frequency proportional to p, we necessarily predict an equal amplitude at the frequency proportional to Np. 5See, for example, G. D. Bergland, A guided tour of the fast Fourier transform, IEEE Spectrum 6: 41 (1969). A good discussion can also be found in W. H. Press, B. P. Flannery, S. A. Teukolsky, and W. T. Vetterling, Numerical Recipes, 2nd ed., Cambridge: Cambridge University Press (1996), section 12.3. ArfKen_Ch20-9780123846549.tex 1008 Chapter 20 Integral Transforms 20.7 L APLACE TRANSFORMS Definition The Laplace transform f.s/of a function F.t/is defined by6 f.s/DLfF.t/gD1Z 0estF.t/dt: (20.126) A few comments on the existence of the integral are in order. The infinite integral of F.t/, 1Z 0F.t/dt; need not exist. For instance, F.t/may diverge exponentially for large t. However, if there are some constants s0,M, and t00such that for all t>t0 jes0tF.t/jM; (20.127) the Laplace transform will exist for s>s0;F.t/is then said to be of exponential order. As a counterexample, F.t/Det2does not satisfy the condition given by Eq. (20.127) and isnotof exponential order. Thus, Ln et2o does notexist. The Laplace transform may also fail to exist because of a sufficiently strong singularity in the function F.t/ast!0. For example, 1Z 0esttndt diverges at the origin for n1 . The Laplace transform Lftngdoes not exist for n1 . Since, for two functions F.t/andG.t/for which the integrals exist, Ln aF.t/CbG.t/o DaLfF.t/gCbLfG.t/g; (20.128) the operation denoted by Lislinear. Elementary Functions To introduce the Laplace transform, let us apply the operation to some of the elementary functions. In all cases we assume that F.t/D0fort<0. If F.t/D1; t>0; 6This is sometimes called a one-sided Laplace transform; the integral from 1 toC1 is referred to as a two-sided Laplace transform. Some authors introduce an additional factor of s. This extra sappears to have little advantage and continually gets in the way; for further comments, see section 14.13 in the text by Jeffreys and Jeffreys (Additional Readings). Generally, we takesto be real and positive. It is possible to let sbecome complex, provided <e.s/>0. ArfKen_Ch20-9780123846549.tex 20.7 Laplace Transforms 1009 then Lf1gD1Z 0estdtD1 s;for s>0: (20.129) Next, let F.t/Dekt;t>0: The Laplace transform becomes Ln ekto D1Z 0estektdtD1 sk;for s>k: (20.130) Using this relation, we obtain the Laplace transform of certain other functions. Since cosh ktD1 2.ektCekt/;sinhktD1 2.ektekt/; (20.131) we have Lfcosh ktgD1 21 skC1 sCk Ds s2k2; (20.132) LfsinhktgD1 21 sk1 sCk Dk s2k2; (20.133) both valid for s>k. From the relations cosktDcosh ikt;sinktDisinhikt; it is evident that we can obtain transforms of the sine and cosine if kis replaced by ikin Eqs. (20.132) and (20.133): LfcosktgDs s2Ck2; (20.134) LfsinktgDk s2Ck2; (20.135) both valid for s>0. Another derivation of this last transform is given in Example 20.8.1. It is a curious fact that lims!0LfsinktgD1=kdespite the fact thatR1 0sinkt dt does not exist. Finally, for F.t/Dtn, we have L tn D1Z 0esttndt; which is just a gamma function. Hence L tn D0.nC1/ snC1;s>0;n>1: (20.136) ArfKen_Ch20-9780123846549.tex 1010 Chapter 20 Integral Transforms Note that in all these transforms we have the variable sin the denominator, so that it occurs as a negative power. From the definition of the transform, Eq. (20.126) and the existence condition, Eq. (20.127), it is clear that if f.s/is a Laplace transform, then lims!1 f.s/D0. The significance of this point is that if f.s/behaves asymptotically for large sas a positive power of s, then no inverse transform can exist. Heaviside Step Function In Exercise 1.11.9 we encountered the Heaviside step function u.t/. Because of its utility in describing discontinuous signal pulses, its Laplace transform occurs frequently. We therefore remind the reader of the definition u.tk/D( 0; t<k; 1; t>k:(20.137) Taking the transform, we have Lfu.tk/gD1Z kestdtD1 seks: (20.138) Example 20.7.1 TRANSFORM OF SQUARE PULSE Let’s compute the transform of a square pulse F.t/of height Athat is on from tD0to tDt0; see Fig. 20.13. Using the Heaviside step function, the pulse can be represented as F.t/DAh u.t/u.tt0/i : Its transform is therefore LfF.t/gD1 s.1et0s/:  Dirac Delta Function For use with differential equations one further transform is helpful, namely that of the Dirac delta function. From the properties of the delta function, we have Lf.tt0/gD1Z 0est.tt0/dtDest0;for t0>0: (20.139) t=0 t=t0A− FIGURE 20.13 Square pulse. ArfKen_Ch20-9780123846549.tex 20.7 Laplace Transforms 1011 Fort0D0we must be a bit more careful, as the sequences we have used for defining the delta function involve contributions symmetrically distributed about t0, and the integration defining the Laplace transform is restricted to t0. Consistent results when using Laplace transforms, however, are obtained if we consider delta sequences that are entirely within the range tt0, which is equivalent to Lf.t/gD1: (20.140) This delta function is frequently called the impulse function because it is so useful in describing impulsive forces, that is, forces lasting only a short time. Inverse Transform As we have already seen in our discussion of the Fourier transform, the taking of an integral transform will ordinarily have little value unless we can carry out the inverse transform. That is, with LfF.t/gDf.s/; then it is desirable to be able to compute L1ff.s/gDF.t/: (20.141) However, this inverse transform is not entirely unique. Two functions F1.t/andF2.t/can have the same transform, f.s/, if their difference, N.t/DF1.t/F2.t/, is a null function, meaning that for all t0>0it satisfies t0Z 0N.t/dtD0: This result is known as Lerch’s theorem, and is not quite equivalent to F1DF2, because it permits F1andF2to differ at isolated points. However, in most problems studied by physicists or engineers this ambiguity is not important and we will not consider it further. The inverse transform can be determined in various ways. 1. A table of transforms can be built up and used to identify inverse transformations, exactly as a table of logarithms can be used to look up antilogarithms. The preceding transforms constitute the beginnings of such a table. More complete sets of Laplace transforms are in several of the Additional Readings, and a relatively short table of transforms appears in the present text as Table 20.1. Many functional forms not in Table 20.1 can be reduced to tabular entries using a partial fraction expansion or other properties of the Laplace transform presented later in this chapter. Of particular value in this regard are the translation and derivative formulas. There is some justification for suspecting that these tables are probably of more value in solving textbook exercises than in solving real-world problems. 2. A general technique for L1will be developed in Section 20.10 by using the calculus of residues. 3. Transforms and their inverses can be represented numerically. See the work by Krylov and Skoblya in Additional Readings. ArfKen_Ch20-9780123846549.tex 1012 Chapter 20 Integral Transforms Table 20.1 Laplace Transformsa f.s/ F.t/ Limitation Equation 1: 1 .t/ Singularity atC0.20:140/ 2:1 s1 s>0 .20:129/ 3:0.nC1/ snC1tns>0;n>1.20:136/ 4:1 skekts>k .20:130/ 5:1 .sk/2tekts>k .20:176/ 6:s s2k2cosh kt s >k .20:132/ 7:k s2k2sinhkt s >k .20:133/ 8:s s2Ck2coskt s >0 .20:134/ 9:k s2Ck2sinkt s >0 .20:135/ 10:sa .sa/2Ck2eatcoskt s >a .20:159/ 11:k .sa/2Ck2eatsinkt s >a .20:158/ 12:s2k2 .s2Ck2/2tcoskt s >0 .20:177/ 13:2ks .s2Ck2/2tsinkt s >0 .20:178/ 14: .s2Ca2/1=2J0.at/ s>0 .20:182/ 15: .s2a2/1=2I0.at/ s>a Exercise 20.8.13 16:1 acot1s a j0.at/ s>0 Exercise 20.8.14 17:1 2alnsCa sa 1 acoth1s a9 >>= >>;i0.at/ s>a Exercise 20.8.14 18:.sa/n snC1Ln.at/ s>0 Exercise 20.8.16 19:1 sln.sC1/ E1.x/ s>0 Exercise 20.8.17 20:lns slnt s>0 Exercise 20.10.9 a is the Euler-Mascheroni constant. ArfKen_Ch20-9780123846549.tex 20.7 Laplace Transforms 1013 Example 20.7.2 PARTIAL FRACTION EXPANSION The function f.s/Dk2=s.s2Ck2/does not appear as a transform listed in Table 20.1, but we may obtain it from the tabulated transforms by observing that it has the partial fraction expansion f.s/Dk2 s.s2Ck2/D1 ss s2Ck2: The partial fraction technique was discussed in Section 1.5, and the present example was the subject of Example 1.5.3. Each of the two partial fractions corresponds to an entry in Table 20.1, and we can therefore take the inverse transform of f.s/term by term: L1ff.s/gD1coskt: Remember that the range of the inverse transform is restricted to t0.  Example 20.7.3 A STEP FUNCTION This example shows how Laplace transforms can be used to evaluate a definite integral. Consider F.t/D1Z 0sintx xdx: (20.142) Suppose we take the Laplace transform of this definite (and improper) integral, naming itf.s/: f.s/DL8 < :1Z 0sintx xdx9 = ;D1Z 0est1Z 0sintx xdx dt: Now, interchanging the order of integration (which is justified),7we get f.s/D1Z 01 x2 41Z 0estsintx dt3 5dxD1Z 0dx s2Cx2; (20.143) since the factor in square brackets is just the Laplace transform of sintx. The integral on the right-hand side is elementary, with evaluation f.s/D1Z 0dx s2Cx2D1 stan1x s 1 0D 2s: (20.144) Using entry #2 in Table 20.1, we carry out the inverse transformation to obtain F.t/D 2;t>0; (20.145) 7See Chapter 1 in Jeffreys and Jeffreys (Additional Readings) for a discussion of uniform convergence of integrals. ArfKen_Ch20-9780123846549.tex 1014 Chapter 20 Integral Transforms π 2 tF(t) π 2− FIGURE 20.14 F.t/DR1 0sintx xdx, a step function. in agreement with an evaluation by the calculus of residues, Eq. (11.107). It has been assumed that t>0inF.t/. For F.t/we need note only that sin.tx/Dsintx, giving F.t/DF.t/. Finally, if tD0;F.0/is clearly zero. Therefore, 1Z 0sintx xdxD 2T2u.t/1UD8 >>>< >>>: 2; t>0 0; tD0  2;t<0:(20.146) Here u.t/is the Heaviside unit step function, Eq. (20.137). Thus,1R 0.sintx=x/dx, taken as a function of t, describes a step function (Fig. 20.14), with a step of height attD0. The technique in the preceding example was to (1) introduce a second integration, namely the Laplace transform, (2) reverse the order of integration and integrate once, and (3) take the inverse Laplace transform. This is a technique that will apply to many problems. Exercises 20.7.1 Prove that lims!1s f.s/Dlim t!C0F.t/: Hint. Assume that F.t/can be expressed as F.t/DP1 nD0antn. 20.7.2 Show that 1 lim s!0LfcosxtgD.x/: 20.7.3 Verify that Lcosatcosbt b2a2 Ds .s2Ca2/.s2Cb2/;a26Db2: ArfKen_Ch20-9780123846549.tex 20.7 Laplace Transforms 1015 20.7.4 Using partial fraction expansions, show that (a)L11 .sCa/.sCb/ Deatebt ba;a6Db: (b)L1s .sCa/.sCb/ Daeatbebt ab;a6Db: 20.7.5 Using partial fraction expansions, show that for a26Db2; (a)L11 .s2Ca2/.s2Cb2/ D1 a2b2sinat asinbt b : (b)L1s2 .s2Ca2/.s2Cb2/ D1 a2b2fasinatbsinbtg: 20.7.6 Show that (a)1Z 0coss sdsD 2.1/Wcos.=2/;0<< 1. (b)1Z 0sins sdsD 2.1/Wsin.=2/;0<< 2. Why isrestricted to (0, 1) for (a), to (0, 2) for (b)? These integrals may be interpreted as Fourier transforms of sand as Mellin transforms of sinsandcoss. Hint. Replace sby a Laplace transform integral: L t1 =0./ . Then integrate with respect to s. The resulting integral can be treated as a beta function (Section 13.3). 20.7.7 A function F.t/can be expanded in a Maclaurin series, F.t/D1X nD0antn: Then LfF.t/gD1Z 0est1X nD0antndtD1X nD0an1Z 0esttndt: Show that f.s/, the Laplace transform of F.t/, contains no powers of sgreater than s1. Check your result by calculating Lf.t/g;and comment on this fiasco. 20.7.8 Show that the Laplace transform of the confluent hypergeometric function M.a;cIx/is LfM.a;cIx/gD1 s2F1 a;1IcI1 s : ArfKen_Ch20-9780123846549.tex 1016 Chapter 20 Integral Transforms 20.8 P ROPERTIES OF LAPLACE TRANSFORMS Transforms of Derivatives Perhaps the main application of Laplace transforms is in converting differential equations into simpler forms that may be solved more easily. It will be seen, for instance, that coupled differential equations with constant coefficients transform to simultaneous linear algebraic equations. For the study of differential equations we need formulas for the Laplace trans- forms of the derivatives of a function. Let us transform the first derivative of F.t/: L F0.t/ D1Z 0estdF.t/ dtdt: Integrating by parts, we obtain L F0.t/ DestF.t/ 1 0Cs1Z 0estF.t/dt DsLfF.t/gF.0/: (20.147) Strictly speaking, F.0/DF.C0/;8and dF=dt is required to be at least piecewise continuous for 0t<1. Naturally, both F.t/and its derivative must be such that the integrals do not diverge. An extension to higher derivatives gives Ln F.2/.t/o Ds2LfF.t/gsF.C0/F0.C0/; (20.148) LfF.n/.t/gDsnLfF.t/gsn1F.C0/ F.n1/.C0/: (20.149) The Laplace transform, like the Fourier transform, replaces differentiation with multi- plication. In the following examples ODEs become algebraic equations. Here is the power and the utility of the Laplace transform. But see Example 20.8.7 for what may happen if the coefficients are not constant. Note how the initial conditions, F.C0/; F0.C0/ , and so on, are incorporated into the transform. This situation is different than for the Fourier transform, and arises from the finite lower limit ( tD0) of the integral defining the transform. This property makes the Laplace transform more powerful for obtaining solutions to differential equations sub- ject to initial conditions. 8This notation means that zero is approached from the positive side. ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1017 Example 20.8.1 USE OF DERIVATIVE FORMULA Here is an example showing how the derivative formula has uses even in contexts not involving the solution to a differential equation. Starting from the identity k2sinktDd2 dt2sinkt; (20.150) we apply on both sides of the equation the Laplace transform operation, reaching k2LfsinktgDLd2 dt2sinkt Ds2Lfsinktgssin.0/d dtsinkt tD0: Since sin.0/D0andd=dtsinktjtD0Dk, the above equation has solution LfsinktgDk s2Ck2: This result confirms Eq. (20.135).  Examples involving the solutions of differential equations follow. Example 20.8.2 SIMPLE HARMONIC OSCILLATOR As a physical example, consider a mass moscillating under the influence of an ideal spring, spring constant k. As usual, friction is neglected. Then Newton’s second law becomes md2X.t/ dt2Ck X.t/D0: (20.151) We take as initial conditions X.0/DX0;X0.0/D0: Applying the Laplace transform, we obtain mLd2X dt2 CkLfX.t/gD0: (20.152) Letting x.s/denote the presently unknown transform LfX.t/gand using Eq. (20.148), we convert Eq. (20.152) to the form ms2x.s/ms X 0Ckx.s/D0; which has solution x.s/DX0s s2C!2 0;with!2 0k m: From Table 20.1 this is seen to be the transform of cos!0t, which gives the expected result: X.t/DX0cos!0t: (20.153)  ArfKen_Ch20-9780123846549.tex 1018 Chapter 20 Integral Transforms Example 20.8.3 EARTH’S NUTATION A somewhat more involved example is the nutation of the Earth’s poles (force-free pre- cession). We treat the Earth as a rigid (oblate) spheroid, with z-axis through its direction of symmetry. We assume the spheroid to have moments of inertia IzandIxDIyand to be rotating about its x,y, and zaxes at the respective angular velocities X.t/D!x.t/, Y.t/!y.t/,!zDconstant. The Euler equations of motion for XandYreduce to d X dtDaY;dY dtDCaX; (20.154) where aT.IzIx/=IzU!z. For the Earth, the initial values of XandYare not both zero, so the axis of rotation is not aligned with the symmetry axis (see Fig. 20.15), and because of this lack of alignment, the axis of rotation precesses about the axis of symmetry. For the Earth, the deviation between the rotation and symmetry axes is small, only about 15 meters (measured at the Earth’s surface at the poles). Our first step in solving these coupled ODEs is to take their Laplace transforms, obtaining sx.s/X.0/Day.s/; sy.s/Y.0/Dax.s/: Combining to eliminate y.s/, we have s2x.s/s X.0/CaY.0/Da2x.s/; or x.s/DX.0/s s2Ca2Y.0/a s2Ca2: (20.155) Recognizing these functions of sas transforms listed in Table 20.1, X.t/DX.0/cos atY.0/sin at: Y X ωz FIGURE 20.15 Earth’s rotation axis and its components. ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1019 Similarly, Y.t/DX.0/sin atCY.0/cos at: This is seen to be a rotation of the vector .X;Y/counterclockwise (for a>0) about the z-axis with angle Datand angular velocity a. A direct interpretation may be found by choosing the xandyaxes so that Y.0/D0. Then X.t/DX.0/cos at;Y.t/DX.0/sin at; which are the parametric equations for rotation of .X;Y/in a circular orbit of radius X.0/, with angular velocity ain the counterclockwise sense. For the Earth, aas defined here corresponds to a period .2= a/of some 300 days. Actually, because of departures from the idealized rigid body assumed in setting up Euler’s equations, the period is about 427 days.9 These same equations arise in electromagnetic theory. If in Eq. (20.154) we set X.t/DLx;Y.t/DLy; where LxandLyare the x- and y-components of the angular momentum Lof a charged particle moving in a uniform magnetic field Bzez, and then assign athe value aDgLBz, where gLis the gyromagnetic ratio of the particle, then Eq. (20.148) determines its Larmor precession in the magnetic field.  Example 20.8.4 IMPULSIVE FORCE For an impulsive force acting on a particle of mass m, Newton’s second law takes the form md2X dt2DP.t/; where Pis a constant. Transforming, we obtain ms2x.s/ms X.0/m X0.0/DP: For a particle starting from rest, X0.0/D0. We shall also take X.0/D0. Then x.s/DP ms2; and, taking the inverse transform, X.t/DP mt; d X.t/ dtDP m;a constant: The effect of the impulse P.t/is to transfer (instantaneously) Punits of linear momen- tum to the particle. 9D. Menzel, ed., Fundamental Formulas of Physics, Englewood Cliffs, NJ: Prentice-Hall (1955), reprinted, Dover (1960), p. 695. ArfKen_Ch20-9780123846549.tex 1020 Chapter 20 Integral Transforms A similar analysis applies to the ballistic galvanometer. The torque on the galvanometer is given initially by k, in whichis a pulse of current and kis a proportionality constant. Sinceis of short duration, we set kDkq.t/; where qis the total charge carried by the current . Then, with Ithe moment of inertia, Id2 dt2Dkq.t/; and transforming, as before, we find that the effect of the current pulse is a transfer of kq units of angular momentum to the galvanometer.  Change of Scale If we replace tbyatin the defining formula for the Laplace transform, we readily obtain LfF.at/gD1Z 0estF.at/dtD1 a1Z 0e.s=a/.at/F.at/d.at/ D1 afs a : (20.156) Substitution If we replace the parameter sbysain the definition of the Laplace transform, Eq. (20.126), we have f.sa/D1Z 0e.sa/tF.t/dtD1Z 0esteatF.t/dt DL eatF.t/ : (20.157) Hence the replacement of swith sacorresponds to multiplying F.t/byeat;and conversely. This result can used to check some entries in our table of transforms. From Eq. (20.157) we find immediately that L eatsinkt Dk .sa/2Ck2; .s>a/; (20.158) and L eatcoskt Dsa .sa/2Ck2;s>a: (20.159) These are entries 10 and 11 of Table 20.1. ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1021 Example 20.8.5 DAMPED OSCILLATOR Equations (20.158) and(20.159) are useful when we consider an oscillating mass with damping proportional to the velocity. Equation (20.151), with such damping added, becomes m X00.t/CbX0.t/Ck X.t/D0; (20.160) in which bis a proportionality constant. Let us assume that the particle starts from rest at X.0/DX0, soX0.0/D0. The transformed equation is mTs2x.s/s X0UCbTsx.s/X0UCkx.s/D0; with solution x.s/DX0msCb ms2CbsCk: This transform does not appear in our table, but may be handled by completing the square of the denominator: s2Cb msCk mD sCb 2m2 Ck mb2 4m2 : Considering further only the case that the damping is small enough that b2<4km, then the last term is positive and will be denoted by !2 1. We then rearrange x.s/to the form x.s/DX0sCb=m .sCb=2m/2C!2 1 DX0sCb=2m .sCb=2m/2C!2 1CX0!1.b=2m!1/ .sCb=2m/2C!2 1: These are the same transforms we encountered in Eqs. (20.158) and(20.159), so we may take the inverse transform of our formula for x.s/, reaching X.t/DX0e.b=2m/t cos!1tCb 2m!1sin!1t DX0!0 !1e.b=2m/tcos.! 1t'/: (20.161) Here we have made the substitutions tan'Db 2m!1; !2 0Dk m: Of course, as b!0, this solution goes over to the undamped solution, given in Example 20.8.2.  ArfKen_Ch20-9780123846549.tex 1022 Chapter 20 Integral Transforms CR L FIGURE 20.16 RLC circuit. RLC Analog It is worth noting the similarity between the damped simple harmonic oscillation of a mass (Example 20.8.5) and an RLC circuit (resistance, inductance, and capacitance). See Fig. 20.16. At any instant, the sum of the potential differences around the loop must be zero (Kirchhoff’s law, conservation of energy). This gives Ld I dtCRIC1 CtZ I dtD0: (20.162) Differentiating Eq. (20.162) with respect to time (to eliminate the integral), we have Ld2I dt2CRd I dtC1 CID0: (20.163) If we replace I.t/with X.t/,Lwith m,Rwith b, and C1with k, then Eq. (20.163) is identical with the mechanical problem. It is but one example of the unification of diverse branches of physics by mathematics. A more complete discussion will be found in a book by Olson.10 Translation This time let f.s/be multiplied by ebs, with b>0: ebsf.s/Debs1Z 0estF.t/dt D1Z 0es.tCb/F.t/dt: (20.164) 10H. F. Olson, Dynamical Analogies, New York: Van Nostrand (1943). ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1023 Now let tCbD. Equation (20.164) becomes ebsf.s/D1Z besF.b/d: (20.165) Since F.t/is assumed to be equal to zero for t<0, so that F.b/D0for0 <b, we can change the lower limit in Eq. (20.165) to zero without changing the value of the integral. Then renaming as our standard Laplace transform variable t, we have ebsf.s/DLfF.tb/g: (20.166) If instead of relying on the assumption that F.t/D0for negative twe insert a Heaviside unit step function u.b/to restrict the contributions from Fto positive arguments, Eq. (20.165) takes the form ebsf.s/D1Z 0esF.b/u.b/d: For this reason the translation formula, Eq. (20.166), is often called the Heaviside shifting theorem. Example 20.8.6 ELECTROMAGNETIC WAVES The electromagnetic wave equation with EDEyorEz, a transverse wave propagating along the x-axis, is @2E.x;t/ @x21 [email protected];t/ @t2D0: (20.167) We want to solve this PDE for the situation that a source at xD0generates a time- dependent signal E.0;t/starting at time tD0and propagating only toward positive x, with initial conditions that for x>0, E.x;0/D0;@E.x;t/ @t tD0D0: Transforming Eq. (20.167) with respect to t, we get @2 @x2LfE.x;t/gs2 v2LfE.x;t/gCs v2E.x;0/C1 [email protected];t/ @t tD0D0; which due to the initial conditions simplifies to @2 @x2LfE.x;t/gDs2 v2LfE.x;t/g: (20.168) ArfKen_Ch20-9780123846549.tex 1024 Chapter 20 Integral Transforms The general solution of Eq. (20.168) (which is an ODE inx) is LfE.x;t/gDf1.s/e.s=v/xCf2.s/eC.s=v/x: (20.169) To understand more fully this result consider first the case f2.s/D0. Then Eq. (20.169) becomes LfE.x;t/gDe.x=v/sf1.s/; (20.170) which we recognize as of the same form as Eq. (20.166), meaning that E.x;t/DF tx v ; where Fis the function whose Laplace transform is f1, namely E.0;t/.11Since Fis assumed to vanish when its argument is negative, this formula can be written in the more explicit form E.x;t/D8 >< >:F tx v DE 0;tx v ;tx v; 0; t<x v:(20.171) This solution represents a wave (or pulse) moving in the positive x-direction with velocity v. Note that for x>vtthe region remains undisturbed; the pulse has not had time to get there. If we had decided to take the solution of Eq. (20.169) with f1.s/D0, we would have obtained E.x;t/D8 >< >:F tCx v DE 0;tCx v ;tx v; 0; t<x v;(20.172) which we must reject because (for propagation toward positive x) it violates causality. Our solution to this problem, Eq. (20.171), can be verified by differentiation and substi- tution into the original PDE, Eq. (20.167).  Derivative of a Transform When F.t/, which is at least piecewise continuous, and sare chosen so that estF.t/ converges exponentially for large s, the integral 1Z 0estF.t/dt 11Consider Eq. (20.170) with xset to zero. ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1025 is uniformly convergent and may be differentiated (under the integral sign) with respect tos. Then f0.s/D1Z 0.t/estF.t/dtDLft F.t/g: (20.173) Continuing this process, we obtain f.n/.s/DL .t/nF.t/ : (20.174) All the integrals so obtained will be uniformly convergent because of the decreasing expo- nential behavior of estF.t/. This technique may be applied to generate more transforms. For example, Ln ekto D1Z 0estektdtD1 sk;s>k: (20.175) Differentiating with respect to s(or with respect to k), we obtain Ln tekto D1 .sk/2;s>k: (20.176) If we replace kbyikand separate Eq. (20.176) into its real and imaginary parts, we get LftcosktgDs2k2 .s2Ck2/2; (20.177) LftsinktgD2ks .s2Ck2/2: (20.178) These expressions are valid for s>0. Example 20.8.7 BESSEL’S EQUATION An interesting application of a differentiated Laplace transform appears in the solution of Bessel’s equation with nD0. From Chapter 14 we have x2y00.x/Cxy0.x/Cx2y.x/D0: This ODE cannot be solved by the method illustrated in Example 20.8.2 because the deriva- tives are multiplied by functions of the independent variable x. However, an alternate approach depending on Eq. (20.174) is available. Dividing by xand substituting tDxand F.t/Dy.x/to agree with the present notation, we see that the Bessel equation becomes t F00.t/CF0.t/Ct F.t/D0: (20.179) ArfKen_Ch20-9780123846549.tex 1026 Chapter 20 Integral Transforms We need a regular solution, and it appears possible for F.0/to be nonzero, so we scale the solution by setting F.0/D1. Then, setting tD0in Eq. (20.179), we find that F0.C0/D0. In addition, we assume that our unknown F.t/has a transform. Transforming Eq. (20.179), using Eqs. (20.147) and(20.148) for the derivatives and Eq. (20.173) to append factors of t, we have d dsh s2f.s/si Cs f.s/1d dsf.s/D0: (20.180) Rearranging and simplifying, we obtain .s2C1/f0.s/Cs f.s/D0; or d f fDs ds s2C1; a first-order ODE. By integration, lnf.s/D1 2ln.s2C1/ClnC; which may be rewritten as f.s/DCp s2C1: (20.181) To confirm that our transform yields the power-series expansion of J0, we expand f.s/ as given in Eq. (20.181) in a series of negative powers of s, convergent for s>1: f.s/DC s 1C1 s21=2 DC s 11 2s2C13 222Ws4C.1/n.2n/W .2nnW/2s2nC : Inverting, term by term, we obtain F.t/DC1X nD0.1/nt2n .2nnW/2: When Cis set equal to 1, as required by the initial condition F.0/D1, we recover J0.t/, our familiar Bessel function of order zero. Hence, LfJ0.t/gD1p s2C1: (20.182) This simple, closed form is the Laplace transform of J0.t/. After making a scale change to form J0.at/using Eq. (20.156), we confirm entry 14 of Table 20.1. Note that in our derivation of Eq. (20.182) we assumed s>1. The proof for s>0is the topic of Exercise 20.8.10.  ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1027 It is worth noting that this application was successful and relatively easy because we took nD0in Bessel’s equation. This made it possible to divide out a factor of x(ort). If this had not been done, the terms of the form t2F.t/would have introduced a second derivative off.s/. The resulting equation would have been no easier to solve than the original one. This observation illustrates the point that when we go beyond linear ODEs with constant coefficients, the Laplace transform may still be applied, but there is no guarantee that it will be helpful. The application to Bessel’s equation, n6D0, will be found in the Additional Readings. Alternatively, given the result LfJn.at/gDan.p s2Ca2s/n p s2Ca2; (20.183) we can confirm its validity by expressing Jn.t/as an infinite series and transforming term by term. Integration of Transforms Again, with F.t/at least piecewise continuous and xlarge enough so that extF.t/ decreases exponentially (as x!1 ), the integral f.x/D1Z 0extF.t/dt is uniformly convergent with respect to x. This justifies reversing the order of integration in the following equation: 1Z sf.x/dxD1Z sdx1Z 0dt extF.t/D1Z 0estF.t/ tdt; DLF.t/ t ; (20.184) where the last member of the first line is obtained by integrating with respect to x. The lower limit smust be chosen large enough so that f.s/is within the region of uniform convergence. Equation (20.184) is valid when F.t/=tis finite at tD0or diverges less strongly than t1(so that LfF.t/=tgwill exist). For convenience we summarize the definition and properties of the Laplace transform in Table 20.2. Included in the table are formulas for convolution and inversion that will be discussed in Sections 20.9 and Sections 20.10. ArfKen_Ch20-9780123846549.tex 1028 Chapter 20 Integral Transforms Table 20.2 Laplace Transform Operations Operation Equation 1. Laplace transform f.s/DLfF.t/gD1Z 0estF.t/dt (15.99) 2. Transform of derivative s f.s/F.C0/DL F0.t/ (15.123) s2f.s/sF.C0/F0.C0/DL F00.t/ (15.124) 3. Transform of integral1 sf.s/DL8 < :tZ 0F.x/dx9 = ;Exercise 20.9.1 4. Change of scale1 afs a DLfF.at/g (20.156) 5. Substitution f.sa/DL eatF.t/ (15.152) 6. Translation ebsf.s/DLfF.tb/g (15.164) 7. Derivative of transform f.n/.s/DL .t/nF.t/ (15.173) 8. Integral of transform1Z sf.x/dxDLF.t/ t (15.189) 9. Convolution f1.s/f2.s/DL8 < :tZ 0F1.tz/F2.z/dz9 = ;(15.193) 10. Inverse transform,1 2i Ci1Z i1estf.s/dsDF.t/ (15.212) Bromwich integrala a must be large enough that e tF.t/vanishes as t!C1. Exercises 20.8.1 Use the expression for the transform of a second derivative to obtain the transform of coskt. 20.8.2 A mass mis attached to one end of an unstretched spring, spring constant k(Fig. 20.17). Starting at time tD0, the free end of the spring experiences a constant acceleration a, away from the mass. Using Laplace transforms, (a) find the position xofmas a function of time. (b) determine the limiting form of x.t/for small t. ANS..a/ xD1 2at2a !2.1cos!t/; !2Dk m; .b/ xDa!2 4Wt4; ! t1: ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1029 m xx1 FIGURE 20.17 Spring, Exercise 20.8.2. 20.8.3 Radioactive nuclei decay according to the law d N dtD N; with Nthe concentration of a given nuclide and its particular decay constant. This equation may be interpreted as stating that the rate of decay is proportional to the num- ber of these radioactive nuclei present. They all decay independently. Consider now a radioactive series of ndifferent nuclides, with Nuclide 1 decaying into Nuclide 2, Nuclide 2 into Nuclide 3, etc., until reaching Nuclide n, which is stable. The concentrations of the various nuclides satisfy the system of ODEs d N1 dtD 1N1;d N2 dtD1N12N2;;d Nn dtDn1Nn1: (a) For the case nD3findN1.t/,N2.t/, and N3.t/, with N1.0/DN0andN2.0/D N3.0/D0. (b) Find an approximate expression for N2andN3, valid for small twhen12. (c) Find approximate expressions for N2andN3, valid for large t, when (1)12, (2)12. ANS. (a) N1.t/DN0e1t,N2.t/DN01 21.e1te2t/, N3.t/DN0 12 21e1tC1 21e2t . (b) N2N01t,N3N0 212t2. (c) (1) N2N0e2t N3N0.1e2t/;  1t1. (2)N2N0.1=2/e1t, N3N0.1e1t/;  2t1. ArfKen_Ch20-9780123846549.tex 1030 Chapter 20 Integral Transforms 20.8.4 The rate of formation of an isotope in a nuclear reactor is given by d N2 dtD'h 1N1.0/2N2.t/i 2N2.t/: Here N1.0/is the concentration of the original isotope (assumed constant), and N2is that of the newly formed isotope. The first two terms on the right-hand side describe the production and destruction of the new isotope via neutron absorption; 'is the neutron flux (units cm2s1);1and2(units cm2) are neutron absorption cross sections. The final term describes the radioactive decay of the new isotope, with decay constant 2. (a) Find the concentration N2of the new isotope as a function of time. (b) For original isotope153Eu,1D400barnsD4001024cm2,2D1000 barns D10001024cm2, and2D1:4109s1. If N1.0/D1020and'D 109cm2s1, find N2, the concentration of154Eu, after 1 year of continuous irra- diation. Is the assumption that N1is constant justified? 20.8.5 In a nuclear reactor135Xe is formed as both a direct fission product of235U and by decay of135I (another fission product), half-life 6.7 hours. The half-life of135Xe is 9.2 hours. Because135Xe strongly absorbs thermal neutrons, thereby “poisoning” the nuclear reactor, its concentration is a matter of great interest. The relevant equations are d NI dtD' I.fNU/INI; d NXe dtD'h Xe.fNU/XeNXei CINIXeNXe: Here NI,NXe,NUare the concentrations of135I,135Xe,235U, with NUassumed to be constant. The neutron flux 'in the reactor causes fission of235U with cross section f and removes135Xe by neutron absorption with cross section XeD3:5106barnsD 3:51018cm2. Neutron absorption by135I is negligible. The yield of135I and135Xe per fission are, respectively, ID0:060 and XeD0:003 . (a) Find NXe.t/in terms of neutron flux 'and the product fNU. (b) Find NXe.t!1/ . (c) After NXehas reached equilibrium, the reactor is shut down: 'D0. Find NXe.t/ following shutdown. Note the short-term increase in NXe, which may for a few hours interfere with starting the reactor up again. Hint. The half-life t1=2of a radioactive isotope is the time required for decay of half of the nuclides in a sample. For a decay rate d N=dtD N, the half-life has the value t1=2Dln 2= , socan be computed as Dln 2=t1=2D0:693= t1=2. 20.8.6 Solve Eq. (20.160), which describes a damped simple harmonic oscillator, for X.0/D X0,X0.0/D0, and (a) b2D4mk(critically damped), (b) b2>4mk(overdamped). ANS..a/X.t/DX0e.b=2m/t 1Cb 2mt : ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1031 20.8.7 Again solve Eq. (20.160), which describes a damped simple harmonic oscillator, but this time for X.0/D0,X0.0/Dv0, and (a) b2<4mk(underdamped), (b) b2D4mk(critically damped), (c) b2>4mk(overdamped). ANS. (a)X.t/Dv0 !1e.b=2m/tsin!1t, (b)X.t/Dv0te.b=2m/t. 20.8.8 The motion of a body falling in a resisting medium may be described by md2X.t/ dt2Dmgbd X.t/ dt when the retarding force is proportional to the velocity. Find X.t/andd X.t/=dt for the initial conditions X.0/Dd X dt tD0D0: 20.8.9 Ringing circuit. In certain electronic devices, resistance, inductance, and capacitance are placed in a circuit as shown in Fig. 20.18. A constant voltage is maintained across the capacitance, keeping it charged. At time tD0the circuit is disconnected from the voltage source. Find the voltages across each of the elements R;L, and Cas a function of time. Assume Rto be small. Hint. By Kirchhoff ’s laws IRLCICD0and ERCELDEC; where ERDIRLR;ELDLd IRL dt;ECDq0 CC1 CtZ 0ICdt; q0Dinitial charge of capacitor. 20.8.10 With J0.t/expressed as a contour integral, apply the Laplace transform operation, reverse the order of integration, and thus show that LfJ0.t/gD.s2C1/1=2;fors>0: LR C FIGURE 20.18 Ringing circuit. ArfKen_Ch20-9780123846549.tex 1032 Chapter 20 Integral Transforms 20.8.11 Develop the Laplace transform of Jn.t/fromLfJ0.t/gby using the Bessel function recurrence relations. Hint. Here is a chance to use mathematical induction (Section 1.4). 20.8.12 A calculation of the magnetic field of a circular current loop in circular cylindrical coordinates leads to the integral 1Z 0ekzk J1.ka/dk;<e z0: Show that this integral is equal to a=.z2Ca2/3=2. 20.8.13 Show that LfI0.at/gD.s2a2/1=2;s>a: 20.8.14 Verify the following Laplace transforms: (a)Lfj0.at/gDLsinat at D1 acot1s a , (b)Lfy0.at/gdoes not exist. (c)Lfi0.at/gDLsinhat at D1 2alnsCa saD1 acoth1s a , (d)Lfk0.at/gdoes not exist. 20.8.15 Develop a Laplace transform solution of Laguerre’s equation, t F00.t/C.1t/F0.t/CnF.t/D0: Note that you need a derivative of a transform and a transform of derivatives. Go as far as you can with a general value of n; then (and only then) set nD0. 20.8.16 Show that the Laplace transform of the Laguerre polynomial Ln.at/is given by LfLn.at/gD.sa/n snC1;s>0: 20.8.17 Show that LfE1.t/gD1 sln.sC1/; s>0; where E1.t/D1Z te dD1Z 1ext xdx: E1.t/is the exponential integral function, first encountered in this book in Table 1.2. ArfKen_Ch20-9780123846549.tex 20.8 Properties of Laplace Transforms 1033 20.8.18 (a) From Eq. (20.184) show that 1Z 0f.x/dxD1Z 0F.t/ tdt; provided the integrals exist. (b) From the preceding result show that 1Z 0sint tdtD 2; in agreement with Eqs. (20.146) and (11.107). 20.8.19 (a) Show that Lsinkt t Dcot1s k : (b) Using this result (with kD1), prove that Lfsi.t/gD1 stan1s; where si.t/D1Z tsinx xdx;the sine integral: 20.8.20 IfF.t/is periodic (Fig. 20.19) with a period aso that F.tCa/DF.t/for all t0, show that LfF.t/gD1 1easaZ 0estF.t/dt: Note that the integration is now over only the first period ofF.t/. a 2a 3atf(t) FIGURE 20.19 Periodic function. ArfKen_Ch20-9780123846549.tex 1034 Chapter 20 Integral Transforms 20.8.21 Find the Laplace transform of the square wave (period a) defined by F.t/D(1;0<t<a=2; 0;a=2<t<a: ANS. f.s/D1 s1eas=2 1eas: 20.8.22 Show that (a)Lfcosh atcosatgDs3 s4C4a4; (b)Lfcosh atsinatgDas2C2a3 s4C4a4; (c)LfsinhatcosatgDas22a3 s4C4a4; (d)LfsinhatsinatgD2a2s s4C4a4: 20.8.23 Show that (a)L1n .s2Ca2/2o D1 2a3sinatt 2a2cosat, (b)L1n s.s2Ca2/2o Dt 2asinat, (c)L1n s2.s2Ca2/2o D1 2asinatCt 2cosat, (d)L1n s3.s2Ca2/2o Dcosatat 2sinat. 20.8.24 Show that Lf.t2k2/1=2u.tk/gD K0.ks/: Hint. Try transforming an integral representation of K0.ks/into the Laplace transform integral. 20.9 L APLACE CONVOLUTION THEOREM One of the most important properties of the Laplace transform is that given by the convo- lution, or Faltung, theorem. We take two transforms, f1.s/DLfF1.t/gand f2.s/DLfF2.t/g; and multiply them together: f1.s/f2.s/D1Z 0esxF1.x/dx1Z 0esyF2.y/dy: (20.185) ArfKen_Ch20-9780123846549.tex 20.9 Laplace Convolution Theorem 1035 If we introduce the new variable tDxCyand integrate over tandyinstead of xand y, the limits of integration become .0t1/ ,.0yt/. Noting that the Jacobian of the transformation from .x;y/to.t;y/is unity, we have f1.s/f2.s/D1Z 0estdttZ 0F1.ty/F2.y/dy DL8 < :tZ 0F1.ty/F2.y/dy9 = ; DLfF1F2g; (20.186) where, similarly to the Fourier transform, we use the notation tZ 0F1.tz/F2.z/dzF1F2; (20.187) and call this operation the convolution ofF1andF2. It can be shown that convolution is symmetric: F1F2DF2F1: (20.188) Carrying out the inverse transform, we also find L1ff1.s/f2.s/gDtZ 0F1.tz/F2.z/dzDF1F2: (20.189) Convolution formulas are useful for finding new transforms or, in some cases, as an alterna- tive to a partial fraction expansion. They also find use in the solution of integral equations, as is illustrated in Chapter 21. Example 20.9.1 DRIVEN OSCILLATOR WITH DAMPING As one illustration of the use of the convolution theorem, let us return to the mass mon a spring, with damping and a driving force F.t/. The equation of motion, Eq. (20.160), now becomes m X00.t/CbX0.t/Ck X.t/DF.t/: (20.190) Initial conditions X.0/D0,X0.0/D0are used to simplify this illustration, and the trans- formed equation is ms2x.s/Cbs x.s/Ck x.s/Df.s/; with solution x.s/Df.s/ m1 .sCb=2m/2C!2 1; (20.191) ArfKen_Ch20-9780123846549.tex 1036 Chapter 20 Integral Transforms where, as in Example 20.8.5, !2 0k m; !2 1!2 0b2 4m2: (20.192) We identify the right-hand side of Eq. (20.191) as the product of two known transforms: f.s/ mD1 mLfF.t/g;1 .sCb=2m/2C!2 1D1 !1Lfe.b=2m/tsin!1tg; where the second of these is a case of Eq. (20.158). Now applying the convolution theorem, Eq. (20.189), we obtain the solution to our original problem as an integral: X.t/DL1fx.s/gD1 m!1tZ 0F.tz/e.b=2m/zsin!1z dz: (20.193) We go on to consider two specific choices for the driving force F.t/. We first take the impulsive force F.t/DP.t/. Then X.t/DP m!1e.b=2m/tsin!1t: (20.194) Here Prepresents the momentum transferred by the impulse, and the constant P=mtakes the place of an initial velocity X0.0/. As a second case, let F.t/DF0sin!t. We could again use Eq. (20.193), but a partial fraction expansion is perhaps more convenient. With f.s/DF0! s2C!2 Eq. (20.191) can be written in the partial fraction form, x.s/DF0! m1 s2C!21 .sCb=2m/2C!2 1 DF0! m" a0sCb0 s2C!2Cc0sCd0 .sCb=2m/2C!2 1# ; (20.195) with coefficients a0,b0,c0, and d0(independent of s) to be determined. Direct calculation shows for a0andb0 1 a0Db m!2Cm b.!2 0!2/2; 1 b0Dm b.!2 0!2/b m!2Cm b.!2 0!2/2 : The terms of x.s/containing a0andb0lead upon inversion of the Laplace transform to the steady-state component of the solution: X.t/DF0 Tb2!2Cm2.!2 0!2/2U1=2sin.! t'/; (20.196) ArfKen_Ch20-9780123846549.tex 20.9 Laplace Convolution Theorem 1037 where tan'Db! m.!2 0!2/: Differentiating the denominator, we find that the amplitude has a maximum when !D!2, with !2 2D!2 0b2 2m2D!2 1b2 4m2: (20.197) This is the resonance condition.12At resonance the amplitude becomes F0=b!1, showing that the mass mgoes into infinite oscillation at resonance if damping is neglected .bD0/. This calculation differs from those used for the determination of transfer functions (com- pare Example 20.5.1) in that a steady-state solution at a fixed frequency is not assumed. Use of the Laplace transform (rather than the Fourier transform) permits solution for tran- sient as well as steady-state components of the solution. The transients, which we will not work out in detail, arise from the terms of Eq. (20.195) involving c0andd0. These terms contain the quantity .sCb=2m/2in the denominator, and its presence will generate terms of the inverse transform that contain the exponential factor ebt=2m. In other words, these terms describe exponentially decaying transients. It is worth noting that we have had three different characteristic frequencies: Resonance for forced oscillations with damping: !2 2D!2 0b2 2m2; Free oscillation frequency, with damping: !2 1D!2 0b2 4m2; Free oscillation frequency, no damping: !2 0Dk m: These frequencies coincide only if the damping is zero.  Recall that Eq. (20.190) is our ODE for the response of a dynamical system to an arbi- trary driving force. The final response clearly depends on both the driving force and the characteristics of our system. This dual dependence is separated in the transform space. In Eq. (20.191) the transform of the response (output) appears as the product of two factors, one describing the driving force (input) and the other describing the dynamical system. This is a factorization similar to that we found when discussing the use of Fourier trans- forms in signal-processing applications in Section 20.5. Exercises 20.9.1 From the convolution theorem show that 1 sf.s/DL8 < :tZ 0F.x/dx9 = ;; where f.s/DLfF.t/g. 12The amplitude (squared) has the typical resonance denominator (the Lorentz line shape), found in Exercise 20.2.8. ArfKen_Ch20-9780123846549.tex 1038 Chapter 20 Integral Transforms 20.9.2 IfF.t/DtaandG.t/Dtb,a>1,b>1, (a) Show that the convolution FGis given by FGDtaCbC11Z 0ya.1y/bdy: (b) By using the convolution theorem, show that 1Z 0ya.1y/bdyDaWbW .aCbC1/WDB.aC1;bC1/; where Bis the beta function. 20.9.3 Using the convolution integral, calculate L1s .s2Ca2/.s2Cb2/ ;a26Db2: 20.9.4 An undamped oscillator is driven by a force F0sin!t. Find the displacement X.t/as a function of time, subject to initial conditions X.0/DX0.0/D0. Note that the solution is a linear combination of two simple harmonic motions, one with the frequency of the driving force and one with the frequency !0of the free oscillator. ANS. X.t/DF0=m !2!2 0! !0sin!0tsin!t : 20.10 I NVERSE LAPLACE TRANSFORM Bromwich Integral We now develop an expression for the inverse Laplace transform L1appearing in the equation F.t/DL1ff.s/g: (20.198) One approach lies in the Fourier transform, for which we know the inverse relation. There is a difficulty, however. Our Fourier transformable function had to satisfy the Dirichlet conditions. In particular, we required that in order for g.!/to be a valid Fourier transform, lim!!1g.!/D0; (20.199) so that the infinite integral would be well defined.13Now we wish to treat functions F.t/ that may diverge exponentially. To surmount this difficulty, we extract an exponential factor, e t, from our (possibly) divergent F.t/and write F.t/De tG.t/: (20.200) 13We made an exception to deal with the delta function, but even in that case g.!/was bounded for all !. ArfKen_Ch20-9780123846549.tex 20.10 Inverse Laplace Transform 1039 IfF.t/diverges as e t, we require to be greater than so that G.t/will be convergent. Now, with G.t/D0fort<0and otherwise suitably restricted so that it may be represented by a Fourier integral, as in Eq. (20.22), we have G.t/D1 21Z 1eiutdu1Z 0G.v/eiuvdv: (20.201) Inserting Eq. (20.201) into Eq. (20.200), we have F.t/De t 21Z 1eiutdu1Z 0F.v/e veiuvdv: (20.202) We now make a change of variable to sD Ciu, causing the integral over vin Eq. (20.202) to assume the form of a Laplace transform: 1Z 0F.v/esvdvDf.s/: The variable sis now complex, but must be restricted to <e.s/ in order to guarantee convergence. Note that the Laplace transform has extended a function specified on the positive real axis onto the complex plane, <e s .14 We now need to rewrite Eq. (20.202) using the variable sin place of u. The range 1<u<1corresponds to a contour in the complex plane of s, which is a vertical line from i1to Ci1; we also need to substitute duDds=i. Making these changes, Eq. (20.202) becomes F.t/D1 2i Ci1Z i1estf.s/ds: (20.203) Here is our inverse transform. The path has become an infinite vertical line in the complex plane. Note that the constant was chosen so that f.s/would be nonsingular for s . It can be shown that the nonsingularity of f.s/extends to complex sprovided that<e s , so the integrand of Eq. (20.203) can have singularities only to the left of the integration path. See Fig. 20.20. The inverse transformation given by Eq. (20.203) is known as the Bromwich integral, although sometimes it is referred to as the Fourier-Mellin theorem orFourier-Mellin integral. This integral may now be evaluated by the regular methods of contour integration (Chapter 11). If t>0andf.s/is analytic except for isolated singularities (and no branch points), and is also small at large jsj, the contour may be closed by an infinite semicircle in the left half-plane that does not contribute to the integral. Then by the residue theorem (Section 11.8), F.t/DX .residues included for <e s< /: (20.204) 14For a derivation of the inverse Laplace transform using only real variables, see C. L. Bohn and R. W. Flynn, Real variable inversion of Laplace transforms: An application in plasma physics, Am. J. Phys. 46: 1250 (1978). ArfKen_Ch20-9780123846549.tex 1040 Chapter 20 Integral Transforms Possible singularities of est f(s)s-planeγ FIGURE 20.20 Possible singularities of estf.s/. It is worth mentioning that in many cases of interest f.s/may become large in the left half- plane or have branch points, and evaluation of the Bromwich integral may then present significant challenges. Possibly this means of evaluation with <e s ranging through negative values seems paradoxical in view of our previous requirement that <e s . The paradox disappears when we recall that the requirement <e s was imposed to guarantee convergence of the Laplace transform integral that defined f.s/. Once f.s/is obtained, we may then proceed to exploit its properties as an analytic function in the complex plane wherever we choose. Perhaps a pair of examples may clarify the evaluation of Eq. (20.203). Example 20.10.1 INVERSION VIA CALCULUS OF RESIDUES Iff.s/Da=.s2a2/, then the integrand for the Bromwich integral will be estf.s/Daest s2a2Daest .sCa/.sa/: (20.205) From the form of Eq. (20.205), we see that this integrand has poles at sDa, and the value of for the integral must be larger than jaj. Since these are simple poles, it is easy to verify that the residue at sDamust be eat=2, while the residue at sDawill beeat=2. The form of the integrand also permits us to close the contour in the left half-plane. We find, in accord with Eq. (20.204), ResiduesD1 2 .eateat/DsinhatDF.t/: (20.206) Equation (20.206) is in agreement with entry #7 of our table of Laplace transforms, Table 20.1.  ArfKen_Ch20-9780123846549.tex 20.10 Inverse Laplace Transform 1041 Example 20.10.2 MULTIREGION INVERSION Iff.s/D.1eas/=s, the Bromwich integral then has integrand estf.s/Dest1eas s ; (20.207) and the possibilities for closing the contour depend on the relative magnitudes of tanda. Considering first t>a, we may close the contour for the Bromwich integral in the left half-plane without changing its value. Our integrand is an entire function (analytic every- where in the finite s-plane; note that the sin the denominator cancels when the numerator is expanded in a Maclaurin series). Since no singularities are enclosed, we conclude that fort>a,F.t/D0. Fortin the range 0<t<a, a different situation is encountered. Expanding the inte- grand into the two terms est ses.ta/ s; we see that the first becomes small in the left half-plane (but large in the right half-plane), while the second terms behaves in an opposite fashion (large in the left half-plane, small in the right). The obvious solution is to use different contours for the two terms, each of which is individually singular, with a pole at sD0. We therefore close the contour for the first term in the left half-plane, but close that for the second term in the right half-plane. Since the vertical portion of the contour is at <e sD >0, we see that the integral of the first term encloses the singularity, while the integral of the second term does not. Therefore the first integral will have a value equal to the residue of the integrand at the singularity (this residue is 1), while the second integral will vanish. These contours are illustrated in Fig. 20.21. Finally, for t<0, the entire integrand becomes small in the right half-plane, the contour (for the entire integrand) surrounds no singularities, and the integral is zero. Summarizing these three cases, F.t/Du.t/u.ta/D8 < :0;t<0; 1;0<t<a; 0;t>a;(20.208) a step function of unit height and length a(Fig. 20.22).  FIGURE 20.21 Contours for Example 20.10.2. ArfKen_Ch20-9780123846549.tex 1042 Chapter 20 Integral Transforms t=aF(t) t1 FIGURE 20.22 Finite-length step function u.t/u.ta/. Two general comments may be in order. First, these two examples hardly begin to show the usefulness and power of the Bromwich integral. It is always available for inverting a complicated transform when the tables prove inadequate. Second, this derivation is not presented as a rigorous one. Rather, it is given more as a plausibility argument, although it can be made rigorous. The determination of the inverse transform is somewhat similar to the solution of a differential equation. It makes little difference how you get the inverse transform. Guess at it if you want. It can always be checked by verifying that LfF.t/gDf.s/: Two alternate derivations of the Bromwich integral are the subjects of Exercises 20.10.1 and(20.10.2). Exercises 20.10.1 Derive the Bromwich integral from Cauchy’s integral formula. Hint. Apply the inverse transform L1to f.s/D1 2ilim !1 Ci Z i f.z/ szdz; where f.z/is analytic for<e z . 20.10.2 Starting with 1 2i Ci1Z i1estf.s/ds; show that by introducing f.s/D1Z 0eszF.z/dz ArfKen_Ch20-9780123846549.tex 20.10 Inverse Laplace Transform 1043 we can convert our integral into the Fourier representation of a Dirac delta function. From this derive the inverse Laplace transform. 20.10.3 Derive the Laplace transformation convolution theorem by use of the Bromwich inte- gral. 20.10.4 Find L1s s2k2 (a) by a partial fraction expansion. (b) Repeat, using the Bromwich integral. 20.10.5 Find L1k2 s.s2Ck2/ (a) by using a partial fraction expansion. (b) Repeat using the convolution theorem. (c) Repeat using the Bromwich integral. ANS. F.t/D1coskt: 20.10.6 Use the Bromwich integral to find the function whose transform is f.s/Ds1=2. Note that f.s/has a branch point at sD0. The negative x-axis may be taken as a cut line. See Fig. 20.23. FIGURE 20.23 Contour for Exercise 20.10.6. ArfKen_Ch20-9780123846549.tex 1044 Chapter 20 Integral Transforms Hint. A portion of the path needed to close the contour will yield nonzero contributions to the contour integral. These will need to be taken into account to get the proper value for the Bromwich integral. ANS. F.t/D.t/1=2: 20.10.7 Show that L1n .s2C1/1=2o DJ0.t/ by evaluation of the Bromwich integral. Hint. Convert your Bromwich integral into an integral representation of J0.t/.Fig- ure 20.24 shows a possible contour. 20.10.8 Evaluate the inverse Laplace transform L1n .s2a2/1=2o by each of the following methods: (a) Expansion in a series and term-by-term inversion. (b) Direct evaluation of the Bromwich integral. (c) Change of variable in the Bromwich integral: sD.a=2/.zCz1/. 20.10.9 Show that L1lns s Dlnt ; where D0:5772:::is the Euler-Mascheroni constant. s=i s=−i FIGURE 20.24 A possible contour for the inversion of J0.t/. ArfKen_Ch20-9780123846549.tex Additional Readings 1045 20.10.10 Evaluate the Bromwich integral for f.s/Ds .s2Ca2/2: 20.10.11 Heaviside expansion theorem. If the transform f.s/may be written as a ratio f.s/Dg.s/ h.s/; where g.s/andh.s/are analytic functions, with h.s/having simple, isolated zeros at sDsi, show that F.t/DL1g.s/ h.s/ DX ig.si/ h0.si/esit: Hint. See Exercise 11.6.3. 20.10.12 Using the Bromwich integral, invert f.s/Ds2eks. Express F.t/DL1ff.s/gin terms of the (shifted) unit step function u.tk/. ANS. F.t/D.tk/u.tk/: 20.10.13 You have a Laplace transform: f.s/D1 .sCa/.sCb/;a6Db: Invert this transform by each of three methods: (a) Partial fractions and use of tables, (b) Convolution theorem, (c) Bromwich integral. ANS. F.t/Debteat ab;a6Db: Additional Readings Abramowitz, M., and I. A. Stegun, eds., Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables (AMS-55). Washington, DC: National Bureau of Standards (1972), reprinted, Dover (1974). Chapter 29 contains tables of Laplace transforms. Champeney, D. C., Fourier Transforms and Their Physical Applications. New York: Academic Press (1973). Fourier transforms are developed in a careful, easy-to-follow manner. Approximately 60% of the book is devoted to applications of interest in physics and engineering. Erdelyi, A., W. Magnus, F. Oberhettinger, and F. G. Tricomi, Tables of Integral Transforms, 2 vols. New York: McGraw-Hill (1954). This text contains extensive tables of Fourier sine, cosine, and exponential transforms, Laplace and inverse Laplace transforms, Mellin and inverse Mellin transforms, Hankel transforms, and other more specialized integral transforms. Hamming, R. W., Numerical Methods for Scientists and Engineers, 2nd ed. New York: McGraw-Hill (1973), reprinted, Dover (1987). Chapter 33 provides an excellent description of the fast Fourier transform. Hanna, J. R., Fourier Series and Integrals of Boundary Value Problems. Somerset, NJ: Wiley (1990). This book is a broad treatment of the Fourier solution of boundary value problems. The concepts of convergence and completeness are given careful attention. ArfKen_Ch20-9780123846549.tex 1046 Chapter 20 Integral Transforms Jeffreys, H., and B. S. Jeffreys, Methods of Mathematical Physics, 3rd ed. Cambridge: Cambridge University Press (1972). Krylov, V. I., and N. S. Skoblya, Handbook of Numerical Inversion of Laplace Transform (translated by D. Louvish). Jerusalem: Israel Program for Scientific Translations (1969). Lepage, W. R., Complex Variables and the Laplace Transform for Engineers . New York: McGraw-Hill (1961); Dover (1980). A complex variable analysis that is carefully developed and then applied to Fourier and Laplace transforms. It is written to be read by students, but intended for the serious student. McCollum, P. A., and B. F. Brown, Laplace Transform Tables and Theorems. New York: Holt, Rinehart and Winston (1965). Miles, J. W., Integral Transforms in Applied Mathematics. Cambridge: Cambridge University Press (1971). This is a brief but interesting and useful treatment for the advanced undergraduate. It emphasizes applications rather than abstract mathematical theory. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics. New York: McGraw-Hill (1953). Parseval’s relations are derived independently of the inverse Fourier transform in Section 4.8 of this comprehensive, but difficult text. Papoulis, A., The Fourier Integral and Its Applications . New York: McGraw-Hill (1962). This is a rigorous development of Fourier and Laplace transforms and includes extensive applications in science and engineering. Roberts, G. E., and H. Kaufman, Table of Laplace Transforms. Philadelphia: Saunders (1966). Sneddon, I. N., Fourier Transforms. New York: McGraw-HiII (1951), reprinted, Dover (1995). A detailed com- prehensive treatment, this book is loaded with applications to a wide variety of fields of modern and classical physics. Sneddon, I. N., The Use of Integral Transforms. New York: McGraw-Hill (1974). Written for students in science and engineering in terms they can understand, this book covers all the integral transforms mentioned in this chapter as well as in several others. Many applications are included. Titchmarsh, E. C., Introduction to the Theory of Fourier Integrals, 2nd ed. New York: Oxford University Press (1937). Van der Pol, B., and H. Bremmer, Operational Calculus Based on the Two-sided Laplace Integral, 3rd ed. Cambridge, UK: Cambridge University Press (1987). Here is a development based on the integral range 1 toC1, rather than the useful 0 to 1. Chapter V contains a detailed study of the Dirac delta function (impulse function). Wolf, K. B., Integral Transforms in Science and Engineering. New York: Plenum Press (1979). This book is a very comprehensive treatment of integral transforms and their applications. ArfKen_Ch21-9780123846549.tex CHAPTER 21 INTEGRAL EQUATIONS 21.1 I NTRODUCTION With the exception of the integral transforms of Chapter 20, we have for the most part been considering relations between an unknown function '.x/and one or more of its deriva- tives. We now proceed to investigate equations containing the unknown function within an integral. As with differential equations, we shall confine our attention to linear relations, which are called linear integral equations. These integral equations are classified in two ways: If the limits of integration are fixed, we call the equation a Fredholm equation; if one limit is variable, it is a Volterra equation. If the unknown function appears only under the integral sign, we label it first kind. If it appears both inside and outside the integral, it is labeled second kind. Here are some examples of these definitions. In each of the following equations, '.t/is an unknown function whose value we seek. K.x;t/, which we call the kernel, and f.x/ are assumed to be known. When f.x/D0, the equation is said to be homogeneous. This is a Fredholm equation of the first kind, f.x/DbZ aK.x;t/'.t/dt: (21.1) Next we have a Fredholm equation of the second kind, which is an eigenvalue equation withthe eigenvalue, '.x/Df.x/CbZ aK.x;t/'.t/dt: (21.2) 1047 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch21-9780123846549.tex 1048 Chapter 21 Integral Equations Here we have a Volterra equation of the first kind, f.x/DxZ aK.x;t/'.t/dtI (21.3) and a Volterra equation of the second kind, '.x/Df.x/CxZ aK.x;t/'.t/dt: (21.4) Why do we bother about integral equations? After all, the differential equations have done a rather good job of describing our physical world so far. However, there are several reasons for introducing integral equations. First, we have placed considerable emphasis on the solution of differential equations subject to particular boundary conditions. For instance, the boundary condition at rD 0determines whether the Neumann function Yn.r/is present when Bessel’s equation is solved. The boundary condition for r!1 determines whether In.r/is present in our solution of the modified Bessel equation. To the contrary, an integral equation relates the unknown function not only to its values at neighboring points (derivatives) but also to its values throughout a region, including the boundary. In a very real sense the boundary conditions are built into the integral equation rather than imposed at the final stage of the solution. It will be seen later in this section that if we construct an integral equation that is equivalent to a differential equation with its boundary conditions, the form of that integral equation depends on the boundary conditions. A second feature of integral equations is that their compact and completely self- contained form may turn out to be a more convenient or powerful formulation of a problem than a differential equation and its boundary conditions. Mathematical problems such as existence, uniqueness, and completeness may often be handled more easily and elegantly in integral form. And finally, whether or not we like it, there are problems, such as some diffusion and transport phenomena, that cannot be represented by differential equations. If we wish to solve such problems, we are forced to handle integral equations. Example 21.1.1 MOMENTUM REPRESENTATION IN QUANTUM MECHANICS The Schrödinger equation (in ordinary space representation) for a particle of mass msub- ject to a potential V.r/is Nh2 2mr2 .r/CV.r/ .r/DE .r/; (21.5) and we previously found, extending the 1-D result from Eq. (20.97), that in momentum space the equivalent equation (for the Coulomb potential in hartree atomic units) is k2 2m'.k/C1 .2/3=2Z4 jkk0j2'.k0/d3k0DE'.k/: (21.6) This is an integral-equation eigenvalue problem. Note that the kernel of Eq. (21.6) is a function of kk0; this functional dependence, which arises from the convolution theorem, ArfKen_Ch21-9780123846549.tex 21.1 Introduction 1049 is typical of an ordinary potential in which the direct-space wave function is multiplied by a function that depends only on position.  Transformation of a Differential Equation into an Integral Equation Often we find that we have a choice. The physical problem may be represented by a dif- ferential or an integral equation. Let us assume that we have the differential equation and wish to transform it into an integral equation. Starting with a linear second-order ordinary differential equation (ODE), y00CA.x/y0CB.x/yDg.x/; (21.7) with initial conditions y.a/Dy0; y0.a/Dy0 0; we integrate to obtain y0.x/DxZ aA.t/y0.t/dtxZ aB.t/y.t/dtCxZ ag.t/dtCy0 0: Integrating the first integral on the right by parts yields y0.x/DA.x/y.x/xZ ah B.t/A0.t/i y.t/dtCxZ ag.t/dtCA.a/y0Cy0 0: Integrating a second time, we obtain y.x/DxZ aA.t/y.t/dtxZ aduuZ ah B.t/A0.t/i y.t/dt CxZ aduuZ ag.t/dtCh A.a/y0Cy0 0i .xa/Cy0: (21.8) To transform this equation into a neater form, we use the relation xZ aduuZ af.t/dtDxZ af.t/dtxZ tduDxZ a.xt/f.t/dt: (21.9) ArfKen_Ch21-9780123846549.tex 1050 Chapter 21 Integral Equations Applying this result to Eq. (21.8), we obtain y.x/DxZ a A.t/C.xt/h B.t/A0.t/i y.t/dt CxZ a.xt/g.t/dtCh A.a/y0Cy0 0i .xa/Cy0: (21.10) If we now introduce the abbreviations K.x;t/D.tx/h B.t/A0.t/i A.t/; f.x/DxZ a.xt/g.t/dtCh A.a/y0Cy0 0i .xa/Cy0; Eq. (21.10) becomes y.x/Df.x/CxZ aK.x;t/y.t/dt; (21.11) which is a Volterra equation of the second kind. Note that f.x/in Eq. (21.11) has a form that includes the initial conditions from the original differential equation. Another method for obtaining an integral equation equivalent to a differential equa- tion plus its boundary conditions was presented in Section 10.1, where we found that the Green’s function for a differential equation appeared as the kernel of the equivalent integral equation. Example 21.1.2 LINEAR OSCILLATOR EQUATION Let’s find an integral equation equivalent to the linear oscillator equation y00C!2yD0 (21.12) with boundary conditions y.0/D0; y0.0/D1: This corresponds to Eq. (21.7) with A.x/D0; B.x/D!2; g.x/D0: Substituting into Eq. (21.10), we find that the integral equation becomes y.x/DxC!2xZ 0.tx/y.t/dt: (21.13) ArfKen_Ch21-9780123846549.tex 21.1 Introduction 1051 This integral equation, Eq. (21.13), is equivalent to the original differential equation plus the initial conditions. A check shows that each form is indeed satisfied by y.x/D .1=!/ sin!x. Let us reconsider the linear oscillator equation, Eq. (21.12), but now with the boundary conditions y.0/D0; y.b/D0: Since y0.0/is not given, we must modify the procedure. The first integration gives y0D!2xZ 0y dxCy0.0/: Integrating a second time and again using Eq. (21.9), we have yD!2xZ 0.xt/y.t/dtCxy0.0/: (21.14) To eliminate the unknown y0.0/, we now impose the condition y.b/D0. This gives !2bZ 0.bt/y.t/dtDby0.0/: Substituting this back into Eq. (21.14), we obtain y.x/D!2xZ 0.xt/y.t/dtC!2x bbZ 0.bt/y.t/dt: Now let us break the interval T0;bUinto two intervals,T0;xUandTx;bU. Since x b.bt/.xt/Dt b.bx/; we find y.x/D!2xZ 0t b.bx/y.t/dtC!2bZ xx b.bt/y.t/dt: (21.15) Finally, if we define the kernel K.x;t/D8 >< >:t b.bx/; t<x; x b.bt/; t>x;(21.16) ArfKen_Ch21-9780123846549.tex 1052 Chapter 21 Integral Equations we have y.x/D!2bZ 0K.x;t/y.t/dt; (21.17) a homogeneous Fredholm equation of the second kind. Our new kernel, K.x;t/, illustrated in Fig. 21.1, has some interesting properties. 1. It is symmetric, K.x;t/DK.t;x/. 2. It is continuous, in the sense that t b.bx/ tDxDx b.bt/ tDx: 3. Its derivative with respect to tisdiscontinuous. As tincreases through the point tDx, there is a discontinuity of [email protected];t/=@t. Comparing with the discussion in Section 10.1, we identify K.x;t/as the Green’s func- tion for this ODE with the specified boundary conditions. Note in particular Eq. (10.30), which corresponds exactly to what was found here.  The above example shows how the initial or boundary conditions play a decisive role in the conversion of a linear second-order ODE into an integral equation. Summarizing, If we have initial conditions (only one end of our interval), the differen- tial equation transforms into a Volterra integral equation. But if we have a boundary value problem (boundary conditions at both ends of our inter- val), the differential equation leads to a Fredholm-type integral equation with a kernel that will be the Green’s function appropriate to the given boundary conditions. In closing, we call attention to the fact that the reverse transformation (integral equation to differential equation) is not always possible. There exist integral equations for which no corresponding differential equation is known. xbtK(x, t ) FIGURE 21.1 Kernel, Eq. (21.16), for linear oscillator boundary-value problem. ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1053 Exercises 21.1.1 Starting with the ODE, integrate twice and derive the Volterra integral equation corre- sponding to (a) y00.x/y.x/D0I y.0/D0; y0.0/D1. ANS. yDxZ 0.xt/y.t/dtCx: (b) y00.x/y.x/D0I y.0/D1; y0.0/D1 . ANS. yDxZ 0.xt/y.t/dtxC1. Check your results with Eq. (21.11). 21.1.2 Starting with the given answers of Exercise 21.1.1, differentiate and recover the original ODEs and the boundary conditions. 21.1.3 Given'.x/DxxZ 0.tx/'.t/dt, solve this integral equation by converting it to an ODE (plus boundary conditions) and solving the ODE (by inspection). 21.1.4 Show that the homogeneous Volterra equation of the second kind .x/DxZ 0K.x;t/ .t/dt has no solution (apart from the trivial solution D0). Hint. Develop a Maclaurin expansion of .x/. Assume that .x/andK.x;t/are dif- ferentiable with respect to xas needed. 21.2 S OME SPECIAL METHODS It is well known that general methods are available both for differentiating functions and (compare Chapters 7 and 9) for solving linear differential equations, while there is no general direct method for evaluating integrals. Integrations are carried out using a variety of tools of limited applicability, and the process is ultimately one of pattern recognition and the application of experience. Similar observations apply to the solution of integral equations. We consider here some special methods that work when the integral equation under study has suitable characteristics. ArfKen_Ch21-9780123846549.tex 1054 Chapter 21 Integral Equations Integral-Transform Methods When the kernel of an integral equation (and its integration limits) match the specification of an integral transform for which we have an inversion formula, we can use that identifi- cation to solve the integral equation. Formulas based on four integral transforms are listed here for reference, in each case with f.x/a known function and '.x/to be determined. If our integral equation is f.x/D1p 2R1 1eixt'.t/dt;then its solution is '.x/D1p 21Z 1eixtf.t/dt (Fourier transform): (21.18) If our integral equation is f.x/DR1 0ext'.t/dt;then its solution is '.x/D1 2i Ci1Z i1extf.t/dt (Laplace transform). (21.19) If our integral equation is f.x/DR1 0tx1'.t/dt;then its solution is '.x/D1 2i Ci1Z i1xtf.t/dt (Mellin transform): (21.20) If our integral equation is f.x/DR1 0t'.t/J.xt/dt;then its solution is '.x/D1Z 0t f.t/J.xt/dt (Hankel transform): (21.21) Note that these formulas can also be applied “in reverse,” i.e., with '.x/known and f.x/ to be determined. This observation, however, is of somewhat limited utility since nothing significantly new appears for the inverse Fourier and Hankel transforms, while the integra- tion limits for the inverse Laplace and Mellin transforms make them unlikely to appear in an integral equation. Actually the usefulness of the integral transform technique extends a bit beyond these four rather specialized forms. We illustrate with two examples. Example 21.2.1 FOURIER TRANSFORM SOLUTION Let’s consider a Fredholm equation of the first kind with a kernel of the general type k.xt/, where kis a function (not a constant), f.x/D1Z 1k.xt/'.t/dt; (21.22) ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1055 in which'.t/is our unknown function. Assuming that the needed transforms exist, we apply the Fourier convolution theorem, Eq. (20.71), to obtain f.x/D1Z 1K.!/8.!/ei!xd!: (21.23) The functions K.!/and8.!/ are, respectively, the Fourier transforms of k.x/and'.x/. Taking the Fourier transform of both sides of Eq. (21.23), the formula for which is Eq. (21.18), we find K.!/8.!/D1 21Z 1f.x/ei!xdxDF.!/p 2; (21.24) where F.!/ is the Fourier transform of f.x/. Since8.!/ is the only unknown in Eq. (21.24), we may solve for it, obtaining 8.!/D1p 2F.!/ K.!/; (21.25) and, using the inverse Fourier transform, we have the solution to Eq. (21.22): '.x/D1 21Z 1F.!/ K.!/ei!xd!: (21.26) A rigorous justification of this result is presented by Morse and Feshbach (see Additional Readings). An extension of this transformation solution appears as Exercise 21.2.1.  Example 21.2.2 GENERALIZED ABEL EQUATION The generalized Abel equation is a Volterra equation of the first kind: f.x/DxZ 0'.t/ .xt/ dt;0< < 1; withf.x/known; '.t/unknown:(21.27) Taking the Laplace transform of both sides of this equation, we obtain Lff.x/gDL8 < :xZ 0'.t/ .xt/ dt9 = ;DLfx gLf'. x/g; ArfKen_Ch21-9780123846549.tex 1056 Chapter 21 Integral Equations the last step following by the Laplace convolution theorem, Eq. (20.186). Then, evaluating Lfx gfrom entry 3 of Table 20.1, Lf'. x/gDs1 Lff.x/g 0.1 /: (21.28) In principle, our integral equation is solved, since all that remains is to take the inverse transform of Eq. (21.28). A clever way of obtaining the inverse transform proceeds as follows, with its initial step being to divide Eq. (21.28) by s.1We get 1 sLf'. x/gDs Lff.x/g 0.1 /DLfx 1gLff.x/g 0. /0.1 /: Combining the gamma functions according to Eq. (13.23) and applying the Laplace con- volution theorem again, we discover that 1 sLf'. x/gDsin L8 < :xZ 0f.t/ .xt/1 dt9 = ;: Inverting with the aid of Entry 3 of Table 20.2, we get xZ 0'.t/dtDsin xZ 0f.t/ .xt/1 dt; and finally, by differentiating, we have the solution to our generalized Abel equation: '.x/Dsin d dxxZ 0f.t/ .xt/1 dt: (21.29)  Generating-Function Method Occasionally, the reader may encounter integral equations that involve generating func- tions. Suppose we have the admittedly special case, f.x/D1Z 1'.t/ .12xtCx2/1=2dt;1x1; (21.30) where f.x/is known and '.t/is to be determined. We note two important features: 1..12xtCx2/1=2generates the Legendre polynomials. 2.T1; 1Uis the orthogonality interval for the Legendre polynomials. 1This division converts s1 , which cannot be inverted when 0< < 1, into s , which is the transform of x 1=0. / . ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1057 These features make it possible to expand the denominator in Legendre polynomials, sug- gesting that it may be useful also to represent '.t/as an expansion in these same functions. Thus, we introduce the expansions 1 .12xtCx2/1=2D1X nD0Pn.t/xn; '.t/D1X mD0amPm.t/: Substituting these expansions into our integral equation, Eq. (21.30), f.x/D1X nD01X mD0amxn1Z 1Pn.t/Pm.t/dtD1X nD01X mD0amxn2nm 2nC1 D1X nD02an 2nC1xn: (21.31) If we now insert into Eq. (21.31) the Maclaurin series expansion for f.x/, f.x/D1X nD0f.n/.0/ nW; we may equate powers of x, reaching, for each n, f.n/.0/ nWD2an 2nC1; so the solution to our integral equation is '.t/D1X nD02nC1 2f.n/.0/ nWPn.t/: (21.32) Similar results may be obtained with other generating functions (see the list in Table 12.1). This technique of expanding in a series of special functions is always avail- able. It is worth a try whenever the expansion is possible (and convenient) and the interval is appropriate. Separable Kernel We consider here the special case that the kernel of our integral equation is separable, in the sense that K.x;t/DnX jD1Mj.x/Nj.t/; (21.33) where n, the upper limit of the sum, is finite. Such kernels are sometimes called degener- ate. Our class of separable kernels includes all polynomials and many of the elementary transcendental functions. For example, K.x;t/Dcos.tx/is separable: cos.tx/DcostcosxCsintsinx: ArfKen_Ch21-9780123846549.tex 1058 Chapter 21 Integral Equations Integral equations with separable kernels have the desirable property that they can be related to eigenvalue equations and permit the application of methods of linear algebra. Let’s consider a Fredholm equation of the second kind, Eq. (21.2), with a separable kernel of the form given in Eq. (21.33). Inserting this formula for K.x;t/and bringing the summation outside the integral, we have '.x/Df.x/CnX jD1Mj.x/bZ aNj.t/'.t/dt: (21.34) We now see that the integral with respect to twill for each jbe a constant (with values that are currently not known): bZ aNj.t/'.t/dtDcj: (21.35) Hence Eq. (16.71) becomes '.x/Df.x/CnX jD1cjMj.x/: (21.36) Once the constants cihave been determined, Eq. (21.36) will give us'.x/, the solution to our integral equation. Equation (21.36) further tells us that the form of '.x/will consist of f.x/plus a linear combination of the x-dependent factors in the separable kernel. We may find the ciby multiplying Eq. (21.36) byNi.x/and integrating to eliminate the x-dependence. Use of Eq. (21.35) yields ciDbiCnX jD1ai jcj; (21.37) where biDbZ aNi.x/f.x/dx; ai jDbZ aNi.x/Mj.x/dx: (21.38) It is perhaps helpful to write Eq. (21.37) in matrix form , with AD.ai j/: bDcAcD.1A/c; (21.39) or cD.1A/1b: (21.40) Equation (21.39) is equivalent to a set of simultaneous linear algebraic equations .1a11/c1a12c2a13c3D b1; a21c1C.1a22/c2a23c3D b2; (21.41) a31c1a32c2C.1a33/c3D b3;and so on: ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1059 If our integral equation is homogeneous, so f.x/D0, then bD0. To get a solution in that case, we set the determinant of the coefficients of ciequal to zero: j1AjD 0; (21.42) exactly as for any matrix eigenvalue problem. The roots of Eq. (21.42) yield our eigen- values. Substituting into .1A/cD0, we find the ciand then Eq. (21.36) gives our solution. Example 21.2.3 HOMOGENEOUS FREDHOLM EQUATION To illustrate this technique for determining eigenvalues and eigenfunctions of the homo- geneous Fredholm equation of the second kind, we consider '.x/D1Z 1.tCx/'.t/dt: (21.43) Writing the kernel of this equation as M1.x/N1.t/CM2.x/N2.t/, we have M1.x/D1; M2.x/Dx; N1.t/Dt; N2.t/D1: Using the notation of Eqs. (21.33) to(21.42), we find from Eq. (21.38): a11Da22D0; a12D2 3;a21D2I b1Db2D0: Equation (21.42), our secular equation, becomes2 12 3 2 1 D0: (21.44) Expanding, we obtain 142 3D0; Dp 3 2: (21.45) Substituting the eigenvalues Dp 3=2intoEq. (21.39), we have c1c2p 3D0: (21.46) Finally, with the choice c1D1,Eq. (21.36) gives the two solutions '1.x/Dp 3 2.1Cp 3x/; Dp 3 2; (21.47) 2This equation would look more like our usual secular equations if each row of the determinant were divided by . Then we would have the secular equation in a familiar form, but with 1=identified as the eigenvalue. ArfKen_Ch21-9780123846549.tex 1060 Chapter 21 Integral Equations '2.x/Dp 3 2.1p 3x/; Dp 3 2: (21.48) Since our equation is homogeneous, the normalization of '.x/is arbitrary.  If the kernel of an integral equation is not separable in the sense of Eq. (21.33), there is still the possibility that it may be approximated by a kernel that is separable. Then we can get the exact solution of an approximate equation, which we can treat as an approximation to the solution of the original equation. Exercises 21.2.1 The kernel of a Fredholm equation of the second kind, '.x/Df.x/C1Z 1K.x;t/'.t/dt; is of the form k.xt/.3Assuming that the required transforms exist, show that '.x/D1p 21Z 1F.t/eixtdt 1p 2K.t/: F.t/andK.t/are the Fourier transforms of f.x/andk.x/, respectively. 21.2.2 (a) The kernel of a Volterra equation of the first kind, f.x/DxZ 0K.x;t/'.t/dt; has the form k.xt/. Assuming that the required transforms exist, show that '.x/D1 2i Ci1Z i1F.s/ K.s/exsds; where F.s/andK.s/are, respectively, the Laplace transforms of f.x/andk.x/. (b) In terms of the notation of part (a), show that the Volterra equation of the second kind, '.x/Df.x/CxZ 0K.x;t/'.t/dt; 3This kernel and a range 0x<1are the characteristics of integral equations of the Wiener-Hopf type. Details will be found in Chapter 8 of Morse and Feshbach (1953); see the Additional Readings. ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1061 has solution '.x/D1 2i Ci1Z i1F.s/ 1K.s/exsds: 21.2.3 Using the Laplace transform solution (Exercise 21.2.2), solve (a)'.x/DxCxZ 0.tx/'.t/dt. ANS.'.x/Dsinx. (b)'.x/DxxZ 0.tx/'.t/dt. ANS.'.x/Dsinhx. Check your results by substituting back into the original integral equations. 21.2.4 Reformulate the equations of Example 21.2.1 for integrals on the range .0;1/using Fourier cosine transforms. 21.2.5 Given the Fredholm integral equation, ex2D1Z 1e.xt/2'.t/dt; apply the Fourier convolution technique of Example 21.2.1 to solve for'.t/. 21.2.6 Solve Abel’s equation, f.x/DxZ 0'.t/ .xt/ dt;0< < 1; by the following method: (a) Multiply both sides by .zx/ 1and integrate with respect to xover the range 0xz. (b) Reverse the order of integration and evaluate the integral on the right-hand side (with respect to x) by recognizing it as a beta function. Note. zZ tdx .zx/1 .xt/ DB.1 ; /D0. /0.1 /D sin : ArfKen_Ch21-9780123846549.tex 1062 Chapter 21 Integral Equations 21.2.7 Given the generalized Abel equation with f.x/D1, 1DxZ 0'.t/ .xt/ dt;0< < 1; solve for'.t/and verify that '.t/is a solution of the given equation. ANS.'.t/Dsin t 1. 21.2.8 A Fredholm equation of the first kind has a kernel e.xt/2: f.x/D1Z 1e.xt/2'.t/dt: Show that the solution is '.x/D1p1X nD0f.n/.0/ 2nnWHn.x/; in which Hn.x/is an nth-order Hermite polynomial. 21.2.9 Solve the integral equation f.x/D1Z 1'.t/ .12xtCx2/1=2dt;1x1; for the unknown function '.t/, if (a)f.x/Dx2s, (b)f.x/Dx2sC1. ANS. (a)'.t/D4sC1 2P2s.t/, (b)'.t/D4sC3 2P2sC1.t/. 21.2.10 Find the eigenvalues and eigenfunctions of '.x/D1Z 1.tx/'.t/dt: 21.2.11 Find the eigenvalues and eigenfunctions of '.x/D2Z 0cos.xt/'.t/dt: ANS.1D2D1 ; '. x/DAcosxCBsinx: ArfKen_Ch21-9780123846549.tex 21.2 Some Special Methods 1063 21.2.12 Find the eigenvalues and eigenfunctions of y.x/D1Z 1.xt/2y.t/dt: Hint. This problem may be treated by the separable-kernel method or by a Legendre expansion. 21.2.13 Use the separable-kernel technique to show that .x/DZ 0cosxsint .t/dt hasnosolution (apart from D0). Explain this result in terms of separability and symmetry. 21.2.14 Given'.x/D1Z 0.1Cxt/'.t/dt, solve for the eigenvalues and the eigenfunctions by the separable-kernel technique. 21.2.15 Knowing the form of the solutions of an integral equation can be a great advantage. For '.x/D1Z 0.1Cxt/'.t/dt; assume'.x/to have the form 1Cbx. Substitute into the integral equation. Integrate and solve for band. 21.2.16 The equation f.x/DbZ aK.x;t/'.t/dt has a degenerate kernel K.x;t/DPn iD1Mi.x/Ni.t/. (a) Show that this integral equation has no solution unless f.x/can be written as f.x/DnX iD1fiMi.x/; where the fiare constants. (b) Show that to any solution '.x/we may add .x/, provided that .x/is orthogo- nal to all Ni.x/: bZ aNi.x/ .x/dxD0for all i: ArfKen_Ch21-9780123846549.tex 1064 Chapter 21 Integral Equations 21.2.17 A Kirchhoff diffraction theory analysis of a laser leads to the integral equation v.r2/D ZZ K.r1;r2/v.r 1/d A: The unknown, v.r1/;gives the geometric distribution of the radiation field over one mirror surface; the range of integration is over the surface of that mirror. For square confocal spherical mirrors, the integral equation becomes v.x2;y2/Di eikb baZ aaZ ae.ik=b/.x1x2Cy1y2/v.x1;y1/dx1dy1; in which bis the centerline distance between the laser mirrors. This can be put in a somewhat simpler form by the substitutions kx2 i bD2 i;ky2 i bD2 i;andka2 bD2a2 bD 2: (a) Show that the variables separate and we get two integral equations. (b) Show that the new limits,  ; may be approximated by 1 for a mirror dimen- siona. (c) Solve the resulting integral equations. 21.3 N EUMANN SERIES Many and probably most integral equations cannot be solved by the specialized techniques of the preceding section. Here we develop a rather general technique for solving integral equations. The method, due largely to Neumann, Liouville, and Volterra, develops the unknown function '.x/as a power series in , whereis a given constant. The method is applicable whenever the series converges. We solve a linear integral equation of the second kind by successive approximations; let’s take as an example the Fredholm equation '.x/Df.x/CbZ aK.x;t/'.t/dt; (21.49) in which f.x/6D0. If the upper limit of the integral is a variable (Volterra equation), the following development will still hold, but with minor modifications. Let us make the following initial approximation to our unknown function: '.x/'0.x/Df.x/: (21.50) This choice is not mandatory. If you can make a better guess, go ahead and guess. The choice here is equivalent to saying that the term of the equation containing the integral is ArfKen_Ch21-9780123846549.tex 21.3 Neumann Series 1065 small relative to f.x/. To improve this first crude approximation, we feed '0.x/back into the integral in Eq. (21.49), getting '1.x/Df.x/CbZ aK.x;t/f.t/dt: (21.51) Substituting the new '1.x/back into Eq. (21.49), we obtain a second approximation to '.x/: '2.x/Df.x/CbZ aK.x;t1/f.t1/dt1 C2bZ abZ aK.x;t1/K.t1;t2/f.t2/dt2dt1: This process can be repeated indefinitely, defining after nsteps the nth order approximation 'n.x/DnX iD0iui.x/; (21.52) where u0.x/Df.x/ u1.x/DbZ aK.x;t1/f.t1/dt1 (21.53) u2.x/DbZ abZ aK.x;t1/K.t1;t2/f.t2/dt2dt1 un.x/DbZ abZ abZ aK.x;t1/K.t1;t2/K.tn1;tn/f.tn/dtndt1: We expect that our solution '.x/will be '.x/Dlimn!1'n.x/Dlimn!1nX iD0iui.x/; (21.54) provided that our infinite series converges. We may conveniently check the convergence by the Cauchy ratio test, Section 1.1, not- ing that jnun.x/jjnjjfjmaxjKjn maxjbajn; ArfKen_Ch21-9780123846549.tex 1066 Chapter 21 Integral Equations usingjfjmaxto represent the maximum value ofjf.x/jin the intervalTa;bUandjKjmax to represent the maximum value of jK.x;t/jin its domain in the xt-plane. A sufficient condition for convergence is jjj Kjmaxjbaj<1: (21.55) Note thatjun.max/j is being used as a comparison series. If it converges, our actual series must converge. If this condition is not satisfied, we may or may not have convergence, and a more sensitive test would be required to determine the convergence. Of course, even if the Neumann series diverges, there still may be a solution to our integral equation obtainable by another method. To gain more understanding of our iterative manipulation, we may find it helpful to rewrite the Neumann series solution, Eq. (21.54), in operator form. We start by rewriting Eq. (21.49) as 'DK'Cf; where Krepresents the integral operatorRb aK.x;t/TUdt. Solving symbolically for ', we obtain 'D.1K/1f: Binomial expansion leads to Eq. (21.54). The convergence of the Neumann series is a demonstration that the inverse operator .1K/1exists. Example 21.3.1 NEUMANN SERIES SOLUTION To illustrate the Neumann method, we consider the integral equation '.x/DxC1 21Z 1.tx/'.t/dt: (21.56) To start the Neumann series, we take '0.x/Dx: Then '1.x/DxC1 21Z 1.tx/tdtDxC1 21 3t31 2t2x 1 1DxC1 3: Substituting'1.x/back into Eq. (21.56),we get '2.x/DxC1 21Z 1.tx/tdtC1 21Z 1.tx/1 3dtDxC1 3x 3: ArfKen_Ch21-9780123846549.tex 21.3 Neumann Series 1067 Continuing this process of substituting back into Eq. (21.56), we obtain '3.x/DxC1 3x 31 32; and by mathematical induction (Section 1.4), '2n.x/DxCnX sD1.1/s13sxnX sD1.1/s13s: (21.57) Letting n!1 , we get '.x/D3 4xC1 4: (21.58) This solution can (and should) be checked by substituting back into the original equation, Eq. (21.56). It is interesting to note that our series converged easily even though Eq. (21.55) is not satisfied in this particular case. Actually Eq. (21.55) is a rather crude upper bound on . It can be shown that a necessary and sufficient condition for the convergence of our series solution is that jj<jej, whereeis the eigenvalue of smallest magnitude of the corresponding homogeneous equation (that with f.x/D0). For this particular example, eDp 3=2. Clearly,D1 2< e.  The technique illustrated by the Neumann series occurs in a number of contexts in quan- tum mechanics. For example, one approach to the calculation of time-dependent perturba- tions in quantum mechanics starts with the integral equation for the evolution operator U.t;t0/D1i NhtZ t0dt1V.t1/U.t1;t0/: (21.59) Iteration leads to U.t;t0/D1i NhtZ t0dt1V.t1/Ci Nh2tZ t0dt1t1Z t0dt2V.t1/V.t2/C: (21.60) The evolution operator is obtained as a series of multiple integrals of the perturbing poten- tialV.t/, closely analogous to the Neumann series, Eq. (21.52). A second and similar relationship between the Neumann series and quantum mechanics appears when the Schrödinger wave equation for scattering is reformulated as an integral equation. See Example 10.2.2. The first term in a Neumann series solution is the incident (unperturbed) wave. The second term is the first-order Born approximation, Eq. (10.51). The Neumann method may also be applied to Volterra integral equations of the second kind, corresponding to replacing the fixed upper limit bin Eq. (21.49) by a variable, x. In the Volterra case the Neumann series converges for all as long as the kernel is square integrable. ArfKen_Ch21-9780123846549.tex 1068 Chapter 21 Integral Equations Exercises 21.3.1 Using the Neumann series, solve (a)'.x/D12xZ 0t'.t/dt, ANS. (a)'.x/Dex2. (b)'.x/DxCxZ 0.tx/'.t/dt, (c)'.x/DxxZ 0.tx/'.t/dt. 21.3.2 Solve .x/DxC1Z 0.1Cxt/ .t/dt by each of the following methods: (a) The Neumann series technique, (b) The separable-kernel technique, (c) Educated guessing. 21.3.3 Solve '.x/D1C2xZ 0.xt/'.t/dt by each of the following methods: (a) Reduction to an ODE (find the boundary conditions), (b) The Neumann series, (c) The use of Laplace transforms. ANS.'.x/Dcoshx: 21.3.4 (a) In Eq. (21.59), take VDV0, independent of t. Without using Eq. (21.60), show thatEq. (21.59) leads directly to U.tt0/Dexp i Nh.tt0/V0 : (b) Repeat for Eq. (21.60) without using Eq. (21.59). ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1069 21.4 H ILBERT -SCHMIDT THEORY Symmetrization of Kernels The Hilbert-Schmidt theory deals with linear integral equations of the Fredholm type with symmetric kernels: K.x;t/DK.t;x/: (21.61) The symmetry is of great importance, both because we will find it leads to results parallel to those found for the Sturm-Liouville theory of differential equations, and also because many problems of physical relevance can be written as Fredholm integral equations with symmetric kernels. Before plunging into the theory, we note that some important nonsymmetric kernels can be symmetrized. If we have the equation '.x/Df.x/CbZ aK.x;t/.t/'.t/dt; (21.62) the total kernel is actually K.x;t/.t/, clearly not symmetric if K.x;t/alone is symmetric. However, if we multiply Eq. (21.62) byp.x/and substitute p .x/'.x/D .x/; we obtain .x/Dp .x/f.x/CbZ ah K.x;t/p .x/.t/i .t/dt; (21.63) with a symmetric total kernel K.x;t/p.x/.t/. Orthogonal Eigenfunctions We now focus on the homogeneous Fredholm equation of the second kind: '.x/DbZ aK.x;t/'.t/dt: (21.64) ArfKen_Ch21-9780123846549.tex 1070 Chapter 21 Integral Equations We assume that the kernel K.x;t/is symmetric and real. Perhaps one of the first questions we might ask about the equation is: “Does it make sense?” or more precisely, “Does an eigenvaluesatisfying this equation exist?” This question can be answered in the affir- mative. Courant and Hilbert (in their work cited in the Additional Readings, chapter III, section 4) show that if K.x;t/is continuous, there is at least one such eigenvalue and possibly an infinite number of them. It is useful to recognize that Eq. (21.64) represents a linear-operator eigenvalue problem: The integral on its right-hand side converts 'into (in general) some other function, which we can indicate symbolically by the equation .x/DbZ aK.x;t/'.t/dtK'.x/; (21.65) so our eigenvalue problem is K'.x/D1 '.x/: (21.66) We do not have to worry about the possibility that D0, since we can read directly from Eq. (21.64) that in that case the solution to our integral equation will be uniquely '.x/D0. The integral operator Kislinear, since it is obviously true that K a'1.x/Cb'2.x/ DaK'1.x/CbK'2.x/: In addition, if we define the scalar product as an integral on the range .a;b/: h j'ibZ a .x/'.x/dx; (21.67) we then see that our requirement that the kernel K.x;t/be real and symmetric will make Ka self-adjoint operator: h jK'iDbZ a .x/bZ aK.x;t/'.t/dt dxDbZ adtbZ adx K.t;x/ .x/ '.t/ DhK j'i: (21.68) The linearity and self-adjointness indicate that we can expect to confirm that Khas the key properties of self-adjoint operators, namely that its eigenvalues are real and (except in the case of degeneracy) its eigenvectors are orthogonal. While the above constitutes a complete demonstration of the orthogonality of our solutions to the homogeneous Fredholm equation, let’s confirm these properties more explicitly. ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1071 We can start from the two equations, 1 i'i.x/DbZ aK.x;t/'i.t/dt; (21.69) 1 j'j.x/DbZ aK.x;t/'j.t/dt: (21.70) If we multiply Eq. (21.69) by' j.x/andEq. (21.70) by' i.x/and then integrate with respect to x, the two equations become4 1 ibZ a' j.x/'i.x/dxDbZ abZ aK.x;t/' j.t/'i.x/dtdx; (21.71) 1 jbZ a' i.x/'j.x/dxDbZ abZ aK.x;t/' i.t/'j.x/dtdx: (21.72) Since we have demanded that K.x;t/be real and symmetric, we may take the com- plex conjugate of Eq. (21.72) and then interchange the roles of xandtin the integral, reaching 1  jbZ a'i.x/' j.x/dxDbZ abZ aK.x;t/'i.x/' j.t/dtdx: (21.73) Subtracting Eq. (21.73) from Eq. (21.71), we obtain 1 i1  j!bZ a' j.x/'i.x/dxD0: (21.74) Just as in our earlier derivation from Sturm-Liouville theory, we conclude that if iDjthe integral in Eq. (21.74) is necessarily nonzero; so 1= iD1= i, meaning that imust be real. But ifi6Dj, bZ a' i.x/'j.x/dxD0;  i6Dj; (21.75) proving orthogonality. The derivation can also be completed if K.x;t/is Hermitian, mean- ing that K.t;x/DK.x;t/. See Exercise 21.4.1. Since we are mostly concerned with real 4We assume that the necessary integrals exist. For an example of a simple pathological case, see Exercise 21.4.4. ArfKen_Ch21-9780123846549.tex 1072 Chapter 21 Integral Equations K, it is appropriate to assume also that 'is real, and for the remainder of this chapter we will often omit the complex conjugate asterisks that occur, for example, in Eq. (21.75). If the eigenvalue iisdegenerate,5the eigenfunctions for that particular eigenvalue may be orthogonalized by the Gram-Schmidt method (Section 5.2). Our orthogonal eigen- functions may, of course, be normalized, and we assume that this has been done. The result is bZ a' i.x/'j.x/dxDi j: (21.76) It can be shown that the eigenfunctions of our integral equations form a complete set,6 in the sense that if a function g.x/can be generated by the integral g.x/DZ K.x;t/h.t/dt; with h.t/a piecewise continuous function, then g.x/can be represented by a series of eigenfunctions, g.x/D1X nD1an'n.x/: (21.77) The series in Eq. (21.77) can be shown to converge uniformly and absolutely. Let us extend this to the kernel K.x;t/by asserting that K.x;t/D1X nD1an'n.t/; (21.78) andanDan.x/. Substituting into the original integral equation, Eq. (21.64), and using the orthogonality integral, we obtain 'i.x/Diai.x/: (21.79) Therefore, for our homogeneous Fredholm equation of the second kind, the kernel may be expressed in terms of the eigenfunctions and eigenvalues as K.x;t/D1X nD1'n.x/'n.t/ n: (21.80) Equation (21.80) is not actually a new result. In the Green’s function chapter, Section 10.1, we identified K.x;t/, there called G.x;t/, as the Green’s function appearing in Eq. (10.30), with the expansion given in Eq. (10.14). However, it is possible that the 5As for differential operators, if more than one distinct eigenfunction of Eq. (21.64) corresponds to the same eigenvalue, that eigenvalue is said to be degenerate. 6For a proof of this statement, see Courant and Hilbert (1953), chapter III, section 5, in the Additional Readings. ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1073 expansion given by Eq. (21.80) may not exist. As an illustration of the sort of pathological behavior that may occur, you are invited to apply this analysis to '.x/D1Z 0ext'.t/dt: Compare Exercise 21.4.4. It should be emphasized that this Hilbert-Schmidt theory is concerned with the establish- ment of properties of the eigenvalues (real) and eigenfunctions (orthogonality, complete- ness), properties that may be of great interest and value. The Hilbert-Schmidt theory does not solve the homogeneous integral equation for us any more than the Sturm-Liouville theory for differential equations solved the ODEs. The solutions of the integral equation come by the application of techniques such as were introduced in Sections 21.2 and21.3, or perhaps even by numerical methods. Inhomogeneous Integral Equation We now continue with the Hilbert-Schmidt theory by seeking solutions of the inhomoge- neous equation '.x/Df.x/CbZ aK.x;t/'.t/dt: (21.81) We assume that the solutions of the corresponding homogeneous integral equation are already known: 'n.x/DnbZ aK.x;t/'n.t/dt; (21.82) the solution'n.x/corresponding to the eigenvalue n. Note that at this point we are assum- ing nothing about ; it is a constant that has no specific relationship to the eigenvalues n of the homogeneous integral equation. We expand both '.x/andf.x/in terms of this set of eigenfunctions: '.x/D1X nD1an'n.x/(anunknown); (21.83) f.x/D1X nD1bn'n.x/(bnknown). (21.84) Substituting into Eq. (21.81), we obtain 1X nD1an'n.x/D1X nD1bn'n.x/CbZ aK.x;t/1X nD1an'n.t/dt: (21.85) ArfKen_Ch21-9780123846549.tex 1074 Chapter 21 Integral Equations By interchanging the order of integration and summation, we may evaluate the integral by Eq. (21.82), and we get 1X nD1an'n.x/D1X nD1bn'n.x/C1X nD1an'n.x/ n: (21.86) If we multiply by 'i.x/and integrate from xDatoxDb, the orthogonality of our eigen- functions leads to aiDbiCai i: (21.87) This can be rewritten as aiDbiC ibi: (21.88) We now multiply Eq. (21.88) by 'i.x/and sum over i, giving '.x/Df.x/C1X iD1'i.x/ ibi Df.x/C1X iD1'i.x/ ibZ af.t/'i.t/dt: (21.89) Here it is assumed that the eigenfunctions 'i.x/are normalized to unity. Note that if f.x/D0,there is no solution unless is equal to one of the i, thereby confirming that the homogeneous integral equation has only the solutions 'i.x/. In the event that for the inhomogeneous equation, Eq. (21.81), is equal to one of the eigenvalues pof the homogeneous equation, our solution, Eq. (21.89), blows up. It can be shown that the inhomogeneous equation then has no solution unless the coefficient bpvanishes, meaning that there is no solution unless the inhomogeneous term f.x/is orthogonal to the eigenfunction 'p. If the eigenvalue pis degenerate, there will be no solution unless f.x/is orthogonal to all the degenerate eigenfunctions. For the case that bpD0, we can return to Eq. (21.87), which then reduces for apto apDbpCapDap; (21.90) which gives no information about ap. Note that if bp6D0this equation cannot be satisfied, a signal that a solution cannot be obtained. Under the assumption that bpD0we now can rewrite Eq. (21.86), identifying its first two summations, respectively, as '.x/and f.x/, separating the final summation into the single term ap'p.x/plus a sum over all nother than p, thereby reaching '.x/Df.x/Cap'pCp1X iD10'i.x/ ipbZ af.t/'i.t/dt: (21.91) ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1075 In this solution the apremains as an undetermined constant,7and the prime indicates that iDpis to be omitted from the sum. It is of interest to relate Eq. (21.89) to what might be expected if we tried to develop a similar equation by Green’s-function methods. To do this, we start by rewriting Eq. (21.81) as an operator equation of the form K'.x/1 '.x/Df.x/ ; (21.92) where Kis the operator that we introduced in Eq. (21.65). Next, we note that, from Eq. (21.82), the 'nare eigenfunctions of Kwith eigenvalues 1= n: K'n.x/D'n.x/ n: (21.93) Then, applying Eq. (10.39), the Green’s function of the entire left-hand side of Eq. (21.92) will be (assuming 'is real): G.x;t/DX n'n.x/'n.t/ 1n1DX nn n'n.x/'n.t/ DX n'n.x/'n.t/C2X n'n.x/'n.t/ n D.tx/2X n'n.x/'n.t/ n: (21.94) To reach the last line of Eq. (21.94) we used the eigenfunction expansion of the delta function, Eq. (5.27). Applying this Green’s function to the right-hand side of Eq. (21.92), we get '.x/D1 bZ aG.x;t/f.t/dt D1 bZ a" .tx/2X n'n.x/'n.t/ n# f.t/dt Df.x/CX n'n.x/ nbZ a'n.t/f.t/dt; (21.95) which agrees with Eq. (21.89). 7This is like the inhomogeneous linear ODE. We may add to its solution any constant times a solution of the corresponding homogeneous ODE. ArfKen_Ch21-9780123846549.tex 1076 Chapter 21 Integral Equations Example 21.4.1 INHOMOGENEOUS FREDHOLM EQUATION Let’s seek solutions to the inhomogeneous Fredholm equation '.x/Dx3C1Z 1.tCx/'.t/dt; (21.96) for the twovaluesD1andDp 3=2. The corresponding homogeneous equation, treated in Example 21.2.3, has solutions only for the two eigenvalues p 3=2. In norma- lized form, they are: 1Dp 3 2; ' 1Dp 3 2 xC1p 3 I2Dp 3 2; ' 2Dp 3 2 x1p 3 : Taking firstD1, which is not an eigenvalue of the homogeneous equation, we have '.x/Dx3C2X iD1'i.x/ i11Z 1t3'i.t/dt Dx3Cp 3 2 xC1p 3 p 3 21p 3 21Z 1t3 tC1p 3 dt Cp 3 2 x1p 3 p 3 21p 3 21Z 1t3 t1p 3 dt Dx36 5.2xC1/: (21.97) Continuing now to Dp 3=2, we note that it is the eigenvalue 1of the homogeneous integral equation. That means the integral equation will have no solution unless h'1jfiD0. For the present problem, h'1jfiDp 3 21Z 1 xC1p 3 x3dx6D0; so our integral equation will have no solution for Dp 3=2. If in spite of this observation we attempted to generate a solution using Eq. (21.91), the function '.x/we obtained would not satisfy the integral equation, irrespective of the value we might choose to assign to ap. The immediate reason we cannot obtain a solution is that the integral 1Z 1.tCx/f.t/dtD1Z 1.tCx/t3dtD2 5 ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1077 evaluates to a quantity that cannot be represented as a linear combination of the eigenfunc- tions'iother than'p(in the present case, this means that 2/5 is not proportional to '2). There is therefore no way to add an additional component to f.x/to obtain a cancellation of the 2/5.  Exercises 21.4.1 In the Fredholm equation '.x/DbZ aK.x;t/'.t/dt; assume that the kernel K.x;t/is self-adjoint or Hermitian: K.x;t/DK.t;x/: Extend the analysis of the present section to show that (a) the eigenfunctions are orthogonal, in the sense that bZ a' m.x/'n.x/dxD0; m6Dn.m6Dn/: (b) the eigenvalues are real. 21.4.2 (a) Show that the eigenfunctions of Exercise 21.2.12 are orthogonal. (b) Show that the eigenfunctions of Exercise 21.2.14 are orthogonal. 21.4.3 Use the Hilbert-Schmidt method to solve the inhomogeneous integral equation '.x/DxC1 21Z 1.tCx/'.t/dt: The corresponding homogeneous integral equation was treated in Example 21.2.3. Note. The application of the Hilbert-Schmidt technique here is somewhat like using a shotgun to kill a mosquito, especially when the equation can be solved quickly by expanding in Legendre polynomials. 21.4.4 The Fredholm integral equation '.x/D1Z 0ext'.t/dt has an infinite number of solutions, of which one is '.x/Dx1=2; D1=2: Verify that this is a solution and that it is notnormalizable. ArfKen_Ch21-9780123846549.tex 1078 Chapter 21 Integral Equations Note. A basic reason for this anomalous behavior is that the range of integration is infinite, making this a “singular” integral equation. Note also that a series expansion of the kernel extwould permit a solution by the separable-kernel method (Section 21.2), except that the series is infinite. This observation is consistent with the fact that this integral equation has an infinite number of eigenvalues and eigenfunctions. 21.4.5 Given y.x/DxC1Z 0xt y.t/dtV (a) Determine y.x/as a Neumann series. (b) Find the range of for which your Neumann series solution is convergent. Com- pare with the value obtained from jjj Kjmax<1: (c) Find the eigenvalue and the eigenfunction of the corresponding homogeneous inte- gral equation. (d) By the separable-kernel method show that the solution is y.x/D3x 3: (e) Find y.x/by the Hilbert-Schmidt method. 21.4.6 In Exercise 21.2.11 it was found that the integral equation '.x/D2Z 0cos.xt/'.t/dt had (unnormalized) eigenfunctions cosxandsinx, both with eigenvalue iD1=. Show that the kernel of this integral equation has an expansion of the form K.x;t/D2X nD1'n.x/'n.t/ n: 21.4.7 The integral equation '.x/D1Z 0.1Cxt/'.t/dt has eigenvalues 1D0:7889 and2D15:211 . The corresponding eigenfunctions are '1D1C0:5352 xand'2D11:8685 x. (a) Show that these eigenfunctions are orthogonal over the interval T0;1U. (b) Normalize the eigenfunctions to unity. ArfKen_Ch21-9780123846549.tex 21.4 Hilbert-Schmidt Theory 1079 (c) Show that K.x;t/D'1.x/'1.t/ 1C'2.x/'2.t/ 2: ANS. (b)'1.x/D0:7831C0:4191 x, '2.x/D1:84033:4386 x. 21.4.8 An alternate form of the solution to the inhomogeneous integral equation, Eq. (21.81), is '.x/D1X iD1bii i'i.x/: (a) Derive this form without using Eq. (21.89). (b) Show that this form and Eq. (21.89) are equivalent. Additional Readings Bocher, M., An Introduction to the Study of Integral Equations, Cambridge Tracts in Mathematics and Mathe- matical Physics, No. 10. New York: Hafner (1960). This is a helpful introduction to integral equations. Byron, F. W., Jr., and R. W. Fuller, Mathematics of Classical and Quantum Physics. Reading, MA: Addison- Wesley (1969), reprinted, Dover (1992). The treatment of integral equations is rather advanced. Cochran, J. A., The Analysis of Linear Integral Equations. New York: McGraw-Hill (1972). This is a comprehen- sive treatment of linear integral equations intended for applied mathematicians and mathematical physicists. It assumes a moderate to high level of mathematical competence on the part of the reader. Courant, R., and D. Hilbert, Methods of Mathematical Physics, Vol. 1 (English edition). New York: Interscience (1953). This is one of the classic works of mathematical physics. Originally published in German in 1924, the revised English edition is an excellent reference for a rigorous treatment of integral equations, Green’s functions, and a wide variety of other topics on mathematical physics. Golberg, M. A., ed., Solution Methods of Integral Equations. New York: Plenum Press (1979). This is a set of papers from a conference on integral equations. The initial chapter is excellent for up-to-date orientation and a wealth of references. Kanval, R. P., Linear Integral Equations. New York: Academic Press (1971), reprinted, Birkhäuser (1996). This book is a detailed but readable treatment of a variety of techniques for solving linear integral equations. Morse, P. M., and H. Feshbach, Methods of Theoretical Physics. New York: McGraw-Hill (1953). Detailed, rigorous, and difficult. Muskhelishvili, N. I., Singular Integral Equations, 2nd ed. New York: Dover (1992). Stakgold, I., Green’s Functions and Boundary Value Problems. New York: Wiley (1979). ArfKen_Ch22-9780123846549.tex CHAPTER 22 CALCULUS OF VARIATIONS The calculus of variations deals with problems where we search for a function or curve, rather than a value of some variable, that makes a given quantity stationary, usually an energy or action integral. Because a function is varied, these problems are called varia- tional. Variational principles, such as those of D’Alembert, Lagrange, and Hamilton, have been developed in classical mechanics; Fermat’s principle (that of the shortest optical path) finds use in electrodynamics. Lagrangian variational techniques also occur in quantum me- chanics and field theory. Before plunging into this rather different branch of mathematical physics, let us summarize some of its uses in both physics and mathematics. 1.In existing physical theories: a. Unification of diverse areas of physics using energy as a key concept b. Convenience in analysis: Lagrange equations, Section 22.2 c. Elegant treatment of constraints, Section 22.4 2.Starting point for new, complex areas of physics and engineering. In general rela- tivity, the geodesic is taken as the minimum path of a light pulse or the free-fall path of a particle in curved Riemannian space. Variational principles appear in quantum field theory. Variational principles have been applied extensively in control theory. 3.Mathematical unification. Variational analysis provides a proof of the complete- ness of the Sturm-Liouville eigenfunctions, and can be used to establish bounds for the eigenvalues. Similar results follow for the eigenvalues and eigenfunctions in the Hilbert-Schmidt theory of integral equations. 22.1 E ULER EQUATION The calculus of variations typically involves problems in which a quantity to be minimized (or maximized) appears as a functional, meaning that it is a quantity whose argument(s) are themselves function(s), not just variable(s). As a simple, yet fairly general case, let J 1081 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch22-9780123846549.tex 1082 Chapter 22 Calculus of Variations be a functional of y, defined as JTyUDx2Z x1f y.x/;dy.x/ dx;x dx: (22.1) Here fis a fixed function of the three variables y,dy=dx, and x, while Jwill have a value dependent on the choice of y. The square-bracket notation is frequently used to remind the reader that Jis a functional. Because Jis given as an integral, its value depends on the behavior of y.x/throughout the entire range of x(here x1xx2). A typical problem in the calculus of variations is to find (usually subject to some constraints) a continuous and differentiable function y.x/that makes Jstationary relative to small changes in y anywhere (or everywhere) in its range of definition. These stationary values of Jwill in many problems be minima or maxima, but they can also be saddle points. The conditions of physical problems will normally require that variations in ybe restricted to those that preserve its continuity and differentiability It is convenient to introduce a notation that makes our discussions less cumbersome; we usually rewrite Eq. (22.1) in a notation with dy=dx denoted yxand with the arguments xandTyUsuppressed, and we indicate the variation in Jproduced by a (small) variation inyas JDx2Z x1f.y;yx;x/dx: (22.2) Note that we wrote rather than dor@; this distinction reminds us that the variation is that of a function (here y) rather than that of a variable. In visualizing the situation described by Eq. (22.2), it is helpful to think of y.x/as apath or curve connecting the values y.x1/andy.x2/; in fact, a common problem in the calculus of variations will be to determine y.x/subject to the constraint that y.x1/andy.x2/have specified values (and often subject to further constraints that may also be integrals). To illustrate the class of problems represented by Eq. (22.2), here are two simple examples: Determination of the minimum-energy configuration of a rope or chain of given length attached to fixed points at both ends, in the presence of a uniform gravitational field. Determination of the track between two points at different heights that will minimize the travel time of an object that, starting from rest, slides without friction along the track subject only to a uniform gravitational field (this is known as the brachistochrone problem). The problems here under consideration are much more difficult than typical minimiza- tions in differential calculus, where the minimum in a function can be found by comparing its values, say y.x/, at neighboring points (by looking at dy=dx). What we can do, instead, is to start by assuming the existence of an optimum path, i.e., a function y.x/for which J is stationary, and then compare Jfor our (unknown) optimum path with that obtained from neighboring paths, of which there are an infinite number. See Fig. 22.1. Even this strategy may sometimes fail, as there exist functionals Jfor which there is no optimum path. ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1083 y x2, y2 x1, y1δy x FIGURE 22.1 Neighboring paths. Restricting attention to functions y.x/for which the endpoints y.x1/andy.x2/are fixed, we consider a deformation of y.x/, called the variation ofyand denotedy. We describe yby introducing a new function, .x/, and a scale factor that controls the magnitude of the variation. The function .x/is arbitrary except for being continuous and differentiable, and, to keep the endpoints fixed, with .x1/D.x2/D0: (22.3) With these definitions, our path, now a function of , is y.x; /Dy.x;0/C .x/; (22.4) and we choose y.x;0/as the (unknown) path that will minimize J. Relative to y.x;0/, the variationyis then yD .x/: (22.5) Using Eq. (22.4), our formula for Jcan now be written J. /Dx2Z x1f y.x; /; yx.x; /;x dx; (22.6) and we see that we have reached a simpler formulation in which Jis now a function of rather than a functional ofy. This means that we now know how to optimize it.1 We proceed now to obtain a stationary value of Jby imposing the condition @J. / @  D0D0; (22.7) analogous to the vanishing of the derivative dy=dxin differential calculus. 1The arbitrary nature of the dependence of J. /on.x/will come into play later. ArfKen_Ch22-9780123846549.tex 1084 Chapter 22 Calculus of Variations Now, the dependence of the integral is contained in y.x; / and yx.x; /D .@=@x/y.x; /. Therefore,2 @J. / @ Dx2Z x1@f @y@y @ C@f @yx@yx @  dxD0: (22.8) From Eq. (22.4), @y.x; / @ D.x/[email protected]; / @ Dd.x/ dx; (22.9) so Eq. (22.8) becomes @J. / @ Dx2Z x1@f @y.x/C@f @yxd.x/ dx dxD0: (22.10) Integrating the second term by parts to get .x/as a common factor, we convert it to x2Z x1d.x/ dx@f @yxdxD.x/@f @yx x2 x1x2Z x1.x/d dx@f @yxdx: (22.11) The integrated part vanishes by Eq. (22.3), and Eq. (22.10) becomes @J. / @ Dx2Z x1@f @yd dx@f @yx .x/dxD0: (22.12) Equation (22.12), which must be satisfied for arbitrary .x/, is to be understood as a con- dition on y.x/. Occasionally we will see Eq. (22.12) multiplied by  , which gives, upon using.x/ Dy, JDx2Z x1@f @yd dx@f @yx y dxD0: (22.13) Equation (22.13) is to be solved for arbitrary ywithy.x1/Dy.x2/D0. We now take up the solution of Eq. (22.12). That equation can be satisfied for arbitrary .x/only if the bracketed expression forming the remainder of its integrand vanishes “al- most everywhere,” meaning everywhere except possibly at isolated points.3The condition for our stationary value is thus formally a partial differential equation (PDE), @f @yd dx@f @yxD0; (22.14) known as the Euler equation. Since the form of fis known, it will actually reduce (be- cause there is really only one independent variable, x) to an ordinary differential equation (ODE) for ywith boundary conditions at x1andx2. In that connection, it is important 2Note that yandyxare being treated as independent variables because they occur as different arguments of f. 3Compare the discussion of convergence in the mean, at Eq. (5.22). ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1085 to note that the derivative d=dx occurs in the Euler equation, and that it has a meaning distinct from the partial derivative @=@x. In particular, if fDf.y.x/;yx;x/, then d f=dx, which stands for the change in f(from all sources) due to a change in x, has the evaluation d f dxD@f @xC@f @ydy dxC@f @yxd2y dx2; where the last term has the form given because dyx=dxDd2y=dx2. Note that the first term on the right gives the explicit x-dependence of f; the second and third terms give its implicit x-dependence via yandyx. The Euler equation, Eq. (22.14), is a necessary, but by no means sufficient condition that there be a function y.x/that is continuous and differentiable on the range .x1;x2/and yields a stationary value of J.4A nice example of a lack of sufficiency is provided by the problem of determining stationary paths between points on the surface of a sphere (this example was provided by Courant and Robbins; see Additional Readings). The minimum- distance path from point A to point B on a spherical surface is the arc of a great circle, shown as Path 1 in Fig. 22.2. But Path 2 also satisfies the Euler equation. Path 2 is a maximum, but only if we demand that it be a great circle and then only if we make less than one circuit (as Path 2 plus ncomplete revolutions is also a solution). If the path is not required to be a great circle, any deviation from Path 2 will increase the length. This is hardly the property of a local maximum, and that illustrates why it is important to check solutions of the Euler equation to see if they satisfy the physical conditions of the given problem. Sometimes a problem admits a discontinuous solution that has physical relevance and will not be found by straightforward application of the Euler equation. An example is provided by the soap film of Example 22.1.3, where such a solution describes what happens if the film becomes unstable and breaks. Following are examples of the use of the Euler equation. 2A1B FIGURE 22.2 Stationary paths over a sphere. 4For a discussion of sufficiency conditions and the development of the calculus of variations as a part of mathematics, see the works by Ewing and Sagan in Additional Readings. ArfKen_Ch22-9780123846549.tex 1086 Chapter 22 Calculus of Variations Example 22.1.1 STRAIGHT LINE Perhaps the simplest application of the Euler equation is in the determination of the shortest distance between two points in the Euclidean xy-plane. Since the element of distance is dsDT.dx/2C.dy/2U1=2DT1Cy2 xU1=2dx; the distance Jmay be written as JDx2;y2Z x1;y1dsDx2Z x1T1Cy2 xU1=2dx: (22.15) Comparison with Eq. (22.2) shows that f.y;yx;x/D.1Cy2 x/1=2: Substituting into Eq. (22.14) and noting that @f=@yvanishes, we obtain d dx1 .1Cy2x/1=2 D0; or 1 .1Cy2x/1=2DC;a constant: This equation is satisfied if yxDa;a second constant: Integrating this expression for yx, we get yDaxCb; (22.16) which is the familiar equation for a straight line. The constants aandbare now chosen so that the line passes through the two points .x1;y1/and.x2;y2/. Hence the Euler equation predicts that the shortest5distance between two fixed points in Euclidean space is a straight line.  The generalization of this to curved four-dimensional space-time leads to the impor- tant concept of the geodesic in general relativity. A further discussion of geodesics is in Section 22.2. 5Technically, we have only found a y.x/of stationary J. By inspection of the solution, we easily determine the distance to be a minimum. ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1087 Example 22.1.2 OPTICAL PATH NEAR A BLACK HOLE We now wish to determine the optical path in an atmosphere where the velocity of light increases in proportion to the height yaccording tov.y/Dy=b, with b>0some parameter describing the light speed. So vD0atyD0;which simulates the conditions at the surface of a black hole, called its event horizon, where the gravitational force is so strong that the velocity of light goes to zero, thus trapping light. Our variational principle (Fermat’s principle) is that light will take the path of shortest travel time from .x1;y1/to.x2;y2/, namely 1tDZ dtDx2;y2Z x1;y1ds vDx2;y2Z x1;y1b ydsDbx2;y2Z x1;y1p dx2Cdy2 yDminimum. (22.17) The path is along a line defined by the relation between yandx. While we have in previous equations taken xto be the independent variable, there is no inherent requirement to do so, and our work on the present problem will be simplified if we choose yas the independent variable, and we write Eq. (22.17) in the form 1tDy2Z y1q x2yC1 ydy; (22.18) where xystands for dx=dy. Then our Euler equation will be @f @xd dy@f @xyD0; with f.x;xy;y/Dq x2yC1 y: Noting that@f=@xD0and differentiating @f=@xy, we have d dyxy yq x2yC1D0: This equation can be integrated, giving xy yq x2yC1DC1Dconstant; orxyDC1yq 1C2 1y2: Writing xyDdx=dyand separating dxanddyin this first-order ODE, we find the integral xZ dxDyZC1y dyq 1C2 1y2; which yields xCC2Dq 1C2 1y2 C1;or.xCC2/2Cy2D1 C2 1: ArfKen_Ch22-9780123846549.tex 1088 Chapter 22 Calculus of Variations 12y x FIGURE 22.3 Circular optical path in medium. Irrespective of the values of C1andC2, this light path is the arc of a circle whose center is on the line yD0, namely the event horizon. The actual path of light passing from .x1;y1/ to.x2;y2/will be on the circle through those points centered on yD0; the construction of the path can be performed geometrically as shown in Fig. 22.3. Note that light will not escape completely from the black hole with this model for v.y/unless x1Dx2(a path perpendicular to the event horizon). This example may be adapted to a mirage (Fata Morgana) in a desert with hot air near the ground and cooler air aloft (the index of refraction changes with height in cool vs. hot air). For the mirage problem, the relevant velocity law is v.y/Dv0y=b. In that case, the circular light path is no longer convex with center on the x-axis, but becomes concave.  Alternate Forms of Euler Equations Another form of the Euler equation, which is often useful (Exercise 22.1.1), is @f @xd dx fyx@f @yx D0: (22.19) In problems in which fDf.y;yx/, i.e., in which xdoes not appear explicitly, Eq. (22.19) reduces to d dx fyx@f @yx D0; (22.20) or fyx@f @yxDconstant: (22.21) Example 22.1.3 SOAP FILM As our next illustrative example, consider two parallel coaxial wire circles to be connected by a surface of minimum area that is generated by revolving a curve y.x/about the x-axis. ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1089 (x1, y1)(x2, y2)y ds y x FIGURE 22.4 Surface of rotation, soap-film problem. See Fig. 22.4. The curve is required to pass through fixed endpoints .x1;y1/and.x2;y2/. The variational problem is to choose the curve y.x/so that the area of the resulting surface will be a minimum. A physical situation corresponding to this problem is that of a soap film suspended between the wire circles. For the element of area shown in Fig. 22.4, d AD2y dsD2y.1Cy2 x/1=2dx: The variational equation is then JDx2Z x12y.1Cy2 x/1=2dx: Neglecting the 2, we identify f.y;yx;x/Dy.1Cy2 x/1=2: Since@f=@xD0, we may apply Eq. (22.20) and get y.1Cy2 x/1=2y y2 x .1Cy2x/1=2Dc1; which simplifies to y .1Cy2x/1=2Dc1: (22.22) Squaring, we get y2 1Cy2xDc2 1; ArfKen_Ch22-9780123846549.tex 1090 Chapter 22 Calculus of Variations which rearranges to .yx/1Ddx dyDc1q y2c2 1: (22.23) We note in passing that c1had better have a value that causes dy=dx to be real. Equa- tion (22.23) may be integrated to give xDc1cosh1y c1Cc2; and, solving for y, we have yDc1coshxc2 c1 : (22.24) Finally, c1andc2are determined by requiring the solution to pass through the points .x1;y1/and.x2;y2/. Our “minimum”-area surface is a special case of a catenary of revo- lution, or a catenoid.  Soap Film: Minimum Area This calculus of variations contains many pitfalls for the unwary. Remember, the Euler equation is a necessary condition, and assumes a differentiable solution. The sufficiency conditions are quite involved. Again, see the Additional Readings for details. Respect for some of these hazards may be developed by further considering the soap-film problem in Example 22.1.3, with .x1;y1/D.x0;1/,.x2;y2/D.Cx0;1/. We are therefore consider- ing a soap film stretched between two rings of unit radius at xDx0. The problem is to predict the curve y.x/assumed by the soap film. By referring to Eq. (22.24), we find that c2D0because our problem is symmetric about xD0. Then yDc1coshx c1 ; (22.25) and our endpoint conditions become c1coshx0 c1 D1: (22.26) If we take x0D1 2we obtain the following transcendental equation for c1: 1Dc1cosh1 2c1 : (22.27) We find that this equation has two solutions: c1D0:2350 , leading to a “deep” curve, and c1D0:8483 , leading to a “shallow” curve. Which curve is assumed by the soap film? ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1091 Before answering this question, consider the physical situation with the rings moved apart so that x0D1. Then Eq. (22.26) becomes 1Dc1cosh1 c1 ; (22.28) which has no real solutions. The physical significance is that as the unit-radius rings were moved out from the origin, a point was reached at which the soap film could no longer maintain the same horizontal force over each vertical section. Stable equilibrium was no longer possible. The soap film broke (irreversible process) and formed a circular film over each ring (with a total area of 2D6:2832:::). This is known as the Goldschmidt discon- tinuous solution to the soap-film problem. The next question is: How large may x0be and still give a real solution for Eq. (22.26)? Solving Eq. (22.26) forx0, x0Dc1cosh1.1=c 1/; (22.29) we find that x0will be real only for c11and that its maximum value is attained when dx0=dc1D0. A plot of x0vs.c1is shown in Fig. 22.5; it helps to explain the behavior we observed at x0D1 2. We see from the plot (and more precisely from Exercise 22.1.6) that the Euler equation has no solutions for x0>xmax, where xmax0:6627 , and that this x0 value occurs when c10:5524 . For values of x0smaller than xmax, there are solutions for two different values of c1, corresponding to the “deep” and “shallow” curves found earlier forx0D1 2. Returning to the question as to which solution of Eq. (22.26) describes the soap film, let us calculate the area corresponding to each solution. Using Eq. (22.22) to reach the last 0.250.5Deep curve Shallow curvex0 0.2 0.4 0.6 0.8 1 c1 FIGURE 22.5 Solutions of Eq. (22.26) for unit-radius rings at xDx0. ArfKen_Ch22-9780123846549.tex 1092 Chapter 22 Calculus of Variations member of the first line below, we have AD4x0Z 0y.1Cy2 x/1=2dxD4 c1x0Z 0y2dx D4c1x0Z 0 coshx c12 dxDc2 1 sinh2x0 c1 C2x0 c1 : (22.30) Forx0D1 2, Eq. (22.30) leads to c1D0:2350! AD6:8456; c1D0:8483! AD5:9917; showing that the former can at most be only a local minimum. A more detailed investiga- tion (compare Bliss, Additional Readings, chapter IV) shows that this surface is not even a local minimum. For x0D1 2, the soap film will be described by the shallow curve yD0:8483 coshx 0:8483 : This shallow catenoid (catenary of revolution) will be an absolute minimum for 0x0< 0:528 . However, for 0:528<x<0:6627 , its area is greater than that of the Goldschmidt discontinuous solution (6.2832) and it is only a relative minimum. See Fig. 22.6. 8.0 7.0 6.0 5.0 4.0Area 3.0Shallow curveDeep curve Goldschmidt discontinuous solution 2.0 1.0 00.1 0.2 0.3 0.4 x00.5 0.6 0.7 0.8 FIGURE 22.6 Catenoid area and that of the discontinuous solution of the soap-film problem (unit-radius rings at xDx0). ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1093 For an excellent discussion of both the mathematical problems and experiments with soap films, we refer to Courant and Robbins in Additional Readings. The larger message of this subsection is the extent to which one must use caution in accepting solutions of the Euler equations. Exercises 22.1.1 Fordy=dxyx6D0, show the equivalence of the two forms of Euler’s equation: @f @xd dx@f @yxD0 and @f @yd dx fyx@f @yx D0: 22.1.2 Derive Euler’s equation by expanding the integrand of J. /Dx2Z x1f y.x; /; yx.x; /;x dx in powers of . Note. The stationary condition is @J. /=@ D0, evaluated at D0. The terms quadratic in may be useful in establishing the nature of the stationary solution (maxi- mum, minimum, or saddle point). 22.1.3 Find the Euler equation corresponding to Eq. (22.14) iffDf.yxx;yx;y;x/, assuming thatyandyxhave fixed values at the endpoints of their interval of definition. ANS.d2 dx2@f @yxx d dx@f @yx C@f @yD0. 22.1.4 The integrand f.y;yx;x/of Eq. (22.2) has the form f.y;yx;x/Df1.x;y/Cf2.x;y/yx: (a) Show that the Euler equation leads to @f1 @y@f2 @xD0: (b) What does this imply for the dependence of the integral Jon the choice of path? 22.1.5 Show that the condition that JDZ f.x;y/dxhas a stationary value (a) leads to f.x;y/independent of yand (b) yields no information about any x-dependence. We get no (continuous, differentiable) solution. To be a meaningful variational problem, dependence on yor higher derivatives is essential. ArfKen_Ch22-9780123846549.tex 1094 Chapter 22 Calculus of Variations Note. The situation will change when constraints are introduced (compare to Exer- cise 22.4.6). 22.1.6 A soap film stretched between two rings of unit radius centered at x0will have its closest approach to the x-axis at xD0, with the distance from the axis given by c1, with x0andc1related by Eq. (22.26) orEq. (22.29). (a) Show that dc1=dx 0becomes infinite when x0sinh. x0=c1/D1, indicating that the soap film becomes unstable if x0is increased beyond the value satisfying this condition. (b) Show that the condition of part (a) is equivalent to x0 c1Dcothx0 c1 : (c) Solve the transcendental equation of part (b) to obtain the critical value of x0=c1 and show that the separate values of x0andc1are then approximately x00:6627 andc10:5524 . 22.1.7 A soap film is stretched across the space between two rings of unit radius centered atx0on the x-axis and perpendicular to the x-axis. Using the solution developed in Example 22.1.3, set up the transcendental equations for the condition that x0is such that the area of the curved surface of rotation equals the area of the two rings (Goldschmidt discontinuous solution). Solve for x0. 22.1.8 In Example 22.1.1, expand JTy.x; /U JTy.x;0/Uin powers of . The term linear in leads to the Euler equation and to the straight-line solution, Eq. (22.16). Investigate the 2term and show that the stationary value of J, the straight-line distance, is a minimum. 22.1.9 (a) Show that the integral JDx2Z x1f.y;yx;x/dx;with fDy.x/; hasnoextreme values. (b) If f.y;yx;x/Dy2.x/;find a discontinuous solution similar to the Goldschmidt solution for the soap-film problem. 22.1.10 Fermat’s principle of optics states that a light ray in a medium for which nis the (position-dependent) index of refraction will follow the path y.x/for which x2;y2Z x1;y1n.y;x/ds is a minimum. For y2Dy1D1,x1Dx2D1, find the ray path if (a)nDey, (b) nDa.yy0/,y>y0. 22.1.11 A particle moves, starting at rest, from point Aon the surface of the Earth to point B (also on the surface) by sliding frictionlessly through a tunnel. Find the differential ArfKen_Ch22-9780123846549.tex 22.1 Euler Equation 1095 equation satisfied by the path if the transit time is to be a minimum. Assume the Earth to be a nonrotating sphere of uniform density. Hint. The potential energy of a particle of mass ma distance r<Rfrom the center of the Earth, with Rthe Earth’s radius, is1 2mg.R2r2/=R, where gis the gravitational acceleration at the Earth’s surface. It is convenient to describe the path of the particle (in the plane through A,B, and the center of the Earth) by plane polar coordinates .r;/, with Aat.R;'/ andBat.R;'/. ANS. Letting r0be the minimum value of r(reached atD0), Eq. (22.21) yields r2 Dr2R2.r2r2 0/ r2 0.R2r2/(the constant in that equation has the value such that rD0atD0). The solution for the path is a hypocycloid, generated by a circle of radius1 2.Rr0/ rolling inside the circle of radius R. You might like to show that the transit time is tD.R2r2 0/1=2 .Rg/1=2: For details see P. W. Cooper, Am. J. Phys. 34: 68 (1966); G. Veneziano, et al., 34: 701 (1966). 22.1.12 A ray of light follows a straight-line path in a first homogeneous medium, is refracted at an interface, and then follows a new straight-line path in the second medium. See Fig. 22.7. Use Fermat’s principle of optics to derive Snell’s law of refraction: n1sin1Dn2sin2: Hint. Keep the points .x1;y1/and.x2;y2/fixed and vary x0to satisfy Fermat’s principle. Note. This is notan Euler equation problem, because the light path is not differentiable atx0. n1 θ1 θ2(x0, 0) (x2, y2)(x1, y1) n2x FIGURE 22.7 Snell’s law. ArfKen_Ch22-9780123846549.tex 1096 Chapter 22 Calculus of Variations 22.1.13 A second soap-film configuration for the unit-radius rings at xDx0consists of a circular disk, radius a, in the xD0plane and two catenoids of revolution, one joining the disk and each ring. One catenoid may be described by yDc1coshx c1Cc3 : (a) Impose boundary conditions at xD0andxDx0. (b) Although not necessary, it is convenient to require that the catenoids form an angle of 120where they join the central disk. Express this third boundary condition in mathematical terms. (c) Show that the total area of catenoids plus central disk is then ADc2 1 sinh2x0 c1C2c3 C2x0 c1 : Note. Although this soap-film configuration is physically realizable and stable, the area is larger than that of the simple catenoid for all ring separations for which both films exist. ANS. (a)8 >< >:1Dc1coshx0 c1Cc3 aDc1cosh c3(b)dy dxDtan 30Dsinhc3. 22.1.14 For the soap film described in Exercise 22.1.13, find (numerically) the maximum value ofx0. Note. This calls for a calculator with hyperbolic functions or a table of hyperbolic cotan- gents. ANS. x0 maxD0:4078: 22.1.15 Find the curve of quickest descent from .0;0/to.x0;y0/for a particle that, starting from rest, slides under gravity and without friction. Show that the ratio of times taken by the particle along a straight line joining the two points compared to along the curve of quickest descent is .1C4=2/1=2. Hint. Take yto increase downwards. Apply Eq. (22.21) to obtain y2 xD.1c2y/=c2y, where cis an integration constant. It is helpful to make the substitution c2yDsin2'=2 and take.x0;y0/D.=2c2;1=c2/. 22.2 M OREGENERAL VARIATIONS Several Dependent Variables To apply variational methods to classical mechanics, we need to generalize the Euler equa- tion to situations in which there is more than one dependent variable in roles like yin ArfKen_Ch22-9780123846549.tex 22.2 More General Variations 1097 Eq. (22.2). The generalization corresponds to functionals Jof the form JDx2Z x1f u1.x/;u2.x/;:::; un.x/;u1x.x/;u2x.x/;:::; unx.x/;x dx: (22.31) We are now calling the dependent variables uito be consistent with notations we will shortly introduce, and as before we use the subscript xto denote differentiation with respect tox, so that uixDdui=dx and (later)ixDdi=dx. As in Section 22.1, we determine stationary values of Jby comparing neighboring paths for each ui. Let ui.x; /Dui.x;0/C i.x/; iD1;2;:::; n; (22.32) with theiindependent of one other but subject to the continuity and endpoint restrictions discussed in Section 22.1. By differentiating Jfrom Eq. (22.31) with respect to and setting D0(the condition that Jbe stationary), we obtain x2Z x1X i@f @uiiC@f @uixix dxD0: (22.33) Again, each of the terms .@f=@uix/ixis integrated by parts. The integrated part vanishes and Eq. (22.33) becomes x2Z x1X i@f @uid dx@f @uix idxD0: (22.34) Since theiare arbitrary and independent of one another,6each of the terms in the sum must vanish independently. We have @f @uid dx@f @uixD0; iD1;2;:::; n; (22.35) a whole set of Euler equations, each of which must be satisfied for a stationary value of J. Hamilton’s Principle The most important application of Eq. (22.31) occurs when the integrand fis taken to be a Lagrangian L. The Langrangian (for nonrelativistic systems; see Exercise 22.2.5 for a relativistic particle) is defined as the difference of kinetic and potential energies of a system: LTV: (22.36) Using time as an independent variable instead of xandxi.t/as the dependent variables, our conversion of Eq. (22.31) involves the replacements x!t; yi!xi.t/; yix!Pxi.t/I 6For example, we could set 2D3D4D 0, eliminating all but one term of the sum, and then treat 1exactly as in Section 22.1. ArfKen_Ch22-9780123846549.tex 1098 Chapter 22 Calculus of Variations xi.t/is the position andPxiDdxi=dtis the velocity of particle ias a function of time. The equation JD0is then a mathematical statement of Hamilton’s principle of classical mechanics, t2Z t1L.x1;x2;:::; xn;Px1;Px2;:::;PxnIt/dtD0: (22.37) In words, Hamilton’s principle asserts that the motion of the system from time t1tot2 is such that the time integral of the Lagrangian L, or action, has a stationary value. The resulting Euler equations are usually called the Lagrangian equations of motion, d dt@L @Pxi@L @xiD0 (each i): (22.38) These Lagrangian equations can be derived from Newton’s equations of motion, and New- ton’s equations can be derived from Lagrange’s. The two sets of equations are equally “fundamental.” The Lagrangian formulation has advantages over the conventional Newtonian laws. Whereas Newton’s equations are vector equations, we see that Lagrange’s equations involve only scalar quantities. The coordinates x1;x2;::: need not be a standard set of coordinates or lengths. They can be selected to match the conditions of the physical prob- lem. The Lagrange equations are invariant with respect to the choice of coordinate system. Newton’s equations (in component form) are not manifestly invariant. For example, Exer- cise 3.10.27 shows what happens when FDmais resolved in spherical polar coordinates. Exploiting the concept of energy, we may easily extend the Lagrangian formulation from mechanics to diverse fields, such as electrical networks and acoustical systems. Extensions to electromagnetism appear in the exercises. The result is a unification of otherwise sep- arate areas of physics. In the development of new areas, the quantization of Lagrangian particle mechanics provided a model for the quantization of electromagnetic fields and led to the gauge theory of quantum electrodynamics. One of the most valuable advantages of Hamilton’s principle (the Lagrange equation formulation) is the ease in seeing a relation between a symmetry and a conservation law. As an example, let xiD', an azimuthal angle. If our Lagrangian is independent of '(that is, 'is said to be an ignorable coordinate), there are two consequences: (1) the conservation or invariance of the component of angular momentum associated with (conjugate to) ', and (2) from Eq. (22.38), @L=@P'Dconstant. Similarly, invariance under translation leads to conservation of linear momentum. Example 22.2.1 MOVING PARTICLE, CARTESIAN COORDINATES A particle of mass mmoves in one dimension with its position described by a Cartesian coordinate x, subject to a potential V.x/. Its kinetic energy is given by TDmPx2=2, so its Lagrangian Lhas the form LDTVD1 2mPx2V.x/: ArfKen_Ch22-9780123846549.tex 22.2 More General Variations 1099 We will need @L @PxDmPx;@L @xDdV.x/ dxDF.x/: (22.39) We have identified the force Fas the negative gradient of the potential. Inserting the results from Eq. (22.39) into the Lagrangian equation of motion, Eq. (22.38), we get d dt.mPx/F.x/D0; which is Newton’s second law of motion.  Example 22.2.2 MOVING PARTICLE, CIRCULAR CYLINDRICAL COORDINATES Now let us consider a particle of mass mmoving in the xy-plane, that is, zD0. We use cylindrical coordinates ;'. The kinetic energy is TD1 2m.Px2CPy2/D1 2m.P2C2P'2/; (22.40) and we take VD0for simplicity. We could have converted Px2CPy2into circular cylindrical coordinates by taking x.;'/Dcos',y.;'/Dsin', and then differentiating with respect to time and squaring. What we actually did was to recognize that the cylindrical coordinates are an orthogonal system with scale factors hD1,h'D, so the velocity vhas in the cylindri- cal system components vDPandv'DP'. We now apply the Lagrangian equations of motion first to the coordinate and then to': d dt.mP/mP'2D0;d dt.m2P'/D0: The second equation is a statement of conservation of angular momentum. The first may be interpreted as radial acceleration7equated to centrifugal force. In this sense the centrifugal force is a real force. It is of some interest that this interpretation of centrifugal force as a real force is supported by the general theory of relativity.  Hamilton’s Equations Hamilton was the first to show that Euler’s equation for the Lagrangian enabled the equa- tions of motion to be reduced to the set of coupled first-order PDEs called Hamilton’s equations. A starting point for this analysis is the definition of the canonical momentum piconjugate to the coordinate qi, defined as piD@L @Pqi: (22.41) 7Here is a second method of attacking Exercise 3.10.13. ArfKen_Ch22-9780123846549.tex 1100 Chapter 22 Calculus of Variations This definition is consistent with the elementary definition of momentum in Cartesian coordinates, where (in one dimension) TDmPq2=2,pDmPq. From Eq. (22.41) and the Lagrangian equations of motion, Eq. (22.38), we have by direct substitution PpiD@L @qi; (22.42) and this permits us to write the variation of Lin the form d LDX i@L @qidqiC@L @PqidPqi C@L @tdtDX i.PpidqiCpidPqi/C@L @tdt:(22.43) We now define the Hamiltonian as HDX ipiPqiL; (22.44) and compute d HDX i.pidPqiCPqidpi/ X i.PpidqiCpidPqi/C@L @tdt! DX i.PqidpiPpidqi/@L @tdt: (22.45) But from the chain rule for differentiation, we also have d HDX i@H @pidpiC@H @qidqi C@H @tdt: (22.46) Equating the coefficients of dpi,dqi, and dtin Eqs. (22.45) and (22.46), we obtain Hamil- ton’s equations: @H @piDPqi;@H @qiDPpi;@H @tD@L @t: (22.47) In conservative systems, @H=@tD0, and Hhas a constant value equal to the total energy of the system. Several Independent Variables Sometimes the integrand fin an equation analogous to Eq. (22.2) will contain an unknown function, u, that is a function of several independent variables, uDu.x;y;z/. In the three- dimensional case, for example, that equation becomes JDZZZ f u;ux;uy;uz;x;y;z dx dy dz; (22.48) where uxD@u=@x,uyD@u=@y,uzD@u=@z, and uis assumed to have specified values on the boundary of the region of integration. Generalizing the analysis of Section 22.1, we represent the variation of uas u.x;y;z; /Du.x;y;z;0/C .x;y;z/; ArfKen_Ch22-9780123846549.tex 22.2 More General Variations 1101 whereis arbitrary except that it must vanish on the boundary. Our integral Jis now, as in Section 22.1, a function of , and our variational problem is to make Jstationary with respect to . Differentiating the integral Eq. (22.48) with respect to the parameter and then setting D0, we obtain @J @ D0DZZZ@f @uC@f @uxxC@f @uyyC@f @uzz dx dy dzD0: We continue to use a notation similar to that used previously: xis shorthand for @=@ x, etc. Again, we integrate each of the terms .@f=@ui/iby parts. The integrated part vanishes at the boundary (because the deviation is required to go to zero there) and we get ZZZ@f @u@ @x@f @ux@ @y@f @uy@ @z@f @uz .x;y;z/dx dy dzD0: (22.49) We must now digress to clarify the notation in Eq. (22.49). The derivative @=@xenters that equation as a result of the integration by parts, and it therefore must act on all the x dependence of @f=@ux, not just on the explicit appearance of xinf. The reader may recall that this derivative was written d=dxwhen it arose in Section 22.1, but that notation is not entirely appropriate here as the functions involved also depend on yandz. We conclude our analysis with the now-familiar observation that since the variation .x;y;z/is arbitrary, the term in large parentheses is set equal to zero. This yields the Euler equation for (three) independent variables, @f @u@ @x@f @ux@ @y@f @uy@ @z@f @uzD0: (22.50) Remember that the derivative @=@xoperates on both the explicit and implicit xdependence of@f=@ux; similar remarks apply to @=@yand@=@z. Example 22.2.3 LAPLACE’S EQUATION A variational problem with several independent variables is provided by electrostatics. An electrostatic field has energy densityD1 2"E2; where Eis the electric field. In terms of the static potential ', energy densityD1 2".r'/2: Now let us impose the requirement that the electrostatic energy (associated with the field) in a given charge-free volume be a minimum subject to specific conditions on 'at the boundary. The assumption that the volume is charge-free makes 'continuous and dif- ferentiable throughout the volume, and we therefore have a situation to which an Euler equation applies. We have the volume integral JDZZZ .r'/2dx dy dzDZZZ .'2 xC'2 yC'2 z/dx dy dz; ArfKen_Ch22-9780123846549.tex 1102 Chapter 22 Calculus of Variations where'xstands for@'=@ x. Thus, f.';' x;'y;'z;x;y;z/D'2 xC'2 yC'2 z; so Euler’s equation, Eq. (22.50), yields (with uin that equation replaced by ') 2.' xxC'yyC'zz/D0; which in the usual vector notation is equivalent to r2'.x;y;z/D0: This is Laplace’s equation of electrostatics. Closer investigation shows that this stationary value is indeed a minimum. Thus the demand that the field energy be minimized leads to Laplace’s PDE.  Several Dependent and Independent Variables In some cases our integrand fcontains more than one dependent variable and more than one independent variable. Consider fDf p.x;y;z/;px;py;pz;q.x;y;z/;qx;qy;qz;r.x;y;z/;rx;ry;rz;x;y;z : (22.51) We proceed as before with p.x;y;z; /Dp.x;y;z;0/C .x;y;z/; q.x;y;z; /Dq.x;y;z;0/C .x;y;z/; r.x;y;z; /Dr.x;y;z;0/C .x;y;z/;and so on: Keeping in mind that ;, andare independent of one another, as were the iin Eq. (22.32), the same differentiation and then integration by parts will lead to @f @p@ @x@f @px@ @y@f @py@ @z@f @pzD0; (22.52) with similar equations for functions qandr. Replacing p,q,r,:::with yiandx,y,z,::: with xi, we can put Eq. (22.52) in a more compact form: @f @yiX j@ @xj@f @yi j D0; iD1;2;:::; (22.53) in which yi j@yi @xj: An application of Eq. (22.53) appears in Exercise 22.2.10. ArfKen_Ch22-9780123846549.tex 22.2 More General Variations 1103 Geodesics Particularly in general relativity, it is of interest to identify the shortest path between two points in a “curved space,” i.e., a space characterized by a metric tensor more general than that of Euclidean or even Minkoswki space. A path that is a “local minimum” (calculated using the relevant metric), meaning that it is shorter than other paths that can be reached from it by small deformations, is referred to as a geodesic. This definition causes both the two great-circle paths of Fig. 22.2 to be identified as geodesics, because even the longer path is of minimum length relative to small deformations. In practice, it is usually easy to identify which of several geodesics in fact corresponds to the shortest path. The calculus of variations is the natural tool for identifying geodesics, and in fact it was used in Example 22.1.1 to verify that a straight line is the geodesic connecting given points in Euclidean space. To extend the analysis to more general metric spaces, we start by relat- ing the distance between two neighboring points, ds, with the changes in their coordinates, dqi,iD1;2;::: . Note that we distinguish between covariant and contravariant quantities, using superscripts for the latter (coordinate displacements are contravariant; compare with Section 4.3). The distance dsis a scalar, given by ds2Dgi jdqidqj: (22.54) Here gi jis the metric tensor, which is symmetric but in many cases of interest not diagonal. Note that we are using the Einstein summation convention, so iandjinEq. (22.54) are summed, causing ds2to be a scalar. This formula is an obvious generalization of that for Euclidean space, ds2Ddx2Cdy2Cdz2; but differs therefrom in that the coordinates qiare not assumed to be mutually orthogonal, sods2contains cross terms dqidqjwith i6Dj. A path in our curved space can be described parametrically by giving the qias functions of an independent variable that we will call u, and the distance between two points A and B can then be represented as JDBZ Ads duduDBZ Aq gi jdqidqj duduDBZ As gi jdqi dudqj dudu DBZ Aq gi jPqiPqjdu; (22.55) where we are borrowing the dot notation, Pqidqi=du. One could now proceed to find the qi.u/that minimize J, but this is a relatively difficult problem. Instead we rely on the Lagrangian formulation of relativistic mechanics, where, for a particle not subject to a potential (other than a gravitational force whose effect is described by the metric), the Lagrangian reduces to LDm 2gi jPqiPqj: (22.56) ArfKen_Ch22-9780123846549.tex 1104 Chapter 22 Calculus of Variations Here the dot notation refers to derivatives with respect to the proper time (or to any other variable related thereto by an affine transformation (meaning the new variable, e.g., u, is related toby a transformation of the form uDaCb). This means that we can replace the minimization of Jby that of the action: BZ Agi jPqiPqjduD0; (22.57) in effect simplifying our problem by eliminating the radical that was present in Eq. (22.55). The minimization in Eq. (22.57) is a relatively simple standard problem in the calculus of variations; for solving it we note that each gi jis in general a function of all the qk(but not the derivativesPqk). There will be an Euler equation for each k; before simplification they take the form @gi jPqiPqj @qkd du@gi jPqiPqj @PqkD0: (22.58) Starting to evaluate Eq. (22.58), we get @gi j @qkPqiPqjd dugi j@ @Pqk PqiPqj D@gi j @qkPqiPqjd du gkjPqjCgikPqi D0: (22.59) Some simplification is achieved by using the relations dPqj duDRqjanddgkj duD@gkj @qiPqi (remember that the Einstein summation convention is still in use). Equation (22.59) reduces to 1 2PqiPqj@gi j @qk@gkj @qi@gik @qj gikRqiD0: (22.60) As a final simplification, we multiply Eq. (22.60) bygkland use the identity gklgikDl i, reaching (in a more expanded notation) the geodesic equation d2ql du2Cdqi dudqj du1 2gklh@gkj @qiC@gik @qj@gi j @qki D0: (22.61) Comparing with the formula for the Christoffel symbol, Eq. (4.63), we can rewrite Eq. (22.61) as d2ql du2Cdqi dudqj du0l i jD0: (22.62) Note that although Eq. (22.62) gives the differential equation describing geodesics in curved space, it is a long way from that equation to its explicit solution for significant problems in general relativity. The exploration of such solutions is a topic of current re- search and beyond the scope of the present text. ArfKen_Ch22-9780123846549.tex 22.2 More General Variations 1105 Relation to Physics The calculus of variations as developed so far provides an elegant description of a wide variety of physical phenomena. The physics includes classical mechanics, as in Ex- amples 22.2.1 and22.2.2; relativistic mechanics, Exercise 22.2.5; electrostatics, Exam- ple 22.2.3; and electromagnetic theory in Exercise 22.2.10. The convenience should not be minimized, but at the same time we should be aware that in these cases the calculus of variations has only provided an alternate description of what was already known. The situation does change with incomplete theories. If the basic physics is not yet known, a postulated variational principle can be a useful starting point. Exercises 22.2.1 (a) Develop the equations of motion corresponding to LD1 2m.Px2CPy2/. (b) In what sense do your solutions minimize the integralZt2 t1L dt? Compare the result for your solution with xDconstant, yDconstant. 22.2.2 From the Lagrangian equations of motion, Eq. (22.38), show that a system in stable equilibrium has a minimum potential energy. 22.2.3 Write out the Lagrangian equations of motion of a particle in spherical coordinates for potential Vequal to a constant. Identify the terms corresponding to (a) centrifugal force and (b) Coriolis force. 22.2.4 The spherical pendulum consists of a mass on a wire of length l, free to move in polar angleand azimuth angle '(Fig. 22.8). (a) Set up the Lagrangian for this physical system. (b) Develop the Lagrangian equations of motion. θ ϕl y x FIGURE 22.8 Spherical pendulum. ArfKen_Ch22-9780123846549.tex 1106 Chapter 22 Calculus of Variations 22.2.5 Show that the Lagrangian LDm0c20 @1s 1v2 c21 AV.r/ leads to a relativistic form of Newton’s second law of motion, d dt m0vip 1v2=c2! DFi; in which the force components are FiD@ V=@xi. 22.2.6 The Lagrangian for a particle with charge qin an electromagnetic field described by scalar potential 'and vector potential Ais LD1 2mv2q'CqAv: Find the equation of motion of the charged particle. Hint..d=dt/AjD@Aj=@tCP i.@Aj=@xi/Pxi. The dependence of the force fields Eand Bon the potentials 'andAis developed in Section 3.9; see in particular Eq. (3.108). ANS. mRxiDqTECvBUi: 22.2.7 Consider a system in which the Lagrangian is given by L.qi;Pqi/DT.qi;Pqi/V.qi/; where qiandPqirepresent sets of variables. The potential energy Vis independent of velocity and neither TnorVhas any explicit time dependence. (a) Show that d dt0 @X jPqj@L @PqjL1 AD0: (b) The constant quantity X jPqj@L @PqjL defines the Hamiltonian H. Show that under the preceding assumed conditions, H satisfies HDTCV, and is therefore the total energy. Note. The kinetic energy Tis a quadratic function of the Pqi. 22.2.8 The Lagrangian for a vibrating string (small-amplitude vibrations) is LDZ1 2u2 t1 2u2 x dx; whereis the (constant) linear mass density and is the (constant) tension. The x-integration is over the length of the string. Show that application of Hamilton’s ArfKen_Ch22-9780123846549.tex 22.3 Constrained Minima/Maxima 1107 principle to the Lagrangian density (the integrand), now with two independent vari- ables, leads to the classical wave equation, @2u @x2D @2u @t2: 22.2.9 Show that the stationary value of the total energy of the electrostatic field of Example 22.2.3 is a minimum. Hint. Investigate the 2terms of J. 22.2.10 The Lagrangian (per unit volume) of an electromagnetic field with a charge density  and current density Jis given by LD1 2 "0E21 0B2 'CJA: Show that Lagrange’s equations lead to two of Maxwell’s equations. (The remaining two are a consequence of the definition of EandBin terms of Aand'.) Hint. Take'and the components of Aasdependent variables; and x,y,z, and tas independent variables. EandBare given in terms of Aand'by Eq. (3.108). 22.3 C ONSTRAINED MINIMA /MAXIMA In preparation for dealing with problems in the calculus of variations in which an inte- gral is to be minimized subject to constraints (which may either be algebraic equations or fixed values of other integrals), we look now at situations in which we seek a constrained extremum of an ordinary function. A typical constrained problem of the type now under consideration is the minimization of a function of several variables, here illustrated as f.x;y;z/, subject to the constraint that g.x;y;z/be kept constant. Since the equation g.x;y;z/DCdefines a surface, our con- strained problem is that of minimizing f.x;y;z/on a surface of constant g. The presence of the constraint means that only two of the three variables x;y;zare actually independent, and in principle one could solve the constraint equation to obtain zas a function of xand y:zDz.x;y/, after which one could obtain the desired minimum by setting to zero the derivatives @ @xf x;y;z.x;y/ and@ @yf x;y;z.x;y/ : However, it may be cumbersome, or in some cases nearly impossible to solve the constraint equation, and in any case this approach does not treat the variables x;y;zon an explicitly equivalent basis. For these reasons it is useful to employ an alternate procedure, known as the method of Lagrangian multipliers. ArfKen_Ch22-9780123846549.tex 1108 Chapter 22 Calculus of Variations Lagrangian Multipliers Continuing with our three-dimensional illustration in which we seek to minimize f.x;y;z/ subject to the constraint g.x;y;z/DC, our starting point is that the constraint equation implies dgD@g @x yzdxC@g @y xzdyC@g @z xydzD0; where (as indicated here explicitly) the partial derivatives of gare taken viewing x,y, and zas independent. Proceeding as for the derivation of Eq. (1.144), we have @z @x yD@g @x yz@g @z xyand@z @y xD@g @y xz@g @z xy: (22.63) Now setting.@f=@x/yto zero, we have (imposing the constraint dgD0) @f @x yD@f @x yzC@f @z xy@z @x yD@f @x yz@f @z xy@g @z xy@g @x yz D@f @x yz@g @x yzD0; (22.64) where D@f @z xy@g @z xy: (22.65) The quantity is called a Lagrangian multiplier. Now taking Eq. (22.64), its equivalent with yreplacing x, and a rearranged form of Eq. (22.65), we have the symmetrical set of formulas @f @x yz@g @x yzD0; @f @y xz@g @y xzD0; @f @z xy@g @z xyD0:(22.66) ArfKen_Ch22-9780123846549.tex 22.3 Constrained Minima/Maxima 1109 The generalization of Eqs. (22.66) to nvariables and kconstraints is @f @xikX jD1j@gj @xiD0; iD1;2;:::; n: (22.67) Thenequations, Eqs. (22.67), contain nCkunknowns (the n xiand the kj), and they are to be solved subject also to the kconstraint equations. In some problems it is never necessary to evaluate explicitly the Lagrangian multipliers, and for this reason the method is sometimes referred to as that of (Lagrange’s) undetermined multiplier(s). Note that the formulation provided above does not only identify minima; the same equa- tions will locate maxima and saddle points. It is necessary to determine the nature of the stationary points from the specific problem at hand. While the derivation of Eq. (22.66) was asymmetric in that was obtained considering zto be a dependent variable, we could have carried out the analysis with xoryin place ofz. This gives us an alternate route to the final formulas in the special case that .@g=@z/ vanishes, in which case Eq. (22.65) becomes undefined. The method only fails if all the derivatives of a constraint function vanish at the stationary point. Example 22.3.1 MINIMIZING SURFACE-TO-VOLUME RATIO Consider a right circular cylinder of radius rand height h. We wish to find the ratio h=r that will minimize the surface area for a fixed enclosed volume. The relevant formulas are: surface area SD2.rhCr2/, volume VDr2h. Applying Eqs. (22.67) for the case of one constraint and two independent variables, we have @S @r@V @rD2.hC2r/.2 rh/D0; @S @h@V @hD2rr2D0: Eliminatingfrom these equations, we find h=rD2. Because we have not also used the constraint equation, we get only the ratio of the two variables handr(which is the information that is relevant for the present problem). However, if we specify the volume V(i.e., use the constraint equation), we then get individual values of handr. We close with two more observations: (1) Our solution obviously provides a minimum S=Vratio, but in principle this has to be determined by closer study of the problem. In the present case, there is no maximum, as S=Vincreases without limit as h=rapproaches zero. (2) We note that minimizing Sfor fixed Vis the same thing as maximizing Vfor fixed S, and leads to equivalent Lagrangian multiplier equations.  ArfKen_Ch22-9780123846549.tex 1110 Chapter 22 Calculus of Variations Exercises 22.3.0 The following problems are to be solved by using Lagrangian multipliers. 22.3.1 The ground-state energy of a quantum particle of mass min a pillbox (right-circular cylinder) is given by EDNh2 2m.2:4048/2 R2C2 H2 ; in which Ris the radius and His the height of the pillbox. Find the ratio of RtoHthat will minimize the energy for a fixed volume. 22.3.2 The U.S. Post Office limits first-class mail to Canada to a total of 36 inches, length plus girth. Using Lagrange multipliers, find the dimensions of the rectangular parallelepiped of maximum volume subject to this constraint. 22.3.3 A thermal nuclear reactor is subject to the constraint '.a;b;c/D a2 C b2 C c2 DB2;a constant; where the reactor is a rectangular parallelepiped of sides a,b, and c. Find the ratios of a,b, and cthat maximize the reactor volume. ANS. aDbDc;cube: 22.3.4 For a lens of focal length f;the object distance pand the image distance qare related by1=pC1=qD1=f. Find the minimum object-image distance .pCq/for fixed f. Assume real object and image ( pandqboth positive). 22.3.5 You have an ellipse .x=a/2C.y=b/2D1. Find the inscribed rectangle of maximum area. Show that the ratio of the area of the maximum-area rectangle to the area of the ellipse is 2=D0:6366 . 22.3.6 A rectangular parallelepiped is inscribed in an ellipsoid of semiaxes a;b, and c. Maxi- mize the volume of the inscribed rectangular parallelepiped. Show that the ratio of the maximum volume to the volume of the ellipsoid is 2=p 30:367 . 22.3.7 Find the maximum value of the directional derivative of '.x;y;z/, d' dsD@' @xcos C@' @ycos C@' @zcos ; subject to the constraint, cos2 Ccos2 Ccos2 D1: ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1111 22.4 V ARIATION WITH CONSTRAINTS As in earlier sections, we seek the path that will make the integral JDZ f yi;@yi @xj;xj dxj (22.68) stationary. This is the general case in which xjrepresents a set of independent variables andyia set of dependent variables. Now, however, we introduce one or more constraints. This means that the yiare no longer independent of each other. Then, if we vary the yi by writing yi. /Dyi.0/C i, not all theimay then be varied arbitrarily, and the Euler equations would not apply. Our approach will be to use Lagrange’s method of undetermined multipliers. We con- sider first the possibility that the kth constraint takes the form of an equation: 'k yi;@yi @xj;xj D0: (22.69) This will ordinarily not be meaningful unless there is more than one dependent or indepen- dent variable, so that Eq. (22.69) restricts, but does not fully determine yi. Remember that yiandxjare here used to denote setsof variables. To introduce an undetermined multi- plier and remain in harmony with our study of the calculus of variations, we note that the constraint, Eq. (22.69), can be stated in the form Z k.xj/'k yi;@yi @xj;xj dxjD0; (22.70) withk.xj/an arbitrary function of the xj.Equation (22.70) is clearly satisfied if Z k.xj/'k yi;@yi @xj;xj dxjD0: (22.71) Alternatively, we may have a constraint in the form of an integral (now dependent on both theyiand their derivatives throughout the interval on which the problem is defined): Z 'k yi;@yi @xj;xj dxjDconstant: (22.72) The effect of this constraint can be brought to a form consistent with Eq. (22.71) by writing Z k'k yi;@yi @xj;xj dxjD0: (22.73) Note that in this equation kdoes not depend on the xjbut is simply a constant, as it is only the integral of 'kthat is required to be stationary. At this point, our constraints have been written as integrals that are dependent on the undetermined multipliers k, wherekmeans either k.xj/or justk, depending on whether the constraint was from Eq. (22.71) or(22.73). We therefore have our problem in a form suitable for applying the method of Lagrangian multipliers as developed in ArfKen_Ch22-9780123846549.tex 1112 Chapter 22 Calculus of Variations Section 22.3, and may use a formula analogous to Eq. (22.67). In our present notation, we obtain Z" f yi;@yi @xj;xj CX kk'k yi;@yi @xj;xj# dxjD0: (22.74) Remember that the Lagrangian multiplier kmay depend on the xjwhen'.yi;xj/is given in the form of Eq. (22.69). We now continue by treating the entire integrand as a new function whose integral is to be made stationary: g yi;@yi @xj;xj DfCX kk'k: (22.75) If we have Ndependent variables yi.iD1;2;:::; N/andmconstraints.kD1;2;:::; m/, then Nmof theimay be taken as arbitrary. In place of arbitrary variation of the m remainingi, we may instead set the mmultiplierskto the (presently unknown) values that permit the Euler equations to be satisfied. The overall result is that we may require satisfaction of an Euler equation for each of the dependent variables yi, but the mquantities kthat appear in the solution of the Euler equations must be assigned values consistent with the constraints that have been imposed. In other words, it will be necessary to solve simultaneously the Euler equations and the equations of constraint to find the function g (and hence f) yielding a stationary value. Lagrangian Formulation with Constraints In the absence of constraints, Lagrange’s equations of motion Eq. (17.52) were found to be8 d dt@L @Pqi@L @qiD0; with t(time) the one independent variable and qi.t/(the particle positions) a set of depen- dent variables. Usually the generalized coordinates qiare chosen to eliminate the forces of constraint, but this is not necessary and not always desirable. In the presence of holo- nomic constraints (those that can be expressed via mathematical expressions, e.g., 'kD0), Hamilton’s principle is Z" L.qi;Pqi;t/CX kk.t/'k.qi;t/# dtD0; (22.76) and the constrained Lagrangian equations of motion are d dt@L @Pqi@L @qiDX kaikk: (22.77) 8The symbol qis customary in classical mechanics. It serves to emphasize that the variable is not necessarily a Cartesian variable (and not necessarily a length). ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1113 Usually the constraint is of the form 'kD'k.qi;t/, independent of the generalized veloci- tiesPqi. In this case the coefficient aikis given by aikD@'k @qi: (22.78) Then aikk(no summation) represents the force of the kth constraint in theOqi-direction, appearing in Eq. (22.77) in exactly the same way as @V=@qi. Example 22.4.1 SIMPLE PENDULUM To illustrate, consider the simple pendulum, a mass m, constrained by a wire of length lto swing in an arc (Fig. 22.9) under a gravitational force characterized by a constant acceleration g. In the absence of the one constraint, '1DrlD0; (22.79) there are two generalized coordinates rand(assuming the motion to be restricted to a vertical plane). The Lagrangian is LDTVD1 2m.Pr2Cr2P2/Cmgrcos; (22.80) taking the potential Vto be zero when the pendulum is horizontal, at D=2. Noting that ar1D@'1 @rD1; a1D@'1 @D0; the equations of motion obtained from Eq. (22.77) are d dt@L @Pr@L @rD1;d dt@L @P@L @D0; (22.81) or d dt.mPr/mrP2mgcosD1; d dt.mr2P/CmgrsinD0: θ m FIGURE 22.9 Simple pendulum. ArfKen_Ch22-9780123846549.tex 1114 Chapter 22 Calculus of Variations Substituting from the equation of constraint ( rDl,PrD0), these equations become mlP2CmgcosD 1; ml2RCmglsinD0: (22.82) The second equation may be solved for .t/to yield simple harmonic motion if the am- plitude is small .sin/, whereas the first equation expresses the tension in the wire in terms ofandP. Note that since the equation of constraint, Eq. (22.79), is in the form of Eq. (22.69), the Lagrange multiplier 1will be a function of t. Since the second equation suffices to determine .t/(assuming a choice of initial conditions), the left-hand side of the first equation can be evaluated if an explicit form for 1is desired.  Example 22.4.2 SLIDING OFF A LOG Another example from mechanics is the problem of a particle sliding on a cylindrical sur- face, as shown in Fig. 22.10. The object is to find the critical angle cat which the particle flies off from the surface. This critical angle is the angle at which the radial force of con- straint goes to zero, and it will depend on the initial velocity with which the particle departs from a position atop the cylinder. To make the problem well-defined, we seek the maxi- mum value that can be attained by c, corresponding to its limit at low initial velocity. To illustrate the present constrained-minimization method, we take LDTVD1 2m.Pr2Cr2P2/mgrcos (22.83) and the one equation of constraint, '1DrlD0: (22.84) Proceeding as in Example 22.4.1, with ar1D@'1 @rD1; a1D@'1 @D0; we reach mRrmrP2CmgcosD1./; mr2RC2mrPrPmgrsinD0: θr FIGURE 22.10 A particle sliding on a cylindrical surface. ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1115 We have chosen to identify the constraining force 1as a function of the angle , a valid choice sinceis a single-valued function of the independent variable t. Inserting the constrained values rDl,RrDPrD0, these equations reduce to mlP2CmgcosD1./; (22.85) ml2RmglsinD0: (22.86) Differentiating Eq. (22.85) with respect to time and remembering that d f./ dtDd f./ dP; we obtain 2mlRmgsinDd1./ d: (22.87) Combining Eqs. (22.86) and (22.87) to eliminate the Rterm, we have d1 dD3mg sin; which integrates to 1./D3mgcosCC: (22.88) To fix the constant C, we evaluate Eq. (22.88) for D0: mlP2 D0CmgD3mgCC; which shows that C2mg , with CD2mg when the initial velocity P.0/ is zero. Using this value of C(which leads to the largest critical angle), we have 1./Dmg.3 cos2/: (22.89) The particle will stay on the surface as long as the force of constraint is nonnegative, that is, as long as the surface has to push outward on the particle, corresponding to 1./> 0. From Eq. (22.89) we find that the critical angle, at which 1.c/D0, satisfies coscD2 3;orcD48110 from the vertical. At or before this angle (neglecting all friction) our particle takes off. It must be admitted that this result can be obtained more easily by considering a varying centripetal force furnished by the radial component of the gravitational force. The example was chosen to illustrate the use of Lagrange’s undetermined multiplier without confusing the reader with a complicated physical system.  ArfKen_Ch22-9780123846549.tex 1116 Chapter 22 Calculus of Variations Example 22.4.3 SCHRÖDINGER WAVE EQUATION As a final illustration of a constrained minimum, let us find the Euler equations for the quantum mechanical problem of a particle of mass msubject to a potential V, JDZ .r/H .r/d3r; (22.90) with the constraint that is the normalized wave function of a bound state: Z .r/ .r/d3rD1: (22.91) Equation (22.90) is a statement that the energy of the system is stationary, with Hits Hamiltonian operator HDNh2 2mr2CV.r/: (22.92) In Eq. (22.90) and are dependent variables; since they are in principle complex we can treat each as a separate variable; this point was discussed in Chapter 5, footnote 3. The integrand in Eq. (17.121) involves second derivatives, but it is convenient to convert them to first derivatives using Green’s theorem, Eq. (3.86): Z .r/r2 .r/d3rDZ S r dZ r r d3r: We now observe that the surface terms vanish due to the requirement that be continuous, and our variational principle becomes Z" Nh2 2mr r CV  # d3rD0: (22.93) The function gfor our constrained variation is therefore gDNh2 2mr r CV    DNh2 2m.  x xC  y yC  z z/CV    ; (22.94) again using the subscript xto denote@=@x. For yiD , our Euler equation becomes @g @ @ @x@g @ x@ @y@g @ y@ @z@g @ zD0: This yields V  Nh2 2m. xxC yyC zz/D0; or Nh2 2mr2 CV D : (22.95) ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1117 The Euler equation for yiD gives the complex conjugate of Eq. (22.95), and therefore provides no further information. Reference to Eq. (22.92) enables us to identify physi- cally as the energy of the quantum mechanical system. With this interpretation, Eq. (22.95) is the celebrated Schrödinger wave equation.  Rayleigh-Ritz Technique A number of physically important problems can be related to variational principles of the general form JDbZ a p.x/y2 xCq.x/y2 dxD0; (22.96) where y.a/andy.b/have fixed values, and the variation is subject to the constraint bZ ay2w.x/dxDconstant. (22.97) Treating Eqs. (22.96) and(22.97) as a constrained minimization, its Euler equation takes the form d dx p.x/dy dx q.x/yCwyD0; (22.98) whereis a Lagrange multiplier. This situation usually arises in contexts such that w.x/ is a nonnegative weight function and y.a/andy.b/satisfy Sturm-Liouville boundary con- ditions, meaning that p.x/yxy b aD0: (22.99) From the above we conclude that although originally introduced as a Lagrange multiplier, must also be an eigenvalue of the Sturm-Liouville system described by Eqs. (22.98) and (22.99). This identification was already noted in Example 22.4.3. Often problems of the type now under discussion are presented as unconstrained mini- mizations of the form JD0 BBBBBBB@bZ a p.x/y2 xCq.x/y2 dx bZ ay2w.x/dx1 CCCCCCCAD0: (22.100) Equation (22.100) is equivalent to the earlier formulation because py2 xCqy2is homoge- neous in yand the denominator normalizes ywithout otherwise changing its functional form. The Jsatisfying Eq. (22.100) evaluates to the eigenvalue . ArfKen_Ch22-9780123846549.tex 1118 Chapter 22 Calculus of Variations In the frequently occurring case that p.x/is actually independent of x, we can manipu- late the y2 xterm in the integrand of J, causing Eqs. (22.96), (22.97), and (22.100) to assume the useful forms, JDbZ a p yxxCq.x/y2 dxD0; (constrained minimum), (22.101) pd2y dx2q.x/yCwyD0; (22.102) JD0 BBBBBBB@bZ a p y y xxCq.x/y2 dx bZ ay2w.x/dx1 CCCCCCCAD0; (unconstrained, JD). (22.103) The Rayleigh-Ritz technique uses the direct evaluation of any one of the above forms for JD0as a means of obtaining solutions to the eigenvalue problem shown as Eq. (22.98) or(22.102). Application of the technique can be as simple as guessing a form for yand evaluating J, but more accurate results are obtained by taking a form for y.x/that contains adjustable parameters, and then varying the parameters to minimize Jwithin the param- eter space. The quality of the results obtained obviously depends on whether the actual minimum form for yhas been well approximated. Ground State Eigenfunction Suppose that we seek to compute the ground-state eigenfunction y0and eigenvalue 0 of some complicated atomic or nuclear system.9A classic example, for which no exact analytical solution has been found, is the helium atom problem. The eigenfunction y0is unknown, but we shall assume that we can make a pretty good guess at an approximation to it, which we will call y. Although we do not know either y0or any other eigenfunctions yi(iD1;2:::), or the corresponding eigenvalues i, we do know, because the eigenfunc- tions can be chosen to form a complete orthogonal set, that we can write the expansion yDc0y0C1X iD1ciyi: (22.104) We shall assume that we picked ysensibly enough that it is not orthogonal to the ground state, so c06D0. Invoking the orthogonality property, Ey, the expectation value of the 9This means that 0is the smallest eigenvalue. ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1119 energy for wave function y, is EyDhyjHjyi hyjyiD1X iD0jcij2i 1X iD0jcij2; (22.105) where H, the operator defining the Schrödinger equation, typically has the form HDNh2 2md2 dx2CV.x/: The Schrödinger equation and its approximate solution Eyare then seen to correspond to Eqs. (22.102) and (22.103). The final member of Eq. (22.105) results from the substitution ofEq. (22.104). This substitution is similar to that carried out in Eq. (6.30), but note that in that equation the function (there called ) was assumed normalized. As we already observed in Section 6.4, the expression for Eyis a weighted average of the eigenvalues (with all the weights jcij20), so Eymust be at least as large as y0, and in fact must be larger if ycontains any admixture of eigenfunctions whose iare larger than 0. It is useful to scale ysoc0D1and rearrange Eq. (22.105) to EyD0C1X iD1c2 ii 1C1X iD1c2 i; (22.106) a form that makes clear that the error in Eywill be quadratic in the ci, even though the difference between yandy0is linear in the ci. Our analysis therefore contains two important results. (1) Whereas the error in the eigenfunction ywasO.ci/, the error inis onlyO.c2 i/. Even a poor approximation of the eigenfunctions may yield an accurate calcula- tion of the eigenvalue. (2) If0is the lowest eigenvalue (ground state), then Ey>0, so our approximation is always on the high side, but converges to 0as our approximate eigenfunction yimproves ( ci!0). In practical problems in quantum mechanics, yoften depends on parameters that may be varied to minimize Eyand thereby improve the estimate of the ground-state energy 0. This is the “variational method” discussed in quantum mechanics texts. It was illustrated in Example 8.4.1. Example 22.4.4 QUANTUM OSCILLATOR The ground state of a quantum-mechanical particle of mass mconstrained to the region 0x<1and subject also to a potential VDkx2=2is described (in a unit system with ArfKen_Ch22-9780123846549.tex 1120 Chapter 22 Calculus of Variations NhD1) by the lowest-eigenvalue eigenstate of the Schrödinger equation, 1 2md2 dx2Ckx2 2 DE ; (22.107) subject to the boundary conditions .0/D .1/D0. A guessed wave function consis- tent with the boundary conditions is y.x/Dxe x. Let’s find the value of making the approximate eigenvalue a minimum. Our Schrödinger equation is of the type represented by Eq. (22.102), so we can use Eq. (22.103) and find the unconstrained value of Jas given there with pD1=2m ,qD kx2=2,wD1, and integration range .0;1/. Noting that yxxD . x2/e x, we have JD1Z 0 x 2m. x2/Ckx4 2 e2 xdx 1Z 0x2e2 xdxD1 8m C3k 8 5 1 4 3D 2 2mC3k 2 2: (22.108) Differentiating Eq. (22.108) with respect to 2and setting the result to zero, we get 1 2m3k 2 4D0; or D.3mk/1=4: Inserting this value into the expression for J, Eq. (22.108), we find JD.3mk/1=2 2mC3k 2.3mk/1=2Dr 3k m1:732r k m: (22.109) This value of Jis an upper bound to the ground-state energy, the exact value of which is 1:5pk=m. Taking a somewhat more complicated (and flexible) wave function, of the form yD.xC cx2/e x, and optimizing both andc, the approximate energy improves to 1:542pk=m: The approximate wave functions of this example are compared with the exact wave func- tion in Fig. 22.11. Note that the second approximation yields an eigenvalue that is in error 2ExactExact Approx Approx 46 246 FIGURE 22.11 Exact and approximate ground-state wave functions for quantum oscillator, Example 22.4.4, plotted for k=mD1. Left: Single-term approximation yDxe x. Right: Two-term approximation yD.xCcx2/e x. ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1121 by less than 3%, even though the approximate wave function exhibits considerably larger relative errors.  Example 22.4.5 VARIATION OF LINEAR PARAMETERS A frequent use of the Rayleigh-Ritz technique is the approximation of an eigenfunction of a Schrödinger equation, H .x/DE .x/; as a truncated expansion in a fixed orthonormal set of functions. The advantage of this pro- cedure is that the parameters in the wave function all occur linearly, and the optimization reduces to a matrix eigenvalue problem. Given an approximate function (often called a trial function) of the form y.x/DNX iD1ci'i.x/; (22.110) we seek to minimize JDhyjHjyisubject tohyjyiD1: (22.111) Again we emphasize that the 'ihave no specific relation to the eigenfunction we seek; they are simply members of an orthonormal set that has two desirable features: (1) they are such that a few of them can provide a reasonable representation of the eigenfunction, and (2) they are tractable in the sense that it is convenient to evaluate the matrix elements we are about to define. Defining a matrix Hof elements Hi jDh' ijHj'jiand a column vector cwith compo- nents ci, we can restate Eq. (22.111) as the minimization of JDc†Hc subject to c†cD1: (22.112) This formulation, in turn, can be reduced using Lagrangian multipliers to the unconstrained matrix eigenvalue problem HcDc; (22.113) which we can solve (using matrix methods) for (the approximate value of J). This ap- plication of the Rayleigh-Ritz technique is therefore seen to be equivalent to the approxi- mation of an operator equation by a finite matrix equation.  Exercises 22.4.1 A particle of mass mis on a frictionless horizontal surface. In terms of plane polar coordinates.r;/, it is constrained to move so that D!t(accomplished by pushing it with a rotating radial arm against which it can slide frictionlessly). With the initial conditions tD0;rDr0;PrD0V ArfKen_Ch22-9780123846549.tex 1122 Chapter 22 Calculus of Variations (a) Find the radial position as a function of time. ANS. r.t/Dr0cosh!t: (b) Find the force exerted on the particle by the constraint. ANS. F.c/D2mPr!D2mr 0!2sinh!t: 22.4.2 A point mass mis moving on a flat, horizontal, frictionless plane. The mass is con- strained by a string to move radially inward at a constant rate. Using plane polar coor- dinates.r;/,rDr0kt: (a) Set up the Lagrangian. (b) Obtain the constrained Lagrange equations. (c) Solve the -dependent Lagrange equation to obtain !.t/, the angular velocity. What is the physical significance of the constant of integration that you get from your “free” integration? (d) Using the !.t/from part (b), solve the r-dependent (constrained) Lagrange equa- tion to obtain .t/. In other words, explain what is happening to the force of constraint as r!0. 22.4.3 A flexible cable is suspended from two fixed points. The length of the cable is fixed. Find the curve that will minimize the total gravitational potential energy of the cable. ANS. Hyperbolic cosine. 22.4.4 A fixed volume of water is rotating in a cylinder with constant angular velocity !. Find the curve of the water surface that will minimize the total potential energy of the water in the combined gravitational-centrifugal force field. ANS. Parabola. 22.4.5 (a) Show that for a fixed-length perimeter the plane figure with maximum area is a circle. (b) Show that for a fixed planar area the boundary with minimum perimeter is a circle. Hint. The radius of curvature Ris given by RD.r2Cr2 /3=2 rr2r2 r2: Note. The problems of this section, variation subject to constraints, are often called isoperimetric. The term arose from problems of maximizing area subject to a fixed perimeter, as in part (a) of this problem. 22.4.6 Show that requiring J, given by JDbZ abZ aK.x;t/'.x/'.t/dx dt; ArfKen_Ch22-9780123846549.tex 22.4 Variation with Constraints 1123 to have a stationary value subject to the normalizing condition bZ a'2.x/dxD1 leads to a Hilbert-Schmidt integral equation of the form given in Eq. (20.64). Note. The kernel K.x;t/is symmetric. 22.4.7 An unknown function satisfies the differential equation y00C 22 yD0 and the boundary conditions y.0/D1,y.1/D0. (a) Calculate the approximation DFTytrialUforytrialD1x2. (b) Compare with the exact eigenvalue. ANS. (a)D2:5. (b)= exactD1:013 . 22.4.8 In Exercise 22.4.7 use a trial function yD1xn. (a) Find the value of nthat will minimize FTytrialU. (b) Show that the optimum value of ndrives the ratio = exact down to 1.003. ANS..a/nD1:7247: 22.4.9 A quantum mechanical particle in a sphere (Example 14.7.1) satisfies r2 Ck2 D0; with k2D2mE=Nh2. The boundary condition is that D0atrDa, where ais the radius of the sphere. For the ground state [where D .r/], try an approximate wave function, a.r/D1r a2 ; and calculate an approximate eigenvalue k2 a. Hint. To determine p.r/andw.r/, put your equation in self-adjoint form (in spherical polar coordinates). ANS. k2 aD10:5 a2;k2 exactD2 a2. 22.4.10 The wave equation for a quantum mechanical oscillator may be written as d2 .x/ dx2C.x2/ .x/D0; ArfKen_Ch22-9780123846549.tex 1124 Chapter 22 Calculus of Variations withD1for the ground state; see Eq. (18.17). Take trialD8 < :1x2 a2;x2a2 0; x2>a2 for the ground-state wave function (with a2an adjustable parameter) and calculate the corresponding ground-state energy. How much error do you have? Note. Your parabola is really not a very good approximation to a Gaussian exponential. What improvements can you suggest? 22.4.11 The Schrödinger equation for a central potential may be written as Lu.r/CNh2l.lC1/ 2Mr2u.r/DEu.r/: Thel.lC1/term, the angular momentum barrier, comes from splitting off the angu- lar dependence. Compare Eq. (9.80) (divide that equation through by r2). Use the Rayleigh-Ritz technique to show that E>E0, where E0is the energy eigenvalue of Lu0DE0u0corresponding to lD0. This means that the ground state will have lD0, zero angular momentum. Hint. You can expand u.r/asu0.r/CP1 iD1ciui, where LuiDEiui,Ei>E0. Additional Readings Bliss, G. A., Calculus of Variations. The Mathematical Association of America. LaSalle, IL: Open Court Pub- lishing Co. (1925). As one of the older texts, this is still a valuable reference for details of problems such as minimum-area problems. Courant, R., and H. Robbins, What Is Mathematics? 2nd ed. New York: Oxford University Press (1996). Chapter VII contains a fine discussion of the calculus of variations, including soap-film solutions to minimum-area problems. Ewing, G. M., Calculus of Variations with Applications. New York: Norton (1969). Includes a discussion of sufficiency conditions for solutions of variational problems. Lanczos, C., The Variational Principles of Mechanics, 4th ed. Toronto: University of Toronto Press (1970), reprinted, Dover (1986). This book is a very complete treatment of variational principles and their applications to the development of classical mechanics. Sagan, H., Boundary and Eigenvalue Problems in Mathematical Physics. New York: Wiley (1961), reprinted, Dover (1989). This delightful text could also be listed as a reference for Sturm-Liouville theory, Legendre and Bessel functions, and Fourier series. Chapter 1 is an introduction to the calculus of variations, with applications to mechanics. Chapter 7 picks up the calculus of variations again and applies it to eigenvalue problems. Sagan, H., Introduction to the Calculus of Variations. New York: McGraw-Hill (1969), reprinted, Dover (1983). This is an excellent introduction to the modern theory of the calculus of variations, which is more sophisticated and complete than his 1961 text. Sagan covers sufficiency conditions and relates the calculus of variations to problems of space technology. Weinstock, R., Calculus of Variations. New York: McGraw-Hill (1952), New York: Dover (1974). A detailed, systematic development of the calculus of variations and applications to Sturm-Liouville theory and physical problems in elasticity, electrostatics, and quantum mechanics. Yourgrau, W., and S. Mandelstam, Variational Principles in Dynamics and Quantum Theory, 3rd ed. Philadel- phia: Saunders (1968), New York: Dover (1979). This is a comprehensive, authoritative treatment of vari- ational principles. The discussions of the historical development and the many metaphysical pitfalls are of particular interest. ArfKen_Ch23-9780123846549.tex CHAPTER 23 PROBABILITY AND STATISTICS Probabilities arise in many problems dealing with random events or large numbers of parti- cles defining random variables. An event is called random if it is practically impossible to predict from the initial state. This includes cases where we have merely incomplete infor- mation about initial states and/or the dynamics. For example, in statistical mechanics we deal with systems containing large numbers of particles, but our knowledge is ordinarily limited to a few average or macroscopic quantities such as the total energy, the volume, the pressure, or the temperature. Because the values of these macroscopic variables are consistent with very large numbers of different microscopic configurations of our system, we are prevented from predicting the behavior of individual atoms or molecules. Often the average properties of many similar events are predictable, as in quantum theory. This is why probability theory can be and has been developed. Random variables are also involved when data depend on chance, such as weather re- ports and stock prices. The theory of probability describes mathematical models of chance processes in terms of probability distributions of random variables that describe how some “random events” are more likely than others. In this sense, probability is a measure of our ignorance, giving quantitative meaning to qualitative statements, such as “It will probably rain tomorrow” and “I’m unlikely to draw the queen of hearts.” Probabilities are of fun- damental importance in quantum mechanics and statistical mechanics and are applied in meteorology, economics, games, and many other areas of daily life. Because experiments in the sciences are always subject to measurement errors, theories of errors and their propagation involve probabilities. Statistics is the area of mathematics that connects observations on data samples to inferences about the probable content of the entire population from which the sample(s) came. It is an extensive and sophisticated branch of mathematics, and in the present text only a few of the most basic concepts can be presented. The material found here may be adequate to provide a conceptual basis for statistical mechanics, but can at best be an elementary introduction to the ideas needed to 1125 Mathematical Methods for Physicists. ©2013 Elsevier Inc. All rights reserved. ArfKen_Ch23-9780123846549.tex 1126 Chapter 23 Probability and Statistics gain maximum information from data-intensive experimental studies such as those arising from the study of cosmic rays or the data from high-energy particle accelerators. A more complete picture of the role of statistics in physics and engineering can be obtained from a number of the texts in the Additional Readings. 23.1 P ROBABILITY : D EFINITIONS , SIMPLE PROPERTIES All possible mutually exclusive1outcomes of an experiment that is subject to chance represent the events (or points) of a sample space S. Suppose we toss a coin, and record that it lands either “heads” or “tails.” These are mutually exclusive events, so our sample space for a single coin toss can be deemed to be spanned by a discrete random variable x, with possible values xi, which (based on our experiment, called a trial), will have one of the two values x1(for heads) or x2(for tails). Now suppose, with the same sample space, we carry out larger numbers of trials. Some will have the result x1(heads), others x2(tails). It is of interest to define the probability of an outcome in our sample space by the ratio P.xi/number of times event xioccurs total number of trials; (23.1) where it is assumed that the number of trials is large enough that P.xi/approaches a constant limiting value. In the event that we are able to enumerate all the possible events that produce outcomes in our sample space and can also assume that each event is equally likely, we may then define the theoretical probability of an outcome xias P.xi/number of outcomes xi total number of all events: (23.2) An example of the use of this theoretical probability can be illustrated using coin tosses. For example, suppose that we toss a coin twice and take our random variable xto be the number of heads obtained in a two-toss trial. Our sample Snow contains three possible values of x, which we designate x0,x1,x2, where we are now letting xistand for the occurrence of iheads in the two tosses. Obviously, the only possible values of xare 0, 1, and 2. But we also know that the four possible results of two successive tosses are (heads, then heads), (heads, then tails), (tails, then heads), (tails, then tails); these possibilities are mutually exclusive and it is reasonable to assume that they are equally likely. Then, using Eq. (23.2), we conclude that the probabilities of x2(two heads) and x0(no heads) will each be 1/4, while the probability of x1(one heads) will be 1/2. The experimental definition, Eq. (23.1), is the more appropriate when the total number of events is not well defined (or is difficult to obtain) or we cannot identify equally likely outcomes. A large, thoroughly mixed pile of black and white sand grains of the same size and in equal proportions is a relevant example, because it is impractical to count them all. But we can count the grains in a small sample volume that we pick. This way we can check that white and black grains turn up with roughly equal probability 1=2, provided that we put back each sample and mix the pile again. It is found that the larger the sample volume, 1This means that given that one particular event did occur, the others could not have occurred. ArfKen_Ch23-9780123846549.tex 23.1 Probability: De/f_initions, Simple Properties 1127 the smaller the spread in probability about 1=2will be. Moreover, the more trials we run, the closer the average of all the individual trial probabilities will be to 1=2. We could even pick single grains and check if the probability 1=4of picking two black grains in a row equals that of two white grains, etc. There are lots of statistics questions we can pursue. Thus, piles of colored sand provide for instructive experiments. The following axioms are self-evident. Probabilities satisfy 0P1:Probability 1means certainty; probability 0means impossibility. The entire sample has probability 1:For example, drawing an arbitrary card from a deck of cards has probability 1: The probabilities for mutually exclusive events add. The probability for getting exactly one head in two coin tosses is 1=4C1=4D1=2because it is 1=4for head first and then tail, plus 1=4for tail first and then head. Example 23.1.1 PROBABILITY FOR AORB Before proceeding with this example, we must clarify the definition of “or.” In probability theory, “ AorB” means A,B, orboth AandB. The specification “ AorBbut not both” is referred to as the exclusive or ofAandB(sometimes abbreviated xor). What is the probability for drawing a club or a jack from a shuffled deck of cards?2To answer this question we need to identify equally probable mutually exclusive events. We note that because there are 52 cards in a deck, the drawing of each being equally likely (with 13 cards for each suit and 4 jacks), there are 13 clubs including the club jack, and 3 other jacks; that is, there are 16 mutually exclusive draws that meet our specification out of the total of 52, giving the probability .13C3/=52D16=52D4=13 .  Sets, Unions, and Intersections If we represent a sample space by a set Sof points, then events meeting certain specifica- tions can be identified as subsets A;B;:::ofS, denoted as AS, etc. Two sets A;Bare equal if Ais contained in B, denoted AB, and Bis contained in A, denoted BA. Theunion A[Bconsists of all points (events) that are in AorBor both (see Fig. 23.1). Theintersection A\Bconsists of all points that are in both AandB. IfAandBhave no common points, their intersection is the empty set (which has no elements), and we write A\BD;. The set of points in Athat are not in the intersection of AandBis denoted byAA\B, thereby defining a subtraction of sets. If we take the club suit in Exam- ple 23.1.1 as set Aand the four jacks as set B;then their union comprises all clubs and jacks, and their intersection is the club jack only. Each subset Ahas its probability P.A/0:In terms of these set-theory concepts and notations, the probability laws we just discussed become 0P.A/1: 2Note that these events are not mutually exclusive. ArfKen_Ch23-9780123846549.tex 1128 Chapter 23 Probability and Statistics A B FIGURE 23.1 The shaded area gives the intersection A\B, corresponding to the Aand Bevent sets; the dashed line encloses A[B, corresponding to the event set AorB. The entire sample space has P.S/D1. The probability of the union A[Bof mutually exclusive events is the sum P.A[B/DP.A/CP.B/;where A\BD;: (23.3) Theaddition rule for probabilities of arbitrary sets is given by the following theorem: Addition rule: P.A[B/DP.A/CP.B/P.A\B/: (23.4) To prove Eq. (23.4), we write the union as two mutually exclusive sets: A[BDA[ .BB\A/, where we have subtracted the intersection of AandBfrom Bbefore joining them. The respective probabilities of these mutually exclusive sets are P.A/andP.B/ P.B\A/, which we add. We could also have written A[BD.AA\B/[B, from which our theorem follows similarly by adding these probabilities: P.A[B/DTP.A/ P.A\B/UC P.B/. Note that A\BDB\A:The relationships among these sets can be checked by referring to Fig. 23.1. Sometimes the rules and definitions of probabilities that we have discussed so far are not sufficient, and we need to introduce the notion of conditional probability. Let AandB denote sets of events in our sample space. The conditional probability P.BjA/is defined to be the probability that an event which is a member of Ais also a member of B. To understand the need for this somewhat formal definition, consider the following example. Example 23.1.2 CONDITIONAL PROBABILITY Consider a box of 10identical red and 20identical blue pens, from which we remove pens successively in a random order without putting them back. Suppose we draw a red pen first, event R, followed by the draw of a blue pen, event B. One way to compute P.R;B/is to note that our sample space consists of 3029mutually exclusive and equally probable points (each a two-event ordered sequence), of which 1020meet our specifications, leading to the computation P.R;B/D.1020/=.3029/D20=87 . Note that in this example, P.R;B/refers to ordered events. Another way of making the same computation is to start by noting that the initial drawing of a red pen will occur with probability P.R/D10=30 . But now the probability of drawing a blue pen in the next round, event B, however, will depend on the fact that we drew a red pen in the first round, and is given by the conditional probability P.BjR/. Since there are ArfKen_Ch23-9780123846549.tex 23.1 Probability: De/f_initions, Simple Properties 1129 now 29 pens of which 20 are blue, we easily compute P.BjR/D20=29 , and the probability of the sequence “red, then blue” is P.R;B/D10 3020 29D20 87; (23.5) equal to the result we obtained previously.  The generalization of the result in Eq. (23.5) is the very useful formula P.A;B/DP.A/P.BjA/; (23.6) which has the obvious interpretation that the probability that AandBboth occur can be written as the probability of A, multiplied by the conditional probability P.BjA/thatB occurs, given the occurrence of A. Two observations relative of Eq. (23.6) are in order. First, it can be rearranged to reach an explicit formula for P.BjA/: P.BjA/DP.A;B/ P.A/: (23.7) Second, if the conditional probability P.BjA/DP.B/is independent of A, then the events AandBare called independent, and the combined probability is simply the product of both probabilities, or P.A;B/DP.A/P.B/;(AandBindependent). (23.8) IfAandBare defined in a way that neither depends on the other (a condition not satisfied in Example 23.1.2), we can rewrite Eq. (23.7) as P.BjA/DP.A\B/ P.A/: (23.9) Example 23.1.3 SCHOLASTIC APTITUDE TESTS Colleges and universities rely on the verbal and mathematics SAT scores, among others, as predictors of a student’s success in passing courses and graduating. A research university is known to admit mostly students with a combined verbal and mathematics score of 1400 points or more. The graduation rate is 95%; that is, 5% drop out or transfer elsewhere. Of those who graduate, 97% have an SAT score of at least 1400 points, while 80% of those who drop out have an SAT score below 1400: Suppose a student has an SAT score below 1400: What is his/her probability of graduating? LetArepresent all students with an SAT test score below 1400, and let Brepresent those with scores1400 . These are mutually exclusive events with P.A/CP.B/D1. LetCrepresent those students who graduate, and let QCrepresent those who do not. Our problem here is to determine the conditional probabilities P.CjA/andP.CjB/. To apply Eq. (23.9) we need the four probabilities P.A/,P.B/,P.A\C/, and P.B\C/. Among the 95% of students who graduate, 3% are in set Aand 97% are in set B, so P.A\C/D.0:95/.0:03/D0:0285; P.B\C/D.0:95/.0:97/D0:9215: ArfKen_Ch23-9780123846549.tex 1130 Chapter 23 Probability and Statistics Among the 5% of students who do not graduate, 80% are in set Aand 20% are in set B, so P.A\QC/D.0:05/.0:80/D0:0400; P.B\QC/D.0:05/.0:20/D0:0100: Since P.A/DP.A\C/CP.A\QC/, and likewise for P.B/, we have P.A/D0:0285C0:0400D0:0685; P.B/D0:9215C0:0100D0:9315: Now, applying Eq. (23.9), we obtain the final results P.CjA/DP.A\C/ P.A/D0:0285 0:068541:6%; P.CjB/DP.B\C/ P.B/D0:9215 0:931598:9%I that is, a little less than 42% is the probability for a student with a score below 1400 to graduate at this particular university.  As a corollary to the equation for conditional probability, Eq. (23.9), we now compare P.AjB/DP.A\B/=P.B/andP.BjA/DP.A\B/=P.A/, obtaining a result known as Bayes’ theorem: P.AjB/DP.A/ P.B/P.BjA/: (23.10) Bayes’ theorem is a special case of the following more general theorem: If the random events Aiwith probabilities P.Ai/>0are mutually exclusive and their union represents the entire sample S, then an arbitrary random event BShas the probability P.B/DnX iD1P.Ai/P.BjAi/: (23.11) The decomposition law given by Eq. (23.11) resembles the expansion of a vector into a ba- sis of unit vectors defining its components. This relation follows from the obvious decom- position BD[ i.B\Ai/(this notation indicates the union of all the quantities B\Ai, see Fig. 23.2), which implies P.B/DP iP.B\Ai/because the components B\Aiare mu- tually exclusive. For each i;we know from Eq. (23.9) thatP.B\Ai/DP.Ai/P.BjAi/, which proves the theorem. Counting Permutations and Combinations Counting the events in samples can help us find probabilities; this procedure is found to be of great importance in statistical mechanics. If we have ndifferent molecules, let us ask in how many ways we can arrange them in a row, that is, permute them. This number is defined as the number of their permutations. Thus, by definition, the order matters in permutations. There are nchoices of picking the first molecule, n1for the second, etc. Altogether there are nWpermutations of n different molecules or objects. ArfKen_Ch23-9780123846549.tex 23.1 Probability: De/f_initions, Simple Properties 1131 B A1 A2A3 FIGURE 23.2 The shaded area Bis composed of mutually exclusive subsets of B belonging also to A1;A2;A3;where the Aiare mutually exclusive. Generalizing this, suppose there are npeople but only k<nchairs to seat them. In how many ways can we seat kpeople in the chairs? Counting as before, we get n.n1/.nkC1/DnW .nk/W(23.12) for the number of permutations of kobjects which can be formed by selection from a set originally containing nobjects. We now consider the counting of combinations of objects, where the term combination is defined to refer to sets in which the object order is irrelevant. For example, three letters a;b;ccan be combined, two letters at a time, in three ways: ab,ac,bc. If letters can be repeated, then we also have the pairs aa,bb,ccand have a total of six combinations. These examples illustrate the fact that a combination of different particles differs from a permutation in that the particles’ order does not matter. Combinations may occur with repetition or without; the essential point is that no two combinations contain the same particles. The number of different combinations of nnumbered (and thereby distinguishable) par- ticles, kat a time and without repetitions, is given by the binomial coefficient n.n1/.nkC1/ kWDn k : (23.13) To prove Eq. (23.13), we start from the number nW=.nk/Wof permutations in which k particles were chosen from n, and divide out the number kWof permutations of the group ofkparticles because their order does not matter in a combination. A generalization of the above is a situation in which we have a total of ndistinguishable (numbered) objects, and we place n1of these into Box 1, n2into Box 2, etc. We wish to know how many different ways this can be done (this is a combination problem because the objects in each box do not form ordered sets). A simple way to solve this problem is to identify each permutation of the nobjects with an assignment into boxes; the first n1of the permuted objects is placed in Box 1, the next n2in Box 2, etc. However, permutations that differ only in the ordering of objects destined for the same box do not constitute different distributions, so the total number of distributions will be nW(the overall number ArfKen_Ch23-9780123846549.tex 1132 Chapter 23 Probability and Statistics of permutations) divided by n1W,n2W, etc. Thus, our overall formula is B.n1;n2;:::/DnW n1Wn2W:::: (23.14) This quantity is sometimes referred to as a multinomial coefficient; if there were only two boxes it reduces to the binomial coefficient. For a related problem with repetition, suppose that we have an unlimited supply of particles bearing each number from 1 through k. Then the number of distinct ways in which nparticles can be chosen can be shown to be nCk1 n DnCk1 k1 : (23.15) The following example provides a proof of Eq. (23.15). Example 23.1.4 COMBINATIONS WITH REPETITION The physical relevance of the situation giving rise to Eq. (23.15) is that it is mathematically equivalent to the number of ways that nidentical, indistinguishable particles can be placed inkboxes. To see that these problems are equivalent, note that the number on each particle ofEq. (23.15) can be used to identify the box in which that particle will be placed. A simple way to count the possible assignments is to consider the distinguishable ways thatnindistinguishable particles and k1indistinguishable partitions can be placed in a line containing nCk1items. The particles (if any) that occur in the line earlier than the first partition are assigned to Box 1; those between the first and second partitions are assigned to Box 2, etc., with the particles (if any) occurring later than the .k1/th (the last) partition are assigned to Box k. Each different placement of the partitions yields a unique assignment of particles to boxes, and the number of different partition placements is the number of combinations given by the binomial coefficient in Eq. (23.15).  In statistical mechanics, we frequently need to know the number of ways in which it is possible to put nparticles in kboxes subject to various additional specifications. If we are working in classical theory, our more complete specification includes the notion that the particles are distinguishable, and we refer to the probability computation as that given byMaxwell-Boltzmann statistics. In the quantum domain, it is assumed that identical particles are inherently indistinguishable; in fact, we cannot even identify them by their trajectories, as the notion of path is blurred by the Heisenberg uncertainty principle. This indistinguishability leads to the requirement that many-particle states must have symmetry under the interchange of identical particles, and in nature we find two cases: Either the wave function is symmetric under interchange of the coordinates of a pair of identical particles (such particles are said to exhibit Bose-Einstein statistics), or the coordinate interchange causes a reversal in the sign of the wave function (the case called Fermi- Dirac statistics). The symmetry (or antisymmetry) under particle interchange influences the way in which particles can be assigned to states (boxes): In Bose-Einstein statistics ArfKen_Ch23-9780123846549.tex 23.1 Probability: De/f_initions, Simple Properties 1133 any number of indistinguishable particles may be placed in the same box; in Fermi-Dirac statistics no box may contain more than one indistinguishable particle. Application of the various kinds of statistics in general problems is outside the scope of this text; however, the basic case in which we simply count the number of assignments that are possible is easily approached. If we have nparticles and kavailable states: In Maxwell-Boltzmann (classical) statistics, the number of possible assignments of particles to states is kn(each particle can independently be assigned to any state). In Bose-Einstein statistics, the number of possible assignments is given by Eq. (23.15). In Fermi-Dirac statistics, the number of possible assignments isk n . This formula gives the number of ways that nof the kstates can be selected for occupancy. Note that the number of assignments is zero if n>k, indicating that we cannot make any assignment (with a maximum of one particle per state) unless there are at least as many states as there are particles. Exercises 23.1.1 A card is drawn from a shuffled deck. (a) What is the probability that it is black, (b) a red nine, (c) or a queen of spades? 23.1.2 Find the probability of drawing two kings from a shuffled deck of cards (a) if the first card is put back before the second is drawn, and (b) if the first card is not put back after being drawn. 23.1.3 When two fair dice are thrown, what is the probability of (a) observing a number less than 4;or (b) a number greater than or equal to 4but less than 6? 23.1.4 Rolling three fair dice, what is the probability of obtaining six points? 23.1.5 Determine the probability P.A\B\C/in terms of P.A/;P.B/;P.C/;P.A[B/; P.A[C/;P.B[C/, and P.A[B[C/. 23.1.6 Determine directly or by mathematical induction (Section 1.4) the probability of a dis- tribution of N(Maxwell-Boltzmann) particles in kboxes with N1in Box 1,N2in Box 2;:::; Nkin the kth box for any numbers Nj1with N1CN2CC NkD N;k<N:Repeat this for Fermi-Dirac and Bose-Einstein particles. 23.1.7 Show that P.A[B[C/DP.A/CP.B/CP.C/P.A\B/P.A\C/ P.B\C/CP.A\B\C/: 23.1.8 Determine the probability that a positive integer n100is divisible by a prime number p100: Verify your result for pD3;5;7. 23.1.9 Put two particles obeying Maxwell-Boltzmann (Fermi-Dirac, or Bose-Einstein) statis- tics in three boxes. How many ways of doing so are there in each case? ArfKen_Ch23-9780123846549.tex 1134 Chapter 23 Probability and Statistics 23.2 R ANDOM VARIABLES In this section we define properties that characterize the probability distributions of random variables, by which we mean variables that will assume various numerical values with individual probabilities. Thus, the name of a color (e.g., “black” or “white”) cannot be the value assigned a random variable, but we can define a random variable to have one numerical value for “black” and another for “white”; the usefulness of our definition may depend on the problem we are attempting to solve. Having defined a random variable and given its distribution, we are interested in partic- ular in its mean oraverage value, and in measures of the width or spread of its values. The width is of particular importance when the random variable represents repeated mea- surements of the same quantity but subject to experimental error. In addition, we introduce properties that characterize the extent to which the value of one random variable depends on (i.e., is correlated with) those of another. Random variables can be discrete, as for example those introduced in the previous sec- tion to describe the outcomes of coin tosses, or they may be continuous, either inherently so (as, for example, the wave function in a quantum mechanical system) or because they consist of so many closely spaced discrete points that it is impractical to work with them individually. Example 23.2.1 DISCRETE RANDOM VARIABLE The possible outcomes of the tossing of a die define a random variable Xwith values x1;x2;:::; x6, each with probability 1=6; we can denote this by writing P.xi/D1=6,iD 1:::6. If we toss two dice and record the sum of the points shown in each trial, then this sum is also a discrete random variable, which takes on the value 2when both dice show 1with probability.1=6/2; the value 3in either of the two cases in which one die has 1and the other 2, hence with probability .1=6/2C.1=6/2D1=18 . Continuing, the value 4is reached in three equally probable ways: 2C2,3C1, and 1C3with total probability 3.1=6/2D 1=12 ; the values 5and6are reached with the respective probabilities 4.1=6/2D1=9and 5.1=6/2D5=36 ; and the value 7occurs with the maximum probability, 6.1=6/2D1=6. The value 8is reached in five ways ( 6C2;5C3;4C4;3C5;2C6), with probability 5.1=6/2D5=36 , and further increases in xlead to smaller probabilities, finally at xD12 reaching probability .1=6/2D1=36 . This probability distribution is symmetric about xD7, and can be represented graphically as in Fig. 23.3 or algebraically as P.x/Dx1 36D6.7x/ 36;xD2;3;:::; 7; P.x/D13x 36D6C.7x/ 36;xD7;8;:::; 12:  ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1135 P(x) x6 36 5 36 4 36 3 36 2 36 1 36 0 123456789 1 0 1 1 1 2 FIGURE 23.3 Probability distribution P.x/of the sum of points when two dice are tossed. In summary, then, If a discrete random variable Xcan assume the values xi, each value occurs by chance with a probability P.XDxi/Dpi0that is a discrete-valued function of the random variable X, and the probabilities satisfyP ipiD1. We define the probability density f.x/of acontinuous random variable Xas P.xXxCdx/Df.x/dxI (23.16) that is, f.x/dxis the probability that Xlies in the interval xXxCdx:Forf.x/ to be a probability density, it has to satisfy f.x/0andR f.x/dxD1: The generalization to probability distributions depending on several random variables is straightforward. Quantum physics abounds in examples. Example 23.2.2 CONTINUOUS RANDOM VARIABLE: HYDROGEN ATOM Quantum mechanics gives the probability j j2d3rof finding a 1selectron in a hydrogen atom in volume3d3r;where DNer=ais the 1swave function, ais the Bohr radius, andND.a3/1=2is a normalization constant such that Z j j2d3rD4N21Z 0e2r=ar2drDa3N2D1: 3Note thatj j24r2drgives the probability for the electron to be found between randrCdr, at any angle. ArfKen_Ch23-9780123846549.tex 1136 Chapter 23 Probability and Statistics The value of this integral can be checked by identifying it as a gamma function: 1Z 0e2r=ar2drDa 231Z 0exx2dxDa3 80.3/Da3 4:  Computing Discrete Probability Distributions In Example 23.2.1 the overall probability of a particular value of a discrete random variable was computed as a product in which one factor was the number of equally likely ways in which that value could be obtained, and the other factor was the probability of each mutually exclusive occurrence. This type of computation arises sufficiently frequently that we should learn how to deal with it in general. Therefore, consider a situation in which Nindependent events take place (examples of such events include tosses of an individual die, selection of a card from a deck, energy state occupied by a molecule, orientation of the magnetic moment of a particle), and that each such event has one of a set of mmutually exclusive outcomes (e.g., number showing on the die, identity of the card, energy state, or magnetic moment orientation). We assume that the outcomes x1;x2;:::; xmof an individual event will have the respec- tive probabilities p1;p2;:::; pm, with p1Cp2CC pmD1(so that we have included all the possible outcomes). Then, we compute the probability that any n1of the events have outcome x1, any n2events have outcome x2, etc.: P.n1;n2;:::; nm/DB.n1;n2;:::; nm/.p1/n1.p2/n2:::.pm/nm; (23.17) where n1Cn2CC nmDN, and B.n1;n2;:::; nm/is the number of ways that, for each i,niof the events have outcome xi. Now B.n1;n2;:::; nm/is just the multinomial coefficient encountered earlier; in the present context the numbered objects correspond to events numbered from 1 to Nand each box corresponds to an individual-event outcome. Thus, our final formula for the probability of a distribution defined by n1,n2, etc., is P.n1;n2;:::; nm/DNW n1Wn2W:::nmW.p1/n1.p2/n2:::.pm/nm: (23.18) Mean and Variance When we make nmeasurements of a quantity x, obtaining the values xj, we define the average value NxD1 nnX jD1xj (23.19) of the trials, also called the mean orexpectation value, where this formula assumes that every observed value xiis equally likely and occurs with probability 1=n. This connection ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1137 is the key link of experimental data with probability theory. This observation and practical experience suggest defining the mean value for a discrete random variable Xas hXiX ixipi; (23.20) while defining the mean value for a continuous random variable xcharacterized by prob- ability density f.x/as hXiDZ x f.x/dx: (23.21) Other notations for the mean in the literature are NXandE.X/. The use of the arithmetic mean Nxofnmeasurements as the average value is suggested by simplicity and plain experience, again assuming equal probability for each xi. But why do we not consider the geometric mean xgD.x1x2:::xn/1=n or the harmonic mean xhdetermined by the relation 1 xhD1 n1 x1C1 x1CC1 xn or the valueQxthat minimizes the sum of absolute deviations jxiQxj? Here the xiare taken to increase monotonically. When we plot O.x/DP2nC1 iD1jxixj;as in Fig. 23.4(a), for an odd number of points, we realize that it has a minimum at its central value iDn;while for an even number of points E.x/DP2n iD1jxixjis flat in its central region, as shown in Fig. 23.4(b). These properties make these functions unacceptable for determining average values. Instead, when we minimize (with respect to x) the sum of quadratic deviations, nX iD1.xxi/2Dminimum, (23.22) (a)( b)x1x2 x3 x1 x2x3x4 FIGURE 23.4 (a)P3 iD1jxixjfor an odd number of points. (b)P4 iD1jxixjfor an even number of points. ArfKen_Ch23-9780123846549.tex 1138 Chapter 23 Probability and Statistics setting the derivative equal to zero yields 2P i.xxi/D0;or xD1 nX ixiNx; that is, the arithmetic mean. The arithmetic mean has another important property: If we denote byviDxiNxthe deviations, thenP iviD0;that is, the sum of positive deviations equals the sum of negative deviations. This principle of minimizing the quadratic sum of deviations, called the method of least squares, is due to C. F. Gauss, among others. The ability of a mean value to represent a set of data points depends on the spread of the individual measurements from this mean. Again, we reject the average sum of deviationsPn iD1jxiNxj=nas a measure of the spread because it selects the central measurement as the best value for no good reason. A more appropriate definition of the spread is based on the average of the squares of the deviations from the mean. This quantity, known as the standard deviation, is defined as Dvuut1 nnX iD1.xiNx/2; (23.23) where the square root is motivated by dimensional analysis. If we square Eq. (23.23) and expand .xiNx/2, written as.xihxi/2, we get n2DnX iD1x2 i2hxinX iD1xiCnhxi2 Dn hx2ihxi2 : Dividing through by n, we obtain the very useful formula, 2Dhx2ihxi2: (23.24) Note that these two expectation values are equal only if all the xihave the same value; for example, if we have two xi, equal, respectively, to hxiCandhxi, thenhx2iD hxi2C2, so the spread in the xihas causedhx2 iito increase. Example 23.2.3 STANDARD DEVIATION OF MEASUREMENTS From the measurements x1D7,x2D9,x3D10,x4D11,x5D13, we extractNxD10for the mean value and, using Eq. (23.23), Ds .3/2C.1/2C02C12C32 5D2:2361 for the standard deviation, or spread.  ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1139 There is yet another interpretation of the standard deviation, in terms of the sum of squares of measurement differences: X i<k.xixk/2D1 2nX iD1nX kD1 x2 iCx2 k2xixk D1 2h 2n2hx2i2n2hxi2i Dn22: (23.25) The last step in the above equation made use of Eq. (23.24). Now we are ready to generalize the spread in a set of nmeasurements with equal prob- ability 1=nto the variance of an arbitrary probability distribution. For a discrete random variable Xwith probabilities piatXDxi, we define the variance 2DX j xjhXi2pjI (23.26) for a continuous probability distribution the definition becomes 2D1Z 1.xhXi/2f.x/dx: (23.27) We now develop some relationships satisfied by random variables: 1. The variance 2of a random variable Xhas the property 2DhX2ih Xi2: (23.28) This formula, previously derived as Eq. (23.24) only for a discrete random variable with all xiequally probable, is true in general. The proof is left as Exercise 23.2.3. 2. If random variables XandYare related by the linear equation YDaXCb, then Y has mean valuehYiDahXiCband variance 2.Y/Da22.X/. We prove this theorem only for a continuous distribution, leaving the case of a discrete random variable as an exercise for the reader. Directly from the definitions, we have hYiD1Z 1.axCb/f.x/dxDahXiCb; where the integral multiplying bsimplifies becauseR f.x/dxD1. For the variance we similarly obtain 2.Y/D1Z 1.axCbahXib/2f.x/dxD1Z 1a2.xhXi/2f.x/dx Da22.X/: ArfKen_Ch23-9780123846549.tex 1140 Chapter 23 Probability and Statistics 3. Probabilities of random variables satisfy the Chebyshev inequality, P.jxhXijk/1 k2; (23.29) which demonstrates why the standard deviation serves as a measure of the spread of an arbitrary probability distribution from its mean value hXi. We first derive the simpler inequality P.YK/hYi K for a continuous random variable Ywith values yrestricted to y0. (The proof for a discrete random variable follows along similar lines.) This inequality follows from hYiD1Z 0y f.y/dyDKZ 0y f.y/dyC1Z Ky f.y/dy 1Z Ky f.y/dyK1Z Kf.y/dyDKP.YK/: Next we apply the same method to the positive variance integral, 2DZ .xhXi/2f.x/dxZ jxhXijk.xhXi/2f.x/dx k22Z jxhXijkf.x/dxDk22P.jxhXijk/; where we have first decreased the right-hand side by omitting the part of the positive integral withjxhXijkand then decreased it further by replacing .xhXi/2 in the remaining integral by its minimum value, k22. We now divide the first and last members of this sequence of inequalities by the positive quantity k22, thereby proving the Chebyshev inequality. For kD3we have the conventional three-standard- deviation estimate, P.jxhXij3/1 32D1 9: (23.30) ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1141 Moments of Probability Distributions It is straightforward to generalize the mean value to higher moments of probability distri- butions relative to the mean value hXi: D .XhXi/kE DX j xjhXikpj; discrete distribution; D .XhXi/kE D1Z 1.xhXi/kf.x/dx;continuous distribution.(23.31) Themoment-generating function het XiDZ etxf.x/dxD1CthXiCt2 2WhX2iC (23.32) is a weighted sum of the moments of the continuous random variable X, which is obtained by substituting the Taylor expansion of the exponential functions. Therefore, hXiDdhet Xi dt tD0;hX2iDd2het Xi dt2 tD0;:::;hXniDdnhet Xi dtn tD0: (23.33) Note that the moments here are not relative to the expectation value, but are relative to xD0; they are called central moments. Example 23.2.4 MOMENT-GENERATING FUNCTION Suppose we have four cards, numbered from 1 through 4, from which we draw two at random and add their numbers. Letting this sum of the drawn numbers be values of a random variable X, we find that Xhas the following values and respective probabilities P.x/: P.3/D1=6; P.4/D1=6; P.5/D1=3; P.6/D1=6; P.7/D1=6: Verifying these probabilities is the topic of Exercise 23.2.1. The moment-generating function for this system has the form MD1 6 e3tCe4tC2e5tCe6tCe7t ; and its first two derivatives are M0D1 6 3e3tC4e4tC10e5tC6e6tC7e7t ; M00D1 6 9e3tC16e4tC50e5tC36e6tC49e7t : ArfKen_Ch23-9780123846549.tex 1142 Chapter 23 Probability and Statistics Setting tD0, we get hXiDM0.0/D5;hX2iDM00.0/D80 3: Thus, the mean of Xis found to be 5, and its variance is given by 2DhX2ih Xi2D80 325D5 3: In this example we see that the moment-generating function does (in a systematic way) the same thing as direct formation of the moments; in a later example, Example 23.3.2, we see a situation in which the use of the moment-generating function provides an opportunity to compute moments with rather little computational work.  Mean values, central moments, and variance can be defined analogously for probability distributions that depend on several random variables. We illustrate for the case of two random variables XandY, for which the mean values and the variance of each variable take the forms hXiD1Z 11Z 1x f.x;y/dx dy; hYiD1Z 11Z 1y f.x;y/dx dy;(23.34) 2.X/D1Z 11Z 1.xhXi/2f.x;y/dx dy; 2.Y/D1Z 11Z 1.yhYi/2f.x;y/dx dy:(23.35) Covariance and Correlation Two random variables are said to be independent if the probability density f.x;y/fac- torizes into a product f.x/g.y/of probability distributions of one random variable each. The covariance, defined as cov.X;Y/Dh.XhXi/.YhYi/i; (23.36) ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1143 is a measure of how much the random variables XandYare correlated (or related): It is zero for independent random variables because cov.X;Y/DZ .xhXi/.yhYi/f.x;y/dx dy DZ .xhXi/f.x/dxZ .yhYi/g.y/dy D.hXih Xi/.hYihYi/D0: The normalized covariance cov. X;Y/=.X/.Y/, which has values between 1andC1, is often called correlation. In order to demonstrate that the correlation is bounded by 1cov.X;Y/ .X/.Y/1; we analyze the positive mean value QDhTa.XhXi/Cc.YhYi/U2i Da2hTXhXiU2iC2achTXhXiUTYhYiUiC c2hTYhYiU2i Da2.X/2C2accov.X;Y/Cc2.Y/20: (23.37) For this quadratic form to be nonnegative for all values of the constants aandc, its discrim- inant must satisfy cov .X;Y/2.X/2.Y/20, which proves the desired inequality. The usefulness of the correlation as a quantitative measure is emphasized by the follow- ing theorem: The probability P.YDaXCb/will be unity if, and only if, the correlation cov.X;Y/=.X/.Y/is equal to1. This theorem states that a 100% correlation between X and Y implies not only some functional relation between both random variables but that the relation between them is linear. Our first step in proving this theorem is to show that P.YDaXCb/D1(meaning thatYDaXCb) implies that cov .X;Y/=.X/.Y/D1 . For the meanhYi, we simply compute hYiDhaXCbiDahXiCb: For the variance, .Y/2DhY2ihYi2Dh.aXCb/2i.ahXiCb/2 Da2hX2iC2abhXiCb2 a2hXi2C2abhXiCb2 Da2 hX2ih Xi2 Da2.X/2; ArfKen_Ch23-9780123846549.tex 1144 Chapter 23 Probability and Statistics which is equivalent to .Y/Da.X/. We also need cov .X;Y/, which is cov.X;Y/Dh.XhXi/..aXCb/.ahXiCb//i Da hX2ih Xi2 Da2.X/D. X/.Y/; where the last equality was obtained by identifying a.X/as.Y/. This result completes the first step in our proof of the theorem. To complete the proof, we must establish the converse of the relation we have just proved, namely that cov .X;Y/=.X/.Y/D1 implies P.YDaXCb/D1for some set of values.a;b/. We proceed by forming the quadratic expectation value *h ..Y/X.X/Y/D .Y/X.X/YEi2+ ; where the symbolindicates that we choose a sign opposite to that of the correlation cov.X;Y/=.X/.Y/. Our plan is to show that this expectation value is zero. Since the expectation value is that of an inherently nonnegative quantity, we may then conclude that .Y/X.X/Yis (almost) everywhere equal to its expectation value, the value of which is some constant C. We therefore have .Y/X.X/YDC;equivalent to YD.Y/XC/ .X/; the linear relation we seek. It remains to confirm that the quadratic expectation value vanishes. Rearranging it first to the form *h .Y/.XhXi/.X/.YhYi/i2+ and then expanding the square, we reach D .Y/2.XhXi/2C.X/2.YhYi/22.X/.Y/.XhXi/.YhYi/E : Making now the substitutions .XhXi/2D.X/2,.YhYi/2D.Y/2, and .XhXi/ .YhYi/ D. X/.Y/, our quadratic expectation value reduces to zero. Marginal Probability Distributions It is sometimes useful to integrate out (i.e., average over) one of the random variables in a multivariable distribution. When we do so, we are left with the probability distribution of the other random variables. For a two-variable distribution, we can eliminate either of the two variables: F.x/DZ f.x;y/dy;orG.y/DZ f.x;y/dx; (23.38) and analogously for discrete probability distributions. When one or more random variables are integrated out, the remaining probability distribution is called marginal, the name ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1145 motivated by the geometric aspects of projection. It is straightforward to show that these marginal distributions satisfy all the requirements of properly normalized probability dis- tributions. Here is a comprehensive example that illustrates the computation of probability distri- butions and their mean values, variances, covariance, and correlation. Example 23.2.5 REPEATED DRAWS OF CARDS This example deals with independent repeated draws from a deck of playing cards. To make sure that these events stay independent, we draw the first card at random from a bridge deck containing 52cards and then put it back at a random place and reshuffle the deck. Now we repeat the process for a second card. Let’s define the random variables: XDnumber of so-called honors, that is, tens, jacks, queens, kings, or aces; YDnumber of twos or threes. In a single draw the probability of Event a(drawing an honor) is paD20=52D5=13 , while the probability of Event b(drawing a two or three) is pbD2.4=52/D2=13 . The probability of Event c(drawing anything else) is pcD.1352/=13D6=13 . Since that exhausts all the mutually exclusive possibilities, we have aCbCcD1. In two drawings, it is possible to draw zero, one, or two honors (i.e., xD0, 1, or 2). Likewise, we may draw zero, one, or two cards of value 2 or 3 (i.e., yD0, 1, or 2). But because we are only drawing two cards, we have the additional condition 0xCy2. The probability function P.XDx;YDy/, which we will write in the simpler form P.x;y/, is given by a formula of the type presented in Eq. (23.18), with N(the number of events) equal to 2 and with the three individual-event probabilities pa,pb, and pc. The number of events aisx, the number of events bisy, and therefore the number of events c is2xy, and, by Eq. (23.18), P.x;y/D2W xWyW.2xy/W.pa/x.pb/y.pc/2xy D2W xWyW.2xy/W5 13x2 13y6 132xy ;(23.39) with 0xCy2. More explicitly, P.x;y/has the following values: P.0;0/D6 132 ;P.1;0/D25 136 13D60 132; P.2;0/D5 132 ;P.0;1/D22 136 13D24 132; P.0;2/D2 132 ;P.1;1/D25 132 13D20 132: ArfKen_Ch23-9780123846549.tex 1146 Chapter 23 Probability and Statistics The probability distribution is properly normalized. Its expectation values are given by hXiDX 0xCy2x P.x;y/DP.1;0/CP.1;1/C2P.2;0/ D60 132C20 132C25 132 D130 132D10 13D2pa; and hYiDX 0xCy2y P.x;y/DP.0;1/CP.1;1/C2P.0;2/ D24 132C20 132C22 132 D52 132D4 13D2pb: The values 2paand2pbare expected because we are drawing a card two times. The vari- ances are 2.X/DX 0xCy2 x10 132 P.x;y/ D 10 132 TP.0;0/CP.0;1/CP.0;2/UC3 132 TP.1;0/CP.1;1/UC16 132 P.2;0/ D10264C3280C16252 134D425169 134D80 132; 2.Y/DX 0xCy2 y4 132 P.x;y/ D 4 132 TP.0;0/CP.1;0/CP.2;0/UC9 132 TP.0;1/CP.1;1/UC22 132 P.0;2/ D42112C9244C22222 134D114169 134D44 132: The covariance is given by cov.X;Y/DX 0xCy2 x10 13 y4 13 P.x;y/D104 13262 132109 13224 132 1022 1324 13234 13260 132C39 13220 132164 13252 132D20 132: Therefore, the correlation of the random variables X;Yis given by cov.X;Y/ .X/.Y/D20 8p 511D1 2r 5 11D0:3371; ArfKen_Ch23-9780123846549.tex 23.2 Random Variables 1147 which means that there is a small (negative) correlation between these random variables, because if an honor is drawn, that drawing is not available to yield a 2 or a 3, and vice versa. Finally, let us determine the marginal distribution, P.XDx/D2X yD0P.x;y/; or explicitly, P.XD0/DP.0;0/CP.0;1/CP.0;2/D6 132 C24 132C2 132 D8 132 ; P.XD1/DP.1;0/CP.1;1/D60 132C20 132D80 132; P.XD2/DP.2;0/D5 132 ; which is properly normalized because P.xD0/CP.XD1/CP.XD2/D64C80C25 132D169 132D1: The mean value and variance of Xcan be computed from the marginal probabilities: hXiD2X xD0x P.XDx/DP.XD1/C2P.XD2/D80C225 132D130 132D10 13; FD2X xD0 x10 132 P.XDx/D 10 1328 132 C3 13280 132C16 1325 132 D80169 134D80 132: These data agree with our earlier computations of the same quantities.  Conditional Probability Distributions If we are interested in the distribution of a random variable Xfor a definite value yDy0 of another random variable, then we deal with a conditional probability distribution P.XDxjYDy0/. The corresponding continuous probability density is f.x;y0/. Exercises 23.2.1 Verify the probabilities for the outcomes of the two-card draws in Example 23.2.4, and by direct computation of the mean and variance check the results given in that example. ArfKen_Ch23-9780123846549.tex 1148 Chapter 23 Probability and Statistics 23.2.2 Show that adding a constant cto a random variable Xchanges the expectation value hXiby that same constant but not the variance. Show also that multiplying a random variable by a constant multiplies both the mean and variance by that constant. Show that the random variable XhXihas mean value zero. 23.2.3 Using the definition given in Eq. (23.27) for the variance 2of a continuous random variable, show that 2DhX2ih Xi2: 23.2.4 A velocityvjDxj=tjis measured by recording the distances xjat the corresponding times tj:Show thatNx=Ntis a good approximation for the average velocity v;provided that all the errors are small: jxjNxjjN xjandjtjNtjjNtj. 23.2.5 Redefine the random variable Yin Example 23.2.5 as the number of fours through nines. Then determine the correlation of the XandYrandom variables for the drawing of two cards (with replacement, as in the example). 23.2.6 The probability that a particle of an ideal gas travels a distance xbetween collisions is proportional to ex=fdx, where fis the constant mean free path. Verify that fis the average distance between collisions, and determine the probability of a free path of length l3f. 23.2.7 Determine the probability density for a particle in simple harmonic motion in the intervalAxA: Hint. The probability that the particle is between xandxCdxis proportional to the time it takes to travel across the interval. 23.3 B INOMIAL DISTRIBUTION In this and the next two sections, we explore specific random variable distributions that are of importance both in physics and in the mathematical theories of probability and statistics. The topic of the present section is the binomial distribution, which typically occurs in the study of repeated independent trials of random events. Example 23.3.1 REPEATED TOSSES OF DICE What is the probability of three sixes in four tosses, all trials being independent? Getting one six in a single toss of a fair die has probability aD1=6, and getting anything else has probability bD5=6with aCbD1. Let the random variable SDsbe the number of sixes. In four tosses, 0s4. The probability distribution P.SDs/is given by the product of the two possibilities, asandb4s, times the number of ways that ssixes can be obtained from four tosses. This number is given by Eq. (23.18), and our probability is P.SDs/D4W sW.4s/Wasb4sD4 s asb4s: (23.40) ArfKen_Ch23-9780123846549.tex 23.3 Binomial Distribution 1149 We can now check that our probability is properly normalized by verifying that the sum of P.SDs/for all sadds to unity. From properties of the binomial coefficients, we find 4X sD04 s asb4sD.aCb/4D1 6C5 64 D1: (23.41) Writing out the cases of Eq. (23.40) explicitly, we have f.0/Db4;f.1/D4ab3;f.2/D6a2b2;f.3/D4a3b;f.4/Da4; so we can answer our original question: The probability of three sixes in four tosses is 4a3bD41 635 6D5 434; which is fairly small.  This case dealt with repeated independent trials, each with two possible outcomes of constant probability pfor a hit and qD1pfor a miss, and it is typical of many ap- plications, including practical issues such as the random instances of defective products. The generalization to SDssuccesses in ntrials is given by the binomial probability distribution: P.SDs/DnW sW.ns/WpsqnsDn s psqns: (23.42) Figure 23.5 shows histograms for cases with 20 trials and various hit probabilities p. Example 23.3.2 USE OF MOMENT-GENERATING FUNCTION If we view our probability distribution as the result of adding together nrandom variables Si, each having the value siD1with probability pand the value siD0with probability q, we can use the moment-generating function of Eq. (23.32) to obtain more information about the binomial distribution. We write het SiDhet.S1CS2CC Sn/iDhet S1ihet S2ihet SniDh het S1iin ; (23.43) where we have used the fact that the trials are independent to write het Sias a product of single-trial expectation values, all of which are identical. We continue by evaluating het S1i, which is an average for the two values s1D1, with probability p, and s1D0, with probability q. We get het S1iDpetCqe0DpetCq; (23.44) soEq. (23.43) reduces to het SiD.petCq/n: (23.45) ArfKen_Ch23-9780123846549.tex 1150 Chapter 23 Probability and Statistics f(x=n) p=0.1 p=0.3 p=0.50.30 0.25 0.20 0.15 0.10 0.05 02 4 6 8 1 0 1 2 1 4 1 6 1 8 2 0n FIGURE 23.5 Binomial probability distributions for nD20andpD0:1, 0.3, 0.5. Note that the fact the trials were independent enabled us to obtain the moment-distribution function without enumerating all the many-trial possibilities. Now that we havehet Siwe can differentiate it, as in Eq. (23.33), to obtain moments of our distribution. Using @het Si @tDnpet.petCq/n1; hSiDX isif.si/D@het Si @t tD0Dnp; @2het Si @t2Dnpet.petCq/n1Cn.n1/p2e2t.petCq/n2; hS2iDX is2 if.si/D@2het Si @t2 tD0DnpCn.n1/p2; we obtain, applying Eq. (23.28), 2.S/DhS2ihSi2DnpCn.n1/p2n2p2 Dnp.1p/Dnpq: For a given n, we see that the variance is largest when pDqD1=2. This behavior is apparent in Fig. 23.5, where we see that the distribution broadens as pis increased from 0.1 to 0.3 to 0.5.  ArfKen_Ch23-9780123846549.tex 23.4 Poisson Distribution 1151 Exercises 23.3.1 Show that the variable XDx, defined as the number of heads in ncoin tosses, is a random variable and determine its probability distribution. Describe the sample space. What are its mean value, the variance, and the standard deviation? Plot the probability function P.x/DTnW=xW.nx/WU2nfornD10, 20, and 30 using graphics software. 23.3.2 Plot the binomial probability function for the probabilities pD1=6; qD5=6, and nD6 throws of a die. 23.3.3 A hardware company knows that the probability of mass-producing nails includes a small probability pD0:03 of defective nails (usually without a sharp tip). What is the probability of finding more than two defective nails in its commercial box of 100 nails? 23.3.4 Four cards are drawn from a shuffled bridge deck. What is the probability that they are all red? That they are all hearts? That they are honors? Compare the probabilities when each card is put back at a random place before drawing the next card, with the probabilities when the cards are not replaced in the deck. 23.3.5 Show that for the binomial distribution of Eq. (23.42), the most probable value of x isnp: 23.4 P OISSON DISTRIBUTION The Poisson distribution is often used to describe situations in which an event occurs repeatedly at a constant rate of probability. Typical applications involve the decay of radioactive samples, but only in the approximation that the decay rate is slow enough that depletion in the population of the decaying species can be neglected. Other applications of interest include so-called Poisson noise, where fluctuations in a low rate of arrival of particles at a detector cause statistically predictable fluctuations in the detector signal. The Poisson distribution can be developed by considering the probabilities that varying numbers of events are detected over an interval during which events occur at a constant rate of probability. The essential features of the development are that it assumes that (1) the event rate is small enough that there will be observationally accessible intervals in which at most one event occurs (i.e., one can consider intervals containing either zero or one event), and (2) the total number of events is small enough that it is useful to model their occurrence by a discrete probability distribution. Let’s proceed by defining the probability Pn.t/that exactly nevents occur in a time t, and that the probability of one event occurring in a short time interval dtwill bedt, whereis a constant such that dt1. This time interval dtis therefore short enough that we can neglect the possibility that more than one event occurs within it. Based on this hypothesis, we can set up a recursion relation for Pn.t/by considering the two following mutually exclusive possibilities for the occurrence of nevents in a time dCdt: (1) that nevents occur during a time tand no events occur in a subsequent time interval dt, and (2) that n1events occur during the time tand one event occurs during the subsequent interval dt. We therefore write Pn.tCdt/DPn.t/P0.dt/CPn1.t/P1.dt/: ArfKen_Ch23-9780123846549.tex 1152 Chapter 23 Probability and Statistics Then, inserting P1.dt/Ddt andP0.dt/D1P1.dt/and dividing through by dt, we get, after minor rearrangement, d Pn.t/ dtDPn.tCdt/Pn.t/ dtDPn1.t/Pn.t/: (23.46) As a first step in solving this recursion relation, we note that for nD0it simplifies (because the possibility involving Pn1does not exist) to d P0.t/ dtD P0.t/: (23.47) This equation, with initial condition P0.0/D1(meaning that it is certain that no events are observed in an interval of zero length), has solution P0.t/Det. Our solution informs us that the probability that no events have occurred before time tdecays exponentially with t, at a rate dependent on the magnitude of . From this starting point and the further initial conditions Pn.0/D0forn1(again, no detection of events occurs during an interval of zero length), the recursion relation can be solved to yield Pn.t/D.t/n nWet: (23.48) Equation (23.48) can be checked by substituting it into the recursion formula, Eq. (23.46), and by verifying that it satisfies the initial conditions Pn.0/Dn0. Equation (23.48) is taken as the definition of the Poisson distribution, regarded as a func- tion of the quantity t. Replacingtby, we write the Poisson-distribution probabilities given for a discrete random variable Xin the standard form, p.n/Dn nWe;XDnD0;1;2;:::: (23.49) We can check that the probabilities in Eq. (23.49) are properly normalized by noting thatP nn=nWevaluates to e. An example of a Poisson distribution is given in Fig. 23.6. The mean value and variance of a Poisson distribution are easily calculated: hXiD1X nD1nn nWeDe1X nD1n .n1/WD; (23.50) hX2iD1X nD1n2n nWeDe1X nD1n .n2/WCn .n1/W D2C; (23.51) 2DhX2ih Xi2D.C1/2D: (23.52) The moments can also be calculated from the moment-generating function D et XE D1X nD0n nWeetnDe1X nD0.et/n nWDe.et1/: Recall that the procedure for obtaining moments is to differentiate with respect to tand read out the derivatives evaluated at tD0. ArfKen_Ch23-9780123846549.tex 23.4 Poisson Distribution 1153 0 2468 1 0 1 2 1 4 1 6 1 8 2 00.020.040.060.080.10.120.140.16 FIGURE 23.6 Poisson distribution, D5. Relation to Binomial Distribution A Poisson distribution becomes a good approximation of the binomial distribution for a large number nof trials and small probability p=n, withheld constant. Theorem: In the limit n!1 andp!0so that the mean value np!stays finite, the binomial distribution becomes a Poisson distribution. To prove this theorem, we need to find the large- nlimit of the binomial distribution formula, Eq. (23.42). To do so, we apply Stirling’s formula, in the form nWp 2n.n=e/n for large n. See Eq. (12.110). For the quotient of the two n-dependent factorials occurring in Eq. (23.42), we have (keeping sfinite while letting n!1 ): nW .ns/Wn ene nsns n esn nsns n es 1Cs nsns : The factor in the final expression raised to the power nsis, in the limit of large n, an expression of value es(in fact, it is, with nschanged to n, one of the often-used definitions of the exponential). The final result is nW .ns/Wns: (23.53) We use a similar defining expression for the exponential to evaluate the factor qnsin Eq. (23.42). Writing qnsD.1p/nsand replacing pby its limiting value pD=n, ArfKen_Ch23-9780123846549.tex 1154 Chapter 23 Probability and Statistics 00.020.040.060.080.10.120.14 2 4 6 8 10 12 14 16 18 20 FIGURE 23.7 Comparison of binomial distribution ( ND80,pD0:1), wide bars, and Poisson distribution ( D8), narrow bars. we have qnsD.1p/ns 1 nn 1 ns e.1/e: (23.54) Inserting the large- nlimiting values from Eqs. (23.53) and(23.54) into the formula for the binomial distribution, we reach P.SDs/DnW sW.ns/Wpsqnsns sWpses sWe; (23.55) where in the last step we have combined nsandpsintos. Equation (23.55) establishes our theorem, and thereby completes the connection be- tween the Poisson and binomial distributions. This result, which becomes valid in the limit of a large number of trials, each of small probability, is sometimes referred to as an exam- ple of the laws of large numbers. A comparison of the binomial and Poisson distributions is presented as Fig. 23.7. Exercises 23.4.1 Radioactive decays for long-lived isotopes are governed by the Poisson distribution. In a Rutherford-Geiger experiment, the numbers of emitted particles are counted in each ofnD2608 time intervals of 7:5seconds each. In Table 23.1 niis the number of time intervals in which iparticles were emitted. Determine the average number of particles emitted per time interval, and compare the niofTable 23.1 with npicomputed from the Poisson distribution with mean value . ArfKen_Ch23-9780123846549.tex 23.5 Gauss’ Normal Distribution 1155 Table 23.1 Data for Exercise 23.4.1 i! 0 1 2 3 4 5 6 7 8 9 10 ni! 57 203 383 525 532 408 273 139 45 27 16 23.4.2 Derive the standard deviation of a Poisson distribution of mean value . 23.4.3 The number of -particles emitted by the decay of a radium sample is counted per minute for 40hours. The total number is 5000 . How many 1-minute intervals are ex- pected with (a) 2, and (b) 5 -particles? 23.4.4 For a radioactive sample, 10decays are counted on average in 100seconds. Use the Poisson distribution to estimate the probability of counting 3decays in 10seconds. 23.4.5238U has a half-life of 4:51109years. Its decay series ends with the stable lead isotope 206Pb. The ratio of the number of206Pb to238U atoms in a rock sample is measured as 0:0058 . Estimate the age of the rock assuming that all the lead in the rock is from the initial decay of the238U, which determines the rate of the entire decay process, because the subsequent steps take place far more rapidly. Hint. This is not a Poisson distribution problem, but is an application of the decay law N.t/DNet, where, the decay constant, is related to the half-life TbyTDln 2= . ANS. 3:8107years. 23.4.6 The probability of hitting a target in one shot is known to be 20%. If five shots are fired independently, what is the probability of striking the target at least once? 23.5 G AUSS ’ NORMAL DISTRIBUTION The bell-shaped Gauss distribution is defined by the probability density f.x/D1 p 2exp TxU2 22 ;1<x<1; (23.56) with mean value and variance 2:In part because it represents continuous limits of both the binomial and Poisson distributions, it is by far the most important continuous probability distribution and is displayed in Fig. 23.8. ArfKen_Ch23-9780123846549.tex 1156 Chapter 23 Probability and Statistics h=3 h=2 h=1f 1 01x FIGURE 23.8 Gauss normal distribution for mean value zero and various standard deviations (marked by circles). Curves are labeled by hD1=p 2. It is properly normalized because, substituting yD.x/=p 2, we obtain 1 p 21Z 1e.x/2=22dxD1p1Z 1ey2dyD2p1Z 0ey2dyD1: To check the mean value, we can make the substitution yDx, and find that hXiD1Z 1x p 2e.x/2=22dxD1Z 1y p 2ey2=22dyD0; (23.57) showing thathXiD. The zero result in Eq. (23.57) occurs because the integrand is odd iny, so the integral over y>0cancels that over y<0. A check that the variance of this normal distribution is indeed 2is the topic of Exercise 23.5.1. We can compute conditional probabilities for the normal distribution. In particular, making for convenience the substitution yD.xhXi/= , P.jXhXij>k/DPjXhXij >k DP.jYj>k/ Dr 2 1Z key2=2dyDr 4 1Z k=p 2ez2dzDerfc.k=p 2/; we can evaluate the integral for kD1;2;3, and thus extract the following numerical rela- tions for a normally distributed random variable: P.jXhXij/0:3173; P.jXhXij2/0:0455; P.jXhXij3/0:0027:(23.58) ArfKen_Ch23-9780123846549.tex 23.5 Gauss’ Normal Distribution 1157 It is interesting to compare the last of these quantities with Chebyshev’s inequality, which gives 1/9 for the probability that an event falls further than 3from the mean. The 1/9 applies to an arbitrary probability distribution, and is in strong contrast to the much smaller 0:0027 given by the 3-rule for the normal distribution. Limits of Poisson and Binomial Distributions In a special limit, the discrete Poisson probability distribution is closely related to the continuous Gauss distribution. This limit theorem is another example of the laws of large numbers, which are often dominated by the bell-shaped normal distribution. Theorem: For large nand mean value , the Poisson distribution approaches a Gauss distribution. To prove this theorem, in the limit n!1 , we approximate for large nthe factorial in the Poisson probability p.n/by Stirling’s asymptotic formula, nWp 2n.n=e/n, and choose the deviation vDnfrom the mean value as the new variable. We let the mean valueapproach1and treatv= as small, but assume v2=to be finite. Substituting nDCv, we obtain lnp.n/Dlnne nW Dnlnlnp 2nnlnnCn D.Cv/ln Cv Cvlnp 2.Cv/ D.Cv/ln 1v Cv Cvlnp 2.Cv/: We next expand the first logarithmic term in powers of v=.Cv/, reaching lnp.n/D1X tD1vt t.Cv/t1Cvlnp 2.Cv/: (23.59) The first two terms of the tsummation yield nonvanishing contributions in the large-  limit; further terms vanish because the power of vin the numerator is less than twice that ofin the denominator. Replacing Cvby, Eq. (23.59) reduces to lnp.n/v2 2lnp 2; equivalent to p.n/1p2ev2=2: (23.60) This is a Gauss distribution of the continuous variable vwith mean value 0and standard deviationDp. In another special limit, the discrete binomial probability distribution is also closely related to the continuous Gauss distribution. This limit theorem is yet another example of thelaws of large numbers. ArfKen_Ch23-9780123846549.tex 1158 Chapter 23 Probability and Statistics Theorem: In the limit n!1 , with pa finite trial probability such that the mean value np!1 , the binomial distribution becomes a Gauss normal distribution. Recall from Section 23.4 that, when np!<1, the binomial distribution becomes a Poisson distribution. Instead of the large number sof successes in ntrials, we use the deviation vDs pnfrom the (large) mean value pnas our new continuous random variable, under the condition that as n!1 ,jvjpn(sov=n!0) butv2=nis finite. Thus, we replace sby pnCvandnsbyqnvin the factorials of the formula for the binomial distribution, Eq. (23.42). Writing now W.v/as our probability distribution in the large- pnlimit, we apply Stirling’s formula as we have done several times before, obtaining initially W.v/DpsqnsnnC1=2enCsC.ns/ p 2.pnCv/sC1=2.qnv/nsC1=2: (23.61) Next we factor out the dominant powers of nand cancel powers of pandqto find W.v/D1p2pqn 1Cv pn.pnCvC1=2/ 1v qn.qnvC1=2/ : (23.62) Taking the logarithm of W.v/and expanding in powers of v, we retain only the terms throughv2, yielding lnW.v/Dlnp 2pqn v n1 2p1 2q v2 n21 4p2C1 4q2 Cv2 n1 2pC1 2q C :(23.63) Settingv=nto zero, noting that 1 2pC1 2q DpCq 2pqD1 2pq; and dropping all terms vtwith t>2, we obtain our large- nlimit W.v/D1p2pqnev2=2pqn; (23.64) which is a Gauss distribution in the deviations spn, with mean value 0and standard deviationDpnpq. The large values assumed for both pnandqn(and the discarded terms) restrict the validity of the theorem to the central part of the Gaussian bell shape, and exclude the tails. Exercises 23.5.1 Show that the variance of the normal distribution given by Eq. (23.56) is2, the symbol in that equation. 23.5.2 Show that Eq. (23.62) can be obtained by manipulation of the formula Eq. (23.61) for W.v/. 23.5.3 With W.v/the expression in Eq. (23.62), show that the expansion of lnW.v/in powers ofvleads to Eq. (23.63). ArfKen_Ch23-9780123846549.tex 23.6 Transformations of Random Variables 1159 23.5.4 What is the probability for a normally distributed random variable to differ by more than 4from its mean value? Compare your result with the corresponding one from Chebyshev’s inequality. Explain the difference in your own words. 23.5.5 An instructor grades a final exam of a large undergraduate class, obtaining the mean value of points Mand the variance 2. Assuming a normal distribution for the number Mof points, he defines a grade F when M<m3=2; D when m3=2<M< m=2; C when m=2<M<mC=2; B when mC=2<M<mC3=2; and A when M>mC3=2: What is the percentage of As, Fs; Bs, Ds; and Cs? Redesign the cutoffs so that there are equal percentages of As and Fs ( 5%),25% Bs and Ds, and 40% Cs. 23.6 T RANSFORMATIONS OF RANDOM VARIABLES We have already encountered some elementary transformations involving random vari- ables: In Section 23.2 we observed that a random variable YDaXCbwill have mean valuehYiDahXiCband variance 2.Y/Da22.X/. Here we consider more general transformations, with particular focus on continuous probability distributions. First, consider a simple change of random variable from XtoY, where yDy.x/. If the probability distribution of Xisf.x/dx, then the contribution at xto some quantity M.y/is PfMTy.x/UgdxDMTy.x/Uf.x/dx: (23.65) But we may wish to express the probability in terms of the distribution of Y, writing PTM.y/UdyDM.y/g.y/dy; (23.66) foryevaluated at the point corresponding to x, i.e., yDy.x/. To make these equations consistent, it is necessary that g.y/dyDf.x/dx;org.y/DfTx.y/Udx dy: (23.67) For example, if yDx2, then dx=dyD1=.2x/Dy1=2=2andg.y/Df.py/y1=2=2. Let’s now address the transformation of two random variables X,Yinto U.X;Y/, V.X;Y/. Again we treat the continuous case. If uDu.x;y/; vDv.x;y/;xDx.u;v/; yDy.u;v/ (23.68) describe the transformation and its inverse; integrals of the probability density will trans- form by formulas that include the Jacobian of the transformation (see Section 4.4). The transformed probability density becomes g.u;v/Df.x.u;v/;y.u;v//jJj; (23.69) ArfKen_Ch23-9780123846549.tex 1160 Chapter 23 Probability and Statistics where the Jacobian is [email protected];y/ @.u;v/D @x @u@x @v @y @u@y @v : (23.70) This is a generalization of Eq. (23.67). Addition of Random Variables Let’s apply this analysis to a situation in which Zis the sum of two random variables X andY, orZDXCY. We transform to new variables XandZ, so the transformation is xDx,zDxCy; orxDx,yDzx. The Jacobian for this transformation is JD @x @x@x @z @.zx/ @[email protected]x/ @z D 1 0 1 1 D1: If our original probability distribution was f.x;y/, it therefore transforms into g.x;z/D f.x;zx/. We are usually interested in the marginal distribution in Z, obtained by inte- grating over x, and P.ZDz/g.z/D1Z 1f.x;zx/dx: (23.71) In the oft-occurring case that XandYare independent random variables, so f.x;y/D f1.x/f2.y/, Eq. (23.71) assumes the form g.z/D1Z 1f1.x/f2.zx/dx; (23.72) which we recognize as a Fourier convolution, see Eq. (20.68): g.z/D1Z 1f1.x/f2.zx/dxDp 2.f1f2/.z/: (23.73) Equation (23.72) gives us a general formula whereby we can obtain the distribution of ZDXCYfrom the distributions of independent variables XandY, while Eq. (23.73) shows that it may be useful to consider the use of Fourier transforms for evaluating the integral. In fact, the moment-generating function, Eq. (23.32), is (if tis replaced by it) proportional to the Fourier transform of the probability density, and heitXiDZ eitxf.x/dxDp 2fT.t/ (23.74) is known as the characteristic function in probability theory. ArfKen_Ch23-9780123846549.tex 23.6 Transformations of Random Variables 1161 Applying the Fourier convolution theorem, Eq. (20.70), we therefore write TP.ZDz/UT.t/gT.t/Dp 2fT 1.t/fT 2.t/; (23.75) showing that we can obtain g.z/as the inverse Fourier transform g.z/DZ eiztfT 1.t/fT 2.t/dt: (23.76) Connection with statistics texts will be improved by restating Eqs. (23.75) and (23.76) using the characteristic function notation. Equation (23.75) is equivalent to heit ZiDheit.XCY/iDheit XiheitYi: (23.77) Equation (23.76) states that g.z/is the distribution that corresponds to heit Zi; since Fourier transforms have inverses it can be assured that such a distribution exists. Example 23.6.1 ADDITION THEOREM, NORMAL DISTRIBUTION A good example of the analysis for a random variable ZDXCYis provided when Xand Yare taken to be Gauss normal distributions with zero mean value and the same variance. This situation corresponds to a relationship known as the addition theorem for normal distributions. Theorem: If the independent random variables X;Yhave identical normal distribu- tions, that is, the same mean value and variance, then ZDXCYhas normal distribu- tion with twice the mean value and twice the variance of XandY. To prove this theorem, we assume without loss of generality that the normal distributions each have variance D1, so, from Eq. (23.56), each has the form f.x/D1p 2e.x/2=2: From Eq. (20.18) and the translation formula, Eq. (20.67), we find the Fourier transform off.x/to be fT.t/D1p 2eitet2=2: Now, applying Eq. (23.75), we have gT.t/D1p 2e2itet2: Taking the inverse transform, noting that the complex exponential shifts the origin by an amount 2, we get g.z/D1 2pe.z2/2=4; which shows that the mean and variance of Zare twice those of XandY, so the theorem is satisfied.  ArfKen_Ch23-9780123846549.tex 1162 Chapter 23 Probability and Statistics Multiplication or Division of Random Variables Consider now the product ZDXY;taking X;Zas the new variables. This corresponds to the transformation xDx,yDz=x, with Jacobian JD @x @x@x @z @.z=x/ @[email protected]=x/ @z D 1 0 z=x21=x D1 x; so the marginal distribution of Zis given by g.z/D1Z 1f x;z xdx jxj: (23.78) If the random variables X,Yare independent with densities f1,f2, then g.z/D1Z 1f1.x/f2z xdx jxj: (23.79) Finally, let ZDX=Y, taking Y;Zas the new variables, corresponding to xDyz,yDy, with Jacobian JD @.yz/ @[email protected]/ @z @y @y@y @z D z y 1 0 Dy; and the probability distribution of Zis given by g.z/D1Z 1f.yz;y/jyjdy: (23.80) If the random variables X,Yare independent with densities f1,f2, then g.z/D1Z 1f1.yz/f2.y/jyjdy: (23.81) Gamma Distribution Up to this point the only specific continuous probability distribution we have introduced is the Gauss normal distribution. However, if we make a change in the random variable of that distribution from XtoYDX2, there will result a different distribution of significant utility, known as a gamma distribution. Let’s start the present discussion with the now ArfKen_Ch23-9780123846549.tex 23.6 Transformations of Random Variables 1163 quite familiar normal distribution of mean value zero and variance 2. It has probability distribution f.x/D1 p 2ex2=22: As indicated just after Eq. (23.67), a transformation to write the distribution in terms of yDx2leads us to g.y/Dey=22 p 2y1=2 2: However, this equation does not take into account the fact that ymust be restricted to nonnegative values, and that the same value of ywill be encountered for two different values of x, namely xDCpyandxDpy. These considerations make a more proper and complete formula for g.y/the following: g.y/D8 >< >:0; y0; y1=2ey=22 .22/1=2p;y>0:(23.82) This expression for g.y/is normalized (it must be, due to the way in which it was obtained). However, it is instructive to check, which is best done by changing to a new variable zDy=22, in terms of which we have 1Z 0g.z/dzD1p1Z 0z1=2ezdzD0.1 2/pD1; where we have identified the integral as 01 2 and also noted that 01 2 Dp. Because the functional form of g.y/is essentially that of the integrand of the integral representation of the gamma function, the distribution given by g.y/is called a gamma distribution, and in particular, a gamma distribution with parameters pD1=2(the argu- ment of the gamma function) and 2(the variance of the underlying normal distribution). We generalize to gamma distributions of general pand: g.p;Iy/8 >< >:0; y0; yp1ey=22 .22/p0.p/;y>0:(23.83) The gamma distribution often appears in contexts where the random variables involved need to be added together. It is therefore useful to take note of the Fourier transform of g.p;Iy/: Tg.p;/UT.t/D1p 21 .12i2t/p: (23.84) Using the characteristic-function notation as introduced at Eq. (23.74), and defining Xto be a.p;/gamma-distributed random variable, Eq. (23.84) takes the alternative form heit XiD1 .12i2t/p: (23.85) ArfKen_Ch23-9780123846549.tex 1164 Chapter 23 Probability and Statistics Example 23.6.2 ADDITION OF GAMMA-DISTRIBUTION RANDOM VARIABLES Let’s compute the distribution of a random variable YDX1CX2, where X1has gamma distribution g.p1;Ix1/andX2has gamma distribution g.p2;Ix2/. Note that both X1 andX2have the same variance. Using Eq. (23.77) for the characteristic function of X1CX2and Eq. (23.85) to evaluate heit Xji, we get heitYiD1 .12i2t/p1Cp2: Recognizing this result as the characteristic function for a gamma distribution of parameter pDp1Cp2, we see that g.y/Dg.p1Cp2;Iy/: Generalizing this result to an arbitrary number of Xj: The probability distribution for a sum of gamma-distributed random variables Xjof parameters pjbut all of the same is a gamma distribution for that and with pDP jpj. A corollary to the above is obtained if we consider the probability distribution of a sum of the form ZDnX jD1X2 j; (23.86) where the Xjare Gauss normal distributions, all with the same variance 2. Because the quantities being summed are squares of random variables, it is useful first to make the substitutions YjDX2 j, changing each distribution of Xjto that of a Yjgamma distribution with pD1=2, and finally combining the ngamma distributions to form the distribution ofZ; the result will be a gamma distribution with pDn=2and the common value of . Summarizing, The probability distribution for the sum of the squares of nGauss normal random variables with a common variance 2, as in Eq. (23.86), will be a gamma distribution with parameters pDn=2and the common value of .  Exercises 23.6.1 LetX1;X2;:::; Xnbe independent normal random variables with the same mean Nxand variance2:Show thatP iXi=nNx pn is normal with mean zero and variance 1. 23.6.2 If the random variable Xis normal with mean value 29and standard deviation 3;what can you say about the distributions of 2X1and3XC2? ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1165 23.6.3 For a normal distribution of mean value mand variance 2;find the distance rsuch that half the area under the bell shape is between mrandmCr: 23.6.4 IfhXi;hYiare the average values of two independent random variables X;Y;what is the expectation value of the product XY? 23.6.5 IfXandYare two independent random variables with different probability densities and the function f.x;y/has derivatives of any order, express hf.X;Y/iin terms of hXiandhYi:Develop similarly the covariance and correlation. 23.6.6 Letf.x;y/be the joint probability density of two random variables X,Y. Find the variance2.aXCbY/;where a;bare constants. What happens when X,Yare independent? 23.6.7 Obtain an addition theorem for the distribution of a random variable YDX1CX2 where X1andX2are Gauss normal distributions with different mean values jand variances2 j. ANS. Yis normal with mean 1C2and variance 2 1C2 2. 23.6.8 Show that the Fourier transform of the gamma-distribution probability density, Eq. (23.83), has the functional form given in Eq. (23.84). 23.7 S TATISTICS In statistics, probability theory is applied to the evaluation of data from random experi- ments or to samples to test some hypothesis because the data have random fluctuations due to lack of complete control over the experimental conditions. Typically one attempts to estimate the mean value and variance of the distributions from which the samples derive, and to generalize properties valid for a sample to the rest of the events at a pre- scribed confidence level. Any assumption about an unknown probability distribution is called a statistical hypothesis. The concepts of tests and confidence intervals are among the most important developments of statistics. Error Propagation When we measure a quantity xrepeatedly, obtaining the values xj, or select a sample for testing, we can compute NxD1 nnX jD1xj; 2D1 nnX jD1.xjNx/2; whereNxis the mean value and 2is the variance, a measure of the spread of the points about the mean value. We can write xjDNxCej;where ejis the deviation from the mean value, and we know thatP jejD0. Now suppose we want to estimate the value of a known function f.x/based on these measurements xj; that is, we want to assign a value of fgiven the set fjDf.xj/. ArfKen_Ch23-9780123846549.tex 1166 Chapter 23 Probability and Statistics Substituting xjDNxCejand forming the mean value NfD1 nX jf.xj/D1 nX jf.NxCej/ Df.Nx/C1 nf0.Nx/X jejC1 2nf00.Nx/X je2 jC Df.Nx/C1 22f00.Nx/C; (23.87) we obtain the average value Nfasf.Nx/in lowest order, as expected. But in second order there is a correction given by the variance with a scale factor f00.Nx/=2. It is also of interest to determine the spread predicted for the values of f.xj/. To lowest order, this is given by the average of the sum of squares of the deviations. Approximating fjasNfCf0.Nx/ej, we get 2.f/1 nX j.fjNf/2Tf0.Nx/U21 nX je2 jDTf0.Nx/U22: (23.88) In summary, we may formulate somewhat symbolically f.Nx/Df.Nx/f0.Nx/ as the simplest form of error propagation for a function of one measured variable. For a function f.xj;yk/fjkof two quantities xjDNxCuj,ykDNyCvk, where the xj andykare measured independently of each other and we have rvalues of jandsvalues ofk, we obtain similarly NfD1 rsrX jD1sX kD1fjkDf.Nx;Ny/C1 rfxX jujC1 sfyX kvkC f.Nx;Ny/: (23.89) The error inNfis seen to be second-order in the ujandvk. In writing Eq. (23.89) we have used the relationsP jujDP kvkD0and have introduced the definitions fxD@f @x Nx;Ny;fyD@f @y Nx;Ny: (23.90) The variance of fis (to first order) 2.f/D1 rsrX jD1sX kD1.fjkNf/2D1 rsX j;k.ujfxCvkfy/2Df2 x rX ju2 jCf2 y sX kv2 k; ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1167 where we have dropped the zero cross termP j;kujvkDP jujP kvk. Noting thatP ju2 jDr2 xandP kv2 kDs2 y, we reach the final result 2.f/D1 rsX j;k.fjkNf/2Df2 x2 xCf2 y2 y: (23.91) Symbolically, the error propagation for a function of two measured variables may be summarized as f.Nxx;Nyy/Df.Nx;Ny/q f2x2xCf2y2y: Example 23.7.1 REPEATED MEASUREMENTS As an application and generalization of the result given in Eq. (23.91), let’s consider what happens when we regard the mean of nmeasurements xjas a function NxDf.x1;x2;:::Cxn/D.x1Cx2CC xn/=n of the variables x1;:::; xn, each with variance 2. Then we have fxjD1=nfor each j, and, according to Eq. (23.91), 2.Nx/DnX jD1f2 xj2DnX jD12 n2D2 n: (23.92) This result indicates that the standard deviation of the mean value, .Nx/, will decrease with the number of repeated measurements, approaching zero as =pn. It is important to recognize the distinction between the variance of the mean value, denoted 2.Nx/, and the corresponding quantity for the individual measurements (denoted 2). If we refer now to our earlier result that the sum of nidentically distributed, Gauss normal random variables is also a normal random variable with a variance equal to ntimes that of each variable (see Example 23.6.1), and note also that division of the sum by n(to form the mean) causes division of the variance by n2, as discussed following Eq. (23.28), we find a result identical to that developed in the present example, but with the additional feature that the mean is also normally distributed.  The arithmetic mean Nxwill, because of the distribution in the xj, differ from the true (but unknown) value , withandNxdiffering by some amount , orNxDC . However, as the number nof measurements increases, we expect that the error will tend to zero, and that, according to Example 23.7.1, we can estimate to fall in the range .=pn< <=pn/. We can refine this estimate by considering the spread of the xjmeasured with respect to the true value , meaning that we compute the variance using the average of v2 j, wherevjDxj, instead of that of e2 j, where, as before, ejDxjNx. Calling this version of the variance s2, we write s2D1 nnX jD1v2 jD1 nnX jD1.ejC /2D1 nnX jD1e2 jC 2; (23.93) ArfKen_Ch23-9780123846549.tex 1168 Chapter 23 Probability and Statistics where the term linear in ejvanishes becauseP jejD0. Inserting now an estimate of , in the form 2s2=n(this approximation good to first order), Eq. (23.93) rearranges to s2 11 n D1 nnX jD1e2 j; (23.94) equivalent to sDsP j.xjNx/2 n1: (23.95) The quantity sis referred to as the sample standard deviation. Equation (23.95) is not well defined when nD1, but that is not an issue because a single data point is insufficient to determine a spread. The presence of n1, in contrast to the factor ninEq. (23.23), allows for the probable error in Nx, and is known as Bessel’s correction to the standard deviation formula. Fitting Curves to Data Suppose we have a sample of measurements yjtaken at times tj, where the time is known precisely but the yjare subject to experimental error. An example would be snapshots of the position of a particle in uniform motion at the times tj. Our statistical hypothesis, motivated by Newton’s first law and the initial condition that yD0when tD0, is that y.t/ satisfies an equation of the form yDat, where the constant ais to be determined from the measurements. To fit our equation to the data, we first minimize the sum of the squares of deviations SDP j.atjyj/2to determine the slope parameter a, also called regression coefficient, using the method of least squares. Differentiating Swith respect to awe obtain 2X j.atjyj/tjD0; which we can solve for a: aDP jtjyjP jt2 j: (23.96) Note that the numerator is built like a sample covariance, the scalar product of the variables t;yof the sample. As shown in Fig. 23.9, the measured values yjdo not as a rule lie on the line. They have the sample standard deviation, computed from Eq. (23.95), sDsP j.yjatj/2 n1: Alternatively, suppose that the yjvalues are known precisely while the tjare measure- ments subject to experimental error. As suggested by Fig. 23.10, in this case we need to ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1169 y t 0 FIGURE 23.9 Straight line fit to data points .tj;yj/with tjknown, yjmeasured. y t 0 FIGURE 23.10 Straight line fit to data points .tj;yj/with yjknown, tjmeasured. interchange the roles of tandyand to fit the line tDbyto the data points. We minimize SDP j.byjtj/2, setting dS=dbD0, and find similarly the slope parameter bDP jtjyjP jy2 j: (23.97) In case both tjandyjhave errors (we take tandyto have the same measurement precision), we have to minimize the sum of squares of the deviations of both variables. It is convenient to fit to a parameterization tsin ycos D0, soy=tDsin =cos Dtan , meaning that is the angle the fitting line makes with the t-axis (see Fig. 23.11). Our task will therefore be to determine . We also see from Fig. 23.11 that the fitting line has to be drawn so that the sum of the squares of the distances djof the points .tj;yj/from the line becomes a minimum. To find dj, we rotate our coordinate system the angle , which moves tj;yj to t0 j;y0 j according to t0 j y0 j! D cos sin sin cos ! tj yj! ; ArfKen_Ch23-9780123846549.tex 1170 Chapter 23 Probability and Statistics y ttjujvjdj yj tPj 0y t• 0 (a) (b)α FIGURE 23.11 (a) Straight line fit to data points .tj;yj/. (b) Geometry of deviations uj;vj;dj. which yields djDy0 jDt jsin Cyjcos , the (signed) distance to the line at angle . The minimum of the square of the distances to the line is found from d d X jd2 jD2X j.t jsin Cyjcos /.t jcos yjsin / Dsin cos X j t2 jy2 j cos2 sin2 X jtjyjD0; which can be reduced to tan 2 D2P jtjyj P j t2 jy2 j: (23.98) This least-squares fitting is appropriate when the measurement errors are unknown, as it gives equal weight to the deviation of each point from the fitting line. Finally, if we have information that permits the assignment of different probable errors to different points, we have the alternative of making a “weighted” least-squares fit called achi square fit, which we discuss in the next subsection. The2Distribution Given a set of ujcorresponding to values tjof an independent variable (which is not nec- essarily a time), we seek to fit these data to a function u.t;a1;a2;:::/ , where the aiare parameters that are adjusted to optimize the fit. The optimization is carried out by mini- mizing a weighted sum of the squares of the deviations, where the weights are controlled by the assumed standard deviations jof the respective measurements uj. The quantity to be minimized is traditionally labeled 2and called chi-square, and its precise definition is 2DnX jD1uju.tj;a;:::/ j2 ; (23.99) ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1171 where nis the number of data points. This quadratic merit function gives more weight to points with small measurement uncertainties j. The key assumptions adopted to analyze the probability distribution corresponding to the chi-square fit are (1) that each data point is an independent Gauss normal random variable Xjwith zero mean and unit variance, with the unit variance assured by the presence of the jin each term, and (2) that the distribution 2is related to the Xjby 2DnX jD1X2 j: (23.100) Making a chi-square fit requires no knowledge of statistics; we simply apply standard analytical or numerical methods to minimize 2for our set of data points. On the other hand, a knowledge of the chi-square probability distribution will be needed to determine whether we are getting the expected quality from our chi-square fit. In particular, if we wish to determine the probability of the occurrence of our data set based on the chi-square distribution (and possibly assess the adequacy of our assumptions regarding the individual- point variances 2 j), we must undertake further analysis. Our earlier discussion of transformations of random variables included the analysis of sums of normally distributed X2 jof the form given in Eq. (23.100), with the result de- veloped in Example 23.6.2. Specializing to the case at hand, we note that 2will have a gamma probability distribution with parameters pDn=2andD1, so g.2Dy/Dy.n=2/1ey=2 2n=20.n=2/: (23.101) Plots of g.y/for several values of nare given in Fig. 23.12. 0.5 0.45 n=2 n=3 n=4 n=50.35 0.250.15 0.050.4Chi-square densities, n =2, 3, 4, and 5 0.3 0.20.1 002 1 2 10 8 6 4 FIGURE 23.122probability density gn.y/. ArfKen_Ch23-9780123846549.tex 1172 Chapter 23 Probability and Statistics Table 23.2 2Distribution nvD0:8vD0:7vD0:5vD0:4vD0:3vD0:2vD0:1 1 0:064 0:148 0:455 0:708 1:074 1:642 2:706 2 0:446 0:713 1:386 1:833 2:408 3:219 4:605 3 1:005 1:424 2:366 2:946 3:665 4:642 6:251 4 1:649 2:195 3:357 4:045 4:878 5:989 7:779 5 2:343 3:000 4:351 5:132 6:064 7:289 9:236 6 3:070 3:828 5:348 6:211 7:231 8:558 10:645 Note: A data set with ndegrees of freedom will have probability vthat its value of 2exceeds the tabulated value. It is also useful to note the moment-generating function for this distribution: het2iD1 .12t/n=2; (23.102) a result that follows directly from Eq. (23.85). Differentiating Eq. (23.102), we find h2iDd.12t/n=2 dt tD0Dn;h.2/2iDd2.12t/n=2 dt2 tD0Dn.nC2/; (23.103) and therefore 2.2/Dh.2/2ih2i2Dn.nC2/n2D2n: (23.104) These results suggest that typical data with realistically assigned individual-measurement variances would yield a value of 2comparable to the number of data points. However, by calculating P.2>y0/D1Z y0g.y/dy; (23.105) where g.y/is the distribution in Eq. (23.101), we can obtain for any y0the probability that a data set would have a larger spread than that corresponding to 2Dy0. Because it is somewhat laborious to compute the integral in Eq. (23.105), its values are generally obtained by table lookup. A short table of these chi-square data are given in Table 23.2. Before closing this subsection, we need to deal with the fact that our random variables Xjdo not really have zero mean values if the function u.tj;:::/ was chosen based on the available data and therefore was not exact. By reasoning similar to that involved in the discussion leading to Eq. (23.95), it can be shown that if a chi-square fit involves ndata points and the determination of rparameters, the effective number of degrees of freedom is nr, with the implication that the inexactness of the fit in rdegrees of freedom corresponds to a chi-square distribution with nreplaced by nr. Finally, it is worth pointing out that the 2analysis does not really test the assumptions that the data points are independent normal random variables. If these assumptions are not approximately valid, it is unlikely that good chi-square fits can be achieved. ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1173 Example 23.7.2 CHI-SQUARE FIT Let us apply the 2function to a straight-line fit of the type shown in Fig. 23.9 , with the three measured points and their individual standard deviations, written as .tj;ujj/, having the values .1;0:80:1/; .2; 1:50:05/; .3; 2:70:2/: Before proceeding to the chi-square fit, we first fit a line assuming the points to be equally weighted, corresponding to using Eq. (23.96) for the slope a. We find aD1.0:8/C2.1:5/C3.2:7/ 12C22C32D11:9 14D0:850: The sample variance of the points from the line is 2D1 2 T0:81.0:850/U2CT1:52.0:850/U2C.2:73.0:850/U2 D0:0325; and the variance of ais 2.a/DX j@a @uj2 2DX j tjP kt2 k!2 2D2 P kt2 kD0:0325 14D0:00232: Thus, the unweighted fit yields aD0:850p 0:00232D0:8500:048 . Turning now to the chi-square fit, we next minimize 2DX jujatj j2 with respect to a. This process yields @2 @aD2X jtj.ujatj/ 2 jD0; or aDX jtjuj 2 j,X jt2 j 2 j: In our case aD1.0:8/ 0:12C2.1:5/ 0:052C3.2:7/ 0:22 12 0:12C22 0:052C32 0:22D1482:5 1925D0:770: The value we obtained for ais dominated by the middle point, the smallest j; if that point were the only one used, we would have gotten aD1:5=2D0:75. The variance of a,2.a/, is now 2.a/DX j@a @uj2 2 jDX j tj=2 jP kt2 k=2 k!2 2 jD1P kt2 k=2 kD1 1925D0:000519: The chi-square estimate of ais therefore aD0:770p 0:000519D0:7700:023 . ArfKen_Ch23-9780123846549.tex 1174 Chapter 23 Probability and Statistics Our fit has for 2the value 2DT0:81.0:770/U2 0:12CT1:52.0:770/U2 0:052CT2:73.0:770/U2 0:22D4:533: Our problem involves three points and one parameter, and therefore its chi-square distri- bution has two degrees of freedom and, according to Eqs. (23.103) and(23.104), has a mean value of 2 and a variance of 4. Our value of 2, 4.533, is significantly larger than the mean value of the distribution and therefore describes a data set with more spread than would normally be expected for the stated values of j. We can obtain a more quantitative measure of the probability that 2would be at least as large as our value by comparing with the entries in Table 23.2. Using the row of the table for nD2, we see that the prob- ability of getting a spread larger than that of our present data is quite small, only slightly above 0.1.  Student tDistribution The Student tdistribution (sometimes just called the tdistribution) is designed for use with small data sets for which the variance is unknown. This distribution was first de- scribed by W. S. Gosset, who published his work under the pen name “Student” because his employer, the Guinness brewery, would not permit him to publish it under his own name. Gosset considered the probability distribution of a random variable T, of the form TDYpnpS=n: (23.106) For the applications under consideration here, YD1 nnX jD1XjDNX; (23.107) SDnX jD1X2 j: (23.108) Here the Xjare a set of nindependent Gauss normal random variables, each of the same unknown variance 2. The quantity is the (unknown) value of the mean of X. An important feature of Gosset’s choice for Tis that (as we shall shortly show) its proba- bility distribution fn.t/is independent of the variance of the Xj. The procedure for obtaining the probability distribution of Tdepends on the fact that YandSare independent random variables. That is so, but proof is beyond the scope of the present abbreviated discussion. We start by noting that Sis a gamma distribution, with probability distribution g.n;Is/, as given in Eq. (23.85). Next, we proceed to the distribution of UDpS=n, which we denote h.u/. Making a change of variable from sto nu2, and observing that dsD.ds=du/du, we find h.u/Dg.n;Inu2/.2nu/: (23.109) ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1175 This is the probability distribution of the denominator of T. To get the distribution of the numerator, we note that Yis a normal distribution with variance 2=n(see Eq. (23.92)), and mean zero. It, therefore (see Eq. (23.56)), has the distribution we denote r.y/, of the form r.y/Drn 22eny2=22: (23.110) The numerator, ZDYpn, will therefore have a distribution k.z/, where zDypn, so k.z/Dr.z=pn/.dy=dz/D1p 22ez2=22I (23.111) the presence of the factorpncauses the numerator to have variance 2. Finally, we use the formula for the ratio of two independent distributions, Eq. (23.81), to obtain f.t/D1Z 0k.ut/h.u/u du: (23.112) The integration only extends from zero to infinity because the gamma distribution in h.u/ is only nonzero for positive u. Inserting expressions for the quantities in Eq. (23.112), fn.t/D1p 221Z 0eu2t2=22g.n;Inu2/.2nu2/du D1p 221 2n=2n0.n=2/1Z 0eu2t2=22.nu2/.n=2/1enu2=22.2nu2/du D2 nC1pnn 2.nC1/=2 1 0.n=2/1Z 0eu2.t2Cn/=22undu: (23.113) To complete the evaluation, we change variables in the integral to zDu2.t2Cn/=22, thereby making the integral identifiable as a gamma function, so 1Z 0eu2.t2Cn/=22unduD1 222 t2Cn.nC1/=21Z 0z.n1/=2ezdz D1 222 t2Cn.nC1/=2 0nC1 2 : (23.114) Inserting this result into Eq. (23.113) and simplifying, we note that the instances of  entirely cancel, and we are left with fn.t/D0nC1 2 pn0n 2 1Ct2 n.nC1/=2 : (23.115) ArfKen_Ch23-9780123846549.tex 1176 Chapter 23 Probability and Statistics 0.4T density for n =2, 10, 20, and 30 0.35 0.3 0.25 0.2 0.15 0.05 −4 −3 −2 −10.1 01234n=2 n=10 n=20 n=30 FIGURE 23.13 Student tprobability density fn.t/fornD2, 10, 20, and 30. Equation (23.115) is the probability density for the tdistribution with ndegrees of free- dom. This equation shows that we have achieved the desired result, namely that the dis- tribution Tis independent of the variance of the input random variables Xj. Since it is our intention to use the tdistribution for the reduction of experimental data of unknown variance, we have achieved our current objective. Figure 23.13 shows densities fn.t/for several n; an important feature of these curves is that they depend very weakly on n. Confidence Intervals Aconfidence interval for a random variable Xis the range within which xwill fall, not with certainty but with a high probability, the confidence level, which we can choose. If Xhas a probability distribution f.x/, the confidence interval for probability pwill be the range of x, usually symmetrically centered about some value x0, that contains the fraction pof its probability distribution. If this range is bounded by x0dxandx0Cdx, it is customary to write that xDx0dxwith.100 p/% confidence. If, for example, a computed value of xis 0.50 and 90% of its probability distribution falls between xD0:40 andxD0:60, we say that xhas the value xD0:500:10 with 90% confidence. Confidence intervals are usually found by what is called the pivotal method, which involves relating the variable for which we desire a confidence interval to a known proba- bility distribution. The identification and selection of pivotal quantities is in general outside the scope of this text, but for a Gauss normal random variable with zero mean, a suitable pivot is its tdistribution. Referring to Eq. (23.106), this means we can estimate a confi- dence interval for Y(the deviations of the observed mean NXfrom its true value) from the equation YDNXDTpS=npn; (23.116) ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1177 where Tis the random variable corresponding to the tdistribution and Sis the single value obtained by inserting the observed values of the Xiinto Eq. (23.108). We use Eq. (23.116) by inserting into it the range of Tthat corresponds to a total probability p, calculating therefrom the corresponding range of Y. Note that we do not insert a probability distribution forS; we use the value of Sarising from our data. The distribution of Tis an even function of twith a maximum at tD0, as is obvious from Eq. (23.115) andFig. 23.13, and our confidence interval for Twill naturally be cen- tered about zero. Therefore, a confidence interval of probability pwill correspond to a symmetric range of t,.Cp<t<CCp/, such that P.Cp<t<CCp/Dp: Because f.t/is even, we also have P.1<t<CCp/D1 2C.Cp<t<Cp/=2; which is equivalent to P.1<t<CCp/D1Cp 2Op: (23.117) Because of the frequent need to use values of Ccorresponding to various values of Opand degrees of freedom n, these Cvalues have been tabulated and appear in many statistics texts. A short table is included here (Table 23.3). Given a confidence interval for T, we may insert it into Eq. (23.116), which when solved forbecomes DNXTpS=npnDNxTpn: (23.118) From the limiting values for T, we get the corresponding range for , which is valid with the probability of the Trange. Note that except for the range of T, all the quantities on the right-hand side of Eq. (23.118) are to be computed from our sample data. In particular, we need the mean value NXfor our sample and the standard deviation of our data points, DpS=n. Note further that, as with the chi-square distribution, when measured data are used to generate the sample mean and sample standard deviation, the appropriate t distribution to use for ndata points is that with n1degrees of freedom, and in using Table 23.3 Student tDistribution Op nD1 nD2 nD3 nD4 nD5 0:8 1:38 1:06 0:98 0:94 0:92 0:9 3:08 1:89 1:64 1:53 1:48 0:95 6:31 2:92 2:35 2:13 2:02 0:975 12:7 4:30 3:18 2:78 2:57 0:99 31:8 6:96 4:54 3:75 3:36 0:999 318:3 22:3 10:2 7:17 5:89 Note: Entries are the values CinRC 1fn.t/dtDOp, where fn.t/is given in Eq. (23.115), with nthe number of degrees of freedom. ArfKen_Ch23-9780123846549.tex 1178 Chapter 23 Probability and Statistics Eq. (23.118) it is customary to take as the sample standard deviation, as defined in Eq. (23.95). These points are explained more fully in several of the Additional Readings. Example 23.7.3 CONFIDENCE INTERVAL Suppose we have the following random data from a population that can be assumed to have a Gauss normal distribution: 7:12 4:95 6:18 5:69 2:90 8:47; and we wish to determine 90% and 95% confidence intervals for the population mean. Since we have neither the population mean nor variance, but have assumed the popula- tion distribution to be normal, we can use the tdistribution as just outlined. As a prelimi- nary to doing so, we need to calculate the sample mean and standard deviation. Since we have six data points, the number of degrees of freedom will be nD5. We have NXD.7:12C4:95C6:18C5:60C2:90C8:47/=6D5:885; D1 5 .7:125:885/2CC.8:475:885/21=2 D1:9035: Considering first the 90% confidence interval that corresponds to the range .C90<t< C90/withOpD.1Cp/=2D0:95, we read from Table 23.3 the value C90D2:02. Thus, D5:885.2:02/.1:9035/p 5D5:8851:720 (90% confidence): For 95% confidence, we need C95, again for nD5. This time,OpD0:975 , soC95D2:57, and D5:885.2:57/.1:9035/p 5D5:8852:188 (95% confidence): A few final observations are in order. First, we see that by demanding an increase in the confidence level, the interval probably containing the true mean becomes wider. Note that at high confidence levels the probable width can become much larger than the sample standard deviation. Finally, note that even the confidence intervals are sample-dependent. Other data from the same population could generate intervals of different widths. Perhaps oversimplifying, these analyses show that there is no way of converting probability data into significant statements that have complete certainty.  Exercises 23.7.1 Let1Abe the error of a measurement of A;etc. Use error propagation to show that .C/ C2 D.A/ A2 C.B/ B2 holds for the product CDABand the ratio CDA=B. ArfKen_Ch23-9780123846549.tex 23.7 Statistics 1179 23.7.2 Find the mean value and standard deviation of the sample of measurements x1D6:0, x2D6:5,x3D5:9,x4D6:1,x5D6:2. If the point x6D6:1is added to the sample, how does the change affect the mean value and standard deviation? 23.7.3 Carry out a2analysis of the fit corresponding to Fig. 23.10 using the same points as inExample 23.7.2, but with the errors now associated with the tirather than the yi. 23.7.4 Using the data from Exercise 23.7.2 (including the point x6), find the 90% and 95% confidence intervals for the mean of the xi. Additional Readings Bevington, P. R., and D. K. Robinson, Data Reduction and Error Analysis for the Physical Sciences, 3rd ed. New York: McGraw-Hill (2003). Chung, K. L., A Course in Probability Theory Revised, 3rd ed. New York: Academic Press (2000). DeGroot, M. H., Probability and Statistics, 2nd ed. Reading, MA: Addison-Wesley (1986). Devore, J. L., Probability and Statistics for Engineering and the Sciences, 5th ed. New York: Duxbury Press (1999). Freund, J. E., and R. E. Walpole, Mathematical Statistics, 4th ed. Englewood Cliffs, NJ: Prentice Hall (1987). This well-regarded text is at a level comparable to the exposition in this chapter. Clear and with many statistical tables. Kreyszig, E., Introductory Mathematical Statistics: Principles and Methods. New York: Wiley (1970). Montgomery, D. C., and G. C. Runger, Applied Statistics and Probability for Engineers, 2nd ed. New York: Wiley (1998). Papoulis, A., Probability, Random Variables, and Stochastic Processes , 3rd ed. New York: McGraw-Hill (1991). Ramachandran, K. M., and C. P. Tsokos, Mathematical Statistics with Applications. New York: Academic Press (2009). Relatively detailed but readable and self-contained. Ross, S. M., First Course in Probability, 5th ed., vol. A. New York: Prentice Hall (1997). Ross, S. M., Introduction to Probability and Statistics for Engineers and Scientists, 2nd ed. New York: Academic Press (1999). Ross, S. M., Introduction to Probability Models, 7th ed. New York: Academic Press (2000). Suhir, E., Applied Probability for Engineers and Scientists. New York: McGraw-Hill (1997). ArfKen_Index.tex INDEX Page numbers followed by ‘ f’ and ‘ t’ indicate figures and tables, respectively. Numbers 0-forms, 233 1-D axial Green’s function, 684 1-forms, 233 2-D integration, 74f region, 70, 71f 2-forms, 233 3-forms, 233 A Abel equation, generalized, 1055–1056 Abel’s test, 23 abelian group, 816, 822 absolute convergence, 13, 23, 29 abstract group, 819, 819 t addition of gamma-distribution random variables, 1164 of tensors, 208 of random variables, 1160–1161 addition by scalar, 255 addition rule, for probabilities, 1128 addition theorem application Laplace expansion, 799–801 spherical wave expansion, 798–799 for spherical harmonics, 797–798 for normal distributions, 1161 adjoint matrix, 105 adjoint operator, 277, 297 and scalar product, 278 basis expansion of, 281–282 finding, 278 affine transformation, 1104 algebraic formula, 51 aliasing, 1006 alternating series absolute convergence, 13 Leibniz criterion, 11–12 angular momentum, 126, 126 f, 299 angular momentum operators, 774–776 angular momentum formulas, 781–782 exercises, 782–784 ladder operators, 776–779 spinor, 779–781coupling, 784–786 Clebsch-Gordan coefficients, 789–791 exercises, 795–796 ladder operators construction, 788–795 ofpanddelectrons, 793–795 spinors, 792–793 vector model, 786–788 angular momentum formulas, 781–782 angular momentum operators, 774–776 angular momentum formulas, 781–782 exercises, 782–784 ladder operators, 776–779 spinor, 779–781 annihilation operator, 882 anti-Hermitian, 277 anti-Hermitian matrices, 108, 319 antiderivation, 239 antisymmetric stretching mode, 323 antisymmetric tensor, 208, 216 arbitrary probability distribution, 1157 arbitrary-vector technique, 167 Argand diagram, 56, 56f, 57, 470, 492 arithmetic mean, 1137 associated Laguerre equation, 894 associated Laguerre polynomials generating function, 892–895 associated Legendre equation, 425, 716, 741–743 exercises, 753–756 magnetic field of current loop, 748–753 orthogonality, 746–748 parity and special values, 746 associated Legendre functions, 425, 744–745, 745t associated Legendre polynomials, 743–744 associative, 96, 97 asymptotic expansions, 581 asymptotic forms, 692–693 of Hankel functions, 688–690 properties of, 693–695 exercises, 695–698 of an integral representation, 690–692 Stokes’ method, 688 1181 ArfKen_Index.tex 1182 Index asymptotic series, 577 Bessel functions, 691 cosine and sine integrals, 580–582 definition of, 582 –583 exercises, 583–584 exponential integral, 578–580 integral representation expansion, 690–692 overview, 577 asymptotic values, Bessel functions, 690, 703 atomic interaction integral, 72 average value, 1136 axial Green’s function, 464 axial vectors, 136, 215 B Baker-Hausdorff formula, 114 band-pass filter, 1002 f baryons, 852, 853 t multiplets, decomposition of, 858–861, 859 f, 860f basis expansion, adjoint, 281–282 basis functions, 252, 253 Bayes’ theorem, 1130 Bernoulli equation, 330, 378 Bernoulli numbers, 556, 562 t contour of integration for, 563 f exercises, 566–567 generating-function, 560 overview, 560–565 polynomials, 565–566 Riemann zeta function, 564 Bernoulli polynomials, 565–566 Euler-Maclaurin integration formula, 567 Bessel functions, 67 asymptotic expansions asymptotic forms, 692–695 exercises, 695–698 Hankel functions, asymptotic forms of, 688–690 of an integral representation, 690–692 asymptotic values, 690 of first kind Bessel’s differential equation, 646–647 confluent hypergeometric representation, 919 cylindrical resonant cavity, 650–653 exercises, 654–661 Fraunhofer diffraction, circular aperture, 648–650 Frobenius method, 643 generating function for integral order, 644–645 integral representation, 647–648 modified, 681 orthogonality, 661 recurrence relations, 645–646second kind, 644 Wronskian, 670–671 Hankel functions contour integral representation of, 676–678 definitions, 674–675 exercises, 678–680 Helmholtz equation, 680, 698, 705 hyperbolic, 683 Laplace equation, 651 modified, 428, 643, 678, 682 f asymptotic expansion, 688 contours, 696 f exercises, 688 Fourier transforms, 684 Green’s function, 684–685 Hankel function, 682 hyperbolic Bessel functions, 683 integral representation, 683–684, 690–692 Laplace equations, 680 recurrence relations, 681–682 series expansion, 681 Whittaker functions, 682 Neumann functions, Bessel functions of second kind coaxial wave guides, 672 definition and series form, 667–669 exercises, 674 integral representations, 669 recurrence relations, 669–670 uses of, 671 Wronskian formulas, 670–671 orthogonality Bessel series, 663 electrostatic potential in a hollow cylinder, 663–664 exercises, 665–667 Neumann boundary condition, 662 normalization, 662 Sturm-Liouville theory, 661 PDEs, 643 recurrence relations, 645–646 of second kind, 667 Schlaefli integral, 653–654 spherical, 643 asymptotic values, 703 definitions, 702 exercises, 709–712 Helmholtz equation, 698 limiting values, 703 modified, 709 orthogonality and zeros, 703 particle in a sphere, 704–706 recurrence relations, 702 waves, 703 ArfKen_Index.tex Index 1183 of third kind, 675 in wave guides, 671–672 zeros, 648–653 Bessel series, 663 Bessel’s correction, 1168 Bessel’s differential equation, 646–647 Bessel’s equation, 344–345, 366–367, 1025–1027 limitations of series approach, 351–353 Bessel’s inequality, 262 beta function, 617 definite integrals, alternate forms, 618 derivation of Legendre duplication formula, 618–619 exercises, 619–622 binomial coefficients, 34 binomial distribution, 1148–1151 limits of, 1157–1158 and Poisson distribution, 1153–1154, 1154 f binomial expansion, application of, 41–42 binomial probability distribution, 1149, 1150 f binomial theorem, 33–36, 493–494, 581, 716 exercise, 36–40 Biot-Savart law, 750, 751, 751 f black hole, optical path near event horizon of, 1087–1088, 1088 f Bohr radius, 897, 1135 Born approximation, quantum mechanical scattering, 465–466 Bose-Einstein statistics, 1132, 1133 bosons, 840 boundary conditions, 381, 405 Cauchy, 412 Dirichlet, 412, 985 Green’s function, 452–454 hollow cylinder, 664 homogeneous, 448 Neumann, 412 ring of charge, 730 specific, 438–439 sphere in uniform electric field, 728 sphere with, 428–430 waveguide, coaxial cable, 671 boundary curve, 406 boundary value problem, 1052 brachistochrone problem, 1082 branch cut (cut line), 500, 508 f exploiting, 534–537 using, 534–535 branch points, 499–503, 499 f,500 f,502 f,503 f, 503t, 536–537 avoidance of, 532–534 of order 2, 500 Bravais lattice, 869 Bromwich integral, 1038–1040, 1040 f brute-force approach, 32C calculus of residues Cauchy principal value, 512–515, 512 f,515 f computing residues, 510–511 counting poles and zeros, 518–519 exercises, 520–522 pole expansion of meromorphic functions, 515–518 product expansion of entire functions, 519–520 residue theorem, 509–510, 509 f calculus of variations Euler equation, 1081–1085 alternate forms of, 1088 exercises, 1093–1096 optical path near event horizon of a black hole, 1087–1088, 1088 f soap film, 1088–1090, 1089 f soap film–minimum area, 1090–1093, 1092 f straight line, 1086 Lagrangian multipliers, 1107–1109 Rayleigh-Ritz variational technique, 1117–1118 ground state eigenfunction, 1118–1119 several dependent variables, 1096–1097, 1102 exercises, 1105–1107 Hamilton’s principle, 1097–1098 Laplace’s equation, 1101–1102 moving particle–Cartesian coordinates, 1098–1099 moving particle–circular cylindrical coordinates, 1099 several independent variables, 1100–1102 variation with constraints, 1111–1112 exercises, 1121–1124 Lagrangian equations, 1112–1113 Schrödinger wave equation, 1116–1117 simple pendulum, 1113–1114, 1113 f sliding off a log, 1114–1115, 1114 f canonical momentum, 1099 Cartesian coordinate system, 47 Cartesian coordinates, 415–420 spherical harmonics using, 758 Casimir operators, 849 Catalan’s constant, 13, 572, 613 catenoid, catenary of revolution, 1090 Cauchy (Maclaurin) integral test, 5–8 Cauchy boundary conditions, 412 Cauchy criterion, 2 Cauchy inequality, 490 Cauchy principal value, 512–515, 512 f,515 f, Cauchy ratio test, 1065 Cauchy root test, 4 Cauchy’s integral formula, 486–487, 554, 591 applications of, 490 derivatives, 488 exercises, 491–492 ArfKen_Index.tex 1184 Index Cauchy’s integral formula ( continued ) Morera’s theorem, 489–490 Cauchy’s integral theorem contour integrals, 477–478, 478 f exercises, 485, 486 f Goursat proof, 481–482, 481 f Laurent expansion, 492–497 multiply connected regions, 483–484, 483 f statement of, 478–481, 480 f Cauchy-Riemann conditions, 471–477 analytic functions, 472–474 derivatives of, 474–475 exercises, 476–477 overview, 471–472 point at infinity, 475 Cauchy-Riemann differential equations, 591 causality, 591 cavities, cylindrical, 650–653 central field potential, Laplacian of, 154 central force, 192 central force problems, 426 central moments, 1141 chain rule, 63 chaotic behaviour, 377 character, 831, 832 t characteristic curves, 405 characteristic equation, 303 characteristic function in probability theory, 1160 characteristic polynomial, 303 characteristics of PDEs, 404–406 charge density, 739 Chebyshev differential equation, 388 Chebyshev inequality, 1140 Chebyshev polynomials exercises, 907–911 generating functions, 899 hypergeometric representations, 914 numerical analysis, 905–906 orthogonality, 906–907 recurrence relations, 901–903 shifted, 908 trigonometric form, 904–905 type I, 900 type II, 899 ultraspherical polynomials, 899 chi square fit, 1170, 1173–1174 chi-square (2) distribution, 1170–1174 Christoffel symbols, 222 evaluating, 223–224 circle of convergence, 493 circular contour znon, 479 circular cylindrical coordinates, 187–190, 188 f, 421, 431 cylindrical eigenvalue problem, 422–424circular disk, rotations of, 818 circular functions, 58–59 circular membrane, Bessel functions, 659 circular optical path, 1088 f circular wave guide, 672 circular wire loop, 931 f classes, 830–835 Clausen functions, 949 Clebsch-Gordan coefficients, 789–791 Clifford algebra, 112 closed loop, 499–501, 499 f,500 f closure relation, 264 coaxial wave guides, 671–672 coefficient vector, 261 colatitude, 72 collinear velocities, addition of, 864 column vector, 95, 123, 125 extraction of, 108 combinations, counting of, 1130–1133 commutation rules, 785–786 commutative, 96, 816 commutative operation, 47 commutator, 97, 276 comparison tests, 3–4 completeness, 255, 262 of Hilbert-Schmidt of integral equations, 1073 complex conjugation, 54, 470 complex exponentials, integrals with, 527–531, 529f complex numbers and functions, 53 Cartesian components, 53 circular and hyperbolic functions, 58–59 complex domain, 55–56 exercises, 60–61 imaginary numbers, 54 logarithm, 60 polar representation, 56–58 powers and roots, 59 multiplication of, 54 of unit magnitude, 57 complex plane, 56 complex variable theory, 53 complex variables, see also Cauchy-Riemann conditions; mapping; singularities algebra using, permanence of algebraic form, 55 Cauchy’s integral formula, 591 causality, 591 dispersion relations ArfKen_Index.tex Index 1185 exercises, 596–597 optical dispersion, 594–595 overview, 591–593 Parseval relation, 595–596 symmetry relations, 593 functions of, 470 computing residues, 510–511 conditional convergence, 13 conditional probability, 1128 conditional probability distributions, 1147 Condon-Shortley phase, 758, 760 t confidence interval, 1176–1178 confluent hypergeometric functions, 912 asymptotic expansions, 919 Bessel and modified Bessel functions, 918–919 exercises, 920–922 Hermite functions, 919 Laguerre functions, 919 Whittaker function, 919 Wronskian, 922 conformal mapping, 549 conjugate subgroup, 820 conjugation, complex, 56, 105 connected, simply, 164 conservation laws, 815 conservative force, 171, 244 constant B field, vector potentials of, 172 constant coefficients, with ODEs, 342–343 constrained minima/maxima, 1107–1109 exercises, 1110 contiguous function relations, 913 continuous deformation, 484 continuous groups, 816, 845–846 exercises, 861 homomorphism SU(2)–SO(3), 851–852 Lie groups and their generators, 846–849 of representation, 824–825 SO(2) and SO(3), 849–851 SU(3), 852 continuous random variable, 1135–1137 contour integral, 477–478, 478 f contour integral representation, 676–678 contour integration, 967, 967 f singularity on, 530–531 methods, 572, 603 contraction, 209–210 contravariant basis vectors, 220–221 contravariant metric tensor, 219 contravariant tensors, 206–207 contravariant vectors, 206, 219 convergence infinite products, 575 infinite series, partial sum approximation, 579 of Neumann series, 1066 convergence in the mean, 262convergence of infinite series absolute, 13, 23 of power series, 29 rate, 16 tests, see also Cauchy (Maclaurin) integral test comparison, 3–4 Gauss’, 9 improvement of, 16–17 Kummer’s, 8–10 uniform and nonuniform, 21–22 convergence, rate of, 16 convolution (Faltungs) theorem driven oscillator with damping, 1035–1037 Parseval relation, 987–990 coordinate transformations exercises, 138–139 of orthogonal, 135 of reflections, 136–137, 137 f of rotations, 133–135 of successive operations, 137–138 coordinates, see also Cartesian coordinates; circular cylindrical coordinates; orthogonal coordinates; spherical polar coordinates curvilinear, 182 correlation, 1142–1144 cosines asymptotic expansion, 581, 582 confluent hypergeometric representation, 920 infinite products, 575 integral of in denominator, 523–524 integrals cos in asymptotic series, 580–582 Coulomb’s law, 447, 730 counting poles and zeros, 518–519 coupling, angular momentum, seeangular momentum covariance, 1142–1144 covariance of Maxwell’s equations, Lorentz, see Lorentz covariance of Maxwell’s equations covariant, 862 covariant basis vectors, 218, 220–221 covariant derivatives, 222–223 covariant metric tensor, 219 covariant tensors, 206–207 covariant vector, 206 Cramer’s rule, 84 creation operator, 882 criterion, Leibniz, 11–12 cross derivatives, 62 cross product, 126–128, 126 f,127 f,see also triple vector product crossing conditions, 593 crystallographic point groups, 869 curl,r, 149–153 circular cylindrical coordinates, 193 in curvilinear coordinates, 186–187, 186 f ArfKen_Index.tex 1186 Index curvilinear coordinates, 182 differential operators in, 185–187, 185 f,186 f exercises, 196–203 integrals in, 184–185 cut line (branch cut), 500, 508 f exploiting, 534 using, 534–535 cylindrical symmetry, 443 cylindrical traveling waves, 694–695 D d’Alembert ratio test, 4–5, 55, 578 d’Alembert’s solution, of wave equation, 436 damped oscillator, 1021 de Moivre’s Theorem, 59 decuplet, 859, 859 f defective matrices, 324 definite integral (Euler), 600–601 definite integrals, 580 evaluation of, 522 exercises, 538–544 degeneracy, 307–308 degenerate, 1057, 1072 delta function, Dirac, 263–265, 1010–1011 -sequence function, 76f Dirichlet kernel, 77 exercise, 80–81 Fourier series, 77 Kronecker delta, 79–80 properties of, 78–79 sequence, 76 spherical polar coordinates, 79 denominator, integral of cos in, 523–524 dependent variable, 329 derivative operators, tensor, seetensor derivative operators derivatives, see also exterior derivatives chain rule, 63 cross derivatives, 62 exercises, 64 mixed derivatives, 401 partial derivatives, 62, 401 stationary points, 63–64 determinants, 295 and linear dependence, 89–90 derivatives of, 102 exercises, 93–94 homogeneous linear equations, 83–84 inhomogeneous linear equations, 84 product theorem, 103–104 properties of, 87 deuteron, 391–393 diagonal matrices, 99 eigenvalues, 313 eigenvector, 312diagonalization matrices, 311–314 simultaneous, 314–315 differentiable manifolds, 233 differential equations first-order differential equations, 331–342 exact differential equations, 333 exercises, 339–342 homogeneous equations, 334–335 isobaric equations, 335 linear first-order ODEs, 336–339 nonseparable ODEs, 333–334 parachutist, 331–332 RL circuit, 338–339 separable equations, 331 Fuchs’ theorem, 355 linear independence of solutions, 358–360 second solution, 362–363 series form of the second solution, 364–366 nonlinear, 377–380 number of solutions, 361 partial differential equations, 329 particular solution, 337 second solution exercises, 370–374 finding, 362–363 for linear oscillator equation, 363 logarithmic term, 668 Neumann functions, 368–369 of Bessel’s equation, 366–367 series solutions, Frobenius method, 346–350, 350f exercises, 355–358 expansion about, 350 Fuchs’ theorem, 355 limitations of series approach, Bessel’s equation, 351–353 regular and irregular singularities, 353–354 symmetry of solutions, 350–351 singular points, 343–345, 345 t differential forms, 232 0-forms, 233 1-forms, 233 2-forms, 233 3-forms, 233 complementary, 235–236 exercises, 238, 243, 248 exterior algebra, 233–235 exterior derivatives, 238–243 Hodge operator, 235 in Minkowski space, 236–237 integration of, 243–248 Maxwell’s equations, 241–243 miscellaneous, 237–238 simplifying, 234 ArfKen_Index.tex Index 1187 Stokes’ theorem on, 245 three-dimensional (3-D), 407 differential operators, 275 differential vector operators, 143 gradient, 143 properties, 153–157 exercises, 157–159 differentiate parameter, 67 differentiation of forms, 238–243 power series, 30 diffraction, 648–650 diffusion partial differential equations, 437–444 digamma and polygamma functions, 610 digamma functions, 610–611 exercises, 614–616 Maclaurin expansion, computation, 613 polygamma function, 612 series summation, 613 dihedral, 818 dilogarithm exercises, 926 expansion and analytic properties, 923–924 properties and special values, 924–926 dimensionality theorem, 831 dipole moment, 738 Dirac braket notation, 265 Dirac delta distribution, 972 Dirac delta function, seedelta function, Dirac Dirac gamma matrices, 112 Dirac half-braket notation, 265 Dirac matrices, 111 Dirac notation, 265–266 Dirac’s relativistic theory, 38 direct product, 108–112, 837–839 exercises, 837 f, 840, 840 t generators for, 857–858 of tensors, 210–211 direct space, 964 Dirichlet boundary conditions, 385, 412, 704 Dirichlet conditions, 936, 985 Dirichlet kernel, 77 Dirichlet series exercises, 573–574 overview, 571–572 discontinuous functions, 937–939 expansions in, 262–263 discrete Fourier transform, 1002–1007 aliasing, 1006 exercises, 1007 fast Fourier transform, 1006–1007 limitations, 1005 orthogonality over discrete points, 1002–1004 discrete groups, 815 classes, 830–835exercises, 835–837, 836 t,837 f other, 835 discrete probability distributions, computing, 1136 discrete random variables, 1134–1135 discrete spectrum, 420 dispersion integral contour for, 592 f dispersion relations causality, 591 crossing conditions, 593 exercises, 596–597 Hilbert transforms, 593, 595 optical dispersion, 594–595 overview, 591–593 Parseval relation, 595–596 sum rules, 596 symmetry relations, 593–594 divergence, r, 146–149, 149 f curvilinear coordinates, 185–186, 185 f divergent series, 4 division, of random variables, 1162 Doppler shift, 37 dot products, 49–50 gradient of, 143, 157 double factorial notation, 35 double series, rearrangement of, 18–19 driven oscillator with damping, 1035–1037 dual tensors, 216–217 E Earth’s gravitational field, 727 Earth’s nutation, 1018–1019, 1018 f eigenfunction, 299 eigenfunction completeness of Hilbert-Schmidt of integral equations, 1073 orthogonal, 1069–1073 eigenfunction expansion of Green’s function, 460–461 eigenvalue problem, 422–424 eigenvalues equations, 299–300 basic expansions, 300 equivalence of operator and matrix form, 300 of Hermitian matrices, 310 of Hilbert-Schmidt theory, 1073 eigenvectors, 300 normalizing, 304 of Hermitian matrices, 310–311 Einstein convention, 207 electric dipole, 738, 737 f electric multipoles, 737–739 electric quadrupole, 738 electromagnetic field tensor, 866 electromagnetic wave equation, 156 electromagnetic waves, 1023–1024 electromagnetism, potentials in, 174 ArfKen_Index.tex 1188 Index electron spin, 846 electrostatic potential for ring of charge, 729–730 in hollow cylinder, 663–664 elementary functions, 1008–1010 elliptic integrals definitions of, 928–929 exercises, 931–932 of first kind, 928 limiting values, 930 period of simple pendulum, 927–928 of second kind, 928 series expansion, 929–930 elliptic partial differential equations (PDEs), 410 empty set, 1127 energy, relativistic, 35–36 entire function, 519–520 equality of matrices, 96 equation of continuity, 148 equations, see also Maxwell’s equations motion and field, 213 equilateral triangle, symmetry of, 817, 817 f,818 f error function, 637 error propagation, 1165–1168 essential (irregular) singular point, 344 essential singularities, 344, 498 Euclidean space, 237 Euler angles, 140, 140 f Euler equation, 1081–1085 alternate forms of, 1088 exercises, 1093–1096 soap film, 1088–1090, 1089 f soap film–minimum area, 1090–1093, 1092 f straight line, 1086 Euler identity, 113 Euler transformation, 43, 44 Euler-Maclaurin integration formula, 566 Bernoulli polynomials, 567 example, 569–570 exercises, 570–571 overview, 567–569 Euler-Mascheroni constant, 7, 33, 367, 675 event horizon, 1087 evolution operator, 1067 exact ODEs, 333–334 expansion, 736, see also Taylor’s expansion Laplace expansion, 760–762, 799–801 pole, of meromorphic functions, 498, 515–518 product, of entire function, 519–520 spherical harmonic, 761 spherical wave, 798–799 expectation value, 283, 285, 297, 1136 in transformation basis, 295 exponential function, of Maclaurin theorem, 27–28exponential integral, 578–580, 634–637 exterior algebra, 233 exterior derivatives, 238–243 exterior products, 233 extrema, 62–64 F factorial function, asymptotic form of, 588 factorial notation, 606 faithful group, 822 Faraday’s law, 168 fast Fourier transform (FFT), 1006–1007 Feldheim’s formula, 885 Fermi-Dirac statistics, 1132, 1133 fermions, 840 FFT, seefast Fourier transform field equations, 213 finite wave train, 971–973, 972 f first-order Born approximation, 1067 first-order differential equations, 331–342 exact differential equations, 333 exercises, 339–342 homogeneous equations, 334–335 isobaric equations, 335 linear first-order ODEs, 336–339 nonseparable ODEs, 333–334 parachutist, 331–332 RL circuit, 338–339 separable equations, 331 first-order partial differential equations, 403 characteristics of, 404–406 exercises, 408–409 general, 406–407 fixed and movable singularities, special solutions, 378–379 flux, 148 Fourier convolution theorem, 1055 exercises, 994–997 multiple convolutions, 990–992 Fourier cosine series, 941 Fourier cosine, sine transforms, 966 Fourier expansions, characteristic of, 951 Fourier integral representation, 969–970 Fourier series, 77 applications of, 949–957 exercises, 952–957 full-wave rectifier, 950–951, 951 f,952 t square wave, high frequencies, 949–950, 949f definition of, 935 general properties, 935–949 discontinuous functions, 937–939 exercises, 945–949 periodic functions, 939–940 sawtooth wave, 937–939, 938 f ArfKen_Index.tex Index 1189 Sturm-Liouville theory, 936–937 summation of a Fourier series, 944 symmetry, 940–941, 942 f Gibbs phenomenon calculation of overshoot, 959–961 exercises, 961–962 square wave, 958–959 summation of series, 957–958 operations on, 942–944 Fourier sine series, 941 Fourier transform, 965–968 aliasing, 1006 convolution theorem, 985–987 of derivatives heat flow PDE, 983 wave equation, 981–982 discrete, seediscrete Fourier transforms exercises, 975–980, 985 fast, 1006–1007 of Gaussian, 969, 968 f inverse, 970–973 limitations on transfer functions, 1000–1001 momentum space representation, 993–994 of product, 992 properties of, 980–984 solution, 1054–1055 successes and limitations of, 984–985 in 3-D space, 973–975 unitary operator, 988 Fourier transforms–inversion theorem, finite wave train, 971–973, 972 f Fourier-Mellin integral, 1039 Fraunhofer diffraction, Bessel function, 648–650 Fredholm equation, 1047, 1052, 1054, 1058, 1064 homogeneous, 1059–1060, 1069 inhomogeneous, 1076–1077 Fresnel integrals, 712 f Frobenius method, 643, 645 series solutions, 346–350, 350 f Fuchs’ theorem, 355, 692 full-wave rectifier, 950–951, 951 f,952 t functions, 143 Chebyshev polynomials exercises, 907–911 generating functions, 899 numerical analysis, 905–906 orthogonality, 906–907 recurrence relations, 901–903 trigonometric form, 904–905 type I, 900 type II, 899 ultraspherical polynomials, 899 confluent hypergeometric functions Bessel and modified Bessel functions, 918–919exercises, 920–922 Hermite functions, 919 Laguerre functions, 919 Whittaker function, 919 dilogarithm exercises, 926 expansion and analytic properties, 923–924 properties and special values, 924–926 Dirac delta, 263–265 discontinuous, 262–263 entire, 498 exponential, of Maclaurin theorem, 27–28 Hermite functions applications of the product formulas, 885–887 direct expansion of products of Hermite polynomials, 884–887 exercises, 876–878, 887–888 Hermite product formula, 884–887 molecular vibration, 882–883 orthogonality and normalization, 875–876 quantum mechanical simple harmonic oscillator, 878–879 recurrence relations, 872–873 Rodrigues formula, 874 threefold Hermite formula, 883–884 values of, 873–874 hypergeometric functions, 911 confluent, 912 contiguous function relations, 913 exercises, 915–916 hypergeometric representations, 913–914 Pochhammer symbol, 912 Laguerre functions associated Laguerre polynomials, 892–895 differential equation–Laguerre polynomials, 890–892 exercises, 897–899 hydrogen atom, 896–897 Rodrigues formula and generating function, 889–890 of complex variables, 470 of operators, 282 orthonormal, 269–271 series expansions, 41–44 excercise, 44–45 series of Abel’s test, 23 exercises, 24–25, 32–33 uniform and nonuniform convergence, 21–22 Weierstrass Mtest, 22–23 square-wave, 263 G Galilean, 862 gamma distribution, 607, 1162–1164 ArfKen_Index.tex 1190 Index gamma function, see also factorial function analytic properties, 604 asymptotic form of, 588–589 beta function definite integrals, alternate forms, 618 derivation of Legendre duplication formula, 618–619 exercises, 619–622 definitions, simple properties, 599 definite integral (Euler), 600–601 factorial notation, 606 infinite limit (Euler), 599–600 infinite product (Weierstrass), 602 incomplete beta function, 634 incomplete gamma functions and related functions, 633–634 error function, 637 exercises, 638–641 exponential integral, 634–637 Riemann zeta function, 626–631 Stirling’s series, 622 derivation from Euler-Maclaurin integration formula, 623–624 gamma function contour, 605 f gamma functional relation, 506 gauge condition, 174 gauge transformations, 174 Gauss elimination, 91–93 Gauss technique, 91 Gauss’ fundamental theorem of algebra, 490 Gauss’ law, 175–176, 175 f Gauss’ normal distribution, 1155–1159 Gauss’ test, 9 Legendre series, 9–10 Gauss’ theorem, 164–165, 165 f, 176, 248 Green’s theorem, 165–166 Gegenbauer polynomials, seeultraspherical polynomials Gell-Mann matrices, 854 general coordinates, tensor in covariant derivatives, 222–223 exercises, 226 metric tensor in, 218–219 general relativity, 862 generalized Abel equation, 1055–1056 generating function, 555, 1056 associated Laguerre polynomials, 892–895 Bernoulli numbers, 560 Bessel functions, modified, 919 Chebyshev polynomials, 899 electric multipoles, 737–739 exercises, 740–741 expansion, 736–737 Hermite polynomials, 555–556, 872 for integral order, 644–645Laguerre polynomials, 889–890 Legendre polynomials, 557–558 physical interpretation of, 735 Taylor expansion of, 565 generators of continuous groups, 846–849 geodesics, 1103–1104 geometric properties, 47 geometric series, 2–3 Gibbs phenomenon calculation of overshoot, 959–961 exercises, 961–962 square wave, 958–959 summation of series, 957–958 Goldschmidt discontinuous solution, 1091, 1092 f Goursat proof of Cauchy’s integral, 481–482, 481f gradient, r as differential vector operator, 143–146 in curvilinear coordinates, 185 of dot product, 157 Gram-Schmidt orthogonalization example, 270–272 exercises, 273–275 orthonormalizing physical vectors, 273 overview, 269–270 physical vectors, 272 vectors by, 269–275 Gram-Schmidt process, 308 Gram-Schmidt transformation, 293 graphene, 869 Grassmann algebra, seeexterior algebra gravitational potential, 172 Green’s function, 447–467, 684–685, 983–984, 1050, 1052, 1072, 1075 advantage of, 452 axial, 464 boundary conditions, 452–454 accomodating, 464 at infinity, 454 initial value problem, 453–454 differential vs. integral formulation, 456 eigenfunction expansion of, 460–461 exercises, 456–459, 466–467 features of, 448, 459–460 form of, 450–452, 461–466 fundamental, 462, 463 t general properties of, 449–450 Helmholtz equation, 463 Laplace’s equation, 462, 464 one-dimensional, 448–459 relation to integral equation, 454–456 self-adjoint problems, 460 spherical, 463, 800–801 two and three dimension problems, 459–467 Green’s theorem, 165–166, 246–247 ArfKen_Index.tex Index 1191 Gregory series, 39 ground state, 391 ground state eigenfunction, 1118–1119 group theory, see also generators of continuous groups; homogeneous Lorentz group definition of, 816–817, 817,818 f,818 t discrete classes, 830–835 other, 835 exercises, 820, 821 f faithfulness, 822 homomorphic, 817 homomorphism and isomorphism, 819 isomorphic, 817 Lorentz covariance of Maxwell’s equations, 866–868 vierergruppe, 820 H Hamilton’s equations, 1099–1100 Hamilton’s principle, 1097, 1098 Hankel functions, 682 asymptotic forms, 692, 693 contour integral representation of, 676–678 definition, 674–675 integral representation of, 698 f series expansion, 675 spherical, 701 Wronskian formulas, 675 Hankel transforms, 965, 1054 harmonic functions, 473 harmonic numbers, 3 harmonic oscillator, 878–879, 1017 harmonic series, 3 harmonics, 799, see also spherical harmonics; vector spherical harmonics Hartree atomic units, 396 heat flow partial differential equations, 437–444, 983 Heaviside shifting theorem, 1023 Heaviside step function, 1010 Heisenberg uncertainty principle, 973 Helmholtz equation, 415, 422, 439, 705 Bessel functions, 680, 698 Green’s function, 463 spherical coordinates, 698 Helmholtz’s theorem, 177–180 Hermite equation, 390–391 Hermite functions applications of the product formulas, 885–887 confluent hypergeometric functions, 919 direct expansion of products of Hermite polynomials, 884–887 exercises, 876–878, 887–888 Hermite polynomials, 872Hermite product formula, 884–887 molecular vibration, 882–883 orthogonality and normalization, 875–876 quantum mechanical simple harmonic oscillator, 878–879 recurrence relations, 872–873 Rodrigues formula, 874 threefold Hermite formula, 883–884 values of, 873–874 Hermite polynomial, see also Legendre polynomials Hermite polynomials, 280, 391, 554, 873 f direct expansion of products of, 884–887 example, 554–556 generating function, 555–556, 872 orthogonality integral, 875 recurrence relations, 872–873 Rodrigues representation, 874 Hermitian matrices, 108, 301 anti-, 319 diagonalization, 311–313 example, 313 exercises, 317–318 expectation values, 316 finding diagonalizing transformation, 313–314 positive definite and singular operators, 317 simultaneous, 314–315 spectral decomposition, 315–316 of eigenvalues, 310 unitary transformation, 313 Hermitian operator, 277, 284 expectation value, 316 self-adjoint ODEs, 384 Hilbert space, 255–256, 278, 279, 289 Hilbert transforms, 593, 595 Hilbert-Schmidt theory homogeneous Fredholm equation, 1069 inhomogeneous Fredholm equation, 1076–1077 inhomogeneous integral equation, 1073–1077 orthogonal eigenfunctions, 1069–1073 symmetrization of kernels, 1069 Hodge operator, 235 homogeneous boundary condition, 448 homogeneous equations, 334–335 homogeneous Fredholm equation, 1059–1060, 1069 homogeneous linear equations, 83–84 ODEs, 330 homogeneous Lorentz group, 862–864 homogeneous ODEs, 335, 338, 351 second-order, 344 homomorphic group, 817 homomorphism, 819 SU(2) and SU(2)–SO(3), 851–852 ArfKen_Index.tex 1192 Index Hooke’s law spring, 342–343 Hubble’s law, 52 hydrogen atom, 896–897 Schrödinger’s wave equation, 896 hyperbolic functions, 58–59 hyperbolic partial differential equations (PDEs), 410 hypercharge, 854 hypergeometric equation alternate forms, 918 singularities, 345, 912, 917 hypergeometric functions, 911 confluent, 912 contiguous function relations, 913 exercises, 915–916 hypergeometric representations, 913–914 Pochhammer symbol, 912 hypergeometric series, seehypergeometric function I identity operator, 277 imaginary axis, 56 imaginary numbers, 54 imaginary part, 56 improper rotations, of coordinate system, 215 impulse function, 1011 impulsive force, 1020 incomplete beta function, 634 incomplete gamma functions, 633–634 of first kind confluent hypergeometric representation, 917 indefinite integral of f.z/, 489 independence, linear, 671 independent variables, 329, 407–408, 411 indeterminate forms, 31 indicial equation, 348 indistinguishable particles, 1133 inertial frames, 815 infinite limit (Euler), 599–600 infinite product (Weierstrass), 602 infinite products convergence, 575 evaluate, 575 exercises, 576–577 overview, 574–575 sine and cosines, 575 infinite series, 1, see also Taylor’s expansion; power series algebra of alternating series, 11–13 convergence, 13 convergence: absolute, 13 convergence: Cauchy integral, 5–8 convergence: Cauchy root, 4convergence: comparison, 3–4 convergence: conditional, Leibniz criterion, 15–16 convergence: d’Alembert ratio, 4–5 convergence: Gauss’, 9 convergence: Kummer’s, 8–10 convergence: Maclaurin integral, 5–8 convergence: test of, 3–11 convergence: uniform, 21–22, 29 divergence of squares, 15–16 double series, 18–19 exercises, 20–21 rearrangement of double, 18–19 exercises, 10–11, 13–14 fundamental concepts geometric series, 2–3 harmonic, 3 of functions Abel’s test, 23 exercises, 24–25 uniform and nonuniform convergence, 21–22 Weierstrass Mtest, 22–23 power series, 29–30 infinity, boundary conditions at, 454 inhomogeneous Fredholm equation, 1076–1077 inhomogeneous integral equation, 1073–1077 inhomogeneous linear equations, 84 inhomogeneous linear ODEs, 375–377 exercises, 377 inhomogeneous Lorentz group, 862 inner product and matrix multiplication, 97–98 integer powers, 59 integers, sum of, 40–41, 41 integral equations boundary condition, 1048 exercises, 1060–1064 feature of, 1048 Fredholm equation, 1047, 1052, 1054, 1058–1060, 1069 generating-function, 1056–1057 Green’s function, 454–456 Hilbert-Schmidt theory exercises, 1077–1079 homogeneous Fredholm equation, 1059–1060 orthogonal eigenfunctions, 1069–1073 symmetrization of kernels, 1069 integral-transforms Fourier transform solution, 1054–1055 generalized Abel equation, 1055–1056 introduction, 1047–1048 definition, 1047 exercises, 1053 linear oscillator equation, 1050–1052 ArfKen_Index.tex Index 1193 momentum representation in quantum mechanics, 1048–1049 transformation of differential equation into integral equation, 1049–1050 linear, 1047 Neumann series exercises, 1068 overview, 1064–1066 solution, 1066–1067 separable kernel, 1057–1059 Volterra equation, 1047, 1048, 1050, 1055 integral form, Neumann functions, 671 integral operator, 275, 1066 linear, 1070 integral representations, 647–648, 964 of dilogarithm, 924 f expansion of, 690–692 of Hankel functions, 698 f modified Bessel functions, 684 integral test, Cauchy, seeCauchy (Maclaurin) integral test integral theorems exercises, 169–170 Gauss’ theorem, 164–165, 165 f Green’s theorem, 165–166 Stokes’ theorem, 167–168, 167 f,168 f integral transforms, 1054 convolution (Faltungs) theorem, driven oscillator with damping, 1035–1037 convolution theorem, 985–987 Parseval relation, 987–990 Fourier transform of derivatives heat flow PDE, 983 wave equation, 981–982 Fourier transform of Gaussian, 968–969, 968 f Fourier transform solution, 1054–1055 generalized Abel equation, 1055–1056 inverse Laplace transform Bromwich integral, 1038–1040, 1040 f exercises, 1042–1045 inversion via calculus of residues, 1040 multiregion inversion, 1041–1042, 1041 f, 1042 f Laplace transform of derivatives, 1016–1020 Earth’s nutation, 1018–1019, 1018 f impulsive force, 1019–1020 simple harmonic oscillator, 1017 use of derivative formula, 1017 Laplace transforms, 1008–1034 definition, 1008 Dirac delta function, 1010–1011 elementary functions, 1008–1010 exercises, 1014–1015 Heaviside step function, 1010 inverse transform, 1012 t, 1011–1014, 1014 fpartial fraction expansion, 1013 properties of, 1016–1034 step function, 1013–1014, 1014 f Laplace, Mellin, and Hankel transforms, 965–966 use of, 964 f integrals, 67, 764, 927, see also Cauchy (Maclaurin) integral test; definite integrals; elliptic integrals containing logarithm, 532–534, 533 f contour, 581, 592, 581 f cosine, 580–582 definite, 580 evaluation of, 65 1-D integral, 66 differentiate parameter, 67 exercises, 74–75 integration by parts, 65 integration variables, 72–74 multiple integrals, 70–72 recursion, 69 trigonometric integral, 69 exponential, 578–580 of meromorphic function, 526–527, 527 f of three spherical harmonics, 803–805 oscillatory, 529–530 range, 525–527, 525 f sine, 580–582 trigonometric, 522–524 with complex exponentials, 527–531, 529 f integrating factors, 334 integration by parts, 65, 568, 578 by parts of volume integrals, 163 contour of, 530–531 of power series, 30, 583 order, reversing, 70–71 technique, 531–532 variables, 72–74 intersections, 1127–1130, 1128 f invariants example, 295 exercises, 296 overview, 294–295 inverse Fourier transform, 1055 inverse Laplace transform Bromwich integral, 1038–1040, 1040 f exercises, 1042–1045 inversion via calculus of residues, 1040 multiregion inversion, 1041–1042, 1041 f, 1042 f inverse matrix, 99–102 inverse transform, 211, 1011–1014, 1012 t inversion multiregion, 1041–1042, 1041 f,1042 f ArfKen_Index.tex 1194 Index inversion ( continued ) of power series, 32 via calculus of residues, 1040 inversion operation, 136 irreducible representations, 822 irreducible spherical tensors, 796 irregular (essential) singular point, 344 irregular sign changes, series with, 12–13 irregular singularities, 353–354 irregular solution, 369 irrotational, 152, 154–155 isobaric ODEs, 335 isomorphic group, 817 isomorphism, 819 isospin, SU(2), 852–861 isotropic tensors, 209 J Jacobi method, 314 Jacobi-Anger expansion, 655 Jacobian, 73 2-D and 3-D, 229–230 definiton, 227 direct approaches to, 231 exercises, 231–232 inverse of, 230–231 Jacobian determinant, 229 Jacobian matrix, 229 Jensen’s theorem, 585 Jordan’s lemma, 528 K Kepler’s laws of planetary motion, 189–190 kernel, 963 of integral equation, 455 kernel equation, 1047, 1052 f separable, 1057–1059 kernel function, 447 Kirchoff diffraction theory, 166 Kirchoff’s law, 338 Korteweg-deVries equation, 413 Kronecker delta, 79–80, 209, 258, 805 Kronig-Kramers optical dispersion relations, 591, 594 Kummer’s test, 8–10 L L’Hôpital’s rule, 31, 517, 516, 576, 662 ladder operators, 776–779 construction, 788–795 Lagrangian equations, 1112–1113 of motion, 1098 Lagrangian mechanics, 63 Lagrangian multipliers, 1107–1109 Laguerre functionsassociated Laguerre polynomials, 892–895 differential equation–Laguerre polynomials, 890–892 exercises, 897–899 hydrogen atom, 896–897 Rodrigues formula and generating function, 889–890 Laguerre polynomials associated confluent hypergeometric representation, 919 generating function, 892–895 integral representation, 895 orthogonality, 895 recurrence relations, 893 Rodrigues’ representation, 895 Schrödinger’s wave equation, 896 confluent hypergeometric representation, 919 differential equation, 890–892 generating function, 889–890 recurrence relations, 890, 893 Rodrigues’ formula, 889–890 self-adjoint form, 894 Laplace convolution theorem, 1034–1038 exercises, 1037–1038 Laplace equation, 154, 433–434, 726, 1101–1102 Bessel functions, 651 for parallelepiped, 417–419 Green’s function, 462, 464 Laplace expansion, 760–761 Laplace series expansion theorem, 762 gravity fields, 762 Laplace spherical harmonic expansion, 799–801 Laplace transforms, 965, 1008–1034, 1054 convolution theorem, 1056 of derivatives, 1016–1020 Earth’s nutation, 1018–1019, 1018 f impulsive force, 1019–1020 simple harmonic oscillator, 1017 use of derivative formula, 1017 definition, 1008 Dirac delta function, 1010–1011 elementary functions, 1008–1010 exercises, 1014–1015 Heaviside step function, 1010 inverse transform, 1011–1014, 1012 t,1014 f one-sided, 1008 operations, 1028 t other properties Bessel’s equation, 1025–1027 change of scale, 1020 damped oscillator, 1021 derivative of a transform, 1024–1025 electromagnetic waves, 1023–1024 exercises, 1028–1034 ArfKen_Index.tex Index 1195 integration of transforms, 1027 RLC analog, 1022, 1022 f substitution, 1020 translation, 1022–1023 partial fraction expansion, 1013 properties of, 1016–1034 step function, 1013–1014, 1014 f two-sided, 1008 Laplacian, 154 development by minors, 88 in circular cylindrical coordinates, 192 of vector, 155–156 Laurent expansion exercises, 496–497 Laurent series, 494–496 Taylor expansion, 492–494, 493 f Laurent series, 33, 494–496, 644 least squares, method of, 1138 Legendre duplication formula, derivation of, 618–619 Legendre equation, 425, 716 Legendre functions, 425, 715, 768 f associated, 744–745, 745 t hypergeometric representation, 914 recurrence formulas for, 745–746, 764 associated Legendre equation, 741–743 exercises, 753–756 magnetic field of current loop, 748–753 orthogonality, 748 parity and special values, 746 generating function electric multipoles, 737–739 exercises, 740–741 expansion, 736–737 physical interpretation of, 735 Legendre polynomials, 716 associated, 743–744 exercises, 722–724 recurrence formulas, 718–720 Rodrigues formulas, 720–721 upper and lower bounds for Pn.cos/, 720 of second kind, 766 alternate formulations, 769–770 exercises, 770–771 properties, 769 orthogonality, 724 Earth’s gravitational field, 727 electrostatic potential for ring of charge, 729–730 exercises, 730–735 Legendre series, 726–730 sphere in uniform field, 727–729, 728 f spherical harmonics, 756 Cartesian representations, 758 exercises, 765–766Laplace expansion, 760–762 properties of, 764–765 solutions, 758–760 symmetry of solutions, 762–763 Legendre ordinary differential equations (ODEs), 716 Legendre polynomials, 270–271, 425, 557–558 719t associated, 743–744 exercises, 722–724 generating function, 557–558, 716, 735 orthogonality of, 726 recurrence formulas, 718–720 Rodrigues formulas, 720–721 Schlaefli integral, 557 upper and lower bounds for Pn.cos/, 720 Legendre series, 9–10, 726–730 Legendre’s differential equation, 276, 388 Legendre’s duplication formula, 604 Legendre’s equation, 389–390 Leibniz criterion, 11–12 Leibniz’s formula, 553, 742 Lerch’s theorem, 1011 level lines, 586 Levi-Civita symbol, 85, 87, 216, 841 , 850 line integrals, 159–160, 160 f linear electric quadrupole, 738 f linear equation, 88–89 linear equation system, 102–103 linear first-order ODEs, 336–339 linear Hermitian operator, 311 linear independence of solutions, 358–360 linear integral equations, 1047 linear integral operator, 1070 linear operation, 329 linear operators, 275, 329, 401 linear oscillator, 347–350 linear oscillator equation, 347, 363, 1050–1052 linear parameters, variation of, 1121 linear vector space, 252 linearly dependent equations, 90–91 Liouville’s theorem, 490 Lippmann-Schwinger equation, 466 logarithm, 60 Lommel integrals, 665 Lorentz covariance of Maxwell’s equations, 866–868 exercises, 868–869 Lorentz gauge, 174 Lorentz group, seehomogeneous Lorentz group exercises, 865–866 Lorentz transformation, 862 of E and B, 867–868 lowering operator, 777 ArfKen_Index.tex 1196 Index M Maclaurin expansion, 64 computation, 613 Maclaurin integral test, 5–8 Riemann Zeta function, 7 Maclaurin series, 27, 44, 253 Maclaurin theorem, 27 exponential function, 27–28 logarithm, 28–29 magnetic dipole, 748–753 magnetic field of current loop, 748–753 magnetic moment, 753 magnetic vector potential, 173–174, 193 manifestly covariant form, 868 mapping, 57 complex variables, 547–549, 548 f exercises, 549–550 conformal, 549 matching conditions, 391 mathematical induction, 40–41 excercise, 41 matrices, 95 addition and subtraction, 96 adjoint matrix, 105 defective, 324 definitions, 95–96 diagonalization, 311–314 Dirac notation in, 266 direct product, 108–112 equality, 96 functions of, 113–114 Hermitian matrices, 108, 315 multiplication, 97, 279 inner product, 97–98 by scalar, 96 normal exercises, 324–326 normal modes of vibration, 322–324 overview, 319–320 null matrix, 96 numerical inversion of, 100 orthogonal matrices, 107 product theorem, 103–104 rank of, 104 symmetric, 105 trace matrix, 105 transpose matrix, 104 unitary matrices, 107, 314 matrix algebra, 95 matrix eigenvalue equation, 300 matrix eigenvalue problems, 301 example, 301–303 2-D ellipsoidal basin, 303–305 block-diagonal matrix, 305–307 exercises, 308–310matrix elements, 279 of operator, 280–281 matrix invariant, 295 matrix products, operations on, 106 Maxwell’s equations, 155, 241–243, 594 Gauss’ law, 176 Lorentz covariance of, 866–868 Maxwell-Boltzmann distribution, 606–607 Maxwell-Boltzmann statistics, 1132, 1133 mean value, 1136–1140 mean value theorem, 26, 62 measurement errors, 1125, 1170 Mellin transforms, 966, 1054 meromorphic, 498, 515 meromorphic functions integral of, 526–527, 527 f pole expansion of, 515–518 metric spaces, 218 metric tensor, 218–219 Christoffel symbols as derivatives of, 223 metric, curvilinear coordinates, 184 Milne’s model, 37 Minkowski space, 236–237, 864 miscellaneous vector identities, 156–157 Mittag-Leffler theorem, 515–516 mixed derivatives, 401 mixed tensor, 209, 210 modified Bessel functions, 678, 680, 682 f asymptotic expansion, 688 contours, 696 f exercises, 688 Fourier transforms, 684 Green’s function, 684–685 Hankel function, 682 hyperbolic Bessel functions, 683 integral representation, 684, 690–692 Laplace equations, 680 recurrence relations, 681–682 series expansion, 681 Whittaker functions, 682 modified spherical Bessel functions, 428 modulus, 56, 470 molecular vibration, 882–883 moment-generating function, 1141–1142, 1149–1150 momentum, seeangular momentum momentum representation in quantum mechanics, 1048–1049 Schrödinger wave equation, 994 monopole moment of charge distribution, 739 monotonic decreasing function, 5 movable singularities, 378–379 moving particle, Cartesian coordinates, 1098–1099 multinomial coefficient, 1132 ArfKen_Index.tex Index 1197 multiple integrals, 70–72 multiplet, 827, 851 multiplication by scalar, 255 of matrices, inner product, 97–98 operator, 275 of random variables, 1162 multiply connected regions, 483–484, 483 f multipole expansion, 738, 739, 801–803 multipole moments of charge distribution, 739, 801 multivalued function, 500 mutually commuting operator, 785 mutually exclusive events, 1126 N Navier-Stokes equations, 190, 377 NDEs, seenonlinear differential equations negative definite operators, 317 neighboring paths, 1082, 1083 f Neumann boundary conditions, 385, 412, 428, 662 Neumann functions, 367–369, 693, 917 Bessel functions of second kind coaxial wave guides, 672 definition and series form, 667–669 exercises, 674 integral representations, 669 recurrence relations, 669–670 uses of, 671 Wronskian formulas, 670–671 integral form, 671 recurrence relations, 669–670 spherical, 700 f Wronskian formulas, 670–671 Neumann series, 1064–1066 exercises, 1068 Newton’s equations of motion, 213 Newton’s law, 331, 342 Newton’s second law of motion, 322, 1106 nodes of standing wave, 435 nonlinear differential equations (NDEs), 377–380 Bernoulli and Riccati equations, 378 exercises, 379–380 fixed and movable singularities, special solutions, 378–379 nonlinear dispersive equation, 413 nonlinear methods and chaos nonlinear differential equations (NDEs) Bernoulli and Riccati equations, 378 exercises, 379–380 fixed and movable singularities, special solutions, 378–379 nonlinear ODEs, 377–380 nonnormal matrices, 322–324 nonuniform convergence, 21–22nonunitary transformations, 293 normal distributions, addition theorem for, 1161 normal eigensystem, 320–321 normal matrices defective, 324 example, 320–321 exercises, 324–328 normal modes of vibration, 322–324 overview, 319–320 normalization, 662 normalization constant, 606 nucleon, 853 null matrix, 96 numerical evaluation, 91–93 O ODEs, seeordinary differential equations Oersted’s law, 168 Olbers’ paradox, 11 one-dimensional problems, Green’s function, 448–459 one-sided Laplace transform, 1008 operators adjoint, 277 basis expansions of, 279–280 commutation of, 276–277 example, 277, 278, 280–282 exercises, 282–283 expression, 285–286 functions of, 282 identity, inverse, adjoint, 277–278 matrix elements, 280–281 overview, 275–276 self-adjoint, 277, 284–285 example, 284–286 overview, 283–284 transformations of, 291 exercises, 294 nonunitary transformations, 293 unitary successive transformations, 290 unitary transformations, 287–288 operators, differential vector, seedifferential vector operators optical dispersion, 594–595 optical path near event horizon of black hole, 1087–1088, 1088 f orbital angular momentum, 782 order 2 branch points, 500 ordinary differential equations (ODEs), 329, 330, 381, 644, 715, 982, 1084 exact, 333–334 Hermite, 554 homogeneous linear, 330 homogenous, 335 ArfKen_Index.tex 1198 Index ordinary differential equations ( continued ) inhomogeneous linear, 375–377 exercises, 377 isobaric, 335 Legendre, 557, 716 linear first-order, 336–339 linear second-order, 1049 initial/boundary conditions in, 1052 nonlinear, 413 nonseparable exact, 333–334 Rodrigues formulas, 551, 552 second order, 452 second-order linear, 343–346 second-order Sturm-Liouville, 551 separable, 331–332 singularities of, 345 t with constant coefficients, 342–343 ordinary points of the ODE, 344 orthogonal, 124 transformations, 135 orthogonal coordinates, R3, 182–184, 182 f,183 f orthogonal eigenfunctions, 1069–1073 orthogonal functions, expansions in, 258–259 orthogonal matrices, 107, 135 orthogonal polynomials, 272 f exercises, 558–560 generating function, 555, 556 Hermite, 555–556 Rodrigues formula, 551–554 Schlaefli integral, 554 orthogonal unitary, 277 orthogonality, 51, 703, 724, 906–907 associated Legendre equation, 746–748 Bessel series, 663 Earth’s gravitational field, 727 electrostatic potential for ring of charge, 729–730 electrostatic potential in a hollow cylinder, 663–664 exercises, 665–667, 730–735 integral, Hermite polynomials, 875 Legendre series, 726–730 Neumann boundary condition, 662 normalization, 662 over discrete points, 1002–1004 sphere in uniform field, 727–729, 728 f Sturm-Liouville differential equations, 1073 Sturm-Liouville theory, 661 orthogonalization Gram-Schmidt overview, 269–270 example, 270–272 exercises, 273–275 orthonormalizing physical vectors, 272–273 orthogonalized Laguerre functions, 892orthonormal set, 258 orthonormalization, physical vectors, 272–273 oscillator damping, 1035–1037 driven, 1035–1037 harmonic, 878–879 oscillatory integral, 529–530 oscillatory series, 2 outward flow, 148 overlap integral, 989–990 overlap matrix, 317 overshoot, calculation of, 959–961 P parabolic partial differential equations (PDEs), 410 parallelepiped, Laplace equation for, 417–419 parity and special values, 746 Bessel functions, 655 Parseval relation, 595–596, 987–990 partial derivatives, 62, 401 partial differential equations (PDEs), 329, 643, 981–982 boundary conditions, 405, 411–413 characteristics of, 404–406 classes of, 409–411 elliptic, 410 examples of, 402–403 exercises, 408–409 first-order, 403–408 heat flow, or diffusion, 983 alternate solutions, 439–441 exercises, 444 special boundary condition again, 441–442 specific boundary condition, 437 spherically symmetric heat flow, 442–444 homogeneous, 402 hyperbolic, 410 nonlinear, 413–414 parabolic, 410 second-order, 409–411 separation of variables, 414, 430–432 Cartesian coordinates, 415–420 circular cylindrical coordinates, 421–424, 431 exercises, 432–433 spherical polar coordinates, 424–430 types of, 402 partial fraction expansion, 42, 43, 767, 1013 partial sum approximation, 579 partial-wave components, 799 particle, in a sphere, 704–706 passive rotations, of coordinate system, 215 Pauli matrices, 112 ArfKen_Index.tex Index 1199 PDEs, seepartial differential equations periodic functions, 939–940 permutation group, 845 permutations, counting of, 1130–1133 physical space, 964 piecewise regular, 936 pivotal method, 1176 plane triangle, 131, 131 f Pochhammer symbol, 35, 699, 912, 917 Poincaré group, 862 Poincaré’s lemma, 239, 240 point groups, 820, 835 point quadrupole, 738 Poisson distribution, 1151, 1153 f exercises, 1154–1155 limits of, 1157–1158 relation to binomial distribution, 1153–1154, 1154 f Poisson noise, 1151 Poisson’s equation, 176–177, 433–434 polar coordinates, 442, see also spherical polar coordinates evaluation, 72 polar vectors, 136 polarization matrix, 212 pole expansion of meromorphic functions, 498, 515–518 poles, 497–498 polygamma function, 612 polylogarithms, 923 polynomials, 701, 748 Bernoulli, 565–566 Hermite, 280, 554 example, 554–555, 556 Legendre, 270–272, 557–558 orthogonal, 272 f exercises, 558–560 generating function, 555 Hermite, 555–556 Rodrigues formulas, 551–554 Schlaefli integral, 554 positive definite operators, 317 potential theory exercises, 180–182 Gauss’ law, 175–176, 175 f Helmholtz’s theorem, 177–180 Poisson’s equation, 176–177 scalar potential, 171–172 vector potential, 172–175 potential, of charge distribution, 988–989 power series, convergence, uniform and absolute, 29 differentiation and integration, 30 inversion of, 32 uniqueness theorem, L’Hôpital’s rule, 31power spectrum, 945 power-series expansion, 670 principal axes, 300, 305 principal quantum number, 897 principal value, 501 probability binomial distribution exercises, 1151 limits of, 1157–1158 moment-generating function, 1149–1150 repeated tosses of dice, 1148–1149 definitions, simple properties, 1126 conditional probability, 1128–1129 counting permutations and combinations, 1130–1133 exercises, 1133 probability for AorB, 1127 scholastic aptitude tests, 1129–1130 Gauss’ normal distribution, 1155–1159 Poisson distribution, 1151, 1153 f exercises, 1154–1155 limits of, 1157–1158 relation to binomial distribution, 1153–1154, 1154 f random variables addition of, 1160–1161 computing discrete probability distributions, 1136 continuous random variable: hydrogen atom, 1135–1136 discrete, 1134–1135, 1137 exercises, 1147–1148 mean and variance, 1136–1140 multiplication or division of, 1162 repeated draws of cards, 1145–1147 standard deviation of measurements, 1138–1140 transformations of, 1159–1160 statistics chi-square (2) distribution, 1170–1174 confidence interval, 1176–1178 error propagation, 1165–1168 exercises, 1178–1179 fitting curves to data, 1168–1170 student tdistribution, 1174–1176 theory of, 1125 probability density, student t,1176 f probability distributions arbitrary, 1157 conditional, 1147 marginal, 1144–1147 moments of, 1141–1142 product expansion of entire functions, 519–520 products, seecross product; direct product expansion of entire functions, 519–520 ArfKen_Index.tex 1200 Index products ( continued ) infinite convergence of, 575 exercises, 576–577 overview, 574–575 sine, cosine functions, 575 pseudoscalar, 137 pseudotensors, 215–216 and dual tensors, 216–217 exercises, 217–218 Levi-Civita symbol, 216 pseudovectors, 136, 137 f, 215 Q QCD, seequantum chromodynamics quantum chromodynamics (QCD), 861 quantum mechanical oscillator wave functions, 880f quantum mechanical scattering, Born approximation, 465–466 quantum mechanical simple harmonic oscillator, 878–879 quantum mechanics momentum representation in, 1048–1049 of triangular symmetry, 829–830 Schrödinger equation of, 330, 1048 sum rules, 594, 596 time-dependent perturbations, 1067 quantum number, 776 quantum oscillator, 1120 f, 1119–1121 quantum particle, 419–420 quantum theory, 704 quarks, 853 ladders, 856–857, 857 f quantum numbers of, 855–856 quotient rule, 211–213 R radius of convergence, 494 radius vector, 48 raising operator, 777 random variables, 1125 addition of, 1160–1161 gamma-distribution, 1164 computing discrete probability distributions, 1136 continuous random variable: hydrogen atom, 1135–1136 discrete, 1134–1135, 1137 exercises, 1147–1148 mean and variance, 1136–1140 multiplication or division of, 1162 repeated draws of cards, 1145–1147 standard deviation of measurements, 1138–1140transformations of, 1159–1160 exercises, 1164–1165 rank, 848 tensor of, 205, 207–208 rapidity, 863 ratio test, Cauchy, d’Alembert, 4–5 Rayleigh formulas, 702 Rayleigh’s theorem, seeParseval relation Rayleigh-Ritz variational technique, 1117–1118, 1121 ground state eigenfunction, 1118–1119 real axis, 56 real part, 56 rearrangement of double series, 18–19 rearrangement theorem, 820 reciprocal lattice, 129 recurrence formulas, 556, 718–720, 764 recurrence relations, 348 Bessel functions, 645–646 spherical, 702 Chebyshev polynomials, 901–903 confluent hypergeometric functions, 918 Hankel functions, 675 Hermite polynomials, 872–873 Laguerre polynomials, associated, 893 modified Bessel functions, 681–682 Neumann functions, 669–670 spherical Bessel functions, 702 recursion, 69 reducible representation, 823–824 reference frame, 868 reflection formula, 603 reflections in spherical coordinates, 196 of coordinate transformations, 136–137, 137 f regression coefficient, 1168 regular singularities, 353–354 regular solution, 369 relativistic energy, 35–36 representation counting irreducible, 832–833, 833 t decomposing a reducible, 834–835 exercises, 825, 826 f of continuous groups, 824–825 of group, 821 reducible, 823–824 unitary, 821–823, 823 f residue theorem, 509–510, 509 f resolution of identity, 266, 297 resonant cavity, 650–653 Riccati equations, 378 Riemann Zeta function, 7, 16–17, 571, 626–631 exercises, 631–633 Riemann’s theorem, 15 Riemannian spaces, seemetric spaces ArfKen_Index.tex Index 1201 RL circuit, 338–339 RLC analog, 1022, 1022 f Rodrigues formula, 551–554, 720–721 for Hermite ODE, 554 Laguerre polynomials, 889–890 associated, 895 Rodrigues representation, Hermite polynomials, 874 root diagram, 857 root test, Cauchy, 4 rotations groups SO(2) and SO(3), 849–851 inR3, 139–142, 140 f exercises, 142–143 in spherical coordinates, 194–196, 195 f of circular disk, 818 of coordinate system, 215 of coordinate transformations, 133–135 Rouché’s theorem, 518–519 row vectors, 95, 123, 125 extraction of, 108 S saddle points, 63, 433 argument, 586 asymptotic forms factorial function, 588 of gamma function, 588–589 for avoiding oscillations, 589 method, 587–588 overview, 585–587 sample space, 1126 sample standard deviation, 1168 sawtooth wave, 937–939, 938 f, scalar, 205 scalar field, 143 scalar potential, 171–172 scalar product, 51, 254–255, 271, 285, 295, 297 and adjoint operator, 278 in spin space, 259 triple, 128–130, 129 f, scalar quantities, 46 scattering cross section, 465 Schlaefli integral, 554, 604, 653 f, 676 Legendre polynomials, 557 scholastic aptitude tests, 1129–1130 Schrödinger equation, 708, 1048, 1116–1117 hydrogen atom, 896 momentum space representation, 994 of quantum mechanics, 330 Schwarz inequality, 51, 257 Schwarz reflection principle, 547 exercises, 549–550 second-order linear ODEs, 343–346 second-order partial differential equations (PDEs)boundary conditions, 411–413 classes of, 409–411 exercises, 414 nonlinear, 413–414 second-order Sturm-Liouville ordinary differential equations (ODEs), 551 second-rank tensor, 207–208 secular determinant, 302 secular equation, 302, 306 self-adjoint matrices, 108 self-adjoint ODEs boundary conditions, 381 deuteron, 391–393 eigenvalues, 389 exercises, 393–395 Hermitian operators, 384 Legendre’s equation, 389–390 self-adjoint operators, 277, 284–286, 1070 example, 284–286 exercises, 286–287 overview, 283–284 self-adjoint poblems, Green’s function, 460 self-adjoint theory, 384 semi-convergent series, 579 separable kernel, 1057–1059 homogeneous Fredholm equation, 1059–1060 separable ODEs, 331–332 separation of variables, 403, 414, 430–432 Cartesian coordinates, 415–420 circular cylindrical coordinates, 421–424, 431 exercises, 432–433 spherical polar coordinates, 424–430 series approach Bessel’s equation, limitations of, 351–353 Chebyshev, 25 hypergeometric, 912 Legendre, 9–10 shifted polynomials, Chebyshev, 908 ultraspherical, 25 series expansion, 681 series solutions, Frobenius method, 346–350, 350 f exercises, 355–358 expansion about, 350 Fuchs’ theorem, 355 limitations of series approach, Bessel’s equation, 351–353 regular and irregular singularities, 353 –354 symmetry of solutions, 350–351 sets, 1127–1130 several dependent and independent variables, relation to physics, 1105 sign changes, series with alternating, 12–13 signal-processing applications, 997–1001 exercises, 1001–1002 similarity transformations, 208, 293 ArfKen_Index.tex 1202 Index simple pendulum, 927–928, 1113–1114, 1113 f simple pole, 498 simultaneous diagonalization, 314–315 sine infinite products, 575 integrals in asymptotic series, 580–582 single-electron wave function, 396 single-slit diffraction pattern, 972 singular points, 343–345, 345 t essential (irregular), 344 irregular (essential), 344 isolated, 497 singularities analytic continuation, 503–507, 504 f,505 f exercises, 507–508 fixed, 378–379 movable, 378–379 on contour of integration, 530–531 poles, 497–498 Slater-type orbitals (STOs), 990 Snell’s law, 1095, 1095 f SO(2) rotation groups, 849–851 SO(3) rotation groups, 849–851 soap film, 1088–1090, 1089 f soap film–minimum area, 1090–1093, 1092 f solar products, 256–257 solenoidal, 149, 154–155 soliton, 413 source term, 447 space groups, 869 special relativity, 862 special unitary groups, SU(3), Gell-Mann matrices, 852–861 special values, 764 parity and, 746 spectral decomposition, 315–316 sphere in uniform field, 727–729, 728 f sphere with boundary conditions, 428–430 spherical Bessel functions, 427, 428 asymptotic values, 703 definitions, 702 exercises, 709–712 Helmholtz equation, 698 limiting values, 703 modified, 709 orthogonality and zeros, 703 particle in a sphere, 704–706 recurrence relations, 702 spherical coordinates, Helmholtz equation, 698 spherical Green’s functions, 463, 800–801 spherical harmonics, 445, 473, 756 addition theorem for, 797–798 Cartesian representations, 758 Condon-Shortley phase, 758, 760 t exercises, 765–766integrals of three, 803–805 ladder, 779 Laplace expansion, 760–761, 799–801 Laplace series–gravity fields, 762 properties of, 764–765 symmetry of solutions, 762–763 vector, 809–813 spherical pendulum, 1105, 1105 f spherical polar coordinates, 79, 183, 190–194, 194f, 424–430 spherical tensors, 796 addition theorem, 797–798 Laplace expansion, 799–801 spherical wave expansion, 798–799 exercises, 806–809 integrals of three spherical harmonics, 803–805 spherical volume, 704 spherical waves Bessel functions, 703 expansion, 798–799 spin operator, adjoint of, 282 spin space, 259–260 of electron, 253 spinor ladder, 780–781 spinors, 213, 779–780, 852 square integrable, 595 square integration contour, znon, 479–481, 480 f, square pulse, transform of, 1010, 1010 f square wave, 949–950, 949 f, 958–959 expansion of, 264 f squares of random variables, 1164 squares of series, divergent, 15–16 standard deviation, 1138 of measurements, 1138–1140 sample, 1168 standing waves, 382–384, 435 star operator, seeHodge operator stationary, 63 stationary paths, 1085, 1085 f stationary points, 433 statistical hypothesis, 1165 statistics, 1125 chi-square (2) distribution, 1170–1174 confidence interval, 1176–1178 error propagation, 1165–1168 exercises, 1178–1179 fitting curves to data, 1168–1170 student tdistribution, 1174–1176 steepest descent method of, 585 asymptotic form of gamma function, 588–589 exercises, 590–591 factorial function, 588 saddle points, 585–588 ArfKen_Index.tex Index 1203 step function, 1013–1014, 1014 f Stirling’s expansion, 589 Stirling’s series derivation from Euler-Maclaurin integration formula, 623–624 exercises, 625–626 Stirling’s series, 624 Stirlings formula, 567 Stokes’ theorem, 167–168, 167 f,168 f, 193–194 on differential forms, 245–248 STOs, seeSlater-type orbitals stream lines, 149 strong interaction, 852 structure constants, 848 student tdistribution, 1174–1176, 1177 t student tprobability density, 1176 f Sturm-Liouville boundary conditions, 892 Sturm-Liouville equation, 1117 Sturm-Liouville system, 746, 892 Sturm-Liouville theory, 384, 661, 936–937, 1071, 1073 SU(2) and SO(3) homomorphism, 851–852 isospin and SU(3) symmetry, 852 SU(3) symmetry, 852–861 substitution, 1020 subtraction of sets, 1127 of tensors, 208 successive applications of r, 153–154 successive operations, of coordinate transformations, 137–138 successive transfer functions, 1002 f successive unitary transformations, 290 sum evaluation of, 544–546, 546 t exercises, 546–547 sum rules, 596 summation of series, 957–958 superposition principle, 402 for homogenous ODEs, PDEs, 330 surface integrals, 161–162, 161 f,162 f symmetric group, 835, 840–844 exercises, 844–845 symmetric matrix, 105 symmetric stretching mode, 323 symmetric tensor, 208 symmetrization of kernels, 1069 symmetry, 815, 940–941, 942 f and physics, 826–830 exercises, 830 of equilateral triangle, 817 f, 817, 818 f of solutions, 762–763 relations, 593–594T Taylor expansion, 492–494, 493 f Taylor series, 560 Taylor’s expansion, 653 binomial theorem, relativistic energy, 35–36 Maclaurin theorem exponential function, 27–28 logarithm, 28–29 tensor analysis, 205–213 addition and subtraction of, 208 covariant and contravariant, 206–207 exercises, 213–215 isotropic, 209 symmetric and antisymmetric, 208 tensor derivative operators curl, 225 divergence, 224–225 gradient, 224 Laplacian, 225 tensors, see also direct product; quotient rule; spinors; pseudotensors direct product of, 210–211 in general coordinates covariant derivatives, 222–223 exercises, 226 metric tensor, 218–219 second-rank, 207 –208 tensors of rank 0, 205 tensors of rank 1, 205 tensors of rank 2, 207–208 three-dimensional (3-D) differential forms, 407 threefold Hermite formula, 883–884 time-independent Schrödinger equation, 300 TM, seetransverse magnetic trace matrix, 105, 210 transfer function, 998–999, 998 f high-pass filter, 999–1000, 999 f limitations on, 1000–1001 transform, derivative of, 965, 966, see also Hankel; Laplace; Mellin transformations Gram-Schmidt, 293 of differential equation into integral equation, 1049–1050 of operators, 291 nonunitary transformations, 293 of random variables, 1159–1165 unitary, 287–290 translation, 1022–1023 transpose matrix, 104 transverse magnetic (TM), 651 traveling waves, 435 triangle rule, 788 triangular pulse, Fourier transform of, 976, 977 f ArfKen_Index.tex 1204 Index triangular symmetry, quantum mechanics of, 829–830 trigonometric form, 904–905 trigonometric functions exploiting periodicity of, 537–538 trigonometric integrals, 69, 522–524 triple scalar product, 128–130, 129 f triple vector product, 130 triplet state, 259 Two and three dimension problems, Green’s function, 459–467 two-sided Laplace transforms, 1008 U ultraspherical polynomials, 388, 899 equation, 903 self-adjoint form, 906 undetermined multipliers, seeLagrangian multipliers uniform convergence, 21–22, 29, 262 uniformly convergent series, properties of, 24 unions, 1127–1130 unique expansion, 494 uniqueness theorem L’Hôpital’s rule, 31 of power series, 30–31 unit cell, 869 unit matrix, 99 unit vectors, 47 unitary matrices, 107 unitary operators example, 289–290 exercises, 290–291 successive transformations, 290 unitary transformations, 287–288 unitary representation, 821–823, 823 f unitary transformation, 297 V variables dependent, 1096–1097 Hamilton’s Principle, 1097–1098 Laplace’s equation, 1101–1102 moving particle–Cartesian coordinates, 1098–1099 moving particle–circular cylindrical coordinates, 1099 independent, 407 –408, 411 separation of, 403 variance, 1136–1140 variation, 1081 with constraints, 1111–1112 exercises, 1121–1124 Lagrangian equations, 1112–1113 Schrödinger wave equation, 1116–1117simple pendulum, 1113–1114, 1113 f sliding off a log, 1114–1115, 1114 f of linear parameters, 1121 of constant, 338 of parameters, 338, 375–376 variation method, 395–397 exercises, 397 vector analysis reciprocal lattice, 130 rotation of coordinate transformations, 133–135 vector fields, 46, 143 vector integration exercises, 163–164 line integrals, 159–160, 160 f surface integrals, 161–162, 161 f,162 f volume integrals, 162–163 vector Laplacian, 155–156 vector model, 786–788 vector potential, 172–175, 175 vector spaces, 253–254, 295 completeness, 255, 262 linear space, 252 vector spherical harmonics coupling, 810–813 exercises, 813 spherical tensor, 809–810 vector triple product, 130 vectors, 123, 205, see also rotations; gradient, r; tensors; Stokes’ theorem addition of, 47f angle between two, 798 basic properties of, 124–125 by Gram-Schmidt orthogonalization, 269–275 coefficient, 261 contravariant, 206, 219 contravariant basis, 220–221 covariant, 206 covariant basis, 218, 220–221 cross product, 126–128, 126 f,127 f differential vector operators, 143 gradient, 143 direct product of, 210–211 dot products, 49–50 exercises, 52–53, 131–133 fields, 123 Gauss’ theorem, 164–165, 165 f Green’s theorem, 165–166 Helmholtz’s theorem, 177–180 in function spaces Dirac notation, 265–266 example, 253–254, 256–263, 265 exercises, 266–269 expansions, 261 Hilbert space, 255–256 orthogonal expansions, 257–258 ArfKen_Index.tex Index 1205 overview, 251–253 scalar product, 254–255, 260–261 Schwarz inequality, 257 irrotational, 154–155 matrix representation of, 106–107 multiplication of, 252 orthogonality, 51 physical, 272–273 radius vector, 48 Stokes’ theorem, 167–168, 167 f,168 f successive applications of r, 153–154 triple product, 130 triple scalar product, 128–130, 129 f unit vectors, 47 vibrating string, 382–384 vibration, normal modes of, 322–324 vierergruppe, 820 Volterra equation, 1047, 1048, 1050, 1055, 1067 volume integrals, 162–163 vorticity, 151 W wave equation, 435, 981–982 d’Alembert’s solution of, 436 exercises, 437wave guides, coaxial, Bessel functions, 671–672 wedge operator, 233 wedge products, 233 Weierstrass, 504 Weierstrass Mtest, 22–23 Weierstrass infinite-product form of, 602 weight diagram, 859 f, 859, 860 f Weyl representation, 121 Whittaker functions, 682, 919 Wigner matrices, 797 WKB expansion, 577 Wronskian determinant, 359–360 Wronskian formulas Bessel functions, 670–671, 694 confluent hypergeometric functions, 922 linear dependence/independence of functions, 360, 671 solutions of self-adjoint differential equation, 670 Z zero matrix, 96 zero-point energy, 705, 879 zeros, Bessel function, 648–653